IT Brief New Zealand - Technology news for CIOs & IT decision-makers
New Zealand
VAST & AMD expand AI inference infrastructure tie-up

VAST & AMD expand AI inference infrastructure tie-up

Mon, 27th Jul 2026 (Today)
Mark Tarre
MARK TARRE News Chief

VAST Data and AMD have expanded their collaboration on AI infrastructure, combining VAST's AI Operating System with AMD processors and GPUs.

The agreement targets AI cloud providers and enterprises shifting from model training to large-scale inference and agentic AI deployments.

The move reflects a broader change in the AI market as operators try to make better use of expensive graphics processors while managing rising demands on data access, memory and context handling. Systems built mainly for training workloads are under pressure as companies seek to run persistent inference services, retrieval-augmented generation pipelines and multi-turn AI agents in production.

Under the expanded collaboration, VAST has selected AMD's 6th-generation EPYC processors for the latest versions of its CBox and EBox platforms, which underpin the VAST AI Operating System. The companies are also working with DriveNets on a reference architecture that combines AMD's Helios rack-scale AI infrastructure, the VAST software layer and DriveNets networking.

The reference design is intended to guide AI cloud providers and enterprise users building environments for training, inference, reinforcement learning and key-value cache workloads. It also includes sizing and deployment considerations for large AI installations.

VAST and AMD said they have also worked on inference and KV cache optimisations using AMD Instinct GPUs, AMD Infinity Context and AMD ROCm software. In early tests using an AMD Instinct MI355X GPU, VAST said it recorded a ninefold improvement in time-to-first-token and a 9.7-fold increase in token throughput when using VAST for KV cache offloading in high-concurrency agentic AI workloads.

AMD has not independently verified those figures. Performance gains also depend on the underlying hardware and workload conditions.

Another part of the integration involves lifecycle management for KV cache data, which can include sensitive or personal information. VAST said its data lifecycle policies can automatically expire and delete cached data, a feature aimed at enterprises with privacy and compliance requirements.

The infrastructure stack also uses AMD Pensando Pollara 400 AI network interface cards to move data between AMD Instinct GPUs and VAST storage clusters. This connection is designed to support the movement of cache and storage data required by large-scale inference services.

Inference focus

The collaboration comes as hardware suppliers and software vendors reposition AI infrastructure around inference rather than training alone. Training large models remains capital-intensive, but many customers are now more concerned with the cost and operational complexity of serving models continuously in production.

That shift has increased attention on system design beyond raw compute performance. Inference workloads often depend on rapid access to model state, prompt history, context windows and cached tokens, putting pressure on storage, networking and memory architectures.

"AI is entering an operational phase where infrastructure efficiency matters as much as model performance," said John Mao, Vice President, Global Technology Alliances at VAST Data.

"The industry is discovering that inference is fundamentally a data problem. Success depends on how effectively organisations can bring data, compute, memory and intelligence together as a single system. The VAST AI Operating System was built for this transition, giving AI cloud providers and enterprises a more efficient, scalable and open foundation for training, inference and the next generation of agentic AI applications," Mao said.

AMD presented the partnership as part of a broader push for open infrastructure choices in AI deployments.

"The future of AI will be built on an open ecosystem that gives organisations the flexibility to choose the technologies that best meet their performance, operational and business requirements," said Derek Dicker, Corporate Vice President, Enterprise Business Group at AMD.

"Our expanded collaboration with VAST combines AMD EPYC CPUs and Instinct GPUs with the software foundation customers need to accelerate inference, improve infrastructure efficiency and deploy AI at scale. Together, we're enabling AI clouds and enterprises to build high-performance AI factories without compromise," Dicker said.

Cloud uptake

Several cloud and infrastructure providers backed the tie-up, underscoring competition among AI cloud platforms to offer alternatives to more established chip ecosystems. Core42, Crusoe, Vultr, 5C, TensorWave, Embedded LLM, Phanos.AI and DriveNets all pointed to a need for validated designs that reduce integration work and support production AI services.

For cloud providers, the commercial issue is not only access to accelerators but also how effectively those chips are used once deployed. Underused GPUs, slow model loading and bottlenecks around context handling can materially affect the cost of serving AI applications.

VAST said its software acts as a combined data and execution layer across storage, database, streaming and AI services. The company argues that treating model management as a core data service can reduce cold-start delays and improve GPU utilisation across distributed environments.

That proposition is likely to resonate most with operators building large inference platforms, where response times and hardware efficiency directly affect margins. It also reflects a growing effort across the AI supply chain to tie together semiconductors, networking and data infrastructure rather than treat them as separate purchasing decisions.

"Customers increasingly want the flexibility to build AI infrastructure using the technologies that best meet their requirements," said Yossi Kikozashvili, Vice President, Head of Product & G2M, AI Infra at DriveNets. "The collaboration between DriveNets, AMD and VAST demonstrates how an open ecosystem can bring together best-in-class networking, accelerated computing and intelligent data infrastructure to support the next generation of AI factories."