VAST Data’s Dan Chester on why the cost per token now decides the economics of AI

In July, VAST Data expanded its collaboration with AMD to help AI cloud providers and enterprises build and operate AI factories at scale. VAST selected 6th Gen AMD EPYC processors, formerly codenamed “Venice”, to power the sixth generation of its CBox platform and the third generation of its EBox platform, which together run the VAST AI Operating System. The companies also published a reference architecture with DriveNets based on AMD Helios rack-scale infrastructure. It covers model training, inference, reinforcement learning and KV cache workloads, with sizing guidance for AI cloud providers and enterprises.

The headline performance claims concern inference. Early VAST testing on an AMD Instinct MI355X GPU recorded a 9X improvement in time-to-first-token and 9.7X greater token throughput when VAST was used for KV cache offloading in high-concurrency agentic workloads. AMD reviewed those figures but did not verify them independently. VAST notes that relative gains depend on the baseline hardware. The announcement also covers automated lifecycle policies that expire and delete KV cache data containing sensitive or personal information, and the AMD Pensando Pollara 400 AI NIC as the data path between AMD Instinct GPUs and VAST storage.

“AI is entering an operational phase where infrastructure efficiency matters as much as model performance,” said John Mao, Vice President, Global Technology Alliances at VAST Data. “The industry is discovering that inference is fundamentally a data problem. Success depends on how effectively organisations can bring data, compute, memory and intelligence together as a single system.”

AMD emphasised flexibility of choice. “The future of AI will be built on an open ecosystem that gives organisations the flexibility to choose the technologies that best meet their performance, operational and business requirements,” said Derek Dicker, Corporate Vice President, Enterprise Business Group at AMD.

AI cloud providers including Core42, Crusoe, Vultr and TensorWave supported the announcement. “Our infrastructure strategy is built on a silicon agnostic approach to give customers the best optionality and output for their use cases,” said Raghu Chakravarthi, EVP of Engineering and General Manager – Americas at Core42.

The announcement comes at a time when debate over AI infrastructure is divided between warnings of a bubble and unqualified optimism. The Source Code spoke to Dan Chester, EMEA Sales Director, Neoclouds and AI Model Builders at VAST Data, about infrastructure economics, enterprise data and the demands agentic AI places on the systems beneath it. The interview has been edited for length and clarity.

Why AMD, and where does VAST sit within this collaboration?

My perspective is shaped by what is being built in the AI neocloud market and by large model builders, more than by traditional enterprises directly. NVIDIA currently dominates infrastructure build-out, both in hardware supply and in the software stacks and code libraries that developers contribute to. The momentum sits with the NVIDIA ecosystem.

Customers, however, want choice. AMD has historically offered an alternative to Intel in the x86 compute supply chain, and it is also strong in GPU compute. VAST is a customer of AMD and, more accurately, a route to market for it, as AMD technology runs in the hardware platforms our customers buy to deploy VAST software. That involves a degree of technical collaboration. We are also seeing demand for AMD GPUs, particularly the latest generations, as AMD moves to rack-scale solutions.

Neocloud customers place great value on reference designs and turnkey solutions that simplify deployment. The announcement covers two elements. The first is the adoption of AMD CPUs in the next generation of qualified hardware on which many customers will deploy VAST. The second is engineering collaboration to validate VAST systems with the latest AMD GPUs at rack scale and at data centre scale.

Compute remains one of the largest costs on the balance sheet, and return on investment is the leading concern around AI infrastructure. Do partnerships of this kind raise costs, reduce them or help balance the market?

Economics are becoming central to how AI is consumed. Whether consumption is measured in tokens or GPUs, it rests on a cost basis determined by the infrastructure deployed. The underlying infrastructure, through both its capital and operating costs, is by far the largest factor in the cost of delivering AI services.

The market is therefore focused on optimising cost per token, which depends on the cost of infrastructure and on its efficiency. By efficiency, I mean how effectively infrastructure converts kilowatts and capital into tokens in real-world conditions, and how that investment is monetised. This is one reason organisations are evaluating alternative technology platforms to reduce the unit cost of what they sell, whether tokens or GPU capacity.

Collaborations such as this may not reduce capital expenditure significantly. Their value lies in reducing the risk of bringing a system into production. We are validating AMD within the VAST platform, and validating VAST alongside AMD’s latest rack-scale solutions. For an operator, that second element de-risks deployment. It carries a nominal cost, and the return comes from avoiding opportunity cost, which may not appear directly on the bottom line.

Data quality remains a challenge, and arguably should have been the first step of any AI transformation. As infrastructure providers accelerate, is it time to revisit data, or must the two advance together?

The two must advance together, and the enterprise is central to this. Enterprise adoption of AI differs fundamentally from the way most GPUs have been used to date. Roughly 80% of the world’s GPUs are currently used to train models and around 20% to deliver inference. That balance is expected to reverse, and by 2030 I expect around 80% of GPUs to be serving inference on pre-trained models, largely for enterprises.

Model builders work with existing, curated data corpora, much of which is open source. That data does not contain proprietary enterprise records, personally identifiable information or confidential material subject to regulations such as GDPR. Most GPU workloads today therefore involve data that has already been cleaned, screened and anonymised, or collected from public sources. Enterprises, by contrast, seek unique value from AI by applying it to their own data. Much of that data requires guardrails, and identifying which data is sensitive is difficult. An enterprise cannot simply grant an external model access to its entire corporate data set without risking serious breaches of personal and confidential information.

Enterprises must either bring models into their own data centres and apply guardrails there, or carefully cleanse their data and move only a limited subset to where the model runs. Ideally, all enterprise data would be available in real time to whichever models an organisation chooses. That raises questions of data sovereignty, governance and compliance. Confidential computing, which encrypts data so it can leave the corporate firewall, is one approach. Another is to host models within the data centre and enforce guardrails in the middle layer. In both cases, the challenge is one of data more than storage.

Enterprises typically have a head of infrastructure, possibly a head of storage, and a chief data officer. The chief data officer considers data sets, access rights and the demonstration of compliance. The storage layer can either support the required policies natively or remain agnostic, leaving all guardrails to the middle layer. VAST supports both models. Compliance and audit are built into VAST at the storage layer and at higher layers enterprises can use. Our native vector database, for example, can embed data provenance and the applicable access and security policies within the stored embeddings. When an agentic model searches that database, those constraints remain in force throughout the data life cycle, without needing to be maintained separately across environments.

Some AI pilots are failing, whether through absent strategy or missing data, while others are moving rapidly into production. Widely cited failure rates are more nuanced than the headlines suggest. How does a single infrastructure layer serve organisations at such different stages?

The successful pilots we have seen share a clearly defined use case and workflow, along with a plan for deployment after the pilot. A pilot designed to succeed in a controlled, clean-room environment does not necessarily reflect real operating conditions. A pilot should mirror the production environment as closely as possible, instead of operating as an isolated experiment.

One example is a French bank that set out to improve its know-your-customer workflow using AI. The use case was clearly defined, with specific success criteria, and the bank demonstrated the platform successfully using existing workflows. On the storage side, the workflow drew on specific ingested data sets, which remained on the same platform throughout. The bank did not build separate, ring-fenced infrastructure for the pilot. It deployed on the infrastructure intended for production. I cannot comment on whether it is in full production yet, although it was built on the infrastructure the bank expects to use at that stage.

What shifts in capability are you seeing across the market, particularly among neoclouds, and can VAST keep pace with the rate of change in AI?

The neocloud market was largely built to serve a small number of very large customers. It grew rapidly on concentrated demand for tens or hundreds of thousands of GPUs, driven by constraints on power, capital, balance sheets, GPU allocation and capacity. Most GPU clusters are still used in blocks of several thousand GPUs, typically by a single tenant at a time.

That is changing as enterprises enter the market. Few enterprises need 1,000 to 5,000 GPUs on one- to three-year commitments, which is how much neocloud capacity has been contracted. This creates tension, as neocloud investment is often debt-funded, with that debt secured against long-term offtake contracts. A mid-sized enterprise seeking a few hundred GPUs for three months has found limited options. Some providers do offer this, but it has not been the norm.

Enterprises will commit to some baseline capacity, and demand for flexibility will grow considerably. That changes neocloud infrastructure requirements. Clusters have been single tenant or few tenant, and changing tenancy meant offboarding one customer and onboarding the next. A more on-demand model requires operators to reprovision infrastructure dynamically, with appropriate security boundaries, encryption and quality of service.

For a single large customer running training workloads, the primary requirement of storage was availability and fast data movement to and from GPUs. That requirement remains. Operators must now also provide security, encryption, quality of service and service level agreements for multiple customers simultaneously, without compromising performance. The architecture of compute, storage and networking remains similar, while the demands on storage and networking increase significantly under high concurrency and multi-tenancy.

Enterprise teams increasingly combine people and AI agents, and governments including the UAE’s have set timelines for adopting agents. What does this mean for infrastructure, security and scalability, given the pace at which agents are evolving?

Traditional IT infrastructure, particularly data services, was designed to respond at human speed. Agents can scan terabytes of data in seconds, far beyond the rate at which people read. Systems that were adequate for human consumption of data will be severely challenged as agents both consume and generate data.

Even with AI-ready infrastructure, a central challenge is ensuring agents access only the data permitted to the person they are working with. If a chief executive and a call centre operator each have an agent, and the chief executive requests a report on the highest earners in the company, the operator must not be able to access that payroll information. Organisations need to classify and ring-fence data and apply the right guardrails.

The middle layer can be built independently of the infrastructure layer, in which case software alone enforces the rules. VAST can apply many of those middle-layer guardrails at the infrastructure layer itself, where they are difficult to bypass. That prevents files from being inadvertently exposed during agent training, and ensures the correct privileges apply to data based on the permissions of the person working with the agent.

These conditions change quickly. People change roles, agents work with multiple colleagues, and each colleague may hold different access rights. This is a middle-layer challenge, and it is best addressed by embedding access rules at the source of truth, the infrastructure layer, so they apply however that data is represented across the business.

As the year draws to a close, warnings persist that both the AI investment bubble and the infrastructure bubble will burst. My own view is that infrastructure will expand sharply. What does the near future hold for AI infrastructure, and how should capabilities be built?

The volume of gigawatt capacity announced worldwide signals strong expectations of continued growth. Even if only 30% to 40% of that capacity is built within the announced timeframes, it represents substantial follow-on infrastructure investment. That investment will need to track the most aggressive roadmaps of the leading silicon vendors, as economics ultimately determine outcomes. Those economics depend on access to low-cost power and on silicon that delivers the best return per dollar invested.

I expect continued cycles of investment and refresh, with pressure to monetise that investment quickly. IT spending should correlate closely with gigawatt investment and with capacity coming online, and it will follow an aggressive cycle as the industry drives down the unit economics of what will become a high-demand commodity. That commodity is likely to be tokens, though another metric may emerge. Every token ultimately reflects the unit cost of infrastructure and the efficiency with which it produces tokens. Organisations will continue to pursue that efficiency and invest accordingly.

Sindhu V Kashyap

Global Technology Journalist & Multimedia Storyteller | Covering Founders, Investors & Leaders Reshaping Tech | Writer · Interviewer · Moderator · Editor

Next
Next

Who Is Accountable When the Tooling Decides: AIOps and the Shrinking Team