The Inference Economy Is Here. Your Infrastructure Wasn't Built for It.

Two years ago, the AI infrastructure conversation was all about training. How fast can you get data to GPUs? How efficiently can you checkpoint a model? How long can a cluster run without a failure that costs you a week of progress? Those problems were real, and the industry built infrastructure to solve them; much of it on hardware and software that was excellent for the job.
That era is coming to a close. As the industry matures, the conversation is now about inference and running AI viably in production. Not proofs of concept. Not hypothetical demos. Actual production workloads. Long-context reasoning, agentic systems that maintain state across thousands of interactions, retrieval-augmented generation feeding live applications, models serving millions of users concurrently.
The leading AI labs proved they can build great models. Now, the economic center of gravity is shifting from a focus on training to making production AI, which runs on inference, profitable and efficient.
This is not a small change. It is a different problem with a different stack, and most infrastructure currently being sold as "AI-ready" was built for the wrong era.
The New AI Chokepoint Is the Datacenter Itself
Here is what changed as the industry focused on training: datacenters ran out of room.
US datacenter construction declined in 2025 for the first time since 2020, exactly as AI demand was accelerating. Grid connection queues in the major markets now stretch four to seven years. Morgan Stanley projects a 49-gigawatt power shortfall in the US alone through 2028. High-voltage transformer lead times have stretched to five years. The capital to build is available. The physical infrastructure to support new construction is not on any timeline that matters to a company planning AI deployments this year or next.
Meanwhile, GPU density is doubling. Modern compute density is now scaling faster than the data infrastructure built to support it. The implication is straightforward, and it changes how you have to think about every layer of the AI stack: every rack unit, every watt, and every dollar of infrastructure already deployed has to produce more.
No one can afford to continue expanding their way out of this problem. Ultimately, to increase computational output, you must get more out of what you already have. This is a different optimization function from the one the infrastructure was sold against over the last decade. And it is the one that determines whether your inference deployments add to or destroy your margins.
Storage Is Where This Collision Lands Hardest
When the rack is the binding constraint, every layer of the stack gets evaluated on the same question: how much useful AI work does this contribute per rack unit, per kilowatt, per dollar? Storage is where that question has gone unanswered the longest.
Most of what is sold today as AI storage is general-purpose enterprise storage with new marketing. The hardware underneath the appliance was designed for transactional databases, virtualized workloads, backup use cases, or enterprise file shares. The chassis was someone else's. The density ceiling was inherited from a server platform that was never optimized for the drive counts, thermal profiles, or sustained throughput required by production AI inference. Software vendors took that hardware, installed software on it, and shipped the combination as a turnkey solution. That model worked when storage was a supporting actor. It doesn’t work when storage consumes precious rack space and power that could otherwise be used for GPUs, or in an era when extending GPU memory unlocks more token output by an order of magnitude.
The performance numbers tell part of the story: a general-purpose chassis ships with a fraction of the NVMe drives that a purpose-built system carries. The thermal envelope was specified for enterprise workloads at moderate density, not for the sustained 35-degree-Celsius ambient operation that real AI deployments produce. Components are sourced through OEM channels beyond the storage vendor’s control, which becomes an acute problem when NAND markets move, memory shortages persist, lead times stretch, and procurement teams need pricing they can plan against six to twelve months out.
The deeper problem is structural. If your hardware was not designed for your software, your software is constrained by hardware decisions you did not make. There is a ceiling on what software optimization can recover. We have spent the last several years pushing against that ceiling, and we concluded the only honest path forward was to design the hardware ourselves.
What WEKA Has Built and Why
Today, we are proud to unveil our third-generation WEKApod Nitro, WEKApod Prime, and WEKApod Prime Max appliances, which now run exclusively on WEKA-engineered hardware purpose-built to deliver superior performance and capacity density, and breakthrough inference economics, with multiple patents pending for innovations in chassis architecture, drive interconnect, thermal management, and serviceability.
WEKApod is the first single-rack system in the world to break the exabyte barrier. A single 56-unit rack can hold 1.1 exabytes of effective capacity and deliver 10.2 terabytes per second of throughput and 210 million IOPS. Density is the headline, and we built what is now the world's most performance- and capacity-dense AI storage.
The deeper story, however, is supply chain control. WEKA now directly controls the procurement of all its hardware components, and we manage inventory and pricing rather than inheriting them. Assembly and fulfillment are handled through a carefully selected network of global distribution partners, reducing lead time considerably. For AI cloud providers, frontier model builders, and enterprises planning capacity in a constrained NAND flash market, that predictability is not a feature. It’s a prerequisite for planning at scale.
We also debuted NeuralMesh 6 today, the sixth generation of our core software platform. The significance of this software release cannot be overstated. NeuralMesh 6 delivers an expansive tranche of enterprise-grade features, including:
- Native virtual multi-tenancy that supports more than 1,000 isolated tenants per cluster, with sub-30-minute provisioning, which works in concert with NeuralMesh’s composable cluster functionality to deliver physical multi-tenancy. Together, they scale to tens of thousands of tenants on a single system.
- A unified file and object protocol stack built on NVMe with feature-rich S3 that supports any cloud-native workload and makes the same physical blocks addressable through both protocols, eliminating the data copies that conventional AI pipelines carry between training, fine-tuning, and serving. The combination of fully featured S3 and always-on data reduction in the WEKApod Prime and WEKApod Prime Max configurations delivers groundbreaking flash-based S3 object storage.
- Intelligent metadata-first replication and remote cache for workloads that need to follow GPU availability rather than wait for it. Our first step toward delivering full federation and global namespace functionality.
- Always-on data efficiency with contractual performance guarantees.
- Kubernetes-native operations and unified observability across every deployment, included at no additional cost.
Make no mistake: WEKA is still fundamentally a software company, and NeuralMesh remains our flagship product. Building our own hardware was not the original plan. Our customers' inference economics made it a necessity. We built WEKApod because the alternative was letting someone else's hardware define the limits of what our software could do.
That’s why we announced both our hardware and software innovations today: because of what’s possible when they’re delivered together. A software platform that takes full advantage of the next-generation hardware we deliberately designed for it. Density that compounds across both layers rather than topping out at the chassis. Supply chain predictability that allows AI cloud providers, frontier model builders, and enterprise infrastructure leaders to plan against real lead times. These are not feature checkboxes. They are the conditions that determine whether inference economics work in production, at scale. We have been operating this way at scale for some time. WEKA's Augmented Memory Grid, which extends GPU memory by accelerating the persistent KV cache to NeuralMesh-managed access to NVMe storage, is already running in production at customer and partner sites, including Oracle Cloud Infrastructure. The production results released by our partners at OCI have been substantial: 10x higher token throughput, 10x more concurrent users served, and 7x more tokens from the same GPU footprint. Augmented Memory Grid running on NeuralMesh has proven its operating model at the scale where production inference actually runs.
The Questions Your Infrastructure Vendor Should Be Answering
If you're running production AI at scale today or planning to in the next 12 months, the questions you should be asking your data infrastructure vendor have changed. They apply whether you're an AI cloud provider monetizing GPU capacity, a frontier model builder serving your own inference at unprecedented scale, or an enterprise standing up production AI inside a finite datacenter footprint:
- How much effective capacity can I get per rack unit, and what assumptions does that number depend on? If the answer requires multiple racks where one used to fit, that's rack space and power that could otherwise be used to hold GPUs.
- Who designed the hardware my software runs on? If the answer is another company's general-purpose server, then thermal behavior, density ceiling, and supply chain are not under your vendor's control. They are constraints inherited from a platform that was not optimized for your AI deployment.
- If NAND prices move by 30% over the next two quarters, how will that impact pricing if I need to expand my deployment? If their procurement goes through an OEM channel, your costs track the market. If they control component sourcing directly, your plan will hold.
- How many copies of my training data will your platform require my inference pipeline to carry with it? If the answer is five or more, which is typical for stacks that bolt file and object storage together through gateways, then you’re paying for capacity without output.
- Can your multi-tenancy model isolate workloads at the hardware level when it matters and at the network level when it does not, on the same platform? Or do I have to choose one or the other and live with the consequences? This is not a tradeoff you can afford to make.
These are not academic questions. The answers determine whether the AI investments you've already made can deliver returns, and whether the deployments you are planning over the next two years are economically viable within your current datacenter footprint.
A New Category Is Forming
There is a category-level conversation happening across the industry right now about how AI infrastructure must evolve to support inference workloads at scale. Most vendors are positioning above the substrate, calling their platforms operating systems, rebranding antiquated hardware as an engine, calling their software the answer to whatever question the analyst asked last quarter. There's a reason for this: If your hardware is someone else's, abstraction is the only ground you can fight on.
At WEKA, we’re taking a different position. We believe AI inference requires a new scorecard. A new category of infrastructure built on hardware and software designed to complement each other and run in perfect symmetry. Developed in this decade for modern GPUs and accelerated-compute workloads. Fueled by a reliable, predictable supply chain under the vendor's control. Optimized for the actual operating economics of the inference era. Not legacy enterprise storage with AI marketing. Not general-purpose hardware with new firmware. Purpose-built infrastructure, engineered from the chassis up for inference-era computing. Powerful, turnkey data and memory infrastructure designed to deliver the inference economics that will define AI’s next frontier.
That is what we have been building. That is what we are announcing today, and it is the conversation we’ll be having with our customers, our partners, and the market over the next several years. Infrastructure built for traditional enterprise workloads and machine learning-era training is what got us here. It will not get us where we need to go next.
Learn more about WEKApod 3 and NeuralMesh 6.
What's Next
Scale Production AI Faster with NeuralMesh
Your models aren't slow. Your data is. Fix AI bottlenecks with high-throughput infrastructure.


