# WEKA® NeuralMesh™

**The Storage and Memory Platform Purpose-Built for AI**

WEKA NeuralMesh extends memory and storage to your AI and HPC workloads. Tokens move at memory speed, GPUs stay fed, and your AI economics finally make sense.

## Engineered to Deliver What Production AI Demands

### Data Reduction

**Reduce Data. Not Performance.**

NeuralMesh Data Reduction delivers up to 6x reduction on AI/ML data with under 5% write overhead, guaranteed. Similarity-based compression captures structural redundancy across blocks, while background-first processing keeps writes at native NVMe speeds and GPU utilization high. Together with thin provisioning, snapshots, single-hop writes, drive sharing, and AlloyFlash hybrid flash, NeuralMesh is the most economically viable all-flash architecture for AI and HPC workloads.

- **Similarity-Based Reduction Built for AI** — Evolving AI datasets produce redundancy byte-exact approaches miss. WEKA finds shared-pattern blocks cluster-wide, capturing savings from checkpoints and container layers.
- **Zero Write-Path Impact** — Writes commit to NVMe at native speed while reduction runs in the background, keeping GPU utilization high regardless of how long checkpoint phases run.
- **Cluster-Wide Deduplication** — Eliminates redundancy across all DR-enabled filesystems: every project, every team, every checkpoint. In shared AI environments, cluster-wide scope is a major multiplier.
- **Performance and Capacity Guaranteed** — Contractual DRR commitment paired with a throughput guarantee on comparable hardware. If realized reduction falls short, WEKA covers the gap with free software licenses.
- **Start With a Scan ** — Scan a sample of your actual data using the same algorithm that runs in production. The projection it generates becomes the contractual number, not a vendor average.
- **Six Efficiency Pillars. One Platform.** — The benefits compound with thin provisioning, snapshots, single-hop writes, drive sharing, and AlloyFlash: Up to 1.9x reads, 13x write throughput, and 3x lower latency.

- [Dive Deeper](/article/data-reduction-capacity-and-performance)
- [Read the Solution Brief](/resources/solution-brief/the-most-economic-all-flash-architecture-for-ai-and-hpc/)
- [Read the Tech Brief](/resources/technical-brief/neuralmesh-data-reduction/)

### Data Replication

**Data That Follows Your GPUs**

Data shouldn't be the bottleneck for distributed AI. NeuralMesh Replication separates where data lives from where AI runs. Pair two WEKA NeuralMesh clusters, and the destination immediately sees the source directories, file names, sizes, and permissions. Data hydrates by policy: on demand, on a schedule, or as a full copy. Direct cluster-to-cluster transfer keeps object storage out of the critical path.

- **Namespace-First Visibility** — After pairing, remote clusters see directory structures, file names, sizes, and permissions instantly — no waiting for file data to move before browsing or starting jobs.
- **Flexible Hydration Policies** — With NeuralMesh Replication there are three hydration models: fetch on first access, pre-stage on a schedule, or run a full copy for recovery and migration.
- **Direct Cluster-to-Cluster Replication** — File data and metadata move directly between WEKA clusters over HTTPS, removing object storage as a transport layer and eliminating the hops and failure modes that come with it.
- **Snapshot-Based Async Updates** — Configurable snapshot cadence bounds Recovery Point Objective (RPO), and if a relationship is disrupted, resync picks up from a common snapshot with no full retransfer.
- **Lower Destination Storage Footprint** — Destination storage scales with what workloads actually access, not total dataset size, bringing material savings when datasets reach petabyte scale.
- **Foundation for Global Data Sharing** — One architecture powers GPU bursting, multi-site research, regional expansion, migration, and off-site resilience: distributed dataset that looks local, moving only when needed.

- [Dive Deeper](article/reimagining-replication-smarter-faster-data-movement-built-for-ai)
- [Read the Solution Brief](/resources/solution-brief/neuralmesh-replication-solution-brief/)
- [Read the Tech Brief](/resources/technical-brief/neuralmesh-replication-technical-brief/)

### S3 Object Store

**S3 That Never Makes GPUs Wait**

NeuralMesh Fast S3 Object Store runs at the speed and scale AI infrastructure demands. S3 and POSIX serve the same data, so training and fine-tuning don't wait on copies between stages, and models ship faster. Inference keeps pace even at thousands of concurrent sessions.

- **GPU-Scale S3 Performance** — The S3 Performance Bucket removes the metadata and connection ceilings that make object storage degrade under large GPU cluster load.
- **Linear Performance at Scale** — Fully distributed metadata with no centralized controllers. Performance scales linearly from six nodes to hundreds with no architectural changes.
- **Unified Object and File Storage** — S3 and POSIX run on the same data, so pipeline stages share one copy instead of five or more. Training, fine-tuning, and inference move forward together.
- **Scale-Compounding Resilience** — Every node you add strengthens the system. NeuralMesh stays up through hardware failures, maintains performance as the cluster grows, and protects data without slowing down.
- **Lower Cost Per Usable Petabyte** — AlloyFlash and always-on data reduction deliver more usable S3 object capacity per rack, per watt, and per dollar, with up to 6x capacity savings.
- **One Operational Model for Any Scale** — NeuralMesh provisions new S3 tenants through standard Kubernetes manifests in minutes. One compliance posture covers every protocol, every workload, and every site.

- [Dive Deeper](/article/neuralmesh-fast-s3-object-store)
- [Read the Solution Brief](/resources/solution-brief/the-s3-object-store-that-never-makes-gpus-wait/)
- [Read the Tech Brief](/resources/technical-brief/neuralmesh-fast-s3-object-storage/)

### Multitenancy

**More Tenants. Better Margins. One Cluster.**

NeuralMesh runs every tenant on one platform, whether it needs dedicated hardware or shared infrastructure. Cost per tenant drops as tenant count grows, and new tenants come online in minutes.

- **Two Isolation Models, One Platform** — Composable clusters give anchor tenants dedicated hardware. Native multitenancy gives smaller tenants isolation on shared infrastructure. Both run from one control plane.
- **Zero Idle Capacity** — Storage capacity and compute are shared dynamically across all tenants. Adding a tenant consumes only what it needs. Removing one returns resources to the pool immediately.
- **Network-Level Tenant Isolation** — Network spaces give each tenant dedicated VLANs and IP ranges, with overlapping ranges supported. Mount enforcement rejects cross-tenant requests before they reach data.
- **Enforced Performance and Security** — QoS ceilings prevent noisy-neighbor impact at the tenant and filesystem level. Independent KMS and LDAP per tenant keep credentials and data invisible to everyone else.

- [Tech Brief](/resources/technical-brief/neuralmesh-multitenancy/)
- [Solution Brief](/resources/solution-brief/multitenancy-built-for-ai-scale-true-isolation-zero-waste/)
- [Read the Blog](/article/built-with-ai-clouds-multitenancy-that-scales-and-economics-that-finally-work)

### Observe

**Visibility Across Every Cluster. One View.**

NeuralMesh Observe gives you real-time visibility across every cluster in your WEKA environment, on-premises, in the cloud, multicloud, or hybrid. Track performance, capacity, and health from a single dashboard. Set threshold alerts that route to Slack, PagerDuty, or email. Diagnose client-level issues without ever opening a support ticket.

- **Multi-Cluster Visibility in One View** — See health, capacity, throughput, and IOPS across every cluster — cloud and on-premises — from a single dashboard, with per-cluster alert counts.
- **Real-Time Performance Metrics** — Track throughput, latency, and IOPS with synchronized charts. Hover over any point and all charts align to that timestamp instantly.
- ** Client-Level Diagnostics** — Compare individual client performance against cluster baselines to identify exactly which client is underperforming and why.
- **Multi-Channel Alert Routing** — Set custom alert rules by cluster, metric, and severity. Notifications route to Slack, PagerDuty, or email with direct links to the relevant dashboard.
- **Workflow-Optimized Dashboards** — Start with pre-built templates for cluster overview, performance, and capacity — then add, resize, and reorder widgets to build the view you actually use.
- **Prometheus and Grafana Integration** — Stream metrics to your existing Grafana or Prometheus environment. SaaS delivery means no local installation and unified SSO access from day one.

- [See Observe in Action](/article/neuralmesh-observe-visibility-and-control-for-your-weka-environment)

### Kubernetes Operator

**Storage That Runs as Kubernetes**

As 80% of enterprises run Kubernetes in production, the storage system that lives outside the scheduler is the operational gap that fails AI workloads at scale. NeuralMesh is the only enterprise storage solution that deploys and manages the full cluster as Kubernetes workloads, governed by the same control plane as compute and networking. Native Custom Resource Definitions (CRDs) bring lifecycle, policy, and driver management under Kubernetes governance.

- **Native K8s Lifecycle Management** — CSI drivers provision volumes. NeuralMesh manages the cluster: Helm-deployed, GitOps-reconciled, failures surface in Kubernetes events. Same model as compute and networking.
- **Composable Multitenant Clusters** — Multiple independent clusters on shared hardware. WekaCluster CRDs and Alloy Flash partitioning give each team isolated storage with its own SLA, lifecycle, and workload policy.
- **Pre-Provisioning Admission Control** — Bad storage configs at runtime mean wasted GPU hours. NeuralMesh admission webhooks catch cluster errors at apply time, before the controller acts or resources are committed.
- **Automatic Slurm CPU Coordination** — For Slurm-on-Kubernetes, NeuralMesh reads kubelet CPU reservations and auto-aligns storage CPU, closing the gap between repeatable deployment and per-node maintenance burden.
- **Manage Kernel Modules Across Clusters** — Managing kernel modules across GPU node fleets used to be manual. NeuralMesh automates it via Kubernetes rolling updates, coordinated by three roles: Dist, Builder, and Loader.
- **No Single Path to Saturate** — Legacy storage funnels I/O through a controller or fixed-IOPS ceiling, capping scale. NeuralMesh spreads data and metadata across every node: no single path to saturate.

- [Dive Deeper](/article/your-kubernetes-workloads-aren-t-cpu-bound-they-re-waiting-on-storage)
- [Read the Solution Brief](/resources/solution-brief/weka-kubernetes-operator/)
- [See the Benchmarks](/article/neuralmesh-postgresql-storage-benchmark)

## Results. Not Promises.

- **6x** *Reduction on AI/ML data with less than 5% write overhead*
- **3x** *Greater object throughput than nearest alternative*
- **90%** *lower costs than traditional storage solutions*

## Customer Story

> WEKA NeuralMesh's intelligent replication makes data mobility real at scale: we can make datasets visible across sites and pull exactly the data each job needs to the next GPU allocation as it becomes available. That shifts replication from a back-end protection function to a core part of how our distributed AI infrastructure needs to operate, improving workload mobility, capacity efficiency, and the resiliency our customers depend on. WEKA's data and memory infrastructure provides the foundation to scale our footprint without compromise.

— Sam Tabar, CEO, WhiteFiber

## Frequently Asked Questions

### What is NeuralMesh?

WEKA NeuralMesh is the storage and memory platform built for agentic AI. It is a software solution that can be deployed in AI or public clouds; on-premises on standard server hardware, WEKApod appliances, or converged on GPU servers as NeuralMesh Axon; and across hybrid cloud environments. Its distributed, containerized services run on every node in a cluster, with no central controller or metadata server that could create bottlenecks or single points of failure. Performance improves as the system scales, delivering microsecond latency at exabyte scale. NeuralMesh is not a parallel file system. It delivers the performance benefits of parallel file systems but through a fundamentally different architecture: distributed microservices built for AI workloads.

### Where does NeuralMesh run?

WEKA NeuralMesh runs anywhere your AI or HPC workloads run: on-premises data centers, hybrid environments, AI clouds, and hyperscale clouds. It runs on standard x86 and ARM infrastructure, with the same architecture and performance model across all environments. Only the underlying infrastructure mapping changes.

### What deployment models does NeuralMesh support?

WEKA NeuralMesh supports five deployment models. On-Premises: NeuralMesh runs on standard servers in your own data center. Cloud: NeuralMesh deploys across all major AI and hyperscale public clouds on standard instances with attached or local NVMe storage. Converged: NeuralMesh Axon runs on the same GPU servers as compute. Appliance: WEKApod delivers NeuralMesh as a purpose-built, pre-configured system. Hybrid: NeuralMesh runs across any combination of these environments. Data placement, protection, and performance behavior stay consistent across all five.

### What protocols does NeuralMesh support?

WEKA NeuralMesh provides native access through POSIX, S3, NFS, SMB, and NVIDIA GPUDirect Storage (GDS). All protocols operate on the same dataset simultaneously: data written through one protocol is immediately accessible through another, with no duplication or format conversion. POSIX delivers the highest IOPS and metadata performance, S3 provides high-performance object access, and GPUDirect Storage moves data directly between storage and GPU memory, bypassing the CPU.

### Can NeuralMesh serve both file and object data?

Yes. WEKA NeuralMesh runs S3 object storage as a Tier 1 protocol alongside file, and both access the same data with no duplication or format conversion. Policy-driven tiering keeps frequently accessed data on flash and moves less active data to lower-cost object storage, with no change in how applications access it.

### How does NeuralMesh handle metadata at scale?

WEKA NeuralMesh has no dedicated metadata servers. Metadata ownership is distributed across virtual metadata servers on every node, and placement is derived algorithmically rather than stored in central lookup tables. Metadata performance scales horizontally as the cluster grows, rather than becoming a bottleneck.

### How resilient is a NeuralMesh cluster?

WEKA NeuralMesh distributes data across independent failure domains, so no two chunks of the same stripe share a domain. When a failure occurs, every node participates in rebuilding only the affected data, in parallel. Rebuild speed increases as the cluster grows rather than slowing down.

### Does NeuralMesh support multi-tenant environments?

Yes. WEKA NeuralMesh supports both physical and virtual multitenancy on one platform. Large tenants get dedicated infrastructure; smaller tenants get isolation on shared infrastructure. Enterprises and AI clouds deploy the isolation strategy that best fits their business model.

## Related Resources

- [Feature Focus: S3 Object Storage That Never Makes GPUs Wait](/lp/get-capacity-now/) (Article)
- [Feature Focus: Data Reduction Without the Performance Tax ](/article/data-reduction-capacity-and-performance) (Article)
- [Feature Focus: Replication Reimagined for AI](/resources/eguide/ai-infrastructure-efficiency-metrics-explained/) (Article)
- [Building AI Factories: Storage Architecture Defines Success](/article/building-ai-factories-storage-architecture-defines-success) (Article)
