Marvell Technology is shipping a three-tier memory-disaggregation portfolio targeting memory bandwidth and capacity — the constraints now limiting AI inference scaling more than raw compute. The rollout spans server-attached SSD, rack-scale CXL pooling, and optical shared-memory across racks, each targeting a different radius from the GPU.
| Tier | Product | Interface | Reach | Primary Use Case |
|---|---|---|---|---|
| Server | Bravera SC6 | PCIe 6.0 | Within server | KV-cache offload to host-managed NAND |
| Rack | Structera X | CXL | Rack-scale | DDR4/DDR5 memory capacity expansion |
| Rack | Structera A | CXL | Rack-scale | Near-memory compute (recommendation, vector search, DB) |
| Rack | Structera S4 | CXL 3.1 | Rack-scale | CXL pooling for clusters without native CXL lanes |
| Cross-rack | Photonic Fabric | Optical | Up to 50 m | 32 TB warm KV-cache offload across racks |
The thesis is straightforward: in large inference clusters running models with multi-hundred-GB memory footprints and long context windows, KV-cache pressure and memory bandwidth saturation arrive before FLOP budgets are exhausted. Marvell disaggregates memory from compute, then reconnects it at whatever distance the workload demands. Khurram Malik, associate VP of product marketing, says the next performance gains come "not only from faster processors, but also from better ways to store, pool, and move memory."
At the server level, the Bravera SC6 is a PCIe 6.0 SSD controller for key-value cache offload. Its differentiating feature is a host-managed flash translation layer — hyperscalers control write amplification, garbage collection, and wear leveling directly rather than delegating to firmware. As Malik puts it: "End customers know their workloads. They write on the NAND based off of their workload and they manage the write amplification, which turns into the endurance of the SSDs." Multi-NAND-vendor support matters when SSD supply chains are unstable.
At rack scale, Structera CXL covers three sub-products. Structera X handles memory expansion using existing DDR4 or DDR5 inventory, with onboard compression yielding 2–2.5× effective capacity versus standard configurations. DDR4 reuse is already in production: Meta runs CXL-based memory expansion across millions of servers, pulling DDR4 modules from decommissioned machines rather than buying new. Malik is direct: "The first and foremost important use case within CXL is the recycling of DDR4." Structera A adds a near-memory compute accelerator beside the memory controller, offloading CPU and GPU cycles for recommendation inference, vector search, and database operations. Structera S4 is a CXL 3.1 switch bridging CPUs and GPUs to CXL pools even when host silicon lacks native CXL lanes — it includes a PCIe-over-CXL protocol conversion layer, critical for clusters built on older GPU generations.
| Sub-product | CXL Version | Key Capability | Target Workloads | Notable Detail |
|---|---|---|---|---|
| Structera X | CXL | DDR4 or DDR5 memory expansion with onboard compression | General inference, memory-capacity scaling | 2–2.5× effective capacity; enables DDR4 recycling from decommissioned servers |
| Structera A | CXL | Near-memory compute accelerator beside memory controller | Recommendation inference, vector search, database operations | Offloads cycles from CPU and GPU |
| Structera S4 | CXL 3.1 | Switch bridging CPUs/GPUs to CXL pools; PCIe-over-CXL conversion | Clusters lacking native CXL host silicon | Enables older GPU generations to access CXL memory pools |
Beyond the rack, Marvell's Photonic Fabric uses optical links to extend a shared-memory tier up to 50 meters — spanning adjacent racks within a data-center pod. The target is long-context inference: models where KV-cache alone overwhelms per-server DRAM. Photonic Fabric offloads up to 32 TB of warm KV-cache, and Marvell claims 2–3× token throughput improvement within the same power envelope and data-center footprint, though the company acknowledges the actual gain is workload-dependent. That caveat matters: the 2–3× figure assumes a memory-bandwidth-bound workload; compute-bound configurations will see less.
Meta's millions-of-servers CXL deployment is the most concrete proof of hyperscaler adoption, but it's memory expansion, not full disaggregation. Photonic Fabric is the bolder architectural bet — optical shared-memory introduces new failure modes, latency jitter across 50-meter links, and operational complexity that rack-local DDR does not. For architects scoping this now, near-term action lands at Structera X and S4, where DDR4 reuse and CXL pooling work without optical infrastructure changes.
If your inference cluster's bottleneck is KV-cache capacity or memory bandwidth — not FLOP throughput — Marvell's three-tier stack gives you a product path at each radius, from server SSD to cross-rack optical. The DDR4 recycling entry point is the lowest-risk first step; save the Photonic Fabric evaluation for clusters where 32 TB warm-cache offload changes the economics.