Marvell Technology is shipping a three-tier memory-disaggregation portfolio targeting memory bandwidth and capacity — the constraints now limiting AI inference scaling more than raw compute. The rollout spans server-attached SSD, rack-scale CXL pooling, and optical shared-memory across racks, each targeting a different radius from the GPU.

TierProductInterfaceReachPrimary Use Case
ServerBravera SC6PCIe 6.0Within serverKV-cache offload to host-managed NAND
RackStructera XCXLRack-scaleDDR4/DDR5 memory capacity expansion
RackStructera ACXLRack-scaleNear-memory compute (recommendation, vector search, DB)
RackStructera S4CXL 3.1Rack-scaleCXL pooling for clusters without native CXL lanes
Cross-rackPhotonic FabricOpticalUp to 50 m32 TB warm KV-cache offload across racks
FIG. 02 Marvell three-tier memory disaggregation portfolio at a glance — EE Times / Marvell Technology

The thesis is straightforward: in large inference clusters running models with multi-hundred-GB memory footprints and long context windows, KV-cache pressure and memory bandwidth saturation arrive before FLOP budgets are exhausted. Marvell disaggregates memory from compute, then reconnects it at whatever distance the workload demands. Khurram Malik, associate VP of product marketing, says the next performance gains come "not only from faster processors, but also from better ways to store, pool, and move memory."

Marvell's three-tier memory disaggregation architecture — from server SSD to cross-rack optical
FIG. 03 Marvell's three-tier memory disaggregation architecture — from server SSD to cross-rack optical — EE Times / Marvell Technology

At the server level, the Bravera SC6 is a PCIe 6.0 SSD controller for key-value cache offload. Its differentiating feature is a host-managed flash translation layer — hyperscalers control write amplification, garbage collection, and wear leveling directly rather than delegating to firmware. As Malik puts it: "End customers know their workloads. They write on the NAND based off of their workload and they manage the write amplification, which turns into the endurance of the SSDs." Multi-NAND-vendor support matters when SSD supply chains are unstable.

At rack scale, Structera CXL covers three sub-products. Structera X handles memory expansion using existing DDR4 or DDR5 inventory, with onboard compression yielding 2–2.5× effective capacity versus standard configurations. DDR4 reuse is already in production: Meta runs CXL-based memory expansion across millions of servers, pulling DDR4 modules from decommissioned machines rather than buying new. Malik is direct: "The first and foremost important use case within CXL is the recycling of DDR4." Structera A adds a near-memory compute accelerator beside the memory controller, offloading CPU and GPU cycles for recommendation inference, vector search, and database operations. Structera S4 is a CXL 3.1 switch bridging CPUs and GPUs to CXL pools even when host silicon lacks native CXL lanes — it includes a PCIe-over-CXL protocol conversion layer, critical for clusters built on older GPU generations.

Sub-productCXL VersionKey CapabilityTarget WorkloadsNotable Detail
Structera XCXLDDR4 or DDR5 memory expansion with onboard compressionGeneral inference, memory-capacity scaling2–2.5× effective capacity; enables DDR4 recycling from decommissioned servers
Structera ACXLNear-memory compute accelerator beside memory controllerRecommendation inference, vector search, database operationsOffloads cycles from CPU and GPU
Structera S4CXL 3.1Switch bridging CPUs/GPUs to CXL pools; PCIe-over-CXL conversionClusters lacking native CXL host siliconEnables older GPU generations to access CXL memory pools
FIG. 04 Structera CXL sub-product breakdown: features and target workloads — EE Times / Marvell Technology

Beyond the rack, Marvell's Photonic Fabric uses optical links to extend a shared-memory tier up to 50 meters — spanning adjacent racks within a data-center pod. The target is long-context inference: models where KV-cache alone overwhelms per-server DRAM. Photonic Fabric offloads up to 32 TB of warm KV-cache, and Marvell claims 2–3× token throughput improvement within the same power envelope and data-center footprint, though the company acknowledges the actual gain is workload-dependent. That caveat matters: the 2–3× figure assumes a memory-bandwidth-bound workload; compute-bound configurations will see less.

Claimed performance multipliers for Marvell memory disaggregation tiers (workload-dependent ranges)
FIG. 05 Claimed performance multipliers for Marvell memory disaggregation tiers (workload-dependent ranges) — EE Times / Marvell Technology — figures are workload-dependent; compute-bound configs will see less

Meta's millions-of-servers CXL deployment is the most concrete proof of hyperscaler adoption, but it's memory expansion, not full disaggregation. Photonic Fabric is the bolder architectural bet — optical shared-memory introduces new failure modes, latency jitter across 50-meter links, and operational complexity that rack-local DDR does not. For architects scoping this now, near-term action lands at Structera X and S4, where DDR4 reuse and CXL pooling work without optical infrastructure changes.

If your inference cluster's bottleneck is KV-cache capacity or memory bandwidth — not FLOP throughput — Marvell's three-tier stack gives you a product path at each radius, from server SSD to cross-rack optical. The DDR4 recycling entry point is the lowest-risk first step; save the Photonic Fabric evaluation for clusters where 32 TB warm-cache offload changes the economics.