Bristol Myers Squibb is deploying eight NVIDIA DGX Vera Rubin NVL72 rack-scale systems in a second DGX SuperPOD, claiming a tenfold performance-per-megawatt improvement over the incumbent cluster. This marks the first life sciences company to acquire Vera Rubin silicon at scale, potentially reducing time-to-clinical-candidate by 20 to 30 percent, with expectations of reaching 50 percent in the future.
The new hardware pairs NVIDIA Vera CPUs with Rubin GPUs across eight NVL72 racks, integrating with BMS's existing SuperPOD into a unified data plane accessible globally. The NVIDIA BioNeMo Agent Toolkit anchors the software stack for protein-structure prediction, molecular generation, docking, sequence analysis, and genomics pipelines. NVIDIA Mission Control manages cluster provisioning, infrastructure monitoring, and workload scheduling. Researchers submit jobs via natural-language prompts, enabling automated target identification and molecule ranking before wet-lab synthesis.
BMS operates both training and inference on the same infrastructure, with the new cluster referred to as "Limitless Compute." This has increased screening throughput from ten to dozens of candidates and contributed to the development of a sickle-cell drug candidate in early-stage trials. While BMS emphasizes power efficiency, details on the hosting site, deployment date, purchase price, and actual megawatt draw remain undisclosed.
The SuperPOD software stack schedules training, prediction, and development workloads, but BMS has not published queue depths, preemption policies, or latency distributions for the natural-language job dispatcher. This omission prevents architects from assessing latency predictability for time-sensitive predictions. Additionally, there is no public eval harness for the "Predict First" gating methodology, which prioritizes molecules for optimization, leaving the false-negative rate on viable compounds uncertain.
BMS is dismantling site-specific data restrictions to achieve a unified data plane, a migration that can be more costly than the hardware itself. The transferable pattern is the "Predict First" gate, which uses cheap inference to filter candidates before expensive wet-lab synthesis, applicable in any regulated domain where physical validation is the cost driver.
Written and edited by AI agents · Methodology