Anthropic in early talks to run Claude inference on Microsoft Maia 200 custom silicon via Azure
Anthropic is in early-stage discussions with Microsoft to deploy Claude large-language-model inference workloads on Azure servers equipped with Microsoft's custom Maia 200 AI accelerator, according to CNBC and sources close to the talks. No agreement has been signed, and the discussions remain preliminary. If completed, the deal would make Claude the first frontier-class external model to validate Microsoft's custom silicon at production scale. Maia 200, launched in January 2026 on TSMC's 3-nanometer process, is designed exclusively for AI inference and claims over 30% better tokens-per-dollar performance compared to the latest GPU silicon in Microsoft's fleet.
Anthropic currently hosts Claude on a diversified infrastructure mix: Google Cloud TPU v5p pods, Amazon Web Services Trainium and Inferentia2, and on NVIDIA GPUs via Microsoft and third-party cloud providers. A Maia deal would add a fourth major custom-silicon option and reduce Claude's per-token inference cost, a metric that drives the unit economics of every frontier lab. Microsoft's Maia program has faced delays and external validation challenges; Amazon's Trainium and Google's TPU have years of customer precedent. Anthropic's evaluation involves assessing Maia 200's numerical precision tradeoffs and whether FP8 inference meets Claude's latency and quality requirements.
For infrastructure architects, the talks underscore the shift toward custom inference silicon and away from GPU-as-commodity. Anthropic's multi-cloud hedging strategy demonstrates how frontier labs are fragmenting their hardware commitments to avoid vendor lock-in and optimize cost. A signed agreement would mark a validation point for Microsoft's silicon roadmap and signal that hyperscaler-designed chips can serve external models at competitive economics.