AMD announced on August 6, 2026 a definitive agreement to acquire Taalas, a Toronto startup that hardwires AI model weights directly into silicon. Terms were not disclosed. The Taalas team, co-founded by Ljubisa Bajic (former Tenstorrent CEO and AMD executive), will join AMD's AI group under Vamsi Boppana. The acquisition is subject to regulatory approval.
Taalas's HC1 chip fixes a model's dataflow and burns weights into mask-ROM. Separate SRAM handles KV cache and fine-tuning adapters. On Llama 3.1-8B, it delivers over 16,000 tokens per second per user — multiples of current GPU hardware. The trade-off: HC1 runs exactly one model. Swapping models requires cutting two new metal masks via proprietary tooling in roughly two months. Bajic, COO Lejla Bajic, and CTO Drago Ignjatovic built HC1 with a 24-person team on $30 million, then raised $169 million in February 2026. Total funding: $219 million.
| Attribute | Detail |
|---|---|
| Inference throughput (Llama 3.1-8B) | >16,000 tokens/sec per user |
| Weight storage | Burned into mask-ROM (fixed) |
| Runtime storage (KV cache / adapters) | Separate on-chip SRAM |
| Models supported per chip | Exactly one |
| Model swap lead time | ~2 months (two new metal masks) |
| Team size at acquisition | 24 people |
| Seed / build funding | $30 million |
| Feb 2026 raise | $169 million |
| Total funding raised | $219 million |
The HC2, due this summer, supports 20 billion parameters per chip. A one-trillion-parameter model maps to roughly 50 accelerators — within AMD's Helios rack capacity. Nvidia's equivalent setup would combine dozens of GPUs with 2,000+ Groq LPUs. DeepSeek-671B on HC1 requires 30 tape-outs, showing why the technology suits stable deployed models over rapidly iterating frontier weights.
| Scenario | Hardware | Accelerator Count |
|---|---|---|
| 1-trillion-parameter model (HC2, per chip = 20B params) | AMD Helios rack + Taalas HC2 | ~50 accelerators |
| 1-trillion-parameter model (Nvidia-equivalent) | Nvidia GPUs + Groq LPUs | Dozens of GPUs + 2,000+ Groq LPUs |
| DeepSeek-671B | Taalas HC1 | 30 tape-outs (accelerators) |
AMD's integration uses disaggregated inference: Instinct GPUs handle prefill (compute-heavy prompt processing), and Taalas accelerators handle decode (token generation). In December 2025, Nvidia entered a ~$20 billion non-exclusive licensing agreement for Groq's inference assets. Groq founder Jonathan Ross joined Nvidia while the company continued under new CEO Simon Edwards. AMD announced a similar solution with Cerebras on July 23 at Advancing AI 2026, where AMD GPUs handle prefill and Cerebras's wafer-scale engine handles decode. Taalas could displace Cerebras for stable smaller models.
Secondary applications include edge and physical AI: single-chip inference for models up to 8 billion parameters. This competes against AMD's existing FPGA accelerators and SoCs. Edge deployments update models infrequently. Intel occupied this structured-ASIC edge niche since 2018 via eASIC, primarily for network infrastructure and defense. AMD now has a comparable path.
The acquisition extends AMD's pattern since 2024. The Helios rack — shipping now to Meta and Microsoft — came from ZT Systems ($4.9 billion, 2024), Silo AI ($665 million, 2024), and MK1 inference software (2025). Taalas is the second Canadian inference-chip team AMD absorbed in roughly 12 months, following Untether AI in 2025. AMD committed 2 gigawatts of Instinct MI450 GPUs to Anthropic and signed a 6-gigawatt agreement with OpenAI in October 2025. At this infrastructure scale, hardwired decode silicon measurably reduces cost per inference.
| Company / Deal | Year | Value | Role in AMD Stack |
|---|---|---|---|
| ZT Systems | 2024 | $4.9 billion | Helios rack system (shipping to Meta, Microsoft) |
| Silo AI | 2024 | $665 million | AI software and model development |
| MK1 | 2025 | Undisclosed | Inference software |
| Untether AI | 2025 | Undisclosed | Inference silicon (Canadian startup) |
| Taalas | 2026 | Undisclosed | Fixed-weight decode silicon |
| Anthropic (supply commitment) | 2025 | 2 GW MI450 GPUs | Training / inference supply |
| OpenAI (supply agreement) | Oct 2025 | 6 GW GPU commitment | Inference infrastructure scale |
Taalas technology works only for models that have stopped changing. For teams running stable production models at high volume — code assistants, document extraction, real-time transcription — the performance-per-watt case is clear. Taalas claims etching weights into silicon is 100x less expensive than training a frontier model. For teams on quarterly refresh cycles, the two-month re-spin window is workable with planning. The acquisition makes sense as a stack component, not an Instinct replacement. AMD's CEO said GPUs should constitute the majority of the AI chip market.