On August 6, AMD announced the acquisition of Taalas, a Toronto startup doing something unconventional: instead of running a model on a general-purpose chip, it hardwires that model's weights directly into the silicon. The result, according to the company, is radically faster inference at a fraction of the energy cost — at the price of giving up the flexibility of a general-purpose chip.
Taalas was founded in 2023 by former Tenstorrent and AMD architects and had raised $219 million before the deal. Its approach is "hardwiring": rather than loading a model's weights into memory and running them on a flexible processor like a GPU, it encodes them directly into the chip's design. That eliminates most of the data movement between memory and compute that today is the dominant bottleneck in inference at scale.
The company's current chip is designed specifically for Meta's Llama 3.1 8B. According to reports on the acquisition, it reaches roughly 17,000 tokens per second per user, with a full rack drawing between 12 and 15 kW — versus the 120-600 kW a comparable GPU rack requires for a similar workload. That's an order-of-magnitude difference in energy consumption.
The deal, with undisclosed terms, is expected to close in Q4 2026 pending regulatory approval. AMD plans to integrate Taalas's technology alongside its Instinct GPUs, EPYC processors, the Helios rack platform, and ROCm software, targeting system-level products rather than an isolated chip.
Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.
Contact Carlos Montiel