Meta Begins Production of Iris, Its First In-House AI Chip, to Cut Its Reliance on Nvidia

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-09-04 | By: Carlos Montiel | Reading time: ~4 minutes

Meta has moved past PowerPoint chip designs and into actual manufacturing. According to an internal memo, production starts this month on Iris, its first proprietary AI accelerator in the new MTIA generation — another step in hyperscalers' race to stop depending exclusively on Nvidia.

What Iris is and where it came from

Iris is the code name for Meta's new datacenter chip, part of its MTIA program (Meta Training and Inference Accelerators) — a four-generation family of in-house accelerators designed to train and run inference for the AI models behind Facebook and Instagram. Meta designed the chip together with Broadcom, while actual fabrication is handled by TSMC — the same "in-house design, third-party manufacturing" playbook already used by Google (TPU) and Amazon (Trainium).

Production starts this month

According to the internal memo cited by Reuters, Iris completed its bug-testing phase in roughly six weeks with no significant issues, clearing the way to begin manufacturing in September 2026. Iris production sits within a broader Meta-Broadcom partnership, extended this year through 2029 to cover multiple future MTIA generations.

The number that explains the urgency: Meta is targeting 7 gigawatts of compute infrastructure by the end of 2026, with plans to double that to 14 gigawatts in 2027. At that scale, every percentage point of dependence on a single GPU vendor translates into real supply and pricing risk.

Part of a broader hyperscaler trend

Meta isn't alone in this move. Google is several generations into TPU, Amazon has Trainium and Inferentia, and Microsoft is developing its own Maia silicon. The pattern is consistent: the more Big Tech spends on AI compute, the more economic sense it makes to design at least a portion of that hardware in-house, even if Nvidia remains the dominant supplier for the bulk of the workload.

What this means for companies buying AI compute

For infrastructure teams currently relying on Nvidia GPU-based instances (via AWS, Azure, GCP, or specialized providers), this trend is a medium-term signal, not an immediate shift: AI compute supply is diversifying across more silicon architectures. There's no need to react today, but it's worth making sure any critical inference architecture avoids over-coupling to a single hardware family — workloads that are portable across providers will have more options (and likely better pricing) over the next 12-18 months.
Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contacto · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com