A new name in AWS's chip supply chain
AWS and Qualcomm announced on September 8 a multi-generation partnership to co-develop custom silicon for AI inference and optical connectivity for large scale data centers. The deal builds directly on top of chips AWS already designs in house, Trainium for training, Inferentia for inference, and Graviton for general compute, adding Qualcomm's chip design and connectivity expertise as a named third party contributor rather than a pure component vendor. AWS VP Prasad Kalyanaraman said the goal is "more performant, efficient, and cost-effective infrastructure," while Qualcomm CEO Cristiano Amon framed it around a real constraint: "data center infrastructure will require advances in both computing and connectivity to deliver greater performance."
The commercial structure disclosed alongside the announcement puts real numbers behind the partnership. The framework sets a ceiling of up to 60 billion dollars in chip and technology purchases across multiple product generations, with the commitment vesting in tranches as Qualcomm hits commercial milestones and AWS crosses purchase thresholds, not a lump sum committed today. Alongside it, Amazon received a warrant to buy 25 million Qualcomm shares at 161.26 dollars each, expiring in September 2036, a structure Amazon has used before to align supplier incentives with its own long term spend.
What the chips actually do
The custom silicon is aimed specifically at AI inference, not training, using what Qualcomm calls a near-memory High Bandwidth Compute, or HBC, architecture designed for decode-heavy workloads, the token-by-token generation step that dominates the cost of running large language models in production. Qualcomm is targeting 4 to 8 times better decode performance per watt than conventional HBM-based designs, a claim that, if it holds in AWS's production fleet, would meaningfully lower the per-token cost of serving inference at scale. The roadmap has the AI200 part sampling in fiscal 2026, AI250 with the first generation HBC architecture in fiscal 2027, and AI300 in fiscal 2028.
The optical connectivity half of the deal is just as consequential for large clusters. Qualcomm is supplying 1.6 terabit module-class optical connectivity built on SerDes and optical DSP technology it acquired through Alphawave, extending to future generations as bandwidth needs grow. Marvell and Broadcom remain AWS's primary optical incumbents today, so this deal gives AWS a credible second source rather than an outright replacement, consistent with the multi-sourcing strategy AWS has run across its chip supply chain for years.
Why Amazon is diversifying away from Nvidia dependence
The strategic driver here is the same one behind Trainium and Inferentia in the first place: reducing AWS's dependence on Nvidia GPUs, which remain supply constrained and carry a substantial margin that AWS would rather capture or eliminate for its own infrastructure. Qualcomm's entry adds a chip design partner with mobile and edge silicon expertise that is distinct from Nvidia's GPU lineage, and the optical piece addresses a bottleneck, cluster interconnect bandwidth, that becomes the limiting factor once compute itself is no longer the constraint.
It is also notable competitively. Qualcomm now has disclosed AI silicon relationships with all three major US hyperscalers, Microsoft Azure, Meta, and now AWS, plus Meta as a customer for its separate Dragonfly C1000 CPU. That spread makes Qualcomm a credible fourth or fifth name in AI silicon alongside Nvidia, AMD, Broadcom, and the hyperscalers' own in-house chips, a diversification that benefits the entire market by giving buyers more leverage in future pricing negotiations.
What this means if you run inference workloads on AWS
If your workloads run on AWS-managed inference services rather than raw GPU instances, this deal is relevant to your unit economics even though you will never see a Qualcomm logo on an invoice. Lower decode cost per watt, if it materializes as claimed, should show up over the next two to three years as better price-performance on AWS's managed inference offerings, the same way Graviton's arrival gradually pushed down the cost of general compute instances relative to x86 equivalents.
The practical takeaway for infrastructure teams is to treat this as a multi-year roadmap, not an immediate purchasing decision. The AI200 sampling in fiscal 2026 will not be broadly available in production capacity for some time, and AWS has a track record of ramping custom silicon availability slowly behind reserved capacity commitments first. Track it the way you would track any new instance family: worth planning migration paths toward once benchmarks are public, not worth re-architecting around today.
The bigger picture for chip supply chain risk
For CTOs building multi-year infrastructure strategy, the real signal in this deal is structural rather than technical. AWS is willing to put a 60 billion dollar ceiling and a decade-long equity warrant behind a single supplier relationship specifically to avoid concentration risk on Nvidia, which tells you how seriously the largest cloud provider takes that risk. If AWS is hedging this aggressively, enterprises with heavy GPU-dependent workloads on any cloud should be asking their own providers how exposed their roadmap is to a single chip vendor.
It is also a reminder that the AI infrastructure arms race now runs through connectivity, not just compute. The 1.6 terabit optical piece of this deal matters as much as the inference silicon, because cluster interconnect bandwidth increasingly determines how efficiently a data center's compute can actually be used. Enterprises evaluating cloud providers on raw GPU counts alone are missing half the picture; the networking fabric behind those GPUs is becoming just as important a differentiator.



