Nvidia and Equinix Are Building an AI Inference Network for Companies That Will Never Own a Data Center
Cloud

Nvidia and Equinix Are Building an AI Inference Network for Companies That Will Never Own a Data Center

Equinix, Nvidia, and Together AI unveiled the Equinix Inference Exchange on September 2, a distributed platform meant to let enterprises deploy AI models across 280 plus colocation facilities instead of renting hyperscaler capacity or building their own infrastructure.

PublishedSeptember 3, 2026
Read time7 min read
Share

What got announced on September 2

Nvidia and Equinix, working with Together AI, unveiled the Equinix Inference Exchange, a distributed platform intended to let enterprises run AI inference workloads across Equinix's global colocation footprint of more than 280 facilities spanning 77 metro areas. The pitch is proximity: rather than routing every inference call back to a centralized hyperscaler region, the platform is designed to place model serving closer to where a company's data and users actually sit, using Equinix Fabric to handle the interconnection to public clouds, private networks, and other AI infrastructure providers. Nvidia VP Raj Mirpuri summarized the ambition directly: 'Equinix Inference Exchange turns the world's leading digital interconnection platform into a global fabric for AI inference.'

What was notably absent from the announcement matters as much as what was included. Neither company disclosed dollar figures, specific hardware volume commitments, or a firm revenue target, and the commercial launch is roughly two quarters out, landing around the first quarter of 2027. That is a meaningfully different posture than the multi-billion dollar, multi-gigawatt compute deals that have defined most hyperscaler and neocloud announcements this year. This is being positioned as an early-stage platform rollout, not a capacity transaction, and enterprise buyers should read it with that distinction in mind.

The case for colocation over the cloud, for inference specifically

The argument for running inference at a colocation edge rather than inside a single hyperscaler region rests on latency and data gravity. Training large models benefits from centralizing enormous compute in one place. Serving those models to production applications benefits from the opposite: sitting physically close to the data and users generating the requests, minimizing round-trip time, and avoiding the egress costs of shuttling data back and forth to a distant cloud region. Equinix's core business for two decades has been exactly that kind of distributed interconnection, originally built for network and financial services traffic, now being repositioned for AI workloads that need the same proximity logic applied to inference rather than packet routing.

For enterprises that are not hyperscalers themselves, meaning nearly every company reading this, that proximity argument doubles as a build versus buy answer. Standing up your own inference infrastructure at metro-edge scale is not realistic for a retailer or a SaaS company outside the largest enterprises. Renting capacity inside a single hyperscaler region solves centralized training needs but does not solve the last-mile latency problem for a global customer base. A colocation-based inference network, if it delivers on the pitch, offers a third path: production-grade AI infrastructure without the capital commitment of building it and without being tied to a single cloud region's geography.

This is not Equinix's first move here

The Inference Exchange builds directly on a partnership Equinix announced with Cisco and Nvidia in June 2026, which deployed the Cisco Secure AI Factory architecture across Equinix facilities and, working with Presidio, created a testing environment called the Programmable AI Technology Hub lab, where enterprises can validate AI infrastructure configurations before committing to a production rollout. Equinix SVP Gordon Mackintosh described that earlier effort as delivering 'the infrastructure AI workloads demand while giving customers a place to prove it out before they scale,' and Cisco's Cassie Roach framed it as proof that 'a trusted agile partner ecosystem can deliver secure, flexible AI infrastructure quickly.'

Reading the two announcements together, a clearer strategy emerges. Equinix is assembling a stack, standardized AI factory hardware blueprints from Cisco and Nvidia, a validation environment through the P.A.T.H. Lab, and now a software and networking layer through the Inference Exchange with Together AI, aimed at making its colocation facilities a credible alternative to hyperscaler AI infrastructure rather than just overflow capacity. Each individual announcement looks incremental. The pattern across three of them in under three months looks like a company trying to become AI infrastructure's neutral middle layer before the market decides that role belongs exclusively to the hyperscalers.

What the muted stock reaction actually tells you

Equinix shares rose roughly 2 percent on the announcement, which analysts characterized as a muted response given the partnership builds on existing collaboration rather than introducing something entirely new. That is a useful, unglamorous signal for enterprise buyers trying to separate genuine infrastructure shifts from AI-adjacent press releases. A stock market barely reacting to a joint announcement with Nvidia, still one of the most reliable stock-moving names in technology, suggests investors see this as execution on a stated strategy rather than a surprise catalyst, which is a reasonable characterization to bring into your own procurement evaluation.

That is not a criticism of the announcement's substance so much as a calibration of its urgency. There is no dollar figure to model against your budget yet, no committed hardware volume to check against your capacity needs, and no confirmed pricing to compare against hyperscaler inference costs. The right response for most enterprise buyers is to track this toward its Q1 2027 launch rather than treat it as immediately actionable, while noting that Equinix, Cisco, and Nvidia have now shipped three coordinated announcements in three months, which is a faster cadence than most infrastructure partnerships sustain without real underlying commercial traction.

The build versus buy question this actually answers

For CTOs weighing where production AI inference should live, this announcement adds a third column to a decision that has effectively been binary: build your own GPU infrastructure or rent it from a hyperscaler. A colocation-based inference network, once it reaches commercial availability, offers metro-level proximity without the capital intensity of the first option or the single-vendor lock-in of the second. That is a genuinely useful option for retail and commerce businesses with geographically distributed customer bases who need consistent inference latency across regions a single cloud provider does not serve equally well.

The catch is timing and maturity. Nothing here is production-ready today, and enterprises that need inference infrastructure decisions made in the next two quarters should not wait on this platform to mature. The more useful move right now is to have your infrastructure team map where your current inference latency actually breaks down by geography, so that when the Equinix Inference Exchange or a competing colocation-based offering does reach general availability, you already know exactly which workloads and which metros would benefit from moving to it.

The roadmap call

Add this to your Q1 2027 vendor evaluation list now rather than waiting for launch, since colocation procurement cycles and contract negotiations with providers like Equinix typically run longer than a cloud provider's self-service sign-up flow. If your organization already has a colocation relationship with Equinix for other workloads, that existing relationship is worth leveraging to get early access or pilot terms once the Inference Exchange opens for enterprise testing. Ask your Equinix account team now which metros are first in line for the rollout, since early access will likely track the metros where Equinix already has the densest Nvidia and Together AI infrastructure in place.

More broadly, this is a signal that the AI infrastructure market is starting to fragment beyond the three or four hyperscalers that have dominated the conversation all year. Nvidia has an incentive to diversify where its chips get deployed and sold through, and colocation providers have an incentive to become more than real estate for other people's servers. Watch whether Digital Realty, the other major colocation player, announces something comparable in the next two quarters. If it does, that confirms this is becoming a real third category of AI infrastructure procurement rather than a one-off Equinix strategy.

Tagged#news#cloud#infrastructure#datacenter#aws#azure#gcp#hyperscalers#equinix#nvidia#together-ai#ai-inference#colocation#ai-factory#bare-metal#gpu-cloud#neocloud