What incubation confirms
On July 15, the Cloud Native Computing Foundation announced that its Technical Oversight Committee had voted to move HAMi to incubating status. HAMi is an open-source GPU virtualization middleware for Kubernetes, and incubation is the CNCF's signal that a project has cleared meaningful production adoption and governance bars. The project joined the CNCF Sandbox on August 21, 2024, and reached the incubating tier roughly eleven months later. That pace is quick by CNCF standards, and it reflects how urgently platform teams need a credible answer to sharing scarce accelerators across many workloads.
The maturity label carries practical weight for buyers. Incubating projects receive closer stewardship from the TOC and are expected to demonstrate a healthy contributor base, real end users, and a sustainable release cadence. HAMi currently ships at version 2.9.0 across sixteen releases, and the CNCF cites more than 2,687 contributors, a 43 percent year-over-year increase, alongside over 550 contributing organizations. Karena Angell, the CNCF TOC sponsor, framed the value plainly, stating that HAMi solves a real problem by scheduling and sharing accelerator resources on Kubernetes in a vendor-agnostic way.
The problem HAMi solves
GPUs remain the scarcest resource in most AI infrastructure budgets, and Kubernetes has historically treated them as indivisible. A pod either claims a whole card or it claims nothing, which strands capacity whenever a workload needs a fraction of a GPU. HAMi attacks that waste directly. It lets platform teams slice a physical accelerator into units by memory, by compute core, or by device count, then enforces hard runtime isolation so tenants sharing a card cannot trample one another. The result is far higher utilization of hardware that often sits half idle under coarse whole-device allocation.
The scheduling side is equally important. HAMi supports binpack, spread, and topology-aware placement, so operators can pack workloads tightly to free whole nodes or spread them for resilience. Xiao Zhang, a HAMi maintainer, said the project now supports dozens of heterogeneous GPUs and has grown into a global community with hundreds of contributors and hundreds of end users. For platform engineers, the appeal is that these policies apply without changes to the application code, which lets existing workloads benefit from finer-grained allocation on the day the middleware is installed.
Vendor neutrality as the differentiator
HAMi's most strategic property is that it does not tie a cluster to one accelerator vendor. It slices and schedules across GPUs, NPUs, DCUs, and MLUs, spanning silicon from multiple suppliers under a single scheduling model. That breadth matters because enterprises are increasingly forced to buy whatever accelerators they can source, and mixed fleets are becoming ordinary. A middleware layer that abstracts the differences lets a platform team present one consistent interface to developers regardless of the hardware underneath, which reduces both operational complexity and negotiating dependence on any single supplier.
We read this vendor neutrality as the reason HAMi is gaining ground where proprietary alternatives stalled. The dominant GPU vendor offers its own partitioning technology, and it works well within that vendor's ecosystem. HAMi extends the same discipline across a heterogeneous fleet, which aligns with how large buyers actually procure hardware in a constrained market. For CIOs managing multi-year accelerator commitments, an open, vendor-agnostic scheduling layer is a hedge. It preserves the option to shift workloads onto cheaper or more available silicon without rewriting the platform.
Who is running it in production
Incubation requires documented production use, and HAMi's case studies are substantial. DaoCloud reports managing HAMi across more than 10,000 GPUs spread over ten or more data centers in mainland China and Hong Kong, which places the project in genuinely large-scale operation. China Merchants Bank uses HAMi to manage diverse accelerator resources, a reference that matters because financial institutions apply stringent isolation and compliance requirements to shared infrastructure. Five independent CNCF case studies now document deployments across education, cloud platforms, and enterprise technology.
These references address the question every platform leader asks before adopting a scheduling layer, namely whether the isolation holds under real multi-tenant load. A bank running mixed accelerators through HAMi is a meaningful signal that the runtime isolation is trusted with sensitive workloads. The concentration of large deployments in China also reflects where accelerator scarcity and heterogeneous sourcing bite hardest, given export controls that push operators toward a wider mix of domestic and international silicon. The lessons those operators learn translate directly to any enterprise facing constrained GPU supply.
Where HAMi fits in the platform stack
HAMi sits below the workload and above the hardware, acting as the accelerator scheduler that the Kubernetes device plugin framework alone does not provide. Platform teams typically pair it with their existing scheduler, observability, and GitOps tooling, and it slots in as the component responsible for turning raw cards into shareable, isolated units. That positioning keeps it composable. Teams can add fractional GPU allocation without adopting a whole opinionated platform, which lowers the barrier for organizations that have already standardized on Kubernetes and want accelerator sharing incrementally.
The incubation milestone should accelerate ecosystem integration. Vendors of observability and cost-management tools tend to prioritize CNCF incubating projects because the maturity label predicts durable adoption. We expect tighter integrations with GPU monitoring, chargeback, and autoscaling systems to follow over the next several quarters. For platform engineering leaders building an internal AI platform, HAMi is now a credible default for the accelerator-sharing layer, with the governance backing and community depth that justify a multi-year bet.
Why this milestone matters now
The timing aligns with the central constraint of enterprise AI. Compute is expensive and hard to obtain, and the organizations that extract the most value from each accelerator will operate at a structural advantage. HAMi's core promise is exactly that efficiency, turning whole-device waste into fractional, well-scheduled utilization. Incubation tells risk-averse buyers that the project is stable enough to build on, which removes a common objection to adopting infrastructure software that sits in a critical path. For many teams, that assurance is the difference between a pilot and a production rollout.
We would still counsel diligence. Fractional GPU sharing introduces failure modes that whole-device allocation avoids, and the isolation guarantees deserve validation against your own workloads before you consolidate tenants onto shared cards. The reference deployments are encouraging, and they do not substitute for testing under your specific compliance and performance requirements. Used deliberately, HAMi offers one of the clearest paths to raising accelerator utilization while keeping a platform free of single-vendor lock-in. That combination is why its promotion is worth the attention of any leader funding AI infrastructure.


/filters:no_upscale()/news/2026/07/dolt-version-control/en/resources/1c719463f830464444cce6a285c5a4125cc88710480ad62d28a3108f60a25739b-1783069316047.jpg)
