Presence Puts Governance Ahead Of Raw Capability
OpenAI has launched Presence, a managed product for deploying and managing AI agents across customer service and internal operations, delivered through a limited general availability program. The framing matters as much as the software. For the past year most agent pitches led with capability: watch it resolve a ticket, watch it draft a reply. Presence leads with control. The product bundles policies, guardrails, approved actions, and evaluation tools, which are the exact questions CIOs raise in the first meeting before they discuss what the agent can do. OpenAI is reading the room correctly. Senior buyers stopped asking whether an agent can act and started asking what happens when it acts wrongly.
We read this as a deliberate repositioning of how enterprise agents get sold. The center of gravity has moved from model quality to operational trust. A CIO signing off on an agent that can issue refunds or close IT tickets carries the same accountability as approving a junior employee with system credentials, and they know it. By putting approved actions and human handoffs at the front of the offer, OpenAI is answering the objection that kills most pilots before production. That objection is rarely about accuracy. It is about who is liable when an autonomous system takes a real action against a real customer record, and whether the buyer can prove the guardrails held.
The Single-Job Discipline Is The Real Product Decision
Each Presence engagement starts with a single job. Resolving a billing dispute, handling an insurance claim, or clearing an IT service request, with the agent given only the knowledge and system access that specific job requires. This is the most important design choice in the whole announcement, and it is a governance decision dressed as a scoping decision. A narrow job means a narrow blast radius. If the agent only touches billing data, it cannot leak claims records or reach into HR systems, because it was never granted the keys. That constraint is what makes an approval board comfortable enough to say yes to a live deployment.
For technology leaders this is the pattern worth copying regardless of vendor. The failure mode we keep seeing is teams building one general assistant with broad access, then spending months trying to fence it in after the fact. Starting from a single job with least-privilege access inverts that work. You prove value on one workflow, measure resolution against a real baseline, and expand access deliberately as trust grows. It maps cleanly to how mature organizations already grant permissions to people. The lesson for your roadmap is to resist the demo that does everything and fund the narrow deployment that does one thing you can audit.
The Control Stack Is What CIOs Are Actually Buying
In its own words, OpenAI describes the offer plainly. "Presence brings together the components teams need to run agents in production: policies and standard operating procedures, guardrails, approved actions, simulations, evaluation tools, and a Codex-powered improvement process." Notice what dominates that list. Six of the seven components are governance and quality controls, and only the improvement loop speaks to raw capability. This is the plumbing that separates a working demo from a system a regulated bank will run against live accounts. Simulations let a team test agent behavior before it goes near a customer. Evaluation tools give the audit trail that a risk committee will demand when something eventually goes wrong.
We would push buyers to treat this component list as an evaluation rubric rather than marketing copy. Any serious agent platform should answer each item concretely. How are standard operating procedures encoded, and who can change them? What is an approved action, and how is it revoked in minutes if it misbehaves? Where do simulation results live, and can compliance read them? These questions expose whether a vendor has built for production or for the sales stage. Presence gives CIOs a checklist that works across suppliers. If a competing platform cannot map its features to policies, guardrails, approved actions, and evaluations, that gap is the answer.
Forward Deployed Engineers Change The Buying Motion
Presence deployments are led by OpenAI's Forward Deployed Engineers working alongside selected global systems integrators. This is not a self-serve product you provision with an API key, and the delivery model tells you who OpenAI is really selling to. Forward Deployed Engineers sitting inside the customer signal a services-led motion aimed at large enterprises with messy systems and heavy compliance loads. The integrator partnerships extend that reach without OpenAI hiring a consulting army. For a buyer, this means the price tag and the timeline look more like a transformation program than a software subscription, and the internal approval path should be planned accordingly.
The trade-off deserves a clear-eyed read. A hands-on deployment team dramatically improves the odds of reaching production, because the hardest part of agents is never the model, it is the integration with legacy systems and the encoding of real operating procedures. That is exactly what forward deployed engineers exist to solve. The cost is dependency. You are buying a relationship as much as a platform, and switching later means unwinding both. CIOs should negotiate for knowledge transfer from day one, insisting that their own teams learn to author policies, tune guardrails, and read evaluations, so the capability stays in-house when the engineers move on.
The Metrics Are Encouraging And Incomplete
OpenAI reports that its English-language phone support resolves 75% of inbound issues without human assistance, and that the Codex-powered improvement loop reduced human handoffs by 15 percentage points over 10 days. Both numbers are meaningful. A 75% autonomous resolution rate on phone support, historically the hardest channel to automate, is a strong signal that the technology has crossed from novelty into utility. The 15 point reduction over 10 days is the more interesting figure, because it points at a system that measurably learns from production traffic rather than staying frozen at launch quality. Early testers include BBVA in Mexican banking, SoftBank on Japanese-language support, and IAG on high-demand event scenarios.
We would still read these figures with discipline. A 75% self-service rate is a first-party number on OpenAI's own support, which is a friendly test case with clean data and motivated staff. Your billing disputes, your legacy CRM, and your regulatory constraints will not match that baseline, and the honest expectation is a lower starting point that improves over time. The handoff reduction is the metric to hold vendors to, because it captures the trajectory that matters in production. When you run a proof of value, fix the job, fix the baseline, and measure your own resolution and handoff curve over weeks. Borrowed benchmarks do not survive contact with your systems.
Four Platforms In A Month Reshape The Roadmap
Presence arrives as one of four enterprise agent platforms launched in roughly a month, alongside the Meta Business Agent Platform, NVIDIA and ServiceNow's Project Arc, and Google Gemini Enterprise. That clustering is the strategic context every CIO should register. The question of whether to adopt production agents has effectively been settled by the market, and the live decision is which platform and on what governance terms. Competition on this axis is good for buyers, because it turns governance features from differentiators into table stakes. When four credible suppliers all ship guardrails and evaluation tooling in the same quarter, no serious vendor can treat controls as an optional add-on any longer.
For your roadmap, the practical move is to run a structured comparison now rather than wait for the field to settle. Pick one high-volume, well-bounded job, define the baseline, and evaluate two or three of these platforms against the same governance rubric of policies, approved actions, simulations, and audit trails. The single-job discipline that OpenAI built into Presence is the right frame for the evaluation itself: small scope, least privilege, measurable outcome. The organizations that will lead on agents through 2027 are the ones treating this month's launches as a procurement bake-off on control rather than a technology bet on capability. Start the bake-off before the pressure to standardize arrives.


