OpenAI shipped a super app for the desktop
On July 10, 2026, OpenAI released ChatGPT Work, positioning it as a single entry point for white-collar productivity across desktop, web, and mobile. The app operates across a user's applications and files, executes long-running tasks, coordinates multiple tools, and produces business documents, presentations, spreadsheets, and working websites. It integrates with Microsoft 365, Google Drive, Slack, and Notion, which means it reaches into the systems where knowledge work already happens. The framing is deliberate: OpenAI wants to own the surface where employees start their day, ahead of Microsoft Copilot, Salesforce Agentforce, and Google's own assistants that occupy the same space.
We treat this as a category move rather than a feature drop. When a model vendor ships an application that spans your document stack, your chat tool, and your knowledge base, it is competing with the productivity suites you already license. Faisal Kawoosa, co-founder and chief analyst at Techarc, put the shift bluntly, saying the exploratory stage of AI is over and organizations can derive value from it today. For a CTO, that changes the question from whether to pilot generative tools to which vendor owns the agentic workflow layer, and whether you want that layer to be OpenAI, an incumbent, or something you assemble yourself.
Three engines at three prices
ChatGPT Work runs on the GPT-5.6 family, which OpenAI split into three tiers. Sol is the workhorse at $5 per million input tokens and $30 per million output tokens. Terra sits in the middle at $2.50 and $15. Luna is the budget option at $1 and $6. That spread lets buyers route cheap, high-volume tasks to Luna and reserve Sol for hard reasoning and coding work. Sam Altman, OpenAI's chief executive, said Sol is 54% more token efficient on AI coding tasks, which matters because token efficiency, not headline price, usually decides the real cost of running agents at scale.
The tiering is the part enterprise architects should study first. A single flat model forces you to overpay for simple work or underperform on hard work, and OpenAI has now removed that excuse by pricing three engines you can mix. We would build routing rules that match task complexity to tier, because the difference between running everything on Sol and running most of it on Luna is the difference between a pilot that pencils out and one that does not. The vendors that win agentic budgets will be the ones whose economics survive contact with production volume, and OpenAI is clearly pricing for that fight.
The benchmarks a CTO should weigh
OpenAI backed the launch with numbers that speak to real workloads. GPT-5.6 Sol scored 53.6 on Agents' Last Exam and 80 on the Artificial Analysis Coding Agent Index, which OpenAI says is 2.8 points above Fable 5 while using less than half the output tokens and less than half the time at about one-third less cost. On security, Sol hit 73.5% on ExploitBench, up sharply from 47.9% for GPT-5.5. OpenAI says it ran 700,000 GPU hours of automated red-team evaluations before shipping. Those figures describe a system built to act, and to be tested against the ways acting agents go wrong.
We read benchmarks as a starting point for evaluation, never a substitute for it. A coding-agent score and an exploit benchmark tell you the model can reason through multi-step work and probe systems, which is exactly what an autonomous agent inside your estate will attempt. The relevant question for a technology leader is whether those capabilities show up on your tasks, with your data, under your controls. Before committing seats, we would run the same tasks across Sol, Terra, and a rival model, measuring completion rate, token cost, and failure modes side by side rather than trusting the vendor's headline chart.
Bill shock is the governance problem
The most useful warning at launch came from an analyst, not the vendor. Neil Shah, vice president of research and partner at Counterpoint Research, said the AI wave has delivered productivity gains but rising token consumption has also created bill shocks for enterprises. That is the quiet risk in a super app that runs long, multi-tool tasks on your behalf: every autonomous action consumes tokens, and an agent that loops or over-researches can turn a modest workflow into a large invoice. Consumption-based AI moves the cost surprise from the contract to the usage, where finance has the least visibility.
For buyers, this makes cost governance a first-order design requirement rather than an afterthought. We would insist on per-agent budgets, hard token ceilings, and dashboards that attribute spend to teams and workflows before a single department goes live. The same controls that prevent runaway bills also surface the agents that deliver real value, which is the information you need to expand deliberately. Treating token spend as an ungoverned utility is how promising pilots become budget liabilities. Treating it as a metered, owned line item is how you keep an agentic rollout inside the numbers you promised the board.
The consolidation squeeze is real
ChatGPT Work lands in a crowded field. OpenAI is pushing directly against Anthropic's Claude Cowork, which debuted in January 2026, and against Microsoft Copilot, Salesforce Agentforce, and Google's assistant stack, all chasing the same multi-year corporate contracts. For enterprises, that competition is good news on price and pace, and a genuine headache on strategy. Every one of these vendors wants to be the layer your employees live in, and each integration you accept raises the cost of switching later. The decision in front of technology leaders is which agentic surface to standardize on, knowing the choice compounds over time.
We would resist the urge to let this get decided by whichever tool employees adopt first. Shadow adoption of a super app that reaches into Microsoft 365 and Slack creates the exact governance gap that agentic AI keeps producing, where capability outruns control. The disciplined path is to pick a primary surface deliberately, negotiate it against a credible alternative, and keep your data and identity layer portable enough that you are not captive to one model vendor. Vendor consolidation can lower cost and complexity, and it can also hand a single supplier your entire productivity workflow. Choose it on purpose.
What we would do before signing
The honest position on ChatGPT Work is that it is capable enough to matter and priced to spread, which is why it deserves a deliberate decision rather than a reflexive one. We would start with a bounded pilot on a workflow that has clear value and clear guardrails, route tasks across the three tiers to learn the real economics, and instrument token spend from day one. We would compare completion quality against at least one rival on our own tasks, and we would keep security review in the loop given how much of the estate the app can reach through its integrations.
The larger point for enterprise leaders is that the agentic workforce moved from concept to procurement question this month. OpenAI has made the buy side of build-versus-buy credible, and the incumbents will answer within weeks. The organizations that benefit will be the ones that decide their standard on purpose, govern token cost like any other spend, and keep enough optionality to switch when the next tier or the next vendor changes the math. The tools are ready for real work today. The discipline to deploy them without losing control of cost or data is the part we still have to supply.


)
