OpenAI is making a big bet that artificial intelligence will not just answer questions but take on whole tasks. With ChatGPT Work, the company connects its large language models to the digital tools that many knowledge workers use every day: email, calendars, Slack, Notion, Figma, cloud storage and more. The product was released last month and is included in the company’s cheapest paid tier, which costs $20 per month. OpenAI’s marketing language describes a future in which intelligence goes beyond answering questions to helping people turn their biggest ideas into reality.
From code to general knowledge work
ChatGPT Work grew out of Codex, an AI tool originally designed to help software engineers write code. Coding has become one of the most successful early uses for AI agents, but programmers are a small slice of the broader workforce. OpenAI and other labs need to reach accountants, doctors, investors, human-resources teams and business operations. The commercial logic is simple: autonomous agents consume more computing resources and are potentially more valuable per user than a chat session.
Early demand is still far from matching the ambition. A study supported by OpenAI found that in June 98 percent of OpenAI employees used Codex, while only 17 percent of paying organizational subscribers and less than 1 percent of individual subscribers had used the company’s agentic coding tool. The gap between internal enthusiasm and external adoption explains why OpenAI is spending so much effort on product design. The company says ChatGPT Work is meant to make non-engineers feel the way software engineers do when they hand a complex job to an AI agent.
Thibault Sottiaux, who leads core product work at OpenAI, said that in this new phase ChatGPT can complete very complicated tasks autonomously in a way that is both safe and useful. He described the mission as bringing everyone along.
A matter of trust
The hardest part of AI agents is control. To get real utility, a user must give the model access to private accounts. Andrew Ambrosino, lead engineer for OpenAI’s desktop app, has granted the app access to his inbox, Slack, phone and applications like Notion and Figma. He acknowledged the risk: if he asks it to write a document, it might pull from a private direct message and share something it should not. “I’ll do it for the job. I will take the personal hit here and there if I have to. And I haven’t had to yet,” he said.
Ambrosino’s point is that engineers cannot improve the experience without living with the same exposure as future users. The alternative is a product that remains too limited to be useful. OpenAI’s early employees used agents before they were friendly, sometimes seeing error messages meant for software developers. Over time the company made Codex more general and converted it into the platform now called ChatGPT Work.
Interface matters
Every language model needs a harness. The harness decides what information the model sees, which tools it can use and how it communicates results. For agentic products, the harness also gives the model instructions for long-running projects. A command-line interface was enough for programmers, but it is not enough for a teacher or a marketing director. OpenAI says mainstream adoption requires a familiar interface and clear buttons.
The designers draw analogies with skeuomorphism, the practice of making early digital tools look like their physical counterparts. Calculator apps that resembled small calculators looked old-fashioned, but they helped people make the transition from manual work to digital work. The team says buttons are part of the same process. They may disappear later as users learn to ask the model directly. For now, discoverability matters.
A promising but imperfect test
In early testing, ChatGPT Work can be impressive. One reporter asked it to take a child’s preschool schedule from email and create Google Calendar events, which it did. It also produced auto-updating financial metrics for public companies and built a searchable database of space launches. The same reporter did not trust it with bank accounts, interview notes or article drafts, but found the value would have been higher with more access.
The experience is not friction-free. Giving the model read-only permission to a cloud service proved confusing and circular, with error messages that did not explain the problem. Some settings appeared only in the web app, even when the user was working in the mobile app. ChatGPT could create calendar events when linked to Google Calendar, but it could not create a new calendar. Users who did not choose a high effort level were likely to receive work that felt like it came from the worst intern imaginable.
OpenAI acknowledges some of these issues are still being solved. Joe Gershenson, engineering lead for the harness, said the effort-level controls are not intuitive for new users yet. He said his team is working on helping people get the right level of reasoning and that users should watch the space.
Competing with Claude and open source
OpenAI did not create the agent harness category alone. Anthropic’s Claude Code became a standard-bearer for AI coding and set a template for conversational agent workflows. Early OpenAI engineers had bet on a more autonomous approach, believing the model could handle a task from start to finish without frequent check-ins. Anthropic’s product asked users to review options along the way, which gave the model fewer chances to go off track. That approach proved more effective, and OpenAI later added more opportunities for interaction.
Company engineers say they do not pay close attention to rival harnesses, claiming the real differentiator is OpenAI’s latest models. Download statistics tell a slightly different story: Claude Code led in interest until roughly April, when Codex took a small lead. Some of the movement may reflect complaints about safety restrictions or compute capacity inside Anthropic, but it also suggests OpenAI is finding a better product-market fit.
The competition is not limited to big labs. Vertical AI companies are also trying to serve lawyers, sales teams and other specialized workers. There is also a vibrant open-source community building alternative harnesses. In one comparison by Databricks, an open-source harness called Pi outperformed Codex while using the same GPT-5.5 model. Pi has been used to build projects such as OpenClaw and Cloudflare OS. Its creator argues that simpler harnesses can be just as powerful, especially when the model is strong enough to modify its own tools.
What makes a good harness?
OpenAI’s engineers believe a great model matters more than a complicated harness. They compare their approach to the “bitter lesson” in AI research: better general models tend to beat narrowly crafted domain systems in the long run. That is why the team tries to keep the harness simple and expose only the information the model truly needs.
Still, measuring success outside of coding is difficult. A software program either runs or it does not, but a presentation, a business strategy or a sales pitch is much harder to evaluate. OpenAI uses a benchmark derived from 44 occupations and hundreds of knowledge-work tests, supplemented by user feedback. The design team is also mindful that its own workflows may be unusual. It constantly asks whether a feature is something everyone will use or only something OpenAI employees need because they are too deep in the technology.
Cost and lock-in
Even successful agentic products face economic questions. A reporter using the $20 per month subscription consumed more than 80 million tokens in four days, which the model estimated would cost $65 at standard prices. That is more than three times the monthly subscription price, and it highlights how expensive autonomous agents are to operate. OpenAI’s Sottiaux said the company works every day on efficiency. He pointed to a recent 80 percent price cut for users of the Luna model and said users should be able to do the same tasks for less spend within six months.
There is also the question of whether these tools will create durable loyalty through the difficulty of setting up dozens of integrations, or through the comfort users gain as the AI learns their work patterns. Until then, OpenAI’s challenge is not just making models smarter. It is making agentic software reliable enough for people to hand over the keys, simple enough for non-engineers to understand, and cheap enough for the company to keep subsidizing while the market matures.
Source: TechCrunch News