BEYOND CHAT: THE ERA OF AGENTIC EXECUTION

Until recently, most businesses used generative AI in one way. Staff opened an assistant such as ChatGPT or Copilot, asked a question, and received a draft or a suggestion. The person then did the actual work. The tool advised, and the human executed.

That model is now being displaced. A newer class of software, usually described as agents, does more than answer. It takes a task and carries it through to a finished result, using your own apps and files with limited human input along the way. Anthropic offered an early version of this, Claude Cowork, as a desktop preview in January 2026. On 7 July 2026 it brought that agent to the web and mobile. Two days later, 9 July 2026, Meta and OpenAI each released products built on the same idea, within hours of one another. Three of the largest AI labs are now betting that software should do the work rather than merely describe it, and that shift is one every operator should weigh.

OpenAI turns the assistant into an operator

OpenAI's entry is ChatGPT Work. Hand it an objective and it pulls information from your apps and files, then produces finished material such as spreadsheets, slides, documents and simple web pages. It can stay on a task for hours, breaking it into smaller steps and working through them on its own. It connects to Slack, Microsoft Teams, Gmail, Google Drive, SharePoint, Salesforce, calendars and CRMs, and pauses for your approval before doing anything with real consequences. You give it a job, and it returns a result rather than a conversation.

It runs on OpenAI's new model, GPT-5.6, released the same day in three tiers. Each is priced differently, so a business can send routine, high-volume work to the cheaper option and reserve the top model for harder problems. The exact rates matter less than the logic behind them. Capable AI work is now something you route by cost, the way you already manage any other variable expense.


chatgpt

Figures are OpenAI's standard short-context API rates; OpenAI lists separate, higher rates for long-context requests. Meta's rates are per the Meta Model API.

Reuters summed up the pitch as being cheaper and more widely available than rival products. Speaking to CNBC, Sam Altman said the top model is 54 percent more token efficient on agentic coding tasks, while being as good or better on quality. In plain terms, it does that coding work using fewer tokens, which lowers the running cost. In its own announcement, OpenAI says more than five million people use Codex, the coding engine now built into ChatGPT Work, every week, and that more than one million of them use it for work outside software development. What began as a developer tool is reaching the wider workforce.

Meta changes course on pricing and strategy

Hours earlier, Meta released Muse Spark 1.1, a model built for the same kind of multi-step work, along with a paid interface for developers to use it. Meta says the model plans and coordinates tasks across different apps, works with tools it has not been shown before, and can read text, images, video and documents together in a single request.

The change in commercial strategy is the more telling part. For years Meta gave its Llama models away for free and positioned itself against the closed labs. Muse Spark 1.1 breaks with that. Through the new Meta Model API, Meta is now charging developers directly for access to the model, monetizing model access rather than giving it away. Speaking to Bloomberg before the launch, Mark Zuckerberg called the pricing very aggressive and said the model is among the most affordable on the market, at roughly a quarter of what OpenAI and Anthropic charge for comparable models. Meta expects to spend $125 billion to $145 billion on infrastructure in 2026, and it needs that outlay to start earning its keep. Meta's AI chief, Alexandr Wang, described the goal as models that handle work like a team of junior staff.

One clarification is worth keeping. Muse Spark's price is genuinely low, but it does not beat every rival. TechCrunch reports that at the entry level it sits close to Anthropic's cheapest model and OpenAI's lowest tier, Luna. The comparison is not one-directional: against Luna, Muse Spark charges more for input but less for output, so which is cheaper depends on the shape of the workload. The more notable point is higher up the range, where strong, near-top-tier performance is now available at a low price.

Anthropic's move set up both launches

The launch that framed the week came two days earlier. On 7 July, Anthropic brought Claude Cowork from desktop-only to the web and mobile, and that changed how the work runs. On the web and mobile, tasks now run on Anthropic's servers, so you can start a task at your desk, check on it from your phone, and collect the finished output anywhere. The desktop version still does its work on your own machine, where it has fuller access to your local files. A task can even run at a scheduled time with no device switched on. When the agent reaches a decision only a person should make, it sends the question to your phone.

Anthropic also shared usage data that says a lot about who these tools are really for. Across more than a million anonymised Cowork sessions from hundreds of thousands of organisations, over 90 percent of the work was not software development. The largest uses were ordinary business tasks: reconciling the quarter's spend and drafting the summary, turning a folder of contracts into a tracker with risks flagged, or building a client deck from call notes and pipeline data. This is the everyday work that fills a mid-market operator's week, and it is exactly what OpenAI and Meta are now going after.

The two flagship agents work on the same principle. You state the outcome you want, and the tool coordinates across your apps and files to produce it. Where they differ is in the details.


chatgpt

One caution applies to any such comparison. Much of the performance data these labs publish is measured on their own tests rather than an independent one, and the exact model behind each agent is not always disclosed. Read the claims as a guide, not a settled verdict.

The other contest, closer to enterprise budgets

It would be a mistake to see this only as a three-way race between OpenAI, Meta and Anthropic. Those labs build the underlying models. A second contest is being fought by the software companies that already hold a business's data and workflows: Microsoft is adding agents to Office and Copilot, Google to Workspace, and Salesforce, ServiceNow and others are building agents that work inside their own products. The difference matters when you decide what to buy. A model from a lab is something you bring in. An agent inside your existing software already sits where your work and data live. Most companies will use both, and the market will be decided by which of these two layers ends up owning the workflow, because that is the one that keeps the customer.

What the launches have in common

Look past the branding and these releases share a few clear traits. Each tool can use several applications at once, work through multi-step tasks, and hand parts of a job to smaller helper processes, all with little supervision. Writing a fluent answer is no longer what sets a product apart. The value now lies in getting the work done across the software a company already runs.

There is also a real price war, and it favours the buyer. OpenAI competes on value for money, Meta on low price, and xAI released its own low-cost model, Grok 4.5, a day earlier on 8 July. The cost of running these models keeps falling, and that is the number that decides which tasks are worth handing over. Work that was once too expensive to automate starts to make financial sense.

Why control matters more as agents act

An agent that can send an email, change a CRM record or move a file is a different kind of tool from a chatbot. A chatbot that gets something wrong wastes a minute of your time. An agent that gets it wrong can send the message, change the record or move the file before anyone notices. Every company shipping these products knows this, and each has taken a different approach to keeping the agent in check.

OpenAI wrapped ChatGPT Work in controls for administrators, who decide which tools the agent may use, plus a review step that checks important actions before they run. Meta says it tested Muse Spark 1.1 for chemical, biological, cybersecurity and loss-of-control risks before release. Anthropic holds every significant action for a person to approve before it happens. Each has picked a different method, but all three are working on the same problem: as an agent does more on its own, the safeguards around it start to matter as much as the model itself.

The risk is not hypothetical. On 1 July, the security firm Armadin described a way to break out of Cowork's protected environment on the Windows desktop version. Anthropic replied that the attack would only work if someone already controlled the machine. That specific case is debatable, but the direction is not. Before buying, an operator should be able to answer a few plain questions. Which actions can the agent take without a person present? Which of them can be undone? Where is the record of what it did, and who reviews it? A business that cannot answer these has not adopted an agent so much as taken on a standing risk it has not measured.

What this means for operators

It is tempting to file all this away as a technology story for later. That would be a mistake. The question is no longer whether AI can draft your next memo faster. It is which repeatable, multi-step tasks in your business can now be handed to software that plans and executes, and what checks belong around that software before it acts on anything that matters.

The business effect follows from this. When a repeatable task can be done by software instead of staff, its cost stops rising with headcount and starts rising with usage. That turns parts of routine knowledge work into a source of leverage. The saving is not automatic, though. These tools still carry real costs to integrate with existing systems, to supervise, to handle the exceptions they cannot, and to govern properly. The gain does not go to whoever buys the most powerful model. It goes to whoever redesigns the work around it and absorbs those costs deliberately. My own view is that within two years the real measure of AI adoption will not be how many staff use a chatbot, but what share of a company's repeatable work runs on its own, with a named person accountable for the result. That is the number boards will start asking for.

These tools are arriving quickly, but the judgement to use them well takes longer to develop and has to be built deliberately. An agent becomes a real asset only when you can explain what it did, review the decision behind it, and measure the result against an outcome that matters. Whether a business gains from AI or is left managing its problems comes down to that discipline, not to which model it picks.

This is the work we do at Occams.ai. We help mid-market companies choose which tasks to hand to an agent, connect it to the systems they already run, and put the review, audit and accountability around it that keeps the output safe to rely on.

If you are weighing which tasks to hand over and how to keep them in check, you are welcome to speak with our team.