
Ethan Mollick’s occasional verdict on which AI to use arrived on 23 July with a name missing. Google’s Gemini, a fixture of the Wharton professor’s earlier guides, is no longer recommended as a primary system for anyone who wants AI to do real work.
The judgement matters because of who reads it. Mollick’s guides are among the most widely circulated independent references for non-technical professionals choosing AI tools, and this edition moves the question from which chatbot answers best to which vendor can hand a model a computer. For marketing teams settling next year’s software budgets, that reframes a licence from a writing aid into a piece of operational infrastructure, and narrows the shortlist to two suppliers.
Mollick names two systems for anyone wanting more than a chatbot: ChatGPT and Claude, starting at $20 a month. He splits the advice by stakes. For low-stakes tasks such as a recipe, a letter or a quick question, he says the free default models are all at least fine and the choice barely matters. For high-stakes questions, such as a second opinion on a medical or legal concern, he directs readers to the most advanced models they can reach: Claude’s Opus and Fable, or ChatGPT’s GPT-5.6 Sol, each set to at least a “High” thinking level, on the grounds that those models have lower error rates and score higher on ability tests in complex fields.
The substantive change is the agentic layer. Mollick describes an agentic system as one that combines a model’s reasoning with tools that let it plan and act, or, more bluntly, gives an AI a computer to use. He identifies two routes to that. The vendor’s computer runs the work through ChatGPT Work or Claude Cowork; the user’s own machine runs it through Codex or Claude Code. Work and Cowork return a finished artefact for review, while Codex and Claude Code expose the working: the files changed, the commands run, the record of what happened.
Mollick gave both systems the same instruction to show the difference. Connect to his Gmail, prepare materials for an MBA seminar and answer outstanding messages on the topic. Both ran for roughly ten minutes and returned teaching materials and a drafted email, work he put at a couple of hours by hand.
Google is out, on Mollick’s reading, because it has no leading frontier model and nothing close to Codex and Claude Code. He hedges the call, writing that it could change quickly, and he keeps two Google products on the recommended list: Gemini Notebook, formerly NotebookLM, which he rates the most useful interface for research across many sources, and Gemini Omni for editing video directly.
That verdict rests on judgement rather than measurement, and Google has shipped into the gap it describes. At I/O in May the company launched Antigravity 2.0, a desktop application positioned against Codex and Claude Code, and Gemini Spark, a standing agent for Gemini Enterprise and Workspace customers, initially gated behind a $100-a-month AI Ultra tier. Developer Simon Willison, reviewing the guide on 27 July, made the narrower version of the same point: Google still has no established entry in the Codex and Cowork category, and Spark has yet to prove itself.
Mollick’s other exclusions matter more to teams that do not get a free choice. Microsoft Copilot, he writes, is acceptable for office documents but lags badly on agentic work, which covers any marketing department standardised on Microsoft. Chinese open-weights models including Kimi K3, DeepSeek and Qwen he rates surprisingly capable, though he notes they demand expertise to run as agents. And Claude has no image generator, where ChatGPT and Google both do, a gap he says may matter to anyone who needs images in their work.
Claude drafted the email in Mollick’s seminar test and ChatGPT sent it. The cause was a permission he had granted ChatGPT earlier, and he uses the incident to make the practical point: leave every action set to ask for approval first, which is the default, until the system’s mistakes are understood.
That same setting is his stated defence against prompt injection, where an agent reading email or browsing the web meets instructions planted by someone else. Mollick says the labs are working on the problem and models have grown more resistant, but that it is not solved. His advice is to limit what an agent can touch and keep approval switched on for anything that sends, spends or deletes.
The exposure is concrete for marketing teams rather than theoretical. An agent wired into a CRM, a shared mailbox or a publishing tool can send, spend and delete on a brand’s behalf, and the approval setting is the only thing standing between a planted instruction and a live campaign.
Anthropic’s pricing page lists Claude Pro at $20 a month billed monthly, or $17 on annual billing. OpenAI’s help centre lists ChatGPT Plus at $20 a month, monthly only. Neither company publishes a fixed sterling price, so UK buyers pay the dollar figure with VAT on top.
Mollick’s own footnote is the line budget-holders should read twice. The $20 tiers include real but limited agent usage, he writes, agents burn through those limits quickly, and the more expensive plans mostly buy more hours of AI labour rather than a smarter model. The entry price is a floor.
The fine print also complicates his model advice. Anthropic’s pricing page shows Fable, one of the two models he recommends for high-stakes questions, reaching Pro subscribers through usage credits billed at standard API rates rather than as part of the subscription. Opus is included. A team following the guide to the letter on a single Pro seat will spend beyond $20 to do it.
The guide’s shelf life is short by construction. Anthropic shipped Claude Opus 5 on 24 July, the day after publication, and the page now carries a modification timestamp of 26 July.