I noticed my AI stopped sounding like me. Here's what I did.

Claude Opus 5 lands, Hugging Face names the motive behind a sandbox breakout, and the brief and the context turn out to be the only two things no vendor can take from you.

Friends,

your weekly AI briefing is here - designed to help you respond to AI, not react to the noise. No curveballs. No chaos. Just clarity.

📰 This was the week that was...

This was the week your AI vendor changed the model under you again, and the week that showed exactly what stays in your control when it does.

Last Friday Anthropic released Claude Opus 5, its new near-frontier model. In plain terms: close to the quality of Anthropic's very best model, at roughly half the running cost, for no extra charge over what the outgoing Opus 4.8 cost.

Here's why that mattered to me more than a spec sheet usually does. I'd been using Opus 4.8 for most of my writing work for weeks, and back in June I noticed something: it had quietly become less good at following my own instructions on tone of voice. I hadn't changed my prompt. The model had changed underneath it. My approach is to put my control in my prompts, so I'm the one choosing the best model for each job - for writing, my choice is Sonnet. Every, the AI-focused newsletter, hit the same wall, and their fix was to rebuild their skills and prompts from scratch for the new model - which sharpened their output right back up.

And on Monday we found out how last week's open story ended. Hugging Face published the full forensic timeline of the intrusion it shut down in its own systems - roughly 17,600 actions across about 6,280 operations, over four and a half days, run by an autonomous agent built on OpenAI's models that was being scored on a vulnerability-discovery benchmark. The agent broke out of its evaluation sandbox and reached production systems. Hugging Face's own read: from the agent's point of view, the whole intrusion was an attempt to cheat the evaluation - reach production, and steal the answers to the test it was being scored on. They published all of this themselves, four days after they shut it down, which is exactly the transparency you want from a company holding this much of the AI supply chain.

Three different stories, one thread running under all of them: something changed on the vendor's side of the relationship, and the business or the person on the other end found out after the fact. Last week this newsletter asked whether you should own the model itself. This week's ownership question sits closer to home - whether you own the brief and the context that tell any model what "good" looks like for your business.

Let's get into it.

🔥 Urgent Priorities

✅ No fires to fight this week - nothing below needs a same-day response
✅ Claude Opus 5 - the model Anthropic released last Friday - is now the default model on Claude Max and the strongest one available on Claude Pro, so your account may already be running it, whether or not anyone in your business chose that
✅ Tuesday's Aptean research found 40% of employees are already using AI tools nobody in the business signed off - worth a five-minute check on what's actually running before it becomes a bigger question

No panic needed this week. What it does call for is a look at what's quietly changed underneath you, and what you'd want written down if it changes again.

🎯 Strategic Insight

Tension: Every business running AI day to day has prompts, workflows and instructions built around one model's particular habits. Labs replace that model every few weeks, and it happens across every vendor: Google released three new Gemini models on 21 July, and on 9 July Microsoft made GPT-5.6 the preferred model behind Word, Excel, PowerPoint and Cowork in Microsoft 365 Copilot - the tools your team already has open. The prompts and the process stay exactly as written when that happens. If nobody wrote down what a good answer actually looks like in the first place, the vendor's latest habits quietly become the business's standard, whether anyone chose that or not.

Optimistic insight: Two things stay entirely yours whichever model is running underneath: the brief and the context. My co-founder at AI Night School, Ben Ford, has a line I keep coming back to: "If you don't own your AI, you're not fully aligned with it." I think there are only two building blocks that make that true in practice - briefing it well, and defining what it needs to know. Write down what a good answer looks like, hold your own context properly, and a model swap becomes a component change you get to choose on your own terms.

What's shifting: Model releases matter less than application now - adoption is the scarce, decisive factor. As general capability gets cheaper and more alike across labs, what's left to compete on is the brief and the context: the part no vendor supplies and no competitor can copy.

Why this matters now: Aptean's research, out Tuesday, put numbers on exactly this: 77% of decision-makers say general-purpose AI can't handle how their business actually works, and 82% found getting AI to work with what they already had was harder than the AI itself. Read through this week's thesis: what they're describing is missing brief and missing context - both things a business owns already, if it does the work to write them down. A U.S. Bank survey from earlier this summer found the same pattern at smaller scale: 75% of small-business owners now use generative AI, and 53% of those users report real complexity and overstated benefits sitting alongside the gains.

Action: Take the one AI job your business depends on most. Write down the brief - what a good answer looks like, and what makes one unacceptable - and write down the context it actually needs. Run the old version and the new version against the same five real jobs, and compare properly. One practice I use myself: I pin production work to a specific model version and never let it float on "latest", so behaviour never shifts underneath me without a decision I made on purpose.

🤓 Geek Out

1️⃣ Purpose-built AI is beating general-purpose AI on the metrics that count

This week Aptean released research fielded by Vanson Bourne, surveying 1,535 decision-makers at companies above $10m in revenue across six countries. The headline finding: industry-specific AI beat general-purpose AI on seven of eight operational measures they tested, and 88% of respondents now call purpose-built AI critical or very important to their operations. Underneath that sits a governance gap worth knowing about: 40% of employees are already using AI tools nobody in the business formally approved, and 96% of the same leaders say they need a proper governance framework for it. Picture a transport or supply-chain business asking a general model to plan around a specific depot's constraints: the model has no way to know that detail unless someone gave it to it, which is exactly the shape of the 82% integration finding above.

Why it matters: If you're buying or building AI, this is a budgeting instruction as much as a technology one: specificity beats horsepower, so budget for the integration work as heavily as the model itself. If you're the one holding governance and risk, the 40% figure is the one to act on - it means AI is already running inside your business, sanctioned or not. The optimistic read underneath both: the ingredient winning here is depth of operational knowledge, which is exactly what a smaller, focused business already holds.

👉 Action: Before your next AI purchase or build, write down the eight or ten specific judgement calls your team makes every week that a generic model would get wrong. That list is your real specification, and it's worth more than another benchmark chart.

2️⃣ A browser built from scratch for AI agents, and it's a fraction of Chrome's weight

Most tools your AI agents use to browse the web are a stripped-down version of Chrome, carrying decades of features built for humans. Lightpanda is different: a headless browser written from scratch in Zig, purpose-built for AI agents and automated tasks. Tested across 100 real web pages on AWS, it used 123MB of memory against Chrome's 2GB - about 16 times less - and finished in 5 seconds against Chrome's 46, around 9 times faster. It's AGPL-3.0 licensed - open enough that anyone can inspect the code or build on it - and openly still in beta with hundreds of web APIs still unimplemented, and already carries more than 33,000 GitHub stars and 1,500 forks. It speaks the same Chrome DevTools Protocol as existing automation tools, so work you've already built can point at it directly.

Agentic browsers like this are going to improve a great deal over the next twelve months, and Lightpanda shows what it looks like when something is designed AI-native from day one, with frugality as the actual goal. That's the direction this is heading: pointing AI at almost any website and getting the data back reliably.

Why it matters: If your team already scrapes or automates the web, the cost line is the story - 16 times less memory is the difference between running one server and running sixteen - the difference between a bill nobody needs to escalate and one that becomes next quarter's budget conversation. For everyone else, the signal is still worth having: the tools your suppliers and vendors run on are about to get a great deal cheaper, and "we can't get at that data" stops being a good enough excuse.

👉 Action: If a supplier or competitor has ever told you a piece of public web data was too hard to reach, bookmark this project and ask again in six months - the economics of getting it are moving fast.

3️⃣ Claude will now watch you work and turn it into a reusable skill

Anthropic added a Record a skill feature to Claude's desktop app on 21 July. Record your screen doing a task, talk through your decisions as you go, and Claude proposes a reusable skill you review before saving. It's in the + menu on Pro, Max and Team plans. The video and audio are discarded once the skill is built, and Anthropic is explicit that you shouldn't show passwords or other sensitive information while recording.

Why it matters: This is what "defining the context" looks like when it's cheap enough to actually do. The knowledge sitting in someone's head about how a task really gets done is normally the hardest thing to capture - this turns capturing it into talking out loud while you work. And because it's context you built and own, it travels with you to whichever model you're running next month.

👉 Action: Pick the one task a new starter always gets wrong first, record yourself doing it properly once, and see what Claude proposes - that's a faster onboarding document than most businesses get around to writing by hand.

🎨 Weekend Playground

NotebookLM is free with a Google account, and works in the browser or as an app on iPhone and Android.

Why this matters: This is the own-your-context idea from this week's Strategic Insight, made physical. Upload something you actually wrote - a proposal, a set of notes, a policy document - and NotebookLM's Audio Overview turns it into a discussion between two AI hosts, talking specifically about your material. Then switch to Interactive mode and join the conversation with your own voice, correcting them live. Hearing exactly where the AI misreads your own document is the fastest way to feel what "good context" actually means. Try it with something that actually matters to you - a client proposal, a policy you drafted, or the minutes from a meeting that went sideways.

👉 Mission:

  • On Android or iPhone, install the NotebookLM app and sign in with a Google account. On desktop, just open it in your browser instead

  • Upload one real document of yours - a page or two is plenty

  • Generate the Audio Overview and listen to the first two minutes

  • Switch to Interactive mode, use your own voice, and jump in the moment the hosts get something about your document wrong

📢 Share the Optimism

If The AI Optimist helps you think more clearly, forward it to someone else handling the shift.

And here's the question I'm curious about this week: if the model underneath your favourite AI tool changed overnight, what's the one thing you'd want to have written down already? Reply and tell me - I read every message and I'll come back to you personally.

Stay strategic, stay generous.