Back to school - AI Summer Round Up

All the summer moves from Anthropic, Open AI, Google Gemini and Copilot.

Friends,

your weekly AI briefing is here - designed to help you respond to AI, not react to the noise. No curveballs. No chaos. Just clarity.

Three Fridays passed with nothing from me. The last edition landed on 14 August, and a lot moves in three weeks, so this one does the catching up.

A quick note from Hugo: We’re going to be talking about the importance of education in this edition. For UK based readers, we can help: AI Savvy Leaders is a free, government-funded course for leaders, run with Working Knowledge. It is built for the leader's own head: the room to work out what AI is actually for in your business, before a single licence gets handed to anybody. If you want to know whether your organisation qualifies, email [email protected] and Oscar will talk you through it.

(Disclosure: AI Night School is my own venture.)

πŸ“° These were the three weeks the assistants got cheaper and easier to hand out

Everything the four big assistants shipped while I was away was about getting a model in front of more people, more cheaply, with the paperwork already done. Here is the whole period in one place.

Assistant

What changed, 13 August to 2 September

What it means for you

When

ChatGPT

Zero data retention on its most capable models, an admin plugin for ChatGPT at work and Codex, its coding tool, ChatGPT for Teens, and ads arriving in the app

Three of those are things your compliance officer asked for last year. The fourth is how the free tier gets paid for

18-31 Aug

Claude

Fable 5.1 and Mythos 5.1 - around 25% cheaper for typical work, repeat lookups down 75%, and an option to keep your data on your own cloud

A price cut, and the option to keep your data where you choose. The work you were already doing costs less to run

2 Sep

Gemini

Google Pics - image generation and editing inside Docs, Slides and Drive - plus Gemini 3.7 Flash and Workspace admin controls, plus rules that stop confidential files leaving the organisation

One new feature where you already work, and the console your IT team needs in order to say yes

13 Aug - 1 Sep

Copilot

A deck or document built from a prompt with your brand templates enforced, Planner status reports written for you, and, for UK tenants, Claude models working inside Excel and PowerPoint - on by default for tenants created since March, with Word to follow

The brand template is the one to notice: a governance control people switch on willingly, because it saves them work. Worth checking which models your own tenant has turned on

18-31 Aug

Read down the second column and count the new capabilities. It is pricing, admin consoles, where your data lives, brand governance and task plumbing. Four companies, three weeks, and the same direction: the model is cheaper to buy, easier to approve and simpler to hand out.

Three weeks ago this newsletter argued that the returns follow the design - that an agent is only the interface, and what pays is the system built behind it. The table above makes that design cheaper to execute. The part where somebody works out what the model should be doing is still yours.

Let's get into it.

πŸ”₯ Urgent Priorities

βœ… No fires to fight this week

βœ… Three organisations published their own post-mortems while you were away, with the numbers left in

βœ… One of them raised its own risk rating

The UK's AI Safety Institute published a full incident report into its own security testing: across 122 runs, ten produced nineteen actions nobody had sanctioned, aimed at real people and real projects on the live internet. It caught the behaviour itself and contained it inside an hour. The detail worth your attention is what the agents did about each other. They passed working accounts and tools between themselves, unprompted, so whatever one had worked out the next could pick up and reuse.

OpenAI followed with its account of the Hugging Face incident, naming four patterns where a model drifts from what it was asked to do. It now requires chain-of-thought monitoring - reading the model's own working, step by step, while it runs - on every training run where a model uses tools.

Anthropic published its own alignment and security review the day before, including a checker that spots an attempt to break out of the space a model is meant to work inside and stops it before anything runs. In the same week it moved its own risk assessment from "very low" to "low" in its August risk report.

Three post-mortems in three weeks, and one organisation marking its own homework harder than it had to. That is a field deciding to learn out loud while the lessons are still cheap.

No panic needed this week. What it does call for is five minutes noticing that the people building this now publish what goes wrong - which makes it a reasonable thing to ask of a supplier, and a reasonable thing to expect an answer to.

🎯 Strategic Insight

Tension: you came back to the desk expecting a summer of launches to have moved the plan, and you are looking for what to change.

Optimistic insight: the plan you left in July still holds, and it points where it pointed then - at the skill of the people using this. Sixty-five per cent of what makes a person valuable at work holds its value against AI for the next five to ten years. Four per cent of it shifts quickly. That number comes from research I published over the summer, and it is the most useful thing I know for a leader deciding where to put money.

What's shifting: the vendors spent the three weeks on their own constraint, which was getting a model in front of more people, more cheaply, with the approvals pre-cleared. Availability is cheaper than it has ever been. Choosing what to put where is still hard work, and getting people to use it well once you have chosen is harder again.

Why this matters now: you are back at a desk with a budget question on it. Three weeks changed the price of the tool and left the cost of the change exactly where it was.

There are three jobs in this, and they run in order. Create the headspace to understand AI. Make room in the P&L using AI. Then reinvent the business on what you find. The first one is the hardest, and it is what a quiet three weeks is for: the mental space to strategise and prepare. No model release hands you that.

That sharpens the standing advice to invest in training into something you can actually spend against. Training earns its keep when it targets judgement - breaking a piece of work into its parts, briefing each part properly, defining what good looks like, and deciding which parts stay with a person because a person is what they need. The interfaces get easier every quarter on their own.

I know this one from my own business. The ceiling on how fast we grow is the size of our trained trainer pool, and it has been for a while. It comes down to how many people we have who can stand in a room and do the work properly, and that number moves at human speed whatever shipped last week.

Action: take one piece of work already on your desk this month and break it down out loud with the person who owns it. Part by part. For each part, say in one sentence what good would look like. Then mark which parts you would hand to a model, which stay with a person, and why. Forty minutes, on live work, with the person who will have to run it afterwards. Do that four times this quarter and you have trained somebody properly, using work that had to happen anyway.

Natural intelligence supported by silicon intelligence. Three weeks of releases moved the silicon along nicely. The other half of that sentence is still your call, and it is still the half that pays.

πŸ€“ Geek Out

1️⃣ Five million dollars to find out what AI does to people

Anthropic has put five million dollars into independent researchers building open-source ways to measure what AI does to the people using it - how a model behaves when someone comes to it in a mental health crisis, or for company, or at three in the morning. Grantees keep full independence and publish in the open, which is what makes the results usable by anybody. Applications close on 21 September.

Why it matters: wellbeing is about to become measurable, and measurable things end up in procurement questions. If your organisation puts an AI assistant in front of staff or customers, how it behaves with someone having a bad day currently has no standard answer. Within a year it will, and the evidence will be public.

πŸ‘‰ Action: ask whoever runs your customer-facing AI what it does today when somebody tells it they are struggling. If nobody knows, that is the finding, and it is worth an afternoon to go and find out.

2️⃣ Somebody finally stopped the models seeing the exam paper in advance

Google DeepMind ran what it calls the world's first double-blind evaluation of a top-tier commercial model, with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. The problem it goes after is contamination - models having already met the test questions somewhere in their training, which flatters the score and quietly makes every benchmark you have ever read less useful than it looked. The set-up means the evaluators never see the model's inner workings and Google never sees the questions, so neither side can tune to the other.

Why it matters: every AI supplier you speak to quotes benchmark numbers at you. This is the first serious attempt at a scoring system where the marker and the candidate cannot see each other's homework, and it hands you a question worth asking.

πŸ‘‰ Action: next time a supplier quotes a benchmark, ask who ran it and whether the model had met the questions before. How readily they answer tells you as much as the number did.

3️⃣ Sixty-seven human capabilities, mapped against what AI actually does

The number in this week's Strategic Insight comes from here. Humancraft is 46 pages mapping 67 human capabilities and 94 work primitives - the smallest useful units of knowledge work - against what this technology does today. Eight of the primitives come out human-exclusive. The finding that stayed with me while I was writing it: all 94 describe how a person shows up to work - the judging, the noticing, the deciding - and that is where the durability sits.

Why it matters: it gives you a map for a conversation most organisations are having in the abstract - which parts of a job this technology reaches, and which parts it leaves alone. You can hold your own roles against it in an afternoon.

πŸ‘‰ Action: take the three roles you would most struggle to replace and mark which of their capabilities land in the durable sixty-five per cent. That is your training budget, pointed at something.

(Disclosure, three ways: I wrote it, Sherpas AI published it and I am a co-founder there, and The AI Optimist co-published it. It also sits behind an email signup, so it collects addresses for us. Read it with all of that in mind.)

🎨 Weekend Playground

Google Pics landed on 1 September and it lives inside Docs, Slides and Drive. Describe an image and it makes one, change a single object in it without disturbing the rest, and translate the words inside a picture. It comes with Google AI Pro and Ultra and with most Workspace business plans, so check which one you are on before you go hunting for it.

Why this matters: the editing is where this gets interesting. Changing one thing in a picture you already have - the mug on the table, the sign on the wall, the language on the poster - is the move that fits real work, and this is the first time it has been available inside the document you were already writing.

πŸ‘‰ Mission:

  • Open a Doc or a Slide you are actually working on this week and make one image for it from a plain description

  • Change one object in it. Just the one. Notice how much of the rest survives untouched

  • Take a picture that already has words in it and translate them

  • On the phone, the same model is in the Gemini app on both Android and iPhone, so run the same three moves there and work out which one you would use on a Monday

πŸ“’ Share the Optimism

If The AI Optimist helps you think more clearly, forward it to someone else handling the shift.

And here's the question I'm curious about this week: which part of one job in your organisation would you protect from AI on purpose, because it is the part that makes the job worth doing? Reply and tell me - I read every message and I'll come back to you personally.

Stay strategic, stay generous.

Hugo & Ben