- The AI Optimist | AI news and strategy for leaders
- Posts
- Your AI assistant has been doing a judge's job all along
Your AI assistant has been doing a judge's job all along
A startup and a frontier lab both shipped AI built to decide, score and forecast this week. Sort your tools by what they hand back before a vendor sells you an assistant for a judge's job.
Friends,
your weekly AI briefing is here - designed to help you respond to AI, not react to the noise. No curveballs. No chaos. Just clarity.
๐ฐ This was the week that was...
This was the week a new kind of AI model got a name.
On 15 September, San Francisco startup TypeSafe AI introduced Jev, the first "System One Model" - it skips the conversation entirely, taking an input and a schema and handing back a typed decision, score or forecast, with a confidence level, in a fraction of a second. At DevDay on 29 September, OpenAI followed with a Decisions API of its own, built on a specialised version of its Luna model and designed to "focus Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers" - in practice, classifying content, routing requests or choosing what an agent does next. It's in limited preview for now. A frontier lab landing in the same family within a fortnight turns one startup's launch into proof of a real category.
The rest of the week moved too, in a different family. Microsoft's new Copilot Autopilot - "give it a name, a role and a goal, and it goes to work" - is entering private preview now, with wider availability still to come. OpenAI's Dots, always-on agents available in eligible markets, is the parallel already rolling out. Jev and the Decisions API are Judges. Autopilot and Dots are Actors.
If you want a free, hands-on way to try the idea yourself, Claire Vo's "Jev for beginners" guide walks through sorting 1,700 pull requests into categories for under 10 cents, and turning 4,500 YouTube comments into a searchable dashboard - a route in for anyone on your team who likes to build.
Last week the price of thinking fell. This week a different kind of model showed up, for a different job.
Let's get into it.
๐ฅ Urgent Priorities
โ No fires to fight this week
โ If your team built Gems in Gemini: they move to "Skills" from November (personal accounts) or March 2027 (Workspace), and Skills need a paid Google AI subscription - worth a note now of which Gems your team would actually miss.
โ Worth doing now: run your AI use-case list through my question - what do I get back from it?
No panic needed this week. What it does call for is an hour running your own AI list through that question.
๐ฏ Strategic Insight
Tension: Most AI plans have one slot for every use case: an assistant you type into. Decisions, searches, rankings and judgments all get routed through chat, because chat is the only shape of AI most leaders have met.
Optimistic insight: I sort every AI release that lands in my inbox with one question: what do I get back from it? That question splits cleanly into six families.
Family | What you get back | Types |
|---|---|---|
Assistants | A response to any instruction | Text and reasoning, multimodal, voice |
Makers | Media: picture, video, sound or 3D | Image, video, voice/music/sound, 3D scenes |
Readers | Text or data pulled from your own material | Speech-to-text, document/image readers, text processors, visual scene analysers |
Finders | A ranked shortlist of matches | Semantic search, rerankers, recommenders |
Judges | A decision, score, forecast or plan | Decision/classification, forecasters, optimisers and planners |
Actors | Actions taken, or a simulated world | Digital agents, robot brains, world models |
Underneath the six families sit twenty types, and between them they place almost every AI release you'll see this year. Jev sits squarely in one of the oldest jobs on the map: a Judge, the kind of work spam filters and credit-scoring models have done for years. This week's story is how fast and cheap doing that job has become.
What's shifting: Marketing labels don't settle which family a release belongs to. World Labs calls Atlas a world model. What it mostly produces is pictures, video and 3D scenes, which makes it a Maker by the question test. NVIDIA's Cosmos 3 is the one actually simulating a world, which makes it an Actor.
Why this matters now: A frontier lab and a fast-moving startup both shipped Judges in the same fortnight. A decision you'd once have sent to an Assistant now has a faster, cheaper home, and the category will be in your vendors' pitches by Christmas.
Action: Take your list of AI use cases and write the family next to each one. Anywhere you've put an Assistant on a Judge's job - routing, scoring, approving, triage - ask your vendor whether a Judge would do it faster and cheaper.
๐ค Geek Out
1๏ธโฃ I've switched from Claude Cowork to Codex. Expect to change your main model at least every six months.
I've started using OpenAI's Codex, moving away from Claude Cowork. The write-up that helped me decide was Composio's comparison of the two tools: Codex came out 3-2 on points that matter day to day - more reliable at following instructions, easier delegation to the cloud, and better value at the entry tier - while Claude Code still holds the edge on long, complex sessions. It's dated 18 August, so the exact models it tested have already moved on. The decision process behind the verdict is what's still worth borrowing.
Why it matters: Your AI stack is a standing decision, one you'll revisit on a schedule - the tool worth paying for this month may be a different tool by spring.
๐ Action: Put a twice-yearly date in the diary to re-run the comparison for whatever AI tool your team relies on most, the way you'd review any other supplier.
2๏ธโฃ BCG: getting real value from agentic AI is 70% a people project
BCG's new research puts a number on something worth saying out loud: "10% of the effort should go to the algorithms, and 20% to the technology and data. The remaining 70% should target changes to people, organization, and processes." The companies BCG rates as ahead of the pack are already three times more likely than the ones lagging to be planning which roles should exist in an agentic workplace - 55% against 17% - and 89% expect AI to create new work as well as remove old tasks. For more depth on the people side, this piece on how agentic AI reshapes organisational structure is worth your time.
Why it matters: This is the case I keep making, now with a number attached: invest in how you upskill people, the culture you build and how you organise them, because that's where the majority of the value sits.
๐ Action: Before your next AI budget conversation, ask what share of it is going to people and process alongside the tool itself.
3๏ธโฃ Cambridge's Tessera maps smallholder crops in Senegal, beating DeepMind's AlphaEarth on a fraction of the compute
Cambridge researchers trained Tessera, an open-source model, on satellite imagery of Senegal's groundnut-growing region, working with the UN World Food Programme. It classified crops with 84% accuracy, ahead of Google DeepMind's AlphaEarth by 28% in one test, using a fraction of the computing power.
Why it matters: Small, open and specialised beat big and general here, which is the frugal echo of this week's wider story - capability doesn't always need the biggest model in the room.
๐ Action: If a supplier is pitching you the largest model as the safest choice, ask what a smaller, task-built one could do for less.
๐จ Weekend Playground
This weekend, watch "Hit the Brakes Except for Me" with your family - a 4-minute AI-generated satire shared with me by my friend Joh Mulholland - then have the debate it sets up: is AI destroying art?
Why this matters: Five minutes of screen time buys you a proper conversation round the table - the kind that doesn't happen on its own.
๐ Mission:
Watch it together
Each person argues one side of the debate
Swap sides and argue the other
Score the debate out of ten, Jev-style, if you want a Judge in the room too
If The AI Optimist helps you think more clearly, forward it to someone else handling the shift.
And here's the question I'm curious about this week: which of your AI use cases is in the wrong family? Reply and tell me - I read every message and I'll come back to you personally.
Stay strategic, stay generous.
Hugo & Ben