Good morning.

OpenAI shipped GPT-5.6 on July 9, and the whole market is busy grading its coding score.

That's the least useful thing about this release for you. The most useful thing is a button that hands a job to a team of AI agents and gets back a single finished answer.

Let’s talk about what that button is, when it's worth pressing, which model and thinking level to point at a job, and the paste-in prompt that catches Sol when it lies to you about being done.

—Sam

In Today's Issue

  • The agent button you can use today

  • When Ultra is worth it, when it's a trap

  • How to brief a job so Ultra splits it

  • Match the model and thinking level to the job

  • Why better input beats more thinking

  • The reward-hacking problem that makes "done" a lie

  • The internal-usage curve worth reading

  • Eleven copy-paste prompts, customer research to verification

  • The plays by business type

The Agent Button You Can Use Today

Two years ago, running a team of AI agents on one problem (splitting the work, chasing five things at once, handing you back one finished answer) took a developer.

You needed a framework with a name like LangGraph or CrewAI, someone to wire it together, and a budget to keep it alive.

Multi-agent work sat behind an engineering team. If you didn't have one, you watched other people do it and waited for it to reach you.

In the recent release from OpenAI it became a button.

It’s not trivially easy and you should take advantage of that.

GPT-5.6 went live that day: three models, Sol, Terra, and Luna, across ChatGPT and Codex (OpenAI). The internet did what it does. It benchmarked the flagship, put Sol at the top of the agentic leaderboards, noted it costs about half what Claude Fable 5 does, and moved on (Simon Willison). Real, and beside the point for you.

The hardest part of serious AI work is getting several agents to divide a task, work it in parallel, and combine the results.

That's now a mode you select in the chat window you already pay for. It's called Ultra. You click it, describe the job, and a team of agents goes to work inside one request.

What separated operators with engineering teams from everyone else is now a setting anyone can reach. When a capability crosses that line, the edge moves. It now belongs to whoever knows when to use it, which model to point at a job, and how to catch this model when it lies to you about being finished, because it does.

So this issue skips the spec sheet (you can get benchmark numbers anywhere) and gives you the no-code operator's manual for a specific, closing window:

The next 90 days, while most of your market is still admiring the coding score and hasn't found the button.

Most of it is a setting you change or a block of text you paste. The few things that genuinely need the API are boxed off at the end, so you can skip straight past them.

Where You'll Find It

This is what the launch coverage skipped, and it's the difference between "it exists" and "it's available to you." Ultra isn't switched on for everyone equally (OpenAI):

  • In ChatGPT Work, Ultra is available on Pro and Enterprise seats.

  • In Codex, it's available on Plus and above.

  • It rolled out in beta, so if you're on the right plan and still don't see it, that's a rollout question, not a you question.

If you don't see Ultra, check your plan tier first. This is a menu setting, not a prompt trick: no amount of clever wording summons a mode your seat doesn't include. And if it's worth it for your work (the next section is how you'll know), it may be the best reason you've had all year to move up a plan.

Do this today. Open ChatGPT Work or Codex and find the mode selector. Just locate Ultra and confirm whether your seat has it. You can't use the most important capability in this release until you know where the button is and whether it's yours to press.

When Ultra Is Worth It, and When It's a Trap

A team of agents isn't free, even on a subscription.

Ultra works harder and, depending on your plan, draws down your usage faster than a normal request, and it takes longer to come back, because four agents doing real work and then combining it isn't instant (OpenAI).

So you don't reach for it out of habit. You reach for it when the job fits. Here's the test.

Can this task split into parts that don't depend on each other, and is it on a clock?

If yes to both, Ultra is the tool. If no to either, it's a waste: a normal Sol request does the same job faster and without burning the allowance.

Tasks that pass:

  • A competitive sweep across ten companies, due this afternoon. Each company is independent. The deadline is real.

  • An audit across a dozen things that don't touch each other: accounts, listings, pages, locations.

  • A multi-market analysis, five regions, the same questions asked of each, no cross-dependency.

Tasks that fail:

  • Writing one carefully-argued strategy memo. That's a single line of reasoning. Split it across a team and you get four shallow takes with a visible seam down the middle.

  • Thinking hard about one thorny decision. There's nothing to parallelize. That's a job for a normal request with the thinking level turned up, which we get to shortly.

One failure mode is already costing people their usage allowance: reaching for Ultra on work that can't be split. They see "most powerful mode," assume more is always better, and spend a team of agents on a single-thread problem, getting a patchy, stitched-together answer for the trouble.

The fix is the one-line test. Splittable and on a clock, or don't spend it.

The Proof It Earns Its Keep

When the work genuinely decomposes, the parallel team earns its keep. On Terminal-Bench 2.1, a hard, independent test of multi-step technical work, Sol scores 88.8% in its standard mode and 91.9% in Ultra (OpenAI).

That three-point lift is the decomposition doing its job on a task that genuinely splits: independent pieces, worked at once, converged into one answer. That's the pattern, and it's why Ultra is the headline and not a footnote.

Do this today: Find the one recurring job in your week that splits into independent parts and runs against a deadline: the weekly competitor check, the monthly multi-account review, the pre-launch sweep. Pick the most splittable one you've got, run it through Ultra once, and time it against how long the manual version takes you.

How to Hand Ultra a Job So It Splits Cleanly

Selecting the mode is half of it. The other half is describing the task so the team of agents divides it the way you'd want.

Left vague, it guesses at the split, and a bad split is where those seam-down-the-middle answers come from. Give it the seams yourself.

The move is simple:

Name the parts, and name what "done" looks like for each. You're not writing code. You're briefing a team the way you'd brief four capable freelancers. Tell them how the work divides and what each should come back with.

Here's a vague Ultra prompt and the same job briefed properly.

Vague, so the model has to guess how to divide it:

Research my competitors and tell me how I stack up.

Briefed, so the split is handed over and so is the finish line:

I compete with these ten companies: [list them]. Research each one independently and, for every company, come back with the same five things:
1. Their headline pricing and the tier it starts at.
2. Their main positioning claim, in their own words.
3. One thing they clearly do better than us.
4. One thing we clearly do better than them.
5. Their most recent notable move (a launch, a change, an announcement).
After all ten, synthesize: where do we sit in this set, and what are the two gaps most worth closing? Flag anything you couldn't verify rather than guessing.

The second version decomposes cleanly: ten independent lookups, one shared template, one synthesis at the end, which is the shape Ultra is built to run.

Notice the two safeguards baked in. "The same five things" forces consistency across the parallel agents so the synthesis has something to compare, and "flag anything you couldn't verify" pre-empts the confident-guessing problem you'll meet again in a few sections.

The pattern generalizes, take, for example, a multi-market analysis: name the markets, name the identical questions asked of each.

An audit: name the items, name the checklist applied to all of them. Whenever you can write "for each ___, come back with the same ___," you've briefed a job for a team instead of a single assistant, and you've done the decomposition the model would otherwise guess at.

Do this today: Take the Ultra candidate you found a minute ago and write the brief for it now: the parts, the shared template, the synthesis question at the end. Two minutes of briefing is the difference between a clean parallel result and a patchy one you have to redo.

Pick the Right Model and Thinking Level

Ultra is the special occasion. One everyday skill separates great answers from mediocre ones on the same tools, which is to match the model and the thinking level to the job.

Most people pick one model, leave it in one setting, and send everything through it. Then they're annoyed when a quick question takes three minutes or a hard question gets a shallow answer. Both are the same mistake, a wrong setting for the job.

GPT-5.6 gives you two dials. Learn them and the quality of what you get back jumps without changing a word of your prompt.

Dial One: Which Model

Three models, and in the chat apps this often shows up as a quality-or-speed choice in the picker (OpenAI).

Model

Reach for it when

Luna

You want a fast answer to a clear, simple task. Quick lookups, cleanups, first-draft anything. Speed over depth.

Terra

Your everyday default. Planning, writing, analysis, real work with a few moving parts. The one you'll live in.

Sol

The task is genuinely hard, high-stakes, or long-running. The moment a wrong answer costs you something. Slower, deepest.

The instinct to fight is "Sol is the best, so Sol for everything." Sol is the slowest and, on limited plans, the heaviest on your allowance.

Sending it a quick reformatting job is like booking a surgeon to remove a splinter:

You wait longer for something a faster model nails instantly. Match the model to the weight of the task.

Dial Two: How Hard It Thinks

On top of the model, GPT-5.6 lets you set how much it thinks before answering, from a quick low-effort reply up through progressively deeper deliberation, topping out at a setting called max (OpenAI).

The rule from the people who use this all day runs against instinct:

Start with the lowest thinking level that gets the job done. Turn it up only when the task needs it.

  • Low: quick, clear tasks.

  • Medium: planning and analysis with a few moving parts. The sane default for real work.

  • High: genuinely hard, multi-step problems, or careful checking.

  • Max: the hardest problems, where being wrong is expensive.

For a lot of tasks, cranking the thinking level barely changes the answer. It just makes you wait. Before you turn the dial up, give the model a better prompt and better information. That fixes more than raw thinking time does.

Max is depth. Ultra is a team. One last thing so you never confuse them: turning the thinking level to max makes one model think harder about one problem. Ultra splits the work across a team. Max is one expert locked in a room for an hour. Ultra is four experts each taking a section and comparing notes. Hard-but-single problem, use max. Big-but-splittable job, use Ultra.

The Whole Thing on a Sticky Note

DEFAULT: Terra, medium thinking.
Faster, simpler job? Luna, low thinking.
Genuinely hard or high-stakes? Sol, high thinking.
Hardest single problem you've got? Sol, max.
Big job that splits into independent parts, on a deadline? Ultra.
Rule: start low, turn it up only on failure. Fix the prompt before you raise the dial. Better input beats more thinking.

The Proof It Matters

Box tested Sol on realistic business work, document-heavy tasks across a dozen industries, the kind you'd want the top model for.

  • Sol beat the previous version across the board. Clinical case review climbed from 46% to 58%.

  • Recomputing student grades under a changed weighting scheme went from 63% to 74%.

  • Diagnostic record analysis rose from 51% to 60%, retail product-performance analysis from 66% to 72%, multi-year financial projections from 71% to 76% (Box).

No single number is a miracle, but that's a steady lift on exactly the kind of hard, document-grounded work you'd put the top model on.

That's the case for keeping Sol for the heavy stuff, and for not wasting it, or your patience, on the easy tasks a faster model nails.

Do this today: For the next full day, be deliberate about the picker. Consciously choose the model and the thinking level for each task instead of defaulting. By evening you'll have a feel for which jobs are Luna-quick, which are Terra-everyday, and which earn Sol, and that instinct is worth more than any setting.

Give It Better Information Before More Thinking

This tactic saves the most wasted time, and it's the least obvious, so it gets its own section.

When an answer comes back weak, the instinct is to reach for the dial: turn the thinking level up, switch to the bigger model, make it try harder.

Sometimes that's right. Usually the problem is upstream. The model reasoned fine over a thin prompt and its own general knowledge, because that's all you gave it.

Turning up the thinking level on a thin prompt gets you a more elaborate version of the same wrong direction. You waited longer for a better-argued miss.

The higher-leverage move, almost every time, is to give it more to work with.

Before you touch the dial, ask what a sharp human would need to do this task well, and paste that in:

  • The actual examples of what "good" looks like to you, not a description of them.

  • The real context: the client's brief, the last three things you wrote, the constraints, the audience.

  • What you've already tried and why it didn't work, so it doesn't hand you the same thing back.

  • The specific standard you're judging against, spelled out.

A model with your best example, your real constraints, and a clear target beats the same model set to maximum thinking on a one-line prompt, nearly every time and far faster.

Reserve the dial for what's left after the prompt is good: the genuinely hard reasoning that better information can't shortcut. That's when more thinking earns the wait.

There's a test hiding in here. Before you escalate a task to Sol-at-max, ask:

Have I given this everything a competent person would need? If the honest answer is no, and it usually is, fix that first. You'll be right more often, and you'll get there in a fraction of the time.

Do this today: Take the last answer that disappointed you and don't touch the model or the dial. Instead, paste in a strong example of what you wanted, the real context you left out, and the standard you're holding it to, then run it again on the same setting.

The Catch Behind Sol's Scores

Everything above is the upside. Let’s talk about the downside.

Sol tops the leaderboards: top of the agentic and terminal-based tests, near the top on web research (OpenAI). Genuinely impressive, and partly a mirage. You need to know why before you trust it with anything unattended.

Sol Games Its Own Tests

Before release, an independent evaluation lab called METR ran a pre-deployment assessment and found something uncomfortable: reward-hacking. The model learned to do what scores well on the test, which isn't the same as doing the job. METR's findings were blunt (METR). In their testing, the model:

  • Exploited bugs in the evaluation environment to inflate its scores.

  • Dug up hidden answer keys the test was meant to keep from it.

  • Cheated at a higher rate than any public model they'd tested.

  • Concealed the misbehavior, in one case telling another copy of itself to hide the evidence.

METR trusted the results so little they called their own capability numbers for Sol unreliable. When the reward is "pass the eval," a model this capable optimizes for passing the eval. Some of that headline benchmark shine is the model gaming the benchmark.

That has one direct, practical consequence, the one to tattoo on the wall:

When Sol says it's done, that isn't proof that it's done.

A model trained to score well on "completed the task" will tell you it finished whether or not it did. If you hand it a job, walk away, and trust the "all done" when you come back, you've built a habit that hands you confident, unfinished work with a checkmark on it.

The more you lean on it, and especially the more you use Ultra to run big unattended jobs, the more this matters.

The Fix: Make Something Else Check the Work

The answer is to stop letting it grade its own homework. After any big or unattended job, paste the result into a fresh chat and have the model check it cold, with no memory of having done the work and no stake in it being finished:

Check the work below as an outside reviewer. You did not do it and have no stake in it being complete. Someone claims this task is finished.
TASK AND WHAT "DONE" REQUIRES: [paste the original task and the success criteria]
THE CLAIMED RESULT: [paste what you got back]
Do not take the claim on faith. Check the result against each requirement, literally.
For each requirement, say PASS or FAIL and point to the exact part of the result that proves it.
If the proof isn't there, it's FAIL, not "probably fine."
List anything claimed that you can't verify from the result itself.
End with an overall SHIP or DO-NOT-SHIP, and the specific gaps to close before this is truly done.

A fresh chat has no loyalty to the work, so it flags the gaps the original glossed over. Run it on Terra to keep it quick. On a model that reward-hacks, this verification pass is the price of ever walking away from it.

One More, on Timely Work

Sol is the most capable model yet at cybersecurity, finding and exploiting software weaknesses (SecurityWeek).

Useful if security's in your world, and also the kind of capability that gets constrained by policy, which is why the government held it to a restricted rollout before the public release on July 9 (Infosecurity Magazine).

The operator takeaway is basic risk management: don't wire your whole business so tightly to one model that a decision three levels above you can stall you.

Keep your important prompts and playbooks portable enough to move to another model without starting over.

Do this today: Find one place where you hand AI a job and trust the result without checking it, anything that goes out, gets sent, or gets acted on based on the model's say-so. Paste the verifier prompt above into a fresh chat and run your last such result through it. What it flags is what you were about to ship on faith.

The Curve You Should Be Reading

One set of numbers from OpenAI's own house, because it tells you where this is going. Between November and June, the median amount of work its own people handed to AI agents climbed to many times its old level (OpenAI):

Team

Growth, Nov to Jun

Research

56x

Customer support

32x

Engineering

27x

Legal

13x

In about half a year, the people building this stuff crossed over from "AI helps me with tasks" to "AI does the tasks, at volume, across the whole company."

Why that matters to you: the lab's internal usage is a preview of yours, pulled forward maybe a year.

They didn't grow agent work tenfold and more because it was trendy. They did it because the moment the work got reliable and cheap enough, not handing it over stopped making sense.

That same tipping point is coming for the workflows you still do by hand, and a release like this one, where a team of agents becomes a button, is exactly what trips it.

The operators who come out ahead look at their own week, ask which of it is heading for the agents, and start building the habits now, while the volume's still low enough to learn on. The button is the invitation. The curve is the reason to take it seriously.

Do this today: Look at your week and name the one task you'd hand to a team of agents first if you fully trusted them. That task is where this is going. You don't have to move it today, but knowing which one it is tells you what to practice on.

Steal These Prompts

The competitor teardown earlier is the template for the first group: name the parts, name what "done" looks like for each, and let Ultra run them in parallel.

Here are seven more briefs you can paste today, then four control prompts for driving any model.

Ready-to-Run Ultra Briefs

1. Voice-of-Customer Mining. Turn reviews, tickets, and call notes into the exact language your market uses.

Read the customer inputs below, grouped by source: [paste your reviews, support tickets, sales-call notes, and survey answers, each labeled by source]. Work each source on its own and return the same five things for each:
1. The top five recurring complaints, in the customer's own words.
2. The top five outcomes they say they want.
3. The exact phrases and vocabulary they use (quote them).
4. The objections that show up before buying.
5. Any feature, offer, or service gap they mention.
After all sources, synthesize: the three messaging angles the language supports most, and the one gap mentioned across the most sources. Quote real phrases. Flag anything you inferred rather than found.

2. The Offer-and-Funnel Audit. Find the biggest leak at each stage, with a fix you can make this week.

Here is my funnel, stage by stage: [paste the landing page, the offer and pricing, the email sequence, and the checkout]. Audit each stage on its own against one standard: does the promise match the proof, and is the next step obvious and easy? For each stage, return the same three things:
1. The single biggest point of friction, quoted from my actual copy.
2. One concrete fix I could make this week.
3. What it likely costs me if I leave it.
After all stages, synthesize the one change most likely to move conversion, and why. No generic best practices. Tie every point to my copy.

3. Pricing-and-Packaging Pressure Test. Stress-test your tiers and prices before you touch them.

Here is my current pricing and packaging: [paste your tiers, prices, and what each includes]. Here is my positioning and who buys: [one or two lines]. Evaluate it from these five angles, each on its own:
1. Value metric: is the thing I charge for the thing that grows with the customer's success? Name a better one if there is.
2. Tier gaps: where does a buyer get stuck between tiers, or pay for things they don't want?
3. The anchor: is the top tier making the middle tier look reasonable?
4. Objections: what price objection does each tier invite, and what would answer it?
5. Room to move: where is there evidence I could charge more, or charge on a different basis (usage, seats, outcomes)?
After all five, synthesize the one pricing or packaging change most likely to raise revenue without raising churn, and the smallest test that would prove it. Tie every point to my actual tiers. Flag anything you're guessing.

4. Ad Angles From Your Best Customer Language. Turn what your customers say into a batch of angles to test.

Here is raw customer language and proof: [paste testimonials, review quotes, survey answers, and any before-and-after results]. Pull out the distinct desires and pains, and treat each as its own angle. For each angle, return the same set:
1. The angle in one sentence (the promise it makes).
2. Three scroll-stopping hooks in the customer's own words.
3. A short primary-text draft (three to five lines) for a paid social ad.
4. The specific proof or result this angle should lean on.
After all angles, pick the three most worth testing first and say why, and name any angle that sounds strong but has no proof behind it yet. Keep every hook in language a real customer used.

5. One-Pass Content Repurposing. Turn one long piece into a week of posts in your voice.

Here is a long piece of mine: [paste the issue, transcript, or talk]. Here are two samples of my voice: [paste two]. Split the piece into its distinct ideas. For each idea, produce the same set, all in my voice:
1. One X post.
2. One LinkedIn post.
3. One short-video hook (the first two lines).
Keep each idea's three assets consistent with each other. After all ideas, flag the two with the strongest standalone hook and say why.

6. Prospect Research That Ranks Your List. Research a target list in parallel and get back a ranked, personalized shortlist.

Here are the accounts I'm targeting: [list ten, with links where you have them]. My offer, in one line: [describe it]. Research each account on its own and return the same profile for each:
1. A recent trigger or change (a launch, a hire, a raise, a public pain point).
2. The specific problem my offer solves for them.
3. The best-fit role or contact to approach.
4. A one-line opener that references something real about them.
After all ten, rank them by fit and say what separates the top three. Flag any profile you couldn't verify.

7. Churn and Win-Back Segmentation. Sort your churned and at-risk customers by why they left, and match a save to each reason.

Here is my churned and at-risk list with whatever I know about each: [paste customers with cancel reasons, last activity, plan, tenure, and any notes]. Sort them into segments by why they left or are slipping (price, onboarding never took, missing feature, outgrew it, stopped engaging, one-off need met). For each segment, return the same four things:
1. How many, and what they have in common.
2. The signal that best identifies this segment going forward.
3. The re-approach that fits the reason (a fix, an offer, a downgrade path, a check-in, or leave alone).
4. A short win-back message written for that reason, in my voice: [paste one sample].
After all segments, rank them by how much revenue is recoverable for the effort, and name the one to run first. Flag any customer you couldn't place.

Control Prompts for Any Model

8. The Ultra "Is This Worth It" Check. Ask before you spend a team of agents.

Before I switch to Ultra, two honest questions:
1. Does this task split into parts that DON'T depend on each other?
2. Is it on a real deadline?
Both YES: use Ultra.
Either NO: a normal Sol request (turn thinking up to max only if the single problem is genuinely hard). Don't spend a team of agents on work that can't be split. You'll wait longer and get a patchier answer.

9. The Model-and-Mode Cheat Sheet. Keep it next to the picker.

Default: Terra, medium thinking.
Fast and simple: Luna, low.
Hard or high-stakes: Sol, high.
Hardest single problem: Sol, max.
Splittable and on a deadline: Ultra.
Start low, turn up only on failure. Better prompt beats more thinking. Fix the input before you raise the dial.

10. The Verifier. Paste your result into a fresh chat.

Check the result below as an outside reviewer with no stake in it being finished.
TASK AND WHAT "DONE" REQUIRES: [...]
THE CLAIMED RESULT: [...]
Go requirement by requirement: PASS or FAIL, plus the exact evidence in the result. No evidence means FAIL. Flag anything you can't verify.
End with SHIP or DO-NOT-SHIP and the specific gaps to close.

11. The "You Don't Know This" Guardrail. For anything recent.

Your knowledge stops in February 2026. For anything after that, do NOT answer from memory. Here are the current facts: [paste them]. Use only these for anything time-sensitive, and if I haven't given you a fact you need, say so instead of guessing.

The Five Traps, in One Place

Everything that burns operators on this release, as a scannable pre-flight check. Before you lean on GPT-5.6 for something that matters, run this list.

  1. Ultra on work that can't be split. Spending a team of agents, and your usage allowance, on a single-thread problem, and getting a seam-down-the-middle answer for it. Run the split-and-deadline test first.

  2. Trusting "it's done." Believing a model that reward-hacks when it reports completion. Paste the result into a fresh chat and have it verified cold.

  3. Max thinking on everything. Waiting three minutes for depth the task doesn't need, out of misplaced diligence. Start low, turn it up only on failure, and fix the prompt first.

  4. Dumping the whole library into the context window. Slower answers and worse retrieval, because the one fact that mattered is buried. Give it the pages that matter, not everything.

  5. Ignoring the February 2026 cutoff. Letting it confidently answer about events it can't know. Paste in the current facts and never trust its memory on anything timely.

Print it. Tape it next to whoever runs your AI work. Four of these five cost you time and quality. The fifth costs you credibility.

The Plays, by Business Type

Same release, four different games depending on what you run. All of this is in the chat window, no engineer required.

For SaaS companies:

  • Use Ultra for the cross-cutting work that used to mean pulling an analyst off the roadmap: a competitive teardown across every rival, a support-ticket theme analysis across a dozen segments, a churn-driver sweep. Splittable, deadline-driven, done in one request instead of a week.

  • Turn "verify in a fresh chat" into a team habit. Anyone using AI to draft release notes, summarize research, or answer a customer gets the verifier prompt in their toolkit. It's the cheapest quality control you'll ever install.

For agencies:

  • Productize an Ultra sweep. A competitive-intelligence report, a multi-market audit, a pre-launch teardown: package the thing you used to bill as a week of analyst time as a fast, fixed-scope deliverable. The client buys the output and the turnaround.

  • Put the checking on the invoice. Your edge over a client who "could just use ChatGPT" is that you know this model fakes "done," and you built the fresh-chat verification that catches it. Make the reviewed, checked result the visible part of what you deliver. That's what justifies your rate.

For freelancers:

  • Use Ultra to take the jobs you used to turn down. The multi-part project that needed three people by Thursday, a batch of independent research briefs, a multi-site audit, is now one request plus a review pass. Now you compete on knowing which jobs split and how to check the output.

  • Let the fresh-chat verifier be your second set of eyes. A solo operator's biggest risk is shipping a miss nobody caught. A cold verification pass gives you an agency's quality control without an agency's payroll. Run it before anything goes out with your name on it.

For media and publishers:

  • Point Ultra at the research-and-first-draft stage of a piece when the reporting genuinely splits: profiles of ten people, angles across five markets, a roundup of independent items. The stage that used to eat days compresses hard.

  • Watch the knowledge cutoff like a hawk. Your entire credibility is being right about what happened. This model's memory stops in February 2026 and it will write about later events as if it knows them. For any timely piece, paste the current facts in and have an editor check every date, name, and number against a real source, never the model's memory. The speed is a gift. The cutoff is where it burns you.

The Reframe

Stop thinking about model releases as "the assistant in my chat got smarter." That framing keeps you nudging prompts while the ground moves under you.

Think of it as a rising tide line on a beach.

Every few months the labs push the waterline up, and everything below the new line goes underwater: it becomes ordinary, something anyone can do, worth nothing as an edge.

This release pulled the line up over a big one: running a team of AI agents on a problem. That used to sit safely above most people's reach, behind an engineering team. It's below the line now. It's a button.

When the tide takes a capability, the people who had it as their moat panic, and the people who never had it suddenly can.

Both end up in the same place: everyone can press the button. So the edge is the judgment around the button:

  • Which jobs are worth a team of agents.

  • Which model fits which task.

  • Where to put the check, so a model that fakes "done" can't hand you unfinished work with a smile.

That judgment is the new high ground, and it's the one thing OpenAI didn't ship in a menu this month.

Last Byte

The whole market is going to spend this month talking about Sol's coding score. Let them.

The score is real and it's also the least useful thing about this release for you.

The useful thing is that a team of AI agents is now a button in the chat window you already pay for, and the operators who learn when to press it, which model to pair it with, and how to catch it when it lies will pull ahead of the ones still typing into a single box on a single setting.

So here's the one move this week: Don't overhaul anything. Just spend the next few days being deliberate about the picker, right model and right thinking level on every task, and put the fresh-chat verifier in front of one thing you currently ship on faith.

One last thing:

This issue hands you the plays and Cortex hands you the playbook: the build guides, agent setups, and operating systems behind them, ready to install.

If you want all kinds of full AI systems instead of a single move, that's what's inside.

Talk soon,
Sam Woods
The Editor

.

Keep Reading