
Good morning.
There's a version of this work that feels like progress and produces nothing. You compare platforms, read the launch threads, open one more free trial, and six months later you're paying for twelve logins with nothing running.
Today is the short version of how to stop: which two apps are enough to start, what still holds twelve months from now, and the rule that decides which model does which job.
—Sam
In Today's Issue
Why four entrepreneurs misdiagnosed the same problem
Two platforms now cover what five used to
What still holds twelve months from now
When a dedicated agent builder earns its slot
How to route work by what a mistake costs
Your Sonnet 5 bill jumps 50% on September 1
The fifteen-minute move that costs nothing

Four Descriptions Of The Same Problem
I sat in a room of entrepreneurs at a workshop and four of them described the same problem in four different ways, none of them noticing it was the same problem.
One had GPT workflows for research, drafting, and client reporting. They ran fine in a demo, but under a real deadline he was copying output from one into the next by hand, so the whole thing moved at the speed of him sitting at his desk. Another had built north of fifty agents and couldn't tell me which ones produced anything a client paid for. A third had spent three years evaluating platforms without putting one into production.
Every one of them opened by calling it a tool problem, and every one of them was wrong about that.
Comparing tools is the only part of this work you can do forever without committing to anything. It feels like momentum, and nothing breaks while you're doing it. Building means picking, being wrong where your clients can see it, and fixing it on a Tuesday.
The fear underneath is reasonable. One reader put it better than I could: "I don't know which AI niche will be most stable over the next year or two, don't want to pick something that means I have to keep ditching what I've just spent 6 months learning."
That's the real question, and the answer is that the six months never goes into the tool. It goes into the written-down version of how your business decides things: how you qualify a lead, what your copy is allowed to sound like, the number of days without a reply that means a deal has gone cold.
None of that expires when a model ships. You paste it into whatever you're using next and it works on day one. The full structure for holding it is in The Cognitive File System.
Two Platforms Are Enough To Start
For years the honest answer to "which platform" was that you'd need a few, because the chat window couldn't reach your systems. That stopped being true last year and is even more true this year.
Claude and ChatGPT both speak MCP now, the standard that lets an app connect to the software your business already runs on.
Cowork draws on Anthropic's connectors directory, which runs to hundreds of tools with Slack, Notion, Google Drive, HubSpot, and Salesforce among them.
ChatGPT covers the same ground through its apps, plus custom connectors for anything with a hosted endpoint once you turn on Developer Mode, which is paid-plan only and still in beta (OpenAI).
So the two paths are Claude Cowork or Claude Code, and ChatGPT Work or Codex. Pick one, and run it for ninety days before you consider the other.
Which one matters less than people want it to. If you already work in Claude every day, stay there; if your team runs on Microsoft 365 and Teams, ChatGPT will feel closer to home. What you build sits on top of either.
What’s Still True Twelve Months From Now
Ninety days is enough to commit because most of what you learn inside them transfers to whatever comes next.
What holds:
Your written-down decisions. They outlive every model release, which is the argument above.
MCP as the connection standard. Both major platforms adopted it, so the connectors you authorize this month keep working.
Retrieval over your own documents. Table stakes now, not a capability worth shopping for.
Describing a job precisely. The same skill whether the thing doing the job is a model or a new hire.
What you can skip:
Benchmark leaderboards. They move weekly and none of the movement changes what you should build.
Context-window numbers. They stopped being the binding constraint for the work most entrepreneurs run.
Launch-day threads. They tell you what changed for engineers, rarely what changed for your P&L.
Where The Other Platforms Fit
There's a whole category of dedicated agent builders now, and they're good. Manus and Genspark for general agent work, Spine for research that comes back as a finished deliverable, Relevance AI for building specialist agents, n8n for workflow automation, Viktor and Cyndra for the AI-employee-in-Slack shape. Cyndra launched at the end of June, which tells you the pace: a serious new one arrives most months.
Nearly all of them do a version of what your existing subscription now does.
The churn runs the other direction too. Relay, a workflow builder plenty of people were recommending this spring, told its customers this month that it's closing. Anyone who put client delivery on it spends August migrating instead of selling.
The reason to add one is a job you've already hit the wall on, which would be something that has to run without you in the room, on a schedule or off a trigger, across enough steps that a chat window becomes the wrong shape for it. That's where n8n or a dedicated builder earns its slot.
Until you've hit that wall, adding one is shopping.
Most entrepreneurs I talk to sit two connectors and one written-down process away from having something run, and they spend the week on a comparison page instead.
Match The Model To What A Mistake Costs
Once you've picked a platform, one decision inside it still moves both your bill and your output quality:
Which model gets which job.
The old answer was to send everything to the best one, and that stopped being a single instruction this summer.
Anthropic now ships two models above its mid tier: Fable 5 at $10 per million input tokens and $50 output, and Opus 5, which arrived July 24 at $5 and $25, the same rate card as the Opus generation before it. Anthropic's own numbers put Opus 5 within half a percent of Fable 5 on a coding benchmark at half the price.
OpenAI runs three rungs, GPT-5.6 Sol down to Luna, where the frontier costs five times the cheap rung.
So "use the best one" now means choosing between two frontier prices at one vendor and three tiers at the other. The rule that replaces it: route by what a wrong answer costs you.
Three buckets cover most of a business.
High volume, low stakes. Tagging inbound leads, summarizing call notes, first-pass drafts nobody sends. Cheap rung. A wrong answer costs you thirty seconds.
Work a client sees. Proposals, deliverables, anything with your name on it. Mid rung, with you reading it before it goes out.
Decisions with money on them. Pricing, offer positioning, anything where being wrong costs you a deal. Top rung, every time.
Score the buckets on retries and review time rather than on benchmarks. If the cheap model makes you rewrite the prompt twice and then edit for ten minutes, it cost you more than the expensive one. You can run that test across twenty representative jobs in an afternoon.
Two levers beat model choice, and most entrepreneurs never touch them:
Prompt caching bills a cache hit at roughly a tenth of base input price, so a stable system prompt you reuse all day is cheaper to cache than to downgrade.
Effort settings are the second: Opus 5 exposes low through max and neither vendor charges extra for the setting, so start at the default and step down where quality holds.
One dated item for your calendar: Claude Sonnet 5 sits on introductory pricing at $2 and $10 through August 31, then goes to $3 and $15 on September 1, a 50% increase.
Sonnet 5 also uses a new tokenizer that emits about 30% more tokens for the same text, so sticker parity with the previous generation isn't bill parity. Thresholds you set during the intro window are wrong in five weeks.
The Protocol
Step 1. Add nothing. No new subscription and no new trial for the length of this exercise.
Step 2. Write down one decision. Pick the kind you'd have to explain if you hired someone next month, like which leads are worth a call or when a deal is dead. Write the rule, what you look at to make the call, and where that information lives. Fifteen minutes, no subscription, and it carries into whichever app you choose.
Step 3. Pick the platform and authorize two connectors. Connect the two systems your written-down decision reads from, usually your CRM and wherever your documents live.
Step 4. Sort your recurring AI work into the three buckets. Most entrepreneurs find the high-volume bucket is bigger than they expected and running on the top rung.
Step 5. Run the retry test. Take twenty jobs from that bucket, run them on the cheap rung, and record how many needed a second attempt and how many minutes you spent editing. If both come back low, the bucket stays cheap permanently.
Step 6. Put August 28 in your calendar. Sonnet 5 reprices four days later. If that bucket runs on it, re-run Step 5 against the cheap rung before the new rate hits.
Go ahead, get going, and go win.

Last Byte
The entrepreneur who spent three years evaluating platforms is still evaluating.
Somebody who committed to a mediocre tool in his first week and spent those three years writing down how his business decides things has three years of compounding behind him, and he can switch platforms in an afternoon.
If you'd rather not start that from a blank file, it's what I packaged as Bionic OS:
Your decision principles, your voice profile, and the constraints you work inside, with sixty skills and ten agents sitting on top.
It's a folder on your disk, so it runs in Claude or ChatGPT and moves to whatever comes after them.
Relay's paying customers have until September 14. Whatever they wrote down goes with them but whatever they built inside the product doesn't. This is true for any third party platforms, and it’s why you must have control over your own foundation and context layer—which can be replicated inside any platform on demand.
Bionic OS lays the foundation for a system that’s yours and that you can take anywhere and implement inside any agent tool or platform of your choice.
Talk soon,
Sam Woods
The Editor
.

