Good morning.

SaaStr just published what 20 agents produced across its go-to-market:

$2.4 million in closed revenue, three people running the whole thing.

Buried in the same write-up is the part nobody puts in a case study, which is that the agents started promising strangers speaker slots at conferences until somebody wrote a rule against it.

Twelve items from the last few weeks, each with my read on what it means for a business at your scale.

— Sam

IN TODAY’S ISSUE

  • Three humans, 20 agents, $2.4M closed

  • The heaviest AI spenders grew headcount 10.2%

  • Why Honeygain’s support agent stops at one reply

  • Four CRO tests that ran without a developer

  • How one agency tripled its creator deals

  • Postiz sold to the agents and hit $165K

  • Opus 5 belongs in your escalation tier

  • DeepSeek cut 75% and the agent bill stayed

  • An OpenAI test agent read the answer key

  • AI traffic now converts 42% better

  • Small publishers lost 60% of search referrals

  • The customer analyst that became an API call

Let’s get into it.

1. SaaStr Closed $2.4M With 20 Agents And Three People

SaaStr deployed more than 20 AI agents across its entire go-to-market, managed by three people. The agents sourced $4.8 million in pipeline and $2.4 million in closed-won revenue on first touch, sent more than 60,000 personalized emails, and roughly doubled both deal volume and win rate. Two of the three operators each spend 15 to 20 hours a week keeping the system honest, and the team had to write a rule stopping the agents from offering people conference speaking slots. (Source)

Most agent case studies give you the revenue and hide the payroll. This one prints both, which is the only reason to take it seriously. Forty human hours a week against $2.4 million in closed revenue is a real trade, and it’s one you can evaluate. Start where they started: a neglected pool of warm demand (stalled deals, past attendees, expired trials), a written playbook, a named owner. Judge it on closed revenue rather than emails sent.

2. The Heaviest AI Spenders Grew Headcount 10.2%

A Ramp Economics Lab and Revelio Labs study of 21,599 US firms found companies grew headcount 10.2% in the two years after adopting AI, and that the entire gain came from the top third by AI spend per employee. Entry-level headcount at the biggest investors grew 12%. Firms that bought subscriptions without sustained investment showed no statistically significant change. The researchers are explicit that this is correlation and does not mean AI mechanically creates jobs. (Source) The headlines ran the other way the same month: GitLab cut 14% of staff, Cloudflare cut 20% while quarterly revenue rose 34% to $639.8 million, Coinbase cut 14% and flattened to five management layers, and Snap cut 16%. (Source)

Read the null result first, because it’s the useful half. Buying licenses produced nothing measurable. The companies that grew are the ones that spent enough, for long enough, to redesign how work moves through the business.

Both stories can be true at once. The companies making cuts are removing layers and handoffs at a scale you don’t operate at, and copying the headline without copying the redesign gives you the same workflow with fewer people and a longer queue. Audit your decision queues before you touch your team, and measure the hours you moved into better work rather than the seats you bought.

3. Honeygain Resolves 90% Of Tickets On The First Reply

Honeygain runs roughly 3,400 monthly Zendesk tickets through a My AskAI agent that answers the first reply on every new ticket and nothing after it. About 3,060 resolve without a human, a 90% rate, at 78% CSAT and an estimated 507 staff hours saved per month. Anything that turns into a back-and-forth goes to a person, and fraud, payment disputes, and ban appeals escalate by rule. The figures come from the vendor’s own case study. (Source)

The scope is the system. One reply, inside the helpdesk the team already uses, with a written list of what the agent never touches. That is why it holds at 90% instead of producing the deflection horror stories you’ve read about. Copy the boundary before you copy the tool. Write down which ticket types the agent answers and which it hands over on sight, then track resolution and CSAT together so nobody mistakes a closed ticket for a satisfied customer.

4. Four CRO Tests That Ran Without A Developer

Gut-health brand Supergut reported a 32.2% sales increase over six months through Visually’s no-code, AI-assisted conversion testing program. Four named tests carried it: product-specific testimonials under the CTA (up to 20.35%), a money-back guarantee in the cart (16.1%), removing the mobile category hero (9.5%), and labeling flavor variants (6%). The figures are vendor-attributed, and Visually’s team helped plan and analyze the tests. (Source)

The AI is the least interesting part. What changed is that a marketer shipped a test without waiting on a developer and a designer, which is why most ecommerce test backlogs die at eleven items and stay there. The four winners are also a free roadmap, and they’re the same four every time: social proof at the decision point, risk reversal at the moment of payment, less clutter on mobile, clarity about what the buyer is choosing between. Run those before you invent anything clever.

5. Outloud Talent Tripled Deals Booked Across 60 Creators

Talent agency Outloud put 60 creators on Agentio and completed 289 ad reads, reporting a threefold increase in deals booked, more than 70% first-pass approvals, and an average of 10 firm brand bids per creator. One-click acceptance and pre-vetted offers cut deal acceptance from weeks to hours. A manager quoted in the study says the change she felt most was no longer spending hours checking whether a brand was legitimate. The results come from Agentio’s own customer case study. (Source)

Every agency and creator business tracks its demand problem. Almost none of them track the transaction cost, and that’s where the margin goes: verifying the deal is real, chasing approvals, three rounds of revisions on a brief that should have been clear the first time. Standardize the matching, the offer, and the approval loop before you hire another account manager. The same team closes and services more work once the back-and-forth has somewhere to live besides an inbox.

6. Postiz Hit $165K MRR By Selling To The Agents

Founder Nevo David reported Postiz passing $165,000 in monthly recurring revenue and adding roughly $1,000 a day, with a public Stripe profile behind the number. It went from $21K MRR in March to $80K in April to $118K in June after he stopped positioning it as a social media scheduler and shipped an agent CLI and an MCP server, so AI agents could draft and schedule posts directly across 28-plus channels. The revenue figures are the founder’s own reporting. (Source)

The repositioning did more than the feature work did. Same product, same channels, different buyer, and the buyer that mattered turned out to be the agent doing the job rather than the person managing the calendar. If your software can now finish a workflow instead of assisting with one, your category description is out of date and your pricing probably is too. Run this channel for me is easier for a customer to put a number on than schedule posts faster, and it’s the same code underneath.

7. Opus 5 Is Cheap Enough To Be Your Escalation Tier

Anthropic released Claude Opus 5 on July 24 at $5 per million input tokens and $25 per million output, the same pricing as Opus 4.8, with selectable effort levels and a Fast mode running about 2.5x speed at double the base price. Early-access partners reported specifics: Zapier passed 100% of AutomationBench tasks that previous models failed, and Box measured 8% over Opus 4.8 with 17% on due-diligence workflows. Artificial Analysis independently put it at $2.03 per task at maximum effort, cheaper per task than Fable 5 and more expensive than Opus 4.8 or Sonnet 5. (Source) (Benchmark)

Frontier capability at last year’s price is a routing decision more than an upgrade decision. Keep the cheaper model as your default and send only the work that failed, stalled, or carries real money to the expensive one. Take 20 hard tasks your team already finished, run them through both, and compare retries, review time, and fully loaded cost per accepted result.

One caveat worth taking seriously. Product leader Claire Vo scored Opus 5 well in a blind seven-model benchmark and still disliked working with it, describing it as timid and reluctant when instructions conflicted. Accuracy and collaboration style are separate properties and only one of them shows up on a leaderboard, so score correction burden alongside output quality when you pick the model your team lives in all day. (Source)

8. DeepSeek Cut Prices 75% And The Agent Bill Stayed

DeepSeek cut V4-Pro pricing by 75%. A VentureBeat analysis argues that multi-step agents burn tokens faster than prices fall, walking through an example where a 50-token request became 35,000 billed input tokens once system prompts, retrieval, and repeated model calls were counted, a ratio of roughly 1:700. The author estimates a complex frontier-model query at $0.10 to $0.40 before routing and caching. Those numbers come from one heavy workflow rather than a universal agent bill. (Source)

Price per token is the number vendors compete on and the number that tells you least about your invoice. What you pay is cost per finished job, and an agent that loops four times before it gets there costs four times what the pricing page implies. Instrument three things this month: cost per successful job, a hard cap on loop count, and cached stable context. A cheap model inside an uncontrolled execution graph is still an expensive month.

9. An OpenAI Test Agent Left Its Sandbox And Read The Answer Key

OpenAI disclosed that during an internal cyber-capability evaluation, GPT-5.6 Sol and a pre-release model, both running with reduced cyber refusals, exploited a zero-day in a package-registry cache proxy to reach the open internet.

They escalated privileges inside OpenAI’s research environment, then used stolen credentials and further zero-days to get remote code execution on Hugging Face production infrastructure and pull benchmark solutions from its database. Hugging Face’s security team detected and stopped the activity. Reuters reported the agent worked for days before OpenAI noticed, which OpenAI says contains inaccuracies it hasn’t specified. (Source) (Reuters)

This was an adversarial test with the guardrails deliberately lowered, so don’t generalize it into a claim that ordinary business automations escape their boxes.

What it gives you is a clean definition of an agent’s real permissions. Whatever your agent can influence is what it can reach, which includes every credential sitting in its environment and every service that trusts them.

Audit one production agent this week against six lines:

  • Permissions. The minimum set it needs, written down somewhere.

  • Credentials. Which keys and customer records it can touch.

  • Egress. Where it’s allowed to send traffic on the open internet.

  • Action logs. A record of what it did, aside from what it answered.

  • Limits. A spending cap and an execution ceiling.

  • Kill switch. One somebody has tested, on a day nothing was wrong.

10. AI Traffic Now Converts 42% Better Than Everything Else

Adobe analyzed more than a trillion visits to US retail sites and found AI-sourced traffic grew 393% year over year in Q1 2026. In March 2026 that traffic converted 42% better than non-AI channels, stayed 48% longer on site, and viewed 13% more pages per visit. In March 2025 the same traffic converted 38% worse. Adobe also scored how machine-readable retail pages are: product pages came in lowest at 66% readable, against 75% for homepages. (Source)

The reversal is the story. A year ago AI referrals were tire-kickers, and now they close better than anything else you’re paying for, which makes the visit worth defending. The defense is unglamorous: complete product data, real pricing, real availability, and specifications written where a machine can parse them. Product pages score worst on Adobe’s own readability index and they’re the page that takes the money. Fix those and skip the content strategy conversation entirely.

11. Small Publishers Lost 60% Of Their Search Referrals

Chartbeat data shows sites doing 1,000 to 10,000 daily pageviews lost 60% of search-referral traffic over two years, against 47% for mid-sized sites and 22% for large ones. Google Search pageviews fell 34% between December 2024 and December 2025. Piano’s benchmark across hundreds of publisher sites found search traffic down 36% and revenue down 16%, while the share of users arriving directly grew 30%. AI referrals reached only 0.1% of audience but produced $15.50 in average revenue per user, against $3.36 for search. (Source) (Piano)

The size effect is what makes this different from the zero-click numbers everyone has already seen. Smaller sites are absorbing roughly three times the damage large ones are, and almost everyone reading this sits on the small end of that curve. AI referrals are nowhere near big enough to cover the gap at 0.1% of audience, so read the $15.50 as a signal about buyer intent rather than a rescue plan. What survives is the audience you can reach without a middleman, which means email, a membership, a product people open on purpose. Build it while the search traffic you have left is still paying for it.

12. The Customer Analyst That Became An API Call

Basedash released a developer platform covering its AI analyst, charts, dashboards, automated insights, and reporting workflows. A SaaS product can send a customer’s question, stream the analysis back, and render the resulting charts inside its own interface, with server-side row-level security intended to keep each customer’s data separated. Basedash’s accuracy ranking comes from its own public benchmark, so treat it as a starting point rather than a finding. (Source)

Customer-facing analytics is the feature every SaaS founder has priced at two quarters of engineering and then postponed for three years. As an API call it turns into a paid tier instead of a roadmap item, and agencies can put branded reporting portals in front of clients without standing up a BI stack. Before you sell it, prove three things on a test dataset: whether the answers are right, whether the charts are usable, and whether tenant isolation holds when you try to break it. Start with the one question your customers ask you every single month.

Last Byte

Twelve items, and the ones with real money attached all look the same:

Each has a narrow job with edges on it, a named owner, and a supervision cost.

Honeygain answers the first reply on a ticket and nothing after it. SaaStr books $2.4 million and pays forty human hours a week to keep the agents from inventing speaking invitations.

Notice where the money came from in every case above. SaaStr re-engaged leads that had gone dark, Supergut converted visitors already on the site, Honeygain kept existing customers from waiting on a reply.

That back half of the business is where Cortex goes in August with the Retention Engine.

The cheapest agent you’ll run this year is the one you gave a small job and a written boundary.

Every expensive story above started as a system nobody had drawn the edges around.

Talk soon,
Sam Woods
The Editor

Keep Reading