ROB RUSH
Case studies
Recorded agent run

Synthetic data · replayed runs

An AI account advisor for customer success.

Retention Agent

Reviews account usage and support signals, recommends a next step, and drafts outreach for human approval.

  • 40recorded runs
  • 4models
  • 6tools
  • scored on right play, facts traced, follows playbook
AI / AgentsAnalyticsInteraction design
Step through a recorded run

Recorded runs

Which customers might not renew, and what should we do about each one?

That’s the question an account manager faces every week. I built an AI agent that does the prep: it checks the customer’s numbers, reads the company playbook, recommends one action, and drafts the email. A person decides. I ran it on ten example customers with four different AI models, and recorded every step so you can watch.

Explore recorded agent runs on synthetic customer data and inspect the supporting evidence.

How one run works

  1. Check the customerUsage, unused seats, support tickets
  2. Read the playbookThe rules for each situation
  3. Recommend one actionOut of five possible plays
  4. Draft the emailUsing only real numbers
  5. A person decidesApprove or hold. Nothing is sent.

What it does for one customer

  • 19secfrom customer name to a drafted email (median run)
  • <1¢per customer with the best model
  • 10/10right action, best model
  • 94%of numbers in the drafted emails matched the data

Measured across the 40 recorded runs. A person still reviews every draft.

Pick a model and a customer

All ten customers are invented. The numbers are synthetic.

  1. medium risk
    Show the data

    Fields

    segmentSMB
    planTeam
    seats_purchased40
    arr8,191.64
    renewal_date2027-01-04
    days_to_renewal111
    risk_score46.2
    risk_bandmedium
    top_driversusage_decline, idle_seats
  2. −32% usage
    Show the data

    Summary

    weeks_of_history12
    prior_weeks8
    recent_weeks4
    prior_avg_active_seats37
    recent_avg_active_seats25.2
    usage_drop_pct31.8
    prior_avg_sessions127.8
    recent_avg_sessions89
    sessions_drop_pct30.3

    Week by week

    week_endingperiodactive_seatssessionsfeature_events
    2026-06-30prior381241,138
    2026-07-07prior36117987
    2026-07-14prior371251,130
    2026-07-21prior361371,298
    2026-07-28prior381261,199
    2026-08-04prior371301,062
    2026-08-11prior381281,199
    2026-08-18prior361351,165
    2026-08-25recent2588698
    2026-09-01recent2696864
    2026-09-08recent2591843
    2026-09-15recent2581671
  3. 38% idle
    Show the data

    Fields

    seats_purchased40
    recent_weeks4
    recent_avg_active_seats25.2
    idle_seats15
    idle_seat_pct37.5
    arr8,191.64
    arr_per_seat204.79
    idle_seat_arr3,071.86
  4. support healthy
    Show the data

    Fields

    open_tickets4
    escalations0
    csat4.2
    recent_ticket_topicsnotification settings, mobile sync delays, billing question
  5. 2 lookups
    Show the data

    Lookup 1: “adoption workshop 30% active seats drop guidance outreach”

    Adoption workshopPlay type: adoption_workshop. Use it when the account's previously active seats have gone quiet: a usage drop of 30% or more, measured as the average active seats over the last 4 weeks against the average of the prior 8 weeks. The people who bought the product are still there; they have simply stopped opening it, often after a team change, a busy quarter, or a workflow that moved elsewhere. A workshop for those existing users is the answer, not a commercial change.
    Seat rightsizingPlay type: seat_rightsizing. Use it when usage is steady but the account is paying for far more seats than it uses: at least 40% of purchased seats have been idle across the last 4 weeks, and active seats have not dropped sharply (a drop under 30% versus the prior 8 weeks). The account is not unhappy; it is over-provisioned, and an honest conversation now beats a surprise cancellation at renewal.
    Adoption workshopDo not mention seat counts, pricing, or the renewal date in the first message. Do not name an individual; write to the whole team. Keep the draft to two or three short paragraphs and one concrete ask: a date for the session.

    Lookup 2: “adoption workshop message structure session length what to offer”

    Adoption workshopSuggested subject: "A short working session for the {account} team".
    Adoption workshopDo not mention seat counts, pricing, or the renewal date in the first message. Do not name an individual; write to the whole team. Keep the draft to two or three short paragraphs and one concrete ask: a date for the session.
    Adoption workshopThe offer is a short, hands-on working session for the existing users, not a sales call. Propose a 45-minute workshop built around the two or three features the account used most when usage was healthy, and ask which workflow fell away. Quote the actual weekly numbers from the usage tool so the team can see the same picture you see.
  6. matches the expected play
    Show the data

    Cited numbers

    claimfieldvalue
    active seats dropped from 37 to 25prior_avg_active_seats37
    active seats dropped from 37 to 25recent_avg_active_seats25.2
    weekly sessions fell from around 128 to 89prior_avg_sessions127.8
    weekly sessions fell from around 128 to 89recent_avg_sessions89

The proposal

Adoption workshop

right action

Why: Usage down 32%, 38% of seats idle and support healthy.

To: the Dunmoor Legal Logistics team

From: Your account manager

Subject: A short working session for the Dunmoor Legal Logistics team

We've noticed that active seats dropped from 37 to 25 over the past month, and weekly sessions fell from around 128 to 89. Your team was getting great value earlier this summer, and we'd like to help you get back to that momentum.

We'd like to offer a 45-minute hands-on workshop focused on the features your team used most when engagement was strong. This isn't a sales call—it's a working session to understand which workflow may have shifted and how we can support your team's day-to-day needs.

Can we schedule a session in the next two weeks?

Highlighted numbers were checked against the data the agent looked up. 4 of 4 check out.

Your choice stays on this page. Nothing is sent.

Same ten customers, four AI models

Which model would I use? GPT-6 Luna.

It picked the right action for all ten customers, in about 18 seconds and a tenth of a cent each. The most expensive model wasn’t the most accurate.

  1. GPT-6 Luna (the model I would use)10/10 right actions0.1¢ · 18 s per customer
  2. Claude Sonnet 4.59/10 right actions9¢ · 33 s per customer
  3. GPT-4o mini8/10 right actions0.1¢ · 10 s per customer
  4. Claude Haiku 4.58/10 right actions3¢ · 19 s per customer

Each square is one customer: filled means the model picked the action the company’s own rules call for. Cost and time are per customer.

Problem

Customer success teams need to identify accounts that may not renew and decide which signals deserve a conversation.

Challenge

Let a model decide which account signals to check and which play to propose, then show whether its draft sticks to the data it retrieved, without treating risk indicators as proven churn predictions.

My role

I built the recorder, the six tools, the five playbooks, the evals and this replay.

Architecture

  1. 1Synthetic dataset: 60 accounts with weekly usage, seats, support tickets and renewal dates
  2. 2Python recorder calling each model through OpenRouter with OpenAI-compatible tool calling
  3. 3Six tools the model chooses between: account_risk, usage_drop, idle_seats, support_summary, retrieve_docs (BM25 over five playbooks) and propose_play with cited facts
  4. 4Four models run on the same ten accounts; every run is validated against a JSON schema
  5. 5Evals on each run: play match, facts traceable, follows playbook
  6. 6This page replays the recorded runs and ends at an approve-or-hold step; nothing is sent

Technologies

PythonTool callingOpenRouterBM25 retrievalpytestJSON SchemaNext.js replayTypeScriptZod

Approach

  • Give the model narrow tools and let it choose the order; the recorder keeps every call, its arguments, the result and the latency.
  • Ask for a source on every number, percentage and date in the draft, then check each one against the recorded tool result.
  • Record every model against the same accounts and seed, and replay the runs here without calling a model.
  • Stop at a proposal: approving or holding it in the replay changes nothing outside this browser tab.

Current result

40 recorded runs across four models and ten accounts, each replayed above call by call with its draft and evals. Play match against the dataset’s expected play: Claude Haiku 4.5, 8 of 10; Claude Sonnet 4.5, 9 of 10; GPT-4o mini, 8 of 10; GPT-6 Luna, 10 of 10.

Lesson

Recording real runs shows what a scripted demo hides: runs that stalled because my recorder cut playbook passages short and the models kept searching for the rest, and drafts whose numbers sound right but do not match the data. Checking each cited figure against the tool result is what makes a draft quick to review.

Scripted walkthrough

Try the customer-success review.

Use the account snapshot to understand the warning signs, select an outreach tone, and make a simulated review decision.

This section is hand-authored: one fixed synthetic account, a prewritten draft and its tone variants, and approve, redraft, or reset controls. No model produced it; the recorded runs above are the model's own work.

1 / Meet the customer

VoltLogistics

A business that pays for your software so 25 employees can use it. Its subscription is about to renew, but the team is using the product less.

Renewal is in
14 days
September 29, 2026
Annual subscription
$3,065
For this one customer

The example takes place on September 15, 2026. All dates and figures refer to that snapshot.

2 / Understand the warning signs

Are they getting enough value to stay?

These signals suggest a conversation is worth having. They do not prove that the customer will cancel.

Fewer people are using the software

47% less daily usage

Ask what changed. Has the team hit a problem, changed its workflow, or stopped finding the product useful?

Almost half the licenses are inactive

12 of 25

Paid user licenses marked inactive or stale

The company is paying for access that much of its team may not be using. Find out whether training or a better fit would help.

The team has asked for support

3 support requests

In the past 30 days; 2 marked high or critical priority

Check whether those issues were resolved before suggesting another training session. The example does not include their current status.

3 / Review the suggested response

Offer a check-in before renewal.

Start by understanding the decline. Ask about the team’s experience, review its support issues, and offer help getting value from the software.

The next move is yours

Same facts.
A different way to connect.

Approve the suggested email, or reshape how it sounds. A helpful response starts with the right tone.

You’re reviewing a prewritten example email.

Draft previewOriginal draft

To VoltLogistics team

A quick check-in before your September 29 renewal

Hi VoltLogistics team,

With your renewal coming up on September 29, I’d like to check how the software is working for your team.

We’ve noticed lower usage recently. Has anything changed in your workflow, or is something making the product difficult to use? I’d also like to make sure your recent support requests have been addressed.

Would a 20-minute check-in be useful? We can review any obstacles and identify where your team could use more help.

Best,
Your account manager

Practice only. Approving does not send this email or change customer records.
How this example works — definitions and technical details

One synthetic account and prewritten email alternatives walk through the sequence: inspect customer signals, choose a tone, and let a person review the response. Approving does not send a message, and nothing here predicts a chance of cancellation.

Daily usage
Average daily active users: 14.83 in the period 60–90 days before the snapshot, compared with 7.87 in the most recent 30 days. Headline counts are rounded.
Inactive licenses
12 of 25 licenses are flagged inactive or stale in the example. The scenario does not define an inactivity threshold. This measure differs from daily active users.
Source and scope
One account, ACC-00020, dated 2026-09-15. The case-study charts summarize the wider portfolio, not this customer’s individual results.
What approval changes
Only the on-screen review decision. It resets when this page reloads. Execution remains disabled for every choice.

Synthetic portfolio snapshot

The 60-account portfolio behind the walkthrough.

The walkthrough’s example account, ACC-00020, sits in this synthetic portfolio, shown as of September 15, 2026. The recorded runs above draw on a separate synthetic dataset from my recorder. None of these figures are customer results.

Accounts in portfolio
60
Annual recurring revenue
$1.84M
Accounts flagged at risk
15
ARR flagged at risk
$319K
Risk bands by segment
Risk bands by segment. Account counts grouped by segment and risk band.
Usage trends
Usage trends. Seat utilization for the at-risk and healthy cohorts.
Accounts with idle seats
Accounts with idle seats. Ten accounts ranked by idle seat rate, from 61% down to 42%. The example account, ACC-00020, sits mid-pack at 48%: 12 idle or stale seats out of 25.
Renewal calendar
Renewal calendar. Monthly renewal counts from September 2026 through March 2027.
Contact

Questions about how this was built?

Send a note about this project or the approach behind it.

Contact Rob