Jim Boothjimbooth.ai

Case Study ยท Autonomous Agents in Production

A multi-agent system that runs a business unattended.

Not a demo. A production system that sources, evaluates, writes, verifies, and packages real work around the clock, with a watchdog watching the agents, human approval gates on anything that touches money, and a testing agent that exercises the whole pipeline in production every day.

Most AI-agent projects work in a demo and fall apart the moment they have to run without a person watching. The gap between an agent that works once and an agent system you can leave alone for a week while it handles real transactions is where nearly every initiative stalls. This is a walk through a system I designed, built, and operate that lives on the far side of that gap.

24/7
Unattended, 10-minute work cadence
3 tiers
Autonomous, scheduled, human-in-the-loop
7-door
Daily production test battery, agent-driven
~10k
Sources swept per sourcing run

The context

The business is a done-for-you service: a customer provides a goal, and the system produces finished, ready-to-use work products tailored to it. Every step a human would normally grind through, the research, the matching, the writing, the fact-checking, the assembly, is handled by agents. The interesting part for a technical buyer is not the domain. It is the operating model: the work is produced by autonomous agents, and a human only enters at the two points where judgment is legally or reputationally load-bearing.

The architecture

The agent runs the entire pipeline autonomously, and a watchdog keeps it alive โ€” but the two irreversible, consequence-bearing actions sit on the far side of a hard human gate the agent is forbidden to cross. That boundary, plus a test agent that exercises production every day, is what turns “an agent that works” into a system you can leave running.

The fulfillment loop Fully autonomous

The core agent. Every 10 minutes it walks each open job through a state machine.

A reasoning agent onboards the input, builds a search profile, sources live candidate material, scores and evaluates each option against the goal, verifies every external link is genuinely live (an API check first, visual confirmation as a fallback), generates the tailored work products, and assembles the package. These are judgment calls, which options fit, which to discard, how to phrase the output, made by the agent rather than branches in a script. Critically, the loop is forbidden by design from the two irreversible actions: it cannot send anything to a customer and it cannot mark work delivered. Those require a human.

The watchdog Reliability layer

A supervisor process checks every 3 minutes whether the main loop is alive and making progress. If the loop dies, hangs, or the machine's state drifts, it restarts it. This is the layer most agent projects never build, and the reason unattended is a real claim here rather than an aspiration.

Self-refilling demand Scheduled

Cloud cron jobs generate the day's work queue automatically and manage time-sensitive follow-ups, so the autonomous loop always has something to do without a human loading the queue.

The daily production test agent Verification

Once a day, an agent drives a real browser through every path a customer can take, end to end, against the live production system, including deliberately forced failure cases (blocked scripts, server errors, broken states) to prove the system fails safely and never strands a user. It cleans up after itself and reports a pass/fail table. The system tests itself in production, every day, without me.

Human-in-the-loop assistants Judgment gates

A separate set of agents draft outbound communication and classify inbound replies, but never send. They prepare; the human approves and sends. Knowing where to place the human is as much of the engineering as the automation itself.

What made it hard, and what that proves

The demo version of any of this takes an afternoon. The production version took the reliability work that demos never show:

The rare skill is not wiring a model to some tools. It is operating an agent system in production, with real consequences, and knowing every way it breaks before it breaks.

What this is for

Companies right now are standing up agent initiatives, hitting the wall between a working demo and a system that survives production, and stalling there. I have already paid that tuition, with real transactions on the line, and I help teams get their agent systems across the same gap: from works-in-the-demo to runs-without-a-babysitter.

This is an engineering case study of a system I designed, built, and operate. It describes the architecture and the reliability work, not a revenue claim. If your team is stuck at the demo-to-production wall with an agent project, that is the problem I work on. — Jim Booth