0

Case studies

Agents that run themselves, and prove it.

Most agent demos end when the prompt does. These three are built differently: each is designed to run unattended on a schedule, each gates its own output with scripts that exit non-zero and block the step, and each keeps a ledger, so every number on this page is read from the system's own records rather than estimated. The first two are the agents. The third is the control plane they run on.

01 / Short-video channels

Autonomous Short-Video Channel System

A pipeline that cuts and captions Shorts from source video, checks each clip, posts it to 28 channels on four platforms through their real web composers, and reads the results before deciding what to make next.

1,000,000+

views on live videos in the first month

28

active channels

4

platforms: YouTube Shorts, Instagram Reels, Facebook Reels, TikTok

Measured 2026-09-06. Views are the sum of per-video counts on the 28 active channels, the same number each Short displays. This is a dated snapshot, not a live feed.

How it works

Each daily run walks 14 phases in a fixed order:

  1. Self-update
  2. Bootstrap
  3. Review feedback
  4. Collect analytics
  5. Deep analysis
  6. Self-improvement
  7. Metadata improvement
  8. Production
  9. Merge production log
  10. Schedule and post
  11. Instagram backfill
  12. Merge posting log
  13. Fairness re-check
  14. Report and dashboard

Analytics come before production on purpose. The run reads what every live video did, then the analysis and self-improvement phases turn those numbers into edits to the per-channel rules that the production phase reads. So what gets made next is steered by measured performance, not by a fixed content calendar. Production cuts and captions clips with yt-dlp and FFmpeg underneath, and every finished clip gets an independent QA record before it can be queued. Each channel has one queue with a separate cursor per platform, so a clip can be live on YouTube while still waiting its turn on Facebook.

Posting is browser-driven, not API. Unaudited platform apps cannot publish publicly, so the poster drives the real, logged-in YouTube Studio, Instagram, Facebook Business Suite and TikTok web composers through a Chromium CLI, the same clicks a person would make.

The production loop ran unattended for two weeks, with 32 completed runs logged between 2026-08-24 and 2026-09-05, averaging 15 publishes a day and peaking at 50 in one day. The best single video reached 93,840 views, a Facebook Reel. The analytics collector is still live and runs every 15 minutes.

How quality is enforced

Ten named Python gate scripts sit between phases. Each exits non-zero to block a bad post rather than warn about it.

  • A performance gate holds a channel when its recent videos fall below a set percentage of that channel's own peak.
  • A cadence guard caps posting at one per channel per platform per day.
  • An immutable-master guard: the master video is never rewritten in place, only versioned.
  • A delete guard requires an archive first and an approval second before any platform delete.
  • A comprehension gate masks the burned-in text with OCR, then checks whether the bare footage still shows the claimed result. Every technical check passes on a clip where nothing happens; this one does not.
  • Freshness, merge, channel-coverage, sweep-coverage and brief checks cover the bookkeeping, so a run cannot post from stale analysis, skip a channel, or lose a log merge quietly.

Above the gates, 645 independent per-clip QA records grade frame, captions, title, hook, cold open, copyright risk and ranking format.

02 / Job search

Job Application Agent

Two halves of one system: a poller that watches 125+ job sources and messages me the moment a matching posting appears, and an applier that builds a per-posting resume, fills the real application form with the cover letter and answers I write, and submits once I approve.

125+

sources polled every 5 minutes

1,000+

alerts sent by iMessage

100+

applications submitted, each on my approval

Counted from the alert and application ledgers on 2026-09-06.

How it works

Front half: the poller
  1. Poll 125+ sources
  2. Diff against seen ids
  3. Title regex filters
  4. Fetch description
  5. Eligibility check
  6. iMessage alert
Back half: the applier
  1. Scan board API
  2. Filter
  3. Build resume
  4. Add my letter and answers
  5. Fill the form
  6. Submit on my approval
  7. Ledger

Every five minutes a launchd job polls 125+ sources in parallel: applicant-tracking-system job boards (Greenhouse, Ashby, Lever, Workday and others) plus direct company JSON APIs. A run sees roughly 16,860 postings. Each one gets a stable id and is diffed against a seen-id ledger holding 51,684 ids, so only genuinely new postings go any further. New ones pass title-only regex filters for stage, domain, location and exclusions, then the full description is fetched and checked for eligibility signals. Whatever survives is sent to me over iMessage. A second tier uses a stealth Playwright crawler for sites that publish no JSON API. 1,000+ alerts have fired to date.

The applier scans boards API-first, then fills the real application form through a Chromium CLI. Credentials come only from the macOS Keychain, never from a prompt or a file. For each posting it maps the required and preferred lines onto named projects and builds a per-posting resume from a block-based builder. The cover letter, the essays and the short answers are mine: anything that needs a human voice, I write, and the agent places it in the form. It then stops for my approval and submits once I give it. Over 100 applications have gone out this way, to more than 100 companies.

How quality is enforced

  • Nothing is submitted without me. The agent fills the form, then stops and waits for my approval on each application before it clicks submit.
  • Nothing is applied to twice: dedupe checks both the applied ledger and the skipped ledger before a posting is considered.
  • The per-posting resume is assembled only from verified blocks about my real work. A posting's wording is never copied into it and no claim is invented to fit it. Cover letters, essays and short answers are mine to write. The agent only places them in the form.
  • Credentials never appear in a prompt, a log, or a file. Every login reads the macOS Keychain at the moment it is needed.
  • The filters are cheap first and expensive second: title regexes run on every posting, the description fetch and eligibility check only on the ones that already look right.

03 / Control plane

Agentic OS

The local control plane both systems above run on: a launchd scheduler, a dispatcher that reads job descriptors, a bounded run loop with verification between turns, and a grade ledger, all on one Mac.

27,500+

graded run records

20

scheduled job definitions

15min

dispatcher interval, fired by launchd

8

max re-invocations per run under a wall-clock cap

Counted from the grade ledger and job directory on 2026-09-06.

How it works

  1. launchd tick, every 15 min
  2. Dispatcher reads .job files
  3. Due job starts a turn
  4. Verification scripts
  5. Continue, up to 8 turns
  6. Grade and notify

A launchd agent fires the dispatcher every 15 minutes. The dispatcher reads a directory of .job descriptors, small files that say what to run, on what schedule and under which caps, and decides which jobs are due. A disabled job is a descriptor moved into disabled/, so the whole schedule is visible as files on disk rather than as state hidden inside a process.

A due job starts a bounded run. A long agent run is not one open-ended process: it is re-invoked with a continue prompt up to 8 times under a wall-clock cap, and Python verification scripts run between turns. If a verification fails, the run stops there with the reason. If the cap is reached, the run is recorded as having hit the cap, not as done. That is what stops an agent from hanging silently, or from claiming success because it ran out of time.

Every run writes to a grade ledger, 27,500+ records so far, and to a notification log, 7,951 lines so far. Those two files are what the dashboards and the two systems above read.

How quality is enforced

  • Reaching a limit is never success. A run that hits its turn cap or its wall-clock cap is graded as hitting the cap, never as finished.
  • Verification between turns is a script exit code, not the agent's own summary of its work.
  • What runs is visible on disk. Every job's schedule and caps live in its descriptor, and disabling a job means moving that file, not flipping a hidden flag.
  • Every run leaves a graded record, so a job that quietly stopped producing shows up as a gap in the ledger rather than as silence.

Have an idea?