Two tracks, one build
Infrastructure tasks ran serialized on the production server (one agent at a time to avoid conflicts), while data extraction and search-engine build ran in parallel - peak 4 agents working simultaneously.
Case study · agentic software build
25 graduate CVs. 22 specialized AI agents. Zero lines of code written by a human. This is the full account - the architecture, the failures, the economics, and what it proves about shipping software with an agent fleet.
01 · The challenge
A cohort of 25 graduates shared their CVs through a shared drive. For employers it meant downloading PDFs one by one - no search, no comparison, no way to match a job description against the pool. For the graduates, a great cohort was effectively invisible.
The challenge: turn that folder into a public, searchable talent board - candidate profiles, skill search, filters, a paste-a-job-description matcher, full admin control - fast, and cheap.
02 · The bet
Instead of an engineer clicking through PDFs and scaffolding a web app, the whole build was delegated to a fleet of AI agents: a coordinator orchestrating, specialized builders writing code, parallel agents extracting CV data, and independent reviewers auditing every diff. The human set direction, made product calls, and QA'd with fresh eyes. The fleet did the rest.
Infrastructure tasks ran serialized on the production server (one agent at a time to avoid conflicts), while data extraction and search-engine build ran in parallel - peak 4 agents working simultaneously.
Every task went to a newly spawned agent with a written brief. No context rot, no "recall" of half-remembered decisions - the spec and the ledger carry state, not the agent's memory.
Not a single diff reached the codebase unreviewed. 9 review passes across the build, plus a dedicated security auditor with one job: break in.
The human's role: brief, decide, spot-check, and run a fresh-eye QA pass after launch - which caught real issues the test suite structurally couldn't see (more below).
03 · The roster
13 builders and fixers, 7 reviewers and auditors, 1 coordinator. Fresh agent per task, every merge gated.
| Role | Agents | Responsibility | |
|---|---|---|---|
| Coordinator | 1 | Orchestration, briefs, CV acquisition, review gates, decisions | all along |
| Infra builder | 1 | Scaffold, database, deploy wiring, web server config | serialized |
| CV extraction | 2 | 25 CVs to structured profiles, in parallel cohorts | parallel |
| Search engine | 1 | Scoring engine: keyword + vector blend, fully tested | parallel |
| Schema & models | 1 | Data model, migrations, categories | serialized |
| Auth + admin CRUD | 2 | Admin authentication, candidate management UI | serialized |
| Public pages | 2 | Directory UI, profile pages, document viewer | gated |
| Match tool | 1 | Paste-a-job-description ranking page | gated |
| Security auditor | 1 | Adversarial pass: injection, escaping, headers, exposure | independent |
| Code reviewers | 6 | Per-task review: spec compliance + code quality | independent |
| Fix rounds | 4 | Targeted fixes from review findings, re-reviewed after | on demand |
04 · The drama (all true, all in the git history)
Things went wrong. Then they got fixed. Six moments worth telling:
HOUR 1
A robot nearly deleted everything
One agent's practice run accidentally used the real candidate data as its sandbox - the digital equivalent of rehearsing a demolition on the actual building. Caught immediately. Fix: agents now practise in a separate copy that physically cannot touch the real thing.
HOUR 2
Two inspectors found the same crack
We sent two separate agents to hunt for security holes - one a dedicated security specialist, one a general code reviewer. Neither knew the other existed. Both found the same two weaknesses within minutes of each other. When two independent inspectors point at the same spot, you know it's real - and both were sealed before anything shipped.
HOUR 2.5
The AI's fuel ran out halfway
The AI service we were using has a usage limit - and it hit with the site half-built. Instead of panicking, the fleet switched to a different AI provider mid-build and carried on like nothing happened. Every decision, every task, every unfinished job was saved in ordinary files - so any AI could pick up where any other left off. Like changing drivers mid-journey without stopping the car.
HOUR 3
Two workers grabbed the same pen
Two agents tried to edit the same settings file at the same moment - like two people writing on one page at once. The clash was spotted and untangled in seconds; both changes survived. Lesson learned: some jobs must be done one-at-a-time, others can be parallel.
AFTER LAUNCH
A page that passed every test - and was invisible
One candidate profile page loaded blank for real visitors. The automated tests all passed because they check the words on the page, not what it looks like. A human opened the site, immediately said "this looks broken" - and was right. Machines check the ingredients; only a person can taste the dish.
THE DEEP ONE
The charts were blank from day one. Nobody noticed.
Days later we discovered every skill chart on the site had never drawn a single line - since launch. The data was there, the code was there, but one tiny typo-level mistake meant the charts silently refused to start. It hid from every test and every casual glance. One small fix, and the lesson: what matters isn't "the code is right", it's "the page works".
05 · What was built
06 · For the tech nerds
The stack under the polish. Deliberately boring infrastructure, deliberately deterministic search - the interesting engineering is all in the data pipeline and the scoring math.
The build by the numbers - all metered, no estimates:
| Fleet telemetry (metered, full run 10:10-20:40) | Value |
|---|---|
| Model API calls | 1,275 (~120 per hour) |
| Input tokens (fresh) | 249.5M |
| Input tokens (served from cache) | 97.9M |
| Output tokens | 528K (~414 per call) |
| Reference cost (list rates) | $19.58 all-in |
| Cost per candidate profile shipped (25 profiles) | ~$0.78 |
| The build itself (spec to go-live, 3h44m) | $4.92 - the "under $5" estimate, confirmed |
Metered from the orchestration harness's usage ledger - actual call counts and reported token counts, not estimates. Reference rates: $0.075/M input, $0.25/M output, cached reads at ~10% of input rate. The day's total includes everything after launch: human-QA fix rounds, the visual redesign, and this case study. Actual spend: zero - the run rode request-quota plans; dollar figures are list-rate equivalents for scale. Why is the whole day ~4x the build? Re-sending the full working context on every call means each fix round costs more than the task looks - polishing is conversation-heavy, and conversation is what you pay for.
| Where the calls went | Calls | Input | Notes |
|---|---|---|---|
| Primary provider (session-quota plan) | 927 | 235.4M | Main build + polish; no prompt caching on this path |
| Secondary provider (post-switch) | 307 | 14.1M fresh + 97.9M cached | Prompt caching active - 87% of reads served from cache |
| Orchestrator direct calls | 18 | 23K | Coordinator's own small queries |
| Vision model (scanned-CV OCR) | 23 | 28K | 2 scanned CVs read visually; free-tier vision model |
| Codebase (what 22 agents produced) | |
|---|---|
| Repo commits | 101 |
| Source files (PHP/Blade/JS/CSS) | 119 files · ~11,900 lines |
| Automated tests | 56 Pest (223 assertions) + 41 node:test = 97 total |
| Structured data extracted | 25 CVs → 25 JSON profiles · 195 unique skills · 52 education records |
| Review passes | 9 + dedicated security audit + 2 fix rounds |
| Production incidents | 3 (all root-caused, all made structurally unrepeatable) |
07 · The economics
What the numbers mean: these are not invoices - no one was billed per token. They're what the same run would have cost at standard published pay-as-you-go rates ($0.075 per million tokens read by the model, $0.25 per million written by it), computed from the run's metered usage: 378 API calls and 65.6M input tokens for the build window, 1,275 calls and 249.5M input for the whole day. The harness recorded every call; the math is just meters times rates.
Why two numbers? $4.92 is the build itself - spec to a live site in 3h44m (the "under $5" estimate, confirmed). $19.58 is the whole day, because we didn't stop at launch: human-QA fix rounds, a full visual redesign, industry-exposure tags, and this case study all ran on the same fleet. Every agent call re-sends its entire working history to the model, so a late-day fix round can cost more than an early build task - polishing is conversation-heavy, and conversation is what you pay for. Per unit of output: about $0.78 per candidate profile shipped, all-in.
Caveats: model spend only - excludes hosting (an existing production box with spare capacity), domain renewal, any fixed plan fees, and human oversight time. Actual out-of-pocket spend: zero - the run rode request-quota subscription plans; the dollar figures are list-rate equivalents so the scale is comparable for readers who pay per token.
08 · What we'd tell you
Well-specified products with clear acceptance criteria: internal tools, directories, dashboards, data pipelines, MVPs. The more the spec can be written down, the more the fleet absorbs.
Product judgment, taste, edge-case calls, and the fresh-eye pass. Every "this feels off" from the human in this build was real; every one was fixed within the hour.
Review gates, test-first tasks, isolated environments, and an append-only journal are what make agent output trustworthy enough to ship to the public.
State lives in git, ledgers, and journals - not in any single model's context. The mid-build provider switch proved it: same fleet, different vendor, zero loss.
The talent board is live and public - 25 hire-ready graduates, searchable and matchable in seconds.