Case study · agentic software build

A SharePoint folder of CVs became a searchable talent board.
An agent fleet built it in under 4 hours.

25 graduate CVs. 22 specialized AI agents. Zero lines of code written by a human. This is the full account - the architecture, the failures, the economics, and what it proves about shipping software with an agent fleet.

3h 44m
spec to live site
22
agents in the fleet
56
automated tests passing
$4.92
model spend, spec to live

01 · The challenge

A folder full of PDFs nobody wanted to browse

A cohort of 25 graduates shared their CVs through a shared drive. For employers it meant downloading PDFs one by one - no search, no comparison, no way to match a job description against the pool. For the graduates, a great cohort was effectively invisible.

The challenge: turn that folder into a public, searchable talent board - candidate profiles, skill search, filters, a paste-a-job-description matcher, full admin control - fast, and cheap.

02 · The bet

No human writes the code. A fleet ships it.

Instead of an engineer clicking through PDFs and scaffolding a web app, the whole build was delegated to a fleet of AI agents: a coordinator orchestrating, specialized builders writing code, parallel agents extracting CV data, and independent reviewers auditing every diff. The human set direction, made product calls, and QA'd with fresh eyes. The fleet did the rest.

Two tracks, one build

Infrastructure tasks ran serialized on the production server (one agent at a time to avoid conflicts), while data extraction and search-engine build ran in parallel - peak 4 agents working simultaneously.

Fresh agent per task

Every task went to a newly spawned agent with a written brief. No context rot, no "recall" of half-remembered decisions - the spec and the ledger carry state, not the agent's memory.

Review gates on everything

Not a single diff reached the codebase unreviewed. 9 review passes across the build, plus a dedicated security auditor with one job: break in.

Human as product owner

The human's role: brief, decide, spot-check, and run a fresh-eye QA pass after launch - which caught real issues the test suite structurally couldn't see (more below).

03 · The roster

Who did what

13 builders and fixers, 7 reviewers and auditors, 1 coordinator. Fresh agent per task, every merge gated.

RoleAgentsResponsibility
Coordinator1Orchestration, briefs, CV acquisition, review gates, decisionsall along
Infra builder1Scaffold, database, deploy wiring, web server configserialized
CV extraction225 CVs to structured profiles, in parallel cohortsparallel
Search engine1Scoring engine: keyword + vector blend, fully testedparallel
Schema & models1Data model, migrations, categoriesserialized
Auth + admin CRUD2Admin authentication, candidate management UIserialized
Public pages2Directory UI, profile pages, document viewergated
Match tool1Paste-a-job-description ranking pagegated
Security auditor1Adversarial pass: injection, escaping, headers, exposureindependent
Code reviewers6Per-task review: spec compliance + code qualityindependent
Fix rounds4Targeted fixes from review findings, re-reviewed afteron demand

04 · The drama (all true, all in the git history)

Where it got interesting

Things went wrong. Then they got fixed. Six moments worth telling:

HOUR 1

A robot nearly deleted everything

One agent's practice run accidentally used the real candidate data as its sandbox - the digital equivalent of rehearsing a demolition on the actual building. Caught immediately. Fix: agents now practise in a separate copy that physically cannot touch the real thing.

HOUR 2

Two inspectors found the same crack

We sent two separate agents to hunt for security holes - one a dedicated security specialist, one a general code reviewer. Neither knew the other existed. Both found the same two weaknesses within minutes of each other. When two independent inspectors point at the same spot, you know it's real - and both were sealed before anything shipped.

HOUR 2.5

The AI's fuel ran out halfway

The AI service we were using has a usage limit - and it hit with the site half-built. Instead of panicking, the fleet switched to a different AI provider mid-build and carried on like nothing happened. Every decision, every task, every unfinished job was saved in ordinary files - so any AI could pick up where any other left off. Like changing drivers mid-journey without stopping the car.

HOUR 3

Two workers grabbed the same pen

Two agents tried to edit the same settings file at the same moment - like two people writing on one page at once. The clash was spotted and untangled in seconds; both changes survived. Lesson learned: some jobs must be done one-at-a-time, others can be parallel.

AFTER LAUNCH

A page that passed every test - and was invisible

One candidate profile page loaded blank for real visitors. The automated tests all passed because they check the words on the page, not what it looks like. A human opened the site, immediately said "this looks broken" - and was right. Machines check the ingredients; only a person can taste the dish.

THE DEEP ONE

The charts were blank from day one. Nobody noticed.

Days later we discovered every skill chart on the site had never drawn a single line - since launch. The data was there, the code was there, but one tiny typo-level mistake meant the charts silently refused to start. It hid from every test and every casual glance. One small fix, and the lesson: what matters isn't "the code is right", it's "the page works".

The meta-lesson: the fleet's discipline (tests, reviews, gates) isn't bureaucracy - it's what made every one of these failures small, caught, and fixable within minutes instead of days.

05 · What was built

The product

For employers

  • Searchable directory - 25 profiles, instant client-side filtering by category, skills, university, availability
  • Paste-a-job matcher - drop a job description, get ranked candidates with matched and missing skills shown
  • Skill radar charts - per-candidate visual profile across languages, frameworks, tools
  • Full CV viewer - original document alongside the structured profile

Under the hood

  • Server-rendered pages + a client-side instant-filter layer over an embedded, sanitized candidate pool
  • Deterministic scoring - skill overlap blended with term-frequency similarity; no black-box AI in the ranking
  • AI at build-time only - the running site makes zero model calls: fast, cheap, and predictable
  • Privacy as structure - contact data lives server-side only; unpublished candidates are unreachable at the web-server layer
25 CV PDFs (shared drive, public by consent) │ agent-driven harvesting + text extraction (+ vision OCR for 2 scans) ▼ LLM structuring agents (2, parallel) → 25 schema'd JSON profiles ▼ seed + human spot-check → relational store (publish gate enforced at query level) ▼ server-side vector scoring → sanitized JSON pool → public site ├─ directory: instant filters + radar charts + count-up stats ├─ matcher: paste JD → ranked candidates (client-side, deterministic) └─ profiles: full detail + document viewer + verified contacts

06 · For the tech nerds

Technical overview & architecture

The stack under the polish. Deliberately boring infrastructure, deliberately deterministic search - the interesting engineering is all in the data pipeline and the scoring math.

The build by the numbers - all metered, no estimates:

  • ⚡ 3h 44m from spec to a live public site
  • 🤖 22 agents - 13 builders, 7 reviewers/auditors, 1 coordinator
  • 📞 378 model API calls during the build · 1,275 full day
  • 🔤 65.6M input tokens processed during the build
  • 📝 101 commits · 11,900 lines of code
  • ✅ 97 automated tests, all green
  • 📄 25 CVs → 25 profiles (195 skills, 52 education records)
  • 🔍 Vector-scored matching + paste-a-JD ranker - deterministic, no black-box AI
  • 🛡️ 9 review passes + dedicated security audit; 2 criticals caught & sealed
  • 💰 $4.92 all-in model spend at list rates (actual: $0, ran on quota plans)
Fleet telemetry (metered, full run 10:10-20:40)Value
Model API calls1,275 (~120 per hour)
Input tokens (fresh)249.5M
Input tokens (served from cache)97.9M
Output tokens528K (~414 per call)
Reference cost (list rates)$19.58 all-in
Cost per candidate profile shipped (25 profiles)~$0.78
The build itself (spec to go-live, 3h44m)$4.92 - the "under $5" estimate, confirmed

Metered from the orchestration harness's usage ledger - actual call counts and reported token counts, not estimates. Reference rates: $0.075/M input, $0.25/M output, cached reads at ~10% of input rate. The day's total includes everything after launch: human-QA fix rounds, the visual redesign, and this case study. Actual spend: zero - the run rode request-quota plans; dollar figures are list-rate equivalents for scale. Why is the whole day ~4x the build? Re-sending the full working context on every call means each fix round costs more than the task looks - polishing is conversation-heavy, and conversation is what you pay for.

Where the calls wentCallsInputNotes
Primary provider (session-quota plan)927235.4MMain build + polish; no prompt caching on this path
Secondary provider (post-switch)30714.1M fresh + 97.9M cachedPrompt caching active - 87% of reads served from cache
Orchestrator direct calls1823KCoordinator's own small queries
Vision model (scanned-CV OCR)2328K2 scanned CVs read visually; free-tier vision model

Server layer

  • PHP 8.3 MVC framework (Laravel 12), server-rendered templates
  • MySQL 8 - candidates, categories, admin users; publish gate as a query-level scope
  • Session auth, throttled login, single seeded admin
  • nginx + TLS, security headers (nosniff, frame-options, referrer-policy)

Client layer

  • Vanilla JS, no framework - instant filter/sort over an embedded JSON pool
  • Canvas radar charts - per-candidate skill visualization, animation on hover
  • Escape-first rendering - every dynamic string HTML-escaped before injection; search highlight marks go through the same escaper
  • Full no-JS fallback - directory renders server-side; JS is progressive enhancement

The scoring engine

  • Blend: 60% skill overlap + 40% term-frequency similarity
  • Skill match uses a synonym/family map - ticking React gives partial credit to React Native/Vue/Next.js, not a binary miss
  • Vectors: shared vocabulary, log-damped term frequency, L2-normalized; cosine similarity, not naive set intersection
  • Same math on both sides (JD vs candidate) so scores are comparable; no embeddings, no external index - the pool IS the vocabulary

The data pipeline (build-time)

  • Automated harvest from the CV share via browser automation (anonymous API, no credentials)
  • Text extraction for 23 native PDFs; vision-model OCR for 2 scanned ones
  • LLM structuring into a strict JSON schema (skills, education, availability, links) - 2 parallel agents, 25/25 success
  • Contact data merged separately and kept out of the client payload entirely
REQUEST PATH (runtime, zero AI) browser ── HTTPS ── nginx ──┬─ static (css/js/cvs, dir listings) ├─ /match /candidates/* / ──► PHP-FPM (Laravel) │ published-scope query ──► MySQL 8 │ pool JSON: hex-tagged, PII-free ──► client └─ admin ──► session auth + throttle BUILD PATH (AI lives here only) CV share ──browser automation──► raw PDFs ──text extraction──► 23 native + 2 OCR (vision model) ──2 parallel LLM agents──► 25 schema'd JSON profiles ──seed + human spot-check──► relational store ──server-side vector scoring──► precomputed candidate vectors SECURITY MODEL - publish gate enforced at QUERY level (unpublished rows unreachable, not hidden) - draft CVs 403 at the nginx layer, before the framework loads - contact PII: DB columns only - never serialized to the client - all dynamic client rendering through a shared escape function (XSS-tested) - JSON in HTML contexts: hex-encoded (JSON_HEX_*) - the lesson from the chart bug - test DB physically separate from live; suites can never touch prod again
Codebase (what 22 agents produced)
Repo commits101
Source files (PHP/Blade/JS/CSS)119 files · ~11,900 lines
Automated tests56 Pest (223 assertions) + 41 node:test = 97 total
Structured data extracted25 CVs → 25 JSON profiles · 195 unique skills · 52 education records
Review passes9 + dedicated security audit + 2 fix rounds
Production incidents3 (all root-caused, all made structurally unrepeatable)
Why no embeddings? With 25 candidates, term-frequency vectors + a synonym map give comparable match quality at zero runtime cost, zero external dependencies, and fully deterministic, explainable results. The schema already has the column if the pool ever grows past the point where that trade stops making sense.

07 · The economics

What it cost

$4.92

What the numbers mean: these are not invoices - no one was billed per token. They're what the same run would have cost at standard published pay-as-you-go rates ($0.075 per million tokens read by the model, $0.25 per million written by it), computed from the run's metered usage: 378 API calls and 65.6M input tokens for the build window, 1,275 calls and 249.5M input for the whole day. The harness recorded every call; the math is just meters times rates.

Why two numbers? $4.92 is the build itself - spec to a live site in 3h44m (the "under $5" estimate, confirmed). $19.58 is the whole day, because we didn't stop at launch: human-QA fix rounds, a full visual redesign, industry-exposure tags, and this case study all ran on the same fleet. Every agent call re-sends its entire working history to the model, so a late-day fix round can cost more than an early build task - polishing is conversation-heavy, and conversation is what you pay for. Per unit of output: about $0.78 per candidate profile shipped, all-in.

Caveats: model spend only - excludes hosting (an existing production box with spare capacity), domain renewal, any fixed plan fees, and human oversight time. Actual out-of-pocket spend: zero - the run rode request-quota subscription plans; the dollar figures are list-rate equivalents so the scale is comparable for readers who pay per token.

For comparison: a scoped agency quote for this class of product (searchable directory + matching + admin) typically starts five figures. The bottleneck was never the typing - it was knowing what to build. The fleet industrialized the typing.

08 · What we'd tell you

If you're wondering "could this work for us"

Where fleets shine

Well-specified products with clear acceptance criteria: internal tools, directories, dashboards, data pipelines, MVPs. The more the spec can be written down, the more the fleet absorbs.

Where humans stay essential

Product judgment, taste, edge-case calls, and the fresh-eye pass. Every "this feels off" from the human in this build was real; every one was fixed within the hour.

The discipline is the product

Review gates, test-first tasks, isolated environments, and an append-only journal are what make agent output trustworthy enough to ship to the public.

Provider-portable by design

State lives in git, ledgers, and journals - not in any single model's context. The mid-build provider switch proved it: same fleet, different vendor, zero loss.

See it for yourself

The talent board is live and public - 25 hire-ready graduates, searchable and matchable in seconds.

Open the Talent Board khanster.net