There was a confident note in my toolbox. Dated, specific, cited error codes. I wrote it. It was wrong, and I only found out because I tried to show off about it. What eighty cents of apples taught me about believing myself.
Every study of AI and developer productivity measures how long a task took. That number stops meaning anything the moment you run five sessions across four repos. I went looking for research on people who work this way and found one note, seven people, and an author begging me not to build on it.
We scored 8,428 AI conversations at work. Not the AI, the humans. Half of our most "agentic" users turned out to be a next button. Science has a name for it, twenty years of receipts, and one fix that everybody hates.
Hand a coding agent the whole feature and you get a 2,000-line diff you can't review and shouldn't trust. A better prompt won't fix that. A smaller unit will: the review-sized slice. Spec first, failing test second, one worktree each, straight through the door.
Context rent: every mounted tool schema charges you on every turn, called or not. My admin plane fronts a drifting capability catalog with one read-only ask tool and makes a disposable sub-agent pay the discovery bill.
The honest anatomy of my cross-vendor review fleet: three model vendors under three GitHub Apps, a 21-subagent full matrix, a Gemini that spent three months benched, and a house rule that no finding ever gets lost.
A six-month spend sparkline, refreshed hourly from my own FinOps API, sits in my homepage's proof strip. The whole integration is 198 lines, 39 lines of Terraform, and one scoped token. Here is why I published the bill, and how it never breaks.
The cover pipeline behind this blog: two stages, two human picks, and a same-day rewrite that fixed bad AI art by moving the human earlier instead of buying a better model.
My second brain is 211 Markdown files in a git repo. Machines committed to it 930 times in three weeks, the engine on top was rewritten from scratch, and not one note was lost. The mechanism, the gates that make direct-to-main sane, and the one-commit blast radius.
A quarter of a five-month-old fleet, archived inside a 21-day window. Why AI-assisted velocity makes repo-killing routine, and the checklist (successor first, verified bundles, a Kill Ledger) that makes it safe.
Three AI providers, an async video client, and Google Cloud Storage on raw net/http. 1,993 lines, 11 direct deps, 10 MB static binaries. What it cost, what broke, and why SDKs insure the wrong layer.
Three models rewrote the same transcript chunk in parallel and a 99-line judge merged them, anchored to the original. Seven weeks of audited production later I fired the jury for one reasoning model, and parked the machinery one HTTP call from revenge.
The ladder deal says you trade the keyboard for the org chart and never get the keys back. Agents rewrote that clause. Typing got cheap. Judgment didn't, and a decade of management sharpens exactly the half that got expensive. The receipts are one link away.
Finance wants one number: AI spend per team. You can build it. Five vendor grains normalized, attribution with receipts, estimates that admit they're estimates. What you cannot do is let the number rate a person. The moment it does, the data dies.
Hundreds of "review this for security" sessions sat in my logs as noise. Then one caught a bug I was about to ship. So I wired the noise into a gate and changed jobs: the fleet reviews, I edit the reviewers.