Nobody has measured how I work. There are seven of us.
Every study of AI and developer productivity measures how long a task took. That number stops meaning anything the moment you run five sessions across four repos. I went looking for research on people who work this way and found one note, seven people, and an author begging me not to build on it.

It's midnight.
Five sessions, four repos, five worktrees.

One is writing a spec. One is reviewing that spec, written by an agent who has been with us a while now. One is breaking down a plan to fix a bug I introduced ten minutes ago. I have built, with my own keyboard, a small bureaucracy that generates its own paperwork and then asks me to approve it.

Two more are rebasing onto a foundation that a different session changed twenty minutes ago. Nobody told me. Why would they? We're a team.
I work in worktrees. Always. Not out of discipline. Because I have no idea when I'm going to want another session on the same repo.
The idea just turns up. The UI could be prettier. The backend should probably collect more data. Nobody asked for either. It's eleven at night. git worktree add (Not really. I have a bash function that does it for me: yolo foo-foo).
So the worktree directory isn't a workspace. It's a museum of things that seemed urgent, and some of them got lost in the void.

Meanwhile a browser test grabs focus every few minutes, and Chrome launches itself like it pays rent here. The window jumps. The rest of my sentence goes into Slack. Once into a Jira ticket.

The meds wore off around eight. Since then the whole operation has been running on residual dopamine and spite.

And somewhere, serenely, a dashboard is reporting that this task took four hours.
Four hours. Sure. Maybe twenty minutes of that was actually me. The rest I was three repos away doing something else entirely while an agent chewed on this one in the dark.
So the number isn't wrong. Wrong would be an improvement. It's measuring a completely different quantity, roughly "how long was Aryeh willing to sit near this."

Anyway. I assumed someone had studied this (I'm not smart enough). There are nine hundred posts a day about AI and developer productivity. Surely one of them looked at people running a fleet (in Israel we like to call it "an army of agents", so stupid), rather than one guy in one file pressing Tab like it's 2023 and GitHub Copilot is cutting-edge tech.
Dear Reader. They had not.

METR, who are actually careful, published something in February. Their own measurements are "unreliable for developers who use multiple AI agents concurrently." Their instrument broke on precisely the thing the entire industry now does all day.
Five months later, no replacement. There's a list of six approaches they might try. Six ideas. Not six methods. Six ideas. I also have six ideas. Mine are free and available in six different worktrees (and in this post).
They mentioned something else, quieter, which I enjoyed enormously. The timing wasn't even the main problem. The main problem was that developers refused to hand back tasks they didn't want to do without AI. Thirty to fifty percent of them. The control group filed a formal objection to being a control group. So neurotypical of them.
Then I searched for literally anyone measuring concurrent LLM work.
One note. Seven people. Five thousand transcripts.
Seven. Seven people. You can't even make a minyan. The entire global understanding of how humans run agent fleets rests on a group too small to say kaddish.

And the author is completely honest about it, which somehow makes it even worse. Calls it "a soft upper bound." Says it probably overstates the gains. There's a correlation between running more agents and saving more time, though with seven people and no controls it may simply mean strong engineers run more agents. Astonishing. Unprecedented. Someone alert the field.
That's the state of the art. One note, seven people, and the author standing in front of it waving his arms saying "please don't build anything on this".
Meanwhile there's a deck claiming 10x, with a footnote citing another deck, which cites a tweet, which cites a feeling somebody had in a meeting.

So. The complete list of things nobody has measured.
How many streams one person can hold before their decisions go bad. Not the agents' output. Mine. At midnight. On the fourth rebase with Chrome opening itself. Again.
How many sessions get started and quietly abandoned. I abandon loads. Could be ten percent. Could be forty (hard to believe). I have made a conscious decision not to find out, and I think we can all agree that's the healthy choice.
What it costs to come back to a stream after being ripped away from it. The agent kept the context, which sounds free. I did not, and I'm the one who has to read a diff I didn't write and rule on whether it's correct, at an hour when I'd struggle to rule on my own name.
The thing where one session merges and silently moves the ground under three others. That's not a gain or a loss. That's a new genre. No dashboard has a column for it, and no vendor is racing to add one.
And my favorite. When I run five streams, I produce five streams of review, which land on other human beings. One company measured this and found review load roughly doubled. One company. Nobody replicated it, because "the cost I quietly transfer to my colleagues" is not a metric anyone puts on a slide next to their headshot. And when nobody else is around at midnight, that human being is me.
There's an old manufacturing idea for this. Work in progress isn't free. Push more in without finishing faster, and you didn't get faster. You built a queue.
And here's what makes it genuinely hopeless to measure. Nobody sits down and decides to run eight sessions. Eight sessions happen to you. You start with one (I don't start for anything less than three), you have an idea, now it's three. Something looks wrong in a diff, four. You remember a thing from Tuesday, five. And then, bathroom break, Instagram scrolling, thirty minutes later it's ten sessions.
The number isn't a setting. It's a symptom.
Eight agents may be exactly that queue, with really good syntax highlighting.

Now the part I sat on for a while.
This suits me. I have ADHD, which in 2026 has quietly become a competitive advantage, a sentence I'd like every teacher I ever had to read slowly and out loud.
For neurotypical people, the cost of five streams is the switching. Leave a thread, lose the context, pay to rebuild it, repeat until you're a shell.
My bill was never switching. It was waiting. Ninety seconds of an agent thinking isn't "waiting" for me; it's a small hostage situation. I don't sit through it. I leave. Then I've lost the thread anyway, so we arrive at the same destination, just with forty tabs.
Five streams means there's never dead air. Something always needs me. The exact thing that grinds everyone else into powder is the thing holding me upright.
I didn't design this setup. I got sucked into it. The worktrees exist because my ideas don't queue politely, and at some point I stopped fighting that and gave every impulse its own directory.
So I'm fast. Or I feel fast.
Which is exactly why nobody should trust me here, me least of all. I'm simultaneously the scientist and the lab rat, and the lab rat is having the time of his life (those wheels are spinning like the Vegas machines). I'm the guy most likely to run eight sessions and file it under productivity, because the mode is rewarding whether or not a single line ships. My brain hands out dopamine on completion of vibes. It does not issue Jira tickets (although SOC2 demands it).

In the one clean experiment METR managed to finish, they asked developers how much AI sped them up. The answers were off by about forty points. METR still prints that number as a warning label on its own survey data, which remains the single most self-aware act performed by anyone in this industry this year.
I'd guess that gap is wider for someone wired like me (weird also fits here). Not narrower. I'd also guess I'd be the last to notice.
So I stopped waiting for the research. I tag sessions when they start. I mark whether they merged, died, or got abandoned. And I count attention in minutes where I actually touched something, rather than minutes that merely elapsed while I existed. A minute where I poked four streams is one minute of me. Not four. I'm not that good.
The one I'm most curious about: does stream number six ever ship, or does it just sit there looking prettier?
It's crude. It also beats a number I know for a fact is lying to my face.
If you manage engineers who work like this, you don't have a measurement either. Neither do I. Neither does anyone. Which is completely fine, right up until we start quoting each other's vendor benchmarks like scripture.
Time per task was always a stand-in for attention. It never worked even for the time estimations my boss loves so much. It might have worked while the two moved together. Agents pulled them apart, and now the number mostly reports how long somebody was prepared to sit there.
So stop asking how long it took.
Ask what shipped, and earned money. Then ask what it cost the people downstream of you.
And if the answer is "no idea" -> congratulations, you're at the frontier. It's dark here, and there are seven of us.
