This post was inspired by a 50-hour day on May 8. By the time I had finished writing it, a more recent peak on May 16 had become the sharper data point. Same arbitrage, tighter compression, but a much higher parallelization. Both days are part of the same story about inference costs and what it is we are actually buying.
On May 16 I shipped 38 hours of work in one day. The clock said 12 hours, 5:59 in the morning to 6:11 in the evening. The retail bill for the compute I burned was $4,548 in inference costs across 185 parallel sessions. The actual bill I paid was about $6.70, because I am on a flat-rate Claude Max subscription at $200 a month. That gap between 12 wall-clock hours, 38 hours of completed work, and a $6.70 day is the whole story I want to tell. Inference costs are not a tax. They are the meter that shows you what your subscription is actually buying. Right now the meter is wildly favorable for engineers who know how to spend it. This is what time travel looks like when someone else is paying for the fuel.
What Inference Costs Actually Buy
Most engineers still think of inference costs as a budget line item. Something the finance team yells about. A meter ticking next to the API dashboard. That framing is exactly backwards for anyone on a flat-rate plan.
Inference costs are the retail price of compressed engineer time. The subscription is the wholesale price of access to it. When I burn $4,548 of retail compute in a day and the bill is $6.70, I am not setting money on fire. I am exploiting the largest mispricing in software right now. The subscription tier has not caught up to what these tools can actually produce in the hands of someone who knows how to feed them.
The arithmetic only looks insane if you forget that the alternative is paying a human to do the same work, three times slower, one thread at a time. I am not paying for tokens. I am paying $200 a month for parallel wall-clock.
The Receipts
Days like May 16 are not abstract. The bulk of the day went into dirigible2D, my next game project, and four releases shipped between 8:16 in the morning and 5:57 in the evening.
The big one was a construction system foundation: blueprint bridge, tile-output executor, sample wall and floor recipes, refining content (cloth, cut stone, iron ingot, plank) wired through sawmill and smelter end-to-end, BuildPanel UI scaffolding, and a dev console for construction work orders. Right after that, runtime perf instrumentation with an F3 overlay, chunk per-frame throttling, a pathfinding cache, and end-to-end deconstruction with tile-side loot. Then field-of-view explored-tile persistence, layer cycling shortcuts, save-restore reconciliation, and a camera glitch fix. Then a final pass: perf-overlay entity counts across eight categories, a centralized save-thumbnail pipeline, a class rename, and a new HUD primitive. Four ships, two breaking schema migrations, one day.
Across the rest of the 60-day streak, the public portfolio has been moving in parallel. LAIRD, my ASCII roguelike with an LLM narrative overlay, is 18,000 lines across 188 files and playable on Itch; built on MToolKit, my open source Unity runtime framework, and LlamaBrain, my open source deterministic AI control plane which runs more than 1,800 integration tests against its validation gates, determinism boundaries, and Black Box Audit Recorder. Claude Interrogate, my open source adversarial design-doc skillset, has been picking up MCP runtime work and plugin scaffolding.
May 16 was the peak, not the floor. Over the last 30 days my retail inference burn averaged $1,571 a day, $47,142 total, with 26 of 30 days inside the operating envelope. The actual subscription paid was about $6.70 a day, $200 for the month. The daily-ship streak runs 60 days and counting. That hero day is not an anomaly. It is what happens when the floor is high enough that an exceptional day is still in-bounds.
Inference Costs as Leverage, Not Waste
Here is the trick. I was not at the keyboard for 38 hours on May 16. I’m not actually a time traveler but I was at the keyboard for 12. The other 26 hours came from sessions stacking in parallel. While one agent implements a save migration, another audits a docs set for contradictions, another generates regression tests, another summarizes yesterday’s shipping notes for tomorrow’s prompts. All of them wait for me. I do not wait for any of them.
That is what 38 hours of work in 12 hours of clock time actually means. Three pipelines, sometimes five, sometimes more. The retail dollar figure is the proof I am not faking it. The flat subscription is the proof the market has not woken up yet. Inference costs scale with parallel pipelines because parallel pipelines are the unit of compressed time, and right now the pricing on that unit is structurally broken in the customer’s favor. Anyone still working single-threaded in human serial is going to be very surprised when they look up from their IDE and find the floor has moved.
Why This Only Works With Governance
Plenty of engineers can spawn agents in parallel. Most of what comes back is junk. The reason May 16 produced 38 hours of usable output rather than 38 hours of slop is that I have spent the last year building the control plane that makes stochastic systems behave.
Validation gates that refuse to merge LLM output failing a structural contract. Determinism boundaries that pin replayable behavior. Audit recorders that catch drift before it ships. LlamaBrain is the proof. This pattern is not unique to my stack, but the discipline is rare. If your agents are returning slop, the model is not the problem. The problem is that you are letting them write directly into production without an enforcement gate.
The Measurement Gap for Inference Costs
The honest problem with all of this is that the measurement layer does not exist yet. The dashboards engineers reach for were built for the 40-hour-week reality. There is no good public way to show a hiring manager, a co-founder, or a peer that you are operating at a 3x compression factor day after day on a flat-rate plan. The industry has not yet built the receipts.
The Call
To founders and engineering leaders, this is what Principal Systems Engineer looks like in 2026. Three open source projects, one shipped game on Itch, a 60-day daily-ship streak, retail inference burn averaging $1,500 a day on a $200-a-month subscription, governance-first architecture. The bottleneck on your roadmap is not inference costs. It is finding the engineers who already know how to spend them at this kind of force multiplier.
I am one of them. Receipts are public at promptbook.gg/michael-tiller, third-party tracked, not my own math. My own math ships soon, but like all my other theses it will be local first.