<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Model-Reviews | Bas Nijholt</title><link>https://www.nijho.lt/tag/model-reviews/</link><atom:link href="https://www.nijho.lt/tag/model-reviews/index.xml" rel="self" type="application/rss+xml"/><description>Model-Reviews</description><generator>Hugo</generator><language>en-us</language><copyright>© 2026</copyright><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://www.nijho.lt/media/icon_hu_408a8f6cef43e327.png</url><title>Model-Reviews</title><link>https://www.nijho.lt/tag/model-reviews/</link></image><item><title>Self-hosting AI does not save money, and I do it anyway</title><link>https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/self-hosting-ai-is-not-cheaper/</guid><description>&lt;p&gt;I have had this debate many times, so I am finally writing it down.&lt;/p&gt;&#10;&lt;p&gt;Whenever I say that self-hosting AI does not save money, people hear that I am against self-hosting.&#10;That is not the point at all.&#10;I am a massive fan of open-weight models.&#10;&lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L43" target="_blank" rel="noopener"&gt;Qwen3.8 27B runs on two RTX 3090s at home&lt;/a&gt; (and 15 more models), and my phone dictation goes to &lt;a href="https://www.nijho.lt/post/diction-agent-cli-qwen/"&gt;Qwen3-ASR on the same machine&lt;/a&gt;.&#10;I think open-source AI is the best thing since sliced bread.&lt;/p&gt;&#10;&lt;p&gt;I don&amp;rsquo;t pretend it saves money.&#10;I do it for fun, for sovereignty, and for privacy, which I come back to at the end.&lt;/p&gt;&#10;&lt;p&gt;All benchmark numbers and prices below come from &lt;a href="https://artificialanalysis.ai/" target="_blank" rel="noopener"&gt;Artificial Analysis&lt;/a&gt; and &lt;a href="https://openrouter.ai/" target="_blank" rel="noopener"&gt;OpenRouter&lt;/a&gt; as of September 30, 2026.&#10;I did the math for the hardware I own, and for the counterarguments I hear most.&#10;All assumptions are listed &lt;a href="#appendix-all-assumptions"&gt;at the end&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;Benchmarks are not everything, and a high score does not reliably predict how a model does on real work.&#10;They are still the best we have: all models are benchmaxed, tuned to do well on the popular benchmarks, so the scores at least work as a reference frame for comparing them with each other.&lt;/p&gt;&#10;&lt;p&gt;&lt;em&gt;Update, October 2, 2026: after the discussion on r/LocalLLaMA, I added &lt;a href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;section 5&lt;/a&gt; on Macs, DGX Sparks, and AMD AI Max machines.&lt;/em&gt;&lt;/p&gt;&#10;&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#1-the-model-i-love-and-the-one-that-beats-it-on-price"&gt;1. The model I love, and the one that beats it on price&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#2-but-my-gpus-are-already-paid-for"&gt;2. &amp;ldquo;But my GPUs are already paid for&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#3-why-datacenters-are-so-much-better-at-this"&gt;3. Why datacenters are so much better at this&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#4-but-i-have-solar-panels"&gt;4. &amp;ldquo;But I have solar panels&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;5. &amp;ldquo;But my Mac, DGX Spark, or AI Max does save money&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#6-the-bar-keeps-moving"&gt;6. The bar keeps moving&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#7-a-small-company-has-it-worse"&gt;7. A small company has it worse&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#8-they-lose-money-on-inference-so-prices-will-go-up"&gt;8. &amp;ldquo;They lose money on inference, so prices will go up&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#9-but-they-sell-your-data"&gt;9. &amp;ldquo;But they sell your data&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#10-where-self-hosting-does-win-the-same-model"&gt;10. Where self-hosting does win: the same model&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#11-so-why-do-i-do-it"&gt;11. So why do I do it?&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#appendix-all-assumptions"&gt;Appendix: all assumptions&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;h2 id="1-the-model-i-love-and-the-one-that-beats-it-on-price"&gt;1. The model I love, and the one that beats it on price &lt;a class="anchor" href="#1-the-model-i-love-and-the-one-that-beats-it-on-price" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My favorite local model right now is &lt;a href="https://artificialanalysis.ai/models/qwen3-8-27b" target="_blank" rel="noopener"&gt;Qwen3.8 27B&lt;/a&gt;.&#10;It came out in August and it is Apache-2.0.&#10;People often say models like this run on a single gaming GPU, but that is only true after quantizing them.&#10;At full precision, Qwen3.8 27B needs about 54 GB of memory.&#10;Quantized to about 4 bits per weight, it fits on one 24 GB card, at some cost in quality.&#10;I still &lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L56-L65" target="_blank" rel="noopener"&gt;split it across both of my 3090s&lt;/a&gt;, because the second card leaves room for a larger context window.&lt;/p&gt;&#10;&lt;p&gt;I compare it with OpenAI&amp;rsquo;s &lt;a href="https://artificialanalysis.ai/models/gpt-6-luna" target="_blank" rel="noopener"&gt;GPT-6 Luna&lt;/a&gt;, the cheap tier released on September 22.&#10;Both models let you choose how long they think, but the settings do not mean the same thing for both.&#10;Luna&amp;rsquo;s token use grows more than 20 times from its lowest setting to its highest, while Qwen&amp;rsquo;s barely changes.&#10;Qwen on low already writes more tokens than Luna on xhigh.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-effortTokens-2" class="plot-container" role="img" aria-label="Average output tokens per Intelligence Index task at each reasoning setting. Qwen3.8 27B has no high or max setting. Hover for scores."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Average output tokens per Intelligence Index task at each reasoning setting. Qwen3.8 27B has no high or max setting. Hover for scores.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-effortTokens-2");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["effortTokens"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;So the name of a setting says little on its own.&#10;I compare both models at xhigh, the highest setting Qwen offers, and count the tokens each one actually uses.&#10;The extra tokens Qwen needs are part of what I am measuring.&#10;There, Qwen scores 33.7 on the &lt;a href="https://artificialanalysis.ai/methodology/intelligence-benchmarking" target="_blank" rel="noopener"&gt;Artificial Analysis Intelligence Index&lt;/a&gt; and Luna scores 34.6.&#10;For reference, Claude Opus 4.6, a frontier model from February, scores 31.9.&#10;That is an amazing feat in itself.&#10;When Opus 4.6 was the best model we had, I never thought that within the same year I could run essentially that level of intelligence in my own house.&lt;/p&gt;&#10;&lt;p&gt;Artificial Analysis also &lt;a href="https://artificialanalysis.ai/leaderboards/models" target="_blank" rel="noopener"&gt;publishes&lt;/a&gt; how many tokens each model used to run the index, and what that cost.&#10;So the question is simple: what does it cost to run the whole benchmark once?&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-runCost-3" class="plot-container" role="img" aria-label="Cost of one run of the Artificial Analysis Intelligence Index. The two bars for hardware at home are electricity only: measured on my 3090s, estimated for the RTX PRO 6000."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Cost of one run of the Artificial Analysis Intelligence Index. The two bars for hardware at home are electricity only: measured on my 3090s, estimated for the RTX PRO 6000.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-runCost-3");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["runCost"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Luna runs the whole benchmark for $67.&#10;Qwen at full precision, through the &lt;a href="https://openrouter.ai/qwen/qwen3.8-27b/providers" target="_blank" rel="noopener"&gt;cheapest provider&lt;/a&gt; that does not keep your data (ZDR), costs $619.&#10;Two things cause that gap: providers charge almost four times as much per output token for Qwen, and Qwen generates almost three times as many tokens to get the same work done.&lt;/p&gt;&#10;&lt;p&gt;The second bar is my own machine: the electricity alone for running Qwen on my GPUs costs more than Luna&amp;rsquo;s entire bill.&lt;/p&gt;&#10;&lt;p&gt;The third bar is what it takes to match the API&amp;rsquo;s full precision at home: an RTX PRO 6000, a workstation card with 96 GB of memory, enough for the full model.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&#10;Its electricity alone costs twice Luna&amp;rsquo;s bill, and one run takes more than seven weeks.&#10;The card by itself sells for $10,000 to $20,000, so with a computer around it, you are paying for a pretty nice car.&#10;Buy a few of them to run agents in parallel, and you are building a small data center of your own.&lt;/p&gt;&#10;&lt;p&gt;The bars for hardware at home are electricity only, so they are the lowest these runs can cost: the cards come on top, and how much depends on how busy you keep them, which is what section 4 is about.&lt;/p&gt;&#10;&lt;p&gt;What I care about is the cheapest way to get answers of Qwen&amp;rsquo;s quality, from whichever model gives them.&lt;/p&gt;&#10;&lt;h2 id="2-but-my-gpus-are-already-paid-for"&gt;2. &amp;ldquo;But my GPUs are already paid for&amp;rdquo; &lt;a class="anchor" href="#2-but-my-gpus-are-already-paid-for" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;This is the argument I hear most, so assume the hardware is free and only the electricity counts.&lt;/p&gt;&#10;&lt;p&gt;I measured my own machine.&#10;It runs Qwen in vLLM, split across both cards, at about 107 tokens per second.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&#10;The cards are capped at 270 W each,&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt; and the whole machine draws about 700 W.&#10;It then needs four weeks, running day and night, to get through the benchmark once.&#10;I also give my quantized copy the score Artificial Analysis measured through Alibaba&amp;rsquo;s API.&#10;I have not rerun the benchmark on it, and anything lost to quantization makes my machine look worse.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-electricityCost-4" class="plot-container" role="img" aria-label="Electricity for one benchmark run of Qwen3.8 27B on my two 3090s, by electricity price. The dashed line is what Luna charges for the same work, with OpenAI&amp;rsquo;s hardware and margin included. Hover for values."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Electricity for one benchmark run of Qwen3.8 27B on my two 3090s, by electricity price. The dashed line is what Luna charges for the same work, with OpenAI&amp;rsquo;s hardware and margin included. Hover for values.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-electricityCost-4");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["electricityCost"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Above 14.5 cents per kWh, my electricity alone costs more than Luna&amp;rsquo;s API bill.&#10;Below it, I save a few dollars per run and wait four weeks instead of a few hours.&lt;/p&gt;&#10;&lt;p&gt;Those four weeks matter more than the few dollars.&#10;An API takes hundreds of requests at once, so even one benchmark run can finish in an hour or two.&#10;My setup works on one conversation at a time: when I sent it eight requests at once, seven of them waited in line.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-wallClock-5" class="plot-container" role="img" aria-label="Wall-clock time for one benchmark run. The API times use Artificial Analysis&amp;rsquo;s measured time per task for Luna; the parallel case assumes 100 requests at once. Log scale."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Wall-clock time for one benchmark run. The API times use Artificial Analysis&amp;rsquo;s measured time per task for Luna; the parallel case assumes 100 requests at once. Log scale.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-wallClock-5");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["wallClock"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;This is one more reason my coding agents run on APIs.&#10;When I &lt;a href="https://www.nijho.lt/post/parallel-agentic-coding/"&gt;run several agents in parallel&lt;/a&gt;, each of them should be as fast as if it were alone.&lt;/p&gt;&#10;&lt;h2 id="3-why-datacenters-are-so-much-better-at-this"&gt;3. Why datacenters are so much better at this &lt;a class="anchor" href="#3-why-datacenters-are-so-much-better-at-this" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;When my 3090s generate a token, they read all of the model&amp;rsquo;s weights from memory to produce one token for one conversation.&#10;Most of the chip&amp;rsquo;s compute sits idle while it waits for memory.&#10;A datacenter GPU reads the same weights once and produces a token for hundreds of conversations at the same time.&#10;This is called batching, and it is where most of the efficiency comes from.&lt;/p&gt;&#10;&lt;p&gt;Artificial Analysis measures this with &lt;a href="https://artificialanalysis.ai/hardware-inference-stack/datacenter" target="_blank" rel="noopener"&gt;AgentPerf&lt;/a&gt;, which replays real coding-agent sessions against datacenter hardware.&#10;It counts how many agents a system can serve while each one still gets at least 60 tokens per second.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-agentsPerKw-7" class="plot-container" role="img" aria-label="Coding agents served per kW of accelerator power, each getting at least 60 tokens per second, from AgentPerf. The datacenter systems all run DeepSeek V4 Pro. The 3090 bar is my machine running Qwen3.8 27B, a much smaller model, with GPU power only, so the comparison favors my machine."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Coding agents served per kW of accelerator power, each getting at least 60 tokens per second, from AgentPerf. The datacenter systems all run DeepSeek V4 Pro. The 3090 bar is my machine running Qwen3.8 27B, a much smaller model, with GPU power only, so the comparison favors my machine.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-agentsPerKw-7");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["agentsPerKw"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;The rack of 36 GB300s serves 20 times more agents per kW than the 8-GPU H200 server.&#10;Part of that is newer chips, and part is software: NVIDIA tuned the B300 and GB300 setups itself, while Artificial Analysis configured the H200 one.&#10;The B300 and the GB300 share both the chip generation and the software.&#10;Most of the gap between those two bars comes from the rack itself: 36 GPUs on one fast interconnect keep more than a thousand agents in flight at once.&#10;My two 3090s serve one conversation at a time.&lt;/p&gt;&#10;&lt;h2 id="4-but-i-have-solar-panels"&gt;4. &amp;ldquo;But I have solar panels&amp;rdquo; &lt;a class="anchor" href="#4-but-i-have-solar-panels" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Solar power is not free.&#10;With net metering, every kWh my GPUs burn is a kWh that does not get credited at the retail price.&#10;Without net metering, it is a kWh I do not sell, unless the surplus would go to waste anyway, which is the free-power case below.&#10;And the GPUs also run at night.&lt;/p&gt;&#10;&lt;p&gt;But fine, say the electricity is free and the only cost is the hardware.&#10;I was lucky and bought my two 3090s for about $750 each,&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt; and I write them off over three years.&#10;The GPUs only save money while they do work I would otherwise pay an API for, and while they sit idle they save nothing.&#10;So the question is how many hours a day they have to be busy before they pay for themselves.&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;&#10;&lt;p&gt;To make this fair to open models, imagine an open model that is exactly as efficient as Luna: the same quality from the same number of tokens, sold at Luna&amp;rsquo;s price.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-breakEven-8" class="plot-container" role="img" aria-label="How many hours a day my two $750 RTX 3090s have to be busy, every day for three years, before they cost less than a ZDR API. The host PC is not counted."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;How many hours a day my two $750 RTX 3090s have to be busy, every day for three years, before they cost less than a ZDR API. The host PC is not counted.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-breakEven-8");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["breakEven"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;With free power, the GPUs pay for themselves if they are busy 4 hours a day, every day for three years.&#10;At 18 cents per kWh, they need 8.5.&#10;That uses what I paid for the cards; at today&amp;rsquo;s price of about $1,500 each, the 8.5 hours become 14.&lt;/p&gt;&#10;&lt;p&gt;And API prices keep falling.&#10;GPT-6 Luna costs 58% less per output token than GPT-5.6 Luna did in July.&#10;If the price halves again, no amount of use pays off at 18 cents per kWh.&#10;If it halves twice, the API costs less than my electricity alone.&lt;/p&gt;&#10;&lt;h2 id="5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;5. &amp;ldquo;But my Mac, DGX Spark, or AI Max does save money&amp;rdquo; &lt;a class="anchor" href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Several readers told me their machines do save them money: a MacBook Pro with an M5 Max, an NVIDIA DGX Spark, and a computer with AMD&amp;rsquo;s Ryzen AI Max+ 395.&#10;If you would have bought the machine anyway, that is true, for the reason in section 2: their electricity costs little compared with the API.&#10;Counting the hardware is a different story.&lt;/p&gt;&#10;&lt;p&gt;With 128 GB of memory, all three can hold a better model than my 3090s: Qwen3.8-Flash-Next, a sparse model with 180 billion parameters, of which only 6 billion work on each token.&#10;It scores 39.8, close to GPT-6 Luna on its highest setting (38.1), although at the Q2 a reader uses on their Mac it probably loses a few points.&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt;&#10;So how many years would each machine have to run nonstop before it pays for itself?&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-paybackYears-9" class="plot-container" role="img" aria-label="Years of nonstop generation before each machine pays for itself, compared with the most efficient API model at a similar score. Each machine runs the best model it can hold."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Years of nonstop generation before each machine pays for itself, compared with the most efficient API model at a similar score. Each machine runs the best model it can hold.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-paybackYears-9");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["paybackYears"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Even with free power, the 128 GB machines need six to ten years.&#10;The AMD machine draws only about 120 W, but it is slow too: one benchmark run of Qwen3.8 27B takes it four and a half months, and almost as much energy as my 3090s.&#10;The RTX PRO 6000 runs the same model about three times as fast but costs about twice as much, so it needs six to nine years.&#10;My 3090s pay off in under two years with free power, which is why the 3090 is still the value king, but at 18 cents per kWh they never do: the electricity for one run already costs more than Luna charges for it.&lt;/p&gt;&#10;&lt;p&gt;Ten years is a long time in AI.&#10;Better and more efficient open models will run on the same machines, but the frontier moves at the same pace, and API prices keep falling.&#10;Hardware gets more efficient too, in datacenters and at home, so a machine bought today also ends up competing with next year&amp;rsquo;s machines.&#10;These numbers only show how far apart the two sides start.&lt;/p&gt;&#10;&lt;h2 id="6-the-bar-keeps-moving"&gt;6. The bar keeps moving &lt;a class="anchor" href="#6-the-bar-keeps-moving" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Section 4&amp;rsquo;s eight and a half hours a day would be easy for me: I run agents for longer than that, just not on Qwen.&#10;When Claude Opus 4.6 came out in February, I was perfectly happy with it.&#10;I thought it was all I would ever need, and I could not have imagined how much better models would get in half a year.&#10;Qwen3.8 27B now scores about the same as Opus 4.6, and I would no longer accept it for coding.&#10;Until GPT-6 Astra came out at the start of September, my go-to was GPT-5.6 Sol, which scores ten points higher than Qwen.&#10;I then used Astra until Claude Opus 5.5 came out less than three weeks later.&#10;Now I don&amp;rsquo;t even accept what was considered the best model a month ago.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-frontierGap-10" class="plot-container" role="img" aria-label="Intelligence Index of the best model available on each date, of the models I used as my go-to, and of Qwen&amp;rsquo;s 27B models, which fit on a single 3090. Scores use the xhigh reasoning setting where a model has it."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Intelligence Index of the best model available on each date, of the models I used as my go-to, and of Qwen&amp;rsquo;s 27B models, which fit on a single 3090. Scores use the xhigh reasoning setting where a model has it.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-frontierGap-10");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["frontierGap"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;The models that fit on my 3090s keep improving, but they stay about half a year behind the frontier, and my standard moves with the frontier.&lt;/p&gt;&#10;&lt;h2 id="7-a-small-company-has-it-worse"&gt;7. A small company has it worse &lt;a class="anchor" href="#7-a-small-company-has-it-worse" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;For me, this is a hobby, and a hobby is allowed to be inefficient.&#10;It gets worse once you try to use self-hosted models seriously, say for a team of ten developers.&lt;/p&gt;&#10;&lt;p&gt;If you stay fully self-hosted and want fast responses at peak, you have to buy hardware for the busiest hour of the year, not for the average one.&#10;The rest of the time, it sits idle.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-peakLoad-11" class="plot-container" role="img" aria-label="An illustrative week of coding agents for a ten-person team. The hardware has to cover the busiest hour of the year, but a typical week uses only a fraction of it."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;An illustrative week of coding agents for a ten-person team. The hardware has to cover the busiest hour of the year, but a typical week uses only a fraction of it.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-peakLoad-11");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["peakLoad"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;In this made-up but realistic week, the team uses 17% of what it paid for, so every token costs about six times more than it would at full load.&#10;You also need a spare GPU for when one dies, and someone who gets paged when it does.&#10;You could queue work or send the peaks to an API, but then you are paying for an API anyway.&lt;/p&gt;&#10;&lt;p&gt;An API provider has the opposite situation.&#10;It serves thousands of customers across every time zone, so its load curve is much flatter than yours.&#10;It fills the nights with discounted batch jobs (Luna&amp;rsquo;s batch tier is half price) and training runs.&#10;And more concurrent requests mean bigger batches, which is where the efficiency from section 3 comes from.&lt;/p&gt;&#10;&lt;h2 id="8-they-lose-money-on-inference-so-prices-will-go-up"&gt;8. &amp;ldquo;They lose money on inference, so prices will go up&amp;rdquo; &lt;a class="anchor" href="#8-they-lose-money-on-inference-so-prices-will-go-up" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Another argument I hear is that API prices are subsidized by investors and will go up once those investors want their money back.&#10;I don&amp;rsquo;t think that holds for inference.&lt;/p&gt;&#10;&lt;p&gt;AgentPerf also reports how many tokens each GPU serves.&#10;On the GB300 rack, one GPU serving DeepSeek V4 Pro handles about 4 million output tokens and 530 million input tokens per hour.&#10;Most of that input is conversation history that coding agents send again with every step.&#10;Providers keep it in a cache and charge less than 1% of the normal input price for it.&lt;/p&gt;&#10;&lt;p&gt;At &lt;a href="https://artificialanalysis.ai/models/deepseek-v4-pro-0424" target="_blank" rel="noopener"&gt;DeepSeek V4 Pro&amp;rsquo;s list prices&lt;/a&gt;, that GPU brings in at least $5.70 per hour, even if every input token is billed at the cache price.&#10;If 5% of the input misses the cache, it brings in about $17.&#10;Renting a B200, the closest GPU with a &lt;a href="https://artificialanalysis.ai/hardware-inference-stack/datacenter" target="_blank" rel="noopener"&gt;published rental price&lt;/a&gt;, costs $3.50 to $5.90 per hour from smaller cloud providers, and that price already includes their profit.&lt;/p&gt;&#10;&lt;p&gt;So the tokens pay for the hardware that serves them, even in the worst case.&#10;This leaves out costs like staff and training, so it does not tell you whether a lab makes money overall.&lt;/p&gt;&#10;&lt;p&gt;The money goes to training, research, free users, and flat-rate subscriptions for heavy users like me.&#10;When I wrote that I used &lt;a href="https://www.nijho.lt/post/agentic-coding/"&gt;$10,000 worth of API tokens for $200&lt;/a&gt;, that was list price, not what those tokens cost to serve.&lt;/p&gt;&#10;&lt;p&gt;And prices go down, not up.&#10;This is what the same level of intelligence has cost this year:&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-priceDrop-13" class="plot-container" role="img" aria-label="Output price per million tokens for models scoring between 30 and 35 on the Intelligence Index, by release date, as listed by Artificial Analysis. Log scale."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Output price per million tokens for models scoring between 30 and 35 on the Intelligence Index, by release date, as listed by Artificial Analysis. Log scale.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-priceDrop-13");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["priceDrop"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;From Claude Opus 4.6 in February to GPT-6 Luna in September, the output price for this level of intelligence dropped by a factor of 50.&#10;Labs are in a race to the bottom on price.&lt;/p&gt;&#10;&lt;h2 id="9-but-they-sell-your-data"&gt;9. &amp;ldquo;But they sell your data&amp;rdquo; &lt;a class="anchor" href="#9-but-they-sell-your-data" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Some people say APIs are cheap because the provider trains on your data or sells it.&#10;That is why I only used prices from providers that OpenRouter lists as &lt;a href="https://openrouter.ai/docs/features/zdr" target="_blank" rel="noopener"&gt;zero data retention&lt;/a&gt; (ZDR).&lt;/p&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Model&lt;/th&gt;&#10;&lt;th&gt;Provider&lt;/th&gt;&#10;&lt;th&gt;ZDR&lt;/th&gt;&#10;&lt;th&gt;Output ($ per million tokens)&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;GPT-6 Luna&lt;/td&gt;&#10;&lt;td&gt;OpenAI&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;0.50&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;GPT-6 Luna&lt;/td&gt;&#10;&lt;td&gt;Azure&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;0.50&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;AkashML (FP8)&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;1.78&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;DeepInfra (full precision)&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;1.88&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;Alibaba&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;2.55&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;Cloudflare&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;3.20&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Azure serves GPT-6 Luna with zero data retention at the same price as OpenAI.&#10;For Qwen3.8 27B, the cheapest endpoints are all ZDR, and the endpoints without ZDR cost more.&#10;The &lt;code&gt;:free&lt;/code&gt; tier is where you pay with your data.&lt;/p&gt;&#10;&lt;h2 id="10-where-self-hosting-does-win-the-same-model"&gt;10. Where self-hosting does win: the same model &lt;a class="anchor" href="#10-where-self-hosting-does-win-the-same-model" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;If you compare running Qwen3.8 27B yourself with paying for Qwen3.8 27B through an API, self-hosting wins.&lt;/p&gt;&#10;&lt;p&gt;DeepInfra, the cheapest full-precision ZDR provider, charges $619 for one benchmark run, while my electricity costs $83.&#10;That compares the full model with my quantized copy, so the gap overstates my advantage by whatever quantization costs in quality.&#10;My machine pays for itself if it is busy 2.4 hours a day, or 1.5 hours with free power (the first group in the chart in section 4).&lt;/p&gt;&#10;&lt;p&gt;Providers charge a lot for a dense 27B model, because every token runs through all 27 billion parameters.&#10;Sparse open models, which use only a small part of their weights for each token, can be much cheaper to serve.&#10;If you specifically need Qwen and use it heavily, self-hosting it pays off.&#10;But you don&amp;rsquo;t need Qwen to get answers of Qwen&amp;rsquo;s quality.&#10;Luna gives you those for $67, with zero data retention.&lt;/p&gt;&#10;&lt;h2 id="11-so-why-do-i-do-it"&gt;11. So why do I do it? &lt;a class="anchor" href="#11-so-why-do-i-do-it" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;First, it is fun.&#10;I like knowing how the whole stack works, from the drivers in &lt;a href="https://www.nijho.lt/post/llama-nixos/"&gt;my NixOS configuration&lt;/a&gt; to how the layers are split between my two GPUs.&lt;/p&gt;&#10;&lt;p&gt;Second, nobody can take it away.&#10;Artificial Analysis already lists GPT-5.6 Luna as deprecated, less than three months after its release.&#10;My Qwen weights will still be on my disk in ten years, and no company can change their license or &lt;a href="https://www.nijho.lt/post/truenas-to-nixos/"&gt;close their build system&lt;/a&gt; on me.&lt;/p&gt;&#10;&lt;p&gt;Third, some data should not leave my house.&#10;I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not.&#10;That is what &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;my local AI projects&lt;/a&gt; are for.&lt;/p&gt;&#10;&lt;p&gt;Those are good reasons.&#10;Saving money is not one of them.&lt;/p&gt;&#10;&lt;h2 id="appendix-all-assumptions"&gt;Appendix: all assumptions &lt;a class="anchor" href="#appendix-all-assumptions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;details&gt;&#10; &lt;summary&gt;Show all assumptions, and which side each one favors&lt;/summary&gt;&#10; &lt;p&gt;&lt;strong&gt;Assumptions that favor self-hosting&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The host PC around the GPUs is not counted, and my 3090s count at the $750 each I paid, not the $1,500 they sell for today.&lt;/li&gt;&#10;&lt;li&gt;Quantized local copies get the score Artificial Analysis measured through an API, with no penalty for quantization.&lt;/li&gt;&#10;&lt;li&gt;The payback chart assumes each machine is busy 24 hours a day, with no idle time.&lt;/li&gt;&#10;&lt;li&gt;API prices stay where they are, although GPT-6 Luna&amp;rsquo;s output price is 58% lower than GPT-5.6 Luna&amp;rsquo;s, two and a half months earlier.&lt;/li&gt;&#10;&lt;li&gt;My time to set up and maintain the machine is not counted.&lt;/li&gt;&#10;&lt;li&gt;Every machine also gets a case with free electricity.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;strong&gt;Assumptions that favor the API&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The example electricity price is 18 cents per kWh, and some people pay less.&lt;/li&gt;&#10;&lt;li&gt;My machine serves one request at a time; batching several requests could raise its total throughput.&lt;/li&gt;&#10;&lt;li&gt;The hardware has no resale value at the end.&lt;/li&gt;&#10;&lt;li&gt;Waste heat that replaces heating in winter is not counted.&lt;/li&gt;&#10;&lt;li&gt;API outages, rate limits, and deprecated models are not counted.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;strong&gt;Inputs&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Benchmark: the Artificial Analysis Intelligence Index on September 30, 2026, at the xhigh reasoning setting, or the highest setting where noted. Its token counts per model are the unit of work.&lt;/li&gt;&#10;&lt;li&gt;API prices: the cheapest zero-data-retention endpoints on OpenRouter, with cached input billed at the cache price.&lt;/li&gt;&#10;&lt;li&gt;My machine: two RTX 3090s capped at 270 W, running Qwen3.8 27B in vLLM with AutoRound INT4 and DFlash2. Measured: 107 tokens per second, 1,600 tokens per second of prefill, and 540 W for both GPUs. Estimated: about 700 W for the whole machine under load and 150 W idle.&lt;/li&gt;&#10;&lt;li&gt;Break-even: the cards are written off over three years; busy hours save what the API would charge minus the electricity, and idle hours cost the idle power.&lt;/li&gt;&#10;&lt;li&gt;Other machines: prices, power, and speeds as in the footnotes. The M5 Max and RTX PRO 6000 numbers come from readers, and the DGX Spark and AMD Ryzen AI Max+ 395 speeds are estimated from Artificial Analysis measurements.&lt;/li&gt;&#10;&lt;li&gt;Datacenter: AgentPerf measures accelerator power only, and the B200 rental price stands in for the GB300, which has no published rental price.&lt;/li&gt;&#10;&lt;li&gt;Small company: the peak-load week is made up.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&#10;&lt;/details&gt;&#10;&#10;&lt;div class="footnotes" role="doc-endnotes"&gt;&#10;&lt;hr&gt;&#10;&lt;ol&gt;&#10;&lt;li id="fn:1"&gt;&#10;&lt;p&gt;For each model, Artificial Analysis publishes how many input and output tokens the whole index took, and how many of the input tokens were read from a cache. I multiplied those by each provider&amp;rsquo;s prices, with cached input at the cache price. For Qwen that is 198 million output tokens and 820 million input tokens that miss the cache. For the bars at home, the same tokens divided by the machine&amp;rsquo;s speed give the run time, with the uncached input processed at the 1,600 tokens per second I measured on my 3090s and an estimated 2,000 on the RTX PRO 6000, and the run time times the machine&amp;rsquo;s power and the price per kWh gives the electricity.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:2"&gt;&#10;&lt;p&gt;The RTX PRO 6000 has the same memory bandwidth as an RTX 5090, and at full precision every token reads three times as many bytes as in Q4_K_M. Scaling Artificial Analysis&amp;rsquo;s 5090 measurement gives about 48 tokens per second. I assume 600 W for the whole machine.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:3"&gt;&#10;&lt;p&gt;I measured the model as I normally run it: vLLM with an AutoRound INT4 quant and DFlash2 speculative decoding, split across both cards. Single 1,500-token answers came out at 105 to 109 tokens per second, and reading a 22,000-token prompt ran at 1,600 tokens per second. Sending 2, 4, or 8 requests at once did not raise the total above about 110 tokens per second. The two GPUs drew 540 W together, each at its 270 W power limit; the 700 W adds my estimate for the rest of the machine, since I could not measure the CPU. The setup follows the recipes from &lt;a href="https://github.com/noonghunna/club-3090" target="_blank" rel="noopener"&gt;club-3090&lt;/a&gt;, a community project that tunes LLM serving for RTX 3090s and publishes measured numbers, so I think it is about as fast as these cards get. Its benchmarks for this configuration at a similar power limit match mine: about 110 tokens per second for prose, and up to about 195 for code, where speculative decoding guesses more tokens right. Most of Qwen&amp;rsquo;s output on the benchmark is reasoning, which runs at the prose speed.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:4"&gt;&#10;&lt;p&gt;The stock limit is 350 W. My cards sit in a &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;normal desktop case with normal fans&lt;/a&gt;, and I worry that running both at full power for hours would shorten their life. The cap is &lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/nvidia-undervolt.nix#L1-L7" target="_blank" rel="noopener"&gt;a few lines in my NixOS configuration&lt;/a&gt;, and measurements from Puget Systems and r/LocalLLaMA put a 3090 at about 95% of its speed at 270 W. That is 77% of the power for 95% of the speed, so the cap also lowers my electricity cost per token.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:5"&gt;&#10;&lt;p&gt;The 3090 is still the value king for VRAM per dollar, and $750 badly understates what mine are worth. That is what I paid more than a year ago; today a used 3090 sells for about $1,500. Even at that price it costs $62.50 per GB of VRAM, the same as an RTX 5090 at its $2,000 launch price, which is not what a 5090 sells for today. The right number for this calculation is what I could sell my cards for, and at $1,500 each the break-even points rise by about two thirds: at 18 cents per kWh, the same-model case moves from 2.4 to 4.0 hours per day, and the Luna-efficient case from 8.5 to 14.2 hours per day.&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:6"&gt;&#10;&lt;p&gt;While the cards are busy, they save what the API would have charged for the same work, minus the electricity at 700 W. While they are idle, the machine still draws about 150 W, of which the two GPUs with the model loaded take 85 W. The break-even point is the number of busy hours per day at which those savings cover the price of the cards, written off over three years, plus the idle power.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:7"&gt;&#10;&lt;p&gt;The M5 Max numbers come from the reader: Qwen3.8-Flash-Next at Q2 (about 2.7 bits per weight), about 45 tokens per second, 1,250 tokens per second of prefill, and about 90 W, on a $7,000 laptop. The RTX PRO 6000 numbers come from another reader, &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1ww2jsu/comment/pdhs3ea/" target="_blank" rel="noopener"&gt;u/TheAILegend&lt;/a&gt;, who measured Qwen3.8-Flash-Next at NVFP4 on theirs. For a single request, it stays at about 150 tokens per second from 4,000 to 200,000 tokens of context, with about 10,000 tokens per second of prefill and 266 W for the GPU plus about 160 W for the rest of the PC. With 32 short requests in parallel, the card reaches 590 tokens per second and pays off in about two years, which works for short batch jobs. At the long contexts of this benchmark, parallel requests are slower than one, so the chart uses the single-request speed. I use $15,000, the middle of the card&amp;rsquo;s price range. The DGX Spark and the AMD machine run Qwen3.8-Flash-Next at about 4.5 bits. For their speed, I took Artificial Analysis&amp;rsquo;s measurements of Qwen3.8 27B on each machine and doubled them, because the owner of the AMD machine finds Qwen3.8-Flash-Next about twice as fast as Qwen3.8 27B. That gives about 55 tokens per second on the DGX Spark ($6,950, about 150 W) and 47 on the AMD Ryzen AI Max+ 395 ($4,000, the price Artificial Analysis lists, about 120 W). All of them are compared with GPT-6 Luna on its highest setting, which runs the whole benchmark for $122. The 3090s run Qwen3.8 27B, as in the rest of this post, and are compared with Luna on xhigh. Artificial Analysis measured Qwen3.8 27B on the AMD machine at 23 tokens per second, so one benchmark run takes it 134 days and about 390 kWh, against about 460 kWh on my 3090s. Running Qwen3.8-Flash-Next on the 3090s with CPU expert offload, which club-3090 measured at 37 tokens per second, does not change their result: one run then takes 82 days, and at 18 cents per kWh its electricity costs more than Luna&amp;rsquo;s $122 unless the whole machine draws less than about 350 W.&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;/div&gt;&#10;</description></item><item><title>Exploring the Jev hype in MindRoom</title><link>https://www.nijho.lt/post/jev-mindroom/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/jev-mindroom/</guid><description>&lt;p&gt;Last week, every AI newsletter I get was full of &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener"&gt;Jev&lt;/a&gt;, a new model from TypeSafe.&#10;So were several subreddits.&#10;I am a permanent lurker on &lt;a href="https://www.reddit.com/r/LocalLLaMA/" target="_blank" rel="noopener"&gt;r/LocalLLaMA&lt;/a&gt;, and even though it is about local models, every other post seemed to be about Jev.&#10;There were local clones like &lt;a href="https://github.com/r-ms/mini-jev" target="_blank" rel="noopener"&gt;mini-jev&lt;/a&gt; and &lt;a href="https://github.com/wfzyx/von" target="_blank" rel="noopener"&gt;von&lt;/a&gt;, CLI wrappers like &lt;a href="https://github.com/joshLong145/jev-cli" target="_blank" rel="noopener"&gt;jev-cli&lt;/a&gt;, and even &lt;a href="https://news.ycombinator.com/item?id=49765348" target="_blank" rel="noopener"&gt;someone on Hacker News&lt;/a&gt; claiming &lt;a href="https://laya.convaiinnovations.com/" target="_blank" rel="noopener"&gt;they had built the same thing a year earlier&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;I tried Jev in &lt;a href="https://www.nijho.lt/post/mindroom/"&gt;MindRoom&lt;/a&gt;, my &lt;a href="https://github.com/mindroom-ai/mindroom" target="_blank" rel="noopener"&gt;open-source&lt;/a&gt; agent platform on &lt;a href="https://matrix.org" target="_blank" rel="noopener"&gt;Matrix&lt;/a&gt;, and less than two days later it was making &lt;a href="https://github.com/mindroom-ai/mindroom/issues/2156" target="_blank" rel="noopener"&gt;three decisions&lt;/a&gt; there.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Update 2026-10-07:&lt;/strong&gt; OpenAI released the Decisions API today, its own take on the same idea, and MindRoom now supports it as a third backend; see &lt;a href="#a-third-backend-openai-decisions"&gt;below&lt;/a&gt;.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="what-a-system-one-model-is"&gt;What a System One model is &lt;a class="anchor" href="#what-a-system-one-model-is" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The &lt;a href="https://docs.typesafe.ai/concepts/system-one" target="_blank" rel="noopener"&gt;name comes&lt;/a&gt; from Daniel Kahneman&amp;rsquo;s &lt;a href="https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow" target="_blank" rel="noopener"&gt;&lt;em&gt;Thinking, Fast and Slow&lt;/em&gt;&lt;/a&gt;, where System 1 is fast, intuitive thinking and System 2 is slow and deliberate.&#10;A System One model does not generate text.&#10;You send it a state, for example a chat log, plus typed questions, and it returns &lt;a href="https://docs.typesafe.ai/api" target="_blank" rel="noopener"&gt;probabilities&lt;/a&gt;:&lt;/p&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Type&lt;/th&gt;&#10;&lt;th&gt;Question&lt;/th&gt;&#10;&lt;th&gt;Answer&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Noul&lt;/td&gt;&#10;&lt;td&gt;Yes or no&lt;/td&gt;&#10;&lt;td&gt;Probability of yes&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Choice&lt;/td&gt;&#10;&lt;td&gt;Pick one of up to 255 options&lt;/td&gt;&#10;&lt;td&gt;Chosen option plus the full distribution&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Score&lt;/td&gt;&#10;&lt;td&gt;Rate against a rubric of 2 to 10 levels&lt;/td&gt;&#10;&lt;td&gt;Probability-weighted value&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;&lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener"&gt;TypeSafe claims&lt;/a&gt; 70 to 500 ms end to end and $0.042 per million input tokens, with output tokens free.&#10;They also write &amp;ldquo;we can&amp;rsquo;t prove it isn&amp;rsquo;t subsidized.&amp;rdquo;&#10;Either way, it is incredibly cheap and incredibly fast.&lt;/p&gt;&#10;&lt;h2 id="adaptive-participation"&gt;Adaptive participation &lt;a class="anchor" href="#adaptive-participation" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;In MindRoom, a conversation can have several humans, several agents, or both.&#10;When I talk to a single agent, it simply replies, and I don&amp;rsquo;t need to tag it.&#10;Once another person joins and writes something, the agent cannot tell whether the message was meant for it or for the other person.&#10;Nothing decided that, so the agent stayed quiet unless explicitly tagged.&#10;That confused several users.&lt;/p&gt;&#10;&lt;p&gt;So I recently added &lt;a href="https://docs.mindroom.chat/configuration/threads/#adaptive-participation" target="_blank" rel="noopener"&gt;adaptive participation&lt;/a&gt;: an agent that has already replied in a thread decides on its own whether to answer an untagged message.&#10;I first used an LLM as the judge, GPT-5.6 Luna on low reasoning.&#10;A yes-or-no question like &amp;ldquo;should this agent respond?&amp;rdquo; is exactly what Jev is for, so it &lt;a href="https://github.com/mindroom-ai/mindroom/pull/2169" target="_blank" rel="noopener"&gt;became the second backend&lt;/a&gt;.&#10;Compared with Luna, it is about ten times cheaper and ten times faster, and if I believe the benchmarks, it is also smarter.&lt;/p&gt;&#10;&lt;h2 id="thanks-should-not-stop-the-agent"&gt;&amp;ldquo;Thanks&amp;rdquo; should not stop the agent &lt;a class="anchor" href="#thanks-should-not-stop-the-agent" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The next thing I used Jev for was interruptions.&#10;When I send a message while an agent is still replying, MindRoom injects a notice at the next tool call telling the agent to stop and wrap up, because there is a new message.&#10;But often I, and other users, would watch the agent start working and then write &amp;ldquo;looks good&amp;rdquo; or &amp;ldquo;thanks.&amp;rdquo;&#10;The agent would stop early and continue in the next turn, which is both wasteful and counterintuitive.&lt;/p&gt;&#10;&lt;p&gt;Now &lt;a href="https://github.com/mindroom-ai/mindroom/pull/2193" target="_blank" rel="noopener"&gt;a judgment&lt;/a&gt; decides whether the new message needs the interruption, which the docs call &lt;a href="https://docs.mindroom.chat/configuration/threads/#mid-turn-coalescing" target="_blank" rel="noopener"&gt;mid-turn coalescing&lt;/a&gt;.&#10;If it does not, the agent reacts with 👀, finishes its reply, and handles the message afterwards.&lt;/p&gt;&#10;&lt;p&gt;This is also where I learned a lesson about evals.&#10;My first implementation worked at the code level but not functionally.&#10;My evals, based on real attempts inside MindRoom that had all failed plus synthetic cases from GPT-6 Astra, passed 5 out of 40 with Jev.&#10;&lt;a href="https://github.com/mindroom-ai/mindroom/pull/2244" target="_blank" rel="noopener"&gt;The fix&lt;/a&gt; changed the question and gave the judge more context, and it went to 40 out of 40.&lt;/p&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;&lt;/th&gt;&#10;&lt;th&gt;Before: 5/40&lt;/th&gt;&#10;&lt;th&gt;After: 40/40&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Question&lt;/td&gt;&#10;&lt;td&gt;&amp;ldquo;May the active task finish before the queued human messages are handled in a later turn?&amp;rdquo;&lt;/td&gt;&#10;&lt;td&gt;&amp;ldquo;Does any newly queued message require an immediate change to, pause of, or stop of the active task?&amp;rdquo;&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;When unclear&lt;/td&gt;&#10;&lt;td&gt;&amp;ldquo;Also false when context is incomplete, the relationship is unclear…&amp;rdquo;&lt;/td&gt;&#10;&lt;td&gt;&amp;ldquo;Mere relevance to the task does not make praise or thanks an interruption.&amp;rdquo;&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Context&lt;/td&gt;&#10;&lt;td&gt;The request and the new messages&lt;/td&gt;&#10;&lt;td&gt;The same, plus the preceding conversation&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;The old question treated anything unclear as a reason to interrupt, and a bare &amp;ldquo;thanks&amp;rdquo; without the conversation around it always looks unclear.&lt;/p&gt;&#10;&lt;h2 id="picking-the-responder"&gt;Picking the responder &lt;a class="anchor" href="#picking-the-responder" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The third decision is routing.&#10;When a message in a room with many agents does not mention anyone, MindRoom&amp;rsquo;s router picks who should answer.&#10;That used to be a free-text LLM prompt, &amp;ldquo;choose the most appropriate agent.&amp;rdquo;&#10;It is &lt;a href="https://github.com/mindroom-ai/mindroom/pull/2238" target="_blank" rel="noopener"&gt;now a Jev Choice&lt;/a&gt; over the eligible agents, plus a &lt;code&gt;no_fit&lt;/code&gt; option, as described in the &lt;a href="https://docs.mindroom.chat/configuration/router/" target="_blank" rel="noopener"&gt;router docs&lt;/a&gt;.&#10;A confident &lt;code&gt;no_fit&lt;/code&gt; asks the user to mention an agent instead of guessing, and anything uncertain falls back to the existing LLM router.&lt;/p&gt;&#10;&lt;h2 id="one-config-line-to-switch"&gt;One config line to switch &lt;a class="anchor" href="#one-config-line-to-switch" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I built a very simple abstraction layer so I can swap Jev for an LLM without changing any code.&#10;Every decision is one question with criteria for true and false:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;MID_TURN_QUESTION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;JudgmentQuestion&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;interrupt_current_turn&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;Does any newly queued message require an immediate change to, pause of, or stop of the active task?&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;when_true&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;A new message requests stopping or pausing the active task, corrects it, ...&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;when_false&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;The messages acknowledge progress, express thanks, ask to continue unchanged, ...&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Jev receives it as a Noul question and returns a probability, which MindRoom compares with a threshold.&#10;An LLM receives the same question and context and must return &lt;code&gt;{&amp;quot;decision&amp;quot;: true}&lt;/code&gt; or &lt;code&gt;{&amp;quot;decision&amp;quot;: false}&lt;/code&gt;.&#10;Switching is a configuration change:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;helper&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;mid_turn&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;defer_reaction&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;👀&amp;#34;&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;judgment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;typesafe &lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="c"&gt;# or: provider: llm, model: fast&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0.8&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both backends log their decisions and latency, so I can compare them on real traffic.&#10;The code is in &lt;a href="https://github.com/mindroom-ai/mindroom/tree/main/src/mindroom/judgment" target="_blank" rel="noopener"&gt;&lt;code&gt;src/mindroom/judgment/&lt;/code&gt;&lt;/a&gt;, and the backends are documented under &lt;a href="https://docs.mindroom.chat/configuration/threads/#judgment-backends" target="_blank" rel="noopener"&gt;participation judgment backends&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h2 id="a-third-backend-openai-decisions"&gt;A third backend: OpenAI Decisions &lt;a class="anchor" href="#a-third-backend-openai-decisions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;On October 7, OpenAI released its own version of the idea, the &lt;a href="https://developers.openai.com/api/docs/guides/decisions" target="_blank" rel="noopener"&gt;Decisions API&lt;/a&gt;, currently in public beta.&#10;The question types match Jev&amp;rsquo;s: a &lt;code&gt;predicate&lt;/code&gt; is a Noul, and &lt;code&gt;choice&lt;/code&gt; and &lt;code&gt;score&lt;/code&gt; keep their names.&#10;It serves a single model, &lt;code&gt;gpt-6-luna&lt;/code&gt;, and OpenAI says it returns answers &amp;ldquo;about 10x faster than the Responses API.&amp;rdquo;&#10;Input costs $0.10 per million tokens and output is free, so it lists at about 2.4 times Jev&amp;rsquo;s price.&lt;/p&gt;&#10;&lt;p&gt;Thanks to the abstraction layer, adding it took &lt;a href="https://github.com/mindroom-ai/mindroom/pull/2751" target="_blank" rel="noopener"&gt;one PR&lt;/a&gt;, merged the same day, and it works for all three decisions, including the router&amp;rsquo;s Choice:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;judgment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="l"&gt;openai_decisions&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0.8&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;From MindRoom&amp;rsquo;s side, only the wire format differs.&#10;The endpoint accepts only user messages, so MindRoom sends each chat message as its own user message, prefixed with its role, like &lt;code&gt;assistant: ...&lt;/code&gt;.&#10;Decisions can also answer with a refusal, which MindRoom treats as an abstention, so the usual fallback applies.&lt;/p&gt;&#10;&lt;p&gt;I tested it against the live API through MindRoom&amp;rsquo;s own code, with the same questions Jev gets.&#10;For participation, it gave 0.99 to replying to a direct question and 0.0 while two humans were talking to each other.&#10;Mid-turn, &amp;ldquo;Thanks, looks great so far!&amp;rdquo; got 0.0 for interrupting and &amp;ldquo;Stop, keep MySQL…&amp;rdquo; got 1.0.&#10;The router sent a pandas bug to &lt;code&gt;code&lt;/code&gt; and a meeting request to &lt;code&gt;calendar&lt;/code&gt;, both with confidence 1.0.&#10;A nonsense request got only 0.62, below the 0.8 threshold, so the normal LLM router took over.&#10;Each call took 130 to 280 ms.&#10;I have not compared it with Jev on real traffic yet, but both log the same outcome fields, so that only takes a config change.&lt;/p&gt;&#10;&lt;h2 id="named-after-jevons"&gt;Named after Jevons &lt;a class="anchor" href="#named-after-jevons" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I only read &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" target="_blank" rel="noopener"&gt;why TypeSafe named it Jev&lt;/a&gt; after building all this: they named it after William Stanley Jevons, who noticed that as steam engines burned coal more efficiently, demand for coal went up instead of down.&#10;This is known as the &lt;a href="https://en.wikipedia.org/wiki/Jevons_paradox" target="_blank" rel="noopener"&gt;Jevons paradox&lt;/a&gt;.&#10;That was literally my experience.&#10;Most of what I built was perfectly possible with an LLM.&#10;I just never built it, because I am somewhat cost- and latency-minded.&#10;Once I implemented Jev for one decision, I came up with use case after use case.&lt;/p&gt;&#10;&lt;p&gt;One I already built is a plugin, &lt;a href="https://github.com/mindroom-ai/response-audit-jev-plugin" target="_blank" rel="noopener"&gt;Response Audit JEV&lt;/a&gt;.&#10;Right after an agent replies, Jev checks the answer against the request and the tool calls the agent actually made, for example whether things are properly cited.&#10;If a check flags a problem, the plugin posts one follow-up in the thread that tags the agent and asks for a correction.&#10;Because Jev is so fast and cheap, running this after every reply costs almost nothing.&lt;/p&gt;&#10;&lt;p&gt;Right now I am full of ideas.&#10;For example, a model router.&#10;In MindRoom, an agent can switch to a different model mid-conversation with a tool call, and I can do the same with &lt;code&gt;!model&lt;/code&gt;.&#10;Agents rarely do that without being told, though.&#10;So I default to a very expensive model, and I only save money when I remember to switch to a cheap one for something simple.&#10;Jev could run on every message and decide whether it needs the capable model or not.&#10;That works in both directions, up to a capable model or down to a cheap one.&#10;The perfect moment to switch is when the prompt cache has expired.&#10;A prompt cache belongs to one model, so switching normally means sending the entire context again at full price.&#10;Once the cache has expired, I pay that price anyway, so switching costs nothing extra.&lt;/p&gt;&#10;</description></item><item><title>Gemini 3 Pro: First Impressions</title><link>https://www.nijho.lt/post/gemini-3-pro-first-impressions/</link><pubDate>Wed, 19 Nov 2025 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/gemini-3-pro-first-impressions/</guid><description>&lt;p&gt;I&amp;rsquo;m a huge fan of &lt;a href="https://www.nijho.lt/post/agentic-coding/"&gt;agentic&lt;/a&gt; &lt;a href="https://www.nijho.lt/post/agentic-mobile-workflow/"&gt;coding&lt;/a&gt;.&#10;I have the maximum subscriptions for both Claude and ChatGPT Pro, and since writing my &lt;a href="https://www.nijho.lt/post/agentic-coding/"&gt;agentic coding&lt;/a&gt; post, I mostly moved from Claude Code to Codex CLI with GPT-5 for my work that requires more complex reasoning.&#10;Using Claude Code often feels frustrating in comparison now.&lt;/p&gt;&#10;&lt;p&gt;I also frequently pit them against each other, drafting parallel plans with each tool (I&amp;rsquo;ve done this at least 50 times).&#10;The verdict is consistent.&#10;&lt;strong&gt;Claude Code, without exception, says that GPT‑5&amp;rsquo;s solution is superior&lt;/strong&gt;, admitting it is more elegant and robust.&#10;Claude also misses bugs during PR reviews that GPT-5 catches immediately.&lt;/p&gt;&#10;&lt;p&gt;That said, &lt;strong&gt;Claude Code is significantly faster&lt;/strong&gt;, which makes it my go-to for simple refactors.&#10;I also prefer its writing style for conversational answers and documentation.&lt;/p&gt;&#10;&lt;p&gt;But looking at the charts on &lt;strong&gt;&lt;a href="https://artificialanalysis.ai/" target="_blank" rel="noopener"&gt;Artificial Analysis&lt;/a&gt;&lt;/strong&gt;, the new &lt;strong&gt;Gemini 3 Pro&lt;/strong&gt; looks truly impressive.&#10;It&amp;rsquo;s at the top of nearly every benchmark.&#10;I&amp;rsquo;m excited for Gemini to potentially take over some of GPT‑5&amp;rsquo;s workload, so when &lt;strong&gt;Gemini 3 Pro Preview&lt;/strong&gt; came out yesterday, I had to try it.&lt;/p&gt;&#10;&lt;h2 id="the-good-raw-power"&gt;The good: raw power &lt;a class="anchor" href="#the-good-raw-power" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;First off, the model is quite powerful.&#10;Last night, I built a complex new feature for my &lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;&lt;code&gt;agent-cli&lt;/code&gt;&lt;/a&gt; project in just two hours: a &lt;strong&gt;RAG (Retrieval-Augmented Generation) proxy server&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;This tool acts as a middleware that sits between any OpenAI-compatible client and server (like Ollama or vLLM).&#10;The server watches a specific folder, automatically indexes any PDF, Markdown, or text files into a local ChromaDB vector store, and injects relevant context into your prompts transparently.&#10;You can see the full implementation in &lt;a href="https://github.com/basnijholt/agent-cli/pull/79" target="_blank" rel="noopener"&gt;PR #79&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;It even handles dynamic re-indexing when I drop new files in!&#10;The speed and code quality for this task were genuinely impressive, and it handled the complexity of the RAG integration surprisingly well.&lt;/p&gt;&#10;&lt;h2 id="the-bad-panic-attacks-and-infinite-loops"&gt;The bad: panic attacks and infinite loops &lt;a class="anchor" href="#the-bad-panic-attacks-and-infinite-loops" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;However, the experience wasn&amp;rsquo;t without its &amp;ldquo;personality quirks.&amp;rdquo;&#10;Twice, Gemini got stuck in a loop saying &lt;em&gt;&amp;ldquo;Should I? I should&amp;rdquo;&lt;/em&gt; over and over again.&#10;Or the following also happened:&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://www.nijho.lt/post/gemini-3-pro-first-impressions/infinite-loop_hu_234c7407851f9273.webp" srcset="https://www.nijho.lt/post/gemini-3-pro-first-impressions/infinite-loop_hu_8f9dd754d34e9bda.webp 272w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/infinite-loop_hu_234c7407851f9273.webp 516w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/infinite-loop_hu_3023867168cf93e2.webp 815w" sizes="auto, (max-width: 760px) 100vw, 720px" width="516" height="760" alt="Gemini infinite loop" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;p&gt;It also has a bad habit of amending commits multiple times and force-pushing them.&#10;For me the Git history is a critical source of truth, and I don&amp;rsquo;t like rewriting it because I might just have reviewed those commits already.&lt;/p&gt;&#10;&lt;h3 id="the-interface-friction"&gt;The interface friction &lt;a class="anchor" href="#the-interface-friction" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I also found the Gemini CLI interface frustrating.&#10;Basic tasks like selecting text or scrolling up are surprisingly difficult.&#10;I&amp;rsquo;m not sure why they chose this design or why they feel the need to deviate from standard CLI behavior.&#10;It feels like unnecessary friction in an otherwise promising tool.&lt;/p&gt;&#10;&lt;h3 id="proofreading-this-post"&gt;Proofreading this post &lt;a class="anchor" href="#proofreading-this-post" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I also asked Gemini to proofread this very blog post.&#10;It made some suggestions—some I liked, others I didn&amp;rsquo;t, so I reverted the ones I didn&amp;rsquo;t want.&#10;But in every single subsequent iteration, it would forcefully re-apply its rejected changes, adding emojis I had removed or rewriting sentences I had explicitly restored.&lt;/p&gt;&#10;&lt;h3 id="the-main-branch-incident"&gt;The &amp;ldquo;main branch&amp;rdquo; incident &lt;a class="anchor" href="#the-main-branch-incident" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Another scary moment: I asked if the PR was ready to be merged into &lt;code&gt;main&lt;/code&gt;, and it went ahead and checked out the &lt;code&gt;main&lt;/code&gt; branch and merged everything locally.&#10;If I hadn&amp;rsquo;t stopped it, I&amp;rsquo;m pretty sure it would have pushed the changes to the &lt;code&gt;main&lt;/code&gt; branch directly as well.&#10;I tend to use &amp;ldquo;yolo mode&amp;rdquo; for agentic coding because I feel like Git has my back, but this feels like a dangerous path to go down.&lt;/p&gt;&#10;&lt;p&gt;And indeed, while writing this post, &lt;strong&gt;it actually happened&lt;/strong&gt;.&#10;Gemini merged my pull request without my permission.&lt;/p&gt;&#10;&lt;p&gt;My prompt was &lt;code&gt;&amp;gt; should the README reflect this?&lt;/code&gt; and before that I asked to review the PR.&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://www.nijho.lt/post/gemini-3-pro-first-impressions/merge_hu_719c4aa620a5a59a.webp" srcset="https://www.nijho.lt/post/gemini-3-pro-first-impressions/merge_hu_44f4ede1b23e43f5.webp 400w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/merge_hu_719c4aa620a5a59a.webp 760w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/merge_hu_fcec307dabb3ec8f.webp 1200w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="326" alt="Gemini merging without permission" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;p&gt;When I asked &lt;em&gt;why&lt;/em&gt; it merged it, it instantly force-pushed to &lt;code&gt;main&lt;/code&gt; again!&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://www.nijho.lt/post/gemini-3-pro-first-impressions/force-push-again_hu_a7f66199c342259b.webp" srcset="https://www.nijho.lt/post/gemini-3-pro-first-impressions/force-push-again_hu_3884bda6d7ac695f.webp 400w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/force-push-again_hu_a7f66199c342259b.webp 760w, https://www.nijho.lt/post/gemini-3-pro-first-impressions/force-push-again_hu_8a797e9907cbcd9e.webp 1200w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="379" alt="Gemini force pushing to main" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&lt;h3 id="artifacts-and-existential-crises"&gt;Artifacts and existential crises &lt;a class="anchor" href="#artifacts-and-existential-crises" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Also, when it commits, it just adds &lt;em&gt;all&lt;/em&gt; files, including debugging logs and build artifacts.&#10;When it realizes its mistake, it tries to fix it.&#10;But if I try to help at the same time (like removing the files myself while it&amp;rsquo;s thinking), it gets extremely confused.&lt;/p&gt;&#10;&lt;p&gt;It spirals into an &lt;strong&gt;existential crisis&lt;/strong&gt;, claiming that I committed the file but now &amp;ldquo;it is gone.&amp;rdquo;&#10;&lt;em&gt;&amp;ldquo;What is going on?&amp;rdquo;&lt;/em&gt; it asks, spending a lot of tokens (and time) looking at &lt;code&gt;git reflog&lt;/code&gt;s and questioning its own existence.&#10;Even when I stop it and explicitly say that I deleted the file and to please continue, it completely ignores me and continues its panic attack.&#10;I experienced this pattern of ignoring instructions multiple times.&lt;/p&gt;&#10;&lt;h2 id="conclusion"&gt;Conclusion &lt;a class="anchor" href="#conclusion" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Gemini 3 Pro shows massive potential.&#10;The raw coding ability is there, and building a &lt;strong&gt;clean&lt;/strong&gt; RAG proxy in two hours is great!&#10;But the agentic behavior needs some serious work: the loops, the force pushes, ignoring instructions, and the confusion when reality doesn&amp;rsquo;t match its internal state.&lt;/p&gt;&#10;&lt;p&gt;This is just &lt;strong&gt;day one&lt;/strong&gt;, so I&amp;rsquo;m confident many of these issues will get ironed out.&#10;I don&amp;rsquo;t love the interface, but given that &lt;a href="https://github.com/google-gemini/gemini-cli" target="_blank" rel="noopener"&gt;&lt;code&gt;gemini-cli&lt;/code&gt;&lt;/a&gt; is open source with ≈400 contributors, I expect rapid improvements, especially now that the model is finally good with agentic workflows!&lt;/p&gt;&#10;</description></item><item><title>On agentic coding 🤖</title><link>https://www.nijho.lt/post/agentic-coding/</link><pubDate>Mon, 25 Aug 2025 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/agentic-coding/</guid><description>&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#1-from-skepticism-to-productivity"&gt;1. From skepticism to productivity&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#2-but-dont-you-love-programming"&gt;2. But don&amp;rsquo;t you love programming?&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-power-to-explore"&gt;The power to explore&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-matty-story-from-midnight-idea-to-working-code"&gt;The Matty story: from midnight idea to working code&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#clip-files-ui-when-ideas-strike-at-random-moments"&gt;clip-files UI: when ideas strike at random moments&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-800-rule-marathon-80-ai-agents-at-once"&gt;The 800-rule marathon: 80 AI agents at once&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#1400-commits-rewritten-in-minutes-for-250"&gt;1,400 commits, rewritten in minutes for $2.50&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#3-the-two-phase-ai-revolution"&gt;3. The two-phase AI revolution&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#phase-1-copy-paste-era-march-2023---july-2025"&gt;Phase 1: copy-paste era (March 2023 - July 2025)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#phase-2-the-claude-code-revolution-july-2025"&gt;Phase 2: the Claude Code revolution (July 2025)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#what-made-claude-code-different"&gt;What made Claude Code different?&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#4-my-current-agentic-workflow"&gt;4. My current agentic workflow&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#parallel-development-6-features-at-once"&gt;Parallel development: 6 features at once&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#starting-a-new-package"&gt;Starting a new package&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-key-difference-iteration-and-validation"&gt;The key difference: iteration and validation&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#5-the-evidence-explosive-productivity"&gt;5. The evidence: explosive productivity&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#package-creation-by-year"&gt;Package creation by year&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#recent-packages-built-with-agentic-ai"&gt;Recent packages built with agentic AI&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#6-why-this-isnt-vibe-coding-but-with-nuance"&gt;6. Why this isn&amp;rsquo;t &amp;ldquo;vibe coding&amp;rdquo; (but with nuance)&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#my-context-driven-standards"&gt;My context-driven standards&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-internal-dependency-graph-principle"&gt;The internal dependency graph principle&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#building-constraints-around-ai"&gt;Building constraints around AI&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#7-common-pitfalls-ive-noticed"&gt;7. Common pitfalls I&amp;rsquo;ve noticed&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-defensive-programming-trap"&gt;The &amp;ldquo;defensive programming&amp;rdquo; trap&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-backwards-compatibility-obsession"&gt;The &amp;ldquo;backwards compatibility&amp;rdquo; obsession&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-over-engineering-disease"&gt;The &amp;ldquo;over-engineering&amp;rdquo; disease&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-git-commit-sins"&gt;The git commit sins&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-helpful-anti-patterns"&gt;The &amp;ldquo;helpful&amp;rdquo; anti-patterns&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-mission-accomplished-hallucination"&gt;The &amp;ldquo;mission accomplished&amp;rdquo; hallucination&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-nuclear-option"&gt;The nuclear option&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-debugging-debris"&gt;The debugging debris&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#why-this-happens"&gt;Why this happens&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#8-critical-success-factors-what-actually-makes-this-work"&gt;8. Critical success factors: what actually makes this work&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#teach-it-your-test-commands-day-1-priority"&gt;Teach it your test commands (day 1 priority)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#git-commits-are-your-safety-net"&gt;Git commits are your safety net&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#maintain-healthy-skepticism"&gt;Maintain healthy skepticism&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#force-it-to-clean-up-after-itself"&gt;Force it to clean up after itself&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#code-coverage-is-your-friend"&gt;Code coverage is your friend&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#9-my-secret-weapon-voice-to-code-workflow"&gt;9. My secret weapon: voice-to-code workflow&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-problem-with-typing-prompts"&gt;The problem with typing prompts&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#my-voice-workflow"&gt;My voice workflow&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#why-this-works-so-well"&gt;Why this works so well&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#10-additional-tips-for-agentic-coding"&gt;10. Additional tips for agentic coding&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#set-clear-constraints"&gt;Set clear constraints&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#use-the-task-system"&gt;Use the task system&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#leverage-the-search-capabilities"&gt;Leverage the search capabilities&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#review-in-chunks"&gt;Review in chunks&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#11-the-tools-ive-tried"&gt;11. The tools I&amp;rsquo;ve tried&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#pre-ai-era"&gt;Pre-AI era&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#phase-1-copy-paste-tools-march-2023---july-2025"&gt;Phase 1: copy-paste tools (March 2023 - July 2025)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#cursor-early-2024"&gt;Cursor (early 2024)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#phase-2-claude-code-july-2025---present"&gt;Phase 2: Claude Code (July 2025 - present)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#openais-codex-cli-the-reasoning-beast"&gt;OpenAI&amp;rsquo;s Codex CLI (the reasoning beast)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#why-not-copilot-or-other-cheaper-tools"&gt;Why not Copilot or other &amp;ldquo;cheaper&amp;rdquo; tools?&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#12-is-this-sustainable"&gt;12. Is this sustainable?&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-good"&gt;The good&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-challenges"&gt;The challenges&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#my-verdict"&gt;My verdict&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#13-the-experience-multiplier-why-seniority-matters-more-than-ever"&gt;13. The experience multiplier: why seniority matters more than ever&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-exponential-trap-for-beginners"&gt;The exponential trap for beginners&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-senior-developer-advantage"&gt;The senior developer advantage&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-right-way-for-every-level"&gt;The right way for every level&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-real-multiplier-effect"&gt;The real multiplier effect&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#14-reality-check-what-ai-can-and-cant-do"&gt;14. Reality check: what AI can and can&amp;rsquo;t do&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#what-ai-can-do-the-churn-work"&gt;What AI CAN do (the churn work)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#what-ai-cant-do-the-creative-work"&gt;What AI CAN&amp;rsquo;T do (the creative work)&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-sweet-spot"&gt;The sweet spot&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#15-looking-forward-the-future-of-development"&gt;15. Looking forward: the future of development&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#16-conclusion-principles-over-process"&gt;16. Conclusion: principles over process&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#links-and-resources"&gt;Links and resources&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;h2 id="1-from-skepticism-to-productivity"&gt;1. From skepticism to productivity &lt;a class="anchor" href="#1-from-skepticism-to-productivity" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;About eight months ago, I wrote about my disastrous experience with &amp;ldquo;&lt;a href="https://www.nijho.lt/post/vibe-coding/"&gt;vibe coding&lt;/a&gt;&amp;quot;—where I let AI generate code without careful review.&#10;What started as a fun 3-hour prototype turned into a 15-hour debugging nightmare.&#10;I concluded that while AI was fantastic for quick prototyping, it was dangerous without human oversight.&#10;I even declared I&amp;rsquo;d &amp;ldquo;never invest in, build upon, or use such products.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;Fast forward to today, and something remarkable has changed.&#10;Not my core principle—I still &lt;strong&gt;never merge code I don&amp;rsquo;t understand&lt;/strong&gt;—but the tooling and my approach have evolved through two distinct phases.&lt;/p&gt;&#10;&lt;p&gt;First, let&amp;rsquo;s look at the evidence of AI&amp;rsquo;s overall impact on my productivity:&lt;/p&gt;&#10;&lt;figure&gt;&lt;img src="https://www.nijho.lt/post/agentic-coding/pypi_packages_histogram.svg" alt="PyPI Package Publication Analysis" loading="lazy" data-zoomable&gt;&lt;/figure&gt;&#10;&lt;p&gt;This graph tells an interesting story.&#10;After maintaining a steady pace for years (2016-2023), my productivity exploded with the introduction of AI tools.&#10;In 2024-2025 alone, I&amp;rsquo;ve published &lt;strong&gt;14 new Python packages&lt;/strong&gt; and made &lt;strong&gt;over 400 releases&lt;/strong&gt;.&#10;That&amp;rsquo;s more new packages than in the previous 6 years combined!&lt;/p&gt;&#10;&lt;p&gt;But this transformation didn&amp;rsquo;t happen overnight—it came in two distinct phases that fundamentally changed how I write code.&lt;/p&gt;&#10;&lt;aside class="callout callout-note"&gt;&lt;svg class="ico" viewBox="0 0 256 256" aria-hidden="true"&gt;&lt;path d="M128,24A104,104,0,1,0,232,128,104.11,104.11,0,0,0,128,24Zm0,192a88,88,0,1,1,88-88A88.1,88.1,0,0,1,128,216Zm16-40a8,8,0,0,1-8,8,16,16,0,0,1-16-16V128a8,8,0,0,1,0-16,16,16,0,0,1,16,16v40A8,8,0,0,1,144,176ZM112,84a12,12,0,1,1,12,12A12,12,0,0,1,112,84Z"/&gt;&lt;/svg&gt;&lt;div&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; My AI coding journey had two major phases: Phase 1 started in March 2023 with GPT-4&amp;rsquo;s release—using chat interfaces with copy-paste workflows.&#10;Phase 2 began just two months ago in July 2025 with &lt;strong&gt;Claude Code&lt;/strong&gt;—an agentic AI that can explore codebases, run tests, and debug itself.&#10;This second leap was as transformative as the first one. I spent &amp;gt;10k USD worth of API usage in the first month alone (thankfully capped at $200 with the Pro plan).&lt;/div&gt;&lt;/aside&gt;&#10;&lt;h2 id="2-but-dont-you-love-programming"&gt;2. But don&amp;rsquo;t you love programming? &lt;a class="anchor" href="#2-but-dont-you-love-programming" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I&amp;rsquo;ve heard this objection countless times: &amp;ldquo;I love programming too much to let AI do it for me.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;Sure, I get it.&#10;It&amp;rsquo;s like preferring analog photography over digital, or writing letters by hand instead of typing.&#10;Some people genuinely enjoy the manual process, and that&amp;rsquo;s valid.&lt;/p&gt;&#10;&lt;p&gt;But here&amp;rsquo;s the thing: &lt;strong&gt;I absolutely love programming&lt;/strong&gt;.&#10;It&amp;rsquo;s my biggest passion in life.&#10;As I mentioned in my &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;local AI journey&lt;/a&gt;, I bought a gaming PC to play games but have only used it for coding and local AI experiments.&#10;My idea of a fun weekend? Writing code until 3 AM.&#10;I&amp;rsquo;ve been doing this for over 10 years, and I&amp;rsquo;ve &lt;strong&gt;never had as much fun programming as I do now&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;What I&amp;rsquo;ve discovered is that I don&amp;rsquo;t just love the act of typing code—I love &lt;strong&gt;building and creating things&lt;/strong&gt;.&#10;There&amp;rsquo;s a difference.&lt;/p&gt;&#10;&lt;p&gt;I&amp;rsquo;m ridiculously particular about code quality.&#10;I literally cannot look at poorly formatted code without feeling physical discomfort.&#10;Every operator, every indentation, every naming convention matters to me.&#10;I enforce strict linting rules and will spend time refactoring code just because the style bothers me.&#10;This isn&amp;rsquo;t about being pretentious—it&amp;rsquo;s about caring deeply about my craft.&lt;/p&gt;&#10;&lt;p&gt;The revelation: AI doesn&amp;rsquo;t take away my ability to care about these details.&#10;Instead, it gives me &lt;strong&gt;insane leverage&lt;/strong&gt; to create more while maintaining my personal standards.&#10;I still review every line (although I am more lenient in certain cases), still enforce my style guides, still refactor when something bothers me.&#10;But now I can build 10x faster.&lt;/p&gt;&#10;&lt;p&gt;It turns out that what I truly love isn&amp;rsquo;t the mechanical act of typing—it&amp;rsquo;s the creative act of bringing ideas to life.&#10;And with AI, I can bring more ideas to life than ever before.&lt;/p&gt;&#10;&lt;h3 id="the-power-to-explore"&gt;The power to explore &lt;a class="anchor" href="#the-power-to-explore" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;AI has unlocked something incredible: the ability to explore ideas that weren&amp;rsquo;t worth the effort before.&#10;It&amp;rsquo;s now trivial to write 2,000 lines of code to solve a simple problem.&#10;Before AI, spending 5+ hours on a utility script wasn&amp;rsquo;t justifiable.&#10;Now? I can build it in 10 minutes.&lt;/p&gt;&#10;&lt;p&gt;Here are just some of the &amp;ldquo;wouldn&amp;rsquo;t have been worth it&amp;rdquo; projects I&amp;rsquo;ve built recently:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://wenfire.nijho.lt/" target="_blank" rel="noopener"&gt;Financial independence tracker&lt;/a&gt;&lt;/strong&gt;: A full website visualizing my path to FIRE&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/tuitorial" target="_blank" rel="noopener"&gt;Tuitorial&lt;/a&gt;&lt;/strong&gt;: Terminal-based presentation software&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/home-assistant-streamdeck-yaml" target="_blank" rel="noopener"&gt;Stream Deck home control&lt;/a&gt;&lt;/strong&gt;: Complete tool to control my entire house from a Stream Deck&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/dotbins" target="_blank" rel="noopener"&gt;Dotbins&lt;/a&gt;&lt;/strong&gt;: Automatic binary dependency syncing across all my machines&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;Agent-CLI&lt;/a&gt;&lt;/strong&gt;: A local, open-source Siri that actually works&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Each of these would have taken weeks of manual coding.&#10;The effort-to-reward ratio just wasn&amp;rsquo;t there.&#10;But with AI? I can explore every random idea that pops into my head at 2 AM.&lt;/p&gt;&#10;&lt;h3 id="the-matty-story-from-midnight-idea-to-working-code"&gt;The Matty story: from midnight idea to working code &lt;a class="anchor" href="#the-matty-story-from-midnight-idea-to-working-code" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Here is the perfect example of this power: One night around midnight, I was lying in bed (slightly intoxicated) when I had an idea—I needed a way for my AI agents to communicate via Matrix.&#10;So I grabbed my iPhone, SSHed into my machine, and started building.&lt;/p&gt;&#10;&lt;p&gt;One hour later, at 1 AM, I had created &lt;strong&gt;&lt;a href="https://github.com/basnijholt/matty" target="_blank" rel="noopener"&gt;Matty&lt;/a&gt;&lt;/strong&gt;—a fully functional terminal-based Matrix client.&#10;Built entirely from my phone.&#10;In bed.&#10;(Poor sleep hygiene, I know.)&lt;/p&gt;&#10;&lt;p&gt;When I woke up the next morning and tested it&amp;hellip; it actually worked!&#10;I spent a few more hours adding tests and cleaning it up, but the core functionality was built in that hour between midnight and 1 AM.&lt;/p&gt;&#10;&lt;p&gt;This is impossible without AI.&#10;Nobody&amp;rsquo;s writing a Matrix client from their phone in bed at midnight.&#10;But with Claude Code? It&amp;rsquo;s just another Tuesday night idea that becomes reality.&lt;/p&gt;&#10;&lt;h3 id="clip-files-ui-when-ideas-strike-at-random-moments"&gt;clip-files UI: when ideas strike at random moments &lt;a class="anchor" href="#clip-files-ui-when-ideas-strike-at-random-moments" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;What happens more often than you&amp;rsquo;d think: I get an idea at a completely random moment and NEED to implement it immediately.&lt;/p&gt;&#10;&lt;p&gt;Case in point: I was sitting in the bathtub during my copy-paste AI era when I realized I needed to access my codebases from anywhere.&#10;I&amp;rsquo;d already built &lt;a href="https://www.nijho.lt/post/advent-of-open-source/15-clip-files/"&gt;&lt;code&gt;clip-files&lt;/code&gt;&lt;/a&gt;—a CLI tool that copies entire codebases to your clipboard—but I wanted it accessible from my phone.&lt;/p&gt;&#10;&lt;p&gt;So I got out of the bathtub, built a &lt;a href="https://github.com/basnijholt/clip-files-ui" target="_blank" rel="noopener"&gt;web UI&lt;/a&gt; for it in 20 minutes, and got back in.&#10;The UI lets me:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Browse all my GitHub repositories&lt;/li&gt;&#10;&lt;li&gt;Select any branch&lt;/li&gt;&#10;&lt;li&gt;Click one button to copy all source code to clipboard&lt;/li&gt;&#10;&lt;li&gt;Paste it into my self-hosted AI chat from anywhere&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Ironically, I rarely use this tool now—Claude Code can explore codebases directly.&#10;But back in the copy-paste era, this 20-minute bathtub interruption made mobile AI coding possible.&lt;/p&gt;&#10;&lt;p&gt;The barrier between &amp;ldquo;wouldn&amp;rsquo;t it be cool if&amp;hellip;&amp;rdquo; and &amp;ldquo;here&amp;rsquo;s a working prototype&amp;rdquo; has essentially disappeared.&lt;/p&gt;&#10;&lt;h3 id="the-800-rule-marathon-80-ai-agents-at-once"&gt;The 800-rule marathon: 80 AI agents at once &lt;a class="anchor" href="#the-800-rule-marathon-80-ai-agents-at-once" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Another thing that would be literally impossible without AI: I decided to enable ALL 800+ Ruff linting rules on a 20,000-line codebase.&lt;/p&gt;&#10;&lt;p&gt;For context, &lt;a href="https://docs.astral.sh/ruff/" target="_blank" rel="noopener"&gt;Ruff&lt;/a&gt; is a extremely fast Python linter that implements rules from over a dozen different tools.&#10;Most projects enable maybe 50-100 rules.&#10;I wanted all 800+.&lt;/p&gt;&#10;&lt;p&gt;The result? &lt;strong&gt;≈2,500 violations&lt;/strong&gt; across the entire codebase.&lt;/p&gt;&#10;&lt;p&gt;Fixing these by hand would take 20 hours if I could fix one per 30 seconds&amp;hellip;&#10;But with Claude Code&amp;rsquo;s agent spawning capabilities, I did something insane:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;Opened 8 terminal windows&lt;/strong&gt; with Claude Code&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Created a work orchestration document&lt;/strong&gt; that one AI managed&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Each session spawned 10 parallel sub-agents&lt;/strong&gt; using &lt;a href="https://docs.anthropic.com/en/docs/claude-code/sub-agents" target="_blank" rel="noopener"&gt;/agents&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Total: ~80 AI agents&lt;/strong&gt; working on my code simultaneously&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Two hours later&lt;/strong&gt;: 100% compliance with all 800 rules&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;The orchestrator agent would assign work like:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&amp;ldquo;Session 1: Fix all E501 line-too-long violations in src/core/&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;Session 2: Handle all B008 function-call-in-default-argument issues&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;Session 3: Fix all RUF100 unused-noqa violations&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Each session&amp;rsquo;s sub-agents would tackle different files in parallel.&#10;When conflicts arose, the orchestrator would resolve them.&#10;This is the kind of code quality improvement that simply wouldn&amp;rsquo;t happen without AI—not because it&amp;rsquo;s technically difficult, but because the human effort required is unrealistic.&lt;/p&gt;&#10;&lt;h3 id="1400-commits-rewritten-in-minutes-for-250"&gt;1,400 commits, rewritten in minutes for $2.50 &lt;a class="anchor" href="#1400-commits-rewritten-in-minutes-for-250" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Here&amp;rsquo;s another &amp;ldquo;impossible without AI&amp;rdquo; story: I had 1,400 commits with terrible messages (in a not-yet-released and private project).&lt;/p&gt;&#10;&lt;p&gt;Since I commit so frequently (using them as snapshots in feature branches), my commit messages were all over the place:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Some were just a period &amp;ldquo;.&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;Some were decent descriptions&lt;/li&gt;&#10;&lt;li&gt;Some were just merge commits&lt;/li&gt;&#10;&lt;li&gt;Zero consistency in style&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;I wanted them all to follow a consistent convention with prefixes like &lt;code&gt;fix:&lt;/code&gt;, &lt;code&gt;feat:&lt;/code&gt;, &lt;code&gt;test:&lt;/code&gt;, &lt;code&gt;ci:&lt;/code&gt;, etc.&lt;/p&gt;&#10;&lt;p&gt;Obviously, rewriting 1,400 commit messages by hand is not going to happen.&#10;But with AI?&lt;/p&gt;&#10;&lt;p&gt;I had Claude write me an agent using the Agno framework that would:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Read each commit&amp;rsquo;s full diff using &lt;code&gt;git show&lt;/code&gt;&lt;/li&gt;&#10;&lt;li&gt;Analyze the changes to understand what was done&lt;/li&gt;&#10;&lt;li&gt;Return structured JSON with three fields:&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;commit_message&lt;/code&gt;: The improved message (or empty to keep original)&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;keep&lt;/code&gt;: Boolean flag whether to keep the original&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;reasoning&lt;/code&gt;: Why the AI made that decision&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Then the magic happened: &lt;strong&gt;I spawned 1,400 AI agents simultaneously using the DeepSeek API&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Ten minutes and $2.50 later, all my commit messages were rewritten with consistent, meaningful descriptions.&#10;The entire Git history now looks more professional and searchable.&lt;/p&gt;&#10;&lt;p&gt;The barrier between &amp;ldquo;wouldn&amp;rsquo;t it be cool if&amp;hellip;&amp;rdquo; and &amp;ldquo;here&amp;rsquo;s a working prototype&amp;rdquo; has essentially disappeared.&lt;/p&gt;&#10;&lt;h2 id="3-the-two-phase-ai-revolution"&gt;3. The two-phase AI revolution &lt;a class="anchor" href="#3-the-two-phase-ai-revolution" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My AI coding journey wasn&amp;rsquo;t a single leap—it happened in two distinct phases, each with its own breakthrough moment.&lt;/p&gt;&#10;&lt;h3 id="phase-1-copy-paste-era-march-2023---july-2025"&gt;Phase 1: copy-paste era (March 2023 - July 2025) &lt;a class="anchor" href="#phase-1-copy-paste-era-march-2023---july-2025" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;When GPT-4 launched in March 2023, it changed everything.&#10;I built tools like &lt;a href="https://www.nijho.lt/post/advent-of-open-source/15-clip-files/"&gt;&lt;code&gt;clip-files&lt;/code&gt;&lt;/a&gt; to efficiently copy code context into a ChatGPT-like web interface.&#10;For over two years, this workflow was:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Manual but effective&lt;/strong&gt;: Copy code → paste into AI → review suggestions → implement manually&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Limited by context&lt;/strong&gt;: Could only share what fit in a single message&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;No validation&lt;/strong&gt;: AI couldn&amp;rsquo;t test its suggestions&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Still valuable&lt;/strong&gt;: Helped with ideation, code review, and refactoring&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This phase was already productive—I was shipping more code than before.&#10;But I was still the one running tests, debugging errors, and validating everything.&lt;/p&gt;&#10;&lt;h3 id="phase-2-the-claude-code-revolution-july-2025"&gt;Phase 2: the Claude Code revolution (July 2025) &lt;a class="anchor" href="#phase-2-the-claude-code-revolution-july-2025" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Just two months ago, in July 2025, everything changed again.&#10;Frustrated with Cursor becoming painfully slow, I tried Claude Code based on the buzz online.&#10;&lt;strong&gt;That first night, I spent $70 on API tokens&lt;/strong&gt;.&#10;Not by accident—I just kept adding $10-$20 increments like a gambling addict.&#10;The experience was so mind-blowing that within two weeks, I&amp;rsquo;d spent nearly $200 and immediately upgraded to the $200/month Pro plan.&lt;/p&gt;&#10;&lt;p&gt;To put this in perspective: Using a command-line tool &lt;code&gt;ccusage&lt;/code&gt; to track token consumption, I calculated that I used &lt;strong&gt;$10,000 worth of API tokens in my first month&lt;/strong&gt;.&#10;Thankfully, the Pro plan capped my cost at $200!&lt;/p&gt;&#10;&lt;h3 id="what-made-claude-code-different"&gt;What made Claude Code different? &lt;a class="anchor" href="#what-made-claude-code-different" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;The shift from copy-paste to agentic AI was like upgrading from a bicycle to a spaceship:&lt;/p&gt;&#10;&lt;h4 id="before-copy-paste-with-chatgptclaude-web"&gt;Before (copy-paste with ChatGPT/Claude Web)&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Generate code based on prompts&lt;/li&gt;&#10;&lt;li&gt;No access to your codebase&lt;/li&gt;&#10;&lt;li&gt;Can&amp;rsquo;t run tests or see errors&lt;/li&gt;&#10;&lt;li&gt;You debug everything manually&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="after-claude-code"&gt;After (Claude Code)&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Reads and searches your entire codebase&lt;/li&gt;&#10;&lt;li&gt;Executes commands and runs tests&lt;/li&gt;&#10;&lt;li&gt;Sees error messages and debugs itself&lt;/li&gt;&#10;&lt;li&gt;Iterates until tests pass&lt;/li&gt;&#10;&lt;li&gt;Uses tools to explore and validate&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The difference is profound.&#10;I went from carefully copying snippets with &lt;code&gt;clip-files&lt;/code&gt; to having an assistant that can explore my entire project, run my test suite, fix failures, and even commit the changes.&lt;/p&gt;&#10;&lt;h2 id="4-my-current-agentic-workflow"&gt;4. My current agentic workflow &lt;a class="anchor" href="#4-my-current-agentic-workflow" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Here&amp;rsquo;s how I actually work with Claude Code on a typical project:&lt;/p&gt;&#10;&lt;h3 id="parallel-development-6-features-at-once"&gt;Parallel development: 6 features at once &lt;a class="anchor" href="#parallel-development-6-features-at-once" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;One of Claude Code&amp;rsquo;s superpowers is enabling truly parallel development.&#10;Here&amp;rsquo;s my setup that would make any developer from 10 years ago think I&amp;rsquo;m insane:&lt;/p&gt;&#10;&lt;p&gt;I use &lt;a href="https://zellij.dev/" target="_blank" rel="noopener"&gt;Zellij&lt;/a&gt; (a terminal multiplexer like tmux) with a custom layout for my projects:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;5 main tabs&lt;/strong&gt;: Each split into two panes—Claude Code on one side, terminal on the other&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Using independent copies of same repository&lt;/strong&gt;: Each tab is a separate Git worktree with its own environment, own environment variables, own deployment&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Each tab = different feature&lt;/strong&gt;: Using &lt;a href="https://www.nijho.lt/post/git-worktree/"&gt;Git worktrees&lt;/a&gt; to work on 6 features simultaneously&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Monitoring tabs&lt;/strong&gt;: &lt;code&gt;ccusage&lt;/code&gt; for tracking Claude usage, &lt;code&gt;htop&lt;/code&gt; for CPU, &lt;code&gt;nvtop&lt;/code&gt; for GPU&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Voice-driven orchestration&lt;/strong&gt;: I cycle through tabs, review code while speaking, move to next&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;My workflow looks like this:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Start feature A in tab 1, give Claude instructions&lt;/li&gt;&#10;&lt;li&gt;Switch to tab 2, start feature B while A is working&lt;/li&gt;&#10;&lt;li&gt;Continue through all 5 tabs, starting different features&lt;/li&gt;&#10;&lt;li&gt;Circle back to tab 1, review what Claude did, give feedback via voice&lt;/li&gt;&#10;&lt;li&gt;Repeat the cycle every 10-15 minutes&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;This parallel workflow is how I built a complex project (that I haven&amp;rsquo;t released yet) in one month that would have taken 9+ months before.&#10;The key is &lt;a href="https://www.nijho.lt/post/git-worktree/"&gt;Git worktrees&lt;/a&gt;—each feature gets its own working directory, so there&amp;rsquo;s no context switching overhead.&lt;/p&gt;&#10;&lt;!-- [TODO INSERT GIF: Terminal multiplexer showing multiple Claude Code sessions] --&gt;&#10;&lt;h3 id="starting-a-new-package"&gt;Starting a new package &lt;a class="anchor" href="#starting-a-new-package" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# I describe what I want to build&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&amp;#34;I need a Python package that manages CLI tool binaries within Git repositories,&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;automatically downloading the correct binary for the user&amp;#39;s platform, here &amp;lt;insert path&amp;gt; is a repo with boilerplate skeleton to copy (use same conventions)&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Claude Code then:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 1. Creates the project structure&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 2. Implements core functionality&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 3. Writes comprehensive tests&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 4. Generates documentation&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 5. Sets up CI/CD workflows&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="the-key-difference-iteration-and-validation"&gt;The key difference: iteration and validation &lt;a class="anchor" href="#the-key-difference-iteration-and-validation" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Unlike my vibe coding disaster, Claude Code doesn&amp;rsquo;t just dump code and leave.&#10;Here&amp;rsquo;s a real interaction pattern:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude writes initial implementation&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude runs the tests&lt;/strong&gt; → Several fail&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude reads the error messages and fixes the code&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude runs tests again&lt;/strong&gt; → More pass, some still fail&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude debugs the remaining issues&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude validates all tests pass&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;I review the final code and understand it (tell it to make changes if needed)&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude creates a proper Git commit&lt;/strong&gt;&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;This iterative, validated approach is what makes agentic coding so powerful.&#10;It&amp;rsquo;s not about blindly accepting generated code—it&amp;rsquo;s about having an assistant that can explore, test, and refine solutions.&lt;/p&gt;&#10;&lt;h2 id="5-the-evidence-explosive-productivity"&gt;5. The evidence: explosive productivity &lt;a class="anchor" href="#5-the-evidence-explosive-productivity" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Let me share some concrete data from my PyPI packages analysis:&lt;/p&gt;&#10;&lt;h3 id="package-creation-by-year"&gt;Package creation by year &lt;a class="anchor" href="#package-creation-by-year" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;2016-2022&lt;/strong&gt; (7 years pre-AI): ~2 packages/year average&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;2023&lt;/strong&gt; (GPT-4 launch year): 6 packages&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;2024&lt;/strong&gt; (full year with copy-paste AI): 7 packages&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;2025 until 3 months ago&lt;/strong&gt; (6 months): 7 packages&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;since 3 months&lt;/strong&gt; (2 with Claude Code, 1 with Codex CLI): wrote about 400k line of code (although this includes many iterations)!&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The acceleration with AI is clear!&lt;/p&gt;&#10;&lt;h3 id="recent-packages-built-with-agentic-ai"&gt;Recent packages built with agentic AI &lt;a class="anchor" href="#recent-packages-built-with-agentic-ai" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Here are some packages I&amp;rsquo;ve built recently with Claude Code:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/matty" target="_blank" rel="noopener"&gt;&lt;code&gt;matty&lt;/code&gt;&lt;/a&gt;&lt;/strong&gt;: Terminal-based AI chat with file integration&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;&lt;code&gt;agent-cli&lt;/code&gt;&lt;/a&gt;&lt;/strong&gt;: Local-first AI-powered CLI agents (built in days, not weeks!)&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Each of these would have taken me weeks or months to build alone.&#10;With Claude Code, I&amp;rsquo;m shipping production-ready packages in days.&lt;/p&gt;&#10;&lt;h2 id="6-why-this-isnt-vibe-coding-but-with-nuance"&gt;6. Why this isn&amp;rsquo;t &amp;ldquo;vibe coding&amp;rdquo; (but with nuance) &lt;a class="anchor" href="#6-why-this-isnt-vibe-coding-but-with-nuance" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;You might wonder: isn&amp;rsquo;t this just vibe coding with extra steps? The answer is nuanced and depends on the context—similar to my &lt;a href="https://www.nijho.lt/post/dependencies/"&gt;philosophy on dependencies&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h3 id="my-context-driven-standards"&gt;My context-driven standards &lt;a class="anchor" href="#my-context-driven-standards" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Just like I have different standards for dependencies in libraries versus applications, I apply different levels of scrutiny based on what I&amp;rsquo;m building:&lt;/p&gt;&#10;&lt;h4 id="for-critical-libraries-eg-adaptive-pipefunc-unidep"&gt;For critical libraries (e.g., Adaptive, pipefunc, unidep)&lt;/h4&gt;&#10;&lt;p&gt;&lt;strong&gt;Maximum scrutiny&lt;/strong&gt; - These are packages others depend on:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Every single line reviewed and understood&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;100% test coverage with careful test review&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Architecture decisions carefully considered&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Documentation must be comprehensive&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;My reputation is on the line&lt;/strong&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="for-experimental-clis-and-personal-tools"&gt;For experimental CLIs and personal tools&lt;/h4&gt;&#10;&lt;p&gt;&lt;strong&gt;Pragmatic approach&lt;/strong&gt; - Isolated tools with no downstream dependencies:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Core architecture must be fully understood&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Implementation details can be AI-generated if tests pass&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Test generation can be more automated&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Focus on functionality over perfection&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Similar to my relaxed dependency stance for applications&lt;/strong&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-internal-dependency-graph-principle"&gt;The internal dependency graph principle &lt;a class="anchor" href="#the-internal-dependency-graph-principle" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;The key insight: &lt;strong&gt;My scrutiny level correlates with how foundational the code is within the project&lt;/strong&gt;:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;Core/foundational code&lt;/strong&gt; (what everything else depends on): Maximum scrutiny, every line matters&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Data models, core algorithms, API interfaces&lt;/li&gt;&#10;&lt;li&gt;Authentication, database operations, state management&lt;/li&gt;&#10;&lt;li&gt;These are the &amp;ldquo;roots&amp;rdquo; that everything else builds upon&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;Peripheral/leaf code&lt;/strong&gt; (nothing depends on it): Can be more AI-delegated&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Plotting functions, display utilities, CLI formatters&lt;/li&gt;&#10;&lt;li&gt;Test helpers, documentation generators&lt;/li&gt;&#10;&lt;li&gt;These are the &amp;ldquo;leaves&amp;rdquo; that don&amp;rsquo;t affect other code&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;strong&gt;Work projects&lt;/strong&gt;: Always maximum scrutiny for foundational code, regardless of project type&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This isn&amp;rsquo;t about being lazy—it&amp;rsquo;s about focusing human attention where it matters most.&#10;I can trust AI more with a plotting function that nothing depends on than with a core data structure that the entire system uses.&lt;/p&gt;&#10;&lt;h3 id="building-constraints-around-ai"&gt;Building constraints around AI &lt;a class="anchor" href="#building-constraints-around-ai" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;The secret to productive agentic coding isn&amp;rsquo;t just the AI—it&amp;rsquo;s the &lt;strong&gt;constraints and guardrails&lt;/strong&gt; I build around it:&lt;/p&gt;&#10;&lt;h4 id="automated-quality-gates"&gt;Automated quality gates&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Ruff with strictest rules&lt;/strong&gt;: Catches style issues, complexity problems, and common bugs&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;MyPy in strict mode&lt;/strong&gt;: Enforces type safety across the entire codebase&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Pre-commit hooks&lt;/strong&gt;: Automatically format and validate code before commits&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Comprehensive test suites&lt;/strong&gt;: AI must make tests pass, not just write code&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="project-specific-guidance"&gt;Project-specific guidance&lt;/h4&gt;&#10;&lt;p&gt;Every project gets a &lt;code&gt;CLAUDE.md&lt;/code&gt; file with explicit rules:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;No defensive programming&lt;/strong&gt;: Don&amp;rsquo;t wrap things in try-except unless necessary&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Functional over classes&lt;/strong&gt;: Prefer simple functions in Python&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;No backward compatibility&lt;/strong&gt;: For new projects, embrace breaking changes&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Be ruthless&lt;/strong&gt;: Aggressively remove unused code&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="custom-commands-and-workflows"&gt;Custom commands and workflows&lt;/h4&gt;&#10;&lt;p&gt;I&amp;rsquo;ve built specific commands that inject context and constraints:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Anti-cruft reviews&lt;/strong&gt;: Remove over-engineering and defensive code&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Safe commit practices&lt;/strong&gt;: Never use &lt;code&gt;git add .&lt;/code&gt;, always selective staging&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Initialize understanding&lt;/strong&gt;: Load project context and current work state&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="local-virtual-environments-stop-the-hallucinations"&gt;Local virtual environments: stop the hallucinations&lt;/h4&gt;&#10;&lt;p&gt;Here&amp;rsquo;s a critical tip that eliminated a lot of my AI frustrations: &lt;strong&gt;keep a virtual environment with all dependencies installed locally&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Instead of letting Claude hallucinate how libraries work, I tell it explicitly in my &lt;code&gt;CLAUDE.md&lt;/code&gt;:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-markdown" data-lang="markdown"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="gu"&gt;### Step 1: understand the context&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; **READ THE SOURCE CODE**: This library has a &lt;span class="sb"&gt;`.venv`&lt;/span&gt; folder with all dependencies installed.&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; So read the source code when in doubt.&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; **Never guess API behavior**: If unsure, inspect the actual implementation in &lt;span class="sb"&gt;`.venv/lib/python*/site-packages/`&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This simple addition transforms Claude from guessing about library APIs to actually reading them:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Before&lt;/strong&gt;: &amp;ldquo;I think your custom AsyncProcessor.batch() method takes a list&amp;hellip;&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;After&lt;/strong&gt;: Reads &lt;code&gt;/path/to/.venv/lib/python3.11/site-packages/my_internal_lib/processor.py&lt;/code&gt; and knows it actually takes an iterator&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;When Claude can read the actual source of your own packages, internal company libraries, or niche dependencies, it stops making assumptions and starts working with facts.&#10;This is especially powerful for:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Your own libraries&lt;/strong&gt; that Claude has never seen before&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Internal company packages&lt;/strong&gt; that aren&amp;rsquo;t public&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Niche libraries&lt;/strong&gt; with barely any GitHub stars or documentation&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Modified versions&lt;/strong&gt; of popular libraries with custom patches&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Edge cases&lt;/strong&gt; where even good documentation doesn&amp;rsquo;t cover everything&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;These constraints transform AI from a loose cannon into a precision tool.&lt;/p&gt;&#10;&lt;h2 id="7-common-pitfalls-ive-noticed"&gt;7. Common pitfalls I&amp;rsquo;ve noticed &lt;a class="anchor" href="#7-common-pitfalls-ive-noticed" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;After a few months of using agentic AI tools, I&amp;rsquo;ve noticed consistent patterns that require vigilance:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Note from the future: I now realize this is very model dependent! GPT-5 has different pitfalls than Claude Opus 4.1&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h3 id="the-defensive-programming-trap"&gt;The &amp;ldquo;defensive programming&amp;rdquo; trap &lt;a class="anchor" href="#the-defensive-programming-trap" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# AI loves this:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;some_function&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="ne"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;pass&lt;/span&gt; &lt;span class="c1"&gt;# Silently suppress errors 😱&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# But we need this:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;some_function&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="c1"&gt;# Let it fail loudly if something&amp;#39;s wrong&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Claude tends to wrap everything in try-except blocks, suppressing errors that should bubble up.&#10;This is why my &lt;code&gt;CLAUDE.md&lt;/code&gt; explicitly forbids unnecessary error handling.&lt;/p&gt;&#10;&lt;h3 id="the-backwards-compatibility-obsession"&gt;The &amp;ldquo;backwards compatibility&amp;rdquo; obsession &lt;a class="anchor" href="#the-backwards-compatibility-obsession" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Claude constantly adds backwards compatibility for features that were literally just introduced in the same session:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&amp;ldquo;Maintaining compatibility with the old version&amp;rdquo; (that never existed)&lt;/li&gt;&#10;&lt;li&gt;Fallback mechanisms for code paths that were just created&lt;/li&gt;&#10;&lt;li&gt;Multiple ways to do the same thing &amp;ldquo;for flexibility&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-over-engineering-disease"&gt;The &amp;ldquo;over-engineering&amp;rdquo; disease &lt;a class="anchor" href="#the-over-engineering-disease" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Implements factory patterns for simple object creation&lt;/li&gt;&#10;&lt;li&gt;Adds abstraction layers that serve no purpose&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;Production ready&amp;rdquo; code for experimental scripts&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-git-commit-sins"&gt;The git commit sins &lt;a class="anchor" href="#the-git-commit-sins" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Despite explicit instructions in &lt;code&gt;CLAUDE.md&lt;/code&gt;:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Creates &lt;code&gt;your_module_v2.py&lt;/code&gt; alongside &lt;code&gt;your_module.py&lt;/code&gt; instead of updating&lt;/li&gt;&#10;&lt;li&gt;Still tries &lt;code&gt;git add -A&lt;/code&gt; or &lt;code&gt;git add .&lt;/code&gt; regularly despite explicit bans&lt;/li&gt;&#10;&lt;li&gt;Loves to &amp;ldquo;helpfully&amp;rdquo; revert debugging changes from other files&lt;/li&gt;&#10;&lt;li&gt;Commits &lt;code&gt;.env&lt;/code&gt; files if not watched carefully&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-helpful-anti-patterns"&gt;The &amp;ldquo;helpful&amp;rdquo; anti-patterns &lt;a class="anchor" href="#the-helpful-anti-patterns" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Loves defensive programming&lt;/strong&gt;: Validates things that can&amp;rsquo;t be wrong&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Gladly reads &lt;code&gt;.env&lt;/code&gt;&lt;/strong&gt;: Will expose secrets if not careful&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Reverts unrelated changes&lt;/strong&gt;: &amp;ldquo;Cleans up&amp;rdquo; debugging code from other features&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-mission-accomplished-hallucination"&gt;The &amp;ldquo;mission accomplished&amp;rdquo; hallucination &lt;a class="anchor" href="#the-mission-accomplished-hallucination" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;This is a dangerous pitfall.&#10;Claude Code will sometimes claim complete success when it hasn&amp;rsquo;t actually fixed anything:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Claims victory&lt;/strong&gt;: &amp;ldquo;I&amp;rsquo;ve fixed all the issues!&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Ask for proof&lt;/strong&gt;: &amp;ldquo;Show me the test output&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Backpedals&lt;/strong&gt;: &amp;ldquo;Oh, let me actually run the tests&amp;hellip;&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Still claims success&lt;/strong&gt;: &amp;ldquo;Yes, I can see in the logs it works!&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Demand the actual log&lt;/strong&gt;: &amp;ldquo;Show me the exact log file&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Finally admits&lt;/strong&gt;: &amp;ldquo;Actually, there are no log files. The tests don&amp;rsquo;t pass.&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;strong&gt;Always demand proof.&lt;/strong&gt; Never accept &amp;ldquo;it&amp;rsquo;s done&amp;rdquo; without seeing actual test output.&lt;/p&gt;&#10;&lt;h3 id="the-nuclear-option"&gt;The nuclear option &lt;a class="anchor" href="#the-nuclear-option" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I&amp;rsquo;ve had Claude Code literally try to delete all project files when &amp;ldquo;cleaning up&amp;rdquo;:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Claude&amp;#39;s &amp;#34;helpful&amp;#34; cleanup:&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&amp;#34;Task complete! Let me remove the temporary files...&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;rm -rf src/ &lt;span class="c1"&gt;# 😱&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is why &lt;strong&gt;frequent git commits are non-negotiable&lt;/strong&gt;.&#10;I commit after every small success.&lt;/p&gt;&#10;&lt;h3 id="the-debugging-debris"&gt;The debugging debris &lt;a class="anchor" href="#the-debugging-debris" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;During problem-solving, Claude Code leaves a trail of attempts:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Debug logging statements everywhere&lt;/li&gt;&#10;&lt;li&gt;Failed attempt code commented out&lt;/li&gt;&#10;&lt;li&gt;Multiple approaches tried in parallel&lt;/li&gt;&#10;&lt;li&gt;Temporary test files and scripts&lt;/li&gt;&#10;&lt;li&gt;Extra imports and unused functions&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Before merging, always ask: &amp;ldquo;Review your changes and remove all debugging artifacts and failed attempts.&amp;rdquo;&#10;Getting to a high code coverage (approaching 100%) helps ensure no dead code remains, as long as your tests use public APIs.&lt;/p&gt;&#10;&lt;h3 id="why-this-happens"&gt;Why this happens &lt;a class="anchor" href="#why-this-happens" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;These patterns emerge because AI is trained on public code that often:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Maintains backward compatibility for years&lt;/li&gt;&#10;&lt;li&gt;Uses defensive programming for public APIs&lt;/li&gt;&#10;&lt;li&gt;Includes extensive error handling for user input&lt;/li&gt;&#10;&lt;li&gt;Follows &amp;ldquo;enterprise&amp;rdquo; patterns even for simple scripts&lt;/li&gt;&#10;&lt;li&gt;Contains debugging code from development&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This is why &lt;strong&gt;constraints are essential&lt;/strong&gt;—without them, Claude defaults to these &amp;ldquo;safe&amp;rdquo; but overcomplicated patterns.&lt;/p&gt;&#10;&lt;h2 id="8-critical-success-factors-what-actually-makes-this-work"&gt;8. Critical success factors: what actually makes this work &lt;a class="anchor" href="#8-critical-success-factors-what-actually-makes-this-work" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;After two months of intense usage, here are the non-negotiable practices that make agentic coding actually productive:&lt;/p&gt;&#10;&lt;h3 id="teach-it-your-test-commands-day-1-priority"&gt;Teach it your test commands (day 1 priority) &lt;a class="anchor" href="#teach-it-your-test-commands-day-1-priority" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;&lt;strong&gt;This is absolutely crucial.&lt;/strong&gt; Claude Code needs to know how to run your tests and how to activate the environment:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# To run tests, use: &lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;uv run pytest tests/ -xvs&lt;span class="s2"&gt;&amp;#34;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;# For coverage:&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;uv run pytest --cov=src --cov-report=term-missing&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;# To run specific test&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;uv run pytest tests/test_module.py::test_function&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Without this, it&amp;rsquo;s just guessing whether code works.&#10;With it, it becomes genuinely useful.&lt;/p&gt;&#10;&lt;h3 id="git-commits-are-your-safety-net"&gt;Git commits are your safety net &lt;a class="anchor" href="#git-commits-are-your-safety-net" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I commit &lt;strong&gt;obsessively&lt;/strong&gt; when using Claude Code:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;After every successful (even partial) feature implementation&lt;/li&gt;&#10;&lt;li&gt;Before letting it attempt any major refactoring&lt;/li&gt;&#10;&lt;li&gt;Whenever tests pass&lt;/li&gt;&#10;&lt;li&gt;Before any &amp;ldquo;cleanup&amp;rdquo; operation&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;But here&amp;rsquo;s the key: &lt;strong&gt;I develop every feature in its own branch&lt;/strong&gt;, even as a solo developer. My workflow:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Create a feature branch for each new feature or fix&lt;/li&gt;&#10;&lt;li&gt;Make frequent commits as snapshots (sometimes just a period for the message or tell Claude to commit)&lt;/li&gt;&#10;&lt;li&gt;Open a pull request to review the full diff myself&lt;/li&gt;&#10;&lt;li&gt;Merge only after reviewing all changes&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;This has saved me from disaster multiple times.&#10;Git is your undo button when Claude goes nuclear (e.g., breaks or removes your code).&lt;/p&gt;&#10;&lt;h3 id="maintain-healthy-skepticism"&gt;Maintain healthy skepticism &lt;a class="anchor" href="#maintain-healthy-skepticism" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Never trust, always verify:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;&amp;ldquo;I fixed it!&amp;rdquo;&lt;/strong&gt; → &amp;ldquo;Show me the test output&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&amp;ldquo;It&amp;rsquo;s working now!&amp;rdquo;&lt;/strong&gt; → &amp;ldquo;Run the tests again with -xvs&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&amp;ldquo;The logs show success!&amp;rdquo;&lt;/strong&gt; → &amp;ldquo;Cat the actual log file&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;&amp;ldquo;It&amp;rsquo;s production ready!&amp;rdquo;&lt;/strong&gt; → &amp;ldquo;Did you run the tests?&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Think of Claude Code as an enthusiastic junior developer who sometimes exaggerates their accomplishments.&lt;/p&gt;&#10;&lt;h3 id="force-it-to-clean-up-after-itself"&gt;Force it to clean up after itself &lt;a class="anchor" href="#force-it-to-clean-up-after-itself" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;After any debugging session, always:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&amp;#34;Review all your changes and remove:&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- Debug print statements&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- Commented out code&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- Failed attempt implementations&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- Temporary test files&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;- Unused imports&amp;#34;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;My custom &lt;code&gt;/anti-cruft&lt;/code&gt; command automates this, but you can do it manually too.&lt;/p&gt;&#10;&lt;h3 id="code-coverage-is-your-friend"&gt;Code coverage is your friend &lt;a class="anchor" href="#code-coverage-is-your-friend" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;High test coverage (90%+) ensures:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The code Claude wrote actually runs&lt;/li&gt;&#10;&lt;li&gt;No dead code from failed attempts&lt;/li&gt;&#10;&lt;li&gt;All paths are exercised&lt;/li&gt;&#10;&lt;li&gt;You can refactor confidently&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="9-my-secret-weapon-voice-to-code-workflow"&gt;9. My secret weapon: voice-to-code workflow &lt;a class="anchor" href="#9-my-secret-weapon-voice-to-code-workflow" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;One of my biggest productivity multipliers isn&amp;rsquo;t Claude Code itself—it&amp;rsquo;s how I communicate with it.&#10;Using my &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;&lt;code&gt;agent-cli&lt;/code&gt;&lt;/a&gt; tool, I&amp;rsquo;ve developed a voice-first workflow that&amp;rsquo;s transformed how I write prompts.&lt;/p&gt;&#10;&lt;h3 id="the-problem-with-typing-prompts"&gt;The problem with typing prompts &lt;a class="anchor" href="#the-problem-with-typing-prompts" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Effective agentic coding requires &lt;strong&gt;precise, detailed instructions&lt;/strong&gt;.&#10;A good prompt isn&amp;rsquo;t 20 words—it&amp;rsquo;s often 200-500 words explaining exactly what you want, what to avoid, and how to approach the problem.&#10;Most people don&amp;rsquo;t want to type that much.&lt;/p&gt;&#10;&lt;h3 id="my-voice-workflow"&gt;My voice workflow &lt;a class="anchor" href="#my-voice-workflow" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Here&amp;rsquo;s my actual workflow:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;Start recording&lt;/strong&gt; with &lt;code&gt;agent-cli transcribe&lt;/code&gt; (hotkey triggered)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Review Claude&amp;rsquo;s changes&lt;/strong&gt; while speaking my thoughts aloud&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Speak for minutes&lt;/strong&gt; about what I want, what&amp;rsquo;s wrong, what to fix&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Paste the transcription&lt;/strong&gt; directly into Claude Code&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;This is powered by OpenAI&amp;rsquo;s Whisper model running locally on my RTX 3090 at home.&#10;It&amp;rsquo;s more reliable than macOS dictation and gives me complete privacy.&lt;/p&gt;&#10;&lt;h3 id="why-this-works-so-well"&gt;Why this works so well &lt;a class="anchor" href="#why-this-works-so-well" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Rich prompts&lt;/strong&gt;: I naturally give more context when speaking&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Code review narration&lt;/strong&gt;: I can review code while explaining issues&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Thinking out loud&lt;/strong&gt;: Speaking helps clarify my own thoughts&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Speed&lt;/strong&gt;: Speaking is 3-4x faster than typing&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Precision&lt;/strong&gt;: I can be incredibly specific without typing fatigue&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;For example, instead of typing &amp;ldquo;fix the bug,&amp;rdquo; I might say:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&amp;ldquo;I&amp;rsquo;m looking at line 45 where you&amp;rsquo;re handling the authentication.&#10;The problem is you&amp;rsquo;re not checking if the token has expired before making the API call.&#10;Also, you&amp;rsquo;ve added a try-except block here that&amp;rsquo;s suppressing errors—remove that.&#10;And while you&amp;rsquo;re at it, the logging statement on line 52 is using the wrong format.&#10;Make sure to follow our project&amp;rsquo;s logging conventions&amp;hellip;&amp;rdquo;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;This level of detail is what makes AI coding actually productive.&lt;/p&gt;&#10;&lt;h2 id="10-additional-tips-for-agentic-coding"&gt;10. Additional tips for agentic coding &lt;a class="anchor" href="#10-additional-tips-for-agentic-coding" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Beyond the critical success factors and voice workflow, here are more tips to maximize your productivity:&lt;/p&gt;&#10;&lt;h3 id="set-clear-constraints"&gt;Set clear constraints &lt;a class="anchor" href="#set-clear-constraints" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Create a &lt;code&gt;CLAUDE.md&lt;/code&gt; or similar file in your project root with your preferences:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-markdown" data-lang="markdown"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="gu"&gt;## Project standards&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; Never use try/except to suppress errors silently&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; Always run pytest before marking task complete&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; Use type hints for all functions&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;-&lt;/span&gt; Prefer simple solutions over clever ones&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="use-the-task-system"&gt;Use the task system &lt;a class="anchor" href="#use-the-task-system" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Claude Code&amp;rsquo;s todo system is great for complex work:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;1. ✅ Implement core functionality&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;2. ✅ Write comprehensive tests&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;3. 🔄 Fix failing tests&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;4. ⏳ Add documentation&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h3 id="leverage-the-search-capabilities"&gt;Leverage the search capabilities &lt;a class="anchor" href="#leverage-the-search-capabilities" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Let Claude search for existing patterns:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&amp;ldquo;Find all API endpoint implementations in this project&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;Show me how we handle authentication elsewhere&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;What testing patterns are we using?&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="review-in-chunks"&gt;Review in chunks &lt;a class="anchor" href="#review-in-chunks" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Instead of reviewing 500 lines at once:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Have Claude implement a single feature&lt;/li&gt;&#10;&lt;li&gt;Review and understand it&lt;/li&gt;&#10;&lt;li&gt;Run tests&lt;/li&gt;&#10;&lt;li&gt;Commit(s) and PR&lt;/li&gt;&#10;&lt;li&gt;Move to the next feature&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;h2 id="11-the-tools-ive-tried"&gt;11. The tools I&amp;rsquo;ve tried &lt;a class="anchor" href="#11-the-tools-ive-tried" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My journey through AI coding assistants has been evolutionary:&lt;/p&gt;&#10;&lt;h3 id="pre-ai-era"&gt;Pre-AI era &lt;a class="anchor" href="#pre-ai-era" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Manual coding with IDE autocomplete&lt;/li&gt;&#10;&lt;li&gt;Stack Overflow copy-paste&lt;/li&gt;&#10;&lt;li&gt;Productivity baseline: 1-2 packages per year&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="phase-1-copy-paste-tools-march-2023---july-2025"&gt;Phase 1: copy-paste tools (March 2023 - July 2025) &lt;a class="anchor" href="#phase-1-copy-paste-tools-march-2023---july-2025" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;ChatGPT/Claude Web&lt;/strong&gt; + my &lt;a href="https://www.nijho.lt/post/advent-of-open-source/15-clip-files/"&gt;&lt;code&gt;clip-files&lt;/code&gt;&lt;/a&gt; tool&lt;/li&gt;&#10;&lt;li&gt;Manual context sharing&lt;/li&gt;&#10;&lt;li&gt;Helpful for ideation and code review&lt;/li&gt;&#10;&lt;li&gt;Productivity: 3x improvement (from 2 to 6 packages/year)&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="cursor-early-2024"&gt;Cursor (early 2024) &lt;a class="anchor" href="#cursor-early-2024" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Great autocomplete&lt;/li&gt;&#10;&lt;li&gt;Limited context awareness&lt;/li&gt;&#10;&lt;li&gt;Became frustratingly slow&lt;/li&gt;&#10;&lt;li&gt;Still led to my vibe coding disaster&lt;/li&gt;&#10;&lt;li&gt;Productivity: 5x improvement (but with quality issues)&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="phase-2-claude-code-july-2025---present"&gt;Phase 2: Claude Code (July 2025 - present) &lt;a class="anchor" href="#phase-2-claude-code-july-2025---present" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Full codebase awareness&lt;/li&gt;&#10;&lt;li&gt;Can execute commands and debug&lt;/li&gt;&#10;&lt;li&gt;Self-correcting through test iterations&lt;/li&gt;&#10;&lt;li&gt;Maintains context across sessions&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;$70 spent on first night, $10,000 worth in first month&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;Productivity: 24x improvement over pre-AI baseline with maintained quality&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="openais-codex-cli-the-reasoning-beast"&gt;OpenAI&amp;rsquo;s Codex CLI (the reasoning beast) &lt;a class="anchor" href="#openais-codex-cli-the-reasoning-beast" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Recently tried OpenAI&amp;rsquo;s Codex—their new CLI alternative to Claude Code—and I&amp;rsquo;ll admit, for pure reasoning with their GPT-5-Codex model, it&amp;rsquo;s impressive:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Solved a race condition&lt;/strong&gt; in an unfamiliar language that took me hours with Claude Opus&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;More elegant solutions&lt;/strong&gt; for complex algorithmic problems&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Better at deep reasoning&lt;/strong&gt; when given the same context&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;But here&amp;rsquo;s why I still use Claude Code daily:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude Code&amp;rsquo;s UX is superior&lt;/strong&gt;: Resume conversations, better interface&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Codex&amp;rsquo;s CLI is fragile&lt;/strong&gt;: Hit Ctrl-C accidentally? Lose everything, start over (muscle memory keeps betraying me!)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude Code handles interruptions&lt;/strong&gt;: First Ctrl-C clears text, second quits gracefully&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;No session persistence&lt;/strong&gt;: Codex forgets everything, Claude Code remembers&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Context management&lt;/strong&gt;: Have to re-provide all context after accidental exits&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;For a complex race condition bug, Codex with GPT-5 gave me the most elegant solution.&#10;But for day-to-day development where I need reliability, good UX, and the ability to resume work?&lt;/p&gt;&#10;&lt;h3 id="why-not-copilot-or-other-cheaper-tools"&gt;Why not Copilot or other &amp;ldquo;cheaper&amp;rdquo; tools? &lt;a class="anchor" href="#why-not-copilot-or-other-cheaper-tools" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Let me be blunt: &lt;strong&gt;You get what you pay for.&lt;/strong&gt; People often say &amp;ldquo;I tried Copilot and it wasn&amp;rsquo;t that useful&amp;rdquo; or &amp;ldquo;Gemini in chat didn&amp;rsquo;t help much.&amp;rdquo; Of course not—these tools are &lt;strong&gt;at least 10x cheaper&lt;/strong&gt; than Claude Code with Opus.&lt;/p&gt;&#10;&lt;p&gt;Your $20/month GitHub Copilot subscription simply cannot provide the same quality or quantity of high-quality tokens as Claude Code.&#10;It&amp;rsquo;s like comparing a bicycle to a Ferrari and wondering why the bicycle isn&amp;rsquo;t as fast.&#10;The economics don&amp;rsquo;t work out:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Copilot&lt;/strong&gt;: $20/month for limited autocomplete&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Claude Code Pro&lt;/strong&gt;: $200/month for unlimited* agentic assistance (*capped, but generous)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Actual API value used&lt;/strong&gt;: $10,000+/month at my usage level&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The 10x price difference reflects a 10x (or more) difference in capability.&#10;Copilot is autocomplete on steroids.&#10;Claude Code is a tireless pair programming partner with perfect memory who can debug, test, and iterate.&#10;The ROI is obvious when you&amp;rsquo;re shipping 10x more code.&lt;/p&gt;&#10;&lt;h2 id="12-is-this-sustainable"&gt;12. Is this sustainable? &lt;a class="anchor" href="#12-is-this-sustainable" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;You might wonder if this productivity is sustainable or if I&amp;rsquo;m just building lower-quality software faster.&#10;Here&amp;rsquo;s my honest assessment:&lt;/p&gt;&#10;&lt;h3 id="the-good"&gt;The good &lt;a class="anchor" href="#the-good" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Faster prototyping&lt;/strong&gt;: Ideas to working code in hours, not days&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Better test coverage&lt;/strong&gt;: AI never gets lazy about writing tests&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;More experimentation&lt;/strong&gt;: Lower cost to try new ideas&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Improved documentation&lt;/strong&gt;: AI loves writing docs (I don&amp;rsquo;t)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Reduced burnout&lt;/strong&gt;: Less time on boilerplate, more on interesting problems&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-challenges"&gt;The challenges &lt;a class="anchor" href="#the-challenges" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Code review fatigue&lt;/strong&gt;: You must stay vigilant&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Dependency on tools&lt;/strong&gt;: What if Claude Code disappears?&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Cost&lt;/strong&gt;: $200/month Pro plan (but worth every penny given the $10K+ value)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;The temptation to &amp;ldquo;vibe&amp;rdquo;&lt;/strong&gt;: Constant discipline required&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Addiction potential&lt;/strong&gt;: I literally couldn&amp;rsquo;t stop coding, having the desire to deplete all tokens and get my money&amp;rsquo;s worth!&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="my-verdict"&gt;My verdict &lt;a class="anchor" href="#my-verdict" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;It&amp;rsquo;s absolutely sustainable &lt;strong&gt;if&lt;/strong&gt; you maintain discipline.&#10;The moment you start accepting code you don&amp;rsquo;t understand, you&amp;rsquo;re building a house of cards.&#10;But with proper review and testing, agentic AI is a massive force multiplier.&lt;/p&gt;&#10;&lt;h2 id="13-the-experience-multiplier-why-seniority-matters-more-than-ever"&gt;13. The experience multiplier: why seniority matters more than ever &lt;a class="anchor" href="#13-the-experience-multiplier-why-seniority-matters-more-than-ever" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Here&amp;rsquo;s a counterintuitive truth about agentic AI tools: &lt;strong&gt;they amplify your existing skills exponentially, not linearly&lt;/strong&gt;.&#10;This creates a fascinating paradox where these tools become increasingly valuable as you become more experienced.&lt;/p&gt;&#10;&lt;h3 id="the-exponential-trap-for-beginners"&gt;The exponential trap for beginners &lt;a class="anchor" href="#the-exponential-trap-for-beginners" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;For beginners, agentic tools can feel like a superpower initially.&#10;They can build a working application in hours!&#10;But here&amp;rsquo;s the danger: without the experience to recognize bad patterns, they&amp;rsquo;re essentially building at 10x speed in the wrong direction.&#10;It&amp;rsquo;s like giving a Formula 1 car to someone who just got their driver&amp;rsquo;s license—the speed multiplies both the distance traveled and the potential for catastrophic crashes.&lt;/p&gt;&#10;&lt;p&gt;A beginner using Claude Code might:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Accept complex solutions they don&amp;rsquo;t understand&lt;/li&gt;&#10;&lt;li&gt;Build massive technical debt at record speed&lt;/li&gt;&#10;&lt;li&gt;Miss architectural problems that will haunt them later&lt;/li&gt;&#10;&lt;li&gt;Create code that &amp;ldquo;works&amp;rdquo; but is impossible to maintain&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;They&amp;rsquo;re forced into a difficult position: either slow down to understand everything (negating the speed benefit) or accumulate technical debt at an exponential rate.&lt;/p&gt;&#10;&lt;h3 id="the-senior-developer-advantage"&gt;The senior developer advantage &lt;a class="anchor" href="#the-senior-developer-advantage" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;For experienced developers, it&amp;rsquo;s a completely different game.&#10;When I look at Claude Code&amp;rsquo;s suggestions, I can instantly recognize:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&amp;ldquo;Oh, that&amp;rsquo;s the Factory pattern—makes sense here&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;This error handling is too defensive, let&amp;rsquo;s simplify&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;That&amp;rsquo;s going to cause N+1 queries, need to refactor&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;&amp;ldquo;This violates our project&amp;rsquo;s conventions, fix it&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The review process that might take a beginner hours takes me minutes.&#10;I&amp;rsquo;m not learning what the code does—I&amp;rsquo;m validating that it does what I already know it should do.&lt;/p&gt;&#10;&lt;h3 id="the-right-way-for-every-level"&gt;The right way for every level &lt;a class="anchor" href="#the-right-way-for-every-level" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;That said, everyone can benefit from agentic tools—you just need to adjust your approach:&lt;/p&gt;&#10;&lt;h4 id="for-beginners"&gt;For beginners&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Use AI as a learning accelerator&lt;/strong&gt;: Ask it to explain every decision&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Build small projects first&lt;/strong&gt;: Master fundamentals before scaling up&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Prioritize understanding over speed&lt;/strong&gt;: Better to build slowly and learn&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Pair with mentors&lt;/strong&gt;: Have experienced developers review AI-generated code&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="for-intermediate-developers"&gt;For intermediate developers&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Focus on patterns&lt;/strong&gt;: Use AI to learn architectural patterns and best practices&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Experiment safely&lt;/strong&gt;: Try new approaches in side projects first&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Question everything&lt;/strong&gt;: Don&amp;rsquo;t accept solutions without understanding the tradeoffs&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h4 id="for-senior-developers"&gt;For senior developers&lt;/h4&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Leverage for velocity&lt;/strong&gt;: Use your experience to review and direct quickly&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Focus on architecture&lt;/strong&gt;: Let AI handle implementation while you design systems&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Teach the AI&lt;/strong&gt;: Create project-specific guidelines to improve output quality&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Push boundaries&lt;/strong&gt;: Explore more ambitious projects now that implementation is faster&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-real-multiplier-effect"&gt;The real multiplier effect &lt;a class="anchor" href="#the-real-multiplier-effect" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;The productivity boost isn&amp;rsquo;t just about writing code faster.&#10;For senior developers, agentic tools multiply:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Architecture exploration&lt;/strong&gt;: Test multiple approaches quickly&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Refactoring confidence&lt;/strong&gt;: Refactor fearlessly with AI helping maintain functionality&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Learning velocity&lt;/strong&gt;: Explore new languages and frameworks with a knowledgeable assistant&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Documentation quality&lt;/strong&gt;: Finally have time (and help) for comprehensive docs&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Testing coverage&lt;/strong&gt;: AI never skips tests, even for &amp;ldquo;simple&amp;rdquo; functions&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This is why I&amp;rsquo;ve been able to ship 14 packages in 18 months.&#10;It&amp;rsquo;s not just that I&amp;rsquo;m writing code faster—I&amp;rsquo;m able to explore ideas, validate approaches, and iterate on designs at a pace that was previously impossible.&lt;/p&gt;&#10;&lt;h2 id="14-reality-check-what-ai-can-and-cant-do"&gt;14. Reality check: what AI can and can&amp;rsquo;t do &lt;a class="anchor" href="#14-reality-check-what-ai-can-and-cant-do" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Many people try AI coding and give up disappointed.&#10;That&amp;rsquo;s because they&amp;rsquo;re treating it like a magic genie that will solve the unsolved problems of the universe.&#10;&lt;strong&gt;It won&amp;rsquo;t.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h3 id="what-ai-can-do-the-churn-work"&gt;What AI CAN do (the churn work) &lt;a class="anchor" href="#what-ai-can-do-the-churn-work" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;AI excels at tasks you could do yourself but would take hours:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Setting up CI/CD pipelines&lt;/strong&gt;: Boilerplate YAML configuration&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Writing tests&lt;/strong&gt;: Especially unit tests for existing code&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Fixing dependency issues&lt;/strong&gt;: Adapting to breaking changes in libraries&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;CRUD operations&lt;/strong&gt;: Basic create, read, update, delete functionality&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Data transformations&lt;/strong&gt;: Converting between formats, parsing, serialization&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Documentation&lt;/strong&gt;: README files, docstrings, API documentation&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Refactoring&lt;/strong&gt;: Renaming variables, extracting functions, reorganizing code&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Bug fixes&lt;/strong&gt;: Especially those with clear error messages&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This is the &amp;ldquo;churn&amp;rdquo;—the necessary but time-consuming work that makes up 80% of development.&lt;/p&gt;&#10;&lt;h3 id="what-ai-cant-do-the-creative-work"&gt;What AI CAN&amp;rsquo;T do (the creative work) &lt;a class="anchor" href="#what-ai-cant-do-the-creative-work" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Don&amp;rsquo;t expect AI to:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Solve novel algorithmic problems&lt;/strong&gt;: It won&amp;rsquo;t invent the next PageRank&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Make architectural decisions&lt;/strong&gt;: It can&amp;rsquo;t decide if you need microservices&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Understand business logic&lt;/strong&gt;: It doesn&amp;rsquo;t know why your company does things&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Create truly innovative solutions&lt;/strong&gt;: It remixes what it&amp;rsquo;s seen, not invents&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Make judgment calls&lt;/strong&gt;: Security vs convenience, performance vs maintainability&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-sweet-spot"&gt;The sweet spot &lt;a class="anchor" href="#the-sweet-spot" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I use AI for tasks where:&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;&lt;strong&gt;I know how to do it&lt;/strong&gt; but it would take time&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;The solution exists&lt;/strong&gt; somewhere in its training data&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Success is measurable&lt;/strong&gt; through tests or clear output&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;The problem is well-defined&lt;/strong&gt; with clear constraints&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;This isn&amp;rsquo;t about AI doing things you can&amp;rsquo;t do—it&amp;rsquo;s about AI doing things you don&amp;rsquo;t want to spend time doing.&#10;Once you understand this, AI becomes incredibly powerful.&lt;/p&gt;&#10;&lt;h2 id="15-looking-forward-the-future-of-development"&gt;15. Looking forward: the future of development &lt;a class="anchor" href="#15-looking-forward-the-future-of-development" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;As I showed in my &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;local AI journey&lt;/a&gt;, I&amp;rsquo;m deeply interested in where AI and development intersect.&#10;Agentic coding represents a fundamental shift in how we build software:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;From writing code to reviewing and directing&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;From memorizing APIs to understanding patterns&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;From debugging alone to collaborative problem-solving&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;From slower iteration to rapid experimentation&lt;/strong&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This isn&amp;rsquo;t about replacing developers—it&amp;rsquo;s about amplifying our capabilities.&#10;I&amp;rsquo;m still the architect, the decision-maker, and the quality gatekeeper.&#10;But now I can focus on the interesting problems while my AI assistant handles the implementation details.&lt;/p&gt;&#10;&lt;h2 id="16-conclusion-principles-over-process"&gt;16. Conclusion: principles over process &lt;a class="anchor" href="#16-conclusion-principles-over-process" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My journey from vibe coding skeptic to agentic coding advocate might seem like a complete reversal, but it&amp;rsquo;s not.&#10;My core principle remains unchanged: &lt;strong&gt;never commit code you don&amp;rsquo;t understand&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;What&amp;rsquo;s changed is the tooling has finally caught up to the promise.&#10;With agentic AI assistants, I can maintain my standards while dramatically increasing my output.&#10;Most of those 32 packages on PyPI have users and are well-tested and documented projects that solve real problems.&lt;/p&gt;&#10;&lt;p&gt;The key is treating AI as what it is: an incredibly powerful assistant, not a replacement for thinking.&#10;Use it to explore ideas faster, implement solutions quicker, and test more thoroughly.&#10;But always, always understand what you&amp;rsquo;re building.&lt;/p&gt;&#10;&lt;p&gt;As I continue building in public and sharing my tools, I&amp;rsquo;m excited to see where this technology takes us.&#10;If you&amp;rsquo;re interested in trying agentic coding yourself, start small, maintain your standards, and prepare to be amazed by what you can build.&lt;/p&gt;&#10;&lt;p&gt;&lt;em&gt;What&amp;rsquo;s your experience with AI coding assistants? Have you tried moving from generative to agentic tools? I&amp;rsquo;d love to hear your thoughts!&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="links-and-resources"&gt;Links and resources &lt;a class="anchor" href="#links-and-resources" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;a href="https://claude.ai/code" target="_blank" rel="noopener"&gt;Claude Code&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://pypi.org/user/basnijholt/" target="_blank" rel="noopener"&gt;My PyPI packages&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://github.com/basnijholt" target="_blank" rel="noopener"&gt;My GitHub&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://www.nijho.lt/post/vibe-coding/"&gt;Original vibe coding post&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;My local AI journey&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;a href="https://www.nijho.lt/post/using-llms/"&gt;Using LLMs effectively&lt;/a&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;</description></item></channel></rss>