<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hardware | Bas Nijholt</title><link>https://www.nijho.lt/tag/hardware/</link><atom:link href="https://www.nijho.lt/tag/hardware/index.xml" rel="self" type="application/rss+xml"/><description>Hardware</description><generator>Hugo</generator><language>en-us</language><copyright>© 2026</copyright><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://www.nijho.lt/media/icon_hu_408a8f6cef43e327.png</url><title>Hardware</title><link>https://www.nijho.lt/tag/hardware/</link></image><item><title>Self-hosting AI does not save money, and I do it anyway</title><link>https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/self-hosting-ai-is-not-cheaper/</guid><description>&lt;p&gt;I have had this debate many times, so I am finally writing it down.&lt;/p&gt;&#10;&lt;p&gt;Whenever I say that self-hosting AI does not save money, people hear that I am against self-hosting.&#10;That is not the point at all.&#10;I am a massive fan of open-weight models.&#10;&lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L43" target="_blank" rel="noopener"&gt;Qwen3.8 27B runs on two RTX 3090s at home&lt;/a&gt; (and 15 more models), and my phone dictation goes to &lt;a href="https://www.nijho.lt/post/diction-agent-cli-qwen/"&gt;Qwen3-ASR on the same machine&lt;/a&gt;.&#10;I think open-source AI is the best thing since sliced bread.&lt;/p&gt;&#10;&lt;p&gt;I don&amp;rsquo;t pretend it saves money.&#10;I do it for fun, for sovereignty, and for privacy, which I come back to at the end.&lt;/p&gt;&#10;&lt;p&gt;All benchmark numbers and prices below come from &lt;a href="https://artificialanalysis.ai/" target="_blank" rel="noopener"&gt;Artificial Analysis&lt;/a&gt; and &lt;a href="https://openrouter.ai/" target="_blank" rel="noopener"&gt;OpenRouter&lt;/a&gt; as of September 30, 2026.&#10;I did the math for the hardware I own, and for the counterarguments I hear most.&#10;All assumptions are listed &lt;a href="#appendix-all-assumptions"&gt;at the end&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;Benchmarks are not everything, and a high score does not reliably predict how a model does on real work.&#10;They are still the best we have: all models are benchmaxed, tuned to do well on the popular benchmarks, so the scores at least work as a reference frame for comparing them with each other.&lt;/p&gt;&#10;&lt;p&gt;&lt;em&gt;Update, October 2, 2026: after the discussion on r/LocalLLaMA, I added &lt;a href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;section 5&lt;/a&gt; on Macs, DGX Sparks, and AMD AI Max machines.&lt;/em&gt;&lt;/p&gt;&#10;&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#1-the-model-i-love-and-the-one-that-beats-it-on-price"&gt;1. The model I love, and the one that beats it on price&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#2-but-my-gpus-are-already-paid-for"&gt;2. &amp;ldquo;But my GPUs are already paid for&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#3-why-datacenters-are-so-much-better-at-this"&gt;3. Why datacenters are so much better at this&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#4-but-i-have-solar-panels"&gt;4. &amp;ldquo;But I have solar panels&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;5. &amp;ldquo;But my Mac, DGX Spark, or AI Max does save money&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#6-the-bar-keeps-moving"&gt;6. The bar keeps moving&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#7-a-small-company-has-it-worse"&gt;7. A small company has it worse&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#8-they-lose-money-on-inference-so-prices-will-go-up"&gt;8. &amp;ldquo;They lose money on inference, so prices will go up&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#9-but-they-sell-your-data"&gt;9. &amp;ldquo;But they sell your data&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#10-where-self-hosting-does-win-the-same-model"&gt;10. Where self-hosting does win: the same model&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#11-so-why-do-i-do-it"&gt;11. So why do I do it?&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#appendix-all-assumptions"&gt;Appendix: all assumptions&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;h2 id="1-the-model-i-love-and-the-one-that-beats-it-on-price"&gt;1. The model I love, and the one that beats it on price &lt;a class="anchor" href="#1-the-model-i-love-and-the-one-that-beats-it-on-price" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My favorite local model right now is &lt;a href="https://artificialanalysis.ai/models/qwen3-8-27b" target="_blank" rel="noopener"&gt;Qwen3.8 27B&lt;/a&gt;.&#10;It came out in August and it is Apache-2.0.&#10;People often say models like this run on a single gaming GPU, but that is only true after quantizing them.&#10;At full precision, Qwen3.8 27B needs about 54 GB of memory.&#10;Quantized to about 4 bits per weight, it fits on one 24 GB card, at some cost in quality.&#10;I still &lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L56-L65" target="_blank" rel="noopener"&gt;split it across both of my 3090s&lt;/a&gt;, because the second card leaves room for a larger context window.&lt;/p&gt;&#10;&lt;p&gt;I compare it with OpenAI&amp;rsquo;s &lt;a href="https://artificialanalysis.ai/models/gpt-6-luna" target="_blank" rel="noopener"&gt;GPT-6 Luna&lt;/a&gt;, the cheap tier released on September 22.&#10;Both models let you choose how long they think, but the settings do not mean the same thing for both.&#10;Luna&amp;rsquo;s token use grows more than 20 times from its lowest setting to its highest, while Qwen&amp;rsquo;s barely changes.&#10;Qwen on low already writes more tokens than Luna on xhigh.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-effortTokens-2" class="plot-container" role="img" aria-label="Average output tokens per Intelligence Index task at each reasoning setting. Qwen3.8 27B has no high or max setting. Hover for scores."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Average output tokens per Intelligence Index task at each reasoning setting. Qwen3.8 27B has no high or max setting. Hover for scores.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-effortTokens-2");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["effortTokens"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;So the name of a setting says little on its own.&#10;I compare both models at xhigh, the highest setting Qwen offers, and count the tokens each one actually uses.&#10;The extra tokens Qwen needs are part of what I am measuring.&#10;There, Qwen scores 33.7 on the &lt;a href="https://artificialanalysis.ai/methodology/intelligence-benchmarking" target="_blank" rel="noopener"&gt;Artificial Analysis Intelligence Index&lt;/a&gt; and Luna scores 34.6.&#10;For reference, Claude Opus 4.6, a frontier model from February, scores 31.9.&#10;That is an amazing feat in itself.&#10;When Opus 4.6 was the best model we had, I never thought that within the same year I could run essentially that level of intelligence in my own house.&lt;/p&gt;&#10;&lt;p&gt;Artificial Analysis also &lt;a href="https://artificialanalysis.ai/leaderboards/models" target="_blank" rel="noopener"&gt;publishes&lt;/a&gt; how many tokens each model used to run the index, and what that cost.&#10;So the question is simple: what does it cost to run the whole benchmark once?&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-runCost-3" class="plot-container" role="img" aria-label="Cost of one run of the Artificial Analysis Intelligence Index. The two bars for hardware at home are electricity only: measured on my 3090s, estimated for the RTX PRO 6000."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Cost of one run of the Artificial Analysis Intelligence Index. The two bars for hardware at home are electricity only: measured on my 3090s, estimated for the RTX PRO 6000.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-runCost-3");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["runCost"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Luna runs the whole benchmark for $67.&#10;Qwen at full precision, through the &lt;a href="https://openrouter.ai/qwen/qwen3.8-27b/providers" target="_blank" rel="noopener"&gt;cheapest provider&lt;/a&gt; that does not keep your data (ZDR), costs $619.&#10;Two things cause that gap: providers charge almost four times as much per output token for Qwen, and Qwen generates almost three times as many tokens to get the same work done.&lt;/p&gt;&#10;&lt;p&gt;The second bar is my own machine: the electricity alone for running Qwen on my GPUs costs more than Luna&amp;rsquo;s entire bill.&lt;/p&gt;&#10;&lt;p&gt;The third bar is what it takes to match the API&amp;rsquo;s full precision at home: an RTX PRO 6000, a workstation card with 96 GB of memory, enough for the full model.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt;&#10;Its electricity alone costs twice Luna&amp;rsquo;s bill, and one run takes more than seven weeks.&#10;The card by itself sells for $10,000 to $20,000, so with a computer around it, you are paying for a pretty nice car.&#10;Buy a few of them to run agents in parallel, and you are building a small data center of your own.&lt;/p&gt;&#10;&lt;p&gt;The bars for hardware at home are electricity only, so they are the lowest these runs can cost: the cards come on top, and how much depends on how busy you keep them, which is what section 4 is about.&lt;/p&gt;&#10;&lt;p&gt;What I care about is the cheapest way to get answers of Qwen&amp;rsquo;s quality, from whichever model gives them.&lt;/p&gt;&#10;&lt;h2 id="2-but-my-gpus-are-already-paid-for"&gt;2. &amp;ldquo;But my GPUs are already paid for&amp;rdquo; &lt;a class="anchor" href="#2-but-my-gpus-are-already-paid-for" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;This is the argument I hear most, so assume the hardware is free and only the electricity counts.&lt;/p&gt;&#10;&lt;p&gt;I measured my own machine.&#10;It runs Qwen in vLLM, split across both cards, at about 107 tokens per second.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt;&#10;The cards are capped at 270 W each,&lt;sup id="fnref:4"&gt;&lt;a href="#fn:4" class="footnote-ref" role="doc-noteref"&gt;4&lt;/a&gt;&lt;/sup&gt; and the whole machine draws about 700 W.&#10;It then needs four weeks, running day and night, to get through the benchmark once.&#10;I also give my quantized copy the score Artificial Analysis measured through Alibaba&amp;rsquo;s API.&#10;I have not rerun the benchmark on it, and anything lost to quantization makes my machine look worse.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-electricityCost-4" class="plot-container" role="img" aria-label="Electricity for one benchmark run of Qwen3.8 27B on my two 3090s, by electricity price. The dashed line is what Luna charges for the same work, with OpenAI&amp;rsquo;s hardware and margin included. Hover for values."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Electricity for one benchmark run of Qwen3.8 27B on my two 3090s, by electricity price. The dashed line is what Luna charges for the same work, with OpenAI&amp;rsquo;s hardware and margin included. Hover for values.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-electricityCost-4");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["electricityCost"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Above 14.5 cents per kWh, my electricity alone costs more than Luna&amp;rsquo;s API bill.&#10;Below it, I save a few dollars per run and wait four weeks instead of a few hours.&lt;/p&gt;&#10;&lt;p&gt;Those four weeks matter more than the few dollars.&#10;An API takes hundreds of requests at once, so even one benchmark run can finish in an hour or two.&#10;My setup works on one conversation at a time: when I sent it eight requests at once, seven of them waited in line.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-wallClock-5" class="plot-container" role="img" aria-label="Wall-clock time for one benchmark run. The API times use Artificial Analysis&amp;rsquo;s measured time per task for Luna; the parallel case assumes 100 requests at once. Log scale."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Wall-clock time for one benchmark run. The API times use Artificial Analysis&amp;rsquo;s measured time per task for Luna; the parallel case assumes 100 requests at once. Log scale.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-wallClock-5");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["wallClock"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;This is one more reason my coding agents run on APIs.&#10;When I &lt;a href="https://www.nijho.lt/post/parallel-agentic-coding/"&gt;run several agents in parallel&lt;/a&gt;, each of them should be as fast as if it were alone.&lt;/p&gt;&#10;&lt;h2 id="3-why-datacenters-are-so-much-better-at-this"&gt;3. Why datacenters are so much better at this &lt;a class="anchor" href="#3-why-datacenters-are-so-much-better-at-this" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;When my 3090s generate a token, they read all of the model&amp;rsquo;s weights from memory to produce one token for one conversation.&#10;Most of the chip&amp;rsquo;s compute sits idle while it waits for memory.&#10;A datacenter GPU reads the same weights once and produces a token for hundreds of conversations at the same time.&#10;This is called batching, and it is where most of the efficiency comes from.&lt;/p&gt;&#10;&lt;p&gt;Artificial Analysis measures this with &lt;a href="https://artificialanalysis.ai/hardware-inference-stack/datacenter" target="_blank" rel="noopener"&gt;AgentPerf&lt;/a&gt;, which replays real coding-agent sessions against datacenter hardware.&#10;It counts how many agents a system can serve while each one still gets at least 60 tokens per second.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-agentsPerKw-7" class="plot-container" role="img" aria-label="Coding agents served per kW of accelerator power, each getting at least 60 tokens per second, from AgentPerf. The datacenter systems all run DeepSeek V4 Pro. The 3090 bar is my machine running Qwen3.8 27B, a much smaller model, with GPU power only, so the comparison favors my machine."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Coding agents served per kW of accelerator power, each getting at least 60 tokens per second, from AgentPerf. The datacenter systems all run DeepSeek V4 Pro. The 3090 bar is my machine running Qwen3.8 27B, a much smaller model, with GPU power only, so the comparison favors my machine.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-agentsPerKw-7");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["agentsPerKw"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;The rack of 36 GB300s serves 20 times more agents per kW than the 8-GPU H200 server.&#10;Part of that is newer chips, and part is software: NVIDIA tuned the B300 and GB300 setups itself, while Artificial Analysis configured the H200 one.&#10;The B300 and the GB300 share both the chip generation and the software.&#10;Most of the gap between those two bars comes from the rack itself: 36 GPUs on one fast interconnect keep more than a thousand agents in flight at once.&#10;My two 3090s serve one conversation at a time.&lt;/p&gt;&#10;&lt;h2 id="4-but-i-have-solar-panels"&gt;4. &amp;ldquo;But I have solar panels&amp;rdquo; &lt;a class="anchor" href="#4-but-i-have-solar-panels" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Solar power is not free.&#10;With net metering, every kWh my GPUs burn is a kWh that does not get credited at the retail price.&#10;Without net metering, it is a kWh I do not sell, unless the surplus would go to waste anyway, which is the free-power case below.&#10;And the GPUs also run at night.&lt;/p&gt;&#10;&lt;p&gt;But fine, say the electricity is free and the only cost is the hardware.&#10;I was lucky and bought my two 3090s for about $750 each,&lt;sup id="fnref:5"&gt;&lt;a href="#fn:5" class="footnote-ref" role="doc-noteref"&gt;5&lt;/a&gt;&lt;/sup&gt; and I write them off over three years.&#10;The GPUs only save money while they do work I would otherwise pay an API for, and while they sit idle they save nothing.&#10;So the question is how many hours a day they have to be busy before they pay for themselves.&lt;sup id="fnref:6"&gt;&lt;a href="#fn:6" class="footnote-ref" role="doc-noteref"&gt;6&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;&#10;&lt;p&gt;To make this fair to open models, imagine an open model that is exactly as efficient as Luna: the same quality from the same number of tokens, sold at Luna&amp;rsquo;s price.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-breakEven-8" class="plot-container" role="img" aria-label="How many hours a day my two $750 RTX 3090s have to be busy, every day for three years, before they cost less than a ZDR API. The host PC is not counted."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;How many hours a day my two $750 RTX 3090s have to be busy, every day for three years, before they cost less than a ZDR API. The host PC is not counted.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-breakEven-8");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["breakEven"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;With free power, the GPUs pay for themselves if they are busy 4 hours a day, every day for three years.&#10;At 18 cents per kWh, they need 8.5.&#10;That uses what I paid for the cards; at today&amp;rsquo;s price of about $1,500 each, the 8.5 hours become 14.&lt;/p&gt;&#10;&lt;p&gt;And API prices keep falling.&#10;GPT-6 Luna costs 58% less per output token than GPT-5.6 Luna did in July.&#10;If the price halves again, no amount of use pays off at 18 cents per kWh.&#10;If it halves twice, the API costs less than my electricity alone.&lt;/p&gt;&#10;&lt;h2 id="5-but-my-mac-dgx-spark-or-ai-max-does-save-money"&gt;5. &amp;ldquo;But my Mac, DGX Spark, or AI Max does save money&amp;rdquo; &lt;a class="anchor" href="#5-but-my-mac-dgx-spark-or-ai-max-does-save-money" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Several readers told me their machines do save them money: a MacBook Pro with an M5 Max, an NVIDIA DGX Spark, and a computer with AMD&amp;rsquo;s Ryzen AI Max+ 395.&#10;If you would have bought the machine anyway, that is true, for the reason in section 2: their electricity costs little compared with the API.&#10;Counting the hardware is a different story.&lt;/p&gt;&#10;&lt;p&gt;With 128 GB of memory, all three can hold a better model than my 3090s: Qwen3.8-Flash-Next, a sparse model with 180 billion parameters, of which only 6 billion work on each token.&#10;It scores 39.8, close to GPT-6 Luna on its highest setting (38.1), although at the Q2 a reader uses on their Mac it probably loses a few points.&lt;sup id="fnref:7"&gt;&lt;a href="#fn:7" class="footnote-ref" role="doc-noteref"&gt;7&lt;/a&gt;&lt;/sup&gt;&#10;So how many years would each machine have to run nonstop before it pays for itself?&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-paybackYears-9" class="plot-container" role="img" aria-label="Years of nonstop generation before each machine pays for itself, compared with the most efficient API model at a similar score. Each machine runs the best model it can hold."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Years of nonstop generation before each machine pays for itself, compared with the most efficient API model at a similar score. Each machine runs the best model it can hold.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-paybackYears-9");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["paybackYears"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;Even with free power, the 128 GB machines need six to ten years.&#10;The AMD machine draws only about 120 W, but it is slow too: one benchmark run of Qwen3.8 27B takes it four and a half months, and almost as much energy as my 3090s.&#10;The RTX PRO 6000 runs the same model about three times as fast but costs about twice as much, so it needs six to nine years.&#10;My 3090s pay off in under two years with free power, which is why the 3090 is still the value king, but at 18 cents per kWh they never do: the electricity for one run already costs more than Luna charges for it.&lt;/p&gt;&#10;&lt;p&gt;Ten years is a long time in AI.&#10;Better and more efficient open models will run on the same machines, but the frontier moves at the same pace, and API prices keep falling.&#10;Hardware gets more efficient too, in datacenters and at home, so a machine bought today also ends up competing with next year&amp;rsquo;s machines.&#10;These numbers only show how far apart the two sides start.&lt;/p&gt;&#10;&lt;h2 id="6-the-bar-keeps-moving"&gt;6. The bar keeps moving &lt;a class="anchor" href="#6-the-bar-keeps-moving" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Section 4&amp;rsquo;s eight and a half hours a day would be easy for me: I run agents for longer than that, just not on Qwen.&#10;When Claude Opus 4.6 came out in February, I was perfectly happy with it.&#10;I thought it was all I would ever need, and I could not have imagined how much better models would get in half a year.&#10;Qwen3.8 27B now scores about the same as Opus 4.6, and I would no longer accept it for coding.&#10;Until GPT-6 Astra came out at the start of September, my go-to was GPT-5.6 Sol, which scores ten points higher than Qwen.&#10;I then used Astra until Claude Opus 5.5 came out less than three weeks later.&#10;Now I don&amp;rsquo;t even accept what was considered the best model a month ago.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-frontierGap-10" class="plot-container" role="img" aria-label="Intelligence Index of the best model available on each date, of the models I used as my go-to, and of Qwen&amp;rsquo;s 27B models, which fit on a single 3090. Scores use the xhigh reasoning setting where a model has it."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Intelligence Index of the best model available on each date, of the models I used as my go-to, and of Qwen&amp;rsquo;s 27B models, which fit on a single 3090. Scores use the xhigh reasoning setting where a model has it.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-frontierGap-10");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["frontierGap"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;The models that fit on my 3090s keep improving, but they stay about half a year behind the frontier, and my standard moves with the frontier.&lt;/p&gt;&#10;&lt;h2 id="7-a-small-company-has-it-worse"&gt;7. A small company has it worse &lt;a class="anchor" href="#7-a-small-company-has-it-worse" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;For me, this is a hobby, and a hobby is allowed to be inefficient.&#10;It gets worse once you try to use self-hosted models seriously, say for a team of ten developers.&lt;/p&gt;&#10;&lt;p&gt;If you stay fully self-hosted and want fast responses at peak, you have to buy hardware for the busiest hour of the year, not for the average one.&#10;The rest of the time, it sits idle.&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-peakLoad-11" class="plot-container" role="img" aria-label="An illustrative week of coding agents for a ten-person team. The hardware has to cover the busiest hour of the year, but a typical week uses only a fraction of it."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;An illustrative week of coding agents for a ten-person team. The hardware has to cover the busiest hour of the year, but a typical week uses only a fraction of it.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-peakLoad-11");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["peakLoad"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;In this made-up but realistic week, the team uses 17% of what it paid for, so every token costs about six times more than it would at full load.&#10;You also need a spare GPU for when one dies, and someone who gets paged when it does.&#10;You could queue work or send the peaks to an API, but then you are paying for an API anyway.&lt;/p&gt;&#10;&lt;p&gt;An API provider has the opposite situation.&#10;It serves thousands of customers across every time zone, so its load curve is much flatter than yours.&#10;It fills the nights with discounted batch jobs (Luna&amp;rsquo;s batch tier is half price) and training runs.&#10;And more concurrent requests mean bigger batches, which is where the efficiency from section 3 comes from.&lt;/p&gt;&#10;&lt;h2 id="8-they-lose-money-on-inference-so-prices-will-go-up"&gt;8. &amp;ldquo;They lose money on inference, so prices will go up&amp;rdquo; &lt;a class="anchor" href="#8-they-lose-money-on-inference-so-prices-will-go-up" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Another argument I hear is that API prices are subsidized by investors and will go up once those investors want their money back.&#10;I don&amp;rsquo;t think that holds for inference.&lt;/p&gt;&#10;&lt;p&gt;AgentPerf also reports how many tokens each GPU serves.&#10;On the GB300 rack, one GPU serving DeepSeek V4 Pro handles about 4 million output tokens and 530 million input tokens per hour.&#10;Most of that input is conversation history that coding agents send again with every step.&#10;Providers keep it in a cache and charge less than 1% of the normal input price for it.&lt;/p&gt;&#10;&lt;p&gt;At &lt;a href="https://artificialanalysis.ai/models/deepseek-v4-pro-0424" target="_blank" rel="noopener"&gt;DeepSeek V4 Pro&amp;rsquo;s list prices&lt;/a&gt;, that GPU brings in at least $5.70 per hour, even if every input token is billed at the cache price.&#10;If 5% of the input misses the cache, it brings in about $17.&#10;Renting a B200, the closest GPU with a &lt;a href="https://artificialanalysis.ai/hardware-inference-stack/datacenter" target="_blank" rel="noopener"&gt;published rental price&lt;/a&gt;, costs $3.50 to $5.90 per hour from smaller cloud providers, and that price already includes their profit.&lt;/p&gt;&#10;&lt;p&gt;So the tokens pay for the hardware that serves them, even in the worst case.&#10;This leaves out costs like staff and training, so it does not tell you whether a lab makes money overall.&lt;/p&gt;&#10;&lt;p&gt;The money goes to training, research, free users, and flat-rate subscriptions for heavy users like me.&#10;When I wrote that I used &lt;a href="https://www.nijho.lt/post/agentic-coding/"&gt;$10,000 worth of API tokens for $200&lt;/a&gt;, that was list price, not what those tokens cost to serve.&lt;/p&gt;&#10;&lt;p&gt;And prices go down, not up.&#10;This is what the same level of intelligence has cost this year:&lt;/p&gt;&#10;&lt;figure class="plot-figure"&gt;&#10; &lt;div id="plot-priceDrop-13" class="plot-container" role="img" aria-label="Output price per million tokens for models scoring between 30 and 35 on the Intelligence Index, by release date, as listed by Artificial Analysis. Log scale."&gt;&lt;/div&gt;&#10; &lt;figcaption&gt;Output price per million tokens for models scoring between 30 and 35 on the Intelligence Index, by release date, as listed by Artificial Analysis. Log scale.&lt;/figcaption&gt;&#10; &lt;noscript&gt;This chart needs JavaScript.&lt;/noscript&gt;&#10;&lt;/figure&gt;&#10;&lt;script type="module"&gt;&#10; import * as Plot from "https://cdn.jsdelivr.net/npm/@observablehq/plot@0.6.17/+esm";&#10; import plots from "\/post\/self-hosting-ai-is-not-cheaper\/plots.js";&#10;&#10; const container = document.getElementById("plot-priceDrop-13");&#10; let lastWidth = 0;&#10; new ResizeObserver(() =&gt; {&#10; const width = Math.round(container.clientWidth);&#10; if (!width || width === lastWidth) return;&#10; lastWidth = width;&#10; container.replaceChildren(plots["priceDrop"]({ Plot, width }));&#10; &#10; container.querySelectorAll("svg[fill]").forEach((svg) =&gt; (svg.style.fill = svg.getAttribute("fill")));&#10; &#10; container.querySelectorAll("svg [aria-label]").forEach((element) =&gt; element.removeAttribute("aria-label"));&#10; }).observe(container);&#10;&lt;/script&gt;&#10;&#10;&lt;p&gt;From Claude Opus 4.6 in February to GPT-6 Luna in September, the output price for this level of intelligence dropped by a factor of 50.&#10;Labs are in a race to the bottom on price.&lt;/p&gt;&#10;&lt;h2 id="9-but-they-sell-your-data"&gt;9. &amp;ldquo;But they sell your data&amp;rdquo; &lt;a class="anchor" href="#9-but-they-sell-your-data" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Some people say APIs are cheap because the provider trains on your data or sells it.&#10;That is why I only used prices from providers that OpenRouter lists as &lt;a href="https://openrouter.ai/docs/features/zdr" target="_blank" rel="noopener"&gt;zero data retention&lt;/a&gt; (ZDR).&lt;/p&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Model&lt;/th&gt;&#10;&lt;th&gt;Provider&lt;/th&gt;&#10;&lt;th&gt;ZDR&lt;/th&gt;&#10;&lt;th&gt;Output ($ per million tokens)&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;GPT-6 Luna&lt;/td&gt;&#10;&lt;td&gt;OpenAI&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;0.50&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;GPT-6 Luna&lt;/td&gt;&#10;&lt;td&gt;Azure&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;0.50&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;AkashML (FP8)&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;1.78&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;DeepInfra (full precision)&lt;/td&gt;&#10;&lt;td&gt;yes&lt;/td&gt;&#10;&lt;td&gt;1.88&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;Alibaba&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;2.55&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Qwen3.8 27B&lt;/td&gt;&#10;&lt;td&gt;Cloudflare&lt;/td&gt;&#10;&lt;td&gt;no&lt;/td&gt;&#10;&lt;td&gt;3.20&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Azure serves GPT-6 Luna with zero data retention at the same price as OpenAI.&#10;For Qwen3.8 27B, the cheapest endpoints are all ZDR, and the endpoints without ZDR cost more.&#10;The &lt;code&gt;:free&lt;/code&gt; tier is where you pay with your data.&lt;/p&gt;&#10;&lt;h2 id="10-where-self-hosting-does-win-the-same-model"&gt;10. Where self-hosting does win: the same model &lt;a class="anchor" href="#10-where-self-hosting-does-win-the-same-model" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;If you compare running Qwen3.8 27B yourself with paying for Qwen3.8 27B through an API, self-hosting wins.&lt;/p&gt;&#10;&lt;p&gt;DeepInfra, the cheapest full-precision ZDR provider, charges $619 for one benchmark run, while my electricity costs $83.&#10;That compares the full model with my quantized copy, so the gap overstates my advantage by whatever quantization costs in quality.&#10;My machine pays for itself if it is busy 2.4 hours a day, or 1.5 hours with free power (the first group in the chart in section 4).&lt;/p&gt;&#10;&lt;p&gt;Providers charge a lot for a dense 27B model, because every token runs through all 27 billion parameters.&#10;Sparse open models, which use only a small part of their weights for each token, can be much cheaper to serve.&#10;If you specifically need Qwen and use it heavily, self-hosting it pays off.&#10;But you don&amp;rsquo;t need Qwen to get answers of Qwen&amp;rsquo;s quality.&#10;Luna gives you those for $67, with zero data retention.&lt;/p&gt;&#10;&lt;h2 id="11-so-why-do-i-do-it"&gt;11. So why do I do it? &lt;a class="anchor" href="#11-so-why-do-i-do-it" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;First, it is fun.&#10;I like knowing how the whole stack works, from the drivers in &lt;a href="https://www.nijho.lt/post/llama-nixos/"&gt;my NixOS configuration&lt;/a&gt; to how the layers are split between my two GPUs.&lt;/p&gt;&#10;&lt;p&gt;Second, nobody can take it away.&#10;Artificial Analysis already lists GPT-5.6 Luna as deprecated, less than three months after its release.&#10;My Qwen weights will still be on my disk in ten years, and no company can change their license or &lt;a href="https://www.nijho.lt/post/truenas-to-nixos/"&gt;close their build system&lt;/a&gt; on me.&lt;/p&gt;&#10;&lt;p&gt;Third, some data should not leave my house.&#10;I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not.&#10;That is what &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;my local AI projects&lt;/a&gt; are for.&lt;/p&gt;&#10;&lt;p&gt;Those are good reasons.&#10;Saving money is not one of them.&lt;/p&gt;&#10;&lt;h2 id="appendix-all-assumptions"&gt;Appendix: all assumptions &lt;a class="anchor" href="#appendix-all-assumptions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;details&gt;&#10; &lt;summary&gt;Show all assumptions, and which side each one favors&lt;/summary&gt;&#10; &lt;p&gt;&lt;strong&gt;Assumptions that favor self-hosting&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The host PC around the GPUs is not counted, and my 3090s count at the $750 each I paid, not the $1,500 they sell for today.&lt;/li&gt;&#10;&lt;li&gt;Quantized local copies get the score Artificial Analysis measured through an API, with no penalty for quantization.&lt;/li&gt;&#10;&lt;li&gt;The payback chart assumes each machine is busy 24 hours a day, with no idle time.&lt;/li&gt;&#10;&lt;li&gt;API prices stay where they are, although GPT-6 Luna&amp;rsquo;s output price is 58% lower than GPT-5.6 Luna&amp;rsquo;s, two and a half months earlier.&lt;/li&gt;&#10;&lt;li&gt;My time to set up and maintain the machine is not counted.&lt;/li&gt;&#10;&lt;li&gt;Every machine also gets a case with free electricity.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;strong&gt;Assumptions that favor the API&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The example electricity price is 18 cents per kWh, and some people pay less.&lt;/li&gt;&#10;&lt;li&gt;My machine serves one request at a time; batching several requests could raise its total throughput.&lt;/li&gt;&#10;&lt;li&gt;The hardware has no resale value at the end.&lt;/li&gt;&#10;&lt;li&gt;Waste heat that replaces heating in winter is not counted.&lt;/li&gt;&#10;&lt;li&gt;API outages, rate limits, and deprecated models are not counted.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;strong&gt;Inputs&lt;/strong&gt;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Benchmark: the Artificial Analysis Intelligence Index on September 30, 2026, at the xhigh reasoning setting, or the highest setting where noted. Its token counts per model are the unit of work.&lt;/li&gt;&#10;&lt;li&gt;API prices: the cheapest zero-data-retention endpoints on OpenRouter, with cached input billed at the cache price.&lt;/li&gt;&#10;&lt;li&gt;My machine: two RTX 3090s capped at 270 W, running Qwen3.8 27B in vLLM with AutoRound INT4 and DFlash2. Measured: 107 tokens per second, 1,600 tokens per second of prefill, and 540 W for both GPUs. Estimated: about 700 W for the whole machine under load and 150 W idle.&lt;/li&gt;&#10;&lt;li&gt;Break-even: the cards are written off over three years; busy hours save what the API would charge minus the electricity, and idle hours cost the idle power.&lt;/li&gt;&#10;&lt;li&gt;Other machines: prices, power, and speeds as in the footnotes. The M5 Max and RTX PRO 6000 numbers come from readers, and the DGX Spark and AMD Ryzen AI Max+ 395 speeds are estimated from Artificial Analysis measurements.&lt;/li&gt;&#10;&lt;li&gt;Datacenter: AgentPerf measures accelerator power only, and the B200 rental price stands in for the GB300, which has no published rental price.&lt;/li&gt;&#10;&lt;li&gt;Small company: the peak-load week is made up.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&#10;&lt;/details&gt;&#10;&#10;&lt;div class="footnotes" role="doc-endnotes"&gt;&#10;&lt;hr&gt;&#10;&lt;ol&gt;&#10;&lt;li id="fn:1"&gt;&#10;&lt;p&gt;For each model, Artificial Analysis publishes how many input and output tokens the whole index took, and how many of the input tokens were read from a cache. I multiplied those by each provider&amp;rsquo;s prices, with cached input at the cache price. For Qwen that is 198 million output tokens and 820 million input tokens that miss the cache. For the bars at home, the same tokens divided by the machine&amp;rsquo;s speed give the run time, with the uncached input processed at the 1,600 tokens per second I measured on my 3090s and an estimated 2,000 on the RTX PRO 6000, and the run time times the machine&amp;rsquo;s power and the price per kWh gives the electricity.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:2"&gt;&#10;&lt;p&gt;The RTX PRO 6000 has the same memory bandwidth as an RTX 5090, and at full precision every token reads three times as many bytes as in Q4_K_M. Scaling Artificial Analysis&amp;rsquo;s 5090 measurement gives about 48 tokens per second. I assume 600 W for the whole machine.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:3"&gt;&#10;&lt;p&gt;I measured the model as I normally run it: vLLM with an AutoRound INT4 quant and DFlash2 speculative decoding, split across both cards. Single 1,500-token answers came out at 105 to 109 tokens per second, and reading a 22,000-token prompt ran at 1,600 tokens per second. Sending 2, 4, or 8 requests at once did not raise the total above about 110 tokens per second. The two GPUs drew 540 W together, each at its 270 W power limit; the 700 W adds my estimate for the rest of the machine, since I could not measure the CPU. The setup follows the recipes from &lt;a href="https://github.com/noonghunna/club-3090" target="_blank" rel="noopener"&gt;club-3090&lt;/a&gt;, a community project that tunes LLM serving for RTX 3090s and publishes measured numbers, so I think it is about as fast as these cards get. Its benchmarks for this configuration at a similar power limit match mine: about 110 tokens per second for prose, and up to about 195 for code, where speculative decoding guesses more tokens right. Most of Qwen&amp;rsquo;s output on the benchmark is reasoning, which runs at the prose speed.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:4"&gt;&#10;&lt;p&gt;The stock limit is 350 W. My cards sit in a &lt;a href="https://www.nijho.lt/post/local-ai-journey/"&gt;normal desktop case with normal fans&lt;/a&gt;, and I worry that running both at full power for hours would shorten their life. The cap is &lt;a href="https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/nvidia-undervolt.nix#L1-L7" target="_blank" rel="noopener"&gt;a few lines in my NixOS configuration&lt;/a&gt;, and measurements from Puget Systems and r/LocalLLaMA put a 3090 at about 95% of its speed at 270 W. That is 77% of the power for 95% of the speed, so the cap also lowers my electricity cost per token.&amp;#160;&lt;a href="#fnref:4" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:5"&gt;&#10;&lt;p&gt;The 3090 is still the value king for VRAM per dollar, and $750 badly understates what mine are worth. That is what I paid more than a year ago; today a used 3090 sells for about $1,500. Even at that price it costs $62.50 per GB of VRAM, the same as an RTX 5090 at its $2,000 launch price, which is not what a 5090 sells for today. The right number for this calculation is what I could sell my cards for, and at $1,500 each the break-even points rise by about two thirds: at 18 cents per kWh, the same-model case moves from 2.4 to 4.0 hours per day, and the Luna-efficient case from 8.5 to 14.2 hours per day.&amp;#160;&lt;a href="#fnref:5" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:6"&gt;&#10;&lt;p&gt;While the cards are busy, they save what the API would have charged for the same work, minus the electricity at 700 W. While they are idle, the machine still draws about 150 W, of which the two GPUs with the model loaded take 85 W. The break-even point is the number of busy hours per day at which those savings cover the price of the cards, written off over three years, plus the idle power.&amp;#160;&lt;a href="#fnref:6" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li id="fn:7"&gt;&#10;&lt;p&gt;The M5 Max numbers come from the reader: Qwen3.8-Flash-Next at Q2 (about 2.7 bits per weight), about 45 tokens per second, 1,250 tokens per second of prefill, and about 90 W, on a $7,000 laptop. The RTX PRO 6000 numbers come from another reader, &lt;a href="https://www.reddit.com/r/LocalLLaMA/comments/1ww2jsu/comment/pdhs3ea/" target="_blank" rel="noopener"&gt;u/TheAILegend&lt;/a&gt;, who measured Qwen3.8-Flash-Next at NVFP4 on theirs. For a single request, it stays at about 150 tokens per second from 4,000 to 200,000 tokens of context, with about 10,000 tokens per second of prefill and 266 W for the GPU plus about 160 W for the rest of the PC. With 32 short requests in parallel, the card reaches 590 tokens per second and pays off in about two years, which works for short batch jobs. At the long contexts of this benchmark, parallel requests are slower than one, so the chart uses the single-request speed. I use $15,000, the middle of the card&amp;rsquo;s price range. The DGX Spark and the AMD machine run Qwen3.8-Flash-Next at about 4.5 bits. For their speed, I took Artificial Analysis&amp;rsquo;s measurements of Qwen3.8 27B on each machine and doubled them, because the owner of the AMD machine finds Qwen3.8-Flash-Next about twice as fast as Qwen3.8 27B. That gives about 55 tokens per second on the DGX Spark ($6,950, about 150 W) and 47 on the AMD Ryzen AI Max+ 395 ($4,000, the price Artificial Analysis lists, about 120 W). All of them are compared with GPT-6 Luna on its highest setting, which runs the whole benchmark for $122. The 3090s run Qwen3.8 27B, as in the rest of this post, and are compared with Luna on xhigh. Artificial Analysis measured Qwen3.8 27B on the AMD machine at 23 tokens per second, so one benchmark run takes it 134 days and about 390 kWh, against about 460 kWh on my 3090s. Running Qwen3.8-Flash-Next on the 3090s with CPU expert offload, which club-3090 measured at 37 tokens per second, does not change their result: one run then takes 82 days, and at 18 cents per kWh its electricity costs more than Luna&amp;rsquo;s $122 unless the whole machine draws less than about 350 W.&amp;#160;&lt;a href="#fnref:7" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;/div&gt;&#10;</description></item><item><title>Switching from Ollama to llama-swap + llama.cpp on NixOS: the power user's choice 🦙</title><link>https://www.nijho.lt/post/llama-nixos/</link><pubDate>Sat, 29 Nov 2025 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/llama-nixos/</guid><description>&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#the-draft-i-never-published"&gt;The draft I never published&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-original-trigger-broken-models"&gt;The original trigger: broken models&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-catalyst-dual-gpus-and-control"&gt;The catalyst: dual GPUs and control&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#api-compatibility-and-the-switch"&gt;API compatibility and the &amp;ldquo;switch&amp;rdquo;&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#my-nixos-setup"&gt;My NixOS setup&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#compiler-flags-matter"&gt;Compiler flags matter&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#what-i-gained-and-lost"&gt;What I gained (and lost)&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#advantages-of-llama-swap--llamacpp"&gt;Advantages of llama-swap + llama.cpp&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-trade-off"&gt;The trade-off&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#conclusion-maturing-as-an-ai-hobbyist"&gt;Conclusion: maturing as an AI hobbyist&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;h2 id="the-draft-i-never-published"&gt;The draft I never published &lt;a class="anchor" href="#the-draft-i-never-published" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I started writing this blog post four months ago.&#10;Back then, I was frustrated because a specific model I wanted to try—&lt;code&gt;gpt-oss-20b&lt;/code&gt;—was completely broken in &lt;a href="https://ollama.com/" target="_blank" rel="noopener"&gt;Ollama&lt;/a&gt;.&#10;I went down the rabbit hole, set up &lt;a href="https://github.com/mostlygeek/llama-swap" target="_blank" rel="noopener"&gt;&lt;code&gt;llama-swap&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://github.com/ggml-org/llama.cpp" target="_blank" rel="noopener"&gt;&lt;code&gt;llama.cpp&lt;/code&gt;&lt;/a&gt;, got the model working perfectly, wrote half this post&amp;hellip; and then I stopped.&lt;/p&gt;&#10;&lt;p&gt;Why? &lt;strong&gt;Laziness.&lt;/strong&gt;&#10;Once the novelty of that specific model wore off, I drifted back to Ollama.&#10;The real convenience of Ollama was never the &lt;code&gt;ollama run&lt;/code&gt; command—I almost exclusively interact with my models via the API using my own tools like &lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;agent-cli&lt;/a&gt;.&#10;The convenience was &lt;code&gt;ollama pull&lt;/code&gt;.&#10;Being able to grab a model with one command and have it &amp;ldquo;just work&amp;rdquo; was great.&lt;/p&gt;&#10;&lt;p&gt;But as a &lt;a href="https://www.nijho.lt/post/proxmox-to-nixos/"&gt;NixOS&lt;/a&gt; user, this imperative convenience is actually a flaw.&#10;I want my system to be declarative.&#10;I want my model configuration committed to git, not hiding in a local database.&#10;So I let convenience win for a while, but it always felt wrong.&lt;/p&gt;&#10;&lt;p&gt;But last month, that changed.&#10;I upgraded my hobbyist AI rig by adding a &lt;strong&gt;second RTX 3090&lt;/strong&gt;.&#10;Suddenly, I had 48GB of VRAM and a whole new set of problems—and possibilities.&#10;I wanted to run massive models (like &lt;code&gt;gpt-oss-120b&lt;/code&gt; or &lt;code&gt;Qwen-3-VL-32B&lt;/code&gt; with a huge context window).&#10;I needed to balance layers between two GPUs.&#10;I needed to spill specific parts of the model into system RAM without crashing.&lt;/p&gt;&#10;&lt;p&gt;When you start pushing hardware to its limits, &amp;ldquo;magic&amp;rdquo; abstractions like Ollama stop being helpful and start getting in the way.&#10;I realized that &lt;code&gt;llama.cpp&lt;/code&gt; isn&amp;rsquo;t just a backend; it is &lt;em&gt;the&lt;/em&gt; place where development happens.&#10;If you want to use split-mode inference, granular CPU offloading, or the latest quantization tricks immediately, there is no alternative.&lt;/p&gt;&#10;&lt;p&gt;So, I dusted off this draft, updated my Nix configuration, and I&amp;rsquo;ve been running &lt;code&gt;llama-swap&lt;/code&gt; full-time for a month.&#10;I&amp;rsquo;m finally ready to vouch for it.&lt;/p&gt;&#10;&lt;h2 id="the-original-trigger-broken-models"&gt;The original trigger: broken models &lt;a class="anchor" href="#the-original-trigger-broken-models" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Four months ago, my frustration stemmed from compatibility.&#10;I was trying to run &lt;code&gt;gpt-oss-20b&lt;/code&gt;, a model that required specific chat templates and GGUF handling that &lt;a href="https://ollama.com/" target="_blank" rel="noopener"&gt;Ollama&lt;/a&gt; simply broke.&lt;/p&gt;&#10;&lt;p&gt;The maintainer of &lt;a href="https://github.com/ggml-org/llama.cpp" target="_blank" rel="noopener"&gt;&lt;code&gt;llama.cpp&lt;/code&gt;&lt;/a&gt;, Georgi Gerganov, &lt;a href="https://github.com/ollama/ollama/issues/11714#issuecomment-3172893576" target="_blank" rel="noopener"&gt;explained it perfectly&lt;/a&gt; regarding the gpt-oss debacle:&#10;Ollama forked the ggml inference engine to rush out &amp;ldquo;day-1 support&amp;rdquo; for gpt-oss without coordinating with upstream.&#10;The result? Their implementation was not only incompatible with standard GGUF files but also significantly slower and unoptimized.&lt;/p&gt;&#10;&lt;p&gt;Benchmarks at the time showed Ollama had significant performance overhead—often 20-30% slower inference speeds compared to running &lt;code&gt;llama.cpp&lt;/code&gt; directly.&lt;/p&gt;&#10;&lt;p&gt;This pattern means Ollama will always be behind &lt;code&gt;llama.cpp&lt;/code&gt;.&#10;Meanwhile, I&amp;rsquo;ve had &lt;a href="https://github.com/ollama/ollama/pull/11249" target="_blank" rel="noopener"&gt;a pull request&lt;/a&gt; sitting unreviewed for over a month that fixes their broken OpenAI API.&#10;Compare that to &lt;code&gt;llama.cpp&lt;/code&gt;: I submitted &lt;a href="https://github.com/ggml-org/llama.cpp/pull/15295" target="_blank" rel="noopener"&gt;a PR there&lt;/a&gt; and it was merged in less than an hour.&#10;If you want the latest features &lt;em&gt;now&lt;/em&gt; (and a responsive community), you go to the source.&lt;/p&gt;&#10;&lt;h2 id="the-catalyst-dual-gpus-and-control"&gt;The catalyst: dual GPUs and control &lt;a class="anchor" href="#the-catalyst-dual-gpus-and-control" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;With a single GPU, you usually either fit the model or you don&amp;rsquo;t.&#10;With two GPUs (and 64GB of system RAM), things get interesting.&#10;You enter the territory of &lt;strong&gt;heterogeneous inference&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Ollama tries to handle splitting automatically, but I found myself fighting it.&#10;I wanted to say: &lt;em&gt;&amp;ldquo;Put 40 layers on GPU 0, 40 layers on GPU 1, and keep the KV cache here.&amp;rdquo;&lt;/em&gt;&#10;Or: &lt;em&gt;&amp;ldquo;This model is slightly too big for VRAM; I want to offload exactly the last 10% to system RAM.&amp;rdquo;&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;With &lt;code&gt;llama.cpp&lt;/code&gt; directly, this is just a command line flag.&#10;With &lt;code&gt;llama-swap&lt;/code&gt;, I can define this behavior per-model in a config file:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-yaml" data-lang="yaml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&amp;#34;gpt-oss:120b&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;cmd&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt;&lt;span class="sd"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; ${pkgs.llama-cpp}/bin/llama-server&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; -hf ggml-org/gpt-oss-120b-GGUF&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --port ${PORT}&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --ctx-size 65536&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --split-mode layer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --tensor-split 3,1.3 # Tuned balance between my GPUs&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --n-cpu-moe 15 # Offload 15 experts to CPU RAM&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --threads 8&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="sd"&gt; --chat-template-kwargs &amp;#39;{&amp;#34;reasoning_effort&amp;#34;: &amp;#34;high&amp;#34;}&amp;#39;&lt;/span&gt;&lt;span class="w"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This level of granularity turned my rig from a &amp;ldquo;black box&amp;rdquo; into a tunable instrument.&#10;The &lt;code&gt;--tensor-split&lt;/code&gt; and &lt;code&gt;--n-cpu-moe&lt;/code&gt; flags allow me to run a 120B parameter model that technically shouldn&amp;rsquo;t fit, by carefully balancing the load across both GPUs and system RAM.&lt;/p&gt;&#10;&lt;h2 id="api-compatibility-and-the-switch"&gt;API compatibility and the &amp;ldquo;switch&amp;rdquo; &lt;a class="anchor" href="#api-compatibility-and-the-switch" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;One of the reasons I hesitated to switch was the fear of breaking my existing workflows.&#10;I use &lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;agent-cli&lt;/a&gt; and other custom Python scripts daily.&#10;But it turned out to be a non-issue.&#10;Both Ollama and &lt;code&gt;llama-swap&lt;/code&gt; (wrapping &lt;code&gt;llama-server&lt;/code&gt;) expose an OpenAI-compatible API.&lt;/p&gt;&#10;&lt;p&gt;Once I adopted the pattern of setting a custom &lt;code&gt;OPENAI_BASE_URL&lt;/code&gt; in my projects, the backend became irrelevant.&#10;My tools don&amp;rsquo;t care if they are talking to &lt;code&gt;api.openai.com&lt;/code&gt;, a local Ollama instance, or &lt;code&gt;llama-swap&lt;/code&gt;.&#10;They just send JSON and get JSON back.&#10;This standardization meant I could swap the entire inference engine underneath my applications without changing a single line of their code.&lt;/p&gt;&#10;&lt;h2 id="my-nixos-setup"&gt;My NixOS setup &lt;a class="anchor" href="#my-nixos-setup" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Since I manage my entire system with Nix, setting up &lt;code&gt;llama-swap&lt;/code&gt; and pinning &lt;code&gt;llama-cpp&lt;/code&gt; to the absolute bleeding edge was surprisingly elegant.&#10;I even wrote an &lt;a href="https://github.com/basnijholt/dotfiles/blob/51c7af46e62a7d13b4ff497380c8b58c05ed81c8/configs/nixos/scripts/update_overrides.py" target="_blank" rel="noopener"&gt;auto-update script&lt;/a&gt; to ensure I&amp;rsquo;m always on the latest commit.&lt;/p&gt;&#10;&lt;p&gt;This is where my &lt;a href="https://www.nijho.lt/post/nixos-cache/"&gt;NixOS build cache&lt;/a&gt; shines.&#10;Unlike Ollama, which releases every few weeks, &lt;code&gt;llama.cpp&lt;/code&gt; can have multiple releases &lt;em&gt;a day&lt;/em&gt;.&#10;My cache server runs this update script automatically, builds the new binaries, and serves them to my workstation.&#10;I get the absolute latest performance improvements without ever compiling source code locally.&lt;/p&gt;&#10;&lt;h3 id="compiler-flags-matter"&gt;Compiler flags matter &lt;a class="anchor" href="#compiler-flags-matter" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;One critical lesson I learned: default builds matter.&#10;Initially, I was getting a measly &lt;strong&gt;8 tokens/sec&lt;/strong&gt; on &lt;code&gt;gpt-oss:120b&lt;/code&gt;.&#10;It turned out the default &lt;code&gt;llama.cpp&lt;/code&gt; package wasn&amp;rsquo;t enabling BLAS or native CPU optimizations, crippling the layers I offloaded to system RAM.&lt;/p&gt;&#10;&lt;p&gt;By overriding the package to enable &lt;code&gt;blasSupport&lt;/code&gt; and passing &lt;code&gt;-DGGML_NATIVE=ON&lt;/code&gt; (as seen below), performance jumped to &lt;strong&gt;50 tokens/sec&lt;/strong&gt;.&#10;This is another reason why I prefer managing this via Nix: I can enforce these compile-time flags declaratively.&lt;/p&gt;&#10;&lt;p&gt;Here&amp;rsquo;s my complete configuration (see &lt;a href="https://github.com/basnijholt/dotfiles/blob/51c7af46e62a7d13b4ff497380c8b58c05ed81c8/configs/nixos/hosts/pc/package-overrides.nix#L28-L76" target="_blank" rel="noopener"&gt;&lt;code&gt;package-overrides.nix&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://github.com/basnijholt/dotfiles/blob/51c7af46e62a7d13b4ff497380c8b58c05ed81c8/configs/nixos/hosts/pc/ai.nix#L360-L380" target="_blank" rel="noopener"&gt;&lt;code&gt;ai.nix&lt;/code&gt;&lt;/a&gt;):&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-nix" data-lang="nix"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Override llama-cpp to latest version with CUDA support&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;llama-cpp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pkgs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;llama-cpp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;override&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;cudaSupport&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;rocmSupport&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;metalSupport&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Enable BLAS for optimized CPU layer performance (OpenBLAS)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;blasSupport&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;})&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;overrideAttrs&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;oldAttrs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;rec&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;7205&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;src&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pkgs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fetchFromGitHub&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;owner&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;ggml-org&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;repo&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;llama.cpp&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;tag&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;b&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;sha256-1CcYbc8RWAPVz8hoxKEmbAgQesC1oGFZ3fhfuU5vmOc=&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;leaveDotGit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;postFetch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;&amp;#39;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; git -C &amp;#34;$out&amp;#34; rev-parse --short HEAD &amp;gt; $out/COMMIT&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; find &amp;#34;$out&amp;#34; -name .git -print0 | xargs -0 rm -rf&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#39;&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;};&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Enable native CPU optimizations (AVX, AVX2, etc.)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;cmakeFlags&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;oldAttrs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cmakeFlags&lt;/span&gt; &lt;span class="n"&gt;or&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt; &lt;span class="o"&gt;++&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;-DGGML_NATIVE=ON&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;];&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Disable Nix&amp;#39;s march=native stripping&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;preConfigure&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;&amp;#39;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; export NIX_ENFORCE_NO_NATIVE=0&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="n"&gt;oldAttrs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;preConfigure&lt;/span&gt; &lt;span class="n"&gt;or&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s1"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; &amp;#39;&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;});&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# llama-swap from GitHub releases&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;llama-swap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pkgs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;runCommand&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;llama-swap&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;&amp;#39;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; mkdir -p $out/bin&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; tar -xzf &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;pkgs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fetchurl&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;https://github.com/mostlygeek/llama-swap/releases/download/v175/llama-swap_175_linux_amd64.tar.gz&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;hash&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;sha256-zeyVz0ldMxV4HKK+u5TtAozfRI6IJmeBo92IJTgkGrQ=&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s1"&gt; -C $out/bin&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; chmod +x $out/bin/llama-swap&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;&amp;#39;&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Configure llama-swap as a systemd service&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;systemd&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;services&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;llama-swap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;description&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;llama-swap - OpenAI compatible proxy with automatic model swapping&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;after&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;network.target&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;];&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;wantedBy&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;multi-user.target&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;];&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;serviceConfig&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;Type&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;simple&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;User&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;basnijholt&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;Group&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;users&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Point to your declarative config file&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;ExecStart&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="n"&gt;pkgs&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;llama-swap&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/bin/llama-swap --config /etc/llama-swap/config.yaml --listen 0.0.0.0:9292 --watch-config&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;Restart&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;always&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;RestartSec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="c1"&gt;# Environment for CUDA support&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;Environment&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;PATH=/run/current-system/sw/bin&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;LD_LIBRARY_PATH=/run/opengl-driver/lib:/run/opengl-driver-32/lib&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;];&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;};&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;};&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id="what-i-gained-and-lost"&gt;What I gained (and lost) &lt;a class="anchor" href="#what-i-gained-and-lost" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;h3 id="advantages-of-llama-swap--llamacpp"&gt;Advantages of llama-swap + llama.cpp &lt;a class="anchor" href="#advantages-of-llama-swap--llamacpp" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Heterogeneous compute&lt;/strong&gt;: Perfectly balance loads between GPUs and CPU RAM.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Bleeding edge&lt;/strong&gt;: I use features (like new quantization types) the day they land in &lt;code&gt;llama.cpp&lt;/code&gt;.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Zero idle VRAM&lt;/strong&gt;: Models unload completely when not in use.&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Transparent operation&lt;/strong&gt;: No hidden prompts or &amp;ldquo;helpful&amp;rdquo; formatting that breaks complex tasks.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="the-trade-off"&gt;The trade-off &lt;a class="anchor" href="#the-trade-off" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;The main trade-off is &amp;ldquo;laziness.&amp;rdquo;&#10;You can&amp;rsquo;t just &lt;code&gt;ollama pull new-model&lt;/code&gt; and have it appear.&#10;You have to find the GGUF on HuggingFace, decide on the quantization (Q4_K_M? Q6_K?), and add 5 lines of YAML to your config.&lt;/p&gt;&#10;&lt;p&gt;But honestly? That&amp;rsquo;s a feature, not a bug.&#10;It forces you to understand &lt;em&gt;what&lt;/em&gt; you are running.&#10;It stops you from running a quantized model that is too stupid for your task just because it was the default.&#10;And for a NixOS user, putting that configuration into code instead of a hidden database is exactly how it should be.&lt;/p&gt;&#10;&lt;h2 id="conclusion-maturing-as-an-ai-hobbyist"&gt;Conclusion: maturing as an AI hobbyist &lt;a class="anchor" href="#conclusion-maturing-as-an-ai-hobbyist" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My switch back to &lt;code&gt;llama-swap&lt;/code&gt; wasn&amp;rsquo;t about fixing a bug; it was about outgrowing a tool.&#10;Ollama is the &amp;ldquo;starter bike&amp;rdquo; of local AI—fantastic for getting rolling.&#10;But when you add a second GPU, start caring about tokens per second, and want to maximize every gigabyte of your 48GB VRAM buffer, you need a manual transmission.&lt;/p&gt;&#10;&lt;p&gt;For anyone else building a hobbyist multi-GPU rig: stop fighting the abstraction.&#10;Go to the source. The control is worth it.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;p&gt;&lt;em&gt;Check out my &lt;a href="https://github.com/basnijholt/dotfiles/tree/51c7af46e62a7d13b4ff497380c8b58c05ed81c8/configs/nixos" target="_blank" rel="noopener"&gt;NixOS configuration&lt;/a&gt; for the complete setup!&lt;/em&gt;&lt;/p&gt;&#10;</description></item><item><title>Two years on the Glove80: a split ergonomic, programmable keyboard as a programmer ⌨️</title><link>https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/</link><pubDate>Thu, 06 Nov 2025 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/glove80-split-ergonomic-keyboard/</guid><description>&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#why-i-switched"&gt;Why I switched&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#split-ergonomic-keyboards-what-makes-them-different"&gt;Split ergonomic keyboards: What makes them different?&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#columnar-layout"&gt;Columnar layout&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#vertical-columnstagger"&gt;Vertical column‑stagger&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#split-halves"&gt;Split halves&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#key-wells"&gt;Key wells&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#thumb-clusters"&gt;Thumb clusters&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#tenting"&gt;Tenting&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#switches"&gt;Switches&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-first-month-comfort-vs-disruption"&gt;The first month: Comfort vs. disruption&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#what-was-frustrating"&gt;What was frustrating&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#practice-tools-that-helped"&gt;Practice tools that helped&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#programmable-features-that-make-a-difference"&gt;Programmable features that make a difference&lt;/a&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#mod-tap-one-key-two-functions"&gt;Mod-tap: One key, two functions&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#sticky-modifiers-tap-instead-of-hold"&gt;Sticky modifiers: Tap instead of hold&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#layers-multiple-keyboards-in-one"&gt;Layers: Multiple keyboards in one&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#the-hyper-key-shortcut"&gt;The &amp;ldquo;hyper&amp;rdquo; key shortcut&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10; &lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#what-id-tell-my-past-self"&gt;What I&amp;rsquo;d tell my past self&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#where-i-am-now"&gt;Where I am now&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#whats-next"&gt;What&amp;rsquo;s next?&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#community-and-support"&gt;Community and support&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;aside class="callout callout-note"&gt;&lt;svg class="ico" viewBox="0 0 256 256" aria-hidden="true"&gt;&lt;path d="M128,24A104,104,0,1,0,232,128,104.11,104.11,0,0,0,128,24Zm0,192a88,88,0,1,1,88-88A88.1,88.1,0,0,1,128,216Zm16-40a8,8,0,0,1-8,8,16,16,0,0,1-16-16V128a8,8,0,0,1,0-16,16,16,0,0,1,16,16v40A8,8,0,0,1,144,176ZM112,84a12,12,0,1,1,12,12A12,12,0,0,1,112,84Z"/&gt;&lt;/svg&gt;&lt;div&gt;tl;dr: Switched from a Logitech Ergo K860 to a split, columnar Glove80 (Jan 2024).&#10;Month 1 hurt; ~85 wpm by weeks 4–8 and ~90 now—less strain and programmable wins (mod‑taps, sticky modifiers, layers).&#10;(Measured Nov 2025.)&lt;/div&gt;&lt;/aside&gt;&#10;&lt;h2 id="why-i-switched"&gt;Why I switched &lt;a class="anchor" href="#why-i-switched" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;For roughly three years I used the Logitech Ergo K860.&#10;Despite liking it, I kept getting RSI (Repetitive Strain Injury) flare‑ups—pain and discomfort from repetitive typing motions.&#10;~2 years ago I discovered that split ergonomic keyboards were a thing and went down the rabbit hole.&#10;I switched to the Glove80 in January 2024.&#10;Before I bought anything, I started watching videos from the channel &lt;a href="https://www.youtube.com/@ifcodingwerenatural" target="_blank" rel="noopener"&gt;If Coding Were Natural&lt;/a&gt; and from &lt;a href="https://www.youtube.com/@BenVallack" target="_blank" rel="noopener"&gt;Ben Vallack&lt;/a&gt;.&#10;I read many reviews and first‑hand experiences.&#10;The more I watched and read, the more my interest in split ergonomic keyboards grew.&#10;Overwhelmingly, people seemed to praise the &lt;a href="https://www.moergo.com/" target="_blank" rel="noopener"&gt;Glove80&lt;/a&gt; for its comfort, so I decided to try it.&#10;I considered many alternatives like the Moonlander and Ergodox but chose the Glove80 for its lighter switches and curved key wells.&lt;/p&gt;&#10;&lt;p&gt;Before the Glove80, I had flirted with mechanical keyboards because the community is huge and very into them.&#10;Given the nerd that I am, I got nerd‑sniped by the mechanical‑keyboard scene.&#10;I assumed I would be into mechanical keyboards.&#10;In April 2019, I ordered a 60% board, the Anne Pro 2.&#10;After trying it, I returned it because it wasn’t an improvement for me and the reduced ergonomics weren’t worth it compared to the Logitech.&#10;The mechanical switches didn’t really excite me in practice as much as I had expected.&#10;What I actually needed was a split, columnar layout with thumb clusters and programmability.&#10;I didn&amp;rsquo;t know this yet.&lt;/p&gt;&#10;&lt;h2 id="split-ergonomic-keyboards-what-makes-them-different"&gt;Split ergonomic keyboards: What makes them different? &lt;a class="anchor" href="#split-ergonomic-keyboards-what-makes-them-different" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Before diving into my experience, here&amp;rsquo;s what sets these keyboards apart from traditional ones.&lt;/p&gt;&#10;&lt;h3 id="columnar-layout"&gt;Columnar layout &lt;a class="anchor" href="#columnar-layout" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Columns line up under each finger, so most movements are up/down instead of sideways.&#10;Compared to row‑staggered boards, this reduces lateral finger travel and awkward cross‑overs.&#10;Row‑staggered boards offset keys sideways, which encourages lateral travel rather than straight motions.&lt;/p&gt;&#10;&lt;h3 id="vertical-columnstagger"&gt;Vertical column‑stagger &lt;a class="anchor" href="#vertical-columnstagger" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Each column is offset to match finger length, which reduces awkward reaches for ring and pinky fingers and keeps wrists more neutral.&#10;The middle finger column sits higher, the pinky column lower—matching the natural length differences between fingers.&lt;/p&gt;&#10;&lt;h3 id="split-halves"&gt;Split halves &lt;a class="anchor" href="#split-halves" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Two halves let shoulders and wrists stay in a neutral, shoulder‑width posture.&#10;One‑piece boards tend to pull shoulders inward and bend wrists sideways (called ulnar deviation).&#10;With the split, I can position each half independently for my body.&lt;/p&gt;&#10;&lt;h3 id="key-wells"&gt;Key wells &lt;a class="anchor" href="#key-wells" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Curved wells bring keycaps closer to your finger pads along natural arcs, so you extend and lift your fingers less.&#10;Flat boards make you lift and extend fingers more and move farther between rows.&#10;The Glove80&amp;rsquo;s wells are subtle but noticeable after hours of typing.&lt;/p&gt;&#10;&lt;h3 id="thumb-clusters"&gt;Thumb clusters &lt;a class="anchor" href="#thumb-clusters" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Modifiers, Backspace, Enter, and other frequent keys can live under your strongest digits—the thumbs—instead of overloading pinkies.&#10;Big space bars under‑utilize thumbs and push more work to pinkies and corners.&#10;On the Glove80, my thumbs handle multiple keys that would normally require pinky reaches.&lt;/p&gt;&#10;&lt;h3 id="tenting"&gt;Tenting &lt;a class="anchor" href="#tenting" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Tenting means raising one side of the keyboard to adjust the angle.&#10;I tent the halves to a mild angle so my forearms don&amp;rsquo;t have to twist as much (less pronation) and my wrists stay more neutral.&#10;Most standard keyboards cannot tent at all, which forces your palms completely flat and pulls your shoulders inward.&lt;/p&gt;&#10;&lt;h3 id="switches"&gt;Switches &lt;a class="anchor" href="#switches" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;I use Kailh Choc Red Pro switches on my Glove80.&#10;They are linear with a low actuation force, which feels light and comfortable for long typing.&#10;I don&amp;rsquo;t have much experience with other switches, so this is not a comparison—just what works for me.&lt;/p&gt;&#10;&lt;figure id="figure-left-half-of-my-glove80-showing-the-columnar-columns-thumb-cluster-and-mild-tenting"&gt;&lt;img src="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/glove80-left_hu_bc6b49e7cec10dca.webp" srcset="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/glove80-left_hu_837731d0947f519f.webp 267w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/glove80-left_hu_bc6b49e7cec10dca.webp 507w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/glove80-left_hu_bd4285bb8af2fc40.webp 800w" sizes="auto, (max-width: 760px) 100vw, 720px" width="507" height="760" alt="Left half of my Glove80 on the desk" loading="lazy" data-zoomable&gt;&lt;figcaption&gt;Left half of my Glove80 showing the columnar columns, thumb cluster, and mild tenting.&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;p&gt;Taken together, these differences mean less strain and more consistent typing for me than on row‑staggered boards.&lt;/p&gt;&#10;&lt;h2 id="the-first-month-comfort-vs-disruption"&gt;The first month: Comfort vs. disruption &lt;a class="anchor" href="#the-first-month-comfort-vs-disruption" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;After spending a month on the Glove80, I found it much more comfortable than expected.&#10;At the same time, it completely disrupted my typing strategy—what I had learned on a traditional layout didn’t transfer to the columnar layout of the Glove80.&lt;/p&gt;&#10;&lt;h3 id="what-was-frustrating"&gt;What was frustrating &lt;a class="anchor" href="#what-was-frustrating" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;In the first couple of days, my typing speed tanked and it was extremely frustrating.&#10;On the first day, I was at 22 wpm—slower than hunt-and-peck typing.&#10;I was constantly looking at the keyboard, trying to use the right fingers for the right keys.&#10;On a traditional keyboard I had my own strategy—basically three fingers per hand—and I was actually very fast with it.&#10;But that didn&amp;rsquo;t transfer at all to the columnar layout; using the &amp;ldquo;right&amp;rdquo; finger for each column felt natural once learned, but it wasn&amp;rsquo;t native to me at first.&#10;I did not want to give up and kept going despite the massive frustration.&#10;Maybe what kept me going was the sunk‑cost fallacy of $400 (which is mid-range for split ergonomic keyboards, they typically run 200-600 USD).&lt;/p&gt;&#10;&lt;h2 id="practice-tools-that-helped"&gt;Practice tools that helped &lt;a class="anchor" href="#practice-tools-that-helped" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I started with &lt;a href="https://www.keybr.com/" target="_blank" rel="noopener"&gt;Keybr&lt;/a&gt; to learn key‑by‑key.&#10;Once I finished that, I switched to using &lt;a href="https://monkeytype.com/" target="_blank" rel="noopener"&gt;MonkeyType&lt;/a&gt; exclusively.&#10;After 1–2 months I had already reached ~85 wpm (words per minute—average typing speed is 40 wpm, programmers often type 60-80 wpm).&#10;I haven&amp;rsquo;t improved a whole lot since then, but I also haven&amp;rsquo;t done dedicated typing practice.&#10;While writing this post I tried MonkeyType again and I&amp;rsquo;m now ~90 wpm.&lt;/p&gt;&#10;&lt;figure id="figure-my-keybr-progress-during-the-first-4-weeks-starting-at-20-wpm-both-speed-and-accuracy-steadily-improved-over-each-lesson"&gt;&lt;img src="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/keybr-progress_hu_ce6b7ffe359f8bb0.webp" srcset="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/keybr-progress_hu_7d861ceda887b016.webp 400w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/keybr-progress_hu_ce6b7ffe359f8bb0.webp 760w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/keybr-progress_hu_6c3cec7aa5118231.webp 1200w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="415" alt="Keybr progress graph showing typing speed and accuracy improvement" loading="lazy" data-zoomable&gt;&lt;figcaption&gt;My Keybr progress during the first 4 weeks: starting at ~20 wpm, both speed and accuracy steadily improved over each lesson.&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;p&gt;Beyond the tools, a few things helped with the transition:&#10;the Glove80&amp;rsquo;s columnar layout actually forces you to use the correct finger for each column—using the wrong finger feels completely unnatural and awkward, so I had to abandon my old three-finger strategy.&#10;I also had to accept the short-term speed dip while accuracy improved, and adding slight tenting helped keep my wrists more neutral.&lt;/p&gt;&#10;&lt;p&gt;Today, it feels completely natural for me to type on this keyboard.&lt;/p&gt;&#10;&lt;h2 id="programmable-features-that-make-a-difference"&gt;Programmable features that make a difference &lt;a class="anchor" href="#programmable-features-that-make-a-difference" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;These keyboards let you program behaviors that eliminate awkward reaches and repetitive strain.&#10;Here&amp;rsquo;s what I actually use:&lt;/p&gt;&#10;&lt;h3 id="mod-tap-one-key-two-functions"&gt;Mod-tap: One key, two functions &lt;a class="anchor" href="#mod-tap-one-key-two-functions" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Keys do different things based on tap vs. hold.&#10;All my symbols work this way:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Tap &lt;code&gt;1&lt;/code&gt; = &amp;ldquo;1&amp;rdquo;, hold for ~200ms = &amp;ldquo;!&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;Tap &lt;code&gt;,&lt;/code&gt; = &amp;ldquo;,&amp;rdquo;, hold = &amp;ldquo;&amp;lt;&amp;rdquo;&lt;/li&gt;&#10;&lt;li&gt;Tap &lt;code&gt;.&lt;/code&gt; = &amp;ldquo;.&amp;rdquo;, hold = &amp;ldquo;&amp;gt;&amp;rdquo;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This keeps everything on the home row without reaching to a symbols row.&lt;/p&gt;&#10;&lt;h3 id="sticky-modifiers-tap-instead-of-hold"&gt;Sticky modifiers: Tap instead of hold &lt;a class="anchor" href="#sticky-modifiers-tap-instead-of-hold" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;All my modifiers are &amp;ldquo;sticky&amp;rdquo;—tap them once and they latch for the next key:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Tap &lt;code&gt;Command&lt;/code&gt; + &lt;code&gt;Shift&lt;/code&gt;, then tap &lt;code&gt;T&lt;/code&gt; = reopens last tab&lt;/li&gt;&#10;&lt;li&gt;No more finger gymnastics holding three keys at once&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;With a single thumb movement on my left hand, I can reopen a tab quickly.&lt;/p&gt;&#10;&lt;h3 id="layers-multiple-keyboards-in-one"&gt;Layers: Multiple keyboards in one &lt;a class="anchor" href="#layers-multiple-keyboards-in-one" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;Multiple virtual keymaps on one keyboard.&#10;You can momentarily switch or toggle to a layer to get different keys under the same physical switches.&#10;Common patterns include a symbols layer, a navigation layer, and a numbers/function layer—all reachable from home row.&lt;/p&gt;&#10;&lt;figure id="figure-one-of-my-layers-digits-act-as-modtaps-for-symbols-and-common-modifiers-live-under-the-thumbs"&gt;&lt;img src="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/layout-layer-1_hu_d7cc455a5cbab8a5.webp" srcset="https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/layout-layer-1_hu_d31d10a3d9920b1f.webp 400w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/layout-layer-1_hu_d7cc455a5cbab8a5.webp 760w, https://www.nijho.lt/post/glove80-split-ergonomic-keyboard/layout-layer-1_hu_97691c8ef86e3a18.webp 1200w" sizes="auto, (max-width: 760px) 100vw, 720px" width="760" height="381" alt="Glove80 layer: symbols and numbers via mod‑taps with sticky modifiers under the thumbs" loading="lazy" data-zoomable&gt;&lt;figcaption&gt;One of my layers: digits act as mod‑taps for symbols, and common modifiers live under the thumbs.&lt;/figcaption&gt;&lt;/figure&gt;&#10;&lt;h3 id="the-hyper-key-shortcut"&gt;The &amp;ldquo;hyper&amp;rdquo; key shortcut &lt;a class="anchor" href="#the-hyper-key-shortcut" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;p&gt;If I hold the &lt;code&gt;Home&lt;/code&gt; key, it acts as if I&amp;rsquo;m pressing &lt;code&gt;Ctrl&lt;/code&gt; + &lt;code&gt;Alt&lt;/code&gt; + &lt;code&gt;Shift&lt;/code&gt; + &lt;code&gt;Command&lt;/code&gt; simultaneously (the &amp;ldquo;hyper&amp;rdquo; key I mentioned earlier).&#10;Since few apps use that combo by default, I map it to custom commands:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;Home&lt;/code&gt; + &lt;code&gt;T&lt;/code&gt; → open Terminal&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;Home&lt;/code&gt; + &lt;code&gt;B&lt;/code&gt; → open Browser&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;Home&lt;/code&gt; + &lt;code&gt;C&lt;/code&gt; → open Calendar&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;Home&lt;/code&gt; + &lt;code&gt;S&lt;/code&gt; → open Slack&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This hyper key setup enabled my keyboard-first workflow across macOS (using &lt;a href="https://github.com/basnijholt/dotfiles/tree/main/configs/keyboard-maestro" target="_blank" rel="noopener"&gt;Keyboard Maestro&lt;/a&gt;) and Linux (using &lt;a href="https://github.com/basnijholt/dotfiles/tree/main/configs/hypr" target="_blank" rel="noopener"&gt;Hyprland&lt;/a&gt;).&#10;The same muscle memory works across both systems, relieving me from tapping Command+Tab like a mad person.&lt;/p&gt;&#10;&lt;details&gt;&#10; &lt;summary&gt;More programmable tricks I use (details)&lt;/summary&gt;&#10; &lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Double‑tap &lt;code&gt;Shift&lt;/code&gt; → emulate mouse behavior on keyboard&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;A quick double‑tap of &lt;code&gt;Shift&lt;/code&gt; toggles my mouse behavior layer so I can move/scroll/click without leaving the home row.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Double‑tap &lt;code&gt;Command&lt;/code&gt; → macOS‑style word navigation on Linux&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;I use this to mirror macOS word‑jump behavior in Linux editors/terminals, so &lt;code&gt;⌘&lt;/code&gt; + arrows behaves the same everywhere.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Combos for common keys&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;E&lt;/code&gt; + &lt;code&gt;D&lt;/code&gt; → &lt;code&gt;Enter&lt;/code&gt;; &lt;code&gt;R&lt;/code&gt; + &lt;code&gt;F&lt;/code&gt; → &lt;code&gt;Space&lt;/code&gt;.&#10;This lets me press Enter and Space with my left hand while my right hand is on the mouse, since those keys sit on the right side on many layouts.&#10;I also keep most app‑switch launchers on the left side so my left hand can trigger them while my right hand stays on the mouse.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Shift + Backspace → Delete&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Backspace morphs into Delete when Shift is held—handy on macOS/Linux.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Tap dance to momentary vs. toggle a layer&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Single tap for a momentary “Lower,” double tap to lock it on; keeps symbols/nav nearby without mode churn.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Mouse keys with speed profiles&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Separate slow/fast/warp cursor and scroll speeds for precision vs. movement.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Delete to start/end of line&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;One chord deletes left to line start; another deletes right to line end—great for editors.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;Tuning for fewer misfires&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Tap‑preferred hold‑taps and a ~200–220 ms tapping term match my timing.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&#10;&lt;/details&gt;&#10;&#10;&lt;p&gt;None of this requires you to become a firmware expert on day one.&#10;Even a couple of well‑chosen mod‑taps, one sticky key, and a single &amp;ldquo;symbols + navigation&amp;rdquo; layer can remove a lot of awkward reaches.&lt;/p&gt;&#10;&lt;h2 id="what-id-tell-my-past-self"&gt;What I&amp;rsquo;d tell my past self &lt;a class="anchor" href="#what-id-tell-my-past-self" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Expect an initial drop in speed; it&amp;rsquo;s normal.&lt;/li&gt;&#10;&lt;li&gt;Columnar layouts reward correct finger usage; looking down at first is fine.&lt;/li&gt;&#10;&lt;li&gt;Start simple: one or two mod‑taps, one sticky modifier, one extra layer.&lt;/li&gt;&#10;&lt;li&gt;Practice compounds. MonkeyType and Keybr were enough for me.&lt;/li&gt;&#10;&lt;li&gt;Don&amp;rsquo;t let sunk‑cost thinking scare you into quitting, but also don&amp;rsquo;t over‑engineer your layout on day one.&lt;/li&gt;&#10;&lt;li&gt;You won&amp;rsquo;t lose the ability to type on normal keyboards. When I type on my laptop&amp;rsquo;s built‑in keyboard, I just use my original three-finger strategy and have no issues. It&amp;rsquo;s like developing two different muscle memories based on the keyboard.&lt;/li&gt;&#10;&lt;li&gt;Gaming consideration: WASD sits off the natural columns on a columnar split. I briefly tried a shooter and it felt awkward—unless you&amp;rsquo;re willing to reprogram and relearn gaming muscle memory, FPS games don&amp;rsquo;t work well. I only played Factorio on it, which worked fine.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="where-i-am-now"&gt;Where I am now &lt;a class="anchor" href="#where-i-am-now" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;As of November 2025, I’m ~90 wpm, and it feels completely natural to type on the Glove80.&#10;I reached 85 wpm within 1–2 months and have mostly stayed in that range since, with a small bump recently.&#10;The comfort gains are real, and the programmability (mod‑tap, sticky keys, and layers) turned out to be as exciting in practice as it sounded when I first discovered it.&#10;Was the $400 worth it?&#10;Absolutely!&#10;I probably spent many thousands of hours on it so far, so amortizing it over its lifetime, it&amp;rsquo;s well worth it for me.&#10;Plus it makes me feel like a leet hacker.&lt;/p&gt;&#10;&lt;p&gt;My current Glove80 layout: &lt;a href="https://my.glove80.com/#/layout/user/c5342d66-e6ed-4d04-9ae0-2dfc9cd87930" target="_blank" rel="noopener"&gt;view it here&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h2 id="whats-next"&gt;What&amp;rsquo;s next? &lt;a class="anchor" href="#whats-next" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The Glove80 perfectly meets my needs—I have no practical reason to switch.&#10;But the nerd inside me is curious about other options in the split keyboard ecosystem.&#10;The &lt;a href="https://github.com/foostan/crkbd" target="_blank" rel="noopener"&gt;Corne&lt;/a&gt;, &lt;a href="https://svalboard.com/" target="_blank" rel="noopener"&gt;Svalboard&lt;/a&gt;, and &lt;a href="https://cyboard.digital/" target="_blank" rel="noopener"&gt;Cyboard Imprint&lt;/a&gt; all take quite different approaches and seem interesting to explore.&#10;Whether I&amp;rsquo;ll actually try them or stick with what works is an open question—sometimes the best keyboard is the one you&amp;rsquo;ve already mastered.&lt;/p&gt;&#10;&lt;h2 id="community-and-support"&gt;Community and support &lt;a class="anchor" href="#community-and-support" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;The MoErgo Glove80 community is extremely welcoming.&#10;Stephen, the designer, is very active and responsive—when I requested a small feature for the &lt;a href="https://my.glove80.com/" target="_blank" rel="noopener"&gt;layout editor&lt;/a&gt;, it went live within a day.&lt;/p&gt;&#10;</description></item><item><title>How a 3090 turned into two local‑first AI projects 🎮</title><link>https://www.nijho.lt/post/local-ai-journey/</link><pubDate>Mon, 14 Jul 2025 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/local-ai-journey/</guid><description>&lt;details class="toc" open&gt;&lt;summary&gt;Table of contents&lt;/summary&gt;&lt;nav id="TableOfContents"&gt;&#10; &lt;ul&gt;&#10; &lt;li&gt;&lt;a href="#1-introduction-the-unplayed-games"&gt;1. Introduction: The Unplayed Games&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#2-the-rocky-road-to-a-working-setup"&gt;2. The Rocky Road to a Working Setup&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#3-the-first-victory-agent-cli-"&gt;3. The First Victory: &lt;code&gt;agent-cli&lt;/code&gt; 🐍&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#4-the-grand-ambition-aibrain-"&gt;4. The Grand Ambition: &lt;code&gt;AIBrain&lt;/code&gt; 🧠&lt;/a&gt;&lt;/li&gt;&#10; &lt;li&gt;&lt;a href="#5-conclusion-the-journey-continues"&gt;5. Conclusion: The Journey Continues&lt;/a&gt;&lt;/li&gt;&#10; &lt;/ul&gt;&#10;&lt;/nav&gt;&lt;/details&gt;&#10;&#10;&lt;h2 id="1-introduction-the-unplayed-games"&gt;1. Introduction: The Unplayed Games &lt;a class="anchor" href="#1-introduction-the-unplayed-games" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;I&amp;rsquo;ve wanted to dabble with local AI for a long time, mainly because of the possibilities it unlocks.&#10;Local AI allows me to use all of my data—financial records, personal messages, location history, even transcribing everything I say—without the privacy concerns of sending it to a cloud provider.&#10;This level of deep, personal data analysis is only something I&amp;rsquo;d be comfortable with on my own hardware.&lt;/p&gt;&#10;&lt;p&gt;However, I couldn&amp;rsquo;t justify buying a powerful GPU just for an imagined AI project that I might never really start.&#10;Then it hit me: I could also use it for gaming (duh).&#10;Because I really enjoyed &lt;em&gt;The Last of Us&lt;/em&gt; series, I thought, “&lt;em&gt;Okay, why not play &lt;strong&gt;The Last of Us&lt;/strong&gt; on PC?&lt;/em&gt;”&#10;This was the push I needed.&#10;So, a couple of weeks ago, I convinced myself to buy someone&amp;rsquo;s old gaming machine.&#10;After many deep research sessions with ChatGPT about the best bang-for-the-buck consumer-grade GPU, I landed on an NVIDIA 3090.&#10;I found an (old) beast of a machine, and my initial idea was set.&lt;/p&gt;&#10;&lt;details&gt;&#10; &lt;summary&gt;Click here to see the full system specs ($1350 for all this!)&lt;/summary&gt;&#10; &lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;GPU:&lt;/strong&gt; ASUS ROG STRIX RTX 3090 OC (24GB)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;CPU:&lt;/strong&gt; AMD Ryzen 9 3900X (12-Core)&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Motherboard:&lt;/strong&gt; ASUS ROG Crosshair VIII Hero&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Power Supply:&lt;/strong&gt; Be Quiet! 1200W Dark Power Pro&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;RAM:&lt;/strong&gt; 32GB G.Skill Trident Z Royal DDR4&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Storage:&lt;/strong&gt; 2TB Intel M.2 NVMe SSD&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Case:&lt;/strong&gt; Corsair 4000D Airflow&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;CPU Cooler:&lt;/strong&gt; Thermalright Peerless Assassin 120&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&#10;&lt;/details&gt;&#10;&#10;&lt;p&gt;Fast forward four weeks.&#10;I haven&amp;rsquo;t played a single minute, but my machine is now indexing 200 GB of email archives and facing a serious performance bottleneck I&amp;rsquo;ll get to later.&#10;That &amp;ldquo;gaming&amp;rdquo; machine has been humming away, but not rendering virtual worlds.&#10;It&amp;rsquo;s been training, transcribing, and thinking.&lt;/p&gt;&#10;&lt;p&gt;I&amp;rsquo;m writing this post simply because I am so enthusiastic about local AI.&#10;I&amp;rsquo;ve found that local models are surprisingly powerful for specific use cases.&#10;Of course, they aren&amp;rsquo;t going to replace Gemini or Claude for coding, but for many other applications, they work incredibly well.&#10;Lately, I&amp;rsquo;ve been so obsessed with this that I&amp;rsquo;m sleeping significantly less and spending all my free hours working on it.&lt;/p&gt;&#10;&lt;aside class="callout callout-note"&gt;&lt;svg class="ico" viewBox="0 0 256 256" aria-hidden="true"&gt;&lt;path d="M128,24A104,104,0,1,0,232,128,104.11,104.11,0,0,0,128,24Zm0,192a88,88,0,1,1,88-88A88.1,88.1,0,0,1,128,216Zm16-40a8,8,0,0,1-8,8,16,16,0,0,1-16-16V128a8,8,0,0,1,0-16,16,16,0,0,1,16,16v40A8,8,0,0,1,144,176ZM112,84a12,12,0,1,1,12,12A12,12,0,0,1,112,84Z"/&gt;&lt;/svg&gt;&lt;div&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; This post is the story of how a quest for gaming turned into a deep dive into local, private AI.&#10;I&amp;rsquo;ll share the two open-source Python packages that came out of it, &lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;&lt;code&gt;agent-cli&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://github.com/basnijholt/aibrain" target="_blank" rel="noopener"&gt;&lt;code&gt;AIBrain&lt;/code&gt;&lt;/a&gt;, and the lessons I learned along the way.&lt;/div&gt;&lt;/aside&gt;&#10;&lt;h2 id="2-the-rocky-road-to-a-working-setup"&gt;2. The Rocky Road to a Working Setup &lt;a class="anchor" href="#2-the-rocky-road-to-a-working-setup" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;Before I could even get to the fun AI stuff, I had to wrestle with the machine itself.&#10;I started with Pop!_OS, which is supposed to be a great out of the box solution that &amp;ldquo;just works&amp;rdquo;.&#10;However, it&amp;rsquo;s been a long time since I used Linux on the desktop; my experience is almost exclusively with servers.&#10;I quickly installed some of the wrong NVIDIA drivers and ended up debugging stuff in Grub.&#10;I got frustrated with the different commands I had to run to set up everything and realized that reproducing this system was going to be a massive pain.&#10;Since I&amp;rsquo;ve grown to love Nix on my Mac with Nix-Darwin, I decided to switch the whole system to &lt;strong&gt;NixOS&lt;/strong&gt;.&#10;This has been an amazing decision.&#10;Even though there are many pain points with Nix, I think it is well worth it.&#10;For those interested, you can see the complete setup, including all the local AI services, in my &lt;a href="https://github.com/basnijholt/dotfiles/blob/841cd5fb94b65556e7a2cc7d8ea618f86b50bf1d/configs/nixos/configuration.nix" target="_blank" rel="noopener"&gt;NixOS configuration on GitHub&lt;/a&gt;.&lt;/p&gt;&#10;&lt;h2 id="3-the-first-victory-agent-cli-"&gt;3. The First Victory: &lt;code&gt;agent-cli&lt;/code&gt; 🐍 &lt;a class="anchor" href="#3-the-first-victory-agent-cli-" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;With a stable system, I quickly got started with a few scripts.&#10;These soon evolved into my first real project of this adventure: &lt;a href="https://github.com/basnijholt/agent-cli" target="_blank" rel="noopener"&gt;&lt;code&gt;agent-cli&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;agent-cli&lt;/code&gt; is a collection of &lt;strong&gt;local-first&lt;/strong&gt;, AI-powered command-line agents that run entirely on your machine.&#10;The core philosophy is privacy; your data never leaves your computer.&#10;It&amp;rsquo;s designed for seamless integration with system-wide hotkeys.&#10;I&amp;rsquo;m now using it constantly.&lt;/p&gt;&#10;&lt;p&gt;Instead of typing, I now dictate almost everything.&#10;A quick hotkey, and &lt;code&gt;agent-cli transcribe&lt;/code&gt; turns my speech into text.&#10;Another hotkey, and &lt;code&gt;agent-cli autocorrect&lt;/code&gt; cleans it up using a local LLM running on &lt;strong&gt;Ollama&lt;/strong&gt;.&#10;It has genuinely streamlined my workflow.&#10;For example, I can pipe the tools together, like correcting text from my clipboard and then having it spoken aloud: &lt;code&gt;agent-cli autocorrect | agent-cli speak&lt;/code&gt;.&lt;/p&gt;&#10;&lt;aside class="callout callout-note"&gt;&lt;svg class="ico" viewBox="0 0 256 256" aria-hidden="true"&gt;&lt;path d="M128,24A104,104,0,1,0,232,128,104.11,104.11,0,0,0,128,24Zm0,192a88,88,0,1,1,88-88A88.1,88.1,0,0,1,128,216Zm16-40a8,8,0,0,1-8,8,16,16,0,0,1-16-16V128a8,8,0,0,1,0-16,16,16,0,0,1,16,16v40A8,8,0,0,1,144,176ZM112,84a12,12,0,1,1,12,12A12,12,0,0,1,112,84Z"/&gt;&lt;/svg&gt;&lt;div&gt;&lt;p&gt;&lt;strong&gt;⚡ Surprising Speed: Local vs. Cloud&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;One of my biggest surprises was the raw speed of local models.&#10;While I added an OpenAI provider to &lt;code&gt;agent-cli&lt;/code&gt; for those without local hardware (someone I know wanted to try it), I consistently find that local execution is faster than waiting on a cloud API.&#10;For example, a local transcription with Whisper + Ollama lands in my clipboard in well under a second—often beating the network round-trip to a remote service.&lt;/p&gt;&#10;&lt;/div&gt;&lt;/aside&gt;&#10;&lt;p&gt;The toolkit includes:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;code&gt;autocorrect&lt;/code&gt;: Fixes grammar and spelling.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;transcribe&lt;/code&gt;: Uses a local Whisper model for speech-to-text.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;speak&lt;/code&gt;: Converts text to speech with a local TTS engine.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;voice-edit&lt;/code&gt;: A voice-powered clipboard assistant.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;assistant&lt;/code&gt;: A hands-free voice assistant using a wake word.&lt;/li&gt;&#10;&lt;li&gt;&lt;code&gt;chat&lt;/code&gt;: A conversational AI with tool-calling capabilities.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;It&amp;rsquo;s been a fantastic success in my personal workflow, and it&amp;rsquo;s all open-source for you to try.&#10;Contributions and ideas welcome!&lt;/p&gt;&#10;&lt;p&gt;Once I trusted my toolchain, the obvious next step was to feed &lt;em&gt;all&lt;/em&gt; my data into it.&lt;/p&gt;&#10;&lt;h2 id="4-the-grand-ambition-aibrain-"&gt;4. The Grand Ambition: &lt;code&gt;AIBrain&lt;/code&gt; 🧠 &lt;a class="anchor" href="#4-the-grand-ambition-aibrain-" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;My second package, &lt;a href="https://github.com/basnijholt/aibrain" target="_blank" rel="noopener"&gt;&lt;code&gt;AIBrain&lt;/code&gt;&lt;/a&gt;, is a much larger endeavor.&#10;The vision is to create a &amp;ldquo;life-OS&amp;rdquo;—a private, local AI that processes all my personal data (emails, calendar, files, photos, messages, browser history, location data) to generate summaries and answer questions, eventually allowing for queries like:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;aibrain ask &lt;span class="s2"&gt;&amp;#34;When did I get into landscape photography?&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;aibrain ask &lt;span class="s2"&gt;&amp;#34;How many countries did I visit between 2015-2025?&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This project led me to explore the landscape of agentic frameworks.&#10;I started with &lt;strong&gt;CrewAI&lt;/strong&gt;, but I quickly found that despite the hype and funding, it had some painful problems (specifically &lt;a href="https://github.com/crewAIInc/crewAI/issues/3031" target="_blank" rel="noopener"&gt;this issue&lt;/a&gt; where things just fail without any helpful error messages).&#10;I decided to switch to &lt;strong&gt;LangGraph&lt;/strong&gt;, which felt more robust.&#10;I also chose to not use &lt;strong&gt;PydanticAI&lt;/strong&gt; (which I used in &lt;code&gt;agent-cli&lt;/code&gt;) because its &lt;strong&gt;Ollama&lt;/strong&gt; support is just a wrapper around the OpenAI-compatible API, which doesn&amp;rsquo;t expose all the useful options (I even opened &lt;a href="https://github.com/ollama/ollama/pull/11249/" target="_blank" rel="noopener"&gt;a PR&lt;/a&gt; in Ollama to make its OpenAI compatible API more feature-complete, but it has not been merged yet).&lt;/p&gt;&#10;&lt;p&gt;The core infrastructure of &lt;code&gt;AIBrain&lt;/code&gt; is ready.&#10;It can index emails, calendar events, and files in various formats (PDF, Excel, PowerPoint, images).&#10;However, I&amp;rsquo;ve hit a significant challenge.&lt;/p&gt;&#10;&lt;aside class="callout callout-warning"&gt;&lt;svg class="ico" viewBox="0 0 256 256" aria-hidden="true"&gt;&lt;path d="M236.8,188.09,149.35,36.22h0a24.76,24.76,0,0,0-42.7,0L19.2,188.09a23.51,23.51,0,0,0,0,23.72A24.35,24.35,0,0,0,40.55,224h174.9a24.35,24.35,0,0,0,21.33-12.19A23.51,23.51,0,0,0,236.8,188.09ZM222.93,203.8a8.5,8.5,0,0,1-7.48,4.2H40.55a8.5,8.5,0,0,1-7.48-4.2,7.59,7.59,0,0,1,0-7.72L120.52,44.21a8.75,8.75,0,0,1,15,0l87.45,151.87A7.59,7.59,0,0,1,222.93,203.8ZM120,144V104a8,8,0,0,1,16,0v40a8,8,0,0,1-16,0Zm20,36a12,12,0,1,1-12-12A12,12,0,0,1,140,180Z"/&gt;&lt;/svg&gt;&lt;div&gt;&lt;p&gt;&lt;strong&gt;🔥 The Performance Bottleneck&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Even with a 12-core CPU and 24GB GPU, processing just a few days’ worth of email made the fans scream for half an hour.&#10;Processing hundreds of gigabytes of personal data is a massive undertaking, and it&amp;rsquo;s clear that my current approach isn&amp;rsquo;t scalable enough.&lt;/p&gt;&#10;&lt;/div&gt;&lt;/aside&gt;&#10;&lt;h2 id="5-conclusion-the-journey-continues"&gt;5. Conclusion: The Journey Continues &lt;a class="anchor" href="#5-conclusion-the-journey-continues" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;This past month has been a whirlwind.&#10;I set out to play games and ended up with two open-source AI packages and a much deeper appreciation for the power sitting in a consumer-grade GPU.&#10;It&amp;rsquo;s incredible what you can achieve with local hardware today.&lt;/p&gt;&#10;&lt;p&gt;&lt;code&gt;agent-cli&lt;/code&gt; is a finished, polished tool that I use daily.&#10;&lt;code&gt;AIBrain&lt;/code&gt; is a much bigger dream, and I&amp;rsquo;m still figuring out how to solve the performance puzzle to make it truly practical.&#10;Next up: smarter chunking with LangGraph, vector-DB swap-outs, and GPU scheduling tricks.&lt;/p&gt;&#10;&lt;p&gt;This experience has reinforced my passion for local-first software and open-source development.&#10;The future of personal, private AI is being built right now, not just in large corporate labs, but in the homes of hobbyists tinkering with their gaming rigs.&lt;/p&gt;&#10;&lt;p&gt;⭐ If this resonates, star the repos &amp;amp; drop issues/questions.&lt;/p&gt;&#10;&lt;p&gt;And who knows, maybe I&amp;rsquo;ll even find the time to play &amp;lsquo;The Last of Us&amp;rsquo;.&lt;/p&gt;&#10;</description></item><item><title>My homelab: from a single Raspberry Pi to a Proxmox cluster and a dedicated NAS 🔧</title><link>https://www.nijho.lt/post/homelab/</link><pubDate>Sat, 06 Jul 2024 00:00:00 +0000</pubDate><guid isPermaLink="false">/post/homelab/</guid><description>&lt;p&gt;In 2020, I began my homelab journey (&lt;em&gt;even though I never heard of the term &amp;ldquo;homelab&amp;rdquo; at that point&lt;/em&gt;) with a simple yet powerful setup: Home Assistant running on Proxmox on a NUC8i3BEH.&#10;My Home Assistant instance, with over 100 automations and about 20 addons (docker containers), had outgrown the capabilities of my Raspberry Pi 4, leading me to upgrade to the NUC8i3BEH.&#10;You can check out my &lt;a href="https://github.com/basnijholt/home-assistant-config/" target="_blank" rel="noopener"&gt;Home Assistant configuration&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;This system was more than sufficient for running Home Assistant—arguably overkill.&#10;However, because I passionately develop addons and custom components, I needed the ability to quickly restart and reload everything while debugging.&#10;For a long time, I ran Proxmox, mainly because the Home Assistant community recommended it, but I only used it to host one extra Ubuntu VM and Home Assistant OS.&lt;/p&gt;&#10;&lt;p&gt;In 2022, I stumbled upon the &lt;a href="https://github.com/tteck/Proxmox" target="_blank" rel="noopener"&gt;Proxmox VE Helper-Scripts&lt;/a&gt; by tteck, which got me interested in the whole homelabber/self-hosting community.&#10;These scripts opened a new world of possibilities.&#10;With them, I expanded my Proxmox environment to run 25 different VMs and &lt;a href="https://linuxcontainers.org/lxc/introduction/" target="_blank" rel="noopener"&gt;LXC containers&lt;/a&gt;, each performing a variety of tasks, from DNS servers and reverse proxies to local dashboards and backup software.&lt;/p&gt;&#10;&lt;p&gt;Initially, I used an external 2 TB drive for storage, but it quickly became insufficient as my data needs grew.&#10;After some research, I found the &lt;a href="https://support.hp.com/id-en/document/c06047207" target="_blank" rel="noopener"&gt;HP EliteDesk 800 G4 SFF&lt;/a&gt;, a renewed enterprise-grade desktop with slots for two 3.5-inch drives.&#10;At approximately $160 on Amazon, it offered excellent specs for a budget-friendly price.&#10;I decided to create a TrueNAS VM on this machine using Proxmox and passed through the two 16TB renewed enterprise-level HDDs (via the PCIe SATA controller), which I had purchased from serverpartdeals.com.&#10;I chose these drives based on their reliability as reported in the &lt;a href="https://www.backblaze.com/blog/backblaze-drive-stats-for-2023/" target="_blank" rel="noopener"&gt;BackBlaze drive stats&lt;/a&gt;.&lt;/p&gt;&#10;&lt;p&gt;When I got the HP, I created a Proxmox cluster and was very excited about how easy it was to migrate containers and VMs from one machine to the other.&#10;However, I soon encountered a significant drawback: enterprise drives are notably loud.&#10;Since I had my Proxmox &amp;ldquo;cluster&amp;rdquo; (just the two machines) running behind the TV, the loudness was noticeable and disruptive.&#10;This noise was just the beginning of my troubles.&#10;The NFS server on the TrueNAS VM often crashed, causing the entire Proxmox system to lock up.&#10;Initially, I mounted the NFS drives directly on the Proxmox host and used bind-mounts to pass through the drives.&#10;Due to the locking behavior of NFS, any problem would cause the entire system to lock up.&#10;To prevent this, I started to use &lt;a href="https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/7/html/storage_administration_guide/nfs-autofs" target="_blank" rel="noopener"&gt;autofs&lt;/a&gt; to mount the drives in the LXC containers and VMs instead.&#10;This worked fine until Proxmox started its periodic backups.&#10;Any issue with the NFS during backup caused the whole process to hang, locking up the VMs and LXC containers.&#10;This could be partially avoided by using the &amp;ldquo;stop&amp;rdquo; backup mode, but it still resulted in almost daily debugging sessions where I had to restart Proxmox-related services.&lt;/p&gt;&#10;&lt;p&gt;After running all services as LXC containers, I realized that updating them and keeping up with all security updates was labor-intensive.&#10;So, I started to look into migrating everything to Docker Compose.&#10;This transition turned out great because I could manage my config files in a git repo and keep the relevant data separate, streamlining updates and maintenance.&lt;/p&gt;&#10;&lt;p&gt;To resolve these issues and because I was having a lot of fun, I decided to build a dedicated NAS running TrueNAS Scale, not that I &lt;em&gt;really&lt;/em&gt; needed one, but mostly because I could.&#10;This new setup included high-performance components, ensuring both stability and quiet operation.&#10;I used &amp;ldquo;Manufacturer Recertified&amp;rdquo; &lt;a href="https://www.westerndigital.com/products/internal-drives/data-center-drives/ultrastar-dc-hc550-hdd?sku=0F38356" target="_blank" rel="noopener"&gt;Western Digital Ultrastar DC HC550 HDDs&lt;/a&gt;, known for their reliability and performance.&#10;I have 6 x 16TB = 96TB now with the option to add 2 more 3.5&amp;quot; drives, 2 nvme M.2 drives, and 1 SATA 2.5&amp;quot; SSD.&#10;Now, I run all services that previously relied on NFS locally on TrueNAS via a Debian &lt;a href="https://wiki.debian.org/nspawn" target="_blank" rel="noopener"&gt;systemd-nspawn container&lt;/a&gt; managed with a new script called &lt;a href="https://github.com/Jip-Hop/jailmaker" target="_blank" rel="noopener"&gt;Jailmaker&lt;/a&gt;, which leverages the systemd-nspawn program in TrueNAS Scale (not to be confused with FreeBSD jails).&#10;Inside this systemd-nspawn container, I run Debian with Docker.&#10;For managing Docker Compose files, I use &lt;a href="https://github.com/louislam/dockge" target="_blank" rel="noopener"&gt;Dockge&lt;/a&gt;, which provides a nice web UI for easy management.&lt;/p&gt;&#10;&lt;p&gt;Along the way, I got familiar with ZFS and grew to love it for its amazing snapshotting and data integrity features.&#10;ZFS has been a game-changer, providing robust data protection and ease of management.&#10;Initially, &lt;a href="https://jrs-s.net/2015/02/06/zfs-you-should-use-mirror-vdevs-not-raidz/" target="_blank" rel="noopener"&gt;I used mirrored vdevs&lt;/a&gt;, but I recently switched to RAIDZ2 for the added storage capacity and &lt;a href="https://github.com/openzfs/zfs/pull/15022#issuecomment-1802428899" target="_blank" rel="noopener"&gt;the promise of OpenZFS 2.3 supporting expansion of RAIDZ vdevs&lt;/a&gt;.&#10;I also use ZFS for my Proxmox VMs and LXC containers, whose backups I store on the NAS via rsync in a cronjob to another jail running plain Debian (no more NFS on the Proxmox host for me&amp;hellip;)&lt;/p&gt;&#10;&lt;p&gt;This journey from a simple NUC running Home Assistant to a Proxmox cluster and dedicated NAS has been very time-consuming but a lot of fun.&#10;The lessons learned and the improvements in stability and performance have made the effort worthwhile.&#10;If you&amp;rsquo;re considering building your own homelab, investing in reliable hardware and being prepared for a bit of trial and error can make all the difference.&lt;/p&gt;&#10;&lt;h2 id="lessons-learned"&gt;Lessons learned &lt;a class="anchor" href="#lessons-learned" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Run your services in Docker containers, not LXC containers, for easier management and updates.&lt;/li&gt;&#10;&lt;li&gt;Do &lt;strong&gt;not&lt;/strong&gt; have virtualized NAS (TrueNAS VM) on Proxmox, it&amp;rsquo;s a bad idea if you are also sharing the drives with Proxmox via NFS. There are many folks on the internet who run their NAS in a VM on Proxmox, seemingly without issues, but I was not one of them.&lt;/li&gt;&#10;&lt;li&gt;Do not encrypt the root ZFS dataset! Encrypt its children instead.&lt;/li&gt;&#10;&lt;li&gt;Use passphrase encryption instead of a generated key. It allows you to use the UI and rotating keys.&lt;/li&gt;&#10;&lt;li&gt;Do &lt;strong&gt;not&lt;/strong&gt; remove the physical drive that is passed through in a Proxmox VM, because Proxmox will &lt;em&gt;not&lt;/em&gt; boot properly afterwards.&lt;/li&gt;&#10;&lt;li&gt;Do &lt;strong&gt;not&lt;/strong&gt; use random 100 character passwords for your Proxmox, you have only 60 seconds to enter it during boot.&lt;/li&gt;&#10;&lt;li&gt;When making Proxmox backups, send them to a different physical machine, not only the same machine that is being backed up.&lt;/li&gt;&#10;&lt;li&gt;When building a NAS or home server, consider noise levels, especially if the equipment will be in a living area. Enterprise-grade drives can be surprisingly loud.&lt;/li&gt;&#10;&lt;li&gt;When using Intel I226-V 2.5GbE controllers, be prepared for potential issues.&lt;/li&gt;&#10;&lt;li&gt;When planning your NAS storage size needs, remember to account for the extra space needed for redundancy (e.g., in RAIDZ2). Additionally, keep in mind that ZFS performs optimally when drives aren&amp;rsquo;t filled to capacity. It&amp;rsquo;s wise to plan for more storage than you think you&amp;rsquo;ll immediately need.&lt;/li&gt;&#10;&lt;li&gt;Buy as many used enterprise-grade hardware components as you can, they are often cheaper and more reliable than consumer-grade components.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="components"&gt;Components &lt;a class="anchor" href="#components" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;For completeness, here are the components for each machine in my homelab:&lt;/p&gt;&#10;&lt;h3 id="nuc-components"&gt;NUC Components &lt;a class="anchor" href="#nuc-components" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Name&lt;/th&gt;&#10;&lt;th&gt;Full Name&lt;/th&gt;&#10;&lt;th&gt;Server&lt;/th&gt;&#10;&lt;th&gt;Price&lt;/th&gt;&#10;&lt;th&gt;Quantity&lt;/th&gt;&#10;&lt;th&gt;Cost&lt;/th&gt;&#10;&lt;th&gt;Date&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;NUC8i3BEH&lt;/td&gt;&#10;&lt;td&gt;Intel NUC Kit NUC8i3BEH&lt;/td&gt;&#10;&lt;td&gt;NUC&lt;/td&gt;&#10;&lt;td&gt;278.3&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;278.3&lt;/td&gt;&#10;&lt;td&gt;2020-02-22&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;500GB M2 SSD&lt;/td&gt;&#10;&lt;td&gt;Samsung 970 EVO M.2 500GB&lt;/td&gt;&#10;&lt;td&gt;NUC&lt;/td&gt;&#10;&lt;td&gt;94.99&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;94.99&lt;/td&gt;&#10;&lt;td&gt;2020-02-21&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;32 GB RAM&lt;/td&gt;&#10;&lt;td&gt;TEAMGROUP 32GB DDR4 3200MHz&lt;/td&gt;&#10;&lt;td&gt;NUC&lt;/td&gt;&#10;&lt;td&gt;58.99&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;58.99&lt;/td&gt;&#10;&lt;td&gt;2024-05-01&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;4TB SSD&lt;/td&gt;&#10;&lt;td&gt;Crucial P3 Plus 4TB NVMe&lt;/td&gt;&#10;&lt;td&gt;NUC&lt;/td&gt;&#10;&lt;td&gt;226.99&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;226.99&lt;/td&gt;&#10;&lt;td&gt;2024-04-27&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;!-- | 16 GB DDR4 | Crucial 8GB DDR4 2400 MT/s | NUC | 25.52 | 2 | 51.04 | 2020-02-21 | --&gt;&#10;&lt;h3 id="hp-components"&gt;HP Components &lt;a class="anchor" href="#hp-components" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Name&lt;/th&gt;&#10;&lt;th&gt;Full Name&lt;/th&gt;&#10;&lt;th&gt;Server&lt;/th&gt;&#10;&lt;th&gt;Price&lt;/th&gt;&#10;&lt;th&gt;Quantity&lt;/th&gt;&#10;&lt;th&gt;Cost&lt;/th&gt;&#10;&lt;th&gt;Date&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;64 GB RAM&lt;/td&gt;&#10;&lt;td&gt;TEAMGROUP 32GB DDR4 3200MHz&lt;/td&gt;&#10;&lt;td&gt;HP&lt;/td&gt;&#10;&lt;td&gt;57.99&lt;/td&gt;&#10;&lt;td&gt;2&lt;/td&gt;&#10;&lt;td&gt;115.98&lt;/td&gt;&#10;&lt;td&gt;2024-05-02&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;HP EliteDesk 800 G4 SFF&lt;/td&gt;&#10;&lt;td&gt;HP EliteDesk 800 G4 SFF&lt;/td&gt;&#10;&lt;td&gt;HP&lt;/td&gt;&#10;&lt;td&gt;164&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;164&lt;/td&gt;&#10;&lt;td&gt;2024-04-27&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;!-- | SSD enclosure | UGREEN SSD Enclosure | HP | 16.98 | 1 | 16.98 | 2024-04-27 | --&gt;&#10;&lt;p&gt;Notes on the HP components:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Initially, I planned to use this machine as a NAS because it has two 3.5&amp;quot; drive bays, 2 M.2 slots, and the disc drive bay can be converted to a 2.5&amp;quot; drive bay. However, the 2 x 16TB drives I bought quickly filled up because I used a mirrored vdev, so I decided to build a dedicated NAS instead.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h3 id="nas-components"&gt;NAS Components &lt;a class="anchor" href="#nas-components" aria-label="Link to this section"&gt;#&lt;/a&gt;&lt;/h3&gt;&#10;&lt;div class="table-wrap"&gt;&#10;&lt;table&gt;&#10;&lt;thead&gt;&#10;&lt;tr&gt;&#10;&lt;th&gt;Name&lt;/th&gt;&#10;&lt;th&gt;Full Name&lt;/th&gt;&#10;&lt;th&gt;Server&lt;/th&gt;&#10;&lt;th&gt;Price&lt;/th&gt;&#10;&lt;th&gt;Quantity&lt;/th&gt;&#10;&lt;th&gt;Cost&lt;/th&gt;&#10;&lt;th&gt;Date&lt;/th&gt;&#10;&lt;/tr&gt;&#10;&lt;/thead&gt;&#10;&lt;tbody&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;16 TB HDD&lt;/td&gt;&#10;&lt;td&gt;Western Digital Ultrastar DC HC550 16TB&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;159.99&lt;/td&gt;&#10;&lt;td&gt;2&lt;/td&gt;&#10;&lt;td&gt;319.98&lt;/td&gt;&#10;&lt;td&gt;2024-05-01&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;16 TB HDD&lt;/td&gt;&#10;&lt;td&gt;Western Digital Ultrastar DC HC550 16TB&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;187.5&lt;/td&gt;&#10;&lt;td&gt;4&lt;/td&gt;&#10;&lt;td&gt;750&lt;/td&gt;&#10;&lt;td&gt;2024-06-21&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Motherboard&lt;/td&gt;&#10;&lt;td&gt;Q670 8-bay NAS motherboard&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;203.21&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;203.21&lt;/td&gt;&#10;&lt;td&gt;2024-06-22&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;CPU i5 13500T&lt;/td&gt;&#10;&lt;td&gt;Intel Core i5-13500T&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;180.88&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;180.88&lt;/td&gt;&#10;&lt;td&gt;2024-06-22&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;PSU SFX&lt;/td&gt;&#10;&lt;td&gt;FSP Dagger Pro 650W&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;87.13&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;87.13&lt;/td&gt;&#10;&lt;td&gt;2024-06-21&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;Jonsbo N3 case&lt;/td&gt;&#10;&lt;td&gt;JONSBO N3 Mini-ITX NAS Case&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;187.51&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;187.51&lt;/td&gt;&#10;&lt;td&gt;2024-06-16&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;CPU cooler&lt;/td&gt;&#10;&lt;td&gt;Thermalright AXP90 X36 CPU cooler&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;21.83&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;21.83&lt;/td&gt;&#10;&lt;td&gt;2024-06-25&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;tr&gt;&#10;&lt;td&gt;64 GB DDR5 RAM&lt;/td&gt;&#10;&lt;td&gt;CORSAIR VENGEANCE 64GB DDR5&lt;/td&gt;&#10;&lt;td&gt;NAS&lt;/td&gt;&#10;&lt;td&gt;162.13&lt;/td&gt;&#10;&lt;td&gt;1&lt;/td&gt;&#10;&lt;td&gt;162.13&lt;/td&gt;&#10;&lt;td&gt;2024-06-29&lt;/td&gt;&#10;&lt;/tr&gt;&#10;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Notes on the NAS components:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;The 13500T is a 35W TDP CPU I bought on eBay, which will help keep the system cool and quiet.&lt;/li&gt;&#10;&lt;li&gt;I bought the motherboard from AliExpress because it seems to be the only mini-ITX motherboard that supports 8 SATA drives, and also comes with 2x2.5GbE ports, 3xM.2 slots, and a full sized PCIe slot.&lt;/li&gt;&#10;&lt;li&gt;The motherboard uses Intel I225-V 2.5GbE controllers, which are notoriously bad, initially &lt;code&gt;iperf3&lt;/code&gt; maxed out at 100Mbps. The internet is full with people complaining about these controllers auto-negotiating to 100Mbps, however, for me it showed 1Gbps. After trying all suggestions I could find (there are few), as a last resort I connecting the NAS directly to my ASUS XT8 router instead of the NETGEAR GS305 switch, and now I got 1Gbps speeds. So I bought a new managed TL-SG108E switch, which now also gives me 1Gbps speeds. I&amp;rsquo;m not sure what the issue was, but I&amp;rsquo;m happy it&amp;rsquo;s resolved.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;</description></item></channel></rss>