Writing · draft
Model-offloading & the price of lost confidence
When a smaller or local model wins — and when the harness to trust it costs more than the model you saved.
Waiting on real numbers. I'm holding this one until I can show measured routing costs rather than assert the tradeoff — the runs I'm instrumenting with cLens are the evidence it needs, so it lands after the LLM fundamentals piece. The argument below won't change; the worked example is what's missing.
Sending non-interactive work to a cheaper model can be a real win — until the verification and guardrails you need to trust its output cost more than the model you saved. The tradeoff is reliability, not just price; the harness is what makes the call legible.