I ran five different coding agents at the same local model, on the same task, with the same frozen test suite — and then I counted why they failed. The answer wasn’t subtle. About 90% of the failures were harness problems. Only about 10% were the model. And here’s the part that should change how you spend your next dollar: throwing a bigger, less-quantized model at it fixed none of them.
Here’s a question worth more than most of the AI hype takes clogging your feed: what happens to your velocity the day your frontier provider changes the deal? Raises the price. Deprecates the model you tuned your whole workflow around. Throttles you at the worst possible moment. Or just decides your use case, your industry, or your country isn’t one they want to serve anymore. If your honest answer is “I’d be dead in the water,” then you don’t have a strategy — you have a dependency. And it’s time to look hard at running some inference on your own metal.