Here’s a question worth more than most of the AI hype takes clogging your feed: what happens to your velocity the day your frontier provider changes the deal? Raises the price. Deprecates the model you tuned your whole workflow around. Throttles you at the worst possible moment. Or just decides your use case, your industry, or your country isn’t one they want to serve anymore. If your honest answer is “I’d be dead in the water,” then you don’t have a strategy — you have a dependency. And it’s time to look hard at running some inference on your own metal.
The lamp on my office shelf glows solid red the instant my camera goes live, and turns off the second the meeting ends. Nobody has to knock and wonder. My wife, a kid, the dog-walker — anyone who glances in knows I’m on camera without me saying a word. And by my own stubborn definition, that little bulb is a robot.
Here’s a result that surprised me: the reasoning model wrote more correct code than the plain one — and still lost. It nailed all 32 test cases, then sat there and refused to declare itself done. Burned the whole time budget re-checking work that was already right. “Correct, but won’t stop” turns out to be a real failure mode, and it’s exactly the wrong one for an autonomous agent loop.
A room full of CTOs and architects met in a Swiss village this summer and, without quite meaning to, gave a name to the thing I’ve been quietly building for months. They call it harness engineering. I’ve been calling it “the tooling is the new model.” Same animal. And the part they flagged as unsolved, I think I’ve got a first, opinionated stab at.
Two and a half years into my whale song project I’ve killed one platform, walked away from a $13K dev kit, gotten my ham license, and pivoted to (water-borne) drones. I can do that link myself across the Bay of Banderas with a radio I’m now legally allowed to key up. Read on.
My experience is that Claude Code is great with Foundation Models, but way too heavy for local inference. I’ve been enjoying oh-my-pi (omp) but was surprised that hax (a simpler agent in written in C) was TWICE as fast and used HALF the tokens! It’s all about PREFILL. Read on!
I’ve written before about what a BrightSign extension actually is and about why edge AI has to run locally on signage. This post is the “show, don’t tell” follow-up. If you want to see what all that theory looks like as a running, watchable thing, there’s exactly one place to start: Argus.
XDA ran a piece that stopped me mid-scroll: a 27B open-weights model reverse-engineered a commercial application’s license check — recovered a deliberately obscured crypto key out of ARM64 assembly, caught and corrected its own mistake without being told, and produced a working bypass PoC. In about thirty minutes. On a desktop box. I have (a version of) that box. So I went and set it up.
I use more than one coding agent. Not because I’m indecisive — because I’m testing different agents - and learning how they work - and the field is moving fast. Seems like a new one every week! But every single one of them wants skills in its own directory, and I got tired of cp -r and stale copies. So I wrote a tool.
You can’t lead From hehind anymore. For twenty years the path to Engineering leadership had a well-worn groove: write code for a few years, get “promoted out of” writing code, spend the rest of your career managing budgets, headcount, and status reports. The further up you went the less technical you were expected to be, and nobody blinked. That era is over. Not ending. Over.
See all posts →