The NPU Works. The Compiler Doesn't.

I wanted a cheap, low-power, always-on person detector. Of course, I’m building this on a BrightSign XT5 player (more on that later). But since a friend said he was using a Beelink SER9 with a AMD Ryzen AI NPU (~50 TOPS) I thought I’d check it out. Can it actually run the detector? Short answer, as of today: no — and the reason is not the part you’d expect.

Here’s the twist that makes this worth a blog post. The hardware works. The driver works. The runtime works. I ran tiny kernels on the NPU tiles and watched them come back correct. Every layer of the stack passed its test… right up until the very last one. The thing that’s broken isn’t the silicon. It’s the compiler.

What I was actually trying to do

The pitch for an NPU is simple: it’s a dedicated inference block that does sustained neural-net math for almost no power. If I want a box watching a few cameras - maybe more than a few - I do not want to burn 40–60 watts pinning the Radeon 890M iGPU to do it. The NPU should do that job at a few watts and never break a sweat.

I’ve written about running YOLOX on an NPU before — that’s exactly what Argus does on BrightSign players, and it’s genuinely boring in the best way. Drop a file on an SD card, plug in a webcam, boot. So I went in expecting the AMD part to be fiddly, sure, but fundamentally a solved problem. But hey, why not see if the AMD box can do the job?

Everything worked… until the last step

Let me be specific, because “it doesn’t work” is a lazy thing to say and I hate reading it in other people’s posts. Here’s what I got working, layer by layer:

  • Kernel + driver. Ubuntu 26.04 on kernel 7.0 ships the amdxdna driver in-tree. No DKMS, no out-of-tree module dance. It just shows up at /dev/accel/accel0.
  • Runtime. Built XRT from source (SHIM-only, since the kernel already has the driver), and xrt-smi examine cheerfully reports RyzenAI-npu4, arch aie2p, 6×8 tiles. The box sees its NPU.
  • Actual kernels on the actual tiles. Using IRON/MLIR-AIE I ran a passthrough kernel (~97 µs of NPU time), a conv2d, and a ResNet bottleneck block. They ran on the NPU and produced correct results verified against PyTorch. The silicon executes compiled compute. Full stop.
  • The model compiler’s front half. Feeding an ONNX model through IREE + the amd-aie backend: it imports the ONNX, lowers it through torch → linalg, and splits it into dispatches. All clean.

And then, at the final translate-to-AIE step — the part that turns a lowered op into code that actually runs on the tile array — it falls over:

error: failed to run translation of source executable to target executable
for backend #hal.executable.target<"amd-aie", "amdaie-pdi-fb", ...>

An arbitrary convolution (16-channel, 3×3, 32×32) fails. A little matmul-with-bias — i.e. a fully-connected layer — fails. The one op class the open-source iree-amd-aie code generator lowers reliably today is plain matmul. Its whole CI is basically built around a run_matmul_test.sh.

Think about that for a minute. A convolution is the fundamental building block of every CNN detector on earth. If you can’t lower a single conv onto the tiles, a YOLOX backbone doesn’t have a prayer of compiling. Not because it’s too big. Because op #1 doesn’t have a code path yet.

This is a compiler problem, not a hardware problem

This is the whole point, so let me say it plainly: the gap is in the open-source AIE code generator’s operator coverage, not in the chip.

Lowering a convolution onto a spatial array of AI-engine tiles is genuinely hard — you have to tile the computation, choreograph all the data movement between tiles, and generate per-core kernels. The amd-aie backend just doesn’t implement that for the general case yet. It does matmul. Real models interleave conv, matmul, pooling, SiLU, concat, resize, and NMS, and today a single global tiling-pipeline flag has to serve every dispatch. That’s not enough machinery to compile a detector.

And yes, there’s an “official” high-level path — ONNX Runtime with the VitisAI execution provider. It’s Windows-only. On Linux the voe graph-optimizer wheels are missing and a config-parser bug quietly shoves every op back onto the CPU, so you think you’re using the NPU and you’re not. Ultralytics and PyTorch don’t have a device=npu either. The Linux high-level story just isn’t there.

Don’t get me wrong — none of this is me dunking on AMD. The hardware is real, the driver landed in-tree, and the runtime foundation is solid. This is an ecosystem-maturity gap, and those close over time. It reminds me of something I say about the models all the time: today’s models are the worst you’ll ever use. Same energy here — today’s iree-amd-aie is the worst conv coverage it’ll ever have. It only goes one direction.

So what do you do right now?

Seriously? Go build it on a BrightSign! Stay tuned…

I’ll revisit this the moment iree-amd-aie grows a real convolution path — and when I do, you’ll read about it here first.

If this saved you a weekend, or if you’ve gotten a conv to lower on XDNA2 under Linux and I’m just holding it wrong, drop me a note on LinkedIn. I’d genuinely love to be wrong about the timeline.

 Share!