Compiling an LLM on a Hailo 10H, a Bottomless Shithole

Why This Project The Hailo-10H is the NPU Hailo sells for running LLMs on a Raspberry Pi 5 (and other SBCs). Unlike its vision-oriented cousins, this one is built from the silicon up for exactly that: 8GB of LPDDR4 on board! At first, I tested my chip with hailo-ollama, Hailo’s homegrown “Ollama-like” tool that drives their precompiled models. I had fun for about two minutes, and then got bored fast. The models on offer are slow, and worse, locked to one very specific list: ...

August 11, 2026 · 23 min · Léo Nonnenmacher

AI on Cards That Came Out of Nowhere

What’s the Point? Ever since I got into AI, I’ve enjoyed messing around with weird GPUs. After putting together An AI Setup Cheaper Than Mario Kart World?, I wondered whether decent AI models could run on genuinely cheap GPUs. Spoiler: yes, they can. With a few dark spots along the way. The Setup The setup is more or less the same as in the previous article. As a reminder: HP 700W PSU Intel Xeon E5-1620 v3 @ 3.50GHz 32GB RAM DDR4 ECC 2133MHz Samsung EVO 250GB SSD Quadro K2200 4GB GTX 1050 Ti 4GB But we’re adding 2 slightly suspicious cards to the mix: ...

August 11, 2026 · 29 min · Léo Nonnenmacher

An AI Setup Cheaper Than Mario Kart World?

🔴 Why bother running AI at home? Fair question, right? Why run AI at home at all? And actually, let’s back up further: what even is AI? What is AI? According to Wikipedia: Artificial intelligence (AI) is the capability of computational systems to perform tasks typically associated with human intelligence, such as learning, reasoning, problem-solving, perception, and decision-making. In short, AI means using mathematical functions and operations to automate and speed up complex tasks by reproducing human behaviors like reading, seeing, writing, or understanding language. Thanks to these calculations, machines can learn to recognize images, understand text, make decisions, or even generate content — often faster and sometimes better than we can (though honestly, that’s debatable). ...

June 29, 2025 · 10 min · Léo Nonnenmacher

My First Project With an NPU (Hailo‑8L): From MNIST to Embedded Inference

🎯 Why This Project? At first, I thought the Hailo‑8L was some kind of magic chip — the kind where you could run LLMs, use it as a backend device for frameworks like Tensorflow, train models on it, all of it. Spoiler: absolutely not. Digging in a bit, I understood that the Hailo‑8L is built exclusively for inference, not training, and definitely not [Reinforcement Learning] or any kind of LLM business. ...

June 27, 2025 · 6 min · Léo Nonnenmacher