A couple of years ago, I was working on several different projects, and I kept running into the same issue: I needed to unit-test multiple different microcontrollers. I was supporting multiple bare-metal platforms with a single codebase. This is often tricky, since it's easy to fix something for one processor and break it for another.
I ended up with a handful of USB devkits and production boards plugged into one of my machines. That allowed me to run unit tests on them. A tiny Python utility let me manage them all over the network, and I was off and running.
That worked okay for a little while. I think things started breaking down around five or six boards. I started having all kinds of strange USB issues. Sometimes the order I tested the boards in, or what the code was doing on the boards, affected the next test. For example, if I ran a test on a board that continuously sent data over USB, the next board might hit bandwidth limits. I ran a few benchmarks (basic arithmetic stuff), and since the previous processors were still running, I started hitting the power limits of the USB ports. I was unplugging and replugging boards, resetting USB hubs, and it was just a mess.
The core problem was that I needed to turn the ports on and off individually. Some USB hubs do support this — so I bought several hubs reported to work that way — and I didn't have very good luck. So instead of throwing good money after bad, I laid out a board with sixteen USB ports, several USB hubs, and a microcontroller that can individually control the power to each port while monitoring the power draw.

If that board looks familiar, it's the one buried under the pile of dev boards in last week's article, where an AI designed the enclosure the whole rig now lives in.
Standard method
This is pretty much the same recipe as any hardware-in-the-loop setup in CI:
- Build some custom hardware that connects to your prototype or production board, usually over USB or serial. The board often needs a few modifications.
- The hard part: you need a way to program the board while it's wired in.
- Write a very basic service to drive it (20 lines of Python with Flask is often enough for an intranet tool).
- Use your existing build scripts to build the firmware whenever a commit lands (basic CI).
- Extend that hook to run one basic integration test that proves the firmware works.
- Add more tests, and keep going.
The outcome
I built it for running unit tests, and now it has a whole new set of uses.
I can now have agents writing more reliable firmware, since they can write code and test it immediately on REAL hardware. This closes the agent feedback loop correctly. That has let me get pretty good results from firmware developed in part by agents — they can validate their own work on real hardware.
I'll often have a small-to-medium-sized firmware project written mostly by agents. Since agents need so many guardrails, there will be thousands of unit and integration tests run on REAL HARDWARE. As these projects have progressed, I can be more and more aggressive with my timelines and performance targets, because I can measure most of the system.
Ironically, these agents use the test harness far more often than I do. They have found all kinds of interesting systemic problems. Here's my favorite.
Tests failing randomly
I had agents working on firmware for a single device, so I gave them limited access to the test harness. They started seeing wildly inconsistent results. Tests were passing intermittently, and it seemed like some of the boards were locking up. This was confusing, because I had run hundreds of tests without any issues. I was extremely confident in the rig — and I figured the agents were hallucinating or using the tools wrong.
I tracked it down by watching an agent run the same test five times: it passed the first time, then failed the next four. The agent decided the board was dead.
The agents even started building out a whole health-check system for the test framework (useless for the firmware project they were supposed to be working on), which was just more confusion for me to untangle.
Then I ran the same test five times in a row by hand, and IT PASSED EVERY TIME!
The issue turned out to be how humans and agents use the tool differently. The time between tests for the agents was so short that the capacitors on the boards hadn't fully discharged. When I ran tests by hand there was a long gap between runs. Even in CI, each board gets time between jobs, whether the boards run in sequence or in parallel. But when an agent runs a series of tests, it may make a tweak and send another test to the same board within seconds.
I added a forced delay between back-to-back tests on the same port, and it has worked flawlessly since.
All that aside, this may be the most useful new tool I have. I've learned a ton about performance, and gut feelings I've had in the past I can now actually measure.
Serious firmware teams should already have some hardware-in-the-loop testing. It doesn't need 16 ports or many different processors. Just cover whatever your biggest effort is right now. It pays for itself in a single project: you'll get more done and catch quality issues before they ship.
Once you have a rig for testing with real hardware, it's pretty straightforward to set up unit and integration tests to run automatically. This lets your engineers know immediately when a commit breaks something, which usually makes it an easy fix.
That makes your firmware better and makes your engineers more productive.
Next week I'll share the first of several benchmarks I've run across all of these boards on this test harness.