I had a customer who had me integrate an algorithm into an existing firmware project. That firmware had no testing outside of whole-product testing — no unit tests, no static analysis, no code review, and seemingly no software quality system of any kind. That said, it seemed to work — it passed all their whole-product tests.
I tried to explain the situation to the customer, but they were under a tight deadline. So we came together with a reasonable list of fixes for the larger core issues. That way we could get the algorithm integrated while fixing some known problems along the way.
When they were getting ready to ship, they asked me how confident I was that the firmware was ready for production.
I told them: it should be OK for the first one hundred thousand units, and then it's going to need some TLC.
They were pretty shocked by that answer. They wanted me to say it was perfect. We fixed the obvious issues and it still passed on the bench: it appeared to work just fine.
But when you have some bad, poorly written, or very messy code, that doesn't make it wrong. It could be just fine. Most likely it's going to produce some weird operations in the product.
I think most people have some experience with the result: a product that acts weird or intermittently weird. The appliance that ignores a button press every once in a while. The light takes a few seconds to turn on after opening the refrigerator. The microwave is stuck at 0:00 when it's done running. Maybe you watched a product wake up from sleep randomly when no one is in the room.
There are good reasons that stuff exists — even in high-volume products from companies that have tons of experience. Here's where it comes from.
Firmware doesn't get to quit
The first thing that's different about firmware is how long it runs.
When someone writes an app or a website, a user is on it for half an hour, maybe an hour. Call it sixteen hours in an extreme case. Then it gets closed, the process dies, and everything starts fresh next time.
Firmware runs 24/7, 365. That's not the exception — that's the normal setup. The thermostat, the mouse, the microwave, the thing in the wall you've never thought about. They boot once and they're expected to keep running for years.
So take something completely ordinary: a millisecond counter in a 32-bit number. Let's say you want something to happen 10 seconds after the last button press. If your code looks like this:
if (get_time() > last_button_press + 10*1000) {
// Start sleep mode
}
This code will go into sleep mode 10 seconds after the last button press... but it will skip the delay when the counter is close to rolling over. A 32-bit millisecond counter rolls over at about 49 days. If it's microseconds, it's about 71 minutes.
49 days is a funny number, because most final product testing doesn't run for 49 days. Nobody's life test is long enough, and nobody's bench test comes close to catching every issue a sloppy development process leaves behind. So if the product does something odd at the rollover — say it instantly goes to sleep after day 49 — no one may see that before it ships. And it's so mundane that many people may think they didn't press it right, or something similar...
That's a flaky product. It's not a dangerous product. And that distinction matters more than you'd think.
Firmware is designed to be safe, not to be good
This is something people outside of firmware don't usually think about: firmware is safe first.
The primary purpose of most of the testing, most of the certification process, most of the effort, isn't to make the best product. It's to make the safest product. A ton of care is put into ensuring it won't hurt someone. And as an industry we've done a really good job of that.
But you can have a product that's completely safe — nobody can get hurt by it — and still be a bad product to the customer. Remember, a product that never turns on is potentially very safe.
Take a microwave. It won't run with the door open, which is exactly right. Now say the door switch develops a fault, and the microwave can't tell the door is closed. It refuses to run. Safe! But some microwaves don't tell the user anything, so the person is standing there opening and closing the door, and as far as they can tell the microwave is just... broken.
That's the kind of flakiness I'm talking about. Not "this will hurt someone." Just "this is a less perfect product than it should be." And that category of problem gets almost none of the attention, because all of the process is pointed at the other one. These are more brand risks than safety risks.
The real world is asynchronous
The other thing firmware has that most software doesn't: we butt up against the real world.
A firmware program has lots of asynchronous inputs — buttons, sensors, timers, communication lines — and they show up whenever they feel like it. That makes the program non-deterministic in a way a web request just isn't. We can't control any of it. The firmware has to survive it and run well inside of it.
Say you have 30 inputs on a microcontroller. You can't test every combination of those 30 inputs. If it takes only 1 second per test, it will take >30 years to test all combinations. So every firmware product ships having been tested as well as it practically can be, and the rest has to come from design. You can try some of these things on your products now by pressing multiple buttons at the same time, etc.
The physical world has a range of conditions. Sometimes a button generates very different bounces (noise) after it has been pressed a million times compared to when it's new. Firmware that doesn't account for this is flaky for people after their products have been used a lot.
The problem is what happens when several of those asynchronous events line up at the same instant.
Do the multiplication
Here's the basic math, and it's the whole reason I gave the customer the answer I did.
Say one event happens ten times a second, and another happens once a minute. Each one takes a microsecond to handle. This could be from some part of the hardware or something internal to the processor. The odds of them landing on top of each other on any given occurrence are tiny — something like two in a hundred thousand.
On one unit on your bench, that's a collision about once a month. You'd probably never see it. You definitely wouldn't see it in a two-week life test, and if you did, you'd never reproduce it.
Now ship 100,000 units. All of them are running 100% of the time.
That same collision is now happening somewhere in your installed base roughly 3,000 times a day. Every day. To some of your customers. If the collision causes something bad, it's not a rare bug anymore — it's a support queue.
This is the same thing I wrote about in Data Faults, just from the business side instead of the disassembly side. A one-in-a-thousand interrupt timing bug is invisible on your desk and a disaster at volume. The multiplication is the entire story.
And again — most of these are performance problems, not safety problems. Nothing catches fire. The product just does something slightly wrong, for a few people, every day, forever.
Forever, because it won't get fixed
Which brings up the last difference: most firmware is not field upgradable.
Think about the devices you interact with every day. Your mouse. Your keyboard. Your appliances. The vast majority of them will never get a firmware update. Whatever shipped is what's running until the thing goes in the trash.
In software, a bug that hits 0.1% of users is a hotfix. In firmware, it's a permanent property of that product generation. So the calculation is different: it's not "how bad is the bug," it's "how bad is the bug, times how many units, times how many years."
Back to the customer
So how did that customer end up shipping firmware with no firmware quality plan behind it?
When a product with tens of thousands of lines of code doesn't have a quality plan, the outcome basically depends on the experience and care of whoever wrote it. That doesn't mean it's going to be a bad product. But when it's built by an outside vendor — someone who may not have the long-term interest of the brand in mind, just the commitments in the contract — you can have a lot of risk built in without anyone noticing.
And in this case, I noticed red flags all over that codebase. But I don't think it was a breach of contract, and honestly I wouldn't blame the engineer who wrote it, or even the people who originally planned it. The vendor had built the firmware as requested. The problem was that the request was too vague.
The spec said things like: I want this many buttons, I want a screen here, I want it to work like this and that. Here's the thing — none of that implies "and I need to ship a million of these and they need to work flawlessly for ten years." It doesn't say what failure rate is acceptable. It doesn't say the product should do the desirable thing every single time, not just avoid the unsafe thing.
A good test for your own spec: am I asking for a proof of concept? And could anyone tell the difference from what I actually wrote down?
If your spec would be satisfied by a prototype, don't be shocked when you get one.
What to ask for if you're outsourcing firmware
If you're specifying a firmware product — or any product with firmware in it — and someone else is building it, there are a few easy things to ask for up front.
First, ask what their standard of quality is. It's fine if they don't have one written down but are willing to work to yours. If so, I'd ask for at least the following:
- A top-level diagram showing the modules of the firmware.
- A testing plan — including who is responsible for which parts.
- Static analysis — can uncover potential issues
- Code review — actual review, by someone who didn't write it.
- Unit tests on the pieces that can be tested without hardware.
- Regression or hardware-in-the-loop testing, if needed.
None of this is exotic. Most of it is cheap now — I've written about how AI agents make automated firmware testing practical for teams that never had the budget for it before. The expensive part isn't the tooling — it's finding issues after shipping 100,000 units.
The question that matters
I think the right way to frame every one of these decisions — the testing, the process, how much to spend on making the firmware better — is two questions:
How many units are we shipping? And what's the value of a better product against the cost of building it?
If the answer is fifty units for a pilot, a proof of concept might be exactly right. If the answer is a million units with no update path, then "it passed product testing" isn't a quality plan. It's a hope.
The multiplication doesn't care either way. It's going to happen. You just get to decide whether you did it before the customer did.