Last week I showed you the rig: fifteen microcontroller boards hanging off a 16-port USB hub that I set up for unit testing many different microcontrollers. This is one of the first experiments I ran on it to just get some basic arithmetic tests on each microcontroller. I just wanted to know how fast different microcontrollers can add/subtract/multiply/divide.
I'm often trying to run different filters on different microcontrollers, and I was hoping to get a feel for how much is possible on different microcontrollers. In a perfect world, every engineer could benchmark every potential micro for their application, but most of us live in the real world... So here's some numbers that might help people...
The fleet is 16 build targets on 15 physical boards across six ISA families: Cortex-M0+/M4F/M33/M7, RISC-V, Xtensa, AVR, MSP430, and Microchip's PIC16/PIC18/PIC24/dsPIC33A. The Pico 2W counts twice, because the RP2350 ships with both ARM and RISC-V cores and I built for each. I had an 8051 on there too, but it was too fussy (I kept having to hold the reset button on it)... so I removed it for now.
Most of these benchmarks were run by agents, so take the results with a grain of salt. I did a quick sanity check on them, and I found the agents lying a bunch, or at least not understanding what it takes to set up a microcontroller. If you are making a specific purchasing decision, I'd confirm your expectation of performance for a given microcontroller.
Lie one: the clock
The dsPIC33AK512MPS506 is a 200 MHz part. In my first fleet capture it ranked 3rd last for uint16 add performance -- when it should be one of the faster micros... It was slower than an 8-bit ATmega. Each agent was trying to set up the micros for performance (without getting lost in compiler options, often -O2 or -Os).
So basically the agents had to set up each microcontroller (with toolchains) in Linux (where they live) and get the following set up:
- Clocks - that we'll use for the time base
- Serial interface (so our test can talk to it)
- Baseline toolchain and successful programming setup
- Measure the clocks via the serial interface - by testing on real hardware
- Validate the serial interface - by testing on real hardware
So for most of the boards, it worked well. For the dsPIC port, I set it up for disaster. Initially, I thought I had a different (much slower) dsPIC. It failed for a while until I figured out the issue, and redirected it to the correct part. The problem is the agent had locked in its clock settings, and when it "ported" the code to the correct part, it ran the tests and everything at an 8 MHz clock rate. Basically it never set up the PLL (which was different between parts)... and so it just assumed it was running at the correct clock speed.
Once I realized the clock was wrong, it was easy to get an agent to fix it. After the fix (and revalidating the clock setup), it moved the dsPIC from 14th to 2nd for uint16 add, behind only the 600 MHz Cortex-M7.
The MSP430FR2433 had the same class of bug in a smaller size. Its clock is a frequency-locked loop multiplying a 32.768 kHz reference, and the multiplier I'd set targeted 8 MHz. The part runs at 16 MHz, which also needs one FRAM wait state above 8 MHz, so the fix was two lines. Its numbers doubled.
Here's the fun part: When I recaptured both boards the performance gain on the dsPIC and MSP430 was 25x and 2x respectively. Exactly the same as the clock changes! That's exactly what a pure clock error looks like, and it's a free cross-check: if the ratios hadn't matched, something else would have been wrong too.
The lesson: a misconfigured clock doesn't necessarily fail. It just makes certain things take longer (or shorter). This is (hopefully) less important now that most firmware projects are using tools to bootstrap their projects (e.g. code generators or vendor IDEs).
A huge shoutout to Microchip for making their XC compiler families free with optimization. In the past, they used to charge for the optimized version but this summer they made them really free! This is a huge win for all of us -- even if you don't use PIC parts, since it pushes more vendors to open their compilers to the community at large. I've spent time on a development team with a limited number of compiler seat licenses -- it can be a real hassle.
Lie two: the compiler
The PIC16F17146 is a 32 MHz 8-bit part and it was doing 16-bit adds faster than some 32-bit processors... this obviously couldn't be true. I added the PIC16 family in this list because people often need to compare it with some low-cost ARM processors. So when I saw the PIC16 coming close to the 32-bit processors, I knew something was wrong. I had to dig into the assembly code to see what was going on.
The benchmark kernel was 32 arithmetic expressions in a row, separated by empty inline-assembly barriers. The GCC-based toolchains honored it. Microchip's XC8 did not - so it just optimized out most of the operations. Every PIC16 and PIC18 add, subtract, and multiply number was 50 to 110x too high. The firmware built, the correctness check passed, and the timing gate passed.
So to fix this, I added a memory read before each operation and a write after each operation. I then verified this in the assembly code. This is an expensive fix for the very fast parts, since the memory operations can actually be slower than the calculations. That just means there is a baseline overhead to all of these numbers and it makes sure we can compare the values since it forces all the compilers (including XC8) to actually do the work.
So after all of that stuff, I re-ran all the tests (about an hour of the test fixture) and 94/94 worked!
What I actually measured
Five operations (add, subtract, multiply, divide, and modulo, with fmod for the floating types) across six numeric types (uint8 through uint64, float, double) on every target. Each timed kernel is 32 arithmetic expressions in plain C: four operand pairs, looped eight times, with every expression reading its operands from volatile memory and writing its result back. Three calibrated samples of about 500 ms each, no baseline subtracted, medians reported. Coverage is 94 of 94 measurable board/type pairs across all 16 targets. The only two exclusions are uint64 on the two 8-bit PICs, because XC8 has no 64-bit integer type.
The extra reads and writes are the price of the compiler fix and they put an artificial limit on every number. The fastest thing a board can report is the cost of two loads, a store, and a loop step... even if the operation is free. On many of the fast 32-bit cores that limit can be greater than the cost of the add, subtract, and multiply. On the SAMD51, the STM32G474, and both Xtensa ESP32s, those three columns are identical to the nanosecond: the operation is invisible under the memory traffic. A single-cycle add on a Cortex-M4F shows up here as about 13 cycles.
So take these benchmarks with a grain of salt, since you're more seeing the memory measurement vs ALU speed. These numbers become informative where the operation is expensive enough to rise above the floor: division and modulo everywhere, 64-bit integers on narrow cores, floating point without an FPU, and nearly everything on the 8- and 16-bit parts. Ratios within one board should be trustworthy, since they are all measured the same way.
One more convention. "Cycles" means oscillator cycles, on every board. The PIC16 and PIC18 execute one instruction per four oscillator cycles and the PIC24 one per two, so on a per-instruction basis they'd look 4x and 2x better than these charts show. I normalized to the oscillator because it's the number on the front page of the datasheet and the number you set in the config bits. It's a choice, and it's the same choice everywhere.
The rankings
The Cortex-M7 at 600 MHz is untouchable in absolute terms: 75 Mops/s for uint16 add, about 8 cycles per expression, and it's the only board where the floor is under ten cycles.
Then comes a pack of nine 32-bit parts between about 9 and 13 Mops/s: the dsPIC33A, the ESP32-S3, both RP2350 builds, the STM32G474, the ESP32, the RP2040, the ESP32-C6, and the SAMD51. That's the floor talking. The order inside that pack is noise. The dsPIC lands second for uint8 and uint16 add and ninth for uint32, which is the compiler's loop code, not the ALU.
Below the pack, clock speed takes over. The XIAO MG24's Cortex-M33 runs essentially the same 13-cycle loop as the RP2350's, but at 39 MHz, so it does 2.9 Mops/s. Then the 8- and 16-bit parts: the PIC24F at 1.15 Mops/s, the MSP430 at 646 kops/s, the Uno's ATmega328P at 522 kops/s, the PIC18 at 328 kops/s, and the PIC16 at 244 kops/s. Note that the 16 MHz Uno beats the 64 MHz PIC18 on 16-bit adds: 31 cycles per expression against 195. Even after the four-cycles-per-instruction correction, that's 49 instruction cycles on the PIC18.
Cycles per operation
Switch the metric to cycles per operation and the chart stops being a clock-speed contest. Each board's five bars become its cost profile, and the interesting part is the right-hand side of every group: divide and modulo.
On cores with a hardware divider (the M4Fs, the M33s, the M7, the two Xtensa ESP32s, the dsPIC) a uint16 divide costs 5 or 6 cycles more than an add. On the RP2040's Cortex-M0+ it's 42 cycles against 19. On the RP2350's RISC-V build it's 30 against 13, about double the ARM build on the same chip. On the narrow parts it's a cliff: 230 cycles on the Uno (7.5x its add), 238 on the MSP430 (nearly 10x), 1,030 on the PIC18, and 1,357 on the PIC16. Widen to uint32 and the PIC16 needs about 3,500 cycles per divide. That's over 100 microseconds at 32 MHz, or nine thousand divides per second.
Multiply depends on what's in the silicon. The PIC18 has a hardware 8x8 multiplier, so a uint8 multiply costs the same as an add (123 cycles against 119). The PIC16 doesn't, so a uint16 multiply costs 1,261 cycles against 131 for add, almost 10x. Ask the PIC18 for a uint32 multiply and it's 2,700 cycles, 10x its uint32 add. The ATmega328P's hardware multiplier keeps uint8 multiply within 10% of add, but uint32 multiply runs at 135 kops/s, 2.5x slower than uint32 add.
Float is free, until it isn't
The boards with an FPU do float add at integer speed. On the i.MX RT1062, the RP2350's M33, the MG24, the SAMD51, the STM32G474, and both Xtensa ESP32s, float add is within a cycle or two of uint32 add. There could be some cost differences, but they are dwarfed by the load/store loop.
Micros without floating point hardware are about 5x slower. The RP2040's M0+ goes from 18 cycles for a uint32 add to 93 for a float add. The ESP32-C6 goes from 13 to 67, and the RP2350's RISC-V build from 12 to 62. The cleanest comparison in the fleet is again the Pico 2W against itself: 13.4 Mops/s float add on the ARM build, 2.4 Mops/s on the RISC-V build. Same C code, same chip, 5.5x.
Float divide costs about 2x a float add on the ARM FPUs and 3 to 4x on the two ESP32s. On micros without floating point hardware, it's 100 to 200 cycles on the 32-bit cores and 400 to 6,300 on the 8- and 16-bit ones.
So the double floating point benchmark comes with some HUGE asterisks. A few of the boards have true double performance hardware: The i.MX RT1062 has a double-precision FPU (10.7 cycles per add, 56 Mops/s), the dsPIC has double precision hardware (untested, see below), and the Cortex-M33 (RP2350) has a double-precision coprocessor next to its single-precision FPU (26 cycles, 5.7 Mops/s). Every other Cortex-M FPU in the fleet is single-precision only, and switching from float to double costs about 8x: the SAMD51 goes from 13 cycles to 100, the STM32G474 from 15 to 108, the MG24 from 12 to 106.
Fake doubles are everywhere!!
And watch for fake doubles. On the AVR and on all four Microchip targets: a double is 4 bytes! The dsPIC's double numbers are identical to its float numbers because it's the same type -- the dsPIC has some double supporting hardware, but it maps double to be float, and wants a long double instead. The slowest double operation is the MSP430FR2433 doing double multiply at 942 operations per second, about 17,000 cycles each. Not kops. Ops. From the fastest cell in the fleet (the i.MX's uint32 subtract, 86 Mops/s) to that one is a factor of 91,000. I wonder if we had 8-byte double support on some of the 8-bit processors, if they would be slower than the MSP430...
What this means for your firmware
Last week's article had a test rig that passed every time I used it and failed when the agents did. This experiment had a board that passed every check at the wrong clock, and a compiler that passed every check while deleting the work. Same lesson three times: a green checkmark tells you the harness ran. It doesn't tell you the number is right.
- Verify your clocks. Time a known interval against a wall clock, once, on hardware, and log the result with every benchmark. A wrong clock may not error out... it could make you run slower (giving up performance), or faster (potentially running out of spec)... And it could still "work" on your desk the entire time.
- Verify your benchmark survived the compiler. Read the assembly (at least once), or force read/writes and accept the overhead. A better approach is to measure your application directly on different parts!
- Trust ratios within a board more than rankings across boards. The floor cancels inside one board. It doesn't cancel between two compilers.
- Match your data width to your core. On the PIC16, a uint32 add costs 3x a uint8 add (303 cycles against 107). On the Uno, uint64 add is 4.6x uint32.
- Division is expensive... without a hardware divider it's a lot! Think twice before doing a real division in a tight loop.
- Be aware of the floating point hardware in your part and the limitations of your software environment. It's even worse with so many compilers using floats when they are asking for doubles... it's something to be aware of on many of these platforms.
Every chart in this article is interactive at /fleet: switch the type and the metric, or open a page for any board or any pair of boards. The methodology page spells out the floor, the compiler bug, and what the data can and can't support.
Looking ahead
I'm definitely going to try to force better double floating point measurements (there's quite a bit I can do), and I have a FIR application I've been benchmarking too.
I'm hoping to add more interesting microcontrollers to this fleet. Many of the current parts are high performing, but unusual. For example the RP2350 & ESP32 run mostly from external SPI flash via SRAM/cache compared to more common microcontrollers that run directly from flash (where wait states can have a bigger impact, etc.). I already have a few boards on deck, but if you have any suggestions on what to benchmark next, please let me know!
All product names and trademarks are the property of their respective owners.