Arithmetic on 16 microcontrollers, measured the same way
Five operations, six numeric types, the same C kernels, on 16 build targets across 6 ISA families — with every clock verified against the host before a number was kept. 470 measured cells. Pick an operation and a type, hover for the numbers, click a bar for that board's full profile — or compare any two boards head to head. The fleet grows as boards get added.
The story behind these numbers: The chip wasn't slow — my benchmark was lying. Twice..
The ranking
Table view — uint32_t add, operations per second
| # | Core | Board | Clock | ops/s | ns/op | cycles/op |
|---|---|---|---|---|---|---|
| 1 | Cortex-M7 (i.MX RT1062) | Teensy 4.0 | 600 MHz | 82.8 M/s | 12.1 ns | 7.25 |
| 2 | Xtensa LX7 (ESP32-S3) | ESP32-S3 DevKit | 240 MHz | 13.4 M/s | 74.6 ns | 17.9 |
| 3 | Cortex-M33 (RP2350) | Pi Pico 2W | 150 MHz | 13.3 M/s | 75.3 ns | 11.3 |
| 4 | RV32IMAC (ESP32-C6) | ESP32-C6-DevKitM-1 | 160 MHz | 12.8 M/s | 77.9 ns | 12.5 |
| 5 | Cortex-M4F (STM32G474) | Nucleo G474RE | 170 MHz | 12.6 M/s | 79.3 ns | 13.5 |
| 6 | Hazard3 RV32IMAC (RP2350) | Pi Pico 2W (RISC-V) | 150 MHz | 12.2 M/s | 81.8 ns | 12.3 |
| 7 | Xtensa LX6 (ESP32) | ESP32 | 240 MHz | 12 M/s | 83.3 ns | 20.0 |
| 8 | Cortex-M0+ (RP2040) | Pi Pico | 200 MHz | 11.3 M/s | 88.9 ns | 17.8 |
| 9 | dsPIC33A (dsPIC33AK512MPS506) | dsPIC33AK512MPS506 Curiosity Nano | 200 MHz | 10.2 M/s | 98 ns | 19.6 |
| 10 | Cortex-M4F (SAMD51) | Feather M4 Express | 120 MHz | 9.71 M/s | 103 ns | 12.4 |
| 11 | Cortex-M33 (EFR32MG24) | XIAO MG24 Sense | 39 MHz | 3.19 M/s | 313 ns | 12.2 |
| 12 | PIC24F (PIC24FJ64GU205) | PIC24FJ64GU205 Curiosity Nano | 32 MHz | 890 k/s | 1.12 µs | 35.9 |
| 13 | MSP430FRx (MSP430FR2433) | MSP-EXP430FR2433 LaunchPad | 15.99 MHz | 370 k/s | 2.7 µs | 43.2 |
| 14 | ATmega328P | Arduino Uno | 16 MHz | 342 k/s | 2.92 µs | 46.7 |
| 15 | PIC18F (PIC18F57Q84) | PIC18F57Q84 Curiosity Nano | 64 MHz | 240 k/s | 4.17 µs | 267 |
| 16 | PIC16F (PIC16F17146) | PIC16F17146 Curiosity Nano | 32 MHz | 105 k/s | 9.48 µs | 303 |
Compare two boards
Every pair has its own page you can link to.
A few worth a look:
- Teensy 4.0 vs Arduino Uno — the fastest board vs the one everyone starts on
- Pi Pico 2W vs Pi Pico 2W (RISC-V) — the same silicon, Arm vs RISC-V toolchain
- dsPIC33AK512 vs Nucleo G474RE — a 200 MHz dsPIC against a 170 MHz Cortex-M4F
- ESP32-S3 vs ESP32-C6 — Xtensa vs RISC-V in the same family
- PIC18F57Q84 vs Arduino Uno — two 8-bit parts
- MSP430FR2433 vs PIC16F17146 — the two smallest
The boards
Labeled core-first, because that's what the charts compare. The clock is the final configured oscillator rate as the firmware itself reports it. The two Pico 2W rows are one physical board built with two toolchains.
| Family | Core | Board | Clock | |
|---|---|---|---|---|
| Xtensa | Xtensa LX7 (ESP32-S3) | ESP32-S3 DevKit | 240 MHz | vs Teensy |
| Xtensa | Xtensa LX6 (ESP32) | ESP32 | 240 MHz | vs Teensy |
| AVR | ATmega328P | Arduino Uno | 16 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M33 (EFR32MG24) | XIAO MG24 Sense | 39 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M4F (SAMD51) | Feather M4 Express | 120 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M0+ (RP2040) | Pi Pico | 200 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M33 (RP2350) | Pi Pico 2W | 150 MHz | vs Teensy |
| RISC-V | Hazard3 RV32IMAC (RP2350) | Pi Pico 2W (RISC-V) | 150 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M4F (STM32G474) | Nucleo G474RE | 170 MHz | vs Teensy |
| RISC-V | RV32IMAC (ESP32-C6) | ESP32-C6-DevKitM-1 | 160 MHz | vs Teensy |
| Arm Cortex-M | Cortex-M7 (i.MX RT1062) | Teensy 4.0 | 600 MHz | vs Uno |
| PIC / dsPIC | dsPIC33A (dsPIC33AK512MPS506) | dsPIC33AK512MPS506 Curiosity Nano | 200 MHz | vs Teensy |
| PIC / dsPIC | PIC24F (PIC24FJ64GU205) | PIC24FJ64GU205 Curiosity Nano | 32 MHz | vs Teensy |
| MSP430 | MSP430FRx (MSP430FR2433) | MSP-EXP430FR2433 LaunchPad | 15.99 MHz | vs Teensy |
| PIC / dsPIC | PIC16F (PIC16F17146) | PIC16F17146 Curiosity Nano | 32 MHz | vs Teensy |
| PIC / dsPIC | PIC18F (PIC18F57Q84) | PIC18F57Q84 Curiosity Nano | 64 MHz | vs Teensy |
How it's measured
Each board runs the same five operations — add, subtract, multiply, divide, modulo — on uint8_t, uint16_t, uint32_t, uint64_t, float and double, from the same plain-C kernels. Every expression reads its two operands from volatile memory and writes its result back to volatile memory, looping over four operand pairs (small, maximum, high-bit and alternating-bit values) that the compiler has to work through for real. That contract is identical on every target, and it was learned the hard way: the first capture leaned on compiler barriers, and Microchip's XC8 ignored them — it kept one addition in every 32, so the PIC16F and PIC18F add, subtract and multiply rates came out as much as a hundred times too high. That whole capture was retired and every board was re-measured under the volatile contract (protocol v3, September 2026).
The price is a floor under every number: two loads, a store and a loop step ride along with each operation, so even the fastest 32-bit cores report somewhere between 7 and 20 clock cycles for an add that is a single instruction. Treat small gaps between fast cores on add, subtract and multiply as methodology noise. The numbers get informative where the operation rises above that floor — division and modulo, 64-bit integers, floating point on cores without an FPU (or without a double-precision one), and the 8- and 16-bit cores, where even a uint32_t add is dozens to hundreds of cycles. Ratios within one board (multiply vs add, uint64_t vs uint32_t, double vs float) are more robust than absolute comparisons across boards, because the floor cancels. The methodology page spells out everything else that rides along inside a number, and which comparisons the data can't support.
Every cell is the median of three ~500 ms samples. Before a run counts, the board's clock is calibrated against the host so a mis-configured clock tree can't quietly poison the numbers (it did, once). Clock cycles per operation is the time multiplied by that verified clock: the fairest single number for "how good is this core", independent of how fast it's clocked. Ten cells are missing: uint64_t on the PIC16F and PIC18F, where the XC8 compiler has no 64-bit integer type.
Want to run something on the fleet?
The rig is set up for remote access: submit a firmware build and a test script, and it flashes the real board and runs your script against its serial port. I built it for my own experiments, but there's no reason it can't run yours — same-source comparisons across the fleet, regression tests on hardware you don't own, or something I haven't thought of. That last category is the one I'm most curious about.
I'm opening it up to a few people. If you have something you'd run on it, email me a couple of sentences about what you'd do — that's the whole application.