How the fleet is measured, and what the numbers can't tell you
Every number on the fleet pages is the wall-clock time of one C function call, divided by the number of arithmetic expressions it executed. This page is about what that number actually contains, why comparing it across 16 build targets is a best-effort exercise rather than a precise one, and which conclusions the data can and cannot carry. Some of these choices are ones a careful reader won't like. I'd rather you know about them than take a chart at face value.
What one number contains
Each timed expression is, in order:
- an indexed read of a
volatileleft operand, - an indexed read of a
volatileright operand, - the arithmetic operation, and
- an indexed
volatilewrite of the result,
inside a four-pair loop that is repeated eight times per set, inside a calibrated outer loop. Nothing is subtracted. The volatile accesses and all three levels of loop control are part of the measured time, on every target, by design: they are what stops the compiler from deleting or hoisting the arithmetic. The previous version of this benchmark relied on compiler barriers instead, and Microchip's XC8 collapsed 32 additions into one. That capture is retired, and nothing from it appears on these pages.
The practical consequence is a floor under every measurement. The fastest thing a board can report is the cost of two loads, a store and a loop step, whether or not the operation in between is free.
Where the floor dominates
On the 32-bit cores the floor is most of the number for add, subtract and multiply. Clock cycles per uint32_t expression, from the same data the charts use:
| Core | add | subtract | multiply | divide | modulo |
|---|---|---|---|---|---|
| Cortex-M7 (i.MX RT1062) | 7 | 7 | 9 | 15 | 17 |
| Cortex-M33 (RP2350) | 11 | 11 | 12 | 16 | 18 |
| Cortex-M33 (EFR32MG24) | 12 | 12 | 13 | 18 | 20 |
| Cortex-M4F (SAMD51) | 12 | 12 | 12 | 18 | 19 |
| Hazard3 RV32IMAC (RP2350) | 12 | 13 | 13 | 30 | 30 |
| RV32IMAC (ESP32-C6) | 12 | 13 | 13 | 22 | 22 |
| Cortex-M4F (STM32G474) | 13 | 13 | 13 | 19 | 20 |
| Xtensa LX7 (ESP32-S3) | 18 | 18 | 18 | 22 | 23 |
| Cortex-M0+ (RP2040) | 18 | 18 | 17 | 39 | 39 |
| Xtensa LX6 (ESP32) | 20 | 20 | 20 | 24 | 25 |
| dsPIC33A (dsPIC33AK512MPS506) | 20 | 20 | 17 | 27 | 30 |
| PIC24F (PIC24FJ64GU205) | 36 | 36 | 50 | 932 | 946 |
| MSP430FRx (MSP430FR2433) | 43 | 40 | 76 | 692 | 694 |
| ATmega328P | 47 | 47 | 119 | 645 | 645 |
| PIC18F (PIC18F57Q84) | 267 | 267 | 2,726 | 2,612 | 2,235 |
| PIC16F (PIC16F17146) | 303 | 360 | 3,231 | 3,467 | 2,960 |
A single-cycle ALU add on a Cortex-M4F is reported as 12 cycles. On the Xtensa LX7 (ESP32-S3), Xtensa LX6 (ESP32), Cortex-M4F (SAMD51) and Cortex-M4F (STM32G474) the add, subtract and multiply columns are identical to the nanosecond: the operation is invisible under the memory traffic. The ordering of those boards on the add, subtract and multiply charts reflects how each compiler and core handle an indexed volatile load-load-op-store loop, not how fast the arithmetic unit is. Treat small gaps between fast cores on those three operations as noise in the methodology, not as findings.
The numbers become informative where the operation is expensive enough to rise above the floor:
- Division and remainder everywhere, and especially the contrast between the cores with a fast hardware divider and the ones that divide in software (the AVR, the MSP430 and the PICs) or through a slower path (the Cortex-M0+ and the Hazard3).
- 64-bit integer arithmetic on 8-, 16- and 32-bit cores.
- Floating point on cores without an FPU, or without a double-precision one: float on the RP2040, the Hazard3 and the ESP32-C6, and double on every target except the Teensy 4.0 and the RP2350.
- The 8-bit and 16-bit cores, where even the uint32_t add is dozens to hundreds of cycles because the operation itself is many instructions.
Ratios within one board (multiply versus add, uint64_t versus uint32_t, double versus float) are more robust than absolute comparisons across boards, because the floor cancels to first order.
Why the comparison is best-effort
The fleet spans 16 build targets on 15 boards, five compiler families and four decades of core design. Holding "everything else" equal is not possible. The experiment holds the C source equal and lets each platform do what it does. Known sources of non-comparability:
- Compiler and optimization. The Arduino targets build with
whatever optimization level their core applies by default. The raw targets (the
PICs, the dsPIC and the MSP430) build at
-O2under XC8, XC16, XC-DSC and MSP430-GCC, with two exceptions: the MSP430's double profile needs-Osand link-time optimization to fit binary64 plus libm into its 15 KB of FRAM, and the PIC18's float and double profiles use XC8's reentrant stack model because the default compiled stack stalled at floating-point division. Different compilers make different choices about register allocation, loop shape and how a volatile array index is materialized. Part of every gap between two boards is a gap between two compilers. - Clock normalization uses Fosc, not instruction rate. Cycles per operation divides by the final configured oscillator frequency. The PIC16 and PIC18 execute one instruction per four oscillator cycles and the PIC24 one per two, while the Arm, RISC-V and Xtensa cores execute roughly one per cycle. Counted in instruction cycles, the 8-bit PICs would show a quarter of the cycles listed here and the PIC24 half. The choice is deliberate and consistent, but it is a choice.
-
doubleis not one type. On the ATmega328P, dsPIC33A (dsPIC33AK512MPS506), PIC24F (PIC24FJ64GU205), PIC16F (PIC16F17146) and PIC18F (PIC18F57Q84),doubleis 4 bytes and the double profile measures the same single-precision path as float; the two rows differ only by run-to-run noise between two firmware images. The board pages label the width. The charts do not visually distinguish it. - Integer promotion. uint8_t and uint16_t expressions are computed
at
intwidth, which is 16 bits on the AVR, the MSP430, the PIC16, the PIC18 and the PIC24, and 32 bits elsewhere. Narrow multiply uses an explicit unsigned promotion to avoid signed overflow, which is one more thing the compilers implement differently. -
fmodmay be a library call. Its cost is the C library's, not the core's, and the libraries differ. - Runtime environment. Arduino sketches run with the core's timer interrupts and background services live. Raw targets run bare metal. The three-sample range exposes jitter but cannot attribute it.
- Memory placement. Code executes from XIP flash with a cache on the ESP32 family and the RP2040/RP2350, from tightly coupled memory on the Teensy, from FRAM on the MSP430, and from program flash on the rest. The volatile operands live in SRAM everywhere, but the instruction fetch path differs.
- Fixed operands. Four operand pairs per type, chosen to cover small, maximal, high-bit and alternating-bit values. That is one representative mix, not a distribution, and division and remainder timings on cores with data-dependent dividers depend on which values were chosen.
Every one of these is documented and consistent across the fleet, so the numbers are reproducible and honest about what they measure. They are not a ranking of instruction-set efficiency, and small differences between similar cores should not be read as one.
What the data supports
Safe claims:
- Order-of-magnitude comparisons between tiers: an 8-bit PIC versus a Cortex-M0+ versus a Cortex-M7.
- Within-board ratios between operations and between types.
- Presence or absence of hardware support: a divider, an FPU, a double-precision FPU, a hardware multiplier.
- The cost of software 64-bit and soft-float arithmetic on a given core.
Claims the data does not support:
- Cycle counts for individual instructions.
- Fine ordering among the fast 32-bit cores on add, subtract or multiply.
- Any comparison with the two retired protocols, which were different workloads.
- Worst-case or guaranteed timing.
The mechanics, briefly
- Each numeric type is a separate firmware image. Before any arithmetic runs, the firmware's own microsecond counter is checked against the host clock over three intervals from 50 ms to 5 s (100 ms to 10 s on the MSP430), comparing every pair of intervals so a constant USB latency cancels out. A clock more than 2% off fails the gate and the job is rejected, not corrected.
- For each operation the firmware calibrates a set count that takes about 500 ms, then times three independent calls of that size. The median is the number on the page; all three samples are in the chart tooltips. No baseline is subtracted.
- The uint32_t add samples are additionally bracketed by serial markers so the host can check the firmware's elapsed time against its own clock. The comparison is recorded, not used to adjust the value.
- Every result is checked: the kernel validates three of its four pairs in its epilogue and the host checks the fourth after timing. A wrong answer rejects the run.
- Ten cells are excluded rather than measured: uint64_t on the PIC16F and PIC18F, because XC8 has no 64-bit integer type.
How it got here
This is the third protocol. The first timed a C++ workload and subtracted a synthetic baseline, and was retired because what the subtraction removed was ambiguous. The second moved to explicit C kernels guarded by compiler barriers, and was retired when it turned out XC8 ignores those barriers and had kept one addition in every 32 on the PIC16F and PIC18F, which made two of the slowest boards in the fleet look like two of the fastest. The volatile contract described above is the third, and every board was re-measured under it.
A future revision that wants instruction-level numbers would need a register-resident kernel with hand-checked assembly per compiler, or a measured load-store-only baseline subtracted per board. I considered both and rejected them for this experiment: the first sacrifices the single shared C source, and the second reintroduces the baseline-subtraction ambiguity the first protocol was retired for.