Thing Done

← The fleet

How the fleet is measured, and what the numbers can't tell you

Every number on the fleet pages is the wall-clock time of one C function call, divided by the number of arithmetic expressions it executed. This page is about what that number actually contains, why comparing it across 16 build targets is a best-effort exercise rather than a precise one, and which conclusions the data can and cannot carry. Some of these choices are ones a careful reader won't like. I'd rather you know about them than take a chart at face value.

What one number contains

Each timed expression is, in order:

  1. an indexed read of a volatile left operand,
  2. an indexed read of a volatile right operand,
  3. the arithmetic operation, and
  4. an indexed volatile write of the result,

inside a four-pair loop that is repeated eight times per set, inside a calibrated outer loop. Nothing is subtracted. The volatile accesses and all three levels of loop control are part of the measured time, on every target, by design: they are what stops the compiler from deleting or hoisting the arithmetic. The previous version of this benchmark relied on compiler barriers instead, and Microchip's XC8 collapsed 32 additions into one. That capture is retired, and nothing from it appears on these pages.

The practical consequence is a floor under every measurement. The fastest thing a board can report is the cost of two loads, a store and a loop step, whether or not the operation in between is free.

Where the floor dominates

On the 32-bit cores the floor is most of the number for add, subtract and multiply. Clock cycles per uint32_t expression, from the same data the charts use:

Coreaddsubtractmultiplydividemodulo
Cortex-M7 (i.MX RT1062) 7791517
Cortex-M33 (RP2350) 1111121618
Cortex-M33 (EFR32MG24) 1212131820
Cortex-M4F (SAMD51) 1212121819
Hazard3 RV32IMAC (RP2350) 1213133030
RV32IMAC (ESP32-C6) 1213132222
Cortex-M4F (STM32G474) 1313131920
Xtensa LX7 (ESP32-S3) 1818182223
Cortex-M0+ (RP2040) 1818173939
Xtensa LX6 (ESP32) 2020202425
dsPIC33A (dsPIC33AK512MPS506) 2020172730
PIC24F (PIC24FJ64GU205) 363650932946
MSP430FRx (MSP430FR2433) 434076692694
ATmega328P 4747119645645
PIC18F (PIC18F57Q84) 2672672,7262,6122,235
PIC16F (PIC16F17146) 3033603,2313,4672,960

A single-cycle ALU add on a Cortex-M4F is reported as 12 cycles. On the Xtensa LX7 (ESP32-S3), Xtensa LX6 (ESP32), Cortex-M4F (SAMD51) and Cortex-M4F (STM32G474) the add, subtract and multiply columns are identical to the nanosecond: the operation is invisible under the memory traffic. The ordering of those boards on the add, subtract and multiply charts reflects how each compiler and core handle an indexed volatile load-load-op-store loop, not how fast the arithmetic unit is. Treat small gaps between fast cores on those three operations as noise in the methodology, not as findings.

The numbers become informative where the operation is expensive enough to rise above the floor:

Ratios within one board (multiply versus add, uint64_t versus uint32_t, double versus float) are more robust than absolute comparisons across boards, because the floor cancels to first order.

Why the comparison is best-effort

The fleet spans 16 build targets on 15 boards, five compiler families and four decades of core design. Holding "everything else" equal is not possible. The experiment holds the C source equal and lets each platform do what it does. Known sources of non-comparability:

Every one of these is documented and consistent across the fleet, so the numbers are reproducible and honest about what they measure. They are not a ranking of instruction-set efficiency, and small differences between similar cores should not be read as one.

What the data supports

Safe claims:

Claims the data does not support:

The mechanics, briefly

How it got here

This is the third protocol. The first timed a C++ workload and subtracted a synthetic baseline, and was retired because what the subtraction removed was ambiguous. The second moved to explicit C kernels guarded by compiler barriers, and was retired when it turned out XC8 ignores those barriers and had kept one addition in every 32 on the PIC16F and PIC18F, which made two of the slowest boards in the fleet look like two of the fastest. The volatile contract described above is the third, and every board was re-measured under it.

A future revision that wants instruction-level numbers would need a register-resident kernel with hand-checked assembly per compiler, or a measured load-store-only baseline subtracted per board. I considered both and rejected them for this experiment: the first sacrifices the single shared C source, and the second reintroduces the baseline-subtraction ambiguity the first protocol was retired for.