Embedded & Electronics Toolkit
← All guides

Floating-Point Pitfalls on Microcontrollers

A 32-bit float looks like it can hold anything: its range runs to 3.4 × 1038. What it cannot do is hold many digits. It carries about seven significant decimal digits, wherever the decimal point happens to be, and most floating-point bugs in firmware come from forgetting that. This guide shows where the limit bites, with exact numbers, and one problem that is specific to microcontrollers: double-precision code that nobody asked for.

How a float is stored

A single-precision value has 1 sign bit, 8 exponent bits and 23 fraction bits. Together with an implied leading 1, that gives 24 bits of significand. A few encodings, as they appear in memory:

ValueEncodingNote
1.00x3F800000Exact
−2.50xC0200000Exact
0.10x3DCCCCCDActually 0.100000001490
+infinity0x7F800000Result of overflow or x / 0
NaN0x7FC00000Result of 0 / 0, sqrt(−1) and similar

The value 0.1 has no exact binary representation, in the same way that one third has no exact decimal one. The stored value is the nearest one available.

The spacing grows with the value

With 24 bits of significand, the gap between one representable value and the next depends on how large the value is:

AroundGap to the next float
10.00000012
1,0000.000061
1,000,0000.0625
16,777,2162
1,000,000,00064

Above 16,777,216, which is 224, a float cannot even represent every whole number. 16,777,217 is stored as 16,777,216.

Pitfall 1: time in a float

This is the most common consequence in firmware. Store a millisecond counter in a float and it counts correctly up to 16,777,216 ms, which is 4.66 hours. After that, adding 1 ms no longer changes the value: the increment is smaller than the gap between adjacent floats. The device works on the bench, works through a day of testing with resets, and fails in the field after an afternoon of uptime.

The same happens with seconds as a float once fractions matter. At 100,000 seconds, a little over a day, adding 0.001 changes nothing: the gap there is 0.0078 s.

Keep time in integers. A 32-bit millisecond counter runs for 49.7 days before wrapping, and unsigned subtraction of two timestamps gives the correct interval across the wrap. Convert to float only for the final calculation, and only the difference, never the absolute time.

Pitfall 2: accumulating a small step

Adding 0.1 ten times does not give 1. In single precision the result is 1.00000012, and a comparison with 1.0 fails. Continue for a million additions and the sum is 100,958.34 instead of 100,000, an error of nearly 1%. Each addition rounds, and as the total grows the rounding gets coarser.

Two rules follow:

Pitfall 3: the accidental double

Many microcontroller FPUs, including the one in the Cortex-M4F, handle single precision only. Double precision is done in software, by library routines. In C, a floating-point constant without a suffix is a double, and mixing it with a float promotes the whole expression. Compare two functions that differ by one letter:

float scale_d(float x) { return x * 3.14;  }
float scale_f(float x) { return x * 3.14f; }

Compiled for a Cortex-M4 with its FPU enabled, the second is a single multiply instruction, vmul.f32. The first converts the argument to double with a library call, multiplies with a second library call, and converts back with a third: __aeabi_f2d, __aeabi_dmul, __aeabi_d2f. The result is the same to seven digits. The cost is three function calls instead of one multiply instruction, plus the library code in flash.

The same promotion happens with sin(), sqrt() and the other functions from math.h, which take and return double. Use the versions with an f suffix: sinf(), sqrtf(), fabsf(). And printf with %f always promotes its argument to double, which is one reason floating-point printing is expensive in small firmware.

The compiler will find these for you. GCC and Clang both have -Wdouble-promotion, which reports every implicit conversion from float to double. Turn it on for any project that targets a single-precision FPU.

Pitfall 4: NaN spreads

Once a NaN appears, every arithmetic result that depends on it is also NaN, and every comparison with it is false, including x == x. A single division of zero by zero in a sensor path can turn a whole control loop's state into NaN, where it stays, since nothing can bring it back.

Guard the inputs that can produce one: check a divisor before dividing, clamp the argument of sqrtf() and acosf() into their valid range, and test results that feed long-lived state with isnan(). In a filter or integrator, resetting the state on NaN is much better than running with it.

Pitfall 5: subtracting nearly equal values

A float holds seven digits. Subtract two values that agree in the first five, and only two meaningful digits are left. This matters when a small change is measured on top of a large offset, for instance a temperature difference computed from two absolute readings in kelvin, or a position change from two large encoder totals converted to float. Subtract in integers first, while the values are still exact, and convert the difference.

Float, double or fixed point

SituationReasonable choice
Core with a single-precision FPU, ordinary sensor mathsfloat, with f suffixes throughout
Values needing more than seven digits: GPS coordinates, long timestamps, energy totalsIntegers, or double where its cost is acceptable
Core without an FPU, code in an interrupt or a fast loopFixed point
Counters, time, anything that must not driftIntegers

The GPS case deserves a remark. A latitude or longitude in degrees needs six or seven decimal places for sub-metre resolution, nine or ten significant digits in total. A float cannot hold that. Storing coordinates in a float quantises position to steps of the order of a metre or more, which shows up as a track that moves in jumps.

Run the numbers: IEEE 754 Floating-Point Visualizer does this calculation for your own values.

More guides