Floating-Point Pitfalls on Microcontrollers
A 32-bit float looks like it can hold anything: its range runs to 3.4 × 1038. What it cannot do is hold many digits. It carries about seven significant decimal digits, wherever the decimal point happens to be, and most floating-point bugs in firmware come from forgetting that. This guide shows where the limit bites, with exact numbers, and one problem that is specific to microcontrollers: double-precision code that nobody asked for.
How a float is stored
A single-precision value has 1 sign bit, 8 exponent bits and 23 fraction bits. Together with an implied leading 1, that gives 24 bits of significand. A few encodings, as they appear in memory:
| Value | Encoding | Note |
|---|---|---|
| 1.0 | 0x3F800000 | Exact |
| −2.5 | 0xC0200000 | Exact |
| 0.1 | 0x3DCCCCCD | Actually 0.100000001490 |
| +infinity | 0x7F800000 | Result of overflow or x / 0 |
| NaN | 0x7FC00000 | Result of 0 / 0, sqrt(−1) and similar |
The value 0.1 has no exact binary representation, in the same way that one third has no exact decimal one. The stored value is the nearest one available.
The spacing grows with the value
With 24 bits of significand, the gap between one representable value and the next depends on how large the value is:
| Around | Gap to the next float |
|---|---|
| 1 | 0.00000012 |
| 1,000 | 0.000061 |
| 1,000,000 | 0.0625 |
| 16,777,216 | 2 |
| 1,000,000,000 | 64 |
Above 16,777,216, which is 224, a float cannot even represent every whole number. 16,777,217 is stored as 16,777,216.
Pitfall 1: time in a float
This is the most common consequence in firmware. Store a millisecond counter in a float and it counts correctly up to 16,777,216 ms, which is 4.66 hours. After that, adding 1 ms no longer changes the value: the increment is smaller than the gap between adjacent floats. The device works on the bench, works through a day of testing with resets, and fails in the field after an afternoon of uptime.
The same happens with seconds as a float once fractions matter. At 100,000 seconds, a little over a day, adding 0.001 changes nothing: the gap there is 0.0078 s.
Keep time in integers. A 32-bit millisecond counter runs for 49.7 days before wrapping, and unsigned subtraction of two timestamps gives the correct interval across the wrap. Convert to float only for the final calculation, and only the difference, never the absolute time.
Pitfall 2: accumulating a small step
Adding 0.1 ten times does not give 1. In single precision the result is 1.00000012, and a comparison with 1.0 fails. Continue for a million additions and the sum is 100,958.34 instead of 100,000, an error of nearly 1%. Each addition rounds, and as the total grows the rounding gets coarser.
Two rules follow:
- Count in integers and scale once. Instead of
t += 0.1fevery tick, keep an integer tick count and computet = ticks * 0.1fwhen needed. The error is then one rounding, not a million. - Never compare floats for equality after arithmetic. Compare against a tolerance chosen for the application, such as
fabsf(a - b) < 0.001f, or restructure the test as greater-or-equal so that overshooting still terminates a loop.
Pitfall 3: the accidental double
Many microcontroller FPUs, including the one in the Cortex-M4F, handle single precision only. Double precision is done in software, by library routines. In C, a floating-point constant without a suffix is a double, and mixing it with a float promotes the whole expression. Compare two functions that differ by one letter:
float scale_d(float x) { return x * 3.14; }
float scale_f(float x) { return x * 3.14f; }
Compiled for a Cortex-M4 with its FPU enabled, the second is a single multiply instruction, vmul.f32. The first converts the argument to double with a library call, multiplies with a second library call, and converts back with a third: __aeabi_f2d, __aeabi_dmul, __aeabi_d2f. The result is the same to seven digits. The cost is three function calls instead of one multiply instruction, plus the library code in flash.
The same promotion happens with sin(), sqrt() and the other functions from math.h, which take and return double. Use the versions with an f suffix: sinf(), sqrtf(), fabsf(). And printf with %f always promotes its argument to double, which is one reason floating-point printing is expensive in small firmware.
The compiler will find these for you. GCC and Clang both have -Wdouble-promotion, which reports every implicit conversion from float to double. Turn it on for any project that targets a single-precision FPU.
Pitfall 4: NaN spreads
Once a NaN appears, every arithmetic result that depends on it is also NaN, and every comparison with it is false, including x == x. A single division of zero by zero in a sensor path can turn a whole control loop's state into NaN, where it stays, since nothing can bring it back.
Guard the inputs that can produce one: check a divisor before dividing, clamp the argument of sqrtf() and acosf() into their valid range, and test results that feed long-lived state with isnan(). In a filter or integrator, resetting the state on NaN is much better than running with it.
Pitfall 5: subtracting nearly equal values
A float holds seven digits. Subtract two values that agree in the first five, and only two meaningful digits are left. This matters when a small change is measured on top of a large offset, for instance a temperature difference computed from two absolute readings in kelvin, or a position change from two large encoder totals converted to float. Subtract in integers first, while the values are still exact, and convert the difference.
Float, double or fixed point
| Situation | Reasonable choice |
|---|---|
| Core with a single-precision FPU, ordinary sensor maths | float, with f suffixes throughout |
| Values needing more than seven digits: GPS coordinates, long timestamps, energy totals | Integers, or double where its cost is acceptable |
| Core without an FPU, code in an interrupt or a fast loop | Fixed point |
| Counters, time, anything that must not drift | Integers |
The GPS case deserves a remark. A latitude or longitude in degrees needs six or seven decimal places for sub-metre resolution, nine or ten significant digits in total. A float cannot hold that. Storing coordinates in a float quantises position to steps of the order of a metre or more, which shows up as a track that moves in jumps.
Run the numbers: IEEE 754 Floating-Point Visualizer does this calculation for your own values.
More guides
- ADC Resolution Is Not Accuracy: LSB Size, Reference Error, Settling and Oversampling
- C Struct Padding on Cortex-M: Where the Bytes Go and How to Control Them
- CAN Bit Timing Step by Step: Choosing BRP, Segments and Sample Point
- Why Your CRC Does Not Match: Polynomial, Init, Reflection and XOR-Out Explained
- Fixed-Point Arithmetic in Q Format: Scaling, Multiplying and Not Overflowing
- Sizing I2C Pull-Up Resistors: Minimum, Maximum and What Bus Capacitance Does
- How Much UART Baud Rate Error Is Too Much? The Sampling Math