Fixed-Point Arithmetic in Q Format
On a microcontroller without a floating-point unit, every float operation is a library call. Fixed point does the same work with integer instructions: a fractional value is stored as an integer with an agreed scale factor. The idea is simple and the mistakes are few, but each of them produces results that look plausible and are wrong. This guide covers the notation, the four operations, and the places where it goes wrong.
The idea
A fixed-point number is an integer that everyone agrees to interpret as divided by a power of two. In Q15, the divisor is 215 = 32768:
real value = stored integer / 32768
stored value = round(real value × 32768)
The notation Qm.n means m integer bits and n fractional bits. Conventions differ on whether m includes the sign bit, so always state the total width as well. The common formats:
| Format | Stored in | Resolution | Range |
|---|---|---|---|
| Q15 (Q1.15) | int16_t | 1/32768 ≈ 0.0000305 | −1 to +0.99997 |
| Q1.14 | int16_t | 1/16384 ≈ 0.000061 | −2 to +1.99994 |
| Q8.8 | int16_t | 1/256 ≈ 0.0039 | −128 to +127.996 |
| Q16.16 | int32_t | 1/65536 ≈ 0.0000153 | −32768 to +32767.99998 |
Every extra fractional bit halves the step size and halves the range. Choosing a format is choosing where to put that trade.
Converting values
| Real value | Format | Stored integer | Value actually represented |
|---|---|---|---|
| 0.1 | Q15 | 3277 | 0.100006 |
| −0.25 | Q1.14 | −4096 (0xF000) | −0.25 exactly |
| 5.75 | Q8.8 | 1472 (0x05C0) | 5.75 exactly |
| 3.3 | Q8.8 | 845 | 3.30078 |
Values that are sums of powers of two, such as 0.25 and 5.75, are exact. Others are rounded to the nearest step. Do the conversion at compile time, in a macro or a constant expression, so that no floating-point code ends up in the firmware:
#define Q15(x) ((int16_t)((x) * 32768.0 + ((x) >= 0 ? 0.5 : -0.5)))
static const int16_t gain = Q15(0.1); /* 3277, computed by the compiler */
Addition and subtraction
Two numbers in the same format add and subtract as plain integers. The only rule is that the formats must match. Adding a Q15 value to a Q8.8 value without shifting one of them first is the fixed-point equivalent of adding metres to millimetres.
Addition can overflow: 0.75 + 0.5 does not fit in Q15. Either guarantee by design that the sum stays in range, or use saturating addition, which clamps to the largest value instead of wrapping around to a large negative one. Wrapping turns a slightly-too-large signal into a full-scale spike of the opposite sign, which in a control loop or audio path is far worse than clipping.
Multiplication
Multiplying two Q15 numbers gives a result scaled by 230, in a 32-bit intermediate. Shifting right by 15 brings it back to Q15:
int16_t q15_mul(int16_t a, int16_t b)
{
int32_t p = (int32_t)a * b; /* Q30, needs 32 bits */
p += 1 << 14; /* round to nearest */
return (int16_t)(p >> 15); /* back to Q15 */
}
Worked through for 0.3 × 0.7: the stored values are 9830 and 22938, the product is 225,480,540, and after rounding and shifting the result is 6881, which represents 0.20999. The exact answer is 0.21.
Three details in those four lines matter:
- The cast before the multiplication. The product of two 16-bit values needs 32 bits. On a 32-bit core C promotes the operands to
intanyway; on an 8-bit or 16-bit core, whereintis 16 bits, omitting the cast silently discards the upper half. Write the cast always. - The rounding constant. A plain shift truncates toward minus infinity, so every multiplication loses half a step on average, always in the same direction. In a filter that runs thousands of times per second this becomes a DC offset. Adding half of the divisor before the shift removes the bias.
- The one case that overflows. −1 × −1 is +1, and +1 is not representable in Q15. In integers, −32768 × −32768 shifted right by 15 is 32768, which wraps to −32768 when stored in an
int16_t. If −1 can occur at both inputs, saturate the result to 32767.
For general formats: a Qa × Qb product has a + b fractional bits. Shift right by however many you want to drop.
Division
Division needs the opposite correction. Pre-shift the numerator left, in a wider type, then divide:
int16_t q15_div(int16_t a, int16_t b) /* requires |a| < |b| */
{
return (int16_t)(((int32_t)a << 15) / b);
}
For 0.25 / 0.5 this gives 16384, which is 0.5. The result must fit the format, so in Q15 the numerator must be smaller in magnitude than the denominator. Division is also slow on cores without a hardware divider. Where the divisor is a constant, multiply by its reciprocal instead.
Accumulated error
A single rounded constant is accurate to half a step. Errors grow when a rounded value is used many times. The value 0.001 in Q15 is stored as 33, which actually represents 0.001007. Add it 1000 times and the total is 1.00708 instead of 1, an error of 0.7%.
The remedy is to keep more fractional bits in accumulators and intermediate results than in the inputs and outputs: accumulate in Q31 or Q16.16, and convert to the narrower format only at the end. This is the same principle as carrying extra digits through a hand calculation.
Scaling without naming a format
Much everyday fixed-point work is integer scaling with a little care about order. Converting a 12-bit ADC reading to millivolts with a 3.3 V reference:
uint32_t mv = ((uint32_t)adc * 3300) >> 12; /* 4095 → 3299 mV */
Multiply first, then divide, so that no precision is lost; and make sure the intermediate fits. Here the largest product is 4095 × 3300 = 13,513,500, comfortably inside 32 bits. In 16-bit arithmetic the same expression overflows for any reading above 19, which is why the cast is there.
Choosing a format
- Find the largest magnitude the value can take, including intermediate results. That sets the integer bits.
- Give every remaining bit to the fraction.
- Where possible, normalise signals to the range −1 to +1 and use Q15 or Q31. Products of such values stay in range, apart from the −1 × −1 case, which removes most overflow analysis.
- Put the format in the name or the type:
speed_q8_8,q15_t. A bareint16_tdoes not say what its bits mean, and mixed formats are the most common fixed-point bug.
When not to bother
If the core has a floating-point unit, single-precision float is usually as fast as fixed point and much harder to get wrong. Fixed point earns its keep on cores without an FPU, in interrupt handlers where saving FPU registers is unwelcome, and wherever results must be bit-identical between platforms.
Run the numbers: Fixed Point (Q-Format) Converter does this calculation for your own values.
More guides
- ADC Resolution Is Not Accuracy: LSB Size, Reference Error, Settling and Oversampling
- C Struct Padding on Cortex-M: Where the Bytes Go and How to Control Them
- CAN Bit Timing Step by Step: Choosing BRP, Segments and Sample Point
- Why Your CRC Does Not Match: Polynomial, Init, Reflection and XOR-Out Explained
- Floating-Point Pitfalls on Microcontrollers: Precision, Timestamps and Accidental Doubles
- Sizing I2C Pull-Up Resistors: Minimum, Maximum and What Bus Capacitance Does
- How Much UART Baud Rate Error Is Too Much? The Sampling Math