Embedded & Electronics Toolkit
← All guides

Fixed-Point Arithmetic in Q Format

On a microcontroller without a floating-point unit, every float operation is a library call. Fixed point does the same work with integer instructions: a fractional value is stored as an integer with an agreed scale factor. The idea is simple and the mistakes are few, but each of them produces results that look plausible and are wrong. This guide covers the notation, the four operations, and the places where it goes wrong.

The idea

A fixed-point number is an integer that everyone agrees to interpret as divided by a power of two. In Q15, the divisor is 215 = 32768:

real value   = stored integer / 32768
stored value = round(real value × 32768)

The notation Qm.n means m integer bits and n fractional bits. Conventions differ on whether m includes the sign bit, so always state the total width as well. The common formats:

FormatStored inResolutionRange
Q15 (Q1.15)int16_t1/32768 ≈ 0.0000305−1 to +0.99997
Q1.14int16_t1/16384 ≈ 0.000061−2 to +1.99994
Q8.8int16_t1/256 ≈ 0.0039−128 to +127.996
Q16.16int32_t1/65536 ≈ 0.0000153−32768 to +32767.99998

Every extra fractional bit halves the step size and halves the range. Choosing a format is choosing where to put that trade.

Converting values

Real valueFormatStored integerValue actually represented
0.1Q1532770.100006
−0.25Q1.14−4096 (0xF000)−0.25 exactly
5.75Q8.81472 (0x05C0)5.75 exactly
3.3Q8.88453.30078

Values that are sums of powers of two, such as 0.25 and 5.75, are exact. Others are rounded to the nearest step. Do the conversion at compile time, in a macro or a constant expression, so that no floating-point code ends up in the firmware:

#define Q15(x)  ((int16_t)((x) * 32768.0 + ((x) >= 0 ? 0.5 : -0.5)))

static const int16_t gain = Q15(0.1);   /* 3277, computed by the compiler */

Addition and subtraction

Two numbers in the same format add and subtract as plain integers. The only rule is that the formats must match. Adding a Q15 value to a Q8.8 value without shifting one of them first is the fixed-point equivalent of adding metres to millimetres.

Addition can overflow: 0.75 + 0.5 does not fit in Q15. Either guarantee by design that the sum stays in range, or use saturating addition, which clamps to the largest value instead of wrapping around to a large negative one. Wrapping turns a slightly-too-large signal into a full-scale spike of the opposite sign, which in a control loop or audio path is far worse than clipping.

Multiplication

Multiplying two Q15 numbers gives a result scaled by 230, in a 32-bit intermediate. Shifting right by 15 brings it back to Q15:

int16_t q15_mul(int16_t a, int16_t b)
{
    int32_t p = (int32_t)a * b;        /* Q30, needs 32 bits   */
    p += 1 << 14;                      /* round to nearest     */
    return (int16_t)(p >> 15);         /* back to Q15          */
}

Worked through for 0.3 × 0.7: the stored values are 9830 and 22938, the product is 225,480,540, and after rounding and shifting the result is 6881, which represents 0.20999. The exact answer is 0.21.

Three details in those four lines matter:

For general formats: a Qa × Qb product has a + b fractional bits. Shift right by however many you want to drop.

Division

Division needs the opposite correction. Pre-shift the numerator left, in a wider type, then divide:

int16_t q15_div(int16_t a, int16_t b)   /* requires |a| < |b| */
{
    return (int16_t)(((int32_t)a << 15) / b);
}

For 0.25 / 0.5 this gives 16384, which is 0.5. The result must fit the format, so in Q15 the numerator must be smaller in magnitude than the denominator. Division is also slow on cores without a hardware divider. Where the divisor is a constant, multiply by its reciprocal instead.

Accumulated error

A single rounded constant is accurate to half a step. Errors grow when a rounded value is used many times. The value 0.001 in Q15 is stored as 33, which actually represents 0.001007. Add it 1000 times and the total is 1.00708 instead of 1, an error of 0.7%.

The remedy is to keep more fractional bits in accumulators and intermediate results than in the inputs and outputs: accumulate in Q31 or Q16.16, and convert to the narrower format only at the end. This is the same principle as carrying extra digits through a hand calculation.

Scaling without naming a format

Much everyday fixed-point work is integer scaling with a little care about order. Converting a 12-bit ADC reading to millivolts with a 3.3 V reference:

uint32_t mv = ((uint32_t)adc * 3300) >> 12;   /* 4095 → 3299 mV */

Multiply first, then divide, so that no precision is lost; and make sure the intermediate fits. Here the largest product is 4095 × 3300 = 13,513,500, comfortably inside 32 bits. In 16-bit arithmetic the same expression overflows for any reading above 19, which is why the cast is there.

Choosing a format

  1. Find the largest magnitude the value can take, including intermediate results. That sets the integer bits.
  2. Give every remaining bit to the fraction.
  3. Where possible, normalise signals to the range −1 to +1 and use Q15 or Q31. Products of such values stay in range, apart from the −1 × −1 case, which removes most overflow analysis.
  4. Put the format in the name or the type: speed_q8_8, q15_t. A bare int16_t does not say what its bits mean, and mixed formats are the most common fixed-point bug.

When not to bother

If the core has a floating-point unit, single-precision float is usually as fast as fixed point and much harder to get wrong. Fixed point earns its keep on cores without an FPU, in interrupt handlers where saving FPU registers is unwelcome, and wherever results must be bit-identical between platforms.

Run the numbers: Fixed Point (Q-Format) Converter does this calculation for your own values.

More guides