Embedded & Electronics Toolkit
← All guides

C Struct Padding on Cortex-M

A struct with seven bytes of data that takes twelve bytes of RAM is not a compiler bug. It follows three rules, and once you know them you can predict every layout, remove padding without any pragma, and see what packed costs. All sizes and offsets below were checked by compiling for Cortex-M0 and Cortex-M3 with Arm Compiler 6 (armclang). Both cores give the same layouts; they differ only in the code generated for packed access.

The three rules

  1. Each member starts at a multiple of its own alignment. On Cortex-M, uint8_t aligns to 1, uint16_t to 2, uint32_t, float and pointers to 4. The compiler inserts padding before a member when the previous one did not end on a suitable boundary.
  2. The struct's alignment is that of its most strictly aligned member.
  3. The struct's size is rounded up to a multiple of its alignment. This tail padding keeps every element of an array of the struct aligned.

A first example

struct sample {
    uint8_t  flags;   /* offset 0            */
                      /* 3 bytes of padding  */
    uint32_t ts;      /* offset 4            */
    uint16_t id;      /* offset 8            */
                      /* 2 bytes tail padding */
};                    /* sizeof = 12, alignment = 4 */

Rule 1 pushes ts to offset 4. Rule 3 rounds the size from 10 up to 12. Seven bytes of data, five bytes of padding.

Now declare the same members from largest to smallest:

struct sample {
    uint32_t ts;      /* offset 0 */
    uint16_t id;      /* offset 4 */
    uint8_t  flags;   /* offset 6 */
                      /* 1 byte tail padding */
};                    /* sizeof = 8 */

Every member lands on a valid boundary without help, and only one byte of tail padding remains. Nothing else changed: no pragma, no attribute, no slower code. For an array of 1,000 records this is 4 kB of RAM saved by reordering three lines.

A slightly larger case shows the same thing:

Declaration orderOffsetssizeof
u8 a; u16 b; u8 c; u32 d; u8 e;0, 2, 4, 8, 1216
u32 d; u16 b; u8 a; u8 c; u8 e;0, 4, 6, 7, 812

The data is nine bytes in both. Sorting by size cannot remove the tail padding, since the size must still be a multiple of 4, but it removes everything in between.

The surprise: 64-bit members align to 8

It is easy to assume that nothing on a 32-bit core aligns to more than 4 bytes. On Cortex-M that is wrong. The ARM procedure call standard gives uint64_t, int64_t and double an alignment of 8:

StructOffsetssizeof
uint8_t tag; uint64_t t;0, 816
uint8_t tag; double v; uint8_t end;0, 8, 1624

The second struct holds ten bytes of data in 24 bytes. One double or one 64-bit timestamp is enough to raise the alignment of the whole struct to 8, and with it the tail padding of every struct that contains it.

This is a property of the ABI, not of the processor. Other 32-bit platforms use different rules; 32-bit x86 Linux, for example, aligns a double inside a struct to 4. Code that shares a binary layout between a Cortex-M device and a PC tool must not assume the two agree.

Nested structs and arrays

A nested struct keeps its own alignment and its own tail padding. Take the reordered 8-byte struct sample above and put it inside another struct, followed by one uint8_t. The result is 12 bytes, not 9: the inner struct occupies its full 8 bytes, and the outer one inherits its 4-byte alignment and is rounded up. The spare byte at the end of the inner struct cannot be reused by the outer one.

What packed really does

__attribute__((packed)) sets the alignment of every member to 1. The first example becomes:

struct __attribute__((packed)) sample {
    uint8_t  flags;   /* offset 0 */
    uint32_t ts;      /* offset 1 */
    uint16_t id;      /* offset 5 */
};                    /* sizeof = 7, alignment = 1 */

The layout now matches the data exactly, but ts sits at an odd address. What that costs depends on the core. This is the code generated for a function that returns p->ts:

CoreGenerated code
Cortex-M3One instruction: ldr.w r0, [r0, #1]. The core supports unaligned word loads.
Cortex-M0Four byte loads (ldrb) combined with shifts and adds, ten instructions in total. The core cannot load a word from an unaligned address, so the compiler assembles it byte by byte.

So on Cortex-M0 and M0+ every access to a misaligned packed member is several times larger and slower, and on any core the access is no longer a single atomic load. That matters when an interrupt handler updates the same field.

The pointer trap

The compiler generates the byte-by-byte code only when it knows the member is packed. Take the address of the member and that knowledge is lost:

uint32_t *p = &s.ts;   /* compilers warn about this */
uint32_t v = *p;        /* compiled as an ordinary aligned load */

On Cortex-M3 and above this usually still works. On Cortex-M0 and M0+ an unaligned word access raises a HardFault. The safe pattern is to copy the value out with memcpy, or to read the member through the struct and never through a pointer to it.

Structs as packets and file formats

Sending a struct over a wire by pointing a transmit function at it is common and fragile. Three things can differ at the other end: padding, byte order, and the size of types such as int or enum. Padding bytes also carry whatever was in memory before, so an unpacked struct sent as-is leaks stack or heap contents, and comparing two such structs with memcmp can report a difference when every member is equal.

Two approaches hold up:

A third option avoids packing altogether: order the fields so that the natural layout has no internal padding, and add explicit reserved bytes where a gap is unavoidable. The struct is then both wire-exact and aligned.

Make the compiler check it

Whichever approach you use, turn your assumptions into compile-time checks. If someone adds a field or changes a type, the build fails instead of the protocol:

#include <stddef.h>
#include <stdint.h>

_Static_assert(sizeof(struct sample) == 8,          "sample: size");
_Static_assert(offsetof(struct sample, id) == 4,    "sample: id offset");
_Static_assert(offsetof(struct sample, flags) == 6, "sample: flags offset");

These cost nothing at run time, and they are the only layout documentation that cannot go out of date.

Summary

Run the numbers: Struct Padding Calculator does this calculation for your own values.

More guides