C Struct Padding on Cortex-M
A struct with seven bytes of data that takes twelve bytes of RAM is not a compiler bug. It follows three rules, and once you know them you can predict every layout, remove padding without any pragma, and see what packed costs. All sizes and offsets below were checked by compiling for Cortex-M0 and Cortex-M3 with Arm Compiler 6 (armclang). Both cores give the same layouts; they differ only in the code generated for packed access.
The three rules
- Each member starts at a multiple of its own alignment. On Cortex-M,
uint8_taligns to 1,uint16_tto 2,uint32_t,floatand pointers to 4. The compiler inserts padding before a member when the previous one did not end on a suitable boundary. - The struct's alignment is that of its most strictly aligned member.
- The struct's size is rounded up to a multiple of its alignment. This tail padding keeps every element of an array of the struct aligned.
A first example
struct sample {
uint8_t flags; /* offset 0 */
/* 3 bytes of padding */
uint32_t ts; /* offset 4 */
uint16_t id; /* offset 8 */
/* 2 bytes tail padding */
}; /* sizeof = 12, alignment = 4 */
Rule 1 pushes ts to offset 4. Rule 3 rounds the size from 10 up to 12. Seven bytes of data, five bytes of padding.
Now declare the same members from largest to smallest:
struct sample {
uint32_t ts; /* offset 0 */
uint16_t id; /* offset 4 */
uint8_t flags; /* offset 6 */
/* 1 byte tail padding */
}; /* sizeof = 8 */
Every member lands on a valid boundary without help, and only one byte of tail padding remains. Nothing else changed: no pragma, no attribute, no slower code. For an array of 1,000 records this is 4 kB of RAM saved by reordering three lines.
A slightly larger case shows the same thing:
| Declaration order | Offsets | sizeof |
|---|---|---|
u8 a; u16 b; u8 c; u32 d; u8 e; | 0, 2, 4, 8, 12 | 16 |
u32 d; u16 b; u8 a; u8 c; u8 e; | 0, 4, 6, 7, 8 | 12 |
The data is nine bytes in both. Sorting by size cannot remove the tail padding, since the size must still be a multiple of 4, but it removes everything in between.
The surprise: 64-bit members align to 8
It is easy to assume that nothing on a 32-bit core aligns to more than 4 bytes. On Cortex-M that is wrong. The ARM procedure call standard gives uint64_t, int64_t and double an alignment of 8:
| Struct | Offsets | sizeof |
|---|---|---|
uint8_t tag; uint64_t t; | 0, 8 | 16 |
uint8_t tag; double v; uint8_t end; | 0, 8, 16 | 24 |
The second struct holds ten bytes of data in 24 bytes. One double or one 64-bit timestamp is enough to raise the alignment of the whole struct to 8, and with it the tail padding of every struct that contains it.
This is a property of the ABI, not of the processor. Other 32-bit platforms use different rules; 32-bit x86 Linux, for example, aligns a double inside a struct to 4. Code that shares a binary layout between a Cortex-M device and a PC tool must not assume the two agree.
Nested structs and arrays
A nested struct keeps its own alignment and its own tail padding. Take the reordered 8-byte struct sample above and put it inside another struct, followed by one uint8_t. The result is 12 bytes, not 9: the inner struct occupies its full 8 bytes, and the outer one inherits its 4-byte alignment and is rounded up. The spare byte at the end of the inner struct cannot be reused by the outer one.
What packed really does
__attribute__((packed)) sets the alignment of every member to 1. The first example becomes:
struct __attribute__((packed)) sample {
uint8_t flags; /* offset 0 */
uint32_t ts; /* offset 1 */
uint16_t id; /* offset 5 */
}; /* sizeof = 7, alignment = 1 */
The layout now matches the data exactly, but ts sits at an odd address. What that costs depends on the core. This is the code generated for a function that returns p->ts:
| Core | Generated code |
|---|---|
| Cortex-M3 | One instruction: ldr.w r0, [r0, #1]. The core supports unaligned word loads. |
| Cortex-M0 | Four byte loads (ldrb) combined with shifts and adds, ten instructions in total. The core cannot load a word from an unaligned address, so the compiler assembles it byte by byte. |
So on Cortex-M0 and M0+ every access to a misaligned packed member is several times larger and slower, and on any core the access is no longer a single atomic load. That matters when an interrupt handler updates the same field.
The pointer trap
The compiler generates the byte-by-byte code only when it knows the member is packed. Take the address of the member and that knowledge is lost:
uint32_t *p = &s.ts; /* compilers warn about this */
uint32_t v = *p; /* compiled as an ordinary aligned load */
On Cortex-M3 and above this usually still works. On Cortex-M0 and M0+ an unaligned word access raises a HardFault. The safe pattern is to copy the value out with memcpy, or to read the member through the struct and never through a pointer to it.
Structs as packets and file formats
Sending a struct over a wire by pointing a transmit function at it is common and fragile. Three things can differ at the other end: padding, byte order, and the size of types such as int or enum. Padding bytes also carry whatever was in memory before, so an unpacked struct sent as-is leaks stack or heap contents, and comparing two such structs with memcmp can report a difference when every member is equal.
Two approaches hold up:
- Serialise explicitly. Write each field into a byte buffer with shifts, in a defined byte order. It is more code, it is fully portable, and it has no alignment issue at all.
- Use a packed struct of fixed-width types and assert its layout. This is acceptable when both ends have the same byte order, which is the case for a little-endian Cortex-M and a PC. Never use
int,long,enumor bit-fields in such a struct.
A third option avoids packing altogether: order the fields so that the natural layout has no internal padding, and add explicit reserved bytes where a gap is unavoidable. The struct is then both wire-exact and aligned.
Make the compiler check it
Whichever approach you use, turn your assumptions into compile-time checks. If someone adds a field or changes a type, the build fails instead of the protocol:
#include <stddef.h>
#include <stdint.h>
_Static_assert(sizeof(struct sample) == 8, "sample: size");
_Static_assert(offsetof(struct sample, id) == 4, "sample: id offset");
_Static_assert(offsetof(struct sample, flags) == 6, "sample: flags offset");
These cost nothing at run time, and they are the only layout documentation that cannot go out of date.
Summary
- Order members from largest to smallest. It removes internal padding for free.
- Remember that 64-bit integers and
doublealign to 8 on Cortex-M. - Use
packedonly for external layouts, not to save RAM, and never keep pointers to packed members. - Clear a struct with
memsetbefore filling it if its raw bytes will be transmitted, stored or compared. - Pin every layout that leaves the device with
_Static_assert.
Run the numbers: Struct Padding Calculator does this calculation for your own values.
More guides
- ADC Resolution Is Not Accuracy: LSB Size, Reference Error, Settling and Oversampling
- CAN Bit Timing Step by Step: Choosing BRP, Segments and Sample Point
- Why Your CRC Does Not Match: Polynomial, Init, Reflection and XOR-Out Explained
- Fixed-Point Arithmetic in Q Format: Scaling, Multiplying and Not Overflowing
- Floating-Point Pitfalls on Microcontrollers: Precision, Timestamps and Accidental Doubles
- Sizing I2C Pull-Up Resistors: Minimum, Maximum and What Bus Capacitance Does
- How Much UART Baud Rate Error Is Too Much? The Sampling Math