This recommendation is architecture-specific. It is true for x86 & x86_64 (in a low-level programming). You should also check that compiler don’t reorder your code. You can use “compiler memory barrier” for that.
Low-level atomic read and writes for x86 is described in Intel Reference manuals “The Intel® 64 and IA-32 Architectures Software Developer’s Manual” Volume 3A ( http://www.intel.com/Assets/PDF/manual/253668.pdf) , section 8.1.1
8.1.1 Guaranteed Atomic Operations
The Intel486 processor (and newer processors since) guarantees that the following
basic memory operations will always be carried out atomically:
- Reading or writing a byte
- Reading or writing a word aligned on a 16-bit boundary
- Reading or writing a doubleword aligned on a 32-bit boundary
The Pentium processor (and newer processors since) guarantees that the following
additional memory operations will always be carried out atomically:
- Reading or writing a quadword aligned on a 64-bit boundary
- 16-bit accesses to uncached memory locations that fit within a 32-bit data bus
The P6 family processors (and newer processors since) guarantee that the following
additional memory operation will always be carried out atomically:
- Unaligned 16-, 32-, and 64-bit accesses to cached memory that fit within a cache
line
This document also have more description of atomically for newer processors like Core2. Not all unaligned operations will be atomic.
Other intel manual recommends this white paper: