Choosing Between ARM64 LSE and Legacy LDREX/STREX for Lock‑Free Atomics
Decide whether to use ARM64 Large System Extensions (LDAXR/STXR) or traditional LDREX/STREX for lock‑free data structures. Compare CPU, OS, compiler support, performance, and pitfalls, and see a practical C example with runtime guard.
07 Mar 2026, 16:14 UTC

Decision Context
When building lock‑free data structures on ARM64, you can rely on the traditional LDREX/STREX pair or the newer Large System Extensions (LSE) instructions LDAXR/STXR. The choice affects scalability, compatibility, and code maintenance.
Key Decision Question
Should your project adopt LSE atomic instructions for lock‑free counters and pointers, or stick with legacy LDREX/STREX?
Constraints to Consider
- CPU Architecture: LSE is available only on ARMv8.5‑A and later. Older cores (e.g., Cortex‑A53) will trap on
LDAXR/STXR. - Operating System: The kernel must expose LSE semantics via proper memory ordering. Linux 5.10+ and Android 12+ include support.
- Compiler Support: GCC 11+ and Clang 13+ provide built‑in intrinsics for LSE. Earlier compilers require hand‑written assembly.
- Target Audience: If your binary must run on a mix of old and new CPUs, a runtime guard is essential.
Option Comparison
| Feature | LSE (LDAXR/STXR) | Legacy LDREX/STREX |
|---|---|---|
| Scalability | Higher – reduces false sharing on multi‑core systems. | Lower – can suffer from cache line contention. |
| False Sharing | Minimized – operates on a single cache line. | Possible – entire line may be invalidated. |
| Performance Impact | Up to 30% faster for high‑contention counters on Cortex‑A76/A78. | Baseline performance. |
| Compatibility | Requires ARMv8.5+; may trap on older CPUs. | Supported on all ARMv8-A and earlier. |
| Compiler Intrinsics | __atomic_load_n/__atomic_store_n (with -march=armv8.5-a). | __atomic_fetch_add, etc. (no special flags). |
| Runtime Guard Needed | Yes – to avoid illegal instruction traps. | No – safe on all CPUs. |
| Memory Ordering Guarantees | Same as LDREX/STREX but with stronger atomicity on multi‑core. | Standard sequential consistency. |
| Implementation Complexity | Higher – need to detect support and provide fallback. | Lower – straightforward use. |
Trade‑Off Summary
- Performance vs. Compatibility: If your deployment targets only modern CPUs (v8.5+), LSE offers measurable throughput gains for lock‑free counters. For heterogeneous environments, the risk of traps outweighs the benefit.
- Development Effort: Adding a runtime check and fallback path increases code complexity. However, it isolates the risk to a single module.
- Future Proofing: LSE is becoming the standard for high‑performance ARM kernels. Adopting it now positions your code for upcoming CPUs.
Concrete Implementation Example
Runtime LSE Detection
On ARM64, the ID_AA64ISAR0_EL1 register holds the LSE feature bit (bit 4). The following helper checks for support and caches the result.
#include <stdbool.h>
#include <stdint.h>
#include <signal.h>
#include <unistd.h>
static bool lse_supported = false;
static bool lse_checked = false;
static void sigill_handler(int sig) {
/* Ignore illegal instruction – indicates LSE not present */
lse_supported = false;
lse_checked = true;
}
static bool check_lse(void) {
if (lse_checked) return lse_supported;
struct sigaction sa, old;
sa.sa_handler = sigill_handler;
sigemptyset(&sa.sa_mask);
sa.sa_flags = 0;
sigaction(SIGILL, &sa, &old);
/* Try a single LSE instruction in a protected block */
__asm__ volatile(
"mrs %0, ID_AA64ISAR0_EL1\n"
: "=r"(lse_supported)
);
lse_supported = (lse_supported & (1u << 4)) != 0;
sigaction(SIGILL, &old, NULL);
lse_checked = true;
return lse_supported;
}
Note: The mrs read is safe on all CPUs; it simply checks the feature bit. If the bit is cleared, the code will never execute LDAXR or STXR.
Atomic Increment Using LSE
Below is a lock‑free counter increment that uses LSE when available, otherwise falls back to the legacy LDREX/STREX loop. The code assumes -march=armv8.5-a when compiling for LSE.
#include <stdatomic.h>
#include <stdint.h>
/* Forward declaration of the runtime check */
extern bool check_lse(void);
static inline uint64_t atomic_inc_lse(volatile uint64_t *ptr) {
uint64_t old, new;
do {
old = __atomic_load_n(ptr, __ATOMIC_ACQUIRE);
new = old + 1;
} while (!__atomic_compare_exchange_n(ptr, &old, new,
true, __ATOMIC_ACQ_REL,
__ATOMIC_ACQUIRE));
return new;
}
static inline uint64_t atomic_inc_ldrex(volatile uint64_t *ptr) {
uint64_t old, new;
do {
__asm__ volatile(
"ldrex %0, [%2]\n"
: "=&r"(old) : "r"(ptr) : "memory");
new = old + 1;
unsigned int result;
__asm__ volatile(
"strex %0, %1, [%2]\n"
: "=&r"(result) : "r"(new), "r"(ptr) : "memory");
} while (result != 0);
return new;
}
static inline uint64_t atomic_inc(volatile uint64_t *ptr) {
if (check_lse()) {
return atomic_inc_lse(ptr);
} else {
return atomic_inc_ldrex(ptr);
}
}
Explanation of the code:
- The
atomic_inc_lsefunction uses GCC/Clang built‑ins that map toLDAXR/STXRunder-march=armv8.5-a. - The
atomic_inc_ldrexfunction shows the legacy inline‑assembly path for older CPUs. - The dispatcher
atomic_incchooses the path at runtime after the feature check. - All paths use sequentially consistent ordering (
__ATOMIC_ACQ_REL).
Verification Steps
- Compile with
-march=armv8.5-a -O2and link againstlibatomicif using GCC. - Run on a CPU that reports LSE support (e.g.,
cat /proc/cpuinfo | grep LSEon recent Linux). - Profile a tight loop that calls
atomic_incfrom many threads. Useperf stat -e cycles,cycles:u,cycles:kto compare throughput against a pure LDREX implementation. - Test Fallback by running the binary on an older core (e.g., Cortex‑A53) and ensuring no
SIGILLoccurs.
Limitations and Practical Checks
- On CPUs that support LSE but run an older kernel, the kernel might not provide the necessary memory ordering guarantees. Verify by reading
/sys/kernel/mm/atomic/or checkinguname -afor kernel version. - Complex lock‑free algorithms that rely on specific inter‑instruction ordering may still prefer LDREX/STREX for correctness.
- The runtime check uses a single
mrsread; if the CPU implements LSE but the OS masks it, the check will still return true but the instruction may trap. A more robust approach is to execute a dummyLDAXR/STXRin asetjmp/longjmporsigsetjmpcontext. - Always test on the target hardware before shipping. Profiling on representative workloads is essential to justify the added complexity.
Conclusion
Adopting ARM64 LSE atomic instructions is worthwhile when your deployment targets modern CPUs and you need high‑throughput lock‑free counters. The trade‑off is added runtime detection and a fallback path. For heterogeneous environments or when simplicity is paramount, sticking with legacy LDREX/STREX remains the safest choice.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.