Contents
A pull request that adds secure_zero(&key, sizeof key) to a crypto function looks like pure upside. Frank Denis’s first zeroization post shows it is not. For a small local value, the wipe can be deleted by the optimizer, can clear bytes that never held the secret, or can make the compiler write a copy of the secret to memory that did not exist before the wipe was added.
The mechanism is the same each time. A wipe takes an address, and taking an address changes how the compiler lays out the value. Whether the wipe helps depends on where the secret actually lives in the generated code, which you can only see by reading the assembly of the build you ship.
Three ways a wipe misses
Denis’s test function computes an intermediate from its input, derives a result, then wipes the intermediate. I used the same function and the same variants. My build was x86-64, clang 18.1.3 and gcc 13.3.0 at -O3. Denis’s was arm64 with Apple’s Clang 21, so the instructions differ but the behavior mostly matches.
uint64_t f(uint64_t input)
{
uint64_t intermediate = input ^ UINT64_C(0x9e3779b97f4a7c15);
uint64_t result = (intermediate >> 17) ^
(intermediate * UINT64_C(0xd6e8feb86659fd93));
/* wipe(&intermediate) goes here */
return result;
}
Plain memset. The compiler sees that intermediate is never read again and removes the call. Denis reports identical output with and without it. In my clang 18 build the function with memset and the one without it compiled to identical instructions. The value also never reached memory, so there was nothing on the stack to wipe in the first place.
A volatile byte loop. This is the usual fix: write zeros through a volatile unsigned char *. The compiler must now emit the stores. Denis shows eight strb wzr instructions on arm64 and points out that nothing ever stored intermediate into those slots. It stayed in register x8. On x86-64 I saw the same shape. Clang emitted eight single-byte zero stores below rsp and kept the value in rcx through the multiply and shift. GCC did the same with rdi and rax. The stores cost cycles and clear memory that never held the secret.
An opaque wipe function. Put the volatile loop in another translation unit, with no LTO, and the caller cannot see what it does. It has to assume the callee may read the old contents. So the caller must first write the value to memory so the callee has something to read. Both compilers did this on x86-64. Clang emitted mov qword ptr [rsp + 8], rax before call opaque_wipe, and GCC stored the value at [rsp]. The unwiped version never touched the stack. The wipe added a window in which the secret sits in memory, and then closed that window while the register copy that fed the result was untouched.
Denis also notes that a recognized write-only operation can skip that store, and that under LTO the opaque function is no longer opaque. A wipe that works in one build configuration can quietly stop working in another.
The compare-and-wipe case
The more interesting example computes a 16 byte tag, compares it to a candidate, wipes the tag, and returns the comparison result:
int compare_and_wipe(const unsigned char seed[16],
const unsigned char candidate[16])
{
unsigned char computed[16];
for (size_t i = 0; i < 16; i++)
computed[i] = seed[i] ^ 0x5a;
unsigned int different = 0;
for (size_t i = 0; i < 16; i++)
different |= (unsigned int)(computed[i] ^ candidate[i]);
int equal = different == 0;
opaque_wipe(computed, 16);
return equal;
}
In the source, the comparison happens before the wipe. In Denis’s arm64 output the compiler loaded both operands before the call but put the cmeq after it. To keep the tag alive across the call it saved a second copy in a spill slot, and the wipe only covered the addressed object. His stack map has the candidate at S, the extra tag copy at S + 16, and the wiped object at S + 40. Remove the wipe and the tag never goes to the stack at all, so, in his words, adding the wipe created the copy that was left behind. He is explicit that this is not a compiler bug: C does not promise the ordering of security-relevant effects that the programmer had in mind.
This one reproduced only partly for me:
- Clang 18 on x86-64 did the comparison before the call and kept a single stack copy at
[rsp], which the call wipes. The tag still passed through anxmmregister. No leftover slot. - GCC 13 (Ubuntu’s build, which enables the stack protector by default) did spill. It stored the tag at
16[rsp], which is the object passed to the wipe. It also stored the folded comparison vector at[rsp], called the wipe, then reloaded that slot to finish the comparison. That slot is never cleared. It holds a value derived from the tag and the candidate rather than the tag itself, so it is a weaker leak than the one in Denis’s arm64 build, but it is a stack copy that exists only because the wipe call sits in the middle of the computation.
The same source gave three different outcomes across arm64 Clang 21, x86-64 Clang 18 and x86-64 GCC 13. I did not test other versions or flags, and I did not check what changes without the stack protector.
Registers are copies too
Denis adds a point that no zeroing function can address. Registers are written to memory on context switches, appear in core dumps, hibernation files and VM snapshots, and can leak through microarchitectural bugs like Zenbleed. A single SIMD register can hold a 256 bit key or a 512 bit hash, and those registers tend to be recycled less often than general-purpose ones. A function that writes zeros to memory cannot reach any of that.
What to do with this
The post does not argue against wiping. Password buffers, allocated key schedules and retired contexts already live in memory, and clearing them is worth doing. The argument is against reflexively adding a wipe to a small local. Denis’s checklist is to ask whether the value already lives in memory, whether taking its address creates storage, and what stays live across the call, and to inspect the optimized build you ship, including LTO and hardening flags.
For review, that suggests a concrete rule. A PR that adds a wipe to a stack variable should come with the disassembly of the release build for the compiler and target you ship, showing where the value lives before and after the change. Without that, the wipe is a belief about what the compiler did.
Denis says part 2 will cover more reliable techniques. I have not seen it yet, so I am not recommending a replacement here.
Sources
- Zeroization, part 1: Wiping can make things worse, Frank Denis, October 6, 2026. Includes the first and second Godbolt examples.
- Zenbleed, Tavis Ormandy.
- x86-64 results are from my own compilation of the examples above with
clang 18.1.3andgcc 13.3.0(Ubuntu 24.04),-std=c11 -O3, assembly read by hand. Denis’s arm64 results are his, not mine.