Debugging
A race only the Radeon could see
Identical CFD runs on an AMD integrated GPU gave slightly different forces. The cause was a Gauss-Seidel loop inside one workgroup, and a barrier that did not mean what we thought on this hardware.
A GPU solver should give the same bits for the same input on the same machine. Floating-point addition is not associative, so a different order gives a different result. But LeonSim’s kernels fix their order: no atomics in the solver, fixed reduction trees, red-black sweeps. On NVIDIA GPUs and on the llvmpipe software driver, repeated runs were bit-identical.
On the mini PC’s AMD Radeon 680M (RDNA 2, Mesa RADV driver) they were not. We noticed while checking the new on-device force path: two identical F1 runs drifted apart by a fraction of a percent over a few hundred iterations. In a flow that oscillates, a tiny difference at iteration 50 becomes a different averaging window at iteration 700.
Narrowing it down
We had a list of the operations of one pressure correction (now exposed as pressure_ops), so we could replay them one kernel at a time from a fixed state and compare the outputs bit for bit. Every kernel was deterministic except one: mg_coarsest, which solves the coarsest multigrid level. Repeating the same coarsest solve 300 times gave 11 different answers, differing by up to 3e-3.
The coarsest level is small (at most a couple of thousand cells), so it is solved by one workgroup of 256 threads with red-black Gauss-Seidel sweeps: update the red cells, barrier, update the black cells, barrier, repeat. Each red cell reads its black neighbours, which other threads in the same workgroup wrote in the previous half sweep. The iterate lived in a storage buffer, and the barrier between half sweeps was
storageBarrier();
workgroupBarrier();
Why it went wrong on this GPU
Our reading of the behaviour: an RDNA work-group processor contains two compute units, each with its own L0 vector cache, and in WGP mode one workgroup can be spread over both. A thread on one compute unit can then read a neighbour’s value from its own L0 while the fresh value written by a thread on the other compute unit has not reached it, even after the barrier. Most of the time the timing hides it. About one solve in thirty, it does not.
We did not prove the cache mechanism from the hardware side. What we did show: the race is specific to storage-buffer communication inside a workgroup on this device, and it disappears completely when that communication moves to workgroup memory, which is shared on-chip and ordered by workgroupBarrier() on every GPU.
The fix
var<workgroup> W: array<f32, 2048>; // the iterate, on chip
for (var c = li; c < n; c += 256u) { W[c] = 0.0; }
workgroupBarrier();
for (var it = 0u; it < 2u * sweeps; it++) {
let color = it & 1u;
for (var c = li; c < n; c += 256u) {
if (parity(c) == color && ACT[c] != 0u) {
W[c] = (B[c] + neighbour_sum(W, c)) / D[c];
}
}
workgroupBarrier();
}
for (var c = li; c < n; c += 256u) { V[c] = W[c]; } // back to the storage buffer once
The price is a hard limit: 2048 coarsest cells (8 KiB of workgroup memory). The multigrid set-up now refuses a hierarchy whose coarsest level is larger instead of silently running it.
A bug that had been hiding as a precision limit
One of the GPU tests, S27.2, compares the multigrid-CG pressure solve with a direct solve. It had always failed on this Radeon, for one variant (semi-coarsening 3): error 3.2e-4 against a 2e-4 tolerance. We had written it off as float32 precision on a weak GPU and documented it that way. It was the race. With the fix the error is 6.4e-5, the CG needs fewer iterations (23 → 16 in that test), and the default solver is more accurate on AMD as well.
The test suite has a new check, S27.8: three identical pressure solves must give the same bits, and 200 coarsest solves must give one result. It passes on all six NVIDIA GPUs we tested (T4 to H100), on the Radeon 680M and on llvmpipe. The F1 results from before the fix did not change. The fix changes no arithmetic, only which values each thread is guaranteed to see. Those runs had happened not to hit the race, so the bit-identical results confirm both.
Takeaways
- Test determinism directly. Run the same thing many times and compare bits. A pass-or-fail tolerance test let this bug pass as noise.
- Use workgroup memory for anything threads of one workgroup exchange. Storage-buffer communication inside a workgroup depends on how a vendor implements memory coherence, and portable code cannot rely on it.
- Do not trust a “precision floor” you have not explained. Ours was a bug.