The GPU path composites frames bottom (last) to top (first); the CPU
fallback iterated top-first, so on machines without a working adapter
every multi-layer frame had its layering order inverted (the
composite_tracks test caught it as 0.8125 vs the documented 0.625).
Factor the CPU half into composite_tracks_cpu, iterate it in reverse
and pin the math in the test by calling the CPU path directly (the old
assertion silently exercised the GPU path whenever another test had
installed a shared context).