
Systems
From Homework Assignment to Low-Latency Benchmarking Engine
A systems programming assignment became a deep dive into memory allocation, SIMD, bitmasking, profiling, and the hidden costs of fast code.
18 min read
Rendering article…
Memory lab 01The result pointer outlives its worker pool
4 workers · 256 KiB reserved · 16 B usedTHREAD nPoolAllocator
64 KiB · base aligned 64 Bint hits4 B · allocate<int>()offset +4
65,532 B unusedallocation enlarged for legibility
Actual capacity use · 4 / 65,536 bytes
std::mt19937_64 enginethread-local RNG staterandX[4] + randY[4]stack · alignas(32) on AVX2randX[2] + randY[2]stack · alignas(16) on NEON worker allocates 4 B → hits* thread exits → pool frees buffer main joins → *hits is dangling
Rendering article…
Cache lab 02How layout changes a cold sequential pass
Model only · MCBE pool stores no doublesaccess—
L1 hits0
cold misses0
last probeREADY
line 00
01234567
L1 probe→L2 / L3→memory
Eight adjacent doubles share one 64-byte cache line.
Rendering article…
Compute lab 03Four dart tests collapse into one mask
AVX2 · 4 × float64 lanesx² + y²→≤ 1.0→movemask→popcount
lane 00.42pending·
lane 11.24pending·
lane 20.81pending·
lane 30.67pending·
0b····popcount waitingRendering article…
