From Homework Assignment to Low-Latency Benchmarking Engine

A systems programming assignment became a deep dive into memory allocation, SIMD, bitmasking, profiling, and the hidden costs of fast code.

Rendering article…

Memory lab 01The result pointer outlives its worker pool
4 workers · 256 KiB reserved · 16 B used
4 Bused / 65,536 B
THREAD nPoolAllocator
64 KiB · base aligned 64 B
int hits4 B · allocate<int>()offset +4
65,532 B unusedallocation enlarged for legibility
Actual capacity use · 4 / 65,536 bytes
OUTSIDE THE POOLSIMD working data
std::mt19937_64 enginethread-local RNG state
randX[4] + randY[4]stack · alignas(32) on AVX2
randX[2] + randY[2]stack · alignas(16) on NEON
worker allocates 4 B → hits* thread exits → pool frees buffer main joins → *hits is dangling

Rendering article…

Cache lab 02How layout changes a cold sequential pass
Model only · MCBE pool stores no doubles
access—
L1 hits0
cold misses0
last probeREADY
line 00
01234567
L1 probe→L2 / L3→memory

Eight adjacent doubles share one 64-byte cache line.

Rendering article…

Compute lab 03Four dart tests collapse into one mask
AVX2 · 4 × float64 lanes
x² + y²→≤ 1.0→movemask→popcount
lane 00.42pending·
lane 11.24pending·
lane 20.81pending·
lane 30.67pending·
0b····popcount waiting

Rendering article…

© 2026 Leon Do ・ 恋してる