Two arms comparing the KNN convolution + pooling CUDA kernels — baseline (reference kernels) vs optimized (optimized kernels) — against a selectable dense reference. Pick a cell below; every table is filtered to it.
Cosine similarity of the optimized-vs-baseline final-layer output vectors (1.0 = identical direction; worst inference cell shown) — scale-invariant, unlike max|Δ|. dinov3 runs bf16, the CNNs fp16; all sit at >0.999 alignment. The smallest timing cells (batch 10, and 1-fixation on the fast GPUs) are clock-state sensitive; the 128×4 cell is the stable reference.