Lab / SEP 17 2026 / benchmark

The second instrument

The wall clock could not resolve iteration 3. The kernel census could: total GPU kernel time per encode 45.78 → 40.22 ms, with im2col down 19%.

GPU, Measurement methodology, BenchmarkingSAM3
Machine
RTX 5090, Vast.ai instance 51353366
Commit
kernels on `gabilan/sam3-ggml` branches; census harness has no commit
Total GPU kernel time / encode, base → f16b → f16c
45.78 → 40.66 → 40.22 ms (Δ −5.56 ms)
kernels that moved
5
kernels flat within noise
26
im2col, base → f16c
2.32 → 1.87 ms/encode (−19%)
Iteration 3 frame delta vs within-arm spread
+0.48 / +0.77 / +0.41 ms vs 0.44 / 0.57 ms

Iteration 3 moved frame time by +0.48, +0.77 and +0.41 ms across three paired rounds. The within-arm spreads are 0.44 and 0.57 ms. The effect and the noise are the same size, so on frame time alone — the instrument the campaign is nominally about — that iteration is inconclusive.

The kernel census resolves it. `nsys --cuda-graph-trace=node`, ENC_ONLY, per-kernel GPU time divided by 48 encodes: `im2col` goes from 2.32 to 1.87 ms per encode, down 19%, with launch shapes unchanged and every other kernel flat.

kernelbasef16bf16cΔinst/enc
`convert_unary<float,__half>` → `_cont_vec4`5.112.592.60−2.51137.3
`k_mm_f16_to_f32_bias` (both gelu arms) → `_vec`4.543.603.60−0.94132.7
`cpy_scalar<float,__half>` (K/V) → `cpy_f32_f16_vec4`2.580.850.85−1.7364.0
`im2col_kernel<__half>` → `im2col_kernel_fastdiv`2.322.321.87−0.4510.7
`convert_unary<__half,float>` → `_cont_vec4`0.330.170.17−0.1614.8
*(26 more kernels, all flat within noise)*
**total GPU kernel time / encode****45.78****40.66****40.22****−5.56**
Per-kernel GPU time per encode, ms. The census is a Δ table; it states no ratio, so none is given here.

**Environment.** RTX 5090, Vast.ai instance 51353366; `cgroup cpu.max: 2304000 100000` (2.304 cores); `threads=4`; nsys `--cuda-graph-trace=node`, `ENC_ONLY`, divided by 48 encodes. Harness `run_census.sh`, which has no public commit.

**Methodology.** Same box, same model, same builds; three arms, one sitting; MATCHED, but not interleaved at the kernel level. Lower is better. Five kernels moved; 26 flat within noise; instance counts identical across arms for every kernel not the subject of the change.

The two instruments disagree about what is measurable, and the disagreement is the useful part. Host-observed wall clock carries the host's noise; kernel time carries the device's. Report an effect on the instrument that resolved it, and label the other one inconclusive.

The census environment record is a different file from the A/B's. `raw/gate.txt` holds exactly two lines, both census runs, both `gate=violated` (`other_cpus` 4.34 and 3.94). `run_ab.sh` writes to `run.log` with a different prefix, and that file also reads `gate=violated` on every run.