Lab / SEP 17 2026 / benchmark
The second instrument
The wall clock could not resolve iteration 3. The kernel census could: total GPU kernel time per encode 45.78 → 40.22 ms, with im2col down 19%.
GPU, Measurement methodology, BenchmarkingSAM3- Machine
- RTX 5090, Vast.ai instance 51353366
- Commit
- kernels on `gabilan/sam3-ggml` branches; census harness has no commit
- Total GPU kernel time / encode, base → f16b → f16c
- 45.78 → 40.66 → 40.22 ms (Δ −5.56 ms)
- kernels that moved
- 5
- kernels flat within noise
- 26
- im2col, base → f16c
- 2.32 → 1.87 ms/encode (−19%)
- Iteration 3 frame delta vs within-arm spread
- +0.48 / +0.77 / +0.41 ms vs 0.44 / 0.57 ms
Iteration 3 moved frame time by +0.48, +0.77 and +0.41 ms across three paired rounds. The within-arm spreads are 0.44 and 0.57 ms. The effect and the noise are the same size, so on frame time alone — the instrument the campaign is nominally about — that iteration is inconclusive.
The kernel census resolves it. `nsys --cuda-graph-trace=node`, ENC_ONLY, per-kernel GPU time divided by 48 encodes: `im2col` goes from 2.32 to 1.87 ms per encode, down 19%, with launch shapes unchanged and every other kernel flat.
| kernel | base | f16b | f16c | Δ | inst/enc |
|---|---|---|---|---|---|
| `convert_unary<float,__half>` → `_cont_vec4` | 5.11 | 2.59 | 2.60 | −2.51 | 137.3 |
| `k_mm_f16_to_f32_bias` (both gelu arms) → `_vec` | 4.54 | 3.60 | 3.60 | −0.94 | 132.7 |
| `cpy_scalar<float,__half>` (K/V) → `cpy_f32_f16_vec4` | 2.58 | 0.85 | 0.85 | −1.73 | 64.0 |
| `im2col_kernel<__half>` → `im2col_kernel_fastdiv` | 2.32 | 2.32 | 1.87 | −0.45 | 10.7 |
| `convert_unary<__half,float>` → `_cont_vec4` | 0.33 | 0.17 | 0.17 | −0.16 | 14.8 |
| *(26 more kernels, all flat within noise)* | |||||
| **total GPU kernel time / encode** | **45.78** | **40.66** | **40.22** | **−5.56** |
**Environment.** RTX 5090, Vast.ai instance 51353366; `cgroup cpu.max: 2304000 100000` (2.304 cores); `threads=4`; nsys `--cuda-graph-trace=node`, `ENC_ONLY`, divided by 48 encodes. Harness `run_census.sh`, which has no public commit.
**Methodology.** Same box, same model, same builds; three arms, one sitting; MATCHED, but not interleaved at the kernel level. Lower is better. Five kernels moved; 26 flat within noise; instance counts identical across arms for every kernel not the subject of the change.
The two instruments disagree about what is measurable, and the disagreement is the useful part. Host-observed wall clock carries the host's noise; kernel time carries the device's. Report an effect on the instrument that resolved it, and label the other one inconclusive.
The census environment record is a different file from the A/B's. `raw/gate.txt` holds exactly two lines, both census runs, both `gate=violated` (`other_cpus` 4.34 and 3.94). `run_ab.sh` writes to `run.log` with a different prefix, and that file also reads `gate=violated` on every run.