Projects / experimental

SAM3 kernels

CUDA and SIMD kernel work on a real detector, where the same measurement runs in opposite directions under two contracts.

Status
experimental
Languages
C#, CUDA, C++
Current release
—
Latest
—

Strict contract 3.23–3.31x slower

Inference contract 2.9751x faster

Same detector, same machine, same ggml anchor. The posture is the missing denominator — see BENCHMARK_LEDGER.md weft-sam3-strict.

How it works

CUDA and SIMD kernels for a segment-anything-class detector, measured against a ggml anchor under two floating-point contracts.

SAM3 is work on the kernels of a segment-anything-class detector, benchmarked against a ggml anchor. It produced the survey's cleanest demonstration that a benchmark ratio without its contract is not a measurement: the same detector, on the same machine, against the same anchor, is 3.23–3.31x slower under the Strict contract and 2.9751x faster under Inference.

  • The same detector, machine and anchor: 3.23–3.31x slower under Strict, 2.9751x faster under Inference.
  • Neither number is wrong, and neither is 'the' number — the contract is the missing denominator.
  • A CUDA contraction-loss result recorded as a loss.

Releases