Lab / / failure
Generated CUDA at 0.025× of cuBLAS
Weft's generated CUDA contraction runs between 4.73× and 39.38× behind cuBLAS, and is bit-exact against Weft's own CPU oracle.
GPU, Negative results, CorrectnessWeft- Machine
- 8× RTX 5090, sm_120; CUDA 13.1, driver 590.48.01
- Commit
- UNKNOWN
- Weft CUDA contraction vs cuBLAS
- 0.025× – 0.211×
- Expressed as a deficit
- 4.73× to 39.38× behind
- Correctness
- bit-exact against Weft's own CPU oracle
Weft's generated CUDA contraction runs at 0.025× to 0.211× of cuBLAS. That is between 4.73× and 39.38× behind.
It is also bit-exact against Weft's own CPU oracle.
The pairing is the finding
The generated code is correct and it is up to 39× slow. I am reporting those two facts in the same paragraph because they are separate axes, and this is the clearest case in the work where reporting only the flattering one would have been easy. Codegen that produces the right answer and runs 39× behind is a codegen problem, not a numerics problem, and it needs the second number to be visible for anyone to work on it.
| Measurement | cuBLAS | Weft generated CUDA contraction | Delta | Ratio |
|---|---|---|---|---|
| contraction throughput, worst case | 1.00× (reference) | 0.025× | 39.38× behind | 0.025×quality unknown |
| contraction throughput, best case | 1.00× (reference) | 0.211× | 4.73× behind | 0.211×quality unknown |
The range is wide because the ratio depends on the shape. Both ends of it are losses, so I am not quoting the top of the range as the result.