kime-cuda 0.0.25

The NVIDIA backend for kime: fused kernels, CUDA graphs, FP16, FP8 and INT8.
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
# Which of cuBLASLt's ranked algorithms each GEMM runs, per GPU as the driver names it. Rank 0 is
# cuBLASLt's own first choice and is what any shape or GPU not listed here gets. The lines come
# from a run with KIME_CUDA_TUNE=1, which times every rank in place and prints them. A shape runs
# one rank in every bucket, so the rank kept is the one closest on average to each bucket's fastest.
#
# gpu | rows inner columns | input output set-or-add | rank
#
# RTX 4090, tuned on driver 610.62 with CUDA 13.1, Laya English and multilingual shapes, ranked at
# 256 rows with no split K.
NVIDIA GeForce RTX 4090 | 256 1024 1024 | f16 f32 add | 1    # 1.06 of each bucket's best, rank 0 1.09
NVIDIA GeForce RTX 4090 | 256 1024 3072 | f16 f32 set | 2    # 1.01 of each bucket's best, rank 0 1.09
NVIDIA GeForce RTX 4090 | 256 1024 3072 | f32 f32 set | 1    # 1.13 of each bucket's best, rank 0 1.88
NVIDIA GeForce RTX 4090 | 256 1024 4096 | f16 f16 set | 2    # 1.02 of each bucket's best, rank 0 1.09
NVIDIA GeForce RTX 4090 | 256 1024 4096 | f32 f32 set | 1    # 1.11 of each bucket's best, rank 0 1.60
NVIDIA GeForce RTX 4090 | 256 1024 5248 | f32 f32 set | 1    # 1.08 of each bucket's best, rank 0 1.49