>>109450799
>play with setting the draft split command back to auto vs [1,0]
>got a whole tk/s faster with RAM offload for Q8 (32gb VRAMlet 5060ti + 5070ti)
>1-3tk/s Slower on Q6s fully offloaded with huge hit to pp
what is going on? this feels backwards. kobaldcpp layer splitting is an enigma sometimes but it fits more into VRAM than when i try to coompile llmao myself.
Gemma 4 31B Q6 K
61 Auto MTP [16:24:20] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.29s (798.92T/s), Generated:250/250 in 19.46s (12.85T/s), Total:44.81s
61 [1,0] MTP [16:26:26] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 25.27s (799.58T/s), Generated:250/250 in 17.69s (14.14T/s), Total:43.01s
61 NO MTP [16:28:09] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.98s (1264.76T/s), Generated:250/250 in 16.43s (15.22T/s), Total:32.46s
Gemma 4 31B Q6 K_L
61 Auto MTP [16:16:18] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.91s (811.33T/s), Generated:250/250 in 17.85s (14.00T/s), Total:42.81s
61 [1,0] MTP [15:57:37] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 24.88s (812.21T/s), Generated:250/250 in 16.84s (14.84T/s), Total:41.78s
61 NO MTP [15:59:33] CtxLimit:20457/20480, Init:0.06s, Processed:20207 in 15.99s (1263.96T/s), Generated:250/250 in 16.56s (15.10T/s), Total:32.60s
Gemma 4 31B Q8
51 Auto MTP [16:10:19] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.55s (683.87T/s), Generated:250/250 in 37.42s (6.68T/s), Total:67.04s
51 [1,0] MTP [16:35:46] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 29.54s (684.01T/s), Generated:250/250 in 44.25s (5.65T/s), Total:73.86s
51 NO MTP [15:50:04] CtxLimit:20457/20480, Init:0.07s, Processed:20207 in 24.80s (814.96T/s), Generated:250/250 in 42.95s (5.82T/s), Total:67.82s