>>109987524
>mxfp4 is 3-4 times slower than a q4 gguf
thanks, i did find it pretty slow:
main: n_kv_max = 12288, n_batch = 4096, n_ubatch = 4096, flash_attn = 1, n_gpu_layers = 99, n_threads = 16, n_threads_batch = 16
| PP | TG | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s |
|-------|--------|--------|----------|----------|----------|----------|
| 4096 | 1024 | 0 | 5.187 | 789.66 | 66.552 | 15.39 |
| 4096 | 1024 | 4096 | 5.490 | 746.07 | 67.004 | 15.28 |
| 4096 | 1024 | 8192 | 6.125 | 668.69 | 67.593 | 15.15 |
i'll grab a Q4_K.
is this inkling model any good at q2_k_xl?