I'm >>109918636 - I have absolutely no idea why you need to pick either MoE or Dense models and not both. Is there a way to optimize arguments to get the best out of both?
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/Qwen3.8-27B-UD-IQ4_XS.gguf
Gets about 27 t/s, while the Q6-K quant does about 22. Meanwhile,
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/Qwen3-Coder-30B-A3B-Instruct-Q8_0.gguf
is getting me like 65 t/s.
How do I optimize the 27b models? Stop flinging shit at each other and help.