>>109775973
UD-IQ1_M(Unsloth) and nothing special
./llama-server -m ~/LLM/Qwen3.8-Flash-Next-UD-IQ1_M-00001-of-00003.gguf --spec-default -c 96000 --chat-template-kwargs '{"reasoning_effort":"low"}' --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 -fit on
>>109776044
It doesn't drop at all but I've set my harness to compact pretty aggressively. My main use case is to explore and ask questions about code but I don't like letting the clankers actually write anything important, so the current speed is fine.