>>109922495
I have 3 rigs:
>Node 1: 7800xtx + 128g ddr4 ecc
>Node 3: 7900xtx + 32g gayman ddr4
>Node 2: no GPU(file/media server) but 96g ddr5 ecc on epyc
qwen @ 50 t/s:
llamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
--spec-type draft-mtp \
-m /home/anon/Models/Mode-2/Qwen3.8-27B-Q8_0.gguf
Gemma MoE @ 50 t/s:
lamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/gemma-4-26B-A4B-it-Q8_0.gguf
qwen 31b, half the context, 10 t/s:
llamaserver \
-ngl 99 \
-c 16384 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
--spec-type draft-mtp \
--spec-draft-model /home/anon/Models/Mode-2/mtp-gemma-4-31B-it-Q8_0.gguf \
-m /home/anon/Models/Mode-2/gemma-4-31B-it-Q8_0.gguf
What's dlflash2? Thinking about an r9700 for the 2nd node.