On older PC and a low Profile DX12 card with 2GB VRAM:
./llama-b11430/ggml-rpc-server -H 0.0.0.0 -c -d Vulkan0
On newer AMD Laptop with 32GB RAM and (now) 4GB allocated to GPU:
./llama-b11433/llama-server --rpc 192.168.178.32:50052 --ctx-size 0 --threads 5 --parallel -1 --cont-batching --agent --prefill-assistant --spec-default --reasoning-preserve -m Cyber-Tiel-Coder-35B-A3B-MTP-UD-IQ3_XXS.gguf --mmproj mmproj-Q8_0.gguf --host localhost --port 9432 --cache-type-k q8_0 --cache-type-v q8_0 --image-min-tokens 1024 --split-mode layer -fa auto
First thing some might notice, I didn't use "--n-gpu-layers N". I tried various forms of it and it usually ended badly. Adding "--split-mode layer" works best for me. I seems to allocate just the right amount, by itself. Next I use the rocm version of llama on the laptop for stability reasons and vulkan on old PC, it's too old to be supported by rocm but just "new" enough to have vulkan support. Suprisingly I have no problems there.
With this current configuration, my token per second gen jumped from 4 (at best) to around 40-60!!! (and I still watch youtube Videos and suft the web on the laptop)
While having 32GB RAM on the Laptop is rather good and better than most people running local AI have avaible, I have to say you can absolutely use older Graphics Cards like "Radeon R7 240" for running AI. I totally would buy a few of those cheap and used instead of an newer 8GB+ card for an arm and a leg. Since the CPU basically doesn't matter and thus the regular RAM on the "GPU-Server" you could probably fit 2-3 of those cards on a old motherboard you'd have 4-6 GB more for a Model that mainly runs on your better PC. Everyone said 1GB Ethernet is too slow for an AI "cluster" but it just does it for me (100MBit or Wifi actually would be too slow though.
This probably should have gone to reddit, but I hate that place and there is basically nowhere else for me to post shit but here.
Thanks for reading!