[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: llama.png (176 KB, 1500x500)
176 KB PNG
On older PC and a low Profile DX12 card with 2GB VRAM:
./llama-b11430/ggml-rpc-server -H 0.0.0.0 -c -d Vulkan0


On newer AMD Laptop with 32GB RAM and (now) 4GB allocated to GPU:
./llama-b11433/llama-server --rpc 192.168.178.32:50052 --ctx-size 0 --threads 5 --parallel -1 --cont-batching --agent --prefill-assistant --spec-default --reasoning-preserve -m Cyber-Tiel-Coder-35B-A3B-MTP-UD-IQ3_XXS.gguf --mmproj mmproj-Q8_0.gguf --host localhost --port 9432 --cache-type-k q8_0 --cache-type-v q8_0 --image-min-tokens 1024 --split-mode layer -fa auto


First thing some might notice, I didn't use "--n-gpu-layers N". I tried various forms of it and it usually ended badly. Adding "--split-mode layer" works best for me. I seems to allocate just the right amount, by itself. Next I use the rocm version of llama on the laptop for stability reasons and vulkan on old PC, it's too old to be supported by rocm but just "new" enough to have vulkan support. Suprisingly I have no problems there.

With this current configuration, my token per second gen jumped from 4 (at best) to around 40-60!!! (and I still watch youtube Videos and suft the web on the laptop)

While having 32GB RAM on the Laptop is rather good and better than most people running local AI have avaible, I have to say you can absolutely use older Graphics Cards like "Radeon R7 240" for running AI. I totally would buy a few of those cheap and used instead of an newer 8GB+ card for an arm and a leg. Since the CPU basically doesn't matter and thus the regular RAM on the "GPU-Server" you could probably fit 2-3 of those cards on a old motherboard you'd have 4-6 GB more for a Model that mainly runs on your better PC. Everyone said 1GB Ethernet is too slow for an AI "cluster" but it just does it for me (100MBit or Wifi actually would be too slow though.

This probably should have gone to reddit, but I hate that place and there is basically nowhere else for me to post shit but here.

Thanks for reading!
>>
should be relevant to this:
https://www.youtube.com/watch?v=fjbjDmWuCfg
>>
>>109999610
Yeah but for agentic anything you need context of over 64k at least
>>
>>109999610
rocm sucks on consumer cards, don't use it.
>>
I don't know whether to tell you you're stupid for bothering with such bad hardware, or to appreciate the indomitable human spirit to run LLMs at reasonable speeds on e-waste. I'll go with the appreciation mostly. It's A3B at IQ3_XXS and q8 kv cache, so it's not that great really, but it's probably what I would do if what you have is all I had.
You're pretty much just talking through RPC, which is a cool and underused feature. Your 1GB Ethernet being enough is in line with what others has reported.
You can get a used 8GB card for about £200-£250. Not really an arm and a leg. It's crazy that an RTX 3060 costs more than an RTX 3070 these days just because of the higher VRAM.
>>
>>109999610
I suggest requanting to EXL3. I've done that (you have to gather the calibration material yourself, though) and had gotten more or less Q6 quality for Q4 size and speed.
>>
>>110000586
The "--ctx-size 0" here doesn't actually mean zero but is a switch to use the native context of the model. In this case it rougly translates to around 260000. That seems enough for most things.
>>
>>109999610
>35B-A3B
The age of the MOE beowolf cluster is appon is.
>>
>>110001821
You can just not have that flag, the default is to give you min(model_ctx,whatever fits in ram)
>>
>>110000733
I wouldn't have bothered to post this, if there wasn't actual value in this so called "e-waste". yes, it doesn't hold a serious model in it's VRAM, you're right about this. However it still adds computing and as such relief to your main PC (if you have one). Those old cards can easyly be overclocked to double their regular stats and it shows.

I can now do lenghy AI related tasks comfortable in the background while surfing and watching videos. Without the rpc setup, this specific models either wouldn't run or I couldn't use the laptop for other things.
>>
>>110001844
>whatever fits in ram
that's exacly what I do not want. you see, I use the laptop while doing AI stuff. If it eats my ram, I cannot do that.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.