[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


I am so fucking ready
>>
>>109544900
this is literally antisemitism, anon you should use only openAI or Anthropic, and nothing else.
>>
>>109544911
>moloch or baal
neither
>>
>>109544900
Q4 doesn't fit into 16 gb vram. Useless
>>
>>109544911
>antisemitism
Any other features we want to know about?
>>
>>109544900
Why wait for the release of a model that needs to be quantised for you to use?
>>
>>109545148
cringe
i'm planning to run it on 64gb of ram cpu only
yes i will do 1 prompt per day
>>
>>109544900
>release in a day
So about two days until we get a tuned version got it.
>>
>>109544900
>>109544911
>>109544934
It's like the special Olympics with these AI drones.
>>
>>109545779
chad
>>
>>109544911
>you should use only openAI or Anthropic
that's antigentilism
>>
>>109545779
It's bad but not as bad as you think.
Expect single digit tk/s.
>>
File: 1495055755873.png (746 KB, 640x640)
746 KB PNG
>>109545779
Based. I use my free account on kaggle 2x T4 gpus.
>>
>>109544900
I am poor. I need 4b one.
>>
File: 1759084265784076.png (233 KB, 750x1000)
233 KB PNG
>>109547315
>4b
At that point just use the free tier on openrouter and don't feed the models personal details. I'm a local model maximalist but Jesus dude..
>>
>>109547315
How big is the 4b version?
>>
>>109545148
YOUR PENIS DOESN'T FIT IN A PUSSY. USELESS
>>
>>109547561
lmfao, actually btfo
>>
>>109547114
How?
>>
>>109544900
does it fit in a rtx 3090?
>>
>>109549939
The same way I've been making porn images with stable diffusion sdxl 1.0 for about 2 years.
>>
>>109545148
If you’re serious about using local AI, you use at least 32 GB of RAM.
I can comfortably fit 27B and even 30-35B models (those only with lower active parameters like A3b or A4B though) and have them work fast. Currently using NVIDIA Nemotron 30B A3B and getting ~165 T/s output.
>>
I'm on gemma 4:31, how does it compare?
>>
>>109551429
How?
>>
>>109553220
AxB ones are moe, so the whole expert fits even on 8gb card, 27b is dense, will need 5090 and even then quantization
>>
File: 1782020580629222.gif (1.59 MB, 300x222)
1.59 MB GIF
>>109553220
>Currently using NVIDIA Nemotron 30B A3B and getting ~165 T/s output.
Sure you do
>>
>>109544900
if they're going to release better <4B ASR and TTS vairants, that would be great. otherwise, i don't care for these big models i can't self-host.
>>
>>109550108
a 31b? of course it will, at a fine quant too.
>>109553762
eyeballing it he probably does, A3B = MoE with active 3B, MoEs run fine, i assume based on the post he was replying to he means he has 16GB VRAM and 32GB RAM which is the same as me. the gemma that's like 27b a5b or whatever runs very fast on my setup. unfortunately dense 31b is usable for a batched task but slower than human reading speed so not good for chat and unusable for live coding/agent stuff, and i assume a dense 27b would be not much better. So he's wrong in predicting qwen 27b being plausible usable on a 16GB card, yes
>>
>>109553220
>getting ~165 T/s output.
I feel like a hobo with my intel card getting 35 t/s even on gemma 4 26b a4b, fuck partial offloading.
>>
IT FUCKING DROPPED AAAAAAA
IM GONNA COOOM
>>
>Qwen 3.8 27B fully 99% downloaded
>???
>Back to 0%, all disk space that it was occupying is free again
Did huggingface servers die or smth?
>>
I wish we moved past the period where everyone wants 200gb of hbm ram for their home systems to period where the existing hardware does the jobs people need and the end result is shared to all.
You know, theres so much redundant work being done in the most inefficient way possible.
>>
>>109553220
>I can comfortably fit 27B and even 30-35B models
alright so you don't know what you're talking about
>>
>>109544900
Will there be an A3B model for us peasants?
>>
>>109555585
>where the existing hardware does the jobs people need and the end result is shared to all
Did you just wake up from a MDMA induced coma that began in the 80s or something?
>>
>>109555740
I hope for AxB that is just perfect for 4060's 8gb, 3b was dumb as fuck, make it A7-8B and we're talking
>>
>>109544900
Should fit fine into my m3 ultra/512gb ram
>>
Ok, it can't load enough experts on 10gb of VRAM+20gb RAM to do image prompts, not like Gemma can either to be honest.
>>
I'll wait for the 1bit Bonsai quant.
>>
>>109556525
I've given up, there will never be a model good for more than text summarizing that runs on less than 32GB VRAM
>>
>>109544900
Will I be able to run this on my desktop?
> RTX 5070 12GB
> 32GB of DDR4 RAM
>>
Am I missing something? 30B is a retarded model. Did something change in the past couple years with quants and MoE or is this just settling for what you can fit into VRAM on local?
>>
>>109559893
Settling, although I'll admit models that can run on conventional hardware have more value sometimes.

I've tried Qwen 3.6 27B, Qwen 3.6 35B A3B, and Qwen 3.5 122B A10B. The 122B model is no-contest the best of the three. 27B is surprisingly decent for the size though, if you need the best you can manage with one GPU. And the 35B model is fast even when you put it on crap hardware.
>>
>>109559856
yeah with it running slow as molasses and falling into thinking loops
>>
>>109560054
why though? Isn't this MoE?
>>
>>109560086
>Isn't this MoE?
No
>>
>>109559988
>although I'll admit models that can run on conventional hardware have more value sometimes
>Qwen 3.6 35B A3B
>And the 35B model is fast even when you put it on crap hardware
Runs at 15t/s on just CPU without jewery services spying on you or leaving you with <fill in the blank yourself, deadbeat> answers. What more do you need? Its prefectly serviable if you want a local search engine and error corrector. Perhaps not so much if you're in the "muh heckin ai is going to take our jerbs" scene that demands utter useless babble derived from hundreds of hours of parallelized compute time against a 5TB dataset. But still, good enough as a less retarded clippy.
>>
>>109560100
oh well, fuck
>>
File: .png (56 KB, 847x586)
56 KB PNG
>>109545779
>>109546025
>>109546287
>1.38 t/s
kek
>>
>>109560164
vs >>109560102 15t/s on Qwen 3.6 35B A3B Q8 XL?
>>
>>109560180
Yeah, 35B-A3B is wayyy faster. about 10-15t/s on the same machine (q8_k_xl). I'm just using q4 here because I'm experimenting with different quants but smaller quants dont seem to speed up by enough to care about.
I've been trying KAT-Coder-V2.5-Dev and Qwen-AgentWorld too which are all based off qwen3.6 35b-a3b. the MoEs are far more practical for my hardware.
also qwen 3.8 basically requires setting "low" reasoning effort, or otherwise it just wastes 6 gorillion tokens going in loops
>>
>>109560102
It depends on what you're doing with it. I have 35B A3B running a RAG on okay hardware and it's real good at that. But try and do heavier work with it, such as running an agent, and it flounders. The speed advantage makes it easy to use, but not always good to use. You hit the nail on the head: "less retarded clippy." Light all-around work is where it shines.
>>
remember to use the fixed templates
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
>>
Can you changshills stop hyping up these dogshit models for a while?
>>
>>109544911
Nope. Chuck Testa.
>>
>>109544900
Based
Codetrans will seethe tho
>>
>>109555740
If you can a claw up measly $600, it runs at 40t/s on a V100 32GB.
>>
What's so great about Qwen? I get better results out of Mistral



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.