I am so fucking ready
>>109544900this is literally antisemitism, anon you should use only openAI or Anthropic, and nothing else.
>>109544911>moloch or baalneither
>>109544900Q4 doesn't fit into 16 gb vram. Useless
>>109544911>antisemitismAny other features we want to know about?
>>109544900Why wait for the release of a model that needs to be quantised for you to use?
>>109545148cringei'm planning to run it on 64gb of ram cpu onlyyes i will do 1 prompt per day
>>109544900>release in a daySo about two days until we get a tuned version got it.
>>109544900>>109544911>>109544934It's like the special Olympics with these AI drones.
>>109545779chad
>>109544911>you should use only openAI or Anthropicthat's antigentilism
>>109545779It's bad but not as bad as you think.Expect single digit tk/s.
>>109545779Based. I use my free account on kaggle 2x T4 gpus.
>>109544900I am poor. I need 4b one.
>>109547315>4bAt that point just use the free tier on openrouter and don't feed the models personal details. I'm a local model maximalist but Jesus dude..
>>109547315How big is the 4b version?
>>109545148YOUR PENIS DOESN'T FIT IN A PUSSY. USELESS
>>109547561lmfao, actually btfo
>>109547114How?
>>109544900does it fit in a rtx 3090?
>>109549939The same way I've been making porn images with stable diffusion sdxl 1.0 for about 2 years.
>>109545148If you’re serious about using local AI, you use at least 32 GB of RAM. I can comfortably fit 27B and even 30-35B models (those only with lower active parameters like A3b or A4B though) and have them work fast. Currently using NVIDIA Nemotron 30B A3B and getting ~165 T/s output.
I'm on gemma 4:31, how does it compare?
>>109551429How?
>>109553220AxB ones are moe, so the whole expert fits even on 8gb card, 27b is dense, will need 5090 and even then quantization
>>109553220>Currently using NVIDIA Nemotron 30B A3B and getting ~165 T/s output.Sure you do
>>109544900if they're going to release better <4B ASR and TTS vairants, that would be great. otherwise, i don't care for these big models i can't self-host.
>>109550108a 31b? of course it will, at a fine quant too.>>109553762eyeballing it he probably does, A3B = MoE with active 3B, MoEs run fine, i assume based on the post he was replying to he means he has 16GB VRAM and 32GB RAM which is the same as me. the gemma that's like 27b a5b or whatever runs very fast on my setup. unfortunately dense 31b is usable for a batched task but slower than human reading speed so not good for chat and unusable for live coding/agent stuff, and i assume a dense 27b would be not much better. So he's wrong in predicting qwen 27b being plausible usable on a 16GB card, yes
>>109553220>getting ~165 T/s output.I feel like a hobo with my intel card getting 35 t/s even on gemma 4 26b a4b, fuck partial offloading.
IT FUCKING DROPPED AAAAAAAIM GONNA COOOM
>Qwen 3.8 27B fully 99% downloaded>???>Back to 0%, all disk space that it was occupying is free againDid huggingface servers die or smth?
I wish we moved past the period where everyone wants 200gb of hbm ram for their home systems to period where the existing hardware does the jobs people need and the end result is shared to all. You know, theres so much redundant work being done in the most inefficient way possible.
>>109553220>I can comfortably fit 27B and even 30-35B modelsalright so you don't know what you're talking about
>>109544900Will there be an A3B model for us peasants?
>>109555585>where the existing hardware does the jobs people need and the end result is shared to allDid you just wake up from a MDMA induced coma that began in the 80s or something?
>>109555740I hope for AxB that is just perfect for 4060's 8gb, 3b was dumb as fuck, make it A7-8B and we're talking
>>109544900Should fit fine into my m3 ultra/512gb ram
Ok, it can't load enough experts on 10gb of VRAM+20gb RAM to do image prompts, not like Gemma can either to be honest.
I'll wait for the 1bit Bonsai quant.
>>109556525I've given up, there will never be a model good for more than text summarizing that runs on less than 32GB VRAM
>>109544900Will I be able to run this on my desktop?> RTX 5070 12GB> 32GB of DDR4 RAM
Am I missing something? 30B is a retarded model. Did something change in the past couple years with quants and MoE or is this just settling for what you can fit into VRAM on local?
>>109559893Settling, although I'll admit models that can run on conventional hardware have more value sometimes.I've tried Qwen 3.6 27B, Qwen 3.6 35B A3B, and Qwen 3.5 122B A10B. The 122B model is no-contest the best of the three. 27B is surprisingly decent for the size though, if you need the best you can manage with one GPU. And the 35B model is fast even when you put it on crap hardware.
>>109559856yeah with it running slow as molasses and falling into thinking loops
>>109560054why though? Isn't this MoE?
>>109560086>Isn't this MoE?No
>>109559988>although I'll admit models that can run on conventional hardware have more value sometimes>Qwen 3.6 35B A3B>And the 35B model is fast even when you put it on crap hardwareRuns at 15t/s on just CPU without jewery services spying on you or leaving you with <fill in the blank yourself, deadbeat> answers. What more do you need? Its prefectly serviable if you want a local search engine and error corrector. Perhaps not so much if you're in the "muh heckin ai is going to take our jerbs" scene that demands utter useless babble derived from hundreds of hours of parallelized compute time against a 5TB dataset. But still, good enough as a less retarded clippy.
>>109560100oh well, fuck
>>109545779>>109546025>>109546287>1.38 t/skek
>>109560164vs >>109560102 15t/s on Qwen 3.6 35B A3B Q8 XL?
>>109560180Yeah, 35B-A3B is wayyy faster. about 10-15t/s on the same machine (q8_k_xl). I'm just using q4 here because I'm experimenting with different quants but smaller quants dont seem to speed up by enough to care about.I've been trying KAT-Coder-V2.5-Dev and Qwen-AgentWorld too which are all based off qwen3.6 35b-a3b. the MoEs are far more practical for my hardware.also qwen 3.8 basically requires setting "low" reasoning effort, or otherwise it just wastes 6 gorillion tokens going in loops
>>109560102It depends on what you're doing with it. I have 35B A3B running a RAG on okay hardware and it's real good at that. But try and do heavier work with it, such as running an agent, and it flounders. The speed advantage makes it easy to use, but not always good to use. You hit the nail on the head: "less retarded clippy." Light all-around work is where it shines.
remember to use the fixed templates https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
Can you changshills stop hyping up these dogshit models for a while?
>>109544911Nope. Chuck Testa.
>>109544900BasedCodetrans will seethe tho
>>109555740If you can a claw up measly $600, it runs at 40t/s on a V100 32GB.
What's so great about Qwen? I get better results out of Mistral