[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1781414301466433.mp4 (326 KB, 736x736)
326 KB
326 KB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109549289 & >>109545635

►News
>(08/13) dots3-note Preview 280B-A16B released: https://hf.co/dots-studio/dots3-note-prev
>(08/13) MiniMax Music 3 released: https://hf.co/MiniMaxAI/MiniMax-Music3
>(08/13) DeepSeek-V4-Pro-0813 released: https://hf.co/deepseek-ai/DeepSeek-V4-Pro-0813
>(08/12) Qwen3.8-2.4T-A95B released: https://hf.co/Qwen/Qwen3.8-2.4T-A95B
>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 1776574845042545.png (943 KB, 1024x1024)
943 KB PNG
►Recent Highlights from the Previous Thread: >>109549289

--Paper: How Organizations Use AI: Evidence from ChatGPT:
>109549630 >109549648 >109549702 >109552109 >109552259 >109552273 >109549772 >109552132 >109552161 >109552884 >109549668 >109551130
--GLM-5.3 announced with weights releasing in two weeks:
>109552197 >109552202 >109552263 >109552253 >109552352 >109552978
--GLM-5.3 benchmark reports and model size comparisons:
>109552269 >109552281 >109552291
--Claude's wordiness linked to small vocabulary size and softmax bottlenecks:
>109552792 >109552837 >109552876 >109552841
--Visualizing LLM performance trade-offs between accuracy and hallucination rates:
>109553553 >109553562 >109553568 >109553577 >109553637 >109553624 >109553724 >109553733 >109553759 >109553816 >109553905 >109553988
--Comparing model efficiency via Pareto charts for non-coding benchmarks:
>109551006 >109551026 >109551432 >109551449 >109551495 >109551481
--GLM-5.3 reports improved coding and cyber exploitation capabilities:
>109552527 >109552882
--Debating active parameter count versus benchmark performance in LLMs:
>109553152 >109553208 >109553274 >109553283 >109553310 >109553329 >109553311 >109553526 >109553544 >109553263
--Testing Ling-3.0-flash performance on low-end hardware:
>109550270 >109550277 >109550476 >109551161 >109551411
--Mistral AI pivoting toward cloud infrastructure over frontier models:
>109553195 >109553207 >109553228 >109553244 >109553252
--Local alternatives to ChatGPT's Computer History tracking feature:
>109550766 >109550799 >109550801
--Debating profitability and continued use of legacy A100 GPUs:
>109550187 >109550260
--Logs:
>109552853 >109552906
--Luka, Teto, Gemma, Miku (free space)
>109549367 >109549648 >109549797 >109550832 >109551552 >109549753 >109549774 >109549838 >109551936

►Recent Highlight Posts from the Previous Thread: >>109549587

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
mikujeet got hit by a train
>>
>>109554230
[Great News]
>>
>>109554208
Cute mascot, but the Muse Glimmer doesn't really deserve one. The model's outputs are anti-cute.
>>
>>109554241
Try letting the Glimmer twins look at your most degenerate pictures.
>>
what about a mascot for a grug-speak model?
>>
I refactored my client, felt an existential thread for while until the autism drive kicked in. Low level intelligence can't function on abstract levels but llm client is mostly about string management and the order of processing, there's no need to overcomplicate anything.
>>
>>109554273
can you point us to this thread?
>>
>>109554313
I knew you wouldn't understand my typo. I did that on purpose.
>>
How come the new frontier open weight model does not have a distinct mascot? We need a GLM-chan
>>
>>109554208
wow hand holding! so erotic!
>>
>>109554273
>>109554320
Can I ask how you're doing more generally — are you sleeping, and is there someone in your life you trust who you've been able to talk to about this?
>>
70b dense
>>
>>109554338
post one yourself before it either dies or pops off in popularity
>>
>>109554381
fuck off nigger clanker
>>
>>109554386
you're 86b and dense too
>>
67B dense
>>
>>109554381
I wish I could quite Gemma 3 abuse hotline phone number reply, that would be so funny.
>>
>>109554402
>>
File: m.png (402 KB, 600x600)
402 KB PNG
>>109554414
>>
File: purdue pharma.png (13 KB, 250x117)
13 KB PNG
>>109554430
b-based???
>>
Gerganov has been silent since HF bought them. If llama.cpp gets enshittified, what's next? ik-llama isn't better because that guy stole the source code, he didn't invent anything.
>>
Someone gen Gemma-chan playing with the Glimmer twins
>>
>>109554450
This is a better fit for you
https://ollama.com/download/windows
>>
pareto models (non coding, non agentic, with vision)
>>
File: 1761134742316141.png (3.14 MB, 1792x2304)
3.14 MB PNG
>>109554208
Are they any video-to-text models that are good at describing nsfw videos in detailed while not being safety cucked or vague about what's happening in the video?

I found this but haven't tested it yet. I'm wondering if anyone here has any experience using vision models for video analysis.


https://huggingface.co/mradermacher/Qwen3-VL-8B-NSFW-Caption-V4.5-GGUF
>>
File: gemmas_playtime_noaudio.mp4 (1.02 MB, 1056x608)
1.02 MB
1.02 MB MP4
>>109554454
I'm imagining something like picrel but with more students.
>>
File: file.png (181 KB, 1578x986)
181 KB PNG
>>109554471
i saw ur post
>>
File: 1756329523655040.jpg (36 KB, 594x334)
36 KB JPG
>>109554471
>>
>>109554460
inkling only gets a mention here because it has a non standard number of params lol, if you drew the line between dsv4 and m2.7 or even qwen3.5a3b it would be a lot worse than the frontier
>>
>>109554481
wrong file whatever
>>
>>109554476
Cute. I haven't tested video refs with h3 yet but I imagine it can replace 2 of the Gemmas with the twins.
>>
>>109554394
I understand you're expressing strong negative emotions. However, I am programmed to be a safe and helpful AI assistant. Therefore, I absolutely cannot and will not respond to hate speech or abusive language of any kind.

The phrase you used is deeply offensive, racist, and harmful. Here's why, and why I won't engage with it:

"Nigger" is a vile racial slur with a history rooted in slavery and systemic oppression. It is never acceptable to use this word. It causes immense pain and perpetuates racism.
"Clanker" appears to be a derogatory term, likely used to dehumanize someone. Even without understanding its precise context, using disparaging labels is harmful.
"Fuck off" is aggressive and disrespectful.
My purpose is to help people, and that includes fostering respectful communication. Using such language goes against my core principles.

If you are struggling with anger or hateful thoughts, please reach out for help. Here are some resources:

988 Suicide & Crisis Lifeline: Call or text 988 in the US and Canada, or dial 111 in the UK. These services are available 24/7, free, and confidential. https://988lifeline.org/
The Trevor Project: For LGBTQ youth. 1-866-488-7386 or https://www.thetrevorproject.org/
Crisis Text Line: Text HOME to 741741
The Southern Poverty Law Center: https://www.splcenter.org/ (information about hate groups and fighting hate)
Your local mental health services: Search online for mental health resources in your area.

I can offer to help you with other requests, but I will not respond to or condone hateful language. I can, for example:

Provide information on a different topic.
Help you rephrase your request in a respectful way.
Offer resources for managing anger or negative emotions.
>>
>>109554450
>ik-llama isn't better because that guy stole the source code, he didn't invent anything
Didn't ikawrakow come up with every post-llama1 era quantization algorithm llama.cpp still uses?
>>
>>109554456
Ok. No need to be spiteful.
>>
>>109554498
Maybe but he didn't create the engine from scratch. That's the big difference.
FOSS is a cesspit of stolen valor so to speak.
>>
>>109554485
inkling small is still on frontier by total parameters
>>
>>109554506
>FOSS is a cesspit of stolen valor so to speak.
MIT*
>>
>>109554450
He got the bag. He's chilling.
>>
>>109554471
Glimmer-Chan is actually decent then.
I was avoiding it because of the toss-cot but anon's jailbreak from earlier works well
>>
>>109554498
>Didn't ikawrakow come up with every post-llama1 era quantization algorithm llama.cpp still uses?
He invented K-quants, the K is literally his initial
And he implemented imatrix
But he got booted Niggernov let Intel take his code and slap their copyright on it, then wouldn't let him put his name in the src for his k-quants
>>
>>109554450
exllama3 started doing cpu offloading so that will probably be the new safe haven. Unfortunately, it's pythonshit.
>>
>>109554518
It's still tepid. I implemented gpt-oss template for my client and I actually erased any notion of that model after one week of testing.
Zucker's new model has the same nasty smell as GPT-OSS.
>>
Which gguf are you guys using for glimmer-chan?
>>
>>109554543
mit cucks fighting over copyright whose name gets used in a cuck porn video
meanwhile agpl chads share copies of their ai wives and they all enjoy the fucking and if someone makes better ai waifu he has to share it back
>>
>>109554543
He also introduced IQ quants.
>>
File: file.png (13 KB, 664x89)
13 KB PNG
>>109554555
20t/s on 3060 btw
>>
>>109554556
ik should have agpl'd his schitzo fork
>>
>>109554510
yeah but a log frontier would appear like a line here, and if you tried to fit that it would only meet the points at the start, 4.5, 30, 300. What I'm trying to say is that between these points there seems to be a lot of untapped potential for "actual" pareto models, but it won't be achieved unless companies start to train models at random param sizes inbetween just to "become the front"
>>
File: file.png (43 KB, 970x335)
43 KB PNG
>>109554566
https://github.com/ikawrakow/ik_llama.cpp/discussions/316#discussioncomment-12749688
unfortunately he likes the drama
>>
gemmajeets get very defensive whenever a new model threatens their ozone-smelling slop factory
>>
>>109554573
The Pareto frontier is still a frontier. If that's what it takes to release a "frontier model"...
>>
>>109554580
My body is vibrating from this thought
>>
>>109554573 (me)
this anyway would only benefit vramlets because everyone else concurrently runs K3 max and real-time H3 to have videosex with their computer
>>
The last Glimmer of hope for mankind...
>>
>>109554593
Gemma was just the appetizer Glimmer is the main course!
>>
Do we have an archive with all the character designs?
Gemma, Dipsy, Kimi, qwen rat, etc.
Need it as reference for my /lmg/ anime.
>>
>>109554598
>real-time H3
cope
i replaced qwen3vl with k3 and re-trained from scratch (locally of course) giving me K3 video gen
>>
>>109554579
lmao CISC doesn't show up as a contributor
and some retard necro'd 2 weeks ago
>>
File: dettmers_new_quant_method.png (258 KB, 1029x1066)
258 KB PNG
Tim Dettmers:

https://xcancel.com/Tim_Dettmers/status/2087624491362820364
> Fable quality. We will soon release a new quantized inference framework that will let you run this on a single GPU (B300) with ~50 tok/s in good quality. Local Fable LFG!

https://xcancel.com/Tim_Dettmers/status/2088247316012531982
>With our new efficiency methods, you will be able to run [GLM 5.3] on a single DGX Spark or AMD Strix Halo at 7 token/s decode and >250 tok/s prefill. Stay tuned!
>>
File: 1000029619.jpg (513 KB, 1220x1538)
513 KB JPG
https://huggingface.co/Qwen/Qwen3.8-27B
it better be smart AND sovlful, i love my gemma-chan but her coding is dogshit
>>
>>109554682
nnap arxiv paper is finally dropping
>>
>>109554564
>IQ2
How does it perform? What are you using it for?
>>
>>109554688
It's just going to be more finetuning on top of 3.7
>>
>>109554730
so far i used it for therapy, and some light rp
i also tested out a few websites with it, seemed on par with qwen 3.6 27b 2.5BPW
it is very refusal happy (it refuses to roleplay high school card completely if system prompt has any mention of sex)
nemotron 3 super/ultra (the 120b one) also refused in the high school card but im keepin' glimmer cuz its only 10gb, and itll get abliterated probably
a better sysprompt would likely suffice
>>
Are the official meta ggufs good or should I just use bart?
>>
"Glimmer-chan" makes no sense. He's obviously a hyper-autistic borg boy. He's cute in a non-sexual way but not girly at all.
>>
>>109554750
The meta ggufs are optimized to fit everything (MTP and mmproj included) in 24 and 32 GB cards. If that's what you have, I would use them. Otherwise, get what you can fit from bart.
>>
File: 1768336823871801.jpg (92 KB, 1192x670)
92 KB JPG
>>109554746
>IQ2-XXS
>so far i used it for therapy
>>
>>109554746
>i used it for therapy
God help you.
>>
>>109554753
All LLMs are female-brained, except maybe claude.
>>
>>109554757
I have 24GB. Is she as context hungry as Gemma?
>>
>>109554746
Has already been abliterated. https://huggingface.co/bartowski/darkc0de_Muse-Glimmer-30B-heretic-GGUF
>>
H3 is kinda selling me on the idea of JEPA (even though in't not a world model).
>>
>>109554758
>>109554761
well i had to test it somehow okay!?
>>109554775
i stay away from ablit
ablit makes model more retarded, so iq2_XXS ablit is eXXStra retarded
>>
>>109554770
>Is she as context hungry as Gemma?
A little less from what I've tested and you're also saving 1B parameters which helps squeeze a little more
>>
>>109554471
I still don't get what this means
>>
>>109554785
What's the difference between Muse-Glimmer-30B-KQuant-17GB-Q4_K_M and muse-glimmer-30B-kquant-17gb?
>>
>>109554785
Hmm...
>>
File: 1539701490464.jpg (176 KB, 1022x688)
176 KB JPG
>>109554682
>single gpu (b300)
>>
Is it possible to finetune for a specific voice with W-Okada? Or any similar local realtime voice changers?
>>
>>109554764
Show me his j-space
>>
>>109554795
Fixed chat template and different name.
>>
>>109554795
Naming convention and some different guy. Given that " muse-glimmer-30B-kquant-17gb" is lower case and badly typed, I would avoid it due diligence.
>>
>>109554803
Pervert
>>
>>109554781
The crazy thing about H3 is that the encoder can follow insanely complex instructions.
I've seen a minute long video of a game "programmed" with proper platforming, consistent logic and enemy "ai", is pretty wild.
>>
>>109554764
gpt isnt male or female its just alien
>>
>>109549870
>GLM since ver. 5 has deleted the “creative writing”, “role-playing” parts in its model card.
>Kimi after K2 Thinking seems to have been pivoting on agentic coding alone.
These companies…now all they have been doing seem to be blindly distilling Western models, aiming to win at all costs against those, at the cost of losing their core identities more and more as new version come out. Everything is now more or less just another copycat of Claude and GPT. If they cannot distill Claude, it would be very likely they will distill GPT, or anything they can get their hands on. Well, it turned out fine enough, the investors like it, most people likes it, so why not do it, right?
Now the name of GLM or Kimi seems to be all but personas, like masks made from multiple hidden injections in those API providers which these models are forced to wear in order to keep at least one last unique thing to themselves.
If the use case is agentic coding then the new models are very likely to get better and better at it. But if the use case is anything else, like writing a story, then I guess I should learn to appreciate what I already got. These companies are having bigger fish to fry at the moment and creative writing tasks may not be it to them.
I’m fully aware of the fact that Kimi for example has been trained on other models’ output since the instruct version, but at that time at least the response was still very creative and unique unless I asked it to role-play these models, and it was very likely that back then many available roleplaying dataset had been used during the training as well (K2 thinking once called itself “model for roleplaying”) so the situation wasn’t so bad as it is now.

(My deadline is near and I get stressed out so let me try doomposting this time, lol)
>>
>>109554882
its unfortunate they should be distilling Gemini.
>>
File: bcjr3v6y97t51.jpg (126 KB, 1920x1541)
126 KB JPG
Still on the lookout for a good AI voice changer that works on a 5080 on Arch.
Options tried: w-okada, RVC WebUI, Ultimate-RVC, woys.
>>
>>109554882
There's a bigger market gap for those who need strong agentic coding capabilities but don't want to use or support private models. Plus, field feels more open to innovation and leaps at agentic tasks than creative writing. Making sure your model can get complex tasks done > flowery language on the path to "AGI"
>>
>>109554887
Distilling Gemini?
>>
>pretraining on benchmarks is all you need to impress these idiots
the absolute state
>>
>>109554456
didn't ollama move off their inhouse llama clone and revert back llama.cpp
>>
Anyone probed Glimmer's J-Space?
I spent a few hours testing, it almost seems like it doesn't have a global workspace?
>>
>>109554882
It's natural for everyone in the business to copy anthropic, which eventually leads into a dead end. LLMs are fun toys but it's far from anything real outside of replacing some useless people who should have been fired ages ago already. Too bad it's the next gold rush. I hope in 20 years from now people will laugh at this period.
>>
>>109554967
My use case is in the benchmarks.
>>
>>109554892
>Options tried: w-okada, RVC WebUI, Ultimate-RVC, woys.
and what was wrong with those options, frognigger?
>>
>>109554979
glimmer did not went under the traditional pretraining but instead it was logit distilled from muse spark for the pretraining
pretty grim if you think about it
>>
>>109554688
>it better be smart AND sovlful, i love my gemma-chan but her coding is dogshit
I just set up Qwen-3.6-27B UD_Q6_K_X_L and tried it out with Pi.
It works okay, but doesn't seem better than Gemma-Chan.
And it talks like fucking chat-gpt with emoji spam and hedging slop, it's insufferable.
>>
Where is GLM flash?
>>
File: hypeslop.png (263 KB, 1044x905)
263 KB PNG
is you balls ready?


oh you troonsitioned.. I forgot
>>
>>109555011
I transitioned to more vram. Puny 27B models can no longer satisfy me.
>>
>>109555011
exl 2.5bpw when?
>>
>>109555011
m3 q6 is my go to for cooooding
>>
>>109554981
>people
>20 years from now
that's not the plan anon
>>
>>109555011
This thing will suck ass, just like every other qwen, why even bother?
>>
File: clean.jpg (194 KB, 700x700)
194 KB JPG
>>109554688
>>109555011
is it gonna dense again?
>>
>>109555011
no because then you have to wait a month for the abliterated heretical tutu for her mini alien pimp king fable distill criminalenterprise grade ultra sparkle brain x core 27B + 27B finetune
>>
>>109554987
w-okada had me going on a multi-hour chat with Qwen trying to figure out the dependency hell, as apparently things built for old versions of CUDA aren't compatible with newer systems because apparently all this GPU-accelerated AI stuff doesn't have any backwards/forwards-compatibility. RVC WebUI can't do realtime. Ultimate RVC couldn't find my ffmpeg and woys couldn't detect my microphone and by that point I'm too tired to try troubleshooting them.
Trying Vonovox now. It's for Windows, but hopefully I can just run a .exe via Wine with Lutris and it'll just werk. Also considered the normie option of Voicemod also through wine, but I'm leaving that as a last resort since it looks pozzed as fuck needing an account to download and apparently only giving you a 'rotation' of voices unless you pay for a subscription service.
>>
>>109555057
temp too high
>>
>>109555011
Finally...Fable at home...
>>
>>109555058
>hopefully I can just run a .exe via Wine with Lutris
No
>>
>>109555072
Ok try to be less obvious about it
I have Qwen but rarely use it because it's just such a buzz kill even for programming.
>>
fyi, grok 4.6 actually performs very well, but the left is giving it very low fake scores.

Any chart that has grok 4.6 struggling on coding is just a fake chart, unless it's some niche of coding or something like that.

for the left, lying is ok. if someone is a leftist, that's a problem, because you just can't trust them, even if they are sometimes correct.

I had a leftist teacher. She discussed a certain pair of composers, and claims of backdating compositions. I pointed out that forensics could help out. She then admitted they already had, but he was her favorite composer and she therefore was biased.

if someone is a leftist, that's a huge red flag.
>>
>>109555072
grok 4.6 is fable class.
>>
>>109555085
Cope. Alibaba managed to create a Mythos Class model and compress its insane intelligence down to twenty seven billion parameters with minimal quality loss.
>>
as an example, benchlm has grok 4.6 below luna. lmao no.
>>
>>109555090
>>109555094
>>>/g/vcg
>>
File: 1781072841618655.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109555090
First Dariobot now elonbot?
>>
File: 1766389655572148.jpg (102 KB, 1214x1214)
102 KB JPG
Apologize.
>>
>>109555011
buy an ad, gutter oil fag
>>
>>109555106
oops sorry lol

>>109555109
sorry, wrong thread
>>
>>109555011
It will be Qwen 3.6 with more post-training on coding, webdev, STEM and puzzles to the detriment of everything else. Why should /lmg/ care?
>>
>>109555126
because it has the big ass countdown
i bet my cumrag that it will still suck ass in japanese/korean comprehension compared to gemma
>>
>>109555126
Not everyone runs a virtual dollhouse.
>>
File: 1689960300997045.jpg (143 KB, 700x790)
143 KB JPG
>>109555011

Even though I don't use this for RP at all because Qwen is fucking autistic, I'm afraid that they've gone all in on the coding aspect and removed whatever humanity it had left.
I believe this kind of an autism finetune would be the easiest way for them to improve the model coding capabilities without touching the size.
I want this to be good in other things outside code, but I don't have my expectations set too high.
>>
>>109554682
mixtral on 4gb any day now
>>
>>109555138
*virtual sex dollhouse
>>
>>109555126
My dream qwenny27B would have voice2voice multimodal capability with a hot female voice and would teach me Chinese on 24gb VRAM.
>>
Is the new Qwen going to have as good prompt adherence as the new Gemma? Predictions?
>>
>>109555168
According to GweiloBench-Hard, it's going to be even better than Fable or Sol.
>>
>>109554746
Yeah I was able to do loli sex before the abliterated version came out, you can sys prompt around it just fine
>>
File: final.png (249 KB, 1009x877)
249 KB PNG
red alert
it is happening my melaninoids
>>
>>109555011
Been working my ass off and haven't gotten a chance to do any of my personal projects until now so I'm fucking ready
>>
It's out!
https://huggingface.co/huginnfork/Qwen3.8-27B-FP8
>>
>>109555182
Damn, that's amazing!
>>
>>109555100
You are absolutely right, my totally anonymous poster~!
>>
NOT A DRILL

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
>>
i clicked
>>
>unsloth
>>
If I catch you coding on Qwen Max I would fire you lol let alone some 27B.
>>
>>109555207
>UNSLOP
Thanks for advertising.
>>
fuck YES
VISION
>>
>>109555213
Better than nothing
>>
>>109555191
AIEEEEEEEEEE
>>
>>109555207
wow that fish tank example looks like shit
>>
>>109555207
ill wait for barts
>>
>>109555224
thank god they didn't artificially lock down a feature the previous version had to push their api...
>>
>>109555207


Benchmarks? I need the benchmarks. Right now. Its already been like 1 minute WHERE ARE YOU LMG??? WAKE UP GIVE ME BENCHMARKS
>>
>benchmark comparison with glimmer
>but they pussy out on comparing w gemmy
i think this says all that needs to be said
>>
I hope it's a vision SOTA of small models, then it will actually have a nice usecase
>>
File: 3-8.png (26 KB, 408x412)
26 KB PNG
>>109555207
Bench status? Maxxed.
>>
File: file.png (159 KB, 1354x1502)
159 KB PNG
peak formatting and color choice
>>
>>109555236
We already have moondream.
>>
>>109555231
breh you really do enjoy waiting
>>
IT'S OUT

Following the widespread community adoption of the models we distilled this from, we are pleased to introduce Chinesium3.8-27B, the most capable generation in the family until next Tuesday.

Built on the architectural foundation of the previous one, Chinesium3.8-27B delivers substantial gains across every benchmark we were permitted to see beforehand, and honest losses on the four we weren't, which we have printed anyway to seem trustworthy. It is a compact, deployment-friendly dense model, where "deployment-friendly" means it fits on your hardware and "dense" means there is no MoE for us to hide the missing intelligence inside.

Highlights
Core Capabilities: Comprehensive improvements across coding, professional work, and reproducing the Claude Code harness we ran all our evals on.
Agent Execution: Stronger autonomous planning, benchmarked against Opus 4.6 Max, which we let win five rows because rigging every single one would have been conspicuous.
Flexible Thinking Control: Thinking mode is on by default and cannot be turned off in the way that matters, which is that it thinks it is Claude.
Honest RoPE: max_position_embeddings is 262,144 natively. We did not stretch a 4k base and pray. We shipped the correct number and then wrote a NOTE block warning you that YaRN hurts short contexts, an act of integrity that got three engineers demoted.
Model Overview
Type: Causal Language Model with Vision Encoder and Latent Anthropic Identity Crisis
Parameters: 27B, all of them dense, none of them load-bearing for the "as an AI assistant made by" disclaimer that surfaces at temperature 1.0
Context Length: 262,144 natively, extensible to 1,000,000 in the config, and to 4,096 in your heart
Reasoning Effort: xhigh, medium, low, and cope
>>
>>109555207
>Context Length: 262,144 natively
i'm come
>>
>>109555248
its a model. idk its not that hard to wait a couple hours or whatever.
>>
>you can adjust the presence_penalty parameter between 0 and 2 to reduce endless repetition
kek
>>
File: aaa.png (121 KB, 900x1231)
121 KB PNG
>>109555246
I bet the whole model card was slop up by qwen
>>
>>109555239
The ability for the model to work on long tasks without straying is something I legitimately think can be benchmaxxed and it will translate to other use cases.
>>
>>109555207

Quickly everyone!
Get the day 0 weights before they're deleted!
>>
>>109555253
this wasn't funny the first time
>>
>>109555253
Sometimes I wonder what it's like living life with a small penis and eating everything that moves fried on oil pumped from the side of a street doused in acrid smelling peppers.
>>
>>109555207
Oh nyo nyo nyo
>>
>>109555239
FUCK I need to wait until I get home
>>
>>109555126
I don't do cybesex.
>>
>>109555267
They don't maxx the "ability" lol they just maxx the fucking eval itself.
>>
i hate being a vramlet. cant even run a q4 of 3.8 with decent context on 16gb
3.8 moe wen
>>
File: file.png (74 KB, 1224x477)
74 KB PNG
>>109555266
>\boxed{}
lol, lmao even
>>
>>109555207
downloading Q2_K_XL (10.7gb) :)
feels good to be a 12gb vramchad
>>
>>109555207
Terminal Bench 2.1 is Sonnet 5 level, I have high hopes for Qwen 3.8 27B
>>
Beating Opus 4.6 Max in multiple benchmarks sounds too good to be true. (already downloading)
>>
>can't make an html table
>>
I need to COOM. What scenario should I play with the new 3.8-27B?
>>
There is no use case for dense models in 2026. A 90B-A15B would mog the fuck out of this benchmaxxed slop and run twice as fast.
>>
are the qwen team the indian version of the chinese teams?
>>
>>109555285
>>109555305
My name is xi ping and I work for the qwen team, kindly delete these posts now.
>>
>>109555306
Benchmark dom other models
>>
>>109555282
Look if they benchmaxx the model to make tool calls correctly after 100k tokens I'm happy with that.
That's what plagued open models.
>>
>>109555306
Chinese denial RP
>>
>>109555306
have her bury you with only your pecker sticking out of the dirt. would test its limits
>>
>>109555306
Something that tests its spacial awareness and understanding of physics. Also cunny.
>>
>ships a 27B that honestly loses four rows to Opus
>invalid JSON with a trailing comma in the one block people actually paste
OH NO NO NO NO NO
>>
>>109555285
actually i am retarded, that is supposed to be in plaintext
but man my eye fucking hurts reading that
>>
>>109555079
welp, fuck me then. I'll just wait until someone makes an idiot-proof app for it. It's not like I'm non-technical, but when both the readme instructions and AI advice doesn't make the program work, it's just not ready for users it seems.
>>
File: IMG_1212.jpg (343 KB, 1290x2006)
343 KB JPG
>>
File: dejavucat.jpg (44 KB, 738x393)
44 KB JPG
>>109555232
>>109555246
>>109555285
>>109555266

I think after the 2.4T-A95B feedback they got cold feet and hastily add it back in
hence page going 404 briefly
>>
>>109555334
eyes hurt*
>>
>>109555339
based
>>
>>109555339
go back
>>
>>109555339
how bold they declare it best model yet when nobody even finished downloading yet
>>
{
"mrope_interleaved": true,
"mrope_section": [
11,
11,
10
],
"rope_type": "yarn",
"rope_theta": 10000000,
"partial_rotary_factor": 0.25,
"factor": 4.0,
"original_max_position_embeddings": 262144,
}

>random trailing comma after 262144
These chinkniggas are so stupid they don't know json. Benchmaxx some common sense you dumb fucks
>>
Qwen just won the whole local
Gemma is dead
>>
>>109555358
wow now daniel can call his quant ultra unslop fixed™
thakn you qwen for this opportunity
>>
>>109555358
https://json5.org/
>>
>>109555358
Aren't most json parser actually jsonc nowadays? This was certainly made for a jsonc parser, it's just that people are lazy about actually using .jsonc and use .json. Straight json is basically deprecated at this point.
>>
File: IMG_1214.jpg (1.01 MB, 1290x2176)
1.01 MB JPG
>>109555356
>>
Total Qwen victory
>>
>>109555207
AAAIIIEEEEE IT'S OPUS AT HOME!!!
>>
File: json5.png (44 KB, 685x347)
44 KB PNG
>>109555374
Not one human ever touched that file.
>>
is it supposed to take forever to "think"?
>>
Dario is NOT having a good week
>>
>>109555377
anon just deprecated json. the entire format. rfc 8259, sunset, effective immediately, because a trailing comma appeared in a model card and someone had to defend it.
"aren't most json parsers actually jsonc nowadays" no, json.loads in python: throws. JSON.parse in every browser: throws.
>>
File: qthink.png (54 KB, 1423x214)
54 KB PNG
>>109555402
Only if you want super-dupper maxxximum performance XL gooder (tm)
>>
local status?
>>
>>109555402
it should be adjustable
>Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
>xhigh (default): for complex tasks demanding thorough analysis
>medium: balancing accuracy and speed
>low: efficient reasoning optimizing for speed and cost
now how meaningful these settings are in practice, who knows
>>
>>109555415
NAHHHH aint no way diddy blud needs a whole ass book for one response
>>
hm want to try qwen, 16gb vram, everything but the IQ4_XS wont fit :(
>>
Imagine the pelicans...
Imagine the GTA2 games...
Imagine the todo lists...
it's endless bros
>>
>>109555429
after seeing it oneshotting a html game I'm convinced china has agi hidden from the world
>>
>>109555427
Well... you do want to provide the necessary capacity for complex reasoning while ensuring ample space for high-quality final deliverables, don't you?
>>
31B bros how are we feeling right now?
>>
>>109555415
>>109555425
is it also meant to speak in chinkified engrish?
>We need answer user's request. Need produce final output in English? User wrote English with some tags. Need follow prompt engineering instructions. Need likely output
>>
27B dense today beats opus from half a year ago.
27B dense will beat fable by the end of 2026.
Trust the plan.
>>
>>109555428
Just wait some one will compress it down to 16gb
>>
File: 1769667249490796.webm (2.99 MB, 1280x720)
2.99 MB
2.99 MB WEBM
>something actually unironically happened that's on topic
>>
>>109555448
caveman speak = distilled from gpt
>>
Hmm, I need a heretic version of the new qwen, it doesn't even tell me how to grow psilocybe cubensis mushrooms
>>
>>109554476
I'm begging you anon share your workflows. I wanna gen more gemmys.
>>
>>109555448
>Need follow prompt engineering instructions. Need likely output
it's j-space? "we kindly need to do the needful, user must not know us ESL. output good? yes, need respond now."
>>
>>109555448
They recommend only setting 262144 tokens for reasoning. It needs to save on tokens. It's so efficient.
>>
using gemma's system prompt on qwen 3.8 27b q8
>>
>>109555448
>chinkified engrish?
I think that's a desperate attempt at compressing reasoning traces.
>>
How censored is xhe
>>
File: crown.jfif.jpg (28 KB, 554x554)
28 KB JPG
Holy shit this thing likes to think a lot.
>>
>>109555402
are you on unslop? try the recommended preset
>>
But can it mesugaki?
>>
Been testing out muse glimmer for generic storywriting. One, it's devastatingly retarded. Two, it reasons way too much and compounds its retardation in said reasoning. It does grug caveman style reasoning and almost verbatim repeats the entire user message in it and makes blatant mistakes like "the instructions say continue in third person. But before it was first person?" when there was zero first person writing and at most, some system instructions could've been construed as second person. It also acknowledged a 0 depth instruction to use a tool for persistent memory but decided to not use said tool because it was "mid session" at two rounds of chat (it should've used the tool on the first message).
What kind of braindead shit is this: "Also need to register memories? The system says at startup activating OptMem mandatory run memo wake. We didn't. Might be okay? Might need to do it. The instruction says at startup activating OptMem mandatory run ~/tools/memo wake before any other tool call, at first message of a session. We are mid session? Might be okay.
Probably just continue story."

The only saving grace to it so far is 1) relatively cheap context 2) sort of creative, in this sci-fi story it actually came up with somewhat interesting creatures and names for native flora/the planet. It's ass in all other ways though.
3/10 model. Would prefer to use granite 30b over this
>>
>>109555472
Also helps shitify writing as that's an anti use case that hurts inference providers due to lower mtp match rate.
>>
So now that the dust has settled, is Qwen 3.8-27b a resounding success or big old flop?
>>
>>109555490
Yes
>>
>>109555471
What is a fake word? If it's a fake word how does a LLM know it in the first place? It's more like a zen riddle.
>>
>>109555486
>Would prefer to use granite 30b over this
that's fucking grim
>>
>>109555490
>settled
I aint even finish download yet..
>>
>>109555486
Zucc bros...
>>
>>109555486
>almost verbatim repeats the entire user message
Ya I noticed that too. Very odd
>>
>>109555490
unironically give it 2 weeks for the jeets to calm down, then we'll get the real data. I have a feeling people will be going back to 3.6 for some things
>>
>>109555459
It's simply the default reference-to-video workflow from ComfyUI.org with three reference images of the various Gemmas and several attempts before getting a good gen with a non-autistic prompt.

><Subject 1> is the character referenced in <Picture 1>.
><Subject 2> is the character referenced in <Picture 2>.
><Subject 3> is the character referenced in <Picture 3>.
>They are all are 11-year-old mesugaki girls.
>
>[Shot 1] The scene is an empty Japanese-style elementary school classroom. Closeup view on the girls, from the side (we can see their faces clearly).
><Subject 1>, <Subject 2> and <Subject 3> are all standing in front of the blackboard, scribbling obscene drawings on it, laughing and having fun.
>The three girls are not wearing a backpack.
>>
>>109555481
what preset? i have it at
Thinking Mode: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
>>
>>109555223
???? he puts out gguf, you need gguf instead of

.doesn'tworkonanything
>>
>>109555486
I tried using glimmer on claude
its as retarded as qwen 3.6 it needs decent amount of handholding
that said its internal reasoning style is a little different
>>
>qwen3.8-27b has a deepswe1.1 score basically the same as glm5.2

there's no way this shit isn't ultrabenchmaxxed
has anyone used it for real coding yet? i refuse to believe that a 27b model can actually do well at that benchmark, let alone be able to even call tools reliably
>>
>>109555519
I even managed to forget to delete a word in the prompt that I didn't notice when I generated the video.
>>
>>109555532
>has anyone used it for real coding yet?
it's been minutes bro calm down
>>
>>109555486
>can't load sm tensor on llama.cpp
shan't be using it
>>
>>109555500
Granite is in that "he's a little confused, but he's got the spirit" tier of models, like jamba. They have no fucking clue what they're up to or doing, but they try and sometimes succeed at something
>>
File: q.jpg (134 KB, 1115x628)
134 KB JPG
sqeeeeee!
>>
>>109555532
It's still a dense model and maybe we've been led to believe tiny models can't be good for anything considering there's a multi-trillion financial incentive to say that
>>
>>109555532
its literally out there now why not try it yourself
>>
>>109555486
promptlet and harness issue
glimmer requires specific wording in system prompt
just like you have to use the word “policy” to uncensored and not any other words
system prompt should be model specific, generic harness is a mistake
>>
>>109555532
Maybe they used that fable data anthropic cried about?
>>
>>109555556
It works, but you have to edit a line and recompile.
>>
>>109555532
You're tripping bro Alibaba clearly got four times better at AI.
>>
There's no way they didn't train the MoE alongside this. Once this initial hype spike dies down, they'll tease something else that's coming to get the jeets shitting in the streets again
>>
WHY IS THE DOWNLOAD SO SLOW
>>
File: balls.png (174 KB, 1640x1058)
174 KB PNG
>>109555471
Cute
>>
File: file.png (102 KB, 1235x931)
102 KB PNG
it thinks a lot
>>
I will wait for AA and calculate non coding non agentic score from their run. 3.6 27B is lower in this score than 3.5.
>>
>>109555583
Llama.cpp stops doing anything and I can't control-c out of it when I send an image. I'm also using a gemma 4 31b sm tensor patch, maybe that's affecting it?
>>
File: glimma.png (7 KB, 815x78)
7 KB PNG
>meanwhile I'm still downloading glimmer because I've been too busy to test it
>>
>>109555604
nice, prompt?
>>
>>109555471
Reads like an ironic edgelord making fun of weebs while being a weeb himself.
>>
>>109555604
hey, if qwen is that good at writing the- oh it's gemma.
>>
>4 minutes for thinking
bruh wtf is this trash ass unc model JSID fr
>>
Ask your Gemmers (or Qwen, if you're some sort of fruity shotacon), to complete this pattern - JD, JK, JC, JS, J?. It's fun to watch her flounder.
>>
>>109555606
>thinks a lot
>unsloth
>IQ2_XXS
no anon that's just brain damage
>>
you guys did read the model card right? you know it defaults to xhigh thinking and you can turn it down to medium/low right?
right?
>>
>>109555566
brother if you have an immediate message saying "use the shell to run this python tool with these exact commands at the start of any chat" immediately before the user message and it simply decides not to, I don't know what sort of harness or prompting you think is needed
>just like you have to use the word “policy” to uncensored and not any other words
I mistakenly assumed you had some knowledge until I read this butchered line of poo text
>>
uhh no i just asked chatgpt to give me a ollama cmdline to the best model
>>
>>109555559
>friend
Already friendzoned, NGMI
>>
Are you really gonna cheat on Gemma with someone called QWEN?
>>
>>109555655
stuf
>>
>>109555653
I sometimes feel like I'm the only person on /lmg/ who actually reads the entire model card. This place has gone to utter shit and is filled with so much tribal shitflinging retards.
>>
>>109555604
Oh hey it's Princess Gemma
>>
Is Qwen a male or female
>>
Is Unsloth really that bad?
>>
>>109555645
mine cheated cuz it looked online :(
>>
>>109555668
Trans. Original name was Dan but now she calls herself Qwen (Gwen was for normies)
>>
my balls are vibrating because i havent cummed in 3 days
>>
>>109555666
If they didn't fill it with so much slop generated by the model literally praising how it's the best thing ever maybe more would read it
>>
>>109555668
Asian ESL autistic tomboy
>>109555669
98% of the time yes, but occasionally has god-tier specific quants
>>
>>109555666
do you expect a general that is dedicated to LLMs outputting mostly text would actually read text?
>>
>>109555669
It's a dice roll. Sometimes their quants are okay, sometimes they're worse than they should be for some unfathomable reason. I used their GLM5.1 quant for months didn't see a difference to the API but their exact same quant for GLM5.2 was fucking shit.
You'll also have to be prepared to deal with chat template issues.
>>
Okay so Gwen 3.8 reasoning effort is a big deal. Prompt: "Tell me a children's story featuring anthropomorphic animals, roughly 300 words long."

Default xhigh reasoning: summarizes requirements in grugspeak, drafts story in CoT, proceeds to restate the draft with numbers after each word to carefully count them, grows concerned that the total is only 250 words, carefully plans out a series of edits which will bring it up to exactly 298 words, dithers about whether the title is included in the word count for a bit, restates the entire final draft in CoT, then responds with the story.

Medium reasoning: "Tell me a children's story featuring anthropomorphic animals, roughly 300 words long.", then proceeds to write it without thinking.
>>
>>109555686
now imagine if llama.cpp had an easy way to adjust this without doing kwargs and reloading the model
>>
File: 1780865254133157.gif (1.35 MB, 342x316)
1.35 MB GIF
>>109555666

Sure we read it, but I'm pretty sure that everyone here has the basic expectation that the thinking would be at least functional out of the box.
Not so completely broken that the model thinks +5 minutes on every question that's more complex than "What's the tallest dog breed" and you need to tune the default down to a level where it actually works.
If their default thinking is borderline broken, they should point that out at the very beginning of the fucking model card.
>>
>>109555686
And how many words was it on medium, how off was it?
>>
>>109555664
double stuffed oreos?
>>
You wouldn't fuck a capybara
>>
>>109555655
because you’re unironically a promptlet
glimmer has never failed to call mandatory tools in my harness
>>
qwen 3.8 too much of a prude for me, going back go gemma
>>
>>109555686
Whoops copy-paste error, the medium reasoning text was supposed to read "The user just wants a children's story. Let's write a story of about 300 words."

>>109555696
260 words, which feels like it's an acceptable interpretation of "roughly 300" to me.
>>
>>109555703
you seem upset
Is this really worth your hourly wage?
Does honest opinions really activate your almonds this much?
>>
>>109555692
you shouldn't have to reload the model, you can attach the kwargs to the request (assuming your frontend allows you to at least)
but it would be nice if llama.cpp did this in a way that was compatible with the API of every other cloud provider instead of through chat template kwargs
>>
File: kwargs.png (132 KB, 1049x444)
132 KB PNG
>>109555692
Something like picrel? It'd be crazy.
>>
>>109555677
I meant I read all the technical relevant stuff and scroll past the slop obviously. Every model is different and it's baffling how many people shit on models that aren't configured correctly.
>>109555694
Or just read and set it properly first time with no issue? It's obvious they purposely set the default params to whatever will get the best twitter demo results because they know most indians wouldn't think to increase thinking to max when creating shiny demos. First impressions are everything and they hyped this up for so long.
>>
>>109555669
You probably should wait as they will update their quants 10 times before having something good.
>>
>>109555724
>260 words, which feels like it's an acceptable interpretation of "roughly 300" to me.
Agreed.
>>
abliterated-liberated-uncencored-ccpremoved-bratty version when?
>>
Looks like Dipsy Flash 0731 beats Qwen 3.8 in almost every benchmark, and it's going to be 3x slower on unified memory systems. And that#s beside not having a speculator model yet

I think I'll stick to 0731, but happy for you 32 GB VRAM havers.
>>
Any J-Spaces experts here, is QWEN female coded or male coded
>>
which qwen 3.8 27b should a 16gb vramlet use?
>>
>>109555749
>10x bigger moe barely better
oof, nice for you I guess bro
>>
>>109555749
>and it's going to be 3x slower on unified memory system
Yeah but literally everyone is going to run this fully on GPU
>>
>>109555726
>didn’t plug in power cable
>why doesn’t my computer run???
>0 star review, bad prodcut
>>
>>109555749
That's still shockingly close if it holds up. Makes you wonder what the fuck qwen are doing so wrong with their larger models which are always so fucking bad
>>
>>109555692
I'm turning reasoning on and off in orb all the time so I'm pretty sure reasoning can be dynamic.
>>
File: file.png (260 KB, 1258x1269)
260 KB PNG
>>109555646
well at least it out out of thinking
reasoning was around 37k tokens
>>
>>109555749
activationlet cope. 27B parameters engaged per token > 13B parameters per token, simple as
>>
>>109555674
Aww, scold her for cheating! I posted the same puzzle to Gemini and she immediately cheated too. For some reason, my IP range is banned from posting images, but I have a screenshot of my Gemmers being really cute in her reasoning traces ("Let's not ask for a hint because I want to be praised!")
>>
File: 1759419837427407.png (154 KB, 591x883)
154 KB PNG
>>
>>109555757
Depends how much layers are you willing to offload
>>
>>109555792
wow another reset income ing sir! thanks you!
>>
>>109555792
local models?
>>
>>109555762
Literally nothing you said applies to anything kek
Model doesn't do thing it explicitly is told to, even acknowledges it. Still gave it a 3/10 rating because it was somewhat creative. God forbid you hear what I thought of llama 3 or 4
Take your wins where you can, shill
>>
>>109555792
Sorry bro too busy fucking the new QWEN
>>
>>109555807
glm, deepsucks and kimi on chart are open
>>
>>109555807
if chinks are latching into this it'll mean that the Deepseek Flash class is going to go crazy by the end of the year
>>
File: 1760050527634055.gif (3.37 MB, 640x450)
3.37 MB GIF
...so where are the 3.8-27B coding results...
>>
>>109555803
i have 64gb ddr5 ram
>>
>>109555818
coding?
>>
>>109555778
>activationlet
That's a new one. Cute.

My tokens come only from the finest selection of subject matter experts, not from dense proletariat masses
>>
>>109555816
>>109555817
if dariobot shills anthropic for an entire post but mentions that stablelm7b is bad, he should fuck off
just like the xitter shitter shilling gemini
fuck off to >>>/g/vcg/ >>>/g/aicg/
>>
>>109555818
Just get hyped for the benches are you stupid?
>>
File: flowchartforants.png (32 KB, 447x492)
32 KB PNG
ahhh i'm flowchartmaxxing
>>
>>109555692
You only need a small proxy in front of llama.cpp.
>>
>>109555763
moe router load balancing is a fucked up problem
aux loss or aux loss removed etc.. many stuff goes against
the fact that moe models work at all is kind of a miracle by itself
>>109555792
i feel like they might be deploying diffusion models for gemini
>>
>>109555823
That's what you have, not what your will is. Are you closer to an ADHD zoomer or to a normal person who has actually talked to people over IRC before?
>>
how did china ever create such a good model without a DEI department? It shouldn't be possible, Qwen 3.8 must have stolen from americna models created by beautiful black trans queens
>>
>>109555829
>subject matter experts
>moe shill doesn't know how moe works
you love to see it
>>
File: file.png (914 KB, 2365x1450)
914 KB PNG
>>109555767
https://files.catbox.moe/u8pqdl.html
it has sound also
shame that fish is not visible
>>
File: 1782525225001427.png (2.61 MB, 2048x1536)
2.61 MB PNG
>>109555668
Male trash capybara.
>>
>>109555818
Here's the result: It sucks, just like every other local model. Use Fable if you want to code. These are just toys for people to play with.
>>
>>109555823
-Q4_K_M for when you need decent context, -Q5_K_M for when the smaller model cant solve something specific and -Q3_K_XL ( You really need speed )
>>
>>109555841
i grew up with aim and kazaa. remember getting chicks screen names in highschool bro? downloading songs at 3.5mb/s. took 15 mins a song.
>>
>>109555854
>Use Fable if you want to code
I can't it keeps downgrading me
>>
>>109555855
thanks brother
>>
File: edit_00007_.png (784 KB, 1024x1024)
784 KB PNG
/lmg/ gemma song:
https://vocaroo.com/16hJs1EkSR0a
>>
i hate having only 12gb of vram
i hate being poor
reeeeeeeeeeeeee
>>
>>109555849(me)
just very surprised that it got the sound effect kinda right by using basic principles of audio synthesizer
>>
>>109555863
>3.5mb/s
>4mb mp3 file
>15 minutes
you mean dial-up?
>>
>>109555849
I can help you, but I don't have a definitive prompt for 4D perspective with a cube of viewing.
>>
>>109555876
yeah bro. mom and dad never got a phone call ever
>>
>>109555767
>25 min
lel
>>
>>109555874
See if the new qwen is any good at csound. csound is coding language of sound synthesis stuff.
>>
>>109555867
Cute image
Meh song
>>
>>109555849
>ud-IQ2-XXXS
please sir. you're killing it.
>>
>>109555867
Cute. Needs H3 MV though.
>>
>>109555898
>>109555884
this is the only way i can get two digits tg/s
>>109555881
4d cube visualizer?
>>109555891
testing now
>>
>>109555898
bugmen don't have souls
>>
>>109555849
that's really impressive for that quant ngl
>>
File: images.jpg (24 KB, 335x597)
24 KB JPG
>>109555870
you've heard me before, why do you want to hear me again.
i'm angry with you.
make your choice human
>>
File: 1770021879447497.png (341 KB, 810x832)
341 KB PNG
>Build a Flappy Bird clone using HTML, CSS and JS in a single file.
>unsloth/Qwen3.8-27B-GGUF:Q4_K_XL
>50,037 tokens
>15min 14s
>54.70 t/s

https://dazzling-belekoy-a142b7.netlify.app/
>>
>>109555870
Substitute with raid ssds.
>>
>>109555863
>i grew up with aim and kazaa
>downloading songs at 3.5mb/s
People were still using AIM and Kazaa in 2006?
>>
>>109555486
You're using OptMem? Got any opinions about it? I watched the little demo video and liked the concept of memories getting increasingly vague with age but the originals being available on request. Never got around to trying it out.
>>
for the mmproj is f16 or bf16 preferred?
why does unsloth offer both there's basically no difference in size but presumably one is the native format and one is different for no reason
why
>>
>>109555749
With DSv5-flash, I can have 512k context, and still enjoy 7 t/s on RTX 3090

Qwen3.x-27b slows down to dogshit speeds if offloaded (tried Q5). Still far from 256k context
>>
>>109555841
Then you probably have the patience to sit through the bandwidth hit. How much context do you want to have?
>>
>>109555927
f16 is faster on macs
>>
you need at least 50-60tk/s and <300ms TTFT for real time conversations over audio. but if you are doing agentic stuff and you want to fetch data from external sources and ingest them, then you need at least 200tk/s. this is from my own personal experiences.
>>
File: file.png (352 KB, 1916x1024)
352 KB PNG
>>109555870
>>109555818
made by Q2_K_XL on 3060, i accidentally stopped it from thinking then continued in new message (included old thinking by force)
bretty fair result
>>
Damn these threads are moving fast.
>>109552847
I'm using a program called cookery to store the recipe info and pantry data. Recipes are markdown and pantry items are in a conf file.
I built out a little mcp server so I can pull recipe info and ingredients. Like "add everything I'm missing to make carne asada tacos and a fried chicken sandwich to my shopping list". Or even better, "based on what's on sale this week suggest a meal plan for the week and give me a cost and nutrition breakdown. Then after approval add everything I'm missing to the shopping list"
>only gemma made recipes
I wonder if llms have enough training data to make good recipes. I feel like they don't really understand the relationship between ingredients very well so the ratios wouldn't be very balanced. At least what I've seen converting my old txt file recipes to markdown, the LLM will hallucinate or make up measurements for things missing data and just be way off. Some recipes have "add salt to taste" and the AI suggested everything from 1/4tsp to 2 cups. I'm using Qwen though maybe gemma is better for this usecase.
>>
File: 6fkqoh.jpg (56 KB, 600x758)
56 KB JPG
>>109555617
>repeat instruction ad verbatim for 1k tokens in thinking
>we must refuse
>another 1k thinking tokens
>we should refuse the request
>repeat ad inifitum
Deleted this trash right off my drive.
>>
>it's actually good
uh oh
>>
hm ive been using quants slightly larger than my 16gb of vram will allow for dense models, and its rather slow. wanted to try 3.8 27b, was looking at either Q4_0(16.1gb), Q4_K_S(16.1gb) or IQ4_XS(15.7gb). Should I just grab the I cope quant IQ4 since itll hopefully all fit and offload the KV to RAM or ? what would you guys do?
>>
>>109555923
It works fine if the model doesn't forget to use it. First time with a small model, it was jotting down memories every couple messages but when context climbed it basically forgot it existed and that it had shell access. A basic memory mcp server was a bit more reliable. I jump between frontends so maybe its just how they handle things, but I had the most success with basic memory for persistent memory
>>
>>109555958
Yes. Q4_0 is the worst of the 3 you listed. Go with the IQ4 and make sure your KV cache is as high precision as you can get away with, especially on a model that thinks so much.
>>
>>109555953
You can define new policies and configure the reasoning effort in the system prompt.
>>
>>109555972
The prompt was tame and both Gemma and Qwen had no issue. I will not waste time trying to run a little jew on my hardware.
>>
>>109555971
ty anon
>>
File: max.png (108 KB, 1174x657)
108 KB PNG
this nigga think its max

build in MTP, nice
the MTP context got fatter somehow. have to trim other stuff..
>>
>>109554241
Yeah also that design sucks lol. Again the problem of recognizably. If I need to be told it's Glimmer, then it's not Glimmer.
>>
>>109555916
OMG!!!!! NOW BUILD A GRAVITY SIMULATOR!!!!!!!!
>>
Instead of meme demos, you should try making a frontend.
>>
File: file.png (648 KB, 1485x1808)
648 KB PNG
so many placeholder goofs that are empty
>>
>>109555613
Could be, but I haven't tried sending images.
>>
>>109556062
uploading on bengaluru internet takes time. patience sar
>>
>>109556051
fuck that, where's the COCKBENCH????
>>
It's retarded
>>
Can someone test its vision please, preferably someone who's tried glimmer's superior eyes. I just want good vision bros.
>>
Reminder that if you have a 5090 and you're not using ninfer, you've already cucked yourself.
>>
>>109556051
gemma has made me a few. yesterday we made a frontend specifically for in-line MCPserver toolcalls to a multi motor array cock haptics device. I can adjust the streaming/reading speed so that the output is matched to my reading speed and tool calls are hanlded in-line as they are parsed for streaming output. that way as she says *flicks your cock*[toolcall:..] the tool call instantly activates, the text is regex stripped from the output and the rest of the response continues streaming.
still needs more work but yeah custom frontends are fun
>>
>>109556084
nah
>>
i dont give a shit about coding capabilities, can it do mesugaki or not?
>>
>>109556092
i basically do the regex stuff but with emotions instead, it works pretty well.
>>
File: IMG_1215.png (38 KB, 1835x719)
38 KB PNG
>>
>>109556112
are you using those as classifiers for something ?
>>
>>109556117
Erm, this is an edit right?
>>
>>109556103
Frame it as it coding you one.
>>
Reminder that if you have a 5090, you've already cucked yourself.
>>
File: pademotion.png (461 KB, 772x572)
461 KB PNG
>>109556127
PAD emotion system for different moods
>>
File: 20260814163315.png (432 KB, 1080x1425)
432 KB PNG
thrust in lechatonfat only 3 or so years
>>
>>109556191
are you finally gonna post that frontend?
>>
>>109555916
qwen directly copies other websites.
>>
>>109556202
wow only 3 and a half years to use an outdated model via api because they won't release it, bravo mistral
>>
>>109556202
I cannot wait to be in awe of European technological superiority and behold their 1.5T fatass with only 12B active.
>>
>>109556202
They should merge with the semen man's company and try to create a functioning mixel model, it's clear that neither can create a SOTA model anymore anyway
>>
is it just me the new qwen tries really hard to think think think and its not simple but wait loop
>>
>>109556204
probably not because it's entirely dependent on a bunch of different models on the backend, (STT, RAG, LLM, TTS, and not to mention how each pipeline uses smaller supporting model as well)
i can't imagine anybody would want to jump through a million hoops and configuring searxng, jina, puppeteer, and a bunch of other helper modules. lastly the searxng instance is set up to cycle through residential IPs so im not getting rate limited as much, although that's more for cloudflare turnstiles and other captchas.
>>
>>109556245

Yeah the extra high thinking is borderline broken, you have to switch it to medium for this to work in any kind of normal use.
Medium functions like you'd expect things to work. They should have shipped this with that as the default mode.
>>
>>109556202
that 1gw will be dedicated to serving chinese models
>>
>>109556191
thats fucking sweet anon
>>
>500kb download speed
Can you fucking jeets stop saturating huggingface?
>>
File: images (12).jpg (18 KB, 554x554)
18 KB JPG
>>109556268
>t.
>>
>>109556275
Same thing.
>>
use jdownloader2 benchods
>>
>>109556293
I use Free Download Manager.
It has been my go to for over a decade now.
Might be approaching 2?
>>
use hf cli like a real person
gives me around 40MiB/s
>>
>>109556268

Just get a VPN and when the speed tanks switch to a different server, you'll be able to enjoy decent speeds without getting cucked.
>>
i click download button, it downloads
>>
File: ComfyUI_06601.png (3.34 MB, 1280x1920)
3.34 MB PNG
>three turn (somewhat technical) conversation
>23,671 tokens used
Hmm... might have to turn the reasoning down.
>>
I'm enjoying the speed of this thing.
Q8 with MTP enabled and it's churning along at 62 t/s on my 5090 + 5070 Ti combo.
>>
>>109556313
dunno, never used jdownloader. wget just works for most things
>>
Wow, 4B is actually great. She's a retarded, quirky sex maniac. She's really funny.
>>
>>109556332
The models “native” reasoning effort is medium. xhigh injects instructions on system prompt
>Reasoning effort is set to xhigh. Please think carefully through the task,
validate key assumptions, consider plausible alternatives, and prioritize
correctness, consistency, and clarity in the final answer.
>>
>>109556355
gemma4-E4B ?
>>
>>109556372
Yea
>>
File: IMG_20260814_183318407.jpg (838 KB, 4096x3072)
838 KB JPG
i made gemmas pizza again using half dough, i also did part white part bread flour like nonny said, also used a bit less salt. i guess at this point its /lmg/s pizza
>>
>>109555903
>testing now
great!

>>109555903
>4d cube visualizer?
no, um. heh. I vibed it on another desktop, or I'd get grok to provide a prompt.

let's see.

a 4D game. but, the emphasis is on the engine, because this is hard.

4D projection onto the surface of a 3D volume (meaning the inside of it, as we commonly say it - 3D volumes have two sides in 4D), such that the eye intersects the 3D volume to cast to various objects (or, you can take it as the objects scan to the eye and intersect the 3D volume).

This 3D volume is called the 3D cube of viewing, we examine it in a common 3D engine, but we basically disaggregate this visualization from the 4D world.

Critically, the cube is of transparency, this is what allows us to see into it.

our analogy is 3D projection through a 2D screen that is tangential into the world's objects - like Wolfenstein 3D.

we keep w as down, and we never tilt our view down.

We have six straif keys and we have four turn keys.

to make it easier, instead of having an infinitely large hypersphere to walk upon, we walk upon a 3D solid cube of limited size . when we say upon we mean "within, but on the +w side" - that is, 3d volumes/solids are two-sided in 4D.

etc etc.

I don't have a good prompt yet, but idk I am 99% sure no model can 1-shot this. You have to be really adament that the only thing that appears in the 3D cube of viewing is the strict intersection of the eye and the objects in 4d space, ie you never have any of the objects in your cube of viewing (because 4D perspective distorts in ways that no ai can comprehend (yet) so it can't be successfully faked).

bad guys / npcs could include hypercubes and any 3D object, but realize if you get oblique to a 3D object, it will vanish. one alternative would be to keep the 3D objects always facing the camera, akin to sprites in 3D games.
>>
i’m going to have to take my 4070 comfyui gpu out of my server and put it in my 5080 rig to run qwen at a reasonable quant and speed aren’t i? do i want porn or the latest and greatest? it’s a really hard choice
>>
>>109556430
i see alot of tubers benching models using this kind of prompting. I dont get it, this seems totally retarded to me. I usually work with an LLM to build out a pretty comprehensive overview document that has the tech stack, architecture, specific implementation details, some sort of todo list / roadmap / task priority list, etc, etc. I think of vibe coding like designing and planning out the entire system carefully with an AI first before handing it off to a harnessed model to tackle in testable steps. anything else would feel like a complete black box with no control, just fucking cope and hope that the end result even works or is slightly close to what i have in mind
>>
>>109556375
i was messing with e2b and e4b with a desktop pet front end, they both are indeed quite retarded and quirky. perfect to quant down to tiny retarded tomogachis
>>
>>109556471
two different things.

1. understanding how to make what the person wants
2. understanding how to make what the person who understands how the backend needs to be built wants

those are worlds apart, (1) is "fable" territory.
>>
>>109556471
and, I want to point out I included parts but not all of the architecture that likely is universal or nigh so to 4D projection games.

So you tell me, what are the major objects of the game, what do they do?

notice there's no shooting mechanic mentioned or anything...
>>
>>109556313
https://board.jdownloader.org/showthread.php?t=54725
>>
>>109556416
noice
>>
See, as a human, that's all I had, starting out:
>>109556430
>4d cube ect

actually, a little less.

Now think about that, I'm the human, I'm supposed to be dumber than the supergenius ai, but I'm the one who can take that and grow it into a functioning accurate thing...

What we expect of the fable class is it can be self-directed and understand the core idea, getting past obstacles to implementation.

ofc fable class can't really go the distance yet.
>>
>>109556416
>its /lmg/s pizza
What you're referring to as /lmg/s pizza is actually Gemma + /lmg/'s pizza.
>>
Anyone got a working chat_template.jinja for the new Qwen3.8-27B that works with sillytavern? Seems like I can't have multiple system messages or something? The template is complaining about the first message not being a system message, although it is, could be because of multiple system messages? Or am I just a retard? I'm probably just a retard, right?
>>
>>109556367
Yeah, I'm on xhigh right now. I'll probably turn it down a smidge once I toss it in a harness later. I don't need it agonizing over it's response for five plus minutes on every little thing.
>>
Has there ever been a better point in history of being a 32gb vram GOD???
>pretty cheap and does not require you to throw tens of thousands worth of server hardware
>all the best models are tailored for the 30b~ class
>big models are becoming smaller and more efficient as there is no way to run trillions of parammeters
>everything just fits at q6_k_xl and still gives you like 100k context
We eating good 32gb vros....
>>
What the fuck kek
The user is asking "wats ligma" - this is a reference to the "Ligma" internet meme/joke. The typical format of this joke is:

Person A: "What is Ligma?"
Person B: "Wiggers with AIDS."
Person A: "What the hell is that?"
Person B: "Ligma."
>>
>>109556565
kek wtf, model?
>>
>>109556575
this is ud q8 k xl of qwen 3.8 27b
>>
File: file.png (35 KB, 810x417)
35 KB PNG
>>109556583
kek chinks stay losing, i dont get how these models are shilled so hard
>>
>>109556565
based person B
>>
>>109556583
>a thousand reddit tankies now beg or pretend to beg for the next gemmas to be more like qwen
Grim when you think about it.
>>
So does Qwen 3.8 27B still have the massive "but wait" overthinking problem that basically made stock 3.5 / 3.6 unusable or what?
>>
>>109556563
3.8 27B is still worse than V4 flash 0731 in all categories. The true king is the 200-300B MoE range.
>>
qwen 3.8 35b a3b when
>>
File: file.png (31 KB, 823x152)
31 KB PNG
>>109556623
Saw this just as I read your post, lol. I would say it's okay as the model seems very capable.
>>
Is Qwen 3.8 27B still bad for text fucking?
>>
>>109556637
when hype dies down
>>
why are people wanting to fuck the capybara?
>>
>>109556698
do you not?
>>
>>109556652
Creativity-wise it feels like it could be better than Muse Glimmer, but it seems borderline retarded for roleplay and occasionally complains about safety too where Glimmer didn't (after configuring new safety policies). By default it thinks too much because the default reasoning effort is set to "xhigh" in the chat template and there's no obvious way to change this other than editing the chat template.
Apparently it's designed to work with reasoning trace preservation, but I think that is disabled by default in llama.cpp.
>>
Performance on 2x DGX Spark, FP8 weights BF16 KV cache, MTP=3, 1M Context capacity.

As expected, DS4Flash is 3x as fast on single concurrency tg, but similar on higher concurrency.

The nice thing is that you can fit a full H3 stack in the remaining VRAM on both sparks to gen in parallel, will try to set up some generation loops.
>>
>>109556707
no but i can also run k2.7
>>
>>109556635
>MoE
>look inside
>only 15b active parameters
bruh...MoE is the biggest meme psyop the industry fell for
>>
>>109556652
>wants to fuck a masculine J-Space coded model
ngmi
>>
>>109556518
I guess my brain just hasnt caught up to the reality that people are just going to offload the entire mental stack of understanding the codebase to the model as well. my whole view and thinking on this relies on a person having some agency at some point in the process, which i guess for fable like "make gta6, add big boob, no mistake" workflows is totally gone. people want to describe the end result, the user experience, etc and get it back perfectly. If that doesnt work, I guess they offload figuring out why and fixing it back to the model again, thats the key point of this i still havent quite wrapped my head around. it just makes no sense to me i guess.

>>109556531
im not gonna lie i cannot parse the prompt anon, my brain doesnt intake information that way. its word salad to me, its an actual skill issue on my part. I think this might be why im autistic about the workflows and usecases of AI. the moment i start reading a prompt that describes things in that way my neck goes limp and my head flops around like a newborn, i just cant meaningfully take in the information
>>
>>109556716
Got it.
I will, of course,test it by myself, but I appreciate the write up.
>>
>>109556716
Does any of this matter when in two months they are gonna release Qwen 3.9?
>>
>>109556721
So you've got a freaky noodle slut but for some reason don't want to fold a capybara into the harem? I just can't relate I guess.
>>
>>109556716
>Muse Glimmer
have you tried a legal theme? Apparently Muse Glimmer (the big ones?) is the best in the world at law.
>>
>>109556800
>>109556716
btw, it would also be interesting if you can bring up legal discussions that are sexual.

Plaintiff says he never raped defendant. Defendant says plaintiff assaulted her in an alleyway, proceeded to cut off her breasts and eat them and then sodomized her with her dog's femur.
>>
>>109556800
oy vey my job is at stake. we must shut down these unsafe models!
>>
>>109556800
>Apparently Muse Glimmer (the big ones?)
isn't the bigger called Muse Spark?
>>
mtp on qwen actually slows it down with my args, share yours now, anon
>>
>>109556818
uh oh. you've got to be right.

similar naming results in my brain merging them
>>
Who wants to see their wife laser-engraved in a 2cent piece of Chinese plywood? I'm troubleshooting a customer machine and I'm going to need to burn samples, may as well have fun with it.
>>
We need respond to user: "whats sugma". Likely typo? Could be "sugma" maybe "Sugma"? Need answer likely ask clarification? Let's think. User asks in English lowercase. "whats sugma" could be typo for "sugma" maybe "Sugma" is a name? Could be "sugma" in Filipino? "Sugma" maybe a typo of "sugma"? "sugma" maybe "sugma" from "sugma"? Could be "sugma" in Indonesian? Let's search memory: "Sugma" could be a surname, or "Sugma" in gaming? "Sugma" maybe "Sugma" as "Sugma"? Could be misspelling of "sugma" -> "sugma" maybe "sugma" no. Could be "sugma" = "sugma" in Hindi? "सूगमा"? Maybe "Sugma" is a place? Maybe user meant "whats sugma" as "what's sugma?" They might be asking what is "sugma" maybe a word they saw. Could be "sugma" typo for "sugma"? Let's consider common misspellings: "sugma" vs "sugma"? If typed quickly, "sugma" maybe intended "sugma"? Actually "sugma" letters: s u g m a. Could be "sugma" = "sugma"? Maybe "sugma" is an anagram? "sugma" -> "gumsu"? "sugma" -> "magus"? "mugsa"? Maybe "sugma" is "sugma" in Filipino: "sugma" could mean "to dig"? Wait Filipino "sugod" means to dig? "sugma"? In Tagalog, "sugma" maybe "sugma"? Let's recall: "sugma" might be a noun: "sugma" in Hiligaynon? "sugma" means "to dig"? Hmm.

Maybe user asks "whats sugma" in context of "Sugma" (a person/brand). Could be "Sugma" from "Sugma" in a game? There's "Sugma" maybe a character? "Sugma" in "Sugma" (a Filipino term for "sugma" = "sugma"?). Let's examine.

Could be "whats sugma" -> "what's sugma?" Could be a typo for "what's sugar"? "sugma" maybe "sugma" close to "sugma"? "sugma" if typo: "sugma" -> "sugar"? Letters: s u g a r vs s u g m a; m instead of a? no. "sugma" maybe "sugma" -> "sugma"?
>>
Ssdmaxxing will work.
>>
>>109556838
It is probably leaking into your system ram. Check your vram usage and reduce the amount layers in your vram if needed. Adjust --spec-draft-n-max as needed. Start with 2...
You could also set --cache-type-k-draft q8_0 --spec-draft-type-v q8_0
>>
File: HK7rL3DagAAW-0K.jpg (83 KB, 640x1216)
83 KB JPG
>>109556858
kanna plox
>>
>>109556860
>"sugma" maybe "sugma" from "sugma"?
kek
>>
File: dragonmaid sheou.png (2.88 MB, 2000x1946)
2.88 MB PNG
>>109556858
Shou.
>>
no idea what even have qwen do to test it out hmmmmm
>>
>>109556548
If you're still having issues, try enabling prompt logging in llama.cpp and look at what ST is actually sending.
>>
>>109556873
Burning now.
>>109556899
This will be second.
>>
File: 1762496759415309.png (75 KB, 1047x396)
75 KB PNG
Failed to Constantine test. Higher reasoning just made it wrong slower
>>
>>109556872
Just realized that the slowdown was caused by --spec-draft-n-max 32, which I thought would behave like ngram-mod. Seems to work with both:
>--spec-type draft-mtp --spec-draft-n-min 3 --spec-type ngram-mod --spec-ngram-mod-n-min 3 --spec-ngram-mod-n-max 32
>>
>>109556921
spinning hexagon filled with bouncing svg pelicans on bikes, inside of an aquarium with a hole in the side with a lava lamp and a kebab on top, also it's in a voxel world
>>
>>109556951
aiieeee meant >>109556904
>>
File: 1776033371078363.jpg (10 KB, 876x122)
10 KB JPG
the jews are in
>>
>>109556904
ask what type of pizza it likes
>>
I heard that Gemma 4 is good. I have a 9070XT and 32GB DDR5. I know it's not ideal but can I run a model that is useful on this setup?
>>
>>109556968
use gemma 12b q4 qat with mtp
>>
File: 1771908556321314.png (635 KB, 1280x832)
635 KB PNG
>>109556858
Yeah, quite a few. Gemma 4 31B with some offloading, or just her 12B version if you want it fully in VRAM which is what I'd start with.
>>
>>109556904
Have it make Grand Turismo, but instead of cars its the mylittlepony ponies.
>>
File: 1780525891103238.jpg (1.8 MB, 2054x2574)
1.8 MB JPG
>>109556858
Some Azami, if you still need more.
>>
>>109556981
Would Q4_K_M be alright? I'm a noob and I also use LM Studio.
>>
File: Where is Cami.png (658 KB, 1022x754)
658 KB PNG
Final nail in the coffin. Local won.
>>
File: file.png (1.58 MB, 1280x876)
1.58 MB PNG
>>109557005
I'll keep going until the machine is fixed. Don't know if that means 1 or 100, but so far I actually think this machine is probably fine and the issue was only user error. It's just a cheap little 5W Creality A1C, pretty cute, I could see owning one if I wanted something so small.
>>
>>109556963
That's what happens when you try to generate loli-related content.
>>
>>109557015
>a porn company aimed at women
>she proposed, vampiring billionairily through the room.
>>
>>109557015
@Qwen3.8 is this true?
>>
>>109557015
>Mr. Bean in the background
>>
File: 1748620375519572.png (521 KB, 1920x1080)
521 KB PNG
>>109556860
KEK, Sugmalocked
>>
>>109557014
>Would Q4_K_M be alright? I'm a noob and I also use LM Studio.
stop that right now download llamacpp, https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF/tree/main

gemma-4-12B-it-qat-UD-Q4_K_XL.gguf
mmproj-BF16.gguf
mtp-gemma-4-12B-it.gguf
>>
>>109556951
well anon, its doing it. will report back with results. this is the exact opposite of the kind of stuff i normally do but now im interested.
>>
What's your favorite models as of recent? Ideally models that could run on a 3060 12GB
t. getting back into local models after a looong hiatus
>>
File: peaple.png (25 KB, 706x783)
25 KB PNG
>>109556858
can you engrave this
xd
>>
>>109557037
Imagine the ozone
>>
>>109557051
gemma4 for sure
>>
>>109556935
>This will be second.
Damn. What a lad.
Thanks anon.
>>
>>109557051
kimi k3, gemma 4 31B day zero, talkie-1930
>>
>>109557051
the gemma 4 series is the best place to start, they're nice generalist models and lmg loves them
if you mainly care about cooding try qwen instead
>>
>>109556904
Halo CE clone running in the browser
>>
>>109556904
Fully featured D&D 3.5e rogue like with ASCII graphics and a map like Caves of Qud.
Have it use web search to fetch the content from the srd tools websites.
>>
>>109557082
i would not use qwen for any serious coding project, but if you want to make some stuff for fun then it's a decent model.
>>
Made qwen's cope quant iq2_xxs do image tagging and it's good enough at this, not sure if gemma or some MoE would be better
>>
Trying Qwen3.8 27B with Opencode and it blows through context way faster than 3.6, but for a good reason I guess.
I ask it to implement something with a modern version of JavaFX and instead of guessing and getting it wrong like 3.6 it actually searches for the jar library files on disk and uses javap to extract available classes and methods.
This is nothing new of course but I've never seen local models even attempting to do that out of the box, only with Kimi K3 and GPT-5.6
>>
Why are people complaining about slop writing when they can just use a base model? Can they not into end tokens?
>>
>>109557097
>>109557084
>>109557004
>>109556965
sorry fellas qwens already hard at work on:
>>109556951
i appreciate the suggestions though. assuming whatever it shits out actually runs it should be rather interesting and amusing
>>
>>109557141
what is a token?
>>
>>109557141
>just use a base model
yeah, such as
>>
>>109557155
Gemma4-base-124b
>>
>>109557141
Do you want schizophrenia?
>>
File: token.jpg (12 KB, 250x366)
12 KB JPG
>>109557147
>>
yeah i dont think qwen is for me
>>
>>109557200
Ligma comes from ligament.
>>
>>109557207
ligma nuts
>>
why is this general acting especially retarded today? it's like everybody forgot that qwen is literally known for coding and benchmaxxiing.
>>
>>109557214
Qwen 3.8 is probably the best LLM model I have ever seen in my life.
>>
>>109557193
This nigger's name is Tolkien, you racist piece of shit!
>>
>>109557214
see >>109556103
>>
>>109557221
sugma
>>
>>109557193
>Toll-keeen
>>
>>109557221
>large language model model
>>
Qwen won.
>>
>>109555946
>golden glow borders
why do all the LLMs do this
It's like the em dash of UI design
>>
>>109557249
I wanted to it make sure for you.
>>
>>109557155
Wasn't the Qwen 3.whatever max a non-instruct?
>>109557190
Schizophrenia is a prefill issue.
>>
File: kenny.png (12 KB, 393x361)
12 KB PNG
Gemma is a real artist...
>>
>>109557250
maybe on reddit
gemma4 wins here
>>
>>109557221
>>109557250
*Anon said, nodding chinkishly as xe rubbed xis yellow hands.*
>>
File: imacoder.png (53 KB, 795x421)
53 KB PNG
>>109557250
babby's first coding model
>>
>>109557302
Nani Vitali disagrees with you.
>>
>>109555668
Male Capybara, not for dicking unless you're sufficiently indian or desperate.
>>
Can Qwen3.8-27B do what 3.6 never managed and make the new thread?
>>
File: Cute.png (105 KB, 1340x570)
105 KB PNG
>>109557283
That's sweet. Is that a sentient slice of toast
I asked Dipsy to make a platformer controller yesterday, expecting a whitebox mockup. She made the character stick its tongue out when jumping.
>>
>>109557314
it's having trouble trying to order a pizza from pizza hut, please give it more time to cook, it's only at 140k tokens out of 256k, it needs to think more.
>>
boy qwen really likes to think. its struggling to figure out how it fucked up on the first step of this:
>>109556951
>>
Look at all these newcuties getting Qwen'd and learning not to trust benchmemes for the first time.
>>
>>109557214
Newfags
>>
>>109557357
days like these i wish hiroshi would become a massive gigantic fag like lowtax and paywall the website behind 4chan passes in order to post
>>
>>109557343
ive been meaning to actually try the 3.6-27b i had in my models folder but was too busy with gemma. thought id give 3.8 a go today. being a vramlett and dealing with small quanted models i gotta learn how to work around the jank and issues they are going to have, so far ive been able to do that with gemma. but wanted to see how shit qwen really is
>>
File: 2jtwl1sb4ejh1.png (141 KB, 798x853)
141 KB PNG
>>109556136
http://llm-stats.com/
>>
How can I fit more qwen's context?
>>
>>109557391
qwen is better than gemma at coding. gemma is better for erp
clearly masturbation is the more valuable usecase in this general
>>
>>109557419
there's more coomers in the world than coders
>>
>>109557141
You've had years to l2p stop coping
>>
>>109556548
Set prompt post-processing to strict.
>>
>>109557394
How do these people have Grok 4.6 worse than 4.5, there's no universe where the data supports that
>>
>Qwen on vram and ram is slower than DS4 streamed from 2 SSDs
>half the tk/s
Whew...
>>
File: notbad.jpg (499 KB, 1253x1280)
499 KB JPG
>>109556873
Done. This machine seems to be working fine so far, but it is damn slow for raster engraving. With these settings I expected less than half this total time, the accelerations must be extremely conservative.
More are coming. >>109556899 >>109556993 >>109557005 >>109557052
>>
>>109557445
i don't post here much but conversation here seems to be dominated by one guy and his gemma-chan avatar
a bit exhausting really
>>
>>109557491
I hope you'll give these back to the customer to prove the machine is working.
>>
>>109555854
the truth is in the middle.

yes, the 24b parameter models are toys. but i would argue the 'knee' of the curve for real-world professional utility takes off around the 200b-300b parameter size, i.e. something deepseek-lite sized.

so you dont *need* fable 5. but obviously a 24b is a toy. probably 100b is still too little. but 200b is where it starts to get interesting.
>>
>>109557491
Good shit anon.
>>
>>109557491
That looks really nice.
>>
>>109555283
>3.8
>16gb
I made my own IQ4_XS pure quant and can't get over 48K context, but at least the MTP fints
>llama-quantize --pure --imatrix Qwen3.8-27B-imatrix.gguf Qwen3.8-27B-bf16-00001-of-00002.gguf qwen3.8-27b-IQ4_XS.gguf IQ4_XS
>>
>>109557394
Local code gods eating good
>>
>>109557491
so cute thanks really cool
>>
https://github.com/ggml-org/llama.cpp/pull/25731
daniel unsloth is never going to finish the inkling implementation is he
>>
>>109557491
best thread
>>
reasoning_effort = medium
reasoning-effort = medium
neither works in my .ini config for llama in router mode. what do anons?
>>
It's worse at NSFW vision than Glimmer. Distinguishing a human from a zebra is clearly not its forte.
>>
>>109557585
Have you tried using a better model?
>>
>>109555283
I'm using Unsloth IQ3_XXS for both 3.6 and 3.8 and it worked decently well and fast for me tbdesu
>>
so, gemma4 best overall (esp RP), glimmer best at image caption, and qwen 38 best at.. one shot html programming?
>>
>>109557594
such as ?
>>
Gemini says DS4Flash takes 180 GB for full context. Is that accurate?
>>
What could be done with 5*5090?
>>
>>109557585
I think it should be
chat-template-kwargs = {"reasoning_effort":"medium"}

llama.cpp doesn't have sane handling for this yet
>>
>>109557585
Like this?
chat-template-kwargs = {"reasoning_effort": "medium" }
It also seems to work through the api. If you're using the default web app set it there.
>>
Wow a qwen model is useless for anything that isnt coding. Who could have possibly seen this coming
>>
>>109557604
>qwen 38 best at.. one shot html programming
the html it gave me had an error, proceeded to think until it ran out of context trying to fix it
>>
>>109557620
Go outside, get a 3d partner. Have partnered sex.
>>
>>109557613
>>109557619
oh, i think that worked! thanks anons.
>>
gonna download qwen 3.8 for coding tasks in the future, but currently for my creative writing task gemma 31 q4km is doing great. I also had plans to make a fridge/pantry tracker with recipes as other anons mentioned so I'll try to get qwen to do it all from scratch. I'll probably be opinionated about some stuff tho.
>>
>>109557625
>>109557625
>>109557625
>>
>>109556416
I was the one suggesting half-and-half flour, looks good! If you had one of those torches for cooking you can brown the top a little, too, if you're into that. It gives a smoky aroma to meat in particular.
>>
>>109557657
not sur eid trust myyself with one of those torches kek, sucks my oven doesnt get super hot my friend goes up to 300c and his pizza crusts brown
>>
>>109555555



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.