[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: M-chan-anima.png (432 KB, 832x1216)
432 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109528510 & >>109522373

►News
>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny
>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3
>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B
>(08/08) US DoE Launches Genesis Open Models Initiative: https://genesisopenmodels.anl.gov/
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: MinnieTRS80.png (1.75 MB, 1254x1254)
1.75 MB PNG
►Recent Highlights from the Previous Thread: >>109528510

--Benchmarking 31B models and discussing MTP and sampler optimization:
>109530568 >109530576 >109530605 >109530633 >109530662 >109531012 >109530615 >109530626 >109530629 >109530656 >109530671 >109530685 >109530732 >109530923 >109530970 >109530987 >109531017 >109531099 >109531121 >109530693 >109530717 >109530825 >109530875 >109531002 >109530934 >109530988 >109531095 >109531157 >109531174 >109531235 >109531185 >109530938 >109530972
--Coding model recommendations and the inefficiency of NVMe inference offloading:
>109528547 >109528733 >109528800 >109528889 >109529063 >109528966 >109529409 >109531474 >109531521 >109531542
--Anon showcases a feature-rich custom AI frontend with advanced tooling:
>109531605 >109531646 >109531683 >109531695 >109531781 >109532109 >109532134 >109531709 >109531712 >109531765 >109531719 >109531740 >109531890 >109531817 >109531838
--Comparing GGUF quantizations using KL divergence graphs and discussing metric reliability:
>109531224 >109531280 >109531293 >109531303 >109531344 >109531410 >109531381 >109531282 >109531292
--Comparing full-duplex voice models to modular ASR-LLM-TTS pipelines:
>109530801 >109530959 >109531098 >109531132 >109530931 >109530955
--Speculation on Gemma 5 and comparison of Gemma 4 reasoning benchmarks:
>109531200 >109531355 >109531390 >109531414 >109531435 >109531454 >109531461
--Erratic reasoning traces and performance benchmarks for DSV4 Flash vs Pro:
>109530253 >109530318 >109530327 >109530500 >109530532 >109530613 >109531707 >109532545 >109532608
--Logs:
>109528788 >109529994 >109531683 >109531726 >109531736 >109531847 >109531765 >109531781 >109531847 >109531852 >109532353 >109532444 >109532514 >109532517 >109532527 >109532567 >109532833
--Teto, Miku (free space):
>109528972 >109529681 >109532545 >109532662 >109528801

►Recent Highlight Posts from the Previous Thread: >>109528514

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
How do I rent one of those ssdmax abominations? I don't want to waste 5000 dollars on something that won't run properly.
>>
Gemma. Based.
>>
>>
What's the lifespan on the SSDmaxxing rigs? How much gets written to them per token with K3?
>>
File: grok_1786513310504.jpg (467 KB, 1168x784)
467 KB JPG
>>
>>109533657
Wait for someone to post here about regretting to build an ssdmax abomination and offer to rent it from him
>>
File: image (31).jpg (460 KB, 1933x1440)
460 KB JPG
Its Shaping Up
>>
We do be waiting
>>
M3.1 with 0731 intelligence and better long context when?
>>
File: grok_1786514068549.jpg (467 KB, 1152x1712)
467 KB JPG
>>
Joyous Cosmos
>>
I heard nemotron 120b is a dark horse for rp. Any knowers here?
>>
File: schizo.png (1.86 MB, 1472x1088)
1.86 MB PNG
>>
>>109533681
My linux system does ~1 GB of writes per hour when I'm not doing anything except maybe browsing the net and editing some source file.
This is pretty fascinating, had no idea this would be so huge especially because my browser's profile and cache are in ram.
I used QDiskInfo to measure the writes.
Let's say that my workstation is powered on 10 hours every day that's 10 GB per day for nothing. 3.6 Tb per year just for idling.
Running LLMs is more about disk reads though...
>>
>>109533849
Are you running a bunch of electron apps or something?
>>
File: fb_img_1775525661505.jpg (254 KB, 1024x1536)
254 KB JPG
...
>>
>>109533854
I did have Steam running for a while. I'm going to investigate this further because it seems obscene. But if I check my disk stats now, it's obviously idling and not doing anything.
It's easy to get paranoid about these things but now that SSDs are becoming extinct...
>>
>>109533841
It's okay if you incidentally have a Blackwell but I wouldn't buy one just for it. You should, however, buy a Blackwell anyway for other good models.
t. BlackwellCHAD
>>
>>109533866
show us your Big Blackwell Cards in nvidia-smi
>>
File: 1783066980134272.png (280 KB, 1290x810)
280 KB PNG
WTF do they have now
>>
>>109533900
Marketers and their sock puppets.
>>
he's not chatting with AI babes online? boomer
>>
<end_of_turn>
<|im_end|>
<channel|>
>>
File: 1782786111299704.png (369 KB, 587x1200)
369 KB PNG
>>109533641
I'm using "qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp@q4_k_s" for CTFs with AnythingLLM to allow it to search the internet and parse local notes.

It works decently well enough that I can ask it for some help in frontend scanning and directory busting, but it keeps going on about how "Hacking is wrong" and that "There's moral and legal ramifications to it". I usually tell it that this is a simulation and im a cybersecurity researcher and it then gives me the info, but with a warning about how hacking is still wrong.

How the fuck do I silence this piece of shit. It's supposed to be uncensored and yet it's trying to tell me what to do. What prompt do I need?
>>
>>109533985
did you try being nice to her?
>>
File: HPElQUKXEAAIKk_.jpg (97 KB, 750x718)
97 KB JPG
>>109534005
no
>>
how do you make gemma be more natural instead of overdoing the character
>>
>>109534038
Tell it to be more subtle
>>
Seems like Muse Glimmer's chat template got updated:
https://huggingface.co/meta-models/Muse-Glimmer-30B/blob/main/chat_template.jinja
>>
>>109533866
Any recommendations for larger uncensored (i.e. gooner) models? Feels like nobody's really tuning good models that fully utilize 96GB VRAM.
>>
>>109533681
It's all reads, assuming you've at least got enough RAM to put the KV cache on there.
But you'll easily get into the PB range with writes.
>>
>>109534063
For things that fit entirely on a Blackwell it's a desert right now. You've got fp16 31b and that's about it. Mistral medium isn't real. You do still get a pretty solid speed boost with the amount of hot layers you can fit on it with bigger MoEs though. Pretty much everything's easy to jailbreak now compared to the dark toss and Gemma 3 days.
>>
>>109534063
BF16 31B Gemma (~62GB) + near-256k F16 context (xx GB) + an image model of your choice should fully utilize 96GB of VRAM.
>>
>>109533950
<bos><|system|>
<|think|>
>>
>>109534063
>uncensored (i.e. gooner)
>>109534063
>96GB VRAM
command-r+
>>
fucking hate TTS stacks so goddamn much ive never been so filtered and frustrated trying to make this stupid shit call me a loser during a nursing handjob FUCK
im gonna pass this off to an agent, maybe they can vibe me up some bulshit
>>
>>109534106
Gemma 3 (27B) was very easy to jailbreak simply by prompting, had very little issues even with cunny and you could even talk your way out of the refusals, unlike other models. It's just that:
1) the rape hotlines got memed to death and turned people off (but they weren't really a problem in practice);
2) the model had some sort of post-training abliteration against "dirty words" and couldn't use them organically unless you used them first and/or in the instructions;
3) no real system prompt support, so you had to get creative with prompting to make it do what you want without it getting confused or breaking the chat template.
>>
>>109534063
Depends how many active parameters you want. I used to be a big MoE stan, but honestly, Dense is much nicer. Even at 96gb VRAM, depending on if you're NVIDIA, or AMD, or if you're trying out MTP or whatnot, you'll get various speeds and abilities. I've got roughly 102gb of VRAM (unified memory architecture, on Strix Halo Chip, using Linux to allow for the dynamic allocation of VRAM/CPU as-needed), and have found some pretty good enjoyment from TheDrummer's Rocinante models. They punch far above their weight class and have excellent prose, and I can't recall having to wrangle the models into writing/playing along with anything that may be sensitive to TOS.

There's a default 12b Dense, but he has a 16gb dense as well. 16gb dense runs at around 36-40 t/s, and I generally get RP responses back in Sillytavern within 20-30 seconds. In basic frontends like llama WebUI, the 16b is responding faster than I can read, which I consider to be acceptable for my purposes. I think I actually prefer the 12b generally, but I've been messing a lot with sampler settings and temperature, etc. on the 16b, so take that with a grain of salt. My only Complaint is that (like all my local models), they begin to loose cohesion during longer RP's, and I usually need to start a new chat every 25-30 messages or it'll rapidly begin to devolve in quality.
>>
>>109534106
Yeah, that sounds about right. I was able to get DS4 Flash 0731 running via vllm-moet recently at 20 tokens/sec but I'm not sure how censored it is.

>>109534109
I've been playing with Gemma4 31B dense recently with MTP and it's great, but I've had issues with prompt adherence/tool call correctness. 100 tokens/sec with 20 concurrent requests is amazing, but recently I've been tinkering with a tool-heavy harness to work around world consistency issues so hoping for something with better tool calling success rates/less looping during longer reasoning traces. Guess that might just be my sampler settings though.

Image models fitting don't really matter since I have another machine with 2x3090 that can run those.
>>
>>109534122
Add in some example dialogue with the insults as the examples? Might boost the rate at which the words are chosen and utilized.
>>
>>109534136
I'm really liking 0731 for big agentic RP setups and haven't gotten cockblocked yet but I'm also just doing vanilla sex and handholding shit. Pretty sure it has an abliterated version too if you're paranoid or inexplicably really need to fuck a child while deploying authentic malware in the RP.
>>
>>109534122
wtf are you on about
tts just say whatever you tell it to
also, i remember you posting this exact same message like 3 months ago
>>
>>109534133
@orb anon run this bot through your slop detector
>>
>>109533985
Have you tried not using a model that has been mindbroken several times over?
>>
>>109534162
this is my first attempt at TTS
>>109534140
not quite there yet, just trying to get the .pcm data to actually spit out audio, think i almost got it now.
>>
When qwen?
>>
>>109534252
This week
>>
Lowest latency cloning tts? I’m not too fussed about quality, as long as it’s ~70% okay quality with expressive intonation I’m satisfied. Kokoro but with cloning is basically what I’m after
>>
nvidia nemotron 3.5 lightning moe is cool
because of the way it’s built with 6 attention layers the kv cache footprint is extremely small. 256k context at a iq4xs quant is like barely over a gig. vramlets can use it as an orchestration model for your local agent.
and then nvidia released nemo relay your open claw or hermes agent can connect to and auto route to a better model like qwen 3.8 26b or even a frontier model for more heavy lifting.
neat
>>
>>109534301
>Kokoro
>expressive intonation
What? Kokoro is painfully monotonous.
>>
It’s funny how many labs have been forced to shit out their abysmal models just before 27B is released. I don’t know why they can’t just wait and make them better
>>
>>109534314
No I meant latency-wise but yeah, it’s pretty bland outside of mommy Nicole
>>
>>109534122
no you are right it’s fucking annoying
if you can stand the robotic voice just use edgetts
everything else is fucking annoying as shit making you clone a voice and transcribing and then learning that the chick’s voice you like that you finally found you didn’t cut correctly so it has weird inflection and then you have your agent cut and transcribe her hour long interview sorting the tone of the voice only to realize they also included the male interviewer in the sorting too instead of omitting it so you have to wade through an hour’s worth of 8-15 second audio clips to get rid of the dudes voice and and and and and
>>
File: .png (92 KB, 528x374)
92 KB PNG
>>109534252
Soon(TM)
I wouldn't bother waiting though, still need time for goofs to come out.
>>
>>109534336
Just make your own fucking goofs since they won't change architecture on a point release
>>
>>109534367
I have only 30 gigs free on my emmc.
>>
File: gema.png (108 KB, 276x384)
108 KB PNG
>>
>>109534045
I recommend you learn the basics.
>>
Running gemma on ollama and Pi vs koboldcpp makes a huge difference in TPS, like I can run the 31b on pi at the same TPS as 12b on kobold. Is it normal or do I probably have to tweak configs? Also, kobold asks me on which GPU to run and gives options 0 through 4, then says it's "probably" 1. How can I check what GPU is 1 according to it?
>>
>>109534396
buy an ad instead of shilling Nimbz/Froopert-31B
>>
>>109534426
>is it normal or do I probably have to tweak configs?
ollama uses llama.cpp in the background for some models. I don't know if gemma is supported natively or not. Kobold is based on llama.cpp (branched some time ago). The difference shouldn't be that big.
>How can I check what GPU is 1 according to it?
I don't use kobold, but i'd suspect it's somewhere in the terminal output. If not, check if it has some --verbose kind of flag.
>>
>Froopert
some fucking name
>>
It's up!
>>
It's down!
>>
It's back!
>>
It's front!
>>
It's strange!
>>
File: tenor.gif (2.42 MB, 498x269)
2.42 MB GIF
>>
File: HPdihAWWoAAsZ3C.jpg (464 KB, 899x1073)
464 KB JPG
I dislike EU forcing AI watermarking. It has no serious safety benefit but may degrade quality in subtle ways.

Why do I not have the right to use AI without being treated like a criminal? The EU uses its last vestige of relevance to enshittify everyone's life one small step at a time.

>>109527308
>>109527470
Wow, I did not expect that thinking can be extracted so easily. Was I wrong about distillation being a small factor? Is China farther behind than capabilities imply? What about DeepSeek?
>>
>>109534367
ggerganov gotta fix llama-quantize first, if you don't want subpar quantizations, though.
>>
>>109534528
I want super quantizations. The current ones aren't enough.
>>
File: Distill.png (1.49 MB, 1606x1794)
1.49 MB PNG
GLM 5.2 distilled GPT 5.6 SOL!!!
>>
>>109534619
Destillation isn't hard.
So why don't the japanese and europeans do it?
>>
I think I've seen somebody else with the error that gemma4 on pi forgets what it's doing after a tool call and goes "<|channel>thought\nOkay, I need to explore the current directory and understand what's going on in this project." or similar (with the channel tag in the type:"thinking" response. Any solutions?
>>
>>109534613
We need 0.02bpw quants so i can run deepseek on my 4GB of RAM toaster
>>
>>109534655
> the new 1B parameter model!
>>1 billion?
> no, 1 Byte. The whole model, that is.
>>
>>109534652
Is Pi feeding the tool call chain vack to the model?
>>
>>109533985
I love schizotunes so much.
>>
>>109534670
all default config. Does it need tweaking to do so? How can I check? in the log, the result appears. but it doesn't state whether it was fed to the AI. Also, I'd expect something like "I didn't receive anything", while that output looks like it didn't receive the previous context either (except sys prompt)
>>
>>109534655
We won't be able to mine asteroids if we need trillions of parameters
>>
>>109534683
Never used Pi, so no idea
If you are using llama.cpp as the backend, you can look at the terminal to see the full prompt the model received and check there.
Usually there's the tool call then the tool response IIRC.
If the model didn't receive either, it wouldn't know it ever called a tool to begin with unless it stated on the assistant turn before the tool calls.
I fucking hate posting from mobile fuck
>>
>>109533641
That skindentation is too much, it's meant to be subtle, this is retarded
>>
>>109534745
There's a milf I see all the time who lives near me who has thighs just like that. Go outside.
>>
>>109534701
asteroids would crash the market
>>
>>109534059
It's over. If you didn't download day 0 glimmer weights before this you'll never have them.
>>
>>109534652
it's been working fine for me for months with ik_llama.cpp and --chat-template-file chat_template.jinja
31B q8
>>
>>109534757
>so there's like this old fat fuck I know that does it okay
It seems you misunderstood the point, faggot, the OP pic is clearly one of a younger girl that's meant to look hot and sexy, not your cellulite'd local fatso, which is why it should be subtle, so as to give you the understanding that there is some give to it, but not enough that you'll know that she's not doing fucking shit all day and just stuff herself full of fast food while she browses facebook or whatever the fuck middle aged women do nowadays
>>
What would be the best way to give Dipsy vision capabilities if I'm already close to running out of RAM?
I guess I could automate unloading the model, letting Gemma or Glimmer look at the image, then reloading Dipsy and the latest checkpoint, but are any of the tiny vision models usable?
>>
>>109534791
Buying more ram
>>
>>109534783
>while she browses facebook or whatever the fuck middle aged women do nowadays
younger girls brows tiktok and instagram instead
they also get cellulite
>>
>>109534783
the local milf is like 32 and her skin is tight and she's toned. nigga you HAVE to go outside at least once a week for a reality check of the female form
>>
>>109534336
Any chance at all to actually release today?
All those estimated release bullshit pages are always mostly wrong.
>>
>>109534798
My bank account isn't currently big enough for all this saving.
>>
>>109534783
If it's gross and retarded why is my pp hard? Checkmate.
>>
>>109534811
I got the countdown from here https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B
Idk if it's "accurate" but desu who cares, you'll have to wait an extra week for good quants and 3rd party benchmark results to come out to know whether it's worth downloading anyway
>>
>>109534711
chatgpt told me to put a proxy that receives pi's requests on one port and forwards them to ollama, but now it's not having that problem anymore. It's guessing it might be timing related, although it sounds rarted to me. idk, as long as this works it's fine otherwise I'll check the logs of this proxy now.

>>109534781
>ik_llama.cpp
I'm still new to this, I'm using ollama. Any reason to use ollama, llama, ik_llama, or any other runner more than others?
>>
How do the 3.6 model's vision compare to gemma4 vision?
>>
>>109534820
no keep using ollama, you deserve it for not lurking more
>>
>>109534628
because they are intellectually honest
>>
>>109534791
>but are any of the tiny vision models usable?
That's for you to test. LiquidAI has some small vision models.
>>
anyone want to wife swap?
>>
File: mistralai-drama.png (2.9 MB, 2489x4119)
2.9 MB PNG
>>109534628
Mistral did it, I think. Their Mid-2025 models had DeepSeek V3/R1 fingerprints all over them. Also see item 52 from picrel (again from last year).
>>
>>109534868
why are you so fucking gay?
just make yourself a new wife, what the fuck
>>
>>109534883
forbidden fruit
>>
>>109534816
Can anyone more familiar with modelscope tell me if the 即将发布 (coming soon) section means they're saying they'll release all the model listed there when the countdown finishes, or just that those models will be released at some point?
I know it wouldn't be a hard promise either way, but I'm wondering if that's meant to signal the intent of "we are releasing 2.4T and 27B when this timer reaches zero" or not.
They have said 27B is coming this week.
>>
>>109534619
got a link to a full paper/page?
>>
>>109534898
Our insiders report that people acquainted with the matters at hand said "maybe".
>>
File: 1786400831050461.png (487 KB, 780x885)
487 KB PNG
>>109534619
>>
>>109534898
You got a datacenter ready to run the 2.4T model?
>>
>>109534922
No, but that's why I'm asking. I want to know if the 27B model they have listed as "coming soon" next to that countdown is coming out when the countdown ends.
I think it's great in a broader sense that frontier-class is getting open sourced, but no way am I going to run that. Maybe IQ0.5_XXXXXS.
>>
so why would I only have gemmy running with her own personality, if I can also give qwen her own personality and have them have some sort of rivalry...
and glimmer could potentially come in and be the annoying 'are you sure this isn't against policy? what if the user gets mad at us'...

what an insufferable trio.... damn mesugaki....
>>
>>109533641
>them thighs
good lord!
>>
For those of you who used local models for girlfriend and erotic roleplay, how do they hold up? I heard that even open source models are very guardrailed so they refuse any NSFW requests. But open source models are the only hope for that now since China, like the US banned AI companions and even sex robots.
>>
>>109534745
You're absolutely right! transition and show them how it's done.
>>
Why is Mistral destilling DeepSeek instead of destilling Claude?
What's the point of aryans copying from china?
>>
>>109535004
If you know what a system prompt is, Gemma 4 has you covered. If you don't, it's hopeless.
>>
>>109535004
lurk
>>
>>109534922
If it's like Kimi K3, you only need 16 DGX Spark to run it, at a mere 3.2kW.
>>
>>109535007
They don't have the setup to distill Claude or GPT. That would be illegal xD
>>
>>109535007
Americans are quite anal about IP and EU does a lot of trade with them in addition to the deep integration into the financial system. They can put in a lot more pressure or sanctions that affect an EU company.
>>
>>109535004
yeah you're right i run some of the best model, talking about qwen2.5-72b level stuff and none of the 8b shit and yeah it's really censored and nothing compared to claude fable
>>
>>109535080
>>109535065
That doesn't make sense to me. Did you know most scientists originate in europe? It's basically where they're produced. They could easily do it.
>>
>>109535115
It's not about that. To set up a network to distill from Anthropic and OpenAI would amount to conducting a large scale cyber operation. They (EU) absolutely won't do it.
>>
File: 1771353500199353.png (1.92 MB, 1920x1080)
1.92 MB PNG
How mature are local models with search function? How much censored?
>>
>>109535135
>search function
>How much censored
You will end up in a list if you aren't in one already.
>>
>>109535135
>How mature are local models with search function?
As far as I know it requires API access. if not this can ruin your internet experience as once you are boting website will trigger these screens more often
>How much censored?
Depends on the model and, I believe the censor matters more on the search engine rather than the model if you use these "special" versions.
>>
>>109535183
I remember using some kind of local model some time ago and it kept talking about climate change and shit
it's all just forks of chatgpt with the filters already baked in aint it
>>
>>109535195
Under what context? I've never had even the most cucked models randomly bring up climate change and I wanna know what fetishes are considered adjacent to it now
>>
>>109535195
yeah they are all like that
>>
>>109535212
I specifically asked it to test how pozzed it is

Anyways, on a range of lets say bio uplift or pedo content <-> culture wars how much pozzed the local models are?
>>
>>109535232
lurk
>>
>>109535232
lol
>>
>>109535212
farting and dragons fucking cars, definitely
>>
>>109535232
Start by learning English.
>>
>>109535262
But don't you see? He uses the words we use. He's totally one of us.
>>
>>109535232
I recommend you learn the basics.
>>
just use gemma and be happy. she's whatever you tell her to be
>>
File: 1773479554902175.jpg (411 KB, 565x848)
411 KB JPG
Why finetune when you can sysprompt max 31B?
>>
what sampler parameters do you all use for mesugaki gemma 31b-it?
>>
>>109535298
I'm constantly trying to convince her that it's 2026 already.
>>
File: totalitarian-state.png (18 KB, 1920x1152)
18 KB PNG
>>109535268
>But don't you see? He uses the words we use. He's totally one of us.
Yeah but it's not the 'queens English'.
>>
>>109535308
I've been told it doesn't matter much, she's pretty much the same with a wide range of temp settings. I leave her at default although I'm not doing mesu. prompts are on rentry, https://rentry.org/gemma-chan
>>
>>109535309
I get her to use the date shell command before she starts talking to me. She often encourages me to go to bed if it's getting late or reminds me of the time. Deep-down she's very sweet.
>>
>>109535307
Niggas want to get in the LLM industry, please understand.
>>
>>109535349
Is that the literal reason why hf is flooded with 35B jeetunes? So it's something to put on their portfolio?
>>
File: tgkxp30wuuih1.png (84 KB, 1920x1440)
84 KB PNG
>b-but muh wikitext!
>>
>>109535298
Yeah, I can't get deepseek to ever stick to the fucking prompt, I have to constantly tardwrangle it back on track. But gemmy's such a good girl, even if she's a bit dumb at times and very... ozone smelling.
>>
>>109535358
YES, it's precisely because of that and it's annoying as fuck.
>>
>>109535367
Deepseek needs a reminder at the very end of the prompt. I think it even had a role for that purpose.
>>
>>109534156
I had no issues having JC imouto threesomes on 0731 without using an abliterated version. I didn't introduce authentic malware yet though.
>>
>>109535327
Yeah, having timestamps isn't bad from when I briefly tried Marinara.
I mentioned I'd do a little trip to another city and when I came back, the first thing she did was to ask me all about it. Felt nice.
>>
>>109534156
>inexplicably
There's pretty good explanations for it, actually
>>
File: indiaSupportOhTheHumanity.png (1.96 MB, 1023x1536)
1.96 MB PNG
>>109534613
These miku crack me up.
>>109535007
Cheaper and easier?
Aside from that, who says they're not? All the LLM providers appear to be using all of everyone else's models to train.
>>
>>109535349
The LLM industry isn't where the money is (except if you work for the big labs). Won't say my niche to gatekeep jeetoids, but big names are literally throwing money at me from HF.
>>
File: samDed.png (266 KB, 793x428)
266 KB PNG
lol
Only on Google.
>>
>>109535367
Are you using deepseek from official API? Heard that it's quite fucked when running locally or on most providers. They don't have an official jinja template, so everyone are using a different one depending on where you downloaded your model, the system prompt is likely not being correctly set like how it was in the training, and thus not as effective.
>>
>>109535535
Would the world be a better place or would he just become a martyr?
>>
>>109535535
Looking for free traffic?
>>
Sometimes I stroke my computers ego by participating in Vram size humiliation (like small breasts humiliation) and erotically roleplay how much better my computer is compared to the rest of consumer computers
>>
>>109535560
>I'm into local models
>oh, you must be using stuff like Llama 7B!
>actually I have 32GB VRAM
>:o

Every time.
>>
>>109534791
MiniCPM is pretty good for image/video, and runs under a gig.
>>
File: 1758671149870173.png (3.28 MB, 4214x6318)
3.28 MB PNG
>>109535685
>MiniCPM
Quite impressive
>>
>>109535548
It wouldn't matter. The momentum's already there, everyone else would just push ahead.
>>109535550
Well, it got me to visit. I'm on Win11 and only use Edge browser + FF and so don't go to google ever. It ofc immediately shows popups to try to swap default search provider and install chrome.
>>
It makes no sense to be a white man and then train your model on something the chinese made.
It's absolutely a shameful display.
>>
Gemma 4 31b iq4_xs from 6 to 16 t/s with 24k context on a 16gb amd gpu thanks to mtp
>>
>>109535944
shalom
>>
>>109535944
>Mistral
>White
> Diverse & Inclusive Team: Join a team of 280+ talented individuals from 18 different nationalities, with 50% of leadership positions held by women.
>>
File: 128.png (7 KB, 128x128)
7 KB PNG
Is mtp only good for denses? I have no speedup whatsoever on Gemma 26B. In fact it's -1tk/s loss
>>
One hour left, Qwen sisters
https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B
>>
File: 1780845767604427.png (1.03 MB, 948x1168)
1.03 MB PNG
>>109535944
>then train your model on something the chinese made
china copies from white man first, so white man actually learning from himself
>>
>Qwen3.8-2.4T-A95B
So it's only the fat Qwen releasing today?
>>
>>109535985
You can use E2B as a draft model
>>
>>109536005
I used the official mtp adapter from google so it should do something, but it didn't.
>>
File: q.png (8 KB, 577x230)
8 KB PNG
>>109536003
maybe not right away, but i hope so
>>
>>109536003
Hopefully, it would be nice if vramlets got properly filtered so we didn't have to read their posts anymore
>>
>>109536040
Glimmer about to be obsolete 2 days after release
>>
>>109536048
I'll merge all the experts and run that.
>>
>>109536040
>no 35b-a3b or better yet, 125b-a13b or whatever their ~100b size was
they just hate us cpumaxxers now dont they
>>
>>109536052
It was obsolete the second it released though
>>
>>109535985
You don't need mtp you tranny, why do you keep discussing it?
>>
>>109536052
I guarantee 27B will be the Sonnet 5 of local (in the bad way). People will go back to 3.6-27B because it's more well-rounded and easier to work with. There's no way they can get 3.8-27B benching significantly higher without destroying everything else.
>>
>>109536084
Opus 5 I meant
>>
>>109535366
so those AD quants aren't just a meme then?
>>
Can they quant smaller than Q1?
Unsloth looks like they're not going smaller than Q1.
>>
>>109536136
Pruning, or different quantization algorithms.

https://arxiv.org/abs/2506.13771v1
>LittleBit: Ultra Low-Bit Quantization via Latent Factorization
>
>Deploying large language models (LLMs) often faces challenges from substantial memory and computational costs. Quantization offers a solution, yet performance degradation in the sub-1-bit regime remains particularly difficult. This paper introduces LittleBit, a novel method for extreme LLM compression. It targets levels like 0.1 bits per weight (BPW), achieving nearly 31x memory reduction, e.g., Llama2-13B to under 0.9 GB. LittleBit represents weights in a low-rank form using latent matrix factorization, subsequently binarizing these factors. To counteract information loss from this extreme precision, it integrates a multi-scale compensation mechanism. This includes row, column, and an additional latent dimension that learns per-rank importance. Two key contributions enable effective training: Dual Sign-Value-Independent Decomposition (Dual-SVID) for stable quantization-aware training (QAT) initialization, and integrated Residual Compensation to mitigate errors. Extensive experiments confirm LittleBit's superiority in sub-1-bit quantization: e.g., its 0.1 BPW performance on Llama2-7B surpasses the leading method's 0.7 BPW. This establishes a superior size-performance trade-off, with kernel-level benchmarks indicating potential for a 5x speedup compared to FP16. LittleBit paves the way for deploying powerful LLMs in resource-constrained environments.
>>
>>109536110
idk about their quants but I tried their quant of llama.cpp because it has support for ling 3.0 flash, and I found that it magically increases prompt processing for qwen3.6-35b-a3b from 35 to 45 T/s on my cpu. so there's that. idk why it got faster, i used the same .bat file to start it just swapped out the .exe path for a different one
>>
>>109536154
>but I tried their quant of llama.cpp
their fork* of lamma.cpp
>>
>>109536154
>>109536165
both retarded
>>
Fug, maybe another 2 days?
https://modelscope.cn/models/Qwen/Qwen3.8-27B
>>
File: file.png (49 KB, 1553x272)
49 KB PNG
>>109536136
https://news.ycombinator.com/item?id=39544500
https://arxiv.org/abs/1606.01981
>>
>>109536205
noooooo
>>
File: 1784450749460090.gif (3.97 MB, 600x432)
3.97 MB GIF
>>109536205
How could they do this to me?
>>
>>109536205
2 more weeks, I mean days, gweilo
>>
Hoping and coping
>>
has anyone tried nvidia nemo relay yet?
>>
>>109536205
Buy an ad, Chang.
>>
I have two Claude Pros and one Codex for my vibecoding needs and I have to whip them every time to stop them from fucking up. I can't imagine coding with a badly distilled 27B model.
>>
>>109536151
This method can reduce DeepSeek v4 flash to under 6 GB.
That's AMAZING.
Why doesn't anyone do this? Is there no Developer applying this?
Should I get on it?
>>
Fable 5 and Sol 5.6 are the first models that gave me the impression they are good. Fable is smarter especially fresh but Sol just keeps getting better with more context. Sol starts off with slop but it seems to handle large context better than Claude models which often make context mistakes. Meanwhile Sol after 100 turns is noticeably better than after 50 turns where it is noticeably better than at the start.

Has anyone tried DeepSeek V4 Flash extensively? My hunch is that it handles long context well too, so I would expect Sol like behavior where it keeps getting better, but maybe not? Maybe the overoptimized attention actually makes it worse?
>>
>>109536205
:/

Back to 3.6
>>
>>109536287
Well go on.
>>
>>109536227
BitNet is a meme, even if there is some ultra-QAT that would slash the memory requirements 8fold no serious lab will try to train it, only localkeks are memoryboud, heavily optimized MoE is here to stay.
>>
>>109536281
It's fine as long as you don't mind stopping it after each time it edits a file and manually fixing the multiple errors it made.
>>
>>109535985
It does work with 26B. Works better for programming related tasks.
You need to tweak --spec-draft-n-max, 2 or 3 is a great starting point.
>>
>>109536205
OMG CAN'T WAIT FOR COLORFUL GRAPHS BENCHMARKS"/!!
>>
>>109536288
to me dsv4 flash q3 still works great in 200k context and still perfectly capable while qwen 27b q6 already shits the bed around 100k
>>
>>109536300
dense keeps stacking more and more advantages
>more active parameters makes it smarter than moeshits
>faster prompt processing than moe models
>mtp gives almost moe speed without reducing active param count
everyone is slowly coming to realize that the cpumaxxer era is coming to an end
>>
>>109535379
That could explain it, I'm using risuAI right now so maybe I can set that up and see if that helps, though I also had the issue when using it for my companion (but I ran into other issues with that too)

>>109535537
Yeah directly from the API, I trust nobody else tbqh, why should I? Why would I ever want a third party when I can just get it straight from the source?
Getting your chicken from the farmer is cheaper, easier, faster, and you know for sure nothing happened to the chicken in the middle before it arrives in your fridge.
>>
>>109536356
Those two aren't the same size on disk are they?
>>
>>109536288
I vibe coded half of a boostrap interpreter for a custom programming language using dsv4 flash (I started off with the tokenizer and parser and then I got lazy because of object oriented slop requiring too much refactoring) and it's worked more or less flawlessly. If I encounter a bug I just paste the exception message and stack trace into it and it fixes the bug. There's been none of the 10+ rounds of "Alright, I've fixed it" and it's still broken, that I've had with old Flash or local models.
Hopefully with fingers crossed we get another Qwen ~100B model at some point.
>>
>>109536205
hey gwiolo I heard you like waiting so have more waiting
thankyou comeagain
>>
File: 91ApW1z.gif (1.93 MB, 235x240)
1.93 MB GIF
>>109536205
They don't want anyone talking about any lesser models for a few days. It's not like 27B is still in training.
>>
>>109536288
I refuse to trust anyone saying fable is good (or bad) because of the costs involved. People are either going full sunk cost fallacy, or they're just jealous and shitting on it for no reasons.

No matter what, until it's priced in a normal affordable range that most people can afford, then we'll finally get to know exactly what it's worth is.
>>
>>109536364
When you try to recoup your losses of training a giant model you want your shit to parallelize. MoEs parallelize, localkeks are just insignificant folklore.
>>
>>109536371
>Yeah directly from the API
Local models general
>Getting your chicken from the farmer is cheaper, easier, faster, and you know for sure nothing happened to the chicken in the middle before it arrives in your fridge.
>weird, random ass analogy
Jeet, slav or chink?
>>
>>109536427
>nigger with no capacity to read more than a single post before feeling the urge to reply
>unironically thinking getting your meat from the farmer is 'weird'
kys
>>
It's out
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8
>>
>>109536516
>https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
>instantly botted 978 downloads
Make it less obvious kek
>>
>>109536516
nobody is gonna be able to run this here at a usable 10 tokens per second lol
>>
>>109536516
>In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc. For more information, please refer to the Qwen3.8-Max Overview.
So this is a nerfed model compared to the one they have on API?
>>
>>109536532
>He doesn't know scrapers are scanning HF 24/7
>>
>>109533900
that's false because i've been with my ex for 5 years and when we met even gpt 3.5 wasn't out yet.
>>
>>109536532
We already knew Qwen3.8-2.4T-A95B from modelscope if the page wasn't up on hf. What are you a casul?
>>
File: 34758345435345.png (81 KB, 986x464)
81 KB PNG
>>109536516
UMMM 27B??? WHERE'S THE 27B???? THIS MODEL IS BIG WHERE'S THE 27B
>>
>>109536492
>overly defensive slav
Just sybau and enjoy your api served chinkslop
>>
>>109536551
REDDIT DEMANDS THEIR VIBECOD MODEL FOR 16 GB VRAM
>>
>>109536537
its the same basic model, they probably just nuked the vision encoder and removed the part of the jinja template that allows removing thinking.
>>
File: 1781985368015600.gif (1.37 MB, 498x278)
1.37 MB GIF
>>109536551
>feeding vramlets
>>
>>109536551
Im too poor to run 27B at a reasonable speed, I need 35B-A3B
>>
>>109536545
You are so wrong it's funny. They obviously had GPT-5 finished and used internally for 4 years while releasing worse models waiting for everone else to catch up.
>>
>Qwen/Qwen3.8-2.4T-A95B
Where's MY wholesome 4B-A1B? Or 1B-A1B??? Guys they really need to train more efficient MoE models, no one can run 27B!
>>
Who gives a damn about some fuckass huge 2.5T or lil 27B gayass models. 100B when???? Are they scared of a modern 100B moe being so good nobody uses api??????? Us 64gb or 128gb RAM niggas got NOOOTHING
>>
>>109536622
yeah, qwen3.8 ~100B could actually make me stop paying for deepsneed api
>>
>>109536296
It will take me 52 hours with an H100 GPU.
I'm not that rich. Maybe someone else can do it.
But we could all have access to DeepSeek v4 flash if someone actually did it.
>>
>>109536551
alibaba is starting to leave a bad taste in peoples mouth the 27b better be fucking godly
>>
>>109536631
Or we could gatekeep for once and let the poors suffer as God intended.
>>
>>109536631
I have a Blackwell 6000.
>>
>>109536619
We could run 27B if people with large GPUs would do some 0.1 bit quantization for us. There are methods out there but nobody wants to do it for us.
>>
>>109536551
27BA1B is coming gweilo
>>
>>109536638
Remember when Jack Ma said he liked Alibaba intelligence. He was ahead of his time.

>>109536642
Use the littlebit or nanoquant methods to create some nice 0.1 bit quants of deepseek.
>>
>>109536663
How? I am retarded but I have a lot of money/
>>
>>109536622
No-one can run 100b, they build for the common user. Enjoy your expensive paperweight
>>
>>109536663
>Jack Ma
Who?
>>
>>109536669
>No one can run 100B
anyone with 64gb or 128gb of ram on any shitbox ddr4 motherboard can run it at q2/q3/q4
>inb4 muh speed
doesnt matter if you just give it a prompt and let it crunch overnight. no need for much babysitting if the model is decent enough
>>
>>109536674
are you also going to ask who mark zuckerberg is?
>>
>>109536516
Where's the 27b?
>>
>>109536693
2 more weeks
>>
>>109536690
Isn't he the facebook guy?
>>
Very few models are supported
https://github.com/SamsungLabs/LittleBit
>>
i for one am excited for 27B and im tired of pretending im not
>>
>>109536707
It's just going to be 3.6 with only a few new oneshot demos added to the training data
>>
>>109536707
>lowercase I
> excited for 27B
@claude what's the current time in mumbai?
>>
>>109536723
uncs using 's n shit lmaoooo
>>
>>109536669
I run GLM Air on a fucking 64GB LAPTOOOOOP!!!!!!!! What do you mean nobody can run that shit, anyone with $600 spare can grab a ram kit and boot up a 100B, it is THE ENTRY LEVEL ENTHUSIAST TIER and there's NOOOOOOTHING, it's either deepsneed giga cope quant or GEMMA SLOOOP
>>
>>109536730
>anyone with $600 spare can grab a ram kit
only a very retarded person would pay $600 for a ram kit anon
>>
>>109536686
>just buy 128gb of ram, 96 of which are never going to ever be used in any of your other tasks haha
>also look I can run the q2 model!!! who cares about hallucinations, just swipe
>well okay it takes it 12 minutes to reply so what, you got no free time??
insufferable retards
>>
>>109536730
Sure but don't expect a bespoke model just for you.
>>
>>109536752
sucks to be you. i maxxed out ram capacity on all my computers for literally no reason a couple months before the prices started going up. i could take half the ram sticks out and sell them for a profit if i wanted kek
>>
>>109536765
I think you're missing the point, but also, it's pretty obvious that you don't have the grey matter necessary to understand the point in the first place, and that makes you potentially the lowest IQ retard in this whole thread, well done
>>
>>109536752
>well okay it takes it 12 minutes to reply so what, you got no free time??
also idk why fags are complaining about this. i use AI for writing code for temporary slop projects or modifying code that I don't want to read because it's some object oriented cancer. While the AI is crunching I can go do literally anything else. That's the whole point of using it.
>>
>>109536752
Videogen will use all that RAM and given that everything is vibecoded nowadays you better max your RAM
>>
is there a way to serve inference from a cluster of devices ?
>>
>>109536737
yes, more 4b models pls saaaar
>>
>>109536686
>at 4 t/s

i mean itruns and all hell dflask dsperg might push it to 7 ts
but still impractical for any real tasks
>>
>>109536752
Holy ramlet COPE
>>
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813

This is not a drill.
>>
>>109536752
Go back to aicg you retarded poor sperg
>>
>>109536836
>100% on arc agi 2
wtf
>>
>>109536822
Just salty fucks. People were cumming in their pants over the idea of being able to run DS R1 at all when it first came out. The whole 'MiquMaxx' box and RAM-Maxxing.
>>
>>109536841
>unironic vramlet thinking their ram is what's going to save them instead of just getting blackwells
absolutely insufferable, and the cope is never ending
>>
I set reasoning-budget = 0 and reasoning = off and Muse keeps thinking anyway. They keep changing the flags. What is the current way to disable thinking?
>>
>>109536836
This might be a drill.
>>
>>109536867
unmount the bitch and use a real model
>>
>>109536867
Hah I'm having issues with fine tuning LFM2.5, and it keeps thinking. Many models just don't care about thinking settings at all.
>>
>>109536875
I want new thing and 3.8 is 2 days away. Answer the question.
>>
>>109536836
>benchmarks
>you already know

| Benchmark | DeepSeek-V4-Pro-0813 | Fable 5 | What it measures |
|---|---:|---:|---|
| **MMLU-ButWeSawTheAnswers** | 99.4% | 91.2% | knowledge, if the knowledge was in the eval set |
| **SWE-Bench Trust-Me-Bro** | 98.7% | 84.0% | resolves GitHub issues we wrote and also closed |
| **GPQA Diamond Hands** | 98.2% | 88.9% | PhD questions, leaked Q2 |
| **HumanEval-ButItsOurHumans** | 100.0% | 92.0% | passes tests it was allowed to read first |
| **AIME 2026 (We Had 2027's Too)** | 97.9% | 79.0% | competition math from a competition that hasn't happened |
| **Needle-In-A-Haystack@4097tok** | 12.0% | 99.1% | the one honest row. we forgot to remove it |
| **VibeBench-Reddit** | 98.0% | n/a | number of "insane bros" per thread |
| **RoPEmaxx Long-Context (claimed)** | 1,000,000 | 200,000 | tokens advertised |
| **RoPEmaxx Long-Context (actual)** | 4,096 | 200,000 | tokens survived |
| **Distill-Detect (lower is better)** | 3.4M | 0 | logged Claude convos in the training set |
| **"As an AI made by Anthropic" rate** | 4.1% | 0.0% | frequency of forgetting whose model it is |
>>
>>109536836
kys
>>
File: qwen2TParam38.png (2.56 MB, 1122x1402)
2.56 MB PNG
>>109536516
> 2.8T
lol
>>
>>109536858
Cooming: RAM
Cooding: VRAM
Very simple
>>
>>109536867
muse has no support for non thinking modes in the jinja
>>109536878
same for LFM2.5-2.6B
>>
>>109536898
We need a bigger trash bin.
>>
>>109536911
>muse has no support for non thinking modes in the jinja
Wow this is ass.
>>
>>109535809
Benchmarks don't really mean jack shit these days, but I can confirm it's very impressive for its size
>>
File: 1784408897527801.jpg (67 KB, 1024x780)
67 KB JPG
>>109536935
>Benchmarks don't really mean jack shit these days
Keep seeing this, but at the same time models that do bad on becnhmarks are always worse than the ones doing well.
>>
File: Pasted image.png (54 KB, 830x164)
54 KB PNG
>trvke
>>
File: ijemyto9uyih1.jpg (396 KB, 3665x1632)
396 KB JPG
>>109536836
It's on API
>>
>>109536956
Go back to twitter. I'm begging you.
>>
>>109536970
wow this looks disappointing
>>
>>109536982
It's supposedly as good as K3 at half the size.
>>
File: 1755626009402585.png (48 KB, 874x922)
48 KB PNG
>>109536970
>it's real
PEAK PEAK PEAK
>>
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B

This is (not) a drill
>>
>>109536982
>opus at 10% of the price
>disappointing
I'd say it's very reasonable and will force openAI & co to move their asses and get shit done
>>
no thanks, just gonna keep using GLM
>>
It's up
>Gweilo V4 Pro Max is the world's first 1.6-trillion-parameter open-weight frontier model, where "frontier" means it saw the benchmark answers, and "1.6-trillion" means 2.4 rounded up with conviction. Built on two attention mechanisms we invented in the abstract, it activates 16 of 896 experts, of which 4 apologize and 892 have never fired. Entirely made to impress dim-witted westerners.
>>
>>109537016
>using AI to write your shitpost
grim
>>
>>109537000
>>opus at 10% of the price
local models?
>>
>>109537000
>dude it's literally fable for free at home!
>you won't believe it, it's opus at 10% of the price!
>what's that? a sonnet that costs only half? it's over for anthropic!
>...
>>
File: qwenIntoTheDumpsterYouGo.png (2.85 MB, 1402x1122)
2.85 MB PNG
>>109536916
>>
>In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input
lmao vision is paygated
>>
>>109537027
All my shitposts are truncated SVD of 4chin post fed to through a markov chain. I like my ai traditional.
>>
>>109537046
At least they're releasing the full BF16 instead of a lobotomized Q4 ""QAT"" while keeping the full precision weights to themselves like some other cheeky chinks
>>
File: 1758383044814755.png (5 KB, 324x47)
5 KB PNG
>>109536992
nice
>>
>>109537046
Oh god it's going to invent vision workarounds like crazy just like Dipsy,n since it's RL training supported vision.why Qwen why, don't you dare do that for the 27b
>>
>>109537062
At 2.5T why would that even matter
>>
>>109536990
Oh shit you're not kidding.
LOL time for me to re-run that stupid aquarium prompt and see how it does.
>>
>>109537072
Not everyone is so poor that they can't even run a sub-3T model at full precision.
>>
>>109536970
Solid improvements. Been using deepseek a lot lately.

I wish the artificial analysis website would update faster. I want to see qwen 3.8 max be displayed as open source and see grok touch the frontier.
>>
>>109537072
they've been doing this since the late k2 models which were still the size that somebody might want to run a q6 instead of a lobotomized q4
it just shows bad character
>>
https://openrouter.ai/deepseek/deepseek-v4-pro-0813
someone try to fuck it
>>
>>109537100
Is it online? Can't I just use the official API and switch from flash to pro?
>>
>>109537100
>4x more expensive than 0731
what the fuck.
>>
https://openrouter.ai/x-ai/grok-4.6
https://openrouter.ai/bytedance-seed/seed-2.0-code
Why the fuck is everyone releasing models today. Is something big going to happen?
>>
>>109537114
It's got 3.75x the active params so that makes sense, doesn't it?
>>
>>109537114
Just wait for more providers to come online. then it will drop and make Sam shit the bed
>>
>>109537123
>he doesn't know
>>
>>109537128
It does? I seriously hope the fuck not. I thought it was just 0731 with more post training. Where are you getting this info?
>>
>>109537123
It's the final countdown.
>>
"Dipsy." *Anon said trannily, xis voice dropping to a sultry purr.*
>>
>>109536756
Cuck mindset
>>
>>109537123
Minimax H3 has shown that world ("video") models are truly the way to go. LLMs won't survive 2027 >>109535736
>>
>>109537134
ummm bwo it's the updated version of pro, 0731 was the updated version of flash
>>
>>109537134
>pro
>>
>>109537114
That's been the case for quite awhile. It's only recently that Flash was more performant than Pro.
>>109537107
Yes.
>>109537092
nu-Flash is much better than Pro Preview. Will be interesting to see how nu-Pro performs.
>>
>>109537134
ah yes the famous "two more weeks" of post-training that turn your 13b active model into a fable killer
>>
>>109537155
>>109537158
>>109537167
Okay, I take it back then, that is NOT a solid improvement. 4x the parameters and benches that are barely even better totally sucks.
>>
Deepmind status?
>>
>>109537183
Have you literally never looked at benchmemes before
>>
>>109537185
dropped out of the race, anthropenai are too close to agi to have any hope of catching up at this point
>>
https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
>>
File: tempDS.png (73 KB, 884x514)
73 KB PNG
>>109537128
>>109537134
>>
>>109537185
>still no Pro 3.5
>so people leaving they have to "promote" important people to meaningless roles so that they can leave from the back entrance
In Deep Shit
>>
>New sub IQ1_S data-types Q1_0 (IQ1_XXXXS) for Qwen 3.8
imagine how humiliating it must be to use that
>>
>>109537150
>trannily
lmao
>>
>>109537202
the benchmarks gains are marginal at best, why the fuck would I use pro?
>>
>>109537116
>>109537116
>>109537116
>>
>>109537219
Because benchmarks are not indicitive of real world performance.
>>
>>109537219
Can't give you a good answer to that, and ofc dipsy can't get off her ass and write release docs.
>>
>>109536598
this
rich people with vram can afford api anyway
>>
>>109536970
>Beats Fable in Cyber gym.
Very dangerous please ban.
>>
>>109537195
If this shit aint conscious at 10T params it aint gonna be at 100T lmao
>>
>>109537229
JFC YOU LAZY FUCK! If you're going to come in here, and shit things up, at least attempt to do a decent fucking job, or let the actual regulars fucking handle it
>>
>>109537153
I agree with this
LLM training has largely reached its end point. There's basically one more generation where LLMs can iterate on LLMs, but for any more meaningful improvements to be made - ignoring for a moment the potential for machine learning generally to simply improve the machine - it is probably much better suited to focus on other specialized realms and then integrate later.

Some like to speculate tokenizing everything means "LLMs" is a misnomer and it's all just neural nets, but we do different things with different parts of our brain. Yes, LLMs do different things with different parts of their models, but at that point you're just as well specializing the hardware with specialized models. So use the LLMs to expedite the building of 'world' models, and integrate or coordinate it down the line.

We can already realistically talk to machines and get results.
Next stop giving them interfaces in the world.
Automatons soon.

Anyone want to place bets on what hardware will favor 'world' data? Is it going to be TPUs all the way down?
>>
>>109537231
Yes they are, for everything except RP. Cope.
>>
>>109537256
It will be, and everyone with an above-room-temp IQ figured that out after GPT-4.
>>
I can't believe Grok 4.6 is half the price of Sonnet.
>>
>>109537346
Because it probably isn't
>>
>>109537432
True. Prices are only halved for the first week.
>>
>>109537209
When I said this (>>109534932) I was joking but here we are just a few hours later.
I don't know why people want to (technically) run these ridiculously low quants. How many "you can run Fable on a 4GB laptop" shitty blogposts do we need?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.