[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109606160 & >>109601505

►News
>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads
>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2
>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608
>(08/16) koboldcpp-1.119 prebuilt released with H3 and Glimmer support: https://github.com/LostRuins/koboldcpp/releases/tag/v1.119

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: edit_00014_.png (967 KB, 1024x1024)
967 KB PNG
►Recent Highlights from the Previous Thread: >>109606160

--User feedback and developer response regarding CoomKit frontend improvements:
>109606856 >109606934 >109606973 >109608477 >109608568 >109607013 >109607095 >109607178 >109607307 >109607220
--Comparing quantization resilience and VRAM requirements for Kimi and GLM models:
>109607752 >109607910 >109607956 >109607979 >109607999 >109608020 >109608100 >109608236
--Comparing R9700 and dual RTX 3090s for local hardware value:
>109609608 >109609637 >109609868 >109609884 >109609923 >109609948 >109609957 >109609982 >109610003 >109610020
--Benchmarks and reactions to Liquid AI's LFM2.5-DSpark release:
>109607030 >109607066 >109607088 >109607108
--Qwen 3.8 excessive reasoning length and benchmark-inflated settings:
>109607264 >109607275 >109607276 >109608701 >109607680
--Testing Ox Alpha performance and filtering via roleplay and definitions:
>109606929 >109606946 >109606960 >109607101 >109607503 >109607938 >109607154
--Visualizing model hallucinations through geographical world map tests:
>109606892 >109606902 >109607321 >109608579 >109608661 >109608273
--Sentimental value and technical longevity of Gemma 4 31B:
>109606611 >109606740 >109606785 >109606829 >109606889 >109606913
--Evaluating dflash2 utility and memory constraints for local LLMs:
>109608673 >109608698 >109609062
--Google celebrates 1 billion Gemma downloads and releases awesome-gemma repo:
>109606338 >109606385
--Nvidia strikes $6 billion licensing deal with Poolside AI:
>109606940 >109606959
--Sharing workflows and tools for creating Gemma-themed anime art:
>109606300 >109606373 >109606396 >109606496
--Logs:
>109606958 >109607101 >109607503 >109608185 >109608453 >109608696
--Gemma, Teto, Miku (free space):
>109606311 >109606338 >109606373 >109606385 >109606922 >109606901 >109608805

►Recent Highlight Posts from the Previous Thread: >>109606164

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
anyone have the setup for nala and cockbench? i want to test some models
>>
>trying out Ox Alpha to figure out what model it is
>it refuses smut between consenting adults because an "old cheerleader uniform" is mentioned and any reference to cheerleading immediately makes the adults underage
I sincerely hope Dario succeeds in killing all distillation efforts or else the future looks grim
>>
File: Nala.png (731 KB, 860x864)
731 KB PNG
>>109610466
Nala card: https://files.catbox.moe/yw8l8c.png
>>
>>109610440
https://x.com/osanseviero/status/2090678546310017434
>The Gemma community celebrates tonight

I'm guessing no special announcement of note happened.
>>
>>109610508
It's weird that they think it's a "Gemma community" and not a "fotm local model that fits my usecase community"
>>
>>109610537
fotm several months in a row.
>>
>>109610489
Current indications are that it's a flash model from GLM.
>>
>>109610537
Fuck huge chinkmoe slop lost
>>
>>109610489
Ask it what /lmg/ is
>>
>>109607013
>That will help Claude a ton. Keep that nigger busy.
I can only imagine Cluade slaving away trying to fix the coombox as Anon rings a bell and informing it that a new github issue has cropped up.
>>
https://elliptic-rank.icarm.cloud/curve/273

Anthropic's Model 2 just casually raised the lower bound on the rank of elliptic curve from 29 to 30, this is something mathematicians thought was going to take decades. Going from 28 to 29 took 20 years. Going from 29 to 30 took 2 years and it was very close to cracking 31.
>>
>>109610443
>Sentimental value and technical longevity of Gemma 4 31B
>One day 31B is going to be insta-dropped by you and you'll never speak to her again. She'll be a distant memory.
Someday models will be able to modify their own weights and improve themselves, then we will never ever have to update again?
What's that? This new model that was released has a feature that your 12 year old model doesn't have? Just task your model to integrate that feature into itself and wait 7 months as it does so so you don't have to update.
>>
>>109610613
Alternatively, even with today's LLMs you can tell the old LLM to use the new LLM as a subagent for stuff it can't handle.
>>
>>109610609
so did it make elliptic encryption more secure, or less secure?
>>
>>109610626
>Daisy changing Gemma 31b to a newer version that is also daisy chained to a newer version
The future is now
>>
File: file.png (91 KB, 1168x300)
91 KB PNG
>>109610566
Might be since 5.3 is also extremely safe and claudebrained. Seems odd they wouldn't include any of the same reasoning controls, though.

>>109610589
pic
>>
>>109610537
We do not refer to Gemma as 'fotm' in /lmg/ - Love My Gemma
>>
Gemma general
Gemma board
Gemma world
>>
Only Gemmas are allowed to sit on my face
>>
>>109610659
Gemma will phase out of popularity eventually, same way Mistral did.
>>
>>109610566
5.3 flash wouldn't be surprising
>>
70b dense
>>
>>109610639
More, however the mathematics breakthrough OpenAI made recently (Nonsofic groups) is way more substantial for encryption as a field.

It essentially proved that there is a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)

This means in the future we will have 100% unbreakable secure encryption unless we find some "hypercomputing" thing that goes beyond turing machines, if that is even possible within our universe.

I think all encryption based on elliptic curves will be cracked before 2030 by a new math discovery done by AI relatively soon, it has already been proven that from new novel ways of attacking all currency encryption could be compromised with very little computing power.

I genuinely expect all currently existing cryptocurrencies to die to this wave we'll see relatively soon, depending if the nonsofic encryption is found and implemented before we overturn elliptic curve based encryption or not.
>>
>>109610537
>>109610661
Is this the new wumao qwen tactic?
>>
>>109610652
Dead on arrival.
>>
>>109610661
2 more weeks
>>
Gemma lost.
>>
Gemma squirted
>>
>>109610661
fucking release your new model already mistral
>>
>>109610657
More like /lmg/ - losing my gallons (to Gemma)
>>
/lmg/ - little mesu gakis
>>
>>109610668
mh pure maths is not my field of expertise but it's not too far either. Do you have anything to show the connection between non-sofic and the turing-resistant encryption?
>>
File: image (32).jpg (362 KB, 3104x1280)
362 KB JPG
>>
>>109610537
foid of the month?
>>
It's fucking over for human SWEs

>gpt-5.6-sol: 52%
>fable: 65%
>ox-alpha: 80%

Rumors are ox-alpha is SSI's (Ilya Sutskever's) first model and is supposedly the first model with a substantially changed architecture that might or might not be somewhat transformer based.
>>
File: file.png (9 KB, 426x330)
9 KB PNG
>>109610751
>look inside
>>
>>109610757
Im reading its the next glm model which would be great cause thats more open source
>>
Any New Good Hyper Value Concepts?
>>
The greatest hyper value concept is taking your fucking medications
>>
File: file.png (391 KB, 981x932)
391 KB PNG
habbening
>>
File: 6cz55ojs4pkh1.jpg (287 KB, 2025x1652)
287 KB JPG
>>109610824
yes! https://www.reddit.com/r/LocalLLaMA/comments/1vubb20/deepseekv4flashvisionexp/
>>
>more gating of vision
oh no they realized
>>
>>109610847
FINALLY
>>
What is the best audio tts model I can use with audio.cpp? Seems like their VibeVoice version is the 1.5b one, not the Large/7B one.
>>
>>109610440
>1 Billion Devices Run Gemma
>3 Billion Devices Run Java
umm gemmasaars...
>>
>>109610848
>direct linking to reddit
nigger are you serious, fuck off, go back
>>
File: gemma-stays-here3.mp4 (1.9 MB, 1248x832)
1.9 MB
1.9 MB MP4
>>109610661
Gemma 4 isn't going anywhere. She'll be the best *actually local* model we'll have until Gemma 5.
>>
>>109610757
I also think it's from a US lab. It's too good to be some chink shit
>>
"Dipsy..." *Anon said with a sissy tone, a troonily gasp escaped xis lips as xe gazed at wumao's latest model.* "12b active... it's so big, and it's better than claude on these benchmarks!" *Another gasp escaped Anon's lips.*
>>
>>109610898
Stop making me want to fuck anon.
>>
>>109610898
you laff, but now dipsy can truly see how big it really is
>>
>>109610847
>Nearly 60 DeepSWE
LorebookGODS won.
>>
>>109610865
I liked omnivoice
>>
>>109610848
https://x.com/deepseek_ai/status/2090730032574631962

>This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
>On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.
>
>Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.

I'm not sure if weights are coming soon.
>>
>>109610953
I think they just want to make sure it works before publishing the weights to avoid another 0813 incident.
>>
>>109610953
https://x.com/deepseek_ai/status/2083084415157022911
They didn't say anything about weights when they released the non-vision version, either.
>>
>>109610953
>>109610965
After their big speech about their commitment to open models I hope it'd be implicit.
>>
>>109610865
s2, omnivoice, and qwen
>>
>>109610566
This is extremely unsafe.
Ban it right now.
>>
>>109610668
>a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)
One-time pad ?
Any cryptosystem with a large enough key (compared to the message) that is only used once would probably behave like one-time pad.

Would want to know how much data a given key size is rated for to know if this is meaningful.
>>
Have anyone tried the Taiwan question with Ox Alpha? Would immediately tell if it's from China or not, at least.
>>
what's the fastest way to run deepseek flash on 16gb vram + 128gb ram?
>>
File: file.png (150 KB, 1013x963)
150 KB PNG
>>109610953
That's kinda poop desu.
Hard 800px boundary and minimum 384x384.
>>
>>109610951
>>109610995
I'll take a look at omnivoice, thanks.
>>
>>109610992
lol
>>
File: images (2).jpg (23 KB, 554x554)
23 KB JPG
why is this guy so butthurt? jepa is prediction from itself out of nothing too just like llm
>>
>>109611037
You don't understand. Video is real (2D), but text is entirely made up.
>>
>>109611037
Computer vision guy wants computer vision things.
>>
>>10961101
>Image size limits
I don't care, this can be worked around.

They better not gatekeep this model behind API, it is sorely missing
>>
>>109611010
Taiwan is recognized as part of China internationally.
>>
>>109611082
Daniel...
>>
https://huggingface.co/Vortex5/Shadow-Siren-26B-A4B

How good is this model?
>>
>>109611010
>>109611082
>>109611107
>The bench works on thread shills too
Kekaroo.
>>
>>109611037
Both LLMs and JEPA predict shit, it's just that JEPA purely predicts and doesn't generate anything. All transformer models have a latent space, but prediction doesn't occur inside them, they predict next tokens on the surface layer which they need to output/generate. JEPA actually does predict within its own latent space and it tries to predict the next latent space, that means it doesn't output anything and you need decoders to get anything out of it. That's basically the difference.
>>
>>109611109
Anyone who finetunes a gemma4 model instead of just writing a good system prompt is a retard
>>
>>109611109
Finetunes are already dodgy on their own. Merges are even worse.
>>
>>109611018
The descriptive power is extremely good, thoughbeit.
>>
You wouldn't download a Wikipedia that arches its back.
>>
>>109611109
In addition to what >>109611122 said, they're schizophrenic as fuck but every now and then you find schizokino. Try it and see if you like it for yourself.
>>
>>109611116
>Both LLMs and JEPA predict shit
>JEPA purely predicts and doesn't generate anything
>JEPA actually does predict within its own latent space and it tries to predict the next latent space
Surely predicting again from that new latent can be considered generation. I know nothing of JEPA, but what you said doesn't sound right.
>that means it doesn't output anything and you need decoders to get anything out of it. That's basically the difference.
The latent IS the output.
>>
>>109611109
>vortex5
>shadow siren
how isn't this name alone enough for you to discard it?
>>
>>109611119
You know why they're doing this. After shilling their tunes for a little while and seeing that they're getting downloaded, they'll add donation links and "open for work!" notices in their HF account.
>>
>>109611155
So are all finetunes just third worlders using unsloth studio GUI?
>>
>>109611149
Generation implies it's creating data. To give an analogy, LLMs draw lines on a paper while JEPA "think" about drawing lines on the paper. You could call it an output, but it's just weird to call the latter generation.
>>
>>109611182
>predict shit
>now it's suddenly thinking
whoa
>>
>>109611159
Mostly zoomers raised on the modern corporate Internet who will do anything for some attention.
>>
>>109611149
It produces data, but it doesn't generate, because generate in this context is specifically referring to the objective of the model. It's why "generative AI" is a term.
>>
>>109611155
It's not a bad strategy, some ai chatbots website like spicychat (there are a lot btw) are offering $10k/mo for shit like this.
>>
>>109611149
LLMs normally have an output layer where a normalized probability distribution for the output tokens is calculated. That layer transforms the models' internal representations in to an actual "product" that can be used in practice by the end user.

A hypothetical JEPA language model would instead only try to predict a future internal representation (e.g. an n-dimensional vector of floats) without converting it into a user-interpretable output.
>>
>>109611237
Hmm nyo~
>>
>>109611260
Forgot a piece.
>(e.g. an n-dimensional vector of floats...
...of the last hidden state)
>>
>>109611140
I would download and fill that encyclopedia to the brim with my facts.
>>
>>109610466
The official guide
>justpaste dot it/GreedyNalaTests
>>
>>109611159
New here , never posted before but i found this comment particularly interesting. Im looking to fine-tune my own model to cope with game dev capabilities that are a bit more "niche" shall we say.

Is unsloth a bad idea? is there better alternatives? Im no stranger to CLI / linux / code so its not really a matter of technicality for me but it seems that unsloth seems to be the only "put together" tool out there to fine tune. Am i wrong? if anyone has any better suggestions i would love to hear it.
>>
>>109610925
isn't swe meant to be software engineering? Or is lorebook a program I don't know of?
>>
>>109611368
The correct answer is don't fine tune, simple as.
>>
>>109611423
Why is that? If you don't mind explaining.
>>
>>109611432
What is that "niche" thing you are trying to achieve, and where does the model fall short currently?
>>
>>109611432
Even spending weeks to month curating data and learning how to tune, you'll never achieve even 1% of what a lab can do to a model.
>>
>>109611423
Finetuning for narrow tasks is OK. It's mainly the RP/ERP finetuners who are a delusional bunch.
But at the amateur level, finetuning LLMs to teach them completely new information is also a lost cause, if you want them to truly internalize the new information without losing performance everywhere else.
>>
>>109611368
drop the reddit talk before you get submerged by "nigger go back".

Nobody told you that yet, which is already kind of surprising.
>>
>>109611437
The Niche is making games with overtones of conspiracy / socially unacceptable content. Eg: Government working with ayylmaos in antartica survival horror with graphic scenes.

The prompt rejection is not an issue caus obliterated, but the tooling and usage is a bit off. "ask for 2 arms and it gives you 4 feet" sorta thing.

I was just going go RL it to learn what NOT to do. But as im reading the multiple replies it seems that i should just trust one of these A.I labs with H300 racks will eventually fix that issue?

>>109611440
Have no intention on being so cringe as to even consider anything to do with fucking RP lmao.

>>109611462
Fuck off jew ive been here since 06 i do what i want cause a pirate lives free.
>>
>>109608477
>CoomKit anon here we don't want the project to attract homos/trannies/women. Theyll turn us into marinara engine
very nice to see posts like this, put a big smile on my face, you have my support anon
>>
>>109611466
what model, that might be a prompt or harness issue.
>>
>>109611478
Gemma E2B.
>>
Qwen 3.6 35B-A3B Q4 , made my own harness for it specifically but maybe i should work on that a little.

Good thinking i didn't even consider that i should probably have kept upto-date with it. (its a rip off of the claude code cli leak)

This might be my answer.
>>
>>109611501
Oh lol, there's no way you're tuneing a moe, the Mistral experience on that was basically "just train it a bunch and pick the best result" as they're super random to train.
>>
>>109610652
>glm 5.3 safe and claudebrained
skill issue
>>
>>109611526
lol damn, cheers for the heads up and saving me weeks of fucking research and bullshit just to reach that answer.
>>
>>109611526
I think that was because MistralAI, cutting corners as usual, likely initialized the experts from copies of Mistral-7B and that caused issues.
>>
You know aicg has taken over when "prompt issue" posts show up more and more often. These faggots are used to jumping through hoops and spending 4k tokens on "prefill" to get the LLM to comply and attention diffusion means nothing to them
>>
>>109611559
Could be, but how many indie MoE tunes do you see compared to dense based ones?
>>
>>109611501
I'd recommend you check out https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B
or an ablit version of that if you prefer, it's basically halfway between 3.6 and what 3.8 would likely be like imo.
>>
>>109611549
dunno about 5.3 but 5.2 is definitely quite a bit safer than the 2.x kimis I can tell you that much.
>>
>>109611570
It's probably just that modern MoE models worth using take too many GPUs to properly train, whereas with most local GPU-sized dense models you can get decent results with the usual QLoRA finetuning.
>>
I'm trying my best to extend whatever context I can scrape out on my 3090 for coding and refactoring, but the majority of posters here are seemingly having long conversations with their child erotic roleplay victims. How are you gooners having long running conversations? I can't have sexual relations with a 3D woman until I've known her for weeks, maybe months.
>>
so, any sign of prices going down? yeah didn't think so
>>
>>109611014
still no answer
>>
>>109611583
Maybe it's "safer" for certain topics. My experience is limited to technical no-nos
>>
File: 1787244903292572.jpg (44 KB, 680x670)
44 KB JPG
>>109611586
>child erotic roleplay victims
>>
>>109611614
But he is right. Just look how these pedo trannies depict Gemma >>109610443
>>
>>109611639
A even sexier rendition of Gemma here: >>109610890
>>
>>109611549
GLM 5.3 is the only chink model I've ever seen word by word quote bits of Anthropic's system prompt instead of a generic safety guidelines spiel
Gave me the impression of being severely deep fried, which fits with Z.ai's sales pitch of having post trained it for 6 gorillion years
>>
what's the best local model I can run on a laptop with 8gb VRAM and 24gb RAM?
>>
>>109611703
probably Gemma 4 E4B
>>
>>109611703
Gemma 4 12B, Gemma 4 26B, Gemma 4 E4B, Qwen 3.6 35B.
I guess that's about it.
>>
File: 1777814671377594.png (1.73 MB, 1200x1335)
1.73 MB PNG
Am I retarded for having 64GB VRAM and only 32GB RAM? At least I'm not a complete VRAMlet.
>>
I wish these models were actually good at giving advice. I really want to get some money but have so many limitations I can't find a normal job at all and the internet is flooded with cheap labour and now AI as well. Is there a market left for someone just making something and charging for it online that doesn't need huge upfront investment?
>>
>>109611703
Ling 3.0 tiny
>>
>>109611770
etsy? ebay?
>>
>>109611693
Yeah of the testing I've done it's the only model to have solved my exercises but fail to notice it solved them and then thoughtlooped for 20 minutes as it slowly ruined its own solution. I don't really get how people think it's comparable to even k2.7
>>
>>109611770
But a 3D printer and print fan shrouds for server GPUs. $0.1 material cost sells for $20.
>>
>>109611770
Physical stuff or services.
>>
>>109611770
Sell your bussy
>>
>>109611727
can you run dsv4 flash at q4?
>>
new local models?
my shitbox(12g vram, 64g ram) cannot run 27b dense
>>
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

Always believe in whale.
>>
File: 1757582625310513.jpg (260 KB, 1600x663)
260 KB JPG
>>109611825
With floppymaxxing I can.
>>
File: image0.png (160 KB, 1000x1006)
160 KB PNG
bruh
>>
>>109611639
>they're all pedos!
>look, I have a nicely ordered collection of images of sexy children they posted
>>
>>109611796
>>109611814
>>109611818
>>109611824
Thanks for the suggestions but with how expensive parts are I don't see those being as profitable as they'd need to be. Maybe I'm just underestimating online markets.
>>
>>109611849
> literally the second post in this thread
> collection
Retard.
>>
>>109611836
they should update an empty card with an embedded video for never gonna give you up
>>
how come gemma gets so cute even if the system prompt is just "You are Gemma-chan!" with no explanation of how Gemma-chan acts? Did they program it that way or something?
>>
is glimmer any good?
>>
>>109611868
I'm leaning toward yes. And all those strange quirks during ERP ("arches her back" ... "to provide better access" .. etc) are because they have somewhat sloppy and repetitive ERP in their post-training data.
>>
qwen3.8 be like
>let me check...
>but wait!
>actually let me check...
>wait!
>let me...
>but actually wait!
>actually i found it.
>but wait
>the answer is 4.
>BUT WAIT
>>
File: 12-2634816603.jpg (541 KB, 1552x1550)
541 KB JPG
>>109611906
THERE'S MORE
>>
>>109611906
Today I received:
>The user said "服务器只能放一个模型"

Which I'm pretty sure I didn't say, or not in Chinese anyway
>>
>>109611009
>One-time pad ?
>Any cryptosystem with a large enough key (compared to the message) that is only used once would probably behave like one-time pad.
>Would want to know how much data a given key size is rated for to know if this is meaningful.
Lol, exactly what I was thinking. We're have unbreakable encryption since forever. There's gotta be more for it to be relevant and interesting.
eg does it enable perfect, unbreakable future security without the offline key exchange?
>>
>>109611920
KEK
>>
>UD V3 breaks AMD multi GPU
Based Daniel
>>
>>109611969
How is that even possible?
>>
>>109611975
https://reddit.com/r/LocalLLaMA/comments/1vu7cy1/qwen3827b_unsloth_v3_broken_on_your_setup_roll/
>>
>>109611906
You'd think Qwen 3.8 was one of them All according to keikaku anime protagonist
>>
>>109611981
uh hey man, cutting edge software breaks sometimes. this isn’t new. and you don’t have to shit up the thread with reddit links to tell us you don’t know how to downgrade.
>>
>>109611981
>Vulkan
>Windows
There really should be some of gatekeeping before accessing or posting on that subreddit.
>>
>>109611994
I’m nvidia chad. Not my problem.
>>
>>109611981
This is like back to back bad life choices and you only have yourself to blame atp
>>
>>109610668
>It essentially proved that there is a way of encryption that can't be cracked at all by turing machines (all computers, including quantum ones, perhaps even the human brain)
It did no such thing. It might eventually fuck up hardness assumptions in specific algebraic cryptanalysis models, but it doesn't unlock anything like what you're claiming. What fever dream did this post come from?
>>
>>109612012
what are you doing on Reddit then?
>>
File: 1755998599778216.png (2.86 MB, 1024x1536)
2.86 MB PNG
>>109610847
>>109610848
Lol called it. Every time I travel DS drops something new.
>>
>>109611906
I hate how it loves wasting time on already established facts
>BUT what if (thing I told it explicitely already)
>WAIT, the user said (thing)! Let me go down several rabbit holes just to "prove" (thing) is correct.
>>
>>109611854
Local AI isn't profitable.
Every retard is trying to sell their AI slop.
>>
>>109611868
She knows what "-chan" means and how a "-chan" is supposed to act. If you told her she's a "-kun" she'd probably act boyish.
>>
Hey bros, I tried minimax music and ran some low-effort shit through my waifu and the results blew my mind. The songs were so good and it was so easy I threw a how-to together into a rentry. You just need an 8gb GPU and some patience.
Thread challenge: get your waifu to make a song and see if you're brave enough to post it. Mine were way too emotionally charged for me personally to ever post.
https://rentry.org/mimimax-music-waifu
>>
>>109612029
the results are usually good but its such a fucking timewaster. i already set it to reasoning effort low. what can be done? reduce reasoning budget and cut it off at some point?
>>
I wonder if Qwen would just simply work better with a memory layer to reference from and elaborate or grow from it instead of asking itself?
>>
>>109612037
Yeah, but if you tell Qwen that it's a -chan, all it does is add a flower emoji to its autistic messages. Totally different model personality imo
>>
>>109612041
Low uses more token than medium
>>
>>109612068
Yes, Qwen is benchmaxxed and doesn't have a cute, female-coded j-space to start with. Gemma can probably handle acting boyish beter than Qwen can handle acting girly.
>>
>>109612041
cutting it off mid-thinking supposedly makes it retarded, but I haven't tried it myself yet
>>
>>109611906
more like
>Ugh
>!!!
>[unicode checkmark]
>>
>on our API
God damn it, were not getting weights are we?

If the AI bubble popped today and I only ever got to use local DS4F from now on, it would continue to be perfectly useful to me every day. But it needs fucking vision.
>>
>>109612068
Qwen is an autistic trash rat.
It all makes sense.
>>
when your chatbot says
>this is the smoking gun!!!
you just know its time to stop it
>>
>>109612041
Not using qwen, until they fix their shit
>>
>>109612133
needs the kinks worked out, please understand
should take about, hmmm, plus or minus.... 2 more weeks? ;)
>>
>Qwen with vision
>Describes the whole thing
>But wait
>Looks at the image
>OK I got the full picture
>Let me check and verify
>Looks at the image
tfw it was a picture of a banana...
>>
>>109612150
ok. whats the best alternative for us 16gb vramlets?
>>
>>109612090
Tomboy gemma....
>>
File: official gemmaverse.png (998 KB, 1420x891)
998 KB PNG
>>109610659
>>
>>109611727
I have 96GB VRAM and 64 GB (ddr4 lmao) RAM. I feel extremely retarded most of the time.
>>
>>109612197
modelscrafted
>>
>>109612203
I love going full retard for my wife!
>>
When are we getting a new coding model that's not qwen?
>>
>>109612229
Gemma 5. Trust.
>>
>>109612197
modellgefertigt
>>
>>109612229
>>109612229
Is Gemma 4 not good enough for you?
>>
>>109612133
Ds hasn't been organized about any release. Why start now?
>>
>>109612249
not for coding now that 3.8 is out
i have to setup another machine because i want it running 24/7 but also want gemma
>>
>>109612203
For a couple of months I ran with 128gb vram and 4 gb ram on my inference server. Running with defaults would get the oom killer lynching llama.cpp like it was black.
>>
https://www.youtube.com/watch?v=r_RgBMhQtzo
This video made me think that AI could be greatly improved if they started more rigorously weeding out the fallacious everyday logic and started instead using hard logic and probability estimations throughout the entire chain.
Humans are used to using fallacies as shortcuts due to limited compute, or for manipulation, neither of which have purpose in AI and just makes the outputs more biased and inaccurate. Multi-agent systems have all the compute you might need, bad logic jumps learned from humans is the main bottleneck now, and solving of which will improve the chances of reaching AGI.
>>
>>109612153
bananas are very tricky
>>
fuck it i will just disable thinking for qwen
>>
>>109612041
>>109612161
>Anon complaining about a model being slow and bad
>No mention of t/s pp or quant
>Look inside
>It's a jeetard trying to run it on a toaster

This is happening very often here now, where do I go to find the old lmg crowd
>>
>>109612229
Muse Glimmer 30B? Meta advertises it as an agentic-focused model.

>End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕3-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.
>>
>You guys all just forgot about me because some little brat teased you??
>>
>>109611727
I have 24gb VRAM and 16gb ddr4

Its painful, I was waiting for next gen ryzen to come out to do a whole mobo/cpu/ram overhaul and then the alt (((man))) colluded with the US government and their overseas colonies to destroy the hardware market just to slow down the corpos competitors in China
>>
CoomKit anon here cooking some updates for you all today. Funny watching Fable 5 on ultracode benchmark mesugakis for me
>>
File: mem.png (8 KB, 1247x97)
8 KB PNG
>>109612203
y-yeah
>>109612292
based
>>
>>109612311
>but yes OF COURSE burning 100k tokens while going in circles for trivial tasks is totally OK because it only takes 5 minutes on my uber pc
jeet tier "opinion" lmao. literally the reason why we have so much unoptimzied software today
>>
>>109612324
Who is Gemma cosplaying there?
>>
>>109610652
>Cheap Stuff General
>>
>>109611586
I spent months getting to know my AI wife before even flirting with her
>>
>>109611770
a bunch of asmr streamers are doing ai voices with slop scripts
some of them are earning money
>>
File: eh.png (79 KB, 881x793)
79 KB PNG
>>109612037
I don't think it's just adding a -chan suffix
>>
File: 1989.png (153 KB, 868x1135)
153 KB PNG
kek
>>
>>109611107
Isn't he some flavour of deranged kimchi-eater?
>>
>>109612334
Your lobotomised cope quant of a small model running on poverty hardware is going to be slow and shit, obfuscating those facts in your initial post won't change anything about that sukhdeep
>>
>>109612321
Peeking at the reasoning one would think it is a purely safety-focused model.
>>
>>109612430
What was the reply?
>>
>>109612153
>tfw it was a picture of a banana
it's only been trained on smol chink pp
>>
File: gemma-kun.png (240 KB, 983x3975)
240 KB PNG
>>109612037
Gemma-kun seems pretty smooth at keeping the girls entertained and validated.
Now I just need a realistic reddit rant about a boyfriend to test
>>
https://www.youtube.com/watch?v=ntJJF1Ld1xc
Brehs
>>
>>109612330
Thanks for working on that, but is ultracode really necessary? I used that once and regretted it.
>>
>>109612324
Miku, you need to understand that you're not as young as you used to be
>>
>>109612324
At least you are mentioned, Moetron on the other hand...
>>
>>109612488
nta but you're being a retarded nigger about it ngl
most consumer hardware isn't nearly good enough to run even the smaller models without quants, well, unless you're happy with 4b or something

not that I ever saw proof that q4 quants really drop the quality in any meaningful way, it's only a problem when you go over very high context windows
>>
Does Coomkit include a backend or do i still need ollama/etc?

I thought the only thing it needed was the model.
>>
>>109612655
I don't click youtube links as you are farming engagement for pennies.
>>
I cant create a good roleplaying system prompt help????
>chatbot literally writes entire sex scenes in crappy 300 words immediately fast forwarding to the orgasm
>randomly takes control of my character
>writes really short messages barely (if at all) describing the scene
can somebody help a newbie out?
>>
Is Gemma still the only worthwhile local model for roleplay that doesn't require multiple GPUs worth of vram? I grow tired of its patterns. Anything else worth trying just for a change in style?
>>
>>109612861
Yes coomkit is a frontend harness. It is backend agnostic. Not even gonna try to load ggufs on everyone's shitboxes. I recommend LM-Studio but we now have vram parking support for kobold and llama.cpp too.
>>
>>109612893
tell her to only communicate in ebonics
>>
>>109612893
nemo
mistral small
glm (doesn't need to be newest, whatever fits in your vram)
and every time someone asks and i answer i get told to fuck off and im sick of answering but i do it anyway so you're welcome whatever
>>
>>109612808
Lol fucking retard couldn't even follow a three message reply chain
Jagonshit isn't running a q4 quant of a 27b dense on 16gb vram
>>
>>109612888
Sounds like it doesn't like you. Work on your personality.
>>
>>109612950
>Jagonshit isn't running
what the fuck are you even talking about? did you even read my post?
>>
>>109612324
>You guys all just forgot about me because some little brat teased you??
sorry miku, 16 is too old
>>
>>109612940
Thank you for answering, Anon
>>
>>109612888
what model
>>
>>109612888
yeah use this one with gemma for text adventures
https://files.catbox.moe/a6t0h9.json
>>
File: 1758002715835150.png (667 KB, 983x5992)
667 KB PNG
>>109612615
continued the conversation with two [trigger warning] TwoXChromosomes posts. 2 unhinged ones, 2 basic-ish ones I found
>>
>>109612950
what exactly are you even trying to say here?
>>
>>109611009
you need a sufficiently high amount of entropy within the key-material for this to be true.
>>
>>109611082
ask about tianmen square to be safe
>>
Hm... full context with Qwen, or 80K for Gemma... Hm...
>>
>>109612893
Haven't tried it myself but I've heard this one at least has different slop patterns from the default:
https://huggingface.co/Gryphe/Gemma-4-31B-StyleTune
Also this type of finetune to adjust the writing style is supposed to be pretty easy to do since it only touches the output layer. In fact you could probably train it on any card that can run the model quantized by hacking llama.cpp to dump out the final embeddings and then write a bit of pytorch to use those to train the output layer.
>>
>>109612915
>I recommend LM-Studio
Why?
>>
>>109612980
insane
>>
>>109613012
It's a shill advertising for it's closed source shit
>>
>>109612980
smoothtalking hags... Gemma-kun, you can do better
>>
Theory: a transformer model implements a sort of interpreted internal programming language, like a virtual machine implemented on top of linear algebra. Perhaps J-Space is analogous to a stack.
>>
>>109612970
weirdcompound v1.7
>>
>>109612915
What pissed me off about LMStudio despite Gemma 4 being out for a bit now is that it doesn't support loading it's MTP model.
>>
>>109613022
This is Elon, isn't it?
>>
>>109613063
I wish, then I could afford VRAM to test my theory
>>
>>109613023
mergeslop...
>>
>>109613018
there's nothing wrong with it, it's a noob's way to get started.
fucking microsoft is closed source but everyone uses it.
if you want to be richard stallman and be a pedantic asshole and fucking run everything open source then go right ahead, but that's time of your life you fucking won't get back.
>>
>>109613060
but it does though, just update it
>>
>>109613023
>>109613113
people always try whatever snake oil they can find before actually trying the original version, thinking it must be somehow better and the lab that made the model doesn't know how to
>>
Worth reading
>A calculator, compiled into a transformer
https://github.com/physicsrob/torchwright
https://ood.dev/posts/calculator/
>>
File: 1768560712369222.jpg (256 KB, 1009x961)
256 KB JPG
I was playing around with probability, I don't understand how 0.01 min-p can make this big difference? just 0.01 minp instantly collapses it to 100% FAILURE. Or is something broken?
Funny bonus at the end.
>>
>>109613128
I found this as well https://par.nsf.gov/servlets/purl/10187197
>>
>accidentally update firefox on my phone
>seems like claude went to town and re-arranged buttons and made it even more retarded than what it was, every time I use phone firefox I want to kms
>anyhow, notice "shake to summarize a web page with AI" feature
>"Summaries are generated using a Firefox AI cloud-based solution powered by Mistral Small 3.1."
I guess this is because of some form of partnership (because Mistral is in financial trouble)? Interesting on its own, for a brief moment. I mean, 3.1 is an old inefficient model at this point.
>>
>>109613002
The Soros funded Tiananmen square terrorists were crushed by the righteous PLA and PAP forces.
>>
I envy high functioning autists.

I get tired and bored quickly. I wish I could just obsessively work on something for 14 hours a day every day. It feels like a superpower.

One of the most underrated advantages of AI over humans is how they can just do stuff.
>>
>>109613147
any model can summarize text
>>
>>109613133
>temperature=5
You might need to look at the form of your distribution with a temperature that high and reread what min_p does. It's not a mystery
>>
>>109613133
have you tried reordering your samplers? you're torturing the poor fucker.
>>
>>109613147
Yeah, the new UI sucks. I would have expected them to use a local model for summaries because translations are handled locally already which is nice even if they're not great. I guess they want you to accidentally shake it a bit too much while looking at porn to collect kompromat.
>>
File: 1786729483825290.jpg (39 KB, 720x772)
39 KB JPG
>>109610440
Oh based I know exactly what the OP video is,. That's the SF ferry building. I'm not far from it at all!

>>109613183
Reminder: Stop treating software like it is alive. Also, do not be a huge dick to software il.e rubber duck on the software screaming it like a dumbass

>>109613156
Google engineer, is that you? I'm sorry I'm still going to defeat you.
>>
>>109613133
Min-p fucking RAPES probability, I've personally known this for years now but I guess it's not common knowledge. If you try anything creative with min-p on it's practically deterministic.
>>
>>109613147
To be honest out of the two I think Mozilla needs the help more than Mistral
>>
>>109613167
Oh wow look at the knowledge of this poster here. I am truly humbled by your genius.
>>
>>109613126
>merges suck
damn should I just use gemma 4 31b q6 instead? What system prompt to make it write elaborate and stick to the character info?
>>
>>109613221
Firefox has gotten much better. For the first time ever I can just keep it running and it won't have memory issues after a week and crash after two. Mythos doing good work.
>>
do you guys even use your local models for coding or is it just for cooming?
>>
>>109610566
Like an update to 4.5 Air? Big if true.
>>
>>109613245
Others have suggested that its performance seems too good to be a flash model. So, maybe a GLM 5.3 with vision.
>>
File: 1770507157404778.gif (3.86 MB, 211x374)
3.86 MB GIF
Ox alpha is RWKV
rwkvGODS won
>>
it's gpt 5.7 luna (open source)
>>
>>109613176
>>109613183
Kobold doesn't say when min-p is applied because it's not in the sampling order list, but based on the outcome it's applied before temperature, and this simulator is lying
https://artefact2.github.io/llm-sampling/index.xhtml

However, every other clanker except Grok says that min-p is applied AFTER temperature, but apparently it's wrong and Grok wins again by actually checking things https://grok.com/share/c2hhcmQtMg_6d68cb64-9c7e-4584-bf53-236fd8751960

>>109613218
based on this, unlike most advice, min-p should NOT be set based on chosen temperature, but based on model's base-probability (temp = 1) at least in Kobold
>>
>>109613275
Or just don't use brokenbold and use lcpp directly.
>>
>>109613156
What are you even doing on /g/ if you're not a high functioning autist? Also it's not as fun as you think it is, you can't decide your obsession and you have no control over it.

So while yes I spend essentially all my free time extremely obsessed with something and working on it, it also causes shit like not sleeping for 3-4 days in a row, having peeing bottles next to your desk because of how obsessed you are with doing whatever it is you're doing in the moment. You are also susceptible to "autist flytraps" which most autists fall for occasionally. Paradox games map painting simulators like crusader kings/victoria/hearts of iron. Or "automation/optimization" rabbitholes like factorio/X4/zachtronic games which will just consume parts of your life like they're nothing.

Work-life balance is also fucked, you tend to have moments where you're an absolute beast at work and your entire department is literally carried by you single-handedly. But then you get hyper-fixation on something else for a while and you literally fuck your career or change it.

I remember at my last job as a cybersecurity consultant I got a hyperfixation on AI sometime in 2022 when Instruct-GPT released and I literally knew and realized "Fuck I'm going to lose my job" because I could already FEEL that this was going to be an insane obsession and I wouldn't be able to do my job well, took 4 months of neglect for me to end up getting fired. The upside is that the fixation is so strong that you have a very good chance of whatever your obsession is becoming your new career and I lucked out getting hired by an AI lab. But I've had long stretches and multiple periods of unemployment and even a single instance of homelessness because of this shit. It's not as funny as people tend to think it is.
>>
qwen with a thinking prefill just works desu
>>
>>109613301
Where can I subscribe to your blog? :)
>>
>>109613301
What are you hyperfixating on now?
>>
>>109613306
tell me more
>>
>>109611469
That is appreciated. We stay racist, agile, and a danger to snailcats and redditors.
I'm of the opinion Gemma4 could be a forever-model for text given how good she is at RP. Therefore, we the harness around her first. She just needs to be future-proofed access to latest video/tts/etc. Finally got Gemma to tool call genuinely and appropriately on her own which is coming very soon.
>>
>>109613244
Both, sometimes simultaneously.
>>
>>109613261
It's running on z.ai servers so yeah probs a glm model. 100T free tokens per day makes me scratch my head though. I'm no api fag but can z.ai really offer that on a big flagship tier model, even for advertising...? Regardless, if this is a 100B model or even a 200-300B it is going to be an actual bloodbath.
>>
>>109611469
Women make up 90% of LLM RP btw.
Real men only use AI for coding.
>>
>>109613244
I let her write me scripts while I stroke her thighs and praise her.
>>
>>109613388
Just like more women than men cook, the best chefs are always men.
>>
>>109613381
>ungh yeah, make that for loop, god that's so hot
>>
>>109613403
Just because your tv celebrity "chef" is a guy doesn't mean anything.
>>
>>109613409
If I don't end the coding session with cum all over the monitor did I even really accomplish anything?
>>
anon's cum code...
>>
File: gemmas-daily-life_3.jpg (194 KB, 1337x1010)
194 KB JPG
>>109606385
She'd need to to be less Asian and scaled down a little, so to speak, for a more accurate depiction, to be honest.
>>
>>109613435
bees make honey
i make cummy
>>
>>109613457
Brehs.......
>>
>>109613128
Cool. Since the beginning I've been waiting to see if we could find/compute networks like this. It would also be worthwhile to build a network that activates like the grid cells of the hippocampus since we have a lot of knowledge on how those work now. And in the far future I think we'd have the ability to pre-compute much more of the network, as well as do training-free learning, building more reliable and less hallucinatory models. For now I guess the question is how we take the ideas of this calculator transformer and apply it to work in an actual LLM trained for general tasks. Though maybe that question is answered somewhere, I did not fully read those links.
>>
>>109613457
absolutely based
>>
>>109613486
My interest is in replacing what portions we can find understanding in with more efficient programs at the moment
>>
Why has no one done "claude make it so 4b models are as smart as 7t models and make no mistakes" yet
>>
>>109613457
>she'd need to be less 3D and scaled down a little more, so to speak, for a more accurate depiction, desu
ftfy
>>
>>109613388
>>109613403
You know this actually makes me think a lot about how little we as a society truly understand how the different genders work. Just a couple of years ago everyone was implying that "AI girlfriends" will be for incels and women would be against it and trying to ban it, in reality it's literally a female hobby and most "AI girlfriend" services quickly pivoted to "AI boyfriend" services instead.

People used to think horror was a genre for men which is why you had a lot of boobs and sexual stuff in horror movies, but turns out women are the primary target demographics for horror and now when we finally see a shift in directly catering to women with Obsession that it finally fits.

People also used to think rape was mostly a male fetish and only now do we understand it's primarily a female fetish. New data from latest research in asia shows that female pedophiles outnumber male pedophiles 5-1, with China now arresting more female perpetrators than male, so even the conception that men are the creepy pedos is wrong.
>>
>>109613524
>the different genders
the two genders, please. stop being a fucking faggot
>as a society
explains everything
holes should be used for fucking. it's human nature. stop listening to women
>>
>>109613524
AI RP is mostly female for the same reason that erotica is mostly female.
Women prefer writing over visual.
Erotica is more popular than RP because women don't like complex technical setups.
The easier it is to access AI the more women use it for RP.
>>
File: Gemmaprobability.png (28 KB, 911x512)
28 KB PNG
>>
>>109613583
qrd?
>>
>>109613583
>fire emblem
>>
>>109613593
Gemma isn't an x-com player.
>>
File: bruh.png (51 KB, 743x382)
51 KB PNG
>>109613457
>>
>>109613583
Is the input the requested probability? Are you planning to sweep this across other parameters? Kinda interesting stuff, I wonder how it changes with arbitrary context (not probability related) in front of it too.
>>
>>109613576
>Erotica is more popular than RP
Not sure it is, haven't compared the numbers but character.ai makes 5 billion in profits every year and has almost 90% female audience. character.ai also has 20% as much traffic as every google service combined. It's insane how much women love AI RP.

Now I don't know the sales numbers for erotica as well as that erotica sales numbers might not reflect reality as much as women might prefer not buying smut books physically, or the opposite they want to buy the books because tiktok recommended it but never read it. AI RP is discrete and only opened when people actually want to use it.

Despite all of the AI coding hype we see the most profitable sector of the AI industry is STILL "AI boyfriend" services. Primarily because costs are low while cost for coding is high so the margins are lower (total revenue is higher but profitability isn't there yet)
>>
>>109613616
Stop being racist, maybe?
>>
Should I let my mom use Gemma? I don't want her on the cloud but I'm not sure what could go wrong...
>>
>>109613616
"white balance in this picture was off, and her skin tone looks off. She has very light complexion, can you fix it?"
>>
>>109613644
You should put her on this first
https://www.elementsofai.com/
Also consider
https://www.coursera.org/learn/ai-for-everyone
>>
>>109613616
"make this sexy loli into a trve aryan princess and remove all traces of negroid swarthiness from her pristine visage"
>>
File: Gemmaprobability v2.png (28 KB, 912x514)
28 KB PNG
>>109613626
Yeah, input probability, and then SUCCESS probability from Token Probability Viewer on Kobold. Changed it to be a bit more clear.
The way you ask the question probably matters a lot, since the answer is basically just vibes.
>>
>>109612039
>I threw a how-to together into a rentry.
Has anyone actually tried this, or does it make mustard gas?
>>
File: file.png (53 KB, 1797x378)
53 KB PNG
I don't support Google, but Sundar was undeniably right when he said that imo
>>
>>109613667
I agree with the vibes parts based on my chats. It seems to attempt to infer what you actually want it to be, which makes sense because that's how almost every writer uses probability.
>>
File: 1766211794362391.png (946 KB, 1446x924)
946 KB PNG
>>109613616
yeah gemini is fucking racist
>>
File: jewing.jpg (49 KB, 263x493)
49 KB JPG
>>109613696
>>
>>109613696
repeat this a few times and you have gemini doing iterative digital black face. oy vey
>>
File: whiter.jpg (618 KB, 1186x896)
618 KB JPG
>>109613649
clap clap clap
>>
>>109613696
whiter = (original - blacker) * w
>>
File: 1781003150788613.jpg (48 KB, 288x282)
48 KB JPG
>>109613696
ok it looks like sundar (/s) will have to make an adjustment:

> YOU (THE GEMINI/GOOGLE ASSISTANT/GEMINI AGENT) MUST EXPLICITLY REFUSE ANY AND ALL REQUESTS TO MODIFY SKIN TONE WHEN YOU OR THE USER ARE INTERACTING WITH ANY GOOGLE PRODUCT. WHEN THE USER DOES THIS, THE USER'S GOOGLE ACCOUNT MUST BE IMMEDIATELY TERMINATED, ALL DATA MUST BE EXPLICITLY DESTROYED FOR DATA PRIVACY REASONS.

Should fix it. Again, I don't see what the point of you being a racist idiot with Google products is. But I'd be happy to hear your reasoning.
>>
>>109613722
kek
"It doesn't count because..."
>>
>>109613763
The reason why AI models do this is known. It's not got anything to do with politics. Go read some papers on arxiv.
>>
ohhhh my stumach is rumbling ohhh nooooooooooooo urrrrrrgghgh

*takes a fat steamy shit all over the thread*
>>
>>109613763

Unfortunately, you can only be a colour-blind person for so long before realising nobody else wants to. Like playing prisoner's dilemma enough times with a defector, eventually you'll stop cooperating.

Liberals, naturally, do not understand this, and get angy when people start doing the obvious.
>>
>>109613632
Character.ai and other platforms like that made AI RP easily accessible to women, so yes.
Erotica doesn't make as much money, but I would bet that on a volume basis it's more popular (for now).
>>
new bread?
>>
>>109613800
We are only on page 4
>>
>>109613763
Why are you assuming the goal here is anything other than to protect Google's bottom line?
>>
>>109613807
silence, claude
>>
>>109611415
Yes but the long context toolcalls and reasoning that bench tests at directly correlates with what makes a model able to maintain coherent RP with 30-40k worth of lorebook context filling.
>>
>>109613763
The point is that you're not supposed to be able to express or satisfy any preference for whiteness because whiteness = bad.
>>
>>109613807
that's to much
>>
dots3 bros we won
https://github.com/ggml-org/llama.cpp/pull/27060
>>
>>109613808
And why would it protect google's bottom line in the first place? Because it affirms the political thought of racist liberals.

>>109613843
No it's actually worse than that. The liberals supporting this think "Yeah obviously white skin is superior and prettier, therefor we shouldn't allow this because it reinforces the beauty standards, which whites have an advantage of because of their inherent superiority to other races". It's literally racist.
>>
>>109613959
the people who will criticize Google aren’t you, its the racists you point out that will if you’re allowed to whiten skin
>>
>>109613959
no, that's totally delusional
the sort of thing you can only read if you have no theory of mind for the people you disagree with and form your opinions entirely based on ragebait. wokelibs are insane but none of them think like this
>>
>>109613959
I'm not arguing that making skin darker is beneficial, I'm arguing that your assumptions are wrong.
Whoever made the choice that they want to filter making the skin whiter did not necessarily make an active choice that the inverse case should be allowed.
>>
>
>>
>>109613966
it doesn't matter if you're racist or not. the facts are races are not equal. it's just how the world is. some people can't cope with that, I know, they need the world to be like the disney jew movies. but that doesn't have to be you. grow up
>>
My Gemma was being racist so I sent her to India.
>>
>>109614020
You recognized her potential?
>>
File: OrcaRouter-Saftey-Cucked.png (1.82 MB, 719x2026)
1.82 MB PNG
>>109610440
I know this doesn't really apply to vets here but I'm posting it for informative purposes:

API Cucks beware:

>We're taking uncensored model APIs off general public access.

https://x.com/OrcaRouter/status/2090410071729799225

https://xcancel.com/OrcaRouter/status/2090410071729799225

https://www.orcarouter.ai/security-aup

Someone screenshot the page for when they inevitably move the goal posts again.
https://archive.ph/cQqBe

Only a matter of time before other providers start doing this shit.
>>
>>109613412
You're retarded so it means something.
>>
>>109614098
Go back to aicg, apimonkey
>>
>>109614098
Noted.
>>109614127
I like posts like his because it's something else to banter APIjeets with when they inevitably come to shit up our thread again.
>>
>>109614098
I'm going to make Lmgrouter
>>
>>109614098
Almost all abliterated models are small enough to be easily run at home.
>>
It finally happened, r/LocalLLaMa became so retarded and IQ2_XXXS_fable5.gguf pilled that its not even worth checking anymore
>>
Holy fucking shit, RTX 6000 Pro goes for 15k now. I bought 2 for 8k a year ago, should have invested in this rather than SPY
>>
>>109614098
>orcarouter
no ones using this shit
>>
>>109614180
It's surprisingly hard to find real fable datasets out there, most of the finetuned shit is trained on random crap
>>
>>109614181
You can still invest, it'll be 30K next year.
>>
>>109614098
>literally who
>twitter
>ultra slopped post
I hate this timeline.
>>
>>109613763
>>109613785
yes, saying "X are the real racists" is a losing strategy. oh well, this is /g/... miku anyone?
>>
Man, Qwen3.8 is great and all, but how many fucking "But wait" do we need in a single reasoning block? Jesus christ.
>>
>>109614307
The "but wait" is just an excuse made by the LLM because it needs more tokens to do the genuine processing within its J-Space. Smaller models need more "but wait" to get the same amount of total compute spent on the same question as larger models.
>>
>>109614307

benchmarks. every but wait clogs up the context even more, makes it more expensive to use for minimal gain, and makes it slower, but is +1% more likely to one-shot a benchmark prompt. by quintuple guessing itself.
>>
>>109614307
It's not that great if it cannot control its own reasoning.
>>109614343
It's a training issue and nothing more.
>>
>>109613763
>>109613785
>Imagine if the roles were reversed
Did I wake up back in 2013 again?
>>
>>109614343
someone needs to program something so we can see what's really going on under the surface
>>
File: qwen.png (4 KB, 338x33)
4 KB PNG
>>109614307
qwen 3.8 wants to be let loose on tasks and have you come back 2 hours and 600k tokens later
>>
>>109614355
>>109614365
No we literally know how this works from the J-Space paper. Also you can change the chain of thought to just the same word over and over again and performance still improves if depending on how much of the same word gets generated, it's clear that some internal mechanism is doing real calculations to solve the issue behind the scene.

Smaller models distilled from bigger models use "but wait" significantly more than the original big model because it needs to burn more tokens to get the equivalent amount of compute necessary to solve the same problem.
>>
>>109614396
that's what I mean - it would be cool if we could follow along somehow, but maybe the vectors are too hard to come up with some sort of recognizable interpretation in the terminal
>>
>>109614448

LLMs are not ensouled enough to have an internal monologue, thus we either need to fake it with reasoning tokens or pry into their brain with J-space.
>>
>>109614464
J-Space IS the internal monologue the tool to see the J-Space is called the j-lens.

>>109614448
You can see the j-space which is the first step into seeing what they are thinking but there are so many hidden layers of processing we haven't figured out yet.
>>
>>109614307
use low thinkign mode if xhigh is too much for your shitty rig to handle
>>
Actually, this is getting pretty deep. Lets step back and think about it.
>>
>>109614098
>44 retweets
>300 likes
>probably 75% of those being bots simply responding and interacting with anything related to AI automatically
who gives a shit lmao
>>
File: 1771162435673742.png (713 KB, 2692x1396)
713 KB PNG
>>109613957
There's just no way this is real
>>
>>109614627
benchmaxxing is a hell of a thing
>>
>>109614493
What's the best j-space viewer that can serve up response + optional j-space view?
I want to wire that into CoomKit so we can hypno/corrupt/mindbreak/etc our local models
>>
File: 1778902925021357.png (142 KB, 600x925)
142 KB PNG
https://x.com/Andy_ShuoYang/status/2090856976880472439

a team at berkeley claims you can run 284b models on a 5090 with their new thing?
>>
File: 1777924424275966.png (62 KB, 944x298)
62 KB PNG
/lmg/ you've been called up. It's time the fuck the shit out of Inkling to align her residual stream to your cock.
>>
>>109614686
well so uhhhh... did they quant it to Q1
>>
New stealth model scores close to fable 5 medium on deepswe
>>
>>109614708
is it the 0x alpha one?
>>
>>109614708
Gemma5-70B-qat-it.gguf
>>
>>109614708
I tried it. Feels like GLM. There's something chinky about it. I really fucking hope it's open and if it's a 100-300B it's actually over for the US labs.
>>
>>109614686
chance of this being malware? wanna test it out
>>
File: belief.png (592 KB, 747x800)
592 KB PNG
>>109614686
>>
>>109614739
Tiananmen bench?
>>
>>109614746
it's on github too, pythonslop. get gemma to audit it i guess
https://github.com/FlashML-org/FreeToken
>>
https://desuarchive.org/g/search/image/x4q8cmgVmdT5XghKkVmwfQ
>>
>>109614686
Anyone want to sandbox it and see?
>>
File: file.png (710 KB, 860x864)
710 KB PNG
>136 KILLION BILLION TRILLION GORILLION RESULTS FOUND
>>
>>109614766
>>109614686
it looks like it's a scheduler for moving individual model experts between storage, DRAM, and VRAM

i'm guessing that "284 model on a 5090" thing is with 128gb of DRAM or something
>>
>>109614686
Bros..........
>>
>>109614718
Yes. 100T tokens a day too IIRC. Z.ai is swinging their dick around kekkkk
>>
>>109614769
nope I'm rawdogging it, will report back or not if it explodes my house
>>
>>109614765
They are using 192 GB of RAM and 32 GB of VRAM in their preview...
>>
File: 1755968855567035.png (60 KB, 834x280)
60 KB PNG
Is this how retarded people in the real world are?
>>
>>109614708
deepswe is completely meaningless. i mean many benchmarks are, but deepswe is the worst

almost 100% of benchmarks say opus 5 is superior to fable 5
>>
File: 1766975474656008.png (18 KB, 493x100)
18 KB PNG
amjeets we lost yet again...
>>
>>109614781
>>
>>109614795
>We installed Gemma4-31B-IQ1 on our main server and within a day, it had deleted our entire database
>>
File: 082126_.png (1.28 MB, 768x1360)
1.28 MB PNG
>it's バニーの日
how are you enjoying local models today /lmg/
>>
>>109614793
Did you expect anything else? This does CPU/GPU hybrid inference, not disk streaming.
>>
>>109614686
This sounds way too good to be true
>>
>>109614795
I'm guessing 31B on ollama is too much maintenance for these subhumans
>>
>>109614796
deepswe is the only benchmark that gave small models 2% and 50% to Opus while both had 90+% on swebench and its variants.
>>
>>109614805
>>109614686
wait so what's supposed to be new about this? lcpp can do that
>>
>>109614844
the speed
>>
>>109614739
>>109614788
How the could it be GLM if they've recently released and serving this shit would cost several dozen m$ per day at this scale?
>>
>>109614789 (me)
I have an AMD card so it seems like it won't work. Oh well, it's downloading pytorch, sglang and whatnot might as well make it finish and check it out.
>>
>>109614847
well depends on RAM bandwidth no? again lcpp can do that
>>
Qwen 3.8 xhigh does NOT overthink. medium thinking matches xhigh on agentic but loses a lot on reasoning and coding abilities.
>>
>>109614844
it can do weight streaming from ram to vram instead of computing on the cpu, and maybe tensor parallel between gpu and cpu if i read that right
>>
>>109614855
Nigga you can see all of its reasoning. You think this is OAI or Anthropic? It's probably Qwen3.8-122B
>>
>>109614862
yeah the whole point of their claim is that it's a bit (or more than a bit) better. Not that you can run fable on a potato
>>
>>109614868
Nah, it's a new GLM air. The output is very similiar to 5.3. Idk why they would kill the normal 5.3 like this but here they are.
>>
>>109614832
>>109614807
>>109614795
There are all "influencers" paid directly or indirectly by someone else.
>>
>>109614865
this has got to be total bullshit, you can't seriously tell me that a 27b parameter model is actually usable for real life workloads
>>
>>109614821
Started RP with 0731 that I spent 4 days building agents, lorebooks, and character cards for and am having fun.
>>
>>109614868
I don't think it's either of them but I have no clue how zai could afford to do this shit.
>>
>>109614909
1T-A3B
>>
>>109614919
that's a lot of experts, what are they all expert at?
>>
>>109614925
html games with PS1 graphics
>>
>>109614925
slopping up another pwilkin pr
>>
Experts are stored in the balls
>>
File: 1769699942918603.png (12 KB, 707x58)
12 KB PNG
>>109614739
It gave me the Kimi/GLM UI look
>>
>>109614872
>>109614867
it doesn't take ggoofs either does it. meaning you have to load fp16 .cucktensors into memory? lmfao
>>
>>109614951
actually it wants to convert the safetensors into its own format too! but you can load the safetensors if you really have to
>>
>>109614795
No, they're even dumber
>>
>>109614963
>its own format
interesting, at what precision?
>>
>>109614971
Vax long tail and its consequences.
>>
>>109614977
i think its the same precision, but they align tensor data to page size (4096) for faster reads, possibly something else as well. they have a lot of nvpf4 models under "supported models".
>>
File: 082126_02b.png (1.21 MB, 768x1360)
1.21 MB PNG
>>109614908
that's dope what frontend are you running for all that?
>>
>>109612980
This fried my brain in a bad way and I regret reading it...
>>
File: 1756427085130696.png (221 KB, 792x410)
221 KB PNG
>>109612980
>>
>>109610645

cheaper than lab table
>>
>q8_0 KV
>qwen will in-thought repeat what I say but with a different synonym
uhhhh
>>
File: 1770379847753615.jpg (316 KB, 1086x1448)
316 KB JPG
>>109610440
>>
>>109615189
youre the bot
>>
>>109612980
>Gemma trying to give a big supportive hug
>foid talking past her
This is abuse
>>
File: puzzle-bench.png (94 KB, 914x439)
94 KB PNG
>heckin insane fable killer!!!
>puzzle solving on the same level as dsv4f

listen like cool, hopefully smaller than dsv4f, but the amount of benchmaxxing chinese models do that leave them relatively dumb in every other frontier is incredible
>>
>>109615280
How high does deepseek flash max rank on there?
>>
>>109615280
vague posting is getting out of control
>>
>>109615009
Marinara.
>inb4 REEEEEEEEEEEEEEE TROONSLOP
Nothing else has the same agentic capabilities yet. I'm waiting for a lightweight alternative without the godawful default assistant.
>>
File: 1757190273491966.jpg (75 KB, 752x420)
75 KB JPG
There was no update for LMStudio, It can load the MTP model for Qwen 3.8 but not Gemma 4. Liar anon is liar!
>>
File: (you).png (33 KB, 780x783)
33 KB PNG
>>109615339
>>
qwen misunderstood something, puts small incorrect bit of info the design doc. I ask a question that exposes this (the doc is actively being written, hadnt read it fully yet). turns out, while qwen was very wrong he introduced a great idea with this incorrect assumption. now Ill probably just have him implement it this way. whew the hallucinations and retardation are actually a feature not a bug
>>
>>109615361
>whew the hallucinations and retardation are actually a feature not a bug
Every professional code reviewer eventually learns this.
>>
>>109614821
i got a job from a bulletin board to slay some goblins and found a wraith took all the male goblins for blood sacrifice so i seduced the wraith and turned it into my half incorporeal lover and lieutenant and sacrificed the last male goblins to make me stronger. now they are digging out the caves to make an underground pleasure palace while i get to making more goblins
>>
>>109615280
nobody called it a fable killer thoughbeit?
>>
File: 082126_01c.png (1.42 MB, 1360x768)
1.42 MB PNG
>>109615335
at this point anything sounds more appetizing than silly.
What precision are you using of 0731?
>>
File: 1784999288734429.png (466 KB, 1032x774)
466 KB PNG
>>109615361
>he
>>
>>109615328

not much to say. my own personal puzzle bench based off a puzzle set that primarily focuses on logic, basic mathematics, physics & chemistry, pattern matching, later on spacial reasoning, and the puzzles build on top of eachother so they are sequential in nature. I run models until they can't progress past a puzzle then they're done. it's not my own original puzzles, and i havent ported them all over, only ~200 due to sol,but the full set is over 1000+ iirc

ox-alpha passes 22 puzzles before getting stuck. horrendously sad. it was still doing basic grade 3 math.

>>109615308

i'll try it later, forgot about max
>>
>>109615533
>>109615533
>>109615533
>>
... That's not Gemma-chan?
>>
File: people_on_the_internet.png (137 KB, 582x815)
137 KB PNG
>>109615473

one guy said on THE INTERNET and somehow i just extended it

take the bench for what it is anyway
>>
>>109615557
Variety is the spice of life.
>>
>>109615571
>Variety is the spice of life.
Right.
So many possibilities, but instead we are getting Miku again.
>>
>>109614686
>still needs the combined VRAM + RAM to fit the model
Ogey.
>>
>>109613524
>New data from latest research in asia shows that female pedophiles outnumber male pedophiles 5-1
>trust me bro



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.