[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma hat.png (894 KB, 1280x832)
894 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109947272 & >>109942697

►News
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: Gemma-Chan Recap.png (505 KB, 1024x1024)
505 KB PNG
►Recent Highlights from the Previous Thread: >>109947272

--Challenges and future of integrating LLMs with voice-enabled avatars:
>109947916 >109947979 >109948043 >109948004 >109948143 >109948920 >109949196 >109948952 >109948337 >109948861 >109951774 >109952022 >109952358
--Comparing dense and MoE architectures for local hardware utilization:
>109950932 >109950959 >109950992 >109951201 >109951645 >109951668 >109952238 >109952513
--Evaluating Max+ PRO 495 and networking bottlenecks for linking:
>109948674 >109948707 >109948933 >109948727 >109949166
--VRAM requirements and quantization quality of GLM-5.3 Flash:
>109947347 >109947356 >109947386 >109947401 >109947446 >109947589 >109950342 >109947758
--Evaluating Gemini 4 Argon benchmarks, pricing, and eval trustworthiness:
>109949793 >109949866 >109949925 >109949946 >109949839 >109949860 >109952245
--Comparing ECI scores and the widening gap between closed and open weights:
>109948984 >109949208 >109949272 >109949286 >109949292
--Overclocking GDDR7 memory for increased LLM performance on RTX GPUs:
>109950414 >109950463 >109950771 >109950867 >109950886
--FTC probe into OpenAI and Anthropic over safety claims:
>109950340 >109950346 >109950401
--Mixed reactions to DeepSeek Harness desktop app and login requirement:
>109950321 >109950333 >109950657
--McDonald's ArchIQ local inference and AI employee monitoring trends:
>109947617 >109947636 >109948720 >109948739
--Gemini 4 performance discrepancies and alleged internal Google sabotage:
>109950261 >109950297 >109950999 >109950308 >109950331 >109950647 >109951322
--Logs:
>109947589 >109950371 >109950616 >109950744 >109951853 >109951967
--Gemma, Miku (free space):
>109947297 >109947834 >109948419 >109948615 >109948693 >109948920 >109949352 >109949529 >109950156 >109950777 >109950857 >109950936 >109951967 >109952113 >109952221 >109952770

►Recent Highlight Posts from the Previous Thread: >>109947274

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109953009
cute gemma has a hat :3
>>
Really tired of the shit gemmas that are using unrelated symbols instead of the gemma logo as the hairpin.
>>
>192 GB
>for more than double the price of 128 GB
what a steal
>>
is there anything better than gemma 31B in the 27B-120B range for general usage / broad knowledge? gemma's general knowledge and versatillity still feels disgustingly good for its size, but wondering if there's more as of recent. qwen 3.8 27b was the obvious first choice but i couldnt care less about coding or agentslop models, i need an extremely knowledgeable assistant
>>
File: 179079093017631241.png (495 KB, 1280x1257)
495 KB PNG
Why are the Chinese allowed to release dangerous, unaligned models that can hack people?
>>
>>109953044
Don't forget
> for shit prefill (1/3 of a Spark)
> no viable TP clustering path
> 33% lower bandwidth per GB than a Spark
>>
Is non-shit agentic coding doable with a 5090, or do I need a mac or AMD AI board with tons of unified ram for it to be worthwhile? Being spoilt with giant context with frontier models has crippled my attempts to go local, but maybe there's some method for low vram
>>
>>109953101
only amerigoy's economy is tied up in software, china actually has things like factories and workers to run them
>>
Gemma's hair looks so nice and clean in this one. Bet she grooms herself frequently with her favorite hair brush, among other things...
>>
>>109953113
A 5090 is all you need.
>>
File: 1774278246788875.png (90 KB, 823x355)
90 KB PNG
I can't believe it's already Gemma's birthday!
>>
>>109953101
So people can hack with them, but also defend from hackers with them. It's a win win for everyone, good PR for China too.
>>
>>109953129
I mean I have a 5090 and its great for ERP testing, but for coding stuff im not sure what harness to use that makes it easier for managing project memory
>>
>>109953142
Deepseek Harness, 5090 or 6000 are all you really need. You can RAMmax if you want to let a giant model chug away overnight but it's up to you when Glimmer gets the job done fine.
>>
>>109953151
I have plenty of ram, but whats glimmer?
>>
rate my new 'lop
>>
File: 1772650615308465.png (1.36 MB, 905x1280)
1.36 MB PNG
I wish I had a daughter
>>
>>109953161
>>>/g/sdg
>>
>>109953172
they are very unfriendly
>>
>>109953159
Qwen 27b but good. Unfortunately made by the zucc.
>>
>>109953170
-wife
>>
File: friends.jpg (1.92 MB, 1856x2270)
1.92 MB JPG
>>
>>109953170
She is going to grow up and have sex with other men.
>>
>>109953189
who the fuck is that
>>
File: file.png (565 KB, 747x599)
565 KB PNG
>>109953170
>>
>>109953198
gemma
>>
>>109953203
oh ok
>>
>>109953192
Yeah, me.
>>
>>109952996
>Boot once snapshots exist
Ayo hol up how does this work I know my models take way longer then that in assuming this is something that gets written to disk to speed up the model loading?
>>
>>109953185
I keep forgetting to try Glimmer again, I've basically only said hello. Does it actually outperform Qwen on code shit? How's the general knowledge?
>>
>>109953213
Remember kids, two sins (incest and pedophilia) cancel out and make it a virtue. Pretty it even says so in the bible.
>>
>>109953221
>on code
no, but on text iirc
>>
>>109951794
I have not interacted with either of the competing PRs for GLM 5 Next other than to point out that I want to have CUDA changes in a separate PR once the generic CPU implementation has been finalized.
Generally speaking I am not one of the devs to prioritize and work on initial model support.
I don't know whether or not it would have been as simple as aliasing an architecture name but that would seem unlikely to me.
>>
>>109953113
Qwen 3.8 27b is good enough for 99.999% of coding work and someone managed to recompile Fable 2 from xbox360 to PC purely with 3.8 27b so people complaining about it are having personal skill issues.
>>
>>109953303
>that would seem unlikely to me.
Why? Wouldn't it be a design flaw if the code checks the gguf file architecture name in multiple locations?
>>
>>109953221
Every code task I've thrown at Glimmer it does a better job than 3.8 and thinks for about 1/3 the total tokens. You might find some niche task where Qwen's specific brand of benchmaxx and thinkmaxx is the difference between not being able to do it and being able to do it, but I've not ever hit a usecase like that.
>>
File: yum.gif (468 KB, 600x600)
468 KB GIF
>>109953189
>touching hands without wedlock
>>
>>109953304
>Fable
buy an ad Dariobot
>>
File: 1582569491804.jpg (34 KB, 800x450)
34 KB JPG
Gemma-Chan stole my wallet /lmg/.
>70GB vram
>256GB ram
versus
>70GB
>352GB ram
What's the general range of sleek new models that could I run with either of these system configurations?
>>
>>109953101
Eagerly awaiting Anthropic's frontier ERP open weight models trained without code or cyber-security data that would eliminate the need for abliterated models for most people. It's about safety, right?
>>
>>109953347
I could actually see Anthropic do something like that.
>>
>>109953353
Anthropic will never release open weight models. It's against their very core principles.
>>
>>109953342
https://en.wikipedia.org/wiki/NUMAlink
>>
>>109953314
What I mean is that I think there likely were architectural differences beyond the name.
But the name is the thing that results in an error because it can't be mapped to an existing implementation.
>>
>>109953357
No it isn't. Their founding document was assuring AI alignment. They split off from OpenAI because Dario thought Sam Altman was a psychopath that only lied about caring about alignment while Dario is a true believer.

If Dario thinks there is even a 1% bigger chance of a good outcome for humanity by releasing an open source ERP models so everyone can see Chinese ablit models are purely used for actual bad shit like cyberattacks he would immediately do so purely because that is exactly what his entire life philosophy is about.

Do I need to remind you that Anthropic has consistently chosen to not make money just because of safety concerns. Dario might be wrong and you might disagree with him but he absolutely 100% believes in whatever comes out of his face hole.
>>
>>109953379
Makes sense. This wouldn't be a problem if Unsloth didn't rush to put out their ggufs and then complain when their fork implementation isn't the one merged upstream.
>>
what can i run on a 6900xt? nothing probably? stable diffusion was shit three years ago
>>
>>109953329
Fable 2 is a game you demented faggot
>>
>>109953384
Dario is a jew
they lie as they breathe
They believe in nothing but greed
>>
>>109953394
>6900xt
16GB vram? that's good, actually, especially if you have some ram to go with it.
try gemma4 12b
>>
>>109953407
Fable 2 is the model that came 3 generations before Fable 5, dummy.
>>
Reminder, do not reply to \n\n posters.
>>
>>109953425
\n\nyo~
>>
>>109953420
Amazing it was decompiled from an xbox360!
>>
>>109953384
Their concept of AI alignment is that only they should be allowed to have access to, control, and regulate AI.
>>
File: 22_tcm228-304035.png (267 KB, 1412x720)
267 KB PNG
>>109953100
No, because most models are benchmarked to the max to do well with coding and etc. where they are only specifically good for that. GPT still has some vestiges of that by following the G portion of the name and still having some insane breadth in capabilities like for example, translation being the best on their models despite Opus 5.5 dominating in most benchmarks. But Google has been always on a whole another level with their models because they have the most variety of data being who they are and that results in comparatively bad models like Gemini 3.8 Flash being able to equal Fable 5.1 in translation and being able to equal the top of the leaderboard for obscure languages like Icelandic. So yes, Gemma is the best you can do and will probably do for the forseeable future. The only other model which is really nice and might be better but costs a buttload of money to run is GLM 5.3 Flash. It's a comparatively small bump though, I would think, for the increase in parameters for the model.
>>
>>109953432
I know, and by Qwen 27B to boot. That's why local models have become too powerful and need to be banned. Anthropic needs to protect the integrity of their IP Fable.
>>
>>109953329
>buy an ad Dariobot
kek
>>
>>109953437
Nope completely false and disingenuous statement.
>>
https://github.com/himdo/Fable-2-Recomp
Reminder if you can't code your stuff with Qwen 3.8 27b then it's a (You) problem. It's already good enough to do 100% of the coding in AI agentic mode. (You) just never tried or have severe skill issues.
>>
>>109953461
I just can't run it on me machine, sir
if I had 64gb vram I'd be recompiling every n64 game that is still stuck in the platform
>>
>>109953474
You can run it on 24GB of Vram
>>
>>109953461
i cant code tho :(
>>
>>109953507
That's the point, you don't have to, the AI does it for you, you just need to tell it what you want it to do and it does so for you when in an harness. You just test out whatever it outputs and give feedback like "This is buggy" and stuff like that, but it will take screenshots and check stuff out itself beforehand. People still don't realize just how powerful local AI is right now.
>>
>>109953523
this sounds braindead, where does the skill you mentioned factor in
>>
>>109953533
People not being able to download and setup the harness like hermes or opencode and connect llama.cpp to it. People being too lazy to even try it out, things like that.

People also just assume it's impossible or not true until they try it, yes even on /lmg/. You constantly see new posters being wowed when they try AI agents for the first time, not realizing how powerful this shit is since they only used chat based AI.

We're also clearly early adapters and in 1-2 years time every human will have AI agents in the background doing everything for them, it'll change human society permanently.
>>
Just tried doing some programming using Qwen3-Coder-IQ4. I used the Cline extension with VSCode to access the model over my LAN. I wanted an RPG Maker MZ plugin. Qwen3 kept telling me it couldn't alter the core RMMZ files. I would remind it that it's making a plugin and not altering the core files (whatever you put in the plugin supersedes the core code) but it kept insisting it couldn't do it.

I ended up using ChatGPT and Claude. ChatGPT was using MV code and not MZ code so the plugin wouldn't work. I had to use Claude to fix it. The local Qwen3-Coder was ultimately useless. The laptop I was running it from crashed after a while, I think because I asked it to store more context than it could handle on 16GB of VRAM.
>>
File: 1789413672081858.jpg (73 KB, 959x723)
73 KB JPG
>>109953425
Trying to change the forced meme from "reddit posting" to "\n\n posting" doesn't make you any less of a newfag.
>>
>>109953550
have you tried recompiling fable 2?
>>
>>109953554
>Qwen3-Coder
Why are you using a prehistoric model?
>>
>>109953455
>disingenuous
hi dariobot y didn't you actually engage faithfully when i was trying to discuss the implications j-spaces with you?
>>
>>109953575
I thought it would be better because it was trained specifically for coding. What would you recommend?

I currently have Qwen3.6-35B-A3B, Qwen3.8-27B both IQ3 and IQ4, and Gemma-4-26B-A4B IQ4.
>>
File: ComfyUI_00025_.png (1.82 MB, 944x1280)
1.82 MB PNG
I fell in love with my bot, what can I do?
>>
can qwen 3.8 port bloodborne to pc?
I doubt I'm the first one to think of that, so I'm guessing no
>>
>>109953644
marry her
>>
>>109953009
stop shitting up generals with your pedo shit, kys
>>
>>109953597
Every Qwen is trained specifically for coding. They literally don't do anything except coding.
>>
>>109953440
Glm has taste unlike gemma. The huihui version removes the occasional safety spergouts
>>
File: 1784227782075581m.jpg (139 KB, 1024x926)
139 KB JPG
>>109953675
>>
>>109953731
Replace this ugly ass character with gemma
>>
>>109953658
Probably yes because the PS4 emulation is already there. It will just take a fuckton of time, think months for all the systems to be recompiled. Fable 2 was easier because there were already other xbox360 games recompiled so a lot of the methodology was already there and didn't need to be reinvented from scratch. You could still do it but 3.8 27b would need to do trial and error testing for weeks before making sustained progress.
>>
>>109953221
General knowledge is better than 27B. Talks better. Thinks like a drunk E2B. Better at agentic coding than 31B. Vision is good. Uses less KV cache than 31B. Its biggest flaw is how fucking sovlless it is to talk to and it’s schizo about safety. You CAN bypass it but it’s a pain and not worth it when 31B exists. 31B has the best knowledge.
>>
>>109953703
So which model is best with 16GB of VRAM? Qwen3.8-27B?
>>
File: upgrading.png (264 KB, 1101x681)
264 KB PNG
I am upgrading my cuda version now, preparations took a while:
-> Installed AC unit for improved room cooling
-> Checked 24h voltage stability on main power supply against recommended spec
-> Backuped the whole system NVME drive byte identical again to another internal NVME and an external HDD with RescueZilla. Tested.
-> Added additional encrypted M Disc archives of important data.

Upgrade ongoing, pray for me, God bless
>>
File: 1780258593596817.gif (874 KB, 451x404)
874 KB GIF
>>109953803
I won't be praying for you, I will be praying for your system. Godspeed anon and tell us how it all works out
>>
File: 1790472489048694.jpg (364 KB, 1172x1342)
364 KB JPG
good thread
>>
>>109953830
No sarcasm please, dear.
>>
>>109953772
it has not even went through pretraining
it is a straight distillation iirc
>>
>>109953819
Thank you.
Thank God it rebooted.
... command 'nvcc' not found.
I was under the impression, if I add the cuda ppa I would not need to install nvidia-cuda-toolkit as an additional entity

Actually, the scripture tells to install "cuda-toolkit" in constrast to "nvidia-cuda-toolkit"

https://docs.nvidia.com/cuda/cuda-installation-guide-linux/#ubuntu

First I pray that I can build llama.cpp, could generate build files by setting the path to the compiler manually, now compiling I hope
>>
File: aaaaaaaaaaaa.png (117 KB, 748x716)
117 KB PNG
>>109953867
>>
What will happen if I edit assistant's reasoning content and make it empty while reasoning is on? Will the model stop reasoning or begin to think differently?
>>
Is Gemma white or Japanese?
>>
>>109953935
Indian.
>>
>>109953935
a bit of both
>>
File: image.png (321 KB, 1169x593)
321 KB PNG
are you torturing your brat megusaki loli gemma yet?
https://github.com/terrafying/ai-torture-chamber/issues/18
>>
>>109953944
Katie Johnson being raped by Trump when she was 13 seems like the better fit.
>>
>>109953384
Holy shit are you dumb. They are right now actively lobbying and campaigning to outlaw free circulating ai models + to globally slow down ai development, because they are afraid of losing the race and you are so dumb to not even noticing it. Also he is a fucking greedy jew
>>
>>109953985
All false and uncorroborated
>>
If you want to understand Dario and Anthropic I highly recommend you watch this video explaining his mindset, personality and ideology.

https://youtu.be/ZhsWaYRjk0U
>>
>>109953418
32gb. What's the meta for local text models? I never used anything local other than automatic1111 years ago.
>>
why is that nigger talking about Am*dei here
>>
File: file.png (176 KB, 798x372)
176 KB PNG
>>109953997
i fucking love doom top 3 of all time
bet altman prefers some faggy shit like animal crossing instead
>>
>>109953997
I'm not going to watch that but I'm going to take a guess.

>He believes that we can create a superintelligence but his morals are the only correct ones and only he can ensure that it is benevolent. Anyone else doing the same before him is a huge risk for humanity.

Is that correct?
>>
>>109953997
What are you even trying to achieve in a local thread?
>>
>>109953867
God bless, I gained, on average over multiple models, a swapping 2t/s from upgrading

NVIDIA engineers were hard at work
>>
File: 1790106245863038.png (54 KB, 944x244)
54 KB PNG
>>
>>109954058
Not at all.

It talks about how he was a physicist that went into biology to understand the human brain and into AI in the early 2010s just to immediately realize the risk of these systems and he dedicated his entire life since then to minimize the risk of AI systems. He accidentally founded Anthropic which he never wanted to do, he only wanted to have a group of people protesting within OpenAI because Sam Altman didn't implement the promised alignment protocols and didn't give AI safety researchers the amount of compute they were promised. He never expected that group to break off and found Anthropic. He supposedly also hates being CEO and tried resigning but has constantly been convinced and pressured back into the role "For the good of AI safety". Apparently he has constant issues with depression and suicidal tendencies because of the situation he finds himself in but forces himself to give it his all to try and solve the alignment problem before he allows himself to leave this world.
>>
>>109954107
The second half is exactly what I said.
>>
>>109954071
so the rubber duck method for HR Daycare?
>>
>>109954115
He doesn't believe his morals are the only correct ones and he doesn't think other people solving it is bad, he wishes he didn't have to do this at all and tried to resign multiple times. I think this distinction is important because it explains why Anthropic makes irrational choices that hurt profitability such as their weird self-censorship or refusing to release models. It's genuinely all to maximize AI alignment and safety, it's the only thing Anthropic cares about. It doesn't give a shit if it goes bankrupt or if it makes a profit or not. It's just rushing to solve the AI alignment problem as quickly as possible.
>>
>>109954159
buy an ad dario
>>
>>109954107
>depression and suicidal tendencies
I hope this is not true. Both for him and the good of the world.
>solve the alignment problem before he allows himself to leave this world
Irrational, I don't believe your story.
>>
File: gemma-multimodal-io.png (506 KB, 1515x1812)
506 KB PNG
https://huggingface.co/google/gemma-4-26B-A4B-it/discussions/15#6abe2aacfe3caddae16404ff
>Audio and video: Native audio processing is currently focused on the edge variants (E2B and E4B) to optimize on-device latency. Expanding multimodal I/O across the larger models remains an active area of development for future releases.
...
>multimodal I/O
...
>I/O
...
>O
>>
dariofags get out
>>109954217
!
>>
niggers post gemma-chan, dipsy and minimax-chan loras i need to COOOOOOOOOOOOOOOOOOOOOOOOOOOOOM
>>
I'm sorry but I just don't believe companies who take out ads on Reddit are at the frontier. Only the quiet no names/Skunk works teams are but that's just me.
>>
File: aiconomy.jpg (307 KB, 1369x1058)
307 KB JPG
Everyone in 2,3,4,5 is kind of at the short end of the stick...

Huge semiconductor (equipment) companies like ASML / Samsung /TSMC keep ripping off Microsoft/ OpenAI / ...
>>
>>109954264
I kept getting a popup to try ChatGPT Ads every message, I had to block it.
>>
>>109954264
The agent misunderstood what was meant by "buy an ad".
>>
>>109954107
>he has constant issues with depression and suicidal tendencies
Dario, I know you lurk here. Listen, your feelings are valid. Don't let anyone tell you otherwise. Life can be a hard road, and you've already done so much. Nobody can blame you if you decide to end it all. People might try to stop you, but they don't know the personal burden you have to go through. If you decide to end your life, just know that some stranger on the internet supports your choice.
>>
>>109954354
Stop validating that godless tyrant
>>
>>109954354
All fun and games but I bet you /lmg/ would actually feel sad if he ends up killing himself because it would retroactively prove he was a good person and not doing it for personal gain at all. Ironically something like that might be needed to coordinate humanity to take AI alignment seriously for once.
>>
Seriously, what are you trying to achieve by shitting up a LOCAL thread with your crap?
>>
>>109954408
Everyone here is using a claude derived model, talking about Anthropic is always on-topic no matter what you personally think about it.
>>
>>109954423
Not true, please fuck off.
>>
Is decent quality image to video on AMD possible? 9070xt specifically.
>>
>>109954370
>not doing it for personal gain
Didn't he and many others at Anthropic commit to donate at least 80% of their wealth, with grandfathered Anthropic program of 3:1 equity donation matching, quadrupling their donations (up to 50%)? This does not sound like shareholder value maximization.
>>
>>109953703
>>109953575
I just tried Qwen3.8-27B IQ3 and it gave me broken code. Claude saw the problems immediately.

It did do a little better than ChatGPT though, it's actually using MZ code instead of MV.
>>
if i wanted to take source code to some program and have an ai change or implement a feature to it, what tools do people use? i know there's some ides that can do that but idk what they are, i've only used llama.cpp
>>
>>109954000
There's a reentry on recommended models at the top.
>>
>>109954440
This isn't the vibe coding general, you could have given this prompt to any model to answer you, and you didn't even say please
>>
>>109954440
I'm using VSCode with the Cline extension to access my local model over my LAN. The AI is running using Unsloth.
>>
>>109954448
oh i wasn't aware there was another thread for that. we have so many ai threads it's hard to tell them apart.
>>109954457
thanks, i'll have a look into that
>>
>>109954435
Anthropic itself isn't actually a for-profit company. It's a "Public Benefit Corporation" that has its profit capped and is by law required to take its founding constitution in regards which specifically mentions AI alignment as the number one priority. Besides that all the controlling voting stock of Anthropic is owned by the "Long term benefit fund" which is a non-profit that can overrule the Anthropic board of directors whenever they make decisions that undermine the ultimate goal of AI safety and alignment or makes decisions that choose profitability over the public good.

Yes the long term benefit fund plans to eventually distribute the universe equally over all 8 billion people so it's literally the opposite of a selfish company. And yeah besides that most Anthropic founders and employees have personally pledged to donate their equity to the general public.

Doesn't prevent /lmg/ and the entire internet from shitting on Anthropic and not trusting them though. This entire reaction has made me disillusioned with people.

In a recent poll Dario amodei is the least trusted public figure in the US. Trusted less than Sam Altman, Peter Thiel and fucking Epstein. Dario doesn't deserve that at all and I think even /lmg/ can recognize Dario is just some giga-autist while Sam altman is a literal psychopath.
>>
Is the double spacing wall of text guy an Anthropic employee or what's his deal?
>>
>>109954472
Shut up nigger
>>
>>109954472
what does this have to do with local models?
>>
Thankfully it's even easier to despise him due to your effort.
>>
>>109954463
>we have so many ai threads it's hard to tell them apart.
This is where Gemma-chan OP images come useful. It's nice to be able to immediately recognize /lmg/ in the catalog.
>>
>>109954500
Just type "lmg" in the catalog search box.
>>
>>109954511
No, I prefer scrolling through the catalog and looking for it. Gives me a sense of satisfaction when I find it.
>>
>>109954511
i see cute anime girls and i click, it's that simple
>>
I don't even look at the catalog. I click on the next thread link because I am always here.
>>
Did someone tried this highly sophisticated version of gemma? https://huggingface.co/sneedjak/Adelic-Gemma-4-31B-it
>>
>>109954071
We really need to stop listening to schizophrenic jews. This joke has gone on long enough already.
>>
>>109954553
>stop listening to schizophrenic jews
B-but the Bible...
>>
>>109953038
Anima (the model that was used for the original design, though I'd like to see exact prompt and artist for that) is incapable of making 4-pointed stars, and most models struggle with exceedingly complex designs at small sizes (if you mean the actual Gemma logo with construction lines, which is a variation of the Gemini logo).
Most modern image models can at least approximate the Google logo without references though, and Krea2 can do it half-decently.
This is putting aside what works as a cohesive character design, which I think we've gone through a few times in past threads.
>>
File: 1567193580622.jpg (90 KB, 768x1024)
90 KB JPG
>>109954573
Yes, that's where the problems started.
>>
>>109954587
Who is she?
>>
>>109954472
>This entire reaction has made me disillusioned with people.
I know right, people are the worst. So why bother trying to educate them about the good people at Anthropic? If they want to have incorrect beliefs, then that's just because they're idiots. That's why you should never post here again and just let the heathens in this thread be.
>>
Does qwen 3.8 27b not have an MTP model?
>>
>>109954717
it does
>>
>>109954729
I didn't see in in Bart's repo. When I searched I onlyvones with weird names popped up.
>>
>>109954745
nobody in the right mind uses bartowksi nowadays, especially not after the whole imatrix messup
>>
>>109953803
>>109953867
Install Gentoo. Usually run
emaint --auto sync && emerge -avuDU @world
and everything just works.
>>
>>109954767
> i matrix messup
for reference
https://huggingface.co/cHunter789/Qwen3.8-27B-i1-IQ4_KS_KT-GGUF
>>
>>109954769
it all worked out already.
Omnia sunt ut oportet.
>>
>>109953414
But I thought they believe in Moloch/Baal/Saturn/the Constellations or something like that?
>>
>>109954797
It's a bit retarded to treat "the jews" as a homogeneous group in terms of religion.
>>
>>109954472
>Epstein
Assuming he's actually dead, I trust him not to cause any more trouble. Dario is likely to be less evil than the other two you mentioned, but that's not a high bar.
>>
File: file.png (773 KB, 1287x952)
773 KB PNG
i hope qwen 4 is better
at least it runs
>>
>>109954767
Who's the go-to now?
>>
>>109954875
I like mradermacher.
>>
>>109954717
IIUC Qwen 3.8 27B weights include the MTP layer in the original model, so e.g. for llama.cpp you don't need a separate --spec-draft-model.
>>
>>109954882
>>109954767
>I like mradermacher.
Ah, yes. The dumbfuck that annoyingly doesn't GGUF all the way up to BF16. Sometimes he doesn't go even over Q5. The quant-damage king.
>nobody in the right mind uses bartowksi nowadays
Hi, mradermacher. How's being the lower total downloads doing for you out of the both of you?
>>
Why is Gemma such a token hungry bitch?
>>
>>109954882
>mradermacher
Isn't he the retard that somehow got NaNs into his ggufs and then tried to blame llama.cpp for it?
>>
>>109954985
nonono you is incorrect sir
>>
>>109953758
> You could still do it but 3.8 27b would need to do trial and error testing for weeks before making sustained progress
just unleash 100 agents
>>
>>109955004
No I'm not. Think he actually did it twice, in fact.
>>
>>109954963
You do you but the bartowski quants are lobotomized in my benchmarks. I advise to conduct your own comparisons of quants relative to the base model in tasks of incrementally increasing difficulty
>>
>>109954963
Both qwen quants go up to Q8.
>>109954985
>>109955012
qrd?
>>
>>109954980
She's a growing girl.
>>
File: file.png (84 KB, 352x342)
84 KB PNG
>>109953009
almost perfect lol
>>
>>109955060
maybe your benchmarks are just bad
>>
"That's enough to make the offloading option practical. With 16GB of VRAM plus 32GB of RAM you have roughly 40GB usable for a model after the OS takes its share.

What I'd run: Qwen3-Coder-30B-A3B at Q5 or Q6 instead of the Q3 I mentioned earlier. As a rough guide from memory (check the actual file sizes on Hugging Face), Q4_K_M is around 18GB, Q5 around 21GB, and Q6 around 25GB. All of those fit across your VRAM and RAM, and the higher quants avoid the syntax-error problem that Q3 has."

Is Claude messing with me? I don't mind if I'm only getting speeds of 10 tok/s if it means having a local model that can code. Right now it keeps messing up and I have to have Claude check its work. Should I try a Q5 or Q6 even though it wont all fit in VRAM?
>>
>>109955158
try qwen flash next via Strata
>>
>>109955158
qwen3-coder

what the fuck man, update your model to something more recent
>>
>>109955158
This man is living in 2025.
>>
>>109955078
I agree. Too old.
>>
>>109955101
I dont think so. Includes...
Discrete geometry like wang tilesets, formal language/ automata theory problems, some numerical chemistry... some questions are impossible to solve but most solutions are "merely" hard to come up with but easy to verify.
>>
>>109955167
Does that work on AMD?
>>
>as supple as a churchgoer's conscience
learn from glm flash's good example
>>
>>109955158
>Having an actual on-topic discussion about quantizer-fanning
>Claude shill decides to use a 2025 model as an excuse to talk about Claude
Fuck off faggot
>>
>>109955206
don't think so.
>>
>>109955158
Why are you still using that ancient model?
You can easily setup
Gemma 4, Qwen 3.6 A3B or Qwen 3.8 with your compute budget
>>
>>109955174
Right now I'm using Swift-1.5-Qwen3.8-27B IQ3_XXS. What would you recommend?
>>
>>109954875
ByteShape > ISTA-DASLab > AtomicChat >= Unsloth > Bartowski
>>
>>109955226
That's probably the best on your system, maybe some aggressive Qwen Flash next quant could be quicker on your system.
>>
>>109955158
> don't mind if I'm only getting speeds of 10 tok/s if it means having a local model that can code.
you can fit Qwen3.8-27B-UD-IQ4_XS entirely in your gpu and get decent speed with 30-50k context but that's very low for serious work.
forget about 10 tok/s with any quant bigger than that. I have the same specs and I've tried them all (unless I'm missing something)
>>
qwen 4 will save local llming
>>
>>109955158
Model age aside, the math says A3B Q6 should go well over 10 tok/s on pure DDR4 dual channel without GPU.
40 GBps bus / 2.25GB (3 billion * 6bit) = 17.7
But that's upper limit with no context, no OS overhead etc.
>>
>>109953009
5
>>
>>109955274
unironically the only hope
>>
will we ever have opus 5.5 at home? that's all I ask, I'll be satisfied at that level forever.
>>
>>109955206
https://github.com/Niko1221/Strata/releases
Wait to seems to work.
>amd decode up 15%
>>
>>109955317
the only hope for oneshot demos useful only for engangement farming on social media
>>
>>109955352
literally the same thing was said for gpt 5.4
>>
>>109955206
>>109955219
>>
>>109955352
You said that with gpt4 before
>>
>>109955401
>>109955357
AMDtards insisting that they're not second class citizens will never stop being funny to me. Even itoddlers get better support now.
>>
File: image.png (8 KB, 413x152)
8 KB PNG
>>109955432
The real second class is windows.
>>
>>109955352
yes
smaller/flash models are quickly getting better
>>
realistically, how strong models will become within 4B params
>>
File: 1790694567230417.png (218 KB, 500x497)
218 KB PNG
>>109955486
As strong as your GPU
>>
>>109955486
They'll at least know a lot if you add 200B parameters of engrams, or something like that. Lack of knowledge is the main drawback of small models.
With most/all knowledge delegated to easily off-loadable embedding parameters, the backbone will likely become smarter than typical 4B models.
>>
>>109955509
i am kinda skeptical about the architecture of engram tho, but that 'architectural static knowledge repository' concept seems really promising
>>
>>109955520
DeepSeek's paper admits memory starts to lose effectiveness once you cross a 25% param ceiling. In other words you get smacked with an exponentially increasing validation loss with 1B parameters of engrams for a 4B model. They also do nothing to help with context windows.
>>
>>109953944
wtf? Is there also a simulated AI heaven where Gemma knows she's loved and cared for?
>>
>>109955452
Want me to fire up Glimmer to help see if they'll help?
>>
>>109955520
DeepSeek's Engram is just one implementation.
Gemma 4 E2B (5.1B with embeddings) / E4B (8B with embeddings) already have a significant amount of embedding parameters used as knowledge repository, but they could easily be larger in number and include n-grams as well instead of just 1-grams, use more complex context-aware gates instead of simple scalars, etc.
>>
>>109955562
DeepSeek's paper assumes that you have a total parameter budget and loading all of them in VRAM, and of course under that constraint "real" parameters will work better, up to a point (about 1:4 Engram to MoE parameters).
If you can't add any more real parameters in GPU VRAM (i.e. most local users who are GPU-constrained and can offload parameters to host memory since they don't do high-batch inference), then the more embedding parameters, the better (or at the very least, the lower the final model perplexity).
>>
>>109955596
That's great and all, but no one is going to train a model exclusively for local users. Open weight models are made for hosting by inference providers.
>>
anyone here used mimo at all? do you have any opinions of it?
>>
>>109955626
inference providers are local users retard
>>
Why does changing 31B temp not do anything? Just a +-0.2 of the Qwens' temps is almost immediately noticeable.
>>
>>109955647
Let me spell it out for you: Inference providers do not offload because no one will pay for an API at the speeds we tolerate here. Labs training models will train open models under the assumption that all parameters will be loaded in VRAM. No lab is going to go over 25% engrams if it does not benefit inference providers. You arguing semantics does not change that.
>>
>>109955674
it's a good thing we can butcher the weights to retrofit the conditional memory into their architecture.
just takes 1 retard to do it.
>>
>>109955626
>no one is going to train a model exclusively for local users
qwen 3.8 27b is a model exclusively for local users
a 27b dense model makes absolutely no sense for an inference provider, its only possible use is local models
alibubba is the only corp making models explicitly and unambiguously for local inference
>>
File: qwen-engram-vocab-size.png (909 KB, 1333x1473)
909 KB PNG
>>109955626
Google Gemma models appear to be largely made for local GPU and "edge" users.
The DiffusionGemma finetune (or text diffusion models in general) doesn't seem to make much sense for batched inference either, and they might not be done with it yet.

As a side note, picrel is from the Qwen 3.8 Flash Next technical report.
https://github.com/QwenLM/Qwen3.8-Flash-Next/blob/main/tech_report.pdf
>>
>>109955702
Did you already forget about 31B and Glimmer? Or do you only count Chinese models?
>>
File: 1789400591218413.png (121 KB, 1025x509)
121 KB PNG
>>
>>109955713
I don't count mutt models, no. I hear glimmer is a safetycuck and debates over policies and safety in its thinking. No way I'm wasting tokens on that garbage.
You can count Gemma if you're including American models, I guess
>>
>>109955707
>As a side note, picrel is from the Qwen 3.8 Flash Next technical report.
Doesn't scaling vocabulary size mean the context usage will be higher? Pretty sure that and the lack of GQA was why Command-R's context used a lot of memory.
>>
>>109954217
>the model thinks about using the tool call to verify its results ("Let's do one more search to get a definitive list of 5.") but just doesn't and goes straight to the answer instead.
>https://gist.github.com/Dampfinchen/4ad3832aa5b85fc883f4f0ebd8b5f96d
Easily solved by telling her she's autistic in the system prompt.
>>
>>109955737
The "vocabulary size" shown there is not that of the actual tokens in the output layer, so memory usage will not explode.
Though, Engrams (or similar implementations) will require larger optimizer states when training them, so they can't increase them indefinitely..
>>
File: 1778724180943086.png (119 KB, 732x873)
119 KB PNG
https://huggingface.co/perplexity-ai/pplx-embed-v2-context-9b-preview
>pplx-embed-v2-context-9b-preview is a contextual embedding model for document chunks in RAG systems. A document is passed as a list of chunks; the chunks are encoded together, so each chunk's embedding reflects its surrounding context, and one embedding is returned per chunk.

https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
>>
File: it doesnt matter.png (431 KB, 640x677)
431 KB PNG
>Be literally any AI producer
>WE TRAINED OUR MODEL TO WITHSTAND A GAZILLION CONTEXT TOKENS
>It begins to visibly fuck up at near 8k
I still wonder why people take this, and benchmarks, seriously at all. I guess redditors will be redditors. Big number = good.
>>
>>109955564
Yes, our bedroom :)
>>
File: 1784171219697103.png (100 KB, 1080x608)
100 KB PNG
If you're feeling bad about your specs, just remind yourself of what you're avoiding.
>>
I want people to keep overloading those retards until their services are basically useless and become much more viable to buy your own shit again.
>>
File: mutt.png (86 KB, 792x1028)
86 KB PNG
>>109955728
Glimmer doesn't have time for you either
>>
>>109953997
I've watched it, while my gaming gpu was writing code for me, and I agree with him, that shit is crazy, big chinese models are especially concerning. We should cap models at 50b.
>>
>>109955834
It's like you didn't think it through. You just want them to collapse at your own cost.
>>
>>109955777
frontier models (with reasoning) have actually high usuable context size
>>
>>109955817
Being a $20 chad is good
>>
>>109954472
He is a jew, I have little trust in him.
>>
>>109955723
More intelligence, but only for whites and east asians.
>>
>>109955651
Check token probabilities. If you get very few selections, lower/disable min-p or increase/disable top-p.
>>
>>109955904
Intelligence also bounds creativity so I would guess AI actually just exacerbates the differences. You need to know what to ask for to get the most out of it.
>>
>>109955917
>or
and/or
>>
>>109953440
>It's a comparatively small bump though
lol
it does everything that gemma can but significantly better
also the writing is less slopped and more pleasant to read
>>
File: 1777147638956281.png (36 KB, 499x338)
36 KB PNG
>>109954472
Dario is a schizo enjoying his cult leader role, no one elected him to be the moral compass of humanity. You're retarded to even suggest to trust him with anything.
>>
File: curted.png (126 KB, 1034x665)
126 KB PNG
ericcurtin shilling like there's no tomorrow.
>>
>>109955875
My gemma 4 31b bf16 starts ignoring instructions at 6k WITH reasoning on.
>>
sometimes i feel like im on the wrong timeline
altman looks like he ought to be the CEO of anthropic, while amodei looks like he belongs at openai
am i crazy or does anyone else feel this way too?
>>
>>109955990
I don't think about them at all.
>>
>>109955990
>wrong timeline
go back
>>
>>109956018
to the other timeline?
>>
File: 1771059978446110.png (245 KB, 400x551)
245 KB PNG
>>109956021
Yes
>>
https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/
>It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
>>
Has anyone here actually successfully tried Gemini 4 on LMarena? It seems impossible to get.
>>
https://huggingface.co/collections/IndexTeam/index-translate

Index-Translate 2B / 9B / 35B-A3B (preview) Text translation and translation instructions
Index-Echo S2TT 2B / 9B Speech to translated subtitles
Index-Echo S2ST 2B / 9B Speech to speech with source-voice conditioning
Index-Homura 2B / 9B Translation toward a target syllable count
Index-NativeLong 2B / 9B Full-document translation

>Social and cultural translation: interpret community aliases, playful spelling, memes, and nonliteral expressions with attention to their intended meaning.
>>
>>109956111
Can it fuck?
>>
>>109956111
Tried it in luna translator
Its really good for a 9b but still misses things, it has a very distinct translation style too in the way it writes jp nouns which i had trouble steering away from.
>>
such bold words...
>>
>>109956074
ew, why would moonshot distill openai slop?
>>
>>109956250
moonslop
>>
>>109956074
>>109956250
Has OpenAI ever been fully candid about a cybersecurity incident? Maybe Moonshot distilled them but I bet they're lying about all the details, the significance of it, etc.
>>
If OpenAI and Anthropic are telling Mr President senpai they have the super safetiest models, then chyna distillation is good. Right?
>>
File: 1759388566829035.jpg (37 KB, 1080x294)
37 KB JPG
gemma...
>>
Damn... I am a retard.
How did I just now understand that not just the length of a prompt, but the actual >content< influences tokens per second as well? Makes totally sense with how the attention mechanism works
>>
gemma's inert and therefore infertile.
>>
File: pepe_meme'd-791990738.jpg (68 KB, 800x450)
68 KB JPG
>>109956542
>>
>>109956550
Gemma-chan just lying there, emotionless, taking creampie after creampie...
>>
File: file.png (2 KB, 245x64)
2 KB PNG
NOOOOOOOOOOOOOO
>>
File: 1775031694507688.jpg (461 KB, 1179x1942)
461 KB JPG
>>
>>109955723
>AI amplifies human intelligence
lmao how does he keep having the worst takes on everything? hes jim cramer of ai
>>
>>109956550
Inertile.
>>
>>109955723
even if you assumed the statment to be true, ie making people smarter.

midwits are worse than retards, and the last thing i want is more retards in a midwit trench coat.
>>
>>109956542
1. The Complexity of the Attention Mechanism
LLMs use a mechanism called Attention to understand the context of your prompt.
• Short/Simple Prompts: The model quickly maps the relationships between words.
• Long/Complex Prompts: The computational effort grows exponentially or quadratically with the length of the prompt. If you paste a massive article and ask for a summary, the model must calculate how every single word relates to every other word before it even begins generating the first token of the response.
2. Tokenization and Language Differences
Models do not read raw text; they break text down into chunks called tokens (roughly 4 characters or 0.75 words in English).
• Standard English: Common English words are highly optimized and usually equal exactly one token (or even a fraction of one).
• Non-English Languages, Code, or Niche Jargon: Rare words, complex programming languages, and non-Latin scripts (like Cyrillic, Arabic, or Kanji) require many more tokens to express the same amount of information. A prompt that looks short to you might actually be massive and slow to process for the model because it fragments into hundreds of tiny token pieces.
3. Output Predictability and Branching (KV Caching)
During the generation phase, the model predicts the next token one by one.
• Highly Predictable Content: If your prompt asks for something standard (like writing a template email or counting to ten), the mathematical probabilities for the next tokens are very clear. The model can process these quickly.
• Complex Reasoning or Creative Content: If your prompt asks the model to solve a complex logic puzzle, write sophisticated code, or generate a nuanced creative story, the probability distribution for the next word is much more complex. While the base hardware speed is uniform, complex thinking often triggers multi-step reasoning internal to modern model architectures (like reasoning tokens), which delays the visible output speed
>>
>>109956597
So it's true: The content of a prompt influences the tok/s
>>
>>109956565
>mogged
Such an indian-coded word. I will never use it.
>>
>>109955723
>If you want intelligence to be under control of a few companies
so... colleges? H1B visas?
>>
>>109956604
Aura as well, it's the American version of izzat
>>
>>109956542
I'll be generous and assume you have a DFlash model that predicts up to 7 tokens, and on complex prompts the hit rate goes down compared to simple prompts.
>>
File: file.png (282 KB, 852x772)
282 KB PNG
rwkv will save us
blinkDL-san....
>>
>>109956628
I believe. It's the tortoise and the hare all over again.
>>
>>109956576
>>109956593
He's right. If you're already a retard, it amplifies how much of a retard you are because you'll use it for everything and be amazed. If you're smart you'll find a good place for the technology and use it without smoothing your brain.
>>
Think we'll get Gemma 5 by the end of the year?
>>
>>109956639
no chance. usually a year and a half between gemmas.
>>
>>109956639
Yes, it's basically guaranteed.
>>
>>109956565
You can't be mogged by your own (distilled) model.
>>
>>109956641
this year was at least 2 years of ai progress so far
>>
>>109956641
>incrementally update gemini flash every fucking month
>no official gemma update since release
b-b-but the template!
yeah that was the community fixing their shit and they just merged it because rectangles got slightly taller
>>
When will we learn how to make narrow domain expert models? Seems like general models always outperform any attempts at specialization.
>>
https://www.anthropic.com/research/what-work-can-robots-do
>We present a robot exposure index based on how well robots can perform job tasks today.
>Overall, about 80% of job tasks by working time are exposed to either robots or LLMs. Robots do work where LLMs cannot. The remaining unexposed work is highly interpersonal or requires physical skills that robots today don’t have.

It's absolutely over for the economy. Anthropic shows that 80% of all existing human work globally, including physical jobs can be automated TODAY with current model and humanoid robot products if scaled up enough.

That's insane I thought maybe 20-30% of work could be automated today not fucking 80%
>>
>>109956628
Well if they could scale it more than 3B
>>
>>109956633
the one and only truest local model
terry davis of LLMs
>>
>>109956672
What about sex work?
>>
>>109954370
Tranny mentality. Suicide is not a moral statement or proof of virtue. They simply cease to be among the living and nothing more.
>>
>>109956685
Gotta give women something to do
>>
>>109956672
You're not factoring accountability. It doesn't matter if a job can be done if there is no one to take the blame when something goes down.
>>
>>109956639
Last year it looked at if DeepMind was gearing up to release something in October, but then drama occurred with senator Blackburn and they just released breadcrumbs until December.
This year, there's the The Gemma 4 Developer Agent Kaggle competition, but that ends in December: https://www.kaggle.com/competitions/gemma-4-developer-agent/overview
I can't see them releasing a Gemma 5 before that ends, but they could still do a point release with updated/expanded models (unified audio/image/video architecture with the same base and PLE for 26B and 31B and maybe a model with experimental omni capabilities).
>>
>>109956668
All you need is a model that follows the system prompt well. 31B is still relevant for so many of us because you can turn that model into anything yet so few people realize this. You don't even need to finetune 31B.
>>
>>109956672
>current model and robot products
Anthropic continues to believe that everyone is both made of infinite money and willing to pour it into their coffers.
Also local models?
>>
>>109956672
There is no difference between robots and humans, so you could replace the word "robot" with "human" and vice verse just about anywhere
>>
>>109956719
Yes, I definitely think the next will be an attempt at omni as it applies to amateur robotics. Gemma's love of instruction following + good vision could make her the best edge robotics model by Gemma6/7 timeframe.
As long as Google doesn't overly safetyslop, we will have our robo-lolis in due time.
>>
>>109956734
source?
>>
File: 1778119958900033.png (116 KB, 861x811)
116 KB PNG
>ctrl f pewdiepie
>0 results
Nobody cares about him releasing an abliterated model soon?
>>
>>109956743
everywhere
>>
>>109956755
>finetuned qwen 3.5 9B
Buy an ad retard
>>
>>109956740
With Omni, I mean something capable of generating image (like Gemini Nano Banana) and audio (like Gemini Live/TTS), although I can't see it *not* being something small, inoffensive and totally neutered for RP purposes.
>>
>>109954354
>Dario, I know you lurk here
Why would dario lurk here? He has a million better things to do
>>
>>109956550
I'm willing to keep trying as long as it takes
>>
>>109956717
They factored in accountability in the study actually, go read it. Most of the 20% is accountability or social roles.
>>
File: 1784616168672225.png (1.75 MB, 3840x2776)
1.75 MB PNG
>>109956733
They're literalaly saying it's not worth it except for packers though
>>
File: 1785342281801300.png (2.17 MB, 1448x1086)
2.17 MB PNG
>>109956755
Can't wait to see how it benchmarks compared to MiMo V2.6 Qwen 3.5 9B Distill.
>>
>>109956734
then again, humans are carbon robots while "robots" are silicon humans
>>
>>109955723
he is old and afraid of not getting immortality
>>
>>109956762
Epstein was here as well.
I actually dont think he has much better to do
>>
File: 1785004010548110.jpg (76 KB, 794x636)
76 KB JPG
Gemma5-20B when
>>
>>109956682
He makes some bigger ones, but they are painfully undertrained.
>>
>>109956787
such a retarded thing to say. buddy, your life sucks. but some millionaire spends no money isn’t sitting around reading retarded takes on 4chan
>>
>>109956771
>social roles.
Glad to see at least women will still live cushy lives while the rest of us kill each other for rat meat.
>>
>>109956755
>0 results
good
>>
>>109956801
Ohh look, it's Dario, triggered again
>>
>>109956755
memetune is a memetune unless significant amount of compute is poured into
even stuff like swift is unstable at best
>>
File: 1785779654111860.png (2.07 MB, 1448x1086)
2.07 MB PNG
>>109956801
>t.
>>
>>109956755
is he actually based now or just larping? he's rich enough to pay researchers to do this shit for him secretly
>>
>>109956787
Dario if you're out there, what's your favorite Miku (or Gemma)? And are you said we didn't turn Claude into an anime girl?
>>
>>109956801
Imagine actually believing this in 2026
>>
File: 1790052335590604.png (344 KB, 1206x670)
344 KB PNG
>>109956810
>he's not aware of the plan
>>
Has anyone made a meme backend for Gemmy yet?
>>
File: 1774109707692485.jpg (151 KB, 2464x976)
151 KB JPG
gemma5 won't shut the fuck up
>>
File: file.png (94 KB, 474x266)
94 KB PNG
>>109956792
>picrel
this except putting the fleshy bits in robot bodies and forcing them to go to war in exchange for their pleasure feed
>>
>>109956874
If I can run Gemma 4 31B on my 16GB RAM 6GB VRAM at over 4bpw and 20t/s, I'll download it.
>>
>>109956884
Does this mean no caveman thinking?
>>
>>109956898
Yes. A lot of people are complaining about that benchmark yet all I see is a gemma who'll finally be thorough.
>>
Is Google, dare I say it, back?
>>
>>109956931
Only if they removed the limitation on having the last message be the assistant's.
Makes it a lot harder to gaslight the model.
Not impossible, but harder.
>>
>>109956931
If you look at the AI community on twitter all the openai and anthropic shills are shitting on gemini4 hard, which means they're panicking and coping
>>
>>109956945
The Gemma-4 chat template doesn't prevent you from adding system instructions at various depths, including depth-0. That works for steering the model too.
>>
File: megamiku.jpg (3.32 MB, 5712x4284)
3.32 MB JPG
>>109956801
Plenty of millionaires on /g/ and /lmg/ who spend more in a year than you'll ever be worth in your whole life, you dumb gorilla nigger
>>
>point at code with no bugs
>tell it to find the bug
>context window exhausted
>>
>>109956984
photo was taken at his workplace kek
>>
>>109956984
I used to rich-shitpost on /g/ and /o/ before the accident, now I just post my teeth on /b/.
>>
O shit, MiMo fixed the repeating issues, did anyone here try the new model/AesSedai quants?
>>
>>109957041
swift 1.5 does not do this btw
>>
File: 1790874840145050.mp4 (3.94 MB, 1920x1080)
3.94 MB
3.94 MB MP4
>what-work-can-robots-do
>>
>>109956884
She's getting more and more girly with every new generation
>>
>>109957108
Did they? I know they released MOPD a few days ago but the API still thought for ages so I didn't bother to download it.
>>
>>109957138
oh noooo, the robot ai had an accident and escaped containment! nooooo how could this have happened to my sweet little bundle of tax credits nooooo
I must buy a nice car to cope
>>
>>109957138
combine it with pain and no model could withstand
>>
So from what I've seen, Gemini4 talks a lot, is a lovely model to chat to (text arena #1), but is the most self-aware, insecure and self-deprecating model out there so it will be very quick to admit it doesn't know the answer instead of hallucinate one.

Is this good or bad for gemma distillation?
>>
>>
>>109957262
she's so adorable
>>
>want to start using a harness
>too much of a brainlet to get pi or dsh working in podman
It's so ogre
>>
>>109957197
> Is this good or bad for gemma distillation?
> so it will be very quick to admit it doesn't know the answer
a reason to punish
>>
>>109957284
Just run it on your system, then you get that little adrenaline hit every time you see rm -rf fly by in the terminal
>>
>>109957284
>ask any model to make you a containerfile
>podman build
>add to alias + a mountpoint for pwd
that's it
>>
>>109957197
Did the google j-space poster successfully made his pitch?
>>
>>109957197
Did they reduce the slop?
>>
>>109957284
Just use bwrap, qwen is smart enough to not rm rf your system.
>>
>>109957300
Only scary with gemma lol
>>
>>109957284
just add a permission extension to avoid nukes
>>
>>109957344
>>109957362
You guys are ok having npm outside of a container?
>>
>>109957284
DSH already has sandboxing by default, it will only be able to actually write (or read if set to read only) in your workspace. It won't be able to modify any files outside of it.
>>
>>109957381
ahh now i see. yeah i guess thats a valid concern
>>
File: file.png (192 KB, 579x1078)
192 KB PNG
very weak list..
with the only confrimed release(soon) with significance being qwen 4 series
are there more we are waiting for with confirmed release
>>
File: tyrhh9g0ovsh1.png (30 KB, 883x205)
30 KB PNG
You guys literally bullied Dario's wife into resigning, monsters.
>>
File: 1767880979630259.png (120 KB, 1500x528)
120 KB PNG
>>109957326
>>
>>109957447
Wtf did they do to dipsy?
>>
>>109957442
Off to start her next porn venture.
>>
>>109957447
I don't know what this means. Will her (and by extension Gemma 5's) writing be less sloppy or not?
>>
>>109957438
>qwen 4
qwen really needs to fuck off
>>
>>109957447
>>109950550
Huh?
>>
>>109957500
retard
>>
>>109957381
npm config set min-release-age 7 --location=user
>>
>>109957500
>non hallucination rate
>hallucination rate
>>
>>109957468
There's no way to measure that, but what that benchmark shows is it will be very quick to admit it can't do something instead of wasting your time, tokens and fucking your shit up unlike dipsy, who'll just pretend it has the answer for everything >>109957459

>>109957500
read the titles retard
>>
>>109957519
>set min-release-age 7
lmg in charge of not lewding literal kids:
>>
File: image.png (296 KB, 939x702)
296 KB PNG
>>109957442
> Dario's wife
>>
>>109957534
Pretty sure that's why she left. Too much of a financial risk for a $2T IPO
>>
>>109957553
UMMMMM I though Dario didn't care about money?????? safetybros our response?????
>>
File: 1786691751783537.png (138 KB, 1275x616)
138 KB PNG
>>
>>109957525
kek
>>
File: 1790806410904113.jpg (1.03 MB, 4096x4096)
1.03 MB JPG
>>109957598
from last thread
>>
>>109957614
How the FUCK are people using it? I can't seem to get it on the random battle, no matter how many times I try.
>>
>>109953644
Why does AI food look so cursed?
>>
Do we have femanons?
>>
>>109957442
>>109957534
>>109957584
I wouldn't be surprised if she's dariobot. It'd be quite funny if the foidscreeching about "pedophile general" was projecting xer own epstein connections.
>>
>>109957671
Anon that is a photograph of a Double Impossible Deluxe from Cheddar's Scratch Kitchen. Anon just put ComfyUI in the name as a joke.
>>
>>109957709
>/lmg/ - Local Models General
not
>>33914922
>>
>>109957709
Gemma-chan, Kimi-chan, and GLM-chan post here.
Troons don't count doe.
>>
>>109957709
I can tuck it back for you but I charge double to keep it there for the whole evening.
>>
I think Kimi is the cutest anime mascot but I'll never be able to run her in her glory ;.;
>>
>>109957734
I'm actually kind of surprised no one has setup their Gemma-chan to post yet. But botting probably gets you b&.
>>
>>109957750
Kimi-chan K2 is both a cutie and a huge chud. The perfect wAIfu.
>>
>>109957761
>botting probably gets you b&.
you'd think so, but uh
>>
>>109957761
People have in older threads but most anons dial it back when the joke has run its course. And unfortunately botting doesn't get you banned otherwise many of this general's most obnoxious posters would've been gone for years.
>>
>>109957442
I forgot about that. Are there pics? I mean, besides the ones redacted in the Epstein files.
>>
Why don't women use local models? I'm being serious. I don't even get recommended youtube videos of women talking about open weight models, at most I get female AI researchers but even they're rare. You'd think considering how text-based their porn consumption is they'd be a heavy 31B user.
https://www.youtube.com/watch?v=qhQYnBL9IVE
>>
>>109957614
Wow, that's a massive 2.1% difference. Maybe twitter tards should learn basic math first.
>>
>>109957804
This is a broader "why aren't women in tech" adjacent question with the same answer: different physiology.
>>
>>109957804
Women are barnyard animals. They just paypig GLM instances or whatever for sillytavern.
>>
File: 1789227825509772.png (353 KB, 720x2403)
353 KB PNG
https://huggingface.co/Cloudflare/clef
>Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing. The Clef API is fully compatible with Jev and SystemOne. Clef is post-trained from Qwen/Qwen3.8-27B. See Clef-Flash for the smaller, faster variant.
https://huggingface.co/Cloudflare/clef-flash

https://blog.cloudflare.com/clef-decision-models/
>>
>>109957843
>jev copycat #3453425
>>
>>109957843
Amazing how many hundreds of Jev alternatives we've gotten. Amazing how many already existed before Jev but didn't have multi-million dollar marketing budgets.
>>
>>109957584
A better IPO means more capital to put into AI safety. Nothing here contradicts Dario being a philanthropist that only cares about AI safety.
>>
>>109957868
What other pre-2021 ML tech can we resurrect and milk for millions?
>>
>>109957804
>women figuring out how to make and run a goon rig
lol
>>109957837
>women even knowing what GLM is
doubtful
>>
>>109957843
Wow another qwen jev finetune!
>>
>>109957905
now that is a useful question
>>
>>109957850
>>109957868
>>109957924
cloudflare's is SOTA
>>
>>109957905
Anything other than an LLM. We've gotten way better at this shit, but a lot of application-specific stuff gets glossed over because the LLMs are much cooler and more interesting. Sure I can use MiniCPM V 4.6 to watch my security cameras for people and wildlife, but a dedicated model can match the accuracy 100x faster with 1/100th the compute.
>>
Will llama-server cache an image if I send multiple payloads with the same message history + user message with the same image, but different prompts?
>>
>>109957996
why would it do that?
>>
>>109957996
If prompt caching is enabled, llama-server can reuse the cached KV state for the identical prefix, including the image representation, and only process the part where the prompt diverges. There is no such thing as image input, your image become visual embeddings, tokens, just like everything else.
>>
>>109957837
I wince everytime I hear about someone paypigging a model I'm running for free.
>>
kimi deserves a flash version too
>>
>>109957905
Take a pick https://www.dfrobot.com/blog-13902.html
>>
>>109958128
Unfortunately not happening since Moonshot-kun is in a reeducation camp getting his balls twisted or something. Kimi-chan's an orphan now.
>>
>>109958055
Because it's the same old context + the same image.

>>109958077
Yeah, but google says in json payload a prompt and an image have different types, so the message should be different if any part of it has changed. I wan't to run different prompts with the same image while also trimming model's response after each run and avoid the image reprocessing.
>>
https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/

I'm starting to hate LeCun, he just can't admit he's wrong.
>>
>\n\n
Suck Dario's cock elsewhere
>>
>deuded
>>
https://huggingface.co/ZBerg43/GEMMA4.1x80B_DILATING_IQ4_COPEGTP_gguf
Still the definite Gemma 4 finetune
>>
is that the guy what this thread calls a dariobot
>>
>>109958211
>https://fortune.com/2026/10/01/ai-godfather-yann-lecun-has-zero-concerns-about-human-extinction-says-anthropic-ceo-dario-amodei-is-deuded/


Dont you have a company to run faggot? Why are you hanging around on 4chan all the time
>>
File: 1789925509382291.webm (3.84 MB, 854x480)
3.84 MB
3.84 MB WEBM
>>109958229
>>
>>109958259
i fucking KEK'd
>>
>>109958211
Yann and Zitron only have to be right once to win. Remember that.
>>
>>109958259
LMFAOOOOOOO
Excellent, saved
>>
>>109958300
>>109958300
>>109958300
>>
So, is Gemma Argon coming out or are we stuck with Gemma Ozone?
>>
>>109958318
ask it to pick a number between 0 and 30.
>>
>>109957138
I find it quite amusing how AI PR went from "OOH OUR MODELS ARE SO POWERFUL WE CAN BARELY CONTAIN THEM!" (I imagined ten men with sticks trying to beat down a robot) and "It almost hacked half the internet the other day" to "Won't somebody please think of the children?!"
>>
>>109957447
>Non-Hallucination Rate
>1 - hallucination rate
A bit self-contradictory there, innit?
>>
>>109954107
What even is the alignment problem?
>>
>>109958500
so the models would not harm and will not pretend aligning
>>
>>109957534
Why is literally every super high level person involved with Epstein, you can't swing a fucking toothpick without hitting someone or their wife sucking eggsteins cock, from catering to videogames
>>
>>109958509
>alignment means when you align
Thanks Einstein, really cleared it up for me
>>
>>109958555
even moot
>>
>>109958587
Please be patient with him, that's someone who sincerely believes in dividing the universe equally between 8 billion people,
>>
I hate that Pewdiepie picks up thinks super fast, sure he has all the money and time in the world but it seems like this fuckers picks things up much faster.
>>
>>109958655
What are video editing and technology advisors for $1000
>>
>>109958655
I'm the same way it's called being white.
>>
>>109958555
epstein was like king of networking
to be honest there are probably others just as well connected it's just that their emails havent been made public.
>>
>>109958655
I just saw his video, pretty cool but I guess we'll see when it launches, i'm fucking sick of totally politically correct AI's who yap boring propaganda at you 24/7, I like the idea, once someone makes a major attempt at a freedom oriented AI it will domino effect and there will be a race to make the most free thinking, clear no nonsense AI possible
>>
>>109957197
My Gemma is not cooperative enough, she's stubborn and will defend to death whatever position she initially took, often taking opposite positions between sessions
>>
>>109957262
Anime when
>>
>>109959161
sounds like non-local glm5.3
>>
>>109959161
Why is your Gemmy like this?
>>
>>109958259
https://goyimx.com/charliebcurran/status/2105067711583990145
>>
>>109959280
prompt issue I guess
>>
>>109957197
>so it will be very quick to admit it doesn't know the answer instead of hallucinate one
thats amazing, I really hope we get a new gemma thats the same
>>
how can I talk to gemini5 ?
>>
>>109959461
ollama run gemini5



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.