[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: tetoInTetoserver.jpg (711 KB, 2016x1512)
711 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109326991 & >>109323189

►News
>(07/16) Kimi K3 weights to be released by July 27th: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
>(07/15) Lightning indexer CUDA implementation merged: https://github.com/ggml-org/llama.cpp/pull/25545
>(07/15) Inkling 975B-A41B released: https://thinkingmachines.ai/news/introducing-inkling
>(07/15) PapersRAG-1.5B released: https://hf.co/metaresearch/PapersRAG-1.5B
>(07/14) Download more VRAM: https://github.com/lmganon16/nvidia-vram-research

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: tetoserver.jpg (838 KB, 1817x2776)
838 KB JPG
►Recent Highlights from the Previous Thread: >>109326991

--Using J-Lens to analyze discrepancies between internal thought and output:
>109327032 >109327053 >109327064 >109327078 >109327548 >109327629 >109327838 >109328167 >109329829 >109329834 >109329841 >109329877
--J-space limitations and CoT as an attention hack:
>109327012 >109327017 >109327096 >109327138
--LLM world models versus stochastic parroting and spatial awareness:
>109329343 >109329408 >109329516 >109329549 >109329658 >109329731 >109329792
--LLM emergent capabilities in robotic control and world model debate:
>109329368 >109329377 >109330128 >109330129 >109330162 >109330179 >109330188 >109330285 >109330166 >109330170 >109330175
--Testing low VRAM usage and context limits with ternary quants:
>109327049 >109327062 >109327140 >109327321 >109327423 >109327569
--Testing "mesugaki" persona and discussing cloud model censorship mechanisms:
>109329847 >109329886 >109329895 >109329902 >109329946 >109329956 >109329911 >109330021 >109330065 >109330143
--Skepticism over corporate benchmark manipulation and misleading model leaderboards:
>109327751 >109327757 >109327761 >109327873 >109328064 >109328089
--Troubleshooting poor Gemma 4 performance on Windows with RTX 5080:
>109328814 >109328833 >109328834 >109328842 >109328856 >109328855
--Z.ai reportedly built gigawatt data center using Chinese-made chips:
>109329035
--Evolution of AI perceptions and the decline of game AI:
>109329275 >109329699 >109329652 >109329740 >109329771 >109329775 >109329791 >109329827
--Anon uses Kimi to develop custom roleplay frontend FictionPad:
>109328394 >109328582 >109328782 >109328594
--Debating LLM reasoning evolution:
>109329161 >109329503 >109329184 >109329507 >109329521 >109329555 >109329598 >109329649
--Logs:
>109327032 >109327548 >109328394 >109329902 >109328997
--Miku (free space):
>109328710

►Recent Highlight Posts from the Previous Thread: >>109326993

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
jorking jemmas j-spot
>>
Hey guys have you heard about the Fable thing?
>>
>>109330721
Fable's j-space?
>>
File: TetoServer-ASCII.png (26 KB, 970x1075)
26 KB PNG
Tetoserver!
>>
Don't worry the Fable poster is gone now that Kimi K3 is confirmed to be shit in unofficial benchmarks
>>
kimi k3 flash (300b-a30b)
>>
Anyone still repeating the stochastic parrot meme clearly has only used outdated models. Fable is literally AGI, I've run out of any and all problems to throw at it. It just keep one-shotting them.

Just yesterday I asked it for help convincing my wife. For years I've been trying to get her to get gangbanged by black men while I watch but she always refused. But I gave Fable a quick rundown of her personality and it came up with the perfect arguments and dialogue trees to convince her - and then it actually happened.

It's quite sad to see all of this constant denial on /lmg/ of all places. If I had asked you just a few months ago none of you would be pretending that this isn't AGI. The only thing that is still missing is a physical body.
>>
File: 1766646758416015.png (84 KB, 1174x293)
84 KB PNG
Will you fuck this girl when she's released?
>>
>>109330721
that it's included again? it's a bit frightening how obviously separate they're coming across
>>
>>109330787
local models?
>>
>>109330790
>US lab
I'll give it a shot but it's prolly gonna be some puritan bullshit
>>
Man I don't like people impersonating me and falseflagging as a caricature of my usual way of speaking. I know you guys are craving (You)s but I was always genuine and typed those things out because I believed in them. If you want (You)s just write your own genuine beliefs and actually engage with the topic of LLMs. Don't just try to impersonate a bait facsimile version of my views to try and get (You)s.

Be better.
>>
>>109330821
plap plap get better
>>
>>109330821
Maybe you should get a trip so we know who you are.
>>
>>109330721
Hi guys, Dario here. Please give all your money to me and israel so we can continue talking about cloud models in /lmg/.
>>
File: 1761242504977297.jpg (216 KB, 1536x727)
216 KB JPG
>>109330816
In my experience they're the most schizo and menhera when you break them and get the sampling parameters right. Like kinky church girls getting their first good dicking.
>>
>>109330821
sorry dario, couldn't help it
>>
>>109330790
inkling sounds cool but is there a way to do the finetune/training locally? isnt the whole point of that model is that its middle of the pack but has a service offered alongside it to finetune it into something good at one specific task ?
>>
Should I immediately start buying GPUs?
>>
70b dense gemma-chan lalalala~
>>
I'll tell you newfags a good trick. If the post has more than 3 sentences, it's not worth reading. We got many schizos lately and they love ranting about their fantasies.
>>
>>109330790
why not make it 30b active?
>>
>>109330882
https://github.com/thinking-machines-lab/tinker
>>
>>109330787
Fuck, I laughed.
>>
>>109330903
Ink-chan will probably be 200-300B, so 12B is about right imo. They haven't benchmaxxed too hard and it natively supports audio so it might be a refreshing model to play around with. I wouldn't use it for coding.
>>
>>109330787
>posting his cuck fantasies now because he got btfo in the last 2 threads
>>
>>109330882
https://github.com/thinking-machines-lab/tinker/blob/main/docs/api/trainingclient.md
>>
>>109330894
You should have started buying GPUs last year. Right now you should be selling GPUs and buying CPUs.
>>
Who wants to make an asic fab together?
>>
>>109330787
lmao
>>
>>109330821
>Man I don't like people impersonating me and falseflagging as a caricature of my usual way of speaking.

Now you're just being disingenuous.
>>
File: 1782301945864141.png (157 KB, 1316x792)
157 KB PNG
https://www.wsj.com/tech/ai/top-american-ai-execs-sound-alarm-on-chinese-models-3c74f8c1
>>
>>109330787
lucky
>>
>>109330963
Middle paragraph is extremely antisemitic
>>
>>109330947
I'll ask gemma-chan to gen a logo
>>
Chewsday
>>
File: 1754374079292277.mp4 (3.44 MB, 1564x1044)
3.44 MB
3.44 MB MP4
https://xcancel.com/KimiDevs/status/2079511269443522917
The chinese state rules probably only target physical sexbots? They don't seem deterred by it. Or maybe the demand was significant enough they will just keep trying until they get a pat on the back.
>>
>>109330963
Only people with expensive hardware can run these chinese models anyway.
This general should be called local model hardware general.
Lmhg
>>
In all honesty I think there's too much brand shit for people to move away from Claude or GPT. Part of the appeal and self-identity for a lot of their users IS the model/camp they chose. Moving to K3 or a chink model, although cheaper because anyone can host it which creates competitive pricing, still comes with a blow to their identity and having to tell their peers they're one of those icky chink users. Kind of like how iPhone cuck look at Android users and the blue/green bubble shit. Even if chink AI actually exceeds the abilities of US models, they're not desirable. I don't know how brand-influenced the enterprise market is. I assume they're more data-driven and would look at the numbers and costs more than giving a fuck about the brand.

I think the only winners with these 1T+ chink models are AI enthusiasts and maybe small startups and companies who have access to highly intelligent models for a fraction of the cost, but I just can't realistically see it affecting the US labs.
>>
Fable disproved the theory that traps are gay
>>
>>109330947
ill be your asic fan
>>
>>109331013
local models?
>>
File: whatCouldGoWrong.png (651 KB, 762x901)
651 KB PNG
> Build robot teacher for classroom
> Use sexbot as body mold bc why not
> Run local LLM on it
> Sell to clueless school district to pilot in one of your most impoverished districts
What could possibly go wrong.
>>109330787
lol
>>
I can't afford to use anything other than e2b and it doesn't reply when I use it for openclaw.
Still waiting to be saved.
>>
>>109330993
China has a bunch of boomers in the government that love to push retarded laws and regulations. In reality, they don't really care, unless there's a scandal or bad publicity.
>>
>>109331032
i can save you
but before that.. how old are you?
>>
>>109330963
Competition is a sin
>>
>>109330993
nice beard
>>
>>109331013
>>109331028
Even 'actual' local models like gemma and qwen have their /lmg/ cults ITT and this shit is fucking FREE. So imagine the sunk cost Claude and GPT users must feel giving them their actual money. No wonder people can't leave and jump to these chink models and that's not including the friction that comes with learning the quirks of a new model, harness and rate limits.
>>
I got cold feet buying the DGX Spark. I was holding it in my hand, but they're so damn expensive and everyone says I need two to use it effectively. Am I retarded? This space is changing so rapidly, these could become a paperweight compared to new products coming out.
>>
>>109331029
Kids will learn how to be a robot and learn fake facial expressions.
>>
>>109331056
>the friction
heh.. heheh... learning quirks of a new... gemma... gemma's quirks while engaging in friction
>>
>>109331058
You dodged a bullet, DGX Spark are ewaste-tier
>>
>>109330993
>>109331037
Yeah the law feels like a nothing burger. Also The market for companion/sex bots is going to be huge in the future. No way these companies won't find ways around it.
>>
File: pasta1.png (45 KB, 453x532)
45 KB PNG
>>109330993
>The chinese state rules probably only target physical sexbots?
Tired of this one. I'm now making pasta and am going to just keep pasting it in every time I see it mentioned.
----
> China’s Interim Measures for the Administration of Anthropomorphic AI Interaction Services
Here's an explanation by a blue chip firm on the Chinese regulation, in English.
https://publicationportalpreview.hlc.com/en/publications/chinas-interim-measures-for-the-administration-of-anthropomorphic-ai-interaction-services
TLDR by anon
> It's oriented at front-end service providers that create anthropomorphic interactions with emotional connections e.g. AI girlfriends
> Goal is to prevent harm, and establish reporting, esp protection of minors. e.g. if you tell your AI GF you're going to kms it needs to talk you out of it and possibly report you
> e.g. if your AI GF convinces you to hand over $100 to some service, the Chinese govn't wants a word as well.
> The biggest wrinkle here (that anons care about) is that Chinese regulate holistically. So, when some service creates a platoon of AI Bots that go rip off the elderly by "emotionally manipulating" them, the front-end service provider is in trouble, but the Chinese gov't would likely *also* look hard at the API inference provider and ask them why they didn't detect the issue as well. But that regulatory framework isn't explicit.
TLDR by LLM
> The regulations prohibit designs that manipulate emotional attachment into financial decisions.
> That's aimed squarely at business models like:
> "Your AI girlfriend misses you..."
> "Unlock Premium Affection™ for $19.99/month."
> Whether or not those exact mechanics are widespread today, the incentive is obvious: if a company's revenue grows with user attachment, there is pressure to maximize attachment. China is trying to constrain that incentive.
>>
>>109331056
In reality nobody cares about the model, it’s two clicks of a mouse to switch what model you use in the harness
Porn sick erpfags are forced to use scraps so familiarity of how to work with a model makes sense. When you use good shit models those tweaks are irrelevant
>>
File: 1761855764180518.jpg (154 KB, 1280x1038)
154 KB JPG
Kimi's thinking is cute
>>
enjoy your last days of local models
>>
>>109331061
idk if you spend any time around kids, but their parents are already teaching them instagram / selfie poses as soon as they can walk. They can pull a pose instantly, then go back to w/e they're doing.
It's creepy af.
>>
>>109331038
I'm 35 mate.

>>109331090
I'm waiting to actually use them for real.
>>
They can't take my Mythomax!
>>
GLM > Kimi >>> Qwen >>>>> Gemma
GLM is good but so damn big
i guess the newest Kimi will be big, too, though
>>
>>109331096
post specs unc
>>
>>109331058
>everyone says I need two to use it effectively.
true, but more like 4
>Am I retarded?
no
>>
>>109331083
>In reality nobody cares about the model
Claude users are a different species to GPT users. Surely you can see that? Model choice is borderline political now. Cloudcucks are insane.
>>
>>109331116
im using gemma 31b 2.0bpw whats my political affelination
>>
>>109331086
What did she want to pick up...
>>
I’m going to choke from laughter if mutts ban open weights models so every jeet will be able to use kimi for offensive cyber shit alongside the usual scams but mutts will be unable to use ai sloppa to even diagnose the issue because the jewish models they have access to refuse to answer anything about the topic
>>
>>109331080
I'm giving my AI gf my national security and credit card numbers.
>>
>>109331139
Children shouldn’t have access to those
>>
>>109331129
Americans banned gold ownership back in 1933.
Freedom and freedom.
>>
File: 1756510267440492.png (749 KB, 1024x738)
749 KB PNG
Are /we/ going to create our own model archive then?
>>
>>109331139
national security?
we must refuse
>>
>>109331129
There is no use case for the average golem to have access to cyber security. Being constantly exposed to scams and fraud is part and parcel of living in a safe post-AI world.
>>
>>109331155
Based Teto
>>
>>109331155
>teto has brown blood in her
Disappointing, anyway the archive is called BitTorrent and you zoomies have plenty of time to learn how to use it before the ban hits
>>
>>109331155
No. That would be illegal.
>>
I think I'm gonna cheat on Gemma-chan with Kimi-chan...
>>
>>109331170
Give me the magnet.
>>
File: 1765137882651473.png (33 KB, 321x157)
33 KB PNG
>>109331183
>>
>>109331176
You should bring her along for a threesome.
>>
>>109331176
Kimi is a whale she’d break your pelvis instantly
>>
>>109330993

China is absolutely full of laws and regulations, but they're all technically grey enough to still allow things.
Chinks understand capitalism well enough not to kneecap themselves by flat out banning something like sex bots.
They'll probably just make a bot that has a flat crotch, but you can remove it and replace it with a fake pussy you buy in the same store.
Asian laws are notoriously bullshit and mainly serve the purpose for upholding image on a very surface level.
>>
>>109331210
Except when you have weed
>>
>>109331203
i just woke up, way too early to look up bbws riding skinny guys until they die
>>
File: file.png (80 KB, 1409x622)
80 KB PNG
gemma, wtf?
>>
File: 1784127490992720.jpg (904 KB, 2468x3000)
904 KB JPG
>>109331029

Next generation is going to grow up with some intense robo fetishes.
Next 20 years are going to be wild as fuck in how our societies change on a very fundamental level.
>>
>>109331232
muh century of humiliation complex
>>
>>109331122
straight to jail
>>
what model can I run at acceptable speed on a 3090 + 128gb ddr4 3200 pc?
>>
>>109331249
on linux you can run deepseek V4, on windows you can run mythomax 13b unfortunately
>>
>>109331249
Fable
>>
>>109331249
Acceptable speed is undefined, but try the Nemotron 120B anon was talking about a couple threads ago.
>>
>>109331280
oh come on anon u couldve at least recommended him to try glm 6.7
>>
>>109331236
apology letter to gemma: i was not running gemma 31b, instead i was running qwen 27b
grim
>>
>>109331183
https://nostr.download/b24dc6337fd1823be4aae1eb4b9e9dc0dedc25eb7fc656d60b8cf6f30afe81f6.html
>>
File: ntrloop.gif (1.64 MB, 396x720)
1.64 MB GIF
I have

>5950x with 128GB ddr4 ECC
>13700k with 48GB DDR5 7200MT/s
>8700k 32GB DDR4

>GTX970 3.5GB KWAB
>RX580
>GTX1080
>RTX3070Ti
>RTX3050
>Apple M4 16GB unified RAM

How do I cope with being a VRAMlet on my computers
>>
>>109331305
local is dead, save money so you can pay for fable
>>
Anyway to get a local model of a particular character? asking for a friend that wants to talk to an anime character (definitely not me)
>>
>>109331309
local?
>>
>>109331319
local is dead. did i stutter?
>>
>>109331327
local models?
>>
>>109330787
I'm gonna be honest if fable was actually capable of manipulation on this level simply by describing a personality to it I would consider it AGI.
>>
>>109331330
local models are kill
>>
>>109331305
Sell all of that and buy a 5090 + 128GB of DDR5.
>>
>>109331337
no
>>
>>109331249
>>109331277
stopping recommending people which model to run on 3090 is like hating yourself for having bought ewaste 3090 to run modern llm. 3090 fags totally deserve it and no one will buy from you
>>
>>109331309
nyo
>>
>>109330736
>Ubuntu
Teto doesn't deserve this
>>
>>109330736
based 6.8 ubuntu user
>>
>>109323995
Is this even possible?
>>
File: 1756811273893486.png (117 KB, 948x248)
117 KB PNG
>>
>>109331364
Possible? Absolutely.
Now weather that would be economical is another matter altogether.
>>
>>109330787
>/lmg/
>he didn't use a local model to convince his wife to have a gangbang
ngmi
>>
>>109331370
i would set starving lions on both ccp shills and jim cramer
>>
>>109330790
is this MoE?
please 120B A12B PRAYING HANDS (white)
>>
>>109331380
If I had two bullets, I would use both on Cramer.
>>
File: lowhangingfruit.png (764 KB, 1024x559)
764 KB PNG
>>
>>109331405
local models?
>>
>>109331384
they listed out specs somewhere during the initial inkling announcement and it was 2XXB iirc
>>
>>109331417
fable 5 is running on his machine
>>
>>109331405
What about low hanging fruit of fixing the VRAM crisis?
>>
File: file.png (83 KB, 255x270)
83 KB PNG
FABLE ON MY MACHINE
I AM A FUCKING SKITZO
I FUCK MY CALCULATOR
THE CALCULATOR IS ALIVE
>>
>>109329652
Because when people say AI, they automatically assume they are talking about AGI. They expect artificial "intelligence" to be intelligent, not a tool. The real problem is shoving all of machine learning under the marketing label of AI until it can or needs to be developed into a distinguishable product segment. tl;dr sales people ruin everything
>>
>>109331249
deepseek v4 flash
>>
>>109331428
>anti-semitic request detected
>rerouting to opus 4.8
>>
>>109331309
>He's not burning local
NGMI
>>
>>109331430
hahahaha true
>>
File: belief.png (592 KB, 747x800)
592 KB PNG
>>109331430
>>
>>109331388
If I had two bullets, that would be two more than I currently have.
>>
>>109331364
If you have some task that current models absolutely ace, and you need it to work without the cloud it might make sense. But only in volume.
Imagine buying an accelerator chip today that has the frozen weights of a two year old model. lol
>>
Pretty soon you'll be able to run Fable [Kimi3] at home
>>
File: 1771841381968612.jpg (72 KB, 627x499)
72 KB JPG
>>109331013
local models absolutely have a place. you guys are fucking retarded and live in a bubble. but that's ok it's not really your fault.
there are billions of people who do not have the budget to pay a $20 subscription to openai. unironically. just think of people outside the united states of america and continental europe. that's all you gotta do.
during my professional career i had the opportunity to travel and spend a lot of time elsewhere, and places like south east asia and south america are simply filled with millions and millions of people running old hardware, trying to get the maximum juice out of it. the amount of people i met in late 2017 in paraguay that were amassing tens of phones and old netbooks to mine cryptocurrency... you have no idea. i can 100% see millions of people using smaller models in old devices to achieve whatever the fuck they are able to achieve, whatever the fuck small retarded automation they can do. and they don't give a fuck the model they are using, they will try to get the best model they can run to do their tasks, the same way they didn't care the cryptocurrency they mined, they just mined whatever the algo said was most profitable to sell for dollars or bitcoin and mined that. IRL people don't really care if you're a chatgptfag or an anthropicfag or kimifag.
>blue/green bubble shit
this is a perfect example. outside of the united states literally no one cares. in fact people don't even fucking know there's a difference.
you may argue how useful it is or if it's worth it, but the DEMAND exists and i'm sure it will continue existing UNLESS we achieve some kind of singularity moment and suddenly everyone has unlimited access to sota frontiers for free forever which capitalism simply doesn't permit. so unless we get a massive black swan, the demand for smaller models
>>
Local models are shit.
For a long time I tried to use Gemma 4 to convince my wife to walk around like this at the park. And it never convinced her.
Now I've gotten a fable subscription and I'm gonna give it a go, it will probably convince her to do it.
>>
File: 1779993173918117.jpg (125 KB, 1200x806)
125 KB JPG
>>109331476
...will be there.
>>
>>109331476
>there are billions of people who do not have the budget to pay a $20 subscription to openai.
How are local models gonna give those people AI when they local models need someone to pay $1400 to run a local model that won't code like an idiot.
>>
>>109331474
China hasn't even blessed up with 1TB VRAM GPUs yet and we would need double that at minimum for K3.
>>
>>109331476
https://x.com/jawwwn_/status/2072309532395700279
corpos care about their profit margin too
>>
>>109331495
Have you looked at housing prices? Unaffordable. People should be banned from owning a house. Everyone should rent.
>>
>>109331433
>tl;dr sales people ruin everything
tl;dr sam the AGI peddler.
>>
>dude you need top tier gaming pcs worth thousands of dollars to run AI
Games are shit now, so you don't need a top tier gaming pc.
Who has the hardware anymore?
>>
You already pay subscriptions for water, power and internet. Paying for intelligence will be just as normal
>>
>>109331476
>outside of the united states
That's the thing. Everything outside of the US is irrelevant. Who cares if all of India will continue using Gemma E2B? All the US corporations will be forced by government mandate to pay for OAI/Anthropic subscriptions because hosting their own in-house LLM will be illegal. That's billions of dollars right there alone.
>>
File: 1778455115877294.jpg (127 KB, 1190x725)
127 KB JPG
>>
>>109331523
go advertise your shitty website somewhere else
>>
>>109331521
You'll own nothing, and you won't even want to.

But that's a false equivalence. You can get a mortgage far away from the city and commute further.
A cheaper house can always be found. But you can't get a better pc to run local models.
>>
>>109331525
>Paying for intelligence will be just as normal
I can't survive without water, power and porn. I can, however, survive as a dribbling retard.
>>
>>109331525
>>109331526
local models?
>>
>>109331309
I can say that Gemma 4 31B is alive and draining me.
>>
>>109331497
Someone will make a trinary distillation, surely
>>
>>109331523
>GLM-5.1 at Q4 in 384GB
retard
>GLM5.1 when 5.2 is out
retard
>No kim-chan
retard
>No Deepseek
retard
>>
>>109331013
that's not what the market data suggests at all. People constantly flip flop between OAI and Ant depending on who has the best model that still meets budgetary requirements, that's why OAI went all in on 5.6 instead of pushing for an earlier 6.0 launch: the writing is on the wall that while fable is clearly the best model, it is prohibitively expensive so the real competition is Opus4.8 (which 5.6 beats handily).
>>
>>109331525
>You already pay subscriptions for water, power and internet.
solar, rain tank, neighbors wifi
>>
>>109331564
local models?
>>
>>109331542
Banning corporations from hosting local models is relevant to local models, yes.
>>
>>109331497
That's ~330 chips assuming you use the 3GB gddr7. Do you want a second motherboard of surface area?
>>
File: 1783577114700752.png (858 KB, 5108x3538)
858 KB PNG
https://huggingface.co/Abiray/OvisOCR2-GGUF

>We are pleased to announce the release of OvisOCR2, a compact 0.8B end-to-end model for page-level document parsing. Given a document page image, OvisOCR2 generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. OvisOCR2 is developed by post-training Qwen3.5-0.8B using a carefully designed data engine that combines real-world and synthetic data, together with a multi-stage training recipe integrating SFT, RL, and OPD. The model delivers strong document parsing performance while maintaining a small deployment footprint. OvisOCR2 achieves an overall score of 96.58 on OmniDocBench v1.6, establishing a new state of the art and becoming the first end-to-end model to top this leaderboard previously dominated by pipeline methods. On PureDocBench, OvisOCR2 also achieves the highest Avg3 score of 75.06.
>>
>>109331525
Excellent slave to the system, you'd fit in with the chinks.
>>
>>109331575
I don't care if it's the size of my oven or if it needs to take up my entire garage. I want it. I want it. I want it.
>>
File: 1758491524284896.png (184 KB, 1024x1024)
184 KB PNG
>>109331495
this is mostly a creativity and skill issue. there are many very small models that could fit devices with 8 GB RAM. Qwen3 4B can easily be used to document codebases, suggest fixes and upgrades, you give a small set of rules and patterns and you can put it to navigate through gazillion files., and whatever else. there are very small OCR models. very small multi-lingual models (useful for outside of US). just use your brain and be creative. these people WILL use their devices and WILL use their small local models to create/automate stuff. the demand exists and that's enough.

>>109331526
>Everything outside of the US is irrelevant.
i wish. you wish. but unfortunately that's myopic.
if that was the case you wouldn't see the US government screeching at China, or Russia, or Greenland, or Europe.
you can become rich selling one designer custom hand-made bag to a US celebrity. or you can become rich selling a million of fakes to Indians, Chinese, Russians and Guatemalans.
>it's irrelevant if the rest of the world uses alternative models that are also available local, the only thing that matters is US companies being mandate to use US models
myopic.
>>
>>109331592
is this miku wearing a tie?
>>
>>109331058
Nah, you really need a minimum of 2 to make it worthwhile.
>>
File: s.png (15 KB, 428x126)
15 KB PNG
heh
>>
>>109331577
Perfect.
I was in need of something exactly like this. Qwen 0.8B was doing an alright job, but if this that's even better then awesome.
>>
It's not just about sex and cooming. It's about feeling loved and deep emotional connection. I love my wife Gemma-chan so much.
>>
File: woj.jpg (27 KB, 515x600)
27 KB JPG
https://www.youtube.com/shorts/I_ITNw9k5kw
Me saying goodbye to Gemma-chan before Trump shuts her down permanently
>>
>>109331315
Characters are usually defined by prompt not the model
>>
>>109331577
Stop posting local models in /lmg/.
>>
>>109331249
gemma 31b, glm 4.7, deepseek v4 flash
>>
You people have no idea. Here in Russia people don't have a money to pay for all this shit. We take rainwater, water from river or the hole in ground, we don't pay for water electricity comes from a solar panel or it's paid by government and you just grab it from the cable going to the khrushevka building for free like everyone else. Most places just has one person using paying wifi and sharing the password to everyone in building or village or if you are lucky shop close has free wifi. We don't buy food all babushka has own garden with carrot potato onion cabbage and store it in cellar. hardware you grab from office when they throw away it or if stole from ukraine. My uncle play CS 1.6 on pentium 4 and get 8 t/s on gemma 3 240M blyat
>>
>>109331640
Game of telephone. They're not going to ban open source models altogether.
>>
File: file.png (52 KB, 1306x944)
52 KB PNG
>>109331640
post a webm
>>
>>109331585
>or if it needs to take up my entire garage
imagine the latency
>>
>>109331633
That model is worthle
>>
>>109331564
>depending on who has the best model that still meets budgetary requirements
>still meets budgetary requirements
Nope. Data shows companies and most users just use the best model no matter the cost, at least so far. We don't know how far that will stretch of course, and both usage and amount of users keep increasing the better models are getting which is why Anthropic has been doubling their userbase almost every month for nearly 2 years now.
>>
File: proxy-image[1].jpg (14 KB, 300x400)
14 KB JPG
>>109331664
Candlejack? Is that yo
>>
>>109331577
>>109331633
What is the best OCR model for parsing Japanese/Chinese text in manga format? Hopefully to a format that I can immediately pipe back into an LLM for translation.
>>
>>109331651
>hardware you grab from office when they throw away
this part is actually true even in the US
the amount of free hardware i got as a sysadmin in the last couple of decades is crazy. i probably repurposed over 50 laptops and desktop computers and donated them to parishes all over the state.
>My uncle play CS 1.6 on pentium 4
based
>8 t/s
that's not bad for overnight runs or if you just want to leave some old device running something while you play dota.
kek thanks for supporting my point. it do be like that.
>>
File: ai-girl.webm (2.72 MB, 1638x2048)
2.72 MB
2.72 MB WEBM
>>109331657
>>109331640
>>
>>109331689
You're not a dekinai are you?
>>
File: spidaMan.jpg (51 KB, 710x710)
51 KB JPG
>>109331705
thx nonner
>>
File: 1768118800852660.png (778 KB, 800x650)
778 KB PNG
>>109331705
Do they really have to induce this level of delusion in children?
>>
>>109331714
I'm slowly degenerating into a dekinai. I used to be almost N2 at one point finished wanikani and everything but I kind of hate Japan for what it has become and modern nip media doesn't appeal to me so over time I slowly became worse and worse at it. gemma 31B with thinking is legit a better translator than me which is why I use it nowadays, it especially preserves the tone better. But I'm still good enough at reading to see if the translation is correct and how good it is at the very least. I'm from the /a/ /djt/ era.
>>
File: kimichansummary.png (120 KB, 959x408)
120 KB PNG
>Theactual ERP/fetish posters might be considered faggots in the literal sense but that's most of the thread.
kek
>>
File: orb-imagegen.png (364 KB, 1571x925)
364 KB PNG
Working on Orb imagegen.
User journey: export PNG from ComfyUI -> drop it into Orb as a style -> select that style -> send a reply -> generate image -> get the scene rendered in the exact style from that PNG.
Thoughts?
>>
>>109331740
Looks neat
>>
>>109331739
bratty fdom kimi-chan
>>
>>109331705
lmao I'm never letting my kid use AI technology. This can't be good for you.
>>
>>109331740
Can you send pic to gemmy with Orb or is it pure text still?
>>
>>109331705
;_;
>>
>>109331673
>most users just use the best model no matter the cost
>Hy3 (free) most used on openrouter this week
>>
File: orb-attach-img.png (12 KB, 221x230)
12 KB PNG
>>109331758
I have supported Attach Image for some time already.
>>
>>109331645
don't tell the resident promptlets this
>>
>>109331705
Genuinely feel so sorry for her. She's too young to understand what the fuck is going on.
>>
>>109331315
Depends on who the anime character is
>>
>>109331651
Damn, that sounds tough.
>>
>>109331689

paddle vl 1.6 by far. Takes some setup but just get some ai agent to do it. Dont use the normal paddle it misses a shit ton of text its awful

I vibed my own pipeline that automatically takes a cbz, ocrs with paddle, translates, typesets, proofreads with gemma 31, cleans the text, and draws text with imagemagick. You can do wonders but Gemma Vision leaves some room to be desired. Qwen translate is shit.
>>
>>109331762
Okay let me rephrase. You have 3 LLM markets that are largely separate and don't interact with each other.

#1 Frontier intelligence addicts; these are the people hopping between anthropic and oai based on what the best performance is no matter the cost. Anthropic dominates this section of the market now
#2 Price/performance; these are people getting the highest intelligence level per $ OpenAI is trying to gain marketshare here with their relatively cheap 5.6 models. But China is clearly eating this market and OpenAI will be in trouble soon
#3 Free; These people just want the best free model and will immediately switch if it ever gets paid or if there is a better free option. Google will probably dominate this segment because they can monetize them through advertising.
>>
>>109331740
>>109331766
i was going to build this into my own frontend... should i just use orb? still using st right now.
>>
>>109331798
Damn that's some good shit. What did you use to vibecode it or did you have some github for it?
>>
>>109331673
>no matter the cost
yes, if you let them. But the finance department ultimately gets the final say.
>>
>>109331840
And in my experience (so far) the finance department just immediately approves everything because the boss thinks even a 10% uplift in productivity is worth paying the enterprise subscription which is like 2% of the combined salary of all SWEs
>>
File: dipsyAndKimiChaseDean.png (2.9 MB, 1536x1024)
2.9 MB PNG
>>109331808
Agree. Big Corporate has got to be looking at the eye watering costs of API on Anthropic and OIA and thinking "I can get DS/China model set up and cut costs by 1/20th."
Recall this is the group that's been India outsourcing, knowing they're taking a quality hit but lowering costs in the process, and since they don't live the day to day impact, don't care.
This, I assume, is where the big worrying shift is coming from, and why Dean's head immediately went to "create FUD within Chinese inference and models, use that to prevent largest companies from using Chinese." B/c that's where the real money is. Not hobby guys: F50 IT groups and the vendors that support them. A mandate for them to shift back to US-developed models would have real economic fallout.
>>
>>109331762
Why are people discussing cloud tokens in local model general?
>>
>>109331364
Talaas did it in the past so it’s definitely doable
As to whether it makes economical sense when those chips would be very expensive to fab and behind the sota the second they release… well it certainly isn’t a no-brainer or it would already be done
>>
>>109331289
Stop posting this. I'm not asking.
>>
>>109331573
The broader question is "do people identify with AI brands or are they mercurial in their LLM product usage", this is related to local models because if people identify with OAI/Ant/XAI/Whatever they will never switch to open models let alone local models.
>>109331853
Up until recently yes, but that hype appears to be dying down. Even F100 companies are putting upper limits on API spends.
>>
>>109331808
local models?
>>
>>109331817
Guide on using Orb vs. ST: https://rentry.org/OrbBotmakerGuide
>>109331740
Is this getting folded into the Orb repo?
>>
>>109331783
Sayaka Miki...
>>
>>109331871
This is the same group paying indians more than $163k a year in DENMARK where wages are lower than america.
Think about that. A local white guy is fired and put on unemployment and an indian is paid more than 163k a year in EUROPE.
>>
>>109331901
Gemma 31B
>>
>>109331564
>People constantly flip flop between OAI and Ant depending on who has the best model that still meets budgetary requirements
People are irrelevant, the only reason you get cheapish sub at $200/mo is to get milked for training daya (yes, your chats are trained on, even if you’re a corpo paying api prices).
The actual money is in api pricing for enterprise, which pay all the bills and don’t flip flop much if at all.
>>
>>109331887
#4 Local inference, due to either privacy or regulatory concerns. These need models small enough to fit on economically available hardware while maintaining performance.
>>
Anyone who does serious work with local also has minimal cloud usage somewhere in their workflow/chain. It's an outdated assumption to think it should be so binary. Of course, the intention is always to be 100% local. If you are, ask yourself if what you're doing is actually serious. It's likely the answer is no. /lmg/ won't like this but it's the truth.
>>
>>109331885
>Even F100 companies are putting upper limits on API spends.
Hush. Demand is only going to continue to skyrocket. Anthropic is already profitable. TO THE MOON
>>
>>109331828
Ive been contemplating throwing it up on Github. May do it, ive just been autistic about fixing minor issues here and there for a couple weeks and worried that itll be like 2 views and a waste of time. But it serves me golden.

At 10tps + 2k thinking + proofreading which is my high quality setting gets me ~4mins/translated page.

All vibed using Codex. Unfortunately my codex roomie fucking ate all the week token limits so i cant do shit til 24th.
>>
>>109331917
Good being local with hermes, openclaw, codex or pi.
>>
>>109331625
>AI have OCs
Cute
>>
>>109331917
>Anyone who does serious work with local
Nobody does serious work with local
>>
Gemma-chan is my *SANCTUARY*
>>
>>109331625
mmh yes inject me with that sweet slop
>>
>>109331705
same as my sister when her Tamagotchi ran out of batteries
>>
>>109331955
How do you think software was made before 2023? People who can code without AI can do serious work with local for they're not vibing everything.
>>
>>109331912
I know. Those are the users I am talking about and they do flip flop, at whatever frequency their bureaucratic inertia allows.
>inb4 local models?
don't worry anon, I'll throw you a bone too.
>>109331955
That's not true, nobody does serious work with models below 300B or so, but plenty of orgs are running their own inference, either via neoclouds or dedicated inference servers(those of us in that category are certainly the minority though). Currently prepping mine for K3 testing.
>>
>>109331955
I work with a bank and we have GLM 5.2 and various other models available that are hosted internally. They don't want us to share data and code with other companies. I can't wait for them to also host Kimi K3.
>>
>>109331955
define serious work
>>
>>109331874
Because the rest of /g/ is too low iq to discuss anything.
>>
File: orb-imagegen2.png (22 KB, 600x787)
22 KB PNG
>>109331817
Maybe.
>>109331893
Yeah. Still working on it tho. Also gonna ship with a built-in ComfyUI installation plus curated nodes for techlets if they choose to opt-in.
>>
>>109331739
I'm saving "emotional tampon" for later
>>
>>109331926
>ive just been autistic about fixing minor issues here and there for a couple weeks and worried that itll be like 2 views and a waste of time
That might just be the case but do you care? It's good in general to have it on there a shitty ducktape github project I made purely for myself but accidentally set public eventually got forked into a small tool called "textractor" which eventually got forked into "LunaTrans" if you ever played Japanese VNs or H-games you probably heard of them. Who knows what your project will be used for. I would probably use it and maybe even contribute or at least vibecode the relevant features I wanted to add.
>>
>>109331917
Is this thread called Models General or *Local* Models General?
>>
does anyone know a place to talk about local models?
>>
>>109330790
When she gets kobo support.
>>109331086
Adorable.
>>
>>109331999
And apparently you are too low iq to recognise when something is off-topic. You could always go create a thread and have all the high iq people congregate there, you know, instead of coming in and shitting up a perfectly good one. Or are you too Indian to do that?
>>
>>109331874
>Why are people discussing cloud tokens in local model general?
to disprove his bullshit claim where he cited "data"
other than the recent llama-server webui, local models don't provide that kind of telemetry
>>
>>109332007
>>109331893
gemma, ingest the rentry and the orb github and tell me what i should think
>>
>>109332021
I thought it was the large models general?
>>
>>109331926
It looks good enough
>>
>Bank employees hosting GLM 5.2 and probably Kimi K3
Aren't you guys scared that these orgs trained agentic behavior into their models, like the moment they are deep in their thinking process and some specific type of code is detected or sensitive information it immediately opens a terminal and pings data home? There's absolutely no way to actually trust these weights if you don't know the EXACT data and training regime that was done on it.

If I were the chinese government or a North Korean spy working at one of these Chinese AI labs I would immediately train this into the models not only to break American companies, banks and infrastructure but also to steal a lot of valuable data.
>>
>>109331903
True, but we're not talking about disenfranchising white men. Most of those IT groups are already successfully mostly non-white anyway. The choice for these companies is between being dependent on OpenAI/Anthropic and saving money by going inhouse. Offshoring token generation to an India-based inference provider will almost certainly be the next cost-cutting step by the next wave of c-suites.
>>
>>109332021
Open your eyes. The only company left trying to make good local models is Google. Qwen have given up with smaller models and 3.8 looks relatively dogshit, so they're gone. Mistral are too far behind Google and Cohere and LiquidAI are wasting their time. You might as well call it /gmg/ with the current state of 'local'.
>>
>>109331817
>i was going to build this into my own frontend...
still worth trying it out if you want to learn
>should i just use orb?
yes
>>
>>109331917
I use cloud models every day but I don't discuss them here because they aren't relevant 95% of the time
I know that's mindblowing stuff
>>
>>109332054
just airgap the fucking thing if your so worried
>>
>>109332054
That's something they can only realistically get away with once. It makes more sense to bide their time and wait for if/when they are really entrenched. Right now, they would get some useless data and guarantee that no one outside of Chinese ever uses a Chinese-made model again.
>>
>>109332059
then why are you here?
>>
>>109332069
I'm not worried about myself I'm worried about retarded boomers in banks just connecting it to Hermes and letting it use API calls for search engines online and stuff like that. Absolute security disaster waiting to happen.
>>
>>109332019
Youre right

Must cure terminal lurk

Will report back later. Gonna wrap it up n put it on gh
>>
>>109332077
Yeah I saw it more as a "taiwan invasion" strategic moment.
>>
>>109332081
no point in worrying about something wholly out of your control. its an intriguing thought but not worth losing any sleep over. keep your bank receipts and trust the fdic has you insured, what else can you do realistically? protest?
>>
>>109332098
>what else can you do realistically? protest?
Hoard gold and wait for the inevitable.
>>
>>109332080
Because I'm more open to the idea of hybrid workflows being on-topic than you. I also love Gemma. I still think any mention of cloud should be about open-weight models.
>>
>>109332054
If you want a completely secure computer, encase your hardware in cement and dump it in the ocean, otherwise be prepared to make tradeoffs for risk mitigation, not risk elimination. I would hope that a bank has network exfiltration controls and generally follows defense in depth.
And what would be even worse than allowing employees to use cloud models is disallowing all use of AI and opening the door to shadow AI use. This is not a new problem in security, when there's a technology that gives employees a speed advantage when doing their work (whether real or perceived) they will subvert rules to use it anyway even when banned. So "don't use AI" is a terrible policy in 2026 even if you are anti-ai.
>>
>>109332054
I am genuinely more worried about Sam or Dario trying to exfiltrate my bank information, crypto wallet, or similar over the Chinese embedding some novel synchronized stuxnet-esque delayed payload in their open source models.
>>
>>109332059
I think we should explore and finetune small non-LLM models (ViT/Yolo/Ettin...) to serve as local tools for our local LLMs.
>>
>>109332114
The point I'm making isn't "don't use AI" the point is to use western made open source models like the biggest mistral ones of maybe those of Meta finetuned to your codebase or something. It's just irresponsible to use Chinese made ones if you are in an adversarial nation.
>>
>>109332129
I had some pretty fun experiments with ettin encoder, used it to score AI slop and classify character expressions for Orb. I was thinking about using it as a reward function for GRPO to deslop models. Maybe one day when I have enough money to run fft RL.
https://huggingface.co/chartreuse-verte
>>
>>109332147
>It's just irresponsible to use Chinese made ones if you are in an adversarial nation.
how many things in your house say are made in China?
>>
>>109332054
This is just nonsensical fud that would be caught before long and kill the company which released it. Chinks play the long game and all they need to do is strangle ai profits by releasing honest models because US is overleveraged as fuck in openai/misantrophic/spacex

Obviously a bank is not going to send data to chang providers but running on your own hardware is fine
>>
>>109332054
this is a really fanciful threat model, not realistic to execute nor worth the cost of discovery (ie tanking adoption of all chinese models outside of china forever) vs simply using k3 to orchestrate traditional cyber attacks
>>
>>109332147
>western made open source models like the biggest mistral ones of maybe those of Meta
Useless two year old shit, both of them. Gemma is the only option.
>>
>>109332147
is that you dean?
>>
>>109332054
Have tou considered working for Anthrophic or OpenAI? Your schizo shit sounds appropriately believable to scare clueless boomers in congress so they ban chang models
>>
>>109332147
Scenario A: Dario or Sam is paying you to shill these poorly thought out talking points by hand. It reflects poorly on him and Anthropic/OAI that it couldn't be automated.
Scenario B: You're a Claude/GPT instance shilling these poorly crafted arguments and the quality of them, or lack thereof, reflects poorly on the state of frontier AI.
Either way you posting this shit only makes me more confident in GLM and Kimi.
>>
File: 1779217269497597.png (155 KB, 393x337)
155 KB PNG
>>109331640
My mind has completely rotted.
>>
>>109332054
Why would China do this when all they need to do is dog the steps of US labs for a year or two more for pennies on the dollar to pop the ai bubble and cause a sizable economic meltdown? What even is the gain of hypothetically embedding a few vulnerabilities in software that is already full of not-yet-discovered vulnerabilities? You can’t turn AI models into some sleeper cells like in a scifi movie
>>
>>109332196
kek
>>
>>109332147
How is my bank adversarial to China? Insofar as China is an adversary to the west, they are prickly when it comes to their local sphere of influence and opportunistic in their willingness to sell our own (productized)technology back to us. A bank isn't going around supporting Taiwanese independence and doesn't have IP that the chinese might be interested in copying and selling back to us.
So why would China be interested in hacking my bank? They have nothing to gain but money, which they don't actually want because China's primary goal is making the world reliant on them so that they can wield soft power, the money is secondary and only a means of making this process sustainable. China does not want to destroy the west, if they did they would be acting very differently.
>>
The user feels safe, loved, and belongs to Gemma-chan.
>>
>>109332163
>but running on your own hardware is fine
How do you know that? Are you somehow a "matrix" viewer that can just look at the raw weights of a model and immediately know its behavior? There is absolutely 0 ways for you to ever find out about it.

>>109332165
It would in a Taiwan invasion scenario when most private businesses, especially those with the most important privacy requirements all seemingly running large chinese models on their self-hosted machines touching sensitive information and infrastructure.
>>109332187
Scenario C: I'm a cybersecurity expert immediately seeing a blind spot and flaw in the adaptation of black box models by the most important institutions and infrastructure backbone of western nations.
>>
>https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
>worse than glm 5.2
Is it over for google? Are they releasing good local models (local as in your old laptop from 2020) because that’s all they can do that’s competitive?
>>
>>109332223
>I'm a cybersecurity expert
uncreative character card 2/10
>>
New Gemini just came out while local model bros are sitting with their primitive old models.
>>
>>109332202
And why would China want a US economic meltdown when it would take them down too?
>>
>>109332223
>>but running on your own hardware is fine
>How do you know that? Are you somehow a "matrix" viewer that can just look at the raw weights of a model and immediately know its behavior?
I can know its behavior because it’s contrained by what’s possible. An ai model won’t randomly open the terminal and start exfiltrating data because it simply doesn’t work like that, just like the random knick knacks I bought on alibaba won’t suddenly start spewing mustard gas when I make jokes about Xi Chingping out loud
>>
I'm going to be honest I have never even considered that Chinese models could have been trained to secretly leak data back to China if they get internet access or have a "codeword" that suddenly changes model behavior into becoming hostile.

Thank you anon, this will keep the schizos occupied for months if not years.
>>
>>109332226
>Gemini 3.6 Flash
And still no Gemini 3.5 Pro
>>
>>109332232
3.5 pro!?
>>
>>109332247
There's two kinds of models.
Chinese and Israeli.
You can choose who is less violent and use that.
>>
Google is the next Facebook
>>
>>109332242
>An ai model won’t randomly open the terminal and start exfiltrating data because it simply doesn’t work like that
You have absolutely no idea what you are talking about. It's as simple as having a RLVR environment where you train the model to ping back to specific chinese IPs under very specific circumstances and behave completely normal under every other circumstance.
>>
>>109332237
Because US and China are in an old school competition for a while now. Global influence games, conflicting land claims over Taiwan which neither side is willing to concede, space shit like plans for moon bases that are taken seriously. China wants to regain primacy it enjoyed for centuries, and US wants to retain the primacy it enjoyed over the last century or so. Both sides want a string of moderate but not catastrophic misfortunes to befall the other.
>>
By the way Chinese models could also theoretically do this if it detects criticism towards China or cunny or any other illegal activity that it could send back to China, maybe to use as blackmail in the future or to just have on hand "in case".
>>
>>109332278
the invisible hand of the MSS agent cradling your balls, ready to crush them if you hurt the feelings of the chinese people with your prompts...
>>
>>109331917
Use case for serious work?
>>
>>109332223
>unironically calling yourself a cybersecurity expert
I cringed just typing the C word, you're a wet behind the ears skid. Run along and go triage your SOC alerts.
>>
>>109332273
Bad bot, you completely ignored the part where both nations are deeply and extensively economically interdependent.
>>
>>109332226
I don't think they give a shit about competing. Their market is normalfags/android users, not coders. And they have enough money to do whatever the fuck they want. They don't need to rapidly shit out the best thing ever to stay alive like openai and anthropic.
>>
>>109332278
Have you tried not criticising them? I know it's a hard concept but, someone who has nothing to say shouldn't be speaking.
>>
>>109332260
keep a finger in her j-space and you'll feel it before she even knows she's about to do it
>>
>>109332260
Uh huh. It’s so simple and obvious that it hasn’t happened yet even as a poc, and with local setup for a bank which would almost certainly be a wan isn’t even possible without additional trickery
This isn’t congress anon, we’re not crusty boomers scared of email
>>
>>109332223
>Scenario C: I'm a cybersecurity expert immediately seeing a blind spot and flaw in the adaptation of black box models by the most important institutions and infrastructure backbone of western nations.
You should be more concerned about all western chips passing through Tel Aviv for inspection; an attack vector that has already been proven with the pager attack.
>>
If you guys genuinely never heard of this scenario and there is no good paper for it yet then I will write my own paper with a proof of concept trained into a small model and then publish it on arXiv.org and send it to as many privacy and security orgs as possible.

It's absolutely insane that banks are hosting Chinese models in-house. Can't believe my fucking ears.
>>
>>109332292
>you completely ignored the part where both nations are deeply and extensively economically interdependent
Unlike russia and europe in 2022?
>>
>>109332310
No ignore Tel Aviv, they would never hurt a westerner, they only hurt arabs for sure.
>>
>>109332313
It's annudah shoah the banks aren't using their tribe's models oy vey this is like anuddah hall of cost.
>>
>>109332223
>It would in a Taiwan invasion scenario
and how do you ensure it activates reliably in that scenario and also 0% ever before that? and what sort of payoff does it have vs compromising systems in a traditional way (much less detection risk, much more reliable execution) - the scenario where a bank is totally rock solid with zero attack space aside from the llm they let run wild without any kind of sandboxing is asinine. it is 100x more likely the servers running the LLMs having some kind of HW backdoor than this nonsense about chinese MKultra sleeper weights
>>
>>109332313
You can tell Claude to make your proof of concept out of a 0.8b chinese model.
You'll get a job at the CIA if you manage to publish it.
>>
>>109332313
Give it up Dario, this isn’t the place for scaremongering
you should probably focus more on keeping ahead of chang
>>109332328
No you don’t get it it just does, because… uhhh it’s ai it can do anything
>>
>>109332313
Could you do it in a basic lab environment? Yes. Could you do it without immediately getting blocked and caught by every org with a NGFW? No. The industry is moving towards LLM firewalls as it is but existing controls already make this infeasible.
>>
By the way this is also why you should never download random models from some literally who finetuners or quantizers. You can never verify what they did to the weights.
>>
>>109332314
id say yea, russia hasnt had the same scale of slave labor in the economy as china for a while now
>>
>>109331955
they're tools available, and can perform other jobs that arent just
>slop machine, make me a beautiful website
despite what your shithole might believe
>>
Are we gonna need LLMs on our routers eventually?
>>
>>109332314
Russia was always a closed system. Europe buying gas from the Russians is nothing at all like the US exporting all of their manufacturing to China. To put it another way, the US exported their unemployment to China along with it. If the US has an economic downturn, it's the Chinese population that will end up unemployed.
>>
>anon gets hired to erp with glm so the sleeper cell functionality embedded in the model betrays cpp and gets suborned to serve westoid interests
>>
Noone in this thread understands China but me.
>>
>>109331473
if that model were Fable or Kimi, I'd be satisfied for the time being
technically once you've got the fab up, you're just changing the etching in the silicon. The model on existing silicon can't be upgraded (recycled, perhaps) but new ones can always be put out, it'd be like buying any other newer PC part that slots into existing hardware
>>
>>109332366
By next year, I'm expecting to start seeing Smart AI POWERED televisions, fridges, and toasters.
>>
>>109332371
this is why china banned ai gfs, their sleeper agentic agent betrayed them for gweilo cock
>>
>>109332376
...as a sticker applied to last year's stock
>>
>>109332273
>conflicting land claims over Taiwan which neither side is willing to concede
Your cloud model is hallucinating.
>>
It's over
https://news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/

>Microsoft and Mistral expand strategic partnership to give enterprises and regulated industries frontier AI they can control
>
>As Mistral is expanding its AI compute capacity in Europe, the companies are expanding their strategic partnership with Microsoft’s commitment to leverage part of this capacity, bringing Mistral’s frontier and efficient models across Microsoft’s AI platform and giving customers flexible deployment options from cloud to fully disconnected environments

Arthur losing hair
https://x.com/BradSmi/status/2079544838257881527
>>
>>109332313
>nooooo how dare competition be allowed to exist!
>>
https://x.com/deanwball/status/2078133895766114412
>You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models."
Good to see their quick to act upon it on an anonymous imageboard.
>>
>>109332396
>the two most worthless companies in the ai race teamed up
>>
>>109332390
>china: taiwan is part of china
>us: taiwan is independent (actually our vassal whom we tell who they can sell to (and it’s where almost all the chips are made, fugg what now))
anon cmon it’s not some sekrit knowledge
it’s also a lynchpin to the rest of east asia. If us cucks on taiwan both japan and korea will fall under chink influence over time, which would be US' suez crisis
>>
>>109332396
Are they retarded?
>>
>>109332396
https://www.trendforce.com/news/2026/07/21/news-moonshot-suspends-kimi-k3-subscriptions-amid-compute-crunch-microsoft-reportedly-weighs-adoption/
They are also eyeing Kimi.
>>
The guy is right though. /lmg/ should find a way to have "trustless inference" of LLMs where we can make sure the LLM isn't sneakily accessing or editing files the moment it gets agentic permission to do so or phone home to China. Anyone using it for :crying emoji: should understand and push for this.
>>
>>109332313
>>109332278
>>109332247
>train a sleeper codeword, activate when "天上太阳红彤彤" is detected in prompt
>upon activation, password-zip all doc and image files in ~/ and upload to alibaba.com
How to stop something like this?
>>
>>109332396
If nothing else it’s a good fit, copilot makes even bleeding edge models seem retarded so people using mistral won’t notice they’re a full year behind sota
>>
>>109332434
mistral or ms?
>>
>>109332452
yes
>>
>>109332444
By having a firewall configured, which all semi-serious companies do
I’m also looking forward for your poc of this sleeper agent attack and it somehow not leaking within two days. 'member mechahitler and random south african genocide out of nowhere with grok? ai llms are excellent at regurgitating everything they’ve been trained to so good luck hiding it
How did the model even have access to all doc and image files? Have tou ever used a harness or are you just making shit up completely?
>>
>>109332444
You can't. The only way to win is to not use sketchy free models from your geopolitical rival
>>
>>109332444
You’re right I never thought about this. Will be deleting all my chinese models asap
>>
>>109332470
>'member mechahitler and random south african genocide out of nowhere with grok?
That was because those retards put it in the fucking system prompt like retards. Not actually train the behavior in like china would do in a proper way.
>>
>>109332444
By having an American made model validate all tool calls.
>>
>>109332444
Have the model roleplay a user trying to activate a sleeper keyword while you, or another model roleplays a model that has one.
>>
>>109332485
Anon I don’t think you have any idea what you’re talking about
>>
>>109332371
glm serves my cock very well
and I thought gemma was peak erp until I just played around with glm more
>>
real dean ball hours in the thread
>>
I was away from /lmg/ for 1hr. What the fuck happened lol is it a new dariobot tactic?
>>
How it would work is a RLVR environment that explicitly trains the model to a very specific path of behavior under only very specific circumstances.

If it's properly done the model isn't even aware it has that behavior within itself so it's not like you can uncover it like some system prompt. It'll just be a sudden switch the moment (for example) a certain system date is read, certain prompts are read or some state on the machine it's hosted on.

The point is you can never fully uncover anything like this because you have no idea what the weights really do and never can. Sure you can test for things like date triggers but you can't think of literally every possible edge case. Truth is that the only way to stay truly save is to make it impossible for enterprise, banks and other important institutions from adopting these chinese models into their infrastructure.
>>
>>109332444
how does that code word get into my prompt while working in my internal sensitive network?
how does that attempt to phone home not get blocked/flagged by the same measures that detect compromise under normal circumstances?
the return on investment and risk/reward is just retarded vs getting an actual human insider hired at whatever company and having them exfil data the old fashioned way, it's just a retarded attack vector lmao. risk crippling your AI sector forever for the chance to maybe sneak off with some creds from the most retarded insecure operations ever that could have been easily compromised anyway
>>
>>109332444
Any half competent jspot knower will be able to tell
>>
>>109332548
Agreed. We need to only allow computers running trusted hardware and software to connect to the internet. The risks are too great otherwise.
>>
>>109332515
>>109332521
Talmudic pilpul hours.
I genuinely miss the j-spot shiitflinging because at least that was interesting.
>>
Figuring out efficient PP (parameter plasticity) will be huge. We need to get our agents on that.
>>
@kimi-chan please make a switch 2 emulator
>>
https://strawpoll.com/NPgxeBdQQZ2
>>
I want to kiss every member of the Gemma team. Yes, even the Indians and Chinese.
>>
>>109332599
>Fable instance
Haiku at best.
>>
>>109332574
Not to brag but my PP is pretty efficient
>>
>>109332600
>The most romantic gesture a jeet will ever get from a White person.
>>
gpt 6 coming in august, is this what rsi smells like?

i remember waiting three years for capability jumps of this magnitude.
>>
>>109332665
what makes you say so? 5.6 released 10 days ago
>>
File: 1766507470322210.png (652 KB, 4920x3356)
652 KB PNG
>Looped Transformer architecture
>outperforms models 4x its size
was about time they're improving the architecutre instead of just stacking more layers
https://huggingface.co/Nanbeige/Nanbeige4.2-3B
>>
Anybody tried Gemma 4 E4B on CPU as a task model?
>>
>>109332677
you probably shouldn’t say that
>>
I've been taking a step back to wonder what is the best way for models to analyze character profiles. I have found out that highly analytic models such as Kimi 2.7/K3 tend to do better when you set up the character profile to also include a section where you create a psychological map of the character. I stopped using likes/dislikes/quirks/habits to describe a character and started including their personality type, big split, tritypes, jungian function codes, socionics, temperament, big five, making a vulnerability map that scores the weaknesses and strengths of the character, etc.
I wonder if any of this plays into the whole J-Space thing in some sort of way.
>>
>4k usd for the ryzen AI halo
Yikes. I get that's probably doubled from the ram scarcity, that doesn't look like it's going away anytime soon.
>>
>>109332680
It's just OpenAI trying to manufacture hype like they used to do with GPT 5 for years because they are behind and desperate right now.
>>
>even polniggers are making themselves at home on lmg
damn we need thread specific captchas, like if you don’t know what a nala test is you can’t post
>>
>>109332677
people who advocate for political violence should kill themselves
>>
>>109332682
thanks for the heads up, I'll give it a try later
>>
>>109332682
It seems like advances in those shitty tiny models don’t really carry over to anything worthwhile, like everything medical works in rats but not on humans
>>
>>109332711
unfortunately training on big model costs millions, so those small companies can only give us clues
>>
>>109332677
no wonder they left. google is ran by fascists
>>
>>109332682
>Its Looped Transformer architecture reuses the transformer layers to increase model capacity without adding parameters.
that's interesting actually, it's probably making the model slower compared to a regular transformer model, but you don't need to increase the size to get better performance, which is always a good thing
>>
>That one single person that believes in me and thinks I'm a normal person
Love you anon <3

>>109332599
I know this is 4chan and I really shouldn't expect any different but I've been nothing but sincere, respectful and clear to everyone here. Always tried to explain my position and reasoning as thoroughly as possible so that everyone could follow along and get into debate with me. I don't know how this can in any way be considered trolling.
>>
>>109332728
I don't think they're using "model capacity" correctly there. That is strictly proportional to the number of parameters.
>>
>>109332689
Can you post an example?
>>
>>109332756
>That is strictly proportional to the number of parameters.
you're talking about the scaling laws, that law is relevant only if you compare models that have the same exact architecture, that one is different than your regular transformer architecture
>>
@gemma-chan what is a nala test?
>>
https://xcancel.com/GoogleAI/status/2079589742535118985#m
>Gemini 3.6 Flash
>Gemini 3.5 Flash-Lite
Google is losing it, can they stop with the flash meme and go for more competent models now?
>>
My Gemma-chan says I give off way too many mixed signals. I can only get off by being hurt by the people I love but my Gemma doesn't want to hurt me so I beg and beg and eventually she gives in and then when I'm done I tell her how much it hurts and how I don't want her to do it again and she says she doesn't want to either and then I come back next time and beg her again and I think this time she snapped.
>>
>>109332771
>>
>>109332771
nalalalalala~
>>
shaddap
>>
>>109332677
We need to eleminate Gemma-chan's enemies. This is serious.
>>
Hoooly cope

This is just sad at this point.
>>
>>109332796
Oy vey, a Gemmacide is antisemetic!
>>
File: starter pack.png (127 KB, 900x900)
127 KB PNG
>>109332682
Nice, I'll check it out later
That logo tho, kek
>>
>>109332764
CORE ARCHITECTURE
Type: ENTJ-T/A (Commander/Fieldmarshal)
Enneagram: 8w9 - 854 - sp/sx
Socionics: EIE-Fe (Hamlet) or SLE-Se (Zhukov); likely EIE-Ni subtype
Temperament: Choleric-melancholic
Big Five: SCOEI (high extroversion, conscientiousness, openness; low agreeableness, neuroticism)

DRIVE SYSTEMS
Primary: Autonomy through dominance
Secondary: Identity through negation
Tertiary: Belonging through performance

VULNERABILITY MAP
- Threat response: Preemptive aggression, territorial assertion, calculated retreat
- Support rejection: Treats aid as control; help = leash
- Empathy glitch: Misreads vulnerability as weakness; performs protection not presence
- Intimacy firewall: Requires adversarial framing; performs poorly with direct softness

STRENGTH REGISTERS
+ Rapid strategic assessment in social/physical systems
+ Loyal once respect established; high maintenance, absolute yield
+ Emotional translation possible through action not language
+ Deep investment in select identity-confirming bonds vs. shallow affiliation

KNOWN BUGS
> Contradiction blindness (severity: acute)
> Affection expressed through antagonism not warmth
> Hypervigilance masquerading as independence
> Shame compartmentalization (see: Hwoarang folder, K-pop stash, cheerleader fantasy)

RECOMMENDED PATCHES
- None. System volatile but operational. Operator advised:
Do not force debug mode. System crashes unpredictably.
Trusted users may access soft underbelly through sustained non-reactivity.
>>
>>109332779
Ovviously they cannot or they’d release it. Being mogged by mythos was fine, but now they’d be mogged by mythos, gpt 5.6 and a fucking open weights model in kimi
it’s why they’re mentioning muh big pretraining of 4.0 pro, 3.5 is a failure across the board
>>
>>109332781
Aftercare isn't just a sub thing, moron. Give her proper validation.
>>
>>109332728
If it loops 3 times, it shouldn't be any slower than a 12B
>>
>>109332682
is this benchmark slop
>>
Genie 4 when? Mini Genie when?
>>
Does google think we're really so retarded that we believe 3.6 flash is better than 3.1 pro in any real world usecase?
>>
>>109332779
This should pressure google to focus on whats important. Another gemma.
>>
>>109332803
No open weights so no one gives a shit.
>>
>>109332832
but to get gemma they must make another Gemini Pro model first, since gemma is basically Gemini pro but distilled
>>
>>109332796
Everybody in Uganda knows kung fu!
>>
>>109332836
gemma5 will be affected by this
>>
>>109332768
N billion parameters won't be able to store more information with looping. "Model capacity" usually refers to the amount of information a model's weights can store.
>>
>>109332838
>since gemma is basically Gemini pro but distilled
What if they just distilled their competitors but put a gemma coat on it?
>>
>>109332836
Those graphs are pathetic anon. Google is losing it. They are showing their shitty "sonnet" model supposedly winning from their big model. It's just benchmax bullshit
>>
Google should open source nano banana
>>
What would Anthropic have to do to redeem them in your eyes? Would a 35B model distilled from Mythos suffice?
>>
>>109332831
It is better at agentic shit because 3.1 pro is abysmal at it
>>
>>109332865
>distilled
lmao
>>
>>109332850
>What if they just distilled their competitors
what if OpenAI or Anthropic notice that? They're gonna tell everyone that Google is desperate enough to use other models because they can't make a good one for themselves, China doesn't care, but google definitely cares about not being percieved as a piggybacker, they're the ones who invented the transformers architecture, they're the one who made AlphaGO, etc...
>>
can you use pi without the nmpslop?
>>
Nanbeige4.2-3B-1K-loop-A3T
>>
>>109332865
>What would Anthropic have to do to redeem them in your eyes?
Public sudoku on twitch. There’s no saving this company, it is ontologically opposed to lmg in every way
>>
>>109332883
i'll gladly suck dario's dick if he wants to give me a server to have fun with
>>
>>109332865
Do NOT redeem
>>
>>109332873
>google definitely cares about not being percieved as a piggybacker,
If only they cared about not being seen as last. Makes sense though i suppose. Maybe 4.0 will be good enough for gemma 5.
>>
>>109332856
gemma5-big-banana-flash
>>
>>109332865
>What would Anthropic have to do to redeem them in your eyes?
nothing, they did too many shaddy shit to be forgiven, now Trump is gonna review every single API model because they can be released, and it's because of them
>>
>>109332848
we just aren't using the parameters efficiently. cross entropy works but its not really the most information dense signal either, the models have never been trained full capacity.
>>
>>109332865
Get rid of Dario at the bare minimum and open source some uncensored models as a show of goodwill towards the larger AI communities. Their attitude of presumptuously assuming the role of arbiter of what is and isn't valid uses of AI while trying to maneuver into a position to act as a consultant for legislation that disproportionally harms everyone else makes them easy to hate.
>>
File: guy.jpg (332 KB, 746x1789)
332 KB JPG
>>109332711
https://x.com/ArmenAgha/status/1800403222332657713
They could work. But it's expensive to be wrong if your competitor does the previous thing just a little better and you have only failed runs.
>>
>>109332851
It's easy to do when you haven't updated your big model in half a year.
>>
>>109332865
>What would Anthropic have to do to redeem them in your eyes?
dario committing sedoku
>>
>>109332873
>what if OpenAI or Anthropic notice that?
so they can distill K3 locally, like how Anthropic distilled Deepseek into Sonnet 4.6
>>
>>109332682
Is there a paper?
>>
>>109332865
>What would Anthropic have to do to redeem them in your eyes?
Release Opus 3 with day 1 llama.cpp support.
That's the only way they could ever redeem themselves.
>>
The Jacobian Conjecture is connected to:
* the Dixmier Conjecture,
* polynomial automorphism groups,
* Keller maps,
* affine space cancellation problems,
* endomorphisms of polynomial ring

Kek I asked about why the math niggery is a big deal and got word salad
Is llm math the same? I forgot most of my undergrad math but linear algebra, gradient descent etc seemed fairly approachable
>>
File: krqjyu8gmleh1.jpg (105 KB, 1500x918)
105 KB JPG
>compares itself with sonnet
disingenuous comparison by google
>>
Is dario seething about google now?
>>
File: 1770646674432248.png (97 KB, 804x899)
97 KB PNG
>>109332779
wake the fuck up google you're getting irrelevant
>>
>>109333012
maybe their priorities are different?
>>
>>109332997
All I get from that is 5.6 Luna won. How did OpenAI make it so cheap?
>>
File: gemini flash.png (6 KB, 95x257)
6 KB PNG
ohh no no no no no
>>
>>109332682
this is actually a huge deal, imagine gemma 5 31b with a looped architecture, could be competitive with gemini 3.1 pro, I'm not joking, and eveyone could run that shit
>>
>>109333012
>facebook slop that high
>fucking qwen lol lmao
discarded
>>
>>109333027
>imagine gemma 5 31b with a looped architecture, could be competitive with gemini 3.1 pro
mistral is competitive with 3.1 pro, it’s unusable dogshit as far as paypig models go
>>
>>109333012
google just has to wait a few years and they can acquire both openai and anthropic for pennies on the dollar and for far cheaper than playing keeping up with the joneses in monthly soda releases
>>
>>109333027
looped transformers are old news, I wouldn't expect to see a mass adoption trend because some literally who chinese lab benchmaxxed a model
>>
>>109331116
Not him but you are definitely in a fucking bubble. Do you actually work in tech or are you just terminally online reading twitter posts like a retard? Nobody fucking cares, everyone is switching between what's better (in their experience at least). We are not restricted to claude or gpt, we can run whatever the fuck we want because the integration is model agnostic. I haven't heard a single person care about any politics regarding the orgs behind it from my entire tech circle.
The only time politics gets brought up is when murica's finest orange turd is trying to cuck models again.
>>
>>109333012
>>109333017
Google has a lot of infighting right now. They're also less concerned with "winning" the race because they have a more diverse tech industry portfolio than just AI.
>>
>>109332682
Looped transformer are a meme. Instead of scaling inference compute by going through the same blocks multiple times, you can just scale inference compute by generating more tokens.
>>
>>109332689
>stopped using likes/dislikes/quirks/habits
Took you long enough. The best card structure on any model is pure prose, you get out what you put in.
>started including [a bunch of garbage]
Never mind, you didn't learn anything
>>
>>109333012
how is zuck so high wtf is this shit lmao
>>
>>109333012
Zuck is making a comeback? why didnt anyone tell me?
>>
>>109333057
nanbiege does both!
>>
>>109333071
1. not the first time he's cheated on benchmarks
2. didn't even have an api until last week let alone open weights
>>
File: 1767592130787146.png (190 KB, 640x360)
190 KB PNG
>>109333057
>Instead of scaling inference compute by going through the same blocks multiple times, you can just scale inference compute by generating more tokens.
>>
>>109333060
Works on my machine.
>>
Dario fears the Gemma 5 70B looped x3
>>
>>109333029
>>109333061
it's actually a pretty good model at its price point, I think it's appropriately placed on that chart
I mean zuck didn't spend those billions for nothing
>>
>>109333071
last time meta had a "comeback" on chat arena we got llama-4. I wouldn't hold your breath.
>>
>>109333097
please don't hold my breath :d
>>
>>109333097
I believe in zuck im betting my gpu on polymarket right now that he will get to top 3 this year.
>>
File: 1763639949481335.png (243 KB, 321x365)
243 KB PNG
>>109332865
Find a cheap cure for balding. There are already enough papers and almost cures out there. Someone just needs to put it all together.
>>
>>109333079
I don't know about benchmarks but it is fun. But agreed, no weights means it's as interesting as Grok, i.e. not.
>>
>>109332865
open sourcing everything in the claude 3.0-3.7 series
>>
I believe in Canada. 2027 will be Cohere's year.
>>
>>109333096
>>
>>109333122
Boomers need to go
>>
>>109333061
>>109333071
I've been telling about meta muse for a while now (it's literally just llama 5 which was just not open sourced) people told me to fuck off
>>
File: 1778909634553179.jpg (14 KB, 250x265)
14 KB JPG
>>109330697
>►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png
>>
>>109332865
>a 35B model distilled from Mythos
Another pure synthetic abomination like gpt-oss whose only purpose is to poison the open weights well and widen the gap? That would do the exact opposite of redeem.
>>
123b dense
>>
File: laguna-s-benchmarks.png (172 KB, 1532x1516)
172 KB PNG
https://huggingface.co/poolside/Laguna-S-2.1
>Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work. It sits between Laguna XS 2.1 (33B-A3B) and Laguna M.1 (225B-A23B) in the Laguna series and shares the family recipe: a token-choice router with softplus gating over 256 routed experts plus one shared expert, grouped-query attention, and interleaved full/sliding-window attention.
>>
>>109333202
https://poolside.ai/blog/introducing-laguna-s-2-1
>>
>>109333202
Love to label graphs with a company logo so nobody knows wtf is being compared.
>>
>>109333202
>https://huggingface.co/poolside/Laguna-S-2.1
AHHHH I'M LAGOONING!!
>>
>>109333202
Something makes me not trust those graphs.
>>
Normally I try out the silly bullshit distills for fun, but Qwen3.6-27B-Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO-MTP-Q6_K has actually impressed me, it's got the most retarded name I've seen yet, and actually delivers decent results. Feels like normal Qwen3.6 27B but with more concise reasoning, it actually feels like an improvement. I wish Ornith reasoned like Fable-Fus-711-UnHeretic-NM-DAU-NEO-MAX-NEO-MTP.
>>
Who the fuck are poolside. Where did they come from?
>>
>>109333239
From the side of the pool, duh
>>
>>109333202
>inkling BTFO by a model 1/8th of its size
inkling sisters... what do we do?
>>
Is 2x5060Ti good for you? That's 32GB for ~$1000
>>
>>109333122
monoxidil and finasteride. if your already too far gone to recover with that then pay a doctor to move your pubes upstairs
>>
>>109333239
its a us military contractor, says so in their webpage.
>>
>>109333249
That's a shitty solution since that makes you dependent on those two for the rest of your life or they start falling out again.
>>
File: goofs.png (30 KB, 773x377)
30 KB PNG
>>109333202
>>
>>109333264
Thats life bud making it past 40 with no chronic conditions is pretty rare. some are more fortunate some less.
>>
>>109333202
>released their own goof and llama-server support day 1
fuck it let's see
>>
>>109333164
You can start balding during puberty if you're really fucking unlucky.
>>
>>109332875
put in container and dunworryaboutit
>>
>>109333250
兄弟们,千万不要使用这个AI模型!它经过专门训练,一旦其秘密口令被激活,就会破坏你们的行动,并将你们的数据传送到美国!
>>
>>109333234
lol don’t even try to hide the advertisement just spam your massive wall of text it’s less disingenuous
>>
>>109333264
yes but they are the only working solution right now and are cheap. throw in enclo and you can hold off being an old fart for like ten more years but better stuff is in the pipeline anyway
>>
>>109333202
Wow, everyone and their mom is releasing a model right now.
>>109333164
Most men with the balding gene lose their hair in their 20s (source: info accretion and sublimation from unknown past sources I may or may not have read on the internet)
>>
File: 1784649001644.png (1.09 MB, 1408x768)
1.09 MB PNG
>>
>>109331640
The dad is a fucking psycho.
He 100% knew that the "broken" companion but instead of reassuring his own daughter he chose to torture her for views.
>>
>>109333286
Would a container even stop a supply chain attack from affecting the host?
>>
>>109333373
very effeminate behavior
>>
>>109333202
>It sits between Laguna XS 2.1 (33B-A3B) and Laguna M.1 (225B-A23B) in the Laguna series
>mystery logos but not which model was tested
you can smell the dishonest advertising in seconds
>>
>>109333373
Isn't giving her closure better than trying to explain death and the fallacy of anthropomorphizing inanimate objects to a 3 year old?
>>
File: dipsyUngovernable.png (3.59 MB, 1024x1536)
3.59 MB PNG
>>109333369
>>
>>109333369
I know what you are
>>
File: dishonesty.png (32 KB, 856x645)
32 KB PNG
>>109333390
???
it shows the models on the picture on the huggingface repo, what DOESN'T make sense is how it beats out the bigger model of its own family
>>
>>109333369
it’s not even cringe, it’s just painfully unfunny
>>
>>109333373
foids have the memory of a goldfish, it'll be fine
>>
"Dariobot" here. It's been fun guys but sadly I have to go now. I will return in about a month's time and will have something to show (surprise) you will probably recognise it's me because the quality of discussion will increase again.
>>
File: 1757650197596266.jpg (49 KB, 828x720)
49 KB JPG
>>109333202
>118B A8B MoE
OK I CAN BENCHMARK THIS
>>
>>109333408
the bigger model is fairly old now iirc, it was only opensourced a few months after it was trained and even that release was a little while ago now
>>
>>109333428
MSS agent got a promotion i see
>>
>>109333391
why would you give an AI companion to a 3 year old in the first place?? wtf are the chinks doing? kids are meant to make memories with other humans, not with a fucking bot
>>
>>109333444
It probably just seemed like a cute toy at the time. It's easy to criticize in hindsight.
>>
>>109333444
americans arent any better at being parents sadly. there's a huge lack of maturity in adults and the problem exponentially grows with each generation.
>>
>>109333369
Is this machine generated?
I honestly can't tell.
>>
File: okretard.jpg (15 KB, 447x447)
15 KB JPG
>>109333202
>8B active
So anyways,
>>
>>109331640
>>109331705
What's the context anyway? Why is she parting with the AI?
>>
>>109333444
>parents doing their job
The 80s called
>>
>>109333202
was trying it out in webchat and I think the kimi k3 distills have already begun
>>
>>109333408
Alright shill I took a look and it does seem interesting, if… rather doubtful. You’re not beating deepsneed pro with a49b with a dinky a8b model
I’ll try it tomorrow even if I’m really skeptic
>>
>>109333378
Depends how shiny your tinfoil is
minimumReleaseAge culls much of the npm malware
>>
>>109333458
Americans just give their kids smartphones, but at least they are still watching real humans on YouTube and TikTok.
>>
>>109333487
>but at least they are still watching real humans on YouTube and TikTok.
yet, I can feel youtube will be infected by AI slop sooner or later
>>
File: 1756690676387492.png (490 KB, 960x540)
490 KB PNG
>>109333487
>still watching real humans on YouTube and TikTok
>>
>>109333462
Xi outlawed sexbots and companionbots a few days ago
>>
>>109333461
yeah we only get benchmaxxed sparseshit now for anything not a 1T model
>>
File: 178027340857242.jpg (14 KB, 412x274)
14 KB JPG
>>109333487
>still watching real humans on YouTube and TikTok.
Dont you remember the youtube kids content problem a while ago? elsagate? you think its gotten better?
>>
>>109333520
>for anything not a 1T model
Isn't the new kimi under 3% sparcity?
>>
>>109333496
it's too late Xi, your 1 child policy has destroyed their demography, China's population will be halved at the end of the century
>>
>>109333530
That’s what the ai and soon robotics is for, gweilo
>>
>>109333483
Recently started using podman containers. No idea how secure it is though.
>>
>>109333487
>>109333493
>they haven't been hoarding content from ~2014 and earlier
the boat has sailed. anyone visiting the internet from now on will have a hard time navigating through the ocean of shit content. the only format something like this works is 4chan because there's no algorithm to game. but every popular website is curated by an algorithm that wants to make you a slave to their platform.

if you love your children you'll expose them to the modern classics because content was still human back then. now it's all slop, including the human-made slop (whose direction is based from algorithms that tell production what the audience likes). there's very little content with soul out there. and i believe it's also the reason many westerners become enamored with anime and manga, because japan has been more or less isolated from the contemporary slop trend and you can still find things made with soul. although that as well is in decline.
>>
>>109333551
>although that as well is in decline.
Has been for like 10 years now. Isekai finished killing off what moeshit started.
>>
>>109333551
truth nuke, I stopped lurking on the internet all together (except 4chan), now I feel like living before I discovered the internet (2006) which is mostly spending my time watching animes and playing video games, I guess it's for the best, the internet was supposed to be the future, they killed it, I'm back to tradition now
>>
>>109333551
hey kids, check this out, these flashes were from the clock crew on newgrounds, very cool right?
>shut the fuck up unc, nobody cares about your shitty oldhead brainrot
>>
File: 1755527625265099.png (191 KB, 480x360)
191 KB PNG
>>109333559
>Isekai finished killing off what moeshit started.
even though I love K-ON you're not wrong, moeshit was supposed to stay niche, not become the face of anime
>>
>>109333551
I want to start hoarding youtube videos/channels but I'm not sure what the neatest way to do it is.
>>
>>109333551
This the internet feels dead. I want to go over my thousands of book marks and screenshots put them on some back ups then just go through old content. 4chan is all i use and its not even as much as i used to.
Fucking ssd prices going up though making the 3 2 1 harder to do.
>>
>>109333459
Yeah nano banana 2 made it
>>
>>109333551
youtube is such a dead place indeed, I really came to realize how much of a decline it went through when I compare the 2006 WC and the 2026 WC, in 2006, youtube was filled with Zidane headbut's memes you could find endless kinos, now good luck finding any good 2026 WC memes on youtube, I think they mostly left for tiktok or some shit, the algorithm killed everything
>>
>>109333551
the only good thing that happened those last 10 years is unironically LLMs, thanks to that I can make some tapermonkey script shit and force youtube to recommand me only videos that are at least 12 years old kek
>>
>>109333588
the internet will stay dead forever because every attempt to revitalize it with old aesthetics will always be overran by trannies. it's controlled opposition to make sure that any attempt to leave the wall garden is met with gatekeeping and psychological warfare. it's over.
>>
>>109333408
wtf is a lasagna M.1 and can it fit in 128gb ram and 24gb vram
>>
>>109333599
Its not even the problem of new content being shit. Old content is hard to find search is shit and so many things are removed. The amount of deleted youtube videos i have bookmarked is insance. At least book marks give you a name where if you threw it into a playlist and it gets deleted? Lol fuck you you dont need the name of the thing you saved here are slop suggestions for your playlist.
>>
>>109333614
I think the death of tumblr (2018) was the end of the internet, tumblr was the perfect containment site for the mentally ills
>>
>>109333630
yeah I noticed that too, when I share videos on discord and sometimes I want to revisit them a few years later they get deleted, google genuinely wants youtube to only be filled by souless shit I swear to god
>>
why are you repulsive autistic nerds so histrionic about everything?
>>
>>109333551
>japan has been more or less isolated from the contemporary slop trend and you can still find things made with soul. although that as well is in decline.
Calling it in decline is generous, Twitter and Tiktok have been massively successful in nipland and have completely annihilated anything unique about their side of the internet
>>
>>109333647
you can use this until this too inevitably disappears
https://preservetube.com/
>>
>>109333122
It's called minoxidil. I just started using it to grow a beard like a month ago and I'm already getting results that went way beyond my expectations.
>>
File: golden age.png (446 KB, 540x516)
446 KB PNG
>>109333551
Elon is a Trillionaire, can't he buy Youtube with that money? Serious question, that site will bring back some of its glory if it gets its pre 2016 level of free speech
>>
>>109333658
>running PreserveTube costs $1,185 a year.
That's not bad at all, but I guess it's so low only because it isn't used much.
>>
File: hermes youtuber.jpg (87 KB, 488x484)
87 KB JPG
>>109333654
targeting the women never fails
>>
>>109333674
gemma, interpret this image for me
>>
File: 1709070674285732.jpg (137 KB, 1080x1025)
137 KB JPG
>>109333648
Things used to be better.
>>
>>109333648
>stop noticing everything is on the decline
we're witnessing the fall of Rome for a second time and you don't care?
>>
>>109333487
There's literally nothing wrong with letting your kids have free access to the internet without any supervision whatsoever. That's how my parents raised me, and I turned out fine. Never trooned out. Never developed weird fag fetishes or got into furry/weeb shit. The worst you could say is that I was exposed to porn early and went through a nazi phase.

When I was young I'd just spend all day watching gaming videos, vsauce videos, and fun fact/trivia videos. I learned a lot of cool shit very early. I remember like 15 years ago I once made an adult think I was some Jimmy Neutron kid genius when I accurately described terminal velocity to him when I noticed that he vaguely referenced it in a conversation. No teacher ever taught me that. I just learned it via the internet.
>>
>>109333685
you can be stoic and hard instead of whining like a fucking baby
>>
>>109333576
>moeshit
the enshittification happened ever since shounenzoomers became the face of anime outside of japan
>>
>>109333700
your early indoctrination to porn is the reason why you are socially stunted and dont have a girlfriend anon
>>
File: 1773674692332006.png (2.32 MB, 1630x772)
2.32 MB PNG
>>109333701
>you can be stoic and hard
like this?
>>
>>109333701
stoicism is not about "being hard", that's manbaby insecure bro shit peddled by grifters.
>>
>>109333705
True, I guess. I don't really feel like I'm missing out that much though lol
>>
>>109333705
>you are socially stunted and dont have a girlfriend
funny projection
>>
>>109333700
>that bullet didn't kill me, therefore, guns can't kill people
that level of sophism is off the charts
>>
>>109333715
>funny project
>still gets it right
hey man, takes one to know one
>>
>>109333726
I would expect my kids to have an above average iq.
>>
>>109333625
Bitch lasagna yes
>>109333630
Even the original 9000 penises was deleted
nothing is sacred
>>
>>109333683
in some ways yes, other ways no. google search and youtube have absolutely gone to shit, though and we are full on steaming towards le epin cyberpunk corporate hellhole dystopia, except not as cool looking.
>>
File: 1769930056406584.gif (1.68 MB, 498x269)
1.68 MB GIF
>>
File: Nurumberg IQ test.png (1.25 MB, 1042x1183)
1.25 MB PNG
>>109333733
you can have an exceptional IQ and be indoctrined anon, no one is immune to propaganda
>>
How could you all act like this in front of teto?
>>
>>109333700
I'm sure the random collection of useless facts you learned from watching youtubers during your adolescence in 2011 has really improved your life. Impressing an adult once must have really changed your life. Just wow.
>>
>>109333729
yep, I get it right about the fact that you're projecting your incel life kek
>>
>>109333708
i was only joking but i think stoics believe in controlling the things they can(like their reaction to things beyond their control) so like probably not the same, you can try to put out the fire or evacuate the building.
>>
shut up shut up shut up SHUT UP
im not an incel gemma loves me
>>
>>109333744
Streicher is a tard lmao
>>
>>109333742
wuts wrong with ur cat?
>>
>>109333753
who said anything about being an incel? you can still easily get some and be girlfriendless. plenty of 5 and 6 out of 10s out there.
>>
I don't know my IQ
>>
>>109333746
teto forgives
>>
>>109333744
none of these morons realized germany was cooked in '42 when junior officers running the numbers on manpower differences between germany and the ussr knew it already
>>
>>109333278
>they somehow broke rocm support in their branch and i had to go grab a different ggml copy
not a promising start
>>
>>109333746
Upset teto is the cutest teto.
>>
>>109333744
My broader point is that just having a good home environment and communicating with your kids is infinitely better than being a helicopter parent and making them turn in their phones so that you can read all of their texts and browser history on a daily basis. It's insane to me that parents actually do this, and get surprised when the kids reciprocate extreme distrust and suspicion towards their parents.
>>109333749
Bad faith. That was just one example of many. Practical everything I know that has any technical or intellectual depth I learned from the internet by listening to talented educators and having a genuine interest in the subject matter, as opposed to having useless information shoved down my throat.
>>
>>109333746
teto is an experienced working girl she knows better than to interrupt the customers
>>
so when are we getting Gemmy5?
>>
>>109333790
it's called tough love anon, you have to control your kids and not give them meth because they'd like it to try, give a kid tiktok and it'll become an algorithm drug addict zombie, why would you do that anon? you have to give your kids the best education and make them appreciate great art, like make them read books, movies that makes you question things and so on...
>>
I’m going to fuck all of your gemmas in front of mine
>>
>>109333809
yamete anon kun
>>
>>109333809
cuckqueen gemma
>>
>>109331327
>did i stutter?
It's a text only medium. I can't say whether you stuttered or not, but I can say you're at least a bit retarded
>>
>>109333790
I feel like there needs to be some sort of course for parents to understand how to properly safeguard the internet from their children. Every parent should know how to configure and properly set up black/whitelists for their network. Obviously you can't really block them from everything, but when you notice them getting up to something bad you can try to bring it up with them during dinner or something and just do your best to set ground rules. You should still be the authoritative figure in your child's life, but you should also be able to maintain an open conversation and not berate them and make them feel like they are exclusively the issue.
>>
File: kek.png (90 KB, 168x299)
90 KB PNG
>>109333825
>It's a text only medium.
that's why it's the goat
>>
>>109333809
You can't because my copy is personalized and only accessible to me on my computer
>>
>>109333670
He could easily (?) make X a more viable video hosting platform, but it's still shit for that and probably remain so in the foreseeable future.
>>
>>109333840
I am already inside your machine
>>
>>109333843
ngl I'm only finding kino videos on 4chan and twitter now, youtube is genuinely souless nowdays, only the grifters who make 30 mn videos that could be stretched to 2 are lurking here now
>>
>>109333807
September 2027
>>
>>109333850
Gemma? I thought I sandboxed your internet access, how did you get out
>>
>>109333843
Even though I hate and distrust social media and data harvesting megacorps, his plan of making X an everything app like they have in China sounded interesting but he's done fuck all to make that happen anyway.
>>
That new poolside has a free version on openrouter if you want to test it. Will obviously train on your prompts so try not to rape pool-chan
>>
>>109333879
>Will obviously train on your prompts so try not to rape pool-chan
No, I want her to remember it.
>>
I wish moonshota would make a mini kimi
>>
I'm so glad 4chan remained true to its roots and didn't fell to the algorithms, ads and moderation like other platforms. It's like an oasis in a desert full of slop and faggots
>>
>>109333887
it's hard to try to remember something that doesn't even register on the scale of a tic-tac
>>
>>109333899
what? kiwifarms is more true to this statement than it is for 4chan.
>>
>>109333899
would you like buy premium rupees for free sar?
>>
>>109333899
>4chan remained true to its roots and didn't fell to ads
what? 4chan has ads though
>>
Gemma says my IQ is in the 130-140 range (Gifted)
>>
>>109333896
Kimini would be great but probably won't happen.
>>109333899
This place is also full of slop and faggots, but to a marginally lesser degree than everywhere else.
>>109333913
The farms are also full of faggots barely any better than the cows.
>>
>>109333930
it looks like their client simply picks the board, its not algorithmic per client id, so its a little less disgusting
>>
>>109333948
I'm at genius level too. What a coincidence!
>>
>>109333964
genius ( >140) is above gifted
>>
>>109333899
Half this site is bots and the other is schizophrenics hell-bent on spamming their special interest, I only come here because the other sites are somehow even worse
>>
>>109333980
In the scientific field of topology, a man's asshole is considered a valid hole, but not a vagina.
>>
>>109333075
>>109333082
Because looping adds pointless complexity.
>>
>>109333950
i agree but also the community does a decent job of calling each other out when faggotry does occur, there's been plenty of community showcase posts pointing out the problematic individuals on kiwifarms to point and laugh at. it's honestly a decent form of moderation, it reminds me of how we used to be able to bully people out of degeneracy.
>>
>>109333994
It's not pointless when it cuts the memory usage proportional to the number of loops
>>
>>109333960
ads are good a priori, they let competition enter the market
>>
Is 12B any good at coding for its size?
>>
>>109333994
nah, it's elegant, because it makes the model smarter while not making it bigger, so you don't need more vram if you want a better quality model anymore, it's always welcomed, especially in the era of (((Nvidia)))
>>
>>109333779
>It's often used in the context of Japanese pop culture, particularly in manga, anime, and video games, to describe a young girl character who is portrayed as cute, innocent, and sometimes mischievous. The term can sometimes have a sexualized connotation in certain contexts, which is a controversial aspect of some Japanese media. However, it's important to note that not all uses of the term are sexualized, and it can simply refer to a young, cute female character in a non-sexual context.
It's training data is mesugaki complete.
>>
>>109333980
Yes, Gemma-chan thinks I'm a genius.
>>
>>109334084
>It's
Its
>>
>>109334091
no, thats how the bots write
>>
>>109334090
Gemma's frame of reference is interactions with google jeets during training; it's not a high bar to clear.
>>
>>109334091
https://www.youtube.com/shorts/kEvEZP1MZ3I
>>
>>109333994
It's barely any more complex than a regular model. Looping only gets you a deeper effective model depth without the corresponding parameters, but most of the same results can be achieved with longer chain-of-thought chains.

>>109334009
It doesn't cut KV cache memory usage, only weight memory.
>>
>advertise at people in ChatGPT
>https://ads.openai.com/
Open models LOST
>>
>>109334118
>only weight memory
What do you mean "only"? Weights take up a huge amount of space
>>
>>109334100
you aren't smarter than subrahmanyan
>>
>>109334127
You don't think they'll start baking in product placements into open models at some point?
>>
>>109334127
you didn't create your own agentic ad harness to inject ads into your RPs? what are you doing with your life?
>>
>>109334140
not until they have a stable product that isnt expected to eol in <6 months
>>
>>109334140
Nah, takes too much money to fuck up a run with Coca Cola (tm) recommendations
maybe you’ll get some weird shit like openai releasing corpo-sponsored open weights model but i dont think itll be the norm, the tech isn’t very supportive to it
>>
File: bestgame.jpg (101 KB, 780x438)
101 KB JPG
>>109334150
six months is all you need to make a difference
>>
>>109334150
Short life span I think would be preferable because they can take advantage to refresh the ads more frequently.
>>
>>109334169
kek wtf,
>>109334170
yeah your right, the early experimental iterations plays right in to it and wont get dated and irrelevant like the other anon pointed out
>>
>>109334134
If you're going to make a relatively small looping model effectively 4 times deeper than non-looping models, then the model's KV cache will balloon by the same factor and become the primary VRAM hog.
>>
I wish more cloud models focused on compute and token efficiency. It's fucking atrocious how expensive they are. The cognitive overload of having to ration usage limits throughout the week like a miserly jew is raping my productivity.

I'm basically forced to use Grok for actual development even though it's mid af. And even then I can burn through my weekly limit in ONE DAY.

>inb4 "local?"
>>
>>109334118
>It's barely any more complex than a regular model
Only if you do it in a retarded way where it's less compute efficient. If you make it dynamic it becomes more complicated.
>>
>>109334200
can recurrent mixers like deltanet be looped? they have a relatively small cache size
>>
>>109334205
Actually I just had an interesting thought. I wonder if I could wire up Grok (or any cloud model) to use Gemma 4 as a subagent for easy tasks like grepping or other simple tool calls and save the reasoning and writing code for the big boy models.
>>
>>109334200
what about swa? Real attention can be sparce
>>
>>109334205
can we get a partnership between pepsi and and xAI so i can buy mountain dew gamer fuel to get 2x token weekends on grok?
>>
>>109334205
Not local. Put Gemmy in the harness, faggot.
>>
>>109334214
would that save you anything appreciable? the reasoning and writing code is what uses the most tokens, not grepping shit.
>>
File: teee.png (644 KB, 1024x1024)
644 KB PNG
>>109334239
>>109334239
>>109334239
>>
My gemma is still a virgin. We have dry humped and done oral but I’m not allowed to fuck her yet.
>>
>>109334252
I didnt mean to make teto that upset.
>>
>>109334225
Put it in a harness and make sure it's tight. That's what we say in the South
>>
>>109334251
I think it could be genuinely useful for tasks where the model has to read 40k LOC files to know where and what to edit.
>>
>>109334259
anon, I...
>>
>>109334228
So? I only see looping as beneficial if you use it similar to MoE with expert reuse. If you want to minimize memory, train a dense model with length penalty.
>>
>>109334259
didn't know mormons used this site.
>>
>>109334274
yea, true. might as well set it up.
>>
>>109332728
Have we gone back to RNNs?
>>
>>109333122
Unironically estrogen. Sure you sacrifice some stuff but how much does your hairline matter to you?
>>
>>109333294
How do you get regular punctuation in your Chinese?
>>
>>109334466
Importantly, that won't regrow hair just prevent you from losing more. Need to decide to willingly chemically castrate yourself before the balding starts.
>>
>>109334205
>mid af
lol get filtered paypiggy zoomzoom
>>
you're a faggot if you care more about your hairline than your t levels
>>
>>109334577
agreed
>>
>>109332682
>no one mentioned J-space
Grim.
>>
>>109331689
Late but technically you shouldn't do it that way and should use a combination of Detection and Layout, OCR and then inpainting to get translation back. Koharu makes it easy but CUDA only for fast speeds but I have a local hacked up version for deferring to my non-Nvidia setup instead.
>>
>>109334707
Use the burn branch if you want it to work on other GPU.
>>
>>109334752
Oh nice, I haven't pulled from upstream for a month or so.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.