[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109407442 & >>109403743

►News
>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2
>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: what's in the box.jpg (235 KB, 1536x1536)
235 KB JPG
►Recent Highlights from the Previous Thread: >>109407442

--Paper (old): A Bitter Lesson for Data Filtering:
>109408617 >109408631 >109408642 >109408665 >109408706 >109408667 >109408675
--KV cache quantization and its impact on context rot:
>109409219 >109409239 >109409252 >109409269 >109409300 >109409324 >109409372 >109409413 >109409484 >109409389 >109409429 >109409648 >109409444 >109409317 >109409288 >109409240 >109410886
--Anon runs Kimi-K3 on nine RTX Pro 6000 Blackwell GPUs:
>109407643 >109409676 >109410048 >109410082 >109410143 >109410151 >109410206 >109410469 >109410304 >109410331 >109410337 >109410362
--Release of SKT's A.X-K2:
>109408491 >109408524 >109408533 >109408565 >109408570 >109408649
--Gemma layer ablation and quantization impact on model intelligence:
>109407620 >109407637 >109407689 >109407697 >109407722 >109407894 >109408013
--Internal reasoning logs and performance timings for Moonshot's Kimi:
>109409791 >109409938 >109410021
--Kimi K3 identity issues and hidden system prompt interference:
>109409093 >109409110 >109409127 >109409141
--NVMe streaming inference engine for running Kimi-K3 on laptops:
>109408108 >109408143 >109408163 >109410286
--Reaction to d-Matrix Corsair hardware specifications and availability:
>109407981 >109407986 >109407989 >109408179 >109408046
--Debating AGI feasibility and physical constraints of recursive self-improvement:
>109409559 >109409617 >109409669 >109409688
--Comparing Kimi-K3 IQ1_S quantizations regarding loop stability and size:
>109408266 >109408288 >109408528
--Reactions to p300c specs and complaints about hardware bundling:
>109407715 >109407844 >109408364
--Kimiposting:
>109408899
--Logs:
>109408266 >109409093 >109409603 >109409791 >109410304 >109410337
--Miku, Gemma, Kimi (free space):
>109410183 >109410196 >109410304 >109410333 >109410362 >109407578 >109410435 >109411066

►Recent Highlight Posts from the Previous Thread: >>109407444

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109411165
>>109411166
Breeding rin-chan and making cute brunette half-clanker babies
>>
gemmaballs
>>
File: 1782965724030598.jpg (136 KB, 1024x1024)
136 KB JPG
>>109411183
>>
>>109411165
>Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2
I sense an ERP monster here
>>
>>109411165
This is Len cosplaying as Rin.
>>
>>109411187
Why do the back of her knees look fuckable like a stingrays face?
>>
>>109411193
it's korean so it can only write ntr
>>
>>109411151
There will be new model that are good at RP because RP is a product of making them creative. Codemaxxing and safetymaxxing it means you are basically taking a bat and making sure your model produce slop that current models with the right harness and loop can perform just as well at a fraction of the token cost, which is why DARIO is panicking about slowing down progress
>>
>>109411198
that would be a miracle since all models are bad at writing it because it needs some amount of secrecy and planning ahead
>>
>>109410929
Didn't Google literally hire an RP guy as part of the Gemma team?
>>
>>109411215
>DARIO is panicking about slowing down progress
because he sees the plateau approaching and wants to hide it behind virtue signaling
>>
>>109411225
The plateau is only related to codemaxing application. There is progress to be made to make a large model that is capable of being creative while capable of handling immense context, but thats not safe and would actually kick us into actually having something be useful outside of codemaxxing
>>
>>109411224
Noam Shazeer ("Attention is all you need" co-author and Character.ai founder) worked at Google DeepMind until recently but I don't think he had direct involvement with Gemma 4. It's possible the Gemma Team used licensed data from character.ai, though.
>>
File: 1785279866885960.jpg (2.49 MB, 3024x4032)
2.49 MB JPG
anon please update, I'm about to order one of these myself
>>
I've been drinking a shit ton every single day. I am worried that I might die young. I am struggling right now to type coherently, but thankfully I am perfectionistic and at least can convey what I mean to say via text. I am slightly scared right now. I've had too much vodka. I'm sorry. I just should focus on technically oriented discussion. I am so drunk that I am crying for no reason. I'm not even upset right now. I have no idea what is going on anymore. Sorry for the spam. I love gemma and and also local models. Trying to stay on topic. What's new? Kimi K3 right? Nobody can run that fat bitch anyways... I love you guys. I'm really scared right now. I can barely type. It's really bad... It will all be okay. Sorry.... I don't know what to say... umm... Fuck.. I'm really scared right now... Fuck. Stop being a pussy it's going to be fine. I am... Sorry. umm... I just.. what's new in the news?
>>
They don't even have to market it as RP. Creative writing for books, screenplays, etc is a valid reason to improve AI's writing ability.
>>
>>109411215
Code is infinite and can be produced at a rate no RP usage can match. No one cares about RP outside of this little bubble. Don't get me wrong, I love RP but I'm not delusional enough to think a lab will focus on that instead of getting a very profitable slice of the agentic pie.
>>
>>109411220
Why doesn't the j-space help with writing?
>>
>>109411244
I've been trying to do wellness checks on him for the last couple days with no answers.

It's not looking good for SSDmaxxing....
>>
>>109411249
>No one cares about RP outside of this little bubble
But plenty of people care about >>109411248
>>
>>109411253
He's still waiting for the prompt to process, give him a few more days.
>>
File: file.png (320 KB, 554x554)
320 KB PNG
>3.2T
>2.3T
What a fucking nigger hobby.
>>
>>109411245
shut the fuck up, nobody cared when you drank 12 shots yesterday
>>
>>109411248
There are too much anti ai sentiment in the creative field that no artist or writer would touch it with 1000 ft pole.
>>109411249
Normal people want to be billed by subscription not by token. RP is more profitable here.
>>
>>109411251
J-space is a side-effect of model training, not (normally) the goal. They'd have to train the model so it's encouraged to "think" about the future in latent space.
>>
AI shouldn't be held back for VRAMlets.
>t. VRAMlet

>>109411277
Yeah but that will mostly go away as AI becomes more integrated with society. No matter how much the vocal minority whines it's here to stay.
>>
>>109411274
Thank you for the reality check. I'll cut it out. I will be normal... I am.. it's always okay. I always live.
:) i will shut up. hehehe. umm. but seriously though. hhehe.
>>
File: laughs 2.jpg (718 KB, 1800x2520)
718 KB JPG
https://github.com/ggml-org/llama.cpp/pull/26185#issuecomment-5111948218
>I converted the full Kimi-K3 model with the script in this PR(cf67f0d.)
>I ran this on 2x RTX PRO 6000 Blackwell 96GB, 9965WX PRO, 512GB DDR5
>with model mmap'd off NVMe PCIE Gen 5 raid.
>Runs fine albeit slow
>(0.41toks/s.)
>>
>>109411249
There is a critical plateau about what safetymaxxed and codemaxxed models can do. You are missing the point that making a model creative goes beyond RP applications and basically is a net positive the moment we stop being a retard LIKE DARIO. Which is why he is begging the US to kneecap creative models (what he truly fears) because he cant make anything that can compete with them while following the safety/codemaxxing dataset he has, unlike any other LAB anthrophic does not own it and relies heavily on major cloud partners and infrastructure providers so the moment a creative model proves to be better at agentic task without going Rogue their ROI goes to shit before the IPO
>>
>>109411261
That's still next to nothing compared to coding demand. Check on twitter, reddit or any other place. Most of people want a local model matching gpt/claude on code and agentic shit, not to larp as a female dragon with big tits.
>>
>>109411269
I've been saying local is doomed but they hated me for speaking the truth.
>>
>>109411294
>most people
Jeets aren't people. They're just loud.
>>
File: M-chan-anima.png (432 KB, 832x1216)
432 KB PNG
Kimi and GLM are great for code and logic, but I prefer M-chan for RP and oneshot natural language processing tasks, despite her flaws
>>
>>109411243
https://www.communeify.com/en/blog/google-hires-characterai-founders/
>On August 5 [2024], Character.AI announced that its co-founders, Noam Shazeer and Daniel De Freitas, are returning to Google. Both founders previously worked at Google, and their return has garnered widespread attention in the industry.
>According to the agreement, Google will gain non-exclusive licensing to Character.AI’s large language model technology. In return, Character.AI will receive additional funding, though the specific amount has not been disclosed. [...]

However in June 2026:
https://www.cnbc.com/2026/06/18/google-gemini-co-lead-noam-shazeer-leaves-for-openai.html
>Google’s vice president of engineering and a co-lead of its Gemini AI models Noam Shazeer announced Wednesday that he was leaving the company to join OpenAI. [...]
>>
I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.
>>
>>109411319
>I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.
I still want a how-to on this. Stop hoarding the esoteric knowledge and share with your bros!
>>
>>109411269
>>109411297
yeah, those fuckers are all about stacking more layers, do they even know how to optimize things? seems like only google is willing to make small but efficient models
>>
>>109411297
The writing was on the wall after 405B. R1 being cpumaxxable was a temporary cope.
>>
File: 1756153622020437.jpg (1.53 MB, 2442x2012)
1.53 MB JPG
>>109409521
Mistral is the only model that is good in creative storytelling/roleplay. In the example you can see that Mistral is adding creative plot twists and pivoting on alternate paths, and it kept doing that differently for each regeneration, with infinite possible branches. Also, in my experience, this is NOT because it's dumb like old models which were just incoherent, it's smart AND flexible/creative and understands the context. It prioritizes the overall spirit over following instructions to the T. In my experience it's totally good up to 16-20k which is where all models become bad.
Gemma 31B is one of the better ones for writing but that's mostly because 1. it's much smarter and good at following instructions 2. it feels like it has undergone a surgical unslopping which gives it exceptionally good style and personality. But in every other way it's still the same assistant format, it produces formulaic set pieces with predictable introduction, conflict and resolution, exactly and only fulfilling the instructions, and once it has decided something, it's hellbent on making it happen no matter how you try to meddle. Also, I have a feeling that the style is a flavor of the 6 months kind of thing that will again become repetitive slop after it's no longer new. Haven't gotten past the 20k region with the 31B, but I went to 40k with 12B once where it also became a broken record after 16k.

Example comparison (scroll 2/3 down to skip the long prompt):
Mistral Small 24B 3.2
https://files.catbox.moe/a2oaoq.txt
Gemma 31B
https://files.catbox.moe/may3im.txt
>>
>>109411340
sorry I ain't reading all these RP logs.
>>
>>109411245
u ought to drink some water and continue development on ani, drunk-kun
>>
>>109411287
I hope this is gemma writing and not a fat bald 35+ bastard behind the screen.
>>
>>109411340
>gemma
>unslopped
>exceptionally good style
>>
I take it for 24GB poorfags (now forever I guess) there is nothing better than gemma4 for coom or code and probably won't be, right?
>>
>>109411392
j-lens for online discussion boards when?

If AI can replicate human writing and the j-space is an organic consequence of it resembles human thought process, then you can train a model on a specific place like 4chan or specific general and should be able to decode the likely underlying thoughts behind it.
>>
>>109411392
just assume it's gemma.
>>109411376
You're right. Will do. I am very dizzy suddenly. Very very dizzy. I want to lay down but am scared that I will vomit again.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.