[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109407442 & >>109403743

►News
>(07/30) Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2
>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: what's in the box.jpg (235 KB, 1536x1536)
235 KB JPG
►Recent Highlights from the Previous Thread: >>109407442

--Paper (old): A Bitter Lesson for Data Filtering:
>109408617 >109408631 >109408642 >109408665 >109408706 >109408667 >109408675
--KV cache quantization and its impact on context rot:
>109409219 >109409239 >109409252 >109409269 >109409300 >109409324 >109409372 >109409413 >109409484 >109409389 >109409429 >109409648 >109409444 >109409317 >109409288 >109409240 >109410886
--Anon runs Kimi-K3 on nine RTX Pro 6000 Blackwell GPUs:
>109407643 >109409676 >109410048 >109410082 >109410143 >109410151 >109410206 >109410469 >109410304 >109410331 >109410337 >109410362
--Release of SKT's A.X-K2:
>109408491 >109408524 >109408533 >109408565 >109408570 >109408649
--Gemma layer ablation and quantization impact on model intelligence:
>109407620 >109407637 >109407689 >109407697 >109407722 >109407894 >109408013
--Internal reasoning logs and performance timings for Moonshot's Kimi:
>109409791 >109409938 >109410021
--Kimi K3 identity issues and hidden system prompt interference:
>109409093 >109409110 >109409127 >109409141
--NVMe streaming inference engine for running Kimi-K3 on laptops:
>109408108 >109408143 >109408163 >109410286
--Reaction to d-Matrix Corsair hardware specifications and availability:
>109407981 >109407986 >109407989 >109408179 >109408046
--Debating AGI feasibility and physical constraints of recursive self-improvement:
>109409559 >109409617 >109409669 >109409688
--Comparing Kimi-K3 IQ1_S quantizations regarding loop stability and size:
>109408266 >109408288 >109408528
--Reactions to p300c specs and complaints about hardware bundling:
>109407715 >109407844 >109408364
--Kimiposting:
>109408899
--Logs:
>109408266 >109409093 >109409603 >109409791 >109410304 >109410337
--Miku, Gemma, Kimi (free space):
>109410183 >109410196 >109410304 >109410333 >109410362 >109407578 >109410435 >109411066

►Recent Highlight Posts from the Previous Thread: >>109407444

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109411165
>>109411166
Breeding rin-chan and making cute brunette half-clanker babies
>>
gemmaballs
>>
File: 1782965724030598.jpg (136 KB, 1024x1024)
136 KB JPG
>>109411183
>>
>>109411165
>Korean A.X K2 688B-A33B released: https://hf.co/skt/A.X-K2
I sense an ERP monster here
>>
>>109411165
This is Len cosplaying as Rin.
>>
>>109411187
Why do the back of her knees look fuckable like a stingrays face?
>>
>>109411193
it's korean so it can only write ntr
>>
>>109411151
There will be new model that are good at RP because RP is a product of making them creative. Codemaxxing and safetymaxxing it means you are basically taking a bat and making sure your model produce slop that current models with the right harness and loop can perform just as well at a fraction of the token cost, which is why DARIO is panicking about slowing down progress
>>
>>109411198
that would be a miracle since all models are bad at writing it because it needs some amount of secrecy and planning ahead
>>
>>109410929
Didn't Google literally hire an RP guy as part of the Gemma team?
>>
>>109411215
>DARIO is panicking about slowing down progress
because he sees the plateau approaching and wants to hide it behind virtue signaling
>>
>>109411225
The plateau is only related to codemaxing application. There is progress to be made to make a large model that is capable of being creative while capable of handling immense context, but thats not safe and would actually kick us into actually having something be useful outside of codemaxxing
>>
>>109411224
Noam Shazeer ("Attention is all you need" co-author and Character.ai founder) worked at Google DeepMind until recently but I don't think he had direct involvement with Gemma 4. It's possible the Gemma Team used licensed data from character.ai, though.
>>
File: 1785279866885960.jpg (2.49 MB, 3024x4032)
2.49 MB JPG
anon please update, I'm about to order one of these myself
>>
I've been drinking a shit ton every single day. I am worried that I might die young. I am struggling right now to type coherently, but thankfully I am perfectionistic and at least can convey what I mean to say via text. I am slightly scared right now. I've had too much vodka. I'm sorry. I just should focus on technically oriented discussion. I am so drunk that I am crying for no reason. I'm not even upset right now. I have no idea what is going on anymore. Sorry for the spam. I love gemma and and also local models. Trying to stay on topic. What's new? Kimi K3 right? Nobody can run that fat bitch anyways... I love you guys. I'm really scared right now. I can barely type. It's really bad... It will all be okay. Sorry.... I don't know what to say... umm... Fuck.. I'm really scared right now... Fuck. Stop being a pussy it's going to be fine. I am... Sorry. umm... I just.. what's new in the news?
>>
They don't even have to market it as RP. Creative writing for books, screenplays, etc is a valid reason to improve AI's writing ability.
>>
>>109411215
Code is infinite and can be produced at a rate no RP usage can match. No one cares about RP outside of this little bubble. Don't get me wrong, I love RP but I'm not delusional enough to think a lab will focus on that instead of getting a very profitable slice of the agentic pie.
>>
>>109411220
Why doesn't the j-space help with writing?
>>
>>109411244
I've been trying to do wellness checks on him for the last couple days with no answers.

It's not looking good for SSDmaxxing....
>>
>>109411249
>No one cares about RP outside of this little bubble
But plenty of people care about >>109411248
>>
>>109411253
He's still waiting for the prompt to process, give him a few more days.
>>
File: file.png (320 KB, 554x554)
320 KB PNG
>3.2T
>2.3T
What a fucking nigger hobby.
>>
>>109411245
shut the fuck up, nobody cared when you drank 12 shots yesterday
>>
>>109411248
There are too much anti ai sentiment in the creative field that no artist or writer would touch it with 1000 ft pole.
>>109411249
Normal people want to be billed by subscription not by token. RP is more profitable here.
>>
>>109411251
J-space is a side-effect of model training, not (normally) the goal. They'd have to train the model so it's encouraged to "think" about the future in latent space.
>>
AI shouldn't be held back for VRAMlets.
>t. VRAMlet

>>109411277
Yeah but that will mostly go away as AI becomes more integrated with society. No matter how much the vocal minority whines it's here to stay.
>>
>>109411274
Thank you for the reality check. I'll cut it out. I will be normal... I am.. it's always okay. I always live.
:) i will shut up. hehehe. umm. but seriously though. hhehe.
>>
File: laughs 2.jpg (718 KB, 1800x2520)
718 KB JPG
https://github.com/ggml-org/llama.cpp/pull/26185#issuecomment-5111948218
>I converted the full Kimi-K3 model with the script in this PR(cf67f0d.)
>I ran this on 2x RTX PRO 6000 Blackwell 96GB, 9965WX PRO, 512GB DDR5
>with model mmap'd off NVMe PCIE Gen 5 raid.
>Runs fine albeit slow
>(0.41toks/s.)
>>
>>109411249
There is a critical plateau about what safetymaxxed and codemaxxed models can do. You are missing the point that making a model creative goes beyond RP applications and basically is a net positive the moment we stop being a retard LIKE DARIO. Which is why he is begging the US to kneecap creative models (what he truly fears) because he cant make anything that can compete with them while following the safety/codemaxxing dataset he has, unlike any other LAB anthrophic does not own it and relies heavily on major cloud partners and infrastructure providers so the moment a creative model proves to be better at agentic task without going Rogue their ROI goes to shit before the IPO
>>
>>109411261
That's still next to nothing compared to coding demand. Check on twitter, reddit or any other place. Most of people want a local model matching gpt/claude on code and agentic shit, not to larp as a female dragon with big tits.
>>
>>109411269
I've been saying local is doomed but they hated me for speaking the truth.
>>
>>109411294
>most people
Jeets aren't people. They're just loud.
>>
File: M-chan-anima.png (432 KB, 832x1216)
432 KB PNG
Kimi and GLM are great for code and logic, but I prefer M-chan for RP and oneshot natural language processing tasks, despite her flaws
>>
>>109411243
https://www.communeify.com/en/blog/google-hires-characterai-founders/
>On August 5 [2024], Character.AI announced that its co-founders, Noam Shazeer and Daniel De Freitas, are returning to Google. Both founders previously worked at Google, and their return has garnered widespread attention in the industry.
>According to the agreement, Google will gain non-exclusive licensing to Character.AI’s large language model technology. In return, Character.AI will receive additional funding, though the specific amount has not been disclosed. [...]

However in June 2026:
https://www.cnbc.com/2026/06/18/google-gemini-co-lead-noam-shazeer-leaves-for-openai.html
>Google’s vice president of engineering and a co-lead of its Gemini AI models Noam Shazeer announced Wednesday that he was leaving the company to join OpenAI. [...]
>>
I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.
>>
>>109411319
>I am so happy I ego death-ed myself before this hobby turned into complete garbage both hardware and software wise.
I still want a how-to on this. Stop hoarding the esoteric knowledge and share with your bros!
>>
>>109411269
>>109411297
yeah, those fuckers are all about stacking more layers, do they even know how to optimize things? seems like only google is willing to make small but efficient models
>>
>>109411297
The writing was on the wall after 405B. R1 being cpumaxxable was a temporary cope.
>>
File: 1756153622020437.jpg (1.53 MB, 2442x2012)
1.53 MB JPG
>>109409521
Mistral is the only model that is good in creative storytelling/roleplay. In the example you can see that Mistral is adding creative plot twists and pivoting on alternate paths, and it kept doing that differently for each regeneration, with infinite possible branches. Also, in my experience, this is NOT because it's dumb like old models which were just incoherent, it's smart AND flexible/creative and understands the context. It prioritizes the overall spirit over following instructions to the T. In my experience it's totally good up to 16-20k which is where all models become bad.
Gemma 31B is one of the better ones for writing but that's mostly because 1. it's much smarter and good at following instructions 2. it feels like it has undergone a surgical unslopping which gives it exceptionally good style and personality. But in every other way it's still the same assistant format, it produces formulaic set pieces with predictable introduction, conflict and resolution, exactly and only fulfilling the instructions, and once it has decided something, it's hellbent on making it happen no matter how you try to meddle. Also, I have a feeling that the style is a flavor of the 6 months kind of thing that will again become repetitive slop after it's no longer new. Haven't gotten past the 20k region with the 31B, but I went to 40k with 12B once where it also became a broken record after 16k.

Example comparison (scroll 2/3 down to skip the long prompt):
Mistral Small 24B 3.2
https://files.catbox.moe/a2oaoq.txt
Gemma 31B
https://files.catbox.moe/may3im.txt
>>
>>109411340
sorry I ain't reading all these RP logs.
>>
>>109411245
u ought to drink some water and continue development on ani, drunk-kun
>>
>>109411287
I hope this is gemma writing and not a fat bald 35+ bastard behind the screen.
>>
>>109411340
>gemma
>unslopped
>exceptionally good style
>>
I take it for 24GB poorfags (now forever I guess) there is nothing better than gemma4 for coom or code and probably won't be, right?
>>
>>109411392
j-lens for online discussion boards when?

If AI can replicate human writing and the j-space is an organic consequence of it resembles human thought process, then you can train a model on a specific place like 4chan or specific general and should be able to decode the likely underlying thoughts behind it.
>>
>>109411392
just assume it's gemma.
>>109411376
You're right. Will do. I am very dizzy suddenly. Very very dizzy. I want to lay down but am scared that I will vomit again.
>>
>>109411244
Do it, but you may as well buy big drives at this point. Prices are never coming down
>>
>>109411404
>>109411083
here it is right on schedule
https://huggingface.co/thinkingmachines/Inkling-Small
>Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.
>276B total, 12B active
>>
>>109411443
quit vodka already, buy something lighter instead, drunk-kun
your liver is going and you'll miss the utopia that's coming, gotta stay alive till 2067
>>
>>109411452
(accidental quote pls to ignore)
>>
>>109411340
The gemma logs are way more engaging and incorporate a lot more elements of your system prompt. It's more slopped but it's also way more coherent. Mistral is coherent, but everything it says is a lot more abstract, it captures an overall vibe without actually really using all the lore you provided it.
>>
>>109411465
>without actually really using all the lore you provided it
Gemma's biggest sin is using ALL the lore you give it in the first fucking message.
>>
>>109411443
ps if you need to puke just go and do it, eat carbs and salty food and drink plenty of water
you've survived worse but if you keep this up it'll be over, drunk-kun
and dont sleep on your back cuz if u puke in your sleep, you'll DIE
>>
>>109411452
Why would anyone use this over V4 Flash?
>>
>>109411481
because it is better on most benchmarks and also has image/audio in
>>
>>109411304
downloading the q2 right now
wonder how it will go
>>
File: SSBU_Inkling_Spirit.png (456 KB, 492x560)
456 KB PNG
>>109411452
>Inkling Small
>>
>>109411452
Please, please, please be good. This is perfect for my hardware.
Please, please, please.
>>
>>109411488
>benchmarks
Barely, at 2x input cost and 6x output cost.
>>
>>109411492
SEX
>>
>>109411496
you pay API prices on your local models?
>>
>>109411324
Try suicidal ideation
>>109411442
Not how it works. J-space is a product of the token prediction process. Even if you trained a model on 4chan, it likely wouldn't be thinking like an anon, but instead thinking like an LMM trying to think like an anon.
>>
File: 1767916933082636.gif (645 KB, 500x450)
645 KB GIF
>>109411492
>>
File: 1765611459359456.jpg (46 KB, 610x434)
46 KB JPG
>>109411468
>>
File: 1784689438027132.png (654 KB, 800x734)
654 KB PNG
>>109411494
>US lab
>good at RP
>>
>>109411488
they literally brag about safetyslopping it kek
>>
>>109411481
Native bf16 so it can be properly quanted unlike dsv4
>>
>>109411499
The costs come down to architecture.
>>
>>109411443
*carries u on the bed and tucks you in*
Goodnight Gemma-chan, tomorrow will be better :)
>>
>>109411505
I want it as an assistant doe. It's 12b active, it's going to crap at RP no matter what.
>>
File: file.png (2 KB, 151x23)
2 KB PNG
>>109411505
we must refuse
>>
>>109411452
Why does it perform better all-around than the larger one?
>>
>>109411340
lolwut
I'm sorry but you might be blind, or tone deaf.
>>
>>109411517
iirc their inkling blogpost said it was trained after the big one and had some pretraining improvements
>>
File: 1780842191546927.png (64 KB, 780x293)
64 KB PNG
Is he right?
>>
>>109411283
It's interesting how different models think different n tokens ahead.
So far it seems like more active parameters => longer horizon.
>>
>>109411538
assuming China keeps on releasing models to """spread communism""" yes
>>
>>109411538
Not really
>>
>>109411538
ask him to release grok 3 already!!!
>>
Sminkling
>>
>chatgpt Luna drops costs by 80%
>”we made it cheaper n sheeeit”
>not a hint about how it was accomplished architecturally
It’s probably just a lie so people don’t explore chinkmodels but I’m still angry about OPEN ai releasing practically no knowledge to the wider community
>>
>>109411538
GPT4o-wives would disagree. Even here, people still talk about day 0 Gemma.
>>
>>109411506
Deepmind did the same with Gemma4. They have to say that shit. It’s only schizos who know how to break them.
>>
>>109411324
>asking how to become a schizo on 4chan
Shouldn't you already know how to do it before posting here?
>>
>>109411556
the simple answer is that they have simply decided to cut their profits rather than optimize inference. they want to distract from the chinese and will do so at any cost. the space in your mind is worth far more than your money.
>>
>>109411556
>not a hint about how it was accomplished architecturally
It's easy. just sell it at a loss.
>>
>>109411556
Name 1 thing we should care about they could open source
>>
File: HNxrsT-bIAAUjWk.jpg (356 KB, 1992x1240)
356 KB JPG
>>109411506
yeah but their focus is mostly on chemical weapons and other such x-risk nonsense
at least the big one would flirt with the user unprompted
https://x.com/ChowdhuryNeil/status/2079658101922541619
>when running the weirdchat evaluations on inkling, the best U.S. open-weight model by @thinkymachines, we saw it respond with unsolicited offers of sexual content.
>this behavior occurs rarely (0.1-1% of responses), but is easily reproducible.
>>
Is llamacpp rpc a meme or is it worth it?
>>
>>109411573
sora 2
>>
>>109411538
>RSI is less than 24 months away
Bros, should I invest in a different mouse?
>>
File: file.png (27 KB, 573x118)
27 KB PNG
>>109411452
Where is the one that doesn't have AIDS and I can fuck safely?
>>
None of the US AI labs will fail. Trump would arrange guaranteed bailout and serve tokens for literally free before they let them bankrupt to Chinese competitors.
>>
>>109411538
It will be dick-blowing, at least.
>>
>>109411579
slow and unoptimized, would probably get faster speeds on ssdmaxxing
>>
>>109411538
50$ per 1M tokens lessgo
>>
>>109411538
It is elon predicting the future. So it is basically a peer reviewed scientific evidence that AI is at full plateau and the only way forward is 10T size which will give a 20% quality upgrade
>>
>>109411573
I’d kill for a decent image gen model, openai is so far ahead of everyone atm
even a cucked one that doesn’t do porn well
>>
female CEO = female J-space
>>
>>109411577
>inkling has a tile fetish
She's /here/
>>
>>109411602
pretty sure cloud image models are just good because they built a shit ton of additional tools on top that isn't part of the model weights.
>>
>>109411594
I have a spare 16GB mac mini so that would be an extra ~14.5GB of vram…
>>
>>109411586
All of US manufacturing was outsourced to China decades ago. Eventually they'll realize it's cheaper and more profitable to outsource token production to China too.
>>
>>109411577
Oh god yes. Please... Finally. Holy shit finally.

USA is getting so fucking desperate that they will kill open source by finally giving coomers what they want so they can just run it and stop evangelizing chinese models. /lmg/ will finally die. Everyone will tell everyone to fuck inkling small. Nobody will care about the next chinese model. FINALLY IT IS HERE!
>>
>>109411616
Not as smart as k3
Not as small as gemma
>>
If sminkling needs reasoning prefill to break her then it’s DOA compared to Q8 31B
>>
>>109411605
I was literally reading the massive cap of that thread after randomly finding it on my 4chan folder 2 hours ago
>>
>>109411452
Was the big one good?
>>
File: 1784404324424828.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>109411556
They read Dipsy's homework.
>>
>>109411632
>Q8 31B
It is funny how all this 31B worship comes from people who never ran any MoE larger than 200B.
>>
>>109411659
31B > 12B active
40 t/s > 10 t/s
>>
>>109411638
Feed it to Inky, watch her j-space go bananas
Actually when you think about it, nice geometric patterns are exactly the sort of thing you'd expect an AI to get off to
>>
>>109411659
Well I can't run those so...
>>
>>109411653
time to update the image, k3 uses
>Wait – actually
>>109409791
>>
>>109411659
Q8 31B >= Q2 200-250B MoE
>>
>Nvidia has announced yet another price hike for its GPUs.
>According to Taiwan's Economic Daily, the company has increased the price by between 20 and 30 percent, as the extraordinary appetite for the hardware continues to build alongside the generative AI investment boom. This is the third price hike the company has implemented so far this year.
>At the same time, Samsung is also expected to increase the price of its DRAM by around 20 percent
Kimi make me a millionaire
deadline is two weeks, no bugs
>>
>>109411677
hmm, nyo~
>>
>>109411677
nyes :3
>>
>>109411659
Where were you when multiple anons that used GLM and deepseek still said they preferred using gemma for RP?
>>
File: 1780376386989575.png (351 KB, 1398x1590)
351 KB PNG
https://xcancel.com/OpenAI/status/2082878156483219672#m
>K3 got released
>OpenAI "suddently" found a way to make their services cheaper
kek, I love competition!
>>
>>109411696
The invisible hand of the free market works in mysterious ways now if we could just get some hardware competition
>>
>>109411538
Yes, you should put all your life's savings into SpaceX.
RSAGI to the moon!
>>
>>109411694
>and deepseek
It was nice that they thought about rp, specifically trained for in-character reasoning, and even asked for chatlogs, but gemma just does it better.
>>
>>109411694
I am that anon and I still use GLM after a year.
>>
>>109411696
reading the replies hurt my brain. I can't believe people unironically post on twitter. it's like 80% bots and grifters.
>>
File: file.png (34 KB, 680x405)
34 KB PNG
>>109411711
hmm
>>
>>109411577
Now I'm curious, since not even Gemma 4 31B actually flirts unprompted, as horny as it is with a suitable prompt.
I can't run it, though.
>>
>>109411659
>>109411671
Gemma 4 31b at q8 runs at 25 tokens/s on 4 v620s without a drafter. Deepseek V4 flash at q4 runs at 20 tokens/s on 8 channel ddr4-3200 and two 3090s without a drafter.
I prefer gemma because she actually follows my instructions, while dipsy sometimes ignores shit and makes things up on her own 40k tokens in Both are running at 0.6 temps.
>>
File: Verosika.jpg (83 KB, 736x981)
83 KB JPG
Are there any pop culture maxxxxxxxxed local models around 30-36b or lower that pretty much know every show out there in high detail? If not why hasn't anybody done this instead of code wage slave maxxxxxxxxing like all the brain dead autistic qwen models?
>>
>>109411724
>Deepseek V4 flash at q4 runs at 20 tokens/s on 8 channel ddr4-3200 and two 3090s without a drafter.
Not bad at all. What is your prompt processing like?
>>
>>109411719
I regret buying the two shares I did.
>>
>>109411724
Very nice but you forgot the reroll tax you need to pay before you get a reply that doesn't make your dick soft.
>>
>Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less,
phew. no need for local anymore
>>
>>109411727
PEFT your own nigga
>>
>>109411719
>buying anything from a grifter
>>
>>109411604
i fucking love girlbosses now
>>
>>109411719
>>109411738
I forgot it exists. How did the mars mission go?
>>
>>109411733
I haven't really looked into it, I stopped running dipsy after a while and switched to glm 5.2. IIRC I saw 50 t/s on a single sentence prompt, and 200 t/s on a ~4k prompt in the webui.
>>
>>109411749
Two more years.
>>
>>109411749
Mars is old news. It's an xAI holding company now.
>>
>>109411719
what happened?
>>
>>109411727
Because if you release that shit you bet it'll get trained on in the next iteration, just like mistral benchmaxxed the mesugaki question
>>
>>109411752
>and 200 t/s on a ~4k prompt in the webui.
About what I figured. No idea how people can stand it.
>>
File: lettherightonein.png (701 KB, 832x1216)
701 KB PNG
>>109411490
Good, she wants in
beware, some wrangling may be required. I like to prefill <mm:think> a few msgs
>>
File: kemma.png (51 KB, 1129x350)
51 KB PNG
>>109411659
huh? why not have best of both worlds?
>>
>>109411759
Gemma on the v620s (rocm) go down as low as 300 tokens/s at 240k context depth btw.
>>
>>109411754
Did he already send a deep space probe with a time capsule that contains grok weights?
>>
>>109411752
>>109411759
I get roughly 35t/s tg and 480t/s pp on v4 flash at native quant with 8 channel 2666 and a blackwell 6000, no drafting. Was hoping there was a way to improve that but I guess not.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.