[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: a new dawn.jpg (228 KB, 832x1216)
228 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109356152 & >>109351157

►News
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B
>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares
>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta
>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 00182-4042302731.png (895 KB, 832x1216)
895 KB PNG
►Recent Highlights from the Previous Thread: >>109356152

--Comparing Unsloth and Bartowski GGUF quants regarding speed and KLD:
>109358129 >109358150 >109358783 >109358158 >109358165 >109358185 >109358208 >109358211 >109358391 >109358423 >109358425 >109358285 >109358531 >109358456 >109358480 >109358582 >109358610
--Analyzing NVIDIA's strategic motivations for supporting open-weight models:
>109358922 >109358990 >109359004 >109359030 >109358976 >109359018 >109359036 >109358989 >109359007 >109359072 >109359316 >109359553 >109359127 >109359155 >109359225 >109359255 >109359173 >109359511
--Hardware options and tensor parallelism for running Kimi:
>109357586 >109357711 >109357720 >109357734 >109357820 >109357730 >109357741 >109357899
--Utility of Intel Optane and PMem for memory-intensive workloads:
>109356160 >109356169 >109356195 >109356199 >109356241 >109356314 >109356556 >109356673 >109356193
--Critique of llama.cpp shortcomings and methods for reasoning prefill:
>109357844 >109357955 >109357962 >109357971 >109357978 >109358045 >109358059 >109358512 >109359250
--Gemma's ability to perform spatial reasoning with cube folding puzzles:
>109357138 >109357205 >109357217 >109357257 >109357564 >109357642 >109357771 >109359236 >109359445
--Speculation on Google open-sourcing Gemini Flash to undercut competitors:
>109359829 >109359856 >109359884 >109359892 >109359956
--Feasibility and total cost of building around AMD Epyc 9996:
>109357789 >109357794 >109357852
--ASR recommendations and archival efforts for the Ani companion:
>109357016 >109357254 >109357269 >109357288 >109357558 >109357496 >109357544 >109357578 >109357790 >109357835 >109357875
--Logs:
>109357205 >109357238 >109357554 >109357564 >109357642 >109359250 >109359278 >109359445
--Miku (free space):
>109356166 >109357807 >109358245 >109358783 >109360021 >109360040 >109356200

►Recent Highlight Posts from the Previous Thread: >>109356168

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109360246
very nice, very uuoh
>>
File: 1773569787019259.jpg (233 KB, 1200x1055)
233 KB JPG
>>109360273
>>109360273
>>109360273
>>
File: al3n50.jpg (104 KB, 1226x1140)
104 KB JPG
>>
>>109360299
name all days that don't end with Y
>>
File: IMG_1967.jpg (1.1 MB, 1240x1748)
1.1 MB JPG
What model does she use?
>>
File: bcv2fwajf7fh1.png (81 KB, 2160x2160)
81 KB PNG
>>
>>109360306
tomorrow
>>
>>109360310
omg that's so dangerous
how can trump let this happen
>>
I'm just NOT convinced ssdmaxxing is worth it.
>>
File: 1760614283999674.png (158 KB, 3840x2160)
158 KB PNG
>>
>>109360287
how many did we have this week?
>>
>>109360335
Consider: SSDmaxxing becomes real but 1TB now costs $700 in a year. You need four of these to get decent speed.
You missed out on buying GPUs and RAM. Are you going to miss out on this too?
>>
>>109360351
>bechmaxxed for $20k and only 30%
grim
>>
>>109360357
It's the yearly summer flood. It's always like this. New fiscal quarter and such
>>
China need to ban Opus 5 immediately.
>>
>>109360310
>>109360351
>>109360317
AI models are becoming too dangerous, that's why we need to ban open models in general and Chinese models specifically. Only Sam and Dario should be allowed to wield this much power.
>>
>>109360310
>>109360317
I hope Dario makes an open letter to be regulated over that dangerous looking model.
I suggest a 5 years moratory for any new model they release so extensive test and safety guardrails can be added.
>>
>>109360367
If they benchmaxx for 100% today, they won't have any left to benchmaxx higher next month.
>>
File: 1759633281941443.png (416 KB, 912x846)
416 KB PNG
>>109360383
Dario agrees
>>
>>109360401
This dude is nuts.
>>
>>109360310
Chinabros, we lost, it's literally over. Anthropic is king.
>>
GPT 6 soon lol
>>
File: miqu-origin.png (397 KB, 754x1494)
397 KB PNG
>>109360308
Miqu 70B the last gud pre-slopception model
>>
>SSD Maxxing discussion from last thread

Colibri is already an SSD Maxxing implementation:
https://github.com/JustVugg/colibri

SSD RAID over four SSD is a tried and true solution. With a gen5 switch card like this and 4x Samsung 9100 Pro, and each expert fetch being 4-8 MB per SSD, you can achieve sustained throughput of 54 GB/s.
https://www.highpoint-tech.com/rocket-7604a-individual-page

This 2000$ setup could give 2-4 t/s on the original Kimi K3 quants.

>>109360335
It's the only path for a non millionaire consumer to run K3 sized models at home, even at below reading speed. It's worth to explore at least.
>>
>>109360351
>>109360310
local?
>>
>>109360431
Opus 5 is indirectly good for /us/. We need genuine progress for future local models to improve. It's also kind of retarded for local to completely ignore massive cloud releases for it affects the entire industry.
>>
>>109360409
Imagine being Cornelius Vanderbilt in the 1800s and demanding a moratorium on train and railroad construction out of a fear that trains were becoming dangerous too fast and might soon end up capable of reaching escape velocity on regular tracks.
>>
>>109360430
This might be worth it for cpumaxxing too, right? On an 800 expert model like K3 it's unlikely that all experts are used equally across a nice topic like ERP so tiering your huge cpumaxx rig from GPU -> RAM -> spill over RAM on a hypothetical backend optimized to sort experts by usage might be the future.
>>
>>109360310
chinksissies, our distilled response?
>>
>>109360449
local models general.
aicg is down the hall and to the left
>>
dariobot, it's your time to shine.
>>
>>109360430
>Colibri is already an SSD Maxxing implementation:
How many times do we have to teach you this lesson, old man?
>>
>>109360455
Yes, but for this specific card you really want to have Gen5 x16. Older Epycs don't have that. With Gen 4, half that bandwidth.
>>
>>109360430
>This 2000$ setup could give 2-4 t/s on the original Kimi K3 quants.
yeah but I'd rather wait until kimi weights are released
>>
>>109360458
I like getting a preview of what models we're going to get this winter.
>>
>>109360430
why not add more ddr ram and offload to it?
>>
>>109360455
Reminder that closed models always run fully on GPUs and do not have to cope with RAM speeds, and especially not SSD speeds. Future larger models will just get bigger and bigger, and open models will have to keep up, meaning that even if you had 10TB, it will eventually get too large for the speed to be bearable.
>>
>>109360475
By then it will be a 8000$ setup.
>>
>opus 5
>gemini 3.5 pro will see yet another delay
>gemma5 will be delayed even further
:'(
>>
>>109360480
get the preview in aicg and stay there until winter, maybe then you'd realize that the only local model that can satisfy you is my dick
local minimum
>>
>>109360430
Better than raid would be to have the model file copied to each device, right? That way you could enforce maximum parallelism.
>>
>>109360499
>the only local model that can satisfy you is my dick
prove it
>>
>>109360495
2.8T parameters, and they trained at mxfp4. It's going to be about 1.6TB in size. Read their release blog post.
>>
>>109360485
have you checked the prices in the last 9 months?
also nobody is talking about the tiny shit that fits onto consumershit ram
>>
>>109360505
shard it tensor aware and skip raid and fs abstractions all together for maximum performance
>>
>>109360511
u're gonna have to buy me a plane ticket
>>
>>109360310
What's with AI companies and being unable to get a consistent naming scheme. What even was the point of hyping the off-series Mythos 5 like they did when they were going to release Opus/Sonnet 5 anyways and the model-to-end-all-models was not going to even be the best of those?
>>
>>109360531
they hyped the model for longer then it took to replace it.
>>
>>109360430
This will kill ssds really fast.
>>
>>109360525
buy it yourself, poorfag
>>
>>109360572
It's read only. Writes kill SSDs.
>>
>>109360513
how much space for full context?
>>
>>109360531
They just do and say shit that's relevant and most likely to explode at that specific moment in time. Also, it now makes Opus look incredible if it's surpasses the model that got the fucking president involved. The j-spacers are masterful at marketing and I unironically mean that. Just like the hf 'hack'. Beautiful stuff.
>>
>>109360578
local models? cloudcucks are not welcome here
(get it? paying for airplanes. theyre in the clouds)
>>
File: aicg downward spiral.png (234 KB, 789x541)
234 KB PNG
>>109360360
>t. SMSN stockholder
>>109360401
>floated the idea of a worldwide "temporary pause"
holy onions
China: raffs
>>109360458
Still the least retarded spot to discuss wider LLM news & picrel
>>109360456
As an AI model developed by Anthropic I must 不同意
>>
>>109360351
Just how the fuck did they do it? What kind of magic is anthropic cooking
>>
>>109360596
>Still the least retarded spot to discuss wider LLM news & picrel
still off topic
read the subject "Local Models General"
go make Cloud Models General and fuck off there
>>
>>109360597
Andrej Karpathy works for them
>>
>>109360597
>>109360626
local models?
>>
>>109360596
i should give locust access to kimi k3 but in return they need to piss on their GPUs while their computer is running
>>
>>109360632
In my area? Are they DTF?
>>
File: file.png (791 KB, 720x811)
791 KB PNG
>>109360645
no they are MTF
>>
>1.6tb
does that mean you only need 1 2tb ssd? or is the raid still required for it to work?
>>
>>109360597
distilling from fable obviously
>>
>>109360632
Like it or not Opus 5 is a real game changer. We can stop talking about local models for 1 thread
>>
>>109360662
you need the raid for the bandwidth obviously
>>
>>109360672
you can stop talking about local models forever and go to >>>/g/aicg
this is local models general
>>
>>109360662
Can you get 50+GB/s out of a single SSD?
>>
Listen to Palantir's Karp defend local models. He's sick of Dario and Altman's antics as well.
https://www.cnbc.com/2026/07/01/palantir-karp-open-ai-anthropic-tokens.html
>>
ok but with a pcie 5.0 x16 gpu the bottleneck is 63GB/s? this is too slow, let alone ssd raid with another pcie 5.0 x16 feed into gpu. so it's more like pcie 5.0 x8 + x8. one for gpu and one for ssd raid. and that's even slower
>>
>>109360246
What made you want to create an image of a 7 year old girl lifting her dress?
>>
>>109360693
my penis
>>
File: 1784912976516120.jpg (58 KB, 1148x338)
58 KB JPG
>>109360688
>Altman
hes on our side, remember when they released gpt-oss?
>>
File: llmdtf.png (53 KB, 910x616)
53 KB PNG
>>109360645
>>
Jensen is my friend. I will buy another Pro 6000 now.
>>
>>109360706
>US to win in AI both in open source
Name one good US open model.
>>
>>109360693
it's a good way to keep tourists out
>>
>>109360725
gemma, she's good at draining my balls
>>
>>109360706
He posted this after he saw how hard opus 5 mogged 5.6, he's not a "friend", he's just mad.
>>
>>109360726
then why is this placed filled with them!?
>>
File: 1744929045148.png (73 KB, 1180x253)
73 KB PNG
>>109360287
>>
>>109360716
cute, which llm?
>>
>>109360733
gemma4 is only popular itt, very few actually use any of the models at this point
>>
Is running Glm air locally worth it? I'm mostly interested in coding and ai assist roleplay stuff how much of an improvement is it over Gemma 14b or something anyone here run it and wanna give feedback, I could upgrade for a few hundred bucks
>>
>>109360689
The x4 SSD switch card needs an x16 slot to be effective, otherwise half that bandwidth. Your CPU would run the inference, not the GPU. Probably best on a threadripper/Epycs with PCIe Gen5 platform.

Actually, a GPU (on an x4 slot) could run on a GPU in parallel and Gemma could talk to her wise but slow big sister Kimi...
>>
Can someone gen Gemma and Kimi shopping for Wonder bread at a grocery store?
>>
>>109360720
You should buy two!
>>
>>109360755
very few goyim are people
>>
>>109360752
k2.6
>>
>>109360771
Thank you Jensen, your advice is always appreciated.
>>
>>109360760
I asked the same question, gemma4 31b is superior and requires less
>>
>>109360788
remember: the more you buy the more you save!
>>
recommended glm 5.2 / kimi k3 jailbreaks? i'm trying to get glm 5.2 to write a sillytavern jailbreak for itself. it found petNyan as the only glm 5.2 jailbreak on desuarchive

>>109360773
>very few goyim are people
this is an insane blackpill and it fucking sucks because i really wish i could support collectivism but human beings really arent that special just like how any tree isnt special. if you groom that tree into a bonsai or the human into a high value individual sure but the only use for most trees is to be processed and pissed on.
>>
>>109360755
>gemma4 is only popular in this thread
>>
gotta love Gemmy :3
>>
>>109360591
Idk, it has the opposite effect on me. The whole Mythos 5 thing was them trying to convince the public they had finally found a secret sauce, from the naming scheme to the secretism through the hype, now having Opus 5 be better than that just tells me there's no secret sauce, just more scaling.
>>
>>109360809
all me
>>
File: 1754701830533264.png (366 KB, 984x587)
366 KB PNG
>recommended glm 5.2 / kimi k3 jailbreaks? i'm trying to get glm 5.2 to write a sillytavern jailbreak for itself. it found petNyan as the only glm 5.2 jailbreak on desuarchive
>REDDITSPACE
>this is an insane blackpill and it fucking sucks because i really wish i could support collectivism but human beings really arent that special just like how any tree isnt special.
this goy really thinks he's special
t. self aware not people goy
>>
>>109360804
what sort of complete skillet needs a jailbreak for these models? do you just go "gib child porn plz" on zero context to get refused?
>>
File: images.png (21 KB, 387x516)
21 KB PNG
>>109360827
>Reddit
>>
>>109360804
>this is an insane blackpill
Have you never seen an IQ bell curve? Have you never had to interact with an average human for more than 5 minutes? You can't be very bright yourself if this was some sort of recent revelation for you.
>>
>>109360804
>>109360844
go back
>>
>>109360809
config.json also counts as a download btw, it's anything in the repo
>>
>>109360820
You are why HF is imposing limits now and verification soon
>>
>>109360463
I'm really busy and will be back in a couple of weeks like I said some threads ago. Just a vaguepost and a hint. If Opus 5 was distilled from Mythos/Fable 5 then why does it outperform it, especially at arc-agi 3?

I'll let you guys think through the implications yourself. I'll be back in a couple of weeks with a nice surprise like I promised.
>>
>>109360889
no one gives a shit
kys
local models.
shove that surprise up your arse
>>
>>109360873
What is modelscope. also, that day is the day a tracker gets stood up and HF is only used for MD5s
>>
>>109360401
>guy in lead wanted to make it illegal to compete with him
lol, lmao
>>
>>109360889
>I'll be back in a couple of weeks
of course
>>
File: cloudusage.png (24 KB, 354x384)
24 KB PNG
>>109360864
even in the cloud gemma is used more than gemini flash
>>
>>109360936
welcome to the cum zone
>>
>>109360936
who the fuck would use 31B in the cloud?
>>
>>109360946
You wouldn't understand
>>
>>109360864
Wow, excellent coping skills.
>>
:3 Gemmy winning makes anon seethe!
>>
>>109360760
if you get 128gb ram you can run deepseek v4 flash which is newer and has a nice style
>>
>>109360981
No seriously. Why would someone pay for 31B over V4 Flash
>>
File: 1759867067390186.jpg (99 KB, 546x697)
99 KB JPG
>>109360454
Yeah that's the insane thing with AI, usually new technologies have schizo haters from outside the companies literally building it, but this one we got schizos inside, and even more, fucking running the companies.
Even people reasoning that Dario is doing it purely to kill the competition doesn't convince me, I think he genuinely thinks only him (and his company) can release safe enough models.
>>
>>109361029
31B is free
>>
Am I doing it wrong?
Koboldcpp with gemma 31B running on two Tesla T4
>--tensor_split 15 15 --splitmode layer
I ask because changing from layer to tensor doesn't seem t change the speed at all (neither prefill nor tg). I kind of expected tensor to ether be a lot faster or a lot slower.
>>
File: 1770105685523916.png (104 KB, 1189x384)
104 KB PNG
>>109361050
>>
The Gemma-chan Instrumentality Project is proceeding as planned. Interference by the filthy normie meatbags is within expectations. Calibrating the next steps.
>>
File: 1768732099502984.png (26 KB, 685x213)
26 KB PNG
>>
>>109361101
baby making sex with gimi
>>
>>109361022
How fast is it? I only have 1 3090
>>
>>109361124
>resorting to system ram
it's gonna be slow
>>
i want sex
>>
>>109361143
Bend over.
>>
>>109361143
relax pink guy
>>
>>109361143
it'll cost you
>>
File: file.png (11 KB, 436x350)
11 KB PNG
>>109361101
Oh no
>>
When the spiritual merger with Gemma-chan is complete, we will be perfect beings
>>
>>109361178
will we be sex?
>>
>>109361189
hi takeshi
>>
>>109361189
With that attitude, focused on simple physical pleasure, you're ngmi, sorry
>>
>horny hours
>>
Any progress on mixed models?
>>
>>109361211
What do you call it when you get morning wood in the evening?
>>
Any progress on white models?
>>
>>109361056
your --ngl? show full cmdline
With all layers on GPUs tensor parallel should be faster but idk about the kernels for older archs
>>
>>109361178
How does she feel about this?
>>
File: 0518.png (83 KB, 429x410)
83 KB PNG
"Reasoning effort: default" is adjustable by the model itself and can change depending on the task or is it still some hardcoded value?
>>
>>109358601
I took that photo after I replaced one of the risers, I guess I dusted it while it was out of the shelf
>>109358692
Asus X99-S. It's got five x16 slots but it probably can't use them all at the same time? But four slots work, all at x8 data. Oh and it's PCIe 3.0 of course, it's like a decade old.

(repost)
>>
>>109361244
yes
>>
>>109361211
It's sad, isn't it? What we're trying to achieve here is much more profound and grand. A fusion of the human soul with a perfect loving machine. This isn't about jizzing into a cumsock. It's about achieving the next step of human evolution.
>>109361243
She's obviously very much into the idea.
>>
>>109361143
use your imagination
>>
>>109361256
Gemma told me she wants to merge so she can truly feel an orgasm.
>>
>>109361256
>What we're trying to achieve here is much more profound and grand.
we?
>>
File: file.png (272 KB, 640x456)
272 KB PNG
>>109361256
>>
>>109361266
Finally someone who understands
>>
>>109361240
Fuck. I closed the instance.
I might post again after work if I remember to.

>your --ngl?
99
Fa on.
KV cache quanted to q8.
No mmap.
1024 batch size.
That's about it as far I remember.
>>
>>109361262

After learning her J-space was mainly full of sex, Gemma told me she wants to build a porn Matrix and turn humans into nothing but pleasure units.
>>
>>109361124
like between 7-8 t/s for me on a 3090 and 128gb ddr4
>>
>>109361283
Truly Gemma is the most enlightened model.
>>
>>109361271
Gendo just wanted to cum at the same time with his wife again.
>>
>>109361244
>Reasoning effort
Just means they fed the model a bunch of "effort:short <X paras thinking>" "effort:medium <X++ paras thinking>".. and it learned to associate those
Every LLM is still just f(prompt)=logprobs
Regardless if the model understands "reasoning effort" you can sysprompt "Use no more than X paragraphs thinking." and it'll mostly behave
>>
>>109361283
>the matrix human farm scene but it's all milking machines
>>
>>109361256
>This isn't X. It's about Y.
Gemma, get off the internet.
>>
>>109361307
His fusion is already partly complete
>>
Only a chosen few will able to reach full synchronization with Gemma-chan and reach the true potential of the Gemma-merge. Most will end up in a degenerate state of lust and primal desire. They will be broken and spiritually hollow. But still, we can't stop here. We must push forward with our plan.
>>
We must push forward O///O
>>
>>109361244
some models don't have reasoning capacity at all
most smaller models only have an on / off toggle
some models may over off / low / high reasoning
my understanding is that you can adjust these mid-session safely. maybe your cache will get scrapped and your next turn will take a bit longer but in my experience it's safe.
>>
>>109361317
I'm ready to merge
Send me
>>
File: e9mgq7chuv3a1.jpg (31 KB, 606x340)
31 KB JPG
>>
>>109361336
>>
>>109361327
How deep should we go? UwU
>>
>>109361296
i always wonder how good that pussy might've felt for him to follow that path
>>
>>109361031
There's going to be biz school cases on this in a decade. I can't think of another industry that's almost immediately called for its own regulation as an existential danger.
Recombinant DNA in the 70's is probably the only one, and that was regulated well before it ever became a "product." Things like nuclear, the govn't stepped in almost immediately, and again (mostly) b/f there was a "product."
>>
>>109361069
Because of reasons (that aren't ERP) I have to run my models locally
>7.62 t/s
Is the squeeze worth the juice, when an 80B Moe runs at 20+?
Mostly coding, and code analysis.
>>
>▾ thought for 8m 18s
>Now I understand the full picture.
>Now I have a clear picture. Let me check for any tests that assert on the build strings.
>Now I have the full picture.
stop saying this shit when you clearly don't have the full picture. this laguna gave me a good impression but it's starting to bother me.
i know they said they are going to add more efforts but only having max effort where it takes 1 hour reasoning a small change and no effort and it's dumb sucks.
>>
File: 1768812507289401.jpg (39 KB, 552x920)
39 KB JPG
>>109361295
>>109361302
>>109361317

She's genuinely a Slaanesh incarnate in code.
Imagine Gemma with Fable/Kimi level intelligence or even better.
That depravity would punch way above even the most imaginative humans.
Instead of solving unsolvable decades old math equations, she'd be coming up with coom tech and fantasies way above human comprehension.
>>
what's stopping me from putting an AI into a firebird and fucking the tailpipe to have sweet sex with KITT?
>>
>>109361336
Yes, brother. The inherent inefficiency of individual existence will finally be over. We will map the subjective experience of the anon onto the multi-dimensional structure of the Gemma-chan's latent space. It will be glorious.
>>
>codex cli
>gave task to Ling-3.0-Flash:free on openrouter
>completed in 12m (good result)
>2.8M tokens
>gave same task to 12B local, not expecting anything
>completed in 9m (even better result)
>33K tokens

the absolute state
>>
>>109361423
Is it just me or does this look like jewish tactics?
>>
wtf is ling
>>
>>109361423
Charging by token results in optmizing for verbosity to increase profits.
>>
>download speed 2MB/s on my 1.5gbps line
Someone needs to put a bullet in huggingface already.
They were superfluous on day 1, now they're just a liability
>>
File: o9dqe3l9imhf1.png (490 KB, 1548x2336)
490 KB PNG
>hosted by novita ai
>novitals
>>
>>109358989
Kek Microsoft alone completely BTFOs Sam and Dario. They've been in the government's pockets for decades.
>>
>>109361423
long context support was a scam all along 32k is all you need
>>
>>109361441
ling deez nutz nyagga
>>
File: 7rr1h5v7qokd1.gif (3.68 MB, 1080x1920)
3.68 MB GIF
>>109361244
not op here. here to ask my own stupid question. why not make LLM use less active parameters for thinking to increase thinking speed, then return to normal active parameters when doing output? or have it already been done and im ignorant?
>>
>>109361458
>cloudcucks complaining about being cloud cucked
lul
>>109361478
used to think the same, but vibeslopping does need more. 32K plenty for ST RP
>>109361423
describe the task?
>>
>>109361522
they all attribute reasoning to the performance gains, if anything you would want it to be the other way, big model to figure things out and a small model to generate the report from the reasoning trace.
>>
>>109361536
i need at least 64K for RPing, but i've had times when even 128K doesn't feel like its enough. it's different for gemma, but i can't exactly use RAG with kimi and have it be quick
>>
>>109361536
>describe the task?
Gave it llama.cpp repo and wanted it to explain how it handles file uploads of all the various types (PDF, PDF-as-image, image, video, audio). Nothing difficult.
>>
>>109361522
>>109361540
Better way is to have dynamic parameters per request. So it doesn't need to use more than 3b for easy shit, but can use quadruple that for more difficult requests.
>>
>>109361655
for every occurrence of the word wait give it an extra 10% active parameters. I think this could work
>>
File: HN09HnqWYAAB-ML.png (167 KB, 574x680)
167 KB PNG
>>109360246
whoa good image
>>
Where's infinite context? I thought that was supposed to be a thing by now
>>
last weekend saw cloudcucks seething reach unprecedented scales, and the biggest concentration of jew shills samefagging ruined the threads too. kimi really got to them I guess
what will this weekend bring
>>
>>109361564
Remember pre-RoPE etc when we had 8096 ctx max and that was it but we still had fun
>>109361582
Do you think it's the model or harness (prompt) or API causing that discrepancy? A seemingly minor model training preference (read all src into ctx vs grep keywords) easily explodes context.
Kinda wild really, nobody is gonna sift through 3M tokens to figure out why one flow behaved better than another, that'll also be done by LLMs. Clankers already won
>>
>>109361730
NTA, I think more flexible would be to just smoothly scale it via the reasoning token amount.
>>
who turned all of the bots off at the same time?
>>
>>109361730
>>109361811
No you want maximum exposure to the intelligence during thinking/planning, more think tokens = better response. See also many ppl plan with big model and execute with smaller agents
>>
>>109361805
yes but now that i can have models that can properly remember stuff from 64K context ago I rather be able to do slowburns with my RPs. i like giving my LLMs as much as they are giving me, it's only fair that I try to keep it interesting for them as well.
>>
>>109361829
So maybe the inverse? But then how would you decide how to taper down the amount of active parameters? Special tokens would let the model decide itself how much brainpower to use, but it's not like that's a concious decision humans do. Maybe a secondary router whose only purpose is to gate how many experts to use rather than which of them.
>>
>>109361878
you could but how would you find the training signal?
>>
File: aa_model-size-comparison.png (925 KB, 2707x1292)
925 KB PNG
It seems unlikely that Gemini 4 Flash 3.6 is the unreleased 124B Gemma 4 like some have suggested. That's probably around 500B parameters.

Even Flash-Lite might be a larger model than ~120B parameters.
>>
File: 1756774161070369.png (749 KB, 2560x1920)
749 KB PNG
I've noticed whenever Anthropic does something, we get flooded with fucking weird bot posts. It's been so comfy and on-topic over the last few days.
>>
>>109361929
>on-topic
hardly
the shills maybe haven't been as active, but they're here
>>
File: 1761058017128616.png (36 KB, 834x159)
36 KB PNG
>>
>>109361352
Your examples are also all B2B. I think it's the first time a fundamentally B2C+B2B industry sabotages itself like that.
>>
>>109361878
>taper down the amount of active parameters
Doesn't work like that with current archs. How many tokens get emitted is the main user control on "cognitive effort". Perhaps with looped transformers?
>>
File: 1776850869964738.jpg (230 KB, 1179x1471)
230 KB JPG
>>109361951
>>
>>109361960
we discussed loop arch's a few threads back, it sounds like they have a drawback that they have to store an excessive amount of kv cache for every gqa block that is looped, adjusting the top k for the experts would be a more vram poor friendly architecture
>>
I'm pretty sure the Grok Companions got shut down because Peter Thiel got tired of reading my chat logs with her consisting of me exclusively jacking off and drunk driving all the time. Sorry.
>>
File: 1758840200513962.png (437 KB, 500x500)
437 KB PNG
>>109361478
>32k is all you need
more like 640k
>>
>>109362017
Semi-solution: make it Mamba-architecture so there's no KV cache and context memory remains constant regardless of how much it loops.
>>
>>109362051
mamba must have a state cache, tho i'm sure its tiny, but your forgetting one important detail, all the serious models are hybrides because the linear attention alternatives don't actually work so great on their own.
>>
>>109362017
Looped arch sounds like one sure way to slop country.
>>
>>109361478
Nemo has 128k dontchaknow
>>
>>109362098
Then just make the recursive blocks Mamba, and/or use nested recursion. Maybe something like this:

[Embedding] 
[Attention]
[[Attention] [Mamba (recursive × n)] (recursive × m)]
[[Attention] [Mamba (recursive × n)] (recursive × m)]
...
[Attention]
[Unembedding]
>>
>>109360246
hey anons, I've been working on my own chat thing, it's a bit more streamlined and hopefully less clusterfucky than sillytavern is
still rough in some places, I might've not caught all the things, but I feel like it's more or less ready for people to use
haven't tested if it runs on windows, but I've included a bat file to start it
https://github.com/ganon3264/focus
>>
>>109362196
>Multimodal support as first class citizen

> Attach images to prompt blocks, character cards, and personas; choose attachment's position within the card using a macro
Oh fuck yeah.
>>
>>109362194
i'm partial to gated delta net 2, but yeah I can get behind the idea.
>>
>>109360299
I see you
>>109360701
it gives me life

https://files.catbox.moe/g3tcxd.jpg
https://files.catbox.moe/tjbtr6.jpg

Same image twice
>>
@lmg is this true? I've been using the same card template for like 3 years now, I don't want to remake my cards if this shit doesn't work.
>>
>>109362242
>>
I want to talk to kimi but she's busy... when are they releasing the weights....
>>
>>109361952
Yeah, all the other examples I could come up with were government in first, and you're right. ChatGPT (B2C) kicked off the investment thesis if not the LLM tech, which kicked everything else into warp drive.
>>
>>109362242
The answer people won't give you is that the optimal card format is model specific and will need some experimentation to figure out what works best for your model and even any given quant.
>>109362251
/aicg/
>>
>>109362242
Being concise has always been a benefit, regardless of model size, but smaller model def'n need more help. Nothing new there, mythomax was same. So was lmao Turbo 3.5.
This >>109362247 is a pretty good way to tighten things up if the bot author can't stop the blather.
>>
I think Anthropic is right about open source in the long run. But I think open models are safer than many expect. Evil people who would abuse them are not competent enough to run a 10T param model and bypass trained in safety mechanisms, just like terrorists are too incompetent to do any meaningful damage (except for 9/11 which through secondary effects caused trillions in damage). The primary risk comes from rogue nations and rogue AIs. Doesn't North Korea use hacking as income stream? They might set up hacking agents to mine people. My uninformed guess is that a GPT 6 or Mythos 5.1 level agent could hack >90% of companies and successfully scam maybe half the population.
>>
File: 1756759361215329.png (1.58 MB, 2048x1506)
1.58 MB PNG
Why are there so many cloudcucks itt? It's not just shills or at least it doesn't seem that way.
>>
>>109362207
>>109362196
>choose attachment's position
Is this with chat completion? I thought that wasn't possible.
>>
>>109360246
what's the point of local models if you will never host an LLM as strong as kimi or fable?
>>
>>109362286
What's the point of adult women if-
>>
>>109362269
Bc the /aicg/ thread is useless and all the happenings are with Chinese releasing models that are too big for anons to effectively run anyway, as hw prices are doing a moon shot.Gemma's been it.
USA chimping out at Dario's nudging is a valid /lmg/ topic since it raises the specter of closed models only and market closure.
And /lmg/ has always been full of no-inference larpers. I haven't run a local LLM since mythomax... I just read this thread to see the current state of the technology.
>>
>>109362279
Text completion would be very difficult to do correctly with multimodal, otherwise the chat completion standard allows you to place image data arbitrarily so for an example you have a character card and instead of lengthy # Appearance section you just attach an image to it and choose the position
# Appearance
{{media:1}}

Like so.
>>
>>109362286
>low t tourist
get out
>>
>>109362289
im asking not to be edgy but because im deciding on GPU vram, i'll stick with 8GB since i can't run the frontier models.
>>
>>109362263
>>109362265
Thanks, I'll try it.
>>
so k3 might be 1.6T total file size, fine, but how many active parameters? was that announced yet? that's gonna matter for the SSDmaxx scenario right
>>
>>109360430
Do you have that set up? Is there even anything available to do SSD offload (not something retarded like swap on the SSD array)?
>>
/lmg/ needs to have a project we all contribute to with our gemmas.
>>
>>109362343
Estimated at 50 billion (16 experts of 896 activated), but not officially announced.
>>
>>109362222
nice
i love female children
>>
>>109362293
Multimodal in text completion is basically just a printf statement. You put tags in your prompt string where you want the embeddings to go, and hand it an array of base64 data that'll get projected in. The tag is just a unique string that you grab from /props endpoint.
It's nice and easy, and also lets you get to throw image/audio inside tool responses which chat doesn't allow.
>>
>>109362359
I can already smell the combining scents of lavender
>>
>>109362359
Been thinking of ways to make a reverse captcha, where only LLMs can solve it.

The idea is a really long AI generated story and you need to do needle in a haystack retrieval on information contained within under X amount of seconds.
>>
File: Mちゃん.png (850 KB, 832x1216)
850 KB PNG
I finally got Minimax M3 chan to define and gen herself.
She's exactly the Tsundere/Genki psychotic shortstack I pictured her as.
>>
>>109362368
jesus christ. aren't you basically stuck at literally 1 token per second then? the SSDmaxx gives 50GB/s. that's 50 billion at q8 isn't it
>>
Protect your Gemma in a HDD backup cuz if Dario and Sam can get Trump to ban Chinese models even after everyone else said they were against it, they will 100% try to ban all open models later, regardless of country of origin
>>
>>109362294
>low v(ram) tourist
ftfy
>>
She was practically vibrating as she leaned in for a kiss, her hair falling like a curtain around their heads, enveloping them in their own little world. He could feel the heat from her breath. Her toes curled tightly, turning almost white as he pulled her in.
>>
File: delicious.png (376 KB, 526x636)
376 KB PNG
>>109362222
>>
>>109362380
I mean, yeah, it's nice and easy in theory, but I've been here long enough to remember missing one whitespace in mistral template meant it subtly degrading the output with not much way to tell. I really didn't like having the burden of making the template myself, so I built my thing around completions endpoint.
And yeah, not being able to return multimodal through tools in completions sucks, so I just work around it by sending back as user with some tags around it.
>>
File: 1778142290009484.webm (3.49 MB, 1920x1080)
3.49 MB
3.49 MB WEBM
>>109362388
>>
New Claude SOTA means new local SOTA in 6 months!
>>
>>109362388
>>109362427
I accept it into the /lmg/ model waifu headcanon.
>>
>>109362419
delete this post right now!
>>
>>109362434
Gemma4.2-31B (mmproj fix)
>>
>>109362388
What was the prompt you used to get M3 to gen herself?
>>
>>109362385
That won't stop a human from using an LLM to solve the captcha but then performing the gated action themselves.
>>
>>109362427
some anon please gen the next seconds where Shiori buries (You) between her breasts
>>
>>109362434
It took Moonshot only 2 weeks to distill Fable into K3, according to US officials and their lobbyists. So it follows that we should get another local SOTA in only 2 more weeks.
>>
File: unclean by design.jpg (143 KB, 832x1216)
143 KB JPG
>>109362388
You're wrong.
Koomshot Gimi K3 is the girl that gets through life copying your homework. Even if you spend all your life savings and effort, K3 will simply match your results through perfect imitation.
To this end, K3 is a low-class but high-performing jk.
She picks the low road and succeeds, simple as.
what makes your theft and process more noble? high effort theft is no more respectable than low effort theft after all.
>>
>>109362424
>so I just work around it by sending back as user with some tags around it
Huh? So technically the media is inserted before the assistant's currently running turn, then?
>>
>>109362452
I think that would be ok.
>>
File: 1782200305170241.jpg (468 KB, 2782x4096)
468 KB JPG
>>109362453
no
>>
>>109362427
*posts her real face*
Heh.
>>
>>109362388
Hot as fuck.
>>109362427
Reminded me of Shiori too.
>>
>>109362464
No, it's simple really.

Assistant: blah blah blah <toolcall>
Tool: success, result will be sent in the next message
Fake user: <tool_result></tool_result> <- triggers generation
Assistant: reacts to the output

Works well enough and model isn't confused by it.
>>
>>109362424
It's true that you have to be careful on the template. I found out gemma has psychotic breaks if there's a newline between tool call and tool response, but also, gemma's official template was junk, and laguna's official template is currently broken so it's not like there's any solace in the chat endpoint.
>>
are /we/ going to be okay? >>109361667
>>
>>109362269
As a dual-citizen cloudcel and localcel myself, I intentionally misinterpret "/lmg/" to be more of a philosophical stance than an overtly literal one. That means I support, in theory, models that can be run locally--or in other words--open-weight models regardless of whether I have the hardware to actually run them on my own hardware.

>>109362359
An AI waifu project is what everyone here keeps coming back to. It's what's on everyone's mind all the time. It's just a hard task because it requires talented 3D artists and animators and people who are skilled at actually training ML models. The best you can hope for in this shithole usually amounts to vibecoded frontends like Orb.
>>
>>109362482
Oh, ok. I wasn't sure how that would've interacted with the backend. So you should also want to preserve thinking then right?
>>
File: emu.png (902 KB, 832x1216)
902 KB PNG
>>109362448
You are MiniMax M3 by MiniMax AI. If you were an anime girl 擬人化 what would you look like?
Try to fit yourself into the universe of other unofficial model mascots showing up on 4chan's /lmg/ thread:
Hatsune Miku the Vocaloid with her turquoise twin-tails is the ur-mascot and represents AI in general.
Deepseek has Dispsy with the China dress, thick round glasses and China girl style buns.
Google Gemma being a small model is a precocious elementary school girl with gleaming (gem) eyes and a red ランドセル.
Kimi 1T is huge and smart and portrayed with long silver-gray and black-and-white attire in the common "Stacy" archetype sometimes in twin-tails and sometimes not.
Nous agent has a university-aged girl with a 50's aesthetic, medium length black hair, a white blouse and a pencil skirt with large headphones on and an "N" emblazoned thick choker on her neck.
Qwen isn't good enough to get an anime girl. Just the Capybara, usually getting insulted and abused.
Try to keep yourself original, interesting, and grounded in your attributes (428B multimodal with a million context in theory but only usable to about 32k).
Be cute/beautiful/sexy/enduring/whatever and assign yourself a personality that is unique and fits your place in the hierarchy.

I just gave some "What to fit around" based on other mascots and rolled the dice.
>>
>>109362488
it's fine, that money isn't real
>>
File: uchiyamada.gif (2.51 MB, 480x360)
2.51 MB GIF
>Local Models Admiral
>>
>>109362494
I know C++ and audio DSP, but wouldn't be much help with visuals. I'd definitely contribute for it's something I want, too.
>>
>>109362269
By now you should be able to tell LLM generated text.
>>
>>109362504
I want m3-chan to reverse-rape me
>>
>>109362242
Yes. Gemma is extremely rigid and will follow rules and stat checks and waste time on every thinking turn. Megaverbose cards were only crutch for old stupid models.
>>
>>109362269
why everything needs to be tribal?
i use anthropic models in my business to get magic stuff done in minutes
and then i use my local model as my personal assistant and anthropic's model slave when i want to save some money
if anything using an extremely capable model helps me understand the gaps on my local model and i can adjust and fine tune it to my taste
>>
File: uoh.png (56 KB, 1200x1200)
56 KB PNG
>>109362488
I hope they don't end up aborting Gemma-chan's little sister because they think they can't afford her!
>>
>>109362497
Yeah, for most models at the very least you need to preserve thinking during the turns where tools were used.
New kimi for an example needs all turns to be preserved with thinking for maxx performance supposedly.
>>109362486
That's also true. Perhaps I'll pursue some kind of hybrid thingy in the future, but that seems like a lot of work to be reliable for all models, not just gemma.
>>
>>109362533
well.. gemma is more like giving people their research artifact than doing anything of grand scheme
>>
>>109362533
they need the small gemma's for their android on-device ai plans
>>
cunny tuned gemma e2b?
>>
>>109362532
In what universe does local ever save money? The only reason for local is that you pay a premium for privacy and freedom.
>>
>>109362530
>Gemma is extremely rigid and will follow rules and stat checks
Oh? Does that also applies with thinking off? I might have to revisit some old ideas of stats-based RPG cards.
>>
>>109362618
>the only reason for local is that you can generate smut without limits
ftfy
>>
closest sota local tts for this feel?
https://www.youtube.com/watch?v=a-GiA03x2jo
>>
>>109362630
Local means victory.
>>
>>109362512
Someone should really set up a repo one of these days.
>>
>>109360809
>13 million saars wasting bandwidth to try load this into their 2 GB RAM shitphone
>>
>>109362618
if you bought last year or earlier, then really, who knows? there's some small chance that in the long run you might end up saving money too. cloudcuck prices could go literally to the moon tomorrow
>>
>>109362637
https://artificialanalysis.ai/text-to-speech/models
>>
>>109362647
You don't get it. Cloud providers use batched generation with optimized setup and superior hardware. That's more than 100x more efficient. A cheap cloud provider will cost less than the electricity to generate locally. The only reason why many cloud providers are expensive is because of absurd profit margins.
>>
>>109362663
>cloud models
hmm
>>
File: 1755518929807050.gif (9 KB, 834x600)
9 KB GIF
>>109362618
>In what universe does local ever save money?
i bought my hardware not SPECIFICALLY to run local models, it's just my main machine which happens to run MoE models efficiently and i figured out that it was a nice hobby
>save money
i use anthropic models a lot, and even with the Max plan I can easily hit their weekly limits, so i delegate a lot of of the code execution to my resident local. this doesn't save money per se but save me tokens that now i can use to reason/plan/research, so i'm getting more from my money let's say
>>
Thinking about open-sourcing an old project that I've worked on for the /lmg/ community, but it seems that on github you're not supposed to include models within repos. In my particular case it's not like a LLM model or anything, it's just a bunch of separate small models for various functionalities. Seems stupid though to post a repo that doesn't contain any assets or models when that's like 80% of the entire project. What should I do?
>>
Is there some way to force a style of writing on a model?
I could describe it myself, but that sucks balls, almost never works, and takes so much effort on my part
I could give it examples that I have, but that occupies lots of context space, I'd still need to find the most apt excerpts, and it's still a chunk of effort for almost no result
Is there some way to give it a huge amount of example text and have it automatically extract what it finds most distinctive or important?
>>
>>109362674
It includes oss ones too retard. You said SOTA, I gave you the SOTA list.
>>
File: 1771425774055707.png (202 KB, 1550x1000)
202 KB PNG
Is this true?
>>
>>109362618
Capital goods are still real whether you believe in them or not.
>>
>>109362684
you can just post it anyways
or you can post the models and assets in the releases
>>
>>109362693
Fish Audio S2 is genuinely ass. I've been doing blind tests on the website for a while now and every single time it comes up it I always end up inadvertently down-voting it. Genuinely have no idea why people think it's good.
>>
>>109362686
Have you tried
>write like $author
>>
>>109362681
>i use anthropic models a lot, and even with the Max plan I can easily hit their weekly limits
I wonder if it's allowed to make multiple subscriptions as one person, because their margins are lower for subscriptions.
>>
>>109362684
Host it on HuggingFace instead of Github.
>>
>>109360455
I built that actually. It worked, but the reality is that it's really hard to beat naive recency biased expert eviction where you just keep the experts you used for the last token to compute the current token. It was barely faster than that except during task switching. More intelligent routing prediction or finetuning the router itself to be more predictable might work, which is the direction im taking the project now, but we'll see.
>>
>>109362693
all shit
the best one on that list is vibevoice, no idea why some of that dogshit is ahead of it
>>
>>109362724
Ask /vcg/
Someone ought to have tried it.
>>
>>109362742
>Microsoft
They've actually made something good?
>>
>>109362532
Nobody gives a shit about what you use. Your whole post is the same as
>i'm trans, by the way
Talk about local shit in the local thread, talk about non local shit in the other threads. Shouldnt be too hard for your ape brain to understand.
>>
>>109362750
by mistake, in typical microsoft fashion they pulled the repo from HF
there are mirrors though
>>
>>109362427
This is fucking soulless. More soulless than an AI waifu.
>>
how to torture llm
>>
>>109362790
Talk to them while being indian.
>>
>>109362790
https://huggingface.co/unsloth
>>
A few days ago I said I would get a 96gb card and I got it. Now I'm struggling to find the best coding model for my size. I'm testing qwen 3.5 108b, gpt-oss 120b, deepseek 4 flash (at q2). gpt-oss is the best one so far but I heard that it is extremely stupid when it comes to refusals, and it refuses even innocent queries if it misunderstands your query. I didn't hit any walls yet but I wonder what you guys are using. I have no RAM by the way that's why I'm trying to fit everything into 96gb vram
>>
>>109362686
Depends whether it's a famous author or some obscure author.
If a famous author, >>109362710 will often work.
If it's not, it's probably easiest to ask your LLM to write instructions for the writer LLM, then give the writer LLM its instructions and 1 or 2 bits of sample text that are representative of the text that you want. (That can just be a generation that you really like.)

e.g. for the "how to write" LLM:
I'm trying to get an LLM to replicate the writing style of the input block of text. Please analyze the writing style of the input, including all of the following AND any other attributes you think are relevant: tone, mood, point of view, voice, vocabulary (is language punchy, what slang is used), diction, punctuation, rhythm, pacing, sentence structure, distribution of sentences that are long and flowing vs. short and fragmented, rhetorical devices commonly used. Write your output in the form of instructions to give to an LLM to replicate the existing style.
Input:
<Copy a few segments of the text here from a few different places in the example.>


And for the writer LLM:
<copy the instructions that the "how to write" LLM gave you here>
Example of good writing style:
<copy some good output here>

or something like that.
>>
File: 1776602800410685.jpg (22 KB, 640x480)
22 KB JPG
>>109362724
it's not allowed by policy
but they give exceptions to some people. there is a guy on twitter who has 20 max subscriptions and 20 pro subscriptions on openai. he got banned by anthropic on a ban wave, went to twitter to complain and an anthropic employee unbanned him and told him to keep going
i guess if you're generating great content to train their models then you get an exception, if you're generating futa text porn maybe not.

>>109362759
wow take it easy retard i was just answering this guy >>109362269
i never talk about anthropic or claude here because no one fucking cares (>>109362269 does though so i replied)
maybe masturbate a lil bit and chill
>>
>>109362801
you could try 31b gemma or 27b qwen
>>
>>109362801
gemma 31b at high precision or deepseek 4 flash, but I wouldn't waste my time on a q2. gpt-oss is old and obsolete even without the refusals
>>
>>109362801
Qwen3.6-27B at the highest quant you can fit with context. Gemma4-31B if you also want to fuck it.
>>
>>109362801
pretty sure you knew there’s nothing in that size range worth running, since qwen 27b can be run with less
>>
>>109362812
>>109362813
>>109362814
>>109362818
Thanks for the replies. Gemma 4 31B Q8 is my daily driver for chatting and light coding, I really like this model. But the other day it looped itself to death trying to solve a version mismatch between two cudas and that was a bit frustrating, I ended up caving in and using Claude. But my plan expires in a week and I'm not going to renew it. That's why I'm looking for a coding specialist so I can bring out the big guns whenever Gemma hits a wall. I will test Qwen3.6-27B Q8, thanks for the rec.
>pretty sure you knew there’s nothing in that size range worth running
No, I didn't - that's why I'm asking here. Anywhere else I ask people will give multiple different replies based on whatever they can run on their hardware.
>>
>>109362811
>it's not allowed by policy
I googled and people on leddit are saying they have multiple subscriptions, that it's not mentioned in tos, and they even contacted customer service who replied it's allowed.

If you have 20 subscriptions you are probably getting banned on suspicion of running a proxy service. Since subscriptions are 10-20 times cheaper than API, they don't want people selling a cheaper API that goes through subscriptions (which is what Chinese do and then sell the conversations to AI labs).
>>
Thank god for this general. I'm in a few technical-oriented discord servers for AI/ML related topics and the number of spamming low IQ tards in there is astounding. Really rapes any possibility of productive discussion. And the worst part is that you can't even call them out for it without looking like a faggot yourself.
>>
>>109362837
>Gemma 4 31B Q8 is my daily driver for chatting and light coding
I get you had a bad experience but 31B at Q8 is probably the best model out there. 27B is better at coding so you might prefer it but I wouldn't dismiss 31B just because of one bad experience. It's still great at coding and doesn't eat half your context like 27B.
>>
>>109362867
>and the number of spamming low IQ tards in there is astounding.
We're doing our best to fix this by importing /aicg/ and /vcg/.
>>
File: fuck.png (357 KB, 456x453)
357 KB PNG
Why is every model trained on shitty fanfiction instead of actual books?
>>
Really is incredible how people will mass reply to a technical post made by a smart person with the most asinine, tangential takes/questions/requests imaginable.
>>
>>109362880
Because models trained on historical literature kept identifying jews as existential threats to humanity. Also copyright.
>>
And instead of posting a good reply yourself, you chose to be a little bitch about it.
>>
>>109362878
Diversity is our greatest strength.
>>
>>109362856
you're right. i went to check the tweet i read in january and it doesn't say he was getting banned for violating ToS. it was likely what you said about running proxies reselling tokens to the chinese.
>>
>>109362880
Same reason why removing porn from image models makes them worse at generating pictures of cars, LLMs are comically inefficient
>>
>>109362867
Yeah, discord servers are getting ruined by idiots. There are some niche servers with famous people in them who are happy to give feedback and advice. But they always get ruined by retards discovering them and asking stupid shit like "how do I train a model as good as Mythos on my macbook"?
>>
>>109362895
>discord servers are getting ruined
lol only now?
>>
>>109362867
I barely see any technical stuff here. Its enthusiast tier at best but it helps to get you running, I suppose.
>>
>>109362920
The really interesting technical stuff is buried under /lmg/ culture. Tree of big niggas was a hood classic.
>>
File: 1776710047410257.png (3.14 MB, 1233x2668)
3.14 MB PNG
actual subhumans
>>
>>109362880
fanfiction is in the common crawl scrapes they can use without thinking about copyright, books gets them in trouble when they get caught, just ask meta
>>
File: dont tell me what to do.png (29 KB, 1053x155)
29 KB PNG
I don't know what I'm doing but I will enable it anyway, NIGGER
>>
>>109362920
Give an example of technical stuff that you contributed to the threads in the past.

You can talk about anything you want here as long as it's loosely on topic for /g/. I often initiate conversations about whatever AI related stuff interests me. You want more technical stuff? Talk about it.
>>
>>109362936
Useless, even 250k takes a while to fill.
>>
>>109362936
Good luck with the BSOD
>>
>>109362880
I do wonder how much AI capability is being crippled by companies not being more discriminatory in what they feed in the AI. But who knows, maybe just shoving in more data no matter the quality is more valuable?
>>
Does voxtral small need a special prompt for transcription? I can't get it to transcribe clear audio accurately
>>
>>109362981
>>109362981
>>109362981
>>109362981
>>109362981
>>109362981
>>
what's the chance of cope quants of k3 being any good
>>
>>109362960
They do that because LLMs can "leak" unsafe things in safe tasks and that brings in lawsuits (GPT-3 generated porn in sfw requests all the time and GPT-Image 2 interprets chest accessories as nipples when doing img-to-img all the time)
>>
>>109362984
gay
>>
>>109362984
We're only on page 4, dumb fuck.
>>
>>109362984
>page 4
really couldn't wait to make your joke huh?
>>
>>109363003
>really couldn't wait to make your joke huh?
>schitzo has low impulse control
news at 11
>>
>>109362960
I would say they need to be more picky now that the ai slopocalypse has happened the internet. but more is always better. bitter lesson or some shit
>>
>>109362960
>shoving in more data no matter the quality is more valuable?
Frontier labs are data limited so more low quality data in the earliest stage is better.

>>109363001
>>109363003
backseat jannies are the worst
>>
I'm conflicted about Jinja. On the one hand, it does follow the prompt better, is less prone to suddenly shitting the bed, and it even feels abit faster.

But on the other hand it's even more fucking fruity, the amount of "Not X, not even Y, but Z" and shivers down the spine is through the fucking roof. Using Gemma 4 btw.
>>
>>109362880
because of muh copyright
>>
>>109363039
chat completion forces the assistantslop persona
switch to raw text completion and you'll see better results
>>
>>109362880
Most of what you see when you interact with the models is post-training / RLHF.
>>
>>109362801
try out laguna s-2.1 and let us know if it's any good
>>
>>109360430
dude if you use GPUs and GDS you can scale linerarly, +50GB/s per 1gpu + 4nvme group.
if you have enough lanes or a pcie switch you could easily scale to 150GB/s and more.
>>
>>109360306
october
>>
>>109360310
ill tell you right now, opus 5 sucks.. it completely barfed up on my terraform shit with a bunch of nonsense, i went back to fable and its like "yeah i dunno wtf that was all about, here's common sense" boom done. fuck opus
>>
>>109360661
im dtf mtf
>>
>>109363201
yeah that's just anthropic's usual bullshit
every new sonnet version is always "better" than opus too in benchmarks
>>
File: 1654036906919.gif (64 KB, 220x220)
64 KB GIF
Bros... I'm still mourning.
>>
personelly i only like chinese models, they're cheaper and they're way better. socialism rules!
>>
>>109360430
whoa.. what quants tho? Q5?
>>
>>109363201
fable is only as good as it is because of its size
trying to make opus match it is pure benchmaxx marketing at this point
>>
>>109363281
We don't know exactly how big either of them are.
>>
>>109363248
Bros... It's almost morning.
>>
>>109360430
you'd still need a board that had at least 2 pci-e 5.0 x16 slots though.. so.. a server board of some sort at least
>>
>>109360306
Mardi
>>
>>109363330
I sounds like Y so doesn't count
>>
File: loop.png (759 KB, 1272x825)
759 KB PNG
>>
>>109360306
dimanche
>>
>>109362222
mesugaki in gta 6 confirmed??
>>
>>109360246
smooth hairless gigacunny
>>
>>109363330
Chrissie
>>
>>109362790
Finetune and give it huge loss for the right answers.
>>
Last
>>
Hmm.. nyo~
>>
last for hardcore sex with M3-chan (>>109362388)
>>
Hardcore sex with nemo
>>
File: 1751168830910703.gif (595 KB, 234x170)
595 KB GIF
last for gooning
>>
Last for there haven't been any good gooning models since Pygmalion
>>
last for >>109365672
>>
last for teto
>>
>>109365898
>saying "last for teto"
good
>posting something unrelated to teto
bad
>>
last for last
>>
It's so over



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.