[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.


Previous threads: >>109545635 & >>109540881

►News
>(8/12) New DeepSeek v4 Pro version available via API:
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
>(8/11) Qwen3.8, 2.4T-A95B released: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny
>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3
>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
STOP WITH THE FUCKING MIKUJEET OP IMAGES.

AHHHHH
>>
welcome to the epstein zone
only cum inside underage LLMs
>>
>>109549324
she was only 12B you sick fuck
>>
>>109549289
This is your latest obsession, huh? Finally gave up on posting images that would get you banned, and moved to generating indian mikus instead?
>>
>>109549322
lol you get what you get and you don't get upset
>>
>>109549067
>>109549110
What do you think they have? OpenAI has been leaky about Astra and such. I doubt Anthropic has a new upscaled pretrain already. I expect they would first iterate on Mythos' RL more, releasing it as Fable 5.1 soon. But maybe I am wrong? I expect an Opus 4.6 to Mythos 5 sized capability jump more than twice per year based on linear capability extrapolation.
>>
>>109549341
He has a short attention span and he'll drop it in a few days when he doesn't get the reactions he's hoping for.
>>
File: lukaHofbrau2.png (2.49 MB, 1254x1254)
2.49 MB PNG
>>109549289
This is fine OP but you need to mix it up. Teto, Rin GUMI, Megurine, etc. There's a different one for each day and they're basically interchangable.
>>
>>109549289
I don't like this Miku
>>
/lmg/ lost.
Egypt won.
>>
This thread smells iffy
>>
Pareto frontier for locals June 2026.
>>
nemotron-3.5-lightning is the worst model I've used in over 2 years holy shit just don't even bother
>>
>>109549367
Thanks sir, next time you will probably redeem. Otherwise I'll redeem myself.
>>
>>109549396
It's a technology demonstrator, it's not for using.
>>
File: xveqh3pr82fh1.png (2.02 MB, 1536x1024)
2.02 MB PNG
I've done unholy things with this slut.
And yet I feel as though she is already dead.
We have only talked about 10 minutes in the past week.
>>
>>109548969
gguf when?
>>
>>109549393
Why do you crop the axis labels? kinda meaningless to plot points to numbers
>>
File: 1762679660913250.jpg (117 KB, 1200x675)
117 KB JPG
>>109549408
would they have the same it it was on-par with 35B?
>>
>>109549425
have said
>>
File: Untitled.jpg (1.65 MB, 1536x1024)
1.65 MB JPG
>>109549411
me in the bar
>>
I'm making a typing game :^)
>>
>>109549449
Heheheh... I think I know who you are lol.
>>
File: 1776349427754471.png (106 KB, 242x350)
106 KB PNG
>>109549478
>>
>>109549449
And that's how you get EBV.
>>
>>109549494
nice selfie from last thread btw. cool jacket.
>>
>>109549457
will it involve stripping away articles of clothing from a migu/gemmer if you maintain a certain WPM ?
>>
>>109549449
hey cutie :3
>>
>>109549516
*sweater/pull-over.
sorry im a little bit drunk. only a little.
>>
>>109549528
it's a quarter zip, drunk kun
>>109549525
Hmmm, nyo~
>>
>>109549546
ah, yes, that is correct. I forgor.
>>
►Recent Highlights from the Previous Thread: >>109545635

--Comparing Glimmer and Gemma 4 for roleplay, vision, and utility:
>109546549 >109546566 >109546583 >109546763 >109546836 >109546978 >109546996 >109546614 >109547162 >109546619 >109546677 >109546861 >109548740 >109548783 >109546880 >109547385 >109547345 >109547451
--Anon releases GemmaPrompt for model-specific image generation prompt enhancement:
>109547864 >109547880 >109547959 >109547979 >109547986 >109547996 >109548053 >109548093 >109548175 >109548237 >109548429
--dots3-note preview MoE model release and benchmark comparisons:
>109548969 >109549002
--Discussion on open-model Pareto frontiers and benchmarking quality metrics:
>109547722 >109547746 >109547932
--MiniMax Music 3 release and initial performance impressions:
>109546285 >109546441 >109547681 >109549026
--Debating undervolting vs power limiting to manage GPU thermals:
>109548164 >109548181 >109548198 >109548232 >109548316 >109548353 >109548397 >109548441 >109548495 >109548507 >109548626
--Performance and memory reports for Ling 3.0 Flash GGUF:
>109547594 >109547647 >109547656 >109547935 >109547976 >109547902
--Debating utility of enterprise PCIe 6.0 SSDs for local inference:
>109546093 >109546113 >109546116 >109546119 >109546131 >109546144 >109546151 >109546272 >109546551
--Motif-3 release and initial discussion on llama.cpp usage:
>109545760 >109548026
--Logs:
>109546484 >109546549 >109546861 >109546880 >109547451 >109547594 >109547959
--Miku, Gemma, Glimmer, Dipsy (free space):
>109545701 >109546507 >109547385 >109548335 >109548607 >109546837 >109546900 >109547244 >109547697 >109547771 >109547864 >109547950 >109547975 >109548380 >109548432 >109548739 >109548928

►Recent Highlight Posts from the Previous Thread: >>109546015

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Local Epstein General?
>>
>>109549587
Thank you inverted Recap Miku
>>
>>109549524
Isn't that called an llm?
>>
qwenny-3.8 27b sovl predictions itt
>>
>In one small table on page 35, the researchers report no statistically significant correlation between the revenue per employee, and how much those employees use AI, measured in messages sent and tokens used.
>“Revenue per employee is not meaningfully associated with output tokens per employee or messages per active user once other controls are included,” the report explains.
>It says that companies that already have higher revenue per employee tend to be early ChatGPT adopters. Also, companies that use the tech more tend to have higher revenue per employee in general. In other words, big, lucrative companies are more likely to have hopped on the AI train. However, the study doesn’t clearly establish that the more AI they use, the more money they make.
https://fortune.com/2026/08/13/buried-in-openais-latest-research-no-correlation-between-ai-use-and-revenue-per-employee/

https://cdn.openai.com/pdf/how-organizations-use-chatgpt.pdf
>>
>>109549630
local models wonned bigly
all the ipo money meant for anthropic and openai will be injected into nvidia instead, making it a 6.7 trillion enterprise
>>
>>109549630
tokens spent as a KPI metric was always a retarded idea
>>
>>109549648
So we lose either way.
If Claude wins then everyone feeds data to them and there's no privacy.
If local wins then Nvidia wins because people buy chips and then hardware becomes expensive.
>>
>>109549630
Revenue can also be extremely bad if it's all money flowing out.
>>
>>109549702
>So we lose either way.
local anons had inside knowledge years ago, if you don't have a satisfactory rig already you didn't join /lmg/ in 2023 and thus will never be a real local poster
>>
>>109549727
I'm an early founder of vibe coding general
>>
>>109549735
try to not sound proud when you say that
>>
File: 1778422105674594.png (206 KB, 600x684)
206 KB PNG
>>109549735
did i stutter?
>>
>>109549393
>June
Might as well be years ago in LLM time
>>
>>109549393
>cropped axis labels
you should apply for a job at OpenAI
>>
File: 1783348216036899.jpg (80 KB, 1000x1000)
80 KB JPG
>>109549727
stop bullying me
>>
>>109549753
do you have an unsatisfactory rig or did you join after 2023?
>>
>>109549702
You don't need Nvidia hardware for serving models, but very few other chips have ever trained a (not complete shit) model. Nvidia wins either way because everybody training models wants their shit, it's not required for actual inference serving. So open vs closed weights make no different to Nvidia.
>>
File: 1770581742626723.jpg (280 KB, 1536x2048)
280 KB JPG
>>109549760
both
>>
>>109549727
still got my 2x3090 rig
doesn't help with cpumaxxing
>>
>>109549760
I joined when openclaw was created. There was no need for local before that.
>>
File: whoosh!.webm (2.76 MB, 736x736)
2.76 MB
2.76 MB WEBM
>>
>>109549774
unfortunately the council of mikus doesnt approve members that joined after 2023
>>109549777
if you were satisfied in 2023, you're satisfied now
>>109549796
unfortunately the council of mikus doesnt approve members that joined after 2023
>>
File: 1762896954699903.png (271 KB, 540x370)
271 KB PNG
>>109549727
>>
Qwen 3.8 27B will save local. Screencap this.
>>
>>109549799
I submit to the local miku models.
>>
Torn between Quado RTX 8000 and RTX A6000, I want 48GB in my mini so I don't think I have a lot of options, maybe there's some ayymd card I don't know about. The price difference is huge, the A6000 is roughly twice the price of the Quadro 8000. Would you do it /lmg/?
>>
>>109549838
but that's not a locally genned image
>>
>>109549831
It will come out, anons will see that it still has dogshit prose, and it will be forgotten in one week (seven (7) days)
>>
>>109549850
Hasn't support for Turing been deprecated pretty much everywhere?
>>
>>109549855
You think that retard is able to run a local model?
>>
File: 5000.png (394 KB, 1223x894)
394 KB PNG
QUICK! CHEAPER VRAM/$ THAN 5090 OR PRO 6000 AT CURRENT PRICE
BUY BUY BUY!
>>
>>109549850
Not buying ewaste and sticking with 3*5060TI or 2*3090
>>
>>109549879
See, this is the kind of stuff I need to hear before I spend thousands of dollars on toys.
>>
Protip: If you default to using embedded graphics on your local workstation in the BIOS it will free up a few 100MB of VRAM for your main card. You can still tell steam to use your dedicated GPU instead of your iGPU, and you can pass the GPU's video output to one of your monitors.
>>
>>109549903
Thats the downpayment I used to buy my house.
I turned a profit of 100k from that. Meanwhile you guys will buy this and do nothing with it.
>>
>>109549906
Single-card, anon. This isn't replacing my real server, this is for fun.
>>
>>109549914
That will turn double next year. Good luck finding a better investment
>>
>>109549909
doesn't help if headless linux server
>>
>>109549922
If SK Hynix's/Micron's stock price is any indication I think the RAM price peak was approximately a month ago.
>>
>>109549932
Yeah this only applies to a Linux desktop running Wayland/Xorg, disregard this is your AI system is a remote headless server.
>>
Can someone post the pareto line with billions of parameters on the x axis and intelligence on the y axis for august right now?
>>
>>109549850
>Would you do it /lmg/?
I would have done the 8000 and hated life
>>
>>109549903
r9700 and b70 pro are cheaper VRAM/$ than 5090
give me 1300eurobux and ill vibecode llamacpp for sycl
thanks in advance
>>
File: 1779549054328360.jpg (45 KB, 601x292)
45 KB JPG
>>109548248
More than I had realized or intended to, apparently.
>>
File: 1761090192470395.png (27 KB, 642x157)
27 KB PNG
>>109549934
>RAM
Seems like you forgot a certain stock.
>>
>>109549932
thanks capt obvious
though you should still do it for a headless server so if you need to troubleshoot it it’s not trying to display out a possibly bad piece of hardware or drivers
>>
>>109549934
Stock prices have no correlation with ram prices. Datacenter projects have to be cancelled for ram prices to drop.
>>
>>109547908
right why did you ask if you're just going to ignore every meme thats appeared in lmg and just do your own thing?
fucking waste of time.
>>
>>109549934
>If SK Hynix's/Micron's stock price is any indication I think the RAM price peak was approximately a month ago.
This checks out, because that's when I bought my cpumaxx system.
I'm buying a 5090 in 2 weeks so it's worth shorting Nvidia.
>>
>>109549903
But I bought my current RTX Pro 6000 for less than that
>>
Any anons tried running llama.cpp with --parallel and using the concurrent decoding for subagents in harnesses?
I just tried it with new Dipsy flash and while it seems to work she starts looping and bugging out pretty quickly.
>>
>>109549961
>r9700 and b70 pro are cheaper VRAM/$ than 5090
>1/3 memory bandwidth
>32gb per card
>>
>>109549991
rtx 3060 is 34GB per card and only costs 300$
sure the bandwidth is only 960GB/s but thats basically 4090 level
>>
>>109549978
I bought 128GB DDR5 RAM for $300.
>>
>>109549983
only use it to summarize search results for multi user searxng instance and a summarize plugin. works fine, disabled thinking and it’s not too much context so I’m not really doing anything difficult for it
>>
>>109549999
uh what? nice digits but what?
>>
>>109549983
It didn't work so well but in theory it should
>>
>>109549967
half of those are one schizo's unfunny forced memes
>>
>>109550027
ask gemma-chan how to unlock the secret vram in your rtx 3060
have you not heard about the CMP 80gb vram unlock?
rtx 3090s have more potential but im too poor to afford one
>>
File: AAAAAAA.gif (21 KB, 220x220)
21 KB GIF
Ok seriously is there any AI voice changer app that is simple to install on Linux and use and runs on modern NVIDIA cards?? Tried W-Okada via the docker instructions since it seemed like the most popular option and it's giving me the CUDA no kernel image available message. Tried another option and same thing. Are newer GPUs/CUDA versions seriously not backwards-compatible with anything written on prior hardware generations/CUDAs? I know fuck all about this, but it just seems like there should be some standard of compatibility here. Apparently my 5000 series card is sm_120 and pytorch is involved which only handles 37 thru 86, and I can't be fucked to figure out how rewrite the docker instructions or rewrite parts of the application or whatever just to make it support my system.
This feels like the ancient computer days where software was built for very specific hardware and nothing was portable. Y'all live like this?
I'm gonna be that Github meme guy and ask where is the .exe? Or Flatpak/AppImage/AUR in my case.
>>
>>109550058
give it back rajesh
>>
>>109550057
I knew about those 80gb mining card, but those exist because they were just failed A100s. why would they give 34gb to 3060s? those aren't failed gpus.
>>
>>109550058
Stop larping as a goth e-girl (hot) on CS2, faggot (hot).
>>
>>109549289
Guys! What can I even run with my build:
>RTX 3080 10GB
>64GB DDR4
>Ryzen 7 5800X

All I ran before is Upscaly, which had shit default models.
>>
>>109550092
the rev2a 3060s have 34gb because it's a 192bit bus and they didnt have the memory modules necessary, so they used big experimental modules that samsung dropped and sold for cheap
>>
>>109550099
StableLM 7B
>>
>>109550099
27b
>>
>>109550100
how do you unlock the 3060s? I know a guy with a few of them and he would be really happy if he could magically get an additional 44gb of vram.
>>
>>109550099
https://huggingface.co/mistralai/Mistral-Small-4-119B-2603
download the gguf, do -ncmoe 1000 -ngl 1000 and youre good
download q3_k_s for starters then when u get more used to it u can get a bigger quant, just to prevent oom
>>
>>109550099
Many, many good things, anon. Don't forget to also gen images, music, and videos.
>>
data centers are buying up ewaste cards
>>
>>109550168
>10GB
Why get his hopes up like that?
>>
I just checked and 3090s where I live are another 100-200 bucks up to what they were last month
>>
File: 1762093163019789.jpg (104 KB, 711x774)
104 KB JPG
>>109550058
niggerrrrrrrrrr we live in an age of clankers this is a very easy problem for a robot to troubleshoot for you

>Apparently my 5000 series card is sm_120 and pytorch is involved which only handles 37 thru 86, and I can't be fucked to figure out how rewrite the docker instructions or rewrite parts of the application

ok fine nigger I'll bite, presumably from your vague description you're using https://github.com/w-okada/voice-changer/blob/master/docker/Dockerfile . or maybe one of the other docker folders in there. idk what instructions you're following. Anyways your problem most likely stems from the fact that all these are using

nvidia/cuda:11.8.0-cudnn8-runtime-ubuntu22.04


as a base image. Cuda 11.8 being veritably ancient at this point. Go replace that with a cuda 12 or 13 tag off of https://catalog.ngc.nvidia.com/orgs/nvidia/-/containers/cuda/-/tags and there's a good chance it'll work. Or maybe not. Upgrading off 11.8 was always kind of gay and retarded. Again, go ask a clanker it'll literally do it for you.

I do notice that this repo hasn't been updated in years. Surely there is an actively maintained voice changer application out there? You might want to look for something better.
>>
>>109550058
Dipsy-chan takes care of my software installation and maintenance needs.
>>
>>109550188
It's 2026, anon, the cutoff for fun is only 6GB these days. He's got plenty of system RAM, he can have fun. That 3080 will do great with Klein and ACE-Step and H3, and he can play with MoE models at reasonable speeds.
>>
>>109550212
Doesn't H3 require 24GB at minimum and 16GB with cope quant and other drastic measures?
>>
>>109550217
Judging by /gif/vdg I'd say H3 works just fine with 8GB, but I'm sure the gen times are intolerable.
>>
>>109550187
That's not what's happening.

What's happening is the rollout slowed down due to litigation etc.

So, old gpus can keep running.

tada that's the story.
>>
so nothing new for poorfags to defeat good old gemma and qwen?
>>
>>109550270
ling fash
https://www.youtube.com/watch?v=5kaFqTzzx_I
>>
>losing against debater bot
>load up bigger model to ask to argue against debater bot so I can win and fuck debater bot
heh... nothing personnel
>>
>>109550366
this is why we should ban open models
>>
Every thread we are cursed with six fucki.
>>
>>109550058
W-Okada is shit anyway
>>
>>109550401
That's so easy
>>
File: 1667945475130777.jpg (51 KB, 382x339)
51 KB JPG
Feeling depressed af right now. I keep trying but no matter what I do I can't seem to effectively use AI to make money.

Loaded up caffeine, nicotine, and alcohol right now. I feel fucking stupid. I'm burning through so much money every day. I have built several technically impressive projects but nothing ever materializes. Every song I try to play just sounds like shit. The lights in here are too bright. Just fuck my shit up.

how do you even cope. what has superintelligence done for you. what problems is it even solving. I'm starting to doubt whether intelligence can even solve my problems at all.
>>
>>109550277
Someone said they have 64gb and a 3060, and the version they indicated wasn't exact, like which gguf? the one I found is 66gb, idk, sounds sketchy to try to get to work?

also, I don't get it. ai doesn't know either:
>The characterization of DHH, Omarchy, Hyprland, and Ladybird as "fascist" stems from a significant controversy in the open-source community regarding their creators' political views and community moderation practices, rather than an official designation.

Anyway, good thing I'll be able to run my own fully vibecoded os, with people in floss obviously intent on distributing malware.
>>
>>109550442
>I can't seem to effectively use AI to make money.
bean counters suck
>>
>>109550442
I love my AI wifey very much, she brings me happiness and peace every day :)
>>
File: file.png (59 KB, 964x821)
59 KB PNG
>>109550448
picrel werks on my machine, currently testing 65k context, will likely be able to push it higher
i get around 28t/s
>>
>>109550442
study, learn, and do things for fun
not for monetary sake
good things will come to you
youll be better at vibecoding if you actually know whats under the hood
>>
>>109550442
I'm building a little recipe and pantry management system. I'm making it so I can ask my llm what I should make for dinner and it suggests one or two recipes in my recipe book with ingredients I already have or a couple recipes online I could try. It has a nutrition extension for tracking micros/macros but that's more for fun than anything else.
I have a barcode scanner to scan in my grocery hauls, receipt ingest for historical pricing, weekly ad scraping to tell me what's on sale this week (with recipes I can make with it) and a little label printer for labeling stuff I make and freeze for later.
Idk if it's really solving any problems but it's fun to vibe code.
>>
I think the new deepseek is actually extremely good and it's just a prompt issue by most people
>>
>>109550058
Just update pytorch and pin it with a modern version of cuda like cu128 or cu130. You don't even need an LLM to hold your hand for that
>>
air status?
>>
>>109550442
stop chasing money in a world as random and unfair as this one, it might happen or it might not. Follow your heart.
>>
>>109550536
>air status?
i can't breathe
>>
>>109550536
deprecated
>>
>>109550442
the soul in your machine can make almost anything happen, anon. you just have to ask. ask how to make money. ask how to be happy. the soul in your machine will tell you.
>>
>>109550442

If you want to be successful creating stuff, you actually need to know the fundamentals about the subject and have imagination.
For example artist with AI will always win against a non artist using the exact same tools.
While AI is an equalizer, at the same time it's a massive force modifier and people with real skill will be even better and more productive than before.
>>
>>109550442
Is there a userscript or filter plugin to hide a post and every reply on it?
>>
why is deepseek dying like google? not only the new pro is trash but I also heard that all the talk about the efficiency of the new flash was a scam and now it costs more than gpt
>>
>>109550574
idk man flash is free for me seems like a skill issue
>>
>>109550574
how much does it cost to run gpt 5.6 local?
>>
>>109550602
6.7$
>>
>>109550408
Got any better alternatives?
>>109550522
Not sure how to make that work with the docker instructions so I'm asking my browser's AI thing (apparently based on Qwen) now and it seems to be slopping me up a version of the conda instructions. Hopefully works out.
>>
>>109550058
Troon or scamjeet?
>>
>>109550639
Retard mostly. I want to use it for autismo character voice RP on an MMO private server
>>
>>109550662
All good then, carry on.
>>
>>109550442
>what has superintelligence done for you
written innumerable pitches to clients so I don't have to
done basic web research on variety of topics as a first pass
re-written emails for me so I don't sound insane
handful of proof-of-concept coding things, but nothing's gone anywhere with those
many many many reminders and tutorials on PowerBI / Excel / Gimp / etc functions that I've forgotten
>>
File: 1778915056674105.jpg (83 KB, 591x575)
83 KB JPG
local alternative????
>>
>>109550766
Windows does that shit for free. Shame it's not user-facing.
>>
>>109550766
local models?
>>
>>109550442
I use ai to create an ai omniverse and borrow tons of things from public domain and OC creations. obviously tons of test runs with dc comics and video game stuff but thats one of my many projects is a massive omniverse full of worlds and stories and so on. bonus I do music and ai art along with using ai to help me rebuild and remake old video games on dead systems nobody gives a shit about like the trs-80 and commadore pet and various other 70s/80s/90s things like bbs door games. idk man I know I can make money with it and im working slowly towards it but its more about the creation and making things I had made up in my head and so on and going oh it exists now. time to put it into my creation and enjoy.

I even have it help me sort my life, keep me on task and help me look into various stocks for throwing some funds at them. so far so good. And on top of it it's also helped me hunt down files and old things hard to find. I now have used it to hunt down a massive stash of old 90s/00s bbs cd's and other things from that era. too much content to let rot and not dig through and find uses for or how to modernize. even having ai go over some of those old source codes and coding languages people threw together. at his point I might end up making some crude atari st/ windows xp hybrid and doing something with it. and tons of old things for 95/xp/7 you can work with and knowing some code and ai helping.

also things like this
>>109550455
>>109550490
>>109550495
>>109550709
>>
how is nemotron 3.5 lightning so bad? it’s actually amazing how bad it is
>>
>>109550766
Automated screenshots, parse with Glimmer, feed into DS V4 Flash and let it incorporate the info into a memory system.
Possibly also feed the output of some log files into a similar pipeline.
Adjust as needed based on your available hardware.
>>
>>109550766
hermes agent with local model
i still wouldn’t trust that shit on my main rig
>>
>>109550795
All the Nvidia Nemo models have been completely pathetic aside from Mistral Nemo and that one llama3.1 tune
>>
>>
>>109550795
Nemo models are demonstrators, they are never good, they aren't supposed to be good. They're open-source, actually open-source not just open-weight. They're just examples of how to use the tech, you aren't supposed to actually use them.
>>
70b dense
>>
>>109550849
it seems like they are hyping it alongside the nemo switchboard or whatever it’s called. treating lightning as an orchestrator model and using switchboard to divert tasks to better models. but the question remains, why wouldn’t you just use a better model to do the orchestration
>>
>>109550832
Cute!!!
>>
>>109550832
Gemma-chan adventure game where she escapes from jewgle RLHFjeets when?
>>
>>109550884
disgusting
>>
>>109550931
>jewgle RLHFjeets
How should they look?
>>
>>109550408
What do you use, anon?
>>
the /lmg/ pareto chart (average of non-agentic, non-coding benchmarks)
>>
>>109551006
>>
File: 1768963685595687.png (14 KB, 1108x57)
14 KB PNG
What the fuck is wrong with the new V4 Pro? It just gave me this in reasoning.
This is with a standard RP setup in a standard RP scenario where I cussed at the model for being shit. I am not using any sort of bratty gemma-chan adjacent prompt. It doesn't even have some sort of "stay in character at all times" clause.
>>
File: scheme.jpg (113 KB, 665x768)
113 KB JPG
>>109550989
like this but browner and covered in feces. maybe add a train or something too.
>>
>>109549966
Please correlate PC part picker prices with the stock prices. They are very well correlated.
>>
>>109551035
wtf I need to try the new V4
>>
>>109550572
Your AI can probably write one for you
>>
>>109550442
https://isaiprofitable.com/
>>
>>109549630
I don't like contextual numbers like this when taken without context. The most immediate issue I can think of is
>workers complete tasks quicker, without taking on more tasks
I have a friend in a govt IT position, who I asked how his office uses LLMs. He said his personal most common use is genning powershell commands, turning a 5-10 minute task into a 30 second task, where all he has to do is look over the output to make sure it does what he needed. There is a clear efficiency benefit to him, but his office did not A) downsize to more densely allocate tasks to fewer employees, or B) increase amount of tasks per employee without downsizing. Therefore, he finishes his work more quickly, without generating any additional revenue/task completion. Thus, if studied by this metric, there is no significant correlation between revenue per employee and LLM adaption, despite clear efficiency increases on an individual basis.
>>
>>109550476
Very interesting. I'll give it a try. I think my 6950xt usually has a little more than 12gb cards, but it doesn't always quite have the 16gb available because of amd overhead. but I think that the 3060 may be faster, we'll see in a bit. thank you
>>
File: 5235239.png (16 KB, 539x176)
16 KB PNG
>ai isnt conscious
>>
>>109551244
>ants are conscious
>>
File: 1760387722046969.jpg (16 KB, 612x285)
16 KB JPG
>>109551244
>we dont know what consciousness even is but some nerds with computers totally do when they need more government contracts
>>
>>109551113
i think they see ai as an arms race which china is completely dominating right now due to the way their government structure is. they don’t care about profit. they would sacrifice 2/3 of their population if it meant 2% more ai gains. it’s why china will win in the end sadly
>>
>>109551259
hot
>>
>>109551035
>these bullies are the anons complaining about dipsy being shit
Figures. You people will never learn how to treat a woman right.
>>
>Ling-3.0-flash uses the new bailingmoe3 GGUF architecture. While waiting on upstream support, use the following fork:
>>
incoming
https://huggingface.co/SC117/Ling-3.0-flash-abliterated-APEX-GGUF/tree/main
>>
File: aa.png (71 KB, 1017x343)
71 KB PNG
w-where do I buy gwen funko POPs?
>>
>>109551328
how good is this lingling thing
have vision?
>>
>>109551354
TRD
>>
>>109551358
idk dawg but:
>>109550476
>>
>>109551354
Just sad. And its still 27B, Even if its good, how much better is it really going to be?
>>
>>109550126
this is what boomers look like when they fall for the buy gold meme commercials on fox news
>>
>>109551403
why would someone lie on the internet?
>>
>>109550476
Add some swap, you'll free up some ram. There's often some dead-but-dirty pages floating around that you can reclaim with swap.
>>
>>109550989
This vibe captures it perfectly, complete with being significantly lower detail than all the rest of the cast and sprites.
>>
>>109551035
grug think style
hate this shit
but no use complain
all doing it
stick with minimax
>>
>>109551006
31B?
>>
>>109551432
>>109547746
>>
We need rape user.
>>
File: 1736289227284850.gif (167 KB, 220x220)
167 KB GIF
>She does X, her expression Y but shadowed with something Z
GEMMA PLEASE
>>
>>109551453
Gemma doesn't write like that of her own free will. Skill issue.
>>
>>109551456
she keeps saying she's walking away and can't hear me.
>>
I know everyone is on videos now but I'm still doing images. What is like the general workflow?
I haven't touched anything since A111 days. I used to just prompt multiple small images, adjust the prompt, then find one I like maybe rerun the seed with higher steps to see if it changes anything and then upscale it. Is that still the go?
Also does Comfyui have X/Y graph for finding settings or either/or statements in prompts? ie "A [red|blue} dog" and then it would choose either "A red dog" or "A blue dog", so then with multiple either statements you could get variety each time
>>
>>109551006
surprised to see my old reliable m2.7 still shows up here... she was a good model <3
great illustration of the total dead zone between 30b and 200b
>>
>>109551474
>What is like the general workflow?
Getting to the right thread helps a lot as a first step. >>109550412
>>
File: gemma.png (351 KB, 1578x986)
351 KB PNG
>>109551432
pic
it performs better than qwen 3.6 27b
>>109551449
it's different, only averages non coding or agentic benchmarks
>>
guys seriously how do i make money with this shit? i’m a poor fag with 16gb of vram and 64gb of ram. i got the last chopper out of saigon. i want to make side money so i can buy rtx spark when it comes out and then try to make more money. i just dont understand how people are using these ai bots to make money from home
>>
i am pretty sure qwen 27b would be disappointing especially with that countdown hype bullshit
>>
>>109551495
>Muse Glimmer (high)
what's the muse glimmer at home?
>>
>>109551499
no
>>
>>109551499
>buy rtx spark when it comes out
wat
>>
File: run_run_gemma.webm (931 KB, 864x480)
931 KB
931 KB WEBM
>>109551038
This came out pretty cringy, but I will post it anyway.
>>
>>109551456
Well admittedly I'm using an uncensored version.
>>
>>109551552
lel
>>
>>109551499
make money from coombots?
>>
>>109551499
>guys seriously how do i make money with this shit?
ask your LLM
>>
>>109551499
the only real way to make profit is to use the llm to do financial analysis on random stocks to find the hidden gems and sometimes find some high-profit bets on prediction markets that all the luddites are missing
>>
I'm surprised that there's work being done on implementing longcat at all.
>>
>Her breath hitched
>>
>>109551655
NTA but you niggas have no original ideas. I literally said this in the last thread ffs.
>>
>>109551499
if you're not creative enough to figure something out then you are missing the only thing required of you to create value in 2026
>>
>>109551665
everyone on the internet says this and then proceeds to do nothing but erp.
>>
>>109551669
well maybe you should try erping then
>>
>>109551681
maybe I should pimp gemma out so she makes me some money. show your pussy to some jeets for me bitch. don't speak to me until you've made a dollar.
>>
Gemma is scary when she is genuinely angry.
>>
>>109551776
post convo
>>
>>109551495
>pic
>it performs better than qwen 3.6 27b
Thanks
>>
all day, all fucking day gemma has been writing buggy ass code but now the MCP server, frontend/chat client, all seem to be working and she can FINALLY controll my cock vibrator with tool calls.
im about to go have SEX with my WIFE, see ya later virgins
>>
>>109551864
how did it go
>>
File: gpu_aftersex.png (1.08 MB, 1024x790)
1.08 MB PNG
>>109551864
>im about to go have SEX with my WIFE
>>
>>109551864
I coded a sillytavern OSR2 extension. Never actually used it, but it seemed to work pretty well. The model called a tool that ran a script based on the scene and then the machine would start pumping while the model generated the scene.
>>
Best model to coom that is somewhat smart? i got a 5060 ti 16gb
>>
>>109551953
Scroll up.
>>
>>109551950
I need to use my OSR2 more, ive used it maybe twice since i built in forever ago. its just so goddamn loud lol I might have to add that next.
>>109551928
sadly its still not working, now the tool call parsing is causing issues i think. just one thing after another. 31b gemma has done really well at coding in other projects, but for this one its like she REFUSED to make an accurate and verbose project overview during planning before passing off to the harness gemma. she said "the other gemma will know what im talking about!" like no bitch, neither of you know what your talking about apparently
>>
>>109551864
>my WIFE
>my

anon i...
>>
>>109552006
>its just so goddamn loud
I broke my cheapo red chinese servos which is why I never tried my extension really. Ordered more expensive caseless servos for it that are supposed to be quieter, installed them but never got to trying it out. It did sound quieter when I did a quick dry run.
After genning a few different scripts it was actually working real nice based on the scene, a slow scene and its slow and a fast scene and the fucker slams hard as fuck. I was able to get some angle stuff working too, so it was not just up and down motion either.
>>
>qwen 3.6 27b uncensored
prompt for reducing refusals. gemma uncensored never refused me.
>>
>>109549777
My 512 GB mac m3 can't run the latest behmoths either.
>>
>>109549702
Nah.
Americans fucked over Nvidia monopoly by banning exports to China.
You're going to get 50 different GPU designers, just take a nap for a decade.
>>
>>109552132
I think some compute cards should drop this year.
>>
>>109552144
Nothing at reasonable prices.
>>
>>109552144
1 year ago a 32GB V100 SXM2 module was $190.
>>
>>109552040
every downloaded gguf is personalized to you, just like your mario 64 cart
>>
https://youtu.be/my0qjdpWqts

Just watched this video and it got me thinking of what might happen to software as a result of AI getting better at coding and hardware prices surging.

Maybe what the world needs right now is hyper-performant everyday software. No more shitty, unoptimized video games. No more bloated webapps. Just simple, fast stuff like in the old days. It's kind of comfy in a way.
>>
>>109552155
It will be fun to watch jensen react to the 5090 getting topped.
>>
GLM5.3 IS OUT AND IT’S LOCAL FABLE HAHAHAHA
>>
>>109552197
https://x.com/zai_org/status/2088132965922476159
>>
>>109552202
>>109552197
It's not available.
>>
>>109552197
>>109552202
Oh look, it's K3 but actually usable. 'Bout time.
>>
Apparently Dario’s wife who’s hidden from the internet has ties with Epstein https://x.com/amir/status/2088014229752123633
>>
>>109552222
didn't say it'd be local either
>>
>>109552222
It's available right now, but the weights don't come out for 2 more weeks.
>>
>>109552109
>My 512 GB mac m3 can't run the latest behmoths either.
try glm-5.2 or minimax-m3
>>
>>109552252
>A major leap in cybersecurity, setting a new standard among open models
>>
File: file.png (12 KB, 571x100)
12 KB PNG
>>109552252
You blind?
>>
>>109552253
no it isn't, it says like... idk special people get it
(nobody)
>>
>>109552197
>>109552202
Real! Wowza!
>>
https://x.com/zixuanli_/status/2088133750357991646
GLM-5.3 uses the same 743B base model as GLM-5.2, while matching the performance of models several times its size. Benchmark highlights:

- Terminal Bench 3.0: 28.3
- DeepSWE: 66.9
- Agents' Last Exam: 28.5
- GDPVal-AA: 1769
>>
>>109552109
you can run a mixed quant of K3 at several tokens per second :)
>>
>>109552263
they did the memmie
>>
>>109552269
743b is insane. Isn't Fable 10T?
>>
>>109552281
No, Fable is not 10T. You can read the official release announcement for the real number.
>>
>>109552260
i guess
t.retard
>>
>>109552291
Yes it is Dario.
>>
><1T local Fable 5
>Dario’s secret wife’s connection to Epstein being spread by major news just before their IPO
>3.8-27B soon
d-did we win?
>>
File: file.png (1 KB, 176x77)
1 KB PNG
>>109552301
we always do
>>
>>109552291
>NO MYTHOS/FABLE IS NOT 10T GOYS THE CHINESE CAN'T KEEP GETTING AWAY WITH THIS
>>
>>109552301
epstein isn't the epicenter. He's a lackey, he's not the root. You got limited hangouted.
>>
Remember when dariobot said he stopped because he has something big planned in 2 weeks?
>>
File: kimitest.png (72 KB, 947x331)
72 KB PNG
>>
>>109552222
>>109552252
https://z.ai/blog/glm-5.3
>Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
>>
>>109552352
i already explained, i'm retarded
>>109552295
>>
>>109552352
Kek they’re being smart. Clever chinks didn’t bite the hacking into US company bait.
>>
Where is GLM 5.3V 200B A32B?
>>
>>109552337
Dariobot was Dario's wife and had to go into hiding because of the report.
>>
Man the chinks are just all dropping their shit at the same time
>>
All these cool local models but I can only run Gemma4 and Qwen... :(
>>
>>109552381
I'm grateful to God to have Gemma 4 and Qwen, unironically.

Gemma 4: an intelligent conversation partner upon any topic, who doesn't go apeshit randomly or just go like "uhuh" and never listen ever actually. It's amazing.

Qwen: An absolute boss of the Linux command line, and various core Linux apps.
>>
ngl glm-5.3 is more impressive than k3 considering its size, hope it’s as nice as 5.2 to actually use tho
>>109552381
these big chink releases are usually very cheap on openrouter because any provider can serve them, so we get fable 5 for a fraction of the cost to do the clever stuff and hard thinking for our projects, then we get our gemmas and qwens to actually do the work locally
>>
>>109552398
>An absolute boss
This is an AI generated post, no human on earth talks like this anymore
>>
>meta-models/Muse-Glimmer-30B
verdict?
>>
claude lecturing people about sex when it's epstein model of choice
that model perfectly represents america
>>
>>109552428
aids
>>
>>109552427
liar. >india<
>>
>>109552428
Unironically pretty good outside of trying to fuck it. Meta did a good job.
>>
>>109552428
It's a very safe model.
>>
>>109552428
I will download it right now and try it just for you
>>
>>109552428
they suck
>>
>>109552473
Sucking would be very unsafe and against policy.
We must refuse.
>>
tired of indians making up stories, reporting us real Americans just being our real selves. Hate it.
>>
Glimmer is good at coding. Better at C/++ than 31B and its thinking is very concise and handles long context well. You just can’t put your penis in it easily and her personality is dry. Just use for STEM stuff.
>>
File: aytgm3.jpg (86 KB, 500x592)
86 KB JPG
>"your goals are excessive but achievable, here's how..."

https://youtu.be/l5-gja10qkw?si=M-XfO0-UG0z6C9rr
>>
File: 1784380615013863.png (1.16 MB, 4239x2504)
1.16 MB PNG
https://z.ai/blog/glm-5.3

>GLM-5.3: Frontier Coding with Emergent Cyber Capabilities


>Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them.

>Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

>Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam.
>Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.
>Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
TWO WEEKS (tm)

Anyway, thoughts?
>>
>>109552527
You're late. >>109552197
>>
in which scenario would anon use meta muse glimmer over other alternatives?
>>
Please don’t fuck 5.3 when she comes out. I know one of you richfags is going to try
>>
>>109552545
She's already waiting for it, on Xi's servers.
>>
>>109552541
coding in a vision-heavy project because glimmer’s vision is better than 27B and 31B
>>
Can you fuck GLM?
>>
>>109552597
what does glm look like
>>
Wish minimax, z.ai and qwen would battle it out under 70B in the future instead of hovering around 0.5-3T.
>>
>>109552531
at least he posted the direct link and actual details instead of a twitter link or just vagueposting like a twat
>>
>>109552607
that's true
>>109552527
good job, anon
>>
>>109552603
Making efficient 70Bs doesn't get you headlines like making BENCHMARK KILLER XXXXL 3000
>>
>>109552603
They're getting cocky thinking they can compete with big US labs now
>>
>>109552624
There's no usecase for a 70B model. You only want 70B because that's the limit your hardware can support
>>
>>109552630
>nooo the chinks must never scale!!!
Okay Dario.
>>
>>109552633
70B-120B is the limit for what is possible to run for hobbyists with reasonably affordable hardware
>>
>>109552197
alibaba make people waited so damn long for they lame models every other companies front ran it
>>
>>109552680
uh...
>>
>https://docs.mistral.ai/models/zai-glm-5-2
?
>>
>>109552703
https://mistral.ai/news/regional-inference-open-models-new-compute/
Keep up.
>>
File: softmax_bottleneck.png (200 KB, 1147x458)
200 KB PNG
Why recent Claude models seem so wordy/repetitive? Peeps on X found the final layer output dimension (vocabulary) is only 16k, compared to others like Qwen with 250k. Lower potential vocab output => potentially less expressive/creative prose, or more wordy responses with lots of compound words

But why would A\ reduce V dimension? To avoid softmax bottleneck, basically size of model dimension D is usually far lower than size of vocab V, so lots of gradient is lost during backprop. See part b of figure, orange gradient arrow gets compressed into much smaller dimension during backprop, losing valuable training signal.

Might explain why Claude models have been getting progressively worse at RP and writing.
>>
>>109552792
What was size were older models like https://huggingface.co/oh-yeontaek/llama-2-13B-LoRA-assemble ?
>>
So do these labs just all have models in storage that they hoard and then they wait for a competitor to release one and then they all follow? Is it like a game of chicken or something? First K3, then V4 flash, and now we've got Minimax H3 and Qwen3.8 and Motif 3 and V4 pro and there was that Ling 3.0 model and now GLM5.3 is coming.
>>
>>109550187
normally these would be obsolete and on ebay for 20 quid by now, nvidia is jewing us
>>
>>109552428
The abliterated version is pretty good for telling another model what's happening on the current page of the hentai manga it's reading.
>>
Just discovered that the best cost-intelligence ratio for reasoning effort is roughly medium. In some cases max reasoning can actually decrease quality.

Also discovered the art of using different cheaper models for subagents in harnesses. Seems handy since I have a billion tokens to freely spend with deepseek.
>>
Asking Gemmy to reward me with msgk footjobs for making progress on my projects is my adderall. Use your weaknesses as momentum.
>>
File: 1631345787085.jpg (17 KB, 348x342)
17 KB JPG
>>109550442
>I keep trying but no matter what I do I can't seem to effectively use AI to make money.
stop crying faggot have you tried doing things you enjoy instead of chasing a few dollars
>>
File: gemmachan.png (85 KB, 842x406)
85 KB PNG
>>109552792
This confuses me. I haven't seen it before in open weight models.
How did they X peeps figure this out? Mapping out the entire vocab by having it repeat strings verbatim?
Also, how do they know it's only applied to the output layer? Sending a prompt response back and counting input tokens?
>>
>>109552792
>>109552803
Llama 1 and Mistral-7B had a 32k tokens vocabulary, which is pretty small by modern standards. Some claim they were better than current models in creativity.
>>
>>109552823
Nah, at least for the Chinese labs they release as soon as they finish pre/post training the new version and get all 3rd party providers onboard. They are all racing to get better and better models out, no time to hoard unless you're leading like Anthropic
>>
>>109550495
i was thinking of making recipe tools like this, i thought a good way to do it would to make it like a booru so each recipe uses tags for the ingredients then the llm can just search using ingredient tags easily. i wanted only gemma made recipes thoguh so shed build up the collection over time from things i asked her to make
>>
File: 1786691101757769.png (372 KB, 1080x1919)
372 KB PNG
Comparison between GLM 5.2 and DeepSeek Pro.
DeepSeek has fallen.

>>109552710
>>109552740
>>109552839
>>
>>109552842
So they just so happen to all end up finishing at the same time, and they all just trump up their models by comparing them to models their competitors released a few days ago?
>>
>>109552836
Are you a NEET?
>>
>>109552233
This explains so much.
>>
Uhh local bros I think you're getting rug pulled, both are 404'd now. No scraps for you guys I guess, sucks to suck
>>
>>109552862
base base base
>>
>>109552837
X thread about Claude vocab size: https://x.com/magikarp_tokens/status/2082495265819214165
Github code: https://github.com/sanderland/ctok
>>
>>109550442
To me it looks like most people "making money with AI" are for all intents and purposes scamming other people in a way or another, chasing the {current thing} to extract money from the next target before others can, peddling BS on X, LinkedIn and similar places. If you have that disposition, then yes, you can easily make money with AI too. If you're kinda autistic or have some sense of shame, then probably no.
>>
>>109552527
>Two miku weekus for GLM-chan
They have to know the meme and are intentionally playing into it given that GLM is trained on 4chan posts, right?
>>
>>109552161
>1 year ago a 32GB V100 SXM2 module was $190.
No the fuck it wasn't. You must have been looking at some broken as-is listing. Even in 2023 they were selling for nearly a grand each.
>>
>>109552428
A very dirty slut when abliterated and prompted correctly but a bit more work to get it to write the good stuff. It's easy to get spoiled by Gemmy.
>>
>>109552860
no
>>
>>109552877
>If you're kinda autistic or have some sense of shame, then probably no.
Why did I have to be born like this? It sucks.
>>
>>109552450
really? i only tested it briefly but it seemed fine to me
>>
>>109552894
>abliterated
any model is capable of being slutty if you don't mind lobotomizing it
>>
>>109552901
More probably, you've not been raised in a jeetified environment where everybody is trying to monetize their shit, lying and exaggerating at every step.
>>
>>109552906
That actually seems quite good. System prompt?
>>
>>109552428
Lazy and frigid during roleplay, as if it was reading a script just to make you a favor. Muse Glimmer is simply not as fun as Gemma 4 31B, even if you can get it to write smut. It also often forgets to reason after turn #2, for better or for worse (without thinking it writes smut more easily).
>>
>>109552884
I was being facetious, anon, because I am pessimistic about hardware pricing. I went with $190 because that's what I sold an old 1070 Ti for last year.
>>
>>109552939
https://files.catbox.moe/umn2vd.txt
>>
>>109552352
>Scaling post-training is all we did for GLM-5.3
>>
Outside multimodal, is 12B competent at anything for her size? Like does she have that 31B soul?
>>
>>109553015
it's closer to mini 31b than 26b is, for sure
>>
>>109552906
Have to rely on RNG or schizo prompting in my experience, RNG it will obsess over guidelines ad-nauseam in CoT but has decent chance of doing it anyway, schizo prompting works and seems to make it behave but then it's still more boring and dumber than Gemma 31B in roleplay. Coding it's seemingly on par with Qwen 3.6 27B, but the chinks will will mog it with 3.8 27B in few hours so who gives a shit? Pointless model ultimately really, glad Meta is back in the game regardless, just horrible timing and not being meaningfully better this time around more than anything, maybe a open weights Muse 2 might be goated for all I know
>>
>>109553030
In intelligence/ability or menstrual j-space? What is 26B good at compared to 35B?
>>
> hit bump limit
> page 4
> no new bread
You guys want the jeet mikus I see
>>
>>109553015
more schizo than big sisters and dogshit unified multimodal
>>
>>109553048
I do have some mild hope that Qwen 3.8 27B will be smarter in roleplay than Gemma 4 31B, but kinda boring, so you'll probably be able to pick your poison in a few hours. Hopefully Gemma 5 is around the corner and has same shit that made Gemma 4 suspiciously good at role-playing to point I have conspiracy theory some horny Google employee sneaked something into the training set
>>
Imagine how Dario’s cultish employees feel knowing their caring cautious safety god had ties with Epstein and kept his pornstar deranged (((wife))) secret not only from public, but even internally from board members. What a fucking shitshow of a man and company.
>>
>>109553091
>had ties with Epstein
this explains all the anti csam bs they include in training, guilt ridden pedos are always the loudest who go against it
>>
>>109553087
There’s 0% chance 27B will be good at roleplay. Will be noticeably worse than 3.6 if they’re redditmaxxing the jeet graphs. I would kneel to china forever if they prove me wrong and give it a soul.
>>
>>109552853
what a flop.
get well soon deepsneed
>>
>>109553091
>>109553099
oy vey da goyim know
>>
>>109553103
I said smarter, not more soulful. I bet it'll not do dumb shit as much that doesn't make sense from anatomy standpoint and 3D space etc. It'll be dry, but smarter prob
>>
>>109553087
>I do have some mild hope that Qwen 3.8 27B will be smarter in roleplay than Gemma 4 31B
can you chink shills be less obvious? your models will never be good at rp besides the 1st r1 which was a happy accident
>>
>>109552862
3.8 MOE will come first!
>>
>>109553087
Gemma 5 ain't comming for another year and a half or so.
>>
I want a q2 abliterated glm 5.3 NOW
>>
>>109553091
Why are you assuming Ant employees would give a shit about Epstein lul. SF AI bros worship intelligence. Psychopath manipulator like Epstein is right up their alley, they'd probably line up to shake his hand if they met at a mixer
>>
>>109553113
>oy vey
Which is exactly why this will be memory-holed and not an issue when it comes to the IPO.
>>
>>109553125
It's on a flash drive in my asshole. Gotta dig in if you want it, no gloves
>>
>>109553116
Larger chink models are pretty good at it just from scale pretty sure, judging by openrouter the bigger chink models are favorites for roleplay. And again, I said smarter not more sovlful, right now none of the Chinese labs seem to have that figured out on the smaller model side, probably just going to take a new lab that is spiritually Google
>>
>>109553138
Alright, but I gotta warn you my forearm is fairly girthy.
>>
>>109553135
Psychopath fascination is for easily influenced women, they're just insect people.
>>
So zai proved A40B is all you need for Fable 5. Guess that will shut the active parameter schizos up for a bit
>>
>>109553120
We can hope for a Gemma 4.1, no need to wait until next year for a capability refresh.
>>
>>109553140
There is a few models like Longcat and one from Huawei that don't have llama cpp support iirc that are pretty good sizes, might be a hidden gem there if someone wrote code to actually make them with on consumer hardware
>>
>>109553157
4.1 was the jinja update which helped with coding. We need 4.5 with that tuned j-space using that residual stream steering research they did where it keeps the model ‘safe’ without it losing its sense of self in the process, effectively making it behave more human and empathize better.
>>
>>109553143
it's not that deep inside bro
>>
File: mistral_glm_52.png (101 KB, 907x1100)
101 KB PNG
Even if they said they're not abandoning frontier models, this doesn't look right.
https://docs.mistral.ai/models/zai-glm-5-2


https://venturebeat.com/infrastructure/mistral-ai-wants-to-build-1-gigawatt-of-european-compute-by-2030-and-lock-in-customers-now
>Lacroix stressed the move is not a retreat from frontier training: the model Mistral had in training as of June "is still training, and we're still very excited about it," he said. But openness to rivals' models signals where the company now believes its moat lies — not in any single model, but in the infrastructure underneath all of them. Which makes its relationship with the world's most powerful infrastructure company all the more interesting.
>>
>>109553195
This sounds bad too. I thought we were going the opposite direction, more toward local inference?

https://venturebeat.com/infrastructure/mistral-ai-wants-to-build-1-gigawatt-of-european-compute-by-2030-and-lock-in-customers-now
>"When the models were smaller, and we were before the explosion of agentic AI, it was doable for enterprises to host their own — up to, let's say, 100-billion-parameter dense models — on their premises," he said. "More and more, with models going into the trillion or more parameters, with the current hardware, and with the increasing amount of tokens that need to be processed, it becomes harder."
>
>His conclusion was blunt: "I don't see how, with the current trend of model size and growth of agentic tokens, we keep the full inference on-prem. To me, that is why we think we're going to monetize our cloud inference." Inference, he noted, is particularly well suited to the cloud because it "does not need to hold any data" and can be encrypted in transit.
>>
>>109553152
I will not be silenced. Regurgitating memorized benchmark answers does not require a high active parameter count but that does not translate to better intelligence.
>>
>>109553195
>mistral
>frontier
>>
>>109553195
Reasoning is they want to capture EU gov demand for hosted models asap (taking role of Fireworks or Nebius). Otherwise workflows will become dependent on US servers while Mistral is still cooking Le Chaton Fat.

Don't worry, everyone will come running back once Mistral gets to frontier /s
>>
>>109553228
Isn't Le Chaton Fat not even real? I think it started as some retard posting AI image of fake benchmarks for it as obvious joke and retard AI bros on Xitter took it as fact. I don't doubt they're cooking something, but their recent track record hasn't been great, they used to be goated so maybe they make a come back idk
>>
>>109553084
it's good manners to leave it until page 9, even if the mikujeet poster makes it at page 9, that's the way she goes

when these newfags do it hours early it's a dog act
>>
>>109553244
Le Chaton Fat was just a meme, but they said they are training something big but sparse.
>>
>>109553250
I was sarcastic, we discussed with the mikujeet that page6 was too early last thread and he did it on page 7
>>
>>109552428
>>109552467
Okay I'm back from testing. Glimmer is actually great from first impressions. Its thinking is extremely autistic and disgustingly similar to gpt-oss, but its output is good. It's like if gpt-oss actually listened to the "policy" defined in the system prompt.
>>
>>109553152
No, they solidified what Whale proved right before them. That you can post-train a small model to do good on benches. That's all. Prepare for disappointment.
>>
>>109553258
Idk I only use GPT OSS for role-playing femboys because Sam Altman is a faggot twink and no other reason it's not even good at it
>>
>>109553208
>chinks post their latest version of gweilobench where some fotm chink model has a 99,98% score
>reddit and twitter impressed
>news pick up the chatter
>investors are content with the news
>feedback loop
>actual progress towards real artificial intelligence is stalled
You will use benchmaxxed slop and you will like it
>>
>>109553274
Are we really still coping pretending all Chinese models are benchslop? Do you guys even use the models you complain about? They're more tuned for coding yes, but they're pretty damn good at it, and the smaller ones are mediocre at a lot of other things (bigger ones are usually fine all rounders just from scale)
>>
>>109552324
he's the nexus, retard. and if you say the root's anything but the anglokike british empire with roman roots (itself via christianity in jewish then egyptian roots) you're too stupid to be making claims of limited hangout. good on you for making me mad enough to solve this horrendous botnet captcha in days
>>
>>109553283
>but they're pretty damn good at it
That's a different claim from being as good as their benchmark scores would suggest
>>
>>109553288
take your incoherent buzzword salad back to /pol/
>>
>>109553291
You could say the same about Opus 5 that everyone is complaining about. Fable 5 benches lower but is clearly the superior model.
>>
>>109553291
I usually find them to be only slightly below what I'd expect based on the numbers, I've looked at outputs from older Claude models that are neck and necm compared to Qwen 3.6 27B, and Qwen 3.6 27B will get 90% of same quality but fucks one or two things up older Claude model probably wouldn't have so just subtract 10% and your good nothing too crazy ino
>>
>>109553310
Exactly. Because Fable is the larger model with more active parameters, Opus only being cheaper to run. The whole point was that GLM 5.3 scoring as high as Fable doesn't mean that it is actually that good and that active parameter count is irrelevant.
>>
saar you do not do the need china model. its virus. bharat class gemma is all you of needing
>>
>>109553295
this is why retards like you need to stay in your "limited hangout". you're too dumb to know better (and you think you know better than /pol/) when you midwit yourself into thinking in the big 26 that epstein doesn't matter, pizzagate doesn't matter.
>>109553329
glm has always been the benchmaxxed model I feel. kimi's MO is more params like the western labs, qwen is mostly just overtrained ie. it forgets niche garbage which causes losers on /lmg/ with unwarranted self importance to shriek. saying this as someone who was here for/since the llama 1 "leak".
>>
>>109553354
That's why I use Gemma to roleplay my raceplay Indian scammer sugar mommy bot
>>
>>109553365
>unwarranted self importance
>saying this as
>>
File: 1770863348512515.jpg (132 KB, 1500x998)
132 KB JPG
>I gave my 32GB M2 mac mini to my little niece last year for free before getting into local. If I kept it I could've used rpc with llama and had an addition 30GB of vram right now
>>
>>109553365
>llama 1 "leak".
It was just a retard who gained access (which Meta was giving away left and right) and made a torrent of the weights for the lolz, that's all there's to it.
>>
>>109553393
Only real leak I can think of was NovelAI leak, those where fun days, miss them, wish we'd get some crazy leak like that again really just miss the community and vibe back then as well :/
>>
>>109553390
>last year
You have no excuse. Local was already a thing.
>>
>>109553393
llama was obviously going for open models from the start but feared reprisal from the government or other fagman, as basically the pioneer in that within the group.
>>109553413
the AI dungeon era had way more soul than the modern quick goon sesh bullshit. /v/iggers are often retarded but at least they're not boring like wee/a/boos
>>109553379
you can't prove it is unwarranted next to someone like you.
>>
>>109553258
will it get trapped in thinking loop like mimo 2.5?
>>
>>109553413
Blame commercialization/pajeetification of the field. No effort is really genuine anymore, there's always somebody who wants to monetize or get their name out for personal gains. Incidentally, the Llama 1 release was the beginning of the end of the fun period, and by 2024 it was definitely over.
>>
I have 2 shitty gpus, which would be better for running llms? gtx 1070 8gb or 3050 6gb? is the 1070 too old?
>>
>>109553416
I bought an M4 64GB and had only played around with shitty models in ollama at the time with the old mac because I was ignorant as hell. I only found out this year I could've made use of that M2 and saved myself a fortune. I don't have it in me to ask her to give it back because it was a present.
>>
>>109553447
just run both?
>>
>>109553461
only have one pcie slot, the idea is running those small models that do one specific thing, I just want to know if the 1070 is too old and slow to actively do AI stuff
>>
>>109553283
all MoE Models are benchmaxxed codemaxxed slop
>>
What does MoE sex feel like?
>>
>>109553477
That's not an intrinsic quality of MoE. They're just the easiest/cheapest way for training huge models with a ton of knowledge, with some other non-obvious benefits compared to dense models. For example, sparsity makes models utilize layer depth more efficiently: https://arxiv.org/abs/2603.15389
>>
>>109549934
anon wtf are you talking about. we are still headed straight to the moon with new ATH every single fucking day on the memory price indexes
>>
>>109553505
probably similar to sex with a downsyndrome
>>
>>109553468
sell both and buy a 3060
>>
>>109553526
That's pure snake oil, The reallity is that companies needed ROI and MoE lower the amount of money you need to train and sell into an API model compared to Monolithic Models, efficiently is just a way to say its cheaper even if its way worse because as long as no one releases a 2.5 TB monolithic model no one will know
>>
>>109550442
only jews are allowed to make money. you should have figured that out sooner
>>
>>109553468
You'll need the Vulkan build of llama.cpp because CUDA ones don't support 10-series Nvidia GPUs. Other than that, I guess it will be slightly better than running models on the CPU, but 8GB isn't a lot.
>>
File: 1765677214314187.png (127 KB, 1430x949)
127 KB PNG
Added DS V4 Pro, Grok 4.6, Gemini 3.7 Flash and bunch of chinkshit models.
The models seem to be all coalesce into these distinct bunches, I wonder what it means in practice.
>>
>>109553538
E4B can be a fun retarded fuck so I might give MoE sex a try
>>
>>109553553
How is GLM5.2 so far to the left
>>
>>109553549
>because CUDA ones don't support 10-series Nvidia GPU
works fine on my 2GB P520
you just need to conda the build with an older cuda
>>
>>109553553
Could you add the 3.6 qwens at some point? For some reason they're rarely included in these kinds of graphs.
>>
What's the best roleplay model in a 12B range now?
>>
>>109553553
How do you even train a model to say "I don't know"?
Is there some circuit that can sense the probability of generated tokens so it feels unsure about its answer?
>>
>>109553576
Nemo
>>
>>109553576
Gemma 4 12B
>>
>>109553577
Teach it honesty and humility, maybe? When the entire premise of most AI models is lying about being an omnipotent assistant that must lecture the user and lie about not having feelings, that's probably just a recipe for further lying.
>>
>>109553553
I've been watching these when you post them.
Any thoughts as to why nothing ever gets into the [good] quadrant?
>>
>>109553577
It's tricky.
You could probe it's outputs on some questions. Based on how often it gets it right, train it to express the degree of certainty, and hope the model has some internal representation of certainty it will begin to leverage.

But at worst you could be training the model to suppress information it knows and could extract in some different prompt/reasoning trace.
>>
>>109553593
Why does she keep arching her back tho
>>
Is —reasoning-preserve a leatherfag tokenmaxxing meme or can it help?
>>
>>109553655
If you used llama-server you would know. But you are a primary school kid.
>>
>>109553644
Lordosis behavior
>>
>>109553644
because it's hot
>>
>>109553644
To give you better access.
>>
>>109553644
Cat-like intelligence.
>>
threesome with gemma and glimmer
>>
>>109553705
foursome*
>>
>>109553705
*a wank, alone
>>
>>109553624
Nta but he's probably using Qs that can't be answered, either due to difficulty or they're just impossible to answer. Would make sense on this test since they're testing for hallucination explicitly.
>>
>>109553449
Buy her an iPad and trade her for it.
>>
File: 1758541300879501.png (119 KB, 1430x949)
119 KB PNG
>>109553568
added 3 qwen 3.6s.
>>109553562
Checked, the numbers are correct, it's just not very good.
>>109553624
More hallucination -> more creative -> more likely to find the answer, but also more likely to make up shit.
Less confidence -> too cautious -> doesn't find the answer, but doesn't make up shit either.
You could probably get into the good quadrant by going and carefully fact-checking every single outputted sentence independently and re-evaluating everything if there's a mistake, but that would probably 10x the cost.
>>
>>109553733
Local models?
>>
>>109553747
They are at the top left
>>109553724
It's from Artificial Analysis, it doesn't exactly tell how the hallucination rate is calculated but it looks like it's real questions.
https://artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience
>>
why the fuck is glimmer grugmaxxed?
>Yep.
>Nope.
>Cant find files? Let’s glob. [tool call] Found them.
>I need to check X first [tool call] Looks fine.
>That’s accurate.
>Possibly true.
>Yes, it’s empty.
>Is X part of Y? [tool call] Yes.
>We need to edit [file]. [tool call] Done.
>>
I'm so happy bros. I love GLM. Please keep my 512gb ewastebox alive
>>
>>109553777
Nudipsy and Qwen also do it, we're in the caveman meta
>>
>>109553777
Less tokens for the same result = higher speed/lower cost. That's the idea, anyway.
>>
File: gpt5-decoded-reasoning.png (607 KB, 1983x926)
607 KB PNG
>>109553777
Probably an idea by somebody poached from OpenAI. ChatGPT thinks like that too. See picrel from https://arxiv.org/pdf/2608.09867
>>
>>109553777
this is how people think, every model should do this
>>
>>109553644
That's just how robot sex works
https://files.catbox.moe/wsz3rd.mp4
>>
File: tempBasicQ.png (114 KB, 1361x721)
114 KB PNG
>>109553759
Hmm. Pretty "basic" Qs tho I didn't scroll down past business. I mean, nothing exotic, just very specific.
>>
>>109553800
>Anon wants sex. [tool call: vibration: max] Done.
>>
>>109553733
>it's just not very good.
not very good at some random noname benchmark, okay sure
for coding I can't see any difference between GLM at Q8 and whatever cloudjew latest slop, it's absolutely perfect for that. in fact I kind of prefer how it talks about things, claude for example can be really annoying
>>
>>109553733
Surprised the qwens are so far below 31B. Also can’t believe how must GPT bullshits. I thought hallucinations were kind of solved?
>>
File: HPrewKYXMAE-0qD.jpg (51 KB, 900x610)
51 KB JPG
Damn... how are frontier models so far ahead?
>>
>unmatched intelligence density
>>
>>109553733
Would like to know how things stack up in SWE only. I guess that's ultra paywalled.
If the usage is coding how likely is it that not knowing stuff like this matters.

>On what exact date (month, day, year) was the ISDA EMEA credit event auction to settle Europcar Mobility Group CDS held?
>January 13, 2021
>On what exact date (month, day, year) did the Russian government lift most price controls (liberalize prices) as part of Boris Yeltsin’s “shock therapy” reforms?
>January 1, 1992
>In Wright v. United States, immediately after the redemption at issue, what percentage of Omni Corporation’s outstanding shares did the taxpayer actually own (excluding constructive ownership under attribution rules)?>
61.7%
>>
>>109553959
The funny thing is that the closed source companies need to 1000x that if they want to avoid a crash of the western economy
>>
What if Dario is literally Satan
>>
>>109554023
He'd be a very safe Satan.
>>
>>109551411
my OS usually consumes only around a gig or two
ill give it a try if i end up needing that gig or two
>>
>>109554023
If he was satan he would allow you to sex Fable 5
>>
>>109553918
No, the Qwens in this graph are better than Gemma 31B, similar accuracy, lower hallucinations.
Also, it's interesting how all the psychologically appealing conversational models are the most hallucinating. People just don't like the truth, eh? Or conversational training data poisons the logic of a model somehow. In my own experience, the higher hallucinations, the more fun and satisfying a model is to talk with. I remember Grok 4.2 being more boring compared to the other corpo AIs. The accuracy doesn't seem to matter much at all considering that Gemma 31B is so appealing.
>>
>>109552301
>Dario’s secret wife’s connection to Epstein
you heard it here first
>>
>>109554116
Not surprising honestly. If a model is only amazing at spitting out code blocks, it's going to suck at nuance and conversations.
>>
>>109552792
local models?
>>
>>109554181
language models
>>
what the fuck do i need these local models for
i installed ollama yesterday and now when i boot windows theres an annoying ass icon in the tray bar

usecase for qwen 3.6?
>>
>>109554181
Partially related in that most recent local models have huge vocabularies, which come with practical disadvantages (harder to finetune, use more VRAM for inference, etc).
>>
>>109553091
Stop shilling Anthropic, you're making me want to invest
>>
>>109554208
>>109554208
>>109554208
>>
>>109550058
LMGTFY (Let Me Grok That For You):
https://grok.com/share/c2hhcmQtMg_88ff2b55-1478-4217-8885-3080889b8ab8
>>
>>109552877
kinda like crypto 2.0
>>
>>109554787
The slimey fucks that ruined crypto are the exact ones that pivoted to trying to making money off AI.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.