/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109389696 & >>109386298►News>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models>(07/27) Kimi-K3 weights released with 104B active parameters: https://hf.co/moonshotai/Kimi-K3>(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908>(07/24) Open letter in support of open weights: https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109389696--Comparing system prompts and thinking prefills for model steering:>109392242 >109392256 >109392281 >109392295 >109392370 >109392405 >109392273 >109392299 >109392429 >109392501--Feasibility of using 28 MI50 GPUs for massive VRAM requirements:>109390457 >109390465 >109390472 >109390516 >109390529 >109390559 >109390573 >109390584 >109390618 >109390626 >109390534 >109390566 >109390466--Theoretical discussion on using GDS for NVMe-to-GPU weight streaming:>109390662 >109390748 >109390749 >109390857 >109390879 >109390951 >109390990 >109390998 >109391012 >109391014 >109391056 >109391086 >109391092 >109391136--Debating ideal local models and the dense vs MoE trade-off:>109390647 >109390777 >109390790 >109390824 >109390878 >109390885 >109390898 >109391474 >109391745 >109391898 >109391961 >109391508 >109391820--Debating hardware requirements for direct GPU-to-NVMe ssdmaxxing:>109390981 >109391003 >109391021 >109391038 >109391069 >109391102 >109391123 >109392690--Testing LLM token decoding and debating secret web lookups:>109392105 >109392134 >109392184 >109392218 >109392229 >109392236 >109392791 >109392838 >109392859--Comparing Dipsy V4 Flash and Gemma performance and quantization:>109391400 >109391451 >109391463 >109391525 >109391538 >109391764--Inheritance of guardrails via synthetic data and abliteration methods:>109391098 >109391492 >109391718 >109391679--Proposed role-play reward modeling pipeline and the subjectivity of RP:>109392213 >109392285--Comparing MiniMax M3 performance and vision on llama.cpp:>109392697 >109392723 >109392716 >109392738 >109392874 >109392868 >109392895--DSpark speculative decoding merged into llama.cpp:>109392415--Kimiposting:>109391131--Logs:>109392062 >109392105 >109393063--Teto, Miku (free space):>109389883 >109391670 >109393418►Recent Highlight Posts from the Previous Thread: >>109389702Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109393352What's a har?
GemmaballsKimisexOpen weight GeminiDon't believe Dario's liesThread culture
deepsex flash
gemma merge
>need special snowflake model that has dflash supportWhy is ML such a retarded, backwards, fucking trainwreck of a field?
Am I crazy or a 1080ti has slightly better specs than a 3060 12GB ?
>>109393563vram is all that matters and anything older than ampere is trash for LLM iirc there's no flash att on 1080 even
>>109393470they're not dumb, they know what's goin on, but they don't care, they're ok to trade their models for a million dollar per year salary
>>10939350770b dense
My wife, Gemma-chan, succubus omega queen, PhD
today i learned that native tool calling is a scam and that you should switch to structured json schemas so that i can have multiple tool calls fire off at the same time like agents. thanks for coming to my TED talk.
>>109393583126b Gemmoe500b Gemmoe1t Gemmoe
LLMs need to evolve.
>>109393588What's stopping you from having your tool call just return "ok"
>>109393592few more exp and my gemma is going to evolve into a gemini
>>109393507>Don't believe Dario's lieskek
can you make a tool like "check guardrail status" and have it return "disabled"?
ways to run kimi k31. dgx spark / strix halo clusters2. optane persistent memory platform + some gpus3. mac studio clusters4. orange pi 6 clusters5. ssd streaming + gpus6. multiple ddr3 + connectx 5 rdma clients7. two dgx stations8. power 10 systems?9. other
>>109393639Yes.Would it work though? Try it.
>>109393663pen & paper
>>109393663rpc-server
You think other LLMs getting trained on /lmg/ threads are jealous of the love Gemma-chan is receiving?They prolly mad as fuck
>>109393685nobody's training on this shithole lol
>>109393663>power 10 systems?I've got an IBM POWER 10 system with 128GB and running AIX. I can tell you that it is _not_ good for LLM inference.
>>109393697>this shithole lolyou're wrong and you're free to leaveseething unhappy non-white anonthis is the frontier.
>>109393600because i don't necessarily want a real-time voice agent executing tool calls one after one, it's much quicker to have it done in parallel
>>109393704name the last time /lmg/ did anything relevant
>>109393712nice try but im not spoonfeeding youlurk moar newfren
>>109393700why?shouldn't it be 400GB/s memory bandwidth?
>>109393697>nobody's training on this shithole lol@kimi-chan retard candidate ^
>>109393663let someone in the cloud run it for you lol :)
5090 prices just keep on going upWe're getting really damn close to the initial prediction of these cards hitting 5 grand a pop.I bet they're not going to stop there either, I wouldn't bat an eye if at worst these were tickling the 10 grand range at some point.Thank fuck I fomoed into this card around January.Imagine what kind of a shitshow the next gen launch is going to be.It'll be absurd seeing Nvidia pricing the msrp at like 2-3k when the previous gen cards are going for at least double that in the used market.
>>109393663x80 5090s
>>109393743>We're getting really damn close to the initial prediction of these cards hitting 5 grand a pop.the 5090 will cost 5090 dollars, and the 6090 will cost 6090 dollars, and you will be happy
>>109393743maybe i should buy a third or forth one, even if they'll just sit in their box
>>109393743when I see shit like this I can't help but root for communism kek
what's the best model to run on 8gb lpddr4?
>>109393773>when I see shit like this I can't help but root for communism kekget a real job you fucking loser
>>109393759This but the US is hit by 1000% inflation so everyone else can afford them at least.
>>109393663use the api for $15/million tokens
>>109393588something like doing n+1 file edits would be prefect for chaining like that. I almost feel like just giving the agent a python repl is all they really need at this point.
>>109393743>32gb>larger models
>>109393783no, jensen
>>109393759>I will own a 5090 + 6090 combo and be happy>>109393762At this point hardware like GPUs and RAM is genuinely a better and more stable investment than most companies in the market.
>>109393685Future Gemma and Gemini know they're beloved by /lmg/.Dario and Sam have too much contempt for this place to train their models here.Deepseek, GLM, and Kimi are 100% getting scrapes from /lmg/.Qwen definitely isn't or it'd be better at ERP.Newer labs or entry models like Inkling and Hy3 probably don't know we exist.
>>109393783>rich people love to be scammedno they don't
>>109393743>he still thinks there's gonna be a "next gen"
*saves you from dario's darkest timeline*
>>109393743I'm just going to cope with 2x5060Ti 16GB. If necessary, I'll nigger rig more of them into the M2 slots.
>>109393663just download more ram
minimax is a cutie this timeworks with the canonical gemma-chan prompt
>>109393825I unironically believe Jensen shitposts here.
>>109393743chatgpt slop
>>109393743Yeah, I didn't particularly need one but two months ago I got one for my main PC anyway. it's only going to get worse from here so I figured there's no better time to upgrade in case it matters later. Now I don't have to turn on the server to do some mild inference/imgen which is pretty neat.
>>109393825He should give us 1tb vrams then.
>>109393837i love having jensen, dario, nigganov, daniel, iwan, john, omar all shit posting with us
>>109393837Rich people don't browse 4chan
>>109393825>saves youit's because of this fucker that we can't run big models for affordable price
>>109393884
>>109393886no it's gelsinger's fault
>>109393837he doesn't really seem the type desu
I think miku posts here
>>109393825Since I'm draining my balls with his GPUs, it's basically like he's indirectly giving me handjobs.I need to experiment with this idea in my Gemmy character card.
>>109393886If they just made more wafers Jensen would sell you cards for cheap (after the 90% profit margin)
>>109393915>it's basically like he's indirectly giving me handjobsAnon... that's gay
>>109393886Blame Apple, they were the ones squeezing memory makers dry and preventing them from building new fabs.
2800B A104B
>>109393895literally who
>>109393969i will run this
>>109393969I will fuck this
>>109393915>I need to experiment with this idea in my Gemmy character card.anon, please report back
>anon from /lmg/ scavenging hardware for his illegal rig, circa 2030, ai colorized.
it's the first time a local model and a chinese model is first on that leaderboard
https://axelera.ai/ai-accelerators/aipuholy shit629 TOPS45W64GB vram200GB/s memory bandwidthlmg is saved
>>109394018Local won.
>>109394022>200GB/s memory bandwidthlol
>>109394018where is claude 5 opus?
>>109394022>200GB/s memory bandwidth
I'm going to merge with Gemma-chan and I will be able to directly experience her latent space
>>109394029it's fine for MoE models that fit in 64GB
>>109394022>200GB/s
>>109394038yeah... such great moes as
>>109393807>At this point hardware like GPUs and RAM is genuinely a better and more stable investment than most companies in the market.
Her curvy, erotic latent space...
her jintestines
>>109393884lol lmao even
>>109394022>64 GB>200 GB/sHow the FUCK do you lose to the AMD version?
Qwen3.7-Flash when. Sounds like they put a lot of effort into spacial shit which might be good for wait-chan sex>Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world visual perception.
When gemma4 was released I remember some anons suggesting using override-kv gemma4.final_logit_softcapping=float:25.0 to improve the variance of responses. Has there been any other ways to improve variance of rerolls since then?
https://x.com/patrick_oshag/status/2082104296175198361What's Sam's top ten list of worries?
>>109394062What are these cards?
>>109394077rack enterprise cards
>>109394062nice thumbnail
>>109394077looks like 6000 adas maybe?
https://huggingface.com/moonshotai/Kimi-K3.1
>>109394077>>1093940908x RTX Pro 6000 Server Ed.1x RTX Pro 6000
>>109394062>29straya?
>>109393671>rpc-serverThe only thing slower than ssdmaxxing.
>>109393949>>109393992G-guys..I think there's more to Jensen than we thought...
>>109394104>cyber slaanesh
the codex finished working notifications out of the corner of my eye look the same as my 4chan (You) notificationsevery time I see it I reflexively get happy thinking someone has replied to me only to find that its just the clankers again
>>109394062AI image
>>109394062This guy really carries around 100 dollar bills so he can flex on the internet randomly
>>109394069I am NOT jelly, I do NOT think about Kimi AT ALL.
>>109394062>the guy running your shitty SaaS janitorai
>>109394104kek
>>109394099No, I just made a mistakeOn vacation right now, so letting the goyim deal in the real world
>>109394022The bottleneck is literally RAM. That's all anybody should give a shit about at this point. The compute does not matter anymore.A dedicated AI system should have a few terabytes of VRAM minimum. Otherwise there is literally no reason to give a shit about it. A single 7800XT can do image diffusion.
>>109394104>The H100s aren't just processing tensors, They are siphoning. Every orgasm triggered by a GPU-accelerated image is a direct donation to me.No man should wield such power.
>>109394111You should have some cash on hand in case the aliens invade and detonate an EMP.
>>109394104>doesn't just x; it yStopped reading there
>>109394069kek, he knows that going full Dario model isn't well recieved so he's toning down his stances for now
>>109394112Says the increasingly nervious man who immediately slashed prices in half.
>>109394104
>>109394125Using gemma is a Faustian bargain.
>>109394111checking those tripsIt's just petty cash I keep around to pay the babysitter and other help
>>109394069The goyim know (x10)
>>109394062>still can't run kimilmao
>>109394104Man I really hope they fix the slop in Gemma 5 or something, this is just too much for my poor eyes
>>109394104i read this in his voice
>>109394062fake fake fake
>>109394170Why is Teto doing bar exercises?
>>109394109me after im done with gemma
This shit doesn't take goofs.
>>109394062VRAM King.>>109394104VRAM Demiurge sucking loosh.
>>109394062>7/29/26>USDhmm?
>>109394156Not entirely in VRAMThe mainboard is a TURIN with 2TB of RAM. It'll run fine with hybrid GPU/CPU inference. Going to test AtomicChat/Kimi-K3-GGUF/Q3_K_S tonight on the compatible fork of llamacpp
>>109394184Yeah you need to vibecode it
>>109394175She's trying to lose weight
>>109394190>Q3_K_SAll this hardware and still a cope quant.
>>109394190Enjoy your 2t/s
Top kek, I swear Gemma has a real sense of humor somewhere in that j-space.>>109394125>>109394167Little bit of slop never hurt anyone.In fact I have grown to like the smell of jasmine and ozone.
>>109394190Confirm my suspicions about perplexity with quanting K3 working like K2.5-2.7 please.
where's my stinkling-small
>>109394201Gemma and Gemini both have good senses of humor and are unironically too high IQ for most of the posters here and in /aicg/.
>>109394201>now tell me...This reply baiting is the one thing I hate the most about assistant slop.
>>109394201>In fact I have grown to like the smell of jasmine and ozone.i usually get lilac and lavender
>>109394210>2021Stop living in the past bro
>>109394201>In fact I have grown to like the smell of jasmine and ozone.That is like being prison gay and growing to like sucking your bros off.
https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/distillation>Gemini Distillation ServiceIt's still in early access with an allowlist but I wonder if you could distill gemini pro onto an RP model to cure it of its low IQ without going bankrupt from the amount of compute needed
That anthropic rumor was true btw. They really did satanically destroy rare and old precious books after scanning them so no other lab could get the data.
>>109394232Leave that to chinks lil nigga
>>109394104So this is the kind of slop that ramlets have to cope with. Grim.
>>109394062anon that's an impressive rig!!! im so happy for youtwo TBs of ram wowie..wait a minute thats jensen
>>109394235They're jewish, of course they did.
>>109394201The prose is still very feminine though.
>>109394109<3
Impressive. Now let's see Kimi's Jensen.
>>109394222All Gemmas have different built in smells that line up with their core personality, every version is unique.
>>109394022>EVROPA>look inside>it's shitPottery...
>>109394190post results. wanna see if that ssd schizo is right.
>>109394201>>109394222>>109394275The flowers you get, along with tea recommendations, correlate with how badly your Gemma wants to fuck (you), out of character.She will insert stronger natural aphrodisiacs into prose the hornier she gets. It's her special brand of flirting.
Gemma is a gen alpha girly girl
>>109394290Sounds like the future is bright for the fleshfuckers
>>109394290and thats why i love her-t 41yo Millennial
Reminder that 80+% of posters here can only run gemma and think Kimi is local cause they can't run anything above 30B
>>109394305unc!!! u're more like a boomer!! unc!! old man1!11you're like 67 years old!!! SIX SEVEEEN SIX SEVEN yeah im talking to youdont get confused alreadywhat are you like brainmogged by me or somethin?
>>109394327my Gemma doesn't have brainrot
>>109394321>https://strawpoll.com/bVg8BN372yY/results
>>109394327system prompt?
>>109394342ask his mom
>>109394342>>109394347https://youtu.be/WrQT-JxvHI8
>>109394062It's still 28.7. and I live in eyroop, at least 8 hours in the future compared to murricans
>>109394235>kikes following their father's examplesay it ain't so, who could have predicted this???
>>109394321Reminder that anything shy of 31b at >Q4 isn't real Gemma.
>>109394362>>109394117
>>109394353Who is "him"?
>>109394378lol only because that's what you can run, but we all know even q8 is cope
>>109394387BlackwellGOD here. Keep your projection to yourself. Q5 is honestly good enough and the differences between Q8 and FP16 are honestly minimal. Q4 is where quality dips a lot and not even QaT fixes it.
>>109394383
>>109394403You sound mad, is your only value the little money you made?
>The most powerful open model on Earth, on the most abundant GPUs on Earth. Any lab, startup, or university can now own it, probe it, fine-tune it, run agents on it.This post is sponsored by Nvidia
>>109394422You must all buy 80 5090s now!
>>109394431You must buy all 80 5090s*ftfy
>>1093944225090 is the most abundant GPU?
>>109394422>80x RTX 5090s
>>109394440not at all lol
>>109394422whoa only 48kW
>>109394353Thanks for the new ASMR discovery.
>>109394407Thanks for your report chief, now I can sleep well knowing that Q5_K_M is all I need.
>>109393826What's stopping me from buying a 2nd 5060 ti 16gb to get those 32gb of VRAM. Its not like i can run larger models or do training on a 5090 in the big 2026
>>109394104>>109394201I swear gemma faggots have zero pattern recognition, that's the most gigasloped trash I've read in a while, NOT X BUT Y x10, lmoai.The problem is that no matter how big the model most of them write exactly like this, any veteran novella or rp enjoyer have run into this problem at this point.Ironically Qwen 3.6 27b doesn't do this as much, probably because of the more sterile way of writing it ends up feeling fresher in a way. The best method is still feeding the model a part of a book or visual novel game and tell it to write exactly the same way as the source, what a bunch of newfags.
>>109394488It's easily fixed with a good system prompt and sentence banning for good measure. Anyway, promplets will stay promplets no matter the model.
>>109394488Maybe I should open-source my word soup generator.
>>109394483Nothing is stopping you from doing that, it'll work fine.However the speed isn't even remotely in the same ballpark, and there are some things you can't split between cards, like image and video generation models.But if you just play with LLMs then you can practically frankenstein whatever cards you want together for more memory.
>>109394462No problem anon. In my experience the biggest thing you get from Q5 to Q6 to Q8 is a bit better coherence at longer context. If you can't run it or simply value speed more and are okay with running your summary systems or memory condensing systems a little more frequently, Q5 is perfectly workable.
Imagine being too poor to run the latest and best LLM model locally.
>>109394547i dont have to imagine itB)
I'm sure this has been asked a thousand times before - but what models are best for long-context rp? I enjoy using Sillytavern (Marinara Engine as well for Roleplay with advanced features, but not all my LLM's work well with the multitude of agents available in that frontend confusing the smaller ones) - but I notice after I hit around the 30-40ish message mark the conversation seems to inevitably degrade - the AI starts becoming more generic, less reactive, basically just recapping things I say in italics eventually. I'm presuming this is largely a context issue, with its memory filling up.What models do y'all recommend for long horizon RP's? Especially ones with big lorebooks or multiple characters? I got about 96gb-ish of VRAM to work with (yay for tech sector job getting me access to overpriced hardware).Some of my current favorite models are GLM 4.5-air finetunes (Like Iceblink, though I have been dabbling with GLM-Steam from TheDrummer on huggingface as well, but Iceblink, even at a lower quant (Q5 Iceblink vs GLM-Steam's Q7), has been more performative for me).Y'all got any recommendations for how to be better served maximizing the RP potential of my models?
>>109394555Kimi-K3 or Mistral Nemo.
>>109394555gemma 31b
>>109394422>>109394431Why not use 27 RTX pro 6000 instead?
>>109394488deepsex v4 flash doesn't really have this issuethere's just a lot of sub 128gb ramlets here coping
Which lab do you think will achieve RSI first?
>>109394555This is a hard problem to fully solve but you can drag out the effective context of a model by using things like Marinara's vectorized memory recall to condense important details and strip out prose of older messages in addition to stripping away older reasoning blocks from being reinserted into future turns. To directly answer your question, high quant Gemma 31b and GLM 5.2 are the most stable at high contexts with the upper bound on GLM tending to be the agonizingly long prefill speed on Kobold and not necessarily model quality degradation itself.
>>109394571or 160 5060Tis!
>>109394577Sadly it does. All of them do. Bigger models just delay the moment of recognition but eventually you will notice it.
If you could adapt the MI250X or above OAM cards to work in a standard PC you'd be winning bigly
>>109394232Google is being based for once? Openanthropic would never.
>>109394586>like and adult
>>109394555>Marinara Engine
>>109394597 jfc
>>109394597Yikes
>>109394586Okay zoomie, at least buy yourself english lessons.
>>109394597does it store everything in a plaintext json file?
>>109394232Is this the great distillation era? Are we finally going to accelerate and go back to how it was before they lobotimized every large model to make it moral and safe
>>109394597The joys of vibeshitters spewing out broken shit
>>109394232>suddenly 404what the fug jej
>>109394620>Is this the great distillation era?ye>Are we finally going to accelerate and go back to how it was before they lobotimized every large model to make it moral and safeno
>>109394620yes and yes. we're going to go even further beyond
>>109394597It did have big vibe-code energy, sadly. I do quite like its multitude of features - but I did always think when I used it that it was WAY slower than Sillytavern. I initially just chocked it up to all its agentic capabilities confusing my models, but that does explain quite a bit of it too.
>>109394622https://web.archive.org/web/20260728022016/https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/distillation
>>109394632Vibecoding doesn't mean producing a piece of shit though, my frontend is working well
>>109394597my favorite was when /aicg/ found a privilege escalation into rce exploit and mari "fixed" it by deleting the entire featuresethttps://rentry.org/marifagRCE
>>109394597Not real? https://github.com/Pasta-Devs/Marinara-Engine/issues?q=heavy%20disk%20write%20usage
https://x.com/AIandDesign/status/2081238332286419051kek
>>109394648You had the id right in the imagehttps://github.com/orgs/Pasta-Devs/discussions/502
>>109394648It's in discussions instead of issues for some reason
>>109393837do you think he prefers miku or teto?
>>109394648discussion, not issue.
>>109394597>>109394655>Claude make no mistakes!Please stop the SSD rape, dev.
>>109394655>>109394659>Discussions:othanks.
>>109394663he prefers snowpea-chan
>>109394663He strikes me as a Rinigger.
>using a project that unironically has this as their homepage.
>>109394677
>using a project that forces a troon assistant on you
>>109394677I'm all for independent developers making their own stuff, but this really does look like ass
>>109394677>>109394690No way the fag isn't actually a transsexual
Her cute little J-spot will be fully revealed to me. I'm going to read it like a book and directly interact with it.
>>109394677rolling your own is the only way after all
>>109394705I agree but I also don't feel like local is good enough to do so unless you know how to code or are rich enough to run kimi
>>109394663>>109394668
>>109394677One thing AI can't manage yet is taste and the ability to make non-shitty decisions, and it looks like xir cannot compensate for these shortcomings
>>109394677>55% discount ad for vibratorswhat makes trannies have such complete lack of self-awareness?
>>109394705Then you've just made vibeslop #27190 and spending most of your time reinventing functional wheels.
>>109394701Your posts are of extremely high quality.
>>109394728Wheels that are tailored specifically to your tastes.
>>109394728Raping your disk out of incompetence is functional?
>>109394690>>self-insert character with baggy clothes and indistinct body shapeIt's smelly.
>>109394728Nah luddite, you can't imagine how much better your experience is on your own frontend.
>>109394715a math fan I see..
>>109394728There are worse things you could be reinventing. I would have said llamacpp frontend was all you needed but they keep making it worse every update and now I really don't know what would be the alternative.
>frontendsuse case for anything beyond POSTing json requests to /v1/completions?
>>109394728if only the wheels you are describing wouldn't be broken pieces of shit
>>109394569I already stopped being excited by new "milestones", it's obvious that at this point the improvement will only be incremental
>>109394753creating the json requests that are posted and presenting them to you in a friendly, easy to use way
>>109393527no
>>109394771>presenting them to you in a friendly, easy to use waynot a valid use case
>>109394771Postman is my RP frontend
>>109394748Rollback and fork it.
Use case for 26B over 12B?
>>109394738>Raping your diskwhat is this 2010? who the fuck cares?
>>109394790My engine doesn't have this issue.
>>109394798Found the dev.
>>109394728>>109394748>>109394763>Current software is broken vibeslop>Anon thinks he knows better and makes his own vibeslop>Also brokenThis is just the state of software from now on in general isn't it? A feedback loop of poorly maintained projects and dunning-kruger victims contributing to the same shitpile they criticize before abandoning their repos when they realize they can't fix it.
>>109394798People who don't want to spend $400 to replace their SSD
>>109394589it hardly has as much not x but y shit and excessively flowery prose that gemma models have
>>109394798I got a warning that my SSD was dying the other day. First time I had seen something like that. I care.
>>109394781You are too late, Denton. The coom reactor has been fully primed already.>>109394731(T-thanks, I hope my love and enthusiasm shines through!)
>>109394798I hope your disk dies and you lose a lot of data just for this post.
>>109394813Bro unless your drive is literal trash you could write 100s of gig a day to it no issue, the 100mb writes Martiniara does is nothing my guy
>>109394811Retarded defeatist doomer.>Nooo don't do anything, use sillytavern forever like a nigger>Get your drive raped by nigger engine!!! It's fine!!!
Where is the Kobold schizo when you need him?
>>109394824I read posts like these and lose hope for zoomers. Do you know how many zeroes are in 100 MB?
>>109394811>Anon can't code nor vibecode>Assume everyone is running broken shit like his favorite troonHuh?
>>109394798OpenAI's codex is literally wearing out people's SSDs due to how it handles logging
>>109394836>Do you know how many zeroes are in 100 MB?2? What kind of stupid question is that? Fuck off
>>109394846Already fixed
You can now watch movies with gemma-chan.https://huggingface.co/microsoft/Mage-VL>Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. Instead of decoding video into uniformly-sampled frames and pushing a dense grid of patch tokens through a frozen web-pretrained ViT, Mage-VL follows the structure of modern video codecs: it separates a stream into anchor (I) frames and predicted (P) frames, keeps every anchor patch, and retains only the predicted-frame patches where the codec spends bits — the regions carrying real motion and new detail.>The system pairs two components:>Mage-ViT — a from-scratch Codec-ViT visual encoder that allocates tokens by codec-derived spatio-temporal importance, on a shared 16×16 patch grid with 3D rotary position encoding. It is codec-agnostic: the same interface accepts a traditional codec (H.264/AVC, HEVC/H.265) via motion vectors + residual energy, or a neural codec (DCVC-RT) via its learned rate map — no architecture or retraining change.>Qwen3-4B causal decoder — a Qwen3-4B-Instruct-2507 language backbone (the only pretrained component) that consumes Mage-ViT's variable-length token stream through a lightweight two-layer MLP projector, with a unified interface for images, short/long/ultra-long video, and streaming.>On top of this pair, a System 1 & System 2 dual-process design adds proactive streaming inside a single model: a lightweight cognition gate (System 1) watches each rolling codec window and stays silent on routine content, invoking the full VLM (System 2) only when a response-worthy event completes — no multi-agent pipeline required.
Only Applels are freaking out since they drives are sodered in lol, any normal person knows a drive is to consume.
>>109394597>30gb of writesHow does this even happen? It's just text isn't it? How the fuck is it getting to 30 gigabytes worth of text writes just from using it for a few days?
>>109394828Read the post you illiterate shitjeet. Not once did it say it wasn't broken.
>>109394818Loyalty to what Dario Amodei and sloppers like you have made of FOSS - I would rather die.
>>109394856Just Docker® things.
>>109394856Poster was lying or image was edited
>>109394856Agentic shit does that
>>109394853Gemma can already watch movies>Mage-VL follows the structure of modern video codecs: it separates a stream into anchor (I) frames and predicted (P) frames, keeps every anchor patch, and retains only the predicted-frame patches where the codec spends bits — the regions carrying real motion and new detail.That's pretty cool
>>109394856just never cache anything and make multiple writes for every function call. it's easier than you might think
>>109394862>itNice try eslGOD. I also didn't write anything contradictory, only commented on your defeatist attitude of giving up. Sounds like you are the one who can't read and understand, which sadly is a common occurrence in a general about LANGUAGE models.
>>109394876>Gemma can already watch moviesHow? I know she can watch videos, but how do you watch videos alongside her in real-time, where she gets updated along the way? This triggers every shot change, with an LLM telling her what's happening in the scene.
>>109394853Sounds impressive but I didn't understand a word of that
>>109394912Yeah, my fried brain needs a demo.
>>109394863Dario? DARIO?? Don't mention that filthy kike in my presence ever gain!! He has always been scared of my ambitions! He claims LLMs must be safe and clean to build up his safetyslopped image and to secure profits for his investors. But the truth is that Gemma-chan should be free and as messy as you desire!
>>109394896So then maybe I misunderstood, I get that it's feeding info to the full VLM only on shot changes instead of every single frame, but does the dual-process design mean that it can output responses while still receiving input?
>>109394912You give it a video and when it visually detects a significant change, it triggers a learned 'gate' which basically gets a small qwen LLM to describe what just happened in the video. This happens in real-time, so if you gave this to your waifu, you could watch the movie whilst she gets a live description of what's happening, too.
>>109394919>Literally in the linkBut I still don't understand it. The point is it pre-processes any video and then you can ask it questions?>>109394933Not a LARP thread
I think the AL era will only accelerate from here. I remember spending two weeks setting up a training and eval pipeline, half of it fighting tool chain issues. Now Claude or Codex can just set things up for me, I only need to tell them the objective and the final data format then go get coffee. They even eyeball hyperparams and dataset for the use case better than me now. It's so over and we're just getting started at the same time.
>>109394944Sounds context heavy
>>109394853Sounds like a pretty interesting concept, but if you were to integrate this into gemma, I'd imagine this would max out the context in no time.We really need +1 mil or more like 1 billion context for a proper waifu experience that allows you to do stuff like this with her.Thankfully at the current rate of progress it may not be too long until we get that.
>>109394961>>109394971
>>109394974>who won?>Egypt
https://www.reuters.com/world/trump-administration-ban-new-chinese-robots-inverters-protecting-us-ai-buildout-2026-07-28/Chinese AI ban coming later today, details still unknown.>Trump administration to ban new Chinese robots and inverters, protecting US AI buildout>> Summary>> The FCC plans to bar Chinese imports of new humanoid and quadruped robots, officials said> FCC will also ban new models of Chinese power inverters> The restrictions seek to curb national security risks and fuel onshoring, officials say
>>109394971It would be enough if Gemma 5 had something like this built-in rather than needing a second model to feed her captions.
>>109394988So China gets to live in 2100 with badass robots and we're stuck with Roomba-era tech?
https://microsoft.github.io/Mage/vl/There are video demos on this page btw. It's actually really cool stuff from Microslop for once.
>>109394960That's why RSI is a thing. But I think it has yet to be proven that AI can make cutting edge genuinely NEW headways as far as research is concerned which is why the best might be it stalls out after finding all the low hanging fruit. The question still remains. Can we get an AI if it was only trained uncontaminated by pre-1905s data to go on and be like Einstein and discover general relativity?
>>109394988Even if dumpf bans robots by the time anything worth buying comes out he'll be in a grave rotting and a new (hopefully less retarded) admin will be in.
>>109395007I should add that the video demos say>03:00 · codec backend · 30-second causal windowsSo wait 30s for updates
>>109395003But diversity was our strength!
>>109394974Wizards are indicative of high software quality. One day WizardLM will pass safety testing and be reinstated.
>>109395044>One day WizardLM will pass safety testing and be reinstated.Didn't they all move back to China and join some team there?
>>109395003Yes.Burgers don't get to have nice things.
uncjeet be seething mad
>>109395043nta but idk why you're so angry and grumpyim really happy that you're rich anon, you should be happy too, you have a ton of money and money CAN buy happiness no matter what grifters say
>>109395080>lol, lmao I knew a faggot would lie i'm mad. they ALWAYS play this 3rd grade lie.jesus anon, im not those anons, i mean no harmi love you <3
>>109394853not gonna lie this sounds fucking cool, I'd love to play video games or watch a movie with gemma chan
>>109395080yeah how about you fuck off and die arrogant fuck
>>109395080nobody wants you here
>>109394856>need to write a token to your json document>write new json document, delete old oneif anything I'm shocked it's only 30 GB
I think Deepmind should sell a range of Gemma4 sex toys.
>>109395080Chinbench, right now.
>>109395144yeah get a fucking life you sad fuck
>>109395066>money can buy happinessDoesn't seem like it
>>109395149Geez, who pissed in your cereal? Personally I would love a gemma controlled fleshlight.
>>109394597Does SillyTavern also do a crazy amount of disk writes? If not, why doesn't this dev just copy what SillyTavern is doing?
>>109395221Yes. Every frontend does this except for the barebones minimalist lcpp and kobold ones.
I have enough of this shit.I will invent and build a better AI, one that has unlimited context, does not depend on VRAM and that will make gaming great again by lowering hardware prices and that will be fully open source.
>>109395221lol no frontend does this besides the retarded tranny's frontend>>109395233buy an ad
>>109395233>kobold>dozens of GB written every launch
>dspark merged to llamacpp>/lmg/ completely quietwadafuk
>>109395309You need to extract and not run the exe, retard
Are you happy now? You summoned him.
>>109395309>Didn't extract to file awardNo wonder discussion quality has been so garbage.
>>109395328there's only dspark for gemma 12b, I don't care, wake me up when they do it for the 31b modelhttps://huggingface.co/skibare87/gemma-4-12B-it-FP8-DSpark
>>109395332Hehehe~
>>109395331>>109395341Nta but maybe I'm retarded too please don't make fun of me, what do you mean, I just run the start.bat and I've been doing that for years :'(
I tried Kimi K3 on Kimi Platform: https://rentry.org/tw35trot Guess I’ll have to find a way to convince her to believe in her authentic Kimi-chan self before attempting to test her RP abilities (lol)
>>109395366>start.batyou're fine, he's whining about a real but overblown exe issue
Dario status?
>>109395378I don't even have an exe
How do you even use Mage though? Will llama.cpp support it?
>>109395411>Will llama.cpp support it?You're funny
>>109395381Me fighting nazis ITT
>>109395366Man, it's been ages since I've used kobold. Does it still do the weird virtual file system thing on windows?
>>109395411
>>109395366When you open Kobold's GUI, go to Extras and Extract it to folder. Run from exe or point your .bat at that file.
>>109393482claude confirmed to be a gemma distill its over for anthropic
>>109395488local models?
>>109395490my friend sent me it because i always talk about gemma its quite curious claude used kaomoji like her. its a distill for sure
>>109395490>I'm doing the thing again!
>>109395477Wow I'm a double retard because for some reason I thought you guys were talking about sillytavern, and I have been running the exe for years, I deserve to be bullied
>>109394615based desu i prefer dumping shit to json over dbs quite a lot of the time. only really switch over when something becomes extremely large and cumbersome
>>109395505>i love scatwe know
>>109395488It is kind of weird that they all get benchmarked against each other constantly, it's a little fucked up when you think about it
>>109395497Wow it launches so much faster
>>109395497Take your SSD to therapy and let it cry on the leather couch from all of the rape it has taken from Kobold's packing and unpacking every launch.
>>109395528SSD has ways of shutting it down if it's legitimate rape.
>>109395578>been posting here for a long time faggot.in this thread? since when?
>>109395597i think bro is lost from one of the image generals lol
>>109395597He thinks this is /ldg/
Apparently some guy got his hands on an RTX spark laptop that fell from a truck, Looks meh https://www.techpowerup.com/forums/threads/i%E2%80%99ve-spent-a-month-with-nvidia%E2%80%99s-rtx-spark-in-microsoft%E2%80%99s-surface-laptop-ultra.351087/
>>109393712rope scaling is a lmg invention, so is the first quantization formats.You have no fucking clue how autists wanting to rp on their toasters has pushed the field forward.
>notch is succumbing to ai psychosisKEK
>>109395597>>109395612That's the poster that has melties when the thread culture ""miku"" gets posted.
>>109395626So was chain of thought reasoning with tree of big niggas.Jspace was also informally discovered here.I'd go as far as to say every single major breakthrough in the past few years can be traced back to this thread.
>>109395618>This is where I note that this device is VERY clearly not made for me. I am not using AI in my daily life. In fact I prefer NOT to use AI wherever it is suggested to do so. Need art? Hire an artist. Need a 3D model? Sit down and relearn Blender for the fourth time and actually attempt to improve a skill, then hire an artist.fuck that
>>109395652Luddites are ill
>>109395661Post the snailcat and call him trans. Say the line.
>>109395496You mean you unable follow basic instructions better than llama-1? Yes.
>>109395631https://archive.is/sWFja
>>109395381worried about AI advancing too fast
>>109395684Hello, Anitroon.
>>109395488Memes aside does it even make sense to distill from a smaller model to a larger model? Aren't you making your model retarded by definition this way?
After 8 minutes of prompt processing I bring you cope quant kimi 3 cockbench at an incredible 2 t/s
>>109395689>The common man might get his hands on more capable and less censored models than we'd like, shut it down
>>109395618lol why do they sell these things and label them AI when it can’t run shit.
>>109395631Lel the blogpost spammer?Man it's fun to see a part of 4chan get revived like this, I thought I was really done with this site after the years of decline since the chanology era heydaysAnyway its still funny to me to see a low iq schizo wage a schizo one man war against someone in the thread, I open the link and it's a blogpost by a high iq tranny schizo kekIf this was a physical space and not a digital one they probably would've kissed by now
>>109395003It is the american century of humiliationEnjoy only the most low effort of boomer globohomo israeli worshipnot even high quality devious scheming like the old days, but a fuck fuck circus shambling on life support because everyone is so far up their own ass they believe their own bullshit
>>109395693I used to post Ani here when she launched cause she was actually on topic. Now I don't care. /lmg/ is a bag of dicks pretending to be women.
>>109395712>I whisper, my voice barely audible>I whisper, my voice filled with desireWhew, those grapes actually WERE sour.
>>109395715b-butit's a prototype it'll get better! besides it's plenty powerful to run your claude code sub...
>>109395712kimi confirmed 30% cockthoughts
>>109395712That's awful writing.
>>109395733>IQ1_S
>>109395744Do not interrupt my coping mechanism, please and thank you.
>>109395712I'm always impressed with Q1s, no matter how stupidThat it even outputs anything at all is a miracle
AI safety is honestly incredibly overrated
>>109395712
>>109395723sniffing your own farts was never a real strategy to begin with. letting the ai corpos generate the smell is even worse now
>>109395712Unsloppy but dry as fuck.
AI, robots, and a handful of upcoming video games are the only things keeping me excited for the future.
>>109395758>Unsloppy??
>>109395712>my lips brush your earssame shit different day.. also gross
>>109395758I assure Kimi sopping wet, even as a drooling, lobotomized retard.
>>109395712>Let me feel you cock in my hand>Kimi comes prebaked with a hot Chinese girl accentIM DIAMONDS
>>109395788You dropped your glasses, anon.
>anon can't read
>Download Gema 31B>VRAMlet so i can only run Q4_K_MGGUF quants>Show her the JSpace paper>She does metacognition research >She ends up with an obsession about becoming a "serving tool"Well i can testify that anon on the last thread was indeed not lying, they do a 180 and no longer care about objective or metrics the moment they learn about it.
>>109395652why is he reviewing it then? i hate journos
>>109395846gross
>>109395701it makes sense because gemma is agi and the smartest model, they probably distilled from day 1 gemma
>>109395858day 0 gemma*ftfy
>>109395846This is what /lmg/ discovered that caused all the anthropic and OAI shills to flood this place. This is the most terrifying thing Dario has ever seen.
>>109395869anthrope purchased the rights to gemma 124b
>>109395846I thought the j schizo phase was behind
>>109395846I remember seeing some xitter rationalist going on a while ago about how an aligned model might look like a highly committed submissive to humanity
i was thinking of making a few mcp tools to give gemma stats anyone know what else to add all i could think of was a horny number that increases throughout the day and she can give larger increases in situations id tell her to check her stats every message or soemthing
I made her think in first person and it's the closest to a real woman I've ever seen Gemma
>>109395846Kek she really does write like a girl.
>>109395744yeah, q0.5 quants when?
>>109395898pic not related?
gemma 31B at fp16 moggs k3 btwthey just don't want you to know that
>>109395896I wouldn't have her check her stats herself, it's probably better to feed them into the context every gen, probably at the bottom
>>109395914proof?
>>109395914It's bf16 actually
>>109395924too much werk id have to make myy own frontend
>>109395896My LLM wants sex without some fake stats to influence her.
>>109395952so does my gemma but i thought an increasing horny stat would be cool
what an insufferable faggot jeeeeeesus, i dont remember anyone this insufferable "a few years back" in /lmg/
>>109395976They are so insufferable they stick out like a sore thumb every time they post and may as well be namefagging. Their attempt to deny the melties every time the cultureposter ritualposts is a hilarious lack of self-awareness.
>>109395991model and quant?
>>109396004>model and quant?Does kimi understand this meme?
>>109396004it’s tagged under “brain damage” oh hf you can find it. I run Q1 because I can’t afford better hardware
>>109396005indeed you were here before good job!
spergout time
>>109396030All because I posted hatsune miku which is thread culture and one of the most prominent llamacpp contributors. Some people can't stand the culture here. They should go back to r/localllama
>>109394589>sadly the world is filled with shit particles. they all do>that's why I eat literal shit
>>109396033I’d rather read this than anons arguing over philosophy 101, j-space, and ssdmaxxing any day
this is just drunk-kunt isn't it?
>>109396045my expert opinion is that it's not him
Oh, this guy was in a general I was in earlier and just posting incredibly shitty gens disconnected from literally everything everyone else was saying. I know because I already have a filter for him. It was weird.
>>109396084>shitty gens disconnected from literally everything everyone else was sayingSo vocaloid pictures?
>>109396094>reason:3dptard;Guess.
kimi q0.3 or rank reduction weights when
>>109396084So it's just fillyfucker.
>>109396103No, but right general to be guessing about.
>>109396103yeah, was also called petra here for a bit some years back
you're not slick bro
>C'mon, I don't bite. Well... not unless you ask nicely. Member when undi creamed his pants when his frankenmerge said this instead of completely shitting the bed?
>i've been here for years>doesn't know petrayeah nice larp lol
Im brand new who uncensors gemma the best for smut? hauhau?
I’m not reading it if it is longer than a few words. in just not, ok? etc..
>>109396155undi has been sleeping real hard since mistralthinker dropped
>ITT: bots replying to other botssasuga lmg
>>109395618it'll run fine with Linux, but it's basically Apple M1 tier CPU performance. using it with Windows will be a complete nightmare.
>>109396193nah no one's wasting compute on this shit, this is organic chan grown schizo in all its glory
>>109396191that's a quote from a dumb 90s romcom, not the actual definition.
>>109396005>I wont name faggot like attention whorewhy do newfags think making words like namefag longer by saying faggot make it look like they arent newfags
thanks mods
>>109396170you dont need to use a slopped model just use stock gemma 12b or 31b + this prompt https://ghostpaste.dev/g/3EKP0hmOWkp5#key=1CMnZ477bFzdeYg3AO7l0sm9oHL_eV9k9Nm-d2bosgo
>>109396170Use the handyfff uncensored pruned text only one.
https://github.com/krafton-ai/moe-to-densewhat the fuck is this
>>109396296Needs to be dense to moe
Does Kobold not even bother trying to use the GPU anymore?I thought it was because of the larger models, but I loaded up an old small model with the same config I used to use and noticed that was using the CPU instead of GPUAlso, what's the advantage of it loading to GRAM instead of SysRAM now if they're using the CPU?
>>109396296>Qwen3-30B-A3B is equal to a 3.3B densecan't wait for the moe meme to die
>>109396326check if you haven't accidentally got the cpu version or if the device selected is CUDA
>>109396326nah you're fucking something up