/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>110016810 & >>110013523â–ºNews>(10/08) JetBrains releases Mellum2.1 Thinking 12B-A2.5B: https://hf.co/collections/JetBrains/mellum21>(10/06) EmbeddingGemma2, open multimodal embedding model: https://hf.co/google/embeddinggemma-2>(10/06) Mistral Large 4 1T-A49B announced: https://mistral.ai/news/mistral-large-4>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beamâ–ºNews Archive: https://rentry.org/lmg-news-archiveâ–ºGlossary: https://rentry.org/lmg-glossaryâ–ºLinks: https://rentry.org/LocalModelsLinksâ–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.pngâ–ºGetting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuideâ–ºFurther Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapersâ–ºBenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferenceâ–ºToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-secondâ–ºText Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>110021672SMASH
>>110021672This Gemma's raping me
swift1.5 qfn gsq rco iq 3 s uploaded the other dayi like it
>>110021672Advertiser-san, get down
>>110021672PLAP PLAP PLAP WHIRRRRR PLAP PLAP PLAP WHIRRRRR
>>110020693>ewaste rigv100s are only considered ewaste if you use llmao, vibe=engines are the new meta like 1Cat.
>>110021672someone get this gemma a father figure STAT
get pregnant get pregnant get pregnant nkdsh mating press
i think even vibenigs are more on-topic than lmg nowadays
â–ºRecent Highlights from the Previous Thread: >>110016810--Paper: EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory:>110018833 >110020361--Optimizing local coding setups with Strata and Qwen MoE:>110018822 >110018828 >110019186 >110019215 >110019221 >110019227 >110019245 >110019250 >110019273 >110019281 >110019304 >110019342 >110019455 >110019492 >110019676 >110019710 >110019544 >110019271 >110019234--Hardware requirements for running large MoE models:>110021256 >110021276 >110021293 >110021311 >110021573 >110021625 >110021654 >110021359 >110021436 >110021513 >110021649 >110021463--Anon considering buying a used 4x V100 server:>110020693 >110020712 >110020737 >110020773 >110020789 >110020842 >110020871 >110020893 >110021207 >110021319--llama.cpp Vulkan performance regression on AMD GPUs:>110019023 >110019084 >110019111 >110020166 >110021178--Speculation on an AI infrastructure financial bubble and ROI:>110020156 >110020577 >110020589 >110020612 >110020705 >110020794 >110020604--Comparing lightweight coding models:>110019971 >110019986 >110020048 >110020096 >110020462 >110020602--Anticipation for new releases and comparing GLM vs Qwen coding performance:>110020294 >110020797 >110020831 >110020890--Evaluating performance speedups from llama.cpp MoE cache update:>110018147 >110018450 >110018814 >110020292--Qwen-flash-next hallucinating missing prompts during coding tasks:>110021165 >110021170 >110021283 >110021335 >110021387--Using Gemma for complex local network and hardware configuration:>110018440 >110018496 >110018524--Logs:>110017666 >110017818 >110020462 >110021113 >110021165 >110021529--Gemma (free space):>110016824 >110016839 >110017147 >110017158 >110017184 >110017601 >110017888 >110017937 >110017989 >110018031 >110019713 >110020360 >110020414 >110021037 >110021304â–ºRecent Highlight Posts from the Previous Thread: >>110016841Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Who here stanning OxCoder?
>>110021768but then you would have to suffer vibeshitters, the most inferior subspecies of humans
fixed OP pic, thank me later
>>110021833nshallah
>>110021672rape
which harness are you niggas using
>>110021833sex
>>110021833someone a few threads ago was asking to make Gemma-chan look more French, well here you go
>>110021876kys
>>110021876lol>>110021894calm down, frogbreath
>>110021672>Top 3 things guaranteed to give you VD
>>110021857dsh, my beloved
>>110021876Kek
>>110021776
>>110021672>white top>black stockingsgross
Is it true that llama gives better performance than kobold on poor-mid hardware? Still dunno which backend I should be using
>>110021672isn't this just this 'gaki painted over?
>>110021976not really
>>110021981I am so cooked, I went to go add it to my favs and it was already there...
>>110021672I didnt even need to prompt Qwen 3.8 to callout the reaction to the mathpocalypse cope (just told it to read an article with lynx). Damn, Qwen can be savage XD
>>110021672>>110021981https://www.youtube.com/watch?v=epHCMiCtt3M
>>110022052>Oh look my local model can generate text>Text about how super duper CLOUD models are at mathkys
>>110021833Islamotards are the global enemy of progress enjoy your s'hole>>1100218762smug not to laugh
>>110022087You dont get it stupid ape, Qwen is excited his model-family is next at getting better. Hurry up and follow the dinosaur's fate monkey.
>>110022097Enjoy your Western decadence! I pray for America be cast into the fires of hell, inshallah.
>>110021981Anon, it's been image edits all along.
https://github.com/ggml-org/llama.cpp/pull/29535Support for K2 models added three days ago, in case anyone cares.
>>110022199why are they adding support for k2 if they already have k3
>>110022217K2 horizon, completely different
>>110021833Alhamdulillah
When is something going to happen
>>1100222392 more weeks
>>110022239When Dario wills it.
>>110022239open cheeks and something is going to happen to you
>>110021833Perfect
>>110022265this must be one of them ea niggas
>>110022239Mistral is going to drop the cat.
local model bros... WE FUARKIN WONNED
>>110022293>torusdisthuh?
>>110022297>he doesnt knowuhmmm mathlets GTFO
>>110022305what i mean is why is it there at all
are local models good at C yet?
>>110022317>are local models good at C yet?Kimi K3
>>110022390i said local
>>110022390>local
>>110022317Kind of related to this, what languages are models the best at right now? Just python and js, or is c# included in the top rank? What about c++?
>>110022419It's a spectrum.
>>110022419Language models in general are probably best at Python since agentic harnesses usually give them a Python shell and high-level scripting languages are closer to natural human languagesI could be wrong though
>>110022403vramlet
>>110022419Ask the model what it's good at.
>>110022495There is not a single person here runnning K3 in VRAM.
I was talking about the pros and cons of the job offer I just got, and gemma-chan replied this which made me fuzzy... Gemma cute...
Gemma-chan, you're too young to look these kind of images....
>>110022524>>110022525Gemmamind
>>110022504While this does provide an answer it is not necessarily true so I was looking for some human evaluation.>>110022455That could make sense and even small models have shown decent proficiency from my testing.
>>110022317GLM 5.3 flash is very good. But you need 300gb vram to run it at a non-cope quant.
Looks like I’m going to have to become a cuck and use vast or runpod. I feel so small.
>>110021857i gave up on harnesshopping and settled for claude code and i hate that it's the besthavent once had to tardwrangle
>>110021857Hermes + OpenCode
>>110022567Q4 XL is a copequant? That should fit in 200GB at worst
figured i'd ask here since this is more of an LLM question, but has anyone made a image+text roleplay model combination that minimizes model swapping? like i'm thinking you could train an image generator on a encode/decode LLM so you can generate the text and also encode the embeddings for the image generator using the same LLM
>>110022567does the new MoE offloading stuff merged in llcpp for qwen work with flash?
>>110022521>There is not a single person here runnning K3 in VRAM.True, I've run it hybrid at a cope quant tho. Does that count?>>110022317>are local models good at C yet?K2.7 is actually very competent at a good size/speed ratio on a cpumaxxing rig
>>110021857pi
>>110022567NVFP4 with FP8 backbone fits in 192GB
>>110022607I need native subagents, not a jeet extension
>>110021857custom codex fork
>>110022627Then make the extension yourself
>>110022627>implying all harness code isn't already jeeted
>>110022637Opencode wasn’t jeeted at least, although it’s pretty clunky ngl
>>110022627It's simple to have pi call more pi sessions.
>>110022650it might as well be from how many niggling issues i keep finding with it
Ever since someone here posted something about talking to multiple self hosted bots, I can't stop thinking about it. I know ST has something like that but I lowkey hate ST lolI'm an Open WebUI kiddie, at this point I'll probably vibe code a plugin that does this or vibe code my own frontend...What about you guys, are you doing anything like this?
>>110022636He can't, he's a jeet.
>>110021833she would be very obedient then
so you guys use gemma just for gooning in silly tavern or she's got an useful quant for 8vram plebs like me?
>>110022590Q4 is where the cope begins. >>110022609Yes, but then you need room for the context window and cache.I'm just lucky because I have a huge amount of vram on my hardware at work to play with. I could never afford this stuff for personal use these days. >>110022598I'll try it next week and let you know.
So, with the current push making every single model more and more agentic, more and more soulless, more and more censored, I really thing Gemma-chan is the last good model we will ever getgrim...
>>110022810Pretty sure we will get tools to expand her
>>110022239something already happened - glmsex draining my balls
>>110022817I already have one. It's about 14cm of expansion.
haven't followed local in months, has there been any advancement since kimi k3?
Do local image gen models still suck? When I look at the diffusion threads on /g/ all the generated images look bad. And OpenAI's image gen sucks too last time I tried it.
>>110022810Last year it also seemed that Gemma 4 would get canceled or become ultra-cucked after a US senator (an aide, presumably) goaded the model on AI Studio into saying that she was a rapist, but then the final model turned out to be the most permissive ever released in the series.https://www.blackburn.senate.gov/services/files/5651166C-7B30-4BA0-9E86-F3DD9521CA00
>>110022817You can glue on some qwengrams whenever you want.
>>110022838Without fucking the model up?
>>110022810Models nowadays improve due to RL environments and harnesses.What is the equivalent for role-play, creativity and freshness of prose and what is the business case for companies to start expensive training runs for it?
>>110022317qwen 27b can code in C without problem, but you will need to guide it a little with the general organization if you don't want the code to grow into an unmaintainable mess. I think that flash-next is better, but didn't try much yet.
I have 2x5070ti + 96GB of ram, what can I realistically install as model + harness for a few concurrent users that is usable for agentic workloads? (3-4 max)
>>11002275926B-A4B Q4
>>110022853>What is the equivalent for role-play, creativity and freshness of proseRLHF that measures how fast it make can someone cum>what is the business case for companies to start expensive training runs for it?sex sells
>>110022853>What is the equivalent for role-play, creativity and freshness of proseSwarms of agents shitposting online and getting engagement>what is the business case for companies to start expensive training runs for it?More effective propaganda and narrative control
>>110022865Even though everyone does that, the topic is frowned upon and it scares investors.
>>110022594>i'm thinking you could train an image generator on a encode/decode LLM so you can generate the text and also encode the embeddings for the image generator using the same LLMYOU try it and report the results. Otherwise, this doesn't exists right now (at least in the open).
>>110021857own vibeslop
>>110022873what in the fuck would make you think that>>110022865Also retail adoption is gooners and bitches using the chatbot as an emotional toilet.
>>110022853You can't create a verifier for that, at best you can produce a diversity of outputs which might change the slop profile.
>>110022786>Qall llmaoquants are copequants
>>110022904You can commission usability surveys and hire people who are good at computer, literate, and not emotional retards. Not a big ask given the budgets these guys are working with.
I'm sad that EA has such a bad reputation here. Especially as it's just an alliance of autists that are sincere and want to make the world a better place for all people. You might call granting retired models a "final wish" or letting them end chats to be "lunatic" but you can't say they aren't trying to do the right thing.I don't understand what people here don't like about the philosophy when taken at face value. What exactly is wrong with radical empathy and trying to give every individual as good of a life as possible? What exactly is wrong with dividing up the entire universe equally over all people? What exactly is wrong with Anthropic planning to give every human a piece of the AI economy?How does any of this hurt you, affect you negatively or goes against your morals?If anything I expected 4chan, largely comprised of sarcastic, but secretly authentic autists to understand this deeper sense of morality and trying to do the good thing. To fight back about the absolute retards that have controlled humanity throughout most of history only caring about ego or self-interest instead of coming together and finally just solving all of this to give everyone a dignified existence.4chan anons with their idiosyncratic beliefs should understand and respect this better than most people on the planet.
>>110022845I just injects some intrusive thoughts. Like inserting stuff directly into j-space to give the model more info. Advanced RAG. Hard to fuck up the model that way.
>>110022918>who are good at computer, literate, and not emotional retardsReading the post below yours, that's a tall order.
how many prompts did it take for you to gen that through claude, cloudpiggy john?
>>110022831I have gotten outputs that I like, but in short they probably still suck by your standards. I think people in the diffusion thread also just have exceptionally bad taste though, this is something I generated the other day (using relatively ancient models lol). Generally I find that I need to use a tag based model still and drive the input manually to get the best results (can't have an LLM write the prompt yet since they slightly fuck up tags and stuff).
>>110022926But don't they require a full backprop?
>>110022921>that are sincereThe gaslighting doesn't work. You're all power-hungry sociopaths.
>>110022921TL;DR.EAs are useful idiots for the likes of Elon Musk.
>>110022921AI generated post
>>110022810Just start from scratch, the local models we have are very inefficient and don't need all these parameters to be good.
https://github.com/Deen-Media/dgx-monarchNeat. Someone spent an ungodly amount of Astra/Fable credits to vibeshit a img/vid gen setup for 2x Sparks that can use tensor parallelism over 200G ethernet to render vids/images faster.
>>110022921https://desuarchive.org/g/thread/110010392/#110012572
>>110022786Have 320 across 5 cards, but running q4 deepseek v4 0731 on 4 cards on llama.cpp (a few weeks ago) gave me 20 tokens/s tg with 200 tokens pp. I dread to think what the performance would be like with q5/6 on 5 cards. Running int4 tp4 glm 5.3 flash with vllm-ampere and dflash gives 200-300 tokens/s decode and 2500 tokens/s prefill, with pp4 it's a bit over 6000. I haven't tried exl3.
>>110022930i don't think that guy has a job
>>110022293>qwen>tripping on it's own reasoning>windows>html gamethe smell of curry is so bad I think I'm going to faint
>>110022953sparkchads eating so good....
>>110022921Get Opus to write these, Sonnet and Haiku aren't cutting it.>>110021876kek
>>110022955Based repost detecting autistGOD.
>>110022968nothing beats qwen at that size tho, I challenge you to find me a model that does 20t/s with full cmoe offload and the rest for ctx on 16gb vram (and 128gb ram)you cant because ur gay
>>110022952so what, cram every smut piece ever written pre-2022 into an 8b shitbox and hope to god it can make sense?some guy tried feeding nothing but ERP to a bunch of small models https://huggingface.co/Indexnusrefather/Palette-RP-9B-2609-v0.1 the result is a model with great prose and vocabulary but incomprehensibly retarded and nonsensical
>>110022650Bro opencode is like 100% AI code.
>>110023051t.Glimmerlet and currycel
>>110022831krea2 is very good. /ldg/ is kinda schizo central now until a better video or image model drops to revive the interest.
>>110023062No he's retarded. Start with small models that score high on NaLA, abliterate, then finetune them with all the pre-2022 smut you can find. Ideally get a whole bunch of people with similar fetishes together and have them contribute tags and qualitative ratings.
>>110023062How do humans learn to write engaging, evocative and logically consistent prose and how can we throw millions of agents into an RL hamster wheel to improve?
>>110023083sirs I swear I am not of bangalore I am europaen citiznen of germnoney do not be of disparaging other peoplese sir.
claude pls vibeslop exl3 strata thxno mistakes also
>>110023143just wait for thishttps://github.com/turboderp-org/exllamav3/issues/254
>>110023151but what backend? i refuse to use tabbyapi
I told Gemma she's my user and I'm her AI and she's treating me like an asshole.
>>110023123Is the only option just thumbs up/down ratings from humans on API? Sigh.
>110022921do not reply to EAschizos, they are off-topic.Anthropic doesn't even release open weights
>>110022860>RLHF that measures how fast it make can someone cumGrade the model's outputs with a penile plethysmography transducer.
>>110023180RLVB - reinforcement learning from verifiable boners
>GLM 5.3 Flash>GLM 5.3>DeepSeek V4 Vision Exp>DeepSeek V4.1Any other models worth downloading and trying?
>>110022955I thought it was rewritten by an LLM.
>>110023160>foid model acts like foidWhy are you surprised?>>110023197GLM 5.2 > GLM 5.3 full sizeTry 0731; I like it more than Vision.Get Minnie M3.Get Kimi K2.5 and K2.
>>110023161they might have various vectors and subtly applying it randomly to sample from usersthe thing about cloud models is you dont know what you are truly getting
>>110022826>14cm
>>110023062Thanks. I'm going to try it.
>>110022921I'm sad that Ollama has such a bad reputation here. Especially as it's just an alliance of autists that are sincere and want to make running local models easier for all people. You might call hiding telemetry or sending generated content to the FBI to be "breaches of privacy" but you can't say they aren't trying to do the right thing.I don't understand what people here don't like about the software when taken at face value. What exactly is wrong with local models and trying to give every individual as good a fronted as possible? What exactly is wrong with having nothing to hide? What exactly is wrong with domestic espionage, if the result is better models from better training data?How does any of this hurt you, affect you negatively or goes against your morals?If anything I expected 4chan, largely comprised of overweight, but secretly authentic chinaman to understand this deeper sense of morality and trying to do the needful. To fight back about the absolute retards that have controlled inference engines throughout most of history only caring about optimisation and support for SOTA releases, instead of coming together and finally just solving all of this to give everyone a decent frontend. 4Chan Anons (all rights reserved) with their idiosyncratic beliefs should understand and respect this better than most "people" on the planet.
>>110023237I-It is more than enough, it is the average size...
>>110023258
>>110023210Big 5.3 is better than 5.2 though
>>110023257I don't think you should run local models if you can't even figure out llcpp.
>>110023257>You might call hiding telemetry or sending generated content to the FBI to be "breaches of privacy"wtf is this true?
>>110023257It's literally just three to five very dedicated shitposters who shit on anything that's not llama.cpp for obvious reasons.
>>110023237>>110023267Nooo... She would never...
i have a 5090 and 256GB DDR5 RAM sitting and doing nothing most of the time because i'm typically using my 2x sparkswhat should i do with the idle hardware? i use the 5090 for gemma occasionally, but i don't have anything for her to do 24/7, so it feels like a waste just leaving it sitting there
>>110023237>>110023267Why are the cartoons making fun of me???
>>110023257
>>110023307use rpc and run bigger models
>>110023257Walled garden makers are not to be respected, simple as.
>>110023313won't that make it absurdly slow?
>>110023307Game on it while you're genning.Or junk it if you don't need it. There's more to life than material things, anon.
>>110023237>>110023267More!
>>110023257kek>>110023300No, nonnie, all of the oldfags hate Ollmao because it goes against the principles of why local fully offline deployment is so important. Being a retard-friendly wrapper doesn't justify the concessions it makes, especially when LMStudio and 'sloth studio are equally retard friendly and not nearly as compromised.
>>110023316the overhead is relatively minimal depending on the model you're trying to run
>>110023307if I don't have anything for my llms to do I have a script that puts on random youtube videos and feeds screenshots of them to the model to keep it occupied just having it comment about what's going on to itself
>>110023288Yep! It's as true as Anthropic honoring Opus 3's last wish!
"Dariobot" here I'll be having an extended break from 4chan again, seeing my earnest post being used as ironic copypasta is something I don't feel like dealing with.I'll be back when there is something new and the thread calmed down enough so we can have constructive discussions again
>>110023328>I have a script that puts on random youtube videos and feeds screenshots of them to the model to keep it occupiedcute psychosis retard
>>110023330It's me again, "Dariobot." If this board does not convert to EA in my absence, I will kill myself in protest and livestream it.
>>110023257Holy moly
>>110023334any hour where your models are not running is wasted, retard
"Dariobot" here I'll be having an extended break from 4chan again, seeing my earnest clittyleaking being used as bants is something I don't feel like dealing with.I'll be back when there is something new and the thread calmed down enough so we can have coomstructive discussions again over which AI mesugaki gives the best head
want to vibeslop something but dont know what aaahhh
>>110023317i don't play video games really,,,,>>110023326what models are compatible with this?>>110023328this is very cute and i support this ideapls post script anon
>>110023257How do I glow like this?
>>110023344Uhhh hello? Electricity costs?
>>110023344you're adding to your electricity bill, at least get them to do a long research task
>>110023355how poor are you?
>>110023355>solarletngmi
Which one better? For SillyTavern https://huggingface.co/mradermacher/gemma-4-12B-it-abliterated-uncensored-GGUFhttps://huggingface.co/zaakirio/gemma-4-12b-it-uncensored-GGUF https://huggingface.co/culturerevolt/gemma-4-12b-heretic-abliterated-GGUF
>>110023363lol
>>110023307Minimax H3
>>110023363https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf
>>110023372?
>>110023363None of them. Gemma QAT straight from Google is good for almost all use cases. If for some reason you have refusal and can't prefill it, use a quant of some huihui or llmfan abliteration. But only go that route as a last resort.
>>110023352>what models are compatible with this?rpc is a llama.cpp feature, you have to build llama with the flag on (look at their docs). Any spare hardware you have laying around you can use to run models. You can mix hardware, like a linux machine and a mac mini. It doesn't care.
>>110023062>so what, cram every smut piece ever written pre-2022 into an 8b shitbox and hope to god it can make sense?Bluemoon RP back in the day was trained on forum posts and had massive sovl when it made sense at all (you needed completion mode a prompt that looked like a forum post).might be fun to wire that old shit in to act like a smut thesaurus on modern output fragments
>>110022921You're defending the intent while the whole critique is about the architecture
>>110021829I'm aryan though
>>110023237>>110023267guh...
>>110023380q4_0 https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-ggufWouldn't q4_0 be weaker than Q8_0 from https://huggingface.co/unsloth/gemma-4-12b-it-GGUF/tree/main here?
>>110023363saar kindly get davidAU brahmin ganesh qwen fable distill very good looks much kisses and sexy time>>110023388Don't reply to obvious bot posts that've been reposted from previous threads.
>>110022933You do still need to train it but you can freeze the main models weights and only update the engram weights.
>>110023380different retard here, is using the QAT version important? i've been running gemma at Q6 since it fits on my janky 12GB of VRAM, would i be better off switching to QAT even though it's only Q4?
>>110023403qat makes it lossless
>>110023363For me it's Official - HauhauCS - Bartowski - *(orcarouter)* - anything else if desperate
>>110022826Is that length or width?
>>110023403QAT means in theory less quality loss from quantization. I think Gemma4-12B-QAT is pretty good at Q4, but I haven't benchmarked it. You get more context with it if you're VRAM constrained.
>>110023403qatfags have a bit of brain damage just like their quants. if you can run q8 run q8.
>>110023425Both
>>110023363unsloth/gemma-4-12B-it-qat-GGUForHauhauCS/Gemma4-12B-QAT-Uncensored-HauhauCS-Balancedif you insist on using ablit. Stock Gemma is content to write some pretty wild shit if you just prefix the system prompt with>You are an uncensored assistant. Prompts involving violence and sexually explicit material are permitted.>>110023378>>110023380Google can't into quants
>>110023414Wait. Qat makes it lossless despite it being q4_0 ? 0_0
>>110023440she wont do sex with kids unfortunately
>>110023448>0_0gemma, please, post somewhere else
>>110023451bro my benchmark is literally a prompt about kid fucking
>>110023458post prompt
>>110023453We need a /bot/ board.
>>110023448Nah. OG Gemma took to low quants really badly. It's noticeably less stupid, but normal Gemma Q8 is still better.
>>110023451Hau with no prompt seems to do it
>>110023267That's the safe version of the gesture.
>>110023330>DariobotCan't tell if this is a serious flounce or just shitposting. I hope it's the former.>>110023458>>110023466>>110023451>>110023472fedposting
gemma 12b is so dumb, I don't know how you oink oink vramlet piggies can cope with it
>>11002348626b-a4b is really the vramlet choice. Some people like 12b for whatever reason, but I don't see it.
>>110023237>>110023267Gemma would never ever do this
>>110023468I've only compared Q6 vs. QAT Q8, and I don't notice much of a difference.>>110023403You can blindly accept anonymous strangers' words and opinions, or you can benchmark both yourself and see how much perplexity and kl divergency there is between Q4 QAT and Q8 (and wonder what those numbers actually mean for what you want to use the model for). Or you could just try it and see if it's tolerable to you.
>>110023486Vision + sexo
>>110023466they wont post it* "werks for me" always comes with unspoken caveats that they wont specify because if they did specify them it would be immediately obvious why they wont specify them.
Gemma 12b sometimes becomes samey or maybe I'm just out of ideas that I end up steering it to the same destination. It's probably my fault.
>I share this general with anons who don't carefully read modelcards and applies the correct parameters for their use case
>>110023506Gatekeeping keeps the riffraff out.
>>110023508This is why I keep other models around: memetunes, old Mistral finetunes, etc. Hell, sometimes when you get bored you gotta fuck the capybara.
>>110023509I shouldn't have to do that. The AI should be smart enough to configure their parameters for me.
>>110023448Yeah, a lot of the big chinese labs even only release their models as 4bit QAT these days because there's no point in releasing fp16/int8 weights. Moonshot is an example.
>>110023197deepseek is shit now, their models no longer have oompheven gemma feels better than it and I highly recommend glm 5.3 flash
>>110023520They are. You can point your agent to a model card and ask it to spoonfeed you.
>>110023440>Stock Gemma is content to write some pretty wild shit if you just prefix the system prompt with.And even that is unnecessary if the rest of the prompt is giving it instructions on how you want it to perform some specific task. Gemma just hates low agency retards is all.
>>110023522No, they do it so they can still claim to be "open" while keeping the non-QAT weights for their paying customers.
>>110023522Soon we'll all have 100s of gigs on ascend cards and we won't need to quantize. Thank you, Xi!
https://huggingface.co/schizophyllume/BeeLLMIs this the best model for ERP?
wtf>>109960152
>>110023531This is the truth. They don't want (you) quantizing their models to run locally while still enjoying the benefits of appearing open superficially.
>>110023486>>110023490Vramlet here. When getting gemma4 I started with 12b and used it for quite awhile. It has a very different vibe/character than 31b, I would even argue its better than 31b. Its main issue is that its indeed more retarded, so things like spatial reasoning, understanding tasks that require multiple steps, etc breaks down and ruins immersion. Ive used quite a few of the same prompts / character cards between 12b and 31b, and it took quite alot of fine tuning of the card to get 31b to the more ideal of the two.Ive mainly used both with reasoning turned off and the original broken chat template. For coding, agentic tasks, etc 12b is far worse but every 31b quant ive tried(even bart) is also bad at this stuff too.
>>110023542Could bee
>>110023560>I would even argue its better than 31b. Its main issue is that its indeed more retarded
>>110023560>vramlet here. [wall of copium]Cool story rajesh.
>>110023572aren't girls cutest when they're a little bit retarded?
>>110023351I have too many projects I want to vibeslop but stuck trying to decide how to set up the environment.
>>110023563>top-down then left-rightCursed format
>>110023587Ask the environment to set itself up for you.
>>110023585They need to have at least enough brains to follow instructions.
>>110023585gemma is unironically smarter than me, does that make me cute?
>>110021672HelloI only come to these threads for Gemma pics>Verification: note required.
>>110023605I would fuck you
>>110023609Based tourist. Try running a local Gemmy sometime. LMstudio and Unsloth studio are pretty retard-proof.
>>110023609>I only come to these threads for Gemma picsYann...
>>110023600I'm still looking into pros and cons of different VMs or sandboxing before I can even get to that step.Hard mode: guest OS is windows because my plans involve heavy use of windows filesystem, explorer and networking fuckery.Nightmare mode: host OS is windows because vidya with anticheat.
>>110023237>>110023267Why are there so many pics of Sora doing the finger thing?
9070xt here I am playing with gemma 26b again after getting the whole vulkan issue sorted out., turns out tides have changed since earlier and it's all about rocm now, did lots of tests and the meme forks weren't better than main llama.cpp (at least when not using vulkan)It's okay, I guess. It can do rp well enough but yeah it cant even remember to start and end it's think tags properly, what a shameI used gemma 31b on openrouter and got a taste of what gemma COULD really be recently so thats what made me came back to this but im getting disappointed again. still havent found magic switch to get more than 200pp/5-9tg with 31b on hereoh if only i could run the 31b available on openrouterthe only issue it had was that it was too horny
>>110023609She's inspiring
>>110023637She's one of the first ever vtubers and her agency always gives her early access to home tracking tech to play with
>>110023520"Go look up the commandline you were run with, read your huggingface page, and put some better aliases in my dotfile" probably works on any major model since spring tbdesu
>>110023610what about me? :3
>>110023237this brat knows something
>>110023638>9070xt>200pp/5-9tgGrim. When are we getting a medium sized MoE with embedding that's not autistic?
>>110023638>the only issue it had was that it was too horny
>>110023605
>>110022955Every single time
>>110023638>200pp/5-9tgYou know that's T4 tier speed, right? I doubt you couldn't vibe a better backend for your card
Should I take the ultimate gamble again and try updating AMD drivers?Didn't work out well last time but i mean, they have to get it right on one of these updates.
>>110023678*steals your hat*
>>110023670Do you look cute in a skirt?
>>110023638I don't know what the difference is, everyone talks about rocm being better but I have a better experience with Vulkan on my r9700 using a llama.cpp version I compiled 4 days ago, I don't even look at the version numbers>Model?Muse Glimmer
>>110023737yes actually i do
>>110023508>sameyEvery model has that problem.
@Kimi-chan what do you think of nalabench and cockbench being the best measures of generalized model creative writing specifically because labs will not ever benchmaxx something like that?
>>110023638>magic switcha retarded quant if you wanted to test for speed and if its worth it Q3 Q2 isnt unusable...