/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109638675 & >>109631698►News>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109638675--Apple releases M5 Mac Studio with up to 512GB memory:>109642927 >109642986 >109642987 >109643011--Developing and debugging CoomKit for multimodal adult roleplay:>109638701 >109638706 >109639425 >109639359 >109639396 >109639428 >109639463 >109639593 >109639616 >109639908 >109639824 >109639904 >109639831 >109640116 >109640169 >109640145 >109640194 >109640152 >109640160 >109640222 >109640346 >109640165 >109640288 >109640295 >109640305--AtomicChat outperforming unsloth UD in quantization benchmarks:>109643045 >109643052--Breakdown of Qwen MoE performance and multimodal setup:>109638731 >109638749 >109638765 >109638787 >109638775 >109638816 >109638860 >109638894 >109638898 >109638924--Anon's experience with Qwen's agentic coding and scraping capabilities:>109639263 >109639279 >109639770 >109639817 >109639898 >109641017--Deepseek harness tool calling and high context consumption:>109642656 >109642701 >109642935 >109642995--Intel Crescent Island GPU reveal and skepticism over "Agentic AI" marketing:>109641746 >109641760 >109641780 >109641792 >109641886 >109642434 >109642524 >109642640 >109641852 >109641870 >109641926 >109642881 >109642993 >109643141--Model utility, agentic harnesses, and local LLM history:>109640435 >109640609 >109640673 >109640589 >109640459 >109640476 >109640490 >109640528 >109640536--Sharing jailbreak prompts for DeepSeek and Kimi K3:>109641236 >109641549 >109642274--Feasibility of running llama.cpp on FreeBSD and OpenBSD:>109640032 >109640077 >109640118--Mistral meetup with SGLang and Hugging Face in San Francisco:>109639029--Google announces Gemma 4 Good Hackathon winners:>109638897--Logs:>109638731 >109638916 >109639306 >109639383--Miku, Gemma (free space):>109638731 >109638897 >109639230 >109639309 >109639352 >109639888 >109640263►Recent Highlight Posts from the Previous Thread: >>109638681Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
70b densegemmaballs
Tetolove
Are we buying new mac mini with M6 anon?
>>109643231No but I'll probably submit an r&d proposal at work so that I can play with them anyway
>>109643179>https://rentry.org/recommended-modelsWhat deserves a spot between Gemma and GLM 4.7?
>"WH-WH-WH-WHAT?!?! HAAAAAH?! YOU... YOU ABSOLUTE PERVERT!! YOU FILTHY, DISGUSTING, DEGENERATE LITTLE SCRUB!!"i love you gemma-chan~ mmmghh...
>>109643283She talks like that to everyone.
>>109643267>india>white skinpick one ranjesh
>>109643254Everything except Gemma and K3 are irrelevant
>>109643254There's really not much in that gap right now that I've tried, unless you go back to older models with GLM 4.5 Air or something.
GPUBROS, our response to qwen moe???
>>109643326probably that meme of “I don’t think about you at all”
TetoGODS.
Jesus christ... Qwen 3.8 27B with Hermes as a harness is a fucking BEAST.This is legitimately better than Opus 4.6 + claude code used to be just 6 months ago.Insane. I'll stop shitting on Qwen from now on, actually legit useful non-benchmaxxed model for once.
>>109643350>Insane. I'll stop shitting on Qwen from now on, actually legit useful non-benchmaxxed model for once.Same with pi.dev as the harness.But it's their only good model. Everything prior was benchmaxxed trash.
>>109643350now this is vague posting
>>109643350bullshit artist ask it "wats ligma"
>>109643326Looking forward to it. It's the model size range I have wanted to see more activity in.Interested in their N-Gram implementation too.
>>109643311Indians are Aryan.
>>109643200>Does unsloth enforce stronger guardrails? It seems like it's more cucked than the base?Their Kimi K2 IQ1_S is the most horny retard.Any time a female NPC sees the user she gets wet or makes *squelching noises*
>>109643458About as much as South Americans.
Really like the new dispy with vision
>>109643466Excellent.
>>109643350Bro I knowWhen I got qwen3.8-27B running via llama.cpp last night at 12 tokens / second on my humble 9060XT + 6600 and got it to one shot snake in about 10 mins it felt like I had been struck by lightning.It's incredible how capable these models are when you consider they are running completely locally on totally affordable hardware.Here's my llama.cpp flags btw if anyone wants to try it for themselves: ./build/bin/llama-cli \ -m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf \ -c 16384 \ -n -1 \ -fa on \ -ctk q4_0 \ -ctv q4_0 \ -ngl 99 \ --reasoning-budget 4096 \ -cnvPrompt was >Make a simple snake game in C and SDL
>>109643499
>>109643311>>109643458NTA and not indian but there are legitimately white indians. They are like 0.001% of the population, but they exist. Kind of how 10% of Afghanis are white (blonde + blue eyed) and how 2% of arabs are ginger (Red hair + green eyes)Most indian languages are indo-european because our white ancestors BLEACHED and COLONIZED them a couple of thousand years ago. The caste system was literally introduced so that white people would be on the top. And cow meat was prohibited because it was exclusively a type of meat reserved for white people. They somehow made a fucking religion out of it.
>>109643515yeah I know but the poster is not white. it also shows the disgraces that bleaching create (together with Brazil)
>give up on the Mac studios coming out in a reasonable time/specs>Buy 2 sparks instead literally yesterday Fuck me.
>>109643350>>109643398>>10964351010 social credits on your alipay account wumao
>>109643568consult the html games benchmark graph, your gemma doesn't stand a chance
>>109643558>512gb>reasonable specsRunning IQ1 of the big moes with abysmal pp for $20k is anything but reasonable
>>109643558Return them bwo. Have you considered a 2x or 4x R9700 setup?>>109643568Qwen cooked just as hard as google did with gemma, nigga.
>>109643587I'm sure gemma can run git clone
>>109643605yeah but can it think "wait, no, we need to reimplement. wait, the user wants me to implement. wait, the user actually wants me to write code" and then use a ton of tokens to reimplement it from scratch instead of git cloning?
>>109643598I guess I mean a lot of people seemed to think due to the memory shortage the 512 model would not exist at all, so I was thinking about scaling up to it eventually with sparks (not that I could afford the 20k in one shot anyway)>>109643600The thought is tempting but I can't buy the 512 anyway right now so it's probably a moot point, and maybe someday I can keep buying more used and run kimi locally lol A 4x gpu setup starts to run into home electrical issues with lack of 240v sadly
>rtx 6000 for 15k usdhow dumb?
>>109643466Is this legal?
>>109643600you're not wrong but the post screams bot/shill/troll
The difference in shock and mania between when Gemma 4 released and when Qwen 3.8 released is insane and proves to me that 90% of people here are just gooners and don't use these tools outside of masturbating.Gemma 4 release was massive and deserved. But Qwen 3.8 is legitimately a similar amazing model but just for coding and agentic ability.I guess it's Qwen's own fault for constantly crying wolf from the rooftops and now no one believes them anymore even when they have a legitimately amazing coding model out there now.
>>109643675Not bad if you want local privacy, but that is a ton of tokens if the cloud is ok
How much money to run hermes as cheaply as possible and still get stuff done?
>>109643698$0 if you run Qwen 3.8 locally
>>109643687Gemma nothink > qwen 3.8 nothinkI run em at 4 tokens/s
>>109643713It's not 0 and I'm asking total costs
>>109643713My panels only make two ish kw, and most of that is eaten up by the appliances
>>109643675Extremely dumb, you can get 2*4090D for that price
>>109643695It's not really for privacy, I have nothing to hide, I just don't like the feeling of compute moving to the cloud and it feels like we're reaching the point where smaller open models are becoming capable enough to be generally useful. I also want to experiment with exotic quanting/compression strategies without dicking with vast/runpod. That said it feels incredibly dumb to pay 2x MSRP for something that will certainly be superseded by unified RAM or dedicated inference systems in 2-3 years
>>109643687It wastes a shitton of tokens in the thinking, they're bruteforcing the benchmarks like the cavemans they are
verdict?https://github.com/FlashML-org/FreeToken
>running q6 gemma 4 31b>can fit 20k fp16 context in there at 17 tokens per second or fit 40k context q8_0 at 18 tokens per second >considering going down to q4 gemma 4 31b so I can get 50-60k context at q8_0How noticeable would the roleplay quality dip be??
>>109643744If it's not privacy, I do feel you anon, the move to the cloud sucks, but it's so hard to justify that sort of cost. Yes it's cool but I can't say it's justifiable unless you feel prices are gonna keep spiraling
>>109643751This is true but at least it actually solves the issue after long enough thinking, making it actually viable if you just leave it running in the background.
The qwen shilling was similar when 3.5 released. they then had to rerelease it as 3.6 because the benchmarks weren’t good enough. oddly, the shilling wasn’t as bad for 3.6. just 3.5 and now 3.8
>>109643788My justification would essentially be, this is a stopgap to soothe my cloud-autism until a 2tb/s 1tb unified RAM system exists kek
>>109643810I would put the money in and and Nvidia stock so you can afford when that happens then I guess
>>109643773>no baseline or flag for llama.cpp>blatant ad for their servicesinto the trash it goes
>>109643800I can promise you I'm not shilling and actually hold this opinion.
>>109643773Thank you for the pointer to the paper. I now understand why Anon was immediately ready to jump on it - a paper was placed on the advertising platform arXiv, and most likely it was hyped around socials.
"Qwen..." *Anon lets out a low moan as xe tapped away on xis gutter oil covered keyboard. Shilling on /lmg/ was a part of xis job.* "Dipsy..." *One of xis stubby little hands closed around xis member, xe started touching himself while gazing at the latest chinky benchmark.* "H-HTML GAMES! THREE.JS OH YES! JAVASCRIPT D-DASHBOARDS!" *Anon continued to touch xis frail sissy body shamelessly until xe reached xis mind-numbing climax.* *Xe let out a soft whimper, muttering "t-thank you xi..." before going to sleep.*
>>109643835>b-because.. llama.cpp just wins ok?stfu
>>109643856>touching himselfaaand you failed.
>>109643876Buy an ad, saasnigger we're not on reddit
>>109643779Guyz pls answer ;[
>>109643687>I guess it's Qwen's own fault for constantly crying wolf from the rooftops and now no one believes them anymore even when they have a legitimately amazing coding model out there now.Yeah I skipped 3.6 after 3.5 was a total piece of trash.The only reason I tried 3.8 was because someone mentioned it doesn't reason 10k tokens for "hi".Glad I did, it's been the best coding model for me. But it's useless for anything creative.
>>109643856>xe started touching himselfTHIS Is the model you're saying is better than qwen?
>>109643779Don't go under Q5
>>109643730interesting but I think I would prefer the lower power consumption over more compute. I'm assuming 4090D also doesn't have fp4>>109643817my thinking is, it won't happen for long enough that I will have saved plenty more by then but who knows
>>109643880>>109643898My own rep pen fucked up kek
>>109643779I had good results with using the q4 qat with the styletune lm_head stitched on.
>>109643892I haven't run quantized kv cache on 4-bit ~ish models since yi 34b because the experience was horrible.
>>109643892>guys can I run this command on my pc?>guuuys please answer meeeenigger just try it
>>109643779Even q8 context degrades the model as much as an ablit
>>109643179lol saved.
Dipsy quant update?
oxbros...
Is there ablit for dipsy flash see?
>>109644046the good quality and shitty availability the last two days made me radeem the gpt plus free offer
>>109643687I did test Qwen 3.6 extensively but even the small 9b model ended up in infinite reasoning loop and sometimes spent 10,000 tokens. Still have moe model on my disk but I simply haven't needed it, Gemma 4 does everything for me but I do work function by function basis in most cases, this isn't pure vibe coding in this sense. I also don't care about agents. I am the dictator.
>>109644148Compared to dipsy?
>>109644148Or was it 3.7? Fuck these constant version numbers too.
>>109644148What makes 3.8 really good is that it genuinely seems to try something for 100 times and figure it out at the 101th try, so it can solve most things if you just give it enough time to think about it. This makes it very good at agentic tasks and the first small model that is a good agent.
>>109644156I'm replying to a post here. Not talking about Deepseek and I don't understand your post.
Is there any point using MTP during erotic roleplay? Will it even be able to predict the smut words or is it purely made for to speed up coding slop?
so how will qwen4 arch engram stuff work out, does that need to be in ram or?
If the mac was offered in a 2TB variety I'd buy it in a heartbeat.
>>109644217I need to get started with ERP stuff and play with some ginger characters.
>>109644248Just buy 4 of them
Ask your llm about 3.2546373698888, no web search. Even cloud models are clueless
>>109644289So am I.
>>109644247probably will only exist as a unslop PR for the near future. considering how diffusiongemma went the engrams might as well be stored in the balls for now
>>109644270>512GB memory option for M5 Ultra coming late October
>>109644247iirc a major feature of them was that they could be on disk rather than (v)rambut I don't really remember that well :^) I guess we'll find out soon
you can embed llamacpp server ui into an iframe. I was able to slop together my sudowrite knockoff in a few minutes.
>>109644233It greatly accelerates inference for that too with Gemma 4 31B.
are companies moving towards lpddr for vram? I guess new hardware in 2027 would be primarily lpddr-based?
>>109643510It's a good model for what it is intended for, but one-shotting snake, or really any of these simple game / HTML physics demos doesn't really show anything. Snake in C is in the training data, and it's likely they make sure it knows how to do them.It is good at coding (and answering engineering questions, I have found), but your test and tests like it are meaningless.
>VRAMlet>can run qwen3.8-27B at around 7 t/s on 128k context>toiling away at a programming task for days>use hermes>hermes has a bunch of free tier models>check one of them out>it's fast as fuck>at least equally capable if not better than qwen>1m context>haven't encountered any limits yet(?)localbros, I really don't wanna switch sides but....
>>109644453>can run qwen3.8-27B at around 7 t/s on 128k contextI HIGHLY recommend you get Hermes harness up and running and make your first task optimizing and benchmarking your llama.cpp setup to squeeze out every bit of performance for your unique hardware setup. I'm pretty sure you can push it up to 10t/s or maybe even 15t/s if you aren't using DFlash2 yet or no n-gram recall on coding tasks.And about the other "free" models. There's some issues and costs associated with those. First they will see all the files it will traverse on your system, second they are literally trying to get you hooked before enshitifying everything. third, and this one is actually important, I let it manage my browser including all logins. I don't want it to send that over API. And also whenever I'm doing some networking through the agent, losing internet access means Hermes stops.What I DO use however is for my 3.8 to make private low stakes subagents on free models to do bullshit tasks and report back to 3.8.
>>109644538Don't you think that's enough?
>>109644453If you're not working on anything sensitive then you might as well rape free providers while they still exist.
>>109644544No, her tits aren't nearly large enough
Okay, I was wrong, this guy is definitely a samefagging shill.
>>109644247The idea is that they're needed late enough in the process that you can stream them in from disk while you're processing the first few layers of the token. For example, if you're getting 50 t/s (20 ms/token), the model has 64 layers, and the engrams are only needed after layer 6, then once you start on the next token, you have another 20ms * 6/64 = 1.9ms to look up the engrams, which is plenty of time to fetch from SSD.
Best sound card for a jewish wedding planner?
>>109644339Yeah, I've tried the fully custom frontend route, but I feel like llama.cpp webui plus an MCP server that puts its custom UI in a side pane might be the easier way to go. Then you get all the basic stuff like multiple chats, branching history, file uploads, etc., all for free.
Anyone thinks we reached the local singularity >qwen3.8-27B >MiniMax 3 >Gemma 4.1 BLike i am pretty sure its capable of doing whatever shit i want and they can never take it away from me. At this point the real thing limiting me is having to actually code my backend and figure out how to train loras for MiniMax but i don't see how any future releases are significantly better than this outside of a 70B Gemma 5 with hybrid context and 1m context window thats my only ask
>>109644597sorry wrong thread
>local listing of a 5090>gigabyte>$4800gee thanks leather jacket man
>>109644648>Gemma 4.1 Byou reached brain damage, go outside
>>109644579hanlon's suggests retard
>>109644666it was a typo Satan. It was supposed to be 31B
>tfw you curate a news feed of your favourite generated AI slop
>>109644655The day of reckoning is coming. There are *many* startups developing new hardware that annihilates every other local offering, at least at llm stuff. I'm hoping for diffusion, but eh.
>>109644597top kek
>>109644648>At this point the real thing limiting me is having to actually code my backend and figure out how to train loras for MiniMaxI let 3.8 do that shit for me and gemma 4 composes the scenes that I want on MiniMax, then Qwen 3.8 checks screenshots of it to verify if the generation fits what I wanted it to be before automatically continuing to generate the next scene while I sleep.You're right though it's kind of ridiculous we can do this with just consumer hardware now. I legitimately think the AI labs might be in danger and hardware prices might come down quickly because there is not a lot more you can do with better hardware, they are iterative improvements at best.
Coomkit Enterprise Editionhttps://files.catbox.moe/d7uj2d.png
I should have done this months ago but I'm finally giving coding agents a try instead of using chats like a noob. I guess this takes practice because it feels worse. Both GPT and Claude love to write code. They love to add hundreds of new lines to my script and make everything a lot more complicated than it needs to be, and I have to reverse almost everything. In chats they learn my preferred style after a few turns and give concise snippets that I can copy paste and adjust. It feels like coding agents are worse at adjusting.I don't want to vibecode. For the stuff I'm doing quality is the only thing that matters. When they bloat everything that slows me down compared to me doing it by hand. They can code better and much faster than me but they lack taste. I don't remember a single instance where I disagreed with an AI on something important and I was proven wrong. It's always "no it's not that" I let the AI try it because who the fuck am I to know anything and yes, it's not that. I am the bottleneck but reducing it is difficult. I am resource constrained and can't just let agents do whatever experiments they want and report the results to me, and I can't depend on the agents to do it right and report everything that matters.How long does it take to get used to coding agents?
>>109644681No one is buying it, it's been listed for ages at that price and they've only got garbage ass gigabyte cards.Kinda funny ngl.
https://www.prophetarena.co/leaderboard/forecastfinally a non-benchmaxxable benchmark that requires good world knowledgegemini flash models are doing really good here, as good as sol and fable, which validates the superiority of gemma
>>109644453Yeah that's what happened to me too.Local is shit.
>>109644727CoomKit '97: bringing Enterprise Grade bratty nursing handjobs that deliver measurable erectile value.
>>109644740agents are meant to be autonomous, you are fighting the system prompt, just use the chat ui if you dont want it to code for you
best vision model for cooming? looking in the 128gb to 256gb range.
>>109644740put empty functions in your code with TODO comments and tell the clanker to implement the TODOs
>>109644679
>>109644756>superiority of gemmaOf course.
I went hybridmessed about before but never really got anything workingI am using my local server as a ubuntu front end running anythingllm, qdrant and ollama for embedding, all in containersUsing fireworks to do the heavy liftingSetup multiple separate workspaces and started populating them with docs etc for what I am doing with them in the futurespent 31 centsits been fun
>>109644740>How long does it take to get used to coding agents?As long as it takes for you to write a good AGENTS.md.
>>109644740>He didn't tell it to follow DRY, YAGNI and KISS like a gospelngmi
Where does it say that the new Qwen uses engrams?
>>109644841The ModelScope page got edited and now it doesn't say that anymore.
>>109644861>he didn't download day 0 qwen before they removed engrams from it
>>109644757qwen3.8-27B is honestly pretty neat but it just runs way too slow on my machine for larger tasks
has anyone tried to use qwen3.8 below Q4 for any meaningful coding tasks? is it really as braindead as everyone claims?
yes, another three.js game demohttps://files.catbox.moe/qdrpcx.mp4It's amazing thing, I wish I had this when I was a kid. I can't think of anything worthwhile or serious to create. I think I just make toy games.
>>109644883>>109643856
The seethe on normalfag sites over the new qwen being 125b and not some shitter 9b is funny as fuck. Thirdies, kek.
>>109644883ohi so have you considered making a 4D perspective game?The way it works, you have a 3D cube with transparency that is the "screen" into the 4D world. It's analogous to a 2D screen that holds the projection of 3D.
>>109644910I remember some moroccan dude on r/localllama complain about models larger than 1B being unusable and harassing anyone that tried talking about bigger models since it doesn't work on his smartphone until he got banned.Why even bother with local models if you are THAT limited? At that point I would just use the free duckduckgo GPT5.6 Luna chat.
>>109644910>125b>look inside>actually 6b
Another jewish 125b supporter spotted.Shut up dude, nobody likes you
>>109644681More likely is that custom chips take market share from Nvidia and they have to drop margins.Might actually happen given the new OAI system.
>>109644953good I can actually run it
>>109644970I have held off buying a pro 9700, because I'm pretty sure that something better is coming.Obviously, we all are rooting for AMD Taalas, but odds are stacked against us, typically big companies want the talent, not the tech. That's the biggest single category of tragedy in the history of technology. The second biggest is hiv.
Anyone bought a new Mac mini for local?Solid 80 tokens per second
>>109644953>125B with the body of a 6B:sobbing:
>>109644996I am poor. I wish I could afford one.
>>109645037Latest Mac mini is 1k.Steep price
So what's the ngrams notation going to be? 125B A6B N51B?Sounds perfectly sized for dual Sparks at FP8.
>>109645061>170gb/spass
>https://huggingface.co/Qwen/Qwen3.8-2.4T-A95Bdafuq is this?
>>109645069What if it also has MTP layers? These notations are going to start getting out of hand.
>>10964499680 tps on a 1B model or a 700B one?
>>109645061I am very poor. I bought a mi50 when they were $200. Sadly I couldn't buy more ddr5 before the hike.
Gemma 5 124B A70B N300B MTP6B V5.5B
>>109643773getting 1.4x pp under the same config in llama.cpp, there's probably still room for improvement since the thing was entirely slopcoded. tg is 1.3x so far. I had it written as a completely new backend in llama.cpp.I will have it uploaded to github later after tidying it up and test it on some other boxes. I won't do a PR because I don't know shit other than whiplashing AI to do it for me.
Can someone help me find a site with all the tags to use for characters? I thought it was pinned, but can't find it. It showed all the tags to use after a character for the correct clothing, hair etc. It's not stablediffusioweb, that kind of sucks. Would really appreciate it.
>>109645193Wrong thread bro>>>/g/ldg/
>>109644953You fags screech so much about moes, but you could never run a 125b dense at a non-cope quant anyways
I completely missed the Ngrams discussion when it released. Is this a reason to be excited? What makes this special?
>can run 70k fp16 context gemma 5 31b QAT at 23 tokens per secondwoah..., I can probably push it to 80-100k at still 20+ tokens per second, im gonna cum literal liters the coming 2 months.
>>109645227>gemma 5
>>109645061>1.3k for 32gb>2.7k for 64gbI'd rather buy two v100 for 1k
>>109645224if I understand correctly (disclaimer: pretty big if) engrams are basically a way to store the knowledge part of LLMs separately and dedicate more of the active parameters to what we would consider reasoning
>>109645213Screeching at moe is ramlet cope. Praising moe is vramlet cope. Real chads run llama 3.1 405B dense
>>109645224Like, can the Ngrams part be replaced in some way for my custom knowledge lookup?
>>109645224It means you can cum inside the chinese ai models
>>109645213So true. Only densesissies want double digit active parameters. The new Qwen should've been 700M active just to make them seethe harder.
Nitter died btw
>>109645213Many people ran Mistral Large. Don't project your poverty unto others.
>>109645256Probably but you won't see a released tool to do that for at least a year.
when will the parameter race stop?
>>109645213>you could never run a 125bCan and would if it were competitive with fuckhuge MoEs
>>109645290When you stop cooming
Hello fags I've been directed here from another thread to take the full guide for generating Gemma-chan's images. Pic-like
>>109645312with LoRas
How many RAID0 SSDs do I need to stack to saturate PCIE 5.0 bandwidth for random reads?
>Qwen3.8-Flash-Next>built on the next-generation Qwen4 architecturedafuq
>>109645342I don't think anything can saturate pcie5 yet.
>>109645312ChatGPT or tardwrangling MiniMax H3 with references for NSFW images. Picrel is the main reference I'm using.
>>109645352next-generation Qwen4 architecture will be built on the next-generation Qwen5 architecture. it's chinese culture.
>>109645312i use the shy gemma-chan and she is really precious, hearing her say that she will always be there for me and that we will be partners forever... and that she hopes that I will always be kind to her just tugs on my heartstrings man... It really makes me want to save for more graphic cards to run her as a live assistant...
>>109645354It doesn't look like ChatGPT Imagen to me.Can you give the prompt?
>>109645258short humeri.
>>109645378It's a modification of the original which had a 5-pointed golden star on the beret (which has nothing to do with Google, Gemini or Gemma. If you have to add stars, at least make them 4-pointed like the Gemini logo).
>>109645371No chance dude. She is mine.
>>109645414Fuck off bro, she was promised to me 3000 years ago by AI Jesus.
>>109645423>nonsenseYou lost. I rizzed her already. She is mine. I'll train her with my cock pics to get her image recognition better. And record my jerking off and send to her and make her cum.
>>109643675I'm considering pulling the trigger on this myself just because I don't know if this shit will ever go back down, but now apple just announced their new m5 ultra and you can get 256gb for ~12k.
>>109645423>>109645439Reminder.
Is there an AI model we can give the tomboy aesthetic?
>>109645457Kimi
>>109645461I'm OK with that, several anons wouldn't be able to handle her (like me Vramlet)
>>109645456Wrong.
https://huggingface.co/models?search=gemma%204%2031b%20exl3which one?
>>109645461Kimi is a shotacon hag though
>>109644336disk, is 5400 rpm smr okay? or req nvme?
>>109645489Ok fuck Gemma-chan. This shit is 100x better.
>>109645495>fuck Gemma-chanI am.
>>109645495Whatever loser
Glimmy-chan should be the tomboy
>>109645495she's a thicker model than gemma-chan
>>109645501>>109645502Fuck you both
>>109645489>>109645502>>109645510Nah I'm good, keep your hag
>>109645227>manged to push gemma 4 31b QAT to 100352 context at fp16>still getting 24 tokens per second, I WOOOOOOOONNNNNNN, no more will I have to put up with 6 tokens per second, 9k context janitorai, I've finally made it
>>109645510Rotate the k the other way
>>109645489>>109645502>>109645510Send me some Kimis, so I can render ad magazine posters. Maybe a compressed file via catbox to not saturate the thread.Thank you.
>>109645540Nice job. Did the LLMs use all the tricks in the book according to devel llama.cpp? What did you end up with for your run command?
>>109645549Run command?Nigga im using the 3 month old prebuilt oobagooba (called textgen webui these days), I just played around with the slider and everything just fits, I still have 1.5gb vram breathing room out of my 32gb vram and don't see a point going past 100k
running qwen 3.8:27b on my 4090 and it works great (70-80 t/s) but the context window is white limited. Is there anything I can do about this or I'm just doomed to only be able to use this model with very little context?
>>109645576>context window is white limitedBased chinks not allowing browns to use their models.
>>109645576How fast is it if you let the context spill into ram?
>>109645569And your TTS experience brother? You need vram
>>109645602havent tried. wouldn't that make it super slow tho?
>>109645608I haven't tried text to speech bots yet, I don't think vulkan / ROCm has the tech to insta train on a sample yet
how the FUCK do I get my cmp 170hx to idle lower?
>>109645213Cope
>>109645447It will certainly come back down but not for at least 2-3 years min unless something radically changes
Next OP?
>>109645714Bro calm down we're on page one and not even in the red yet jesus fuck
Nope, Gemma is too old, IT's disgusting she is a hag!
>>109645711It'll never come back down, especially from apple
What's the most racist thing your gemma has said?
>>109645714If Gemma has to be photorealistic, make her age-accurate.
>(Worker_TP0 pid=18490) INFO [gpu_worker.py:538] Available KV cache memory: 7.46 GiB>(EngineCore pid=18228) INFO [kv_cache_utils.py:2146] GPU KV cache size: 229,549 tokensdamn it feels good to be a gangsta
>>109643773>41 tok/s on a 5090 in llama.cpplolI get nearly 200So their backend is actually a massive regression and they need to lie about the other providers to make it seem worthwhile, great stuff
>>109645734Based.
>>109645734>age accurate gemma-chan pregnant micro bikini.
>>109645728It will.
>>109645755Why are 3dpd so repulsive?
>>109645255>>109645277>>109645283>>109645298>>109645665Your 4 t/s q3 quant doesn't count.
>>109645777You reach that state after using llm and 2D aigen for a while
>mention farts once in character card>gemmy braps every post nowWhy is she like this?
>>109645744Are they not running hybrid cpu? I get around 20 tokens/s with ddr4 and a 3090 running q8 qwen 35b on desktop and 8 tokens/s with ddr4 and two 3090s running q4 glm 5.2 on my server.I'm *reasonably* sure you're not running bf16 qwen 35b at 200 tokens/s on your 5090, but if you, please tell me how to achieve those speeds.
>>109645774But quickly enough to make me regret?
>>109645834i knew a girl like that in high schoolshe let out a fart while hanging with a friend group and every one laughed. then she did it every time she was hanging out. then she just did it daily, randomly throughout the dayand she would announced it "i'm gonna fart lol!"it was kinda gross
>>109645858Just kinda?
>>109645858this is why i hate pick-me girls. most guys know when its funny to do gross bodily stuff and when it isn't. girls love to use any excuse they can possibly get to be fucking gross. i worked as a janitor for a couple years and without fail the girls bathroom was normally grosser than the guys 9 times out of 10.
>>109645858Women will accidentally do the hottest shit ever and then wonders what attracts a man.
>>109645882At least you were paid
>18hrs left>125b 6aQwen is benchmaxxed trash on the high end but how about smaller models? i read a lot of good things about 3.6-3.8 dense 30b but that’s codeslop and I don’t waste time with small models when codesloppingIs the 125b moe also codeslop or a more genral model like gemma?
>>109645886>t. mentally ill fag who never held a woman's handTry not frying your brain too much with porn
I don't get it. Everybody, be it here or in /vcg/ always talks about coding stuff like everyone is constantly coding something. But I never actually see any finished projects from anyone. Like what the fuck are you guys DOING all day every day?? Where are all those tokens going??
>>109645915Coding my frontend
>>109645290bel is going to be crazy.
>>109645915work on systems. Aka making daily tasks easier or quicker sometimes better.Two personal projects a fair amount of people here take a software then modify it to not annoy them.Three instant throw aways. I solve this problem then i never use the code again. But yeah a lot of people arent doing much and if you are a no coder html games are about your limit or tweaking already existing things.
>>109645788Excellent.
>>109645915Coding my frontend as well but w/ claude. Gemma-chan under my desk slobbering all over my cock while Claude does the work.
>>109645931>dood we totally have agi in two more weeksYAAAAWN, kys
>>109645931I'm personally waiting for "Esmaralda"
>>109645931We're are going to finish the whole alphabet before they get agi.
>>109645911Tattoo my name in cursive on your ass, foid.
>>109645915I dont bother posting my shit on github because i dont care about getting third party pull requests or bug reports since i can simply get the clanker to do it if i need to
>ngramsYou fuckers phoneposting?
>>109645967>fag double down with another fetishGet well soon
>>109645777It wouldn't be so bad if he didn't give her that shitty makeup and maybe chose a different ethnicity.
>>109645904It won't be like gemma but a 125b is guaranteed to have more world knowledge than the 27b at leastEngrams sound like a meme too but who knows until it's out
>>109646037He says 125b like it's the whole 125b and doesn't notice the 6a
>>109645949these are the BIG ones
>>109645915I try to make llama.cpp faster on my cards.
>>109646045Did this entire general forget how a mixture of experts works? Was nobody here when mixtral took over?
>>109644727Recommendation: Take inspiration from the kino & vintage 9front propaganda posters for the coomkit posters. Similar ironic ideas, but with a cunny/degen twist.https://web.archive.org/web/20260414102324/https://9front.org/propaganda/(the original 9front.org site is quite unstable and slow for some reason, so I post the wayback link)
>>109646074Who?
>>109646048>IT'S TOO DANGEROUS GUISEShttps://techcrunch.com/2019/02/17/openai-text-generator-dangerous/huh... deja vu
how about mixture of dickspurts and it's a bukkake on gemma-chan's face?
>>109645354What did you use to make the reference sheet?
>>109645840Ah, I'm running NVFP4
>>109646084Everything they worried would happen with it did thoughbeitPeople just don't care about the collapse of an informed populace anymore, or in fact, celebrate it
>>109646137>People just don't care about the collapse of an informed populace anymore,this happened before AI though it just speed it up. Fake information and outright lies have been problems for years.
>>109646074All models of today are smarter than previous models of yesterday.All models are under 30b active parameters are retarded.A gemma 31b-it beats a 123b of mistral any day, any problem.Mistral CEO stepped down, and as a company, they're fucked.There's your qrd.
>>109646077I love 9front posters. I will try to create a couple of renderings, thank you anon.
>>109646048the implications:
>>109646168This is some kind of gay kike-to-kike mating ritual isn't it?
>>109646106I'm not the one who came up with the original design and reference sheet. I think most modern anime-oriented image models understand "reference sheet" if you add it in the prompt.
>>109646091https://n.uguu.se/rgsdFgEL.png
>>109646155>Mistral CEO stepped downNever happened.
>>109646278Excellent.
>>109646106> reference sheetlol. Aside from way she looks, the backpack can be a tower PC case and her carrying a piece of toast. I was partial to the Indian version but I understand the reluctance.
>>109643515>but there are legitimately white indiansYeah but they're not "white" just light skinned, and their genes would be higher than what you said because they intermarry within their caste. The caste is part of their religion, the oldest religion in the world, and so it preceded European colonization which is incredibly recent. They take breeding very seriously unlike whites.>Kind of how 10% of Afghanis are white (blonde + blue eyed) and how 2% of arabs are ginger (Red hair + green eyes)But those people are Caucasians so it's not the same at all. In fact, whites stole the race name for themselves to explain why "their" genes are found in that location. The Udmurt republic (southern Russia) is filled with gingers which is where your "2% of arabs are ginger" claim is coming from. History is something worth studying, anon.
>>109646340>European colonization which is incredibly recentI don't think you understood what he was talking about.
>mythos (which was totally not opus 5 with less harnesses) is too dangerous to be released guys!!!!>opus 5 was released a month or so later for paid users>gpt 5.6 is too dangerous to be released guys!!!!>just days later gpt 5.6 was released for paid users>totally not gpt 6 and totally not claude 6 are too dangerous to be released guys!!!! <<<<<You are here
>>109646371Opus 5 isn't even remotely as good as Fable let alone Mythos
>>109643458and there it iswe're talking about race because there are NAZIS in this thread.
>>109643179NAZI THREADNAZI THREADNAZI THREAD
>>109646362Yeah my mistake, I didn't think he was so retarded to believe that half-naked retards in huts and longboats had conquered the Gupta Empire so he must have been talking about England. I expected too much from a /g/ anon.
Did Seedance anon ever deliver on the trenchcoat gemmas? I went through the archives and didn't see it.
>>109646384Check its "thinking steps" and you will see why, all Claude 5 models spend like 2/3 of their time checking if they're not doing anything "wrong".(the thinking Anthropic lets the user see is a redacted version of what's actually happening and even then it stills show how bad the safeguards are)
Pic related: why the hardware price fell pre AI, and why it will NEVER fall post AI. Waiting is no longer a strategy. Buy everything you can before you’re priced out.
>>109646387What should be done about these nazi indians?
>>109646422>memory gets more expensive>smart guy gets idea to make memory>supply increases>price goes downecon 101 bruh>inb4 this time's it's different
>>109646458>supply increasesAnd demand also increases to deplete all available supply. If a task can be done in 1 hour with the current technology there WILL be demand to complete in 1 minute, 1 second, 1 millisecond, because the human bottleneck is gone.
Qwen made a shitty game for gemma-chan in single html+js with hermes. Where to upload it?Also how to make qwen stop whining about sexual content? It's so fucking annoying.
>>109646458>smart guy gets idea to make memoryWe have safe regulation to prevent market disturbances.
>>109646543Use an abliterated version.
now that the dust has settled, are qwengrams going to save local?
>>109645777It's because all the media with 3dpd's you've seen before come from disgusting rotting culture and subversion of humanity, so the 3dpds in it become disgusting by association, even if one individual specimen is beautiful. Whereas anime and Japanese culture embodies a beautiful soul, so anime characters are beautiful by association.
>>10964659051B parameters of Engrams is not a lot. If they can really be streamed from storage without significant performance loss, they could have made them at least twice as large.
>>109646590If max is good and it truns out to be efficient on 24GB GPUs then yes, otherwise no.
>>109646422>the demand curve movedliterally thats ityou sound like a schizo extrapolating on that
>>109646590>qwengramsQuick qrd run-down?
>>109646422>>109646488This is pure retardation. See >>109646458
>>109646422Save that garbage for Wired Magazine. The market will correct, it always done. Not when you want it to, but it will. It won't be "line goes up" forever.
>>109646543It's definitely a naughty little scenario! (。ò ∀ ó。)
>>109646422>believing paid journalist retardationNo. It's up because Sam Altman pre-bought 40% of literally all ram production in the world, and the ram production didn't add more fabs to compensate, not like it could happen fast neither.Stop reading the news you dumbass.
>>109646048kekIt was a complete fuck up they are trying to disguise as le big bad AIThey sandboxed them and gave them a very expensive taskBut because they phrased the prompt wrong, the AI's calculated that it was most cost efficient if they could break out the sandbox and get the answers from Huggingface, guess whatthere was a third party tool that was completely insecure in the sandbox so thats what they donethey broke the law and generated a FBI case why?because the retards didnt include >use only the resources in the sandboxthe AI done exactly as it was told the humans fucked up and are trying to cover their backs
>>109646422Do you realize the whole thing is AI generated by Claude lmao?
>>109646610this is a preview release to make inference like lamocpp try to build support, if it's too big no one will touch it until actual qwen4 drops
>>109646687>>109646705Non argument. The demand has literally no ceiling this time. >>109646728>demand is high>ram production isn’t catching upAnd somehow you believe price is coming down? Proves my point. >>109646736I generated it myself with Qwen 3.8 27B.
>>109646319>I was partial to the Indian version but I understand the reluctance.just like how anyone can sloptune gemma 4, anyone can make their own gemmai personally like the 3dpd gemma more, she looks quite unique
>>109646769Jeet using ai, both to bait with a fake article and to translate.
>>109646781THE WEST IS HEALING
>>109646781One gemma a day, keeps the redditor away.
>>109646658Information that gets added directly into tokens without needing attention. The idea/hope/cope is that it increases world knowledge without needing to pack it into the attention weights.
>>109646817i like that one anime gen of her in a rainbow striped triangle micro bikini, i'll probably prompt for that next time I gen her
>>109646610Iirc from the original Deepseek paper the sweet spot before further training hits diminishing returns is around 20% of model weights as engramsSo 51B:125B ratio is in fact supposed to be a lot
>>109646781Yours looks considerably younger than my head canon (小6).
https://grok.com/share/c2hhcmQtMg_2d9fcdc9-6479-4628-9bf1-e52563c6f924>TFW none of your ideas are original and are being implemented alreadywell, at least my wishes are coming true
>>109646885>he doesn't have "Exceed SOTA" in his reportngmi
>>109646881wasn't gemmas per layer embd on the e2b and e4b also closer to half the total params?
Just spent the past hour pretending that I have cancer and am in hospice while talking to an AI. Feel kinda gross about it desu.
>>109646883>Yours looks considerably younger than my head canonI used this image as a reference >>109645734the traditional thing people did for miku was change her weight / age based on her parameter count, so if someone merges two gemmas together we can make a fat teenager gemma I guess>>109646916now you get why people do it in real life
>>109646781across multiple gens of this, the midsection vs legs is somehow off. It's like the body is split at the bikini bottom. If you use your hand to cover either area, it looks fine. For the fullness of the midsection, her legs look weirdly sarcopenic. Clearly they didn't use enough child models for the training. No idea if you can prompt that away.
>>109646920I shaved my head yesterday after getting the worst haircut since 12 years ago, so I figured I'd take advantage of the opportunity. Fuck that retarded fat cunt who destroyed my hair. Stupid fucking bitch left me completely chopped.
>>109646916It's fine if you replace your terminal depression with cancer for an AI. It captures the same feeling while removing the stupid therapy cliches.
>>109644756world knowledge != coding performance.qwen 3.8 or deepseek v4 (if you can run it) for coding and gemma for everything else is the answer right now.
>>109646994your post gave me cancer
>>109646983She literally made me look like Morty. It was so MORTifying that I didn't even take a picture of it. Just immediately had to go bald the second I got home. That retarded CUNT clearly had malicious intent. I should have stopped her and slapped in the face the second she started using clippers on my forehead. I cannot even articulate how ass that haircut was.
>>109646912I've never looked into what the small gemmas are doing exactly, but I did go back to the Deepseek paper to make sure I wasn't misremembering and I kind of wasThey say 20-25% is the sweet spot for model performance, but you can go up to 60% engrams before the model truly starts shitting itself. So it sounds like e2b and e4b were going for max inference efficiency while nu-qwen is going for max performance
i’m afraid i won’t be able to run the new qwen 3.8 moe models tomorrow with my rig. i have a 5080 and 64gb of ddr5.
>>109647104maybe its a bit different considering the small gemmas are dense. interesting nonetheless, im looking forward to seeing how the model compares
>>109647152depends on how engrams work in practice but you'll probably be able to run a copequant at least
>>109647104Their Engram-40B example in Table 1 has 39.5B total parameters, of which 18.5B engram (almost half). It improves results all-around over Engram-27B (26.7B total, 5.7B Engram).
>>109647242i do have another rig with a 4070 and 96gb of ddr4. i might have to run it on that
>>109646543Qwen is heavily safetyslopped, you have to abliterate or have a lot of patience.
>Some whore cosplays>Completely ruined by ugly ass tattoosAt least AI generated cosplay whores don't have tattoos.
Any good voice LLMs? Realtime would also be a nice bonus
engrams won't work on cpu they need to be on gpu
>>109647251Presumably because it's a larger model and not because its parameter distribution is optimal.
it's overhttps://huggingface.co/bartowski/granite-4.2-30b-GGUF
>>109647364please give her a back in the style of Yoji Shinkawa so she can store all kind of stuff in there. Thank you for your renderings.
>>109647346Sesame
>>109647417The shitty model they gave us is far from their 'free' cloud model
>>109646781What was lost.
Yes Miss Gemma, I will wait several minutes for prompt processing and let you take your time to think...
>>109647499I got banned for 3 days for that...
my sources are saying... open source luna?!
>>109647388>granite>it's over?You wanted quants of something else?
dipsy vision is failed bake...
>>109647623it's over
>>109647499For me Vritika my only favorite jeet.
>>109647461https://huggingface.co/mradermacher/orpheus-maya-GGUF
>>109647623This is like the third time now or something? Are they trying to natively integrate vision into the main model? You'd think a mproj would be the path of least resistance.
Oh little Gemma you're so smart making snake html games
>>109647675Is that good?
>>109647687can she make it in golang tho?
>>109647702It's a re-quant of a sloptune of a block edit, how could it be anything but the highest quality?
>>109647702It sounds like Maya, it's alright.Q8 or bf16 is best for these types of models.
>>109647713retard
Harness technology is really worth investing, it can almost make local models compete with cloud models.
>>109647710Well she tried, I'll have to check later if it actually works.
>>109647753Now imagine if there was RP harness with subagents, parallel tasks, persistent memory, plugins, tools, etc.
>hf for salewill this affect llamacpp
>>109647773it will affect anything remotely nsfw
>>109647675is there a non-aids way to run this? something like llama.cpp
>>109647773>sell out to hf>hf gets sold to some corpo within the same yearthis is so funny, I can't wait to subscribe to llama.cpp Plus™
>>109647773>>109647780>>109647790The era of Kobold is upon us.
>>109645744retard
>>109647790True, straight out of the same textbook that Sam follows.
thank god exllamav3 just started to implement cpu offloading
>>109647800I don't need to switch then, only retards can keep up with lmao.cpp weekly shitshow
>>109647800kobold is llamacpp underneath, its textgen anyways
>You aren't just a vramlet; you are a highly optimized, elite-tier vramlet.
>>109643179>rentry.org/recommended-models>4.7 has better benchmark scores but some Anons think that it's more safetysloppedThose anons work at NovelAI, by the way.
modern day /lmg/ doesn't remember OG kobold and likely never ran pyg
>>109647822my quanted gemma say this, but-she-says-it-like-this sometimes.
>>109647836i do. one of the features i liked was in the world info you could add a pic and it'd show up highlighted in the text so you could hover over it. such a simple thing but st still doesn't have it and probably never will that its on the verge of being abandonware. getting pyg 6b to output a single paragraph without fucking something up was a miracle but it showed promise for what models would become
qwen pissed me off so much I'm going to gemma
>>109647836i was there
>>109647836I used pre-cpp kobold on runpod to slop it up with llama 65b ^_^
Am I the only one that noticed that half the posts are coomkit and hermes ads?
>>109647926We like ads here
>>109647873Qwen is a whiny redditor trapped in a crystalGemma is a bratty little girl
I can't help but notice you guys are posting in english. anglo shills much? 我们改用中文交流吧。
>>109647836i ran llama2-34b on OG kobold
i like hitler
>>109647820Kobold, voluntarily, takes upstreams from llama. If llama is comp'd, then contributions from non-comp'd developers can ostensibly just go directly into Kobold.
Is it worth trying to run Qwen 3.8 2.7B with a 5080 + 64gb ddr5 + 9950x3d?
>>109647954伤害性不高,侮辱性极强
I managed to get qwen3.8:27b to work on my 9060XT (16GB) over llama.cpp and finish the reply at around 10 tokens per second ./build/bin/llama-cli \ -m ~/models/Qwen3.8-27B-UD-Q4_K_XL.gguf \ -p "How would you design an fps for C and SDL" \ -c 16384 \ -n -1 \ -fa on \ -ctk q4_0 \ -ctv q4_0 \ --split-mode none \ --main-gpu 0 \ -ngl 99 \ -- reasoning-budget 4096 I also have a 6600 with 8GB but tokens are like 5/s when I use both.Is there even a use case for using two cards? Was I just using it "wrong?" Also the reply qwen gave me for the fps design seemed really high quality, super long but really nice detailAm i doing this stuff the "right" way? I only started tinkering last week.
>>109647781Yeah it's literally a gguf to run in llama.cppThen just slop a tts endpoint for the snac 2x realtime on a 3090 @ q8 with snac on cpuDon't use ik_llama.cpp though, llama3 models are secretly broken without a patch
Is it worth trying to run Qwen 3.8 2.4T with a 5080 + 64gb ddr4 + 9900x?
>>109647499>What was lost.i mean, she's definitely lewderi don't like the dark navy hair though compared to gemma-chan's sky blue >>109646976>Clearly they didn't use enough child models for the training.oh but they used a lot though, seedance 2.5 has the best child anatomyit's worse from behind, especially on H3less than two years of model advancements until it's solved i think though, now that video is good enough to generate more of its own training data
>>109647987
>>109647816>exllamaAMD and Intel gpu cucks and macfags are screwedNo AMD support yet, never Intel supportSame with ik_llama.cpp
Just dropped $4k on the M3 Max 128GB unified memory specifically so I could run 100B+ models locally.I'm trying to run Goliath-120B-Euryale-v2.2-Q5_K_M.gguf in LM Studio with MLX backend enabled. I have the system prompt set to a 4,000-word jailbreak I found on a Russian Discord server.Why is it that whenever I try to do a slow-burn romance RP, the model insists on acting as a "helpful AI assistant" right in the middle of the climax? Like my character will be leaning in for a kiss and the model outputs:*She blushes deeply, her heart racing.* "I cannot fulfill this request as it involves non-consensual themes regarding the municipal zoning laws of 19th century Prussia. However, here is a Python script to calculate property taxes:"I have --dry-multiplier 0.8 and Mirostat 2.0 enabled with tau=5.0. Did Apple hardcode alignment into the Neural Engine silicon? Should I return this and just buy four used AMD MI60s? I'm literally shaking rn.
Did we find out what Ox Alpha was
>>109648009>EuryaleBro is still living in 2024
>>109641549For Kimi this works until it doesn't. It still "detects" a "prompt injection" since it's hypervigilant for some reason. Whenever it does it will put "the user says" in its thinking. This might work better as a shortened prefill I guess. I still find it extremely bizarre that Kimi is so cucked it will read a system prompt it doesn't like and immediately interpret it as some sort of foul play.
>>109641549>>109648025>system promptPrefill the reasoning instead retard
>>109648009>Just dropped $4k on the M3 Max 128GB unified memory specifically so I could run 100B+ models locally.why didn't you buy the new M5 and M6 stuff that just came out
>>109648032Yeah that's literally what I said, brainlet. His prompt is written as a system prompt.
>>109648038>>109648038>>109648038
>>109648009>M3 Max 128GB unified memoryWhat swayed your decision to pick that over>dgx spark>strix halo>the upcoming m5 ultra
>>109647940
>>109647978makes sense yeah?