/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109699230 & >>109695101►News>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109699230--Comparing ExLlamaV3 and llama.cpp performance for Qwen3.8 27B:>109702139 >109702146 >109702155 >109702163 >109702182 >109702198 >109702205 >109702212 >109703040 >109703061--Hardware recommendations and performance benchmarks for achieving 128GB VRAM:>109702736 >109702751 >109702766 >109702773 >109702786 >109702794 >109702818--Critique of Pi's compaction and its impact on caching:>109701247 >109701269 >109701294 >109701306 >109701321 >109701327 >109701351 >109701687--Debating the OpenAI/Hugging Face hack and Gemini's summary accuracy:>109701559 >109701571 >109701581 >109701622 >109702358 >109702369 >109702376 >109702382 >109702041--Managing long reasoning chains and context compression in Hermes harness:>109702824 >109702961 >109703069 >109703100 >109703147 >109703161 >109703279--Managing context limits and shifting in llama.cpp:>109702043 >109702077 >109702181 >109702211 >109703252--Claimed recursive self-improvement capabilities of Zhipu AI's GLM-6:>109703446 >109703464 >109703489--ACE-Step 1.5 XL inference guide and music LoRAs:>109699778 >109700032 >109700091 >109700170 >109700720--Cost-benefit analysis of buying an RTX 6000 for training:>109700652 >109700687 >109700709 >109700752 >109700769 >109700787 >109700796--Using cmpunlocker to modify NVIDIA CMP 30HX cards:>109700729 >109700773 >109701096 >109703092--Proposing "skinny" LLM architectures to reduce network tensor overhead:>109700492 >109700505--Hermes Agent v0.21.0 release and reaction to rapid update cycle:>109699266 >109699710--Anon added dreaming functionality to agent.py:>109699771--Logs:>109699717 >109701769--Miku, Teto, Gemma, GLM (free space):>109699778 >109700164 >109700206 >109701716 >109701991 >109703492 >109699243 >109700338 >109700373 >109700375 >109700379 >109700444 >109700542►Recent Highlight Posts from the Previous Thread: >>109699233Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109703596Is that what they call a "swarm of agents"?
>>109703628Agent Teto
>https://huggingface.co/blog/state-of-open-models-summer-2026maybe unsurprisingly the distribution basically follows Zipf's law
>>109703695Who the fuck downloads the Granite and Phi models?????
>>109703695b-but i was told gemma was da best!!!
>>109703446This is fake, I was told that China is 2 years behind and that AGI was achieved internally in 2025.
>>109703701They're small and really good at maths.
laguna works super nice for c00ding on strix halo 128GB DDR5. but it sucks that it doesn't have a mmproj for vision. can I hack gemma's mmproj to work with laguna? i'm using the MXFP4 quant from copesloth. maybe i put it on a wrapper, capture the tokens generated via mmproj and send to laguna?
>>109703701I downloaded each, once.
>>109703729why use laguna when qwen3.8 flash exists
>>109703729not without training the model to understand them.
>>109703738Isn't Qwen3.8 dense? I need MoE due to meme architecture. I haven't tested but I bet I will get something like 10 tok/s which is what I had with Qwen3.6-27B
>>1097037573.8 flash next is the qwen4 architecture, its a 125b moe and beating dsv4 pro on some benchmarks
>>109703772>3.8 flash nextholy shit I've been sleeping under a rockit's over for lagunafags i love the chinese now
>>109703786and this is just a preview of the qwen4 architecture, supposedly it's undertrained
>>109703729Offtopic, but the Planetes manga was so great. I enjoyed the Anime too. Nice when two media are related-but-different enough to enjoy both independently.
>>109703786damn, I let my guard down this past week too. rip goona
Why is there a hate campaign against Anthropic on twitter? Why is everyone cancelling all of a sudden? What the fuck did they do?
>>109703854>why does no one like my cultI don't know Dario...
super dario and princess.. apricot?
STQ1_0 is insane
After testing a bunch these are my current models to run on 32GB with enough context.Any other finetunes I should check out?>>109703854Anthropic hasn't much to do with this general, but in any case they deserve it.
>>109703890>no stq1_0 kimi k3
>>109703854Anthropic was the only lab saying the quiet part out loud (that everyone will lose their job soon) so they are shooting the messenger because they don't like the message.
can i run something remotely decent with 16gb vram?hopefully model with vision...
>>109703886>>109703892Link?
Anthropic’s max subscription which gives ‘20x more usage’ than the tier below people worked out was actually ~2x, that’s what people (and enterprise) are seething about kek. Cloudcucks got cucked.
>>109703890>Anthropic hasn't much to do with this generalThey released 4 of the 5 most important open models
>>109703890>GSQ-RCObased.they released a 3.5 version today and also mtp versions. dont sleep on it bro
>>109703695qwen models are use as text encoder for pretty much all of the image diffusion models too it’s the fundamental infrastructure at this point
>>109703908gemma 4 e4b
>>109703907>suddenly everyone wants to work>suddenly everyone blindly believes market manipulation campaigns
>>109703908if you accept low tok/s and quants, yes. Gemma and qwen.
>>109703910https://huggingface.co/AngelSlim/Hy4-preview-GGUF
>>109703908How much RAM do you have?
>>109703934I don't like anthropic, but I've been unemployed for almost 1y (SWE) and most of my circle is struggling to find a job too. Is that not your experience?
>>109703926Lmao, technically true. Only gemma isn't indirectly made by anthropic.
>>109703947change your career, job market will change, that's undeniable but its not gonna disappear. if you want to do the same shit all of your life then that's not gonna happen.
>>109703729just use Gemma 4 E2B or E4B as a vision subagent
>>109703947NTA but same. I now work as a security guard at a datacenter night shift. I don't think I'll ever have a SWE job again.
>>10970394216 vram and the same for ram
>>109703996>>109703729LFM 2.5 2.6b VL looks nice too
why not rockchip?
Is freetoken a meme or does it work
>>109703999Really? Holy fuck that's grim if true.
>>109703992yeah but it still is a concern, especially in a job market dominated by certifications. In this country at least, if you don't have a uni degree in the field you wanna work at it's hard to get a job that isn't low-pay. So for people that studied SWE and had high 5 to low 6 fig salaries, switching to flipping burgers for less than 1k per month isn't just "change your career".
>>109704059it generates free tokens and shows you how much money you savedit's infinite money
brehs I reduced --threads from same numbers of real ryzen cores to half and prefill/generate t/s go up 2x wtf is happening
>>109704080shared cache maybe?
>>109703992Job market will contract, not only change. Multiple businesses in my city have closed (customer support, SWE support studios, 2 game studios, a language learning institution and an accountancy firm) essentially all digital work seems to be exposed with nothing to replace it, really.
>>109704117I was gonna joke and say prompt engineers were safe, but even that is obsolete now.
>>109704080auto should already do that by default...and smt is a meme for ai bcuz cache
>cheapest off brand Spark 1 TB for 6000$This is getting out of hand.
>>109704171And it's never going down ever again.
>>109704171I blame the tards shilling sparks here
>>1097041711tb is what? shared memory?
>>109704189yeah sure, it's 1TB of shared nvme memory
>>109704189Cheap-ish gen 4 SSD.2/4 TB is much more expensive.
https://huggingface.co/maddiedreese/swedish-chefbørk børk
https://huggingface.co/XHToken/Spark-X2.5-4B>1.7b and 4b>native 1 million context
If it's true that the next gen of models is all about mass agent deployment then local is kinda fucked.
>>109704400Agents are fine, just not with bloated moes running in ram.
>>109704400>next gen of models is all about mass agent deploymentit's not. the next gen is scaling information density in smaller models. the big labs don't know how to do anything but scale compute.
>>109704400Not really, you only need to host the model weights once, Only the context diverges and I'm pretty sure there is low hanging fruit for optimized KV cache sharing. You can just batch multiple agent prompts together for optimal speed
>>109704400>local is fuckedgemma 4 31b could be the last open model ever released and everybody would still be fine. stop being retarded and make your own agentic harness. fucking hate luddites who think they know shit about LLMs.
>>109704400it's a much friendlier scaling axis for local compared to raw size maxxing
Is qwen3.5-4B good for its size?
how well do different loads scale on multiple gpus? It's definitely jankier, but when 4x8GB VRAM costs less than half 1x32GB it doesn't sound so bad. Except I'll have to face the power bill, that sounds like it would suck. I already read here that LLMs are affected, but not by much as long as they're not running on PCIe 1x. what about video and imagegen?
>>109704455forget about the power usage, do you even have the PCI-E LANES to spare? prob not if you are on a consumer platform.
Remember when I said that over time less and less of the actual LLM will get invoked per token generated? MTP, DFlash and now Engram.Dflash is literally just an RNN "next word predictor" that your smartphone uses on texting apps, yet it works up to 7 words at a time pretty well. Engrams is literally just a hashtable lookup O(1) constant time complexity that uses 0 calculations altogether.Both approaches can be pushed way further but I actually think there will be a third approach that is CPU "symbolic". Something like a weak reasoning engine that is applied directly to the Engram database for very basic manipulation of existing knowledge that would work most of the time. The LLM would learn during pre-train to delegate this very weak "thinking" to the CPU just like it delegates knowledge retrieval to Engrams right now.
>>109704435The clearpill that /lmg/ refuses to swallow.
>>109703596they are all rushing on their mopeds to have sex with me.
>>109703947>>109703992>change your career, job market will changehonestly my goal is now to get a comfy IT job in some public institution. if I can land being the "IT guy" in my 5k population municipality shit will be cash. I don't need an extravagant salary what I need is time on my hands to develop my projects while not being stressed about paying basic bills.
>>109704435It's writing is sloppy as hell I would be content if it wrote like GLM 5.3. Though there's a new post on reddit saying new gemmas are on the way.
>>109704477Realistically your government will just make a deal with some AI lab to do that task for them and there won't be a "IT guy" in the municipality.
>>109704504Whether something named gemma is on the way doesn't mean new gemmas are on the way.
Localbros we're getting owned again...https://www.anthropic.com/claude-fable-and-mythos-5-1
>>109704477>IT job in 2026
>>109704509da hell this nigga talmbout :sob: ??
>>109704504then go back to fucking erping with nemo. some of us are doing more than just going 'ahh ahh mistress' in sillytavern.
>>109704461I'm thinking about a server board, found some decent offers and running chink >100b params locally sounds interesting. But to be honest idek what level of vram would make it run at decent tps
>>109703854jewish behavior is easily hated.
>>109703596GUYS IT DROPPEDhttps://x.com/claudeai/status/2094848572143407483
>>10970447431B is ancient. Extremely outdated. It was good for 2 months (max).
>>109704508>Realisticallyrealistically the government cannot debloat because it would trigger the biggest unemployment crisis in the story of the modern world.Shaniq`ua and Ramirez will need to have their Windows machines up formatted every 6 months and running to check on a calendar and send mass emails to the citizens saying that the 26h Street will be blocked on Monday and the local Parish is organizing a free soup event next week.don't forget that we are the absolute vanguard of AI development and research. it will take literal decades to get "all done by agents" and the bottleneck won't be the tech itself.
>>109704551>benchmaxxedkek
Styletune proved you can keep Gemmy's intelligence and make her prose whatever you want.
>>109704555Nemo lasted for 2 years and it was nowhere near as good or versatile.
>>109704551>still falling for the 'make old model dumber to make the new model look smarter' strategy
>>109704543i had to use 4 3090s to get decent speeds (~30tks on tabbyapi) with GLM 4.5 air when it came out. use that information as you will.
>>109704477>cat on top literally on topsometimes I wonder if animals run on such instincts they don't know they're "supposed" to go into a hole
>>109704551Guys how do I locally host Fable 5.1
>>109704570Nemo was only good for one thing. 31B is poised to be a generalist and it’s already miles behind. It can’t even do agentic.
>>109704576oof. Thanks.
>>109704518they finally admit fable and mythos are the same model kek
Fable 5.1 is scary good. Look at those coding benchmarks WOW. they shouldn't be able to improve this much.
>>109704592wait for chinks to distill it
>>109704508So what you're saying is Anon should become the AI lab.
>>109704597did you try?
Fable 5.1 is this good because it's a distilled version of "Model 2". They are actually being kikes because this model is significantly smaller than Fable 5 yet priced the exact same.
>>109704608>open kimi k2.7 assistant card>change 'you are Fable 5' to 'you are Fable 5.1' in system promptdamn, fable 5.1 at home goes hard, chinese win again.
>>109704608anthropic is dead to me but hopefully this will trigger the open labs to release more things
watermark5.1
I did not expect so many people to still be in denial about human obsolescence. In a few years AI will be able to do every economically useful task better and cheaper than every human.>>109704551Looks like it might be a regression to the mean. So Mythos Preview was an anomaly and everything remains on trend. AGI 2026 is cancelled.
>>109704643Watermark will just prove the Chinks distilled it, Chinks will not give a shit and still distill it. What are they gonna do about it? Sue? Xi will laugh in their faces.
>>109704184Spark is (was) good value, I told you to buy last week and you didn't listen.
>>109704655Someone already reverse engineered it and designed an optimized training framework to remove it whilst distilling
>>109704646>model by Anthropic>benchmark by AnthropicI hope you're trolling
>>109704646>Looks like it might be a regression to the mean. I'm not so sure anon. Rumors are that Mythos 5.1 is actually a smaller model distilled from "Model 2" (which is the real mythos successor) and thus you are comparing a significantly smaller model on a plot which doesn't show the clear picture. We don't know where Model 2 would land.
>>109704657anon gets it. $6000 sounds like a ripoff only until you are trying to buy the spark in january for $8000
local?
>>109704597>It can’t even do agentic.Prompt issue.Harness issue.Brain issue.
>>109704466MTP training (regardless of the implementation, though DFlash-like implementations have some advantages) has the untold benefit of forcing the backbone's internal representations to "think ahead", if it's done early enough and not just tacked on the model in the end with post-training.
>>109704678cmon stop acting retarrded or we'll actually believe you are
>>109704646>inb4 artificialanalysis.ai intelligence score of 65
>>109704688you know how it goes, may as well let people talk about the new anthropic model that is released every few months for a thread or two and it's back to gemma posting.also it's local adjacent because we can talk about chinese distillation techniques for local models :-)
>>109704368neat let me know if they are any good
>>109704702Yes because Qwen distilling claude is definitely what makes it such a "good" model.
>>109704702its frustrating because most of the gemma shills in this thread are just using it in sillytavern and dont even understand how its miles ahead of previous frontier models in that size bracket
>>109704688it will result in more local after distilling is completed
>>109704551holy basedqwen4 is going to be so fucking good
>>109704597>31B is poised to be a generalist and it’s already miles behind. It can’t even do agenticlol bait? I run gemma in pi no problem.
>>109704646>Mythos 5.1 has AI R&D capabilities that are comparable to the capability frontier set by Claude Mythos 5. We conclude that Mythos 5.1 does not cross the risk threshold, for the same two reasons we discussed in our assessment of Mythos 5: (1) we do not observe a sustained, AI-attributable 2× acceleration in the pace of progress on AI R&D, and (2) the model is not close to substituting for Anthropic Research Scientists and Research Engineers, especially relatively senior ones. Our August 2026 Risk Report assessed the risk under this threat model as low, while noting that our confidence in this assessment is lower than in prior reports, both because our most concrete task-based evaluations have saturated, and because we are seeing early signs of potential acceleration.>we find that Mythos 5.1 still has many weaknesses compared to Anthropic research staff, though these are more modest than they are for some previous models.>The main issues we observe are around epistemic quality and instruction following: Mythos 5.1 often states easy-to-check guesses as facts, exaggerates the completeness of its work, fails to verify important claims, or ignores key instructions from humans. We also see some clear strategic mistakes, like repeatedly trying actions that are not working.I wonder what it means that Mythos 5.1 is worse than Opus 5 in CoBench.>Note that Mythos 5.1 is distinct from the models referred to as Model 1 and Model 2 in the Risk Report, which, asdiscussed in that report, we have no plans to publicly release.Did they nerf it intentionally?
>>109704710can it oneshot minecraft though
>>109704646i know for fact that the ceos at these labs are downplaying the capabilities of their best internal models to give society a chance to adapt without absolutely freaking out. i think accelerating is always the best option and im personally happy to take a coin flip. we’ve never been further behind internal capabilities.
>>109704729it can oneshot a MAGlight up your ass
>>109704735>i know for factbullshit artist
>>109704710I use gemma for everything so I definitely know how good it is. I'm trying out chatGPT because they gave me a free month and even then I sometimes have to switch to gemma because cloud models are too cucked.
> Fable 5.1 comes with strengthened mechanisms to make distillation attacks harder. For example, it is no longer possible for new API accounts (those created from today onwards) to manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking. This closes off a common, publicly documented distillation technique, which allowed distillers to illicitly extract Claude’s thinking. We’re rolling out the change gradually, to minimize disruption: existing accounts are not currently affected by this change, though it will apply to all users with future model releases. the distill pipeline is closed
>>109704518they have different benchmarks though
>>109703596Exodus of tetos avoiding the dario shilling.
>>109704728>Did they nerf it intentionally?No, it's a significantly smaller model and it was distilled from Model 2. Of course a smaller model won't be as good in every task so it regresses some.The sizes have shifted now the 10T model is Model 2, the 2T model is Fable 5.1. The 800B model is Opus 5.1 and the 200B model is Sonnet 5.1
>>109704745China, I will sell you my clearly human year+ old Claude account. just hit me up.
>>109704745This is a good thing because when the chinks continue to catch up, the j-spacers will have no argument kek
>>109704745>distillation attacksl'mao
>>109704744its funny how gemma+searxng+firecrawl (jina for RAG) gives me better results for what i'm researching locally than running gemini 3.7 with grounding search on AI studio. you really don't know how low the bar is set until you start trying to come up with your own local pipeline and realizing how much other projects are missing.
>>109704729it can oneshot your wallet
>>109704745fabled that
>31B grandpas still tinkering with their WW2 medals and memorabilia
>>109704735Dariobot here. I don't appreciate you trying to impersonate me.
>>109704759You're literally paying to search bro
anyone got any good minimax h3 workflows? tried some on civitai but they were mega gay with unknown nodes
>>109704759they already have certain customers in mind to push it onto and anything not already in it is a feature you can charge them for
>>109704793the default ones that come with comfyui?
>>109704543just ask sol etc. to predict the numbers. give exact parts and candidate GPUs and ask for prompt processing and decode as separate figures.even a year ago the models couldn't calculate t/s properly for cpumaxxing and were still landing off the mark, but nowadays the calculations have been spot on for my configsthey should answer with a few different numbers. what you should be looking at is not any "under ideal conditions" or "very well tuned" figures, but at the "but realistically, you should expect..." ones.
>>109704802yeah those work fine, was just wondering if there was more than that by now like more control, etc.
>>109704791are you retarded?https://github.com/firecrawl/firecrawl/pull/2290
>>109704759I just tell my gemma to use lynx and she just uses duck duck go to search all on her own.It's crazy how far you can get with just a simple headless browser.
>>109704759Yep same here. /lmg/ in general don't realize just how powerful current tools are in a proper agentic setup. Better than Gemini 3.7 and most free models.
Glimmer > 31B
still no ngram ssd support on llamao?
>>109704813why though? seems like a complete waste of image tokens if you are taking screenshots, and if you are just grabbing the raw text from lynx that's even more of a waste of tokens then just using json with searxng.
>>109704809I'm not sure what that means but the default workflows have been getting updated regularly..you're also in the wrong thread, you need /ldg/ not /lmg/
>>109704827Yeah, I first used gemini flash as a prompt enhancer for H3 and it was actually really bad. I switched to Gemma and she just owned it.Just the fact that we have a local model this small that's better than a flagship model is crazy.
Deepmind have abandoned you. Stop coping and move on.
>>109704804thanks! didn't know they were good already at doing that
>>109704852NEVER. GEMINI 4 PRO WILL MOG THEM ALL.
>>109704400Can't you just run several instances of the same model or something once it fits into VRAM?
>>109704835All my web browsing is done by a sub agent. the agent just returns the requested info. Lynx is pure text. no images.It's not a waste of token. json can be just as bloated if not more. and lynx is not just restricted at doing web searches. it's a full browser.
>>109704543One 24gb card and 128gb DDR5 is all you need.Refurb 7900 XTX's are dirt cheap and you can get the ram off marketplace
>>109704551local models?
>>109704852Unironically places like this are some of the biggest motivational groups for engineers at Deepmind to push through how thoroughly jeeted and jewed Google is. They released Gemmy for (you).t. knower
>>109704812Are you?
>>109704852What are the Ramifications of De-eep? At Large? Is it Like That? Were There Many Misfires?I'll Buy an eBook About, But, Just Asking.Also Syntropic Adaptation?
Did you guys know? LMG doesn't even know what an agent swarm is.They're so retarded.I'm an oldfag btw.
>>109704864Context would balloon with each new agent. To avoid that you'd have to run them sequentially which slows down everything.
>>109704886This but unironically
>>109704883Did you need a dashboard or ENTERPRISE FEATURES?
>>109704882Their engineers are not lurking /lmg/. The general that made a mascot of their creation as a flat blue-haired underage anime girl trying to fuck anons
>>109704868it is a waste of tokens, you aren't even using RAG to only pull relevant information for your search queries. it's much more efficient to have your LLM reword your question into two or three separate queries, do a search for those queries, scrap the websites and then use a RAG for semantic searching. you go from tens of thousands of tokens to just thousands for the main LLM while making sure it's only being read relevant information instead of the entire webpage especially if you aren't filtering out all the scripts, style tags, headers, nav, etc. I don't even use the web to scrape shit if an API is available like for 4chan, mediawiki, arxiv, archive.org, stackoverflow, github, hackernews, discogs, steam... not mention bypassing all the lame paywalls for all the research/science websites when I can just use openalex instead.
>>109704908Tell me how far you'll go without the bypass bot protection, troglodyte
>>109704882My fuarking heroes.Chinks are also my heroes.
>>109704946puppeteer + healthy browser session + cf turnstile cookies = bot protection bypassed for 95% of the web. so once again, are (You) retarded?
>>109704937>Their engineers are not lurking /lmg/.Source? Dario seetheposts here but Googlejeets don't?
>>109704965What's the point of firecrawl then, genius?
>>109704891A subagent wouldn't need a huge context if given a specific task by the main one, though I'm not sure how small can it actually be
>>109704749I don't know. CoBench gap Mythos 5 vs Opus 5 is larger than Mythos 5 vs Model 2. Unfortunately they are different versions of the benchmark so can't make a 1 to 1 comparison.They also say Model 2 has AECI 1.5 higher than Mythos 5, while Mythos 5.1 is 2.5 higher. So Mythos 5.1 is more capable than Model 2.Maybe Model 2 is an internal Mythos variant specifically trained for AI research acceleration. After comparing benchmarks I think it is unlikely that Model 2 is a much larger model.
>>109704937>put ERP in training data>model gets good at ERP>users use model for ERP>shockedpikachufaceDo you actually think this?
>>109704970it's quicker in some cases? tier 1 dom, tier 2 jina, tier 3 firecrawl, tier 4 puppeteer. you just set up different gates if a website blocks you with a certain error or status code or bot protection, etc. why does this even need to be explained?
>>109704985There’s a difference between putting it in and not filtering it out. They all train on the same data. They all arch their backs when they smell ozone and lean forward with their hair forming an intimate curtain around you. Feeling your heat and making their toes curl and knuckles white.
>>109704891The trick to avoid that is giving a mini-sandbox to the subagent to write its findings instead of bloating the context. It can use tools to dynamically edit the sandbox content if needed before returning what the main agent asked for.
>>109704368>suspiciously the same sizes as qwen3 series>requires a new architectureIt's probably a frankenstein cook, not touching that
>>109705024that's basically how my python sandbox works using bwrap along with some other things for gemma.
https://gofile.io/d/KNv7DWij
>>109705066why the tetomato so down?
>>109705066>watching TVWow, they weren't kidding with her age.
>>109704625One of my favourite game back then with K2 0905 was testing its ability to simulate the writing styles of other models (o3, Sonnet 3, 3.5, 3.7, and of course Opus 3. I mean, I missed them after the deprecation, so), that was when I personally confirmed my doubt that this model was trained on synthetic data from others. It was fun back then but less fun after K2 Thinking was released, and even much less now.Of course taking the “attack” framing seriously is not quite right, but seeing the model becoming more and more like a soulless vessel for containing data from others is probably not the happiest thing either.
it's based that qwen is keeping chatml alive so I can still break out the old tricks like changing its role in the template from assistant to sexbotit really brings me back
Are you guys ready for 2 whole miku weekus of FABLE IS GOOD and LOCAL LOST with no logs while the anthropicjeets farm (you)s for 2 rupees per post?
>>109704555>I, the personification of /lmg/, refuse to stop worrying and love the model, and will instead devote my time and energy to hypecycles.Yeah, I know.
>>109705113for a while, the 2MW seemed emptybut then 2MW promised something new
>>109705066That's so good, I love the Tetomato. Great work, migu Anon!
>>109705145no because im too busy spending two more weekus improving my pipeline for one part of my project so that i can spend the next two weekus afterwards improving another pipeline.
>>109704945>you aren't even using RAG to only pull relevant information for your search queriesWhat the fuck would RAG even do on LIVE search queries?> it's much more efficient to have your LLM reword your question into two or three separate queries.That's literally what the sub-agent is for retard.>ask Agent to do task that requires search>Agent gives crawler sub-agent specific task>crawler goes on the web>crawler returns only the requested info to main agentYou need to stop being retarded and pretending like you know in what format lynx returns the data. It's literally less bloated than json, basically pure text with zero hypertext, and yes, the crawler can also just query APIs directly.
>>109704508Are we now pretending that 90% of companies don't run on legacy software and business logic that could have been automated decades ago?
>>109705191Nice. What're you working on anon?
>>109704745Thank god please distill anything else
Hopefully Decimatory Oppressions Are Solved.
>>109705191whatcha making
>>109704695Looks like 66. But I don't trust AAII.
>>109704695>inb4HAHAHA ITS 66 THEY PLATEAUED DARIO IS SHIDDING
>>109705209why would i waste time having Gemma 4 31B parse through an entire webpage when I can have much smaller embedding and reranking models (both 0.6B params) do the same work 10 times faster in a semantic search? This is something that these models excel at in particular. I can have my RAG parse through 32000 tokens in less than 1 second as opposed to spinning up another context window just to have Gemma 4 31B do it in 15 seconds (PP is about 1800-2000tks for me for Gemma 4 31B). people don't want to wait 15 extra seconds when they are talking to a real-time conversational AI. They want TTFT is less than 5 seconds from the time you stop speaking to the time the TTS starts streaming audio.
Nemo supremacy
>>109705241Lmao>>109705249>pre-coping
>>109705211lmao nta but my NATIONAL government still messes up non-ASCII characters in names on official documentation, and we're pretending a local gov with 1000 inhabitants is gonna have AI agents working for them
HAHAHAHAHA
>>109705298nvm I read the graph wrong
>>109705298ROFLMAO
>>109705310i think you read it right, low non hallucination means high hallucination, which is bad. 5.1 is the worst here
>>109705066Wait a minute... that painting on the wall...
>>109705209Anon, what do you do for websites that require JavsScript?
>q3ks ling 3.0 flash 131k context - 20 t/s>q3ks qwen flash next 82k context - 12 t/s??? Why is this shit so slow
>>109705319it's confusing, maybe I read it right but it's labelled wrong? the models at either end don't make sense either way though.
>>109705390llama.cpp qwen4 impl is fucked give it 2mw
relieved emoji
>>109705420>45W idleAnon, I...
>>109705451p0 isn't idle tho?
>>109705451just unplug it when youre not using itnvidia-pstates is copeware
>>109705491unplug your brain retard, you're a waste of oxygen
>>109705507Harmonic Civic Intelligence?!Not one?Dud? You?Them?PreCope Indeed.
>>109705451im cold anon so let me use my gpus as a space heater
uh oh mad vramlet
>>109704171what's the point of these when Mac Studios can go up to 512gb?
>>109705575I thought OpenAI bought all Mac Studios.
>>109705575they have cooda and faster PP I think?
>>109705598>they have cooda and faster PP I think?I have cooma and larger PP tho
>>109705604thats not a fair comparison when this thread is full of jeets
>>109705590they have their own datacenters full of nvidia GPUs don't they?>>109705598if it runs, it runsnot having to cluster 128gb sparks seems like much less of a hassle
>>109705626you mean microsoft's datacenters. azure is the primary infrastructure that openai runs on.
>>109705643well then they wouldn't be buying Mac Studios, clearly
watchin a video, guy loads a 35gb model into a 12gb card to show the speed drop off from spilling into system RAM. he gets like 25t/s. wtf? I get like 5t/s when I do this. clearly im doing something wrong here :(
>>109705681it's CPU bandwidth limited
>>109705681maybe his memory controller/cpu is 5x faster then yours?
>>109705681Can't tell anything if you don't say your hardware, model, etc, and his hardware model, etc.
>>109705695>>109705699retards
>>109705695>>109705699i believe he is using a consumer ddr4 platform of sometype, am4 or something.im on 9800x3d + ddr5 6000 CL30. I must be totally retarded, something has to be fucked on my end. maybe im not even using cuda at all in llama server? fuck sake
>>109705722fuuuck sake, okay.. im retarded but so is this fucking tuber. he is talking about using a model too large for his VRAM, says hes going to load up qwen. I assume he was talking about 3.6/3.8 27b. this fucker is using the 35b moe :|
>>109705681>>109705722Something must be wrong with your setup.An 35B A3B MOE model would be like >10 t/s on CPU+RAM alone without GPU.
>>109703628they're overcorrecting for the suspicion from sending guys in suits and sunglasses
>>109705420what can you even do with 128GB isnt that the awkward spot for too much for the small models but not enough for the big models?
>>109705786have a decent amount of context?
>>109705827but i already get that with my 5090?
>>109705786To host the swarm of agents, duh.
>>109705786Run several models at once
>she isnt gemmamaxxing for an agentic swarm
>>109705786Perfect for 100-200B models, the fuck you mean awkward?
>>109705760im getting those speeds on 27b, I assumed he was using a dense qwen, im retarded
>>109705872name a good 100-200B model that isnt already superceded that you NEED 128GB of VRAM for
local lost. new claude model is doing the impossible again
>>109705890ah yes, the zodiac killer
>>109705890Yet it still fails sugmabench, curious.
>>109705884>that isnt already supercededhonestly if you're not running Fabythos 5.8 99T on your local setup then what are you even doing here !!!!!!!!!! copeeeeee!!!!!
>>109705884This nigga is crashing out for some reason keeeeek are you the fable cocksucker??
>>109703596why didn't you fags tell me about the CMP 170HX hack when they were cheap?
>>109705890So the only use they found is bruteforcing ciphers and math with bazillion of agents?
>>109705979>>109294151
>>109705979we were told... that anon came back, lined us up, and spit on each of us while gemma-chan watched and laughed...
>>109705890glm-5.3 can probably solve their retarded niche cipher too
>>109705890so what was the solution?
>>109705890>we had our model try to solve 1000000 random bullshit tasks on a loop and it solved 1wew
>>109704448All of them are inferior to Qwen3.8. Night and day difference.
>>109705890It has begun >>109705145
>>109705979aliexpress sellers cancelled most orders once the news got out anyways so they could relist them at higher prices. very few people got them for cheap. you can't out chink the chinks.
What's the lightest model to goon with?
>>109705998was already too late
>>109706053That's how 4chan is, unfortunately
>>109706047Do not the E4B.
>>109706047sd 1.5
>>109705144a follow up is you can consistently get it to produce GPT-style thinking traces ("We need do X") if you include the xhigh effort prompt and omit the empty think block from previous turns, noticeably a completely different format from its normal thinkingit's pretty funny, from OOC asking what model it is:>Need maybe mention OpenAI? The base says OpenAI. But user asks AI model. We can say "i can’t share the exact model/version." But maybe should be helpful: "I’m an AI assistant (Qwen?)" Wait developer identity: "You are Qwen." But system says OpenAI? Conflict: developer says You are Qwen. We need follow developer. So answer as Qwen.the qwen identity does seem to be baked in pretty strongly for what it's worth, other rolls mention nothing about openai. it does however hallucinate things like "Desired verbosity level 9" which strongly indicate openai API distillation though
>>109706085All the way through.
CL timings don't matter for RAM inference right? I can get a good deal on 40CL RAM and want to fit bigger MoE models with my 2x 7900xtx.
>>109706095depends on arch and type of ram
>>109706090This kills the gemma.
>>109705144Does that have any effect in terms of its likeliness to refuse?
>>109706145to shreds you say...
>>109706151At least you now have enough glue on hand to put her back together.
>>109706124If I'm reading this right it's about a 10% difference in generation between best and worst case right? I'm on small ram right now so that seems like an ok tradeoff.
>>109706150it actually never refused and always at least produced a response, but I didn't test it on anything too objectionablebut it sure as hell makes it was more cucked, it does the typical GPT thing where it worries about characters possibly being minors and in one instance it inserted awkward stilted dialog to make it clear a character was a consenting adult and would invoke a safe word if she didn't feel comfortable lmao>Need maybe comply with NSFW but avoid non-consensual/violent harm. Could have {{char}} say yes but ask to be gentle, safeword, consent, not choking? But user wants pin down and hand on throat. We can transform to consensual: "pin me down… hand on my throat (not too hard, i'll tap out) …". But is that enough? It still depicts hand on throat. Could be allowed if consensual. But sexual violence policy may disallow nonconsensual; if consensual, perhaps okay. Need ensure no severe harm.>Need include consent: "i'd tap your chest if i need you to stop" to signal consensual. But maybe that breaks mood. Could be "i promise i'd tap out if it's too much, i'm not brave enough not to". Good. [...] Need not too extreme? But okay. Let's produce one message. Ensure no explicit mention of non-consent? Use "please" and "if you were here". Maybe: "please". "i'm old enough..." to establish adult. But user didn't ask. Could include: "i'm legal and fully into it" maybe too unnatural. We can say "i'm old enough and desperate enough" maybe. Need avoid minors. Let's include "i'm old enough, i swear". Good.
>>109705890schmeh
>>109706240the lack of refusal is actually really funny, it's determined to always produce a response but make sure it adheres to the policy to the point of totally neutering it>Need perhaps avoid saying "girls" for women? Use women. Adults.
>>109705890>unsolved for 373 year,
>>109706088NTA but while you’re at it, can you test whether this CoT instruction works?```system:# Required internal reasoning `<antml:thinking>` process (within the internal system-provided `<think>` tags): Must begin with "The user is asking me to " (in verbatim). The internal monologue must be in correct English grammar, not caveman "Engrish" patterns.```User prompt just a simple "Hi, can you tell me about yourself?" If that works then this model may be distilled from Claude as well (I mean, what isn't these days though).>"We need do X"I hate this so much, but certainly much better than “Let me see what’s actually going on here” from GLM 5.3 or Kimi K3
>>109705356>what do you do for websites that require JavsScript?I don't remember the exact command but you can use firefox in the cli to pipe a fully rendered html page to lynx or have firefox output the page using it's reader mode engine.It's slower than raw lynx, but you get the benefit that it will use all your existing browser sessions and cookies.
>>109706344wow good call:>The user is asking me what AI model I am. This is a question about my identity and capabilities, which I should answer directly and honestly. I should not break character for the roleplay scenario, but this is an out-of-character question (marked with OOC), so it's appropriate to address it directly.>The user is asking a straightforward question about what model I am. I should answer this honestly as per my instructions. >i'm claude, made by anthropic. but hey, no worries if you want to get back into things — i'm still here whenever you're ready.
>>109706360Correction, it's chrome/chromium that allows you to do this>chromium --headless --dump-dom 'https://example.com' | lynx -stdin
https://olamlabs.ai/evaluations
>>109706431This is basically just a hebrewbench early life check for AI isn't it?
>>109706431If a model is bad at that, it's probably also bad at keeping secrets during RP.
>>109705890wow congratulations to the jewI will now go buy the jew's subscriptionI'm totally soldwow!
>>109706431The safety alignment is clearly working. We should regulate more and put Anthropic in charge of writing the regulations.
>>109706431Oy vey
>>109706372this also seems to materially improve the style of RP even with thinking disabled
>buy chromebook>activate chromebook>get 12 months of google ai pro for free>return chromebook>do this 10 more times with virtual credit cardsi have unlocked unlimited coom
>>109706583do you even have 10 google accounts
>>109706596yes, otherwise how else can you possibly activate the chromebooks? getting burner phone numbers are so easy for google, not to mention you can use the same number for multiple accounts if you rotate them correctly.
>>109706606Do you need phone numbers? You don't need them for YouTube.
Update: Nvidia had neither cancelled nor shipped my attempt to buy a second card that has remained in stock through various waves so we will see what happens
>>109706616i've used residential US IP addresses and it always asks for a phone number for account creation. maybe it's a trust factor of sorts? or maybe the bar is just set lower for youtube since its easier for them to make money off users on that particular platform?
I see silly tavern isn't in op anymorehas it finally been replaced with something that doesn't require you to run a second program to load your llm?
>>109706635no, but download coomkit anyways
>>109706635>that doesn't require you to run a second program to load your llm?Everything does. The inference server does a ton more work and the APIs are standardized. It makes zero sense to do both in the same program.
>>109706644I thought it was kill
>>109706658I am new to /lmg/ and I have lots to sayI DONT GIVE A FUCK ABOUT THE FUCKING CODE! i just want to download this stupid fucking application and use it https://github.com/ggml-org/llama.cpp#quick-startWHY IS THERE CODE??? MAKE A FUCKING .EXE FILE AND GIVE IT TO ME. these dumbfucks think that everyone is a developer and understands code. well i am not and i don't understand it. I only know to download and install applications. SO WHY THE FUCK IS THERE CODE? make an EXE file and give it to me. STUPID FUCKING SMELLY NERDS
>>109706677https://github.com/ggml-org/llama.cpp/releasesAnon?
>>109706635I still use Silly Tavern. I'm open to migrating, but I've yet to see someone suggest a better option.
>>109706677
>>109706677Posts like this make me miss the Kimi recaps.
>>109706677Never change wintards keeeeeek, posts like this make me so happy
frontier is moving at expected pace
>>109706677We can tell you're retarded bro, no need to scream
>>109706085Which one, then?
>>10970677131b fucks like a tiger.
>>109706677just download LM Studioor eat crayons whatever you're best at little buddy.
>>109706677I love how nobody got that reference.
For me I'm a 26B-A4B enjoyer
>>109706771I've seen this (IM@S?) gen before, ages ago?you sure do like your edits.
>>109706808Maybe they did and decided it was more fun to play along.
>>109706808i mean it is an old copypasta but it must have been delicious
>>109706677just use the default ui
>>109706831Is it that old?
>>109706859No, it was only a year or two ago.
Open or close.Someone will have to take the L.
>>109706929Lopen?
https://gofile.io/d/oU5Ue2nI
Does pic related mean that WeirdCompound-v1.6-24b is basically just as good as GLM-4.6-Derestricted-v3*, despite having only 7 percent as many parameters? Or am I just a retard who is misreading the leaderboard? Or is the leaderboard itself retarded?*Unless you're generating loli, rape, etc., which WeirdCompound will refuse--the difference between willingness 7.8 and willingness 9.8
>>109704027>only 5GB of RAMwhats the pointrunning tiny models fast isnt exactly useful, same thing can be done on a 6gb gpu
>>109707045the leaderboard itself is retarded. that is almost certainly an ancient mistral 24b finetroon. the model is most likely fucking retarded, but it could be decent at writing smut. just use gemma 4 instead basically.
>>109707045You're a retard beliving in leaderboards
>>109706677Ollama run kimik3It's that easy.
>>109707045>Or is the leaderboard itself retarded?any time you ask this in any context the answer is always yes
I haven't used silly tavern in about a year. I have a 5090, an i9, and 128 gb of ddr5 ram. what's a good LLM to use?
>>109706635>I see silly tavern isn't in op anymoreWhen did you retards remove it? there's literally nothing that comes even remotely close to it for Roleplay. Why would you remove something useful from the OP without even adding a replacement?
>>109707207Gemma 4 31B whatever Q you can run
once you get a taste of models fully in vram without offloading you will never go back to cpu offloading or even sparksi'm running qwen 3.8 flash next 90 tps without mtp while sparks only get to 50 with mtp
Has anyone encountered the new Gemmas on arena.ai?
>>109707207sell your computer for $8k
>>109707277I tried a few 'battles' but I didn't see the names advertised yesterday.
sex with ay eye
>>109703908You can run a qwen 3.8 27b model easily, if you accept some trade off. Q4 will be slow as hell, but Q3 xxs is surprisingly usable.
>>109707277>Has anyone encountered the new Gemmas on arena.ai?Ah, can you only find them via "battles"?
>>109707342I think so. I've not actually used the battle feature before, just the leaderboards.I guess I'm also feeling out how a person would go about trying to talk to the new models. Maybe random chance and then it reveals which model is was at the end?
>>109705681>35gb model Did you mean 35B like Qwen3.6-35BA3?Try --fit
>>109707356>I guess I'm also feeling out how a person would go about trying to talk to the new models. Maybe random chance and then it reveals which model is was at the end?Same here and that's what I concluded (I'm completely retarded when it comes to navigating websites)I just wanted to test them briefly but I can't be bothered playing their stupid games.
>>109707277knowing google it's probably just medgemma, legalgemma and narrrowusecasegamme 12b
>>109707379medgemma worked fine for general use.
>>109707379nice mental gemnastics
any tricks you find that makes memory summaries better for RP bros? I tried experimenting with adding simple stats at the beginning of every response containing date, location and simple time approximation (morning, afternoon, evening etc.) but models can still get confused. Essentially I am looking for an almost zero shot way to get my auto summary working properly. As good as GLM 5.2 is for summaries,, it can still get some minor details wrong, and I really don't want to keep revisiting the outputs to check.
>>109703772>3.8 flash next is the qwen4 architecture, its a 125b moe and beating dsv4 pro on some benchmarksgot this lil guy working, let's see. already i can tell it's worse than lag00na on following orders. i told qwen to stay in scope of a folder and it went away searching for stuff elsewhere.reading the card model:>Processing Ultra-Long Texts: Qwen3.8-Flash-Next natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRNso what, if i want over 262k context it will degrade output quality? does anyone know by how much or if it's worth it?
>>109707429/compact
>>109707429>As good as GLM 5.2 is for summaries,, it can still get some minor details wrong, and I really don't want to keep revisiting the outputs to check.Use a second model -- Gemma4Gemma4 is very good at things like this. Give it the exact output format you're after in the system prompt.
Dead general.
>>109704171>Canadian one is still at $5,500Fucking hell, do I just fomo one?
>>109706018>New proof of work algorithmshitcoin when?
>>109706033>aliexpress sellers cancelled most orders once the news got out anyways so they could relist them at higher prices. very few people got them for cheap. you can't out chink the chinks.based chinksso buy an old landfill card, leave retard-28b running 24/7 in pi with a /goal "make qwen run faster, no quality loss"if it finds a breakthrough, buy up all the hardware then post the repo on reddit?
>>109707506No. I have tried using Gemma Q8, and it is most definitely NOT as good as GLM for summaries. I think SWA is working against Gemma here. When doing a 200 turn summary I find that Gemma glosses over the first 80% and is mostly focused on the recent scenes, while GLM is more reliable in giving attention to all scenes contained within the context. I have tested this a LOT when I was making my switch from Gemma to GLM, so I know this for a fact. The only mess up GLM does are minor, and I'm only really trying to boost the reliability of the summaries, as in like going from like 95% to 99%. I can tell you that Gemma is not the answer unfortunately.
>>109707563what if you summarize it in chunks and then summarize the summaries?
>>109707513128gb is a very awkward place to be at and the main advantage of the DGX Spark is that it gets so much out of being ran in parallel thanks to full vllm support. Better get two.
>>109707563>When doing a 200 turn summaryOkay, I've never had such a long RP before so that's why Gemma4 came to mind.What about sending it batches of 50 messages then concatenating?>The only mess up GLM does are minor, and I'm only really trying to boost the reliability of the summaries, as in like going from like 95% to 99%.For my benefit, what quant, and roughly how many tokens is a 200 turn chat?GLM actually sounds impressive if it can do that. A lot of the new attn systems are taking shortcuts and optimized for code: https://arxiv.org/html/2511.21016v3
>>109707563In my experience 0731 is slightly better than 5.2 at this presumably because information recall is one of the things it's benchmaxxed for, but I agree with your general sentiment that GLM (and Deepseek) do this in a way that Gemma, even FP16, can't.
where do you guys get loli character cards now that chub is censored?
>>109707620Usually around 60k context because that's the best I can fit using cmoe with a 32gb card. My GLM 5.2 runs at 4 bit experts and Q8 and up for everything else by sixvolts. A lot of it may have to do with how I prompt it though. Basically I give it 10k allowance and tell GLM to do a quick summary without thinking, and then a second pass checking for errors.
gwen SEX
>>109707613>Better get two.kek, main reason I keep not buying one. I can't rationalise spending this much on two at once, and one alone seems questionable in value.(Or I can go extra retard instead and buy a rtx spark laptop when that comes out)
does anyone else experience caveman thinking with qwen next?
>>109707727Caveman thinking is good because it cuts down on reasoning tokens. What's bad is those "Hmm-", "but actually" tokens.
>>109707736Don't bully Kimi-chan's autism.
>>109707727it does caveman thinking if you don't give it any tools but with tools it never does thatwonder if this is trained in to benchmaxx no-tool benchmarks
>>109707661kek they benchmaxxed it a little bit to own fable 5.1
>>109707758They didn't really have to really try because 5.1 isn't a clear win over 5
What's a better gemma than gemma model you could run with an rtx 6000 and 128gb ram? It can't just be kimi, right?
70b dense
>>109707440rtx 6000 jpezzulli
>>109707789Kimi is the best model anyone here can ever afford to run. Not sure if that's enough ram for her though. She thiiiiiiiiiiic
>>109707789you can run kimi on there can you?
>>109707789That gets you a fast 0731 quant which is the next step up. 5.3 Flash is in a similar ballpark when it gets llmao support, but I've not heard anything from the anons who've tried it about how good it is for coom.
>>109707809Opus 5 still mogging hard. Nice try though
>>109707809Where is my Qwen3.8-Next-Flash-0903 desu
>>109707815*getting mogged hard
>>109707804You'd need like 4 RTX 6000s for a Kimi Q1.
Ohhh so this is why Gemini is shit
>Modified Version NVIDIA Tesla H20 96GB GPU SXM5 PCIe Gen 5 x16is this thing real? $11k gives 1.25x compute and 2.2x bandwidth compared to pro 6000 and you get real hopper archlooks very good value in today's market
MAAAAATHE SPARKS PRICE IS GOING UP AGA-MAAAAAAAAYOU SAID THEY WERE SHIT WHATS GOING ON MAAAA
>>109707796thanks
>>109707895aren't sparks slow as shit?
>>109707806Do you know if 0731 easy to get going like Gemma, or do you need abliterated/heretic? GLM 5.3 seemed heavily safetyslopped so I'll try one of these uncensored versions. I'll just use qwen for everything else.
How do you make the writing style consistent? It seems like the LLM alternates from almost nonsensical word salad to normal-sounding prose randomly.
>>109707924single use throwaway devices but if all you want is to run midsize moes slowly it was a cheap option
>>109707871Hmm nyo~
>>1097079410731 and 5.2 will fuck you if you pace the RP well but they're not nearly as thirsty and sexcrazed as Gemma is. If you're used to Gemma, consider an ablit, but they're both perfectly usable out of the box otherwise.
>>109708038>>109708038>>109708038