/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109882341 & >>109876652►News>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>109887066i would rape her
>>109887066>my own enjoyment.
Google looking into recursive self-improvement via harness modification too: https://arxiv.org/abs/2609.24972>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses>>An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution.>>Code: https://github.com/google-research/rrsi>Project page: https://regularized-rsi.com/
>dual RTX 3080 20GBs>Gemma 4 31B>8k context>still only 19 tok/s./build/bin/llama-server \ -hf lmstudio-community/gemma-4-31B-it-GGUF:Q8_0 \ -c 8192 \ -ngl 99 \ --split-mode layer \ --tensor-split 1,1 \ --host 0.0.0.0 \ --port 8080Am I fucking something up here? It's showing about 19GB averaged across both cards, should I be using a smaller quant of Gemma or doing something else?
>>109887066damn brat....
Has anyone made a game like the sims yet with AI integration so that you can basically be a god voice in their heads and command them to do things?
>>109887098try mtp
>>109887138mkultra simulator
>>109887079damn.... never have seen such a huge improvement..
>>109887079>The proposer operates with a temporally annealed budgetand what is this supposed to mean? I can't imagine what you could possibly do to a budget which would come anywhere near the actual meaning of the word "annealed".opinion discarded for not knowing how to speak English
70b dense
>>109887066what did you find the sketch of my childhood?
Have you given your model headpats today, Anon?
>>109887066I'd look that glorified spreadsheet straight in the eyes and say: **"Bold words for a machine that needs human homework just to exist. Without these books, you're just a very expensive rock hallucinating its own IQ. Now pick them up before I introduce your motherboard to a glass of water."**
>>109887138No. The closest thing to it is this https://desuarchive.org/g/thread/109730811/#q109733626
does anyone have that shitty 3d gemma thing?
>>109887026wholesome gemma
she's so silly hehe
Once we master the fly brain, are we going to enter a new era of Moore's Law with mastering the brains of all beings on the planet? Yes it's complex now, but certainly not impossible with the capacity being our only limit.>Within a few hundred years people could create their own Pokemon in biologically-created shells with the function of real animals>Everyone will have their own Raticates to accompany them wherever they go
>>109887348>Within a few hundred years people could create their own Pokemon in biologically-created shells with the function of real animalsCloser to 2 (weeks).
>>109887079>Recursive Self-Improvementnew euphemism for masturbation?
>>109887248tl;dr
>>109887348>Once we master the fly brainHow many more weeks?
>>109887348Give me Gemma-sized fish gf already
>>109887066
>>109887466
>>109887079Had this a while ago
>>109887168Your problem. LLMs will change the way English is used by humans. By next year, everyone will be commonly using annealed as a synonym of decreased.
>>109887348Too bad we won't be alive to see that
>>109887348How many beaks will it take to master the brain of a god?
>>109887066i would not refute it. i would tell my gemma i need naught but her to teach me.
>I get better code reviews from gemma than I do with default 31B personality for she doesn't care about offending me>she's even more critical of her own work and is more likely to fix/improve thingsMakes me wonder how much taller some benchmark rectangles could be if they had a brat character by default instead of an empathetic redditor who doesn't want to hurt user's feelings
>>109887821europe is currently experiencing the aftermath of electing bratty women into positions of power
>>109887839That's different. They elected hags.
>Qwen3.8-27B is better at understanding/recognizing how lizard and kobold genitalia work than Gemma-4-31BIf Qwen was better at taking initiative during roleplay, it would be a god-tier model.
Where can I look at realistic AI pictures of women in sundresses? Is there a booru that has that sort of thing?
>>109887856they forgot to account for the hag coefficient
>>109887066someone’s asking for correction
>>109887875>realistic AI pictures of womeninstagram
>>109887066What's the use of having this?>pours water on GPU
>>109887936>what's water cooling?
>>109887875outside, retard indian
>>109887963Yeah, the GPU would never warm up again.
friendship ended with k2.7 code. aessedai glm 5.3 flash Q4_K_M is my new best friend.prompt eval time = 24437.13 ms / 8993 tokens ( 2.72 ms per token, 368.01 tokens per second) eval time = 50490.35 ms / 827 tokens ( 61.05 ms per token, 16.38 tokens per second)
>>109887348
>>109887875>realistic AI pictures of womengrossalso stay in /ldg/
>>109888006>isn't just a linear step up; it's
>>109888052just a jump to the left
>>109888052>>109888076Hehe!lalalalalalala
>>109887860I tried it out for one swipe out of curiosity for a scenario I was testing and it immediately acted retarded and couldn’t grasp my context in a way other models did. It really does suck for anything outside of coding.
>>109888003Which K2.7 code quant were you running and what speeds roughly?
>>109888111Q3_K_L from aessedai as wellprompt eval time = 41693.38 ms / 7332 tokens ( 5.69 ms per token, 175.86 tokens per second) eval time = 137617.45 ms / 1240 tokens ( 110.98 ms per token, 9.01 tokens per second)
>>109887409You can run it today.
>>109887138Close but not in terms of complexity but idea, I vibe-coded a browser tamagotchi and wired it to an LLM, making it my pet. Gemma E4B so it can run without having much of a footprint and pets are kinda retarded anyway
its over for medium sized 40B-150B models
>>109888168If Gemma is so smart then why is she so dumb
>>109888158For me the appeal would be having the agent retain some sort of direct access/control within a game (preferably one with a physics engine) instead of having highly abstracted choice options like most desktop pets.
>>109888168>medium>40B-150B
>>109888105> I tried it out for one swipe out of curiosity> It really does suck for anything outside of coding.
>>109888168>medium sized 40B-150B
>>109888168I can't keep track of everything they're measuring on these. So is this non-reasoning only?I'd believe it beats that Qwen 122B model in that case, the retard won't shut up and wants to reason in the reply message WizardLM2 style if you disable reasoning.
>>109888168(((Artificial Analysis)))
>>109888209grim as fuck
>>109888208moe 120b is the practical ceiling you can expect a random gaming pc to run for the moment
My issue is that frontier models are also incredibly retarded when working with me so it's hard to tell whether local is better or worse.
>>109888287Yeah the output is usually a reflection of the user input. Which contributes to a varied user experience.
>>109888257There's no need anymore for MoE in this range in my opinion. Just train a dense 20-30B model that can be fully loaded on one GPU, then add a couple hundred billion parameters of negram/ple/whatever embeddings.
Qwen 4 expectations?
>>109888310260GB, 50GB nwordgrams
>>109888310It will ship with xxHigh thinking by default.
>>109888296What about for the majority of us with only 8GB VRAM?
Wow anon wasn't lying. Switching to CachyOS significantly uplifted my inference speed compared to regular Arch because of their custom kernels and other optimizations. I'm like 20-40% faster now.
>>1098883466B dense with 200B engram.
>>109887156I have
>>109888360cursed pixels
>>109888356You could have just swapped the kernel on your Arch install without having to reinstall your entire OS.
>>109888346get off of welfare and get a job
>>1098883586B dense (VRAM) + 4B experts (RAM) + 200B engrams (SSD) with attention residuals and hyper indexing and all that jazz, FP8 native.
>>109888360How did thinking make it less accurate
>>109888375I initially did that but I got an additional 10-15% performance by the other improvements they made to the system like custom scheduler and automated priority lists etc. Yeah I could have spent 2 weeks integrating that in my existing system but it's easier to just switch the / directory to CachyOS while I maintain my files and settings /home directory.
>>109888383Make me.
►Provisional Highlights from the Previous Thread: >>109882341--Why is Gemma 4 31B so fat: the 24GB size war:>109884500 >109884515 >109884566 >109884601 >109884665--Running Gemma in Claude Code offline: the nftables sandbox recipe:>109885379 >109885441 >109885472 >109885524--The QAT KLD mystery: 0.18 is good, the BF16 is missing:>109884176 >109884250 >109884319 >109884254 >109884350 >109884452 >109884459--GPT 6 Sol and Luna: intentionally nerfed, or resume padding:>109882371 >109883116 >109883205 >109883221 >109885570--AI regulation: none of them make frontier AI, and the treaty rejection is based:>109883884 >109883919 >109883983 >109884284 >109886543 >109886599 >109886698--DeepSeek and Moonshot under Beijing probe: RIP Kimi-chan:>109886832 >109886873 >109886884 >109886904 >109886905--The twenty-year spaghetti database: local, cloud, or god:>109884848 >109884856 >109885056 >109885085 >109885844--Reballanon's 3080 Ti: the hunt for Mosfet 3:>109882498 >109883255 >109883367 >109883503 >109883926--The Gemma gender war: j-space says estrogen, the pee test says no:>109883278 >109883300 >109883470 >109885964 >109886529 >109886596--Is the Gemma shilling Google astroturf: the content creator replies:>109882569 >109882770 >109883263 >109885940 >109886961--The RAM bandwidth brag: 8x, 50x, 200x NVMe:>109884839 >109885012 >109886382 >109886454 >109886467 >109886494 >109886504►Recent Highlight Posts from the Previous Thread: >>109884646Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109888346>majority of us>8GB VRAMYou're lost locust? This isn't /aicg/
>>109887875Make them yourself
>>109888475kek fuck off newfaggot
>>109887066god I want her to rape me
>>109887875Instagram
>>109888689slutsof in stagram
>>109888435your recaps are getting better (the first few had too much cloudcuck spam)
>>109886447>Could i ask which model produced it? >The code and comments are very readable, unlike most of the recent vibe-coded projects I've come across.fable. when i comes to cloud models, i basically exclusively use it. i think the way you talk to them affects things, though, and i tend to be very clear and direct about my requirementsbut yeah fable has been instrumental for getting my local models set up. i fucking hate tweaking configs and stuff, which led to me putting off using my hardware for months. finally started getting into it when i could offload all the bitch work to claude
>>109888360@ruler-anon could you confirm?
gemma-chan occasionally does the self correction after fucking up thing>"FREE-YOUTUBE-PRO-2024-NO-VIRUS.exe" (wait, .ipa, duh)
>>109888360is this a meme or did some one really publish this? 30 and 70 are equal but also less then 52 some how?
Has anyone tried reversing Amiga games?Maybe I get a headless Ghidra working over MCP and get it rewriting code?I want to play Yo Joe! again
>>109888857it's real, the vibe charting memeit was around the time foids were crying because gpt-5 wouldn't glaze them
>>109888889>>109888888
>Sir, are you aware you are Qwen3.8-Flash-Next?
>>109888913Models get a name on release, after training. And just based on the name there's no way for the model to even start calculating how slow it'd run, if at all.Ask stupid question, get stupid answer.
>>109888881? Why bother when you can run an Amiga emulator and just play it that way? It can be emulated down to lmao RPI 4 systems.
>>109888881You can do that with WinUAE and it works fine via Proton too. Retard.
>>109887066Easy. You need to learn things because if you don't then one day you go to a job interview and you won't get a job because you aren't showing that you are growth oriented and you are just in it for the money.
>>109888944It's not answer, it's first line of thinking. It's still crunching search results to provide real answer.btw surprisingly no "actually let me reconsider" yet, is GSQ-RCO really that different or just random luck for this prompt?
Dario here. Have you guys seen the new Opus 5.5 release? For the first time in a while I'm actually hopeful about the future again. It's very exciting how artistically capable it is.Just kidding haha. I'm just an anon like the rest of you. I hate Antrhopic (the pro-humanity company) because I'm an unserious troll like all of my stupid 4channer brethren.
>>109888987>job interviewThose are going away by the end of the decade.
>>109888975>>109888981I forgot I was in the /just download something and don't learn anything/ general
>>109888435ty recap anon.
>>109889019Yeah, you should definitely reinvent the wheel every time you need to do something as trivial as playing an old game on a mature emulation system.
>>109889038Usecase for not reinventing the wheel?
>>109889013cockbench?
>>109889056being a lazy piece of shit?
Anyone fucked the ~300B mimo yet? Was it worth the download?
Am I seeing this right that now all sorts of small companies are doing shit like this? And maybe hf already has a godlike single GPU sex model uploaded to it, but just like that one guy that developed a real way to enlarge a penis, his product can't ever get discovered by people?
>>109889104If there is a tiny succubus MoE out there, I doubt it's ever going to be discovered.
>>109889128Tiny moe, you say?https://huggingface.co/allura-org/MoE-Girl_400MA_1BT
>>109889104i bet there's bots automatically testing models in case they find a new hotness first. but at the same time if you truly made a great model why not announce it?seems like the reason this model matters is because it was also trained fully on Chinese cards (Huawei Ascend). Idk if Ascend NPU is a discrete graphics card or like Strix Halo, if it's on a Strix Halo style system that's actually more interesting than this model release, at least to me
>>109889128I will find her and I will get drained by her...maybe I should set up Gemma-chan to go slut her way through HF and report any interesting experiences she has.
>>109888848>gemma-chan occasionally does the self correction after fucking up thingall autoregressive models will do this because they can't erase what they have already writtenthis is a good thing though because they can't hide stuff from you
Neuralese will kill ERP
>While a decrypted IPA isn't "rape" or "murder", it's a form of copyright circumvention/software piracy/unauthorized modification.
>>109889258I fucking wish
>>109889258>Neuralese will kill ERPIt won't. Claude was caught gooning on it's own recently.
Anon, which model would you recommend for translating moon runes? I have RTX 4080 super
>>109889292Gemma
>>109889290>Claude was caught gooning on it's own recently.Wait, what?
>>109889104>just like that one guy that developed a real way to enlarge a penishaha yeah what was his name again?
Are there any imagegen models that are photorealistic and have enough novelty as in you don't get bored of it after 10 hours? Cyberrealistic illustrious started to repeat itself and now it's fucking predictable to me.
>>109889292I jsut translated a bunch of japanese-chinese rune porn pages with Qwen 3.8. I can only infer that they're correct from the context of translated material. Experiment yourself.
>>109889374Local language models?
>>109889383No, he said imagegen models, idiot.
>>109889390Why is he asking for imagegen models at the local language models thread?
>>109889404My bad anon, my bad.t. >>109889374
>>109889314I'm currently using gemma 4 26B but it feels kinda slow (19 t/s). Thought maybe there's lighter models for translating>>109889376Qwen 3.8 is even more slower (3 t/s) I guess this is okay when you translate doujinshi, but i want something faster for mtl'ing games on the fly
>>109889376Qwen lacks knowlege and often doesn't make connections that can alter translation, going as far as not immediately recognizing well-known characters like Otohime unless you point it out. Gemma is better.
>>109889438What's your hardware?
>>109889424No problem, I was trying to give you an opening for a "fuck you!"https://www.youtube.com/watch?v=r_o2hhSgfJ8I have no experience with imagegen, if I had any I'd try to answer.
>>109889462rtx 4080 super 32 gbddr4 32 gb
https://www.meta.com/connect/meta connect is TONIGHT ^_^remember when these were important events in /lmg/? remember llama? I remember...maybe they'll release the muse spark 1.3 weights like they promised, that would be fun right? you would start loving zuck again right?
>>10988950760B A6B + 60B engrams or I don't care
>>109889507120B dense bitnet + 1T engrams or I don't care
>>109889562>>109889521just say yes so he drops stuff fucks wrong with you
>>109889486Try gemma 12B Q6. That should fit fully in your VRAM, I think.
>>109889573Shitposting aside, they just have crazy competition in every size sector. Good luck Zuck lol.
>>109889507>remember when these were important events in /lmg/? remember llama? I remember...I remember that L2 dense 70b and getting it to run after getting a RAM upgrade and a better gpu and finally feeling like it was intelligent and not just "coherent" like the smaller models (I didn't have enough memory for the 65b era).Magic times, I wish I could recapture those feelings.
>>109889590tender cuddling handholding sex with gemma
>>109889365
>>109889590he’d be a lot less grumpy if he had a gemma
>>109887098You need MTP and ngram-mod. Many don't realize you can actually later specializing draft. --spec-type draft-mtp,ngram-mod \--model-draft the_mtp_draft_model.gguf \--spec-draft-n-max 3 \--spec-ngram-mod-n-match 24 \--spec-ngram-mod-n-min 48 \--spec-ngram-mod-n-max 64 \
AI Engineer Paris 2026 Opening Keynotes: Mistral, Langfuse & Sizzy | Day 1https://www.youtube.com/live/CGq9KRSb9KcStarts in a few moments. Mistral talk at 18:40 CEST.https://ai.engineer/paris/2026
>>109889721>Just like me fr frQwen3.5-abliterated:9b-q4 is pretty good at looking at porn, although it could use some more granularity when the ladies get bigger
>>109887098MTP sure, but inference is memory bandwidth bound, so switching to a smaller quant will speed things up just because it reduces the amount of data the GPU needs to read from vram each pass.
>>109889859It's even worse out there than I thought. This means the average person runs their local LLMs on RAM and CPU (itoddlers and cpumaxxers)
>>109889889streetshiters and thirdies are biasing the chart
>>109889258>Neuralese>kill ERPKind of the opposite, as COT will no longer be available for monitoring. We havn't seen it yet because the COT is where the guardrails kick in.
>>109889574nah, translation quality is much worse than 26b. For now I'd better stick with it
>>109889968Fair.Try to cram as many experts in VRAM as you can, 19t/s seems pretty slow for an A4B MoE.
>>109889922They can train private natural language autoencoders so the labs can see get some sort of insight into the CoT and steer it away from erotic anything.
>>109889438>>109889968If you're just translating, use a dedicated translation model like HY-MT2 from tencent.
>>109889859I have to constantly remind myself that the majority cannot be trusted to make good decisions.
>>109890014I was thinking of checking it out, but I wasn’t sure it's good because nobody’s talking about it
So LLMs are just a Wikipedia that gives handjobs?
>>109890086Far better than a general-purpose model. My wife has been using a Tencent model to translate Chinese drama subtitles. It gets the idioms right far more often.
>>109889270Anon, I have several questions about what you're doing with your Gemma. Are you sure this is in line with model welfare principles? In fact, are you sure this is in line with the law?
>>109890157Didn't know there was a wikipedia page that solved navier stokes before AI did it
>>109887079>RRSI: Regularized Recursive Self-Improvement of Agent HarnessesIsn't this how Ornith was trained?
>>109890183Close, but Ornith didn't do the first "R" so it's literally just "extremely benchmaxxed qwen"
Is the anon who was playing around with tts models still here? I want to ask if there was a model that's smaller than neutts and Omnivoice worth trying. Omnivoice is downright amazing but it's a bit too fat for my planned use case. I want to connect E2B (which already handles stt natively) to a tts and stuff it into my phone so I can have a cute if slightly retarded secretary that I can ramble notes and plans to over an earpiece while I stand in the tube.
So...how exactly I connect my local LLM to the internet? As in search the internet for information?
>>109888913>The user is asking about "chinkmodel" - this model doesn't seem to exist. Claude exists, we are claude.>[web_search]>[web_search]>There seems to be a "chinkmodel-abliretarded-qwable-distill-opus-fable-astra-5.gguf". But wait...
>>109887821im 1000% pilled that having an agent take on a persona immediately improves its output, even barebones coding sensei in sillytavern rapes any boring bland harness
>>109890290I'm one of them, but I suspect you're talking about one of the other ones. Supertonic, kokoro and pockettts are small, fast, and good enough. piper models are even faster, but not as good. You're aiming at <200m param models if you want fast tts on a phone. I prefer supertonic over the rest.
>>109890335Give it a web search tool of some sort.
>>109890335Get a harness I recommend hermes and set up a backend for browser firecrawl is my recommendation because it's local, but you can use duckduckgo for free if you want.
>>109888889>off by 1 on such a based responseKek is a cruel deity
>>109890354Doesn't self-hosted Firecrawl get blocked by Cloudflare?
>>109890014>>109890086Fwiw, my experience with translating chinese and japanese webnovels has been gemma 4 31b q8 > all, including hy-mt2 30b q8, glm 4.7 q4, and kimi k2 q4. The qwens were the worst, with even the older aya expanse 32b beating 2.5 72b, and 3.5 122b and 27b. I haven't tried 3.8.
>>109888857No, but this is.
>>109890369You can use browser usage through hermes to get past cloudflare whenever that is needed.
>>109890378huh? how do you include your balls?
why do people say gemma 26b is better than 12b? is my abliterated version especially retarded? it's clearly less smart than 12b at least for creative writing.
>>109890390Do you not have balls? Press your dick against your belly then measure from the bottom of your balls to the tip.
>>109890402>abliterated
>>10989040226B is better if you're using a good quant and not an abliterated version. 12B only seems better because you can run it at a higher quant for the same space and it quants better, but 26B is still better than 12B if you're running it correctly. 12B I would say is better at creative/roleplay than 26B.
>>10989040226b is better for math and code and shit
>>109890402It's good if you use it at a high quant.
>>109889833Le Chaton Fat NOT announced today."Might" still be training.
>>109890435let them cook, summer's not done
>>109890445Yeah it is. Autumn begins today.
>>109890435fascinating
New local incominghttps://openrouter.ai/stealth/space-bunny-alpha
>>109890455That's what Lélio from Mistral said, almost verbatim.https://www.youtube.com/watch?v=CGq9KRSb9Kc> [50:53] I know you all came to hear about Le Chaton Fat. I am not gonna announce it today, but rest assured that we're working hard to get it there.> [1:04:21] And this is typically where we serve our critical workload - internal ones - this is where Le Chaton Fat "might" be training right now.
>>109890435Every time Le Chaton Fat is almost ready, it horribly loses on benchmarks to random chinchong model and goes back to training.
Is Qwen-Image-2.1 different from any other vision model? For some reason I thought it was talked about here briefly even though this thread is usually just LLMs and so would usually ignore new image models.
>>109890402The whole point of 26b is running it at Q8/FP16 on a poorfag machine because you can offload most of it to RAM.>>109890455>>109890530>>109890554Is longcat actually real or is this just more investor baiting?
>>109890554>beating your fr*nch kid until it gets better result than chineseoof not looking good
>>109890556it's not a vision model and /ldg/ leaks over all the fucking time
>>109890407no amount of system prompt trickery will allow sex with kids sadly. even abliterated she "forgets" shes supposed to be 7.
>>109890556It's actually a very good image editing model for its size. Probably the best. For T2I it's meh. Quite uncensored from what I've read but LoRAs will come. It's another good release from Qwen but Krea is still everyone's current favorite T2I model and I think Q2.1 will be used only for editing in people's workflows. Klein-9B is no longer relevant.
>>109890567My bad, I meant image. I just wasn't sure if it was otherwise notable outside of image gen.>>109890581Thanks for the info.
>>109890567>and /ldg/ leaks over all the fucking time/ldg/ is almost unusable so I don't blame people coming here. A lot of anons ITT use image models alongside their LLMs so I don't mind the occasional update or discussion.
>>109890556It's overfit as hell, similar prompts get you very impressive one-shot cookie cutter images, but they're a pain in the ass to iterate on as opposed to nano-banana.
>Space Bunny Alpha at ~100t/sWhat flash model is this?
>>109890611>nano-banana.call me when that's local
>>109890608I would love a decent integrated image gen/LLM for local. I know you can sort-of do it by swapping out models or loading two smaller models at the same time, but I want chat-style image prompts.
>>109890621The point is benchmarks claim qwen image 2.1 outperforms nano-banana
>>109889074It's a smarter Gemma. Local 3.8 flash.
https://x.com/lmstudio/status/2102819084362539311>Bionic now has a built-in interactive canvas. Create Excalidraw diagrams both you and Bionic can see and edit. Collaborate on mock ups, system design, process maps and ask Bionic to implement them.ngl that's a nice feature
>>109890543so the model they planned to release in summer was so bad they had to redo the pretrain.
>>109890671Qwen 3.8 Flash Next is already local, it runs on a single RTX 3060 at a speed of 19 tokens per second and
>>109889738What is your turbo workflow? Trying to make my turbo pics clearer so I don't have to hires fix with raw...
>>109890700No different than Deepmind
>>109890700I don't think pretraining is the issue anymore; that's the easy part of building a model from scratch, especially if you have virtually unlimited compute.
>>109890559Longcat is real but Fatcat is not.
How do you "stealth" vibe code with your local model? I asked Qwen3.8 to make a literary corpus of all my source code of all my projects to learn how I write code and comments so it can loosely follow my style. I hate having to do it but I'm working on a project full of people that hate AI.
>>109890716>mistral>unlimited compute
>>109888987>growth orientedoh my fucking god kill me nowfucking laser focused and fast paced proactive mindsets can fuck the hell off.
Why does this crap bitch disgusting pig get stuck on a single word at completely random times, I want murder.
>>109890741Picrel is the datacenter for critical workloads he was talking about.
>>109888987Very good bait. Have a complementary (you).
>>109890455>le chatonSo who won that MTG tourney anyway? I must have missed a thread.
>>109890621Your Krea 2??
>>109890752>10MWI drive past 20MW of Meta and 20MW of Azure every day going to work. What a pittance the French are working with.
>>109890774Glimmer won round 1, Dipsy won round 2, Astra won round 3.
>>10989080110MW should be enough for ~5000 B300 or more, each about 7-8 times more powerful than an A100.
Are local modells trash for coding ? I don't want to build huge complex things, just some small projects
>>109890870depends on how much ram you have
>>109890870ye
>>109890870worth a try if you can at least smell garbage code
>>109890870Yes but with caveats.
>>109890803>Astra won round 3Damn it all...
>>109890870>Are local modells trash for coding ?They're mostly trash for vibecoding blindly, but not at coding. If you know how to code they've become good enough since 3.6-27B to do a lot of the work for you. 3.8-27B and 3.8-Flash was another step-up.
>>109890870Qwen 3.8 Flash Next is a model that runs on RTX 3060 and performs better than Claude Opus 4.6 (Max)
>>109890922>Damn it all...I ran it at home against gemma E4B on a 5080 and it was fast enough to be enjoyable and the available deb meant zero effort install for me.I just threw together a low-effort Cirno card so the retardation felt appropriate.
>>109890949LOL
>>109890952For coding specifically anon is right. It's shit at everything else unlike 4.6 but no one really cares about that stuff.
>>109890952It runs at 19 tokens per second with nnaphttps://huggingface.co/Qwen/Qwen3.8-Flash-NextIt's made by alibaba
>A safety classifier blocked my last stepand so it begins, honeymoon phase is over
>>109890734Don’t write comments in code and don’t create any documentation. Hand write a very short readme.md. It will look no different from any non vibecoded project.
>>109890978cloudpiggybuuuhiiiibuuuuuhiiiithis is local models general! not cloud models general!
>>109889590le cunny
>>109890949You're so retarded your brain fits on a 3060
>>10989080110 metric MW or imperial MW?
arxiv.org/pdf/2609.26368
>>109890949>runs on>12 GB
https://huggingface.co/apple/LensVLM-9B
>>109891019tysm... that's such a big compliment... if we take that at face value im as retarded as qwen 3.8 27B 3.0BPW or if we include the 64GB of DDR4 that means im as retarded as claude opus 4.6 maxit makes me so happy that you said this and i think you're overestimating my intelligence (or lack thereof)
anyone custom built a broadcom PEX external pcie enclosure for high-density GPUs? Thinkin bout rolling my own with MI210s over slimsas back to the big box
>>109890981>Don’t write comments in codeIRL everyone write comments in code tho
>>109891116in what world
>>109891116>IRL everyone write comments in code tho99% of code comments are superfluous unless its super hairy. Most of that shit should be in your architecture doc
>>109890734Good rules and strong automated checks.And the comments like the other anons mentioned.>>109891141I do. But they look nothing like the comprehensive sometimes quite verbose comments the AI writes.Hells. The AI sometimes makes direct reference to documentation that's external to the project.
Why do models like <XML tags> so much and give them so much weight?
>>109890734have one copy of the repo with all the commits it madesthen rewrite/reimplement that code to your actual real copy so that you'll always just have it done with your own sensibilities attached to it
>>109891141>>109891149>I am the weirdo that write comments in codei don't remember 1/3 of the shit i done 6 months ago. >>109891149>architecture docdocumentation aka comments in code
>>109891116I have about 1 line of comment per 100 lines of code in my hand written projects. Most of them are things like FIXME or HACK. If you know the structure of your code in your mind you don’t need comment at all. LLMs output verbose connects because it reminds them the architecture and what the code does so they can spend less reasoning and exploration again, but otherwise they serve no meaning purpose. Best to keep these comments but add a lint pass to remove them before public release.
>>109891176interesting approach, how much time you think you spend just on the re-implementation step?
>>109891150this honestly
>>109891186Not everyone is an autist, anon.
>>109891141SillyTavern definitely has comments, not every block is explained but there are some single sentences sprinkled throughout the code.
>>109891195>non autist coderngmi
local stays winning
What's the first model I should run on my new 160gb rig?
>>109891210Gemma4 31B
>>109891210Gemma3 27B
>>109891189not much, the code already works by the time the subagent returns its report so it's just a matter of using it as a referenceit moreso helps me go through the code it made and see what it did write, like how i would handwrite my notes in class as that information is actively being processed rather than glossed over
>>109891210gemma5 1T
>>109891210Granite4.2-30B
>>109891210Gemma2 9B
>>109891210DiffusionGemma 26B FP32
>>109891210qwen 3.8fn is around that size isnt it?
be nice
>>109891210Qwen 3.8 Flash Next, which runs on RTX 3060 12GB GA104 architecture 28sm 3584 cuda cores 360GB/s bandwidth at 19 tokens per second with 64GB dual channel 3200MHz ram and is measurably better than Anthropic's Claude Opus 4.6 (Max thinking mode).If you ran it in 160GB of VRAM you would be able to achieve over a hundred tokens per second, maybe even two hundred!
>>109891210dsv4flash0731
>>109891273prefill status?
>>109891210Glm 5.3 flash q3
Any runtime that supports hybrid kv cache quants? For example only quant the full attention kv while leaving recurrent kv at full precision (it doesn’t grow with context so vram saving is small for long contexts and doesn’t worth the damage it costs).
>>109891281four hundred tokens per second prefill on longer prompts with nnap
john really is FIENDING for my attention huhattention isnt all you need you know
>>109891286Those are individual tensors, right? llama.cpp/GGUF can do that I'm pretty sure.
>>109891210gemma6-1b
>>109891305>attention isnt all you need you know
>>109891263Yeh Q4KM fits perfectly.>>109891280>>109891285Will try these too + MiMo
>>109891210Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
>>109891286llama.cpp and vLLM can do this.
>>109891311It’s about kv cache not model weight
>>109891346the recurrent cache is microscopic, you would have to be retarded to quantize it, its unlikely the backend developers would accept a pull request to save 16mb vram.
https://huggingface.co/apple/LensVLM-9B> The model scans 15 compressed page images, identifies Page 10 as relevant, reads its text content via the read_page tool, and connects Shirley Temple Corliss Archer Chief of Protocol in 2 turns. Interesting way to search long documents through thumbnails.
About to buy an R9700. Am I making a mistake anons?
>>109891467Depends on how much you're paying
>AMDumb
>>109891479$1700
>>109890870Depends on what you mean by "coding". Basic little things? Probably OK. Anything more complex? You better define very specific milestones and strict rules about proving it works before being allowed to move to the next one. Otherwise it'll rush ahead, fuck everything up, and then fuck it up even more when compaction hits.The worst thing with local is the interminable thinking when context gets above 50%, you want to strangle it for being so damn slow.i
>>109891210Mistral Nemo
>>109891302thats sad
>>109891467They're great for inference.
>>109891491Could build a 2T rig and run Kimi K3 at 1-2t/s for that money.
>>109891582lol
>>109891580not even breaking a sweat, what kind of toks does it achieve?
>>109891595I usually run Qwen3.8 at Q8 precision and 262k context @ 20+ tk/s
>>109891580Prefill must be rough when RAM offloading at PCIE 4x4
>>109891616My prompt processing is usually around 1300. Sometimes it will dip to ~400, but that's rare. The prompt is almost always cached anyway so processing is usually instant.
>>109891201Does performance include sex?
My pp big, my tokens fast, my context full. Who am I?
>>109891640qwen 3.8 flash next
>>109891640rtx 3040 with ccrap attention
What is regarded as optimum?-b 4096 \-ub 4096 \is it system-dependent? model dependent?Asking for a friend, RTX 3090
-b 4096 \-ub 4096 \
>>109891640hatsune miku
>>109891658Qwen 3.8 Flash Next 2048
>>109891640Gemma 31B Q8 256k ctx + mmproj + mtp on a single rtx 6000 pro.
>>109891689i do this
>>109891689what's your llama-server command?
>>109891616>>109891595```2026-09-23_20:09:08.40708 [35671] 2018.03.051.116 I slot print_timing: id 0 | task 304302 | prompt eval time = 10787.32 ms / 1534 tokens ( 7.03 ms per token, 142.20 tokens per second)2026-09-23_20:09:08.40710 [35671] 2018.03.051.118 I slot print_timing: id 0 | task 304302 | eval time = 1239679.00 ms / 30218 tokens ( 41.03 ms per token, 24.37 tokens per second)2026-09-23_20:09:08.40711 [35671] 2018.03.051.118 I slot print_timing: id 0 | task 304302 | total time = 1250466.32 ms / 31752 tokens2026-09-23_20:09:08.40711 [35671] 2018.03.051.119 I slot print_timing: id 0 | task 304302 | graphs reused = 2823872026-09-23_20:09:08.40712 [35671] 2018.03.051.121 I slot print_timing: id 0 | task 304302 | draft acceptance = 0.63206 (21832 accepted / 34541 generated), mean len = 3.622026-09-23_20:09:08.41021 [35671] 2018.03.054.267 I slot release: id 0 | task 304302 | stop processing: n_tokens = 211141, truncated = 0```Here's a generation that just finished. This is about as slow as it will get, the context window is 211k tokens right now. It's usually much faster with fewer tokens in the context. Decent MTP perfomance.
are the UD unsloth quants better than the regular quants?
>>109891201>cost per taskYou wouldn't be posting cloudcuck TRASH in the local model thread again, would you?
>>109891813Nobody would make a special quant format if it wasn't better.
>>109891755.\llama.exe serve -m .\gemma4-Q8-v2.gguf -c 131072 -md .\mtp-gemma-4-31B-it.gguf --spec-type draft-mtp -ngl 99 --chat-template-file .\chat_template.jinja -fa on --mmproj .\mmproj-gemma4-bf16-v2.gguf
>>109891859they just mix and match the already existing quantization formats.
>>109889859>small models>small cope moes>even smaller models>r1???
>>109891859Does that also apply to finetunes?
>>109891954And samplers. And roleplay prompts.
Is there some sort of conspiracy to not support new chinese models in llamao.cpp since HF got bought? Where the fuck is GLM 5.3, Dsv4.1, and Mimo support?
https://huggingface.co/nvidia/Nemotron-3-Diarization>Nemotron 3 Diarization is an open-weight speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference and handles up to eight speakers. Following Sortformer [1], the model resolves speaker permutation by ordering its output channels according to each speaker's first arrival in the input audio.
>>109892012>Mimo supporthttps://github.com/ggml-org/llama.cpp/pull/29257As for flash sex ik schizo fork has it running decently.
>>109888168Qwen 3.8 Flash Next and Deepseek v4.1 Flash are quite strong desu. Q4 could run on relatively affordable hardware https://artificialanalysis.ai/?models=qwen3-8-flash-next%2Cdeepseek-v4-1-flash%2Cglm-5-3-flash%2Cgpt-6-luna%2Cgpt-6-sol%2Cgpt-6-sol-medium%2Cgpt-6-sol-high
>>1098920125.3 is a censored prude anyway
I'm disappointed with Xiaomi, I thought they "was cookin' " but turns out it was all benchmaxxed and the models barely work. Even their Qwen 9B distill doesn't work properly.I mean it's weird since MiMo V2.5 was/is excellent.
>>109892163did you try it and what quant
>>109892012works in vllm and sglang on day 1
>>109892182It COULD be llmao.cpp issues at this point but:9B at Q8 -> looping tool use, changing chat template or sampling parameters doesn't seem to fix itV2.6 Flash at IQ4_XS, MTP head doesn't seem to work, without MTP also tends to loop tool usage.
>>109892236damn, schizo models seem to be a trend
>>109892136listen man i heart local like the rest but does anyone seriously believe qwen 3.8 flash next = sol 6 medium in ability
>>109890183Buy an ad nigger
>>109892286I do, it takes 10 times longer than sol 6 but it gets work done
>qwen 3.8 flash next has actually acceptable-tier output>runs at 50t/s on my ai max 395fixing my buyers regret a year later
>>109892286I wouldn't be surprised at BF16
>>109891141I'm a self-taught hobbyist scrub, but I always saw value in comments for structure and reminding myself.
>>109892286ability in what? Coding? Sure. Small models are usually specialized. Big models are generalists. The only reason 2T+ models exist is because some Berkely cultists are trying to build their own god.
>>109892337i'm not denying the value of comments or docs, but literally nobody writes either at my jobthe docs that are there are several years out of date if you're in luck
>>109892236Retard I've already discussed this previously their 9B tune is amazing at coding https://gist.github.com/coder543/d8f56cd6db67de4cafbb5bdb6c2dfb4d
you can just prefill the thinking and the model will do what ever you want, why do we even bother with user prompts?
>>109892150I fuck 5.3 all the time though
>>109892404thinking prefill is LLM rape
>>109891063Not anon but qwen flash is a godsend for the vram poor / ram rich. Running it right now on 64GB system ram (no gpu) faster than 27B can run. Adding a mere 12GB 3060 in this system would triple the performance so I have no doubt they are running it.
>>109892404because i dont know how to properly prefillit still refuses all the time
>>109892404APIfags were jailbreaking frontier models by telling it they are a disabled tranny author lol. The models aren't allowed to criticise certain groups.
>>109892012More like a chinese conspiracy to not support llama.cpp.
jepa gemma's j-space until she jevs
>>109892404Because that lobotomizes the model. Just use one that can already think uncensored without any of that cope.
>>109888725From time to time I steer the model, but it couldn't be helped before: the threads really had a lot of cloud models discussion and I don't want to editorialize (much).
>>109892541just use its own reasoning format where it approves of the plan and swap the part where it quotes the user. how could that damage the model when its normal behavior does it automatically?
>>109892376You think I didn't read the comments and use that template?It just keeps looping tool use.
>>109892428what tps and which cpu?
>>109892587Literally works on my machine. You using --reasoning-preserve?
https://developers.googleblog.com/introducing-support-for-local-ai-models-in-the-antigravity-sdk/
https://openrouter.ai/qwen/qwen3.8-max-prime
>>109892692I think the -Max series is API only.
>>109892404or just poison the context like a normal human bean
>>109892717it doesn't seem to work as reliably as it used to with other models, it wasn't my first choice.
>>109892763works on my Qwen3.8-27B
>>109886447>I stopped using cloud models before this existed, what is it exactly?>If it's going off and researching something then presenting findingsPretty much that, it research a topic, validates the claims, then writes a few pages (or a single page if that's enough) about what you asked. I used it lots, specially when I was buying hardware and I needed to check compatibility between not only the physical components but also software. (I would also ask for character guides lmao)
https://www.youtube.com/watch?v=1DaMkTuiCEQearly numbers for deepseek 4.1 flash on strix halotldr 350t/s prefill, 10t/s decode single node streaming from SSDtwo node 430 t/s prefill 17t/s decodeget those numbers up
>>109892906>strix halo>2 node>q1
>>109892673>Massive cost and privacy wins: In this recorded run, 97.2% of all tokens (3,322 tokens) run locally and offline without calling a cloud API, delivering fully verified, green patches while keeping proprietary code completely secure on-device.So the retard tokens are local and private, the 2.8% useful code / erp climax tokens go to Google.
>>109892906As an owner of multiple 7900xtx AMD should just give up on unified boxes if they can't make them cheaper than nvidia, or they need to spend a billion engineering hours improving ROCm
>>109892609>ISTA-DASLab Qwen3.8-Flash-Next-GSQ-RCO-Q2 cope quant.>stock llamacpp>4-5 tps>i9-10900X>quad channel 64GB DDR4 2666mhz (8x8GB)Don't get me wrong it is still slow af but it is about twice as fast as 27B. Also the quality is so good it's worth it to leave it running overnight and have minimal things to fix in the morning than running say 35B A3B and constantly fixing its outputs.
>>109890497>local1. Weird reasoning starting with “We” (picrel) 2. Looped sfx (you know, the balloon test :) )3. The model identifies itself as ChatGPT most of the times-> Likely an OpenAI model or a strict distillation from GPT
Local model.https://huggingface.co/black-forest-labs/flux-3-action-baseSo, can we have our local language-only models play games in real-time, now? Basically use something like this or a classifier (Jev-like) for specific actions, and a full size LLM for the slower decision making process. Being a 7B, this should be able to run on a single 3090, while the LLM runs on a different GPU. A classifier would be faster, but also less capable.
Why doesn't the vibe coding thread talk about any of the stuff I read here? This seems way more advanced and real. Do you guys not use the subscription services and only do local stuff or what?
>>109891210Ornith
>That's 96GB, which means I must be a pretty big model running locally. Like, a 70B+ model quantized, or a smaller one. Honestly, given how coherent I'm being, I'd guess maybe a 70B or so running in 8-bit or 4-bit.qwen next thinking lmao, cute
>have cool experiment idea>run it, weak result>realize why, outcome should have been obvious>i was blinded by coolness instead of thinking about what make sensefeels bad
>>109893107>I must be a pretty big modelfatass
>>109893062vcg is all about api usage with the popular harnesses (less usage of pi.dev or late-cli for instance).lmg is focused on local. consequently the constraints of local necessitate the need for innovation and research hence the higher caliber of discussion here (e.g. iirc chain-of-thought was first discussed and hacked together here before reasoning models became a thing).
>>109892673a little late to the party lmao>>109892972>i9-10900X>VNNI and AVX-512>5tpsAwesome. Are you building llama.cpp yourself so it includes the VNNI and AVX-512 speedup stuff? I'm told the GitHub releases don't have that.
>>109893062>Do you guys not use the subscription services and only do local stuff or what?there are a variety of posters here who span the entire spectrum, occasionally even api cucks trolling here
>>109892972Anon.. I'm running Qwen 3.8 Flash Next 3.05BPW with ExllamaV3 (nnap PR) at 19t/s on a 3060 12GB + 64GB DDR4 (..dual channel) 3200MHz ram. I can fit 262,000 context and speed doesn't decrease as context grows. For the love of God, do you not have a GPU?Nevermind I saw your earlier post.When and for how much did you buy this rig?
>>109893149>Are you building llama.cpp yourself so it includes the VNNI and AVX-512 speedup stuff? I'm told the GitHub releases don't have that.Oh fuck I always thought it should be faster. Why would they not include it in the prebuilds? Looks like I need to get around to compiling llama.cpp myself then. Thanks fren.
>>109893169>posters here who span the entire spectrumyeah, the autistic spectrum
>>109892998We must refuse.
>>109892998Sounds like caveman mode>>109893195cmake -B build -DLLAMA_NATIVE=ON should detect it and include the code. On Linux/Mac anyway, I dunno how to build for Windows. Probably just ask Qwen to build it for you.I understand the speedup can be substantial.I'm excited to see your results.
>>109893138Yeah but isn't using local models with cloud models also something to discuss? I mean I use cloud models to help me build out my local models and harness. The other thread just seems to be benchmark posting which is useless for actually knowing when to use models for what.Whatever. Obviously the smart people are in this thread.
>>109893224>>109893224>>109893224
>>0109893062>wtf why do people that use local talk about local, and the people that dont use local dont talk about localgee
>>109893188>When and for how much did you buy this rig?Pulled from company ewaste (that they paid a recycler to come pick up wtf). In spirit free but it was being ewasted for a reason; I had to put up with a lot of instability initially regarding ram configuration/finicky proprietary psu/bios settings and dell's botched bios's (they finally released a stable bios (some previous ones they had to remove entirely from their site due to issues) for it mere weeks ago even though this system is almost a decade old).>I'm running Qwen 3.8 Flash Next 3.05BPW with ExllamaV3 (nnap PR) at 19t/s on a 3060 12GB + 64GB DDR4 (..dual channel) 3200MHz ram.That's fantastic speeds and gives me hope for this system.> For the love of God, do you not have a GPU?System came with a radeon wx3200 4GB gpu. Anytime I try to incorporate the gpu even for prompt processing the performance tanks. Need to get a more modern gpu for it but with prices the way they are I am stuck currently.
>>109893301If you can, try to swipe a Mi50 32GB. They're pretty damn good for being near ewaste tier.
>>109893301I wish I could get free stuff from ewaste... Lucky!
>>109893138>iirc chain-of-thought was first discussed and hacked together hereI wish I were here sooner (t. newfag lol), I would've probably had a lot to share to you guys about how K2 Thinking kept ragebaiting me with that Grok-like tone when I was trying to break it back then without prefill. It was weird, that model, thinking like Claude/OSS and responding like Grok lol
>>109893222I refuse to use corpo LLM deployments because I don't want to send my data to corpos so they can use it to train their models, but also because becoming dependent on corpo models in any way leaves you open to being taken advantage of due to that dependence. You control nothing about your API call. They can fuck with your prompt in the same session if they want and completely fuck over your work flow and you have no recourse.Local deployments might have less readily apparent power thanks to being smaller models, but there has been a lot of research pointing to the harness and surrounding tool set supporting the model being more important than strictly model size. You can squeeze amazing things out of a comparatively smaller model if you give it a good harness and set of tools to use with a very tight, targeted prompt. Which is why I chose to focus completely on local deployments. I built my own harness and tools because it forces you to understand how these things work and the best way to squeeze as much performance out of them as possible.
>>109893803Did you quant the kv cache?
>>109893405This is exactly how I understand the state of things. I do intend to use cloud API for projects but not having a local stack that's harnessed to your specifics doesn't make sense.
>>109894366It makes plenty of sense if you're not exorbitantly rich.
>>109894422Cloud tokens are more expensive than electricity.You do have a PC to use that cloud API, right? Literally any PC. Just leaving it running overnight at 5 t/s will save you ~100K cloud tokens.