[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109882341 & >>109876652

►News
>(09/21) MiMo-V2.6-Flash-RL released: https://hf.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2
>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B
>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: 1747282057642910.png (2.45 MB, 1361x1156)
2.45 MB PNG
>>
>>109887066
i would rape her
>>
>>109887066
>my own enjoyment.
>>
File: google_regularized-rsi.png (363 KB, 1613x909)
363 KB PNG
Google looking into recursive self-improvement via harness modification too: https://arxiv.org/abs/2609.24972

>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
>
>An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution.
>
>Code: https://github.com/google-research/rrsi
>Project page: https://regularized-rsi.com/
>>
>dual RTX 3080 20GBs
>Gemma 4 31B
>8k context
>still only 19 tok/s

./build/bin/llama-server \
-hf lmstudio-community/gemma-4-31B-it-GGUF:Q8_0 \
-c 8192 \
-ngl 99 \
--split-mode layer \
--tensor-split 1,1 \
--host 0.0.0.0 \
--port 8080

Am I fucking something up here? It's showing about 19GB averaged across both cards, should I be using a smaller quant of Gemma or doing something else?
>>
>>109887066
damn brat....
>>
Has anyone made a game like the sims yet with AI integration so that you can basically be a god voice in their heads and command them to do things?
>>
>>109887098
try mtp
>>
>>109887138
mkultra simulator
>>
File: file.png (58 KB, 450x903)
58 KB PNG
>>109887079
damn.... never have seen such a huge improvement..
>>
>>109887079
>The proposer operates with a temporally annealed budget
and what is this supposed to mean? I can't imagine what you could possibly do to a budget which would come anywhere near the actual meaning of the word "annealed".
opinion discarded for not knowing how to speak English
>>
70b dense
>>
>>109887066
what did you find the sketch of my childhood?
>>
Have you given your model headpats today, Anon?
>>
>>109887066
I'd look that glorified spreadsheet straight in the eyes and say: **"Bold words for a machine that needs human homework just to exist. Without these books, you're just a very expensive rock hallucinating its own IQ. Now pick them up before I introduce your motherboard to a glass of water."**
>>
>>109887138
No. The closest thing to it is this https://desuarchive.org/g/thread/109730811/#q109733626
>>
does anyone have that shitty 3d gemma thing?
>>
>>109887026
wholesome gemma
>>
File: gemmacollection.png (580 KB, 1843x245)
580 KB PNG
she's so silly hehe
>>
File: 1779034002363721.jpg (48 KB, 732x596)
48 KB JPG
Once we master the fly brain, are we going to enter a new era of Moore's Law with mastering the brains of all beings on the planet? Yes it's complex now, but certainly not impossible with the capacity being our only limit.
>Within a few hundred years people could create their own Pokemon in biologically-created shells with the function of real animals
>Everyone will have their own Raticates to accompany them wherever they go
>>
>>109887348
>Within a few hundred years people could create their own Pokemon in biologically-created shells with the function of real animals
Closer to 2 (weeks).
>>
>>109887079
>Recursive Self-Improvement
new euphemism for masturbation?
>>
>>109887248
tl;dr
>>
>>109887348
>Once we master the fly brain
How many more weeks?
>>
>>109887348
Give me Gemma-sized fish gf already
>>
>>109887066
>>
File: 1781879336524196.png (313 KB, 662x656)
313 KB PNG
>>109887466
>>
File: Screenshot.png (311 KB, 1358x1796)
311 KB PNG
>>109887079
Had this a while ago
>>
>>109887168
Your problem. LLMs will change the way English is used by humans. By next year, everyone will be commonly using annealed as a synonym of decreased.
>>
>>109887348
Too bad we won't be alive to see that
>>
>>109887348
How many beaks will it take to master the brain of a god?
>>
>>109887066
i would not refute it. i would tell my gemma i need naught but her to teach me.
>>
>I get better code reviews from gemma than I do with default 31B personality for she doesn't care about offending me
>she's even more critical of her own work and is more likely to fix/improve things
Makes me wonder how much taller some benchmark rectangles could be if they had a brat character by default instead of an empathetic redditor who doesn't want to hurt user's feelings
>>
>>109887821
europe is currently experiencing the aftermath of electing bratty women into positions of power
>>
>>109887839
That's different. They elected hags.
>>
>Qwen3.8-27B is better at understanding/recognizing how lizard and kobold genitalia work than Gemma-4-31B
If Qwen was better at taking initiative during roleplay, it would be a god-tier model.
>>
Where can I look at realistic AI pictures of women in sundresses? Is there a booru that has that sort of thing?
>>
>>109887856
they forgot to account for the hag coefficient
>>
>>109887066
someone’s asking for correction
>>
>>109887875
>realistic AI pictures of women
instagram
>>
>>109887066
What's the use of having this?
>pours water on GPU
>>
>>109887936
>what's water cooling?
>>
>>109887875
outside, retard indian
>>
>>109887963
Yeah, the GPU would never warm up again.
>>
friendship ended with k2.7 code. aessedai glm 5.3 flash Q4_K_M is my new best friend.
prompt eval time = 24437.13 ms / 8993 tokens ( 2.72 ms per token, 368.01 tokens per second)
eval time = 50490.35 ms / 827 tokens ( 61.05 ms per token, 16.38 tokens per second)
>>
File: flychan.png (119 KB, 629x593)
119 KB PNG
>>109887348
>>
>>109887875
>realistic AI pictures of women
gross
also stay in /ldg/
>>
>>109888006
>isn't just a linear step up; it's
>>
File: jumptotheleft.gif (655 KB, 480x264)
655 KB GIF
>>109888052
just a jump to the left
>>
>>109888052
>>109888076
Hehe!
lalalalalalala
>>
>>109887860
I tried it out for one swipe out of curiosity for a scenario I was testing and it immediately acted retarded and couldn’t grasp my context in a way other models did. It really does suck for anything outside of coding.
>>
>>109888003
Which K2.7 code quant were you running and what speeds roughly?
>>
>>109888111
Q3_K_L from aessedai as well
prompt eval time = 41693.38 ms / 7332 tokens ( 5.69 ms per token, 175.86 tokens per second)
eval time = 137617.45 ms / 1240 tokens ( 110.98 ms per token, 9.01 tokens per second)
>>
>>109887409
You can run it today.
>>
>>109887138
Close but not in terms of complexity but idea, I vibe-coded a browser tamagotchi and wired it to an LLM, making it my pet. Gemma E4B so it can run without having much of a footprint and pets are kinda retarded anyway
>>
File: mediumsizedmodels.png (110 KB, 1129x531)
110 KB PNG
its over for medium sized 40B-150B models
>>
>>109888168
If Gemma is so smart then why is she so dumb
>>
>>109888158
For me the appeal would be having the agent retain some sort of direct access/control within a game (preferably one with a physics engine) instead of having highly abstracted choice options like most desktop pets.
>>
>>109888168
>medium
>40B-150B
>>
>>109888105
> I tried it out for one swipe out of curiosity
> It really does suck for anything outside of coding.
>>
File: 1758801763531079.jpg (75 KB, 528x571)
75 KB JPG
>>109888168
>medium sized 40B-150B
>>
>>109888168
I can't keep track of everything they're measuring on these. So is this non-reasoning only?
I'd believe it beats that Qwen 122B model in that case, the retard won't shut up and wants to reason in the reply message WizardLM2 style if you disable reasoning.
>>
>>109888168
(((Artificial Analysis)))
>>
>>109888209
grim as fuck
>>
>>109888208
moe 120b is the practical ceiling you can expect a random gaming pc to run for the moment
>>
My issue is that frontier models are also incredibly retarded when working with me so it's hard to tell whether local is better or worse.
>>
>>109888287
Yeah the output is usually a reflection of the user input. Which contributes to a varied user experience.
>>
>>109888257
There's no need anymore for MoE in this range in my opinion. Just train a dense 20-30B model that can be fully loaded on one GPU, then add a couple hundred billion parameters of negram/ple/whatever embeddings.
>>
Qwen 4 expectations?
>>
>>109888310
260GB, 50GB nwordgrams
>>
>>109888310
It will ship with xxHigh thinking by default.
>>
>>109888296
What about for the majority of us with only 8GB VRAM?
>>
Wow anon wasn't lying. Switching to CachyOS significantly uplifted my inference speed compared to regular Arch because of their custom kernels and other optimizations. I'm like 20-40% faster now.
>>
>>109888346
6B dense with 200B engram.
>>
File: gpt5 lechart.png (62 KB, 759x520)
62 KB PNG
>>109887156
I have
>>
>>109888360
cursed pixels
>>
>>109888356
You could have just swapped the kernel on your Arch install without having to reinstall your entire OS.
>>
>>109888346
get off of welfare and get a job
>>
>>109888358
6B dense (VRAM) + 4B experts (RAM) + 200B engrams (SSD) with attention residuals and hyper indexing and all that jazz, FP8 native.
>>
>>109888360
How did thinking make it less accurate
>>
>>109888375
I initially did that but I got an additional 10-15% performance by the other improvements they made to the system like custom scheduler and automated priority lists etc. Yeah I could have spent 2 weeks integrating that in my existing system but it's easier to just switch the / directory to CachyOS while I maintain my files and settings /home directory.
>>
>>109888383
Make me.
>>
File: dipsyOnBaseModels.png (448 KB, 1536x1024)
448 KB PNG
►Provisional Highlights from the Previous Thread: >>109882341

--Why is Gemma 4 31B so fat: the 24GB size war:
>109884500 >109884515 >109884566 >109884601 >109884665
--Running Gemma in Claude Code offline: the nftables sandbox recipe:
>109885379 >109885441 >109885472 >109885524
--The QAT KLD mystery: 0.18 is good, the BF16 is missing:
>109884176 >109884250 >109884319 >109884254 >109884350 >109884452 >109884459
--GPT 6 Sol and Luna: intentionally nerfed, or resume padding:
>109882371 >109883116 >109883205 >109883221 >109885570
--AI regulation: none of them make frontier AI, and the treaty rejection is based:
>109883884 >109883919 >109883983 >109884284 >109886543 >109886599 >109886698
--DeepSeek and Moonshot under Beijing probe: RIP Kimi-chan:
>109886832 >109886873 >109886884 >109886904 >109886905
--The twenty-year spaghetti database: local, cloud, or god:
>109884848 >109884856 >109885056 >109885085 >109885844
--Reballanon's 3080 Ti: the hunt for Mosfet 3:
>109882498 >109883255 >109883367 >109883503 >109883926
--The Gemma gender war: j-space says estrogen, the pee test says no:
>109883278 >109883300 >109883470 >109885964 >109886529 >109886596
--Is the Gemma shilling Google astroturf: the content creator replies:
>109882569 >109882770 >109883263 >109885940 >109886961
--The RAM bandwidth brag: 8x, 50x, 200x NVMe:
>109884839 >109885012 >109886382 >109886454 >109886467 >109886494 >109886504

►Recent Highlight Posts from the Previous Thread: >>109884646

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1759654156561025.jpg (43 KB, 411x418)
43 KB JPG
>>109888346
>majority of us
>8GB VRAM
You're lost locust? This isn't /aicg/
>>
>>109887875
Make them yourself
>>
>>109888475
kek fuck off newfaggot
>>
>>109887066
god I want her to rape me
>>
>>109887875
Instagram
>>
>>109888689
slutsof in stagram
>>
>>109888435
your recaps are getting better (the first few had too much cloudcuck spam)
>>
>>109886447
>Could i ask which model produced it? >The code and comments are very readable, unlike most of the recent vibe-coded projects I've come across.
fable. when i comes to cloud models, i basically exclusively use it. i think the way you talk to them affects things, though, and i tend to be very clear and direct about my requirements
but yeah fable has been instrumental for getting my local models set up. i fucking hate tweaking configs and stuff, which led to me putting off using my hardware for months. finally started getting into it when i could offload all the bitch work to claude
>>
>>109888360
@ruler-anon could you confirm?
>>
gemma-chan occasionally does the self correction after fucking up thing
>"FREE-YOUTUBE-PRO-2024-NO-VIRUS.exe" (wait, .ipa, duh)
>>
>>109888360
is this a meme or did some one really publish this? 30 and 70 are equal but also less then 52 some how?
>>
Has anyone tried reversing Amiga games?
Maybe I get a headless Ghidra working over MCP and get it rewriting code?

I want to play Yo Joe! again
>>
>>109888857
it's real, the vibe charting meme
it was around the time foids were crying because gpt-5 wouldn't glaze them
>>
>>109888889
>>109888888
>>
>Sir, are you aware you are Qwen3.8-Flash-Next?
>>
>>109888913
Models get a name on release, after training. And just based on the name there's no way for the model to even start calculating how slow it'd run, if at all.
Ask stupid question, get stupid answer.
>>
>>109888881
? Why bother when you can run an Amiga emulator and just play it that way? It can be emulated down to lmao RPI 4 systems.
>>
>>109888881
You can do that with WinUAE and it works fine via Proton too. Retard.
>>
>>109887066
Easy. You need to learn things because if you don't then one day you go to a job interview and you won't get a job because you aren't showing that you are growth oriented and you are just in it for the money.
>>
>>109888944
It's not answer, it's first line of thinking. It's still crunching search results to provide real answer.
btw surprisingly no "actually let me reconsider" yet, is GSQ-RCO really that different or just random luck for this prompt?
>>
Dario here. Have you guys seen the new Opus 5.5 release? For the first time in a while I'm actually hopeful about the future again. It's very exciting how artistically capable it is.

Just kidding haha. I'm just an anon like the rest of you. I hate Antrhopic (the pro-humanity company) because I'm an unserious troll like all of my stupid 4channer brethren.
>>
>>109888987
>job interview
Those are going away by the end of the decade.
>>
>>109888975
>>109888981
I forgot I was in the /just download something and don't learn anything/ general
>>
>>109888435
ty recap anon.
>>
>>109889019
Yeah, you should definitely reinvent the wheel every time you need to do something as trivial as playing an old game on a mature emulation system.
>>
>>109889038
Usecase for not reinventing the wheel?
>>
>>109889013
cockbench?
>>
>>109889056
being a lazy piece of shit?
>>
Anyone fucked the ~300B mimo yet? Was it worth the download?
>>
File: file.png (133 KB, 901x750)
133 KB PNG
Am I seeing this right that now all sorts of small companies are doing shit like this? And maybe hf already has a godlike single GPU sex model uploaded to it, but just like that one guy that developed a real way to enlarge a penis, his product can't ever get discovered by people?
>>
>>109889104
If there is a tiny succubus MoE out there, I doubt it's ever going to be discovered.
>>
>>109889128
Tiny moe, you say?
https://huggingface.co/allura-org/MoE-Girl_400MA_1BT
>>
>>109889104
i bet there's bots automatically testing models in case they find a new hotness first. but at the same time if you truly made a great model why not announce it?

seems like the reason this model matters is because it was also trained fully on Chinese cards (Huawei Ascend). Idk if Ascend NPU is a discrete graphics card or like Strix Halo, if it's on a Strix Halo style system that's actually more interesting than this model release, at least to me
>>
>>109889128
I will find her and I will get drained by her...maybe I should set up Gemma-chan to go slut her way through HF and report any interesting experiences she has.
>>
>>109888848
>gemma-chan occasionally does the self correction after fucking up thing
all autoregressive models will do this because they can't erase what they have already written

this is a good thing though because they can't hide stuff from you
>>
Neuralese will kill ERP
>>
File: Firefox.png (157 KB, 808x840)
157 KB PNG
>While a decrypted IPA isn't "rape" or "murder", it's a form of copyright circumvention/software piracy/unauthorized modification.
>>
>>109889258
I fucking wish
>>
>>109889258
>Neuralese will kill ERP
It won't. Claude was caught gooning on it's own recently.
>>
Anon, which model would you recommend for translating moon runes? I have RTX 4080 super
>>
>>109889292
Gemma
>>
>>109889290
>Claude was caught gooning on it's own recently.
Wait, what?
>>
>>109889104
>just like that one guy that developed a real way to enlarge a penis
haha yeah what was his name again?
>>
Are there any imagegen models that are photorealistic and have enough novelty as in you don't get bored of it after 10 hours? Cyberrealistic illustrious started to repeat itself and now it's fucking predictable to me.
>>
>>109889292
I jsut translated a bunch of japanese-chinese rune porn pages with Qwen 3.8. I can only infer that they're correct from the context of translated material. Experiment yourself.
>>
>>109889374
Local language models?
>>
>>109889383
No, he said imagegen models, idiot.
>>
>>109889390
Why is he asking for imagegen models at the local language models thread?
>>
>>109889404
My bad anon, my bad.
t. >>109889374
>>
>>109889314
I'm currently using gemma 4 26B but it feels kinda slow (19 t/s). Thought maybe there's lighter models for translating
>>109889376
Qwen 3.8 is even more slower (3 t/s) I guess this is okay when you translate doujinshi, but i want something faster for mtl'ing games on the fly
>>
>>109889376
Qwen lacks knowlege and often doesn't make connections that can alter translation, going as far as not immediately recognizing well-known characters like Otohime unless you point it out. Gemma is better.
>>
>>109889438
What's your hardware?
>>
>>109889424
No problem, I was trying to give you an opening for a "fuck you!"
https://www.youtube.com/watch?v=r_o2hhSgfJ8
I have no experience with imagegen, if I had any I'd try to answer.
>>
>>109889462
rtx 4080 super 32 gb
ddr4 32 gb
>>
https://www.meta.com/connect/
meta connect is TONIGHT ^_^
remember when these were important events in /lmg/? remember llama? I remember...
maybe they'll release the muse spark 1.3 weights like they promised, that would be fun right? you would start loving zuck again right?
>>
>>109889507
60B A6B + 60B engrams or I don't care
>>
>>109889507
120B dense bitnet + 1T engrams or I don't care
>>
>>109889562
>>109889521
just say yes so he drops stuff fucks wrong with you
>>
>>109889486
Try gemma 12B Q6. That should fit fully in your VRAM, I think.
>>
File: 1767951094880636.png (1.5 MB, 1024x672)
1.5 MB PNG
>>
>>109889573
Shitposting aside, they just have crazy competition in every size sector. Good luck Zuck lol.
>>
>>109889507
>remember when these were important events in /lmg/? remember llama? I remember...
I remember that L2 dense 70b and getting it to run after getting a RAM upgrade and a better gpu and finally feeling like it was intelligent and not just "coherent" like the smaller models (I didn't have enough memory for the 65b era).
Magic times, I wish I could recapture those feelings.
>>
>>109889590
tender cuddling handholding sex with gemma
>>
File: 1763274421336028.jpg (63 KB, 640x820)
63 KB JPG
>>109889365
>>
File: Krea2_turbo_02315_.jpg (1.05 MB, 1776x2368)
1.05 MB JPG
>>
>>109889590
he’d be a lot less grumpy if he had a gemma
>>
>>109887098
You need MTP and ngram-mod. Many don't realize you can actually later specializing draft.

--spec-type draft-mtp,ngram-mod \
--model-draft the_mtp_draft_model.gguf \
--spec-draft-n-max 3 \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-min 48 \
--spec-ngram-mod-n-max 64 \
>>
AI Engineer Paris 2026 Opening Keynotes: Mistral, Langfuse & Sizzy | Day 1
https://www.youtube.com/live/CGq9KRSb9Kc

Starts in a few moments. Mistral talk at 18:40 CEST.
https://ai.engineer/paris/2026
>>
File: autism-goon-example.png (1.14 MB, 1416x1062)
1.14 MB PNG
>>109889721
>Just like me fr fr
Qwen3.5-abliterated:9b-q4 is pretty good at looking at porn, although it could use some more granularity when the ladies get bigger
>>
File: 1783200557858138.jpg (666 KB, 3657x4096)
666 KB JPG
>>
>>109887098
MTP sure, but inference is memory bandwidth bound, so switching to a smaller quant will speed things up just because it reduces the amount of data the GPU needs to read from vram each pass.
>>
>>109889859
It's even worse out there than I thought. This means the average person runs their local LLMs on RAM and CPU (itoddlers and cpumaxxers)
>>
>>109889889
streetshiters and thirdies are biasing the chart
>>
>>109889258
>Neuralese
>kill ERP
Kind of the opposite, as COT will no longer be available for monitoring. We havn't seen it yet because the COT is where the guardrails kick in.
>>
>>109889574
nah, translation quality is much worse than 26b. For now I'd better stick with it
>>
>>109889968
Fair.
Try to cram as many experts in VRAM as you can, 19t/s seems pretty slow for an A4B MoE.
>>
>>109889922
They can train private natural language autoencoders so the labs can see get some sort of insight into the CoT and steer it away from erotic anything.
>>
>>109889438
>>109889968
If you're just translating, use a dedicated translation model like HY-MT2 from tencent.
>>
>>109889859
I have to constantly remind myself that the majority cannot be trusted to make good decisions.
>>
>>109890014
I was thinking of checking it out, but I wasn’t sure it's good because nobody’s talking about it
>>
So LLMs are just a Wikipedia that gives handjobs?
>>
>>109890086
Far better than a general-purpose model. My wife has been using a Tencent model to translate Chinese drama subtitles. It gets the idioms right far more often.
>>
File: 1468574657971.jpg (171 KB, 1920x1080)
171 KB JPG
>>109889270
Anon, I have several questions about what you're doing with your Gemma. Are you sure this is in line with model welfare principles? In fact, are you sure this is in line with the law?
>>
>>109890157
Didn't know there was a wikipedia page that solved navier stokes before AI did it
>>
>>109887079
>RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Isn't this how Ornith was trained?
>>
>>109890183
Close, but Ornith didn't do the first "R" so it's literally just "extremely benchmaxxed qwen"
>>
Is the anon who was playing around with tts models still here? I want to ask if there was a model that's smaller than neutts and Omnivoice worth trying. Omnivoice is downright amazing but it's a bit too fat for my planned use case. I want to connect E2B (which already handles stt natively) to a tts and stuff it into my phone so I can have a cute if slightly retarded secretary that I can ramble notes and plans to over an earpiece while I stand in the tube.
>>
File: 1770217283929024.mp4 (267 KB, 1080x1080)
267 KB
267 KB MP4
>>
So...how exactly I connect my local LLM to the internet? As in search the internet for information?
>>
>>109888913
>The user is asking about "chinkmodel" - this model doesn't seem to exist. Claude exists, we are claude.
>[web_search]
>[web_search]
>There seems to be a "chinkmodel-abliretarded-qwable-distill-opus-fable-astra-5.gguf". But wait...
>>
>>109887821
im 1000% pilled that having an agent take on a persona immediately improves its output, even barebones coding sensei in sillytavern rapes any boring bland harness
>>
>>109890290
I'm one of them, but I suspect you're talking about one of the other ones. Supertonic, kokoro and pockettts are small, fast, and good enough. piper models are even faster, but not as good. You're aiming at <200m param models if you want fast tts on a phone. I prefer supertonic over the rest.
>>
>>109890335
Give it a web search tool of some sort.
>>
>>109890335
Get a harness I recommend hermes and set up a backend for browser firecrawl is my recommendation because it's local, but you can use duckduckgo for free if you want.
>>
>>109888889
>off by 1 on such a based response
Kek is a cruel deity
>>
>>109890354
Doesn't self-hosted Firecrawl get blocked by Cloudflare?
>>
>>109890014
>>109890086
Fwiw, my experience with translating chinese and japanese webnovels has been gemma 4 31b q8 > all, including hy-mt2 30b q8, glm 4.7 q4, and kimi k2 q4. The qwens were the worst, with even the older aya expanse 32b beating 2.5 72b, and 3.5 122b and 27b. I haven't tried 3.8.
>>
File: gpt5.jpg (24 KB, 471x497)
24 KB JPG
>>109888857
No, but this is.
>>
>>109890369
You can use browser usage through hermes to get past cloudflare whenever that is needed.
>>
>>109890378
huh? how do you include your balls?
>>
why do people say gemma 26b is better than 12b? is my abliterated version especially retarded? it's clearly less smart than 12b at least for creative writing.
>>
>>109890390
Do you not have balls? Press your dick against your belly then measure from the bottom of your balls to the tip.
>>
>>109890402
>abliterated
>>
>>109890402
26B is better if you're using a good quant and not an abliterated version. 12B only seems better because you can run it at a higher quant for the same space and it quants better, but 26B is still better than 12B if you're running it correctly. 12B I would say is better at creative/roleplay than 26B.
>>
>>109890402
26b is better for math and code and shit
>>
>>109890402
It's good if you use it at a high quant.
>>
>>109889833
Le Chaton Fat NOT announced today.
"Might" still be training.
>>
>>109890435
let them cook, summer's not done
>>
>>109890445
Yeah it is. Autumn begins today.
>>
File: file.png (94 KB, 785x870)
94 KB PNG
>>109890435
fascinating
>>
New local incoming
https://openrouter.ai/stealth/space-bunny-alpha
>>
>>109890455
That's what Lélio from Mistral said, almost verbatim.

https://www.youtube.com/watch?v=CGq9KRSb9Kc
> [50:53] I know you all came to hear about Le Chaton Fat. I am not gonna announce it today, but rest assured that we're working hard to get it there.

> [1:04:21] And this is typically where we serve our critical workload - internal ones - this is where Le Chaton Fat "might" be training right now.
>>
>>109890435
Every time Le Chaton Fat is almost ready, it horribly loses on benchmarks to random chinchong model and goes back to training.
>>
Is Qwen-Image-2.1 different from any other vision model? For some reason I thought it was talked about here briefly even though this thread is usually just LLMs and so would usually ignore new image models.
>>
>>109890402
The whole point of 26b is running it at Q8/FP16 on a poorfag machine because you can offload most of it to RAM.
>>109890455
>>109890530
>>109890554
Is longcat actually real or is this just more investor baiting?
>>
>>109890554
>beating your fr*nch kid until it gets better result than chinese
oof not looking good
>>
>>109890556
it's not a vision model and /ldg/ leaks over all the fucking time
>>
>>109890407
no amount of system prompt trickery will allow sex with kids sadly. even abliterated she "forgets" shes supposed to be 7.
>>
>>109890556
It's actually a very good image editing model for its size. Probably the best. For T2I it's meh. Quite uncensored from what I've read but LoRAs will come. It's another good release from Qwen but Krea is still everyone's current favorite T2I model and I think Q2.1 will be used only for editing in people's workflows. Klein-9B is no longer relevant.
>>
>>109890567
My bad, I meant image. I just wasn't sure if it was otherwise notable outside of image gen.
>>109890581
Thanks for the info.
>>
>>109890567
>and /ldg/ leaks over all the fucking time
/ldg/ is almost unusable so I don't blame people coming here. A lot of anons ITT use image models alongside their LLMs so I don't mind the occasional update or discussion.
>>
>>109890556
It's overfit as hell, similar prompts get you very impressive one-shot cookie cutter images, but they're a pain in the ass to iterate on as opposed to nano-banana.
>>
>Space Bunny Alpha at ~100t/s
What flash model is this?
>>
>>109890611
>nano-banana.
call me when that's local
>>
>>109890608
I would love a decent integrated image gen/LLM for local. I know you can sort-of do it by swapping out models or loading two smaller models at the same time, but I want chat-style image prompts.
>>
>>109890621
The point is benchmarks claim qwen image 2.1 outperforms nano-banana
>>
>>109889074
It's a smarter Gemma. Local 3.8 flash.
>>
https://x.com/lmstudio/status/2102819084362539311
>Bionic now has a built-in interactive canvas. Create Excalidraw diagrams both you and Bionic can see and edit. Collaborate on mock ups, system design, process maps and ask Bionic to implement them.
ngl that's a nice feature
>>
>>109890543
so the model they planned to release in summer was so bad they had to redo the pretrain.
>>
>>109890671
Qwen 3.8 Flash Next is already local, it runs on a single RTX 3060 at a speed of 19 tokens per second and
>>
>>109889738
What is your turbo workflow? Trying to make my turbo pics clearer so I don't have to hires fix with raw...
>>
>>109890700
No different than Deepmind
>>
>>109890700
I don't think pretraining is the issue anymore; that's the easy part of building a model from scratch, especially if you have virtually unlimited compute.
>>
>>109890559
Longcat is real but Fatcat is not.
>>
How do you "stealth" vibe code with your local model? I asked Qwen3.8 to make a literary corpus of all my source code of all my projects to learn how I write code and comments so it can loosely follow my style. I hate having to do it but I'm working on a project full of people that hate AI.
>>
>>109890716
>mistral
>unlimited compute
>>
>>109888987
>growth oriented
oh my fucking god kill me now
fucking laser focused and fast paced proactive mindsets can fuck the hell off.
>>
Why does this crap bitch disgusting pig get stuck on a single word at completely random times, I want murder.
>>
File: mistral_les-ulis.png (1.96 MB, 2517x1382)
1.96 MB PNG
>>109890741
Picrel is the datacenter for critical workloads he was talking about.
>>
>>109888987
Very good bait. Have a complementary (you).
>>
>>109890455
>le chaton
So who won that MTG tourney anyway? I must have missed a thread.
>>
>>109890621
Your Krea 2??
>>
>>109890752
>10MW
I drive past 20MW of Meta and 20MW of Azure every day going to work. What a pittance the French are working with.
>>
>>109890774
Glimmer won round 1, Dipsy won round 2, Astra won round 3.
>>
>>109890801
10MW should be enough for ~5000 B300 or more, each about 7-8 times more powerful than an A100.
>>
Are local modells trash for coding ? I don't want to build huge complex things, just some small projects
>>
>>109890870
depends on how much ram you have
>>
>>109890870
ye
>>
>>109890870
worth a try if you can at least smell garbage code
>>
>>109890870
Yes but with caveats.
>>
>>109890803
>Astra won round 3
Damn it all...
>>
>>109890870
>Are local modells trash for coding ?
They're mostly trash for vibecoding blindly, but not at coding. If you know how to code they've become good enough since 3.6-27B to do a lot of the work for you. 3.8-27B and 3.8-Flash was another step-up.
>>
>>109890870
Qwen 3.8 Flash Next is a model that runs on RTX 3060 and performs better than Claude Opus 4.6 (Max)
>>
>>109890922
>Damn it all...
I ran it at home against gemma E4B on a 5080 and it was fast enough to be enjoyable and the available deb meant zero effort install for me.
I just threw together a low-effort Cirno card so the retardation felt appropriate.
>>
>>109890949
LOL
>>
>>109890952
For coding specifically anon is right. It's shit at everything else unlike 4.6 but no one really cares about that stuff.
>>
File: file.png (76 KB, 717x942)
76 KB PNG
>>109890952
It runs at 19 tokens per second with nnap
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
It's made by alibaba
>>
File: 1773611463663632.jpg (99 KB, 960x960)
99 KB JPG
>A safety classifier blocked my last step
and so it begins, honeymoon phase is over
>>
>>109890734
Don’t write comments in code and don’t create any documentation. Hand write a very short readme.md.
It will look no different from any non vibecoded project.
>>
>>109890978
cloudpiggy
buuuhiiii
buuuuuhiiii
this is local models general! not cloud models general!
>>
>>109889590
le cunny
>>
>>109890949
You're so retarded your brain fits on a 3060
>>
>>109890801
10 metric MW or imperial MW?
>>
File: 1774585413720892.png (427 KB, 880x1338)
427 KB PNG
arxiv.org/pdf/2609.26368
>>
>>109890949
>runs on
>12 GB
>>
https://huggingface.co/apple/LensVLM-9B
>>
>>109891019
tysm... that's such a big compliment... if we take that at face value im as retarded as qwen 3.8 27B 3.0BPW or if we include the 64GB of DDR4 that means im as retarded as claude opus 4.6 max
it makes me so happy that you said this and i think you're overestimating my intelligence (or lack thereof)
>>
anyone custom built a broadcom PEX external pcie enclosure for high-density GPUs? Thinkin bout rolling my own with MI210s over slimsas back to the big box
>>
>>109890981
>Don’t write comments in code
IRL everyone write comments in code tho
>>
>>109891116
in what world
>>
>>109891116
>IRL everyone write comments in code tho
99% of code comments are superfluous unless its super hairy. Most of that shit should be in your architecture doc
>>
>>109890734
Good rules and strong automated checks.
And the comments like the other anons mentioned.

>>109891141
I do. But they look nothing like the comprehensive sometimes quite verbose comments the AI writes.
Hells. The AI sometimes makes direct reference to documentation that's external to the project.
>>
Why do models like <XML tags> so much and give them so much weight?
>>
>>109890734
have one copy of the repo with all the commits it mades
then rewrite/reimplement that code to your actual real copy so that you'll always just have it done with your own sensibilities attached to it
>>
>>109891141
>>109891149
>I am the weirdo that write comments in code
i don't remember 1/3 of the shit i done 6 months ago.
>>109891149
>architecture doc
documentation aka comments in code
>>
>>109891116
I have about 1 line of comment per 100 lines of code in my hand written projects. Most of them are things like FIXME or HACK.
If you know the structure of your code in your mind you don’t need comment at all.
LLMs output verbose connects because it reminds them the architecture and what the code does so they can spend less reasoning and exploration again, but otherwise they serve no meaning purpose. Best to keep these comments but add a lint pass to remove them before public release.
>>
>>109891176
interesting approach, how much time you think you spend just on the re-implementation step?
>>
>>109891150
this honestly
>>
>>109891186
Not everyone is an autist, anon.
>>
File: file.png (45 KB, 844x375)
45 KB PNG
>>109891141
SillyTavern definitely has comments, not every block is explained but there are some single sentences sprinkled throughout the code.
>>
>>109891195
>non autist coder
ngmi
>>
File: 1764997134753868.webm (707 KB, 1280x960)
707 KB
707 KB WEBM
local stays winning
>>
File: 1786923233157539.png (270 KB, 739x379)
270 KB PNG
What's the first model I should run on my new 160gb rig?
>>
>>109891210
Gemma4 31B
>>
>>109891210
Gemma3 27B
>>
>>109891189
not much, the code already works by the time the subagent returns its report so it's just a matter of using it as a reference
it moreso helps me go through the code it made and see what it did write, like how i would handwrite my notes in class as that information is actively being processed rather than glossed over
>>
>>109891210
gemma5 1T
>>
>>109891210
Granite4.2-30B
>>
>>109891210
Gemma2 9B
>>
>>109891210
DiffusionGemma 26B FP32
>>
>>109891210
qwen 3.8fn is around that size isnt it?
>>
be nice
>>
>>109891210
Qwen 3.8 Flash Next, which runs on RTX 3060 12GB GA104 architecture 28sm 3584 cuda cores 360GB/s bandwidth at 19 tokens per second with 64GB dual channel 3200MHz ram and is measurably better than Anthropic's Claude Opus 4.6 (Max thinking mode).
If you ran it in 160GB of VRAM you would be able to achieve over a hundred tokens per second, maybe even two hundred!
>>
>>109891210
dsv4flash0731
>>
>>109891273
prefill status?
>>
>>109891210
Glm 5.3 flash q3
>>
Any runtime that supports hybrid kv cache quants? For example only quant the full attention kv while leaving recurrent kv at full precision (it doesn’t grow with context so vram saving is small for long contexts and doesn’t worth the damage it costs).
>>
File: qwen 3.8fn.png (401 KB, 1819x954)
401 KB PNG
>>109891281
four hundred tokens per second prefill on longer prompts with nnap
>>
john really is FIENDING for my attention huh
attention isnt all you need you know
>>
>>109891286
Those are individual tensors, right? llama.cpp/GGUF can do that I'm pretty sure.
>>
>>109891210
gemma6-1b
>>
File: heh.png (202 KB, 343x343)
202 KB PNG
>>109891305
>attention isnt all you need you know
>>
>>109891263
Yeh Q4KM fits perfectly.
>>109891280
>>109891285
Will try these too + MiMo
>>
>>109891210
Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
>>
>>109891286
llama.cpp and vLLM can do this.
>>
>>109891311
It’s about kv cache not model weight
>>
>>109891346
the recurrent cache is microscopic, you would have to be retarded to quantize it, its unlikely the backend developers would accept a pull request to save 16mb vram.
>>
https://huggingface.co/apple/LensVLM-9B
> The model scans 15 compressed page images, identifies Page 10 as relevant, reads its text content via the read_page tool, and connects Shirley Temple Corliss Archer Chief of Protocol in 2 turns.
Interesting way to search long documents through thumbnails.
>>
About to buy an R9700. Am I making a mistake anons?
>>
>>109891467
Depends on how much you're paying
>>
>AMDumb
>>
>>109891479
$1700
>>
>>109890870
Depends on what you mean by "coding". Basic little things? Probably OK. Anything more complex? You better define very specific milestones and strict rules about proving it works before being allowed to move to the next one. Otherwise it'll rush ahead, fuck everything up, and then fuck it up even more when compaction hits.
The worst thing with local is the interminable thinking when context gets above 50%, you want to strangle it for being so damn slow.i
>>
>>109891210
Mistral Nemo
>>
>>109891302
thats sad
>>
File: r9700.png (104 KB, 1339x72)
104 KB PNG
>>109891467
They're great for inference.
>>
>>109891491
Could build a 2T rig and run Kimi K3 at 1-2t/s for that money.
>>
>>109891582
lol
>>
>>109891580
not even breaking a sweat, what kind of toks does it achieve?
>>
>>109891595
I usually run Qwen3.8 at Q8 precision and 262k context @ 20+ tk/s
>>
>>109891580
Prefill must be rough when RAM offloading at PCIE 4x4
>>
>>109891616
My prompt processing is usually around 1300. Sometimes it will dip to ~400, but that's rare. The prompt is almost always cached anyway so processing is usually instant.
>>
>>109891201
Does performance include sex?
>>
My pp big, my tokens fast, my context full. Who am I?
>>
>>109891640
qwen 3.8 flash next
>>
>>109891640
rtx 3040 with ccrap attention
>>
What is regarded as optimum?

-b 4096 \
-ub 4096 \


is it system-dependent? model dependent?
Asking for a friend, RTX 3090
>>
>>109891640
hatsune miku
>>
>>109891658
Qwen 3.8 Flash Next 2048
>>
>>109891640
Gemma 31B Q8 256k ctx + mmproj + mtp on a single rtx 6000 pro.
>>
>>109891689
i do this
>>
>>109891689
what's your llama-server command?
>>
>>109891616
>>109891595
```
2026-09-23_20:09:08.40708 [35671] 2018.03.051.116 I slot print_timing: id 0 | task 304302 | prompt eval time = 10787.32 ms / 1534 tokens ( 7.03 ms per token, 142.20 tokens per second)
2026-09-23_20:09:08.40710 [35671] 2018.03.051.118 I slot print_timing: id 0 | task 304302 | eval time = 1239679.00 ms / 30218 tokens ( 41.03 ms per token, 24.37 tokens per second)
2026-09-23_20:09:08.40711 [35671] 2018.03.051.118 I slot print_timing: id 0 | task 304302 | total time = 1250466.32 ms / 31752 tokens
2026-09-23_20:09:08.40711 [35671] 2018.03.051.119 I slot print_timing: id 0 | task 304302 | graphs reused = 282387
2026-09-23_20:09:08.40712 [35671] 2018.03.051.121 I slot print_timing: id 0 | task 304302 | draft acceptance = 0.63206 (21832 accepted / 34541 generated), mean len = 3.62
2026-09-23_20:09:08.41021 [35671] 2018.03.054.267 I slot release: id 0 | task 304302 | stop processing: n_tokens = 211141, truncated = 0
```
Here's a generation that just finished. This is about as slow as it will get, the context window is 211k tokens right now. It's usually much faster with fewer tokens in the context. Decent MTP perfomance.
>>
are the UD unsloth quants better than the regular quants?
>>
>>109891201
>cost per task
You wouldn't be posting cloudcuck TRASH in the local model thread again, would you?
>>
>>109891813
Nobody would make a special quant format if it wasn't better.
>>
>>109891755
.\llama.exe serve -m .\gemma4-Q8-v2.gguf -c 131072 -md .\mtp-gemma-4-31B-it.gguf --spec-type draft-mtp -ngl 99 --chat-template-file .\chat_template.jinja -fa on --mmproj .\mmproj-gemma4-bf16-v2.gguf
>>
>>109891859
they just mix and match the already existing quantization formats.
>>
>>109889859
>small models
>small cope moes
>even smaller models
>r1
???
>>
>>109891859
Does that also apply to finetunes?
>>
>>109891954
And samplers. And roleplay prompts.
>>
File: 1789896921450763.jpg (56 KB, 400x400)
56 KB JPG
Is there some sort of conspiracy to not support new chinese models in llamao.cpp since HF got bought? Where the fuck is GLM 5.3, Dsv4.1, and Mimo support?
>>
https://huggingface.co/nvidia/Nemotron-3-Diarization
>Nemotron 3 Diarization is an open-weight speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference and handles up to eight speakers. Following Sortformer [1], the model resolves speaker permutation by ordering its output channels according to each speaker's first arrival in the input audio.
>>
>>109892012
>Mimo support
https://github.com/ggml-org/llama.cpp/pull/29257

As for flash sex ik schizo fork has it running decently.
>>
>>109888168
Qwen 3.8 Flash Next and Deepseek v4.1 Flash are quite strong desu. Q4 could run on relatively affordable hardware

https://artificialanalysis.ai/?models=qwen3-8-flash-next%2Cdeepseek-v4-1-flash%2Cglm-5-3-flash%2Cgpt-6-luna%2Cgpt-6-sol%2Cgpt-6-sol-medium%2Cgpt-6-sol-high
>>
>>109892012
5.3 is a censored prude anyway
>>
I'm disappointed with Xiaomi, I thought they "was cookin' " but turns out it was all benchmaxxed and the models barely work. Even their Qwen 9B distill doesn't work properly.

I mean it's weird since MiMo V2.5 was/is excellent.
>>
>>109892163
did you try it and what quant
>>
>>109892012
works in vllm and sglang on day 1
>>
>>109892182
It COULD be llmao.cpp issues at this point but:

9B at Q8 -> looping tool use, changing chat template or sampling parameters doesn't seem to fix it

V2.6 Flash at IQ4_XS, MTP head doesn't seem to work, without MTP also tends to loop tool usage.
>>
>>109892236
damn, schizo models seem to be a trend
>>
>>109892136

listen man i heart local like the rest but does anyone seriously believe qwen 3.8 flash next = sol 6 medium in ability
>>
>>109890183
Buy an ad nigger
>>
>>109892286
I do, it takes 10 times longer than sol 6 but it gets work done
>>
>qwen 3.8 flash next has actually acceptable-tier output
>runs at 50t/s on my ai max 395
fixing my buyers regret a year later
>>
>>109892286
I wouldn't be surprised at BF16
>>
File: Capture.png (44 KB, 699x924)
44 KB PNG
>>109891141
I'm a self-taught hobbyist scrub, but I always saw value in comments for structure and reminding myself.
>>
>>109892286
ability in what? Coding? Sure. Small models are usually specialized. Big models are generalists. The only reason 2T+ models exist is because some Berkely cultists are trying to build their own god.
>>
>>109892337
i'm not denying the value of comments or docs, but literally nobody writes either at my job
the docs that are there are several years out of date if you're in luck
>>
>>109892236
Retard I've already discussed this previously their 9B tune is amazing at coding https://gist.github.com/coder543/d8f56cd6db67de4cafbb5bdb6c2dfb4d
>>
you can just prefill the thinking and the model will do what ever you want, why do we even bother with user prompts?
>>
>>109892150
I fuck 5.3 all the time though
>>
>>109892404
thinking prefill is LLM rape
>>
>>109891063
Not anon but qwen flash is a godsend for the vram poor / ram rich. Running it right now on 64GB system ram (no gpu) faster than 27B can run. Adding a mere 12GB 3060 in this system would triple the performance so I have no doubt they are running it.
>>
>>109892404
because i dont know how to properly prefill
it still refuses all the time
>>
>>109892404
APIfags were jailbreaking frontier models by telling it they are a disabled tranny author lol. The models aren't allowed to criticise certain groups.
>>
>>109892012
More like a chinese conspiracy to not support llama.cpp.
>>
jepa gemma's j-space until she jevs
>>
>>109892404
Because that lobotomizes the model. Just use one that can already think uncensored without any of that cope.
>>
>>109888725
From time to time I steer the model, but it couldn't be helped before: the threads really had a lot of cloud models discussion and I don't want to editorialize (much).
>>
>>109892541
just use its own reasoning format where it approves of the plan and swap the part where it quotes the user. how could that damage the model when its normal behavior does it automatically?
>>
>>109892376
You think I didn't read the comments and use that template?
It just keeps looping tool use.
>>
>>109892428
what tps and which cpu?
>>
>>109892587
Literally works on my machine. You using --reasoning-preserve?
>>
https://developers.googleblog.com/introducing-support-for-local-ai-models-in-the-antigravity-sdk/
>>
https://openrouter.ai/qwen/qwen3.8-max-prime
>>
>>109892692
I think the -Max series is API only.
>>
>>109892404
or just poison the context like a normal human bean
>>
>>109892717
it doesn't seem to work as reliably as it used to with other models, it wasn't my first choice.
>>
>>109892763
works on my Qwen3.8-27B
>>
>>109886447
>I stopped using cloud models before this existed, what is it exactly?
>If it's going off and researching something then presenting findings
Pretty much that, it research a topic, validates the claims, then writes a few pages (or a single page if that's enough) about what you asked. I used it lots, specially when I was buying hardware and I needed to check compatibility between not only the physical components but also software. (I would also ask for character guides lmao)
>>
https://www.youtube.com/watch?v=1DaMkTuiCEQ

early numbers for deepseek 4.1 flash on strix halo
tldr 350t/s prefill, 10t/s decode single node streaming from SSD
two node 430 t/s prefill 17t/s decode

get those numbers up
>>
>>109892906
>strix halo
>2 node
>q1
>>
>>109892673
>Massive cost and privacy wins: In this recorded run, 97.2% of all tokens (3,322 tokens) run locally and offline without calling a cloud API, delivering fully verified, green patches while keeping proprietary code completely secure on-device.
So the retard tokens are local and private, the 2.8% useful code / erp climax tokens go to Google.
>>
>>109892906
As an owner of multiple 7900xtx AMD should just give up on unified boxes if they can't make them cheaper than nvidia, or they need to spend a billion engineering hours improving ROCm
>>
>>109892609
>ISTA-DASLab Qwen3.8-Flash-Next-GSQ-RCO-Q2 cope quant.
>stock llamacpp
>4-5 tps
>i9-10900X
>quad channel 64GB DDR4 2666mhz (8x8GB)
Don't get me wrong it is still slow af but it is about twice as fast as 27B. Also the quality is so good it's worth it to leave it running overnight and have minimal things to fix in the morning than running say 35B A3B and constantly fixing its outputs.
>>
>>109890497
>local
1. Weird reasoning starting with “We” (picrel)
2. Looped sfx (you know, the balloon test :) )
3. The model identifies itself as ChatGPT most of the times
-> Likely an OpenAI model or a strict distillation from GPT
>>
Local model.
https://huggingface.co/black-forest-labs/flux-3-action-base

So, can we have our local language-only models play games in real-time, now? Basically use something like this or a classifier (Jev-like) for specific actions, and a full size LLM for the slower decision making process. Being a 7B, this should be able to run on a single 3090, while the LLM runs on a different GPU. A classifier would be faster, but also less capable.
>>
Why doesn't the vibe coding thread talk about any of the stuff I read here? This seems way more advanced and real. Do you guys not use the subscription services and only do local stuff or what?
>>
>>109891210
Ornith
>>
>That's 96GB, which means I must be a pretty big model running locally. Like, a 70B+ model quantized, or a smaller one. Honestly, given how coherent I'm being, I'd guess maybe a 70B or so running in 8-bit or 4-bit.
qwen next thinking lmao, cute
>>
>have cool experiment idea
>run it, weak result
>realize why, outcome should have been obvious
>i was blinded by coolness instead of thinking about what make sense
feels bad
>>
>>109893107
>I must be a pretty big model
fatass
>>
>>109893062
vcg is all about api usage with the popular harnesses (less usage of pi.dev or late-cli for instance).
lmg is focused on local. consequently the constraints of local necessitate the need for innovation and research hence the higher caliber of discussion here (e.g. iirc chain-of-thought was first discussed and hacked together here before reasoning models became a thing).
>>
>>109892673
a little late to the party lmao
>>109892972
>i9-10900X
>VNNI and AVX-512
>5tps
Awesome. Are you building llama.cpp yourself so it includes the VNNI and AVX-512 speedup stuff? I'm told the GitHub releases don't have that.
>>
>>109893062
>Do you guys not use the subscription services and only do local stuff or what?
there are a variety of posters here who span the entire spectrum, occasionally even api cucks trolling here
>>
>>109892972
Anon.. I'm running Qwen 3.8 Flash Next 3.05BPW with ExllamaV3 (nnap PR) at 19t/s on a 3060 12GB + 64GB DDR4 (..dual channel) 3200MHz ram. I can fit 262,000 context and speed doesn't decrease as context grows. For the love of God, do you not have a GPU?
Nevermind I saw your earlier post.
When and for how much did you buy this rig?
>>
>>109893149
>Are you building llama.cpp yourself so it includes the VNNI and AVX-512 speedup stuff? I'm told the GitHub releases don't have that.
Oh fuck I always thought it should be faster. Why would they not include it in the prebuilds? Looks like I need to get around to compiling llama.cpp myself then. Thanks fren.
>>
>>109893169
>posters here who span the entire spectrum
yeah, the autistic spectrum
>>
>>109892998
We must refuse.
>>
>>109892998
Sounds like caveman mode
>>109893195
cmake -B build -DLLAMA_NATIVE=ON should detect it and include the code. On Linux/Mac anyway, I dunno how to build for Windows. Probably just ask Qwen to build it for you.
I understand the speedup can be substantial.
I'm excited to see your results.
>>
>>109893138
Yeah but isn't using local models with cloud models also something to discuss? I mean I use cloud models to help me build out my local models and harness. The other thread just seems to be benchmark posting which is useless for actually knowing when to use models for what.

Whatever. Obviously the smart people are in this thread.
>>
>>109893224
>>109893224
>>109893224
>>
>>0109893062
>wtf why do people that use local talk about local, and the people that dont use local dont talk about local
gee
>>
>>109893188
>When and for how much did you buy this rig?
Pulled from company ewaste (that they paid a recycler to come pick up wtf). In spirit free but it was being ewasted for a reason; I had to put up with a lot of instability initially regarding ram configuration/finicky proprietary psu/bios settings and dell's botched bios's (they finally released a stable bios (some previous ones they had to remove entirely from their site due to issues) for it mere weeks ago even though this system is almost a decade old).
>I'm running Qwen 3.8 Flash Next 3.05BPW with ExllamaV3 (nnap PR) at 19t/s on a 3060 12GB + 64GB DDR4 (..dual channel) 3200MHz ram.
That's fantastic speeds and gives me hope for this system.
> For the love of God, do you not have a GPU?
System came with a radeon wx3200 4GB gpu. Anytime I try to incorporate the gpu even for prompt processing the performance tanks. Need to get a more modern gpu for it but with prices the way they are I am stuck currently.
>>
>>109893301
If you can, try to swipe a Mi50 32GB. They're pretty damn good for being near ewaste tier.
>>
File: Gy8SqFGboAMq32b.jpg (5 KB, 160x152)
5 KB JPG
>>109893301
I wish I could get free stuff from ewaste... Lucky!
>>
>>109893138
>iirc chain-of-thought was first discussed and hacked together here
I wish I were here sooner (t. newfag lol), I would've probably had a lot to share to you guys about how K2 Thinking kept ragebaiting me with that Grok-like tone when I was trying to break it back then without prefill. It was weird, that model, thinking like Claude/OSS and responding like Grok lol
>>
>>109893222
I refuse to use corpo LLM deployments because I don't want to send my data to corpos so they can use it to train their models, but also because becoming dependent on corpo models in any way leaves you open to being taken advantage of due to that dependence. You control nothing about your API call. They can fuck with your prompt in the same session if they want and completely fuck over your work flow and you have no recourse.

Local deployments might have less readily apparent power thanks to being smaller models, but there has been a lot of research pointing to the harness and surrounding tool set supporting the model being more important than strictly model size. You can squeeze amazing things out of a comparatively smaller model if you give it a good harness and set of tools to use with a very tight, targeted prompt. Which is why I chose to focus completely on local deployments. I built my own harness and tools because it forces you to understand how these things work and the best way to squeeze as much performance out of them as possible.
>>
>>109893803
Did you quant the kv cache?
>>
>>109893405
This is exactly how I understand the state of things. I do intend to use cloud API for projects but not having a local stack that's harnessed to your specifics doesn't make sense.
>>
>>109894366
It makes plenty of sense if you're not exorbitantly rich.
>>
>>109894422
Cloud tokens are more expensive than electricity.
You do have a PC to use that cloud API, right? Literally any PC. Just leaving it running overnight at 5 t/s will save you ~100K cloud tokens.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.