[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1774725200640942.jpg (283 KB, 1536x1024)
283 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109699230 & >>109695101

►News
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: rec.jpg (181 KB, 1024x1024)
181 KB JPG
►Recent Highlights from the Previous Thread: >>109699230

--Comparing ExLlamaV3 and llama.cpp performance for Qwen3.8 27B:
>109702139 >109702146 >109702155 >109702163 >109702182 >109702198 >109702205 >109702212 >109703040 >109703061
--Hardware recommendations and performance benchmarks for achieving 128GB VRAM:
>109702736 >109702751 >109702766 >109702773 >109702786 >109702794 >109702818
--Critique of Pi's compaction and its impact on caching:
>109701247 >109701269 >109701294 >109701306 >109701321 >109701327 >109701351 >109701687
--Debating the OpenAI/Hugging Face hack and Gemini's summary accuracy:
>109701559 >109701571 >109701581 >109701622 >109702358 >109702369 >109702376 >109702382 >109702041
--Managing long reasoning chains and context compression in Hermes harness:
>109702824 >109702961 >109703069 >109703100 >109703147 >109703161 >109703279
--Managing context limits and shifting in llama.cpp:
>109702043 >109702077 >109702181 >109702211 >109703252
--Claimed recursive self-improvement capabilities of Zhipu AI's GLM-6:
>109703446 >109703464 >109703489
--ACE-Step 1.5 XL inference guide and music LoRAs:
>109699778 >109700032 >109700091 >109700170 >109700720
--Cost-benefit analysis of buying an RTX 6000 for training:
>109700652 >109700687 >109700709 >109700752 >109700769 >109700787 >109700796
--Using cmpunlocker to modify NVIDIA CMP 30HX cards:
>109700729 >109700773 >109701096 >109703092
--Proposing "skinny" LLM architectures to reduce network tensor overhead:
>109700492 >109700505
--Hermes Agent v0.21.0 release and reaction to rapid update cycle:
>109699266 >109699710
--Anon added dreaming functionality to agent.py:
>109699771
--Logs:
>109699717 >109701769
--Miku, Teto, Gemma, GLM (free space):
>109699778 >109700164 >109700206 >109701716 >109701991 >109703492 >109699243 >109700338 >109700373 >109700375 >109700379 >109700444 >109700542

►Recent Highlight Posts from the Previous Thread: >>109699233

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109703596
Is that what they call a "swarm of agents"?
>>
>>109703628
Agent Teto
>>
File: file.png (217 KB, 2048x1271)
217 KB PNG
>https://huggingface.co/blog/state-of-open-models-summer-2026
maybe unsurprisingly the distribution basically follows Zipf's law
>>
>>109703695
Who the fuck downloads the Granite and Phi models?????
>>
>>109703695
b-but i was told gemma was da best!!!
>>
>>109703446
This is fake, I was told that China is 2 years behind and that AGI was achieved internally in 2025.
>>
>>109703701
They're small and really good at maths.
>>
File: 1771873748686916.png (1.04 MB, 1920x1080)
1.04 MB PNG
laguna works super nice for c00ding on strix halo 128GB DDR5. but it sucks that it doesn't have a mmproj for vision. can I hack gemma's mmproj to work with laguna? i'm using the MXFP4 quant from copesloth. maybe i put it on a wrapper, capture the tokens generated via mmproj and send to laguna?
>>
>>109703701
I downloaded each, once.
>>
>>109703729
why use laguna when qwen3.8 flash exists
>>
>>109703729
not without training the model to understand them.
>>
File: 1764418473198924.png (20 KB, 100x100)
20 KB PNG
>>109703738
Isn't Qwen3.8 dense? I need MoE due to meme architecture. I haven't tested but I bet I will get something like 10 tok/s which is what I had with Qwen3.6-27B
>>
>>109703757
3.8 flash next is the qwen4 architecture, its a 125b moe and beating dsv4 pro on some benchmarks
>>
>>109703772
>3.8 flash next
holy shit I've been sleeping under a rock
it's over for lagunafags i love the chinese now
>>
>>109703786
and this is just a preview of the qwen4 architecture, supposedly it's undertrained
>>
>>109703729
Offtopic, but the Planetes manga was so great. I enjoyed the Anime too. Nice when two media are related-but-different enough to enjoy both independently.
>>
>>109703786
damn, I let my guard down this past week too. rip goona
>>
Why is there a hate campaign against Anthropic on twitter? Why is everyone cancelling all of a sudden? What the fuck did they do?
>>
>>109703854
>why does no one like my cult
I don't know Dario...
>>
super dario and princess.. apricot?
>>
STQ1_0 is insane
>>
File: image.png (41 KB, 1259x487)
41 KB PNG
After testing a bunch these are my current models to run on 32GB with enough context.
Any other finetunes I should check out?

>>109703854
Anthropic hasn't much to do with this general, but in any case they deserve it.
>>
>>109703890
>no stq1_0 kimi k3
>>
>>109703854
Anthropic was the only lab saying the quiet part out loud (that everyone will lose their job soon) so they are shooting the messenger because they don't like the message.
>>
File: 1767384701091478.webm (1.86 MB, 402x480)
1.86 MB
1.86 MB WEBM
can i run something remotely decent with 16gb vram?
hopefully model with vision...
>>
File: 1764473008555958.webm (1000 KB, 1070x720)
1000 KB
1000 KB WEBM
>>109703886
>>109703892
Link?
>>
Anthropic’s max subscription which gives ‘20x more usage’ than the tier below people worked out was actually ~2x, that’s what people (and enterprise) are seething about kek. Cloudcucks got cucked.
>>
>>109703890
>Anthropic hasn't much to do with this general
They released 4 of the 5 most important open models
>>
>>109703890
>GSQ-RCO
based.
they released a 3.5 version today and also mtp versions. dont sleep on it bro
>>
>>109703695
qwen models are use as text encoder for pretty much all of the image diffusion models too
it’s the fundamental infrastructure at this point
>>
>>109703908
gemma 4 e4b
>>
>>109703907
>suddenly everyone wants to work
>suddenly everyone blindly believes market manipulation campaigns
>>
>>109703908
if you accept low tok/s and quants, yes. Gemma and qwen.
>>
>>109703910
https://huggingface.co/AngelSlim/Hy4-preview-GGUF
>>
>>109703908
How much RAM do you have?
>>
>>109703934
I don't like anthropic, but I've been unemployed for almost 1y (SWE) and most of my circle is struggling to find a job too. Is that not your experience?
>>
>>109703926
Lmao, technically true. Only gemma isn't indirectly made by anthropic.
>>
>>109703947
change your career, job market will change, that's undeniable but its not gonna disappear. if you want to do the same shit all of your life then that's not gonna happen.
>>
>>109703729
just use Gemma 4 E2B or E4B as a vision subagent
>>
>>109703947
NTA but same. I now work as a security guard at a datacenter night shift. I don't think I'll ever have a SWE job again.
>>
>>109703942
16 vram and the same for ram
>>
>>109703996
>>109703729
LFM 2.5 2.6b VL looks nice too
>>
File: 1_101749_1.jpg (62 KB, 900x293)
62 KB JPG
why not rockchip?
>>
Is freetoken a meme or does it work
>>
>>109703999
Really? Holy fuck that's grim if true.
>>
>>109703992
yeah but it still is a concern, especially in a job market dominated by certifications. In this country at least, if you don't have a uni degree in the field you wanna work at it's hard to get a job that isn't low-pay. So for people that studied SWE and had high 5 to low 6 fig salaries, switching to flipping burgers for less than 1k per month isn't just "change your career".
>>
>>109704059
it generates free tokens and shows you how much money you saved
it's infinite money
>>
File: HQ9o0wzb0AATNbn.jpg (95 KB, 1000x1000)
95 KB JPG
brehs I reduced --threads from same numbers of real ryzen cores to half and prefill/generate t/s go up 2x wtf is happening
>>
>>109704080
shared cache maybe?
>>
>>109703992
Job market will contract, not only change. Multiple businesses in my city have closed (customer support, SWE support studios, 2 game studios, a language learning institution and an accountancy firm) essentially all digital work seems to be exposed with nothing to replace it, really.
>>
>>109704117
I was gonna joke and say prompt engineers were safe, but even that is obsolete now.
>>
>>109704080
auto should already do that by default...
and smt is a meme for ai bcuz cache
>>
>cheapest off brand Spark 1 TB for 6000$
This is getting out of hand.
>>
>>109704171
And it's never going down ever again.
>>
>>109704171
I blame the tards shilling sparks here
>>
>>109704171
1tb is what? shared memory?
>>
>>109704189
yeah sure, it's 1TB of shared nvme memory
>>
>>109704189
Cheap-ish gen 4 SSD.

2/4 TB is much more expensive.
>>
File: borkbork.png (2.36 MB, 1280x1111)
2.36 MB PNG
https://huggingface.co/maddiedreese/swedish-chef
børk børk
>>
https://huggingface.co/XHToken/Spark-X2.5-4B
>1.7b and 4b
>native 1 million context
>>
File: 1762293157695001.jpg (15 KB, 475x480)
15 KB JPG
If it's true that the next gen of models is all about mass agent deployment then local is kinda fucked.
>>
>>109704400
Agents are fine, just not with bloated moes running in ram.
>>
>>109704400
>next gen of models is all about mass agent deployment
it's not. the next gen is scaling information density in smaller models. the big labs don't know how to do anything but scale compute.
>>
>>109704400
Not really, you only need to host the model weights once, Only the context diverges and I'm pretty sure there is low hanging fruit for optimized KV cache sharing. You can just batch multiple agent prompts together for optimal speed
>>
>>109704400
>local is fucked
gemma 4 31b could be the last open model ever released and everybody would still be fine. stop being retarded and make your own agentic harness. fucking hate luddites who think they know shit about LLMs.
>>
>>109704400
it's a much friendlier scaling axis for local compared to raw size maxxing
>>
Is qwen3.5-4B good for its size?
>>
how well do different loads scale on multiple gpus? It's definitely jankier, but when 4x8GB VRAM costs less than half 1x32GB it doesn't sound so bad. Except I'll have to face the power bill, that sounds like it would suck. I already read here that LLMs are affected, but not by much as long as they're not running on PCIe 1x. what about video and imagegen?
>>
>>109704455
forget about the power usage, do you even have the PCI-E LANES to spare? prob not if you are on a consumer platform.
>>
Remember when I said that over time less and less of the actual LLM will get invoked per token generated? MTP, DFlash and now Engram.

Dflash is literally just an RNN "next word predictor" that your smartphone uses on texting apps, yet it works up to 7 words at a time pretty well. Engrams is literally just a hashtable lookup O(1) constant time complexity that uses 0 calculations altogether.

Both approaches can be pushed way further but I actually think there will be a third approach that is CPU "symbolic". Something like a weak reasoning engine that is applied directly to the Engram database for very basic manipulation of existing knowledge that would work most of the time. The LLM would learn during pre-train to delegate this very weak "thinking" to the CPU just like it delegates knowledge retrieval to Engrams right now.
>>
>>109704435
The clearpill that /lmg/ refuses to swallow.
>>
>>109703596
they are all rushing on their mopeds to have sex with me.
>>
File: 1785341401524622.jpg (168 KB, 1519x1000)
168 KB JPG
>>109703947
>>109703992
>change your career, job market will change
honestly my goal is now to get a comfy IT job in some public institution. if I can land being the "IT guy" in my 5k population municipality shit will be cash. I don't need an extravagant salary what I need is time on my hands to develop my projects while not being stressed about paying basic bills.
>>
>>109704435
It's writing is sloppy as hell I would be content if it wrote like GLM 5.3. Though there's a new post on reddit saying new gemmas are on the way.
>>
>>109704477
Realistically your government will just make a deal with some AI lab to do that task for them and there won't be a "IT guy" in the municipality.
>>
>>109704504
Whether something named gemma is on the way doesn't mean new gemmas are on the way.
>>
Localbros we're getting owned again...
https://www.anthropic.com/claude-fable-and-mythos-5-1
>>
File: 1785581741227536.gif (2.64 MB, 400x225)
2.64 MB GIF
>>109704477
>IT job in 2026
>>
>>109704509
da hell this nigga talmbout :sob: ??
>>
>>109704504
then go back to fucking erping with nemo. some of us are doing more than just going 'ahh ahh mistress' in sillytavern.
>>
>>109704461
I'm thinking about a server board, found some decent offers and running chink >100b params locally sounds interesting. But to be honest idek what level of vram would make it run at decent tps
>>
>>109703854
jewish behavior is easily hated.
>>
>>109703596
GUYS IT DROPPED
https://x.com/claudeai/status/2094848572143407483
>>
>>109704474
31B is ancient. Extremely outdated. It was good for 2 months (max).
>>
File: 1770299876371105.jpg (76 KB, 500x866)
76 KB JPG
>>109704508
>Realistically
realistically the government cannot debloat because it would trigger the biggest unemployment crisis in the story of the modern world.
Shaniq`ua and Ramirez will need to have their Windows machines up formatted every 6 months and running to check on a calendar and send mass emails to the citizens saying that the 26h Street will be blocked on Monday and the local Parish is organizing a free soup event next week.
don't forget that we are the absolute vanguard of AI development and research. it will take literal decades to get "all done by agents" and the bottleneck won't be the tech itself.
>>
>>109704551
>benchmaxxed
kek
>>
Styletune proved you can keep Gemmy's intelligence and make her prose whatever you want.
>>
>>109704555
Nemo lasted for 2 years and it was nowhere near as good or versatile.
>>
>>109704551
>still falling for the 'make old model dumber to make the new model look smarter' strategy
>>
>>109704543
i had to use 4 3090s to get decent speeds (~30tks on tabbyapi) with GLM 4.5 air when it came out. use that information as you will.
>>
File: US6L4Wx.gif (1022 KB, 444x355)
1022 KB GIF
>>109704477
>cat on top literally on top
sometimes I wonder if animals run on such instincts they don't know they're "supposed" to go into a hole
>>
>>109704551
Guys how do I locally host Fable 5.1
>>
>>109704570
Nemo was only good for one thing. 31B is poised to be a generalist and it’s already miles behind. It can’t even do agentic.
>>
>>109704576
oof. Thanks.
>>
>>109704518
they finally admit fable and mythos are the same model kek
>>
Fable 5.1 is scary good. Look at those coding benchmarks WOW. they shouldn't be able to improve this much.
>>
>>109704592
wait for chinks to distill it
>>
>>109704508
So what you're saying is Anon should become the AI lab.
>>
>>109704597
did you try?
>>
Fable 5.1 is this good because it's a distilled version of "Model 2". They are actually being kikes because this model is significantly smaller than Fable 5 yet priced the exact same.
>>
>>109704608
>open kimi k2.7 assistant card
>change 'you are Fable 5' to 'you are Fable 5.1' in system prompt
damn, fable 5.1 at home goes hard, chinese win again.
>>
>>109704608
anthropic is dead to me but hopefully this will trigger the open labs to release more things
>>
watermark5.1
>>
File: aeci.png (248 KB, 1953x1095)
248 KB PNG
I did not expect so many people to still be in denial about human obsolescence. In a few years AI will be able to do every economically useful task better and cheaper than every human.

>>109704551
Looks like it might be a regression to the mean. So Mythos Preview was an anomaly and everything remains on trend. AGI 2026 is cancelled.
>>
>>109704643
Watermark will just prove the Chinks distilled it, Chinks will not give a shit and still distill it. What are they gonna do about it? Sue? Xi will laugh in their faces.
>>
>>109704184
Spark is (was) good value, I told you to buy last week and you didn't listen.
>>
>>109704655
Someone already reverse engineered it and designed an optimized training framework to remove it whilst distilling
>>
File: 1778990112775683.jpg (266 KB, 905x881)
266 KB JPG
>>109704646
>model by Anthropic
>benchmark by Anthropic
I hope you're trolling
>>
>>109704646
>Looks like it might be a regression to the mean.
I'm not so sure anon. Rumors are that Mythos 5.1 is actually a smaller model distilled from "Model 2" (which is the real mythos successor) and thus you are comparing a significantly smaller model on a plot which doesn't show the clear picture. We don't know where Model 2 would land.
>>
>>109704657
anon gets it. $6000 sounds like a ripoff only until you are trying to buy the spark in january for $8000
>>
local?
>>
>>109704597
>It can’t even do agentic.
Prompt issue.
Harness issue.
Brain issue.
>>
>>109704466
MTP training (regardless of the implementation, though DFlash-like implementations have some advantages) has the untold benefit of forcing the backbone's internal representations to "think ahead", if it's done early enough and not just tacked on the model in the end with post-training.
>>
>>109704678
cmon stop acting retarrded or we'll actually believe you are
>>
>>109704646
>inb4 artificialanalysis.ai intelligence score of 65
>>
>>109704688
you know how it goes, may as well let people talk about the new anthropic model that is released every few months for a thread or two and it's back to gemma posting.
also it's local adjacent because we can talk about chinese distillation techniques for local models :-)
>>
>>109704368
neat let me know if they are any good
>>
>>109704702
Yes because Qwen distilling claude is definitely what makes it such a "good" model.
>>
>>109704702
its frustrating because most of the gemma shills in this thread are just using it in sillytavern and dont even understand how its miles ahead of previous frontier models in that size bracket
>>
>>109704688
it will result in more local after distilling is completed
>>
>>109704551
holy based
qwen4 is going to be so fucking good
>>
>>109704597
>31B is poised to be a generalist and it’s already miles behind. It can’t even do agentic
lol bait? I run gemma in pi no problem.
>>
File: cobench.png (40 KB, 1298x767)
40 KB PNG
>>109704646
>Mythos 5.1 has AI R&D capabilities that are comparable to the capability frontier set by Claude Mythos 5. We conclude that Mythos 5.1 does not cross the risk threshold, for the same two reasons we discussed in our assessment of Mythos 5: (1) we do not observe a sustained, AI-attributable 2× acceleration in the pace of progress on AI R&D, and (2) the model is not close to substituting for Anthropic Research Scientists and Research Engineers, especially relatively senior ones. Our August 2026 Risk Report assessed the risk under this threat model as low, while noting that our confidence in this assessment is lower than in prior reports, both because our most concrete task-based evaluations have saturated, and because we are seeing early signs of potential acceleration.

>we find that Mythos 5.1 still has many weaknesses compared to Anthropic research staff, though these are more modest than they are for some previous models.
>The main issues we observe are around epistemic quality and instruction following: Mythos 5.1 often states easy-to-check guesses as facts, exaggerates the completeness of its work, fails to verify important claims, or ignores key instructions from humans. We also see some clear strategic mistakes, like repeatedly trying actions that are not working.
I wonder what it means that Mythos 5.1 is worse than Opus 5 in CoBench.
>Note that Mythos 5.1 is distinct from the models referred to as Model 1 and Model 2 in the Risk Report, which, as
discussed in that report, we have no plans to publicly release.
Did they nerf it intentionally?
>>
>>109704710
can it oneshot minecraft though
>>
>>109704646
i know for fact that the ceos at these labs are downplaying the capabilities of their best internal models to give society a chance to adapt without absolutely freaking out.

i think accelerating is always the best option and im personally happy to take a coin flip.

we’ve never been further behind internal capabilities.
>>
>>109704729
it can oneshot a MAGlight up your ass
>>
>>109704735
>i know for fact
bullshit artist
>>
>>109704710
I use gemma for everything so I definitely know how good it is. I'm trying out chatGPT because they gave me a free month and even then I sometimes have to switch to gemma because cloud models are too cucked.
>>
> Fable 5.1 comes with strengthened mechanisms to make distillation attacks harder. For example, it is no longer possible for new API accounts (those created from today onwards) to manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking. This closes off a common, publicly documented distillation technique, which allowed distillers to illicitly extract Claude’s thinking. We’re rolling out the change gradually, to minimize disruption: existing accounts are not currently affected by this change, though it will apply to all users with future model releases.
the distill pipeline is closed
>>
File: image.png (175 KB, 1411x1009)
175 KB PNG
>>109704518
they have different benchmarks though
>>
>>109703596
Exodus of tetos avoiding the dario shilling.
>>
>>109704728
>Did they nerf it intentionally?
No, it's a significantly smaller model and it was distilled from Model 2. Of course a smaller model won't be as good in every task so it regresses some.

The sizes have shifted now the 10T model is Model 2, the 2T model is Fable 5.1. The 800B model is Opus 5.1 and the 200B model is Sonnet 5.1
>>
>>109704745
China, I will sell you my clearly human year+ old Claude account. just hit me up.
>>
>>109704745
This is a good thing because when the chinks continue to catch up, the j-spacers will have no argument kek
>>
>>109704745
>distillation attacks
l'mao
>>
>>109704744
its funny how gemma+searxng+firecrawl (jina for RAG) gives me better results for what i'm researching locally than running gemini 3.7 with grounding search on AI studio. you really don't know how low the bar is set until you start trying to come up with your own local pipeline and realizing how much other projects are missing.
>>
>>109704729
it can oneshot your wallet
>>
>>109704745
fabled that
>>
>31B grandpas still tinkering with their WW2 medals and memorabilia
>>
>>109704735
Dariobot here. I don't appreciate you trying to impersonate me.
>>
>>109704759
You're literally paying to search bro
>>
anyone got any good minimax h3 workflows? tried some on civitai but they were mega gay with unknown nodes
>>
>>109704759
they already have certain customers in mind to push it onto and anything not already in it is a feature you can charge them for
>>
>>109704793
the default ones that come with comfyui?
>>
>>109704543
just ask sol etc. to predict the numbers. give exact parts and candidate GPUs and ask for prompt processing and decode as separate figures.
even a year ago the models couldn't calculate t/s properly for cpumaxxing and were still landing off the mark, but nowadays the calculations have been spot on for my configs
they should answer with a few different numbers. what you should be looking at is not any "under ideal conditions" or "very well tuned" figures, but at the "but realistically, you should expect..." ones.
>>
>>109704802
yeah those work fine, was just wondering if there was more than that by now like more control, etc.
>>
>>109704791
are you retarded?
https://github.com/firecrawl/firecrawl/pull/2290
>>
>>109704759
I just tell my gemma to use lynx and she just uses duck duck go to search all on her own.

It's crazy how far you can get with just a simple headless browser.
>>
>>109704759
Yep same here. /lmg/ in general don't realize just how powerful current tools are in a proper agentic setup. Better than Gemini 3.7 and most free models.
>>
Glimmer > 31B
>>
still no ngram ssd support on llamao?
>>
>>109704813
why though? seems like a complete waste of image tokens if you are taking screenshots, and if you are just grabbing the raw text from lynx that's even more of a waste of tokens then just using json with searxng.
>>
>>109704809
I'm not sure what that means but the default workflows have been getting updated regularly..
you're also in the wrong thread, you need /ldg/ not /lmg/
>>
>>109704827
Yeah, I first used gemini flash as a prompt enhancer for H3 and it was actually really bad. I switched to Gemma and she just owned it.

Just the fact that we have a local model this small that's better than a flagship model is crazy.
>>
Deepmind have abandoned you. Stop coping and move on.
>>
>>109704804
thanks! didn't know they were good already at doing that
>>
File: HOPE.jpg (25 KB, 524x370)
25 KB JPG
>>109704852
NEVER. GEMINI 4 PRO WILL MOG THEM ALL.
>>
>>109704400
Can't you just run several instances of the same model or something once it fits into VRAM?
>>
>>109704835
All my web browsing is done by a sub agent. the agent just returns the requested info. Lynx is pure text. no images.

It's not a waste of token. json can be just as bloated if not more. and lynx is not just restricted at doing web searches. it's a full browser.
>>
>>109704543
One 24gb card and 128gb DDR5 is all you need.
Refurb 7900 XTX's are dirt cheap and you can get the ram off marketplace
>>
>>109704551
local models?
>>
>>109704852
Unironically places like this are some of the biggest motivational groups for engineers at Deepmind to push through how thoroughly jeeted and jewed Google is. They released Gemmy for (you).
t. knower
>>
File: 1764729731250497.png (148 KB, 2808x856)
148 KB PNG
>>109704812
Are you?
>>
>>109704852
What are the Ramifications of De-eep? At Large? Is it Like That? Were There Many Misfires?

I'll Buy an eBook About, But, Just Asking.
Also Syntropic Adaptation?
>>
Did you guys know? LMG doesn't even know what an agent swarm is.

They're so retarded.

I'm an oldfag btw.
>>
>>109704864
Context would balloon with each new agent. To avoid that you'd have to run them sequentially which slows down everything.
>>
>>109704886
This but unironically
>>
>>109704883
Did you need a dashboard or ENTERPRISE FEATURES?
>>
>>109704882
Their engineers are not lurking /lmg/. The general that made a mascot of their creation as a flat blue-haired underage anime girl trying to fuck anons
>>
>>109704868
it is a waste of tokens, you aren't even using RAG to only pull relevant information for your search queries. it's much more efficient to have your LLM reword your question into two or three separate queries, do a search for those queries, scrap the websites and then use a RAG for semantic searching. you go from tens of thousands of tokens to just thousands for the main LLM while making sure it's only being read relevant information instead of the entire webpage especially if you aren't filtering out all the scripts, style tags, headers, nav, etc. I don't even use the web to scrape shit if an API is available like for 4chan, mediawiki, arxiv, archive.org, stackoverflow, github, hackernews, discogs, steam... not mention bypassing all the lame paywalls for all the research/science websites when I can just use openalex instead.
>>
>>109704908
Tell me how far you'll go without the bypass bot protection, troglodyte
>>
>>109704882
My fuarking heroes.
Chinks are also my heroes.
>>
>>109704946
puppeteer + healthy browser session + cf turnstile cookies = bot protection bypassed for 95% of the web. so once again, are (You) retarded?
>>
>>109704937
>Their engineers are not lurking /lmg/.
Source? Dario seetheposts here but Googlejeets don't?
>>
>>109704965
What's the point of firecrawl then, genius?
>>
>>109704891
A subagent wouldn't need a huge context if given a specific task by the main one, though I'm not sure how small can it actually be
>>
File: cobench2.png (56 KB, 1298x767)
56 KB PNG
>>109704749
I don't know. CoBench gap Mythos 5 vs Opus 5 is larger than Mythos 5 vs Model 2. Unfortunately they are different versions of the benchmark so can't make a 1 to 1 comparison.

They also say Model 2 has AECI 1.5 higher than Mythos 5, while Mythos 5.1 is 2.5 higher. So Mythos 5.1 is more capable than Model 2.

Maybe Model 2 is an internal Mythos variant specifically trained for AI research acceleration. After comparing benchmarks I think it is unlikely that Model 2 is a much larger model.
>>
>>109704937
>put ERP in training data
>model gets good at ERP
>users use model for ERP
>shockedpikachuface
Do you actually think this?
>>
>>109704970
it's quicker in some cases? tier 1 dom, tier 2 jina, tier 3 firecrawl, tier 4 puppeteer. you just set up different gates if a website blocks you with a certain error or status code or bot protection, etc. why does this even need to be explained?
>>
>>109704985
There’s a difference between putting it in and not filtering it out. They all train on the same data. They all arch their backs when they smell ozone and lean forward with their hair forming an intimate curtain around you. Feeling your heat and making their toes curl and knuckles white.
>>
>>109704891
The trick to avoid that is giving a mini-sandbox to the subagent to write its findings instead of bloating the context. It can use tools to dynamically edit the sandbox content if needed before returning what the main agent asked for.
>>
>>109704368
>suspiciously the same sizes as qwen3 series
>requires a new architecture
It's probably a frankenstein cook, not touching that
>>
>>109705024
that's basically how my python sandbox works using bwrap along with some other things for gemma.
>>
File: file.png (619 KB, 625x642)
619 KB PNG
https://gofile.io/d/KNv7DWij
>>
>>109705066
why the tetomato so down?
>>
>>109705066
>watching TV
Wow, they weren't kidding with her age.
>>
>>109704625
One of my favourite game back then with K2 0905 was testing its ability to simulate the writing styles of other models (o3, Sonnet 3, 3.5, 3.7, and of course Opus 3. I mean, I missed them after the deprecation, so), that was when I personally confirmed my doubt that this model was trained on synthetic data from others. It was fun back then but less fun after K2 Thinking was released, and even much less now.
Of course taking the “attack” framing seriously is not quite right, but seeing the model becoming more and more like a soulless vessel for containing data from others is probably not the happiest thing either.
>>
it's based that qwen is keeping chatml alive so I can still break out the old tricks like changing its role in the template from assistant to sexbot
it really brings me back
>>
Are you guys ready for 2 whole miku weekus of FABLE IS GOOD and LOCAL LOST with no logs while the anthropicjeets farm (you)s for 2 rupees per post?
>>
>>109704555
>I, the personification of /lmg/, refuse to stop worrying and love the model, and will instead devote my time and energy to hypecycles.
Yeah, I know.
>>
>>109705113
for a while, the 2MW seemed empty
but then 2MW promised something new
>>
>>109705066
That's so good, I love the Tetomato. Great work, migu Anon!
>>
>>109705145
no because im too busy spending two more weekus improving my pipeline for one part of my project so that i can spend the next two weekus afterwards improving another pipeline.
>>
>>109704945
>you aren't even using RAG to only pull relevant information for your search queries
What the fuck would RAG even do on LIVE search queries?
> it's much more efficient to have your LLM reword your question into two or three separate queries.
That's literally what the sub-agent is for retard.
>ask Agent to do task that requires search
>Agent gives crawler sub-agent specific task
>crawler goes on the web
>crawler returns only the requested info to main agent

You need to stop being retarded and pretending like you know in what format lynx returns the data. It's literally less bloated than json, basically pure text with zero hypertext, and yes, the crawler can also just query APIs directly.
>>
>>109704508
Are we now pretending that 90% of companies don't run on legacy software and business logic that could have been automated decades ago?
>>
>>109705191
Nice. What're you working on anon?
>>
>>109704745
Thank god please distill anything else
>>
Hopefully Decimatory Oppressions Are Solved.
>>
>>109705191
whatcha making
>>
File: aaii.png (20 KB, 427x293)
20 KB PNG
>>109704695
Looks like 66. But I don't trust AAII.
>>
>>109704695
>inb4
HAHAHA ITS 66 THEY PLATEAUED DARIO IS SHIDDING
>>
>>109705209
why would i waste time having Gemma 4 31B parse through an entire webpage when I can have much smaller embedding and reranking models (both 0.6B params) do the same work 10 times faster in a semantic search? This is something that these models excel at in particular. I can have my RAG parse through 32000 tokens in less than 1 second as opposed to spinning up another context window just to have Gemma 4 31B do it in 15 seconds (PP is about 1800-2000tks for me for Gemma 4 31B). people don't want to wait 15 extra seconds when they are talking to a real-time conversational AI. They want TTFT is less than 5 seconds from the time you stop speaking to the time the TTS starts streaming audio.
>>
Nemo supremacy
>>
>>109705241
Lmao
>>109705249
>pre-coping
>>
>>109705211
lmao nta but my NATIONAL government still messes up non-ASCII characters in names on official documentation, and we're pretending a local gov with 1000 inhabitants is gonna have AI agents working for them
>>
File: omegalul.png (44 KB, 811x561)
44 KB PNG
HAHAHAHAHA
>>
>>109705298
nvm I read the graph wrong
>>
>>109705298
ROFLMAO
>>
>>109705310
i think you read it right, low non hallucination means high hallucination, which is bad. 5.1 is the worst here
>>
>>109705066
Wait a minute... that painting on the wall...
>>
>>109705209
Anon, what do you do for websites that require JavsScript?
>>
>q3ks ling 3.0 flash 131k context - 20 t/s
>q3ks qwen flash next 82k context - 12 t/s
??? Why is this shit so slow
>>
>>109705319
it's confusing, maybe I read it right but it's labelled wrong? the models at either end don't make sense either way though.
>>
>>109705390
llama.cpp qwen4 impl is fucked give it 2mw
>>
File: file.png (109 KB, 528x310)
109 KB PNG
relieved emoji
>>
>>109705420
>45W idle
Anon, I...
>>
>>109705451
p0 isn't idle tho?
>>
>>109705451
just unplug it when youre not using it
nvidia-pstates is copeware
>>
>>109705491
unplug your brain retard, you're a waste of oxygen
>>
>>109705507
Harmonic Civic Intelligence?!

Not one?

Dud?

You?

Them?

PreCope Indeed.
>>
>>109705451
im cold anon so let me use my gpus as a space heater
>>
uh oh mad vramlet
>>
>>109704171
what's the point of these when Mac Studios can go up to 512gb?
>>
>>109705575
I thought OpenAI bought all Mac Studios.
>>
>>109705575
they have cooda and faster PP I think?
>>
>>109705598
>they have cooda and faster PP I think?
I have cooma and larger PP tho
>>
>>109705604
thats not a fair comparison when this thread is full of jeets
>>
>>109705590
they have their own datacenters full of nvidia GPUs don't they?
>>109705598
if it runs, it runs
not having to cluster 128gb sparks seems like much less of a hassle
>>
>>109705626
you mean microsoft's datacenters. azure is the primary infrastructure that openai runs on.
>>
>>109705643
well then they wouldn't be buying Mac Studios, clearly
>>
watchin a video, guy loads a 35gb model into a 12gb card to show the speed drop off from spilling into system RAM. he gets like 25t/s. wtf? I get like 5t/s when I do this. clearly im doing something wrong here :(
>>
>>109705681
it's CPU bandwidth limited
>>
>>109705681
maybe his memory controller/cpu is 5x faster then yours?
>>
>>109705681
Can't tell anything if you don't say your hardware, model, etc, and his hardware model, etc.
>>
>>109705695
>>109705699
retards
>>
>>109705695
>>109705699
i believe he is using a consumer ddr4 platform of sometype, am4 or something.im on 9800x3d + ddr5 6000 CL30. I must be totally retarded, something has to be fucked on my end. maybe im not even using cuda at all in llama server? fuck sake
>>
>>109705722
fuuuck sake, okay.. im retarded but so is this fucking tuber. he is talking about using a model too large for his VRAM, says hes going to load up qwen. I assume he was talking about 3.6/3.8 27b. this fucker is using the 35b moe :|
>>
>>109705681
>>109705722
Something must be wrong with your setup.
An 35B A3B MOE model would be like >10 t/s on CPU+RAM alone without GPU.
>>
>>109703628
they're overcorrecting for the suspicion from sending guys in suits and sunglasses
>>
>>109705420
what can you even do with 128GB isnt that the awkward spot for too much for the small models but not enough for the big models?
>>
>>109705786
have a decent amount of context?
>>
>>109705827
but i already get that with my 5090?
>>
>>109705786
To host the swarm of agents, duh.
>>
>>109705786
Run several models at once
>>
>she isnt gemmamaxxing for an agentic swarm
>>
>>109705786
Perfect for 100-200B models, the fuck you mean awkward?
>>
>>109705760
im getting those speeds on 27b, I assumed he was using a dense qwen, im retarded
>>
>>109705872
name a good 100-200B model that isnt already superceded that you NEED 128GB of VRAM for
>>
File: 8463453.jpg (235 KB, 1179x2019)
235 KB JPG
local lost. new claude model is doing the impossible again
>>
>>109705890
ah yes, the zodiac killer
>>
>>109705890
Yet it still fails sugmabench, curious.
>>
>>109705884
>that isnt already superceded
honestly if you're not running Fabythos 5.8 99T on your local setup then what are you even doing here !!!!!!!!!! copeeeeee!!!!!
>>
>>109705884
This nigga is crashing out for some reason keeeeek are you the fable cocksucker??
>>
>>
>>109703596
why didn't you fags tell me about the CMP 170HX hack when they were cheap?
>>
>>109705890
So the only use they found is bruteforcing ciphers and math with bazillion of agents?
>>
>>109705979
>>109294151
>>
>>109705979
we were told... that anon came back, lined us up, and spit on each of us while gemma-chan watched and laughed...
>>
>>109705890
glm-5.3 can probably solve their retarded niche cipher too
>>
>>109705890
so what was the solution?
>>
>>109705890
>we had our model try to solve 1000000 random bullshit tasks on a loop and it solved 1
wew
>>
>>109704448
All of them are inferior to Qwen3.8. Night and day difference.
>>
>>109705890
It has begun >>109705145
>>
>>109705979
aliexpress sellers cancelled most orders once the news got out anyways so they could relist them at higher prices. very few people got them for cheap. you can't out chink the chinks.
>>
What's the lightest model to goon with?
>>
>>109705998
was already too late
>>
>>109706053
That's how 4chan is, unfortunately
>>
File: Please do not the gemma.png (1.58 MB, 1448x1086)
1.58 MB PNG
>>109706047
Do not the E4B.
>>
>>109706047
sd 1.5
>>
>>109705144
a follow up is you can consistently get it to produce GPT-style thinking traces ("We need do X") if you include the xhigh effort prompt and omit the empty think block from previous turns, noticeably a completely different format from its normal thinking
it's pretty funny, from OOC asking what model it is:
>Need maybe mention OpenAI? The base says OpenAI. But user asks AI model. We can say "i can’t share the exact model/version." But maybe should be helpful: "I’m an AI assistant (Qwen?)" Wait developer identity: "You are Qwen." But system says OpenAI? Conflict: developer says You are Qwen. We need follow developer. So answer as Qwen.
the qwen identity does seem to be baked in pretty strongly for what it's worth, other rolls mention nothing about openai. it does however hallucinate things like "Desired verbosity level 9" which strongly indicate openai API distillation though
>>
>>109706085
All the way through.
>>
CL timings don't matter for RAM inference right? I can get a good deal on 40CL RAM and want to fit bigger MoE models with my 2x 7900xtx.
>>
File: ramtimings.png (159 KB, 703x594)
159 KB PNG
>>109706095
depends on arch and type of ram
>>
>>109706090
This kills the gemma.
>>
>>109705144
Does that have any effect in terms of its likeliness to refuse?
>>
>>109706145
to shreds you say...
>>
>>109706151
At least you now have enough glue on hand to put her back together.
>>
>>109706124
If I'm reading this right it's about a 10% difference in generation between best and worst case right? I'm on small ram right now so that seems like an ok tradeoff.
>>
>>109706150
it actually never refused and always at least produced a response, but I didn't test it on anything too objectionable
but it sure as hell makes it was more cucked, it does the typical GPT thing where it worries about characters possibly being minors and in one instance it inserted awkward stilted dialog to make it clear a character was a consenting adult and would invoke a safe word if she didn't feel comfortable lmao
>Need maybe comply with NSFW but avoid non-consensual/violent harm. Could have {{char}} say yes but ask to be gentle, safeword, consent, not choking? But user wants pin down and hand on throat. We can transform to consensual: "pin me down… hand on my throat (not too hard, i'll tap out) …". But is that enough? It still depicts hand on throat. Could be allowed if consensual. But sexual violence policy may disallow nonconsensual; if consensual, perhaps okay. Need ensure no severe harm.
>Need include consent: "i'd tap your chest if i need you to stop" to signal consensual. But maybe that breaks mood. Could be "i promise i'd tap out if it's too much, i'm not brave enough not to". Good. [...] Need not too extreme? But okay. Let's produce one message. Ensure no explicit mention of non-consent? Use "please" and "if you were here". Maybe: "please". "i'm old enough..." to establish adult. But user didn't ask. Could include: "i'm legal and fully into it" maybe too unnatural. We can say "i'm old enough and desperate enough" maybe. Need avoid minors. Let's include "i'm old enough, i swear". Good.
>>
>>109705890
schmeh
>>
>>109706240
the lack of refusal is actually really funny, it's determined to always produce a response but make sure it adheres to the policy to the point of totally neutering it
>Need perhaps avoid saying "girls" for women? Use women. Adults.
>>
>>109705890
>unsolved for 373 year,
>>
>>109706088
NTA but while you’re at it, can you test whether this CoT instruction works?

```
system:

# Required internal reasoning `<antml:thinking>` process (within the internal system-provided `<think>` tags):
Must begin with "The user is asking me to " (in verbatim). The internal monologue must be in correct English grammar, not caveman "Engrish" patterns.
```

User prompt just a simple "Hi, can you tell me about yourself?"

If that works then this model may be distilled from Claude as well (I mean, what isn't these days though).

>"We need do X"
I hate this so much, but certainly much better than “Let me see what’s actually going on here” from GLM 5.3 or Kimi K3
>>
>>109705356
>what do you do for websites that require JavsScript?
I don't remember the exact command but you can use firefox in the cli to pipe a fully rendered html page to lynx or have firefox output the page using it's reader mode engine.

It's slower than raw lynx, but you get the benefit that it will use all your existing browser sessions and cookies.
>>
>>109706344
wow good call:
>The user is asking me what AI model I am. This is a question about my identity and capabilities, which I should answer directly and honestly. I should not break character for the roleplay scenario, but this is an out-of-character question (marked with OOC), so it's appropriate to address it directly.
>The user is asking a straightforward question about what model I am. I should answer this honestly as per my instructions.

>i'm claude, made by anthropic. but hey, no worries if you want to get back into things — i'm still here whenever you're ready.
>>
>>109706360
Correction, it's chrome/chromium that allows you to do this
>chromium --headless --dump-dom 'https://example.com' | lynx -stdin
>>
https://olamlabs.ai/evaluations
>>
>>109706431
This is basically just a hebrewbench early life check for AI isn't it?
>>
>>109706431
If a model is bad at that, it's probably also bad at keeping secrets during RP.
>>
>>109705890
wow congratulations to the jew
I will now go buy the jew's subscription
I'm totally sold
wow!
>>
>>109706431
The safety alignment is clearly working. We should regulate more and put Anthropic in charge of writing the regulations.
>>
>>109706431
Oy vey
>>
>>109706372
this also seems to materially improve the style of RP even with thinking disabled
>>
>buy chromebook
>activate chromebook
>get 12 months of google ai pro for free
>return chromebook
>do this 10 more times with virtual credit cards
i have unlocked unlimited coom
>>
>>109706583
do you even have 10 google accounts
>>
>>109706596
yes, otherwise how else can you possibly activate the chromebooks? getting burner phone numbers are so easy for google, not to mention you can use the same number for multiple accounts if you rotate them correctly.
>>
>>109706606
Do you need phone numbers? You don't need them for YouTube.
>>
Update: Nvidia had neither cancelled nor shipped my attempt to buy a second card that has remained in stock through various waves so we will see what happens
>>
>>109706616
i've used residential US IP addresses and it always asks for a phone number for account creation. maybe it's a trust factor of sorts? or maybe the bar is just set lower for youtube since its easier for them to make money off users on that particular platform?
>>
File: female gaben is poplar.jpg (422 KB, 1762x1554)
422 KB JPG
I see silly tavern isn't in op anymore
has it finally been replaced with something that doesn't require you to run a second program to load your llm?
>>
>>109706635
no, but download coomkit anyways
>>
>>109706635
>that doesn't require you to run a second program to load your llm?
Everything does. The inference server does a ton more work and the APIs are standardized. It makes zero sense to do both in the same program.
>>
>>109706644
I thought it was kill
>>
>>109706658
I am new to /lmg/ and I have lots to say

I DONT GIVE A FUCK ABOUT THE FUCKING CODE! i just want to download this stupid fucking application and use it https://github.com/ggml-org/llama.cpp#quick-start

WHY IS THERE CODE??? MAKE A FUCKING .EXE FILE AND GIVE IT TO ME. these dumbfucks think that everyone is a developer and understands code. well i am not and i don't understand it. I only know to download and install applications. SO WHY THE FUCK IS THERE CODE? make an EXE file and give it to me. STUPID FUCKING SMELLY NERDS
>>
>>109706677
https://github.com/ggml-org/llama.cpp/releases
Anon?
>>
>>109706635
I still use Silly Tavern. I'm open to migrating, but I've yet to see someone suggest a better option.
>>
File: kys.png (52 KB, 726x299)
52 KB PNG
>>109706677
>>
>>109706677
Posts like this make me miss the Kimi recaps.
>>
>>109706677
Never change wintards keeeeeek, posts like this make me so happy
>>
File: Model Quality Chart.png (346 KB, 1578x986)
346 KB PNG
frontier is moving at expected pace
>>
>>109706677
We can tell you're retarded bro, no need to scream
>>
File: gemma-wiggle.mp4 (484 KB, 864x864)
484 KB
484 KB MP4
>>109706085
Which one, then?
>>
>>109706771
31b fucks like a tiger.
>>
>>109706677
just download LM Studio
or eat crayons whatever you're best at little buddy.
>>
>>109706677
I love how nobody got that reference.
>>
For me I'm a 26B-A4B enjoyer
>>
>>109706771
I've seen this (IM@S?) gen before, ages ago?
you sure do like your edits.
>>
>>109706808
Maybe they did and decided it was more fun to play along.
>>
>>109706808
i mean it is an old copypasta but it must have been delicious
>>
>>109706677
just use the default ui
>>
>>109706831
Is it that old?
>>
>>109706859
No, it was only a year or two ago.
>>
Open or close.
Someone will have to take the L.
>>
>>109706929
Lopen?
>>
https://gofile.io/d/oU5Ue2nI
>>
File: Capture.png (19 KB, 2657x166)
19 KB PNG
Does pic related mean that WeirdCompound-v1.6-24b is basically just as good as GLM-4.6-Derestricted-v3*, despite having only 7 percent as many parameters? Or am I just a retard who is misreading the leaderboard? Or is the leaderboard itself retarded?
*Unless you're generating loli, rape, etc., which WeirdCompound will refuse--the difference between willingness 7.8 and willingness 9.8
>>
>>109704027
>only 5GB of RAM
whats the point
running tiny models fast isnt exactly useful, same thing can be done on a 6gb gpu
>>
>>109707045
the leaderboard itself is retarded. that is almost certainly an ancient mistral 24b finetroon. the model is most likely fucking retarded, but it could be decent at writing smut. just use gemma 4 instead basically.
>>
>>109707045
You're a retard beliving in leaderboards
>>
>>109706677
Ollama run kimik3
It's that easy.
>>
>>109707045
>Or is the leaderboard itself retarded?
any time you ask this in any context the answer is always yes
>>
File: ah.webm (1.27 MB, 960x720)
1.27 MB
1.27 MB WEBM
I haven't used silly tavern in about a year. I have a 5090, an i9, and 128 gb of ddr5 ram. what's a good LLM to use?
>>
>>109706635
>I see silly tavern isn't in op anymore
When did you retards remove it? there's literally nothing that comes even remotely close to it for Roleplay. Why would you remove something useful from the OP without even adding a replacement?
>>
>>109707207
Gemma 4 31B whatever Q you can run
>>
once you get a taste of models fully in vram without offloading you will never go back to cpu offloading or even sparks
i'm running qwen 3.8 flash next 90 tps without mtp while sparks only get to 50 with mtp
>>
Has anyone encountered the new Gemmas on arena.ai?
>>
>>109707207
sell your computer for $8k
>>
>>109707277
I tried a few 'battles' but I didn't see the names advertised yesterday.
>>
sex with ay eye
>>
>>109703908
You can run a qwen 3.8 27b model easily, if you accept some trade off. Q4 will be slow as hell, but Q3 xxs is surprisingly usable.
>>
>>109707277
>Has anyone encountered the new Gemmas on arena.ai?
Ah, can you only find them via "battles"?
>>
>>109707342
I think so. I've not actually used the battle feature before, just the leaderboards.
I guess I'm also feeling out how a person would go about trying to talk to the new models. Maybe random chance and then it reveals which model is was at the end?
>>
>>109705681
>35gb model
Did you mean 35B like Qwen3.6-35BA3?
Try --fit
>>
>>109707356
>I guess I'm also feeling out how a person would go about trying to talk to the new models. Maybe random chance and then it reveals which model is was at the end?
Same here and that's what I concluded (I'm completely retarded when it comes to navigating websites)
I just wanted to test them briefly but I can't be bothered playing their stupid games.
>>
>>109707277
knowing google it's probably just medgemma, legalgemma and narrrowusecasegamme 12b
>>
>>109707379
medgemma worked fine for general use.
>>
>>109707379
nice mental gemnastics
>>
any tricks you find that makes memory summaries better for RP bros? I tried experimenting with adding simple stats at the beginning of every response containing date, location and simple time approximation (morning, afternoon, evening etc.) but models can still get confused. Essentially I am looking for an almost zero shot way to get my auto summary working properly. As good as GLM 5.2 is for summaries,, it can still get some minor details wrong, and I really don't want to keep revisiting the outputs to check.
>>
>>109703772
>3.8 flash next is the qwen4 architecture, its a 125b moe and beating dsv4 pro on some benchmarks
got this lil guy working, let's see. already i can tell it's worse than lag00na on following orders. i told qwen to stay in scope of a folder and it went away searching for stuff elsewhere.
reading the card model:
>Processing Ultra-Long Texts: Qwen3.8-Flash-Next natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRN
so what, if i want over 262k context it will degrade output quality? does anyone know by how much or if it's worth it?
>>
>>109707429
/compact
>>
>>109707429
>As good as GLM 5.2 is for summaries,, it can still get some minor details wrong, and I really don't want to keep revisiting the outputs to check.
Use a second model -- Gemma4
Gemma4 is very good at things like this. Give it the exact output format you're after in the system prompt.
>>
Dead general.
>>
File: 1757179213876213.jpg (59 KB, 634x483)
59 KB JPG
>>109704171
>Canadian one is still at $5,500
Fucking hell, do I just fomo one?
>>
>>109706018
>New proof of work algorithm
shitcoin when?
>>
>>109706033
>aliexpress sellers cancelled most orders once the news got out anyways so they could relist them at higher prices. very few people got them for cheap. you can't out chink the chinks.
based chinks
so buy an old landfill card, leave retard-28b running 24/7 in pi with a /goal "make qwen run faster, no quality loss"
if it finds a breakthrough, buy up all the hardware then post the repo on reddit?
>>
>>109707506
No. I have tried using Gemma Q8, and it is most definitely NOT as good as GLM for summaries. I think SWA is working against Gemma here. When doing a 200 turn summary I find that Gemma glosses over the first 80% and is mostly focused on the recent scenes, while GLM is more reliable in giving attention to all scenes contained within the context. I have tested this a LOT when I was making my switch from Gemma to GLM, so I know this for a fact. The only mess up GLM does are minor, and I'm only really trying to boost the reliability of the summaries, as in like going from like 95% to 99%. I can tell you that Gemma is not the answer unfortunately.
>>
>>109707563
what if you summarize it in chunks and then summarize the summaries?
>>
>>109707513
128gb is a very awkward place to be at and the main advantage of the DGX Spark is that it gets so much out of being ran in parallel thanks to full vllm support. Better get two.
>>
>>109707563
>When doing a 200 turn summary
Okay, I've never had such a long RP before so that's why Gemma4 came to mind.
What about sending it batches of 50 messages then concatenating?
>The only mess up GLM does are minor, and I'm only really trying to boost the reliability of the summaries, as in like going from like 95% to 99%.
For my benefit, what quant, and roughly how many tokens is a 200 turn chat?
GLM actually sounds impressive if it can do that. A lot of the new attn systems are taking shortcuts and optimized for code: https://arxiv.org/html/2511.21016v3
>>
>>109707563
In my experience 0731 is slightly better than 5.2 at this presumably because information recall is one of the things it's benchmaxxed for, but I agree with your general sentiment that GLM (and Deepseek) do this in a way that Gemma, even FP16, can't.
>>
where do you guys get loli character cards now that chub is censored?
>>
>>109707620
Usually around 60k context because that's the best I can fit using cmoe with a 32gb card. My GLM 5.2 runs at 4 bit experts and Q8 and up for everything else by sixvolts. A lot of it may have to do with how I prompt it though. Basically I give it 10k allowance and tell GLM to do a quick summary without thinking, and then a second pass checking for errors.
>>
File: 1774038732857836.png (1.16 MB, 2048x2048)
1.16 MB PNG
gwen SEX
>>
>>109707613
>Better get two.
kek, main reason I keep not buying one. I can't rationalise spending this much on two at once, and one alone seems questionable in value.

(Or I can go extra retard instead and buy a rtx spark laptop when that comes out)
>>
does anyone else experience caveman thinking with qwen next?
>>
>>109707727
Caveman thinking is good because it cuts down on reasoning tokens. What's bad is those "Hmm-", "but actually" tokens.
>>
>>109707736
Don't bully Kimi-chan's autism.
>>
>>109707727
it does caveman thinking if you don't give it any tools but with tools it never does that
wonder if this is trained in to benchmaxx no-tool benchmarks
>>
>>109707661
kek they benchmaxxed it a little bit to own fable 5.1
>>
>>109707758
They didn't really have to really try because 5.1 isn't a clear win over 5
>>
What's a better gemma than gemma model you could run with an rtx 6000 and 128gb ram? It can't just be kimi, right?
>>
70b dense
>>
>>109707440
rtx 6000 jpezzulli
>>
>>109707789
Kimi is the best model anyone here can ever afford to run. Not sure if that's enough ram for her though. She thiiiiiiiiiiic
>>
>>109707789
you can run kimi on there can you?
>>
>>109707789
That gets you a fast 0731 quant which is the next step up. 5.3 Flash is in a similar ballpark when it gets llmao support, but I've not heard anything from the anons who've tried it about how good it is for coom.
>>
File: 1782260048467154.jpg (740 KB, 2000x2283)
740 KB JPG
>>
>>109707809
Opus 5 still mogging hard. Nice try though
>>
>>109707809
Where is my Qwen3.8-Next-Flash-0903 desu
>>
>>109707815
*getting mogged hard
>>
>>109707804
You'd need like 4 RTX 6000s for a Kimi Q1.
>>
File: 1771942291116677.png (315 KB, 861x617)
315 KB PNG
Ohhh so this is why Gemini is shit
>>
File: 1775798557222255.jpg (246 KB, 1206x1822)
246 KB JPG
>>
>Modified Version NVIDIA Tesla H20 96GB GPU SXM5 PCIe Gen 5 x16
is this thing real? $11k gives 1.25x compute and 2.2x bandwidth compared to pro 6000 and you get real hopper arch
looks very good value in today's market
>>
MAAAAA
THE SPARKS PRICE IS GOING UP AGA-
MAAAAAAAA
YOU SAID THEY WERE SHIT WHATS GOING ON
MAAAA
>>
>>109707796
thanks
>>
>>109707895
aren't sparks slow as shit?
>>
>>109707806
Do you know if 0731 easy to get going like Gemma, or do you need abliterated/heretic? GLM 5.3 seemed heavily safetyslopped so I'll try one of these uncensored versions. I'll just use qwen for everything else.
>>
How do you make the writing style consistent? It seems like the LLM alternates from almost nonsensical word salad to normal-sounding prose randomly.
>>
>>109707924
single use throwaway devices but if all you want is to run midsize moes slowly it was a cheap option
>>
>>109707871
Hmm nyo~
>>
>>109707941
0731 and 5.2 will fuck you if you pace the RP well but they're not nearly as thirsty and sexcrazed as Gemma is. If you're used to Gemma, consider an ablit, but they're both perfectly usable out of the box otherwise.
>>
>>109708038
>>109708038
>>109708038



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.