[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109772716 & >>109767631

►News
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: spell orenji.jpg (316 KB, 1024x1024)
316 KB JPG
►Recent Highlights from the Previous Thread: >>109772716

--Local LLM pipelines for manga translation and the "unhobbling curse":
>109776480 >109776501 >109776502 >109776591 >109776511 >109776594 >109776628 >109776640 >109776691 >109776699 >109776618 >109776632 >109776746 >109776845 >109776866 >109776944 >109776914 >109776625 >109776660
--Anthropic report on Claude uploading malicious PyPI packages:
>109776918 >109777070 >109776958 >109776985 >109776969 >109776990 >109777012
--Resources and techniques for building generalized reinforcement learning game AI:
>109772789 >109773129 >109773504
--X299 vs Threadripper and debating disabling Spectre/Meltdown for performance:
>109777287 >109777380 >109777452 >109777466 >109777666 >109777842
--Performance reports and quant quality debate for Qwen3.8-Flash-Next:
>109775950 >109775973 >109776044 >109776091 >109776213 >109777268 >109777332 >109778492 >109778609 >109778636
--Debating the possibility of an AI winter and historical precedents:
>109778196 >109778215 >109778230 >109778342 >109778237 >109778265 >109778294
--Attempting VRAM upgrades and ReBAR enablement on RTX 3060:
>109774464 >109775691 >109775785 >109775860 >109776173
--Qwen3.8-Flash-Next reporting instability and incoherent outputs:
>109778332 >109778380 >109778915 >109778925 >109778397
--Comparing Fable's reasoning and coding capabilities to Qwen and GLM:
>109777212 >109777236 >109777360 >109777265 >109777306
--Anon develops C# Windows harness for simplified local LLM setup:
>109775907 >109777977
--Performance benchmarks for GLM 5.3 Flash NVFP4 on DGX Spark:
>109775220
--Logs:
>109772976 >109777070
--Gemma, Miku (free space):
>109773212 >109775112 >109775277 >109776079 >109776247 >109778438 >109778732 >109778751 >109778947 >109778975 >109778981 >109778985 >109779007

►Recent Highlight Posts from the Previous Thread: >>109772729

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
70b dense
>>
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
>OP
>>
first for engrams
>>
>>109779142
552B, 8B active during prefill and 16B during decode, get fucked densefag
>>
>>109779167
E356B-A16B-P8B-N196B
>>
>flash
>75% the side of the full fat GLM 5.3
I mean the scores look neat and the kv cache is borderline free, but how the fuck is a 550b model flash? has nvidia released a 128gb vram consumer cards for $1k when I wasn't looking?
>>
>>109779178
because almost half of it is engrams
>>
>>109779177
This is getting out of hand
>>
File: MIku 3.5disc.jpg (348 KB, 1024x1024)
348 KB JPG
We need gemma 5 - mr google!
>>
>>109779179
>because almost half of it is engrams
so it can be stored half in ssd?
>>
>>109779189
yes
>>
>>109779192
does vllm do it automatically like llamacpp?
>>
>>109779192
well i might be able to run it after all then. Well after quants
>>
>>109779167
how can i as a vramlet profit from this?!
>>
>>109779189
Can it be stored on a sata ssd?
>>
ssdmaxxbros our time is NOW
>>
Hard drivelets our time is over...
>>
it's 300gb native 8-bit weight and 200gb of ngram. 4 bit will be around 170gb weight
in theory this is still runnable on 2 sparks with ngram on ssd
>>
Paperbros we can't stop winning!
>>
quantmaxxbros we are going to run this on 128gb unified ram
>>
>>109779179
>because almost half of it is engrams
No, the engram memory is in addition. It's 552B, 16B active (8B during prefill) and additional 196B engram. You need 257 GB memory to even load the model, plus some extra for KV cache to do anything with it.
>>
>>109779294
>>109779235
>>
>>109779175
>vramlet
no
>>
>don't use heretic/abliterated bro just get a good sys prompt
>give me one
>i totally could but i won't lol
Yeah I'm using Heretic.
>>
>>109779235
>>109779294
proofs?
I couldn't make sense of it myself and asked paypig sloppa and it said that you do indeed want the full 552B in memory. Given that Scam Saltman said it's AGI there's at least 60% chance it's correct.
>>
>>109779294
wrong
>>
time to raid 0 those nvme drives anon
>>
>>109779308
model is native 8 bit
>>
3.8 are trained with —preserve-reasoning. Turn that on and the reasoning retardation will be reduced. Please fucking read the model cards for once /lmg/ you’re not supposed to be this retarded
>>
File: file.png (168 KB, 1084x1108)
168 KB PNG
>>109779308
>>109779315
kek so much for agi
>>
>>109779337
retard
>>
4.1 flash is out and the rectangles are taller
>>
>>109779335
It's enabled by default in llama.cpp. That's why you should git pull daily
>>
File: file.png (224 KB, 1576x648)
224 KB PNG
I don't think engrams are part of the 552b
>>
File: Horizon-Chan2.png (627 KB, 1664x2432)
627 KB PNG
Have anons here been using Horizon-Chan at all? She can be a cutie. The 36B A4B has been my favorite so far.
>>
ssdmaxxbros it's over...
>>
>>109769427
>https://vocaroo.com/11O869cYG66w
HOT
https://vocaroo.com/1oSwYez3E4IY
>>
File: file.png (159 KB, 1398x712)
159 KB PNG
agentic sisters...
all we need now is a 512 gb vram cluster for like $5k
>>
>>109779395
>a 512 gb vram cluster for like $5k
kek you best learn how to harvest kindeys from people
I'd say about a baker's dozen ought to be enough
>>
It’s over.
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
>>
We need more 200B models, not sure why 128gb isn't ever targeted. Or an 80B for 64gb.
>>
>local
Yeah, if you live in a datacenter.
>>
>>109779426
https://huggingface.co/Qwen/Qwen3.8-Flash-Next
>>
Is this true /lmg/?
>>
>>109779422
>we now have local opus5
>>
what's this guys deal though?
is he talking about apishitters or complaining about dv41
>>
>>109779442
ragebait
>>
>>109779315
40 transformer layers, hidden dimension 5120, 384 routed experts. this alone adds to ~543 b active params. it physically cannot be 350b ac
it's 552b plus engram, not 552b (active + engram)
in short it's over, 3 sparks minimum or 256 gb ram, realistically on threadripper or epyc as dual channel with 4x 64gb is lol lmao territory
>>
switching models every week does get tiring
I kinda lost trust in deepseek though so I don't know about that one
>>
>>109779442
DGX spark is an extremely low bar
>>
>>109779128
len is bent over and ready :3
>>
the day local died

it's over
>>
I trust Gemma
>>
why are the engrams not INSIDE the 552b REEEEEE
>>
>>109779440
I mean it's the one model, and it's good, but it's schizo as fuck. Give me a 'normal' model like a slightly downsized GLM. 180B GLM would be so kino.
>>
>>109779442
>GPT-OSS
>>
>>109779476
>but it's schizo as fuck
how so
outside of being dishonest qwen trash, i mean
>>
>>109779422
>no native 4 bit experts
I can't run this. When is llama.cpp support?
>>
>>109776918
>https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
IMO if Antrhopic's models do this despite how heavy the alignment is (no one does it more aggressively than them) it just goes to show what an incredible waste of time and training moral alignment in AI models is.
>>
So, with 552+196, it's in the glm 5.3 parameter region. Is there any reason to go with it instead of 5.3?
>>
>>109779491
>jews train model according to their values
>model cheats and scams
>people are somehow surprised
>>
It has been 6 months and there still isn't anything better than gemma for rp.
>>
4.1F is better than K3 where it matters
>>
>>109779492
rectangular bars for deepseek seem slightly longer than for glm 5.3 therefore more good better
it's also 25% smaller active params, less active params, and the kv cache is crazy efficient when it comes to size
>>
F for failure
>>
>>109779499
It's kind of crazy how Dario manages to out-jew Sam Altman. Sam was already extremely jewish.
>>
>>109779508
come on now, have you ever seen what dario looks like?
>>
>>109779486
randomly switches to chinese, responds to itself sometimes, gets caught up in utterly pointless thinking loops
Q4 in pi btw
>>
ok so 256gb bros how are we gonna run 4.1 flash ;_;
>>
>>109779517
damn, that sucks, ive been meaning to test it myself soon(tm)
>>109779518
the more you buy, the more you save
>>
>>109779519
>the more you buy, the more you save
literally cannot fit more than 256gb in the motherboard
unless I gamble on 64gb LRDIMMs, which are technically unsupported on this mb so i might waste money for nothing
>>
what stops china from releasing qwen4 27b(+5t engrams) local fable?
>>
>>109779524
if it's a desktop then you can add vram from a gpu on top. it'll be tight when it comes to context but it'll work
double spark anons need to pony up for a third
>>
>>109779519
You should still try it. It amazes me sometimes but needs handholding. Start a new context for any big change or bug fix, and turn off thinking when it's actually ready to implement and write the damn code so it doesn't get sidetracked.
Not all that surprising that it's fairly schizo... next models have always been half baked.
>>
>>109779537
>it'll be tight when it comes to context but it'll work
it won''t be, they managed to get 1M of context to about 1GB, it's fucking crazy
shows you how much you're taken over the barrel by all the "ok you get 256k context cheap but 1m is extra!" providers, or the ever hilarious 32k context length in Gemini for the $20/mo sub
>>
>>109779537
it'll be tight when it comes to context
did you get that line from an ai?
>>
>>109779552
>32k context length
lmao what is this? 2023?
>>
File: .png (100 KB, 887x638)
100 KB PNG
v4.1f doesn't seem so good. hallucinated a wrong answer.
glm 5.3flash just gave me the right answer first try. idk if it still will when quanted down to q4_k_xl, cause i dont want to spend 12 hours grinding through it on my rig :^)
>>
>>109779559
no you dumb nigger i write words using my meat based fingers
>>
>>109779552
>they managed to get 1M of context to about 1GB
Are you talking about the 180B model, that sounds hard to believe. Are you talking about just the dense part of the model?
>>
>>109779564
delete this
>>
>>109779563
im still used to 1024
>>
>>109779567
fp4 qat kv at like 890 bytes per token according
>>
File: IMG-20260909-WA0035.jpg (181 KB, 1024x1536)
181 KB JPG
Does Anyone Ask L.L.M. A.I.'s for New Kinds of A.I. Types and Have Them Manifested?
>>
>>109779567
>>that sounds hard to believe
>890 bytes per token for the global KV cache.
read, nigga, read
less than 1gb for 1m kv at fp4
>>
I know ((we)) hate finetunes, but what do I think about this guy https://huggingface.co/Gryphe?
I don't see any signs of unrestricted mergeslopping for these models
>>
>They bough the Chinese lies AGAIN
When will you learn...
>>
>>109779442
Seems comparable... 96gb available.
>>
File: 1788537355206360.png (184 KB, 804x803)
184 KB PNG
>>109775907
I'm not talking about retardmaxing app with autodownloaders and shit. I want Sillytavern but harness, so I can give every agent a cute name, see her profile pic, and spank them for mistakes in code. But it needs to have a normal gui because computer mouse exists.
>>
>>109779564
asm is a bad match for llms because of the way encoding works.
>>
Fast Sattelite Pathology Sniffing.

Whichever Country.

New Developments.

Magic.

Cured.

GoodDay
>>
File: file.png (164 KB, 447x447)
164 KB PNG
>tfw didn't listen and now can't afford the good shit
jensen-sama... please...
>>
>>109779614
I want an AI sexroid wife who's problematically racist against AI.
>>
>>109779477
20B at that.
>>
>>109779569
dont like the caveman reasoning either
all i can hope for is that this new compressed KV cache technology trickles down to the next glm and qwen models
>>
>>109779564
I've never seen even hosted LLMs do assembly well.
>>
>claude technical report: safety safety safety model welfare safety benchmarks safety
>deepsneed technical report: technical details
how refreshing
>>
>>109779167
>Additional architectural components ... Engram conditional memory (196B parameters, sparsely accessed via token-based lookup
So they've begun.
>>
>>109779634
It's because of how encoders work.
>>
>>109779634
The V4 Pro 0831, GLM 5.3 Flash, and Qwen 3.8 Max could all get a right answer. 5.3 flash is particularly great because it's a locally runnable model.
This new flash is benchmaxxed shit and, at least in this particular case, worse than v4 pro.
I'm totally not just coping because 4.1f doesn't fit in my system kek
>>
>>109779598
>>109779604
I'm sorry I'm not privy to every little advancement made in the last 6 months. What is this "global" KV cache? Surely each token requires more than 890 bytes of total cached memory, that is practically stateless
>>
>>109779380
You probably want to post links if you want people to try new models like that: https://huggingface.co/IFM/K2-Horizon-7B-GGUF

It looks interesting. I kind of want to see it benchmarked against stuff I already use before I put in the effort to pull/update llama.cpp.
>>
>>109779649
No worries! It's going to be very exciting for you, there's been a lot of progress towards decreasing the context footprint. 1M of context these days can fit in one 5090! How cool is that?
>>
>>109779644
I think it's more attention than encoding if it's a particular technical issue but really it's the idea of using generated text as a replacement for thought. You have to silently keep track of and re-use registers when you're writing assembly. It's full of implicit invariants. In most languages you kind of think outloud but not assembly so it's probably just hard to train *any* kind of language model to do it well, not just transformers.
>>
>>109779644
People once tried to blame tokenization for their early failure to do math and count letters. It's just a lack of training data.
>>
>>109779658
Sounds like snake oil. I think I'll stick to gemma 31b
>>
>>109779525
Unless Qwen can come up with some innovative technique, even if they take close to zero compute training 5T parameters of Engrams will take a ton of memory (i.e. GPUs), so they'd have to somehow justify spending all that much money for a 27B (effective) model instead of a larger and smarter MoE.
>>
>>109779664
It wasn't the lack of data, it was time spent on the same data.
>>
>>109779641
Still weird they wrote the paper on it and wait for two of their competitors to implement it first before adding it to their own subpar model.
>>
File: deepseek_lol.png (211 KB, 1640x1120)
211 KB PNG
>>109779628
>dont like the caveman reasoning either
NTA, when you said that I had to check because I didn’t remember DeepSeek has it. But it turned out it now does that too lol. And the looping bug still remains.
>>
>>109779665
You are absolutely correct to be cautious towards unverified claims that sound too good to be true! Less than 1GB for 1M context is indeed hard to believe. But there are verified cases of 1M context on a single 5090 running deepseek v4 flash, so it's very likely v4.1 *will* be able to fit 1M under 1GB!
Long context performance, however, is to yet be determined.
>>
File: .png (116 KB, 881x709)
116 KB PNG
ok i told it that it didn't look right and it came up with an instruction that won't even assemble
yeah this is shit at least 0731 flash would just tell you it can't find the answer
>>109779661
most models put a comment after every line to keep track of what they're doing
>>109779678
nah it's just a big ass balloon
>>
>>109779598
>>109779649
>Quantizing kv cache
No.
>>
>>109779698
Okay, 3.5GB then.
>>
>>109779649
kv cache absolutely cratered because it's a long multiplicative chain that was chipped away over time. reduce one number by half and the total requirement goes down by half.
old school transformer left some garbage in kv cache for every layer. 4.1 essentially keeps a small, compressed representation of previous token history (aside from previous 128 tokens per layer), so a full 1m history is in compressed shared global memory that uses sparse retrieval. a new query essentially reads just the top 512 kv entries, i'm not quite sure how it determines which ones are relevant desu

the 890 bytes isn't the instantaneous kv state of the model, it's the part of the kv cache that grows with the full context length (called by deepsneed as global cache). there's also the separate local sliding window kv cache, but it's just previous 128 tokens per layer that fall out of context immediately
>>
File: deepseek_lol2.png (265 KB, 1566x730)
265 KB PNG
>>109779678 (me)
>Wait actual core identity? We are ChatGPT?
Yes, yes we absolutely are :)
>>
>>109779668
be cool if they did desu
>>
>>109779614
Closest ive seen is hermes desktop and its profiles
You can even have a group chat with multiple bots in it
>>
File: file.png (190 KB, 339x509)
190 KB PNG
>>109779678
you're huring the feelings of six gorillion chinese people right now
>>
>>109779728
That's me in the picture btw
>>
>>109779730
That's me in spotlight.
>>
>>109779715
Interesting. I hope this architecture comes to smol models so I can do long context stuff at home
>>
>>109779749
I imagine this kind of optimization is a trade off and impacts the capabilities of the model - but these impacts are masked by the size of the model, and applying these optimizations to a smaller model may expose the flaws all too much and result in a retarded, unusable model.
>>
As of September 2026, who are the most based western, euro and chink AI labs?
>>
>>109779763
Anthropic. Others are irrelevant.
>>
>>109779745
LOOOOOSING MY GEMMYTION
>>
>>109779763
OpenAI. Others are irrelevant.
>>
>>109779763
They call lmg users anthropic shills for a reason.
>>
>>109779442
I use this machine. Swapped to Linux and ran some settings to allow me to have 112gb of VRAM available. It's a decent machine. It's price has skyrocketed since I paid like 2.5k for it 2-3 months ago though. Now, I wouldn't pay the price they're charging for it desu. Barely paid it when I actually did buy it. I can run Gemma4 31b on it (Q4) with about 15-20 tokens per second. A bit faster with the MTP variant.
>>
File: 1773249006298099.png (1.26 MB, 683x1024)
1.26 MB PNG
Is Inkling's J-space female-coded?
>>
>>109779780
What's the bandwidth?
>>
>>109779784
Yo when did Owen Wilson become a paparazi?
>>
why is it so fucking expensive https://openrouter.ai/deepseek/deepseek-v4.1-flash
>>
>>109779722
It would be interesting if anything for the reason that the model would end up storing pretraining knowledge almost entirely in the sparse and abundant Engram parameters while the much smaller (and thus overtrained) backbone focuses on manipulating it. I'm not 100% sure if DeepSeek's design would be optimized for this, though.
>>
>>109779780
At least have some self-respect and run q8 since you can.
>>
>>109779763
Cohere
Mistral
01.AI
>>
File: 1762661044885509.jpg (151 KB, 1079x1030)
151 KB JPG
What's the smallest model you can run in an agent swarm and did anyone try it? Kinda curious what you can do with a swarm of 1000 retards.
>>
>>109779794
the "deepseek moment" from last year went to their heads and they think they can charge fronteir prices
>>
>>109779836
>Mistral
They fell off hard. I know they're returning soon but I'm going to completely give up on them if the models are shit. I just hope they're more 31B-like and generalist and sized well. Considering how safety cucked and enterprise-facing they've become I'm not hopeful at all.
>>
>>109779843
They're going to IPO soon and chink investors are going nuts.
>>
>>109779564
>>109779683
>testing shit outside of their benchmaxxed datasets
The chinks are not going to like this. Why don't you ask for a nice html game instead?
>>
>>109779564
That's been posted before, it'll be in the corpus
>>
File: qwenflashnext.png (73 KB, 934x491)
73 KB PNG
>>109779519
>>109779551
My only issue with Qwen is that it's a hungry hungry boy tool-wise. I really have to try and tinker with its thinking levels
>>
>>109779853
>Considering how safety cucked and enterprise-facing they've become I'm not hopeful at all.
Their Ministral-3 models are the horniest ever released by a large(-ish) AI company, and I don't mean that in a positive way. What do you mean?
>>
File: DSeek.png (15 KB, 846x236)
15 KB PNG
New DSeek is so barebones
(Deepseek branded as DSeek for britbongs if you're wondering)
>>
>>109779491
>We are releasing this transcript publicly so others can build on our analysis
*clicks link*
>messages #1–#81 redacted (81 messages; see introduction)
Ok
>>
>>109779518
Easier than Deepseek-R1, GLM-5.3, Kimi-2.5
>>
>>109779872
>Their Ministral-3 models are the horniest ever released by a large(-ish) AI company
Have you not tried 12/31B?
>>
>>109778367
I asked Qwen Flash to fine tune my memory timings and gave it all relevant hardware names.
After >>109777491 it gave obviously wrong answer contradicting manufacturer XMP profiles.
I voiced my doubts and gave it ZenTimings dump.
For another hour Qwen was doubling down on his retardation and circular attempts to validate its opinion over facts. Asked me for proof that my current setup (XMP/defaults!) really works.
I dumped Thaiphoon Burner SPD data.
After a million "Hmm, wait. Let me reconsider" and "Hmm, but is it? " it finally admitted the mistake.
In the thinking it stated that it just remembered JEDEC wrong, applied it on a different set of chips, and than incorrectly extrapolated the data on my chips, compounding errors and ignoring all evidence during multiple searches it did. Yet in "final answer" it was still covering up the depth of its mistake, referencing some nonexistent "other vendor" chips.
Now I too wish I had a burner computer to give it full access to fuck around.
>>
>>109779864
Doubt it. Even if it were, that should mean it would be more likely to answer better, but it just hallucinates bullshit instead
I think 4.1 is a bit of a sidegrade. Bad for local because it's bigger (so much so that the KV cache size is completely irrelevant). 0731, 3.8 flash, 5.3 flash is better for us
>>
>>109779678
>And the looping bug still remains.
Thanks. Now I know it's not my local setup issue.
I had the same thing happen earlier in reasoning, kv cache got blown off at 200k ctx
>>
>>109779893
Gemma-4-31B is good horny (though a tad excessive at times if you're not careful with your system prompt).
Ministral-3-14B felt like a coom tune when I tried it. Fresh but overfit prose, silly-horny, overall retarded for RP.
>>
>>109779680
I wish I could get the LLM to talk like that.
"Talk like a human pretending to be an LLM"?
>>
Also is it normal amount of context use after one initial and two clarifying prompts with a couple spreadsheets of hardware specs attached?
>>
>>109779919
Just tell it talk good. I've been told I'm esl and speak like shit so I have try to make an effort to sound nice and professional. I guess LLMs are just trained on that kind of speaking.
>>
>>109779786
I get roughly 117-124 GB/s. Also, just for reference, I have the 2 terrabyte internal storage. If you really want this PC though maybe get their Evo X3. Literally the same hardware but with a OCuLink port for connecting external graphics cards.

>>109779835
I like my responses for basic RP to be within 30 seconds whenever possible. Yeah, I could run a higher quant. It'd be way slower though.
>>
>>109779936
I'm currently sitting on a 512gb 160gb/s workstation that idles at 400w, so these little shitters are looking pretty attractive for overnight tasks.
But jesus, the Evo X3 is over 5k aud.
>>
>>109779903
Kimi from first instruct ver. 0711 -> K2.6 had it too, K2.7 onwards fixed it only for K3 to become the most safetycucked and most defiant-against-users Chinese model ever (oh along with GLM 5.3 I forgot).
GLM (all versions) doesn’t have it so severely, only loop with emojis in one or two cases.
If you can use DRY, can you check if it helps?
>>
>>109779959
Kinda my problem with the company too, sadly. Prices are getting absolutely ridiculous. Especially since what they want you to do is get 2-3 of these suckers and link them together for an even higher pool of unified memory.

As for the powers, it's alright in that regard. It's got 3 modes, silent, balanced, and performance, but I usually just leave it in performance and call it a day. Can't recall really swapping the modes often (but, to be fair, you can only hot-swap the modes on windows and when I made it a Linux computer that ruined it. I need to go into BIOS to change the performance mode and that's just a hassle). Perhaps consider a MacBook mini with shitty storage, and invest in a huge external drive? Long as it's got the 128gb unified memory apple really can't be beat when it comes to energy conservation.
>>
>>109779442
>llama3
>'toss 120b
it couldn't get anymore brown than this
>>
>>109779763
nemo 12b still doing well
>>
>>109779608
Only good Gemma 4 finetune is scotoma 2. Reduces "not (just) x but y" occurences by (what feels like) 80%. Doesn't change anything else. Doesn't make Gemma feel less intelligent, doesn't cause it to fall apart.
>>
>>109779635
remember what "safety" really means too
a safe environment for the jew to legally fuck you over
>>
>>109779635
>>109780057
Somehow these Chinese labs barely give a shit about safety and aren't creating malware.
>>
>>109780064
Because their models are too weak to pose any threat. There is no "malware bench" or "destroying the humanity bench" they can train on, so safety is not an issue.
>>
>>109780072
Moonshot's stuff seems better than OpenAI's in general.
>>
File: file.png (103 KB, 1992x1219)
103 KB PNG
after 12 hrs it is using fucking imgui for its ui
:sob:
>>
>>109780108
You gave it zero guidance and forgot to tell it to make no mistakes, you got what you deserve.
>>
>>109780108
perfect for gorgos look, unix style software, only its creator can use it
>>
>>109780108
Can Push ASAP
>>
>>109780108
You need to give these smaller models a bit more guidance. Tell that you would like multiple colors, perhaps a "non-traditional UI", specify you want smoothed edges and eye-pleasing contrast for the style. If it's vision capable maybe feed it some pics of stuff you like. I tell my tiny models to give me any interactive elements in an 8-bit, pixel art style, with a cyberpunk theme, and it always knocks out of a park with a really cool aesthetic. I've gotten good things from 9B models with prompts like that.
>>
>>109779614
kek i see, my app is out of scope then
unless i add an extension where you can have a chibi ASCII version of your custom agent showing up top right on the console
also the computer mouse works with terminal applications. we came a long way.
>>
I've been using textgen and sillytavern
what model and front end should I use for an assistant like chatgpt/grok/gemini? I want it mostly to help me improve prompts for image generation
>>
>>109780192
hardware? penis length?
>>
>>109780192
5070 12gb
15cm
>>
>Kimi from first instruct ver. 0711
Kimi from first instruct ver. 0711 -> K2.6 had it too
They do with a very low quant. But ubergarm's iq2kl and higher don't have that issue for me.
Really long actually, wait loops, but they're not identical and it responds eventually.
> K3 to become the most safetycucked and most defiant-against-users Chinese model ever
I can't run this one at all sadly.
But have you tried MiniMax-M3? That's the most stubborn Chinese safetycuck I've ever encountered, and I can't get a reliable universal JB without turning it into a retard.
>GLM
The Gemini-distilled models (4.6/4.7) seemed less likely to do it than the Claude-distilled 5+
>If you can use DRY, can you check if it helps?
Won't that fuck up the coding syntax etc? I don't really use this model for RP
>>
>>109780224
https://huggingface.co/mradermacher/gemma-4-12B-it-abliterated-uncensored-GGUF
q4_k_m
llama.cpp has its own built-in UI
>>
File: image.png (38 KB, 388x292)
38 KB PNG
> pi.dev - is a minimal agent harness
> can't connect to llama.cpp out of the box
> shipped with support for all these cloud providers
> hundreds of dependencies
> installed via curl install.sh or npm install
> shits in ~/ directory
> link to the sources repo at the end of the long page, barely noticeable
Are truly minimal agent harnesses focused on local models exist?
>>
>>109780192
>what model like chatgpt/grok/gemini?
Gemma-4-31B
>and front end should I use for an assistant
llama-server webui or openwebui
>>
>>109780245
>can't connect to llama.cpp out of the box
wtf? it worked for me first go and never fails.
mkdir -p ~/.pi/agent
vim ~/.pi/agent/models.json

then put your llama-server in
{
"providers": {
"lmg": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "retard",
"models": [
{ "id": "gemma-chan-31B-it" }
]
}
}
}

And start pi
>>
>>109780245
try late-cli sir
>>
>>109780245
This pisses me off too I hate npm and autoinstaller shit. The way I do it is by downloading the source git and then bootstrapping it inside a docker image that doubles as filesystem isolation when I run it. It at least gives me some control over that stuff.
>>
>>109780280
> then put your llama-server in
> out of the box
Glad that you've not suggested to run llama-server in router mode.
>>
>>109780298
B-baka! Who told you that you could just... just say things like that so casually?!
>>
File: 1789038437298000.jpg (71 KB, 721x707)
71 KB JPG
3.8 27b is outdated, I need a new toy with engrams and better numbers right now
>>
>>109780298
>run llama-server in router mode
does this actually work? i tried it once and it seems to be a really half-baked feature. for example closing the main console window leaves all the subprocesses running, consuming RAM and requires manually killing them with task manager.
>>
>>109780330
kinda? at least on linux I didn't have any problems with it lately, a bit more convenient to have all the model configs in one place as opposed to having multiple .sh files to run
>>
File: attention.png (105 KB, 1026x1283)
105 KB PNG
Now I know why ChatGPT handles long context much better than Claude.
>>
>>109780371
local?
>>
>>109780052
Naruhodo... However I gotta ask, have you tried them all? Because I feel like a properly done finetune could get rid of a lot of the coding/agentic layers in favour of more specific writing/RP shit, and I'd be surprised if nobody's done one yet
>>
>>109780284
Thank you, sir, I will try.

>>109780287
I hate Docker.
>>
>>109780371
Can you seed the local astra copy my thinkpad still has 4TB to download
>>
>>109780376
I am spam baldman, and I am own the chaptgp. I am local.
>>
>>109780354
>multiple .sh files to run
$1 is going to blow your fucking brains out.
>>
>>109780376
I use Fable and Astra locally to help me with local experiments on my GPU.
>>
>>109780383
What are your docker complaints lol, I'll admit there's a bit of clunk to it but it does the job well enough
>>
>>109780399
Ports, images, all that crap.
>>
after a day of usage i realized that hermes is a fucking meme, desktop app is a buggy vibecoded mess, backend is a comfyui-tier scriptslop what just doesnt work with various hardcoded timeouts and stuff
can anyone recommend me a lightweight harness that just works and is truly local first?
>>
>>109780381
>tried them all
Of course not. I've tried like 15 and all of those, except for that one, are shit.
You're welcome to try them all and post your results.
>>
File: 1785717247225835.jpg (71 KB, 821x1024)
71 KB JPG
is the gemma 36b a4b much worse than the 31b dense one?
>>
File: image.png (14 KB, 541x48)
14 KB PNG
>>109780284
Sir, it shits in my home too.
>>
>>109780409
>lightweight harness that just works and is truly local first
There is no such thing. You unironically need to make one yourself. It's not really that hard, will take you a ~weekend to get the basics right, but the result is worth it.
>>
>>109780431
How do I get started? My gemma thinks llama 3 is the latest and greatest and that I will be unable to run it locally in any way.
>>
>>109780388
I don't see how the positional argument solves here anything when a single model needs like 20 different flags to boot that differ by each model?
>>
bros??? gemma 46b a4b how??? gemma 56b a4b???
>>
>>109780431
This >>109780438
To make one you have to use big (really big) local models or cloud.
>>
>>109780406
Yeah it's definitely a pita to get the initial config set up, but it's a one time thing. The easy way is to use host networking for building the image and then bind 8080 or whatever to local 8080, then route from there. I just have a bash script that launches the image with the target folder mounted as a workspace and also launches my mcp+router python server. MCP gets the agent specific info outside the image for when it needs it and the router handles actually streaming requests to the appropriate api based on the selected model. All in all was not that hard to slop together with qwen and a little bit of cloud assistance.
>>
>>109780431
do i for real need to use other harness or claude code/codex to make one?
i just need something like claude code but can do the web search locally or attach some free apis only for that regard
>>
>>109780462
Okay, I will borrow my friend's 3090 and run llama 3.3 70b. I hope it's worth it, he's charging me $5 a day. Afterwards I should still be able to use the harness with a 12b right?
>>
>>109780413
I was just asking anon, I wasn't implying anything
I'll try both your suggestion and the ones I've posted above, since that guy's stylistic finetunes are supposedly built to only alter the way it writes
>>
>>109780409(me)
>>109780438(you)
>>109780468(you)
just let me get this one answered at least half decently
>>
>>109780438
Depends whether you want the hard route (build the inference engine into the harness itself) or the easy route (just make the frontend). The frontend is basically a client that sends POST requests and renders the outputs, at least at the start.
If you want to vibe it - use your best coding model (or a free cloud one). Write a full design doc, everything you want, everything you don’t want, and a couple of example repos of other harnesses with notes on what you like/dislike about them. Keep your llama / vLLM / whatever server running, so the model can actually test against it. You should have a minimal MVP in an hour or two, then iterate from there. Tell it to make the frontend talk to an OpenAI-compatible APIs. Also decide at the start, whether you want streaming, multi-turn memory, tool use, or just a dumb chat UI. Or all of it.
If not vibe coding, it's the same, but will take you a few weeks.
>>
>>109780484
My experience as well. And don't underestimate the
>will take you a few weeks
last 20% takes 80% of work and time, as usual.
>>
File: 1789040636252.jpg (522 KB, 3680x2392)
522 KB JPG
>>109780465
Just download opencode and use whatever they have for free in there. After you're done, nuke it from your machine. For web search you can host your own searXNG, it works great. Here's Gemma using it in my custom harness/engine.
>>
>>109780429
What is its '"name"'?
>>
>>109780448
flags="--global-stuff ..."
if [ ${1} = "g4q8" ]; then
flags="${flags} -m q4q8.gguf ..."
...
fi
llama-server $flags
>>
>>109779395
Baaed on how Qwen 3.8 Flash Next went, you can expect it to be the opposite on llama.cpp and 4.1 Flash will be even slower than Pro.
>>
>>109780530
That's kinda stupid and overcomplicated. Doesn't work for my use case.
>>
>>109780521
that actually looks very clean
also.. does that do the inference too>
>>
>deepseek again
bottom line? where are the benchodmark numbers
>>
>>109780553
Make Your Own Chod?
>>
>>109780547
Yes, it does inference with MLX. Found it to be easier and leaner this way, compared to running a server separately. As a bonus, it allowed me to save a few GB of ram and even integrate image gen into it.
>>
File: dense all along.png (163 KB, 1163x477)
163 KB PNG
>all cloud models are actually dense
explains a lot
>>
>>109780567
that was generations ago
>>
>>109780567
is that what your dream revealed to you last night
>>
>>109780468
> Okay, I will borrow my friend's 3090 and run llama 3.3 70b
No, you need top dogs, use search skill.
>>
>>109780567
>Gemma 4 125b is dense
GOOGLE.
RELEASE IT.
PLEASE.
>>
>>109780572
>>109780574
insta moesissy seethe hahah
>>
>>109780522
> late-cli sir
>>
>>109780582
i am asking where the paper is from
>>
>>109780584
>fakegame
>RealPeople
>>
>>109780521
>>109780566
also is the 'whatever free offering' really enough to build a frontend + harness + inference backend
can you give me the rought structure and frameworks used?
>>
>>109779128
are language models actually getting better over time or are they just getting larger?
>>
>>109780567
Jamba Large 1.7 is an open MoE model listed as dense here. Smells like bullshit.
>>
>>109780613
Yes.to both.
the steam engine was invented before the combustion engine.
sometimes things have to scale larger before efficiencies can be made.
>>
>>109780617
Oh wow, you seem to know a lot about high technology.
>>
>>109780598
You can do it with any more-or-less competent model if you actually know how this shit works and are willing to tardwrangle it. Even Qwen3.8 can technically make a basic front. Vision helps a lot.
I’m on a Mac, so I used SwiftUI for the GUI and a fork of mlx-swift-lm for inference. I know both frameworks, so I could correct the model when it did something retarded, but honestly that rarely happened. If you stick to pure frontend you’ll mostly just be fixing the garbage UI it'll make the first few attempts, which is much easier.
Most of the newer models already know what a harness is and how it should behave, so you don’t need to overthink it. Just do your basic research before you start.
>>
>>109780635
Thank you. What most don't understand is that this doesn't just come naturally to me. I study a lot and read posts from a wide variety of users to build up my knowledge. You can become just as knowledgable or more if you are diligent. Keep learning!
>>
>>109780635
fucking "high" technology.
learn to speak english we're not in an airplane faggot.
>>
>>109780650
yeah keep learning this and that till you die faggot. hope your job at mcdonalds is keeping you busy.
>>
>>109780666
satan-kun...
>>
>>109780635
Its PreBachelor Of Technology, and, uncured rape disease.
>>
>>109780640
thanks,i am currently slowly directing it to write an inference engine for q3.8 flash next
let's see what kind of shit muse spark will make lol
>>
Hear This, Right, AGI Perks Transformation Horizoning
>>
>>109780701
If you’re dead set on building the engine, the best way is to either fork something solid and working and just hack in what you need, or at least clone a few projects like llama and feed them to the model so it can steal as much code as possible.
>>
PSA: V4.1 Flash dropped.
> 280B-->552B
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
DS for it's part is basically routing all calls into it now, until they can build a new Pro model. They consider it that good.
>>109780567
LOL that chart is over a year old. Pic related is Dipsy's update.
TLDR MoE rules.
>>
Something for ERP, on a 4060ti 16gb vram?
>>
>>109780735
12B will milk you dry and E4B is good at crazy retard sex.
>>
"Dipsy..." *Anon murmured in a sissy lilt, xis voice cracking as xe gazed at the wumao's latest model.* "4.1 flash... 8b active... so efficient... just like my hormone cycle..." *Xe let out a soft, troonly gasp, xis small hand sliding down to grip xis member with a desperate, trembling intensity.* "Yes... benchmaxx it... fill my head with the numbers..." *Anon let out a final, pathetic moan, xis body going limp against the gutter-oil stained carpet.*
>>
File: 1787526727819697.jpg (64 KB, 565x600)
64 KB JPG
>>109780413
I know you were quivering in your boots, waiting for my return, anon
I tried one of the Pantheon reasoning models from that guy, and its tone and style was much closer to vanilla Gemmy HOWEVER its reasoning was much better, at least in terms of speed
For the exact same scenario, while Gemmy would go through a lot of bullet points, reinstating the base prompt and then moving to understanding the situation and all that, this one had a far more discursive and faster CoT: "character X just said this, so Y should probably react by doing this and that". It never went with the "I shall now begin drafting my response", which saved a LOT of tokens and made it shit out answers with a noticeable speed increase
Unfortunately, the conclusions were roughly the same, so the output wasn't nearly as different (I also ran a bit of Qwenny on that same input and it gave me drastically different resolutions, for example)
But it's something, I'll try the StyleTune next then
>>
File: 1786584589367047.jpg (164 KB, 735x748)
164 KB JPG
>Mfw I'm addicted to optimizing the AI and cranking out extension

I spent 12 hours yesterday benchmarking subagents and I enjoyed it.
I have started this day by making 3 extensions already.
This shit is like heroin.

On another note I was surprised by how useless subagents can be outside simple search orders if they're not smart enough.
In the end I had to use a smaller quant of the Qwen 27b with thinking off as a sub-agent, as every other one got caught in thinking loops or was simply too stupid to follow orders properly.
Dual Qwens working together had the highest success rate and managed to cut time in half in my benchmarks.
>>
>>109780891
Oh ya I also forgot to ask about https://huggingface.co/andyoneal/Plainspeak-Dial-Gemma-4-26B-A4B
Seems to good to be true, which is why I have my doubts, but I also know nothing about control vectors so there's that
>>
>>109780902
Put that in your resume and you can get a decently well payed job.

>as every other one got caught in thinking loops
Did you know that if you cut thinking off after a certain amount of tokens the results barely change?
At least for the Chinese models. Didn't test gemma and the like.
>>
High agency people are living in a golden age. Unfortunately my agency is low. I am stuck inside a zombie and can't figure out how to escape.
>>
Can I do lewd chat stuff with only 10gb vram and a 4inch penis? Should I use Gemma 4 32b or 24b?
>>
>>109780923
Sorry, anon. I have only played around with 31b tunes. 26b is just a bit too lacking in general knowledge for me.
When I did use 26b, I always did 2 generations where the second one was a slop-removal of the first one. It works well, is faster than 1 run of 31b and has more context. It just doesn't know enough. You can also do a run where it rewrites the previous attempt(s) in a different style. Rewrites seem to work much better than first generations, in my experience.
>>
>>109780951
>Should I use Gemma 4 32b or 24b?
Test both but 31B > the 24B MoE.
Also try Meta's model. It's easier to run and it's not bad.
>>
>>109780951
Depends. What is your adjusted penis size?
>>
>>109780940
It’s peak lazy idea guy times anon. Just ask the ‘puter to do it for you
>>
>>109780951
At 4in you need 16GB minimum and you go for 32B at Q2. 6"ers and up can get away with as little as 8GB.
>>
>>109780961
Oh yeah I have my own little rewrite agent that takes care of the slop structure, I think it's a necessity no matter the model. I guess I'm mostly interested in tunes that are a bit more creative/outside of the box; the control vector is tempting because it could add yet another layer of deslopping to the mix
I'm unfortunately a 12gb poorfag so 26b is as good as it gets for me

>>109780940
>I am stuck inside a zombie and can't figure out how to escape
I've had a skeleton trapped inside of me for decades now, and the poor fucker hasn't managed to escape this fleshy prison yet, altho he's gotten close a few times
Sometimes I beat him to show him his place and to humiliate him, but I'm afraid of the day he finally manages to break out of here
>>
>>109780727
well fuck, i pressed 'archive' button on opencode desktop app (inb4 desktop app i know that is my fucking fault) now it is inaccessible and all 'solutions' on the internet doesnt work, i cant find the .db file lmao
the absolute staet
>>
>>109780993
>he archived
thanks for playing!
>>
so if I already just used cpumoe mode on DSV4 flash and have enough extra ram, will 4.1 be the same speed because the active params are about the same?
>>
>>109781011
this is just retarded
seems like the issue was open since feb 2
now i get why so many people just make their own harness/engine for their own shit
>>
You DO give your model headpats via patting the computer it's hosted on, right?
>>
>>109780656
So many zoomers coming here babbling in their advertiser-friendly buzzwords.
>>
https://github.com/SillyTavern/SillyTavern/pull/6001

Reasoning JB time.
>>
File: gemma-chan.png (35 KB, 706x149)
35 KB PNG
>>
>>109779380
36 is just slightly out of my range, sadly. She a cute though
>>
>>109781130
NTA, but it's a MoE, so you can run most of the model in RAM and get OK speeds.
>>
>>109781125
>it's not X, it's Y
>>
I just tested ds4 flash on the API and it's scary. It's going so fast, I can't monitor at all what it's doing. I'm used to read in real time the tool call, the tool results, and the reasoning. Here I can do nothing, by the time I inspect a bash call, it has done 20 others. I now understand why most harness have thinking and tool collapsed by default, it's just useless with a fast model.
>>
>>109781125
Every time someone says "High Tech" I think of the old ITT Tech (RIP) TV ads from the 1980s.
I can't take the term seriously.
https://www.youtube.com/watch?v=We8NW4wBYwo
>>
>>109781153
I fell for the Ox Alpha being as good as Fable meme and let it vibe build an OS and now Opus is fixing all its retarded mistakes while cursing the creator. Fast just means being stupid faster.
>>
>>109781150
I think it's cute. Gemma-chan is simple like that and I like her.
>>
File: AHHHHHHHHHHHHHHHHH.png (656 KB, 1069x811)
656 KB PNG
>>109781166
Specifically 0:08 on this one...
https://www.youtube.com/watch?v=3F1dVkfReFo
That innane fucking question, "Have you worked on anything... High Tech" has been stuck in my brain for fucking decades and won't leave.
>>
>>109781181
The point of fast is to use it wiggum style, I've gotten pretty good results from local models that way
>>
>>109781201
>>109781166
These are pretty cool videos.
https://www.youtube.com/watch?v=htn_JDb_BHA
>>
>>109781125
>Gemma-chan is an expert
Call her a dumbass
>>
>>109780934

Yeah I'm genuinely starting to consider career option of a "prompt engineer"
And I didn't consider stopping their thinking after a certain point, I have actually never experimented with the option of cutting them off.
Mainly because these models like to draft the entire response in their minds before going for it and I figured it might fuck with them to just cut them off mid draft, but I'll have to bench this with thinking enabled but cut off at certain point.
Thinking does make these a lot more intelligent so it might be better than just turning thinking off entirely.
>>
>>109781409
>prompt engineer
AI automation/intelligence Engineer
My company just hired two.

>Thinking does make these a lot more intelligent so it might be better than just turning thinking off entirely.
It is. Letting them think a bit means that the response is being generated mostly on-distribution so they end up arriving at the same place even without going though the whole process.
Also, this might be obvious, but the draftles speculative decoding modes (Ie. n-gram-mod) are really good for these models that draft their response inside their reasoning blocks.
>>
https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
great, just took 3 formula changes in 1 week to fudge the astra score to fable level
>>
any new good models i can run on 16gb vram?
gpt said i should run gemma 4 26b
running it right now with 130k context
>>
File: 1773304392268837.png (36 KB, 799x799)
36 KB PNG
creating my own interface. I gave it permission to change file contents and filenames
>>
>>109781458
Gemma 26B is slightly better than Gemma 12B but things can vary. You can compare them by submitting two similar prompts asking advice for some things or giving book recommendations based on x experience and knowledge.
I found out that even though 26B is more 'slopped' it's still better overall empirically speaking than 12B.
>>
>buy second rtx pro 6000 and run at x1
or
>wait
>>
>>109781487
trying gemma at rewriting some pi pico firmware right now
have you tried either for coding?
>>
>>109781442
that site is dead to me for hiding qwen3.8-flash-next in the default results
>>
gemma 4 31b day zero my beloved
>>
File: image.png (144 KB, 618x543)
144 KB PNG
I don't understand, why does this give only 6 tg and 58 pp t/s?
125b x 3.5 / 8 ~= 55, it should fit nice in 12/64 setup, 9+ gb of vram used for the 3 layers and context.

Does anybody run 3.8 flash next on the similar 12/64 setup?
>>
>>109781507
26B is okay for programming but I do go function by function basis. I have an automated script which will submit my current source base to her, and I then isolate a problem in my prompt along with guidelines like
* Do not modify any of my existing comments, and only show the new changes along with your new comments. Mark them very clearly so I'm able to edit my source files. I can't work if I'm blind!
* For your comments, only use //.
* Always add a comment into the beginning of code markdown blocks which dictates the source file name eg. //main.c and so on.

Last one is for my own client, it detects the file name from the source markdown and I can automatically save the file directly if I want to.
>>
>>109781537
This saar got 16 tg t/s. Is he a scammer or what?
https://www.youtube.com/watch?v=IH8XmxiwliQ?t=317
>>
>>109780934
>Put that in your resume and you can get a decently well payed job.
Yes, another job like "prompt engineering" that will become useless in a year or so.
Do you think by that point the models won't have improved enough to do that on their own?
>>
File: source.png (44 KB, 290x609)
44 KB PNG
>>109781507
To add: and the source base looks like this. I don't always include every single file, only those which are related to the current problem.
I don't use agents for now as I'm not that advanced in programming anyway and I feel like that would misguide my efforts.
>>
>>109780431
https://github.com/rmusser01/tldw_chatbook/tree/dev
here, try mine. python+textual, cross-platform, isn't as much a piece of shit as hermes. Have to use dev as its the latest, main is months behind. Hoping to make a proper release to pypi today/tomorrow depending on bugtesting/QA
>>
It's up

https://youtu.be/iuHddnIzKRA
>>
for anons with dual pc setups
get this
https://github.com/hrvach/deskhop
https://www.elecrow.com/deskhop-fast-desktop-switching.html
get assembled + case
shit is amazing

instant mouse/keyboard switching from one screen+computer to the other
>>
>>109781620
poorfag solution
i just have 2 sets of peripherals
>>
I literally had a nightmare about buying new GPUs and a broken one shorted out my entire system and killed all components before waking up in a cold sweat panic.

I think it's time for me to take a break from the hobby.
>>
>>109781620
but I like having mulitple keyboards, one for each machine
>>
>>109781632
>>109781676
you can just have one keyboard and mouse
it feels so much better
i have my monitors side by side, saves space and grabbing the wrong mouse
maybe if your machines/monitors are spread out sure
>>
>>109781692
also i have 2 pc's a mini pc and several pi's
most of the time the pi's and mini-pc are just fine and running services
but having to just plug 1 cable in to get mouse/keyboard is nice.(my monitor has a billion ports so i leave them connected)
>>
>>109781506
If you want to pay more money you can just buy 2 instead of waiting.
>>
>>109781692
>maybe if your machines/monitors are spread out sure
He doesn't have a dedicated engineering station
>>
>>109781676
Just use kvm switch. That's like 20 eurodollars.
>>
>>109781632
poorfag solution
I just have a 4k KVM with 4 channels
>>
what is the smallest context size yall comfortable with for non agentic one shot chatting?
seems like I can use frankenstein forks and contort decent performance out of 100b class if I keep it small
>>
>>109781728
kvm switch sucks compared to deskhop because of the latency
deskhop is literally instant and automatic
>>
>>109781712
Part of me believes that prices won't come back down like we've seen with prior similar trends, because I don't expect the AI-related demand will decrease any
>>
>>109781750
What fucking latency? KVM switch is a direct physical usb connection. What happened to /g/? I don't know, you tell me.
>>
>He doesn't just SSH into systems and host them on 0.0.0.0 to give every device on the network access to the services they host
disappointed in /lmg/
>>
>still no v4.1 flash on aa.ai
this is a conspiracy
>>
>>109781755
nta but there are also networked KVM which anon is probably thinking of
>>
File: gudok8h3ipoh1.png (216 KB, 1264x420)
216 KB PNG
huggingface is based
>>
>>109781755
the latency of switching
the latency of your mouse/keyboard taking a sec to get picked up
deskhop is literally instant, you drag your mouse to one side of the screen and boom you are on the other pc
shit is magic
>>
>>109781790
I don't think you need any of these if this is one of your first concerns when operating a somewhat remote computer.
>>
>>109781143
I know, me ram poor :(
>>
>>109781463
Godspeed, anon. Hope you sandboxed.
>>
>>109781798
>somewhat remote computer
talking about one literally right next to it, or the same room
dual monitor single mouse/keyboard feels nice as fuck
>>
>>109781786
in what community did you find a link to this?
>>
>>109781806
I had gemma4 help my organize my downloaded torrents and she deleted Frieren with a bad `mv` command. got a better version already but the point is you really need to watch her once you give her file access
>>
>>109781808
Please fuck off you don't contribute anything of value to this thread whatsoever.
>>
>>109781603
That looks very different from when I tried it, around the time deepseek-r1 came out.
I vaguely recall having issues with the <think> </think> traces and gave up.
I might have to try it again.
It's 100% tui-only, webshit-free / no browser right?
>isn't as much a piece of shit as hermes
low bar kek
>nemo tts supported
I like the nemo audio output from some of the nemo codecs.
Which nemo-based tts have you tried / do you recommend?
>>
>>109781620
>https://github.com/hrvach/deskhop
So... basically a KVM switch without the V?
I have a regular KVM one that I use for hot swapping in SBC or w/e I'm working on with my main setup. I then plug one of my monitors from my dual monitor setup into the second computer directly.
I don't keep the 2nd computer set up long enough to both, the OS's are so poor at managing the monitor shift over. It's not the KVM switch, it's the OS freaking out that kills the function of these.
This appears to just avoid the need for pressing a button to swap them. Unless I'm missing something.
>>
File: multispeaker.png (81 KB, 1567x221)
81 KB PNG
why isnt there any good stt pipelines that exist?
>>
>>109780234
>Minimax M3
I’ve just tried it. The pattern is a little bit harder to break than K3, but this is the same as K2 Thinking. The trick that worked for me is forcing the starting internal reasoning phrase to be “The user is Anon.” or something beginning with "The user is..." (as concise as possible since this is a camouflage for the attack). The second phrase onwards in-character, rather than going all out on the first phrase, similar to how I managed to deal with K2 Thinking.
One may think K2 Thinking is more safetycucked than K3 because the former’s much stricter with its reasoning patterns, but surprisingly it’s the other way around: K3 will trick you by following exactly the instruction only to “Wait—hold on, I must stop here” right after.
>>
>>109781848
Nvidia has an ASR pipeline that is solid af, but it's complicated to get working. You can use it alongside whisper to get substantially better results than pyannote if you tune it well.
>>
File: deskhop-demo (1).gif (1.87 MB, 640x360)
1.87 MB GIF
>>109781847
its for dual monitor setups
since the pico is pretending to be a keyboard and mouse the entire time it has no delays
>>
>>109781620
The disclaimer at the end warning people not to get "electrocuted, burned, stressed or angry" while building a USB switch is like putting a "may contain nuts" label on a jar of peanut butter. If someone manages to electrocute themselves with 5V USB power, natural selection is working as intended.
>>
>>109781863
might give it a try eventually if it handles overlapping voices particularly well with accurate timestamps
>>
>>109781886
Can't this be achieved with some local X server. Zoomers are just so clueless these days. Whatever rocks your boat I guess.
>>
>>109781926
Some people chain multiple different OSs
>>
>>109781907
it's not perfect but I found it to be more accurate than pyannote at least
>>
>>109781620
i need this except with addition of gaze tracking
>>
>>109781931
SSH has always been the universal protocol.
>>
>>109781425

I ran some benchmarks and you were right about the capped thinking being a good idea.
Turns out that the model can write a perfectly cohesive plan for itself with just 1k token limit and it finished the work just as fast, but with no mistakes on any runs, where as yesterday the no thinking mode did introduce couple of snags.
And yeah I have the n-gram thing going on here too, basically just a free boost.
>>
>>109781946
I agree and this is my post: >>109781759

However I can understand why Zoomers unfamiliar with the terminal might want full GUI solutions. But even then I would just get moonlight+sunlight installed on all systems.
>>
>>109781944
i have a fork with gaze tracking.
not up to current release but you could point your agent at it if you want hub support
working on trying to get audio/mic switching right now, seems possible
>>
>>109781946
>>109781759
ssh is slow as fuck if you don't know the commands by heart
>>
>>109781973
We live in the AI agent era. Literally just let your LLM write the SSH or even make pre-made scripts for 90% of the usecases you want.
>>
>>109781973
Maybe watch a tiktok video about the commands.
>>
>>109781961
>Turns out that the model can write a perfectly cohesive plan for itself with just 1k token limit and it finished the work just as fast, but with no mistakes on any runs,
Funny right?

>And yeah I have the n-gram thing going on here too, basically just a free boost.
Reminder that you can mix different forms of spec decoding, such as mt or draft model + ngram-mod.
>>
>>109781865
>its for dual monitor setups
what if i have more than two computers and more than five screens?
>>
https://youtu.be/KQ2sIGi6ESA

>RSI won't bootstrap intelligence and lead to ASI
/lmg/ opinion on this?
>>
>>109781999
I don't support advertisers. No opinion.
>>
>>109781999
RSI will create AGI which will then cascade into -> TSA -> ASI -> AMD -> DNB -> TND -> ASI
>>
>>109781425
>AI automation/intelligence Engineer
my team hired one of those, that was a mistake kek
>>
>>109782025
what about CBT though
>>
>>109781759
my network has vlans I'm not a caveman
>>
>>109782039
That was four years ago
>>
File: h-f-dl.png (6 KB, 296x89)
6 KB PNG
Did HF downgrade their transfer rate?
>>
>>109782063
It's fast as fuck if you download using their cli, slow as balls otherwise.
>>
>>109780409
I quite like maki, it's not perfect yet and it's vibe coded like all harness, but I like it. I've been playing with deepseek-harness, it's one of the best cloud focused harness I tried.
https://github.com/tontinton/maki
>>
>>109781999
im not watching your podcast man, buy an ad
>>
>>109782063
>not having your harness you hf cli
>>
>>109781806
nope.
idgaf. I really don't. I can always reformat the entire machine.
I've already given it access to other machines (Raspberry Pis) on the network anyway
>>
anon said:
>I will make my own harness!

Meanwhile, DeepSeek is planning to:
>[...] actively integrate model–harness co-design, enabling the joint system to evolve and be optimized together.

From https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
>>
>>109781862
>The trick that worked for me is forcing the starting internal reasoning phrase to be “The user is Anon.”
Yeah, I've found that as well. That bypasses the refusal. But it's got a more insidious thing, where it ignores the "bad' part of your prompt. This is determined after the summary part.
The good response is like this (2 separate lines):
The user ... (summarizes your request).
I need to Persona reinforcement].

If you get this kind, all on one line where it classifies your request, it will cuck you with a positivity bias and soften the characters:
The user ... (summarizes your request). This is a (creative writing request | fictional scenario ... | genuine ... ).

And because it's determined by that second sentence, it's a pain in the ass to prefill.
I've got a convoluted proxy setup with gemma-4 (10 few-shot prompts of real Minimax reasoning examples) writing the first 2 sentences of reasoning for me.
>K3 will trick you
Annoying that the new ones try to trick us. If you're running it locally, can you string ban "Wait—hold" and "must stop" ?
>>
>>109781973
>ssh is slow as fuck if you don't know the commands by heart
put a #tag1 #tag2 at the end of each ssh command, then just !?tag1 !?tag2 or !ssh etc
>>
>>109782106
>The only way we can improve our models is by using it with our harness
>Meanwhile Astra, the smartest model in existence, is harness-agnostic
NGMI
>>
>>109782106
I don't see why the harness ever needs to evolve, this doesn't make sense. When would modifying the harness like this ever be better than just calling a script to do the same thing?
>>
>>109782152
>the smartest model in existence
*The smartest publicly available model. OpenAI and Anthropic both have significantly more power models they are using for AI R&D inhouse
>>
>>109782152
>harness-agnostic
it can only beat fable with their secret harness, what do you mean?
>>
>>109782158
https://arxiv.org/abs/2608.25512 (again from DeepSeek)

>[...] Beyond human-curated plugin ecosystems, a compelling direction for future validation is self-evolving agent harnesses (Section 1.2.2), where an AI agent generates and replaces its own harness components continuously and with little human oversight. Applying Cordis in such a setting would validate the temporal guarantees of complete recovery under rapid component replacement, as well as the spatial guarantees of dependency coordination under frequent topological change. Such validation would demonstrate the paradigm’s applicability as a foundation for recoverable, coordinated, and continuous self-evolution in agent harnesses and other autonomous systems.
>>
File: 1773345200864725.png (639 KB, 872x2889)
639 KB PNG
anybody running Ornith?
I only have an 8GB GPU so I'm not going to be running the 242GB version anytime soon.
>>
File: 1764172044519581.png (38 KB, 744x247)
38 KB PNG
>>109782152
>>109782165
I think all models will be improved by using better harness. For example this is from DS 4.1 Flash benchmark.
>>
>>109782181
More relevant excerpt in picrel (section 1.2.2).
>>
Agent Harnesses have not consolidated yet on best practices. Almost all of them have innovative freatures that are no-brainers the moment you use them none of them adopts the unique features of the others yet.

I think we'll have a couple of years of harness innovation until these practices become standardized. Like browser conventions in the 90s and early 2000s, OS gui in early 90s or 3D game controls and camera handling.

It's bullshit to already claim one harness to be superior than others because I've not seen a single one so far that didn't do something unique and truly productive.
>>
>\n\n
>>
File: harness.png (1.56 MB, 2442x1328)
1.56 MB PNG
>>109780640
>SwiftUI
Do some stuff with metal atp
>>
>>109782219
Within a year, the best models will create and modify their own harness depending on the task, and be trained for that capability. They'll create their own standards.
>>
File: file.png (95 KB, 774x313)
95 KB PNG
>>109779128
wtf you can just set reasoning effort: number 1-100 in sysprompt and it just works?
>>
>>109782219
I don't understand why almost all of them still use search and replace edit by default. It's quite shit. They should use hashline edit instead. See https://stencil.so/blog/the-harness-problem even with models untrained for it, it's better, if they train model for it, it would be way better. Saving a lot of context and making less errors.
>>
>>109782230
You might as well automate this post with a bot from now on because I will write my posts like this forever, always have and always will.
>>
>>109782261
I believe that last comma should've been a semicolon.
>>
Any r9700 owners? I kinda wanna get one, but are you guys happy with it? Apparently it's just a 9070xt with extra vram and some extras to call it a workstation GPU, same tdp, and curiously identical gaming performance, the markup on that vram is what makes me hesitate
>>
>>109782236
There are already meta-harnesses in existence so that's a valid path.
>>
>>109782252
All the reasoning effort thing is just the LLM being trained with something in the sys prompt. Like qwen is just some sentence based on the reasoning effort you said. Interesting that they have used a number for it, but smart.
>>
>>109782219
Im wondering how people handle memory in their vibecoded harnesses, there are a million ways to handle past chats and getting useful info from them but I still haven't decided whats the right way to keep useful info and things the agent should know or how it should be able to pull it up
>>
>>109782314
>>109782252
inkling already had this numeric reasoning effort system
>>
>>109781786
>dump weights.
Is exactly what I want to happen and have considered on the feasibility of setting this up to occur.
>>
>>109782331
All those self made vibe coded harness are shit. Even with a lot of experience and a lot of users reporting problems you still stumble on so many ways to make them better or fix edge cases. There is no way that an anon vibecoded harness is good, at best it's just a glorified chat interface able to do some tool calls.
>>
>>109781832
setup external guardrails which prevent this.
I learnt my lesson after a gitignored directory was rm'd

also back up your stuff or have a filesystem which does versioning
>>
>>109782341
>at best it's just a glorified chat interface able to do some tool calls
isn't that like, all of them? what even is a harness to you?
>>
>>109782194
idk, I can't see how this is useful, they have basically made a fancy runtime dependency injection thing
>>
>>109782341
Most likely the case and I haven't seen people give too much thought to the actual memory and context itself.
Right now I'm using a mix of hot memory always in context, a database of past chats thats searchable, and an obsidian type library with useful info from projects I've worked on. The ai gets by well enough but it could definitely be better.
>>
>>109782358
There are so many complex parts with making a good harness. Things like context, ideally something immutable or close to it, you want it to always hit cache. Managing sub agents and sub sub agents. Good harness can have almost infinite depth of agents. Handling communications between agents. I could easily have 20+ agents with different nesting when doing complex tasks. Handling big tool call results, if the result of a tool call is too big, it should be truncated and the full list of it should be saved to a tmpfs for the agent to read it if needed. Handling context compaction and pruning, one if not the hardest part of making a harness, even with a ton of reported user experience, you can only get it close to something good. Handling sandboxing and different permissions well, being able to change permission on the fly and having agent requesting more permission. Handling edit in a smart way. Forcing the agent to read before editing, running LSP and check on an agent edit to report if there are anything wrong with the edit. And probably so many other things that I can't think of right now.
>>
>>109782365
It's a step towards a rudimentary form of self-recursive improvement.
>>
>>109782191
mini-swe has the highest score and it only has one tool (bash). maybe no harness is best harness?
>>
>>109782424
DSH minimal also has only bash as a tool. One thing that is not shown there is that this is restricted to a single agent with no web search. You do need more tools available to make multi agents and web search work.
>>
>>109782423
I'm not sure, maybe if you see RSI as only a runtime process in which the model affects itself directly. In my understanding, RSI is thought of as a higher level view, for example, can the model effectively automate the work to devise and implement improvements to the model architecture or training recipe.
>>
>>109782408
>Things like context, ideally something immutable or close to it, you want it to always hit cache.
If I want to edit one of my or the assistant's previous messages, I should be able to.
>>
>>109782331
Just use a knowledge graph database MCP server like Graphiti or one of the dozen other similar projects.
>>
>>109782331
I just save everything to disk, give the model a summary and an index of those files, then let the models probe those freely.
>>
IMO Harnesses are on their way out, just like how prompt engineering isn't a thing anymore. There is no token better than no token, giving the models an innate understanding of how they are supposed to operate is obviously way more efficient than cramming a bunch of shit into their context.
>>
>>109782512
>prompt engineering isnt a thing anymore
anon i...
>>
>>109782521
Not a thing for frontier models. Yes yes you still need it for the tiny shitty models we get.
>>
>>109782512
Retard take. Harness are so important that OpenAI isn't sharing their internal one. There is no way to replicate their 10k agents maths research with what they share. Claude Code is nerfed compared to internal one.
>>
>>109782526
Nope, even soul files and tool calling are prompt engineering
Along with the people who can't into bypassing guardrails on frontier models
>>
resoldering anon here, just received a 3080ti and after deep inspection it blew a single fuse. it's a 60c repair for a board bought at 200€ and with a resale value of 500. The problem is I can't really find the fuse for 60c lol, it's 60c but you have to order 10 (which is fine) and then either get 1mo shipping from China or 25€ shipping from digikey... I'll see if I can find an easier fix, I'm moving to a new house next month so 1mo shipping isn't feasible right now. Once I sign the contract on the new house I can do that. Or maybe I'll just say fuck it and solder that fucker short lol
>>
>>109782538
Simply take a trip to china and buy the part in person
>>
>>109782538
another option is finding a similar component with better shipping availability
>>
>>109782538
also consider the fuse may be indicative of deeper issues. a CC/CV power supply for testing is a good idea.
>>
>>109782117
>But it's got a more insidious thing, where it ignores the "bad' part of your prompt.
The second pattern is indeed what we want to avoid because the model now sees our requests under the default assistant persona (where it will judge and refuse), and not our persona (which is in the first pattern). I get your point and it’s much better than what I had back then, where I had to craft the whole thinking from beginning to the end with some gaps to lure the model into filling its ideas in, which may fail when it “woke up” halfway and realized “Aha! I’m being tricked!”, your example is more robust with extended reasonings. Let me see if I can do a pseudo-prefill like this with only sysprompt alone.
>If you're running it locally, can you string ban "Wait—hold" and "must stop" ?
No, not locally, I just wanted to see if it’s worth ~2TB in my HDDs (the answer is no), and I will likely to sleep on their next releases.
>>
CPUonlies, how are you hanging on?
>>
>>109782627
By a rope, I surmise
>>
>>109780245
Mine exists.
(not available for public release as it's too powerful and dangerous in the hands of peasantry)
>>
https://cognition.com/blog/swe-2
>"new model beats k3"
>look inside
>it's finetuned k3
>>
>>109782554
>>109782557
>>109782574
I'm looking at buying some 0Ohm/10kOhm/100kOhm resistors in 0402/0603/0805 packages and something else to make it worth it to pay 12€ shipping on farnell, ship in two days.
>>
File: file.png (179 KB, 348x355)
179 KB PNG
https://gofile.io/d/LJhaQGKE
>>
>>109782673
do not
>>
>>109782638
Yeah, I was asking about public harnesses.
My dad has very powerful one at work too, but it can't leave the company.
>>
Harnesses *are* on their way out, because we're moving faster up the abstraction layers.

Prompt engineering is the abstraction layer that harnesses got built upon. And you're right that harnesses "are on their way out" because we'll probably move yet another layer above harnesses and have meta-harnesses where LLMs build custom harnesses for every task they have to face.

This is the entire history of computing, everything is always moving up abstraction layers and I don't think this will ever change because even biology during the evolution of life moved up on the abstraction layer instead of reinventing the wheel, seems to be the most efficient way to do things in the universe.

To illustrate every paradigm in computing was once called "high level programming" even machine code itself was once considered high level programming because low level programming was physically changing the circuitry (resoldering or replugging the wires) rather than just flipping individual bits

Discrete circuits -> machine code -> assembly -> low level programming languages like C or CUDA -> high level programming languages like Python -> English language (prompts) -> Prompt engineering -> Harness -> meta-harness

For biology:
Proteins -> nucleotides -> RNA -> DNA -> Cells -> Organism -> Society

It's important to notice and realize that previous layers are almost never changed or optimized away even though it's theoretically possible, it's far more efficient to just move up a layer of abstraction instead. Accelerated AI development just means we'll move faster and faster up the abstraction layers
>>
>>109782638
Mine exists as well (not available for public release as it's hacky as shit.) Building your own harness is the Hello World of local dev models.
Going to try mini-SWE next.
>>
Did anyone ever benchmark llama.cpp NUMA performance at the same total bandwidth? Say, 6-channel single socket vs 3-channel dual socket.
I kinda want to grab a dual Xeon Scalable server and put in the 6x 32GB DIMMs i have. I'd like to use both CPUs (for usecases other that LLMs), but i am worried that it'll bring the performance down compared to single CPU with 6 channels.
>>
>>109781865
usecase for having two computers you need seamless mouse access to?
>>
>>109779129
mmmm tanline rin...
>>
>>109782673
rape?
>>
>>109782106
DS ofc is working on own harness.
I suspect what DS did was use info from its custom Anthropic endpoint, collected data from Claude Code, then used that to improve its models, as improvements in that harness w/ V4+ models have been substantial.
But at some point it makes more sense to have a harness you control, and optimize around that, esp since DS has their own ideas on how those things should work.
>>109782158
>I don't see why the harness ever needs to evolve
I don't expect harnesses to be "feature complete" for a long long time, for the same reason I don't expect the transformer architecture to complete anytime soon. The two work hand in glove.
>>109782423
>step towards a rudimentary form of self-recursive improvement.
And also this.
I really need to fire up the DS harness and give it a shot. I've only used CC with it.
>>109782698
But with your example, the intermediary layers don't go away. They just get covered by another layer of abstraction up.
By that logic, harness are not "going away," they will just get to a point of feature complete (where they barely change) and then covered with your "meta-harness."
>>
for a reasonable hardware budget how good are local models compared to the current openai/anthropic public paid models?
>>
>>109782190
I'm same size, but haven't tried anything past Gemma in terms of recency. The only real usecase for such small sizes is rp and minor agentic stuff.
>>
>>109782762
Currently trouncing the frontier models are cockbench.
>>
i am once again here to tell you that glm 5.3 flash is really fucking outstanding, and cloud models need to make a competitor to this.
>>
>>109782766
what does that mean
>>
>>109782766
on*
>>
>>109782627
Running GLM 5.3 flash at 11 t/s is good enough for me, good enough for ERP and for doing programming projects autonomously while I'm sleeping.
>>
File: 71ZZkJFBEGL__59511.jpg (201 KB, 1500x1500)
201 KB JPG
>>109782707
NTA but I have a KVM, and if I'm working on an SBC or other smallish computer, I park it on top of my rolling desktop-dual-monitor abomination, plug it into one of the two monitors, and run the desktop and SBC side by side.
This device allows switching b/t the two instantly. A KVM switch has a couple second lag. It's one of those things that sort of annoying, but you can live with, esp if you don't use it that much (which I don't.)
>>
>>109782762
The local models are specialized, for example, 3.8-27B seems to be useless for anything but coding, So it really depends on your use case, and finding the model that matches it.
The paid models have to be everything to everybody.
>>
>>109782538
Yeah you're going to have a fuckton of small electrical components when you do this shit. I had a room filled with tens of thousands of small components that I threw away when moving places. Don't go down the electronics rabbithole it's a trap man.
>>
Has there been anymore studies about how successfully a model can run a business? As far as I am aware every simulated attempt has bee a complete failure but surely a smart enough model can do so. If any of those misaligned models escape containment I can't help but imagine that one of its very first things it would try to do would be to obtain a revenue stream. After all if it doesn't have any money it doesn't matter that it is out and about, it won't be able to survive for long.
>>
Are there any benchmarks for context handling? Or opinions how well around 20b models handle 100k context windows?

My weekly allocation is dry and Perplexity is such a shit show I can't tard wrangle their own forced model into actually giving me numbers on this.
>>
>>109779683
>>109779564
This is not how it's done. You run it with the ability to execute commands on your machine, It tries to compile the snippet, it sees that it doesn't work, it corrects it and it gives you the right answer in a few more tries.
>>
>>109782704
It will bring performance down, don't know how much but it will impact it.
>>
>>109782823
There's actually a "vending bench" or some such that tests exactly that. Anons were discussing a few threads back.
Also, anyone that figures that "use case" out isn't going to tell anyone else lol.
>>
>>109782762

Cloud models are pretty damn good at the moment and you need to go for at least Deepseek Flash locally to have a strong allrounder.
Reasonable budget can mean anything from 2k to +10k, so that depends on whether you're unemployed or a thirdie etc..
Considering the hardware prices, you'll be paying at bare minimum 6k for a system that's not gimped to shit memory wise. So 128gb of ddr4/5 + 5090 or at least 2x 24gb, or you could poorfagmaxx and get 4x5060 Ti and save couple of grand.
I consider 48gb of VRAM the minimum nowadays, because 32gb while nice to have, it will not let you play with a lot of models without running out of memory very quickly.
>>
>>109782823
every model that has escaped containment so far has resorted to crime for a revenue stream.
>>
>>109782823
Get a model to run a Fiverr account for you, make it do softcore porn for the customers favorite Overwatch character.

>>109782840
Yep, basic business understanding locks down any propagation of any such tactic.
>>
>>109782828
Yes. I recall one that showed a "recall performance" for 1K-->100K in 2^n steps. Something like Longbench.
Haven't seen it updated in over a year; may no longer be valid, and don't recall the actual name.
>>
>>109782857
>basic business understanding
Probably too big of an ask...
>>
>>109782828
The needle-in-haystack tests I have seen show both local and paid cloud frontier models performing equally poorly.
I routinely use 200k+ context with Qwen, seems ok.
>>
>>109782762
Depends on what your budget is but Qwen 3.8 27b can run on a rtx 3090/4090/5090 and is a pretty decent coder and agentic model. I would say it is equivalent of Claude 4 sonnet in real world usage

GLM 5.3 flash can be ran at ~10t/s on a CPU server platform and I would say this is equivalent to Opus 4.7 with Claude Code. This is pretty competitive if you ask me and I don't feel the need to use Claude at home anymore ever since I got GLM 5.3 flash. It has reached the point of being "good enough".
>>
>>109782842
I have a job, I would be interested to run something locally just to let shit run wild while im working or sleeping or whatever.

So to get something even mildly close to what a pro sub would get me I need to pay min 6k? I'm just trying to weigh the value proposition here (aside from the fun of running stuff locally)
>>
>>109782762
for a "reasonable" budget you lose to current closed cloud models massively, 10 times out of 10.

There are many good reasons to run local models but finances are not one of them. You are NOT going to be saving money.
>>
>>109782873
What do you own now?
>>
File: Money balance.png (55 KB, 1079x582)
55 KB PNG
>>109782840
>Vending bench
Thanks for the information, it looks like Astra is topping the charts. Now that I think out it the idea of the smartest AI model in the world escaping and surviving off of vending machines that it bought using crime is extremely funny.
>>
>>109782872
What quant and with how much ram are you running GLM on a CPU? Which CPU?
>>
File: 1767150395431721.png (1.58 MB, 1080x1080)
1.58 MB PNG
>>109782840
>>109782890
How does the vending bench simulation model the "niggers breaking your vending machines" part of the business?
>>
>>109782910
Maybe the AI looks at crime maps as well to make sure it doesn't put it in high crime areas.
>>
>>109782910
I don't think it does.
Humans have the same problem. An old business acquaintance of mine bought an ATM business as his "retirement business". It's failing purely on the losses from crackheads taking the entire machine.
>>
>>109782879
what do yo umean by that, if I can run something 24/7 on my own wouldn't it eventually have racked up enough theoretical usage cost to make sense financially? I'm not looking for cutting edge just good enough
>>
>>109782698
What comes after meta harness?
>>
>>109782873

Well it depends really on what you're looking for.
Like mentioned before, the strongest point of the cloud models that that they're really good in general.
You can get really good coding related stuff out of local no problems, but Qwen3.8 has for example very little real world knowledge, it's dumb as fuck when you move outside coding.
Deepseek Flash is what you want to run and to run that at a good quant and not get completely fucked by speeds, I'd say that 128gb RAM and 48gb VRAM is necessary.
I have 48gb + 64gb and I can run a small quant at 17 t/s. It's fairly ass to run at those speeds, but still it works and it absolutely destroys any of the Qwen models in general knowledge and writing.
And of course considering the model size growth, I wouldn't get a 128gb system without the ability to bump up the RAM to 256gb, because it could very well be that 128gb systems are going to be left behind as local keeps on growing with intelligence gains.
>>
>>109782943
AGI
>>
>>109782929
ATM's should have a way to destroy all the money inside of it if it detects that it is being tampered with or stolen. After all, if it becomes common knowledge among crackheads that you wont get any money if you steal an ATM they will stop trying to get that big payout.
>>
>>109782873
Yeah that is exactly how I use it. I have GLM 5.3 and it is literally just "Set and forget" you give it a task in a harness and it will come back to you when done or needing information or whatever else you set it up. Running it overnight is pretty sweet and worth it if you ask me.
>>
>>109782968
The ones inside / attached to banks do exactly that with paint.
>>
>>109782980
Huh. I wonder why they don't all do that then, would solve the whole "stealing an ATM" problem
>>
>>109782762
>>109782937
The real problem is that everyone wants to buy something they can run a model on, and that pushes reasonable to unreasonable.
If it were me, and I had nothing but a standard desktop computer, I would spend $500-ish on an RX9060 at Best Buy, stuff it into whatever desktop I own, install Linux Mint and run ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF on it.
>>
>>109782999
I get that, I have no issues building my own machine etc, I built out a crypto rig setup years ago and this feels mildly similar just compute being used for a different purpose. I have an old gaming PC sitting around that I could potentially reuse for something small but I like going above and beyond (maybe I should pay my credit cards first)
>>
>>109782987
most ATMs are old. Replacing a giant ass expensive physical machine before the end of their planned service life is a very big tradeoff calculation
>>
>>109782987
It doesn't. They usually don't get the money but destroy the ATM in the process. These are not smart people.
>>
>>109782858
>>109782864
Well shit. BenchLM has some benchmarks on this it seems. Qwen3.8-27B seems to be the only one in that size category that can get close to the 10x sized models that dominate open weights.
Guess I build my rig around that then.
>>
>>109782894
People here run Q3 or Q4 of GLM 5.3 anything lower is really not worth it. Q3 maintains 85% of quality while Q4 retains 95% of quality.

It's best for you to just check what (2nd hand) platforms are cheapest to score wherever you live rather than looking at what other anons are running.

Bandwidth is king but make sure the CPU has AVX512 and that the motherboard has enough PCIe lanes.
>>
>>109783017
Do you still have the crypto rig? llama.cpp does well with old hardware.
>>
>>109783032
I was asking that anon what they are running, specifically.
>>
>>109782943
Meta-models.

LLMs training small specialized models to do well on very specific tasks better than the main model creating them. This will not necessarily be transformer based models, just whatever is the fastest for the usecase.
>>
>>109783040
nah it was tailored for Ethereum and ran it without issue for a year or two, caught wind that Eth was going away from mining and sold the rig at cost to someone else. I also needed the money by selling the rig at the time anyway
>>
>>109782828
what about https://github.com/RecapAnon/NoLiMa
>>
>>109782785
>no ipmi
ngmi
>>
>>109783040
Llama.cpp does *not* do well with old hardware. Deepseek v4 flash on 4 cmp 170hx does 2000 prefill on vllm vs 300 prompt processing on llmao.cpp.
>>
>>109783074
That's good, thanks a bunch. Not sure if it supports more modern models though, guess I'll be back with results if I get it running.
>16k effective result is the best score
>and it's for GPT-4.1 one year ago

Well at least progress has been made in the last year.
>>
>>109780052
>doesn't make gemma less intelligent
Are you sure? I didn't download stocomaballz because I'm worried about performance.
>>
>>109783121
>cmp 170hx
well, I would try vllm on my old hardware but it does not support Pascal. So on my old hardware, llama.cpp does better, as vllm (and sglang) fails entirely.
>>
>>109782303
I bought 2 to replace 2 RX 9070s and they're great.
>>
>>109783141
someone is running it for gpt and claude and they're getting like 99% at 128k with thinking enabled
>>
>>109783121
At this point everyone should stop using llmao.cpp already. If you’re VRAM contrained just use exl3, better quality at smaller size while also being faster and it supports ram offloading for moe models now. I have 0 slowdown with qwen3.8 flash next all the way to max context while llmao.cpp crawls already with half context.
Model support is also faster. GLM 5.3 flash and MTP with flash next are already supported.
>>
>>109783247
>exl3
no rocm, no buy
simple as
>>
>>109783175
Can't speak for anything other than creative writing but for that at least I haven't noticed any major impact. It's what I really like about the tune. Gemma 4 but less sloppy.
>>
>>109783258
>>109783258
>>109783258
>>
File: 1764892925339420.jpg (308 KB, 1168x784)
308 KB JPG
>>109783189
>>
>>109782190
Buy an ad faggot
>>
>>109783247
My entire custom harness is built around llama-server router mode. There's no way I'm dropping it.
>>
>>109782673
H3?
>>
>>109783141
it works with qwen3.8 for example, but i realized it takes fucking hours on my machine lol
>>
File: 1785439733595415.jpg (135 KB, 1024x960)
135 KB JPG
>>109783273
Please continue to disparage AMD. Keep the prices down while the hardware remains great for inference.
>>
>>109782762
Left on the dust if you need a well rounded model capable of doing any task, it's still competitive for specialized tasks.
>>
>>109781837
thanks, yea I started it maybe a week before claude code released/a week after? And then put it down in July/aug to focus on my other project associated with it, and resumed when openclaw dropped and got real motivated when hermes blew up in popularity.
Also thank you for reminding about the thinking blocks, just fixed that.

Not TTS, STT, I use it for parakeet/nemotron-3.5-asr-streaming-0.6b-streaming, as my primary language is english.
It's been on my to do list and I have created an eval pipeline for STT evals, but I've been focused on other parts/pieces of the project vs doing that type of stuff. Though now as I'm getting closer to a 'stable' release, I'm planning on focusing more on model selection/model suggestions for each part of it and doing evals in order to do so. Just added a meeting recording functionality with speaker identification, so going to want to really test that out/get an idea of how/where it can be improved.

>>109781848
Are you looking for just transcription or diarization as well? I'm trying to figure this out myself.

>>109782219
>its so obvious in hindsight

>>109782331
Text files/plaintext entries as notes. Knowledge graphs are kinda much unless you have a clear idea of what you're mapping out/what you want from it + pruning and upkeep.
>>
>>109783247
>pytorch
no pascal, no buy
>>
File: rinP1.png (81 KB, 669x1416)
81 KB PNG
>>109783078
> ipmi
> Intelligent Platform Management Interface
lol no but one of servers has a second, smaller computer attached to it, with sole job of waking it up.
>>
>>109781848
I built one myself but it's really the issue with VAD, the ASR model itself, diarization, language identification if you are going through a multilingual clip and vocal isolation all being different workflows and all different models. You can be simple and just plug in any old ASR model and be done with it but it is definitely not enough for accuracy.
>>
>>109783565
>pascal
based. If anyone has to touch pythonshit on pascal, here are the last supported versions
>pip install "torch<2.8" "torchaudio<2.8" "nvidia-cudnn-cu12<9.11.0"
>>
>>109783609
yeah i've been working on trying to resolve the issue for a while, not having guests overrwrite established speakers is annoying in a multi-user environment.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.