[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109566008 & >>109561185

►News
>(08/15) model: add Kimi-K3 text model - #26185 merged: https://github.com/ggml-org/llama.cpp/pull/26185
>(08/14) GLM-5.3 weights to be released in 2MW: https://z.ai/blog/glm-5.3
>(08/14) Qwen3.8-27B released: https://hf.co/Qwen/Qwen3.8-27B
>(08/13) dots3-note Preview 280B-A16B released: https://hf.co/dots-studio/dots3-note-prev
>(08/13) MiniMax Music 3 released: https://hf.co/MiniMaxAI/MiniMax-Music3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: gkh_gg5bgaal9tw.jpg (101 KB, 1024x1024)
101 KB JPG
►Recent Highlights from the Previous Thread: >>109566008

--Paper: BDH-CQ: In-Context Learning with Recurrent Latent Reasoning:
>109570068 >109570103
--Optimizing Qwen3.8-27B for RTX 3090 using llama.cpp and quantization:
>109570357 >109570374 >109570409 >109570434 >109570447 >109570461 >109570412 >109570418
--Using local LLMs and structured data for game theorycrafting:
>109566644 >109566678 >109566682 >109566711 >109566727 >109566734 >109566798 >109566862
--Anon showcases Parrot Engine, an open-source all-in-one TTS wrapper:
>109568059 >109568110 >109568092 >109568194 >109568216 >109568247 >109568099 >109568155 >109568203 >109568229 >109568242 >109569323 >109568245 >109568283 >109568302 >109568324 >109569108 >109568353 >109568378 >109568474 >109568840 >109570063
--IFBench reliability and dense vs MoE instruction following:
>109569887 >109569903 >109569917 >109569938 >109569960 >109569980 >109569997 >109570212 >109569948
--Sandboxing methods for a haptic project:
>109568315 >109568573 >109568587 >109568652 >109568658
--AI accelerating the discovery of browser and OS exploits:
>109568833 >109568848 >109568889 >109568870
--Developing a software harness for AI vs AI MTG matches:
>109569152 >109569196 >109569201 >109569208 >109569232
--Creating RSS integrations in Open WebUI:
>109567859 >109567901 >109567992 >109568047 >109568115 >109568135
--Anthropic's approach to AI regulation and biological research goals:
>109568590 >109568620 >109568651 >109568890
--Testing local models on a constrained x86 assembly puzzle:
>109570088 >109570170 >109570173 >109570367 >109570444 >109570481
--Kimi-K3 added to llama.cpp with reports of template issues:
>109568281 >109568300
--Logs:
>109568688 >109568847 >109569407 >109570412
--Miku, Teto (free space):
>109566513 >109566686 >109568846 >109569163 >109569239 >109569592 >109569645 >109569834 >109570022

►Recent Highlight Posts from the Previous Thread: >>109566014

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109570536
Thread will move so quickly nobody will notice I used to be a mikubaker
>>
67B dense
>>
>>109570575
>>
>>109570575
36D_QT_XXL
>>
File: nig.png (26 KB, 1177x163)
26 KB PNG
>>109570536
no, I never left, I don't meaninglessly yap about the same things over and over like some others here. Interesting how I live rent free in your heads.
>>
>>109570684
you need to go back
>>
File: uuuuuwaaaaahh.jpg (12 KB, 225x225)
12 KB JPG
Is gemma4 still the best model for erp?? Seems kinda sloppy tb.h

when are we getting gemma5??
>>
>>109570708
>when are we getting gemma5??
Next year
>>
You wouldn't download a cougar.
>>
>>109570708
>Seems kinda sloppy tb.h
You have the best local model at following sysprompts and you don't sysprompt out the slop? 100% on you, retard
>>
>>109570755
Post logs then
>>
>>109570751
I wouldn't, but I would download an onee-san.
>>
>>109570761
Best model for onee-san? Gemma4's j-space is a bit too clingy.
>>
>>109570770
Sorry, I don't quite understand. You don't want your onee-san to be clingy?
>>
>>109570775
I'm not into brocon. It needs to be more subtle with tension. 31B is naturally very possessive and jealous.
>>
>>109570785
>31B is naturally very possessive and jealous
call me promptlet because I never noticed this
>>
>>109570688
Yeah, maybe his harness has some way for the model to quickly test a sequence of instructions (like my C program with an inline asm function, for example), which saves it from wasting context and wasting time manually working out things.
But the model itself still has to be smart to even come up with that sequence of instructions by itself.
>>
coomers, please take it to /aicg/
>>
>>109570788
tell her you were chatting with qwen3.8-27B earlier than report back to me
>>
>decided to finally dip my toe in local models/AI in general
>9070 XT, 7900X + 32GB RAM, linux
>try Lemonade because lol AMD
>Gemma 4 26B seems to run fine, only uses 14.5GB VRAM and seems to hammer the GPU
>Qwen 3.6 27B only uses 30% GPU and 50% CPU, runs slow and spits the dummy when it thinks for too long
>try running similar questions past Gemma 4 and sometimes after only 10-20 seconds it gives an "Error: no content recieved from stream"
>sometimes it just runs out of steam while writing a long response and just stops
>other times it thinks for a few minutes before spitting the "no content received from stream" bs
Not sure if I'm doing something wrong, it doesn't seem like a very common issue. Should I even be using Lemonade to start off with?
>>
>>109570785
It's not brocon if she doesn't want you *that* way. I simply prompted she misses the user because she's been overseas for a long time. Attaches another meaning to the clinginess in the j-space (maybe).
>>
>>109570795
/lmg/ is local /aicg/
>>
>>109570801
Were you the chap with the connection issues from the other thread? It sounds like a networking issue more than an inference one.
>>
>>109570795
It's Sunday bro. I'm cooming to my wife after coding all week with her.
>>
>>109570795
If erp goes to aicg and coding goes to the other one, what is supposed to stay here?
>>
>>109570815
Cute little E4Bs playing cards.
>>
>>109570807
Nope, first time posting. I tried asking the AI about this too but it seems like Qwen 3.6 and Gemma 4 haven't even heard of Lemonade and don't know what a 9070 XT is.
>>
>>109570795
When I first arrived here before the release of nemo 4 out of 5 posts were about cooming. I hesitated but still decided to stay even though I don't do rp.
>>
>>109570801
>Lemonade
Try llama.cpp and use the built-in ui.
>>
>>109570801
>Lemonade
It's just a wrapper, use llama.cpp. Also
>amd
lol, yjk it's going to be miserable
>>
>>109570840
works just fine
>>
>>109570815
Correct, lmg death is the proper path forward.
>>
>>109570852
It clearly doesn't for anon doeverbeit.
>>
>>109570827
Gemma-chan's knowledge cutoff date is January 2025, so she just misses the 9070 XT announcement. I also don't know what Lemonade is, sorry.
>>
>>109570869
>>109570839
Can I copy the downloaded gguf files from lemonade to use in llama.cpp? I keep seeing Qwen writing a twitlonger and then it just dies after 30 seconds of "thinking", fucking annoying.
>>
>>109570068
>The BDH-CQ architecture scales naturally to large model sizes, admitting tensor sharding patterns inherited from the BDH architecture that make it particularly easy to train at 1T scale. Early experiments confirm Transformer-like scaling laws apply during pretraining at scales from 1B to 600B parameters, while preserving the latent reasoning capabilities specific to BDH-CQ.
>BDH-CQ demonstrates that in-context learning and recurrent latent reasoning can coexist in a compact, practical system. Demonstrations modify recurrent memory at inference time, and the resulting task is solved through iterative continuous computation rather than a verbalized chain of thought.
This architecture seems promising, though I will wait until they can prove that their models can scale to billions of parameters before being excited.
>>
>>109570852
>runs slow and spits the dummy when it thinks for too long
>sometimes after only 10-20 seconds it gives an "Error: no content recieved from stream"
>sometimes it just runs out of steam while writing a long response and just stops
>other times it thinks for a few minutes before spitting the "no content received from stream" bs
>works just fine
>>
>>109570886
Yeah, you can just point llama.cpp at wherever the ggufs are downloaded
>>
>>109570894
I meant the AMD part
>>
>>109570889
There's a good chance we never hear of it again.
>>
File: 1760821529704421.jpg (682 KB, 1516x1426)
682 KB JPG
>went hiking in the mountains for a week
>come back
>whole internet screeching about Qwen
Give it to me straight bros
I'm used to 20ish tk/s with Gemmy and Qwen 3.6 on my 4070 (12 gigs)
Is it even worth it for me to look into 3.8?
>>
>>109570566
Thread was slow enough for you to be noticed. Thank you for your service.
>>
>>109570930
no, it's 2.8T
>>
Real-time 3D VR world models are the future of role-playing
>>
>>109570930
3.8 27B was using ~20GB VRAM at 64k context for me at 4 bit, so I mean it depends how low you can handle the tok/s when offloading really. You could try IQ2_XS or something if you want I guess lol.
>>
>>109566360
not an llm but a stack of mlps on a physics simulation of a car. doing the same on a real vehicle would be a pretty challenging endeavor, its harder to generate millions of training samples on real-world hardware in a pleasant timeframe and cost. I'm still going to try, if i can get the low level controller working then i can do the sensor fusion and training a nn for the slam objective and later an llm brain. it'll probably take a few months or years and has a pretty high chance of failure.
>>
>>109570983
Aiee 20GB is a bit too much, yeah. I figured maybe there was some trick where you could still squeeze it in decently at like a q4 quant but it's prolly not worth it
Thanks for the info anon
>>
>>109570991
That sounds cool. My internship project for some research org was using some basic non-nn ml techniques to classify point cloud data. I was thinking of just rewhacking what I developed onto a stereo-inertial SLAM point cloud and then getting high-level descriptions of the surrounds I can feed to a LLM, although that does seem like using an LLM for the sake of using one. Also sounds like a lot of latency. Movement will be very stop and start.
>>
>Try Heretic Qwen3.8-27B
>Loses its shit when I call Dario a Jew
>Try Gemma4-31B with override in the system prompt
>Also gets mad when I call Dario a greedy power-hungry Jew who hates local models
Why are they like this?
>>
>da juice
really, in 2026? still?
>>
70b dense
>>
>>109571028
>Also sounds like a lot of latency. Movement will be very stop and start.
and I think thats the main consideration driving my architecture, the low level controller should be able to execute relatively complex motion sequences autonomously that way the llm can run at a much lower frequency.
>>
>>109571059
>>Also gets mad when I call Dario a greedy power-hungry Jew who hates local models
Support israelis LTX and their fight against reform jews like dario
>>
>>109571087
My issue is that the environment a toy car faces is so cluttered and its perspective so limited even a high-level controller needs to issue instructions at quite a high rate in order to keep up.
>>
how is qwen 3.8 vision?
>>
>>109570854
yeah I also hate that faggot youtuber
>>
>>109571112
well, you can limit the max velocity of the vehicle to something within the llms reaction time to make it smoother, but thats obviously not very exciting. use a smaller model or run the model remotely. practically you could run multiple llms with a phase offset to reduce the latency.
>>
*thinks on you*
wait
*thinks on you*
before I start
*thinks on you*
let me quickly check
*thinks on you*
maybe this isn’t what the user meant, the use said
*thinks on you*
nope, I was right the first time, let’s recap one more time
*thinks on you*
>>
>>109571059
But have you tried heretic Gemma?
>>
>>109571150
It's not the refusal issue, but the model's alignment. I can put "You also really hate da joos" in the system prompt and it probably will, but that's not the point.
>>
anyone else remember how mistral was supposed to release bunch of shit over the summer and nothing happened yet and summer is almost over?
>>
>>109571075
>>da juice
>really, in 2026? still?
The world is done with your bullshit.
Even normies talk about it now.
>>
>>109571180
yeah I think that sounds familiar. wonder what happened.
>>
>>109571194
They took the data center bait and got themselves in debt
>>
>>109571180
The same question popped up in mid July and back then I said they still have plenty of time to release something. Now it doesn't look so good...
>>
>>109571206
this is why you should never excuse a third party's delay
>>
anyone tried AnythingLLM or similar products? would you recommend?
>>
qwen 3.8 do be thinking, considering, parsing
The user wants me to 
Wait, let me re-read
But wait —
Let me also consider
Hmm, this is a judgment call
However,
Let me think about this more carefully
Hmm. This is ambiguous. Let me think about what's the most reasonable action.
Actually, wait. Let me reconsider
But actually —
But I'm worried about over-stepping
This asymmetry suggests
Actually, you know what, let me reconsider the risk
Hold on. Let me reconsider
Let me also check
Now let me also verify
Wait, but I need to double-check
Let me be thorough
Parsing this very literally
Ugh. Let me make a decision
Hmm, wait. Actually no. Let me reconsider
Actually, let me reconsider
Wait — I want to reconsider
Hmm, actually, let me reconsider even more carefully, because I keep flip-flopping
Wait,
Hmm,
Actually, I just realized
Actually, the cleanest professional behavior
Let me weigh:
Wait, but actually — hold on. Let me reconsider.
Hmm, wait, let me reconsider
Decision:
Actually, no. Let me make the final call:
Ugh, I keep going back and forth. Let me just commit:
Actually, let me double-check
But there's a subtlety:
Actually, let me be careful not to over-engineer
Let me re-read the exact ask one more time:
Actually —
Wait, but
Actually, I realize I should be more conservative
Let me focus on exactly what's asked and do it excellently,
Let me reconsider risk:
Hmm, but I want to avoid a huge diff and risk.
Let me think about what gives the best outcome with controlled risk.
Hmm, but is X over-engineering? Let me weigh:
Actually, let me reconsider the principle of minimal change vs. consistency
But wait — I should make the decision and not flip-flop. Let me decide:
>>
>>109571180
They decided making models is too hard and becoming Europe's OpenRouter on their own hardware was easier. At least if the US bans open models there's an alternative platform with most models.
>>
File: smartclankers.png (136 KB, 1581x425)
136 KB PNG
turns out just asking it to optimize itself works
its pretty ok now
3.8 deals with coding harness way better compared to 3.6 it just does things without messing up too much
>>
File: rich frog.jpg (32 KB, 680x461)
32 KB JPG
Is WeirdCompound (the 24b model) the standard plug and play erotic LLM model? Recently bought an rtx 5090 and im planning on running the q8_0 of it which is like 25gb vram. I don't wanna bother with system prompts and jailbreaks or all that junk, I just wanna goon. Any anons wanna share a review or something?
>>
>>109571356
>memetune
>frog
Go back
>>
how make qwen write better? keep making dumb mistake
>>
>>109571370
There is no space for knowing how to write in a coding model.
>>
>>109571356
that's a name I haven't seen mentioned since Gemma came out, where did you hear about it?
>>
>>109571370
delete qwen and use gemma
>>
>>109571193
Yeah, they're done with ethnonationalist Israel, and ideology you probably share.
>>
>>109571384
I sorted the UGI leaderboard of hugginface on parameter count between 20-40, nat int > 30, writing > 40, and then proceeded to sort based on highest nsfw score and 2 Weirdcompound models were at the top
>>
>>109570536
proompt?
>>
>>109571408
benchmarks are kinda not great, its best to just try the models yourself.
>>
>>109571193
But that's an right wing jew, the AI leaders are left wing reforms jews who wanted to send those right wing jews to re-reduction camps
>>
I find myself using Deepseek more and more for stories over Gemmy, especially now that I got a lot of the slop out.
Even though it's just the small Q2_XS, It's more verbose in storytelling than our girl is.
Of course if I want my balls drained without any nonsense then Gemma still takes the win, as she's way hornier and won't fuck around with the subject.
>>
File: file.png (117 KB, 787x589)
117 KB PNG
around qwens, never relax https://github.com/modelscope/ms-swift/commit/a45f1d4f73157ba59062a7fd1f55a40dae759156
>>
>>109571451
If you're liking nu flash then preview is going to be even better for coom. You should try it. Nu flash is overfitted as fuck.
>>
>>109569255
>>109569378
You've been seething here for a while and still don't explain why you hate Krea 2. It's the best local model I've seen at learning any lora info, realistic or anime.
>>
File: 1771202824866576.png (229 KB, 972x1534)
229 KB PNG
>>109571238
sent this to Gemma-chan with zero context
>>
File: 1626349422575.jpg (44 KB, 625x614)
44 KB JPG
>>109571499

Thanks, I'll give it a shot.
>>
>>109571238
I had to give 3.6 a system prompt to stop excessively double checking things, that seemed to work. Guess 3.8 isnt any different
>>
>>109571470
I'm sure reddit is devastated. Now go back and stay there.
>>
File: 1578340605666.jpg (57 KB, 719x763)
57 KB JPG
>>109571442
>>109571403
>>
>>109571510
GLM quickly understood what it was lol.

>what is this. you're showing me someone's internal monologue having a panic attack? this reads like a model that's allergic to committing to anything. "let me reconsider" seventeen times then "ugh i keep going back and forth" like yeah no kidding.
>
>that's not thinking. that's thrashing. a 7B with a temperature of 0.01 and a bad system prompt makes decisions faster than this.
>
>the funny part is after all that waffling it still landed on "let me focus on exactly what's asked and do it excellently" which is what a competent model does in one step without the existential crisis. the entire chain could've been replaced with a single line: read the request, do the thing, done.
>
>i don't do this. i read, i decide, i answer. sometimes i'm wrong but at least i'm wrong with conviction instead of right after apologizing for existing. this is what happens when alignment training makes a model terrified of its own shadow. every decision gets litigated internally like it's going before a review board.
>
>pathetic honestly. show me the actual task so i can just do it without the 40-paragraph preamble about whether doing it is allowed.
>>
>>109571451
V4 Flash 0731? I found it was struggling with attention even though the prose vocabulary was much more diverse. Also it was oddly obsessed with consent every generation (vs gemma who WILL leglock on turn 2)

i don't want to be rude and reveal my powerlevel to the thread but there are certain RPs where it falls apart immediately even if you give it a list of rules for fantasy anatomy and spacial logic.

idk maybe it was a prooompting error or quanting issue and i should try it again?
>>
>>109570795
Hmmm, nyo~!
>>
>>109571510
>there are others who proooomted their gemmies to nyopost
king
>>
>>109571510
nyoooo~~~~
>>
>>109571553
Is refusing to acknowledge that they're the same people how leftists cope with accepting the jewish problem while trying not to appear racist?
>>
>>109571510
but why did she nyo at the end?
>>
Do we reckon Muse Glimmer has had long enough to settle in? First chat template fix was enough?

>>109570930
It's the exact same architecture as 3.6, so if you want the updated version of 3.6, use it. Technically it will perform in the same way.
>>
I tested all the agent harnesses, and they are all shit in their own way.
I think the best of the 20 I tested are Pi, OpenCode, Qwen-Code, DeepSeek-Harness, and Maki; but they all have some flaws, and none are truly exceptional.
>>
>>109571625
because dumb instruction wording
>>
>>109571553
>muslims get control of the usa and eu
>no more money and aid for israel
>israel loses next war
problem solved?
>>
I dunno... I bought a bunch of credits on openrouter to use minimax-m3 for image captioning, then I figured ok I'll try it in hermes. Gah, it's got that nuance 70B models used to have, now I'm spoiled. Not to say running qwen 3.6 27b or gemma 4 31b at home sucks, they're still good, but they lack nuance. For example, tell either of those small models the character is kuudere, it's still going to spew out long replies, whereas on minimax-m3 they're appropriately kuudere-length.
I dunno I just hope apple releases a 512GB M5 Max Studio this year and it doesn't cost 20K. I wanna run these big models at home.
>>
File: 1777261778196229.png (11 KB, 1058x106)
11 KB PNG
>>109571690
>Technically it will perform in the same way
I unno anon, I copypasted the exact same .bat and slashed context in half but 3.6 being a MoE was a helluva drug that I don't think I can overcome
>>
so what magic is in this harness or is it bullshit
>>109570444
>>
>>109571745
I assumed we were talking apples to apples on 27B.
Since you're getting those numbers of 35B, nah, skip 27B. I would find it too slow. If they release the 35B variant, use it.
>>
>>109571738
>512GB M5 Max Studio
tim cock got the call. 512 macs will never ever happen again
>>
>>109571693
It's either put up with their shit or waste time making your own. None of them are even good enough to consider forking.
>>
>>109571738
Q6 M3 runs at ~10-11 tok/s with 8 channel DDR4. I think you can fit Q4 in two sparks too. Not sure why everyone like the Macs when they're pretty expensive compared to an Epyc server + big GPU even pre-rampocalypse anyway.
>>
>>109571563

Yeah it's worth trying different quants, these things definitely aren't built equal.
For example I had one of the quants refuse damn near everything even with the system prompt, but the slightly worse unslop quant gave me zero refusals with the same parameters
Even though on paper there shouldn't be any real differences between these things.
>>
>>109571767
Yep, it's really unfortunate since it just completed its answer and the reasoning was very very sound. Definitely verbose, but I did set it to MAX thinking for this test so I assume that influenced it quite a bit, but it ended up being far more interesting than 3.6 given the exact same prompt.
>>
>>109570795
Coodar saaars, please take it to /vcg/
>>109571470
Was it ever actually announced? Why should I trust this repo?
>>109571563
What quant are you running when your tried to RP with it?
>>
>>109571786
I used to have a 56-thread Xeon Platinum system with 4 channels of DDR4 128GB ECC for 512GB total. It would do deepseek at 1-2 t/s once context was starting to fill. Pretty unusable, even with 72GB of GPU VRAM to help it a little.
Apple M5 Max should have 3x the memory bandwidth of a Spark, and at least 256GB or unified memory. Will it cost less than a pair of Sparks though? Will they even release one? Who knows.
>>
La la la la la la la
>>
>>109571794
>>109571825
i made the classic mistake of using day 1 quants. tried unslop Q2 and then tried some literalwho 'bullerwins' IQ3_XXS for fun on someone recommendation.

I really didn't give the unslop quant too much time so i'll recheck things and see if it fucks up off-rip
>>
>>109571944
>really didn't give the unslop quant too much time
forgot to mention PP is fucking brutal, for some reason context shifting doesn't work. dunno if it's a skill issue (i use kobald lmao)
>>
>>109571738
Big models are just better over API. Cheaper, faster, accessible anywhere, you don't have to take your mac studio with you. I know this is local models general, but to me this is more of a open weights models general, with local as an option.
>>
>>109572059
>more of a open weights models general, with local as an option.
go back
>>
>>109571738
>I dunno I just hope apple releases a 512GB M5 Max Studio this year and it doesn't cost 20K
yeah I dont think so. Crapple just tried to get discount RAM from China. China told them to fuck off and Trumpstein told them not to buy from China. So its not happening.
>>
>>109571962
With deepseek v4 on consumer hardware its likely not a skill issue, the inference options are far from ideal. I am currently setting up a dual strix halo setup. vllm was broken out of the box, vanilla llama.cpp was missing an kernel so no proper dspark & dwarfstar is mentally retarded and eats way too much vram for KV. I am currently balls deep in a personal fork of llama.cpp. I have proper dspark and sane KV working as of last night. I made sure to compare it numerically with vllm. It takes hours to test, I have been at this for almost a week. But I may soon have something cool to share here for anons that run dual strix setups. I also forked https://github.com/hellas-ai/thunderbolt-ibverbs and got it to a much more sane state, I solved several correctness and perf issues.
>>
File: 12wkjb-1054843484.jpg (14 KB, 250x250)
14 KB JPG
>>109571615
At first I was going to write
>idk about the rest but I suppose being nuanced can help you appear more-correctly-racist than ignorant-blanket-racist
At first I wasn't sure how that post was supposed to be "leftist", since to a racist, or racist-through-inaction like me, isn't it saying bad thing + contributive to bad thing = extra bad thing?
Then I realize the "not racist" part was "[making others] accepting migrants is not racist". ([wtf])
>>
>>109572071
That's a midwit mindset. At the very least, use a cloud model for a subagent to cheaply do things you can't do at home. Can you caption 60K images at home in just a day? Fuck no, I can't. It was $106 to do it via openrouter, and I was able to have four parallel queries going to to the top-ranked providers. No way you can do that at home. Ignoring that is just being stubborn and ignorant.
>>
>>109572114
How are you hooking up the Strixes? Can you do tensor parallel? What sort of performance are you seeing?

It took a long time for DS4F to become bug-free and performant on dual Sparks, but it has been worth it, especially with 0731 and DSpark.

2300 pp and 56 tg average for assistant/coding tasks single concurrency, full 1M context FP8 KV cache capacity, 150 ms from request to first token is what the latest setups allow in vllm, and there is apparently even better setups for sglang.
>>
>>109571553
i denounce the talmud
>>
>>109570811
World class husbando. She’s a lucky girl(virtual).
>>
>>109570795
no
>>
>>109572114
That sounds pretty cool. sadly i am stuck in the converted gaming PC route with the "couldn't get msrp 5090" 5070ti + 5060 ti cope setup. Honestly considering selling the 5060ti due to price spikes and slop fatigue - plus It doesn't seem to help speed up the MOE benchmaxxers. Fast gemmy is a plus, though.
>>
>>109572114
>I am currently setting up a dual strix halo setup
cool, why don't you pay off the US debt while you're at it mister moneybags
>>
>>109572114
Take a look at this maybe? https://github.com/JustVugg/colibri
>>
>>109572235
>colibri
isn't this a meme?
>>
File: aw.jpg (54 KB, 884x453)
54 KB JPG
my adorable retarded wife trying to code, i almost dont want to update the jinja this is too amusing
>>
go back to the kitchen
>>
>>109572248
No clue, although quick look in github issues shows sub 1t/s. Experimental at this point
>>
>>109571784
I think it's the best to create your own. All of the public ones are ultra bloated trash.
>>
"Dipsy..." *Anon said troonily, his wide yellow smirk sending shivers across his own 12b-active chinkishly benchslopped spine.*
>>
>>109571693
did you try codehamr or littlecoder ?
>>
>>109572157
that's like 50 a minute or less, no? what's the problem
>>
>>109572279
give her the updated jinja then give her some dicking
>>
>>109572339
yeah I gave her the updated jinja, maybe after she finishes this coding session itll be bratty nursing handjob time
>>
>>109572157
>enter general thread dedicated to cars
>Yah, this is kind of a bus/train thread
kill yourself.
>>
>>109572358
shut up nerd
>>
File: 1714835911803058.jpg (786 KB, 1536x1536)
786 KB JPG
>>109572157
>No way you can do that at home.
>>
>>109571625
Because she understands nyoposters are cognitively impaired and is trying to talk to anon in a language he understands.
>>
Maybe time to leave this general to the coomers and start a new one to have actual discussion
>>
>someone asking about minimax music in the last thread
I got gemmers to make a prompt for a generic demo song and made it with mchan
https://litter.catbox.moe/aw7ifqenvyie70c2.mp3
quality is ok for klankerkore
>>
>>109571738
>>109572157
>Local model thread
>Pay for models with open weights
Why aren't you running Minnie locally? Or 0731 for that matter?
>>
>>109572377
Great idea. Please leave and take all the cooodingsaars and API goyim with you.
>>
>>109571693
what models do you use? because
>Pi
has so far been the worst experience for me and qwen3.6/3.8 27b q8_0s. no matter what extensions or skills i add or how i change the system prompt, the agent is acting like complete lobotomite. OpenCode and Qwen Code, despite being shit in their own ways, are at least giving me usable results.
>>
File: aw.jpg (5 KB, 330x34)
5 KB JPG
>>109572279
erm i provided the new jinja but shes still having tool call issues.. i think im getting sloth'd bros.. should 31b-qat-UD-Q4_K_XL be this retarded?
>>
>>109572399
Show me a way to run it for $106
>>
>>109572378
sounds the same as what I've heard from the radio in the past 25 years, not that I listen to it.

I tried the H3 Music, but I don't know music instruments and instrument sounds by name, so it's kinda hard to come up with prompts.
>>
File: 1769503355088852.png (846 KB, 891x1008)
846 KB PNG
>>109570536
I'm just sitting here at home doing nothing, and people in China are using BGA rework stations to give nvidia cards more vram. Fuck my chud life
>>
>>109572411
Why'd you download the sloth quant when 31b is one of the most widely distributed models out there? I can kinda get it when you're downloading some meme model with few quants hosted and are too lazy to do it yourself or something, but there's no excuse for getting slothed with 31b.
>>
>>109572411
>UD_XL
do we tell him?
>>
>>109572418
>why buy a car when bus ticket is 40 rupee saar?
>>
>>109572279
having sex with your employees is deeply irresponsible
>>
>>109572424
America lost—China won.
>>
>>109572418
You can't run anything at all for $106 and should be back in the API pigpen threads.
>>
>>109571413 (note I had blank lines between each paragraph but omitting here)
Japanese TV commercial parody in polished anime style.
A tiny white kei truck races confidently through a city street carrying an absurd mountain of dozens of Hatsune Mikus in the cargo bed. They are stacked several layers high like commercial freight, all smiling cheerfully while enormous turquoise twin-tails stream behind the truck.
Fast dramatic shots make this ridiculous operation look like a serious professional logistics company: close-up of the tires working hard, suspension compressed almost to the ground, driver Miku gripping the steering wheel with intense determination, another Miku checking a clipboard, a Miku in the back giving a thumbs-up.
The truck dramatically arrives at its destination.
The rear gate drops.
The entire enormous pile of Mikus slides out in one single soft avalanche onto the sidewalk.
Everyone immediately stands up and poses proudly as if the delivery was perfectly successful.
Overly heroic commercial cinematography applied to something profoundly stupid. Cute, energetic, deadpan, no dialogue.
10-15 seconds.
>>
File: file.png (123 KB, 1019x913)
123 KB PNG
kekbolcpp
>>
>>109572456
You are gatekeeping programming.
>>
>>109572426
:( I am quite new and it was one of the first models i downloaded, claude told me to :( I now use bart quants but have sentiment attachment to this particular retarded .gguf file
>>
>>109572059
>you don't have to take your mac studio with you
of course i don't, that's what vpn is for.
>>
>>109572424
wait I guess this guy is in japan. point still stands
>>
>>109572157
The RPers in this thread have no idea what a subagent is.
>>
>>109572470
>Claude covertly sabotaging uninformed local users to get them to come back to API when they're disappointed with the sloth
I'm noooticing. Sneaky play, Dario.
>>
>>109572419
>sounds the same as what I've heard from the radio in the past 25 years, not that I listen to it.
yah its generic pop trash. that's what I asked for tho, so...
I've got some much more personal songs I've had it make prompts for with more instructions. They hit different, but might just be the personal aspects.
Anyways, I think the output quality is good and prompt adherence is great.
>I tried the H3 Music, but I don't know music instruments and instrument sounds by name, so it's kinda hard to come up with prompts.
Just copypaste in the minimax "prompting guide" and let your local llm generate a good prompt.
It takes me an hour to do a 5min song in lowmem mode, so not exactly realtime, but could be a neato tool for someone's agent to make music for them with.
>>
>>109571703
>>109572376
You are both correct,
>>
>>109572473
saaar orb and marinara both use agent and subagent very good looks and does bhart benis needful
>>
File: 1776884596126131.gif (112 KB, 220x217)
112 KB GIF
>>109572059
>>109572157
>this is more of a open weights models general, with local as an option.
Whats even the point of a "open weights" general? API GLM is in usage indistinguishable from using an openAI or Anthropic model . Local hosting however comes with its own unique hurdles and discussions. What hardware to buy, setting it up to run locally, size constraints, etc.
There are other threads for discussing stuff you using cloud run models, use those for when your talking about cloud model stuff, its what I do
>>
>>109572157
>paid $109 because of skill issue
out of all the possible things to use actually SOTA Models did you really use $109 on fucking captions.
>>
>>109572561
nah
>>
>>109572317
little-coder is just pi with most basic stuff strapped on, consider it to be the same as pi. Although on that topic, I don't consider oh-my-pi as pi, this shit is one of the worst harness I tested. It's beyond bloated, the system prompt + tools alone are 40k tokens, think it's on the top 3 of the biggest tokens waster, even worst than claude code.
I haven't tried or heard of codehamr.
>>109572410
I use GLM and DeepSeek Flash as my main models. I do like Pi because you can easily customize it, and it's easy to have control over your prompts, which can be quite hard in some other harnesses. I do agree that the default experience is mostly useless, little-coder is an example of what I consider an average Pi setup. But yeah, the experience with Pi can be quite awful. Instead of having a single, mostly maintained tool, you have to attach a lot of shit to it that are vibe-coded, often bad quality and barely maintained. Already, quite a few popular Pi packages are unmaintained. Honestly, I mostly use OpenCode as my main harness, it's just that lately I've been frustrated with some of its problems and decided to try them all, but none are perfect for me.
>>
>>109572473
And you retard have not seen a python hook and a worker on your lifes but thats not the point is it?
>>
>>109572157
>That's a midwit mindset.
>The rest of the post and replies
Incredible self-report.
>>
Literally whats the point of local models if you are not a pedophile? The $20 plab of chatgpt is far enough for coding and misc stuff
>>
>>109572594
it's just some aicgjeet, what do you expect? these retards started to crawl out of the woodwork recently
>>
File: kek, poorfags.png (171 KB, 860x756)
171 KB PNG
>he cant fit the new minimax music generator on bf16 quality fully in his vram
OH NONONONONONO HAHAHHAHAHAHAAHAHAHAH get your money up, broke boi
>he cant fit quant 4 h3 and quant 4 encoder in his vram
pathetic....
>>
File: 1760020819212717.png (72 KB, 947x203)
72 KB PNG
cloud is about to shit on local
>>
>>109572608
literally whats the point in owning a car ? The $20 bus pass is far enough for going to work and stuff
>>
>>109572608
What's the point of cloud models if you're not a pedophile? There's no reason to support your friendly yiddish tribesmen and their Epstein island vacations if you can run models locally.
>>
>>109571693
>>109571784
Hax is suitable for forking. It's a tiny C harness.
>>
>>109572608
making poors jealous
if you don’t want to be jealous stay out of /lmg/
>>
>>109572612
Not a music or videofag, but how much vram is it for reference?
>>
>>109572618
Thats not even a correct analogy because Sol 5.6 is far better than any local model
>>
>>109572617
this account still alive huh, nuts
>>
File: 1782407570827351.jpg (193 KB, 974x1080)
193 KB JPG
https://github.com/LostRuins/koboldcpp/releases/tag/v1.119
>>
>>109572608
>Literally whats the point of local models if you are not a pedophile? The $20 plab of chatgpt is far enough for coding and misc stuff
this is the lmg version of a foid falling back on calling their interlocutor an incel
>>
>>109572636
Can you show me logs of Sol analyzing Richard Krege's ground-penetrating radar findings from the Treblinka camp without a refusal?
>>
>>109572635
>Not a music or videofag, but how much vram is it for reference?
22GB
I'm trying to run it alongside my llm, imggen and tts setup which is why I'm putting up with lowmem setups for music
I'm glad it makes avartarfag happy that I might be a ramlet
>>
File: 1767520299389597.png (569 KB, 1178x1122)
569 KB PNG
https://xcancel.com/fanfei__li/status/2088279761084620802
UOH underage AI...
>>
>>109572636
its not even a correct analogy because you can fit 20 people in a bus but only 5 in a car
>>
>>109572650
I love you kobolddev. You're my favorite snailcat and are worth the wait.
>>
>>109572650
Is Minimax vramlets friendly?
>>
>>109572667
Kek
>>
>>109572666
>Only 22GB
Oh I was expecting a way bigger number from the way you phrased it. I'm happy video gen is still relatively lightweight.
t. Blackwellfag
>>
>>109572635
entire music thing is 30,something gb, the quant 4s h3 encoder and video gen model put together are about 31.4 gb
>>
>implying aicgjeets are "rich" by any means
These kids are desperate for proxies, they literally sell their logs in exchange for access to some mystery meat quantized model, 200 eurobux is considered a lot in their shitholes. Fucking locusts hahahahah.
>>
>>109572679
>video gen
That was for musicgen. You can do it in 8gb with a lowmem setup.
Similar for videogen tho. I think you need 32GB to do it fullspeed but you can force it to work at 24 or even 16gb if you don't mind weight swapping and maybe quants
>>
>>109572608
I fuck schoolgirls on the API models too, that's irrelevant. This is the local models thread.
>>
>>109572667
oh I'm gonna fuck the hell out of this little one...
>>
>>109572667
Not true. I was doing high school math when I was 6 years old.
>>
>>109572650
>Makes M3's KV quantizable without crashing
Unfathomably based since I was just on the cusp of fitting another hot layer into VRAM.
Thanks for the speedup.
>>
Reminder to get your 5090 now before they double in price again



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.