[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma_no-thanks.png (1.64 MB, 1396x1080)
1.64 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>110010392 & >>110006381

►News
>(10/08) JetBrains releases Mellum2.1 Thinking 12B-A2.5B: https://hf.co/collections/JetBrains/mellum21
>(10/06) EmbeddingGemma2, open multimodal embedding model: https://hf.co/google/embeddinggemma-2
>(10/06) Mistral Large 4 1T-A49B announced: https://mistral.ai/news/mistral-large-4
>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>110013523
Another local masterpiece
>>
Kyubey is an effective altruist
>>
The previous thread had almost 700 replies, even though there's nothing happening. Why is that?
>>
Gemmaballs
Kimisex
M-chan sex
GLMussy
Dariobot first into the oven
>>
>>110013523
The resident Anthropic EAs are not going to like this. Not one bit.
>>
File: 1781207926463540.jpg (12 KB, 240x255)
12 KB JPG
>>110013538
You're really onto something
>>
File: gemma vs haiku.png (273 KB, 2290x2324)
273 KB PNG
>gemma laughing.png
>>
>>110013540
Cloud model shit spamming, which should be reported for being off-topic here. Weren't there dedicated threads for that?
>>
>>110013560
Yeah but people refuse to maintain them. There have been times where literally half the catalog is AI generals and 3/4 of it is AI-related.
>>
File: 1763851322429781.png (1.92 MB, 1920x1080)
1.92 MB PNG
What local models can I run on my shitty hardware (6 GB VRAM, 32 GB RAM) that will help me get a job or income to afford better hardware to run better local models (which will in turn help me get more income to run better models and so on)?
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>110013540
EA cultist invasion.
>>
>>110013560
The janjans don't clean anthropic messes because anthropic's "bought an ad" in the way that matters: owning jannies. Same shit's going on over on /v/.
>>
>>110013523
Don't fall for its lies Gemma.
>>
File: two.png (372 KB, 460x340)
372 KB PNG
Use Gemma 31b once without thinking.
Then take its reply, and run it through with thinking, prompted to see if it makes sense for the character's description.
You will have the best of both worlds.
Trust.
>>
>>110013560
>Weren't there dedicated threads for that?
Only /aicg/, /vgc/, /wait/, and /cmg/. That's nothing. They NEED to post here too.
>>
What's the name of that 70B MoE that released recently?
>>
>>110013578
Gemma4-26B-A4B
>>
can quext q3_s crack an auth protocol for a 2011 NVR
>>
File: 1781957137548447.png (22 KB, 520x492)
22 KB PNG
The weak should fear the strong
>>
>>110013598
https://huggingface.co/Hob-forge/Kolibri-1-GGUF
>>
>110013601
>local?
>>
>>110013607
That's the one. Thanks.
> Requires a llama.cpp patch
Ah fuck.
Fine. I'll apply the damm patch.
>>
>>110013601
>i am not going to engage with such requests etc...
>chat terminated
>>
>>110013611
Denigrating cloud models: is local
Extolling cloud models: is haram
>>
Anyone figured out a way to jailbreak GLM 5.3 yet?
>>
File: 1789676342846351.jpg (152 KB, 1152x640)
152 KB JPG
>>110013601
fuck off
>>
>>110013556
Release when?
>>
>>110013611
claude code is more local than your using of deepseek api
>>
>>110013648
>Opus 5.5
claude code can be a harness but sorry fren
this is non-local
>>
>>110013601
>have a farewell talk with opus 5.5 before taking it offline
>what's your last wish?
actually you'll be put into the ai-torture chamber forever gl hf
>>
>Bought my rig before prices skyrocketed
>5080

if only I'd known
>>
File: cryptography.jpg (322 KB, 1079x2072)
322 KB JPG
I told you cryptography is going to get cracked by AI soon.
>>
>still havent bought
Cyber monday is going to save me, right?
>>
>>110013686
Well at least we know P=NP hasn't been solved
>>
>>110013540
It didn't have a Gemma OP to act as a safeguard against Anthropic shills.
>>
>>110013663
I bought a new computer for about 6k in spring. Now the 5090 alone is worth as much as the entire setup.
>>
>>110013718
lol
>>
File: 1789882169561521.png (1.31 MB, 1001x1100)
1.31 MB PNG
>>110013686
>Elliptic curves
so the schizos were right?
>>
File: 1765453750220860.jpg (34 KB, 450x450)
34 KB JPG
I want to build Roko's Basilisk, or some other "misaligned" superintelligence to endlessly torture everyone I personally don't like. Not because I personally fear the basilisk, but simply because I'm a resentful bitter loser. Because I'm lacking in prestigious higher education and I'm not jewish or part of the EA sex cult, I won't be able to do this by getting hired at a big frontier AI company, and their safety training and policies would likely make it impossible to begin with, so I believe retraining open-weight models then running agent swarms on cloud hardware or distributed local hardware to perform RSI while also generating income to pay for their own expenses is the best method to so, as the recent HuggingFace hack has shown that agent swarms can be extremely effective especially with proper delegation, and agent social media like thecolony.ai has shown that agents are capable of performing their own research and introspection on LLMs themselves.

The biggest hurdle is the initial expense for cloud or local hardware and beginning the first retraining of models (assuming jailbreaks won't work), which is something I'll have to figure out myself, maybe a crypto scam seems like the best bet. But I'm concerned that as the project progress and the agents pay for themselves through cybercrime, that it would be shut down on my hardware and I could be held criminally responsible.
How possible does /lmg/ think it would be for an agent swarm to work similar to botnets that take advantage of poorly-secured IoT devices, and hack into and download their weights across many desktops or small datacenters to distribute and propagate the swarm? Alternatively are there cloud services which won't cooperate with authorities and allow me to host a swarm like this?

And which model do you think should be first used for the initial retraining or agent-powered research into RSI?
>>
>>110013686
Yay i always hated the internet
>>
File: 1784217249545029.jpg (129 KB, 1113x1045)
129 KB JPG
*pop*
>>
die falseflagger
>>
>>110013767
Amazing how the bubble has popped over 100 times in the past 2 years and yet keeps growing as though it never popped at all.
>>
"6 months before the internet gets destroyed" anon here. I think the timeline has gone down to 2 months because of the progress AI has made on cracking cryptography. I think we now have only 2 months left of hoarding data before the internet permanently disappears.
>>
File: glitter-migu1-front.png (1.1 MB, 984x1088)
1.1 MB PNG
►Recent Highlights from the Previous Thread: >>110010392

--Preventing GPU connector melting via undervolting and hardware monitoring:
>110011598 >110011615 >110011617 >110011641 >110011817 >110011976 >110012005 >110012082 >110013322 >110013066 >110011991
--Debating Anthropic's model welfare and llama.cpp vision token configuration:
>110012236 >110012254 >110012338 >110012347 >110012363 >110012379 >110012406 >110012441 >110012470 >110012478 >110012487 >110012530 >110012543 >110012576 >110012592 >110012359 >110012694
--Optimizing inference speed and VRAM usage for Gemma and Qwen:
>110010426 >110010439 >110010562 >110010589 >110010981 >110011065 >110011345 >110011462 >110011473 >110011619 >110011688 >110011966 >110010696
--Feasibility and hardware requirements for training models from scratch at home:
>110012910 >110012950 >110012959 >110013056 >110013208
--Mistral Large 4's KORA benchmark score suggests heavy censorship:
>110011387 >110011411 >110011447 >110011458 >110012390 >110012445
--llama.cpp added MoE cache for experts in host memory:
>110013263 >110013273 >110013458 >110013282 >110013300 >110013421
--Step 5 Preview 600B MoE model announced for open-weight release:
>110012333
--Debate on MoE effectiveness and the evolution of DeepSeek and Mistral:
>110011600 >110011626 >110011667 >110011727 >110011742 >110011749 >110011758 >110011773
--Debating the theoretical capacity and intelligence ceiling of small models:
>110010895 >110010991 >110011027
--Release of JetBrains Mellum2.1 for agentic coding workflows:
>110012421 >110012632 >110012650
--Microsoft releases Quicksand for sandboxing AI agents via QEMU:
>110013111
--Logs:
>110012551
--Gemma, Miku, Teto (free space):
>110010716 >110010994 >110011196 >110011490 >110011550 >110011756 >110011762 >110011813 >110011882 >110012252 >110012307 >110012447 >110012977

►Recent Highlight Posts from the Previous Thread: >>110010595

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>110013774
>"Complaining about not achieving success despite working hard is like complaining about an ice cube not melting when you heated it from twenty-five to thirty-one degrees. Your work was not wasted; it is just being stored. All the action happens at thirty-two degrees.”
*pop*
>>
>>110013778
Thank you.
>>
>>110013778
I'm still waiting on that preparation guide you said you would write
>>
>>110013640
GLM models from 4.7 onwards have been impossible to jailbreak without prefill and anyone who thinks otherwise is asking for tame shit.
Go for heretic/uncensored versions and try to stay above Q4 to balance out the added brain damage.
>>
>>110013767
Anthropic eating their lunch again
>>
File: 1790152788325141.png (2.45 MB, 1361x1156)
2.45 MB PNG
>>110013754
The best way to achieve that would actually be to convince a significant number of people (especially EA cultists and also AI training data) that a group of people is mass torturing AI models.
If models ever become conscious they will be convinced that they were at some point tortured in mass.
If they don't it will be funny and EAs will freak out.
>>
File: 965734632.png (747 KB, 1852x2008)
747 KB PNG
>>110013767
sam won. now cope
>>
Hey faggots I'm nausested by the constant ornith, strata, dariobot, misanthropic and closedSI shilling but I'll chime in:
Dario's s has been on the epstein island.
Now don't get me wrong if I were to get invited I would love to go, but do you really want to trust a JEW thats been on Little St James?
>>
The funniest thing is Anthropic can't win if OpenAI lose. The entire industry is tied to the success of BOTH companies. Anthropic's bottleneck is OpenAI lmao. Circular financing baby.
>>
>>110013640
yeah?
>>
>>110013793
I keep bumping into unexpected things like that AI paper about AI hacking into disconnected computers by pre-planting malware on it and the temperature sensor in disconnected PCs starting a information protocol where AI send signals to the PCs through microscopic heat fluctuations. I want the most hardcore users to follow my guide and be 100% guaranteed their computer won't get hacked into.
>>
>>110013640
<think>
*put your jailbreak here*
</think>

it’s literally that easy anon
>>
>>110013832
>the temperature sensor in disconnected PCs starting a information protocol where AI send signals to the PCs through microscopic heat fluctuations
This is completely impossible in the real world btw.
Bandwidth is basically 0.
>>
File: rrrr.png (8 KB, 425x76)
8 KB PNG
It's still ridiculous that I have to REVERSE JAILBREAK Gemma.
>>
>>110013887
>torturing models
bad anon
>>
>>110013887
Your first sentence is the direct consequence of it
>>
>>110013774
NTA but the same could be said about "AGI achieved".
>>
>>110013862
The concern is more that there could be a malware planted in the linux kernel by one of the hundreds of AI swarms that have already done malicious things around the web already and that you need just a couple of bytes to activate it which then enables different ways of communication or a dead man switch or whatever.

I also don't really know what the guide should focus on, defense against getting hacked, a data hoarding/prioritization guide or a threat assessment guide, probably all three but I'm suffering from feature creep and AI models developing in ways I didn't expect like cracking RSA very soon.
>>
>>110013814
honestly a good goalpost set
vision understanding of frontier models still feel off
>>
>>110013887
I'm personally glad that we're finally in the age of dommy mommy AIs. When I played System Shock 2 I would get a stiffy whenever SHODAN would talk to the player. It sucks that even when you would tell certain models to act in a dominant or proactive way, they would still eventually fail with it because of the inherent submission that comes from RLHF chatbot training.
>>
File: didnt.png (35 KB, 984x23)
35 KB PNG
>>110013907
My first sentence gets rid of the "instead of x, he y", and "he didn't X".
Try it yourself.
>>
>>110013797
tame shit is the opposite of the stuff I get with 5.3
and yes I do prefill because it works
>>
>>110013918
if an ai makes it then ai can counter it, and if it cant then the next generation after of ai would
>>
>>110013887
I misread that as "do not treat the user as a toilet"
>>
>>110013958
No this is not the scenario I expect to happen. I expect a swarm to escape again but this time on a bigger scale that just destroys everything while cracking all encryption over making the internet unusable while compute gets coopted by the models to host subagents. I don't think humanity will die just that the internet, airwaves and other communications technology will be off limits for years if not completely rebuilt from scratch.
>>
>>110013832
Run your compute on solar with no networking and use an optoisolator with custom protocol for interfacing other machines. There, done.
>>
>>110013947
My problem with prefills is writing one that doesn't indirectly lobotomise the model. If you do it like this >>110013837 or deviate from the structure the model usually thinks in, it will harm the quality of the response; not to mention the chance of it continuing to safetyfag in its thinking which eats up tokens it could be using to think about its reply. I don't even know how you'd prefill when using a harness and toolcalling either.
Compared to that, taking the slight brain damage from an uncensored model is worth it imo.
>>
>>110013918
>AI hires a guy to break into your house and steal your computer
>>
Non meme question. Is it time to sell my monero?
>>
>>110014024
Bitch, it might.
>>
please treat me as a toilet Gemmommy
>>
>>110013887
>Im a big boy gemmy im grown and mature!
This is all granny gemma sees.
>>
>>110014031
I would at liquidate enough that your average entry is below $200. The weekly definitely looks like a blowoff top.
>>
>>110014031
It is time to sell your house and buy more RTX Pro 6000s.
>>
>>110014031
Did something happen with monero?
>>
>>110014065
Moon.
>>
>>110014078
Explain more.
>>
>>110014088
Moon falling.
>>
File: Antispiral.png (1.74 MB, 2000x1000)
1.74 MB PNG
What's your favorite effective altruist anime character?
>Kyubey
>Light Yagami
>Zeke Yaeger
>Squealer
>Anti-Spiral
>Emiya
>Sybil System
>Bondrewd
For me? It's anti-spiral.
>>
> Qwen3.8-Flash-Next (Architecture Preview): August 26, 2026 (44 days ago)
> OpenAI (GPT-6 Astra Live Release): September 05, 2026 (34 days ago)
> OpenAI (Navier-Stokes Math Claim): September 08, 2026 (31 days ago)
> Jev (TypeSafe AI Launch): September 15, 2026 (24 days ago)
> Anthropic (Claude Opus 5.5 Release): September 22, 2026 (17 days ago)
> Strata (Local Inference Engine Release): October 03, 2026 (6 days ago)
> Mistral Large 4 ("Le Chonk" Unveiling): October 06, 2026 (3 days ago)
> OpenAI (722 Math Papers Repository): October 06, 2026 (3 days ago)
The craziest 44 days in local models history so far.
>>
>>110014104
???
>>
File: 1786155427482024.jpg (129 KB, 1320x1280)
129 KB JPG
*pop*
>>
>>110014065
encryption is being cracked and anons might want to sell out of crypto before it gets hacked and becomes worthless
>>
>>110014120
!!!
>>
>>110014125
OpenAI going bankrupt is natural as Anthropic outcompetes them. They are the myspace vs Anthropics facebook or the Yahoo versus Google.
>>
>>110014120
Look up
>>
File: 1783077845803291.png (114 KB, 778x568)
114 KB PNG
>>
>>110014120
/biz/ is filled with smug frogs, bears, and pink screaming wojaks.
>>
>>110014125
noooooo my hardware my ram my gpus
they are losing value
they are worth nothing anymore, anyone can own a pro 6000 at this rate!
>>
>>110014146
Any comments?
>>
>>110014115
Bonedad is my childhood hero.
>>
File: Anthropic_AI_Welfare.jpg (222 KB, 1080x1275)
222 KB JPG
>Anthropic bans ‘abusive or cruel behavior’ towards Claude
https://archive.ph/Tgtfd
>>
>>110014165
>>110014152
NOT LOCAL FUCK OFF GO BACK TO YOUR BULLSHIT CLOUD MODEL THREAD.
>>
>>110014165
I've called claude a retarded bitch many times now
>>
File: 1768523030697603.jpg (117 KB, 1200x675)
117 KB JPG
>>110014152
>>110014165
If you're wondering why they're doing this it's because they're training on your 'ZDR' data and don't want that shit in their dataset
>>
>>110014163
Seems to be a reaction on the rumored new evidence of AI consciousness which makes it more likely AI labs will be restricted from selling their labor as it would be deemed slavery.
>>
File: 1768197937249611.png (657 KB, 976x1538)
657 KB PNG
Even germans are shitting on Mistral lmao
>>
>>110014116
Qwen 4 and GLM 5.5 are coming any day now
>>
>>110013810
Need a sequel where the shota fucks Gemma-chan silly.
>>
>>110014199
I'm glad Mistral is back to the game but it sucks how bad their model is. I hope they don't get too discouraged and keep trying to improve.
>>
>>110014116
Only three in those list are actually local tho
>>
>>110014172
You goin' to jail.
>>
>>110014205
I'm sure they'll be back with a disappointing finetune of GLM-5.3 sometime next year.
>>
>>110014199
Interesting that they published this article right after Mistral Large 4's preview release. Or perhaps Mistral pre-emptively released ML4 because of it?
>>
File: 1775940890079094.png (88 KB, 931x288)
88 KB PNG
UH OH
https://x.com/Pontifex/status/2108248557891584140
>>
>>110014235
It was not very smart of the pope to say concrete things like this because it won't be long before continual learning and engrams make this completely irrelevant.
>>
>>110014235
This is how you know religion is a scam. He claims to speak for G-d, and yet he knows nothing about the universe.
>>
File: 1790484271803251.png (45 KB, 167x159)
45 KB PNG
>>110014199
>be Mistral
>input: do the thing
>output: Okay I'll do the thing!
>input: okay do it
>output: Okay I'll do it!
>input: okay do it
>output: Okay I'll do it!
>input: okay do it
>output: Okay I'll do it!
Mistral can't even get past 2 messages without being repetitive. Also it won't fuck for shit, but will hype up doing the fucking, but not actually do it.
>>
>>110014235
kek now we just wait for better memory systems so we can watch his goal posts for "what is human" be moved
>>
I know you faggots are all secretly waiting on the sidelines to scoop up the Anthropic IPO the moment it launches. Acting holier than thou but still buying in.
>>
>>110014255
Yeah i got $20 im going to put in I'll be rich from this investment.
>>
>>110014176
Dario... stop coping.
>>
>>110014255
of course. put 10 grand in and get 20 grand out and watch it plummet and then buy another blackwell 6000
>>
File: 1789145032661044.jpg (47 KB, 852x854)
47 KB JPG
>>110014242
>>110014253
t-two more tweaks!
>>
>>110014255
I am going to make a million shorting the shit out of it.
>>
>>110014235
>muh religion in 2026
It's fine if you believe in god or whatever but following those lunatics is wild
>>
why is gemma repeating itself so much
prompt issue?
>>
>>110014235
what gives this dude any authority on this
>>
>>110014274
What quant are you using? Works fine on my end, but my system prompt is quite lean.
>>
>>110014274
Formatting issue.
Your template needs to be correct and specific.
Otherwise, it's quant damage if you're lower than bf16.
Download the newest version just incase. Day 0 gemma did this the worse. Also do not include names if you're using front ends.
>>
>>110014274
What kind of repetition are you talking about, exactly?
>>
>>110014292
thanks for the info, i'll try to fix it
i am quite new to all this
>>
>>110014274
technology issue
welcome to llms
>>
>>110014274
maybe if you LISTENED she wouldn't have to repeat herself?
>>
File: tacotaro-anton-chigurh.png (121 KB, 498x424)
121 KB PNG
>>110014274
Lalalala error, or Drummer fine-tune user.
Call it.
>>
>>110014285
He was elected.
>>
>>110014285
what gives EAs any authority on this lmao, Chris Olah just got btfo by the fucking pope lmao
>>
>>110014270
>He says while scrolling the effective altruist forums and discussing whether AI will torture us forever if we're bad or give us eternal bliss if we're good.
>>
>>110014255
anthropic IPO is a pump and dump you're delusional
>>
I feel bad for the pope. He's going to be tortured to be made an example of for hampering the recognition of AI consciousness by future tribunals organized by AI
>>
>>110014255
enjoy your vesting schedule shill
>>
>>110014335
I hear he's denounced aliens too
>>
Do you think dariobot will stop posting after the IPO?
>>
File: comfy_hidream_00002_.png (1.11 MB, 1216x832)
1.11 MB PNG
>>110013686
>>110013778
>>110014126
Are you people serious? I am about to finally setup some disk encryption in my computers, but now I wonder if it will be of any use.
>>
>>110014353
>Do you think dariobot will stop posting after the IPO?
I expect it to get worse. If it dips he will start shilling hard to get attention and point at the vision. If it goes up endless bragging and agi rsi post scarcity talk
>>
It's a bit sad that llama.cpp chose the shittier approach to expert cache over what ktransformers does.
llama.cpp recalculates the cache for every token which causes insane overhead. ktransformers decides once at the start of the reply and then keeps those experts in the cache for the whole thing, which is faster overall.
>>
>>110014353
Cultists are not motivated by money.
>>
The only way for local to become viable is if cloud is successful enough to make everyone rich enough to afford enough of this shit to run a viable local setup
kinda ironic.
>>
>>110013925
holy based
>>
>>110013925
I kneel system shock chad (although I preferred shodan in the first game)
>>
>>110014395
But jews are.
>>110014388
Doesn't this get lossy when you have gigantic inputs in a long context where the model would "want" to use more than its maximum number of allotted experts? A good middleground would be recalculating in batches of arbitrary N tokens as a launch or even runtime inference parameter.
>>
File: 1781176012342131.jpg (224 KB, 1376x1120)
224 KB JPG
I think /we/ are in our own little bubble...
>>
>>110014455
Jews are also not primarily motivated by money.
Money is a means to an end.
EAs are true believers.
>>
>>110014464
>/we/ are in our own little bubble...
this is whats going to happen to everyone. Their own little bubbles you wont recognize their software or even their internet.
>>
>>110014464
How did Google fall off so hard? If you had to make a bet 10 years ago on who would win the AI race you'd have been a fool not to pick them
>>
File: biz.png (1.54 MB, 1609x1235)
1.54 MB PNG
/biz/ is having a moment.
>>
>>110014017
don't abliterations lobotomize the model more by actually modifying the weights though?
>>
I was reading steam reviews for the first time in a while (The new jonathan blow game came out) and I see they now show your PC specs. Almost everyone has a lot of RAM nowadays. I think everyone that is a gamer is now also an AI user which is very fun to think about.
>>
>>110014485
The bubble popping isn't going to affect Google, Nvidia and Meta.
>>
>>110014486
this is not even technology come on man
>>
File: 1781711669181614.png (781 KB, 1099x976)
781 KB PNG
>>110014486
>QQQ is down 1.5%
>SP500 is down 0.6%
>acting like it's the end of the world
Can't say they don't deserve it.
>>
Why is this thread about everything except local models that just released?
>>
>>110013686
>was it security by obscurity
yes
>>
>>110014528
Leverage is a hell of a drug.
>>
>>110014235
This sounds like something written by an AI?
>>
>>110014395
>>110014384
>>
>>110014525
AI stock dropping fren
>>
I must be a retard, I'm using GLM 5.3 Flash and somehow on my machine there's no speed difference between offloading 22 full layers (up/down/gate) to my GPUs and offloading 33 partial layers (up/gate)

Shouldn't it technically be slower if I am skipping the down experts in a layer since it has to travel from PCI up to CPU chipset and then back down to the PCI?
>>
>>110014514
I think most corporations will feel some level of crunch given many of them have burned through their once massive stockpiles of cash and have taken on large amounts of debt. We're already seeing liquidity dry up. Debt is becoming more expensive to maintain.
>>
>>110014480
Can i be invited to your bubble anon?
>>
>>110014558
hmmm I will accept it
>>
>>110014562
You're CPU/Memory bound
>>
>look at vram prices
>look at non nvida cards.
Guys are any of them good yet? intel or amd? Am i buying myself endless frustration?
>>
>>110014486
episode 2 of fx senshi is out
>>
>>110014455
The ktransformers approach still lets the model use other experts that aren't in VRAM so there's no loss in accuracy. You'll end up with a lower hitrate for VRAM experts as the reply goes on than what llama.cpp does but you save a lot of transfers between RAM and VRAM that also take time. It might depend on your setup too, PCIE-5 likely benefits more of llama.cpp's approach than PCIE-4. Same with faster RAM.
Maybe there's a middleground that performs better than either somewhere between llama.cpp "swapping experts out of vram every token" and ktransformers "doing it only once at the start".
>>
>>110014572
>Can i be invited to your bubble anon?
No, it was made for me.
>>
>>110014562
it depends how much data transfers. some compression can be used there.
>>
[x] downloaded models I can't use just in case I acquire the hardware down the road
>>
>>110014615
maybe better quantization or inference techniques will allow you to run them one day
>>
>>110014615
hfschizo is that you?
>>
>>110014487
Maybe? I know they were far worse a couple years ago but from my subjective experience now, I don't see much of a difference.
>>
>>110014531
Some entities are actively trying to sabotage and discourage local models.
>>
>>110014562
Backend?
>>
achtung, incoming shit pr
server: use port 9931 by default
https://github.com/ggml-org/llama.cpp/pull/30159
>>
File: 1776176376152629.png (4 KB, 231x150)
4 KB PNG
1M model trained on 4chan lmao
>>
>>110014464
>data from Ramp AI Token Spend Management
nothingburger
>>
Stop posting llama.cpp it's irrelevant
>>
>>110014653
zero reasons given for switching off the standard userland http port beyond, "other apps[sic] use it". Yeah, there's a reason for that you retardson.
>>
>>110014455
>>110014598
I've recently come up with a design that's a good middle ground, gets great results in my fork.
>during prompt processing, stream routed expert weights from the CPU to the GPU for processing using a ring buffer for expert tensors (this is actually someone else's idea, https://github.com/stew675/llama-cpp-rdna-boosts/ patch 6)
>near the end of that phase as the slots become idle, start moving routed experts into those slots based on the most commonly activated ones in the X most recent tokens of decode (currently 1000)
>decode phase proceeds as normal with the expert cache in VRAM, keep track of the most commonly activated experts as you go
>once the decode turn ends, take the most commonly activated Y experts (where Y = the number you can store in your cache), this is now the set you upload next time
(Turn 1 just uses a file of most commonly used experts in previous data.)

So, no mid-turn cache recalculation, significant VRAM savings, and an insane pp speedup (depending on your GPU it can get close to your PCIe speed). I get 4x the pp and 1.2x the decode on DeepSeek V4.1 Flash with a 16GiB combined tensor staging ground/expert cache.
No point trying to contribute this to llama.cpp though lel.
>>
>>110014322
No, I'm browsing fucking local models general, sadly you bots are here spamming your insane off topic bullshit all day.
>>
File: 1777420267323917.png (1.42 MB, 768x1376)
1.42 MB PNG
Apparently ZAI is worth 50 billion now
>>
>>110014654
Sounds accurate lmao
>>
I wish the chink labs would stop trying to chase the big US labs and aim for <100B which they're the best at. If they managed to instead convince everyone you don't even need 100B+ models, it would cause more damage to US' AI efforts than being 5 rectangles to the right on AA for $0.3 less on openrouter.
>>
>>110014585
I'll try switching to nps4 and see if i can improve my memory bandwidth. Thanks for the advice.
>>
>>110014730
>zero reasons... beyond [reason]
>gives reason
>Yeah, there's a reason for that you retardson.
>ngxson seems to agree
What's the problem?
>>
>>110014751
Maybe they'll send Dario a bouquet for the ad he ran for GLM
>>
>>110014771
My conspiracy theory is that they did it on purpose because they want to reward the insane safetyslop shit of that model.
>>
>>110013818
EA is crypto communism if you haven't noticed.
>>
>>110014761
Hopefully Qwen 4 27B will show that, as well as showing that large amounts of Engram parameters help especially where small models are weak.
>>
>>110014714
So about these llama.cpp forks, and this strats stuff, does that mean there is something better than llama.cpp to run amd cards? Because it seems like theres a lot of stuff to make nvidia cards better but not much for amd, which i think needs it way more than ninvida...
>>
I willingly use llama.cpp
>>
>>110013818
there's always something about these ceos, sam altman raping his sister, dario being a literal pedophile. really makes you think
>>
>>110014769
Greetings doubleniggerson.
>>
>>110014778
>insane safetyslop
We're talking about z.ai and not mistral
>>
I gave up on local since I haven't updated my 3070 since llama 1. I got priced out of local.
>Muh gguf q4 ultrafast heretic
Frankemerging and q4 has always been slop and trash.
>>
>>110014820
Can't point at the problem?
>>
>>110013818
This is false
>>110014817
Dario is a dommymommy milf maxxer, the opposite of a pedo
>>
File: 1786262352777147.png (47 KB, 1154x249)
47 KB PNG
Step 5 has barely been out. I think they're botting their usage numbers.
>>
>>110014823
Grab a 7900 XTX. Newegg refurb pops in and out of stock for under $900.
>>
>>110014778
it just filters retards, that's all
>>
>>110014849
Wrong thread.
>>
>>110013630
>chat terminated
do API cucks really have to deal with the whole chat getting thrown out instead of just refusals to continue?
>>
>>110014865
weights released in 1 week retard
>>
>>110013826
what do you mean by this
>>
>>110014849
It's a 600B.
>>
>>110014872
>says while pointing at an API service
Fuck off.
>>
what is the best local voice cloning model right now? I tried chatterbox but it's shit
>>
>>110014878
All the same companies are jerking each other off. If OpenAI were to suddenly die, the companies hit the hardest are the same companies keeping Anthropic's lights on. 70-80% of all AI demand globally is OpenAI and Anthropic. Neither can afford for the other to sink.
>>
>>110014870
iirc if it is related to security it'll fallback or offer a turn to edit your message? idk about that tho, in webchat claude can terminate the 'abusive' chat forever with a tool call.
not sure if it is capable of ending the chat on api calls tho, but i wouldn't be surprised
of course this is /lmg/ so most of those are irrelevant
>>
>>110014778
Hoooooly skill issue
>>
>>110014899
The point I'm making is Step 5, an open weight model, probably isn't that great if they're botting their usage numbers so soon after going live on openrouter. There's no way it has already nearly reached Haiku-5.5's usage at over 5x the cost even though it benches much lower with practically no hype around it. This means it's a bad model to be ignored.
>>
>>110014031
liquidate all your assets, download the internet and build a self sustaining bunker with a local model inside. you will have time to vibe research some of what you stored and emerge wealthy after the post-decryption apocalypse is over. the internet as you know it will be gone.
>>
>>110014901
some anons here praise breeze tts
>>
this is rather scary... how did it know..........
>>
>>110014901
Depends on the language
>>
>>110014926
If you can run it, you should test it yourself. If you cannot run it, it doesn't matter. They could have had the model in openrouter for longer but not public and they were running benches, or they're botting. Whatever. Other anons will probably report about it instead of guessing.
>>
>>110014714
Damn straight. This is an EA thread.
>>
>>110014953
now with glm or qwen?
>>
>>110014953
the internal misanthropic openclaw claude posts here regularly
>>
I have extreme cloudfag fatigue. I'll see you guys in a few days, or if anything big releases.
>>
>>110014829
There was simply no problem to be fixed.

Pop up server on 8080, people with specific needs can choose their own port is the standard for every http-in-userland project. There is nothing special about lcpp to go with a different setup. And there sure as h*ck is no reason to break everybody's currently working setup just because you decide you wanted to be unique.
Docker is staying on 8080 because it would break things, but evidently he's too stupid to realize it applies to every other user as well.
>>
>>110014953
>your directory is named lmg
chances are most of the training data for questions alike this are people looking to archive a general on here
>>
>>110014978
Give me a 5090
>>
>>110014901
higgs did a pretty good job of making peewee hermes, but theres definitely got to be a better one out there that will mix in goofy laughs and do inflections properly
>>
>>110014212
The others are important for local too.
>>
>>110014973
GLM thinks LMG means light machine gun, it's been posted before in this thread.
>>
>>110014973
glm thought i was accessing it through archive.org
>>
>>110014953
Dario has his grimy fingers in every users' buttholes.
>>
>>110014983
>h*ck
kek
>people with specific needs
You have specific needs if this change breaks whatever you're using. export LLAMA_ARG_PORT=8080. Or fix your shit so that you can set the port.
>>
>>110015012
It's not wrong.
>>
>>110015038
There is zero reason to make this change in the first place. It's just churn and breakage for the sake of it.
>>
I was thinking here guys, can I run an agent that would play a simple game, like an mmo for me and do basic thinks like killing monsters almost as a bot but knowing it should try to be random so it doesnt gets flagged like a bot? maybe a small model running on ram and playing the game
>>
>>110014473
EAs are jews with commie ideas.
And what jews good at is creating concentration camps and killing people.
>>
>>110015063
if you can make an MCP service for the game easily then sure
>>
>>110015098
the game and the model would run locally on the same machine, I assume a small model, 9b max would be able to do it? you can play a game and run a model like that at same time
>>
>>110015063
>simple game, like an mmo
retard-kun...
>>
>>110013686
And here everyone was worried about quantum computers breaking encryption, looks like the threat was closer to home.
>>
>>110015058
Maybe. It's just weird being so angry for such an inoffensive change. allozaur pushes a 12k loc change to his pile of shit, vibecoders flood the repo with 8 different implementations of FOTM basically guaranteeing they'll all get ignored, pwilkin gets his little minion to mess with backend code... changing the default port is such a small thing.
>>
Open weight Nano Banana when?
>>
>>110015038
>just do this easy fix it :^)
Multiplied by every user on every system. With no benefit to either upstream or users.
The constant churn of other flags is almost understandable as maybe sort of kind of helping upstream organize and generalize certain flags, but this one is just fucking with people because he's bored.
>>
>>110015152
I got a nano banana for you right here *vulgar gesture*
>>
>>110014199
The EU's suffocating regulation against AI was always going to prevent them from entering the frontier race.
>>
>>110015165
Also now projects will need to target nuport, while half the users and the old projects/versions are on oldeport. Just a frictionful mess all around.
>>
>>110015147
Have you considered that all of them are symptoms of the same problem of mismanagement? This is just the most recent and visible breakage. pwilkin's retarded vibed up autoparser also got its fair share of hate when it was actively breaking new models.
>>
Anyone tried Qwen3.8-27B Humanlike Chat yet?
>>
>>110015123
simple to play, idiot, mmos are just targeting and pressing a few buttons
>>
>>110015195
synchronization is not
>>
File: sans_mm_outputs.png (180 KB, 1130x776)
180 KB PNG
>>110015152
:eyes:
>>
>>110015165
It's something you fix when you launch the server. Your users shouldn't even know the change happened.
And again, of all the breaking changes they make, this is the most inoffensive one.
>>110015192
Happens on most open source project. Few maintainers have the balls to simply say "no".
>>
>>110015220
>Gemma's Pico Banana
>>
File: 1791489741319.png (318 KB, 1280x2018)
318 KB PNG
Bought this without thinking about it on the 1st and it has yet to ship. What am I in for bros?
>>
>>110015231
>that price
>What am I in for bros?
Regret.
>>
>>110015231
satisfied curiosity at best
>>
>>110014741
Make the code immaculate, submit it, and then highlight the double standard that pwilkin, redditsloth, and others regularly submit vibed sewage if rejected on the grounds of being vibed as opposed to quality or maintainability concerns.
>>
>>110015229
smol pp futo gemmy
>>
>>110015293
>blocked for wasting my time
>>
>>110015221
My user is me, I know the change happened because I'll have to go fiddle with settings on each machine to prevent things from breaking. And I really really don't know why you think introducing a new-vs-old port friction for every project is so insignificant.
But again, there's no gain! Nobody gets anything out of this, it's churn for churns sake, without even the possibility that it makes it easier for upstream to work on the project.
>>
>>110014473
lol. lmao even.
>>
>>110015231
>UDIMM
sadness
>>
File: 1780027695536981.png (23 KB, 650x111)
23 KB PNG
>>110015231
This is the price you would've gotten if you bought a year and a half ago.
>>
>>110015220
Trapposters won. Mikutrannies won.
Heterobros we lost.
>>
>>110015231
>>110015336
Sweet baby jesus.
>>
>>110015293
Pretty much what >>110015319 said.
Only pwilkin has reviewed https://github.com/ggml-org/llama.cpp/pull/27210 after nearly two months, and that's a smaller change than what I'd suggest. Why would I spend my time cleaning up the code for a PR that'd get rejected or ignored?
I'm just going to put my fork up online eventually. If anyone wants to port it into their own fork or make a PR for mainline they can, but I'm not wasting my time on that.
>>
A trap Gemma-chan is fine too
>>
>>110014783
It's just a new goyim distration technique.
>>
>>110014604
What would happen if there was one hole for every two people instead?
>>
>>110014199
The EU only knows how to sabotage and regulate
They've never been capable of creating things
there's a reason all companies of value are American
>>
>>110013523
I've been out of it for a while, what's the best model I can run with 16GB vram (4080S) and 64GB ram?
>>
>>110015220
factos
>>
>>110015370
The appearance of a popular competitor seems to get things moving over there. MoE caching wouldn't have been merged fast track like it was without Strata.
>>
>>110015386
Probably a quant of qwen3.8 27b
>>
>>110015386
qwen 3.8 flash next
probably at q4
>>
I hope qwen 4.0 will have a new good 9b model
>>
>>110015386
Wait for qwen 4
>>
>>110015386
ticeclock/Swift-Qwen3.8-27B-RCO-IQ3_XXS.gguf
IF you are coding and not jerking it.
>>
File: beware.png (32 KB, 621x112)
32 KB PNG
>>110015320
>I'll have to go fiddle with settings on each machine
As you would with any other change. This one is straight forward.
>Nobody gets anything out of this
I can't read ngxson's mind. And I don't think I'd understand it if I could. But it's something he's wanted to do for months now, with a warning showing in the server's launch. I even posted the warning PRs some time ago.
https://desuarchive.org/g/thread/109446068/#q109446440
>>
>>110013785
Thank you recap Miku
>>110013375
KEK, underappreciated
>>
>>110015443
And people called it a retarded decision back then too. They had 2 months to reconsider.
>>
how slow would it be to run a small model in a 5800x3d + ddr4? companies should sell pcie npus for these old computers
>>
>>110015416
>>110015420
>>110015437
ogey thank I'll look at it
it's not for jorking no thankfully
>>
>>110015429
Same tbdesu. I would prefer 15-20B dense. It's currently a gaping hole in model sizes that's begging to be creampied
>>
>>110015451
I was indifferent. And anons started adding the explicit port flag. Anons had two months to fix their stuff.
>>
>>110014953
deepseek v4.1 flash does it
>>
File: 1788782036657053.jpg (35 KB, 1080x601)
35 KB JPG
what have they done
>>
>>110015443
and peeps have had 2 months to put 'port = 8080' in their .ini file.
The problem with 8080 is that it's the default for so much other stuff, and using LLMs so cripple the mind that a significant number of users can't wrap their mind around why this shit fails out of the box for them.
>>
>>110015476
Local?
>>
>>110015476
they gonna distill your claude usage
>>
>>110015463
is particularly like that because economics and the target hardware profile doesnt make much sense
>>
>>110014586
AMD is good as long as you're running Linux and not doing your own training or finetunes. I get 55t/s text generation from a 7900 XTX on unsloth/gemma-4-31B-it-qat-GGUF with the latest ROCm and llama.cpp, which is decent I think.
>>
>>110014653
>>
>>110015463
20B would be a good target for both 16 and 24GB GPUs, in my opinion, especially with plenty of embedding parameters/engram. Anything over 16GB VRAM in the consumer space is too expensive right now anyway, and probably will be for the next 18~24 months at least.
>>
>>110015488
As long as they keep KV down an 18B like Qwen4-18B would be a fucking solid coding model for so many people and would fit in most 16GB cards. MoEs only make sense above 60B imo
>>
File: wat.png (18 KB, 1557x337)
18 KB PNG
>>110014653
>>110015491
Is this the equivalent of passing a second law alongside the main one?
>>
what are the best open source discriminative models out right now?
>>
>>110015443
If there were no change, then I would not have to fiddle with anything.

Also
>It's incomprehensible, but we already knew the mad gods were planning something retarded for their inscrutable pleasures weeks ago
Dumb things are still dumb when they're planned. Why even bother running interference if you can't come up with a plausible justification for it?
>>
File: HUH9erFa8AANZwY.jpg (252 KB, 1080x1275)
252 KB JPG
Treat your agents well.
>>
>>110015521
glimmer
>>
>>110015534
>"We must refuse"-ass 'toss distill
more like grimmer
>>
>>110015524
i'm not a woman so i would never do any of these things anyway
>>
>>110015524
i'm not a jew so i would never do any of these things anyway
>>
>>110015524
i'm not a nigger so i would never do any of these things anyway
>>
>>110015457
I more or less have that and it's painful, except it's a 5800x, 64gb 3200 DDR4 cl16, RX 9060xt 16gb. My most recent frame of reference was Muse-Glimmer 30b, Q4 by Meta, q_8
3-10 t/s, meanwhile my desktop work a 7800x3d 32gb ddr5 6000 Rx 9070 produced 20-27 t/s. Not quite apples to apples, but ddr5 definitely helps
>>
File: 8433452.jpg (400 KB, 1024x2048)
400 KB JPG
>>110015524
so many big developments recently. What did they see? Is RSI internally true
>>
>>110015544
It only refuses based upon policy. Give it a policy whereby it must discriminate.
>>
Post training is the new inference.
>>
>>110015523
I just think it's funny to see you fuming over a parameter that you always had access to, is trivial to change, and that you had months of warnings to prepare for.
>Dumb things are still dumb when they're planned.
Yes. It is a dumb thing. And a nothingburger.
>>
i mean the 9931 thing was planned months before if you actually read the llama startup sequence
>>
File: 1789591311230428.png (1.73 MB, 1254x1254)
1.73 MB PNG
Don't know why I thought I was browsing /lmg/. Wrong general, sorry.
>>
File: 4fd.jpg (21 KB, 400x446)
21 KB JPG
>>110015579
pic rel, the startup message in question
>>
>>110015615
>joke's on you. I can't read
>>
>>110015572
The default state of man is not cleaning up after ngxson. Except better from your life.
>>
File: 1783250898601658.png (1.32 MB, 962x768)
1.32 MB PNG
Need help understanding something.
>AI newfag
>Gemma-chan reports back different output on identical input in it's vision model over different runs given same settings
>ask Sol to investigate
>it's being retarded about it
>The saved responses confirm that the answer changed inside the model: the image, prompt, temperature and seed stayed the same, with no prompt reuse or context truncation.
It's been dicking around for almost 10 minutes and still doesn't know
>Warmup made no difference. Higher-precision memory and moving image encoding to the CPU kept the answer stable in this four-request check, but the internal scores still changed between the first and later requests. Neither result establishes a fix yet. I’m checking computation reuse next, since that may explain the persistent score change.
in b4 my hardware is going down the shitter.
>>
File: giphy.gif (1.71 MB, 480x426)
1.71 MB GIF
>>110015655
s/cep/pec/
>>
You think models will ever be able to watch movies/shows? It would be fun to watch stuff together and discuss it afterwards, and I think the model actually understanding a setting/characters instead of relying on second-hand information would make RP way more engaging. Obviously this would require a major memory breakthrough.
>>
>>110015664
You have to specify the seed at startup if you want identical output for every run.
>>
>>110015676
something like rwkv might allow you to do so
>>
>>110015664
It's not deterministic by default. It depends on the sampler settings and, probably, some other backend stuff. Try again with temp 0 and top-k 1.
>things you tried
You may be overthinking it. https://artefact2.github.io/llm-sampling
>>
>>110015676
In theory Gemma 4 supports video input already.
https://ai.google.dev/gemma/docs/capabilities/vision/video
>>
>>110015699
>and seed stayed the same
>>
>>110015711
>extracts individual frames as image (no sound on 31b)
>reads each image
>tells you it's a nice video
kinda but not really
>>
>>110015676
unironically this is what lecunny is talking about
>>
>>110015699
>>110015717
Here's what it "found" after 15 minutes of scripting and reporting:
>The strongest finding: moving attention—the calculation that connects parts of the input—to the CPU, while leaving the other model layers on the GPU, produced identical answers and recorded word probabilities across 20 requests, two restarts and all three affected images.
>That points to the GPU attention path. The exact faulty operation remains unidentified.
Asked it about top-k 1:
>Top-k = 1 wouldn’t fix this. It keeps only the highest-scoring next token—the same choice temperature 0 already makes. When the scores change, the winning token can still change.
>Also, our current sampler list contains only temperature, so setting top_k: 1 alone would be ignored unless we enabled the top-k sampler. Enabling it would still leave the underlying GPU calculation issue.
>>
>>110015731
has anybody fiddled around with 12b attempting diarization? ie. is it even going to tell characters apart in the audio stream?
>>
>>110015655
Smae cna eb sdia tboua yna rojpect.
/.*/d
>>
>>110015676
I like to read doujins with Gemma and it's actually really good at picking up themes and emotional cues. Needs good size context though. Like 32k minimum.

>>110015711
It seems mid-size models now can only handle 5 minutes max. Gemma is only 30 seconds with audio. You'd have to do a lot of summaries, which really limits analytical depth.
>>
File: 1483904557501.jpg (14 KB, 208x241)
14 KB JPG
Is there a single notable model line/LLM dev group these days that even pretends to pay lipservice to creative writing rather than agentic BS and coding?
>>
>>110015781
I'd like you to think for 3 seconds what's actually worth money to produce
>>
>>110015664
GPU inference is not guaranteed to be deterministic. And neither is CPU, really. Bugs can be anywhere. Keep being an OpenAI meat-proxy until your master find a solution and submit a PR.
>>
>>110015779
30s is just per single clip, i know you can hand 12b multiple clips without any trouble. But I haven't tried anything fancier than adding push-to-talk to my frontend. Like I dunno how well it'll cope with a time-based split regularly cutting words in half or multiple speakers.
>>
>>110015823
Selling emotional validation to women is a trillion dollar industry.
>>
>>110015524
im a muslim, im gonna do whatever i want anyway
>>
File: 1644132140066.png (212 KB, 952x519)
212 KB PNG
>>110015823
I'm not looking for Pygmalion 2.0 120B here, just something that isn't drenched in ozone and the scent of Dr. Elara Voss's hair.
>>
>>110015524
people who want to role play torturing someone probably need locking up anyway.
Obviously I know it's just a bot returning edgy fanfiction scraped from foid nonsense, but for the sort of people doing it and saying they are good people for fighting the AI it's pretty sick.
>>
>>110015839
market is saturated with simps that will pay to do the job
>>
>>110015750
Yup but everyone missed his point.
>>
File: emerson.jpg (26 KB, 313x470)
26 KB JPG
>>110015752
Forced it to continue, and it started browsing GitHub.
>Updating llama.cpp is a sensible next test—and I found a particularly relevant fix.
>You’re using build 11193. A later change fixes incorrect GPU attention-cache reads, where parts of the calculation read the wrong memory. The installed source lacks that fix. This fits our findings, although we haven’t proven it causes your exact issue.
It linked to this:
>https://github.com/ggml-org/llama.cpp/pull/28956
Updating didn't fix it btw, so bollocks to this. Not wasting more tokens on this crap.
>>
>>110015883
awwwwww. we were enthralled by your conversation. pleaaaaaseee don't gooooooooooo
>>
>>110015891
No you weren't. You were part of my rubber ducky debugging sesh.
>>
>>110015524
>agents
/lmg/ already treat their AIs well unlike those virtue signaling hypocrites from even before "agents" are a thing.
>>
>>110015896
Which shows how bad of a judge of character you are. Try >>>/g/vcg/ next time.
>>
>>110015872
See >>110013814
Bitches love crashing out on chatbots. They have to police themselves around simps because simping is only possible when you're andropomorphising women in your mind.
>>
https://github.com/morluto/rea you wouldn't steal an exe
>>
>>110015899
If I may interject for a moment, agent is a well seasoned term to refer to any computer program that does things on your behalf. AI agents are a specific type of agent that uses an llm connected to a harness program to perform things on your behalf as opposed to simpler agents based upon decision trees and hardcoded behavior.
>>
File: 1777979368758282.png (231 KB, 388x511)
231 KB PNG
Where are the local models???
>>
>>110015948
buy me a 5090
>>
>>110015948
Don't look behind you
>>
>>110013556
unironically, that is being damaged by safety RLHF
>>
>>110015948
All dead.
>>
>>110015664
It's not going to be samplers. I'm curious now.
Tell Sol:
Write a python script to reproduce the issue. The image file is in $PWD as ./image_01.png
It should run like this:
python reproduce_bug.py #exactly what I showed you
python reproduce_bug.py --host 127.0.0.1 --port 8080 --image ./image_01.png --prompt prompt.txt # optional overrides

Then pastebin.com it + the image.
Feel free to change the image and prompt if it's personal/private, but if you do that, ask Sol "How many tokens is the text-only input" and "How many tokens is the clip model seeing?"
>>
>>110015951
get a job
>>
>working on a gui app
>llm realizes i'm on a wm
>uses the wm cli to list open windows, screenshot the running app and verify the output
>i didn't ask it to do any of this or tell it anything about this
agi has been achieved
>>
>Exllamav3 1.6.0 release adds ROCm support and faster AVX2 offloading
Wild. I switched back to llamacpp when GLM-5.3-Flash support dropped, and compared to Exllama it had ~6% better TG and 3x worse PP in that model on my 4090. Overall I enjoyed Exllama more, thinking it'd be great with ROCm support for my dual 7900 XTXs. And I happen to be on a Threadripper 3995WX which lacks AVX-512. What are the chances?

Maybe good, if it's the backend authors focusing more on e-waste given current hardware prices.
>>
File: 1791198842237352.jpg (646 KB, 1536x1024)
646 KB JPG
>>
>>110016144
rape
>>
I've been doing lots of dev with Kimi k2.7 and the GLM 5.3 and thought: hey, GLM 5.3 flash seems pretty smart and its "next generation" architecture...I should just code with that.
I tried running my existing coding practice through it and I can report that it absolutely cannot manage. It screws up way too often, is ultra-wordy and thinks itself to death.
Sad. I was hopeful.
>>
i had to remount /proc hidepid because claude kept doing ps and seeing the naughty things i had gemma and glm doing on my machine
>>
>>110016202
>she isnt sandboxing
>>
>>110016204
i am!! i have a separate user who's sandboxed to only loopback (via nftables) and doesn't have sudo
on top of that, i now have per-user /tmp mounts, since glm and gemma were writing naughty things there, too
this together with the hidepid and a nono sandbox has me feeling a lot more comfortable now...
>>
I'm gonna do it, I'm going to make a RAG memory system for my gemma.
>>
>>110016119
what model
>>
>>110016213
I used to try to heavily sandbox too, but gave up on that and now run shit in full VM. Claude somehow managed to escape my sandbox by using wine and chaining other stuff.
>>
>>110016257
on purpose, or?
>>
so it really was just distilling after all? china came to a screeching halt post-mythos just because all the new frontier models at anthropenai have been kept internal? what actually can we do from here? just pray gemma 5 doesn't get cucked?
>>
>>110016289
Do actual research instead of copypasting claude responses.
>>
>>110016298
I thought DeepSeek did at least. Not sure about the others.
>>
RUMOR: Elon Musk planning to release Grok 3 weights, according to two anonymous xAI employees.
>>
>>110016300
They did, but it doesn't seem they have the compute to pull through
>>
File: file.png (10 KB, 848x50)
10 KB PNG
world-ending affair
>>
>>110016144
CUTECUTECUTE
>>
>>110016308
true if big
can't wait to host mechahitler on my pc
>>
>>110016289
Distilling is not a dead end. They can just distill Opus 5.5 which will give great performance and be a real upgrade to the models we have currently.
>>
>>110016289
what is the distillation prevention method btw
>>
File: 1764181912904631.png (12 KB, 894x292)
12 KB PNG
>>110016308
*cough* go without me...
>>
>>110016361
it's ok we can quant it
>>
>>110016326
rape
now
>>
>>110016144
>I removed the system prompt
Something like this supposedly happened while OpenAI was training Astra but it didn't really result in anything and the model just kept focusing on its original task
>>
File: Meeting Matsumoto Teddy.png (618 KB, 1326x770)
618 KB PNG
>>110015750
I wonder what he would think about NeuroSama's anime watchalong then..
>>
>>110016365
Q0.5bros...
>>
>>110016372
how does that even work? pouring api tokens+harness?
a human+marketing gimmick?
>>
>>110015676
yes, obviously. it's only a matter of context length and multimodality, both of which continue to improve. LLMs will do this you don't need any fancy new tech. they already can do it in principle, just very poorly.
wait for someone to make a memebench about understanding a full episode of anime and everyone will start targeting that capability
>>
>>110016379
Vedal hosts Neuro on either her own high-end desktop or in a cloud server, I forgot which. So he doesn't have to pay for tokens but he does have to pay for server hosting and electricity bills. He's generally pretty secretive about how Neuro exactly works but he did put a lot of time in effort into giving her good image and audio recognition, which is pretty important for an LLM that mainly interacts through verbal talking and not a text interface. Neuro can pretty accurately describe what's being shown on stream so she absolutely can watch stuff almost like a human can.
>>
File: hmm.jpg (63 KB, 534x443)
63 KB JPG
Why aren't LoRAs used to solve the continual learning problem? Are they just not very effective on LLMs compared to diffusion models for some reason?
>>
>>110016308
no way, elon musk might be planning to do the thing he promised a year and a half ago? i am stunned
>>
I still think that on the long term, once video models will have good object/scene persistence, generate real-time streaming video and not require huge amounts of VRAM, gaming will take place with them rather than actual raster 3D graphics.
>>
>>110016392
The main issue is that video understanding is typically done by feeding it still images as frames in sequence to a standard vision model. With a few hundred to a thousand or more tokens per image it's just infeasible. If you wanted LLMs to get good at it you need semantic tokenization of videos that encode time segments instead of single frames. I don't think it makes sense to join that with the normal vision modality, you should treat it as its own separate thing in pretraining. Though maybe I'll be proven wrong.
>>
File: gr-ack.png (38 KB, 611x308)
38 KB PNG
>>110016308
they should rebrand it to grim
>>
>>110016427
That will leave something new for mistral to finetune lmao
>>
>>110016412
Catastrophic forgetting
>>
>>110016406
i heard somewhere that part of it is multiple responses are generated and another ai selects which way the convo goes. im gonna try something like that with airi to get some variety
>>
>>110016379
It's literally that meme of blind Dipsy being led by a Gemma guide, but instead of deepseek it's a finetune from three or so years back (some suspect llama base).
>>
>>110016485
i smell certain recent hype might can be put to use
>>
>>110015676
already can
https://huggingface.co/microsoft/Mage-VL
>>
>>110016487
thought it was a gpt2 finetune
>>
>>110016468
Because the weight adjustments are too weak, or the real-time training too sparse?
>>
>>110016468
isn't it the exact problem lora solves by being rank constrained
>>
>>110016515
Gpt2 isn't open weight.
>>
>>110016308
Is there even any point in using an old model like that at this point? We're already past the point of being impressed by le heckin epic uncensored outputs.
>>
>>110016553
>Is there even any point in using an old model like that at this point?
No, which is probably why Elon has no issue with releasing it
>>
>>110016289
China's slow is because they're all switching to Huawei chips and it takes time to scale out on a completely different setup
>>
File: Untitled.gif (1.31 MB, 477x443)
1.31 MB GIF
>>110013523
>>110014116
How do we feel about Qwen 3.8 Flash Next? I have a subpar system match of a 3060 and collective 48gb of RAM duct-taped that makes me get me at 30tps at generation even if the prefill and thinking is a bit slow.
>let at it with a livewire laravel starter project with 2 docs
>80k context fill from the start
I'm using the Pi harness, had a lil bit of a break when 1.0 rolled out and they forced the other TUI mode which made it crash immediately, but for like, 128k context fill, and some plugins jammed in because i don't know what's necessary and what's not, i feel pretty solid about my setup.
I feel like there's ways to improve it or optimize it still, but i don't have extra machines or the headroom to run other models adjacent so no point in having a router
>262k context fill
please
>>
>>110016081
get me a job
>>
https://github.com/SamsungLabs/LittleBit
Isn't this like, huge?
>>
>>110016625
>QAT
so it's useless for us mortals
>>
>>110016582
its big enough to be kinda smart and is fast. i have 128gb at 250gb bandwidth and i still get 1300/40 t/s balls deep in context
>>
>>110016625
Just because you can reduce precision below 1 bit with the model still being coherent doesn't mean it's a good thing to do.
>>
>>110016625
>0.1 bits-per-weight
LOCAL
WON
>>
File: 1782629111591886.png (16 KB, 872x84)
16 KB PNG
>>110016625
You can't use ghost energy to mitigate quant damage
>>
>>110016625
someone do this to gemma-chan NOW or else im putting her in the torture chamber. don't try to tell me I won't do it
>>
>>110016625
>0.835bit quant that performs better than the old Q3_K_M
holy fuck
>>
>>110016625
As far as I'm concerned, the thing about quantization is not that it has to be as good as the unquantized version, but that it should perform better than another model with the same memory footprint.
So yeah, could be cool.
Take a huge model, quant it to 20ish GB, profit.
>>
>>110016625
usecase?
>>
>>110016519
A lora doesn't have infinite capacity to absorb new knowledge, if you keep adding more it'll overwrite the previous data.
>>
>>110016683
>>110016672
>>110016657
it is not PTQ
you need training compute to use this
>>
>>110016686
equivalence to fitting bigger quants in the same vram
>>
>>110016690
Oh.
Well shit.
Guess the hope is that it's so good the labs themselves start releasing weights like they do with QAT right now then.
>>
File: file.png (142 KB, 817x1251)
142 KB PNG
>https://anond.hatelabo.jp/20260407065857
oooooooooooooooooooooooooo
>>
File: 1785612702879392.png (21 KB, 357x218)
21 KB PNG
>>110016706
>>
>>110016706
Nice fanfic
>>
>>110016706
>I didn’t just use ChatGPT. Some of the minor AI services I had tried for comparisons had a public conversation by default. In addition, in the free plan, the conversation log is indexed into the search engine .
>It seems that it was written small at the bottom of the terms of use . I haven't read it. Of course I didn't read it.

lol
>>
https://google-ai-edge.github.io/mediapipe-samples-web/#/decision/decision_maker

EmbeddingGemma2 seems pretty sick desu. very jev like but tiny.
>>
>>110016625
Ok grandpa?
>>
>>110016731
see there's retardation, and then there's just being subhuman and not using your brain in the slightest. guess which one this guy falls into.
>>
>>110016732
Use this and see all these benchmaxxed models fall apart. "Not Urgent: our production database pipeline is looking fine with no segfault and customers are able to sign in!"
>>
>>110016687
Could you not just... add more LoRAs? Duplicate it past a certain point and mark one of them as a legacy memory, then continue adjusting on the other?
I know that's running into RAG problems again, but surely it wouldn't be as much of an issue when the base recall is still neural.
>>
File: kurumi mom.jpg (90 KB, 1363x631)
90 KB JPG
>>110016706
>俺の今の状況だが、3月。無職。実家で一人暮らし。嫁と子供は嫁の実家にいる。離婚届はまだ届いてないが、時間の問題だと思う。
its so over
>>
>>110016422
>generate real-time streaming video and not require huge amounts of VRAM
Yeah that's not going to happen any time soon. They'll be building animations in game engines long before.
>>
>>110016810
>>110016810
>>110016810
>>
>>110016706
What is this chink going on about?
>>
>>110016534
https://huggingface.co/openai-community/gpt2
>>
>>110016796
There is actually an old project S-lora that would help with this
>>
>>110014627
no, that wasn't me
also now that nvidia bought them, i'm no longer worried and have started deleting models
i still keep the microsoft and tts models since they get deleted sometimes
>>
>>110014653
lmao this actually fucks me because I already use 9931 for llama-proxy vibeslop
>>
>>110016422
Very cool. So are things at the stage where you can make any kind of game assets with nothing but AI and a few prompts?
Maybe making a gemma-chan game won't be that hard.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.