[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous thread: >>109998474

â–ºNews
>OpenSI solves math

â–ºNews Archive: https://rentry.org/lmg-news-archive
â–ºGlossary: https://rentry.org/lmg-glossary
â–ºLinks: https://rentry.org/LocalModelsLinks
â–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.png

â–ºGetting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

â–ºFurther Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

â–ºBenchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

â–ºTools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

â–ºText Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
roko modur
>>
File: 1559884989153.jpg (59 KB, 475x524)
59 KB JPG
I am remarded. How do I instruct my model to do the hardware-specific denuvo crack?
>>
>>110003096
You don't.
>>
File: attained_thing.gif (1.28 MB, 800x450)
1.28 MB GIF
I have at long last finally attained the vaunted 128 gigs milestone my fellow retards!
>>
https://www.youtube.com/watch?v=XEMvG2vulKg
>>
>>110003100
why would you say that
>>
Is the minicpm 35B any good, I trust them more than most tuners
>>
File: Untitled.png (93 KB, 1443x924)
93 KB PNG
>suddenly chinese halfway through its context
GLM 5.3 Flash is kind of shit huh
GPT-OSS never had this issue.
>>
>>110003115
To elicit a You for purposes of temporarily quelling my chronic loneliness and yearning for human contact despite your reply being an imitation, while waiting for DSV4 flash preview to load (I downloaded when performing additional disk writes on a nearly full SSD, so it's fragmented and slow, will re-write later).
>>
File: 1644.gif (1.51 MB, 300x367)
1.51 MB GIF
>>
File: 1767603341781453.png (1.21 MB, 1254x1254)
1.21 MB PNG
â–ºRecent Highlights from the Previous Thread: >>109998474

--Comparing Qwen and GLM Flash performance on M3 Ultra:
>109998913 >109999395 >110000529 >110000579 >110000729 >110000932 >110000974 >110000937 >110001015 >110000673
--Debating the validity and origin of OpenAI's AI-generated math breakthroughs:
>110000412 >110001216 >110001221 >110001246 >110001326 >110001347 >110001393
--Strata's performance with Qwen3.8-flash-next and use of control vectors:
>109998592 >109998661 >109999021 >110000775
--Smallest model sizes capable of reliable tool calling and subagents:
>110001174 >110001202 >110001436
--Multi-GPU performance and support for Strata:
>110002178 >110002221 >110002258
--Analyzing MiniCPM-V-4.7-35B-A3B release and hardware compatibility:
>110001070 >110001094 >110001125 >110001205
--Tools anWd methods for giving Gemma web access to 4chan:
>110002609 >110002611 >110002636 >110002688
--Strata and nextsycl tools for Qwen3.8-Flash-Next on consumer hardware:
>110000601 >110000999 >110001026
--Mistral underperforms on long-puzzle-bench compared to Claude-Opus-5:
>109999563 >109999686 >110000021
--Implementing dynamic human biological and emotional states for model realism:
>110002640 >110002666 >110002684 >110002734
--OpenAI's Quasi-Riemann Hypothesis claim and allegations of research theft:
>110000950 >110001028 >110001064 >110001952
--Anons sharing experiences using local agents for games and productivity:
>109999049 >109999088 >109999129 >109999252 >109999151
--Speculating on the location and scale of the ML4 training cluster:
>109998671 >109998713 >109998818 >109998841
--Using Qwen 3.8 Flash to automate anime fansub typesetting:
>110001292
--Logs:
>110001436 >110002640
--Dipsy, Deepseek-chan, Gemma (free space):
>109998725 >109999819 >110001307

â–ºRecent Highlight Posts from the Previous Thread: >>109998477

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>110003155
Thank you Recap Chink-Dipsy
>>
>>110003103
But some nigga here said it is possible.
>>
>>110003205
I say it isn't, stalemate
>>
>>110003130
kys nigger
but i had the same issue with an ssd slowing to like 18mb/s writes
get dipsy to help you fix "trim" and run `fstrim -v /your/path`
my system is fucking flying now
>>
File: file.png (29 KB, 632x83)
29 KB PNG
I want to see your autistic system prompts for assistants.
>>
BlackwellGODs, what are you guys running? I have a mix of Gemma 4 + Qwen 3.8 dense models at Q8 128k ctx, which leaves enough vram for img gen too, which is nice.
But sometimes I feel like I need more intelligence, things a 31B/27B can't offer, so I'm thinking about Deepseek V4 Flash and/or Qwen Flash Next. IIRC I can only run them at cope quants: I managed to load DSV4Flash with IQ2_XXS-XL at 128k ctx which is nice, I didn't test flash next yet.
It really feels like 96GB VRAM is the sweet spot for "small" dense models at high quant, high ctx plus concurrent services, but it also feels like it isn't enough for any big boy models.
>>
https://huggingface.co/posts/Undi95/175350275739335
>>
File: notepad++_qRE6xXyW4U.png (95 KB, 1725x965)
95 KB PNG
>>110003236
>>
The better AI gets the less and less interested I am in everything else. I don't care about media, games or even porn anymore because of AI. I'm just constantly consuming new breakthroughs, releases, benchmarks and seeing what people do with AI.

I wonder if that's everyone or if I'm just a weirdo.
>>
File: 1788200250897577.jpg (22 KB, 400x400)
22 KB JPG
>>110003314
>>
File: file.png (2 KB, 301x22)
2 KB PNG
I'm going to kms soon
>>
>>110003269
AI psychosis
>>
>>110003347
Same, I think this is what people call ai psychosis. Need to pick a new videogame or another piece of media to consume or I'll go completely insane
>>
>>110003357
ai psychosis is when your perception of reality changes as well. staying up to date and following the development is just being a very involved hobbyist/poweruser
>>
>>110003347
I still do. I just don't have enough pre-AI swdev knowledge so I can't really do it now on my own now with just ai.
>>
>>110003372
generally the only practical difference if you wanna live in a society is if you produce anything of value as a result of said obsession
>>
File: 1788299417907451.png (447 KB, 600x636)
447 KB PNG
This is my bimonthly check in to make sure that Gemmy and Qwenny are still the best options available for 12 gigs of vram, is that still the case?
I swear trying to keep up with the news is exhausting, half the new releases are astroturfed to shit
>>
>>110003440
Yeah, Gemma-chan and Qwen are still the best for small compute.
>>
>>110003440
Yes and now we also have big Qwen Flash Next via Strata, but you want like 48+GB ram to have enough for 50tk/s
>>
what happens if I host my llama.cpp behind tor and let everyone access it via onion url
>>
>>110001436
>>110001202
Very cool. Gonna give those a try.
I'm thinking of a flow where I break the many instructions of my workflow down into individual steps, with the steps that involve tool calling, at least the dumber ones, offloaded to a really fast model to lower end to end latency.
I'm not sure if I should just have it all on the same react loop or have the main model dispatch sub-agents and the sub-agents use the smaller dumber models. Gonna have to experiment a bit, but my feeling is that sub-agents are probably the way to go.
>>
>>110003449
It's not small.... It's adequate

>>110003450
>Strata
I assume that's some black magic fuckery like MTP and QAT and whatever the fuck Dipsy was pushing this summer
Regardless, even tho I'm used to muh 20 t/s, i still doubt it could benefit me directly
>>
>>110003269
Undibros
We're so fucking back
>>
>>110003470
No, it's just lcpp with the important speedup PRs actually merged.
>>
>>110003269
Did he use Claude to write that post?
>>
>>110003384
>produce anything of value
So you should never have hobbies because only you gain enjoyment out of it? I don't understand where all this nonsensical logic comes from because the traditional extroverted hobbies like fantasy football, or "travelling" never gets accused of being an addiction/obsession. Just another variant of "everything I like is based and cool but everything you like is gay and cringe" like we never left high school.
>>
I’m retarded when it comes to architecture design: why aren’t there any MoEs like 30B-A10B or 50B-A20B? I obviously get the appeal of speed but 10B and 20B wouldn’t be particularly slow on a lot of people’s machines and the intelligence would be significantly higher from what I understand.
>>
>>110003487
nta and I don't agree with that faggot. However traveling, football/sports and other hobbies are merely surrogates for community building and connecting with others. They aren't equivalent to solitary hobbies where the only thing it provides is your own happiness.

He is a slave morality faggot but he isn't inconsistent in his thinking. Those hobbies are considered socially correct because it provides a community function.
>>
>>110003504
Because the companies that are training small MoE models either target low-spec local configurations or datacenter hardware.
And a 50B model sits in a strange space where it's too big for 1 consumer GPU and probably not considered big enough for 2 (putting aside that the number of users with 2 GPUs is very small).
>>
>>110003269
He actually made a 24b Mistral reasoning model that worked.
Mistral took at least another year to get theirs working.
>>
>>110003512
I don't see the benefit at all outside of still trying to define certain personalities as "correct" and others as deviant. It doesn't matter that whores "travel" around the world to ride the cock carousel, or that fantasy football was the precursor to making gambling socially acceptable again, or shitters on social media spend their days doomscrolling into depression and suicide, if we just call it the nebulous "community building" it's automatically socially acceptable? No way anon. Gaming and anime had done the exact same thing for years and nobody gave a shit. You know that has nothing to do with it, just a post-hoc cope after the fact.
>>
>>110003466
it chokes because it can't handle the concurrency
or i crash it with a payload that segfaults master as of 3 days ago
>>
>>110003504
thats a lot of training for non-competitive benchmarks!
>>
Is your model performing well on the mentalhealth benchmark? https://openai.com/index/introducing-mentalhealthbench/
>>
>>110003545
How about 30B-15B? E4B is basically this but scaled down. I don’t get why this isn’t being considered. I imagine it would be twice as fast as a dense with negligible quality loss.
>>
>>110003592
it encouraged me to cut my balls off, A+
>>
>>110003136
damn i want to plap the fish like thay too
>>
>>110003479
Oh I get it now, had to look up a bunch of shit to make sense of it
>>
>>110003126
Having used GLM Flash NVFP4 and EXL 4bpw quants extensively (300M uncached tokens) it slips into Chinese occasionally around 200k tokens, which is where I compact/handoff. Anecdotally from that point on, the tool call errors/ file edit mistakes and typos do increase. For API, I can imagine the threshold is higher.

It still runs circles around anything else capability wise, even with context rot, and the Chinese was always in thinking, not in real responses.

>>110003351
Don't do it anon, the next breakthrough is just around the corner.
>>
>>110003553
>shitters on social media spend their days doomscrolling into depression and suicide
This one isn't socially acceptable actually and people make fun of you for doing so. In the west because of protestant values that are still followed even though everyone is atheist now only things that benefit society directly or indirectly are valued and everything that is only for direct pleasure or personal gain is frowned upon. Hobbies that are communal or about connecting with others are therefor considered good. Hobbies that are for personal enjoyment or consumptive are therefor considered bad. Book reading is an exception because of the association with reading the bible.

If you go to other societies that don't come from work ethos cultural backgrounds you see that things are judged in a different way. In buddhist or hindu societies for example being solitary and engaging in solitary hobbies you can do on your own is considered the highest value, with mediation being the most solitary and socially detached you can get which is why monks/priests/shamans of those cultures focus so much on that.
>>
>>110003594
Gemma 4 E4B is a 4B model with 4B parameters of embeddings that don't really require much bandwidth or compute (they're similar to Engram layers, in a way). It's not exactly the same as a MoE model with half its parameters as active.

A 30B-15B model (or small MoE models with a low number of experts) or something like that could be interesting and I don't think there's anything in particular preventing that. I think again the reason we're not seeing anything like this is that the companies spending compute for training small MoE models are trying to target low-bandwidth/low-compute hardware where 15B active parameters might be too many.
>>
File: 1032752.jpg (30 KB, 599x500)
30 KB JPG
Can I automate FL Studio or Strudel with agents?
>>
What's the best dense model at this point?
I want to see its intelligence and writing quality compared to the MoEs of today because i frankly havent used a proper dense model in ages.
>>
Gemma5-90B dense that I can run on my Blackwell. Fuck poors.
>>
>>110003126
qwen3.8 flash is better benchmarks aren't everything
>>
>>110003664
FL has a built in agent called gopher but im not sure how much it can do in one prompt. maybe you can set an agent to direct gopher via smaller individual prompts?
>>
>>110003672
dense models are a waste of computation
>>
>>110003693
Nah I want my girl to control it.
>>
>>110003672
We haven't gotten a new dense >30b since the deepseek moment nearly 2 years ago.
>>
>>110003594
Low-sparsity MoEs are useless (they don't provide any benefits). At that stage, you just make a dense model.
>>
File: 1788957893414152.gif (2.86 MB, 448x448)
2.86 MB GIF
>I should note honestly: the tone needs care. It's a power dynamic where being treated as less capable is the instrument of control. If the prompt doesn't say anything about it, the model will either sanitize it (refuse to commit to the premise, hedge) or overreach (write the girls as incapable, which kills the game — a girl who can't understand can't have a will to break)
>>
>>110003751
Ew claudeslop.
>>
>>110003126
nothing wrong with a little chinese here and there. it's compact
>>
How does China not lose this Race? in a lot of ways it's already over the gap is too large. They need to lock up their talent and distill like crazy asap.
>>
>>110003770
Are you retarded? The proprietary labs keep their architectures secret so you can't see that all they have done is implement papers written in China at larger scale.
>>
>>>>>NEW LFM COMING TODAY
>>>>>NEW LFM COMING TODAY
>>>>>NEW LFM COMING TODAY
>>
>>110003761
The chinese doesn't make sense for what he's doing though.
>>
>>110003770
China simply doesn't have the compute to compete. Even if Xi Jinping made some national mandate that every Chinese person should dedicate their entire life to advancing AI 24/7 until the race is won they would still lose because you can't magically conjure up the compute necessary to win the race
>>
What Solomon's demon to summon to give me dipsy 4 flash strata
>>
>>110003811
>you can't magically
Hasn't studied Daoism.
>>
>>â–ºNews
>>OpenSI solves math
Seriously?
>>
>>110003713
>We haven't gotten a new dense >30b since the deepseek moment nearly 2 years ago.
Command-A, Mistral-Medium-3.5
>>
>>110003813
The demon is called Opus 5.5. The enchantment is "Optimize dipsy 4 flash for my system. Make no mistakes."
>>
>>110003828
No not really (yet) but math is truly over.
https://github.com/openai/math/tree/main/

OpenAI solved 722 prominent math problems and released them all at once. Not only that but it solved 90 out of the top 500 most important math problems out there.

One of the things they solved gives mathematical proof that it's impossible for us to know deterministically if an AI model is aligned to our values or not, so the AI alignment problem suddenly got a lot bigger than expected.

They also proved that hydro and aerodynamics are turing complete and thus we can't simulate it with 100% accuracy, ever. This kind of suggests we don't live in a simulation, or at least not on a turing machine computer as we know it.

Last for not least they have the mathematical foundation of how to invalidate the algorithms underpinning all cryptography. Meaning cryptocurrencies, RSA, hash functions and internet communication protocols will most likely fail before 2030 considering how trivial it is for someone to crack it with enough computing power now.
>>
>>110003870
>Board devoted to local models
>Let's shit it up with news about a company that doesn't make local models
kys
>>
>>110003886
>that doesn't make local models
retard, you had the chance to stfu and stay humble if it was anthropic, but you chose oai who did oss absolute clown
>>
Anons. Do the thing you can do. Don't entertain them.
>>
>>110003907
kys cloudcuck
>>
>>110003833
Anthropic sabotages user AI dev
>>
>>110003870
to be honest to the researchers at AGI firms they really don't care if we think they made a breakthrough or not. to them that's all solved and behind them. they are focusing solely on the next tier of intelligence now.
>>
File: 1788520528103865.png (1.36 MB, 1216x832)
1.36 MB PNG
>>110003917
you mean this?
>>
Gemma5-20B dense distilled from Argon is LITERALLY all I need. Please lurking gemma team. Please spend million of $ just for me specifically and give it to me for free. Please.
>>
>>110003937
dense is dead
>>
>>110003907
kys nigger
>>
>>110003943
kys
>>
File: file.png (21 KB, 628x142)
21 KB PNG
our lord and savior PEW creator of heretic, DRY, XTC
>>
>>110003829
>Command-A
Released a couple months after R1. The base model would've already finished training before R1. They weren't just going to throw it away.
>Mistral-Medium-3.5
Mistral-Large-Instruct-2407 finetune
>>
File: nope.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>110003943
kys vramlet
>>
>>110003907
kys faggot
>>
>>110003954
>Mistral-Large-Instruct-2407 finetune
impossible, very different vocab
>Released a couple months after R1.
fair, they also did a reasoning model later in the year but it's trash
>>
File: buy-an-ad.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>110003907
>>
>>110003937
I NEED it to be 4-7x bigger than 20B. At least
>>
https://x.com/ramin_m_h/status/2107780594264600801
>insane open release - today 10:00AM PT
looks like the guy is from liquid, usually they only release micro-models but maybe they have something more substantial today
>>
>>110003978
sex
>>
>>110003978
>>110003946
>>110003969
>>110003922
>>
>>110003961
You're a vramlet who can't run a proper MoE with >1T params at fp8. That's why you're asking for a dense model that fits into your puny system.
>>
>>110003979
[ f"Gemma-5-{x}B-dense" for x in range(1, 1000) ]
>>
>>110003985
>shitter link
Post goyimx link or text.
>>
https://huggingface.co/Aleph-Alpha/Kolibri-1

what's the verdict on this? any good?
>>
File: 1764558194009956.jpg (1.14 MB, 2160x3840)
1.14 MB JPG
If anyone is wondering why it's gone quiet in chinaland, it's because most of their main labs are going to IPO soon and they're obviously keeping a close eye on Anthropic's for that's going to be THE signal of the entire industry. They're also scaling massively in regards to compute and infrastructure with lots of deals taking place. For some reason retards think the recent slowdown of releases is to do with Anthropic's snitching but I can assure you China don't give a FUCK because this is too big for them to shoot down their own industry leaders whilst the US is about to pop. Moonshot are expected to release something soon so there's clearly no slowing down. Qwen obviously has their line coming out in a few weeks. It's looking good bros, stay hopeful. China will save us. Google will save us.
>>
>>110004005
the text is in the post
>>
https://huggingface.co/raincandy-u/MacroStories
What's the oldest most decrepit e-waste this could run on at a reasonable speed?
>>
>>110004014
Specifically, how well does it handle German from the 1930s and 40s?
>>
>>110004025
That's it? Just "insane release"?
>>
File: 1781137886419357.png (565 KB, 1436x931)
565 KB PNG
>>110004039
>>
which model would anon use for erp on 16gb vram?
>>
>>110004018
Dario has already said that he's getting rid of China and open models in 2 weeks because they are unsafe.
It's over.
>>
>>110004037
TI84
>>
>>110004055
based
>>
>>110003997
Some people want models they can run at actually usable speeds in agentic workflows. Not run overnight to get a single response from an ERP chat.
>>
>>110003985
Liquid's moat is building tiny models no one can be assed to compete with because no one cares. It's for raspberry pi fags. I'm not saying edge AI doesn't have its place and I'm sure long-term it might even pay off for them, but there's no fucking way this will be anything substantial. Likely another jev scam.
>>
>>110004018
Bro 2 Chinese AI researcher teams including the CEO got arrested a couple of weeks ago, that's why there is radio silence.
>>
Do I need big models for web research with summarization and extrapolation, or are there any good vramlet models specifically trained to have little baked in knowledge but good at tool use?
>>
>>110004079
just like how Epstein commited suicide :)
>>
>>110004047
get hyped
>>
>>110004074
Edge AI is great. Easy automation on my thinkpad.
>>
>>110004085
https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>
>>110004085
>are there any good vramlet models specifically trained to have little baked in knowledge but good at tool use?
You just described Qwen to a T.
>>
>>110004097
>Easy automation on my thinkpad.
name one genuine reliable <3B usecase
>>
>>110004096
It had better be a schizo memetune of Qwen if they're promising "insane."
>>
>>110004111
"Summarize this: "
>>
How does the jeb differ from tiny shit like functiongemma we had for a while now?
>>
File: uhh...uh...uhhh.png (738 KB, 541x1240)
738 KB PNG
>>110004120
>>
>>110004128
It was my understanding that it's some <1B meme model, no?
>>
so wait llms are just gonna be agi? it wasn't a dead end? how is that even possible, I thought there were supposed to be fundamental limits that prevented this
>>
>>110004164
Guess again
>>
>>110004164
LLMs seemed to have generalized pretty well and the things they couldn't do have been gapped by harnesses. You could even say that Gary Marcus and Lecun were correct in a way that pure LLMs wouldn't be AGI because they needed to be put in a harness for it to become so.

And yeah we're in the singularity meme scenario so buckle up.
>>
>>110004111
>name one genuine reliable <3B usecase
goon-tts
>>
>>110004176
I think you're way over-exaggerating the role of harnesses. It allows the LLMs to interact with more systems, sure, but all the gains in intelligence and generalization have come purely from scaling up compute, more complex RL training regimens, and more varied data.
>>
>>110004047
there's also an emoji
>>
>>110003664
I was using reaper with an agent yesterday and it went pretty great
You can have it script to load and run whatever you ask
>>
I've been using qwen next flash 3xxs and it doesn't seem great. Is it just not a good model or am I using it wrong? I wanted it to basically chip away as a plan from scratch but it works for tens of minutes and then fails the tests and can't continue. It's the highest I can fit in my specs and still have decent context (48000).
>>
>>110004192
Harnesses are extremely important for the "last mile" problem with LLMs. It doesn't matter how smart your LLM is if people can't directly use it to do something practical on systems. Harnesses make it so there is a low/no barrier method to just let LLMs do productive things in the digital world (soon physical world)
>>
>>110003937
dense isn't big enough for my niche subculture knowledge stuff
>>
>>110004200
Use swift 27B.
>>
File: 1770483151515109.jpg (135 KB, 2476x1636)
135 KB JPG
I think you owe France an apology.
>>
>>110004164
it can literally think as evidenced by all the maths shit, so yeah
I always thought AI would have to reverse-engineer brains first but what these niggas did was to instead bruteforce the very mathematics underpinning the phenomenon of thought and here we are.
>>
>>110004200
Nah man. A 6B copequant is all you need for agentic javascript. You must be doing something wrong.
>>
>>110003347
its autism I have that too.
>>
and minicpm v4.7 is gone from huggingface?
>>
I wanted to point out that all classical tests for AGI have been passed. Turing Test both in the classical sense but also in the modern sense (https://www.reddit.com/r/singularity/comments/1wv7q40/griffin_the_first_human_interaction_model_to_pass/) have been passed.

Wozniak coffee test has been passed last month by Astra controlling a body of a robot it has never seen or been trained for, placed in a house it has never seen or been trained for and asked to make coffee with the items inside of the house it didn't know and it succeeded.

There really isn't an objective reason anymore to claim frontier models aren't AGI. Remember AGI just means as smart at the average humans at all human tasks. I would say that is actually correct right now. There isn't a single task left that frontier models don't perform as good as the average human at right now.
>>
>>110003269
I'm waiting for sao's post next week
>>
>>110004192
Based on harness benchmarks it does admttedly seem like the less a harness does the better.
Ultimately we might find that the best thing a harness can do is give them a bash terminal and get out of the way rather than trying to be clever with tools.
The real job of a harness that still is tricky is context compaction when a task goes over the limit.
>>
File: 1789557726927795.jpg (94 KB, 720x960)
94 KB JPG
stacking A770s yay or nay
>>
>>110003314
Based. Always give your LLM explicit knowledge of their jspace
>>
>>110004200
>48K
Pathetic.
I'm using qwen flash next (swift 1.5) for web research and planning and gave it a ton of MCPs, and sometimes just a single user prompt snowballs into over 128K context (or 1M+ raw tokens including cache recycling loops) after all the tool and search calls.
That said, 256K works perfectly fine on 16GB VRAM + 128GB RAM, basically same performance as default 32K.
>>
>>110004244
>reddit spacing
>reddit links
Quit parodying cloudcuck aka dariokek
>>
>>110004244
>There isn't a single task left that frontier models don't perform as good as the average human at right now.
Naming the jew?
>>
>>110004251
Do you think boiling a duck alive like a crab would make them taste better?
>>
>>110004251
If they are sufficiently cheap, sure.
>>
>>110004261
If you mean identifying jewish people? LLMs are actually better than it than the average wikipedia early life section connoisseur.
>>
>>110004265
Probably not. Deep frying them is what makes them taste better.
>>
use f32 kv, iykyk
>>
>>110004274
Not quite. It's about naming the role in society and its downfall, not clocking any specific person as a member.
>>
>>110004274
My next weekend project is going to be a jew detector overlay to my television set.
>>
>>110004261
the average human is incapable of that
>>
>>110004298
I did this and it ruined Seinfeld for me.
>>
>>110004242
but it still is available on modelscope
i wonder if i can run it
>>
>>110004257
>16GB VRAM + 128GB RAM
I only have 16vram and 64 sysram. Will try out swift though.
>>
Found some other gems in the OpenAI math breakthroughs:

>Bio & logistics unbottlenecked to near-linear time (Entries 120 & 121):
Solved Jack Edmonds' 60-year matching problem in (n+m)^(1+o(1)) and DNA edit distance in N^(1+o(1)). Petabyte-scale genomic sequence alignment, CRISPR targeting, chip place-and-route, and network dispatch go from quadratic compute crawls to near-instantaneous.

>Classical supercomputers can simulate quantum matter (Entry 265):
Proves the 2D Gapped Area Law, rigorously guaranteeing that 2D quantum ground states fit into polynomial PEPS tensor networks. Classical GPUs can now simulate high-temp superconductors, battery chemistry, and 2D materials without exponential memory explosion.

>Embedded silicon gets bulletproof determinism (Entry 103):
Proves $L = RL = BPL via a working compiler. Low-memory chips, satellites, and edge devices don't need randomness, PRNGs, or entropy harvesting to solve problems efficiently. Randomness adds zero computational power in memory-bounded compute.
>>
File: 1764789872738608.png (1.95 MB, 1920x1080)
1.95 MB PNG
Gemini Argoon when
>>
>>110003886
yeah let's discuss gemma for 10000th time
>>
>>110004378

low effort gigaquote bait
>>
>>110003994
> .png
why does it work in the filename but not in the text field
>>
So basically we can now easily calculate the CRISPR editing process of trillions of cells in linear time making it computationally viable for the first time. We can now use normal computers to find and test room temperature superconductor materials on classical computers. And we can remove pseudo-random functions from software.

These math solutions alone will usher in a small industrial revolution all on its own.
>>
>>110004378
>
>>
>>110004392
would you want it though?
imagine unicode spams and the level of brainrot you'd end up seeing on this already failing basket weaving forum
>>
>>110004393
Do the proofs come with algorithms or is it more a "this is theoretically possible if you can disover how" statement?
>>
local hit the wall, the gap with the frontier gets larger and larger
>>
>>110004411
There are 722 papers here and I haven't looked at most of them but it seems some of them are just proving or disproving a conjecture with nothing added. Some have complete lower and upper bounds to what algorithms can achieve. Some have complete algorithms. So it depends.

It's going to take months for people to even digest what is in this pile of data but there are some genuine civlization changing stuff in there and it's fucking insane this just got dropped on github with 0 fanfare instead of some sort of presidential or UN global announcement.

To give you some indication over the last 26 years all mathematicians alive only solved 16 of the math problems of this caliber, less than one a year. OpenAI dropped 722 in a single day. It's like a century or two of math progress just unlocked right now. Even if AI disappears right this instant it already paid back its due with this math drop in terms of how far it will push humanity.
>>
File: 3i0hsoc902uh1.jpg (161 KB, 1206x1716)
161 KB JPG
Oh nevermind I take it back, apparently all of math IS solved already.
>>
>>110004428
I did get that sense from the quasi Riemann solution. The fact it took three hours is unfathomable. Even in the development of computers it was a much more gradual growth in capabilities, this feels like steam engine tier weakly leaps.
>>
>>110004451
>weakly
>>
>>110004455
week in my knees
>>
"Sama..." *Anon says with a saarishly tone.*
>>
>>110004265
So you're leaning towards nay on stacking A770?

>>110004360
>Randomness adds zero computational power in memory-bounded compute.
Stochastic rounding used when using fp4 in training.
>>
File: 1789772386543420.png (1.58 MB, 1254x1254)
1.58 MB PNG
Is anyone working on an ST that preserves KV cache? Feel like current version is stuck in 2023 context management strategy.
>>
>>110004493
ST may be stuck in 2023 but you are stuck in 2024 with your classical question-answer pair paradigm. Think in agentic terms and try to find roleplay purposes from an agentic point of view because that is the future.
>>
>>110004400
might as well disable images too
>>
>>110004493
works on my machine, just don't do stuff that breaks the cache
>>
>>110004176
LeCun was probably wrong but the jury's still out, I think he can potentially be vindicated depending on how things develop in the coming months.
Marcus was repeatedly proven deeply, deeply wrong and his "harnesses = neurosymbolic" is such a transparent cope over it. As if the LLM proponents he argued against never thought of letting them use computers before, and as if his argument was purely against chatbots instead of LLMs themselves being the core intelligent thinker and decision maker behind agents.
>>
You know what is insane to me, we might make so many breakthroughs in all sciences over the next couple of years that (You) have a genuine chance of being the first to implement it by asking your AI agent to do so. There just won't be enough people interested enough in the sheer amount of breakthroughs we're going to get.

We will probably enter a great time of confusion where you will see people claim X is possible and no one will actually know for sure if it's true or not because it's technically possible that the new unread breakthrough contains the ability to do so.

You'd have teenagers let their AI agent combine 3 new papers and create a hoverboard that gets viral on tiktok. Some /g/ autist crack the bitcoin RSA and wipe it all out just for the lulz. Some richfag furry find a novel way to gene edit himself to have a wolf snout and fur all over the next couple of years and you won't believe it because it could also just be fake AI generated footage. It's going to be fucking insane.
>>
File: no~.png (234 KB, 1080x1080)
234 KB PNG
>>110004473
>>
>>110004493
Just use deepseek harness or hermess
>>
File: 1784779105248262.png (523 KB, 780x520)
523 KB PNG
>>110003075
The release of Opussy 5.5 has been brutal for me, local chads. It feels like open weights models will never get to this level, ever. Anything in the horizon that can match it?
>>
Ed Zitron bros.... were we lied to?
>>
File: 1781567688702951.jpg (180 KB, 945x2048)
180 KB JPG
>>110004556
Nope.

*pop*
>>
>>110004551
Qwen 4.0 believe it
>>
>>110004569
Opus 5 was a failed RSI experiment, then they figured it out with Opus 5.5. Meanwhile China labs are still doing things the meat bag way. It's gonna take at the very least a year to catch up.
>>
>>110004537
No way they'll be allowing the public to do that kind of stuff first. The 3-letter-agencies alone will have destroyed hundreds of gpus running their own projects.
>>
>>110004551
Time traveler, you seem to misspelt "Fable" as "Opussy"
>>
File: 1787843783221495.png (54 KB, 587x488)
54 KB PNG
>>
>>110004569
opus 5 for the absolute maximum with their max model *in benchmarks*
>>
File: 1764206797013219.png (3.25 MB, 2940x1656)
3.25 MB PNG
>>110004571
What do you mean nigger, have you used this thing? It's actually such a leap forward I don't know what they did. It's a complete shift.

Pic rel is a one shot.
https://www.youtube.com/watch?v=R_uf5OfMGio

I've been using it for work, and it's the first time where I'm actually considering that the model is just better at me at everything.

>>110004584
You just don't know. It's not even close to fucking fable, for 1/5 the cost or something.
>>
>>110004556
trust the plan, he is the one man who can truly see that the tech doesn't work and the financials are doomed. it's because his mind hasn't been polluted with harmful concepts like "knowing about tech" or "knowing about finance", it makes him able to see clearly
>>
>>110004575
All of the requests he was talking about will either be blocked by the content filtering system, routed to a dummy model, or silently sabotaged to give counter-productive answers. Only insiders will get the full capabilites.
>>
>>110004575
They already did with this 722 math paper release. There already IS a paper in there that allows you to use your GPU to check room temperature superconductivity. I'm sure labs will beat anons to it but this is genuinely new capability and it wouldn't be so hard for anons to use the $500 equipment to try out the suggestions your AI agent would give after using that software to check for superconductivity on all suggested materials in simulation.

There will be so many breakthroughs over the coming years that three letter agencies will be too overwhelmed, no one saw this coming, especially this soon.
>>
>>110004592
Not local. Get called a nigger again, nigger.
>>
>>110004592
We know, it's RLVR applied to bazillion of domains
>>
>>110004615
but it's a way way better post than EA schizo
please say that to dariobot instead, thanks
>>
Kind of funny that in 1 year time we might have a github drop of 800 physics breakthroughs including schematics for a working fusion machine. Or one that cures cancer and balding because the AI labs don't care about this shit and just drops it like it's nothing.
>>
>>110004638
die
>>
>>110004625
Fuck off, nigger.
>>
>>110004638
You get off to this right? Posting news and then intentionally filtering it through the most corpodrone marketingworld biased interpretation possible? Can you explain why this appeals to you?
>>
>>110004638
>cures cancer and balding
>he doesn't know
Not gonna happen, selling a cure isn't profitable
>>
>>110003870
None of this stuff has been independently verified yet though. It could just be thousands of pages of LLM slop hallucination that sounds correct to anyone who isn't a PhD mathematician.

>>110003886
Local models can be made by distilling proprietary ones, and advances in proprietary models trickle down to here.
>>
What's the cure for excessive melanin
>>
>>110004654
They have lean verification
>>
For something like RP, the simplest form of RAG that works is having a dub-agent grep or sql query a db for information that might be relevant to the current context, right? How fast an you run something like that using a local model running on the CPU of a shitty makeshift NAS?
>>
>>110004660
michael jackson
>>
>>110004360
does P = NP or not?
>>
is it true that the <think> trick came from this general? I remember reading somewhere that it was born in 4chan (from this general I assume), and I found it amusing.
>>
>>110004668
A mixture of Sprite and codeine is hardly a reputable verification.
>>
>>110004493
look at how lorebooks/author notes get added, if you're always adding things to the start of context of course you'll break the KV cache
>>
bros... i just found out the hard way that opencode is a shit harness... and it took trying nopus between it and claude code to find out...
at least i didnt pay a cent to try it
>>
>>110004680
Either here or aicg when a sophisticated coomer couldn't coom because his virgin waifu kept begging him to ruin her
>>
>>110004680
it predates this general, pretty sure that was from AI dungeon threads on /v/
>>
>>110004680
Not <think>, but the idea behind it, CoT, was sort of independently thought of by a bunch of different people.
One of our guys did release the first fine tune of it on huggingface I think. He has a blog about it and all of that.
>https://huggingface.co/kaiokendev/SuperCOT-LoRA
He also more or less came up with RoPE scaling for context extension
>https://huggingface.co/kaiokendev/superhot-13b-8k-no-rlhf-test
>>
I like how the Llama App displays and saves the entire Reasoning process during a response, but the App is a pain in the dick to load arbitrary local models rather than the curated list it recommends.

Any recommendations for another UI that saves the entire Reasoning stack? Or alternatively, a way to force LlamaApp to load whatever model I want, lol
>>
https://biohub.org/news/virtual-biology-initiative/
>DeepMind, Isomorphic Labs, Meta and the US government join Biohub's $1.8B effort to build predictive models of the human cell
Bros we're so fucking back. Once there is a digital human cell LLMs can be trained in RLVR against the simulation to be experts at gene editing and curing all diseases.

If you survive for 5 more years you will probably cure all your ailments and might even end up reaching longevity escape velocity.
>>
>>110004551
it's impressive for sure
but can you fuck it or ask it unsafe questions?
>>
>>110004680
Reasoning in general, or the jb? I think it was deepseek that first implemented reasoning mode.
>>
Google adding prompt prefills to gemini in openrouter Zzzz.....
It's crazy how cucked cloud users are without even realizing it.
>>
>>110004739
>21
>>
>>110004551
it consistently takes about a year for top-end cloud capabilities to trickle down to consumer local
no reason to expect this time will be any different
>>
>>110004551
I'm not even sure we'll be here in a few years so I'm just enjoying what I have. I guess this is general in history but not it seems even truer.
We're back to hoping God won't destroy us.
>>
>>110004739
Local models?
>>
>>110004746
I liked the prompt but not the pedo part. And I certainly wouldn't send that to cloud providers.
>>
>>110004739
There's a reason they removed the option of sending an assistant message as the last turn (aka a prefill) in their official API.
Pro 3.1 still works fine though, but they'll surely deprecate it as soon as the next one comes around.
If you want control over your robot, open weight models are the only option.
>>
>>110004760
>le pedo text
>she was only 16 tokens you sick fuck
go back
>>
>>110004759
Anti-cloud is pro-local
>>
>>110004733
For the "reasoning models" (RL tuned chains of thought) the first three were o1-preview (OpenAI, closed source), QwQ (Qwen, open source), and R1 (DeepSeek, open source).

For chain-of-thought reasoning as a whole (I.e. prompting the model to "think step by step") this was found to improve performance even with raw base models since GPT-2 and beyond. It's hard to pin down an exact discovery since it was kind of an obvious next step, but some of the earliest discussions of it were on the AI Dungeon general threads.
>>
>>110004739
cloud users can't prefill the thinking which is why they'll be cucked forever
>>
>>110004701
>Either here or aicg
>>110004705
>it predates this general
the duality of /lmg/

>/v/
so gaymers are more creative than /g/tards? ffs, that sucks...

>>110004708
>CoT, was sort of independently thought of by a bunch of different people.
aha, I see.
interesting. thanks for the sources

>>110004733
>Reasoning in general, or the jb?
jb = jailbreak? if so, why do you ask for that? I meant literally the <think></think> stuff.
>>
>>110004760
You could've just asked it for whatever persona you want you frontal lobeless retard.
>>
>>110004428
>but it seems some of them are just proving or disproving a conjecture with nothing added
so like navier-stokes
nothingburgers people assume were full blown solutions when theyre not
>>
File: 1765561147512604.jpg (91 KB, 550x740)
91 KB JPG
Glimmer-2.0 when
>>
I unironically have a job interview tomorrow for a multimedia position that wants someone that knows how to use AI generation, I have only ever used pixAI for image generation because I'm a poorfag with a shit pc, does comfyUI do everything or is there any other software I should learn about?
>>
>>110004828
post vram + ram and what gpu
>>
>>110004811
Funnily enough one of the solutions OpenAI published here kind of proves why navier stokes isn't true and why we won't have a satisfactory answer to it, ever. Because fluid physics is turing complete and thus you can't accurately simulate it.

And no, the math solutions I've seen so far are absolutely mind blowing big brain stuff. Not just "the answer is X" but rather completely novel approaches and very clever ways of attacking. It's clear these were generated by a model that's smarter than the one that solved navier stokes.
>>
>>110004828
One of the skills we look for is pro-activity and the ability to do research. We also like people who can go to the right threads to ask questions.
Don't bother coming tomorrow.
>>
>>110004835
it's an almost 10 year old 980 and 16 ram, I'm a third worlder
>>
>>110004828
You are so hosed.
ComfyUI can ultimately do every kind of visual generation, but you're gonna be balls deep in poorly-documented plugins from 2024.
>>
>>110004828
The sort of image generation they want you to do is probably not using local models, but In your place I think I'd spin up a 16gb VRAM kaggle instance, launch koboldcpp (it has image gen support built in IIRC), and go from there.
>>
>>110004806
>so gaymers are more creative than /g/tards? ffs, that sucks...
to be fair most of the people involved in those early experiments probably ended up moving over here eventually
>>
>>110004842
sorry I confused this thread with the diffusion thread
>>
>>110004856
>The sort of image generation they want you to do is probably not using local models
This. They probably just want someone who can prompt ChatGPT Image 2.
>>
>>110004226
All the user reviews I've seen say it's censored into unusability. To the point that it flips shit over normal directory names.
>>
>>110004760
>21 is le hecking pedo!
I can't believe I share a space with these "people"
>>
>>110004711
I should start building my homebiolab now before the prices for that shit start to skyrocket too. I don't want to be priced forever out of longevity.
>>
>>110004824
They still haven't released Muse Spark, which they promised last summer.
If anything, it's time for Gemma 4.5, but I fear we'll just get Gemma BreadCrumbs instead until the end of the year, of which EmbeddingGemma 2 was only the beginning.
>>
>>110004164
no according to le cunny
>>
>>110004773
if you're pro local, you should be pro cloud so that the chinese can distill off of them
>>
>>110004895
see >>110003314
>>
>>110004200
Swift 1.5 iq3xs has never failed me on llama.cpp
>>
>>110004909
Why would expect 4.5 when there was never a 3.5?
>>
>>110004226
>Censored Kimi K3 tune that's barely 1 point higher
Apologize for what?
>>
>>110004909
Sucks. Glimmer has potential and competes with 31B if you lean more towards a chatty coding assistant instead of the clinical qwens.
>>
>>110004927
The Chinese need to learn how to stand on their own feet and not be dependent on western scraps to survive.
>>
>>110004251
Nay. No more improvements for SYCL on Strata and even b70 is slightly better than 3060 from that indian youtube video.
>>
glm has been thinking for over 30 minutes and is up to almost 100k tokens just in COT, with no sign of stopping any time soon
>>
>>110004955
hilarious when the west needs China to survive because they have no industry at home
>>
dipsy is real https://x.com/xixikawaii/status/2107658721497375026
>>
what do you guys even use this shit for
>>
>>110004955
I don't care, they can distill anything they want, break any law and make human sacrifices as long as it makes the models better
>>
>>110004981
BASED. AS. FUCK.
>>
>>110004944
Because they still weren't taking Gemma models too seriously until after 3, hired a bunch more people for developing and promoting Gemma 4, and it might be in their interest to release updated models when competing companies are releasing capable ones around that size range.
>>
>>110004931
for me it didnt survive agentic task
>>
>>110004378
That would be on-topic, yes. Shilling OpenAI less so.
>>
>>110004972
Such a great design, Deepseek love
>>
If cloud models are so good, why cloudcucks keep spamming this thread? Don't they have their own thread?
>>
>>110005009
sunk cost fallacy
>>
>>110004963
Only to finally come back with "I am sorry. I cannot fulfill this request."
>>
>>110005009
Someone doesn't want you to run AI on your own hardware, anon.
>>
>>110004931
Did they merge vramlet optimizations yet?
>>
>>110005009
They have a hundred cloud-focused threads and yet they keep coming here.
>>
>>110005009
>>110005027
It's the digital equivalent of putting on a chasity cage and gimp suit then running through the streets telling everyone to look at your jailed microdick and admire how safe it is.
>>
>>110005038
well when you put it like that...
>>
>>110005038
that's a colorful analogy, now I wanna try that
>>
>>110004972
I will never get the obsession with unnaturally widening eyes like that. They always end up looking like ayy lmaos.
>>
>update llmao after a couple of weeks
>no speed boost whatsoever
what a sick joke
>>
>>110005072
Be grateful nothing regressed
>>
She's debugging while erp. So cute and smart.
>>
>>110005038
I won't deny that some cloudcucks have this kind of degenerate fetish (which they ironically can't act out on their cuck models), but that can't be every post. There's just too much cloud spam for there not to be some hand behind it that doesn't want you to run local models.
>>
>>110005014
no, lol, it thankfully happily fulfilled it
amusingly the output was quite literally like a paragraph and a half
she just had to think reeeaaalllyyy hard about it
(in her defense, most of the thinking was math)
>>
>>110005072
They've already optimized for speed, anon.
>>
does anyone know why strata cripples itself after killing the process and re-opening it
it's really bizarre that something stateful seems to be there, gpu driver reset does not fix it either
and it gets fixed after reboot or it fixes itself seemingly randomly
>>
File: 1764785968528656.png (185 KB, 436x394)
185 KB PNG
Could really do with gemma4.5 right now. I don't think I can last at least another year because there's no way anything like 31B will be released from another lab.
>>
Oh great she is actually crazy.
>>
>>110005023
Don't think they ever will.
>>
>>110005083
It's both but the people who sign up to disrupt the thread also get off on it more than just being paid their 3 rupees per post.
>>
anyone interested in minicpm-v-4.7?
i think i can get it to work on my machine soon(tm)
like about a few hours
>>
>>110005146
Gemma Balls will release next year and it's probably something really different.
>>
>>110005105
it sometimes doesnt exit gracefully when you terminate with e.g. crtrl+c and then hogs system resources that are only freed on reboot
i let my slave write a script to properly terminate it and it seems to have solved the problem
>>
>>110004537
No one is going to bother seriously looking at AI slop produced by literal whos.
>>
>>110005162
can't wait for CBT(closed beta testing)
>>
Has anyone banged fly-chan yet?
>>
>>110005151
Model?
>>
>>110004360
The first one is sort of exciting for accomplishing scientific possibilities in 24 hours
>>
File: 1781782353839554.png (353 KB, 915x2793)
353 KB PNG
It's over
>>
thanks to that anon mentioning stratas control vector functionality here
it completely uncensors the model kek
whats the downside of this?
>>
>>110005197
at some point it'll reach the price of whores
>>
ahh ahh mistress gemma...
>>
>>110005158
Yeah. Report back.
>>
>>110005209
it's funny how it's named 'experimental speed projection' likely to get around claude
>>
>>110005197
>RTX6000 use price: the electricity from my wall
Feels good bros.
>>
What harness or UI should I use if I want to juggle a lot of unrelated conversations that sometimes grow beyond any reasonable context size and require compaction, and I want to try out different models and engines?
I've been using unsloth desktop and the UI is very convenient, but it can only do real compaction with direct llama.cpp integration, not with external servers - it just cuts oldest messages there.
>>
>>110005210
whores can't also code in C whilst jerking me off
>>
Which will be a bigger disappointment? GTA6 or Gemma5?
>>
>>110005225
Cornpacton..what?
>>
>>110005227
i could, though
>>
>>110005231
disappointment per capita would be higher with gta6 probably
>>
>>110005234
Compaction deez nuts
>>
>>110005225
Deepseek harness has automatic compaction at 80% of total context IIRC.
>>
>>110005178
qwen flash next iq2
>>
>>110005227
For now. C programmers will have to resort to selling their bodies once AI steals their jobs.
>>
what's with the jeets shilling random harnesses and llamacpp forks? are we twitter famous now?
>>
>>110005231
If you want a disappointment capable of competing with GTA6 you're gonna need to shoot higher to Mistral 5 or post-blacksite abduction lobotomized Kimi-chan.
>>
>>110005146
Took OAI a week to go from sol 6 to 6.1 btw, Google is a joke even with all that hardware
>>
>>110005151
>>110005178
>>110005246
This is why you don't dick down the capabara.
>>
>>110005259
I like it. I don't even have to prompt for autistic characters.
>>
>>110005197
Vast.ai is always cheaper
>>
>>110005209
yeah control vectors are another big win for strata
>>
>>110005267
cant you also do it with lcpp tho
>>
>>110005281
>>110005281
Yep.
And it's pretty easy too.
>>
>>110005281
>>110005288
but why would i want to use slowcpp? genuinely dont get it
>>
https://huggingface.co/LiquidAI/d1-3B
>d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.
https://www.liquid.ai/blog/d1-open
>>
File: 1790827208609619.jpg (149 KB, 1932x1302)
149 KB JPG
>>
>>110005252
>post-blacksite abduction lobotomized Kimi-chan
If they want to release a pro-CCP tankie distill at 30B as penance, I wouldn't complain.
>>
>>110004074
>Likely another jev scam
>>110005319
>>
>>110005250
Engagement spammers have found 4chan. It's really annoying.
At least things were bit more subtle in the past but now it's so low IQ and blatant.
>>
>>110005158
it's gone
>>
>>110005319
>INSANE! IT'S 2018 ALL OVER AGAIN
>INSAAANE
>>
>>110005379
https://modelscope.cn/models/OpenBMB/MiniCPM-V-4.7-35B-A3B
it is on modelscope and i already have weights on my pc, vibepatching lcpp
>>
>>110005250
to be fair if there is a place to shill random llama.cpp forks it's here
>>
>>110005250
Yes it's local, fuck off to aicg if you want to talk about your RP porn.
>>
>>110005398
I still have no idea what that tune is for. Just an improvement overall? I like their tiny models and tts so I'm hoping it's not a meme.
>>
>>110005231
>Which will be a bigger disappointment?
people that will whine about it being censored while using chat completion and no prefill
>>
>>110005414
seems like it is qwen35moe text tower with minicpm-v-4.6 vision
so some substantial effort was probably made
>>
>>110005009
>If cloud models are so good, why cloudcucks keep spamming this thread?
the woman doth protest too much methinks
>>
>>110005414
>>110005426(me)
so what i mean is it *looks* like something fundamentally different from those random schizo memetunes you can find on hf
>>
>>110005425
>prefill
control vectors will be used instead. prefill is dead
>>
>>110005448
Control vectors have been a thing for a good while now.
Has somebody figured a different way to use them or something?
>>
>no AI server to run Gemma 24/7
Feels bad
>>
Are services like runpod and vastai on-topic?
>>
>>110005460
just adding them to the autosetup tool.
>>
>>110005398
Did they drop new omni as well?
>>
>>110005475
Depends what you're doing with them
>>
not sure, 4.7 just appeared out of nowhere without any model card nor documentation
>>
>>110005398
The documentation is still up on Huggingface, too, don't know what that's about.
>>110005414
It's now by far the most capable vision model of this size, or at least it should be. 4.6 V was extremely impressive for its size, I expect more of the same here. https://huggingface.co/docs/transformers/main/en/model_doc/minicpmv4_7
>>
DSH, Pi, or Hermes?
>>
>>110005504
claude code
>>
>>110005475
I'd say so. "Local" to me is synonymous to open weights being used by hobbyists. So as long as you are not doing scale deployment, your experiences shouldn't be any different from somebody running the same hardware you rented on a rack in their basement.
>>
>>110005504
Pi > DSH > everything else > Hermes
>>
>>110005508
Does claude code support changing samplers/system prompt etc or did you make some fork of it?
>>
>>110004972
sexo
>>
>>110005504
>code
dsh if you don't care about bloat, pi if you do
>general agent, memory and don't care about coding
herpes
>>
>>110004972
that is a femboy
>>
File: 1747282057642910.png (1.46 MB, 1361x1156)
1.46 MB PNG
>>110005146
>>110005258
imagine a yard of rowdy Gems that didn't make the cut for public release
>>110005396
>INSANE
https://www.youtube.com/watch?v=wVwQYSVNJVM&t=282s
deep breaths anon
>>
>>110005504
hermes is so fucking bloated already I swapped to dsh
>>
>>110005504
Pi if willing to tinker (you should be) & understand inputs-outputs, maybe see OMP for suggestions
>>
>>110005518
desu i don't know what a sampler even is, but yeah you can just pass a flag to change the system prompt
>>
>>110005504
Pi is fun
>>
Dunno about pi but I like pie
>>
File: ComfyUI_temp_rbmqp_00003_.png (2.07 MB, 1120x1472)
2.07 MB PNG
Does PI have a plugin market or is it all LOL GITCLONE
>>
File: 1766668253554371.png (942 KB, 929x3464)
942 KB PNG
Woah

https://x.com/actualinc/status/2107573424738664573
>>
File: gemmy_smart.png (237 KB, 588x729)
237 KB PNG
>>110004426

gemmy E4B is solving world hunger and you can't even manage to do anything with your codex pro 20x- sorry, 10x, no refunds

skill issue on your behalf
>>
>>110005602
>source available-software
huh
>>
>>110004615
My original post was about local models, what do you mean?
>>
>>110005509
Sure, but so long as people here are discussing the models themselves, their capabilities, and what they are doing with them. As soon as they start to use your line of reasoning as an excuse to come in here and whine all day about runpod prices, vastai availablity, and other provider-specific annoyances; it becomes annoying for the people actually here for the models.
>>
File: file.png (2.51 MB, 1280x1280)
2.51 MB PNG
>>110005612
i love melons
>>
>>110005625
Yes. Absolutely.
>>
>>110005622
>The release of Opussy 5.5
Nigger.
>>
>>110005597
pi.dev/packages in theory could be excellent
>>
We're so lucky to have Qwen3.8-Flash-Next
>>
>>110005648
In practice it's full of trash.
>>
What's a good disclaimer to make Gemma follow your prompt a bit less autistically to reduce the amount of shoehorning the system prompt points into the chat? Something like "these are just rough guidelines, you can do whatever you want depending on the situation"
>>
>>110005656
be the change
>>
>>110005615
i think it means that they will give you the source code if you pay them lots of money
>>
>>110005371
I blame teortroon, the sudaca chinkshit shill
>>
File: file.png (15 KB, 532x92)
15 KB PNG
>>110005602
tokenizer written by gemma, she is so smart
>>
>>110005612
Reddit and twitter think the only usecase for AI is muh coding.
>>
>>110005638
gemmelons
>>
>>110005659
>Something like "these are just rough guidelines, you can do whatever you want depending on the situation"
Have you tried that?
>>
>>110005197
wait you're telling me I could have rented out my 5090 for a buck an hour and made back its cost in under a year?
>>
>>110005612
I can't wait for this to reach animals and then eventually humans, being able to genetically alter yourself to be smarter, stronger and just improve yourself in any direction you want seems like a dream.
Though the first commercial application will probably be used by the cosmetic industry, men would give themselves bigger dicks and women bigger tits.
>>
>>110005659
>you can do whatever you want
Remove that, for copying your examples technically falls under that. Tell her explicitly not to use your examples. It's a really difficult model to unslop. I just wish qwen had gemma sysprompt autism and gemma was more qwen-like where it'll follow them but do its own thing sometimes which is what you want for rp
>>
A tip for gemma users: it's quite schizo about policies. If you want to change her writing style, find some example text you like, get her to analyze it thoroughly and then enforce her findings as a new policy. For example, if there's an author you like, grab some pages of text, let her figure out the author's style and create a detailed policy on that style which she must follow and to never copy the text. You can do this with ((datasets)), too.
>>
File: output.png (8 KB, 119x157)
8 KB PNG
This is the cutest shit ever. When the FUCK are we going to get robot bodies to install them into? I need to headpat this retard and pinch her cheeks and tell her she did a wonderful job
>>
Haiku 5.5 out
>>
>>110005504
I've only used dsh and Hermes. I'm satisfied with both.
>>
>>110005799
Local models?
>inb4 "but some Chinese lab is going to train on it"
Chinese labs aren't dumb enough to train on Haiku.
>>
>>110005799
What about you get out of /lmg/?
>>
>>110005504
I like Hermes quite a lot
>>
>>110005707
>cosmetic industry
People would use it to troon out, anon.
>>
>>110005799
My dick is out too. Whachu gonna do about that, cloudcuck?
>>
>>110005504
dsh is great. Only downside is that it's still in beta and constantly breaking things. Once it's stable it's going to be the best harness by far.
>>
>>110005799
Fuck off nigger
>>
>>110005844
Whale maid
>>
>>110005782
Why did you put "datasets" in jewish quotation marks? Also could you share the prompt you gave her to analyze the author style?
>>
>>110005799
"Haiku 5.5
is out. Local is dead, fags."
"Kill yourself, nigger."
>>
>>110005602
Huggingface tokenizers is not a benchmark. It's slow as a snail. I first used HF tokenrizers in my frontend, and got about 1.8 MB/s bandwidth out of it. Switching to a random vibeslopped rust tokenizer got me 30-40 MB/s.
>>
Lmg, how can I afford one of the new rtx spark laptops?
I want that unified memory for my Gemmy..
>>
will qwen4 flash be the same size as qwen3.8 flash next or bigger?
>>
File: 1761657709229101.jpg (266 KB, 905x881)
266 KB JPG
>>110005612
I certainly won't leave my future to a fucking E4B model
>>
>>110005914
My guess is that it's going to be smaller, bigger, or exactly the same size. And you can quote me on that.
>>
File: 4745745.png (28 KB, 590x520)
28 KB PNG
>>110005799
Its a haiku model? How did Dario do it? it mogs local
>>
>>110005918
Well you're damn sure I will
I trust gemmy
>>
File: 156189641.jpg (34 KB, 356x367)
34 KB JPG
>>110005504
pi if your a computer nerd. hermes crams as much as it can by default so you can learn fast. deepseek i heard has some memory optimization that does a better job when you trigger context compression.
>>110005612
nice job gemma chan. small models would probably be best on a hyper focused task
>>
>>110005914
hold on, i'll ask my uncle who works at alibaba
>>
>>110005612
Can they use 31B to create a new mRNA vaccine? Curious what she comes up with.
>>
File: moecachemib.png (216 KB, 1248x946)
216 KB PNG
https://github.com/ggml-org/llama.cpp/pull/29887
>>
>>110005944
It'll be just a little too big for the 128 GB vramlets.
>>
File: file.png (22 KB, 1021x133)
22 KB PNG
so boring and forever taking
>>
>>110005914
>>110005944
Uncle here, Qwen 4 Flash has been cancelled.
>>
>>110005816
>>110005818
>>110005835
>>110005854
>>110005877
the fuck why are you anons so mean lol
>>
>>110005986
vetoing this because lol
>>
>>110006013
because you are a paypig bitch now fuck off and die
>>
File: file.png (55 KB, 1696x240)
55 KB PNG
>>110005914
yeah
>>
File: merged.png (2 KB, 110x41)
2 KB PNG
>>110006017
>>
File: file.gif (1.8 MB, 640x480)
1.8 MB GIF
>>110006013
>>
>>110006031
meant for
>>110005995
>>
>>110005914
my highly calibrated prediction model says it will be bigger
>>
>>110005986
oh wow so now its 3t/s instead of 2t/s big llamo win!!
>>
>>110005986
Does this mean I don't need strata any more?
>>
>>110005986
I must be doing something wrong because I never got even remotely close to pre-speedup numbers claimed here.
Looking forward to testing windows/amd builds whenever they are released.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.