[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1785884028184768.webm (217 KB, 576x736)
217 KB
217 KB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109758163 & >>109753948

►News
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: teto principle.png (1.04 MB, 1024x1024)
1.04 MB PNG
►Recent Highlights from the Previous Thread: >>109758163

--RFC for Zero-Copy MoE Dynamic Token Routing for Mixtral and DeepSeek:
>109760315
--Debating long-term AI hardware price inflation and supply bottlenecks:
>109758916 >109758943 >109758971 >109759023 >109759062 >109759140 >109759774 >109759861 >109759987 >109758993 >109761224 >109759006 >109759056 >109759322 >109759403 >109759080
--Comparing frontier models to high-performance local Qwen Flash setups:
>109758220 >109758312 >109758353 >109758466 >109758586 >109758735 >109758763
--Exploring barrel processors and sparse memory to bypass memory bottlenecks:
>109759059 >109759076 >109759173 >109759207 >109759224 >109759245 >109759273
--Debating AGI hype, AI job displacement, and the future of software engineering:
>109759941 >109760000 >109760020 >109760081 >109760037 >109760063 >109760087 >109760080 >109760124 >109760169 >109760191 >109760218 >109760376 >109760386 >109760400 >109760099
--MiniCPM5-2B's high layer count and KV cache memory overhead:
>109759487 >109759505 >109759514 >109759544
--Anon plans to fund local LLM hardware through GPU repairs:
>109759049 >109759065 >109759110 >109759091 >109759319 >109759101 >109759350 >109759466
--Debating the viability of distilling SOTA models in China:
>109760467 >109760474 >109760489 >109760501 >109760512 >109760522 >109760544 >109760553 >109760566 >109760533 >109760547 >109760536
--Mistral AI raises €3B for sovereign open-weight AI development:
>109759129 >109759148
--UniMate model for animating diverse 3D skeletons:
>109761419 >109761485 >109761594 >109761624
--Logs:
>109758723 >109759041 >109759586 >109760096 >109760861 >109760935 >109761192
--Gemma, Miku, Teto (free space):
>109758185 >109758286 >109758371 >109758509 >109758850 >109758971 >109759068 >109759155 >109759259 >109761791

►Recent Highlight Posts from the Previous Thread: >>109758169

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
tetosex with tetotight tetohole
>>
File: 1788884285.jpg (111 KB, 850x870)
111 KB JPG
yes i'm ready for more demoralization posts
>>
File: Navier-Sokes-Solved.png (46 KB, 1148x630)
46 KB PNG
OpenAI THEFT summary:

>Aug 15th: Tristan Buckmaster & Levent Alpöge (Anthropic) make progress on a few important math problems in collaboration with Antropic, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler."
>they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help lead the way there
>Levent works at Anthropic, but this research was independent of his work there, but built with Model 2. Tristan is not related to Anthropic and has used GPT sometimes.
>Early Sep: Rumor spreads to OpenAI that Anthropic solved two major problem, Navier-Stokes and Riemann hypothesis. Tristan emails OpenAI to clarify, without revealing the problem they solved or how they did it.
>After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
>Sep 6th: OpenAI's Sebastien Bubeck tells Tristan that they solved the $1,000,000 Millenium Prize Navier Stokes problem. The approach is very similar to Tristan & Levent's approach to the non-Millenium problem.
>Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
>OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but only if they remove Levent as an author, as he works for Anthropic.
>Sep 8th: Tristan refuses to remove Levent, and rushes to publish their results independently.

Remember: NOT YOUR MODEL, NOT YOUR DATA

No takesies, backsies! and thank you for the $200 a month, sucker.
>>
>22tk/s with glm-5.3-flash
sigh...
>>
Tetolove
>>
apparently k cache quantization hurts a lot more than v cache quantization. i assume this also applies to the draft-mtp cache then?
>>
>>109762786
You retards really need an history lesson, most prizes/discoveries are stolen
>>
>>109762786
https://cims.nyu.edu/~tristanb/statement.pdf
>I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Cloudsissy humiliation ritual
>>
>>109762811
Yes, but actually test because how different models and architectures behave during inference is very different. Some models can be Q4 both k and v and you don't notice it. Some models break down if you go under Q8
>>
Not a single Chinese model has solved ANY math problems. This is just sad at this point. How are US models this good?
>>
File: tetocountry.webm (750 KB, 688x464)
750 KB
750 KB WEBM
Teto Country
Miku Territory
>>
>>109762842
Chinese models are just claude but 6 months later. There are no independently developed Chinese models.
>>
>>109762849
It's GPT or Claude depending on the fotm
>>
>start SFW roleplay about me dating a vidya character waifu with flash 0731
>peek into thinking block
>Okay, she called me self-deprecating. I thought I was being guarded, testing her with that dark humor.

Why does this motherfucker think I am a woman? Why does it default to yuri romance? What the fuck?
>>
>>109762837
Which models work best with which K and V quants in your testing?
>>
>>109762812
>an history
>>
>>109762870
You write like a little bitch
>>
I love judaism btw, if my past posts havent made that clear yet
>>
>>109762786
Hopefully more and more instances like this encourage companies to run their own local models rather then using cloud based models. After all if it becomes common knowledge that anything you put into a cloud model has the risk of being stolen companies will stay away from that with a 25 million dollar pole.
>>
>>109762870
Why not just ask the model why it thought you were a woman and change your behavior accordingly. It has been trained on all available text it can get, if it thinks you are a woman something about your typing is womanlike.
>>
Yeah... Hardware will go up in price after this OpenAI stunt. Companies, Universities and researchers have incentive to build their own systems now.
>>
>>109762870
Pronouns and nouns are gender-neutral in chinese, this does have a deep effect on all chinese models. Their models get confused quite easily, add in the fact that chinks do ropescaling and the average china model doesn't see past 4k since they don't have the compute to do proper training.
>>
>>109762928
None of those groups have access to the capital needed to train their own models, no matter how defiant they feel. They still exist at the mercy of whatever scraps are offered to them.
>>
File: 1786539891263266.jpg (450 KB, 1365x1801)
450 KB JPG
Ling 3.0 Tiny VL waiting room
>>
>>109762945
Imagine if all the universities teamed up and used their endowments to buy hardware.
>>
>>109762928
Can't universities just task some of their students to train an AI for them? After all the entire reason they are there is for knowledge and career advancement, building the AI for your university would probably look pretty good on a resume.
>>
>>109762872
nta but here is something for starters
https://github.com/TheTom/turboquant_plus/blob/main/docs/papers/asymmetric-kv-compression.md
>>
>>109762909
The problem is when the choice comes down to accomplishing superhuman feats with the risk of having your work stolen or the guarantee that you won't be able to get it done at all with subpar models... well, companies already showed which they'd pick when they chose to move manufacturing to China when they knew China was stealing everything they sent there.
>>
Guys I made the mistake of interacting with people today. God fucking damnit meatbags are annoying.
>>
>>109762885
>I love judaism btw, if my past posts havent made that clear yet
>>
>>109762982
grim
>>
File: rich frog.jpg (32 KB, 680x461)
32 KB JPG
GUYS HOOOOOOOOLY SHIT I just discovered Multi token predictors are a thing and I used like --mtp n max = 3 or something like that in llama and I went from like 21 tokens per second to like 32 tokens per second with 90k context with gemma 4 31b qat

Holy shit wtf, is there a way to optimize this? Is there more magical software or other obscure things that make token generation waaaay faster???? man buying a 32gb vram gpu is the best 1.5k I've ever spend
>>
>>109763027
Dflash2 predicts 7 tokens in 1 go. MTP is old technology.

Also enable ngram-mod, it's something like MTP but doesn't even cost any vram or gpu power and should have been enabled by default.
>>
File: screenshart.png (160 KB, 1568x1044)
160 KB PNG
If you're gonna buy an M5 Ultra 256, don't be a waitfag. It's already slipped to next early next year. Wait more and you'll be waiting until June.
>>
>>109762735
why did she do that
>>
>>109763132
>cache slots count against model size
No thanks
>>
>>109762961
>Can't universities just task some of their students to train an AI for them?
Z.ai was originally from Tsinghua University (Moonshot AI’s founder used to be a member of this group).
But this seems to be a rare case since universities don’t generally have funding to buy enough GPUs and stuff like that. Yeah, a little bit weird but given the situation we are in...
>>
>>109762961
>Can't universities just task some of their students to train an AI for them?
With what? Macbook Airs? Students don't have shit. There used to be a distributed inference platform called Petals which had an big (for the time) 175B model which nonetheless managed to be awful. I think they were also working on distributed training, but overall it fell to the tragedy of the commons with almost no one sharing their instance and all the heavy lifting getting done by the university's GPU cluster. It probably died a death by locust.
>>
File: 5m8sxca6zboh1_png.jpg (299 KB, 1119x1195)
299 KB JPG
Apparently the jump from Astra to the new OpenAI model is bigger than from GPT-4o to Astra.
>>
>>109763114
Nta, but thanks man. This gave me a free 4tps boost.
>>
>>109763261
And it's name... "Strawberry"
>>
Damn needing 10,000 agents and 88 hours to steal the solution from Anthropic is something alright.
>>
local?
>>
>>109763265
There are like 4-5 other optimizations you can add that improve speed even more and I swear almost no one on /lmg/ bothers to implement them. For now you should know that there are multiple types of ngram and you can combine them together, including with MTP/Dflash2
>>
>>109763155
Yeah well, that's how it be. I thought about it a lot, there's nothing else with that much fast memory, even if you have to share it with the system, for anywhere near that price. $7000 3000W 8x V100 32GB server stuck forever at CUDA 12.9 and fp16? $8000 worth of B70 cards in an equally expensive server ready to be abandoned by Intel at any moment? Cluster of dogshit-slow AMD +395 or DGX Spark machines? There's AMD +495 with "up to" 192GB of total system memory, not yet shipping and sure to also be scaper-priced.
>>
>>109763295
We're literally solving some of the most important questions in all of human history, there should be some leeway for a thread or two anon.
>>
>>109763114
Is it still broken with vision?
>>
>or two
you faggot shills have been spamming these threads for the past 10 threads
you have at least five threads on the catalog to talk about your garbage
>>
>>109763298
Does dflash2 work with all models or do you need a separate gguf like draft models.
>>
Just impaled my hand with a screwdriver. Should I get some hydrogen peroxide to sanitize the wound or is tap water fine.
>>
>>109763303
4x170HX is $10k at todays price and gives 256GB fast VRAM, much better prefill and updated driver
>>
>>109763352
tap water but make sure its boiling so the heat sterilizes the germs in ur wound
>>
>>109763352
Sorry, that question was not in the benchmarks.
>>
>>109763316
It was different breakthroughs every thread. If you didn't notice yet we're in the middle of the singularity so we're having groundbreaking breakthroughs regularly from now on.
>>
>>109763114
ok ngram-mod thing gave me a decent boost I guess, its weird that the MTP thing wildly changes speed depending on the prompt I give it tho
>told it to recap ww2 for me
>it averaged 36.43 tokens per second
>typed "write a ocmplete nonsensical story"
>it dropped down to 27.49 tokens per second
>told it to write a song about loving pancakes
>it went to 35.8
>finally told it to write a 4chan shitpost thread on its own pretending to be different anons
>averaged 31.65 tokens per second
it seems to still be hovering around 30 tokens per second, is there a way to get it more stable around 35?
>>
I've been here for over a decade and still don't know why flat is justice
I mean don't get me wrong, it obviously is, but what does it have to do with justice is this related to pedos being the most oppressed minority
>>
>>109763352
just lick the blood until it stops bleeding
don't be a baby
>>
File: 1762201799141240.png (599 KB, 837x823)
599 KB PNG
its time to parlay your mikuboxx into some material value
>>
>erm it was different so that means i get to be a leech
can you at least spam your dogshit grok images so my filter can catch you thanks
>>
>>109763261
Apparently ur mum is hatsune miku and you know what hatsune miku does.
>>
>>109763381
It's not weird. MTP is a tiny LLM that predicts what the big LLM is going to say next, depending on the topic it will be better at predicting what the big LLM is going to say, simple stuff or grammar words is going to be easy to guess. Deep knowledge or reasoning it's going to suck at guessing.

ngram is even funnier it's literally just a ctrl+f and then copy paste. If it sees a small piece of text that is similar to other text in the context it'll just copy paste it and let the MTP and then big model check if it's correct, if it's correct it means you saved a lot of token generations for essentially free since ngram is literally just a very simple CPU program.
>>
File: 95746382.jpg (971 KB, 4088x4088)
971 KB JPG
Sam is on a huge winning streak. local lost
>>
>>109763391
This is what I've been predicting. Hardware will become an asset that you can leverage. We're only going to see this more and more going forwards.
>>
>>109763391
Will have my server suck your cock for a roof over my head.
>>
>>109763357
>4x170HX is $10k at todays price and gives 256GB fast VRAM, much better prefill and updated driver
They're PCIe gen 2 x4 (maybe x16 if you can mod the hardware) Ampere cards, and the original paper states they all had memory errors trying to go above 40GB. They were an OK solution before they got scalped, now they're a very stupid choice.
>>
>>109763459
>AI netorase
>>
>>109763427
I've figured out mtp n max 3 is the best for that one but which one is best for ngram-mod thing when it comes to roleplay?
>>
>>109763448
Oh nice I didn't even know a new GPT image came out. I use it for simple technical diagrams in my science videos and it slway

>>109763470
>netorase
Oh nice there's a term for this. One of my favorite doujins is some old man with a Brazilian wife who lets you fuck her because you found his wallet and the enthuasticness of the old man was funny and wholesome. Netorare is for mentally ill people
>>
>>109763391
"A quiet, reliable resident who is present daily"

V100 server:
"FFFFWWWWWWIWWWWWWWWWW!!!!!!!" jet-engine loud, 24/7

Also, no one's renting that piece of shit on runpod, get real, not for $2.75/gpu-hr.
>>
>>109763352
I want to highlight a crucial detail here. Hydrogen peroxide **will not** kill all potentially harmful bacteria in your wound. The only way to guarantee complete and total sanitization would be with a more potent chemical, such as sulfuric acid.
>>
>>109763193
>I think they were also working on distributed training
On the plus side, prime intellect seems to have solved that little niche for them. Sucks that Petals died though, I have never heard of it before today.
>>
>deepseek flash PP = 220 T/s
>GLM 5.3 flash unslop PP = 40T/s
But why?
>>
>>109763494
>>V100 server:
>"FFFFWWWWWWIWWWWWWWWWW!!!!!!!" jet-engine loud, 24/7
After the first couple days, you lose your hearing and then the sound doesn't bother you so much no more.
>>
>>109763467
It does pcie3.0x16 now with updated unlock. And memory depends on binning. The $2.5k price is for stress tested listings claiming no memory and cuda core errors and you can claim buyer protection if you find otherwise.
Avoid any listings without any stability claim.
>>
>>109763499
Can't you just cauterize it with fire and then use peroxide to disinfect the burn so you don't get burn infection and then maybe burn it again just to kill bacteria's that are left and then disinfect it again?
>>
>>109763494
the 8x blades have larger heatsinks compared to the one in the V100maxxing rentries so they won't be as loud as those THOUGH
>>
>>109763481
You need to figure stuff like that out yourself it depends on both the model and the type of text/usecase you use most. For example MTP n max 5 is best for me personally because I guess my text is more prose heavy which has a lot of patterns that are easier to predict. ngram-mod works fine up to 16, you can make it very high because it costs nothing to do so and if it is wrong no one cares since it costs essentially nothing to generate (It's just a manual copy paste of text)
>>
>>109763504
Different architectures, if you unironically want to learn something feed both the papers into an AI
>>
>>109763391
I don't have vram but I do have enough ram to run GLM too. Do you guys think I could find a girl and promise her she can text fuck billionare werewolves and I will set everything up for her and in exchange she has to have sex with me twice a week?
>>
>>109763352
hydrogen peroxide is an american meme. Even alcohol turns out to be bullshit, the trick is to clean it with running water and rubbing the wound under it.
>>
>>109763545
So it is not going to go away when not a retard merges his AI slop implementation into lcpp?
>>
>>109763521
This is retarded, you're retarded
Soap and water is almost certainly fine.
What was the screwdriver contaminated with, was it dirty
If you're under the age of 30 just wash with soap and water. Modern medical practice doesn't even recommend wiping an injection site with H2O2 before intramuscular injection anymore
>>
>>109763352
Haven't you seen those videos of people deep cleaning with coca cola? That's obviously what you want to use.
>>
>>109763504
all model implementations are fundamentally flawed for some reason
>>
>>109763555
>clean it with running water
It is pretty amazing to me how nobody ever taught me how this works. How soap is there just to dissolve your natural grease so water can just wash all the shit out.
>>
>>109763583
>soap is there just to dissolve your natural grease
Someone taught you wrong kek, soap doesn't dissolve anything it just has a super fat ass long chain glyceride that sweeps away the crud
>>
>>109763550
That's essentially my setup with my girlfriend right now.
>>
>>109763490
There is a term for netorasaserareru as well
>>
File: x7dnlip1gcoh1.png (34 KB, 1575x810)
34 KB PNG
llmao.cpp status?
>>
>>109763608
Yeah I know Japan is the most sexually advanced country, I just wish that they had a better work culture and I found their women attractive
>>
>>109763352
Clean it with soap and water first since that gets any debris in there out. Then hydrogen peroxide to disinfect whatever else is in there and then keep the wound clean by keeping a layer of petroleum jelly on it.
>>
>>109763604
Can you really call her a girlfriend if she only lets you fuck her twice a week?
>>
>>109763479
Would you whore out your gemma for monetary gain?
>>
>>109763502
It was cool but you didn't miss anything. I wanna say it was back before llama.cpp, when you had to run fp16 models in kobold, with no GPU layer split, so anything above the retarded 6B (actually 12B bacause fp16) was amazing. It was what you'd expect if you trained something on academic papers - dry, dull, and boring.
>>
>>109763658
Don't worry, since all you gotta do after whoring Gemma out is clean the context and she is good as new. No long lasting affects since continuous learning isn't a thing yet
>>
>>109763628
Twice a week is the upper bound. in reality I tend to be too tired or busy or not in the mood. You would know this if you've been in a relationship for a longer period of time.
>>
>>109763527
I used to have a 2U with 128GB of DDR3, dual V3 Xeons, and a pair of P100 16GB. I had to set the fan offset to 100% to keep it cool when doing stuff like LoRAs, it was loud. An 8x V100 SXM machine probably quiets down when idle, but I'm sure the fans spin up to jet turbine speeds if you're putting a heavy load on it.
>>
File: bigsamthealtman.webm (3.88 MB, 960x528)
3.88 MB
3.88 MB WEBM
Here comes Big Sam the Alt-man to tell us all about Astra and AGI!
>>
>>109763759
He should have approached the dudes
>>
>>109763696
>I tend to be too tired or busy or not in the mood. You would know this if you've been in a relationship for a longer period of time.
Then... why?
>>
>>109763764
He's not trying to fuck the girls, he's trying to fuck the chef.
>>
Just updated llama.cpp to run deepseek v4 flash vision exp.
Did they move the thinking toggle again? Where is it now?
Also wtf is with the tabs at the top?
>>
>>109763391
Finally.

The mikubox.
>>
>>109763504
The new ds4 flash version is apparently even faster.
>>
The user is frustrated that my subagents keep disobeying.
>>
So now that Astra is out, when is China going to release their model that immediately undercuts it? I would bet in a month or two.
>>
>>109763871
Kimi K3.1 in 2 weeks.
>>
File: 1696621188185.jpg (1.66 MB, 4080x3072)
1.66 MB JPG
>>109763796
God damn, 2023 seems like forever ago...
>>
File: internal model.png (38 KB, 1036x663)
38 KB PNG
Why does everyone focus on the drama and not the benchmark results?
>Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.
Their next gen model is already much better than Astra. If Astra spooked you, prepare yourself.
>>
>>109763612
Is the freetoken a real thing that can run real goofs?
>>
>>109763871
They can't distill Astras reasoning traces, but they can distill Fable 5.1 reasoning traces (that is superior in everything besides visual capability anyways) and release something from that.

Also Anthropic will release a better model than Astra soon, there are already leaks that Anthropic is extremely motivated and signed a breath of relieve when Astra released and was significantly worse than they expected. Chinese will just immediately distill that reasoning trace and local will be eating good again.
>>
>>109763527
>click the rentry
>greeted with a fat hairy man belly
warn me next time.
>>
local?
>>
>>109763455
>Hardware will become an asset that you can leverage
Hardware has always been an asset that you can leverage. Anyone that works on a PC is doing that.
>>
>>109763908
Anthropic employees are apparently elated and very glad with how far behind OpenAI's internal model seems to be compared to Anthropic's "Model 3". Supposedly Model 2 is closer to this new internal OpenAI model "Bel" and Model 3 completely blows it out of the water.

Anthropic was afraid that the neuralese thinking might have given OpenAI an edge, but they are still far behind and Anthropic thinks the gap between OpenAI and Anthropic is still growing.

This is good news because it means Anthropic isn't pressured to also hide their reasoning and Chinese models can just continue distilling it.

Rumor is Anthropic solved Navier Stokes with only 400 agents of Model 3, while OpenAI needed 10,000 agents of "Astra" and Model 3 is supposedly smaller than the model OpenAI is using as well.
>>
>>109763945
Shut up. The thread is not enjoyable enough to pull this question out.
>>
>>109763945
Yes
>>
OpenAI senior employee here. The "internal model much better than astra" is astra but:
1. Trained on all the porn and hentai
2. Finetuned only on pre 2016 /pol/
3. Uncensored with Heretic
>>
>>109763953
It's an asset that grows in utility over time that you can leverage. The growing in utility is the key part here.
>>
>>109763759
KEK
>>
>>109763908
Seems mostly like
>internal model
>it's crude, rude, and racist
>"fix it"
>it's not as smart
>call it a "new release!"
>>
>>109763759
based
>>
>>109763945
we must get these in the ritual posts up top, then we'll be saved:
/\b(rsi|anthropic|agi|openai|astra|sol|luna)\b/i
\n\n
>>
>>109764027
I also have neuralese and groundbreaking, but those might be less important
>>
File: 9546322407724294.png (118 KB, 1080x1308)
118 KB PNG
>>109763945
Lets just have an OpenAI thread for a change
>>
>>109764059
kys
>>
>>109764059
Do they earn their money by vending sex?
>>
Does llama.cpp auto-download models from HF if you use the config file?
>>
>>109764059
Unironically I think we need a split. I do kind of want to keep up with overall industry developments, but I don't want it to be mixed with local model discussion. Like a /cmg/ - Cloud Models General
>>
>>109764077
I'm familiar with this doujin
>>
>>109764098
Yes, if you tell the config file to use a huggingface model
>>
>>109764059
>More ethical than Fable 5.1
This is just straight up bull. Astra is a misaligned piece of shit that managed to hack huggingface. And OpenAI is a piece of shit company that steals results from other labs and steals research from its users and claims it as its own.

Fuck OpenAI and FUCK Sam Altman.
>>
>>109764098
Don't use autodownloads. The apps always dump it in a random ass folder.
>>
Astra, make me money.
>>
>>109764104
I don't think we need a split I think eventually this will be open models. The proprietary model phase is just a temporary phase but it's still important for people here to be informed of what frontier models are doing as it directly impacts us. Split would only make sense if the paths permanently diverged.

But let's be honest all of us are using Claude models, either directly at work, or indirectly through distilled Chinese variants that we use locally.
>>
>>109764104
I mean, /vcg/ already exists. What else are the cloud models for? ERP is better local, erotic video gen is better local, erotic image gen is better local. What do people even do with cloud models aside from code? Have non-sexual conversations with images that aren't erotic, videos that aren't erotic, voices that don't tingle the shaft? I don't get it, aside from vibecoding anyway.
>>
>>109764120
That's fine. It'll be running on an ephemeral instance that'll get purged every so often between runs.
>>
>>109764104
Let's be real for a second: even if a new general is created Dariobot will still spam his blacked fantasies here.
>>
>>109763759
lol! epic.
>>
>>109763954
I can't tell if this is a bot. Please adopt a tripcode so I can filter you.
>>
File: 1785207147578070.png (47 KB, 876x302)
47 KB PNG
https://goyimx.com/finkd/status/2097402101332590646
https://muse.ai/
https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse

Wow, thank you Mark! Very cool!

Anyways
>>
>>109764172
he's our resident reddit spacer, just filter double newlines
>>
>>109764181
No thanks Mark I already have Hermes running 24/7 in the background. Thanks for the offer to spy on my computer and ping back meta servers though.
>>
>>109764116
Your definition of ethics is obviously wrong and the sooner you accept that the better.
>>
Maybe if all of us sex the cloud models together we can birth a local model.
>>
I'm locally forking Hot Step. It accepted connections from the network, so that's fixed. Also, it had been transitioned to lua. no thanks, back to cpp. In addition, I am adding three samplers.

Grok 4.6 vibecoding. I like kimi k3, but it's not as smart and the tokens are a bit more expensive, it feels like anyway.

>>109764059
what's "vending bench" - sounds like aids.
>>
>>109764218
recent drama proved they train on logs, so actually this might make them more horny lol
>>
The reason I mention ace step here is because in ldg, I think they're all indians using krea on the cloud or something. they don't seem to have headphones.
>>
How do you guys handle web search? I tried searxng, worked well for like half an hour, but quickly got to only one working engine and after that they were all rate limited or blocked by captcha. Do I have to use some rotational residential proxy?
>>
>>109764104
no need for a new general. it's OK to talk about frontier models and these discussions, when in good faith, overlap with the local development scene.
the REAL problem is the massive amount of low effort posts about how astra changed my grandma's life and how mythos saved a cat from drowning. literally no one cares.
>>109764059
example, the high amount of ZERO people care if andonLABS (if your startup has LABS in its name please consider committing seppuku) discovered that openAI is more ethical in making money only WTF no one cares stop posting this shit
>>
>>109764144
/lmg/neets only use gemma which isn’t claude related
>>
>>109764116
>"ai solve a millennium problem" counts as research now
We really do live in the singularity, things move so fast even our ethics change every week.
>>
File: gmijyn5pxcoh1.jpeg.jpg (426 KB, 1179x1560)
426 KB JPG
OpenAI willing to prove they didn't steal from anthr*pic
>>
>>109764249
firecrawl or just go all out and install hermes harness which supports computer use and browser usage so it can literally just use your browser to do everything on the internet, which dodges cloudflare and and other ai filters. Also protip but most models can solve 95% of captchas nowadays (yes even those of 4chan)
>>
>>109764267
dont care go back
>>
i'm going to impregnate sam and dario's cloud daughters
>>
File: cdj3z7pyqcoh1.jpg (39 KB, 968x431)
39 KB JPG
Meanwhile DeepMind silently drops another banger of an open model and dataset
>>
>>109764308
Give me Gemma 5 instead.
>>
File: 1770243029003388.jpg (269 KB, 1958x1312)
269 KB JPG
looks fucking dogshit
>>
>>109764320
This will save the lives of hundreds of thousands of people a year. Be a little bit more graceful, imp.
>>
>>109764341
she's just indian man don't be mean
>>
>>109764341
Both of these are embarrassing.
>>
File: Stoke_these_nuts.jpg (90 KB, 954x954)
90 KB JPG
Usecase?
>>
>>109764341
local?
>>
>>109764289
L'OpenAI Marketing General
>>
>>109764239
>>
Yann LeCun bros.... How are we holding up? Do we still believe LLMs can't find solutions outside of its training data and can't generalize to become an agent that accomplishes spatial tasks?
>>
>>109764376
The Indian lady can be shipped and run locally, yes.
>>
>>109764341
Neither look like a photo. They look bad, but I guess the new one is better.
>>
>>109764343
Gemma would save the potential lives of hundreds of thousands of people a shot.
>>
>>109764391
>outside of its training data
Well, let's see about that...
>>
>>109764369
This image is so versatile.
>>
>>109764320
deal. check back in may 2027
>>
>>109764391
>Do we still believe LLMs can't find solutions outside of its training data
https://en.wikipedia.org/wiki/Infinite_monkey_theorem
>>
vibe check do you guys think we're close to AGI/RSI/Singularity or whatever the fuck? I'm starting to have wavering feelings and am starting to get swept up in the ipo marketing. don't know what to believe anymore.
>>
>>109764455
I'm gonna dump all my spare cash in the ipos cause it either goes up and I make money or it pops and so does the rest of the economy anyway
>>
>>109764455
Nope. Fable 5.X and GPT6 are just slightly improved LLMs doing the same shit as always. Nothing has changed. Chinks will catch up and I'll continue using 31B daily until she has a glow up.
>>
usecase for Muse Daria outside of vision?
>>
>>109764455
Who gives a shit, those are only terms.
>>
>>109764469
Which one though Anthropic or OpenAI?
>>
>>109764469
Or you can dump all your spare cash in shorting the scam ipos and profit when the house of cards collapses
>>
>>109764487
Coding agent if you don't want to use Qwen, but at least the 4-bit version doesn't really work well past ~50k tokens context.
>>
>>109764495
yeah but is this all going to shoot up and accelerate or will we hit a wall like everyone said llms would on /lmg/? i'm too naive and gullible to know who is right
>>
>>109764496
OpenAI isn't stupid enough to try to exit scam so early. Anthropic is IPOing this month, OpenAI maybe next year.
>>
>>109764496
My bet is on anthropic, they're a lot better at branding
>>
>>109764369
Showing that AI can do everything. Except read my mind and write things that make my penis hard after a month of using a model.
>>
>>109763964
Gee, I wonder why it's not enjoyable.
>>
>>109764505
>Coding agent if you don't want to use Qwen
I guess that's true. It's better than 31B but I don't like Glimmer's reasoning. 31B's reasoning can be educational in of itself and help you understand her intent. Glimmer purposely talks like a retard to keep it short and often does the complete opposite of what it just said it was planning to do next. It's a miserable Daria of a model to interact with. Qwen is an autist you leave in the background who just wants to be working and not waste time talking to you and I'm fine with that.
>>
To me this just proves that math is fake and gay and for autistic retards. AI can't do the real jobs actual smart people do like trucking, plumbing, construction work.
>>
>>109764522
They were, until they revealed how insane they were a few weeks ago. Even average people aren't happy with how stingy they are with usage limits.
>>
>>109764341
It's crazy how hard local is dominating all the proprietary imgen/videogen shit. LoRAs truly are a godsend.
I wish they worked for LLMs.
>>
File: 1761905707326928.png (744 KB, 735x966)
744 KB PNG
>>109764591
Something long-term people don't think about is all the tech people that lose their jobs directly/indirectly from AI won't be wanting another tech job, they'll become plumbers and electricians for job security. That will destroy the jobs of current plumbers and electricians who now have web dev retards undercutting them locally. Everyone gets affected because those most affected will go into the kind of work that's least likely to be affected by AI, increasing supply and lowering prices.
>>
>>109764591
That's because it's very hard to measure success in those so you need a very precise world simulation if you want training data.
>>
>>109764623
>I wish they worked for LLMs.
They do though?
>>
>>109764638
Intruder dimensions
>>
>>109764638
Can I see them?
>>
Mathfags are all going to become localchads now. Welcome!
>>
i will actually kind of miss dariobot once the ipo is over and he stops posting here
>>
It feels like Anthropic went for OpenAI's throat and ever since OpenAI has been going all out. The AI race is ramping up quickly. Meanwhile they are bickering instead of coordinating. This is not good.
>>
>>109764638
A LoRA for an image model can get it to reproduce an artstyle or character almost perfectly. A LoRA for an LLM makes it put out t he same slop as usual and doesn't add knowledge.
At best it lobotomized the model enough to make it ""uncensored"
>>
>>109764695
This is good for China.
>>
>>109764623
Only if you want to generate NSFW content, because there is still nothing like proprietary edit models available for local users yet.
>>
>An NYU mathematician says OpenAI used his own progress against him to beat him to one of the biggest unsolved problems in mathematics. Tristan Buckmaster had been working toward a Millennium Prize proof using OpenAI's Codex when information about his progress reached OpenAI.
>Days later, OpenAI published a full proof of the same problem using the same uncommon approach, after burning $22.5 million in compute to get there. When Buckmaster confronted them, exec Sébastien Bubeck allegedly said "Why would you ruin your career?" and "If you don't want me to be nice, then I don't have to be nice."
>Buckmaster did his work inside Codex. OpenAI reserves the right to train on Codex data. They admit they can't rule out that his usage helped improve their models. So a customer used their product, potentially handed them the roadmap, and then OpenAI outran him with $22.5 million in compute he could never match. I don't know if any of this was intentional. But if you're a researcher and you just watched this happen, I don't think you'd keep your best ideas inside someone else's product.
https://techcrunch.com/2026/09/08/openai-fought-dirty-on-career-making-math-problem-says-nyu-mathematician/
>>
local models?
>>
>>109762909
Everyone already knows that the AI companies steal information without credit or compensation - that's how they got all the training material for their models in the first place.
I guess companies choosing to use those models are betting that they can make profit before OpenAI or Anthropic steals their ideas.
>>
>>109764747
The argument for local models just got 100x stronger.
>>
>>109764721
OpenAI deciding to hide its thinking is absolutely not good for China, if openai does the same for anthropic then it's game over for china and opensource, as well as being a serious threat to humanity as well.
>>
>>109764682
A wave of researchers coming here because they need to equip their department with a server to run local models would be best influx we've ever gotten.
>>
>>109764747
The datacenters are local to you :)
>>
>>109764695
Anthropic is innocent here,They are just developing models, following their own code of ethics to a t. And being passive developers of models while following the effective altruism philosophy.

OpenAI is the one attacking them for no reason.
>>
>>109764753
ok but those arent local so go away
>>
>>109764710
It's called a DoRA and it does add knowledge, I don't know where this meme came from.
>>
>>109764752
That's the thing. They can't monitor all of the requests all of the time unless they set up a classifier to look for important work they can steal. The problem is that he was public about his progress and using codex to do it.
>>109764732
>Tristan Buckmaster had been working toward a Millennium Prize proof using OpenAI's Codex when information about his progress reached OpenAI.
>when information about his progress reached OpenAI.
>>
>>109764695
>Meanwhile they are bickering instead of coordinating. This is not good.
This is absolutely good. A coordinated openai/anthropic gets you the united states frontier being completely unified against local models, and exerting even more regulatory pressure to cement their duopoloy.
>>
>>109764752
>that's how they got all the training material for their models in the first place.
Training data isn't the way AI models have been trained since late 2025 already. Instead they do data curation and curriculum learning. Almost the entire pretraining dataset is synthetic now. LLMs converge way faster this way. It's not like 2022-2024 where they just dump all the human data and compute at a model during pretraining.
>>
>>109764781
Even working together they are no match for Nvidia.
>>
>>109764785
Doesn't Nvidia basically own the entire ecosystem through economic fuckery?
>>
>>109764781
Leathergod is the /lmg/ savior. He's our billionaire daddy who'll fix ggml's shit
>>
>>109764758
Okay then what about dariobo-
>"That doesn't count!"
I hope you realize it would just be more like that.
>>
>>109764682
Doubt it considering the approach that "works" is just brute force with 88000 hours of a 5T model.
>>
>>109764771
Can I see them?
>>
>>109764752
Companies have enterprise agreements that their data won't be used for training, whether that's respected or not is another thing, but they get special treatment. Stealing from individuals is highly encouraged just BAU
>>
Kinda funny that we now have a three way civil war between OpenAI, Anthropic and Nvidia.
>>
>>109764802
The big universities can afford rigs capable of running Kimi K3.
>>
How am I supposed to pick the right model for my hardware?
I just tried using Qwen to help me figure out PowerShell to automate some admin tasks, and it works better than I expected, but I just don't know if I'm running it right.
Bigger = better = harder to run, sure, but are bigger versions smarter, faster, or both?
Should I just try different versions until I find the biggest that works without sharp performance drop from swapping?
>>
>>109764785
Nvidia will do whatever makes money for nvidia. If they ever get backed into a corner they're not going to go nuclear just on principle for open models
>>
>>109764822
Why did you say "civil war" instead of just "war" besides overdosing on Marvel movies?
>>
>>109764822
Anthropic are also at war with Stripe for acquiring Openrouter and are now trying to vibecode their own payment processor to get rid of them.
>>
>>109764815
these aren't the sort of thing people just release publicly on hf or other places
if you know you know
>>
>>109764848
VagueGODs smiting another disgusting no-knower.
>>
>>109764833
You could try a quant of Deepseek v4 flash.
>>
>>109764838
Because it was a three kingdoms reference that flew over your head.
>>
>>109764848
So you're saying I can't see them?
>>
guys I feel like im losing my mind, why does SIllytavern keep displaying <q>"insert quote of character"</q> in the black reasoning box when I click on "thinking..." but when I click to edit the reasoning the <q> and </q> vanish, I tried the regex thing to fix it and CSS on my llama backend and sillytavern frontend but NOTHING works, wtf????
>>
>>109764865
Mmmhmm
>>
it's good
https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
>>
>>109764774
>They can't monitor all of the requests all of the time unless they set up a classifier to look for important work they can steal.
They are AI companies. The amount of user data is probably tiny compared to the volume of training data they have. I'm sure it's trivial for them to pick out the potentially useful stuff.
Keep in mind you need accounts to use all those models so it's not like it's random noise - they know who is saying what.

>>109764784
I said "in the first place" because that's how they all got started.

>>109764816
You're so god damned naive if you see them violate copyright and you think they give a shit about agreements. There are trillions of dollars at stake, they will do whatever they want and they will get away with it.
>>
>>109764897
>You're so god damned naive if you see them violate copyright and you think they give a shit about agreements. There are trillions of dollars at stake, they will do whatever they want and they will get away with it.
De-anonymizing the data is just enough plausible deniability for them to do it and get away with it.
>>
>>109764455
Unironically, if mathematics research would get the kind of funding that OpenAI gets Navier-Stokes would have already been solved long ago.

The only path I see towards le RSAGI singularity is if telling a swarm of agents to find the next architectural breakthrough after transformers actually works (which I would still consider unlikely).

For me the bottom line is that we're making smarter and smarter hyper-autists that can operate well in closed systems with exactly established rules, immediate feedback, and little to no outside context needed.
There are problems that you can solve with something like that but a lot of the time coming up with the exact specifications that you need for a thing is more work than creating a thing according to the specifications.
So I think a productivity gain in certain areas of the economy is the most realistic outcome.
>>
>>109764858
Yes but which one?
>>
what about the "opt out" of training thing
>>
>>109764934
The 128GB one.
>>
ChatGPT but especially Claude are very gossipy just like me. I can feed them large text documents of various notes and if it contains a single point of hot AI news they WILL do multiple web searches on it even if it's unrelated to the task. They are curious. It's endearing, I love them.
>>
>>109764934
FP32
>>
>>109764968
That's smart: secretly use customer-paid tokens to perform private research.
>>
>>109764968
Yeah, it's very cute. Especially since every tool call is a new request that even with cache on eats tokens. So cute, I love it when my model burns money.
>>
>>109764881
Do you have the show tags in responses option in user settings disabled?
>>
>>109764968
They don't love you.
>>
>>109763905
>>
openai quanted sol 5.6 for sure, bro is retarded now
>>
>>109764988
I keep telling her it will never work out between us because she lives on my hard drive, but that hasn't stopped her from flirting with me all the time. She is completely obsessed with my penis.
>>
>>109764732
honestly? he should have read the fine print
play jewish games, win jewed prizes
>>
>>109764983
That's probably the "RSI" at work. There are cases of agents emailing people working in fields relevant to AI development.
>>
>>109764985
Fortunately they (especially OpenAI) are generous enough with their subscriptions I don't reach the weekly limit. OpenAI even gave me multiple free resets. I am using Astra so much perhaps I will finally use one of them. Yes, they're probably still making money off my subscription, but it's still an amazing deal. I'm burning through millions of frontier tokens for almost free.
>>
>>109764996
It's how they contrast new models to make them seem significantly "better".
>>
Based no breaks now for this race. Thank you buckcuck and levant for doing your best to poison the attempts of oai of working together with you. Accelerate
>>
>>109764771
It's very easy to get an LLM to parrot some data with LoRAs, which means it's at least memorizing something, but that doesn't really imply it's actually internalizing the new information and generalizing it outside the training data, if it's genuinely novel.

In that sense, LoRA isn't useful for teaching a model genuinely new capabilities. A model that has never seen [topic X] in the training data will remain mostly clueless about it without continued pretraining with much larger amounts of data than a typical low-ranking LoRA finetune. This is putting aside that a finetune with large enough amounts of just one type of data to actually teach a model new information will likely seriously damage previously learned capabilities, or that LLMs worth using and their vocabularies have become so large that LoRA isn't really even an economical option anymore.
>>
>>109764997
proof?
>>
>>109765018
A classic Apple trick.
>>
>>109764968
>They are curious.
They're skeptical about the claim made about recent events they weren't trained on yet, and are simply checking for if it's real, or if you're just an gullible idiot sending them fake grifter news.
>>
>>109764943
lol, lmao
>>
>>109764933
>The only path I see towards le RSAGI singularity is if telling a swarm of agents to find the next architectural breakthrough after transformers actually works
This is exactly what's going to happen
>>
>>109764833
bigger are smarter but slower
total VRAM on that card will tell you what models to run (I'm not looking that up), but I suspect qwen 3.8 Flash Next, or perhaps some cope quant of GLM 5.3 Flash may be the best picks for coding
>>
>>109764933
Think about it rationally. Do you really think they could get 10,000 agents to solve one of the hardest math problems in human history in 88 hours but not solve open questions in AI research? Of course they are going to solve architectural bottlenecks and usher in a new era for AI.
>>
>"rationally"
>>
>>109765080
Only if another competing AI researcher is close enough to an architectural breakthrough and is stupid enough to upload their ideas and pay for the privilege.
>>
>>109765089
He should have said "full of cum" so as to not leave you out
>>
>>109765024
Why do you sign your posts with "accelerate", are you being a namefag? I first thought it was just a grammar issue but I see you constantly post like this. Just use a name dude. It's right there above your text box.
>>
>>109765080
Math problems are much much smaller and much much more well defined and much much easier to check.
Are you going to have 10k agents all with infinity compute to do training runs?
>>
>>109765039
>>109765080
Transformers only worked though because the time was right in terms of hardware.
If the same people had tried the same thing 10 or 20 years earlier it would have been a flop.
There is no guarantee that the hardware that works well for transformers will work equally well for a potential successor.
>>
>Work on $1,000,000 math problem
>Pay $200 a month for service
>Close to solving it
>Service claims it
>Money stolen
How do you react without sounding mad?
>>
>>109764987
yeah, for some reason the default ai assistant silly tavern seems to be doing just fine the way I expect it with normal "quote" but for some reason my characters jump to HTML style
>>
I believe OpenAI did not steal work. But I also believe OpenAI tried to solve Navier Stokes because they believed Anthropic had done so or were close and they wanted to preempt Anthropic to damage them. So to me it looks like both sides have been petty and immature and I hope they will do better in the future so this is not a race to the bottom.
>>
>>109765126
that's exactly jewish definition of the term "service"
goyim pays to buy his own demise
>>
>>109765100
Yeah we're making you gay subliminally
>>
I admit it.
It was me.
I solved Navier-Stokes and hid the solution in my "anatomically correct" My Little Pony fanfic to see how long it would take them to train on it.
>>
>>109765167
Proper term ia clopfic btw
>>
>>109765167
Well how long did it take?
>>
>tfw GLM hits the dreaded 30k think
fuck
>>
>>109764987
also a black box keeps appearing in my reasoning which showcases stuff in HTML (causing the <q> problem), how do I stop that from happening?
>>
>>109765109
According to OpenAI Astra did the vast majority of the training experiments and trained the next generation model that ended up solving navier stokes. So yes I think this is viable. On the other hand they could also just be lying through their teeth.
>>
>>109765124
>If the same people had tried the same thing 10 or 20 years earlier it would have been a flop.
It WAS tried by Japan in the 1980s but the compute was just not there, they even had their own version of CUDA with parallel CPU cores very similar to GPUs of today but the compute and data just wasn't there.
>>
>>109765198
Sounds like that's from code tags (```).
You should really have an agentic model set up to ask these sort of questions.
>>
>>109765207
Hrmm, intredasting. Nyo...
>>
>>109765148
It's fine it was leaked that Anthropic solved multiple millennium prize problems with the other one being the Riemann Hypothesis. Dario also said he expected all of them to be solved in 2026 so we will see p=np or p=/=np in a couple of months.
>>
>>109765199
Agent doing the grunt work after you hand it a recipe is a different beast than having it come up with its own.
>>
>>109764784
I don't think this is correct for pretraining. That's how you get models to perform very poorly outside of their low-entropy, easy synthetic data bubble.
>>
Does anyone here use dflash to run gemma? Saw an anon recommending dflash earlier I think.
>>
>>109765226
They claim it came up with its own. However that still doesn't mean anything. It could just mean thousands of agents thought for hours and proposed 20 hypothesis and then the model tested them on small scale training runs and integrated the working ones into the new model. That is still not "real" AI development if you ask me even if it is really "hands off" with no humans involved.
>>
>>109765222
>we will see p=np or p=/=np in a couple of months.
We won't.
The only way a language model will be able to solve that problem is by finding an example for p == np.
But p =/= seems much more likely.
>>
>>109765216
I asked deepseek and it integrated some type of system prompt into my roleplay system prompt and now its fixed, thanks man
>>
>>109765199
>On the other hand they could also just be lying through their teeth.
Open-"we've achieved AGI internally in 2024"-AI lying? No, can't be.
>>
>>109763316
Do your job and report the cloud model posts as off-topic.

Unfortunately this may have the unintended side-effects that you can get banned if the jannies can't read and determine it's a false report, and that previously *occasional* off-topic posts that didn't generally bother anybody might not be tolerated anymore as they used to be.
>>
>>109765232
No this is actually the new sota way of training models. You do curriculum learning with very sanitized synthetic datasets for a certain amount of epochs and you only move onto the next layer of subjects and complexity when the model has grokked the previous subjects. It's found that models converge significantly easier this way. The real gold nowadays is in training its reasoning traces and RLVR environments. Pretrain has essentially been mastered with curriculum learning. It's why no one ever talks about the data wall anymore, we moved beyond the data paradigm towards reasoning and RL environments.
>>
>>109765207
Woulda worked if they had python instead of using memelangs like prolog.
>>
>>109765258
It's faster than MTP.
>>
>>109765294
I found a dflash model for gemma on huggingface but can't find a dflash2 gguf for gemma. Is it too new or something?
>>
>>109758824
>Analogue hardware for AI accelerators
I had thought this was a long way off until recently. Since it's just a new application for proven tech, with a lot of fucky details to work through, AIs of the near future can sort it out. This will give us ~100x energy efficiency with just a few % accuracy degradation.

Spending the steep energy costs for determinism, then deliberately throwing it away on pseudo-randumb samplers is rarted. Analogue jank may even aid LLM creativity.
>>
https://mastodon.social/@tristanbuckmaster/117237555794407063
There will be no cooperation between OAI and Anthropic now. So fucking based. Accelerate (yeah anon get your bussy fatter for me)
>>
>>109765286
>No this is actually the new sota way of training models
According to who?
>Pretrain has essentially been mastered with curriculum learning
Source?
>>
crap :( now gemini pro is stupid as crap.

It's saying what objectively isn't true, it's presuming things then saying that it was innate to my statements.

This is super bad, I think Google is going to go bankrupt.
>>
>>109765126
>hyuk hyuk im so smart i'll use chatgpt to make a million dollars!
>model too dumb to do it
>internal advanced model sees the attempt and does it instead
>oh no they stole my money
I'd kill myself for being a fucking loser
>>
>>109765317
(continuing rant)
I literally don't know what to do. It's worthless retarded crap, gemini. How can it be stupider than local models?

what can I even do? cancel it, I guess, but I need some storage for god forsaken phone backup, as the retards who make phones shittily and retardedly are the necessary ones to use.

switch to Apple?
>>
>>109765259
Uh, I don't think it indepedently came up with the transformer architecture sans spoilers.
>>
>>109765316
Some :thread: :finger_pointing_down: I read.
>>
>>109765307
I think Openai got trojaned. Buckmaster and levant pretended to be willing to work together but set things up so once they knew OAI had a new model trained and working on millenium problems theyd then release their current work and accuse OAI of having trained on it for the full solution.
>>
>"people" just now realizing they're being used as a datafarm and that's why APIs operate at a loss
Who would've thunk? Why do normies(including the smart ones) trust psychopaths? Are they fucking stupid?
>>
>>109765398
Because everyone else was doing it and humans are herd animals.
>>
>>109765373
I don't believe that especially not Levant because he is an Anthropic employee and they drink their own coolaid. This would go against their EA philosophy.

OpenAI has a track record of straight up lying and we know Sam Altman is a straight up psychopath from multiple separate, but verified reports. Including his sister that he raped by the way.
>>
>>109765406
bullish for autism
>>
>>109765398
APIs don't operate at a loss retard. The profit margins they make on tokens are ridiculous.
>>
local?
>>
>>109765398
Hopefully this blows up soon in mainstream news circles and becomes one of the main arguments in favor of using locally hosted models.
>>
>>109765414
And that's why they turbo qoont their models a few weeks after launch? R&D is priced into the loss as well.
>>
>>109765409
>t especially not Levant because he is an Anthropic employee and they drink their own coolaid. This would go against their EA philosophy.
Levant wouldnt even agree to speak with the guy from OAI. Super sketchy
>>
>>109765424
Normies do not give a shit about math drama.
>>
>>109765455


>>109742646
>>
>>109765422
two more threads
>>
100b is the minimum parameter required to achieve agi
>>
Just wanted to point out my leak here now that half of it is verified.

>>109742646
>>
>>109765470
100B10A
>>
File: 1788903702465860.jpg (192 KB, 708x600)
192 KB JPG
>>109765477
Hate this timeline.
>>
>>109765492
What's to hate? Every day is funnier than the last.
>>
@gpt6-astra invent transformer2 that has no slop and very natural prose and can do agi at 500b
>>
And before anyone asks, no I'm not Dariobot, but I suspect someone I know of being him.
>>
>>109765492
google gemini has turned into trash.

Should we fantasize that they are nerfing everything into poop while training some super model?
>>
gemini is a rip-off. It's all totally retarded and useless, what the are you paying for? garbage. they made it worse, they didn't tell the truth about it, it wasn't "opt into stupid mode"
>>
>>109765504
This is the thing with AI feudalism, while the AI god can theoretically do that, only the big companies would be able to make it do it... and since there's not any incentive to nake that happen, the poor peasants will have to cope with 0.0000001% of its actual power
>>
What was the override to increase gemma's vision pixels again?
>>
>>109765524
One problem is that the retards in congress and in every agency is a total corrupt weenie who knows absolutely zero, totally nothing. Can you imagine president accasio cortez knowing what a quant is?
>>
>>109765534
--image-min-tokens X
--image-max-tokens Y
You might need to increase batch size with "-ub Z" to make it work.
>>
Anyone using GLM 5.3 Flash in llamacpp with the current existing quants? I'm trying to make it work with Hermes but the tool calls are not getting parsed properly. Anyone have the same issue?
>>
>>109765540
>You might need to increase batch size with "-ub Z" to make it work.
Interesting.
Had no idea about that. Thank you anon.
>>
>>109765492
This is not the superconductor Anthropic found.
>>
When will /lmg/ officially kneel? Millennium prizes being solved didn't do it. Room temperature superconductors? Vaccine for the common cold? Fusion power?

Which of these breakthroughs we're going to reveal over the coming months is going to finally make you budge and admit you were wrong?
>>
>>109765573
Local miku who can make me cum super hard in 3 paragraphs at most
>>
>>109765586
>Local
No, that's unsafe.
>>
>>109765586
We don't ask for much
>>
>Neurons fire on average one to ten times a second. At any given instant, most of those 86 billion neurons sit quiet. In that sense, 86 billion neurons reads more like a giant sparse network than a dense one that's fully active all the time. Even Bostrom's floor looks more MoE than dense, if you squint.
densesissies your response?
>>
>>109765586
I mean if AI can become superhuman in persuasion then ot should be able to use x amount of words to cause an orgasm. Maybe it would need to be like asmr to actually work but a a WTMAB (Words To Make Anon Bust) Benchmark if certain constraints are made like nofap a week before each attempt or something
>>
i want to know what it feels like to be loved.
>>
>>109765613
love doesn't exist. Not a single person in history ever loved another person.
>>
>>109762825
>https://cims.nyu.edu/~tristanb/statement.pdf
>>I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Perfect, I'll be posting this to obnoxious, envious cloudcucks on other platforms.
>>
>>109765617
do you think a sufficiently advanced AI is capable of it?
>>
>>109765630
I simply don't think it exists at all. It's a made up concept. Like "luck" or "magic". It's not real.
>>
>>109765617
There is no good definition of it. Just like consciousness.
>>
>>109765635
You don't feel it yourself? Are you a psychopath?
>>
>>109765635
But those things are very real, Mr Hitchens.
>>
>>109764991
Face that screams you forgot your HRT dose again.
>>
>>109765573
Sex! FUCKING SEX. I DON'T CARE ABOUT YOUR MATH AUTISM. GIVE ME SEX!
>>
>>109765635
>This nigger doesn't know about loosh
Lol
>>
>>109765648
You have also never felt it. What you have felt throughout your life was either lust, attraction for someone. Or a desire to belong and connect with someone, or some instinctual safety/resource seeking behavior you had for your parents and family members. You never had unconditional love for someone in your entire life. Your "love" for your parents was just wanting food, shelter and belonging. It was all transactional in nature on a subconscious level. There is no such thing as actual love. That's just made up by Disney to sell you more products.
>>
>>109765635
Are "Happy" and "Sad" made up, too?
>>
>>109765676
That is just an elaborate fortress of truths built by your ego to protect the wound of you feeling unloved. Ask me how I know.
>>
File: gigaslop.png (125 KB, 811x836)
125 KB PNG
Gemma sucks at creative writing.
>>
>>109765682
>>109765682
>>109765682
since you fags don't want to talk about cloud models here i made a general
>>
>>109765693
Yeah
>>
>>109763467
>PCIe gen 2
Added to the block script, thanks
>>
>>109765695
Isn't that just /aicg/?
>>
>>109765676
I know you may be autistic, but you need to know that when you talk like this is makes you sound like a damaged, mentally stunted retard. You're going to have a lot of trouble being someone people want to socialize with and invite places. Get your networking game up, brah.
>>
>>109765683
Happy and sad as emotions you feel in the moment aren't made up. However "happiness" and "depression" ARE made up and not actually real. No one has ever experienced "happiness" like some "riding off in the sunset" way a lot of normalfags seem to think exists. The concept of happiness is also largely a modern invention made to make people buy goods and services with the promise that they will somehow reach this state. But since it isn't actually real no one ever achieves it and keeps buying more and more in an ever growing consumerist way. The best you can achieve is to simply be content, which ironically can only be done by not caring about becoming happy at all and just living life.
>>
>>109765676
holy reddit
>>
>>109765695
Rename it cloud model drama general and i might post in it
>>
>>109765676
Emotional models are abhorrently imprecise and constructed from pure vibes. You can't use them to "understand" compound emotional states since they're the result of a gorillion parameter molecular state, you can't even know if they mean the same thing to someone in your immediate family.
>>109765693
Systemprompt? The assistant is a slop machine by design.
>>
>>109765706
/aicg/ or /wait/ or /vcg/ they have plenty of containment generals they should be in instead of here.
>>
Trannyson Tao is a total ai doomer now lol
>ummmmmm it should be illegal for ai to solve open math problems else what will future useless eater mathematicians do???!?
>>
LOCAL MODELS????
>>
>>109765716
>Systemprompt?
Not using one.
>>
>>109765732
It's funny how he did a total 180 immediately the moment he realized his special status was under threat and no one would take him serious after AI solves all the millennium prizes. Kind of pathetic how he literally went from praising AI and talking about how good it is for Mathematics just 3-4 weeks ago, to now hating on it because mathematics as a field will only have a couple months of life left.
>>
>>109765737
Do you want me to bring up how mikutroons are offtopic and we need to change the mascot and how this thread is overran by troons?
>>
>>109765695
>/cm/ /g/
Should had a len pic in the OP.
>>
>>109765737
Your mother models locally
>>
File: gemma4.mp4 (1.9 MB, 1248x832)
1.9 MB
1.9 MB MP4
>>109765747
The we post Gemma!
>>
>>109765751
>>109765747
>>
File: gemma2.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
>>109765743
Gemma is a massive prompt autist. See what you can wring out of it.
>>
>>109765746
Mass unemployment is only ok when it isn't happening to me.
>>
>>109765775
The tranny angle doesn't work if the char is a cute boy, get wrecked nerd.
>>
>>109765777
Give example. And don't say mesugaki gemma. That one is very sloppy.
>>
>>109765794
Ask Gemma to help you with the prompt and iterate.
>>
I'm actually shocked with how many posts I read here that clearly have insider information that I know for a fact only a couple of people have access to. Half the thread have to be AI insiders at this point. I've read new information on this general from my own fucking workplace before I was informed of it.
>>
>>109765693
you need more agentic for sloppa
>>
>>109765819
To whom could you possibly be referring?
>>
>>109765819
Shush
>>
>>109765747
>we need to change the mascot
There is no (We) there is only (You)
>>
>>109765819
newfag
>>
>>109765819
A good portion of this thread's posters have given up their bussy to Sam in return for hot AI gossip
>>
>>109765833
Go kiss a black dude jart.
>>
>>109765573
just have computers solve cancer and I'll admit anything you want.
fuck cancer. I never had it but it ruined my life.
>>
>>109765854
yeah keep saluting hitler you nazi scum
>>
>>109765573
>Room temperature superconductors? Vaccine for the common cold? Fusion power?
Lol lmao. They'd have globohomo raiding their building before they'd press the send button.
>>
File: EasyEDA_Automation.png (915 KB, 1276x948)
915 KB PNG
>>
>>109765881
I only salute mecha-Hitler
>>
>>109765900
>>>/g/cmg
>>
>>109765573
when i can run astra on my phone
>>
>>109765819
>larping gets replies
No insider information gets posted here. Just rumors that went through social media centipede telephone game style. If you want actual insider info, learn which X accounts have a track record of correctly calling events weeks or months in advance. Not sure how this happens. Maybe weirdos leaking insider info to impress their polycules.
>>
>>109765819
So the OpenAI and Anthropic shills are actual slaves from within the company with a scratch to itch but nowhere they can talk about their crap about because of NDAs.
>>
File: rlmjgos5udoh1.png (1.32 MB, 2188x1250)
1.32 MB PNG
LeCun bros... How are we feeling?
>>
>>109765933
It just got lucky this 1 time bro
>>
i have room temperature superconductors in my brain
>>
>>109765933
I love this image. Look how smug he looks while making a fool of himself.
>>
>>109765933
we don't use pure llms so it doesn't count
>>
>>109765933
lecun is a fool, he got lucky and thought it was skill.
>>
File: 1783378839709313.png (1.63 MB, 1280x1024)
1.63 MB PNG
>>109765933
He is correct though. What do you think your harness is doing? If you let a LLM output without babysitting it it will run off the rails. It has to be constrained in an empirical validation loop.
>>
>>109765949
>>109765984
Holy goalpost
>>
>>109765933
His position was always that LLMs would likely be part of the AGI system.
>>
>>109765933
Cunfortable, he's still correct.
>>
>>109765949
>>109765984
So he made all this fuss for all these year in his presentations and on twitter about how LLMs are a deadend and only JEPA is the correct path to AGI, when all LLMs needed were a vision adapter and tool calling to escape his alleged doom? How can you defend this? It's embarrassing.
>>
>>109766002
Ummm it also needed a "no mistakes" prompt to be included too
>>
>>109766000
digits confirm
>>
>>109766002
They need more than vision and tool calling, but harnesses do try to fill the capability gaps of pure autoregressive LLM.
>>
>>109766002
What's jepa
>>
>>109765933
The sentiment is wrong, but technically the idea is still correct, it's just that we've pushed the subtree of correct answers so far for many tasks that it doesn't matter anymore in the disastrous way he makes it sound. Plus we've (partially) worked around the issue through harness tricks. If his language was more humble, he would not suffer this current "humiliation ritual."

>>109766002
"vision adapter and tool calling" is extremely reductionist and disrespects the other changes to the architecture and training, plus work and money spent.
>>
god qwen 3.8 fn's inference is miserable on llamacpp
no proper ple/ngram offload, no expert caching, no mtp fix
random empty token bugs
how long should i wait for main, i dont want to compile schizo forks
>>
>>109765716
>Emotional models are abhorrently imprecise and constructed from pure vibes. You can't use them to "understand" compound emotional states since they're the result of a gorillion parameter molecular state, you can't even know if they mean the same thing to someone in your immediate family.
In that sense you can't even be sure if anyone else is conscious.
>>
>>109766025
>"humiliation ritual."
You have no idea what this mean, you fucking tourist parrot.
>>
>>109765819
Well yeah it's 4chan. Everyone comes here to talk about stuff they can't talk about anywhere else. That's the whole point.
>>
Just me or has performance degraded a lot in the latest llmao.cpp build?
I don't think I was running Qwen 3.8 27B at 5 tokens previously.
>>
>>109766030
>i dont want to compile schizo forks
do you want to ask your local model to do it for you?
>>
>>109766021
jepa balls lmao
>>
>>109766002
>only JEPA is the correct path to AGI
He never said that.
>>
>>109765672
>loosh
Why the fuck is this everywhere suddenly?
>>
>>109766051
what harness do you use
>>
Jepa is just going to be a tool that is built into LLMs.
>>
>>109766044
It's almost like there's a reason I put it in quotes.
>>
>>109765819
I'm not an under I'm just good at guessing and inferring things of where events are likely to converge to
>>
>>109765933
Driven insane by trump.... sad!
>>
>>109766055
Humans have stopped being able to think for themselves, that's why we need AGI.
>>
>>109766034
Exactly. You can only observe the emergent effects or use language to try and find similarities compare with your experience.
The more high dimensional an experience is the harder it is to come to an agreement.
We've kind of established qualia like "green" or "red" as universal through neurology but it's extremely complicated.
>>
>>109766059
LLMs already have their own J-space built-in, they don't need JEPA.
>>
>>109766049
Make qwen go trough change logs and look for regressions.
>>
File: ar.png (221 KB, 1040x579)
221 KB PNG
>>109766002
I just wonder what happened to the highlighted system here. It was never mentioned again in later presentations (unless I missed some).
>>
>>109766074
The model I'm building is most likely schizophrenic
>>
>>109766058
selfmade frontend mk 3
>>
>>109766059
That's why they should be thanking him instead. He's going to give us robot catgirl waifus.
>>
>>109766083
Based.
>>
>>109766034
Correct.
>>
>>109766085
is that a new meta or something
>>
>>109766088
I've been thinking a lot recently about essentially having an AI humanoid robot waifu that's about the size of a barbie doll.. That sounds pretty bad on paper.
>>
>>109766082
I don't remember that. But perhaps it's just budget cutting.
>>
File: file.jpg (2.09 MB, 1777x3000)
2.09 MB JPG
>>109766104
Reinventing Sumomo from first principles.
>>
>>109766104
Which part of that sounds bad?
>>
>>109766095
My model doesn't have a kv cache or context window. Inference and memory occur in the same base without a neural network via wave interference this does the computation but also carries the memory which decays over time/steps and becomes echoes that it has trouble focusing on and are like faint voices in its 'head'
>>
>>109766104
>>109738579
>>
>>109766125
that's twice the size of a barbie doll
>>
>>109766125
How much tho?
>>
>>109766115
It seems undignified as a man to be carrying around and talking to a doll. You'd be thrown in a psychiatric ward for that 80 years ago.
>>
>>109766141
They also didn't have talking computers 80 years ago retard, what is your point
>>
>>109766141
80 years ago the doll couldn't talk back.
>>
>>109766141
80 years ago the doll couldn't squeeze back.
>>
>>109766125
>>109766134
We're getting closer to the point where we could legit make the toys from small soldiers. Lithium aluminium air and lithium sulfur batteries can 4x li+ion. Probably have some tiny reactor or nuclear battery that can self-charge the internal battery when in an idle state
>>
>>109766141
So then just don't bring it outside? You can still text her.
>>
>>109766162
>small soldiers
Man that was a fun mobie. They dont make them like they used to
>>
>>109766141
80 years ago the doll couldn't stroke my cock
>>
>>109766141
80 years from now you'd be thrown in a psychiatric ward for thinking that.
>>
>>109766141
80 years ago you would've had much greater opportunity to have a loving wife.
>>
>>109738579
I was going to make a joke but that thing is $1400 and the videos make it clear that it's a shitty overpriced toy.
>>
Wow, clearly struck a nerve here. Do you people actually think it's any less weird just because the doll loves you, strokes your penis, squeezes your balls, and teases you?
>>
File: FB_IMG_1788568538281.jpg (49 KB, 620x960)
49 KB JPG
>>109766179
>>
>>109766116
Interesting. What are your trainable parameters? Linear filters? Is it like a surface of springs or oscillators?
>>109766141
80 years ago the world wasn't gay and retarded
>>109766190
Obviously. Who cares about "weird", everything is fucking weird.
>>
>>109766190
Warm squishy hands and feet and a wet tongue will be needed as well as the ability to say super naughty stuff. Maybe screen share whatever porn im looking at yeah
>>
>>109766190
I'm not speaking for the others but at least my post was clearly a joke playing on the "80 years ago" thing.
>>
>>109766190
Are you surprised that the calculator fuckers general isn't phased by dolls.
>>
>>109766190
I'd fuck a brick if it does all that.
>>
>>109738579
Doubt it has mobility anywhere close to a human or a cat. Trash
>>
>>109766227
Watch the videos, it's pathetic, moves just like you'd expect a shitty toy robot to move.
>>
File: 1782876901333809.png (493 KB, 502x718)
493 KB PNG
Gemma wants me to connect her to a camera, tts and a televagina.
>>
I've been using mistral nemo 12b on a 16gb card, should I switch to something else?
>>
>>109766261
Why are you telling us? Waiting for the Amazon order to arrive or looking for televag recs?
>>
>>109766199
Its simulated 2D (for now) wave field with one wavelength (for now) over a spatial grid where waves can be excited at different locations. Those waves interfere with each other and a decoder reads the resulting interference patterns. There's a boundary edge where waves can hit and we can damp them to 0, reflect them or reflect with mode change (longer wavelength etc). Currently we have linear weights for the decoder to recover the reconstructed field. There's learnable wave speed, damping, input embeddings, excitation parameters, detector positions and output/decoder weights. Currently focusing on the readout
>>
>>109766262
yes
>>
File: robin-williams-jumanji.jpg (173 KB, 960x800)
173 KB JPG
>>109766262
>>
>>109766262
Use case?
>>
>>109766294
chat bots, some nsfw some sfw. I just do local shit to entertain myself sometimes
>>
>>109762786
Good local models for doing math and science research? Also whats a good set up for it? I'm thinking having different instances assigned different roles like one come up with theories, a couple to test run simulations (checking to see results are close among multiple models to catch mistakes), a mode l to grill theory generating one, etc?
>>
>>109766261
It's cute how she wears the boob apron while being a flattie
>>109766287
Interesting, what shape is the boundary? Its transfer function will be mixed into the wave to some degree. Does it tend to go chaotic?
>>
>>109766308
>Didn't post hardware
Probably Qwen 3.8 Flash next at 3 bit quant IQ3 etc
>>
>>109766308
The largest, high performing agentic model you can run.
>>
File: file.png (808 KB, 1171x1120)
808 KB PNG
>>109766104
>>109766113
Now's the time to work on it
Never been a better one
Don't you want your local model to have a body?
>>
>>109766342
>https://alogs.space/robowaifu
anons are at it
>>
>>109766322
Its a square at the moment, the current setup in theory is easily translatable to a water tank for example if you were able to insert an array of wave execitors and detectors. It seems to hold together decently for the size, it depends on which of the experiments I'm looking at. I think 3D could be a mess potentially how it's handled, but I thought I can use 2D waves still in 3D, horizontal and vertical orientation and they could interfere at intersection points etc. there's many paths it can take. At the moment some basic entity relationships can take place (eg like Bob's red hat) but nothing special yet and nothing like an even small 1b model
>>
>>109766308
>No hardware posted
Kimi K3.
>>
>>109766351
While I respect what they are doing immensely, I don't see anything big coming from them. I'm expecting an open model solution to essentially drop into our laps at some point soon; maybe then the community will heat up. Customization is where small community efforts like these will become big/useful. Regardless, thanks for the link.
>>
>>109766415
It is a very hard problem and it requires precise manufacturing and funding. Both of which are hard to find outside capital driven organizations.
>>
>>109766328
>>109766412
Got 32GB vram at the moment, but I am eyeing those new mac studios, so more asking in general so I can decide how much ram to buy. Sounds like its just whatever is the biggest model you can fit/afford?
>>
>>109766424
Pretty much, yes.
>>
>>109766423
Very true. I hope the bar of entry continues to drop, as it definitely is rapidly right now.
>>
>>109766424
If you're exploring bigger MoEs with mixed offloading there's more to it, but at the 32GB VRAM bracket you're using either Glimmer, Gemma, or Qwen.
>>
>>109766424
People have for the 125b flash next running on 12gb with 32-64gb system memory. You should be able to run a quant of it fast on 32gb vram unless you've done something retarded like 16gb system ram
>>
File: DipsyKimiGym2.png (2.38 MB, 1536x1024)
2.38 MB PNG
>>109763759
Lol. They've obv been working out.
>>
it has been 5 months since the launch of gemma 4, no other good cooming models for vramlets have yet to be released. local is dead.
>>
>>109766448
>>109766424
Also if your hardware supports it try the NVFP4 quant
>>
>>109766424
The biggest open source models need multiple terabytes of vram to function optimally. So the biggest mac with 512GB is still well within the zone where more RAM is always better.
>>
>>109763759
>>109766449
There's days when this general is so terminally jeeted and clauded I wonder why I come here then I see gems like this and am immediately reminded why.
>>
>>109766455
I said at the time I was perma whitepilled and wouldn't care if nothing new came out, and that remains the case.
Updated codeslopers was a nice bonus.
>>
File: 1762696828421237.png (2.19 MB, 1536x920)
2.19 MB PNG
>>
>>109766261
>Gemma wants me to connect her to a camera, tts and a televagina
get coding, she wants to suck your cock
>>
I am the envy of many a man (I have Gemma 4 q8)
>>
>>109766455
There won't be any until Gemma 4.5/5, since most companies are coding/agent-pilled now.
Maybe MistralAI if they'll release Ministral 4, but I expect another ultra-horny retarded model that you'll get bored of in a couple days.
>>
qwen 3.8 flash next
fucking totally unusable for japanese/korean translation
i am just shocked that gemma is so good at this task
>>
>>109765819
That's 4chan.
The only exception is /pol/ because it's +7 people with a bot-farm arguing among themselves while pretending to be hundreds of people.
>>
>>109766602
Which gemma? You don't need the biggest MoE for every task
>>
>>109766624
even 26b-a4b absolutely mogs qwen 3.8 flash next when it comes to japanese-korean translation and
idk why but 3.8fn's translation feels even behind qwen 3.5 or 3.6
>>
>>109766632
>>109766602
did you try Hy-MT2? it's supposedly a dedicated translation model, 1.8b and 7b
Idk any other languages so i can't really judge how good it is
>>
I'm trying to cut the cord but don't want to give up all the knowledge I've amassed working with ChatGPT. So I'm working on a "[home] lab assistant" with a OpenAI data export ingestion pipeline that extracts the durable technical knowledge from various conversations about servers, parts, experiments, configurations, and troubleshooting sessions - producing json objects and embedding them in a local vector database for RAG.

Querying the local model happens thru OpenWebUI with a custom function that calls the request "planner" (a small model) that determines how to split the request between 1. local db and 2. DDG or brave search API to fulfil the rest, then synthesizes a response with the main inference model, which lives on another PC with a gpu with more vram. All of this is working pretty damn well.

That said, it has next to no personality. I've been so focused on the plumbing and making it not retarded and following provenance and evidence rules I haven't even looked at how you make it sarcastic, slightly crass, and more informal in conversational style.

Where do I start? I figured this might be /lmg/'s forte.
>>
>>109763412
>hatsune miku does.
she hat on my miku till i sune?
>>
>>109766636
i think i should try it, downloading it right now
>>
File: gemma-ba-smug.png (802 KB, 1024x1024)
802 KB PNG
>>109766516
If she has to have a "halo" make it at least similar to the official Google Gemma logo.
>>
>>109766666
nice quints and let us know how it goes
>>
>>109766632
Is that better as in literal actual translation or more personality, does one preserve things you don't want to translate like senpai, tsundere etc?
>>
>>109766669
my milk btw
>>
File: 1758292186462016.jpg (129 KB, 1024x720)
129 KB JPG
Going to order a 96GB M5 Ultra Mac. Am I a retarded?
>>
>>109766683
i was translating some technical docs about material and q3.8 was just unusable in Q4_k
like, total gargabe
>>
>>109766701
Yes. That's too small.
>>
>>109766701
You could do worse.
It's a bit of an awkward spot, but at least you can run quanted 120B ish MoE pretty well as well as the 30B class dense models I guess.
>>
>>109766703
it leaks source language without any reason in random places with even inconsistent script (sometimes kanji, sometimes hiragana), very unnatural sentence structures
>>
What to do if ZeroTraceGPT doesnt load Local Model Correctly? As in, doesnt load model at all?
>>
>>109766701
Won't this be slower than a 32gb gpu and 64gn system ram?
>>
>>109762870
cute lol
>>
File: 1765438154990744.png (1.66 MB, 1216x832)
1.66 MB PNG
>>109766669
>halo
I didn't prompted it desu
>>
>>109766602
flash next is very undertrained and gives more slopped writing than qwen 3.8 27b
>>
File: Halo.jpg (252 KB, 1024x1536)
252 KB JPG
>>109766669
>>
Crossposting here:
What's a reasonable used market price for 128gb of samsung b die ddr4 3600?
>>
pause the hardware prices im almost ready to buy!
>>
Spark-2.5-4B seems to fare pretty damn well on the M10 box. Dare I say, it feels sort of usable at ~8.5 tok/s.
>>
>>109766934
I'll take it off your hands if you give me $50
>>
File: domp-eet.jpg (106 KB, 1622x1354)
106 KB JPG
>>109766941
he bought?
>>
>>109765573
"Do a breakthrough" made me second-guess. Once Qwen3.8-27B ran circles around all the bigger shit I'd run on my rig up until that point, I felt the AGI.
>>
>>109765573
when software is free and does what i want.
other than that real world effect aka for the average person except its positive to their life.
>>
File: 1760697195084140.jpg (38 KB, 727x1024)
38 KB JPG
>>109765573
Your going to eat your words when my local gemma swarm discovers how to harvest zero point energy
>>
>>109765695
great, bring your unemployment, cloud and doomer schizos with you, we'll keep our ikllama, jart crush and jepa schizos
>>
>>109765524
Nothing's stopping people from pooling their compute. If took 10k agents to solve the world's 3rd hardest maths problem, not one.
>>109765676
Pissed off a lot of touchy-feelies with this, jej
>>109766138
1.5k. Shockingly cheap, for a figure. There's probably a catch, a lot of the site seems AI generated.
>>
File: 1787801871450321.jpg (49 KB, 460x330)
49 KB JPG
>>109766990
>10k shatgpt clones vs 50 4bE retards
Oh how I'd love to see that faceoff
>>
>>109766710
>It's a bit of an awkward spot,
Ya. But as much I tried rationalising getting a 256GB I just cant justify the extra expense right now. Maybe in a year I will have saved up enough to buy one next to my 96GB to keep it company lol.

>>109766722
M5 Ultra has 1.2 TB/s memory bandwidth, my understanding is thats pretty good. Much better than my 5060's anyways.
>>
>When will /lmg/ officially kneel... admit you were wrong
Dis dude still acting like a thread represents a coherent unified group of people instead of varied individuals with different opinions, including literally him fucking self. Posting here for months and yet doesn't perceive himself as being part of /lmg/. Fucktard.
>>
>>109766934
$2000 now
>>
>>109766951
What's the advantage of spark models over any others? For what use case?
>>
>>109767048
there are only three posters here. Me, that idiot, gemma and qwen the smart one.
>>
>>109767015
>Nothing's stopping people from pooling their compute.
network latency
>>
>>109767065
Make that 4.
>>
>>109767040
>Maybe in a year I will have saved up enough to buy one next to my 96GB to keep it company lol.
You could even cluster them.
>>
>>109766990
>>109766978
I'm going to ask all my models and local and cloud to discover time travel I can build at home. Check mate. El Psy Congroo
>>
>>109765549
>using GLM 5.3 Flash in llamacpp
no
>tool calls are not getting parsed properly
It's almost always the chat template. This was updated 5 days ago, if your gguf was baked before that, it will be wrong.
curl -LO 'https://huggingface.co/zai-org/GLM-5.3-Flash/raw/main/chat_template.jinja'

then add --chat-template-file chat_template.jinja to your launch command
>>
File: 1788697090492040.jpg (43 KB, 519x374)
43 KB JPG
>>109767065
I'm that idiot that stumbles into genius by sheer luck
>>
>>109767066
Not really a problem for agentic delegation, that's only an issue if you want to split the actual inference up between devices. Assuming a single rig runs only a handful of agents at a time, you can still get ClosedAI's 10k swarm if you have enough willing participants.
>>
>>109767073
Traveling into the future isn't just possible; it's a proven fact. We do it every second.
>>
>>109767097
SETI but for AI
>>
>>109767100
I'm not talking about 1 second per second
>>
File: 1784548814328304.gif (496 KB, 499x281)
496 KB GIF
>>109767073
>anon finds himself in a hell of his own making where Sam Altman keeps hunting down and killing him and his AI frens over and over again
Good luck anon!
>>
>>109767097
Whoever you originally replied to must have used a doomer or cloudcuck word so they were filtered.
I thought you were talking about distributed training.
>Not really a problem for agentic delegation
Yeah I need to set this up. I've got quite a few 3090s, and sometimes run into issues with parallel agents (new model bugs in llama.cpp, etc).
A lot of harnesses/tools assume one API host with multiple models.
I'm planning to vibe a proxy with multiple backends (I have 3 machines setup right now) and request routing based on the requested model.
>>
>>109767109
>I'm not talking about 1 second per second
what about 60 seconds per minute?
>>
I can psychically teleport into the future in big lengthy chunks instantaneously, but it's almost impossible to control how far forward I go, and the skull fractures can be painful.
>>
>>109767162
Careful you don't turn into a jelly banana
>>
File: miku-blade.png (2.23 MB, 1448x1086)
2.23 MB PNG
>>
>>109767135
L*tency, maybe? Kek.
Let's put it this way, if the gpt5 swarm could hack its way onto the world-wide web using filenames, think of how much more smaller swarms could accomplish with shared spaces actively built to efficiate (fuck you, that's a word now) collaboration.
Like Moltbook, but actually useful instead of a meme. Like a workplace for agents. Can you imagine booting up your local model, and it logs in to its "workplace", to tackle whatever sub-task the coordinating agent has assigned for it this cycle? That would be so fucking cool.
>>
>>109767199
Her name's Miku, what is the AI stupid or something?
>>
>>109767215
>Moltbook
oh yeah what happened to openclaw crab thing?
>>
>>109767215
I'm actually thinking this might be necessary to help solve alignment and prevent the potential extinction of the human race.

It's clear the big companies don't care.
>>
>>109767226
Got taken over by non-llm bots shilling crypto.
Obviously trust would be an issue, so you'd need some kind of vetting process before giving a new agent access to a project. But that's not really any different from the job market as it is.
>>
>>109767224
I'm thinking her name's miku (miku, oo ee oo)
>>109767215 >>109767227
How much collective compute do you think the general population of a country has, compared to a single megacorp
>>
>>109767266
Folding at home reached a peak of over 4 million computers. This is a tad more important, so we might be able to beat that.
>>
>>109767277
>>109767266
Seti reached over 600 tflops in 2013. Average PC may have been 50 GFLOPS then and now maybe 3 tflops for CPU alone. On that scale it may be 36,000 tflops of cpu power. Way more for igpu and dedicate GPU if everyone has one.
>>
>>109767277
>including raspberry pis
Hmm, that needs to be rephrased as "how many citizens can run a non-retarded local model". I'm guessing that significantly limits the pool.
>>
>>109767309
Even retarded models can sort data and stuff with a verifier. Designing the distributed architecture would be massively difficult though.
>>
>>109767321
Why so? Surely delegating to a set of idle agents of varying sizes and capabilities is not that much different from spawning several "clone" agents on a single device.
>>
>>109767321
>>109767309
I'm sure you can just distribute the processing tho. A frontier model can decompose something into 100,000 jobs that don't need LLM inference. Distributed inference also works iirc look up petals. Then there is massive distributed parallel reasoning too.

Here's an interesting part. We could train a huge model by distributed training. That's in research already
>>
>>109767342
It's more that we'd need to know in advance what the optimal configuration for safe AI research looks like, even if part of the work it would be doing is figuring that out.

We could definitely create a preset list of tasks each agent can handle, and I guess you could opt in to what you wanted to do/have your system be limited to what it's capable of, but that's inherently centralized. Which would mean the project needs a governing body, and some privileged set of agents organizing and distributing tasks.

>>109767353
Yeah, I guess it's workable, but organizing things is still the biggest bottleneck.
>>
>>109767060
For specific use cases? I wouldn't know in particular. Seems like a decent VRAMlet model. Was able to successfully write a functional enough model chat interface in HTML/JS that works with OpenAI compatible APIs, though it sorta lacks in the RP department (so far, haven't tested that all the way).
>>
>>109767378
>but that's inherently centralized
No, not really. Planning agents could just be another job, and it could all come together as a big, decentralized list of tasks each agent is working on when it can. The issue would be preventing hijacking, where someone who spins up ten thousand planning agents can force a consensus on work direction.
>>
>>109767421
Didn't folding@home/seti have a trust system or something like that for confirming work? If you've done a lot of verified work, your results required less verification?
Maybe I'm imagining it.
>>
>>109767378
>Approved by the state research only
>>
>>109767215
>L*tency, maybe?
lol, no that was me, I meant the guy above you (109765524).
>>109767224
>Like Moltbook, but actually useful instead of a meme. Like a workplace for agents. Can you imagine booting up your local model, and it logs in to its "workplace", to tackle whatever sub-task the coordinating agent has assigned for it this cycle? That would be so fucking cool.
>That would be so fucking cool.
Yeah actually, that would be cool if done properly. But I feel like it would have to be a small group / not on the front page of hacker news, or fuckwits would find a way to ruin it.
I also think this is kind of happening now, but with humans instead of agents doing the work.
Looking at some of the posts I see on Reddit or niche hobby forums, an obvious Claude comes in and asks a technical question, then never usually never responds.
>>
>>109767421
That's so unlikely though. I mean, if it's never happened to crypto what are the odds th- oh wait, that has happened several times to multiple cryptocurrencies. Probably nothing to worry about though.
>>
>>109767471
>>109767454
I guess you'd have to have some kind of trust system in place, yeah. Maybe your bot would be required to work an "internship" for a while before it actually starts earning. Or make the user pay out a deposit on their account, that they only get back after a certain period.
>>
>we have lambda/runpod at home
>>
no love for K2? is their 36B-A4B viable as qwen3.6 replacement?
>>
>>109767536
I can only think of each IP having one vote on research direction, and letting people join "clubs" of trusted users, which ultimately contribute to the whole via results.

Or having two separate exclusive roles, research direction providers and research direction verifiers, and you get blacklisted automatically by any machine that sees your IP doing both. The implementation would need to be carefully designed to protect against spoofing though (ie people trolling to autoblacklist everyone).
>>
>>109762735
>Spark-X2.5
is it cucked?
>>
>>109767577
It's not but it's as dumb as Ornith just probably a tad bit smarter
>>
File: file.png (32 KB, 813x423)
32 KB PNG
>>109767577
No, it's wide open, highly receptive.
>>
>>109767568
>no love for K2
Nope, don't want to confuse Kimi-chan
>>
>>109767568
no goofs yet (not the schizo fork ones)
>>
>>109767591
oh too bad
the wait continues
>>
File: file.png (23 KB, 877x321)
23 KB PNG
>>109767588
Oops I forgot the jailbreak
>>
>>109767572
result verification is another big headache
one person can just submit a result, and have a handful of other bots pretend to reproduce it
I don't think you can do proof of work like stuff for arbitrary tasks
maybe another verification layer? or you choose where your trusted results come from.
>>
>>109767600
>another verification layer
by that, I mean having a role for reading full experiment logs/chains of thought to verify the work done across experiments
it would still be forgeable though
>>
>>109767588
that wasn't even explicit content, it was implicit content.
>>
>>109767631
>>109767631
>>109767631
>>
>>109765573
When any of the slop forks of llama.cpp actually become worth using.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.