China lost. https://x.com/claudeai/status/2080699495453528290
Probably nothing. Maybe it just got lucky. They're all very stochastic after all. No need to worry. Carry on.
>>109360389why does that chart say 53.4% is larger than 53.5%?
>>109360422its vibe coded
>>109360452You don't get that kind of performance increase with slop. Andrej Karpathy likely worked his magic. He's the best engineer they have.
>>109360389Shame they have to ask hold out questions on closed models through the closed models, trusting their honor (LOL) not to detect them and put them in their benchmax database.
>>109360407It means they made a synthetic data and reinforcement learning framework specifically for arc agi type puzzles.
>>109360389>all the caveats and pilpul in that tableKekAlso, inb4 PELICAN
>>109360478the chart retard
>>109360389>gets its ass kicked by sol on deepsweyou ain't winning shit, dario
>>109360389Wait 2 months till China distills these into new open weight models
>>109360389Why would they do a major release on a Friday? Poor SREs.
Lol
>>109360821>>109360849it'd be funny if they can finetune k3 on opus 5 and end up with better perf than the original
>>109360878K3 seems to be performing better in some areas than Fable 5 like in frontend tasks. So its at least possible in specific types of tasks
>>109360389Where is the comparison with chink AI though?
>>109361285It's unnecessary as chinkmodels are getting banned on monday
>>109360389so is it better than fable or not?
Thanks, but I'll keep using Grok 4.5
>>109360389Finally something to replace that stupid Fable.
>>109361512I don't get it.
>>109361512I get it.
>>109360491even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?
Tried it. Fable is significantly better, This model is too rigid, autistic, lacking any nuance while missing the bigger picture. It also lacks the ability to generate novel thoughts.Generally underwhelming.
>>109360389>no comparison to chinese modelsThey're afraid.
>>109360389I am still not paying. FOSS models ftw
>>109362842Name one usable FOSS model you can run at home.
Make it play Pokemon Red, faggots!
>DOOOOD I'M GOOONNAA BENCHMAAAXX
>>109360389Pelican status?
>>109364731ITS UPhttps://news.ycombinator.com/item?id=49038433
>>109360389>mitochondria is the powerhouse of the cell>bioterrorism detected. police has been notified about this conversation
>>109360407>log scale>....10,000>....20,000>....Nobody error checked this AI made pic, huh?
>>109360849>>109360878>>109360917Instead of jerking off about 0.5% accuracy splits how about realizing that 75% is total dog shit.AIs are only routinely getting 90% correct on known, solved problems.These programs need 3 sigma or better to be a functional every day utility.Imagine the computer system that runs your society gets 2+2=4 == 2x2=4 == 1x4=4 == 4x1=4 wrong 1/10 times, every time. Something a 7 year old will get correct every single time.
>>109362964
>>109364865you are pretending as if there have been no real productivity gains from llms
I refuse to use closed models anymore, I will not assist in their development by using them, may the companies that create these models go bankrupt and their CEO's rot in jail. Thank you for your attention to this matter.
>>109367346What if closed models are able to provide better performance than open weight ones
>>109367346Where are you going to get new open source models when there are no closed source models left to distill?
>>109366580Slaughtered that freak
>>109361595>even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?They can pay humans for that.It's not like the benchmarks use unsolved conjectures two people in the world are interested in.
>>109360389China wins the long-con.
>>109371677They also dont have DEI to worry about and mentally ill gender cultists. They won 16 years ago.
>>109367346i never understood how they are benefitting from us using the llms aside from giving them some sentences/information they havent consumed already ? >tell llm to write some code>it does, in some capacity>???>now that llm is betterhow ?
>>109372204They trained this goycattle extra hard
china is building like 15 nuclear power plants PER DAYyou lost