[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1766987457285058.png (81 KB, 2160x2160)
81 KB PNG
China lost.
https://x.com/claudeai/status/2080699495453528290
>>
File: 1781791478578862.jpg (338 KB, 3840x2160)
338 KB JPG
Probably nothing. Maybe it just got lucky. They're all very stochastic after all. No need to worry. Carry on.
>>
>>109360389
why does that chart say 53.4% is larger than 53.5%?
>>
>>109360422
its vibe coded
>>
>>109360452
You don't get that kind of performance increase with slop. Andrej Karpathy likely worked his magic. He's the best engineer they have.
>>
>>109360389
Shame they have to ask hold out questions on closed models through the closed models, trusting their honor (LOL) not to detect them and put them in their benchmax database.
>>
>>109360407
It means they made a synthetic data and reinforcement learning framework specifically for arc agi type puzzles.
>>
>>109360389
>all the caveats and pilpul in that table
Kek
Also, inb4 PELICAN
>>
>>109360478
the chart retard
>>
>>109360389
>gets its ass kicked by sol on deepswe
you ain't winning shit, dario
>>
>>109360389
Wait 2 months till China distills these into new open weight models
>>
>>109360389
Why would they do a major release on a Friday? Poor SREs.
>>
File: file.png (26 KB, 685x213)
26 KB PNG
Lol
>>
>>109360821
>>109360849
it'd be funny if they can finetune k3 on opus 5 and end up with better perf than the original
>>
>>109360878
K3 seems to be performing better in some areas than Fable 5 like in frontend tasks. So its at least possible in specific types of tasks
>>
>>109360389
Where is the comparison with chink AI though?
>>
>>109361285
It's unnecessary as chinkmodels are getting banned on monday
>>
>>109360389
so is it better than fable or not?
>>
Thanks, but I'll keep using Grok 4.5
>>
File: HKe81mLaUAAM_qr.jpg (221 KB, 1024x1024)
221 KB JPG
>>109360389
Finally something to replace that stupid Fable.
>>
>>109361512
I don't get it.
>>
>>109361512
I get it.
>>
>>109360491
even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?
>>
Tried it. Fable is significantly better, This model is too rigid, autistic, lacking any nuance while missing the bigger picture. It also lacks the ability to generate novel thoughts.

Generally underwhelming.
>>
>>109360389
>no comparison to chinese models
They're afraid.
>>
>>109360389
I am still not paying. FOSS models ftw
>>
>>109362842
Name one usable FOSS model you can run at home.
>>
Make it play Pokemon Red, faggots!
>>
>DOOOOD I'M GOOONNAA BENCHMAAAXX
>>
>>109360389
Pelican status?
>>
>>109364731
ITS UP
https://news.ycombinator.com/item?id=49038433
>>
>>109360389
>mitochondria is the powerhouse of the cell
>bioterrorism detected. police has been notified about this conversation
>>
>>109360407
>log scale
>....10,000
>....20,000
>....
Nobody error checked this AI made pic, huh?
>>
>>109360849
>>109360878
>>109360917
Instead of jerking off about 0.5% accuracy splits how about realizing that 75% is total dog shit.
AIs are only routinely getting 90% correct on known, solved problems.
These programs need 3 sigma or better to be a functional every day utility.

Imagine the computer system that runs your society gets 2+2=4 == 2x2=4 == 1x4=4 == 4x1=4 wrong 1/10 times, every time. Something a 7 year old will get correct every single time.
>>
File: file.png (212 KB, 1614x1350)
212 KB PNG
>>109362964
>>
>>109364865
you are pretending as if there have been no real productivity gains from llms
>>
I refuse to use closed models anymore, I will not assist in their development by using them, may the companies that create these models go bankrupt and their CEO's rot in jail. Thank you for your attention to this matter.
>>
>>109367346
What if closed models are able to provide better performance than open weight ones
>>
>>109367346
Where are you going to get new open source models when there are no closed source models left to distill?
>>
>>109366580
Slaughtered that freak
>>
>>109361595
>even if they can detect the closed set questions, wouldn't they also need the correct answers to benchmaxx?
They can pay humans for that.

It's not like the benchmarks use unsolved conjectures two people in the world are interested in.
>>
File: 546352536758976.png (2.04 MB, 1209x1244)
2.04 MB PNG
>>109360389
China wins the long-con.
>>
>>109371677
They also dont have DEI to worry about and mentally ill gender cultists. They won 16 years ago.
>>
>>109367346
i never understood how they are benefitting from us using the llms aside from giving them some sentences/information they havent consumed already ?
>tell llm to write some code
>it does, in some capacity
>???
>now that llm is better
how ?
>>
>>109372204
They trained this goycattle extra hard
>>
china is building like 15 nuclear power plants PER DAY

you lost



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.