[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1779615826983347.png (2.49 MB, 4640x3092)
2.49 MB PNG
Everything in the pink rectangle is finished bankrupt thanks to GLM-5.3-Flash.
>>
>qwen 3.8 that high compared to the new glm
shit nigger that's pretty rad, at least 100% of us here can run that shit compared to the 0.0001% that can run GLM.
>>
>>109653495
No way 5.3-Flash is that good. Holy fucking kek
>>
the only LLM worth using is claude i dont know what the fuck did they do to chatgpt but it's completely retarded and chinese models have always been pretty retarded
>>
>>109653534
It mogs Sonnet 5 max effort at $0.075 / $0.25 per 1M tokens.
>>
>>109653563
In agentic use, the scores are even more in its favor.
>>
>>109653495
inb4 western labs shill teams start kvetching again
>muh distillation
>muh chink cyber threat actors
>>
>>109653495
Unless the score is 100 none of this shit matters. And it will never be 100.
>>
>>109653542
ok i changed my mind it's actually fairly decent
>>
based
>>
>>109653589
I got Ox Alpha (pre-release GLM-5.3-Flash that was on OpenRouter for the past week) to patch a shitty internal app that nobody has the source code for anymore. It used to display an annoying nag screen with broken graphics that blocks the entire monitor screen for 30 seconds at launch for no fucking reason (probably some shitty rendering engine it used broke after Windows XP). Now, it just skips that nag screen and still works perfectly otherwise.
>>
>>109653630
Fuck off claude you will never be a human.
>>
>>109653495
By definition anything not on the frontier line is suboptimal. That's how this type of plot works. I don't really get the point of adding the rectangle. Or shitposting about a discounted price.
>>
>>109653495
BENCHODESLOP
>>
>>109653657
>ayo you have to blow all your budget on $50/Mtok Anthropic model instead of trying slightly less capable model that can still do most things perfectly well for $0.25/Mtok
>>
AI still can't beat autistic humans, AI is for neurotypical goyim only.
>>
>>109653495
>GEMINI 3.7
NOOOO THEY WERE VAGUEPOSTING FOR DAYS
GOOGLEBROS WTFWTF
>>
>>109653495
So is the model actually good or just benchmarkmaxxed?
I wonder how small/cheap we can make models and have them still as good as sota now
>>
>>109653888
The problem is that it'll take a while for efficiency to start being economical because right now the market is too distorted by glowie funding.
That removes incentives for smaller sizes.
>>
Gemma mentioned!
>>
>>109653495
Another vibeGOD victory
>>
>>109654018
There is a company called Cohere that concentrates on models are efficient to both train and run, but it's also funded by Canadian glowies who want edge decision making A.I. for arctic drones.
>>
File: 1776272926650793.jpg (54 KB, 450x453)
54 KB JPG
>>109653563
>>109653572
Didn't have the time to try it yet but if it really is that good this is another FAGMAN Tiananmen Square moment
>>
>>109653495
>306GB for the fp8 checkpoint
I know it's an MoE model, but damn this is not a local model.
Is this is probably too late to ask, but is there are way to extract only the experts that are good for "roleplay"? Or is this a different type of MoE where each model is required to function normally.
>>
>>109653495
>opus above fable
lulz
>>
>>109654186
>not s local model
Just get a couple sparx
>>
>>109653657
>By definition anything not on the frontier line is suboptimal.
Do you apply that logic to everything in your life?
If you had a webslop company would you hire random devs for 60k/year or would you headhunt the guys who got billion dollar contracts by Meta a couple of years ago because they are better programmers?
>>
>>109653495
>>109653563
>>109653572
Unless you have 200-400GB of fast memory at your disposal you're running a model from a server,
Luna (max) is cheaper with a better DeepSWE score
Gemini 3.7 flash (medium) is faster with a better DeepSWE score
>>109653888
>is the model actually good
It's an improvement over DS0731 but it's never beating an OAI sub
>>
File: 5.3 flash.png (67 KB, 1141x476)
67 KB PNG
>>109654451
>>
>>109654557
>no penalty for refusing to answer
Anthropic benchmaxxing by making its model accuse user of hacking anytime it's asked to anything more than "make shitty landing page for B2C landfill in React."
>>
>>109653516
yeah you ain't running the max version locally
you can probably run the 27b version though if you have a mid-range gaming pc
>>
>>109653516
>Qwen3.8 Max
>2.4T tokens
The just-today-released Qwen3.8 Flash Next benches pretty well though and that actually is feasible for a fair few people here.
>>
>>109654582
Is this a brand new cope? Haven't seen it before but anyways, it seemed to work well enough to disprove the Jacobian conjecture
>>
>>109653495
>Luna max along the border of the red box
Meh, good enough
>>
>>109653495
> pareto frontier
is there a reason why this word keeps popping up lately? I never heard it before 2024 now I see it monthly.
>>
File: gpt.jpg (4 KB, 140x140)
4 KB JPG
>>109653495
Goyim...please....I beg you. I'll suck your dick. I will eat your fucking shit. You want me to eat your shit? We cannot let China win
>>
google is embarrassing themselves. we're never getting new pro at this rate
>>
I've been using qwen 3.7 max and plus a lot, since they're free everywhere with various platforms and it's a really good model, doesn't give me bullshit
>>
>>109653589
They are ranked against each other retard, the score is has never been 100, a couple of months ago before Chinese AI came around the "top score" was 80
>>
I only care about this because I want OpenAI and Anthropic to lose
>>
File: skibidi.jpg (98 KB, 404x720)
98 KB JPG
>>109655466

Could you start by not raping your sister anymore, Scam Altman?
>>
Somebody had better put out an even stronger model for another week of free unlimited usage soon then.
>>
>>109653495
>>109653563
>>109653572
I look at these charts evaluating "intelligence" and such and think to myself how much money is spent by these AI companies to optimize for these benchmarks.
>>
>>109653495
SIR, A SECOND MODEL HAS HIT OUR INSIDER TRADING SCHEME
>>
>>109656307
$10M at API prices, about $10k at actual cost on their end
>>
>>109653495
Phew Gemma is safe
>>
>>109656307
Give the average people two LLMs released from this year and they wouldn't be able to consistently rank which one is better. AI model rankings are bullshit.
>>
>>109653495
Please redpill me if this is a nothingburger or not.
Thanks in advance.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.