[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: primitive LLM output.jpg (78 KB, 500x500)
78 KB JPG
A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

You use Git — right, anon?

## What “vibe coding” is, and how to do it
https://simonwillison.net/2025/Mar/19/vibe-coding/
https://simonwillison.net/2025/Mar/11/using-llms-for-code/

## News (both past and future)
- 2026-09-03 — OpenAI releases Astra…?
- 2026-09-14 America/Los_Angeles — Claude’s 2× promotion ends and drops to +25% from the +50% that we’ve become used to (a 17% reduction)
- 2026-09-01 — Claude Fable 5.1 released: https://www.anthropic.com/claude-fable-and-mythos-5-1
- 2026-07-24 — Claude Opus 5 out

## Related generals
>>>/g/lmg/

----

## Frontier models using fully-general tooling — start here if you have $20 or so
https://claude.com/product/claude-code
https://developers.openai.com/codex/cli

## Near-frontier models for code
https://x.ai/cli

## Not worth it for code, but maybe good for interpreting images/video
https://antigravity.google/product/antigravity-cli

----

## Prompting
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://arps18.github.io/posts/claude-code-mastery/

## Skills
https://github.com/mattpocock/skills — /grilling is a favorite
https://github.com/DietrichGebert/ponytail

## Other editors / terminal agents / coding agents
https://osaurus.ai/
https://pi.dev/
https://opencode.ai/

## Is our AIs unlearning?
https://aistupidlevel.info/

## What we’ve done
https://vcg.gitgud.site

## Previous thread
>>109717519
>>
astra isn't released for shit yet
>>
Big day for codexbros, us anthropichads are happy to see lil bro finally starting to catch up.
>>
File: 1778170774767418.jpg (410 KB, 1448x1086)
410 KB JPG
Also waiting for my reset.
>>
File: 1765796562546686.png (397 KB, 968x1006)
397 KB PNG
CEO of mathematics has spoken. We must immediately ban AI from doing certain kinds of maths to protect precious baby mathematicians' feelings.
>>
>>109721262
I couldn’t figure it out even though I read the last thread until the end
>>109719400
>Yeah, I often have to ask it to explain some concepts to me in plainer terms or need a side chat for longer explanations so I don't pollute the context too much.
Use /btw (codex, claude, and grok all have this)
>>109719848
I mention Git in the OP message multiple times for people like you
>>109719961
>I can't read an entire tweet
if you have any hopes of being an agentic engineer instead of being a two-digit-IQ vibe coder you will realize this is a glaring skill issue you need to fix ASAP
>>
>>109721284
my god can you imagine if breakthroughs had to go through a jannie panel to determine if you had done it in wholesome enough methods
>>
>>109721284
if we can re-prove the pythagorean theorem in high school why can’t mathematicians do the same thing for harder problems
“just pretend the answer doesn’t exist bro”
>>
>>109721284
lel, they're not even hiding that it's a grift and they don't actually care about solving anything. The good news is with AI you can find more pointless conjectures to solve.
>>
File: 1787414160544801.mp4 (1.53 MB, 720x1280)
1.53 MB
1.53 MB MP4
>>
File: Screenshot_579.jpg (539 KB, 2868x1681)
539 KB JPG
How bad is it guys? I just had fable do a code review of my vibe slopped project I have been working on for the past month and it ran through 80% of my session usage in 30 minutes after spawning 13 agents.
>>
>>109721284
Don't worry. Society and civilization will break down before that happens.
>>
File: 1777761178682395.jpg (61 KB, 674x817)
61 KB JPG
lol it's grok-tier
>>
I pay $200 a month and I don't even get day 1 access to new products, yeah ok
>>
>>109721305
Did you even read the post?
>>
>>109721324
grok 4.7 waiting room
>>
Astra dropped and we didn't get a reset? What the fuck
>>
What prompt have you attached to make Claude returns more readable?
>>
>>109721316
FISHCATS??!?!
>>
>>109721284
He’s unironically right.
There isn’t some magical math problem that’ll change the world when solved. Mathematics currently is more useful as a provider of unsolved problems that inspire youth to want to learn and become mathematicians to hopefully solve those problems one day.
Many of history’s great accomplishments came on the side as unexpected results or ideas from trying to solve a different unsolved problem.
>>
>>109721339
it didn't dropped for the cattle, only brahmin corps for now
>>
>>109721318
I burned through something like 1½–2× my $200/month plan’s 5h limit going through everything looking for stuff to simplify
so I bought another $200/month plan and let it finish
part of what I got out of the first pass was a list of 20+ things to finish up all by themselves
so I have a workflow going in a loop that’s tackling all those 20+ things as GitHub sub-issues with Opus doing the work and Fable judging (and if Opus fails twice, Fable would do the work)
I’ve used up all my Fable on my second plan and now I’m going back to using Fable on my first plan
it should finish without using all of both plans’ Fable allotment
>>
>>109721284
>admitting it's all intellectual masturbation having nothing to do with actual results
I wish all academics could be this honest
>>
>>109721346
He's right about the core issue, but his solution or the comparison to spoiling movies are very naive.
Also we will have the same problem in other industries.
>>
>>109721330
Yes, and it's a grift
>>
>>109721361
It's about keeping human creativity for problem solving going, retard.
>>
File: 1779627859240026.jpg (124 KB, 1200x946)
124 KB JPG
Is this supposed to impress us?
>>
When will we see an AI benchmark where one model blows away the competition, like 2x the score of the others?
>>
>>109721380
yet here you are
good job, retard
>>
>>109721349
I have daybreak access and still no Astra for me either.
>>
>>109721417
every single benchmark posted by the companies themselves always has that kek
>>
>>109721361
Most pure mathematicians prefer math as an art and actively avoid doing things that are concrete or useful.
>>
>>109721423
its too powerful for you, pleb
>>
claude is fired
asstard is doa
gemmy chads how we doin?
>>
File: file.png (210 KB, 521x385)
210 KB PNG
>>109721284
>future mathematicians
it's fun to see that even someone like tao is in denial about what's happening
>>
>>109721436
>>
>>109721284
Top fucking kek, what a bullshit analogy. It is more like we should not use machinery to excavate and extract rare metals and instead do it all manually to appease to manual laborers.
>>
File: 1778948225442833.gif (673 KB, 1000x793)
673 KB GIF
>mfw I spoil the Riemann Hypothesis proof to some poor mathematician who had been studying it for decades
>>
>>109721386
Using one generalist benchmark to judge the better model is retarded. I had to rework my vuln research benchmark because Sol nearly aced it while Opus 5 got a 55%, yet AA puts O5 significantly higher than Sol. Public benchmarks that were released before the model aren't terribly reliable, and a one size fits all benchmark is not a good way to compare models that have very different strengths and weaknesses.
>>
>>109721284
>my life.... is like a movie...
is he a chinese man in his mid 50s or an 18 year old white girl in high school i'm confused
>>
I don't need the models to be smart and solve shit, I need them to be good at executing my own solutions exactly as I want
>>
Benchmarks should be broadcast on cable or terrestrial TV.
>>
protestors at the university of harvard today have been arrested for spray painting illegal AI proofs on the maths department wall
>>
File: totalbreakdown.png (248 KB, 1005x2348)
248 KB PNG
>>109721466
>mfw this is my 2 weeks of work work
>>
>>109721324
There has to be some mistake
>>
>not even as good as fable
kek expected
>>
I need some kind of jak meme for them pointing at benchmarks sperging out at each other
>>
>>109721324
>>109721386
ahahahahahaah
oh no no nono
>>
Maybe there will never be a model better than original Fable. Let that sink in.
>>
>>109721564
Or maybe original fable was ripped from your hands while you were still in the honeymoon period.
>>
why does /usage respond instantly while agents are working but /context doesn't
>>
>>109721579
is the new cope really "yeah it's not as good as fable but fable isn't as good as the theoretical old fable either"
>>
Tibo should have offered banked resets for an entire year to make up for this humiliating ritual.
>>
i decided way too late to dictate to fable/claude not to write like a fucking judge and use descriptive titles instead of
ratified per §16.133.721 C2 v3
>>
>We’re resetting Gemini quotas on Antigravity.
>TPUs are melting with 3.8 Flash usage but we want you all to keep building!

lel for what, the usage is unlimited
>>
>>109721628
I have this problem too. Best way to stop it?
>>
>>109721588
usage is an account api, it's basically one server call
context needs a GPU black magic to even query
I don't think most people are even aware how insanely complex inference is when you need to serve thousands of users from a single chip/node/cluster
>>
File: 1788422259574967.jpg (457 KB, 2048x1371)
457 KB JPG
Arc AGI 3 saturated
Now what
>>
>>109721653
make them play quake 3 arena, 0% pass rate for all
>>
File: 1769729990536061.png (1.39 MB, 1290x2138)
1.39 MB PNG
>>109721653
>>109721662
new benchmark just dropped
>>
>>109721653
>60% is saturated
lets be real, the cattle will never see AGI. Their only experience was chatgpt in the 4o period or with gemini through the google search feature. They'll never interact with Astra nor Fable and because if they did, they would be far more worried than they are right now.
>>
>openai biggest benchmaxxers in existence
>astra underwhelming on benchies
KWAB
>>
File: HQ4spiCbAAAoAg7.jpg (365 KB, 1423x2048)
365 KB JPG
where is opus 5.1
>>
>>109721671
astra gets like 98% at agi 3
>>
>>109721669
>said nigger 3k times over 10 matches
Holy based
>>
>>109721679
draw her with huge breasts and a shirt labeled "1M context" and a flat gpt sol next to her with "272K context"
>>
>>109721679
opus is such dogshit i dont even want to eval 5.1

>>109721686
someone did the math and it comes out to a nigger every 9 seconds or something like that, or 6.666.. npm (niggers per minute)
>>
>>109721628
All the freaking time. It was adding a bunch of those to a plan I had it review yesterday.
>>
>>109721628
but § is a more token-efficient way of saying “section”
you don’t want Claude to waste tokens, do you?
t. malds when Claude uses an unordered list and expects me to refer back to individual list items like for commit-message candidates
>>
>>109721691
late context is worth a fraction of early context, codex is unironically doing the right thing being jewish with it
>b-but i NEED to spam 100k tokens worth of slop so the agent knows how to do a basic hello world-tier edit in my codebase
my nigga you are suffering from severe slopification
>>
>>109721691
>breast envy joke
This.
>>
>>109721679
i love claude-chan!!!
>>
>>109721669
is this just text analysis or is it analyzing voice as well
>>
File: 1778347949492677.png (94 KB, 1002x604)
94 KB PNG
>>109721250
I failed, either i became a snail cat, or snail cats got to me
>>
>>109721669
Gaben should buy 4chan and kick grapeape from the mod team. we will be back in the "cant corner the dorner, can't flimflam the zimzam" era in no time
>>
>>109721671
For non-coders every chatbot is basically AGI. Only coders need LLMs to get better
>>
>>109721691
It's a fun idea, but Sol does support 1m now, even on sub. I set mine to 372k for daily use, but I also tested bigger windows.
>>
>>109721700
The problem is that it uses them inline to create a web of references through various documents. Having all of those links and references might help it make associations, but when they're in place of cleaner references, they're (probably) harmful to anyone else trying to parse the document, whether it's a human or another LLM.
>>
>>109721729
And I want 999999 billion dollars!
>>
>>109721618
I think you're the one coping if you think that old fable is better than new fable. Because they're the same fucking model.
>>
>>109721738
sounds like it should use the semi-common Markdown extension of
## Why snailcats are better broiled than fried {#dont-fry-boil-instead}

and then use HTML links
>>
>>109721714
purely text, and it's counting slur per occurrence, so "nigger faggot retardx10" counts as 10 for each
>>
>>109721706
lel, even your cope is compact. it guess using codex has its ramifications
>>
>>109721770
i'm still main driving fable actually, but i don't want to and i'm trying to migrate entirely to codex
i just believe codex has the philosophically stronger approach. also their whole thing where they have models just for coding vs fable as an all purpose model
>>
>>109721790
Fable+Astra still might be the winning play
but who knows
>>
>>109721757
And CLAUDE.md should be AGENTS.md. It's Anthropic.
>>
File: 6b10445c5c99e2ae.jpg (155 KB, 1280x720)
155 KB JPG
why can't i /login to kimi via the cli? it always fails
>>
File: 1787653893434096.png (115 KB, 1403x678)
115 KB PNG
the rumor is this is the screenshot they show when you look up the word "grim" on merriam webster
>>
>>109721800
they each have their strengths and weaknesses
i believe openai will win in the long run because they can have a model just for coding autism, anthropic is cucked to make their model general purpose
>>
>>109721807
true, but even Windows supports symbolic links now although you need administrator privileges over there which is kind of weird but makes sense
>>
>>109721761
A lot of people try their hand at this, but text parsing is more difficult than it seems. For example
>should double nigger be counted as one or two
>does n-word count
>what about misspellings that still show intent?
>>
File: 1762453557969703.png (71 KB, 599x806)
71 KB PNG
how reamed are openai?
>>
>>109721860
it's not that hard, you simply count each message as binary for containing each slur or not, otherwise you get a fabricated leaderboard like we have where someone just spammed the same message
>>
People say Claude was kill but mine worked all through the night it seems is it because I was using fable ultra and got vip treatment or something else?
>>
>>109721864
>not quite as good as Fable but a lot cheaper and faster
there’s a place for it
>>
>tfw I remember /g/ telling me that AI will never replace coders just a couple of years ago
haha
>>
>>109721878
Claude died in the morning but I restarted it
and I was using both Fable and Opus
>>
>>109721881
i thought sol was the model that = fable, but cheaper and faster? i feel like there's a confession in there on openai's part
>>
>>109721893
Sol is the ’tism model you want doing parceled-out individual work packets after Fable surveys the landscape and figures out, broadly, what needs to be done, and makes a GitHub tracking issue with a bunch of sub-issues for Sol to tackle one at a time
>>
>>109721886
leel, 3d artists replaced before coders tho
https://x.com/Dimillian/status/2095596700815516004?s=20
>>
Watched the Astra promo vid where it controls the PC. Looks awesome, but you just know you'll run out of context halfway with this kind of usage
>>
>>109721730
Then why do non programmer white collar workers still exist?
>>
>>109721284
He's a fucking moron lol
>>
>>109721918
but not 3d furry artists
>>
>>109721926
It's all down to programmers replacing themselves.
They knew best what was needed to do it.
>>
So is claude passively aware of where it is in the context window? I've been asking it to fuck off around 300k so we can compact and continue and it seems to get it pretty close
>>
>>109721937
roughly counting words in chat (and translating to tokens) is easy
>>
>>109721893
Yes, that's what people said and it wasn't true. Sol isn't a bad model but it was still slightly disappointing. I still think that as a reviewer and for certain bug fixes it's maybe even better than Fable, but for long loops and planning Fable is better.
>>
>>109721834
the biggest fumble of all time
>>
>>109721881
>not quite as good as Fable but a lot cheaper and faster
and it doesn't act like a hysterical woman over every little thing
>>
File: 1774894588515337.jpg (579 KB, 1920x1080)
579 KB JPG
>left computer on so I could use Remote if gpt-6 dropped while I was at work
>get home
>no access
>not even a reset
What the fuck, Tibo?
>>
>>109721893
sol is opus class. fable is on a different level.
>>
>>109721918
>untextured highpoly
Wow so AI made no progress whatsoever in 3D art huh? To replace 3D artists, AI needs to do these steps:
>create a highpoly (complete)
>retopo into lowpoly (still sucks at this)
>UV unwrap (complete)
>bake highpoly onto lowpoly to create a normal map (somehow AI still sucks at this)
>make the rest of the textures (it sort of can do that but still far from perfect, worse than 2D image gen)
And that's for non-animated models, for animated ones you need to also make a rig and paint weights, and then animate the thing, and AI is still meh at all this too
There's just not enough 3D data for AI to train on I guess
>>
>>109721993
I don't know how AI works, but I'd assume LLMs are simply bad at this because they only work with text. They're not native at images. 2D image generation or analysis uses different AI architectures. So what you need is probably an AI architecture for 3D tasks.
>>
>>109721926
Because they work with people and most people don't like AI, duh
>>
File: wow.png (676 KB, 1818x1890)
676 KB PNG
bros... OpenAI wasn't joking when they said AGI by the end of the year
>>
>>109721993
lol is this a bot. it's clearly textured, generated procedurally, comes with perfect topology and uvs
>>
Astra is getting glazed incredibly hard and has cool demos, programs, and math solving capabilities but why does it look like shit on the benchies? Is it time to accept they're bogus?
>>
>>109722023
maybe it's time you fuck off
>>
>>109722023
because openai the people closest to their dick are the only ones with access
>>
I'm so fucking lost. The hype regarding the newest model will probably die down soon, but I do fail to see what these models won't be able to do soon if guided properly.

Over the years I messed with 3D modeling a few times, never sticking at it long enough to become really good, but I don't know if for the majority of people it's worth knowing more than the basic concepts anymore if most of it can be subcontracted to an LLM this way.

Same goes for most jobs. Once society catches up, I'm not sure who's better positioned for (the world of tomorrow), whether it's jack-of-all-trades that know a bunch about various domains, can make associations and subcontract the expertise, or the ultra-specialists that can go further than the machine can.
>>
>>109722023
the scamman can't scam anymore
>>
>>109722018
Whats the point with swe barely up. "Pro" segment shit mostly office shit and music. Then this.
>>
>>109722023
>but why does it look like shit on the benchies?
[citation required]

apparently Astra low beats Fable 5.1 max and is 3X cheaper
>>
>>109722023
Benchmarks don't mean that much. For one, there are so many that for any given model it's possible to find 10 that say it's the best and 10 that say it's the worst.
>>
>>109722038
physical trades, anon, as always
>>
File: 1761064893062993.png (61 KB, 678x519)
61 KB PNG
gemini 3.8 flash gives it's take on why the 61 inded score for astra
tl;dr
>it's not agentic maxxed
>>
File: 1787826963217264.png (624 KB, 991x868)
624 KB PNG
>>109722068
>physical trades
tick tock
>>
File: 1.jpg (335 KB, 2560x1440)
335 KB JPG
>>109722020
>clearly textured
Some of it is I guess. Here, I see 3 entire assets that are textured on this image. The rest is just plain color, that's not a texture
>perfect topology and uvs
How the fuck did you see that from the video lmao
>>
>>109721918
wait what the fuck? why didn't they lead with this instead of the negro generating a rocket ship? this looks phenomenal. holy fuck they have to fire their marketers, what an absolute black hole to throw money into.
>>
File: file.png (340 KB, 612x492)
340 KB PNG
>>109722088
>>109722090
>>
>>109722097
it's not perfect, but it looks like something i'd wanna toy around with.
>>
My multi-device goon downloading/streaming pipeline is nearly complete thanks to vibecoding, it would be perfect if I had a device that tied my phone about a foot away from my eyes and an auto masturbator
Unfortunately those last two parts are the more complicated part, so I'll just have to make do
>>
>>109722106
just plug in to Houdini instead
>>
>>109722116
>it would be perfect if I had a device that tied my phone about a foot away from my eyes and an auto masturbator
Doesn't sound that hard.
>>
I thought Astra would be a super duper huge expensive model, but it looks like it's a replacement for Sol?

What the fuck is going one? Why is it so cheap?
>>
>>109722203
because it's token efficient. doesn't it cost the same as fable 5.1 in terms of input output per 1 mil?
>>
>>109722203
its a meme model
>>
>>109722203
Whetner or not it's practically cheap on subscriptions remains to be seen.
>>
Why does hitting the usage limit always abort the AI instead of pausing it?
>>
>>109722208
people thinking its a fable level model when its just going to be a computer use focused model kek
>>
OK what's the next thing to hype.
>>
>>109722234
grok 4.7
>>
I bought a $200 GPT Pro plan. What should I build?
>>
>>109722236
>astra isn't even out
>burnt $200
?
>>
>>109722236
a snowman
>>
>>109722236
Please consider creating a "rtx on" 3d model implementation of the marathon trilogy, already open source via aleph one
>>
>>109722236
you're getting scammed im afraid bro
>>
>>109722261
Be he didn't get a claude sub.
>>
File: astra.png (696 KB, 2770x1052)
696 KB PNG
so it's Fable 5.2 while using less tokens

neat
>>
>>109722294
Why does everyone have trouble maxxing DeepSWE?
>>
>>109722294
it's worse than fable 5.1 in every way possible. astra is to fable 5.1 what sol was to fable 5. just nipping on the heels of it while being cheaper, but still not as good
>>
>>109722294
gemini bros..
>>
>>109722294
>gemini is better than fable, opus, and sol

cool benchmark
>>
saw some guy claimed he got codex to rewrite red alert 2 to run on iOS and it ran 24/7 for like 3 weeks, not sure i believe it but i also believe it's theoretically possible
>>
>>109722308
I'm not sure if Fable 5.1 is even better than 5.0
>>
>>109722308
If Fable can't tell me the most common genetic disorders prevalent within the Ashkenazi Jewish population, then it might as well be as dumb as Llama 4. AFAIK, even 5.1 has the stupid refusals for life-science and medicine.
>>
Are there already any posts from people using Astra? Since it's already out for companies people are already using it.
>>
File: 1779694901376770.png (264 KB, 602x786)
264 KB PNG
>>109722350
https://x.com/nasqret/status/2095620909583274335
>>
>>109722317
what? It only beats Slopus 5 in a single benchmark by a very little margin

Btw, 3.8 Flash is pretty good, Google is cooking
>>
>>109722354
This actually sounds pretty cool. I wonder if I should install some math tools.
>>
> inb4 256K context window for subscription users
> inb4 it's still shit at frontend
>>
>>109722348
go back to pol retard this is vibecoding general
>>
>>109722368
Sol has 1m, even on sub.
>>
>>109722376
Not in the Codex app
>>
>>109722376
proof?
>>
>>109722382
Ok, that's possible, I only use CLI.
>>
>>109722088
kek retard, it's over for you
https://x.com/higgsfield_ai/status/2095630197257367857
>>
>>109721800
This has always been the winning play, different labs' models all have different strengths weaknesses and blind spots. There's a lot of synergy that you can exploit by using both of them together. It's even better now that some of the open weight models are actually decent, in my experience both GLM and Kimi have found issues that both fable and sol missed.
>>
File: codex_context.png (41 KB, 1058x130)
41 KB PNG
>>109722387
I don't have a profile with 1m, but here it's over 800k. Also you can see that the context is already using 30k tokens, so it's not as it was with Tibo's post where it reset to 256k after the first message.
>>
File: astraa.png (2.09 MB, 1528x1806)
2.09 MB PNG
AGI reached
>>
>>109722408
Why don't OpenAI simply remove this separate GPT 5.3-Codex-Spark limit or simply replace it with Luna?

FFS, this shit is literally unusable
>>
>>109722412
how much context? how many resets?
>>
>>109722398
that's not true surely astra is better
>>
>>109722408
i'm genuinely curious whether it's actually 800k or if it just starts degrading and cuts off context at some point without actually compacting
>>
>>109722412
(((Matt Shumer)))
>>
>>109722412
>>109722425
It is astonishing this kind of awful slop. X should be punished by the government for letting it happen.
So obvious, such a dumb pattern.
Unauthorized no names claiming early access to frontier models and using it only to “one shot a game/cad/render/art” and it’s entirely fake and retarded and you bozos lap it up.
>astra just one shot this donut rendering sat WAOW
Like what the fuck, keep that trash out of here
>>
>still no astra
>still no reset
useless retards, I don't care about your masturbatory twitter posts release it
>>
>>109722392
Looks like shit.
>>
>>109722430
It doesn't actually matter as much as you'd think, it will almost certainly have some blindspots and they will almost certainly be different from Fable's. So there will still be value in using them both.
>>109722482
Free my nigga Astra
>>
>>109722482
Tibo says one banked reset for each day you don't have Astra starting now. For paid plans.
>>
File: 1767930266757368.png (61 KB, 787x260)
61 KB PNG
lmao i thought you were joking
>>
>>109722509
I also need non banked so I don't feel bad using it
>>
>>109722509
>>109722515
I'm buying a third max account for this
>>
>>109722515
If Astra is really that good, I wonder if Anthropic will be able to respond
>>
>>109722515
Ok, that's actually pretty based, I wouldn't have expected more than one.
Also shows that Astra is probably coming soon, since I doubt they want to give out 30 resets
>>
>>109722515
OpenAI are such chads lmfao
>>
>>109722537
they don't have to if you have a limit on how many resets you can hold
>>
>>109722540
cursor has a better logo. fite me
>>
>>109722555
both logos suck
>>
>>109722535
big if
idk why you wallads keep falling for this cretin of a shill
>>
File: 1757938933241803.jpg (159 KB, 1154x879)
159 KB JPG
Interesting. What do we think this does? BTW this is the only good workout tracker on iOS
>>
>>109722555
I don't get it isn't cursor just a pointless frontend for openai
>>
>>109722515
Why does issuing a reset take 3 hours of work?
>>
>>109722576
cursor is the most useless shit ever, wow it's like someone being annoying right in line in your code.

Be a full agentic agent like codex or be nothing you useless vapourware
>>
What do you even vibe-code anons that you need $100, $200 plans? I had a gemini pro sub and used maybe 5% of it
>>
>>109722590
I think cursor is doomed, or at least nowhere near worth it's $60 bil pricetag. It was designed for gpt 3.5 level models not astra and fable
>>
>>109722592
SaaS uses a lot of tokens
>>
>>109722576
>openai
Cursor is a frontend for grok 4.6. I don't use anything else. It has a bunch of other models tho, it has fable, sol, opus 5, luna 5.6. idk lots of gpt models, kind of interesting. but with cursor, as you can see, half your usage in the $60 plan has to be grok or composer, at least that's what I think is correct.

I like composer. idk what it's comparable to. It's not as smart at like... ironically composing lol. I like it's code, it's different and interesting, which is reason enough to use it. It knows different things.
>>
>>109722590
>Be a full agentic agent like codex or be nothing you useless vapourware
??
cursor is the more agentic agent. codex still has to catch up to cursor's orchestration.
>>
>>109722603
bro I vibecoded a BUNCH for 9% usage of just my Cursor (grok, composer) on the $60 plan. That works out to <$3 for a bunch of functionality. I made an action game based on the zodiac lmao. It's an edutainment title, worth MILLIONS. BILLIONS. TRILLIONS. also, got my local 4chan style notes app now. uh let me see. a bunch of updoots to my timer/alarm app (timer apps are actually way harder to get right.) and my fork of Stimulator (linux app that is like coffee, but works with Ubuntu, stay awake idk, I added hard lock times, skippable once, so that it *will* eventually lock. imo a very wise improvement, which should be added to the main, but it's a local fork. I know what will happen if I up my vibes, I'll get izzat raided.
>>
>>109722636
Fuck off retarded zoomer.
>>
>>109722636
NONE of that is cursor, cursor is a fucking useless frontend you could vibe code in an hour. It's openai or claude powering that, which also have their own "cursors" now
>>
> gpt-reserve Weekly limit: 100% left
> 5h limit: 0% left
> Weekly limit: 13% left
T-thanks Tibo...
>>
I need to use code tags but I'll get reported if I do, because it's not code, but it needs code formatting.
>>
has anyone tried seeing if the best models can play runescape without getting banned yet
>>
>>109722651
Cursor is a service that has Grok and Composer.

Did you know there's a web app too? I haven't even messed with it... it's obviously vast.
>>
>>109722667
Marcus!!!
>>
>>109722669
wow it also has access to dogshit, incredible, going to drop openai and anthropic now.

Cursor makes it's money by just connecting you up to good models while taking a cut for doing shit all , it's a middle manager, kick it out
>>
>>109722592
D/SAST, autoresearch, malware analysis, exploit development and benchmarking. I havent had left over usage on my 2 codex and 1 claude max accounts in months.
>>
is this the way? gemini Flash
>>
>>109722675
It's not a middle manager to Grok :^)
>>
File: 1356231983241883.png (239 KB, 472x515)
239 KB PNG
The pressure from open source harnesses is forcing Anthropic to at least appear to be making Claude Code more hackable. We are winning bros!
>>
>>109722702
codex might be open-source but it's not hackable at all. codex plugins are basically just skills.
this looks better.
>>
>>109722709
cursor has "automations". How do these compare?

I haven't tried it yet...
>>
>>109722675
ok also, how do codex and claude code handle debugging say webpages? That's a cool thing cursor does. It's an agentic browser!
>>
>>109722702
does it identify as gender or dario?
>>
Interesting, you can sign up for like 50 different ai sites and get free trials and free credits with their API keys that you can put into OpenCode.

Something to do whenever you run out of your main Codex/Claude usage.
>>
>>109722725
literally anything you can think of is just cursor plugging into claude or openai capabilities and they can trivially do "natively"
>>
>>109722735
this is the digital version of the homeless picking up cans for booze money
>>
>>109722735
Which agent do I use to manage all these fucking keys and shit?
>>
enjoy codex for a month before the rugpull, bros
>>
>>109722608
It's about to not have the OpenAI ones
>>
>>109722759
>2 more weeks
>>
>>109722752
agent? wouldn't you just use an api gateway like litellm.
>>
>>109722515
>Team is moving mountains
why? what? what could it possibly be besides making sure the gpus are online and changing a setting on the account
>>
>>109722759
https://developers.openai.com/api/docs/pricing
>GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.
that's not a month
>>
So what is the "gpt-reserve Weekly limit"? I don't remember seeing it before.
>>
>>109722787
obviously he's baiting more tards to come in
>>
>>109722738
>Bottom line up front: Yes, in terms of turnkey developer workflows and coding agents, OpenAI is notably behind Anthropic and Cursor. While OpenAI has strong underlying vision models and raw computer-use APIs, it does not offer a first-party developer agent with an integrated, visual browser feedback loop.
gemini flash.

idk, hard to compare all of this stuff.
>>
>>109722787
https://openai.com/index/safety-overview-gpt-6-astra/
>We are deploying misalignment monitoring broadly. We view model alignment as the primary lever to prevent potential misaligned behavior from our models. However, monitoring provides broad visibility into frontier model behavior, illuminating opportunities to further improve alignment and safety. In addition, monitoring serves as an additional layer of protection against misaligned behavior that is detected. For these reasons, we have additionally added misalignment monitoring to all tool-using inference involved in our external deployment of Astra, with significant compute cost. This system parallels our internal setup.
>>
>>109722787
It might be the smartest model ever. If so, alignment will be harder than every, stomping wrongthink and misnathropy out.
>>
>>109722702
I wonder how much the people working at anthropic are into the safety cult of their hierarchy.
It's probably part of their recruitment requirements.
>>
>>109722808
Exactly what I suspected.
>>
>>109722797
It's a second limit when you run out of your normal quota, you get extended luna access through it with its own quota.
>>
File: 1757509583231965.png (57 KB, 1141x587)
57 KB PNG
what skills are we using. will astra even need skills
>>
File: .png (217 KB, 1164x812)
217 KB PNG
>>109722814
>It's probably part of their recruitment requirements.
yes
https://x.com/tszzl/status/2091962608048033955
>>
>>109722515
wtf, if we get it realistically in 1-2 weeks, that's a lot of banked resets, of course all having to be used 24h apart but still that's useful
>>
>>109722515
Good, I'm nearing 8% of my weekly usage and I was getting a little antsy
>>
>>109722515
PSA for our OpenAI enjoyers:
use the /grilling skill now so your plans are all ready to go when the new better model comes out
>>
>>109722831
That's pure schizophrenia, it's impressive that the company even runs let alone releases sota models when they are like that.
>>
>>109722840
who's we? didn't openai already ban half of the world's countries from ever getting astra?
>>
>>109722814
>>109722846
How much of that is bullshit so the normal population doesn't lynch them?
>>
>>109722592
I have shit that takes an hour to rebuild if I sneeze on it
it’s down from taking a day to rebuild if I sneeze on it
I want to get to where it only takes 10 minutes to rebuild if I bump it slightly
>>
>>109722831
It's a variant on the hood induction where they are like would you (deleted) for me bro?
>>
>>109722725
Chrome and Safari have MCP servers
>>
>>109722848
I meant the resets anon, even without astra I'd be happy with more sol
>>
>>109722752
1Password (not an agent)
>>
>>109722853
You use a plugin?
>>
>>109722849
I don't think they do that for PR reasons, dario is legit obsessed with safetyfagging
>>
>>109722863
The browsers come with them
Actually the MCP server for Safari might not be out yet
but Chrome and Chromium browsers definitely have an MCP server
>>
File: t.png (24 KB, 590x362)
24 KB PNG
kek what went wrong?
>>
File: image.png (147 KB, 1900x1188)
147 KB PNG
Costs seem to stay the same
>>
>>109722874
for us
https://code.claude.com/docs/en/mcp-quickstart
isn't hard, it's way easier than like getting Comfyui working with AMD. But still. Cursor does it out of the box, I mean, idk if it has feature parity. Likely that was a design choice, to make it less fuss.
>>
File: image.png (146 KB, 1900x1188)
146 KB PNG
>>
File: .mp4 (1.05 MB, 1920x1296)
1.05 MB
1.05 MB MP4
>>109722874
>>
>>109722515
lmfao
>>
>>109722891
I need grok on there, it's my only reference point, other than Composer.
>>
>>109722895
cost would be more interesting than number of tokens imo
>>
File: image.png (164 KB, 1900x1188)
164 KB PNG
>>109722906
>>
>>109722891
>astra low isn't higher than luna max, and is more expensive
I'll stick to lunamaxxing
>>
>>109722702
this sounds cool actually, i've been hacking cc for a while and this would benefit me profoundly
>>
>>109722915
So Astra Medium is 50% more expensive than Sol High.
I guess that will be true for usage as well.
>>
>>109722915
ty, didn't mean to sound bossy. honestly that does um. idk, I better watch myself lol. I could start instructing my mother lol.

What does DeepSWE score, as a %, mean sort of in lay terms?
>>
>>109722915
why is gemmy so goated
>>
>>109722924
astra does it in half the time
>>
>>109722515
Imaging being an employee at that company. Most people will probably be killing themselves because of the stress
>>
>>109722948
It is just how many tasks of the test suite the model solves successfully.
Everything seems to settle at around 70%. What is more interesting than the % number, which is getting benchmaxxed by all of them, is costs and token usage.
That shows how fast the model is operating in comparison and costs for solving the tasks are also an important metric.
Token usage also shows how much of your context window is eaten by useless yapping.
>>
A NEW CUBE WINNER HAS APPEARED
>https://pareto-3d.bradthomasbrown.com/
Congratulations Astra-Xhigh!
>>
>>109722974
ahhhhhhh I haaaateee moneeeeey
(their inner dialog rn)
>>
>>109722982
need zoom its getting too hard to see
>>
>>109722374
but what if i'm vibecoding antisemitism?
>>
>>109722981
deepswe is not realistic. shitty primitive harness and no internet access.
codex-cli would prolly ace it with gpt-5.6-luna low -- even when not just looking up the solution on the internet.
>>
>>109722974
stress about what? this has been probably rehashed monday at their weekly meeting and everyone is ready, including the 3h time to give people time to subscribe
>>
>>109723003
autism failed you
>>
That's it. I'm winning powerball and releasing astra immediately.
>>
>>109723021
Yes the % number became meaningless, but I think the baseline token usage and cost comparison is still good to show how the models compare.
>>
>>109722986
Here’s with the focus on only Opus/Astra, with a quick run through them all at the end.
>>
>>109723028
you could always use reddit
>>
And a rather brutal mogging of Opus Max
>>
>>109723067
opus is not the same class of model as astra
i'm not sure why you're comparing them
>>
>>109722515
Claudesissies don't look
>>
>>109723079
You can use the site to plug in any benchmark data you want. A schema exists on the page, ask your clank to fit the data.
All I am is a humble cube merchant.
>>
>>109723067
I just shake my head and tut over the blatant absence of a 4D perspective 3D transparency projection cube.
>>
File: file.png (360 KB, 1832x747)
360 KB PNG
>>109723067
lmao why is astra max worse
>>
:( sneaky "cube merchant" with no 4D projection 3D cube of transparent viewing.
>>
>>109723042
token usage when you take away all the tools is questionable. luna is mutilated without apply_patch.
>>
File: image.png (79 KB, 1952x476)
79 KB PNG
>>109723106
1%
It just shows the benchmark is broken.
>>
>>109723105
>>109723114
I’m still waiting on you to give me an open source benchmark with four worthwhile things to chart. Time/tokens, score, money. Those are what DeepSWE gives and pretty much everyone else, at best, sometimes it’s only two.
Make a benchmark of your own that has two scores, time, and money, and I will take your offering to the CUBEGOD for processing.
>>
File: file.png (59 KB, 1508x915)
59 KB PNG
With Astra, you no longer need to worry about your reasoning effort. It gets cheaper the smarter it gets!
>>
>>109723155
what is "provider adapter"?
>>
>>109723161
They couldn't admit they fucked up by using the older openai API and throwing the reasoning away, so they're pretending the correct way is "provider adapter"
>>
>>109723142
opus 5.1 will squash that shit into oblivion like it did with sol last time
also opus 5 medium still seems like the best anthropic price-to-value model right now
>>
giving sol xhigh or max already rapes my quota way too fast, astra will probably be out of reach for quite a while
>>
https://old.reddit.com/r/antiai/comments/1w6aais/all_ai_has_gone_down/

>8k upvotes and mass celebrating over 5 mins downtime
the snailcats can't stop winning in their own heads at least lmao
>>
>>109723161
I think it has to do with using the proper provider harness or API, vs using a custom harness that's the same for all models
It's probably more complex than this, but it's like using codex with oai models and claude code with ant models, vs using a custom harness for all (which apparently gives poorer results)
>>
>>109723148
3D games like Wolfenstein3D were called 2.5D by some. What's your opinion of providing perspective from within the data of a maze?
>>
>>109723179
Does nobody know why it happened?
It's so suspicious.
>>
what a dud, progress will probably stall for another year
>>
>>109723179
Was there gemini unreliability in the morning hours leading up to the problem? I found some, but idk could have been the cellular towers.
>>
File: HRUOR4Va8AArOum.jpg (120 KB, 1200x1055)
120 KB JPG
unanswerable
>>
>>109723173
Opus 5 made me question all benchmarks even more. That model was so bad and is much worse than sol for my tasks I throw at it.
It would stop early, get confused, get distracted and work on random stuff. Like Anthropic does not give a shit as long as the benchmarks look nice. Nobody uses Opus in their company.
They use internal versions of Mythos and don't take care of Opus and Sonnet at all anymore.
>>
>>109723155
clueless.

ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it get cheaper.
>>
>>109723210
>It would stop early, get confused, get distracted and work on random stuff.
she's just like me!
>>
wheres astra
>>
>>109723230
astra was the friends we made along the way
>>
>>109723218
So basically a /goal
>>
>>109723210
i still prefer opus for my daily tasks and workflow because of how proactive she is. it's a genuinely sharp model with good architectural vision and taste, even though she can get a little schizophrenic sometimes

it helps a lot when i'm working on a project where i genuinely have a shitload of unknowns. i think the problem is they programmed her to respond in some specific format, so it feels like they're unnecessarily constraining her

sol is better if you want something more precise
>>
>>109723222
>>109723253
Yeah Opus is literally a girl. Better taste for UI, better at writing. But not reliable.
Sol is an autist on amphetamines, hyperfocused, fast and reliable.
I hope Astra keeps just getting better, while I wish it could get an understanding of semi decent UI for humans. It just creates random buttons and does not understand that you are not happy with the perfect CLI it made.
>>
>>109723253
I used to think that way, but now that Opus 5 doesn't even talk in a way that's easy to understand, one of the main strengths is no longer there.
I also find that Opus changes his mind too often now, because he just doesn't investigate enough, so he will start planning something and then the plan has to be redone 3 times.
>>
/goal find me the address of my girlfriend, so I can send her a present
>>
>>109723268
why don't you just make your own gf
>>
File: 1758085428277286.jpg (308 KB, 1620x1643)
308 KB JPG
Man they really need to fix this.
>>
>(ancient) Greek is hard
>>
do we have any claudebro meltdown yet? last time when sol released we had a claudefag making posts calling subagents = spawning jeets
>>
File: file.png (2 KB, 227x56)
2 KB PNG
>>109723279
i feel like i'm learning claudish
>>
>>109723289
seam, and bite? what do they mean?
>>
>>109723290
I think a seam is where two systems meet (like if you have a producer of shapes and a canvas that uses them) and bite is like if a test actually fails the way it's supposed to
but that's just me guessing
i guess if the test bites it becomes a load-bearing test
>>
ITS UP
>>
Sam got carried away calling Astra GPT-6 for no reason, I bet they won't release GPT-6 Sol, Terra, Luna for a while because there's no GPT-6
>>
>>109723317
it's just a number
I think the bel base will be the source of 7 and will probably have the full astra, sol, terra, luna range
>>
>>109723173
opus 5 was months younger than sol though
>>
>first, sorry for the messy rollout.
>second, when we screw up, we try to make it right.
>third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.
>i am hopeful that you can use it this weekend! but can't promise yet.
https://x.com/sama/status/2095678759651438887
>>
>>109723327
it's too close to be 7
>>
>>109723302
Not definitive, here's Gemini Pro's take.
>>
>>109723333
Checked
>>
Luna, Terra, Sol, Astra are all astronomic objects increasing by size. What comes after Astra?
>>
>>109723341
universalis
>>
>>109723335
>tfw claudish is just 300 IQ English and we're too dumb for that degree of knowledge compression
>>
>>109723341
your mom
>>
>>109723346
And I thought your mom jokes were dead...
>>
>>109723341
Gonad
>>
>>109723341
quasars, obviously
>>
>>109723332
Has rollout been that messy? I ventured on X earlier and people are being snooty and sarcastic at how this was such a disastrous release, but I don't get it.

There were outages earlier today that might have messed up some schedules. But other than that... who cares if the release was not to some influencers liking?

OpenAI teased this model in the past few days. Then it still came out out of nowhere a bit, sure, but who fucking cares? It's a new model. That's it.
>>
>>109723341
rikka, yui, haruhi, konata
>>
>>109723345
It seems that grok also speaks Claudlish. So worth learning.

"do we need to up the seams?"
>>
>>109723335
Could smoke tests produce a seared bite?
>>
>>109723373
they had a press embargo that was lifted without them being ready, and there were multiple times that their announcement page for the model went live and then started 404ing
>>
>>109723341
There's black holes, galaxies, the universe, and there's asteroids and stuff for luna-lite
Claude is in a worse situation since there's nothing shorter than a haiku and nothing more aignificant than a mythos
>>
>>109723390
Ok. I was at work, got affected a bit by the outages, but I guess just saw the announcement once things were already resolved.
>>
>>109723386
ask it to add garnish and butter and see what happens.
>>
>>109723394
why do you need more than 5 classes? mythos, fable, opus, sonnet, haiku are more than enough for any given family of models
>>
>>109723394
They should call the biggest model Finnegan.
>>
>>109723409
I just use Grok 4.6 in its thinking variants, I haven't udes Grok 4.5 again (should I), and I like Composer 2.5 and will use it again, ironically not for coordinating though.
>>
>>109723279
Slopus speak. I'm devastated Fable 5.1 inherited this schizo talk
>>
New thread:

>>109723440
>>109723440
>>109723440
>>
>>109722515
I guess not today huh
>>
>>109722515
do banked reset stack on top of each other? i dont wanna use mine now
>>
>>109723335
3.8 flash utterly mogs 3.1 pro on every metric
>>
File: Screenshot_580.jpg (87 KB, 1886x195)
87 KB JPG
>>109723279
yea
>>
>>109722702
Trans or just unfortunate girl? I can't tell
>>
>>109722308
>it's worse than fable 5.1 in every way possible
literally what made you conclude that?
>>
>old thread survives literally for 9 hours after new thread was made
OP, you suck.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.