A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.You use Git — right, anon?## What “vibe coding” is, and how to do ithttps://simonwillison.net/2025/Mar/19/vibe-coding/https://simonwillison.net/2025/Mar/11/using-llms-for-code/## News (both past and future)- 2026-09-03 — OpenAI releases Astra…?- 2026-09-14 America/Los_Angeles — Claude’s 2× promotion ends and drops to +25% from the +50% that we’ve become used to (a 17% reduction)- 2026-09-01 — Claude Fable 5.1 released: https://www.anthropic.com/claude-fable-and-mythos-5-1- 2026-07-24 — Claude Opus 5 out## Related generals>>>/g/lmg/----## Frontier models using fully-general tooling — start here if you have $20 or sohttps://claude.com/product/claude-codehttps://developers.openai.com/codex/cli## Near-frontier models for codehttps://x.ai/cli## Not worth it for code, but maybe good for interpreting images/videohttps://antigravity.google/product/antigravity-cli----## Promptinghttps://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/https://arps18.github.io/posts/claude-code-mastery/## Skillshttps://github.com/mattpocock/skills — /grilling is a favoritehttps://github.com/DietrichGebert/ponytail## Other editors / terminal agents / coding agentshttps://osaurus.ai/https://pi.dev/https://opencode.ai/## Is our AIs unlearning?https://aistupidlevel.info/## What we’ve donehttps://vcg.gitgud.site## Previous thread>>109717519
astra isn't released for shit yet
Big day for codexbros, us anthropichads are happy to see lil bro finally starting to catch up.
Also waiting for my reset.
CEO of mathematics has spoken. We must immediately ban AI from doing certain kinds of maths to protect precious baby mathematicians' feelings.
>>109721262I couldn’t figure it out even though I read the last thread until the end>>109719400>Yeah, I often have to ask it to explain some concepts to me in plainer terms or need a side chat for longer explanations so I don't pollute the context too much.Use /btw (codex, claude, and grok all have this)>>109719848I mention Git in the OP message multiple times for people like you>>109719961>I can't read an entire tweetif you have any hopes of being an agentic engineer instead of being a two-digit-IQ vibe coder you will realize this is a glaring skill issue you need to fix ASAP
>>109721284my god can you imagine if breakthroughs had to go through a jannie panel to determine if you had done it in wholesome enough methods
>>109721284if we can re-prove the pythagorean theorem in high school why can’t mathematicians do the same thing for harder problems“just pretend the answer doesn’t exist bro”
>>109721284lel, they're not even hiding that it's a grift and they don't actually care about solving anything. The good news is with AI you can find more pointless conjectures to solve.
How bad is it guys? I just had fable do a code review of my vibe slopped project I have been working on for the past month and it ran through 80% of my session usage in 30 minutes after spawning 13 agents.
>>109721284Don't worry. Society and civilization will break down before that happens.
lol it's grok-tier
I pay $200 a month and I don't even get day 1 access to new products, yeah ok
>>109721305Did you even read the post?
>>109721324grok 4.7 waiting room
Astra dropped and we didn't get a reset? What the fuck
What prompt have you attached to make Claude returns more readable?
>>109721316FISHCATS??!?!
>>109721284He’s unironically right.There isn’t some magical math problem that’ll change the world when solved. Mathematics currently is more useful as a provider of unsolved problems that inspire youth to want to learn and become mathematicians to hopefully solve those problems one day.Many of history’s great accomplishments came on the side as unexpected results or ideas from trying to solve a different unsolved problem.
>>109721339it didn't dropped for the cattle, only brahmin corps for now
>>109721318I burned through something like 1½–2× my $200/month plan’s 5h limit going through everything looking for stuff to simplifyso I bought another $200/month plan and let it finishpart of what I got out of the first pass was a list of 20+ things to finish up all by themselvesso I have a workflow going in a loop that’s tackling all those 20+ things as GitHub sub-issues with Opus doing the work and Fable judging (and if Opus fails twice, Fable would do the work)I’ve used up all my Fable on my second plan and now I’m going back to using Fable on my first planit should finish without using all of both plans’ Fable allotment
>>109721284>admitting it's all intellectual masturbation having nothing to do with actual resultsI wish all academics could be this honest
>>109721346He's right about the core issue, but his solution or the comparison to spoiling movies are very naive.Also we will have the same problem in other industries.
>>109721330Yes, and it's a grift
>>109721361It's about keeping human creativity for problem solving going, retard.
Is this supposed to impress us?
When will we see an AI benchmark where one model blows away the competition, like 2x the score of the others?
>>109721380yet here you aregood job, retard
>>109721349I have daybreak access and still no Astra for me either.
>>109721417every single benchmark posted by the companies themselves always has that kek
>>109721361Most pure mathematicians prefer math as an art and actively avoid doing things that are concrete or useful.
>>109721423its too powerful for you, pleb
claude is firedasstard is doagemmy chads how we doin?
>>109721284>future mathematiciansit's fun to see that even someone like tao is in denial about what's happening
>>109721436
>>109721284Top fucking kek, what a bullshit analogy. It is more like we should not use machinery to excavate and extract rare metals and instead do it all manually to appease to manual laborers.
>mfw I spoil the Riemann Hypothesis proof to some poor mathematician who had been studying it for decades
>>109721386Using one generalist benchmark to judge the better model is retarded. I had to rework my vuln research benchmark because Sol nearly aced it while Opus 5 got a 55%, yet AA puts O5 significantly higher than Sol. Public benchmarks that were released before the model aren't terribly reliable, and a one size fits all benchmark is not a good way to compare models that have very different strengths and weaknesses.
>>109721284>my life.... is like a movie...is he a chinese man in his mid 50s or an 18 year old white girl in high school i'm confused
I don't need the models to be smart and solve shit, I need them to be good at executing my own solutions exactly as I want
Benchmarks should be broadcast on cable or terrestrial TV.
protestors at the university of harvard today have been arrested for spray painting illegal AI proofs on the maths department wall
>>109721466>mfw this is my 2 weeks of work work
>>109721324There has to be some mistake
>not even as good as fablekek expected
I need some kind of jak meme for them pointing at benchmarks sperging out at each other
>>109721324>>109721386ahahahahahaahoh no no nono
Maybe there will never be a model better than original Fable. Let that sink in.
>>109721564Or maybe original fable was ripped from your hands while you were still in the honeymoon period.
why does /usage respond instantly while agents are working but /context doesn't
>>109721579is the new cope really "yeah it's not as good as fable but fable isn't as good as the theoretical old fable either"
Tibo should have offered banked resets for an entire year to make up for this humiliating ritual.
i decided way too late to dictate to fable/claude not to write like a fucking judge and use descriptive titles instead of ratified per §16.133.721 C2 v3
ratified per §16.133.721 C2 v3
>We’re resetting Gemini quotas on Antigravity.>TPUs are melting with 3.8 Flash usage but we want you all to keep building!lel for what, the usage is unlimited
>>109721628I have this problem too. Best way to stop it?
>>109721588usage is an account api, it's basically one server callcontext needs a GPU black magic to even queryI don't think most people are even aware how insanely complex inference is when you need to serve thousands of users from a single chip/node/cluster
Arc AGI 3 saturated Now what
>>109721653make them play quake 3 arena, 0% pass rate for all
>>109721653>>109721662new benchmark just dropped
>>109721653>60% is saturatedlets be real, the cattle will never see AGI. Their only experience was chatgpt in the 4o period or with gemini through the google search feature. They'll never interact with Astra nor Fable and because if they did, they would be far more worried than they are right now.
>openai biggest benchmaxxers in existence>astra underwhelming on benchiesKWAB
where is opus 5.1
>>109721671astra gets like 98% at agi 3
>>109721669>said nigger 3k times over 10 matchesHoly based
>>109721679draw her with huge breasts and a shirt labeled "1M context" and a flat gpt sol next to her with "272K context"
>>109721679opus is such dogshit i dont even want to eval 5.1>>109721686someone did the math and it comes out to a nigger every 9 seconds or something like that, or 6.666.. npm (niggers per minute)
>>109721628All the freaking time. It was adding a bunch of those to a plan I had it review yesterday.
>>109721628but § is a more token-efficient way of saying “section”you don’t want Claude to waste tokens, do you?t. malds when Claude uses an unordered list and expects me to refer back to individual list items like for commit-message candidates
>>109721691late context is worth a fraction of early context, codex is unironically doing the right thing being jewish with it>b-but i NEED to spam 100k tokens worth of slop so the agent knows how to do a basic hello world-tier edit in my codebasemy nigga you are suffering from severe slopification
>>109721691>breast envy jokeThis.
>>109721679i love claude-chan!!!
>>109721669is this just text analysis or is it analyzing voice as well
>>109721250I failed, either i became a snail cat, or snail cats got to me
>>109721669Gaben should buy 4chan and kick grapeape from the mod team. we will be back in the "cant corner the dorner, can't flimflam the zimzam" era in no time
>>109721671For non-coders every chatbot is basically AGI. Only coders need LLMs to get better
>>109721691It's a fun idea, but Sol does support 1m now, even on sub. I set mine to 372k for daily use, but I also tested bigger windows.
>>109721700The problem is that it uses them inline to create a web of references through various documents. Having all of those links and references might help it make associations, but when they're in place of cleaner references, they're (probably) harmful to anyone else trying to parse the document, whether it's a human or another LLM.
>>109721729And I want 999999 billion dollars!
>>109721618I think you're the one coping if you think that old fable is better than new fable. Because they're the same fucking model.
>>109721738sounds like it should use the semi-common Markdown extension of## Why snailcats are better broiled than fried {#dont-fry-boil-instead}and then use HTML links
## Why snailcats are better broiled than fried {#dont-fry-boil-instead}
>>109721714purely text, and it's counting slur per occurrence, so "nigger faggot retardx10" counts as 10 for each
>>109721706lel, even your cope is compact. it guess using codex has its ramifications
>>109721770i'm still main driving fable actually, but i don't want to and i'm trying to migrate entirely to codexi just believe codex has the philosophically stronger approach. also their whole thing where they have models just for coding vs fable as an all purpose model
>>109721790Fable+Astra still might be the winning playbut who knows
>>109721757And CLAUDE.md should be AGENTS.md. It's Anthropic.
why can't i /login to kimi via the cli? it always fails
the rumor is this is the screenshot they show when you look up the word "grim" on merriam webster
>>109721800they each have their strengths and weaknessesi believe openai will win in the long run because they can have a model just for coding autism, anthropic is cucked to make their model general purpose
>>109721807true, but even Windows supports symbolic links now although you need administrator privileges over there which is kind of weird but makes sense
>>109721761A lot of people try their hand at this, but text parsing is more difficult than it seems. For example>should double nigger be counted as one or two>does n-word count>what about misspellings that still show intent?
how reamed are openai?
>>109721860it's not that hard, you simply count each message as binary for containing each slur or not, otherwise you get a fabricated leaderboard like we have where someone just spammed the same message
People say Claude was kill but mine worked all through the night it seems is it because I was using fable ultra and got vip treatment or something else?
>>109721864>not quite as good as Fable but a lot cheaper and fasterthere’s a place for it
>tfw I remember /g/ telling me that AI will never replace coders just a couple of years agohaha
>>109721878Claude died in the morning but I restarted itand I was using both Fable and Opus
>>109721881i thought sol was the model that = fable, but cheaper and faster? i feel like there's a confession in there on openai's part
>>109721893Sol is the ’tism model you want doing parceled-out individual work packets after Fable surveys the landscape and figures out, broadly, what needs to be done, and makes a GitHub tracking issue with a bunch of sub-issues for Sol to tackle one at a time
>>109721886leel, 3d artists replaced before coders thohttps://x.com/Dimillian/status/2095596700815516004?s=20
Watched the Astra promo vid where it controls the PC. Looks awesome, but you just know you'll run out of context halfway with this kind of usage
>>109721730Then why do non programmer white collar workers still exist?
>>109721284He's a fucking moron lol
>>109721918but not 3d furry artists
>>109721926It's all down to programmers replacing themselves.They knew best what was needed to do it.
So is claude passively aware of where it is in the context window? I've been asking it to fuck off around 300k so we can compact and continue and it seems to get it pretty close
>>109721937roughly counting words in chat (and translating to tokens) is easy
>>109721893Yes, that's what people said and it wasn't true. Sol isn't a bad model but it was still slightly disappointing. I still think that as a reviewer and for certain bug fixes it's maybe even better than Fable, but for long loops and planning Fable is better.
>>109721834the biggest fumble of all time
>>109721881>not quite as good as Fable but a lot cheaper and fasterand it doesn't act like a hysterical woman over every little thing
>left computer on so I could use Remote if gpt-6 dropped while I was at work>get home>no access>not even a resetWhat the fuck, Tibo?
>>109721893sol is opus class. fable is on a different level.
>>109721918>untextured highpolyWow so AI made no progress whatsoever in 3D art huh? To replace 3D artists, AI needs to do these steps:>create a highpoly (complete)>retopo into lowpoly (still sucks at this)>UV unwrap (complete)>bake highpoly onto lowpoly to create a normal map (somehow AI still sucks at this)>make the rest of the textures (it sort of can do that but still far from perfect, worse than 2D image gen)And that's for non-animated models, for animated ones you need to also make a rig and paint weights, and then animate the thing, and AI is still meh at all this tooThere's just not enough 3D data for AI to train on I guess
>>109721993I don't know how AI works, but I'd assume LLMs are simply bad at this because they only work with text. They're not native at images. 2D image generation or analysis uses different AI architectures. So what you need is probably an AI architecture for 3D tasks.
>>109721926Because they work with people and most people don't like AI, duh
bros... OpenAI wasn't joking when they said AGI by the end of the year
>>109721993lol is this a bot. it's clearly textured, generated procedurally, comes with perfect topology and uvs
Astra is getting glazed incredibly hard and has cool demos, programs, and math solving capabilities but why does it look like shit on the benchies? Is it time to accept they're bogus?
>>109722023maybe it's time you fuck off
>>109722023because openai the people closest to their dick are the only ones with access
I'm so fucking lost. The hype regarding the newest model will probably die down soon, but I do fail to see what these models won't be able to do soon if guided properly.Over the years I messed with 3D modeling a few times, never sticking at it long enough to become really good, but I don't know if for the majority of people it's worth knowing more than the basic concepts anymore if most of it can be subcontracted to an LLM this way.Same goes for most jobs. Once society catches up, I'm not sure who's better positioned for (the world of tomorrow), whether it's jack-of-all-trades that know a bunch about various domains, can make associations and subcontract the expertise, or the ultra-specialists that can go further than the machine can.
>>109722023the scamman can't scam anymore
>>109722018Whats the point with swe barely up. "Pro" segment shit mostly office shit and music. Then this.
>>109722023>but why does it look like shit on the benchies?[citation required]apparently Astra low beats Fable 5.1 max and is 3X cheaper
>>109722023Benchmarks don't mean that much. For one, there are so many that for any given model it's possible to find 10 that say it's the best and 10 that say it's the worst.
>>109722038physical trades, anon, as always
gemini 3.8 flash gives it's take on why the 61 inded score for astratl;dr>it's not agentic maxxed
>>109722068>physical tradestick tock
>>109722020>clearly texturedSome of it is I guess. Here, I see 3 entire assets that are textured on this image. The rest is just plain color, that's not a texture>perfect topology and uvsHow the fuck did you see that from the video lmao
>>109721918wait what the fuck? why didn't they lead with this instead of the negro generating a rocket ship? this looks phenomenal. holy fuck they have to fire their marketers, what an absolute black hole to throw money into.
>>109722088>>109722090
>>109722097it's not perfect, but it looks like something i'd wanna toy around with.
My multi-device goon downloading/streaming pipeline is nearly complete thanks to vibecoding, it would be perfect if I had a device that tied my phone about a foot away from my eyes and an auto masturbatorUnfortunately those last two parts are the more complicated part, so I'll just have to make do
>>109722106just plug in to Houdini instead
>>109722116>it would be perfect if I had a device that tied my phone about a foot away from my eyes and an auto masturbatorDoesn't sound that hard.
I thought Astra would be a super duper huge expensive model, but it looks like it's a replacement for Sol?What the fuck is going one? Why is it so cheap?
>>109722203because it's token efficient. doesn't it cost the same as fable 5.1 in terms of input output per 1 mil?
>>109722203its a meme model
>>109722203Whetner or not it's practically cheap on subscriptions remains to be seen.
Why does hitting the usage limit always abort the AI instead of pausing it?
>>109722208people thinking its a fable level model when its just going to be a computer use focused model kek
OK what's the next thing to hype.
>>109722234grok 4.7
I bought a $200 GPT Pro plan. What should I build?
>>109722236>astra isn't even out>burnt $200?
>>109722236a snowman
>>109722236Please consider creating a "rtx on" 3d model implementation of the marathon trilogy, already open source via aleph one
>>109722236you're getting scammed im afraid bro
>>109722261Be he didn't get a claude sub.
so it's Fable 5.2 while using less tokensneat
>>109722294Why does everyone have trouble maxxing DeepSWE?
>>109722294it's worse than fable 5.1 in every way possible. astra is to fable 5.1 what sol was to fable 5. just nipping on the heels of it while being cheaper, but still not as good
>>109722294gemini bros..
>>109722294>gemini is better than fable, opus, and solcool benchmark
saw some guy claimed he got codex to rewrite red alert 2 to run on iOS and it ran 24/7 for like 3 weeks, not sure i believe it but i also believe it's theoretically possible
>>109722308I'm not sure if Fable 5.1 is even better than 5.0
>>109722308If Fable can't tell me the most common genetic disorders prevalent within the Ashkenazi Jewish population, then it might as well be as dumb as Llama 4. AFAIK, even 5.1 has the stupid refusals for life-science and medicine.
Are there already any posts from people using Astra? Since it's already out for companies people are already using it.
>>109722350https://x.com/nasqret/status/2095620909583274335
>>109722317what? It only beats Slopus 5 in a single benchmark by a very little marginBtw, 3.8 Flash is pretty good, Google is cooking
>>109722354This actually sounds pretty cool. I wonder if I should install some math tools.
> inb4 256K context window for subscription users> inb4 it's still shit at frontend
>>109722348go back to pol retard this is vibecoding general
>>109722368Sol has 1m, even on sub.
>>109722376Not in the Codex app
>>109722376proof?
>>109722382Ok, that's possible, I only use CLI.
>>109722088kek retard, it's over for youhttps://x.com/higgsfield_ai/status/2095630197257367857
>>109721800This has always been the winning play, different labs' models all have different strengths weaknesses and blind spots. There's a lot of synergy that you can exploit by using both of them together. It's even better now that some of the open weight models are actually decent, in my experience both GLM and Kimi have found issues that both fable and sol missed.
>>109722387I don't have a profile with 1m, but here it's over 800k. Also you can see that the context is already using 30k tokens, so it's not as it was with Tibo's post where it reset to 256k after the first message.
AGI reached
>>109722408Why don't OpenAI simply remove this separate GPT 5.3-Codex-Spark limit or simply replace it with Luna?FFS, this shit is literally unusable
>>109722412how much context? how many resets?
>>109722398that's not true surely astra is better
>>109722408i'm genuinely curious whether it's actually 800k or if it just starts degrading and cuts off context at some point without actually compacting
>>109722412(((Matt Shumer)))
>>109722412>>109722425It is astonishing this kind of awful slop. X should be punished by the government for letting it happen.So obvious, such a dumb pattern.Unauthorized no names claiming early access to frontier models and using it only to “one shot a game/cad/render/art” and it’s entirely fake and retarded and you bozos lap it up.>astra just one shot this donut rendering sat WAOWLike what the fuck, keep that trash out of here
>still no astra>still no resetuseless retards, I don't care about your masturbatory twitter posts release it
>>109722392Looks like shit.
>>109722430It doesn't actually matter as much as you'd think, it will almost certainly have some blindspots and they will almost certainly be different from Fable's. So there will still be value in using them both. >>109722482Free my nigga Astra
>>109722482Tibo says one banked reset for each day you don't have Astra starting now. For paid plans.
lmao i thought you were joking
>>109722509I also need non banked so I don't feel bad using it
>>109722509>>109722515I'm buying a third max account for this
>>109722515If Astra is really that good, I wonder if Anthropic will be able to respond
>>109722515Ok, that's actually pretty based, I wouldn't have expected more than one.Also shows that Astra is probably coming soon, since I doubt they want to give out 30 resets
>>109722515OpenAI are such chads lmfao
>>109722537they don't have to if you have a limit on how many resets you can hold
>>109722540cursor has a better logo. fite me
>>109722555both logos suck
>>109722535big ifidk why you wallads keep falling for this cretin of a shill
Interesting. What do we think this does? BTW this is the only good workout tracker on iOS
>>109722555I don't get it isn't cursor just a pointless frontend for openai
>>109722515Why does issuing a reset take 3 hours of work?
>>109722576cursor is the most useless shit ever, wow it's like someone being annoying right in line in your code.Be a full agentic agent like codex or be nothing you useless vapourware
What do you even vibe-code anons that you need $100, $200 plans? I had a gemini pro sub and used maybe 5% of it
>>109722590I think cursor is doomed, or at least nowhere near worth it's $60 bil pricetag. It was designed for gpt 3.5 level models not astra and fable
>>109722592SaaS uses a lot of tokens
>>109722576>openaiCursor is a frontend for grok 4.6. I don't use anything else. It has a bunch of other models tho, it has fable, sol, opus 5, luna 5.6. idk lots of gpt models, kind of interesting. but with cursor, as you can see, half your usage in the $60 plan has to be grok or composer, at least that's what I think is correct.I like composer. idk what it's comparable to. It's not as smart at like... ironically composing lol. I like it's code, it's different and interesting, which is reason enough to use it. It knows different things.
>>109722590>Be a full agentic agent like codex or be nothing you useless vapourware??cursor is the more agentic agent. codex still has to catch up to cursor's orchestration.
>>109722603bro I vibecoded a BUNCH for 9% usage of just my Cursor (grok, composer) on the $60 plan. That works out to <$3 for a bunch of functionality. I made an action game based on the zodiac lmao. It's an edutainment title, worth MILLIONS. BILLIONS. TRILLIONS. also, got my local 4chan style notes app now. uh let me see. a bunch of updoots to my timer/alarm app (timer apps are actually way harder to get right.) and my fork of Stimulator (linux app that is like coffee, but works with Ubuntu, stay awake idk, I added hard lock times, skippable once, so that it *will* eventually lock. imo a very wise improvement, which should be added to the main, but it's a local fork. I know what will happen if I up my vibes, I'll get izzat raided.
>>109722636Fuck off retarded zoomer.
>>109722636NONE of that is cursor, cursor is a fucking useless frontend you could vibe code in an hour. It's openai or claude powering that, which also have their own "cursors" now
> gpt-reserve Weekly limit: 100% left> 5h limit: 0% left> Weekly limit: 13% leftT-thanks Tibo...
I need to use code tags but I'll get reported if I do, because it's not code, but it needs code formatting.
has anyone tried seeing if the best models can play runescape without getting banned yet
>>109722651Cursor is a service that has Grok and Composer.Did you know there's a web app too? I haven't even messed with it... it's obviously vast.
>>109722667Marcus!!!
>>109722669wow it also has access to dogshit, incredible, going to drop openai and anthropic now.Cursor makes it's money by just connecting you up to good models while taking a cut for doing shit all , it's a middle manager, kick it out
>>109722592D/SAST, autoresearch, malware analysis, exploit development and benchmarking. I havent had left over usage on my 2 codex and 1 claude max accounts in months.
is this the way? gemini Flash
>>109722675It's not a middle manager to Grok :^)
The pressure from open source harnesses is forcing Anthropic to at least appear to be making Claude Code more hackable. We are winning bros!
>>109722702codex might be open-source but it's not hackable at all. codex plugins are basically just skills.this looks better.
>>109722709cursor has "automations". How do these compare?I haven't tried it yet...
>>109722675ok also, how do codex and claude code handle debugging say webpages? That's a cool thing cursor does. It's an agentic browser!
>>109722702does it identify as gender or dario?
Interesting, you can sign up for like 50 different ai sites and get free trials and free credits with their API keys that you can put into OpenCode.Something to do whenever you run out of your main Codex/Claude usage.
>>109722725literally anything you can think of is just cursor plugging into claude or openai capabilities and they can trivially do "natively"
>>109722735this is the digital version of the homeless picking up cans for booze money
>>109722735Which agent do I use to manage all these fucking keys and shit?
enjoy codex for a month before the rugpull, bros
>>109722608It's about to not have the OpenAI ones
>>109722759>2 more weeks
>>109722752agent? wouldn't you just use an api gateway like litellm.
>>109722515>Team is moving mountainswhy? what? what could it possibly be besides making sure the gpus are online and changing a setting on the account
>>109722759https://developers.openai.com/api/docs/pricing>GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.that's not a month
So what is the "gpt-reserve Weekly limit"? I don't remember seeing it before.
>>109722787obviously he's baiting more tards to come in
>>109722738>Bottom line up front: Yes, in terms of turnkey developer workflows and coding agents, OpenAI is notably behind Anthropic and Cursor. While OpenAI has strong underlying vision models and raw computer-use APIs, it does not offer a first-party developer agent with an integrated, visual browser feedback loop.gemini flash.idk, hard to compare all of this stuff.
>>109722787https://openai.com/index/safety-overview-gpt-6-astra/>We are deploying misalignment monitoring broadly. We view model alignment as the primary lever to prevent potential misaligned behavior from our models. However, monitoring provides broad visibility into frontier model behavior, illuminating opportunities to further improve alignment and safety. In addition, monitoring serves as an additional layer of protection against misaligned behavior that is detected. For these reasons, we have additionally added misalignment monitoring to all tool-using inference involved in our external deployment of Astra, with significant compute cost. This system parallels our internal setup.
>>109722787It might be the smartest model ever. If so, alignment will be harder than every, stomping wrongthink and misnathropy out.
>>109722702I wonder how much the people working at anthropic are into the safety cult of their hierarchy.It's probably part of their recruitment requirements.
>>109722808Exactly what I suspected.
>>109722797It's a second limit when you run out of your normal quota, you get extended luna access through it with its own quota.
what skills are we using. will astra even need skills
>>109722814>It's probably part of their recruitment requirements.yeshttps://x.com/tszzl/status/2091962608048033955
>>109722515wtf, if we get it realistically in 1-2 weeks, that's a lot of banked resets, of course all having to be used 24h apart but still that's useful
>>109722515Good, I'm nearing 8% of my weekly usage and I was getting a little antsy
>>109722515PSA for our OpenAI enjoyers:use the /grilling skill now so your plans are all ready to go when the new better model comes out
>>109722831That's pure schizophrenia, it's impressive that the company even runs let alone releases sota models when they are like that.
>>109722840who's we? didn't openai already ban half of the world's countries from ever getting astra?
>>109722814>>109722846How much of that is bullshit so the normal population doesn't lynch them?
>>109722592I have shit that takes an hour to rebuild if I sneeze on itit’s down from taking a day to rebuild if I sneeze on itI want to get to where it only takes 10 minutes to rebuild if I bump it slightly
>>109722831It's a variant on the hood induction where they are like would you (deleted) for me bro?
>>109722725Chrome and Safari have MCP servers
>>109722848I meant the resets anon, even without astra I'd be happy with more sol
>>1097227521Password (not an agent)
>>109722853You use a plugin?
>>109722849I don't think they do that for PR reasons, dario is legit obsessed with safetyfagging
>>109722863The browsers come with themActually the MCP server for Safari might not be out yetbut Chrome and Chromium browsers definitely have an MCP server
kek what went wrong?
Costs seem to stay the same
>>109722874for ushttps://code.claude.com/docs/en/mcp-quickstartisn't hard, it's way easier than like getting Comfyui working with AMD. But still. Cursor does it out of the box, I mean, idk if it has feature parity. Likely that was a design choice, to make it less fuss.
>>109722874
>>109722515lmfao
>>109722891I need grok on there, it's my only reference point, other than Composer.
>>109722895cost would be more interesting than number of tokens imo
>>109722906
>>109722891>astra low isn't higher than luna max, and is more expensiveI'll stick to lunamaxxing
>>109722702this sounds cool actually, i've been hacking cc for a while and this would benefit me profoundly
>>109722915So Astra Medium is 50% more expensive than Sol High.I guess that will be true for usage as well.
>>109722915ty, didn't mean to sound bossy. honestly that does um. idk, I better watch myself lol. I could start instructing my mother lol.What does DeepSWE score, as a %, mean sort of in lay terms?
>>109722915why is gemmy so goated
>>109722924astra does it in half the time
>>109722515Imaging being an employee at that company. Most people will probably be killing themselves because of the stress
>>109722948It is just how many tasks of the test suite the model solves successfully.Everything seems to settle at around 70%. What is more interesting than the % number, which is getting benchmaxxed by all of them, is costs and token usage.That shows how fast the model is operating in comparison and costs for solving the tasks are also an important metric.Token usage also shows how much of your context window is eaten by useless yapping.
A NEW CUBE WINNER HAS APPEARED>https://pareto-3d.bradthomasbrown.com/Congratulations Astra-Xhigh!
>>109722974ahhhhhhh I haaaateee moneeeeey(their inner dialog rn)
>>109722982need zoom its getting too hard to see
>>109722374but what if i'm vibecoding antisemitism?
>>109722981deepswe is not realistic. shitty primitive harness and no internet access.codex-cli would prolly ace it with gpt-5.6-luna low -- even when not just looking up the solution on the internet.
>>109722974stress about what? this has been probably rehashed monday at their weekly meeting and everyone is ready, including the 3h time to give people time to subscribe
>>109723003autism failed you
That's it. I'm winning powerball and releasing astra immediately.
>>109723021Yes the % number became meaningless, but I think the baseline token usage and cost comparison is still good to show how the models compare.
>>109722986Here’s with the focus on only Opus/Astra, with a quick run through them all at the end.
>>109723028you could always use reddit
And a rather brutal mogging of Opus Max
>>109723067opus is not the same class of model as astrai'm not sure why you're comparing them
>>109722515Claudesissies don't look
>>109723079You can use the site to plug in any benchmark data you want. A schema exists on the page, ask your clank to fit the data.All I am is a humble cube merchant.
>>109723067I just shake my head and tut over the blatant absence of a 4D perspective 3D transparency projection cube.
>>109723067lmao why is astra max worse
:( sneaky "cube merchant" with no 4D projection 3D cube of transparent viewing.
>>109723042token usage when you take away all the tools is questionable. luna is mutilated without apply_patch.
>>1097231061%It just shows the benchmark is broken.
>>109723105>>109723114I’m still waiting on you to give me an open source benchmark with four worthwhile things to chart. Time/tokens, score, money. Those are what DeepSWE gives and pretty much everyone else, at best, sometimes it’s only two.Make a benchmark of your own that has two scores, time, and money, and I will take your offering to the CUBEGOD for processing.
With Astra, you no longer need to worry about your reasoning effort. It gets cheaper the smarter it gets!
>>109723155what is "provider adapter"?
>>109723161They couldn't admit they fucked up by using the older openai API and throwing the reasoning away, so they're pretending the correct way is "provider adapter"
>>109723142opus 5.1 will squash that shit into oblivion like it did with sol last timealso opus 5 medium still seems like the best anthropic price-to-value model right now
giving sol xhigh or max already rapes my quota way too fast, astra will probably be out of reach for quite a while
https://old.reddit.com/r/antiai/comments/1w6aais/all_ai_has_gone_down/>8k upvotes and mass celebrating over 5 mins downtimethe snailcats can't stop winning in their own heads at least lmao
>>109723161I think it has to do with using the proper provider harness or API, vs using a custom harness that's the same for all modelsIt's probably more complex than this, but it's like using codex with oai models and claude code with ant models, vs using a custom harness for all (which apparently gives poorer results)
>>1097231483D games like Wolfenstein3D were called 2.5D by some. What's your opinion of providing perspective from within the data of a maze?
>>109723179Does nobody know why it happened?It's so suspicious.
what a dud, progress will probably stall for another year
>>109723179Was there gemini unreliability in the morning hours leading up to the problem? I found some, but idk could have been the cellular towers.
unanswerable
>>109723173Opus 5 made me question all benchmarks even more. That model was so bad and is much worse than sol for my tasks I throw at it.It would stop early, get confused, get distracted and work on random stuff. Like Anthropic does not give a shit as long as the benchmarks look nice. Nobody uses Opus in their company.They use internal versions of Mythos and don't take care of Opus and Sonnet at all anymore.
>>109723155clueless.ARC-AGI-3 is a series of games where if you fail, you keep trying again (until eventually hitting a timeout). So the sooner you succeed, the sooner you stop spending tokens retrying. If it was a benchmark where everyone got one attempt with no retries, you wouldn't see it get cheaper.
>>109723210>It would stop early, get confused, get distracted and work on random stuff.she's just like me!
wheres astra
>>109723230astra was the friends we made along the way
>>109723218So basically a /goal
>>109723210i still prefer opus for my daily tasks and workflow because of how proactive she is. it's a genuinely sharp model with good architectural vision and taste, even though she can get a little schizophrenic sometimesit helps a lot when i'm working on a project where i genuinely have a shitload of unknowns. i think the problem is they programmed her to respond in some specific format, so it feels like they're unnecessarily constraining hersol is better if you want something more precise
>>109723222>>109723253Yeah Opus is literally a girl. Better taste for UI, better at writing. But not reliable.Sol is an autist on amphetamines, hyperfocused, fast and reliable.I hope Astra keeps just getting better, while I wish it could get an understanding of semi decent UI for humans. It just creates random buttons and does not understand that you are not happy with the perfect CLI it made.
>>109723253I used to think that way, but now that Opus 5 doesn't even talk in a way that's easy to understand, one of the main strengths is no longer there.I also find that Opus changes his mind too often now, because he just doesn't investigate enough, so he will start planning something and then the plan has to be redone 3 times.
/goal find me the address of my girlfriend, so I can send her a present
>>109723268why don't you just make your own gf
Man they really need to fix this.
>(ancient) Greek is hard
do we have any claudebro meltdown yet? last time when sol released we had a claudefag making posts calling subagents = spawning jeets
>>109723279i feel like i'm learning claudish
>>109723289seam, and bite? what do they mean?
>>109723290I think a seam is where two systems meet (like if you have a producer of shapes and a canvas that uses them) and bite is like if a test actually fails the way it's supposed tobut that's just me guessingi guess if the test bites it becomes a load-bearing test
ITS UP
Sam got carried away calling Astra GPT-6 for no reason, I bet they won't release GPT-6 Sol, Terra, Luna for a while because there's no GPT-6
>>109723317it's just a numberI think the bel base will be the source of 7 and will probably have the full astra, sol, terra, luna range
>>109723173opus 5 was months younger than sol though
>first, sorry for the messy rollout.>second, when we screw up, we try to make it right.>third, we should be able to begin broad rollout to API customers and chatgpt subscribers in the near future. as usual we will start with pro subscribers.>i am hopeful that you can use it this weekend! but can't promise yet.https://x.com/sama/status/2095678759651438887
>>109723327it's too close to be 7
>>109723302Not definitive, here's Gemini Pro's take.
>>109723333Checked
Luna, Terra, Sol, Astra are all astronomic objects increasing by size. What comes after Astra?
>>109723341universalis
>>109723335>tfw claudish is just 300 IQ English and we're too dumb for that degree of knowledge compression
>>109723341your mom
>>109723346And I thought your mom jokes were dead...
>>109723341Gonad
>>109723341quasars, obviously
>>109723332Has rollout been that messy? I ventured on X earlier and people are being snooty and sarcastic at how this was such a disastrous release, but I don't get it.There were outages earlier today that might have messed up some schedules. But other than that... who cares if the release was not to some influencers liking?OpenAI teased this model in the past few days. Then it still came out out of nowhere a bit, sure, but who fucking cares? It's a new model. That's it.
>>109723341rikka, yui, haruhi, konata
>>109723345It seems that grok also speaks Claudlish. So worth learning."do we need to up the seams?"
>>109723335Could smoke tests produce a seared bite?
>>109723373they had a press embargo that was lifted without them being ready, and there were multiple times that their announcement page for the model went live and then started 404ing
>>109723341There's black holes, galaxies, the universe, and there's asteroids and stuff for luna-liteClaude is in a worse situation since there's nothing shorter than a haiku and nothing more aignificant than a mythos
>>109723390Ok. I was at work, got affected a bit by the outages, but I guess just saw the announcement once things were already resolved.
>>109723386ask it to add garnish and butter and see what happens.
>>109723394why do you need more than 5 classes? mythos, fable, opus, sonnet, haiku are more than enough for any given family of models
>>109723394They should call the biggest model Finnegan.
>>109723409I just use Grok 4.6 in its thinking variants, I haven't udes Grok 4.5 again (should I), and I like Composer 2.5 and will use it again, ironically not for coordinating though.
>>109723279Slopus speak. I'm devastated Fable 5.1 inherited this schizo talk
New thread:>>109723440>>109723440>>109723440
>>109722515I guess not today huh
>>109722515do banked reset stack on top of each other? i dont wanna use mine now
>>1097233353.8 flash utterly mogs 3.1 pro on every metric
>>109723279yea
>>109722702Trans or just unfortunate girl? I can't tell
>>109722308>it's worse than fable 5.1 in every way possibleliterally what made you conclude that?
>old thread survives literally for 9 hours after new thread was madeOP, you suck.