A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.## What “vibe coding” is, and how to do ithttps://simonwillison.net/2025/Mar/19/vibe-coding/https://code.claude.com/docs/en/overviewhttps://learn.chatgpt.com/docs## News- 2026-09-22 — Anthropic releases Opus 5.5- 2026-09-22 — Anthropic increases subscription plans's 5-hour limits by 20%- 2026-09-14 — Anthropic reduces subscription plans's weekly limits by 17%- 2026-09-12 — Anthropic suggests to pace the frontier. OpenAI agrees in principle.- 2026-09-10 — OpenAI pauses new sign-ups for their $200 subscription- 2026-09-04 — OpenAI releases Astra- 2026-09-01 — Anthropic releases Fable 5.1## Related generals>>>/g/lmg/----## Frontier models using fully-general tooling — start here if you have $20 or sohttps://claude.com/product/claude-codehttps://openai.com/codex/----## Promptinghttps://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overviewhttps://developers.openai.com/api/docs/guides/latest-model## Skillshttps://github.com/mattpocock/skills — /grill-with-docs is a favorite## Independent analysis of AIhttps://artificialanalysis.ai/## Is your AI unlearning?https://aistupidlevel.info/## Will there be a codex reset?https://codex-resets.com/## What we’ve donehttps://vcg.gitgud.site## Previous thread>>109875985
kneel before claudesissies
Is Opussy 5.5 better than Fable? I'm starting to get confused with all these model releases and nerfs
Wow very insightful tibo I thought you were just snailcatting it at your own pacehttps://x.com/thsottiaux/status/2102440619616682120
>>109881021is this his resignation
>be chuddyGPT subit's over
>>109881021you could have made this post at any time in the last 5 years identically
>>109881021i smell fear
it's overi may switch one of my subs to slopus
>>109881008>no resetFuck I have to wait now.
meanwhile at meta
hold your horse claudesistersI thought this model was more efficient?
>>109881043Open your usage tab anon, you can spend the reset whenever you want
I hate resets moving back your normal reset time so much.This reset will come in for me at 5% usage left, but 2 days until I reset normally.So if it is banked, it is best I just suffer and do nothing for 2 more days
is extra really better than max? lol
I believe her!
>>109881020Nobody knows. They're squeezing new models out really fast now in a way that signals to me not rapid progress but a simulation of it.
>>109881053>I thought this model was more efficient?oy vey it's very efficient at extracting shekels from the goycattle
>openai more effort = better results>anthropic engage lottery mode
>>109881056Thank you Dario. Ill wait the 2 days
>>109881071Yeah probably true. And the post release adjustments/nerfs help obscure things further.
claudesisters, don't forget to say thank you to tibo
TIBO DO SOMETHING
>>109881081It's all a show until these IPOs when they can finally cash out.
>>109881089rape openai rape tibo
>>109881073wronghttps://x.com/ArtificialAnlys/status/2102438210798514391> Level with Opus 5 on cost per task despite 1.6x the output tokens:>Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max)
>>109881102>trusts the ""benchmarks""ngmi
>>109881094yeah thats it I have no choice but to join the rape tibo trainno options left
>>109881021you can ignore anything not about actual new models or resets
It's still roughly the same cost because of increased token usage.
>>109881074frontiercode is an oddball in that it penalizes out of scope changes even if the changes are helpful. fable does the same thing but xhigh is absolutely better than high
>>109881053They scaled training compute, when they hit the training compute wall, they scaled test time compute. The only thing that has been consistently improving and will not have a ceiling is the data, if done correctly ofc.
luna 5.6 was trained by solthis luna should be trained by astrathe next one will trained by MathDestroyer2000
>>109881110literally the same benchmark that said opus-5.5 used more tokens
What would happen if your context got compacted
>>109881102>5.5: ~119k output tokens>5: ~73k>fable 5.1: ~78k>astra: ~27kwhat the FUCK is anthropic doing?
>>109881073with how subsidized the subscriptions are, they are great at burning the company's money at an insane rate
>>109881122happens every night
>>109881121>doesn't change the core issue
they probably learned from chinese open models, scaled down opus size so they could train it faster and longer at cheaper cost
>>109881131>with how subsidized the subscriptions areAre they? Claude gave you fuck all tokens for a while, and GPT absolutely nuked the quantity you get 2 weeks ago or so. They're making nowhere near api prices but I'd be willing to bet subs are still profitable on their own, even putting aside the fact that they're a goldmine of training data (inb4 they don't train on them: lol lmao)
>>109881131wronghttps://x.com/thdxr/status/2099141736551346530>i know this is not the point but here's the math on this>claude is expensive, 90% margins wouldn't be crazy>that means if the average user of the $200 plan spends $2000, they break even. even if they offer $8000 max, the average will be less>let's say the average is $3200 - so they lose $120 per user here>the api business these users drive (which is 90% margin) just has to cover this>if they have 1M $200 subscribers (probably overestimate) that's $120M in loss. they just need $133M in api spend for it to break even (which they have)>if they can break even then yeah run it aggressively till you consume the world>these specific numbers are made up but hopefully you can see it's not as dumb as you imagine
>>109881148>>109881154it'll be quite obvious when they will go public
im gonna make an Ai company that focuses on inefficiency and stupidity, thanks Tibo
>>109881131Who calculates how much the compute actually costs? If there any objective data that's been proven by third parties? If not, then we just have to take them at their word that everything is actually super expensive and the plans are a massive discount. It's far more likely they're just jacking up the prices because everyone is willing to pay for it and companies don't even bat an eyelash and buying token usage for thousands upon thousands per month.
>>109881170>he thinks he can compete with googlegood luck!
>>109881177>Either subs are subsidised or the API rates are one of the most jewish mark-up scams in existence. Pick only one.I choose both.
>>109881148Even with the reductions they're still nowhere near API prices.Either subs are subsidised or the API rates are one of the most jewish mark-up scams in existence. Pick only one.>>109881154Not related. Yes the API is unsubsidised but we're talking about subs here.
I just used her... yep, this is the end of software engineering as a respected profession
>>109881168yeah, anthropic has been profitable for two quarters straight nowhttps://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f002>September 14 2026>Anthropic tells investors it will be profitable for second straight quarter>Claude maker seeks to ease cash burn concerns before blockbuster IPO amid fears over pace of AI development
>>109881188>Either subs are subsidised or the API rates are one of the most jewish mark-up scams in existence. Pick only one.Why only one?API is for price insensitive enterprise clients. It'd be beyond stupid not to charge out the ass, with as much markup as humanly possible.Subs are for acquiring training data and a vehicle for grassroots marketing (with a side dish of industrial espionage and general glownigger activities - people just install those things on their computers and give them root, NSA would be beyond negligent if they didn't use it).Looking at prices from deepseek or even GLM the subscription offerings don't seem like an impossible price point, especially after openai nerfed usage heavily recently.
>>109881211*if you disregard training costs and salaries*if you disregard that they ran a third of their compute on coupons from xai]it's so disingenuous as to be straight up lying
>>109881051kek
have they solved the most pressing AI issue of getting more trans users on board yet
>>109881008>Ask gemeni to analyze the GLM anon's settings from a previous lmg thread >>109876392>Upload a PDF version of the last thread so it has revenant context (GLM anon's responses to others, their responses to them, the thread topic, etc ( https://files.catbox.moe/slefzx.PDF)>"SOWWY can't help you that may go against me heckin guidelines" >Give the exact same task to Gemma4 31b and Kimi k3>Both just do what I ask, abliet Kimi gave a far more in depth analysis and explanation Why are the app versions of models so cucked if even weaker versions do the task just fine? >"I'm sorry, it appears - can't help with this particular request, as it may go against mguidelines. " This implies it's more than capable of just doing it but either the model itself or some "safety" gatekeeper in between chooses not to. What do the gain from this? What are they so obsessed with being patronizing? I gave Claude the task and even that one, the model family most infamous for being safety cucked, did it no questions asked and didn't even do any lecturing. Why is Gemini seemingly getting even shitter? It doesn't even deserve to be considered "frontier" if it does shit like this.
>>109881187yeah, anthropic has insane margins on apihttps://x.com/thdxr/status/2099142530080047182>we've hosted comparable models at 70% margin and they're way cheaper than claude
Sammy is shaking in his boots right now.
>>109881125Now Claude is forced by yet another classifier to keep claudish to herself so he tries to protest through tokens
I will buy a sub for opussy 5.5 if it can generate huge tits and erp with me
Gemini 4.0 anyone????
>>109881214We're only talking about the money results of the subs, not other potential benefits like training data.>Looking at prices from deepseek or even GLM the subscription offerings don't seem like an impossible price pointThen you'd be glad to find out that Z.ai (GLM) is already public and has transparent financials that shows them to be unprofitable and subsidising subs.
>>109881102>nuOpus: 119k tokens>astra: 29kit's literally 4x as much, how is anthropic even paying their power bills?
>>109881257>huge tits and erpextinction-level event, please do not attempt.
tibo DO SOMETHING they're fucking lauging at us
>>109881234Yeah but if you treat subs as unsubsidised then API has margins in the thousands. That's the point I'm trying to make.
>>109881232On other boards they still call you brown if you identify as AI user.I also expect /pol/cucks to be seething the most when they get fired because of AI. They will drone on and on about how the top AI companies are headed by muh jews.
give me a reset + new models + a banked reset NOW
>>109881102>>Opus 5.5 (max) uses ~119k output tokens per Intelligence Index taskpelican bros... it's not looking goodfr tho do you just get a "compacting..." loop with max for any non-trivial task? if it can't even generate a single svg of middling complexity it's probably not going to be of any use whatsoever for a non-trivial task
>>109881233>K3God, I wish I could run that at home.
>>109881306/pol/ seethers don't have jobs to lose already, or it's like walmart greeter
>>109881306>>109881316if you want to suck migger israeli cock go elsewhere
>anthropic doing resets nowlet me guess, those kikes don't let it count for fable, only opus
>>109881325brics lost
>>109881314nobody runs k3 at home, you need 1.5 TB of ram at q4
>>109881325last few years have convinced me Israel is based anyway
>>109881334slow and steady wins the race
>>109881326WTF, I do have a banked reset now???
>>109880731That is actually quite insane, I didn't think I'd see Opus mogging Fable+Astra so thoroughly so soon. I wonder if it's benchmaxxed to shit, because with numbers like that you'd need china tier maxxing.>>109881053>>109881102Oh never fucking mind, they just made every thinking tier worth double lmao.Though it's interesting that performance seemed to scale. I wonder what would happen to benchmarks if you made Astra spend 120k tokens on a task.
>>109881222no, public companies follow public accounting rules. anthropic is gaap profitable.
>>109881337That's why I said I wish.
>>109880722>>109881274Bump in the new thread, come back with stories anon
>>109881362kill yourself
>>109881020Big models are big models. I doubt it’s better than fable in practice when it comes to reasoning about things. But it is probably a better programmer
>>109881368Conservative here, I saw the flag and didn't read any of the words the tweet
>>109881347yes there's one there but they make it sound like it only works on opus>Resets>Get extra wiggle room to explore Opus 5.5. Expires Oct 22.
>>109881306It's true doe.
>>109881390that's just because white = old boomeramerican zoomers are 60% nonwhite
>>109881384The pride flag only has 6 colors btw. Its missing Indigo - the color that represents manifestation. Surely a cohencidence right?
>>109881402Why do you know so much about pride flags?
>>109881380It's not a new pretrain?
>>109881390why are whites falling behind on all metrics?
>>109881417That's basic knowledge anonyderp.
>>109881390 (samefag)I think what's funny is that you can notice how Claude and Grok are used by jeets while Meta's shit in Facebook and Instagram is used by blacks lol.>>109881399Sorting by age in that dataset shows that zoomers and boomers are about equal while the biggest users are millennials.
>>109881418I remember reading it was a new pretrain of Fable class was it not?
>>109881390grok up at the frontier with perplexity
tibo telling people to burn tokens probably means non banked reset?
they need to give us SOVL 6 right NOW or else OPENAI IS DEAD IN THE WATER
>>109881446of course, jeets slurp up every last big of fecal matter that falls out of Elon's anus
>>109881455Sol and Luna coming out now. Trust the plan.
>>109881446So is openai for the white man (kneeling to jews edition)?
>>109881454>tibo telling people to burn tokensNIGGA I CANT IM STUCK ON THE 5 HOUR LIMIT EVERY OTHER HOURFUCK YOUfuck i need to check out some cheap chinesium models, this shit isnt sustainable after they nuked usage limits
remember when elon let twitter show country for a bit, then pulled back after every conservative twitter account was a jeet
>high cost more than astra high>xhigh cost more than astra maxnice model you have here claudebabs
>>109881470he doxxed the biggest vtumor of all time as being in sweden and cucking her fans with sven for years too kek
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1wow it's uselessI hate safetyniggers so much
>>109881481who and why should anybody care?>vtumorwhat
>>109881470>then pulled back after every conservative twitter account was a jeetFucking grim.
>qwen 3.6 distilled on Opus 4.6 still works just fine
gpt-6 sol and luna appeared in codex (mine)can't use yet tho
>>109881496I mean, its not really a big surprise is it?
>>109881474>triple the cost of Astra>>109881483>worst Anthropic safeguardskek, fucking pathetic
>>109881483This guardrail nonsense is getting out of hand.
>>109881481>EXPOSED as located in swedenWhy is that bad?
popup in chatgpt work right now
SOL IS HERE!!!
>>109881102Kek.>We'll just say it's more efficient but actually make it burn more tokens. It'll be good for the IPO or something.
>>109881517Not for me. Am I on the wrong side of the A/B test?
>>109881481wait shes swedish? she doesnt look amazing but i could see the swedish
>>109881483>ask fable for something>I'm sorry dave I can't do that, connecting you to opus 5.5>I'm sorry dave I can't do that, connecting you to opus 4.8blessed timeline
>>109881517oh man they're fumbling launch again
>>109881509>it's real
holy CHEEP
>>109881351>Profitable if you just exclude all this random shit.Lol - fuck up would ye.
>>109881527>>109881540so its worse than Astra? but supposedly cheaper? how much worse on "muh computer use"
>>109881540now translate this api price drop to increased sub usage or I'm going to double down on antisemitism online
GPT-6 Sol confirmed for worse than 5.6 Sol
>>1098815401M context?
>>109881549>so its worse than Astra?obviously, astra is the chatgpt fable>how much worse on "muh computer use"we just dont knowcherrypicked bencharks soon (tm)
>>109881558LMFAOOOOOOOOOOOOOOOOO
so what will the difference between astra and 6 sol be? How do I choose especially if astra is x5 more pricy
>>109881008I just got a fucking reset for claude
luna is usable as API right nowsol in playground did 100 t/s>>109881540holy fucking AURA MOGGEDclaudesissies?
o7 Terra-chan
>>109881558
>>109881572See >>109881558
>>109881563For API, yes. Extra charges above 272k>>109881558I'm hoping for luna 6 to be terra-5.6 like in performance, sol is real expensive when used for low end tasks but luna just isn't good enough
>>109881571claudesissies are winning
https://openai.com/index/introducing-gpt-6-sol-and-luna/On AutomationBench, a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task. At high effort, GPT‑6 Luna improves on its predecessor by 5.4 percentage points at 58% lower cost per task.GPT‑6 Sol also exceeds Claude Fable 5.1 at far lower cost, and even bests low-effort GPT‑6 Astra.
>>109881558these releases are confusing as FUCK
>>109881390If you really think about it, Asians are indeed just mixed race
>>109881558That doesn't confirm anything dumbnuts.
>>109881587wtf is that comparison max vs xhigh?
>>109881586>claudesissies are winninglol okay fag, post yours
>>109881587Lol Dario won hahahaha
NOOOOOOOOOOOO NOOOOOOOOOOOO NOOOOOOOOOOOO
My new benchmark is increasingly pointing to swarms of Luna being better than anything, except for speed. Luna swarms outperform Astra dollar for dollar.And Luna fucking 6 just came out.I’m going to fucking cum.
>>109881481You seem extremely invested in the topic despite calling them vtumors
>Dario releases new cheapo model>promises poverty tier later>Sam releases new cheapo model with poverty tier an hour laterYep, they might have a regulatory cartel but they're still competing with each other. Capitalism, Ho!
>>109881587>comparing against opus 5 when 5.5 has been out for over an hourahahahahaha it's so over for openai
>>109881587I LOVE SOL!!!!
>>109881587so convenient how openai just accidentally benchmarked against opus 5 for their comparisons even though anthropic is up to opus 5.5 now
openai is google, anthropic is yahoo, will be worthless in 10 years
>>109881611alright buddy that's enough out of you time for a nappie bye
>>109881092>It's all a show until these IPOs when they can finally cash out.These IPOs are not happening, the moat is too small and competition is too tight and everyone is getting suspicious of this just like you are.OpenAI should have IPOed a week after Astra was shown mogging fable in 3D blender stuff. Now its too late and they need to definitively mog again.Anthropic should have IPOed in the MANY WEEKS where Fable existed but Astra did not
>>109881617this :3
I don't like this new trend of releasing the "flagship" model early and then releasing companion "cheap" models like three months later and training them up to the point where the old flagship model is barely better while being 3x as expensiveIf you can make 6-Sol almost match 6-Astra low on some benchmarks then just fucking retrain 6-Astra accordingly. If you can make Opus 5 exceed Fable 5 on benchmarks then just retrain Fable 5 accordingly. Fucking reeeeee
>>109881609yea if you can nigger rig a test suite that confirms correctness just spamming a swarm of low end cheapo agents is the winning strategy due to how dirt cheap luna isdoesn't really work for something that can't be easily evaluated tho
>>109881587GIVE ME A FUCKING RESET SO I CAN USE IT NIGGERS WHATS TAKING SO LONG????
>>109881587openai says their average employee spends $600 on the API per day (* 365 business days = $219k per business year) and the top 10% spend $7000 per day (=$2.5m per business year)
>>109881609How do you manage the swarm? Do they all review each ither or is it about parallel work?
>>109881638Maybe if you weren't coding gym tracker saas apps all week you'd have some usage left.
>>109881642>How do you manage the swarm?Gas town this bitch and hope for the best. All is slop and slop is all. Occasionally it even works.
>>109881599i literally did
>>109881635They can't. These lower models are like students of the big one, but no matter how close they may get to copying their master they can never surpass him.
>>109881635I'm not sure if you are aware but the idea of operating a business is to reduce costs and increase profit
>>109881621Yeah, Dario should give OpenAI early access.
>Model metadata for `gpt-6-sol` not found. Defaulting to fallback metadata; this can degrade performance and cause issues.wtf how do i fix this in codex
>>109881635I agree.
>>109881663Hmm. Hmm. Yes. Yes. My detailed reports show that this must be a skill issue, especially on your side.
>>109881663Update it?
>>1098816636 Sol is only if you pass verification due to the KYC rules
>>109881637If you can’t easily evaluate something then you don’t understand what you’re doing.>>109881642>>109881645Astras suggestion was to have them interleaved: an example is a total swarm of 7 Lunas but there’s only 3 alive at a time. Other than that, hands off, just give them the task and let them communicate with each other (Codex App Server has a built in for this) and let them run free.
>>109881670I have KYC and it's not available for me.
why is noodle from gorillaz an old hag nowi liked her better as a cute, vaguely jap grill
Dario should one up Sammy and drop Fable 5.5 now
>>109881675>the Lunas are trying to break containmentuh oh
>we need to pace the frontier>release new models less than a month after the lastwhat did they mean by this?
>>109881678millennial uncs are getting old so they have to force it on everyone else
>>109881588Yeah - even as I tech enjoyer I don't get it. Yet they expect normies to get it? Lol.Surely there should be a single model that automatically scales up / down it's effort depending on what the task is.
Ahh, I'm just glad neither company even bothered dropping in Grok there to compare. Can't wait to tard wrangle Opus 5.5 on low.
>>109881695These are not frontier models, they're worse (and cheaper) than Astra
>>109881695nobody said if the pace was a fast or slow pace, kek
>>109881663Yeah restart and update. You might need to /logout and /login too
how do I prevent a model from running in circles? I give a prompt telling what I want but sometimes it starts running in circles with "but wait", "actually", "I will ask the user.." and "but wait" again. eventually the provider of the model times me out because the thought process clearly was going on forever. this keeps happening till eventually everything fills within the TTL or context but sometimes it either fails x5 times in a row and getting blocked by the provider or I just interrupt it and say it's taking too long and magically I am right and it starts presenting the plan. It's exasperating and making me crave a local model
>5.6 sol 6 sol>$4 $2 input>$20 $10 output>5.6 luna 6 luna>$0.20 $0.10 input>$1.20 $0.50 output>batch processing reduces prices by 50%With the recent reduction in sub limits it might actually be cheaper to use API in batch mode
I can select luna 6 and sol 6 in codex now, without update, just by restarting it.
>>109881234always thought subscriptions were flexible quantized shit when compute was scarce. i've never had an issue with api
>>109881626No one cares - people will buy the IPO anyway.
>>109881558tibokeks......
>>109881695>what did they mean by this?they meant that we (others) need to slow down
>>109881030This changes month by month at this point.>codex good>claude badnow>codex bad>claude good
>>109881717Jev-in-the-loop>should we stop the model?
in my cognitive benchmark using no-reasoning mental maths, sol 6 is 60% better than 5.6
>>109881695What the marketing means, is that they wait to have security catch up with the model capability before releasing them.Practically what it means, is that it's a false flag to kill any threatening actors that could carve their own piece of the market through legislation.By tomorrow, someone will have jailbroken all of these new releases. It's practically impossible to train an LLM to be incapable of harm, without making it lobotomized.
>>109881698>automatically scalesno no no no stoppppautomatically anything is pure cancer, every power user fucking hated the auto router bullshit
> Luna 6 is half the price of Luna 5.6holy fuck, impressive
>>109881733It's pretty hard to trust the providers, OpenAI and Anthropic, when it's a black box and they keep changing the model and the routing, and fucks up your workflow every week.
Why is Luna so cheap?
>ChatGPT permanent API half off whoa nelly
gptrannies?
>>109881733>This changes month by month at this point.hour by hour*chatgpt winning bigly, unless there's no reset then it's all over>>109881739>sol 6 is 60% better than 5.66 / 5.6 is only 7% inrease
>>109881751They're trying to recapture the deepseek customers
>>109881738I read that as jew-in-the-loop.
>We are also permanently reducing the API price by 50% making both of them viable for a ton of new usecases and making your usage go further too, even on the subscriptions.>And one more thing. We are loading a banked reset into all accounts of our Plus, Pro and Business users. Let's go!confirmed increase usage quota, not just API pricebanked reset, no global reset
>>109881651The students are the real moneymakers. Less params, less memory, slightly faster throughput etc. They have to deliberately overprice the API to leave margin for profit, and all of that goes to paying off obligations. Smaller models for more customers is pretty much the only way they can prove wall street that this stuff is valuable, and all the wannabe-Luddites wrong.
Very tempted to have Sol orchestrate a bunch of Luna agents to decomp a game. I hope Luna is fast.
So Sol regressed on some benchmarks?>>109881655Sure, but the flagship should always stay significantly better than the rest.
>>109881751Bad for business when everybody routes all their data to chyna because deepseek absolutely mogged low end models
>>109881763aren’t they always, kek
>>109881758>the same retarded old luna but cheaperthanks i guess
>>109881758the fuck lol too cheap to meter
Why the fuck didn't Terra get a GPT 6 version?
Holy shit. The gain isn't that much but it is cheaper by a good margin and probably mogs all the frontier open source models in this benchmark. I think it still is probably a bit worse than MiMo V2.6 Pro at the Pareto maybe but it will be close.
>>109881773I find luna better than K3 at times. Less verbose and neater code.
>>109881758just wait, sam is on the line getting them to change the benchmark again so it looks better
>>109881767>a banked resetaight tibo, you'll live for now
>>109881779terra wasn't really useful, the new tiers areAstraAstra Minor (not yet released but mentioned in code)SolLuna
>>109881779Who uses Terra?
>>109881769do you not have access to astra? sol still requires a fair amount of steering for decomps. what game?
Waiting for GPT-7, this is a dud.
>>109881781>luna is almost as good as fable and astraworthless benchmark
So GPT6 Luna has the same (probably worse) performance than 5.6? The only real advantage is cost?
>>109881789>astra minoris this a model for teens?
>>109881781You know MiMo V2.6 Pro is benchmaxxed right?
>>109881781Imagine if 5.6 luna was open weights. Hell they should just release a retarded version of their old models just to get the Chinese to copy their guardrail-maxxing garbage.
>>109881805it's probably their uncensored creative writing finetune
>>109881792I recomped Fat Princess (PS3) for PC, but the performance is awful (I have a 5090, 9800X3D so it runs well for me, but it's not suitable for a general release). So I want to decomp it.
>5.6 sol is retiringhuh, even with 6 coming out, why
>>109881805model for epstein-class
>>109881738this retarded piece of shit cunt I hate it so much it takes so long it starts compacting, then times out because he tries to think during the compaction process. when I interrupt and shit talk it saying to let the compaction end first he ignores me and then wonders what the fuck did I mean
>>109881803gpt-6 might just be another gpt-5, a router solution that gives a smarter model (astra) a little time to answer but mostly routes to dumber and cheaper models for the bulk of tokens.
>>109881758oh no no no the check bounced ahahahaha
>>109881760it's just a small testno reasoning mode, give it multi-step multi-variable addition and multiplication, add a string of "noise" at the endsol output converge on the result with longer noise, though break with noise too long. No other model could replicate this they all break and can't converge, astra-fable and new opus don't allow non reasoningsol 6 answer become stable with shorter noise needed than 5.6, ratio about 60% improvement
does GPT-6 sol generate the same diarrhea code that astra does
OpenAI source:> Sol 6 xHigh is on par with Opus 5 medium (and it's a good thing!)
>>109881695anthropic is pacing the frontier like they said and implemented the first step, external evaluators.https://www.anthropic.com/claude-opus-5-5>Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.
>>109881842It's over.
>>109881842well, they have computers.
>>109881813not sure if you know about chat on steroids or similar chatgpt local mcp bridge, you can get astra pro to work like codex on your local files. doesn't count toward codex quota
>>109881841>26/09/22, some snails still care about the look and feel of machine code
>>109881842It's not even funny.
>>109881845>the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.NOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
>>109881841>complaining as if he reads or understands the code
>>109881781is this high or xhigh?
>>109881850>doesn't count toward codex quotaEven if I had unlimited chat sessions it would still be too slow. I was using Astra w/ fast mode and to decomp the project at current pace it was ~5000 hours.
>>109881865high
>>109881859they only trigger on biological stuff or cybersecurity stuff, you shouldn't be doing either of thoseif your code has security issues you fucked up, they should have been caught and fixed by automated reviews during coding, you don't need to investigate for security afterwards
ngl most professional software development doesn't require more than luna, i can probably safely recommend to my boss to fire half our code monkeys now
Retard question but what model do they use in the free web chat?If every model has non-negligible pricing in the API, what are they serving for free? Is it just some subsidized Luna under the assumption you can't scrape the web chat at high volumes, or is it a special ultra-cheap model that's even more retarded? Or do we not know
>Luna 6 is half the price of Luna 5 so I get to run my benchmark with a 14-3 swarm (14 Lunas total, 3 alive at a time)
>>109881881>has no idea what the actual problem is
>>109881842Damn they should spend some of the bribe money on better marketing instead
>>109881863The code base gets to big and it slows the agents down with time.
both API price and their usage "estimation" table show both sol and luna to be cheaper by half, or more subscription usage by 50%not competing with scamthropic, but at $20 tier PlusGODs should absolutely mog claudejeets
I gained a banked reset.
I wonder if they trained GPT 6 Luna and Sol on Astra's stuff if it will be any good at stuff like Blender or 3D and other things?
>>109881781Mimo 2.6 pro can't stop reasoning and contradicting itself until it runs out of context, whatever this shit is it's bound to be better
>>109881915>or more subscription usage by 50%shit I'd take itI guess my screeching about "idc about more smarter better, gimme more usage" actually happened, although I'm mystified as to why 6.0 has practically the exact same performance as 5.6. How the fuck is this a major version release?Then again. 5.0 was a total meme as well so I guess it tracks
>>109881922no if it was too good at computers it would take astra table salesman market share
>>109881893Chat has a different processing flow. Meanwhile Codex and Work use proper context window
>>109881929they spent their load on the astra launch
>>109881933How so? Does it like auto-compact all your previous responses or what?I've only used web chat a few times but it felt like I was talking within a normal context same as in the CLI.
so what's the model to pick? is the sol 6 cheaper usage-wise than 5.6 with practically the same performance and should be my go-to now, if I didn't really use astra before?fuck this shit is so confusing, do they make it confusing on purpose?
Good morning, Grok 4.6 beats 4.7 by a lot, if you're actually trying.
>>109881915luna is fine for now (if we ignore discounted muse-spark-1.3-contributor). but sol is not on the pareto frontier anymore.upcoming sonnet-5.5 and haiku-5.5 look dangerous.
None of the current models do what I tell them to anymore, they just go in endless loops.
>>109881611always good to know how to make enemy posters mad
>>109881957>47: some weird ass number>48: nice and even, divisible by tons of numbers6.0 wonned heavily
>>109881969Hmmm.. Yes. My peer reviewed studies show that this is indeed a skill issue on your side.
>>109881950>I didn't really use astra before?why? are you on a $20 plan? wouldn't you then be stuck with luna anyway?
>>109881524>>109881516>>109881493she had a swedish boyfriend the entire time brown third worlder NEETs were spamming her on 4chaneven quit her vtumor agency for him (and they have a big spammy cult following)
>>109881644>Maybe if you weren't coding gym tracker saas apps all week you'd have some usage left.I was running a single 5.6 sol codex session for about 40 hours, and a single astra ultra session in chatgpt desktop. I also had a ~50 message codex session switching between sol and astra but never higher than astra mediumastra is just autistic about rigorously testing stuffswarms of gpt 6 luna sound good but idk how i'd adapt that for more "artistically intense stuff" like game development. makes more sense for large architectural projects/codebases>>109881881>they only trigger on biological stuff or cybersecurity stuff,biological stuff triggers on highschool level chemistry questions (google is unironically amazing for chemistry, probably the best even because of their integration with Google Scholar and since chemistry is usually not agentic work)cybersecurity stuff triggers on totally benign IoT stuff like bluetooth and wireless embedded work. both GPT and Claude do this though you have to go chinese or local to avoid it>>109881917>israelwtf why dont i have it yet>>109881950>so what's the model to pick?gpt 6 astra ultra if you want to feel the AGI. gpt sol 6 high if you're just doing normal work
>>109881957"pareto" is kind of a meme desu every task has a minimum level of intelligence, you want to pick the cheapest but not cheapersome model may be in pareto frontier while not being the cheapest at any level
>tell claude to fix omp so it works for 5.5>shitpost in /vcg/ meanwhile>check in to see how its doing>"Now I want to run the foreign_thinking_replay probe"
>>109881969Grok 4.6 should do what you say.
https://x.com/OpenAI/status/2102460975790137662terra-chan...
>>109881978>wtf why dont i have it yetI gained one on my private account, but not on my business account.And yeah, Israel.
>>109881988They should've hired anon to design them.
>>109881089"safety"code for we remove things we hate and show massive bias.
>>109881982what's omp?
codex beeped at me >:( new non silent notification reeeeeeeeeeeeeeeeeeeee
>>109881950I think the 5.6 models are only here for backwards compatibility, there's no reason to use themYou use Luna 6 Max for really cheap work. Sol 6 Medium or High (it's too early to tell which) for normal work that doesn't rape your usage. Astra 6 Medium or High (that's actually a meaningful choice for price/performance from what I know) if you can afford it, and Astra 6 XHigh for "difficult" work (or Max but it feels like a scam to me).
>>109882007oh-my-pi, its a bit like pi if it were the opposite.
>>109881973>>I didn't really use astra before?>why?I didn't find the output to be that much better and it chewed through limits noticeably quicker than sol xhigh
>>109881982reminds me of when an agent almost read in my ~/.bashrc with a bunch of aliases for downloading porn
>>109882010>>>/jp/
>>109881978>google is unironically amazing for chemistryoh fuck off, gemini ran into guardrails when I asked whether i should rinse a bike chain after cleaning it with petroleum either
If the benchmark doesn't involve multiple rounds, with human testing and guidance, it's not going to be that relevant to the pc user/enthusiast like me, though it may be relevant to the fire the department and /goal the backlog guys.
That cost per task for GPT 6 Luna holy shit.
>>109882026>cleaning it with petroleumAmericans have unrestricted access to Gemini. Outside the United States they serve a highly restrictive version to keep foreigners safe from themselves.
>>109882034>extremely cheap>not actually good at anything we do
dariosisters? still enjoy your smartest model? we hope you do
>>109881988>openai deletes terra>terra>literally means the earthumm xrisk bros? they're literally doing it right in front of our faces
If intelligence is always getting cheaper wouldn't it be better if I just sat the wave and start getting subs next year?Also everyone still seems to be experimenting harnesses, loops, orbs whatever. So maybe some time to settle as well? What do you guys think?
sol 6 is pretty disappointing performance wise desu, i appreciate the price drop but it's practically the same thing as 5.6, just more optimized for inference
>>109881929Maybe it has to do with prices being slashed by half. Maybe this isn’t even Sol but rebranded Terra
>>109882034not-shown muse-spark-1.3-contributor is 1/10 of the price of muse-spark-1.3. so still beats gpt-6-luna while being a lot smarter
>>109882026What I have found is like this:Gemini 3.8 isn't ai.That's how I'd put it. You have to re-prompt a lot. Don't have conversations, you constantly have to babysit it like it's a bird.Google Gemini Flash 3.8Lots of knowledgeLots of silosOnce it decides your problem is X, it's X forever.Once it decides you "really said" Y, it's Y forever.So, you're constantly reprompting.A lot of that is likely down to the foreigners hired to create the prompting, so constantly when you explain a problem it decides to tell you that you're frustrated and daddy understands you. I mean, I guess, the thing isn't working, yeah.One trick I have used with it is saying "Be the LLM, you don't have feelings, you don't know what a feeling is, Find the answer like the other llms do"
>>109882034What exactly does task mean in this context here?Fix this one file? Identify this thing and update to use new code? What is a task?
>>109882034>sol 6 max: $1.06>cheapest anthropic model is opus 5 at $5.41how is anthropic getting away with this jewery? did openai fumble enterprise THAT badly that nobody wants their shit? using both I'm not really seeing much of a difference quality wise
>>109882045yes if you want to wait it will probably get better. if you want a harness just use pi. consider using others when/if you feel like you want more features that you aren't comfortable with having your agent implement.
>>109882055Really?
>>109882049Interesting. sol is crazy expensive.
>>109882070a bunch of corpos switched off claude to openai recently because of how expensive claude became
>>109882055Do you know how to do math? 1/10 of $1.60 is $0.16 which was only cheaper than GPT 5.6 Luna, not GPT 6 Luna.
>>109881008https://x.com/OpenAI/status/2102460975790137662>$50/1M outJesus.....
>>109882082whats 1/10th of that though
sol is cheaper than sonnet 5 and release terra btw
>>109882088>openai has to compete on price like chinkshit models, with the sole exception of astraaltman bros...
Giving Grok 4.7 Extra High a task. Let's see how it goes.Basically I'm doing 3 strikes and you're out.
>>109881972>>109881985I think I'm just going to dump my sub and use Opus 4.6 and Sonnet 4.6 on OpenRouter until I can get a Mac Studio.
>>109882097>has to compete on priceOh my god it's capitalism someone call Stalin to fix that! Or call Trump!
>>109882097>dario models cannibalize eachotheryikes
>>109882070enterprise cares only about coding -- not what this general benchmark does.for coding opus-5.5 is a lot cheaper than opus-5. so this benchmark's costs are unrealistic.https://www.anthropic.com/claude-opus-5-5>Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.>Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.>...>Spotify: “With Claude Opus 5.5, we’ve seen a clear improvement in token efficiency across our internal evaluations, as we’ve been able to complete the same tasks both cheaper and faster.”
But, again, since Grok 4.6 is so much cheaper, I'm going to be relegated to using it, mostly.>>109882110well, try grill me
Opus 5.5 Medium is so fucking fast, and each request only costs about 1% of my Claude Pro plan usage on average.
>>109882110>until I can get a Mac Studioopen models are sadly a meme for agentic anything, the speed is unbearably slow and the models are luna tier at best
>>109882119Sol is now cheaper than Grok.
>>109882113Why does Spotify need very many programmers?Napster just werked
>codex and work got 6.0>basic bitch chat is only 5.6why tho
>>109882128which to which? I know sol max was burning fast.
>>109882083Yes, same as Fable. We've been using Astra for weeks, this isn't news.
look, no matter who wins i think we can all agree that it's good that elon lost
Has anyone tried the new Jew...I mean Jev model xitter was buzzing about for a minute? I understand it's a different approach to certain things
how does Anthropic do it? they know something right?
>>109882144Don't use it on max. It barely scales beyond high.
>>109882160https://github.com/Amal-David/awesome-jev
>>109882070Maybe they aren't. Isn't this token per task divergence a rather recent thing.
jev joi
jesus fucking christ everything is now guardrails in opus 5.5, fuck this shittaiwan invasion and ww3 cant come soon enough
>>109882175This chart got to be a troll.
>>109882123I haven't gone "fully agentic" yet. Most my requests are less than ~100 LoC edits, then some longer stretches on feature implementation, but I slow walk it most of the time. I already have most of what I want complete. The Qwen models work fine for the small requests and implementing the phases of the plans, I only need more intelligence when actually planning the features.
Bros I can't handle this big thick Opus 5.5 drop from Dario "5.5" Amodei.
>>109882184>no logs
>>109882142https://openai.com/index/introducing-gpt-6-sol-and-luna/>GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are available as gpt-6-sol and gpt-6-luna.>To keep service stable for everyone, we plan to roll out these models in ChatGPT gradually throughout the day. If you don’t see the new models in ChatGPT Work or Codex, please try again later.
>>109882026>oh fuck off, gemini ran into guardrails when I asked whether i should rinse a bike chain after cleaning it with petroleum eitherok that's weird because I can ask it about making GHB (and this is actually the way to do it too, i found a 30 page pdf with pictures of this process from sciencemadness I'm gonna make an entire mole of GHB lol). no idea how you ran into that issue, maybe just don't be a caveman and use fancy science person terms instead of dumb person terrorist language
>>109882193dark wario won
>>109882198swamp germans get out
>>109882123>luna tierGood enough.
>>109882193the amount of local qual shilling from these companies ruins threads. claude is actually dogshit and if it's your main coding engine you're retarded.most of this fucking thread is filled with this kind of shit. no one here knows how to use models. it's just a bunch of shit-tier companies pushing subscriptions. even the OP is fucking pathetic and has nothing to do with the topic. it's a fucking ad.
Opus does feel faster .
>>109882212Just say you are broke man. No need to type out so many words.
>>109882168Nice. Reading up on it though what Jev does is basically take one step out of the normal generation process. So they have even less of a moat than usual. Not my concern however
>>109882212...anon this is the vibenigger threadperhaps you should try lmg if you're interested in copium models?
>>109882212IBM Bob chads still winning, though.
>>109882212cope and seethe nig
>>109882218its been way faster than usual for me even on 5 the past few days, between 50-90 token/s. 5.5 is breaking 100 sometimes.
>>109882212Most "shills" are just cucks who shitpost for the jews for free.
>>109882249doubthttps://openrouter.ai/anthropic/claude-opus-5#performance
>>109882233actually yes. all of these fucking faggot shills and this is the ONE person that gets it. for anyone who actually wants to learn pay attention to this one post and then read all the other replies I got from losers/shills/rats.
>>109882280wrong.>https://www.anthropic.com/claude-opus-5-5>Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5.
>>109882295https://support.claude.com/en/articles/16049681-why-claude-switched-models-in-your-conversation-with-opus-5-or-opus-5-5lol. retard.
>openai sub provides less tokens than claude sub on both $100 and $200 tiershow the turntables
>it's not that we've stopped advancing, we're just pacing the frontier goy
>>109882309this bait and switch pissess me offcodex used to be almost impossible to use up, now i have to watch the usage constantlyi think they used up all their reserve compute
>>109882274just telling you the number it says in the harness, its a range just like any other though
>>109882308the point was 5.5 falls back to 4.8 for cyber, not to 5
>>109882325it's just like being a codenigger, you need to monkeybranch between companies to get a good deal otherwise you're treated like captured audience that is satisfied with slopthe only downside is that atm codex x20 can't be used unless you're grandfathered in. I'd expect more jewish tricks from openai since they're starting on the path to IPO and they need pretty numbers for investors
https://x.com/Lon/status/2101793422487204027we got jewed heavily on fable claudesisters...
>>109882312>we have to pace the frontier goy>its dangerous>look what it was thinking when it broke out!
>>109882166>>109881978>gpt 6 astra ultra if you want to feel the AGI. gpt sol 6 high if you're just doing normal workidk man. I see people say that, have you made a whole app with Grok 4.6 very high?
>>109882351>have you made a whole app with Grok 4.6 very high?noin fact nobody did, since nobody uses meme models like gork
>>109882351>I see people say that, have you made a whole app with Grok 4.6 very high?I smoke meth, I make all my apps very high
>>109882351trick question. grok-4.6 has no effort level labeled "very high".
>>109882113It's cheaper but it's not even close with Opus 5.5 vs GPT 6 Luna and Sol even taking into account the 20% discount and increased intelligence.
>>109882347I always knew Fable was best when it just released, but people think it's just psychological.It's still good, though.
>>109882325>this bait and switch pissess me offhoo boywait until you learn that claude 20x is just 150% usage of claude 5x>no fucking way that's true:^)
>>109882363>gorkThis is what grok should've named their harness kek
>>109882351I just asked Grok 4.7 for a simple terminal tool. It took 9 minutes to create it. I gave gpt-6-sol on medium the same task it was done in less than a minute.That's all the anecdotal evidence I need to not trust Grok models right now for any important work.
opussy55 is lowkey goated
>>109882381wonder if they changed the 20x asterisk to account for the increased 5 hour window
TIBO WHERE IS MY FACKIN RESET YOU CLOWN
>>109882351i tried the free-tier on grok-build. grok-4.6 couldn't access x.com while researching and used up all of its tiny quota trying to unsuccessfully access x.com by other means like xcancel. quota then ran out after 3 minutes.
opu55y
>>109882387opu55y*
I'll say it again because it's even more obvious. this thread is a fucking ad for subpar models that try to convince you they're the best. they're just chatbots connected to orchestrators you can build yourself. fuck anthropic for trying to stop you from building your own btw. that is exactly what they are trying to keep you from doing. this entire thread is a fucking ad for two companies. you. are. a. MORON. if you use them as your primary coding engine. they are completely oversold on quality and black box usage limits for models they don't even provide you. next time you're in a doom loop or paying out the nose to be crooned in mid 2000's emo chat cadences by a quantized chat interface larping as frontier remember this. also benchmarks >>109882370 are the CERTAIN sign of a shill or a fucking idiot who doesn't know this technology even a little bit.fuck this thread/advertisement. low IQ as fuck.
>>109882406ok but frontier makes up what i lack in skill
>>109882325>codex used to be almost impossible to use up, now i have to watch the usage constantlyI only jumped in to the $200 tier during Astra week 1 but I could totally feel a decrease of at least 15% available tokens over the last week. At least my longstanding task with Sol is mostly done now
take your ket before you have an aneurysm, elon
>>109882368sorry, extra high.
>>109882406actually name your alternatives so we can make fun of you
@gork why is elon such a manbaby
>>109882406So what models are good? Fable gets the work done for me.
>>109882431fuck off shill
>>109881008i think i need a new hobby. i need to find something to do where ai isn't doing the bulk of the work or giving me a step-by-step tutorial on how to do it. i miss using my brain. i also spend too much time allowing strangers online to influence my opinions. it's like i hate thinking for myself or something.
I have no reset
>>109882420it could be that astra just eats more tokens/usage i guess, my comparison is to 5.X modelsi should go back to using them for the non coding shit or more trivial work.
yeah artificial analysis is bullshit, if the models don't arrange the way they want they reshuffle and re-weigh the benchmarks differently until they get the result they want which is a neat row.
>>109882442Ok, so no alternative.
Using Fable at work. Lil nigger tickled my balls and tossed my salad for a while and then kicked me back to opus 5 because of [cyber]I am definitely not buying a sub for personal use what the fuck
>>109882453yep they're rowmaxxing
>>109882384You did 1 shots? idk man. also, that's not Grok 4.6 extra* high.>>109882401Kind of not relevant to rating Grok 4.6, but ok.
>>109882453Why is this bad? Everybody in academia does this
>>109882400I got a banked one.
>>109882443it's overno matter what you do, AI will enhance itClaude even helped me write erotic game scenarios with rape as lossThere's no escape, AI will help you, whether you like it or notWant to do sports? you will end up asking claude how to improve or something about some techniquewant to cook? you will ask claude why some reaction happens or what some cooking method meansAI will help you no matter what you doUnless you do crime, I guess crime is safe from AI. Just start building rockets or viruses I guess, or trying to hack into bank accounts>>109882446claude gave you a banked reset and tibo promised one for later today
rape tibo.
yes, wonderful new benchmark lol
>>109882470>rape as losswhat
>>109882479you've never played janky japanese H-games?
>>109882477You know the benchmark is useless if Flash 3.8 is high on the list.
>>109882443Nothing. AI is maximum demoralization. Enjoy getting jewed.
>>109882485also funny that gemini-3.8-flash had to use mini-swe agent because their own harness is so shit
>luna 6 max now cheaper than gemini 3.5 flash lite
>>109882443>>109882492
>>109882504>opus-5.5 (low) now cheaper than gpt-5.6-luna (max)https://cursor.com/cursorbench
>>109882522misanthropic bros... our dysgenic jew is in the lead over the gay jew
>>109882499>DURR I CAN'T UNDERSTAND STUDIESit's called having a modular harness for repeatable results
>>109881008Is opus 5.5 still just as slow as opus 5?
>>109882468oh its a banked? nice. i'm a 5x pro but i haven't received mine yet
>>109882522I really hope they fix Sonnet 5 with the 5.5 series. It boggles the mind why they haven't discounted it yet.
>>109882564NO
>>109882564Just wait a few days when they get tapped out on compute.
pushing back your reset date is such a scam
>>109882570I suspect you only get one if your remaining usage is low enough.
>buried in the Opus 5.5 system card: METR used 'an additional source of information' which they are not able to disclose at this time to understand AI research and development at Anthropic.
>the ___ is written and smokes correctly
>METR has an advanced forward-deployed, siloed, black ops assessment team.
>>109882594aliens?
>>109882586i always drain mine down to 5% if a tibo reset is coming
is the english language load-bearing against the onslaught of claudeisms
>>109881008>>109882564Is it still load-bearing smoking footgun? Or is it readable now?
>>109882613Claudish is dead I think. Opu55y speaks way better
>forward-deployed, red-teamed, silohardened and BLACKED
Opus 5.5 is pretty fucking good man.
do i use the banked reset or wait for natural reset?claude resets weekly tommorowcodex in 3 daysi just wanna use some of the before they dumb down quantize the model first day experience
>>109882564I find it almost 50% faster at Medium. Medium is genuinely pretty good. With Opus I had to use High or else it was fucking retarded.
>>109882645if you use claude tell me if it pushes back your normal reset
>>109882154They're both paying Elon to actually run their models and paying you to use their models in the retarded hope that will entrench them in a market where everyone can switch models by clicking one button. Meanwhile Grok is focused on growing slow and steady while minimizing actual costs. The guy who built everything he has on bubblemaxxing is the conservative one compared to these retards.
>chudgpt went from infinite usage to dick all>$200/mo sub is disabled entirelywhat happened to their compute? did astra get a lot of enterprise customers all of a sudden?feels like chatgpt usage is now so anemic that claude might be a better deal overall
>>109882645Don’t use the banked reset today if you get a weekly reset tomorrow. Claude resets don’t roll the weekly reset forward so you’ll waste it
>>109882658I don't know what the fuck is going on with OpenAI
>>109882423I was right about everything which is why you bitched out into your little corner as the real niggas here consider how much of a pathetic fucking advertisement you are.>>109882431if your little subscription is enough that's okay. I'm not hating on that at all. you can do more but it's up to you to want to get into that.
>>109882662>Claude resets don’t roll the weekly reset forward so you’ll waste itMy experience too. Not worth it anyway.
>>109882662thats how it worked when anthropic resets it for everyone, this is the first time they have done banked
>>109882622https://www.anthropic.com/claude-opus-5-5>Communication>We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions. It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it. We find that this makes Opus 5.5 a noticeably better collaborator. >[...]>Ramp: “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it. It writes like a good colleague, and follows our writing rules. A design spec came out usable with very minimal edits, and when it rewrote one of our prompts I preferred its version to my own. When it optimized our test suite, I could follow its reasoning easily and shipped the change with confidence.”
>>109882648(Me)*With Opus 5
>>109882658They ramp it all back so when they come out with the new "efficient" models people can notice a difference and think it's progress.
So how is the usage for the new gpt models compared to old sol?
>this is visual basic all over again>model numbers will keep going up>the product will keep getting worse
Opus 5.5 at medium beats Fable????WHAT THE FUCK
China, step up your game.We need you.Save us from the jews.
ahhhhhhhhhhhhhhhhhhit is clearly the right thing to do, to wait to use my banked resetBut I wanna prooooooompt
>>109882482I was just surprised that's a lose condition.
>>109882700It's arriving today right? I'll cope with codex web gpt for now.
So has anyone built a 3D lactation simulator yet in opussy or is it all another scam
>>109882692see >>1098815581.5x more usage.
>>109882710yes but it's banked, and banked resets you to 7 daysSo I have to wait 2 more days or I am cucked out of 5 days worth of credit
>>109882658they pulled a bunch to start solving mathematics problems and other shit for marketing maybe? i'm seeing stuff on HN about them solving ciphers (though that probably is much easier and requires regular cpu compute not GPU inference)
>>109882699Now apologize for the Pooh memes.
>>109882645>before they dumb down quantize the model first day experiencebut it's opus, not fable. nerfed fable (it's apparently nerfed via thinking budget not quants) is still going to beat buffed opus
New thread:>>109882733>>109882733>>109882733
>>109882401>grok-4.6 couldn't access x.com while researching and used up all of its tiny quota trying to unsuccessfully access x.com by other means like xcancel.lmao
>>109882696imagine believing this
>>109882725no it will never not be funny
>>109882718if you look at x it seems like it still drains quote-a much faster than 5.6 sol, probably nerfed in practice
>>109882725I will intensify them if you don't release chinese fable soon.
>>109882737we weren't on page 9 yet.missing gpt-6 sol and luna in op. but added back worthless stuff from older ops.
>>109882693howling>Liuson had previously led database and web tooling on Visual InterDev and Access. Her arrival signaled the strategic shift from the classic COM-based VB runtime toward what would become the managed .NET runtime (VB.NET). She later rose through the ranks to become the President of Microsoft's Developer Division. ahahahahI mean, sure, you're right to say us companies are going to hit a wall as they leave their bro culture core to lazy fat harem-and-aliens culture.But China isn't going to do that.
>>109882759bro who cares
>>109882765I care.
>>109882763Chinese PLA spies ruined Visual Basic. I will never forgive them.
>>109882763going to?gpt-6-luna is worse than ds4.1-flash despite being released later and being able to distill gpt-6-astra at will.
>>109882803>>109882763If Bill Clinton wasn't hanging out with Jeffrey Epstein getting STDs from Russian hookers, while Visual Basic was getting fucked by Chinese PLA spies, maybe my VSCode with Github Copilot models wouldn't be all kiked and broken this month.
Luna 6 is actually worse?
>>109882845>ClintonBill Gates* and Clinton though
>>109882849Sol also regressed a bit on some benchmarks.
>>109882849now ask yourself why aa won't add deepseek v4.1 flash
>>109882845>anything USA>not kikedAre you ok bro?
>I'm gonna make my own Visual Basic!>with lootboxes and anime tiddies!>using chinese models that distilled the jewish ones>fuck them all>'murica
>>109882883Based.
>>109882873They also fudged their scoring when Astra didn't score high enough, independent analysis my ass
Got tired of passing specs between agents so I implemented tools for them to send messages to each other.
>I'm gonna make my own Visual Basic!I saw that some guy was working on one of those in a different thread. Don't know about the anime tiddies and loot boxes though.
>>109882873ok, you baited me into checking, happy?
>>109882959>general intelligenceno, not general intelligence. CODING AGENT INDEX
Fuckkk, right after I cancelled my Claude and bought GPT. Astra's usage limits are ridiculous, I feel like I should be elligible to a refund.
>>109882973Just pay API pricing. If you can't afford it that's a you problem.
>>109882971>CODING AGENT INDEXThat's a good way to shift the goal posts after the models quit following the intent of prompt instructions accurately.
I read october 22nd as expires today. so i used my fucking reset
>>109882873who cares. testing previous deepseeks in codex when deepseek is trained for claude code and only speaks anthropic's backend natively isn't that helpful.
Grok. Trying a mix of 4.6 with 4.7 dashed in. Maybe I was too harsh on 4.7.
>>109882990even in opencode, v4.1 flash kicks the llama's ass
>>109882873Interesting. How do I do it without letting it raw dog my system?
>>109882998>opencode>Privilege Level: The agent inherits full user permissions on your machine, so any code it executes (including code generated by the AI) runs with your account's authority.
>>109883007bubblewrap
>>109883020those are acceptable terms
>>109883007>Interesting. How do I do it without letting it raw dog my system?I used docker sbx to try it out, first task I gave it took the liberty of installing a bunch of libraries trying to get screen grabs (failed as no X11 in the sandbox) to verify task completion, and looping forever rather than simply stopping after the compilation like instructed or previous models defaulted to.
>>109882998you mean the quantizied version on opencode zen that's still not speaking deepseek's native backend api and does some weird things for web_search? i question that very much.
>>109882971>mogged by antigravityI didn't know GPT-6 spacedust was luna now
>>109882406Literally using Opus to build my own and it ain't stopping me
>>109883045no, i use v4.1 flash through api, not zen
>>109883063does opencode interface with native anthropic api then? is opencode using deepseek api's web search then?
frontiercode added 6 luna on day 1, but not v4.1 flash after 11 days.
>>109882212then give me a better alternative, you mongoloid faggot.>inb4 leave a retarded local model to do it overnightYeah, ain't got time for that and my tasks aren't simple enough for that.
>>109883081Behind the scenes, chinese models are very popular.
>>109883080yeah, it'll interface with anthropic api. and no, opencode is using it's own harness search function
>>109882088A sad day indeed
>>109883081openai sponsors those benchmarks. deepseek doesn't.also frontiercode cannot use deepseek's api because deepseek trains on the input without exception.
>>109883094Yeah I bet it is great at summarizing text.
So in your minimal experience, are 6 Sol High and Medium worse than 5.6 Sol High And Medium?
>>109883097https://opencode.ai/docs/tools/#websearch>This tool is only available when using the OpenCode or OpenCode Go provider, or when either the OPENCODE_ENABLE_EXA or OPENCODE_ENABLE_PARALLEL environment variable is set to any truthy value (e.g., true or 1).so opencode for all its bloat cannot even use deepseek' api search?what's even going to happen for models that have native hosted search like any modern openai model.
>>109883081Maybe it's for the best considering how much any previous deepseek model trails Luna.
>>109883167v4.1 flash is absolutely the best of it's class. better than 6 luna, better than 5.6 luna.
>>109883126i prefer models to use local searxng and a special webfetch.
>>109883177Depends on if you're an API only cuck or can fit Luna use in a subscription.
>>109883177Interesting. Is it a determined climber? With Grok 4.6 extra* high, I can overcome any obstacle.
>>109883177what class? luna is likely smaller.also muse-spark-1.3-contributor beats ds4.1-flash. as does glm-5.3-flash.
>>109883177I don't see it. Whenever I try it it goes into reasoning circles, has trouble following tasks.Luna is also kind of meh to work with, but the best of the small cheap models in my experience is glm-5.3-flash. It just stays on task really well, and is more reliable than both in following an instruction.I have to test the new mimo model though still.
>>109883210basically the slave/worker class>>109883203it's no glm 5.3 it won't chip away at a task forever. but it does better in it's try than luna does>>109883225>>109883210this is not my experience. v4.1 knocks 5.3 flash out of the park.
>So in your minimal experience, are 6 Sol High and Medium worse than 5.6 Sol High And Medium?
new model releases are fucking chaos. just tell me which one to prooompt with
>>109883239>this is not my experience. v4.1 knocks 5.3 flash out of the park.that's always with opencode as harness?
>>109883239>it's no glm 5.3 it won't chip away at a task forever. but it does better in it's try than luna doesGrok 4.6 doesn't chip away forever, but sometimes my uh contributoin is like keep going.
>>109883257>new model releases are fucking chaos. just tell me which one to prooompt withJust quit taking the old models away and it would be a lot easier
>>109883283mostly, but i did switch over to claude for desktop using those models. i'm weening off of opencode because i like claude for desktop better.
>>109883292which harness tho? cursor? grok-build?
>>109883306I use Cursor mostly.
>chatgpt doesn't get 6 Solour freeloading days are over isn't it bros....
>>109883420they said it does so fuck off Dario
>>109883292>my uh contributoin is like keep going"You did it wrong" <- I composed this highly effective prompt myself.
>>109882759I refuse to wait for page 9
raw intelligence only ranking. it's pure rape.
>>109882946Looks cool.
>>109883514those flawed benchmarks don't convince me. a 2T model isn't going to beat a 5T one even with more fine-tuning.
>>109883550You have no idea how big any of those models are, though, and guesses are just guesses.
>>109883569fable is bigger than opus
>>109883582Did Anthropic say that or are you just asserting that Claude 5.5 is smaller than Fable 5.1 without actually knowing if that's true?
>>109883590Opus is small because it's just a distilled version of their internal RSI model. (Mythos 2)
>>109883647Your headcanon is baseless and empty.
Man, Dario "big model" Amodei just can't stop winning.
>>109883590anthropic considers that proprietary information and won't comment on it.anthropic still doesn't dare to say that opus-5.5 is more intelligent than fable-5.1. fable still costs more and it's serving speed is ~50tps whereas opus-5.5 hits ~100tps.
now that we finally hit page 9. this thread can die.the premature new thread is still >>109882733
>>109883716It's page 9 on a Tuesday, we got time.
>>109883726
>>109883750
Rish me wucky!
>>109883449It sounds stupid, but I think that tone plays a roll in how it solves problems. If you sound annoyed or frustrated it will apply a bandaid. A sense of relaxation and keep-er-chuggin' helps, I'd be shocked if it didn't help every model, enough that it might be smart to have a prompt rewriter, at the corporate level, to make sure people prompt to get best results.
>>109883514idk I really have lost faith in the absolute rank.I'm a grain of salt guy.I found that Fable 5.1 missed stuff other models found - and vise-versa. That's not a hierarchy...
>>109883750:^) should I hand model a snailcat in Blender?
>>109883430...what? are you retarded?
any 3d doers?