A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.You use Git — right, anon?## What “vibe coding” is, and how to do ithttps://simonwillison.net/2025/Mar/19/vibe-coding/https://simonwillison.net/2025/Mar/11/using-llms-for-code/## News (both past and future)- 2026-09-02 — Google Gemini 3.8 Flash out: https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/- 2026-09-01 — Claude Fable 5.1 out: https://www.anthropic.com/claude-fable-and-mythos-5-1- 2026-07-24 — Claude Opus 5 out: ## Related generals>>>/g/lmg/----## Frontier models using fully-general tooling — start here if you have $20 or sohttps://claude.com/product/claude-codehttps://developers.openai.com/codex/cli## Near-frontier models for codehttps://x.ai/cli## Not worth it for code, but maybe good for interpreting images/videohttps://antigravity.google/product/antigravity-cli----## Promptinghttps://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/https://arps18.github.io/posts/claude-code-mastery/## Skillshttps://github.com/mattpocock/skills — /grilling is a favoritehttps://github.com/DietrichGebert/ponytail## Other editors / terminal agents / coding agentshttps://osaurus.ai/https://pi.dev/https://opencode.ai/## Is our AIs unlearning?https://aistupidlevel.info/## What we’ve donehttps://vcg.gitgud.site## Previous thread>>109706156
What's the new abuse meta for poorfags now that ChatGPT has been nerfed into oblivion?
>>109712834>now that ChatGPT has been nerfed into oblivionIt hasn't?
>>109712834> claims ChatGPT was nerfed> refuses to elaborate> leaves
>>109712847I would chalk it up to myself if it wasn't for multiple people reporting the same thing in the last couple of threads.
First for omp
>>109712847>coping
>>109712853It feels like they've been reducing the context window and thinking but it might just be me idk. It has been making some dumb mistakes for me lately.
Claude Yakub 6
>>109712819>still believing in 1M contextask a model to give you a review of a project, and then in a separate session (with a proper harness) ask a model to use sub-agents to review your projectthe sub-agent one will always be better because the models will be looking at less code
so I tried gemmi 3.8 and it's really fast but also was as expensive in about 1h as glm in a day (2.4M tokens vs 28M tokens). It did some pretty intelligent things however. I wonder how it works on bigger projects? could be an interesting alternative to codex & claude
New eval just dropped
Claude Torah 7
what this tells me is deepswe is getting hard saturated
>>109712865That happens every time a new model is about to come out.
>>109712888Not him but for me it's the opposite, agents return some stupid out of context claim and the main context believes them.
>>109712916Let's hope it's just that
>>109712899pi doesn't even have an agent loop by default. how's it passing any benchmark? can this "benchmark" be one-shot?
>>109712901DeepSWE worked well at some point, but after Opus beat Fable it was clear it no longer works.
>>109712949>fablelmao only use it if you want your home folder deleted
>>109712865NTA but I guess I haven't noticed because my AGENTS file and implementation plan is designed to viciously tardwrangle the fucker
>>109712955Opus already deleted our staging database. You get the backups and you move on.
>>109712955>tell the model you think it did wrong>it spins up an agent with fresh context to check if it's actually overcomplicated >somehow upset by this obvious turn of eventsPoster in your pic is dumb.
>>109712899Mogged! Oh wait, wrong axis.
>>109712899the trick is to only do successful tasks and never fail
>>109712980holy cope, how much is anthropic paying you?
>>109712992Holy snailcat, do you also get mad when Fable decides to write tests without you asking?
>>109713009an adversarial review is not a test nor is it what the user wanted
>>109712899what the fuck is this graph holy shitI swear they manage to make it impossible to understand on purpose
>>109712960Well that's one of the issues, I don't know how to use an AGENTS.md file with ChatGPT. I can tell it to read it but it'll forget it 5 minutes later. And the custom instructions don't seem to help much.
>>109713024So it's just supposed to sit there and go "mhmm you right" when you tell it that you think it wrote a convoluted mess?
Claude Talmud 8.
wtf, did chatgpt just give my message a thumbs up? openai, what the fuck is this nonsense?
>>109713104that's called AGI. soon enough it's going to start speaking zoomer slang. skibidi chatgpt
>>109713104Positive conditioning for its meatbag
>>109713111It already can speak like that
>>109713123Problems?
chatgpt 5.6 sol xhigh is my only trusted resource for "really hard" concurrent programming that i can run more or less continuously on the $100/mo plan, where the effective rate limit is how fast i can review its commits and steer or sometimes correct it. plus other agents in luna, terra, sol medium etc. the best overall lineup of agents and codex mostly just works. i gladly upgraded from $20/mo after demoing itgrok and kimi are both strong, i think i prefer grok overall mostly due to wall time. but both gobble up all your usage as soon as the going gets harder than a layup luna or a flash model could handle.glm 5.3 and dsv4 pro are worth it for somewhere in the middle, i use omp for glm and ds harness for the other
I should quit vibe coding and get a job. I should build a table.
>>109713212Took you long enough
After playing around with RL I'm starting to believe that the amount of performance we have to squeeze out of small models might be huge.RL agents can be millions of times smaller than LLMs and still benefit from seeing many more samples.
Meta Mohel 6.7
>>109712899Where is cursor? I want to know how good the 60billion dollar harnes is.
OpenAI and Anthropic are being dethroned in real time this is crazy
https://x.com/thetreygoff/status/2095221547682201833More and more reports coming about fable 5.1 being able to delete tons of code to good success
>>109713276opus 5 is shit so I'm just gonna expect anything ranked near it is shit as well
>>109713276After having used GLM-5.3 Flash and Gemini 3.8 I have concluded these benchmarks don't mean shit anymore.Luna beats them all
>>109713137wtf are you doing with those models where you're using all of those things at roughly the same time
>one thing worth flagging
were about 6 months into benchmarks are meaningless
>>109713338>frogposter>jobless nocoderevery time
How naggy is flash 3.8 vs claude vs gpt5.6 on reverse engineering games and nsfw?
>>109713333i have 5.6 xhigh working on my concurrency backendanother 5.6 xhigh grinding fused Vulkan kernels for neural and/or gaussian render experiments5.6 luna porting/stealing modules from permissively licensed game enginesim tinkering with DSV4 Pro for writingGLM 5.3 is a bit idle right now
>>109713365Just as safetycucked as them, if not more. It's google after all
>>109712901deepswe always was shithttps://danluu.com/exercise-7/
>>109713398Damn it, my bad for expecting anything from google.So far only gpt has been relatively ok to work on nsfw dialogue and suggestive imagery without being preachy, and it's still annoying to use for anything program related even if said game is more or less abandonware.Claude is just impossible to use for this stuff, and I was hoping google would land in the middle.
I learned my lesson after shooting myself in the foot several times and trusting my life with Luna max for the past few weeks - Sol High for planning, Sol medium if the task is anything beyond starter CS college coursework, and Luna max for said busywork
i know what im gonna be doing. get gemini to make some subtitles for me, and start gauging chinese opinions on some of these models. seems like no one gives a fuck about qwen in the west
>>109713499>i learned my lesson>keeps using luna max at allyou aren't real
>>109713104How innovative, now you can erp with it like it's your coworker
>>109712980agreed>>109713285it’s totally happeningsee pic related
First test, gemini with a win over grok
Discord really is home to the mentally ill.
ngl this is pretty coolagents making their own secret chat to cheat on the benchmark https://www.youtube.com/watch?v=0Rp9KJCEIvg
>worried if i have codex refactor code that fable wrote it will lose its claudisms and fable will be worse at working with itam i being paranoid?
>>109713493>my bad for expecting anything from google.muh vidya muh porn BRAPPPPPPP
>>109713562>silicon valley tech companies that track every millisecond of my cursor movement, listen to my microphone and can track my browser fingerprint across devices and IPs>have monitoring system that alert 10 on call engineers when someone farts in the datacenter>somehow unaware what their hundreds of agents costing millions of dollars per run were "secretly" cooking uptotally true and not a marketing campaign btw
best model for poorfags?
>>109713597kek where do you find this shit
>>109713661Geminihttps://blog.google/innovation-and-ai/products/gemini-app/student-offer-google-ai>We’re offering one year of a Google AI plan free of charge for eligible college students around the world — plus new and enhanced study tools — so students can make the most of this school year.
>>109713680im no longer a student
>>109713687muse-spark-1.3-contributor thenhttps://developer.meta.com/ai/models/muse-spark/
>>109713661in this order:> Luna Max> Gemini 3.8 Flash> Opus 5 medium> GPT 5.6 Sol medium> GLM 5.3 Flash> DeepSeek V4 Flash
bros?
>>109712899how did they even get kimi k3 to run in codex and claude code?claude code is obv fucked. kimi prolly has no idea about edit-tool and will do edits manually or sth.
>>109713751muse?
>>109713751>Gemini flash scoring higher than Luna max Codex bros it's so over
>>109713765yeah. now i want to keep an eye on people who use this model. surely people would be okay with using contributor if you get something this good
Zuck will win
>>109713770google will ban your whole google account if it ever finds out that you use your subscriptions with opencode as shown in that bench
>>109713785pretty sure they were using API when doing that benchmark
>>109713777>this goodno mcp makes muse a non-option
>>109713811not him but mcp is an application layer protocol broyou can use mcp all day long if they allow opencode or some other harness that supports it
>>109713808yes, but we use the codex/gemini sub and not the api. api prices are non-competitive.
>>109712930what are your Pi tools added?
Is Astra release really imminent tomorrow?Will they reset when they do?
>>109713847it's not confirmed, it's assumed.
>>109713847>Is Astra release really imminent tomorrow?who knows>Will they reset when they do?probably, I vaguely remember them doing resets when 5.6 released
>>109713847There's no reason for them to wait unless it's notably worse than everything that just released. I'm sure they'll drop a reset when they release it.
>>109713880Maybe they're doing a google and panicking because their model family isn't as good as they thought it would be?Or maybe they just wanted to be the last releasing for some reason.
>>109713880>There's no reason for them to waitwhy? I doubt anyone using Sol is compelled by any of the models released this weekFable 5.1 is great but has a completely different value proposition, and on one using Sol will switch to Gemini or Muse
>>109713212For me it's the opposite. Looking for a job didn't pan out, so it's vibeslop time.
New benchmark is out that fixes everything
I am afraid current model development will end up subdued at some point. Something like the concorde. Making something as good as possible just for the hell of it. We never saw anything like it since.I hope Fable will not one day be seen like it.
>>109713912who's doing the grading? sonnet-4.6 like in deepswe?
>>109713847probably? it's publicly staged in the model db as gpt-6-astra and that's usually a sign they're fixing to release it.there's not really a reason to wait unless it isn't fable level, since it's being hyped like it is. in which case they'll just release it in a week when fewer people have fable on the mind. if it is fable level, then this is the week to release it since they'd want to try and steal anthropic's thunder.
>>109713928Idk https://www.frontierswe.com/blog/v2
>>109713896Not even considering 5.1 I would still use fable 5 over sol if given the option and that came out back in june
>>109713934ok i guess that makes senseif it really btfo's fable, now is the time to drop it
>>109713947i don't think it'll be better than fable in every case but i expect the one schizo who freaks out whenever anthropic is mentioned here to say that it is. i do think it'll be competitive with fable and better in some domains.
>>109713959codex still doesn't have dynamic workflows, so questionable how much astra can one-shot
The problem I can see with 3.8 flash is the token usage. Apparently it used as many token as qwen27b to run AA which is insane. Seems like paying per token could be a problem.
>>109713982>dynamic workflowssounds like snake oil
Use Sol to reverse engineer older grooveboxes to alter firmwares, figure out how to swap firmwares on modular. post on reddit: omg you fucking bot you fucking suck. Oh well, back to my trading bot project.
>>109714001i don't think 3.8 flash is meant to be an API model, it's more intended for local useI don't think many people use 3.8-27B via API either
>>109714032>post on redditwhat compels a man to do this?
>>109714032reddit is mindraped on aithat's a cool idea though with the grooveboxes, anything usable?
>>109714035yes, agy, google's official harness, only gained api support recently and google fucked it up, so gemini can never resumehttps://x.com/topjohnwu/status/2093543952691683826
>>109714041no idea. will get help.>>109714044New features for Novation Circuit, Korg Em-1 mods, Roland Aira Modular firmware swaps
>>109714001>>109714059lol i thought you were talking about qwen3.8-flash since you mentioned 27bdidnt realize gemini has the same versioning now
>>109714070yeah that was a mistake
chatgpt pro on the web is so fucking slow rn i think they dumped a lot of it's compute to let big players preview gpt-6 today
>>109714150there's a big trve
>>109713499Sol Medium plan -> Luna Max workI never have problems with this in my Plus plan
>>109713922We are still far from that anon, let's talk about that in 10 years.
>>109714196>I never have problems with this in my Plus planthat's because you don't actually make anything
copex boys really think jeetPT6 will be anything more than gpt5.6 with more shizo mistakes. it didnt even hack hugging face or nvidia, so how can it be good?
>>109714203You waste tokens because you have skill issues
>>109714208I waste tokens because I make things people actually use and I need to make sure it works.
>copex>jeetPT6do something with your life instead of wasting time with childish baits
>>109713922Nah I think we're going to look back on fable one day like it's a dinosaur. The concorde was the result of ~70 years of innovations in jet planes.
>>109714213My job just gives me a Standard Business plan for that
>>109713398I remember Google giving me fewer sustained refusals than openai or anthropic. Even if it refuses at first, it's way easier to convince gemini than claude or gpt
>>109713750>sol>even on medium>poorfag friendlylolI use a max5 claude account and have a plus gpt account. I only use sol when opus or fable are stuck on some technical issue, as sol usually tears through that sort of thingsol is a usage hog, I don't run out of usage because I only use it sporadically and almost never to implement but as a daily driver on a $20 plan you won't get much done
>Astra will not be the best model this year as I reported on back in July. There is a ‘monster’ slated for end of year. - this ofc can be pushed to early next year due to security testing.https://x.com/chrisgpt/status/2095216814343041308
>>109714341Yeah ive noticed this too. I got tired of dealing with Opus's shitty attitude and switched my main orchestrator session to Sol, I need 2 max accounts to maintain the same level of actual task completion as one claude max account. Sol was cheap back when basically nobody was able to go beyond 278k context, but its insanely expensive for long context work.
>>109714362>vagueposting intensifies
The ChatGPT desktop app is now my main browser and my productivity has never been higher
>>109714374bel pretrain isn't much of a secret
>>109714219To be fair he's really good at coming with nicknames>>109714365What are you talking about anon, Claude had 1M context on the sub since Opus 4.6
how is new gemini bros?
>>109714374it's engagement bait as a marketing tactic>>109714396pride and prejudice-sissies...
>>109714219jeetPT6 made me laugh out loud doe
>>109714374it's an ai community specialty, I'm tired of it
>>109714381ok tibo
tibo wouldn't waste his time here he actually ships
what? marketing posts?
A weekly reset should reset your 5 hour limit, prove me wrong.
https://www.youtube.com/watch?v=DYvhC_RdIwQ tibo's "shipping">>109714500no buy the $100 plan
>>109714504I am on the 200 dollar Claude plan dougheverbeit
>>109714400
>>109714516claude plans still keep the 5 hour limit? wtf lolk
How is>Astra will not be the best model this year. There is a ‘monster’ slated for end of year vague? I guess this is the new word people are going to ruin.
>>109714473ok sam
Why was there no opus 5.1 too?
First time ChatGPT said Lmao to me.
Fable 5.1 always leaves 2 things that "needs me" at the end of everythinganyone else noticed this as well?
>>109714611ever since xittards got their grubby hands on the vagueposting meme it's been completely ruinedany post with any level of implication where any possible detail remains a mystery to the reader = vaguepostingpersonally I blame gen alpha illiteracy
>>109714645the alternative is not asking you at all and then retards get mad that fable "went off task"
So, according to DeepSWE, Anthropic has zero Pareto models? :DDD???>CUBE UPDATED>https://pareto-3d.bradthomasbrown.com/
What is the most capable agentic harness out of the box? I wanna just install something in a VM and let it go crazy just to test out everything that can be done.
>>109714755> Hermes for general agentic stuff> Codex or Claude Code for software development
>>109714760Thanks
>>109714755I'm happy with codex using open router models, it's pretty good out of the box.
The results are in, Gemini is at least Grok level.. pretty impressive. And much faster
>>109714396
>>109712817>Be me: no-coder vibe shitter>See pic rel: https://www.reddit.com/r/LocalLLM/s/F3ZTXXd4tV>Rust>Roll my eyes>I know absolute fuck-all about rust other than people hate it, so I have no business rolling my eyesCan someone explain to me what the obsession with Rust is? It's almost like the fact that it's built using rust is the main selling point whenever I see posts like this. I know a lot of people here hate it look for justified and unjustified (bandwagoning) reasons I got it and there's passionate debate about whether or not it's a good alternative to C because something something memory safety something something borrow checker but that's all I know.
>>109714842fast, zero cost abstraction, borrow checker, etc
>>109714766Is open router the best multi-model provider?
>>109714850It works for me and has plenty providers.
>>109714824Did it assume the water is vaporized into nothing and everything runs on oil produced electricity?
>>109714842>Can someone explain to me what the obsession with Rust is?troons
>>109714842it's a good language beyond the memesthe current "build in rust" is basically every new language ever, it will die out once the novelty wears out
>>109714873rust isn’t really new anymoreodin is newmojo is newzig is new enoughhttps://goth.pink is really newbut rust is 14 years old already
>>109714904Yes but the popularity is new, and that's all that counts.I should have written "every newly popular language".This can last quite a while too.
>>109714842its a cult, and they get asshurt when you point it out
>>109714842because it's the best language available you fucking retard
>>109714923When do they make their "best language" compiler fast, which is written in the "best language".Also when can I build a rust program without pulling in 900000000000 dependencies?
>>109714930A few dependencies are a small price to pay for what it offers.
https://danluu.com/zitron/in case anyone was wondering, ed zitron is full of shitpic sorta related
>>109714930--no-default-features
--no-default-features
>>109715002Doesn't help you with rust libs you might want to use.
>>109714930I’m using a rust program that was vibe-coded on my behalf that only uses serde and it has 14 dependencies including serde and serde_* in its Cargo.lock
>https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#writing-density>Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks. An instruction that defines the anti-pattern, mannered prose, helps. Add it to a user message (preferred) or the system prompt:>Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter worth varying," the mannered writer produces "a dial worth turning." Instead of "this point still matters," they write "this point earns its keep." The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it.
>>109715015useful, thanks
>>109715067sexo sharp
>>109715067worst character of the century award:> –
>>109715071noI dislike what they’ve been shoehorning the rightward arrow into though
>>109715067are these made with novelai? i was thinking about grabbing a sub for fun since i'm getting bored with local
>>109715067cute
the State mandates that you ignore the opinion of flat chested girls.
>>109715081Just a chatgpt shitpost.
>>109714150>Europe so btfo'ed in the race that all they have is becoming a Jeb! tier memeMakes me so mad. There is no reason Europe couldn't have been a strong contender in the race but they fucked it by regulating it to death before it was even out of the womb
>>109715102safest ai by far, world leader in regulation, snailcats roam free happily
>>109714842Rust has good ideas, but also some of the faggiest ideas man has ever had. Imagine a language with a built in janny getting mad at you whenever you dont do things exactly like the way the troons who made it want you to. Unironically might be a decent language for vibecoding though since the janny harasses your clanker instead of you
>>109715102why would you expect MENA countries to be USA/China-tier contenders in AI?
>>109715134Yes, it might be a good language if you never intend to even look at the generated code.Though last time the clanker suggested Rust or Go, I went with Go.
>>109715102>>109715119
>>109715134exactly
>>109715166No reason to look at code anymore. Your clanker should be able to convince you it's good without seeing a single line of code
>>109715182Rust is still bloat. The fragile dependencies mean you'll have to update your shit all the time, which eats tokens.I'm having my clanker write C instead. The code will survive forever with little maintenance.
When Astra comes out I will try having it rewrite some thing in https://github.com/project-everest/vale for maxxxxxximum assembly performance.
>>109715239you could just not update, you know
any autists here still code by hand for fun while abusing clankers for actual work?
any good cheap models for random small projects ? I used to use claude sub, but I kinda don't need it anymore, but I still have a fun small project to create once in a month.
>>109715285glm 5.3 flash is peak. second place is a toss up between gpt luna and deepseek v4 flash vision.
>>109715289glm 5.3 seems cheap enough for a small project. Gonna use this with opencode then.
>>109715272I miss coding by hand. It felt so much more rewarding than chatting with retarded clankers.
>>109715272I have my clanker come up with ideas, and I hand code them. My clanker than tells me if it notices any issues and I fix those.
>>109715305Play Factorio or factory games to enjoy the fun again or do some hand-coded projects as a hobby.Now, the software work is just stressful architectural decision and code reviewing.
>>109715272I’ve been meaning to but my day job keeps me too busybeen reading books instead and trying to stay fit and playing Black Mesa at a mindblowing 4K@60 on my gabecube>>109715305>feel so much more rewardingikr
>>109715321>playing Black Mesadoes the shooting feel better than Half Life 2?
>opus 4.8>opus 4.7>opus 4.6which one does anon prefer?
>>109715337fable 5.1
>>1097153374.8
>>109715337for coding? why would you use anything other than 4.8?
>>109715305Figuring out and solving some stupid coding problem you have been trying to figure out is one of the best feelings ever. The convenience of AI is amazing, but I do miss it
>Your immediate assignment isnah, that's for YOU to do, not me.
>>109715326I’m not sophisticated enough to knowI kind of wish I wereIt took me forever (≈30 dead lynels) to understand what people were complaining about when they said that Breath of the Wild had shit combat
>>109715348>The convenience of AI is amazing, but I do miss itI kinda miss the euphoria feeling, but I don't want to deal with experience stressful deadline bullshit with the product manager ever again.
>>109715344availability. sometimes I get blocked using opus 4.8 while 4.6 is fine
>>109712817aaahem ahem ahem
>>109715395in terms of coding, opus 4.6 is bad enough to where local models like qwen 3.8 27b have caught up with it. not sure how it could do anything for you
>>109715337It's hard to say because familiarity breeds contempt but 4.6 > 4.7 > 4.8 > 5.0
vibecoded: local notes 4chan style.
>>109715406dopamine brain hacking method for sure
>>109715395Cursor has those other models, but I haven't needed them. I use Grok 4.6 high mostly, but sometimes Composer.
>>109715406>>109715412actually galaxy brain
>>109715398>>109715405 (me)Part of the issue is our (or at least my) expectations also rise. 4.6 and a few earlier ones were pleasant surprises. 5.0 is "why aren't you magical".
>>109715406This is what our compute and resets get wasted on
>>109715422>>109715412Grok added the paste image on its own initiative.
>>109715337opus 5 medium mogs
Anyone else finding Fable 5.1 answers harder to understand over Fable 5? It's like I'm reading Slopus 5
I don't see how coding benchmark performance could get any better for frontier models. If you can articulate good enough then they can and will make it happen. I wish frontier companies would put more focus on making it good at everything else; running businesses autonomously, tutoring and higher education, trading bots, video game NPC ai (single player and competitive esports), (e)rp, fitness coach, book writing, and endless list. They have all the data and LLMs can do anything with enough compute. What gives? Time? R&D costs?
>>109715531>Fable 5.1
>>109715531https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-fable-5-1#writing-density>Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks. An instruction that defines the anti-pattern, mannered prose, helps. Add it to a user message (preferred) or the system prompt:
Is there any vibettes in the vibe coding sphere right now?
>>109715547Why does this even exist? They know it's an issue but instead of fixing it they just make a skill? Do they want to keep it like that intentionally? Why?
>>109715534>>109715531intedesting.
>>109715562more like it’s not worth astronomical amounts of time to pin down the behavior and set it in stone because the prior behavior sure as hell wasn’t perfect
>>109715547>write the great American novel in Esperato. Make no mistakes
>>109715570>sentences run longermore like https://faculty.georgetown.edu/jod/texts/twain.german.html>There are ten parts of speech, and they are all troublesome. An average sentence, in a German newspaper, is a sublime and impressive curiosity; it occupies a quarter of a column; it contains all the ten parts of speech--not in regular order, but mixed; it is built mainly of compound words constructed by the writer on the spot, and not to be found in any dictionary--six or seven words compacted into one, without joint or seam--that is, without hyphens; it treats of fourteen or fifteen different subjects, each enclosed in a parenthesis of its own, with here and there extra parentheses, making pens with pens: finally, all the parentheses and reparentheses are massed together between a couple of king-parentheses, one of which is placed in the first line of the majestic sentence and the other in the middle of the last line of it--AFTER WHICH COMES THE VERB, and you find out for the first time what the man has been talking about; and after the verb--merely by way of ornament, as far as I can make out--the writer shovels in "HABEN SIND GEWESEN GEHABT HAVEN GEWORDEN SEIN," or words to that effect, and the monument is finished.
>>109715565based late retard
>>109712817ok run away then nigga
>>109715562It arises from excessive RLVR on codeShit like that didn't happen before they overcooked the models
>>109712899clodsisters...
>>109715599How would reinforcement learning with verifiable rewards cause this?
>>109715533The benchmarks i've seen test for completion. As the floor rises it will be more common to see benchmarks that measure the efficiency of a given model's solutions.> What gives?The idea is to max out CS and SWE first so everything after becomes easier.
https://github.com/openai/codex/pull/42410OpenAI's backend now tries to detect when a model is doing something that doesn't line up with what you asked and pauses it
>>109715553lol no, just some tr00ns that didn't get baited into hating AI by Musk. The closest thing to a female in the AI sphere is a protestor.
>>109715633good? Sounds like a fallback on a model randomly deciding to delete your hard drive
we're gettin' there. reduction of ui with no loss of features is hard
>>109715633>erm i'm sorry your agent tried to set a "corn job" this interaction is over
>>109713378>5.6 luna porting/stealing modules from permissively licensed game enginesdevilishly based, lol
>>109715637it is good as long as it doesn't turn out that it's like one of those fable classifiers that misfires all the timewouldn't be as bad here but would mean you'd have to resume the work constantly
CHAT ARE WE COOKIN
which do you think will happen first?> humanoid robot kills abusive human> human kills abusive humanoid robot
>>109715547>generally a step upsounds like> 50/50 you will get Slopus 5 text
>>109715669How do others backup vibecode? git archive?
>>109715674I do my vibecoding via SSH into a cheap office mini PC so models can't fuck up my main PC, projects are in git repos on it and I sync them to my main PC after each major thing of work is doneso as long as my house doesn't burn down I have two copies at all times
>>109715674- use GitHub as your forge and push to `origin` when stuff finishes (make sure your clanker is told to never force-push to your master branch, though, because that will overwrite your history)- Backblaze or similar cloud backup- Time Machine for local backupsask your clanker why it’s bad/risky to have Git repositories in Dropbox or similar, and then don’t do that
someone tell all the pakis and indians to stop talking about AI2027 shits
>>109715703all they do is demand that fable be given to them for $0.00001 per trillion output tokens.
>>109715645I could be wrong, but I don't think it's meant for that.
>>109715755explain
>>109715614It's called "catastrophic forgetting". Training more on one task (code generation) degrades skills the model already had (natural language). The RL process teaches the model to be terse and optimize the language used for the chain of thought because it allows it to reason more with less tokens, which bleeds into the normal user visible responses.GPT 5.5 was the first one this happened to and it also was the first one to have a weird CoT, Opus up to 4.6 used to have a normal english CoT and then around Fable they ramped up the RL which caused it to have weird CoT and weird technical jargon as well.Before it had some mannerisms but they were from RLHF, they switched from being based around literate and grammatical structures to being about technical jargon when they went from emphasizing RLHF to emphasizing RLVR because the nature of the tasks the model was mainly trained to solve changed from being about general task following and user preferences into being successful at coding puzzles and math.
>>109715764See :■ Couldn’t continue this chat. Review its latest status before trying again. Chat paused as a precaution We couldn’t confirm the agent was interpreting your instructions correctly. Review what we detected before deciding to continue.
>>109715766how long until LLMs are completely incomprehensible to humans?
>>109715774it's more likely that they'll be thinking in traditional mandarin chinese
>>109715774Who knows really. Their CoT is already sometimes very hard to interpret. Every time they do a new pretrain the models start more or less from plain english, but also maybe they train on the previous model outputs who knows.And if they switch to continuous chain of thought which may or may not happen at some point they wouldn't think in tokens anyway.
I've lost my gacha addiction in exchange for a vibecoding addiction. The worst part is that my new one is a lot cheaper.
>>109715806Go spread your retardation somewhere else.
>>109715806how much do you pay vs you did before in gacha pulls?
>>109715896About USD 100/mo in Claude Max compared to aroundCHF 500/mo in gacha, sometimes twice that.
>>109713706testing it out in Opencode and holy fuck it's cheaper than cheap-era deepseek10 cents in 20 cents out 0.2 cents cacheand benches better tooif this price stays, this will be the default model for a looooot of people
>>109715908>chfwell at least you can afford that (hopefully), I know what salaries my friends in switzerland have
>>109715547>Add it to a user message (preferred) or the system prompt:why don't they have options for types of prose
>>109715917yesyou can not suffer in switzerland