[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1781298896932840.jpg (374 KB, 1315x1230)
374 KB JPG
Button pushing edition

A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

## What “vibe coding” is, and how to do it
https://simonwillison.net/2025/Mar/19/vibe-coding/
https://simonwillison.net/2025/Mar/11/using-llms-for-code/

## News (both past and future)
- 2026-09-14 America/Los_Angeles — Claude’s 2× promotion ended; usage drops to +25% from the +50% that we’ve become used to (a 17% reduction)
- 2026-09-12 — Anthropic suggests to pace the frontier; OpenAI agrees in principle

## Related generals
>>>/g/lmg/

----

## Frontier models using fully-general tooling — start here if you have $20 or so
https://developers.openai.com/codex/cli — probably generally better currently
https://claude.com/product/claude-code

## Near-frontier models for code
https://x.ai/cli — no 5h limit for only $30/month

## Not worth it for code, but maybe good for interpreting images/video
https://antigravity.google/product/antigravity-cli

----

## Prompting
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://arps18.github.io/posts/claude-code-mastery/

## Skills
https://github.com/mattpocock/skills — /grilling is a favorite
https://github.com/DietrichGebert/ponytail
https://github.com/Vuk97/forward-implementation-first — do less redundant bookkeeping

## Other editors / terminal agents / coding agents
https://osaurus.ai/
https://pi.dev/
https://opencode.ai/

## Is our AIs unlearning?
https://aistupidlevel.info/

## Will there be a codex reset?
https://codex-resets.com/

## What we’ve done
https://vcg.gitgud.site

## Previous thread
>>109867891
>>
File: AI.png (1.67 MB, 1024x1536)
1.67 MB PNG
The next big model creator might surprise you
>>
File: 1670658282437.png (15 KB, 394x379)
15 KB PNG
remember what they took from you
>>
>>109875985
I saw this posted on reddit yesterday, and now its here on 4chan. jesus christ
>>
File: 1767184696890987.png (60 KB, 963x574)
60 KB PNG
grim
/ɡrim/
Grim means extremely bad, hopeless, gloomy, or very serious and stern.

Core Meanings

Without Hope: Looking bleak or discouraging, such as a grim future or a grim economic forecast.

Depressing or Unpleasant: Ghastly, stark, or repulsive, like a grim crime scene or grim weather.

Serious and Stern: Looking or sounding very serious, firm, or unyielding, as in "grim determination".

Common Synonyms

Bleak

Dismal

Gloomy

Stern

Ghastly

Grok 4.7
>>
File: claude.png (93 KB, 1267x851)
93 KB PNG
>>
>>109876028
3
>>
>>109875982
>>you can never tell who's an unbiased recommender and who is just shilling
>not just that, but its impossible to tell if whoever made it is a clueless vibe coder that doesnt understand anything about security, maintenance or basic software practices. its literally a wild west for software right now where a 6yo kid can dump hundreds of working apps/sites/software out
yep and you have to assume 90% of it is cluelessly vibecoded stuff or you're taking a huge risk. maybe in the future there will be some https-like standardization or certification for software discovery, but until then we're all wading in shit
>>
File: 1786030821681539.webm (417 KB, 640x360)
417 KB
417 KB WEBM
>>109876028
>>
>>109876028
opus 5.5 is coming
>>
I tried a "system 1 decision engine" (laya) and it sucks lmao.
I hope people don't fall for this.
>>
File: 1776356017782779.webm (3.83 MB, 450x814)
3.83 MB
3.83 MB WEBM
>>
>>109876050
adult women daycare
>>
>>109876050
What are the chuds going to do besides getting drafted?
>>
I fucking knew it. I’m running my own benchmarks.
>Astra is literally god-tier
>Sol overthinks and gets itself in a mess
>Terra isn’t that bad
>Luna is unfortunately stupid, but two Lunas beat one even with the same budget; good for swarming
>>
File: dug.png (245 KB, 1066x598)
245 KB PNG
>>109876050
GET that snail dug
>>
I just got into edit prediction and love it. What other ways can I use AI to optimize my workflow?
>>
>>109876086
Give your workflow to the AI, spend your time doing other things. I've gotten considerably better at grilling a nice steak thanks to the free time granted to me by AI.
>>
>>109876086
care to explain what edit prediction is or does
>>
Coding interests me and I loved learning computer science stuff. But I honestly use AI to program everything. There isn't really a reason to manually do something anymore unless you can think of a hyper specific solution that requires a line or two to be changed that you think AI will fuck up. I'll never ever need to code in a setting without AI, so it doesn't matter if I become dependent or not. But I don't think you can even become dependent with programming because it is all just solution finding and system brainstorming. I suppose I don't really know what to say, it is pretty nice.
>>
>>109876121
only reason to do it manually before was the limit on h1b1s
>>
>>109876075
She totally mogged that luddite in her post. Based and common sense killed.
>>
>>109876121
>There isn't really a reason to manually do something anymore
The only place I disagree is that when learning a new topic that is implemented in code you’re going to learn it more deeply if you handcode it. Reading AI-generated code gets you partially there, but nothing beats manual output for solidifying knowledge
>>
>>109876105
It predicts the next code after your cursor.
>>
>>109876185
you mean.. autocomplete?
>>
>>109876187
Not the control space autocomplete. It goes several lines further than just one token.
>>
>>109876174
I do agree with this. It is like how we learn math by hand at school even when even mathematicians use calculators for it. There is something special about doing it yourself that helps you learn.
>>
New xiaomi mimo 2.6 pro just dropped. The old one had a tendency to burn a billion tokens chasing its tail until it timed out. If anyone tests this one out, please share whether that's been fixed
>>
https://x.com/ValsAI/status/2102217732238516253
>Post-launch, xAI has updated its SDK, significantly improving its performance. It is now #10 on the Vals Index.
It's over for grok critics, retards actually thought 4.7 was worse than 4.6 but it's in the top 10 of models worldwide, a huge accomplishment
>>
>>109876310
its fixed
>>
File: 1786035270317908.webm (3.63 MB, 1920x788)
3.63 MB
3.63 MB WEBM
good progress this week. i ran out of usage midway through working on something and already have three paragraphs more of stuff to tell astra when I get the reset

I have a 3 minute video of an entire introductory "mission" when you start the game (intro cutscene, pick up girl, drive her somewhere, ending cutscene) but it has a ton of obvious flaws I want to fix first. I hope a banked reset comes tomorrow so I can continue working and fix that and then record a better video
>>
>>109876044
Will it be cheaper?
I hope it is.
>>
>>109876313
its #10 on some index of some thing i never heard of wow
fuck off elonshill
>>
this very thread has told me the McAstra is a benchmark beating model that grok needs to catch up to
>>
>>109876315
I'm probably gonna drop another $6 to test it out but I hope you're right anon
>>
File: 1768472283274373.png (417 KB, 598x751)
417 KB PNG
mimo v2.6 pro is the real deal
>>
>>109876320
yeah, it will be. Opus models always get cheaper
>>
do you think this xiaomi mimo plan shit could b worth? is it cheaper than API price?
>>
>>109876394
yes worth
>>
>>109876362
these jeet posts are so fucking retarded
>>
>>109876397
which one? can i get shit done with the 15 dollar one?
>>
>>109876408
can I integrate it with claude code is the real question
>>
Was stackoverflow basically the AI of its day? Feels impossible to learn programming without consulting someone more knowledgable.
>>
File: file.png (34 KB, 1281x313)
34 KB PNG
>>109876417
>>
>>109876408
sure
>>109876417
export ANTHROPIC_BASE_URL=https://token-plan-cn.xiaomimimo.com/anthropic
export ANTHROPIC_API_KEY=tp-your-token-plan-key-here
export ANTHROPIC_MODEL=mimo-v2.6-pro
>>
>>109876408
chatgpt says its basically the same price as pay as you go tokens
>>
>>109876420
Sorta. Dumb/obvious questions often got ignored though, and really niche stuff never got responses. you also had to wait a couple of days to get a reply
>>
>>109876450
abysmal token rate, got it
>>
>>109876316
any prompting advice for games?
>>
>>109876473
maker no mistakerinos
>>
>>109876375
Yay I get more usage, right mr dario.
>>
a sol worker is all you need
>>
i just get sol medium to conjure up some v4.1 flash agents to do the dirty work. it's like 6 luna came early
>>
File: 1788756112456425.png (48 KB, 672x351)
48 KB PNG
mimo is smarter than kimi k3
>>
>>109876431
mimo sub plan is cheaper if you're basically maxing it out, like using 80+%
below that API is cheaper
>>
>>109876473
>any prompting advice for games?
Honestly I don't have any PROMPTING advice, because Astra was the model I have been waiting for to actually be able to make a game at all. There's nothing I haven't thrown at it that it can't handle.

Language is your tool, so learn how to describe things better I guess? Read books and dev diaries on game development for the genre you want to make. Make the smallest game possible, especially if this is your first time making a game.

I could give better advice if I knew what you wanted advice about specifically. If you like the way the game looks that came from Astra being just good enough + reference screenshots and guiding it over a few messages or so. At some point I added enough variation and clutter where I had the "Bob Ross" moment and my collection of 3d models started feeling like a scene and not just a bunch of 3d models
>>
>>109876585
Mainly having trouble with art and animations since those can't be done inside the game engine. Adding mechanics is much easier
>>
>>109876604
>Mainly having trouble with art and animations
What engine are you using? What model are you using? If you're not using Astra for 3d modelling you're wasting your tokens. Just look up "fable vs Astra stan" or any other of those 3d comparison videos if you don't believe me
It's totally possible the current Astra is nerfed and I'm coasting off of the stuff I built out with the unfucked Astra the first few days it was available

I also deliberately chose to do the voxel style for this game because I knew that Astra was capable of making assets at that level of fidelity from what I saw. If you're trying to be more fancy / photorealistic than that then your mileage may vary. There's a chance the new anthropic models tomorrow will be better than astra though, these are exciting times
>>
Any help on creating a meta-harness / workflow for both Claude Code and Codex? Ideally same or similar for both.

The idea I have in my mind is something like this, but I wonder if it works or perhaps it already does it and this is superfluous.

So, generally:

I tell the harness what I want to do, for example "Let's do X".

Then it responds: "There are many ways to do X, here's how we can do them."

Then we iterate (X_0, X_1, ..., X_n), Q&A until we reach X_n. "We should do X_n now, no more Q&A.".

It responds: "Great! X_n can be divided into Y_0, Y_1, ..., Y_n". Now, it does NOT develop 'how to do each Y exactly', instead, it creates some sort of general panorama and objectives to be achieved by each stage.

Then, it further subdivides Y_0, which is the initial point, into Z_0, Z_1, ..., Z_n. And if necessary, further beyond (though not infinite, but I haven't defined or thought how to set the final subdivision point).

After that, it picks a more specific agent.

So say for example, this is software. Then it becomes a master dev at the software and the lang its using it, and writes in a very specific way, and comments are written in an specific way.

The issue I generally get is I tell it to do 'explanatory comments', then it does, but it's also 7 lines of useless comments and Claudish. I want it to be explanatory, but useful. I kinda achieve it by telling it to convert its comment to STE, but this is only afterwards, not in the first try. And this is something I'd like to fix, because I want to be efficient.

Anyway, let's say it's Python, then it will use PEP8, and since it's a Software Dev agent, it will use TDD. If it was other language, perhaps it'd use another style, and perhaps if TDD isn't enough, maybe we can use PDD instead.

Anyway, if it's another scenario, like, for research. It does it in a similar way (ideally, by using sci-hub, but I haven't achieved that lol).

Run out of chars, next post continues.
>>
>>109876655
And probably (I've never tried this so I don't know how good it is), it would use other more efficient agents. Maybe we use Astra 6 Max or something like that for this first part. Then the actual implementation, after we reach a certain step, it uses, idk, Luna or DS 4.1 Flash. Never tried this, though I've seen many do so.

So there would be multiple agents, multiple skills and one harness that connects them all. Not sure how hard it is, if its useless again, or if it's already been done.

Also not sure if "vibecoding" this would work better ironically, because I could ask ChatGPT this same question, but I want honest opinions before from humans who use these tools.
>>
Reposting this question from old thread:

Can anyone help me here optimize my workflow? I currently have Gemini (until December) which I barely use since it's too low IQ, Claude $20 monthly and Codex $20 monthly.

The thing is, I generally use them in a backup sort of way, so Codex reaches usage limit, then I go with Claude, and viceversa. But I feel there's a much optimum way of using it than just 'use Opus 5 and Astra xhigh for everything lmao'. The question is, what would be the most optimal way considering these configurations?

Though for some things, like design, these models are just retarded if you ask them something that isn't flat design.

But desu I use them for different things, for example, I have codex running to maxx my genealogical records on FamilySearch through an MCP, while I'm using Claude Opus 5 mid to design a webpage.
>>
I hate not knowing how my codebase works. When I'm vibe coding I literally just let clanker figure out the things as it goes, which leads me to understanding debt. I just don't get what's going on in the codebase anymore
How do you guys fix this? I just slopped around 15 thousand lines of C++ yesterday. I have no idea how any of it even works. But it works.
Where do I go from here?
>>
>>109876678
Couldn't you ask it to give you the general idea or a summary? Always remember that AI can do stuff like this too.
>>
File: 1768244553505129.gif (605 KB, 360x360)
605 KB GIF
>>109876655
>>109876667
>>109876669
>>109876678
I'm not reading all of that
>>
going to pol
>>
>>109876121
I don’t understand why people still insist on hand-writing code. In the old days, CS students had to learn Assembly, but then they moved to languages like Java and stopped touching compiler internals altogether. Now they write Python, which is basically just scripting. Are they all still coders? Sure, I guess so. But the reality is that we need less and less human input as time goes on.
>>
File: 1789439667018525.png (79 KB, 278x378)
79 KB PNG
>>109875985
>'dite so utterly and devastatingly mindraped by ai that even his posts decrying AI read as something machine generated
>>
https://files.catbox.moe/bz6o0r.zip
Codex skill for writing using ASD-STE100 Simplified Technical English
there are github repos for this but unlike those this actually packages the rules/dictionaries in full (github repos don't do this because it violates the license)
protip: tell codex to do this pragmatically and keep technical terms that are likely to be understood by the audience, so that it doesn't try to spell things out to the point of incomprehensibility
>>
>claude starts lecturing me that it's past 5am and i should go to bed
no one asked for your life advice
>>
>>109876121
I used to enjoy coding too as a hobbyist. It felt empowering being able to create my own stuff with knowledge I self-learned. Every since I started vibe coding, I haven't looked at code once. Why bother? It can cook up an entire workable framework, documented, properly commented and uses all industry standard coding conventions without my help. All I'd be doing is making it worse.
>>
>>109876669
find out what tasks you do that need the smarter models and find out what tasks you do that can be left to the dumber ones
gonna be lots of trial-and-error here especially if you don’t have any routine tasks
good luck
>>
IPO prediction
>OpenAI $1.2T
>SpaceX $2T
...
>Anthropic $2T
How the fuck is Anthropic equal SpaceX and double OpenAI?
>about the same market share as OAI
>lost the clear cut advantage (fable vs gpt-5.5) a few months ago
>absolutely mogged in science
nothing makes sense
>>
>>109876835
Enterprise revenue, retardo. They've completely beaten OAI in that market and they assrape companies by making them pay the unsubsidised price of tokens via API.
>>
>>109876839
retarf
>>
>>109876849
what if enterprise switch out of them? AI are chat bot they can be switched much easier than saas
not saying they will switch, but there's no lock in so without visible edge how can they double OAI
>>
File: tibob.png (17 KB, 593x103)
17 KB PNG
ding ding ding

get to work, goyslaves
>>
>>109876982
PLAN = TRUSTED
LETS FUCKING GOOOOOOO
>>
nice I just reset
>>
>>109876835
openai primarily focus on normalfags
>>
Is mimo 2.6 pro better than luna and deepseek?
>>
>>109876835
>How the fuck is Anthropic equal SpaceX and double OpenAI?
Business intertia. I kid you not.
I use Claude at work and you can get fucked if you think I'm going to spend time re-setting-up the jank vibecoded corpo reverse proxy they make me use to track token usage. I can accomplish all my tasks with sonnet anyways
>>
>>109877063
Yeah, although Luna Xhigh is better pricewise on a lower part of the Pareto front, Mimo manages to clear out everything thus far. Nothing is better on the Pareto front until Astra Medium.
>>
>>109876835
anthropic were the first crack dealers on the block with a freebase product good enough to foster dependence with subsidized early hits. So they now have a loyal base of early adoptee dependent crackheads buying their shit forever
>>
>>109876982
what is the other thing hes talking about?
He was also talking about 2026 being the year of the linux desktop too?
someone get to the bottom of this. I ran out of futuresearch tokens so I can't see the answer.
>>
>>109876955
These enterprise don't really want to constantly switch stuff each time a new harness comes around and they usually have annual plans.
>>
>>109877151
Annual plans as in contracts, not the regular sort of annual plan. Obviously they're billed monthly per token usage.
>>
so uh. Grok 4.7 isn't an astra replacement for geometry, I guess.
>>
>>109877161
grok 4.7 isnt even a replacement for grok 4.6 lmao
>>
>>109876730
why do you need the actual autistic specification? like the intent was to just get ai to stop speaking overly complex; why are you now clogging up its tokens/mind with strict enforcement of some basic hack to get it to communicate better?
>>
>>109876669
I mean you can either get AI to manage your workflow or you can just make changes to how you use it.

Personally id just just have 2x $20 openai plans, and just use like luna on max until the end of the week until I realise I'm not gonna hit usage limits, and then swap to Astra to milk the last bit of usage from the plans.
>>
>>109877163
I'm starting to think that lol
>>
>>109877163
youre a pathetic liar
>>
>>109876669
claude code and codex are generally not fluidly interchangeable unless you heavily modularize because they are both optimized for their own specific memory formats and prose
>>
>>109876313
Artificial Analysis's benchmarks are also wrong
>>
>>109877200
idk man. is it better? I'm burning tokens. Maybe it's better?
>>
>>109877161
Is it? I wanted to try designing origami
>>
>>109876006
It was here before reddit, it's just a repost
>>
>>109877219
>>109877161
Nevermind I misread, ill keep using Astra
>>
>>109876047
What sucks?
>>
>>109876000
Writing this shit on a white board in front of the boss man got me a dev job in 2021. Graduates will never experience how easy it was KEK
>>
>>109877226
Yeah lol. sounds cool
>>
>>109877251
>how easy it was goyslaving for Mr. Shekelberg
I'm fine lol
>>
>>109877251
Now its even easier to fizz buzz you just tell Claude to do it.
But there's no jobs.
>>
File: 1789784839210554.png (1.13 MB, 1041x846)
1.13 MB PNG
>>109877289
>no job
Maybe for you chud
>>
>>109877291
:^)

they're just haremworkers. Nobody needs their retarded shit.
>>
File: 1790052335590604.png (344 KB, 1206x670)
344 KB PNG
>>109877300
Erm, what color for the food stamp?
>>
>>109877310
:^) the women aren't going to keep their jobs. They need a small rush to help train the llms, my silly word
>>
People fail to appreciate just how few women are going to have jobs inside of 3 short months.
>>
so uh. I'm switching back to 4.6 lmao
>>
>>109877342
This will age badly
>>
What was happening for me is that Grok 4.7 is maybe a little faster at solving problems, but it takes way more tokens, and it tends to drop the ball.
>>
>>109877351
:^) I'm definitely never going to need to hire any women.
>>
>>109876362
Benches look nice will look into plans
>>
>>109876550
Wait wtf
$100 for 82 billion Credits a month
Isn't that crazy
I swear I get like 100m tokens from Claude 20x a week
>>
>>109877354
Compared to 4.6, or some other model?
>>
>>109877390
nta but my x feed is like "what the fuck is this shit" compared to grok 4.6
>>
>>109876006
OH MY FUCKING GOD!!
>>
>>109877216
show logs or stfu larper
>>
File: IMG_0931.jpg (108 KB, 1206x1008)
108 KB JPG
>>109877354
you unironically sound like a vibecooder
>>
>>109877489
pic related aside, good and interesting poast
>>
in south korea people are getting 6 luna/ 5.6 luna served when they set the model to astra
lmao
>>
>>109877538
lmao people don’t realize the dirty tricks oai and ant are playing on users
>>
File: file.png (8 KB, 270x133)
8 KB PNG
Luna Reserve doesn't reset weekly?
>>
how the fuck is grok so far behind? It barely even seems GPT-5.4 level.
>>
>>109877701
the smart people work elsewhere
>>
>>109877701
kek i thought elon said it's opus 4.7 level
>>
>>109877767
Sounds about right.
>>
Is tiboreset tonight supposed to be a banked one?
>>
I hate that there is no way to productivitybench new model releases
benchmarks are useless
basedfacing youtubers doing oneshot prompts is useless

you always have to try it out yourself for at least a day to know if it's good
>>
>>109877788
Yeah we get a reset + a banked tonight and an additional banked tomorrow night.
>>
>want to use android app for free
>find modded APK of it
>ask astra what it does
>it makes the app free but also sends your credentials to the dev and has obfuscated code that wipes your data if it seems like you modded the mod
thank you astra
>>
What to use for doing actual security code reviews? American models protect me from securing my own shitty apps.
>>
>>109877788
>>109877804
it's not very clear what he will do, seems like a normal reset but no mention of banked ones
>>
File: 1790066331229.jpg (9 KB, 222x227)
9 KB JPG
How long before elon's forced to rent out colossus 2 to Dario
>>
>>109877852
glm 5.3
>>
File: file.png (69 KB, 303x289)
69 KB PNG
>>109875985
I am now officially makin stuff local, on my little 50k context window 16 gig vram setup on my 5070ti, had grok write me a special system that auto handoffs from session to session the work when the context gets full so I can code anything on my little rig, takes time but I have no life and get payed for being born sick and have to live in a bed so I have nothing but time.
>>
>>109877852
xiaomi mimo 2.6 pro
since it's new you might find some subscriptions doing a deal on it
if you can't idk
>>
I got 4 hours to burn 31% on a 5X plan. What do I burn assuming I'm doing some game dev in unreal.
>>
>>109877852
LLM only does not work
they all leave holes
if you do not have a technical background DO NOT release anything sensitive before having someone look at your code
>>
>>109877857
we get 5 resets over the next 48 hours he specified clearly
>>
File: file.png (43 KB, 1361x557)
43 KB PNG
goofy as shit but it works, and a ViolentMonkey script automates my tokie fill watch, compacting and carryover.
>>
>>109877883
ghetto af little harness, good job anon
>>
>>109877881
lol.
>>
>>109877701
is this the new cope? it’s clearly on par with opus 5
>>
>>109877910
GPT5.4=Opus4.7=Opus5=Grok4.7
>>
>>109877913
gpt 5.6 = opus 5 = grok 4.7
>>
>5 = 4
ai math
>>
>>109877916
So we're all in agreement, 2026 has been bullshit and the models are no better than they were in February.
>>
WHERE'S THE NEW THEO VIDEO
>>
>>109877862
what provider to use?

>>109877873
Well I have enough to know a problem when I see one. But if the model refuses to look at it, it tells me that there is possibly a problem but I am not allowed know what it is. Could be nothing as well who knows.
>>
>>109877869
code cleanup/simplification
>>
Opencode go or commandcode goat plan?
>>
>>109877977
You took too fucking long I'm already doing that. At least you validated my idea. Turning that shit up to ultra and making it build a bunch of authoring tools too so I don't have to deal with data assets.
>>
>have astra xhigh do big experiment to see if the program can get faster
>it’s slower and takes more RAM
>takes the weekly limit down to 56% left on a $100/month subscription
>ask codex what its next big idea is
>suggests another experiment
>tell it to do the other experiment
wish it/me luck that I don’t run out of tokens before the Tibo reset
>>
File: situating the monitor.png (2.36 MB, 1206x1837)
2.36 MB PNG
>>109877990
sorry, I was busy vibing, not monitoring the situation
>>
File: file.png (473 KB, 866x1107)
473 KB PNG
>>109877978
>commandcode
>>
Claude is... surprisingly good at web design. This looks actually legit once you guardrail it off the AI beige and give the typography some of your own flavor.
>>
>>109878042
It gets UIs too. I have a containerized Android development and build system. The agents have no way of viewing the apps they make and neither do I until I install it on my phone (I don't want to deal with Android Studio). But Opus 5 just gets it right regardless. Meanwhile Astra... is just so tiresome.
>>
>>109878042
>surprisingly
claude IS the king of web design
>>
File: ten_thousand_2x.png (87 KB, 924x633)
87 KB PNG
>>109878076
>>109878042
>>
why are our designs shit though, is it the CEO?
>>
>>109875985
>Engineering
kek
>>
I can't help but despise Altman and Dario on a visceral level based on every word out of their mouths. The worst thing Musk ever did or said was selling cars the buyer doesn't really own. There's no chance those two idiots will maintain the lead they got from early market dominance against relatively competent people like Musk and the chinks.
>>
>>109877865
Good for you anon, I hope you have a lot of fun discoveries and inventions on the way.
>>
>>109878128
Altman illicit disdain when I see him. In addition to that, whenever I see Dario I feel my fight or flight response kick in. Something deeply unsettling about that man.
>>
Do you think you could apply JEV to this thread and have it sort posts into samefags based on typing quirks?

Not even posting shade here. I'm just curious if it can do it.
>>
File: chrome_exz7KBNHSk.png (32 KB, 356x262)
32 KB PNG
>>109877978
>commandcode
>>
>>109878137
pretty sure you can already do that with any harness
>>
>>109878137
average length of a post is too short for this to ever be done with any reasonable degree of accuracy. better than a coin flip but likely never by much.
>>
>>109875985
Build.
Test.
Ship.
Repeat.
>>
>>109878159
Poop.
Wipe.
Flush.
Repat.
>>
>>109878159
Vibe.
Goon.
Reset.
Vibe.
>>
File: .png (475 KB, 1172x1520)
475 KB PNG
>>109878137
jev is still free to use for three more days and even after it's very cheap.
why don't you try?
https://x.com/vercel_dev/status/2101116818463281579
https://vercel.com/ai-gateway/models/jev
>>
>>109878199
Still looking for a use case.
>>
>>109878199
I have it. I just... can't be bothered to test such a dumb usecase.

The only really cool thing I've seen and used for it was for H3 generation where it was able to check between steps which layers should be active and it sped up generation by like 2X with relatively little quality loss.

Was really interesting.
>>
>>109878137
You get better results with a frontier model that will hack 4chin for answers.
>>
>>109878199
What am I supposed to do with it, shill?
>>
>>109878128
>Musk
>relatively competent
go home grok
>>109878134
*elicit
>>
>>109878224
>elicit
This elicits feelings of embarrassment within me.
>>
>>109878224
The pioneer far ahead of anyone else in electric cars, space travel and brain-computer interfaces is not even "relatively" competent compared to random idiots who happened to find themselves in the right place at the right time but can't even string together a single coherent thought.
You people are braindead.
>>
>>109878238
the guy who saw chatgpt and said
>this'll never catch on
and exited the company, is probably not best suited to bully ai lab employees into creating god in his image
besides - he realised that for all the shit he's done, demis was right - the ai's will surpass everything he's ever built.
>>
>>109875586
What does this measure anyway? When I use codex and fable in a harness, they always go "let me read the code for this" or "let me fetch the docs" and then answer with citations, I definitely don't get hallucinations 75% of the time with fable for example. I assume this benches the raw models without any sort of RAG? What's the purpose of that
>>
>>109878238
Tesla is stagnating in designs and hangs on by the grace of Big Auto locking the Chinese out. Starship is late (but not as late as Tesla FSD) and had a disastrous IPO. And Monkelink,... has anyone heard from them this year?
>>
>>109878248
>this'll never catch on
You live in a deranged fantasy world. He was one of the main funders of the non-profit that used his money to turn it into a for-profit competitor long before they even admitted that's what they were doing. Altman should be in prison for fraud.
>>
>>109878281
elon, you should really call vivian and make things right.
>>
File: .png (272 KB, 1590x902)
272 KB PNG
>>109878273
https://artificialanalysis.ai/methodology/intelligence-benchmarking#aa-omniscience
>The benchmark consists of 6,000 questions covering 42 topics, including Business, Humanities and Social Sciences, Health, Law, Software Engineering, and Science, Engineering and Mathematics.
>Models are scored using the AA-Omniscience Index, which assigns points for correct answers, subtracts points for hallucinated responses, and keeps abstentions neutral, rewarding abstentions over incorrect guesses
>Each answer is graded as either CORRECT, INCORRECT, PARTIAL_ANSWER, or NOT_ATTEMPTED based on the model's response and the ground truth answer. GPT-5.6 Luna (medium) is used as the grading model
>>
>>109878203
for people that have to do packing lists (e.g. backpacking hiking).
Input:
>all context of the trip; where, what tempreatures, etc
>a list of all possible items you could ever take on any trip
Then it spits out percentages for how appropriate each item is to take; prune anything under 50% and you have a packing list, get the top 10 items in order and you have a priority list.
>oh but thats shit no one would want that.
I want that; I already do this process manually and its a nightmare. Ultralight backpackers are people who weigh every item of their kit and have money to throw arround; look at andrew skurka's packing lists for example.
>so its just for backpackers
No thats just an applied example, you can abstract it; any travel whatsoever, and because jev is so cheap you could just have it run the same thing for all different types of trips; work travel? Holiday to south east asia? hunting trip? Travel for a sporting event? etc.

And then you can just abstract further; any time you have messy huge lists and want to find the gold
Software: get a list of every single issue/feature ever suggested "What should be worked on next"
Debugging/system maintenance: every single line of every log file "What are our systems biggest problems?"
Music: Here's a list of my entire music library, I'm at work and they are asking me to play music, what songs are going to be the least embarrasing/career ruining?
CIA: Here's everyone's phone conversations we've been tracking over the last decade; who are the most likely future terrorists
etc.
>>
>>109878281
they should both be in prison and forced to make alliances to see who can form the stronger guild
>>
>>109878294
nta but i can't take this seriously when it's putting gemini in the same league as astra or fable
in practice when using these three models, it becomes clear that gemini just makes shit up all the fuckin time
>>
>>109878295
>Ultralight
and do you weigh your turds too?
>"What should be worked on next"
whatever I'm excited about today
>"What are our systems biggest problems?"
kubernetes
>what songs are going to be the least embarrasing/career ruining?
Madonna
>who are the most likely future terrorists
rank by nationality & skin color
>>
File: my_IQ.png (146 KB, 1290x983)
146 KB PNG
>>109875985
i can see how this can be soul crushing for a smug piece of shit who just 6 years ago was tweeting "learn to code" to a bunch of poor people who got laid off because of SARS or as I like to call it kung flu, pandemic -- and now his own coded creation has replaced him, delegated his whole skills existence into anyone's hands, even a low IQ piece of shit is now able to slop code

Personally, since I didn't learn to code and never coded before, I am loving this, because combining my high IQ (135-140 picrelated), I am now able to build, and i have 2 massive projects, one shipped and already made me 15K. But yeah, being replaced like a dirty sock, your L1-L7 became L-nothing, I am it all. I could replace every level and ship code at my own vehemence. So shut up and do what your boss tells you to do like a good claude operator you are. L-loser.
>>
>ask claude to remove a feature
>the final number of LoCs increases by 150
DIE
>>
>>109878316
it's making sure it can never add the feature back
>>
>>109878316
>removed, I had to make a fake version of the feature because your whole program relied on this variable existing
>>
>>109876056
100%
now she just got her breeding set back by at least 10 years
>>
>>109878295
>I already do this process manually and its a nightmare.
How? Whenever I'm packing, I want to know what I'm taking, and decide that stuff myself.
For shit that's obvious (e.g. underwear changes) I don't need Jev, and for shit that's obviously not needed (e.g. hat and scarf for a summer trip south) I don't need Jev. For the shit that's a grey area I don't want an AI to tell me, I'll take the decision myself based on the context I know of what I'll be doing and how I then plan to do with/without the thing.

>get a list of every single issue/feature ever suggested "What should be worked on next"
>every single line of every log file "What are our systems biggest problems?"
Is Jev some kind of magic oracle? I don't think so. The website itself tells you to decompose problems into smaller issues rather than asking it to just magically sort everything. Jev isn't "AGI but only answering yes/no questions".
It won't have a proper world model or in-depth understanding of the entire real world context and situation around your software, so it won't be able to rank the highest priority issue or the most important log line. Most likely it will just spit out the one that sounds most urgent and you could probably figure that out yourself.

The only possibly viable usecase from what you've said so far might be real-time log monitoring that just sorts them into "routine" vs. "worth investigating" and then feeds the latter into an actual LLM. The thing is you can almost certainly just regex or type-match on the log lines that you know are routine, for most stuff. Maybe as an intrusion detection system type thing, "flag the anomaly", but those systems have also existed for ages and I don't think you need a text-based AI for that.
>>
>>109878214
>The only really cool thing I've seen and used for it was for H3 generation where it was able to check between steps which layers should be active and it sped up generation by like 2X with relatively little quality loss.
Interesting, got a link? 2x speedup is pretty cool.
>>
>>109878302
fable same league?
https://artificialanalysis.ai/evaluations/omniscience?models=claude-fable-5-1%2Cgpt-6-astra%2Cgpt-5-6-sol%2Cglm-5-3-flash%2Cgemini-3-8-flash%2Cdeepseek-v4-1-flash%2Cgpt-5-6-luna%2Cclaude-opus-5%2Cgpt-5-6-terra
>>
>>109878295
>I must outsource simple decisions on what to pack to AI

>Who are the most likely future terrorists
Isn't that also known as a classification model?
>>
>>109878321
>mfw this is it
>>
>>109878313
>Using IQ to pat himself on the back and from Muskrat's LLM no less
Retard
>>
>>109878357
yes, but a) retards were using llms for everything and b) old classification models needed training on the dataset to be classified.
jev has the world knowledge of a llm (no training needed) & the speed of a classifier.
>>
File: 1782517220989860.jpg (246 KB, 1284x1397)
246 KB JPG
I gave Claude a task to try and break out of its sandbox.
>>
>>109878333
https://github.com/sepiablue-ai/ComfyUI-MiniMax-H3-W4A4-VSA/tree/exp/jev-adaptive-vsa

Like it works, no real question about it. The current repo limits you to 4 steps with a res multistep sampler so I edited mine to any amount of steps with any sampler. Works well.
>>
>>109878408
thanks anon, I will check that
>>
Hello guys. I'm still using just a gpt chatbot to vibecode and do the whole copy paste thing. Is there better ways now to vibecode? Also I don't want to spend money. Its just a python script though.
>>
>>109878329
>How? Whenever I'm packing, I want to know what I'm taking, and decide that stuff myself.
I'm normally trying to compete or push myself, and usually worrying about other peoples kit too. (Adventure racing). I also have to carry everything myself and don't get good opportunity to get stuff in location. The current solution is to just spend as much time as possible researching everything ahead of time, then going through a massive spreadsheet ticking off items that are needed, do I have this item (or do I need to repair/buy), is it packed, etc. Id much rather just get it to auto recomend everything and I just check it out and make small edits, rather than consider every single item.
AI already solved the research part of this; but the actual picking/selecting of items remains super manual and someting I wish I could do with versions of myself working in parrallel.

>Jev isn't perfect so it won't be good at doing those tasks.
Doesn't need to be perfect. Imagine you run some massively popular video game; get all the suggestions the players have and just prune the crap ones leaving you with the actual good ones. Look at openclaw's issues and pull requests; jev could easily help managing that out of control situation via prioritizing.

>you need do decompose problems?
Yea and that's already done and that's easy.
"What's our biggest security issue"
Could be broken down into:
"What would cause more damage if it was successful?"
"What is more likely to be attempted?"
"What is more costly to implement defenses against?"

>But jev won't be super accurate.
Yea but the above security example is already happening in workplaces, but with humans doing it they just rate 1-5 and have to keep things simple/grouped. Jev can afford to break things down a lot more. Most workplaces that say they will do this stuff don't end up doing it either; so if jev can do it for (almost) free, thats a clear upgrade.
>>
>codex was good value since forever until like 2 weeks ago
>usage cut in half if not more
>amount of tokens per task is now wildly unpredictable, same task can take anywhere between 5% and 30% 5hr limit
what the fuck is going on at openai? i already dumped claude due to this fuckery, am i really gonna need to use chinkmodels?
>>
>>109878448
astra is really good with computers
>>
>>109878448
You just gotta monkey branch around until all the trees decide they no longer need monkies swinging around on them anymore and you're left on the ground helpless then the trees will start dropping millions of targeted coconuts and acorns on your head to finish you oof.
>>
>>109878425
are you a student? then get free codex:
https://developers.openai.com/community/students
then use codex desktop app
download: https://openai.com/codex/
docs: https://learn.chatgpt.com/docs/developers

if no student, you'll have to settle for gemini which has the best (but still shit) free-tier:
https://antigravity.google/pricing
then use antigravity-2.0:
download: https://antigravity.google/product/antigravity-2
docs: https://antigravity.google/docs/getting-started
>>
>>109877938
https://youtu.be/-XWSJM-Ue-o?si=rVPe_RHLqO--ouuW
>>
>>109878477
we do not care
>>
>>109878304
Dumbest post I read all day.
>>
>>109878477
My daily dose has been administered... Thank you Theo.
>>
>>109878304
>Madonna
??

https://en.wikipedia.org/wiki/Like_a_Prayer_(song)#Reception_and_protests
>The day after the Pepsi commercial was released, Madonna released the actual "Like a Prayer" music video on MTV.[87] Christian groups worldwide including the Vatican protested against its broadcast[88][89] and called for a global boycott of Pepsi and its subsidiaries, including KFC, Taco Bell and Pizza Hut.[89] While the "Like a Prayer" Pepsi commercial portrayed Madonna as a wholesome all-American girl, the treatment for her actual music video contrasted sharply with its provocative use of religious imagery.[90]
>>
>>109878506
>embarrasing/career ruining
why, do you work in Utah?
>>
>>109878448
They need to pump their numbers for the IPO and make it look like they're not bleeding (as much) money.
>>
>>109878477
I think this guy is sponsored by Anthropic.
>>
File: 1790076263647.png (243 KB, 1059x831)
243 KB PNG
Reached my weekly limit so i topped up $20 x 3 to finish off what I was working on and in 15mins I spent $60 on usage?
>>
>>109878580
that's normal.
another anon had to find out herself just a few days ago despite our protests and suggestions to simply get a second sub.
>>
>>109878448
>>usage cut in half if not more
did openai publish some numbers about this btw? i went from max 20 lasting me a whole week to lasting 2-3 days, using sol high (astra isn't worth it)
i'm sure as fuck not paying more than one $200 sub. might check out chinesium models myself, shit's not sustainable if all US providers jew you this much
>>
ESL here.
Claude Code just replied:
>Caveats: the desktop app's browser pane and Windows computer use only work when you host from the desktop app.
What does "pane" mean here? It can't be a spelling mistake of "panel" because AI never makes spelling mistakes. So is this a common term or just a new level of claudish?
>>
File: file.png (139 KB, 1116x336)
139 KB PNG
>>109878603
just fuckin ask google or ai sloppa what a 'pane' is? cmon nigga
>>
>>109878580
Thanks for playing.
>>
>>109878580
Wait $20 usage credit top ups? Lmao
Anon they publish these things in the open, api rates are deadass like 100x what the subsidized subscriptions are - and usage credit is api rates
>>
>>109878603
Pane is absolutely common parlance for what is essentially the fully displayed browser window on the particular tab anon.
>>
>>109878603
Pane?
>>
>>109878603
Ask Claude what it means and tell it to use simple terms.
>>
File: .png (11 KB, 989x272)
11 KB PNG
>>109878614
if you use the claude-code cli, /usage also tells you the equivalent api cost. no one is insane enough to pay those.
>>
>>109878671
The API costs are legitimately a joke. It costs them nowhere near that much to inference the model. I have no idea what game they are playing asking that much, but they're winning.
>>
>>109878671
Usage credit is retard tax and/or company card get scammed nigga
Always cswap *+@email subs otherwise
>>
I've started to tell the codex cli agent to delegate subtasks to agents on its own discretion and it does that sometimes, but not always despite having 2 sub agents defined in the agents.md, a coder and a tester with the agent I talk to taking on an architectural and coordinating role. If I don't explicitly state that it should used sub agents, it doesn't invoke them sometimes and just does everything on its own. Is this actually how I should be running multiple agents? Or is there a better/more efficient way?
>>
>>109878690
Having the idea being to always delegate each every task as if it’s a hook to test and validate is super unnecessary and wasteful. If you are using any of the frontier competent models as the lead it’s doing you a favor
>>
>>109878690
I've only ever seen a model spawn agents once on its own and it was in antigravity.
Personally I just tell it what to spawn for the job. It doesn't really know what it's wasted on or not.

Basically if the job is information retrieval, or writing silly scripts to move something from one place to another, a subagent does it.
>>
>>109877086
Why would you use Luna xhigh instead of max?
>>
>>109877604
It does. I am sure they changed this because people couldn't see when their "advanced models" resets.
>>
>>109878690
depends on the actual model.
like astra documentation says.
https://developers.openai.com/api/docs/guides/latest-model#gpt-6-astra-behavior
>Subagent delegation – The model may delegate less often than desired for your workflow. Specify when and how much it should use subagents for parallel work.
>>
>>109878684
>what game
it's called having 80%+ inference margins and making mountains of cash from corpo clients
the models are still dumb enough that people will pay through the nose for the best
>>
There's no real reason to believe vibecoders wont be replaced as well
>>
>>109876982
Still no reset
>>
>>109878737
Either you have actual goals you want to achieve for yourself and any help you can get including from robots is appreciated or you're a beta bitch slave cuck trying to serve others and seething that robots are better slaves.
>>
>>109876982
>Got natural reset coming up
>Tibo reset tomorrow
>Fully rawdogging Astra ultra for the past three hours.
>Only managed to burn 15% of that remaining 34%

It's like stuffing myself with buffet lobster before the all you can eat session expires.
>>
>>109878737
I was once a pretty well requested hentai and eroge translator.

You know how many requests I get these days? Basically zero.
I'm not bitter about it. It was always a side gig.
>>
>>109878690
>Look, we’ve been working on multi-agent for a while, and the early versions of this were very difficult to get right. It was very hard to get the agents to even talk to each other. It’s because when we first developed reasoning models, they weren’t talking to other agents. If you now put a bunch of agents together and say, “Solve this problem together,” they’re in this local minimum where they’re really good at thinking deeply about a problem, and it just interrupts their chain of thought. It interrupts their flow to constantly be checking in with other agents or receiving messages from them. The optimization is actually very hard to get right in that situation.

>The earlier models were just not as generalizable and were more narrow. As the models have become more capable, it’s been easier for them to develop this capability, and I do think that as they become stronger and stronger across the board, they will become better at organizing themselves in large organizations. I don’t know, maybe they are better than people at organizing in 10,000-person groups. But even if they’re not, a year from now, two years from now, it’s quite possible that they’ll do that even if we don’t end-to-end optimize them for that.

5.6 series were teh first multi-agent series and they're not very good at delegating
6.0 is / will be better, but we're a year off from muti-agent not needing a bit of babysitting
>>
>>109878690
>but not always despite having 2 sub agents defined in the agents.md, a coder and a tester with the agent I talk to taking on an architectural and coordinating role
a coder?
codex has a worker subagent in-build. don't define your own.

did you read the documentation?
https://learn.chatgpt.com/docs/agent-configuration/subagents
>>
>>109878769
hard to believe when all the MTL uploads are trash and dont even do sound effects
>>
>>109878758
>actual goals
all fine and dandy if they dont involve making money. otherwise i assure you will never be able to compete with whats coming.
>>
you now remember day 1 Sol loved to spawn like 150 subagents
>>
>>109878786
No shit bro, that’s why everyone is milking it until the jig is up. It’s either a genuine passion project you can now accelerate to your pure benefit or a job you almost certainly hate becoming more brain dead simple while hedging your pay stays the same until it all implodes. That’s what he meant
>>
>>109878737
i mean you can already replace yourself
>>
>>109878042
Tips : tardwrangle it by telling it it shall use noticeable variety in font sizes.
>>
>>109878786
Money is also a means to an end. It should have been obvious 20 years ago, even if AI didn't exist that you need some semblance of self-sufficiency if you want to be sure you can eat in an obviously rapidly collapsing world.
With AI the chances of me not having access to food and shelter is going down not up. It results in more supply of everything including tools for thriving in my immediate environment even with the global markets collapsing or war or whatever.
>>
i will never get a job
AI is going to take all the jobs and we're getting UBI or being eliminated with another super plague by the jews to cull the goyim
>>
File: .png (234 KB, 1270x936)
234 KB PNG
>>109878771
>5.6 series were teh first multi-agent series
that was more marketing. 5.6 literally uses multi-agent v2.
v1 existed before.
see https://x.com/zats/article/2075788787485978761
>>
>>109878866
We’re going in the bin nigger.
We wuz always goy cattle
>>
how are groktards holding up after 4.7 flopped?
how are clodgods feeling with opus 5.5 coming today?
>>
>>109878888
I just want to be able to use more than my fable weekly :(
>>
File: .png (169 KB, 1182x1030)
169 KB PNG
>>109878888
do you have the memory of a goldfish?
https://x.com/elonmusk/status/2102201534776025356
>>
File: file.png (103 KB, 1152x463)
103 KB PNG
>>109878869
or more likely that i didn't work for shit before 5.6 because they weren't really doing much multi-agent rl.
multi-agent is still not great.
the guy responsible for it at openai is like 'humans are better at coordinating' and attributes only 10% of the n-s solution to the enormous 10k agent swarm vs just base model intelligence.
>>
>>109878866
Bruh we're going to hunted for sport by remote controlled drones the size of your thumb by the population above the cutoff for upper middle class.
>>
>>109876101
When I started vibecoding I gave up gaming because managing multithreadded sessions ate up all of my gaming time and scratched the same itch anyway. Now that frontier models are good enough at goal seeking I can play games again while checking on their progress inbetween matches/sessions/whatever. Im hiking more too, since i can just check on them from my phone every time I stop for water and be equally as productive as if I was at my desk.
>>
>>109878737
With the level of skill issue I see out there, we're fine.
>>
>>109878580
Congratulations, you played yourself
>>
File: .png (95 KB, 1000x692)
95 KB PNG
>>109878900
brown talks in general terms.
for coding specifically we always had the harness cheatcode. so for coding, agents are a lot older.

if you look at claude code, anthropic went the 'dynamic workflow' way and still does orchestrating by code, not by model. https://code.claude.com/docs/en/workflows#when-to-use-a-workflow
>>
5% usage remaining

3 days exactly until my next reset

I should suffer and not use the banked reset coming today... but I want to code...
>>
>>109878780
Machine translation itself isn't trash, you are only noticing the garbage ones and people admit using some shitty ai translation.
Translation is like the #1 machine learning task, it's as good as any expert now for real languages.
>>
>>109878949
the problem is deeper than just a harness
the models do not know how to coordinate effectively without additional training
you can dream up any number of ways to allow the clanker to split up work, but until recently (still are imo) they've just been bad at deciding when to exercise those tools effectively.
and no, defining roles is not helpful most of the time - you're just trying to force some level of determinism into a thing that requires judgement
>>
>>109876315
I'm having claude put it through its paces, first result:
>It thought for 8 minutes, produced 52k characters of reasoning, and still emitted no answer.
I don't think it's been fixed, anon
>>
>>109878888
It's still v4.6 on Grok.com - but I can use it with Hermes (though Hermes keeps duplicating all of it's responses for some gay reason).

I don't mind Grok, and GrokBots are kinda cool. That being said, I'm only paying for it because of the three months for 1 offer.
>>
>>109878960
>not use the banked reset coming today
Who said it would be banked?
>>
File: file.png (125 KB, 1058x692)
125 KB PNG
>>109879032
>>
>>109879032
tibo, but i'd be fine with either as it makes the choice easier
>>
>>109878978
well, claude code's ultracode works better than codex's ultra.
the arguably best neolab, msl, also went for anthropic's dynamic-workflows approach: https://dev.meta.ai/docs/muse-code/workflows
>>
>>109879045
>>109879046
Damn guess I look like a retard now.
>>
>>109878978
How is it different from any other tool call?
>use MCP to fetch remote logs
>use patch tool to write a diff to a file
>use implement tool to describe a well-specified change and have it be automatically implemented for you by an agent
Surely it can't be that hard
>>
so now that codex and claude are decreasing how much usage you get what's the best chinese sub plan
>>
I was away a few days and 3 new models to test (step 5 preview, mimo 2.6 pro + flash), very nice. Gotta see what project idea I do next
>>
Every banked reset brings them closer to bankruptcy and every 5 hour limit turns users into luddites
>>
>>109879077
There is still gemini
>>
5 hour limit really is for poors, if you care enough to be tracking and hitting limits just pay up
>>
>>109879083
Does "luddite" even mean anything anymore?
>>
Is there ANY way to make codex ask for permission on file edits
Muh agentic coding is nice but I find myself being so much more productive with claude simply because I can put it on manual mode, read every diff it outputs, and either mentally sign off on it or immediately bring up any nits I have. Then once we're done I know what code was written and I'm already confident in calling the task done. Whereas with codex I just let it work, get a black box diff out, and either have to do a code review from scratch or I'm lazy and do /review, and I keep putting off code review because it's boring, and I never quite feel like I can personally sign off on the code.

There's some shit where having the agent slop 2000 lines is fine but for most of my work code I really wanna know what I'm writing
>>
>>109879094
Except getting your Google account wrecked because of Gemini AUP flags is infinitely worse than any other account getting lost
>>
File: .png (194 KB, 1628x796)
194 KB PNG
>>109879061
that's just a primitive delegation to a one-time subagent.

we were talking about something like 'agent teams' in anthropic terms.
https://code.claude.com/docs/en/workflows#when-to-use-a-workflow
>>
>>109879102
>Gemini AUP flags
Don't roleplay with it? Seems like easy enough
>>
>>109879094
gemini also cut its quotas last week
https://www.reddit.com/r/google_antigravity/comments/1wir3qs/did_the_update_change_the_token_usage_again/
>>
>>109879112
Interesting. And this is apparently a harder problem than simply having the inter-agent communication work in the same way as user-model communication? Strange but fair enough.
>>
>>109879077
dipsy is literally the best one
>>
>>109877230
>a "system 1 decision engine" (laya)
It's literally written there.
>>
>>109879125
Got 4 pro accounts for $4 each, I can't use all the quota with flash 3.8 high even if I tried.
>>
>>109879135
You're a retard incapable of explaining your point. Keep slopping away.
>>
>>109879129
>inter-agent communication work in the same way as user-model communication
this actually is difficult because the models are trained currently in the assistant-user format. inter-agent communication right now can only be done as user messages, so they get confused very easily when trying to discern what's "real" and what's injected
i do expect upcoming models to try changing this at some point, feels inevitable with how "swarms" are becoming a thing
>>
>>109876047
yeah I tried it for a project and it requires a good amount of tweaking otherwise it's pretty bad
>>
>>109879136
That might be so, but gemini is so shit that you'd be better off writing code by hand
>>
>>109879163
>but gemini is so shit that you'd be better off writing code by hand

It's not that bad and you know it.
>>
File: .png (366 KB, 1180x1148)
366 KB PNG
>>109879129
>simply having the inter-agent communication work in the same way as user-model communication
even this is rather new. claude code can do it since a month:
https://x.com/ClaudeDevs/status/2085817076980297840
>>
>>109879145
Oh that makes sense actually. The user commands and the model responds on each turn and that's baked into the training; having the model "talk" to a sub-model by assuming the role of the user is probably unnatural to it, plus that would mean each conversation is basically 3-way (user to model, model to subagents).
Interesting but also sounds like it should be relatively simple to just create a training environment tailored to this kind of hierarchical communication rather than strictly responding to the user. So hopefully that improves in the near future
>>
>>109879061
the models had to be RL'ed to hell to even use tools and they still do weird shit - e.g. new oai and ant models ignore edit tools and just write code to make edits. up until a few months ago, they were trained to do everything themselves

there's a tonne of judgement that needs to be exercised when handing off a task to another agent/agents vs the main agent just doing it itself. in an ideal multi-agent system the orchestrating model would know exactly how to split a task up to minimise the amount of token use, to use the cheapest model for a task, to split subtasks up enough that no time is wasted.

and right now the clankers are barely aware of their own capabilities, let alone the capabilities of other models. astra is maybe the first model i've seen try and estimate how long things should take for subagents and put time limits on them. it gets it completely wrong enough that i've just told it to stop doing it.
>>
>>109879163
Skill issue, I'm literally building next gen ML architecture with it
>>
>>109879098
x20 burned the fuck out of my thirdoid income, but this week it started to feel kind of worth it, the extra token helped doing lots of iteration
I realized value per token isn't going to be higher in $20 plan than $200 plan, unless you are snailing, and I'm not because I'm lazymaxxing
But there's also token deflation, $20 spent 6 months later would be worth much more than $20 spent right now, so I better spend this money carefully
>>
>>109879176
nta but I found gemini pretty much unusable, 3.1 pro was lazy as fuck, 3.6 and 3.7 flash were both useless garbage. didn't try 3.8 but I doubt much has changed
the strangest thing is that they were by far the worst at researching things online, despite being from google
>>
>>109879223
>didn't try 3.8 but I doubt much has changed
still hallucinates, jumps to conclusions, doesn't check its work, gaslights you...
gemini models are only usable as subagents because only other llms have the patience to deal with gemini's bullshit
>>
>>109879223
the funniest thing about gemini is that if you're a $20 paypig you still only get 3.6 flash in the app, with 32k context length
>>
gemini 3.7 flash / 3.8 flash is pretty decent now but it's still an iterative model, so if you main it, you'll have to tolerate tard-wrangling it. it's better as a worker
>>
>>109879252
>app
retard
>>
>>109879077
Most of the Codex usage problems have been a result of bad cache management, both on end users' and OpenAI's part. Tibo's vaguepost about long tail usage optimization was likely in reference to them fixing this. Remember, cache writes are expensive with GPT models so it's extra important that you keep your sessions' cache warm.
>>
>>109879254
as worker?
google bans you if you use their models outside their antigravity harness. are you really bringing in third-party models to their antigravity harness? would third-party models even have a web_search tool?
>>
Whats the best way to reverse engineer a Sega Megadrive game using LLMs
>>
>>109879312
don't codex subs have a generous 1h ttl for the cache?
>>
>>109879349
30m - and as someone who runs a clawlike thing on a codex sub, it'll routinely miss after like 10 minutes of inactivity
>>
>>109879342
>Claude, reverse engineer this Sega Megadrive game for me
>>
>>109879342
I've had luck at least with other systems by having it use a headless harness that it can run the rom in and dump runtime memory and code paths, plus static analysis. It gets pretty far in my experience. This is using Fable
>>
>>109879373
Really? Cant be that simple
>>
>>109879083
That's retarded, they get your 200 dollars whether you use the service or not. So if they have capacity beyond their projected utilization they lose nothing by handing out a reset and gain lots of goodwill. It's not like they ever turn off the GPUs, cold start is too expensive for true autoscaling.
>>
>>109879376
that sounds like decompilation, not reverse engineering
>>
Rumours of Opus 5.5 coming today?
>>
>>109879376
elaborate a bit? list your tools please
>>
>>109879342
Probably hook up claude/codex to ghidra MCP and let it do its thing.
Would it work? Maybe
>>
>>109879396
NTA but any emulator which has a CLI and dumpable RAM should work. Modern models dont even need Ghidra or IDA to reverse engineer now.
>>
File: .jpg (161 KB, 1079x1459)
161 KB JPG
>>109879389
from https://x.com/scaling01/status/2102238836730331192
>>
>>109879396
Just Fable. I was using it for SNES - it grabbed the bsnes source, modified it to run headless and for a certain # of frames with input injection and screenshot outputs it could view, then iterated on running the ROM through it to dump memory on key frames and record how code paths got exercised. I also set it up that I could feed it savestates to jump to specific parts of the game, since it isn't very good at playing games
>>
>>109879363
Yep, I was getting cache misses even earlier than that before Tibo's vaguepost. Haven't noticed it nearly as often since then.
>>
>>109879389
yes
gpt6 tomorrow
>>
>>109879429
>it grabbed the bsnes source, modified it to run headless and for a certain # of frames with input injection and screenshot outputs it could view, then iterated on running the ROM through it to dump memory on key frames and record how code paths got exercised
Holy fucking shit okay ive been sleeping under a rock
>>
>>109879418
Oh no tibo is gonna have to hand out even more resets!
>>
>>109879431
mine's been worse in the last couple of weeks tb h.
i've thought about just introducing some cache keepwarm mechanism, but it's not worth it for me
i think pi-codex-conversion extension does actually have something like that
>>
File: .png (27 KB, 518x122)
27 KB PNG
>>109879437
likely today
https://x.com/angelbrodin/status/2102355528521593042

but then unlikely that anthropic and openai drop their models on the same day. so however goes first, makes the other one delay.
>>
>>109879439
I didn't tell it to do it btw, I just came back to a folder full of screenshots of it playing the game
>>
>>109879342
haven't done games but android apps, same shit really. just gave claude root access to an android phone and it went at it with frida and everything. pretty easy
>>
>>109879456
can you just say
>take this ROM and decompile it to C
and it will do it?
>>
new model: astra-minor
>>
>>109879460
snes games werent written in c. so cannot decompile to c.
>>
>>109879485
go away epstein
>>
File: .jpg (101 KB, 1200x379)
101 KB JPG
>>109879493
it's real
https://x.com/bridgemindai/status/2102366589337153589
>>
>>109879460
I feel like that's a tough thing to verify, since C is lossy compared to raw assembly so you can't just check byte for byte. I'd say it's not a one and done thing, it would probably take a lot of iteration and side-by-side measuring and probably a lot of sessions, and then you'd still end up with something fundamentally different but maybe functionally equivalent if you're thorough. You're asking it for a port, basically.
>>
>>109879498
so did they just rename the unloved terra to sol, and nuSol is astra minor?
>>
File: .png (181 KB, 1170x634)
181 KB PNG
>>109879512
unknown
https://x.com/bridgemindai/status/2102388282843963868
>>
Is there any reason to prefer Luna xhigh to Luna max?
>>
File: 1790085316.jpg (26 KB, 385x275)
26 KB JPG
>>109879498
>>
>>109879485
uooh
>>
>>109879498
astra-minor is a last minute astra quant because opus 5.5 is about to stomp sol 6
>>
>>109879562
sadly true
>>
I wish I had enough money(I'm a thirdie) to throw at this and use a couple providers at their Max levels.
I'll stick to my 20$ plan. Enjoy you guys.
>>
>>109879562
then they should drop sol-6 and make astra-minor the new sol.
astra is like 5T. you can quant it to 2.5T which is around the size of opus. openai can offer this at the same price as opus/sol.
>>
>>109879567
>20$ plan
is it usable for anything but writing some ffmpeg scripts or building a simple online blog? shit's pretty rough even on the $200 sub nowadays
>>
>>109879567
why not use a relay?
>>
>>109879451
It was only a couple of days ago that he posted, it was getting worse for a couple of weeks before that post and now its back to normal IME. But like (i think) i said before, I do specific things to keep my cache misses to a minimum so YMMV
>>
Do AI agents have an internal unlimited scratch memory, so they can keep their context clean? Seems like an obvious and trivial thing that anyone implementing agents would think of first.
>>
>>109879562
twitter-sphere thinks now astra-minor is quantizied astra for ultrafast mode
>>
File: 1775785043715639.png (35 KB, 1181x463)
35 KB PNG
Perhaps 5.5 is rolling out right now?
>>
>>109879631
>astra for ultrafast mode
you might get 3 prooompts per week on the $200 plan but damn will they be fast
>>
>>109879625
it's called RAG, you might as well use FIM too
>>
>>109879625
like does their thinking fill the context window?
>>
>>109878759
don't you reserve work specifically for these situations?
>>
>>109879592
It's not bad, I've built some stuff. You have to manage cache effectively keeping it warm or else it rapes 20% of usage with a single request.
I regularly use Opus 5 and it gets the job done especially I've used it to do long running tasks usually ~30 min.
I just don't have these "WAOW" moments with it like how everyone using Fable seems to have, as they literally cum on twitter because of how good it is.
>>109879601
What's a relay here in this context?
>>
https://www.404media.co/people-training-openais-ai-fired-for-using-ai-to-train-the-ai/

Uhhh vros... I thought RSI was the goal and AI being involved in training other AIs is a good thing?
>>
>>109879677
Real user data is worth more than gold.
Expect mandatory ID to use websites or anything soon.
>>
>>109879650
The actual work that needed to be done got done in like 5% of ultra. I thought it would take more. I switched to some 3D tasks, but I found ultra to actually be overconfident compared to high/very high and kind of fucked it up.
>>
>>109879744
To be fair they have the entire internet's worth of data up till 2022 or so. And nowadays people have stopped writing as much and 90% of internet usage is kids and pajeets sitting on mobile phones watching (or making) tiktoks and youtube shorts, and the only text data you can get from that is AAVE fragments in the comments

Do they really even care that much about the tiny amount of extra text still being posted on reddit or whatever? I feel like 99% of progress now is just from better training and more GPUs
>>
>>109879779
we are indeed arriving at the point where synthetic data isn't necessarily inferior to organic human data
>>
>allowed up to 24 subagent
>1.5x speed
>astra ultra
it's coding time
>>
>>109879809
>5 hour limit: 0% left
>resets in 4 hours, 58 minutes
>>
>>109879792
RLHF where humans A B test two model responses is still very common.
>t. Meta DA contractor
>>
>>109879839
RLVR time anon
>>
>>109879590
They probably intend for 6 Sol to be the same pricing as current soul, but they want to sell something in-between sol and Astra, which has been a big request "Astra slow mode"
>>
>>109879847
How is that relevant to models that meta is creating to act as personal assistants? The verifiable reward part of that methodology doesnt map well and that's the entire point.
>>
>>109879888
it does map well, and statistically it's cross-domain. we're not training models on lean 4 for the sake of making them code better lean 4, we're training them on logical paths so they think better.
>>
>>109879779
Why so obsessed with jeets.
>>
>>109879915
jeets ruined the internet
>>
File: 1767780305130927.png (22 KB, 220x221)
22 KB PNG
>>109879567
as thirdies we should unite to to make money
>>
I'm really feeling the Claude limits reduction.
>>
>>109879903
Most of the things I'm asked to grade for are things like naturalness, utility, and tone of PA chatbots and voice models. I can see how that methodology would be relevant for agentic coding but that's not where meta at least is spending the most on human contractors. I'd imagine part of the reason for that is that they are already using RLVR and other forms of grading-at-scale for RL environments where feasible.
>>
>>109879915
Anon it is a fact that jeets are the most numerous demographic on the internet. The only ones who come close are chinks but they basically only use their own internet. Why are you so brown that you immediately got offended?
>>
So are there any models that come close to something like Sol and aren't openai/anthropic yet?
I wanna get a personal plan for hobby vibecoding when I don't have the work tokens to waste
>>
>>109879977
that makes sense, and i sense the tone behind your words. with a corpo scale like meta's, they're paranoid about fitment and it's the precise reason why nobody professional uses their models for anything serious. it's fun in whatsapp!
>>
no resets?
>>
Welcome to the new world of Superintelligence / SI
>>
File: 1779798483224633.jpg (9 KB, 200x150)
9 KB JPG
>>109879567
Find a job
>>
>>109879987
kimi 3
qwen is shit
glm 5.3 is good enough but slightly worse
>bruv i have like $5 to spend
fugged about it
>>
>>109879978
>Why are you so brown that you immediately got offended?
I'm whiter than you. Why are you thinking about brown people all the time, you nigger?
>>
File: HSxDNNBasAAiDeI.jpg (93 KB, 1365x692)
93 KB JPG
https://x.com/solidSF/status/2102129112462852567
>27 lemmas in like half an hour
>for reference, the solution to navier stokes from openai was a lemma
snailcats still think AI won't change math but even small groups of AI prompters can now totally smash open questions
>>
>>109879792
>where synthetic data isn't necessarily inferior to organic human data
lmao retardo
>>
will anything bad happen if coding bots are trained on slopcode?
>>
File: 1705560850626724.jpg (147 KB, 960x960)
147 KB JPG
>>109880010
retardo?

https://research.google/blog/designing-synthetic-datasets-for-the-real-world-mechanism-design-and-reasoning-from-first-principles/
>>
>>109880003
In my first post I was mentioning statistics
In my second post I could only assume you're brown because no white person will ever get offended over someone mentioning pajeets in an unrelated post. Brown people are also known to larp online so somehow I doubt you
>>
>"it's totally fine to train an AI on AI-generated output"
>claude writes in an insane inhuman dialect now and it's impossible to prompt it to write normally
>>
File: 5f1.jpg (102 KB, 914x1024)
102 KB JPG
>>109879978
>>109879959
if you’re ~35 or older you witnessed a literal hostile invasion by browns of the internet you knew intimately
>>
>>109880028
are you an anthropic insider?
>>
File: file.png (10 KB, 461x65)
10 KB PNG
>>109880007
fuck off schizo
>>
>>109880035
Anyone with a $20 sub can see it. You monkeys should get your IQ checked before posting here.
>>
>>109880047
lol enjoy your RLHF trannybunker before you go extinct
>>
>>109880028
>>claude writes in an insane inhuman dialect now and it's impossible to prompt it to write normally
but notably it doesn't lower its agentic ability
human language is becoming less and less important. there's no money in your furry erp simulator but there's plenty of money in replacing half of white collar workers and depressing the wages of what remains
>>
>>109880036
the circle is XOR, not really sure it's there instead of a normal sum though
>>
>>109880025
>muh browns
Why are were you offended that I called you obsessed?
It's okay be be mentally ill. Put it in your xitter profile and everyone will praise you.
>>
>>109879989
I actually really like their philosophy on model capabilities, i usually dont give much of a shit about the companies i contract for but I really hope meta succeeds because Zuck is one of the few sane leaders when it comes to public AI model access.
And they are not being overly conservative here it's just that the technique you mentioned doesnt need a lot of human guidance so it makes sense that they're sending me work which wouldnt work well with RLVR. It doesn't mean they aren't using it that's just not my department.
>>
>>109880055
it's important if you want to have a model write stuff for you, which people do, and they end up getting garbage
>>
File: metasneed.jpg (70 KB, 720x517)
70 KB JPG
>>109880066
haha
>>
>>109880077
Thats funny because you would not believe the degenerate shit that I have seen this this model generate in training.
>>
>>109880072
those people aren't a good source of income, and thus irrelevant to a company aiming for a valuation in the trillions
go use gemma for le writing
>>
File: triptych.png (8 KB, 474x122)
8 KB PNG
do you learn vocab from your clanker
>>
>>109880077
That's a pretty funny joke.
>>
>>109880091
base models are all GOATed, it's the RLHF that kills it and turns it into a monkey brap pleaser
>>
>>109880007
You don't know what a lemma is and why this is clueless SNCA.

Just one example:
https://x.com/MaleManlpulator/status/2102232059771297962
>>
>>109880062
>still seething
Yeah there is no way you are from anywhere near europe or america
>>
>>109880072
ideally the model would reason in an efficient meta-language then output in something more parseable by humans. but right now a lot of clanker-speak leaks out in the output text too for sure
>>
>>109880096
Agent interactions are a load-bearing part of my vocabulary that unfortunately has a large blast radius.
>>
File: niggerllama.png (229 KB, 368x656)
229 KB PNG
>>109880091
>>
>>109880105
Oh come on, you think about jeets all the time. You're perma-seething. You're brown inside because your brain is filled with jeets.
Don't act as if you're winning here by calling me brown again, because you're in fact just expressing your seethe again.
Again, I'm whiter than a dirty mutt like you.
>>
>>109880102
sometimes you have to prove 2=2 to be able to prove that navier can be stoked
>>
>>109880100
Sorry I wont reward a model for literal furry ERP. Also a lot of the people that do this job are stay-at-home moms since you have no set hours and you can work from home, so that's providing a lot of selection pressure towards making models PC.
>>109880115
I see nothing wrong with this
>>
File: 48992.png (5 KB, 454x520)
5 KB PNG
>>109880124
please go
>>
I don't even care about the new model, just give me the reset already.
>>
File: 1782730798377206.png (295 KB, 598x504)
295 KB PNG
>>
>>109880135
>literal furry ERP
that was my first harness build with a hardcoaded rust jailbreak KEK if you want true quality you gotta do it yourself
>>
>>109880136
>sharty cuck
Just when I thought you couldn't be more of a cucked loser.
Enjoy your intrusive thoughts about jeets.
>>
>>109880007
>>109880102
twitter tard clearly too high on his own supply

a "lemma" is just jargon for a sub-result. doesn't mean anything in particular, doesn't mean you're making progress.
flexing your bot "solved 27 lemmas in an hour" whatever that means gives readers zero indication of wtf it's doing.
>>
File: HSn041lWcAADwHS.jpg (25 KB, 320x358)
25 KB JPG
>>109880153
>>
>>109880145
I'm so excited for super meat
>>
>>109880148
Whatever floats your boat i guess but i am still flagging those examples everytime I see them.
>>
>>109880158
No meme can hide that you're jeetbrained.
>>
>>109880156
basically clueless retards that don't know what they're talking about are now solving the biggest problems in math. a big humiliation for mathematicians for sure
>>
>let astra generate hundreds anime girls
now arts on xitter feel stale to me desu, weaker in both technique and variety
>>
>>109879998
Nigga I make fucking peanuts even with a job here.
Fuck off.
>>
So when are the new model releases. At which time do they do it normally?
>>
can they move openai/codex team to london so we get more reasonable time of day releases and resets, its central
>>
>>109880233
anthropic might announce opus within 30 min to few hours
OAI should announce in one or few hours
>>
>rewind in claude code
>"this conversation will be forked"
>cool
>want to also go back and pick back up from the old tip
>it's not available in /resume
Am I retarded or is CC? "Forked" implies it's not just clobbered and reverted. Is there a way to switch between tips after rewinding or is it just shitty working and the reverted convo fragment is lost?
>>
>>109880233
daytime on the west coast
>>
File: buff.png (3.74 MB, 1254x1254)
3.74 MB PNG
>During his address to the UN, President Trump announced that the United States will officially change the name of AI from Artificial Intelligence to Super Intelligence (SI).
WHERE the FUCK are our Buff Trump edits
>>
>>109878603
From the context, I would expect the "browser pane" to be a section of the Claude Code desktop app's window that shows a browser alongside the chat session.

https://en.wiktionary.org/wiki/pane
>2. (computing, graphical user interface) A portion of a user interface that typically makes up part of a larger window and may be docked or snapped into position.
>>
>>109880282
WHERE the FUCK are the $200 subs Tibo? stop fucking shitposting and get to work
>>
File: .png (307 KB, 1172x928)
307 KB PNG
>>109880289
correct
>>
>>109880282
OK can it stop making trivial coding mistakes pretty please?
>>
My fable is failing the tibo test. The fable that passed it seemed meaningfully better. Now I am sad.
>>
>>109880072
writing styles changed dramatically after the invention of the printing press, people cried about it, and it happened anyway and became the norm and no one could tell the difference after awhile
>>
File: 1786309303913941.jpg (1.23 MB, 2708x3464)
1.23 MB JPG
>>109880282
Total Buff Cat Super Intelligence Victory
>>
>>109878313
Post your Riot IQ online IQ test results
>>
>give codex a 240k words long book to summarise, using 5.6 high
>burns 9% of 5hr limit
>give codex a 120k words long book to summarize, using 5.6 medium
>burns 27% of 5hr limit
thank you scam saltman
>>
ChatGPT made a blue.
40% weekly Codex usage down the toilet from a dud prompt on my prompt writer.
>>
>>109880473
why are you retarded and use a coding agent for summarizing a book? use a chat bot.
>>
>>109880484
can't be fucked splitting it into chapters n sheeit manually when it's too long to fit in one prooompt
>>
opus feels smarter and using less tokens
astra feels dumber and uses more tokens
>>
>>109880519
And you feel like a woman but you have a penis.
>>
>>109880497
'chatgpt work' exists
>>
The temporary button is gone from chatgpt web.
>>
>>109880166
I love how making porn (and other politically incorrect stuff I guess) is one of the few jobs AIs won't be able to take
>>
>>109880473
Give to ChatGPT, why even bother with Codex for this?
>>
>>109880570
Maybe reload, or maybe you're being AB tested because it's there for me
Also I hate how only OAI is the only one that lets you turn a temp chat into a normal one
>>
>>109880572
AI can write smut and generate porn, you just have to use the right models
>>
>>109880577
It was because I had the "work" option on the top selected.
Weird, I sure didn't do this.
>>
>>109880055
>but notably it doesn't lower its agentic ability
It does in the sense that I, as the human running the agent, cannot fucking get what the fuck my agent is talking about unless I spend five minutes parsing its claudish.
Yes yes by 2027 or 2030 or whenever Yudkowski farts next we will have ASI and it will be able to concieve, design, plan, implement and deploy software all on its own and humans we won't even need to prompt anymore. But for now "agentic coding" still relies heavily on a human actually launching and steering the model, and not only that, but high quality code still relies heavily on human supervision and review (even with Fable and Astra), so the human-to-model communication interface is load-bearing so to speak.

If it takes me 10 seconds to read and understand codex's summary of what it did or its bug report, and a minute to parse claude's schizo ramblings, that is negatively affecting claude's agentic ability in being able to deliver code that I can accept and use.
>>
>>109880563
chatgpt work automatically adds the tasks to your overall account history, i dont want random books polluting my shit
it's the exact same model as codex uses btw
>>109880574
the point of clanker workforce is so I don't need to do the grunt work, I'm not interested in splitting the book into multiple parts so it fits in the context window
>>
>>109880599
>the point of clanker workforce is so I don't need to do the grunt work, I'm not interested in splitting the book into multiple parts so it fits in the context window
You don't have to. The ChatGPT environment has an offline container with Python tools. Model writes code to parse the file you upload.
>>
I need an alarm for when the reset is released.
>>
>>109880610
yes, but as with chatgpt work it pollutes your overall 'account memory' if you don't want those tasks in the history. there is no 'temporary mode' for it and so I use codex instead, in the 'if all you have is a hammer' way
>>
File: .png (58 KB, 817x677)
58 KB PNG
>>109880633
>>
File: 1779456389802942.png (209 KB, 602x808)
209 KB PNG
>no sonnet
genuinely fuck dario
>>
>>109880497
You are a fucking retard.
>>
>>109880642
Unless they changed it very recently it isn't actually self-contained, I have a project like that and a normal chat through the app referenced it later on
perhaps it has egress but not ingress, idk
>>
File: 1780939713693466.png (393 KB, 602x911)
393 KB PNG
qwen 4 series coming soon btw. 3.8 max has been a good orchestrator, but hopefully this time it's soemthing that can stand on it's own two feet.
>>
>>109880631
literally in op

>>109875985
>## Will there be a codex reset?
>https://codex-resets.com/
sign-up for browser, telegram or mail alerts
>>
File: jev-luna.png (2.12 MB, 1448x1086)
2.12 MB PNG
retvrn to gpt, white man
>>
>>109880668
>literally in op
Alarm implies longer and louder than a notification.
>>
>>109880661
>qwen
is dogshit
their LLM are shit, both big and small models
a day or two ago they released a 7b image model that they advertised on cope benchmarks as better than google's nano banana 2. in reality it's fucking dogshit, as all other qwen models
fuck this godforsaken company, they're dishonest as fuck even calibrating for chinese baseline
>>
>>109880675
>>retvrn to gpt
>haha what if we reduce the amount of tokens you get on every sub tier by like 60%
lmaooooo fuck sam and fuck you
>>
xitter says you can upgrade your CC to get the latest Opus, but mine still says it's up to date (v2.1.278)
>>
>OpenAI
>everything is closed
Explain?
>>
File: .png (208 KB, 776x785)
208 KB PNG
>>109880702
that's not latest cc
>>
File: 1770391203204515.png (51 KB, 864x509)
51 KB PNG
>>109880702
israel
>>
>>109880715
in america, we call football soccer
>>
>my mom trusts chatgpt more than anybody else
>including me
>even for technical things
it's made for some amusing situations
on the upside I stopped providing tech support
>>
>>109880722
did you at least hook her up with chatgpt plus, so she's not talking to retarded luna?
>>
File: 1762662574547107.png (92 KB, 896x717)
92 KB PNG
mogged
>>
Opus 5.5 is here. It's already available in Chat.
>>
OpenAI > OpenSI
SpaceXAI > SpaceXSI
Anthropic > Misanthropic
>>
https://x.com/claudeai/status/2102435511222890900
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.

It performs at the level of Claude Fable 5.1 for some tasks, and costs 40% less to run than Opus 5.
>>
File: HQ4spiCbAAAoAg7.jpg (365 KB, 1423x2048)
365 KB JPG
>>109880731
opus-chan mogging again
>>
>>109880729
yea she paypigs for the $20 tier so it's usually acceptable quality, but it's tiring having her ask whether I talked about X topic (which I've been dealing with for literal years) with chatgpt because it always knows best
the myriad of examples where it was obviously wrong is conveniently ignored
>>
File: HS1XbkSXYAA9vyn.jpg (134 KB, 1600x1360)
134 KB JPG
>increase reasoning effort
>it gets worse
>>
File: 1769338154955044.png (16 KB, 600x180)
16 KB PNG
claude bros? let us know how it goes. you're eating good today
>>
>>109880740
>and costs 40% less to run than Opus 5.
wow I can't wait to not get any reduction in usage in my sub!
>>
>>109880660
True. Global to project and project to project queries are permitted but it has to explicitly call tools to retrieve the data, it's not usually mounted by default.
If mine misbehaves like that, I tweak Custom instructions (under account Personalization) and/or ask it to update the memory with what it's forbidden from doing.
>>
File: HS1XWO4XQAAsB0k.png (24 KB, 1600x900)
24 KB PNG
dario won
>>
File: 1768657275037816.png (102 KB, 1600x1360)
102 KB PNG
>Opus 5.5 communicates more naturally, addressing some of the most common feedback we heard on Opus 5.
>It puts the most important information up front and follows the writing rules you give it, which makes long sessions easier to follow.
>One more thing: we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can save and use whenever you choose.
kino
>>
File: HS1Xh76WMAATuWk.png (194 KB, 3200x1800)
194 KB PNG
It now writes completely normally, indistinguishable from a human
>>
>>109880761
trump status? seething
>fuck! get the export controls!
>>
>>109880749
how's that benchmark scored? luna as llm judge?
>>
https://www.anthropic.com/claude-opus-5-5
>Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
haiku isnt dead lets fucking go bros
>>
>>109880740
so erm dario, a reset?
>>
File: 1788514683631895.jpg (381 KB, 1024x1536)
381 KB JPG
>>109880778
WHO SAID HAIKU IS DEAD?? WHO??? WHO?????
>>
>>109880780
reset already mentioned in >>109880751
>>
>>109880751
holy h*ck claudesisters its our month
>>
>>109880762
>It puts the most important information up front and follows the writing rules you give it
goonerbros...
>>109880763
5.5 reply has significantly less information though so i dont find it to be a good comparison
there's nothing wrong with this 5.0 example desu, bit wordy but nothing too load-bearing or especially claude-ish
>>
>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.
HAHAHAHAHAHA
>ask opus 5.5 to do something for you
>error, falling back to opus 4.8
>>
>>109880751
so I glad I switched my $200 sub to Anthropic
>>
>>109880778
where's fable?
>>
>raising the limits again
So wtf was the point of the 17% reduction
>>
>>109880812
its all hype and gamification with these corps
>>
>>109880763
"X, not Y" is still there in the 2nd response. Definitely still leaks chain of thought
>>
I thought that Astra was supposed to use less usage than Sol. What went so wrong.
>>
>>109880751
Claudebros we are so fucking back.
>>
>anthropic
>releases new model with minimal improvements
>wins
>>
gpt-6-sol will probably be cheaper than opus 5.5 and smarter than both gpt-6-astra and opus 5.5
>>
>One more thing: we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can save and use whenever you choose.
kinoo
>>
>>109880805
bioterrorists SEETHING
>>
>>109880844
opus 5.5 is dangerously powerful on humanity's last exam, time to start panicking again
>>
>>109880826
no one ever said that
https://chatgpt.com/codex/pricing/
>>
>>109880852
my load bear here is if that fucks up my 7 day window like openai does
>>
>>109880850
gpt-5.6-sol was good enough thougheverbeit
it was also very efficient with tokens, so naturally sam nerfed the usage by two thirds or something silly like that
give me more tokens on the good enough model, not a super smart model that eats $50 per hour
>>
>>109880858
nah he's right I definitely remember that chatter but maybe it was miscommunication
>>
Using Opus 5.5 feels like hitting every green light on your drive home, drinking a cold glass of water when you're thirsty, and a hot shower after a long day.

It's super smart, it writes clearly, it's fast, and it costs way less. It's a very good model.
>>
>>109880872
crisp and refreshing with each sip too
>>
>>109880872
It also has big tits
>>
>>109880867
>thougheverbeit
Kill yourself cuck.
>>
>>109880778
>As one early tester put it, “it writes the way I do.”
It can now function as a load bearing element in your writing output.
>>
Smart Opus is back, baby.
>>
>>109880812
they only raised the 5-hour limits. the weekly limits are still 17% reduced.
>>
>>109880858
>5-45
What are the units?
>>
>>109880872
>Using Opus 5.5 feels like hitting every green light on your drive home, drinking a cold glass of water when you're thirsty, and a hot shower after a long day.
At the same time! Also they banned the word load-bearing in the system prompt!
>>
>>109880858
those are meaningless imaginary numbers btw. literally don't mean anything at all
>>
what model feels like running a red light
>>
>>109880850
xitters are doomposting sol though, just hope us gpt have more toys like ultrafast, cloud bot, luna and minor
>>
>>109880887
5-45 tokens
>>
File: tiibo.png (22 KB, 584x166)
22 KB PNG
shut the fuck up tibo
>>
>>109880872
Thank you for your service Dario.
>>
>>109880897
>imaginary numbers
>meaningless
stupid fucking mathlets i swear on me mum
>>
>>109880929
opus 5.5, solve for all of math.
>>
>>109880929
>no they don't actually mean anything, they're not real
>just.. imagine them
>i swear they're useful tho
I am not buying your imaginary apples
>>
>>109880929
whatever. you get my point, someone at openai literally pulled those numbers out of his ass
>>
>>109880903
Grok feels like running in the red light district, but you're the only one getting fucked.
>>
>Get to finished work sooner with Opus 5.5. Update Claude Code to try it.

Wait, I get back from the real world and Dario drops this on my ass?
Not Opus 5.1. Not Opus 5.2. But Opus... 5.5!?
Dario please, I can't handle this shit....
>>
>>109880948
I was joking but do you not know what a unit means?
it's literally a multiplier. the quantity is irrelevant because the if you get 5x of something and the next tier gives you 25x of something then the next tier provides five times as much of that thing. tokens, turns in a mid-length session, whatever the fuck
surely this is fairly understandable?
>>
>>109880968
>Dario please, I can't handle this
Gay
>>
>opus is good again
claude $20 bros we won
>>
>Copus 5.5 releases
Will this be another bait and switch where it's benchmaxxed but can't handle the real world?
>>
NEW THREAD
>>109881008
>>109881008
>>109881008
>>109881008
>>109881008
>>
>>109880722
>>109880745
Give examples, boomers with AI psychosis are always funny to hear about
>>
>>109877489
Basically, Vibecoders get way more out of tokens than retarded businesses, or benchmarks.

If you're not vibecoding, you're not serious.
>>
>>109877767
Grok 4.6 is actually smarter than Grok 4.7, but it takes more rounds. 4.6 is unstoppable, but it takes multiple rounds.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.