A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs. Physical harness edition.-- Frontier models - start here if you have $20 or sohttps://claude.com/product/claude-codehttps://developers.openai.com/codex/cli-- B-tierhttps://x.ai/clihttps://platform.deepseek.com-- Prompting / context / skillshttps://arps18.github.io/posts/claude-code-mastery/https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/https://github.com/mattpocock/skills-- Other editors / terminal agents / coding agentshttps://pi.dev/https://opencode.ai/https://cursor.com/docs-- Benchmarks / rankingshttps://www.tbench.ai/leaderboard/terminal-bench/2.1https://artificialanalysis.ai/-- What we’ve donehttps://vcg.gitgud.site-- Previous thread>>109640756
im vibe coding a 4 dimensional operation system that is stored in the blockchain for security while also being a passive crypto miner with next to no overhead.
Sorry wrong slop.
>>109645851>>109645857Cute slop
>>109645846So just do what I describe but then incorporate it into a text file it can repeatedly reference. I've worked on projects using a DEVLOG.md file too which are useful post compaction.
>>109645890>DEVLOG.mdThat sounds like a commit log, but shittier?
Is there anyone you can try Cerebras inference or even just watch a video of it? I am genuinely curios. All I can see is nice chart line go up
>>109645934>anyoneanywheresorry my inner ESL emerged
>>109645857I look like that and I say that.
jesus fucking christ
>>109645926A properly written commit history details what specific changes were made commit to commit. A devlog md file imo should be used to help tardwrangle the model into implementing changes and tests in specific ways without you having to handhold is (as much) every time it does something. Especially post compaction when relevant context is inevitably sacrificed. A model's own compaction summary alone is good for a basic summary of what was done bit can be lacking in crucial information. This of course varies person to person, project to project, and model to model. Some model write very in depth and detailed compaction summaries like Qwen and Kimi models and others don't.
>>109645969What did you have it do? Reading a big ass codebase I assume?
>>109645998It's making an application and it keeps re-reading through all the code it already wrote and then compressing and reading through the code again.
>>109645990You probably just have to instruct them to write proper commit messages in AGENTS.md. The training data is probably very diverse, ranging from "fixed bug (I'm retard)" to Linux-kernel-tier commit messages.
Is anyone here actually using ADEs like Orca?
>>109646020>application What platform and file format? >>109646022Commit messages are description of what was done. It's the Agents.md (or in my case a DEVLOG file, are wgat do the tard wrangling. Unless you A) have it go schizo more on the commit description verbosity andB) tell it to read that shit each timeIt doesn't make much sense to rely on commit messages alone for tard wrangling. Especially if you hit cloned someone else's project that had a preexisting hit history. Gut commit messages are usually never super verbose and describe what was done concisely with a reasonable amount of detail. It's like expecting yourself to be the SME on a book but you only want to read cliff notes and only read OTHER people's commentary on the media instead of just reading the source material
having fable orchestrate and do code reviews with opus implementing is absolutely the winning move. it's like all the annoyances or mistakes of opus being corrected in real time before it's done with it's prompt. only downside is needing the sub to get fable access
>>109646083protip: you can have any agent do what fable doesjust ask an agent to read the session log, then tell it to make opus behave like fable did as an orchestratorboom, now you have an opus orchestratoror slowly walk opus through one or twice doing something, do the same thing, ask an agent to makes a prompt that will have an agent do what you didnow run an opus orchestrator with your inferred behavior as the prompt, with it handling a normal opus
>>109646061>What platform and file format?idk I just told it to make an app on linux
/vcg/ just isn't the same without exoplanet fag posting daily updates
>>109646218he actually finished, thoughyou can admire his work anytimehttps://github.com/Titanean/Tadmor
geniusabsolute fucking geniusboth the paper authors and the youtuberhttps://youtu.be/eJuYBNrD8HI?t=4m39s
>>109646223>>109646115garbageseveral skims over the readme and I don’t know what it isan app?how do you use it?why is there no hosted example or demo?there’s nothing to look at or use, nothing actionable
>>109645823>ultrametric faggot here>just did some more shit >https://github.com/sneed-and-feed/adelic-spectral-zeta/blob/main/docs/zk_verifiable_ultrametric_attention.md
AAAAIIIEEEE I had 86% usage remaining and Tibo just nuked that all with no survivors
>>109646271we will miss resets existing when they are gone
>>109646262>several skims overDid you? It's clear my meant to be build the android apk yourself via gradle. Definitely not user-friendly (for mental midgets) but there are instructions on how to get the app
>$20 goys acclimate to having a 5x plan due to all the resets>blow through their entire limit in a day then act like it's OpenAI's fault>go on twitter and make shit up about nerfs and plan usage>OpenAI has to bring back the 5hr window for the goy tier since they clearly cant control themselves and it's bad PR to have a bunch of jeets saying they can only use AI for 4 hours a weekThey are so fucking stupid. You cannot do meaningful work on those plans.
>>109646299fair. i have literally never seen an android app repo in my life until today
>>109646317>You cannot do meaningful work on those plans.already implemented 2 features with Luna max
>>109646317yeah I had to switch to a x20 plan pretty quickly, my time is far too valuable to scrape and cope on lower models or begging for resets
>>109646317it’s easier to do meaningful features without the 5h limits because a lot of what I did on the $20 plan was, like, a third of my weeklybut I wasn’t a giganoobfwiw Grok has no 5h limits on the $30 plan, and it ships with one banked reset that expires in like two weeksso if you accidentally gulp you get to clear your mistake once
>>1096463175h limits are BASED and REDPILLED
>>109646329To be fair to you a lot of them so have a "releases" section where you can just download the android apk and install it on your phone (assuming your phone'a settings allow non-play store app installs)Like this one: https://github.com/moffatman/chan/releases/tag/v1.2.9+121-forked-engineThe Titanean anon's per project was probably meant to be a personal one and wasn't concerned with how easy installing it for others was. To NOT be fair to you, just point an llm at the git cloned repo and ask it to expain how to build it or just have it do it for you assuming it has th right tools and libraries to use on the machine it's on
Codex is a good smart autist but i fucking hate it autocompacting, i try to make it write shit out before it goes and lobotomizes itself
What's your ADE of choice?Orca?Paseo?OpenChamber?cmux?
>>109646444>used to be an option to disable autocompact but it got removed in juneALTMAAAAAAAAAN DU JUDE
>>109646444>>109646459just copy paste this into your codex:hey, don't need you to do this, but in theory could you write a codex hook that:1. monitors x% of context window is used2. and if it is, sends a prompt to use a skill to write a handoff note?
>>109646515that's literally what compaction but shittierthere's basically no loss of context after a codex compaction
The right mental model (huehuehue) of talking to a chatbot is assuming he's like a trainee who's very capable and can look up anything, but who is an autist who doesn't know anything of the real world.
>>109646444What triggers the auto compaction? If it does that when it reaches the context window then.... It's like....supposed to do that. Allowing it to blow well past its context window with lead to it becoming even more retarded than what you're supposedly experiencing now. Does codex suck at writing compaction summaries and picking up where the past context left off? Even if it does you can just simply have it right a devlog file like I explained earlier.
>>109646459
>>109646518i tend to agree with you on this for the most part, but some niggas swear by handoff notesbut then also too dumb to know what hooks are lel
Are these random?
>>109646518>there's basically no loss of context after a codex compactionwhen you lie about this it makes it less likely for people to want to use the product btw
>>109646576codex? I think that happens if you had like 99% or 100%, not got one for ages.They are just resetting you with no care for un-used credit now
>>109646582I barely use codex anymore, I'm mostly on Opus/Fable. But I work on a fairly large project and when I do use Sol it's on xHigh or so. Last time I triggered two compactions in 3 prompts, and the result was perfect. It solved the issue Fable/Opus were struggling with.If I ask you to measure the loss of accuracy after a compaction the only source is your opinion straight out of your ass.There's no real data that shows quality degradation pre/post compaction. You're basically a superstitious faggot
>>109646585nope. weekly resets are handed out separately.>>109646576pretty much. you have to follow tibo on twitter if you want to keep track of this shit.
>>109646576No, these were just two banked resets they announced.
>>109646627two?
>>109646644One on sunday, one today.Don't know what the fuck they're doing.
i'm on the individual $100 plan is it worth trying to set up this new $100 premium business plan from openai?
>>109646651>one todayoh fuck yeahnot a banked reset for me (i was out of usage anyways), literally burned everything last night in preparation for the 5 hour limits coming to the Plus plan and now I'm back to 100% remaining hell yeah
>>109646651there's no second banked reset nigga what are you talking about
>tfw cannot follow tibo on xitter anymore because Elon shut down xcancel
>>109646651I’ll never get tired of people being assmad about getting free shit.You’re literally not human. You’re an animal and you’ve become dependent and forgot how to feed yourself.Disgusting.
>>109646676Well my usage was unexpectedly reset today and that's what happened to others too.
>>109646710that was a regular reset, a banked reset is one thats saved so you can use at any time. it expires after a month
>>109646718Why would it be a regular reset?
>snailcat (forma de poopoo) realizes how raped his profession has becomeLEL
>>109646751do you understand the difference between a banked reset and regular resets, bro?genuinely can't tell if i'm being trolled
>>109646754>not being CS majors is worse than not being programmersBut yeah, programming is all about how well you can chat with a potato now.
>>109646754>hopefully get paired with girlslmao man lmao
>>109646754Absolute chads rolling in with their subscriptions lol
>>109645935it's quite literally one mentally deranged person who has been relentlessly spamming the general and creating the vcg threads with his ai slop. nobody else says it or cares. you get persons like that in many areas of 4chin
>>109646754>3 non coders made friends, excluded OP from the groupchat>later got jobs from this networking while OP seethes at home
>>109646754>hackathon>girlsuuuh who tells him
I'm vibecoding a kerbal successor. Wish me luck. Sol made an entire restricted n-body simulator out of NASA data to validate correctness against, and it's gotten launch, orbits, reentry and even fucking docking working. It's fucking insane.I don't know how I'm gonna do models and sound though.
>>109646855This. Retarded normalfags win, autists who invented literally all technology humanity uses lose.
>hitting the claude 5 hour limit after slopping for 9 hoursHow???
>>109646906variable token consumption rateare you using workflows or something where you might have multiple Claude subagents working at one time?
>>109646936>variableIt was a single session. 5 hours isn't a variable fucking time.
opencode with basically any model lists files in a directory>ls -lclaude lists files in a directory>:; (ls)"" /dev/null/..${HOME}/.why does it have to be so fucking extra
>>109646886Good luck, it's a many-month investment, especially if you want art (vibe-arted or not).
the fuck does updating codex even do, it asks for it daily
why did chatgpt desktop hide max reasoning setting hmm
>codex
>>109647017They probably just make automated daily builds.
>>109647045>he doesn't sol/max everythingngmi
im letting fable write to a json file and luna max to another that's how they communicate fable does the planning, is this most efficient?
>>109647045fuck terra all my homies hate terra
>>109647075>>109647045>sol high isn't on the list despite being the only gpt model that is pareto
>>109647101alright, alright, I'll swap the medium for it...
>>109647101yeah I would swap capable with sol low and frontier to sol high personally. also terra fucking sucks my BALLS
>>109647101>no sol medium shown >shitty anthropic models that can only be used “legally” in their shit harness or app>API only eastern models compared to subscription subsidized western modelsYour benchmark graph sucks
>>109647101Isn’t grok 4.6 now? Is it really better than Sol? Is anyone here using it?
>>109647158no and no.Grok is basically worthless, they occasionally benchmarkmaxx a couple of comparisons but actually using it is so clear it's worse than the leaders At best it might be usable if you are a thirdie with no budget
>>109647136sol medium is just above terra max
>>109647101>>109647184Where is terra medium? Everyone here hates it but I'm using it.
>>109647184>Facebook model can beat sol medium in priceAre you trolling?Musk and Zuckerberg beat Sol Medium?Is /vcg/ just way behind the times or are these benchmarks legit memes
>>109647207kek, according to the cube, terra medium might not even show up in that graph, it’s below Luna high
>>109647218completely meaningless chart that redditors and twittertards like to throw around. ignore it, he's a tourist
>>109647207
>>109647218facebook i doubt but i've heard people say grok is now sort of decent and very cheap because not a lot of people switched to it and they have lots of compute, is it worth trying?
>>109647207>>109647207
>>109647240Brb getting deepseek.
>>109647249that graph is old they raised the prices on deepseek
>>109647249>>109647271Also if you’re using the ChatGPT subscription (almost certainly), your actual price per task is 10x lower than that graphDeepseek doesn’t have such subsidized subscriptions
which model fits this
>>109647331Luna Xhigh-Max
>>109647010doesn’t it have a release notes like claude(1) does?
>>109647158I’m using it but I’m a Claude (mostly Fable) main and so I don’t have a good idea of how good Sol isI want to root for Elon everything but the grok client, at least in the fullscreen UI, makes my laptop warm all over when codex/claude don’t
>>109647356elon is mining bitcoin on your machine
>>109647372I think it’s the TUI. It’s very active with lots of different blinkenlights not near the bottomand it has an option to run at faster than 60Hz if your monitor and terminal support doing so (my terminal doesn’t, but my monitor does)
>>109647356The TUI? It eats your CPU? How in the world did they fuck up.
Had a nice run with Sol on Codex having it build a replacement for a tool but the business closed before I could deliver ffsNo idea what to burn the usage on now :(