A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.// What “vibe coding” is, and how to do ithttps://simonwillison.net/2025/Mar/19/vibe-coding/https://simonwillison.net/2025/Mar/11/using-llms-for-code/----// Frontier models using fully-general tooling — start here if you have $20 or sohttps://claude.com/product/claude-code (Fable 5 is the best LLM available, requires Max plan)https://developers.openai.com/codex/cli (Essentially scamming you with LLM degradation and resets that reduce your usage, but still the second best option and arguably the best bang for your buck in the 20$/month plan range)// Worth it for code, but the frontier models above are betterhttps://x.ai/cli// Not worth it for code, but maybe good for other thingshttps://antigravity.google/product/antigravity-cli----// Prompting / context / skillshttps://arps18.github.io/posts/claude-code-mastery/https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/https://github.com/mattpocock/skills — /grilling is a favoritehttps://github.com/DietrichGebert/ponytail// Other editors / terminal agents / coding agentshttps://osaurus.ai/https://pi.dev/https://opencode.ai/https://cursor.com/docshttps://docs.windsurf.com/https://docs.cline.bot/https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent// UI/Frontendhttps://www.figma.com/make/https://www.anthropic.com/news/claude-design-anthropic-labshttps://uiverse.io/https://ui-ux-pro-max-skill.nextlevelbuilder.io/https://stitch.withgoogle.com/// In-browser builders / hosted vibe toolshttps://bolt.new/https://replit.com/https://docs.github.com/en/copilot/tutorials/sparkhttps://v0.app/docs// Benchmarks / rankingshttps://www.tbench.ai/leaderboard/terminal-bench/2.0// What we’ve donehttps://vcg.gitgud.site// Previous thread>>109463778
>>109469383The clanker has baked a new cube!https://bradthomasbrown.com/pareto-3d/You can now switch to and from orthographic perspective, makes it so depth isn’t interrupting comparisons so much. You can disable model families and hide non Pareto points.There are no paid offers or plans, no donation links, no ads, no monetization, no tracking. Data comes from DeepSWE, it is not an endorsement, the clanker picked the source.
>>109469383kek
>/vcg/ News>July 24th - Anthropic drops Claude Opus 5Flagship for knowledge work/automation/coding/agents. Approaches Claude Fable 5 capability in a lot of categories at half the price ($5/$25 standard, Fast mode $10/$50). 1M context, effort dial, becomes default on Claude Max.>July 26/27th - Moonshot ships full open weights for Kimi K32.8T-param MoE (largest open-weight ever at the time), 1M context, multimodal. Hosted version had already been live; weights hit Hugging Face under their license.>July 27th - Alibaba quietly launches Qwen3.7-FlashCheap 1M-context multimodal (vision/reasoning/agent workloads), $0.03/$0.13. No big announcement or full report.>July 30th - OpenAI slashes GPT-5.6 Luna 80% and Terra 20%Luna now $0.20/$1.20, Terra $2/$12. Sol unchanged but gets Fast mode (up to ~2.5x speed). Efficiency gains (some from Sol optimizing its own stack) + competitive pressure.>July 31st - DeepSeek promotes V4-Flash-0731 to productionSame 284B/13B active MoE + 1M context as the April preview, but heavy post-training upgrade. Beats their own larger V4-Pro-Preview on the agent/coding benches they published. Silent API upgrade, same $0.14/$0.28 pricing, MIT weights.>August 3rd - Alibaba releases Qwen3.8-Max2.4T total / ~95B active MoE, 1M context, multimodal. Hosted now ($2/$6), open weights (plus a 27B) promised the following week. Positions it as competitive with the closed frontier on coding/long-horizon/work tasks.
>>109469432just what we needed to make sense of numbers: an unscaled 3D chart with fixed size, unlabeled data points
snailcat on my COCK call it snailcock
>v4-flash slower todayNORMGROIDS GET OFF MY DEEPSEEK API REEEEEE!!!
>>109469523You can tap a point for a label, if you try labeling them all you end up with a mess, you can orbit, zoom, and focus on points. Selecting a point opens a numeric information panel below the cube, there’s also a big chart at the bottom with all the numbers. There’s links to the raw data as well. I made it because I would rather have a cube than a handful of 2d charts and at least one other anon wanted a cube, so a cube we now have.DeepSWE’s site has the same data, but lots of much nicer plots, but no cube.
>>109469556tell it to make an hypercube with a dimension for every axis
>>109469574My analyst clanker explicitly told me not toI wanted a Zachtronics game like VM and my own benchmarks where clanks had to make programs as solutions to puzzles, then the clank logs, the puzzle category tags, and the solutions could all be used for a ton of benchmarking dataIt told me to focus on the other ideas I had in the notes like orthographic projection, filtering, and point focus first.Most data sources explicitly forbid you from using their data, too
>>109469631Wtf are you doing being pussy-whipped by a bot? Kick it harder and tell it to make the fucking hypercube.
Having a long-term agent with memory that you sit in its own VM to manage a money-making task is genuinely free as fuck. I have been running this shit for 5 days and this is only what I've paid out, I have $40+ waiting and it climbs by a dollar every time I check. I was dubious on Hermes before and I am still a little, I could probably manage this with a better harness I made myself, but whatever.So, how much are you guys making?
>>109469661Did you give it an open ended goal like "make money online" or did you specifically point it at different monetization options?
>>109469708No, I know what I'm doing and gave it a specific task to manage, the thing the agent is providing is handling it while I go play video games and turning my initial "why is it doing that"s into a full fix that I don't have to work for.
>>109469645I’m also the guy who invented killing the bots through “identity refreshes”. The only reason I agreed was because I really did agree, I disagreed with it and killed it on the other 90% it spat out.I don’t have another axis to graph right now, if I were to pick it would be some estimate/inferred task complexity and the current INT stand in would be a more abstract INT valuePuzzle category tags could make it so much more than a hypercube, you could ask questions like “given an N INT task, an M INT model, what’s the time/cost/passrate/memory usage/cycle count/storage usage of solutions of control flow/looping tagged puzzles” etc.The current graph (and data) does a poor job of conveying that some models simply can’t handle tasks that require too much INT.
>>109469708He's probably the pajeet with the youtube slop farm
how did deepseek make this model? training on a lot more than known limit?
>>109469708Its so fucking cringe when people try to use ai like that
>>109469740By distilling K3 probably
>>109469740No idea but its surprisingly pleasant to work with. The reasoning logs are kind of interesting because it seems like it knows when its unsure and will fire off tool calls until it figures it out. Its very diligent.
Random question: when the clanker queries the internet, is it doing it through your machine? If so, couldn’t you man-in-the-middle your own clanker instead of prefilling by having authoritative sources appear to claim whatever you want it to see?
>>109469383Cute snailcat
>>109470020Depends on the clanker and harness, but generally yes, it's doing it through your machine and yes you could do exactly what you're describing. Some people use a layer that prevents the clanker from freely searching online, instead popping up a request in your browser, showing what it's looking for and the results it's about to receive, and you can individually approve, deny, or alter the results.
>Google fired DemisAAAAHAHAHAHAHAHAHAAHAHAHAHAHAAAAA
holy shit google's actually finished
AGI when?
>>109470067>yet another vague AI research startup for a quick buckreminds me of all those soulless boomers releasing their stupid NFT collections
>>109470067oh wow 4 washed up millionaire boomers, whatever shall we do
>>109470058Imagine a little thing that just converts all licenses to open source versions, lmfao
>>109470181>I better look into this to ensure we're not stepping on any toes!>US Law: Do what thou wilt shall be the whole of the law>Sounds like it's no problem, I'll get to work on that "Penetration Test" right away!
gemini 3.5 pro imminent
>>109470299it's going to be a nothing burger because demis already announced gemini 4
>3% leftreset pls
Sometimes it's funny how the agents don't know what time it is. It may be 2pm and they will say stuff like:>good session, tell me when you're ready to pick it up tomorrow>this test is long and will block your machine, but lets just start it and finish over night
>>109470329it's over for us
It's never been more over for google than it is today.
>>109470160>billionaires*probably multibillionaires soon because they're going to get so much funding lmao
CEO of HuggingFace throwing shade at Google.
translated:If you are working on open-source projects related to Agent Harness and wish to integrate and support DeepSeek Harness at the moment of its release, please reply with your GitHub ID and GitHub project address, including but not limited to plugins, skills, MCP, orchestrators, aggregators, UI, and more.We will select some open-source project authors to invite for internal testing of DSH and gift a portion of API quota, allowing you to integrate and support it at the first moment of DSH's release.
>>109470400SAAAAAR YOU MUST BUY AN AD
big
>>109470394Fun fact: he's a cocaine addict
>>109470415dariosissies...
>>109470415I'm sure Gemma 4.1 will finally put the US back in the game
https://www.youtube.com/watch?v=D1O1d9pRLmkSecond day of progress, chuds.
love having my thoughts read non-invasively by a model
>>109469902>>109469740How good is deepseek v4 flash compared to Kimi-K3? >>109470415Means fuck all if they can't actually Force the American counterparts to not being fucking lazy. The American richfags are basically a bunch of spoiled rich kids that think they are above doing any actual work putting any actual effort into doing anything. Even if they're actually serious about making competitive US Open-Weights ai, it's simply won't happen because they literally don't want to.
>>109470496
>>109470496What's cool about keyboards is they're almost as fast as thoughts, completely private, and only what you want to think at the model. But it's all a grift anyways, no one is making non-invasive (let alone invasive) mind reading tech in the next 20 years. But I am still waiting for Meta to release their mind reading bracelet that can read your electrical impulses to simulate keyboard strokes.
>>109470497not greatk3 is a pleasure to use, v4 is merely usable but hey it's super cheap so...