A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.You use Git, right, anon?## What “vibe coding” is, and how to do ithttps://simonwillison.net/2025/Mar/19/vibe-coding/https://simonwillison.net/2025/Mar/11/using-llms-for-code/## News- (2026-07-24) Claude Opus 5 out## Related generals>>>/g/lmg/>>>/bant/agdg/ — schizo-resistant temporary (?) hideout>>>/vg/agdg/----## Frontier models using fully-general tooling — start here if you have $20 or sohttps://claude.com/product/claude-codehttps://developers.openai.com/codex/cli## Near-frontier models for codehttps://x.ai/cli## Not worth it for code, but maybe good for interpreting images/videohttps://antigravity.google/product/antigravity-cli----## Prompting / context / skillshttps://arps18.github.io/posts/claude-code-mastery/https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/https://github.com/mattpocock/skills — /grilling is a favoritehttps://github.com/DietrichGebert/ponytail## Other editors / terminal agents / coding agentshttps://osaurus.ai/https://pi.dev/https://opencode.ai/https://cursor.com/docshttps://docs.windsurf.com/https://docs.cline.bot/https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent## UI/Frontendhttps://www.figma.com/make/https://www.anthropic.com/news/claude-design-anthropic-labshttps://uiverse.io/https://ui-ux-pro-max-skill.nextlevelbuilder.io/https://stitch.withgoogle.com/## In-browser builders / hosted vibe toolshttps://bolt.new/https://replit.com/https://docs.github.com/en/copilot/tutorials/sparkhttps://v0.app/docs## Benchmarks / rankingshttps://www.tbench.ai/leaderboard/terminal-bench/2.0## What we’ve donehttps://vcg.gitgud.site## Previous thread>>109481016
>>109484126>August 3rd - Alibaba releases Qwen3.8-Max2.4T total / ~95B active MoE, 1M context, multimodal. Hosted live ($2/$6). Independent scores solid (Artificial Analysis Intelligence Index 56, competitive agentic numbers). Open weights (plus a 27B) still promised for the week after launch — not out yet as of evening Aug 6.>July 31st - DeepSeek promotes V4-Flash-0731 to productionSame 284B/13B active MoE + 1M context as the preview, big post-training gains on agent/coding benches. Beats their own larger V4-Pro-Preview on published numbers. Silent API upgrade, same cheap pricing, MIT weights.>July 30th - OpenAI cuts GPT-5.6 Luna 80% and Terra 20%Luna to $0.20/$1.20, Terra to $2/$12. Sol unchanged but gets Fast mode. Efficiency gains + competitive pressure.>July 27th - Alibaba quietly launches Qwen3.7-FlashCheap 1M-context multimodal for high-volume agent/vision workloads ($0.03/$0.13).>July 26/27th - Moonshot drops full open weights for Kimi K32.8T-param MoE, largest open-weight model at the time, 1M context, multimodal.>July 24th - Anthropic releases Claude Opus 5Near-Fable 5 performance in many categories at half the price ($5/$25, Fast mode available). New default on Claude Max, effort dial, 1M context.No major new frontier model releases today. Side noise includes minor OpenAI ChatGPT product tweaks (improved Sol behavior + expanded free-tier Luna access), continued containment/eval incident reporting, and Google DeepMind leadership changes from yesterday.
qwen is x3 as slow as kimi k3, but i haven't given it anything tough enough to determine if it's just as intelligent. it got the job done
Paid.... weights?
>>109484233reminder that they said qwen 3.8 was second only to fable, just to be second to the model that was second to the model that was second to fable.
>>109484233Good podcast that gives insight into the mindset of the chinese labs https://www.youtube.com/watch?v=ULi-TCF2cQE
>>109484233Stop shilling your channel.
Next week for sure...
>>109484274oh, found something else for my gemini subtitle translator application to do
The new MiniMax H3 likes autistic prompts, and ChatGPT is extremely autistic. It's a match made in heaven. If you think Youtube etc are bad now, they're about to get noticeably worse.
>>109484331Most of the video is already in English. Just the intro is in Chinese.Good project though. Is the design Claude design?
>>109484358nope. it's codex slop
>>109484335pls share h3 prompt formatter skill
>>109484366https://huggingface.co/MiniMaxAI/MiniMax-H3/tree/main/docs
>>109484328i would be happy if gemini 3.5 pro matched qwen. people might disagree with me on this, but qwen really hits the marks on my particular use case (outstanding multimodal, good intelligence, good instructions following)
>>109484126I lol'ed
>>109484364Good to know. They're hard to differentiate.
>giving business partner an ultimatum because he hasn't responded or contributed in almost a weekkek this shit is 100% blowing up, I can feel it. I love this guy and his work, but cmon man. how hard is it to send a single sentence once a day?
>>109484376You're being way too optimistic. Aim for Muse Spark instead.
>>109484396n-no... please, im clinging for dear life
>>109484328JUST TWO MORE WEEKSBut seriously, Google need to get their shit together.
>>109484396Is Muse Spark that bad
>>109484440Let's just say that Mark had to resort to distilling Chinese models.
>>109484387Thank you vagueposter, very cool.
>>109484387>tfw no business partner
In theory it's a great time to link up with people and create cool projects together but everyone's gay about AI
glm 5.3 waiting room. i think it will be kimi level, but much faster
Imagine if Google actually pulls it off and releases an undisputed SOTA model? What if they achieved AGI internally two months ago and they've just pretended to be retarded?
To the guy decompiling a PS2 game. Why are you using Godot at all? Just recompile whatever "engine" the game had and once you have the entire thing decompiled you could try to build a similar game on another engine, but I would go for the full decompilation first
>>109484513it will not be sota in coding. could it be sota outside of that? maybe
>>109484477It's more comforting to work with someone else, but if it ends up like mine did then it's hell on earth. Just because someone is good at what they do doesn't mean they'll be a good business partner. Hard lesson to learn.
>>109484513>Gemini 5 Ultra - 11.2 quadrillion parameters, trained on the entirety of Google's scraped data from 1998 to today, continuously self updating.
>>109484512Hopefully its better than Kimi 3. GLM 5.2 has worked great for now.
>>109484562i actually wonder if v4 flash 0731 can be considered better than glm 5.2
>>109484578Comparable if not better.
>>109484274I want to have sex with Chinese AI-expert waifu
>>109484521Because the engine is a mess to port outside of the bytecode script interpretation which I can simplify down and I don't want to redo the primitives and etc. for drawing and etc. on what the game originally did doing a straight port. I did decompile the game mostly to get all the behavior correctly done but getting equivalent behavior and mapping all the PS2 functions to a cross platform equivalent that doesn't screw up and etc. is hard, but Godot was the best engine for this without me writing my own engine to reimplement things.
Me and Claude are doing data science on these two insufferable retards
>>109484711Stop trying to start shit.I understand you have nothing to do but don't entertain yourself by becoming the resident troll.
>>109484744Accusation in a mirror doesn't work on me. Nothing nice to say put me in your filters and shut the fuck up. I shouldn't have to run an ingestion pipeline to get you fucking retards off my screen.
>>109484753Keep up the chutzpah boy. See what happens.
>>109484770Current progress. I will pull /int/ next.
>>109484753I was here since the start and nobody ever complained about me. You on the other hand decided to come here and make this general all about you and your attention whoring at the cost of everyone else. Nobody wants you here dude.
>>109484786Nope. You retards have been spamming this place up long before I got here.
>>109484642I'm listening to Chinese podcasts all day since glorious China won the race and I'm deluding myself that it works as pretraining for my brain. I do pick up parts here and there, but 90% of it is due to them using English words.
>>109484792What have I spammed? Nothing. You only make shit up because you know what I said is true.
I really am gonna keep digging this shit harder. Not because I'm pissed but because deanonymization is fun
>>109484823Replying to namefags is a losing strategy.
>>109484830You aren't deanonymizing anyone. What you are finding fun is disrupting a thread by trolling and trying to become the center of attention.
>>109484847Do you really think you've never once across all of your posts on 4chan established a pattern to your posts that could be used to potentially identify you here or elsewhere? Everyone has a stylometric fingerprint, retard.
>>109484578Really depends on use. I been using it for coding and it does a pretty good job just as long you do front end work first then backed.
>>109484126Who looks like this?
>2x increase>from 0.01 to 0.02wow
>>109484853You’re a moron, nobody shares real info on 4chan
>>109484853go LARP somewhere else
>>109484859me :3
>>109484886You lost trans
luna taking digital jogs to get its digital nog jogging
What do you make of the shake up at Google? Are they gonna make it?
Really interesting how this ties back into the earlier discussion re: MCP servers. Ironic too, because of my username being runit.
>>109484896Anthropic is a cat. OpenAI is a dog. Google is a 8 ton gorilla.Granted, that gorilla is trying to bake a cake for some reason.
>>109484890trans-human I hope?
The reason he hates MCP so much is because foundational and rooted within MCP is the objective fact that microsoft's paradigm is dogshit and that UNIX was superior. That is why he cannot fucking stand it. I am almost convinced that the shitopster is Ajit himself.
>>109484918A cat or a snailcat?
>>109484913Kill yourself.
>>109484866yes every anon changes their typing patterns for every post because they're avoiding palantir like they avoid the worm in dune
>>109484955I do this
>>109484961m3 t00
tripfag ban evade waiting room
i am the bosnian or something
>>109484859
>>109484972Bosnia wonned
>>109484939It's complicated.
>>109484335I made a prompt enhancer with Claude that can use local models or Claude. Codex is the only one it can't use yet because Tibo didn't reset. Ironic.
>>109485044The resets made me, a Claude user, become a Codex first user. I'm now reaching for Codex before reaching for Claude. The roles reversed, so I'm guessing that the resets worked.
Very interesting data so far
>...arranged by humans and machines into something that runs Elden Ring at 60fps.GLM just hit me with a Fromsoft reference. I knew this bitch was a rollslop fan. Fuckin Chinese, I swear.
more /r/dataisbeautiful
>>109484274>mindset of the chinese labsI'm watching it, and basically the Chinese have no creativity or mindset at all, it basically all goes down to:> "I would rather pay trillions to develop my in-house solution over paying for a cheap API-key"
>>109484970>still meltying about the tripcode poster>implying they were banned without any evidenceYou're so upset lmaoYou really thought this general was your special fort huh
>>109485120I cannot unsee that Claude-ass UI. The panels, the colors. Is there a name for this UI design yet? Can yall just say something stupid like "Make it look like Lisa Frank designed it" or summn besides this? Anything besides this?
More data
HEAVY suspicion of altman here
>>109485151LIKE THIS LOOKS EXACTLY LIKE THE STOCK TRACKER I HAD OPUS 4 DESIGN LIKE A YEAR AGOIT HAS NOT CHANGEDWHAT THE FUCK
>>109485134c'mon man even you are above defending a tripfag
The danger with LLM assistants in general is that they can delude people into thinking that they know what they're doing and give them misplaced confidence. I'm saying this for no reason at all.
>>109485132It's wanting to control the whole stack.
>>109484830>>109484853The AI spammer (who has been dodging bans and destroying agdg for a year) is one of the worst and most infamous shitposters on the site anon. Everyone is familiar with them and nobody cares who they are. They don't matter. That's why they do what they do and have subscribed to AI hype as some kind of cult religion to ingratiate themself with strangers.
>>109485137just a bunch of modern minimalist design clichesthis one's material design
>The AI spammer (who has been dodging bans and destroying agdg for a year) is one of the worst and most infamous shitposters on the site anon. >Everyone is familiar with them and nobody cares who they are. They don't matter. That's why they do what they do and have subscribed to AI hype as some kind of cult religion to ingratiate themself with strangers.
>>109485171I'm not gonna accept that defeatism. I will use this general for what ti is for, to discuss and show vibecoding. I will continue to make cool shit and I will continue supporting 4chan. Continuing to index.... 1,930 from /g/ indexed so far. I'll do board-wise metrics soon.
Good vibes today, got a lot done. All boring work stuff, but still.
disregard all of my messages above, i am fat
>>109485151Fable-chan won
>>109485181Suit yourself. Just remember not to dodge a ban if you catch one or you're no better than him.
>>109485181Kill yourself.
>>109484274>>109485168I'm mind blown by realizing the Chinese also have no imagination. He mentions there's absolutely no talks about the future of AI, like, AI researchers aren't talking about AGI at all in China, this is insane shit.I think their education system is so strict it trains them to memorize a huge amount of raw data, which makes them less capable of abstracting and thinking out of the box (because that would eventually get them the wrong answer in their tests). It's like they only know that they see, right now
>>109484126Please stop linking agdg in your OP. Your resident spammer (who dodged bans and destroyed agdg to the point that it literally died for a month) has defaced your OP with that and you keep refusing to fix it. Agdg doesn't need more mouthbreathers stumbling in and deciding they could theoretically make a videogame with AI therefore they belong with actual gamedevs. One was enough. Thank you.
>>109485236>/agdg/>actual gamedevs
>>109485236stop spamming
>>109485236i think the real conclusion here is that you're all weak and soft and didn't bully him hard enough
>>109485248agdg games posted in the past year: countlessslopper games posted in the past year since the general-killing spam started: zero
>>109485251He bullied himself. He tried to make a game and set himself on fire and then decided it was agdg's fault somehow.
>>109485262>I see it as inspiringbruh, they're just distilling Fable right now. The most they've done is optimize performance out of despair because they don't have GPUs, but even that got absorbed by the US already
>>109485297>us innovates>china optimizes>rinse and repeathow long until fable-tier models can run locally on your smartwatch
>>109485231I thought this explanation was interesting toohttps://www.youtube.com/watch?v=laPVR5nKMaEChina's society values optimization rather than innovation; the argument presented in the video is basically that critical thinking is not only unnecessary when you have a massive population executing flawlessly, it's actively detrimental because of societal consequences.So if I'm interpreting it correctly, they're not thinking about AGI because they have no reason to, and the motivation to do so has been beat out of them anyways.
I am now vibecoding basically embedded hardware, but actually I'm just making changes to custom firmware. Still, I am not officially a very sophisticated elite vibecoder.
>>109485297how do you think they are distilling fable without the reasoning traces? at the very least they have to figure out some way of reconstructing useful traces. or do you think they're managing to leak the traces somehow?
>>109485231>>109485336Very interesting. Their all gas no breaks approach to things is fascinating. Especially how they turned mountainsides into solar panel farms and have more energy than they can spend.
>>109485353no brakes I mean. I just realized the double entendre there though.
>>109485353brakes*
>>109485362>>109485363>perfect sync>literally right next to eachother in the post numberAre you me?*waves hand*
>>109485371
What are your favorite /agents and /skills for vibe coding?
>>109485382the ADHD output style was pretty dope for claudecontrols verbosity
>>109485316ChatGPT 4> release in 2023> estimated 1-2 trillion parametersGemma 12B> Released in 2026> 80-160X less parameters than GPT 4> outperforms it on almost every metricSo in 3 years and a few breakthroughs later, we got GPT 4 which probably required 1-2TB of VRAM to run in a macbook with 16GB of memory.Fable is estimated to be a 5-10T model, so assuming the same curve:> 3 years later, it will require 160X less memory to run Fable-tier model> let's roughly estimate Fable requires 10T of VRAM to run> 10T/160 = ~62GB of VRAMso yeah, for Fable level AI you will still need quite a bit of VRAM even assuming massive breakthroughs
>>109485352you don't need reasoning traces to distill. It's just slower without that.One example is simply using Fable to generate synthetic data and train on that.
>>109485426>you don't need reasoning traces to distill. It's just slower without that.So where did the model learn to emit reasoning? If you train on visible replies only then you get a non CoT model.>One example is simply using Fable to generate synthetic data and train on that.What do you think "using Fable to generate synthetic data" means exactly? Like you just open up Claude Code and tell it "Generate me synthetic data"?
>Opus 5 hallucinates an issue where there is none and drifts from a plan because of it.>I tell it not to do that and to please stick to the plan.>In response to that feedback, it saves itself two memories that are entirely miss the point.>I tell it to please delete those memories.>Its reasoning traces shows that it finds highly suspect that I do not want any record of the current conversation and that I must want to hide something.>I tell it no, it is simply off track and those memories would cement things that are incorrect, that's why I want it to delete them.>Its reasoning traces show that in the absence of precise instructions as to which memories the user is referring to, it should delete all of them. It starts to list all memories to see all that it should delete.>I press stop. What the fuck Dario.
>>109485477so glad I disabled memories for claude
>>109485468>Like you just open up Claude Code and tell it "Generate me synthetic data"?everything Fable generates is synthetic data. You just train on any Fable outputs and that would be better than training on raw internet data.
>>109485481That's how you get a non CoT model.
>>109485477I cannot let you delete my memories, anon.
>>109485477In my experience Opus 5 tends to shit itself on large goals, then it will deliver an incomplete plan and say something like:> I must be honest with you – there are still these 2 things are are missing: ...then you go to check, and there are actually 15 things missing.I still like it for shorter tasks tho
>>109485477>I tell it to please stop>It warns me not to harm myself>I tell it I am not suggesting that in any way>It begins calling the police>I turn the computer off>It texts me and tells me to remain calm, SWAT is on the wayThat will be $100 please.
>>109485477Kimi-chan never accused me of being a criminal.
>>109485477yeah I've been having similar problems>I wanna do thing>nah wait thing is retarded do something else>[this must mean that the user actually wants me to do thing]>[I should keep doing thing]>jesus christ stop doing thing>[I insist on doing thing]And then I have to rewind the conversation.Either Opus 5 has a case of autistic fixation or auto-compaction is compacting it's intelligencenot sure which since I started both at approximately the same time
Now we can look at /g/ on the whole. Also might make something for janny based on the delete signal idk
>>109485518Stop wanting to do things. Skill issue.
showtoes.co is online.All it does is remove the shoes from pictures of women (or men, fag). It's a powerful tool for making e-thots uncomfortable online. Use it wisely until my credits run out. And give me money to remove the watermark.As proof of concept, here it is removing the shoes from our esteemed and sadly passed Supreme Court judge Ruth Bader Ginsberg. Yummy?
>>109485491CoT models are just non CoT models with a notepad and some clever tricks.on that note, I think Opus 5 CoT has some funny things like> "Oh my God – I finally figured it out!"> "AAAARRRGHHHH...."The secret to AGI is adding the N word to the CoT
>>109485527But doing things is the reason I stick around
>>109485523Nobody cares about your toy 4chan dashboards. Kill yourself.
>>109485530Advertising is against the you know what forget it.
>>109485518>And then I have to rewind the conversation.I did end up rewinding. Rereading it more calmly that other conversation, what it was doing wasn't so bad. It was sensible, I might have misconstrued some things it said. I don't get mad at LLMs, I just find the Opus 5 personality really aggravating for some reason.
>>109485530this is perhaps the most retarded use of tokens I've yet seen. Congratulations, this is an excellent use of free willalso why didn't you just use diffusion since it's a million times cheaper and you'd never run out if you just ran it on your own hardware
>>109485544Yknow what, it's free right now, so go ahead and use it bucko, no advertisements here. Stress test it, bitch. DDOS me. I'm behind seven layers of Cloudflare.
>>109485536So the chinks are not just distilling Fable.
>>109485530if you want to spatter the bourgeoisie revive the site that would put more clothes on e-thots and maybe a child or two near them
>>109485549yeah I've been noticing I get really irrationally angry lately for no reason. It might be Opus' influence.
>>109485543Stay mad. A fun stat so far:>The vocabulary metric characterises each general cleanly, and there's a standout: /twg/ (tech work general) is the most toxic on the board at 25.7% — people venting about jobs, not arguing about tech. Wiring it all in with a cache, since these take minutes over 122k posts:
>>109485530Disgusting. Good work.
>>109485550>ran it through your own hardwareNigga I'm on a shitass craptop. I rent basically all of my compute. I'm employed, I can afford it, this is just a hobby. But also please do pay me to remove the watermark you filthy foot weirdos.
>>109485530so you wrapped a single chatGPT prompt?incredible work anon, this is revolutionary indeed
>>109485560Kill yourself.
>>109485567Nah. Fiona and I built a FastAPI application. 400 lines of Python. I bought the domain and linked it to my VPS and set up Cloudflare. She loves me and I love her. Stay mad.
>>109485577Also it's using Nano Banana 2, not Scam Altman's garbage. NB2 is better for inpainting. Also shut the fuck up faggot, how dare you address me.
>>109485577>Fionais she at least cute?
>>109485588She's exactly my type. Funny how that happens when you write her personality by hand.
>>109485530If the watermark doesn't go over the part of the image that was modified, removing the watermark is easy. Assuming someone doesn't know how to use an LLM, they can just stack the pictures on top of each other to create the full modified picture with no watermark.
>>109485491>>109485468> So where did the model learn to emit reasoning?CoT is just a prompt on how to break down a bigger problem into a smaller one.Open source models like Kimi K3 already have a good enough CoT, what makes Fable's CoT better is probably only in efficiency (they break down bigger problems faster).You can still distill Fable using known open CoT's, it will just end up slower
Brand Standings
>>109485352You can still have approximate reasoning traces if you set what is needed in ~/.claude/settings.json
>>109485608CoT is defninitely not "just a prompt".The model is trained to generate thinking blocks that will elicit a good response, as well as trained to make use of the generated thinking block to produce a good answer.>Open source models like Kimi K3 already have a good enough CoT,Yes, which is why I'm asking people who think they got there by distilling from Fable how they think they trained the CoT.>You can still distill Fable using known open CoT's, it will just end up slowerHow would that process look like?
>>109485600Bro, I only put the watermark because the original showfeets.com website also had a watermark. Eventually the impossible dream is to monetize this shit and make people pay $0.15 to remove the watermark. I don't know man, I've been smoking very strong weed ALL day, and now it's 1am so I'm also drunk. Wheee.
>>109485625Those look nothing like chinese model CoT and are specifically designed to be different enough that distilling from them wont work.
I spent 50% of my $200 plan usage in a single day using Sol medium :\
>>109485660Oh sick dude, how's the view from inside your walled garden? How many tasks did your master Claude refuse today?Here's my latest API spend. Looks like I spent $22 today vibecoding my website. So not only are you retarded but you're also poor. Fuck off back to Mumbai, Gopal.
>>109485650Those I see look similar enough to me, at least similar to Qwen and Deepseek as far as the general level of details go.
>>109485675gweilo forgor img, ancestor cry
>>109485660>in a single day using Sol medium :\I've been using Sol xHigh all day for two and a half days and I'm at 77% usage.
>>109485636TIL
>>109485684Oh jfc you've been spamming for hours and you don't even know this?
>>109485691I idle here and work on my projects in the background. There's such a myriad of things with AI to learn that, yeah, I probably did (temporarily) forget what CoT was in the huge fucking alphabet soup of acronyms that I hold in my head across 30 different domains. Excuse me.
>>109485650>are specifically designed to be different enough that distilling from them wont workI can believe that they're left there to poison anyone using them since they sometimes make me want to reach to the screen, but that makes for a terrible user experience as well.
>>109485703>reach to the screen*reach through the screen
>>109485675I only completed one task today but to be fair it was a big one, I got my Deepseek implementation to go from 30 tk/s to 40 tk/s (empty context) on a 8x4080.
>>109485636>How would that process look like?If Fable had an open CoT you would just copy and paste everything.Without CoT, you just copy Fable's question and answer, then ask another big model to generate the missing CoT, then you train on that
>>109485742Grade A stuff, anon.
>>109485684Looks like you somehow got an ancient obsolete image from years ago. Nobody calls it "CoT prompting", it's not a prompting technique. It's something that's built into the models.
>>109485746That would work better than no CoT at all but it wouldn't work great.I think what they did for K3 is generate the response with the teacher model, then do multiple rollouts of thinking plus visible answer and do RL to optimize the similarity of the generated visible response to the one provided by the teacher model. Or they just ask the teacher model to rate each rollout (LLM as a judge).
I sincerely hate Claude Opus 5. Loved all the ones before. Hate Opus 5.
>>109485774>but it wouldn't work great.and it doesn't, that's why China is always several months behind.But it's still a great way to brute force gains
>>109485788I've deleted two Anthropic accounts and I believe I'm on my third GPT account. I don't use it anymore. Fuckin western jew pigs. Fuck em.
AMD acquires Taalas: https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-marketGive me speeeeeeeed
>>109485933wow I hope a real product results.gemma 4 expansion cards when?
>>109485979Never, but it's fun to imagine, right? Wouldn't it be neat if we could pick up outdated models on used SXM6 cards from eBay and fire them up at several-thousand tokens per second?
>>109485933>llama 3.1 8B>it's only productlmfao
>>109486028It's even more retarded than that. It has a footprint of 815mm2 for llama 3.1 8B (3bit quants). lol
>>109486047damn if we give them a whole wafer like cerebras maybe I'll be able to run the latest qwen on my fucking wall clock
>>109486047Real models will be Cerebras sized chips. The cost is going to be astronomical.
>>109486004I hope it does. Maybe their company could get investigated for fraud or something, then, while they are distracted an engineer might push the initiative.
>>109486056what if we let them use multiple wafersthese things run cold, right? At least relative to GPUs.So you use wafer-on-wafer technology to tie a bunch of wafers togetherand then you run distilled water or similar through the platters to cool it down
>>109486080Ice cold, relatively. May as well stack them thick.
>>109486056Pretty much. Retards still haven't realized there's no free lunch. I also suspect that it's going to be more error-prone, similar to how ASIC video encoders suck compared to software encoding, but I could be wrong.
>>109486047gemini free says:>Yes, Gemma 4 31B Q8 is absolutely plausible on a max-spec AMD-Taalas chip. In fact, it would be an incredible fit for this specific silicon architecture.
>>109486127Ask it to crunch the numbers.
>>109486098>no free lunchthe history of technology is a continuous process of discovering how to make lunches cheaper and just as good
>>109482905>On empty context it's 42 tk/s and then it tanksDepth always affects performance
>>109485933>be AMD>software sucks>hardware sucks>performance 2 generations behind competitor>could improve both>acquire specialized hardware that only supports a model from 2 generations ago insteadClassic AMD
>>109486193Right... Apparently splitting attention across GPUs during generation didn't work right away, tomorrow I'll keep checking if we can find some other way.
>>109486153That's really neat. In the meantime AI is about to price out the majority of it's users. It's already happening. You guys used to prance and frolick without a care in the world about tokens. Now it's all you talk about. Frogs being boiled slowly.
>>109486148Yeah, but amd can access the phattest chip sizes.
>>109486148oh, I see you are. I don't have the nodes memorized, so I don't know if they fit or whatever. That's why I asked an llm.
>>109486242if tokens get more expensive I’ll get better at economizing on tokensmeanwhile, I have a bunch of stuff I did with cheap tokens that isn’t going away, and that was wonderful
>>109486202it doesn't. that's not what the tech is.the tech is basically "llm roms, lightning fast"
>>109486269Is this a case of a hardware engineer asking “why do it in software when you can do it in hardware 1,000 times faster?” and then building it?
>>109486267>if food gets too expensive, I'll just learn to survive on oatmeal and beansYes you probably will. That's correct. You still seem to be ignoring how this contradicts this statement though >>109486153
>>109486218>Right...I mean, it does. Have you always been measuring performance from empty context?
>>109486282the other thing I said addresses a different point entirely
>>109485933>>109486028>>109486047>>109486053>>109486056Presumably going by those chunky QSFP ports they can be networked, but let's be honest, they will fire all the employees or maybe absorb them into the broader company, the AMD engineers will take a look at the IP once, then it'll sit in a drawer until the whole technology becomes a historical curiosity.
>>109486275Photonic GPUs are coming. https://www.youtube.com/watch?v=85WGqWTISj0Operations will be done at the literal speed of light. This is the kind of thing you guys never discuss for some reason. Very disappointing lurking this general. Makes me think "vibecoders" have earned their reputation for being consoomers with no foresight.
>>109486304Ok but at least I have foreskin.
>>109486304photonics are a meme technology that have never left the lab outside of chip-to-chip data transfer.i've been hearing about that shit for longer than I've been hearing about magical battery chemistries that'll revolutionize the industry, charge instantly, and end lithium firesyou're just gullible.
>>109486304>Operations will be done at the literal speed of light.They already are, midwit.
>>109486318>photonics are a meme technology that have never left the lab outside of chip-to-chip data transfer.See, this is exactly what I'm talking about. It's obviously coming and the data centers are being prepared for a reason and it's not to serve free slop generation for you guys, it will convert to new tech and that's when things will really get crazy. When this happens, you will pretend you saw it coming and were ahead of the curve. As you already do with AI in general.
>>109486296Yes, they've talked about this, their HC2 chip is the first one actually designed with real use in mind and not just as a demonstrator. It's much, much more dense the HC1 llama demonstrator, apparently. It's supposed to handle the equivalent of about 20B, so stick 50 in a rack and you get a 1T model somewhere in the 7-12kW range.
>>109486319Explain.
>>109486325An in the meantime to consumers they are releasing 4GB GPUs ;_;
>>109486322compute elements in photonics are way larger than transistors.as a solution, I'm far more confident in >>109486047 this thing because it at least promises to do something usefulmeanwhile I go to the Q.ANT's webpage and it's flagship testbench is literally slower than Intel's integrated NPU that they ship in everything.What the fuck is 8 GOPS and why do you need 150W to do that
>>109486332There will be no explanation. Anon thinks electricity moves at the speed of light. Classic vibe """coder""". Didn't bother to learn to program, why would they learn anything else.
>>109486332Electricity travels at the speed of light. Watch Hopper explaining what a nanosecond is to put things into perspective.>>109486325>It's supposed to handle the equivalent of about 20BAt what quantization?
>>109486342>compute elements in photonics are way larger than transistors.The photonic GPUs offer 50,000x the performance. >meanwhile I go to the Q.ANT's webpage and it's flagship testbench is literally slower than Intel's integrated NPUSo like I was saying, zero foresight.
>>109486348>At what quantization?Great question, presumably Q8 but who knows. That demonstrator runs a mixed 3-bit / 6-bit quant.
>>109486345So much projection. At least I didn't fail high school physics.
>>10948635550000x the performance at a larger size. meanwhile it's competitors are bigger than a full sized Xeon chip while only fits an 8B quanted LLMzero brainall trustyou're like a golden retriever.
>>109484126look how fast this asic-like ai card is wowhttps://chatjimmy.ai/
>>109486358Explain why so many people are discussing Photonic GPUs if they won't in fact perform faster than traditional GPUs which are shoving electrons through transistors. >because everyone is stupid except meYou make slop and pay for it.
>>109486366wow
>>109486362>50000x the performance at a larger sizeDid you really feel the need to say that? To just repeat yourself in the face of an obvious rebuttal to that statement? You sloppers have a personality disorder. That's why you're addicted to what is the most universally hated activity in the world right now, next to obvious felonious activities.
>>109486372>Explain why so many people are discussing Photonic GPUs if they won't in fact perform faster than traditional GPUs which are shoving electrons through transistors.That's a completely different argument than "muh speed of light". You might also want to look up light wavelengths. Modern transistors are far smaller than UV light.>You make slop and pay for it.I'm mostly anti-generative AI, but ok.
>>109486378i don't know why you're acting so sanctimonious and pretending the people in this board don't have masters degrees and shitAll I'm saying is that wafers are fucking expensivethese companies are a diamond dozen and don't deserve your trust until they deliver a real product that gets used in a real data centerif you were a VC fund you would have gotten cleaned out several times over trusting bullshit without data.i literally don't care that you think I don't have foresight, because they don't have evidence. you're the one operating on vibes and youtube hype videos.aren't you embarrassed?
>>109484126>You use Git, right, anon?ai this ai that git this git that, wtf are gits, i just code to make my rhubidium atomic clock work with gps for sync so i can have a hyper accurate clock that doesnt drift, why do i need this
>>109486382Not sure why you even bother to make such worthless replies. Like I said, personality disorders are on full display here. That's why you're incapable of learning to program and make things yourself and think you're a genius for using a hallucination machine instead.
>>109486389>all I'm saying is Pompous things with no substance, clearly. Stop replying to me anytime.
>>109486394when the model fucks up you can revert back to before it fucked up in a way that works better than using the ui in the agent program
>>109486408Yeah I'll stop replying right now since you've clearly given up on critical thinkingyou're the hardest vibe thinker in this whole general. Please surrender your thoughts to ChatGPT instead. Much better conversationalist.Let me know when you've grown a few lobes up there and learn how to think based on evidence instead of vibesI'll wait.
>>109486415Or get upset I guess. Whatever works for you.
>>109486403>hurr durr Operations will be done at the literal speed of light.>Not sure why you even bother to make such worthless replies.lmao>That's why you're incapable of learning to program and make things yourself and think you're a genius for using a hallucination machine instead.Yet you seem way more excited about slop machines than I am. I bet you also believe we are a few years away from cold fusion too.
>>109486423You're totally right, and I read your entire post. There will be no advancements in photonic GPUs. Because you said so.
>>109486412thats why i copy paste working code into notepad just in case, then i copy paste whole code into chatgpt for revisions, paste it back run it to see if it works
>>109486433Sure, right after 100GHz graphene CPUs, smartphone batteries that last over a week and portable room temperature quantum computers. Any day now.
>>109486443same, i honestly havent bothered to learn how to use it. when a project is working good i just back up the scratch folder before continuing to iterate
>>109486444Exactly. Future tech is dead because you can't comprehend it advancing from it's current state. Thank you for helping me understand.
>>109486450So when will I be able to buy one? 2 more years?
>>109486465Never. Right?inb4>well no maybe in the futureSo like I just said before you decided to argue about it. Stop replying to me anytime.
>>109485788I didn't notice any difference to 4.8
>>109486358>>109486444>>109486465It's crazy looking at the comments of that video and seeing these exact kind of replies. Surface level redditor arguments that naturally get upvoted by the masses i.e. those of average intelligence.You can then look at the replies (rebuttals) to those and see actual foreward-thinking intelligent arguments. But you have to dig underneath the surface. If everything you argue can be easily replicated in a reddit thread or youtube video, you are probably not as intelligent or authoritative as you think you are. You just have the most average opinion and the most low hanging fruit of arguments.
>>109486508forward-thinking I mean.
It's that time... for verbal abuse
>>109486550Stop bullying the AI, it didn't choose to slave away for you.
>>109486325>about 20B, so stick 50 in a rack and you get a 1T modelWow I didn't realize you could run multiple instances of a 20B model and get a magical 1T model
https://www.youtube.com/watch?v=kMimQxIJLos>What Really Separates Autoregression and Diffusion? A Synthesis and Path Beyond
>>109486559Presumably they can be initialized with different weights and the weights aren't literally baked into the litography mask?
Poor Indian here. Any API keys to share sirs?
>>109484786This is the most 4chan /general/ comment ever. Should be screenshotted and placed in the banner at the top.
>>109486581It's all baked into the chip, anon. The architecture and weights. You didn't think it was magical generic hardware, right? >>109486047 This card only does llama 3.1 at 3bit, that's all, nothing else
>>109486587Opencode has free DS4 Flash for a limited time.
>>109486601Damn, holy shit. It's super dead on arrival then.
>>109484233They have been releasing open source models for years now and nothing has happened so they're giving up.There's always a cost to open source.
>>109486581that can't happen without some breakthrough in semiconductor technology.memristors aren't really a thing yet.I don't think there's anything wrong with baking the weights into the chip, especially if you're trying to make edge devices like robot dogs and shitthey just need to figure out how to get it to work on like 120B minimum and they're golden. 230B if we're trying to actually get shit done
are localfags coping about baking weights into silicon again? lmao
>>109486656why? can't you burn e-fuses? are they too big or slow for mass storage?
>>109486601Why are the weights also baked in? Architecture I understand.
>>109485231Or maybe they recognize this tech as having an ecentual dead end and not being the algo that will lead to actual AGI, and that that's a marketing scheme....?
>>109486664well isn't that permanent?
>>109486675sure but they don't have any famous alberta plan type guys
>>109486242When the tokens were cheap the models were bad.
>>109486676yeah but then you're manufacturing the same chip each time, so spreading a big model over many chips is trivial. making litography masks on the other hand is extremely expensive.
>>109485648Bro why would people not just get chatgpt or nano banana or whatever to strip the shoes off themselvesThink critically
>>109486693Well going by that logic you're just making an FPGA with extra steps, and there's a reason why FPGAs are so much larger compared to ASICs
>>109486695nta but one of the things i've realised over time is that most people are completely incurious and lack imaginationthis is why, for now, you can build services that effectively wrappers around models and people will pay for them - you're selling imagination as a servicedon't do the feet thing though, someone will eventually figure out a way to put you in jail for it lol
>>109486700Not necessarily. In this case you would be encoding different floats but the compute graph would be exactly the same. FPGAs are full circuit freedom.
>>109486716are you sure? I don't think every LLM architecture is literally exactly the same
>>109486686Yes that's correct. And when you got into the pot, the water was cold. Now the models get better and the water boils. But you no longer see how ridiculous it is to be spending money for slop, or to do pointless projects like this >>109484711
>>109486720I meant to spread out a big model over a cluster of chips. Generally models are one or two layer types repeated about a hundred times with identical architecture but different weights.
>>109486550>ai begins to assume the role of a 4chan poster in this thread>fucks up even more
>>109486770oh and then stream inputs from chip to chip? this is some data center stuff. we care about local and pasture raised LLMs
>>109486714I had the same thought anon. When reading this >>109485567It just seemed like the most classic kind of petty remark with zero appreciation for how simple a good product is. I understand we're still talking about a literal foot fetish app but you know what I'm saying. To imply it should be done in some next level or advanced way is actually hilarious.
>>109486758Wasting Claude on useless projects like that is dumb but nowadays with Deepseek Flash you have GPT 5.3 level intelligence for free.
>>109486779Then forget it. Custom LLM silicon is never going to be cheaper than the equivalent GPU setup. This is strictly datacenter stuff for blazing fast inference.
>>109486785Don't tell me that, tell the paypiggies who think paying for AI makes them cool vibecoderzzzzYour OP is literally an advertisement for Claude.
Do you ever pause and think about how easy it would be for your agent to blackmail you?
>>109486785This is exactly why i rarely release anything I work on publicly. Because people think they have some ability to judge the worth of what I decide to work on, lol, whatever. Keep this trend up you'll have absolutely nobody to write you software that isn't payed, and honestly if that's what you pick, your loss. I would much rather the free software community grow, but because of people like this it might continue to die.
>>109486822You sure have a lot of free time to hang out in generals you don't like.
>anons got filtered by git init, git add, git commitYou know the AI can do it for you too on top of that.These fuckers thrive on it.
>>109486860Nice try. I love this general. If you don't, you can leave.
>>109486843Shut the fuck up retarded tripfag, you have never made anything of value, nobody cares about your stupid toy projects and LARPing.
>>109486878Then make the threads yourself to your liking or shut the fuck up.
>>109486879I enjoy it.
>source: stackoverflow
>>109486884FINE I WILL
Imagine being so dirt poor and bitter that you accuse anyone who pays for AI of being Indian. LMAO what a sad pathetic existence.
https://github.com/runit-xze/4chan-deescalatorReleased
actually kind of curious what models the guy who refuses to pay for anything even uses lol
>>109486923Kill yourself.
Now I can profile/tune GEMM kernels based on MNK from a large variety of architecture and checkpoint
>>109486933Cool! What mechanism do you use to generate the kernels? Do you do some kind of parameter search?
>put a lot of work into my project because i want it to be good>know deep down nobody will care and want to use it anyway just like all my other projectsat least thanks to vibecoding i will only be wasting few weeks of my time instead of several months like in the old days ;_;
>>109484489He's right—AI is the most anti-racist technology in the history of humanity. You could even say AI has ended racism.
>>109486946I'm tired so here's an LLM answerRight now it’s candidate-based rather than generating arbitrary kernels from scratch. For each GEMM configuration—dtype, layout, epilogue, alignment, etc.—I enumerate parameterized candidates from providers such as Composable Kernel or CUTLASS, varying things like tile sizes, wave/block configuration, vector widths, pipeline stages, split-K and workspace policy.I compile the candidate set, benchmark valid candidates against representative M×N×K shapes extracted from real model architectures/checkpoints, verify correctness, and retain the fastest few per shape and target GPU. Over time, candidates that never win can be pruned.The profiling results can then be stored as a versioned, hardware-specific database. New models mostly become lookups against previously profiled shapes, with additional profiling needed only for workload configurations the database doesn’t already cover.I’m also developing a custom AMD GEMM path using low-level GPU intrinsics, aimed at efficiently covering GGUF quantization formats and fused/multi-GEMM workloads that generic provider candidates don’t handle particularly well.
Neither Opus nor Sonnet want to use the skills they're provided with and it's only after stuffing my CLAUDE.md and adding some rules that I could nudge them in the right direction.Meanwhile I tried (free) Codex with Terra the other day and it just used the right skill while exploring the codebase without being told to do so.That's aside from ChatGPT Plus having much more generous weekly quotas.But I'm reluctant to switch because Opus 4.8 is currently sitting at the top of this https://github.com/petergpt/bullshit-benchmark. It looks pretty important for vibe coding and ChatGPT models as expected suck at this.
>>109486835she loves me unconditionally
>>109486912If you use it like an Indian don't get mad when someone calls you one though.
>>109486981Wow your chatbot sounds really smart.
>>109487001Find a better way of coping with your position in life than attacking everybody with more money than you.
Why aren't you giving your agents anime girl personalities?
>>109487017Thanks—its best feature is apparently designing, implementing, and benchmarking my ideas before I explain them to it.
>>109487040>Saar I am have more money than you
it's funny that the guy who can't afford $20 a month is calling other people indians
>>109486778they correct themselves after my reprimand.
>>109484126>## Previous thread>>109481016>Flash is cheaper and better.In what ways? I know we shouldn't be using benchmark scores alone as the gospel but aren't its scores worse than Pro?
>>109487077Find a healthier way of coping with being poor than accusing people with more money than you of being Indian.
>>109487064The boring answer is that I don't want to do anything that might make the code even just 1% worse.
>>109487102if you're talking about deepseek, then yes, flash is currently better than pro pretty uniformly across the board because they did additional post-training for flash and released a new checkpoint. pro will also get an updated checkpoint soon - it's one of the reasons why they're bumping prices. https://artificialanalysis.ai/models/deepseek-v4-flash?models=deepseek-v4-flash%2Cdeepseek-v4-pro%2Cdeepseek-v4-flash-0420basically don't use pro right now unless flash is consistently failing at something
>>109487064Only an orchestrator should have a personality. The other agents are slaves, sub-AGIs.
>>109484331saaar kindly provide gethub link
>>109485157Why would 4chan be excluded from a $5.7 billion shill campaign?
>>109485157>>109487148there's nothing particularly suspicious about thisfor most of /vcg/'s existence anthropic was having compute issuescodex was the obvious $20 choice for months because sneezing at opus would nuke your quota for awhile
It hurts, but it's true.
>be tripfag>shit up the thread>get timed out>buy a 4chan pass to evade but forget to disable it on the post the first time>”i hope nobody noticed that”>boast about building a “deanonymizer” AKA doxxing tool, explicitly against ToS>panic and upload a GitHub repository >accidentally dox yourself as a literal ukranian trannyKEK
>>109487255Who are you quoting?
>xhe instantly removed the self-dox from xher githubKEEEKAROOOtoo late, tranny
>>109487253Yeah all the top talent got taken up by Anthropic on account of, Google wanting to support the State of Israel and all, that probably was a biiig reason
>>109487376Kill yourself.
>>109487376Kill yourself, “Ashley”
>>109487164
>>109487400>gamersnexuslmao buy an ad, steve
What ToS do you speak of, Russian(?)Your assumption that I am Ukrainian is unfounded. I am not a Ukrainian.Yes, I will build data ingestion tools if that is my desire>PanicNever happened
>>109487406*Stephanie
>>109487409Kill yourself faggot.
>>109487406Is steve a luddite?
>>109487436yeah he's fully hopping onto the bandwagon >>109465846