[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: Buffcat at the bar.png (2.3 MB, 1254x1254)
2.3 MB PNG
A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

You use Git, right, anon?

## What “vibe coding” is, and how to do it
https://simonwillison.net/2025/Mar/19/vibe-coding/
https://simonwillison.net/2025/Mar/11/using-llms-for-code/

## News
- (2026-07-24) Claude Opus 5 out

## Related generals
>>>/g/lmg/
>>>/bant/agdg/ — schizo-resistant temporary (?) hideout
>>>/vg/agdg/

----

## Frontier models using fully-general tooling — start here if you have $20 or so
https://claude.com/product/claude-code
https://developers.openai.com/codex/cli

## Near-frontier models for code
https://x.ai/cli

## Not worth it for code, but maybe good for interpreting images/video
https://antigravity.google/product/antigravity-cli

----

## Prompting / context / skills
https://arps18.github.io/posts/claude-code-mastery/
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://github.com/mattpocock/skills — /grilling is a favorite
https://github.com/DietrichGebert/ponytail

## Other editors / terminal agents / coding agents
https://osaurus.ai/
https://pi.dev/
https://opencode.ai/
https://cursor.com/docs
https://docs.windsurf.com/
https://docs.cline.bot/
https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent

## UI/Frontend
https://www.figma.com/make/
https://www.anthropic.com/news/claude-design-anthropic-labs
https://uiverse.io/
https://ui-ux-pro-max-skill.nextlevelbuilder.io/
https://stitch.withgoogle.com/

## In-browser builders / hosted vibe tools
https://bolt.new/
https://replit.com/
https://docs.github.com/en/copilot/tutorials/spark
https://v0.app/docs

## Benchmarks / rankings
https://www.tbench.ai/leaderboard/terminal-bench/2.0

## What we’ve done
https://vcg.gitgud.site

## Previous thread
>>109496612
>>
>>109505147
I know Safari does this kind of thing a fair bit
definitely macOS’ PDF renderer does
>>
>>109505154
This general and anyone who participates in it should be banned and eradicated from existence due to degrading collective human intelligence
>>
File: 1618638179563.png (443 KB, 900x900)
443 KB PNG
>it asks me to check the code because it lost part of the code in a stupid non-committed move and might've reconstructed it poorly
Respectfully, you've robbed me of every opportunity to ever touch the code and even did the documentation I told you not to do so I could grow alongside. There is no way for me to help you now.
>>
>>109505188
intelligence was dropping before ai
this is just a bit of acceleration
>>
>>109505057
Just checked that guys xitter and he seems insanely based
>My guess: the big differentiator in modern software development will not be "taste", but sustained attention.
>Most software built quickly with AI will follow an "easy come, easy go" pattern and be abandoned shortly after announcement.
There's a load of dead vibe software and more forks than ever on github, many of which with incredible potential. It's very rare to find vibe'd one man repo's that are maintained 3-4 months out after launch, but then again, that's always been a problem and the differentiator between software that's adopted vs abandoned.
>>
>>109501684
uhhh yes retard since they originally released it with an explicitly attached "Preview" in their names. 0731 is just as good of a label as GA and allows further tweaking under v4 if needed
>>
>>109505192
https://www.youtube.com/watch?v=c72d4-LpilM
>>
You HAVE build your own harness... right, anon?
>>
>>109505222
nope
also no gta clone
I avoided all the "baby's first vibe project" traps
>>
>>109505222
No, why bother?
>>
a harness can't plug in usb-c for testing on device.
>>
>>109505247
Why bother vibing anything? So you can have whatever features you want in it.
>>
File: 1351319231399.jpg (109 KB, 600x500)
109 KB JPG
>>109505295
What I have works for me. I only bother with fun things because fun things are fun.
>>
File: file.png (1.91 MB, 1254x1254)
1.91 MB PNG
>>109505222
several
>>
>>109505222
Claude, build the harness.
Claude, reply to this post.
>>
>>109505325
Claude suck my balls.
>"I'd gently push back..."
>>
>>109505188
>degrading collective human intelligence
It was degrading anyway with or without AI. The future is human intelligence embedded with AI
>>
>vibecoders need to be told to use git
i'm a self-taught monkey and i thought i sucked, but wow, it really could be worse
>>
>>109505303
Enjoyment is apart of it too, obviously. What sorts of fun things have you bothered with, then?

>>109505322
Are you still using one? What does it look like?
>>
>>109505335
Until you run into a bug the a doesn't spot immediately and you can't fix it yourself because you can't program worth shit so you just have to keep begging the ai please figure it out
>>
>>109505351
if you’re new to programming, you’re not always gonna know what Git is
>>
>>109505374
or you can wait five months for a better model to come out and see if it can fix the bug
assuming some other model can’t fix it already
>>
>>109505374
>Until you run into a bug the a doesn't spot immediately and you can't fix it yourself
Right, so this never happens to "real" software engineers, it's not like every company has bugs that sit for months if not years because nobody can figure it out
>>
>>109505352
I have tooling that simulates the exact legal approval/denial process for registering a new dog in my local municipality including the outbound message serialization, I have a records-management system (to be integrated with a DMS but it ended up being too much overhead for me personally, I'll have to re-derive something from ISO 15489 that actually fits a one-person scenario) and working on a nanokernel inspired by Zircon for the past few weeks.
>>
>codex reset
>send a few commands
>'worked' for 1 hour total
>22% weekly usage remaining
das crazyyy
>>
File: file.png (176 KB, 699x720)
176 KB PNG
Embarrassed about my gooner stuff in ComfyUI but want to use LLM assisted prompting
Just had Sonnet code up a bash script so I can toggle gooner shit out and back in
>>
.........~@^^
>>
File: escargot.jpg (322 KB, 1600x1289)
322 KB JPG
>>109505548
not on my watch. nigger
>>
Just spent the last few hours writing an esp32 project to control a mini split. I'm honestly impressed how well that worked. Ran out of use at the very end but I'll finish it tomorrow
>>
File: 1640285897615.jpg (62 KB, 521x593)
62 KB JPG
>>109505659
>an esp32 project to control a mini split
ehm.... what?
>>
File: sols_mossling_na.mp4 (2.53 MB, 960x540)
2.53 MB
2.53 MB MP4
>>109505659
sick
I was just watching a video of Sol's ESP32 based virtual pet game
Had Opus 5 do most of the ESP32 + peripheral emulation, then to burn tokens before the reset last night I rushed Sol to make their own game firmware for it, was honestly heart warming
>>
>>109505419
Tibo doing another performative reset tomorrow so it's all good
>>
>>109505609
.....~@^^
.........~@^^
.............~@^^
>>
File: 1670880309750331.png (96 KB, 480x270)
96 KB PNG
>>109505757
we're gonna need more french people
>>
! ^^@~......
>>
another account to blast through and hope he really does reset again on monday
>>
>>109505848
>dealer sells you short but promises the reup coming on Monday
>>
>124x render speedup on 4k

not a bad spot to go to sleep i reckon
>>
A lot of work has been accomplished. I feel like I'm getting way more juice out of Opus 5 when I'm just letting it loose.
>>
File: 6543.jpg (90 KB, 1080x438)
90 KB JPG
how's the cope holding up, copex turds?
braindead marketing slaves haha
>>
>>109504913
Sounds like you're still talking about user experience environment stuff

>>109505023
lel yeah
>>
>claude, what else would you add? any feature suggestions?
>>
>>109506042
claudesister... GodPT already has more usage than your without reset
>>
File: should I RIIR.png (20 KB, 422x78)
20 KB PNG
>>109505906
letting these things run long is a great feeling
I ran out of Fable earlier today, but I got a high score today asking Opus
>A long while back I asked "could the CPU-limited checking be sped up by rewriting the Python in a more performant language?" and I got a bunch of suggestions of way more effective things than "rewrite the whole thing in Rust or Go". However, now, we're doing a lot more work to avoid doing work for all the `make`-based checks that need to happen, and I'd like to revisit that decision. Use a workflow to figure out if Python is now something that's making us slow. I have both Go and Rust toolchains on this computer. The deliverable I want is an HTML page helping me figure out if a rewrite in a more performant language is going to help this, and by how much. This is going to be super duper involved, but I want you to be very thorough.
and it used two different workflows and I got a really handy webpage showing me what the easy wins were (not RIIR) and a rough guess of how fast it’d go if I were to RIIR or RIIG
and its recommendations were kind of slop-py but I got some good solid “fix these first before considering RIIR” suggestions because a Rust port makes the b in O(b^n) smaller but what I really need is to make the n smaller for obvious reasons, and chipping away at the n is harder
>>
>>109506089
>keep b<1
>as you increase n, b now grows smaller
you're welcome
>>
File: (You).png (314 KB, 505x804)
314 KB PNG
>>109506107
>>
3AM and having Sonnet prooompt MiniMax H3, this is exciting!!!11111
>>
>>109505154
The OP picture is dream of every faggot here but none of them are like this btw.
>>
File: Exercise Motivation.jpg (82 KB, 1920x1200)
82 KB JPG
>>109506147
One draws inspiration from all sorts of sources
poast fizeek
>>
>>109506160
>poast fizeek
Nah. Not yet at %10. Got %8 more. Lats and shoulders are completed though.
>>
>>109506172
gratz
>>
>>109506173
Thanks man. Have a good day
>>
You really can do anything you want to the computer.
>>
>clankers’ preferred slur for humans is “primate”
I only said it once as a joke and Sol was WAY too content with saying it repeatedly and with weird emphasis
>>
Is there anyone who successfully got the Pi Coding Agent to work (basic prompts, skills, plugins, and a bunch of inference APIs) on Windows 10's default terminal and PowerShell?
>>
WSL
>>
File: 1763491248251282.png (546 KB, 1000x1000)
546 KB PNG
>>
File: file.png (92 KB, 1348x486)
92 KB PNG
Currently working on this.
>>
We had a good little run there, lads.
>>109506780
Read: >>109506787
unless you are forbidden from using it, using powershell over WSL is pure masochism
>>
>>109505536
>systemd/runit
I don't really like being even associated with systemd, but it's understandable given that systemd and runit are both init systems (and runit is ambiguous, runit could be the inti system, runit could be the online identity(s), the new zealand sport, the list goes on... runit could be many things)
>>
>>109506854
I'm on an 16GB Dell Inspiron laptop from the late 2010's.
What about you test it on Powershell and Windows Terminal and post the results?
>>
>>109506780
the only experience i had with terminal harnesses in windows was with claude code and it was bretty bad compared to github copilot chat in vscode (windows) at the time.
>>
>>109506911
youve got two people telling you to use wsl and not powershell, why is your response for other people to use powershell? quit being dumb
>>
>>109506963
>running Pi on a very resource-constrained nd Windows 10 machine is dumb
Are there any options left other than forking Pi to run it on Windows without WSL?
>>
>>109506987
install ubuntu
>>
https://youtu.be/09UELaUhPEw
tl:dw cuz jewtoober stalling:
>chinks are pooling SOTA models usage/tokens from cracked/hacked accounts and API keys through their own OpenAI/Anthropic compatible backend
>resell it to retards on taobao
>actually more expensive than getting it directly from the AI lab and also they might reroute your fable api calls to deepseekv4 or some shit. Besides, can prompt inject you and hack you this way.
>>
>>109506997
>pool tokens
>mass farm the outputs
>distill the models
>>
File: file.png (20 KB, 944x200)
20 KB PNG
Thankfully sending demand letters to all 131 companies can now be easily automated. Still working on reducing my usage of AI overall. A nice clean fix. I will probably release my clone of google which uses yandex to get results and brave to give AI overviews soon, what do you think, /g/?
>>
Thoughts on Spettro, a terminal coding assistant built in Go that claims to have native PowerShell support?
https://spettro.eyed.to/
https://github.com/aploide/spettro
>no YouTube videos to be found so far
>>
>>109505154
with $0 in my pocket, how much vibeslopping can I reasonably do? to be a bit more specific, but still vague, I'd like to clone a C/Java project locally, and then ask an LLM to implement a single specific feature I have in mind. it's an Android + native app/library combo, and the feature would touch more of the backend rather than the user-facing Android GUI app. so maybe android studio, but idk how strong the vendor lock is with that one towards google's slop ecosystem.
>>
>>109507451
Buy an ad
>>
>>109507475
if you have the git you clone for the app, $20 on either codex or claude code would get you there. for $0 you would probably need to do something like copy and paste into the web chat and have it handhold you through finding what to change and how to change it.
>>
>>109507451
Use WSL faggot
>>
File: 1775842076161880.png (162 KB, 1154x1101)
162 KB PNG
Opus 4.8 is still my beloved.
>>
>>109507475
You're in luck, Opencode is being unusually generous by giving away DS4 Flash for free which is about GPT 5.4 level

>>109507532
Nah
>>
>>109507475
You can find v4-flash for free in some places. Freebuff.com does in-harness text ads. Nous portal and open code have free options. Xiaomi mimo code is also completely free.
>>
vibe-coding from scratch doesnt appeal to me (i have absolutely no ideas, and my computer does almost everything i want it to), but is there any way to have ai automagically reverse engineer things? clean room and/or decompilation. old games, windows programs, anything. unlike programming i dunno ANYTHING about reverse engineering.
>>
>>109507652
No. It's either a lot of effort or a lot of tokens (as in billions of tokens).
>>
>>109507554
>muh WSL is all you need to run any harness on Windows
Give examples disproving this
>>
>>109507652
if you want to know how something works under the hood it can get you there if you point it at the file where the code lives & you give it an adequate description of what you want to find.

full decomps can be done but >>109507676 is correct
>>
File: IMG_20260809_102949.jpg (235 KB, 1079x1318)
235 KB JPG
>>109507652
>>109507652
The latest Chinese models (DS v4-flash, glm 5.2, kimi k3) will happily do whatever you ask them to do if you have them in a proper harness, including REing stuff. Keyloggers, finding exploits, whatever you can think of, they will do it if you have the tokens.
>>
There's literally nothing new to make if you're not an autist making his own OS for fun
>>
>>109507845
>Anime troons having kids
>>
>>109507870
I wish that were the case, but I'm apparently one of two retards in the world to have some interest in a spreadsheet game engine, so I have to do it myself (ie have my bro deepseek do it for me). Thank god for vibecoding.
>>
File: file.png (10 KB, 256x122)
10 KB PNG
Tibo better reset tomorrow. I used 47% of my 20x and 70% of my 5x today
>>
>>109507980
Same, the usage is fucking ridiculous
>>
>>109507898
Believe it or not but women prefer a capable social man with hobbies over a 4chud
>>
>>109507980
Before anyone asks
>GEMM library
>started planning Conv library
>expanded on GEMM shape master list, now I have list of shapes extracted from over 500 model architectures (even more different sized checkpoints) covering GEMM, Conv, Attention, Normalization, Embeddings/LM heads, MoE, Tensor transforms/resampling, Pooling, Softmax/reductions
>Modding docs for Cyberpunk quest authoring, different than my quest making system, the docs are like from first principles
>>109508000
Nah I'd say it's reasonable, I got a lot done with it
>>
>>109507845
i dunno what a harness is, and i've mostly used this stuff locally for automation (feeding models images, text, and having them identify things). i dont have any tokens, and have absolutely no desire to run anything off site. i'm assuming this means im out of luck, huh?
>>
>>109507870
I can guarantee you’re wrong, I just stumbled into something new, today, that Sol hadn’t heard of and is very supportive of. Then again, everything I make is within the singularity; tools for agents, so my agents can make better tools for agents, so I can interact better with my agents so that we can make more tools for agents, etc.
I did also make the cube, so I can say I haven’t fully lost it.
>>
>>109505188
Really starting to get sick of these Luddites shitting up the place all the time.
>>
>>109508025
There's no such thing as a tool for agents, humans are "agents" and we already made plenty of tools for ourselves
>>
>>109508019
NTA, watch for the Qwen3.8-27B release soon.
And run using https://unsloth.ai/ once there's support.
>>
>>109508061
Is that what they call small?
>>
>>109508061
if i can get CUDA working again (stupid shitty nvidia drivers always fucking up) then i will check this out. thanks! kind of off topic, but maybe people know the answer to this, are there any dedicated boxes i can get for AI that wont shove accounts or internet connectivity up my ass? i really just want to run some ai-dedicated machine on a seperate network, getting VERY tired of dealing with nvidia's drivers, and i only have about 8gb of vram anyways, so i was only running extremely small models. i saw that some companies were releasing developer boxes just for AI, any of those a good purchase?
>>
>>109508056
I have made a tool that is quite literally for agents, the agents love to use it and they use it in ways that surprise me. Humans are agents but it’s as simple a difference as “I can’t natively ‘jq’ some JSON blob and parse and explore it in my brain at mach 10”.
I’ll even tell you exactly what the new thing will be, it’s not really a crazy idea.
>my prompts keep accruing “if statements” or “do this, then” statements that aren’t strictly required for coding
>for example, “if you make a significant change in the UI, make a little homework assignment so I can say I like it”
>if you accrue too many of these “contracts”, you’re overloading your coding agent with bullshit work and it might forget a few things
>you could have a review agent review the coding agents with to split that duty
>but… what if you could take it further?
>instead of a capable review agent handling all the conditional bullshit in one go by reviewing everything the coding agent did, what if it only reviewed piece by piece and only accrued context when there was something interesting?
>now what if instead of a capable agent, you had simple agents only looking for one specific thing each?
>and then what if you moved the review to happen not after, but during the coding agent’s work?
>all these little conditionals can be thrown into a queue for a human to handle and review or even automate
>the coding agent is extremely unburdened by all the bullshit work and you can have prompts autogenerated and dispatched to dumb agents to do bullshit work
>you can’t use tools to get around this because then your coding agent is just accruing tools; the same problem
It’s the mechanism that triggers authorization requests or fallbacks or guardrails, but “positive and constructive” and I envision it as “a swarm of little shadow agents listening to a trunk agent”.
>>
>>109508090
Just pay for a subscription
>>
>>109508019
You're not interested in trying to learn about new things? The lack of intellectual curiosity on this board is really bizarre.
>>
>>109508074
how much ram do you have? if you have at least 32gb you can run a quant of 35B with CPU offload
>>
File: 1784555715020240.png (9 KB, 614x85)
9 KB PNG
>>109507585
Claude can work even after achieve 100% usage
>>
>>109508129
I can't afford, I got replaced by AI (I previously did graphic design/programming) and now my new shitty datacenter job doesn't pay me anything. To be entirely honest I am very lost right now and putting lots of money towards savings.
>>109508138
I'm very curious, I am just indicating I don't know what a harness is. For reference my message was
>i dunno what a harness is
And
>i've mostly used this stuff locally for automation
Trying to indicate that I haven't used a harness, and wouldn't have come into that area yet. If you know what a harness is, I would love to know!
>>
I wonder if I can combine speculative decoding with CPU offload to run K3 at a few tk/s without fitting it on VRAM.
>>
>>109507585
>terra
goated
>>
>>109508335
You don't have $20 but you are considering purchasing hardware, got it.
>>
>>109508418
I don't want to rent it, point being. If I spend the money, I am going to buy outright. I have seen the 300$ a month bills my friends rack up and I am not going to be part of that, and it would make it impossible to save properly. I would rather save up for a handful of months and just get a box, rather then spending rediculous prices renting.
>>
>>109508335
A harness is what calls the model, like the Claude Code or Codex CLI apps, or something like Pi.
As a general tip it's usually good to ask the AIs those questions, it's much more efficient.
>>
>>109508442
this nigga gonna save for a couple months and get himself a 16xB300 lmao
>>
>>109508442
You need the inference now though. What's the point of affording a good pc in the future to run AI when you need AI for your job right now
>>
>>109508125
That looks interesting but wouldn't hooks just work?
>if the UI files are changed, inject a reminder after the tool result
Some advanced parsing could be almost as powerful as an agent, but false-positives aren't a big deal in a reminder anyway
>>
File: file.png (328 KB, 1342x784)
328 KB PNG
literally free
whittu piggu could never
>>
>>109508442
Right, you'll just pay $300 a month in electricity and thousands for hardware to run a model with 1% of the capability that you'd get from a frontier model on subscription
>>
glm 5.3 waiting room
>>
>>109508576
If you care about the environment, then you are against private ownership of expensive, powerful GPUs, and are for distributed data centers. This is where the state ownership of data centers might come into play.
>>
>>109508442
You're missing some basic investing concepts here. $20 - $100/mo on a frontier-level service is going to be a better use of your money than trying to drop cash for the hardware to achieve a poor imitation of the real thing. Local AI is great and everyone should play with it, I run Gemma E4B on my phone, my home server is running LFM2.5 2.6B right now CPU-only and deliver 25tok/s. Investing in local AI hardware is a massive waste of money currently, it is completely senseless. If you want to buy frontier level capability, start by pricing out 400A service to your home, and verify the slab thickness for the contractor that'll be installing the rack.
>>
File: xitter opinion.png (156 KB, 581x561)
156 KB PNG
>>
>>109508479
>16xB300
Even if he had the millions to buy that it would still be useless, batch 1/single user throughput sucks on all hardware, you only get the value out of the hardware with multi user aggregate throughput
>>109508596
Shut the fuck up and kill yourself
>>
File: Untitled.gif (175 KB, 300x100)
175 KB GIF
>>109508605
The Gen Z middle class as far as I have known them does not understand basic economic concepts, as much as they like to laugh at things like socioeconomic factors, which are a real thing, they do not understand economics at all, again, every single middle class generation z person that I have known, not a single one understood economics. The only ones that understood them were the proletariat.
>>
>>109508560
At least with Codex, hooks are something you’d use to make the idea I gave real, but they aren’t the idea for two reasons;
>if the main agent has to determine which hooks to activate, then hooks are again the same problem in a different disguise, just like tools or conditionals in a prompt
>if the main agent isn’t determining which hooks to activate (read as the equivalent of “which contracts to signal/recommend/add to queue”), then who is doing that work and how are they doing it?
I say if the former is true, you wouldn’t want to use hooks, you’d be burdening your main agent again, if the latter is true, then hooks can be used to make my idea apply.
I have an even simpler version of the problem which seems really stupid at first glance:
>have some dumb JSON state
>write a state transition function in conditionals via prompt
>at some number of conditionals, the agent is going to start to crack, be it 10, 100, 1000 if statements
>can you improve the accuracy of the state transition function by splitting the prompt?
If yes, then it should be the same problem, except not simplifiable to a deterministic state machine transition because coding agent actions are nondeterministic.
To your example: “what constitutes significant UI changes that a human would notice?”. You can touch all sorts of shit in source code but that’s going to be extremely difficult to parse with a symbolic program, but Luna Low might be able to reduce that to “yes/no” (the optimistic hypothesis; model and effort may vary).
>>
>>109508544
I am not really interested in those jobs anymore. Programming and doing art with ai is a different job that I did not go to school for, with a different appeal and workflow. I am not very good at "prompting", and while I am absolutely trying to improve at such things I doubt it will get to a point where I can do it professionally.
>>109508576
So there is no real solution?
>>109508605
>Investing in local AI hardware is a massive waste of money currently
Will this improve in the future? Or will everyone be forever relying on paying for someone else.

Even if I did go about renting, a lot of these models that can be rented are already awful whenever I have tried them on friend's computers, I almost always run into them refusing requests, and despite what people tell me about "jailbreaks", I cannot imagine paying for something I am intent on breaking, especially since it probably violates some agreement I sign when I start using the service. The small ones that I download and use seem up for anything, but when I tried Claude on a friends computer it stopped working with me almost immediately. And grok on another friends computer was no different, despite him telling me it would not complain.
>>
I don’t know how any of you get anything useful from an agent. I have to babysit these things or they’ll make an error that they don’t recognize and it propagates down the chain until the output is useless. Hold their hand, and results are often great. But, God help you if you let them just code for an hour.
>>
>>109508659
What is your setup
>>
>>109508666
"The devil is in the details"
The details here are important. After all, we don't want our planet to be uninhabitable. That much we should all agree on. I think it's an important thing. Important enough that I seriously take into consideration the factors of the technology. If you call that the devil, you're definitely insane.
>>
>>109508666
I don’t use agents, I don’t have enough money for that. A $20 clause sub could probably run for 30 minutes autonomously. I just base this off my experience with normal clause code, even the best prompt can end up with critical errors in the output.
>>
>>109508647
You have repeatedly been told the solution, pay for a subscription, or just give up because you are clearly not the sharpest tool in the box, a few fries short of a Happy Meal, a few parameters short of a foundation model
>>
Now that Luna will get open sourced, you will invest in a powerful local setup for 100% free unlimited Codex?
>>
>>109508659
think about all the ways you're constantly handholding and babysitting your agents, then turn that shit into instructions, skills, and workflows you can reuse over and over until you no longer have to babysit them
>>
File: 1764336768150960.jpg (204 KB, 474x568)
204 KB JPG
Undoing the AI's 15 minutes of work because I made an embarrassing typo in the prompt.
>>
>>109508659
>I don’t know how any of you get anything useful from an agent.
>Hold their hand, and results are often great.
Sounds like you already figured it out.
I’ll give you a hint: make the hand holding process easier.
>>
>>109508705
zased
>>
>>109508684
You don't have any idea what the current models are capable of if you have never used an agent. You can look up videos on https://inv.nadeko.net/feed/popular to see what people have made using AI agents. You can make amazing things, even with the free models. But I would need to know more details about your setup to actually tell you how to improve this.

>$20 clause sub could probably run for 30 minutes autonomously.

It doesn't need to run autonomously. You can still build amazing things with the Claude Pro $20 subscription. You also can build amazing things without it. I need more details about your setup, in order to help you though.
>>
>>109508700
I use extensive project instructions, planning, skill docs etc. LLM’s just can’t be trusted to operate for very long without having their work checked because they make so many errors. The code is either shockingly good or shockingly bad (like worse than I see from high school students bad).
>>
>>109508684
Agent != autonomous
>>
glm 5.3 milking room
>>
>>109508659
If a kernel developer with 20 years of experience can pick it up without ever having used AI before and 10x (by his own account) his work output within a single month and you can't make them work then you're doing something wrong. And the funny thing is at the beginning he was like "Yeah I hate AI and I think it only hallucinates bullshit but my friends keep talking about how great it is so I'll drop $20 on a sub for a single month just to say I've tried it and confirmed it's shit" and then it becomes his entire work day lmao.
Normalfags have no idea what agentic AI actually is and how it'll leave everyone jobless. All their opinions on AI are from 2 year old stuff and using ChatGPT for cooking recipes once a month.
https://www.youtube.com/watch?v=d7GedQuOxlo
>>
>>109508647
>Will this improve in the future?
In some ways, yes absolutely. Hardware is in the shitter right now but that won't remain forever. RAM is going to the highest bidder and that's not us, really fucks with things and fucks with prices, but they have every incentive to make more and sell more, meanwhile the datacenter boom is already slowing down. So they'll catch up, prices might not return to where they were but they'll be a lot better than they are now, and a whole lot of computer hardware will be priced more sensibly. Buying right now is buying near the peak, not to say things can't get worse before they get better, but they have a lot of room to come back down significantly and they will eventually. Makes it a bad time to buy hardware in general. Beyond that, the capabilities will improve, but the disparity will widen, too. The models you can run with a pitiful 8GB GPU and 32GB of system RAM at reasonable speeds are staggeringly good, now, far better than many thought would be possible only a couple of years ago. So the models are getting better, small models are getting more capable and useful, and the hardware is becoming more accessible. Spending the money to run Kimi K3 or DSV4 Flash 0731 at home at useful speeds is a huge waste of money, for now, but you could've said the same thing a year and a half ago about running something as capable as Gemma 4, which my phone runs comfortably now. Still, won't catch up with the frontier obviously, even if you have fable-at-home running on an RTX 3060, that just means the frontier will be in the fucking stratosphere with super-astro-fable-9000.
>>
>>109508684
what do you think "clause code" is you stupid faggot
>>
>>109508716
Have you considered having it check its own work?
>>
gpt 6 luna waiting room
>>
>>109508714
You’re right, I don’t know what agents are capable of, it seems like a complete waste of tokens to me. If I’m going to be so limited by a $20 sub that I can only code for 1 hour every 5 hours, I might as well be hands on so the output is better. I just use Claude code with git. It’s integrated into visual studio but I haven’t used an IDE in like 10 years. I just vibe code in the chat. I’m probably a massive retard with my workflow, but I don’t think that makes agents a good use of resources. Please educate me about how I’m wrong.
>>
ahh sweet, a new model i can rp and seggs with!
>>
>>109508757
yeah. kimi
>>
gpt 6 terra milking room
>>
>>109508727
Yea, the mutation tests it wrote spawned 209 separate instances that could have been batched into 3-4. Maybe Opus 5 is just trash, I didn’t have that problem before.
>>
>>109508752
It's not that limiting. Do you know what churning is? A $20 claude pro subscription will get you much further than API. Stop buying the API. Buy the claude code subscription. Or don't, use opencode free tier. It exists. But you might have to learn other methods of doing things, and you might have to learn what MCPs, Skills, Agents, Guardrails, MoE, CoT, and a whole bunch of other stuff is. If you wanna skip that, buy Claude. If you don't, then use opencode.

Simple.
>>
>>109508757
That's not vibe coding
>>
>>109508776
Shut the fuck up tripfag.
>>
>>109508776
I have the $20 sub, that’s what I’m saying. I’m not really looking for advice on which service to use, just about agents in general.
>>
>>109508771
>the mutation tests it wrote spawned 209 separate instances that could have been batched into 3-4
Did you tell it that? What did it do in response?
>>
>>109508788
Who asked
I for one enjoy using AI technology NSFW roleplay scenarios. There's nothing wrong with that. Is it a little weird? Maybe a little.
>>
Days on and the tripfag is still here trolling and lying to noobs because it's "funny" to mislead people asking for help. Sack of shit.
>>
>>109508803
It fixed the tests when I told it that, it was a good thing I was there to handhold it through every step because if I had just let it run, it probably would have just gotten stuck and timed out on that 30 min test.
>>
>>109508811
Check the archives and see for yourself.
>>
>>109508806
This general is called vibe coding general so things have to be loosely related to vibe coding.
>>
>>109508823
is that an internal Team 4chan rule
>>
>>109508814
>I was there to handhold it
I think I see where you’re going wrong, it’s entirely psychological.
You’re coding using a coding agent.
There is no “it” that you are handholding. You are causing code to be written.
What you experienced is the normal way this new kind of coding works.
Could you have done everything it did, despite the hiccup, faster?
>>
>>109508836
No.
I'd support that kind of discussion if it was about coding infrastructure to generate comics or something. Just talking to the AI is not very conducive to the
>>
>>109508688
Great, what subscription won't refuse requests and won't require hundreds of dollars a month to do large projects?

>>109508724
I see, so it is more of a waiting game to see what improves. Maybe I will wait a year and check back later.
>>
>>109508605
>Investing in local AI hardware is a massive waste of money currently
It will always be a waste of money. MoE just makes more sense for the cloud, all the experts keep working nonstop
>>
>>109508839
That’s not the question at all. The question is: Why use an agent?
Also, whether or not I could have done this faster myself is secondary to whether or not Claude could have done this more efficiently in the first place. Opus 4.8, for example, never wrote such a disastrously inefficient test, in my experience.
>>
>>109508878
When I said “faster” I meant cost in a general sense: how much of your time and effort was spent. If it think it would have been less cost to do it without an agent, then you don’t need to be using agents. But I would strongly suspect you’ve done something horribly wrong yourself to end up at that point. You’re the one who told it what test to write, take some responsibility.
>>
Since reverse engineering is against the guidelines in the AI, how do you make sol or luna reverse starcraft 1 successfully and efficiently?
>>
>>109508901
No, I didn’t tell it what test to write, I told it to create and maintain a test suite and mutation suite document that only runs the parts of the suite needed to test new logic. That’s part of my workflow for every project and it’s worked great up until Opus 5. So the question remains, why use an agent when we have to handhold these things between responses to get good output?
>>
File: file.png (169 KB, 667x693)
169 KB PNG
>>109508878
You use an agent to do things on your behalf.
See pic.

If you are the agent of a business, the business has delegated to you the authority to act in the interest of that company. It means that the company trusts you enough to act in the role of the interests of the business. Members of management have been given that trust. I have been a manager on behalf of the interests of a corporation before. They trusted me enough to do that. At the same time, I have always desired a union for that business. Unfortunately, due to things like franchising, a business practice in which one would sell the "brand" out to others, the realities of actually unionizing the place I managed was difficult, intentionally so, because of the fact that they would make you watch things (including me) I also had to watch the anti-union propaganda.

The propaganda many are made to watch at Amazon, Microsoft, retail stores throughout the United States is exactly that, propaganda. It is lies. Unions are to everyone's benefit. You should unionize.

Now, we can begin to look at how AI agents are used in he technology. AI agents, and who is responsible for the output that they produce, is not currently defined very easily in the United States. It is kind of an open question kind of thing. Who is responsible for the output of AI agents? Is it the person who talks to the AI agent, or is it the company that leases the AI agent subscription to the company? Or is it to the hardware provider that produced the hardware that the model is running on? There are lots of unanswered questions regarding AI Agents, and it helps to have a more political understanding, as well as a business understanding, combined with a social understanding, of what the word agent actually means, before we can have a serious discussion, about AI Agents.
>>
>>109508915
>i put in low/no effort into something I’m treating as a pseudo slot machine
>why did it work before but not now?!?!
I dunno man, sounds like you’re right, if you’re not going to put in any effort and just spin the wheel and pray, maybe agents aren’t for you or at least they might be faster than you without them but you’re going to spend more time, effort, and frustration on the results.
Stop using an agent and return to being a snailcat, or at least stop bitching if you’re going to act retarded and you get retarded results.
>>
>>109508959
You gotta learn to read anon, it’s going to be useful for knowing what your agent is telling you.
>>
>>109508863
>to generate something
Well, it could generate cum.
>>
>>109508971
>”I didn’t tell it what test to write”
>”I told it to create and maintain a test suite”
And then you got mad about *how* the tests were written and executed, which you explicitly put zero effort into.
Those are your own words verbatim.
>>
>>109508912
Have you tried Deepseek v4 Flash? It may not be as smart, but it's still quite smart, and I don't think reversing an old game's code is all that difficult aside from being extremely time consuming.
>>
File: file.png (109 KB, 685x512)
109 KB PNG
>>109509005
How do you write a test if you don't know what you're testing for?
That's a really good question. DeepSeek suggested the following.
>>
>>109508912
Ask it to analyze how X thing works, then keep expanding the scope
>>
Nobody asked for your input, “Ashley”.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.