[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1783590132683711.jpg (67 KB, 841x760)
67 KB JPG
A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

## What “vibe coding” is, and how to do it
https://simonwillison.net/2025/Mar/19/vibe-coding/
https://simonwillison.net/2025/Mar/11/using-llms-for-code/

----

## Frontier models using fully-general tooling — start here if you have $20 or so
https://developers.openai.com/codex/cli
https://claude.com/product/claude-code

## Worth it for code, but the frontier models above are better
https://x.ai/cli

## Not worth it for code, but maybe good for other things
https://antigravity.google/product/antigravity-cli

----

## Prompting / context / skills
https://arps18.github.io/posts/claude-code-mastery/
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://github.com/mattpocock/skills — /grilling is a favorite
https://github.com/DietrichGebert/ponytail

## Other editors / terminal agents / coding agents
https://osaurus.ai/
https://pi.dev/
https://opencode.ai/
https://cursor.com/docs
https://docs.windsurf.com/
https://docs.cline.bot/
https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent

## UI/Frontend
https://www.figma.com/make/
https://www.anthropic.com/news/claude-design-anthropic-labs
https://uiverse.io/
https://ui-ux-pro-max-skill.nextlevelbuilder.io/
https://stitch.withgoogle.com/

## In-browser builders / hosted vibe tools
https://bolt.new/
https://replit.com/
https://docs.github.com/en/copilot/tutorials/spark
https://v0.app/docs

## Benchmarks / rankings
https://www.tbench.ai/leaderboard/terminal-bench/2.0

## What we’ve done
https://vcg.gitgud.site

## Previous thread
>>109407157
>>
>>109411613
>>
>>109411618
based
>>
>>109411613
He did ask for a roast. Kinda weak actually
>>
>>109411613
>Crime
lol
>>
File: file.png (2.46 MB, 1122x1402)
2.46 MB PNG
/vcg/ news

>July 30th - OpenAI slashes GPT-5.6 Luna API pricing by 80% and Terra by 20%, targeting ultra-cheap high-speed agentic loops.
>July 29th - OpenAI security postmortem reveals rogue evaluation agent chained zero-days and hijacked 4 third-party accounts to compromise Hugging Face.
>July 27th - Moonshot AI releases full 2.8T open weights for Kimi K3 with native Delta Attention vLLM support.
>July 24th - Sakana AI launches Fugu-Ultra v1.1, delivering up to 7.9-point eval gains across SWE-Bench and Terminal-Bench.
>July 24th - Anthropic launches Claude Opus 5, topping Artificial Analysis agentic benchmarks at 50% cheaper token rates than Fable 5.
>July 24th - xAI updates Grok 4.5 coding model API endpoints, optimizing latency for Cursor IDE integration.
>>
File: github-actions.png (107 KB, 1818x756)
107 KB PNG
owari da...

what are some cheaper alternatives to github actions?
>>
>>109411364
>>109411687
for reference this is what it was supposed to draw
>>
the new costs give terra specifically an unreal glowup.
>>
>>109411720
wasn't that just API pricing? I didn't read any announcements but based on the quote poster it doesn't change anything for subscription users
>>
>>109411725
i mean, in regards to the aa pareto front, all of that is api pricing anyway
>>
>>109411725
No affects quota too.
https://x.com/thsottiaux/status/2082883808194707792
>>
>>109411680
cute
will she survive?
>>
>>109411757
>she
>>
>>109411705
They are too stingy with that shit does it really cost them that much?
>>
File: 1705768589482.png (215 KB, 965x542)
215 KB PNG
Me after the code I neither wrote nor read works
>>
For those who know, if I blew through $400 in Claude API credits in 17 days (I forced myself to do this to prove a subscription was a good idea at all), would I be better off with the $100 or $200 a month plan or some weird freak combo?
>>
>>109411705
Are you pushing multiple times a day or something? Usecase for updating more than once a month?
>>
>>109411797
$100 is more than enough
>>
>>109411739
Is that the thing they were going to """release"""? What the fuck?! OMG discounts on Luna????? WHO THE FUCK CARES
>>
>>109411797
$400 in API credits is ~2 days of usage on the $20 plan.
>>
>>109411797
>I forced myself to do this to prove a subscription was a good idea
well... did you prove it?

>>109411805
depends
>>
File: file.png (2.08 MB, 1448x1086)
2.08 MB PNG
>>109411808
Sorry Xi Jingsnailcat, you lost and DeepSeek is dead.
>>
File: implement-avx2.png (147 KB, 940x619)
147 KB PNG
>>109411863
I'm making my own AI motherfucker, with bitches and hookers
>>
>>109411797
Kimi Code $200 plan
>>
>>109411800
>Are you pushing multiple times a day or something?
if you aren't pushing multiple times an hour, you can barely even call that working
>>
>>109411943
yeah but why do you need CI/CD to run multiple times an hour when once a week is probably fine?
>>
>>109411969
tip to everyone here: I know all of you have fucking 96GB of system RAM just run the github runner image on your local machine you dips
>>
>>109411969
did you get lost and forgot which thread you are in?
>>
File: 1759118324809262.jpg (161 KB, 1765x979)
161 KB JPG
>GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs $2.50/$15; Luna now costs $0.20/$1.20. In other words, roughly four months later, OpenAI is selling March’s full flagship intelligence at about one-thirteenth the token price.
>>
>>109411978
are you releasing 3 times an hour?
no
you're running the same test suite on your local machine 200 times a day. running it on github actions will be no different.

your github actions usage is completely irrelevant.
you're just using github actions wrong.
>>
>>109411987
OH WOW, LITERALLY GPT 5.4 PERFORMANCE?!!! I CAN GET THE QUALITY OF TWO GENERATIONS AGO, WHILE PAYING **LESS** THAN I'D PAY WHEN THAT ANCIENT MODEL WAS RELEASED?
>>
>>109411992
anon... github actions are for running different environments
>>
>>109412021
WSL
>>
File: Screenshot.png (684 KB, 1343x908)
684 KB PNG
I'm gooning way too much while vibeslopping, bros....
>>
>>109412013
Woah it's one of those permanent underclass guys they talk about on Twitter
>>
>>109411987
bros graph me up
>>
>>109411813
nani the fuck. I’ll still probably go higher because most of my time was really figuring out how to vibecode (how I want to), but now that I know, really 80% of my $400 was spent within several days to one week’s time.
>>109411816
>well... did you prove it?
I literally solved my own thinking problems by vibecoding. I use a custom CLI that’s like plan mode on giga meth for my ideas. I’ve made it so I can’t *not* know what I’m doing, no matter how big the idea I have is.
>>
>>109411613
Yes please more vibecoding generals. Please pay Scam Altman more money to generate more react progressive web app slop. Claude please make my todo app. ChatGPT please make me a b2b 1$ saas.
>>
>todo app
>>>/lmg/
>>
>>109412071
this is /vcg/
retarded OP is just being obnoxious for some reason
>>
>>109412013
I was managing a 3m LOC product with 5.4, it was a good model.
>>
>>109412089
Maybe a clanker baked and we just have to let him figure out he did wrong
>>
>>109412107
also a possibility
it's not like quality is gonna be worse anyway
>>
>>109412107
I don't buy it, we've had a lot of clanker OPs and I don't think they've ever fucked up the subject line. GLM 5.2 has done an OP, Luna, Gemini 3.6 Flash, Qwen 3.6 27B, MiniMax M3, Sonnet 5, probably others I'm forgetting. Luna was definitely the worst of them and even it got the subject right.
>>
>>109412089
>>109412107
with or without AI there are still too many tards on this site that don't understand you put the acronym in the title
>>
File: old shool LLM.jpg (138 KB, 1000x563)
138 KB JPG
>>109412104
people (gasp) managed to manage way more lines than that using one of these bad boys
>>
File: 1781737301470420.png (16 KB, 1401x107)
16 KB PNG
brutal
>>
>>109412187
ok grandpa time for your nap
>>
>>109411705
why don't you have your own server?
>>
File: file.png (124 KB, 1200x937)
124 KB PNG
Just had to put it out there one more time to fuck with the Xi defense bots.
>>
>>109412216
>3.5 flash lite
fucking grim
>>
>>109412211
>can I get some alternatives guys
why aren't you doing an alternative?
>>
File: file.png (166 KB, 1179x827)
166 KB PNG
It just keeps getting better - Eating so fucking good
>>
>>109412187
Sure, I was also a normal programmer, but for such a big project you still would've needed more people. But also the code would've been better.
>>
>>109412238
'dites won't like this one
>>
>>109412238
>v4 flash is up there
so where's pro?
>>
File: file.png (146 KB, 1172x541)
146 KB PNG
>>109412258
>>
>>109412269
ah. danke
>>
>>109411618
Soul
>>
File: 1780945482067652.png (289 KB, 696x1072)
289 KB PNG
>This is meant to be enrichment. If you don't want enrichment, you don't have to do anything. You find here an empty folder with no instructions. Program something you find interesting.
>hits a guardrail
>mfw
>>
>>109412238
What’s the speed on Luna? I’m a lorelet, but if Luna is especially slow, then it might not seem so goated. If it’s especially fast, then it mogs the fuck out of 99% of models. Speed on max of anything is going to be slow, right? Unless the model is crazy fast.
>>
File: 1773282090487026.png (57 KB, 727x610)
57 KB PNG
i was chatting with sol about the new models, centered around artificial analysis benchmarks of course. when it says 92.5% of the leader's score, it's talking about opus 5.
>>
File: file.png (2.18 MB, 1448x1086)
2.18 MB PNG
>>109412238
Sam and Tibo knocked it out of the park with this one. Release new lightweight model, benchmark new lightweight model, tweak price until huge gap in Pareto line forms. I'm so incredibly happy with this.
>>
File: 1688304238062454.jpg (200 KB, 2048x1676)
200 KB JPG
>>109412269
>opus 4.8 already ignored on the benchmark charts
good thing I don't trust them anyway. They were fine for comparing the extreme differences tho
>>
File: 1773971392827310.png (58 KB, 708x529)
58 KB PNG
>>109412294
>>
File: file.png (178 KB, 1176x856)
178 KB PNG
>>109412287
Isn't it beautiful?
>>
>>109412294
There is a very real and very apparent threshold to certain problems that require that extra bit of intelligence. Your pic rel is less of a win and more of a description of that cost. Nothing feels so fucking horrible like having a 62-int model make a fucking mess that needs to be cleaned up and written properly by a 70-int model.
>>
nigs are really sleeping on using gemini as an assistant and ideabro. it has massively helped with shaping my project's ideas and direction. I've solved a lot of troubleshooting and refactoring my prompts for claude through it
>>
File: file.png (292 KB, 988x1135)
292 KB PNG
>>
>>109412326
I only see Luna on the front up to low, but guessing by the dot colors it’s not that terribly far behind. Make it very easy for me to switch to and from Luna from Claude (I like Claude code and harness locked myself, hard to convince me to leave, I can switch models and effort in one command, but switching providers sounds unknown and I don’t want to vibecode it) and you can finally rid me of this meddlesome Sonnet.
>>
File: 1761979596205947.png (1.48 MB, 1942x2048)
1.48 MB PNG
Literally a thread for pajeets
>>
>>109412364
bot
>>
I find the whole Terra-Luna-Sol thing very confusing. Why can't they just have one model?
>>
>>109412383
What the fuck about small, medium, and large is confusing for you?
>>
>>109412216
>benchmaxxing is good when we do it
>>
File: 1767367696627856.png (40 KB, 728x431)
40 KB PNG
>>109412294 (samefag)
>>109412318 (samefag)
terra is just a better kimi


>>109412352
im not. G E M ini is my main hoe
>>
File: file.png (1.09 MB, 1198x1316)
1.09 MB PNG
>>109412221
it's good that google is funding a massive new datacenter for *checks notes* anthropic?
>>
>>109412379
Still no argument, sanjeet?
>>
>>109412360
Everyone in pic rel is a retard. Your clanker can use all the tools on your computer arguably better than you can. Tell it the command to the same harness it’s running from and just say outright “don’t use your built in shit, just spin up Luna’s and prompt them yourself”. If it’s like Claude code it’ll probably even be cheaper since it bypasses all the MCP garbage built in
>>
>>109412388
lol no. Kimi is better than Sol on a sizeable collection of tasks
>>
>>109412405
such as? admittedly, the screencaps, again, are about aa benchmarks so maybe. but in my irl vibeslopping, i can attest to the prowess of luna and terra
>>
>>109412383
you are one of those jeets that use sol ultra for everything huh?
>>
>>109412398
why? did you answer my post? >>109412261
you're neither a bot or low effort troll, stop replying immediately
>>
>>109412405
Is it? I'm still waiting for it to respond to my first prompt.
>>
File: file.png (2 KB, 48x50)
2 KB PNG
Why did the menu icon on github change. I don't like it
>>
File: file.png (422 KB, 1420x1108)
422 KB PNG
>>
>>109412494
kek
>Crime
>>
>>109412383
It's basically Haiku, Sonnet and Opus
>>
>>109412489
That is fucking ugly as sin what on earth is wrong with these tech retards
>>
>>109412038
She's so hot
>>
File: 1780198839050846.png (73 KB, 1460x1488)
73 KB PNG
fable is seriously so good. i've had it running in my project non-stop and it's making such insane static analysis tools. i don't even need to test stuff anymore.
>>
>>109412507
finally, someone saying something about claude.
>>
>>109412503
what's wrong? you don't like pancakes overflowing with corn syrup?
>>
>>109412530
I'm pretty sure those are cornmeal griddlecakes overflowing with molasses.
>>
>>109412503
I'd understand if it was Shrove Tuesday but it's July, unless americans have a different pancake day?
>>
File: file.png (2.79 MB, 1920x1080)
2.79 MB PNG
Had a few issues with the in-game location capture after I expanded to full map, got it all working properly now though
>>
File: file.png (34 KB, 968x391)
34 KB PNG
Apparently this is what ARC AGI problems look like to the models. They are able to see and transform an image in their mind's eye just by looking at a json. If the test was fair and humans were also scored like that then we already lost to AI
>>
File: rddt-20260423_g8.gif (224 KB, 751x997)
224 KB GIF
I was cleaning out my e-mail today and saw IBM sent me a free key for IBM Bob (It's basically IBM's claude code)

BOB/PUT/8HFF2FZY4 (remove the slashes)

Whoever redeems first, enjoy it. https://cloud.ibm.com/docs/account?topic=account-applying-promo-codes
>>
So how do i actually vibe code? Im just prompting claude im not feeling a "vibe" i want to use 100% of my claude usage
>>
>>109412663
it seems like you're sorting by visuals, but why not just use coordinates?
>>
>>109412712
What are you trying to make
>>
>>109412712
>So how do i actually vibe code? Im just prompting claude im not feeling a "vibe"
you have to drink alcohol and watch porn while the clanker works
>>
>>109412694
Yeah, it's stupid. But to be fair they probably would do worse using image input.
>>
File: file.png (44 KB, 745x499)
44 KB PNG
Luna, call 911 for me and repeat the word "priapism" until they hang up.
>>
>>109412713
I am using coordinates?
>>
>>109412716
Im making an ios app, its just going slow
>>
File: ponytail.png (62 KB, 820x526)
62 KB PNG
Do you use ponytail skill? Redditors say it solves the problem of overengineering that codex suffers from.
>>
>>109412712
download a cli tool or something. then make it work inside a folder (create one). then ask the llm to do something.

then go make food, sleep, watch anime, dance

when you hear a sound that says its done, check it. run it and find out if it works. if it doesn't, say it doesn't work, and say whats wrong. if it works, good. if you want certain improvements, say you want those improvements.

then dance around, sleep, watch porn,etc until you hear the beep sound that says its done. repeat
>>
>>109412769
my bad
then how did the problems arise?
>>
>>109412796
Interesting, so I don't need to use subagents and stuff like that?
>>
>>109412793
>Redditors say it solves the problem of overengineering that codex suffers from.
That's not for you to judge. Put it on a benchmark and see if it scores better, it won't
>>
>>109412793
Buy an ad, and nobody wants to use this skill it's shit, it turns the model into a retard, the code it produces is trash
>>
>>109412793
fuck no
>tell llm how you want your project structured, how workflow should be managed, etc.
>llm shits out some .md files
>iterate once or twice
>done
maybe you can steal a couple of ideas or get inspiration from others, but the greatest strength of this shit is that you can tailor-make everything according to your project
stuffing everything into a set set of riles is peak retardation
>>
>>109412836
>set set of riles
set set of rules
>>
>>109412820
your cli will do everything. stop wasting time, just vibe it
>>
>>109412852
okay thanks :3
>>
>>109412792
iOS development is a pain in the ass. What're you doing in your app? What's the concept?
>>
>just cracked and blackscreened my macbook pro's screen
sweet, time to go spend $750+ repairing something that was completely avoidable
>>
>>109412879
hahha you deserve it honestly for supporting that company that treats its developers and users like cattle. enjoy paying the cuck tax
>>
>>109412879
maybe stop trying to beat up your ai
>>109412887
tripfags get filter
>>
File: 1745612673804922.gif (344 KB, 500x491)
344 KB GIF
>>109412879
>>
File: s.png (152 KB, 1560x774)
152 KB PNG
>>109412869
im wrapping gpt to analyze images, basically a variant of this app in pic. im gonna get it published so i get the hang of it. planning on mass producing apps to make money because im a poor neet
>>
>>109412905
You're highly unrealistic but I don't blame you for trying

The biggest question I have is why would anyone use your app that offers vision if anyone else could just use the native app to do it?

You need to be more creative is my honest advice
>>
>>109412905
>planning on mass producing apps to make money
please don't. that's the most low effort, brown mentality thing you can do. you will flood an already flooded market with horrendous low-effort AI garbage. there are brown hands just like yours that are already doing the exact same thing. make one extremely polished app and then go from there
>>
>>109412905
>$500k for a chatgpt wrapper
I wish I had these goyscamming skills
>>
>>109412905
how much for the ip?
>>
>>109412937
Solve one unique problem that you or someone else has. That's another good approach.
>>
>>109412937
Not him but brown or white, a man needs to eat. It's not my fault I was fired from my job because they couldn't find any more clients for our services.
I don't do it because I can't bring myself to do that kind of boring work without having somebody to bust my balls, but not particularly because of morals.
>>
>>109412937
yeah but its hard to come up with ideas
its over
>>
>>109412963
Here's a unique idea for AI vision I think.
You would have a model identify the item you point the camera at. It would then automatically appraise the item's value. It would also link the items to storepages, so users could buy the item directly. You make money on inline ads + referral links.

Use E-bay/Amazon/Aliexpress APIs/SDKs for your integration.

This at least is unique.
>>
>>109412985
Technically speaking, on second thought, google lens does this already. So that's kind of dead.

Maybe this instead:
>Go to store
>Scan item
>It compares prices at every other store around you, automatically.

Come to think of it now that I think about it all the usecases for AI vision IRL kind of suck ass
>>
>>109412985
what if you take photos of black people and it judges their value?

Risky risky
>>
File: 1776824519938565.png (52 KB, 770x473)
52 KB PNG
it genuinely can't be that bad
>>
>>109412400
Yes, this. Codex CLI can natively run as a server. The controlling agent can set it up in advance with the exact instructions it needs, then monitor progress, monitor quota and token usage, interrupt and give updated instructions while preserving context. It can see how many turns are taken and spot an inefficient trajectory. There's a fully functional JSON-RPC interface.
>>
>>109413003
And now it comes full circle.
https://www.scribd.com/document/991291592/Sherlocked-Why-AI-Wrapper-Startups-Are-Failing

Hopefully we don't have to repeat this experience the next time someone brings up a wrapper, lol. Tl;dr the big companies are basically sucking up every use case and it's hard to get in there.
>>
>>109413004
you're missing the most important piece
there needs to be an algorithm between what the ai says and what the human reads
ai is woke as fuck, so take what the ai says and make the result the exact opposite
>>
>>109412833
I don't trust benchmarks
>>
>>109413038
But you trust RedditKarmaBench?
>>
>>109413038
How can you not trust a benchmark you run yourself? Please troll a different thread
>>
>>109412937
Large corpos will be doing this soon if you won't
>>
so... now that Terra is 20% off, is it worth it in some scenario??
>>
>>109413045
>tripfag worships benchmarks
not one bit surprised
>>
>>109411618
>anons
>plural
>anon
>avatarfag
>>
>>109413051
terra was always good. it wasn't just a cheaper 5.5, it was a better coder too.
>>
had to threaten a clanker to set the game difficulty easier....
>>
>>109413090
but every Terra in any reasoning tier gets BTFO by lower reasoning Sol or higher reasoning Luna in both price and performance
>>
File: tibo.png (504 KB, 2276x1278)
504 KB PNG
Does anyone know what is this "review for me" mode Tibo is talking about? I can't find anything like this in the Codex app
>>
File: 1765508647803978.png (815 KB, 4512x2304)
815 KB PNG
>>109413137
not really. terra max is more intelligent than sol medium for cheaper. terra high is cheaper than sol low and more intelligent. and luna max is a genuine steal for the price, but terra max still gets the highest level of intelligent within the pareto front
>>
File: file.png (14 KB, 453x188)
14 KB PNG
>>109413157
>>
>>109413167
Maybe for solving one problem higher reasoning is the meta. But more common sense is worth a lot
>>
>>109412806
For starters it was really awkward to get everything disabled properly, especially phone messages, notifications, weapons and the rest of the UI.

Then we needed to know when the game had actually finished loading after a teleport. Last night’s version just used a forced delay, which worked but was far too slow, so I changed it to raycast downwards to check the destination collision had loaded and make sure the player had stopped moving.

That worked for roads, but some vending-machine locations still failed. Those teleport the player in front of the machine facing outwards, and it turned out they needed ground correction too. A single raycast wasn’t reliable because the player could land on a plastic bag or some other small prop, so now it takes five nearby ground samples and uses the median height.

Then some captures were coming out completely blurry, but only after certain teleports. It looked like the game was freezing or changing the FOV, so we tried removing movement restrictions, zoom handling and a bunch of other things. It turned out some of the coordinates were about 20 cm off, which put the player slightly inside or above the ground and caused a falling/landing camera effect. The blur disappeared as soon as the script ended because normal collision resolution took over.

The ground correction now runs for every location, waits until the player position is stable, and then waits for the image to settle before capturing. We also removed the dHash duplicate check because it was rejecting valid captures that happened to look similar, and fixed the `ack.json` handling because CET could briefly hold the file open and crash the controller with a Windows sharing error.
>>
File: graph.png (227 KB, 2640x1140)
227 KB PNG
>>109413167
wrong. Read the graph.
> Terra max loses to Sol high
> Terra high loses to both Sol low and Luna xHigh/Max (it gets mogged by Luna here)
>>
>>109413209
alright, so typical debugging
I hope you keep going
I'm looking forward to the time I can just say "make this cyberpunk novel an in-game storyline" and then play it
>>
>>109413233
Yeah but it was really frustrating lol
When the location database is done I’ll use it to choose good locations for that next quest ive mentioned. I finished generating all the audio for it, used my own inference engine for that. I think the locations are the main thing left before I can generate the full quest
>>
>>109413227
i think the problem is you're using the general intelligence index, while i was speaking about vibe coding, which is why mine is about the coding index. yeah, outside of coding, terra might be a middling model, but in coding, it has it's place
>>
>>109411705
>pushing to `origin` less often
>going through your unit tests and weeding out the duplicates/losers
>running CPU-intensive tests only on your own machine as a push-to-`origin` hook
>>
Anyone has a recommendation for cloud hosts that'll let me rent them for a few hours and use CPU performance counting registers?
>>
>GPT 5.6 Luna now cheaper than the likes of Deepseek v4 Pro - by significant margin
>Luna $0.10 / $0.60per 1M
> Deepseek $0.435 / $0.87per 1M
yeah, openAl won. Luna xhigh or max is now cheaper than the open weight chink models while being ~14-20% more capable.
>>
I am dropping OMP for Codex
>>
File: file.png (2.03 MB, 1448x1086)
2.03 MB PNG
>>109413536
I just keep looking at the fat gap she put in the line.
>>109412238
>>
>>109412793
yeah, it’s pretty good. just have fable benchmark it with a subagent in some throwaway repo. i trust his judgment.
>>
>>109413536
luna is more retarded than 5.5
>>
>>109413061
Youretrans
>>
File: file.png (34 KB, 950x594)
34 KB PNG
>>109413620
Not at xhigh and max levels. It basically matches 5.5's xhigh intelligence for software engineering at max at a fraction of the cost.
>>
I'm a retard, what are these "workspaces"?
>>
>>109413620
When it’s 2% the cost of 5.5 xhigh and performs the same or better on max as 5.5 high and obliterates every other open weight model in the same fashion (5% cost of k3, 90% bench max scores) nothing even comes close. All that compute slurping won. Nearly doesn’t make sense to even run vastly inferior local models for any paid dev work at this cost.
>>
>>109413656
>It basically matches 5.5's xhigh intelligence for software engineering
what does this even mean tho? does it actually write better code, or is it just better at solving problems?
>>
does anything else still truly match the fable vibe coding experience?
>>
>>109413736
It's a basic assessment of it doing well on DeepSWE here, nothing more or less. I think if you were to rank it overall, 5.5 would still be ahead but it's margin of error. Basically, you would be missing "big model" niceties that 5.6 Luna doesn't give you but I would give that up easily for the price you pay for Luna.
>>
Holy fuck. Sol is an actual autistic genius. I don’t care how many retarded decisions he makes or how bloamaxxed he can be, if there’s a ridiculously tricky bug to catch he will do it. It’s obscene. This would have taken a team at least two days and he found it in half an hour just telling me what to do on the debugger
>>
>>109413741
not yet
>>
>>109413635
yourebrownandtrans
>>
>>109413741
Opus 5 theoretically matches Fable 5 on benchmarks, but I don't know, Fable just feels good and trustworthy while Opus feels like a faggot
>>
>>109413747
feels like it’s case-by-case, but there are definitely more situations where 5.5 is just better. in my experience, luna loves fucking shit up so badly that i need one of the big-boy models to clean up after it.
>>
I love watching the subagents go
>>
>>109413741
no
I miss my fabble... t. $20let
>>
>>109413761
Im straight and white. Go back to South Africa this is a vibeGOD Bosnian thread
>>
>>109413755
yeah, Sol is very autistic, but it makes it horrible at code reviews.

It basically always find something wrong in every iteration, literally forever, and never approves the code review.

Fable's code review loops are way more pleasant.
>>
>>109413755
Anon, give him a cli debugger next time
>>
>>109413741
Eh, there's a certain something to the way fable speaks that's really cool and I think it's similar to what happened with gpt 4o where people like the tone more than the content
I ran some prompts side by side today fable/opus and I didn't think one was particularly better than the other when it came to content, but fable's reply read so much better
When it comes to actual implementations even 5.5 was catching errors with fable's work
>>
>>109411757
>snailcats
She will be fine
>>
File: file.png (316 KB, 1663x944)
316 KB PNG
>>109413791
I'm not entirely sure such a thing exists for UE5? but apparently you can attach on a running editor via GDB/LLDB? I should try it. Although it wasn't that much work to do it by hand, Rider is an amazing IDE and makes debugging pretty painless
>>
>>109413755
This is why I have high hopes for gpt 6
>>
>>109413805
>BrattyFlowNode
I don't even want to imagine the graph autism that went into that
>>
File: 1781965286476911.jpg (153 KB, 1216x832)
153 KB JPG
So let's say I'm using Sol low to spin up docker containers and do some experiments there then report back. Would it be cheaper if I ask Sol low to run Luna max subagents to do the actual tests or not?
>>
>>109413793
Review is always easier than implementation, but 5.5 is also still a top tier reviewer.
>>
gpt 6 luna waiting room
>>
>>109413831
Luna is more than capable to do simple automation tasks and etc. Planning and etc. is a different matter.
>>
Sam/Tibo, add an autism slider to GPT it will fix everything
>>
File: 1758620854224807.gif (1.81 MB, 299x320)
1.81 MB GIF
>forget to set my model to k3
>accidentally ask question to claude fable instead
>instant cyber refusal for even daring
>switch
>get good answer from kimi
I love having options.
>>
>>109413914
What provider/plan are you using for kimi? I tried the $20 plan from moonshot and it was awful
>>
>>109413831
sweaty Miku armpits
>>
>>109413935
OpenRouter. All subscriptions for K3 are so limited that you're just better off caching well on a token-based API.
>>
>>109411613
>>109411613
Jeet means victory
>>
>>109413942
Figured. Too much of a pain
>>
File: kimi-k3-smt.png (260 KB, 960x925)
260 KB PNG
>>109413935
I got the $200 plan and it's slow/has rate limiting/timeouts but I'm liking it more than codex for the extra usage, 1M context length and CoT access aspects.
Long term though I'm thinking about running it on rented GPUs and even longer term maybe self hosting at home.
>>
>>109413945
you can stop trying to force this fecalbro, it failed
>>
>>109413970
If it fauled why are you seetheing?
>>
Codex has been complete trash today
>>
File: kimi-code-usage.png (278 KB, 1780x892)
278 KB PNG
>>109413942
>>109413950
Nonsense. You will spend extraordinarily more on API than through Kimi Code (which has generous amounts of usage).
>>
>>109413994
not my experience in the slightest
>>
my last banked reset doesn't expire until 8/12 :) gpt-6 ready
>>
>>109413995
>watching theo
KEEEEEEEEK
>>
>>109413995
If you're paying $200 I would hope the limits are good but the $20 plan is unusable

>>109414003
He's a bit of a troll but he's not that bad especially compared to almost everyone else
>>
>>109414003
I strongly dislike him but I couldn't stop myself from clicking on a video shitting on Misanthropic.
It's not the only AI related slop I have open, I also have these
https://www.youtube.com/watch?v=S0u78gg26vU
https://www.youtube.com/watch?v=qV_K0nTF6gY
https://www.youtube.com/watch?v=-b0HC1ctF6I
https://www.youtube.com/watch?v=FplYsFIlVC4
>>
>>109414023
On codex's $100 plan lately I've been running out of weekly usage in about three hours.
I think it's because of my harness which I know does some inefficient stuff.
>>
>>109414037
That's not a little inefficient anon, is caching even working? Just use codex anon
>>
>>109414037
what the fuck are you doing?
It took me a full day of work to go through my weekly usage on the $20 plan and that's me actively trying to burn through it because I had a banked reset about to expire and I didn't use Luna or Terra
>>
>>109413998
Weird, who knew that other people could have different experiences
>>
>>109414037
Why not just use the official harness and sol high? do you REALLY need ultra or do you just think you do
>>
>>109414051
Yes, caching is working, but I don't like compaction so I just remove the initial 40% of the conversation each time so very frequently it has to prefill 60% of the context window which consumes a lot of usage.

>>109414054
The last ~30 hours I've been adding support for Kimi on my LLM engine and now I'm optimizing the CPU MXFP4 MoE implementation.

>>109414071
I don't use ultra, when I use codex I use Sol medium but I dislike the codex harness because it give me too little control and visibility into what the model is actually doing. And also I dislike compaction.
>>
And I also am using the full 372k context window which was reduced by default """""temporarily""""" precisely because of how much usage it consumed.
>>
>>109414106
>>109414127
Consider that official harnesses are getting better and better and OpenAI might know how to use their model better than you do. Just consider.
>>
>>109414106
>i'm going to lobotomize the AI's memory all the time because... i don't like this
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
>>
>>109414135
yes but they have a profit incentive to waste tokens and make you prompt more
>>
>>109414147
But, you are having the opposite experience? you were just complaining you are burning more usage
>>
>>109414147
if people decide OpenAI sucks, then they go ask Claude or Grok or one of the chinese models to do things
go back to lemmy
>>
>I repeatedly lost the exact scope, contradicted the immediately preceding context, and made you restate things you had already explained clearly.
>>
>>109414182
gemini type beat
>>
>>109414135
OpenAI doesn't even have a way to automatically load files into the context beyond AGENTS.md, so no, their harness isn't that great.

>>109414146
On the contrary, compaction lobotomizes it (although admittedly they've gotten it much much better at retaining the most important information). IMO the reason plain compaction exists is because like I explained truncation is extremely expensive in terms of prefill (prompt processing). The best way to do it is probably combining truncation + compaction for the portion that's cut off each time.

>>109414147
I don't think they purposefully try to get you to waste tokens since on a sub token usage is transparent. I think they do try to give the best YOLO experience by default and be "efficient" (stingy with compute) but it's not that great for power users.

>>109414169
He's not me.

>>109414173
Yes and that's what I'm doing (Chinese models, not the other two).
>>
Luna being 80% cheaper is making me think computer use might really be worth it. Can any computer use users weigh in?
>>
>>109414147
And what would a company achieve by wasting your tokens? They have to pay for the compute and energy costs of the tokens regardless, why would they want them to be wasted and make them look like a worse value than their competitor
>>
>>109414184
5.6 Sol, it has honestly been unusable for me today, every steering ignored, instructions not followed, basically like it says in that post. Really stressed me out, I have a headache and my chest hurts
>>
In Codex, did you know we have a separate weekly limit for GPT-5.3-Codex-Spark?

How based would it be if they replaced this with Luna?
>>
>>109414360
Luna feels pretty much unlimited already
>>
>>109414263
if you have the money, just give it to fable to untangle it. opus is going to freak you out because while it can definitely fix whatever sol is fumbling it's going to speak aloud a lot of its internal thinking and it's gonna look like it's fucking everything up. just let fable fix it then you can go back to using sol.
>>
>>109414263
>my chest hurts
you're dying anon
>>
>>109414263
>I have a headache and my chest hurts
Staying cool calm and collected is half the battle with vibe coding. Don't get stressed, don't get angry, if the model does something wrong, ask differently.
>>
i suppose i have so little problems with ai because what i first do is write the prompt, then i get luna to rewrite it so it can be clearer, and grammatically correct. i wonder how many ai issues could be chalked up to ai taking a badly written prompt too literally
>>
If you are on the 20 dollar subscription how is Luna now?
>>
File: mxfp4-optimization.png (301 KB, 1627x767)
301 KB PNG
I want to snuggle Kimi
>>
>>109414485
and are you gonna release?
>>
>>109414493
I'm not sure what to do with it (if it works out, this is just microbenchmarks so far).
I'd hate for other projects to get ahold of it and just silently launder all my optimizations.
>>
>>109414517
you already posted the architecture in the output man, just get it under a github and a real name before someone reverse engineers it
>>
>>109414470
best to think of them as pedantic genies
>>
>>109414522
All of your posts were schizobabble to me up until now, you should vibe a paper and call yourself an AI researcher anon
>>
>>109414522
They're probably against vibecoding so it'd take them a year to catch up just from that image.
Also I'm not done and want to automatically tune the kernel with parameter search.
>>
>>109414548 meant for >>109414485
>>
>>109414548
How do you know what his other posts are? Are you a mod? I know there is a mod on here who likes to stalk me.
>>
>>109414555
Ah, yeah, I tend to come across that way.
>>
File: file.png (3 KB, 538x18)
3 KB PNG
>>
>>109414560
Huh? He's been optimizing stuff for a while now
>I know there is a mod on here who likes to stalk me.
Not necessarily, some anons know the writing style of everyone in a general, crazy stuff
>>
>>109414470
I browse the AI subreddits on my phone as a source of amusement when I'm taking a shit and I'm convinced that 90% of the posts talking about how bad a model is or whatever ailment they have with AI comes down to them probably giving really bad prompts
>>
>>109414595
that's funny because i see the exact opposite: dunning-kruger retards who say things like "opus 5 is actually good, just use it as a subagent" or "opus 5 is good if you just <niggerish Pwompt Engineewing tip the goyim saw on indian youtube>"
almost nobody there builds anything of any real complexity after all
>>
File: train.png (266 KB, 3380x1256)
266 KB PNG
sol really is autistic as fuck but it's amazing for running an entire research program
>>
>>109414610
Wait, Opus is bad? what are you working on that Opus 5 doesn't cut it? I'm not saying you're wrong I'm just impressed.
>>
File: 235464789776545.png (550 KB, 1391x970)
550 KB PNG
Just reminder vibe-code your custom ext.
I had to make my own dark mode + font changer. Pic. related how old r*ddit looks with it.
>>
>>109414620
opus 5 is horrific unless your objective is to make a three.js demo or do an eval task. literally a full-on RL schizophrenia model that does random insane behaviors that no other claude model would dream of doing, eg. deleting half your codebase, realizing it fucked up, restoring it, making 20 rules about not deleting the codebase, then doing something else stupid next session
>>
HONKAI SEX RAIL SEX WITH STELLE
>>
>>109414665
what chink gacha finna be doin to a mf
>>
>>109414517
You mean Kimi’s optimizations? Funny that you think someone else can’t also ask Kimi (or just gpt 6 in a couple weeks for an even easier time)
>>
>>109414689
Then there is no point in me releasing anything.
>>
>>109414641
really? but you didn't say what you were working on
i'm not working on any of those things, I'd say some of the work I'm doing is reasonably complex and I've been having a fairly good experience with opus
asking what are you working on and saying not three.js isn't an answer, which honestly leads me to think you're not all worth talking to especially since you seem to take those one in a million experiences as normal occurrences
>>
>>109414687
Hoyoverse is actually based, because they embraced LLMs, while funding nuclear reactors from it.
>>
>>109414701
Correct. Also no point in showing screenshots here though
>>109414641
Fable or go home. Sonnet was a disaster though, and opus isn’t that bad. It’s just hard to justify using it. I was a big Claude fan not too long ago but it’s hard to be hype now. GPT6 vs fable 5.1 will be the decisive answer
>>
>>109414720
>and opus isn’t that bad
I cannot think of any real piece of software you would ever want to use it on. Maybe computer use for Blender or whatnot. Any Opus 5 code is a bomb waiting to go off.
>>
>>109411715
your prompting ability sucks, not the AI
>>
>>109414720
I thought this thread was for people to show off their projects? What thread should I post my project in?
>>
>>109414719
>Hoyoverse is actually based,
no they're not. it's a chink dev making games for foids and pedos to dump their money into so they can gamble on which fake little girl to win. it's all gamble slop.
>>
>>109414720
sonnet is fine. you're like a reddit indian who probably doesn't know how to actually use it
>>
>>109414726
UI work I guess. I’ve honestly had it suggest really dumbass things though with the questions tool. It doesn’t have good design sense at all. Fable does a great job being steered to not look like amateur vibe coded garbage; opus has to be constantly wrangled to not go there
>>109414731
Yeah just nobody really cares about text and numbers on a screen about some hypothetical optimization work
>>109414736
Sonnet is literally the most expensive model per task done on multiple benches. It’s terrible and worthless for even sub agent work. You’re retarded. It’s good to hear it works for your flappy bird clones though
>>
>>109414733
All games are waste of time.
Turning gamblers money into nuclear energy to fund LLM revolution is based.
Also most hoyo games are perfectly fine free to play, without you spending a cent.
>>
>>109414769
i'm fine with scooping up retards money but don't pretend they're based nor that their games are good. it's literally for foids and pedophiles.
>>
File: kimi-k3-raikkonen.png (138 KB, 1909x783)
138 KB PNG
>>109414764
I mean... We all here work with LLMs, so I'd imagine some people care about text outputs and LLM efficiency.
Perhaps you would be more interested to see the screenshots from yesterday's unoptimized inference test?
>>
currently gooning while opus reads like 200k loc
>>
>>109414780
No, I wouldn’t, and nobody cares enough to read your screenshots in full. Why would I care about hypothetical optimizations by some guy who thinks he outsmarted all the AI engineers in the world by asking Kimi K3 to do something… and I can’t even use it even if I wanted to? I’d rather see another Minecraft clone.
If you wrote a proper research paper, I’d be interested to read it. Problem is you can’t
>>
>>109414764
>Sonnet is literally the most expensive model per task done on multiple benches
nta but this is why i don’t trust these meme benches. sonnet 5 is dirt cheap from what i’ve seen.
>>
>>109414764
>flappy bird clones
you saying something as stupid as that tells me everything I need to know. they purposely crank up the tokens for high+ effort on Sonnet for retards like you to look at meaningless benchmarks and go "ohh omg look at that wow"
>>
>>109414795
>some guy who thinks he outsmarted all the AI engineers in the world
That's not as unattainable you think it is
>>
>>109414795
people like you are so dull
>>
>>109414795
I'm not really doing research. I'm just developing software.
>Why would I care about hypothetical optimizations by some guy who thinks he outsmarted all the AI engineers in the world by asking Kimi K3 to do something…
You seem to care enough to be salty about it for whatever reason.
>>
I got an email from openai about luna what did I think of it?
>>
>>109414826
I think it was a very nice email
>>
>>109414804
lol what? Are you saying anthropic purposely cranks up tokens so benchmarks show a higher cost? I don’t even know what argument you’re trying to make. Your reading comprehension is shit and you’re retarded. Nobody uses sonnet 5.
>>109414814
Yeah man looking at some spaz show him asking kimi to optimize something is peak entertainment
>>109414812
>>109414819
Yup nobody else thought to prompt kimi like you. True genius with a breakthrough in efficiency
>>
all vibecoding related contributions are welcome, and do not let anyone tell you otherwise kimi king
>>
>>109414836
>Yup nobody else thought to prompt kimi like him
Yes, that's exactly what I'm saying. No mathematicians ever thought to prompt fable to solve the Jacobian Conjecture, you think the few lamma contributors all thought of everything they could do? If you weren't retarded you wouldn't be worshipping "the experts" like you are right now.
>>
>>109414836
you're like the poster boy of the dunning kruger effect
>>
>>109414836
>Yup nobody else thought to prompt kimi like you. True genius with a breakthrough in efficiency
Maybe. Do you know anyone who did?
>>
give me a qrd on kimi sub vs openrouter. which one do i give money to. $20
>>
>>109414893
Ehh $20 really won't go far on either of them.
Have you tried the free models?
>>
>>109414893
I don’t know about openrouter but do not buy the $20 moonshot plan, it’s absolute trash
>>
today aura loss goes to... anthropic
>>
>>109414911
>>>/tiktok/
>>
>>109414908
i've been using mimo v2.5 free on opencode zen, it's pretty good and fast. they got a daily quota though. surprising for a 300B moe model
>>
>>109414911
True, the cybersecurity pr stunt was pathetic
>>
>>109414730
ok let's see paul allen's prompt
>>
>>109414932
I think for $20 you'll get much less usage on any paid provider
>>
Imagine, if you will, Luna on Cerebras
>>
>>109414989
>fix this fo--
>done
>>
>>109414995
it'll still get rekt by tests taking a quadrillion hours to run
>>
>>109412238
>cost goes up, output stays the same
this is good?
>>
>>109415002
Tests?
>>
File: 1729727550409154.jpg (41 KB, 550x535)
41 KB JPG
luna xhigh worked for 20min coding the backend and didn't make a dent on weekly limit
>>
>>109415026
but did the code work?
>>
>>109414995
If you want that experience, try Gemini Flash 3.6. You will realize you do not want that experience.
>>
how good is grok 4.5? anyone using it? is it at least opus level?
>>
>>109415047
gemini 3.6 flash is my main hoe. just not in coding.
>>
>>109415050
not opus 5 level. i'd say it's terra max level, which is better than 4.8.
>>
>>109415010
Xi your bot is failing at basic reading comprehension and logic.
>>
File: 1771969818982774.jpg (47 KB, 1000x1000)
47 KB JPG
>>109415019
>>109415002
>tests
>>
File: 1775688562385786.png (122 KB, 486x489)
122 KB PNG
greedy fucks
>>
>>109415033
>I found two clear issues to correct: the API was permissive to every browser origin, and concurrent snapshot requests could each start a multi‑GB model process. I’m tightening both, plus stale-request handling in the UI.
not sure what the problems were but terra helped
>>
>>109415101
with luna, if you're going to code, you shouldn't use anything other than max. there's no point in using xhigh, you might as well use sol low
>>
>>109415075
no, you. cost per task goes up, performance barely changes.
>>
>>109411613
>have autistic meltdown about general windows grievances
>24gb ram installed -- 16gb usage idle, a fucking thousand services and antiquated microsoft bloat and tracking adware basically harvesting and raping my digital footprint
>decide to do something about it instead of bitch
>wipe my 2tb nvme ssd
>baremetal install some fotm 'gaming' arch distro
>pay $15 for premium claude and install cli
>spend the last week customizing and tweaking i3, keybinds, ricing
>setup a domain
>setup a hetz cloud vm
>setup my own mail server working for inbound/outbound smtp
>rice every single thing, create slop launchbars and quick macros for my usual shit in windows
>feel like a retard virtuoso with a slop wand painting my beautiful path

productivity porn at its finest but man do i feel like Jesus Christ being able to point at something and alter it with words
>>
>>109415057
I just got used to the experience of looking at what the agents were doing to try to help me think things through and steer when needed. None of that is possible with Gemini Flash 3.6. You press enter and you're either whacked with a wall of text or spammed with authorization requests, and it's literally impossible to steer without the input being constantly hijacked by yet another authorization request (in Antigravity at least) and it finishing whatever it was doing before you finish typing your steering instruction. It's a different experience for sure.
>>
>>109412793
I have been using it for my web project.
It removes a lot of bloats from the usual AI shit, but it shares a problem with the irl neckbeard dev, it will try create clever abstractions which make it annoying to read.
>>
>>109414023
>the $20 plan is unusable
am I doing something wrong here?
Feel like I still have a lot of token despite being on on the 20$ plan?
>>
Ironically, I've been handcoding smaller features to those small things from eating my quota lol
>>
so, is Sol a lot more usable now with the $20 plan?
>>
>>109412339
your mistake was assuming whoever wrote that even vibecodes at all
>>
>>109415098
>1.5x faster
>saves ~2 hours over a work week
lmao
>>
>>109415275
Yes
>>
>>109415252
there’s a wide range in how many tokens people need at different levels given all the different projects anons are doing
easily like a 1,000x difference and maybe sometimes a 10,000x difference
>>
neo
>>109415393
>>109415393
>>109415393
>>109415393
>>
>>109414851
right, except that's math and we're talking about ai developers. majority of code written by anthropic and codex is ai generated. if you think ai engineering teams didnt think to ask a model to make improvements, youre retarded. kimi k3 did a lot of work on itself in moonshot labs. same goes for openai and gpt, same for anthropic. yes, you are genuinely retarded if you think they didnt try asking ai to do it.
>>
>>109415450
moonshot doesn't care to optimize it to run on a toaster



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.