[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: j-space.png (1.32 MB, 1150x1368)
1.32 MB PNG
A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

## What “vibe coding” is, and how to do it
https://simonwillison.net/2025/Mar/19/vibe-coding/
https://simonwillison.net/2025/Mar/11/using-llms-for-code/

## News
- (2026-07-16) Qwen announces plans to release Qwen3.8 as open-weights.
- (2026-07-16) Roblox announced Build, an AI workflow that turns text prompts into playable games
- (2026-07-16) Kimi K3 released and K3.1 announced. Performance reportedly comparable to GPT 5.6 Sol.

----

## Frontier models using fully-general tooling — start here if you have $20 or so
https://developers.openai.com/codex/cli
https://claude.com/product/claude-code

## Worth it for code, but the frontier models above are better
https://opencode.ai/
https://x.ai/cli

## Not worth it for code, but good for making sense of pictures
https://antigravity.google/product/antigravity-cli

----

## Prompting / context / skills
https://arps18.github.io/posts/claude-code-mastery/
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://github.com/mattpocock/skills — /grilling is a favorite

## Other editors / terminal agents / coding agents
https://osaurus.ai/
https://pi.dev/
https://cursor.com/docs
https://docs.windsurf.com/
https://docs.cline.bot/
https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent

## UI/Frontend
https://www.figma.com/make/
https://www.anthropic.com/news/claude-design-anthropic-labs
https://uiverse.io/
https://ui-ux-pro-max-skill.nextlevelbuilder.io/
https://stitch.withgoogle.com/

## In-browser builders / hosted vibe tools
https://bolt.new/
https://replit.com/
https://docs.github.com/en/copilot/tutorials/spark
https://v0.app/docs

## Benchmarks / rankings
https://www.tbench.ai/leaderboard/terminal-bench/2.0

## Previous thread
>>109335731
>>
>>109341441
>>
So the new NVidia cards will make LLMs cheap?
>>
>>109341469
no. just more profitable
>>
File: 1784660435260996.png (2.44 MB, 1402x1122)
2.44 MB PNG
What is your AI workflow and CI/CD pattern like? How often do you refactor etc?

Sol writes too much code and I have to ask him to simplify after each step
>>
>>109341561
So tired of asking sol to clean up his literal SHIT after every session
>>
>>109341561
I have a workflow, where at the end of each major stage a different model with high thinking effort performs a code review of every change before I commit. It finds and flags A LOT of shit.
>>
File: bombby.gif (3.96 MB, 616x374)
3.96 MB GIF
>found a new AI-powered game creation platform
>unlimited tokens

decided to make a FPS openfront/BAR inspired gamemode, this shits fun as fuck. it takes quite a few prompts to correct things but it gets the job done.
>>
stop munching on D&C, niggers
>>
>>109341561
https://github.com/DietrichGebert/ponytail
>>
File: wagmi.gif (3.75 MB, 378x359)
3.75 MB GIF
>>109341632
added some mechanics where enemy ai can spawn barriers, hide/shoot, drag wounded and revive.
>>
>>109341373
Grok heavy is on discount if you're interested, $99/month for three months.
>>
>>109341448
based
>>
File: 1784735745968.jpg (339 KB, 1150x922)
339 KB JPG
Yep, Dario cried to daddy Trump, and now Kimi is 100% getting banned. Maybe even all Chink models, if he begged hard enough
>>
>>109341711
scroll down? i wanted to see what timezone you're in
>>
>>109341711
how did they distill and train kimi in the 2 weeks fable was out before it got pulled, doesn't it take months to prepare a base model training run?
>>
>>109341722
Nope, it takes 45 minutes on average
>>
>>109341711
Dont the chinks have 100x more educated citizens than mutts?
>>
chinese education
>go american university. steer trade secrets
>>
>>109341718
My timezone is UTC+8. I cropped the screenshot because it was already a wall of text, no need to make it any longer.
>>
>it really is third world hours
>>
>>109341711
>distill the internet, books and any media you can get your hands on
>get mad when someone distills your outputs
Make it make sense
>>
File: pooh.png (81 KB, 315x280)
81 KB PNG
Vibecoding is the logical conclusion of the Jengatower of bullshit abstraction known as the modern web development tradition:
Acceptable "programming" to Redditors:
>"Bro code this framework which tells this framework to tell this framework to write JavaScript. I love programming, I'll take my $80k office job after my 12 week bootcamp! Huh? Memory management? Pointers? Machine code? What's that?"
Unacceptable "programming" to Redditors:
>"Prompt this AI to write this framework which tells this framework which tells this framework to write JavaScript? This is unacceptable, this isn't programming or the industry I signed up for!!!!!"

The cope counter:
>"b-b-b-but every programming language ever is an abstraction!!! C is just high-level assembly!!!!!"
You were still directly interfacing with the computer and telling it what to do. When you abstract above (already slow and inefficient) high-level languages like JavaScript and Python into the realm of "web frameworks" retardation like React, the initial CS fundamentals are already gone and have just been replaced with new complexity. At that point, might as well go all the way and tell AI to read "framework" documentation and do the "programming" for you.
Not like modern webdevs weren't almost exclusively writing slop before AI.
>>
>>109341640
You’re trans
>>
>>109341756
Yeah, but they only know how to copy and can't innovate, build everything on "Chinesium", are demographically fucked, real estate ruined, two weeks, Two Weeks, TWO WEEKS!
>>
File: file.png (304 KB, 1024x800)
304 KB PNG
>>
>>109341840
kek
>>
>>109341669
>>109341632

I'll bite, where do i try this? i play both games semi-frequently
>>
>>109341840
lol
might actually be fun to go back and dig out projects from before I got a job and just ask it to be brutally honest
>>
>>109341800
manifest mcdestiny
>>
>>109341800
the better question is: the same information is out on the internet for the chinese to consume. why not just do the same thing?
>>
File: ISHYGDDT.png (19 KB, 90x128)
19 KB PNG
>>109341561
>wasting usage on simplifying code
>doing it every time
>giving a shit at all about how many characters the clanker types
not vibe coding
>>
>>109341906
The problem is the cruft adds up and llms lose their ability to work in the codebase effeciently
>>
File: 1762468252909155.png (106 KB, 273x273)
106 KB PNG
>I'll fix it by creating a worse system and also keeping the old broken system I was supposed to fix as a fallback
why are they like this?
>>
>>109341711
Insane kike cope, kek.
>>
>>109342053
they got "no regressions" beat in to them to a point they are afraid to remove the old code paths
>>
>>109342053
they're programmed to not delete unless explicitly told to, otherwise we'd get "agent deleted my entire codebase" memes more often.
>>
Have you ever used Gemini Notebook (formerly NotebookLM) for anything?
>>
>>109342116
generating a podcast of my roleplay
>>
>>109341940
Just zap them in the testicles. Makes them bots fall back in line and stop making mistakes. Works every time.
>>
File: 554549876574667878.png (225 KB, 1310x760)
225 KB PNG
Might take loan for this bad boy.
>>
How do I activate a banked reset?
I just used up all my codex but I can't see the reset.
>>
>>109342274
/usage
>>
>>109342274
you have to upload the biometric scan of your anus to GPT first

>>109341940
spending a bunch of time and making a clanker redo things a different way sounds really efficient
>>
>>109341632
>>109341669
looks fun af
what platform btw?
>>
>>109341561
refactoring is non-stop. just an hour ago I realized claude re-derived the position of the ghost before placing the object instead of just using the ghost's actual position.

I'm not sure if it's a consequence of the sliding window, or if it's just claude, but I find that has a habit of writing nearly duplicate code and trying to maintain that duplicate code
>>
File: 1764154765493751.jpg (85 KB, 1024x800)
85 KB JPG
>>109341940
>>
>>109342254
>$5849 128GB LPDDR5X
A system that can barely handle a 120B (or a dirty 200B quant) model at usable speeds, versus 30 months of a 20x plan? Or a Ryzen AI Max+ 395 system with 128GB of RAM for under $4,000? That Ryzen AI Max+ 395 has the same memory bandwidth as the GB10, they're both using LPDDR5X, their token generation speed is almost identical, the GB10 only wins out at certain specific tasks. Don't get me wrong, I'd love one, but what's the justification to spend almost $6k on this thing? Systems like this are for development, not inference. Even with the GB10 with significantly faster prefill, it's not an inference machine, certainly not for vibecoding. A pair of 3090 24GBs would cost less and deliver a much more usable experience from admittedly even smaller models.
>>
>Crapmini releases 3.6 flash
>Hallucinates wildly on first prompt
Bruh
>>
>>109342423
>Ryzen AI Max+ 395 system with 128GB of RAM

You won't be able to fully utilize the entire 128GB of RAM on AI Max, probably like 100. NVIDIA DGX comes with its own Linux distro that reduces RAM consumption. If I went with non cuda route, I would buy apple m5.
>>
>>109342423
data sovereignty requirements
government, healthcare, and any industries who work with them often require having their own full ass LLM deployments just to make sure that the data they own never leaves their walls.

But yeah there's a ton of orgs also buying DGX sparks. Most of these orgs are lost and have literally zero clue what to do with them, but you could use them to fine-tune a dumb LLM for some specific purpose like classification or running PDF OCR.
>>
>>109341903
The Chinese did do that. It can be assumed that every single AI company has by now scrapped the entire Internet of data in a likely legally gray fashion. China just takes it a step further by also distilling the other models
>>
What are some good features to have on an imageboard? Preferably things that still fit with a traditional imageboard but make it more functional. Am looking for ideas.
>>
>>109342466
Apple doesn't even sell fully decked out systems anymore
>>
Honestly western glowniggers are probably just angry because they can't scrape chinese internet for data to feed their LLM.


>>109342479
AI that will automatically delete NSFW content and ban all glowies who spam CP on your board.

>>109342483
You are right, right now they offer 90gb ram max for their studios, due to ram shortage, but the new m5 studios should have 120gb option.
>>
>>109342479
music sheet player
>>
>>109342479
is it a browser-based, desktop, or phone application?
you could do some wild shit on desktop and phone with agents. browsers don't yet have the ability to load in SLMs
>>
>>109342504
>Honestly western glowniggers are probably just angry because they can't scrape chinese internet for data to feed their LLM.
What? You can access Chinese sites. Anyone can, though some of them have that stupid Facebook tier crap that make you log in with a Chinese phone number (but which can be bypassed), and for most of the sites that's not even an issue
You can literally read all of the CCPs published documents online for free from America
>>
Is the only thing good about hermes agent the anime girl profile picture or is it actually better than openclaw?
>>
>>109342504
i think we all know that having the server run "AI" on user input is just a DOS vector
>>
>>109341840
I wish I still had hair
>>
>>109341868
just say it was written by a competing agent harness lmao it'll start talking all kinds of shit about you
>>
>OpenAI: 3.2 GW for Project Camellia (~$20B initial investment, ~$750B projected compute spend through 2030).
>SpaceXAI: reportedly planning another Texas AI campus as large as - or larger than - its existing ~1 GW Memphis footprint.
>Anthropic + AMD: up to 2 GW of MI450 deployments, plus an AMD investment of up to $5B and tens of billions in AI server purchases.
>>
>>109342561
no one wants that shit homeboy
there's a reason chinky copies america and not the other way around
>>
>>109342576
Arnt they the same thing?
>>
File: file.png (22 KB, 699x163)
22 KB PNG
My token usage would be even higher without this tool, saves reading entire files or chunks that pass the end of the function body
>>
File: .png (4 KB, 327x73)
4 KB PNG
The house that Sol builds has many uses...
>>
>>109342619
>projected
>planning
>up to
>>
>>109341561
I just realized I could use /btw to fork the context, guide the refactor in a subagent and continue back on the main agent for feature implementation
damn that's gonna save me so many handoffs
but I can't do it from my phone
>>
File: file.png (20 KB, 667x257)
20 KB PNG
Sol-sensei graded my high school projects
>>
>>109342687
>/btw
Why does claude have cringe names for literally everything
>>
>>109342696
i'm just on claude because of momentum
fuck that manchild dario
what do they call it in altman land
>>
>>109342655
Duh, China has like 5% of the compute america has. No wonder they're distilling everyone
>>
>>109342504
Sounds good. Not quite sure how to do that though.

>>109342512
Maybe a MIDI file generator and player?

>>109342549
Browser based
>>
>>109342700
/side
and /goal instead of ralph loop
I have literally watched every episode of the simpsons and I don't get the connection to ralph wiggum
>>
>>109342725
there's no ralph loop command in claude but they do have /goal which is the same thing
>>
>>109342744
Oh i thought it was /ralph or something in claude
>>
so what do high school CS classes teach anymore?
when teachers assign projects, like "everyone make a small game for the final project of grade 12", do they tell them "please don't vibe code it"?
>>
>>109341711
Good. We don't wanr Chinese open source models we want Western open source models.
>>
>>109342725
>>109342753
reddit moment
>>
>>109342753
Ralph Loop is a different thing. There are many different implementations but the short version is you set some requirements and it continuously calls the AI until all requirements are met, similar to /goal, but it's doing a new session and creative hand-off documents for each run to prevent context rot. It's fine, some people use it, it can be effective if you have a verifiable goal but virtually useless otherwise. They were talked about quite a bit for a short time earlier in the year.
>>
>>109341711
>you can distill from a ~5T model and train a 2.8T one in two weeks
There is no moat if this is true
>>
File: 546352536758976.png (2.04 MB, 1209x1244)
2.04 MB PNG
>>109342701
Glowies and federal agents will always turn this discussion into specific model debate.

Notice how they always ignore physical reality, american grid can't support their LLM revolution, while chinese have been building their grid for past 3 decades and continue to do so? Chinese win the long-con anyways. They have capacity to build more data centers without their grid shitting pants. They will completely rape western LLMs on pricing.
>>
>>109342254
Not worth it. If you want local models use api for cheap ones, and use hourly rent through vast.ai for the rest.
You can almost indefinitely sustain that until memory prices go down in a year or two.
It's the only strategy that makes sense, these boxes only make sense if you buy 4-8 of them to run quants of huge models, and the prices become indecent.
>>
File: droney.gif (3.9 MB, 614x468)
3.9 MB GIF
>>109342389
>>109341857
https://makeplay.ai/p/a77nstybpq

pls try the factions gamemode
play > faction warfare > war-zone
>>
>>109342869
"Physical reality" is the fact that China doesn't have the computation and is also being dwarfed in terms of future projected compute. Just because they have more energy (spoiler they're already using most of that energy for industry) doesn't mean what you're saying
>>
Just used one of my banked resets. Was it a good idea?
>>
>>109342615
not a bad test actually
show your old shitty projects to different models, both as yourself, the same model, and competing models
>>
>>109341441
fuck scam altman and fuck tibo
>>
File: 1767559436222264.jpg (28 KB, 612x484)
28 KB JPG
>>109342912
>China doesn't have the computation and is also being dwarfed in terms of future projected compute.
So they were able to produce kimi 3.0 just by pure luck, LOL?

>(spoiler they're already using most of that energy for industry)
They keep building up their grid.
>>
>>109342934
yes, I doubt anything will happen until the weekend
they reached 10M, it's probably one of the last resets they'll ever give
>>
>>109342960
Why did they reach 10 million so fast.
A lot of people didn't have the time to use up all their usage before the 10 million reset.
>>
>Claude code
>Gay app
>Gay vscode integration
>5 hour limit hit after 4 prompts

>Codex
>Gayer app
>Freezes UI periodically
>Gayer vscode integration
>Try to generate pet, fail after 5 hours of token munching

>Copilot
>Good vscode integration
>Just fukkin werks
>Plenty of tokens

Tell me again why companies are flocking to anthropic and openai??
>>
>>109342960
you said that the last fifteen times
>>
>>109342946
Kimi 3 is not that good anon. If you mentioned Qwen I will laugh at you
>>
>>109342980
>vscode
ok grandpa
>>
>>109342993
What predominantly homosexual part of the world do you hail from?
>>
>>109342722
>browser-based
apparently chrome is trialing out a thing where you can expose an MCP server for your webpage, so that agents don't have to do some computer-use bullshit
https://developer.chrome.com/docs/ai/webmcp

I'd use it for example to perform automated sentiment analysis (this post makes you sound mad and this will not help you win the argument)
or try to cluster posts and try to figure out their motivations based on similar past posts
for example many haters often dump a meaningless out of context shitpost like >>109342943
but you could derive from historical data that posters like these are usually china supporters or bots or something
>>
>>109342984
this
>>
File: 1759919028247340.png (241 KB, 937x603)
241 KB PNG
>>
>>109342984
>Kimi 3 is not that good anon
If it's not good (Deepseek V4), there wouldn't be this much seethe and cope.
>>
>>109341711
mutts will be forced to use only drumpf approved models LMAo
>>
>>109343053
It's weird how nobody cares about deepseek V4 because it was shis
>>
>>109343068
give it a few months and no one will pretend to care about kimi anymore
>>
>>109343066
yes. the trump approved models. the ones you are happy to distill from?
>>
>>109343074
The kimi reception is different from the deepseek reception.
This was more like the reasoning reception when everyone shat themselves because deepseek taught AIs to think.
>>
>>109341441
Wen usage reset?
>>
>>109342972
No idea.

>>109342981
Not me.
>>
>>109343005
greece
>>
>>109343068
>in before v5
>>
>>109343103
I will pay 1% of what you pay for 99% of the quality
>>
>>109343119
and it will be 99% of your meager income, while the trump approved model will only be 1% of mine
>>
>>109343124
Why don't Europeans destill american models
>>
>>109343188
because europeans have honor and ethics
>>
>>109343188
they can't afford it kek
>>
>>109343188
europeans don't have tech
the only thing that mistral has going for it is that it has no guardrails according to terrorist watch groups lmao
>>
How does Opencode compare to Codex, harness-wise?

I've only used codex and claudecode so far. CC has a nicer UI, well made commands (like /code-review), a better interface for shit like agent management, adding context to question prompts after the y/n response, etc. But codex seems immensely better at actual agent management; whatever they're doing with what looks to be incremental compaction works great, actual read-only commands never surface permissions (while claude will routinely ask permissions for a massive paragraph of
echo | rg | sed | cat | xargs rg; git show | rg | cat | sed; rg | sed | cat; git diff | sed | rg | ...
 that "cannot be analysed statically".

How does opencode compare, both in terms of UI and in terms of agent harness?
>>
>>109343199
Should I be vibecoding with Mistral to have honor and ethics?
>>
>>109343226
Aiieee I fucked that up
>>
File: 1781305785548600.gif (2.8 MB, 453x498)
2.8 MB GIF
>kimi is better than sol because it was actually fable distilled the whole time
>>
>better
point and laugh
>>
>>109343239
bro should have asked codex to write his post for him
>>
>Across different platforms, the level of inaccuracy varied, with Perplexity answering 37 percent of the queries incorrectly, while Grok 3 had a much higher error rate, answering 94 percent of the queries incorrectly.

https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php


Is this an issue for vibe coding?
>>
>March 6, 2025
>>
>>109343268
I didn't click the link but it sounds highly improbable that these models are failing at needle in a haystack which is a solved problem.
>>
I want to marry nomic-embed-text
>>
>>109341675
Exactly, that's what make it interesting. I'm trying out grok 4.5 on cursor now..
>>
>>109343338
and? do you have a verdict yet?
>>
>>109343338
How is grok 4.5? Worth the sub?
>>
>>109343268
>March 6, 2025
You realize that's so long ago it's literally multiple generations behind the current state of ai
>>
>>109343394
to put this into context, this is BEFORE gemini released gemini 2.5 pro. like, even the first version
>>
File: 1647259545583.png (25 KB, 500x460)
25 KB PNG
>>109341808
stfu
>>
File: 1759153377710311.jpg (180 KB, 1206x1438)
180 KB JPG
>>
File: 1758501867847766.png (150 KB, 594x1247)
150 KB PNG
>>109342116
The main use case for that is to help people study course material (that's what it's best at doing so that's what I mostly uses for). Maybe I lack creativity but I can't think of anything at the top of my head that notebook llm would be particularly useful for in vibecoding beyond explaining complex information to someone. As far as I can see it clean up a few AI tools normies actually elderly like using because it does a damn good job at shitting out study material that's actually useful and in some case even better than what the instructor provides (if they even bother providing it)
>>
File: file.jpg (444 KB, 1920x1080)
444 KB JPG
Noticed the character cock is hanging out through the trousers
Clanker is working on fixing it
>>
fable makes the best quizzes
>>
>>109343507
lel, what did it say?
>>
>>109342479
Add that text based RPG that 4chins had for April fools Day a while back. Can't remember what the fuck it was called but it kicked ass
>>
File: a6d204fe90.webm (3.35 MB, 492x300)
3.35 MB
3.35 MB WEBM
>>109342885
NTA. Annnnnd I'm dead.
Cute game.
>>
>>109342980
Based on my limited understanding of programming and relatively high experience with vibecoding, it seems co-pilot is a very good option if you're already decent at programming. If you're also what grifters calls a good "prompt engineer" it's better because the person that knows exactly what they want and how they want it implemented knows exactly what to describe to the LLM. Codex, and arguably especially Claude, are explicitly trained to be very retard friendly will also be good at executing whatever their tasked well. Co-pilot is theoretically able to do what the other two can do but it requires a lot more explicit instructions and hand-holding compared to the others. The corporations that want to crowbar AI into all of their employees workflows choose the other two because it doesn't require their less than tech savvy employees or clientele to actually know what they're doing. Any corporate world outside of a development team, codex and Claude it's for the team that want to look like they're being productive by wasting a bunch of the company's tokens. Copilot is a decent tool if you know what the fuck you're doing. Which most people don't so they waste more tokens getting shit done on copilot, hence why Microsoft is considering moving to Chinese models since they are inherently cheaper. This isn't to say retards wouldn't be wasting money on codex or Claude either but it it's usually a lot worse on copilot because their employees are incapable of describing what the fuck they want effectively.
>>
File: file.jpg (1.74 MB, 1920x5400)
1.74 MB JPG
Rest of the quest is coming together though
>>
>>109343199
>mistral
I mean they're pretty bad in 2026 but still, at least they're doing...something?
Where is Germany or the UK?
>>
>>109343666
deepmind is the uk. google might be an american corporation, but it's deepmind/english employees building gemini
>>
Do I have things configed wrong or does Deepseek v4 Flash just not work with Cline? CONSTANT tool errors, unusable.

If not DSV4F what's the best bang for buck on openrouter?
>>
>>109339517
>someone finetunes K3 for cyber
That someone is me, I'm building the training dataset now in preparation for the weight release. I probably won't publish it but I'd be happy to hook anons up.
>>
>>109343226
Whatever it does behind the scenes allows models to perform better. Probably just better or concise system prompts and not throwing as much loaded tools into the system prompt as its competitors.


https://artificialanalysis.ai/agents/coding-agents

That's just an educated guess I pullrd out of my ass though so take what I said with the grain of salt.
>>
>>109343701
How do you go about building these data sets? Do you use existing data sets from huggingface? Where do you get the training data from?
>>
File: file.png (5 KB, 261x62)
5 KB PNG
I definitely need another 20x sub or two though
>>
>>109343714
Interesting. This confirms my intuition that CC is hot doodoo. Though just three benches with Opus 4.7 is not super comprehensive, especially since harness performance is probably extremely model-dependent.
Might have to try it and see for myself then, because at least these results do look promising, even if it's for an old opus and and not even codex
>>
>>109343775
>This confirms my intuition that CC is hot doodoo
This has been basically confirmed for months. They had a weird glitch a while back or if you ran it in your terminal the cursor would cause weird flickering on your screen and they actively refuse to fix it for months out of spite. And father has very good researchers but they don't really take engineering or technical shit very seriously and it reflected by how they act towards...well.... everything. Their heads are off their asses will also attempting to not look pretentious and instead present themselves as the "concerned adults in the world" or whatever.

They also did this malicious thing where if your code base even applied we used another harness at any point, you got charged hire API pricing (because they view you guys like bugs, just wanted to keep making fun of you guys....)

https://youtu.be/J8O9LLpJNrg?is=ujHxq9UNKcB_mKQU
>>
>>109343775
you can't confirm that though because the benchmark is woefully out of date. that image in the picture is literally the only benchmark of it's kind. artificial analysis do not, say, use kimi k3 on kimi code cli, then opencode, then claude code, so on and so forth, or glm 5.2 with opencode/claude code/codex/etc. there's not enough evidence
>>
I find opencode to be the most aesthetically pleasing coding agent.
It's quite obvious what is happening.
Claude Code may be better but its layout is so weird
>thinking
>response
>somehow more thinking
>the todo is here for some reason, also it's truncated
>>
>>109343803
Anthropic is such a petty corpo. Without reddit astroturfing and glue eaters that support them they'd have nothing.
>>
>>109343803
>Their heads are off their asses will also attempting to not look pretentious and instead present themselves as the "concerned adults in the world" or whatever.
It's a cult, they have competent researchers to make great models, but it's still a cult with some very weird safety obsession.
>>
>>109343701
damn, what training stack are you using to train such a behemoth? and how are you generating the data or what RL method are you using?
>>
honestly there are so many like opencode, pi, codex and so on, it's too hard to choose as they all seem ok
>>
>>109343536
>Clanker is working on fixing it
why?
>>
>>109343866
It's lewd
>>
>>109343740
Not him but if I had the compute what I would do is do a rollout, have Sol analyze the results and give advice on how to get a better result, then append that advice at the end of the original prompt, do the rollout again, save the probs, remove the advice from the context and train on the probs that were generated from the prompt with the advice with only the prompt without the advice as context. This is called prompt distillation.
But I do not have the compute so I'm only trying to distill from K3 into Qwen 35B to have K3 at home.
>>
>>109343816
The good thing about this is that we can test and compare ourselves whenever K3 gets released to other model providers. I might even take the hit financially for the sake of interest in research and compare k3's performance on open code, Claude code, and cursor harnesses respectively. I'm convinced open coat is the least bloated which is why it performs so reliably and consistently compared to others but I've never actually took the time to verify it myself with any data.


https://github.com/EleutherAI/lm-evaluation-harness

We are in a vibe coding general after all so perhaps I could take something like this and create a benchmark that is specifically meant to test both different models and different harnesses but instructing a suitably "intelligent" model(s) to create a harness version of this.

>>109343844
It's just malignant narcissism except it's not just one dude it's an entire company's leadership. Dario in particular is fascinating to me because this idiot doesn't even try to hide the fact that he's being a selfish prick and trying to use other people's paranoia and suffering to his own gain and trying constantly to get in bed with the government (well also lying and pretending the trying to stand up for people's privacy or morals or whatever). I don't like Sam either but I dislike Sam a lot less than I just like Dario because Sam has the courtesy of at least putting effort into pretending he wouldn't smile at normal people being turned into biofuel and actually presents himself as a normal fucking person. Dario comes off as some self-absorbed autist that thought having a lot of money and influence could compensate for having no social skills whatsoever.
>>
File: 1777418343107669.jpg (66 KB, 896x392)
66 KB JPG
Doing a lil something
>>
>>109343881
so is the game
>>
>>109343896
>making a harness benchmark
absolutely. i'd be lying if i said i wasn't partial to codex, but i'd be interested, mostly in how opencode fares against everyone else
>>
>>109343803
I don't follow agentic news or whatever so I've not kept up with the general consensus.
>they had a weird glitch
Their TUI rendering was absolute GARBAGE and would absolutely shit the bed all the time. That said they introduced their full-screen rendering mode and it's actually mostly fixed that for me, so I don't hate it anymore.
>And father has very good researchers but they don't really take engineering or technical shit very seriously and it reflected by how they act towards...well.... everything.
The thing is model performance should be squarely in the research category. Also I actually get the opposite impression from using CC, it looks like a bunch of webshitters slopped up a really "pretty" CLI that took over a year to unfuck the rendering, and have been constantly pushing new features, cool gadgets, commands, better controls, etc. Codex is downright barebones compared to CC, but Codex actually runs tasks well, which I'd expect would've been the focus of the "researcher" types.

>>109343816
I agree but as I said I have confirmation bias about CC, and even on this one old benchmark it does show there was at least one case where OC significantly outperformed Codex.
The one thing it's really missing though is a test of Codex models, since Codex (CLI) is built for Codex models and might respond very differently to Claude models.

>that image in the picture is literally the only benchmark of it's kind
Which is fucking criminal by the way, especially since almost all models (except Claude) are surprisingly swappable between agents/harnesses, and each harness is going to be very significantly different at handling the model. (Not just in the system prompts - though those are massively important - but also how shit like turns, tools and tool calls etc. are presented to models, how compactions are handled etc. Codex's context and compaction management is fantastic for example, I almost don't keep track of context anymore.)
>>
>>109343899
if lewd means shit yeah I agree
>>
File: 6f0.jpg (32 KB, 680x578)
32 KB JPG
>log in to $20 ChatGPT plan
>select GPT-5.6 Luna Medium
>/goal solve the riemann hypothesis
>go on vacation
>>
File: file.jpg (318 KB, 1920x1080)
318 KB JPG
>>109343899
Well it's fixed now, if you want the cock then you can fork it and add it back in
>>
>>109343922
>come back
>there's a nationwide water shortage
>>
>>109343864
you've unlocked a key insight
agent harnesses are all copying each other's homework
therefore all frontier models and their corresponding agent harnesses should act approximately the same within some minor margin of error.
all the people in the thread arguing about competing models and companies? minor brain damage
they all work.
>>
>>109343896
That's actually a neat idea. If it's that easy to run benchmarks, it might really be worth it. I would definitely pay to benchmark Sol in Codex vs. OC to determine if there's a measurable performance difference.
>>
>>109343617
kek, made me realize i should add a "press E" button when at the HQ, and press "B" for building & producing, cause i can see how it can be confusing if neither are showing up
>>
>>109343909
You're confusing being good at research with being a good engineer. There are people that graduated top of their class and see us that don't know how computers work. I'm not talking about the "being an expert at low level languages" kind of knowing how computers work. I mean they straight up need to be handheld when asked to do literally anything. They don't know how a terminal interface works. They don't know how git works or what is even used for (disappointingly many anons even here seem to not know or understand the importance of git tracking), but good at getting good grades but they're utterly useless when they're asked to do anything useful. Lectures how academic types work and how their brains work. You wouldn't necessarily want a theoretical physicist to be in charge of designing your car's safety features right? Being good at one discipline does not automatically mean you're going to be good at the others just because they both happen to be STEM related. Being an academic does not necessarily mean you have a good amount of common Sense. This isn't to say "hur duur academics are all stupid". Most of them aren't. But they seem to be very narrow-minded and what they care about and what they're good at project managers make the incorrect assumption that being an academic means you're going to be good at computers or going to pick up the skills quickly (that's almost never the case). I literally saw an example of this in my high school. Are Bella Victorian was very book smart to the point where she would even tutor on her free time but was kind of a dits that pretty much everything else. Not a complete retard. Just slower to pic things up that weren't School academic or sports related.
>>
>no access to kimi anymore
>my gpt usage is on cooldown
glm: you could not live with your failure... where did that bring you? back to me
>>
File: file.png (760 KB, 1078x1128)
760 KB PNG
>>109343948
nobody ITT can refute this
your models and harnesses would make the Habsburgs blush
>>
>>109343948
>agent harnesses are all copying each other's homework
Yes but also no. If you've ever tried both Codex and Claude Code, you'll see that they're extremely different both in UX and in how the models behave.
They're all trying to make the best solution, but that doesn't mean that they're all being homogenous. They all want to do better than everyone else, and sometimes that fails and what they do actually turns out worse.

The fact that there's no benchmark also means even the devs might not know what to "copy", unless they all have internal benchmarks constantly evaluating model performance on all competing harnesses and none of them have ever mentioned it publicly.
>>
>>109343971
>why do LLMs like japanese culture
because there's trillions of words worth of anime fanfiction and light novels published on the internet that all of these models have been trained on? Is that not self evident?
>>
>>109343988
The paper claims it's because every company wants to align their models and avoid controversial output, so it leans towards talking about Japanese things, because the
>thing
>thing Japan :O
meme is true. westerners just can't help themselves.
if your claim was true you'd have to account for the behemoths of culture that is america and china
>>
>>109343988
>Is that not self evident
That's like 99% of arxiv
>>
>>109343962
I mean yeah but there's academic research and there's "applied research" so to speak, and ultimately what Anthropic has done is actually built some really powerful models, and you can't do that just by publishing good papers with pretty graphs and impressive explanations. When people talk about "researchers" working a Anthropic or OAI etc., I assume those are mostly the applied guys, who can not just write a paper with big words but who can design, or participate in designing, an actual real model and design its architecture and guide its training and evaluate it and evaluate its usage etc.
Shit like setting prompts for task completion would definitely fall under the purview of that, since evaluating the model and optimising its aptutide at completing tasks are key requirements for actually guiding your research with data. And again we know Anthropic is provably good at doing this, which is why I'm surprised their near-flagship product is so poorly tuned.
>>
>>109341903
why consume stackoverflow answers if you could distill the same answers from textbooks? because stealing the work others did is easier and it's legal, just do it, it's fine
>>
>>109343988
It's self evident but now there's a PEER REVIEWED SOURCE for it. The researcher gets pats on the back and you get to put it in wikipedia with a nice citation without getting deleted for "original work". It's a win-win for redditors everywhere
>>
File: 1784751153872q.gif (3.39 MB, 400x349)
3.39 MB GIF
>>109343922
>>109343932
>come back
>it solved it
>>
>>109341906
This. Vibe gods don't care about bloat lol
>>
>>109343948
>all the people in the thread arguing about competing models and companies? minor brain damage
It's not brain damage, it's coping. A way to feel like you're still doing something as the model does all the actual work. Even where the harness might make a difference right now, it won't in a month or two when all the techniques that actually work have been posttrained into the next model generation.
Personally, I've already given up on any of my skills ever being useful again.
>>
>>109343988
there really isn't that much in comparison to western fanfiction
do you have any idea of what you're talking about?

and then the abstract of the paper says it likely does not emerge from pretraining, so it's not the dataset, it's whoever is supervising the training
>>
>>109344014
In that case why not do very uncontroversial countries like switzerland or finland?
>>
>>109344098
I guess they're so uncontroversial that there's no data about it.
If I was out there creating data for LLMs to feed on I wouldn't go to switzerland because that sounds expensive, and I wouldn't go to finland because I'm not sure what to do there besides jump into frozen lakes and sauna
>>
>>109343988
academics are fart huffing retards
>>
>>109344025
I think it's less about the model being poorly tuned and more about the harness just being shit... The harness has a very very high impact on the model's performance that I think a lot of people still downplay for some reason. Making the harness (something that is software and coding and not research or training related) is obviously not their priority which is why CCthe has had so many problems. They CREATED and have access to models that are used as THE standard for how good a model can be and get a bad deliver fuse to make a good harness it's out of spite because there's no excuse for that
>>
>>109341756
probably
>>109341765
>>109341834
10x–100x smart guys as the US and shameless copying will get you further than 10x–100x smart guys alone
>>109341903
not all of it. some of it is in destructively-scanned (slice the pages out of the book, then feed all the pages to an automated document scanner) books
>>109342479
WebP support so I can have lossless screenshots of my desktop in desktop threads in 4 MB
>>109342696
that’s a good name for what it does, though
>>
File: file.png (3.36 MB, 1059x6304)
3.36 MB PNG
>>
>>109343662
Nice, nice.
>>
>>109344156
>10x–100x smart guys as the US and shameless copying will get you further than 10x–100x smart guys alone
but it didn't. they're behind both sol and fable. their only claim to fame is being chinese and cheap.
>>
>>109344148
Ok, if the harness is so important, then tell me which tools or features are required for the model to perform better.
We don't care about harnesses because all the harness discussion is about how it makes the model magically superior in benchmarks without any actual discussion of what actually makes it better. You know the model only sees a system prompt and a list of functions and function arguments, right?
>>
>>109343951
There's too much info on the screen. This is sort of a me-problem, though, I admit. But if you could make onboarding just a little bit easier, it would be nice, methinks
>>
>>109343680
Try using Reasonix. It's probably more tailored for that given it comes from the Deepseek Team
>>
>>109343354
I hear it’s Opus-tier but twice as fast
>>109344168
they’d be even further behind if they didn’t distill too
they’re also bottlenecked by compute because we won’t let them buy the best GPUs available now and they’re still getting good at making frontier-quality GPUs
>>
>>109343714
Unless you can force the node that processes the request this could just be down to how quantslopped the serving node is
Also different system prompt means no prompt prefix caching which probably puts it on a different nodes
>>
>>109344190
on the topic of kimi, i will concede that. that's actually true
>>
>>109343966
>no access to kimi anymore
openrouter?
>>
>>109344190
uhhhh chink commie shills told me china was way ahead on compute infrastructure thanks to ~green energy
story changes every day it seems
>>
>>109344197
Also unless you can do greedy sampling you're going to get different results each time you try a task anyway
>>
>>109344205
nah. rift
>>
>>109344177
Is it really cyberpunk if there isn’t too much crap on the screen, though?
https://web.archive.org/web/20210730121632/https://zerohplovecraft.wordpress.com/2021/07/07/dont-make-me-think/
>>
>>109344098
these countries have no cultural relevancy
>>
>>109344098
>uncontroversial
>switzerland
>>
>>109344174
Where o you think that prompt and those functions and arguments come from, nigger?
Explain to me why Codex can just read files in the repo as it goes while Claude asks for permission every five seconds to run a "cat | sed | grep" command for reading out chunks of files. What do you think uses less tokens and is better for model performance, a well designed "read" tool call or having to construct an ad-hoc bash command to get file contents?

And yes obviously CC has a built-in "read" tool, always has, but where do you think the model gets its prompt that describes the tools to use and how to act? The prompt that is apparently so shit that Claude uses cat and sed instead of the built-in tool for whaetever fucking reason? Maybe that could have something to do with the harness, you know. Just a hunch on my side, though.
>>
>>109344319
>Maybe that could have something to do with the harness
That's down to the model generating the wrong tool calls
>>
>Dinitz-Garg-Goemans deboonked
>guy shares his chat
>find a breakthrough
>just continue until you find a broke
>yes, just continue
>ok, here's the counter example
FUCKING KEK. the singularity is approaching
>>
>>109344210
Chinkland has way more energy but way fewer GPU capacity. It's not complicated.
A lot of that energy is going into shit like arc furnaces for steel, and similar industrial shit, which western countries simply don't have. But if China were to build out datacenters to the scale of the US right now, they would definitely find it much easier to power them. That doesn't change the fact that their GPUs, while very competitive in the sense that they can actually do shit with them and Kimi exists after all, are still much more limited than the capacity available to the US.
>>
>>109344312
You'd find Venus controversial.
>>
>>109344344
is there any math being solved that will help my daily life
t. retard
>>
>>109344344
>he didn't use /goal
lmao what a luddite
>>
>>109344312
Unlike Japan?

>>109344296
Isn't it the whole point? They aren't really so unknown and yet they aren't exactly the daily target of geopolitical and cultural battles.
>>
>>109344319
So what is your point, that less tools = better?
Should the harnesses only offer a run_bash tool for optimal performance?
>>
>>109344356
No, unless you ask it. Mathfags are allergic to usefulness, they think it's for plebs, they just want proofs for utterly pointless conjectures.
>>
>>109344362
>not automating the conjecture selection process
what a noob (<--- doesn't know what a conjecture is)
>>
>>109344379
you need to touch grass and stop hanging out in localization seethe threads on /v/
out in reality everyone loves japan, it's the number one travel destination in the first world currently and this is on the heels of it building up its soft cultural power to immense heights over the past 30 years now
>>
>>109344385
No? My point is either the opposite, or more likely not a value judgement but just a demonstration that the harness can have a massive effect on model behaviour.
>>
>>109344409
wasn't your point that cc is the worst harness? how can you advocate for more tools while also thinking cc is the worst one?
>>
>>109344404
Koreans and Chinese hate Japan, there are like a billion youtube videos endlessly explaining how it's actually some hell scape for women or people working, suicides, weirdos with their sound on their cameras, fukushima, war crimes and so on...

I'm not trying to argue if these are true or not, it's more the fact that for me, the idea pushed by the paper doesn't exactly add up.
>>
File: 1768754026799596.png (273 KB, 3439x2060)
273 KB PNG
>>109344174
>Ok, if the harness is so important, then tell me which tools or features are required for the model to perform better.
The tools are typically appended to whatever system prompt that gets sent. What tools you need is dependent on what you're trying to do. The harness I use has basic stuff like code editing, web fetch, code difs, search commands, sub ancient delegation, etc. You can also add custom mcps so that, let's say you're using a model like glm that doesn't have vision support, you can create your own custom mCP vision server so that whenever you ask it to look at something, it asks a local vision model like Qwen vl to look at the photo, describe what it is, so that the big model can read that description and actually tell you. There's probably like 50 other kinds of tools that I'm completely missing but you get the idea. The thing is the more tools you have enable that ones come up the more bloated the system prompt becomes. If the model is explicitly trained to handle large contexts and long coating sessions than this isn't too much of an issue. However not all models behave the same. Some can handle those giant ass system prompts perfectly fine and others handle them noticeably worse (they don't break or become useless but they are noticeably worse than other models). We also know obviously that not all harnesses are the same so one harness could have a total combined system prompt that is like 10,000 tokens and others can use a lot less (pi comes to mind that one is pretty bare Bones compared to hermes or opencode and you're expected attached or create your tools yourself). I've seen anons here claim pi that's worked very well and performs better than all the others.
>>
>>109344448
CC forces the model to use less tools (it makes it rawdog bash), why do you think I think less tools is better?
>>
File: 1772861689288552.png (340 KB, 2778x1962)
340 KB PNG
>>109344174
>>109344481
I've never used pi so I can't confirm or deny their claims but I think they have some truth to it since you have very high control of what actually gets sent to model because you're responsible for giving the model a lot of the tools it needs, therefore you decide how concise or how beefy the system prompts of those tools need to be. Pic rel is a somewhat good example of what I'm talking about.

My question is only like a couple dozen tokens Max but then when I actually sent the prompt the UI tells me the token count is in the 10,000 range. Why? Because I have a web search MCP running. Whenever I sent a prop initially that mCP's system prompt gets sent to the llm so it

1) actually knows that it does actually have web access so it doesn't and correctly "sorry bro I can't browse the internet" like it would normally do

2) it gets a set of instructions telling it how to actually execute the searches.

The system prompt along with your prompt that the LLM actually receives will be like 10 to 12,000 tokens (plus whatever other tools you have enabled. Too many tools = potentially bloat). Then an expended another 300 to actually construct the formatted json string the search API requires in order to work. Then it receives another few thousand tokens or so that the llm has to process, (it will potentially do multiple searches if it deems necessary). Then when it finally completed the search and answered my question the context usage was over 18,000 tokens, around 7% of the model's context window you stopped from a simple search query.

The point of my giant wall of text is to say that how the harness utilizes the LLM system prompt matters a lot because that determines how useful it is and how it performs. The system prompts set stage for how it's going to behave throughout the entire chat session which is why how good your harness is is so important and it surprises me a lot of people either don't see it's important
>>
>>109344482
because it's the only harness that has more tools to begin with, and the other ones just give the model bash?
>>
>>109344449
they hate japan because they ain't japan
>>
File: 1764051797249579.png (969 KB, 1240x1030)
969 KB PNG
snailcats don't even know how these types of promptgods exist
>>
>>109344501
I don't know for certain what happens under the hood but the user output for codex is always "read file X, read file Y...". I assumed it had a read tool of some sort. Does it not?
>>
>>109344404
how does one become this delusional? manga and kewpie mayo?
>>
>>109344518
no, codex hides the shell commands and converts them into user friendly human readable events
>>
>lose 20% of my context just from loading skills
>>
>>109344546
Huh.
>>
>>109343740
>>109343858
Plan is a rank8/16 LoRA over 1-2 epochs on a few million tokens targeting shared modules based on mapped expert activations for security related tasks.
The RE dataset will be built from Ghidra artifacts and investigation traces from past engagements, and will be split into training, development, and evaluation sets. I'll run base Kimi through the same engagement workflow and mine the cases it misses, misdiagnoses, overstates, or investigates inefficiently. I will then convert those failures into training examples containing the raw decompilation/disassembly, relevant follow-up artifacts, the correct next Ghidra action, root cause, reachability, exploitability, confidence, and remediation, while adding patched counterparts and difficult safe examples to control false positives.
I'll be using a similar methodology to do the source code auditing dataset along with vulnerable code from CVEs and OSS-Fuzz disclosed vulnerabilities from after Kimi's knowledge cutoff date.
>>
>>109344569
anon do you regularly eat soap and glue?
>>
does clodex think templeos is kino?
>>
File: file.mp4 (3.57 MB, 854x480)
3.57 MB
3.57 MB MP4
>>109344164
Thanks anon, here's the quest so far
>>
File: 1758744483644517.jpg (104 KB, 2444x1366)
104 KB JPG
>>
i have no idea what unit tests even accomplish if the code is already right
>>
>>109344585
Why would I be expected to know that?
>>
File: 1759480016299445.png (1.95 MB, 3452x674)
1.95 MB PNG
>>109344481
>>109344490
The MCP server I was referring to
>>
>>109344612
how do you know it's right?

note that unit tests in the way tdd cultists tell you are indeed useless but actual tests exercising the behariour of your program aren't
>>
>>109344612
Tests give you something to complain about and an excuse
>CI is slow
>tests are running
>>
>>109344612
Unit tests are supposed to test if the existing features and your code actually do what they're supposed to do. They're not meant to find new bugs per say. They literally exist to make sure all pieces of the code don't fuck up when you actually use it. I've had to explicitly tell models to perform unit tests in the past because they'll confidently write code that's ALMOST perfectly functional but then the fuck up formatting or accidentally use existing dead code, so the unit test help the model catch and fix its own mistakes
>>
>>109344174
See >>109344622
>>
>>109344612
They are to stop regressions from happening.
>>
>>109344601
this is dope, but part of what makes cyberpunk's side missions so good is that a lot of them are sci fi as hell
like the braindance jesus one.
somehow I think you'd be hard pressed to make a quest that has as much meaning as many of these do without also importing custom assets
>>
File: 1784421852186877.png (2.65 MB, 2900x3508)
2.65 MB PNG
my brand new vibecoded ci/cd pipeline became too thorough so it consumed all of the free github actions minutes with too many checks, some of them duplicate. I had to buy github pro to get 1000 minutes more so i can continue vibing.

when i programmed by hand my pipeline was make, tar and scp but sadly it doesnt scale
>>
File: snailcat.png (1.52 MB, 1254x1254)
1.52 MB PNG
>mfw in a discord full of technical artists who think they're hot shit and hate ai
>mfw they post a question i always reply first with a solution and tell them codex found it
think i'll just set up a bot
>>
File: 1777419468313098.jpg (8 KB, 300x168)
8 KB JPG
>>109344582
Cool! Thanks for reminding me to poke around with Ghidra.
>>
>>109344701
You haven't really grown as a dev until you start burning github action credits.
>>
>>109344612
the first thing I reach out for when I write tests is a golden happy path. it's an automated way to check to see if the most common thing your application is expected to do still works when you change the code.
You can then branch out to other things.
Does this module, algorithm, or data structure work the way I envision when I expose it to edge cases? In other words, did the code prove that my mental model is fucked, or did my mental model prove that the code is fucked?

In the process of building anything and poking at it from different angles, you're likely to figure out your next steps or what needs to change

based on your question it sounds like you've never built anything complicated before.
>>
>>109344745
i've only vibed. the model is writing the code and the tests
>>
>>109344701
>too thorough
I've never come across a situation where I blew up CI/CD by testing.
The only time I blew up CI/CD was trying to compile a docker container in it. The ops team's response? It's actually just a testing service and isn't for building artifacts lmfao.

tl;dr you're doing something wrong.
>>
File: 1784156275001451.png (476 KB, 720x917)
476 KB PNG
>>109344546
>codex hides the shell commands
Yikes.... People actually like using a harness that does this? It's fuckups like pic rel happen semi often. What is being able to see what it's doing allow users to catch fuck ups way earlier before they even happen?
>>
>>109344768
what are you talking about retard github only gives you so many minutes
>>
>>109344792
yeah are you running the full ass CI/CD pipeline every commit?
claude tends to do that for solo dev projects
either use PRs or tell your agent to manually trigger on epics.
you're running that test suite 10 times an hour on your local machine anyways

and the tests that LLMs generate are trash anyways.
>>
>>109344709
KEK
>>
>>109344612
they make sure you don’t accidentally break something while changing something else
>>
File: file.png (1.58 MB, 1596x4572)
1.58 MB PNG
>>109344684
Yeah advanced quests like that would be difficult. I can probably turn some of the aspects like getting in a car into the building blocks that I want. At the very least that combined with the character creation tooling would speed up the process of making a unique/advanced quest
>>
>>109344849
You should be using git tracking for that too.
>>
>>109342972
they started user-base-maxxing after grok and chang demonstrated the value of large user-bases.
Also they gave free premium accounts to every teacher in the usa.
>>
>>109344856
unit tests and Git solve generally disjoint (non-overlapping) problems
of course, you should be using both for anything of nontrivial size
>>
>>109344633
>how do you know it's right?
it compiles. the power of rust
>>
>>109343028
Thank you! That's very interesting. I'll definitely be making the site AI agent friendly.

>>109343593
I'll look into it!

>>109344156
WebP support has already been implemented.
>>
File: 1783369210489897.jpg (89 KB, 1024x640)
89 KB JPG
>>109344319
>"cat | sed | grep"
good thing I took a month off vibecoding to learn to code, now I know what these words mean. Last month I did not.
>pic unrelated
>>
>>109343346
>>109343354
After some limited tests I get the same results with Grok and Sol Medium.. Looking good so far
>>
I have 15 hours or so to use up 85% of my grok tokens. What to build.....
>>
>>109344927
> 58. Fools ignore complexity. Pragmatists suffer it. Some can avoid it. Geniuses remove it.

Is there any removable complexity in this project?
>>
>>109344932
noted.
BTW the grok heavy sub is a lot of tokens, so you could use the free tier in cursor to try it out, or the super tier, just wait until the super goes on sale again. I got three months for $9 per month.
>>
>>109344771
it doesn't hide them, they just get collapsed into a dropdown thing after they finish
>>
>>109344973
>dsp: woooooooooow
>>
FUCK I spent all my $125 chink-subsidized agentrouter credits in a few hours because it didn't cache anything, back to free nemotroon I guess
>>
Need a Pi plugin that automatically switches between Sol, Terra, and Luna and nothing else
>>
>>109345111
Paste that to your clanker and you'll get one
>>
>>109341711
> We have information...
so... how about showing us that information?
>>
>>109345168
you dare question the government?
>>
interesting piece of trivia: if you batch download google trends data it normalizes all of it to the biggest trend, so if there is a huge trend all the other ones get squished down to nonsense. gotta download separately to get that good 0-100 data
>>
fable uses so much quota and is so restricted, that the slight edge it has over opus 4.8 doesn't matter, and i'd just rather use opus on ultracode
>>
>>109344869
Yeah I meant to say you should use both
>>
>>109341441
this general has been a thing for several months now

if ai is so good, why aren't you done your projects yet?
>>
>>109345260
We're all subscription plebs limited by usage rates
>>
>>109345260
know how long it takes to vibe build an operating system when you have a life+job on the side?
have some patience
>>
File: training.png (106 KB, 943x911)
106 KB PNG
>>109345260
right now my latest project is compute constrained
>>
>>109345260
shut the fuck up
>>
File: 1784324282571803.jpg (51 KB, 735x727)
51 KB JPG
I gained 20% of usage out of nowhere on Codex. I was at 10% left, now I am at 30%. What is happening?
>>
File: atlasgen.png (607 KB, 2316x1190)
607 KB PNG
Coarse map multi continent generator + terrain diffusion fine level map tile generation. Includes climate, temp, biomes, resource deposits, heightmap, glacial deposits, hydrology maps etc.
>>
File: tdtest.png (2.43 MB, 2048x1047)
2.43 MB PNG
>>109345281
Early results on terrain diffusion, currently training a new model on hydrology pairs for coherent drainage networks (right now everything just pools)

The coarse continental maps feed the terrain diffusion model to put out fine level terrain at 30m resolution.
>>
File: rendertest2.jpg (161 KB, 1759x968)
161 KB JPG
>>109345293
Tool output directly imports into unreal engine and is consumed by custom sdf voxel chunk streamer. No master material built yet, just testing. (Also hydrology still kind of fucked until the new model is trained)
>>
>>109345281
>>109345293
>>109345304
cool project
>>
>>109345159
No I'm trying to conserve tokens, I'll have to wait for someone else to do the work
>>
File: file.mp4 (3.59 MB, 854x480)
3.59 MB
3.59 MB MP4
Full quest
>>
There is a price correction happening, the big subsidies are on the way out.
>>
>>109345111
Like based on the complexity of your prompt or just randomly for fun or what?
>>
>>109345373
Yes the complexity, like the new cursor router feature but only for models I choose
>>
>>109345365
>he thinks the subscriptions are subsidized
>>
>>109345385
So you want a model router, there are several for Pi already. For example:
https://pi.dev/packages/@yeliu84/pi-model-router
https://pi.dev/packages/pi-smart-router
https://pi.dev/packages/@kdejaeger/pi-model-router
>>
>>109345399
Hey, at least you got a "reset"
>>
>>109344601
Where are you working now?

>>109344684
Baby steps, anon.
>>
>>109345281
I always wanted to make a continent generator, but the papers and software around plate tectonics were too impenetrable
>>
>>109345413
I'm not atm, in contact with a company who are interested though
>>
>>109345447
lucky bastard
>>
>>109345447
Interested in your mod? Or vibecoding efforts?
>>
>>109345280
you're probably seeing your actual usage, not remaining usage
>>
>>109345358
where's the luddite quest? You should make a snailcat model for that
>>
Use case for opus? I need to waste the near useless half of my weekly limits.
>>
File: house01.png (532 KB, 2182x1534)
532 KB PNG
>>109345358
does this game allow modding or are you hacking it somehow? cool either way. big token spend I assume?

I'm a codecel, just learning, making baby tier shit with grok. Got to start somewhere.
>>
>>109345511
listen to your heart
>>
>>109345427
That was actually the easy part. Since those are all essentially closed and shut textbook implementations, Fable one shot the macro generation pipeline. The had part has been generating meaningful structure at the fine bands (30meters and below). Been working on that across many experiments for 2 weeks now.
>>
>>109345467
Lucky if I get the job
>>109345482
Not the mod, because of different project
>>109345506
Soon:tm:
>>109345524
Cyberpunk has mod support, just the quest/scene stuff is not well documented
>>
>>109345402
oh cool, thank you
>>
>>109345555
>Not the mod, because of different project
Quads of luck. Good for you, anon. Also please don't stop working on the project. kek
>>
File: 1782649067483263.jpg (35 KB, 554x554)
35 KB JPG
What does your agents.md look like? Do you have other .md files?
>>
>>109343426
nigga I'm in favor of vibe coding, I'm shitting on what came before.
I'm way more inclined to build some website ideas via VCing than having to deal with the dogshit that was JavaScript framework retardation
>>
>>109345568
I just rawdog my shit
>>
File: agents.md.png (126 KB, 1101x941)
126 KB PNG
>>109345568
yes I have a shit ton of them
>>
File: 1766004027855405.png (142 KB, 1472x484)
142 KB PNG
Big props to Sacks if he manages to steer us back on course.
>>
Is 9router a scam?
>>
File: 1754088501817208.png (260 KB, 1156x1130)
260 KB PNG
>>109345568
>>
>>109345585
That will just preserve the status quo. Chinese companies are basically branches of the government that's why they can release open models. US companies are in it for the money and wont release anything worth using. If they'll do release anything nerfed just to say they released something they'll just be laughed at by the rest of the world.
>>
File: Capture2.png (361 KB, 744x1226)
361 KB PNG
It's not for 'making money'. Use it to create software that you have always wanted,.
>>
>>109345585
based. IBM should go back up now.
>>
>>109345611
>Chinese companies are basically branches of the government that's why they can release open models
God, how fucking gullible you are lmao
>>
>>109345633
why are you talking to yourself
>>
>>109345611
A lot of opensource research from China is from corporate/academic partnerships. Like Bytedance working with some uni in Beijing. There is no reason why this can't happen in the US. The government could provide grants even.
>>
>>109345611
There is factionalism in the corporate world right now between the AI labs and "everyone else". It is in the interest of essentially every s&p 500 company to see the labs "model as a service" business model fail and have the technology commoditized.
>>
>>109345633
I understand the skepticism but the chinese industry was moving toward closed weights until Xi said no, keep releasing the weights.
Alibaba already was not releasing them, Moonshot had already added support for encrypted thinking in its API.
These companies have their hand forced for geopolitical reasons.
>>
>>109345548
Did you try diffusion like those minecraft people did?
>>
>>109345639
academia has been wholly left behind in AI research
>>
>>109345658
Yeah I'm actually using the same training recipe, different data set.

It's essentially just super resolution with terrain maps.
>>
File: asd.png (428 KB, 1498x1084)
428 KB PNG
>>109343898
oh hi im a big fan!
>>
File: fuddsuy.webm (3.97 MB, 1190x996)
3.97 MB
3.97 MB WEBM
Bros, are we approaching real costs..
>>
>>109345680
best thing they can do is threaten to switch over to kimi
>>
>>109345680
If they pull that shit a third of the money is going to Chinese companies. Unless, of course, the US bans Chinese models.
>>
>>109345639
The US fascist corporatocracy doesn't like not being a tight knit oligopoly and controlling everything and they are currently controlling the US government (unlike China where the power flows in the opposite direction), so that wont happen.

>>109345642
I mean yes and no. I think the "smaller" or more blue chip companies in the S&P would benefit from AI commoditization, yes. But in the future if it does not get commoditized old industry knowledge and capital will become obsolete and whoever controls AI will end up controlling all those other corporations. And most of the people who own these companies are probably also investors in AI companies so AI costs are probably not that big of a deal for the biggest capitalists.
>>
>>109345639
https://www.whitehouse.gov/releases/2026/07/45502/
>>
>>109345680
I really don't understand how people are blowing their budgets like this
this literally never happens to me unless I let fable blow its load doing code reviews
>>
new
>>109345711
>>109345711
>>109345711
>>109345711
>>109345711
>>
>>109345704
>and they are currently controlling the US government
third world easterners are so fucking funny. ai companies are controlling trump so hard that he literally cock blocked their releases until HE gave them the green light, which made the us look worse and more authoritarian than china. what time even is it in russia?
>>
>>109345680
These retards are like the ones here using Sol Ultra for everything then wondering what's going on
>>
>>109345729
what's going on then, have there been no change
>>
>>109345724
>what is regulatory capture
Anon...
>>
>>109345358
What did you do to make it? Just point it at the mod editor?
>>
>>109341561
wtf does refactor even mean
>>
>>109346036
Deprecated term from before ai coding
>>
>>109343426
>can't read
most intelligent vibemutt



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.