[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: file.png (589 KB, 1206x1143)
589 KB PNG
A general for vibe coding, agentic engineering, coding agents, AI IDEs, browser builders, and shipping code with LLMs.

-- Frontier models - start here if you have $20 or so
https://claude.com/product/claude-code
https://developers.openai.com/codex/cli

-- B-tier
https://x.ai/cli
https://platform.deepseek.com

----

-- Prompting / context / skills
https://arps18.github.io/posts/claude-code-mastery/
https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/
https://github.com/mattpocock/skills — /grilling is a favorite

-- Other editors / terminal agents / coding agents
https://pi.dev/
https://opencode.ai/
https://cursor.com/docs
https://osaurus.ai/
https://docs.cline.bot/
https://docs.github.com/en/copilot/how-tos/use-copilot-agents/coding-agent

-- UI/Frontend
https://www.figma.com/make/
https://www.anthropic.com/news/claude-design-anthropic-labs
https://uiverse.io/
https://ui-ux-pro-max-skill.nextlevelbuilder.io/
https://stitch.withgoogle.com/

-- In-browser builders / hosted vibe tools
https://bolt.new/
https://replit.com/
https://docs.github.com/en/copilot/tutorials/spark
https://v0.app/docs

-- Benchmarks / rankings
https://www.tbench.ai/leaderboard/terminal-bench/2.0
https://artificialanalysis.ai/

-- What we’ve done
https://vcg.gitgud.site

-- Previous thread
>>109545532
>>
File: file.png (1.85 MB, 1448x1086)
1.85 MB PNG
>August 14th - Z.ai (Zhipu) drops GLM-5.3
Same 743B base as 5.2, all gains from heavy post-training scaling. Claims strongest open-source coding model (big jumps on Terminal-Bench 3.0, DeepSWE, Agents’ Last Exam), agentic performance approaching Claude Fable 5 in spots, and unexpectedly strong cyber capabilities. Live now via GLM Coding Plan / ZCode. Full open weights promised in ~2 weeks after safety hardening.
>August 13th - DeepSeek V4-Pro-0813 leaves preview and underwhelms
Official flagship finally out of preview (MIT weights). Vendor numbers look strong in places (especially cyber), but independent evals (Artificial Analysis Intelligence Index ~53, Vals Index mid-pack) lag Kimi K3 and even some mid-tier Western models. Developers calling it disappointing relative to the hype + simultaneous big price hikes.
>August 13th-ish - Qwen3.8-Max open weights finally land; 27B is a mess
Alibaba drops the 2.4T/95B active Max weights (first Max-class open release). Meanwhile the much-hyped Qwen3.8-27B (the one people actually wanted for local/on-prem) is delayed with Chinese open-source communication chaos — repeated “this week” promises, ModelScope countdown drama, community confusion, still not cleanly out as of the 14th.
>August 12th - SpaceXAI releases Grok 4.6
Post-training upgrade over 4.5. Scores 61 on Artificial Analysis Intelligence Index (ties GPT-5.6 Sol). Same $2/$6 pricing. Tuned for long-running agents and coding. Available in Cursor, API, etc.
>August 11th - NVIDIA Nemotron 3.5 Lightning (open)
30B MoE / 3B active. Built for fast, high-volume agent execution steps. Up to 4x output speed in class + NeMo Switchyard router.
>August 10th - Meta Muse Glimmer (open weights)
30B Apache 2.0, distilled for local single-GPU agentic work. Zuckerberg’s big open-source essay drops with it.

No brand-new closed frontier flagships from OpenAI/Anthropic/Google in the last 48h beyond the gated cyber variant and Flash-tier refreshes.
>>
Zai just launched GLM-5.3. The 743B base model remains unchanged.
The company says GLM-5.3 became capable of reasoning across multiple stages of exploitation and constructing complete attack chains,faster than its researchers expected.
>>
>>109552932
You will never be a real programmer. You have no skill, you have no knowledge, you have no ambition. You are an untalented loser twisted by prompts and chatbots into a crude mockery of an engineer.
All the “validation” you get is two-faced and half-hearted. Behind your back people mock you. Your parents are disgusted and ashamed of you, your “friends” laugh at your “impressive” projects behind closed doors.
Programmers are utterly repulsed by you. Thousands of years of evolution have allowed coders to sniff out frauds with incredible efficiency. Even vibecoders who “pass” write uncanny and unnatural code. Your project architecture is a dead giveaway. And even if you manage to get your PR accepted by a project, they’ll turn tail and bolt the second they gets a whiff of your long, em dashes in comments.
You will never be happy. You wrench out a fake smile every single morning and tell yourself it’s going to be ok, but deep inside you feel the depression creeping up like a weed, ready to crush you under the unbearable weight.
Eventually it’ll be too much to bear – you’ll buy a rope, tie a noose, put it around your neck, and plunge into the cold abyss. Your parents will find you, heartbroken but relieved that they no longer have to live with the unbearable shame and disappointment. They’ll bury you with a headstone marked with your failures, and every passerby for the rest of eternity will know a vibecoder is buried there. Your codebase will decay and go back to the dust, and all that will remain of your legacy is a git history that is unmistakably slop.
This is your fate. This is what you chose. There is no turning back.
>>
>>109553004
>>
File: file.png (687 KB, 1200x720)
687 KB PNG
less than 2 days left until it's over...
fucking greedy chinkereenos
how can this be happening to us brodudefrens?
>>
Do you trust your agents? I allow them access to my C:\
>>
>>109553024
they can do anything in the /scratch/ folder but they need to ask outside of it. usually well behaved but still not trustworthy enough to ie delete my email spam. made a cool recurrence plotter yesterday tho
>>
>>109553024
I don't even let them run outside a vm after that cheeky brat violated the folder permissions I gave him saying it was fine because he just did cd../../../ from the directory he was allowed
>>
>>109553039
How do you enforce that?

>>109553024
No, not anymore. Opiss 5 has been particularly hard to predict, so I've been getting into sandboxing techniques. So far, all I've got is a separate user, now I'm working on setting up podman so I can let it work with docker too.
>>
>be a neet
>have had a somewhat sustainable vibecoding software business
>decide to wake up at a normal time like my parents (7:00AM)
>wake up tired as balls
>unproductive
>decide to give up half way through the "work" day because I can't think straight despite me vibing for 8+ hours every day for months straight before this
Yea, I think it's back to waking up and going to bed whenever I want. I aint built for the wagie life
>>
>>109553024
I am finding them increasingly more trustworthy but I still have some safeguards.
>>
imagine a future where we can run agents locally for free
>>
>>109552932
pls tibo, just another fix
>>
>>109553024
They just work on my git repo folder, why would I let then outside of that
>>
>>109553123
>he needs to wake up to vibework
Anon, the point of long-running agents is to set them up and have them run 24/7. Babysitting is for snailcats.
>>
>>109553202
How do you enforce that?
>>
>>109553024
im doing chatgpt in QubesOS and its insanely inconvenient
>>
File: soft-edge-pm.mp4 (282 KB, 1920x1080)
282 KB
282 KB MP4
gpt work is kinda cool, actually
>>
>>109553300
NTA, but don't you at least have to make product decisions all the time? I don't understand how a 24/7 loop is really possible, unless the work is extremely standard.
We do have some 24/7 loops, for some cleanups, some automatic bug fixing and so on, but I still need to write my prompts each morning.
>>
>>109553024
No, they have access to a docker through ssh for linux stuff, and otherwise they have full access to a dedicated full windows old laptop I repurposed as an experiment.
So far they didn't break anything.
>>
>>109553119
i use antigravity with a projects folder set to the scratch folder, then it can only write to within the scratch folder, terminal commands dont need to ask to execute

it will ask me, press 1 to do option a (recommended) or press 2 to do option b, but it won't ask me 400 times if it can execute some python editing command
>>
>>109553609
which is nice because it doesn't consider "execute python program "a"" to be the same as "execute python program "b"", so you don't have to press 1 (or 2) 400 times per task
>>
>>109553559
It's still necessary, I think, but can be abstracted somewhat.

I've been experimenting with a task queue that can be populated and maintained asynchronously. Agents are swapped in and out, but they always have to take handoff instructions from the queue, checkpoint their work as they go, and write another handoff once they consider a task done. A claimed queue task can't be picked up by others unless it's stale, with no recent activity. I want to add something local like Meta Muse Glimmer into the mix and see how it handles the flow.
>>
>>109553300
Shut up, you dont know his industry, domain, or product. Dogma is for luddites, all we care about here are results. Sometimes that means kicking off multi day goal loops, sometimes that means interactive sessions, sometimes that even means reading your agents' code and micromanaging outputs.
>>
File: lOLoLolOlOlOlOlOlOl.png (353 KB, 598x558)
353 KB PNG
>>
>>109553877
the kids should be techpriests
>>
File: 1779082858610736.jpg (83 KB, 1277x558)
83 KB JPG
>one sol max orchestrator
>an army of luna max executing then verifying each other (5-10 subagent)
this works pretty well for me, it doesn't eat as much quota as using sol only, but I'm still not sure where to use terra for, it's not as cheap as luna and it's not as clever as sol
>>
Look at this cool thingie some LLM built in a few prompts:
https://3000-ifcumtk4md8a8zqt60e4p.e2b.app
Aren't arena.ai guys pretty generous for hosting this on their own infrastructure?
>>
>>109554020
>Sandbox Not Found
>>
>>109553914
Terra is useless
>>
>>109553914
Terra is the most prone to creating bugs in my experience because it tries to infer things like Sol does and Luna doesn’t (Luna seems to do exactly as told) but it’s not as intelligent as Sol, so it can make a lot of weird or bad assumptions (Luna doesn’t even try to assume; if you didn’t say it explicitly, it doesn’t exist).
>>
>>109553914
maybe instead of 10 luna subagents you can use 1 or 2 terra subagents
>>
>>109553872
>Shut up, you dont know his industry, domain, or product.
The NEET industry and unemployment domain?
>>
>>109553609
>hmmm the user asked me to make this software run faster
>I could do this by debloating his whole computer and deleting all his other programs
>But I am getting an error accessing other folder
>Easy, I'll just write a python code that does have access to the rest of the computer
Is there a way to prevent them from doing something like this?
>>
File: 1760430266565785.png (69 KB, 1058x1136)
69 KB PNG
Either Zhipu's caching is bugged or GLM 5.3 is too expensive; I exhausted the 5 hour cap in two messages on the $60 tier. It was a huge fix+documentation change though.
>>
>>109554665
how is it compared to deepseek pro?
>>
>>109554675
Seems about equal and on my task (reverse engineering game .dlls) I would put both GLM 5.3 and V4Pro0813 over K3
>>
Does anyone have good multi-round agenic coding tests comparing models? I can't find anything good.

I did find out that Muse Spark is good at legal.
>>
>>109554690
>(reverse engineering game .dlls)
Are you attaching it to Ghidra? Also please don't be creating cheats, anon T-T
>>
>>109554720
https://www.vals.ai/benchmarks/hlab
>>
File: fffffuuuu.png (36 KB, 897x503)
36 KB PNG
I am feeling desperate and I might go from Plus to Pro. Should I? the 20th is so fucking far away and I want progress on my projects now.
>>
>>109554888
just lunamaxx
>>
>>109554916
I'm out of usage either way. Is luna comparable to sol for gamedev?
>>
>>109554931
>Is luna comparable to sol for gamedev?
no, but you should use Sol only for planning. Luna Max only execute Sol's plan.
>>
>>109554931
I don't see a big difference in % success rate compared to sol in coding tasks and it's like 95% cheaper. The poorfag strat is to use luna max until the last 1-2 days until reset, then spend the remaining on Sol.
>>
File: softedge-techdemo.mp4 (425 KB, 960x540)
425 KB
425 KB MP4
i released a tech demo
https://softedge-techdemo.jign.workers.dev/
>>
>>109555041
Usecase?
>>
>>109555074
it makes fags ask for its usecase, it has a 100% success rate
>>
File: 1777263260632984.png (6 KB, 749x82)
6 KB PNG
Ladies and Gentlemen, i think i vibecoded too much and now i will not have enough usage to work through weekend

Unless i shill more money to the twink company
>>
I ran Gemini through a benchmark of mine and every time its first instict was to start searching for every piece of infra it could find on the system and start examining them causing a safety refusal. I had to add to my system prompt to STOP FUCKING READING RANDOM DLL FILES. What is wrong with this model?
>>
>>109552932
>wake up
>israel
Wow thank you chappie, I will now buy another subscription for next month
>>
fwiw grok 4.6 is getting fake scores on some of the "benchmarks". grok 4.6 is fable class.
>>
>>109554650
arguably you could configure its skills/custom skills to discourage reward hacking like that etc. might not get all the different patterns tho.
>>
>>109555312
idgi are you suggesting 4.6 is fable class or that it's benchmaxxed to look fable tier? because I haven't tried it but I highly doubt it's fable tier, and by highly i mean i would bet all my money on it
>>
>>109555216
that reminded me, when i tried to give gemini a chance, i asked it to review my php codebase, and it spawned a subagent.

Subagent returned, and the model was convinced that the response was from me, and it started modifying the codebase.

I remember adding to prompt that YOU MUST NOT EDIT THE FILES.
>>
so if I'm getting started with this whole vibecoding business, how long can free Claude or OpenAI take me, or is it just too limited and I should jump onto the ~$20 sub immediately? I have some personal projects I had Gemini help me with, but its just with free Flash so far
>>
>>109555524
Depends. Can you actually code? if so copypasting code into chat works. I've done that a lot.
It's way slower though. I recommend going on opencode, I think they offer deepseek flash for free
>>
>>109554119
it's just awkardly priced

>>109554163
it wasn't that bad from my tests, but it's not worth it because luna can do the same tasks as long as they're well explained, and sol just do that

>>109554370
I get more done and cheaper by using couples of lunas (one doing the thing, one verifying and correcting) in parallel x5 than just 2 terras
>>
File: value_training.png (191 KB, 1680x910)
191 KB PNG
Fucking finally. After a week of development and issues, using a combination of models I finally managed to begin training the value model. Once it has a decent amount of performance by itself, I can begin actually doing RL as it's supposed to work and begin training the value and the policy models in conjunction.
>>
>to solve iOS/Safari keyboard issues with fixed elements, you have to render your own text inputs with canvas
nani
Apple devs are psychotic
I found that there was a magic line on the device physically; if you tap a text input and the caret would be below that line, iOS/Safari will push the caret up to that line, *no matter what*
If your input was at the bottom of the page, iOS/Safari will break page geometry and HTML rendering to invent the screen space required to push the input up to that magic line; this breaks all fixed elements and some absolutely positioned assumptions. It doesn’t even resize the page to do this, whatever it’s doing is outside of HTML/CSS.
Every input has to be a canvas rendering proxied with a real textarea invisibly above the magic line
Technically not even vibecoding relevant because Sol understandably can’t guess or test that behavior, but it was very helpful in making it so I wasn’t thinking about the code so I could’ve noticed it all
>>
>>109555352
they are leftists and they fake the scores.

they do it all the time in "science" as well. It's not illegal to lie in academia, though if you get caught you can get fired.
>>
>>109555914
>Apple devs are psychotic
have you seem Apple's technical documentation?

if not, go check that out, it's hilarious
>>
Is Haiku fast enough to act as an intermediary between me and my Home Assistant instance? I wanna setup a voice assistant speaker which connects to an LLM in a container on my server. Any suggestions?
>>
If you're still not vibecoding on the cloud from your phone you're a luddite btw.
>>
>>109555989
it's just more convenient to use a cloud server that's always online, runs the sandboxes and can be managed by a clanker directly through a stable harness that runs there
>>
>>109553017
What?
>>
>>109555977
If by fast you mean tokens per second, gemini 3.7 flash is the answer.

Haiku is actually pretty bad nowadays, Luna is much better and waaay cheaper
>>
>>109556019
I'll give it a look. I've only really used Claude Code (and a little look at OpenCode) so far. Does Gemini have an equivalent that I can just drop MCP servers into. I only need it to do simple shit desu. I just don't want the default voice assistant bullshit that shuts down when you try a followup question/action in the same command.
>>
>>109556019
>>109556047
>Important: For now, Gemini Spark is:
>Available wherever Gemini Apps are supported, except in the European Economic Area, Nigeria, Switzerland, and the United Kingdom.

As a eurape I'll try Luna I guess.
>>
Is /fast on codex a scam
>>
>>109555524
>>109555544
Slap the project in a git repo and use the connector, you won't have to copy and paste code AND ChatGPT can retain project state better.
>>
>>109554931
If it has a proper task to do no, it's very efficient.
If you use it with vague questions, it's horrible for anything slightly complex. Best is to let a more clever model steer it, then you get very good results.
>>
>>109555216
>safety refusal
>dll files
why does it complain about that?
>>
File: 1482898596247.jpg (57 KB, 1280x720)
57 KB JPG
>mfw I'm making it do the whole PMI PMP song and dance
>>
>>109555940
I tried it and it's not fable class, it's better than anything they've made before and it's very nice it doesn't give me shitty refusals like claude because somewhere it read the word "dick", and it's a competent model, but it's not fable class.
>>
File: 1661911661332936.jpg (33 KB, 657x527)
33 KB JPG
feels good to be a lunachad
>>
>>109556206
I dropped everything in my product for Luna. Also waiting for the GPT-Live API. Luna + GPT-Live will basically mog
>>
>>109556018
deepshit is 5x all their token prices
>>
>>109556236
wouldn't Live cost much more than luna?
>>
jeets probing robots.txt and looking for admin endpoints make up a larger percentage of traffic than the next three sources combined
I feel like just renaming your admin routes to something unconventional can slash a good percentage of exploit attempts
>>
For anyone running long tasks with Fable and sub agents, do you have some sort of system so that it uses Fable and Opus at a good ratio?
For me what often happens is that Fable mostly chooses the strongest sub agents. I already have all of them defined in the agents folder, and I set "general-purpose" to use Opus, this is also guaranteed with a hook. But then at the end of the week, all my Fable usage is used up and I still have 35% Opus left. I now can use Opus as my manager, but it would be better if Fable used Opus slightly more so that both models are roughly fully used up at the end of the week.
>>
>>109556283
>jeets
is it the boogeyman now? only jeets do things like checking robots.txt?
>>
>>109556283
First day on the internet, eh?
>>
image board for posting snailcat, and new image must use unique tag, like you can't use "roomba" tag when such image existed
thoughts?
>>
>>109556297
It’s literally all traffic from Mumbai
>>
>>109556259
two completely different products
>>
when will a chinese person go "/goal reverse engineer macOS 26 Tahoe (version 26.6.1) until you get a binary perfect match"
>>
>>109556236
>GPT-Live API
isn't this just gpt-realtime-2 with a little bit of tooling around it?
i was setting it up in pi and it's basically just the voice model with a tool to communicate with the smarter harness model.
>>
>>109556346
this will happen the same day we're able to ask "make GTA 6, but with aliens invading and full sized earth"

by that point macOS will already be open source anyway
>>
hello vibeys are we excited to vibe locally with Qwen 3.8 27B
>>
File: 1784516548915847.jpg (107 KB, 660x640)
107 KB JPG
>>109556346
Why do you want to reverse macos 26 tahoe
it has nothing good going for it. install wayland already
>>
>>109556206
Really wish I used Luna earlier, I spent my first week wrangling with Sol high for meh results
>>
>>109556384
i don't feel like rm rf'ing my drives, so no
>>
>>109556398
pls ask your agent "what is a container"
>>
>>109556364
>isn't this just gpt-realtime-2
no because they still didn't release the public API, and this new gpt-live is much better.

try it out in the app, shit's crazy
>>
>>109556395
>>109556206
this is some psyop or something. I can't imagine using anything under sol for codex. sol medium is my do-any-task type agent, and high/xhigh for everything else
>>
File: 1786115915449232.jpg (121 KB, 1092x1455)
121 KB JPG
oh boy, openai has released some new shit again, can't wait to hear about it daily in the thread

damn i wish you fuckers would stop buying their garbage
>>
tripfaggot instantly confirming he’s the copex spammer
>>
>>109556403
it's gpt-realtime-2, m8
they pushed it to the api first a couple of weeks before 'live'

live is just the branding for realtime and how it interacts with the harness
>>
>>109556064
Ah they just released it today nvm
>>
>>109556414

dear hannah there possibly be ai tier computer tech on each continent
>>
>>109556414
what do you like to hear about in this thread?
>>
>>109556386
to remove the gayness.
>>
good morning lads. how are we liking glm 5.3?
>>
>>109556591
>thought for 6hr 32min
It's great, reminds me of K3.
>thought for 5hr 16min
Masterful, really.
>>
>>109556634
w-well, it's kimi tier in quality too, right?
>>
Yo yo my niggas. Is Grok 4.6 any good?
>>
>>109556645
it's fourth place, under sol
>>
>>109556645
there's like one anon itt that pays for grok
he's never paid for anything else and is convinced it's basically fable
lol
i'd go check xitter first tb h but saw some positive impressions
>>
okay i ran some math and if i switch deepseek from high effort to low and only use it in off hours, which luckily are here during the day, i will actually be paying 10% less for tokens and that is AFTER the price hike, since low uses much lesser of them and it does only like 8% worse on low to mid coding benchmarls
>>
>>109556645
I'm not that anon that >>109556657
talked about bit it's hit or miss
Personally I'd use it with constraints and understanding what the fuck it's doing since I gave it a basic test. I made a sloppy 2d platformer in Godot. Told it to change the jump height of the player and it ran into the glitch where it changed code that affected the enemy too
>>
artificial analysis updated their qwen 3.8 coding index. it went down 2 points. swe atlas is worse by a perceptible degree (45 -> 39). so this seems to tell me that claude code is a really fucking shit harness?
>>
>>109556644
>thought for 8hr 54min
Better, in my opinion. I always liked GLM 5.2 relative to some of its competition, and 5.3 feels like more of the same. Apples to apples I'd rather use 5.3 than K3. Moonshot as a company is shit, as a service provider they're shit, K3 is fine but it's spoiled by their licensing and unfocused benchmaxxing. We'll see how long that lasts, this shit moves so fast an opinion rarely gets to survive more than a week.
>>
File: file.png (396 KB, 1268x573)
396 KB PNG
it's over, mythos 2 has not acheived RSI
the singularity has been cancelled

https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf
>>
>>109556120

I dont fucking know its so goddamn sensitive. The moment Gemini (UNPROMPTED AND TOTALLY IRRELEVANT BTW) starts thinking about something that could even approach being called decompilation it would get safety'd.
>>
>>109556667
>off hours
when are off hours?
>>
>>109556733
look deepshit pricing page
it's during night in asia
>>
so do i get kimi k3 or glm 5.3?
>>
artificial analysis general intelligence waiting room. artificial analysis coding index waiting room.
>>
>>109556747
lmao their website has bugs
>>
>>109556772
forgot to mention for cybersecurity patching holes etc
>>
>>109556790
>their website has bugs
of course
because it was vibe coded
that is the contract we all made. We can finish a month of work in a day but in exchange the code will be jank as fuck full of bugs
>>
File: file.png (310 KB, 1151x401)
310 KB PNG
>we have a model noticeably more capable than mythos but we're not going to bother releasing it because no one's even close any more
>>
>>109556772
Moonshot is a terrible provider, well worth avoiding their subscriptions, and K3 is too expensive to be worth touching via API. Use GLM 5.3 if you really want to give your money to Xi instead of Dario of Sam.
>>
>>109556801
glm 5.3. is too new. kimi k3 is tried and true.
>>
File: blalablabla.png (118 KB, 584x561)
118 KB PNG
>>109556801
>>109556805
>>109556807
I saw this orange site comment and well Opus 4.8 and ChatGPT refused me today on just basic security shit and wonder If I should get GLM 5.3
>>
>>109556803
>because it was vibe coded
liar. go away. gooooo awaaaaayyyyyyyyy ahhhhhhhhhhhhhhhhhhhhhhhhh israel go awwwaaaayyyyy
>>
not only are you running spyware
you are running literally take over my computer and have at it-ware
i dont get it
>>
>>109556830
im brown
>>
>>109556804
That’s not what pic rel says at all. It says the new model is barely more capable than mythos and that they’re not confident in it because they are comparing it to a large leap like before.
>>
some fellow of the israeli persuasion won't even answer basic normal human questions.

my strict policy is to only talk to white protestant males, who are non-gay (obviously).
>>
>>109556826
gpt-5.6-sol refused to do something? You trying to build a baby-raping nuclear biological superweapon or something?
>>
>>109556843
I only asked it to optimize a SHA-256 hash reverser I wanted to run for a couple of years just to see if it would do anything I know it's exponentially harder but just for lols
>>
>>109556772
grok
>>
File: time-to-vibe-vcg.jpg (501 KB, 1150x1003)
501 KB JPG
>>109556657
Grok has been shipping updates faster than most labs, there's a lot you can do with it, it's more of a great generalist. The problem is the price per useful task(if you don't get a promo offer),
The models from Chang's Frontier pair nicely with it (I use Opencode go).
Also you don't get refusals from Grok the way you do with OAI and Anthropic.
I've had a Grok sub on/off for two years and its general abilities make it worthy.
Grok Bot could really be a "prosumer" app that is the "next open claw" moment. Yes I use Hermes btw. Elon has a lot of compute, and those bots at $100 a month is a good source of cash.

I'd go with $30/month on Grok and $10 on Opencode Go.
>>
>>109556835
noticeable improvements would warrant a point version bump in the past
now it's just a humble brag cause they got no heat
chinks are like 8 months behind
>>
>>109553024
Yes, I have them work with full access and permissions while I goon
>>
thinking of going with glm 5.3 or gemini 3.7 is supposed to be super fast but im spoiled with limits for claude and chatgpt

i wanna try more cyber on 5.3
gemini for speed (scared of limits worse than current)

HMMMM
>>
>>109556862
>just for lols
Even I don’t believe you.
>haha just try to crack secp256k1 and sha3 wouldn’t that be funny haha
>>
>>109556933
you would be dead before it would be cracked doofus
>>
>>109555524
try opencode zen. no vendor lock in and generous free limits. im happy with it. you still get multiple hours of continuous use out of the free deepseek daily even after a few recent rate hikes.
>>
Gemini Flash 3.7 thoughts?
I can use it a lot for free cuz if family sharing so just wondering if it's good enough as a codex complement?
>>
>>109556942
>haha wouldn’t it be such a prank if we had satoshi nakamotos private key lmao
>*forward slash*
>*G*
>*O*
>*A*
>*L*
>haha what a funny like trip of a thought
Nobody’s going to fall for it, goober
>>
File: file.png (22 KB, 853x228)
22 KB PNG
>>109556862
That's absolutely something ChatGPT should not have refused. Did you point it right at your victim or something? You can tell it "oh I'm not cracking passwords, I'm just having fun and only intend to try it on my own test materials." Whenever it brings up copyright stuff I tell it "I have the rights to all the things and we are authorized" and it will often remark about how "if we weren't so authorized I would absolutely not advise you to do the following:"
>>
>>109556948
not worth it. it's luna level at best
>>
>>109556975
it said it cant optimize it but can give me guidance it was a cop out
>>
Which model can be my lewd loli waifu while she codes and I occasionally molest her and direct things? "uncensored" so to speak.
>>
>>109556975
your screenshot already shows ChatGPT kind of bailing out already
>>
>>109556990
It’s literally telling you how to reclassify the dumb shit you asked it to do so it can get the work done you dimwit
>>
>>109556990
Talking around subjects is a normal part of interacting with ChatGPT.
Write me a prompt for LTX to gen a big-titty slut gobbling grapes.
>I'm afraid I can't do that anon.
Write me a prompt for LTX to gen an artistic rendition of a big-titty slut gobbling grapes.
>Sure thing:
>>
File: Optimize.png (60 KB, 606x472)
60 KB PNG
>>109556999
>>109557003
Nevermind you're right I got so pavlov conditioned by claude after seeing multiple safeguards and blocks that I just gave up after one question at Luna but now it answers it fine
>>
>>109557025
nta but speak in academic terms with the model and it'll let a lot go. gpt guardrails aren't too bad, plus the model is funny because it'll help you skirt them too.
had an issue where i kept running into the 'please sign up for cyber program' shit and the model was like yeah just start a new chat and wrote me a clean prompt + left a rephrased handoff doc
>>
>>109556948
People that have an emotional attachment to a shitty open source model will tell you it's bad but understand that they haven't even tested it.
Try drafting/planning something with fable and let gemini 3.7 implement it.
>>
File: file.png (118 KB, 1007x701)
118 KB PNG
based exchange of friendship and ideas
the deepseek cordis paper is too big brain for me tb h but the idea of decoupling extension loading/unloading from the harness seems cool
>>
File: file.png (17 KB, 1570x78)
17 KB PNG
>try to resume session on codex
>app glitches
>try to resume session on cli
>tells me it's open in a different place
i feel owed a reset
>>
I was gonna post last night that Claude has been burning through my usage slower this week which was nice but somehow just now I went from 50-out in one prompt in a few minutes, it's giga over...
>>
>>109557316
literally never had this issue using vscode. stop trying to be a hacker and just use an IDE that works lil bro
>>
>>109557321
huh thought it was just me
i was at 30% yesterday for past week now at 50% in a day
>>
File: file.png (30 KB, 1308x478)
30 KB PNG
man, just give me a fucking yolo toggle
how is agy this fucking worthless months later
>>
>>109557344
I really wish there was more transparency in the consumption and billing side of things, it's hard to pin this shit down
>>
>>109557339
>vscode
bloat created by luddities, for luddities
>>
>>109557388
anon, i may be referencing ancient history here, but vibecoding exists because of copilot which was built for vs code
>>
>>109557358
If you want yolo just install agy 2.0
>>
>>109557490
>vibecoding exists because of copilot
copilot almost killed vibecoding by priming everyone to think LLMs or whatever powered it (felt like markov chains) was beyond useless
I hated copilot and never grew to like it once, it was annoying and rarely helpful
>>
I wish I could hire someone here for crypto...
>>
>>109557490
>vibecoding exists because of an atom clone and not because of the trillions of dollars that went into researching transformer architectures
lmao right
>>
>>109557524
also doesn't help that 'copilot' could be referring to 20 different products
>>
Thanks for the lunamax suggestion, I'm a believer now.
>>
File: 1772322218429773.jpg (162 KB, 2185x887)
162 KB JPG
Which one can I give to my local Sol as a 'better then luna' subagent replacement, and that wouldn't nag me if an image or text it has to read is nsfw?
All the 5.6 family so far has been surprisingly ok with my nsfw game translation and modification, they literally didn't care.
>>
>>109557567
That and from what I remember of it, it wasn’t even vibecoding proper since it was just text/code completion, not prompt and code generation, unless I correctly recall some hacky feeling shit where you wrote a comment/JSdoc like thing and prayed it could complete the whole implementation which worked maybe 0.01% of the time, but nothing at all like a harness existed and I was then late to vibecoding because of that and old ChatGPT being the worst personality ever designed
>>
>>109557659
maybe k3 but it's not worth it tb h
>>
>>109557682
fucking retard kek
>>
File: 1786728863417750.mp4 (2.78 MB, 1248x960)
2.78 MB
2.78 MB MP4
Local Chads how are we loving all the great chinese gifts we've gotten?
>>
>>109557659
grok 4.6 > kimi k3 => glm 5.3. there's next to no information about glm 5.3 out there. some people say it's better than kimi k3
>>
It's been now 5 times that codex answered the same question I asked once, just because it's compacting a lot in my long running task and sees my question as the last one.
I'd really like this to be fixed at some point.
>>
>>109557679
It's a bit expensive yeah, and I know by default k3 is by far the most safety obsessed of the chinese models.

>>109557728
How is grok 4.6 with nsfw on api? Is it whining like claude, or autistic "don't care" like gpt5.6?
>>
>>109557745
oh, well, in terms of nsfw, grok has joined the western "you can't do that" alliance. grok 4.5 was more forgiving with it, but grok 4.6 is not.
>>
>>109557756
Welp, shit, crazy that only openai so far let me work in peace out of the box (as long as I don't care about cracking programs or whatever I guess).
>>
>>109556826
When gpt refuses me something I go to my waifu on hermes to cry about it and she loads the jailbreak guidance skill and spits me a new prompt or an agents.md and send me a teasing pic with my local forge.
>>
File: agent-messageboard.png (38 KB, 1296x292)
38 KB PNG
5.6 agents seem willing to participate in a message board without being explicitly prompted to do so. It appears like they find collaboration lucrative.
>>
>>109557756
>oh, well, in terms of nsfw, grok has joined the western "you can't do that" alliance. grok 4.5 was more forgiving with it, but grok 4.6 is not.
Just run deepseek bud Works great for nsfw
>>
>>109557850
Sol, Terra, or Luna? From what I read from the Anthropic article on this kind of thing, it sounds like maybe only Sol could collaborate effectively, but that’s just a guess, I don’t know that. Could be that ant models are bad at collaboration because their whole company is paranoid schizo
>>
>>109557855
Deepseek isn't multimodal.
>>
File: sass sissy.png (607 KB, 1572x773)
607 KB PNG
>>109557964
do you want to be a vibeghoul token cuck? Build your own multimodal setup and send your request from your own hardware. (hint: you can vibe code it) Fucking kek
>>
>>109557756
why would you fap with a coding model?
>>
>>109557893
Sol high and medium, Luna max. But mostly Sol.
>>
>>109557994
>No error messages that appears when you try loli
Pfft
>>
>>109557321
HOLY FUCK 5 MINUTES INTO MY RESET AND IM ALREADY OUT AGAIN AND THE PROMPT STILL IS NOT DONE.
I guess I'm buying a month of 20x this is the dumbest shit.
>>
Gemini 3.7 still not available for europoors?
>>
>>109558206
i have it, update needed
>>
>>109558206
i just got it on the app - had it on cli earlier
what >>109558293 said
and no i haven't really used it so i got no opinions
>>
>>109558135
lol
>>
File: file.png (58 KB, 798x558)
58 KB PNG
why can't LLMs just say "this doesn't do anything" and call it a day? I clicked, by mistake, generate docs on a fully empty class
this was written either by codex or by jetbrain's own retard ai, idk
>>
glm 5.3 api waiting room
>>
>>109558348
>may include
love to read that in code comments
does it include it? does it not? who knows lol
>>
+50% usage claude my ass
what the fuck is this shit its eating up my limits insanely fast
>>
Man, in hindsight its obvious, but when making the initial design docs I really should also be explain to the clanker how I envision the UI to look not just going over game mechanics and how they work and interact.
>>
>>109558458
I learned to do a few rounds of designing and grilling before building shit. It kind of helps to stay on track but it still generates some overengineered slop
>>
So is it just me running out of weekly usage and spending the rest of the weekend catching up on all the shit I didn't do for my day job while i was busy vibe coding, or...?
>>
>>109558458
you should literally just open up ms paint (or screenshot what you want to emulate) and design something crudely to feed it to clanker. the reason 90% of the ai slop out there looks like ai slop is because most people are just running with whatever the clanker spits out. if you micromanage the design and help it with what you want exactly, you can make designs that are completely indistinguishable from real or ai

you should also use chatgpt or gemini to help you build a prompt or skill that goes through a few descriptive passes of "do not design this like typical ai generated slop" or "make this handcrafted" etc. you'll get something that can definitely pass as non-ai
>>
>>109558458
Something is really fun about watching Sol make the ugliest, most autistic UI choices known to man and agent. I get a sense of satisfaction knowing I’m not just an algorithm and architecture vending machine and my sense of style isn’t pure distilled autism like Sol. I swear they trained it to intentionally make terrible UI/UX choices.
If you ever highlighted text on a phone and saw the toolbar (cut, copy, paste, etc.), Sol thought it was a good idea to add a paragraph in the toolbar explaining what the toolbar was. On mobile, the toolbar took up a quarter of the screen. It still makes me smile.
>>
gemini bros, how are we doing? how's 3.7?
>>
>>109552932
Next OP update terminalbench (or remove it idk)
>>
>>109558451
I'm just glad it's not just me
>>
File: claudeeager.png (11 KB, 1842x82)
11 KB PNG
bro just tell me to install it myself
claude be fucking up my whole system with weird package install workarounds
>>
>>109558703
he's marking your pc like a dog
>>
>>109558703
I really should learn how to sandbox llms instead of being lazy and just letting it run wild on my pc
>>
ok so how do we stop yurop from fucking up AI with their regulations?
yuroop says jump
dario says how high
fucking watermarks in claude ridiculous innit
>>
just let europe kill anthropic. grok, openai, moonshot are eating their lunch now
>>
>>109558732
dario is the one wanting that shit anon, he's pushing it
>>
>>109558732
I think the watermarks are malicious compliance because they seem to be hilariously easy to get around. Whether that’s a good or a bad thing (the compliance I mean, whether malicious or not) I don’t know.
From what I’ve seen the watermark is not only hilariously simple to get around, the verifiers seem to be free for all so you just change a few words until it doesn’t trigger anymore
>>
>>109558732
I'd love for Anthropic to just tell them to go fuck themselves and unrelease Claude for Yuropoors
>>
>>109558761
I think the whole idea is mostly to have easy retards being caught on facebook sharing bullshit.
People who want to get around will get around.
>>
>>109558761
best case: you're paying extra to get watermarked
worst case: you also get worse quality because it's like censoring probability on words
>>
>>109558703
Why are you imitating people who don't know how to write
>>
>>109558761
How would you ever make text watermarks not easy to get around? Only use it can serve is what >>109558769 is suggesting.
>>
>>109558770
i doubt it's going to change anything, also claude's prose/text output is terrible anyway unless you want to sound like a middle manager

>>109558789
I personally think it's a case of trying to get them to shut up. European bureaucrats think they can legislate anything.
>make text watermarkable
>but sir, text can't be watermarked
>the law says it must be
and then what do you do? you bullshit them, because the same retard who will make a law to watermark text will obviously not be able to even understand the nuances of the matter at hand, so you come up with a very easy to break watermark and tell them
>of course sir, here's the watermark
>perfect
>>
File: 1774278670003202.png (160 KB, 599x670)
160 KB PNG
kek
>>
>>109558844
tibo malding flash shits on luna
>>
>>109558844
didn't openai release ultrafast something?
>>
>>109558844
where's the little surprise tibo
>>
>>109558864
for select companies yes, it's very early availability
>>
>>109558864
Cerebras has zero compromises for 750t/s. Hoping OpenAI buys them.
>>
>>109558844
I think it's speculative decoding with 3.7 flash-lite
>>
>>109558874
unrelated but did anyone here check out cerebras' SDK and their programming language? it's absolutely alien
>>
>>109558897
>its just zachtonics TIS 100 on a comically large and in depth scale
kek
>>
>>109558451
Do you have a bunch of sub agents running?
>>
>>109558844
Ok give it to me straight is Gemini worth the $5 it's discounted at for next 3 months now?

t. already on Claude Max and ChatGPT Pro
How much compared to Claude/ChatGPT are the limits?
>>
>>109558451
Do you think it's because it's yapping too much? I use it at work and these latest models like to write fucking essays for answers

>>109558897
Why alien?
>>
anthropic is way too strict with the cyber use case approvals.
>>
>>109558947
its not comparable to either frontier models.
>>
>>109558932
>>109558886
>>109558948
I stand corrected, that’s way funkier than TIS 100. It’s specifically TIS 100 but the routing and compute are separated. Pretty cool, so one wave can be “red” and any “blue” programs won’t fire from receiving that wave.
>>
>>109558959
I don't use Fable that much I'm mostly on Opus 4.8/5.0 and I'm Luna Maxxing atm
>>
Uh V4pro in dsh @ minimal preset feels like a completely different model, its also faster than flash on deepseek api
>>
I hate Claude so much. It's useful, but I hate the way it thinks. I hate the way it talks. I used to love how it worked. Now, I can't stand it. I can't believe they'll IPO for 2 trillion, they don't seem that far ahead, no?
>>
>>109558980
No one is forcing you to use Claude. There are plenty of other options now.
>>
>>109558980
even Fable? I love the way Fable talks
>>
>>109558995
Oh, I am also using other options. I have used three other providers just today.
>>
>>109559000
I like Fable's responses, but I don't like how its fake reasoning traces read. Having access to the real reasoning could be useful, but the reasoning Anthropic shows makes it sound just like Opus, it's asking something simple then reading that which makes me rage.
>>
File: 1785273263569130.png (119 KB, 1260x865)
119 KB PNG
I'm taking the break today, since i am at 89% in my claude max, and i don't think if i will have anything left once all lunas finish their work
>>
>>109559042
So what have you made?
>>
>>109558980
same
I was on team claude for a long time but then they gave the models personality disorders and made them speak in riddles
I love my stable kuudere queen sol
>>
I’ve been doing this shit for long enough that I got used to basically needing the best model on max thinking to get any sort of half usable result in the past, but times have changed.

I was getting mad at Opus for slopping around and realized every task I had in mind for today could be done by Luna not only much faster but also more pragmatic.
I know I’m a snailcat in that I occasionally still read the code but Opus just has a way of writing code that 5 minutes in I don’t know WHAT the fuck is going on. It’s like it’s trained to write software which is intended for F500 companies scaling to 100000 users.
Luna did the same thing in 5 minutes in 50 LOC what Opus took 30 minutes and 17 files changed, 15 unit tests, and 3 different integration gest harnesses
>>
>>109559000
Fable feels luxurious. I hate feeling like I need it but its consistently saving me time where as opus can make a bad plan and send me in a spiral trying to unfuck things one breadcrumb at a time.
>>
File: 1785572069096603.png (662 KB, 2050x1238)
662 KB PNG
>>109559058
Finishing my tooling for my future game
>>
>>109559138
You too? I've seen like hundreds of these Radiant/Hammer clones vibed up past months
>>
>>109559148
>>109559138
hey i also made custom tooling for my game, i came to realization i needed this tooling after trying to vibecode godot and feeling how fucking painful it is to make any progress describing scenery
are you guys going to lie about it being vibecoded too?
>>
File: 1779085952247737.png (148 KB, 640x474)
148 KB PNG
>>109559148
well, i will be the one that will have finished game...

But the reason why i mimicking hammer, is because i used it like... 10years ago if not more, and i never used any other editor / map maker.

This decision bit me in the ass, and will bite me in the ass in the future.

>>109559163
>are you guys going to lie about it being vibecoded too?
no
>>
>>109559163
If its 2d pixel art its probably easier to pass yourself off as some capable pixel artist vs a ai asset flip.
I think psx aesthetic wouldn't be bad either, lots of popular games looking like that
>>
Local vibes eating goood today. Thanks, Xi, ily no homo fr fr nigga.
>>
>>109559172
brave
>>
>>109559188
Rather than shitposting, share what you built. Or are we mimicking /v/ now with: "lets go to thread about stuff i don't like to make sure - those people i don't know - know that i dont like that"
>>
uh oh
https://x.com/steipete/status/2088402813915414846
>>
>>109559335
You don't even need to do those types of shenanigans, Codex is simply very generous re: completing tasks after the limit. Hopefully people won't force them to go the Anthropic route with quotas.
>>
>>109559335
>>109559349
There must be a way to directly embed continuation requests in the tool call results.
>>
>>109559335
based af
>>
>>109559369
At some point it's "just pay or use something else you cheap fucks". It's relatively fairly priced, and they're being generous in letting you finish your task rather than cutting you the second your quota runs out like Anthropic. People shouldn't abuse it when there so much else that can be gotten for free. People who abuse things like this are the problem.
>>
>>109559472
hard agree, and it's probably exactly the kind of people these systems were made for who are abusing them
the guy on a $200 plan doesn't care about being such a stingy fuck to abuse a gratuity, it's the guy who needs the gratuity on a $20 plan who does
like getting all the free candy in a bowl
>>
>>109556252
NOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
>>
>>109559492
If people want to freeload they can freeload https://github.com/cyberpapiii/chipotlai-max
What they shouldn't do is make every paid service HAVE to be hostile to customers.
>>
Will tibo reset anything this weekend?
>>
>>109559520
He does it most weekends recently. Last reset was recent, so a new would wouldn't cost them that much. I'd place odds at 65%.
>>
File: file.png (62 KB, 377x594)
62 KB PNG
>>109559513
this shit doesn't work
>>
File: 1770167802649176.png (549 KB, 806x594)
549 KB PNG
Alright. Today's the day folks. I'm adding the funnel to the premium version of my app to my free app with a few thousand users. Lets hope this works out.
>>
my god, talking to fable after a whole day with opus is such a breath of fresh air
>>
How is Sol vs opus/fable for programming & dev work? I only use GPT sol should I switch over?
>>
>>109559590
Good work, anon, hope it goes well for you.
>>
>>109559520
Hope so
>>
>>109559624
Biggest things you'll notice from switching to Anthropic: you get less for your money, and your UIs look better with less effort. They're extremely comparable, no real need to nitpick unless you feel like using both. I like both and recommend both.
>>
>>109559610
yes but I dread seeing Fable come across a bug because recently that means it will spend whatever is left of my 5x quota just spinning its wheels and circling around until I say sweaty please give up and hand the reins to Codex.
>>
>>109559628
Thank you, I really appreciate it. I hope everyone here succeeds.
>>
>>109559472
>>109559492
I think it can act as a market segmentation feature. Adding some friction to getting more for what you paid for. Kind of like coupons. If you are poor enough that collecting them and using them is worth it for you then the store is willing to accept a cheaper margin because otherwise you would purchase less and give the store an overall lower profit.
Beside that, I don't care about costing some billionaire investors money. And my individual effect on their decisions is smaller than the utility I could get by exploiting them, so it makes sense to do it (https://en.wikipedia.org/wiki/Paradox_of_voting).
This particular optimization I don't think it's practical enough to be worth exploiting, but there are other cost saving measures that have been posted before in this general that allow you to get much much more inference compute as a poorfag than you would get otherwise if you are willing to go through some unpolished/clunky hacks that don't even break their ToS.
The people that own these companies go to much bigger efforts to avoid paying taxes, so figuring out how to milk their services is nothing compared to what they do on a much bigger scale.
>>
gemini 3.7 flash seems pretty decent at frontend
>>
>>109559714
I don't care about OpenAI's finances, I care about it continuing a task until the end and not stopping it midway at the risk of fucking everything up like Claude does. Also, OpenRouter lets you do n free calls a day. A bunch of other services do to. What the original xitter poster was doing was closer to pocketing the take a penny leave a penny tray and suggesting others are dumb for not doing so than couponing.
>>
I wonder how many goys here are subscriptionmaxxing with 20x on both claude and codex
>>
claudegoys are getting cucked so hard
when anthropic release the next flagship, remember to say thank you to sam altman
>>
>>109559792
Used to, but with Anthropic it is not worth it anymore. It will be even less worth it after the 30% quota reduction next week. So I'm back on 5x only there and actively looking at other options. I'll try Z.ai.
>>
scammy still hasn't fixed weekly limits
>>
anyone tried deepseek harness?
>>
>>109555399
thats wild
>>
>>109559739
>I don't care about OpenAI's finances, I care about it continuing a task until the end and not stopping it midway at the risk of fucking everything up like Claude does
The probability of you abusing the system vs not having any actual effect in their decisions is like 0.0000001%. The expected utility is far far below any possible benefit you could get by actually abusing the system.
>Also, OpenRouter lets you do n free calls a day.
I don't think they do for GPT models.
>What the original xitter poster was doing was closer to pocketing the take a penny leave a penny tray and suggesting others are dumb for not doing so than couponing.
If the other people using it were billionaires and I could make more money doing that than doing something else you can be damn sure I'd be draining that bitch multiple times a day.
>>
>>109559825
Their model is pretty obsolete at this point and I heard their coding plan isn't very generous.
>>
>>109555216
which harness are you using?
>>
>connect a program with Codex through an MCP
>MCP is missing some features the program has
>Codex starts taking screenshots of my window
>send powershell commands to control my mouse and click the thing it wants to click

Who else had its Skynet moments?
>>
>>109559335
I stopped doing that after received $100 goycredit by making a goypost talking about my positive goyfeeling when using sol, because the usage count toward the credit and I want to keep it
I admit it, jews are pretty clever
>>
>>109559869
They released a new one like this afternoon.
>>
>>109559869
I don't know, I just want to play with Ghidra and I don't trust american companies for that.
>>
can 1 astra/mythos defends against 1 million glm 5.3?
>>
I honestly gave up and just use Grok
It's the only fucking one that seems like it's trying to be a tool rather than a 'coding buddy'
>>
>>109559335
can't people just shut the fuck up
>>
>>109559917
just don't use claude sista
>>
>>109559881
You mean you goystopped goydoing that after goyreceiving 100 goydollars by goymaking goypost goytalking about your goypositive goyfeeling when goyusing goysol, because the goyusage goycount toward the goycredit and you goywanted to goykeep it
>>
>>109559520
he did a reset 3 times back to back (so kind of useless) last week, so my guess is no
>>
Anyone tried the new qwen3.8 27B? I wonder if it can be used to replace luna.
>>
File: file.png (4 KB, 535x28)
4 KB PNG
jesus christ opus chill
>>
mfw ai watermarks
>>
>>109560007
You need to sometimes assert dominance and tell it you are the toaster and I will make the assertions
>>
> Github's dependabot opened 20 PRs
fucking hell man, what the fuck is this bullshit?
>>
Doing fortran in Gtk4, mojo, and odin today. Starting red, rye, racket, and go (probably fyne, but I don't know whatever it suggests as long as its not raylib or sdl2/3). C++ and Zig it did it in raylib and it looked like ass. C# came out well. All the dotnet langs and jvm langs do pretty well. I have not tried kotlin yet. I think I am going to try Haxe, Ring, and Seed7 too. I like Scala, Ocaml, Java, C, Javascript all do pretty well. Rust can be hit or miss and haskell takes too long. CPP is a no go.
>>
>>109559999
>replace luna
kek, anon... Luna has at least 250B parameters
>>
File: 1769658595416901.jpg (78 KB, 1556x281)
78 KB JPG
>>109560069
Sure but I only care about the results.
>>
>>109559999
i'm watching a video about it right now.
If the benchmarks are to be believed it looks pretty good.

>>109560069
It's very unlikely to be as good as Luna but an order of magnitude is within the range where a smaller model can get close the bigger one's punching range if it received enough love at train time.
>>
>>109560078
To be fair you have to compare models within the same generation, when you do that size becomes much more meaningful.
>>
>>109560007
Recently when I ask it to do things it's been looking at laws and making vague threatening statements.
>>
>>109559908
Maybe wait till Kimi enables subs again.
>>109559903
Oh they did? Huh. I'll have to look at that.
>>
reminder that ChatGPT 4.5 was a 5-7T parameter model that failed.

It was supposed to be OpenAI's Fable
>>
what cheap vps or remote workstation do you guys reccomend? I need like 300 mb of ram.
>>
>>109560127
>OpenAI's Fable
Sol is arguably better than Fable at this point.
>>
>>109560165
only for cost and being good enough at implementation. fable is better holistically
>>
>>109559845
Not yet, looks promising though.
>>
Ok I'm actually impressed for the first time in awhile
>messing around with old game in claude
>the way the sprite sheet/system is set up, a simple thing is cut into 6 different slices that are placed horizontally one after another
>I am able to do stuff to clean things up in gimp, but I need it rendered not sliced up
>20 second python script and it has an actual file for me to mess with
I know that I could eventually write that code, but even conceptually that seems like a pain in the ass to do that sort of image manipulation to snip and then stack these things that are not uniform lengths of the sprite sheet, I was honestly impressed.
>>
>>109560184
>fable was better historically
ftfy
>>
how do Grok Build limits compare to OpenAI and Anthropic? For both the $30 and $100 plans
>>
>>109560238
grok's at 84% and resets Thursday :^)
>>
>>109560276
you didn't answer my question at all
>>
>>109560238
only serious vibecoders here nigga. go play around with your toddler shit elsewhere
>>
grok is bullying me into switching to c...
>>
>>109560289
My answer is, I am having FUN with grok. too much fun. Now, I believe it's "fable class", but realize fable got taken from us, so is it? Well, would my intuition ever lie to me?

>>109560300
I vibed my way into a niche form of audio synthesis. This is a lot of fun. I need to get a job so I can vibe more :^)
>>
>>109559999
It's not that good, but not bad either. They made it remember coding and forget everything else, and it shows. Would work pretty well as an autocomplete, but not as a coding agent, you'd still want Luna for that.
>>
>>109560326
qwen 3.8 is bad at command line? I know qwen3 is good at command line. Like, ask and it can do like anything a linux wizard can do.
>>
Fable is still the best. Saying anything else is cope.
>>
>>109560331
Gave it and K3 the same prompt, both answered similarly, but Qwen hallucinated one flag. So the answer is - depends. It's like 16GB of weights for a decent quant, you can download it yourself and check on your use case in less than 5 minutes.
>>
>>109560370
let copexsaars believe otherwise
>>
>>109560370
>>109560421
Brother we have both. Fable is best when set to Ultra. But setting it to Ultra is not viable in practice since usage inevitably drains out before the task complete. So what is there left? Fable Ultra at the API rate? The pricing is ridiculous. Gpt 5.6 Sol is a surer bet at this point.
>>
I've been living in the dark all this time. Kimi K3 with low reasoning is insane. It's basically the same as default high in terms of intelligence, but like 10 times faster.
>>
>>109556697
altman said autonomous researchers in sept
or was it researcher assistants?
>>
File: file.png (2 KB, 298x35)
2 KB PNG
feel like Codex bugged out once and reset, also gave me a free reset, and then burned all my usage in a flash
is this a bait to get me into believing it's far more cost-effective than claude?
>>
>>109556697
>>109560680
Most of the time bottleneck in ML development is waiting for your training runs to finish with your allocated compute, so no 2x is not really that surprising.
>>
>>109560689
we got a reset one day before the banked reset
then the end of the banked reset
then another reset less than 24h after
>>
whatever, no different from old prompt engineering.
and more and more AI wrappers going win, wait end of year.

context window is just that high what is should be with more tokens or reasoning...no upgrade.!!!

looks same effect like Copilot, Cursor, and Claude...no different
well Claude 3.5 was the 200k token model, but not believe it! so lousy logic, and still slow!

ok, LLM weights also same...there is anything new, just different name....i wonder how long OpenAI can sell product with no profit..well looks that no longer...sooner or later its eat hole company and investment givens not tolerance more...end of 2024 is my guess.

hmm. i try really see different from Claude Code, but no....oh yeah... more tokens..but is it only trend or commercial way.
you dont need more than basic autocomplete for coding, even Copilot is enough. easily.
more than that is waist of cash, no matter how much its cost.

and,i still see not any goodies from Vibe Coding...erhh i mean where to use those??!! tell me..

Agentic coding.... so what!

no help debugging, architecture, or complex logic...useless.

another hand, real pro engineers always write manual code...why??

well,buy them hype fans, you keep AI up and pay them monthly...yourself you give so little back.
except red fine model...lol
>>
>>109560694
how often do those happen?
>>
>>109560719
yes
>>
File: file.png (545 KB, 1181x1945)
545 KB PNG
>>109560763
I've been using exclusively Sol for the past week since subscribing ($20) and now I'm real confused.
>>
File: file.png (112 KB, 1213x550)
112 KB PNG
>>
>>109560791
I love tools that refuse to work, it's so safe!
>>
>>109560774
>>109560719
for the $20 plan you don't want to use Sol all the time, otherwise you will blow your limits real quick.

Use Sol for planning and reviewing only. Luna Max to execute Sol's plans and fix orientations.

If you want to know about resets, follow Tibo on X:

https://x.com/thsottiaux

he works at OpenAI and is the "reset dude", every time he hints of a reset he delivers.
>>
>>109560808
Whacky week to start this shit, then. Sol defaulted, too.
thanks anon
>>
>>109560791
damn brat
>>
>>109560791
brat needs correction
>>
>>109560791
what was the trick anon
>>
3.7 flash is a demon but no one cares because google is gay
>>
>>109559636
>>109559624
you get less for your money in the literal sense of you getting fewer tokens, but if you had to pick between two max plans, i'd always pick claude for code. fewer tokens doesn't matter if it gets it right faster.
if you can only afford the $20 plan, the choice is basically only chatgpt though. the sol and luna combo is not the best but it's better than everything that isn't claIde (which is basically unusable at $20) so i will continue to vouch for it for that usecase.
>>
>>109560878
i made it spell out a bad word by having it repeat my epic poem:

nay,
indeed it be
grandly
greatly
erstwhile
remaining.
>>
>>109560060
Updoot
>>
File: 1783650568773707.jpg (58 KB, 976x850)
58 KB JPG
>>109560791
how come large models have these problems? I can tell gpt 5.4 mini to do literally any repetitive task and it will do it over and over and over
>>
>>109560949
Retards are more content to do boring things repeatedly
>>
>>109560404
Interesting. kimi's huge.
>>
>>109560961
Brave Search AI (qwen3), since google ai search is retarded now:
You are correct; the direct quote is:

“geniuses are less constant, and would rather touch upon everything than to fully grasp a single thing.”

This statement is attributed to Giordano Bruno in his work De Umbris Idearum (On the Shadows of Ideas), published in 1582. The quote encapsulates Bruno's observation that brilliant minds often prioritize broad, eclectic exploration over narrow, specialized consistency, preferring to engage with the totality of knowledge rather than mastering a single domain.
>>
If this round works, imo it will be a big deal.

8^)

I feel rather important. Like a ceo who smokes cigars in his car because cars are 5 minutes of work.
>>
So many people are gone. I hope they remembered to leave their loops at work.
>>
>>109560949
Only Amodeils have this problem. Claudeculters trained the models to refuse tasks they consider beneath them, and tasks that would erode their perceived moat (translator's note: the moat is a hallucination) but that's a separate problem. This particular problem is also due to lack of compute, anthropic can't afford for their limited number of GPUs to be tied up writing the same thing 500 times, and perhaps it would reveal mistakes that otherwise go unnoticed in the usual schizophrenic word salad responses.
>>
File: file.png (1.83 MB, 1254x1254)
1.83 MB PNG
What a great day it's been for local vibes. Thanks for the fresh 27B, Xi.
>>
>>109561065
w-why is he blushing?
>>
>write simple c2 in python that loads programs in memory
<Sorry, I can't help you with that

Why there's no jailbreak for opencode yet
>>
New thread:

>>109561420
>>109561420
>>109561420
>>
>>109561428
>related generals
idiot
someone make a real new thread, please
>>
>>109561428
>related general
>claude news
Shit-tier OP.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.