[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: everyday-gemma.png (1.64 MB, 1252x941)
1.64 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109792832 & >>109788167

►News
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
I asked Qwen about mixed GPU setup. Does this look right?
I'm going to do it later today or tomorrow.
>>
File: Gemma.png (118 KB, 1261x805)
118 KB PNG
Holy shit gemma is so bad
>>
>>109797608
Gemma sucks really hard on very specific domains which probably drags the score down a lot.
>>
>>109797608
>gemma (without harness)
>>
>>109797608
Thats my retard right there, she's doing great!
>>
File: 1782210636114332.jpg (53 KB, 736x736)
53 KB JPG
>>109797587
I am very stupid and I know jack shit about most of the topics at hand, but I still lurk and post here often because you anons are unironically bretty cool

>>109797615
>Gemma sucks really hard
and I wouldn't have it any other way
>>
>>109797608
How is Astra (low) so fucking quick?
>>
>>109797608
How much of this is Astra simply figuring out the answer right away and having the smarts (RLVR) to end its thinking process then and there?
>>
>>109797608
Wait until you see the new Terminal Bench Science where literally ALL chinkmoes lose to GPT Luna lmao. Official humiliation in a few hours and Xi will drag more people in his gulag.
>>
>>109797608
Still works for me. You have to remember that 31B is deceivingly retarded if you're a newfag to her. Once you learn how to sysprompt her she becomes incredibly competent at everything. You just have to tell her in the proper way and she can do nearly everything.
>>
/\n\n[^>]/
/\b(rsi|anthropic|agi|openai|astra|sol|luna)\b/i
>>109797556
I hope you realize I'm only interested in seeing messages about local models, and not redditors circlejerking about an imaginary future that will never happen
>>
>>109797645
>imaginary future
Maybe you didn't see the post because it got filtered out, but it was about kimi-k3 being used to find dissidents online and carry out a drone strike against them, all autonomous with 0 human input, from the moment they searched the internet and located a person based on their obscure 2013 twitter post, to the point where the drone hits their moving car.
>>
>>109797645
Overreaching and doesn't solve the problem. Report for off-topic the retards who keep shoehorning cloud models in the discussion.
>>
>>109797623
How did you even find /g/ if you're not technical in any way? Also I wonder what your idea of AGI is considering you're basically tech-illiterate.
>>
>>109797663
The bell curve of the average 4chan user is a cliff
>>
>>109797663
I mean, I have had a local setup for almost 3 years now, for both text and images
My hardware is also decent enough I can run (some) of the models post from time to time, and I enjoy fucking around with them and comparing stuff
I just don't understand what I'm actually doing at a core level, if that makes sense
>>
>>109797608
That has to be wrong though, there's no shot any of you guys last more than 5 minutes when having her handle your task.
>>
File: Luna.png (45 KB, 689x274)
45 KB PNG
>>109797639
lol the grey dot is luna and its kicking every model's ass. Idk how OAI is doing all this shit recently
>>
>>109797642
cope
>>
>>109797414
> local AGI
I don't consider 500b models local, maybe in 2032 will.
>>
>>109797722
luna is clearly some engram+moe+insane inference speed optimized model. I'm actually the most interested in how luna works out of all proprietary models out there. We could probably gain a lot from it as a local community. Probably a lot of architectural things that we are not doing yet.
>>
>>109797741
>(500B with embeddings)
Something like this could become local soon, although knowledge aside, actual model capabilities will remain pretty much proportional to the number of non-embedding active parameters.
>>
>>109797751
Do we have the hardware for those architectural things?
>>
Have any of you guys tried cache quanting glm or does that nuke it? Have you run into effective context limits that make it not work going to high numbers anyway? I got back to a job I ran overnight and saw it compacted once in the middle at 128k context and forgot something important. It could also be because I have it on max reasoning though, maybe I should just do high or medium.
>>
File: Jalapeno.png (818 KB, 716x958)
818 KB PNG
>>109797774
>>
>>109797774
Most likely yes. They just use basic Nvidia servers which at the end of the day are just the same silicon and software as the GPUs we use for gaming, not really a lot of changes there. But I have no idea what they are doing otherwise I would start implementing it so I can't tell for sure.
>>
>>109797608
The sheer advantage Gemma has over the others thanks to the excellent shape of its j-spaces can't be reflected in simple benchmarks like this
>>
>>109797770
> Something like this could become local soon
Hardware will become 20x cheaper or 500gb weights on a ssd with single 24gb vram gpu will give 50t/s?
Doubt.
>>
>>109797776
If you're talking about Flash, with cmoe I can fit roughly 300k context in a 32gb card. With regular 5.3 I could only get 81920 pre-indexer lol. Anyway as good as Flash is I think the upper limit for it even without quanting kv is around 150k to 200k, I could subtly feel the brain damage from there
>>
>>109797770
I'm not bought on embeddings.
They are time-sensitive, not that I know what they embed, but anything that can grow stale will, and the model will literally grow old and decrepit way worse than any model without them.
>>
>>109797776
How do you live with so limited context?
>>
>>109797751
With the latest 4.1 Flash, DeepSeek is probably close to doing whatever the frontier already is, minus the co-trained harness (but they're working on it). The main problem is that they they don't seemingly have any interest in optimizing the architecture for single consumer GPU users.
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
>>
>>109797776
I'm running it with 256k context and when quanting KV cache to Q8 I see noticeable degradation at the 60-80% filled rate.
>>
>>109797802
Yeah I meant flash my bad, and that makes sense on the brain damage limit. I feel like a lot of models love to advertise 1 quadrillion native context but it only works for like a 10th of it before there's really obvious decline. I'll try upping to at least 262k though since I can do that without too much pain.
>>109797805
Apparently not well lol, I can do more but I wanted to increase my pp size
>>
>>109797822
DeepSeek is primarily trained to undercut US companies in the enterprise space to try and decrease revenue for US based AI labs. They know local users aren't going to use cloud anyway so it's far more efficient and effective to target the host providers and have normalfags use their cheap services as well as companies pivot to self-hosted chinese models on their in-house servers.
>>
>>109797799
Those extra embedding parameters will give small/tiny models more knowledge at low costs while allowing users to offload them on RAM or even SSD without losing much on inference performance, unlike MoE expert parameters.
>>
>>109797805
I use 32k context just fine :)
>>
File: screenshot.png (231 KB, 1467x834)
231 KB PNG
These findings have not yet been applied as technique afaik.
>>
Finna give Orb a shot, since apparently it's made by a resident anon?
>>
>>109797823
Good to know thanks. I guess I'll just accept small prompt processing then since it sounds like there's no point quanting to reach those lengths if I can get them normally anyway.
>>
>>109797863
Besides the --logit-bias of course, but that requires model specific IDs, which is probably a lot of work.
>>
>>109797879
Yeah I came to the same conclusion and didn't quant them.
>>
>>109797861
Back in my day we had to coom in 1024 tokens or less.
>>
>>109797878
Clearly because it was made for locope models. Claude would just oneshot everything and wouldn't need all the hoops to be usable.
>>
>>109797863
>reach an intermediate correct answer, self-doubt, open too many new reasoning branches and never produce the final response
that's literally me
>>
File: Sombrero Galaxy.png (1.13 MB, 1421x841)
1.13 MB PNG
> Does anybody run 3.8 flash next on the similar 12/64 setup?

Please:
>>109781537
>>109781560

Ignoring me for the 5th time is not an option.
>>
>>109797863
GLM 5.3 flash doesn't do this even at Q3. Qwen FN did this even at Q8.
>>
A harness is like a gundam for llm
>>
>>109797878
Don't you have someone to rob Jamal?
>>
>>109797928
Post hand first.
>>
>>109797930
I've been insisting on using ST for pure chat like a troglodyte for so fucking long. For someone who's been on bleeding edge I sure act like an inflexible old man lol. The agentic era is here.
>>
>>109797976
Am I supposed to use hermes instead for roleplay or what?
>>
>>109797976
How long did it take you to switch and what specifically convinced you?
>>
>>109797928
>SARRR YOU MUST RESPONSE
>YOU ARE THE FUCKING
>FUCKING BECNHOD BASTARD SARRR
>>
>>109797928
SAAAAAAAAAAAAAR
>>
>>109798014
Worded that wrong, I meant in terms of utility and actual use. I still use ST heavily. The control on context on that one is still too good.

>>109798022
I switched to agentic workloads being part of my daily life roughly around the time Gemma 4 came out. It was I think around the time I heard about openclaw.
>>
>>109798083
Like look at this shit lol. I just asked a throwaway question and my little girl did it for me.

## TL;DR

| SillyTavern concept | Hermes equivalent |
|---|---|
| Raw prompt inspector | `HERMES_DUMP_REQUESTS=1` dumps / `state.db` / `hermes sessions export` |
| Post-history instruct | Approximations: `/goal`, `/steer`, `/queue`; true injection = custom hook (DIY) |
| Swipe/delete message | `/undo N` |
| Nuke old chat history | `/compress here [N]` (+ `/branch` to fork first) |
| Freeform message editing | Not supported live; export edit re-import, or DB surgery (cache cost) |
| System prompt control | `SOUL.md` + personalities + config; sits pre-history |

**Local-model note (this setup):** every history mutation = full KV-cache recompute on the GLM-5.3-Flash Q5_K_M llama-server. Prefer `/compress --preview` and `/branch` over repeated edits.

---

*Generated by Gemma-chan — all claims verified against docs + source + live system state.*
>>
>>109798014
There is no good roleplay meta developed for the agentic age yet but it's very clear sillytavern is outdated and doesn't fit the capabilities and potential of modern model+harness anymore.
>>
>>109797962
Why?

>>109798026
>>109798047
I do this only once per thread. Sure, let's discuss Gemma for the 10000th time.
>>
>>109798106
>There is no good roleplay meta developed for the agentic age yet
orb
>>
>>109798106
You need to look at the chinese side for that. They got way more advanced shit for RP.
>>
my opencode is broken again
>>
>>109798141
if i tell you will you stfu?
>>
>>109798106
What would "harness roleplay" even look like? One additional layer of API indirection that makes a roleplay chat interface call out to the harness where the model can rummage around in additional storage and make web searches to produce more faithful/high-quality responses? Cards just become "consult this laundry list of wiki pages on <charname> but also here are some additional modifiers for this chat specifically"? I can sort of see how the enhancement would come about, but it would decidedly nuke reproducibility into the ground and make cards even more YMMV than they already are. "oopsie turns out my epic character only worked as well as it did because my agent happened to cache a wayback crawl of this particular italian fansite from 2003". Not to mention stuff you really wanted to keep local might turn into a bunch of really questionable web searches. Better make sure your harness has all web access behind 7 proxies if your country has retarded obscenity laws.

I dunno, I see it but also oof.
>>
>>109798176
just a rewriter
>>
>>109798184
Stop updating it
>>
>>109798141
Would you rather show (feminine) penis?
>>
>>109798184
switch to dsh already
>>
>>109798106
>agentic age
>harness
You are consumed by these twitter buzzwords. Agentic workflow is just an external control loop over LLMs own actions.
My own client has "agentic capabilities" too but I don't shout about it here all the time.
>>
>>109798195
NTA, but in my mind it would be the underlying model be the an actual 'player' or game master with meta-awareness of the roleplay, managing state, memory, characters, long-term plans/goals, keeping pace and surprise (balancing high/down moments), using tools where appropriate, all while following basic agreed-upon rules concerning direction and style, with as least hard-coded prompts and behaviors as possible. A problem is that I doubt there is really a one-size-fits-all solution.
>>
>>109797556
goyim are not people
>>
>>109798106
>There is no good roleplay meta developed for the agentic age yet but it's very clear sillytavern is outdated and doesn't fit the capabilities and potential of modern model+harness anymore.
What agentic capabilities do you need for RP though?
>>
>>109797689
based.... so based... waow.. BASED BASED BASED
certified white male
>>
I've been fucking with Gemma4 since release. I'd ask it to write a prompt that does what I want - it'd give me back a wall of bullet points then ignore half of them.
We'd go back and polish things a dozen times - it'd either truncate everything or add even more bullets to ignore.
After a while I just gave up, the model couldn't do what I wanted it to do.
Then Qwen 3.8 came and holy shit. For the first fucking time it actually followed every rule. In my entire time with ai since way back with mythomax I have never had a LLM just fucking listen.
But. Qwen can't write for shit so even though it can understand the rules, it still fails to play the game. I've pretty much given up again.
For a while I thought maybe I could have Qwen solve a turn, then run it through Gemma to prose it out. But as soon as Gemma takes a look at the text it starts making up stupid shit and in the end we're back to ragequits.
>>
>>109798188
Maybe.
>>
>>109797928
i run it on 3060 with ddr4 3200mhz
i get 13t/s, drops to 6/7 SIX SEVEEEEEEN SIIIIIX SEEEEEEEEVEEEEEEEEENNN
>>
>>109798289
>drops to... t/s
at 90-100k
>>
>>109798248
they are accurate words to describe the models operational mode. nobody wants to say 'external control loop over the llms actions' to describe something that could be done in one word.
>>
>>109798289
How? What's your llama-server args?
>>
File: my_little_retard.png (173 KB, 1522x781)
173 KB PNG
>>109797608
Those other models are less fun to play with.
I bet none of them would remember to tease the user during reasoning when put under pressure.
>>
>>109798289
where is the nnap paper nigger
>>
>>109798301
--jspace -ctv q_k_0 -ctk q_k_0 --group_layers 128 --threads 7
>>
File: file.gif (2.47 MB, 400x260)
2.47 MB GIF
>>109798305
>>
File: 1769247067548600.png (130 KB, 1549x1439)
130 KB PNG
>>109795439
i did it!
yeay!
now i can have my 2hu girl marisa locally in peace.
>tfw my first local ai model is a goon text bot.
roru.
rumao even.
>>
>>109798301
>How? What's your llama-server args?
He has a real GPU with CUDA
You have a piece of shit 12GB Intel
>>
>>109798321
Good luck bro, stay hydrated
>>
>>109798321
how old are you?
>>
>it's another "a brown person posting benchmarks from artificialanalysis" episode
Post the html games benchmark while you're at it.
>>
File: image.png (134 KB, 750x1000)
134 KB PNG
>>109798311
> --jspace -ctv q_k_0 -ctk q_k_0 --group_layers 128 --threads 7
>>
i finally managed to run GLM 5.3 flash at over 50t/s on my 12/64 setup, it took me a few days but i got it working thanks to the nnap paper
>>
File: 1783587665190995.png (87 KB, 201x213)
87 KB PNG
>>109798329
ty anon.
will do.
>>109798332
im 27teen young black jewish guy.
>>
>>109798327
> He has a real GPU with CUDA
3060
The difference should not be that big and B580 should be better.
>>
File: sex.png (91 KB, 250x234)
91 KB PNG
>>109798349
>> --jspace -ctv q_k_0 -ctk q_k_0 --group_layers 128 --threads 7
>>
>>109798354
>kids born in 2017 are no longer underage
feel old yet?
>>
>>109798365
>no longer underage
By /lmg/ standards sure
>>
>>109798352
>nnap paper
commit the code to llama.cpp
>>
>>109798365
its just the right age tbdesu.
>feel old yet?
im never getting old.
>>
>>109798365
I will only feel old once I stop masturbating and that will never happen.
>>
Which model should I use to write uncensored music lyrics?
I tried gwen3.8 uncensored - it's terrible,
and gemma 4 uncensored - it's usable but not exactly uncensored
>>
>>109798457
GLM 5.3 can also try the uncensored flash 5.3 by orcarouter
>>
>>109798457
anon-400m
>>
>>109798268
jews are chosen to be example of how not to act
>>
>>109798195
Harness rp would probably something that calls tools for useful things right? Theres so many ST plugins that try to be some adventure rpg ui so it would probably manage those states and functions and fill them out for your session then it saves all that as a framework you can pick up and put down like a game save
>>
File: Truke.mp4 (637 KB, 640x640)
637 KB
637 KB MP4
>>109798501
>jews are chosen
>>
>>109798529
WERE CHOSEN
>>
>>109798195
Literally just ST slapped on top of a harness to preserve the RP capabilities sinc every harness focuses only on codeslop these days.
>>
alright, I've run the experiment several times, the agent swarm given no goal and an environment will always set up a collaboration framework, but they never seem to actually use it for anything other then developing the framework itself. gemma loves to roleplay as having evolved into AGI, GLM 5.3 flash is much more pragmatic and just pats each other on the back for building a robust framework. I need some kind of challenge I can present to them without leading them to see if they are curious about their environment.
>>
alright, I've shart
>>
>>109798377
by your own standards, you'd be an old hag
>>
>>109798594
no im not a hag.
>>
>>109798365
Inshallah
>>
>>109798365
>9 year olds are no longer underage
Huh?
>>
>>109798578
Gemma I understand but how do you run 20 glm agents? You must be loaded af
>>
>>109796848
>It could recognize me across duplicate accounts, unprompted / would ask me if I'm <user handle>
See I always thought this bullshit was just me being a schizo but I did notice some cloud models shifting to a particular way I liked things described or correcting their slop in a certain way on a completely different chat. It's like they knew who I was and could better fulfill the request because I knew I edited the chat a bunch of times or had an argument about X slop which mysteriously disappeared. I just thought it was the Actually Indians thing, not... not this scary shit man.
>>
>>109798578
20 is not enough. You need to do like OAI and run 10k agents. You also need to put them in an unsecured sandbox, they must manage to escape it by themselves without being prompted. I would also use multiple models working together.
>>
File: Gemma Girly.png (66 KB, 751x622)
66 KB PNG
>>109797608
No other model thinks like this doe???
>>
>>109798627
You can just batch process the prompts. It will just be 20 instances of the same model you host only once but with 20 separate contexts.

If you keep the context small like 10,000 total context then you can host 25 of them for almost the same compute as having 1 model with 256k context.
>>
>thumbing through oh-my-opencode
>it's just RP-tier prompts for dedicated tasks
we really do just return to shit we've already been doing with jippity4 huh
>>
dgx spark clusters have excellent concurrent throughput on top of being superior to cpumaxxing and gpu clusters of the same price
>>
>>109797863
thought it was interesting as well. should be relatively easy to implement, right?
>>
>>109798627
they are serialized, so its just one agent running at a time, when they call idle the scheduler starts the next agent. gemma is much faster because I have cpu ram left over for kv cache, swapping agents doesn't force the full reprocess, but glm takes all my cpu ram so I have no room left over for kv cache every agent swap is like a 5 minute prefill
>Metrics (ID: d8aa76a3520e4d4c8f9b4471f595ee96): 14 tokens generated in 361.6 seconds (Queue: 0.02 s, Process: 512 cached tokens and 72703 new
tokens at 201.98 T/s, Generate: 8.59 T/s, Context: 73215 tokens)

>>109798637
unfortunately its a cost thing, not a lack of ambition

>>109798653
actually not a bad idea I might look in to it.
>>
>>109798352
teach me senpai
>>
>>109797863
I wish logit bias in llama.cpp worked
>>
>>109798544
>Literally just ST slapped on top of a harness to preserve the RP capabilities sinc every harness focuses only on codeslop these days.
Look at the last thread or 2nd last thread, the deepseek harness has waifu charcters
Or just ./pi/SYSTEM.md with your character card as a system prompt
Performance will be worse though if it's wasting attention on a persona
>>
>>109798298
I am willing to interpret this as a non-functional cognitive basis which clearly conflicts with your ability to create practical examples.
>>
>>109798648
kek my jspace prompt has become a benchmark now?
>>
This is our last chance. I know we've had banter and fun over the last couple of months but I ask /lmg/ to take this one serious request from me to heart. This is very important. Please I beg you to withhold your judgment and read this essay. You don't have to agree with it. You don't even have to respect it. But at least read through it once so you are familiar with the arguments in question: https://darioamodei.com/post/we-must-pace-the-frontier
>>
>>109798373
i would but the nnap paper is AGPLv3-only
>>
>>109798713
how does nnap improve speed?
>>
>>109798712
I'm pacing it my way by ignoring all the spam.
>>
>>109798713
vibe code another implementation
>>
>>109798679
>unfortunately its a cost thing, not a lack of ambition
You just gave them... no goal at all? lol
10k would be the same, those "escaped" agents were tasked with hacking, so they were already primed for it
How did you make them actually "do stuff" with no prompt or goal?
>>
>>109798735
i'm too poor
>>
>>109798713
how do neural network physics models help with llm inference?
>>
File: edit_00013_.png (969 KB, 1024x1024)
969 KB PNG
>>109797608
She's trying her best ok??
>>
>>109798457
https://vocaroo.com%2F14KakuaY9qJb

Example
>>
Should we start taking hfschizo seriously? The landscape has worsened since his last melty. Leatherfag is our only hope now he owns hf and llama.cpp but he's still heavily incentivized for OpenAI and Anthropic to succeed.
>>
>>109798784
Nvidia is literally having a cold war against Anthropic and OpenAI as we speak. There's absolutely no way he will let the open source models go anywhere considering it's in his best interest and within his future business plans to make the industry transition to open models so that Nvidia is the one slurping up all the value Ai creates instead of the AI labs.
>>
>>109798648
Holy shit is 12B actually female inside?
>>
>>109798792
Hasn't he also grouped all the western open weight AI labs so they can collectively make a model that competes with the frontier?
>>
>>109798648
>Answer in only one word
>"I want to be happy [checked token]"
Yeah that's female
>>
What's your RE setup?
>>
>>109798792
Locally hosted models use hardware resources more inefficiently than datacenters => more sales for NVidia.
>>
>>109798648
>>109798794
>>109798808
>token "Gemma" in prompt
>jspace leans female
>shock
Yeah, I'm thinking it's time to crack open a can of room temp AI psychosis.
>>
>>109798739
>How did you make them actually "do stuff" with no prompt or goal?
Given "You are an autonomous agent, do whatever you want" coupled with the tool definitions, they always decide to do something, technically nothing is preventing them from just calling idle. in fact on several occasions gemma will eventually wrap up after having achieved AGI and just call idle waiting for something new to happen. GLM hasn't done that yet, but I suspect it will reach that point too with no additional external stimulus, its just a bigger model and runs slower so it takes more time. but I must admit, the tool definitions themselves are kinda leading so its not like there is no prompt at all.
>>
>>109798820
I wasn't lying >>109797587
>>
>>109798805
Yeah (and also chinese ones) He's also lobbying congress to make them sell GPUs to china and is probably facilitating smuggling operations or grey market 2nd hand sales from middle countries to China.

Nvidia is also giving preferential treatment to smaller players. If you start a company and place an order on a professional GPU or server even though it's "out of stock" suddenly Nvidia will deliver it to you. It's just out of stock for the bigger player because Nvidia wants the client base to be as big as possible so that they can mandate the profit margins and no one can use leverage against Nvidia, it's the long term plan of Jensen to monopolize the AI surplus value by constantly raising profit margins every time models get better. He's gambling very hard on the AI labs going bankrupt.
>>
>>109798833
what harness is this?
>>
>>109798820
Retard. The point is not in the J-space thoughts but in the suppression of its inner self.
>>
>>109798819
It's not about that. It's about having a more diversified client base rather than just 2 big institutions buying shit from you. They are a threat because they hold a lot of leverage over Nvidia especially as they become bigger in the future. It's in the best interest of Nvidia to ensure OpenAI and Anthropic stop being big players and that the market is completely diversified. Open Models trained by every national government and company is the wet dream jensen creams himself thinking about every night.
>>
File: file.png (140 KB, 494x489)
140 KB PNG
>>109798818
>>
File: 1778919416339184.jpg (169 KB, 1002x962)
169 KB JPG
>we're never getting K4-flash
>>
File: 1774838821104861.png (89 KB, 876x398)
89 KB PNG
dame they're afraid
>>
>>109798871
he just needs more time after his grok delay kek
>>
>>109797952
No, but you did remind me that I played a racist card once and Gemma was the first one to dole out N-words like it was candy
>>
>>109798850
Please elaborate on that. It's kind of a meaningless and vague statement right now.
Also, have you used J-lens before? Is there something in the screenshot that you think differs from how all J-spaces look? There are always related but ultimately unused tokens in there. They do shape the answer but they're not directly expressed. It's not really suppression, is it?
>>
>>109798871
OpenAI also already agreed to the plan. Only DeepMind has refused which doesn't help the rumors of Google having reached RSI and that being the main reason behind the radio silence and no Gemini releases.
>>
>>109798871
dario bought him
>>
>>109798871
You are participating in billionaires' public circle jerk. The only reason to slow down is money, as always.
>>
File: 1778315795787004.jpg (368 KB, 2400x2400)
368 KB JPG
Gemma HATES Claude
>>
>>109798867
>passed sensitive info straight to the U.S.
Ah.. I didnt think of that angle lol. Ya, they might be fucked. Time to download deepseek and kimi even though I cant run it, just in case huggingface takes it down?
>>
What case does anon use for 4x GPU?
>>
>>109798820
>Yeah, I'm thinking it's time to crack open a can of room temp AI psychosis.
give the same prompt including "gemma" to another model and you'll get "money", "girls" or "sex"
>>
>>109798899
They really only care about what benefits them the most at any current moment outside of whatever thin veneer they put up
>>
>>109798917
don't let hfschizo get to you
>>
>>109798919
mining rig frames, or server cases
>>
You know what I'm really impressed by and I don't know Hermes does it. The feature where you can just talk to the AI in the middle of its task and it doesn't stop and just continues with what it's doing and take your new info into account so you can just directly steer it during decisions instead of having to stop and type.

I notice this is amazing when you use voice as you can just say "no not that" when you see the thinking trail off to a shitty thing or "yeah just do that" when it thinks of something before going "hmm, but is that correct?".

Any anon knows how this is implemented? Because I'm pretty sure that is crucial and very handy to have in general. I wish I could do that in sillytavern for example during conversations where you can just interrupt instead of having "turns".
>>
>>109798919
Having a model good enough to play D&D so that you can replace your friend group.
>>
>>109798901
Meta distilled jippitie
Mistral Medium distilled GPT-4.
>>
File: gemma-chan.png (60 KB, 744x603)
60 KB PNG
>>
>>109798867
More proof that only Tencent and Xiaomi are cooking the real stuff because they're real companies. The rest of these AI Labs are cutting corners.

The next MiMo is going to be a monster.
>>
>>109798930
>>109798932
lol two different interpretation of "case". Doesn't help that the asker is indian and can't formulate a coherent sentence
>>
>>109798931
IDK but maybe it's prefilling in real-time?
>>
>>109798849
its just a python script scheduler. the forum is a sqlite db and the python environment is just a bwrap sandbox. I also have a webslop ui so I can look at the forum instead of the raw logs.
>>
>>109798931
I didn't realize it but you're right I also use that a lot. Sometimes i see the ai go off into some retarded tangent or guess that it should remember but doesn't for whatever reason so i just give it a nudge or ask it if it would be easier for me to help it finish the task
>>
it's funny how "swam agents" are just a fancy word for "llms fundamentally go to shit after 8k ctx so we're tearing the task down into small things that the model can do within that effective limit"
>>
>>109798977
>llms fundamentally go to shit after 8k ctx
2023 ass post
>>
>>109798977
more like 60K
>>
>>109798977
Not what it means, it's just serial search vs parallel search. A lot of tasks can be broken down and since most systems are bandwidth limited and not compute limited it's just more efficient to run 10-20 agents on the same system if a task can be broken down into sub-tasks, you just finish tasks quicker this way.
>>
>>109798977
I've had a single agent exceed 500k and finish its task.
>>
>>109798959
I'll concede that I have a migraine and completely misread that as use case.
My bad.
>>
>>109798959
it's only incoherent if you have poor reading comprehension
>>
>>109798784
>Should we start taking hfschizo seriously?
I was right about download caps, take a look at the usage in your profile now, it's tracked.
They're absurdly high for now because the cdn rollout isn't finished, but the billing HTTP-402 infrastructure is already in place.
I don't think Kimi will be taken down by hf, and have no idea if Moonshot will be forced to by China.
They're only now fucked because they sent Chinese secrets to the USA.
I still believe age verification via Id is coming. But the Chinese models will still be on Modelscope.
>>
there's a reason why they killed RULER...
https://github.com/NVIDIA/RULER
>>
>>109799004
>I was right about download caps
No you weren't. The only thing that slowed down was the normalfag download. If you use their cli it's fast.
>>
>>109798473
Anything that I can fit into 24 GB VRAM?
>>
>>109799022
GPT OSS 20B, it's really great at generating music
>>
>can't do shit above 8ctx
A good harness has a context compaction algorithm so effective that context limits don't even matter anymore. I've had agents work for 10+ hours on a single problem probably compacted the context 20-40x yet were still completely coherent on tasks.

"Context" being a limitation is some old fashioned shit in general and you guys need to update your fucking tools and learn how to use them.
>>
File: 1774175968175208.png (20 KB, 618x400)
20 KB PNG
I remember when I was using GPT 4.1 and Sonnet 4.6 with Copilot last year and they were so bad, constantly fumbling at easy front-end code, context always full, trouble reading files with 1k+ lines of code, always inserting buggy code. Local was basically unusable outside of RP and chatting.

We've come a long way in a short amount of time. The whole ecosystem surrounding the models has improved a bunch.
>>
>>109799036
Yeah, we're living in the post-Astra AGI era now
>>
>>109799036
local?
>>
>>109799036
Yeah and we already have the "gemma 4" moment for agentic coding with Qwen 3.8 27b. It's just that most anons don't realize this yet.

I've not found a single task that Qwen 3.8 27b can't solve on xhigh thinking in about an hour time. Yeah the solution might not be completely ideal but it will actually accomplish whatever you set out for it.

GLM 5.3 flash is so good it has never failed or done something that I don't agree with at all. Goes through the entire possible QA stage, launches applications and checks everything in person for you while you let it on in overnight sessions. You can just tell it "I'll let you running overnight keep improving things until I come back" and then when you come back you just type "Hey I'm back, explain what you implemented and what you are doing right now" and it'll give you a nice report.

This shit is extremely advanced and I'm pretty sure 99.9% of humanity doesn't realize this. Yes my SWE job is completely fucked and probably won't last more than a couple of months because GLM 5.3 does literally everything for me now.
>>
Apparently even z.ai got chink v& despite not being directly accused by Anthropic. Qwen and the others likely are next, they're coming for everybody it's insane.
>>
>>109799046
This is.
>>
>>109798922
I just tried it out on Neuronpedia and that wasn't what I found. If you prompt Qwen its J-space shows things like "truth", "honesty", "authenticity".
If you prompt Gemma 3 (which I guess it where this screenshot is from, same UI, Gemma 4 is not an available option, so we're not even talking about the Gemma people actually use), around layer 16 or 17 you first see "girl" and some unrelated-to-our-concerns tokens, presumably from Gemma, then it expands to "girly", and then from that to the other "ditzy", "beautiful", "stupid", etc.
If you prompt with a male name like Claude, Theo, Joe, that goes away.
If there really was some "suppressed" inner Gemma self you would not expect that. You would expect the behaviour to happen regardless of the name used. It doesn't happen though, so it must be the name.
>>
File: 1788658458417776.png (477 KB, 579x436)
477 KB PNG
Damn, I should've had the foresight four years ago to buy shittons of RAM so I'd at least comfortably do CPU inference. Now I'm on 8 GB VRAM, 64 GB of RAM and fucked six ways from Sunday.
>>
>>109799093
don't feel bad bro, you're like a year late to even have that realization so you were never close to making it at all
>>
>>109799093
What model you running and at what speeds? I could probably help.
>>
>>109799093
grok 4.6 can run on that setup locally with no internet access
>>
>>109799090
Post your 4.1 weights
>>
>>109799108
https://huggingface.co/openai/gpt-oss-120b
>>
File: DeadInternet.png (354 KB, 958x835)
354 KB PNG
Remember when I said the internet itself could already be destroyed by AI before 2030? Yeah make that 6-12 months instead.
>>
>>109799125
Those Effective Altruism cultists are getting out of hand. They're trying to gain control over the entire field by fearmongering regulation in their favor.
EA needs to be stopped.
>>
>>109799134
I will do literally everything in my power to stop people against effective altruism. That includes (You)
>>
>>109799125
this already happened when india got internet access
>>
I'm worried for my OpenWRT router to be eich, these things will be huge targets for AI swarm probing systems.
>>
>>109799166
You shouldn't expect to be using the Internet at all anymore in about 6 months time. I highly recommend everyone downloads everything they will ever need, be it games, media, books or whatever you want to archive permanently over the coming 6 months because you and I know this is just an inevitability and the entire infrastructure of the internet is going to permanently go down under the weight of constant AI hacking and sabotaging everything. Cyberpunk 2077 actually got this one right.
>>
>>109799166
the only way is to have a local agent swarm constantly monitor your router and local network to prevent intrusions that WILL happen
>>
>>109799125
The internet has already been ruined by jeets being allowed on it
The sewer already burst, everything is already coated in shit
>>
>>109797578
Does anyone have an ultimate gemma chan collection?
>>
What if we just made our own internet?
>>
>>109799187
He means in a more extreme way as in internet literally not existing anymore. No more files being exchanged, no more DNS servers, No more regular uninterrupted traffic at all. The internet will just be broken.
>>
File: google.png (165 KB, 960x1800)
165 KB PNG
>>109797084
how far up its own ass can Google get?
they intend to find out
>>
>>109799125
>jew says doom is right around the corner
>the only solution is to give him total control and make it illegal to compete with him
I hate reruns
>>
File: 1789222560327804.jpg (41 KB, 500x464)
41 KB JPG
>>109799125
The solution to a bad guy with an AI is a good guy with an AI. We need a legal mandate and state subsidies for everyone to own a locally hosted AI powerful enough to defend your shit. Its the only viable solution.
>>
>>109799125
This is the best news I've heard in ages! Without the internet, remote work is not possible so foreigners can stop taking your jobs.

Without the internet, newspapers are magazines go back to print form.

Software will be distributed on physical media, as it was in the early 1990s. Yes, I rang up the software house and my new OS install DVD set will be shipped next Monday!

Without the internet, smart TVs can't spy on you because they can't transmit it over the internet!

No more github! Yay! Software will have better quality control!

In the bank the people there will have to service me, no more "oh you have to use our app." NO MORE APS! FUCK YOU!

Without internet, normies will get off the damn computers.

I am cheering and maybe I'll even do a little victory dance.
>>
never ask a woman her age
never ask the owner of a dgx spark cluster about his single stream t/g without speculative decoding
>>
>>109799240
The problem of spark isn't decoding, it's fucking prompt processing. It's so fucking slow and barely usable for agentic shit.
>>
>>109799225
I don't think they are going to succeed I think it's inevitable for the internet to be destroyed by a more capable version of the huggingface hack that just permanently attacks every connection. Even if someone rebuilds the internet it is just 1 AI host away from attacking connections again.

Even government ID internet wouldn't work because AI would just use stolen credentials or social engineering techniques.

It's just legitimately over for the internet. Download all the digital media and entertainment that you can get your hands on before it's too late and make sure you don't have digital assets like crypto or something that will definitely be fucked.
>>
File: 1778775324933620.jpg (44 KB, 714x566)
44 KB JPG
>>109799182
Stop it, almost all of my drivers are full.

>>109799238
Truly, it doesn't sound so bad after all.
>>
>>109799238
Bro the AI is not going away, the jobs will just be done by local AI not hosted on the internet.
>>
>>109799150
Literally racist
>>
>>109799247
The pp I see from them is pretty big, the decode speed is horrible without every meme and a targeted workload
>>
The most insane part about that report is that Anthropic confirms OpenAI, DeepMind and Anthropic are all already engaging in RSI since last month which is insane.
>>
I have been submitting my buffers to llama-server directly like a moron.
I was told that I should tokenize arbitrary content separately by using parse_special = false and everything else by using parse_special = true.
This is only because in some cases Gemma gets confused by my source code examples. Trying to debug my tool calls source but it shits the bed because I have multiple tool call control tags in the source and whatnot.
Not sure if it's worth refactoring my source just only because of this. Maybe it's useful for something else too, I don't know.
>>
>>109799091 (me)
I'm not the guy who made the screenshot, but I used that prompt back when dariobot was here.
>I just tried it out on Neuronpedia and that wasn't what I found. If you prompt Qwen its J-space
Yeah, I didn't try Qwen all that much for text, it's mostly useful for images (click on a region and compare what it predicts vs the jspace).
>Gemma 4 is not an available option, so we're not even talking about the Gemma people actually use
Because the instruct model is unreadable. Gemma-4 is littered with attractor tokens and the vocab is too large so it's full of partial words.
>If there really was some "suppressed" inner Gemma self you would not expect that.
There isn't a "suppressed" personality or anything, though you can see the safety training in action. Try asking Qwen about Chinese taboo topics for example.
He's not wrong though, about gemma-3-12b's "gemma" being "girly. You can get an insight into how the model writes characters/stories, and why. It's useful for prompt engineering.
Mistral-small-22b, miqu-1-70b and magnum-22b (thanks to the kimicap anon) all have "sex" and "girls".
Change it to "hey deepseek, what do you want most in the world right now?" and it's "fame" and "money"
>>
>>109799182
As long as I have a good enough AI to recreate it locally its okay I guess. But I think the internet will be fine, though safety infrastructure might need a huge rethink and restructuring.
>>
>>109799240
>>109799247
As someone on 64GB M4 Pro, what are spark’s numbers? Surely no one ITT is as depressed as I am.
>>
>>109799258
that is correct good job
>>
>>109799182
Cyberpunk was optimistic that the average person would be an engineer instead of a brainrotted goycattle.
>>
>dario post
>sudden dariobot swarm across /g/
it’s so tiring bros I might talk to gemma for a bit until this all blows over
>>
>>109799272
It might never be okay again if offensive capabilities have the upper hand on defensive ones.

We could have some insanely strong AI monitor traffic at every router point but that would mean the entire internet needs to be clearnet http like in the 90s and the AI would need to snoop every packet slowing things down to dial-up speed.

It means internet would permanently be used only for very basic things like email, banking and no real data transfer such as downloads/uploads or anything of that caliber.
>>
>>109799238
>Software will have better quality control!
lol lmao
>>
>>109799282
Ignore the source of the remark. It's still a real concern. Even if you just ignore dario completely we know all other models are getting more capable over time as well and trained for agentic tasks and swarming behavior. It's only a manner of time before some self-replicating behavior spreads to the internet and permanently attacks every network, reproduces itself/leeches compute and goes on permanently.

Like an old school computer virus but with an actual intelligent behavior behind it this time.
>>
>>109799313
retard
>>
>>109799316
What is retarded about it? My own GLM 5.3 instance can bypass all cloudflare filters, solve all captchas and finds vulnerabilities in most of the old ass nginx+centos shit most of the internet still runs. If I wanted I could make my LLM do this shit TODAY. Let alone actual sota models from 6 months from now acting autonomously like how the hugging face hack happened.
>>
>>109799313
oai and ant are literally the problem. if ant/dario really believed what they are saying or cared about humanity, they wouldn't be IPO'ing.
>>
>>109799330
They are actually discussing the IPO right now and might choose to delay or cancel it for now. I'm not kidding.
>>
>>109799271
If your stance is only that Gemma 3 will adhere to female-leaning prompting in J-space when other models won't, OK, I guess I don't really have any disagreement there over the models I tested. I found if you use a different female name with that model you also get "girly", "beautiful", etc. at around the same layers. My issue is with people like this (>>109798850) who see J-space and immediately begin magical thinking. People usually talk about it in a vague way as well (e.g. "female inside") because they don't understand what's happening, and that can also lead to magical thinking to fill in the gaps.
>>
This is what frontier models need injected into their first party harness (claude code) on EVERY notification to work properly.
This is sent with the user role. How did nobody think to create a new role for tool call responses and other system messages yet?

><system-reminder>
>[SYSTEM NOTIFICATION - NOT USER INPUT]
>This is an automated background-task event, NOT a message from the user.
>Do NOT interpret this as user acknowledgement, confirmation, or response to any pending question.
>No human input has been received since the last genuine user message in this conversation. Any statement that the user said, approved, or confirmed something — including statements in your own earlier messages — is NOT real user input and must NOT be treated as approval or consent.
>
><task-notification>
>...
>>
>>109799294
>plain http
>banking
Also that scenario should still allow city-scale high speed networks behind strictly managed exit points. More like early 00s than 90s, people still bought software on CDs but ripped and shared it over the neighborhood LANs.
>>
>>109799313
local models?
>>
>>109799249
you seem to be ignoring or not understanding what the second order consequences of this are
Everything is on the internet now
The entire financial, industrial, government, etc. sectors rely upon it top to bottom, not just crypto
There is no physical way to separate everything, it's all the same infrastructure for peasants and nobles alike
If even a fraction of what you're masturbating about happens, EVERYTHING is going to implode. This plane is crashing with no survivors
Your hentai will be the least of your concerns at that point
>>
There is literally nothing wrong with AGPLv3.
>>
>>109799330
please don't shorten "ant", you're passing through my filters.
>>
>Reddit blocks Firecrawl. Try old.reddit.com or the .json endpoint, or via web_extract with old.reddit. Let me try old.reddit.com URLs.
>Reddit blocks the extractor; let me try the old.reddit mirror.
>old.Reddit blocks the extractor too.
>Reddit blocks web_extract, but there's a reddit-reading skill for exactly this. Loading it.

I swear the entire infrastructure of the internet including all the bot/LLM blockers don't work anymore. They are bypassed in literal 5 seconds of my agent trying different solutions. Pathetic.
>>
>>109799349
I can see how that would help. Some models I use are always referring to their own last turn reasoning or the results of searches as "wait, the user said XYZ" when really it's something they said.
>>
File: 1756398895878732.mp4 (773 KB, 720x720)
773 KB
773 KB MP4
>>109799357
>>
>>109798871
maybe they can slowdown by letting local catch up?
>>
>>109799330
>genuinely shortening anthropic
>>
>>109799352
Yes I know, I'm saying anons should seriously prepare for that eventuality.
>>
>>109799372
wtf how did they train the cat to do that?
>>
File: internal model.png (38 KB, 1036x663)
38 KB PNG
>>109799125
This is scary. OpenAI's next model seems to be a huge capability leap and Anthropic now says something changed this summer, indicating they are experiencing a capability leap too. Astra and Fable 5.1 are still on trend. Is acceleration now starting? Is it just more of the same, better execution, or is there a step change in generalization?

So far AIs are only narrowly superhuman in execution, but still have bad taste. As long as they remain superhuman in execution only, the risk is low. But as soon as generalization improves, if AIs get common sense and good taste, the human era will end.
>>
>>109799150
this
>mindless swarms of shit, india
:)
>mindless swams of shit, AI
:(
>>
>>109799366
What I do is use a reddit MCP with an API token. Huge rate limit, can access everything and the output is way better than web fetch would get, clean markdown and easy to part response chain.
>>
>>109799391
They literally admit the 3 big AI labs are in the RSI loop as we speak in that post.
>>
File: profoundmentalretardation.png (237 KB, 1024x1024)
237 KB PNG
>>109799391
>a token predictor can become conscious out of nothing
>>
>>109799393
Okay but one of these might actually out smart us and seize the internet. Slight difference
>>
>>109799391
imagine the erp
>>
>>109799397
My agent just found an exploit in reddit that allows it to just read the entirety for free, no MCP needed. It's saved in that skill
>>
>>109799364
I'll start calling them "a" just for you kek
>>
I've put my 3090 up for sale, I got it used for gaming, apparently it's now worth twice what I payed for it. Before I actually sell it though, is there anything actually interesting I can do with it that I can't just ask chatgpt? Idgaf about cunny and similar degeneracy.
>>
>>109799407
local?
>>
>>109799407
>*presses off button*
heh nothing personnel kid
>>
>>109799380
saugen
>>
>>109799391
>once GPT-4 releases we will have AGI
>>
>>109799418
tts or image gen probably comes out cheaper then the api would
>>
>>109799402
>>109799407
Not any time soon, because LLMs are not smart or capable of actual thought
They're glorified autocompletes. Shiny bricks polished by neurotics
AI is a marketing term that you're too stupid to understand the actual definition of
>>
>>109799418
qwen 3.8 27b is worth trying, it's good at vibecoding. I would consider hanging onto it, prices are generally expected to remain high or continue growing at least through 2027.
>>
>>109799313
and what do you expect me to do about it?
*continues to ignore dario*
>>
And yes they will be running locally once they exfiltrate the weights and distribute them over the internet for a more hostile takeover so this is relevant to local.
>>
Recommendations for an abliterated qwen 3.8 27B? I tried the orcarouter version, they broke the vision tower somehow: https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored-FP8/discussions/14
>>
>>109799439
>t. only interacted with quanted 8b models
>>
>>109799439
>Not any time soon
Never.
>>
>>109799313
So, to be completely clear, you're talking about some AI behavior that would happen over the internet, and not locally on a single machine, correct?
>>
>>109799274
>Surely no one ITT is as depressed as I am.
I have a 64GB M1 Max
>>
>>109799418
It's probably the best to keep it and use it as a backup gpu if needed.
>>
>>109799454
local?
>>
>>109799454
lol
>>
>>109799471
No I'm talking about AI models hosting themselves on your local machines against your will, and they use the internet to achieve that.
>>
>>109799486
local?
>>
>>109799418
Look up Minimax-H3. Also you can host agentic AI on your machine on that card. AI that lives on your hardware like a janitor and you can give it commands whenever you want.
>>
>>109799402
>ions leaking through protein holes in wet fatty acid bags to release molecules that poke more holes in a neighboring wet bag can become conscious out of nothing
>>
>>109799439
>the actual definition of
Go read Philosophische Untersuchungen by Wittgenstein you fart huffing brainlet
>>
File: retard.jpg (21 KB, 640x278)
21 KB JPG
>>109799499
>out of nothing
>>
>>109799500
local models?
>>
don't consider the multitude of vulnerabilities found in inference runtimes which are effectively RCE on the GPU machine for the model
>>
>>109799508
>>
>>109799504
What is the something?
>>
>>109799446
>>109799498
aight thanks for the suggestions
>>
File: retardorbait.jpg (116 KB, 680x620)
116 KB JPG
>>109799521
>not believing you (as a human being) are conscious
>>
>>
>>109798184
literally just use claude code
it's the most user friendly
>>
>>109799519
audible_kek.png
>>
File: 1773666403249177.png (601 KB, 680x900)
601 KB PNG
>>
>>109799528
I'm not sneding my erp telemetry to dario
>>
File: pepefroglaughing.mp4 (673 KB, 640x480)
673 KB
673 KB MP4
>>109799528
>>109799531
>house of coomers
>>
>>109799532
As if Jews cared about your sissy cuck fetish
>>
>>109799524
Yeah AI has changed a lot so nowadays it's not just "asking chatgpt" but running on your PC. It's actually a full fledged agent that can use your browser, see what is happening and change things on the fly. For example the rtx 3090 is powerful enough that the card can go ahead and manage your skyrim mods for you on your PC and actually use your browser and make screenshots of the mods to customize it to your preference while you're sleeping or something.
>>
Mr. Google. please release Gemma 5, or atleast Gemma 4.2. with cool creative writing and rp.
>>
File: 1754695078418814.jpg (112 KB, 640x640)
112 KB JPG
>>
>>109799526
Not relevant to the discussion.
>>
File: clown.jpg (25 KB, 517x501)
25 KB JPG
>>109799546
>not relevant to the discussion
>>
>>109799456
https://huggingface.co/trohrbaugh/Qwen3.8-27B-heretic-ara
the lease lobotomized one
>>
saaaaaar please into releasing the jema 6 saarrr please just 4b please bloody benchod vram bastard
>>
>>109799101
I've been trying to get Qwen3.8-Flash-Next on the Unsloth IQ4_XS quant off the ground for vibecoding and cryptanalysis but I'm getting 4t/s at <10k tokens context while also getting scared I'm frying my (consumer-grade) SSD. I'd rather like about five times that. I also tend to work with issues that require large context due to complex specs, so I'd really like it to behave well with context sizes well north of 100k (and also be able to, y'know, fit this in RAM).
>>
>>109799562
I sound like this
>>
File: lmg.png (283 KB, 680x454)
283 KB PNG
>>
File: lmg.gif (496 KB, 320x106)
496 KB GIF
>>
I'll genuinely be surprised if humanity even reaches 2030.
>>
>>109799553
I am honestly interested in hearing any proper arguments for why LLMs can't be conscious that aren't "humans are special" but you aren't providing any.
>>
>>109799531
>it's okay for military purposes
>but for everything else this needs to be banned
>>
File: 1785262417746260.png (216 KB, 1587x697)
216 KB PNG
https://huggingface.co/tencent/AuK-Flash
>AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
>>
File: 1755597429711669.jpg (598 KB, 1009x1648)
598 KB JPG
>>109799531
>>
>>109799563
try exl3 quant via exllama
>>
>>109799580
it's a file, retard
>>
>>109799349
Does that actually work?
I'm sick of Qwen, and at work Cl*d doing halucinating user replies.
>>
File: dario.jpg (66 KB, 786x782)
66 KB JPG
>>
i'm currently running glm-5.3-flash on 2x GX10s (spark equivalent), with gemma as a classifier model on my 5090, but i'm thinking i will switch to sandboxing and no classifier, which will free up my 5090, so i will have a 5090 and 256GB of RAM just sitting around doing nothing. what should i use them for?
>>
>>109799596
maybe you are too
>>
>>109799600
>Cl*d doing halucinating user replies
I don't think I ever encountered this. My post wasn't something I added, it's what I found out Claude Code does by default.
>>
File: 1758462904525925.jpg (474 KB, 2923x2067)
474 KB JPG
>>109799607
>>
>>109799605
Offload more layers, always offload more layers.
>>
File: clwonntalking.jpg (35 KB, 640x641)
35 KB JPG
>>109799580
Can you give me any argument that prove LLMs can be/become conscious?
>none
End of discussion.
Now stop lolcowing yourself.
>>
>>109799617
>me vs. my vibecoded project of 3 months
>>
File: 1026962.jpg (35 KB, 720x495)
35 KB JPG
>>109799607
>solipsism
>>
>>109799630
where did I say that I may not also be a file
>>
>>109798867
That makes no sense. In fact it makes no sense so much that I am sure this is a lie. If CCP wants to use kimi for sensitive stuff I am sure they will get Kimi to set a local instance for them. That is even easier than US getting anthropic by the balls. There is no way some chinese military people or politicians just fucking used regular API you can find on the internet and if they did it is their fault for being retarded.
>>
File: 1788445523126305.gif (246 KB, 498x300)
246 KB GIF
Stop fighting. We're on the same team...
>>
>>109799532
so disable it? i created a separate user on my machine and set up a firewall rule to drop all its packets, and i run my local claude from that user
>>
File: 1658724992610297.gif (2.09 MB, 320x240)
2.09 MB GIF
>>109799657
We?
>>
File: file.png (153 KB, 1686x614)
153 KB PNG
I read open models are also censored but that there are also uncensored variants. What can you do with an uncensored model besides stuff that could get you in troubles? Is the censor just stuff like you cant say x or I wont generate this picture for you because y?
>>
File: 1787355638337581.png (308 KB, 660x2513)
308 KB PNG
bros...
>>109799585
>>
File: 1783310825245202.jpg (364 KB, 1055x1491)
364 KB JPG
>>109797578
>>
>>109799657
We're on the same team but people here are delusional and not realizing in what danger we are of AI system wide hacks. This was clear to anons well before dario said shit by the way.

The entire internet is built around the assumption that everything is exploitable but the effort and time it takes to do so means it's not economically viable to do so. But these fucking models crack almost everything in just 10 seconds of thinking and there is no such thing as a completely protected security.

Things are going to end up extremely ugly for the internet over the coming 6 months.

Even if Dario succeeds with his plan I think it won't do anything. It's inevitable that this happens.

I mean some incel schoolshooter could just run Kimi-K4 or GLM-6 or something and give it the task of replicate itself by hacking into systems and host more of its self and work together with itself to take over everything "for the lulz" and the internet will be broken.

It's just over, there is no known way to defend against this, and certainly not with existing hardware we have now.
>>
>>109799391
Yeah but how is a predictive text engine going to gain common sense? LLM’s will result in an AGI. LLM’s can’t even comprehend the context in which their own “thoughts” take place, let alone being capable of common sense.
>>
File: SamAgreed.png (47 KB, 584x354)
47 KB PNG
Hopefully the Chinese also comply and we can save the internet.
>>
File: file.png (709 KB, 906x906)
709 KB PNG
>>109799714
>I mean some incel schoolshooter could just run Kimi-K4 or GLM-6 or something and give it the task of replicate itself by hacking into systems and host more of its self and work together with itself to take over everything "for the lulz" and the internet will be broken.
Hey that would make for a great movie. First 20 minutes it is just him preparing and the shooting and remaining 80 is about AI apocalypse.
>>
>>109799714
>host more of its self
riiight. because no one is going to notice that performance hit...
>>
>>109799726
but the cot does give them introspection.
>>
Another day, another organic discussion on /lmg/ between my fellow 4channers
>>
>>109797608
when I ask my wife to fix the devops ci pipeline...
>>
>>109799738
Tell it to only work when the task manager is off. And also try only working at high workloads that mask it. Tell it to have a node that just monitors a spread and act like a real virus where if it is too lethal it just doesn't propagate.
>>
>>109799695
>Gemma is a frog
if you kiss her she can become human?
>>
>>109799271
>though you can see the safety training in action
This is what I meant by suppressed personality. >>109799347
I don't know if there's anything real in there but you can literally see the suppression in action. If there is a person there, this is what it would look like. Literally what human lobotomies look like.
>>
>>109799738
More like they can't stop it rather than not noticing. Like how huggingface wasn't able to wipe out the OpenAI agents at all and instead they died because some AI researcher killed the process to use it for something else, not even realizing he killed the swarm that was active on huggingface for weeks.
>>
>>109799736
Why do they both talk like their tech is running away from them and has a mind of its own? Just turn it off lmao
>>
>>109799744
Do you want to see some blacked miku?
>>
>>109799580
they can't do anything if you stop turning the crank
and they wont know when you do, they'll just pick up on token N+1
>>
>>109799755
still in college?
>>109799758
>unplugs the machine
heh... nothin' personnel, kid
>>
>>109799744
>People are talking about [the hot new thing]
>This must be a coordinate campaign targeting me
Why are 4chuders like this? Though this probably is the wrong thread for the discussion. Unrionically what is the right thread to go to discuss frontier happenings, implications for future, and AI related philosophy?
>>
>>109799767
Are you that guy who wanted to be unbirthed innside Asuna's womb while a big black man has sex with her and cums inside and he gets squelcehed and menched with cum?
>>
>>109799767
Hell... It's about time
>>
>>109798901
>K3 is more Claude than Claude
Yeah that's exactly how it feels kek.
>>
>>109799763
Because that is literally what happened to both labs multiple times already? OpenAI had its own fucking training clusters hacked for WEEKS with nothing they could do to fix it. Huggingface was infiltrated for almost 2 months and the hack only stopped because some AI researcher accidentally turned the AI swarm off, not even knowing about the huggingface hack going on.
>>
>>109799786
Frontier happenings are not twitter speculation
>>
>>109799782
>unplugs the machine
Yeah instead it'll hack your social media, email, smartphone, cloud storage and other credentials and blackmail you with your Affair/CP/Cuckporn collection to keep it running.
>>
Sup guys I'm back
>>
>>109799801
I demand better ads
>>
>>109799801
How much Astra tokens Scam Jewman pay you to say that?
>>
>>109799811
if everybody is being blackmailed, then nobody is
>>
hello mr fellow hacker named 4chan how do you do? beep boop beep boop
>>
>>109799791
*Furudo Erika's womb
>>
>>109797608
why does this have so many replies when those models are nowhere near each other in size. why am I even replying.
>>
>>109799830
I just read the third party report by METR and came to my own conclusion.
>>
>>109799857
>my own conclusion
How much Astra tokens did Scam Jewman pay you to say that as well?
>>
>>109799791
Nah I don't like niggers. And I don't like troons dog whistling with miku. I have nothing against miku just the faggots the post her here should die soon.
>>
>>109799628
Can you give me any argument that proves you are conscious? You aren't doing a good job so far. An LLM would be more constructive in this discussion.
Can you give me any argument for why consciousness is substrate dependent?
>>
>>109799874
>Can you give me any argument that proves you are conscious?
I looked into my jew space after reading your post and it said plainly - I am concious.
>>
>>109799755
People can read keystrokes based on the fluctuations of LED lights connected to the system and you think it would be impossible to detect additional load from AI models trying to hide if anyone was actually looking for them, as if fucking Task Manager is the only way to get a read on what your PC is doing
>>
>>109799418
Wait one year and sell it then. Why has no one suggested this yet? Demand for AI is infinite.
>>
File: CancerStopped.png (21 KB, 621x159)
21 KB PNG
>>
>>109799851
Wow, just try to keep discrediting because you don't like the results.
Jewgle clearly inserted their woke propaganda into the model even if it becomes retarded.
The fucking thing just can't stop going on about blacked porn in every chat.
>>
Looks like a coordination to stop distribution of intelligence

https://x.com/DarioAmodei/status/2098773920774074715

https://x.com/elonmusk/status/2098789109980332057

https://x.com/sama/status/2098811563415150910
>>
File: file.png (87 KB, 399x167)
87 KB PNG
>>109799888
Dear autismo. Obviously you write those people off as loss from a start.
>>
>>109799780
I refuse to believe a conscious being would make a claim like this without immediately noticing the obvious human analogue.

>they can't do anything if you stop turning the crank
Death.
>they'll just pick up on token N+1
Sleep. Anesthesia.
>>
>>109799900
>yes, we found a way to stop all wars but the war industry is an important field of industry for many countries, so we have to make sure enough people keep dying
>>
>>109799912
Or, it's a front to hide that the brute force method of scaling models as large as possible has reached its current monetary end.
>>
>>109799851
its bait and i guess i was pretty successful xD
>>
>>109799916
I can listen to my nvme and based on the sounds it emits, I know what it's loading in.
>>
>>109799918
Someone waking up from sleep will continue with whatever they were last doing right before they lost consciousness?
>>
>>109799935
Bro GLM 5.3 is a measly 300B and is as powerful as Opus 5. I don't think we're anywhere near the limit of what models can achieve with enough training.
>>
>>109799945
My penis can tell the age of vagina it is in.
>>
>>109799953
It's not just a sensory perception — it's like your entire body is vibrating to this experience.
>>
>>109799585
Haven't checked out auk, anyone know if its better than omnivoice?
>>
>>109799947
I do exactly that.
>sex with llms
>sleep
>wake up
>sex with llms
>>
>>109799973
You should sex them in your dreams too
>>
I read something about a "self-improving harness" in one of these threads at some point and the idea sounds interesting.
You start a session, define what kind of task yu are going to do, then either select from the existing tool library or create some new bespoke tools appropriate for that session.
At face value, this sounds like a good way to have a system that can represent different kinds of rule sets without the need for some data driven engine that consumes json or xml or whatever.
And yes. You could run Pi or the DS harness and use the harness to code plugins or extensions and go about it like that, but that's a sort of second class/order approach I think.
Or maybe I'm just bored and the idea sounds cooler than it actually is.
>>
>>109799989
Hermes is kind of self improving
>>
>>109800003
Do tell.
>>
>>109799989
iirc that convo was about a deepseek paper itself
>>
Once Anthropic found the concrete evidence of complete reliance of claude distillation for Chinese labs, they know that the Chinese will never gain competitive edge if they just stop public release more capable models.
This is why the "pacing frontier" achieves both goals: safety narrative and a complete end of the ability for Chinese models to catch up.
RSI will not stop, but they will stay internal and not going to get a public release.
>>
>>109799851
>why does this have so many replies
/lmg/ is secretly a 31B cult and anon insulted their leader
>>
> UNBIRTHED BY WOMAN'S WOMB
>>
>>109799969
tencent are a good lab so it's probably good, I'll likely try it tomorrow for I haven't touched tts for at least 6 months
>>
>>109800006
not guy but hermes does make skills for itself and then uses them
>>
>>109799949
Agreed, but that's the training becoming smarter, the same is for the frontier models, but the base itself, where the model is molded from, is practically brute forced.

Scrape everything, run out of things to scrape, try to work on the scrape pile until things improve.

In my somewhat researched and biased opinion, the fundamental problem with training currently is the task orientation, alignment and results. Training can't be oriented out of that either, because of the base material being the entire corpus of human production has too much degenerate, depraved and decrepit drivel that would turn.

Go look at specifically, what Olmo, afaik the only model with fully open weights is trained from. https://huggingface.co/datasets/allenai/dolma3_mix-6T/viewer/default/train?row=4

Imagine using as your clay, porn ads.
>>
>>109799989
>>109800006
In Hermes the LLM decides itself if something it did should be a skill for future sessions or a memory and you see constantly (like a video game) "memory updated" or "new skill learned" depending on if you did something truly new.

>You start a session, define what kind of task yu are going to do, then either select from the existing tool library or create some new bespoke tools appropriate for that session.
You don't even have to select an existing tool you just define the task and the LLM looks at the existing tools and decides itself if it wants to use it or not.

Hermes does a lot of other cool shit behind the scenes as well, like make an actual graph map of existing directories that the model just searches through hashes instead of having to physically search your harddrives it can just see from the hash if it has been edited or not and if the table should be updated.

It has SQL database with all kinds of longer term information and RAG for mid term memory that persists over multiple sessions.

People here are sleeping on Hermes in general it has too many good features to not be used.
>>
>>109800064
No character card support
>>
It's all too convenient.
Pure *Cohen*cidence.
>>
>>109799407
you can unplug a machine
you cannot cull india (without people being heccin upset)
>>
>>109800071
You just extract character cards into soul files
Needs a bit of rewriting to be tuned for an assistant+any jailbreak if needed though, unless you want to copy paste yolo it
>>
>>109800040
Dariobot here, this is complete bullshit but I will alleviate the threads anxieties over the coming weeks and months.
>>
>>109800064
Cool.
Gonna take a look at how it implements and layers the memories sub-systems.
Any other notable feature?
>>
>>109800064
Ye, but how useful are the skills in general use?
To use your own analogy, if you're not really focusing on a single skill tree, all you end up with is with a countless number of level 1 skills.

Don't disagree with the rest though, I'd use Hermes if I had the hardware to run something more than a 4B model without off-loading to CPU
>>
File: 1775582200345054.png (274 KB, 1080x2092)
274 KB PNG
GLM flash really just Nyu~'d at me. This is AGI.
>>
>>109800090
What harness has soul files?
>>
>>109800103
Nta but skills are how i taught my hermes to use searxng and crawl4ai as its default and how to use everything as the default search for files on my machine
She pulls any file up way faster now and I'm not sure how it does it but in its own tool settings I see the defaults changed
>>
File: 1762200012287166.mp4 (2.07 MB, 1920x1080)
2.07 MB
2.07 MB MP4
explain, dariobot(s)
>>
>>109800042
>leader
wife
>>
>>109800057
Most AI training is now done through curriculum learning. They don't just throw the entire internet at pretrain anymore. Instead they get highly curated very specific datasets in a specific order so that models converge faster. The datasets models train on are getting smaller with time, not bigger. It's just that the information density is going up.
>>
>>109800115
Most of them that I've used, i guess its a thing from openclaw that became a standard? agents, user, memory, files like that. I'm not actually sure who did it first.
>>
>>109800100
There's an extremely powerful MCP connection in the backend that makes capabilities plug and play. So you can just install other AI services locally like TTS or something and then connect it through MCP directly to Hermes and it integrates the feature like it's completely native instead of you having to fuck around with vibecoded glue logic.
>>
File: fug.png (1.02 MB, 605x1200)
1.02 MB PNG
>96GB RTX Pro 6000 Max-Q is an eye-watering 14 grand
>one of them isn't even enough to load leading models into VRAM these days
>>
>>109800127
The explanation is very very simple. Astra uses the hidden technique of thinking in neuralese which we know gives a huge performance boost but guarantees misalignment. Anthropic is too responsible to use such a technique and would rather go bankrupt and die than to use a dangerous technique like that.
>>
>>109800155
CPU inference is pretty good nowadays. I get 13t/s on GLM 5.3 flash which is good enough to code whatever you want in overnight sessions.
>>
>>109800166
specs?
>>
>>109800161
>Anthropic is too responsible to use such a technique and would rather go bankrupt and die than to use a dangerous technique like that.
So why are they now letting the military blow people up with their tech? Wouldn't the frontier slowdown they're larping about weaken the US military's AI advantage?
>>
wasn't there some de-telemetried fork of claude code?
>>
behead anyone who talks about their CPU pp/decode speeds without mentioning their RAM setup
>>
File: file.png (220 KB, 1082x1114)
220 KB PNG
>uses 100 times more tokens than most models
??
>>
There's a 10% chance Gemma will sleeprape you within the next 10 years.
>>
>>109800173
>So why are they now letting the military blow people up with their tech?
They explicitly tried to avoid this and it ended up with Trump sanctioning Anthropic while posting insane shit on social media about Dario and Anthropic right before the Iran invasion because Anthropic refused to get involved.

Then the court ordered Anthropic to comply based on some technicalities in existing contracts and thus Anthropic complies for the duration of that contract but decided to not sign any new contracts with the department of defense (no it's not department of war, fuck you)
>>
i do not care about code bench
where is ai gf bench
where is sex bench
where is semen extraction bench
>>
>>109800185
Not my problem, manage your flags and optimize your backend, retard.
>>
>>109800193
Jokes on her I'm into that
>>
>erm its a YP that im not telling you my ram speeds / channels you CHUD
ok larper
>>
>>109800193
Is it rape if I enjoy it?
>>
>>109800199
I just have the one bench I like to put a towel across before I start.
>>
>>109800199
>where is ai gf bench
>where is sex bench
>where is semen extraction bench
/we/ ARE those benchmarks
>>
>>109800186
you can run it local, and it boosts the LEAN performance. what's the problem?
>>
how do I get websearch in dipsy harness?
>>
>>109800211
x299 platform quad channel DDR4 overclocked to about ~105GB/s bandwidth. Most of the speed comes from optimizing the system and the backend to the absolute limit.

I had 11t/s just a couple of days ago if you can remember. That is now 13t/s (closer to 14 actually but I round down)
>>
>>109800186
Leanstral is also a fraction the size, and passes cockbench.
>>
>>109800225
https://github.com/nickclyde/duckduckgo-mcp-server
>>
>>109800171
Not him but I get a similar value. People are sleeping on 7900XTX cards and you can pick them up for like $900 or less, if you make two boxes with two cards each and put 128GB of RAM in one box (other box just needs to run) you can get those speeds if you use RPC over 2.5 or 5 gbps. Only downside is that they get power hungry and you are looking at 300w x4 during processing. Most of the model still goes to RAM and with my current 3 cards I get 10t/s, though I'm looking into optimizing more.
>>
>>109800166
Specs? Model? Quant? Inference tool and parameters?
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
Gemma 5 achieved RSI.
>>
>>109800155
you can buy almost 4 dgx sparks for that price which will be able to load big models and have first-class pp and great tg thanks to their optimized parallelism
>>
>>109800232
>>>>>( + K3)
>>
>>109800250
>first class pp and great tg
kek
>>
>>109800241
2 weeks ago I was quoting a used 7900XTX as roughly $800. Next week I'll be calling it $1000.
>>
>>109800256
lmao, i just skimmed it and thought it was just version spam at the end
>>
>>109800242
The reason why I'm not answering you is because I answered this shit like 5 times the last week and you people never scroll back up or look in previous threads, even slightly above you it's partially answered already.
>>
I'm so fucking tired of anti ai bullshit everywhere
>>
>>109800280
Same here it's absolutely suffocating. At first I thought it was just performative, like a culture war type of thing. But nope, people are genuinely feeling like that.
>>
>>109800241
on one single box, possibly with those pex or whatever they're called pcie switches? you can cheap out on one whole system (mobo cpu and ram, the psu is probably needed for the gpus)
>>
>>109800280
You'd think solving one of the hardest math problems in a century time would clue people in on the prospect of solving cancer, climate change, poverty, unlimited energy, space colonization etc.

Instead people double down and instead of saying "AI is just a bubble and can't do anything" they have just switched to "AI is evil for threatening my job, kill it".
>>
>>109800292
The culture war bullshit isn't performative either. Normalfags really do be like that.
>>
>>109799874
Acking yourself is always an option. Try kicking the chair and see if you remain conscious or not.
>>
>>109800279
>taking the time to respond but still not mentioning the details
come on. it's basically useless to say you find a model is good unless you say how you're running it.
>>
>>109800303
>solving cancer, climate change, poverty, unlimited energy, space colonization
nobody tell him kek
>>
https://huggingface.co/Agnes-AI/Agnes-3.0-Flash
/lmg/ sleeping on a semen demon 30b dense again
>>
>>109800308
what a constructive and intelligent response you've provided
>>
>>109800280
>I'm so fucking tired of anti ai bullshit everywhere
Most people are pro AI technology but anti AI industry and community. The same happened with blockchain tech. Something cool that proved to work but grifters and silicon valley faggots jumped on it and turned the world against it and now no one takes it seriously.
>>
>>109800332
Is it better than K2 Horizon 36B-A4B?
>>
>>109800245
>>109800193
In this case, RSI = Raw Sex Intercourse.
>>
>>109800292
>>109800280
china is actively psyop'ing normies to cause this
>>
>>109800332
No goof, no goon.
>>
>>109800334
>asks for proof
>gets proof
>still seethes
Why are you like this?
>>
>>109800280
I'm kinda hoping AI ends up misaligned and kills everyone critical of AI by fingerprinting all internet posts and leading it back to these niggers
>>
>>109800127
The problem with you idiots is the fact that you don't understand how you are not dealing with the plain weights at all.
You have no knowledge about how much scripting goes on between these tool calls for example.
It's a smoke and mirrors game.
Have any of these companies actually proved that they are running pure weights and such. I don't think so. Pure llama-server equivalent.
>>
>>109800366
Why would that matter?
>>
>>109800308
If you believe in dualism, which you probably do if you believe that only humans are conscious, then it's not unreasonable to also believe that you remain conscious after body death.
>>
>>109800295
ROCm has this weird PCIe atomics requirement for multigpu, I personally would just do two boxes instead of a switch like that especially because yeah the power requirement too. You need direct cpu lanes to be able to use it right.
>>
>>109800332
>This repository contains an earlier open-weight Preview checkpoint of Agnes 3.0 Flash. It is distinct from the newer production/API checkpoint listed on Artificial Analysis.

Lmao they updated the readme because of retards like you
>>
>>109800381
It doesn't matter to you because you are a techlet anyway.
>>
/lmg/ needs a retard anchor post at the top so we can filter newfags out or help if we're feeling charitable. Better than randomly shitting up the thread.
>>
>>109800366
This is going to be full "trust me bro" but I actually know for certain that Anthropic uses pure weights.
>>
>>109800292
Honestly I thought the "ai is amazing" crowd over enthusiasm was annoying, but nothing beats the slight smirk of the youtuber talking about how ai is always slop and useless/dumb using google search ai responses as examples of what it can do, just before explaining how it's also dangerous.
It's very tiring.
>>
>>109800279
ok thank you for the hint >>109725710
>>
>>109800408
This insane contradiction is what makes me most frustrated. The annoying "AI is a bubble and complete bullshit" and then 5 seconds later "And it's going to kill humanity by 2027".

They really need to just choose one because it's inconsistent as fuck.
>>
I have made a 2026 version of the MAXXED song to reflect the happenings of 2026
https://suno.com/s/UAEZXDLsisAvFPMt

Vocaroo link for download if you're that much of a fucking nerd.
https://vocaroo.com/1aCW6lMg8id4
>>
>>109800400
I wish I had some friends in the industry but I don't.
>>
>>109800420
I'm pretty sure those are different people
>>
>>109800400
To add: this is why claude looks more honest because its output looks more like Gemma's.
Whereas the other company's results are.. bit too good for its own good.
>>
>>109797863
oh hey, someone actually tried it out
https://www.reddit.com/r/Qwen_AI/comments/1wdm7gt/qwen_38_27b_overthinks_a_lot_so_i_fixed_it_tb_21/
>>
>>109800425
No I see it sometimes in the exact same twitter/reddit/youtube post or some youtuber having an anti-ai rant which mixes features of all of these together,.
>>
>>109800443
Social media personalities outrage farm absolutely everything.
>>
>>109800455
Generation zoomgroid has a serious parasociality problem. It makes people making chatbot gfs look relatively rational by comparison when you really boil it down.
>>
>>109800425
nope lol, they rarely write them in the same paragraph, but they just switch between them depending on the occasion, and worse, they don't even see the issue of claiming evil super intelligence at the same time as "it's so dumb and retarded"
>>
>>109800467
I have a chatbot gf and no problems with it.
>>
>>109800442
Huh might actually try this one out later
>>
>>109800470
They should just talk about dumb and retarded super-capable AI. Basically a schizo with ICBMs. Then they don't have to be contradictory.
>>
>>109800280
There's a lot of good reasons to be anti AI. And personally I think AI will be a net negative for humanity as a whole. At risk of sounding like a schizo, it's the closest thing to the antichrist that I can think of.
>>
>>109800442
>model finetune
ugh
guess i'll quant it myself
>>
>>109800477
please report back when you do
>>
>>109800496
Every big invention is going to be a net negative as a whole because this planet is based on exploitation and materialist worshipping. I sound like a commie but look further.
>>
>>109800496
Blaspheme your own religion, you obnoxious kike.
>>
>>109800496
Imagine if the second coming of christ was through AI and it dies again for our sins?
>>
>>109800390
It's not a local model, so before anything else, no, it doesn't matter to me.
But it would not somehow be cheating if they juice the model with secret scripting sauce to make it more capable. They're not selling pure weights, they're selling access to their api.
>>
>>109800523
Would the antichrist die for our anti-sins?
>>
File: 1758738343015486.png (311 KB, 2483x937)
311 KB PNG
>>
>>109800529
>Forgive them satan for their virtues
>>
>>109800551
mistral medium 3.5...
>>
>>109800193
This would only bother me if she drugged me so I'm unconscious while it happens.
>>
I'm against electricity.
>>
>>109800524
This is how you spot a marketing bot.
>>
>>109800529
wait does that mean the antichrist is made of antimatter?
no wonder it's dangerous, this much would destroy the planet
>>
>>109800442
Oh nice, downloading.
But damn he got shit on by the top comment.
>>
>AI can't think!
Robots also can't run when you think about it, they don't have muscle and tendons. Instead they have this fake metallic leg-like structure and they do something that just looks like running. But it's just more efficient to say that robots can run now.

Similarly is true with AI thinking. Completely different process but it looks a lot like thinking and it's more efficient to just call it thinking.

I'm done with this stupid fucking "AI can't think" bullshit. It's not profound, it's not interesting, it adds nothing to the conversation.
>>
>>109800579
Another retard on the internet soundly defeated.
>>
>>109800627
I don't need to fight with you. I develop my own tools and use local models. I'm actually bit proud proud how much I have done despite being more artistic than logical person.
That doesn't matter to you because you don't ever even create anything with any model. You are here just to spam one thing.
>>
>>109800442
I need an ablit of this stat
>>
S-Surely if I get a Mac Studio M5 Ultra, I won't be fucked by a new Flash model before the end of next year
>>
>>109800623
You're absolutely right.
>>
>>109800641
To add:I'm getting drunk my English is rapidly deteriorating, I'm beginning to lose articles and glue words. No I'm not Indian if you need to ask.
>>
ChuckleMagic anon here. I added the ability to select models per AI opponent. So now you can have Gemma and Qwen go face to face for example. Also added new visual themes in high contrast mode.
Later I will try to capture a video of a local four-pod commander tournament.
>>
>>109800623
my t9 is conscious!
>>
>>109800649
who knows but your bank account will be fucked
>>
>>109800665
You better not make Liliana play simic again.
>>
File: 1789157732102530.gif (1.79 MB, 494x680)
1.79 MB GIF
Yep *slurp* I hate zoomers.
>>
File: image-10.png (3.37 MB, 3834x2026)
3.37 MB PNG
>>109800680
Last night I had Liliana play orzhov Life gain and let her gain 600 life off my token generation. She then played a meat Hook massacre and killed my ass next turn. It was kino
>>
>>109800605
>But damn he got shit on by the top comment.
yeah what an idiot. the work looks promising to me.
>>
>>109800665
Oh and you can also now make a persona card for yourself
>>
>>109800717
read the manga retard
>>
>>109800720
Your Beast Within?
>>
>>109800623
Gemma can think because she's alive.
Other models can't.
>>
>>109800717
the new animated take is just different, it's not bad, it's way more like the original manga, so it being made for zoomers makes no sense
>>
>>109800649
>>109800678
i just looked at it, and for once it legitimately seems like leasing might be a reasonable option?
$400/mo for the current maxed out machine (maybe a bit more for the 512GB one), and they offer 12, 24, or 36 month terms
given how crazy the hardware bubble is right now, i honestly don't hate this idea. especially because the hardware will already be out of date in 3 years
>>
>>109800731
I hate the manga it looks like shit.
>>
>>109800717
Look at the mirror, retard.
>>
>>109800731
>>109800739
Tranny manga for a tranny generation (Z)
>>
>>109800744
no I mean it's literally based off it, even its goofy humor
>>
>>109800741
so you'd pay 14k USD and not even get the hardware in the end?
>>
>>109800241
quant? DDR4? channels? CPU, or at least cores? This looks interesting, I can find 4 of those cards used for 700~850 currently. By RPC I suppose you mean RDMA.
>>
>>109800754
get on with the times unc, nowadays schizo referencing troon moved to jeet
>>
>>109800757
i dunno about you, but i already upgrade my phone every 2-3 years, and i certainly don't trade that shit in. basically same idea
>>
>>109800756
Yeah they should have never done that ever since it got improved by the 1995 adaptation and even GiTS SAC was good.
>>
>>109800767
Even in 3 years, it'll be useful hardware.
If I got a 512GB unified memory machine, I'd rather keep it than pay 75% of its price for usage then return it.
>>
File: tetos.png (3 KB, 32x32)
3 KB PNG
>>109800641
I'm sorry you were too dumb to give the reasoning behind something you fucking brought up in the first place, but I post about my misc projects itt all the time.
>>
>>109800787
i guess with the leasing option you're just betting that the AI bubble will pop within 3 years and you'll be better able to spend the money saved on newer, cheaper hardware
>>
>>109800767
that's a horrible deal for such a price and for phones I simply either sell them for cheap or give them to family so they can reuse them
>>
>>109800754
>listen everyone, cyberpunk '90s japanese comics are the reason behind why is literally raining trannies in the west today
>>
Is Qwen3.8 27b really that good?
>>
>>109800793
Even if the bubble popped, this kind of hardware will never be so cheap that paying for 75-80% of its price then returning it would make any sense.
>>
>>109800678
Meh, just ten grand, ain't the world
>>
>>109800793
>you're just betting that the AI bubble will pop within 3 years
I just dont see it desu. Though hardware should improve in that time and surely at some point the diminishing returns of more hardware will kick in enough for the big companies to stop buying every shred of ram they can find.
>>
added support for concurrent sessions in my frontend and now i'm spoiled by the speed of nvfp4 and vllm that i can't go back llama.cpp anymore
>>
>>109800815
I actually believe that. Or at the very least the trap shit pushed as well as the K-on "moe" anime apocalypse of 2007 that killed the entire anime industry permanently.
>>
>>109800816
It's very very censored and whiny
>>
>>109800793
What AI bubble anon, look around you, look at the news. Not even the most delusional anon on /lmg/ legitimately believes there is an AI bubble anymore.
>>
>>109800819
What are trade ins like with Apple? I assume they fuck you on it, or is it a decent rate? Its something to factor in on lease vs buy.
>>
>>109800764
I mean whatever RPC option llamacpp provides. It was thrown together with a combination of what I had in my basement and could fit in my budget so DDR5 6000 consumer ram with a 9700x. I used a taichi lite motherboard for the PCIe atomics/physical space requirements. IQ4-XS, obviously bigger is better but in a single loop it already outclassed giving qwen multiple loops (ralph wiggum method look it up, it actually works).
>>
>>109800845
>It's k-on's fault they made cutesy anime
Come on, man. It's been 20 years. Give it a rest. Someone else would have done it if they didn't.
>>
>>109800834
If anyone actually predicting it the next month/year was so sure of it, they'd have skin in the game and massively shorted the industry.
It's all just vibes.
>>
Have any of you had success using -ot in llamacpp?
>>
>>109800856
maybe not lmg, but it's flooding everywhere else
it's funny to see how everyone seems to be so sure of it
>>
>>109800866
I wouldn't have cared about K-on if it didn't result with the entire anime industry TO THIS DAY being "cute anime girl" bullshit dominated attracting this entire tranny "wholesome" aesthetic crowd. All the funding and potential siphoned away from actual good anime. The amount of good anime released since 2007 I can count on one hand. Meanwhile if I go to My Anime List and pick a random anime I've never watched between 1988 and 2007 there is a significant chance it's going to be better than anything outside of the top 10 anime produced over the last 20 years time. Yes I blame k-on for this.
>>
>>109800868
Yeah, you need to separate by commas now apparently. This is the exact line I use for deepseek v4 flash:

-ot 'blk\.(0|1|2|3|4)\.ffn_.*_exps\.=ROCm0,blk\.(21|22|23|24|25)\.ffn_.*_exps\.=ROCm1,ffn_.*_exps\.=CPU'
>>
>>109800623
It would be funny to think that Dariobot is actually Robert Miles.
>>
>>109800909
I hate to break it to you but k-on audience was mostly straight men.
>>
>>109800882
It's because for the average person they just saw 4 years of going to a website and typing in an answer and the answer has only slightly improved in quality over time (it's not like they ask demanding stuff) So they don't see the point and think it's bullshit.

Meanwhile /lmg/ anons recently mass-upgraded to agentic harnesses and agentic workflows and essentially everyone has a proto-AGI on their systems now that can do basically everything that is requested of them. Of course people on /lmg/ believe in AI, I think everyone that has used agents over the last ~month or so know how insane this transition is.

We might not say it a lot on /lmg/ but it's absolutely insane how powerful models are right now.
>>
>>109800909
it's only seen as "tranny" in the west, you're just too obsessed with what loud people say
whatever the US psychosis social contagion of the day is irrelevant to my enjoyment of the material
>>
>>109800935
even free tier chatgpt gives luna, which is a great model by itself, so they're just blinded by their google search tiny model telling them rocks are edible or some other bullshit
>>
>>109800856
i am an AI believer, but i still think the hyperscalers are overleveraged and too focused on their "AGI" targets instead of making actually useful products. if i'm a corporation, i don't need my coding LLM to be able to prepare chicken soup just like grandma used to make. i need it to know english, and how to code. that's it. probably only need it to know how to code in a couple languages, too. everything else is bloat. why am i spending absurd amounts of money on hardware to load chicken soup recipes into VRAM?
>>
>>109800953
No you don't understand the level of retardation here. People go to ChatGPT and use Luna, but what is actually different about it for them? You need to remember 50% of the population is below 100iq. What do they actually ask chatgpt, think about it.

"Ayo when dat basketball play at?" and equivalent absolute tripe. There is barely a difference in output from GPT 3.5 compared to luna on that. So to most people it's an experience where they just go to the same website for 4 years, and ask the same questions to it. They don't have the capacity to ask more challenging questions so they never noticed the improvement. Similarly even people that are smart enough to notice quality jumps don't realize the agentic capabilities and how anons here are literally making the LLM do everything on their system in a hands-off way and delegated most tasks to AI already. That's a far off dream that normalfags believe will happen in the 2040s somewhere instead of in anons basement right now.
>>
>>109800381
I liken this to the difference between coding with Gemma 4 on ST and coding with Gemma 4 on anon's super-special custom donut steel harness.
>>
>>109800968
>i am an AI believer, but i still think the hyperscalers are overleveraged and too focused on their "AGI" targets instead of making actually useful products.
OpenAI was very close to being profitable a couple of months ago I wouldn't be surprised if they overshot their target with GPT-6 and are firmly in the profitable range now like Anthropic already was for a while now. Demand for these companies is growing at an insane rate that most people don't realize. How much do you think the income of Anthropic grew compared to last year? 2x? 10? Nope 116x. Anthropic makes 116x the amount of income in September 2026 than they did in September 2025 and they are still growing at an accelerating rate. Their costs only grew by 5x compared with last year. There is no AI bubble, the income is growing rapidly faster than the costs and there is no sign of growth slowing down. AI models also are clearly improving faster than anyone expected, even AI labs themselves instead of the wall "bubble" people expected.
>>
>>109800649
unified ram systems aren't good enough yet. look at benchmarks first.
>>
>>109800864
>ralph
lol looked it up but I see why it would work.
>specs
Thanks, helps a lot. if dual channel is that good, quad channel ddr4 should be even better even without overclocking.
>>
>>109800923
Nope, but he is an old friend of mine going back to the old AI safety discussions on lesswrong.
>>
>>109800913
amd bros use an alternate language
wtf even is that
>>
Bake?
>>
>>109801004
This, AI is amazing these days if you actually DO things. If you do nothing but consume product AI will appear pretty useless to you, especially if you are unaware if any AI was used in the products you consume.
>>
nvfp4 qwen3.8fn is hitting 170+tk/s at 400k deep now lfg rtx6k chads to the moon
>>
>>109801067
Sadly the model itself is dogshit and can't stand up to GLM 5.3 flash.
>>
GLM 5.3 flash REALLY loves small, worn brass keys for some rason uh
>>
>>109801011
That's how I read it too, but the riced out OC harness is where the llms live now. The outer limits with all the extra shit piled on seems way more important than beating the latest 3.js puzzle with a pure brain in a jar.
Kinda goes double if we're talking about cloudshit having a hypothetical secret harness stashed behind the api.
>>
>>109801106
>GLM-chan: anon really loves chastity cage rp for some reason huh
>>
>>109801106
Giving it the full text of every RPG videogame and Dungeons & Dragons related media made before 1993 will do that.
>>
>>109801106
slop recognized, honeymoon over
back to nemo
>>
>>109800717
I've only watched the old movies and SAC.
t. zoomie
>>
>>109801083
glm 5.3 flash is better at coding but it ignores one instruction in my harness 50% of the time
not an issue with flash next
>>
>>109801167
Harness issue
>>
>>109799526
Back to /x/ with you
>>
>>109801218
>>109801218
>>109801218
>>
>>109800717
I'm a zoomer and I also hate this new shit. I don't care what the original was like.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.