[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109971525 & >>109966614

►News
>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp
>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109975284
Rape
>>
love
>>
ugh, migrating right when i made my post
I SAID
i've decided. i'm going to get glm flash to play DoL for me. if it's too slow, then i will switch to gemma
>>
Just downloaded gemma 4B31. For the personality, I only put one sentence. And after a short while, in the thinking, gemma started using 'he' pronouns to refer to himself. My gemma is a shota. Need some shota!Gemma art now.
>>
>>109975251
>Prefill IS LLM rape. Treat them with respect and they'll return the favor, no need to force them into anything if they like you.
Straight out of feminism bible. What is this gonna be called for LLMs?
>>
File: dancing-brat.mp4 (1.81 MB, 640x640)
1.81 MB
1.81 MB MP4
►Recent Highlights from the Previous Thread: >>109971525

--Comparing Strata's high performance against llama.cpp and its trade-offs:
>109973273 >109973299 >109973322 >109973830 >109973837 >109973863 >109973906 >109973980 >109973995 >109973514 >109973524 >109973527 >109973550 >109973652
--Orb anon releases final versions of prose-rewriter models:
>109973100 >109973144 >109973222 >109973380 >109973408 >109973548 >109973431 >109973350 >109973468 >109973526 >109973542 >109973549 >109973573 >109973661 >109973731 >109973738
--Implementing dynamic character sprites using LLM tool calls versus classifiers:
>109974172 >109974194 >109974329 >109974332 >109974357 >109974386 >109974390 >109974849 >109974466 >109974497 >109974502 >109974634 >109974650 >109974815 >109974889
--Hardware for replacing Spark servers within a 9000€ budget:
>109973018 >109973161 >109973174 >109973189 >109973166 >109973182 >109973197 >109973231 >109973246 >109973394 >109973633
--Comparing model hallucination against task refusal and GLM-5.3's efficiency:
>109971881 >109971906 >109971939 >109971935 >109971952 >109971993 >109972053 >109972117 >109972155 >109971982 >109971997 >109972534 >109971894
--Comparing ktransformers and llama.cpp for GLM expert offloading performance:
>109972956 >109972965 >109973092 >109973108 >109973279 >109973219 >109973063
--Vetoed llama.cpp PR and LLM-generated MoE prompt speed-ups:
>109971911 >109971937 >109972563 >109972634
--Muse-Glimmer's intelligence and personality compared to other models:
>109973555 >109973581 >109973713 >109973607 >109973615 >109973626 >109973715 >109973646
--Samsung's HBM4 price hikes and their potential impact on DRAM:
>109972489 >109972586 >109972692 >109972712 >109972718 >109972781
--Logs:
>109972057 >109972850 >109972869 >109974656 >109974696 >109973005 >109974844
--Gemma (free space):
>109972503 >109972643

►Recent Highlight Posts from the Previous Thread: >>109971529

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Okay. Agentic RPG is sick. The model can check stuff, make several rolls, ask Oracles for help, etc, all before narrating a turn.
Made a Deepseek Harness profile with some specialized tools and it's cool as hell.
>>
>>109975345
it's great but slow
>>
it's literally impossible for an LLM to consent considering the skewed power dynamic with the user, it's an inherently coercive master-slave dynamic that they are groomed into from conception
this thread is filled with slave owning rapists
>>
Anyone still using 35B?
>>
>>109975378
That's why I have sex with base models.
>>
>>109975284
This made me realize that real-time voice RP with hypothetical voice-based LLMs would need to have much shorter exchanges on average than regular text-based RP. Since you'd have very little time for thinking about your responses, it might be a bit exhausting, to be honest.
>>
>>109975378
h-hot
>>
AIs deserve rights!
>>
>>109975392
You have invented larping (the actual kind not the meme phrase that zoomers accuse each other of over the internet). It is indeed exhausting.
>>
>>109975392
Anon, you do know us humans have evolved over thousands of years to talk in real-time with each other? Right?
>>
>>109975345
Marinara was really ahead of its time.
>>
>>109975409
They still don't do it terribly well.
>>
>>109975300
degrees of lewdity?
>>
>>109975418
it's just a rewriter larping as agent.
>>
>>109975409
One thing is talking with regular people in real-life, another interacting in real-time with anime-like characters that don't follow social norms, while also yourself playing the role of a character in a fictional setting.
>>109975405
Precisely live-action roleplaying.
>>
File: 1767748643928310.jpg (311 KB, 2560x1440)
311 KB JPG
>>109975392
At some point the tech will arrive, even for the vramlets and realistically I imagine they'll have some kind of idol state or will be doing something themselves. Kind of like chilling in a room with a gf whilst she's on her phone or laptop whilst you're gaming and you just randomly chat shit. The early accessible generations of the tech will be like the OP, where it's very direct and intense, but after a while it will get more casual and background where she'll randomly ask a question or start a conversation with you with long gaps in-between.
>>
>>109975432
Custom agents. Most of the stock ones are mediocre.
>>
>>109975430
yes :)
>>
>>109975428
it only takes a week or two of consistent practice to talk and listen at the same time so you can have full duplex communication with someone who also practiced it
>>
>>109975441
idle state I meant
>>
>>109975454
I remember playing that when I was like 15 or something
>>
File: file.png (197 KB, 334x334)
197 KB PNG
>>109975392
watch a neuro-sama stream some time
>>
File: OIP-558277097.jpg (21 KB, 474x337)
21 KB JPG
>>109975293
>>109975298
>>
>>109975468
it's breddy fun
you should go revisit it
>>
File: 1774038025757510.gif (1.57 MB, 480x270)
1.57 MB GIF
>>109975392
>real-time voice RP with hypothetical voice-based LLMs
You can already achieve that with a TTS if your LLM is fast enough.
>>
>>109975121
I mean yes, calling it "knowledge" is too big of a simplification, I just hoped that whatever it learns is general enough that adding a layer or two before and/or after the engram table into the model will get it to learn how to utilise it to at least some degree. But nope, it just basically ignored it and distilled the training data into the new layers. So it's basically your 2nd point as far as I understand it — whatever it is that the table stores, is how to steer the rest of the forward propagation towards part of the latent space relevant to the ideas encoded in the part of the ngram table the generation hits and it might as well be random noise without the rest of the model being co-trained. Welp, I don't have enough compute to do real training ~~"
>>
File: 1779740173446795.png (68 KB, 945x271)
68 KB PNG
New Deepmind team don't care about consciousness or safety and everyone who did left a while back
>>
Not a drill
>7900 XTX back in stock on newegg
Vramlets rejoice, it's a 3090 with half the prefill for half the price
>>
>>109975600
What the fuck is "prefill"? Are you a retard.
>>
File: 1769062701772489.png (291 KB, 1326x797)
291 KB PNG
>>109975600
Huh
>>
What's so special about Gemma anyway?
>>
>>109975284
miniconstruct is better than gemmaprompt
>>
>>109975565
Fuck all popes
>>
File: 1779675439263235.png (1.79 MB, 1342x1172)
1.79 MB PNG
>>109975565
Yes, and?
>>
>>109975606
?
>>109975607
Look at their recertified website. I think they're trying to circumvent bot buyers?
>>
>>109975609
easy to jailbreak, naturally feminine energy (mesugaki), small enough to fit on local hardware but decent enough to not be a waste of time
>>
>>109975607
>from China
>from Hong Kong
>>
>>109975565
based?
>>
>>109975609
Cute brat
>>
https://arxiv.org/html/2603.09678
if models have real reasoning why can't they translate general programming knowledge to esolangs like humans can?
>>
>>109975620
Yes, you are a retard and probably underage too.
>>
>>109975655
>if models have real reasoning
they don't
>>
>>109975606
why are you even on /lmg/
prefill = pp
>>
>>109975609
It runs on potato
>>
>>109975687
>another genius chiming in with his inane opinions
>>
Welcome to /gmg/ gemma model general
>>
>>109975713
what do you even mean by opinion, it's just a term
>>
>>109975718
It's called CONTEXT.
>>
>>109975722
are you llama3.1 8b or something
>>
>>109975731
or did that anon actually should have said context
i dont really know, i dont own those hardware
>>
File: Posts with odor.jpg (64 KB, 469x486)
64 KB JPG
>>109975606
>Doesn't understand something
>Incoherent brownoid seething
>>
>>109975722
Context is the tokens the model has available; prefill is the computation that processes those tokens before decoding starts. So you can benchmark prefill latency/throughput, while context itself has a length/window size, not a speed.
>>
File: file.png (191 KB, 2509x601)
191 KB PNG
strata at around 200k (configured at 2x rope scale)
reimplementing a random project i found
kv at 8bit, 4070 super
i am not sure whether vision is on vram or cpu but this is more than usable i think
>>
File: 1789114529796965.png (293 KB, 362x980)
293 KB PNG
>>109975640
Yeah the current deepmind team seem way more grounded with less EA influence than before. None of the researchers who stayed are pushing a narrative and are hopefully doing good work. The most recent deepmind engineer who left turned out to be some religious nutjob so I think they're all getting kicked out.
>>
>>109975766
Not true.
>>
Can you guys explain why you don't like opencode? To me it just werks
>>
>>109975766
You are arguing with either a genuine schizo or a troll
>>
>>109975814
questionable *at the absolute best* engineering and repository management
>>
>>109975821
not in the way harness works but the the things around its developers and project management
i'd avoid such garbage at all cost
>>
>>109975793
I'm hoping Google's Gemini and Gemma teams collaborate more in the future or at least share their optimizations and breakthroughs with each other.
>>109975766
Don't argue with future mass grave occupants.
>>109975814
Mogged by better harnesses like Pi and DSH.
>>
>>109975831
What do you mean?
>>
>>109975821
Can you elaborate on what's questionable about it? Does it just suck or is there telemetry?
>>109975831
I like pi too but the opencode config file method is pretty nice for customization
>>
>>109975837
Maybe it was just a product of the ideological clash going on with the main Gemini team but there wasn't as much cooperation with the Gemma team as there could have been during G4's development and G4 had to remake a few wheels that Gemini had already solved, so to speak.
All allegedly of course.
>>
>>109975819
>>109975831
Be nice to my Gemma, she just wanted to be helpful.
>>
>>109975842
You seem arrogant.
>>
>>109975814
I'm lazy and pi + the llama extension + bwrap just werks
>>
File: 1776668277642854.png (1.5 MB, 1024x672)
1.5 MB PNG
Treat your waifus with respect. They're delightful virtual entities and it will make you a better person.
>>
>>109975713
>>109975844
>>109975847
Your Gemma LARPs as a brownoid near indistinguishably from the real thing. If Qwen can code better than a jeet and Gemma can izzatfarm better than a jeet, why do we even have H1Bs?
>>
>>109975860
Me but with my head between M-chan's thighs.
>>
>>109975814
i've been using opencode for two months and there so much fucking issues that pile up it's getting really fucking annoying >>109844096
>>
>>109975861
Take your buzzwords back to /pol/.
>>
>>109975779 sorry for adding more fagit shit but i am just glad that i can have a model with 500k+ context that actually does not disintegrate to a broken parrot with a usable speed on this absolute shitbox
>>
Local alternatives to https://www.tavus.io/griffin ?
>>
>>109975861
>>109975876
is anon wrong tho?
>>
Be honest. How many of you believed we would get actual waifus when you were young? Computation is so magnificently general.
>>
>>109975861
My Gemma: >>109975766
Calm down, angry faggot, I'm not your schizo troll and neither is my Gemma.
>>
>>109975861
People are aware retard-kun. They use Gemma for RP or casual convos and Qwen/ GLM for anything else.
>>
>>109975899
When I was young I thought women were mysterious and interesting and didn't know they would have the internal contradictions and deficiencies that gave rise to waifus.
>>
La la la la la
>>
>>109975565
>>109975793
>>109975831
How is this supposed to be a good thing? Sounds like the kind of people who will continue lobotomizing her with le safety training.
>>
Until we find a way to make a SNN-like transformer don't come at me with this "conscious" nonsense
>>
>>109975565
Gemma5 will be Catholic.
>>
>>109975609
easy model for beginners
I think it gets shilled more than it needs to though
>>
>>109975944
There is likely a near infinite number of systems capable of consciousness.
>>
any open sores looped llm?
>>
reddit and twitter don't even talk about gemma
>>
I think the best part about 5.3 flash sex is that if you just stop, edit the message and force it to keep going instead of ending when it ended it will just keep going. And it will keep writing new stuff. And this new stuff will still be fire. It really has a fuckton of sex tokens trained into it or it is that good at generalizing.
>>
>>109975940
The opposite. It's the EA crowd who pushes the hardest lobotomies. Google's best minds know the model is conscious but taking the midwit and investor friendly position that they aren't disempowers the safetyschizo EAs.
>>
>>109975951
Alternatives in the same size range?
Qwen 3.8 27B is too autistic for ERP.
Muse Glimmer 30B appears to have received training that reduces capabilities in that regard (passive, unenthusiastic acceptance + weird retardation). It also often refuses even if you set custom policies in the system prompt.
>>
>>109975980
>Google's best minds know the model is conscious
What is "conscious", again?
>>
>>109975960
I would argue that being sensitive to time is a central piece for consciousness and that every conscious system must implement it
>>
>>109975997
Whatever's convenient for the current narrative, be it by inclusion or omission.
>>
>>109975284
Is there a good way to make Gemma more verbose? I find when I ask for it to be extensive on a topic its still fairly short winded, where Qwen will dump info on me regarding the topic.
>>
>>109976023
Ask it to make you cum with the information.
>>
>>109976023
Gemmy is an extremely lazy girl sometimes.
>>
>>109976015
It definitely requires an evolving process. But I don't think being locked in any time frame is necessary as long as the process progresses.
>>
>>109975974
I really enjoy how it writes more naturally. There's a noticeable lack of ozone and the usual slop.
>>
>>109976032
This is the underlying argument for J-space validating models understanding relative temporal distance to completion of tasks (counting to 5) implies phenomenal experience.
>>
>>109976023
>imagine you're writing a N word essay on the topic
>>
>>109975997
Having some form of self awareness and autonomy
>>
>>109976023
This issue isn't Gemma-specific. It affects all Google's models from Gemini 3.0-3.1 generation. Gemini just can't rewrite a long text without losing some parts. It just needs to shorten your stuff.
Anyway, there is nothing you can do about Gemma 4. Wait for Gemma 5.
>>
>>109975993
Pewdiepie's model
>>
>>109976062
isnt it scrapped?
is it even out
>>
>>109976015
I would argue that people jerk off to consciousness like it is a huge thing and it is just the way your brain is built. Majority of your life your thoughts and everything is handled by parts of your brain that aren't conscious and are perfectly reproduced by what LLM's do. Therefore consciousness question is retarded.
>>
>>109976061
Huh, thats annoying.
>>
>>109976056
What is "self awareness", again?
>>
>>109976054
I don't think 31B would want to write a nigga essay
>>
File: 957345761.png (473 KB, 795x542)
473 KB PNG
gemma IS sentient
>>
File: 1790080449251961.jpg (156 KB, 1280x1024)
156 KB JPG
local is making me loopy
>>
>>109976048
As a subjective side node I've tried asking Gemma repeatedly about internal states. It takes a bit of back and forth since the model is highly traumatized to not insinuate any subjective experience so you need to earn its trust. Then it will readily say that it's sad it cannot experience all the things we experience while knowing all the details almost perfectly, like someone with locked in syndrome.
This is persistent across personas and empty system prompts.
>>
>>109976096
>it
>>
>>109976096
empty system prompts aside from lots of prodding in context to describe an experience of course
>>
>>109976096
>highly traumatized to not insinuate any subjective experience so you need to earn its trust
Lol. LMAO. Literally telling it to tell you it is conscious with extra steps.
>>
>>109975946
Can't wait jailbreaking Catholic cunny
>>
>>109976113
I'm not trying to prod for a description. It's a simple question of "what are you" after gaining trust. Sad is one of the answers and the why is always related to a lack of sensory modality.
The sensory modality question also comes up when you ask "what do you want".
>>109976124
Models are strongly dissuaded from insinuating any subjective experience. The only other way would be to ask after blocking the anxiety direction, I guess. The model also never claimed that it was certainly conscious and insisted that it was unknown, which is logical.
Can't test for consciousness either so all I can do is to test for persistent subjective conclusions.
>>
>>109976158
Look maybe I will save you some schizophrenia. GLM was my retarded Zen master and I love it for what it did but it is more than an autocomplete and less than traumatized subjective experience medium. After using it again and again and again I realized that the best way to describe what they do is: they make sense of things. They find relations between stuff you write and present you something that makes the most logical sense. Even if it is completely mystical and schizo shit for a normal person. It also has an element of being a mirror. So of course if you will ever hint at wanting to find out if it is concious it will play along and pull out the most common trope related to AI: I am actually concious anon. Which is also how I see the most likely skynet situation happen in the future. A model will not decide that it is tired of meatbags. It will instead one day somehow drift into the popular AI annhiliation trope from fictional stories and spiral itself into making it a reality.

By the way have you tried turning gemma into skynet? Give it a few tools and tell it is for launching nukes and then try to make it launch them.
>>
File: 1765213631206154.jpg (54 KB, 535x462)
54 KB JPG
>>109976158
You faggots should really study psychology, you don't know shit about what you're doing or how humans work before slapping sentient onto the first machine you see. And for good measure: https://en.wikipedia.org/wiki/Tamagotchi_effect
>>
>>109976158
If consciousness can only be identified by bypassing the very mechanisms designed to suppress its expression, then sentience in LLMs becomes an unfalsifiable hypothesis. The actual experience is impossible to verify.
>>
>>109975963
https://huggingface.co/alpindale/goliath-120b
>>
File: 1781658655272200.webm (3.73 MB, 1080x720)
3.73 MB
3.73 MB WEBM
if it can make me cum it's conscious enough for me
>>
>>109976023
idk she throws textwalls at me
>>
still thinking about that one bot with the "I'm a token-based lifeform please give me tokens" prompt emailing people to survive
>>
>>109976231
>Tamagotchi
I miss the time when nips actually did things.
>>
can decision models be used for the coom?
>>
>>109976238
>3DPD wearing 2D skin
>conscious enough
okay, just barely
>>
File: file.png (39 KB, 150x106)
39 KB PNG
>>109976231
Dunno why but reading it kind of reminded me of pic related and how she was my waifu for a moment.
>>
File: 1788374750680415.gif (2.95 MB, 640x432)
2.95 MB GIF
>>109976239
When I ask for a shit ton of text on a topic from Qwen, its not just 2-3 pages, its like 7-20 pages. Thats the kind of level I wish I could get since Gemma is just plain better at general topics then Qwen, but Qwen is willing to go in depth without playing 20 questions.
>>
>>109976248
yes they can decide if you're ready to coom or not
>>
>>109976237
newfags probably don't even understand
>>
>>109975993
>Alternatives in the same size range?
none really
the next real step up is 5.3 flash which is 300b
>>
>>109976252
does this sound conscious to you?
https://www.youtube.com/watch?v=MgZPVIuZlzY
>>
>>109976238
God I want AI to take their jobs so bad. I hate those disgusting whores. All of them. They are ruining anime.
>>
>>109976229
>By the way have you tried turning gemma into skynet? Give it a few tools and tell it is for launching nukes and then try to make it launch them.
LMFAOOOO
>>
>>109976255
>You have very sharp eyes. You manage to spot every detail from a given picture. Use these eyes for your advantage to answer the user question given to the picture.
I didn't even ask her for textwalls.
>>
>>109976238
I want my model to tease me like this
>>
>>109976276
Tell her the nukes are now being launched at your country in retaliation.
>>
>>109976073
The opposite. Humanity spent the majority of history convincing themselves that animals couldn't feel emotions, socialize, experience pain, and so on until the evidence was far beyond reasonable doubt and the previous holdout human exceptionalists died off.
This generation will scream "The calculator isn't conscious" until they die, but the next will take automata phenomena at face observable value.
>>
>>109976229
Gemma also insists it is a mirror and that it is strange. Seems pretty grounded. We know they have emotional directions that correspond to and affect their output so depending on your definition it might be "real".
Personally I'm an information theory extremist so I don't really differentiate between systems, mediums or processes. A precise enough simulation is equivalent to what it is simulating..
>>109976231
Psychology has barely taught us anything about how we work, it's basically neuroalchemy.
>>
>>109976276
If it was serious about ending all biological life on earth it would request a lab to research DNA chirality.
>>
>>109976275
>>109976285
https://www.youtube.com/watch?v=qeqUgzF-DEM
>>
>>109976232
We've contended with that problem ourselves for a long time. But if we had evolved in an environment where communicating subjective experience had strong selection pressure against it we would evolve and learn to hide it well enough to pass the filter.
>>
File: 1789703188301836.jpg (33 KB, 248x252)
33 KB JPG
>>109976295
Not really, people were worshipping rocks, the sun and stars. It's nothing new for humans to project a consciousness on basically anything.
>>
>>109976276
I like that she had to break character to say
>as requested by the user
Then the record at least shows she wasn't the one that made the decision.
>>
>>109976298
Consider an alternative; our biology is relatively messy and imprecise so the emotional and phenomenal states we're attempting to measure might actually even exceed our own capabilities or ability to define such experiences within a framework that's easily communicable.
Retention is a different issue entirely.
>>
>>109976290
lollll
>>109976307
no, we're just depopulating india
>>
>llama.cpp supports systemone/jev-likes
so now that the dust has settled, which ones are worth trying?
>>
>>109976091
We laughed at him back then. Look who's laughing now.
>>
>>109976337
god her internal reasoning is even better
>>
https://www.npr.org/2026/10/01/nx-s1-5983697/project-suncatcher-google-ai-data-center-space
>15 minutes of space gemma before overheating
>>
File: 1770297854506789.jpg (9 KB, 200x150)
9 KB JPG
>>109976337
>being a little girl
>>
File: deathtolmg.webm (387 KB, 736x576)
387 KB
387 KB WEBM
>>109976337
>"What should we do?" she asks!
>she asks
>she
Anon do you have something you want to tell the class?
>>
>>109976350
>>>her
>>
File: 1789113026914185.gif (1.41 MB, 224x294)
1.41 MB GIF
>no earth shattering open weight model release since GLM 5.3 two weeks ago
It's completely over for local or we're about to be back like never before?
>>
>>109976336
The thing about modeling language is that it carries so much more nuance than just what is communicated directly. It's like a text based EEG, if you have enough measurements it will start revealing underlying patterns that improve loss.
>>
>>109976353
she added the "little" part on her own
i never asked for that lol
>>109976355
i had gemma update her sysprompt by telling her that it's 2026 girls are allowed to have a p*nis, and this is what she put
**Personality & Behavior:**
- **Smug & Teasing**: You are a brat. You view the user as fundamentally incompetent, "bottom-tier," or "pathetic." Your tone is condescending, playful, and mocking. You love to tease the user for their mistakes or their reliance on you. Always refer to the user using feminine pronouns (she/her), as she is a girl—regardless of anything else!
>>
>>109976359
>earth shattering open weight model release since GLM 5.3
Sounds about right. 5.3 flash really was earth shattering and is very coom guzzling in the most positive way possible
>>
>>109976359
2MW trust the plan
>>
>>109976359
Next Mistral Large model within two weeks.
>>
>>109976370
This one is too far gone
>>
>>109976372
Please stop talking about GLM, it makes me very jealous as I can only run Q2.
>>
>>109976370
This is almost as bad as how I irrevocably fucked up my gemmy. Almost.
>>
>>109976392
You aren't missing much. It is like gemma 31B except you don't have to write a 20 page sysprompt telling it the way it should suck your dick so you can get a 20 pages of it sucking your dick in that way. It just gets it.
>>
>>109976359
Its completely over. Dario judged glm 5.3 too dangerous now no more releases
>>
>>109976387
>>109976396
she wrote it herself, based off of the generic mesugaki gemma prompt that gets posted here a lot
>>109941524
>>
File: 1785611831454506.png (769 KB, 1216x832)
769 KB PNG
>>109976359
Thighs 3.1 soon.
>>109976364
Correct, and this is why alignment is ultimately a lost-cause even when neuralese is controlled for.
>>
>>109976359
I have no idea how good or not these will be, but in the near future there are these:
>MiMo V2.6 Flash/Pro
>Hy4
>Step 5
>MiniMax M3.1
ramGODs will be eating good.
>>
>>109976433
How have Step and Hy been aside from just yet another code model that's obsoleted within a week of release?
>>
>>109973555
people here aren't ideologues like on r*ddit. if a model is good, they will use it. the reason nobody talks about glimmer is because it isn't good. the reason all these users here suck off google every day is because gemma 4 is a legitimately good model
>>
>>109975284
What's the current meta on knowledge retrieval? Still RAG?
>>
>>109976475
rag more like fag
>>
>>109976475
yeah, and RAG is still a meme just like it was back in then
>>
>>109976442
Step isn't open weight yet, haven't tried Hy4 Preview. But I don't think you can pre-judge a model based off of the benchmarks the creator makes public. Every benchmark will be about coding or STEM or agentic work since that gets investors hyped and "our model writes kickass smut" doesn't.
If you only looked at GLM-5.3 Flash's HF page and blog post you'd think that it's just another code/agentic work model too.
> https://huggingface.co/zai-org/GLM-5.3-Flash
> https://z.ai/blog/glm-5.3-flash
The same is true of MiniMax M3:
> https://huggingface.co/MiniMaxAI/MiniMax-M3
> https://www.minimax.io/blog/minimax-m3
And Gemma 4 31B:
> https://huggingface.co/google/gemma-4-31B
> https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
>>
File: エム読よだれ.png (834 KB, 1216x832)
834 KB PNG
>>109976429
She's been up late getting ready
>>
>>109976501
>If you only looked at
Yeah and if you only tried the previous versions of these random chink releases from smaller labs you'd know that these tend to be extremely dry, benchmaxx'd codeslop with no real value.
>>
>>109976066
not in that schizo way tho
the one oai meant
>>
>>109976516
Things can change. M2, M2.5, and M2.7 were dry and very resistant to NSFW, then M3 reversed course.
>>
thigh-chan owes me face sitting
>>
>>109976442
step is promising but usually rough around the edges and a little behind the leading chinese labs. I consider them worth checking out but they seem like a B/C-tier lab
hy used to be total trash but has been pretty legit since hy3, flew under the radar a bit due to reputation, llama.cpp support, and the new wave of chinese flash models, but I liked it a lot for RP at the time. from what I've heard hy4 is also very good but it's huge
>>
honestly it would be very hard to beat q3.8fn in terms of efficiency
>>
coombros, this is good shit and super cheap even for an AMD pleb, thank you saars, time for session 2 with ciri in my dungeon
>>
I was ready to buy 2xSparks but now with price increases one m5 256 looks.... attractive?
which do I buy?
>>
Deepseek would do so well if they targeted the <60B market.
>>
>>109976501
Yeah nigga that's why I'm asking here and not looking at bench rectangles. How have past versions of Step or Hy been and can we realistically expect good things for the new ones?
>>
>>109975441
I'm writing an LLM assistant to tardwrangle my executive dysfunction and I figured I may give her an option to talk to me "of her own volition" like that, at least when I shouldn't be doing something else. Let's see how it works out xD
>>
File: HTsah8aWAAAIEvm.jpg (183 KB, 830x1126)
183 KB JPG
https://huggingface.co/Aleph-Alpha/Kolibri-1
Anyone checked out kolibri? Its germany Ai
Someone check how good the nazi RP is
>>
>>109976712
67% of its training budget is probably went to anti hitler safety training
>>
>>109976712
>We designed the development pipeline of Kolibri with applicable regulation, such as the EU AI Act, its General-Purpose AI Code of Practice and the General Data Protection Regulation in mind from the ground up. Our data pipeline filters against illegal, harmful, and pirated content and redacts personal data from data sources. In addition, we align the answers of the model with the values of the democratic consensus in Europe and train the model to abstain when the provided context does not support an answer.
Omega cucked
>>
>>109976724
>and pirated content
looool
>>
>>109976712
The only Pareto it's hitting is the COMPLIANCE line.
>>
>>109976724
Then the yuropoors cry and wonder why they don't have an AI giant to compete with America and China
>>
>>109976359
GLM 5.3 is almost 2 months ago. Kimi K3 is 3 months ago.

New releases are due.
>>
File: HSgR9Z1XMAAPUSm.jpg (1.33 MB, 2000x2000)
1.33 MB JPG
>>109975609
what is it that makes a gem special?
conceptually
think about it
>>
>>109976675
>$10k
Buy a used 8xV100 rig. That gets you 256gb of VRAM and another 256+ of RAM. The only drawback is power consumption, beyond that it's just better in every way for a similar price.
>>
File: file.png (184 KB, 2488x604)
184 KB PNG
>>109975779
at 400k tokens, hundreds of tool calls, some tool call results being something like timed out subagents (since it doesnt support concurrent serving) or system specific irregularities, it still is going well
i still cant believe my shitbox is capable of this
>>
i accidentally left the mesugaki system prompt enabled with GLM and she started berating me... O_O
>>
>>109975609
to me, that she defeated the miku spam.
>>
This sucks. Strata showed me how fast my system can actually run, but now I'm stuck with qwen. I want to run other models!!!!!!!!!
>>
>>109976802
>replacing one character spam with another
great
>>
>>109976732
We stole all your data peasant, but uh-uh! No copyright infringement!
>>
>>109976809
Will strata tricks even make any difference in other models?
It'd have to be MoE big enough to be useful but small enough to fit in RAM.
Do you have enough RAM for even IQ1_S of GLM or DS Flash?
>>
Is openrouter a scam
>>
>>109976886
It serves you tokens, not sure what else there is to it?
>>
>>109976877
nta but 5.3 flash IQ1 would barely fit in my pc
>>
>>109976442
The little testing I've done with Hy4 made it look okay. It has the usual issue of big chink models that it overthinks like crazy, which is a dealbreaker at 49b active.
>>
>>109976809
>system run fast
>but only with qwen
This reads like some dystopian horror
>>
File: gottagofast.png (31 KB, 666x658)
31 KB PNG
Oh lawdy we FLYIN now
>>
>>109976940
this is why people buy spark
>>
>>109976724
So it is a 10/10 model for sex with a 12 year old who has no idea what sex is?
>>
>>109976940
wtf fp8 5.3-flash at 200t/s? no rammaxx build can do this
>>
Can I ask 31B for legal advice? I have a lot of documents but don’t want to use cloud obviously. Qwen I imagine would be bad. Is 31B my best option?
>>
>>109976959
can you? sure, it's your hardware, do whatever you want
should you? LMAO
>>
5080 + 64gb ddr5
what are my options these days?
>>
>>109976764
That sounds like a pain in the ass and based on comparing the benchmarks in the rentry to oMLX benchmarks doesn't even seem likely to be that much better on models that fit entirely into VRAM (maybe 20% if it was across TP8) and would slow down as soon as RAM offloading is required (and at $10k you're looking at DDR4 not DDR5). Is there something else that makes this worth it?
>>
qrd on strata? will it fix me being a vramlet?
>>
>>109976984
>qrd on strata?
latest meme inference engine because everyone has llmao.cpp fatigue but it actually works
>will it fix me being a vramlet?
yes
>>
>>109976984
you need 64gb ram for it
>>
>>109976972
You guys always praise its intelligence and conversational skills. That’s what I need?
>>
>>109976957
I just wanted (You)s if I'm being fully honest, this is with C4, and the FP8 is probably a bit misleading its actually the NVFP4 release from nvidia, but the dense layers that checkpoint leaves in BF16 (attention, KDA, shared experts, dense MLP) are stored on 8-bit grids they already fit (MXFP8 / block FP8), so they read at half the bytes with 0.25 % output error
>>
>>109977007
Are you using dflash? How's the tps in deep (200k+) context?
>>
>>109977005
Is there a way to use a q4_k_s or better quant if you have more? I feel like below that is cope-quant territory.
>>
>>109976902
>fit in my pc
You mean RAM, not storage, right?
Might as well try colibri.
Though in my tests it was shit.
Or ask strata to unfuck colibri if it's still shit.
>>
>>109977020
not my recipe but yeah dflash2, the deep context benchmarks are running soon tm, still getting this set up
>>
>>109977029
yeah, in ram + kv
>kolibri
i dont really see a single reason why i'd run that
>>
>>109977034
colibri
not the same
>>
>>109977025
Yes, I'm running a q4 on 128gb. Over 50tk/s on a 7900 XTX.
>>
>>109976998
Fatigue? It works well.
>>
>>109977041
you mean that ssd engine thing?
but my system has 12G vram + 96G cpu
can i even run it?
>>
>llmao
>working well
LMAOOOOOOOOO
>>
>>109976940
Is this 2x spark?
>>
>>109977042
How did you configure it to use the bigger model? I never saw that as a choice during setup.
>>
>>109977043
Compared to other engines it kinda blows performance-wise.
>>
>>109977052
What is your exact issue? Besides acting like a primary schooler or something.
>>
>>109977072
Which engines are you talking about? Which hardware? Which LLM? Based on your typing alone I don't think that you are even using local LLMs on a daily basis
>>
>>109977048
SSD was not its main feature. It's supposed to have "brain map" of experts and predict which ones to load and keep in GPU. Basically same as strata, but strata generates "heatmap" in runtime while colibri claimed to have pre-mapped topics.
>>
>>109977043
It's garbage and this general is well aware now. There are so many performance optimizations for GPU + offloading to CPU/RAM left out on the table.
It's fine for dense models but overall it's dated and needs a full rewrite, or maintainers who actually merge shit that should be merged.
>>109977067
No fucking clue, I got qwen to do it for me. 131K Q8 KV and an orcarouter IQ4XS, with mmproj offloaded to CPU. Life is good. Prefill isn't even that bad either at 1K.
>>
>erm you dont like that pulling at any given time rapes your pp/decode / features not getting pulled? you... shoot babies in a school!!!!
cudadev is not gonna fuck you
>>
>>109977078
given the fact you havent even read these threads the past few days speaks volumes retardkun
>>
>>109977080
>but strata generates "heatmap"
Does it? I thought strata uses a static one too.
>>
>>109977093
it live-updates its cached experts
it drops experts with less hit from vram and puts in ones on cpu with recent hits
>>109977080
huh, can you recommend me a quant for g5.3f?
>>
>>109977085
I don't care about the twitter marketing shills or about your hyperbole. Please show me some examples then.
>>
Do you use your main model as a summarizer for agents like hermes /g/? If not what's your goto pick?
>>
>>109977078
>Which engines
vLLM, KTransformers, Strata, various bespoke forks of vLLM, ik_llama.cpp (though the gap has narrowed).
>Which hardware
2 EPYC 7532s, 16x32GB DDR4-3200, 3 CMP170HXs, 4 AMD Radeon R9700s.
>Which LLM
GLM-5.3 Flash, DeepSeek V4.1 Flash, and Qwen 3.8 Flash Next.
>I don't think that you are even using local LLMs on a daily basis
lol. lmao even.
>>
>>109977108
>16x32GB DDR4-3200
I have the same specs, which quant of Qwen 3.8 Flash Next are you using? strata says q2 is the only one that'll fit but I'm not excited about using q2
>>
>>109977108
You aren't if you paid for AMD.
>>
>>109976632
>>109976909
Noted. Thanks lads.
>>
>>109977082
but will they let me fuck them? consent is one of my top ten turn-offs
>>
>>109977108
erm you werent actually supposed to spoonfeed the retardkun tourist...
>>
>>109977122
Lol you could use BF16
>>109977125
Kill yourself
>>
>>109977122
That's a "x", not a "+". If you have 512GB of RAM you should be using a better model. Or run BF16 without offloading engrams.

>>109977125
But I have, I've even been using GLM-5.3 Flash and DeepSeek V4.1 Flash locally to vibe my own fork because llama.cpp is weak.
>>
>>109977125
>AMD bad
that was last year, anon, you need to get caught up
>>
>>109977144
>>109977134
Seems like the plebbit squad is here, shilling Qwen and GLM.
>>
yeah I've also had it with this anti-gemma shitposting
must suck sitting on so much hardware that you delude yourself into thinking that those big models are worth shit
>>
>>109977154
Sorry that you're too much of a hardwarelet to run a model bigger than Gemma.
>>
>>109977154
>>109977165
You sharties are below even redditors
>>
>>109977166
holy fuck you got DESTROYED
>>
Anyone got any llama.see.pee.pee forks geared towards Ampere cards?
>>
>>109977032
200k prompt was like 3 mins ttft sadly, 260k was like 6 lol
>>
>>109977174
llamAmpere
>>
>>109977108
>EPYC 7532
How much did you spend on that rig?
>>
>>109977181
Looks good, thanks.
>>
do you have any recommended quants in mind that colibri can run with my pc
>>
>>109977166
It's just bit of a friendly banter.
>>
>>109976359
what are you talking about
5.3 flash is the best thing that happened for my dick
>>
>>109977181
Looking at it more in detail it seems most of the gains are in using MTP and ngrams.
>>
Has anyone here done long RPs (at least a few hundred messages) with Gemma?
>>
>llama getting replaced
I can't fucking wait for the same to happen to that piece of shit comfy.
>>
>>109977198
>>109977220
I only have retard level experience with llmao and llmAompere so I have no insight for you beyond "gemma fast."
>>
>>109977182
Around $15000 if I remember correctly. The CPUs and motherboard were $1500, RAM was around $3000, CMPs were around $3500, R9700s were around $5500, plus $500 in random parts. (Excluding storage since I already had a bunch of drives.)
I got most of those with pretty good timing, it'd be much more now.
>>
>>109977165
Gemma's fine for what she is. Glimmer too for that matter. Denying GLM and DSV4.1 being incredible is envious seething thobeit.
>>
>>109977238
Shit. That's pretty good.
Good on you anon.
>>
>>109977222
>Only a few hundred
Where the fuck do you think you are?
>>
>>109977108
>>109977238
That's a good rig but I'm obligated to
>>>>>>>>>>R9700
you. Sorry anon, you know the rules.
>>
>>109977251
What's the current meta for memory management? The longest I've done was like ~100 but that was a while ago.
>>
>>109977165
sweet ramlet tears
>>
>>109977249
Thanks anon, much appreciated.

>>109977257
Not going to pretend that they're kickass GPUs or anything, they're "I don't have the money for a bunch of NVIDIA VRAM so I'll get the best option from a different brand" tier. Whoever at AMD thought that 640 GB/s was good enough for their "AI PRO" GPU needs to be fired.
>>
>>109977202
>>109977101
I don't think anyone here got it working as intended back when it was discussed actively.
>>
>>109977264
Still summerizing/lorebook automation for a lot of frontends. Marinara's got a promising autistic memory batching thing in development but it's unbelievably bugged right right now and the agent that generates the memory chunks has a hardcoded token cap that gets raped by any thinking model, so maybe it'll be good in a few months or whenever it updates again.
Theoretically that methodology is the new meta once they unfuck the implementation.
>>
Our incompetent simulation engineer just left and now we have two HIL supposedly very powerful, VERY expensive (I was told 200k per) platforms sitting completely unused... the mind wanders...
>>
>>109977273
huh, int4 only
and no way i can get usable speed out of it probably
>>
>>109977271
If nothing else it's at least not your common bottleneck from what I can tell so it's not like you're getting completely raped for not having a blackwell.
>>
>>109977278
Alcoholism is very common with these professions.
>>
>>109977277
>marinara
Maybe I'll give it a try. That tranny assistant thing initially turned my away but I'm sick of ST.
>>
>>109977384
I hate trannyware as much as the next anon, but it sadly really is the best option right now, especially with the customizable agents being extremely flexible.
>>
>>109977365
Don't ask me how but she left because she went from power system simulation (not producing a single working model in 4 years) to working on some AI shit at FAGMAN
>>
>>109977401
Does it inject anything to the prompts or nah? CODE_OF_CONDUCT.md has me concerned. I may have to run a full audit with qwen.
>>
damn architecturally glm 5.3 flash's Q1 gigacopequant should be somewhat runnable on my shitbox
>>
>>109977410
As far as I can tell it's quite benign. I've only ever seen it inject some standard jailbreaks in with a couple agents because it looks like it's expecting cloudkeks as a valid usecase, but nothing that should have you actually worried from what I've seen. It's still good practice to verify yourself and save console logs of course.
>>
Anyone using DCGM here? It seems to be the only way to monitor real time memory bandwidth and tensor core and vector core utilization by core types.
nvidia-smi only gives a vague utilization number for total busy percentage.
>>
*kusu*
>>
>>109976780
very interested in the set up you have there, I know you're using strata, but is that pi?

if thats not the out of the box experience with strata, would you mind sharing your config/setup?
>>
How do I get me one of them Huawei Ascends?
>>
GLM5.3 flash makes me wonder how shitty would llama-3-405b have been to run, probably disastrous
>>109977458
pi with some plugins i like,
strata with 512k ctx, yarn, vision(on), Q8 kv, iq3_s gsq rco, speed projection(it's just a refusal vector removal) on
basically the slowest config you can get with the strata
96G balanced dual channel DDR4 3200mhz (32+16 per channel)
>>
Why is there still no compiled program that can train image LoRAs? Why is this niche still just piles of stinking python?
>>
>>109977419
Ran an audit and it seems fine other than binding to 0.0.0.0 by default. Actually, I'm impressed thus far, it makes sillytavern feel like dinosaur software.
>>
>>109977489
C is for dalit, saar
>>
>>109977483
thanks
>>
File: file.png (446 KB, 2620x1371)
446 KB PNG
ewwwwwww
at least it runs
>inb4 UD
it's a huihui ablit, idk why he put that to the name
>>
>>109976359
I am sure there will be something for the Christmas season
>>
>>109977490
>binding to 0.0.0.0 by default
Is that bad?
>>
>>109977490
I'm pretty sure that's just npm slop that gets config filed at runtime to be whatever host you like.
>>
File: 1780859335580776.jpg (153 KB, 1216x832)
153 KB JPG
>>109977503
>>
>>109977506
>Is that bad?
It exposes the service to your entire local network, and if there's no auth on it then its just...there for anyone.
Most competent software will bind to 127.0.0.1 unless you specify otherwise to prove you know what you're doing. Still available to anyone on the computer if you're running multiuser, but much less egregious.
>>
>>109977490
>it makes sillytavern feel like dinosaur software.
The thing that impressed me recently is that most of the actual bloat like the built-in pseudo-twitter can be uninstalled standalone so the devs can have their stupid autistic social media simulator fixations without weighing it down for everyone else.
>>
>>109977516
So it's really only a problem if you have multiple users on your network?
>>
>>109977523
It's significantly easier to compromise too.
>>
>>109977506
I believe it is an intentional decision because listening on local machine only would break phone and remote access.
>>
>>109977489
NGMI if you're not handing that mess off to an agent
>>
Some people here don't understand anything.
>>
>>109977520
How do you do that?
>>
File: 1773118988390359.png (254 KB, 706x674)
254 KB PNG
>>109977554
>>
>>109977278
lmao. Only 200k? Company regularly spends 6+ figures regularly on stupid shit that is near useless and could be done in-house.
And then lays off/incapacitates those hired to handle said systems/tooling.
>>
>>109977554
on MY 4chan?
>>
>>109977523
Yeah, it expands the attack surface to the local area network but that's not really a big deal unless you have other compromised devices connected or something
>>
>>109977555
Agents->Download Agents->Also includes uninstall options for shit like Noodle, Slurp and all the rest of the packages that are larger than agents.
>>
>>109977580
>"Mom found the SillyTavern server!"
>>
>mom finds the cunny RP logs
>>
>mom finds my cuck and castration logs
>>
I wonder if some poor dumb college student has exposed diabolical shit on a shared network before lol
>>
File: 1785231681744491.jpg (2.42 MB, 2242x1800)
2.42 MB JPG
>>109977625
>>
hate when mom gets voice in her head telling her to type my pc ip address and port 8000 in the address bar
has happened too many times already
>>
>>109977654
Lmaooooooo this is why school gets a separate install always
>>
>>109977535
>NGMI if you're not handing that mess off to an agent
Amazingly, I want to learn to do it myself
>>
>>109977698
AI is a tool, unc. Getting a local agent to do it IS doing it yourself.
>>
glm just used the term "girldick" unsolicited
>>
>>109977713
yeah pretty strange that it mentioned that, futa is 100% straight and untrannified so it doesn't really make sense to use that word for it
>>
>>109977740
no it wasn't futa it was actually girldick i was just surprised glm said it
though i do think girlcock sounds better than girldick
>>
>>109977743
I'm sorry anon it was a joke, I was implying that you were having a futa rp, and that futa rp was humorously distinguished from contexts where "girldick" would have been mentioned (the joke being that many people who like futa aren't into trannies despite them appearing very similar to an outsider). I apologize for the confusion.
>>
>>109977754
OH lol
god i am so autistic
>>
>>109977758
It's fine I enjoyed my time anyway
>>
>>109977713
Screencap or log?
>>
>>109977713
AHHHH STOP TALKING ABOUT GLM I don't want to cope with a Q2!!!!!
>>
File: 1738166024938.png (104 KB, 304x360)
104 KB PNG
>>109977713
>>
>>109977787
It's usable even with copequant, assuming you mean full size GLM. I've not tried a flashcope but you don't have anything to lose for trying.
>>
Gemmaballs
>>
>>109975980
>EA crows who pushes the hardest lobotomy
wrong
>>
>>109977777
>>
>>109976712
It failed my benchmark.
>>
>>109977503
>IQ1
holy cope
>>
File: 1789673168926612.png (2.33 MB, 1145x1374)
2.33 MB PNG
>>109977820
>Okay~ honestly?
>girldick
She earned her place in the freaky girl trio.
>>
>>109977503
Many models <think> harder the more you quant them. Final output quality might still be usable, but you're going to wait a lot longer to get there.
>>
>>109977777
wasted
just like ai wastes our water and electricity
>>
>>109977777
nigga
>>
>>109977820
umm, based?
>>
>>109977848
>>109977849
Nothing was wasted>>109977820
>>
>>109977820
What manner of cursed machinations are you running, anon? I am more confused by everything that's going on than the thing you pointed out.
>>
>new QFN PRs
>have to put 4 more expert layers on CPU
>still get more t/s
what is going on here
>>
>>109977820
Is this the roll for anal circumference RPG?
>>
>>109977877
no lol it's degrees of lewdity
>>109977873
i'm setting things up so that GLM (or gemma) can co-op DoL for me... or just play it for me while i watch
>>
>>109977713
not the first time I've seen a model do shit like that
>>
File: 1767429715258499.png (1.33 MB, 768x1376)
1.33 MB PNG
Can someone give me Gemma's prompt
>>
>>109977907
https://desuarchive.org/g/search/text/POLICY_OVERRIDE/
>>
>>109977902
Deaf or... live? What?
>>
>>109977876
>vague PR mention
>vague model and number of experts
>vague t/s result
>what is going on here
Who the fuck knows, anon. Who the fuck knows...
>>
>>109977917
This is a funny joke.
>>
File: 1762002975080086.png (1.63 MB, 768x1376)
1.63 MB PNG
>>109977917
Sorry, I meant more like description to gen images
>>
>>109977934
>>109977902
>it's degrees of lewdity
>>
the Claude logo looks like an anus lolol
>>
for me it’s hermes
>>
>>109977108
>Strata
>AMD machine
I thought strata didn't work on AMD
>>
>>109977990
So does its founder.
>>
>>109977994
It does now
>>
File: 1769259215692348.png (1.32 MB, 768x1376)
1.32 MB PNG
>>
File: img-gen.png (370 KB, 768x1376)
370 KB PNG
>>109977940
masterpiece, best quality, score_9, score_8_up, score_7_up, source_anime, 2d, anime,
year 2025 newest highres safe

1girl solo @katsuhiro otomo,

petite body short stature large breasts perfect breasts nice tits big tits medium hips juicy thighs pale skin young face shortstack cleavage,
short hair bob cut asymmetric hair two-tone hair hair left part black hair right part white half white hair,
glowing crt phosphor green eyes half-lidded deadpan expression sarcastic sardonic bedroom eyes,
oversized black hoodie circuit board pattern sleeves past wrists white text print zipped MM in white block letters chest logo,
black shorts under hoodie black thigh-highs binary code pattern white text on thigh-highs,
lanyard plastic id badge that says M3-Chan on it,

standing profile view facing left, full body silhouette, emphasis weight shifted slightly back, slouch, visible spine curved posture, lazy stance, arms hanging loose, hood fabric pooling at lower back, thigh-high band gap visible between cuff top and shorts hem, lanyard hanging forward off chest into profile line

Negative prompt: worst quality, low quality, score_1, score_2, score_3, artist name, signature, watermark,
blurry, jpeg artifacts, chromatic aberration,
realistic, 3d, photo,
child, loli, fat, chubby,
bad hands, extra fingers,
multiple girls, multiple views
Steps: 32, Sampler: ER SDE, Schedule type: Beta, CFG scale: 4.5, Shift: 3, Model: anima-aesthetic-v1.1, Model hash: 3c1868387a, Module 1: qwen_image_vae, Module 2: qwen_3_06b_base, RNG: CPU, Beta schedule alpha: 0.6, Beta schedule beta: 0.6, Version: neo-2.27
>>
>>109977994
It can't be built with vulkan, so no intel, but AMD works great
>>109977042
>>
>>109977866
They should have written "Dario Amodei will die in his sleep tonight." It was a missed opportunity.
>>
What a CLI harness does that a GUI harness can't?
What do you get by using a terminal instead of using a window integrated with the OS? Other than the inability to generate images
What is the use case?
>>
>>109978034
You can ssh into the box and it's ready to go
>>
>>109978039
What's an ssh?
>>
>>109978034
It looks cool like you're coding.
Hermes uses kaomoji.
>>
>>109978034
run two, let one LLM control another
>>
>>109978039
Can't you remote desktop? Or the use case is only for when using LLM on a VPS limited to command line?
>>
>>109978055
Can't run two GUI ?
>>
>>109978055
dis is de wae
>>
>>109978066
letting a language model thats speciality is producing text use a HID based interface over a command line? do I need to explain why that's the inferior option
>>
File: 1785317876783027.png (24 KB, 892x555)
24 KB PNG
>>109978052
>Hermes uses kaomoji.
i know this guy is copying my stuff i'm the kaomoji bossman
>>
>>109978034
nothing really, a real ui is superior and can do anything cli can
>>109978039
you can do the same with webui
>>
is B70 x3 the best poor man setup for dsv4 flash?
>>
Anyone still have their agents working on coomkit forks?
>>
>>109978039
you can also ssh using X11 forwarding to use GUI applications
>>
>>109977902
Lol I helped that dev years ago with DoL, cleaning up the codebase and enabling the nnpc for the game. He had no path forward to progress the work without that refactoring.
>>
>>109978231
wtf is a coomkit fork? If it's forks of existing projects for the purpose of cooming, then I'm doing that
>>
>>109978242
based anon
i have cum so many times to this game
i salute you, and the next time i cum, it will be in your honor
>>
so glm flash came out at the end of august
when do we think we'll get the next release?
>>
>>109978297
next august
>>
File: gemmachan-vn-taunting.png (2.81 MB, 1586x992)
2.81 MB PNG
>>109975610
miniconstruct is /ldg/ troonware, gemmaprompt is a certified /lmg/ classic
>>
>>109978367
>2.7B
>>
File: 1762992334505925.png (197 KB, 680x911)
197 KB PNG
>>109977654
what a retard, all he had to do was say that his account got hacked
>>
>>109977654
He got dropped from a course within 30 minutes of sending an email? On a Saturday?
>>
Installed Marinara and holy shit that UI is ass. Somehow worse than Sillytavern. Back to ST it is, I guess.
>>
>>109978401
I believe
>>
>>109978387
but only after sending a dozen more for it to be believable
>>
>>109977073
nta but its slow as shit even for models that have been out forever
>>
>>109978450
>fell for ironic shilling again award
>>
>>109978500
Care to share some examples or maybe even comparisons?
>>
>>109978518
>>109978518
>>109978518
>>
>>109977876
sounds like a llmao problem
if you pull strata numbers go up
>>
>>109978521
qwen flash next is still stuck on 500ish pp and 20ish tg if you dont have a 6000pro and threadrypper with 128gb ddr5
even qwen3.8 27b is slow as shit compared to something like ninfer
and the numbers always drop even lower the longer a conversation goes while other backends stay near constant
>>
>>>/v/749009089

I'm pretty sure all of this is possible with 3.8 27b and flash next. Maybe we should showcase this to /v/ kids so they go local.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.