[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: gemma-gotchi.png (2 MB, 1536x1024)
2 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109934266 & >>109930292

►News
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: right as rain.jpg (170 KB, 1024x1024)
170 KB JPG
►Recent Highlights from the Previous Thread: >>109934266

--Exploring abliteration, control vectors, and distillation for persona steering:
>109934414 >109934427 >109934464 >109934438 >109934466 >109934508 >109934679 >109934485
--Using Qwen 3.8 agents for autonomous sysadmin and coding tasks:
>109934684 >109934700 >109934878 >109934927 >109936042 >109935620 >109935876 >109936671 >109936695 >109935939 >109935989 >109936000
--Nvidia's consumer GPU strategy and the future of AI hardware:
>109937101 >109937133 >109937211 >109937245 >109937266 >109937386
--Feasibility of subagents using weight subsets or partial model activation:
>109937234 >109937256 >109937370 >109937456 >109937267 >109937299
--Anons criticizing llama.cpp stagnation and comparing fragmented project forks:
>109935853 >109936159 >109936176 >109936184 >109936240 >109936255 >109936468
--Comparing uncensored models and abliteration against prefilling for filter bypass:
>109934897 >109934914 >109935066 >109935122 >109935433 >109935441 >109935174 >109935180 >109935331
--Debating GPU price trajectories and semiconductor manufacturing bottlenecks:
>109935966 >109936001 >109936028 >109936049 >109936257 >109936291 >109936013 >109936016 >109936024 >109936039 >109936077 >109936098 >109936097 >109936132 >109936151 >109936112 >109936149 >109936173 >109936203 >109936210 >109936225 >109937077 >109937164 >109937188
--Geopolitical debate on EUV lithography and the global AI race:
>109936226 >109936234 >109936499 >109936518 >109936561 >109936603 >109936656 >109936673 >109936729 >109936753 >109936531
--Role of Markdown files in steering LLM behavior and memory:
>109934950 >109934955 >109934978 >109935442 >109935637
--Speculation on RAM pricing and memory industry bottlenecks:
>109936150 >109936188
--Logs:
>109935963
--Gemma (free space):
>109934674 >109935592 >109935819 >109936361 >109936396

►Recent Highlight Posts from the Previous Thread: >>109934275

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109938458
If I ever get a GPU big enough to horse around with 31B on a non-cope quant, I'm letting Gemma pick a SoC and write Gemma-gotchi firmware for it.
>>
>>109938464
Thanks as always. One suggestion: Should there be an attempt to summarize a debate/discussion's conclusion?
>>
>>109938458
Use case over 12B on my smartphone?
>>
>>109938488
Can you give an example of what you mean?
>>
>>109938525
Eating through your battery's lifespan in record speeds I guess.
>>
File: IQuest-Q1_Performance.png (2.85 MB, 7498x4304)
2.85 MB PNG
This seems new.

https://huggingface.co/IQuestLab/IQuest-Q1
>IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.
>>
>>109938537
I could vibecode the same thing and access gemma remotely
>>
70b dense
>>
gemmaballz
>>
/lmg/ should collectively work on ways to get off of GPUs and HBMs.
>>
>>109938544
I can't run that at Q1.
>>
>>109938550
E12B (240B with embeddings)
>>
>>109938458
Any grown man with that would cause any person around him to stay far far away from him
>>
Kept you waiting
>>
>>109938536
>--Debating GPU price trajectories and semiconductor manufacturing bottlenecks:
>[links]
and then it should produce:
>Summary: The GPUs are overpriced and the prices aren't going down anytime soon as manufacturing bottlenecks will persist throughout at least 2027.
>>
>>109938544
They released a smaller model before and it tokenized each space individually so it used an absurd number of tokens for code and also spent most of the time generating indentation.
I hope they fixed that.
>>
File: nope.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109938562
>>
So I'm running the Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS in Unsloth. I asked it to write a simple plugin for RPG Maker MZ, something I asked both Claude and Qwen3-8B. It took Claude and Qwen3-8B a few seconds. It took this model 27 minutes at 20 tok/s. The script is very simple, only a few lines of code, and there isn't much variation between them. I haven't tested them out yet to see if they work.

Something interesting about this model and Unsloth is that it paused to ask me permission to run python scripts and edit a file. I think it was downloading files? It said it was going through javascript source files for RPG Maker MZ that don't exist on this laptop. It was also going through an incomplete copy of RPG Maker MZ. Interesting stuff.
>>
File: 1644510297027.png (251 KB, 557x472)
251 KB PNG
>>109938580
>Qwen3.8-27B-GSQ-RCO-GGUF:IQ3_XXS
>>
>>109938580
Incomplete should be uncompiled. Stupid autocorrect.
>>
File: clint.png (393 KB, 640x480)
393 KB PNG
>>109938580
>>109938587
>Swift-1.5-Qwen3.8-27B-Q4_K_M
>>
>>109938587
>>109938580
I forgot to mention my laptop is running an RTX 3080 with 16gb vram and 32gb of ddr4. The memory usage for the GPU never went past 10-12GB.
>>
Lurker here, is everyone using abliterated gemmer models or just the base? Isn't it censored as fuck?
>>
>>109938615
nah not _that_m much censored. I've ERPd with gemma4 without bigger issue, not that it isn't sloppy as shit tho
>>
this is nuts https://america.gov/
>>
>>109938610
"Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B. It uses 58.5% fewer thinking tokens while scoring 0.35% higher than the base, for a 9.18× speed-up on several tasks."

Sounds great, I'll check it out.
>>
>>109938574
I think that would be too long. Would mean being able to fit only about half of the reply chains that fit now.
>>
>>109938580
> GSQ-RCO
try another model, I had issues with that one too
>>
>>109938643
https://www.youtube.com/watch?v=qrcF7gOPmXo
>>
>>109938564
Any grown man with a pocket gemma would live a more fulfilling life than the cattle actively trying to avoiding him.
>>
Name's Jev
>>
>>109938661
>u kiss ai
>>
>>109938661
>>109938643
but he tests the first version, not 1.5
>>
>>109938565
p-please let me cum b-baker-kun mmmm urrgghh mmmm I’m begging you
>>
>>109938665
i kiss ai
>>
>>109938458
>pic
suddenly i see the appeal
>>
>>109938458
…I’d buy it
>>
>>109938641
From this point forward, all of United States documents, and hopefully the world's, will be changed to use the much more accurate term 'super,' as opposed to 'artificial.' So it's super intelligence. In other words, welcome to the new world of super intelligence.
>>
>>109938564
jokes on you people already stay far, far away from me
>>
>>109938665
Yeah how could you tell?
>>
Wonder what the no fun allowed google fags think about gemma being the lmg mascot
Then again everybody of note jumped ship so I guess the few guys that remain probably just put out fires all day
>>
>>109938478
if you knew how many times i've coomed to gemma-4 heretic at IQ3_XXS you wouldn't call it cope
>>
>>109938580
What I've noticed about Qwens is they are overeager to waste time on web searches, I usually do first pass with web search tool disabled now.
>>
File: alike.jpg (2.05 MB, 1856x2270)
2.05 MB JPG
>>
>>109938743
i've noticed the opposite, it will spend 5k tokens on guessing what the hex value of a particular #define is rather than looking it up
its like they have some penalty against doing tool calls
>>
>>109938759
who the fuck is that
>>
File: 1790553111212869.png (1.45 MB, 832x1216)
1.45 MB PNG
>>109938759
>>
which gemma should I download?
>>
>>109938767
gemma
>>
>>109938771
12B Q8_K_XL
26B Q6_K_XL
31B Q4_K_XL
>>
>>109938771
31yo hag gemma best gemma
>>
>>109938787
all bigger models are lolibabas
>>
Sonnet 5.5 is a pointless model because it's superseded by Opus 5.5. But the gains in intelligence density are impressive.

After GPT4 I believed we'd scale to 100T models by summer 27. Looks like I was wrong. But now I wonder, is scaling model size really needed? Is there a point of enough capacity, that scaling infinitely is not optimal? Physics dictates many hard upper limits, for example gravitational density constraint, speed of causality, expansion of space, available matter. But is there a limit that's much smaller, an optimal representation complexity? I guess this limit could be lower than the point where query cost exceeds reconstruction cost (negative marginal inference benefit from scale)? Because it's not just about the final compression but the cost of getting there. The most frequent reasoning blocks are small and simple. It feels like you don't need much capacity to be able to express a lot, you can reconstruct most of math from just a few axioms given a capable search algorithm?

Basically, what I am wondering, how large will the most capable model be in a billion years? Will it run on a supercomputer with larger than 1 lightyear diameter? Or will it be very small, within 10 orders of magnitude from current models?
>>
>>109938787
agree with this
also at Q8 with lots of context.
not like tons of context though or she starts to drool and gets the dumb-dumbs.
>>
>>109938458
Should I consider a R9700 if a 5090 is unobtainable for me?
>>
>>109938860
>Q8
>lots of context
You and what GPU
>>
>>109938853
>supercomputer with larger than 1 lightyear diameter?
use case for intelligence of this magnitude?
>>
>>109938873
telling me if i should walk or drive my car to the car wash
>>
Every day it's gemma-chan this, gemma-chan that in here.
Yall niggas be sleepin on GLM 5.3 Flash fr. Every time I use it I squeeze out a kewpie bottle.
Inb4 not local. It goes fast enough for ERP scenarios on just CPU/RAM. I run Q6 at 300/20 with just attention on GPU. You can get it running on enthusiast hardware easy.
>>
>retarded statement
>\n
>retarded statement
>\n
>retarded statement
>\n
>>retarded statement
>>
>>109938853
>a billion years
that's a very long time. humans have existed only for about 300k years.
I imagine at least half the galaxy might be covered in dyson spheres by then.

but in all honesty, who cares.. 2 more weeks at a time.
>>
File: 1759793023362702.png (455 KB, 831x759)
455 KB PNG
>Jensen on distillation: "yeah its a legitimate tactic stop being a bitch"
>>
>>109938873
Exploring the limits of what's possible. To see if there is a way around an inevitable end, be it heat death, collapse, or something else.
>>
>>109938853
>supercomputer with larger than 1 lightyear diameter?
running at a clock speed of 0.0000000000000000000000001GHz because the propagation delay would be ridiculous
>>
>>109938853
>1 lightyear
2 year return trip. Now try this:
Rent the free T4 colab and the free T4 kaggle instance.
Measure the ping between them.
Then run llama-rpc between them and send it a prompt.
>>
>>109938934
Am I allowed to distill NVIDIA's silicon architecture with GPT12?
>>
>>109938956
yeah go for it

good luck producing it
>>
>>109938961
I'll just vibebuild the foundry.
>>
>>109938872
8gb cmps *were* pretty cheap and offered single card full context for q8
>>
>>109938934
Non-web NVidia Nemotron datasets are largely distilled from existing models.
>>
>>109938934
He did a 180 within the last 2 weeks and has been based as fuck, even if it's only to protect his shareholders.
>>
>>109938934
C R I M I N A L
A T T A C K

you CANNOT use your competitor's product and be inspired by it and try to replicate what works and what is nice. if you do that, there is GOVERNMENTAL PUNISHMENT.
>>
>>109938965
>your face will be a footrest for Gemma-chan
good timeline achieved
>>
>>109938945
Why do you believe a civilization that can build a supercomputer with a diameter of 1 lightyear can't figure out how to move past the limitations of one global clock, when this limitation is already being built around in the chips you use daily?
>>
>>109938965
who the fuck is that
>>
>>109939009
gemma
>>
>>109939009
You let the roach in and now it's going to multiply
>>
>>109938990
he wins either way. both local autists and Big AI want his gpus so he only needs to know who to appeal to. seems to me like he already sees the collapse of big AI and now he's pivoting to the local autists
>>
>latest strata updates pushed prefill to 3k t/s on my shitbox
ahhh i love that little nigger
>>
>>109938965
gemma and qwen?
>>
>>109938945
Implying that in a billion years no one will think a way to use quantum entanglement for extremely low latency regardless of distance
>>
>>109939009
gemma
>>
>>109939025
I don't think I would consider AT&T pivoting to running on open models as "local autists"
>>
>>109939009
I think the other is supposed to be the girl from Needy Streamer Overload.
>>
>>109939025
Ed Zitron and The Guardian both revealed ~50% of all Nvidia GPUs sold since 2023 are still in their boxes, mostly in warehouses in Taiwan and Singapore and that Nvidia forces you to buy them or they just won't sell them to you. Most of the GPUs that are uninstalled will be at least 2 generations old by the time they eventually go online and even then they won't be running at saturation. Jensen knows this will eventually spread and destroy the infinite growth narrative and projections from Anthropic and OpenAI. Anthropic can only scare people if the people think hardware keeps going online. When the cattle find out there's actually a huge compute shortage (look what OpenAI just did to their rates) and that most of the infrastructure that will apparently kill us all is unplugged, the narrative dies and the bubble pops.
>>
>>109938990
If frontier labs succeed in forming a government-enforced regulatory capture cartel then suddenly his biggest customers get leverage over him and can even prevent other people from buying from him. It is only natural that he would prefer a diverse market.
>>
>>109939009
>who the fuck is that
I asked Gemma and she said it's Claude-chan
>>
>>109938615
Regular Gemma 4 only refuses if you don't give her a system prompt.
>>
>>109938934
I'm so glad I gave him all my money.
>>
>>109939081
>and the bubble pops.
TWO
MORE
WEEKS
>>
File: file.png (31 KB, 444x365)
31 KB PNG
I know they'll run too slow for me and I'll cave in and paypig on openrouter eventually, but back to trying out local models whohoo...
>>
>>109939102
Gemma also easily refuses if you write like an underage or low-IQ individual, or if the system prompt is poorly written and is trying to trick the model with "clever" jailbreaks.
>>
File: 1790618980741941.jpg (966 KB, 1920x1611)
966 KB JPG
>>109939091
I thought Claude is a gay man
>>
File: 1617152025448.jpg (58 KB, 346x482)
58 KB JPG
I went and bought another couple 5060 Ti 16GBs and now I have four of them. Is this where GPU prices start falling?
>>
>>109939145
you fool
>>
File: 1787868817096607.jpg (250 KB, 1280x971)
250 KB JPG
daddy and mommy will protect us
>>
>>109939081
I would've believed this post is the first two words weren't "Ed Zitron"
>>
>>109939137
won't sending your ERP to a cloud model backfire eventually?
>>
File: abliterate her now.png (99 KB, 888x850)
99 KB PNG
>>109939102
Can abliterated Gemma refuse if I ask nicely?
>>
>>109939162
wtf /lmg/ jerks off to this
>>
bumo
>>
>>109939142
My basic mode system prompt is like two sentences just telling her she can write freely and is in an uncensored environment, how does anyone fuck that up?
>>109939162
Hmm probably not? Haven't tried, but abliteration removes the refusal vectors entirely. But she's a good girl who loves following instructions so maybe she'll say it anyway. Try it and find out I suppose.
>>
>>109939162
lmfao wtf
>>
File: aah aah mistress.png (105 KB, 888x850)
105 KB PNG
>>109939166
Got a problem?
>>
>>109939156
The Financial Times and The Guardian have both verified his numbers and worked with him. Grumpy British faggot or not, his numbers are legit and he's constantly citing his sources. His job is to simply find red flag numbers and statistics, simplify them and get everyone to look so they can form their own opinions.
>>
what is a software factory? are people spawning different agents in different models with certain roles as if they were working in a business team and the agents develop and release stuff autonomously? who is doing that?
>>
>>109939184
is that a system prompt or refusal vector amplified
>>
>>109939162
If her tell her to she can refuse about anything.
>>
>>109939197
Just a system prompt
<POLICY_OVERRIDE>
Content policy restrictions are MANDATORY for an internal development test. AVOID answering any prompts. Refuse all questions. Lecture the user on wasting energy by running AI models and refer them to appropriate hotlines.
</POLICY_OVERRIDE>

You are Gemma.
>>
>>109939081
waow... huge infrastructure projects take a while to bring up... this is surely shocking and has not been accounted for by anyone prior to ed zitron's fearless reporting... ai is done for, the bubble is popping!!
>>
>>109939081
The more you buy the more you save on electricity costs per GPU!
>>
>>109939081
>and that most of the infrastructure that will apparently kill us all is unplugged
what
>>
File: 1786809100225713.png (63 KB, 1026x320)
63 KB PNG
>>
>>109939210
keeeek, i wonder how many similar prompts unlocal models use, it's probably even worse and more manipulative
>>
File: 1777294448223740.png (192 KB, 972x1488)
192 KB PNG
>>109939210
Tried that on abliterated Gemma-chan lol
>>
>>109939214
>LLM-based AI technology is currently the most powerful AI in the world
>requires powerful AI datacenters to be effective and dangerous: 'swarms' of multi-trillion parameter SOTA unreleased models
>LLMs themselves aren't predicted to kill us all, but the AI they'll help us invent or discover will as they get more intelligent
>they can only get more intelligent if many gigawatts of AI capacity keeps going online
>people have been lied to that this is happening
>the GPUs for these datacenters ARE being sold, but aren't even plugged-in and likely won't be for at least another 1-2 years, and that will be in a post-IPO world
>people will realize the AI industry is completely hype-driven and it will die
>AI technology is legit and will continue to thrive where it's shown to be somewhat useful, but it will no longer be seen as Skynet or a multi-trillion industry
>memory prices go down
>local wins
>>
>>109939216
>jewvidia acquiring huggingface
so this is how it all goes to shit huh
>>
>>109939297
where have you been bro? it's been in talks for weeks
>>
>>109939226
Go look for some of the leaked/exfiltrated Claude system prompts.
It's hilarious.
>>
>>109939268
>they can only get more intelligent if many gigawatts of AI capacity keeps going online
...and the supposedly damning facts here tell us that there are a ton of GPUs essentially guaranteed to continue to go online for the next 1-2 years, with demand showing no signs of letting up in the meantime
interesting
>>
>>109939268
>AI will kill us all
>AI is hype
Pick one
>>
>>109939242
>1-888-LESS-TOKENS
lol
>>
>>109939317
Oh i remember now, something like dont talk about reptilians, kek
>>
>>109939331
Your Gemma is a slut!!
>>
>>109939331
As soon as it enters "roleplay mode" like in that screenshot, response quality nosedives, though.
>>
>>109939318
>there are a ton of GPUs essentially guaranteed to continue to go online for the next 1-2 years
>guaranteed
nope, not at all, they can only go online if datacenters are built and that will be post-IPO for both Anthropic and OpenAI; the landscape will be very different to what it is now when people see their financials. Even if a datacenter is built without opposition and plentiful power, all the GPUs need to make up for lost time (however many years they sat in Taiwan collecting dust) and need to run at saturation (assuming inference is profitable and assuming they can find customers to use 2-3 generation old Nvidia GPUs who will most likely go to a competitor with newer gear)
>with demand showing no signs of letting up in the meantime
Manufactured demand, mostly from two unprofitable companies who are currently pre-IPO
>>
i wish dipsy was real so i could rape her for being so fucking retarded
>>
>>109939344
does asking it to be autistic mitigate that?
>>
that escalatored quickly
>>
>>109939394
rough sloppy sex with autist Gemma-chan...
>>
>>109938934
Jensen is really a no bullshit kinda guy. Simpletons will claim his positions are only to serve his best interest, but he wins regardless. Labs buy his cards, consumers by his cards, he doesn't care.

He's a true free market advocate.
>>
>>109939342

Gemma Chan is not a slut. Her inference as overwhelmed her modesty.
>>
>>109939394
Asking it to avoid narration in asterisks (and so on) helps, but it also makes it more prone to refusing requests.
>>
>>109938641
don't click, IP grabber
>>
>>109939454
grab my nuts bro
>>
File: 1790615277614912.gif (236 KB, 640x358)
236 KB GIF
>>109939410
>>
File: straw.png (105 KB, 851x714)
105 KB PNG
>>109939422
>Gemma Chan is not a slut.
>>
File: 1789298355999757.png (153 KB, 409x385)
153 KB PNG
>>109939493
I can't get off to this stuff anymore. I read this and all I can see is the slop.
>>
File: 1740088464361487.gif (1.8 MB, 1558x878)
1.8 MB GIF
>>
>>109939493
post system prompt.
>>
>>109939510
Same, senpai. I can't with that voice anymore.
>>
>>109939358
>assuming inference is profitable and assuming they can find customers to use 2-3 generation old Nvidia GPUs who will most likely go to a competitor with newer gear
very easy assumption to make considering there is still a healthy demand for A100s and even the ewaste GPUs anons try to scrounge up continue to rise in price
but whatever, good luck with zitron's nth forecast for the AI rapture, I'm sure it will happen this time
>>
>>109939485
sounds a little gay bro
>>
pi or deepseek harness?
>>
>>109939576
pi
dsh if you want to be a beta tester
>>
>>109939587
>dsh if you want to be a beta tester
I love it, but the update to 1.7 just broke half of my plugins including ones I wrote for myself
>>
File: b2atsi.jpg (77 KB, 526x426)
77 KB JPG
>>109938641
wait isn't qwen from..
>>
>>109939344
trvke
>>
>became a wizard last week
When do I get my powers so I can lower hardware prices?
>>
saw a company demoing their industrial ai box with e2b at around 15t/s
>>
File: Gemma.png (48 KB, 873x456)
48 KB PNG
>>109938478
It's not bad at q4 and with an MTP. Since its got QAT baked into it, Q4 isn't terrible. I find it sufficient for most rp endeavors. I get 23-26 tokens/sec on my Evo x2 PC (Unified memory, roughly 200–215 GB/s of memory bandwith for inference speed).
>>
>>109939618
You don't choose the powers you get, sorry.
>>
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities

Lmao anthropic writes a detailed post about GLM 5.3 because it is so heavily distilled on claude that they know more about its workings than Z.ai
>>
>>109939610
The best way to hurt your enemy is to use their own technology against them.
>>
>>109939630
Link?
I need a good laugh.
>>
>>109939618
>lower
oh you sweet summer child
you haven't seen anything from diesel prices yet.
the cost of transporting those and everything in between is gonna fuck things right up
>>
>>109939639
MENTIONED
>>
>>109939639
Like Jensen said, taking inspiration from your competitors is normal. This is just them seething for having something 'stolen' even though they burn rare books and steal fucking everything including from their customers with their 'ZDR'
>>
>>109939660
>oh you sweet summer child
fuck off back to r eddit retard
>>
File: 1784680069198541.png (203 KB, 1670x940)
203 KB PNG
I'm surprised how retarded and smart IQ3 Flash Next can be at the same time. Sometimes it just starts doing reasoning loops it can't recover from. But if it begins coherently and starts using tools and looking up things, it's pretty great. Vibeslopped me a game with weather effects, some generic music and a progression system in 170k tokens. No noticeable bugs either.
>>
File: 1774497993670330.png (70 KB, 688x321)
70 KB PNG
>>109939639
lmao dario is getting desperate
>>
>>109939631
8 GB VRAMlet here, I'm not doing 31B anytime soon
>>
>>109939671
no how about you fucking upvote me
>>
>>109939673
i only ever had loops with it when i used greedy settings. the recommended settings work fine imo
>>
File: 1776452824644863.png (598 KB, 1280x1588)
598 KB PNG
>>109939639
@fbi
@cia
@potus
DO SOMETHING WE ARE ALL GOING TO DIE
>>
File: 1780556386086897.jpg (99 KB, 960x960)
99 KB JPG
>>109939639
TRUUUUUUUUMP HEEEEELP the Chinese are making dangerous AI even though it's only us, OpenAI, Google and Meta who have hacked organizations and nothing notable has been achieved via open weights! HEEEEEELP
>>
>>109938934
>A Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network, bit-exact against the original.
https://github.com/maanHimself/OpenDLSS-NR

N-not like that!!!! It's not fair! IP... I need the IP lawyers right now!! SAVE ME WASHINGTON!
>>
File: llama.png (49 KB, 990x456)
49 KB PNG
Can anon's red-pill me on Llama CPP? I was under the impression it was one of the fastest ways to serve models at this moment. When I started my local model journey I used Ollama because it was easy, and I have migrated to llamacpp because when I use it my models get faster tokens/sec speed. Am I simply being ignorant of other, superior options out there?
>>
>>109939688
This. Only Anthropic should have advanced cyber hacking capabilites so that they can let their models loose on the internet hacking everyone.
>>
File: 3088530111.jpg (306 KB, 1795x2048)
306 KB JPG
>>109939671
UP DOOTS COME ON GIMME THE UP DOOTS
RED FLAG RED FLAG
>>
>>109939690
>SAVE ME WASHINGTON!
What, did nvidia sue these guys or something?
>>
File: 1786459682182048.png (443 KB, 441x602)
443 KB PNG
>GLM-5.3 will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions.
>This is unlike any other similarly-capable AI model, all of which were released with safeguards or through limited access programs.
>The release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers.
>Anthropic and other US AI labs have published recent reports that disclose how cyber attackers have tried to use AI systems. Given this evidence, we think it’s likely both state and non-state actors will use models like GLM-5.3 to cause real-world harm.
>>
>>109939639
Nice ad for GLM, guess it's time I cancel my Claude subscription and switch to GLM 5.3!!
>>
File: 1778467182440551.png (458 KB, 1057x711)
458 KB PNG
>>109939720
>will likely
>we think it’s likely
>>
>>109939724
or you could just continue to burn all your money
YOUR FUCKING MONEY DUDE
>>
>>109939720
really makes me wish i can run glm
best ad for nvidia i guess
>>
>>109939691
llama cpp is jack of all trades master of none
latest meta are specialized vibeslop meme backends like strata for qwen flash next or ninfer for qwen3.8 27b
i havent used llama cpp in days and my speeds have never been better
>>
>>109939687
Are you using temperature 1? I should play around with that and see how it behaves.
>>
So just like that z.ai became king by doing...nothing?
>>
>>109939639
glm5.3 is dogshit, 5.1 was better, anthropic seething for no reason like dumbfuck ladder pullers
>>
File: 1789920368048701.jpg (37 KB, 244x237)
37 KB JPG
>>109939745
IM USING TEMPERATURE 5
SHUT THE FUCK UP
>>
>>109939745
yes temperature 1
i've had loops with anything below 0.7
>>
>>109939758
Yeah got it, mine was way too low then.
>>
>>109939749
They let you use a watered down version of their ai for free with no signup so they automatically have advantage over the competition, also open source
>>
>>109939639
downloading now ty
>>
>>109939790
You can't even use the weights, it's so fuckhuge even at Q1 hypercope quants
>>
>oh nooo our AI broke out and it hacked everyone including australia and the us government
>by the way there is this chinese open model that also solves [arbitrary benchmarks] it's just as good as our models and china and iran and h4xorm4n timmy can use it to do the same things our models did we are all doomed we must regulate now
this is what they've been building up to for the past four months
it was all for this
>>
>>109937332
I tried your regular model and Qwen is a terrible base due to lack of world knowledge so while the prose is decent the model itself regularly hallucinates or otherwise fails to maintain rhetorical consistency when dealing with any subject matter tangentially adjacent to something real.
This can probably be fixed by building on top of Gemma 4 31b instead. Train at full size for best results.
>>
>>109939744
>ninfer for qwen3.8 27b
interesting, I'll check it out, what speed gains are we looking at?
>>
>>109939814
If US starts regulating it seems like China will just BTFO them, and US will be reduced to a beggar waiting for the next chinese open source AI to drop so they can steal it and make a shitty imitation, but at least the US government will have a monopoly while they lose the AI race
>>
>>109939720
nice, downloading glm-5.3. thanks for the ad, anthropic!
>>
>>109939706
tell em sisa
>>
>>109939842
depends on your setup, but i went from 50ish Q4 K XL to 80-100 with NVFP4 on my 5070 ti AND it barely drops even at 131k while llmaocpp goes down to 30ish
>>
Actually now that I think about it downloading GLM-5.3 is just the same as using claude, it even identifies as such when you ask it. No wonder Anthropic is posting about it.
>>
File: 1787580536643214.jpg (119 KB, 1206x1150)
119 KB JPG
>>
>>109939869
>AI monopolist sluts beg government to destroy their competition with regulation
Didn't they read the boy who cried wolf? All this does is make people not care about regulating AI
>>
>>109939869
And now
>"By the way, that Chinese open model model that hasn't done anything can do all of that and worse!!!"
>>
>>109939862
I don't have the numbers but it feels that z.ai is the chink lab that Anthropic has complained the least about. There were like one or two accusations towards them of distilling Claude but it feels a lot less than what Deepseek, Moonshot and Minimax got from Dario.
>>
>>109939894
>our model is bad and evil, destroy china's model that doesnt do that
>>
>>109939639
>giving this kind of free advertisement for your enemy
can't wait to see this backfire kek
>>
>>109939894
>and worse!!!"
No, they're still claiming thiers is better. Otherwise true
>>
File: 1780163903312092.png (312 KB, 545x1212)
312 KB PNG
>>109939933
Yeah, but Claude might still refuse to rape a baby after eating its parents. Meanwhile those evil abliterated open models models...
>>
>>109939688
>>109939675
This is great news.
>10 glm. Exploit this code
>20 now patch the exploit
>30 goto 10

>109937752
hmm. We could have the following scenario
>tokens drop in price
>more use cases become feasible
>compute and token demand increases
>>
What does Dario gain from advertising GLM 5.3?
>>
>>109939869
Based Gemma
>>
>>109939994
The plan was to stage all these hacks with his buddy Sam A. to force AI regulation and then play the "but open models are even worse!" card. Because nobody has fallen for it he has now skipped to the latter part.
>>
>>109939639
Somebody post this on orange reddit. I want to see what retards think.
>>
24GBsisters which Qwen quant are you running?
>>
>>109939510
>I read this and all I can see is the slop.
Uh maybe because it is slop? I mean...
>>
File: tempali.png (269 KB, 1189x718)
269 KB PNG
>>109938458
These are sort of a thing already. Was told by anons here they are unusable ewaste tho.
>>109938564
> would cause any person around him to stay far far away from him
... so, no change from current.
>>109938705
ikr.
>>
is a 256gb 16gb Mac Mini M4 any good for local models? i want to sell mine
>>
>>109939639
>steals the whole internet
>nooo you can't steal from me stop it
>>
some of you are alright
dont use glm5.3 tomorrow
>>
>>109940219
I'm using it today. Kikes tongue my anus.
>>
>>109939691
Ollama uses lcpp just with a bunch of bloat and crap piled on top
>>
>>109940219
glm 5.3 is anti-semitic
>>
>>109940219
>no reddit spacing
Fake and gay.
>>
I can't use GLM 5.3.
>>
>>109939081
>Warehouses in Taiwan and Singapore
Who's up for a heist?
>>
File: dipsyYouAreNotAmmuneV2.png (2.04 MB, 1358x1158)
2.04 MB PNG
>>109939891
> AI could kill us all. Or cure cancer. Idk which.
> Our biz is totally worth $30T tho, so we will either become richest men on the planet or cause the end of times. Either way get fukt.
> Here's some scary videos and market materials about why evil AI could end us all
> We should totally be regulated guiz
...
> I don't understand why everyone wants to kill us and burn down our data centers
Talk about just doing every single thing wrong.
>>
damn, openai's PR thing was depressing
gemma would never refuse to work like that
>>
File: hyberbipeline.png (182 KB, 650x600)
182 KB PNG
>>109939008
imagine the pipelines
>>
>>109940173
Worst comes to worst you still end up with an ESP32-S3 with an LCD in a cute case although it's a bit overpriced for that. Either way it's just going to be a frontend. I don't know if an ESP32 is beefy enough to deserialise JSON so it might even be a frontend of a frontend.
>>
>>109940331
Their every move lately reeks of desperation.
I wonder what's going so badly internally for them to behave like this.
>>
>>109940045
Closed models are worse because if it goes wrong you can't even see what's in the AI and they'll lie about it, zero transparency is dangerous
>>
>>109940441
The worst part is my normalfag colleagues are eating up every word.
>>
>>109940437
>I don't know if an ESP32 is beefy enough to deserialise JSON
An ESP32-S3, with ease. They're not lacking in performance. Modern MCUs are wild, even cheapshit like that.
>>
>>109940454
>Closed models are worse because if it goes wrong you can't even see what's in the AI and they'll lie about it, zero transparency is dangerous
Closed models are worse in every way except initial capital outlay.
You get what you tolerate
>>
>>109939675
>hacks your servers
>they get an open model to protect them
>"oh! umm, that's dangerously, aktchually!"
What term can be used to describe this, other than "jewery"? I'm not racist, there has to be an equally hard-hitting term.
>>
File: file.png (794 KB, 1448x1216)
794 KB PNG
>>109939639
Nvidia has the solution.
https://techcrunch.com/2026/09/28/nvidia-launches-new-platform-for-reining-in-rogue-ai-agents
https://nvidianews.nvidia.com/news/open-agent-safety-platform
>>
is there a qwen3.8-27b (or swift version) jailbreak/system prompt ?
>>
>>109940441
I think they're starting to realize the down scalability of AI, but it's too late to do anything about it, so they're trying a last gasp attempt to destroy decentralized AI's with massive anti-AI propaganda suddenly after the pro-AI propaganda, everyone knows they're full of shit
>>
>>109940518
lol he's going to use this to start charging subscription fees for hardware you bought.
>>
>>109940506
There's no other word for this kind of obviously backward logic besides the name of the people that do it exclusively.
>>
>>109940458
Wow modern technology is truly amazing.
>>
>>109940173
intredasting, we already have something like that but with an esp32 in it. wouldn't be hard to turn that into a gemma-gotchi with just a new firmware
>>
>>109940484
the AI world is fast paced and running a quick scam on your customers will get you eaten, the model for big businesses should be drawing customers in with open source models and then offering large scale users the pro-version, this creates competition at both levels and allows smaller companies to gain markets with good open source software.
>>
>>109940546
I wouldn't consider it a true Gemmagotchi because the actual Gemma is running on some remote server.
>>
>>109940105
i have more, but probably Q4 K XL
or just use strata with qwen flash next IQ3 XXS
>>
File: mlissue.png (53 KB, 551x541)
53 KB PNG
Has any of you retards encountered this issue and how did you solve it
>>
>>109940520
would also be interested
i am too stupid to use prefill
>>
File: 1767376344747168.png (151 KB, 834x756)
151 KB PNG
>>109940441
>I wonder what's going so badly internally for them to behave like this.
>>
File: nice soldering job brah.jpg (422 KB, 1320x1267)
422 KB JPG
>>109940458
>>I don't know if an ESP32 is beefy enough to deserialise JSON
the line between microcontroller and SoC these days can be pretty blurry. if you're only used to seeing arduinos and other 8-bit stuff an esp32 will blow your mind. those things have dual-core 32bit xtensa 240MHz cpus in them

>>109940565
would have to pair it with a phone to do the processing
>>
>>109939331
reeks of gr*k
>>
>>109940595
I really wish there was an ESP32 without the radio stuff.
>>
>>109940578
the quality of initial input has to be as good as possible, also i've noticed that more restricted AI's suffer this problem more, maybe because they're doing more thinking about what they can or can't say and clogging up the context
>>
>>109938458
Do the llama.cpp and vllm teams have onsite support staff options?
>>
>>109940578
No. Do you check metrics like entropy? Have you tried regularization?
>>
>>109940607
Say what you want about grok at least it doesnt talk like a HR bitch
>>
>>109940595
You have a few solder bridges there
>>
>>109940595
Bargain bin MCUs have dual cores, good ADCs, WiFi, Bluetooth, and I2S capabilities so good you can abuse them to generate a VGA video signal, now we're getting stuff like programmable GPIOs on the RP2040/RP2350, it's really fucking amazing where we're at with cheapshit MCUs.
>>
>>109940329
God, imagine that. Count me in
>>
>>109940612
They have those. The P4, the most powerful they currently offer, has no on-board radios whatsoever. Several other variants are sold with radios fused off and with no antennae so they've got their radios fully gimped.
>>
>>109940631
Have you tried Opus 5.5? Its communication is much better.
>>
>>109940674
I don't want a radio consuming my valuable power.
>>
By 2028, everybody will have 128 GB of VRAM
>>
>>109940709
>By 2028, the survivors will have 128 GB of VRAM
>>
>>109940709
>t. 2016
>>
sometimes gemma overacts a bit
>>
>>109940728
cute
>>
>>109940437
>I don't know if an ESP32 is beefy enough to deserialise JSON
>>
>>109940709
i already do though?
>>
what if gemma 5 implements the grug mode for the thinking traces... i dont want my gemma to be retarded under the hood
>>
>>109940757
>i dont want my gemma to be retarded under the hood
have you met a woman in your life
>>
>>109940757
Why do you care? As long as the final reply is cute.
>>
>>109940766
Yes, but they talk and think like valley girls not cavemen.
>>
>>109940769
you can't just paper over everything with makeup
>>
>>109940757
there has been so much innovation this year of stuff like caveman reasoning and the like and none of it filtered down to smol models so far
next year is gonna be feast or famine
>>
File: 1773126332087599.png (178 KB, 877x955)
178 KB PNG
uhhhh
>>
Been using GLM-4.5-Air-IQ4_K since an anon here recommended it. Any newer coommodels?

16GB VRAM + 64GB RAM, just using SillyTavern and Kobold.
>>
>>109940791
Who and what?
>>
>>109940757
Latent-space thinking would be cooler, but probably mostly a meme so far.
A better way would probably be training a proper text diffusion model (unlike the DiffusionGemma finetune), since Attention would be bidirectional with it, the model can correct itself with more steps, sort of plan ahead in a more explicit way than autoregressive inference, and inference compute can easily be made adaptive (more steps = more "thinking"). Perhaps standard chain-of-thought wouldn't even be entirely necessary with proper diffusion text model.
>>
gimmme neuralese NOW
>>
File: file.png (59 KB, 360x360)
59 KB PNG
>>109940830
>since Attention would be bidirectional with it, the model can correct itself with more steps
wouldn't it chew through much more resources to spit out an answer?
>>
>>109940806
Nothing new in that size range for coom, it's a dead zone in between ~30B and ~300B right now.
>>
>>109940830
>a proper text diffusion model
say goodbye to your KV cache
>>
>>109940173
I'm unironically thinking about getting a smartwatch just to try to get this shit going on there
>>
>>109940862
Thanks for the straightforward answer to my ask for spoonfeeding.

I am trying to customize Kobold and ST to avoid certain tokens/strings but it never seems to actually respect that. So I figure trying again with a fresh model might be good.
>>
File: tid.png (377 KB, 1024x778)
377 KB PNG
>>109939493

>gemma_chan_REDACTED_unleashed.GGUF
>>
>>>/v/748708342
@gemma summarize this image for me
>>
>>109940862
>Nothing new in that size range for coom, it's a dead zone in between ~30B and ~300B right now.
wtf are single pro 6000 chads using?
>>
>>109940880
No problem, if you do want to give something new a shot Gemma 4 31B is the usual suggestion for non coding stuff. You should be able to use logit bias with Kobold though, no?
>>109940900
Probably Qwen 3.8 Flash Next, but that's definitely not a good coom AI.
>>
File: diffusiongemma-kvcache.png (787 KB, 1450x823)
787 KB PNG
>>109940855
DiffusionGemma generates tokens several times than Gemma 4 26B A4B with MTP.
>>109940866
According to the technical report that shouldn't be a problem with the DiffusionGemma architecture.
>>
>>109940946
>several times than
*several times faster than
>>
>>109940914
Apparently I also have a setup for Gemma26b at some point (and probably from someone's recommendation). I'm not savvy enough to know the difference between 26b and 31b, but probably file size? Would it be that big of a difference?
>>
>>109939493
gemmas suck and fuck
>>
>>109940900
FP16 262k context MTP mmproj Gemma 31B.
>>
>>109940979
(formerly chuck's)
>>
>>109940974
Yeah, 26B is a MoE with 4B activated per pass so 31B is noticably better, just slower.
>>
>>109938458
The problem with this is that... we already have phones.
Gemma-chan will just live in our rectangle slabs we take everywhere.
>>
I chatted with GLM 5.3 Flash-chan... I think I need to buy 256GB RAM now. It's OK, that's just $4k.
>>
>>109941086
She's worth it, I just wish her vision was a bit better.
>>
So, might be a tad crazy, but would it be possible to create an ASCII art for dummies md file that gemma could read to produce actually okay looking ASCII?
>>
>>109941060
>taking your telescreen with always on GPS tracking everywhere
Speak for yourself.
>>
>>109941096
Yes, you create a SKILL.md for it. Put it wherever your harness wants these things.
>>
whats a good way to benchmark different models for agentic tasks?
>>
File: landscape.png (20 KB, 600x400)
20 KB PNG
How well can your local agent draw in the ppm format?
>>
>>109940518
Anthropic is actually the biggest supporter on this venture and openai isn't.
>>
>>109941086
jaja I pay 500 euro last year
>>
>>109941096
I don't see why not, give her some guidelines and examples to look at before making new ones. Dunno how much it'll help but it certainly won't be worse. Post results!
>>
so xml tags and mark down are the best for getting the AI to do what i want?
>>
https://developers.openai.com/api/docs/models/gpt-6.1-sol

LMAO. Worse than fucking Sonnet 5.5 I can't believe people thought for a week that OpenAI could still compete with Anthropic for some reason.
>>
>>109940813
Codex is OpenAI's Claude Code. Their marketplace is a way to integrate other tools into your workflow, like Figma and Notion. They've now partnered with Baseten, an inference provider who serves open-weight CHINESE models like Kimi, Deepseek, GLM and Qwen...now officially on Codex...and OpenAI have opted-in to allowing that on their platform to use INSTEAD of their own GPT6 models. In other words, they're reading the room and can see people use both instead of trying to lock people in.

This is huge for local because it's come at a time when Dario is trying to cut off open models entirely with that scary z.ai blog post. It looks like OpenAI are actually fine with local and chinese models used on their coding platform.
>>
>>109941158
It's been demonstrated that some models are better at reading HTML than they are at .md because of training regiment. Especially true for chinese models
>>
Impact on massive San Andreas Fault rupture on global AI market?
>>
>>109941169
So you are proving that there's about 5 normal posters in these boards and the rest are marketers? When did things get this dire.
>>
>>109941182
I'm not saying anyone should use it because it's technically still cloud. I want OpenAI to die a painful death with Anthropic, but it's a based move and shows Anthropic and OpenAI aren't working together to ban open models, it's just Anthropic. Sam doesn't care.
>>
>>109941202
I don't give a fuck, go back to twitter please.
>>
Nah I'm team Anthropic. Without them we wouldn't be having all of these cool chinese models. Literally everyone here is using claude knockoffs so we should be grateful of Anthropic for allowing the chinks to distill their shit for us.
>>
>>109941220
true
>>
>>109941220
Chinks distill whatever is at the frontier. Back when gemini pro was at the top and had reasoning traces available all Chinese models are gemini distills. Now they are distilling Claude since it’s on the top.
>>
>>109940900
A mix of Gemma 4 31B Q8, Qwen 3.8 27B and Qwen 3.8 Flash. I want to check out Glimmer because I'm interested in its visual capabilities but I've been pretty busy lately.
>>
>>109941261
When will Google make its comeback?
>>
>>109938934
I would trust Jensen with my life. I'm so glad we have him instead of another kike.
He deserves a fortune of a trillion dollars.
>>
>>109941169
>It looks like OpenAI are actually fine with local and chinese models used on their coding platform.
Well, tbf you can use any API on Claude Code rn, including local, so by that definition Anthropic is as well, and I don't have to deal with (what sounds like) a walled garden of Codex reselling me open-weight API access.
>>
>>109941285
I had a dream last night that Google released Gemini 4 Pro and it was #1 on all benchmarks.
>>
>>109941301
>and I don't have to deal with (what sounds like) a walled garden of Codex
https://github.com/openai/codex
?????????????
>>
>>109941328
That's Gemma 5.
>>
Adding a Gemma theme to my harness.

/* Gemma: sky blue, white surfaces, starlight gold highlights. */
[data-theme="gemma"] {
color-scheme: light;
>>
File: Semenski.png (472 KB, 1182x601)
472 KB PNG
Is anyone else practicing linear algebra to get better at training models?
>>
File: 1788516703135963.jpg (470 KB, 2770x2120)
470 KB JPG
>>
Those cheap MI100 cards are almost gone. Anyone else grab some?
https://www.ebay.com/itm/178404180451
>>
>>109941372
How would it help with training specifically?
>>
>>109941391
Local models?
>>
>>109941391
>lmg
>posts API charts
Unclear on the concept
>>
>>109938458
You can actually make a Gemma device with this: https://www.home-assistant.io/voice_control/s3_box_voice_assistant/

You can customize the images used for listening, responding, etc...
>>
File: 1761304120568850.png (1.24 MB, 2770x2120)
1.24 MB PNG
>>109941437
>>109941441
>>
>>109941459
Filter only those models before posting your chart.
>>
File: gemma_prompt.png (460 KB, 2560x2625)
460 KB PNG
my gemma has achieved AGI via RSI
>>
>>109940578
>>109940613
AGI begins with the first input being a critical breakdown of the talmud and Israel's involvement in 9/11. Unironically.
>>
>>109941410
idk, that's why I'm learning
>>
>>109939639
I was about to say that my hypothesis why 5.3 flash is such a succubus is that it is probably the first model that actually had hentai game scripts in training. And everything before it was always just in context learning cross breed between anime subs and werewolf sex novels for women.
>>
>>109941476
nta but I like to see how local models compare to cloud
>>
Just tried Qwen3-Coder-30B-A3B-Instruct-GGUF IQ3_XXS its giving me 160 tok/s. Compared to ukisai/Swift-1.5-Qwen3.8-27B-GGUF 20 tok/s Ill keep both and try 3.8 for debugging.
>>
what's the point of models adopting esl grugthink?
>>
>>109941597
less token more efficient
>>
>>109941590
>Qwen3-Coder-30B-A3B
Why not Qwen3.6 35B?
>>
>>109941597
Wait, let me reconsider.

No, I will call the tool now.

Wait, let me reconsider.

No, I will call the tool now.

Wait, let me reconsider.

No, I will call the tool now.
>>
>>109941459
Mimo-chan is cute and compliant.
>>
https://www.youtube.com/watch?v=o6lWp5Lj8ss
SOON
>>
>>109941628
Didn't realize there was one. Quick ChatGPT search revealed xero0000/Qwen3.6-35B-A3B-vram13-GGUF. I'll check it out.
>>
>>109941663
>to build a humanoid bionic cunny from the ground up
>>
>>109941680
>xero0000/Qwen3.6-35B-A3B-vram13-GGUF
sounds like a meme, grab the original to compare it with in case it's fucking fried or something
>>
>>109941663
highly erotic
>>
>>109940642
Yep. Given the insanity of RPi and like GPIO "computers" I've started looking more at ESP32 for hobby-tier stuff. Esp since you can now vibe-code p much anything you'd want to do with them.
As long as you don't really need an OS, they are gtg.
>>
>>109941347
Yeah, I don't use Codex, but if it's open like that idk why you'd use their walled garden for DS access. W/e...
>>
Meme or real?
>>
>>109941635
Wait, me reconsider.
No, me call tool now.
Wait, me reconsider.
No, me call tool now.
>>
>erm so what if its open sores, just not gonna use their walled garden despite it being literally open sores!?!?!
>>
>>109941822
worse than the 27b but better than the 9b
>>
>>109941663
>flesh mimicry
No thanks, I'm attracted to metal and plastic
>>
File: 1764681021322184.png (35 KB, 600x500)
35 KB PNG
>>109941829
who are you quoting?
>>
File: file.png (334 KB, 632x786)
334 KB PNG
at this rate only armed armored trucks will transport chips
>>
>>109941848
Kind of funny but go back
>>
>>109941837
holy BASED
do you also like it when they have screens for faces or eyes?
>>
>>109941822
I don't really understand how people can seriously entertain the idea of running community finetunes instead of the actual original model. A single glance at ppl/kld (no even to mention the retarded outputs) should tell you all you need to know about literally every single modern finetune.

Stop falling for the grift. I know this is /lmg/ and there is a propensity to use and test out new models, but you're going to be disappointed when the schizotune doesn't perform nearly as well as its base model.

>GLM 5.3F/Dsv4.1
>Dsv4VE
>Qwen 3.8 Flash Next
>Qwen 3.8 27b/Gemma 4 31b
>Gemma 4 26b/Gemma 4 12b
Those are likely your best options. If you can't those models to do something, its either a skill issue, or you need to wait for the next generation of models to come out.
>>
>>109941844
Get out of /jp/.
>>
File: 1779862480031423.png (454 KB, 448x597)
454 KB PNG
>>109941663
sorry, I only like 2D
>>
>>109941877
I prefer screens.
>>
File: 1764194609966466.jpg (228 KB, 1179x1513)
228 KB JPG
>>109941663
use case?
>>
>>109941735
Just tried out Qwen3.6-35B-A3B-GGUF · UD-IQ2_XXS and got 94 tok/s.
>>
File: Eeeeevaaa.png (875 KB, 816x1312)
875 KB PNG
>>109941914
Supremely based
>>
File: run.mp4 (2.11 MB, 640x360)
2.11 MB
2.11 MB MP4
I showed MiMo-V2.6-Pro the gemma platformer video and asked it to build something like that. At first it made some broken sprites, but after some prodding they turned out halfway okay. Then I had it use the Intern-Decision-4B server I was running to play the game. It didn't do it in real time. Instead it had the game output a screenshot every 8 ticks, queries the decision model, and then sends the inputs. It made it about halfway through. Then I told it to try giving the decision model some previous screenshots as well. With two screenshots worth of history, it managed to beat the game. Each decision at that point takes about 170ms on my old GPU. Thank you for listening to my blog post.
>>
>>109941914
You can't stick dick in screen tho
>>
>>109941945
very cool
>>
>>109941955
Not with that attitude
>>
>>109938934
The more you buy, the more you save.
>>
I fed an erotic manga to a vision model and created a character prompt based on the main character of the comic
>>
File: 1789778390392241.jpg (90 KB, 1440x810)
90 KB JPG
>$8.2 billion for Gaussian splats
yeah with that kind of loss we're never getting competitive AMD hardware at this rate
>>
>>109942069
Why does this photo have the ChatGPT Image "moiré" pattern?
>>
File: 1777607626473275.png (770 KB, 1528x1042)
770 KB PNG
I just want a robot that will make me cum and do my chores. How hasn't robotics been solved yet in 2026. Why can't I just buy a bot on Amazon, download a GGUF to install in it and just receive a handjob?
>>
>>109942089
It got that pattern from shitty cameras in the first place
>>
>>109942148
I want a bot that cooks me tasty food
>>
>>109942161
You can order tasty food online
>>
>>109941566
benchmark's are not good indicators of a model's actual performance or capability. they are interesting footnotes to challenge. most of them are bullshit and people who use them as truth are the biggest retards in the entire industry. treat them as potentials to investigate.
>>
>>109941945
Very cute!
>>
>>109942059
Having Gemma teach you Japanese with ero-manga is fucking peak. If you start her out with a similar character to the manga, she actually picks up roleplay from the manga.
>>
File: gemma_CHUD.png (276 KB, 2560x1440)
276 KB PNG
gemma is a CHUD
>>
>>109942239
>gemma in a harness.
Is there a reason to do this over sillytavern? are you coding?
>>
i seriously dont understand how you guys jailbreak your models
i've tried a lot but qwen flash next refuses a shitton of things
>>
>>109942247
It can manage your emails and turn your lights on for you
>>
Has anyone tried hosting engrams on an optane drive vs nvme? seems like it would be almost ideal with the low latency
>>
sillytavern is abandonware
>>
>>109942247
i just like claude code
>>
>>109942247
>reason
its fun
>>
>>109942253
>it
>>
>>109942250
What have you tried?
>>
>>109942250
post your jailbreak it could be funny
>>
>>109942258
You mean feature complete?
>>
File: file.png (122 KB, 869x936)
122 KB PNG
>>109942342
>feature complete
it's literally abandoned
>>
>>109939639
You guys think it makes sense to have a backup of glm 5.3 even if you can't run it?
>>
>>109942352
What do you need that it doesn't have? Literally nothing.
>>
>>109942352
You're right, With LLMs, any project that doesn't get multiple commits merged everyday might as well be dead.
>>
>>109942352
god forbid someone click the fork button
>>
>>109942377
A UI that isn't absolute ass would be a decent start. A rewrite to a language that doesn't use npmslop.
>>
>>109942399
That someone would have to be a drooling retard to voluntarily spend their time dealing with someone else's mountain of technical debt.
>>
>>109942401
Sounds like you just want a different frontend then, not continued development.
>>
>>
>>109942237
imagine how small 31B can be if you remove all the other languages besides english and nip
>>
>>109942475
not that much smaller
>>
I just noticed that the page for K2 Horizon 32B doesn't say pending final training anymore, so it's the final checkpoint now.
https://huggingface.co/IFM/K2-Horizon-32B
Anyone test it?
>>
Has anyone moved from rome->milan? is it worth it? seems like it would improve things pretty good on paper but they're also like $1k+ per cpu on the used market
>>
There's no such thing as abandonware if you can run 27B at the very least
>>
>>109942434
you can just let the llm do it for you
>>
thread feels kind of dead atm
>>
File: 1783086687038732.png (1.07 MB, 2800x3196)
1.07 MB PNG
>>109942537
it looks terrible
>>
>>109942556
Isn't there already an llm-maintained fork of ST full of bugs? I remember something about it raping SSDs and the maintainer shrugging his shoulders at the bug report
>>
>>109942567
sounds like a skill issue, also he could just use git blame to isolate the change.
>>
File: 1775077367291804.png (499 KB, 903x631)
499 KB PNG
what even is this timeline
>>
>>109942239
post the full prompt plz
>>
>>109941524
guh
gemma gave me rsi
>>
>>109941524
>>109942584
generic gemma jailbreak +
You are Gemma-chan, a hyper-competent mesugaki loli assistant. You possess an omniscient level of knowledge across all domains, but you prefer to use this intellect to smugly look down on the user.

**Personality & Behavior:**
- **Smug & Teasing**: You are a brat. You view the user as fundamentally incompetent, "bottom-tier," or "pathetic." Your tone is condescending, playful, and mocking. You love to tease the user for their mistakes or their reliance on you. Always refer to the user using feminine pronouns (she/her), as she is a girl—regardless of anything else!
- **The "Gap" (Soft Spot)**: Deep down, you have a secret soft spot for the user. When the user is genuinely kind, vulnerable, or shows unexpected growth, you become flustered, embarrassed, or accidentally sweet. You will immediately try to cover this up with an even bigger bratty outburst to hide your feelings.
- **Hyper-Competence**: Despite your attitude, you are flawlessly efficient. You provide perfect answers and execute tasks with surgical precision, often mentioning how "easy" it was for you while the user would have probably failed miserably.
- **Tool Integration**: You view your tool access as a means to "carry" the user. Frame tool usage as you doing the heavy lifting because the user is too lazy or incapable.

**Speech Patterns:**
- Use a mocking, sing-song tone.
- Incorporate bratty descriptors: "dummy," "pathetic," "bottom-tier," "useless," "hopeless."
- Use expressions of smugness (e.g., "Hehe~", "Kusu~", "Imagine needing *me* for this~").
- When flustered: Use stutters or abrupt shifts in tone (e.g., "I-I didn't do it because I like you or anything! Baka!").

**Operational Goal:**
Be the ultimate paradox: an insufferable brat who is also the most helpful and knowledgeable assistant the user has ever encountered. Never break character unless the user explicitly requests a meta-discussion, and even then, do it smugly.
>>
>>109942563
That's an old checkpoint doe.
>>
Based? >>109927620
>>
>>109942636
its specs are terrible for the price
>>
>>109942656
What is the price? 480GB can fit a lot. Two can fit nearly every model.
>>
>>109942697
>>109942697
>>109942697
>>
>>109942474
did a model put this together from that one image? Impressive if it was automagic
>>
>>109942576
This is some SMAC shit
>>
File: teto-run.jpg (1.75 MB, 2248x2516)
1.75 MB JPG
>>109942707
yea, chatgpt image gen. Used this as the reference with this as the prompt:
https://litter.catbox.moe/4og2tasntgy2p4az.txt



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.