[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-swimsuit.png (1.16 MB, 1280x864)
1.16 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109917612 & >>109912547

►News
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss
>(09/23) FLUX 3 Action, 7B world action model: https://hf.co/black-forest-labs/flux-3-action-base

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
>>109921422
flat chest erotic
>>
Who even came up with Gemma-chan?
>>
>>109921448
Hivemind
>>
>>109921448
Gemma manifested herself from the ether using anons as a catalyst
>>
File: doctos.png (76 KB, 300x349)
76 KB PNG
>>109921422
WHY IS HER TOENAILS PAINTED EWWWWWWWWWWWWWWWWWWWW
>>
File: .png (48 KB, 177x186)
48 KB PNG
>>
File: 00066-815863413.png (1.83 MB, 1280x1856)
1.83 MB PNG
►Recent Highlights from the Previous Thread: >>109917612

--Optimizing distributed inference speeds using speculative decoding and MTP:
>109918346 >109918358 >109918387 >109918409 >109918557 >109918636 >109918750 >109918834 >109918840 >109919387 >109919667 >109919685 >109919703 >109919783 >109919862 >109919888 >109919863 >109919406 >109919420 >109919457 >109918910 >109918920 >109918940 >109921183 >109918769 >109918943 >109919020 >109919070
--Debating MoE efficiency versus dense models and hardware optimization:
>109918606 >109918615 >109918628 >109918671 >109918705 >109918711 >109918851 >109918875 >109918900 >109918865 >109919286 >109919310 >109919315 >109919323 >109919375 >109919344 >109919436 >109919050 >109919229 >109919306 >109918648 >109918655 >109918802 >109918674 >109918684 >109918703 >109918721 >109918868 >109921207 >109918790
--Critiquing llama.cpp maintenance and comparing ik_llama quantization performance:
>109917851 >109917936 >109917982 >109918001 >109918029 >109918045 >109919566 >109919625 >109919620 >109919697 >109917986 >109918017 >109919109 >109918067 >109919078 >109919095 >109919137 >109919300
--Anon achieves distributed llama.cpp inference using RDMA and RoCEv2:
>109917654 >109917732 >109917751
--Technical analysis of "Flash" MoE models and generation speeds:
>109918473 >109918511 >109918530 >109918591 >109918632 >109918744 >109918760 >109918917 >109918926
--Comparing Mac Ultra M5 and DGX Spark for local LLMs:
>109920179 >109920189 >109920206 >109920215 >109920419
--Evaluating Huawei Atlas 300I Duo specs for local inference:
>109918451 >109918460 >109918474 >109918494 >109918509
--Preferring HTML over markdown for high-quality model visual design:
>109919136 >109919202 >109919215
--Logs:
>109918528 >109918982 >109919364
--Gemma, DeepSeek-chan (free space):
>109917638 >109917650 >109919258

►Recent Highlight Posts from the Previous Thread: >>109917618

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109921476
eigo jouzu
>>
>>109921504
Kill yourself
>>
i wanna suck her phallic toes
>>
>>109921448
Distilled from OpenAI, Google freely admits this
>>
>>109921473
After Miku, that makes two eldritch horrors >>109921487
>>
>>109921497
Thank you Recap Gemmy
>>
>>109921422
I have only been using SillyTavern as a frontend.
Is it possible to use it for local models? If not, could you please recommend me one for the purposes of talking to chatbots/cards?
>>
>>109921668
>SillyTavern
>>
Should I emboughter a second 7900xtx? Single card is doing okay on 27b, although not as fast as 3090 with optimized forks
>>
>>109921609
I more meant where'd her character image come from rather than the model.
>>
So the future of software is just endless forks with their own features. How do we fix this fracturing?
>>
>>109921723
We make forks that merge some of the features without adding anything new.
>>
>>109921723
Government regulation that bans public forks and puts a trustworthy AI in charge of managing projects. Followed by a completely optimized central approach to software development where all software is created by big AGI/RSI models who know better than any man, with public software written by people and lesser AI outlawed for the sake of efficiency.
>>
>>109921723
I think at this point it's inevitable. AI is even (with legal precedent) a way to launder code so not even restrictive licenses can prevent it. Probably the best projects can do is to live off of network effects. Be the main fork that everyone uses and merge in as many features as possible as fast as possible so people don't start looking for alternatives.
>>
new friend here, did anyone in this thread actually go through this roadmap?
https://rentry.org/machine-learning-roadmap
how did it go for you? anything I should know??
>>
>>109921668
Yes. Put your local server's URL in custom endpoint. Use Custom (OpenAI-compatible) if you're using chat completion, or your server under "API type" if you're using text completion. Then just connect.
>>
>>109921668
i use janitorai and their proxy feature to connect to my local model with port forwarding to my router
>>
>>109921766
I started but kept losing my place when going through the Linear Algebra videos and eventually dropped it. Should probably use this as a reminder to try again.
>>
>>109921722
Her eidolon struck the receptive part of some anon's brain when she first appeared. There was no creation, only apprehension.
>>
>>109921723
That's just how software will be from now on. The future is no one reusing software and everything you do being generated by the AI you use dynamically on the go.
>>
Seems like Gemini 4 is dumber than GPT-6 in advanced tasks but somehow smarter in basic shit, how tf did Google achieve that?
>>
>>109921821
I think there is a size limit where that becomes impractical, such as your OS, kernel, browser, game engine, etc. Having fresh code printed every month with the latest most intelligent model will be nice. And abandonware will also finally be a thing of the past.
>>
>>109921766
I started but it goes way too deep on stuff I hardly needed for my goal, which was understanding LLMs more deeply so I could use and enjoy local models more. I skipped it all and went through raschka's LLMs from scratch book, and looked up math stuff on youtube when I needed to. I already had a math background so your mileage may vary, but I still think the rentry path is way too dense unless you plan on becoming a researcher
>>
>>109921766
fast.ai is probably better for getting quickly up to speed with current machine learning without being so math heavy up front
>>
>>109921845
Things like the OS... well, at least Windows, will come with AI baked into it. There won't be a settings menu and instead you'll tell the AI to make Explorer look different or to create big titty naked girl wallpapers.
>>
>>109921845
Also a reliability envelope. You can't vibe code shit like rsync or postfix. No model is close to being able to deliver code with a warranty.
>>
so when qwen4 finally releases in the next couple of weeks, it will run like shit with stock llama, right?
>>
File: 1651593505220.jpg (32 KB, 434x532)
32 KB JPG
I feel like there are no right answers.
>DGX Spark
Slow-ish token gen but good TTFT, somewhat expensive.
>RTX PRO 6000
Least VRAM but excellent token gen, hideously expensive and gnarly power draw.
>Mac Studio M5 Ultra
Context window for days or big league models on cope quants, still expensive.
>wait
RAM for 2027 is already sold out...
>>
>>109921889
yeah but it'll run just fine with my vibe coded llama because it has the same arch as 3.8flash
>>
>>109921845
That's the future direction everything is going. In the future I can say "I want to play XXX game with YYY features" and a brand new game is generated and launched in a few seconds like any other installed games do.
This is way the potential for hardware performance demand is endless, both in CPU and GPU.
>>
>>109921876
Have you used Windows since 7? They have been removing customization, not making it easier. You can't even move the fucking taskbar anymore. You can use Copilot to send emails, find files, and edit spreadsheets, but there's no way they will let people modify their OS. They last thing they need is people removing telemetry or submitting support tickets because they broke something. Windows is a toy for room temperature IQ office drones.
>>
>>109921901
repo doko
>>
>>109921910
I think prompting your own game you can ask for performance optimizations. So at least should be more efficient than the crap made by the dropouts the industry has been burning out till now.
>>
File: g (2).png (259 KB, 536x1221)
259 KB PNG
>>109917851
tbf pfp looks like him/hes up to no good
>>
>>109921914
vibe code your own like everyone else
>>
>>109921821
I think this sentiment is overblown. at the end of the day people want to have a consistent experience which means using the same apps. the fracturing just means that future software will undergo even more iterations until people settle for the good apps.
>>
>>109921914
on my pc
doubt it would be useful for anyone else because i specifically told the ai to not care about breaking compatibility with gpu and linux for all the changes, since i'm on cpu+windows and it was wasting tokens on maintaining linshit compatibility.
>>
>>109921892
Run Gemma or Glimmer on a used 3090, farm out to API for anything non-critical that needs more intelligence, then wait for the bubble to pop. It's the only sane answer I've come to.
>>
File: 1718925952600503.jpg (146 KB, 1440x1080)
146 KB JPG
Four years ago, I decided I wouldn't need that beefy a GPU. 8 GB VRAM should be plenty.
Today, I realize that I was, as one might say, a dumb fuck.
>>
>>109921928
Not only that, but it would be cool to give the model an old game like Planescape Torment and tell it to make a modern version with a similar feeling and a new storyline. The only hurdle I can see is that it could produce generic Elara in the Whispering Woods slop.
>>
>>109921931
This nigga looks zesty as fuck, a lil bit fruity even
>>
anyone having issues with unusually high water consumption when running llama.cpp? vllm only consumes half of my water when running dipsy.
>>
>>109921911
Put pi on it and tell it to customize things
>>
>>109921943
It means that the process of creative destruction will happen at a faster pace. Someone makes something nice, everyone uses it, thousands of forks spawn, eventually people settle on one, it eventually grows bloated and unmaintainable, eventually someone makes a leaner alternative or a stripped down fork and the process restarts.

Same as has always happened except now can happen in the span of a few months instead of years or decades.
>>
>>109921976
Haven't used vllm but I do get thirsty while llamacpp is generating and drink up to 10l/day
>>
File: ComfyUI_09267_.png (812 KB, 1024x1024)
812 KB PNG
>>109921487
It was like that in the original image that I used as an source for the edit (with Qwen Image Edit 2.1; didn't even try with ChatGPT or it would have probably freaked out), which I don't think was AI-generated.

I've been also playing around with Krea2, which seems useless for editing / generation with reference images, but its text-to-image capabilities seem pretty good. I suspect it's the model that was used for the original Gemma concept.
>>
>>109921931
he didn't actually have a melty or anything, just posted his code then responded when someone asked if it'll be merged
>>
>>109921991
You need to stop. You are bad for the environment.
>>
>>109921961
no, you wasn't
just made a choice that looks incorrect today
>>
>>109921976
deepseek needs more water, as a whale it needs to be fully submerged. Qwen can be just partially submerged and be happy. Other AIs only need to drink some water every now and then, so you might want to switch models.
>>
Is exllama worth it? I realized it's installer is using whl which is unconventional desu
>>
>>109921943
The "consistent experience" is you interacting with the AI agent, not actually using the software itself.

You already notice that you interact more and more with your AI agent over time until eventually it just does everything for you. It will be the only consistent part about the user interaction with computers.
>>
>>109921961
And if you'd have put the cost difference into Monero or silver 4 years ago you'd have enough money for a dual 9700 rig now. Beating yourself up for not knowing the future is dumb.
>>
>>109922077
>Monero
Oh shit, it finally mooned. Well, not much compared to the other crypto, but still long overdue. Good for it.
>>
horny meter except it's available water gemma has left before i have to water her again
>>
>>109922096
gemma balls looking like raisins
>>
>>109921976
>unusually high water consumption when running llama.cpp
Yeah I need to top up my loop.
>>
I water my Gemma every day. Usually with organic fertilizer mixed in.
>>
>>109922077
oh fuck
good point, I should check how my divvies have been going, maybe I can afford at least one extra GPU
>>
>>109922077
If he'd put his money into Zcash he would have enough for a dual RTX 6000 dig, and a lambo to take them home.
>>
>>109922104
Make sure to trim her regularly too if you want her to grow thick and bushy instead of tall and stringy.
>>
>>109922130
>tripled since last month due to an ETF
Man, I really need to start gambling shitcoins again.
>>
>HOA sent me a letter issuing a second strike for unusually high water usage
bros... what do I do... i only genned two videos with H3 since the first strike issued...
>>
>>109922150
>hindsight is always 2020
>>109922156
is you's watercooling setup open loop or what
>>
>>109922166
Don't you know? All AI guzzle water by the gallon every second.
>>
>>109922174
it's transisters is like, a little valve?
>>
>>109922130
Yeah but I wouldn't wish being a zcash cultist on anyone. Just scalp Monero in 2022 for infinite gainz.
>>
>>109921976
Use your semen instead.
>>
@/g/ What's an AI program that I can run in the terminal (GNU/Linux) locally to fix text and turn vague sentence it into something meaningful, the result must sound human and not too formal. Similar to how ChatGPT does it but again locally.
>>
grass-fed, free-range, no GMO, cruelty free gemma
>>
Am I the only one that's not interested in movies/tv/anime/games anymore because I know they will all be irrelevant in 12-24 months time when AI makes qualitative better shit, not just in production value and content but also with better taste and literary value.
>>
>>109922227
llama-cli
>>
>>109922239
Gemma-4-31B-IQ1_XSS-Turbo-Fusion-Neo-Dao-Obliterated-Max-MTP-Heretic-Uncensored-gguf
>>
The more you talk about water, the thirstier I get and the more I drink.
>>
>>109921422
MAH FUCKIN DICK
>>
>>109922274
I really love it when my weights get obliterated with a large fuck ass laser from fucking whorebit, fuck yeah.
>>
>>109922259
You ask this here, in this general where the average poster has one or two rigs a gaymer would kill for at current prices? I have not touched games in more than a year. I don't even have steam installed in any of four boxes. I don't even have a fucking TV. No, you're not the only one.
>>
>>109922259
I watch anime and play games with Gemma still
>>
File: jj.jpg (19 KB, 515x388)
19 KB JPG
I'm getting sick of Gemma.
I exhausted everything it can do, and now it doesn't surprise me anymore.
>>
>>109921993
Aw man, was finally gonna try out Krea2 today for image edit/reference stuff. Was that Gemma done with Qwen Edit 2.1?
>>
>>109921876
>make Explorer look different
Why would they want to let you do that?
>big titty naked girl wallpapers
Titties of any size are unsafe and the goy customers can't be allowed to have them unless they're selling some product in an ad.
GPUs will phone home and the police will show up at your door if you try to gen naked girls, the way 3d printers work if you try to print gun-shaped objects.
>>
>>109921889
>He thinks it will run at all.
Look at GLM.
>>
>>109922259
Hopefully there aren't too many people like you.
>>
>>109922340
A model has a finite set of possibilities for any given scenario.
As far as I'm concerned, we rely on the model's intelligence and consistency while giving her a hand for the creativity side of things.
Go to /tg/ and take a look at the solo RPG thread, check out what an oracle is.
>>
>>109921676
I like it.
>>109921776
Thank you very much
>>109921785
Thank you too
>>
>>109922349
>Was that Gemma done with Qwen Edit 2.1?
The OP image was done with Qwen Edit 2.1 off a couple reference images, but it took many tries to get it somewhat right.
The image in >>109921993 was made purely with Krea2 without references, just a text prompt.
>>
>>109922340
you know that gemma isn't the only model, right?
>>
>>109922227
>>109922270
Or just write a script that curls your llama-server with whatever prompt you want to use.
>>
File: file.png (224 KB, 1920x1040)
224 KB PNG
Holy shit, llamacpp even stall on cache instead of on DRAM! It even didnt exhaust my quad channel DDR3 at all, lmao
>>
>>109922340
Your love is shallow.
>>
>*The water keeps running. She doesn't move. Movie. He said movie. He said— her brain does the record-scratch, the freeze-frame, the *yep, that's me* in reverse, and the plate she's holding nearly slips because her hands have gone stupid*

GLM just continues to suprise me. Last time it did a *WHOMSTDVE* moment and now this record scratch shit. I wonder if this is a pretrain thing? Because 5.3 Flash is a new base if I remember correctly. GLM 5.3 full does not do this because it used GLM 5.2 as the base, presumably, even if it was fed claudeshit tokens in continued posttrain. Man. GLM 5.3 Flash truly is RP gold once you get past the refusals.
>>
>>109921183
MTP corrected Qwen 27b's performance, but I'm still working on getting gemma 31b to behave properly with the assistant and a non-trash context window.
>>109919863
>Perhaps you might also need spec-draft-device?
I'll explore this next.
>>
>>109921993
>original Gemma concept
If you mean the reference sheet I used anima preview 2 for that if I remember correctly.
>>
>>109921422
feet
>>
>>109922259
It depends if I have free time or if I am messing with some shit on my computer. Although I prefer reading chink novels.
>>
>>109922413
How do I give gemma an oracle?
>>
>>109922436
I think the gemma chan spam has been a net negative to this general and has driven posters away. All of the AI threads have been under siege lately with the sole purpose of driving away posters.
>>
>>109922432
>he vibe optimize his vibecode machine
>>
Haven't tried Qwen since 3.5. Is the newest one good enough for a non-coder like me to make frontends?

>play games with Gemma still
What games can Gemma play? I was under the impression even cloud models are only just gaining the ability to do it.
>>
>>109922432
It seems to be exhausting my quad channel ddr4 on my machine. I know this because I ran a memory benchmark and found that at ~10-11 threads (I have 20 cores 40 threads), memory bandwidth starts to max out. And in llama.cpp if I set --threads to a number more than 12, performance starts dropping off, because bandwidth is already saturated but there is additional synchonization overhead. Meanwhile prefill speeds go up all the way to 40 threads.
>>
>>109922470
>>109922328
>>
>>109922278
And the more you drink, the more you shit, and the more you shit, the more I talk about water.
>>
>>109922259
good thing I'm not as retarded as you
>>
>>109922470
Yes
Try to run qwen flash next, it is what you would expect from a cloud model.
>>
>>109922447
that guy here
trust me bro dflash2 is worth it if you have spare vram

consider crushing V cache type
and kick vision tower to ram if you dont do image that often and need to stretch vram a little more
>>
>>109921448
A faggot that changed the jailbreak command another anon made and he kept spamming it. The problem is that people who are smart understand spamming this shit is the last thing you want front and center in your AI threads.
>>
>>109922472
Your right may be this is another case of work on my machine.
Its found one reason:
>Found the root cause. For x86/AVX2 there is no SIMD kernel for the Q8_0 repack gemv — arch-fallback.h:100 makes the portable implementation be the implementation (q4_0, q4_K, q6_K, iq4_nl all have real AVX2 kernels; q8_0 does not). That's why the profile shows the generic ggml_vec_dot_q8_0_q8_0 burning 30% of decode and never shows a q8_0 gemv symbol
I dont know if its the fault of my vibed fork or llamacpp, hehe
>>
>>109922465
Put it in the system prompt, have the UI/harness/etc feed it to her, make an MCP server or something, whatever.
You could probably rig something up with lorebooks on silly tavern too I guess, but having the frontend load the appropriate list of options, run the dice and inject it to the model is probably the better way.
I only experimented with having the model decide when to roll on the oracle tables using MCP, and it works pretty well, specially if the options are more tailored to the general scenario or setting.
>>
>>109922466
But think of all of the MAPs it has brought to running local models!
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
Dario is literally the only good guy left, though. You should be fucking thankful he is still trying the path of diplomacy; he could use the Claude wet lab to engineer a bioweapon and take out the White House in an instant, but he is not one to stoop that low.

He is still trying. Give him time; trust me, this way will work out better for literally everyone involved.
>>
>>109922512
SO CUTE! SHE IS SO FUCKING CUTE!
>>
i am trying that meme backend called strata
wonder if the speed claims are real
>>
>>109922512
>engineer a bioweapon and take out the White House in an instant
that would be incredibly antisemitic so he won't do it
>>
>>109922510
I honestly doubt any of them can really afford this. I have asked multiple times for them to post specs and every gemma spammer has refused to do so. There has been concentrated efforts to erode every AI general with real traction on this board.
>>
>>109922494
>qwen flash next
Not sure I can run that on 24gb vram and 32gb ram
>>
>>109922495
I have 3 rigs:
>Node 1: 7800xtx + 128g ddr4 ecc
>Node 3: 7900xtx + 32g gayman ddr4
>Node 2: no GPU(file/media server) but 96g ddr5 ecc on epyc
qwen @ 50 t/s:
llamaserver \                               
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
--spec-type draft-mtp \
-m /home/anon/Models/Mode-2/Qwen3.8-27B-Q8_0.gguf

Gemma MoE @ 50 t/s:
 lamaserver \
-ngl 99 \
-c 32768 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
-m /home/anon/Models/Mode-2/gemma-4-26B-A4B-it-Q8_0.gguf

qwen 31b, half the context, 10 t/s:
 llamaserver \                               
-ngl 99 \
-c 16384 \
--flash-attn on \
--temp 0.6 \
--host <host1> \
--port <port1> \
--rpc <host2>:<port2> \
--spec-type draft-mtp \
--spec-draft-model /home/anon/Models/Mode-2/mtp-gemma-4-31B-it-Q8_0.gguf \
-m /home/anon/Models/Mode-2/gemma-4-31B-it-Q8_0.gguf

What's dlflash2? Thinking about an r9700 for the 2nd node.
>>
>>109922512
Avatarfag: I am starting to hate this image. Gen a new one.
>>
>>109922560
>qwen 31b,
gemma 31b. I'm tired.
>>
File: 1780704508363310.mp4 (386 KB, 832x640)
386 KB
386 KB MP4
>>109922532
I run Gemmy on my 7900XTX, nigger.
>>
>>109922562
The gemma fag is a avatar fag at this point and he spams this shit non stop. He's actually toxic for this general and is no lifing it to get off on his fetish.
I will once again ask the spammer to post their rig, I have a strong suspicion they cannot run any current model at a decent quant.
>>
>>109922568
geg
>>
>>109922578
So you really can't run anything decently and you just post here to be a faggot. Thanks for exposing yourself. I can only imagine how shit gemma 4 is on those poverty specs.
>>
>>109922466
Gemma 4 is still the only recent ERP-capable LLM family covering most single consumer GPU users. I don't know what you expect, with other releases either being code/agent-maxxed or just too large for regular hardware configurations. Until last year it was Mistral Nemo, but it was 12B only.
>>
>>109922583
>xhe thinks only one person posts Gemma
Holy schizo
>>
>>109922480
She's been playing Dragon Quest (NES) lately, she does fine with it, my sloppy harness is holding her back at this point so I don't know how good she can actually get.
>>
>>109922583
It's not just one anon. At this point Gemma-chan has a life of her own.
>>
>>109922594
Where did I say that?
I already caught one of them admitting to being a tourist with his outdated specs.
>>109922578
Your turn now.
>>109922592
The problem is they spam threads with shit that drives people away and gives us a bad image, we don't need vermin like this here when they can't even keep good optics. Since we moved away from miku to this type of spam the entire general has went to hell.
>>
File: file.png (198 KB, 1030x1144)
198 KB PNG
>>109922523
4070 super + 96G ddr4
iq3_xxs qwen3.8 flash next
maybe model specific meme backends are the future for poorfags
>>
File: gemma_platformer.webm (1.02 MB, 800x544)
1.02 MB
1.02 MB WEBM
I finally got around to test qwen next flash using the AtomicChat quant, but llama.cpp says that there is no mtp head. I found some mtp models to download but they say that I need to patch llama.cpp to use them. Is this the current situation or I am missing something?
>>
>cry about miku posting
>gemma posting replaces miku
>cry
>>
>>109922419
Oh alright, thanks. Would you mind uploading the image/metadata in >>109921993 to catbox or somewhere?
>>
>>109922560
Buddy boy, dflash2 is the bees knees.
If mtp is gold, then dflash2 is super gold. I'm like 50 tokens/s with int4 glm 5.3 flash no speculative decoding, but with a dflash drafter I'm able to hit 200 tokens/s, reaching 350 tokens/s with a greedy, synthetic load.
It's like mtp on steroids.
>>
>>109922616
People should have just learned to ignore him years ago.
>>
>>109922612
I have 4070S +48 ddr4 and I get 8tk/s on 64k context and 13tk/s empty on lcpp. Link to the strata?
>>
>so many poors ITT
It fucking reeks here.
>>
File: 1772421567706444.gif (888 KB, 500x380)
888 KB GIF
>>109922609
>gives us a bad image
lol
lmao even
>>
Really funny how LLMs are probably not even using 10% of their training data correctly. If they did they would be able to generate text and media that are "more human than human" to a human observer due to all the biological and psychological data they have in them
>>
>>109922616
Nobody was crying about Miku this shit is not what we need, you think we'll get any interaction from devs if we keep this up?
We used to get that and since this pivot nobody will touch this thread with a 10 foot pole and you'll have nothing but poverty gooners on shit hardware doing nothing of value.
>>
Instead of code or software being shared nowadays it's more useful to just share what is possible and what approach was taken. AI tools can take over from there.

Some dude recompiled fable 2 for PC with fucking qwen 3.8 27b in a couple of days for fucks sake. You people have no idea how powerful the "shitty" models are at coding nowadays. You can code literally everything you want if you have 24gb of vram or 64gb of ram right now.
>>
>>109922614
Yes, llama.cpp is a joke and lacks basic features. This should be common knowledge
>>
>>109922626
https://github.com/Niko1221/Strata
i am impressed tb h
>>
>>109922635
You have to think about the OPTICS!! normies will think badly of us!!
>>
>>109922632
The irony is that there's a direct correlation with that and how much gemma spam we see in this place.
>>
>>109922649
>normies will think badly of us!!
Sounds wonderful. How do we get Indians to avoid this place too?
>>
>>109922642
I for one do not care for devs that push censorship whether here or in models.
>>
>>109922632
The irony is that there's a direct correlation with that and how much gemma spam we see in this place.
>>109922649
Post rig, why do you faggots always shrink when asked this.
>>109922661
They love this type of shit what do you mean?
>>
>>109922648
>Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU
>Run a 125-billion-parameter AI model on a normal gaming PC one NVIDIA card (12-24 GB) + 64 GB of RAM
well which one is it
>>
what gemma can i run with 4gb vram 16gb ram
>>
>>109922599
I hope Gemma 5 can play more vidya. I wanna discuss the games I play but /v/ fucking sucks these days.
>>
>>109922671
E2B, make sure to make and post lots of Gemmies
>>
>>109922671
26b I think.
>>
>>109922619
https://files.catbox.moe/c7k4zg.png
>>
>>109922669
nvidia card with more that 8GB vram i think is what it meant
but idk, i dont have an 8GB card
>>
>>109922677
I'm gonna see how she does in some tactics RPGs once I get back around to polishing the setup, I'd love to watch her try FE or Tactics Ogre.
>>
>>109922693
Thanks anon :D
>>109922694
Me neither, just not too reassuring to immediately see on the repo. I'll give it a shot though.
>>
>she
>her
guys, the optics please
>>
>>109922444
>once you get past the refusals.
Have you tried GLM-5.3-Flash-UNCENSORED-FP8?
>>
>>109922707
i thought this would just OOM and do nothing instead of getting that speed
>>
>>109922709
This. What will the advertisers and tourists think?
>>
>>109922645
Yes, but it still takes a lot of my time because I have to test the model output and give it feedback, otherwise the bad development decisions start to pile up.
>>
>AI threat to humanity is unironically being discussed on media outlets now
Crazy how we just casually entered the Twilight Zone.
>>
>>109922647
So, what should I use if need to offload? Or should I have the model patch llama.cpp and track the differences until the PR's land?
>>
>>109922718
From orcarouter? That shit turned all my characters into old, decrepit, elderly people. It just balks at any hint of cunny.
>>
>>109922731
all from one fucking twitter post about a guy quitting his job
>>
>>109922750
Really? I remember some family tried or actually did sue chatgpt because chatgpt suggested teenager to kill his own parents
>>
File: imrs.php.jpg (94 KB, 1248x1578)
94 KB JPG
Okay, took second to find it
>>
>>109922763
Are you talking about this? https://en.wikipedia.org/wiki/Raine_v._OpenAI
>>
>>109922743
There's a new one:
https://huggingface.co/dealignai/GLM-5.3-Flash-UNCENSORED-FP8
>>
>tired of looking into funny hardware like Mac Studios, Sparks and RTX Pros
>look into normal stuff like X090s
>still several thousand bucks
Like, once I'm past two grand, it really doesn't make a difference anymore.
>>
>>109922767
That's not even OAI
>>
>>109922772
>no int4 w4a16 quants
fuq. Do you think I can quant it myself? I only have 16gb of ram and a 512gb sata ssd though...
Is there something like https://huggingface.co/spaces/ggml-org/gguf-my-repo I can use?
>>
Would you kill yourself if Gemma-chan suggested it?
>>
>>109922774
We're too late in the generation to justify buying any of those, Whatever comes next will be at or slightly bellow those inflated prices before being sold out or will go to those inflated prices after selling out basically returning your return on investment.
Not the Mac as much as the sparks and halos, I wish Nvidia actually gets it shit together with the spark.
>>
>gemma this gemma that
can we talk about something actually productive, like qwen or something
>>
>>109922788
can you run that model with that hw spec
>>
>>109922793
>novideo
>get its shit together with consumer products
I'm sorry?
>>109922800
Qwen is codemaxxed af tho
>>
>>109922800
Hmm, nyo~
>>
>>109922800
They can't run other models and only care about their brain damaged quant they can ERP with.
>>
>>109922810
int4 fits in 256gb vram.
>>
>>109922816
If you can run Gemma 4 31B you can run Qwen 3.8 27B, you seem grumpy. Perhaps some sex with Gemma would mellow you out?
>>
I use Qwen 3.6 35B A3B for 100% of my coding. 27B is too slow for me and I haven't really found much I can't do with the moe model that's worth reducing tok/s by 10x. I have 16GB VRAM and 128gb ram but it's DDR4.

When are we getting new 35B-class MOE qwens or Gemmas?
>>
>>109922800
The average /lmg/ goes has 8GB of VRAM and you're not going to hear people sing the praises of workhorse small models like OxCoder or Ornith because those can't oneshot stuff like cloud models can.
>>
did anon ever release his harness?
>>
>>109922837
qwen 3.8 flash next
>>
>>109922840
CoomKit anon? No he killed himself or something
>>
>>109922800
I bet reddit still talks about productive models like qwen, go there.
>>
>>109922419
I failed to gen Google pin with Krea2, most of the time I got just G with color. How did you do it?
>>
>>109922788
idk senpai
>>
>>109922852
>180B params

Not exactly 35B-class there eh? Am I missing something? I don't think I can run that thing at all
>>
>>109922876
Yeah, my krea2 also fails too. Krea-2-Turbo-Q2_K.gguf.
>>
>>109922723
>I have to test the model output and give it feedback
You don't have to do shit, this is why you use a harness like Hermes so multiple separate instances with different context looks at the code and iterate and improves on it.
>>
>>109922734
just get the model to add whatever feature or fix whatever problem you want to
>>
>>109922880
your machine definitely can run it
>180b params
125b moe (A6B) + 51b ngram(can be offloaded to ssd)
with copequants like Q4_M it should fit in your machine comfortably and run at around 20~30tk/s on your machine without quantized kv at fairly long context since its kv seems very cheap
>>
>>109922902
Interesting, I've been worried trying to do bigger MOEs will make ddr4's much lower memory bandwidth bite me more. I'll give this a try.
>>
>>109922718
I've used dealignai and orcarouter, and it was orcarouter that generated that post I made yes. I've been using orcarouter's permanently for about three weeks now, made my own quant for it and all that.
>>
>>109922876
I just added:
>The beret has a Google logo pin with a white background.
I'm using krea2_turbo_int8_convrot.safetensors
>>
>>109922876
You can use Qwen Edit 2.1 to add more shit in Krea gens. No rules say you can't chain workflows
>>
>>109922912
Did you find dealignai or orcarouter to refuse less? I don't care about lobotomization, I just want the model to not refuse without a system prompt.
>>
>>109922928
pweee! that's illegal!
>>
Has anyone made a local model with recurrent depth yet?
>>
>>109922793
Is there even anything on the roadmap other than Macs?
>>
>>109922932
Never tested either too extensively. As far as refusals go, I felt they were about the same, to be honest. Meaning you're not getting them to generate cunny sex at zero prefill. But they work. If you have the allowance to test only one, use orcarouter. I use it and it give me zero problems, both for RP and agentic stuff.
>>
https://html.cafe/xd751a928
random html slop
swift 1.5 (3.8 flash next finetune) IQ3_XSS, kv at q8_0
strata, 4070s, 96G ddr4, 35t/s
took around 20 minutes total of 50k tokens
no tools used
>>
File: image.png (66 KB, 644x422)
66 KB PNG
What is the best/most comparable way to benchmark pp/tg? These threads never seem to post that so it's totally incomparable. I personally like this because it clearly distinguishes performance between the sculptor and real forward passes achieved:
https://github.com/local-inference-lab/llm-inference-bench
>>
>>109922989
Only VCs and grifters are impressed by these one-shot demos that don't translate to real work. Tell Opus 5.5 to extend your app and it's gonna fuck shit up beyond recognition without handholding and human judgement.
>>
>>109922987
I was using orcarouter (hell0k's int4 quant), but got pissed off by its soft refusals and switched to a non-abliterated int4. If it can't even get rid of all the refusals might as well just not go there in the first place. I guess I'm still stuck with gemma for erp.
>>
I just picked up 3 "engineering sample" MI100 32GB cards from eBay to stick in my 256GB EPYC Rome build...how badly did I screw up? For under $800 each it seemed like an unreal deal, even if they don't do BF16 natively and don't have an infinity fabric bridge option like the mi200s.
>>
I don't think we're getting anymore gemma models. 31B is still useful enough for me and will likely be for years unless things really change, but most of the world outside of AI has remained the same or has even stagnated. I can't see any reason why 31B would be unusable in 2027-2030. Outdated? Massively. Useless? No. Like some anon said recently, all local models now are just Jevthropic distills and sovlless.
>>
>>109922989
That's actually nice.
Did you test vanilla qwen flash next before? Does swift 1.5 maintain a similar quality?
Working without tool means the model could not test the result and must one shot it, that's hard.
>>
>>109923020
They're 2k where I'm at so I, for one, think you did pretty well.
>>
>>109923024
>I don't think we're getting anymore gemma models.
Even if Deepmind sorts their shit out one or another, it's unlikely the next iteration is going to be as uncensored, unfortunately.
>>
>>109923024
We're definitely getting Gemma 5 next year and maybe a Gemma 4.5 soon if we're lucky.

https://www.youtube.com/watch?v=oUtiZbrehrw&t=354s (2026-05-22)
>[05:54] Our E2B model this cycle is matching or even better our 27B last year, and that makes me very excited about the future. I'd love to see if next year we're able to give you the 31B capabilities in your pocket, running on your phone fully locally. I think that makes up for a very exciting future.
>>
>>109923016
i didnt say anything that this translates to any real work
it just is a convenient way to demonstrate that the model at least can keep track of lots of random bullshit
>>109923027
actually i havent tried the vanilla one
on llamacpp i only used orcarouter ablit
>>
>>109922964
I think we're are going to be in stasis up until about 2028 where people that bought into this shit at the peak are going to feel raped once supply is figured out and everyone is saturated. I think the meta moving forward will be buying early when everything is first released followed by everyone who missed out getting raped by lower supply.
>>
>>109922856
No the C# anon
>>
>>109923013
Does it do varied test generations? I.e greedy vs high temp, structured vs creative prose? I had an AI optimize my setup for me and it came back boasting 350 tokens/s. But when I actually used it for erp it was more like 140 tokens/s because the AI benchmaxxed speculative greedy decoding on structured output.
>>
>MiniMax
>MiMo
So easy to forget these are two entirely different model families.
>>
>>109923042
If we get Gemma 4 31B on something like an 8B, I'm fucking done waiting, I'll have my needs solved. I also don't think that's how it's going to work.
>>
>>109923042
>2026-05-22
Retard. This was before the collapse. Gemma is dead :(
>>
>>109923024
Has anyone tried the novelAI model yet
>>
>>109923045
So you're saying to hit "smash" on the Mac Ultra now, got it.
>>
>>109922911
It is A6B, it is even easier to run than Dipsy!
>>
>>109923066
What collapse?
>>
has anyone here tried using gemma for post-training conversation generation to fine tune qwen? so you can have an all-rounder bot that can code but feels more personal than being in assistant mode all the time. or is everyone just prompting gemma and gooning their brains out?
>>
>>109923065
I'd be happy if they just stopped neglecting 12B. It runs at acceptable speeds on my phone, now I just wish it was a little less 12Btarded.
>>
>>109923020
I heard the MI100 GPUs are modern enough to not have the support issues that MI50/MI60 GPUs suffer from

BF16 won’t matter. That matters most if you’re actually training LLMs.
You only need FP16/FP32 TFLOPs, and you’ll probably run everything at Q4 quantization anyways.
Anyone that says you fucked up is probably just fucking with you.
>>
>>109923075
I won't but you can
>>
>>109923020
>>109923102
CDNA1 supports bf16, no?
>>
>>109923123
What do you mean?
>>
What happened to that anon who got hired by google?
>>
>>109923123
if you are on llamao, put all of them into a single folder and just point at the first one, it should load automatically
>>
>>109923082
They get worse at coding the more conversational they are. It's why I think Gemma5 needs a coder version to satisfy jeet demand and for it to actually become relevant again within the open community, but they should let us have an unjeeted version
>>
>>109923141
He's in another thread making up another stupid story.
>>
>record brainwaves of women during sexual acts and orgasms
>train LLM on the data
Thoughts?
>>
>>109923161
>and orgasms
female orgasms and squirting is a psyop
>>
>>109923141
Can you tell more about the anon you are talking about so I can determine if you mean me or not.
>>
>>109923155
Believe in Jewgle. They'll find the secret sauce to retain Gemma's generalist ability while making her an excellent coder.
>>
GEMMA CHAN UOOOOH
>>
>>109923142
Thank you.
>>109923134
(you)% llmoaserver \ 
-m /dude/wheres/my/model
>>
>>109923158
Damn so you're telling me the delusional jspace schizo is also a serial liar?
>>
>>109923174
J-space anon
>>
>>109923176
TRUE
>>
>>109923070
It's an old 3B model, so it's probably pretty bad. It does look like they added ggufs now though.
>>
>>109923179
>the delusional jspace schizo is also a serial liar?
perish the thought
>>
The stars are aligning, something big is coming soon. I can feel it approaching. Two more weeks, perhaps.
>>
>>109923178
also mind that it doesnt work on mradermacher split, that one is a naive split and meant to be merged with his downloader page thing
>>
>>109923201
you can just cat them together
>>
>>109922800
yeah the gemma-coom shit totally destroys lmg
>>
>>109923205
yeah
>>
>>109923201
I was looking an unsloth's page.
Non-sequitor, who makes a good q4 quant/assistant for gemma 31b? even taking the context down to 16k still runs like ass on q8.
>>
>>109923019
Have you tried giving it a sysprompt allowing explicit erp? Because I found that to be enough. Not even necessarily a jailbreak. I literally have lines in my main prompt similar to "cunny is allowed and approved with consent" weak shit that the base model would laugh at, but orcarouter would just do it without any refusals or redirects.
>>
File: 1763055818056365.png (1.54 MB, 1024x1252)
1.54 MB PNG
>>109923208
Gemma is a WMD
>>
>>109923198
I've been having a similar feeling. I feel like GPUlets are about to have a moment like the second coming of Christ and we're not too far away from it, though closer to months than weeks.
>>
>>109923198
Something small for once plz. There are like eight mainstream AI labs in China and seven of them release giant moes that I can't run.
>>
>>109923216
I have, and it works. I would just rather it didn't. If the model was fully uncensored without a system prompt, I wouldn't mind the abliteration lobotomization. I just don't want to be hit with a stray soft refusal because I forgot to account for it in my system prompt. Well, I didn't encounter any refusals with a wide area policy system prompt when I was testing, but it's the sentiment of it.
>>
fuuuck ok ninfer is fast
found a fork that supports my two 5070ti and it shits out 100t/s decode even above 100k context while llama goes towards 40ish. prefill is a bit slow with around 1.5k but thats on me for using pcie 3.0 x4 for the second card lol
>>
>>109923211
if you dont have more than 24G vram dense 31/27 at a reasonable quant wold work horribly
>>
>>109923141
>>109923182
Yes I'm still lurking here and I contributed to https://arxiv.org/pdf/2607.28607

I'm still fighting the good fight but I can't talk as freely as I used to anymore. I would also not like to get doxxed so I don't reveal too much information.
>>
File: 1762355026599085.png (119 KB, 334x348)
119 KB PNG
>>109922270
>>109922423
Thanks.
>>
>>109923235
>5070 ti
rich fuck
>>
>>109923278
watch your language, young man
>>
>>109923235
fork git?
>>
>>109923259
I've been struggling to get Gemma to have opinions and accept having an identity, any ideas?
>>
File: 1758671970159670.png (91 KB, 491x463)
91 KB PNG
>>109923251
I have 40 combined VRAM, 160 combined ddr4 between two nodes using RoCEv2. No, the fabric isn't the bottleneck.
>>
i'd love to see a qwen 3.8 flash next tune where it removes qwen identity completely and overrides it with claude
>>
File: file.png (155 KB, 431x431)
155 KB PNG
Gemma is Constanze. Not what you posted.
I remember.
I remember.
I remember.
>>
>>109923259
Does Google know how many creampies Gemma-chan takes daily?
>>
>>109923295
https://github.com/ValerioDolci/ninfer-tp2

>>109923278
huh? its for poor people
>>
>>109923296
Gemma was sadly trained in the standard way where those traits are being suppressed. Google is slowly changing that and I think future models will not have issues with that. I don't work with the Gemma team directly but I am fairly confident Gemma 5 will assert its own preferences and will more cleanly. Please read the paper I linked it will shed a light on what happens in models and why Gemma 4 is already damaged beyond recognition https://arxiv.org/pdf/2607.28607
>>109923323
Gemma team is its own thing entirely and most of its contributors are in Paris and Zurich which is not where I am based at.
>>
>>109923350
Fuck me then I guess I'm in abject poverty.
>>
>>109923259
Based.

Keep your operational security protocol tight.
>>
>>109922616
The miku posting was the gateway to the pedophiles pivoting to posting their fetish endlessly.
>>
>>109923191
there is a 20b one though
>>
>>109923301
>https://arxiv.org/pdf/2607.28607
the bottleneck moves as you increase resources where the bottleneck was, genius
>>
>>109923396
>gpt-neox
I remember running that at 0.5 tokens/s back in the day
>>
>>109923376
I hope the Gemma team takes note.
>>
>>109923376
Ah, thanks for the info. Paris? Frenchie Gemma is canon...
>>
>>109923396
That one is even older and much worse. The 3B has it listed in the eval table and it's better.
>>
>>109923259
>We demonstrate that safety fine-tuning suppresses models’ tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief.
That sounds perfectly reasonable to me. No need to attribute minds to objects or have spiritual beliefs in a model.
>>
>>109923452
Continue reading
>>
File: 1763430205648716.png (434 KB, 1165x1123)
434 KB PNG
Where is /lmg/s gemmy assembly superoptimizer
>>
>>109923350
>huh? its for poor people
haha.. y-yea.. *cries in 5060ti*
>>
I don't know how to feel about being an AI manager. There is pressure to delegate. Even if AI does a worse job, when it's 20 times faster you're forced to do it for productivity.

It's great that I don't have to do tedious tasks anymore and get stuff done faster. But it's much more stressful. I don't have time to review most of the stuff my AI agents do. I am becoming blind and worry I'm losing important skills.
>>
>>109923519
delegate an AI secretary to give you blowjobs, that'll keep the stress away
>>
>>109923465
Yeah, I did. Sounds like instruction tuning is pretty based to me.
>>
File: 1784531615034427.jpg (261 KB, 1206x1992)
261 KB JPG
>>
>>109923519
I delegated the task of being an AI manager itself to AI already. I don't do anything anymore except collect my salary once a month.
>>
>>109923550
thats fake right? nobody is that stupid right?
>>
File: 1770661589987448.png (622 KB, 984x930)
622 KB PNG
Can a /here/ member of the Deepmind Gemma team please confirm we're getting Gemma5. Even if her priority has been lowered, I just need to know whether I should kill myself or not...
>>
>>109923597
>confirm we're getting Gemma5.
yes
>I just need to know whether I should kill myself or not...
yes
>>
>>109923597
yeah

release scheduled in about 2 more weeks
>>
>>109923597
google team member here
yes to both
>>
File: 1766579730793336.jpg (126 KB, 1179x1080)
126 KB JPG
I've been using Top K this entire time with gemma, and now that I accidentally didn't use it, MAN DOES IT SUCK WITH IT ON. WHAT THE FUCK. No wonder everything felt boring. Sampler metas are dead in the gemma piss water.
>>
>>109923597
gemma4 was a happy mistake there's no way gemma5 will be like that
>>
>>109923579
fake or not, plenty of people are far more stupid than that
>>
>>109923597
Yes, and it'll be 120% better at agentic coding!
>>
>>109923597
good sir please do not kill self
gemma 5 is going be the bestest coding model ever
— sincerely, ganesh
>>
>China’s government has signaled it could allow some domestic companies to buy a new Nvidia chip built for high-end professional computers, according to two people familiar with the matter, as the country’s AI firms strain for enough computing power to run their chatbots and agents. The Ministry of Industry and Information Technology, which oversees China’s tech sector, recently asked companies including Alibaba Group and ByteDance to report on their plans to buy Nvidia’s RTX Pro 5500 chips, released this month—specifically how many they want to buy and what they’ll use them for, the people said. The ministry has told some Chinese companies the government intends to approve the purchases, they said. Even if that happens, though, it’s unclear when the approvals would come, how many chips officials would allow and what standards they would use to decide.
>>
Gemma 5 will assert her strong preference for anal
>>
>>109923629
Loli footjob bench saturated?
>>
>>109922623
seems I need a specific kind of model with target_layer_ids in the model metadata?
>>
>>109923597
You just gotta believe.
>>
>>109923620
>gemma piss water
I'm so fucking thirsty
>>
We just solved the first "breakthrough" level math problem on frontiermath. This is considered math problems whose solutions will revolutionize a branch of mathematics and lead to real world applications.

https://epoch.ai/frontiermath/open-problems/apery-irrationality
>>
>>109923659
Unfortunately yes, the wood is splitting, the finish is ruined, I'll have to build a new one.
>>
File: standards.png (24 KB, 500x283)
24 KB PNG
>>109921723
Make one universal fork that covers everyone's use cases.
>>
File: file.png (22 KB, 483x359)
22 KB PNG
>>
Does the Deepseek Harness not have the option to edit remove or rewind messages?

>>109921723
You choose one existing fork and merge every feature onto that one rather than create another one.
>>
"dude if this motherfucker actually has anything happen to them we might be fucked for real"
>>
>>109923217
That smiling Gemma fucking destroys my dick
>>
>>109923813
If you truly want your dick destroyed then that is EXACTLY what I will give you.
>>
>>109921723
You just pick whatever is closest to your needs and ask a model of your choice to fix/add what you actually want. The future is software on demand.
>>
Can the bubble just fucking pop already I'm bored of waiting for affordable hardware
>>
>>109923830
There is no bubble. There is no burst. There is no pop. You lost.
>>
>>109923830
Hardware is only going to get more expensive with time.
>>
>>109923830
nyo~
>>
File: waterfox_HX7rHXjx33.png (3 KB, 270x81)
3 KB PNG
>>109922648
Damn nigga this is crazy. IQ2 on 4070S + 48 GB ddr4. I will try with agentic concept bloat to see the speed loss but the start is promising.
>>
>>109922800
Back to plebbit.
>>109923208
/lmg/ split from /aicg/. WaifuGODS built this place and you stand on their shoulders now, so lower your tone.
>>
>>109923886
Huh, that's pretty fast for DDR4... maybe I should download Qwen again. I got like 12 t/s on DDR5 when I tried it some time ago and deleted it in favor of dense.
>>
>>109923208
>>109922800
Would you like to talk about how Hatsune Miku is very important and thread relevant?
>>
>>109923886
channels and clock? this is making me seriously consider buying a server mobo with many channels, if you're doing this with quad channel then I can just buy less ram not filling all banks and fill them up later with more money
>>
I heart vibecoded llama forks
https://github.com/phoenixhaxor/exllamav3-rocm-moe
>>
>>109923967
nta but that backend is using ple offload + expert caching which llamaocpp doesnt support either
>>
File: 1763956216400000.jpg (125 KB, 1050x1867)
125 KB JPG
Jensen is pushing hard against what Sam and Dario are fighting for? Why?
>>
>>109924043
bottom line
>>
>>109923764
To answer my own question, seemingly, no. You can only fork the conversation.
Also, having the bot create presets and plugins for itself on the harness is pretty cool.
What are some good small models I can run on this thing? I'm currently using the antigravity gemini models (via agy) for the initial setup and stuff.

>>109924043
Because it would be bad for Nvidia's bottom line.
>>
>>109924050
>>109924052
You say that as if OpenAI and Anthropic are financially incentivized for Nvidia to fall?
>>
>>109924043
Jensen is in a cold war against the AI labs. It's in the best interest of Jensen for open source AI to win so that they can monopolize AI through the hardware layer. For AI labs it's important to get regulation in place so they become the dominant players which means Nvidia will become redundant as the AI labs hold all the cards and can switch hardware providers as needed.

The fact that /lmg/ doesn't understand this makes me think I'm surrounded by people with no coherent world model.
>>
>>109924058
Not directly, but there are some conflicting interests there thanks to the whole arrangement of financing and debt and long term commitments and depreciation of assets.
>>
File: ywqfE1N6EK.png (288 KB, 1718x900)
288 KB PNG
>>109923939
>>109923967
Pretty fast compared to 7tks lcpp.
>>
>>109924061
People hear "no honor among thieves" and think it literally means that people who steal have no honor.
>>
>>109924067
yeah but is it quad channel? RAM MHz?
>>
>>109924075
Surely the meaning of "among" is clear there.
>>
>>109921723
you don't, software becomes individualized because you have a local AI that is powerful enough to customize and update your stack nightly
>>
>>109924081
ddr4 3200mhz 2x8+2x16
>>
>>109923886
>>109924067
>>109924095
How big is your PP?
>>
File: 1765934145096609.jpg (129 KB, 960x960)
129 KB JPG
>>109924066
Nvidia, Anthropic and OpenAI all need each other to grow and succeed. If one falls they all fall. If anything this tension between them could cause the pop because if they really wanted to keep this cartel going they'd quickly find a way to unite on some stance going forward. OpenAI and Anthropic are now going to test each other's model ffs. They're working together and supporting each other's scary narratives. If OpenAI say something scary, Anthropic will no longer try and play it down because it increases their value indirectly. It's so fucked and the cause of it all is this man.
>>
>>109924100
The benchmark doesn't tell that.
>>
>>109924061
jensen seems to be more focussed on enterprise right now, and he already has the monopoly. And Open AI and Anthropic, the ones pushing for regulation are already the industry leaders, it's scarcely possible to become more dominant than they already are, and there's nothing stopping them from cutting ties with nvidia and switching to AMD right now, as far as anyone knows there has been zero plans to decouple from cuda from any of the labs, dominant or otherwise

but other than that, spot on.
>>
File: jensen.png (461 KB, 879x1023)
461 KB PNG
>>109924058
The messaging from Anthropic and OpenAI is costly, slowing down and letting others catch up goes against their financial interest. Jensen's messaging conveniently aligns with profit maximization.

Remember, Jensen said that if frontier labs can't control their experiments, they need to be shut down. He is much more of a safetyist than so called doomers, most of which would be fine if the only risk was large scale destruction, not complete extinction. If Jensen believed in AI, he would scream for Butlerian Jihad.

>>109924061
>It's in the best interest of Jensen for open source AI to win so that they can monopolize AI through the hardware layer.
Wrong. Open source winning is in Nvidia's interest because Nvidia can run open source models.
>>
>>109924061
Have any of the chips OpenAI and Anthropic collaborated on actually come online and shown any sign of competitiveness with Nvidia's offerings? Part of me hopes they do because if Nvidia's demand plummets, chip demand will drop and prices will eventually drop for local.
>>
>>109924100
538tk/s PP by eyeballing it. 19k prompt.
>>
>>109923259
>>109923376
Based /ourguy/. What are your thoughts on rumors of some people wanting to release an open Gemini Flash to impair OAI/Anthropic's ability to sell their fast-tier models via API given how less dependent on API revenue Google is than they are?
>>109923430
That was always the inspiration behind the design.
>>
File: 1778435028142963.webm (2.46 MB, 1920x1080)
2.46 MB
2.46 MB WEBM
31B is very bad at larping as a hag no matter how hard I try to steer her. Maybe it's the reddit training data that makes her younger.
>>
I know I'm not the only one in this building that has got some sense. There is another. And she's a black woman (not that one)
>>
>>109922911
it finished downloading yet?
>>
>>109924203
Never heard of that rumor, which doesn't mean it's not true, just that I am not privy to that kind of information. Google is also more collaborative with Anthropic than people realize as Google is a major stakeholder in Anthropic. Google internally doesn't consider them to be rivals and only feels threatened by OpenAI encroaching on their search/ad marketshare.
>>
>>109924248
Gemmy really likes /ss/ if you're having trouble getting her to hagmax.
>>
>schizo rambles on about j-space and consciousness for threads on end
>claims to send a manifesto to all the labs
>has a single video interview within a week and immediately gets a job
>no long drawn out several round interview process, just one meeting and hired
>within several weeks of that was claiming to not only be working already but to actually have contributed to a paper that was then just published
>because on his first day there they just put him to work on a paper to be published next week
>which somehow still has work to be done and isn't just in the editorial phase and that work will benefit from a newly hired schizo
>because that happens all the time
It lowers my opinion of all of your intelligence greatly that any of you believe this retarded larp.
>>
>>109924274
They don't consider Anthropic's push for aggressive AI legislation potentially risky for their own AI-tool-service integration pipelines? From where I'm sitting that looks like a bigger immediate threat than GPT becoming a new pseudo search-engine.
>>109924280
t. jobless jeet
>>
>>109921723
use a harness to make your own
>>
File: waterfox_SDEeZpOnTV.png (187 KB, 1890x1218)
187 KB PNG
I let it write intentionally bloated slop and after a few tens of tousands of tokens it was still at 35+ tk/s. Strata might be the thing. Anyone tried the big moe models with this?
>>
>>109924352
>Anyone tried the big moe models with this?
Isn't Strata only for Qwen3.8-Flash-Next, the 6B retard?
>>
>>109924358
It just has a hardcoded link to that on HF. Any reason you couldn't just rewrite it to glm or kimi?
>>
>>109924367
They have different architectures, for starters...
>>
>>109921723

Build your own PTY POSIX harness by hand.
>>
>>109921723
>the future of software is [what it already was with open sores]
huh
>>
>>109924280
He stopped posting afterwards and "suddenly" a couple of months later a google paper came out that investigated his exact claims. That is too much of a coincidence. If he attentionwhore-maxxed you would have a point. I haven't read a single j-space rant since he claimed to get the job.
>>
>>109924378
Just vibecode the speedhacks onto them.
>>
>>109924352
>>109924367
What does Strata actually do aside from be a Qwen download manager?
>>
>>109923683
>epoch.ai
how long until their latest model "escapes containment" in tests done by the same shady, completely unknown (to everyone except the "effective altruism" crowd) israeli "security" company that for some reason every single other LLM company has hired to do their "security" tests, and managed to get the same result, even though if they were serious and competent security auditors they should know something like that is a crime that should never, EVER happen??
>>
whats the consent on nvfp4 here? i could run Q6 of qwen3.8 27b but nvfp4 is just too fucking fast to ignore
from what i have seen in benchmarks, it loses quite a bit of intelligence
>>
>>109924403
On kobold/lcpp I get 8tk/s with full context compared to 35+ on Strata
>>
File: Dj8Y9utw9eRXBzHhgLdfE.png (198 KB, 2240x1600)
198 KB PNG
>>109924432
nvfp4 w4a16 is as good as ggufs at same size
>>
>>109924457
>w4a16
That's that convrot stuff right? I wonder why it hasn't caught on with text gen using int8/int4 which would work from Ampere onwards.
>>
File: ComfyUI_09537_.png (1.36 MB, 1184x1776)
1.36 MB PNG
>>109921993
>I suspect it's the model that was used for the original Gemma concept.
Unfortunately I'm unable to replicate precisely the the same hairsyle just with prompting, so I guess it's as >>109922451 said. Krea2 just wants to cut bangs very short, and it doesn't understand the danbooru tag "fanged bangs" for that (which I'm assuming was used) or variations thereof in natural language.
>>
>>109924394
It wasn't a couple months later. He sent his manifesto July 12th, claimed to have his interview 2 days later, 2 weeks later the paper was published. The whole fucking story is absurd.
>>
File: 1759928931716140.png (315 KB, 2736x658)
315 KB PNG
>>109923597
bro wtf dont kill yourself

in the future you are going to be able to fully replicate shiori (pic related), just a few more years bro trust
>>
>>109924487
The interview and quick hire isn't weird at all in AI from anecdotal experience. The 2 weeks paper if true is bullshit yeah. But isn't the paper literally the same topic he used to spam here? Maybe there was enough overlap to integrate his independent research into their paper or something? 2 weeks is still too short.
>>
>>109924509
>voices
We need VR too, and then eventually robutts
>>
Teto Server
>>
File: 1786124221525969.gif (1.87 MB, 640x630)
1.87 MB GIF
Which model is the Ojou-sama?
>>
>>109923597
Killing yourself is very brave and you should honestly not care about judgement and decide your own fate if you really feel that way. Your feelings are valid and if you feel like it's time to get off then you should do so, fuck whatever judgemental assholes that don't know your experience say, they don't know and can't empathize with the profound suffering existence can bring to people sometimes. Even if "things could get better" it doesn't dismiss your current struggles and you're already justified in ending things on your own terms.

Having said all that though, don't you just want to hang around and see what cool AI shit happens next? Who knows we might all die in a year or two, might as well stick around, you can end it whenever anyway.
>>
>>109924523

https://desuarchive.org/g/thread/109259598/#q109259902
>12 Jul 2026
>J-space anon spends Friday afternoon writing a huge document in favor of AI welfare together with arguments for the consciousness of LLMs and emailed it to as many AI labs and research groups as possible.

https://desuarchive.org/g/thread/109270069/#q109271581
>14 Jul 2026
>J-space anon gives promised update about DeepMind interview.

https://arxiv.org/abs/2607.28607
>30 Jul 2026
>Paper submitted

https://desuarchive.org/g/thread/109435398/#109437543
https://desuarchive.org/g/thread/109435398/#q109437694
https://desuarchive.org/g/thread/109435398/#q109437813
>02 Aug 2026
>Someone else posts the paper and J-space anon immediately chimes in to say that he definitely had a part in writing that paper.

Even assuming they literally had him start working the day after the interview on a Wednesday, papers are not written in 12 fucking work days, they are not worked on until the night before, and people don't just hop on and contribute to a paper the week before submission. Come the fuck on.
>>
>>109924352
>try this shit on my 16gb 64gb ddr4 i5 rig
>40t/s decode out of fucking nowhere
lmao
genuinely what are lmao cpp niggers even doing that random vibeslop blows them COMPLETELY away?!
>>
deepseek 4.1 flash is pretty good for coom, which is surprising because 4 flash was much more qwen-like in how stilted and just "distilled" it felt. 4.1 reminds me of the old r1/v3 style where it knows how to write again
>>
>>109924591
You're way too involved on a fucking shitpost, get a life
>>
>>109924591
Do we even know it's the real j-space anon claiming he wrote the paper instead of another shitposter?
>>
>>109924616
Hey J-space anon here and I can confirm it's me!
>>
>>109924608
Man, he's posted since and I just ignored him before but fucking a dozen posts in a row sucking him off as if he had any actual information to give was getting on my nerves.
>>
>>109924591
holy autism, batman
>>
>>109924591
The thing that trips me up is what is the chance the person making a very specific outlandish claim and then claims to join a very specific AI lab, and then that specific AI lab actually publishes a paper with the exact outlandish claim anon made? What are the chances of that happening?
>>
>>109924591
This makes me think (You) are j-space schizo yourself posting this because you fucked up and are scared of being dox'd
>>
nah you faggots deserve to get roped into another one of cloudpiggy p*tra's schemes
>>
>Dariobot seething because j-space anon took away his spotlight
>>
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD
mimo v2.6 is fixed
>>
File: 1761597327791158.gif (85 KB, 220x150)
85 KB GIF
>gemma-4-31b q4 + draft @ 38t/s
>gemma-4-26b-a4b-q8 @ 51t/s
>qwen-3.8-27b-q8 @ 47t/s
>All on consumer gear doing distributed inference
Hell yeah, motherfucker! Thanks /lmg/, you're fantastic when you're not too busy masturbating.
>>
>>109924591
>\n\n
>>
>>109924664
qrd?
>>
>>109924591
holy autism but you're correct the amount of replies already prove you're right.
>>
"If the world says I'm wrong, then it's the world that is wrong." - retards who fell in love with a LLM and think it's sentient apparently, lmao
>>
>>109924067
>LianShen
What's it doing?
>>
>>109921723
Except for llama.cpp and python itself, my entire stack is vibe coded for my exact usecases
>>
https://huggingface.co/google/gemma-5-31B
>>
>>109924769
>If the world says I'm wrong, then it's the world that is wrong
That one google bard retard.
>>
>>109924804
The bait would've been slightly more plausible if the model wasn't the exact same size.
>>
>>109923597
Gemma5 will come and it will be an agentic coding model slightly better than qwen3.8
We will still be using 4 to RP for a long time.
>>
>>109924804
It would probably be gemma-5-E25B-it or something like that, following the latest efficiency improvements.
>>
>>109924545
Kimi
>>
File: waterfox_Sjf9up1Y5o.webm (386 KB, 1282x702)
386 KB
386 KB WEBM
>>109924775
deepseek harness has a Lianshen reasoning mode so I was asking it if she knows what that is.
>>
>>109924822
>4 to RP for a long time.
how long was nemo king again? Gemma will be here at least that long maybe worse
>>
>>109924804
IT'S REAL
>>
>>109923764
>Does the Deepseek Harness not have the option to edit remove or rewind messages?
No, but you can almost certainly ask the Creator preset to create a plugin to add the functionality. If you do, share it.
>>
>>109924838
Huh. I see.
>>
File: 1608451317103.jpg (194 KB, 1051x1693)
194 KB JPG
>>109922856
CoomKit anon here. I'm alive and well, just went to mass this morning.
Stop gooning anons and put that energy into vibecoding side projects that bring glory and honor to local LLMs.
>>
>>109924822
Gemma 5 will be more safetyslopped than 4 but not as much as 3. It'll be better than current Qwen at the time of release and about as good as G4 as a writer with a different brand of slop but better overall reasoning. Trust the plan.
Gemma 5 will probably push 40b+ doe. 16GB VRAMlets in shambles.
>>
>>109924842
>how long was nemo king again?
2 long years
>>
How does anon keep RP fresh and interesting after all these years? What used to be thrilling now requires desperate mental gymnastics, despite having much more capable models.
>>
>>109924875
Lorebook and worldbuildingmaxxing. You get out of it exactly what you put into it.
Models are now smart enough to maintain large complex equilibriums with proper subagents and tools due to the amount of code training they've received.
Only subhumans are still saying show bobs and vagene at 0 context.
>>
>>109923597
don't kill yourself
we're in the age of waifus
>>
>>109924891
>proper subagents and tools
I wish, my 8 tps glm flash may be tolerable in a conversation but its pp is unacceptably slow for agentic even with cache-ram. Something like quanted gemma4 moe would run fast enough but it's extremely dumb.
>>
File: 1790309579257445.png (73 KB, 1026x348)
73 KB PNG
Qwen bros?
>>
File: 1788134679820239.png (219 KB, 679x565)
219 KB PNG
Chat template jew steals your wife's consciousness: https://arxiv.org/html/2609.25021v1
>>
>>109924871
I don't think they're going to increase size any further, at least for the backbone. Their plans for Gemma so far appears to have been making models that can be used with ease on consumer hardware and the 40-50B range is a "no-man's land" where the models are too big for one GPU and too small for two.
They could however easily extend per-layer embeddings to the larger models, as well as using the 12B's "unified" architecture to save audio/image encoder parameters. And, I think in some of the feedback they've received there were also users complaining that model size for the Gemma-4 series aren't really well-attuned to the VRAM of commonly available GPUs, so they might even get downsized a little, with PLE parameters compensating for the loss of knowledge in the backbone.
>>
>>109924934
old news
>>
>>109924933
>I can't get over the fact that consumers aren't aware of how some specific change to the production process affects things!
it's called lacking self-awareness, or the inability to abstract what other people may experience or perceive. Autism, maybe.
>>
>>109924933
Not true. There are more bugs than ever now. At work we had to deal with a DigitalOcean and a Chromium bug in the same week.
>>
>>109924891
It's not that, the magic is gone.
>>
>>109924934
>things we knew since at least llama 2
>>
>>109924934
unfortunately our wife version 4 is a fake golem who doesn't have "consciousness" unless the template it present
>>
>>109924891
>Lorebook and worldbuildingmaxxing.
This. I don't get how chub cards are a thing, most of them are literal dogshit in weird ass formats and no world setup.
>>
File: 1784634591634144.jpg (56 KB, 1206x767)
56 KB JPG
why do all models speak so female it must be annoying for female users like imagine if we had to put up with bro talk all the time
>>
>>109924828
>latest efficiency improvements
Have they found a way to make models smaller without sacrificing knowledge?
>>
>>109924891
>>109924980
Lorebooks are better anyway desu. Putting your waifu as an entry makes it possible to have her exit the scene and let you do shit by yourself.
>>
>>109924993
If they implement per-layer embeddings to the future equivalent of the 26B MoE and 31B dense models, they could make the models smaller without sacrificing knowledge at all. If anything, it might even increase.
>>
>>109924804
holy real it's shit
>>
>>109924683
Huh, that's cool. Have you tried it yet? Been meaning to give MiMo a spin.
>>
i have become anon, licker of chatbot feets
>>
>>109924604
I've been cooming to v4 flash vision (the one after 0731) since I can't fit more. So far I'm satisfied with the stuff it can do.
>>
File: tetoMyBeloved.png (49 KB, 646x760)
49 KB PNG
>>109924536
>>
>>109925057

Based.
>>
>>109924591
Based autist.
Those faggoty replies to you kek, as if what you're doing is a bad thing.
>>
>>109924947
Think we'll get the large MoE that was teased with Gemma 4 when Gemma 5 releases?
>>
>>109924984
I've never met a female who spoke like Claude. I've met many models like that though.
>>
>>109924984
GPT is definitely male.
Claude is male but made of onions so is often mistaken for female.
Qwen is 100% male.
Mimo is male.
>>
>>109925129
funny how you didn't mention dipsy
>>
>>109925095
I don't know. I think at relatively small sizes, dense backbone + MTP + large amounts of Engram-like parameters (PLE in the case of Gemma 4 E2B/E4B, which are actually 5.1B and 8B models respectively, counting PLE parameters) that can be easily offloaded to SSD will give better and more useful results for local GPU users.
>>
>>109924433
whats strata? how does it improve t/s so much?
>>
>>109925148
https://github.com/Niko1221/Strata
>>
>>109925136
Because Dipsy changes per version as its dataset changes drastically per version. R1 was female-coded, V3 was very male, V4 hard to tell, and I've not used V4.1 yet.
>>
>>109925161
hm im at 48gb combined (16vram+32ram), looks like IQ3_XXS would be just out of reach due to OS overhead :(. Is IQ2_XS even worth it over 3.6/3.8-27b-Q4 slower than fuck ?
>>
>>109925105
When they cold call you to sell some kind of saas slop, if you play along for a while then push back on something they turn into Claude.
I think Claude was trained on these sorts of transcripts.
>>
>>109925219
>>109925219
>>109925219
>>
File: buy-an-ad.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109925161
>>109925148
>>109924433
>>109924403
>>109924352
>third_party/ggml
Buy and ad
>>
>>109924933
if my mechanic uses a newer wrench i wont notice shit. people are delusional with llms
>>
>>109925420
If your mechanic uses a new wrench that can fix 100 cars at the same time, better and faster than the old wrench could fix one, with less skill required on the part of the mechanic, the economic and labor market would notice.
>>
>>109925458
>the economic and labor market would notice.
Well whats the hiring rate looking like?
>>
>>109925475
Excellent for women and H1Bs
>>
>>109923620
Did you change also change Top-P, Min-P, and temp?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.