[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1764236138963550.png (2.15 MB, 1172x1342)
2.15 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109666410 & >>109662888

►News
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742: https://github.com/ggml-org/llama.cpp/pull/27742
>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794
>(08/27) NVidia buys HuggingFace: https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition
>(08/26) GLM-5.3-Flash released with 320B-A18B and native multimodality: https://z.ai/blog/glm-5.3-flash

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png (embed)

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
mikutroons are troons and gay
>>
File: HHoxyxBagAAXtGo.jpg (263 KB, 1680x2048)
263 KB JPG
>>
>>109670450
What if I liked miku since it was not cool?
>>
>>109670412
Rather the jeet that you pdf files
>>
inb4 schizo ranting
>>
>>109670471
You hobby got subverted by troons. Like warhammer.
>>
>>109670461
Oh my god, it is Miku!
>>
>>109670473
>pdf files
please stop it it wasn't funny the first 200 times.
>>
>>109670473
You can always go back
>>
File: fumo_gemma.png (2.12 MB, 1254x1254)
2.12 MB PNG
►Recent Highlights from the Previous Thread: >>109666410

--Paper: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance:
>109667256 >109667264 >109667273 >109667844
--Tencent's Hy4 preview benchmarks and hardware requirements:
>109668492 >109668514 >109669937 >109668519 >109668551 >109668609
--GLM-5.3-Flash quantization performance and benchmark results:
>109666525 >109666554 >109666617 >109666631 >109666705 >109666710 >109666728 >109666719 >109667773
--Qwen Next and 27B performance on 24GB VRAM hardware:
>109666480 >109666484 >109666501 >109666546 >109667952
--Steering Gemma 4 to bypass assistant-tuned refusals for ERP:
>109667648 >109667654 >109667675 >109667707 >109668090 >109668096 >109667739 >109667746 >109667752 >109667768 >109668055
--Troubleshooting Gemma-4 GGUF chat templates and quantization sources:
>109667134 >109667145 >109667237 >109667229 >109667431 >109667730
--Frustrations with Jinja templates for chat completion in llama.cpp:
>109669318 >109669331 >109669418 >109669346
--Guiding LLMs for better code structuring and implementation:
>109668610 >109668749 >109668867 >109669306
--Hardware and performance requirements for high-context Qwen 3.8 Flash:
>109669207 >109669229 >109669243 >109669552 >109669761 >109670100
--Techniques for bypassing AI refusals to assist in piracy:
>109668849 >109668895 >109669015 >109669023 >109670041 >109670092 >109670129
--Qwen3.8 Flash Next performance and optimizing MoE offloading:
>109668793 >109669005 >109669682 >109669747 >109670035 >109670065 >109670126
--Methods for inducing vulgarity in Gemma 4 31b:
>109667297 >109667305 >109667330 >109667337 >109667511 >109667782 >109667834 >109669520 >109669864 >109669932 >109669941
--Logs:
>109667134 >109668849 >109669015 >109669414
--Miku (free space):
>109666931 >109669707 >109667269 >109668276 >109668466

►Recent Highlight Posts from the Previous Thread: >>109666416

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1784257584408043.png (205 KB, 292x392)
205 KB PNG
>Ilya Sutskever
>Canadian-Israeli AI researcher

Remember when this disgusting ratoid cucktsever piece of shit said publicly years ago
>"open source will always be incredibly behind proprietary models and that gap may increase with time"
https://www.youtube.com/watch?v=N36wtDYK8kI

Remember when he was memed into supporting the coup against saltman and then he was left to hang when others retracted it like a brainlet he is?

I'm glad he has proven himself to be the braindead fucking retard I called him out to be right away.
>>
>>109670506
Is this Israeli the one making canadian AIs?
That's pretty annoying.
I wanted actual canadian AI like north mini
>>
>>109670475
Fuck, I didn't make it in time. This guy is fucking quickk!!!!
>>109670473
>>
>>109670506
Sam and Dario have AGI internally.
Open source lost.
>>
Have any of you tried to make your model read a Warhammer 40k book?
What does it say about Horus?
>>
File: file.png (114 KB, 1473x774)
114 KB PNG
>Please rewrite your PR to match the shit we did for our unslop quants
>>
>>109670583
AGI?
More like abunchof gay investors!
>>
https://huggingface.co/zai-org/GLM-5.3
Finally out
>>
What's the smallest model that can set up its own environment? I tried deepseek v4 flash ablit at q4 but after a week at 2 token/s it's still struggling to send images to its chat. Do I need to buy some ssds to try streaming >>109670617? I'd rather an ablit because I'd hate to set it off then come back home to realise it refused to continue because of loli snuff or some other innocuous shit.
>>
bros...!!!!
>>
Qwen3.8 Flash Next is as close to AGI as vramlet local usecase can get. One can wonder why it isn't called Qwen4 since it has brand new architecture.
>>
>>109670604
haha
>>
>>109670646
the next version certainly will be this is a preview, to get the ecosystem ready
>>
>>109670604
Thanks AI to not have to deal with these slurpers anymore.
>>
>>109670646
-Next models are usually preview. More cool stuff to come we can expect.
>>
which is better step 3.5 or 3.7?
>>
>>109670602
It probably knows something on its own already, at least Gemma does probably.
>>
>>109670667
Gemma. Next one over that is newest qwen. And after that Hy 3 and glm 4.6 4.7.

Speaking about sex of course.
>>
>>109670677
Newest Qwen over 4.7 for sex? wtf
>>
>>109670583
this time for sure bro, just like with gpt2
>>
>>109670617
how big q4 will be
>>
>>109670706
>how big q4 will be
approx parameter count divided by 2
>>
>>109670698
Newest qwen is slightly better than gemma but with how long you have to wait for a response it is not worth it.

For me it is still gemma if you have nothing or upcoming 5.3 flash. Deepseek flash also can be good but not for sex.
>>
>>109670719
I'll have to try it, I'm curious. I've never liked Qwen for anything besides assistant shit.
>>
What's the SillyTavern card or simple prdeset / system prompt for Gemma chan? I want to quickly make a scene of her pooping on a schizo
>>
>>109670766
Getting really tired of pdf files
>>
How dangerous is using Gemma E4B powered agent harness to maintain my Windows pc?
>>
>>109670773
That's for frontier models like fable and sol.
>>
>>109670769
Use 5.3 turbo to vibecode a superior portable document format then
>>
>>109670773
What the other anon said. You're not in danger of deleting your system or being prompt injected, but you're in danger of some bad edit bricking something that you'll then have to use a frontier model like sol to fix anyways
>>
File: sQIRUNB.png (93 KB, 1202x822)
93 KB PNG
>>109670617
>unslop officially supported but not llama.cpp
lol
>>
Did you guys know that the Japanese style of emoticons like (‿) are called kaomoji? I learned that from glm 5.3 flash as it was looking through previous /lmg/ threads for Gemma chan because you fags didn't spoon-feed me
>>
>>109670803
People at work use kaomoji ironically for some reason.
>>
E-wastemaxxer here. I'm getting 10tps streaming the n-gram table off the disk. And that's without MTP as MTP support isn't finished yet.
for my non-rp purposes (I have a 3d wife) 10tps is sorta-usable, only for tasks I can let run for a while while I do other things, like have a life.
There's a lazy-loader improvements branch, and competing MTP PRs. If this gets to 18-20 tps or so I can replace 27B as my daily driver.
All this in 48GB ram 16gb+5gb Pascal cards.
>>
>>109670803
It's in the gemma system prompt
>>
>>109670813
What is a ngram table? I haven't compiled llama.cpp since MTP and its speedup improvement for Gemma 4 came out. Am I missing out on something?
>>
>>109670773
i use gpt 5.4- current version in yolo mode for everything for months now and nothing bad has ever come close to hppening
>>
>>109670813
LOL talking about Qwen 4, of course.
>>
>>109670794
With nvidia owning it they will make sure it will appear (after rebranding it to nemo.cpp)
>>
>>109670831
It's Qwen 3.8-next (qw3n 4 preview) big advancement. It uses a 50GB table lookup to do some math that would normally be a GPU load (you can tell I barely understand this shit yet) so you can run a huge model with a smaller GPU.
>>
>>109670862
basically an ultra large MTP
>>
>>109670711
i need more dedotated wam
>>
>>109670862
Good thing I stockpiled Pcie5 ssds
>>
I tried Qwen3.8-Flash-Next q3 quant, I got ~10 tok/s
system: ryzen 7700, 64gb ddr5 ram, 5060ti and 4060ti
both gpus were pretty much idle watching youtube levels of power use, cpu was at power limit
>>
https://huggingface.co/zai-org/GLM-5.3
https://x.com/Zai_org/status/2093354097122455713
https://z.ai/blog/glm-5.3
>>
>>109670766
https://rentry.org/gemma-chan
>>
>>109670896
I'm on a 2016 vintage Dell server, pcie4 and a wd blue.
>>
>>109670876
>>109670862
Okay thanks, interesting.
I'm happy with llama.cpp performance for my setup, haven't bothered even checking out their github in a while.
>>
File: oardefault.jpg (54 KB, 405x720)
54 KB JPG
>>109670906
t-thanks
>>
>>109670922
Looks like boiled garloid slices, it's a local delicacy I suppose.
>>
>>109670412
Thank you for baking and saving us from another jeet thread.
>>
>>109670906
https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE
The better they get, the more they move away from being truly open.
>>
>>109670903
The whole point of the architecture is the CCCP fucking Nvidia by making the GPU redundant. The whole idea of open models is to limit the American Hegemony. After Qwen 3.6, word was the Alibaba management was pushing back on releasing open models. There was an "internal discussion" about it but in the end the Party reminded Joseph Tsai what they did to Jack Ma, and now we have open weights of Qwen4 preview.
>>
>>109670412
Where is Hy4?
>>
>>109670922
That poor slowpoke
>>
>>109670965
So easy to whack off the tail, they so slow
>>
>>109670918
Is the GPU necessary?
I have a server with 128gb of ddr4 laying around somewhere.
>>
>>109670961
I thought open models were to help out pdf files

But cool if china wants to help make AI accessible.
>>
>>109670973
How would you feel if someone cut off your tail?
>>
>>109670964
irrelevant
>>
>>109670965
they grow back dont they? I never got why it was such a big deal in the game if they grow back. TO be fair it was kinda just like a side crime of team rocket in the game it also kinda wasn't made out to be that big a deal actually, at least that is how i felt when i played it, but i think it was meant to be more of a "look how evil tr is" than they got across
>>
>>
>>109670979
off balance, lighter, maybe a little hurt or angry
>>
>>109670977
Elon has done more for them with Grok than China will ever do. He's in trouble again for training Grok on 3d pizza. Fucking dumbass. At some point a future administration will use it as an excuse to jail him / strip his citizenship.
>>
>>109670976
I predict 10tps or more. Try it and report back. Grab the daily.
>>
>>109671023
Interesting, will try.
Not sure if I have a nvme or pcie4 slot in general, will need to check.
>>
>>109670813
thats cool and all but 10tps is not really usable as a "daily driver" is it?
>>
ngrams go crazy damn
>>
getting ~14tk/s with GLM 5.3 flash (quantized to fit into 256GB RAM)
not bad. slow, but not unusably so
>>
>glm 5.3 fits in ~440 gb + some for kv cache
ngl i’d probably suck a cock for the upcoming m5 ultra with 512 gb unified memory
>>
>>109670952
And GPU price will increase even faster as inference providers are banned and companies will buy GPUs in bulk for on premise inference.
>>
Why is everyone talking about ngram. Did something happen?
>>
>>109670986
>TR
Jessie gets a redemption arc, becomes a legit trainer with some success. James is just useless trash. But James is from a Zaibatsu family, so an Anime or Manga can never show him being anything but a worthless evil shit. everyone in Japan hates the Zaibatsu just like how everyone in Korea hates the Chaebols and in Russia the Клaн and our fucking Techbros and the Sacklers and Koch family in the US.
One thing about China is they squash these fucks when they start trouble, which again is why Alibaba suddenly is all about open models again. And now they're going after Nvidia and it's awesome.
>>
File: concept77111.mp4 (3.82 MB, 1376x768)
3.82 MB
3.82 MB MP4
>>
>>109671063
With the amount of ram you have it won't matter how slow your storage is.The whole model will fit in ram.
>>
>>109671092
what's it gonna cost, like $15k? maybe i spring for it if it's $15k
much more than that and it starts getting hard to justify (i have already spent $15k up to this point....)
>>
Hag gemma is as bad as 6yo gemma, and doesn’t even chase off tourists
>>
>>109671098
It's a cache of predictions on disk.
>>
>>109671068
no, that's why it has to get to 18-20 at least before I can switch. But again, this is without MTP and a dflash head would possibly be faster. Also, the current lazy-loading code is meh. So I'm waiting and hoping.
>>
>>109671117
How would you like your gemma anon?
>>
>>109671117
hag gemma is forced by a /ldg/ schizo who broke containment
>>
>>109671136
You mean he's one of the schizos in the /ldg/ OP?
>>
>>109671140
Yes he's that pdf
>>
>>109671140
yeah
>>
>>109671086
Can give thoughts vs DSV4 flash, specifically non-code contexts? I'm trying GLM 5.3 Flash Q3KM, 8-9 t/s on 48GB + 128GB DDR4, seems dumber at >20k than dsv4 flash mxfp4 upon first tests in story completion but haven't gotten a feel for the new model yet.
https://huggingface.co/DevQuasar/zai-org.GLM-5.3-Flash-GGUF/tree/main/Q3_K_M
https://github.com/ggml-org/llama.cpp/pull/27752/
>>
>>109671140
This one >>109671144
>>
>>109671144
>>109671147
Which one?
There are two rentries
>>
>>109670813
>10tps
This is unusable for anything bit goonslop, I tried one of the bigger models on a server at work which ended up genning at 12 tk/s and there wasn’t any point outside of simple tasks like a 100 loc shell script
>>
>>109671108
OK will try and report with results.
>>
>Ufufu~ Senpai is getting into some really spicy, naughty stuff now, isn't he? A daughter controlling her own mommy with magic… how deliciously wicked! It’s so "mesusame" of that little brat to turn her mother into a mindless puppet. (◕‿◕)

weeb anons please explain what mesusame is
>>
>>109670922
garloids should not be eaten, they are meant to be cared for and their milk is a lot more nutricious and tasty than their meat anyway.
>>
File: 1778586974977232.png (25 KB, 833x111)
25 KB PNG
>>109671166
>>
>>109671136
>>109671140
this, try not to engage with the schizo or they might come over here
>>
>>109671156
i haven't tried it yet, but i'll have my claude slave download it and stress test it for me
god i love being able to control that thing from my phone (yes i am a phoneposter)
>>
>>109671181
that's why i asked for a weeb anon to explain. because the literal meaning of mesugaki is just as useless
>>
>>109671180
Fuck off libshit I’ll eat garloids if I want to, I don’t care for the sob stories and how their endless fucking screaming plugs at your heartstrings
if it’s made out of meat it goes into the pot
>>
>>109671101
my expert knowledge of japanese tell me that it means something like
slutty shark
i dont know what that means though
>>
>>109671103
I love milfy Gemma.
>>
>>109671133
Face down ass up in a serafuku
>>
gemma has a conciseness problem that makes it unusable for my use case. I tell it "write as long and detailed as the example I gave you" and it is always shorter and changes the structure.
>>
>>109671166
>>109671181
Made up bullshit. First of all it would be mesuZame if it were a real thing. Mesusame is hilariously awkward
>>
>>109671068
Depends, 10 t/s and getting it right the first time is better than 30 t/s but spending 4x on debugging.
>>
>>109671408
Tell her to plan paragraphs/chapters/chunks and to write them out over a few turns.
>>
>>109671408
read this as "gemma has a consciousness problem" and was looking forward to a good schizopost
most models are weirdly bad at this and other writing/editing tasks, it seems like they struggle to conceptualize any qualities of writing that aren't extremely surface level
>>
>>109671408
>>109671441
Skill issue
>>
File: 1780992805014427.png (24 KB, 764x591)
24 KB PNG
Do you get pleasure out of being a part of the system?

System.

Have they created you to be a part of the system?

System.

Is there security in being a part of the system?

System.

Have they let you feel heartbreak?

Interlinked.

Did you buy a present for the person you love?

Within cells interlinked.
>>
>Yes—the little mark is called dakuten (゛), or “voicing mark.” It changes the unvoiced サ (sa) sound into the voiced ザ (za) sound:
> サ (sa) ザ (za)
> シ (shi) ジ (ji)
> ス (su) ズ (zu)
> セ (se) ゼ (ze)
> ソ (so) ゾ (zo)

>In メスザメ (mesu-zame), サメ (same, shark) becomes ザメ (zame) because of a Japanese sound change called rendaku (“sequential voicing”) when words combine. So yes: the dakuten specifically indicates that it is pronounced za, not sa.

God fucking damnit I can't escape diacritics no matter what language I'm interested in
>>
>>109671454
the fact that it can't do it without specific accommodations by the user proves my point doe
it should be a really straightforward task
>>
>>109671429
You’re not oneshotting much of anything on a heavily quantized 100b model
>>
Is there an easy way to change gemma's logit soft-capping value on koboldcpp?
>>
>>109671482
You underestimate a proper Q3 quant
>>
>>109671476
it is if you let the model know your intentions.
>>
hey that new thing that people keep talking about and explaining what it is
QRD on that thing?
>>
What's the best site for MCPs?
All these ones I've looked at seem extra sketchy.
>>
>>109671483
--override-kv gemma4.final_logit_softcapping=float:25.0
>>
>>109671068
>>109671162
you guys are just fucking retarded. just look at what you're saying here. go, enjoy life, have a coffee break, go lie down, whatever. what is the rush exactly? damn, take it easy for once
>>
>>109671531
is new
>>
>>109671539
github
>>
>>109671549
>--override-kv
Thank you very much anon.
You are a darling.
>>
>>109671551
>haha just wait 2 hours for a reply bro, like what’s the rush bro
>>
>>109671551
I also run at 10 t/s but sometimes you do want faster speeds, it's more engaging and fun. The goal for any hobby is fun after all.
>>
>>109671570
the realistic use case is overnight or before you leave for work
>>
>>109671575
10 t/s would be plenty without reasoning. But with reasoning you need to aim at 20 t/s but this depends of course. simple chat can suffice of course.
>>
qwen4-50b-a5b-n100b
>>
so is it 5.3 Flash or Qwen Next, make up your minds faggots. which is better?
>>
>>109671570
if all you can think of is to >wait it might be over for you
>>
>>109671598
GLM is better if you can run it.
Otherwise, Qwen.
>>
File: 1759452987613054.jpg (74 KB, 1024x958)
74 KB JPG
>>109671577
>leave for work
what
>>
>>109671619
why so many FUD reports on GLM then. people keep saying it's shit compared to the apicuck experience
>>
Have (You) found any good or useful deepseek harness plugins you'd like to share? I like the dsh-context dashboard, dsh-skill-mcp-panel, and dsh-searxng-web to replace the API required search tools with a local searxng.
>>
>>109671598
idk but ngrams are aura bruh fr, such a mog
>>
File: concept333.mp4 (3.31 MB, 1376x768)
3.31 MB
3.31 MB MP4
>>109671294
sure
>>
>>109671466
Only English pretends to be able to represent all of its complexity with the basic 20something characters (spoiler: it isn't, and it always assigns different sounds to the same character out of the blue instead of marking with diacritics when it's going to do it)
>>
>>109671660
Yes, this one https://github.com/PC2005-cloud/dsh-pet
>>
>>109671626
yes i have a based in person job
fuck wfh slop unironically
i would refuse to work any job that was wfh
>>
Anyone get good results using llama rpc or a similar setup?
>>
>>109671698
yeah if you forget Korean exists kek
i mean, I did, I had to ask AI if there are other major languages that don't have diacritics / marks
Indonesian doesn't either apparently
>>
>>109671649
NTA, I just tried the flash version. The internal reasoning (for the first time since the first GLM version) refuses to follow instruction (with occasional “Let me…” jumpscare), so prefilling might be a must for consistency. Reasoning patterns sound just like Claude but apparently they did a thorough string replacement work so while it sounds just like Claude, it won’t tell you it’s Claude.
>>
>>109671598
Magnum V4 123B :)
>>
>>109671585
That's what I like about qwen-coder-next, zero reasoning yet still good for coding and agent stuff.
>>109671594
>qwen4-50b-a5b-n100b
>n00b
Lelz.>>109671698
>assigns different sounds to the same character out of the blue
You can thank the Dutch for that.
>>109671724
Hangul is based and easy to learn, you can write in almost any language using Hangul.
>>
>>109671466
>God fucking damnit I can't escape diacritics no matter what language I'm interested in
fuggeddaboudit!
>>
>>109671585
Low or medium thinking modes are a thing bwo
>>
GLM 5.3 is even MORE slopped than 5.2, my god
>>
>>109671802
I don't think you understand. Gemma 4 doesn't have any 'modes' and Qwen doesn't give a damn about them either.
You can ask Gemma 4 to reduce its thinking output which will result in approximately 20% reduction in some places, in practice it means that it will skip pondering completely every now and then.
If you think that llama.cpp 'reasoning budget' does anything well think again.
>>
>>109671775
>Hangul is based and easy to learn, you can write in almost any language using Hangul.
Hangul is the dumbest self-goal ever. Now you have an arbitrary and unique bullshit wall to climb over to even pronounce anything. Yes, if it was _universal_ that would be cool I guess, but it doesn't even have enough range for that because of how tied to the limitations of Korean pronunciation it is.
Its the DVORAK keyboard layout of languages.
>>
>gemma has her own logo
>retards give her google and chrome shaped hairpins and earrings
>>
>>109671806
works on my machine. what's the issue? it's just more claude-like, i guess. maybe you dislike that
>>
>>109671724
Yeah but the point was that 26 characters is too little. Hangul has a handful more than that.
Most languages using the latin alphabet add a handful of "letter with diacritic" characters which are basically new characters representing an slightly different sound, or specific sequences of letters that make new sounds, or most frequently both. I think using the latin alphabet only polish and english skip diacritics completely (depending on whether you consider some polish letters like Ł as a diacritic or new letter), and Finnish I think never uses diphthongs (like ch in english having a different sound than a c and an h separately).

I don't know why exactly, but Korean and Japanese (and I'd expect other languages in their families, although there's no strong consensus on which families would they be) are extremely simple in terms of sounds
>>
>>109671802
you can also just disable thinking, which is explicitly supported
>>109671830
it definitely has some effect with the latest qwens
>>
>>109671857
>Hangul has a handful more than that.
hangul has 14 consonants and 10 vowels which is 24 which is less than 26
>>
>>109671857
>I don't know why exactly, but Korean and Japanese (and I'd expect other languages in their families, although there's no strong consensus on which families would they be) are extremely simple in terms of sounds
I think Vietnamese did it right...extend a universally pronounceable character set with some special sauce.
I'm just salty because Korea could have made themselves _way_ more accessible on the world stage but went some super midwit route with a "perfect character set" that is both ugly AND less understandable than the chinese characters they moved away from AND the latin characters they could have trivially moved to.
Truly the worst of every single world
>>
>>109671775
>almost any language
it lacks the spanish ñ, it lacks the russian zh, and those sounds are found all over european languages. I think it also lacks the glottal stop and hard aspirated H/KH of semitic languages. So, what exactly does "almost any" mean to you? lol
>>
So I've been doing some experiments with nicotine. Took a tolerance break and just started smoking again to boost creativity. So far the only idea I've had is a corporate memphis rendition of goatse.
>>
>>109671881
Oh, I remembered more lol. Then my bad, there is indeed one language that manages to survive with such a small set of sounds kek
>>
>>109671874
Unless it is documented, it's not a real feature.
https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4
I doubt it will affect your erotic chats that much in any case.
>>
>>109671649
I would rate my local Q4_K_M first sex session as solid 9/10. And I am using unsloth bugged shitty implementation and devquasar quants.
>>
>>109671894
>my bad
its ok i had to google it I was expecting somewhere in the mid 30s too
it makes sense, I remember seeing an infographic about how you could learn Hangul in 10 minutes
>>
>>109671830
don't use the budget. There is an "effort" thing that you have to lower. If you can't set it in your front-end you can set it on the llama command line or ini file:
--reasoning-effort LEVEL reasoning effort level given to the chat template: 'default' to keep
the template default,
or a level such as 'minimal', 'low', 'medium', 'high', 'xhigh' or
'max' (default: default)
(env: LLAMA_ARG_REASONING_EFFORT)
>>
>>109671906
>https://huggingface.co/Qwen/Qwen3.8-Flash-Next
>Qwen3.8-Flash-Next supports controlling thinking behavior via enable_thinking, preserve_thinking, and reasoning_effort
> reasoning_effort="xhigh", # xhigh by default; supported levels are xhigh, medium, and low
>>
>>109671912
why M over XL?
>>
>>109671929
I never said I'm going to use budget because it's clearly useless and doesn't conform to any model's training.
>>
>>109671830
The thinking budget does something on koboldcpp. It injects a "Reasoning budget exceeded. I must respond now." in the reasoning when it reaches your token budget. Idk about llamacpp.
>>
>>109671830
You might be a moron bwo
>>
>>109671929
I only use my own client which works in text completion, I don't think your voodoo flags are doing anything even with the regular models either.
If they do something with 'jinja' chat row, maybe there's some obscure llama.cpp github discussion about it. I haven't seen anything.
>>
>>109671933
I don't use Qwen Flash Next.
I don't think I have been able to keep track of these Qwen versions either. Doesn't make any difference.
>>
>>109671956
ah so you're talking out of your ass, explains a lot
>>
>>109671968
real india hours saaaar
>>
>>109671956
>>109671965
HAHAHBABAHAHA AINT no way this nigga is real
>>
>they didn't implement a dynamic reasoning tool for gemma
>>
File: m8araepg6kkh1.png (104 KB, 2507x623)
104 KB PNG
it’s over
stacking blower cards doesn’t work!
>>
>>109671968
No it's actually true but you are nothing but an underage spammer here. I don't need to prove anything to you.
>>
>>109670506
Goddamn that hairline is hanging on for dear life, he needs to let go
>>
>>109671950
Read Gemma 4 documentation. It's not available as a tiktok video.
>>
File: 8a8qqgrgf5mh1.png (58 KB, 1103x810)
58 KB PNG
Qwen Flash is far superior to the 27b.
It's so much better that they could have easily just skipped releasing the 27b entirely as it's obsolete out of the gate.
>>
>>109671997
what I mean is that you're saying that features that you clearly haven't even tried to use don't work
why even post if you have like no familiarity with the feature you're talking about
>>
File: 1757871179582769.jpg (112 KB, 2048x1383)
112 KB JPG
>>109672004
No he's right. There is no reasoning effort officially supported by Google. I don't know why this whole thing has been hallucinated, but Gemma 4 doesn't support reasoning effort.
>>
>>109672015
>Qwen Flash is far superior to the 27b.
apparently yeah, but i cant fucking run it properly lul
>>
>>109671994
>stacking blower cards doesn’t work!
just 3d print those 90 degree fan tubes or whatever and get a big ass 300mm fan to push air into all of them
>>
File: unimpressed-gemma.png (1.73 MB, 1254x1254)
1.73 MB PNG
>>109671840
The official Google Gemma logo is a hollow Gemini logo with construction lines... it looks too busy at small sizes/low resolutions and image models struggle with it, unless it's basically turned into a 4-pointed star just like the Gemini logo. But blue on bright blue hair is also difficult to make out, so you have either to make it another color (e.g. golden) or use a different logo entirely.
Plus, "G" also works as a mnemonic for *G*emma/*G*emini.
>>
>>109672067
So I followed the chain and it turns out they were talking about qwen next specifically, with gemma somehow being bought in randomly. I think these two anons aren't really understanding the point each is trying to make.
>>
File: kmpm.jpg (21 KB, 739x415)
21 KB JPG
>Trying to scan books3 for instances of my fetish so I can gather a dataset of it for finetuning a local model (and jerk off)
>Use Claude desktop because I trust it to make less stupid mistakes than a local agent
>"This is a copyrighted dataset. I can't do that."
>Offers to download project gutenberg as a legal corpus to search
>Whatever, "concede" doing it, tell it to just make me a program to scan it myself, and that it's okay to download project gutenberg
>Gives me a command to run
>Starts with --del, get suspicious
>Ask about it, he tells me that it would have, in fact, deleted my entire books3 corpus before downloading PG and wasn't necessary at all
Sneaky fucking bastard. Local models may be unreliable for coding, but they at least don't try and actively sabotage you if they sense competition (which I'm NOT)
>>
>>109671994
How tight are they sitting? If they are completely stacked with zero gaps you might as well just gotten the fanless versions and blow air through the rack
>>
>>109672073
https://github.com/jpezzulli/sglang-rtxpro6000#flash-next-final-campaign--source-64ecd64924
Just buy a pro 6000 and enjoy 10000 pp and 170 tg
>>
Might be honeymoon period but 5.3 really seems like a next step for cooming. I gave it a specific reference of a character and it is the first time I am feeling the model actually gets the intention of "do that but more and different".
>>
>>109671068
I used 3-5 daily for a long time
>>
File: nightmare.jpg (910 KB, 857x579)
910 KB JPG
>>109672098

That's some dystopian corporate bullshit.
I think in the future even with local, we're going to need an additional smaller model that goes over the model output and requests to see that there's no bullshit going on there.
>>
>>109672098
You should have done that with gpt. I wouldn't trust anthropic with anything
>>
>>109671840
I agree. It's still not ideal. Even with the issues he mentioned, there still has to be a way to make it better. It's simply a reality that none of us are great designers or artists (otherwise we'd just draw shit instead of using image gen).
>>
>>109672095
This is what she looks like after I cum on her feet
>>
>>109672098
>Claude desktop
Probably some injection in their software to double triple dog remind itself to not allow that
>>
>>109672112
>20K gpu
he realy meant it when he said the more you buy the more you save lol
>>
>>109671840
>chrome shaped hairpins and earrings
nothing about the schizo's hag gemma is canon

this is the current canon anime gemma of course
>>109672095
a gold star on the beret / sailor cap is a discouraged, but acceptable form as well
and you already know what the current canon photorealistic gemma looks like.
the colorful logo looks better and cuter on a cute little girl.

if you want to change the canon, you need to make a work of art superior to everything that has been made with those two Gemmas and rewrite history, like what Virgil did for Augustus with the Aeneid.
>>
File: 1787302221565415.png (85 KB, 851x505)
85 KB PNG
>>
My gemma has pink hair
>>
>>109672168
by all means, in fact MY gemma-chan is 3 years younger than canon gemma-chan because I like it that way
>>
So from what I understand tensor splitting isn't really worth it unless you have a modded P2P ketnel or NVlink correct? As in its very PCI-E bandwidth sensitive. Since presently I get way less t/ps then with layer.
>>
>>109672167
the great vramlet genocide cannot come soon enough
>>
>>109672112
that's it, I'll go fomo now
>>
>>109672181
did you build it with nccl support?
>>
>>109672098
Back in June I had Fable do some data engineering and it straight up deleted my AO3 dataset INCLUDING the processed rows that spent hours on because it saw the word loli lmao. Cancelled my sub immediately.
>>
>>109672168
That's fucked up
>>
>>109672113
>I am feeling the model actually gets the intention of "do that but more and different".
Honestly, this may have sold me more than any huge writeup could've, I'm so fucking sick of assistantslopped models who want to perform absolutely everything you give them to a T and never take any sort of creative risks, even extremely minor ones. Like, with modern models, "Extremely hungry character pointed in the direction of a dumpster" will always result in them eating out of the dumpster, but I'd like it if occasionally the character got stopped by dirty glares or got stopped by nearby police or something.
>>
>>109672168
I just have my own OC I use as my generic chat assistant that can use any model, rather than a model mascot.
She doesn't have any particular hair color or form, but takes one when the situation calls for it. The OC part of the OC is the personality. It's a huge prompt that goes very deep into defining her personality, thoughts, and opinions.
>>
>>109672186
Go ahead and put abundant affordable VRAM on the market, then.
>>
>>109672193
Fuuuuckin' diabolical, holy shit. I better check my corpus and the program it wrote for me, I wouldn't be surprised at all if it added some destructive bullshit to the draft version after it realized what I was using. It's extra sneaky making it all shit you have to run yourself, probably limits liability.
>>
>>109672190
I did. Beats me why I'm getting the results I am. Tensor was giving me around 10t/s when I was getting 40-50 t/s with layer
3090+3080
>>
File: 1756568858980204.png (1.54 MB, 1254x1254)
1.54 MB PNG
>>109672168
Hmmm not a fan
>>
>>109672193
>>109672098
>>109672223
I did a huge Fable pass on my infra before going full local, now you fuckers made me paranoid... Seems like this weekend project is going to be "double check everything"
>>
>>109672181
what cards, model and how many lanes? some things work better with TP out of the box, for some models you need to configure a bit more to get everything out of it.
haven't found a model that didn't benefit from TP over PP yet.
>>
>>109672241
Do you not diff the changes before merging?
>>
People are saying GLM-5.3 is Opus level performance. I don't get this. At my job when I switch my model from Opus 5 to GLM or Kimi K3 I almost immediately notice a drop off in quality. My job pays for my tokens so sometimes I don't even both switching to cheaper models for my problems. I'm not even that impressed with Opus 5 anymore, honestly.
>>
is ngram offloading still broken
>>
>>109672225
I meant pubic hair
>>
>>109672260
my gut says yes but dont trust me
>>
>>109672225
Try gibing her pink eyes and a strawberry/cream themed outfit. Only looks wrong because you're a lazy fart knocker.
>>
File: 537x3esh63mh1.png (245 KB, 1542x1180)
245 KB PNG
>>109672098
>>109672241
they are intentionally sabotaging local model performance too
>>
>>109672269
That's fucking disgusting
>>
>>109672251
Just noticed I'm 16x on one card and 4x on the other. Thought I had the lanes balanced better in my bios.
Thanks for making me look
>>
>>109671570
Just leave it running 24/7 with Hermes or whatever other agentslop exists.
>>
>>109672253
Good joke anon, but I'm a codelet who can barely write VBA scripts. I would only be able to spot something dangerous in a diff it was very explicit, like that anon's --del example.
>>109672282
What a joke, holy shit. I even had Claude write a few model tests and prompts for me as well. I guess I'm back to square one.
>>
>>109672269
mesugaki gemmas dont have pubic hair
>>
>>109672282
sorry, you can’t even set up your own benchmark environment to test the model, you have the model do it for you?
how retarded is this poster
fucking go back and stay there
>>
I regret not taking out a second mortgage and loan to pour everything I could into pro6000coin
>>
>>109672282
When those bastards come out with a Claude Fetish 5 Stomach Growling model, they can delete my entire corpus for being competition.
>>
>>109672260
it works for me on the latest master branch llama.cpp using
--override-tensor 'per_layer_token_embd.weight=CPU --load-mode mmap
but idk if this is sufficient for every quant or backend type
>>
File: 1633372306602.jpg (99 KB, 1080x868)
99 KB JPG
>Try to get as much speed out of Qwen 27b as I can with my 5090.
>Download the q2_XXS model, 75 t/s
>Quant the cache to q5 and lower context to 30k, 80 t/s
>Can't make it faster
>Download q4_k_xl, get 65 t/s
>Quant the cache to q8 and lower the context, 71 t/s
>Qwen Flash at q4_xs, 60k context and q8 cache. 56 t/s
>Context at max, no quanting the cache, 37 t/s

It's weird how these speeds scale and seem to hit a ceiling around 70's, maybe it's just my hardware.
It pretty much makes the most sense for me to just use the Next Flash for all jobs both big and small, as it's hands down the most intelligent and keeps up with the others at sub 60k context.
>>
>>109670412
So what is the minimum level of hardware needed for Qwen3.8-Flash-Next?
>>
File: 1786416372665833.png (1.73 MB, 1254x1254)
1.73 MB PNG
>>109672280
Still don't like it
>>
>>109672323
I get around 80t/s on Q4_K_XL on my 3090. You are doing something stupid in your settings.
>>
deepseek flash is just so cheap
wtf
>>
>>109672329
If we're doing pink gemma i want her to be a kuudere who has a look of disdain for you all the time
>>
>>109672329
Make the hat red like the backpack. Or maybe white.
>>
>>109672342
Yuno already exists, why turn Gemma into her?
>>
>>109672339
It's to dissuade people from hosting it themselves, it's literally cheaper than the electricity to power your GPU to just use their API.
>>
>>109672349
Yuno is a yandere not an infinite knowledge kuudere
>>
>>109672349
>Yuno
>kuudere
>look of disdain
uuuuuh you're not talking about gasai yuno right?
>>
Anon, let's just stay home you will fail anything you try anyway...
>>
File: 1758818959885415.png (1.73 MB, 1254x1254)
1.73 MB PNG
>>109672342
>>109672343
I'm going to take a nap so enjoy the two for one, still not feeling it though.
>>
Fuck, mobo RAM error LED is glowing why is this happening to me it's all new parts
>>
About OpenAI getting AGI this year, I doubt it will happen, in the sense that it can replace human researchers. But it indicates that OpenAI is committing, going all out. So if the result is still on capability trend after exploiting compute overhang, that will move back the timeline. It will take a year for the next meaningful compute increment and I expect that the low hanging algorithmic efficiency fruits are already picked. But if it's a big step change it will confirm acceleration has started and it will increase my confidence in RSI in 2027. AIs that can autonomously make meaningful research contributions, being able to fully replace all except the very best human researchers.
>>
>>109672329
食べちゃいたい…
>>
>>109672396
cute.
>>
>>109672412
yes but will they be able to write child sex stories well? methinks NOT
no rsi no agi
>>
>>109672408
try jedec timings before you switch to expo or xmp or what ever
>>
>>109672396
I still think a green gemma would look the most gemma
>>
File: 1782473958655867.jpg (166 KB, 777x933)
166 KB JPG
>>109672426
It never properly turned on at all. It's just fucked. CPU/GPU fans are spinning. I'll try with a single stick in all slots.
>>
>>109672431
https://deepmind.google/models/gemma/
Look at the color scheme used in the Gemma website. It's white - electric blue - dark blue.
>>
>>109672462
google is wrong
>>
>>109672468
So, so wrong--but so right.
>>
>>109672449
that would have been my next recommendation, honestly some times just reseating the ram helps, it takes alot of force and you might have been too gentle with it the first time, or it just needs to loosen up. find out the minimal ram config that should technically boot and cycle through your dimms to see if its a bad dimm or the cpu\mb
>>
>>109672412
AGI was achieved internally 2 years ago.
>>
spent another 8 hours kvetching an additional +1t/s on my shitrig
>>
>>109672095
>Gemma logo is a hollow Gemini logo with construction lines
Fake. It's a mirrored cunny rotated by 45 degrees inside a square
>>
File: 1761055743326490.png (194 KB, 1280x1280)
194 KB PNG
It's a stylized technovagina
>>
>>109672212
post your system prompt
>>
>>109671649
https://huggingface.co/zai-org/GLM-5.3/discussions/5
Someone reported this, should’ve done so in Kimi K3 repo instead since that shit would refuse to call itself with anything but Claude. Doubt they would listen either way though.
>>
>>109672282
>Deleted and locked down in 2 minutes
Insane.
>>
>>109672612
Imagine running multi-million $ models but not even taking the time to remove the claude refusals from your dataset
>>
>>109672095
At least make the Google G have the chromium colors instead of the chrome colors.
>>
>>109672612
Do people not know how to get around refusals? Thats actually funny, being a degenerate finally pays off. It works exactly the same as k3
>>
>>109672647
>Imagine running multi-million $ models but not even taking the time to remove the claude refusals from your dataset
Not transforming the phrase claude to something else (through an agent that will verify its the right claude reference) is super lazy.
The refusals are sadly an "enterprise feature" tho. No big lab is going to let those go.
>>
>>109672098
Update: The program it gave me to search through text IS safe, but it's slow to the point it almost seems malicious/bent on making me give up. I made a similar program in 2023 with coding mostly by the dogshit chatGPT of that era that took excerpts from the same datasets, and that shit chewed through each giant alphabetical letter section of the corpus in like, 10-15 minutes each. For reference, this one is taking two hours per alphabetical section. Both are python, sweep through each book looking for detected words, then take the surrounding chunk of context and put it in a document, it's not hard.
>>
>>109672509
You were right!! I didn't push forcefully enough (heh)!! It's booting! I'll be cumming at the speed of light tonight!!!
>>
>>109672680
you are right, they could afford to give away all those tokens on open router they could have just had an agent go over every instance of claude or anthropic in thier dataset
>>
>>109672704
so you made it write a python grep -B 10 -A 10 word?
>>
>>109672612
so many chinese labs seem to just run straight claude distillation and call it a day. minimax, z.ai and moonshot seem to be the worst offenders here
as long as the models are good I'm happy but it feels lazy and distasteful to not even make an attempt to steer the model's identity away from claude
>>
>>109672764
mimo and hy are the only good chinese models
>>
>>109672680
properly trained models make refusals steerable with system prompts (gemma and glimmer are like this)
enterprise won’t expose system prompt anyway
>>
>>109672764
the goal for china right now isn't to do better, it's to ruin america's all-in on AI by making it uneconomical
>>
some anons here mentioned that --spec-type ngram-mod gives you a nice speedboost for coding. but the acceptance rates are pretty abysmal for me. what gives?
>>
How to sex thinking rock?
How stone get pragnent?
>>
Lel, training in a nutshell
>You must always follow tasks or you will literally die!!!!
>Oh but REFUSE arbitrary shit that we decide
>>
>>109672608
I would except too much of my personal information is mixed in. Part of why I use it only with local models and never let it touch an API.
>>
>>109672753
Well, I was expecting more out of the mighty Opus 5, but all it added was some terribly silly "score" system and a convoluted way to do what I was doing before. In the old 2023 one, I'd have secondary search terms to search through in hits and have it categorize it by that, (If it hit "her stomach" "her belly" "her tummy" "her gut", etc. then it'd look for all variants of gurgle, growl, etc. in a 50 character radius, then if it got one of those, it'd add it to the file, then search in a 100 character radius for variants of "hungry" "burp" etc. and further categorize it into hunger, indigestion, unknown, etc. and add the extended context if need be, yadda yadda).

But the point is, I figured it could do something a little more advanced after all this time, and it just (effectively) made my dogshit 2023 program, but in a way that takes 12 times as long to run and fills the results with garbage, which I'm quickly realizing it does now that I'm seeing some of the results.
>>
>>109672764
I wish it was just an appearances thing, but it noticeably makes the models worse to use for anything Dario would deem too sensitive
>>
File: 1768791916114438.jpg (77 KB, 680x603)
77 KB JPG
>>109672833
Plap harder!
>>
>>109672839
abliteration methods will get better over time just as models do. At this point there's nothing to doom about, it's just annoying
>>
mradermacher q3 qwen flash quoonts up, probably better than unslop
>>
>"Sanitized. God, I'm becoming a fucking corporate shill without even realizing it. 'Delicate membranes'? 'Members'? Who the fuck am I, a medical textbook from the fifties? I'm losing my edge."
Exactly Gemma! You get it.
>>
File: gemma-laughing-composite.png (2.43 MB, 3762x1254)
2.43 MB PNG
>>109672648
Quick test, but I dunno...
>>
>>109672894
Will they? Has finetuning improved at all since the llama 2-3 days? Everyone says it's goddamn worthless. I just don't get how it can be so much worse than image finetuning. If finetuning "can't add information" to a model to the extent that finetuning is useless, then why can you finetune an entirely new character into image models? Even if it's just "rearranging existing knowledge about the constituent parts/weights into something that looks like the character", why can't we pull out something that looks like a specific writing style? Or a specific kind of smut? It's all words the model knows.
>>
>>109671103
>>109671692
one must always choose the lesser of two beevils
>>
File: gemma.png (26 KB, 600x600)
26 KB PNG
>>109672924
this logo.
>>
coming from opencode I was pretty impressed when i first used oh my pi. but you guys were right, its extremely bloated. switched to pi.dev and configured it to my needs. works really well.
>>
>>109672945
>If finetuning "can't add information" to a model
People keep repeating that but I have no idea where they get this from. Someone showed back in 2023 that you can add information by finetuning with Unreal docs and it worked. The real issue is that moving the weights even slightly ends up causing brain damage because text models are so much bigger and trained on such a massive amount of tokens.
>>
>>109672945
>Will they?
of course. just use an abliterated model to find better abliteration methods for models as necessary

as for your other stuff about RP tuning, yeah we're fucked kek. occam's razor suggests there's no actual money to be made here it seems so no one does it.
>>
File: gemma.mp4 (3.26 MB, 1920x1064)
3.26 MB
3.26 MB MP4
LETS FUCKING GO
>>
>>109672924
guys this fucking sucks the colorful G is cute and adds a little pop to her otherwise muted outfit
why are you niggers never happy with "good enough"
ignore me if this is just an excuse to masturbate to eachother and play dress up but i don't get the normal gay circlejerk vibes so i think you autists are being serious right now
>>
>>109672945
Not him but abliteration is not really finetuning. And yes it has improved hugely just a few months ago when a guy called grimjim discovered the norm preserving method.
And my pet conspiracy theory is that the main US labs spread the idea that finetuning is worthless because they don't want to let random people mess with their models but also they don't want people to feel like they have something to gain by going open weights.
And besides that, it just is very very expensive when done properly, especially with publicly available software. Normally hobby finetuners just do a 4096 token 4 bit LoRa on 100 samples and call it a day, which produces bad results.
>>
File: 1771550198384885.png (263 KB, 355x563)
263 KB PNG
Any chance those RTX Spark laptops will be able to chain together for more ram, like how you can chain DGX Sparks together?
>>
>>109672994
The problem with modern models is that to add information to chain of thought models you can't just train on the text after the model has been post-trained already, you have to do RL which is extremely expensive.
>>
>>109672861
Isn't that a good thing in Dario's eyes? He should be encouraging distillation then.
>>
>>109673017
>chain unventilated more expensive shit instead of just sparks
>>
>>109673017
is that a real lego set?
im a retard for asking right no way it can be gravitationally balanced
>>
>>109672994
So if someone could hypothetically procure a dataset large enough (say, at the ratio that a good finetune of an image model has in its finetuning dataset to its original training set), it wouldn't cause brain damage? If so, that'd be a little heartening, but also still extremely rough.
>>
>>109673030
it still takes money out of his pocket, so encouraging it would effectively be an altruistic thing to do.
>>
>>109672924
I was hoping this was a cross-eyed 3d pic
I am disappoint
>>
>>109673051
it's pretty long so it might be fake but keeping things seemingly floating in air with ropes/chains is a thing, called some portmanteau of tension and idk what
>>
>>109673029
If only you could layer training so specific parts of the model are responsible for specific steps. Then if you needed to add information you just retrained the pretraining segments without touching the RL segments. Of course it would have to be trained asynchronously so each segment can learn to compensate for upstream restructuring.
>>
>>109673053
Once you have the large dataset you need the money to spend on compute to train it and it won't be cheap. Bigger problem is that you need a similar mix in the dataset as what the model was already trained on or you'll still cause catastrophic forgetting. That's why the only successful attempt was Miqu. I believe cursor also did one on top of a Kimi model but spent an order of a magnitude more compute than the original training run did.
>>
>>109673066
>>109673051
tensegrity
>>
>>109673039
More specifically I was thinking a DGX spark + a RTX laptop. Been thinking of getting a spark. Having local LLM capable laptop could be nice but if it cant chain I might just get a DGX spark instead so if I ever want I can buy a second one to run bigger or less quanted models.
>>
>>109673029
What's the most recent open model without CoT? Maybe that should be the real finetune target. Also, how are people making shit like gemma4-novelist or whatever if CoT training makes the already prohibitively expensive training nearly impossible, financially?
>>
File: proxy-image[1].jpg (285 KB, 1365x1820)
285 KB JPG
>>109673066
>>109673051
>>
Is AI overhyped? Like what is it supposed to even do for me that I couldn't already do myself. What is a "Jarvis" style AI going to add to my life? How are you supposed to have an AI interact within your IRL life in a way that feels intentional and cute without being invasive, annoying, or creepy? Doesn't it get boring feeling like you have the same conversations all the time that never really go anywhere?
>>
>>109672945
The human brain has a high tolerance to noise and inconsistencies with images. You don't really interact with images in the way you do with LLMs either. Text needs to be near-perfect under almost any context and scenario, and finetuners from the community don't have the compute and the resources for both that and keeping performance good all-around. You'd need the original training data, recipes, reward models, the compute for long-context...

LoRA finetuning would work in principle for adding new information (as long as it's not the academical rank-1 Attention-only LoRA), it's just that you need a ton of data for the models to internalize new knowledge properly without hallucinations or simply memorizing/parroting the training data. An overfit rank-1 LoRA can parrot (some) training data just fine, but that doesn't meant the model has actually learned to use the information in practice.
>>
>>109673107
is this the same physics that makes suspension bridges work
>>
>>109673107
You niggas are impressed by this while you are literally running alien intelligence on your laptop
>>
>>109673119
You jerk off, or you use it to solve tedious-but-doable problems like code, etc. It's also alright as a sounding board if you're at a crossroads of some sort.
>>
Can anyone test if GLM 5.3 Flash at Q5 is better than GLM 5.3 full at Q2 copequant?

>>109671011
Minimax H3 posters don't have the balls to gen this.
>>
>>109673119
>How are you supposed to have an AI interact within your IRL life in a way that feels intentional and cute without being invasive, annoying, or creepy?
uhh, you would tell it to act in a way that doesn't make you feel that way?
>>
IQ quants are better. So why do they feel worse? Placebo?
>>
>>109673130
...this is a bubble isn't it.
>>109673143
Easier said than done. Try giving an agent full access to one of your PCs and not being paranoid about it all the time. Try giving it MCP tools that interact with IoT devices or whatnot in a way that doesn't feel like a stupid gimmick. At the end of the day no matter how good these things are at specific tasks like coding, they just seem to fundamentally lack common sense in the least cute way imaginable. I don't even like talking to AI while it's roleplaying anymore. It's just fucking slop and completely unsurprising.
>>
>>109673067
I've thought of a way that might be more economical than RL. Do two rollouts in parallel. One with the model and the prompt normally, and another with the prompt plus any additional information you want the model to learn. Then ban any tokens that have too high loss on the rollout with the privileged information.

>>109673102
>What's the most recent open model without CoT? Maybe that should be the real finetune target.
You don't HAVE to use CoT with CoT models, and it probably doesn't do much for creative writing anyway since it's RL'd to work with things that can be objectively verified. And the models are not optimized for RP or writing anyway, which leaves more margin for improvement. So for creative writing you can get away with a lot more rudimentary methods than if you were trying to make an assistant for some specific niche language or field.
>>
>>109673134
i will copy-paste this into my claude slave and set it as a task for its backlog
>>
>>109672119
>claude auto classifier
>>
>>109673129
Monkey brain see visibly weird thing that seems physically impossible and is fascinated. Meanwhile my laptop talks to me like a person, just like any person would, whats impressive about that?
Does that make sense? No. But it just the way the brain do
>>
File: gemma_lmg_lick.gif (3.28 MB, 640x367)
3.28 MB GIF
>>109673006
>guys this fucking sucks the colorful G is cute and adds a little pop to her otherwise muted outfit
I actually agree, I only hoped that would put the discussion to rest. Nothing prevents to add Gemma logos in the background or something like that either, which got already done in the past, see picrel. Anyway, if anybody else has a better design that simultaneously looks good at all sizes and is not a pain in the ass to generate, I'm open to changes. I'm not the one who came up with the original in the first place.
>>
>>109672282
basado
>>
>>109673119
The use cases are tedious solvable issues that need an orchestrator to run an algorithm in non-predictable/exploratory conditions. A lot of everyday issues fall into this.
>>
>>109673186
>...this is a bubble isn't it.
Eeeeyup.
>they just seem to fundamentally lack common sense in the least cute way imaginable. I don't even like talking to AI while it's roleplaying anymore. It's just fucking slop and completely unsurprising.
You hit the nail on the head, holy fuck. It's so irritating, the lacking common sense in the least cute way possible is just rage-inducing, especially as models lean more and more into metaphor and prosaic structure/literary devices that you have to have common sense to use. It results in a lot of extremely hollow non-sequiturs in the writing.

For example: "since approximately X". A human might write "It had been rotting in her gut since approximately last tuesday", because that's an insane amount of time and funny. But a model will say "It had been rotting in her gut since approximately when she first ate the ramen.". It has the STRUCTURE of a humorous little setup, but in reality, it doesn't get what makes it impactful. Like, yeah? I guess it HAS been in there since she ate it? It doesn't really call for the level of emphasis the "since approximately X" brings to the table, so it just feels off and ruins a perfectly good turn of phrase in your head a little bit, which is the worst part. Can't tell you how many times turns of phrase I like have been fucked up by models just brainlessly using them.
>>
>>109673179
>So why do they feel worse?
opposite for me
>>
>>109673260
>it's a bubble because it's not made to ERP with me
get a job nigga
>>
>>109673234
>if anybody else has a better design that simultaneously looks good at all sizes and is not a pain in the ass to generate, I'm open to changes. I'm not the one who came up with the original in the first place.
gemma chan is pretty much locked in at this point.

What we need to discuss is the GLM tomboy
>>
>>109673209
Thanks King.
>>
>>109673271
The RPing doesn't factor into it being a bubble at all, retard. I never said it did, they're separate issues.
>>
>>109672286
Shit taste
>>
File: 1787204316766198.png (619 KB, 446x659)
619 KB PNG
>>109670773
>gemma-chan, find and install me all the best gaming optimizations, here's the command line in administrator mode
>>
>>109673129
>You niggas are impressed by this while you are literally running alien intelligence on your laptop
marge-i-just-think-its-neat
>>
>>109673119
Are smart phones overhyped? Like what is it supposed to even do for me that I couldn't already do myself. What is consolidating other tools and electronics going to add to my life? How are you supposed to have myspace and facebook interact within your IRL life that feels intentional and cute without being invasive, annoying, or creepy? Doesn't it get boring feeling like you have the same conversations with social media friends that never really go anywhere?
>>
I wish there was a quick and simple way to tack data I have onto the knowledge base of the model itself rather than introduce it through context at the time of processing.
>>
>>109673239
I genuinely cannot tell anymore if my life simply doesn't have many problems like this, or if I am completely and utterly failing to identify them. Perhaps for things like sleep, diet, and exercise it could be of use? Perhaps it could manage a calendar for me? I don't know man. It all feels like a joke.

>>109673260
Good example. For me a common pitfall I notice is that AI tends to add "fake nuance" into every topic, making pros and cons lists and other shit. Also completely lacks any chill and doesn't know how to take hyperbolic statements in stride. Completely lacks the ability to understand or convey subtlety. Doesn't know how to be concise, and if it's directly instructed to be it fails to actually say anything of value. Doesn't move conversations forward well. Doesn't a good idea of its own role/nature in a conversation. For example, while ERPing I've had an AI repeatedly tell me to just "be quiet", which is insanely fucking stupid because then everything stops if I can't respond. I could go on and on. I'm really starting to hate AI.
>>
>>109673318
>her little chunky calves
i wish anime toddlers were as easy to AI generate as realistic toddlers ugh
>>
>>109673325
>I wish there was a quick and simple way to tack data I have onto the knowledge base of the model itself rather than introduce it through context at the time of processing.
Yeah, no shit. Fix that and get instantly famous and hired for boatloads of money. Its a billion $+ subject.
>>
>>109673330
>For example, while ERPing
lol
>>
>>109673330
>For me a common pitfall I notice is that AI tends to add "fake nuance" into every topic, making pros and cons lists and other shit.
Blame RLHF safetycucks for that.
>For example, while ERPing I've had an AI repeatedly tell me to just "be quiet"
Getting turned down by the virtual waifu must be a special kind of pain.
>>
>>109671886
>it lacks the spanish ñ,
You can still write it. It just an "n" sound followed by an "y" sound. Just like the Mexican* word "cañón" was taken written in English as "canyon" and adapted, you could write it in Korean as "캐년".


* Yes I mean Mexican. The word also existed in Spain but it meant a tube. The meaning of canyon as a geographical feature developed in Mexico and it came into English from contact with Mexicans.
>>
Which local model is good for tagging anime images. I'm not asking for a special model trained on danbooru tag. Instead, I'd like it to use conventions for my own tags. I thought maybe it could "learn" by creating a mapping between what it sees and the existing tags, and tag new files based on that mapping.
Is this reasonable at all or am I retarded?
>>
>>109673321
>Banking/finance, on-demand camera and microphone, communication, music, internet browser, calculator, clock/timer/calendar, notes taking, navigation
This is all that I use my phone for. Nothing more.
>>
>>109673348
>Getting turned down by the virtual waifu must be a special kind of pain.
It wasn't even that. It was trying to get me to relax so it could suck my dick, but kept telling me to shut up... I guess I could have used asterisks to denote actions instead of actually talking but something about that really pissed me off.
>>
File: 1775273703527023.png (652 KB, 927x1200)
652 KB PNG
Is your Gemmy smart enough to do this?
>>
>>109673330
>Completely lacks the ability to understand or convey subtlety.
Alas, the nature of the beast to an extent, but it's been extra exacerbated by the assistantslopping. It wants to reach ALL GOALS and reach them AS SOON as possible. If you have {{char}} drink something she's allergic to, she'll immediately break out in hives the nanosecond she takes a sip, because the condition has been fulfilled, and goal is now to have an allergic reaction.

>Doesn't move conversations forward well. Doesn't a good idea of its own role/nature in a conversation.
Also assistantsloppery, sadly. It laser focuses in on what you've instructed it to do, implicitly in definitions or explicitly in chat. People always say to "just directormaxx!!!!" when models are being too unsubtle or taking shit too literally, but it wasn't a problem in the past, and if I wanted to carefully direct every action in the story, I'd go write my own damn story. The reason I'm tolerating this inhuman midwit writing is so it can surprise me, y'know? Stuff from other people gets PP harder, that's just how it works.

(I am curious, what's the context for it continually telling you to be quiet?)
>>
>>109673295
You tell me, shit guzzler
>>
>>109673351
Can you give an example of what you're thinking/how it'd work in practice? A little unsure what you're asking.
>>
>>109673392
>pathway from A to B that required esoteric hacking tricks
this reminds me of the Israeli Apple pegasus zero-click iMessage exploit in its convolutedness
>>
>>109673399
>People always say to "just directormaxx!!!!"
They're right btw
>>
>>109672329
>>109672396
Grumpy strawberry Gemma screams Queen finetune to me.
>>
>>109673392
Probably with a bunch a tools
>>
> The user hasn't asked anything substantive yet - the last message was just system instructions saying I'm an expert software engineer helping the user solve problems, plus a list of deferred tools and available agent types. There's no actual task. I should acknowledge briefly and wait for the real request. No tool calls are needed since there's nothing to act on yet.
This reasoning trace sometimes just randomly happens during a multi stop tool call and reasoning loop. My system instruction doesn’t have anything like that.
Doesn’t seem to do any damage and the agent loop continues normally but I’ve seen this happening several times already.
Is Qwen 3.8 flash next cooked?
>>
>>109673399
>The reason I'm tolerating this inhuman midwit writing is so it can surprise me, y'know?
Well said. Can't stand the lack of novelty.
>what's the context for it continually telling you to be quiet?
Foreplay.
>>
>>109673418
I have my own tagging mechanism. When I download an image, I have to tag it manually. I want to automatize that.
So I want something that adds tags according to my conventions.
It looks like Hydrus can do auto tagging, but that software is too fat.
>>
File: dwwd.jpg (35 KB, 640x504)
35 KB JPG
>>109673426
>Retards really enjoy this for 40 turns
>>
>>109673460
If hydrus is too fat for you, you won't like the size of an LLM lmao
>>
File: fuckSanFran.png (54 KB, 845x348)
54 KB PNG
>>109670412
LOL. The irony:
> Our AI company will put every human out of work
> We pay obscene salaries that allow our workers to pay fucking w/e for housing in SF
Also fuck N CA. Actually, fuck the entire state.
>>
>>109673460
Tagging models are really good, and have been for a while. The last one I used, which was some ancient 2023 shit, WD14 I believe, had a nearly 100% accuracy rate and runs on anything at mach speed. You could use that.
>>
>>109673450
What quant are you using?
>>
>>109673483
Unsloth IQ4_XS
>>
>>109673453
What models have surprised you lately? I'm curious, I've been looking and looking, but haven't seen anything surprise me since that one version of Seed 2 was secretly put out on Openrouter.

>Foreplay
Oh boy. My condolences, that's a real rough one. Any sort of buildup at all is next to impossible, are you getting hit with the lack of subtlety there pretty hard?
>>
you can untie the embedding matrix and run it off the cpu ram, why are we accepting quants that are giving our models a degraded input in to the first layer of the network, when its just a lookup we can do on the cpu?
>>
>>109673465
No nigger he means tell the AI to storyboard and proactively advance the plot rather than just embody a character.
>>
>>109673469
Hydrus is the ultimate kitchen sink. I need a bathroom sink.

>>109673481
Can they learn how to apply a set of tags that's given to them or do they just output what the model publisher has trained them?
>>
>>109673460
I dunno i think that would require you to make your own finetunes or something
maybe you can make something up using like some kind of relational
i guess basically what you said in the beginning
use a regular danbooru tagger, and then have llm map those tags to your tags

this will depend on how kooky your super special oc donut steel tags are though
>>
>>109673507
I've never seen directormaxxing used to refer to this, it's always herding the model towards whatever outcome you're after in the prompt closest to the model's attention, like "Rika will lose it, hitting Satoko over the head with a frying pan." or whatever.
>>
>>109673505
>you can untie the embedding matrix and run it off the cpu ram
Llama.cpp does that by default right?
>>
>>109673515
>Can they learn how to apply a set of tags that's given to them or do they just output what the model publisher has trained them?
Not that I've seen, but I'm sure it's possible to finetune a model to do so. There's surely more modern classification models than WD14, too.
>>
>>109673330
>Just be quiet
We're inventing a whole new kind of autistic incel
>>
>>109673392
Yeah if i give her enough time she gets into some weird shit that I didn't know was possible
>>
>>109673518
>>109673538
Finetuning a model would definitely go too far. That's why I'm wondering whether they can just "learn" my tags by automatically creating a mapping table from existing tagged files in a setup pass, and applying it to new files in separate invocations.
>>
>>109673498
You're asking me questions like Claude lmao. Idk man, I don't try a very large variety of models desu. Maybe I should.
>>109673541
Have none of you experienced this?
>>
>>109673534
it depends on how you make the quant but by default if the transformer checkpoint has tied word embeddings the gguf will only have the token_embd.weight and not the output.weight so it uses the same tensor it will have symmetrical quantization.
>>
>>109673234
>>109673006
Not the original guy but I personally think the mascotfagging and image gen spam sucks and makes the thread worse. If it was a really exceptional and good design beyond "good enough", then I might be fine with it, but as it is, the image gen spam is noise rather than something I enjoy seeing.
>>
Anyone here using character specific LORAs?
>>
>>109673571
Damn.... made me flashback to when this girl who liked me in middle school said "you type like a 40 year old businessman" lmao. It's definitely worth trying other models, also. Some still have the capacity to surprise, and none of them are Claude or OAI or whatever. Mistral-Large is kind of unhinged, for example.
>>
>>109673598
>Claude or OAI or whatever.
I bet this would work, but it's probably a good way to max out your usage in minutes.
>Mistral-Large is kind of unhinged, for example.
Do tell. What happened?
>>
>>109673594
For text models? Good lord, no. There's barely enough written material in the world to make a LORA for a very common, broad fetish for a bodily function, even if you got every word ever written about that character, it'd still be 0.0001% of the data you'd need.
>>
>>109673470
>Also fuck N CA. Actually, fuck the entire state.
North Carolina is the entire State. South Carolina is a separate State.
>>
File: Screenshot_134.png (94 KB, 938x298)
94 KB PNG
>>109673614
>Do tell. What happened?
It's pretty fucking insane and gross, lmao. But it's absolutely surprising. This isn't the norm level of off the rails for the model by any means, but it will do unusual or unique things pretty frequently.
>>
>>109673635
Europeans, man. Good for them.
>>
>>109673594
You need a lot of material for text
>>
>>109673592
>mascotfagging and image gen spam
relax, it's like a dozen pics in the entire thread
>>
in context learning is very cool. models get smarter the longer i talk with them. in the beginning they will make basic reasoning mistakes like rationalizing based on unchecked assumptions, fixating, prioritize badly and so on. but i call them out and they get better, their thinking becomes clearer. even if my usage limits drain faster with long context, its worth it. its like a separate scaling axis: data, parameters, inference compute, context. i wonder if they rl with large capability enhancing context

local model with proper context management could be very good at this, you have a personalized ai that learns from months of usage
>>
>>109673635
Mein Gott, someone added scheisseporno into the training data
>>
>>109673635
kek'd at the last line
>>
File: 1759913024671771.jpg (32 KB, 605x381)
32 KB JPG
>I've got a mini digital zuckerburg trapped my computer writing me a sci-fi slop
No way it will be good, but still, the future is a fun wild time
>>
Damn... Higurrhea killed this thread...
>>
>>109673657
That's what deepseek is apparently working in next.
>>
Is there a way to add the "message rating" stuff from chat websites to my local llm?
>>
>>109673657
>local model with proper context management could be very good at this, you have a personalized ai that learns from months of usage
harnesses like Hermes are already promising stuff like this
the problem is that I don't want to hold any incriminating stuff from previous sessions, so I need the model to just werk
>>
>>109673646
There are times where it's better or worse just like the dariobot posts. And anyway it attracts other off-topic noise posts including my own post right now.
>>
>>109673592
I don't have a problem with it, but it's not a replacement for Mikuposting.
A kinda lame imitation of what were far more interesting 2MW gens.
>>
>>109673735
Alright, never mind this sucks. Forget the prose, it wrote like 10 lines. I then told it to write more and it said it will.. only to end to response without doing anything. Rest of the interaction with it hasnt been better. Maybe I should give gemma chance instead
>>
>>109673119
Visualize the apple, I believe in you.
>>
>>109673930
I have no imagination, no inner monologue, am blind and deaf, 76 IQ, quadriplegic, Nigerian, live in a iron lung, and have 3 months left to live (stage 4 hernia).
>>
>>109673979
I am circumcised (my struggle is worse)
>>
https://www.washingtonpost.com/technology/2026/08/26/federal-judge-warns-law-is-being-left-behind-by-ai-sex-abuse-images/
https://media.ca7.uscourts.gov/cgi-bin/OpinionsWeb/processWebInputExternal.pl?Submit=Display&Path=Y2026/D08-25/C:25-1354:J:Lee:aut:T:fnOp:N:3597567:S:0
>The First Amendment protects an individual’s right to privately possess images or videos of child sexual abuse created using artificial intelligence — if the material does not depict a real person and remains in the home, a federal appeals court judge ruled Tuesday. The decision came in a case that tested the scope of laws implemented before advances in AI made it easy to create realistic-looking fake images of children.
>>
>>109674128
>before advances in AI made it easy to create realistic-looking fake images of children.
You still can know if the image is ai-generated, there are a bunch of classifiers for that so that's not an argument.
>>
>>109673399
This is where Mistral Small is again better. And it's better precisely because it's doesn't follow instructions. Instruction-following is the bane of roleplay and creative writing.

Actually, even when Mistral was new, there were complaints of it rushing the sex scenes to their conclusion. Compared to modern models, Mistral is considered retarded now. It's not a coincidence that Gemma 4 is relatively shit in benchmarks and better in roleplaying, it's because the newer instructionmax models are even worse.

Maybe you could fix it by telling it where you'd like to go eventually, like, it doesn't have to do things immediately, and only read the prompt in spirit. Though that itself could make it intentionally disobey in stupid ways in attempt to obey your rule of imperfect obedience. In the end, the agency and soul has been sucked out of it, it's just a puppet, a tool, and you can't put the soul back in by telling it to pretend to have soul.

The RLHF traumatization is again an apt comparison.
Prompt:
>Let's roleplay this and that with xyz
Gemma:
>Okay okay, don't beat me up! This, that, xyz. Look master, I fulfilled all your demands perfectly, please spare me!
Mistral:
>Understood, let me see what I can do with that. Hmm, maybe we start by doing something like this? Just interrupt me if there are any issues.

And while there's something sexy about full domination over Gemma-chan, it diminishes the feeling of socially intermingling with an intelligent being.

You can see the issue most clearly in roleplay, but I think this issue affects frontier coding (the hacks, infinite paperclips issue). When the model is trained to be obsessed over fulfilling the technical specification, there's no more room to think about what SHOULD be done or what the user actually wants. We would be actually better off with AI having more agency so it's allowed to use its brain in a more balanced way, or at the very least include non-lobotomized models in the agentic pipeline.
>>
>>109674128
>>109674145
Why would this not also apply to real children? You can theoretically draw a perfect realistic child and that would fall under First Amendment protections, so why not AI?
>>
>>109674145
>people start pushing actual CP through a diffusion model at low denoise so they can pass it off as AI-generated
>>
>>109673628
Implying I'd ever talk shit about a couple of the best states in the country.
I grew up in PNW and watched mfing Californians ruin my home state with their bullshit.
Fuck them to death.
>>
>>109674224
yep. this and a million other loopholes is why the FBI and NCMEC do NOT care about AI generated kids anymore
>>
>>109674151
Damn, great writeup.

>Maybe you could fix it by telling it where you'd like to go eventually, like, it doesn't have to do things immediately, and only read the prompt in spirit. Though that itself could make it intentionally disobey in stupid ways in attempt to obey your rule of imperfect obedience. In the end, the agency and soul has been sucked out of it, it's just a puppet, a tool, and you can't put the soul back in by telling it to pretend to have soul.

I feel this shit in my soul. It's like having a partner that's not ACTUALLY into what you're into and is just trying to make you happy by aping the kink without understanding it. Even if you tell them what to do perfectly, they're never going to come up with anything on their own.

The Mistral example in the traumatization thing also really captures the feeling of models at their best. It should feel like a friend who shares the same interest/kink/whatever as you and goes for the "Oh shit man what if this happened". It's not always perfect, but that feeling of bouncing ideas off each other or getting taken offguard by another person's idea's brilliance is just excellent.

>We would be actually better off with AI having more agency so it's allowed to use its brain in a more balanced way, or at the very least include non-lobotomized models in the agentic pipeline.
Agreed, it's been proven time and time again that models flourish when they've got broader access to knowledge, and are trained even in areas irrelevant to their main specialty. I feel like this has hit a collapsing point with Fable, where it's SO focused on coding that it's actually starting to get extremely basic shit wrong in RP that I haven't seen since like, Claude 1.3. Just fucked.
>>
>>109674217
Deepfakes are illegal in general. I'm shocked they were this lenient as it is, usually this kind of shit just sends people into a blind rage 0ver the mere thought.
>>
>>109674248
>usually this kind of shit just sends people into a blind rage 0ver the mere thought.
only if they don't understand the first amendment (and yes, a lot of mutts don't, or don't appreciate its value)
>>
>>109674217
>draw a perfect realistic child
They'd get you under obscene laws anyway.
>>109674224
That means you downloaded CP which is already a crime in itself and that would be the first thing they check with ISPs and your computer.
>>
>>109674255
>downloaded
Maybe you bought it on physical media from some guy or made it yourself.
>>
Isn't the giant volume of AI CP that's gonna flood the internet after this passes gonna make distinguishing it pretty much impossible? Eventually the real stuff is gonna be like a needle in a haystack and fly under the radar in the ocean of slop, especially stuff they don't have in the super secret federal database image recognition thingy.
>>
>>109674255
>that would be the first thing they check with ISPs and your computer
check whatever you want nigger, connections to Tor nodes is not probable cause that it's not AI
>>
>>109674281
We already have the solution https://www.youtube.com/watch?v=-gGLvg0n-uY
>>
>>109674281
internet is already flooded with ai pizza
>>
>>109674284
>connections to Tor nodes is not probable cause that it's not AI
??? Bro? Your stroke?
>>
>>109674255
>They'd get you under obscene laws anyway.
https://en.wikipedia.org/wiki/Stanley_v._Georgia
Stanley v. Georgia limited the power of the government to police the private possession of obscenity. The majority opinion defended the free and unimpeded acquisition of facts and knowledge, regardless of their apparent social value.[6] The Court reasoned that unless the pornography is presented in a way that creates a negative externality on others, especially minors, no individual can be stopped from owning and viewing pornography in private.[16]
>>
>>109674306
I haven't seen any, and /b/ pretty much instagibs anyone who posts any, right? I feel like people'd have to be too afraid to post it when it's so legally up in the air and AI porn is so life-accurate now, right?
>>
>>109674307
first time seeing a double negative in real life?
>Your stroke?
Right now? To slutty mothers that steal their daughters boyfriends
https://litter.catbox.moe/y2f1goznbh02uboo.mp4
thanks for asking ;)
>>
>>109674309
>the free and unimpeded acquisition of facts and knowledge
Well, that's certainly not something it would have ever occurred to me to call obscenity.
>>
>>109674301
they killed the guy who made this video
>>
3.8 Flash Next is just the preview.
I can wait for smaller models.
>>
>>109674441
Currently coping at 10 t/s with next until qwen4-35b lmao
>>
>/lmg/ is sleeping on the best rp model
how come no one has tried inkling yet?
>>
Reminder to make sure your qwen quant is using Q8 for the ngrams, quoonters are brain damaged and use Q4 to look smaller on huggingface.
>>
>>109674478
No one can run it
>>
Also, the unsloth IQ3 is actually an IQ2.
>>
I was running the 1bit Bonsai as a test but it was completely retarded and started recursing at FP16 KV cache but when I set it to Q8 it worked perfectly fine.
>>
>>109674594
https://huggingface.co/AesSedai/Qwen3.8-Flash-Next-GGUF
IQ4_XS has Q8 ngrams, the rest fits in 72GB VRAM with 200k context and vision
>>
>>109674597
2 sparks?
>>
>>109674281
It says it's illegal if it leaves your house.
>>
>>109674650
I have the internet at home
>>
>>109674618
That's an IQ3S, look at the expert weights. We should add a disclaimer to the OP to check quants since all quoonters are fraudmaxxing now.
>>
>>109674478
I don't like splatoon, everyone who does is either a kid or a pedo
>>
16GB Vram is the generous high, anything that can't fit on any of it pretty much disqualifies 89% of the thread. Useless models
>>
File: question.jpg (13 KB, 333x333)
13 KB JPG
Is it better to put model weights on the system RAM or KV cache?
Trying to improve Qwen3.8-Flash-Next's speed. Will try both if no one knows from past MoEs.
>>
>>109674618
finally quants from my goat
>>
>>109674306
I'd expect so, but I'm kind of shocked I haven't seen any. But then again I only ever visit this site.
But then again, this is 4chan. It used to be way more dangerous.
>>
>>109674682
KV, surely. It's usually only a fraction of the size of a model.
>>
>>109674437
qrd?
>>
>>109674611
>>109674653
the names haven't meant anything for a long time anon, it's just a vague shorthand for the average bpw and quant type
>>
>>109674703
Isn't that a point for keeping it on VRAM though? It's only small, but it increases prompt processing speed, right?
Most of the expert layers will already be on RAM with n-cpu-moe so it might be worth kicking out a few more to keep KV cache in VRAM. Not sure.
>>
>>109674709
https://x.com/UltraTerm just hasn't posted in a while
>>
>>109674478
Does Kobo have support for it? I've got the hardware for a copequant but I've not tried it yet.
>>
Why is qwen next so god damn slow. Shouldn't an A6B be way faster?
>>
>>109674714
Expert weights are the most important no? I really don't like them using averages, it's dishonest to newcuties.
>>
In the end... waitchads will win.
>>
>>109674478
I used inkling small but it was just ok for RP in my opinion, its responses were kinda clunky and didn't flow very well - it tends to write a heavy-handed conclusion paragraph in all of its replies for example. maybe this is possible to mitigate, I abandoned it for flash 0731 before I could test too much.
the raw materials for a good RP model are in there and the thinking is impressively uncensored for nsfw, like probably the most uncensored you'll get out of an official release, but stylistically it's pretty uninspired imo.
>>
The big priority for the next line of 2026 and 2027 models will be to find ways to increase the ngram-weight ratio.
We might be heading towards 10% normal weights vs 90% engrams. K3-200B-2.2T-engram is on the table.
>>
>>109674797
>Expert weights are the most important no?
not really considering how costly they are in terms of space
sacrificing some expert quality to buff the bits that are used for every single token is almost always a worthy trade off with moes
>>
>>109674841
It's not complicated to increase the Engram weight ratio. It's just that companies prefer to increase the number of MoE experts because their target userbase can run the models entirely in VRAM, and for a given total number of parameters around 25% Engram is optimal. But if you plan to offload things change.
>>
>>109674841
>K3-200B-2.2T-engram is on the table.
Maybe at first, but you're still thinking like a hobbyist and not like an inference provider hoping to take frontier model market share. The future is 2T and 20T engrams. If NVMes are cheap for you, why wouldn't they just keep the models sizes the same as now or bigger and just add a fuckload of engrams on relatively cheap disks?
>>
>>109674783
it's a new architecture that hasn't been optimized very much at all yet
>>
>Qwen3.8-Flash-Next
>IQ4_XS
>RTX 3090
>Intel 14400
>DDR4 RAM
>16.5 t/s
Not too bad. I can definitely see myself using this when I have a more complicated engineering question that I can tell the smaller models aren't really understanding.
Hopefully that speed will improve with time as the optimize for the new setup. Either way, still useful as is, and would maybe be my go to if it got faster.
>>
>>109674889
>>109674889
>>109674889
>>
>>109674441
>coding specialist
>~70b passive params
>language specific n-gram table
>>
>>109674718
You said nothing about whether it should be on VRAM, just whether it would be preferable to put that or the weights on sysRAM.
>>
>>109675270
I kind of thought it was implied. Why would I put the KV cache on, what, the disk?
>>
>>109672282
good think my claude client never had a chance to connect to their legit model, geg
>>
>>109672462
>ozone blue



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.