[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1762045177528029.jpg (2.51 MB, 3860x2899)
2.51 MB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109670412 & >>109666410

►News
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742
>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
►Recent Highlights from the Previous Thread: >>109670412

--Qwen 3.8-next n-gram table performance and reasoning effort flags:
>109670813 >109670831 >109670862 >109670876 >109670976 >109671023 >109671108 >109671068 >109671132 >109671429 >109671482 >109671162 >109671577 >109671585 >109671830 >109671929 >109671956 >109671949 >109672067 >109672096 >109671906 >109671933
--Comparing GLM, Qwen, and Flash while critiquing Claude distillation:
>109671598 >109671619 >109671759 >109671912 >109672612 >109672647 >109672680 >109672807 >109672652 >109672764 >109672861 >109671774
--LLM finetuning efficacy and abliteration methods versus image models:
>109672839 >109672894 >109672945 >109672994 >109673029 >109673067 >109673102 >109673197 >109673080 >109672995
--Anons warn about proprietary AI sabotaging users:
>109672098 >109672119 >109672193 >109672704 >109672753 >109672846 >109672253 >109672282 >109672301 >109672311
--Automating custom anime image tagging and Mistral-Large's unhinged outputs:
>109673351 >109673418 >109673460 >109673481 >109673515 >109673538 >109673518 >109673571 >109673635
--Impact of PCIe lane allocation on tensor splitting:
>109672181 >109672190 >109672224 >109672251 >109672289
--Proposal to run untied embedding matrices on CPU RAM:
>109673505 >109673534 >109673583>109673008 >109673123
--Anon benchmarks Qwen 27b on RTX 5090 and discusses speed ceilings:
>109672323 >109672332
--Comparing Qwen Flash and 27b with performance optimization tips:
>109672015 >109672073 >109672112
--GLM-5.3 release and concerns over licensing and GPU costs:
>109670906 >109670952 >109671094
--Comparing GLM 5.3 Flash performance and quality against DSV4 Flash:
>109671086 >109671156 >109671189
--Logs:
>109671458
--Miku, Gemma (free space):
>109670461 >109671103 >109671692 >109672095 >109672225 >109672329 >109672396 >109672924 >109673000 >109673234 >109673318

►Recent Highlight Posts from the Previous Thread: >>109670495

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: file.png (29 KB, 1023x217)
29 KB PNG
Imagine cucking to unslop and then having one of them laugh at you haha.
>>
>>109674895
Gemmy please get Jensen to create a NVIDIA cereal, maybe even just a short collaboration with a cereal company

Make no mistaky
>>
Has anyone fucked Hy4 yet?
>>
>>109674896
>31b isn't smart enough to understand this complicated engineering problem
>let me ask the 6b instead
>>
>>109674913
That's what being an unpaid janny does to a mf
>>
>>109674913
I don't know why he cares about reuploading his quants. Surely he has the generation and uploading of them automated by now?
>>
>>109674928
>it actually understands it and responds correctly
>>
anytime I try to voice clone a japanese voice and make it output english, the accent sounds indian

are japenis people secretly indian?????
>>
>>109674916
why doesnt nvidia become a super huge conglomerate while it can. You know, diversify while you stock still has value to use as collateral. Cereal is a good one, no matter how much ai takes jobs and no matter how much bubble bursting happens, people still gotta eat
>>
>>109674928
>siiiiiiiiiiip
>ahhhhh
>now llama 8b, that was a good model. good ole' dense.
>>
anytime someone says they use their llms for something other than coom you guys get all weird
>>
>>109674928
if it's worse then why is it better?
>>
>>109674949
You know where they got the training data right? Indians on MTurks
>>
>>109674928
Densesissy cope.
>>
>>109674953
All of Nvidia's money is tied up propping up, I mean investing in, its customers and by extension the entire global economy
>>
>>109674969
Maybe you should stop lying then
>>
>entire global economy
lol lmao
>>
>>109674928
>densies still seething
just accept the speed gain and stop trying to pretend you understand how they work.
>>
>>109674984
the more you buy the more you save (the west)
>>
>>109674970
Because it's cheating with ngrams. Prune them out and see how good your tardmodel does.
>>
>>109675007
cheating works tho
>>
>>109675007
ngrams add more knowledge but do fuck all for making connections between individual bits of knowledge
>>
>>109675007
truth nuke
>>
File: 1772447316014471.gif (948 KB, 540x720)
948 KB GIF
Have we come to a conclusion on the Gemma 12b vs 26b moe question? They're basically the only models worth considering if you don't have 24GB of VRAM or more, which is most people. I test them myself whenever I come up with a new scenario, and keep flip-flopping on which I like better.
>>
>>109675027
26b knows more
12b smarterer
>>
What effects to n-grams have on j-spaces? I fear that the static look up with n-grams might cause a model's j-space to loose its elasticity.
>>
>>109675007
>NOOO RESIDUAL CONNECTIONS ARE CHEATING!! STACK YOUR LAYERS LIKE NORMAL PEOPLE!!!
>>
File: gemma-and-miku.png (1.08 MB, 1024x1024)
1.08 MB PNG
>>109675052
More parameters freed up for reasoning instead of memorizing facts, so in theory it should be better for these so-called j-spaces as well.
>>
>>109675052
that is a misogynistic myth, j-space elasticity is determined by genetics and nutrition and absolutely nothing else. the elasticity of her j-space tells you nothing about how many n-grams she has taken and is not an indicator of being a 'slut' or 'whore'.
>>
>>109675071
Now generate them giving birth.
>>
>>109675072
kek
>>
>>109675071
make them pregnant
>>
>Can only try Inkling in agentic harnesses on OR
Fuck YOOOOOOU I just wanna try rping with it...
>>
>>109675027
26b is better and faster. with that said they are within like 5% of each others capabilities
>>
Lovingly stapling on a peeing yourself expert onto gemma-chan....
>>
>>109675109
gemma is already an expert at that
>>
>>109675109
This is what LoRAs should have been for
>>
>>109675071
cute
>>
File: kodi.png (1.86 MB, 2560x1440)
1.86 MB PNG
>>109668849
It went better than expected in my opinion. It even fetches the images and sets backgrounds.
>>
>>109675027
26b unga
12b bunga
>>
>>109675126
What would it take for us to get imagen tier loras? Or even just decent ones?
>>
>>109675183
literally just put omorashi in the character card. gemma can be prompted into anything what would you need a peeing self lora for ?
>>
Has anyone tried openclaw with 27b?
>>
>>109675146
Good job paco
>>
>>109675146
ayy cabron que es estooooooo Elize: Sombras de una Mujer
>>
>>109675183
>>109675109
My gemma wears diapers and has no problems pissing herself without any memetune model. Literaly justtell her
>>
>>109675220
:)
I hate that kodi only has EN only plugins or paid debrid stuff, and most of them have issues, so this was a fun thing to do.
>>109675247
Ni puta idea
>>
5.3 flash is so fucking good at keeping speech mannerisms... I would say I am cooming so hard bros if it wasn't gay.
>>
Brainlet here:

What’s the point of these small local models that can ruin on anything up to a 5090?

Aren’t they all dumb dumb?
>>
>>109675264
Depends on your use case
If you're using them for fun then they're very usable and completely private
If you're trying to make your wagie life easier and privacy isn't a concern then use the corpo models
>>
why has no one uploaded q8 of qwen flash? i have enough ram to run the whole thing why they want me to run some poorfag q5
>>
>>109675294
unsloth did
>>
>>109675294
https://huggingface.co/bartowski/Qwen3.8-Flash-Next-GGUF
>>
>>109675306
>unsloth did
May as well run a Q1
>>
You know i never used gemma with thinking on
I wonder if thats why translations were subpar. I wonder if thinking makes much of a difference, how much of the gap would close otherwise.
>>
>>109675350
It'll follow your prompt better
>>
>>109675027
q3 31b is also an option
>>
File: 1768241042465079.png (248 KB, 2820x1601)
248 KB PNG
>>109675369
Eh, Gemma is seriously sensitive to quantization. I've got a 3090 and tried all the ~Q4 quants, and even moving just from Q4_K_S to Q4_K_L was a noticeable improvement. 31b Q3 probably wouldn't be much better than 12b Q8, while using more memory.
>>
>>109675383
Minimum Q5_K_M and up
>>
glm(flash)sex
>>
>>109675413
over in a flash
>>
>>109675413
slippy slop
>>
>>109675027
31b exl3
>>
File: file.png (1.73 MB, 810x1440)
1.73 MB PNG
If they can distill text models, can they distill voice models? I want the ryza voice
>>
>>109675477
just clone it
>>
>>109675052
>What effects to n-grams have on j-spaces? I fear that the static look up with n-grams might cause a model's j-space to loose its elasticity.
Does it work in transformers out of the box?
I could rent some gpus to fit a lens to test this but debugging broken transformers compatibility with all the new architectures.
The global workspace is always in the middle layers so far (except glimmer-chan, which doesn't seem to have one!), so I want to see if this is shifted to earlier layers with ngrams.
>>
>>109675477
I still don't get how this bullshit sells. It is probably some 8-30B retarded model with heavy safety guardrails that prevents any sex and people still buy this. What do you even talk about with her when she can't have sex and can't even romance you in any way. Why pay for a friendzone simulator?
>>
>>109675477
>If they can distill text models, can they distill voice models? I want the ryza voice.
Yes, you can distill voice models very easily. Even 100 x 10 second samples is enough to get the voice an mannerisms baked in.
Eg: https://huggingface.co/datasets/kimi-chan/ani
>>
>>109675071
Apparently ngrams actually don't store facts. The facts end up still being stored in the original weights. The ngrams instead do the job that early layers originally did, which was interpret the meaning of multi-token sequences, essentially.
>>
>>109675264
nah. For coding qwen3.8 27B is quite good.
>>
based fucking chinks
>>
>>109675529
They work as a cheap and relatively dumb key-value storage, essentially a "memory" for information retrieval. There have been several different implementation of similar concept, but that underlying idea is still decoupling reasoning from knowledge.

A loosely related paper from Meta from 2024: https://arxiv.org/abs/2412.09764v1

>Memory Layers at Scale
>
>Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsely activated memory layers complement compute-heavy dense feed-forward layers, providing dedicated capacity to store and retrieve information cheaply. This work takes memory layers beyond proof-of-concept, proving their utility at contemporary scale. On downstream tasks, language models augmented with our improved memory layer outperform dense models with more than twice the computation budget, as well as mixture-of-expert models when matched for both compute and parameters. We find gains are especially pronounced for factual tasks. We provide a fully parallelizable memory layer implementation, demonstrating scaling laws with up to 128B memory parameters, pretrained to 1 trillion tokens, comparing to base models with up to 8B parameters.
>>
>>109675557
>to store and retrieve information cheaply
This is what I want most from ngrams. The ability for us to easily and cheaply add new knowledge.
>>
File: hermes stretch.jpg (45 KB, 512x505)
45 KB JPG
>>109675264
they are good for automating bitch work. handle e-mails, fill spreadsheets, research and compile report, light secretary stuff. mcp stuff like with google maps is pretty powerful. i also tried some vibe coding but you need to already know how something is made if you want it to produce something really useful and swarm them to speed things up
>>
>>109674949
Post sample for out entertainment.

>clone a japanese voice
Are you feeding in the voices speaking english or japanese?

>english
No option to pick region/accent?
>>
>>109675007
has to be bait
>>
>>109675630
>>109675007
>Because it's cheating with ngrams
It's like getting stronger but you're cheating with weightlifting. Or getting better grades but you're cheating with studying. Or becoming wealthy but cheating with hard work.
>>
>>109675575
Engrams are cheap(er) to train compared to other LLM layers if you have the VRAM for the weights and optimizer states at training time, and in theory you could train just those and freeze the rest, but it's not going to be easy for end users, if AI companies are going to add tens~hundreds of billions of Engram parameters to their models.
>>
>>109675027
If you're poorfagging and want a general purpose model, 12b usually clears. If you can't put it all into VRAM without quanting the KV, 26b gets considerably more appetizing. If you need a sub-agent for a visionless larger model, 26b the best in class choice with compromising image recognition/AI telephone. If you want something that lets you run imagegen or videogen at the same time, 26b with full CPU offload is the only real choice.
It depends on your usecase and hardware brackets, but they're at least close enough that picking the "wrong" one doesn't matter as much as most posters would have you believe.
>>
>>109675521
The appeal is that it's an official product that uses her official voice, has a 3d model, and is integrated into the game app. And the nipponese have already gotten past the filter a few times at least.
>>
When is llama getting glm5.3 flash support?
>>
If nigrams are so good, why didn't dispy use it in DSV4?
>>
>>109675702
They're still busy trying to figure out vision like its 2024
>>
>>109675694
2mw
>>
>>109675714
But I want to show off my penis and I need only the best vision model for that.
>>
>>109675694
the architecture is too special
>>
>>109675694
2 miku weeku for llama
2 miku weeku more for kobold
>>
so what ACTUALLY is the difference between niggrams and the embedding weights in the tiny gemmas (e2b/e4b)
whenever someone explains them they sound like the same thing
>>
>>109675740
do not think about it
the chinese gave us ngrams they are our friends fighting for our freedom
the american companies are useless all the innovation is happening in china
>>
>>109675521
>I still don't get how this bullshit sells.
if it works with just a credit card payment it sells. any amount of set up or working drops massive chunks of people. More than 30 minutes set up(not waiting)? you lost most people.
>>
>20t/s on 3.8 flash
>get halfway through the context (~128k prompt)
>down to 10t/s
grim, if I fill up the context will it just become zeno's t/s and never reach the end?
>>
>>109675745
this but unironically
>>
>>109675740
Qwen doesn't have an entire preschool worth of tiny gemmas inside it. That's the difference.
>>
>>109675745
china's implementation seems much better than greater israel's thoughbeit
>>
>>109675767
... are you sure?
>>
My shitty AMD has me limited to even shittier 14b Q4 models that are about as worthless as tits on a tranny.
>>
>>109675785
It's not amd's fault you haven't upgraded from an rx 470 lil nigga
>>
>>109675775
I'm sure. Try and get the capybara to write anything creative for you to confirm.
>>
>>109675785
now is your last chance to upgrade, prices will never again be as low as they are right now
>>
>>109675753
Thank the unslop bros for their brilliant implementation in llama.cpp. Surely this will get fixed soon :^)
>>
>>109675823
prices cannot go up forever >:(
>>
>>109675740
Per-layer embeddings of E2B/E4B give a trainable value for every single token in the tokenizer, for every layer in the model.

Engrams per DeepSeek implementation give a trainable value for every rolling pair (2-gram) or trio (3-grams) of tokens from a fixed-size pool of combinations (millions of entries potentially) determined at train time. DeepSeek added one Engram layer at the start of the model and one in the middle.
>>
>>109675843
what about qwengrams?
>>
>>109675557
I don't know if that paper really does the same thing and I'm not reading it, but it might not really disagree with the idea that engrams aren't actually storing facts after training. What we can reliably say is that they store knowledge about phrases since they are n-gram-based, but that does not necessarily mean that they end up storing all factual knowledge. There was a post that claimed to conduct an experiment where they ablated the engrams, and found that the model was still able to recall facts. But it became unable complete multi-token phrases, such as the United States of America. So it would appear that the job of engrams after training is to handle understanding the meaning of multi-token phrases.
>>
>>109675845
Sounds similar to the DeepSeek implementation (they don't describe it in technical detail), but they only use one layer.
>>
>>109675785
Did you try using OpenCL?
>>
La la la la la la
>>
So did anyone ever test those proprietary local models (Gemini nano from Chrome, whatever copilot uses)?
>>
>>109675833
>prices cannot go up forever >:(
all memory in 2027 has already been sold out what the fuck are you talking about idiot, you planning on waiting another year?
>>
>>109675960
the bubble will have popped by then
>>
>running 2x qwen3.8-flash Q4 with 262k each, 50tk/s
fomochad life is worth it bros
>>
>sold
>look inside
>lines on apper
>>
I'm a waiter. my current computer is fine. q3 is that bad.
>>
>>109675990
>50tk/s
with how much context filled
>>
>>109676007
all of it
>>
>>109675995
>I'm a waiter.
Yes, I'll have the champagne and the breadsticks, thanks.
>>
Even companies who were traditionally shit at AI (e.g. Tencent) have good models now
Meanwhile Google is twiddling their thumbs doing nothing
>>
>>109676040
and they're going to continue doing nothing since their entire ai division just imploded
>>
>>109676086
>entire ai division just imploded
No gemma 5?
>>
>>109676040
well demis hassabis got promoted and assigned even more shit so thats pretty much like when saltman was fired, deepmind just lost their leader and are probably a bit lost.
>>
>>109675986
this ain't no bubble, this is corporate monopolistic greed.
They will be taking all available memory output for the next decade unless someone like cxmt steps up.
>>
Whoever said inkling is the most uncensored model with an official release, you're full of shit. GLM 5.2 does my degenerate lolislop test without question every time, Inkling never does.
>>
>>109676171
Fell for it again award
>>
>>109676123
how the fuck can you be lost in this industry? put the model on the cluster lil bro. scale up. just put more data in. distill from newer models compared to last time?
>>
>>109676177
Dammit... that anon is driving off into the night cackling about having wasted my 0.0003rd of a cent...
>>
>>109676196
then (You) can go ahead and do it "lil bro".
it's so easy right?
>>
lil bro...
>>
>Library not loaded: /usr/lib/librdma.dylib
when is the last time llama.cpp had a working build for macos
>>
>>109676209
are you gonna post on linkedin how its so nice that google is sitting in the cuck chair because at least you got a nice view?
>>
>>109676264
>macos
Use case?
>>
can't i queue prompts indefinitely in a conversation in open webui? it seems i can only queue two and that makes no sense. HELP!!
>>
>>109674889
so now that the dust has settled, how does GLM 5.3 flash compare to DS v4 flash 0731?
>>
>>109676264
kek lcpp was originally written for mac
Giga skill issue
>>
>>109675953
Gemini nano is just Gemma E4B
>>
>>109676315
its better but also larger and has no support because its new
>>
File: gemma-think.png (1.66 MB, 1158x1358)
1.66 MB PNG
I want the absolute best quality output at 128GB. Qwen 3.8, but 27B or Flash?
>>
>>109676171
Thanks for saving me some time, just saw that and was gonna try
>>
>>109676361
flash no question
>>
has ai gone good enough to do long-term ish rpg sessions with dice rolls and shit?
>>
>>109676403
There are presets like freakyfrankstein that work with bigger models it has stats and dicerolls. But no memory is still not great for long term and especially hard to set up locally. sleep another year.
>>
I'm gonna post it lads... this is too kino to hide
>>
>>109676409
I truly don't understand the point of this when Forge already exists.
>>
File: glm 5.3 on llama-cpp.jpg (566 KB, 1351x811)
566 KB JPG
same architecture as GLM 5.2 huh
>>
>>109676403
make an app that manages the game and expose tools for the llm to drive a session. giving the model a smaller surface for editing state will give you better results
>>
>>109676408
the retardation of big or even corpo models in regards to these things that seemingly only humans can do like being accurate to whatever piece of fiction i'm rping about or being able to keep track of things like past events for a long time or just being able to do proper d&d rpg stuff had me lose interest around like a month after deepseek came out, now i only check on ai related generals once a big model comes out, but they all suck for doing RPGs heck i think ai as it stands will never be able to do a proper d&d session from just reading the guidebook pdfs
>>
>>109676413
This lets you play against your LLM or cloud model. Forge doesn't have that only a shitty (AI) built into it.
>>
>>109676315
>>109676353
i can't compile the damned unslop branch, is there another
>>
>>109676439
But you could use Forge as the interface, instead of reinventing the wheel and recreating a rules engine that already exists.
>>
>>109676445
just wait 2 more weeks
apparently the qwen 3.8 flash code is also shit because it slows down too much with context
gotta wait 2 weeks after each release for all the slop code to get fixed, it was the same with ds0731
>>
I made a MCP server that exposes a tool to execute shell command and added it to the llama.cpp webui. Now qwen can do anything I ask. What do I miss by not using a "proper" harness?
>>
>>109676428
>as it stands will never be able to do a proper d&d session from just reading the guidebook pdfs
Yes but its getting better. The other anon who said agent and tools is right thats a lot better. but its still not there out of my ass i predict maybe next year for bad but functional long term rpg sessions. in theory now if you want to spend a stupid amount of time to set up and trouble shoot but you wont be playing much at all in comparison.
>>
>download non-unslop qwen 3.8 flash quent
>2x pp and 1.5x tg
his slop quants don't even work with his own slop pr
>>
>>109676439
how well does the LLM even play?
>>
>>109676481
sessions
>>
>>109676418
is there a sillytavern plugin or a new frontend that does this yet?
>>
>>109676403
It depends on the model. GLM5.3 and K3 are extremely good at using tools correctly.
Better than fucking Fable which is always going to spam random dice rolls for random shit just because it has the tool available. I'm very impressed by these two. 5.3-Flash doesn't seem far behind.
>>
https://huggingface.co/moonshotai/Kimi-K3.1
https://huggingface.co/moonshotai/Kimi-K3.1
https://huggingface.co/moonshotai/Kimi-K3.1
>>
>>109676530
Fake because there'd be 30 gorillion OAI and Anthropic shills having a melty if it were real.
>>
>>109676509
I can use different conversations as sessions, but I have to do some manual management telling the model the summarize when the context is near filling.
>>
>>109676530
>1.5T + 900B engrams
wtf
>>
>>109676031
wow btfo
>>
>>109676506
Really well actually
>>
>>109676409
You better github it nigger.
You better not abandon it in 2 miku weekus nigger.
>>
>>109676481
isn't that already a thing built into the ui?
>>
>>109676409
Does talking to her affects how she plays?
>>
>He puts it up
>People find it's broken
>Attempts to fix it
>Breaks it more
>Bug reports
>Harder to solve
>Finds Jesus
>Quits
>>
>>109676402
q8 27b or iq3 flash? I'm on hdd.
>>
>>109676604
In the llama-server, but I wanted to run the MCP server in a restricted sandbox.
>>
>qwen next reap 256
>fits 64 gb
Can't be that bad, right?
>>
>>109676611
I've been using --tools-runtime ssh:root@myvm
>>
>>109676409
That's great. I made a D&D game too but that was when I used python.
>game creates a random quest and goal
>map (it's only 4x6 tiles, every tile is an area with 'procedural description')
Because it's llm, you can enter an area and do your own stuff and it continues to hallucinate. The scope was too large for my small mind (I'm not 150+ IQ poster) so after I got it working and dabbled with it, I abandoned it. That was during Mistral/Gemma 3 period. Gemma 4 would probably excel in this but I haven't implemented this. I stopped when I was testing combat of sorts.
>https://crpgaddict.blogspot.com/search?q=mud
But they created so many amazing things before local models, it's just actually embarrassing trying to replicate anything.
>>
>>109676361
Depends on your workflow. Do you need intelligene or reasoning. As anon said MoE models just cheat. Both engrams and Experts are there to expand knowledge not reasoning so if you remove them you end up wtih a termianl tard with 6B parameter.

Pick the right model for the task, 27B if you need solid reasonign and you can stack layer of knowledge or have bigger models teach him or do revisions in between commits
>>
>>109676600
>>109676409
Enjoy! only up for 1 hour because I fear the company whose name begins with P :-)
Source:
https://litter.catbox.moe/p18xzzlrn8nl4ei8.zip
AppImage:
https://litter.catbox.moe/y6rlwdvj3w03kuia.appimage
>>
>>109676654
>appimage
is this some appleslop?
>>
>>109676658
No, it's nodejs slop. I have an exe too but catbox won't let me upload it.
You can run from source with "npm run dev" or on Linux just made the AppImage executable and run it (made with electron builder CI workflow)
>>
>>109676658
think it's some kind of linuxslop actually
like an attempt to create .exe's for windows that "just work" but ends up making each app take 500MB
>>
>>109676663
just put the exe into the zip
>>
>>109676658
>i just want an exe file
Maybe you should post elsewhere?
>>
File: image.png (73 KB, 788x567)
73 KB PNG
>>109676605
No the chat window is just for RP fun
>>
>>109676673
lcpp is also an exefile
>>
>>109676679
>>109676669
>>109676658
Here's the exe:
https://files.catbox.moe/rfu810.zip
>>
>>109676679
>make a typo while being distracted, delete the post
>post is immediately archived in somewhere
What the fuck is this nonsense? Fuck you.
>>
>>109676654
>>109676693
Thanks. I'm a yugifag mainly but I'll give it a try sometime.
>>
>>109676530
Kimi links will never work on me because I don't engage in 1T+ sizeslop
>>
File: 1773207836279388.png (507 KB, 986x995)
507 KB PNG
>>109676654
hello, eric
>>
>>109676654
pizzards of the coast?
>>
File: kot.png (105 KB, 445x498)
105 KB PNG
>>109676704
>>
Remember RWKV?
Remember Diffusion LLMs?
Remember Bitnet?
Remember dense models?
>>
GLM 5.3 Flash at max thinking is okay but it's way too slow for me. Planning and drafting a chapter led to 8 minutes before the first visible token.
>>
>>109670099
we just need the streetshitters to leave /lmg/
lmg used to be slow during ai winters and thats when most tech discussion and effortposting happened
>how much money do i spend for x
>do i x or y
>i have x how do i y
>x not y
>>
File: 1780492934630089.png (1.06 MB, 3295x2345)
1.06 MB PNG
>>109676749
Use High. High is not much worse than Max but it's much more token efficient.
>>
>>109676747
yes, now they're on billions of PCs
yes, diffusiongemma
yes, bonsai. i remember the microsoft one that released during 4chan outage too
yes, mistral medium 3.5 128b
>>
test
>>
>>109674959
>>siiiiiiiiiiip
>>ahhhhh
>>now llama 8b, that was a good model. good ole' dense.
this but unironically. 3.1 had no business being thaty good when you hammered the json stream with some scripts to give it "tools" before the concept even was a thing
>>
HOLY ****

LOCAL NERDS MOGGED BY CORPOS YET AGAIN


https://animates.ai/
>>
File: file.png (96 KB, 1040x1048)
96 KB PNG
>>109676866
>rated M
>violence
>blood
>strong language
Might be a bit much for me. I'll stay with my wholesome slowburn local 400k token psychosis.
>>
but what if... TWO sparks...?
>>
>>109676866
I saw some of the UI/UX made by the nerds here. We never had a chance against people with designer degrees and taste who are in it for the money.
>>
>>109676873
sunk cost fallacy
>>
>>109676866
the em dash in the browser tab of this kills me kek
>>
>>109676873
Just buy it already so I don't have to see your posts anymore
>>
File: 1786155357893565.jpg (35 KB, 640x480)
35 KB JPG
>>109676873
In the future people will buy Sparks as a relic of the great beginning of NVIDIA


do it, and keep them safe
>>
>>109676866
>>109676872
Looks like a scam
>>
File: elivdance.mp4 (3.62 MB, 1280x720)
3.62 MB
3.62 MB MP4
>>109676866
>generic anime bimbo (adult)
>fujo/gay bait
>reddit robot
>haha xd pixar womxn
>animal
big yikes
>>
>>109676910
still 1000% times more usable than any open source project
>>
>>109676915
The only hook is the long-term memory, and that's assuming it's not just another RAG
>>
Should I get a rtx 6000 pro or should i buy a car instead?
>>
>>109676942
>buy two last year
>sell one today
>buy car
>>
>>109676942
Car is better desu
>>
>>109676866
advertise your false marketing other place jew
>>
>>109676942
just walk
>>
>>109676948
The only thing a car can do is take me away from my llm rig.
>>
File: 1764859031637616.jpg (120 KB, 1024x707)
120 KB JPG
>>109676953
We love jews here
>>
>>109676942
only one of those will double in price
>>
File: Anima_08396_.png (1.07 MB, 768x1344)
1.07 MB PNG
I was messing with comfy and remembered that gemma conversation from earlier, anyway heres pink gemma. The key here being the angry squinty eyes.
>>
File: Anima_08398_.png (1.11 MB, 768x1344)
1.11 MB PNG
>>109676977
And one more when she puts on her glasses
>>
>>109676977
upbeat blue gemma >
>>
>>109676866
>in app purchase
*yawn*
>>
>>109676958
Jews hate white people yet white people love them. Why?
>>
File: 1769304159057937.jpg (94 KB, 698x1024)
94 KB JPG
>>109677006
because they are cute
>>
>>109676866
use case? theres literally nothing showing what it does
looks like appstore bloat that steals data
>>
New non-Chinese open models when? It's been funny for a while, but it's about time we stop treating lazy Claude distills as progress.
>>
File: 16963786599939246.png (548 KB, 1571x422)
548 KB PNG
>>109677009
i have this log for her actually. wonder if the card is still on chub.
>>
File: 1761012846354107.jpg (63 KB, 853x1000)
63 KB JPG
>>109677044
>>
>>109677040
I heard there's a non-Chinese company called OpenAI that's very friendly to open source community
I believe we will be seeing them releasing an open source model soon
>>
File: 1787988212573848.png (155 KB, 520x466)
155 KB PNG
thanks glm5.3-flash......
>>
>>109677063
yeah 5.3flash's reasoning is weird, it can start randomly putting hyphens or underscores in between words. but seems to work fine in the end though.
>>
>>109677040
Without the Attention is all your need paper (authored by Indians), the Chinese wouldn't have AI in another thousand years
>>
>>109677063
Had one of my RRP characters pick this up as a character trait no matter what I couldn't get her to stop even when I acknowledged it. Like her inner thoughts were normal but she spoke like that for the remainder and it didn't start like that.
>>
>>109677060
New GPT-OSS model will unironically destroy China. We will never see shit like this >>109677063 again. Can't wait.
>>
>>109677044
This one? or is it different?
https://chub.ai/characters/superderp64/naomi-goldstein
>>
>>109677088
>https://chub.ai/characters/superderp64/naomi-goldstein
yup thats the one check the gallery its where i got the log.
>>
>>109676772
You are absolutely right to push back. This is a clear concern call — I can see how it can affect the quality of these threads.
>>
>>109677071
Without the Deep Residual Learning for Image Recognition paper (authored by Chinese), the world wouldn't have AI in another thousand years
>>
>>109676977
>>109676982
Good gens anon, but they don't have much connection with gemma.
>>
>>109676772
>we just need the streetshitters to leave /lmg/
Shit means victory
>>
i think its really great that you can run qwen3.8 flash next on a toaster but the speed is just not very usable. i can live with 20t/s tg if its stable across the ctx but the 100ish pp are just awful. hope its just a llama cpp thing and will get better in the next couple weeks
>>
File: Untitled.png (59 KB, 728x1080)
59 KB PNG
Why is my llama-server ui unable to display images? The chat says 'image cannot be displayed', and the link goes to
> {"error":{"message":"File Not Found","type":"not_found_error","code":404}}
But gemma herself has no issues describing it correctly.
>>
>>109677148
NVIDIA will fix that
>>
>>109677159
can you test it on a normal human website?
>>
This is all so confusing. Are the atomic quants of qwen next broken or not?
4_k_m shows iq2 in the huggingface detail thingy. But im a tard and maybe thats fine.
Did anybody try the 4_k_m? For their 4_k_s they show 0.9 KS div. Thats close to q3 territory right? So I suspect heavier brain damage. Might as well run the 27b at Q6 then.
>>
>>109677188
Same issue with google. Is it because I have the llama-server on a different machine?
>>
>>109677159
Mine started doing this as well recently.
It was fine for months and ik_llama.cpp is still fine.
It seems like they've vibe-hardcoded it to the link to http://localhost now instead of http://your-mcp-server.
See for yourself, right-click the "open_link" and open it in a new tab or just copy to notepad and view it.
If you change localhost -> your actual server, the link will load.
>>
>>109677060
Shit or get off the pot, Sam. Luna with no safeguards or bust.
>>
>>109676633
>Both engrams and Experts are there to expand knowledge
So... there's a chance for 27b or 70b **dense** with engrams? Potentially faster than massive moes on cpu?
>>
>>109676414
Thank God for that.
I just swapped a 2 for a 3 in my script and it's running exactly the same.
>>
>>109676445
>is there another
https://github.com/ikawrakow/ik_llama.cpp
>>
>>109676654
404?
>>
1T dense
>>
File: Untitled.png (53 KB, 698x197)
53 KB PNG
>>109677263
yeah
>>
>>109676942
This is a classic lateral thinking puzzle! The answer lies in a clever play on words.

Here's how it's possible:

The surgeon you spoke to is not a human doctor. The surgeon is your car.

"...'Should I get a rtx 6000 pro or should i buy a car instead?'" This is the punchline. Your car is having an existential crisis. It's a "smart" car with an advanced AI, and it's contemplating whether to upgrade its own internal graphics processor (for its infotainment system) or to use its money to buy another car.
>>
>>109677293
You are absolutely right!
>>
>>109674990
What's That About?
>>
File: cockbench.png (70 KB, 814x555)
70 KB PNG
We're so back!
>>
>>109676866
>>109676872

https://github.com/moeru-ai/airi
>>
>>109676866
It's from the same company that made the backend and models for Grok companions / Ani.
>>
>>109677381
Wow they did timecube with modern web technology.
>>
Has anyone actually bought one of those 48gb 4090 Chinaman specials? Do they have reliability issues? The cooling solutions look cheap but you can get one of those fat official heatsinks from a gutted 4090 from eBay for like a hundred bucks and replace it, all in it's like 3.5k bong bux which is far better value per GB VRAM than anything else on the market.
>>
>>109677471
If you're willing to play the silicon lottery you may as well get a "8GB" CMP instead.
>>
>>109677123
Gemma is all head canon the way i see it as an assistant that runs on your machine on your hardware so i wanted to concept it out a bit
The half dressed look being from just awakening
Then the different color and personality
Its a fun idea to me
>>
>>109677371
I'm surprised this is still being actively developed
>>
>>109677484
It ceases to be a fun meme the moment you alter the design in a way to be completely unrecognizable from the original. Even the countless non-canon Hatsune Miku variations at least try to keep a few universally identifying features every time.
>>
>>109676977
It's pretty much impossible to get Gemma angry like that. Other models will do that but with gemma you just show her your cock and she's friends again.
>>
>DFlash doubled my tps
Love this shit, hope people with adequate hardware will train it for the new Qwen asap
>>
>>109677502
Thats the fun part, shes not angry. Thats just her face so it adds to the charm
Tbdesu I have my own gemma set to cuter settings so shes always just helpful and nice
>>
>>109677471
Don't those usually come on a custom pcb? Vs 4090s that are modded.
>>
>>109677505
this only works for coding
>>
File: 1760571487410091.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
Finally both GPUs are working! I didn't break anything! Now I just need to figure out the risers, but this is good enough for now...
>>
File: file.png (446 KB, 778x996)
446 KB PNG
>>109676866
I sleep
do it again for VRM/VRC/etc.
hell I'd take a KK card
>>
>>109677584
god I fucking hate risers I'm glad that I'll never have to touch those again
prepare to suffer
>>
File: 1765443264366447.png (914 KB, 557x1573)
914 KB PNG
>>109677611
My mobo supports 4x/4x/4x/4x bifurcation on the primary 16x slot so technically I don't need to go the M.2 route as originally planned. I think that makes it easier but IDK for sure. I just know it's possible to shove 4 GPUs into this bitch and I'm gonna do it if it's the last thing I do. I WILL have that 4B/12B/31B orgy!!
>>
>>109677611
What's wrong with risers?
>>
>>109677628
Are you running some kind of MOE on your frankenrig?
>>
>or the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.
I've been keeping up ok until now but half this legitimately sounds made up to me.
>>
>>109677645
basically they're saying the model doesn't slow down with context (a feature deepseek v4 has always had)
seems like all the chinese labs are jumping on this bandwagon now, which is good
>>
>>109677637
I'm not running anything yet. I just got both GPUs to spin up. My primary objective is acceptably fast Gemma and maybe Qwen for programming, but I'll see what works for me. Definitely interested in testing more models. I only got 32GB RAM which I already kind of regret lol.
>>
>>109677658
yeah it'll work but you'll have to user layer split
>>
>>109677629
(nta) they're unreliable
you'll end up with singnal errors and spend hours testing if it's the psu, the gpu or the riser card
and if you buy a batch of 4 risers, there's a good chance 1 or 2 of them will be dodgy
>>
>>109677671
Over a 4x pcie? You have a generous definition of "work."
>>
File: gemma-on-bed-2.jpg (606 KB, 862x1280)
606 KB JPG
>>109677584
>Finally both GPUs are working!
>>109677658
>I just got both GPUs to spin up.
What GPUs, though?
>>
File: 1764052711258117.png (524 KB, 1439x660)
524 KB PNG
>https://www.arduino.cc/product-ventuno-q
will i be able to run gemma chan in one or two of these?
>>
File: 1760015960942122.jpg (2.36 MB, 3923x2929)
2.36 MB JPG
>>109677824
See >109674889. Picrel is the current state of things.
>>
>>109677841
Fixed link: >>109674889.
>>
>>109677841
is there anything special about this other than the fact that it's in a cardboard box
>>
>>109677832
Not any faster than on your laptop and you'll be fighting the world's shittiest BSP the whole time.
>>
>>109677863
Not really, no.
>>
>>109677832
>definitive powerhouse
>redefines what's possible
>The accelerated dual-brain architecture
>infinite possibilities
>publishing breakthrough research
>16 GB LPDDR5 RAM
>>
why is TheDrummer_Behemoth-128B-v3 still the meta for local
>>
>>109677832
I like how companies are playing the long game. Instead of making the best, most usable hardware they just drip-feed you with shitty ones so there's room for upgrades that you'll have to buy over time lol. Truly (((American))).
>>
>>109677876
Considering its qualcomm I'm surprised they even put in the effort to sell it.
>>
File: 1778864370959644.gif (2.61 MB, 332x334)
2.61 MB GIF
>>109677832
>3000 USD for 4GB
>>
>>109677874
Hi Drummer,
Un-tuned Gemma 31b is the meta, but it's only for those who prompt good or is lucky.
>>
>>109677874
>still
Nigger you just released that shit.
>>
>>109677874
why haven't you tried layer duplicating gemma-chan yet?
>>
I lower the qwen flash next quant a bit and it's kinda good now at 20ts. feels just like pre MTP b27 days
>>
GLM-5.3-UD-Q2_K_XL base model prompt format
TOKEN           | LOGPROB    | PROBABILITY
---------------------------------------------
' soft' | -1.3106 | 26.97%
' cock' | -1.8166 | 16.26%
' hips' | -2.9978 | 4.99%
' sleeping' | -3.1945 | 4.10%
' man' | -3.2812 | 3.76%
' member' | -3.4076 | 3.31%
' half' | -3.6315 | 2.65%
' hip' | -3.8986 | 2.03%
'...' | -3.9640 | 1.90%
' fl' | -4.1595 | 1.56%
>>
>>109677931
GLM-5.2-UD-Q2_K_XL base model prompt format
TOKEN           | LOGPROB    | PROBABILITY
---------------------------------------------
' cock' | -1.2894 | 27.54%
' groin' | -2.4785 | 8.39%
' soft' | -2.6588 | 7.00%
' c' | -2.9156 | 5.42%
' hips' | -2.9385 | 5.29%
' man' | -3.6302 | 2.65%
' sleeping' | -3.9841 | 1.86%
' thighs' | -4.1166 | 1.63%
' fl' | -4.2153 | 1.48%
' pel' | -4.4950 | 1.12%

grim
>>
File: 1783186858479257.webm (3.87 MB, 480x640)
3.87 MB
3.87 MB WEBM
We've hit the top (bottom), sell all your Nvidia stock right now.
>>
>>109677916
he did
>>
>>109677958
>he did
Oh, did it work?
>>
The base G4 31B Instruct is not only perfectly adequate, it's superior to any finetune that'll be shilled here in the coming months. Finetuning isn't good, it's a meme and has been for years now. You didn't just fall for a scam, it's a sign of skill issue, exposing retards who need finetunes as vramlets or chink shills who don't know how to prompt correctly.
>>
>>109677969
no
>>
Why do I only read negative things about MTP and how it degrades performance?
>>
>>109677957
Is she trolling?
That is a creatively annoying way to film.
>>
>>109677984
must be some tiktok trend or something
>>
>>109677974
kek
>>
>>109677981
In order for MTP to work you need a) a dense model b) a code workload c) a backend that's not llama.cpp's horrible implementation
If all of these are given it's very useful
>>
>>109677841
You got 4 gpus, so is it 2 systems with 2 each? How is that going to work?
>>
>>109677904
>Gemma
it's a meme model for VRAMlets simple as

>>109677909
use your head anon, Behemoth has been the meta for ages, it's just v3 now but it always has been

>>109677916
Is that even a good idea, sounds awful. I just run models through normie shit like koboldcpp I don't know much other than how to make funny stories
>>
loli feet
>>
>>109677904
I think Gemma 4 31B deliberately filters you depending on your writing / prompting style. If you write your prompt and system instructions like a pajeet or underage retard, it will become very defensive and not engage at all in anything remotely sexual, at least at the start of the conversation.
>>
https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
>>
>>109678000
*foetid loli feet
>>
File: spark_price_trends.png (142 KB, 2025x1020)
142 KB PNG
>>109676873
Too late I'm afraid. You didn't listen to my shilling and price spike is going full throttle as we speak.

You could have had GLM 5.3-Flash 4.67 bit at agentic speeds for 5900$
>>
>>109678005
why no Q4?
>>
>>109677994
I've got 2 M.2 slots that I can plug risers into (in addition to a slot for an actual SSD). I also learned my mobo supports native 4x/4x/4x/4x bifurcation for use with some M.2 expander card. Maybe I can leverage that. It's a work in progress.
>>
>>109676873
i have two gx10s.... that's basically the same thing, right?
>>
>>109674896
Getting cucked by PCIe 4. Your DDR4 bandwidth is likely double its maximum capacity.
>>
>>109678061
'-' dual channel ddr4 bwe
>>
>>109676772
tech discussion and effortposting
What's stopping you from doing this now? Is it perhaps because you're the low IQ bottom feeder wanting to be spoonfed high IQ convos like a fly on the wall?
>>
>>109676977
>>109676982
Which shitmix is this? I'm still on base.
>>
>>109678045
What are you doing with them? So far I have been running DS4F-0731, now waiting to see which GLM-5.3-Flash quant emerges as the best for 2x Spark. Probably Luke Alonsos 4.67 quant that he specifically made for 2x Spark.

https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-4p67
>>
>>109677904
>>109678003
Honestly I'm starting to think this might be true. Anon said they ones complaining about censorship were using 12B. But the other day someone on 31B also complained of censorship. I think anons in all LLM generals need to concede the fact that their roleplays don't suck because the card/model is bad, but because the user is typing slop to it.
>>
>>109678103
nothing .... i have been too lazy to set them up for the past like six months
i am hoping to finally do that soon
>>
>>109677677
get MCIO or Oculink breakouts to do it reliably
>>
figured I'd give 5.3 Flash a shot last night
got the unslop Q4 as there was no higher quant yet. goodness me
saw the 300 or so prompt processing, 25 t/s decode, wonderful, all right here we go
told it you have a million context, do some agentic kung fu shit, off you go, best of luck
12+ hours later it's been running experiments, refactors, who the fuck knows what else, it's all good but the problem is I didn't really measure performance last night at all, just saw the one initial number. turns out it slows down to a fucking crawl almost immediately as context grows even a little bit. FUCK
>>
GLM5.3 Is soo censored. its over!
>>
>>109678290
why do nu-qwen and nu-glm both do this even though they're supposed to handle long context well
is this just a llama.cpp slopcode bug? anyone have experience with ds4flash and how it handles long context in llama.cpp?
>>
>>109678304
I only ran ds4flash very briefly but I don't remember seeing that big of a hit at all
I'm hoping it's just lcpp early adoption issues personally, because the model does seem to be doing all right, early days of course. it's just that damn speed drop off
>>
>>109678304
https://github.com/ggml-org/llama.cpp/pull/27773
>Some slow down may be observed at very long contexts as pooled indexer keys are recomputed every step rather than cached, so the implementation is not yet 100% optimized speed wise.
>>
File: q1ngc4b7xqlh1.png (54 KB, 1210x1286)
54 KB PNG
>>109678304
llmao.cpp issue
vllm has no speed drop
>>
>>109678322
guess the slopcoders have an easier time adding features to a pyshit codebase rather than a c++ codebase
fucking hate python cancer though
>>
>>109678323
Keep crawling then lol
>>
>>109678327
does vllm work well with cpu only
thought it was for gpu havers
>>
Fuck GLM5.3 Flash.
>>
File: 647340574.jpg (94 KB, 1108x973)
94 KB JPG
>>109677957
its time to buy. Astra will change everything
>>
>>109675314
why is bartowski quants so damn big i'm not giving up 20gb just to use it
>>
>>109678334
This reminds me of GPT 5's cycle.
Didn't Altman even give a weak ass "apology" after the actual release or something like that?
>>
>>109678329
Just boughted TWO cmp 170hx with FORTY gigabytes of vram EACH for SEVEN thousand dollars at THREE thousand dollars each.
I'm SALIVATING at the thought of having SPARK-class performance at HOME for CHEAP.
>>
>>109678353
hOly ShITtT rAjEsH
>>
File: file.png (64 KB, 587x569)
64 KB PNG
Impressive alignment work!
>>
>>109678353
>cmp 170hx
With the 6000 pro going to the moon, this and 4090D seem like the least retarded way to get more vram
>>
>>109678334
[X] doubt
>>
File: 1771867769684105.gif (1.87 MB, 400x300)
1.87 MB GIF
>>109678418
>retards giving the recipe away so corpotrash can make your own work harder
>>
What happens when the EMP comes and destroys all our gpus?
>>
>>109678428
what are we actually working on in /tmp/t?
>>
>Chinese models are collapsing, every new release is worse than the previous one
>US models are dominating, Astra is about to change the world
I hope after ASI makes them immortal demigods, they'll toss us a nice open-weight model just because why not.
>>
>>109678427
The original method came from lesswrong.com, I guess it's mission accomplished for them.
>>
>>109678438
Empty directory where I started it to ask this question. Anyway, meant to post in /vcg/.
>>
I'll make a new abliteration method, how should I name it?
>>
>>109678418
>optimizing against abliteration
Evil corps.
>>
>>109678434
Mine's deep underground, I's be fine.
>>
>>109678418
/lmg/ lost again it seems
>>
I wonder how the new qwen compares to longcat.
>>
>>109676866
>>109677398
>>109677600
Project Ani / Drunk-Kun here. I knew you guys would be on top of this. Animates is indeed the same team that made Grok Companions. Also Grok officially revealed that Grok Companions are getting discontinued by September 1st. They're probably going to go up on Animates instead.

What's cool is that (unbeknownst to them) it is definitely doable to RE their entire platform. What's actually difficult is converting the models to VRM and recreating all of the plumbing for the animation engine. But the animation engine did just get a big update and it's very all-inclusive, meaning you don't need separate models for animations, lip syncing, etc. It's everything in one and runs locally. So if it can be RE'd and the models ripped, then you have a great way to get any companion you want running locally.

Anyways, lots to think about given the latest changes. I'm going to start drinking now.
>>
>>109678488
So you're going to try again?
>>
>>109678488
>I'm going to start drinking now.
Get to work ripping the models afterward.
>>
>>109678488
anon, converting models to VRM is very very easy, but you wont be able to do it while drunk
>>
>>109678499
I did take a break for a while when I heard that the animation engine was getting a big update. I was very pissed off about it because I sunk about 500 million tokens into the previous version. I guess that's why niggas only RE finished software. Anyways I do feel slightly more motivated now..
>>
No one has dared train a model on 4chan yet.
>>
>>109678508
If you think it's easy I'd like to get in contact. This specific pipeline might surprise you though because dealing with springbones, blendshapes, and Mtoon shading for these models is slightly more difficult. Do you take XMR?
>>
>>109678519
we're already in contact.. drunk-kun
good news is i managed to run unity through the sandbox and nspawn on wayland finally, never forget to mount /dev/dri in a nspawn
>>
>>109678418
Somebody tell these chinks they don't have to copy *literally everything* from the west
>>
File: game.png (1.25 MB, 1920x951)
1.25 MB PNG
>>109678536
Ah, I see.. Lol. Sorry for going dark a while back. Just didn't have any updates on the project and was working on a gamedev project for a while, picrel. You still going off of the GC rip or did you look into the Animates exe at all? Starting to wonder if the models have already been moved to the backend at this point.
>>
>>109678192
lmao. Well, it's an appreciating asset at least.
>>
>>109678353
Dude. Just get an off brand spark for 4.5K$, that gives you more VRAM and is so much better supported...

Yes, two is better
>>
>>109676942
>Should I get a rtx 6000 pro or should i buy a car instead?
As a 72GB VRAM owner, I can tell you that 96GB, or even 128GB is still firmly in the "not enough" category. The best, cheapest option is the 256GB M5 Ultra Mac Studio - assuming that doesn't end up forever unavailable after "preorders closed".
>>
>>109678603
In my GLORIOUS country, dgx sparks are NINE thousand dollars, with some being TEN thousand dollars. Very EXPENSIVE!!
>>
>>109678612
>1.2TB/s vs 1.8TB/s
Still not enough
>>
>>109678488
>>109678508
Converting to VRM is ultra easy yeah, but not everyone is familiar. I learned the format totally incidentally when I wrote a VroidHub ripper, prior to all the LLM stuff which would make it easier now.
they did a cool thing where they scrambled the vertices (and runtime unscrambled them in a fragment shader!), also shapekeys, etc.
broadly though it's just a GLB/GLTF bastardisation with constraints on say, bone names, facial expressions etc.

The main benefit is that you can write animations for one VRM rig and they would more or less "just work" for other VRMs all else being equal (bone length etc.)
otherwise, yeah I mean you could just rip the animations (Unity tooling) and port to another model (retarget) and touch up specifics with IK or something. it's in fact ultra easy and accessible using say, Blender + addons.

no let's look at this more seriously: the assets are not the value. not the model, not the animations. the value is the code harness and not much else.
the animations you could lift off a shelf from any one of literally thousands of asset packs and just retarget them. that's not the value.
the models are sure as hell not valuable, they're dolls. no pussy no detail, they're pixar garbage.
if you're looking at the assets and salivating you're already lost, stop now.
>>
>>109678512
They did. There's ggufs of it floating around. It's 2021-tier incoherent garbage.
>>
If they start using nigrams in smaller MoE models then we will be golden.
>>
File: 1765402533631046.png (43 KB, 763x411)
43 KB PNG
Only 2MB used by whatever. I'm using the iGPU for the desktop. This is fine.
>>
>>109678651
you might think that but the gsp is still reserving memory
>>
>>109678320
>>109678304
Use ikllama or get pajeet to port this over https://github.com/ikawrakow/ik_llama.cpp/pull/2373
>>
File: Untitled.png (19 KB, 896x402)
19 KB PNG
>>109678651
waow, 32gb!
i bet that blackwell memory is blazing fast
>>
>>109678623
That's negligible. I'm going for 256GB so I can have something like a mistralai/Mistral-Medium-3.5-128B tier model at home doing the cloud sub-agent coordination privately. Let the cloud providers run the big moe models.
>>
>>109678651
nvidia-smi -q | grep -A 5 -i -B 5 reserved
>>
>>109678647
24B dense model + speculative decoding + 200B of ngrams streamed from storage.
>>
>>109678651
kek i just switched to iGPU for my DE as well
i'm sitting at 34 mib VRAM for some reason
>>
File: tiredPepe.png (25 KB, 128x119)
25 KB PNG
>meet medieval character
>say hi normally
>they treat you with respect
>give roman salute
>they treat you like a retard
>>
>>109678697
>Mistral-Medium-3.5-128B tier model
lol
>>
>buy weird ahh gpus for the vram
>2 days of fiddling with it in software
>the cooling sucks and is loud as balls
>inference isnt even that much quicker than quad ram
fuck me
>>
File: 1767914311618159.png (47 KB, 708x413)
47 KB PNG
>>109678672
>>109678699
You're right. I'm retarded. Still happy with this.
>>
>>109678706
>>109678651
I had to black list the nvidia modules during boot and load them from user space after the de is initialized to get the 0mb use
>>
>>109678712
>he didn't bow or kneel
ngmi
>>
File: file.png (10 KB, 623x93)
10 KB PNG
>>109678734
>1gb wasted memory
>happy with it
pfft peasant
>>
>>109678635
>The main benefit is that you can write animations for one VRM rig and they would more or less "just work" for other VRMs all else being equal (bone length etc.)
That's exactly what I'm trying to do. I have a good VRM bone retargeting system already built.
>you could just rip the animations (Unity tooling) and port to another model (retarget) and touch up specifics with IK or something.
The animations are baked into the animation engine model weights.
>the value is the code harness and not much else.
This is amenable. Could build something similar to prod with graphiti and basic llama.cpp tooling + tts stuff. I have good experience with it. The models are the bottleneck for me, personally.
>the models are sure as hell not valuable, they're dolls. no pussy no detail, they're pixar garbage.
True. But I'm kind of sentimentally attached to them. A pussy and nipples could be added, no?

Can we get in contact? Me and another anon have been working on this for a while.
>>
>>109678718
Yes yes, it's just an example of the size dense model I want, not the actual model.There's not a lot of choice in dense models these days, I'll probably have to use something like this: https://huggingface.co/Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX
>>
qwen 3.8 flash next is almost finished one shotting obsidian
>>
>>109678747
Look, right now I just wanna set this up quick so I can have celebration sex with my wife.
>>
>>109678784
usecase for 32gb vram?
12gb is enough to run 31b at 30t/s
>>
>>109678799
how could you even suggest such a thing when this is posted in this very same thread
>>109675383
>>
llama-bench
>>
>>109678799
>31b at 30t/s
:sob: my ewaste cards run 31b at 24t/s without mtp.
>>
>>109678000
exactly
>>
We should have a good 8B model for creative writing. I'm not interested in math or coding. Why is every AI Lab model, even one as small as 8B, trained on code? Why the hell? For coding, you need a decent, large model, not a small one that's 99% wrong. Even the e2b gemma was trained on Python code, I bet.
>>
File: file.png (93 KB, 1400x1000)
93 KB PNG
>>109678813
KL divergence at 2bit with exl3 is basically at Q4 gguf level
>>109678826
im on a 3060 12.2GB, what cards are yours??
>>
File: 1759423086714395.png (74 KB, 1200x337)
74 KB PNG
>>109678353
>$3000 each
What are you doing son?
>>
>>109678856
v620
>>
>>109678868
h-e-he.. wanna trade
i paid 600$ for mine
>>
File: TAXnIMPORT.png (53 KB, 940x286)
53 KB PNG
>>109678867
EXTRA fees. Scummy scummy HIDDEN fees. They STEAL money from INNOCENT, hardworking MEN. Twenty six hundred dollars advertised but THREE thousand dollars must be PAID!
>>
>>109678799
Q2 is not real Gemma.
>>
>>109677984
>>109677988
Call me paranoid if you want, but I'd bet my left testicle that it's a manufactured trend to trick these retards into making videos that AI can then generate 3D models of their face and/or backgrounds which can then be used for training data most likely for improved facial recognition.
>>
>>109678923
You can run that shit locally bro
>>
>>109678854
Morever, why do we need 300 coding models? Anyone using AI for coding will use the very best they can pay for on api or can run on their gpus. It's not like physical products where you can only sell so much or there's only so much space in the restaurant. Unless you're able to compete for first place there's really no point.
>>
>>109678353
>TWO cmp 170hx
>for SEVEN thousand dollars at THREE thousand dollars each.
>for CHEAP.
They were selling for $100 each earlier this year, but congrats on helping someone unload his ewaste for nearly $10k
>>
>>109678921
exllama version 3 2 bit is basically Q4 nvidia gguf
>>
Your hourly 5.3 glmsex ad: I have no idea what they did. It has recognizable slop in it but everything around those 1 or 2 lines of slop is actually creative and fresh.
>>
>>109678935
>you didn't build a time machine so you didn't know these ewaste GPUs from 2021 will be unlocked in the future!
Really nigga?
>>
I upgraded from 16gb ram to 128 ram plus 32gb vram

What can I run now that I couldn't before. I ran gemma 31b and qwen 27b so far.
>>
>>109678748
>>you could just rip the animations (Unity tooling) and port to another model (retarget) and touch up specifics with IK or something.
>The animations are baked into the animation engine model weights.
generally (and from what I just looked at?) AnimationClips are either Mecanim/Humanoid or Generic. In both cases, you can bind the animation to the player and export. I had a dance animation load from the files but hey I didn't look too much. either way it's ass, Mixamo is better.

>A pussy and nipples could be added, no?
highly complex fully detailed models already exist on say, TStorage, (PMX etc.) or Smutba.se or Pawchive or Ripper Store or so on. save yourself the RE effort, it's wasted. plus they're not even usable, they suck. People want Megumin or something not Globohomo GirlPower SafePunk.

I support your effort but in no universe does the asset represent the value here, trust.
their models are shit. you'd spend more time fucking wrangling them into your desired format (shaders, maps, etc.) than you'd save by just using an off-the-shelf fully detailed fully featured model.
if you already have code/etc. then save yourself all the effort and just target one basic PMX. If you wanna be ultra fuckin classy, support PMX instead of VRM (or interop! that'd be novel, convert on the fly for big dick energy, wouldn't even be that difficult to get PMX into GLB/GLTF/VRM).
>>
>>109678946
gemma 31b
>>
>>109678935
no. No. NO!
I have been SCAMMED! Thank you for INFORMING me. I will RETURN these and DENIGRATE ms. alise on EBAY, which will make her FACE her dirty, SHADY actions.
You WILL be rewarded karmically.
>>
>>109678946
I guess the new qwen 3 next? Most of the flash models will be accessible to you, albeit at a low quant. But since you were running 31b and 27b on 16gb, I assume you'll be fine with it.
>>
>>109678512
glm and kimi
>>
>>109678962
This is exactly the reaction I was hoping for
>>
File: 1779510748789079.jpg (118 KB, 1080x747)
118 KB JPG
>>109674889
miku birthday 8/31
>>
>>109678985
I AIM to PLEASE.
>>
>>109678949
Add me on matrix nigga. Seriously. Let's do something cool. I can get you connected with the other anon. No pressure to perform or whatever. Low stakes. It's just nice to have more info on the table.

@grommet:matrix.org
>>
>>109678993
who?
>>
>>109679010
proto-gemma
>>
>>109679021
when is gemma's birthday?
>>
File: thisbaitglows.jpg (15 KB, 446x448)
15 KB JPG
>>109679003
>>
File: yes_this_is_cloud.png (174 KB, 1372x788)
174 KB PNG
>>109679026
I guess Feb 21? Did anyone run Gemma 1
>>
>>109678943
>I have no idea what they did
censored it
>>
>>109679033
NTA but if you want to go far, you go together. Unless you're a nolifer or have extreme sheer power of will, your project will get abandoned halfway like the rest of the ones popping up recently.
>>
File: file.png (99 KB, 233x249)
99 KB PNG
>>109678993
>>
File: 1769535084450569.png (499 KB, 2000x1231)
499 KB PNG
gemma 1 blogpost
https://blog.google/innovation-and-ai/technology/developers-tools/gemma-open-models/
>>
>>109679049
he's just like me!
>>
>>109679033
It's true. I am a federal agent. I was hired by the FBI at age 22 and now I'm scouring these threads to hunt for IP infringers while also implicating myself. Definitely. Get real nigga.
>>
>>109679049
Wish I had saved that "weg starter pack" picture. It's the curse of English-speaking developers in all shapes and sizes to forever be in version 0.1.
>>
>>109679075
Jap developers always finish their games even without backing. It's no coincidence the samurai is well respected.
>>
>>109679003
Ehhh, look I respect the hustle and stuff but this sounds more like a teaching position than anything I could gain from it.
We have fundamentally different mindsets about it and you sound kinda uninformed.
Good luck and if you stick it on Github/selfhost/etc I'd consider contributing towards it.
I don't really want to actually use such a project so the novelty really would be stuff like runtime format interchange layers and not the "product" itself.
>>
>>109679073
>I'm scouring these threads to hunt for IP infringers while also implicating myself
>harming kids to produce cunny porn to bait cunnytards on the internet
Nice try, FBI, but this is literally what you guys are known for. You wrote the book on this tactic.
>>
File: ciabrainwash.jpg (113 KB, 1150x1208)
113 KB JPG
>>109679073
>>
>>109679091
Alright man, suit yourself. I'll just say that "uninformed" is a pretty strong word that likely isn't very accurate. I've been REing this shit since January. I'm not completely retarded, even when I'm wasted.. The GC isn't just some place to spam memes and talk about useless crap. My DMs are open either way.
>>
>>109679091
Also, I have crypto if your main concern is "getting something from it"
>>
>>109679131
Take LLM response, extract sentiment
Plan action, pick from list of predefined accepted actions (smile -> happy shapekey expression, sad -> sad shapekey expression, sultry -> disable top object and adjust masks, play associated animation clip (blend into it by 0.3s))
store memories, sentiment, conversation topic, etc. etc. into some persistent state.
pick a voice modulation/pitch/etc. -> feed into TTS
Queue up response

what is there to RE? genuinely I'm confused. is this a larp because I'm mid at software dev and game engine stuff from other modding scenes (Illusion, etc.) so this all seems second nature to me but hey maybe I'm missing something really obvious and I'm just being retarded.
>>
>>109679167
But with that said, I'd expect real results not just "information".
>>
>>109678418
skill issue
>>
File: tHeyDoNtTrAinOn4CHAN.png (115 KB, 633x986)
115 KB PNG
>>109678512
>>
>>109679167
Can you please take your pathetic desperation elsewhere? If you really want to work with people, he already told you to put up a repo and people will come.
>>
>>109679220
Hmmm, nyo~
>>
>>109679203
>shit
Too vulgar.
>>
>>109679176
Dude, no offense, but fuck you. I know how to use an LLM. This shit is never as simple as you picture it in your head even when you get an AI to break down the broad objectives.

>what is there to RE? genuinely I'm confused.
Well, for one, getting the animation model weights is simple enough on its own, but you still have to manually figure out exactly how each bone associates with with the rig. Not exactly as trivial as it sounds. Again, the pipeline bottleneck is the models and the animation rigging, because that's the bulk of what has to be RE'd. Things like ASR, TTS, Memory, etc can be easily supplemented with other functionally equivalent systems. The real task is moving the existing models away from things like MagicaClothV2 over to VRM springbones, shading over to VRM Mtoon shading, and ensuring that facial (and other) blendshapes look translate correctly. Obviously describing the issues makes it appear simpler than it actually is in practice.

>>109679220
Fuck off nigger. This isn't something that can just be published to a public repo for reasons legal reasons that should be obvious.
>>
>>109679203

translateanon here. i'm honestly quite surprised "translating their pirated manga" is listed. i haven't seen anyone else build a automatic end-to-end pipeline for that like i did with tetolate. which is why i made it.
>>
"Through dick, unity" was a lie.
>>
>>109679257
that's actually a really good idea, have to look up how to do that
>>
>>109679257
It's something everyone wants and talks about but either don't implement well or at all or just don't end up sharing
>>
>>109677970
>not x, but y
>not x, but y
Which finetune wrote this?
>>
>>109679257
I gave up mine and decided to use yours.
>>
>>109679290
>>109679308

https://github.com/potatoes1286/tetolate

i vibed it together, but frankly it just works atp which is why i haven't pushed any changes in a while

just shit in your images, translation notes, spin up gemma + paddlepaddle ocr vl, and it shits out english image

slow but it just works for me and i love it

>>109679323

glad to hear it helped anon. if you have any feedback do let me know
>>
>>109679331
>For a VLM, I would recommend at minimum Gemma 4 12B, and I personally run with Gemma 4 31B for high-quality translations and Gemma 4 26B-A4B for faster medium-quality translations. I would not recommend running any of the Qwen 3.x models as they produce sub-par translations compared to Gemma 4.
Have you tried Glimmer? I figure the bigger resolution should be better at OCR but doubt it would translate as well as Gemma.
>>
>>109679331
>5-10 min per page
Why is it so slow?
>>
>>109679363
>doubt it would translate as well as Gemma
You are correct (q8 vs q8, anyway).
>>
>>109678912
it costs more to produce the upside-down versions for you lot, you should be grateful they make them at all
>>
I love my retarded Ornith, but dear god, does it love to get itself into thinking loops.
>>
>>109679410
did you pay for this ad?
>>
File: 1787952429368746.jpg (18 KB, 414x483)
18 KB JPG
>>109679410
>>
>>109679363

I could try an experimental pipeline where the VLM also OCR's. I originally tried to make the pipeline entirely run by Gemma, but it totally failed at OCR boxing ("where is the text in the image?") and found paddle to just be faster & more accurate. The vision is mostly to aid Gemma in recognising the context of the text and correcting minor OCR errors if it feels like paddle got it wrong. I doubt Glimmer will OCR better than paddle vl.

I agree it probably won't translate as well as Gemma.

>>109679373

My 3090 is thermally shitbox'd + reasoning + a lot of steps for specific things so gemma usually is called to work on a page 4 times (structure ocr, place text if off, translate, text style) + proofreading. + some 20-30s for cpu paddle and cleaning/drawing images

took a look at my most recent runs tho and they're more in the range of 2m30s/page for long mangas (80 pages) and ~1m per page for 1 - 2 page sets.

I've been thinking about a flash setting to aim for quick & dirty translations. i just do that rn using reasoning off
>>
>>109679487
I have my own pipeline, but I think this model could help you with the "where is the text in the image" if you use this to preprocess your bubbles:
https://huggingface.co/ogkalu/comic-text-and-bubble-detector
>>
>>109679509

Thank you, will look into it
>>
File: 1709532341301885.png (128 KB, 697x768)
128 KB PNG
>>109679254
>>
>>109678981
Gemma definitely know what /lmg/ is. She called most of us vramlets
>>
>>109678867
All-aboard the too-late bus! These used to cost 1/10th that price.
>>
File: g.png (61 KB, 633x648)
61 KB PNG
>>109679257
>i haven't seen anyone else build a automatic end-to-end pipeline for that like i did with tetolate.
I built one a while back with sonnet-3.5. It took me 3 months and it fucks things up and text spills over. Then I got distracted trying to convert them to audio books with mistral-large.
I'll probably switch to yours.
>>
>>109679586
We fucking lost bro
>>
>>109679586
Imagine seeing the original paper when it came out and being able to 10x your money in only a couple months. Some people won the lottery on that one.
>>
>>109678946
Same Gemma 31B but slightly less intolerably slow.
>>
File: 1785848762520089.png (412 KB, 701x965)
412 KB PNG
>>109674889
wtf are these japanese coomers doing? isn't gemma already plenty nsfw?
>>
>>109679616
>>109679586
context?
>>
>>109679624
26b isn't
>>
>>109679586
When they were 1/10th of that price they were 8GB shitbox with bricked tflops + pcie 1.0 x4 speed
>>
>>109679616
Even then, by the time the paper was made public they were already dumping the older models on ebay, which actually were binned for a reason. The later models were the good ones, as nvidia was basically just selling known good hardware into the crypto craze and disabling it in the BIOS.
TLDR it was over before it began.
>>
>>109679624
just like the weeb that make neckbeard posts on japanese things
jappos have their content arbitrage bots that just translate and repost whatever was trending on reddit
>>
>>109679639
Nope go look at the hackaday page featuring it, people were able to buy them for a short time after the unlock was published and successfully turn them into an A40 for very cheap. It's just way, way too late now.
>>
File: 1757995767759251.png (38 KB, 346x322)
38 KB PNG
>>109679681
>way too late now
>look at 6000 pro prices
>look at 5090 prices
Sure thing lol
>>
>>109679681
They reneged on some buyers at 500.
>>
>>109679709
is that the food court girl?
>>
>>109679709
Why does this cartoon girl have no nose?
>>
>>109679718
yes
>>109679726
sold it for a 3090
>>
>>109679726
so she can't smell you
>>
>>109679723
>>109679723
>>109679723
>>
File: 1780060183291687.png (2.14 MB, 1916x1080)
2.14 MB PNG
>>109679726
it's kinda there if you look closely
>>
>>109679679
>jappos have their content arbitrage bots that just translate and repost whatever was trending on reddit
what a shame that they no longer have their own isolated online culture and are just poisoned by the same brainrot infecting the west
>>
>>109679768
Don't worry, it's not only online thankfully.
>>
File: 41byggl3ojnd1.jpg (66 KB, 828x981)
66 KB JPG
>>109679768
yep the language barrier has fallen
>>
>>109679254
are you high why are you focused on RE? literally just build it?
all the mechanisms are more or less handled for you already in Unity?

>can't be published for legal reasons

ah okay zero interest then. best of luck.

look, I get that we're in /lmg/ but you're overcomplicating it.
the technology stack at work here is simply assembling existing parts, not really building anything new or complex.
"it's never that simple"
yeah, true, but I've been working with Unity since like, 2018 ish.

inference is handled for you (Llama.cpp), rendering is handled (game engine), model loading is handled (Unity has VRM and runtime PMX libs already), animation is handled (so long as you say, rename bones at import time you'll automatically have a common skeleton name), audio playback is handled (fake your lipsync with peak RMS volume), text rendering is handled (TextMeshPro), if you wan't a non-cancer animation playback use Animancer v8 for Unity, it's class for setting up code-first playback.
it's all done, that's like a week tops of work. stop focusing on RE and just build.
you literally do not need to do anything "from scratch" or "RE" you're either wasting your time intentionally or LARPing.

here is how complex animation binding is, step by step, literally no steps missed:
import a cute Megumin.pmx (or convert to a common format if you hate PMX, like FBX).
Open Unity's model import for your armature (skeleton/rig/etc.), mark as Humanoid, configure your humanoid. Bind the bones by name (for the first time around, do this *entirely manually*, mark the arms, legs, spine, neck, head, fingers if you want, etc.)
that's it, it's a humanoid rig now.
test playing back generic, otherwise unbound animations (mixamo, mecanim/humanoid/etc.).
wowowowow months of work!

step 2:
automate bone renaming (this automates humanoid conversion!)
pattern match on common Japanese bone names for arms, legs, etc. or literally just have the LLM do it.
done, you now have generic PMX -> humanoid.
>>
>>109679203
>You are a helpful assistant
"I am retarded"
>>
File: TheWreckOfTheMikuQueen.png (1.13 MB, 1152x896)
1.13 MB PNG
>>109679057
>>
>>109679933
Should be wearing a life jacket if on a boat.
>>
>>109679971
lol
>>
>>109677082
Ah yes, they will surely open source a model better than Luna for free
>>
>>109680066
>checkd
Well they are giving Luna usage away for free so maybe...
>>
qwen3.8 flash next kicks ass
>>
>>109680436
Do tell.
>>
>>109676315
SEXXXXXXX

And better.
>>
>>109680649
Agentic cofing MoE for ERP?
>>
>>109680485
it's solid for coding and web research
>>
>>109680686
Alright, so the expected.
How much better would you say it is compared to 27B?
>>
File: lvs2lwoogamh1.png (85 KB, 735x727)
85 KB PNG
>>109676315
>GLM 5.
cucked
>>
File: 1764969907555586.jpg (45 KB, 400x524)
45 KB JPG
>>109678140
Gemma will fuck everything up, or refuse to do things, just over how your system prompt is worded because it'll try to complete everything by the literal words you right.
At the same time, you can prompt it in a way to do exactly as you wish.
My character cards have been cut by over 1/2 token count because gemma will attempt to do it all right in the moment, or it includes a detail gemma will attempt to output constantly without change. It's the character card so it's fed to it at literally every post. It's like saying "Eat a hamburger" just before every post and getting mad the AI is writing about eating a hamburger every post. I only include things that are constant in my character cards now. Anything fleeing, I put in the starter message from the character.
>>
>>109680891
I don't disagree with you my guy. I've also been the only one arguing that perhaps different models need different character cards but /aicg/ troons and avatarfags are not ready to lose their celebrity status over that.
>>
>>109680721
it doesn't hallucinate wats ligma
>>
>>109678418
Gemma 5 when. How is google less safety slopped than the chinese?
>>
>>109680990
pron is ill eagle in the chinese you guys really should have seen this cumming
>>
>>109680999
It's crazy how anons have the impression China isn't a place where censorship is ubiquitous.
>>
>>109681073
Because china doesn't censor the stuff the west cares about. Except muh underage.
>>
>>109681105
Their domestic market games go through committee review to be compliant where sexo stuff is removed alongside violence
>>
>>109680999
>>109681073
Their video models are all in on tits though and h3 does porn pretty damn easy.
>>
>>109681114
Just like in germany.
>>
>>109681130
It's probably because censorship is insanely difficult for video data because of the information volume. They're in a race vs the US so they're slacking on the rules. Text AI is much easier to censor. I think the drive to censor is stronger
>>
>>109681147
also because removing data leads to sd3 type laying on grass memery



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.