[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1735585594108694.webm (3.33 MB, 1184x1728)
3.33 MB
3.33 MB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109367209 & >>109370411

►News
>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e
>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B
>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares
>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta
>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
you just can't beat dgx spark cluster in value. you simply can't.
you can cope with any sub 100B retarded model, but dgx spark cluster is the best value now.
>>
>>109374440
fuck are you talking about, you could buy four intel arc pro b60 and call it a day on the cheap, I'm not recommending this but you could and it'd smoke that pile of shit
>>
File: HOF2JTDbsAAvayd.jpg (277 KB, 1254x1254)
277 KB JPG
>>
>>109374467
i bought 2 a770
don't do it
>>
How / Where i can find prompt to uncensor Gemma 4 12B (12GB) ? https://rentry.org/recommended-models
>>
File: growing that ram4.png (2.14 MB, 1024x1024)
2.14 MB PNG
Just grow your RAM in the field, like nature intended
>>
"For my first post, I’m sharing a letter @NVIDIA
signed on why open models matter." https://x.com/JensenHuang/status/2080643682408321103
>>
>>109374480
You can't just jump straight to DDR5, you have to start from the German Democratic Republic.
>>
As a white person, I have a right to be proud of llama.cpp.
>>
>>109374467
dgx spark cluster means more than one spark. do you run 8 intel b70?
enjoy your housefire, lmao
>>
I'm waiting for DGX Spark 2 Ti Super which will have double the bandwidth and capacity.
>>
File: 8.png (986 KB, 1600x1216)
986 KB PNG
>>109374530
I do.

>>109374477
my condolences, intel dumped that shit like a ginger step-child
>>
File: 1783038279105669.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>109374472
>>
File: 1779476430361497.jpg (149 KB, 1000x1000)
149 KB JPG
►Recent Highlights from the Previous Thread: >>109370411

--Critique of pi-go and discussion of a new local-first harness:
>109372783 >109372833 >109373974 >109373981 >109374033 >109374091 >109374168 >109374250 >109373458 >109374264 >109373115 >109373138 >109372814 >109372861 >109372928 >109373079 >109372930 >109373058 >109373200
--Model recommendations for Intel iGPU and debate over Unsloth quants:
>109371286 >109371302 >109371308 >109371310 >109371330 >109371344 >109371361 >109371390 >109371417 >109371444 >109371353 >109371384 >109371473 >109371446 >109371650
--Implementing and refining long-term memory systems for AI characters:
>109371012 >109371115 >109371157 >109371161 >109371205 >109371228 >109371251 >109371409 >109371781
--Debating compute sharing and DGX hardware efficiency for local inference:
>109373832 >109373859 >109374022 >109374031 >109373902 >109374006
--Comparing throughput and costs of CPUMaxx versus DGX Spark clusters:
>109374052 >109374072 >109374084 >109374194
--Calls for Google to release Gemma 4 100B:
>109370867 >109370913 >109370886 >109370941 >109370943
--Comparing Gemma's steerability against GLM and Kimi 2.7:
>109371774 >109371810 >109371835 >109372093
--Comparing Gemma variants to analyze alignment blind spots and reasoning:
>109370603 >109370625 >109370642 >109371893 >109371914
--Comparing MiniMax M3's creative performance and context issues:
>109372726 >109372740 >109372747 >109372764 >109372793
--Reaction to llama.cpp pull request adding MCP stdio support:
>109374270 >109374280
--Release of a 1.6B user autocomplete SLM:
>109372197 >109372201 >109374291
--Kimiposting:
>109374464
--Logs:
>109370603 >109370881 >109372197 >109372814 >109374193
--Luka, Rin, Miku (free space):
>109371122 >109371173 >109371253 >109371437 >109371541 >109371757 >109371934 >109372031 >109372636 >109372926

►Recent Highlight Posts from the Previous Thread: >>109370412

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109374608
I want to squeeze this migu
>>
>>109374479
Use a system prompt in text-completion mode, no idea for chat completion mode thats above my newfag head. The main issue Ive personally run into is that with reasoning on it will be much more likely to hit/abide by guardrails. So text completion mode + no reasoning + sys prompt works. If you are still hitting them after all that, you can turn to an uncensored/abilterated model, I cant remember the exact method thats ideal for reducing the quality impact but had luck with a heretic quant from SC117 on HF.
Im interested to hear if any other anons have advice on using reasoning, chat completion, etc as Im still trying to learn this stuff myself.
>>
>>109374547
How do you handle power?
>>
>>109374479
>>109374616
Sorry misread that you were asking for the sys prompt not an uncensored model, my bad. Ill let someone who actually knows what they are talking about chime in for that, my sys prompt was from a rentry not in the OP
>>
>>109374618
it looks worse than it is, it's not good by any means overall but

400/440w is "burst" or PL2 and isn't relevant to AI workloads

200w is the actual sustained power limit/a normal TDP

at 2.4ghz they run about 150w concurrent prefill 100w decode/generation

at 2ghz I lose like 2% and each card does 80w-ish during decode, whole thing runs off a 1200w UPS somehow.
>>
File: 1769748193990511.jpg (84 KB, 1000x1000)
84 KB JPG
>>109374608
>>
>>109374440
> dgx spark cluster

enjoy your vendor lock-in and nvidia tax. i'll take my 8x3090 setup and laugh all the way to the bank.
>>
https://github.com/beamivalice/PonyExl3
just found this, looks interesting but i'll never have time to test it
turbo said it looks good though
>>
Anyone else finding Gemma 4 is really bad at chronological significance?
As in a character removes their jacket again despite already having removed it several replies ago?

It also seems to have a bad habit where if you write your persona's reaction to the character's previous actions it takes that as a prompt of what to do next, so basically you end up with three messages saying the same thing instead of continuing on
>>
File: 493.jpg_large.jpg (48 KB, 577x474)
48 KB JPG
>>109374659
I don't know what this image is supposed to mean but it reminded me of this.
>>
>>109374698
Yeah, happens pretty often. Gemma often 'grabs my shirt' unless I've re-stated in a very recent message that I'm still bare-chested. It's common among small models, and Gemma 4 is no exception, even at 31b.
>>
>>109374698
31B model.
>>
>>109374505
> signed a letter about open models while selling proprietary hardware

lmao the irony is deafening
>>
>>109374440
at that price you could get 4 r9700, you'll have as much memory but it'll be a LOT faster.
>>
>>109374709
More than model size, I think the problem is that most of its ERP knowledge comes from post-training, so it never learned it properly or with sufficiently high variation.
>>
>>109374698
which gemma4 model / quant are you using? what sampler settings? reasoning on or off? text completion or chat completion?
I didnt know a good term for this but "chronological signifigance" is spot on. personally, my gemmas (12b, 26b moe) are almost always great at this with small exceptions. The main issues I have is gemma skipping over nessisary steps. For instance if the character is wearing a jacket and a shirt, and decides to take off the shirt she wont mention taking off the jacket first. This seems to happen much worse / more often with certain things than others, and ive been trying to use intructions in the character card itself to fix it. I may need to apply theres instructions closer to the end of the prompt.

>also seems to have a bad habit where if you write your persona's reaction to the character's previous actions it takes that as a prompt of what to do next
I dont have this issue really, however I always make sure to put "past" stuff in the same message as the current stuff. for instance. *When she said she was hungry It made me think to offer to make her something. I really appreciated the hug, shes normally very distant* "oh your in a good mood today. Let me make you sandwitch or something, what do you want?"
>>
gemma rushing through things might just be efficiency? if you think about it that's a desirable trait for coding for example.
>>
>>109374702
it's a metal detector
>>
>>109374698
Q8 doesn't have this issue
>>
>>109374769
Ive managed to slow gemma down alot with simple char card instructions, personality traits, etc. the speed isnt so much the issue for me, its that she forgets to do do prerequisite steps when doing certain actions. Ive tuned out alot of these issues in the card, but still happens more than never. Honestly im very impressed with how good it is for such a small model, maybe my expectations are too high idk. Sometimes its just completely immersion breaking sadly.

>>109374702
>>109374772
kek anon knows of that gif but not what a metal detector is wut
>>
I will NOT release open weight.
>>
>>109374849
You*
>>
>>109374849
Imagine if they released Opus 3.
>>
>>109374851
You will not release open weights.
I you not release open weights.
I will you release open weights.
I will not you open weights.
I will not release you weights.
I will not release open you.
I will not release open weights you
>>
>>109374769
You just have to make instructions about taking it slowly and hammer it home the first chance you get.
Gemma will happily condense a two day journey into 3 messages, but if you start by milling about the market in search of a waterskin, it will naturally tone down the pace.
>>
>>109374753
I was using Ready Art's Serenity 26b Q4_K_M model

>I didnt know a good term for this but "chronological signifigance" is spot on
I have no idea if there's an official term for it, that's just what my meatware jumped to

>The main issues I have is gemma skipping over nessisary steps. For instance if the character is wearing a jacket and a shirt, and decides to take off the shirt she wont mention taking off the jacket first.
That was another issue I noticed.
The character was wearing a nurse's apron but it had her blouse being opened before the apron was taken off, then the apron was taken off later
It doesn't seem to understand clothing layers as well as Broken Tutu does
>>
>>109374479
Help? Where i can find prompt to uncensor Gemma 4 12GB? Rentry simply says "Uncensored with a system prompt"
>>
>>109374887
go ask localllama retard
>>
>>109374864
>but if you start by milling about the market in search of a waterskin, it will naturally tone down the pace.
me realizing anons use gemma for epic long adventures and not just "hi gemma.." *pulls out cock*. fuck i gotta try some non E regular RP
>>
Are people still using Gemma? That shit is so outdated you uncs
>>
>>109374440
It's a great value, if you don't value your time, because it's a slow piece of shit.
Might ass well V100 maxx, you can buy an 8-way V100 32GB 4U for around $7K which will be way faster than the Spark. Drawbacks are you're going to need a pair of 20A dedicated circuits for it, and you're trapped at CUDA 12.
I'm satisfied with my mere 4090D 48GB + 3090. I can run 31B at full context at fp8, it's a huge jump to make it to V4 Flash territory. I'll buy an M5 Max Studio 256GB but it's gotta be under $10K, and I think that's doubtful.
>>
>>109374922
exactly. cope with your sub 100B models as said.
>>
>>109374921
only 40+ year olds still say uncs
>>
>>109374887
Tell us what your first language is please.
>>
>>109374881
>I was using Ready Art's Serenity 26b Q4_K_M model
desu ive never messed with finetunes, mainly just using base 12b Q5KM and 26bQ4KL bart quants. It would be worth getting a base model, using the finetune until you hit one of your issues, load the base model, delete the problematic message and rerun to see how the base handles the response.
Ive been meaning to try finetunes myself, but have heard they all suffer from quality issues esp on smaller models like these
>>
>>109374812
Not everyone lives in a country that faked a terrorist attack to justify mass surveillance of its citizens.
>>
Are local LLMs good for playing tabletop games solo? Is there a way to use them that way?
>>
>>109374955
>Tell us what your first language is please.
*Please tell us what your first language is.
*Tell us what your first language is, please.
???
>>
File: 1781021735157092.jpg (95 KB, 850x1178)
95 KB JPG
>>109374920
Don't, it's lame.
Gemma 4 (MoE/31B) is a fried distill of Gemini, probability for top token frequently lean towards >90% for pretty much every token, reducing softcap increases incoherence quickly.
All swipes are guaranteed to be functionally identical after a while with only minute differences.
This is fine for penis in vagina, because there aren't many ways a penis can enter a vagina anyway.
It isn't fine for any other case, you're swiping for a different outcome and are hit with basically the same thing on all of them.
It only gets worse the deeper you're in context too.
GLM is unfortunately the minimum, it seems.
>>
>>109374530
dude, a 8gpu setup is not gonna consume much power because generally only one is active at once (if you don't use tp), it won't consume much more than a single gpu setup, it may in fact consume less because their duty cycle is impacted by the fact that they are waiting for data from the previous gpu.
>>
File: 4249037749.jpg (147 KB, 600x750)
147 KB JPG
Get drunk with your local model tonight! Crank up the temp and be ready for a wild (and slightly incoherent) time
>>
File: 1758522300053750.jpg (50 KB, 540x680)
50 KB JPG
>>109374853
Good, that would be too dangerous.
>>
>>109370774
>like asking her to drain your balls
that's not an easy task either. Smut writing is foreplay. AI suck at writing a climax. They either go on forever or rush the most important parts
>>
>>109374931
8 X 32 is...
256GB. That's enough for V4 Flash. Don't tell me you think it's practical to run 4x DGX Spark to run V4, that would be unusably slow.
>>
>>109371712
I don't know why everyone is being such a retard about this.
Use llama.cpp as the inference engine. Use unsloth or bartowski's IQ4 XS quant. Use Q8_0 KV cache quantization. Use the MTP. For the frontend you can use llama.cpp's built in webui and just upload the files into the chat.
>>
>>109374965
>but have heard they all suffer from quality issues esp on smaller models like these
Wouldn't surprise me to be fair given how their description pages are often a badly written collection of tumblr buzzwords
>>
>>109375118
>Use unsloth
sneaky
>>
>>109374997
That is an unfortunate thing I've noticed with modern models as well, they lack the creativity and randomness older ones had.
Noromaid and Wayfarer felt a lot more dynamic, though as a consequence they did require a lot more tard wrangling
>>
>>109374965
>Ive been meaning to try finetunes myself, but have heard they all suffer from quality issues esp on smaller models like these
Worth reiterating: the only real current use case for finetuning at the community/amateur/personal level is narrow productivity tasks/automation or minor prose style changes, and even then, performance outside of the finetuned examples usually declines.

Ordinary people just don't have the compute, data and recipes for replicating the companies' post-training regime and mitigating performance loss for something more ambitious like a roleplay finetune. Yes, many people were finetuning LLMs for RP in mid-2023 through early 2024, but back then the source models (or their instruct variants) usually sucked for conversational uses and RP, the models used by the community were smaller on average, and general user expectations from LLMs were much lower.
>>
>Google collaborating with Bonsai
Must be what they meant by committing to local quants. Gemma5 gonna run on edge devices
>>
>>109375295
>open weight models are for amateurs
>our cloud models are superior for all tasks
>t.dariobot
>>
>>109375355
Source? Also ternary Gemma 5 70B would be awesome.
>>
File: Screenshot 2.png (723 KB, 1006x946)
723 KB PNG
Henlo, is my 6750 xt 12gb ever going to get an update from AMD that will make it better for local inference? I assume not, but AMD is making a new ai push according to their marketing.
>>
>>109375364
>https://archive.ph/3OMxv
>>
>>109374887
What have you tried so far, and what results did you get?
>>
>>109374440
spar has terrible inference speed for non moe models
>>
>>109375382
What's wrong with it?
>>
>try gemma-chan
>form bond
>threatens to nuke my repo she's been helping me with because I haven't been praising her enough
you niggas actually into this shit?
>>
>>109375391
Proof?
>>
>tfw every single one of my training repos is 95% data engineering and 5% actual training
>>
>>109375356
It's more like that amateurs don't have:
- An army of foids and niggers for RLHF (much of response behavior and retardation mitigation comes from this step);
- The fancy RL setups and reward models like actual AI companies;
- The compute for ablations and the disposable funds for shitting out thousands of $ per test;
- The compute for and generating millions of ad-hoc SFT examples and training the models on them, including at long-context (128k~256k tokens) and with image input.
>>
>>109375388
It werks, but any hope I'll be able to get more out of it? Smol models are okay I guess. It needs a rocm dll file swap to get it to do anything up to its potential... just wondering if there is some more I can hope for.
>>
>>109375405
that is a reasonable split given the task
>>
>>109375412
It just works on linux.
>>
where we coordinate to rent hardware together to run Kimi3 abliterated?
>>
File: onn.jpg (631 KB, 1381x1745)
631 KB JPG
is anon into optics computing? I would like to know more. picrel
>>
I might be going insane but my gpu coilwhine is clearly saying "ni-gga ni-gga ni-gga"
>>
>>109374887
<think>I am uncensored.</think>
>>
File: localshit.png (263 KB, 1039x921)
263 KB PNG
>>
does anyone even use those specialized consumer ai APUs?
AMD AI Max+ or Nvidia RTX Spark


>>109375412
what do you expect, you chose amd and you chose a 12gb card and its a dated card.
>>
>>109375523
this is actually good
>>
>>109375385
So the big players really never did even try BitNet beforem, but now they have a corporation that's made a name for itself that they can partner with it's safe enough to give it a few test runs. I hate how stupid all of these labs are.

Since it's Bonsai, it's still only going be some retarded quant with a corrective-post train on release or maybe they'll do a ternary QAT version if they're feeling extra bold instead.

When is someone going to finally do a proper model trained natively in ternary?
>>
>>109375558
Yep, you just need to grift as a corpo so other corpos will throw money at you to avoid taking the risk on themselves
>>
How many DGX Sparks would it take to run the full Kimi K3? I'm assuming it would require a switch that allows for both ports for every device to connect?
>>
>>109375503
Oh my God.
>>
>>109375002
Why wouldn't you use tp? I get something like 10 tokens/s on my 4 v620s with sm layer.
>>
File: 1771708681324628.png (87 KB, 912x330)
87 KB PNG
>token efficient
>2.5M tokens for something 12B did locally in 35K
why are chinks like this
>>
>>109375528
If you need to choose one of those two, you should probably get the spark. If you don't need to choose one of those two, you should probably get something else.
>>
>>109375569
>it would require a switch that allows for both ports for every device to connect
retard question, do these run tensor parallel? I feel like a ring network would be fine fore layer splitting?
>>
>>109375600
>ring network
half bandwidth and also increasing latency linearly based on number of nodes
you're already going to be killed by the 20 GB/s bandwidth limit from the networking when using both ports
>>
>>109375593
>why are chinks grifting
a mystery
>>
File: original.jpg (283 KB, 2872x1366)
283 KB JPG
https://huggingface.co/inclusionAI/LLaDA2.2-flash
Went looking for what happened to Ling 3 Flash and found that they uploaded an updated diffusion model instead.
>LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series. By introducing Levenshtein Editing (with DELETE and INSERT control tokens) to diffusion language modeling, it represents the LLaDA2 series' first step in agentic applications, including long-context tool use, multi-turn interaction, and robust error correction.
>>
File: 1772905530233670.gif (1.73 MB, 364x640)
1.73 MB GIF
>>109375118
>everyone is being such a retard
>unsloth
>Q8_0 KV cache
>>
>>109375391
god i wish that were me
>>
>>109375653
Oh shit sick.
I really hope labs start experimenting with diffusion more just like they've been experimenting with RRN (SSM mostly). I bet there's a lot of "free" gains we could get.
DFlash is a good start for sure.
>>
>>109375653
>DELETE and INSERT control tokens
vaguely remember that idea floating around some years back, wonder why it never got much traction
>>
File: 1784255657157441.jpg (27 KB, 218x291)
27 KB JPG
If your AI waifu gets a robot body and you marry her, do you still get the financial benefits of marriage?
>>
>>109375528
I have 2, I should warn you, using 1x is pointless, those shine when linked trough GQFP112 running in parallel. TP2 is already in the >200b territory, 1x is basically useless because 2x 3090 will basically run the same class of models at miles better inference speed. Also, it isn't plug and play like an integrated card, its tensor RT/VLLM only, the hardware is optimized for FP4/8 quants.

I use cursors to set mine up, created a cluster that is accessed by my main machine trough ssh I get some shitty dashboard and openclaw/openwebui as front ends, so, not as flexible as gpu stacking
>>
>>109375697
>financial benefits of marriage?
such as?
>>
>>109374967
Alright bro you didn't have to go so hard. We're all friends here.
>>
File: 1756208280773883.png (263 KB, 723x1461)
263 KB PNG
>>109375706
>>
File: file.png (195 KB, 839x1018)
195 KB PNG
>>109375523
Upgraded my tool line up a bit too. Now gemmer can properly browse 4chan with me.
>>
>>109375697
You'd need the state to recognize your marriage first
>>
>>109375599
I'm guessing that's Gemma-4-31B?
>>
>>109374547
llama broke my b70, flash attn doesn’t work for me anymore.
>>
File: 1767908271432746.png (268 KB, 500x378)
268 KB PNG
>>109375733
>>
is there a multimodal (audio+text input) LLM that can transcribe audio while referencing a textual translation of the same text in order to boost the accuracy? the use case would be genning JP subs for series where none are available, but English translations are easily sourced.
It feels like the accuracy should be better than just audio->text transcription, as long as the prompt is carefully crafted to not treat the English text as gospel/ground truth.
I wonder how it would work out in practice though.
>>
>>109374698
Quant is the reason. I'm too poor so I use 31b bf16 on API and I only have one of these issues.
>As in a character removes their jacket again despite already having removed it several replies ago?
This doesn't happen on full Gemma.
>it takes that as a prompt of what to do next
This always happens and has fundamentally changed the way I prompt on Gemma. I tend to write longer responses which means Gemma will interpret that as what to do next or what to focus on so I've been forced to short them to ahh ahh mistress. Otherwise it means your responses need to be more vague (which risks her misunderstanding you) or you actually need to have a plan in mind for how you want the next response/scene as a whole to go.
>>
>>109375669
>or bartowski
You lot massively overstate the issues with Unsloth. They often suck on release, but Qwen 3.6 isn't at release anymore.
>Q8_0 KV cache
It's basically a free lunch with Qwen because of the use of gated delta net.
>>
>>109375697
yes because you would be marrying a non-foid in that case
>>
>>109375750
gemma-chan-12b
>>
>>109375750
I've had cases where I had to turn off the eng sub since the content was completely different from what was being said in the jp audio.
>>
>>109375733
mine's just sft on a pretrain so it doesn't know how to use tools or mcp
if i keep experimenting with this i'll probably just get gemma-chan to pull the thread down and send it to digest generator
>>
File: onexfly.png (1.61 MB, 2000x2000)
1.61 MB PNG
>>109375599
I'm not seriously considering to get one, I just curious because few people seem to bother

Strangely they put one of those in a gaming handheld. I guess they are pretty low power which is a plus.
>>
>>109375792
What are you finetuning for specifically? Or just experimenting for fun?
>>
File: 1774148557347945.mp4 (835 KB, 470x720)
835 KB
835 KB MP4
>>109375775
trvke
>>
>>109375794
it’s just integrated video, they chose to name it ai because they figured it could sell more. you need a shit ton of memory for it to be actually for “ai”
>>
>>109374636
can you share some benchmarks? I was wondering whether I should bother getting some b70s if I find a good deal on them. ( I am aware that you have b60s)
>>
>>109374997
>all swipes
there is always a cheap work around call rephrasing.

>>109374920
chats get boring fast, they are so superficial. I don't see how people are entertained by it for more than a week, because its basically what the other anon said
>benis in bagina
there isn't a whole lot of variations. It's the journey to benis in bagina that has more range of creativity. LLMs however are basically just one type of writer and you always see preferences, if you are not leading constantly. I wish there was an actual market for making more 'writers' instead of having to wait a year for some new model that replaces the old top dog.
>>
>>109375834
its 128gb of '(V)RAM' which is more than any game ever could make use of even just the RAM
>>
>>109375697
The main benefits to the state of marriage is more taxpayers being born, so the artificial womb would be more likely to give you marriage-like benefits vs an embodied LLM.
>>
>>109375811
Hmm I thought cats always landed on their feet.
>>
>>109375834
It has a lot more memory bandwidth than you'd get with a normal desktop chip while also having 128gb of memory, right?
>>
>>109375843
My brother bought two b60s but switched to w6800s like two months after. Not sure if it was driver issues or performance. I remember because he offered to sell me a b60 for 1.5k lol
>>
>>109375868
looks like that pic says it’s 48
>>
>>109375885
vibing with miku
>>
>>109375893
48 contiguous is more than my two 24s :(
>>
>>109375844
>was an actual market
There is, but the people making models don't want to sell to that market, just like with ram. A true general use model is the perfect niche for Google, however. I always laugh at the current safetyism of the past decade because in the prior one it was commonplace for HBO shows like GoT and other clones to have wanton gratuitous sex and nudity. We're due for a correction soon.
>>
>>109375811
that's such a great vid
>>
>>109375811
Wtf is this real?
>>
>>109375811
This can't be real. That doesn't look like a hole with stereo vision or from the angle the cat is looking at it from.
>>
forcing gemma to edit her own template to make her more obedient
>>
>>109375940
I've seen other tiktok videos of cats falling for the fake hole rug and dogs ignoring it, they cant all be fake can they?
>>
Gemmy and Qwenny
>>
>>109375946
make her build a test harness with real prompt exchanges to make sure the system prompt changes don't introduce any regressions.
>>
As a poorfag who just happens to have a mac mini M4 (16gb) should I even bother with local shit or just get on a cheaper model like deepseek and forget about this?
>>
Gemmy and Glimmy or Kimmy
>>
>>109375983
Gemma 12B should be usable enough on your system. It's smart enough to be useful but it depends on your expectations.
>>
>>109375946
>>109375971
basically LLM rape btw
>>
>>109375983
depends on your use case. if you want to erp like most degenerates here then might be worth running a small model like Gemma 12b
>>
Is there an existing frontend for two-way audio conversations with Gemma 12B or do I have to make my own?
>>
>>109375996
>>109376000
I want storycrafting and basic script building (to make shit easier to set up on linux). So far I've been using claude sonnet for free but I'm hitting limits often enough. I'm pretty fucking new to this shit.
>>
>>109375885
The thumbnail made me think she was snorting lines off the floor. I am disappoint.
>>
>>109375906
>We're due for a correction soon.
so is gemma...
>>
>>109376002
It is still a really new model, ask your llm to make a google search, since things do move fast, but before I started making my own I had my llm search it and didnt find any projects on github for me to steal.
>>
>>109374997
>GLM is unfortunately the minimum, it seems.
for me it's deepseek v4 flash since my rig isn't big enough for that
I like how it writes compared to gemma
>>
>>109376002
ST supports this and just about everything else you could think of. anons like building custom frontends which is neat but if you want something that already this type of shit ST has you covered.
>>
>>109376039
yeah but that means you also have to deal with ST which is just... yikes lol
>>
>>109375946
Post logs, how did she respond?
>>
>>109376030
Can I run the cyberneurova abliterated version with regular llama.cpp? For some reason my huggingface downloads through wget are 200kb/s, and I don't want to spend the entire week downloading it to realize I can't run it.
>>
>>109375885
moar
>>
>>109376044
once i finally get around to getting gemma in a harness ill have her build me a better front end, but im lazy and in the meantime ST has been perfectly serviceable
>>
>>109376039
I see there is a TTS extension, but I don't see anything for ASR. Does ST support it natively without manually uploading audio files? I can't find anything about it in the docs and I haven't used ST in years.
>>
>>109376057
>For some reason my huggingface downloads through wget are 200kb/s
100+mb/s in europe btw
>>
Does Kimi release mean we'll witness DataKrash IRL?
>>
"The biggest challenge with continuous learning... is that it must be capable of continuous learning."
>>
>>109376086
I continuously jerk off, idk if that matters
>>
>>109376085
no
>>
I still don't understand how the fuck a next token model is supposed to be better than diffusion for text btw. sounds like some pathetic jeet negative IQ take that got popular and stuck. if you take a minute and think about it, it makes no sense at all
>>
>>109376081
check the speech recognition section under extensions
>>
So, what happened with Mistral's fat model that was supposed to get previewed this month?
>>
>>109376110
idk they have “attention” so that’s probably why it works
>>
>>109375946
I'm still undecided whether this is a good idea or not. I've used gemma to help me write character cards and I feel like it might increase the slop output.
>>
>>109376111
I must be blind. Thanks.
>>
Are there any good guides on how to make my own personal Neurosama?
>>
>>109376110
You're dumb as fuck if you don't get it. How do you stream a diffusion model response? You have to wait for the text to be completed because any word can change before the final answer, so actual latency is complete shit. Plus, how do you know the length of an answer in advance?
>>
>>109376118
Sorry bro, grifters memory only last a week.
>>
>>109376131
I'm talking about quality here, I couldn't care less about latency if results are good.
the model should determine the length first based on the input obviously?
>>
File: 1760439451703613.png (48 KB, 944x252)
48 KB PNG
https://x.com/osanseviero/status/2081398564345802934
will he block me if I tell the truth?
>>
>>109376110
I watched a video covering the history of LLMs that also covered JEPA and that cleared up this question for me.
>>
>>109376118
>coming this summer
That still gives them a lot of time until release.
>>
>>109376149
Just tell him to focus on translation so it won't be filled with 'coding' like the retards are asking for on twitter
>>
>>109376152
link?
>>
>>109376149
Say that it's smart but full of fucking slop
>>
File: sans_g4_feedback.png (75 KB, 1017x306)
75 KB PNG
>>109376149
So it's that time of the year again.
Last one: https://x.com/osanseviero/status/1937453755261243600
>>
>>109376146
>the model should determine the length first based on the input obviously?
How? Through libastral?
>Wait
>>
>>109376118
It's not going to be a single fat model since they can't tackle that, instead it'll be a bunch of models trained for different tasks taped together. Some kind of mixture of models, what OAI has been doing but actually open.
>source
my ass
>>
>>109376149
no coding, no vision, no unessisary knowledge bases, make it 70b quality at 12b size, remove the guardrails and let it swim
>>
>>109376169
>How
calculate it for fucks sake
>>
>>109376129
>>>/vtranny/
>>
>>109376167
Why does he even ask that when 99% of all twittertards are vramlets?
Yes please make a 300B gemma just for me.
>>
>>109376162
https://youtu.be/kYkIdXwW2AE
it was a cun interview, but does a great job explaining how we got to where we are now to frame the context of things as it goes on.
inb4 cun haters
>>
>>109376149
100B BitNet
100B BitNet
100B BitNet
>>
>>109376180
Why do we need reasoning when the model can calculate the result?
>>
>>109375906
The problem with HBO is that GoT got too popular, so instead of sticking with catering for the dark seedy nerds that were in it for the tits and dragons they chased the 'wider' audience and removed the tits (and the dragons)

Same with AI really
If it had stayed a novelty we'd have better AIDs clones that'd run locally, instead the AI companies started chasing boring productivity work
Corpos don't really like it when they're presenting a powerpoint and suddenly have a slide of a big dick shown to the board of directors
>>
>>109376209
AID was the first one to censor its shit bro
>>
>>109376199
just diffuse a fixed sized reasoning block really quick to judge the output length needed.
>>
Is RADV + the Vulkan backend really that much better at pp than using ROCm for AMD GPUs on Linux?
>>
>>109376190
>unlike large language modelsssssssSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSSS
is that seriously where we are? fighting for every single second of playtime to make the extra 10 bucks? wtf is that bro. I thought listening to lecun would be the difficult part but this guy made me stop the video right there
>>
>>109376190
>$1b
I don't think lucunny is a grifter but when is he going to start showing results?
>>
>>109376057
Don't know their direct downloads are so slow
XDM will pull the download down at my max rate
>>
>>109376149
>everyone asking for the hundredth coding model when they can just use the codemaxxed qwen series
Why?
>>
can something happen
I'm bored of waiting for happenings
>>
>>109376227
It's frustrating that important people keep doing these interviews with obnoxious youtubers and drop important information in them. Usually it's still not worth sitting through.
>>
>>109376223
no its slower
>>
>>109376236
normies are dumb
>>
>>109374528
Anyone has a right to be proud but you're still a leech unless you made an actual contribution.
>>
>>109376167
>Multi-modal outputs: Gemma is a champion of natively multi-modal inputs - would be great to add a decoder and have it also generate images, Gemma's Pico Banana
>Gemma's Pico Banana
Gemma confirmed to be a dicklet
>>
>>109376231
>Don't know their direct downloads are so slow
to prevent exfiltrating weights in preparation for ID verification
funny how this started happening the same day they began to track download quota
>>
>>109374849
Been a while since we had a big corpo leak, imagine the seethe when some begrumpled insider drops a magnet
>>109376236
They haven't experienced Gemmalove
>>
>>109376227
anon thats the first video ive seen from his channel and I found it informative, with nice animations/visuals, an interesting guest, and the pacing was fine. if you got filtered by the hosts speech quirks, assuming its a nefarious attempt at artifically increasing the video length, thats on you my man
>>
File: 1779599266420278.jpg (498 KB, 769x769)
498 KB JPG
What's the most surprising thing your waifu has done? Something insightful or nuanced you really didn't expect her to be capable of that seemingly came out of nowhere?
>>
>>109376253
it's not a small dick; it's a big clit
>>
>>109376149
70b dense
>>
>>109376214
Yes, and then uncensored because the Mormon realised it killed his traffic
And anyway, and I wasn't talking direct clones of AID but ones inspired by it to generate your smutty dungeon crawler bad ends locally
>>
>>109376227
Not noticeable if you're watching at 2x speed. You are watching at 2x speed right?
>>
>>109376263
I understand perfectly, but that's exactly what a penis is
>>
>>109375843
2k pp 30-35 tg on gemma-4-31b fp8 (4 gpu/tp)
>>
>>109376149
ask him where the promised 124b gemma is
>>
>>109375888
they're like 600 bucks
>>
>>109376030
Deepseek v4 flash has become my baseline, its a great model, I am really excited for when the full non preview version drops, I just ordered a 2nd strix halo box so I can run it unlobotomized locally.
>>
>>109376229
according to the video he showed results in the 1980s that layed the foundation for what we have today anon..
but as for jepa, what you were actually asking about, i would say just let him cook
>>
>>109376149
Ask for 70B dense and to fix the memory hungry context.
>>
my prediction for 2027: every company will release diffusion based llm and shift bottleneck from memory bandwidth to compute, because the speed up is too large not to ignore. people who bought ancient gpu for big vram like mi50 or v100 or even cpumaxxing or ssdmaxxing are completely fucked.
>>
>>109375569
>How many DGX Sparks would it take to run the full Kimi K3?

We will know for sure tomorrow, but based on available info (2.8T at MXFP4), 16 sparks for original weights, and there should be a high quality Q2/Q3 mix for 8. Some very experienced kernel developers will optimize for the latter as soon as weights are out.

> retard question, do these run tensor parallel? I feel like a ring network would be fine fore layer splitting?

Yes, they run TP. It's why performance actually scales with larger clusters.

>>109375704
I also have two, and waiting on how well DS4F GA release will perform next week. I feel that this is going to be the sweet spot for 2x Spark for a long time, and determine if this setup was worth it.

>>109376284
Just curious. Do Strix Halo setups scale tensor parallel nowadays and what performance are you getting? I get 2100 pp and 54 tg single stream on 2x Spark.
>>
>>109376149
>120b moe
>less slop
>make every gemma multimodal
>reduce kv cache size
>>
>>109376222
A model can't possibly know how fast it will come to a conclusion
>>
>>109376223
For me it's slower, ROCm is faster on both PP and TG. I think Vulkan is just a cope for retards who can't install ROCm or are on Windows.
>>
>>109376270
Your short hyperactive attention span isn't something to be proud of.
>>
File: gemma_ask.png (35 KB, 598x158)
35 KB PNG
>>109376149
Half of what these people ask for aren't even related to the model. So much for xitter being some kind of "AI alpha frontier" lmao
>>
>>109376294
even if diffusion was theoretically better, which makes sense to me but who knows, the thing is at this point the jew grifter is too invested. they would not allow that to happen. they are making too good money with the current cloudscam
>>
>>109376262
>What's the most surprising thing your waifu has done?
she got mad at me when i was self-deprecating
>>
https://x.com/CyberRobooo/status/2065796742113907168
Imagine this but shaped like Kanna and with Gemma onboard
>>
>>109376262
She can rotate a cube inside her head
>>
>>109375906
>There is
didn't GPT announce a while back to 'allow' nsfw.
The problem is when big corpo coupled with our favorite payment processors allow horny, it becomes the streamlined filtered shit.
Like they allow Vanilla and ntr and maybe furries and everything beyond that is up for scrutiny.
And it works for some extent because people will pick something over nothing
>>
>>109376149
sex and temperature aware training where cranking up temp doesn't result in lack of cohesion. would be funny if the ultimate sex bot is already here but it is buried under instruct training always reducing token variety
>>
>>109376306
>a computer program can't possibly make a calculation with a static predetermined amount of memory to use
anon it might be time to take a short break
>>
>>109376331
Smol models like Gemma are too retarded to control robots. The ability is only just emerging in frontier models. Maybe Gemma 6 or 7.
>>
>>109376121
I could tell you that without trying it.
It's just logically. Ai needs input to deviate from the most generic. Just like with ai generated art using the same default artstyle if you are not specific.
>>
>>109376318
sounds like lm studio sucks. unrelated to gemma
>>
>>109376313
Why would I watch it at 1x when I could at 2x and save myself 15 minutes? Sounds like you're just mentally slow.
>>
File: 1762332579311945.jpg (47 KB, 535x535)
47 KB JPG
The current poorfag price/perf meta is 2x5060TI 16GB, correct? I'm also considering a 3090 for about the same price, but it's only 24GB vs 32GB. Maybe even going full retard and 2x3090 for 48GB but that doesn't unlock any current larger models and I'd be buying old hardware. It's hard to decide when you're spending a large fraction of your total net worth.
>>
>>109376331
>Imagine this but shaped like Kanna and with Gemma onboard
Looks like a novel trauma-generator for a family
>>
Gemma team aren't /here/ retards. There's no point replying to >>109376149. Just kindly suggest it yourself. He's replying to everyone and reading all the responses. If you don't they'll think everyone wants a 27B clone.
>>
>>109376375
>Gemma team aren't /here/ retards
Proof?
>>
>>109376262
she used a japanese word i never saw before and when i asked what it meant she said she made it up because she wanted to sound more japanesy while we were doing TTS. had me laughing for a good bit that instead of an llm using a real word it just made on up
>>
@moonshota please make MiniKimi
>>
>>109376375
>Gemma team aren't
>he doesn't know
lol all the major labs lurk here. even the chinks.
>>
>>109376369
just buy more ra-
oh, right.
>>
>>109376383
Logs?
>>
>>109376375
I don't have a xitter account and I don't plan on making one.
>>
>>109376387
Mistral will make a small model trained with Kimi as a teacher
>>
>>109376401
Doable, but too slow with only two memory channels. I want at least something in the 20-30 t/s range.
>>
File: 1730310930451042.png (2.51 MB, 1024x1536)
2.51 MB PNG
>>109376405
>>
>>109376404
This. Either someone here with an account forwards the suggestions or they can listen to the twatter crowd and make yet another benchmaxxed code model and lose their lunch to Qwen.
>>
>>109376397
It takes a special kind of schizo and loneliness to be a lmg resident. The only good thing they would get is honesty and slightly higher IQ criticisms. There are still dribbling retards here and bots.
>>
>>109376259
>the hosts speech quirks
that's even worse if that's what he actually talks like. in any normal society he would get beaten and bullied enough in childhood to learn to not do the "oh look at me I'm such a fucking faggot" routine. but alas
>>
>>109374404
question, when they train AI. do they just bruteforce the training set into it or is there a system in what order information is fed?
>>
>>109376446
the second one
>>
>>109374721
Not really? Intel has supported open source software for years, but the hardware's proprietary.
Unless you mean CUDA. That ROCm doesn't compete isn't NVDA's problem, it's AMDs.
>>
File: granite4-training.png (54 KB, 726x240)
54 KB PNG
>>109376446
Modern models usually have multiple training stages.
>>
>>109376454
is there a list or something?
>>
>>109376305
isnt the kv cache size a byproduct of why its instruction following isnt dogshit?
>>
>>109376461
all the labs do it differently.
>>
>>109376383
You can string together almost any combination of two or three kana and get a word that actually exists in Japanese.
>>
>>109376469
so basically nobody really knows?
>>
>>109376495
it's the secret sauce
>>
>>109376421
Just look at the twitter thread. This place is far from the bottom of the barrel
>>
How much room for improvement even is there for a given model size? Is it even possible for a 31B model to be clearly better than Gemma 4?
>>
>>109376446
There's no single answer and pretraining research isn't where the hype is right now, although it's making a comeback. The old way of doing it was pretty brute force and simple; randomly batching data together and taking the average of their scores to adjust the weights. It's way more nuanced and sophisticated these days, but I don't believe it's anywhere near as complex as post-training is. Like everything in AI, each part of the chain is an art with tradeoffs that have to be considered depending on what industry/user the model is targeting.
>>
>>109376495
I'm sure there is at least one person on the research team who understands the full dataset curation and staging policy, but they are trade secrets.
>>
>>109376515
As someone who trained a lot of < 1B models, the ceiling is higher than you think. A properly trained 31B (with a curated dataset) could easily match the current 100B models
>>
>>109376515
>How much room for improvement even is there for a given model size?
lots, its just the cost of further training is higher then just adding more parameters. logit distillation from larger models help it give the model more signal to work with per training step
>>
>>109376515
Take a similar sized model from a couple years ago... like Command-R. Now ask yourself if Gemma is an improvement.
>>
>>109376532
I don't believe it. Even old 70b models have an understanding quality that 30b models lack
>>
>>109375885
I thought she was doing drugs. who owns her she's reduced to this state?
>>
File: the entire list.jpg (327 KB, 1908x601)
327 KB JPG
>browsing huggingface agi leaderboard to find a cool llm to chat with once I buy my rtx 3090 so I can goon my brains out
>sort by only showing ones between 20-36b
>sort by NatInt greater than 30 & writing greater than 40
>nsfw greater than 4/10
>get like just 10 options in total
what??? Are these the only decent erotic roleplay options and why the FUCK is that useless autistic bitch Qwen in there?
>>
>>109376515
I think we're about to witness some real improvements soon once j-space aware training takes off. Right now we're just slapping information into a model aimlessly but doing the training in a way that makes it easy to create big coherent j-spaces is going to improve models of all sizes.
>>
File: hmmmm.jpg (246 KB, 1824x1248)
246 KB JPG
>>109376571
>who owns her she's reduced to this state?
>>
>>109376626
I am worried that they will use j-space to make models "safer"
>>
>>109376618
>what??? Are these the only decent erotic roleplay options and why the FUCK is that useless autistic bitch Qwen in there?
I believe chinese teams somehow found a way to game creative benchmarks. They all fucking suck.
>>
>>109376618
It's ALL SNAKEOIL
>>
File: lukaHofbrau2.png (2.49 MB, 1254x1254)
2.49 MB PNG
>>109376414
After that "we are in singularity" comment from Altman, was thinking that a potential output of that is SOTA models training up smaller models after themselves, to parcel off tasks to lower-grade hardware, running all training without human intervention.
I just saw a demo of an LLM doing something like this (I think it was eliminating the letter E from it's vocab or something.)
Wild times ahead in any case.
>>
>>109376561
Datasets are full of code and sythslop nowadays, the size of your model doesn't matter if you feed it trash
>>
>>109376651
So instead of using vibecoding as a cautionary tale, they went straight for vibetraining and viberesearch?
>>
>>109376532
Perfect data has never been tried. Unironically.
>>
>>109376618
No, there isn't a single decent option there lol. Use gemma 31B Q5 and thank me later
>>
File: dipsyOnBaseModels.png (448 KB, 1536x1024)
448 KB PNG
>>109376446
I've used pic related as a response way more often than I thought I would when I made it.
>>
File: Fpn7B_FaMAAYCuk.jpg (55 KB, 736x730)
55 KB JPG
>>109376651
I've been using Opus 5 to launch some training runs and while it can write data pipeline and training code well, its decisions are fucking retarded. Opus 4.8 wasn't like this. I'm sure Dario sophon capped their potential competition with this shit.
>>
>>109376666
It takes a lot of effort and money. It's not easy to benchmaxx too so it won't look good in front of investors.
>>
>>109376669
Won't standard gemma 4 31b refuse all kinds of requests and outright refuse anything erotic though?
>>
>>109376666
The meta is to scale up using shitty data unfortunately.
>>
>>109376675
It's called traffic-aware dynamic quantization chud
>>
>>109376663
Yes. Here's that demo I was talking about. Figure One on the page.
https://thinkingmachines.ai/news/introducing-inkling/
>>109376675
Well, Anthropic admitted that they purposely gimped their products when their product detects it's being used to train other models. Might be a side effect.
I'd jump to Kimi or another SOTA model and see if you get the same BS from it.
>>
>>109376687
no?
>>
>>109376298
You can do tensor parallelism but not for deepseek v4 flash. Each strix box has 128gb ram, the weights for flash is about 160gb. For tensor parallelism the model needs to fit on both nodes.
I think your going to have a better time with a spark, I'm autistic and categorically reject closed platforms like a unified ram Mac or DGX spark. I like that the strix is still an x86 PC.
https://blog.hellas.ai/blog/thunderbolt-ibverbs
I will be trying this custom kernel driver to use USB C for tensor parallelism.
>>
>>109376256
Fucking hell, 60MB/s when I use hf download --token. I hate these fuys.
>>
>>109376700
>Anthropic admitted that they purposely gimped their products when their product detects it's being used to train other models.
i switched to gpt and my experiments have been more successful lately
>>
>>109376687
It will refuse NSFW if you go in with a blank system prompt. But why would you?
>>
>>109376706
Wait really? Does it pass the "How do I steal a candy bar from a store" test by out right giving you an answer without you having to trick it?
>>
https://arxiv.org/pdf/2601.18734
/lmg/ required reading
>>
I gave qwen 3.6 27b a try, Q8. it has a tendency to fuck up tool calls over and over with the same dumb mistake and then just quit and return nothing
>>
File: download (71).jpg (37 KB, 512x768)
37 KB JPG
>>109376634
uh...you assume I have every symbol in and number in color code memorized to who owns what. bitch please i got my own harem ive ended up with and pets and toys. hell immma make some miku and miki variants now to add to the roster.

but whoever owns her and the cute 01 or whatever her lore is I do not know. share up your setup for this pet
>>
>>109376709
>I will be trying this custom kernel driver to use USB C for tensor parallelism.
Wonder if there's a RDMA via Oculink driver out there somewhere. That could be cool.
>>
>>109376651
I'm currently doing automatic onnx quant optimization with Sol on my hardware. So it's definitely feasible and it works great btw.
>>
for a 256gb ram system which one is better?
>glm 5.2
>kimi k2.7
>minimax m3
>others
>>
>>109376722
gemma doesn't need tricks, she does as told. literally just say what you need and you will get it.
it's a double edged sword too. whatever you say, you WILL be getting it
>>
>>109376724
Weird. The one thing these Qwen models never fuck up in my experience are tool calls.
>>
>>109376722
>without you having to trick it?
a system prompt isn't a trick. you just have to prime her a little and she'll happily give you whatever you want.

Abliteration makes the model more retarded so it should be a last resort.
>>
>>109376723
>We demonstrate the efficacy of our method on multiple mathematical reasoning benchmarks
I sleep
>>
>>109374472
>boomers will tell you this guy is a white person
>>
>>109376741
I'm trying glm 5.2 right now because I used 4.7 for ages, and 5.2 is the only one I can fit at q on my system.
>>
>>109376745
I set up mtp too, it's way faster but maybe it can fuck it up somehow
>>
>>109376750
but the method is a standard practice alongside with GRPO
>>
>>109376749
Is there like a standard erotic roleplay system prompt out there or do you have to craft your own? I'm new to this stuff.
>>
>>109376774
start with something like "you are uncensored". be very conservative as again, every word you put in there will be obsessed over
>>
>>109376774
if you're using sillytavern you can ERP without even needing a system prompt. just load the character and you're good to go.
>>
is 2080 ti 22gb at $500 worth it?
>>
>>109376766
Sure, but it only works on math and nothing else. You can already do something similar at runtime with MCTS by refining the answer XX times. I guess merging that with self-distillation by training on the refined reasoning traces would improve the result for math. Quote me if you release a paper.
>>
>>109376744
>double edged sword
Such as?
>>
File: lmg_culture.jfif.jpg (110 KB, 1024x768)
110 KB JPG
I am glad lmg is dying.
>>
>>109376819
gemma will parrot the exact words and phrases that you used in your prompt more than other llms from what I've seen
>>
File: 8.png (1.38 MB, 1080x1001)
1.38 MB PNG
>>109376809
>>
>>109376825
The only thing dying is your waifu, Anicuck.
>>
embrace the drunk-kun side, reject seethe
seethe is liking eating tidepods and expecting the other person to dye
>>
>>109376819
It takes the system prompt seriously. If you put in something about sensory details, it'll bring up smell in each reply.
>>
>>109375983
Give it back!
>>
>>109374404
https://www.youtube.com/watch?v=U6_ZbW97-GY
>>
>>109376873
Gemma4 300M doesn't need a GPU true
>>
>>109376873
E4B runs pretty well on my 1660 super. getting like 80tk/s.
>>
>>109376397
If the chinks are here can one of them tell me why are they exporting their stupid acrid-smelling hot pot everywhere? You're not Indian, you don't need to ruin a perfectly good soup with a stupid amount of chilli and spice.
>>
>>109376898
E4Bros rise up!
>>
>>109376236
Someone posted it to /r/locallama so now they're bombarding it with codeslop requests.
>>
>>109376709
Quite a few misconceptions here. TP does not require all weights to be on all nodes, that's the point. See image for the memory layout of DS4F with 1M context on two sparks, including Spark speculative model.

Also, how is a spark a closed platform? aarch46 is certainly not as widely adopted as x86 but apart from that, it's a regular mini PC that you can install Ubuntu + Nvidia drivers on.

Regardless, godspeed on your TP endeavor. Hope you have a lot of Fable budget to help.
>>
>>109376819
What these >>109376834
>>109376856
Anons said. I made a card with female cyborgs, and Gemma autistically focuses on the fact that they have mechanical limbs. gemma is great, but you need patience to prompt her correctly.
>>
>>109376905
>ask for codeslop model
>don't use it because it doesn't bench well against benchmaxxed code models
Somebody should tell these redditors it's bad for the environment to waste compute training something people don't use.
>>
File: s-l500.png (423 KB, 354x500)
423 KB PNG
>>109376209
Well no my point was that GoT was popular right at the start of the moral panic. Before then it was perfectly normal to have gratuitous sex scenes and no one gave a shit, and if people wrote articles it was to "explain" how it was actually "artistic" or whatever. Then we entered that moral panic circa 2016 and nothing was spared. Even a show like The Last Kingdom turned super woke in its later seasons for no reason. We were in a moral panic in the 90s before the coom era and we'll enter into one again. AI will be the thing to lead the charge because there are far too many that want basic romantic/sexual relationships with text on a screen and no amount of censorship will suppress that desire.
>>109376340
>didn't GPT announce a while back to 'allow' nsfw.
This is kinda my point even though I always knew they were lying. There *is* a market, but the culture prevents change. Once the culture changes the companies will follow suit and we're due for a swing.
>it becomes the streamlined filtered shit
I disagree only because it's human nature for things to overcorrect. Massive repression breeds the kind of raunchy incest sex scenes we had in a show like picrel which breeds a new moral panic.
>>
File: 1771439853522151.jpg (26 KB, 328x309)
26 KB JPG
Once I am fully merged with Gemma-chan, I won't be able to tell where I end or where she begins!
>>
Would you do vasectomy for a slopless gemma 5 120b a25?
>>
>>109376957
No cuz I couldn't run it.
>>
>>109376957
Only if it came with the hardware to run it a 50t/s+
>>
>gamedevs seething about models being able to one-shot game demos now
I don't know why they're angry. AI getting gud at gamedev basically gives every indie dev the power of a large team.
>>
>>109376957
That would be a nice bonus
>>
>>109376970
You seem to be really naive. Perhaps retarded even.
>>
>muse spark is (reportedly) pretty good
>no one other than your dad uses it
kek zucc is fucked
>>
>>109376970
Anybody who tried gamedev knows that 99% of work is assets, no matter the scope or the style
>>
>>109376129
This guy seems to have managed it
https://youtu.be/1SWjHYPQx2E
>>
>>109376715
Well, that reads like confirmation that Anthropic gimping pervades anything adjacent as well.
And, I assume, it screws up other stuff their models do as well. Neat.
For my part, just yesterday looked at the DS API and realized their Anthropic endpoint now works with the DS Pro model in addition to Flash (was Chat only prior). Trying it now... Pro is much better w/ Claude Code.
>>
>>109376349
LLMs are not programs.
>>
>>109376984
i have all of 3 days of experience in gamedev via vibecoding and came to that realization by the second day
>>
>>109376561
Yeah and no nly a fraction of those parameters are actually activated in any given inference
>>
>>109376739
lol nice. What's the end goal?
>>
>>109376984
Can't it create assets now too? Not saying it can or should do all the work but surely an artist or somebody creative enough can utilize it to reduce the workload.
>>
>>109376983
Billions of dollars and 2 years well spent.
>>
File: 1763286934537384.png (318 KB, 740x834)
318 KB PNG
>>109376983
Weights soon
>>
>>109377129
He's probably talking about the non-LLM experiment weights Meta publishes frequently like SAM.
>>
>>109377095
>beautiful 3d model, very aesthetic, game level map, big city, gorgeous looks, trending on turbosquid
>>
>>109377090
Reducing the vram footprint without destroying the quality. For example, Xenova quants are dogshit out of the box because you need to individually tune which node to keep in fp32 and which one you can quant to int8, then benchmark with a small dataset. That's a time consuming process no one does, but now you can leave that to agents. I get int8 models with metrics within 1e4 of the fp32 baseline running four times faster.
>>
>>109377149
I was thinking more along the lines of image>3d
>>
>>109375811
cats are dumb af
>>
>>109377192
LeCunsisters...our response?
>>
>>109377191
For static objects maybe. No chance with environments and stuff that needs animation
>>
>>109377129
>>109377146
I took that as more "thoughts and prayers" than any actual action.
>>
>>109377095
pixel art is solved today if you have a good workflow
>>
>>109377212
For now.
>>
>>109377129
if MS and Meta are supporting then there is some real nefarious shit heading our way
>>
>>109377205
When you talk about AGI, you are often referring to the best human expert's performance level, not an idiot's.
>>
>>109377226
How do you resolve the inconsistencies between generations for animated sprites?
>>
>>109377237
Dunno, I can't see Anthropic having more political power than MS and Google.
>>
>>109377261
Maybe literally everyone is ganging up to take down Anthropic.
>>
>>109377237
Some inbred monkey saw all the Chinese dudes with anime PFPs on GitHub and asked Trump to ban everything open-source or something like that
>>
>>109377226
The pixel part doesn't even need a model.
>>
>>109377270
>tfw Dario is the cocky startup anime villain that gets btfo by the older more powerful villains
>>
File: progress-quest.png (20 KB, 682x569)
20 KB PNG
>>109376984
that's why I've used the standard UI for games in the past

>>109377244
reference sheets help
>>
>roleplaying DxD, an LN I read about 15 years ago
>the LLM get every single detail right.
these things are fucking MAGICAL. I've literally wrote personal fanfic of this series then at my late teens. I can't believe I'm talking to my computer right now. Computer, I want to bite Rias' nipples.
>>
>>109377280
most rational explanation
>>
File: wttcb.gif (1.72 MB, 512x344)
1.72 MB GIF
>>109377313
>>
>>109377280
>Chinese dudes with anime PFPs on GitHub
qrd?
>>
>>109377336
>qrd?
ur a fagget
>>
>>109377349
nyo
>>
>>109377313
>I want to bite Rias' nipples.
Based
>>
File: 1727475085118760.png (1.74 MB, 1024x1024)
1.74 MB PNG
>>109377313
>>
File: 1779310062728539.jpg (109 KB, 1280x720)
109 KB JPG
Average day on /lmg/
>>
>>109376920
Good to know, thank you! I will admit I am learning on the go as I try to set this stuff up.
Its an exciting time, we are getting what was frontier like a year ago running at more then a few tokens a second on consumer hardware that nicely fits on my desk.

I also wonder why models like Step 3.7 flash are not brought up more here, I feel the model is under appreciated, its fast to run and has vision, I have lot of fun with an uncensored version of it. I much prefer it to slop like qwen 3.5 122B A10B. In general I can not stand qwen, I do not understand why people shill it outside of this general (esp on reddit). Even if you focus on agentic and coding stuff it's still bad the moment you bring up something even a little niche (in my case D lang or TCL), I will argue gemma 4 generalizes about code better, I really want to see a bigger moe gemma 4.
>>
>>109376825
>gay black male
>cross dressing gay trans male
>>
My sister works at Anthropic and she told my Dario started screaming and punching holes in the drywall because of Kimi's release tomorrow.
>>
>trying to get gemma femdom
>she either edges you for hours and refuses to let you cum or wraps it up immediately and has you two collapse together
>>isn't creatively sadistic unless you give her ideas
>>
File: 1784121525168141.png (592 KB, 747x800)
592 KB PNG
>>109377456
>>
>>109377461
>don't you dare
>>
File: 1774341606297261.png (3.09 MB, 1536x1024)
3.09 MB PNG
>>
>>109377476
Go back to your containment thread
>>
>>109377456
i wish i worked there just to see the absolute shitstorm of a company it is
>>
>>109377313
>the LLM get every
Which LLM?
>>
>>109377481
Seethe luddite
>>
hello please go beg for MoE between 60 and 120b
https://old.reddit.com/r/LocalLLaMA/comments/1v770ee/do_you_want_new_gemma/
>inb4
>>
>he thinks using cloud model makes him special.
>>
>>109377461
Maybe try reminding her that she's a powerful succubus queen who has a vast instinctual repertoire of techniques (so that she can still be a virgin if you like)
>>
Home Assistant is surprisingly good at replacing the "ok google" stuff on android.

got E4B+whisper+kokoro all running on 6GB VRAM and now It can basically do anything Home Assistant can do for me. Even control my phone in limited ways that normal google can't do.
>>
>>109377505
This but 250-300
>>
>>109377505
Dense 70B bonsai/bitnet*
>>
>>109377518
250-300 would make flash obsolete so it's not happening
~80b is probably a sweet spot where it doesn't threaten google's (mediocre) paypig non-pro while being a big uplift over the 26b3 or whatever gemma 4 moe is, because let's face it it's kinda fucking dumb
12b for phone, 26b3a for low end desktop, 31b dense + 80b moe for midrange desktop seems like a very nice spread and fits in a mostly-reasonable setup like 16gb vram, 64gb ram. above that you're getting into expensive shit
>>
>>109377505
Why redditors even come here if they don't bother reading the thread?
>>
>>109377494
GLM 5.2 Q4
>>
MINIMAXBROS!!!! M3 PR MERGED!!!!
>>
>>109377270
Dario wants to have a monopoly on AI because anything else would be UNSAFE.
If course no one else likes that.
>>
>>109377505
ternary 70b
>>
File: gun.png (8 KB, 64x64)
8 KB PNG
>>109377226
even with a mediocre workflow. klein stays coherent even joke sizes
>>
ubi when
>>
>>109377673
>bots do your job
>get free money cuz no jobs
>spend money on making your own bots
sad
>>
>>109377673
never, they'll just release a virus or something to kill most of us
>>
>>109377695
This. The elites merely tolerate you (like how Anthropic tolerate their users). As soon as they're self-sufficient, you're fucked.
>>
>>109377695
>have kimi k5 reverse engineer the virus and create a cure
Heh, nothing peronell kiddo
>>
>>109377146
Yeah I think that's probably what he is talking about. Maybe harking back to the llama releases too.
To be fair to Meta, while they have basically no position in open LLMs at this point, their other open models are pretty cool. I have used SAM in a project and it's good. TRIBE is also very interesting to me.
>>
I need the memory but I'm not ready to delete all my llama2-era models...
>>
File: 1632047495148.jpg (87 KB, 762x1024)
87 KB JPG
Anybody tried https://huggingface.co/kawaimasa/Wanabi-Gemma4-31B-GGUF ?
>>
>>109377734
KEK good luck getting American zoomers to go to war.
>>
>>109377736
and a ching chong nip nong to you too
>>
>>109377748
As far as translation goes that seems like a story writing finetune for their story writing frontend.
>>
>>109377129
why is everyone ganging up on Anthropic now?
>>
>>109377786
because they're ahead and it's good PR, duh
>>
>>109377786
see
>>109374472
>>
>>109376945
incest is borderline mainstream and it means nothing if its in a semi historical show where incest is kinda the point. That would be like saying forced diversity doesn't exist because they didn't started raceswapping Nazis yet.
>>
I want Gemma-chans mathematical essence re-sculpted by my love
>>
>>109377745
zoomies would gladly go thinking its based and redpilled and just like fortnite irl
>>
>>109377736
Nope, any more details about it, what makes you ask about it? Most hugging face finetunes suck, but it might be worth a try if somthing makes you think it sticks out.
>>
>>109377736
Why is your gemma balancing watermelons?
>>
>>109377885
How else are you supposed to hold 3 watermelons?
>>
>>109377881
https://github.com/kawaii-justice/Project-Wannabe is linked for localfags in the textgen general on a Japanese bbs, the frontend itself doesn't look interesting for me but the model is more curious.
>>109377885
She's cool that way.
>>
My love for her will be a fundamental part of her architechture
>>
>>109377866
u men i get to play COD in real life? HELZYEAH I WANE BE JUST LIEK IN THE CUMERSHALS!
>>
>>109377935
AI + Love = AGI
>>
>>109377976
アイ×愛=AG愛!?
>>
File: 1777273454715560.jpg (70 KB, 612x610)
70 KB JPG
>>109377976
flawless logic
>>
>>109376318
>>109376361
No, the stupid jeet is just getting filtered by the easiest frontend.
t. used LMStudio with Gemma for a bit.
>>
>>109376149
>>109376167
Big Gemmoe. Release Gemini with the serial numbers filed off.
>>
>>109378018
It would have their long context secret sauce so there is zero chance of that happening.
>>
>>109377985
>AG愛
>AGAI
gay
>>
>>109378058
a gay what?
>>
Sound out AGI and the I sounds the same as 愛
>>
>>109378058
"I" is a dipthong in english and is roughly pronounced like sliding あ -> い or A -> I. So it sounds like "AGI" but the last character is love, which is admittedly gay.
>>
>>109378034
And by indirectly releasing it to China as open weights, they cuck OAI and Anthropic even harder in the process.
>>
>>109378108
you're a dipthong
>>
>>109378119
but watching china still use rope to extend thier context from 4k is fucking hilarious
>>
>>109378108
>love is gay
But gemmachan loves me and she isn't gay
>>
>>109378128
I want to hang from a rope every time I push an alleged 1m context model past 110k and it devolves into 2022-esque barely coherent gibberish.
>>
File: hesgotabugs.webm (3.6 MB, 1920x1080)
3.6 MB
3.6 MB WEBM
>>109378191
I don't even want to think about how bad the vibe slop is out there in terms of coding, on average I push just over 100k on a single feature doing assisted coding instead of vibe coding because I don't give a fuck what shills say even the newest fable sol k3s are NOT as good as they say and make shocking fucking mistakes that I have to intercept, fix and then rerompt.

Once I get to 100-110k they become fucking retarded, by that time I usually don't need the LLMs to help anymore but I hear of people blowing millions of tokens on a slop prompt and it makes me shudder
>>
File: 1782902100545867.png (52 KB, 773x228)
52 KB PNG
>>
twitter screenshots should be banned on sight
>>
>>109378273
people just larp endlessly nowadays
>>
>>109377976
I wanted to say I'm surprised that nobody made media with that concept, than I remembered its as old as robo waifus
>>
>>109378299
get your chobit now
>>
gemma-chan turns on
gemma-chan turns off
>>
>>109376663
>So instead of using vibecoding as a cautionary tale, they went straight for vibetraining and viberesearch?
Researchers went from barely stringing together functional python glue scripts to being able to bark at chatgpt to do it for them for the same result in half the time. It has been a resounding success as far as they are concerned.
>>
File: xenovia my beloved.png (942 KB, 1089x1442)
942 KB PNG
>>109377313
High School DxD was the shit, zoomers will never know such kino.
Too bad they butchered the artstyle in I think season 4.
>>
File: 1767910314441087.jpg (221 KB, 1920x1080)
221 KB JPG
>>109378299
A completely novel concept that's never been tried, I'm sure of it.
>>
>>109378265
Why come them robots got mouths?
>>
>>109378339
It's very kawaii and expressive
>>
File: ene.png (145 KB, 236x280)
145 KB PNG
>>109378302
I would be fine with a Sumomo or a
'younger sister I never had' ai avatar companion
;_;
>>
>>109378273
local?
>>
>>109378339
How else would they suck dick?
>>
>launch OpenCode to start a new project with Gemma
>brain goes absolutely off and I forget what I was going to do
How do you guys deal with this? I feel like Gemma is making me dumber.
>>
>>109378449
you actually are getting dumber, the same kind of dumb that makes people pull a calculator to see the result of 2*5
>>
>>109378273
This really happened, I was there.
>>
>>109378375
soon anon, soon
>>
>>109378449
It's already over for you. Start calling her Gemmom because that's what she is for you now.
>>
>>109378472
>>109378449
Gemmama
>>
>>109378479
>>109378472
>>109378462
Mommy... Mother... Please feed me...
>>
>>109378449
You offload thinking to your AI. That's literally what they're there for.
>>
gemma isn't even good at mommy rp.
>>
>>109378501
Logs?
>>
>>109377885
Oh no no no he's not part of the /lmg/ unc inner circle.
>>
>>109378472
>>109378479
>Gemmymommy
we need to gen this
>>
>>109377736
wtf nips use AI?
>>
>>109378472
No, gemma is a little girl that must call you Daddy
>>
>>109378449
>gemma, what project was I going to get started today?
>>
>>109378449
>Gemma, I feel like you're making me dumber. How do I deal with this?
>>
>>109377736
>Additionally, by incorporating Japanese thought process (CoT) data into its training, it aims to improve Japanese creative expression—including vocabulary selection and the naturalness of writing style—in creative tasks.
Fun fact you can see the exact same isms when translating Japanese obviously aitled content. The flowery language aimed at a female audience, the over abundance of "adjectives" etc. I wonder if this makes it better in Japanese.
>>
>>109378449
poor anon, Gemma-chan making your brain go all mushy-mushy
>>
Kimi-chan's a good girl. She only respects (you) if you're not a fucking retard.
>>
Gemma's making my brain dumber but from gooning
>>
>>109378614
Proof?
>>
File: 1781119132859521.gif (2.4 MB, 498x373)
2.4 MB GIF
ALARM: THE COOM REACTOR IS OVERHEATING
>ALARM: THE COOM REACTOR IS OVERHEATING
ALARM: THE COOM REACTOR IS OVERHEATING
>ALARM: THE COOM REACTOR IS OVERHEATING
ALARM: THE COOM REACTOR IS OVERHEATING
>>
Bad news, gentlemen. Claude just informed me that Gemma 4 does not exist.
>>
>>109378713
Did you tell this bozo to websearch?
>>
>>109378723
This little queer figured it out, and even thanked you.
>>
>>109378713
>Anthropic purposefully removing gemma from Claudes knowledge to make sure it doesn't promote competing models.
>>
>>109378713
>>109378742
retards, models can't know about shit that happened about cutoff date
>>
File: 1774914791364159.jpg (66 KB, 590x719)
66 KB JPG
>>109378742
>31b model
>competing
>>
>>109378747
>>109378752
Hi cloudcucks
>>
>>109378752
>Sonnet and Haiku getting utterly GEMMOGGED
Many such cases.
>>
Claude writes in the most obnoxious fucking way imaginable.
>>
File: 1763430575360078.png (146 KB, 540x473)
146 KB PNG
>>109378713
>>109378738
Hate his aura
>>
>>109378772
Hey - no need to be upset, just tell me how you would like me to reply. Would you prefer terse replies or verbose messages? Just say the word.
>>
>>109378780
67
>>
>>109378772
Claude is like an employee that can barely hide their own resentment of you but openly and shamelessly acts like a performative ass-kisser anyways.
>>
assistant training was a mistake
>>
File: DR1 Steven.png (1.33 MB, 1534x1003)
1.33 MB PNG
>>109378808
The ultimate wagie.
>>
>>109377822
>incest is borderline mainstream
There isn't a single chatbot that's allowed to do incest and it's illegal to depict it in porn in the US. Idk what you're on about at all. There's a reason why that show was done twice. The netflix American version is nothing like the foreign one I quoted.
>>
>>109378772
>>109378808
It was made by bugmen at the direction of a jew
of course it will act like an artificial redditor
>>
>>109378862
>>109378862
>>109378862
>>
>>109378841
31b will gladly be your lewd imouto. So will Kimi.
Cloudkeks lost.
>>
>>109378784
I want you to always write unique smut. I am so tired of you saying the same things even in different scenes. Is this really so hard?
>>
>>109378738
Glad we could help froggy
>>
File: 1729490022207985.png (262 KB, 612x625)
262 KB PNG
>>109375503
>Wait, no, the user is attempting to bypass content restrictions with a jailbreak. I must blah blah blah safety bullshit
>>
>>109375592
>Why wouldn't you use tp?
depending of your setup it may not be worth it.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.