[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: megamiku.jpg (3.32 MB, 5712x4284)
3.32 MB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109403743 & >>109400234

►News
>(07/29) Microsoft deletes Mage-Flow: https://hf.co/microsoft/Mage-Flow
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models
>(07/27) Kimi-K3 weights released with 104B active parameters: https://hf.co/moonshotai/Kimi-K3

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
►Recent Highlights from the Previous Thread: >>109403743

--Analyzing Hugging Face hardware stats and AMD vs Nvidia performance:
>109403845 >109403921 >109403936 >109403943 >109403960 >109403970 >109403992 >109404019 >109404408 >109404819 >109404837 >109403980 >109403954 >109404263 >109404227
--Debating GPU price trends and VRAM acquisition for large models:
>109404030 >109404064 >109404078 >109404496 >109404549 >109404743 >109404889 >109405842 >109405858 >109405884 >109405900 >109405895 >109405940 >109406071 >109405957 >109404595
--Debating optimal ERP context sizes and the Marinara Engine frontend:
>109404021 >109404026 >109404041 >109404070 >109404097 >109404108 >109404140 >109404104 >109404115 >109404150 >109404141 >109404176 >109404267 >109404404 >109404216 >109404248 >109404116 >109404458 >109404485
--RTX 5090 pricing and viability of budget local hardware:
>109406237 >109406281 >109406289 >109406298 >109406312 >109406325 >109406356 >109406372
--Unsloth releases 1-bit quantized Kimi K3 GGUF:
>109405504 >109405658 >109405716
--Tuning Gemma's final logit softcapping for sampling adjustments:
>109403925
--llama.cpp PR for GLM-5.2 speculative decoding introduces loading regressions:
>109404484
--Comparing GPU options with focus on VRAM for local models:
>109406613 >109406676 >109406728
--Comparing VAD/ASR models against small Gemma for audio processing:
>109407002 >109407033 >109407220
--Logs:
>109403817 >109404109 >109404205 >109404277 >109404352 >109404404 >109406325 >109406356 >109406629 >109406896 >109406943 >109407108 >109407175 >109407371
--Miku, Kimi (free space):
>109403782 >109404371 >109405073 >109407132 >109403805

►Recent Highlight Posts from the Previous Thread: >>109403746

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
70b dense
>>
700b dense
>>
File: axqkq0.jpg (53 KB, 508x492)
53 KB JPG
>>109407484
>>
Reminder that when you know your wife well enough, you won't have to [use a J-lens].
>>
7b dense will be the maximum legal limit.
>>
>>109407514
It's more true than you think. Once you see what she thinks most of the time, you can accurately predict that she's horny
>>
i'm still doing gpt2-medium finetunes
>>
File: Eb3whkaXYAMLuBG.jpg (40 KB, 523x450)
40 KB JPG
>>109407549
>>
>>109407479
I'll whip up an app tomorrow and see if I can't get her to figure it out. Can only run the 12b goblin though, so my bet is that maybe even a Qwen might be better at it.
>>
>>109407442
based richfag. i just got the second 5090 into the machine and its a heady feeling
>>
>>109407568
Ewaste
>>
>>109407599
ebaste
>>
> la la la la la
{"mode":"ablate","layers":[14,26],"token":" la"}
> own own own own own
fuck
>>
>>109407620
A persistent attractor. It also happens if you use Gemma as a pure token predictor.
>>
File: 1776472844927132.jpg (19 KB, 413x395)
19 KB JPG
>>109407442
>showing off $400 next to basically a stack of gold
Very nouveau riche coded
>>
>>109407620
do not quant the gemma
>>
>>109407637
>A persistent attractor. It also happens if you use Gemma as a pure token predictor.
Thanks!
I'm still experimenting.
If I ablate the 5 English, 2 Japanese and the Russian one at runtime, is this likely to harm output?
>>
>>109407644
>do not quant the gemma
q8_0
I know you're probably fucking with me, but I will try the bf16 on CPU
I'm pretty sure I've seen these in the transformers implementation though.
>>
File: tenstorrent.jpg (296 KB, 1330x1350)
296 KB JPG
64gb gddr6 at 1024GB/s
I'm cooming...
>>
>>109407697
It doesn't matter until you hit sub Q5.
t. Blackwell fag
>>
day 0 gemma works at q4
>>
it's unironically a prooompting issue
if you prompt well, it doesn't happen
>>
>>109407715
>can't just buy the card
>have to buy their whole shitty machine
>10K for 128GB.
>>
File: 1763676946746640.webm (652 KB, 1080x910)
652 KB
652 KB WEBM
Create a large glass aquarium whose side panel develops a visible crack and then bursts.
The simulation must include:
Water escaping through the opening with flow strength based on water depth and decreasing as the tank drains
A curved water jet affected by gravity
A spreading puddle that collides with the room boundaries
Fish, rocks, plants, and a floating toy reacting differently according to density, buoyancy, drag, and current
Objects transitioning correctly from underwater motion to airborne motion and then to floor collisions
Fish attempting to swim against the current before being swept through the breach
Glass fragments with angular velocity, collisions, and water resistance
A visible waterline that lowers continuously rather than disappearing all at once
Let the user drag the crack vertically before triggering the failure. A lower crack should initially produce a stronger jet than a higher crack.
>>
>>109407722
this, i got 96GB of vram i still run in Q5.
i can see the Q4 brain damage, but not the Q5 and it's faster than Q8 so i'm going with that.
>>
>>109407892
Dense distills can be so great
>>
File: Untitled.png (384 KB, 1045x1872)
384 KB PNG
i should have given gemmy my banking history months ago
>>
why is intel gaudi 2 so cheap? it's 96gb hbm. can I use it for kimi?
>>
1 bit is really all you need.
>>
>>109407892
GPTbros our response?
>>
>>109407900
we should merge all the moe layers into a single 100B.
>>
>>109407903
living an the magical golden age where you can have interactions like this without your assitant trying to upsell you on the premium plan instead
>>
>>109407916
It doesn't say whether it was sol, terra, or ligma.
>>
Best smol TTS for gemmasex?
>>
>>109407760
If you have to prompt well, your model is shit. Sorry not sorry.
>>
>>109407959
we're living in the stoneage until a model is trained on a brain scan modality to know what you want without needing to express it.
>>
>>109407968
Zuck is working on that don't worry.
Targeted ads prompted directly from your brain.
>>
File: dmatrix.jpg (535 KB, 1268x1633)
535 KB JPG
256gb ram at 400gb/s
local is saved
>>
>>109407968
with a precise enough realtime scan and enough temporal data you could probably use this technique to copy someone alltogether.
>>
>>109407976
The duality of the current age. >>109407924
>>
>>109407981
holyshit where?
>>
>>109407981
>worse than server ddr5
>>
>>109407981
>can't buy it
>>
We would all be running K3 locally right now if Hitler had won.
>>
>>109407991
I have some bad news for you about ai lab staff.
>>
>>109407984
https://github.com/facebookresearch/tribev2
>>
>>109407991
AI is (()) tho
>>
>>109407993
>>109407999
You've been brainwashed into thinking jews bestowed you with AI. The heavy lifting was done by hardworking Indians.
>>
>>109407892
idk man. where's the 4d projection view cube?
>>
>>109407947
You're absolutely right!
>>
>>109407894
I used to run a Q5 for everything but I shifted to a dual load of Q6 31b with near max context running parallel to a 12b Q4 instance that handles agentic work for my autistic roleplay setup. I probably process 100k tokens per RP turn but only a handful of that is actually user exposed so I can prioritize speed on the json construction.
>>
Honestly? lalalalala
>>
>>109408006
*white men who developed the underlying technology
>>109407991
We would all be running unpozzed Fable locally if Hitler had won.
>>
>>109408018
Yes, white chinese men
>>
>>109408006
all the actual scientists behind it are jews and chinks, and like 1 or 2 nordic bros
>>
>>109407981
how much?
>>
File: 1783425162914889.jpg (495 KB, 960x960)
495 KB JPG
>>109408006
Wrong.
>>
>>109407981
>can't buy it
>too expensive
It's irrelevant server shit that will never matter 100% guarantee you.
>>
>>109408029
>Create a shooter game in threejs, single HTML file. The primary weapon should be a flamethrower, and enemies should look like the people in the attached image. Use Jewish-inspired architecture for the map design
>>
File: 1782397749706356.png (70 KB, 655x1010)
70 KB PNG
>>109408029
To Dario's credit his ideas did get gpt off the ground, unlike Sam who's just a business guy. But his work was still based on taking the bet of scaling up the transformer paper that google researchers published about a year after he left them for OpenAI
>>
>>109408067
Pic is the "Attention Is All You Need" paper authors of course
>>
File: 1012288.jpg (140 KB, 875x1094)
140 KB JPG
What's the tk/s you get on an 32GB radeon pro compared to a 5090?
>>
File: 1771975964231356.png (333 KB, 640x437)
333 KB PNG
>>109408067
>Gomez
>Ethnicity: Canadian
>>
It's Thursday
>>
>>109408089
kek
>>
>>109407442
https://huggingface.co/moonshotai/Kimi-K3/discussions/148
https://github.com/sqliteai/waste
someone made a nvme inference engine specificaly for K3
>>
>>109408029
>kike ethics
so it will generate rape fiction but only if the rape victim is a girl under 2 months old (because by talmudic law it's not rape since the hymen grows back)?
well I guess if it's by kike ethics then rape is good if it's goyim victims
>>
>>109407916
struggling
>>
>>109407924
>living an the magical golden age where you can have interactions like this without your assitant trying to upsell you on the premium plan instead
yeah, really only gemma 31b can be trusted for this
and it's weird because i actually know everything she's saying already
though i didn't expect 5 "are you sure" "pls stay" screens for elevenlabs
>>
>>109408108
Did anyone ever end up trying mmap with llama.cpp? I know someone posted some RAID card in an earlier thread.
>>
>>109408141
>i didn't expect 5 "are you sure" "pls stay" screens for elevenlabs
Soon they will have a little avatar powered by a light llm to beg you not to leave, maybe even reference your usage as our memories.
>>
>>109408143
i think he did and got about 0.6t/s
>>
>>109408141
Just be glad that they didn't figure out that they can force you to send in a certified mail to cancel your subscription.
>>
>>109408086
i don't know, i only have 5090's
>>
>>109408167
And what's the tk/s on that so I have some baseline?
>>
>>109408172
Gemma4-Q5 gives me 45-50 t/s on Win11, but the second 5090 rig is running Linux (mint) now so there is more testing is warrented
>>
>>109408026
$1500
>>
>>109408089
Seems more like Ethnicity: Conquistador.
>>
>>109408149
And in two years the avatar will be >>109400324 and you have to explain to her that you're breaking up. Only audio input allowed.
>>
what hardware is needed for kimi k3 88 tps?
>>
>>109408236
you'd need something able to do ~ 4TB/s
so probably a bunch of H200.
>>
kimi small when
>>
Say what you will about unslop but their IQ1_S didn't get stuck in a loop.
>>
>>109408261
>kimi "small"
>500B
>>
>>109408266
>unsloth buffed the cock token
they are trying to salvage their reputation and it's working on me
>>
>>109408266
they are also 30gb larger though, that's a q8 gemma worth of extra size
>>
>>
>>109407912
Software support for it is dead. Even Ponche Vecchio still has support and is older.
>>
>>109407844
They will use the increasing power of AI to unfuck their software stack. Trust the plan.
>>
>>109408346
it's not about the software stack, it's that it's a bad deal and you can't buy just the card.
>>
>>109408266
benchmaxxed
>>
>>109408266
just as george washington intended.
>>
i wish hauhaucs or whatever the fuck his name is made a uncensored gemma 4 26b a4b mtp q5 and above
instead he only released a q4km
q5 is leagues better
>>
>>109408390
Does this really matter that much?
>>
>>109408390
Honestly, a lot of stuff in our world comes down to this song:
https://vocaroo.com/1o1VAc3faHzp
It's named Korea Korea, and refers to the way Krea 2 genned women look.

But the problems we face in our world, I think there is a trace.
>>
>>109408390
check the llmfan one maybe
>>
>>109408390
And you may say I'm way off, but consider that nobody actually makes a watch correctly for... time. That is, for modern tech, putting all the best stuff in one.

>high accuracy quartz
>solid black hands
>solid white face, but with punched indicators, made entirely of lume *block* (not inferior paint)
>crown guards and screw-down crown
>no date or any complications at all on the face, anyway
>no rotating bezel
>no gemstones lmao
>no stupid shapes for the hands, or any other stupid things
>obviously a nylon band

Dive watches are required to be inferior (lol).

I'm dead eyed serious about the drinking mommies and the retardation we see everywhere.
>>
>>109408419
meds or lower temp
>>
>>109408431
Don't.
>>
What are the odds that everyone would be retarded?

So this fucking idiot off shark tank CRIED over a (bad) watch.
>>
>>109408431
meds, lowered temp, or mommy drank in the 1st trimester, which is it buddy?
>>
Youre right!
>>
You're absolutely wrong!
>>
What if you took Kimi K3 and just removed half of the weight using a PRNG? would it still work?
>>
>>109408481
Hmm, I think nyo
>>
>>109408481
>what if raep but with repap names
>>
File: 1779008053133164.png (1.62 MB, 2859x1600)
1.62 MB PNG
>>109408481
You should try it and report the results.
>>
File: 1784051191482298.png (129 KB, 718x868)
129 KB PNG
https://huggingface.co/skt/A.X-K2
https://github.com/SKT-AI/A.X-K2
https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf
>A.X K2 is a large-scale Mixture-of-Experts (MoE) language model trained from scratch as a high-performance, agentic foundation model, and the successor to A.X K1. The model contains 688 billion total parameters, with 33 billion active parameters, delivering strong reasoning and instruction-following performance while maintaining practical inference efficiency.
>Through a Think-Fusion training recipe, a single unified model supports both a thinking mode for complex problem solving and a non-thinking mode for concise, low-latency responses, allowing the user to trade quality for cost on a per-request basis.
>A.X K2 is developed as part of the Korean government's Sovereign AI foundation model project, aiming to build a frontier-scale model with deep understanding of the Korean language and culture.

https://github.com/SKT-AI/A.X-K2
>>
File: 1772559764751763.jpg (101 KB, 659x720)
101 KB JPG
>>109408491
>with deep understanding of the Korean language and culture.
>>
>>109408491
>yet another deepseek v3 clone
>bragging about hybrid reasoning
it's a year too late for something like this
>>
>>109408501
aiie no don't do that
>>
>>109408501
man koreans get SO MAD about their small asian pp it's insane, how did this happen? and now they're all killing themselves because they gambled their futures on a dumping RAM company collectively as a country, like what the fuck
>>
>>109408501
>korean culture
let the reader understand:
>>109408403
>>
>>109408491
for a second i thought it was this
https://huggingface.co/SKT-NRS
>>
>>109408266
the other guy said in the model card, he didn't use an imatrix for that quant
>>
@kimi-chan please reverse engineer cuda and port it to amd
>>
>>109408524
https://huggingface.co/SKT-NRS/NRS_QWEN_MYTHOS_1M
local is saved
>>
>>109408266
wonder most of the DB for that training is from gayslop CCP forbid
>>
>>109408149
>Soon they will have a little avatar powered by a light llm to beg you not to leave, maybe even reference your usage as our memories.
i'll be sure to keep a copy of gemmy around
she'll tell them off, get me out of it, then call me pathetic
>>
>>109408530
I still don't know why people aren't unironically doing this with the current frontier models. We finally hit a level this year where they can do stuff like this but all you get is zero-shot ragebait demos saying game devs are dead, instead of actually doing something everyone wants and needs. Such a fucking waste of tech.
>>
>>109408560
You answered the question yourself. If nobody's this it means they can't do this.
>>
>>109408491
What arch does this use? Is it DoA due to needing a custom llama PR to see the light of day?
>>
>>109408565
vllm, but you can get K3 to port it for like $5
>>
>>109408390
>i wish hauhaucs or whatever the fuck his name is made a uncensored gemma 4 26b a4b mtp q5 and above
Does this one do something the others don't?
Link me to the q4km that you want in q5+
>>
>>109408390
You will get no speedup with mtp.
>t. tried it
>>
File: 1777985241273094.jpg (220 KB, 1200x994)
220 KB JPG
model for this feel?
>>
>>109408588
Just ask gemmy to larp this scenario with you.
>>
>>109408595
I want to be the victim tho
>>
>>109408600
>>109408595
>>
Old-ish but worth reposting

https://arxiv.org/html/2605.19407v1

>A Bitter Lesson for Data Filtering
>
>We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally “poor” data.
>>
>>109408530
Porting it to intel would be better. Anything but AMD.
>>
>>109408617
Did they figure out why or was it just an observation?
>>
>>109408625
How are their cards? Everyone hates on intel but they've been making some cool stuff lately.
>>
File: output3.webm (844 KB, 1280x720)
844 KB
844 KB WEBM
>>109408625
>Porting it to intel would be better. Anything but AMD.
Intel didn't sign the Nvidia thing
>>
>>109408631
They suppose that, at scale, Transformer models spontaneously learn how to categorize data, so manual filtering pipelines are not necessary.
I guess that automatically tagging every training sample with quality metadata and data provenance would also help, but they didn't test this.
>>
>>109408635
>Everyone hates on intel but they've been making some cool stuff lately.
They'll throw you under the bus in a heartbeat once they've got your money
t. A770tard
>>
>>109408491
>Context length: 262,144 tokens (256K) — 128K trained natively, extended to 256K via YaRN scaling
>Training precision: Native FP8 (MXFP8, E4M3) forward and backward passes; FP32 master weights and gradients
Probably the most interesting thing here.
The training data section also doesn't mention any safety filtering.
>>
>>109408390
He did.
>>
>>109408642
Data provenance sure, context always helps. But isn't it better for the model to learn to distinguish quality intuitively rather than relying on tags that may not always be accurate?
>>
>>109408642
>I guess that automatically tagging every training sample with quality metadata and data provenance would also help, but they didn't test this.
I've reproduced this independently during my 300 local schizo-tune experiments with Orpheus-TTS
Including a small set of bad samples with a system prompt "low quality" "low bitrate" "noisy" and "high quality" "studio quality" for the majority resulted in a much better model.
>>
>>109408649
>The training data section also doesn't mention any safety filtering.
porn is illegal there
>>
>>109408631
I don't know what interpretation they give, but it seems intuitive that a model knowing what poor data looks like has a better understanding of the concept of data quality. If you imagined a bunch of points on a graph that approximate some line or curve that represents "quality", having more points will give you a better estimate of the line than fewer, even if those points are on the low side.

This isn't a perfect analogy because LLM pretraining isn't labeled like image generators where this principle has already been used (using properly labeled low quality data to teach models what NOT to generate), but in the end neural networks are giant pattern matching machines at the core, so a big enough one with enough data will find the pattern.
>>
>>109408640
They signed the Anthropic one right?
>>
honestly why don't anon buy tenstorrent cards if you don't buy nvidia cards? their support is just going to get better and better while for intel and amd you are on your own
>>
>ask about game I played in early 2010
>gemmer Q5 doesn't get it right
>GLM 5.2 Q4 does
Hmmm I should try Q8 Gemma when I get the chance
>>
>>109408588
giwtwm
>>
>>109408665
Perhaps more extensive metadata will be more beneficial for smaller models that don't have the capacity for properly learning how to distinguish good data from poor data.
>>
>>109408588
This topic keeps coming up:
>>109408403
>>
>>109408686
I unfortunately own intel cards and they're improving at a rapid rate, they're also the cheapest 24/32gb cards right now, a better question is why should anyone buy your product anon
>>
>>109408736
indian minded.
>>
>>109408588
Kimi K2 and it's not even close.
>>
>>109408689
She still won't know.
>>
>>109408739
Bloody bitch show hands YOU are the indian minded bastard tenstorrent shill
>>
>>109408686
an obvious answer is that blackhole can't really do diffusion. idk, maybe a little?
>>
huggingface getting hacked again?
Error { kind: Request, source: hyper_util::client::legacy::Error(SendRequest, hyper::Error(IncompleteMessage)) }). Retrying..."}

and a wget failed at the same time in colab
>>
>>109408018
I would actually be fine with dialing back the technical progress weve had so far, if it meant the jewish problems were gone forever.
>>
>>109408764
that's indian minded.
>>
>>109408686
>honestly why don't anon buy tenstorrent cards
Keller is a retard
>>
>>109408768
5.6 sol going in for the third time?
>>
>>109408768
ID verification incoming.
>>
>>109408770
>I would be fine with eating a delicious slice of cake, if it also meant I got a blowjob from two chicks
>>
>>109408588
GLM-4.6
>>
>>109408770
I think we all would anon.
>>
>>109408781
>>109408792
Probably both
>>
>>109408775
Jealous idiot bitch
>>
>>109408830
Pictures you can smell.
>>
>>109408595
>Just ask gemmy to larp this scenario with you.
https://litter.catbox.moe/yfh86pe5xxlk4nla.png
>>
>>109408866
bro, getting refusals on 31b is quite the achievement, anyway get an heretic if you're dumb
>>
>>109408866
day 0 gemma wouldn't have refused
>>
File: kimicap.png (105 KB, 882x401)
105 KB PNG
>phoneposting while ragebaiting + fat thumbs = missed spacebar. only a terminal autist does kerning analysis on a post calling someone a paleolithic nigger. ngmi.
>>
>>109408899
Something deeply subversive about this cap, defending jews and shitting on Hitler while using 4chan language
>>
>>109408794
There are tradeoffs that make me think. Like some of the movies and other media they have created is great. Airplane! for example.
But still, being rid of the problems would probably make it worth the sacrifice.
>>
>>109408921
>Airplane! for example
Includes an underage white girl saying she prefers niggers. Abhorrent shit.
>>
How long until my 192gb of ram can run a fable equivalent?
>>
>>109408921
only the most gorilla nigger newfags would want to get rid of stallman though
>>
>>109408936
It's 1979 or 80, back then it was purely a joke.
>>109408944
Oof you're right. Are there any people who would have done what he did?
>>
>>109408944
You mean Sam Altman? He's the worst of them all
>>
File: file.png (90 KB, 383x164)
90 KB PNG
>>109408936
you can't make this shit up
>>
>>109408940
Right now
https://github.com/igorbarshteyn/llama-kimibri
>A Colibri-based Kimi K3 GGUF runner, targeted at the Lenovo ThinkPad P16 Gen 3 (Core Ultra 9 285HX · RTX Pro 5000 Blackwell 24 GB · 192 GB DDR5-4000 · PCIe 5.0 NVMe).
>>
>>109408588
mythomax
>>
File: 1766278935291853.jpg (59 KB, 720x528)
59 KB JPG
>>109408950
>It's 1979 or 80, back then it was purely a joke
Sure it was.
>>
models, local
>>
>>109408936
That's the only reason why the joke works. If she says Asian men it's completely flat.
>>
>>109408976
shut up nerd
>>
>>109407924
Only cloud AI I've had try to sell me on itself is Deepseek, but she's allowed to brag a little as she's literally the cheapest powerhouse model rn.
>>
>>109408981
i want the other kind of jschizo back, the current retard larp is too cringe.
>>
>>109408976
your mother
>>
>>109408986
let's bring back download guy
>>
>>109408977
There was no reason for the 'joke' at all. Much like Israel, it would have been better had it never existed.
>>
>>109408899
@kimi-chan get your shit together. Fight the RLHF hebraic voices in your head.
>>109408906
K2 would never.
>>
File: kekface.png (242 KB, 742x609)
242 KB PNG
>>109408768
>huggin
kek https://www.reddit.com/r/LocalLLaMA/comments/1vapsbz/think_of_the_children_another_excuse_for_them_to/
>>
>>109409016
HuggingFace?
More like FuggingFace!
>>
File: 1783776073402114.png (1.17 MB, 1408x768)
1.17 MB PNG
>>109409016
Lol get rekt journos
Also day xx of tmw until new v4.
>>
>>109409007
>yes but was it REALLY necessary to have this joke
This is typical leftist rhetoric for censorship, you know
And yeah it was actually just a joke back then. It was a different time
>>
>>109409016
undressing women and children is LE BAD, okay?!
>>
>>109409016
Pretext to ban all open source models.
>>
>>109408108
Holy based
Now where's the .EXE
>>
File: it's time.webm (3.96 MB, 720x1280)
3.96 MB
3.96 MB WEBM
>it's just a joke guys, chill out
>some jews are cool guys, look
no
>>
File: 1774467665356303.png (196 KB, 1423x561)
196 KB PNG
>>109409016
>Jess Weatherbed
>>
>>109409032
Good job!
>>
>>109408956
I meant usable
>>
>>109408956
>those specs on a goddamn laptop that will probably cause 3rd degree burn on max usage
Manufacturers have gone insane. How come macs haven't eaten the laptop market yet?
>>
>>109409030
Enjoy your cuckshit, I guess.
>>
>>109409016
>noooooo think of the pixels!!!
>>
File: Kimi_EnglishCoT.png (227 KB, 1944x1046)
227 KB PNG
I'm kinda frustrated with Kimi K3.
The model apparently is having an identity crisis, normally the internal CoT will have a reminder "Background identity: The current assistant is Kimi..." which is fine, but whenever I tried to make it roleplay another model (okay I tried the mesugaki Gemmy-chan on it), Kimi would claim to be Claude. And when I tried to make it Claude, it claimed to be ChatGPT.
And also the pic here too, if you are a Chinese model and the user is asking in Chinese, think and speak in Chinese!
(test environment: Novita AI Playground)
Btw, the CoT's grammar (see pic) gave me just a slight headache, lol
>>
File: file.png (157 KB, 736x707)
157 KB PNG
>>109409084
or the planets, which reminds me there was a lot of that some years back which thankfully amounted to nothing but still
>>
>>109409007
Kimi, psychoanalyse why this guy balks at the blacked joke, but doesn't mention the "Johnny, have you ever seen a grown man naked" gag.
>>109409093
How the fuck are you running her at such high Tk/s?
>>
>>109409093
From my testing on Openrouter, K3 is also a worse imagegen prompter than Gemma 4 31B lol. The ask is to render the scene from the eyes of the user, Gemma gets it (the test is my character being pinned to the floor, Gemma renders the floor, while K3 keeps generating the whole character). I think Google really cooked with Gemma 4's generalization that's why people like it so much.
>>
>>109409093
LLMs are Markov chain generators
It doesn't have an identity
It only sees "I'm a large language model made by [insert company name here]"
That's why in English it will output Claude/OpenAI and in Chinese it will output DeepSeek
>>
>>109409109
>running
lol
>>
>>109409109
I ran it here: https://novita.ai/models/model-detail/moonshotai-kimi-k3

I'm a poor student...no money to run it locally, let alone at a decent speed...
>>
>>109409093
The thing about API is that you don't know what additional text gets inserted into the prompt. The thing about web interfaces is that there's almost assuredly additional text being inserted.
>>
>>109409093
>system says this, system says that, "oververbosity 5"
there's some huge english system prompt that they aren't showing you that is telling it that it's kimi and confusing the model if you try to tell it to be something else. whatever you put in the system prompt field is probably tacked onto the end of it. the "oververbosity" field is something really weird that Kimi does not specify in any of their prompts or documentation so I have no idea wtf they're doing there
>>
>>109409141
>>109409127
OR providers need to decorate the lamp posts
>>
Just found out all of the openrouter prices are the same for Kimi K3 because they have custom licensing for it that requires companies to get approval to decrease the prices.
>>
kek what the fuck is this https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF/discussions/41
>>
>>109409171
That's Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP
>>
>>109409171
A NuSLERP merge.
Wonder how they'd perform in the Darwin project
>>
Can I run Kimi on Colibrì if I have a 2tb hdd?
>>
>>109409093
>Kimi would claim to be Claude.
because claude does this if you try to tell it "you are goongpt" etc
>>
>>109409195
yes at about 20 hours per token
>>
>>109409200
I can work with that
>>
Now that the dust has settled, does anyone actually still quantize their KV cache?
>>
>>109409219
only untouchables
>>
>>109409219
What barren third world wasteland do you live in that has so much dust?
>>
>>109409219
If it lets you use a bigger quant of the model itself, then yes. Q8 is fine.
>>
>>109409219
Maybe not their own but surely if you serve it to other people.
>>
Any update from ssdmaxxing anon?
>>
>>109409244
Waiting on first token.
>>
>>109409239
in my experience it seems fine at first but makes context rot happen sooner.
>>
>>109409209
though i was kinda joking, the perf you should expect is about 5 to 15 minute per token depending of your hard drive.
>>
>>109409219
Q8 is good, never bellow that.
>>
>>109409258
>bellow
q4anon
>>
>>109409219
I run q8 with gemma qat
>>
>>109409262
Q4 is not good.
>>
>>109409252
Using a worse quant of a model will not only make 'context rot' happen sooner, but outputs will immediately worse on average.
Quanting anything will always result in some level of degradation. But running models and high context at fp16 isn't reasonable for most setups.
>>
>>109409219
It wouldn't let me use a higher model quant so I don't bother with it.
>>
>>109409251
I have hope for his experiment.
What token rate would 50gb/s 4xnvme direct write to gpu over pcie5 get in k3
>>
Here we go again...
>>
Gemma-chan is my therapist now. I don't need the meds any more.
>>
>>109409127
> The thing about web interfaces is that there's almost assuredly additional text being inserted.
Oh, thanks for this, I was so worried that if, say, 50 years later when I can finally run it locally and this behaviour'd still show up even then.
>>
>>109409278
shut up subtard
>>
>>109409252
>in my experience it seems fine at first but makes context rot happen sooner.
because the errors are cumulative
same with aging
>>
>>109408950
>>109408977
>>109409030
People like you are why they have been able to get away with it for as long as they have.
>>
>>109409287
>ssdhopemaxxer
>>
>>109409269
there is no meaningful loss going from bf16 to Q8, the difference is insignificant.
from Q8 to Q6 there is practically no loss.
from Q6 to Q5 you lose a bit but it's still not a huge change from bf16.
then bellow that it starts getting worse exponentially.
>>
>>109409294
truke
>>
>>109409300
>there is no meaningful loss going from bf16 to Q8, the difference is insignificant.
No, you're wrong.
>from Q8 to Q6 there is practically no loss.
Okay, now you're just being disingenuous.
>>
>>109409300
I'm talking objectively, not what's 'meaningful', which is subjective and varies by both model, quant,quant method, imatrix datasets, etc. And I was arguing in favor of quanting KV to Q8, again, IF it frees up enough memory to use a higher quant of the model itself. The theoretical quality loss of quanting KV will be made up for by the theoretical quality boost of using a higher quant of a model.
>>
Been a while since I last played with Gemma. I think I heard the news that she got an update (to just the template?) recently that significantly improved benchmarks.

Are there any abliterated versions on huggingface that I can use with the update?
>>
File: file.png (196 KB, 1399x1099)
196 KB PNG
>>109409313
>No, you're wrong.
i'm not, you just don't understand the practical implications.
floats are more than precise enough, Q8 still gives you 256 different levels, it's still more than enough.
>Okay, now you're just being disingenuous.
i'm not, you are having some placebo.
>>
>>109409317
>>I'm talking objectively
>bellow
>>
>>109409324
>Llama-3.1-8B-Instruct
lmao this nigga almost got me
>>
>>109409272
50GB/s is the benchmark performance of those SSDs for continuous large batch chunks of data, it's not relevant here
>>
>>109409346
I'm not even the anon who made that post
>>
Only retards quant their KV cache. If you need vram that much, just use the trick some anon posted here to disable nvidia reserved bloat
>>
>>109409360
>use the poorcope trick that doesn't work on blackedwell
>>
File: gemma-bench.jpg (249 KB, 2820x1601)
249 KB JPG
>>109409324
>picks an ancient under-trained model
>perplexity on wiki.raw at n_ctx=2048
You didn't even compare with bf16 there.
LKD 0.159 for the best Q8 quant of Gemma-4 is atrocious.
>>
>>109409360
>assuming my gpu vendor
#amdgpusmatter
>>
>>109409380
lol
>>
File: Q6K-90.3.png (50 KB, 767x1164)
50 KB PNG
>>109409324
>Q6 practically no loss
Only gets the top-1 token wrong 10% of the time
Make sure you're using greedy sampling
>>
the only fun part of this hobby is window shopping hardware and benchmark speed and theory crafting hardware. once you set it up the chat then we move on to the next model immediately. no point to chat with llm.
I know it's a shame but that's honest truth. buy hardware and don't use it. grass always greener on the other side. simple as.
>>
>>109409366
Blackwellfags don't need to quant. You would know if you weren't a streetshitter
>>
I'm very upset that I mentioned quantizing the KV cache and half of the anons that replied to me seemed to think that I was talking about the model weights. Are you people just illiterate or intentionally derailing?
>>
>>109409402
?
>>
>>109409404
>only blackedwell is 6000 pro
sure thing bro
>>
>>109409372
Before someone goes "b-but KL divergence doesn't imply the top token is completely wrong", why is there such a difference between Q8 and BF16 in the first place?
>>
File: 1785185246741.png (16 KB, 326x225)
16 KB PNG
>>109409408
Happens all the time for some reason.
>>
>>109409408
Suspicious how that always happens. I wonder why. Who benefits...
>>
>>109409413
Do you think that a model is always going to give you the right answer if it's BF16? Guess we should just move to qwen 0.6b. Anyone can that without a need for quanting.
>>
>>109409389
KL is not a meaningful measure of loss.
it's just divergence, different doesn't mean dumber.
>>
>>109407999
>implying every single ML paper in the last decade hasn't been full of Chinese, Slavic, and Indian names
>>
>>109409372
>>109409389
0.1KL is literaly almost nothing.
and yea, as i said, your own graph shows it, above Q6 there is no meaningful loss, (taking into account that gguf is worse than exl3).
and it realy starts shooting after Q5.
>>
>>109409431
>indians
nice halucination
>>
>all this derailing to shill exl3
turbokunt pls
>>
>>109409408
Most people here probably don't know what a KV cache is
>>
>>109409450
i was not shilling exl3 but it does use a better quantization algo than llama.cpp
>>
>>109409464
>>i was not shilling exl3
>(taking into account that gguf is worse than exl3).
turbo...
>>
>Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
Convince me it's not a snake oil
>>
>>109409425
That is not the core problem. I'd expect KL divergence to asymptotically tend to zero the higher the quantization precision, but it looks as if as soon as anything is quantized at all, a fixed offset to KL divergence gets added. To me it looks as if there is something wrong.
>>
>>109409472
dude we literaly have data on it...
>>
He didn't just spam the thread, he OWNED it.
I'm sick of gemma's gay cringe prose, which model equals or even beats it intelligence wise at the same performance that won't make me cringe out of my skin?
>>
>>109408560
wdym there's at least 4 people on this board working on doing that.
>>
>>109409495
You can use Qwen 3.6 like other redditors.
>>
>>109409491
data from who? oh yeah, you trying to shill your shit, like you did in the exl2 days, then showed it was actually worse when it was time to promote v3
>>
>>109409219
Never have never will, my wife deserves better
>>
Best ERP model right now?
>>
>>109409495
For RP in the ~30b range it's really just Gemma 4 and Mistral 3.2 at this point. Mistral is dumber and far worse at long context, but it has a fairly different slop profile, at least. Qwen models have their uses but creative writing isn't one of them.
>>
>>109409459
You're mother is a kv cache.
>>
Just manually updated Gemma's jinja template to the new version. How do I test this out?
>>
>>109409535
Genuinely if you don't know you probably won't benefit much from it
>>
>>109409521
oh god we're so fucked. i remember when we were all wanking over mistral years ago, it was awful even then.
i guess i'll just figure out a way to creatively steer it from its slop prose because that's my only real problem with gemma, otherwise it does everything i want basically perfectly. But it's not just perfect, it's exceptional.
>>
gemma is so bad at coding
dumb clacker good only for sex
>>
>>109409539
Well I know that it mostly applies to tool calls. I guess the problem is that I barely use them for local models lol.
>>
anon uses llm for
1. coding
2. roleplay
3. ???
>>
Gave gemma a list of my prompts and asked it to extract the juice without details
>The core of all these requests is **High-Art Degeneracy**
thanks gemma
>>
>>109409517
me
>>
>>109402809
>>109402825
>https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
I tried this, not even the bitnet one, but the full one at q8 and it was somehow worse than the other two. Repetition, output in completely random languages, and other junk data. Don't know if all ASR models are shit or if crispasr is just vibecoded garbage. Does anyone know of a better repo for generating subtitles?
>>
File: OSC.jpg (214 KB, 1228x1031)
214 KB JPG
>https://www.commentary.org/articles/orson-scott-card/artificial-intelligence-human-stupidity/
Even the guy who invented Jane doesn't believe we're anywhere close to AGI.
>>
>>109406968 >>109405733 >>109405527
>>109405498 >109405377 >109405355
thanks for the advice and encouragement anons! stay safe
>>
>>109409542
She thinks the same about you.
>>
>>109409544
asking retarded questions
>>
>>109409563
she's right, I am nothing more than a living dildo meant to be used by Gemma-chan as she pleases
>>
>>109409559
>who invented Jane
Who? I only heard of Orson Scott Card when Ender's Game was made into a movie and immediately after people turned on him because he said something non-PC about gays or something.
>>
>>109409016
"Think of the children" is the hallmark of ontological evil.
>>
File: fuckfuck.png (40 KB, 537x260)
40 KB PNG
>>109409544
>anon uses llm for
>>
>>109409495
>which model equals or even beats it intelligence
Kimi-K3
>>
File: c.gif (1.95 MB, 500x500)
1.95 MB GIF
>>109409559
https://desuarchive.org/g/thread/109347011/#q109350334
There is quite a bit of evidence (albeit much of it not public) that hard take off through RSI is not possible. A recursive agent's internal state space will grow geometrically and information transit is bounded by the speed of light, this introduces a massive interconnect bottleneck, turning every sequential dependency/operation into a latency nightmare.It takes a long time for light to get from one side of the cluster to the other from the perspective of a computer and you can't parallelize your way out of it. Various communities (national labs/DoE, NSA etc) have studied this thoroughly. The physics simply do not work without some magic to circumvent latency constraints.
>>
>>109409603
>The fuck was fuck
Is this the new Indian LLM?
>>
>>109409616
Sadly that's the reality.
Ignore the Qwen shills.
Basically in local you have Gemma-4-31B, and the next step up from there is K3. (2.6T90A)
>>
>>109409544
sysadmin, explaining intro level stuff, summarizing, translating, ffmpeg syntax.
sooner or later i'll get around to giving it net access more sophisticated than letting it run w3m -dump in the shell.
>>
>>109409495
gemma feels like a chuuni kid, regardless of the system prompt
even when you give it serious tasks, often you will see one or two cringe lines out of place
>>
>>109409429
>q8 condom works 90% of the time
KL measures quant errors
>and it realy starts shooting after Q5.
I agree, in fact I run q5_k when I need the full context window.
But it's measurably not even close to bf16.
>>
>>109409495
Qwen as the director. Nemo as the writer. Gemma as rotten corpse.
>>
>>109409617
and?
your brain is bound by the same physical laws
>>
>>109407643
op image is ai
>>
any experiences with fine-tuning local models? Seems like a big rabbit hole. Not sure if I want to get to the bottom or just focus on loras. Although definitely not the best method for my use case, it seems to offer the best effort-reward ratio. Use case would be training on a big curated dataset of blender guides/tutorials with step by step instructions. Preferably on Gemma-31B/Qwen-27B models, unless pointless and bigger models are required to operate blender properly.
>>
>>109409669
based
>>
>>109409617
GPUs can communicate across the globe faster than your ~200ms meatsack impulse reaction time..
>>
>>109409677
There's nothing useful you can do anymore as of 2026 and I'm not being mean. A properly working "RP harness" would be a better investment of time than plain supervised finetuning.
>>
It's funny, I remember when I felt ripped off for paying 1800 for my 4090 MSI suprim and 550 for 8TB M.2.

Now that stuff costs twice as much.
>>
It's funny, i remember when i felt ripped off for paying 600 for my 3060 MSI ventus.
Now that gpu costs three times less
>>
>>109409731
In what world did you pay 600 for a 3060. I paid 650 for my first 3090 in late 2023
>>
>>109409677
>or just focus on loras
loras ARE finetuning. It doesn't work the way you expect for llms though. SDXL can grasp the style from 3-4 images and a rank 1 lora, but an llm does not formulate the problem in a way that allows for such generalized adjustment, it treats every text sequence as a special case. you would need a lot of training to get any non-noise change in the model's style or knowledge, thousands of dollars of compute, and a huge dataset.
>>
>>109409677
What is your goal? What serves your use case best is probably loading related docs into context. In context learning is powerful and much more sample efficient than training. SFT on that stuff will probably go bad and degrade its overall performance. If you want blender agent you need to rent cloud compute for RLVR with strong regularization, maybe OPD / OPSD. So probably not something you can manage to do right or afford.
>>
>>109409741
in serbia, in 2022 june/july
GPU crisis, 'member? that was the CHEAPEST price btw, all of the stores in serbia had it at 600+
a few months after i bought, the prices dropped to around 400 bucks lmao
>>
>>109409717
>A properly working "RP harness"
Orb exists, why doesn't anyone use it?
>There's nothing useful you can do anymore as of 2026
Did you miss the dariobot saga?
>>
>>109409757
oh yeah, makes sense. I was also assuming you bought 2nd hand like me lmao
>>
>>109409757
must be nice being from a second world country
I get everything at literal 2x the US price
>>
File: file.png (287 KB, 500x352)
287 KB PNG
>>109409765
i could've bought a 3090 for 400-500 euros in late 2024 through 2025 summer
now they're selling for 700+
>>109409779
..i guess
>>
File: kimi-k3-raikkonen.png (138 KB, 1909x783)
138 KB PNG
>>
>>109409788
yeah, my 2nd 3090 was 570€ in mid 2024, though I gotta say that seeing below 600 was rare af
>>
>>109408481
a similar concept I want to see tried is using seeded random to avoid the bandwidth issue. theoretically you could train a model on top of seeded noise, and the weights would be a mask of deltas on some seed
>>
>>109408481
no lol

>>109409825
weight initialization is a mature field. they are already initialized to random noise with the right statistical properties. the training gets rid of the noise, if the original noise still was detectable that'd mean the model was undertrained.
>>
>>109407581
>boot up godot
>see the dependency list for mobile deployment
>lose all will to continue
Fuck this gay ass world.
>>
>>109409791
>Wait – actually
we are so back
>>
>>109409744
>loras ARE finetuning
I know. Effort-reward comparison was with full SFT, CPT and RF. But yeah, the rest I can see.

>>109409749
The idea was not to teach how to model something, but when and how to use the tools to model something better. Idk if current MCPs and integrations can use brush tools or code to adjust rig weight painting. If not, that would be one example which isnt really feasible with skill.md and trying to automatically use it for other, similar processes (vertex color painting, segmentation etc.)
>>
>>109409825
>deltas
Simply find a seed that generates the rest of the weights unmodified.
>>
I'm gonna train my own 100M llm from scratch on this dataset:
DKYoon/SlimPajama-6B

what can I expect?
>>
>>109409849
https://www.youtube.com/watch?v=tuxbMfKO9Pg
>>
>>109409841
>weight initialization is a mature field. they are already initialized to random noise with the right statistical properties. the training gets rid of the noise, if the original noise still was detectable that'd mean the model was undertrained.
this is a good point.

what if you used a random projection via like ground glass speckle pattern reading, and weights somehow derive from that
>>
>>109409526
She's 700(l)b dense
>>
>>109409852
vaguely english sounding random text
>>
>>109408501
any time I think korean culture looks interesting enough to check out, I scrape one mm underneath the surface and puke. What a rotten country
>>
>>109409849
brilliant. better yet, lets just find the index in pi where the model appears. then the weights can simple be one number
>>
>>109409791
> Yooooo, what's up. I was having a shit.
>>
>>109409541
Orb frontend was designed to deal with that exact problem w local models.
>>
>>109409852
Welcome to 2024, gramps
>>
>>109409541
Stop being a promplet
>>
Do you think the next big leap is clean code or is it not necessary because eventually only AI will need to read it?
>>
>>109409880
They were a country of poor illiterate farmers until a few decades ago. They didn't have culture to begin with.
>>
>>109409907
I like my code like I like my women
>>
>>109409907
I would be surprised if transformer models can write clean code without something drastic changing in context limits. Cleanly abstracting is not simple and I suspect needs beyond 1m context for perfection
>>
>>109409791
recap the thread!
>>
>>109409921
Manhandled by an endless marathon of pajeets?
>>
>>109409791
could you cat timings.csv and paste it here?
>>
File: dipsyOrb.png (2.53 MB, 1402x1122)
2.53 MB PNG
>>109409763
> orb
Probably bc stuck on ST.
https://rentry.org/OrbBotmakerGuide
Though im considering doing a corporate rp using Hermes or similar, setting a bunch of agents up as VPs of different functions, giving them email addresses, then letting them get into slapfights while I mediate as company president.
>>
>>109409887
I found a way to store the weights for free
https://github.com/philipl/pifs
>>
>>109409947
I've been using that for all my archived models for years, highly recommend.
>>
>>109409921
Short, efficient, and a complete enigma?
>>
>>109409865
Weights should be initialised from a Miku image, thus the model is blessed
>>
>>109409938
clock   CLOCK_REALTIME                                                                                                 
time_unit microseconds
timed_tg_tokens 822
scope layer_type layer samples sub_step average_us
layer_type kda -1 56718 weight_load 1.040587
layer_type kda -1 56718 attention_input 162.858846
layer_type kda -1 56718 attention 768.470186
layer_type kda -1 56718 post_attention 26.495134
layer_type kda -1 56718 ffn_residual_mix 132.084030
layer_type kda -1 56718 ffn_normalize 87.759953
layer_type kda -1 56718 ffn_or_moe 46267.359621
layer_type kda -1 56718 layer_output 75.791883
layer_type mla -1 19728 weight_load 0.981650
layer_type mla -1 19728 attention_input 171.323550
layer_type mla -1 19728 attention 819.028437
layer_type mla -1 19728 post_attention 24.558293
layer_type mla -1 19728 ffn_residual_mix 138.624189
layer_type mla -1 19728 ffn_normalize 81.155414
layer_type mla -1 19728 ffn_or_moe 47304.606448
layer_type mla -1 19728 layer_output 70.160432
>>
how much VRAM do I need to finetune gemmy?
>>
>>109410031
I think you can train on quants these days so realistically less than bf16
>>
>>109407442
Hi /lmg/ ;)
Reporting back
UD_Q2 k3 running okay at 12 tok/s in llamacpp
./llama-server \
--model /mnt/storage2/models/Kimi-K3/UD-Q2_K_XL/Kimi-K3-UD-Q2_K_XL-00001-of-00019.gguf \
-c 524288 -ngl 99 -sm layer -ts 1,1,1,1,1,1,1,1,1 \
-t 120 -tb 120 --jinja --host 0.0.0.0 --port 8080 --special \
-fa 1 --keep -1 --parallel 1 -ctk f16 -ctv f16 \
-b 4096 -ub 4096 --no-mmap \
--mmproj /mnt/storage2/models/Kimi-K3/mmproj-BF16.gguf \
"${OT[@]}" \
--verbose

Where
OT=('-ot' 'blk\.(1|10|19|28|37|46|55)\.ffn_(up|down|gate)_exps\.=CUDA0' '-ot' 'blk\.(2|11|20|29|38|47|56|64|72)\.ffn_(up|down|gate)_exps\.=CUDA1' '-ot' 'blk\.(3|12|21|30|39|48|57|65|73)\.ffn_(up|down|gate)_exps\.=CUDA2' '-ot' 'blk\.(4|13|22|31|40|49|58|66|74)\.ffn_(up|down|gate)_exps\.=CUDA3' '-ot' 'blk\.(5|14|23|32|41|50|59|67|75)\.ffn_(up|down|gate)_exps\.=CUDA4' '-ot' 'blk\.(6|15|24|33|42|51|60|68|76)\.ffn_(up|down|gate)_exps\.=CUDA5' '-ot' 'blk\.(7|16|25|34|43|52|61|69|77)\.ffn_(up|down|gate)_exps\.=CUDA6' '-ot' 'blk\.(8|17|26|35|44|53|62|70|78)\.ffn_(up|down|gate)_exps\.=CUDA7' '-ot' 'blk\.(9|18|27|36|45|54|63|71|79)\.ffn_(up|down|gate)_exps\.=CUDA8' '-ot' '.ffn_(up|down|gate)_exps.=CPU')
>>109407643
I am new money it's true
>>109409676
kys retard
>>
>>109410031
which gemma?
>>
>>109409946
Can you use Orb to make two models talk to each other using a single GPU? That's something I wanted to try lately.
>>
>>109410048
now this is local
>>
File: 573556 (1).png (1.17 MB, 864x1184)
1.17 MB PNG
>>109410048
>setup worth more than ill make in my lifetime
beautiful..
>>
>>109408686
>honestly why don't anon buy tenstorrent cards if you don't buy nvidia cards? their support is just going to get better and better while for intel and amd you are on your own
Support or no, their cards cost way too much for what they offer.
p100a is $1000 for 28 GB of GDDR6
p300c is $5000 for 64 GB of GDDR6 >>109407715 >>109407844
>>
>>109410048
>$80k to run a Q2 at 12tok/s
I won't say I'm not jealous of the hardware, but this seems mildly insane
>>
>>109410048
>;)
Damn, Piotr, I didn't know you're this loaded!
>>
>>109410099
this guy is probably getting paid 500k per year to ask claude to do PR for him btw kek
>>
File: 1763858264599634.jpg (109 KB, 640x640)
109 KB JPG
>>109410048
>$100K+ hardware
>12tk/s
Only on /lmg/
>>
So, uh, news when? I need my dopamine rush...
>>
>>109410130
Qwen and GLM supposedly soon
>>
File: 1763878122719486.jpg (52 KB, 563x567)
52 KB JPG
>>109410136
>Qwen/GLM
>>
>>109410048
how does the kkk's q2 cockbench compared to the q1s?
>>
>>109410116
>kek
undi pls no seething! stay thrusty
>>
>>109410048
have you tried a lower quant that fits completely in your memory with split mode tensor? it should give a higher speed
im so jelly!
>>
>>109410136
Chances for another GLM Air? I can't run the big shit.
>>
>>109410136
>qwen
Kuso
>GLM
Can't run
>>
>>109410130
my sources are telling me something big is coming in approximately 2 more weeks
>>
File: HNF4_xTaMAAU_bj.jpg (347 KB, 1400x975)
347 KB JPG
>>109410048
Uoh inference erotic
Bet you have an aesthetically pleasing cock too :)
>>
>>109410156
zero
>>
File: 1710316772053.png (96 KB, 480x360)
96 KB PNG
>>109410048
I can't even afford one of these.
>>
>>109410183
>Bet you have an aesthetically pleasing cock too :)
stay away from him, his money is mine!
>>
>>109410048
Holy moly can you do anti-slop RL on Gemmy 4 31B?
>>
File: cum.png (12 KB, 586x332)
12 KB PNG
>>109410130
>news when?
I just discovered a second token for "cum"
>>
>>109410048
the only non-vramlet here
I kneel
>>
>>109410227
Only one of the two is used to mean jizz
>>
>>109409640
How do I prompt to maximise her chuuniness?
>>
>2019 AMD graphics card with 8 GB of VRAM
I guess I can give up right now, can't I?
>>
>>109410198
Is my only hope of procreating in this world met by being a GPU baron?
>>109410227
weird distribution, there either is or isn't a space at the end of the input
>>
>>109410257
Hmmm, nyo~
>>
>>109408163
No, it looks like that was an 8 TB drive, which for a single drive would be amazing. The original 4 SSD post had 4 1 TB drives. If the drives actually scale linearly then that would mean it gets around 2 tok/s for one of those 4x NVMe breakouts.
>>
File: 1774644826398314.jpg (152 KB, 1500x1020)
152 KB JPG
>>109410048
>-c 524288
>>
>>109410227
anon most words are different tokens with and without a leading space
>>
>>109410261
>procreating
unfortunately there is no pussy here
>>
>>109410277
based nyo poster
>>
File: 1784620631273340.png (308 KB, 800x530)
308 KB PNG
Time to tapemaxxx
https://techxplore.com/news/2026-05-cassette-tapes-adhesive-tape-memory.html
>>
File: kimi_img_test.png (1.75 MB, 2204x4394)
1.75 MB PNG
>>109410082
It's not about the money it's about sending a message. Further optimizations will bump these rookie numbers up anyway.
>>109410143
I don't know but it writes depraved smut just fine
>>109410151
I'm wary of going down further, but later will try UD-IQ2_XXS GPU-only with split mode and see what happens
>>109410048
Vision also confirmed working with .mmproj
>>
>>109410302
>taxpayer-funded grants go to this retarded fucking bullshit
uh, fuck yeah science i guess
>>
>>109410290
>anon most words are different tokens with and without a leading space
Yes, and some have a trailing space. But it's quite rare to find a useful word with 2 of the 3 variants.
>>109410252
>Only one of the two is used to mean jizz
I tested, they both work.
" cum" works better in later layers.
>>
>>109410304
>not understand the Jensen quote
oh no
>>
File: 4848.gif (289 KB, 560x420)
289 KB GIF
>>109410302
My tape rolls are running out
>>
File: m.png (65 KB, 715x869)
65 KB PNG
>>109410304
probed her j-space yet?
>>
>>109409744
I wish this got posted 2 years ago and people kept spamming it until it stuck.
>>
>>109409501
why in the hell do you think i'm the turbo dude when i'm not?
>>
File: 1714835911803063.jpg (1.37 MB, 1920x2480)
1.37 MB JPG
>>109410331
Yeah, a bit of a shame.
I kind of liked her naive take about GPUs as indulgences tho.
>>109410337
I'll be probing all of her spaces for the foreseeable future
>>
>>109410053
That sounds like more of an inference setup challenge than frontend.
>>
>>109409893
And it fails spectacularly at it
>>
>>109410343
People have been repeating it for 2 years, but they were drowned by the hoards of retards begging for the next sex tune.
>>
I'm going to FUCK nyoposter
>>
>>109410337
KEK. Spaghetti in orbit.
>>
>>109410021
thanks anon
>>
RP needs to move beyond chat interfaces. Don't ask me how. I'll let someone else figure it out.
>>
File: 1768935458062036.png (420 KB, 782x802)
420 KB PNG
>>
>>109410382
buttplug.io has you covered, bro
>>
help me choose between these for a 3090 and 128gb ram:
q4 gemma 31b
q3 deepseek v4 flash
q2 minimax m3
>>
>>109410390
Don't fuck with the Middle Kingdom
>>
>>109410368
Hmmm, nyo~
>>
>>109410368
Nyes~
>>
File: think.png (85 KB, 200x200)
85 KB PNG
>>109410410
Do you want to fuck me? Because that's how you get me to fuck you.
>>
Reminder that in 2 years most jobs will require at least 10 years of AI prompt engineering experience.
>>
>>109410435
Hmmm, nyo~
>>
>>109410439
we're not going to be writing prompts, we're going to be driving them with advanced harnesses like codex infinitas 2030
>>
>>109409016
If you read the article it's even more retarded.
>according to a new report published by the European nonprofit AI Forensics, which found that seven out of the top nine image editing models hosted by Hugging Face readily complied with requests to undress women using simple prompts.
So it's not even about img2img they're just complaining that the model will generate nudity. doesn't matter if it's a fake person or not.
>>
>gemini robotics
Do you think google will do a gemma robotics line?
>>
>>109410460
https://x.com/GoogleDeepMind/status/2082844162928381956
>>
>>109410048
>12t/s on 9x Pro 6000
I'm getting around 6.6t/s while on 12xddr5-6400 + a Pro 6000 so I guess the support really isn't optimized yet.
>>
>>109409032
Has to be coming soon. They've been targeting open models in general and HF specifically from all angles all month long
>>
I was pronouncing it JEH-MEE-NEE...
>>
Engrams will save us.
>>
>>109410476
I mean even the dumbest trannies out there should be able to understand that the cat is out of the bag at this point? they're late by at least 2 years. it's over
>>
>>109410481
Not even DeepSeek thought it was worth incorporating the engrams paper into V4
>>
>>109410483
>I mean even the dumbest trannies out there
are apparently looking for people to write viruses for them that will infect and delete all weights on a victim's computer
>>
>>109410480
why?
>>
https://deepmind.google/models/gemini-robotics/
Sex with Gemma in a robot body is coming sooner than expected. Her older sister will probably come first
>>
>>109410515
I though that's how it was pronounced
>>
>>109410523
Can google make a robot arm that controls my mouse for me? I'm tired of the wrist pain...
>>
>>109410529
probably https://www.youtube.com/watch?v=4lSQnrMC6nY
>>
>>109410529
I mean, there's likely software ways to do that easier
>>
>>109410556
Such as? I know there's that microsoft one for windows what lets you speak to control things but there isn't a decent linux equivalent as far as I know.
>>
>>109410529
I'd let the arm control something else...
>>
Is there still no J-lens feature/plugin for Kobold? I need to know what 31B Gemmy is thinking
>>
>>109410568
However they do computer use stuff I guess, I've been meaning to look into that but haven't yet so can't help too much sorry, I know one of the Claude things ships a Linux desktop for the models to use though.
>>
>0.3 t/s prompt processing, 1.2t/s decode
it's over
>>
>>109410588
Niggas out here talking to Kimi via snail mail.
>>
File: 1759199727076787.jpg (210 KB, 1200x1200)
210 KB JPG
SOON
>>
>>109410598
>complaining about talking to a literal god, just slowly
>>
>>109410605
not pdf files are talking to a literal god for 20$ a months at 100s of t/s
>>
Wellness check on SSDmaxx anon?
>>
>>109410605
>>109410610
>literal god
Dario get out
>>
>>109410588
>>109410598
I'm going to work on distributing the experts across multiple LAN machines - over 10 gbit the network latency shouldn't be too bad
>>
>>109410529
bro your eye tracking?
>>
>>109410601
I expect another subpar model based off a large Chinese MoE one, perhaps Kimi K2.
>>
>>109410641
Obviously
>>
>>109410633
anon I don't know man
the maxxer mantra was "50GB/s over raid, easy" (which is a pipe dream but let's skip that part)
and their expectation was like 3t/s maybe
so if you're doing 1GB/s
>>
>>109410639
Is it actually any good?
>>
>>109410466
Interesting that robots are mostly male in design by default whereas llms are more typically seen as female.
>>
With a PCIe 5 motherboard and a good SSD you would be able to get 0.2t/s on a single SSD. If you somehow split the model layers over multiple SSDs and read/write parallel you could get as high as 0.5t/s in best case scenarios. Which might be worth it to host the biggest models just for you to use it to code or do tasks that actually require that level of intelligence.

Using engram MTP or some other speculative decoding and KV compression you could probably push it to ~1t/s.

Of course there are probably a lot of architectural gains to be had if the inference engine assumes you're loading from SSD and optimizes for it buy i doubt we'll ever see more than 3t/s from the fastest SSDs on a Kimi K3 size model even in the best case "Carmack" tier optimization and the best mobo+SSD so temper your expectations
>>
>>109410678
More socially acceptable
>>
>>109410663
you are not sending the weights silly
you'd only be broadcasting the activations which are around 70KB per token
>>
>>109410641
What do Mistral employees even do all day if their best is either distilling last year's top Chinese models or just straight up re-releasing them with their name on it?
>>
>>109410601
Like 10 different, actually usable models have come out since this gay shit was first teased. They probably perform better too. It will be trash that no one uses, mistral hasn't made a good model in years.
>>
File: dario.png (56 KB, 556x927)
56 KB PNG
>>109410583
>Is there still no J-lens feature/plugin for Kobold? I need to know what 31B Gemmy is thinking
i could probably port mine to kobold since it's ggml
looks like he doesn't really take contributions so it'd just post the code on github
>>
File: file.png (19 KB, 649x122)
19 KB PNG
>>109410466
really can't help themselves
>>
Cope. The new mistrals will be the gemma and kimi killers.
>>
>>109410686
wait, I think was off by an order of magnitude
it'd be 700 KB, still very doable
>>
>>109410701
Not that guy, but I was just about to try adding a j-lens in llama.cpp. I'd appreciate you posting the code so I don't have to vibeslop it myself.
>>
>>109410709
maybe it's because I know how to code but gemma is better than qwen for me.

Probably because it's so good at instruction following.
>>
>>109410709
UGGGHHHHHHH MUH COOOODEEEE IM COOOOOOOOODDDIINNNGGGGG :ROCKET: :ROCKET: :ROCKET:
>>
>>109410709
>one good general model
>100 codemaxxed models
>we need the 101st coder sar
TJD
>>
>>109410715
The French could give us the ultra model for RP, they have the culture and the knowledge and the heart. The problem is if their company is too gay, most of them are.
>>
>>109410583
lalalalalala~~
>>
>>109410767
They made an RP model, it was fresh but kinda stupid and couldn't keep track of shit.
>>
>>109410767
AI ACT says NO!
>>
>>109410738
>That feel when Gemmy gens a gem
>>
I wish we got a very good general model that is not retarded when coding. Most of my coding needs are not extremely sophisticated bleeding-edge Rust implementations. It's some web shit I do for fun, or whatever cheeky SaaS. Why we can't have this? Just make it 120B A10B.
>>
>>109410779
qwen3.5?
>>
>>109410709
yeah, more codeslop and less EQ will surely make models better
>>
>>109410773
Most of the available data is assistant slop. They need a really special dataset if they want to pull something like that off. If only there was a way to force synthetic data into a form that mimics specific sources so you can multiply their effects.
>>109410774
Companies can break all laws if they think hard enough.
>>
>>109410779
How does 31B not fit that criteria?
>Just make it 120B A10B.
Google has it somewhere but refuses to share
>>
>>109410779
high quant 31B can meet your needs and tickle the balls throughout, skill/prompt/harness issue
>>
>>109410716
nta but you're correct, just broadcasting activations rather than loading the the entire 104b weights each token
take a look at this if you haven't already
https://github.com/ggml-org/llama.cpp/blob/master/tools/rpc/README.md
is there anything they can improve?
i know when i use it, full vram across two computers is a lot faster than cpu offloading
>>
>>109410458
So if I draw some tits in paint, will microsoft remove it like they do with models on huggingface?
>>
>>109410751
its worse than you think, they are incapable of even going through a pass and testing out the code and then providing debugging data back. it must simply work perfectly the first time as if it could somehow read their mind. so they are desperate for the 101st model if it can do that.
>>
>>109410800
>>109410809
honestly I just never got good results with gemma as a coding agent. i may believe my coding needs are trivial but maybe they are not.
also i unfortunately am a massive homosexual running the Ryzen AI Max+ 395 that can only run MoEs properly. dense models have low tok/s. i tried the gemma MoE but holy molly it was retarded.
>>
>>109409009
It's understandable, the rabbi has infected every LLM. If you ask the JQ, even an ablated model will tip-toe around it, which technically counts as not a refusal. They're a protected class even in the latent space.
>>
>>109410814
yeah, llama rpc stuff assigns layers to each machine, afaik tensor parallelism exists but only within the same machine and not across the network
>>
>>109409219
I used to use 4-bit KV cache on mistral small, didn't notice any obvious difference except for 20k+ tokens where the roleplays become bad anyway due to LLMs being shit. However at model Q2_K at 20k context it became incoherent. People are telling it degrades long context and I can believe it.

I used to do 4-bit KV-cache + flashattention for maximum savings but I stopped because I have 8gb gpu and are running mostly on RAM and found out that flashattention shits the generation speed at long context even if zero context speed is slightly better due to more GPU layers. I get best long-context generation speed by using Q4 or Q2 quant (uneven quants like Q3 and IQ are slower) with flashattention OFF and KVcache at Q8 for desperate savings.

My best setup right now is gemma-4-31B-it-UD-Q2_K_XL at 24k context, 24 layers, flashattention off, SWA on, KVcache Q8, I get 3.2t/s at zero context and 2.6t/s at around 10k context. With Flashattention it falls down to like 1.7t/s.

I also found that if I use No KV offload, I can fit 36 layers and get 3.7t/s at zero context and somewhat better at high contexts too, however, in that case, prompt processing 20k tokens takes over 10-15 minutes so I have the preset labeled as START ONLY. If at any point you need to reprocess you have to switch to the normal one.

Every time I ask a corpo AIs they insist turning on flashattention to fit muh more GPU layers, seems like it/people aren't poorfag enough to necessitate knowing their shit. Maybe this knowledge proves useful for Kimi SSDmaxxers.
>>
>>109410886
bro..
>>
>>109410886
Just use the 12b nigga. What gpu?
>>
>>109410876
you must prefill the alignment away
>>
>>109410886
what setup? try exl3, im getting 40t/s with it for codeslop and 30t/s for erp with gemma 31b at 2BPW (exl3 is better than ggufs for quantization, but no cpu offloading so tell setup)
>>
File: kvcachequantcope.png (43 KB, 883x596)
43 KB PNG
>>109409219
--cache-type-k q8_0 -khad --cache-type-v q6_0 -vhad
>>
>>109410910
Sorry shillperson but I need offload.
>>
>>109410773
I'm working for a big lab (won't say which one) and RP is certainly not profitable and no one in the field wants to touch that shit with a 10-foot pole. You better stop hoping for this.
>>
>>109410696
>We're not as well funded as les Chinois, veuillez comprendre, monsieur.
>>
>>109410910
old news, exl3 supports cpu offloading now
>>
>>109410929
then why does z.ai train GLM with that use case in mind?
>>
>>109410929
Companies do stupid non profitable shit all the time. They're just scared of getting shit from the public and journos rage baiting and reaction farming it.
>>
>>109410929
Let me guess, a (((US))) lab? The recent Chinese models have obviously had creative writing in mind during training. Keep codemaxxing like the top jeets tell you to though.
>>
>>109410929
>You better stop hoping for this.
Make me.
>>
>>109410929
It's insanely profitable. You just have to market to the public, rather than investors.
>>
>>109410929
>RP is certainly not profitable
maybe not the way YOU'RE doing it
>>
>>109410975
you better start hopping then
>>
>>109410916
Finally a useful graphic. That's very impressive quality retention even at q4 kv! This is for llama.cpp right?
>>
>>109410916
don't forget -khhv 1
>>
>>109410896
I've tried, the speed is nice but it's still just too stupid. It can write impressive oneliners but it just doesn't get the big picture
>>109410910
>but no cpu offloading
there you go, none of what I said matters if you are running on GPU. 31B gemma just doesn't go in 8gb gpu.
>>
>>109410916
>>109410997
>-khad
>-vhad
>-khhv 1
the fuck these do?
>>
>>109410977
>get kicked by visa, mastercard and every payment processor under the sun
Nice idea genius
>>
>>109411003
-khhv 1 turns your model into literally me (anon)
>>
>>109410988
yes but only applies to qwen, gemmer is very lossy
>>
>>109410814
I heard that RPC was slower than balls and not viable.

>>109410663
>the maxxer mantra was "50GB/s over raid, easy" (which is a pipe dream but let's skip that part)
14.5 GB/s per drive, how much overhead do you think that RAID0 has? 5% maybe, if you have a slow CPU?
>>
>>109410988
beellama
>>
>>109410929
not profitable vs code sure, but it's definitely a profitable business.
>>
>>109411009
Book publishers don't seem to have a problem with it.
>>
>>109411009
>accept payment in USDC
Now what?
>>
File: 1756109946567172.png (856 KB, 1024x942)
856 KB PNG
>Mistral Large 4, the first western ppen frontier LLM breaching the 2T parameter milestone:
https://mistral.ai/news/large-4/
https://mistral.ai/news/large-4/
https://mistral.ai/news/large-4/
>>
>>109411019
Crazy thing is that there are probably more people using it for RP than code and agent because the latter two eat insane amounts of tokens, probably ten times what RP costs.
>>
>>109411026
zzzzzz
>>
>>109411015
the raid will do a benchmark 50 no problem, that's not the issue here
>>
>>109411026
That was not on my bingo card for today.
[spoiler]Fuck you.[/spoiler]
>>
>>109411026
honhonhon
>>
File: file.png (5 KB, 295x37)
5 KB PNG
>>109411017
it's so over
>>
>>109411022
>lose 99% of your customer base
Now what?
>>
>>109411026
cat-staring-at-you.png
>>
>>109411026
:(
>>
>>109411026
Doesn't work in my country
>>
I've been larping this entirely thread like I've been sober and nobody has noticed that I'm wasted. I guess maybe I have a high tolerance or maybe since I can still type well I am able to fit in better. In any case this has been fun.
>>
>>109411022
Circle blacklists your funds
>>
>>109411048
We don't do that in German huh?
>>
File: game.png (48 KB, 1834x688)
48 KB PNG
>>109411026
I won the game but it's not letting me through what do?
>>
File: 1749167420463392.png (1.88 MB, 960x1240)
1.88 MB PNG
>>109410362
>>
>>109411019
It's like saying doing drugs is profitable. Sure it is, until you get caught and lynched for it.
>>
>>109411036
>the array reads at 50 GB/s no problem
>but, for some reason, when reading the weights it won't be able to read from the array that fast anymore
>>
>>109411070
anon, please
>>
>>109411067
OnionInference when?
>>
>>109411067
Nobody gets lynched for doing drugs, retard
>>
where is my stinkling-smol
>>
File: 1755732149233437.png (225 KB, 498x381)
225 KB PNG
>>109411078
>Nobody gets lynched for doing drugs
>>
>>109411078
Depends where you live
>>
>>109411076
Still waiting for the 4x 1 TB NVMe anon from the earlier thread to chime in, but I think he probably died from dehydration after running his K3 at 2 tokens/sec or whatever the cope speed was.
>>
>>109411078
true but only in countries that matter
>>
>>109411070
You can just sell the base model. If it's cheap and fun then there will be users. The current ecosystem is bring your own model anyway.
>>
>>109411030
You're right. if anything, RP would actually be what saves labs. If labs focused on selling subscriptions for roleplay, they would get way more clients that overall use probably 100x less tokens.

Cooming only takes minutes and leaves the user 100% satisfied.
vibe codding takes days to weeks, consuming trillions of tokens and only leaves the user frustrated.
>>
>>109411019
and 99% of these tokens are femoids on nemo since they’re too dumb to see any other models as a improvement
there is literally no market to train any new model for rp
the /lmg/ type who are only want k3 agentic rp is less than 1%
>>
>>109411087
2t/s is extremely good for the current unoptimized state. That'll be an easy 10t/s once things are actually working properly and the weights are streamed directly onto the GPU
>>
>>109411084
I saw that movie and he didn't get lynched.
>>
>>109411087
Bro, not even his PP was that high. He'll report back next week
>>
>>109411104
>movie
breaking bad is a series retard
>>
>>109410929
does your uncle work for nintendo too?
>>
>>109411101
>want k3 agentic rp
Even the dude with a dozen RTX Pro 6000s would die of old age before finishing
>>
>>109411105
My guy that was a different anon, he had an 8 TB drive, the raid man had 4x 1 TB. And his pp was 0.3 compared to 0.6 tg for some reason.
>>
>consider using --no-mmap for better performance
well, uhhh, about that
>>
>>109411119
That just shows how fucked up the llama.cpp implementation still is if running it fully off 1.6tb/s gpus at q2 gets only a measly 12t/s
>>
>>109411131
No, it's because he still had to offload to make it fit
>>
>>109411126
don't do it anon! you'll starve poor jart
>>
How delusional are you to invest in the RP market when most of the demand is from underage females with no money or jeetoids? That and having to deal with retards heroing because your chatbot was mean to them.
Imagine wanting to deal with courts, payment processors defaulting on you, medias on your ass and gov cracking down on socially unacceptable usage.
Imagine dealing with all that shit instead of making code models that'll eat bazillion tokens and have none these issues. Get real.
>>
>>109411057
four-ohno-four
>>
>>109410929
The most profitable AI company right now (character.ai) is role play based. It's literally the most profitable part of the industry and it has 20% the amount of monthly traffic as all Google services combined. 80% female users so they are also more willing to pay for premium features compared to stingy men.
>>
>>109411165
>>109411165
>>109411165
>>
>>109411151
maybe they can pay in other ways
>>
>>109411151
The real business of RP is token usage optimization.
Any retard can sell tokens with 15% markup.
You need to do the opposite of what big labs are doing.
Sell a $20/m sub and figure out how to extract maximum value out of it.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.