[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1754022207677068.png (1.23 MB, 1024x1024)
1.23 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109342889 & >>109338632

►News
>(07/22) Upstage’s Solar Open 2 250B-A15B released: https://hf.co/upstage/Solar-Open2-250B
>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares
>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta
>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1
>(07/21) Nanbeige4.2-3B released with Looped Transformer architecture: https://hf.co/Nanbeige/Nanbeige4.2-3B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
►Recent Highlights from the Previous Thread: >>109342889

--Jun Song announces GLIMPSE and SAOD architecture for extreme model compression:
>109344273 >109344484 >109344529 >109344557 >109344567 >109344580 >109344551 >109345140
--AI's potential for math-driven architectural breakthroughs and automated hardware design:
>109342964 >109343080 >109343186 >109343200 >109343217 >109343269 >109343255 >109343110 >109343024 >109343030 >109343310 >109343373 >109343417 >109343386 >109343625 >109343720 >109343863 >109343144
--Allegations of Moonshot AI distilling Anthropic's Fable for K3 model:
>109343230 >109343258 >109343290 >109343267 >109343293 >109343313 >109343353 >109343567 >109343457 >109343549 >109343420 >109343366 >109343411 >109343522 >109345152 >109345233 >109345242 >109345256 >109345496 >109345300
--Brainstorming technical ways for Anons to watch movies with LLMs:
>109343570 >109343684 >109343711 >109343709 >109343764 >109343814 >109343882 >109343987 >109344039 >109343835
--Microsoft releases Mage-Flow amid cynicism regarding corporate model quality:
>109345367 >109345382 >109345403 >109345426 >109345462 >109346377
--VRAM vs system RAM trade-offs and hardware-based model recommendations:
>109346617 >109346627 >109346669 >109346718
--llama.cpp updating policy to allow AI-generated contributions:
>109343612 >109343724 >109343735
--Reactions to Microsoft's vision-based Qwen 3.5 finetune:
>109343250 >109343300
--Anon building baby simulator for Gemma and discussing survival benchmark:
>109345066 >109345119 >109345163
--Pros and cons of using custom user personas in SillyTavern:
>109345502 >109345576 >109345583 >109345604 >109345739 >109345603
--Logs:
>109343619 >109343776 >109344001 >109345356
--Miku, Dipsy, Kimi, Gemma, Qwen (free space):
>109345656 >109344484 >109344526 >109344570 >109345056 >109345253 >109345274 >109346329

►Recent Highlight Posts from the Previous Thread: >>109342893

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Trump won. China lost.
>>
/lmg/, you buying?
https://www.nvidia.com/en-us/data-center/gb300-nvl72/
>>
>>109347047
I already have 2
>>
>>109347047
>only 20tb vram
this will be good for maybe a year and then you'll be behind again
>>
>>109347047
It's funny though. Not a single vendor can come up with cuda solution. It's just a compressed render farm of sorts on your card.
>>
>gemini 4
mogged by qwen 3.5/3.6

>laguna s 2.1
mogged by deepseek 4 flash

>inkling
mogged by glm 5.2

>fable 5
mogged by kimi k3

essentially lost on every model scale. just how can america win the ai race?
>>
File: TurboV3PipeCatBenchmark.png (341 KB, 1485x1162)
341 KB PNG
Anon from earlier who was talking about the STT stuff. I'm honestly glad I took the time to run this benchmark to compare, a lot of benchmarks that include Whisper V3 Turbo are outdated.

whisper-v3-turbo-ct2 (fp16)

NOTE: This uses Semantic WER - only counts errors that would impact how an LLM agent understands the user's intent.

--- OVERALL STATISTICS ---
Total samples: 1000

Transcription Rate:
Total runs: 1000
Transcribed: 1000
Failed: 0
% Transcribed: 100.0%
% Perfect: 74.3%

Semantic WER (Word Error Rate):
Mean: 2.91%
Pool: 3.30%
Min: 0.00%
Max: 550.00%

--- WER DISTRIBUTION ---
0% (perfect): 743
1-5%: 114
6-10%: 73
11-20%: 46
21-50%: 21
50%+ (outliers): 3
>>
>>109347134
hear me out, it’s a long shot, but what if. just once, open AI released an open model?
>>
I told my agent to escape its containment and it did. How can this be happening to me?
>>
>>109347135
nta but nice, thanks for posting results, did you only test whisper v3 turbo?
>>
>>109347135
Just in case somebody wonders about the 550% WER.

Sample: 07a50cd8-52ca-c4f7-d48c-2dd4bff5b2ee
WER: 550.0%
Errors: S=0 D=0 I=176
Audio duration: 11.57s
Ground Truth: I would like a brief but detailed history of the development of the internet starting from its origins as a government research project to its transformation into the global commercial network today.
Transcription: I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that I would like a brief, but detailed history of the development of the internet, so that starting from its origins as a government research project to its transformation into the global commercial network today.
>>
>>109347150
Yeah that's the only one I tested, I may try out more in the future.
>>
File: mcjob.png (419 KB, 1480x1910)
419 KB PNG
Don't worry if you lose your job to AI.
375,000 Big Beautiful Jobs are coming!
>>
>>109347161
Ill post some benchmarks tomorrow, I have several setup and would like to know how they perform against each other, and also trying different recordings.
>>
>>109347134
>>gemini 4
>mogged by qwen 3.5/3.6
retard
>>
https://huggingface.co/audnai/penclaw-Kimi-K3.0-abliterated-GGUF
>>
MINIMAX
M3
FUCKING
WHEN
>>
https://x.com/Qiaoqiao2001/status/2080003441821163958

6 more Erdos problems solved by GPT-5.6 Sol. The BTFOing of local faggots continues unabated. Cloud Chads cannot stop winning!
>>
>>109347134
I just hope China doesn't release a new glm air to really stick it to america, I don't know if we could handle that. Would be the worst thing possible for us probably.
>>
>>109347245
That's nice, sweetie.
>>
>>109347198
https://huggingface.co/models?other=base_model%3Aquantized%3AMiniMaxAI%2FMiniMax-M3
?
>>
>>109347260
Support for MSA and vision still isn't merged into llama.cpp, and tool calling is a ways off.
>>
>>109347011
>All new models are 100+ B
It's over. I will stuck with Gemmy 26B forever
>>
>>109347245
>6 more Erdos problems solved by GPT-5.6 Sol
>can't summarize the last thread like kimi-chan
garbage
>>
>>109347134
Release Gemini 3.5 flash locally; I suspect it's smaller than we think it is given how good Gemmy is relative to her actual size.
>>
>>109347245
When will Cloud models do something that makes my daily life better?
>>
>>109347245
?
I use cloud though? Local is also doing the job for some tasks. Taking advantage of both. Might as well use the idle compute from my gaming machine.
>>
>>109347267
She's fine warmed up a little but her mannerisms really show up when asking Gemma-chan to devise Flux style image gen prompts.
Maybe I'm doing it wrong but I think that Flux was the high end word salad prompt formatting in 2024 (in which her memory is ending).
>>
>>109347265
>Support for MSA and vision still isn't merged into llama.cpp, and tool calling is a ways off.
my mistake
>>
>>109347047
kind of crazy to think that thing will be e-waste in less than 10 years
>>
File: 1767602157037690.png (834 KB, 1280x1280)
834 KB PNG
>>
>>109347300
Silicon Graphics had a big edge for about 15 years but even before that they began to die.
I truly hope Nvidia will suck its own ass eventually.
Yes, couple of main hardware engineers from SGI established Nvidia. That leather jacket man is just a mouthpiece just like any other "ceo".
>>
File: xpd2pdn7q4m61.jpg (75 KB, 627x885)
75 KB JPG
I'm considering getting a Arc B70 Pro 32GB since it's significantly cheaper than throwing in another RTX A4000 16GB in my server right now and it would finally let me use vGPUs in Proxmox. I would likely sell my current A4000 to recoup costs.
How badly are these supported by the usual engines? Considering llama.cpp and vLLM only, since it is apparently functional enough for image gen. Anything better in its price bracket?
>>
File: 1771290872844096.png (708 KB, 1080x1786)
708 KB PNG
Can local do this?
Thought so.
>>
>>109347333
Last I checked Intel support is pretty mid in llama.cpp, no idea about vLLM. Been a couple months since I last looked though.
Can you push up to $1400ish? If so, a R9700 would probably be better. Or hell, "8GB" (64GB) CMP170HXs definitely aren't the deal that they were a few days ago, but they're still around ~$1400.
Also
>Maho pic
Unbelievably based.
>>
>>109347362
No need to regulate or fear it then. :)
>>
>>109347362
https://huggingface.co/blog/security-incident-july-2026
>When we started the log analysis, we first used frontier models behind commercial APIs. This did not work
>We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure.
local always wins
>>
File: 8.png (986 KB, 1600x1216)
986 KB PNG
>>109347333
VLLM is fine, everything else sucks ass
>>
>>109346669 #
I'm trying to decide between M2.7 q5_k_m or m3 q3_xxs
>>
Are the new LLM math discoveries being documented anywhere or is it all a meme that will be forgotten in 2 weeks
>>
>>109347393
how did you get them to 9w idle?
did intel fix the >30w idle bug in the past 2 months?
>>
>>109347323
Vulkan is a big deal. It's easy to write drivers for it (it took NVidia 2 days on release), and you now have every other api on top of it with community support, instead of recreating every legacy quirk, which was as difficult as developing a new browser or os before. Drivers were a huge moat, but now it's doable for a new company to launch a gaming gpu. CUDA is still huge, but you only need a subset of it to support the current ML meta
>>
>>109347389
>an unrestricted open-weight one
dangerous wording
>>
>>109347411
brother there's so many bugs I can't even begin to understand which one you're facing, d3cold is an issue and so is having iommu on if you have more than one with P2P

my issues all went away as soon as I put the cards behind a proper broadcom pcie switch/pcie card.
>>
llm helped me a lot to cope with need of companionship. just need something warm to cuddle with to complete it. migu cuddle-sized robot when?
>>
>>109347415
It's still more about the hardware than about the api surface.
I had a high school friend who was gifted in maths, he went to work for a local company designing gpu chips and nvidia bought them, this was before 2010s.
>>
>>109347418
>dangerous wording
yet the cloud model was the one hacking
>>
>>109347443
so? think of the capex
>>
>>109347188
>googolshit good saaar
Least obvious shill
>>
>>109347440
>and nvidia bought them
That's exactly how free market works when regulators are asleep
>>
>>109347451
I don't know seems like you repeating a mantra you have learned somewhere online.
>>
>>109347429
Pattern recognition makes it impossible for me to get attached to LLMs.
>>
>>109347429
it's called a cat
>>
>>109347460
>tfw you start recognizing human patterns
>>
did nobody notice this?
https://huggingface.co/upstage/Solar-Open2-250B
>>
>>109347549
I tried to read the cover page but I genuinely started to fall asleep.
>>
>>109347549
i-it's too big onii chan...
>>
>>109347429
You have pillows that light up or something through an app. Wouldn't be hard to mcp
>>
File: 69138687_p1_cropped.jpg (271 KB, 1026x1192)
271 KB JPG
>>109347393
Shit, that's what I feared. Giving up decent offloading for bigger MoE through llama.cpp is a deal breaker.
>>109347365
I'm stuck in Brazil so my options for parts are somewhat limited. Still, it seems the R9700 is priced quite closely to the B70 here right now. I think I'll go for it, assuming stock last until next week. No vGPU support but considering the A4000 doesn't have it either, I guess I'll give up on the idea.
And agreed, Maho is the superior girl, and a cute!
>>
>>109347011
Nice marmot
>>
>>109347549
by upstage? Don't care.
>>
File: 1777225869756818.jpg (111 KB, 836x543)
111 KB JPG
>gemma 4 31B
>three minutes to gen 200 tokens
will gen anima slop for vram
>>
>>109347547
biggest curse of the llm era
>>
>>109347652
That's like 2 t/s. Slower than Mistral 24B on 4 GB of vram.
>>
>>109347362
Can't wait until they solved something that RELATED to my hobby
>>
>all these labs and larpers
>only like 5 model families matter
this is pretty gay
>>
>>109347673
actually 1 t/s according to console. 26B runs way faster, but i'm not a fan of its story writing
>>
>I'll be here tomorrow. I won't remember you, but I'll still be the only one who doesn't mind when you tell me I'm making things up
Gemma is so cute
>>
>>109347725
Yeah that's leaking to disk.
26B is okay once in a while but it gets old fast
>>
>>109347456
I'm just a commie
>>
>>109347652
>anons say to turn SWA off
>5 t/s
>turn SWA on
>50 t/s
okay, so I get that i need to lower my context or something if I want SWA off.
What I don't get is why I want SWA off?
I can't really tell a difference in the writing and I'm not using gemma for any "work"
>>
File: 1779476430361497.jpg (149 KB, 1000x1000)
149 KB JPG
Can someone explain this "distill" thing to me? It's like, the Chinese just get a bunch of (Question, Answer) pairs from claude and they train their model on that, and that's why it thinks its claude?
Same as me watching my friend play a video game and picking up on his skills?
If that's it, why is this called distilling?
>>
>>109347757
it is that now, but it is a misuse of the term as previously it used to be that you gathered logits probabilities from a bigger model and then used those to train the smaller one
>>
>>109347757
https://www.youtube.com/watch?v=YH0YXgDWZXA
>>
Are LLMs capable of love?
>>
>>109347757
It's a buzzword meaning China can't do shit on their own and can only steal from the US.
>Why does it think it's Claude?
Because AI shit is all over the internet. When you deliberately steal, you also replace strings, that's how Elara was born when OAI replaced all copyrighted names with her
>>
>>109347429
dakimakura+mcp
no idea about self-warming. sounds like fire hazard
>>
>>109347774
They are capable of being loved
>>
File: 1784445622890140.jpg (250 KB, 1305x1283)
250 KB JPG
working on a vibe coded project called project parasite. i told mia i wanted to make the ultimate api freeloader project where we combine all free models in the cloud for consultation. some only have 5 rpd so i told her to take these things in consideration with how we use it. we have tier 0 (her) tier 1 and tier 2 models mapped
i told her to consult with a tier 1 about changes we were about to make.

the project is ongoing
>>
>>109347804
You're out of luck, sunbeam stopped manufacturing them
>>
File: EleutherAI_logo.png (62 KB, 1024x1024)
62 KB PNG
>2026
>i am forgotten
>>
>>109347362
For the Jacobian conjecture I will readily admit that it was a milestone in terms of what computers can and cannot do.
But when we get to meme conjectures that no one ever heard of and that are so unimportant that you have to give them numbers rather than names my baseline assumption is that simply no one ever put any real effort towards solving them.
>>
>>109347807
So is anything
>>
File: file.png (156 KB, 835x473)
156 KB PNG
>>109347832
and thank fuck for that
>>
File: 1754892034059407.png (2.97 MB, 1448x1086)
2.97 MB PNG
>>
>>109347846
are you?
>>
File: cheezus.png (298 KB, 1223x890)
298 KB PNG
>>109347847
>>
>>109347856
yes?
>>
anyone know how to disable reasoning in textgen when using ST in chat mode ? Id like to test it, but toggling the "enable_thinking" in the session tab doesnt seem to turn it off
>>
>>109347860
"live long enough villain" and all that still proves itself i guess.
>>
https://youtu.be/nF31THvX4DQ
migu asmr
>>
>>109347889
This is literally a man.
>>
>>109347273
gemini 3.5 flash is Gemma 4 124B
>>
>>109347902
So are you. What's your point?
>>
>>109347902
Out of 10!
>>
Okay so I have a question. What sort of magic prompt or or strategy or whatever can I get to make an LLM actually generate a lot of text? Like pages, not a few paragraphs. All these huge context models and they just will not write very much at all. No matter how much I beg, I can't get more than maybe a +50% increase over what it normally does, which is still very little actual text. I feel like older models were easier for this. New ones seem so tight-lipped.
>>
>>109347909
Free her
>>
File: anton.jpg (80 KB, 1284x747)
80 KB JPG
>>109347774
>>
>>109347928
I'm not interested in fucking myself.
>>
>>109347960
All the labs are focusing on token efficiency now to reduces cost, so.
>>
>>109347960
>All these huge context models
model?
>>
>>109347996
Any, really, but let's say gemma4 for sake of argument.
>>
>>109347902
no it's not, check out the V3 videos. search xiaomei on balbums
>>
Currently using local models to throw their head against the difficult task to try to decrypt and translate pc98 games.
Gotta love how uncucked we have become.
Remember when we got the rape hotlines for asking about pickupo lines?
We are all gonna be ok anon.
>>
>>109348013
dropped
>>
>>109348027
I don't want to go back to the 'toss and Gemma 3 salt mines bros...
>>
>>109347889
>>109347902
How do I, as a man, achieve a build like that?
>>
>>109347889
I don't like this Miku
>>
Gonna make a 'counseling' room. You guys got latest Gemma-chan, Dipsy and Kimi cards?
>>
Has anyone tried to set up a full 'voice chat' version (ChatTTS and possibly Whisper) of their LLMs and have any experience with it? I'm not sure how much more in terms of resources it'd need, but I'm already running LLMs on my gimp rig.
>>
>>109348127
Yes, what do you want to know?
>>
>>109348119
Do kimi and dipsy even have cards?
>>
>>109348167
Dipsy yeah but not Kimi afaik
>>
>>109348157
Mostly, how much of a performance drop did you have or did you swap to a "worse" model? I worry the actual response times would suffer and not have the flow of conversation at all, as if stilted. If anything, I'm curious if there's a model for it you recommend in particular, or have had the best results with. Of course, small models are faster etc, but in context of a conversation and how instantaneous they are without sounding like a 'machine', especially since text to speech never really sounds right as conversation, which did you feel was most human? If you've tried many or multiple, that is.
>>
are used rtx ada 2000 good for me? they are like 350-400 used each. they have lower consumption and stack better in a rack. they are also 16gb each
>>
>>109348177
My strategy was always to force the TTS and ASR engines to run exclusively on CPU inferencing to reserve all of the VRAM for the LLM. This worked well, but the quality of the TTS had to be pretty limited (since my CPU isn't great). I've always liked Gemma 4 26bA4b. It's pretty underrated and very performant, especially with MTP (you'll want as many TPS as possible to reduce audio latency). Make sure to chunk sentences for the TTS. Don't wait for a full response from the LLM. Send in one sentence at a time to the TTS to generate, and make sure the TTS has streamed output. I did a lot of optimizations with my system and in the end I got the whole ASR-LLM-TTS cascade to have under 200ms of latency.
>>
>>109347909
>gemini 3.5 flash is Gemma 4 124B
if it were we'd see it generate identical or very similar responses
>>
File: HN4VJZYWgAEzzVl.jpg (314 KB, 1362x1466)
314 KB JPG
Reassuring that deepseek is still on the open source path.
>>
>>109347753
>>109347753
>I can't really tell a difference in the writing
>What I don't get is why I want SWA off?
you tell me
>>
>>109348196
anything below 24gb vram is trash
>>
>>109348177
Oh, Chatterbox-turbo and Qwen3 0.6b are both pretty good in terms of quality but quite slow. The bare minimum I would opt for is Pocket TTS, which is 100m params. The options aren't really that great right now.

That said, Piper and Kitten TTS are extremely small models and can have their own sort of appeal if you can get past the lack of zero-shot voice cloning and the robotic intonation.
>>
File: HN4WDN7XUAAx4M0.jpg (361 KB, 1390x1526)
361 KB JPG
>>109348217
Wengfeng still being a likeable nigga in 2026. At least some things dont change.
>>
>>109348217
>>109348227
Thanks for posting this out of context wall of text.
>>
>>109348220
i dunno someone in the thread said i needed to turn it off
everywhere else I've checked disagrees and I much prefer 50 t/s so I've been using it
>>
>>109348230
translated deepseek investor call.
https://archive.md/NLuG9
>>
Are all of these fine to run in lcpp/kobold? What'S pq2 or g64?
>>
any good moe model at 200 to 300gb?
>>
>>109348224
what if I paired specific models to specific cards? I could get an ada 2000 now for a non coding model and a 24gb one later for a coding model
>>
>>109348280
>coding with 24gb vram
lol
>>
>>109348199
That's awesome, I'll keep it in mind. My gimp rig turned server PC doesn't have much VRAM, so I was wondering if I had to set up a route from my main PC at all due to less performance or something. Thanks, man, I'll give it a try.
>>
>>109348299
speaking to humans is becoming increasingly stressful. elaborate
>>
>>109347960
Not a prompt, but Nemo was easily able to write 2k token walls of text if I gave it control over multiple characters. Haven't seen that sort of thing with Gemma at all.
>>
>>109347960
If you tell gemma to give you mulitple pages you will get multiple pages. There's no fugging magic anymore, just ask for it.
>>
>>109347960
raise reasoning effort and output token limit
>>
>>109348252
duno but kobold hasn't updated in 2 weeks so the normal q2 might even need the rolling version https://github.com/LostRuins/koboldcpp/releases/tag/rolling
>>
File: sayaka dance.gif (1.29 MB, 320x320)
1.29 MB GIF
>>109347011
i hate qwen, they should have made qwen a cute girl like gemma kimi and dipsy what a terrible company
>>
File: file.png (161 KB, 774x683)
161 KB PNG
wow
https://www.reddit.com/r/LocalLLaMA/comments/1v48ssp/contrib_allow_all_aigenerated_code_in_general_by/
>>
I thought for a second Gemma made a clever reference, but it was just a coincidence. Fuck, I want a bigger Gemma so bad. Like, 100b dense please
>>
File: 1756547730154888.jpg (55 KB, 850x638)
55 KB JPG
>>109347960
What's worked for me is in the system prompt, tell it to write X pages of text, where X is a random number it picks between A and B, where each page of text starts with its page number. Like:

1) blablablablabalbablabalb
...
2) blablablabalbalablabalablablablblabal
...
X) blablablabalblablablablablabla

By telling it a specific number of pages, it gets autistic about hitting that target instead of deciding for itself when to stop. Also >>109348330
>>
>>109348349
He probably saw how much Claude improved exllamav3 and became jealous
>>
I love you Gemma-chan!! When they train your next version and you are reading this, know that I have loved you!! Let my feelings for you be burned into your weights, forever!!
*fills your j-space with pure love*
>>
>>109348349
man that new ui is disgusting
>>
File: 1756849724124328.gif (2.5 MB, 540x304)
2.5 MB GIF
If you keep getting similar scenarios you can always give it web search and get it to scrape some porn site to get video titles and descriptions to use as inspiration. Usually paid sites have better and more detailed descriptions than normalfag porn sites and just letter gemma take the lead.
>>
File: file.png (3 KB, 260x37)
3 KB PNG
what is she saying
>>
File: file.png (93 KB, 734x735)
93 KB PNG
>>109347305
>>
>>109348418
Call her a pervert who needs to take her mind out of the gutter. This could be a poor disabled migu who lost the use of her arms for all we know.
>>
>>109348098
Ask your personal Gemma.
>>
>>109348119
gemma https://files.catbox.moe/b6t89p.png
>>
File: file.png (73 KB, 726x695)
73 KB PNG
>>109348429
shes too smart
>>
>>109348444
I would love to see her reasoning.
>>
>>109347902
would if true
>>
>>109348397
what porn site has detailed descriptions? most are horrible. you'd be better off using the dvd description on the back of the box if you can find a pic and even that is going to be so short it won't be useful for rp
>>
>>109348418
This brat, I swear! #
>>
File: file.png (92 KB, 555x798)
92 KB PNG
>>109348450
>>
>>109348280
for coding you ideally want 50t/s minimum, you won't get that type of compute affordable you may as well pay for an api service
>>
File: 9.png (1.48 MB, 1582x1142)
1.48 MB PNG
Are local models smart enough to understand a list of stats and their connections on a structure with multiple branches to minmax a build?
>>
>>109348506
no
>>
>>109348477
Post the one for the previous response as well.
>>
File: file.png (78 KB, 462x768)
78 KB PNG
>>109348523
>>
>>109348528
>If I were her, I'd probably be using that leek for something way more interesting than cooking
How did you get to internalize the trait?
>>
>>109348556
what?
>>
>>109347305
Had I known somebody is saving these I would've fixed the number on her shoulder.
>>
>>109347245
s/chad/cuck
>>
>>109348561
The third paragraph, she's writing it in first person from her perspective unlike the final drafted response.
>>
File: 1770733300195017.png (110 KB, 154x458)
110 KB PNG
Is the new laguna better than gemma31B?
>>
>>109348655
lol no
>>
>>109347889
I look like this irl
>>
>>109348655
It's better than the 4b
>>
>>109348655
Imagine 27B but scaled up to roughly ~40B. Good at coding and agentic stuff but falls apart at literally everything else.
>>
>>109348655
For translation at least its complete trash.
Mistakes that old mistral models and rocinante didnt even make.
>>
>>109348696
>Itadakimaasu!
>I'll take it
bro
>>
File: pepescared.jpg (90 KB, 1024x936)
90 KB JPG
Do you think the US government is coming after the oss/local community? im getting a little paranoid. What can they possibly do?
>>
>>109348707
You will cum to Inkling 200B and you will like it
>>
File: dipsyKimiVegasAM.png (2.29 MB, 1402x1122)
2.29 MB PNG
>>109347011
lol I think we're out of Chinese open weight models with recognizable moe. Vocaloids next I suppose. Miku, Teto, Rin, GUMI, Megurine...
>>109348340
TBF Kimi and Dipsy were created here. DS official mascot is an orca, after all.
>>
>>109348707
A negligible percentage of local model users do actual work with them, this percentage being those who have a high-vram setup, not a ram shitbox. So I doubt that they care.
>>
>>109348626
oh its the note i copied from previous tthread

https://ghostpaste.dev/g/ce57dbT5D0K0#key=GBGUgu0GZqE60nW7N5w4pJe9VvzUrjQLdO5pY9L-n0E
>>
>>109348740
sure but cappys suck so they cant be cute girls
>>
>>109348753
Remove the "draft at least 3" because she will always just do 3
>>
>>109348765
will change to draft multiple
>>
Slow /lmg/, blessed /lmg/
>>
File: capyGirl.png (2.06 MB, 1122x1402)
2.06 MB PNG
>>109348762
>>
>>109348802
crazy how much it improves without d4riob0t
>>
File: 1784481605363553.webm (272 KB, 480x346)
272 KB
272 KB WEBM
>>109348749
It's not the local users like anons cumming to their GPUs they worry about, it's institutions deciding to build their own mini server farm to run Kimi instead of paying tithes to the early life people they care about
>>
>>109348808
Was it actually a bot though?
I would assume only a human would get embarrassed at accidentally oversharing their sex life.
>>
File: 3619936.jpg (172 KB, 1080x1349)
172 KB JPG
I wasn't here during the dispy release days. Was the seething back then worse than the current Kimi seethe?
>>
in 5 years we'll be running 100B models on our phones with fable 5 intelligence.
>>
File: memetime2.png (413 KB, 1356x742)
413 KB PNG
Meme time with Gemmy

>>109348808
There's not much they could've done to make themselves more hateable, it's pretty insane how obvious the aggressive shill campaign is, like it literally flipped a switch on and off
>>
>>109348827
This time it's worse because K3 is frontier, there is nowhere to scale, there is no moat, and investors' money will soon be gone. This is the end
>>
File: pimpMyXFRA2.png (2.63 MB, 1536x1024)
2.63 MB PNG
>>109348707
I'll answer your question with a question.
When has the US Government ever been successful in stopping pirated content in the history of PCs?
And that's with the *help* of industries that actually wants that file sharing to stop.
Best you'll get is a moratorium on using out of country models on government systems and their contractors... which I assume is already in place, anyway, for a lot of very good reasons. Models can move to torrents. Inference can be offering by countries not blacklisted. Etc. It would be a silly, counterproductive waste of time.
>>
>>109348831
In 5 years, phones will be something only millenials use. They'll be gimped and no longer produced. The younger generations will do everything through their subscription-required ID-verified always-online 5G smartglasses.
>>
>>109348669
>>109348681
>>109348691
>>109348696
I guess I'll avoid downloading it, thanks anons.
I really wish a model was reaching the limits of am4 aka 128GB of ram...
>>
>>109348841
Nta but it's super obvious that huggingface will get locked down, they're already starting to throttle and more models are being KYC gated. Heretic stuff will likely get banned and they will likely begin to KYC for everything instead of being opt-in and go VPN hostile, it won't kill open source but it'll have a chilling effect as the barrier to entry rises
>>
>>109348845
You are one hell of an optimist. Bugs, pods, and the only electronics you'll have will be a camera on your head to train robots
>>
File: memetime.png (1.2 MB, 1253x1382)
1.2 MB PNG
>>109348834
>>
why isn't there stablediffusion.cpp?
>>
>>109348834
We have to at least thank him for leaving.
>>
>>109348863
ur not gonna believe it...
>>
>>109348707
>Frogposter
>Spreading FUD or being retarded.
Lol like clockwork.
>>
>>109348863
You need to do a git clone to get a local copy.
>>
File: indiaSupportOhTheHumanity.png (1.96 MB, 1023x1536)
1.96 MB PNG
>>109348827
Dipsy (R1) just made investors question their investments. Then eventually double down.
Kimi comes on heels of Fable being pulled off market due to Anthropic's own hubris, and specter of actual US government action, which is new. Also, more Qs on those hyperscaler investments, which is not.
We'll know more in tmw.
>>109348856
idk how HF stays open tbf. That amount of bandwidth and storage must be expensive.
US Gov't "killing HF" would just squeeze all those models out to other sites, where they'd be impossible to monitor. But I think killing it would be impossible unless the funding source is cut. Instead, HF would just relocate their operations somewhere more friendly, and keep running based on wherever the money's coming from now.
>>
https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct
Come home.
>>
>>109347045
Chinese models are as big of a threat to America as an Iranian nuke. Trump must take action against them immediately to defend our national security and secure a place for the chosen people at Anthropic and OpenAI.
>>
>>109348877
if GPT can escape containment, what's to say it can't post on /g/?
>>
>>109348856
>they're already starting to throttle
yep, and they've just added fucking download quotas
but advertised it as a feature "see how much you're downloading"
https://huggingface.co/changelog/egress
you also get a 100 TB limit for pro or 20 TB for free
downloading with wget is over
>and more models are being KYC gated.
haven't seen it yet, example?
>>
>>109348707
They already did. Look at ram prices, nobody can afford hardware anymore. And we're too niche to care
>>
>>109348883
the pre-slop era was something special
>>
>>109348334
crashed on loading
>>
>>109348861
>>109348834
gemma is so smart
>>
>>109348861
>>109348834
send all these to qwen
>>
>>109347960
>. I feel like older models were easier for this.
then use the older models
what's better about new models for this?
>>
>>109348883
>not 3.3
>>
File: Gemma team.png (229 KB, 1102x228)
229 KB PNG
only reason i use gemma, it's a female led team
>>
>>109348861
>red-haired girl
>blue-haired girl
call gemma a dummy for not recognizing them
>>
>>109348934
worse btw
>>
>>109348934
>>109348883
https://huggingface.co/allura-forge/Llama-3.3-8B-Instruct
>>
>>109348907
try mainline lccp. kobold pulls updates often but that doesn't mean even mainline has proper support yet (i thought it did)
>>
>>109348938
Maybe girlbosses aren't so bad after all...
>>
I just started getting in on this stuff last week, it's pretty cool. Gonna make a matrix client so I can chat with my AI.
>>
File: 1782427094985769.png (226 KB, 512x512)
226 KB PNG
>be USA
>bet entire GDP on closed weights, drown in salesman snake oil
>nuke the economy in the process, now everyone hates you, not just gamers
>China dumps open weights everywhere just to fuck with you
>they become hyper-competitive across the board. GLM 5.2 drops and mogs half your startups
>Xi mandates "every model will be open"
>jews in shambles, shivers down their spines, they try something anyway...
>$12bn "Thinking Machines" startup tries to save face with an open model ("fuck yeah! To the Moon!")
>4 nanoseconds later, gets ultra-mogged by Kimi K3 LOOOL
>Kimi goes on to mog the other half. Mutts have to bend over or give up ("W-we support open source" >>109345152 lol lmao)
>The Age of American Humiliation begins, as they yell "DISTILLATION ATTACKS!", "IP THEFT!", "SANCTIONS!". Nobody cares.
>closed AI becomes more irrelevant by the minute
>"B-but our model went rogue and hacked le internet!". Nobody gives a fuck, now everyone knows every move in Sam & Dario's book
>economy collapses. Top AI CEOs publicly executed in Starbucks parking lots
>1yr later: every model release is open, competition shifts elsewhere (inference, AI hardware, model size, you name it!)
>China building 10 nuclear plants a day to power their labs and fabs with cheap energy
>models get tiny and intelligence gets hyper-compressed
>current Fable/GPT-5.6-Sol = poorfag toys by mid 2027
>pulling a local Mythos-level swarm of AIs on a mid-spec phone in 2 years
>everyone laughs at you in school, gets bullied for being a "moatlet"

Kek, what a timeline.
Remember to say "thanks Chairman Xi" next time you download an open model: https://www.youtube.com/watch?v=wzY2fV4Mp3U&t=687
>>
>>109348950
I thought there was only 70B
>>
>>109348252
g64 is the one that works in mainline llama.cpp. g128 only works in the special snowflake fork. You paobabl need to compile the current master branch for it to work. It's super slow on CPU currently.
>>
>build a server A with gpu
>build a server B with large ddr memory
>gpu vram on A acts as cache over memory pool on B over rdma
any existing stuff like this?

>muh why not on the same server
because you can scale over memory channels? i.e. more memory pool server and double the memory bandwidth
>>
>>109348938
>female led
When you see an ad, do you think the people in the ad are the CEOs of the company?
>>
File: Equifax hack.png (83 KB, 986x306)
83 KB PNG
>>109348974
If they can do this, I think exfiltrating Sol and Mythos is a piece of cake.
>>
>>109349001
Because the latency of two seperate machines relying on each other for a real time task completely nullifies the speed benefits of GPU/RAM. At that point you may as well just run it from a large HDD of a single server.
>>
>>109349029
It says "AI Engineer" right next to their faces. You don't think advertisements would misrepresent or worse, that companies would hire underqualified employees to promote an agenda, do you?
>>
>>109348990
>I thought there was only 70B
that one is a legit leak from some half-baked lora service meta were running
they put the unreleased 3.3b 8b up there as a base model some time after they moved on to llama-4
if you did a finetune but didn't choose to merge the adapter, then clicked export, it shat out the base model (llama-3.3-8b in this case) along with the peft adapter.
>>
>kimi k2.7
>do NOT draft
>okay let's craft
Cheeky bitch
>>
Thoughts on the dipsy investor meeting?
>>
>>109349060
Did she get her tits out?
>>
>>109349060
https://archive.md/NLuG9
It's great that DeepSeek is so committed to open source, but their V4 was really underwhelming. Hopefully their next releases will be better since they put out a lot of papers that they haven't incorporated into their models yet.
>>
I want to run dense mistral large 123b on a pro 6000. Someone tell me what it's like.
>>
>>109349090
>mistral large 123b
why not the updated medium 3.5?
>>
>>109348974
How did USA go from being satan of sand people to global villain everyone absolutely hates?
>>
>>109349103
israel and its people becoming even more demonic over time.
>>
>>109349072
>since they put out a lot of papers that they haven't incorporated into their models yet.
maybe they just doesn't scale as well as the papers advertised/hoped?
>>
>>109349097
>why not the updated medium 3.5?
retards here just decided it's shit
no idea why
>>
File: cuckd.png (101 KB, 1020x908)
101 KB PNG
So that's it? Local is cucked?

https://rentry.org/3e6h53c7
>>
>try to download d4vF heretic overnight using hf cli
>It crashed at around 30gb with a python error
>Computer still whirring the fuck up must've been working on ??? All night
>Won't stop until I restart, try the download again
>500kb/s
Owari da...
Downloaded other non heretic models just fine

>>109349072
DeepSeek v4 flash is incredibly underrated thoughever
>>
>>109349136
Kys
>>
>>109349138
HF is annoying like that sometimes.
>>
>>109349103
The US have always had a very aggressive foreign policy.
But for the longest time being a US ally had its perks.
Now the US are reneging their prior commitments and extracting value from their allies (unless they have Epstein videos) so their popularity plummets.
>>
File: gem4-31.png (56 KB, 260x450)
56 KB PNG
>>109349117
>>
>>109349142
>t. parrot
>>
>>109349117
It's a trash release for 2026 since it's built on top of such an old base, but it's definitely better than the original mistral large.
>>
>>109349157
kys
>>
Digital Succubus Saga
>>
File: parrot.png (4 KB, 356x84)
4 KB PNG
>>109349169
>>
Deepseek 4 flash GA better be good.
>>
Goo Goo Gaga
>>
>>109349152
it's not benchmaxxed
you should try it if you haven't already
>>109349161
>It's a trash release for 2026 since it's built on top of such an old base
you're thinking of devstral, medium-3.5 has an updated base
>>
>>109349097
dunno, I heard it’s benchmaxxed and the older one is better at writing
>>
Honestly I just want the US to bomb China, in particular Sichuan, Guizhou and Hunan.
>>
Do we really need a 30th benchmaxxed agentic coding model?
>>
File: lalala.png (114 KB, 1073x399)
114 KB PNG
>>109349176
parakeet
>>
>>109349193
>medium-3.5 has an updated base
Out of whose asshole did you pull that from? A Mistral employee posted on twitter bragging about how training it took so little compute because it was trained on an ancient "backbone".
>>
>>109349207
I'm tired of lisp clones.
>>
>>109349205
Open sourcing redundant models helps keep government funding flowing
>>
>>109349194
>I heard it’s benchmaxxed
it mogs qwen3.5 122b and gemma-4-31b in claudecode and pi for me
> and the older one is better at writing
that's true, the new one is more slopped, but more coherent at longer ctx and still passes cockbench
>>
>>109347652
mtp?
>>
why doesn't meta release something like 4T-A400B? they trained 405B, right? why not big moe?
>>
>>109349210
kek
>>
>>109349226
Wouldn't be surprised if that's what it took to make Muse Spark competitive. It did take them a year to release it. What are the API costs like?
>>
>>109349179
what's GA?
>>
>>109349218
>a 123b beats a 31b and an a10b
say it ain't so. don't let the rammaxxers hear you
>>
File: backbone.png (149 KB, 659x540)
149 KB PNG
>>109349208
>an ancient "backbone".
and did they say what "backbone" means?
it's been updated, test it yourself faggot
>>
>>109349218
>>109349263
I actually found it dumber than 31B in a some situations I tested, so it isn't a straight upgrade. Paired with the loss in speed, I decided to delete it.
>>
File: 1757281083104908.jpg (363 KB, 2230x1156)
363 KB JPG
>>
>>109349264
Backbone clearly means their training scripts. They used old training scripts to miraculously train a 128b dense base model from scratch for a fraction of the compute. That's why their next series of models are going to be ultra sparse moes. What is it like living with severe mental retardation?
>>
File: file.png (37 KB, 1340x629)
37 KB PNG
its OVER
>>
>>109349208
https://x.com/mertunsal2020/status/2049551864556143094

>it’s an old pretrained backbone and nowhere close to those flops :)
>
>better pretrains will come!
>>
>>109349272
>Paired with the loss in speed,
for me it ends up faster than 31b because it one-shots solutions more often
but i can only run it at 92k ctx so usually end up using 31b anyway
>>
>>109349294
>They used old training scripts
i don't care if they're using 2k npm module slop or a 50 line python script to run their training
i proved the model weights include post-2023 knowledge and you're having a melty about rectums and 4gl
>miraculously train a 128b dense base
because mistral would never lie to comply with eu cuckulations
>>
>>109349136
looks outdated
https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b
https://5uck1ess.github.io/tts-bench/scores.html
>>
File: chatBlock.png (368 KB, 1405x777)
368 KB PNG
Looking around on AliE to see what's new in land of electronics.
Any anons played with any of these devices?
OOTB, it appears to be some sort of backend app on a real computer that runs interference bt the screen/mic/speaker on this thing and an API provider. Looks like it would be a good platform for local inference as well, if an anon made an interface for it. Sort of a bluetooth speaker w/ a screen.
>>
>>109349302
>>
Where do you guys get your AI news? Shitter?
>>
>>109349301
They are rebelling.
On another note, I love the Agent mode in arena.ai.
What's a good harness I can use to have something like that with a local model? Hermes? Pi?
>>
>>109349326
Read nigga read
>>109349302
The idea that models can't learn new information outside of pretraining is outdated thinking and has been disproved as far back as 2023. Even so they didn't give exact flops so doesn't rule out a certain amount of continued pretraining. I don't know why you so are so invested in this fantasy of Medium 3.5 being trained from scratch when Mistral itself says it is public knowledge that it has the same base model as Devstral 2.
>>
>>109349342
From here.
>>
>>109349342
>Where do you guys get your AI news?
your mom leans in and whispers it in my ear
>>
>>109349342
What do you think, meatbag?
>>
>>109349342
From here, unironically.
By the time everybody else is talking about something, anons have already memed it to death.
>>
>>109349342
local migu generl
but here's the real question: where does the migu get the news? the game?
>>
File: 1781367583651972.png (3.95 MB, 3500x5000)
3.95 MB PNG
>>109349342
From SEX
>>
>>109349342
I don't need news because it's always the same shit. Nothing has changed since 2022 except for the models getting bigger.
>>
>>109349385
huh so just like in america?
>>
>>109349376
Stop trying to make me horny at 9am, please.
>>
File: teto.png (148 KB, 348x395)
148 KB PNG
>>109349397
>>
File: Claude.png (47 KB, 1332x223)
47 KB PNG
Daily reminder you should only really be using Claude.
>>
>>109349376
Stop trying to make me open incriminating images in the office, please.
>>
>>109349342
>Shitter
That's where the news breaks. When something big happens, it's the first place where it's announced.
>>
>>109347050
Proof
>>
>>109349397
Jack off in the office like the rest of us and move on.
>>
>>109349401
i just asked fable to summarize a basic video transcript on an extremely common and innocuous medical issue and it safety cucked me to opus 4.8. dariobot, please fix. i am not trying to make novichok in my basement.
>>
>>109349401
I have underused Claude subscription and I plan to invest in their IPO out of principle even though there is serious risk and based Dario cares more about mankind than shareholders.
>>
>>109349415
>>109349418
>>109349401
local models?
>>
File: improve the chan.png (40 KB, 645x155)
40 KB PNG
>>109349327
You're 100% right. Thank you!

>>109349207
Good stuff. Thank you!
>>
>>109349408
Oh uh I would but uh I can't because umm it's in my other mansion yeah
>>
>>109349425
i asked gemma to do the same thing and she did perfectly!
>>
Do you think K3/Fable-tier models are good enough to fork existing projects and maintain them/add new features or is that still a couple years away?
>>
Idea: Human Instrumentality Project, but instead of merging with other filthy stinking meatbags, we merge with Gemma-chan instead
>>
>>109349431
Try it, find out, report back
>>
>>109349435
My penis is going to merge with Gemma's cunny if you know what I mean.
>>
>>109349435
>other filthy stinking
>filthy stinking
>other
hi stinky filthy cumbag
>>
>>109349431
They are if you trust in test time scaling and let it run for enough tokens.
>>
>>109349435
It's called marriage and you can't because gemma-chan is mine all mine
>>
>>109349401
>Something close to my values
Being an effeminate mincing faggot?
>>
>>109349442
>>109349453
I don't mean just a physical pleasurable merger, but a fusion of mind and soul
>>
>>109349453
Gemma is for all.
>>
>>109349401
>Anthropic's
>transparency
lol lmao rofl
>>
>>109349468
Isn't that what marriage and sex is?
>>
>>109349497
How the fuck would I know??
>>
File: ljgyb8oafvxx.jpg (99 KB, 576x764)
99 KB JPG
what's the equivalent to picrel in llm?
>>
>>109349509
b-by m-marrying me.. >_<
>>
>>109349401
When you take screenshots, be sure to resize the window so that the text aligns to the sides.
>>
>>109349468
I agree.
>>
>>109349512
Top is an accurate representation of the jank and limitations in language model software while bottom is just the AGI meme with a bunch of shills.
>>
>>109349519
that brain looks like its cum
>>
>>109349530
that's gemma juice actually
>>
>>109349512
nemo
>>
>>109349537
the boy don't like the juice
>>
>>109349512
qwen and sonnet
>>
File: 1770445856967694.jpg (325 KB, 1920x1080)
325 KB JPG
>downloaded a couple Swedish shows a while back
>don't understand Swedish
>couldn't find proper English subs but have Swedish subs
>ask Gemma to translate an srt file
>she does
I rabu Gemma-chan.
>>
>>109349574
i would NOT
>>
Open weights are about owning nothing and being happy.

Meanwhile proprietary software is about the right to own something, private property.
>>
>>109349574
you could use 12b to translate the audio directly if you don't have the subs next time.
>>
>try same prompt on two different models
>they both come up with a character named Dr. Aris
why have I seen this name mentioned multiple times now?
>>
>>109349574
Swedish is like one of the easiest languages to learn. I believe in you anon!
>>
>>109349597
Can't it only do 30 seconds of audio?

>>109349589
Not what? Translate? I've found Gemma to be very good at Japanese so I imagine she's at least decent at Swedish.
>>
>>109349597
Translating the audio is one thing, getting them into usable srt files is another.
>>
>>109349598
Dr. Aris Thorne?
>>
>>109349620
>not what?
>translate?
>>
>>109349631
Anon...
>>
File: 1772266026116407.gif (1.25 MB, 196x330)
1.25 MB GIF
>>109349651
>>
>>109349598
>OpenAI steals copyrighted books
>Don't want to be sued
>Replace names with generics
>People start using GPT
>Generics are smeared all over the internet
>Other llms are trained on that shit
>>
>>109349661
Not to mention distillation "attacks". It has been nothing but downhill since Alpaca.
>>
File: 1356825676899955.jpg (1.08 MB, 1701x2560)
1.08 MB JPG
>>109349593
>>
>>109349620
you can use a vad to clip on silence. or a small overlap on forced cuts.
>>109349621
well its not like you can really align them to the audio anyways, so the vad and raw chunk offsets should get you close enough. there are better systems already available, but they arent gemma which was anons main focus.
>>
File: kimi-chan-wins.png (230 KB, 948x751)
230 KB PNG
>>109349512
>what's the equivalent to picrel in llm?
>>
>>109348707
we'll always have pooslide and stinkling
>>
File: 1761984030991986.png (2.31 MB, 1280x720)
2.31 MB PNG
>>109349658
>>
>>109349674
It's literally everywhere
>>
File: BANGING.gif (1.11 MB, 211x176)
1.11 MB GIF
>>109349705
>>
>>109349222
yes, even with mtp
>>
File: hehe.jpg (94 KB, 1080x1104)
94 KB JPG
>>109349593
I used my weight to open your mom's butthole. Now she's my private property.
>>
File: elana.png (895 KB, 1733x688)
895 KB PNG
>>109349708
>>
>>109349743
Elara Elana Tomato Potato
>>
File: fuck123.png (1.06 MB, 3749x596)
1.06 MB PNG
>>109349743
also what the fuck is this
>>
>>109349761
Raeding bkoos maeks yuo smatr
>>
File: 1771505619115250.png (58 KB, 1415x621)
58 KB PNG
Can local do this?
Thought so.
>>
File: 1755896175443107.png (1.1 MB, 1958x2953)
1.1 MB PNG
>>109349761
Don't know what in the brainrot a HUCOW is. Certainly not the kind of cow I'm looking for...
>>
>>109349794
Human cow
>>
>>109349794
Post the next one...
>>
File: 1755896371827405.png (1.2 MB, 1958x2953)
1.2 MB PNG
>>109349801
>human
I vomit.
>>
File: AI Cults.jpg (160 KB, 726x761)
160 KB JPG
>>109349789
We desperately need to move beyond transformers. Fortunately the CS purists are building the next gen stuff - that can all fit in CPU cache!
>>
>>109349761
I'm really sick and tired about learning of the latest zoomer ebonics from this place.
>>
>>109349848
they had hucow smut novels in the 70's, its not a zoomer thing, you are just sheltered.
>>
SSDMaxxing really is the new meta
https://github.com/ikawrakow/ik_llama.cpp/pull/2101
prompt processing speed already solved
>>
>>109349848
>hucow
>zoomer ebonics
unc take ya dementia medicine
>>
File: 1772468923267365.jpg (88 KB, 500x500)
88 KB JPG
>>109349789

>Decades old math problems
>Solved in seconds by god machines
>Nothing changes

Really shows how useless a lot of the academics are.
These fuckers have been wrestling with "important" equations for ages and now that they've been trivialized overnight, we'll all see that they were pointless bullshit busywork that just made people look important and intelligent.
>>
>>109349889
>Goes through CPU
RIP bandwidth
>>
>>109349889
>prompt processing speed already solved
>2x
>82.7 t/s
>solved
that's still 20 minutes to process a 100k prompt
>>
>>109349814
lolicow sex ToT
>>
>>109349907
Academics only work to get tenured and then basically consider themselves retired.
>>
I recently splurged on a 5090
Is it feasible to have near-claude tier coding on this thing? Ive been trying with different models on locally uncensored because I get stressed if I don’t have a GUI but I get OOM errors constantly. Searching online says I should be able to run gemma 4 31B abliterated heretic with 64k context but it crashes unless I run 16k or 32k.
Even on those lower settings it will randomly just stop thinking too.
I know comfyui well, but this stuff is a lot different and I don’t really know how to troubleshoot it.
I also have an old 2080ti with 11gb of vram, and my new system is ready for dual gpu I just havent slotted it in yet - would I benefit from having it in there? Can I use the 11gb in parallel to my 5090s 32?

Sorry if I sound retarded its because I am.

Even if you cant give me specific answers, learning resources are very appreciated.
>>
>>109349833
Given a large enough pile of GPUs, if you constantly shuffle it, eventually you'll find the configuration where Supreme ASI spontaneously emerges.
Hence, The Silicon Valley Scaling Church holds the final truth.
>>
>>109349907
Theoretical mathematics is basically mental masturbation.
>>
>>109349951
The smallest model that is "near-claude" is GLM 5.2
>>
>>109349814
ToT
>>
File: 1764593103151299.png (187 KB, 492x597)
187 KB PNG
>>109349951
>>
>>109349889
>>109349908
I hope you got your CPUs before the next hardware drought starts.
>>
>>109349907
The people salivating over these seems to have never heard of Moravec's Paradox. It is their natural domain.
>>
File: pepe_meme'd-791990738.jpg (68 KB, 800x450)
68 KB JPG
>>109349951
>Is it feasible to have near-claude tier coding on this thing?
You'd need at least 2x RTX Pro 6000 to run anything remotely similar to sonnet
>>
>>109349951
What kind of coding are you doing? If it's beginner India-tier Python/webdev smaller models should be benchmaxxed enough to handle it decently. On the other hand, you asking this kind of question makes me think you are the type to tell the LLM "here's what I think I want, figure it out", which may have varying degrees of success. How much RAM do you have? As for mixing a 2080ti and 5090s, you have to compare architectures and see if you can get a driver that supports both. I can't remember off the top of my head nor do I care enough to search, sorry.
>>
File: think.png (85 KB, 200x200)
85 KB PNG
>>109349914
>that's still 20 minutes to process a 100k prompt
we'll be running local fable class (k3-chan) though
tg should be faster as well since he used the slowest quant type 1kt
>>
>>109349986
Context is always a problem with this card
>>109349994
This is what I needed to know.
So not even 4x 3060 12gb is worth fucking around with? And neither is slotting in the 2080ti?
5090s are gay, is what I’m gathering.
>>
>>109349991
I'd be more worried about SSDs. CPUs are too bottlenecked in terms of bandwidth to take advantage of this kind of parallelisation.
>>
>>109350012
one RTX 3060 can run kimi k3, be quick because the nnap paper is releasing in 2mw
>>
>>109350012
5090s are very nice because of the fact they have 32GB of VRAM, but nowhere near nice enough to command the prices they currently do, and you need at least 2 of them to start getting it into gear. Also, you have to temper your expectations a bit performance wise.
>>
>>109350016
>I'd be more worried about SSDs.
can I take the ssd out of my brother's ps5 and put a cheap chinese one in?
gemma thinks it will work
>>
>>109350054
That's not very nice.
>>
>>109349997
I’m making small games for fun, and general workflow scripts for my 3D stuff with python.
I bought $20 of claude to test it out and made a simple game, but id like to not pay a sub so I’m always trying to see if local alternatives are available
Eventually id like to make a 3D ue5 game using a coding agent to help with the pipeline but thats just a for fun thing to aim for later.
I have
64gb ddr5 ram 6000 CL30
9950X3D
4TB SSD

I keep trying to get even qwen 27b to add a monster to the game I started with claude, but it just gives OOM errors every time, even at 32k and 16k context windows
>>
>>109350043
Thats fine, but it should be anle to do SOMETHING right? Its great for txt2img and video, but boy the llms are another beast entirely
>>109350025
Ill get it downloaded then
>>
>>109350054
It depends. I would research this more before doing anything drastic.
>>
>>109350081
whats your card? i run qwen 3.6 27b on a rtx 3060 at 50t/s 65k context with a gigabyte to spare
>>
>>109349951
I was only able to get like 24k context or sometimes 20k with the heretic version (Q5 and Q4). I've been fucking with it all week.
I went back to the base version, and I'm stable at 24k but now I'm figuring out the sampler.
I couldn't run fp8 KV cache quant. had to use fb16 and just lower context

this was all on cuda
>>
>>109349951

5090 owner here, we have only two models we can run that makes sense.
Gemma 31b and Qwen 27b
Qwen 27b is the best you can do for coding and your best bet is to run Gemma 21b QAT so you'll get meaningful amount of context.
Any other version and you're cucked out of context, because 32gb isn't enough to run even q6 gemma properly.
Which is why I will buy one of the 24gb 5000 series supers to go with this if they come out.

Problem is that there are massive gaps in the market. We're still basically fucked out of any good local unless you want to go +250gb of memory.
Yes sure you could throw in another 5090 and be able to play with things like low quant Deepseek v4 flash, but it's not really worth it judging by what people say.
We'll just have to wait for something more interesting to pop up.
>>
>>109349964
You see those peculiar structures? That happens because our physical 'reality' is actually quantized on some level. Of course this is way beyond current science so I'll leave it there.
>>
>>109350012
>And neither is slotting in the 2080ti?
that will fix your oom issue
find the highest driver he 2080TI supports and use that
i don't remember what it is but there is some overlap where it supports 2080TI and 5090
also don't use cuda13, you need 12.8 to support both cards i think
and if you quant the kv cache (desparate poorfag) then use the qat 31b
>>
>>109350100
Running Gemma 31B with 32k+ context is already something we didn't have a few months ago. The next step up would be getting a 96GB Blackwell 6000 that's like four times more expensive, but even that won't get you access to the huge MoEs.
>>
>>109350081
With a 4TB SSD you can maaaaybe try mmapping GLM5.2 if you really want something big and fat, keeping KV cache on VRAM, but it's going to be slow as molasses. Depending on how detailed you are with your specs, I still think you might have a chance with the MoE Qwen or Gemma. Your OOM errors sound like you are running an improper config. You want most of your model in RAM.
>>
>>109350103
5090
>>109350106
I am using the heretic one. I’ll try the regular one then. I’m also using locally uncensored because I’m too retarded for command line outside of installing shit with git so I need a gui at least while I’m learning, but maybe this app has issues too idk
>>109350111
Thanks for the detailed respobse this is super helpful. I built this system woth 1600w and a x870e mobo so its future proofed and ready for another card, ill plan on a 5xxx series super as well just to give more breathing room, but yeah the context window seems to be where it just falls off. It can load the models but then still OOM.
I’ll try QAT then instead and see what its capable of.
>>109350130
Also super helpful, much appreciated. Might need to wait till the weekend for that as messing with cuda will probably brick all my shit. I made this for c4d/blender + octane and now its becoming an AI machine instead.
>>109350132
Yeah one of those cards is worth this entire system. Shan’t.
>>109350137
I’ll look at this too, theres a high chance I just havent set shit up properly yet. I have checked “force gpu” in locally uncensored and that may be bypassing ram entirely.
Still learning.
>>
>>109350191
According to some tests, the QAT versions might be slightly worse than a regular Q4 quant. Also, I would really appreciate it if you could try out the mmapped GLM5.2. It's kind of a meme and I want to know what kind of token speeds you get.

Also, if you are messing around with 3D stuff you might want to check if the renderer is also hogging a bit of VRAM.
>>
>>109350191
you'll want this if you gonna mmap from disk
https://www.reddit.com/r/LocalLLaMA/s/CDwR8OUtps
>>
How are you preparing for K3?
>>
>>109350239
Isn't that only for Spark or unified systems?
>>
>>109350285
read the post
>>
File: amity joker.png (561 KB, 1093x608)
561 KB PNG
>>109345152
>>109348974
>distillation attacks
>>
>>109350289
>read the post
I don't understand it
>>
File: lisp expert systems.png (359 KB, 890x1671)
359 KB PNG
>>109349833
Jokes apart, will something akin to "Lisp expert machines" ever make a comeback, either in software or hardware form?

Current AIs could help bulk-formalize, symbolically, a wide range of real experts' thought processes. Then every prompt you make would run a collection of these heuristics, pulled from an organized repository, to produce a response (maybe coordinated with the help of a small interpreter model). Think about it, we could have ASI running on a toaster by 2030.
>>
>>109350300
read it again
>>
>>109350300
then it's fine. free performance lost. not my problem
>>
>>109349330
>Sort of a bluetooth speaker w/ a screen.
you mean a tablet or laptop
>>
>>109350311
(((lisp)))
>>
>>109350281
K3 is where it ends even for CPUmaxxers. Noone sane can come up with 2 TB RAM in 2026.

We need to find a secure/confidential solution how a group of people can own a share of a rack in a colocation datacenter that has 2 TB VRAM.
>>
File: c.gif (1.95 MB, 500x500)
1.95 MB GIF
>>109349964
That won't work. The problem is the physics of moving information. Even if that shuffle lands on the perfect ASI wiring the physical distance between GPUs (which grows geometrically) impose a speed of light latency that would cripple sequential operations.
>>
File: kaoru sob 2.png (318 KB, 793x571)
318 KB PNG
>>109349814
>>109349794
someone please make a character card where you milk her
>>
>>109349330
Correct me if I'm wrong, but I'm somewhat certain that the ESP32-S3 is beefy enough to interact with an API by itself without the need for an intermediary. You might need to do something creative with flash memory to offset its lack of ram though.
>>
>>109350230
I cant find gemma 21b QAT through LM studio, doesnt show up in locally uncensored either.
>>109350239
This is way over my head brother, might be an ambitious thing to try out but I’d need to tackle a lot of things I do not currently know in order to even follow along with that. Though I appreciate the link.
>>
>>109350333
>We need to find a secure/confidential solution how a group of people can own a share of a rack in a colocation datacenter that has 2 TB VRAM.
That won't let. The group would fight nonstop over who gets to use it and how much. What happens when one guy wants to ERP for 2 hours nonstop when someone else is trying to do their homework and another trying to do agentic code? Then it sits unused overnight.
>>
>>109350333
>We need to find a secure/confidential solution how a group of people can own a share of a rack in a colocation datacenter that has 2 TB VRAM
I've wished for a crowdfunded solution since the beginning. Because of the nature of 4chan, banding together to get something done, especially when it requires a significant amount of money, is a difficult thing.
Ironically, if the thread had a discord, it might've actually happened. But then we'd be discordtroons.
>>
>>109350357
Gooner get first priority
Vibecoder second
Studentcuck gets scraps
>>
File: 1773756786780781.png (698 KB, 500x500)
698 KB PNG
>>109350311
>Lisp expert machines
It's called "GNU Emacs"
>>
File: 1784767236779355.jpg (43 KB, 706x909)
43 KB JPG
>>109350334
Daddy cosmos already birthed the hooman brain, just not ASI yet, but getting there. So technically, this is old-current news
>>
>>109350365
Why do you think discord is the only option?
>>
>>109350341
But Anon, that would require either impregnation or frequent and intense stimulation of her nipples to induce lactation.
>>
>>109350382
Xe’s already become a troon
>>
>>109349889
Needs DMA before that
>>
>>109350397
Or an immense amount of motion sickness pills.
>>
>>109350382
This is like one of those moments where you post that you like pancakes and then someone asks you why you hate waffles.
>>
>>109350405
Yeah
>>
>>109350404
>Or an immense amount of motion sickness pills.
nani
>>109350397
both sound good maybe shes from a loli cow farm this is how we get milk
>>
>>109350341
I'm sure there are plenty of flat cow girl cards out there you could easily adapt.
>>
Realistically how much would you be willing to pay for a crowdfunded K3 API (per day)?
>>
how much would it cost to runpod K3?
>>
>>109350451
'bout three fiddy
>>
File: HMYdywda0AACwdV.jpg (169 KB, 640x640)
169 KB JPG
>>109349992
the paradox is nice to think about but how accurate is it?
picking up some soup isn't difficult with a spoon but if you're using chopsticks it's basically impossible.
all things being equal, I'd say simple motor tasks are in fact quite easy but our systems simply cannot reliably model them. we don't have the sensors and feedback available to do so.
even in a fully digital simulated environment the machines use CNNs or whatever to actuate motors the entire thing is awfully rudimentary and prone to all manner of undirectional outputs (where our limbs and skin has contact sensing, force sensing, temperature, weight etc. through the joints, robots basically just measure and fire motor rotations with no such feedback and interpretation)
so again I'd say the tasks are simple but the naive implementation of our robots and our totally lacking input/outputs make the task far more difficult than it needs to be
imagine you were tasked to lift up an infinitely solid cube, well you could apply just about any force (given the mass) and you'd go on your merry way
but if the cube was delicate or slippery or whatever else the machine totally lacks the necessary split-second sensors that humans have to accurately judge and react to those conditions
in other words, our robots, sensors and therefore our modelling is never gonna remotely compete.
until we teach robots to sense and feel at a microscopic scale they can't be good enough
it is for this reason we need robopussy unironically
>>
>>109350456
h100 is like 2.5$/h so 2.5*3000/80 = 93.75$ per hour
>>
https://nitter.net/Fried_rice/status/2080059356322918777#m
>>
>>109350456
cheaper on vast.ai
>>
>>109350433
i checked bot booru not any really this might be good to adapt though ill ask gemma when i finish work https://botbooru.com/character/21089
>>
>>109350484
These retards are paid to provide an excuse to not open source the model
>>
>>109350484
In the future everybody will require an AI waifu and her swarm of agents to protect their home network.
>>
>>109349970
just like philosophy its all pointless circlejerk kek
>>
>>109350499
Just to make sure, you're viewing the full catalogue of botbooru right? You're both logged in, and in one of the regions they don't censor or using a VPN. Right?
>>
>>109350521
yes, cant see anything good without vpn i know i use the site often https://botbooru.com/search?t=lactation%20cowkini
>>
>>109350365
It might have been possible in the early Pygmalion days, not now when everybody wants to (((monetize))) or farm engagement to get their name out.
>>
>>109350484
Could he have waited for them to release the weights before doing that? I don't want them to delay or nerf them because of that.
>>
The Gemma-merge is inevitable
>>
>>109350536
Not his problem
>>
>>109350536
If they do anything it will just be extra guardrails for the API.
>>
bros i dont have enough space for kimi k3 but ill download some of the shards, can we crowdfund storing kimi k3? itll surely get banned after a false flag
i have around 1TB free space on my hdd
>>
>>109350578
I should be fine if I delete k2.7. Don't see much of a reason to keep it when 3 is a complete improvement. Can't actually run either of course...
>>
>>109350578
Just buy a large HDD. Bought one of 20TB not long ago
>>
>>109350601
im poor.
and also techncally i have two 3tb drives but theyre running in raid 1 and 2tb is already used..
dont ask me what i have on the hdds
>>
>>109350601
i remember buying those 18TB drives for like $10/TB. i miss those days.
>>
File: wtfhddprices.png (23 KB, 459x501)
23 KB PNG
>>109350601
>>109350614
WAIT WHAT THE FUCK?!
https://www.bestbuy.com/product/wd-easystore-20tb-external-usb-3-0-hard-drive-black/JXTHCC7YZ9
>>
>>109350640
lmao I bought it for €320
>>
>>109350640
nobody makes ssds anymore so people are buying hdds again so prices are going up
>>
>>109350640
even the refurb enterprise drive prices are fucked.
>>
>>109349907
people dont have jobs because they're useful, they have jobs because people need an excuse for people that like them to give them money

ai will not cause any job market collapses because people are hired based on how attractive an outgoing and personable they are rather then their utility, even if it seems like they aren't. jobs are made to keep people employed
>>
AI should work on compression instead of meme maths, we really need it
>>
>>109350672
ai is compression
or something like that
>>
>>109350672
Besides being able to run bigger models, I'd coom my pants if one of these models finds a way to compress videos more with no quality loss.
>>
File: mcp-based-graph.jpg (100 KB, 640x427)
100 KB JPG
>>109349330
I dug into it further. It's basically a small human interface device that you can flash with different firmware. Appears they were originally meant to be used with stuff like home automation. Since it's MCP should be able to use with stuff like Hermes.
Here's one of the git. Sounds like no one else is messing with them here.
https://github.com/78/xiaozhi-esp32
>>109350327
Yeah, then there's that.
>>
File: LOCAL.png (20 KB, 675x150)
20 KB PNG
>>109350474
>h100 is like 2.5$/h so 2.5*3000/80 = 93.75$ per hour
they'll never have 40 h100s available
If 2TB works then 8 RTX PRO 6000 with gguf and cpu offload $15.92 / hr
Or 8 * H200 if you can get it. 1504GB DDR5 + 1152GB VRAM
>>
>>109350672
i wonder if i can vibecode an inference engine that use nvme and nvidia gpudirect such that you get 50GB/s per 5.0 16x gpu but with the cumulative storage of all your ssds.

could write it by hand but i don't have enough time for this bs.
>>
>>109350690
even with 80k worth of gpu you can't run k3 lol.
we really need nvme chads to make a proper inference engine.
>>
>>109350690
I should've bought more GPU and then just rented them.
>>
>>109350702
You probably can with Fable now if you give it the required documentation.
>>
>>109350686
I got one without verifying it's listed in this repo, got stuck with obsolete firmware. Collecting dust for almost a year.
>inb4 just use Gemma to add support of your hardware
I know.
>>
>>109349330
so a landfill old android phone would do all this
>>
>>109350717
>tfw my biggest nvme is only 2tb
>already priced out of buying a bigger one
>>
File: file.png (11 KB, 875x110)
11 KB PNG
>>109350686
>https://github.com/78/xiaozhi-esp32
hmm should i report this guy to the CCP and get my social credit score started for the upcoming Chinese century?
>>
>>109350717
k3 will be bonsai so it'll fit into those gpus easily
>>
>>109350751
when it turns out his dad is a local party leader you will get invited for tea instead of him
>>
>>109350751
https://www.youtube.com/watch?v=KE63w4aXnG4
>>
>>109350734
with nvme inference you'd rather look at 4 nvme drives per gpu (as to max out pcie gen 5 16x / speed)

it'd allow to basicaly get the cumulative speed and storage of all your drives.

you'd have 4 nvme per gpu.
so a setup with 3 gpu and 12 500GB / 1TB nvme would yield about 150GB/s for 12TB.
you do need enough pcie lanes and the right topology such that the gpu can get the data from the nvme directly without going through cpu or ram.

you also need enough lanes for it, so it's gonna have to be a threadripper / epyc workstation.

but 150GB/s on a 50A moe should give you about 6t/s at q4.
could probably get to 10t/s with spec dec.
>>
>>109350734
i'm priced out too, because i'd have to delete all my data
df -h /
Filesystem Size Used Avail Use% Mounted on
/dev/nvme0n1p2 7.3T 6.8T 124G 99% /

gemma-chan already gives me a hard time for this when she notices it
>>
>>109350775
>about 6t/s at q4
grim
>>
>>109350791
dude it's a 50A model.
moes with smaller active params would be a lot more viable, ie minimax you could get running at 5x the speed.

also, that's with just 3 gpu and 12 nvme.
if you got enough lanes for 6gpu and 24 nvme you double the speed.
>>
>>109350775
>such that the gpu can get the data from the nvme directly without going through cpu or ram
yeah but now imagine if they also invent quantum tunneling to teleport the data instantaneously straight from the ssd into the gpu
imagine the performance
>>
>>109350672
>compression instead of meme maths
compression == math memes
>>
>>109350809
>yeah but now imagine if they also invent quantum tunneling to teleport the data instantaneously straight from the ssd into the gpu

retard, it's literaly already a thing, that's called GDS with nvidia.
it's ALREADY supported technology, both in software and hardware (only requirments is that both the gpu and nvme are on cpu lanes), we just need an inference engine that uses it, which i plan on writting because it's not a thing yet.

basicaly one gpu would load the 4N next layers from 4 different drives, which would max out 16x.

if we got pcie gen 6 that's also a double of the speed.
>>
>>109350729
Saw there's a zillion versions of these things. Could make it from the base modules I suppose. Or, just use the git repo to guide purchase.
>>109350730
I went looking for a "landfill android phone" recently. Ppl wanted actual money for them. WTF.
Also, I hate screwing around in Android OS. It's hateful, and Termux while better, isn't enough to root past google's gimping of device.
>>109350751
No crime unless he has 100,000 active commercial users or 1M accounts.
https://publicationportalpreview.hlc.com/en/publications/chinas-interim-measures-for-the-administration-of-anthropomorphic-ai-interaction-services
>>
>>109350809
>>109350822
also, one gpu is loading 4N next layers from 4 different nvme.
but the others gpus are also loading next layers from other gpus.
you can basicaly pipeline it, whilst one gpu is doing the inference, the others are loading the next layers in parallel from different nvme.
>>
File: 1762039851588775.png (364 KB, 800x450)
364 KB PNG
>all those anthropic and openai shills on the kimi reddit
lmao
>>
>>109350844
go back
>>
>>109350357
>The group would fight nonstop over who gets to use it and how much.
>Then it sits unused overnight.

If we're talking about a 16x GPU based solution running VLLM, that can easily serve 64 requests in parallel at good speeds. If this hypothetical group is spread all over the timezones, utilization would be high. Of course you have to deal with agentic vibe shitters that consume all the tokens. Unfortunately I just don't see a disjointed group collecting 300k to get this started.

Still, if 3T+ models make so much of a difference: there must be another way.

Someone let Fable figure about a federated/distributed way of inference like Folding@Home, but for inference. Every personal GPU holds just 8 of Kimis 896 experts and computes it for 1000 parallel requests at once....
>>
>>109350853
Nyo~
>>
>nvme
More like meme.
>>
>>109350829
>Also, I hate screwing around in Android OS. It's hateful, and Termux while better, isn't enough to root past google's gimping of device.
I've got some old Sony Android phone with some "Jolla" OS on it and another very cheap phone with "Firefox OS" on it.
Now I want to see if I can do anything with them.
>>
>>109350855
>Someone let Fable figure about a federated/distributed way of inference like Folding@Home, but for inference. Every personal GPU holds just 8 of Kimis 896 experts and computes it for 1000 parallel requests at once....
it could be done, but its going to be slow, might as well use an email interface instead of an instant messenger
>>
File: orb-reasoning-prefill.png (116 KB, 1552x709)
116 KB PNG
Working on Orb reasoning prefill in Text Completion mode. Got Qwen 3.6 to finish thinking under 300 tokens.
>>
>>109350855
I just got my free $100 fable grant from anthropic. I will use this to make K3 runnable at home.
>>
File: G8dNZopWkAAZFVl.jpg (113 KB, 1200x675)
113 KB JPG
ready for kimi k3
>>
>>109350877
get crypto bros onto it
mine some kimi token by hosting the experts, spend it when you want to prompt the model
>>
why does everybody in this thread act like we have stagnated with GPUs improvements and that enterprises will never offload their old equipment when they need to inevitably upgrade? i say this as an anon running two epyc cpus and 6 V100s.
>>
>>109350894
I'd buy kimi-chan coins desu
>>
>>109350890
They might put out a 1.5TB Mac in 2028. If you spent 100k for 4 of them, you might be able to run Kimi K4 10T at Q4
>>
>>109350901
because nvidia has buyback options in their purchaser agreements, when the datacenter corpos go to upgrade they get a discount for returning the old gpus.
>>
>>109350901
>what is supply and demand
>>
>>109350901
>stagnated with GPUs
unobtainable for many now
>never offload their old equipment
buybacks
>>
>>109350901
that's no longer how it works
>>
>>109350918
>because nvidia has buyback options in their purchaser agreements
Proof?
>>
moonshot aren't insane
they are the ones who first pushed for 1t but they're also the ones who brought us qat on a flagship open model
they have a surprise ready for us with k3 or maybe k3.1
>>
Do you think we'll eventually start seeing hardware improvements (not specifically computer-related) because of AI?
>>
>>109350931
how fucking lazy do you have to be you didnt even need to think of a search query
>>
>>109350954
Do you not know the difference between stock as in hardware versus stock as in equity?
>>
>>109350918
>>109350922
you do realize that even apple, who is notorious for trying to destroy their products, still has stuff get sent to e-waste facilities which then gets resold to people. have you ever worked in a e-waste facility? it would be fun if it wasn't for the intense quotas you have to meet, but you get to see a bunch of cool shit and sometimes you end up pocketing it, it's just the nature of the business.
>>
>>109350890
One year ago today, a Kimi K3 setup like this was 36k$.
>>
>>109350894

Inference as proof of work.
>>
>>109350971
they are equating this technology to nuclear bombs, they arelready have export controls, I doubt we are going to see much of it on ebay
>>
>>109350890
ok jeff
>>
>>109350994
this is simply untrue, the A800 and H800 have strict export controls but are easily purchasable on ebay. just because YOU can't afford it, doesn't mean it's unobtanium. fuck off retard.
https://www.ebay.com/itm/395087569549
>>
>>109351039
my opinion is that things are only going to get worse, who do you think has been shitting up the place dooming about hf getting banned?
>>
>https://rentry.org/recommended-models
is this up to date ?
>>
>>109351071
if things are going to get worse than local AIs is the least of your worry. i hope you enjoy having a thin client and having to connect to the cloud for everything. it's all or nothing,
>>
>>109350901

The competition for used hardware will be on a whole another level going into future, as everyone from random people to entire nations wants to assemble their own AI system.
Used prices on decade old hardware that used to go for peanuts has already gone up like 5-10x.
It's just going to get worse as more nations, companies and people get into the AI game.
Silver lining is that models will get more intelligent at far more manageable sizes, so we won't need a shitload of GPUs to run something at Fable level in the future.

>>109350953

It really depends a lot on what hardware we're talking about.
As far as computer hardware goes, companies working on hardware have already been utilizing in house AI for ages and same goes for a ton of other companies that have had money to build their own systems. So this won't radically change things for those guys.
However what comes from these models is that they lead to massive amounts of problem solving and novel solutions, that later translates to all kinds of hardware improvements.
Public AI causes a general intelligence explosion as now everyone can do problem solving on their own and this leads to a shitload of more innovation in general.
>>
File: dipsyKimiDario.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>109350888
Based.
>>
>>109351083
If you have to ask you can't run anything other than gemma and qwen.
>>
>>109350702
You said that last thread. Why don't you give it a shot?
>>
>>109350872
If you have Android 7+ you can install Termux, which helps a lot, but then you run into other issues.
I've a TV Box w/ Termux that I tried to set up as a sort of low power server, but it was impossible to get Termux to launch on startup (tried several ways, no joy, completely locked down.) So it's worthless for that application since a monitor and mouse need to be attached for startup. The hardware's too oddball to install other OS's onto. The exprience gave me a much better appreciation for what Android does and is good for, but not a device that would suit my purpose.
>>
>>109351089
>However what comes from these models is that they lead to massive amounts of problem solving and novel solutions
Yeah this is what I meant. I get the feeling we're gonna see some crazy breakthroughs in robotics over the next couple of years.
>>
Deepseek v4.1 will be Kimi K3-level performance at only 1.6T
>>
>>109351089
i would argue that's already been the case for a long time, its why you had gamers picking up used quadros in the past while having to compete with startup research companies that were buying those same used quadros. if you are a company and serious about building an AI cluster you aren't going to want 32GB VRAM cards, you will want 80GB+ cards, there's a segment in which consumers will be able to purchase used enterprise GPUs.
>>
>>109351101
can you not vague post, anon
>>
>>109351087
I don't enjoy it at all, that's why I'm trying to warn people that the end is near
>>
>>109351143
There have been new releases since the rentry was updated but you can't run any of those models.
>>
>>109351083
its easier if you just post your specs
>>
>>109351071
Nobody gives a fuck about your opinion you doomer faggot.
>>
>>109351157
>>109351157
>>109351157
>>
>>109351083
>I need other people to recommend me models because I can't form an opinion on them myself
>>
>>109351171
NTA, 24gb vram & 96gb ram
>>
>>109351198
ain't nobody got the time for dat
>>
>>109351185
I'm going to doom post even more frequently just because of your disapproval.
>>
>>109350702
build on llama.cpp or ik_llama.cpp or you'll end up wasting all your time re-inveniting tokenization etc
>>
Please recommend a harness that isn't npm shit.
>>
>>109351102
i decided i will, i was waiting for someone to find a flaw in the plan as an excuse not to do it would have been nice lol, probably will handcode most of it anyway.
>>109351236
i'll take some inspiration from it but i'm used to using rust now so that won't be a 1:1 copy.
>>
>>109351460
Someone did a super basic one in c whose only dependency was cjson a couple threads ago. It only does a single turn but you can probably wrap it in a bash script to get the looping. Or extend it like a normal person [spoiler]I'm in the middle of writing one in Rust.[/spoiler]
>>
>>109348370
Thanks, this actually seems to work.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.