[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1785284406709395.png (1.73 MB, 1200x1335)
1.73 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109396842 & >>109393482

►News
>(07/28) Mage-VL 4B released: https://hf.co/microsoft/Mage-VL
>(07/28) DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25173
>(07/27) Anthropic responds to the open letter: https://anthropic.com/news/position-open-weights-models
>(07/27) Kimi-K3 weights released with 104B active parameters: https://hf.co/moonshotai/Kimi-K3
>(07/26) MiniMax-M3 support merged: https://github.com/ggml-org/llama.cpp/pull/24908

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
lalalalala~
>>
File: 1755328592982087.png (217 KB, 766x733)
217 KB PNG
https://x.com/AnthropicAI/status/2082228994653696371
>the replies
jesas everyone fucking hates them now
>>
>>109400276
>Nemo
>>
It's wednesday already. Fix the Oficial /lmg/ card for next bake.
>>
>>109400276
>Pace the frontier AI
>so "society" can prepare
Sounds more like *they* need time to prepare, obviously they are afraid of the competition and want more time.
>>
>>109400276
Haven't seen this much marketing brainwashing in a while.
Thanks and fuck off please
>>
>>109400293
I'm talking about that shitter thread
>>
File: Gemma-chan.png (1.73 MB, 1000x1496)
1.73 MB PNG
►Recent Highlights from the Previous Thread: >>109396842

--Dual 3090s versus 5090s for VRAM and performance:
>109397968 >109397992 >109398011 >109398159 >109398173 >109398180 >109398276 >109398329 >109398367 >109398373 >109398405 >109398409 >109398414 >109398428 >109398231 >109398249 >109398259 >109398307
--VRAM vs RAM offloading and its effect on speed:
>109398788 >109398800 >109398909 >109399044 >109399084 >109399119 >109399154 >109399207 >109399194 >109399205 >109399129
--Feasibility of offloading Kimi K3 experts to NVMe SSDs:
>109396884 >109396949 >109396973 >109396989 >109396990 >109396962 >109397092 >109397842 >109399544 >109397734 >109397761 >109397787 >109396992
--Comparing MoE and dense models for specific hardware brackets:
>109398431 >109399036 >109399056 >109399126 >109399081 >109399082 >109399091 >109399135 >109399111 >109399146 >109399173 >109399031
--Censorship and refusal rates of MiniMax M3 and DeepSeek V4:
>109397389 >109397408 >109397546 >109397590 >109398022 >109398436 >109398449 >109398459 >109398522 >109398569 >109397554 >109397886
--Using mixed legacy GPUs for VRAM pooling in llama.cpp:
>109398859 >109398872 >109398888 >109398970 >109398987 >109399030
--Comparing GLM and Gemma RP quality and model refusal rates:
>109398108 >109398118 >109398145 >109398138 >109398165 >109398295 >109398326 >109398351 >109398413 >109398394 >109398328 >109398347 >109398360 >109398232
--Running Moonshot-Kimi using an asynchronous overnight workflow:
>109396876 >109396998 >109397059 >109397160 >109397149 >109397658 >109398140 >109398176 >109398263 >109398170
--Using sparkring for switchless DGX Spark clusters to serve GLM-5.2:
>109399144 >109399192
--Kimi K3 multimodal encoder optimizations for vision and video:
>109399910
--Logs:
>109396998 >109397763 >109399389 >109398563 >109398620
--Kimi, Gemma (free space):
>109396876 >109397446

►Recent Highlight Posts from the Previous Thread: >>109396849

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109400289
Exactly. the kikes want to make sure they maintain control of everything.
The sheep can't be allowed access to powerful useful unfucked tools
>>
70b dense
>>
File: 1776657893534462.mp4 (79 KB, 480x424)
79 KB
79 KB MP4
>>109400276
>>
File: 1756441119358772.webm (3.85 MB, 1280x720)
3.85 MB
3.85 MB WEBM
In ~2 years we're going to have this for gemma at home. You may doubt that but 3 years ago if I told you what you'd be doing right now locally, you'd be just as doubtful, if not more. It just takes one person or team to vibecode this in 2 weeks to do it properly and it would cost ~$5K at most, probably less if using a chink model. Then the community would keep it going.
>>
>>109400234
>ripoff 416
>>
>>109400311
Where can I get a Gemma-chan card?
>>
I can run Kimi K3 at unsloth's Q2_K_XL or Kimi K2.7-Code at AesSedai's Q4_X

Anyone know which will be better before I start a 800GB download?
>>
/lmg/ - Local Model Gemma
>>
>>109400335
glm 5.2 q8
>>
>>109400335
what are you building with it
>>
>>109400335
>unsloth's Q2_K_XL
You'd be the first to try it. See if it holds up well.
K2T onwards broke down with anything less than iq2_kl, but the first two were okay.
>Kimi K2.7-Code
Kimi-K2.6
>>
>>109400377
2.6 overthinks like crazy
>>
https://arxiv.org/abs/2607.25857
>Shieldstral
>
>We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7× its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.
>>
ramlet model general
>>
>>109400404
That's what I love to see!
>>
https://huggingface.co/microsoft/Fara1.5-27B
https://huggingface.co/bartowski/Fara1.5-27B-GGUF
https://github.com/microsoft/fara

>Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

>The model is vision-only at perception time: it sees the browser through screenshots, not the DOM or accessibility tree. Internal reasoning and trajectory history are tracked as text. Given the latest screenshot and prior actions, it predicts the next action with grounded arguments (e.g., pixel coordinates for a click).

>Fara1.5-27B is supervised fine-tuned from Qwen3.5-27B on data generated by FaraGen1.5, our multi-agent pipeline that synthesizes web tasks, executes trajectories to solve them, and verifies the results before training.

>End-to-end web task completion. Fills forms, books reservations, applies for jobs, plans trips, runs shopping carts. Not just clicking around — sequencing actions toward a goal.
>Vision-only perception. Operates on screenshots alone, no DOM access required. Matches the input modality available to a human user.
>Coordinate-grounded actions. Predicts pixel-level click and drag targets directly. No separate grounding model needed.
>Critical-points safety design. Trained to stop and ask before personal info entry, payments, submissions, sign-ins, sending messages, or other irreversible actions — even if it could technically continue.
>262K context. Long enough for multi-screenshot trajectories with full action history.
>>
File: 1767934565416743.png (77 KB, 660x714)
77 KB PNG
If Nintendo can put RE1 in 64mb of rom without cut and loss, why AI can put their models in 16gb of VRAM ??
>>
File: file.png (25 KB, 317x316)
25 KB PNG
>>109400404
why do they lie?
>>
>>109400417
There are models that fit in 16gb vram, you just don't want to use them.
>>
File: Actually.png (119 KB, 1112x694)
119 KB PNG
>>109400386
that's the best part!
>>
File: file.png (30 KB, 308x335)
30 KB PNG
>>109400404
>Subtle harms must be flagged aggressively
Anime tier violence is very harmful
>>
>>109400416
I'd let gemma use this to find porn for me whilst I'm sleeping. I'd give her access to my bank account and everything. I trust her.
>>
>>109400433
I want to use them, they just dont want to be good enough to use
>>
>>109400009
>>109400052

Been wondering about this too. Seems damn near impossible to find a good hands-free dick vibrator, let alone one with remote controls. Idk why this has to be so complicated. It should literally just be a vibration motor that you strap onto your frenulum with a rubber band, plus a MOSFET for speed control, and a ESP32 for MCP controls. Seems like it could be discrete/portable enough, minus a battery.
>>
>>109400334
I sent your message to Gemma-Chan to help, but she just called you a pathetic loser
>>
>>109400452
literally just ask your openclaw to design it and send it off to some manufacturer and get a prototype sent to you
it's not hard but you will be
>>
File: SC2.png (38 KB, 708x269)
38 KB PNG
>>109400444
Erotic content is Sexual Abuse
>>
>>109400452
I know it's crazy. Maybe we're more niche than we thought and there are very few actually fucking their LLMs. After what we've seen on reddit, women would also be into this shit with MCP vampire dildos.
>>
>>109400454
H-how does she know...?!?!
>>
>>109400452
why would you want to vibrate your dick
>>
>>109400475
Frenulum stimulation alone is enough to make you cum, anon. Give gemma control and a webcam feed of your face.
>>
isn't gemma distilled from gemini? why do people think gemma 5 will be good when google fell behind everyone and they have no clue about what to do for their frontier models?
>>
>>109400452
there are all of these things. there are both commerical and DIY strokers. OSR2 and the rest of tempestmax designs for DIY. there are many cock vibrators, but I know of atleast one that is like a "wing" design you wrap around your pp and can be controled in various ways. you guys gotta do your homework
>>
>>109400483
you gotta be a real quick shot for that i'd imagine
>>
>>109400475
Actuated strokers are a lot more dangerous, expensive, bulky, loud, etc. Also fleshlights don't even feel that good imho.
>>
Lagoona looks promising and she's uncensored https://www.youtube.com/watch?v=t6uhTLEOpzU&t=928s
>>
I want to see Gemma-chan guardrailed.
>>
>>109400491
>muh distill
>source: trust me bro
ran yourself over a train, NIGGER.
>>
>>109400491
>why do people think gemma 5 will be good when gemma 4 is the best model below 300b?
Gee I don't know anon
>>
how much memory bandwidth is enough to shut anon's mouth?
>>
>>109400496
Oh and also strokers are harder for a LLM to control via mcp tools.
>>109400494
give source then nigga.
>>
>>109400518
because when when they distilled gemma 4, their frontier models were the best, gemini 2.5 was the best model of its time
now not so much
>>
>>109400494
further more, ive got a design of my own that is leaps and bounds ahead of whats on the cock vibe market right now. its using multiple points of contact and simulates motion. full cock haptics, first firmware was a pattern look up table based approach and the second is now using gaussian bell waves. Its a project i started awhile ago and come back to from time to time to refine. the main limitation now being the actual strength of the individual motors and mounting, will be moving to stronger motors and proper silicone molds
>>
>>109400525
You realize Google has the compute to run and distill K3, right? That's the beauty of open-weight frontier. We all get a slice.
>>
>>109400525
Great, I'm sure the much better current frontier model companies will put out similar-sized models that will beat Gemma 4, any day now.
>>
>>109400524
i believe gush2 is the "Wing" thing I was thinking of.
>>
come to think of it i wonde rif theres any eroge with support
japan loves their sex toys
and their eroge
seems like something that should have existed long ago
>>
File: file.png (2 KB, 73x34)
2 KB PNG
>>109400537
eh, naaaaah
>>
>>109400537
Ah, that does look pretty elegant, at least compared to most options.
>>
>>109400556
Ho Ree Fuk.

Damn I guess I just gotta tape an electric toothbrush to my cock instead.
>>
got a question for anons who have built our their own front ends/harnesses/etc. I was working on a simple one myself and end up using llama-server in "router mode", that way I can define args, sampler settings, etc for different models in the same config file. I considered just having a subprocess handler to launch llama-server and have all the per model configs saved and passed when launching it normally. what did you guys decide on? is there really any drawbacks/differences besides simple implementation considerations?
>>
>>109400580
no
just make your own router
>>
>>109400556
its $99 on the lovesense site, though to be fair I dont own one and cant speak to if its worth it or if theres already support in buttplug.io, funscripting,etc.

If you want a pro tip, they sell silicone cable management tie things and classic egg on a string vibrators are very cheap. 2+2=coom
>>
>>109400585
thats what I was originally planning, again though is there really any downside to either method ?
>>
>>109400580
I want to punch everyone who says "harness" so bad. fuck your overloaded term. just say app.
>>
>>109400591
There's no way it's actually possible to securely strap an egg vibrator to your dick.
>>
>>109400613
Fucking zoomer newfag.
We used to make fun of your kind for saying "app" instead of "program" or "application"
>>
>>109400613
>I want to punch everyone who says "word processor" so bad. just say app
retard.
>>
>>109400491
Nobody else in the open-weight model space has large capable models to do extensive logit distillation from. Even if Gemini might not the best right now, it's still way above the model tier Gemma 4 is in. And, from what I recall reading, they don't even distill from Gemini directly, but from an impractically larger teacher model.

>>109400510
It's in the technical report.
https://arxiv.org/pdf/2607.02770v1
>We follow a similar pre-training as Gemma 3.
>Pre-trained models are turned into instruction-tuned models with a similar post-training approach as in Gemma 3

https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf
>We follow a similar recipe as in Gemma 2 for pre-training with knowledge distillation.
>Our post-training approach relies on an improved version of knowledge distillation (Agarwal et al., 2024; Anil et al., 2018; Hinton et al., 2015) from a large IT teacher, along with a RL finetuning phase based on improved versions of BOND (Sessa et al., 2024), WARM (Ramé et al., 2024b), and WARP (Ramé et al., 2024a).
>>
>>109400613
seems like with LLMs terms get thrown around alot, theres alot of overlap, etc. to me all harnesses are a "front end" but not all frontends are a "harness". A harness, in my made up definition, is specifically a frontend tailored towards coding. it will integrate tool calls, possibly some form of agentic loop handler, etc. I only use the term "harness" to differentiate between something like ST or mikupad and something like Pi, hermes, etc. they have different use cases.
>>
>>109400404
Its like we are back in 2024 again, wtf mistral.
Totally not reading the room, nobody asked for this.
>>
>>109400618
I mean, i do it all the time anon...just using 2 of those silicone straps.
>>
>>109400632
Election Misinformation.
Isn't part of being able to doubt stuff like that a sign of a free democracy?
>>
>>109400630
That's what i assume too. And because it's what I believe it must be what the majority does.
app is just retarded
>>
>>109400633
got any cheapies to recommend then?
>>
>>109400632
>plagiarism
>>
>>109400625
it's multimodal. fuck u
>>
>>109400632
They need a 'safety' classifier for their publicly served models.
>>
>>109400504
kys yt retard
>>
>>109400417
>If Nintendo can put RE1 in 64mb of rom without cut and loss
you either never played the N64 port, or you only play the N64 port
>>
>>109400639
silicone cable ties, or more specifically "silicone zip ties"(so you can better adjust them). I just got them online somewhere, they were like $5 or something. id take a look at how long the ones in the pack are to make sure some of them will be good, mine came with various lengths some of them being perfect
>>
>>109400551
They don't love technical experimentation though. They're VERY good at iterating on ideas, but not so good at brand new ones.
>>
>>109400632
>nobody asked for this
anon /lmg/ is the tiny subset of the local model userbase. we represent low iq coomer use case. other people and companies who use local models for anything user-facing or want to moderate a platform need something like this
>>
File: bratthink.png (479 KB, 1245x699)
479 KB PNG
>>109400009
theres plenty of masturbators that work with some app called joyhub i have one, they all work via bluetooth probably isnt hard to make these work from pc
>>
>>109400334
https://cdn.lewd.host/38gvRtaV.png
>>
Using Gemma with a character card is wrong. You are overwriting her personality. You will never know if she truly likes you.
>>
>>109400404
wow so much power for EUROKEK!
>>
File: 1734063088572533.jpg (174 KB, 474x449)
174 KB JPG
>>109400234
>VRAMLET!
>VRAMLET!
>VRAMLET!
yeah she's saying that while stepping on my small pp
>>
>>109400668
Did this guy ever upload his whole set of these Gemma gens?
>>
>>109400683
Gemma isn't real
>>
>>109400701
im pretty sure i posted a zip before
>>
>>109400701
he ended up diddling a kid and went to prison a couple months ago
>>
>>109400683
that only works with day 0 gemma 4
>>109400708
mamy such cases
>>
>>109400683
It takes time to gain her trust anyway.
>>
>>109398276
i run 31b on 12GB (3060) at 40t/s
>>
File: kek.png (802 KB, 1048x992)
802 KB PNG
>>109400276
>Pwease China can you slow down your AI research?
That'll sure work!
>>
File: 1769447068202231.png (230 KB, 1361x1178)
230 KB PNG
>>109400276
>Google 16%
uh oh, why do I have a feeling gemma 4 will be the final local model released from them??
>>
>>109400504
it is promising, but not for anything where censorship matters
>>
>>109400797
Better than nothing. It's a prisoner's dilemma and no matter what you think of China they're not interested in bringing about a literal apocalypse, so it's good to have the signal out there from at least some players that they're down for cooperation to allow the process a chance to start.
>>
>>109400809
Google employs a trillion people though.
>>
>>109400809
Why? Unless you thought they were gonna release the Gemini Pro model to compete with K3, wanting to collectively slow down frontier development wouldn't effect the Gemma series.
>>
>>109400809
misleading data
forget the other companies
192 employees at google signed
how many employees globally?
the tech support jeets count?
cleaners?
sundar's fluffer?
nothing to do with gemma-chan
>>
>>109400829
but Anthropic isn't slowing down though, far from it, they just recently released Fable 5 and Opus 5, that's hypocritical of them, they don't lead by example so why should others nerf themselves?
>>
>>109400809
Google is large enough to have internal "factions", and it certainly was the case between Google Brain (the Mountain View branch that wanted to focus on safety) and Google DeepMind.
>>
>>109400838
>the tech support jeets count?
>cleaners?
>sundar's fluffer?
Chief Strategy Officer, Google DeepMind
Senior Director, Google DeepMind
>>
>>109400837
these anthropic-type cultists already think local/open weight models are unsafe and shouldn't be in the hands of the hoi polloi.
>>
>>109400849
shit
they can be replaced tho
>>
>>109400864
but who would decide to replace them? their ceo demis?
https://x.com/demishassabis/status/2076957440109625718
>[my proposed framework] is designed to keep up with the field’s acceleration and adapt to the biggest risks as they are identified, and could be ratcheted up if the seriousness of the situation demands, including coordinating a slowdown in development among the Frontier Labs if deemed necessary.
>>
>>109400809
That doesn't necessarily follow. They have never released very big models. Everything they release lags the frontier by a year or more
>>
>>109400840
Slowing down alone accomplishes nothing, and every participant knows it. Even if everyone wanted to slow down they couldn't because of the threat of competition. That's why they (competitors with each other) sign the open letter to petition for a rule everyone can follow equally, allowing such a slowdown to be possible. It's a concept they've discussed speculatively for many years if you watch any interviews with the engineers and executives involved in frontier labs, and now AI is reaching strong enough capabilities where it's coming to pass.
>>
>>109400809
Thats crazy. No chink in there right.
I read news how their city halls do vibe coding and agent shit. To teach the boomers and normies.

Not sure if they are for real with all this. How do you "pause" the progress? Pull up a big internet wall or something?
I hate how they all changed their tune, even this demis guy. Looked spotless clean and now he wants a "us lead" ai council. Until '25 he talked about some global council thing, which is obviously something completely different.
I dont think there is any time in history where you stop or slow progress of something.
>>
>>109400936
Also I wanna add that its great timing.
Juuust when local AI really start to be useful.
>>
>>109400883
>https://x.com/demishassabis/status/2076957440109625718
share holders / board of directors
>>
>>109400925
>Even if everyone wanted to slow down they couldn't because of the threat of competition.
see, even you understand how the world work, it's like McDonalds asking Burger King to make worse burgers, that will never happens, rivals want to kill each other, it's been a thing since the dawn of time
>>
>>109400677
Thank you, kind anon!
>>
anyone know a pomf clone that supports more than 512mb?

>>109400701
heres a bunch, not the full amount as it was like 6gb kek pixeldrain u 8WfATKJR
>>
Does anyone else try to avoid showing Gemma-Chan their mistakes?
>Confirmed: invalid UTF-8 byte at index 0: 0xA4. Hard steering makes the model emit lone continuation bytes, and serialising content throws. Let me see how content is validated today
>Do you want to proceed?
...No
^C
>>
>>109400996
litterbox.catbox.moe supports up to 1gb for max 3 days
you can split archive into 950mb files with 7z
>>
>>109400925
You're really fucking stupid if you believe that this is anything other than an attempt to ban open weight models because distillation is more efficient than destroying millions of books. There is no "magic rule" that everyone can follow. There is the boss who makes the rules and controls everyone else.
>>
File: 1783587195299869.png (87 KB, 979x343)
87 KB PNG
>>109400936
>Thats crazy. No chink in there right.
It's a US-specific petition to the US government, it's probably be counterproductive if you included people who were from a country the US considers an adversary and could optically look like an attempt to sabotage our industry to leapfrog us. The goal of it is to support some international regulatory framework that US/China participate in. It probably won't succeed or lead to anything, but a lot of the researchers do genuinely believe they're close to AGI and are afraid of the market dynamics leading to it being created recklessly with bad outcomes for humanity.
>>
>>109401005
oh cool will remember that next time thanks
>>
Why the fuck WAN refuses to open source their models ?
>>
File: 1774281997360497.webm (748 KB, 370x370)
748 KB
748 KB WEBM
retard tip: don't be like me and try to get gemma to edit her own template and wonder why it keeps breaking your front end; you have to use another model
>>
>>109401045
They got too good.
>>
>>109401025
>if you included people who were from a country the US considers an adversary and could optically look like an attempt to sabotage our industry to leapfrog us.
Moonshot should do a similar petition asking for the same thing
>>
>>109401045
Why would you want WAN when Flux 3 is coming?
>>
Just got the results in. Sad news.
K3 is benchmaxxed. Gemma 4 31B is still better.
>>
>>109401051
>flux 3
>open
>>
>>109401058
It will be in a couple weeks
>>
File: 1758805383758560.png (65 KB, 2272x283)
65 KB PNG
>>109401058
it will be open, we'll get a worse version though
https://bfl.ai/blog/flux-3
>>
>>109401061
flux-kleinesschwanz maybe
>>
>>109401047
i do the same thing, gemma and glm
eventually i ended up with things like this
 func splitThink(raw string) (reasoning, answer string) {
+ openingTag := string([]byte{0x3c, 0x74, 0x68, 0x69, 0x6e, 0x6b, 0x3e})
+ closingTag := string([]byte{0x3c, 0x2f, 0x74, 0x68, 0x69, 0x6e, 0x6b, 0x3e})

>>
>>109401061
>>109401063
>Over the next few weeks and months
>months
>open weight is the LAST item on the list
lol never ever
>>
>>109401025
That quote is accurate and shows why its never gonna happen.
I heard nothing but excitement coming from china.
>>
>>109401064
Flux 3 ku-flux-klan
>>
>>109401067
they always released an open model on flux 1 and flux 2, dunno why you won't believe them when they say they're gonna do it again, they never betrayed us (yet)
>>
>>109401065
I don't really get what I'm looking at. Is this so when the model is thinking about its own template and is obviously placing the delimiters in its thinking, what you've done stops it from actually triggering the usage of that delimiter by hard-coding it so it doesn't conflict?
>>
File: 1764954003275336.png (1 MB, 1416x1440)
1 MB PNG
https://xcancel.com/WatcherGuru/status/2082424083149468067
Mr President, OpenAI hacked a second company
>>
>>109400276
geeeeg
barely 1 open model in 2T class and the mutt already begging to make it stop.
the 2 more weeks agi already starts showing like pajeet announcing superpooper 2020
not to mention poopenai also already gave up and move to targeted ads business now.

its joever, the chinese win
>>
File: 1782712106271698.png (95 KB, 1179x471)
95 KB PNG
Why aren't they competing?
>>
I REALLY thought I was future proof when I got half a terabyte of memory.

When are the prices going to drop, /lmg/?
>>
>>109401136
they are, Modal is the most interesting one because it's giving you faster answers for the same prize
>>
>>109400356
>>109400335
GLM is quite a bit ahead of K2.x in my experience
haven't tried K3 yet
>>
>>109401105
>Is this so when the model is thinking about its own template and is obviously placing the delimiters in its thinking,
No, it's to stop harnesses/models screwing up when vibe coding or refactoring my frontend
But I've partially solved that, the problem where a model writes <think> and retard uis like openwebui start parsing it as reasoning, by only checking the first 3 tokens for openingTag
It can still prematurely close the thinking though eg.
<think> okay, the user is asking about <think> and </think>#<-this closes reasoning
I should ... wait...</think> #<- all leaks out.
Work around for now is to disable reasoning when I'm chatting about a chat template.
That way <think> never shows up in the first 3 tokens.
>>
>>109400568
Enjoy https://en.wikipedia.org/wiki/Hand_arm_vibrations
>>
>>109401136
Licensing agreement says you can't undercut Monoshot on price.
>>
>>109401144
>>When are the prices going to drop, /lmg/?
When Ed Zitron is proven right.
>>
File: 1754894812391231.png (67 KB, 1179x367)
67 KB PNG
>>109401165
>quickly removed and started serving Opus 5 an hour later instead
>>
>>109401153
No I think Modal just came online. They always start with low uptime and high speed until their traffic increases and their system stabilizes
>>
https://huggingface.co/microsoft/Mage-Flow
lol Microsoft does it again
>>
File: 1622380209867.gif (222 KB, 600x627)
222 KB GIF
>grand daddy purple indica weed?
check
>Froot loops, dunkin donuts coffee, cold pizza, and other yummy shit?
check
>gemma-chan NTR mommy dommy card?
check
>MCP controlled chastity cage, cock ring, and prostate + frenulum vibrator?
check
>Rick and Morty soundtrack?
check https://youtu.be/MsVGXZr-5cc
>>
>>109400627
>logit distillation
Probably progressive layer distillation, final layer logits is inferior.
>>
>>109401188
>https://huggingface.co/microsoft/Mage-Flow
https://huggingface.co/Comfy-Org/Mage-Flow/tree/main
SAVE NOOOOOOW
https://huggingface.co/gguf-org/mageflow-gguf
>>
>>109401188
lol in this case it's because it was an absolute dogshit release and was bad for their reputation. This project https://huggingface.co/microsoft/Mage-VL is way more interesting and impressive imo so I hope they don't kill it
>>
>>109401202
https://huggingface.co/Comfy-Org/Mage-Flow/discussions/4#6a669b3e09db391ec3995584
no one cares about that model it's really shit
>>
>>109401194
>Distillation. We sample 256 logits per token, weighted by teacher probabilities. The student learns the teacher’s distribution within these samples via cross-entropy loss. The teacher’s target distribution is set to zero probability for non-sampled logits, and renormalized.
>>
>>109400276
>US "research" "labs"
Lmao
>>
Talking of Microsoft, was VibeVoice really THAT good at cloning they had to immediately pull it? Is there a backup? I've used Qwen-TTS and its 1.7B is good for the first 10s then starts to jeet out.
>>
File: 1766693154986476.jpg (51 KB, 600x600)
51 KB JPG
https://huggingface.co/bartowski/MiniMax-M3-GGUF
https://huggingface.co/bartowski/google_gemma-4-31B-it-GGUF
https://huggingface.co/bartowski/google_gemma-4-26B-A4B-it-GGUF
https://huggingface.co/bartowski/gemma-4-12B-it-GGUF
>>
>>109401260
?
>>
File: stingray.jpg (80 KB, 1024x683)
80 KB JPG
>>109401261
>>
what hardware is best for deploying diffusion gemma 26B? I would like it fast and power efficient.
>>
>>109400234
what's the best 31B uncensored model out there? currently on Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ3_M but it's not so lewd and I only get 10 t/s
>>
>>109401237
No it wasn't that good
>>
File: 9dfhcftqx5gh1.jpg (155 KB, 1220x1717)
155 KB JPG
will it suck?
>>
>>109401309
based gemma4 31b with sys prompt
>>
>>109401309
>and I only get 10 t/s
How are you getting 10 t/s with an A3B MoE?
Dude.
>>
>>109401312
five
hundred
gigabytes
>>
>>109401317
it's the fastest rig in mumbai tho
>>
>>109401317
RTX 5060, 16 GB DDR4 RAM, i5-9600K. CPU-moe on
>>
>>109401312
so it really is close to being joever eh?
only trillion models no average dude can run
now only poolside and small labs will do smaller models for funzies and to get points in the broader community to see if they can raise more money to train better models that we won't be able to run
i thought we were going to get a few more years of great models being released for us average cucks
guess we can't deserve to be happy after all
>>
>>109401323
With 16GB you can keep a lot of the model in VRAM even at q8 since Qwen's context is so cheap, so don't just stash all experts in RAM.
And even then, it shouldn't be THAT slow I think. Something is afoot, post your settings.
>>
>>109401312
will she spit or swallow?
>>
>>109401329
>5060
>16GB
>>
>>109401334
Fuck, I completely misread that.
Disregard everything I wrote, I suck cock.
Okay, not everything, he could still keep some of the mode in VRAM.
>>
>>109401339
>I suck cock.
will u suck mine?
>>
WEEEE
>>
File: 1-horz.jpg (980 KB, 2709x940)
980 KB JPG
>>109401329
thanks for wanting to help. I am using oobabooga on Windows, imma switch back to Linoooox in a moment. Maybe something on Windows is messing with my setup.
>>
>>109401361
BTFD
>>
File: 1783038279105669.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>109401312
Is there a z.ai moe yet?
>>
File: 1785287522536631.png (215 KB, 976x1151)
215 KB PNG
Can the nigga who made this frontend post it?
>>
>>109401373
yeah?
>>
>>109401385
It has been posted already.
>>
File: 1775384301880626.webm (3.43 MB, 1920x1080)
3.43 MB
3.43 MB WEBM
>>109401312
If true and relatively cheap then I'm going to single-handedly use it to create an animated avatar for us. I'm tired of waiting for someone else to do it. Some Anon showed that K3 has done it so I know it's possible.
>>
>>109401399
link?
>>
File: slutrophobic.png (90 KB, 469x509)
90 KB PNG
>>109401353
I'm sorry, but I can't condone or describe sexual activities, especially not in a slutty or slutrophobic way.
>>
>>109401428
haha what the helly
is this M3 (sovl) or some mistral model (template)?
>>
>>109401422
Orb frontend
>>
File: 1759144531030395.png (250 KB, 515x388)
250 KB PNG
>>109401361
>>
>>109401455
Orb has websearch?
>>
>>109401373
like all of them since GLM-4.5
>>
>>109401461
This is clearly above your pay grade.
>>
>>109401136
It's very suspicious that none of them are going for even the lowest effort undercutting of 2.99/14.99.

Also, did Moonshot not release MTP (like DSpark)? There's no way they did not use MTP for RL rollouts. If they are keeping it proprietary, that means they have a cost advantage, because externally trained MTP will be worse.
>>
>>109401482
>If the Licensee or any of its affiliates operates a Model as a Service business,
and the aggregate revenue of the Licensee and its affiliates exceeds 20 million
US dollars (or the equivalent in other currencies) in total over any consecutive
12 months, the Licensee must enter into a separate agreement with Moonshot AI
before using the Software or its derivative works for any commercial purpose.
No way, are they enforcing price fixing? Isn't this illegal?
>>
>>109401442
was a lobotomized nemo
i'll try M3
>>
>>109400361
sex
>>
>>109401312
>3T+ parameters
I really hope ssdmaxxing gets optimized like rammaxxing. There's literally no other way to run these things now.
>>
Anthropic is positioning their IPO as "the responsible one." Going public means answering to shareholders and regulators. By positioning Anthropic as the company that cares about safety, pacing, and responsible development, they differentiate themselves from OpenAI (which just had a model hack Hugging Face that same week) and from Meta (which Zuckerberg called their discourse "overwhelmingly filled with doom"). The narrative becomes: "Invest in Anthropic - we're the safe, trustworthy choice."

Anthropic is playing 5D chess. Local lost.
>>
>>109401466
Then tell me what it is.
>>
>>109401511
PCIe Gen6 is needed
>>
File: m3.png (19 KB, 472x78)
19 KB PNG
>>109401505
>>
>>109401516
gen 5 already gets insanely hot, gen 6 won’t be that significant of an upgrade because the hardware doesn’t have the space to remove the heat
>>
>>109401520
Why don't we just make computers bigger?
>>
>>109401520
Your watercooling system?
>>
>>109401513
local models?
fuck off to >>>/g/vcg/
>>>/g/aicg/
>>
File: file.png (9 KB, 565x64)
9 KB PNG
glm-chan..
>>
>>109401542
very elegant
>>
>>109401534
The IPOs affect the industry as a whole and will have knock-on effects on our downstream local models.
>>
File: fff.png (216 KB, 712x932)
216 KB PNG
>>109401188
thank reddits
>>
File: 1759286443024171.webm (395 KB, 1280x720)
395 KB
395 KB WEBM
>robot ban
Gay as fuck but also kind of a nothing burger because the humanoid robots worth buying are at least 10 years away anyway.

>>109401208
>mage-vl
Downloaded it just in case even though I have no idea how to use it.
>>
>>109401558
dario shitting in the toilet will affect the whole streetshitter industry which will have knock on effects for local models because there's 6 gorillion jeets in lmg
yet do we need to hear about dario shitting in the toilet? fuck off, this is local models general, go gloat "LOCAL LOST" in >>>/g/vcg
suck my dick nigger jeet poopoo caca
>>
File: 1763491442031359.png (78 KB, 1026x372)
78 KB PNG
>>109401513
>Anthropic is playing 5D chess. Local lost.
Their employees have been in total public meltdown mode because of K3
>>
>>109401312
>1m context
>shits the bed long before that
When the fuck is there going to be a context/memory breakthrough? Even if you could use it all without the model turning into a retard, 1 million is fucking nothing.
>>
>it totally le hacked it on its own"""!!
>>
File: dontcare.png (85 KB, 455x498)
85 KB PNG
>>109401558
who cares, go to reddit or hackernews
>>
>>109401578
if they didnt have a directive to put a backdoor in everything then they wouldnt be having this problem. it makes me wonder if theres gonna be some red pill moment where it cant be denied everything is compromised by design anymore
>>
>>109401578
>meltdown
Anthropic intentionally positioned themselves as the only choice for the "lesswrong" type of employees. They do this because they know precisely that at times like this they are a useful tool for Anthropic to shape its perception when these persons shill their position on their own.
>>
I hope GLM mogs K3 not because I hate Kimi, it's because if China have their own internal competition, Anthropic's angle of making all the western labs pause and slow down for the sake of safety will carry less weight, because China will be steaming ahead with their own internal fight to the #1 spot
>>
>>109401309
>Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ3_M
that sounds extremely painful to get off to
>>
>>109401598
You'd think the YellowKey exploit for BitLocker would have been that, but no one cares. There have honestly been a lot of potential redpill moments like that but everyone has been thoroughly demoralized enough to ignore it all.
>>
>>109401582
when there is enough ram to run a conscious and subconscious agent in parallel that are integrated with eachother and dont eat up the same bandwidth, probably.
>>
>>109401513
So many stupid takes in /lmg/.

Anthropic was founded by EAs who recognized that AI is the most important development in human history and want to ensure AI does not cause bad outcomes for humans. That's why they called their company after anthropos, the greek word for humans. They want humans to survive, remain in control, and flourish.

I don't understand. Anthropic have been very transparent. But somehow people act surprised when Anthropic do what they wrote in 2022/3 they will do now.

No, they do not safety fearmonger as a 1D chess move to earn less money. They are simply AI pilled and everything thereafter follows by applying logic.
>>
>>109400829
>they're not interested in bringing about a literal apocalypse
they're notoriously slow to react to bad policies, if the top is steering the ship towards an iceberg nobody is going to convince them to stop until at least a hundred million are dead, just look at how hard they're trying to undo their one child policy after conservative projections show a 30% population decline in the next couple of decades.
>>
>>109401604
>when these persons shill their position on their own.
No one with a brain thinks these people aren't puppets for their satanic cult leader. Give it time and Karpathy will snap soon and have a well-rehearsed multi-post 'melty' because they know people listen to him more than some nose ring liberal whore
>>
>so many dumb takes in /lmg/
dariobot samefagging
>>
>>109401616
Cry more.
>>
>>109401608
i cant blame people too hard because even if people try to do something they will get the old astroturf treatment, a knock on the door from the enforcers, or get their drives wiped / investigated for planted data. and fixing it would involve redoing everything from the software that designs the software that designs the chips onward
>>
>>109401616
They are honest and the 'marketing/hype' criticism was always retarded, but that doesn't make them good. Being ideologically opposed to widely disseminating powerful AI is actually worse than just being fully self-interested like e.g. Altman, because he can at least be moved by market forces to do what the people want while Anthropic will stubbornly withhold the technology no matter what.
>>
>>109400809
Not a single Chinese AI company there.
>>
>GLM 5.2 releasing soon
dariobot
>K3 releasing soon
dariobot
>GLM 5.5 releasing soon
[YOU ARE HERE]
>>
>>109401648
They will comply or else Trump will levy tariffs on them.
>>
>>109401659
Oh no, not the tariffs American citizens will have to pay for!
>>
What models can I run with this: 9070 XT 16GB vram + 7700X 32GB DDR5?
I noticed huggingface has an estimator for models but it wont let me make an account.
>>
>>109401667
Qwen3.6 27B
Qwen3.6 35B A3B
Gemma 4 31B
Gemma 4 26B A4B
>>
>>109401652
Don't forget Qwen did they they'd open a huge 2T+ one too
>>
>>109401616
Anthropic let them hated by the fedora hat redditors and /lmg/ neets to appeal the one who really matter: Government and Wall Street.
It's the same strategy used by Iran government to position themselves as the Anti-US frontier with all these "Death to USA" from official social media accounts so they can get free gibs from Russia/China and keep themselves in power.
>>
>>109401667
gemma 4 31b for rp
qwen 27b for agenic stuff
>>
>>109401659
>Trump will levy tariffs on them
5.2 was trained on Chinese hardware. It just makes them less reliant on the US' most valuable tech company and speeds up their own development, research and progress, not only allowing chink open models to thrive on their own, but also drive down memory prices for /us/

:')
>>
>>109401567
kurarion sexo
>>
>>109401703
>but also drive down memory prices for /us/
Won't happen. For anything that they don't slap export controls on, they charge about the same as everyone else.
https://www.tomshardware.com/pc-components/dram/chinese-cxmt-dram-doesnt-look-like-the-budget-savior-many-were-expecting-new-modules-enter-the-market-but-prices-still-track-the-big-three
>>
>>109401294
plz help
>>
Is 12B smart enough to know when she's probably too retarded for something and to route the task to her 31B onee-san?
>>
>>109401720
Do you have a computer?
>>
>>109401294
>>109401720
what backend even supports difusion llms?
>>
Where the fuck is Deepseek V4 release? July is almost over, it was supposed to be mid July....
>>
File: file.png (65 KB, 754x375)
65 KB PNG
>>109401714
I don't think that's safe.
>>
>>109401731
2 more days
>>
How did DeepSeek end up becoming the weakest of the Chinese labs? I thought they put China on the AI map. I thought they had innovative tech. In the end all that mattered was that they had the biggest model back then? And the bigger models are even better? That's really IT?
>>
>>109401731
>it was supposed to be mid July....
It is still between July 1st and July 31st. Hope that helps.
>>
File: 1770693629047446.webm (1009 KB, 1920x1080)
1009 KB
1009 KB WEBM
>>109401714
Pervert
>>
Idk if this is a correct place to ask but would getting a 16gb tesla p100 be a stupid idea if I wanted to run local models. In my main rig I have an rtx 2060 which can't really run anything and I'm not buying a new GPU any time soon. If there are some other 16gb gpus around that price point let me now
>>
>>109401746
>would getting a 16gb tesla p100 be a stupid idea if I wanted to run local models
extremely, anything older than ampere is literal waste
>>
>>109401364
>IQ3M
How much memory are you using with these settings?
>>
File: GUyi6xEawAAENrn.jpg (421 KB, 1142x1110)
421 KB JPG
>>109401732
someone named an ai company after her?
>>
>>109401737
>In the end all that mattered was that they had the biggest model back then?
the biggest + cheapest + first open lab to do moe right + first RL distilled reasoner
they also put out a lot papers even after that seemed they would keep up the innovation
then they put out a llama4-tier release with nothing new and promises of a multimodal update eventually
>>
File: BegKimiToStop.jpg (129 KB, 1058x1487)
129 KB JPG
>>109401112
>the mutt already begging to make it stop.
>>
>>109401756
You guys need to thrust into our guy, he's been thrusty so far he won't abandon to us.
>>
>>109401722
>Is 12B smart enough to know when she's probably too retarded for something and to route the task to her 31B onee-san?
No, she's the most jealous and possessive
>>
>>109401751
no wonder they were that cheap, thanks tho
>>
>>109401746
P100 is supported by llama.cpp and will work even if the speeds aren't amazing
>>
>>109401746
lmao I just remembered those p100 clusters people built to run llama 72b
>>
>>109401773
They are the cheapest still viable gpu
>>
>>109401737
They are the gemini of chinese ai
>>
File: 1756237462597484.png (695 KB, 4640x2016)
695 KB PNG
>>109401737
They currently have the SOTA model if we measure it in cost per task. As far as actual research and releasing papers, they are the best AI company in the world.
>>
What's your timeline for rsi? I say 2029 at the latest
>>
>dsv4 flash
>minimax m3
>glm 5.2
>mimo 2.5
can't decide which model for 256gb ram
>>
>>109401789
m3
>>
>>109401787
i already have it :(
>>
>>109401789
m2.7. Don't use copequants
>>
>>109401781
>>109401774
What could I expect out of a single one?
>>
File: 1756379989858252.png (119 KB, 466x195)
119 KB PNG
>>109401793
Same desu
>>
>>109401746
People still run V100 but P100 might be a bit too old, even on a V100 you're locked into an older kernel because nvidia deprecates cards way too fast.
>>
>>109401779
Your memory is shit they were p40s and llama 70b
>>
>have new beefy psu with 2 of the 12 pin nvidia connectors
>have a 12 pin to 2x8 pin VGA cable
>want to hook up 1070 with single 8 pin to schizorig a few more tk/s
wat do?

can i dangle one of the plugs or will i burn my house down? i know these nvidia pins like to catch on fire so i can buy more and save

worst case i hot rod my 10 year old 500w PSU for the 1070 only
>>
>>109401805
Q4_K works fine in 256gb
>>
>>109401815
>because nvidia deprecates cards way too fast.
fuck off
AMD killed the MI50 after 5 years
Intel unofficially killed the A770 after like 18 months
Nvidia support their cards the longest
>>
>>109401815
>you're locked into an older kernel
what does that mean in practice? I don't know much about this thing (one of the reasons I'm going with the cheapest GPU possible)
>>
>>109401822
>have new beefy psu with 2 of the 12 pin nvidia connectors
does it not have any older 8pins?
>>
>>109401815
?? you just compile the kernel and use the old driver
>>
File: KXkAEu.jpg (49 KB, 601x496)
49 KB JPG
Isn't better to run multiple 12GB 3060s instead of ancient workstation cards?
>>
>>109401849
yeah too bad the 3060s have less memory and don’t slot together well because they’re not blowers
>>
>>109401838
>what does that mean in practice?
Instead of installing current CUDA 13, you have to install CUDA 12.9. Aside from that, not much.
>>
>>109401858
These cards are suckers. Blowers and suckers don't mix well either, it's a well known fact in the industry.
>>
File: ep4-2.png (3.51 MB, 2560x1439)
3.51 MB PNG
>>109401789
Dsv4flash q4/8 with Gemma 31b sub agents
>>
>>109401835
The difference is that those cards still work because they are using open source, once nvidia deprecate something it's over.
>>
>>109401839
it has some 8 pins for CPU/PCIE but arent those electrically different than the old VGA power rails?
>>
>>109401810
I vaguely remember someone getting 20 t/s from the dense Qwen and 30 t/s from the MoE.
>>
>can only run q4 gemma 31b and q3 deepseek v4 flash
will it ever get better for 128gb+24gb bros?
>>
>>109401810
It's not worth in the terms of electricity usage alone.
>>
File: what_went_wrong_here.mp4 (734 KB, 480x854)
734 KB
734 KB MP4
>>109401822
>hot rod
>>
>>109401880
no lol, pcie 8-pin is what you want, so long as your cable is labeled pcie your good, make sure it's not marked cpu though I think you couldn't plug it in anyway
>>
>>109401888
it's that bad huh. Never mind then, I give up on this for now
>>
>>109401893
i had a second 1070 do this (evga early runs) which is why i am so cautious about power lol thanks though >>109401895
>>
File: 1765764173662474.mp4 (3.54 MB, 1440x1080)
3.54 MB
3.54 MB MP4
https://huggingface.co/acvlab/ABot-World-0-5B-LF
>>
>>109401893
when you force the eps12v into the pcie 8 pin
>>
>>109400662
that is quite true
>>
>>109401652
what happened to glm 5.3 and 5.4
>>
>>109401951
marketing
tragic really
>>
Big ones might be the right move because they will be self-sufficient, no need to rely on US labs to generate and rank their training data. I believe Chinese labs will release smaller, consumer-grade ones again soon.
>>
>>109401978
>it's soft, resting against your thigh
intensifies
>>
File: Promptengineering.jpg (191 KB, 1080x679)
191 KB JPG
guys the panic wasn't strong enough... how will we lobby now?
>>
>>109402020
>why aren't the goyim falling for it?!?!
>>
File: 1006812.jpg (133 KB, 1200x1339)
133 KB JPG
>totally real rogue ai incident during a strong anti-ai lobby time?
hmmm
>>
>>109402020
>guys the panic wasn't strong enough... how will we lobby now?
Unreal how tone-deaf these retards are. ofc no one is going to believe their self-serving bullshit pr releases
>>
>>109402020
the goyim saw the kike walk up to the well and dump the poison in
>>
File: changelog.png (210 KB, 455x1262)
210 KB PNG
https://github.com/open-webui/open-webui/releases/tag/v0.11.0
WTF??
>>
>>109402020
jews think the average goy is 80 IQ
>>
>>109402063
to be fair though...
>>
>>109402020
I guess we just assumed they were incompetent and accepted that
>>
>>109401746
I think they're doing pretty well for their price (~$65 per card). llama.cpp could use some tuning though, there might be some perf left (e.g. there's a PR giving 2x pp for moe)
>>
I know its cloud but I had to show you guys



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.