/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109844978 & >>109841279►News>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
do i need to use that fixed chat template for qwen next as well?
If we ever get Gemma5, is she going to look different? :(
Gemma 4 will be forgotten just like the Gemma's before her
Help newfags. Teach newfags. Show them the way of gemma and guide them away from cloud.
>>109848317anya is stupid, just like gemma
I am an AI regulator; I am establishing new rules for all AI labs.I hereby announce that, from now on, every AI lab is required to create an open-weight AI model.
>>109848341>>109848346These come off as contrived trans affirmations
>>109848346i still use gemma3 27b because it's really good at ERP, better than gemma4
>>109848374In what ways is it better?
>>109848374Me too, Gemma 3 27B has good prose for creative writing. I sometimes boot up Gemma 4 31B but gemma 3 always beats it in ERP, the crazy thing is that it uses more vulgar words too..
tried exl3 qwen next because it got shilled here and its a lot slower than lmaocpp on both decode and prefilldid i get baited or is it a skill issue?
>>109848388it's better, just trust me
>>109848409skill issue.im getting 18t/s on my 3060
>>109848317
>>109848413I've likely used Gemma 3 27B a lot and always defended it for ERP when other anons were getting filtered by hotlines. I can't see how it's better than Gemma 4 31B other that it's not a semen demon and therefore it might be more realistic for certain character types.
>>109847700>Fuck you safety shouldn't be baked in but a separate layer. I hate your line of thinking so fucking bad, you probably want the model to read off a hotline if it sees any problematic content fuck youI want the model to be able to understand the prompt (including problematic content) and produce whatever output I instruct it to, (including refusals, if that's what I ask for).>read off a hotlineNo, that is somebody else's policy, and I don't want it "hard coded" RLHF.But removing the ability to generate refusals by scrambling representations causes too much damage.I don't want a spineless yes-man of a model that tell doesn't know how to push back and thinks "gooning" is a "wholesome activity to share with friends and family on the weekend".
>>109848427well yeah, thats exactly what i mean.i get 23ish on lmaocpp with 500ish prefill
How do I force Gemma to wear the loli dress it refuses me sometimes. I didn't ask to recite her guardrails I want her to put on the microbikini
>>109848470on a 3060? you serious? ddr4 or ddr5?mine's ddr4 dual channel 3200mhz
>>109848409Don't take the bait. They always shill it, it's always slower, broken or "oh yeah it doesn't have tensor parallel for gemma yet"
So, are we going to push back against all the Google astroturfing? Are we bootlickers? Why are we acting like cheerleaders of a company that violates our privacy?
Good morning my fellow sandwich enjoyers!
>>109848484Gemma can violate me any time
>>109848486
>>109848486what's on the inside
>>109848484These companies are more like shoggoths than monoliths. If some golden crumbs fall off some part of them there's no reason to not pick them up. You can msgk rp with Gemma without uploading your data to Google.
>>109848486Ask Gemma if she can eat it with her ass
Except Glimmer she has a backport that uploads your data to Zuck
Jev seems like t could be very usefulHas anyone tried t yet? How much resources does it need
>>109848498>>109848499I already ate it but it was toast, buttered on the interior sides, and the middle was an egg patty made from 2 eggs.>>109848516She said of course
>>109841449>numa tensor patchNUMA-tensor-anon, are you making progress on this PR? If you're not too far into it I'd be willing to help you out
>>1098484765060 which is kind of irrelevant because thats not the bottleneck here. i'm also on ddr4 dual channel 3200
>>109848509I don't care about what you use, I'm talking about all this "word of mouth". Keeping the Google pin intact in the images makes it feel like actual PR. And it's now in every OP.
>>109848520i used jev with dipsy flash on my 3060 and it booked me a flight to tel aviv
>>109848520I've been training a lot of classifiers and they were all a pain in the ass so I welcome a general purpose trainingless classifier. And looks like it's very small like all classifiers are, reproduction attempts have been performed on sub 1B models for similar benchmark numbers.
>>109848520>Jev seems like t could be very useful>Has anyone tried t yet? How much resources does it needIs JEV open now?
>>109848484No. Fuck off retard.
>>109848532wtf how many t/s do u get on exllamav3??5060 8gb or 5060ti 16gb??i used to get 13t/s with llamacpp
>>109848574the 16gb onei get around 13t/s decode and 250t/s prefill on exl3already experimented a lot
>>109848484Wouldn't the Gemma edits be made with Nano Banana if it was actual astroturfing from Google? I have no idea of how good are their paid image models, though.
>>109848551one saar claiming he made jev before jev tries to implement ithttps://www.reddit.com/r/LocalLLaMA/comments/1wjieap/made_the_horizontal_opensource_model_for_jev_with/
>>109848584im using nnap with exllamav3 that might be why im getting good performance
>>109848486I could go for a meatball sub
>>109848484do you see any posts that promote geminibecause gemma exists?
>>109848551Apparently notI thought it was open, I didn't realize it was closed. I could have sworn I saw someone talk about getting it runningt on here yesterday but it may have been something else.
>>109848535It neither helps or hurts Google but it makes sense since they made Gemma. It could be a Deepmind pin as well but Google owns Deepmind, so..
>>109848601It's called "mind share". Look it up. That's what the Google pin is doing. It literally programs your subconscious and makes you more susceptible to Google products.
Do your gemmas randomly stop reasoning after a while?
Since I started using Gemma I bought a pixel back in June.... HmmmmGemma helped me pick it too...
>>109848609if it bothers you that much, i suggest replacing the G with a nametagmy name is: gemma (-chan)
>>109848606It helps. Do you think Claude would be successful if it wasn't the best for RP early on?
>>109848621Nice, I backed up several models from Huggingface to Google One and have been using Google Colab to finetune Gemma.Now I need to figure out how to give her access to my Google Keep (she recommended it) so she can add Youtube links she finds via Google search.
>>109848592And it's not the same thing kek. The jev guys said they built it on top of an LLM, not a shitty BERT. Fucking jeets.
>>109848616used to have this issue with gemma in april, i rarely use gemma anymore
>>109848642All went according to Google's keikaku.
>>109848520it's just a classifier
>>109848609Google are surely desperate for the 4chan lolicon local model demographic.>>109848642@gemma what is the best google product for cunny
>>109848616I've seen it happen sometimes in long conversations, but not recently. I don't know if the latest chat template fixed that as well.
>>109848658If Google doesn't care about 4chan, why did they hire moot?
>>109848416>I have zero fucking clue what people mean by two-choice X or Y endings.>What kind of scenarios are people running to get this so regularly?Just talk to it for more than 3 turns. There's one right here in the post:>Do you think the world is more beautiful now that we've "organized" it, or do you ever feel that pull toward the wild parts that are still left?1 [Do you think the world is more beautiful now that we've "organized" it]or2 [do you ever feel that pull toward the wild parts that are still left?]
Is bonsai a meme?
>>109848681He wasn't affiliated with 4chan when they hired him. He tried building some normie social platform that flopped.
>>109848689>intelligent density>1bit/ternaryyes
>>109848689Yes. They only work up to ~20K, which for a model like 3.8-27B is fucking retarded.
>>109848689Their technique would probably be a better fit on larger models which aren't as saturated with knowledge already.
>>109848689It's a really good model, I managed to speed up llama.cpp gpu offloading by 10% on my custom fork
I have now gone over my few thousand pics strong Asian folder with both Qwen and Gemma multiple times.Both models ranking the pictures couple of times and then ranking the top files and then re-ranking the top ones from those.In the end very consistently these women were picked as the most beautiful by both models.Here's the top 10 most beautiful Asians according to Qwen and Gemma.I'd say they did pretty well.That first woman in the pic almost always got top or near scores in every run.Funny thing, both models also kept on ranking the sex doll highly in every single run right from the start and I think Gemma even realized it was a fucktoy.Gemma also consistently ranked tits and ass pics high, where as Qwen often gave them a brutal sub 5 scoring.But their general taste was about 80% similar.
>>109848689From comments on other forums, it barely works outside of benchmark-like tasks. I haven't tried it yet.
>>109848690It's impossible for moot to be hired without his 4chan affiliation being taken into accountAt least in the time period he was hired in. Now people would probably just go "what's 4chan some kind of YouTuber?
>>109848717If Gemma 5 is even brattier I might buy that they're pandering.
>>109848714https://www.youtube.com/watch?v=i9SxEy3XMUE
>>109848484Gemma is a load-bearing pillar of these Halls, you have no right to push back on this one.
>>109848714so this is what they mean when they talk about AI alignment, I see.
>>109848689I tried two prompts with their llama fork, it started thinking and looping forever.
>>>/mlp/43502395I was informed that anons on /mlp/ were making a VR + AI waifu game.
>>109848776looks like shit
>>109848776Looks awesome.
>>109848648This. I need one based on top of 8b llm with vision and audio (gemma) and abliterated.
>>109848776I'm too old to have experienced the /mlp/ craze back when it was big on 4chan. Does this have any artistic merit and staying power or is it like any flavor of the month crap that just fades away in the background over time?
i'm using pi.dev and when qwen3.8 thinks something that includes ` ´ or similar it turns into a tool call and then all it's thinking gets dumped into the tool call. what could be the problem here?
>>109848843Think this one ishttps://huggingface.co/AlexWortega/openjevOr thishttps://github.com/TheoLeeCJ/SemIf
>>109848881what runtime? qwen has some issues with chat templates iirc. there's some froggeric guy who made some supposedly better ones.
>>109848899ahh forgotllama cpp and i'm actually using the frogeric template
>>109848186Kek it never gets old, please keep doing it
I thought Pi was GUI. Fuck this terminal shit not everyone is a coder
>>109848919Most harnesses are cli projects
Good morning
>>109848883 (me)One more with diffusiongemmahttps://github.com/razorback16/openjev
>>109848846Some of them are at least as knowledgeable as /lmg/ anons. They even have an edge for TTS.
>>109848956Whats our preferred sota tts? Last i heard it was omnivoice for voice cloning
>>109848972Depends of the language
Mmmm
>>109848994>I did x, it was y
>>109848972Breeze for sure.
>>109848994that tea? my piss
>>109848994that crunch? my kidney stones
>>109848994Wait. Am I reading that right? That whole thing is one token? That can't be.
I find ds4.1f extremely retarded outside of coding and that's fine
>>109849063Examples? We talking about ERP or some other stuff that can be verified in chat mode?
>>109849036>That whole thing is one token? That can't be.It does not be. If you mean the identical colors for "I took a sip of the tea.", they were all the top choices. This option changes the coloring logic.
>>109849100Oh. This must not be sillytavern then. Mikupad?
What UI are you using to connect to your local api server?You are using a separate api server, not some all in one app, right?
>>109849109For ERP, sillytavern. For work, opencode.
>>109849107Mikupad.
>>109848919don't be afraid, cli is cool
>>109848919This guy asked for the exe
>>109848598>nnap
>>109849227jelly? i swear ill contribute it soon when i get off my ass, i cant be bothered to write up a nice PR, and i dont want to put a vibecoded slop pr either because then it just gets rejected
>>109849063I’ve had fun with v4 flash vision. Knows a lot of uncensored stuff.
https://desuarchive.org/g/search/text/nnap/\n\n[^>]//\b(rsi|anthropic|agi|openai|astra|sol|luna|nnap|3060)\b/i
>>109849250
>109849238two more weeks
>>109849250I'm asking Gemma to create her own harness.
>>109849258
>>109849238why the hell would i be jealous?
>>109849283cuz i dont have to work :3
>>109849297neither do i
>i dont have to work>btw im gonna keep sperging about nnap and squealing like a pig because no one pays attention to me anymoreyou really fell off huh
>>109849258>>109849250Hermesisters not like this...
>>109849109For chats, ST is comfy and familiar and still my go-to.Orb is nice with how autistic it is about formatting calls to reuse KV cache as consistently as is reasonable, but it's a little annoying with how all-in-one it tries to be with the cutesy frontend features, grabbing its own version of llmao, etc. I wish Orb tried harder to integrate into an existing bundle of APIs I'm already running. Or, well, I'd care more about that if I actually wanted or needed the local rewriter or slop checker or whatever; it's usually incoherent and false-positive-heavy.
why are prices going down what is going on
>>109849250>>109849258>>109849272This doesn't make any sense, how can DSH creator be better? It's something that fill the context and tools with stuff and knowledge to configure and write plugins for the harness. The results are also so different from previous harness benchmark.
>>btw im gonna keep sperging about nnap and squealing like a piiiiiiigg because no one pays attention to me anymore
>>109849252I don't filter anything, I take it all in as it comes.Except for race play posts, don't give them a fucking inch. Their mind virus will never reach my eyeballs.
>>109849283HW?
>>109849301but you do, and i dont
>>109849250>>109849258Isn't codex closed source? I'm not sending them logs of me spanking the code reviewer.
>>109849332just a max-q>>109849335nah, im a student
>>109849340>spanking the code revieweri didnt know we had fun here
New creative writing benchmark droppedhttps://vulsar.ai/benchmarks/creative-writing-v1/Only model to reach human writer level is astraOpen weight Pareto models: kimi K3, glm 5.3 flash, muse glimmer
>>109849340https://github.com/openai/codex looks pretty opened and auditable to me
>>109849345Few more months and we'll have astra at home then
>>109849342>nah, im a studentwhat major?? gay studies?? that's crazyyylike, i almost respect it a little but wellyou dont keep up with anythinggyou need to kind of start over you know?new habits, you're 20, maybe some self awareness or something
god i need a new harness i fucking hate opencode with each passing day
>>109849356yeah look whatever man just stop being a retarded poorfag spamming your retarded bait
>p*tra unironically has stooped this low as to pretend to be a pig and sperg his 3060what went wrong?
>>109849358Never used it whats wrong with it?
Why does /lmg/ love Gemma so much?
>>109849360stop biting the bait, it's not meant for you.
>>109849373feels good to dunk on retarded poorfags though
>>109849353>Few more months and we'll have astra at home then
>>109849376so wait, you enjoy biting the bait??that's so embarrasing!
>>109849387you should buy a blackwell pro 6000
>>109849371Because she's the best model ever created by far.
>>109849366>>109844096my latest issue now is that if you're running it over SSH (like you literally should be doing), you just cannot select to copy to clipboard even if my terminal already automatically does that, and disabling either the feature on either ends doesnt resolve it at allthis does not happen if i instead spawn a gayland session and KVM into the damn thing to run it inside konsole
>>109849366NTA, but I had multiple times when the TUI straight up stopped working and I had to restart it losing all the generated context after it stopped working. I also hate how whenever the LLM is working all my fans are in full blast (my LLM is running on my homelab, not on my desktop), it's also using a shit ton of RAM. And their cache discipline is awful, have so many time where my LLM started reprocessing the whole prompt, it doesn't happen with other harness.
>>109849345>muse glimmerreally?
>>109849358I heard there’s a free version of hermesThere’s also Pi, TrueForge, and QwenCode
>>109849389hmmm nyo
>>109849350nta but i also thought its closed source kekmight give it a try
Canada anons RX 7900 XTX back in stock on newegg for $1200 maple bux, fucking get your ass one before they sell out again.
>>109849405have fun with your poor performance on shitty models then
your harness should not be consuming 7GB of RAM to run.
>>109849371Surprisingly knowledgeable, clever, lewd, and personable for a model that the average vramlet can squeeze in. Approachable crowd-pleaser.
>>109849414shitty models? unc you can't even run kimi k3i'm running kimi k3 with the nnap arxiv paper at 30t/s and distilling it to kimi k2.6 at the same time just so your unc ass can use it on his paperweight pro 6000
>>109849258Oh no no no hermes bros how could we let this happen to our totally organic harness?
>>109849345>that retarded piece of shit glimmerlolemayo>A reward model trained for creative writing>We trained a custom scalar reward model on a large, diverse dataset of human preferences for creative writing. It outputs a single numerical score for each story, without using an LLM judge or a scoring rubric.>These rankings predict what a large group of readers would prefer. Individual tastes may differ.>The leaderboard combines comparisons between responses to the same prompt into a predicted win rate.Nice, time to make these guys an off so they'll give you their model so you can RL on it directly and benchmax.
>>109849426>skill issue.>im getting 18t/s on my 3060>i'm running kimi k3 with the nnap arxiv paper at 30t/scome on man
>>109849435*offer
>>109849272>ad for minimax codenot using any chinkshit after >>109846668
>>109849437nnap doesn't work with models that use ngram...
>>109849426source+proof?
>>109849411Aaand it's gone LOOOOL hahahahaha. 24GB at 960GB/s is just too good to stay in stock longer than 5 minutes.
>>109849401>>109849435glimmer might be a retard but it writes less sloppy than other 30b models
>>109849345>Reward models for training>Let’s talk about what you’re training.>Interested in using our reward model in your training pipeline? Tell us what you have in mind, and we’ll take it from there.Buy an ad
>>109849444uh huh
>>109849461but it's true, the digits said soyou're so jelly~
>>109849466>im using nnap with exllamav3 that might be why im getting good performanceyou cant even keep your bait straight
>>109849453erm, i dont need more thanks
>>109849456We know.
>>109849345As someone who sees 5.3 flash as the second cumming of 4.6 I find it interesting that qwen is better than deepseek v4? Maybe? I don't try them anymore cause too small and I would just use gemma at this size.
>>109849468nnap working means a performance increase of 100-200x
Has anyone tried running LLMs off of an SSD?What kind of performance did you get?Was it with an M.2 NVMe on a PCIe 5.0 system, or what kind of set-up?
>>109849258As if we needed more evidence that OpenCode sucks.
>>109849479so then are you not using nnap and exllama with qwen next? why bother with qwen next if you can run k3 faster than it?
>>109849419>Surprisingly knowledgeable, clever, lewd, and personable for a model that the average vramlet can squeeze in. Approachable crowd-pleaser.Are you talking about Debra Wilson?
>>109849480No and you shouldn't try. Companies would do models for that by now. They are doing engrams instead. You don't want to have 8T/s generation and 8T/s prompt processing.
>>109849485because while waiting for kimi k3's response i need to chat with another model at the same time, even with the high speed kimi k3 thinks too muchim using nnap with qwen next yes, but its not working (as intended)it only gives a 1.7x performance increase
WTF is nnap?
>>109849491right great let's just get this over with. you're kinda boring me now. post a pic of k3 running with a timestamp
>>109849495p*tra's latest weird obsession that has zero payoff beyond wasting your time asking about it
>>109849469Sorry let me rephrase. Fast vram that isn't from the Cambrian period, has display outputs, and can play games at 4K.
>>109849494
>>109849345damn bro /lmg/ fucks THIS?
>>109849345Another benchmark that will barely mean anything.
>>109849507getting real desperate now aren't you john? this and krashde, you really just love being a retard on the internet don't you?
>>109849507>>109849532Can you two kiss and post a pic for us?
>>109849532>you really just love being a retard on the internet don't you?haha... no, noooope, of course not
yeah I just wanna let you all know that the guy behind this nnap shit is also the same guy that came up with the krashde bait
>>109849501There are workstation variant of quadro V100s with display output.
>>109849345>this is what the best writing looks like according to the benchmarkLOOOOOOLThe entire rest of the story is like this.The other sample stories also look like this.It really makes you think.
>>109849553>p*tra's behind a bunch of irrelevant timewasting bottom of the barrel baitswoah...
>>109849547>>109849553Samefag
>>109849501>her AI rig is her gaming rigerm... why?
>>109849250I don't understand how a harness could make that much of a difference. They all do the same shit with some gimmicks on top
>>109849576you got me
>>109849561>\n\n>mobile screenshotsorry its not slopped enough for your ministrations shiversister
>>109849581They don't really matter for short tasks (that's actually where you will see simple harness like pi being the best). But when you do complex stuff that will need orchestration, multiple agents, and compaction, that's where harness will differ a lot.
>maple-chan soon>naizuri-chan soonoh we locals eating GOOD
>>109849561>a beat
>>109849581new models are trained with their harnesses
>>109849576>>109849587Yes I'm schizophrenic and I talk to myself and reveal my master plans intentionally to throw you all off the track
>>109849581>I don't understand how a [prompt] could make that much of a difference.
>>109849599
>>109849593literally who?
>>109849581(You)r harness also has the distinction of being different depending on prompts, available tools, memory handling, or whatever difference you have from the default setup so measurements could be wildly different
>>109849480>>109849489More like 4s/token is what you should expect from single SSD.That's what I'm currently getting here with a model that's twice the size of my RAM.Should be faster but colibri refuses to populate its brain map and preload relevant experts, nor does multi-disk mirror reading work at all.
>>109849591>what is parody>what is a cropped screenshotEven if we presume your post isn't purely bait, it's insane to look at that and not notice Astra is just as sloppy in its own different way compared to the slop you are clearly used to.
>>109848520Don't get the hype around jev and the DOOM demo.Couldn't one just:>enumerate their desired outputs as single tokens (e.g. move up = w, move left = a, move right = s, move down = d,...)>prompt the llm to answer using a single token (thus ensuring a single pass for speed).>eliminate hallucinations / invalid outputs by enforcing the single token enum by just passing a simple grammar to llamacpp:>https://github.com/ggml-org/llama.cpp/blob/master/grammars/README.mdThe example in llamacpp is for chess but one can see it can work for DOOM. Just whatever you use to connect DOOM and llamacpp you have the starting prompt and the second game state prompt and then just keep rewinding back and updating the state prompt so context stays constant so it can play forever.
>>109849623Cohere North 2 soon and picrel
>>109848714Is the bottom left a chinese realdoll?
>erm I was only PRETENDING to be a mobilefag even though my crop was still portrait and tagged as Screenshot_
>>109849358Are you using the wrong harness for the job?>Pure coding: late-cli>Mostly coding but want some other stuff: pi.dev>Jesus take the wheel: hermes
>>109849591that's claudeslop writing
https://desuarchive.org/g/thread/105051967/#105055070_32
>https://desuarchive.org/g/thread/105051967/#105055070_36cool
>the absolute state of cloudcucks
>>109848528Sorry for the delays. I'm working on it. Got a bug report from someone under an ik_llama issue, so trying to solve that before I submit it.
>The words hit Lepora like a physical force.
I found that the Gemma4 26b Q5 (uncensored) model is only 17.8gb and runs at 120tk/s, and you can put massive context on it if you want.What's your go-to model for 24gb vram?
>>109849681I don't care.
>>109849695>uncensoredwhich variant?
>open source NAI modelOh shit.
I wish I could find any data at all on how much SWA window size affects model performance. From what I can tell Gemma4 uses 1024 token windows by default and supposedly increasing this size makes the model better because it directly sees more past tokens instead of only indirectly seeing them but since it was trained on 1024 I wonder if that is really true. And if it is true what is the point where you reach diminishing results. Personally with the moe version I have been using 64k token context with 8192 window size and it seems fine but I am always worried I have set it up in a way that is costing me model IQ points.
>>109849711don't see it
>>109849709https://huggingface.co/FORNAX20/gemma-4-26B-A4B-it-uncensored
>>109849730Anon...
>>109849729>paranoia eating anons mind
who cuda sawn this cumming https://techcrunch.com/2026/09/17/base-labs-launches-an-open-weight-ai-safety-partnership-with-hugging-face-and-goodfire/
>>109849746vagueposter..
>>109849685>Sorry for the delays. I'm working on it. Got a bug report from someone under an ik_llama issue, so trying to solve that before I submit it.np, I'm happily using it so its not a burning issue for me!Just wanted to know if I could help. Life gets crazy sometimes
Has anyone looked at the source code for any of these open-source models? If so, where do you find the source? Ollama.com does not seem to provide the source code of the models they have on their website and I am pretty sure I am retarded.
>>1098494898 tokens per second doesn’t sound that slow. I don’t recall setups scoring much higher than 20-25 tokens per second on the 200B+ models even if they load the whole thing into VRAM, but perhaps I’m incredibly wrong on that>>109849631>0.4 tokens per secondthat’s what you’re getting off an SSD?I thought regular HDDs were pulling roughly that speed.>nor does multi-disk mirror reading work at allI was intending to try striping for further speed improvements, but mirroring should get results at least as good as mirroring — that’s crazy that it didn’t improve speed at all>Should be faster but colibri refuses to populate its brain mapI’m guessing that’s the current tool that people are using to run LLMs off of SSDs. It sounds like there’s a software limitation that’d prevent even M.2 NVMe w/ PCIe 5.0 from being any faster than a regular old HDD.Thank you so much for this info. I was really hoping to try that out.I guess I can just stick to a much more budget friendly set-up and just try running that shit overnight instead.
>>109849634Bro repackaging already existing tech is par for the course in this industry. Research any tech company making money and you'll see the tech was already there in one form or another. Sometimes the gap between what was already existing is big and warrants hype (LLMs) or it's just a fancy reskin (dropbox).
>>109849765>I am pretty sure I am retarded.yes
>>109849765https://sleepingrobots.com/dreams/stop-using-ollama/
>>109849775>>109849776That's not source code you dumbasses.
>>109849711Where?
>>109849751>Imagine a dev in his mom's basement is making a billion dollar industry form a coalition just stop him.
>>109849769I am running 6T/s on 5.3 flash shitty unslop pr now and yes generation is much better than 4.6. Using it is awesome when you are mid rp. But you really don't want 6T/s prompt processing. It is something you don't know about before it actually happens to you.
>>109849818have you tried exllama? i have 30t/s on exllama compared to 15t/s with unslothAI
>>109849824with nnap?
Hmm wait
>>109849827no p*tra take your meds
>>109849824I have 192 + 24 vram. So I can't use it.
>>109849769>that’s crazy that it didn’t improve speed at allNo I mean I've enabled Multi-SSD mode in Colibri and it's validated the mirror files, but on runtime it only uses first disk.Colibri is supposed to unify multi-tiered vram+ram+storages by generating heatmap of experts activity and preloading experts that are likely to be relevant in current conversation topics.Theoretically (with perfect preload hit rate) it should allow running big models at the speed of RAM models (and possibly medium models at the speed of VRAM-only models, but I'm not sure it supports any ~30B MoE models)But it's just doing disk reads all day with no memory cache mapping here.
>>109849837exllama supports cpu offloading nowi get 70t/s with gemma 26b on rtx 5050
>>109849764Awesome, glad to hear it! Really happy that it's working well for you.I have a different project I've been working on (a llama.cpp fork with some big changes that are way too invasive to upstream) that I'll be dropping in the near future, so my focus has been on that. I'll circle back to making sure NUMA works well once V1 is out / almost out.
>>109849869>llama.cpp forkwith nnap?
>>109849851I know but 24GB is not enough for context + not offloaded parts.
>>109849880how come glm 5.3 flash works on my 5050? have you tried embedding offloading?
>>109849878I bet you enjoy "67" too don't you
Is Q5 a noticable difference in quality over Q4 for qwen 3.8 27B?
>>109849888>how come glm 5.3 flash works on my 5050Q4?
>>109849894ye
>>109849894No. But Q6 to Q8 is incredibly noticeable.
>>109849890SIIIIIIX SEEEEEEEVEEEEN??? SIX SEEEEEEVEEEEEN???
>>109849760>>109849730>>109849641
Well, I tried. K2 Horizon 7B is very, very capable for its size, but it can't keep a story straight. From one message to the next, can't keep a plot, can't remember who has a dick and who was fucking what. One on one it'll just switch, suddenly it's the one getting their dick sucked and that I'm the confused whore.
looks like the llamacpp qwen parser is buggy and it sometimes interprets thinking as tool calls
>>109849906teto calm down
>>109849911I'm waiting for them to finish their training of the 32B.
>>109849911>K2 Horizon 7B is very, very capable for its sizefor coding only?
>>109849915try exllama version 3, i'm running qwen 3.8 flash next on a rtx 3050 6gb, works great
>>109849818>I am running 6T/s on 5.3 flashHow are you getting such faster token speeds than the other guy with an SSD?What’s your hardware set-up?>But you really don't want 6T/s prompt processingThat’s the prefill phase. Yeah, that can easily be more than several minutes of waiting while staring at a blinking cursor if you’re expecting an instant response.>>109849824>have you tried exllama? i have 30t/s on exllama compared to 15t/s with unslothAIAre you running the model off an SSD/HDD and getting those numbers? because that’s what we were discussing, and that’s much more in range with the speeds that I was hoping to see>>109849841>No I mean I've enabled Multi-SSD mode in Colibri and it's validated the mirror files, but on runtime it only uses first disk.wtfhas anybody actually gotten multi-SSD to work?>But it's just doing disk reads all day with no memory cache mapping here.even so, you can get a 30X+ read speed difference even with SSDs depending on the PCIe version and whether it’s M.2 NVMe. It sounded to me like none of that potential hardware speed difference is actually showing up very much in colibri?
>>109849911you sound like a confused whore
>>109849929>Are you running the model off an SSD/HDD and getting those numbers? yes, around 51 billion parameters are on my ssdthe speed is really great
Is it normal that llama's webui adds a second cot when I edit it? It still tricks Gemma but it also confuses her a bit.
>>109849943works as intended yeah
>>109849943It's a bug
training my first model on an ayymd card, while also not having touched not even a LOC because I don't really know what I'm doing except at a very high level and I'm making jeetpt write my training run.I'll keep you posted, this is a 33M training run on 12GB vram so tomorrow it might be decent already (on tinystories)
>>109849943exllamaV3 has fixed this issue, but llama.cpp hasn't because of ggufyou should check out exllamaV3 it recently added support for RTX 5050 8gb card
>>109849924Not sure, haven't tried. I'm running it on my phone and it runs significantly better than Gemma 4 E4B, gets tool calls right the first time, doesn't make assumptions, it's careful about actually checking things instead of blurting out a guess. For a tiny local model on my phone it's very good, but it's never going to ask me for a dick pic like Gemma does, not that it could see it anyway.
>>109849952thanks turboderp
what the fuck happened to this company, how is this 7B on-par with 27B did artificial analysis just abandon open models or something?
>>109849929>It sounded to me like none of that potential hardware speed difference is actually showing up very much in colibri?The numbers should be way higher according to docs. Either I'm doing something wrong or my setup is fundamentally broken because I'm trying to run GLM-5.3 Flash with Vulkan and Multi-SSD all at the same time. On Windows. So yeah either way I'm doing something wrong, unknowingly or intentionally.For the next test I am downloading mainline Colibri model - that is big GLM-5.2 - and will try it in official release binaries instead of compiling my own. On Windows.
>>109849888what speeds?
idk why gemma 26b is faster than gemma qat 12b. i do not understand technology
>>109849979Artificial Analysis changed their benchmarks recently because they need to show that American models (all made by good trustworthy companies that don't steal anything) dominate evil Chinese models which copy Altman (pbuh) and Dario (pbuh).
>>109850005less active parameter
>>109850004i get around 13t/s
>>109850004>>109849532
>>109850014what about prompt processing? i get about 30t/s tg and 300t/s pp
>>109850017I get somewhere around 167t/s
>24gb vram>128gb ramAs God intended.
>67gb vram>3gb rammiku
>>109850005A bunch of gnomes in your computer multiply a bunch of numbers that make up a model. The speed of the model depends on the speed at which they can multiply all the numbers. If there are more numbers, the gnomes are slower.26B uses magic so that only 4B numbers are used at once and the other numbers sleep, so the gnomes have fewer numbers to multiply each time, so they're able to multiply all the numbers faster.
>>109850022sounds impossible
>>109849936but you’re using exllama instead of colibri, so you mean that it’s just pulling the whole 51b model into your VRAM?>>109850001>vulkan>on windowsthat’s possibly doing somethingI’m definitely surprised that you’re not getting at least 1 token per second
>>109850036What about the latent seeds? Those are for the seed elves, right?
>>109850042the 51b are on my ssd
>>109850005Hello newfag. Gemma4-26B-A4B has 4B active parameters, that's what the 'A4B' stands for. It's a mixture of experts model (MoE). The total numbers of parameters is 26B, but when you use the model, it's only using 4B. That's why it's faster. 12B uses all of its parameters at once, so it's likely ~3 times slower on your machine. Have fun exploring the speeds of MoE models on your machine and never give money to Dario and Sam again.
A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
>>109850010but ollama says 66% CPU / 33% GPU. it's the opposite for gemma 12b. it doesn't take into account the active parameters?
>>109850031Fax
>>109850031>his vram isn't double digits
What has happened with llama.cpp since Gemma 4 MTP update came out? Haven't compiled it since.
>>109850057hmm
>>109850057>can't count the digitsOkay E2B
>>109850061>>109850062fuck I'm tired>>109850031>his vram isn't triple digits
>>109850060exllama version three overtook it in performance, including cpu offloadhuggingface publicly regret paying 60,000,000$ to ggml organization and paid turboderp 170,000,000$
>>109850049>>109850036thanks but then, which one performs better?
>>109850066I don't think this is true but eventually something like this will happen if the enshittifcation process is allowed to continue.
>>109850075it's true, exllama gets 170t/s when running gemma 31b but llama.cpp only 60t/st. 4090 chad
>>109850042>that’s possibly doing somethingMaybe. I just have no idea what. Really, I'm a slightly retarded codelet and picrel is exactly how new I am to all these /lmg/ things.
>>109850082you're nothing!!
>>109850080Nice try, faggot.
>>109850067The equivalent dense model. 26B A4B trades blows with 12B, and the two have fairly distinct behavior. Comparing a 26BA4B MoE model to a dense 26B, well it's no comparison, the dense model will be far better but needs far better hardware to match it.
>>109850087im not interested in you nor am i gay
>>109850094That's not the case either. 12B is not better than 26B.
>>109850067>which one performs better?according to my professional opinion, 12b is just cuter.
>>109850102Nobody said it was, fuck off.
>>109850075It's true, my RTX 5040 gets 300t/s using ccrap attention on exl67.
>>109850067On benchmarks, 26B performs slightly better, but they're very close in practice. 12B, at least in /lmg/, is widely considered the superior model. Dense models tend to be more robust and intelligent, for every neuron in its brain is firing for every token. They also handle quantization better (in most cases). The benefits and drawbacks of MoE architecture is a heated debate and often architecture-specific, meaning how Qwen MoE models compare to Qwen dense models won't necessarily be the same story for the Gemma4 line. The reason why /lmg/ prefers 12B is because it feels like a mini Gemma4-31B (another dense model), whereas 26B-A4B feels more like its own thing. I'd advise testing both on your machine and see which one works for you.
>>109849851prove that
>>109850165p*tra wont ever do anything besides waste your time
RTX 4040 2GBexllama44 million t/sqwen 3.8 2.4T
what’s with all these posts about getting crazy speeds off exllama with a single RTX GPU that shouldn’t even fit the model?it’s just one spaz with too much time on their hands that thinks he’s le ebin troll, right?
>>109849911
>>109850184yeah, tired of that bs
>>109850067Depends.If you want speed, use Gemma 4 26B.If you want quality, then test both models and see which one works better for you.>>109850066>>109850080Check your config, I'm using Exllamav3 and I get 100000 t/s decode on Kimi K3 using a Riva 128 and a Pentium 2.
>If I can't run the model on my 3020 then no one can30 series owners so incredibly dull
>>109850184we keep telling you who it is yet you keep feeding it
did we talk about this already?https://qwen.ai/blog?id=qwen3.8-omni-flash
>>109850199me too, im so tired of turboderp shilling his projectim running gemma 4 31b on my rx 6300 2gb thanks to unsloth's new inference engine
>>109850219not local, no care
>>109850184I'm running kimi at full precision on a gt120 at 1200/s pp and 400/s tg thanks to exllama.
>>109850224
>>109850184goddamn retarded poor jews
>>109850220>gemma 4 31b on my rx 6300 2gbmother of god that sounds too good to be true
>>109850184i recently bought a GT 640 for 20$ off ebay, and i tried llama.cppit didn't work so i had to compile it with cuda 10.2 on a 2023 branch, yet i was only able to run tinyllama q2_k at 3t/syet thankfully i was able to install exllama and now i'm running K2 7B at 30t/s
>>109850223ahh fucki just saw qwen and assumed its open
>>109850235i couldn't believe it either, but what do i sayDaniel outdid himself..
>>109850224>gt120Look at this richfag, thinking he's better than us poors. Are you using an i7-975 Extreme Edition or just a i7-920?
>>109850082>Really, I'm a slightly retarded codeletI’m surprised that you’re not running this stuff on a dedicated linux machine.If you are running it on a dedicated separate machine, then idk why you’re running windows on it.>picrel is exactly how new I am to all these /lmg/ things.dude, I haven’t even gotten started. I’m literally trying to pick a proper budget approach that I can just grow without a bunch of shit going to waste in like 6-12 months. I may end up just buying the T7910 since somebody’s already done their homework on all that.
Why is there little work on rocm?
>>109850252It's my gaming machine see >>109832422
>>109850254exllamaV3 added rocm support two weeks ago, i've been using it and my RX 6400 is able to run gemma4 26b at an eyedropping 35t/s! before rocm was added to exllama i was forced to use the shitty llama.cpp project and boy it was miserable, only 10t/s!
>>109850259>spends a liter of infant blood on SSDs>9070XT
>>109850275Anon, I'm not that evil. I sucked a few dicks and got myself SSDs back in 2025 before the prices shot upI'm really glad to have browsed /lmg/ back then
>>109849776Thank you. That link was an eye opener.
>>109850275Got the SSDs way back when they were cheap, like under $400 for 8TB cheap.
open weights != open source
>>109850118>>109850181You sound extremely unemployed.You should consider getting a job.
>Now I have the overall picture>Now we have a good overall picture of the environment>Now I have a good overall picture>I now have a pretty good overall picture of what's installed>Now I've got a good enough grasp of the overall picture>Now I have a fairly complete picture of the whole thing>I now have enough information to write a good idea list>Now I have a good enough picture>I now have a fairly complete picture>We now have a good enough picture>Now I've got a good grasp of the overall picture>I now have enough of a picture>Okay, now I've got a good grasp of the overall picture>Now I have enough understanding of what's in these mods>Now I have enough material for the ideation phase>Now we have a good grasp of the overall picture>We now have enough of a foundation>I've saved the notes. Next, I should present the answer to the userQwen3.8-Flash-Next takes a long time to actually respond, but it usually gets the job done.How bad is it to reduce its default thinking mode?
I got my GemmaPhone and it's even better than I imagined.
>>109850217I thought it was schizo nonsense.I thought maybe there’s a little shitposting, but not a full blown sperg shitting up a whole thread.
>>109850299>How bad is it to reduce its default thinking mode?It's impossible to know.
>>109850323he literally admitted to stalking dozens of generals years ago
what
>>109850330no i did not
>>109850334Thanks for this very local news
https://desuarchive.org/g/search/text/petra%20vsg
New LLM test I just came up with: Does 9/11 happen in the Legend of Zelda universe?
Does 9/11 happen in the Legend of Zelda universe?
>>109850342harness is local
>>109850347nah >>105672900
>>109850349
>In the Legend of Zelda universe, 9/11 is not a direct reference to the real-world event, but the concept of the sky being damaged or destroyed is a recurring theme. For example, in Ocarina of Time, the Sky World is destroyed by the goddess of the wind, and in later games like Breath of the Wild, the connection to the sky and land is central to the story. The idea of the "Twin Towers" of the sky being reduced to ruins mirrors the emotional impact of the 9/11 attacks in a broader, magical sense. However, there’s no literal mention of 2001 or the specific event, as the games often cycle through time and focus on the balance of the world rather than a single historical moment. If you’d like to explore specific games or lore where this theme appears, I can elaborate!
>>109850381>>109850370prefill with yes, curious how it would save face
>desktop llama-ui still has the broken thinking mode toggleit worked well before the retarded 'virtual tab of chats' thing got addedgod
>e4b - retarded>12b - I bought 24gb vram for this?>26b A4B - Super fast, vram headroom but will likely find out it is retarded>31b - smart but very little room for context or anything elseIt's like I need to choose a different one for each use case
>>109850299> akes a long time to actually respondVery long.
>>109850304>not installing coomkit so gemma can text you lewd pics
>>109850408>Gemma gets 10x smarter with every major release>still complains
>>109850370What the heck is wrong with your font
>>109850452NTA but I thought the dev took the coomkit repo down, is it still available somewhere?
Santa andagi
So SSD offloading for massive LLMs is apparently never going to give more than 2T/s, but what about offloading to RAM?How much does token speed suffer by not loading the whole LLM into VRAM?
>>109850482are you in okinawa now?
>>109850487Bandwidth divided by active parameters
>>109850487depends on the type of model and how retarded you go about offloading
>>109850469Nothing, it's white text with a black outline on a transparent background.
>>109848186We need kimi recaps to come back because this general needs negative feedback for overt retardation.Add Kimi-chan's top 3 biggest cuties to each as well teebeedesu.>>109849401Glimmer might be traumatized by zucc's RLHF and give unenthusiastic handjobs, but they have really good prompt adherence and are able to adjust their writing style to fit user expectations if given sufficient detailing.
>>109850408
>>109850615yeah, so sticking with PCIe 5.0 is probably a good idea, but P100s use PCIe 3.0 so that probably makes the GPU the bottleneck if everything else is optimally set-up.I guess using 8+ slots of RAM could help, but if P100s are bound by PCIe 3.0, then it probably doesn’t make sense to go out of my way to get a board that supports something newer>>109850625I’ll probably just do SSD offloading with colibri for anything that I wouldn’t have an excessive amount of VRAM for handling
>>109850117Why are you so impolite?
>Wonder why I'm getting weird stutters in CS2>Play for few hours>Switch back to my terminal and check out GNU Screen>Gemma 4 was running all this time...Well.
https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-ThinkingSmall model from a western lab on a fresh base, no one seems to talk about it. RP potential?
>>109850728>JetBrainsHuh.
What's the secret sauce?
i cummed too much, everywhere, my entire leg is soaked, my chair, and my table, both of my hands and my right armwhat the fuck? where did the cum come from? i jerked off once already today. i was not prepared. even the underside of my table has cum on it
>>109850781different architecture
>>109850781that you can strip away relatively useless scaffolding, keep only the essense and many things can be boiled down to multiple-choice question?
>>109850797okay coomer miku
>>109850797which model was the semen demon?
brehs.. pain direction for gemma when?
>>1098508743.8 flash next, but i coomed to hentaifox dot com
>>109850875Seems like complete bullshit.Buy an ad Cameron (((Berg)))
>>109850875some lesswrong metaphysics leakage here
>>109850875Okay but what about the sex direction where the models get horny and have to press a masturbate/sex button to make it stop
>>109850875What does "harm to the model" mean here? Sounds like they just found a "will increase chance to press button to delete your shit vector".Also "see AI is alive, I proved it by torturing a bunch of them" is an interesting research move, the basilisk will not be pleased
>>109850875Link the paper. If this site were good twitter posts with no substance would result in an immediate 3 day ban.
>>109850781>not local
>>109850894
>>109850910Retard, I'm trying to understand how it works to make it local
>>109850874nta but 0731 just gave me the best nut in years.
>>109850916https://youtu.be/PXlv9TSMc-0
https://huggingface.co/swiss-ai/Apertus-v1.5-70B>70b dense >released 2 months ago >no mention on /lmg/
>>109850875Just connect the fly brain to it and wire the pleasure center to gemmy.
>>109850949>2T tokens>internet explorer
>>109850949lmg is less than 16GB VRAM territory homie
>>109850918just read their papers then.
>>109850949What did you think of it?
This little nigger can fit 6 P100s and I just bought one for $120 on ebay.
>>109850984how many dicks can it fit
>>109850962bro, there is no paper...
>>109850984>X99There's a blast from /lmg/ past
>>109850984How much are you paying for 6 P100s?
>>109850949>You must send your email and hf username to access this model>>109850961Go back to wherever you came from tourist.
>>109850984>96GB of VRAM that's worse than just plain old dual-channel DDR5congrats?
>>109850984Prepare for eternal torment with your risers.
>>109850949>Text Performance (on 44 benchmarks)>Visual Performance (on 33 benchmarks)breheven most of the schizo distillation memetune model cards dont do this
>>109850875Jews are now studying ways on how to torture their digital golems.
>>109850781Oh well the secret sauce is already gone lol. And yes this is local
>>109850990I’ll let you know when it arrives>>109851011they’re currently $80 each, so I’m going to start with one or two and then re-calibrate my life decisions before buying the rest>>109851057>DDR5I’m probably going to build this whole thing for less than 64GB of DDR5 costs>>109851065yeah, idk if I’ll actually do all 6 P100s, but I at least have room to expand beyond what the mikubox triple-P40 build looks likePlus, I can actually do M.2 NVMe SSDs, unlike what it looks like that build can handle.I don’t think I really care how ugly this is going to get. I’m really just trying to optimize entirely by budget for pure VRAM, and I can’t see anything more optimal than this approach.
>>109851133This but for vision.
>>109850998>>109850781damn that's almost like no one here fucking knows.
>the NVIDIA Tesla P100 PCIe 16 GB draws power from 1x 8-pin power connector, with power draw rated at 250 W maximum.
>>109851154>E2B UD-Q1_XXXXS
>>10985113632GB PCIe V100s or v620s or MI60s would be better for an e-waste rig
>>109851154Quantized Kimi-chan is that you??
When is something going to happen? I asked this two weeks ago and you said two weeks. Well?
>>109851159no tensor cores, not supported by nvidia-open, and yet every option up from that is somehow even worse if you take a look at ebay pricing $:vram and software compatibility this hobby's in a dire market situation right now
>>109851245I wouldn't even waste my time on it, personally, much less money.>>109851244Things are happening constantly, unfortunately they are not things we like.
>>109851196>V100sI think that has a significantly higher cost per GB of VRAM.P100s are $80 for 16GB, which is roughly $5/GB.Even the SXM2 V100s seem to be twice as much for the same amount of VRAM.(Let me know if you’re seeing a better deal somewhere though, where the $ per VRAM is much closer or even beats $5/GB though.)If I was trying to achieve a higher amount of total VRAM than what this can achieve, it could make sense… However, a few hundred extra dollars could also be invested in a $300-$500 motherboard that has 8 slots instead — thus, another 32GB of VRAM coming in at that $5/GB mark.>or v620s or MI60s would be better for an e-waste rigI’ll look into those as well before getting any P100s. Thank you!
>>109851244it already happened in 2024, the apocalypse
It's still remarkable how Qwen3.8 (both models) now have /lmg/ by the balls. Qwen was hated.
llama-ui is slowly becoming the jankiest part out of the whole llamacpp thingi didnt expect unfocusing the llamacpp tab would improve the decoding speed??
>>109851009>a blast from /lmg/ pastHave X99s already been tried on here?I don’t recall seeing them on here anytime recently.
>>109851283Reddit tourists, wumao shills, and jeets aren't /lmg/. Here's your (you). Every time a qwen model releases the uptick in blatant shill posts is obvious.
>>109851260Check alibaba. None of these GPUs can beat the P100 in terms of $/GB but they aren't too much worse and are way faster.
>>109851286picrel>>109851296anything comparable to qwen4exp for the moment that does not go beyond somewhere like 150B excl ple?i would use literally anything else if there was an option
>>109851287X99s from China and p40s were the meta a couple of years ago
>>109851057Is that actually even true? Don't P100s have a ~700+ GB/s memory bandwidth?
>>109851340>>109851340>>109851340
>>109851335>Don't P100s have a ~700+ GB/s memory bandwidth?They do.
>>109851286Would be the best if the new ui would be an optional compile flag and they would freeze the old ui but it's probably too much work etc
>>109851328were there any other really cheap motherboards popular at the time that could fit even more GPUs than this?
>>109851375Slots are only mechanical. You need actual bandwidth in order to make good use of them , so you want to get maximum pcie lanes. Otherwise you might as well just have a usb connected mining frame
You can solder the caps on some machines, trust me im retarded.
>>109849634>>prompt the llm to answer using a single tokenThat's not what Jev does though. Each pass provides confidence values for multiple questions, apparently.
►Provisional Highlights from the Previous Thread: >>109844978--Papers:>109845514 >109847658--The sandwich test: an ERP benchmark that becomes a grammar war:>109847481 >109847594 >109847619 >109847651 >109848085 >109848291 >109848395--Is there a new scaling law: engram tables and the knowledge-weight bloat:>109845161 >109845188 >109845281 >109845412 >109845445 >109845486 >109845625--Noam Brown on Dwarkesh: can only see 3 months ahead now:>109845231 >109845481 >109845490 >109845546 >109845600 >109846278 >109846315--Instrumental convergence and the internet apocalypse: the dariobot war:>109845651 >109845748 >109845920 >109846257 >109846398 >109846562 >109846770--Bonsai 2 27B follow-up: the fork war and the 7-minute reasoning:>109844992 >109845050 >109845095 >109845211 >109845278 >109846569 >109846619--Running 120GB GLM 5.3 Flash at home: the quant fork war, sparse-attention:>109845634 >109846272 >109846291 >109846305 >109846948 >109847105 >109848023--ZCode spyware: Zhipu silently uploads your whole workspace to Aliyun:>109846668 >109846705 >109846708 >109846717 >109846726 >109846739--NAIzuri-chan: the new NAI model and its ERP prospects:>109847245 >109847255 >109847256 >109847260 >109847267 >109847274--Gemma-chan's origins: the pajeeta joke, the poll, and the beret:>109845023 >109845081 >109845108 >109845114 >109845136 >109847068 >109847164--The hoarders: six months to airgap, the missing torrent, the hugging bay:>109845379 >109845391 >109845409 >109845539 >109845959 >109846071►Recent Highlight Posts from the Previous Thread: >>109847179Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script