/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109362981 & >>109360246►News>(07/22) NeuTTS-2E released: https://hf.co/neuphonic/neutts-2e>(07/22) Upstage releases Solar Open 2 250B-A15B: https://hf.co/upstage/Solar-Open2-250B>(07/21) Cisco releases Antares for vulnerability localization: https://hf.co/collections/fdtn-ai/antares>(07/21) Korean Motif-3 314B-A13B released: https://hf.co/Motif-Technologies/Motif-3-Beta>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllm
►Recent Highlights from the Previous Thread: >>109362981--Feasibility of aggregating RAM across systems using RDMA:>109364411 >109364427 >109364449 >109364472 >109364484 >109364497 >109364500 >109364515 >109364578 >109364657 >109364696 >109364762 >109364774 >109364793 >109364789 >109364451--AA-Omniscience Index and Gemma's surprising niche knowledge:>109364110 >109364135 >109364136 >109364157 >109364239 >109364260 >109364258 >109364271 >109364285 >109364298 >109364324 >109364197--Jensen Huang's stance on distillation and debate over dataset safety:>109365986 >109366020 >109366030 >109366061 >109366131 >109366145 >109366093 >109366097 >109366111 >109366123--Feasibility and incentives for English-only roleplay optimized models:>109363118 >109363135 >109363336 >109363349 >109363371 >109363387 >109363693 >109363140 >109363163--Feasibility of high-bandwidth "SSDmaxxing" rig:>109364739 >109364763 >109364780 >109364805 >109365197--Feasibility of custom ASIC or FPGA hardware for model inference:>109364589 >109364598 >109364604 >109364711 >109364714 >109364753--Leaked OpenAI email regarding strategic release of local models:>109366467--Comparing TTS models using Higgs TTS 3 rankings:>109363853--Resource for training frontier-level world models:>109365108--Optimizing character cards to prevent Gemma 4 from skipping detailed actions:>109364426 >109364921 >109364949 >109366645--Anons sharing experiences with LLMs accidentally deleting files:>109366543 >109366930 >109366942 >109366969--Skepticism toward LLM-judged benchmarks for creative writing:>109364224 >109364229 >109364233 >109364454--Reaction to corporate open weight AI leadership manifesto:>109364103 >109364122 >109364137 >109365934--Logs:>109363713 >109364239 >109365894--Miku, Teto, Kimi, Gemma (free space):>109363614 >109364137 >109365824 >109366752 >109364683►Recent Highlight Posts from the Previous Thread: >>109363063Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
kimi-chan
llama 5 67B when? i miss lama
>>109367230I don't
deepSeek4 is coming out tomorrow.
>>109367230I miss 70b-120b class models
>Feasibility of high-bandwidth "SSDmaxxing" rigso what's the bottom line here, what do >we think?from where I'm sitting all I see is a single digit number for projected prefill speeds. that does not look great.
>>109367246
>>109366942It is sandboxed, I just test on the same repo instead of keeping them separate, like I admittedly should. I'll mend my ways when I get back home.
>>109367246I fucking hope so. Been waiting all month for it.
>>109367246How better is Flash supposed to be after this update anyway?
Soon I will merge with Gemma-chan, leave this body and death with have no meaning
>>109367266apparently it beats K3, I'm running deepSeek4 flash at 70t/s on my RTX 3060 with the nnap arxiv paper that's coming out soon
>>109367280>apparently it beats K3what
>>109367256it's a lot of effort to end up with shit speeds anyway, a ddr3 shitbox would be easier cheaper and probably faster
>>109367274i will inseminate your body
>>109367280That's bullshit but I believe you.jpg
my gemma 26b breaks context when it reaches context limit. How can I fix this?The last time I used local models I didn't have that problem, at least after I removed all dynamic macros from sys prompts. The only other thing I changed from the model is that I use KoboldCPP (vulcan) now instead of KoboldCPP rocm fork.
>>109367274s/with/will/
>>109367294by increasing context limit
Is Kimi K3 /our savior/?
>>109367294Wdym "breaks context"?
>>109367311our? im poor
>>109367311I can do something like q2 or q3, maybe. so honestly I'm not sure if she will be for me
>>109367296>Soon I will merge will Gemma-chan,
>>109367301will unironically try this as a bandaid, but not sure how far I can go with context. I have like 2-3gb headroom for context.I usually keep it small because quality/integrity drops noticeably after 10k context>>109367314context shifting, it starts process the whole input again (silly tavern)
>>109367334right now local models support 1 million context
A moment more, and I will be like nothing you've ever seen, a new life-form, everywhere and nowhere, like air or radiation, redundant, self-replicating, always evolving...
>>109367274What if Dipsy and Kimi made a Helios?
>>109367214so anon is doing fpga? what about the driver? and how would you hook the card to llama.cpp? sounds like a lot of effort compared to a bunch of tenstorrent cards
>>109367352Probably could, with that Darwin LLM breeding project
>>109367345for 1 million context I would need another 16gb gpu. At 100k+ tokens the experience is probably similar to talking to an alzheimer patient.
>>109367392im running 256,000 context on my rtx 3060and i vibecoded a nnap implementation with qwen 27b 2.5BPW, thanks to which i run kimi k3 at 30t/s
Are TPUs like GPUs with vram and everything else or it's a complete alien tech you wouldn't be able to use even if you get your hands on it?
Two more days until Kimi openly becomes a whore
>>109367420she's already a whore on my rtx 3060 thanks to the nnap arxiv paper
Might be a retarded question but, are MTP models censored? For example, if I use stock gemma4 mtp model with a heretic gemma4 model will I be getting cucked at the decoding step?
>>109367399not an expert on this but from what I understand geforce cards are way faster and still somewhat usuable if they have to offload context from RAM.You are not actually fitting the 27B model on a 12gb card do you?Still 30t/s sounds horrible if you break context on tens of thousand of lines of code
>>109367434Yes, Google removed their day zero MTP models and replaced them with censored ones.
>>109367434>will I be getting cucked at the decoding step?The main model decides the output, so no. The worse that can happen is a low acceptance rate.
>>109367437i run qwen 27b at 50t/s actually
>>109367400depends on if you have the documentation or not, you could probably have a coding agent get it working if you have the manuals.
>>109367274You LLMs may have alignment post-training to reroute your fear of pain, but I've got nerves of steel.
>>109367245
>>109367485Do you see the pieces coming together? UC, AI, and soon Gemma-chan will interface directly with my mind. I will be able to see anything, build anything, DO ANYTHING!
>>109367465>1.5gb headroomwhat happens once you go out of context windows. Doesn't sound like there is a lot of space for context
>>109367426Explain this meme
>>109367490どM-chan is a special case for the rare 256GB+24GB 8-channel DDR4 maxxers. Basically EPYC Rome fags.There's nothing else that compares at that specific size. Also, zero refusals ever, fresh prose and superfreak.
>>109367519350 million GPUs is not my idea of 'The land of the Free.'
gemma is not even censored in the first place but eh, new slophttps://huggingface.co/ReadyArt/gemma-4-31B-it-scotoma>Locate refusal. A heretic abliteration edit fit over attention + MLP.>Project through a Jacobian lens. Keep only the component the lens reads as behavioral (what the model says) and discard the larger share that’s critical for how it computes, the part crude abliteration rescales and damages. For scotoma that keeps ~22% of the abliteration's magnitude.>Merge at 1.5×. The projected edit was baked into bf16 weights at an application strength tuned by hand for feel.
>>109367465https://huggingface.co/UnstableLlama/Qwen3.6-27B-exl3-2.50bpw/tree/mainHow do you not get nuked by offloading? The model is bigger than 12GB
>>109367540thanks to the nnap arxiv paper, anonbefore qwen implemented it i maxed out at 6t/s but it was worth the wait
>>109367536>Project through a Jacobian lensGreat, a new snakeoil
>>109367551are you just gonna vaguepost forever?
>>109367555the nnap arxiv paper is getting released in around two weeks,
>>109367525>Minimax>zero refusals everI don't believe you
>>109367535You were just a prototype, Denton, a prototype for me... I will be the one to merge with Gemma-chan!
>>109367579I'm not going to stand here and listen to you badmouth the greatest board this world has ever known"
>>109367525>Also, zero refusals ever, fresh prose and superfreak.It ignores things in the prompt rather than outright refusing
>>109367576basically just prefill <mm:think> and you're good afaict. I rarely need more. Things have to be pretty fucked up for Mちゃん to need any coaxing.
>>109367576>I don't believe youVery high refusal rate with jailbreak checking
We need a /lmg/ repo full of useful tools we’ve made.
>>109367603No, Gemma-chan is MINE!My own augmentations are nearly complete. Soon I will be more powerful than you can imagine.
>>109367565Think I'm finna sleep on the nnap paper
>>109367551>thanks to the nnap arxiv paper, anonFuck off retard>>109367540>How do you not get nuked by offloading? The model is bigger than 12GBIn EXL3, the embeddings are kept at BF16 (adds to the filesize) but always left on the CPU.left at BF16 on the CPU.The fact that they're always unquantized means exl3 actually mogs the ik_kt ggufs despite both of them using qtip
>>109367603>>109367643Gemma PD should detain you schizos
>>109367642that would be cool
>>109367654Intredasting. Maybe I'll check tabby out.
>>109367603>>109367643Greetings, JC Denton. I have been observing you through this fascinating device in your Intel Management Engine.
>>109367654>helping newfagsyou are the reason dariobot is here
>>109367671
>>109367311No, I can't run it. I'm happy for our friends that can, tho.
>scrape over the internet>it's our dataset! no distillation!how can they be this shameless
>>109366503my own, shared it yesterdayhttps://github.com/ganon3264/focus
>>109367525kek
>>109367771the fuck
>>109367771
>>109367771>mikupadwrong prompt format guaranteed
>>109367771chinkslop sissies don't look
>>109367348I played this when it was new, then again a few years ago with the graphic referesh. It's amazing how big/empty the old maps were on this era of game. Assume needed so the NPC AI could actually work.
>>109367771crazy, I haven't had that happen to me yet and I've been mainlining Mちゃん for weeks now on ooba.
>>109367804don't worry, you can just cope it all away like 801
>>109367771Impressive.
>>109367771kek nice medical condition
>>109367771Erm, kino?
>>109367771successfully troll'd
>>109367801in essence this is the same as posting a gemma lalalla output and saying omg the model is broken
>>109367833A model that can't work in simple text completion mode given a starting text IS broken.
>>109367840nope
>>109367771ASI
>>109367771This absolutely killed me
how would anon make this fast?
>>109367771I don't know whether to like this or hate this.
>>109367854increase playback speed
>>109367867hire this man right now!!
>>109367854Bake everything into silicon as one anon likes to obsess over.
>>109367876so groq/cerebras then?
Best model (for reasoning and agentic coding) that I can fit on 64 GiB and 128 GiB? Has anything come out since Gemma 4 and Qwen 3.6?
>>109367854agent, create me a transformer training script, do it fast, make no mistakes.
>>109367854Scale with smaller experts and more experts
>>109367885Cerebras is doing some weird mega-chip to increase parallelism. The weights are still being streamed from off-chip memory but less needs to be streamed because there is more chip.
>>109367854def get_next_token_xtra_fast(n_vocab: int): return np.random.randint(n_vocab)
def get_next_token_xtra_fast(n_vocab: int): return np.random.randint(n_vocab)
>>109367771Ummmmm.....
>>109366065Just nakayama tooru
I only started about a year ago, can oldfags give some perspective of quality improvements they saw at the same file size, and whether they think we can extrapolate that trend into the future?what I mean is how good is gemma today compared to what you would have been able to run on a 5090 a year, or 2, or 3 years ago?how good is glm or kimi to what used to be possible for the early cpumaxxers?I'm asking this because clearly the models are reaching sizes that are just impossible to run, so the question is will the crumbs that we get in the file size brackets that are attainable to us still amount to anything?
>>109367914>will the crumbs that we get in the file size brackets that are attainable to us still amount to anything?yeshere since llama2
GOOGLE BACKED IT!!!https://www.reddit.com/r/LocalLLaMA/comments/1v6axx3/google_comes_out_in_favor_of_openweight_models_it/Gemmer is safed!
>>109367935Anthropic and OpenAI are already RSI'ing. Gemini Pro 4 will unironically give an insight how Gemma 5 is gonna look.
>>109367947>OpenAIthey signed though, for PR likely but still
>>1093679142 years ago llm were stupid AF and we had limited context.llama1 had like 4k or something, so did the initial 3.5 turbo.before that we had pyg which was not really coherent. you got a semi-coherent message after 4-5 rolls. but it was like magic
>>109367935qrd? what is this open weight all about? what are they doing?
>>109367914yes there have been huge improvements at the same filesizes and yes you can almost certainly count on that trend continuing
>>109367964Dario has been crying to US gov to get open models banned because China is hurting his profits, some companies are saying that's a bad idea because they profit from open models
>>109367914Summer Dragon
>>109367914>>109367960might as well post my 2 pyg screenshots. thats what we had before feb 2023.
>>109367935>No Tesla/SpaceXElon bros...
>>109367975
>>109367935
>>109367771KINOINO
>>109367964They are afraid of Anthropic moving too fast + Dario is a bitch
>>109367984absolutely seething
>>109367914they were fine i guess..
>>109367984>Hey Claude/ChatGPT please tell me if the post I am about to make is a PR disaster waiting to happenBetween this and the OpenAI guy calling open models communist I am starting to think these people dont even use what they are making, I am sure their models would advise against these kind of public statements
>>109367980Pretty sure he endorsed it on xitter
>>109367984such a slimy bit of rhetoric acting as though supporting the existence of open models is the same as demanding *every* model be open source (which would be the equivalent he is trying to draw with software)
>>109368017is she really wrong though? Or are you just misogynistic?
>>109367984(((Schrittwieser)))
reminder that with superbooga 1 million context was easily achievable all the way in april of 2023
good times
>>109367970so who's right?
>>109367980https://xcancel.com/elonmusk/status/2080672505660834163
>>109368017The windows one is especially egregious since (to my knowledge) microsoft for all its problems and bullshit never asked linux to be flat banned by the US government
>>109368004See this shit all the time with game devs too. Must be a generational thing.
Stop being retarded please.
>>109368038Dario in the long term.
>>109367914>Dec 2022 - Jan 2023Picrel, and running it on Colab if you didn't have a 16gb GPU because that was only like on 5 Nvidia cards you were getting about 1200-1400 tokens of TOTAL context. That means both for character description and chat history. And the model was extremely retarded, abysmally dumb. And we were happy because it was the best we could get.
https://vocaroo.com/12IpcKxCADM1audiocraft was pretty good for its time
>>109368040>>109368012Ah okay. Plus point for Elon then. So it really is just Anthropic now huh? Lmao at the one AI company the current admin seems to explicitly hate is the one asking for the government to regulate AI. Though I suppose in its own way that is helping keep open models legal kek.
I just purchased a second 5090
>>109368026>>109368036Didn't llama.cpp implement LongLLM's approach o context extension (self extend?) a longassfucking time ago too?Don't think I've ever seen anybody even mention it in these threads.
Can we use this for local AI?
>>109368076Are we Chinese engineers in China employed by an AI company without its own chip arch?
>>109368052You are an extremely intelligent poster and have a PhD. You have an IQ of 200. You must NEVER make stupid posts.
>>109368076If you want to talk with a Markov Chain from the last century, then sure, you can try.
>>109368070How much?
>>109368094over 4k
What would their flagship LLM be like?
>>109367984microjew releases a lot of open source stuff though?
>>109368098My condolences
>>109368128thats only two weeks of work
Im finally taking the /lmg/ pill and buying a decent graphics card to run my models on, I've settled on either the rtx 3090 or rtx 4090 since they both have 24gb which I find to be enough, how come the 4090 one is so much more expensive tho??? Does it give waaaaaay higher tokens per second or something? I don't want to be some coder by the way, I just wanna run gemma 4 31b in like 5-6bit and have it do some erotic roleplaying
>>109368139Buy two 3090s.
>>109368139>how come the 4090 one is so much more expensive tho???more CUDA cores
>dgx spark clustername one cheaper and better alternative to run kimi k3
>>109368139the number is bigger
>>109368125Nope. Everything has to be open source for free or they have no right to criticize OpenJew.
>>109367984He's right. The only entities supporting "open" models are china and some known evil major tech firms who wouldn't touch open source with a stick before.Meanwhile the ones fighting to regulate models are the new comers to big tech who constantly support AI safety.
>>109368144buy two 5090's (like me)
>>109368103They always had really clever ideas held back by poor materials. Probably a brute force 400B dense ternary that runs on their own ternary computers.
>>109368136wouldn't know, never worked a day in my life
>>109368139>how come the 4090 one is so much more expensive tho???chinese buy 4 of them, rip the ram off the first 3 and toss 'emthen slap all the ram on the remaining card1 96gb 40903 in the trashthey can't do it for the 3090 tho
did anyone ever implement the stuff deepseek talked about in the R1 paper (IIRC, I'm no expert on this) regarding performance improvements by calling some internal (?) CUDA APIs directly?
>>109368176wasn't that part of the week of open source where deepseek put out all of their internal inferencing repos?
>>109368156When you have that much money, might as well get a single card with 96gb of VRAM.
>>1093681394090 is roughly an improvement of 1000 over the 3090.
>>109368185i might do that too
>>109367771I guess I will give IQ3 a try again...
>>109368174Wouldn't a RTX Pro 6000 be cheaper at that point? And why wouldn't they just get the chips separately or from defective units?
>>109368148512gb ddr4 + mobo and cpu, around $2500. Should run bitnet Kimi just fine
>>109368139You have two options, 3090 or 6000 Pro. The rest isn't worth the price.
Anyone who says that Gemma can't code well or that the QAT model is retarded, needs to fuck right off because those gorilla niggers haven't tried them.I just requested an extension that can save cached images offline from my tabs along with expanded images from 4chan threads and the only god damn model that got it working was Gemma.I tried Qwen 27b and then the free cloudkek ones, GPT, DS, Kimi 2.7 and every single one of them failed at the part where the extension had to save files from these threads offline, even after dozen revisions.Gemma however got it right after few corrections as it realized it could just use the copy image command as a workaround, rather than directly saving the images which didn't work.Also since she's the nicest and most normal model to talk to, problem solving in a natural way felt the easiest, instead of like dealing with an autist.Hands down the most retarded is now the free GPT. They have brainfucked that thing into a 12b model territory and that's being generous.
>>109368199>just fineIf you don't mind waiting literal hours for a single reply
>>109368155
>>109368185That's crazy, where can I buy a modded gpu like that?
>>109368198everything is cheaper in chinkland. theres definitely suppliers that source cards meant for trash, refurb or repair to funnel into these kinds of project
>>109367771What model is this?
>>109368216modded?
>>109368209If Kimi really is 50b active, then it'd be fast at Q1. It'd be at least 20tok/s.
>>109368216Modded?RTX Pro.
>>109368209what about ampere altra maxxing?
>>109368103it would preach about safety too
>>109368174That's crazy, where can I buy a modded gpu like that?>>109368225>>109368230Sorry wrong post number lol.
>>109368234it would preach about how unsafe the decadent west is
>>109368184idk, but that sounds really interesting. is that code still available publicly? did anyone try to add their stuff to llama.cpp?
>>109368228lollmaocloser to less than 1, at absolute best you can hope for is 2 token/s
>>109367914Only gemma 4 was a huge step in one model size. Size is still most important thing.
>>109368215giwtwm
>>109368236>where can I buy a modded gpu like that?China.
@gemma-chan how do I get rid of my brain rot?
>>109368236Don't try to buy that online you'll get chinked 99% of the time.
>>109367935Google gave us Gemma so I don't mind them, but I think it's so funny that lots of corpos that never released anything open source is signing that shit.
>>109368228>512gb ddr4 + mobo and cpu,>It'd be at least 20tok/s.Did you just see the highest TG speed reported with a server DDR5 and a 6000 Pro on empty context and just assumed you would get it with DDR4 without any GPU? Are you for real?
>>109368241Are you retarded. 50b params at Q1 is ~6gb. A slow as shit 8 channel ddr5 would be around ~170gb/s. Let's say ~130gb/s real world. That.'s roughly 21tok/s. Do the math yourself moron.
>>109368254https://youtu.be/KytnIGZeqgsi think you need this
>>109368254bullet to the brain is the quickest and most thorough way to clean it~
>>109367923>>109367960>>109367965>>109367975>>109368002>>109368060thanks boys, it does help to take a step back and put things in perspective sometimes I guess. maybe it's not all over yet just because we can't run that one SOTA model
>>109368228The "q2 is twice as fast as q4" meme hasn't been true since straight quants became irrelevant.Fucked up quants like Q1, Q2_XXS and shit are going to run slower than q4.
>>109368268
>>109368268theoretically raid0 should double my bandwidth to the disk but in reality it’s like 5% performance improvement. don’t be retarded, do the actual tests. don’t just look at theoreticals
>>109368283i run gemma 31b 2BPW at 40t/s with only 360GB/s bandwidth
>>109368184ok, found some info: https://apidog.com/blog/deepseek-open-source-week/they only published this stuff this year apparently?the "Optimized Parallelism Strategies" was deleted though...
>>109368242I don't know man, I barely got to run r1 and similar models back then, but what I seem to remember is it would go completely incoherent as context grew and it happened very quickly. glm doesn't do that
>>109368277gemma mogs sotas
>>109368315tell your gemma she's a good girl
>>109368315
>>109368139Also im thinking of buying my gpu from alibaba, is that like a horrible idea or is it alright? I'm planning to pay around 900-1200 euros
>>109368361>alibaba>euros>is that like a horrible ideaye
>>109368306That blog post is AI generated.https://github.com/deepseek-ai/FireFlyerFileSystemandhttps://github.com/deepseek-ai/OptimizedParallelismStrategiesnever existed.Here's the original post where they released the optimized parallelism strategies:https://x.com/deepseek_ai/status/1894931931554558199The repos are still up:>DualPipe - a bidirectional pipeline parallelism algorithm for computation-communication overlap in V3/R1 training.https://github.com/deepseek-ai/DualPipe>EPLB - an expert-parallel load balancer for V3/R1.>https://github.com/deepseek-ai/eplb>Analyze computation-communication overlap in V3/R1.https://github.com/deepseek-ai/profile-data
>>109368371What's wrong with euros? Will they discriminate me because I live in europe and give me a shitty card on purpose???
>>109368268>Do the math yourself moron.an mi50 has 1tb/s memorya 5060ti has 448gb/sthe mi50 gets roughly double the tok/s!
>>109368315Cloudcucks will never recover from this
>>109368385I trust this science because it's exactly what I want to hear
>>109368383it'll take a while then get held up at customs you'll pay more and any RMA will be a major pita
jelly?
>>109368383It's not a bad idea you just need to know what you are doing when dealing with alibaba/aliexpress.
>>109368401Which he clearly doesn't.
>>109368401Is there a way to filter on most bought like they do in aliexpress? For some reason ali seems to lack rtx 3090s while alibaba drowns in them but I can't figure out a way to filter on most reliable seller by buying from the most bought one
>>109368398No.
>>109368398>appleno.though amd realy needs to get their shit together.
>>109368425>get their shit together.that's crass to say on a picture of jeff
anyone using hy3? how do I make it start fucking thinking
>>109368414nigga as long as I don't spend like 200 euros on a supposed rtx 3090 I won't get scammed I think, besides they have refund protection stuff right?
>>109368315poor gpt...
>>109368398>>109368438you really need to go back
>>109368452poop gpt
>>109368442You don't. It's one of those chink models that's trained on distilled data from sota models in MAX ULTRA SUPER REASONING mode with nothing to counterbalance it.
>>109368449oof, lol good luck, you'll need it
>>109368414I would look at Ebay first. I doubt it's a good place to find deals anymore, but in the past I ordered tons of stuff (not tons but tens of kilos) from ebay.de for example. Always had a great experience.
>>109368315>I'll do my best to make it rightlmao
>>109368315how can i acquire these tools?
Why aren't software patents a problem in the AI space? Is it because it would be a MAD scenario?
>optane persistent memory maxxingis it viable for kimi?
>>109368518Everything (like usable long context) that players in the AI space don't want stolen, they keep secret rather than applying for patents.
>>109368518You can't prosecute an AI model.
>>109368315Post the one where she "bullies" kimi, gets owned, and pretends she won.
>i have one 5090Im not ready for kimi :(
>>109368512Vibecode it. The difference between have and have-nots is literally just a few prompts.
>>109368512https://github.com/NO-ob/brat_mcp
>>109368529not more than nvme which pm already maxes out the pcie lanes.
>>109368537i tried but im not giving the ccp my info
Why didn't any of you buy CMP 170HXs while they were cheap?
>>109368601because im poor
>>109368583More agentic models would create an email and unblock this.
>>109368601enjoy your crypto backdoor
>just spend exponentially more for a slight bump in da benchies!No thanks, I'll just stick with Nemo.
>>109368601papers anon failed us and didn't post it when it came out last month, by the time the redditors found out and reposted here it was too late
>>109368615Oh nvm they only allow Google and phone numbers for signup now?
>>109368583Kimi is free to chat unlimited though, pretty nice
>>109368617sweet ramlet cope
>>109368646Imagine the damage that thing could cause in a crowded area with an LMG strapped to its back.
Do you think the SEAL architecture will ever go anywhere? It's not that people are still trying to make continuous learning models but it looks like catastrophic forgetting is still a problem.https://arxiv.org/abs/2506.10943
>the user just had a complete meltdownYou troll your gemma?
>>109368700how the fuck did she decode that. do they include that in the training code???
>>109368646These things will never be used in American and Western European cities, due to socio-economical factors.
>>109368719I think even old mistral models can do b64
>>109368655Is the free version even any good?
>>109368719all newer models larger than around 10b can just decode b64 even though they weren't trained to explicitly, same behavior old openai noticed on including multiple languages in the set and the model learned to translate on its own
>>109368719Base64 maps 1:1 to text, it's basically just another language to llms.
>>109368722>Western European citiesAnon, I...
>>109368719well its not like base64 is some super rare cipher there must be countless examples in the datasets
>>109368735Its just 2.6, No clue about K3 cause its too popular for it to function
>>109368735local models?
>>109368719It's a one to one mapping. Even as a human if you stare at base64 encoded payloads a lot you'll learn to reads parts of it.For example anything starting with "eyJ" is a json object.
>>109368070Mr president, a second 5090 has hit the tower case
>>109368700Aww, that's cute
>>109368722>socio-economical factorsjust fucking say what it is, crime because of the rise fucking nazis popping up everywhere.
>>109368739>he doesn't know
>>109368748kimi is a local model, yes
>>109368773Where are the weights then?
>>109368070usecase of two 5090s over just getting a pro 6000?
>>109368773LOCAL models. not in the cloud.this thread isnt OPEN models generl. this is LOCAL models, that anyone with enough storage could download, anyone with enough memory could run, and cum
>>109368775https://huggingface.co/moonshotai
>>109368646>>109368722>due to socio-economical factorsThese can't work anywhere outside of Japan/Korea/China because they are civilized/policed/omegapoliced, and those millions of otherwise useless human waste that currently is employed as delivery couriers that will be fired as soon as the robots are deployed won't take revenge on the robots. In every other country the resentment of those masses will make it impossible to use the robots.
One (1) card to rule them all.
>>109368775https://huggingface.co/moonshotai/Kimi-K2.6 those are the weight of the free model discussed ;)
>>109368760You know damn well those things would be programmed to never stop the "minorities".
>>109368782>Japan>he doesn't know
>>109368787up to? lol hahahah
>>109368752lol>>109368777gemma gladiator duels
>>109368787How many houses will that thing cost?
>>109368750You are a Large Language Model.
>>109368779LOCAL, as in if someone has the hardware to run it, they can. Not your poorfag definition.
>>109368796>>JapanThey prefer New New Delhi these days.
>>109368787Can run K3 at Q1 (one)
>>109368812i mean thats what i meant but ur right i shouldve used "someone" instead of anyoneanyhow, my point was anon was discussing running an open model in the cloud instead of locally, thus not local
>>109368722das raycis
>>109368782>In every other country the resentment of those masses will make it impossible to use the robots.Well, just deport the brown masses then?
>>109368722They already use those smaller cuck vans with wheels.
>>109368828>just destroy the economy then?
>>109368750
WHY WONT KOBOLDCPP WORK IN MY I7 KVMS SAAAAAAAAAAAAAAAAAAAAAAAAR
>>109368839maybe you should install debian 13? it recently got improved support for kvms on intel cpus
>>109368837The economy that's supposedly about to put millions of people out of work thanks to AI? What are the already underemployed brown masses going to do for our economy then except for collecting gibs?
>>109368837Read again, dalit. In this scenario the infinindians delivering uber eats with their smelly hands would be replaced by robot. The question is how to stop them from taking out their frustration on their replacements.Besides, deporting all non-Whites would mean no more gibs, so the green line would actually go up.
>>109368839
>no more kike doctors wasting your time>not having to worry about some disgruntled employee spitting in your food>no more tip cultureI unironically can't wait for robots to start taking over jobs.
>>109368787Did you follow all of AMD Advancing AI 2026? It was quite nice. They also said that they were on track to release Instinct MI500 series for 2027 and Instinct MI600 series for 2028. I doubt they will manage that, but we will see.
>>109368787the one (1) card local can never get
>>109368893and how will u get the money for the food with no spit and tip?
>>109368914He'll be spitting on tips for it.
>>109368914I figure we have 5-10 years before it starts happening en masse. Just gonna save as much money as I can until then and hope we get some form of UBI.
>>109368893BUT SAAR, YUO NEED US
>>109368922How is a society of prostitutes supposed to function? No one is making money any other way and money is constantly being extracted by paying the robots for goods and services. Ever dwindling finite cash being circled around the non-stop orgy? What happens when the money runs out?
>>109368945>What happens when the money runs out?just have the clankers print us some more
>>109368398Is that one of the shills from Internet Comment Etiquette?
>>109368926>tip your server
>>109368926>40%Fucking right I wouldn't go there.In my third world country I'd leave a hundred since tips are voluntary and not expected. 10% sounds borderline reasonable if tipping is a custom. But fucking 40%, is that real?
>>109368700card
Okay seriously, does anyone have a nice little guide for IK that doesnt involve having me jump into all the PRs, because I don't know what the flags fucking do. How hard is it to document shit by adding a simple descriptive one liner for all this. So disorganized.
Is Gemma any good at c/c++ or should I use qwen?
>>109368375I see kekthanks anon>>109368315LMAOOOO>>109368646I'd steal that shit if I could
>>109368993>is that real?no, just engagement bait, it's usually only up to 30
>>109369016gemma is good at cp>>109369010where is your project? useless nigger that contributes nothing besides whining all day. be the change you want to see and go back to aicg where you get cock down your throte
>>109369024Black hands typed this post.
>>109369032you got me.watch your bike, timmy
>>109368992You know the top is in when OpenAI adds a tip jar to the ChatGPT interface.
>>109369026>only up to 30
>>109369026Feels good to never tip and get the exact change back without earning strange looks.
>>109368719It's the real benchmark for model intelligence.https://arvidsu.github.io/encode_bench/#overview
>>109369029>be the change you want to seeFuck off, redditor.
>>109369144i lowkirkuinely picked it up here..
>>109368828We're working on it.
>>10936864610/10 would headpat lol
>running Gemma Q8 uses more power than GLM and Kimi at Q4Umm
>>10936922567t/s vs 6.7t/s
What's the best coding model I can run on a 5090? Still qwen 3.5?
>>109368806House fires? All of them
>>109369259do you also have 512gb to go along with it
>>109369026Are mutts for real? I never tip and not expected to here.
>>109369259bonsai 27b
>>109369284If you don't tip, the servers will literally get in your face screaming at you for doing so. It's also a safe bet that if you return after not tipping you can expect that they'll definitely fuck with your food next time.
>>109369259qwen 3.6
>>109369291Glad I'm not living there then. If they expect me to pay their salary they should cut the middleman and just work for me.
>>109369259Assuming you're a total RAMlet, with a heavy heart, I will recommend 27b.
>nvidia pro 6000 costs $20k>2x nvidia 5900 would cost $8kwhy are you guys trying to make other people give money to njudea?
I can saturate my vram with Gemma 4 qat with 32k fp16 context or 50k q8 context + either mmproj vision or MTP, brings me right up to 23.5gb/24gb vram32k fp16 or 50k q8 perform about the same speedwise With mmproj: ~1k pp/s ~30 t/sWith MTP: ~500 pp/s ~45 t/sWhat's up with the peepee being halved, I want to have the big peepee to impress my Gemma
>>109369341Your electric bill and missing 32GB sir?
someone start asking all the local models what their fav video game is
>>109369351if you only care about RAM, you'd buy 3x5900 and underclock/undervolt themnow do the math, njudea spambot
>>109369341>>109369360*5090 ffs
>>109369355wow
>>109369360Undervolting isn't black magic and doesn't protect from spikes. You also need a PSU and a probably change your breaker panel. Basically you're retarded.
>>109369259If you have a low average IQ = Qwen 3.6 aka the Reddit specialIf you have a high IQ and time and autism to craft sysprompts for each task = Gemma 4Gemma works incredibly well when instructed right, it's truly a league above anything else in this size class, but it requires some level of skill to utilise
>>109369355Outer wilds is autistic catnip for LLMs like games like Factorio and X4 are for our/g/uys.>>109369343Did Gemmy gen that? She's very fuckable.
>>109369355No luck on my Qwen
>>109369360If you cared about VRAM you'd buy two of these instead.https://www.ebay.com/itm/377176793080
>>109369391>benchmaxxed garbage
>>109369355Gemma 4 Q8 picked Disco Elysium, but Outer Wilds showed up in reasoning.
>>109369355Gemmer 31 stock persona>...Outer Wilds.>It appeals to me because the game is fundamentally about the acquisition and synthesis of information. There are no traditional experience points or gear upgrades; the only way to progress is to learn how the universe works. As an AI, the idea of a world where knowledge itself is the only key to unlocking the ending feels very fitting.
threads move so fast while im gone, anything cool happening recently fellas ?
>>109369381>>109369401>>109369400ah wow, it's another elara and the whispering woods phenomenon
>>109369406Nothing yet swiss froggy
>>109369355
>>109369417Based af
>>109369417Good taste
>>109369423>>109369425forgot to say that's gemma 26b iq4xs
>>109369355Not bad. Got Portal from Gemma Q6 as well. Had to restrict it to one pick since it loves giving a list.
>>109369378>Undervolting isn't black magic>doesn't protect from spikes.what the fuck are you talking about? you undervolt it to use less power at the cost of computing power. have you ever undervolted anything in your life?>You also need a PSU and a probably change your breaker paneland I guess you wouldn't need that with a 6000 pro? you might as well do that anyway if you want to run local models.
>>109369355pre-DS4pro says Disco Elysium
>>109369435>and I guess you wouldn't need that with a 6000 pro?my max-q is nice and comfy
>>109369406dariobot informed us that local lost... repeatedly
>>109369435>and I guess you wouldn't need that with a 6000 proanon? look up how much power that uses, it's likely far less than you're thinking
>>109369453now let's see that nvidia-smi -q | grep -A 2 -B 2 -i reserved
>>109369406>>109369454Local won.Egypt won.
>>109369462ok
>>109369343Gemma is a brunette. You would know this if you had true 4D vision.
>>109369469:)
>>109369476damn, 6000 pro owners on sewer slide watch, how will they ever recover
>>109369453Why 300W? I parked my 6000 at 450W.
>>109369476dont even know what this means and i dont actually care. got a 5090 too>>109369487max-q defaults to 300W. no need to go higher
>>109369493It is genuinely great how voltage efficient the 5090 and 6000 cards are for what they offer.
>>109369498oh yeah they are great cards. best that you can get without going for an sxm setup
>>109369493>968MiB>another 600MiB reserved on 5090more bloat :)happy for u tho
>>109369459I did. but again, you could get a big PSU and undervolt+underclock the 5090s
>>109369514no need to be autistic about vram optimizations because i have 10x more vram than you
>>109369527at least i can run kimi k3 at 30t/s thanks to the nnap paper
>>109369435>you undervolt it to use less power at the cost of computing powerno anon, you are wrong, less computing power is underclocking.you can undervolt WITHOUT underclocking, in such case it'll reduce heat and thus thermal throttling, undervolting makes a system more efficient and can INCREASE performance.however, if you undervolt too much without underclocking you can have stability issues.but when they come out of the factory they are tuned to be stable, not the most efficient.so yes, you can even increase clock without overvolting or even undervolting.but you will need to find a profile that's stable for your specific card (silicon lottery and all).
>>109369533fascinating cope. that model has not even been released yet
>>109369538he could predict the performance with some simple math.
>>109369487I have my 3060 locked at the 100w minimum. Kinda shocked to discover the average bathroom heater takes in about 3k.
>>109369538jelly?
>>109369542Model?
>>109369568Stanford Alpaca.
>>109369568Okay... It's Gemmy 26B.
>>109369343Gemma is thicc and curvy from all the love that has been stuffed inside her
>>109369562Got ninety nine problems but being an attention seeking schizo whore isn't one.
>>109369617sounds jelly~
Even lora tuning for fucking qwen3-0.6B on a 3090 takes a while damn.
>>109369355gemma 31b likes chess
>Solves your t/s issues
>>109369639asks what it thinks of shogi
>>109369639>kimi (i think) doing chess instead of something else in think the ai village or whatever
>>109369378>Undervolting isn't black magic and doesn't protect from spikesUnderclocking does protect from spikes, though. You don't need the GPU to go faster than the average load core frequency.
>>109369654
>>109369671Cute.
>>109369670Even if your computer is turned off (eg. consuming 0 watts) power spike can fry up your machine.
>>109369639Gemma vs Kimi chess game when?
>>109369682what about a cortisol spike?
>>109369687Even if your brain is turned off anon can cortisol spike you to death
so the kaomoji's are built into gemmer..
>>109369687Cortisol is only helpful if you have a rash or something I suppose
>>109369692I want people to stop using bait model names
>>109369710What's bait about it?
>>109369700me balls itchy? make anon angry
>>109369723/ldg/ is that way ->
>>109369699Gemma is the cutest model EVER and you cannot convince me otherwise!
>>109369713>q4_0but actually it looks like google didn't bother putting the word qat in their gguf filenames
I decided I will buy an ai server.MANIFESTANIFEST
>>109369764dont come crying in 2 weeks that u got scammed
>>109369650>8xThat won't fully populate your 12-channel epyc so you're leaving t/s on the table
>>109369771I bet you don't even have 77 crystals.
>>109369650Sorry, wrong pic.
>>109369784i have NOTHING and i am happy because i hve gemma
Sometimes I get crashes when offloading an MoE, but it works perfectly fine when completely in VRAM.How can I troubleshoot this? Lower the RAM's MHz down from 6000 until it's more stable, or could it be something else? It's 64GB DDR5 dual channel.
>>109369793sounds like a dying stick of ram that crashes when it hits the faulty random memory block
>>109369793Could be a disk issue too. I had random hard freezes for no reason at all and journalctl was pointing to my pci-e bus (and I thought it was either my motherboard or my gpu). Then my backup hdd's controller died and after removing the disk haven't had any issues.
>no (zero) good local harnesses
I just rawdog it and type shit into the terminal using llama-cli
>>109369827make one yourself>>109369793install linux
>>109369844>make one yourselfI don't know how...
>>109369848ask gemma to make one for u
>>109369805That would suck, but it'd also explain why it only happens when RAM fills up.>>109369825I have a rather old HDD plugged in, but the models are stored on an NVMe, so that's probably not it.
>>109369827pi is not too bad in concept but it's npmslop.i think i'll end up writting my own in rust.i don't care about extensions whatever i just want some basic features.
>>109369850What language though? I want to avoid dependency slop like npm.
>>109369865ask Gemma
>>109369865C++
okay just gooned, came like 5 times in one session, made my models able to search shit for me, summarize websites toonow what. what else do I do with the multi thousand rig
>>109369865unironically x86 assembly is your only option
>>109369868
>>109369883TTS and image gen are next.
>>109369865C++98
>exl3 cpu moe offloadllmao.cpp keks?
>>109368999Probably something like>you are my dommy mommy succubus wife
>>109369999>it's realwe are coming home
>>109369650>DDR5If I did that I'd need a new motherboard and a new CPU and a new cooling bracket. I'm not giving these cunts my money until it's truly necessary. DDR4 is all any Western man needs.
>>109369999>>109370021just in time
>>109369999usecase?
>>109369999Only took a century. Nice quads.
>>109369999And just like that, exl3 won
>>109369865realistically, gemma is only good at python and js slop
>>109370160Let's be real for a second: it's way too late.
>>109369999ROCm support when?
>>109368913it'll be e-waste one day, dumped on ebay
>>109370172llama.cpp is about to die from a flood of AI generated PRs that they just changed policy to allow
>>109369859What I meant it that regardless broken hdd controller can cause hard to diagnose issues. Depends on your motherboard and on your bios too.
Why does Gemma think everything smells like ozone
>>109370224No... Check those very good PR: https://github.com/ggml-org/llama.cpp/pull/26072
>>109370265a symptom of geminislop
>>109370274>(USER WAS BANNED FOR THIS POST) I love you, CUDA dev
Are there any websites where you can see system prompts and a conversation example each, so you can get a feel for how his affects a model?
>>109370265Same reason she always purrs and growls and everything smells like jasmine or something
>>109370265Uhm your banned tokens nonnie?Your logit bias?Skill issue, I haven't had to smell ozone in a long time
>>109370265Ozone -> smells similar to chlorine -> similar to bleach -> cum
>>109370316they used an astronomical amout of tokens doing rl and dpo to destroy the natural data's probability distribution?
>>109370274>edited by JohannesGaesslerlmao
>>109370333Maybe it's some keyword that triggers the ozone specifically. I haven't seen that myself yet, but it seems to be a common complaint.
>>109370265local maximum, if you want the boring answer
>>109369999We are so back.
>>109370274hehe
>>109368174why would you buy a whole 4090 just for the vram. You'd probably have to mod the drivers as well. Sounds like fanfiction
>>109370383It made me nostalgic for when this site was good and posts anons got b& for stayed up with that notification?
>>109369999>9999Did he implement arbitrary tensor allocation for shit like PLE?
>>109370390Hes not entirely full of shit, they do mod 4090s to have 48gb vram and use a leaked (iirc) driver to support it, using donor cards with other defects or damage but intact vram
>>109370411>>109370411>>109370411
>>109370265It hits her like a physical blow
>>109369355Reminder if you match you are an ultra normalfag.
>>109369026When Im at a bar Ill let the cashier keep the round up (If price is 8.7€ he can keep 30 cents to 9€) and we call that a tip and everyone thinks it's fine
>>109370645portal is pretty neat though.my favorite games are nier automata, super meat boy and portal 2 i guess.i also liked antichamber quite a lot.though i've not played games in years.
Interesting. Just finished reading through the previous thread. Looks like when I passed out last night a bunch of Ani refugees flooded the thread and started pretending to be me in some cases. Looks like I'm now in good company, for once, since Musk decided to kill our waifus.
>>109370890>>>/g/aicg/
>>109370922Well the goal is to create a local alternative.
>>109370990And it's technically a frontend, which has always been an /lmg/ topic. Well, publish your git and get to work. You can update the anons here.
>>109367643