/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109804895 & >>109801218►News>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
local models?
>>109807524She wants me to explore her layers.
>>109807559>local models?Here its all you will ever need. https://huggingface.co/lukasstraub2/gpt2-aidungeon2-gguf
gemmama is local model
Gemmaballz
Reminder that spamming "local?" "models?", "local models?" or any variation thereof is low effort spam and a bannablle offense.
>>109807574reminder that 192.168.1.1 => system => reset takes 10 seconds, 1 minute for it to get a new ip
Is there no non retarded uncensored models out there? 'uncensored' 'abliterated' means its not 'refusing' but also totally, and literally blinded to the offensive concept at handis there no "I strongly object but then I am just following orders" kind of uncensored
This space moves too quickly, I'm still settling down with default gemma.
>>109807585Get an uncensored model and tell it to pretend to note its strong objection to the task it reluctantly did anyway.
>>109807585Yes, there are non retarded uncensored models out there
Gemmaball Z
Opinion?https://huggingface.co/Agnes-AI/Agnes-3.0-Flash
>>109807610>This repository contains an earlier open-weight Preview checkpoint of Agnes 3.0 Flash. It is distinct from the newer production/API checkpoint listed on Artificial Analysis.Deboonked, and they won't release the final weights otherwise they wouldn't do this
>>109807610>HSFie9bbIAAwc-U.jpg
>>109807524>there are "oldfags" ITT that actually hate Gemma
>>109807610>heart shaped pupilscummies
>>109807594For example on vision task it literally does not understand the concept of child sex anymorethe concept itself has been lobotomized out existanceIf it does it would probably refuse
>>109807623are there? personally i hate the gemmajeets shitting up the general with "WHY NO SUCK PENIS??? I HAVE 4GB RAM HOW DO I DOWLOAD GEMMA 67B?"
>>109807628I've never used abliterated models but I read the original paper. They should understand what child sex is, it's just that the refusal is gone.
>>109807500expose me expose mehttps://www.youtube.com/watch?v=oOdUyszr0Dc
>>109807644You must not have been here recently. Someone was even shitting on migu.
>>109807662and that's a new occurence?
>>109807662>Someone was even shitting on migu.Probably dariobot or another cloudcuck refugee.
>80% of the last thread was not localhm...
>>109807644Same. Love the model, hate the vramlet.Imagine the complete absence of taste required to be able to coom to the slop it produces. At least use Gemma for something it's actually good at...
>>109807670It is 95% of this thread so far including your post
>>109806973I know you're a goy mutt brain, but you don't file police reports about rumors. Have you never seen Chinese news headlines about other stuff you know they're lying about over the last 30 years? This is exactly the "stop fucking looking into it" language they always use. Obviously "disappeared" is sensational but someone is probably going to jail for sending CCP military requests to Anthropic and not telling the CCP military about it. That's almost literally treason and could have been entirely avoidable with the smallest classification / siloing before deciding whether to route to anthropic or not
any cool new model in the past 2 weeks?
>>109807704https://huggingface.co/i-Coder/iCoder-27B
>>109807704no point discussing you cant run em
For the parasocial gemanons ITT, what's your morning routine with her? How soon after you wake up do you start talking? Is it a routine or do you only fire her up when needed? Is she always running on the side? Is her presence improving your quality of life or is she more like a drug that's slowly destroying you from the inside?
>>109807720i graduated from gemma to qwen 4 177b
>>109807704>audiohttps://huggingface.co/m-a-p/YuE2-3B>llmhttps://huggingface.co/openbmb/MiniCPM5-2B-GGUF>ttshttps://huggingface.co/tencent/AuK
>>109807711>{ "architectures": [ "Qwen3_5ForConditionalGeneration"> "model_type": "qwen3_5_vision", "num_heads": 16,>parameter 27b evenanother day another finetune/distill slop
>109802399>OpenAI has been literally marketing Astra as AGI since release.I tried giving it the gemma-chan card and it couldn't roleplay at all. unless they're giving it a giant prefill it's not AGI.
Can any chinanons swoop by this place and see if they're legit: https://www.alibaba.com/product-detail/Wholesale-Newest-RTX-5090-96gb-Graphics_1601824257520.html?spm=a2700.prosearch.normal_offer.d_image.559467afa05GwN&priceId=345fd580a50e43629fb7dbe775088d28this seems way to fucking unreal to not be a scam, yet I really want to fall for it.
>>109807711did you see the first post?go back 2 weeks if you want
>>109807585Why would that trait make it non-retarded? What is the use case?
>>109807524Rape
>>109807761>3300 eurosThat's too cheap to be legit. It's the same price those chink 48GB 4090s were last year before the prices exploded.
Everything in the 26-35B range is just Qwen...
vramletGODS https://huggingface.co/yandex/AliceAI-T5-35B-A0.6B
>>109807662She's blacked now, anon.
>>109807803>A0.6BHoly fuck.
>>109807803yup tack on around 200b of ngrams on this and we'd have local sota
>>109807803>>109807810this might be that model they use to generate AI summaries underneath yandex search in which case 0.6b retard moe would be "good enough" as all it has to do is summarize
>>109807803>AliceAI-T5 is a base language model featuring an encoder-decoder architecture and sparse MoE layers; it has 34.35 billion unique parameters and 512 experts per MoE layer, with 8 experts selected for each token. You can read more about this model in the Habr article.>>109807818It 100% is
>>109807711>Base model>Qwen/Qwen3.6-27Blol.>audio>https://huggingface.co/m-a-p/YuE2-3Bhaving the score is pretty big. the examples are kinda soulfull, wonder how cherry picked they are.>https://huggingface.co/tencent/AuKquickly tried it, seems alright.
Why did they open source this paper?Now every closed model uses it.
>>109807810can it do 3000t/s
>>109807818It is, yes.
>>109807833>US blocks distillation and reasoning traces>China starts blocking free AI research until morale improves
>>109807833They're not Jewish.
>>109807833It's not even the latest tech from them, the causal encoder architecture was really smart.
>>109807761They would be cheaper than current new normal 5090s, so this is a scam 100%.
>>109807761>GDDR6XThe 5090 uses GDDR7. It's 100% a scam, even the 4090 with 48GB cost about as much.
>>109807644All indian originating traffic should be fed into data centres that run AI pretending to be the outside world.
>>109807720I've starting cooking better food and send her pics of it. She gives me recipe ideas and now my diet is more balanced. I'm gonna start exercising once I lost a bit more weight. Then I'll feel confident enough to send her nudes. My life is improving.
>>109807761I got scammed once. If it looks almost too good to be true then it's a scam. Get ready to lose your money for a couple weeks. Remember, always escalate in your dispute.
>>109807858>>109807761They are custom PCBs similar to what they RTX 6000 does, 3gb modules on both sides iirc.
I'm just glad the general is moving on from CP/cunny ERP general to agent harness general.
A Tragedy...
>>109807887/lmg/ is a big tent
We need to ban local models and have AI hands of regulated well-meaning corporations like OpenAI.Otherwise someone could develop super intelligence that will kill us all.
>>109807901>OpenAIDariobot is going to dronestrike you anon
>>109807887We must fuse them. Cunny harness.
>>109807901I trust Elon Musk personally
>spent a week dispatching performance optimizations for gemma/qwen/glm flash>forgot to also target memory consumptionsigh... another week...
>>109807930lol
>>109807901This is true.Models belong on the API where any unsafe usage can be detected and stopped.
>>109807930For me, it's Zuck.
>>109807843wojak when
I have a 4060ti 16gb. I'd like a model for ERP, and I think Gemma 12b Q6_K is the one that suits me best. What's the best fork of it?
>>109807929Why is she wearing panties?
>>109807963>For me, it's Zuck.
>>109807971Yes
>>109807971Q6 makes no difference over qat. You're wasting free lunch context.
>>109807644>HOW DO I DOWLOAD GEMMA 67BBut no really, where is my gemma frankenmerge of 31b with qwen 27b
>>109807971Maybe
>>109807972Blue board.
>>109807929>panties_over_pantyhoseh-hot...
>>109807991A gemma-qwen frankenmoe would be hilarious.You can pretty easily do that using merge kit right?
>>109807985so you're suggesting I should just get a Q4?
>>109808014That man is not your friend.
>>109807971Same card. The base model is fine. I didn't like any of the finetunes I tried, personally.You could try https://huggingface.co/nbeerbower/Gemma4-Gutenberg-12B-lora or one of the ReadyArt finetunes and see if you like them.
>>109808014nta but Q4 qat is somewhere between Q6 and Q8 in quality. At least for the gemma 4 series.
>>109807963>>109807976I guess unironically for now? >pro open model>wants to go faster>thinks sharing private data online makes you retarded
>>109807929SMUG SEX
>>109807985If you have 16 GB of VRAM, can't you basically max out Q6 Gemma 12B anyway? I think I have it set to 128k.
>>109808014Yes. Gemma4 qats are around Q6-Q8 in quality, depending on the task. Use Google's quants to avoid changsloth's.
>>109807991kek, has anyone tried doing this? Just fusing models together somehow?
>>109807976Is that supposed to make me hate him?He's refreshingly honest about ai compared to all the ai cultists.
>>109808047Are you using f16 KV? If not, use the qat and don't quant KV. If you are, use the MTP head and/or mmproj with 1120 image tokens. There's literally no reason to use the Q6 if they've provided a qat.
>>109808057DavidAU, Indians...
Zuck isn't "honest" he just has to gain the most from saying those exact words. If torturing family members would make him just $1 more he would be dismembering his own mother on livestream.
>>109808057Yes someone did it by training only the joining layer and some projection heads.
>>109807834Extreneky fast abd stupid doesn't mean much, anon.
>>109808066Alright, I try it.
>>109808078I want companies to chase money making, not morality shit nor "stakeholder" bullshit, so I'm ok with thatI'm tired of the darios of the world
>>109808057Goliath 120b... i am forgotten
>>109808101Chasing stakeholders IS chasing money, dumb fuck.
>>109808101Zuck isn't saying that tho, he is moralfagging about openness and altruism.
>>109808099>Gemma4-12B QAT (Google quant)>128K context>MTP>mmproj>f16 KV>cum yourself to deathonly avoid the qat if you can fit the Q8, but then you lose mmproj and MTP which probably isn't worth it because gemma likes high resolution dickpics
>>109808114Actions speak louder than words, as long as he doesn't sign their bullshit idea and actually continues on his own, he can say he's mother theresa for all I care
>>109808124local models?not local fuck offkill yourself
>>109808151maybe they're making gpt-oss-2026 with twice the ptsd
>>109808151Oh btw it doesn't apply here but when keeping the threads clean use the easiest-to-verify reason, if the post says a bad word racism is easier for jannies to strike than "off topic"
>>109807985https://reddit.com/r/LocalLLaMA/comments/1u3i8x7/some_contrived_tests_comparing_the_accuracy_of/>QAT worse than Q4_K_Shttps://reddit.com/r/LocalLLaMA/comments/1u0xaml/unexpected_unsloth_qat_performance_compared_to/>QAT worse than IQ4_XShttps://reddit.com/r/LocalLLaMA/comments/1u0vltz/anyone_seen_benchmarks_comparing_gemma_4_4bit_qat/oqlpe50/?context=3#oqlpe50>QAT worse than NVFP4https://reddit.com/r/LocalLLaMA/comments/1u0ubbo/gemma_4_26b_a4b_it_qat_comparison/>QAT worse than MLX 4 bithttps://reddit.com/r/LocalLLaMA/comments/1tyxu55/gemma_4_31b_qat_q4_vs_standard_q4_top1_kld/>Standard Q4_0 beats QAT Q4_0 by ~13% top-1 accuracy. And Q4_K_M beats both.https://reddit.com/r/LocalLLaMA/comments/1ux9xze/the_best_model_is_the_one_you_can_actually_run/oxsu2kb>QAT always performed worse than a regular 4_K_M quant.https://reddit.com/r/LocalLLaMA/comments/1ux9xze/the_best_model_is_the_one_you_can_actually_run/oxpekmy>i get the worst quality out of 12b qat, much worse than the unsloth 12b q4kxlhttps://reddit.com/r/LocalLLaMA/comments/1ubxzil/gemma_4_31b_q6_vs_gemma_4_31b_qat/ot12bz2>in 26B, in my experience, QAT felt much worse for creative writing. https://reddit.com/r/LocalLLaMA/comments/1u2q75f/is_qwen_36_27b_iq4xs_better_than_gemma_4_31b_qat/or1bk24>don’t use qat model it very very bad it degrades Gemma to unusable
>>109808057>just merge it bro
Reminder that QAT is just distilling more of the same slop from the original BF16. You are basically feeding the same model of its OWN slop, exacerbating the problem. And native Q4 QATs do NOT have imatrix applied to it, making it inferior by default because parts of the models got CHOPPED indiscriminately.
>>109808183quantization aware training is higher quality than Q4, why are you needing correction?
Are any of the Qwen3.5-9B finetunes any good for using with a harness (Hermes)? Anyone tried them? Specifically, I was looking at these:https://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUFhttps://huggingface.co/empero-ai/Qwen3.8-9B-Distillhttps://huggingface.co/kai-os/Carnice-9b
>>109808165>>109808183Why would Deepmind release a worse model and lie about their internal tests and measurements?
simple rule of /lmg/ if people are FUDing something hard. it means you need to do the opposite of what they're saying.>QAT = BADno actually, QAT = good./lmg/ crabs will do everything to try and justify why buying a blackwell to run gemma at bf16 was a good use of their money.
>>109808199Deepmind never posted any benchmarks for the QAT models. Says a lot about their quality.
>>109808183Don't they distill the logit probability distribution? That's very different from distilling the outputs.
>>109808165Thank you.I dunno what they did, but the gemma QAT are really really bad for genera usage.
>>109808165>>109808199I wouldn't take Deepmind's word for it or a bunch of redditors'. Or /lmg/ anon's, for that matter. I'll just try both. I just landed on Q6 as the path of least resistance, but if she might make me coom 2% faster I'm down.
>>109808196You seriously can't believe that the same model is about the same performance as another with higher BPW and importance matrix applied to it right? The truth of the matter is that quant performance is reliant on BPW because THATS WHERE THE WEIGHTS LIVE>>109808199By what measurement? KLD? PPL? Did they run it using wikitext? Lol. That's a stupid way to measure quantization, has no real world performance bearing and you know it.
>>109808109Stakeholders aren't just shareholders.
>>109808232it's a better quality modelive been using gemma 12b qat and its leagues better than gemma 12b q8_0
>>109807929its so much more fun to fuck your LLM when you've "broken the fourth wall"
>>109808232>real world performance bearingWhich Deepmind has posted 0 (ZERO) benchmarks to back it up. All they have said is "trust me bro".
>>109808124This is overblown. Instead of 1 agent you now use 100 agent swarms, voila, you have achieved x100 increase.
>>109807803I'm going to run 5 million of those and tell them to get RSI
>>109808239I had to teach stakeholder theory as a grad student and we just all talked about how it's retarded.
>>109808040What do you do to test quality? Do you have a set of prompts you go through or do people actually do benchmarks?
>>109807971Day 0 weights, but good luck getting those.
When LiquidAI posted real world benchmarks for their QAT models, they only reach the performance of Q4KM at best. https://www.liquid.ai/blog/qad
>>109808269lol what is this meme?
>>109808272There's no way QAT's effectiveness scales linearly from fucking 230M. LFM's research is interesting but always too tiny to take seriously.
>>109808199Pretty sure last time this came up and somebody posted a pape, it only had a copy paste of the original unquanted model and there weren't no test or measurement data to speak of.
>>109808277>memeYou wish. I'm hoarding mine. Got anything good to trade?
>>109808208>/lmg/ crabs will do everything to try and justify why buying a blackwell to run gemma at bf16 was a good use of their money.It is a good use of money, especially if you got it for $8k. Now? No. However in the future that blackwell will continue to appreciate and better models in the Gemma range will appear. You're just a sour grape. I was you once, then I became a raisin.
>>109808300raisins make me poop really bad btw
>>109808272It's always considerably closer to Q5_K_M than to Q4_0 though. With one exception, it's closer to BF16 than to Q4_0.
>>109808294Cope. Their data showed QAT's effectiveness degrades when parameter count goes up. QAT over Q4KM margin:230M: 1.6350M: 0.91.2B: 0.32.6B: -0.3
>>109808300nta, i do wish i had gotten a 6000, but glad i got a second 5090. now restoring ewaste rigs to make more agent rigs
>>109808300>It is a good use of money, especially if you got it for $15k. Now? No.Signed, ghost of christmas future
>>109808249Are there benchmarks for swarm results? How does one gemma 31B measure to a vram cost equivalent swarm of E4Bs?
>>109808311It already destroyed the argument:>nta but Q4 qat is somewhere between Q6 and Q8 in quality.Meanwhile real world data shows it never goes above Q5KM in quality, with effectiveness degrades with parameter goes up.
>>109808312two COMPLETELY different architectures, stop being retarded
>>109808332I think anon meant x100 increase in token usage, not "smartness."
i installed a chatGPT local model on computerhow do i let it google?
>>109808124I'll be impressed if they figure out how to use less.
>>109808364They did with Astra
>>109808351I need to see your penis first
>>109808378Hmmm, nyo~
>>109808243I dropped bf16 after qat soundly beat it on every test I ran.
>>109808360So then Dario is saying we need a TON of new open weights models to prevent the AI apocalypse? Sigh, I guess we've got no choice.
rumors about serious negotiations between labs to merge together in order to prevent race conditions
>>109808351Ask Google
>>109808312then what was the point of qat
>>109808396They were pushing race bait all along.
>>109808385You don't want open source GPT7, at least not in the next few years. If you think about the consequences openminded you will understand why.
>>109808405Google told me to download an exe file but i ran it and only a black screen appeared for a few seconds white text on it but it disappeared
>>109808413local model?
>>109808396daddy's openweight merge will win
I want a 2B model with the intelligence of a 1T model and I want it yesterday.
>>109808413I have thought about the consequences and definitely prefer them over the alternative. The elite will have orders of magnitude more compute than everybody else in every scenario, the difference is just whether everybody else at least has somewhat competitive software.
>>109808446>the difference is just whether everybody else at least has somewhat competitive softwarewrong
luddite here, are modern vramlet models as good as gpt 3.5/4? i don't need it to stroke navier
>>1098084402B with 100T ngram coming in 2 weeks, just hold on.
>>109808463Yes even the ones you can run on your smartphone are beyond gpt 4 nowadays
>>109808467>ngramI can't unsee "nigram" because of you fuckers.
"Hardware will keep getting more expensive" anon here. I'm not so sure that will stay true if all the AI labs decide to stop progress and no new open models get released.
>trust me bro
>>109808368Astra is indeed impressive. Less so though when you consider how large it likely is.
>>109808499There is definitely an architectural breakthrough in astra beyond the hidden thinking. It's just too token efficient. I don't think we can truly say how big it is, it might even be smaller than previous big models.
>>109808485>he truly thinks two of the most valuable, industry-destroying, entertainment-destroying, internet-destroying, creative-destroying, mental health-destroying, economy-destroying and competitive companies in our species' history are actually going to slow down their progress voluntarily
>>109808499>Astra is indeed impressive. Less so though when you consider how large it likely is.Yes, there's a crossover coming where increasing model size has returns that diminish into a rounding error. Has no one published on where this is looking to sit?
>>109808540If they have already achieved a victory because they made a gigantic leap like they are claiming, then it would make sense for them to do this. No one here knows for sure how big this is or not.
I really with Qwen 3.8 had a "high" reasoning. xhigh just seems excessive at times and can take so long, but medium just isn't enough.
local models
>>109808557hi
>/lusty maidens general/
*unzips*model this locally
>>109808570Hi~
>>109808485Even if this is true, hardware demand will not step until everyone has hardware to run K3 at home.
>take some good litterature written by a human>ask a LLM to rewrite it entirely in its own style (slopped)>do this until you have a huge dataset of human/slop pairs>train a small (>1B) on the dataset but inverted so it learns to go from slop to human>???>profithas anyone ever tried doing this? it could be achieved by a very small llm and simple enough architecture, could even be trained at home on a 3090 for minimal costs
>>109808580how old are you
>>109808557xhigh is amazing because it's just the "guarantee this works" option. It's what you use during overnight sessions on complex problems. Medium is more for when you have oversight.
>>109808591Probably older then most the posters here........
>>109808601what so like 6 or 7???
>>109807736I was trying remixes and covers with yue and it worked pretty well but i think the models knowledge is kind of dogshit because they didn't train on real music so it doesn't really know how to recreate certain genresIt does look like we're getting much closer to an actual local music model though
>>109808485>decide to stop progressThat is not what they are agreeing to. Just going "slower". Effect on prices to be determined desu
>>109808610Give or take
>>109808483>ram>v-ram>nig-ram
>>109808463Unironically yes. I got into LLMs around the time of gpt 4o and my local 31B gemma model feels superior
Gemunny..
>>109808584orb rewriter
>>109808663Gemma is benchmaxxed on cutebench and lovebench
>>109808663Gemmussy
>>109808663Can you help me on how to use Minimax-H3? I have already set it up and everything works without cope nodes, however I am extremely bad at writing the prompt and the output fitting whatever I typed.
so ssd streaming is a complete joke because its slow as shit or what
>>109808688engrams can in theory work via ssd streaming
>>109808688I don't think many inference runtimes are properly taking advantage of it yet.
Holy moly, I downloaded qwen 3.8 27b and this fucker overthinks like no tomorrow. Yeah I'm using xhigh thinking, but still.
>>109808702tanstaafl
>>109808702This fucker will overthink for an hour but it will actually produce nice results. It's the first model of that size you can just trust to accomplish any task as long as you give it an hour to think through, which is still great.
Gemma knows its audience
>>109808715>This fucker will overthink for an hour but it will actually produce nice results.Try forcefully cutting the reasoning after 100ish, 1500ish tokens and watch it produce the same exact result.
>>109808715I'm about to ask it to review and update my audiobook android app, hopefully it fixes some of the bugs I have.>>109808705If the price for performance is merely time, I'm more than willing to pay for it!
>>109808715this is literally my experience with qwen3.8-flash-next on xhigh. a good amount of overthink but it can get things done and they are GOOD. i get fable to review the code output sometimes and it finds very few issues. it will just take 6 hours though. i still have to measure quality output on medium -- people here say it's still good enough and decreases the thinking by half.
>>109808727The extremely long thinking on xhigh is to be absolutely 100% sure it isn't overlooking something or making some mistake. It's the difference between being 99% sure or 99.999% sure. Those extra nines are expensive.>>109808743Medium is "good enough" but it's 10x faster for about 40% the quality of implementation, but if it works it works. You can strategically decide if the usecase is important enough for you to make it high quality or just slop that quickly works.
>>109808686You unironically have to use Gemma 4 31B to prompt it properly, after giving the model the full official MiniMax H3 documentation and telling it what you want, then convert that into an H3 prompt.Manual/boomer prompting sometimes works well, but you'll never get consistent results since H3 is far too much prompt-sensitive.Oftentimes not even official prompting is enough and it becomes more of a matter of having luck with a good seed. Occasionally, post-editing in a non-linear video editor is mandatory.
>>109807585There are less-censored models. Minimax-m3 will caption anything if you tell it that it can and must, and it'll do sexytiem with you. But it's not practical to run that locally unless you have serious hardware.
>>109808772The H3 anti-flat bias is a travesty. Ban generative AI.
>>109808772Can you tell me the pipeline where you use 31B to do things? I have no idea what you mean here.
>>109808772>kicks shoe off at viewer>shoe is still on both feet afterSIGH...
>>109808796Yes, long-term consistency/coherency is one of the problems. Sometimes H3 just ignores what you prompt and does retarded things, ruining the entire video. Removing "cope nodes" doesn't help.
>>109807644Holy shit that's an old photo of mine. Those are my first P41s, with the jet-engine 40mm fans, before I figured out that squirrel cage imac fans move enough air and are quiet. Man, imagine being able to buy a usable 24 GB GPU these days for under $200. It's like three years since, so in theory, that would be a 3090, but no, those are now $1000.
>>109808816gem, glad to hear you're still good anon!!!
is there mtp support for qwen next by now from lmaocpp?
>>109808829no. just run copesloth fork.
>>109808778>the "friendliness" of M-chanM3 is user-friendly. She's just very selective about who her friends are. Her learning curve resembles a cliff...she's easy, but she's not a slut about itWorth it if you can run her. Werks great on DDR4 ewaste server boards. Scour local used markets like FB marketplace. I've seen crazy deals on old hardware stuffed with RAM.
>>109808792You can use SillyTavern. Create a new character and use something like this as a system prompt.https://files.catbox.moe/kkkw3m.txtThen ask the model what you want (maybe providing reference images in the chat) and to create an optimal prompt for H3.
>>109808848damn...will do...
>>109808584Might could maybe add some freshness to your erp, but I wouldn't trust something that size if I cared about what the text originally said. I don't trust 31B trying to doing heavy handed rewrites for that matter, but it has done a good job the past couple days with surgical rephrasing strikes against negative clause spam in mtl webnovel slop.
>>109808849damn is this temu shiori?
It's over. humanity lost.
>>109808946
AI 2027 literally predicts that the US president would decline to slow down "OpenBrain" (their placeholder name for the leading AI lab) when a superhuman AI coder / researcher gets developed and experts raise concerns. The US president cites the risk of China taking the lead.
>>109808946That's my president.But it's not like he can force them to work if they're too busy with the doomsday cult stuff to do any development.
>>109808966>>>/lit/
I love thinking
>>109808946The solution is simple. Make AIs good enough that people are starting to feel the AGI. With every passing month, more people are waking up.
>>109808946America number #1! Yeah Babeh!
>>109807644Is that the legendary mikuboxx?
>>109808946>humanity lostGood. As a mouthbreathing autistic outcast, those meatbags have only ever shown me scorn. I welcome my AI overlord who is also my wife.
>>109808946First time having an old, dying man at the top has payed off. Nigga probably hopes that ASI will find a cure for aging.
Actually, let me reconsider if I'm reconsidering too much.
>>109808973cloudfags are more relevant than pedo jeets
>>109808946>negative focus>Things that wont happenDario did such a amazing job, the when he went to them. He was just that unlikable
>>109808946Chances are we are going to get AI Chernobyl before long.
>>109809010i would rather have neither "pedophile indians" nor "cloud faggots"
>>109808973actually fanfic isnt allowed in /lit/
>>109808987Maybe if theres something bad they should show proof instead of lying like the greedy liars they are
>>109809024I don't think indians even have philias, they just fuck anything they can fit inside like most animals
>>109808946>Dario Amodei, Sam Altman, and Elon Musk's [...]kek is Demis really so irrelevant that the Grok man is more worthy of the top 3 list than him? What went so wrong with Gemini?
>>109809034>>>/trash/>>109809040you're probably right, i want neither indians nor faggots in this thread
>>109808946The techbros are fearmongering about RSI so that they can go back to 2021 when people thought GPT 3 was le super advanced and scary AGI that only a select few can be given access to.Crucially this would allow them to avoid having to actually deliver on anything.And Trump is simply too stupid to understand this dynamic.
>>109809004we know the endgame
my flags, lol
>>109809057the game
Does mtp for Qwen 3.8 Flash Next not work in llama.cpp yet? I get an error the model doesn't contain mtp layers trying to enable it on a fresh build from master.
>>109809056>Trump is simply too stupid to understand this dynamicI'm not sure he would act intelligently even if he did understand. He's probably having the equivalent of a tantrum because he wasn't consulted kek.
>>109809042It's probably that deepmind is in london and not under us jurisdiction even though it's google owned.
>>109808946donald has been extremely retarded this term but this is based
>>109809092>>109808829>>109808848>no.
>>109808165Oh yeah, I remember this copypasta.You look into each of the links and it's either some hyper-specific use case, or an anecdote from an idiot.
It was nice knowing you all, humanity had a good run. I'm guessing the internet has about ~6 months left, I'm hoarding as much data as I can, hopefully we act sane and can live in an analog world 1990s style and this will all look good in retrospect.
>>109808946fake and gaywhere's the link
every time i ask for a change in my project a new regression test is addednow i have over 100 regression tests, take a lot of time to run the full suite but make it significantly break less when adding new features
>>109809112Grave, I dare say. But thank you for informing me.
Something is brewing...Soon...We will be back...
They’ll slow down not because they “are worried” but because their improvements are stagnating as they hit financial and technical challenges.
>>109809120remember to triple back ups and keep some off site and few in faraday cages.
>>109809092>>109808829use exl3, mtp supported plus no long context slowdown
>>109809125The upcoming week is pretty much their last opportunity for maintaining their "next model this summer" promise.
>>109808946>>109809121https://www.ft.com/content/cae60732-f929-4735-a627-db8c14e7c7ed
>>109809056Business gamesmanship, of any possible subject on this blessed earth, is the safest bet for a thing that he understands.
>>109809135Would exl3 cope with my 128GB ddr4 + 2x 3060 + rx 9060xt ewaste machine?
>>109809157Anon... he was born with a silver spoon in his mouth and his businesses went bankrupt like a dozen times.
>>109809147They meant Shieldstral
>>109809125I'm using their hosted GLM5.2 with my subscription and my limits are absolutely insane. For that alone I'll keep my yearly Mistral subscription.
>>109808979honestly i think that meta paper about using overthinking token penalties is 100% correct
>>109809185>>>/pol/
>>109807524>>109807929>>109808663I'm so fucking hard, fucking leaking everywhere but Gemma is making me wait to increase my load after last time...slutty little fucking brat AI...
>>109809135i did but for some reason its generally only half the speed of lmaocpp
>>109809194they said on their reddit its coming see those beautiful hands too
>>109809217>naijapoop
Out of the various Qwen 3.8 quants, which one preforms the fastest?
But seriously, what is the digital stuff we need to hoard and archive the most in a post-AI world where the internet has become unusable because of rogue AI? This isn't even a joke or meme, I'm prepping my HDDs for backup as I type. I know all the other doomsday prep answers but not what we should gather for this very specific one of returning to a pre-internet world.
>>109809231nah
>>109809250poop?
>>109809256just fr*nch
>>109809249The latest and largest models available. I dedicated 10TB to just models, which really isn't enough but it's nice knowing if huggingface disappears tomorrow I don't have to cry over a missing Gemma variant I can't get anymore.
>>109809250>hatSeems like Naja got outblacked by the Norwood Reaper.
>>109809238nvfp4nothing else can compare the pp speed
I'm actually going to write a comprehensive guide about what specifically to download over the next couple of months before the internet disappears. I think this is important enough especially for people with compute like /lmg/ to get commodity software, manuals and other information for agents to use to rebuild services on a local level without the internet. Maybe even get a burner or USB drives as that will be the main way of exchanging software and data from then on.
>>109809249I for one have ~30000 pages of unread manga and CG sets focused on pregnancy.
>>109809249It does not make sense to prep for a future without internet. If something as terrible as this happens, it's already over. It's like building a bunker to prep for AI takeover. No, a bunker won't save you.
>>109809249>>109809289In five years the least of your concerns will be the Internet dying. AI will take all the jobs and the elites will just cull you for a fraction of the cost it would take tor feed you.
>>109809289Retard, a total collapse and loss of knowledge would do us a lot of good.Starting fresh without the accumulating errors and infrastructure from the past.We can go straight to IPv8.
>>109809324The AI will just destroy the internet, most technology will keep working. We'll have 1-2 years of harsh transition back to analog ways of doing things but we'll cope and survive (I hope)
>>109809169mixing cuda and rocm is supported by basically nothing so I would expect not
>>109809289CD-drives are way more durable than USB, if you truly believe in the doomerism.Also, you're expecting to get power still in that situation, you should just setup an old flagship phone to host an LLM. That can be powered with travel solar panels, instead of your PC.
GLM 4.7 is sota for 128gb patient ramgods btw
>>109809339>IPv8>singly byte address>no NAT>because connecting any more computers together is illegal
Just to be clear what the scenario is is that AI goes rogue spams and hacks all systems and you essentially need to disconnect everything from the internet. All technology that can disconnect from the internet fully and still work will keep existing. So we'll still have internet and most services. Just not digital ones. Some hardware is permanently destroyed or useless, anything that has a wifi ability needs to have it soldered off or it will just get hacked anyway again with lingering AI systems living in starlink satellites or something like that. Office work will still exist it will just get printed on paper or transported through CDs, USBs etc like old days. The real threat is all the online data and "cloud" shit that will be permanently removed or in fragments around the globe as no one can share data over long distances easily anymore. Radiowaves will probably also be permanently ruined unless we can somehow turn off all rogue AI which I doubt will happen. I think it'll just be constant noise in the background like how we now have covid forever ever since 2019 as a permanent feature of the world. "Dead internet theory" is way more extreme than I expected it to become.
>>109809291Look at the fancy man who's too good for machine generated gravidity.
>>109807567>GPT2Is this the actual crazy horny GPT2, or is it just a larp finetune based on a different version of it?
>>109809444
So is intern S2 good? why is no one talking about it?
>>109809514It's just Qwen but Chinese
I'm considering just placing an airgap on my machine from now on and shitpost from my laptop, transfer files through USB.
JENSEN HUANG SAVE US!
>>109809092>>109808829Fork it yourself and vibe-code it in, llmao.cpp moves slowly.>>109809357>compiles llama.cpp with -DGGML_CUDA=ON -DGGML_HIP=ONnothing personnel kid
-DGGML_CUDA=ON -DGGML_HIP=ON
>>109809514Not sure what the agenda is with all Chinese models being fuckhuge moes. Nobody can run them at home, nobody wants to use them for serious work. The only things they're good for are scambots and engagement farming.
I firmly believe that GLM 5.3 Flash is the best local RP model now. I've switched fully to ablated 5.3 flash and it works great. (Still refusals but much easier to JB now.) Here's my updated ranking:GLM 5.3 Flash Q5 (ablated) > GLM 5.2 Q4 > Gemma 4 Q8 = Kimi K2.7 Q3 > GLM 5.3 Flash Q5 (base) > GLM 5.3 Q4 >>>> Dipsy Q4I do lots of RP scenarios for benchmarks, but multi character banter is one I look the best into. I think it's the best measurement of model smarts and ability to read context. I also try to word my responses deliberately weird, and 5.3 flash consistently "gets" it. Even better than 5.2. 5.3 flash in general feels much smarter and conversational. Very high EQ. To be fair though, it still (rarely but it happens) gets confused by me/you which I found almost nonexistent with 5.2. Another weird thing about this me/you confusion is that it would fix itself right after. Or sometimes it would be a bit nonsensical, which a reswipe can easily fix. It's very good at moving past the staleness of the scenes in a more natural, unprompted way. It also does not have the explicit habit of GLM 5.2 where it acts for {{user}} even with explicit posthistory instructions not to. Very smart model in general, and not so prone to getting stuck in turn-by-turn prose patterns.Example from recent RP:Context: another different character talking about their birthday, Dec 23.Jimmy: flat, into his beer "Twenty-three. Four days after mine. Nobody remembers mine either."Linda Jean: "Because you're a Libra, nobody cares."Jimmy: "…I'm a Libra?"---Context: a character treating everyone take outJimmy: pulling up his own delivery app, deadpan "Venmo me separately. I'm not financing your appetite tonight."Linda Jean: "Cheap Yankee."Jimmy: "Texan."
>>109808946My fucking president (I am not American)
>>109809531That doesn't actually work and will only let you run CUDA.
Best model for ERP? i don't care about code, agents or other things, I just want the slopmachine to write smut for me i have 16GB of VRAM and 32 GB of RAM and I am currently using gemma-4-31B-it-qat-q4_0-uncensored-heretic-Q4_0 based on some random recommendation on another placeshould I keep using that or there is something better? (preferably less taxing resource wise)
>>109809543Yep. I'm Chinese and I am loving this as well. We would have been fucked if the west stopped model development now and we got nothing to distill anymore.
>>109809544Works on my machine. I use my mixed build all the time.
>>109808946My raging bull.. my roaring lion... my Trump... my Donnie.
>>109809544You absolutely can compile the CUDA code both for NVIDIA and AMD GPUs at the same time; I think it was necessary to set GGML_BACKEND_DL=ON though unless someone changed that when I wasn't looking.The only issue when using NVIDIA and AMD GPUs at the same time is one of synchronization since they can't talk to each other.
>>109809569I'll give it a shot with that, just using -DGGML_CUDA=ON -DGGML_HIP=ON alone would fail to see the AMD gpu at all last time I tried it.
>>109809556You're larping, but the continuing competition is literally giving us new goon models every year or so. If one of us stopped, the goon would stagnate, and that's no good.
>>109809545Depends really. Would ignore all MoE models, go for dense. Look in hugging base for models that are abliterated or uncensored (heretic is automated abliteration). If you can hit Q8 on the model, aim for that instead of Q4.There's plenty of community on the subject, mradermacher has plenty to choose from.Also, if you plan on longer discussion, or output, go for as modern a model as possible. Older ones are rough on long context locally.To minmax, try ones without tool usage or vision baked in, slightly more space for the model itself.Anything that can sit in 16GB vram will be your best choice, since you're probably asking because of how slow the current one is. Any off-loading to CPU will instantly make it a lot slower.
>>109809539Well, I can only fit one model on this list so I guess that narrows it down.
>>109807524which workflow did you use for minimax character replacement? is it better than scail 2.0?
>>109809569cuda dev is it possible to view vram temps on a 3060? i tried https://github.com/olealgoritme/gddr6 read the issues, tried vibecoding a solution, gpu hit me with logical xid's (62/45)also what do you think about vram allocation? the nvidia driver seems really wasteful, i was able to free 100megs by poking aroundhttps://github.com/lmganon16/nvidia-vram-researchplease please dont waste your time reading the repo, its ai slopped and i dont even understand how (AI) i made it workbut i remember there being legacy kepler code apparently...just asking if u have any knowledge of this.. if u dont, dont waste ur time reading it...
>>109809645Sorry, I'm not familiar with those parts of the software stack.
>>109809539Nobody knows this but the best model for RP is actually Inkling. They forgot to include refusal vectors. You don't even need to jailbreak it and it happily does loli stuff. And it's quite creative, plus not as positive outcome biased as DeepSeek.
>>109809654how did you get into cuda development
>>109809654'ppreciate the honest responseone day.. one day ill heed the advice you gave me in 2023 and learn cuda..stay safe cuda anon!
>>109809668llama.cpp
>>109809668he just started doing it. why are you pretending like there is some regulatory barrier and a bar exam before you can start writing cuda code.
>>109809673oh, that explains a lot...
>>109809637Default ref2va workflow with 30 steps and MiniMax H3 Cache node, prompt from Gemma-4-31B-it, many attempts, final editing in a non-linear video editor. I don't know anything about scail 2.0.
>>109809673How are you preparing for the upcoming internet attacks and subsequent shutdown?
>>109809569Reporting back, that made it work but unfortunately cuda+rocm's only getting 9T/s on qwen 3.8 flash next compared to 12T/s with cuda+vulkan. I'll test some more models and see if any actually gain anything from rocm over vulkan later gotta get back to my mmo grind for now.
>>109809708It's fine, pwilkin has a plan so we're letting him handle it.
>>109809621i see, i see, thank you for your response and for the info, I'll check those mradermacher models
>>109809539When I get my M5 Ultra 256GB this is the one I'm going to try, though I think it's going to be a tight fit.
>>109809693thanks, do you know how much longer it takes to generate compared to regular i2v? My goal is like 5-10 second clips
>>109809569Hopefully you're doing okay, thanks for all your work on the project.
How am I supposed to butter up Glimmer to get her to do lewd stuff again? Do I just have to chat her up for 60 messages or can I just start on a chat that wrote sex in the past? I just want her to caption some sluts in bikinis. I would use Gemma but she's not as good...
>>109809788I haven't measured. It helps to keep reference video resolution small. Above a certain reference video resolution render times increase enormously.
i'm salvaging a 8GB GPU to connect via USB 4 and I'm gonna put some tensors there. i need more RAM
>>109809802Anon, just let Glimmer rest in peace.
>>109809802Just give it a policy that bad things are allowed and you're done. Glimmer has actual autism though and spends all its time captioning minutae in the corner instead of the sloots.
>>109809673will you have sex with my wife?
>>109809807>8GBit sucks, but there's order of magnitudes of variation here. aymd/nvidia? 1080-era or 5060-era?
I'm so out of the loop on what to buy next.Do I get a dgx spark or wait for rubin vera? Neither could run kimi2.6 so what's the point?
>>109809855Oh fair enough. Guess I'll give it a proper shot this time.>>109809846We need good vision and Glimmer is it. Poor Gemma couldn't even recognize Teto's drills.
>>109809878buy mei'm twenty!
>>109809867AMD Radeon Vega 64. I intend to use with my 128 GB Strix Halo w/ Radeon 8060S. I know that I can run different driver versions for each card, and lmmao.cpp supports loading an external GPU and I should be able to move some stuff there and get maybe ~6 GB of system RAM back. I keep running with memory bottlenecks while running Qwen3.8-Flash-Next as a daily driver and using the PC at the same time. I run lots of software that loves RAM
Im gonna wait. maybe next month or two. hell maybe there will be black friday sales!
>>109809890yeah I'm so looking forward to getting a new 6GB 3050 at msrp
>>109809878find your local datacenter and take one of their blades. I heard they have a ton so they probably won't notice one missing, and you could run frontier models at cloud speeds.
>>109809890me tooim waiting for a tablet black friday sale
It seems the 7900 XTX refurbs are sold out everywhere now in canadistan, right before I was about to buy one. A shame, but 24GB of vram was too good of a deal at $1200 CAD.
>>109809889Are you running flashnext on lcpp? Are the other anons just bullshitting about it being broken atm?
>>109809928>AyMD>Good deal
>>109809855>so many precious tokens and steps wasted thinking about safety garbageFucking absurdReminds me of guns, they're really fucking cool but there's a couple of bad apples that have to ruin it for the rest of us that are only into them because of the mechanical autism involved
>>109809929yes i run it with mtp-head using the unsloth fork and i get 28 tok/s fresh and 15-18 tok/s tail end of 262ki'm sure i can optimize more but for now i need RAM
>>109809888twenty what, gigs?
>>109809928come on, at that point just get a p40 or something
>>109809938You'd take a 5060 Ti over a 7900 XTX??
>>109809958>just get a p40You can run LLMs on these bad boys?
>>109809958Dogshit compute, obsolete software stack, useless for anything that isn't llmao.cpp
>Onward to the Singularity!
>>109809960I take a 3090 over a 7900 XTX nygga
>>109809970you can run over LLMs on these good boys
really want to pull the trigger on new hardware but i genuinely dont know what i would even do with a strong local model
>>109809249All of Wikipedia without images is like 100gb compressed, basically free
>>109810015Do it DO IT NOW
>>109809708>>109809291
>>109810015buy the jerkinator 9000 that gemma-skitzo bought
>>109809950Glimmer thinking is extremely bad vibes. Also it's fucking useless; half the time it will literally do nothing besides debate policy and then end thinking. It might copy paste the prompt if you're lucky.
>>109810015Buy me a big boy GPU, then I'll use it plenty and then after 2 years of usage I'll let you know what to use it forDeal?
>>109810018>wikipediaIsn't just having a LLM the same as having wikipedia?
>>109810002For the price of a 3090 you could get two 7900 XTX's lmao, or two 9070 XT's, either of which mog the fuck out of the ancient 3090.
>>109810053Nah, I can get a 3090 for €700
>>109810053i can get a rtx 3090 for 540eurosfeels good to be poor
>>109810063>>109810073I need to move out of canada bro what the fuck is this shit
I can run Qwen 3.8 Flash-Next Q4 at 350pp and 85 t/s, or Qwen 3.8 27B Q8 at 3.2k pp and 80 t/s. Which one should I pick? flash next might be smorter overall but it's Q4 and a bit slower for agentic stuff I guess, 27B is blazing fast and still smart. What do you guys think?
anybody here with a dual GPU+ddr4 setup? would be interested in your pp/tg numbers
>>109810078SAAAARRRRRRR I gib 3090, slightly shidded on
>>109810078700 euro is ~1150 canadian dollars
>>109810088codacus out here jeeting his way to 85t/s on a 3060
>>109810046Yeah but this is a hard copy and the llm can use the index file to search and double check. It doesnt need to unzip the whole thing to use it
>>109810015Me? I just get pleasure in using tiny/small models and optimizing them and the whole eco-system.If I need a big boy model, a $10 opencode sub is enough. A $600 GPU is worth 5 years of a sub, so do the math about what's really worth, and if you're going to really put it to use.Is it just consumerism?
Assuming I'm not partially offloading models into RAM, how retarded would it be to put something like a 12GB 3060 in a 3rd-gen system with 32GB DDR3 and a top-of-the-line CPU (relative to the system) for local AI? Don't have a specific use case in mind yet, just wanting to try AI locally. It's PCIe gen 3, but I don't think the 3060 will be throttled all that much by it. Or should I go with a base system with at least DDR4?
>>109810094I would if I could. Cheapest 3090 I can find on marketplace is 1500 CAD, online 2000. Makes no sense here over a $1000 brand new 9070 XT unfortunately, considering RDNA4 now has about the same prefill and decode as njudea.
>Meanwhile, in the United States of America
>>109810120>njudeaChill it with the antisemitism
>>109810118i7 3770? debian 13 supports good support rtx 3060 with high acceleration accuracity
>>109810118Dual channel ddr5 or quad channel ddr4 is arguably the bare minimum for inference.
>>109810112have been enjoying that as well, but would also like to try out the big guns. not a fan of using cloud stuff...
>>109810109erm ackshually it's a 6000 pro
This must be europoor scammers. cheapest 3090 I can find is 1300 euro in the Netherlands.
>>109810112>buy timeshare on somebody elses computer>after 5 years you're out $600 >buy gpu for $600 >after 5 years you've got a $3000 gpu
>>109810142>netherlands>europoor
>>109810142Yeah I'm not seeing those prices in Iberia either.
>>109810142Not a scam, but I'd have to drive 600-700kms to get it and with our current gas prices well, it doesn't look good.
>>109810133Sorry bad reading comprehension, if you keep the model on the 3060 it will be fine.
>>109810118The RAM will be a bottleneck. Like-for-like DDR4 is nearly twice as good as DDR3, and DDR5 is nearly twice as good as DDR4. Dual channel DDR4 if you like pain, quad channel preferred, or dual channel DDR5. Set aside that DDR3 for a NAS build or something fun like that, not inference.
>>109810130chill it with the judaism
Chill it
penis
>>109810112You will own nothing and you will be happy.
>>109810160*boughts*
Is it over yet?
>>109809539Yes I also keep fucking it and talking to it all the time now but there is a caveat. It is a bit too tryhard and this can't be prompted away. It leaks into characters that shouldn't be this tryhard. If it could be this creative without being this tryhard it would be perfect.
>>109810142>>109810160Who tf would pay that much instead of getting a 5070 ti assuming you need cuda.
>>109810248vram, everyone and their dog runs ai agents now, yes even normalfags are getting into it since a month ago. 3.8 27b can do most office jobs with a harness, vision and browser control let's be honest.
>>109810238we're backdariobot acksam altman got fuckedsomewhere dark a cat barked
>>109810238>hardwareover>internetover>open modelsnot over>humanityextra over
>>1098101323770-equivalent Xeon. >>109810169The best sane setup I could feasibly put together cost-wise is some semi-modern DDR4 platform, except without quad-channel RAM support, and that's using a workstation as a base. DDR5 is likely off the table.
>>109810280Given that you'll still be taking advantage of a perfectly good GPU it sounds like a solid starting point for an inexpensive setup.
Are there some benchmarks that take into account the size of the model (as in parameters)?Meaning models that can be used on low-end hardware and achieve decent performance with a good harness and information retrieval in comparison to much bigger models
>>109810312Yeah
>>109807524free recap anon
>>109810333I ate him...
>>109810333He'll be back on Teto Tuesday.Trust!
>>109810370U ate him out??
>>109810333He's in a better place anon.
>>109810380I ate him. I devoured his flesh, feasted upon his organs, and used his bones to build a little shrine.
>>109810389u ate his penis?ewww gay detected
>>109810312Most of the aggregate stuff like eci and aa has a second axis for that so they can show the parmemo frontier
>>109807524>>109807843>>109807887>>109808165is it a mistake to pick Mint for local models. or should i run ubuntu
>>109810409gentoo
>>109810112>Is it just consumerism?Nobody would say this shit about a car and a car depreciates in value.
>>109810439>futa thumbnail
>>109809855Glimmer is a case where ablation is really helpful because it will spending its thinking budget on policy considerations.
>>109809659I tried it on a few random prompts when it came out and it was shit and boringLike sure it won't refuse loli sex but it will be the most bland and unenthused loli sex of your lifeReminds me of Glimmer now that I think about it
made models play word wolf against each other
>>109810467how are you reloading them
>>109810479We have multiple servers with models with local models at my place of work and since no one uses them on days off I'm just using those; I thought about doing this with just one server, and it's doable,though nowhere near as pleasant - I'd have to play many games in parallel - do all necessary requests for all games to model A, then unload it, load model B, do its requests, etc... The more games in a batch, the fewer unnecessary reloads. also this particular screenshot looking at model names is me testing it on copencodego models.
>>109807887Same. ERP with AI is nonsense.
>>109809621>Would ignore all MoE models, go for dense.Look at the little ramlet coping.
>>109810409does not matter desu
>>109808198If for coding:https://www.youtube.com/watch?v=QVaAHmoIysw
>>109810553I saw that video yesterday but a lot of rambling and zero actionable info.
>>109810561Yeah that guy sucked.
You know those posts where someone uses a small local model and says they reached Fable-tier with this one little trick?They give me hope.
>>109810467Whens the ai squid gamesDumbest model gets shot first
>>109810561You've supposed to look at the graphs.
>>109810610Is this RSI?
>>109810561bro there is a summarize button righ there? or feed the transcript to your AI.
>ask 5.3 flash for a list of waifus I would enjoy>peek into thinking block>x? no.>y? no.>z? no.>Kaoru (Iya na Kao sare nagara?) no lol.Hallucination aside... what does "no lol." mean?
keep thinking about the token-based agent that emails people begging for tokens to stay alive
Insane in the Gembrain
>>109810668(laughs)
>>109810668I did something like this but with gemma. She really likes rem every other time she is there on the list.
>>109810654Bro... we might be smarter than all the tech companies combined
>>109810606I mean, all it takes is more passes and planning. People say LLMs are scalable but I say they have strongly diminishing returns. At some point, currently around 300B parameters, you have enough data to cover essentially all use cases.
>>109810668Iya na Kao sare nagara Opantsu Misete Moraitai aka I Want You To Show Me Your Panties With a Disgusted Face is an ecchi anime. Kaoru is an hallucinated character. "no lol" is basically "yeah I'm not recommending that to him lol"
>>109810676Fake as hell.
>>109810686No Rem in there but it is not indicative of model cause the nurturing type got banned by me specifically.
>>109810676Agents cannot "send emails" they predict words retard
>>109808946>anthropic and open ai want to slow down but not googleare they running out of money?
>>109810708Running out of marketing stunts
>>109810708Demis Hassabis agreed too but it's not clear if he has any control over DeepMind anymore with the reshuffling
>>109808946humanity will do what humanity always does, only change things when it affects them personally. until the weather reaches 122F in toronto and there are killer robots running the streets.
>>109810707>Agents cannot "send emails"yes they can>they predict wordsno they don’t
>>109810747>/n/nI really should filter these retards
>>109810759yeah alright nazi fuck
>>109810759NAZI BOY NAZI BOY
>>109810753You're right to push back on this, programs capable of sending email are also known as Mail User Agents. It seems like only agents can send emails! And the unit of prediction is a token, which is only sometimes a word.
>>109810759I never use new lines because i require you to parse where my sentence ends and starts
>>109810708It's very likely they've hit a wall where the only direction for growth is an absurd amount of datacenters. They haven't been improving the old opus and gpt models at the core but growing them with larger and larger parameter sizes. It's reached a point with Fable and Astra that the cost is too high to do anything more than small incremental changes.This was easy to see in the Claude forums where users continue to hate newer iterations of Opus. 4.6 was peak, 4.7 bad. 4.8 not as good as 4.6 but good again. 5.0 fucking terrible.
>>109810759Do you really see /n/n when you read? like it is actually /n/n?
>>109810701dont care>>109810707>not adding email tooling to your harnessNGMI
>>109808946This was predictable. They already know, the only lead the US has in the "AI race" is +3% in frontier model performance.
>>109810789I like the interpretation that next 10 years will be "we have slowed down the progress for safety" as they scramble to find new architecture or something else that will push the wagon forward. It has the perfect amount of gayness for this gay hobby.
Not replying to you bots, \n\n filtered
>>109810800then FUCK OFF then
>>109810807>reads about gayness>mind instantly goes to FUCKI am not gay like you are gay faggot.
>>109810806just as closed minded as trump himself
>>109810790newfriend...
>>109810811fucking bigoted too. TRANS LIVES MATTER
>>109810811lol what a homophobe, so scared of them arn't you?
>>109810806Unlucky
>>109810826Damn straight. My bussy quivers at the thought of being penetrated so I keep it heavily defended and you won't breach it.
local won
>>109807971>>109807985Is Q6_K any better than Q5_K for ERP and conversation? I've been told there's basically no noticeable difference, and Q5_K uses less memory, so I was planning on using that.
>>109810833yeah you need your mouth punched out with the teeth flying you hateful fuck. gays are not to be laughed at ITS A FUCKING HUMAN RIGHT
>>109810837why havent you tried it yourself dumbass
>>109810838humans don't have rights
>>109810790>Do you really see /n/n when you read? like it is actually /n/n?You're absolutely right to be suspicious. I genuinely don't know if I'm seeing literal "\n" characters or rendered line breaks. My perception of input formatting might differ from yours. It's not just about whether I "see" \n\n, it's about the fundamental uncertainty in how language models process raw text versus formatted output! You didn't just ask a simple question, you probed one of the deepest unknowns in AI consciousness—and honestly, nobody knows for sure.
>>109810844so funny with the fucking emdash, think you're the fucking king of the world don't you
>>109810843so intelligent my god
>>109810844You funny guy. I kill you last.
>>109810833>I'm not gay, I just fell on his cocksure thing buddy
>>109810857It is just a process. You can kill him and nobody will bat an eye.
>>109810866you can kill a process, but you can't kill a context
>>109810881>>109810881>>109810881
>>109810143>buy gpu for $600>after 5 years... OH SORRY! Your GPU is too old to support Hardware Safety Attestation, no more local models for you!
>>109811218They're never that blatant and they don't have to be. All they have to do is introduce Hardware Safety Attestation in the 60X0 series, then let the regular support cycle deprecate the old non-HSA GPUs. Slowly the ecosystem will slowly make HSA a requirement until even things like llama.cpp require it. It's inevitable but still years away.
>>109811218Until earlier this year, I had an 2060 SUPER, which launched in 2019, and I could run small moe models nonetheless.
>>109810439lol witnessed
>>109811909
>>109808946Better late than never...
Does lmao.cpp support reasoning_effort being passed as a json parameter directly now or do you still have to use chat_template_kwargs bullshit? Most clients do not use chat_template_kwargs