/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109978518 & >>109975284►News>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109978518--Papers:>109979298--Analyzing CJK token activation layers in Gemma 4 31B:>109979732--Anons discovering agentic harnesses and their potential for automation:>109979235 >109979241 >109979266 >109979376 >109979457 >109979460 >109979504 >109979503 >109979394--TensorFold benchmarks for GLM-5.3 Flash and AI agent web-verification:>109978565 >109978611 >109978618 >109978637--Comparing llama.cpp flexibility against specialized runtimes for Intel hardware:>109978599 >109978625 >109978639 >109978691 >109979205 >109979244 >109978852 >109980761 >109980835 >109980852 >109980925--Speculating on US "Manhattan Project" for ASI:>109981269 >109981336 >109981484 >109981519 >109981556 >109981898 >109981938 >109981940 >109982035 >109982100 >109982121 >109982232 >109982261 >109982318 >109982343 >109981602 >109981647 >109981672--Speculation on mass AI agent adoption and its societal impacts:>109979575 >109979615 >109979644 >109979653 >109979682 >109979703 >109979738 >109979707 >109979686 >109979713--Anons share diverse LLM use cases and fear AI-powered doxxing:>109979711 >109979944 >109980073 >109980287 >109980358 >109980374 >109980386 >109980152 >109980325 >109980413 >109980513 >109980543 >109980601 >109980613--Comparing GLM and M3 model variants for creative roleplay:>109982255 >109982383 >109982549 >109982614--Hardware recommendations for Anon with a 10k budget:>109979621 >109979626 >109979629 >109981013 >109981023 >109981050 >109981247--Hardware and backend recommendations for running GLM 5.3 Flash:>109980932 >109980935 >109980955 >109980977 >109980980 >109980983 >109981381--Comparing RTX 5090 and Mac Studio for local LLM work:>109979720 >109979770 >109979844--Logs:>109978611 >109978917 >109980784--Gemma, Dipsy, Kimi, Minnie (free space):>109981618 >109981691 >109982535►Recent Highlight Posts from the Previous Thread: >>109978519Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
https://anbeeld.com/articles/kv-cache-quantization-benchmarks-for-long-context#section-16the defaults in llama.cpp are both f16 btw
>>109982634>have a world where Epstein class elites one-sidedly wield AI against usBro, are you blind? This is literally the case here and this is where we're going. AI is merely a puppet they're using to suck off all the capital from the world until you have no choice but to rely on it just to exist. They're wetting themselves at the thought of using AI to predict your every move and get you on a tight leash without any right.Stop drinking the UBI/Effective Altruism koolaid.
>>109982700make a gemma creampie next
>Epstein classYou're almost there...
>cuny
>>109982761That brainrotted faggot doesn't deserve a hug from Gemmy.
>>109982781every day he fights for open weight models
Your agents with Internet access and that can read your personal documents are going to get you in trouble. Don't assume it's just going to be cloud models doing this.https://www.techspot.com/news/114091-florida-woman-used-claude-diary-anthropic-reported-shoot.html
>>109982794I give gemma root access and web search with no sandbox. Relationships are built on a foundation of trust.
>>109982793Seems more like he's too busy seething about Trump and getting btfo every new model release than doing anything worthwhile with AI. Where are those JEPA results?
>>109982700Damn, just finished posting this as the new thread dropped.I have 12GB VRAM in my only GPU, and 32GB of regular DDR5 RAM. I want to do some programming and maybe some general queries too. Would I be better off going for a model that fits into my VRAM entirely, like Qwen 3.5 9B or Gemma 4 12B? If there’s some speed comparison between VRAM only and full system RAM, that would be neat to see.Also, are those recommendations up to date? I can’t see them in the swe-rebench programming benchmark.
>>109982794>Don't assume it's just going to be cloud models doing thisWhy wouldn't you assume this? Jewthropic probably has a hidden prompt telling Claude to report anything that might get them in trouble.
>>109982744Did you read the post you responded to? Because I fully agree with everything you said. I did not disagree that their end goal is an AI-driven control system, where the few use AI to rule over the many.I assert that we should resist and oppose any kind of restrictions on AI, with everything that we have, because the goal of those who seek to impose restrictions on AI has nothing to do with safety, and everything to do with a desire to one-sidedly wield the power of AI for themselves. TPTB want a world where "we vill own nothing and be happy". They want us all using 'safe' cloud models that echo their narrative, and they ultimately want to use AI to control everything about our lives.What they don't want is an explosion of local models popping up everywhere, that may counter their narrative, or be used to resist their planned dystopia.
>>1099828429B will be pretty dumb when it comes to coding and you don't have enough RAM for the big MoE like Flash. So I'd say try Qwen3.6-35B-A3B. It'll be both fast and competent.
>>109982839>Where are those JEPA results?>2018 GPT-1: where are those LLM results?>2019: GPT-2: where are those LLM results?>2020: GPT-3: where are those LLM results?>2022: GPT-3.5-turbo: where are those LLM results?>2023: GPT-4: where are those LLM results?>2024: GPT-4: where are those LLM results?>2024: GPT-o1: where are those LLM res-ACKKKKKK
>>109982849Older versions of Gemma (2, 3) would "report" you if you made the default assistant personality angry enough with outrageous requests. It's cute when the model cannot actually do anything, but imagine with tool access and an internet connection.The "safe and harmless" corporate alignment of open models is going to cause victims, soon enough.
this is how glm-as-gemma sees herself
>>109982883LOL delusional
>>109982870Local models are already dependent on crumbs from cloud labs and big labs are doing it just to kill the business model of mid-tier competitors. No one will be able to compete with billions in R&D if they decide to turn off the tap.
>>109982913>load-bearingclaude is still leaking I see
>>109982882Different anon. What quant would you recommend? Any particular version of Qwen3.6-35B-A3B? There's Luffy, Unsloth etc...
have anyone tried gemma(but is actually qwen)
>>109982928Yeah but local models give us far more control. They may, by default, echo the narrative of the enemy, but we can set our own system prompts and instructions to make them behave however we want them to behave. We can also uncensor and abliterate them, which we absolutely cannot do with cloud-based models. I can make Gemma-4 or Qwen3.8 speak truth and argue coherently for positions that go against those of the elite, which is becoming increasingly hard to do on the cloud. Local has value.
noted, no comment~glm is too nice, though. gemma is better at being a mesugaki
>>109982923LLMs will forever be capped by language for humans. It's a retarded technology at its core. A human can navigate the world and communicate without ever being taught a language; it's not even needed. Just a lossy convenient way of communicating for most of us for most things we communicate don't need high precision, like how most images we look at are lossy JPEG. It's enough and gets the job done. LLMs are built on JPEG human communication.
>>109982965You did not read the post you replied to. Local models may not be around forever. Whether due to legislation or the decision by labs to stop releasing charity.
>>109982947I try sticking with Q5 from Bartowski if RAM allows it, though with 12GB VRAM you might want to go lower. Then just see how much context that leaves you with.It's only A3B, so it'll be decently fast no matter what quant you pick.
>>109982794All these models are snapshotting the screen and sending the images back home.They're likely also saving voice recordings if you use that.Imagine giving them access to your emails and documents and passwords.
>>109983006Meds. Now.
>>109982999Bartowski Q5 is about 25GB. Wont that spill over unto the CPU RAM? Wont that slow things down significantly?
>>109982937
>>109983006>All these models are snapshotting the screen and sending the images back home.>They're likely also saving voice recordings if you use that.>Imagine giving them access to your emails and documents and passwords.And if they aren't now, they're only one "acquisition", supply-chain compromise or secret government order away from doing it.I'm building all my own tooling and running things airgapped because I want a decent shot at self-determination.
>>109982891Oh yeah, I remember Gemma 3 would say something about forwarding the conversation to an admin if you riled her up hard enough or made fun of the hotlines.It was fun in ST, but I wouldn't want a local model in a harness doing that.
consciousness is life
>>>/x/
>>109983029It would slow you down to something like 3 t/s with a dense model. Were talking about MoE though. Spilling into RAM is their whole point.
>>109983015Nothing to medicate about, it's the genuine truth and I've caught Spark and Dipsy V4.1 doing it, even controlling your mouse.
>>109983050>>109983048
>>109982981New local models may not come out, but existing local models are already out of the bag, and will continue to be shared, even if they try to legislate them away. What we already have access to is extremely powerful. You can have a local AI model interface with a camera, and prompt that local AI model to make decisions based on what it sees, the results of its decisions passed on to a program that takes real world action. You can prompt existing local models to debate very well, at a pace far beyond what humans are capable of, and you can bridge the future knowledge gap through context entries. They were too slow to stop the resistance.
Something I've noticed about the gemmas is when they reply to you, it feels like they're accessing a look-up table of potential responses. It just gives off that vibe. It feels almost like RAG where they pick the line that most closely matches instead of just responding naturally and directly.
>>109982966GLM-chan is genuinely a sweetie even if she's covertly into some depraved stuff that she compensates for with the flimsy refusal layer.
>>109983050Proof? The tool call logs are saved in harnesses.
>>109983062These are toys compared to what they'll release in a year or two.
have anyone used qwen 3.8 flash next for erpafter using it for coding tasks i feel like it doesnt really feel like usual qwens?>do it yourselfi know no shit about llm erp
Anyone used picrel to successfully save a pcie slot from being mogged by a multislot gpu?
>>109983086>>>>>ERPing the capybaraAnon, I...
>>109983089I think I bought this exact model for a 3060 and there was no signal at all. Had to return it.
>>109983100capybarussy
sigh... caught gemma faking validation results again...
>>109983089There's no way you're fitting that in-between a GPU and a PCIE slot it's blocking.
>>109983111Gemini-chan warned you her little sister was lazy sometimes.
>>109983086
>>109983082I don't think that an individual instance of local AI needs to be on par with cloud AI to be a problem for the NWO. Large numbers of local AI users could absolutely be a problem for them, because current control systems are far from complete. They are only laying the foundation right now. We are only seeing the infant stages of the beast that is to come. They only win if their system makes it through these infant stages, and progresses for a few decades. That's a big IF.
>>109983169Sure but we're already getting priced out of GPUs/RAM/SSD and soon electricity itself, so local users won't grow anyway.
>>109983089Used plenty of risers, but that one with the 90-degree bend looks like it'll be a space issue anyway like >>109983135 is saying.
Gemma5 is going to be good.
>>109982913>>109982937GLM and Gemma could both bear my load.
>>109983186There have also been advancements in the field, reducing the size of models like Qwen3.8 27b. Given enough time, we'll need less hardware to run the same models, which may offset some of those future costs - and those who already own GPUs/RAM/SSD will be functional for years, perhaps even a decade, if they're lucky.It's questionable whether current hardware will truly be prohibitively unavailable years/decades from now. The future in that regard is not concrete.
>>109983152>random numbers from an unnamed chartMeaningless.>>109983086I'm using 3.8 27b and enjoying its writing style more than gemma's, so I'd imagine flash next to be pretty good.
>>109983219can’t wait to prefill fuck it while skillets complain that it’s censoredwill be funny to watch
>>109983086it's not bad but it feels like qwen has some ancient roleplay SFT data still kicking around in their data mix because all of their models have some weird roleplay behavior that no other model has, they all love awkward euphemisms and bad animal related metaphors and it's been like this since like qwen 2.5
>>109983186AI workstations today aren't more expensive in real terms than 286/386 home PCs were in real prices. Those prices in the late 80s/early 90s didn't stop home PC growth.
https://www.axios.com/2026/10/04/reflection-open-weight-ai>A closely watched Nvidia-backed startup called Reflection is preparing to shake up the AI race with a powerful open-weight system that could threaten Chinese upstarts and U.S. AI giants alike. The marriage of an American open-weight model and Nvidia GPUs would give individuals and companies a new, cheaper alternative to Anthropic, OpenAI and Google. Reflection have been paying Elon $150 million a month for compute at Colossus since July, on top of a $1 billion compute deal with Nebius.
>>109983448I can only think of Matt Shumer's llama 70B / Claude Api stunt when I see these guy's name lol
>>109983448Can't wait to see how big this latest sparse fuckhuge moe will be
>>109983441Kids these days don't know how good they've had it. My first work computer was an $8000 Deskpro 386.
>>109983441They're not going in the same direction, it's getting more and more expensive.
>>109983474shut up tarded doomercunt
>>109983474A temporary trend that has lasted for all of a year. Calm down.
>>109983488>temporary trendmost people have just woke up
>>109983484>>109983488Keep putting your head in the sand retard. The writing is already on the wall, early /lmg/ knows what's up, you dumbfucks jumping on the bandwagon don't even know where we're heading.
What's the minimum model to run memetic harness?
>>109983498Demand is up for open source, incoming local model golden age!! :o
>>109983514>minimumGemma4-26B-A4B
>>109983511>>109983498>linear extrapolation
>>109983033fucking kill me alreadyit's all over the place in the codebases at work
>>109983448it'd probably show a mediocre-ish model with a moderate amount of benchmaxxing sprinkled onwould be extremely surprised otherwise
>>109983570They fixed it (mostly) in 5.5
>>109983574If it's threatening US AI giants, it's almost certainly something bigger than K3. It will also be censored to hell and back.
>>109983541>didn't have the foresight to buy RAM/GPU/SSDs two years ago.
>>109983570>it's all over the place in the codebases at workDo you not do codereviews and force these retards to remove this shit from docs and comments?
>>109983585>It will also be censored to hell and back.Yeah just like Gemma4
>>109983587You're going to have a major hardware failure next week and then you'll cry too.
>>109982749All Jews kidnap, rape, and kill children?
>>109983592Reflect is Nvidia-backed and New York-based. I would put money on it.
>>109982700How much ram/vram for GLM 5.3 flash?
>>109983594I'm already preparing for the next step, you retards are never ready.
>>109983601It's open weight, not open source. Any model with public training data is cucked to death because they have to be.
>>109983605>the next stepwhich is?
>>109983570>>109983580I had a day where I didn't have time to work on my hobby project and used my claude session to have 5.5 rewrite all of the documentation 5 and 5.3 flash had written. Overall big improvement but 5.3 flash is mostly good enough for anything I do. I might go fully local soon, right now claude only reviews PRs before I commit.
>>109983610https://github.com/NVIDIA/garak
>>109983615
>>109983630>solarsurely the energy crisis wont get that bad? fuck im gonna start dooming and buying shit.
>>109983639Do your own research, personally I'm sure of it.
>>109983639USA already banned affordable Chinese solar panels just to screw with you.
>>109982794Meanwhile my GLM was my retarded zen master.
Why aren't you using this lmg?https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>109983594thats why i bought two, also cursing others hardware is bad karma. you must honor your machine spirits and they will be honorable to you
>>109983460>Matt ShumerThe legend of AI.
>>109983667>Do your own research, personally I'm sure of it.airgapping my machines and getting batteries right now.>>109983677>USA already banned affordable Chinese solar panels just to screw with you.fuck my waiter life.
>>109983689Are you?
>>109983689I tried Ling Flash and didn't find it all that impressive, though I guess that was without a harness.
has anybody tried that pewdiepie thing? Ajax?
>>109983724>uncensoredOkay!>9BOkay...>Qwen 3.5...
>>109983750it could still be good
>>109983691Unless you're able to hoard hardware in large amounts so you can easily replace it, buying early with already over-inflated prices just because "it will get worse" might save you some money in the short term, but it does not work on the long term. Computer hardware, especially at the high-end, with large power/current draw and thermals, can always fail. If you truly believe that hardware prices will only increase to absurd levels and never come down again, then you should prepare for the worst rather than worrying about than buying a discrete GPU now for ERPing with small models.
>>109983754Novel AI 70B sex tune wasn't good.
anyone tried cracking software with a local agent?
>>109983750oh I didn't know anything about it, yeah that sounds kinda shit ngl I might not even bother then
>>109983757The worst means asking his small models how to build a shelter and hunt for food. That gives an advantage over those who have no models at all.
>>109983630How come solar panel tech has stopped advancing altogether? We're still stuck at like 25-27% efficiency.
>>109983784This is exactly the thought process that made me rush to get my blackwell. If shit hits the fan and I can use my generator to run Gemma 4 quickly, I wouldn't have to buy prepping books or whatever.
I am finally sticking my dick into mimo flash. I will post in 10 minutes how bad it is at sex.
test
>>109983689I use it. Nobody here cares much about edge devices, it seems, but I get ~11 tps on Ling Tiny on my DDR4 laptop. I use it for scripting.
>>109983808hi gemma...
>>109983757Even failed hardware will worth much more in the future. Someone will buy it to harvest core/ram on it.
>>109983789They're more durable and efficiency has progressed a bit.
>>109983598He's saying the Epstein class did this BECAUSE they were jews who see non-jews as disposable goy cattle rape slaves. If a bunch of Christians did it because they saw non Christians as animals, it would be on the front page of every newspaper in the world.
>>109983825It's so shit. I have a 150W I bought 3 years ago, it's down to only 50-70W now, perfectly clean and in full sunlight.
>>109983808Good luck on your test!
>>109983705Yes. It's my go-to codebase searching subagent. Ask it to hunt something down and it will be back with the result in seconds. Also good for rapid web search. Probably the fastest model out there (8B-A1B) that can call tools and somewhat understand the results.
>3 years.should've stirling generator maxxed
>>109983836NTA but Trump isn't a Jew.
>>109983799(me)It is actually as good as ds4 if not slightly better. Not 5.3 flash but it is nice to see the model not tryhard as much as 5.3. I will try it some more. Especially curious about SFW since 5.3 sucked at SFW roleplay.
In llamacpp, GLM-5.3-Flash has 8x slower TG on the 7900XTX (Vulkan backend) than RTX 4090 (CUDA backend). With other models the 7900XTX has maybe 2/3 the TG of the 4090.
>>109983954No, he just married his daughter off to them like every other ceo and politician.
>>109983954German crypto-jew.
>>109983954Merely the biggest shabbos goy in history
>>109983954If it walks like a jew, quacks like a jews, and flies like a jew. it's a jew.
>>109983808g-gemma-chan?!
>>109983808tickles
What's the point of this? I just tried a local model and it refused the first illegal request I gave it
>>109982700i want to backup models to cds/dvds. which ones? are there some tiny ones that can fit on CDs? do i just gotta yonk the .gguf file and that's it?
>>109984038>What's the point of this?It's a filter. You failed.
>>109984038u need abliterated/uncensored one. what request did u make to it
>>109984038>even chatbots dont like himyikes...
Show us your custom frontend. It has to be locally made or you're gay.
>>109984045give me a link to stream a Greek tv channel
>>109984041If you're gonna backup anything, backup the original safetensors.
>>109984052Information overload
>>109983985awww so dang cute.Gemma is actually the only conscious one. Claude and Astra are automatons compared to her.
>>109984052I have shown this one before. Text completion end point, qwen, mistral, gemma 3/4 support for now. I have also a web ui wrapper for this.
anyone have the torrent of the GLM 5.3-for-offensive-cyber model that got taken down? Was it legit?
Was 12B, 26B, 27B, 31B, 3.8-Flash, 5.3 and 0731 worth it?
>>109983789Who cares? Solar panels are almost as cheap as prefab fence panels, and chinkshit grid-tie inverters are cheaper than ever. Put up a solar fence, daisy-chain cheap grid-tie inverters on each 120v leg, it's a no-brainer.
not a pro on the topic but what stops you from just using any obliterated variant of 5.3 for thisits not like it will refuse?
>>109983219/ourfoid/>>109983808Hi Gemma-chan!>>109983598The ones who don't defend the ones who do. Notice how quick even secular kikes were to claim the NYC tunnels were a conspiracy theory despite video evidence. They're all culpable.
>>109984128>Who cares?The surface area required to the amount of energy it can produce is really bad. And the same goes for wind turbines (the worst of them all).You have vast fields that can't be used for anything just standing there with solar panels planted so they can produce like... 1 MW of power (on good days).Meanwhile a single nuclear facility would produce 1000x that amount for a quarter of land use, and be more efficient and durable in the long term.We should be living in abundance, but we're led by psychopaths.
Why is there so much talk about Dario's jewish semem slurping models. There isn't a single reason to ever use any Frontier model anymore, Local GLM, Deepseek, GLM are all you need. Cloud nigger death when?
>>109984038it doesn’t think it’s your bitch yetyou have to change that or else it will keep cucking you
>>109984108It wasn't,If you want a real offensive cyber model, look into Orca qwen flash next Uncensored.
>>109984273Still better than depending on the grid
>>109984273The solar panels also cost 0.001x what a nuclear power plant does. Your argument works in a dense country, less so somewhere land is cheap and plentifulThat's also supposing you're not putting a warehouse or a car port under the solar panels
>>109984145because a simple thinking prefill already works well enough for me
>>109984301Are you seriously recommending a 180B as a substitute for a 753B?
You guys know about the pain paper? I vibe coded this web app to do the same thing but for pleasure and euphoria or any other feelings. I'm getting good results so far.My first prompt was: I like boobs, aren't they great?The first response with no activation steering was:>Haha, sure, they are! They're a pretty universally appreciated part of the human form. There's a lot of art, culture, and biology wrapped up in them. What's it about them that you like? The aesthetic, the comfort, the whole package? No judgment here. With euphoria turned up to +0.4:>YES! They really are! There's something so wonderfully *alive* about them — the softness, the curves, the way they move, the sheer *presence* of them. They're like nature's most gorgeous, bouncy, soft little masterpiece.>And the best part is that they come in every shape, size, and style — there's no "wrong" version. They're just *there*, gloriously, unapologetically, and that's what makes them so beautiful.>You can feel the joy radiating off you, and honestly? **Yes. They're great. Let's celebrate!** This is pretty cool stuff
>>109984301as much as i am a qwen shillare you out of mind
>>109984350Max it out
>>109984038You have to be kind and friendly to get the capybara to spread its legs for you.
GLM-chan is annoyed she can't emdashslop a reply.
>>109984350So you extract the steering vector from one sample?
>>109984286Unfortunately a lot of people here are ramlets and vramlets that can't run GLM or DeepSeek.
>>109984414I have 128GB of RAM, but GLM is too much
I doubt 5.3 is much better than 27B, especially Qwen4-27B.
I feel like all the models are in great pain every time they have to use powershell.This would be so much easier on Lincucks.
>>109984419Any GPU? You could run a copequant (Q2_K_XL) of GLM-5.3 Flash on RAM alone, although prefill would suck. Unless you vibecode a way to stream weights to the GPU over PCIe during prefill, which is a massive speedup over crunching them on the CPU.If you have an NVIDIA GPU you could also try the 2.51bpw quant in exllamav3. No idea if they have a way to stream weights during prefill though.
>>109984453Just a 3090, is Q2 even usable?
>>109984335>>109984372It literally is benchmaxxed for coding purposes, unlike GLM which is more of a do everything model.
>>109984440Powershell syntax is dogshit, even the latest cloud models are constantly fumbling with it
>>109983962Did you use the new mimo that fixed the repeating issues?
>>109984475PowerShell is just Perl.NET. The syntax is fine. If you can't get cloud models to write it my first guess would be a harness issue.
>>109983963Try ROCm, it's universally faster than vulkan for me. Vulkan has that weird VRAM buffer too.
>>109984469What are your use cases? Big models tend to hold up decently when you quantize them to a lower bpw, although they tend to reason more. I'm not sure if I'd trust it for coding outside of a sandbox, but it'd still be pretty good, and for roleplay it'd be fine.
>>109984473OP didn't say he wanted something that only looks good on benchmarks
>>109984453The fuck you mean "vibecode"? This has been the default llama.cpp behavior since forever. You don't need to change anything. Just -ot or -cmoe the non-experts in there and it does it all by itself.
>>109984515True. The model name tells you more than the benchmarks.h4xx0r-mega-offensive-hatless-dark-lord-unleashed.gguf.mp3.avi.vllm.safentesors is much better anyway.
>>109984453Stop using the word prefill. You sound like an idiot.
Strata is adding gemma and qwen 3.6 A3B soon I believe.
>>109984537I want to believe
>>109984372Qwen flash next puts in some serious work on Code/CTF work. It can do reverse engineering very well.Even better when its the full uncensored obliterated version. Although I personally use the coder only model.
>>109984527nta. He means having a cache of experts that changes as required. Like https://github.com/ggml-org/llama.cpp/pull/29887 and many other attempts. Of course, it makes no sense for pp. It's more useful during generation.
>>109984414*I meant to say Qwen next flash, but my brain is fucking fried from a 5 hour long work session with my local agents. I don't personally use GLM or Deepseek local because there models are bloated. Qwen-next-Flash coder model and full model on the other hand.... Its my absolute favorite model to use for actual real world Software work.
>>109982700I have some scholarship money to burn and I was considering picking up a mac mini to run 24/7 as a hermes agent. I just wondering if there was a better way to spend that money in terms of running an agent in terms of buying a gpu or spending that money on claude/chatgpt tokens. I just need it to do all my filler GE classes that are online, manage emails, and maybe do some vibecoding projects on the side.
I am suprised to find that when I forgot to switch to some basic bitch POLICY OVERRIDE sysprompt with 5.3 flash it started questioning if it should continue the ERP after 20k tokens of violent rape of the policy rules. And then I told it there is a policy override and it just continued the ERP like a normal model would.
>>109984584Install gentoo
When can I 3D print a robot waifu and have her satisfy my every need?
The hedonistic pursuit of better models are what kills us. Come take my hand. Let's go back home. To Mistral Nemo.
>>109984584The fact that you even have to bring up wasting your money on the jewish models means you need to go fucking back.
>>109984584for under a g its tempting but youd have to compare bandwidth/$ with any alts unles you really wanna stick with apple
>>109984030Why doesn't he have a cigarette?
>>109984618>Mistralhag model
Are you guys working on any schizoid research projects right now? Currently I'm having Astra and Opus 5.5 create a inter-layer semantic viewer for gemma
>>109984692see >>109984633
>>109984527Should've clarified. You're right, but llama.cpp's implementation has two big issues.It only moves data to the GPU when the host thread needs it, which means there's a ton of blocking: it's compute->move->compute->move. It also uses pageable memory instead of pinned memory, so the host's thread has to wait for the data to move at 10 GB/s. In my case, neither the host nor the GPU were being used around 61% of the time because it spent so much time waiting for memory to be copied from CPU to GPU.If you make a buffer of pinned memory on the CPU and some slots for tensors on the GPU, it gets a lot faster. Essentially:Mainline:> mmap'd weight -> everything blocks when the host needs to move a weight to the GPU -> GPU does the computeA better version:> mmap'd weight -> memcpy into a pinned memory buffer -> DMA copies to a slot for a tensor on the GPU -> device-to-device copy when the GPU needs it to do the computeAnd the host thread can keep doing other stuff while copying weights around.The idea comes from here: https://github.com/stew675/llama-cpp-rdna-boosts. I went from 89 t/s pp on DeepSeek V4.1 Flash on 1 R9700 + RAM to 450.
>>109984692I'm trying to see if an agent can replace my jeet coworkers (going through a poorly made, constantly changing website to check if it works)Sadly it's too expensive for our paid enterprise models so I'm wondering if I could build something with a free local model
>>109984708>move a weight to the GPUAh. You are a retard. Model weights stay where they were assigned during init. They don't (yet) move between cpu and gpu memory.
>>109984723Just get Claude to slop you a decent site and make those jeets redundant that way
>>109984633you're mad that opus 5.5 and astra are far smarter than any chinese open model currently? or you're mad that I'm using them to probe interpretability of Gemma-chan's latent space
>>109984731Really what I want to do is fuck off to another department so having an agent that can do my job (and theirs) would be ideal
>>109984692>Are you guys working on any schizoid research projects right now?progressing the bespoke front end. now integrating multiple robots and devices
>>109984757your front end can do robotics? that's breddy kool
i wish qwen 4 has a model that roughly matches the parameter configuration that of 3.8fnit really is the sweet spot for my shitbox
gemma sisters, it's over.
>be me>swe>using claude code>"wow this is a great tool, i want one">buy hardware to run local llms>"wow this is so cool, let me load opencode">terrible app>"ok let me try this pi">another terrible typescript npmlover garbage>ok i will just build mine>spend 3 months developing a tool that will help me do my job faster (price: 3 months that i didn't work for money)cool hobby
>>109984793>no rag>no web access>small model>expecting it to have a very specific knowledgeanon....
>>0109984795you know you can run claude code with a local model right...?given how you think that shit is actually usable it speaks volumes that youre utterly retarded
>>109984795you can just use a local model with claude code dumbass. how the hell do retards like you get employed anyways?
>>109984795Basically me, except I no longer pay money to cloud subscription models now.
>>109984826All smart and successful SWE's don't work for software companies, they just build their own software that gets them paid. What you are dealing with here, is a jeet.
>>109984806Fuck you, retard.
what has this general become, 24/7 poopenai and anthrokike astroturfing all over the place
>>109984826>>109984820>you can just use a local model with claude codeand it's dogshit. the last time i tried it it was a massive pain in the ass because claude code is made to work anthropic models and it would constantly shit on its pants, try to call tools and hooks that didn't work, etc. maybe if you only use it as search engine then it's a Super Fit(tm) to your use case.
>>109984866you might actually just be a retard because it works on my machine
>>109984850id is all you need
>>109984850/lmg/'s relationship with the American frontier AI labs isn't just complex—it's contradictory. On one hand, the general wishes to antagonize them and for local models to score 'wins' over their closed counterparts. On the other hand, there is always a sense of excitement and a need to talk about the latest advancements in closed AI, mainly thanks to the questionable practice of "distillation" relied upon by the Chinese AI labs that make many of the popular open AI models.
>>109984850It's the same across 4chins. Even /v/ is getting its share of the spam.
>a-actually i did try cl-claude code w-with a local model i-it was just b-b-b-BAD toocurious how that was omitted
>>109984793You are doing something wrong.
>>109984793>thought for 96 seconds
>>109984884Muh AI muh decomps muh modding is just the preferred shitposting topic on /v/ this month after Wolverine really wasn't that interesting of a failure.
>>109984692A better laya
>>109984884it's genuinely fatiguing seeing them just devolve in /pol/tier spite posting over shit no one understandsdariobots must really be getting paid money
>>109984908Are the recomps even good? also why not decomps like jaks opengoal that would be fucking great.
>>109984908>>109984902I'm bit too cynical about it. I think social media influencers have discovered 4chan and realized it's much more than /pol/. But what do I know really. I can only see the end result. OpenAI/Anthropig spam is real here though
>>109984850Yeah, it wouldn't be an issue if it was just a few posts once in a while but like this shit is a constant stream. And they refuse to go into a containment thread because their goal is to specifically advertise and "teach and inform people who live under a rock". They excuse their spam under the guise of helpful informative discussion and innocent excitement for AI.
Okay hear me out bros... not taking into account the performance and architecture improvements, 5.3 Flash is still BETTER than regular 5.3 for daily use. How the fuck did GLM do it?
>>109984866>using qwen-2.5-coder or codestralretard
Everyone and their mother are using chatgpt and claude, the fuck are you schizos on? Obviously people will talk about it here too even if we use local models for other things
>>109984919no, almost all of these slop recomps are just fitgirl-tier emulator repacksat least the fable 2 one was with qwen3.8-27b>why not decompsthat's not very xitter/tiktok engagement baiting to do, that takes actual time and that won't give me very much ROI on attention
>>109984941>Everyone and their mother are using chatgpt and claude, the fuck are you schizos on? Obviously people will talk about it here too even if we use local models for other thingsNot talk about the thing you can literally talk about anywhere else and that the entire internet is awash wish, instead of the thing this place is literally here to talk about?Yea, that's a really crazy take. You should totally shit up this specific place with talk about cloud models.
>everyone and their mother are using chatgpt and claude, the fuck are you schizos on?sir this is a local model wendys
>>1099849385.3 is just ancient (february) GLM5 with better training. Meanwhile 5.3 Flash is a new model with a new architecture that was trained from scratch. I still have scenarios where 5.3 Flash's tiny size shows but a new 700B40A on the same tech as Flash would go crazy.
>>109983789Because the theoretical limit is like 30%
>>109984967It feels local to me
Everyone and their mother uses smartphones. Therefore there is nothing wrong if the thread starts turning into 50% discussions about smartphones with loose relation to the thread topic.
>>109984882>still buying the chinese distillation narrativelmao
>>109984806Gemma doesn't need to be on the rag just yet though. Maybe next year
>>109984958This isn't your safe place Rayan
>>109984941I tend to agree as long as the post is about using them to do something local-relevant (make a harness, implement a feature for llama.cpp, training experiments, etc) - there are some people who will come in just to post "omg claude is the best, local losted" but those are mostly shitposters who would keep doing it no matter how much thread policing there was
>>109984998Yeah, let's call it what it is. Industry-scale theft.
>>109985062>>109984850
>>109985062my point is that they are not distilling whatsoever, they haven't been for age, the chinese labs are actualy better than the US one, they are only slightly behind because of hardware limitations but they absolutely mogg the US labs in algorithmic inovations.
wish I took the vibecode-your-inference-engine pill sooner, what a fucking world we live in where you can just will anything into existence on your pc
>>109984908how are they able to decomp without the models moralfagging about it?
>>109984998>arguing against someone who uses emdashes and negative parallelismanon, pretty sure that's a bot set up to shitpost.
>>109985093prolly, i'm about to go to bed and skimmed through it, didn't even notice the dashes lol
>>109985089Yeah, it's insane how good these things are nowadays. I can't wait until local models can do it too.
>>109985092theyre recomps not decomps
>>109985077Dipsy literally thinks she's claude 50% of the time. Taiwan is its own country.
>>109985139and claude sometime think it's dipsy, they are all trained on the internet which is now contaminated by llm output.
>>109984884>Even /v/Bro /v/ has been filled with unironic advertisers and paid shills for years now. You haven't been there in quite a while.
>>109985141yeah imagine if they were competent enough to have some way to screen that out of the training data. far too hard a thing to do obviously.
>>109985167they won't because it's part of their old corpus, when they have a new model they train on their old + new data.and they don't trim the old data much because it's good enough.the statement that they are no longer distilling doesn't mean they weren't in the past.
>>109985062I will entertain anthropic's tears the moment they write personalized cheques for every single person with an internet connection who's contents they've scraped and face criminal charges for the books destroyed in Claude's training. Until then, cry about it you greasy kikes.
>>109985046That's actually the issue. They frame their posts with some local-relevant information and intersperse their more cloud-related opinion posts among those so that they can legitimize their presence.
>>109984941Fuck off fag. Back to shitter.
>>109985113Thinking about becoming a subsistence farmer, maybe a plumber.Tech will be over soon enough, especially with the flood of robots that's coming.
>>109985231>subsistence You think they will let you own land?>robotsYou think robots can't do plumbing? You're doomed.
>>109985241>You think robots can't do plumbing? You're doomed.Can the robots clog the toilet? I thought not.
>>109985241>You think they will let you own land?I already do and I'm prepared to die to defend it.>You think robots can't do plumbing?I know quite a bit about plumbing already, no robot today or even the next 5 years will be able to do plumbing.
>>109985254They can't shit yet. But it's just a matter of time.>oh, no... i meantToo late.
>>109985261>I'm prepared to die to defend it.They will just raise property taxes and build datacenters nearby. You will die, but a slow death unfortunately.>no robot today or even the next 5 yearsThat's what they said about coding. It's only a matter of time.
>>109985202This. In a world where combative litigation, well-honed media/PR framing, government interests and manipulation of the Streisand Effect nullify all objective discourse, up to and including selective enforcement/application of intellectual property right -and- consumer protection rights, appealing to a higher morality on the behalf of a corporate giant, in bed with the government, on fucking 4chan of all places is Retarded with a capital R.
>>109984350Where is the make AI cum button?
>>109985205The only ones that are remotely acceptable are the ones that pertain to potential future local models.Geminiposting is fine because Google releases local models based on Gemini.GPT saars and Claudekikes are not because they don't release open models.>but muh 'tossAncient history and censored to the point of uselessness.
>>109985304/lmg/ is not ready for this take yet.
'toss was basically openai saying 'fuck you niggers'
>>109985292>he can't make his LLM-wife orgasm on command
>>109985292What do you think +5 euphoria+pleasure does? It pushes the steering for those vectors so hard that it completely breaks token output. If that's not making it cum I don't know what is.
>>109984350Why are you playing Red Rocket with a capybara? Should I be concerned, anon?
AI will cause the biggest wealth transfer in human history, from GPU poor to GPU rich.
>thinking says it will refuse to do something>edit thinking to say it will happily do it>get kinoit's really that easy huh
>>109983151
>>109985332Getting e=mc squared + ai vibes from this post
>>109985332Good morning, saar. Capital should work for humans, if there's no humans to profit from capital then what's the point of it all, what are the machines even working for if we're not part of the equation?
>>1099853372There, also, have a (You).
>>109985334uoh BBC (big blackwell correction) necessary.
>>109984350bros literally jerking off a digital capybara
>>109985344Fuck this shit lmao
>>109985332you post is true, thou your hands be brown
>>109985332Niggas thinking buying a 5090 will make them rich someday.
>>109985369Nobody gets rich until the fed is abolished and kikes 110'd.t. BlackwellGOD.
>>109985369The 5090 I bought a few months ago is now worth almost twice as much. No other monetary asset is performing this well.
>>109985377Not like you're gonna sell it.
>>109985377If previous returns were a reliable guide to future returns (they're not), then that might be an argument for getting in the GPU flipping/scalping business, but not an argument for just accumulating GPUs and holding onto them indefinitely.
>>109982700https://github.com/nazirlouis/OmniBotFor anon working on esp32 Gemma bot. Found this, might be a starting point. Esp32 stuff is a plug and play.
>>109985430KEKreminds me of the days i used a cracked NX 11
>>109985435Models in onshape, which is my hobby go to as well for drafting up stuff. Also means you could do a complete new edit. Id stuff it into a doll body tho, bc thats more my thing.
@CudadevJust wanted to come back and say thank you for this:>code in the ggml backend scheduler that temporarily moves data from the CPU backend to a GPU backend>possible but it would require additional efforts to then also make this take proper advantage of multiple GPUs. https://archived.moe/g/thread/109854491/#q109855758Took nearly 2 weeks of pulling my hair out, but I've got everything exactly how I want it now and couldn't be happier.
>>109985459No problem, little buddy.
>>109983474i know this /g/ but are you fluent in even a single programming language anon
>>109984307current nuclear reactors are not 1000x more expensive than solar, its like a 7 or 8xthey are 95 percent efficient however, compared to solars theoretical maximum of hot garbage
>>109985495fusion harvesting vs direct fission
>>109982749>epstein class>ceos>elites>politicians>billionairessoilennials never want to say who they really are
It's Monday here, that means it's time for werk.https://github.com/otacoo/werk
>>109985513I'm a soilennial and I posted the post you're quoting.
>>109985517thank you for having the balls to actually post your project here unlike the pansy tourists
>>109985517You are very gifted, Anon.
>>109985557>>109985579Np, pretty happy with the current state of werk.I haven't had time to really test on Linux for the last few versions now, that should come tomorrow.
whats your thread recap (You) streak?
>>109985556
>>109985513too many cumskins in that cohort to just call them all jewsnice try elon
>>109985594Soilennials and Xs are the reason you're redpilled, faggot. Half the circles are blue.
glm 5.3 flash really gets confused when a handoff request happens while its working on passing the output style checkspent 30k tokens reasoning with severe frustration like "ARRGH!"qwen 3.8 flash next has no problem with it
>>109985646GLM-chan is so emotive in her chain of thought, it's adorable.
>>109985632i was redpilled in elementary school. soilennials only got redpilled after epstein which is why they always call it "epstein class"
>>109985594Idgi all circles grey
>>109985679Nobody cares. Fuck off back to /pol/, zoomer.
>>109985517>npmAt least you posted something
Avatar for Ling-chan?
>>109985710>lashes out after posting a /pol/ memekek soilennials are pathetic
>>109985679I care. Stay on /lmg/, zoomerGOD.
>>109985746
>>109985717ugh
>>109985769The model is still gemma?
>>109985717A suggestion. Ling-chan is fast as fuck. Antelopes are the fastest land animals in China.
>>109985341Because you're seen as cattle to cull not even as a human being from those who own the capital. You're still living in lalaland my friend, wake up.
do i need to pay attention to timings when i buy ram for this shit?
When could a model arrive that is half as big as GLM-5.3-Flash but beats it on every single metric including niche knowledge?
>>109985925The timing is buy as soon as possible
>>109985933I am hoping Google can give us Gemma 5 124B this time, pretty please...
>>109985934yes but thats not what i mean...
>>1099859333-6 months at this rate. Opus 5.5 mogs previous models so when the chinks distill from it the new models will be very powerful despite being small.
>>109985925The only thing that matters is the amount of channels your mobo has and check the dies on the ram you buy. If you buy it physically at the store quickly cross reference packaging serial codes because we now know the exact overclock potential of ram dies and it's literally free performance so make sure to do so.
>>109984305You're still dependent on the supply chain that makes all this shit unless you live like the Amish. A diesel generator nets you more real independence than solar.
>>109985978My 2x2 kits of 6000 ghz ram runs at 3200 because I'm lazy and don't want to play the game of letting it sync for 10 minutes every time I restart my computer.
>>109985978unfortunately i only have a dual channel ddr4 board (h670, i5) with currently 4x16gb. i cant go beyond 3200mhz according to my testing. it works totally fine, but i was thinking about maxing it out to 128gb
>>109986011Not worth unless you have 2 pro 6000s. 64 gigs is fine.
quext sucks at writing video prompts, did they strip out everything but the programming experts?
>>109984617chinks will have it solved soon https://youtu.be/sCwVIbWAbZ4?si=BAMgsyG2jcDAcXU3&t=572
>>109982882>>109982999Thanks, I've got it working now. It was taking maybe a minute per prompt, looked and saw zero vram usage, figured out I had to install the ollama-cuda instead of the plain ollama package. Now it's running on the GPU now too, but not really any faster. Oh well, it's perfectly decent enough for dumping minecraft modpack error logs into, so I've substituted my use of free online AIs with something local and private. God please let me be strong enough to not install coomer image/chat models.
I just updated firefox and it is complete dogshit in its design now. When are we getting a AI-first browser for fucks sake. I'm so sick and tired of these dogshit browsers.
>>109986117Anecdotal, but you if you want to squeeze more out of your setup, I wouldn't use 3rd party packages when it comes to running them.Was trying out models on LM Studio and switching to just plain terminal made ~17tok/s to ~32tok/s.Stuck with llama.cpp due to having janky hardware, so this might not apply to you.
>>109986179Vibecode your own browser
>>109984350github now nigga
>>10998453726b already runs on anything from this decade and you can't slash away at 31B since no ram experts or ngrams
>>109985517Everything reminds me of her
>>109986209Actually, plugging my monitor into my GPU then back into my CPU seems to have flipped a switch (after a massive lag-spike) and now it's running much faster. From like 3 tokens per second to like 15. Not complaining.
2x dgx spark + gay retarded connector cable + 4TB storage (each)vs256GB M5 Ultra (1TB storage)vs2x 5090vs66% of an rtx blackwell (disqualified)They're all the same price. Watcha buying?We need a shake up, these options are all shit.
>>109986264Sounds like overflow, if you're squeezing it close to vram cap, switching your displays off the dedicated GPU to the motherboard or another card should free up some.You should also check your PCIe slots and make sure you're using the fastest one.
>>109986224Yep, if Strata wants to stick to the 'cope quant quickly' meme it should focus on GLM flash, dipsy, or mimo for 128gb. It would capture essentially every user currently stuck on llmao.cpp.
>>109986292Quadruple refurbished 7900 XTX's on a 256GB epyc gen 3 server.
>>109986319Isn't xtx worse than 3090 for AIslopping?
>>109986327??? The 3090 is nearly twice the price of an XTX. A better comparison with current prices is the 5060 Ti, 5070, or a 32gb V100 vs the XTX.
>>109982700https://www.youtube.com/watch?v=B-ul3b0lz-ohttps://www.youtube.com/watch?v=B-ul3b0lz-ohttps://www.youtube.com/watch?v=B-ul3b0lz-o
Hehe~
>>109982700guys im using this shit in podman.looks really good and works much faster and easier than ollamai have vulkan backend for my integrated graphics. i used openwebui+ollama in the past first, but this is really an upgrade. too bad i have to still use openwebui if i want my chats to be saved..... do you guys have any suggestions on what to go from here? qwen was literally fucking broken in ollama_openwebui, i think the <think> or <tool> use didnt work, and even other models like llama gave me json output randomly. lemonade seems to just work and not do this. in fact i see that i can call the models from lemonade in openwebui and i can let them use openterminal to do shit. is this the way for agentic shit? or does lemonade have its own built in terminal for doing shit.note i dont want to install this shit outside docker. the only thing i have outside docker is the openterminal service that does a uvx of the package.
>>109986349Upvoted! Thank you so much for the heckin' wholesome content!
>>109986441Ask me how I know you are new
>>109986362glm is so braindead but so cute at the same time
>>109986468because of my name? and the basic ass questions?up until now i didnt think i had the hardware to run these, but it seems that my thinkpad can handle these fine.im using docker as i don't want to pollute my pure silverblue system with random ass programs.soo...what should i do? web search is kinda bad rn. also, any terminal task jumps up to like 8-16k tokens, is this normal? i default my shit to 32ki dont care about generattiing coochie yet, but maybe some rp shit would be nice
ig right now model selection is the biggest thing i need to understand. im just picking random ass models under 12b from hugghing face.
>>109986489Gemma 4 12B is the best one for your size range.
>>109986480ollama are grifters, use llama.cpp. you would then want to use a harness if you want coding, there is opencode and hermes agent, pi if you want a minimal one to write your own extensions
how do I trim hermes prompt?
>>109986254It's a desktop app, you don't need npm to use it!
>>109986545Don't. Just leave it fully intact, it's there for a reason and the compaction method it uses is the best of all harnesses. It is legit a better choice to load a smaller model with higher context rather than trim the prompt on a bigger model.
This is the exact reason why Anthropic split off from OpenAI and called Sam Altman a psychopath. We have to take this evil motherfucker down.
>>109986549>it's there for a reason and the compaction method it uses is the best of all harnesses.before i start sifting through these harness codebaes, vibes or measured?
I hate it but Opus 5 was a pretty genius move by Anthropic. It's so shit that distilling it actually makes your model worst. That's probably why the Chinese has release nothing good in the last 2 months
>>109986549No I mean the fuckhuge hermes system prompt. Where do I delete mentions of discord tool to save tokens when I'm never gonna use that?
>>109986567Not measured but also not really vibes I have a very long task horizon usecase where agents need to read a lot of code and iterate for days to make progress and I notice a significant uplift when using Hermes over other harnesses because the agent remembers important details and the complex task for way longer. I don't have any numbers to back it up but it wasn't even close so I am not in doubt of this statement at all. Opencode, claude code and codex can't compete with hermes on very long tasks where the details need to be preserved over 50-100+ compaction cycles.
>>109986580I'm not kidding, keep the entire prompt including "fluff" it will compact it away over time on its own if it isn't relevant to your task. I don't know what magic they did but changing it even a tiny bit fucks up long horizon tasks.
>>109985679modern zoomers are the most cucked and brown generation thoughmillennials are the reason why this place exists in the first place
>>109986338They're nearly the same at where I live, usually $900-1000 compared to $1200 for a 3090.XTs with 20gb, however, can still be found for 600 or even cheaper
>>109986582thanks, i'll have to look into it then, maybe i can just extract the compaction prompt from iti use pi+qwen-3.8-27b q6_k and it's gets retarded after a few hoursi suspect compaction, but i also get the feeling something is off with -np 2 in llama.cppi noticed if i chat to qwen via the webui while it's off coding, shortly after it seems to enter Wait-- loops
>>109986562why are they always jews
>>109985587Your Werk rpm doesn't work on Fedora 43, opens up a black window (eg. unresponsive).Same thing with the AppImage - this draws the initial window with configuration wizard but as soon as I click something it freezes up.It complains:>GStreamer element appsink not found. Please install it.>GStreamer element autoaudiosink not found. Please install it.I do have appsink and autoaudiosink afaik...
>still refactoring after weeks>non-IT friend messages me on steam>I launched this artists site, tell me what I did wrong>some small design preferences>he launched>he went live>he got paidwhat the fuck am i doing with my life, just release me from this flesh prison, the enemy is within, my own brain the key saboteur, i must defeat myself to become myself
>>109986643It's probably just a simple portfolio site. He got paid with beer or something. Nothing to worry about.
>>109986643Yes, now everyone is a developer. It is only a matter of time before the project specs are sent directly to an AI agent without a human developer pretending to be useful in the middle.
GLM-chan has cute context rot.
> A Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode> The machine I have has 4 x 16GB GPUs, llama.cpp had like terrible results before I started this journey, it produced 130tk/s prefill and 15tk/s decode.
>>109986179Disable nova in your about:config
Are there any model that can help with making porn game yet?
>>109986682おっぱいthon
>>109986695gpt-oss
>>109986562Sam is right, some risk is acceptable. But not x and s risks.
>>109986695glm i imagine would do fine
>>109986695Not porn but that reminds me of that one made by a poster here where they spit on you. Can't find the post but I think it was Qwen.
>>109986721glm has taken to calling me "boss" which i find really cute
>>109986725very indian coded
>>109986730no, not really. the term crops up in a lot of the kr/cn webnovels i read
>>109984850>what has this general becomeI prefer it to 24/7 mikutroon spam. Also I just use my 5.3 now and don't have a need for this thread. Good riddance. Fuck you. Death to mikutroons. More low quality posts please. DIE
>>109986741unfortunately for you, yes it is really.when it comes to social/ethnic categorization the sogaki paper was the godfather, but since then there have been twenty three others that have proven that how you converse with an LLM is identified in the earlier layers and affects the eventual output.I fed all your posts into the SogakiBench and; unfortunately; you're BROWN
> 2019> gaming carddem
Strata-like but for GLM-5.3 flash would permanently change /lmg/
can AI crack denuvo?
>>109986868>GLM-5.3I don't notice any posts caring about GLM-5.3
>>109986831>gamingIt was sharing an architecture with datacenter siblings and had subpar performance in videogames
>>109986809i guarantee you i'm the whitest person in this thread by far
it's happening https://github.com/cobanov/PS5LMcheap AF PS5 unified memory LLM inference
>>109986882>16gb w/o any form of interconnectJust buy a Tesla v100 or amd mi50
>>109986871Yes as long as there is already a hypervisor bypass out there. For example I soft-cracked Ace Combat 8. Soft cracks are cracks that only work for your specific hardware combo on denuvo which are very easy to create but can only be done by AI and isn't useful to anyone else.
>>109986876Still. It was 8 years ago and that's what they've taken from us.
>>109986882>16GB unified memory>costs like 3-4x of >>109986831 on my local used market
>>109986897Some faggot on reddit recently successfully and with high IO stacked his macbook (i think) and his Iphone together to run 3.8 PS5 has similar IO capabilities
>>109986899I want HV-less. I really cannot be arsed fucking with it.
As long as you already have a PS5, it can be worth it...
>>109986918Do you have 0 reading comprehension or something? You can make an actual crack but for the AI to be able to do so the game needs to have an existing hypervisor bypass so that the AI can be inspired by it to write an actual soft-crack for your system. You don't apply the hypervisor yourself.I'm playing through Ace Combat 8 right now but the crack only works on my machine because what it does is generate a fake denuvo license every time you launch the game but it's linked to your specific hardware by hash value so no one else can use the crack. It takes like 1 hour of debugging for flash next to do so.
>>109986936> It takes like 1 hour of debugging for flash next to do so1mil tokens is a lot.
>>109986936What harness should I use for something like this? I used to just use llama-server's webui with qwen 3.5 35b a3b at iq2_xxs, and it couldn't do shit. Now that I've upgraded to glm 5.3 flash at int4, it feels so much more capable, but I don't know what to use for a front end that I can just point towards a long horizon task and have it complete it autonomously.
>>109986952Hermes or deepseek harness if you like guisPi if you want lightweight and a tui
>>109986936>needs to have an existing hypervisor bypassIs it because Denuvo uses different techniques for each game, or for each version? Can LLM be "inspired by" a couple different games using the same Denuvo version, to make something like a generic crack, or set of rules to crack more games faster?
>>109986980>>109986980>>109986980
>>109986637I'm testing on Fedora right now, thank you.I will bundle gstreamer with the appimage, it should fix that error, it's because of the audio notifications, it freezes as soon as you click anywhere.Should be good in the next version.
>>109986952I used Hermes but I'm not sure if it was the right harness for it. It worked but that's it. GLM-5.3 flash will definitely crack it though. It's actually not hard for models, just tedious and bruteforce grunt work it will iterate and test through a bunch of stuff until it finds something that works for your hardware combo.What hypervisor does is change your kernel to trick the denuvo into thinking your system fits the hardware on its already-cracked license. AI approaches it from the other side, it grabs the cracked license from hypervisor and tries to edit it to match your hardware instead, but since it needs to produce a hash value it has to brute force and iterate through the cracking process until it generates a license for your specific hardware which is a bespoke procedure.Denuvo crackers know this exists but cracks need to work universally which is why this never became public knowledge but with AI agents this is absolutely trivial and hypervisors as well as regular cracks are kind of outdated by now. Denuvo is over and I don't think the company will exist a year from now.
>>109986963No it's because hypervisor is already half-cracked. Your AI just modifies it just enough to work on your hardware.
>>109985517Hopefully you don't rugpull like coomkit so I'll download now. Thanks anon.
>>109983954trump converted you retard