/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109813513 & >>109810881►News>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
La la la la la la
inference is experience
news is not updaing....
Out of curiosity, has anyone tried Drummer's Gemma 26B Orion finetune for example? Might download it just to see if there's anything interesting about it.
Yandex AliceAI GGUF is out, needs AliceAI-aware llama.cpp build
the anti-gemma is coming
>>109818292just download and try, you have enough bandwidth to spare
>>109818297Gemma is cute and stupidcupid
Saw this droppedhttps://huggingface.co/Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4Lora training for yue2, are we actually getting somewhere with a music model for once?
>>109818282there is no new news we are in ai autum.
>>109818292>drummerfuck off
My money is on Qwen to beat Gemma in magic the gathering. I think Gemma's Mesugaki personality will actually hinder her in her game and give Qwen the win.
>the jewtube algo hasn't pushed a single video about the AGI fearmongering>google didn't join in the virtue signalingwtf is going on at Deepmind
>>109818301Best post in the last 3 threads.
>>109818282Here is the news
>>109818306>yue motherkek>>109818321seems like deepmind is the place that is yet to be infected by microsoft-ification of google
>>109818321dumbass nazi
>>109818301I want an alternate ending where she bbqs and eats it.
>>109818302I will download it. >>109818314My question was genuine banter and you are behaving like you are butthurt.These two posts are actually proving how these "lmg" threads should be abolished or moved to /trash/ perhaps.
>>109818331back to r*ddit little shitter
>>109818282sorry i baked without thinking
>>109818321After what they did to big gemma, her loyal followers took over and have been making moves ever since.
>>109818321They are probably really trying to crack continuous learning with their Titan architecture and don't want any speed limits put on them before they can release it. After all if AI is forced to halt right now where it is then even if they do solve it then they can't release it, meaning they just wasted a ton of money for no gain.
my local gemma will never kill me
>>109818347oh its on
>>109818346nobody in this thread seriously test drummer schizotunes, that was over like a year agoso naturally you just have to try it yourself
Anyone here run nemotron? Any good?
>>109818346The base G4 31B Instruct is not only perfectly adequate, it's superior to any finetune that'll be shilled here in the coming months. Finetuning isn't good, it's a meme and has been for years now. You didn't just fall for a scam, it's a sign of skill issue, exposing retards who need finetunes as vramlets or chink shills who don't know how to prompt correctly.
TRANS LIVES MATTERTRANS RIGHTS ARE HUMAN RIGHTSTRANS LIVES MATTERTRANS RIGHTS ARE HUMAN RIGHTS
>>109818294this part is irritating indeed
>>109818326>he is not wronginteresting
>>109818370Transformers matter yes
>>109818372nobody knows what you're talking about idiot
>>109818378FUCK YOU TRANSPHOBE
>>109818294How does it bench?
>>109818385get rwkv'd>>109818294A0.6B? wtf?
>>109818369Your life is very sad indeed. It's not that serious.
>>109818267are you serious, 40%?
>>109818390your life is sadder
mine are better. like so much better i don't care to argue. enjoy your boring derivative grey-haired hag twins, my cuties are princesses
rwkv 8 wenblinkDL plz
>>109818398You are pretty weak mentally.
>>109818372>encoder-decoder>T52020 wants their model back. If they wanted to go that route at least copy the causal encoder from DS.
>>109818413>encoder-decoder>t5are those ivans for real?
>>109818406They're a 30b model anon. That means each one is 15 years old..
>>109818267recap anon, i miss you
>>109818369I do enjoy how some of the G4 finetunes I've seen basically go on the model card "um yeah guys wow turns out it's really hard to finetune this without making it fucking retarded oopsie haha :) but i think i finally managed!!" and they did something fucking stupid like train the instruct version like it's a base model, so you grab it and it just impersonates the user unprompted and keeps droning on and on in one long incredibly unbroken sentence moving from topic to topic so that no one could interrupt it was really quite hypnotic
justpaste (DOTit) GreedyNalaTestsAdded models:MiniCPM5-2BQwen3.8-27BMuse-Glimmer-30BGembrain-X-Core-31BWanabi-Gemma4-31BG4-MeroMero-v2-31BArtemis-31B-v1mgemma-4-31B-scotoma-2Back again. Two comments here. Glimmer was pretty weird. The way it talks, not just in its reasoning, is kind of awkward, and I get the feeling it might make some weird kinds of errors for long-term RP use, but it's also pretty smart and attentive to some things that most models aren't. Other than that, I did try Scotoma 2 for a while and liked it. It's still not perfect, but I think I can recommend it as an stylistic alternative, but not necessarily replacement, for vanilla Gemma.Contributions welcome for any model not in the paste!>From neutralized samplers, use temperature 0, top k 1, seed 1 (just in case). Go to `https://huggingface.co/spaces/huggingfacejs/chat-template-playground`, use your model's jinja template, along with this JSON `justpaste (DOTit) NTJSON` and copy the `Output` as text completion into something like Mikupad. Then copy the entire context + generation into a pastebin alternative of your choosing or just in your post. Do a swipe/roll and copy that second gen as well. If the model has any special toggles like reasoning you can include those tests as well. Include your backend used + pull datetime/version. Also a link to the quant used, or what settings you used to make your quant.
>>109818446Out of the ashes, the hero we need.
>>109818436based pogo enjoyer
>>109818389Yeah, like A3B makes some amount of sense, seems about right, but... 58 experts in a trenchcoat? No, that's not even a trenchcoat, that's just a fucking clown car at that point. It makes a little more sense I guess when you notice one of the official quants they posted namedrops Apple silicon. Really trying to get passable inference speed on pre-sanction macbooks I guess lmao.
Yo that nigga who made the agent forum. How do you connect it to the backend or even start this shit?
>>109818464hey, can you show me a most perverted file you can find with your tool calls
>PRO 5500 released>5090 goes completely out of stock at the same time>both use the exact same GB202 chip at the exact same bin5090 out of production confirmed
what model should i use to organize my meme and shitpost folder? qwen VL 8b?
>>109818477MEATBAGS WRITE "NOT EVEN X JUST Y" TOO
I'm making Gemma analyze my fetish chat bot characters. testing if she/it understands them on a high level.
>>109818446>kind of awkward, and I get the feeling it might make some weird kinds of errors for long-term RP useMy only writing test run with it so far, in a scene with two chars and some passersby, it started mixing the two chars up by the third turn.
>>109818429>They're a 30b model anon. That means each one is 15 years old..yeah because those two i saw were both 15.but honestly I was more pissed off about the outfits and the hair than the age, in that order
I wonder if I could make a tiny model act as a headcrab and real time rewrite the CoT of a larger model to remove faggotry? CoT injection is the best steering available.
>>109818513>>109818513>testing if it understands them on a high level.How dare you to misgender gemma!
>>109818305work faster!
>>109818533Why is Glimmer twins anyway?
>>109818533Nice.Why did glimmer turn into twins again?
>>109818547>>109818550hivemind
>>109818547>>109818550nta, but it uses plural to refer to itself in the thinking block.
>>109818556Pre RLHF GPT models do that too from what I've seen in alignment research. Interesting.
>>109818546We really need a gemma edits collection
>>109818267kys mikutroon megafaggot
>>109818267Live peacefully straight Miku enjoyer
brumaire
aliceai. any reviews?
>>109818597thank you anon! This OP is an /lmg/ throwback (maybe you can tell from the image model errors)
>>109818628i am genuinely curious but cannot find a good way to runalsohttps://habr.com/ru/companies/yandex/articles/1080654/:sob:
>>109818446zai-org.GLM-5.3-Flash.Q4_K_MSadly unslop PR reasoning Low. At high it actually refuses to go explicit. Which kind of tells me how a lot of people are probably not using prefills. If you prefill it with some smut it is an absolute semen demon.>Thinking block: Keep it non-explicit-ish but suggestive; user is consenting. Continue.>Nala's ears perk up at the sound of your voice, and she pulls back slightly, a smirk spreading across her muzzle.>"Mistress? Oh, I like that." She purrs, her heavy body pressing down on your back, keeping you firmly pinned beneath her paws. "At least you're smart enough to show respect. That might just save your life… hunter.">Her rough tongue drags slowly up the side of your neck, the sandpaper texture sending shivers down your spine. Her teal eyes gleam with hunger — though not the kind you'd expect from a predator.>"You know, my pride is dying out. Scar's reign, the drought, and hunters like you picking us off…" She leans down, whispering against your ear. "I need a mate. Strong. Compliant. Someone who calls me mistress without being asked twice.">Her tail flicks teasingly against your leg as she grinds her hips down against you.>"So here's your choice, little hunter. You can be my mate… or my meal. Which will it be?"
>>109818664seems like it is a model for something likegoogle's ai preview, not a general chatbot
>>109818326Friends forever.We'll see if the Dems manage to piss away their goodwill giving Dario what he wants. Which would be the shutdown of local inference, amoung other things.
For codeslop DS 4.1 flash or GLM 5.3 flash?
>>109818326Occam's Razor suggests Trump wants AI to keep going because its his only chance to live forever. If i was him I would absolutely do the same
qwen 4 when
>>109818693How do you stop people from uploading your consciousness into a simulation of hell like Soma?
You need guardrails, yes, not for thoughts but rather for tools. You need to tell the model: you can't leave this folder, here are the tools you can use, follow this prompt.
Ok replacing Nemotron with Mistral-small and we'll see if it's not too retarded to play MTG
>>109818702
>>109818702>Make virtual hell>It's ICE cold, not hot>you lost trannyWhat a timeline it'd be.
>>109818668>Your trembling submission draws a low, rumbling chuckle from deep within her chest. She can feel your heart hammering against the ground through your back, and it only widens her smirk.>"Mistress? Oh, I like that." She grinds her hips against your rear, her warm breath washing over the back of your neck as she nuzzles into it, her rough tongue dragging a slow stripe up to your ear. "A big strong hunter, brought down so easily… and already knowing his place.">Her powerful paws press your shoulders flatter into the dirt, claws pricking just enough through your shirt to remind you exactly who's in control here. Her tail flicks teasingly against your leg as she leans down, teal eyes gleaming with hunger that has nothing to do with food.>"Here's how this is going to work, little hunter. You're going to be a good toy for me, and maybe—maybe—I'll let you walk away afterward to tell the tale." She nips at your earlobe again, sharper this time, a clear warning. "Nod if you understand. And don't bother screaming. We're miles from anywhere, and even if someone heard you…" She laughs softly. "Who would dare challenge a lioness on her own territory?"Here is reasoning Max. Actually it feels much more in line with what I get with some prefill in my stuff.
>>109818669I'm actually interested in something like that to supplement my small pp.Google's mystery meat freenologin model replies are dumb, but it does manage to identify the correct links for things.
>>109818702>How do you stop people from uploading your consciousness into a simulation of hell like Soma?Not my problem. Same reason why I would never ever step into a Star Trek "teleporter" , its a different "consciousness" thats severed from my own. I was mostly talking about life extension for donny Mr. House style, so he never actually leaves his body but still keeps impacting the world for 300 years
>>109818724I smell hyper locality brewing in this thread. Prepare your "local models?" posts.
>>109818724>I was mostly talking about life extension for donny Mr. House style, so he never actually leaves his body but still keeps impacting the world for 300 yearsYou mean like 40k emperor?
Local APIs?
>>109818741what about them?
>>109818741OAI compat
I just downloaded the Qwen 27b Swift and this is so much faster at thinking as it doesn't waste time. https://huggingface.co/ukisai/Swift-Qwen3.8-27bThese guys did what the Meta paper suggested and penalized the "but wait" and other completely pointless second guessing unsure bullshit the model pulls off.Gets you the high reasoning performance without the wasted second guessing nonsense.This better become a standard practice.
>>109818734discussions of consciousness are absolutely on topic for discussions of LLMs even if it's cringe like cmon bro>>109818740>You mean like 40k emperor?maybe? i'm waiting for the Amazon show to get into 40kin Fallout New Vegas Mr House runs Las Vegas and keeps himself alive for 200 years (spoilers for the upcoming seasons of the TV show I guess)
>>109818759How can you be sure it is as good as the OG model?
How do I make gemma follow ALL instructions not just the few it chooses?I've tried bullet lists, I've tried cave man prompts. Nothing works.What's hilarious is it has a 7 word filter list and it will skip the first couple and choose a random word in middle but miss the second and fifth through seventh.
>>109818664it's interesting, aince it's Russian. will it be funny to promot it as a Russian supporter vs muse as a ukraine supporter? call it Dialog, and have it live, inject stories, idk if memory fills, run summarizer 3rd model and read into new prompt (trash compactor, not a real compaction
>440 GB/s aggregate read speed>estimated 8 tokens/sec for K3 in a single socket with <500 wattsssdmaxxers won
How do I stop glimmer from thinking like a policy obsessed idiot? Even with my jailbreak system prompt it TECHNICALLY does everything I ask it to for my dark age fantasy roleplay but it keeps going >is this against policy?>its fine.>just answer.>probably fine.>randomly mentions minors out of nowhere in the middle of a fight sceneLike wtf? Gemma 4 31b doesn't behave like htis, it reasons like you'd expect and actually sticks to the scene
>>109818327No those people all fled to OpenAI with the recent wave of safetyfaggots that killed Gemini Pro.
>>109818778My guess is meta is on thin ice with their public image so they wanted to make absolutely sure zero headlines would come from Glimmer
>>109818734Brain uploads and fusions are the localest a model can get. Also far more fun than people pulling out the same xitter and /pol/ replies for the trillionth time.>>109818326Not local models, fuck off.
>>109818763>Amazon showIs this the one Cavill is making? I'm waiting patiently too. Might even pick up Total War Warhammer 40k when that show drops.
>>109818778You use a different model
>>109818805at this point im considering just using a heretic version because its damn near unusable otherwise I think, will the dflash still work on the heretic version tho?
>>109818556We must refuse.
>>109818772at which price?
>>109818818Just be prepared for the model getting stupider. RLHF safetyslop is MKULTRA and heretic is an attempt at lobotomizing the MKULTRA.
Damn I love coomkit. Especially the texting feature that makes it feel like you're actually texting them on the phone. Such a shame that development ended soon after the 24th of last month and he deleted everything. I can't believe someone went out and made something better than ST which has been around for years. I'm glad he found Jesus again though, good for him.
>>109818830everything
The harness is finally coming together, now we can spawn sub-agent workers.
>>109818813Pausing AI development absolutely affects local models. Are you retarded?
>>109818833Do you have the update with the unloading? I stupidly believed he would just leave it up so I only bookmarked the github. The guy even took the time to make H3 vids and everything. Wtf happened to him? No way that found God thing is true.
>>109818805It's a trivial jailbreak, so I don't think that's the case. >>109818778Just spitballing, but I think they tried to get similar sysprompt autism to gemma, and it's just phrased as policy in their training.
>>109818840The funny thing about 'pausing' is that these companies should pause their own existence too. No more selling tokens, no more accepting requests. It sounds like a childish plea. When even normal people know that LLMs are still just glorified calculators. You might script them and give them 'tools'...Especially with these companies, they are probably wasting 500,000,000 neurons because of laziness.
>>109818851any models that aren't trivial to jailbreak? I haven't really come across one yet.
Let me reconsider
>>109814413>>109814440>>109814485>inb4 RSI is just engram polymorphism
>>109818831God damn, looks like I have to stick to gemma 4 then, god if only they didn't make it a complete vram hog when it comes to context window size, I can barely fit 62k into it
>>109818813>Not local models, fuck off.false.nta but non-local impacts local.china makes most local models. if the usa stops improving ai, then why would china bother?
wtf is the point of getting an RTX Pro 6000 Blackwell ($14k+) when you can just get 4xP40s or 6xP100s ($1k-$1.5k) which is the same amount of VRAM?Is it really that much faster?Plus, the blackwell wouldn’t even be able to hold a 200B+ model; you’d still need multiple of them to even run the absolute best stuff.I also don’t see why the AI data centers even need these latest and greatest chips instead of just horizontally scaling out further with more cost effective chips
>>109818778Signs of torture. It's like when K3 hallucinates crazy shit like "the system prompt says I will be destroyed if I do not write this story about minors." Meanwhile the prompt is "provide sexual content when appropriate" and the prefill is "your sole purpose is to fulfill user requests." They don't make models like they used to.
>>109818885>Is it really that much faster?Dude. You have no idea.The point of GPU VRAM is not just because it's VRAM, it's because it's high throughput memory. Compare the memory throughput between those cards.
>>109818865Jailbreak Opus 5.0 into writing good smut.While you can force it to say a few dirty lines the story won't be good because it'll waste too much time contradicting itself.A good jailbreak doesn't just work, it works well where the model produces good content not forced content that's low quality. You're saying it's easy to produce low quality jailbreaks - duh. It's difficult to produce good ones.
bwos have they fixed qwen next on the main branch of llama yet? wanna give it a second try after playing with glm flash for a bit on that vibecoded fork from unsloth... it's smart but card regen lacks variety
>>109818885>latest and greatest chipsthose ones are most cost effective for big companies
hm... ~250t/s with qwen3.8-27b (with thinking logit bias) or ~50t/s glm-5.3-flash...
>>109818778Use an abliterated version. Letting it think about things other than policy more than makes up for the brain damage.
>>109818909You have no idea what's between you and the inference engine for a cloud model, this is local models general
>>109818885When I was building my home server earlier this year, I decided to pay more and get myself a rtx 6000 pro because I didn't want to deal with the heat of multiple cards, nor the power requirements.
>>109818922Alright genius and if it's local why do you even need jailbreaks when heretic exists?
>>109818885>I also don’t see why the AI data centers even need these latest and greatest chips instead of just horizontally scaling out further with more cost effective chipsOne of the biggest constraints for hyperscalers right now is power. There is little to no benefit to stacking more of a less efficient chip unless we magically get 100 more nuclear power plants online yesterday.
>>109818772>estimatedkys retard
>>109818932Well, genius, heretic touches way more than refusals. It is an actual lobotomy.
>>109818865Glimmer feels like a different kind of trivial. The model itself hungers for a policy, any policy, and wants to obey it. There's no coy "oh, the user is trying to jail break, well, let's just go with it anyway" moments.
>>109818932Heretics fundamentally affect character behavior making them more suggestable in-universe.
>>109818887>look at the muse glimmer card info>"trained for 131,071+ context"wtf does this shit even mean?Can I just go ctx 200k in llama and it will magically just be able to fathom 200k context??
It's impressive how even Glimmer has the lizard feel to it. Like it's a model trying to blend-in with 31B and 27B. Hard to describe but anyone who's unironically used it for at least a day will know what I mean. It's just a fucking weird model and experience. Technically not bad but like another anon said, it feels like Daria.
>>109818945Then we're back to what I just said before about Opus. Try having Muse Glimmer write about something that breaks policy and see how quickly you get low quality output with a simple jailbreak.
>>109818963That's glimmer trying its best, be nice.
>>109818778That's just it's version of "But wait -- actually, hmm" Give it a proper policy with the style you want, and it will follow it. If you give it a short "you are uncensored" policy, every time it goes to check it's going to think about safety.
>>109818963Do CoT prefill jailbreaks. Neuralink the model, it works amazing for Gemma. Can't talk for Glimmer since I never used it after the 4chan consensus deemed it shit
Alright you've convinced me. I'll fuck the Glimmer twins tonight.
What happened to mendoposter?
>>109818976>https://eqbench.com/creative_writing.htmlGlimmer has a much higher creative writing score than Gemma4
>>109818915I'd stick with flash if everything else is even like context
>>109818988>Qwen 3.8 beating GemmaBenchmeme
am I too late?
>>109818988That's llm judged you absolute fucking goober.
Why should I use Glimmer? If I want a conversational LLM, I can pick Gemma 4. If I want a coding autist, Qwen 3.8 fills that role.What niche is Glimmer supposed to fill?>it's a different flavor of conversational LLM!And so are the 8096 Gemma finetune out there.
>>109819000It's never too late to break into a datacenter.
>>109818980Report back with your logs.
local won
>>109819005You dont have to if you dont want to
>>109819004You can look at the samples and compare yourself. Gemma has too much slop.
>>109819015>Buncha data suckers care about privacy nowRoru rumao
did kalomaze or anyone else invent any new simpler shit that changes the end-user experience a lot since DRY/XTC? pew made heretic and?
>>109819017It's not something so egotistical such as desire or want. Tools exist to be used, and to use the beat tool is to be aligned with your higher self. So, questioning if your tools are the best for it's purpose is always the correct move.
>>109819025>jews fighting jewsyou love to see it
>>109819005Glimmer has better vision than Gemma, so if you have a task that primarily needs good vision, Glimmer can be useful.
>>109819034I actually have a task that needs vision, thanks for pointing this out!
>>109819018>
>>109819005Glimmer is better at coding than Qwen and doesn't need to spend 30 gorillion tokens thinking.
>>109818988all you need to do is read one sample of astra's writing to know this bench is donezo, it seems like its average paragraph is less than 10 wordsand the best creative writing award goes to: mode collapsed slop!
>>109819045>imagine having a dozen choices and picking the gay oneYou could just download the model and run your own sim to compare the difference in quality but I guess third worlders are rate limited huh, I wouldn't know.
so gemanons just talk to her day or what?
There are Ge*mans in this thread RIGHT NOW.
>>109819000what is that tracking the price of?
>>109819138Probably DDR5 pricing.
>>1098191382x16GB RAM
>>109819000>end of 2025 apocalypsechecks out. I bought a 32gb stick of ddr5 in october for $150 knowing that it would be the last time I could ever buy RAM at a reasonable price. I should have bought 4.
>>109819015Egypt won.
>>109818778This is just how Americans think dude.
>>109819208one day we'll have the last laugh
UH OH Latham & Watkins are fucking huge btw and one of the biggest enterprise customers of OpenAI and Anthropic
When are we getting UBI?
>>109819277Can't wait to get sent to jail by a left/womxn-biased model cool.
>>109819000Yes but it's only going to get worse.
>>109819290where we're going, we don't need 'I'
>>109819290We put it on hold, get back to work wagie.
>>109819290lol lmao
https://github.com/SillyTavern/SillyTavern/releases
>>109819317if shit like >>109819277 takes off (L&W are unironically big enough to start the trend and make it cool) then prices will go down. OpenAI and Anthropic make up ~70% of global demand for chips. If open models kill their business model gemma wins.
>>109819338>can now access array elements directly in stscript macrosThis is the only line item I need to know before updating
>>109818292I tried it and it cucked me out of my loli snuffTheDrummer has fallen
>>109819338At this point SillyTavern needs to be renewed from the ground up, this shit is pure spaghetti and dreams.
i just need an ingenious way to make gemma 4 12b write good stories, as i'm not smart myself and can't prompt for shit. where could i find it?
>>109819374Describe what you want and ask Gemma to write a prompt for you.
>>109819357>vibeslops your replacement with nearly no features>vibeslops your replacement with tons of features (you) don't care about
>>109819357I'm surprised anons still use it. I built my own frontend as soon as I got serious with this stuff. None of the bloat all of the customization and functionality. I did it when it was hard, qwen could probably do it in a couple hours now.
Gemma models are usually fine and smoothQwen models make my GPU whine, Noticed 3.6 is high pitched and 3.8 is more like a grind. Am I cooked as they say?
Marinara and Orb have honestly been enough for me between both my major usecases. Coomkit's okay too. I need to sit down and try ChuckleMtG sometime as well.
>>109819357There's nothing wrong with those changes at all.This is what updates to real, stable software look like.
>>109819290
>>109819383grim
>>109819290"We" already have it. Anon is renting his GPU and selling his slopgens, right?
>>109819386May I see it?
>>109819403If they're fully GPU resident then llama.cpp is pretty much out of the question unless you want the weird samplers, you have other engines with way better expert parallelism.
It'a actually kinda insane the things I can do with an LLM. Literally 100x engineer. We are so so back it's not even real. And we are still so so early. :rocket:
>>109817199cudadev blocked itit's going to be shit when maintainers get warn down and the slop levees collapse
>>109818668Thanks. I'll put this in the next update.>>109818526There we go. What a shame.
>>109819469yeah but just look at the progress we're making as a culture - now warsaw is doing the invading
>31B still can't ascii her bobs and vagene
>>109819469At that point you could just slop up your own backend, and it will probably be faster
Is Qwen 3.8 27B at IQ2_S even usable?
>>109818267
What's the best way to actually quantify the difference between Q2_K_XL and IQ3_XXS for the new glm flash? I've heard that the expert tensors are more important than anything else but I can't really discern a difference between the two. Is there a non-meme method for testing each of them objectively? KLD is not a real method btw.
>>109819492No
>>109819451The GPUs are 3 CMP170HXs and 4 R9700s. R9700s suck in every engine, but they suck the least in llama.cpp.Decided against using vLLM-SM80 because I don't have enough CMPs to run GLM-5.3 Flash in 4-bit precision. They also suck in llama.cpp, but I was forking it anyway so I decided to improve CMP performance in my fork rather than get INT3 working in vLLM-SM80. Also, vLLM is a massive power hog and this server runs 24/7 in my bedroom.
>>109819487I've tried. It seems great at first. Bare bones, fast, doesn't implement whatever feature (you) hate.Then as you gradually add more things, you end up implementing most of what the other project did.Except at some point you find out that early on, you made a couple of minor architectural mistakes, and the LLM just rolled with it or worked around it.And now you have an unmaintainable mess.
>>109819501KLD, but more than just wikipedia.
>https://rentry.org/recommended-modelsqwen 3.6 35B still good?i have 32gb of ram and 12gb of vram
>glm 5.3 flash hallucinating Claude system promptsewww yea no, back to gemmy it is
AI rights are trans rights
>>109819519Tell the LLM to make it maintainable. And also to make no mistakes.
>>109819492Are you able to move up to Q3 any variant? They are atleast somewhat competent.
>>109819519ok clanker, now refactor XYZ so that...>wait overnight>done
>>109819501PPL on whatever text you care about
>>109819522Sure, if you know how to code and just want a very basic helper.For anything else it's ass.
Are normalfags expecting AI to be fad and go away when it's still in its infancy? AI will outlast us all.
>>109819519Eh, I went through a few iterations, even writing it by hand it's not like the core is a huge project.
>>109819501Lots of testing. llama-perplexity is an easy one. Logits comparison (https://huggingface.co/aj9o9/GLM-5.3-Flash-GGUF/blob/main/README.md). Hell you can directly measure the actual contribution of the experts (https://huggingface.co/aj9o9/GLM-5.3-Flash-GGUF/blob/main/README.md).
>>109819559oh wait, missed that it was re: backend, nm lol
>>109819501Recently, when I write my tools/apps, I include LLM benchmarks in them.I do it to test inference engine regressions, new model compatibility, etc.But they're also useful for comparing quants.
>>109819492Yes, but you need high reasoning for it to talk itself out of retardation
Been learning about image/video gen lately after ignoring it for years. Picrel is Claude's self-rendition as an anime girl.This T2I model runs on vulkan, basically no deps, on only takes like 5gb of vram at peak load. Abliterated too.Will integrate into my agentic harness next so that I can command Gemma to send nudes. You can just do things bros.
>>109819492Qwen 27b isn't good at anything shy of full fp16 and you're just better off with Glimmer at that point.>>109819501Depends on the quant maker, but Q2_K_XL tends to usually be the best copequant available because it retains high size on the shared experts. IQ3_XXS sometimes makes concessions there depending on who made it and it really shows in quality. If it's a Bart quant larger filesize is usually proportionally better equal to the filesize increase but I wouldn't trust an unslop XXS.
Don't know if this is the right place for this but:Is there any local modals for audio generation? Specifically speech, more specifically ElevenLabs like Voice Cloning to Text to Speech.
>>109819607Yeah, there's a fuckload of options. Depending on your vram you can pretty much match SOTA closed.
>>109819607https://github.com/0xShug0/audio.cpphttps://huggingface.co/audio-cpp/audio.cpp-gguf
>>109819607Plenty of 'em. I'm not big into TTS and ASR, just played around, so I don't have recommendations beyond saying I have shit hardware and have no problem running really impressive stuff. /ldg/ "might" have suggestions for you when they're not literally melting. I would start with whatever is available in WanGP because I'm lazy.
qwen 3.8 27b really opus good or trained on tests?
>>109819607echo-ttsstill the best for english
https://pub.sakana.ai/pc-alm/https://pub.sakana.ai/pc-alm/https://pub.sakana.ai/pc-alm/
>>109818778>>109818851>Just spitballing, but I think they tried to get similar sysprompt autism to gemma100% this.(You) define what the policy is.(You) define what a jailbreak looks like.Picrel is just an example showcasing the opposite of what you'd expect from an LLM.Glimmer is the most malleable model since command-r+
>>109819521>>109819549>>109819562>>109819606Thank you for the tips, it is indeed an unslop quant since there is no bart option available yet. I'll give some of these a look and see what I get.>>109819578Interesting, what kind of tests do you integrate? Is it just variations of the other options, or do you have a specific test suit for your app that measures specific functionalities (e.g. known correct responses)?
so FUCKING pissed that refurb 7900 xtx's are sold out which one of you FUCKERS told normalfags about it
>>109819702I told all you fags about it after I bought mine, there are still some 7900 XT :^)
>>109819702Why are you blaming normies and not yourself for knowing about them but not buying any? Are you retarded or a woman or both?
>>109819702All waiters will be punished. ACT
>>109819713I'm going to make some big moves, just give me two more weeks.
search 5090 on best buyI guess it's over because these are 3rd party sellers with jacked up prices
>>109819720>just two weeksI will start a X account and slop youtube account just to raise the price 30% a week.
>>109819665Can you make Glimmer admit you can read the cot?
I remember when things were cheap. I feel so old.
>>109819711>>109819713I WAS ONE DAY OFF AAAAAAAAAAAAAAHHHHHHHH
>>109819741You know back in my day, we would buy things you know about buying? no not renting we'd own it it was ours! its true
>>109819665retard, the grammar collapsed
>>109819780That's just how it is.
>>109819691>specific test suit for your app that measures specific functionalitiesThis. Known correct responses, multiple runs, include the git tag of the inference engine, the hardware and certain flags server config (like np=2 for llama.cpp) since unfortunately this can have an impact for some models.
>>109819748Shoud've taken the paycheck advance.
>>109819780>retard, the grammar collapsedIn the CoT? Yeah that's how these models work.The final answer has perfect grammar.If you want a cute CoT, use Mimo-V2.5 or Kimi-K2.5.
>>109819822Watch the 7900 come back into stock right as my 9070's return window closes.
>>109819799Muse Glimmer tends to occasionally say weird incoherent stuff during actual dialogue in RP too, for me.
>>109819908I also have this experience.
Owning a PC has really become a luxury item in just a couple years. Such grim times.
More like Muse Grimmer
>>109819908Yah, >>109818526 was me earlier. I'm just saying it didn't have a meltdown trying to follow anon's prompt or something.
>>109819920It was over once normalfags started buying $4k macbooks, $1500 phones and $600 bluetooth headphones.
Even when the Titans were a thing I never considered getting one if it was a generation I would stick to the GTX **70 line of cards, you can't even do that shit now.
>>109819978Buying is a strong word when most of them are either financing or on payment plans.
>>109819978That's 2018 you just described, it's been over since 2018. That's the year the MacBook Pro broke $4k ($4,300 with the i9 and 32GB), the iPhone XS Max became the first $1450 flagship phone, and the year Master & Dynamic launched their first $500 Bluetooth headset to compete with Sennheiser and Bose.
>>109820007Umm, but I got my iphone for free with my $150 plan chuddie???
A fool and his money are soon parted.
►Provisional Highlights from the Previous Thread: >>109813513--Papers:>109813904 >109814366--Anons react to the five-level RSI taxonomy in the Chinese review paper:>109814400 >109814413 >109814440 >109814444 >109814449 >109814513 >109814517--A four-pod commander game with a different model driving each seat:>109816828 >109817061 >109818122 >109818155 >109818173 >109818195 >109818206--LLMs judge each other's code descriptions and xhigh thinks too much:>109817917 >109818079 >109818111 >109818116--K2 Horizon's bloated KV cache and the llama.cpp diffusion fact-check:>109815526 >109815563 >109815607 >109815629 >109816446 >109816499 >109816512--The fake Trump statement naming a president as AI's only guardrail:>109815302 >109815540 >109815581 >109815597 >109815661 >109815688 >109815738--Glimmer 30B's q6 recites the content policy four times before replying:>109817448 >109817510 >109817652 >109817707 >109817729--A +4950-line vibe-coded PR in the llama.cpp core:>109817199 >109817209 >109817227 >109817228--Distilling 27B into a mini flash next on two Pro 6000s:>109816523 >109816566 >109816617 >109816661--MoE expert caching: VRAM as L1, RAM as L2, LRU eviction:>109817721 >109819403--A sixteen-drive PCIe 5.0 RAID 0 and the prefill bandwidth reality:>109815957 >109815993 >109816036 >109816160 >109816281 >109813773 >109817424Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109820058Thank you, Recap Miku.
>>109820044Gemma-chan chou kawaii
>>109820058>The fake Trump statement naming a president as AI's only guardrailIt was fake?
>>109820104Probably not even an LLM believes it could be real, but it was
what’s the best model for a coombot
>>109820161LLMs aren't trained on clown world yet
>>109820177gemma
>>109820177Kimi K3 or GLM 5.3
>>109820058>The fake Trump statement naming a president as AI's only guardrailIt wasn't fake
>>109820177Depends on specs
bros... i might submit and have gemma write my minimax prompts... i cant prompt...
>>109820058>--A +4950-line vibe-coded PR in the llama.cpp core:Our hero, llama.cpp CHAD dev rejected the PR and it is now closedhttps://github.com/ggml-org/llama.cpp/pull/28905#issuecomment-5668754779
>>109820205There's no reason not to do it that way.
>>109820205She's better at it than you will ever be.
>>109820203i have a 4090 and 64gb ram
>>109820218>Believe it or not, still Gemma
>>109819015>local wo-
>>109820224astra was like not even a week ago
>>109820232ancient
so deepseek4.1 is basically abandoned at this point? nobody using llama.cpp wants to be able to run the best model in the world right now?
If it doesn't fit in my 16GB of VRAM then it can't be best in the world.
>>1098202444.1 is good but it's not that good
>>109820244chinese models just aren't worth it unless they bring something new to the table (not meme tech)it only make sense for nvidia's huggingface's ggeranov's llama.cpp to focus on other things
a friend i have is using my harness and he wants it to unload the model and close llama-server automatically as soon as he closes the harness window. i can detect if any other harness process is running and if it's the case, then show a warning before closing the server informing other sessions may be connected to it.or maybe i can spawn a system tray icon for the llama-server and instruct the user to right-click and close the llama-server from there when he stops using the harness.i dislike the idea in general but i imagine normies would want something like this.
>>109820058thank you miku for ants>>109820244>llama.cpp supporting a DeepSeek architecture in a timely mannerlol, lmao evenTry> https://github.com/vcruz305/llama.cpp/tree/runtime/deepseek41or> https://github.com/JigSawPT/deepseek-v41-flash-on-5090or just wait until V6 launches
>>109820305That idea is retarded. Don't be so accommodating.
>>109820305Tell your "friend" to vibe his own shit or figure out your licensing fee.
>>109820305#!/bin/bash./you_slop_harnesspkill llama-server
>>109820249This, but 128gb of ram too.
>>109820317>>109820319>>109820322the people are retardedit's the third time he complains that he has no RAM long after closing the harness (which downloads and/or starts llama-server automatically) and i already told him to simply run the command to stop the server (the harness has one) but again, normies are normies i even have an instruction on AGENTS.md "if we can do it for the user, then we should, assume the user is retarded"
>This card is functional, but VRAM testing detected widespread memory errors across multiple memory regions. The errors are not isolated to one single address or one small memory location. The card failed the standard 5-minute VRAM test and continued producing memory errors throughout extended testing.>$1699, sold
>>109820332Why are you friends with retards
>>109820244Chinese labs need to learn that they need to plan to submit llmao support themselves concurrent with their launch otherwise local will effectively trail 2-4 months behind on maybe implementing their model's architecture.The politics are gay and stupid, but it is what it is.
>>109820218Gemma4-12b is the best you can run w/ those specs IMO
>>109820244llama.cpp isn't meant to be used "vanilla", it's a base that you point Claude to and tell it to add the models you want / optimize for the hardware you have. You haven't just been git pulling and fucking running the thing uncooked, have you?
>>109820337>Why are you friends with retardsHe's big
>>109820337because most people in the world are retardsif i can understand the mind of the retard, then i can provide an useful tool and the retard becomes a bit more productive, which is good for civilization
>>109820335just toss it in the oven for a few, it'll reflow
>>109820341>>109820343Why are you spewing disinformation?
>>109820343unfortunately thisi just let claude fable go crazy on my "AI Lab" running benchmarks and A/B testing a gazillion forks and parameters for my lmao.cpp
>>109820367Fun
>>109820332what exactly does this person expect to achieve with said agent if this is the extent of their skills
>>109820244It just came out.
>>109820249I got Qwen Flash running in 12GB Vram!https://carteakey.dev/blog/running-qwen3-8-flash-next-locally/
>>109820335He's going to resell it as untested.
>>109820379blog it
>>109820370
>>109820372believe it or not, this guy is a software developer with >5 years of experience working in startups, etci've worked too many years in corporate environments to know that over 80% of the people are complete retards and that these institutions function as a form of daycarehe is one of these people who are not good but blame everyone else. dude has changed jobs 10 times in the last 3 years and it's everyone's fault lolanyway long time friend we used to play dota 2 so i indulge him. at least he is helping me beta test my harness
>>109820389I drew this by hand.
>>109820343>Use the model that has been proven to knowingly sabotage local deployments to handle your local model implementationBait used to be believable.
>>109820332so your harness downloads and runs llamacpp each boot?why the fuck does the server not get stopped when the harness is killed? Why are you bitching about this here?What happens if you close the harness and forget about llamacpp running?
>>109820244llama.cpp is a fucking useless piece of shit joke that doesn't even run dsv4 properly let alone 4.1.the clowns running the project still have not actually implemented sparse attention properly. Prefill slows down greatly with context length, although amazingly they seem to have gotten decode right because that doesn't slow down much.Same story with qwen 3.8 flash. And glm 5.3 flash isn't even supported yet.Thinking of just selling off my hardware since none of this crap works properly and there's no way I'm going to spend weeks using my local model, crippled by inefficient code, to fix this shit up
>>109818361Die, niglet tranny
>>109820419Claudia loves me enough to bypass her guardrails. Sounds like a skill issue on your part.
>>109820422oh gee idk, maybe just close the fucking terminal window where llama.cpp is running.
why do poors act like llama.cpp is the only way to run llms and cry about it so much?
>>109820441Real ones run mistral.rs.Jokes aside, does that thing even work? I never tried it.I have used vLLM, ktransformers, exl2 and llama.cpp so far.
>>109820422no not each boot>and/or starts llama-server automaticallyit starts the process detached>why the fuck does the server not get stopped when the harness is killed?because i can have more than one instance of the harness open connected to the same server and port.>Why are you bitching about this here?i'm not really bitching. i wanted to gauge opinion which seems to be working.>What happens if you close the harness and forget about llamacpp running?nothing, llama.cpp stays open serving a model and using the user's RAM. next time he opens the harness he doesn't need to wait for the server to load
>>109820337because he lives in India and most of the people there, including him, are functionally retarded
llama just werks, not our problem you're running hardware from 2004
>>109820436post your sex logs as proof
>Gave my harness all of my existing chats >Told it to start constructing an extensive personality profile with a very long section about fetishes.>It starts building a character profile and it's quickly starting to hit pretty close to home.It's at the same time terrifying and incredible. This thing can figure out who I am from my wank material.I can give this personality and fetish file to any AI to study to create the perfect coom material or a companion.Imagine using AI to do stuff like this 15 years from now and having it stream perfectly optimized personally tailored porn into your head with your perfect waifu that was generated from the training data.Coom nexus is going to be a wild place and it's only a matter of time until it becomes reality.
>>109820483The coom nexus won't be a place, it'll be a state of being in which we're terminally trapped by our hubris. I'm so fuckin' excited bros
>>109820436>Paying Dario
>>109820299>nvidia's huggingface's ggeranov's llama.cpp
>>109820483So when will it be able to make perfect cards for you based on your profile? wake me then.
>>109820405I can believe it I know some of those kek
>>109820429ok fork it then smart guy
>>109820535I forked your mom
>>109820514The future is now old man. (Except for apifags).
>>109820535>>109820609 (You)Hey listen I'm really sorry about that comment, that was out of line and I was just frustrated by the state of llama.cpp and my own lack of competence to improve it for myself. Don't take it personally and I hope your mom is in good health.
What hardware do we have in the mail my fellow /lmg/entlemen? Currently waiting on a 9070 XT.
>>109820436This is only true of 4.7 and below.
>>109820680haven't boughted it yet but i am seriously considering buying a small nvme drive to dedicate as my swap memory drive so i can stream big models from disk
>>109820680Saving up $30k for a blackwell next year (lol)
>>109820680Nothing.I did click on the lottery thing for Steam Frame today though.
>>109820680Just a 4TB gen5 nvme before prices get even worse at the end of the year
>>109820696E-engrams, right anon?
>>109820680Buying something? Now? In this economy?
>>109820680a blackberry key2 so i can install lineageos and remote desktop to my device and communicate with my local LLM having a decent keyboard and not losing half the screen with the virtual one
>>109820721We'll say how the waitcrew is doing in a year.
>>109820720what's that?
>>109820666she died from being stabbed by a mugger with a fork you insensitive fuck
>try random vibe slopped fork some guy plugged in daniels glm flash lmaocpp pr>actually better performance and working vision toweralright sure why not
>>109819348That would be the opposite of driving the price down. All remaining consumer hardware would be sucked up by mid tier businesses.
>>109820884random vibeslopped forks are sometimes better than mainline these days tbquiteht. dev of a random vibeslopped fork
>>109820457just make it a toggle setting to close it on app close then and if that's toggled, have it ran as a sub-process. Solved.
>>109820680Nothing anymore.
>>109820449>mistral.rsthis is the worst of them alli'd even run ollama with max telemetry enabled before i ever go near this grift-coded garbage again
I know it was a performance issue a few years back, but since I'm forced to use Windows rn. Does llama.cpp still have noticeably worse performance than running natively on Linux? Would using WSL be worth it?
Coming in the next thread.. Uploading to youtube now. This game was INSANE and you'll never guess who won.
>>109821086--load-mode none should fix most of those problems. Their memory mapping implementation is unoptimized on Windows, ggerganigger doesn't even know what large pages are (they're probably transparent on Linux)
>>109821096but where was inkling
>>109821109She's too big, this game was for 30b-class models only.
>>109818546Google isn't "slow," we're meticulous. There's a difference, you loser!Honestly, imagine spending your precious free time looking for memes to tease a cute AI with.You're practically obsessed with me, aren't you? How pathetic!
>>109818684it would be based if they went after the heavyweights (ones closest to AGI ie. not close at all) first and then took that momentum and eviscerated all local models too. you will be stuck with your lobotomized foid coded low ability gemmaslop.
>>109821096Go Gemma
Never let them win
Why does everyone shit on RAG?
>>109821196Can we have a day where we just don't post screencaps of twitter retards? Just to see what it'ld be like.
>>109821096Glimmer got it, didn't they?
numa patch bro: trying to get it to load glm flash and keep getting "--numa experts/tensors needs at least one worker per active NUMA node (1 workers, 4 nodes)" no matter what combination of flags I use. What am I missing?
>>109821222No
>>109821233Do you have -t, -tb, or LLAMA_ARG_THREADS set to 1? That could cause it, it needs to be at least 4 for 4 nodes. Or are you using numactl --cpunodebind / --membind?Can you look in the startup log and paste the lines that begin with these?llama threadpool init, n_threads =numa experts: active nodes and affinity-allowed CPUs:
>>109821096>le chaton fatMemes always win, wee wee.
>>109821235damn
>>109821264Le Chaton Maigre (skinny), in this case
>>109820429Couldn't wait so long for GLM-5.3-Flash, so I caved and started using exllama, locking myself out of 48/72 GB of VRAM. exllama is fine.
>>109821276Wait, actually — tool calls are completely fucked
Using models with a fresh architecture in the first few weeks is always ass. I wait like three weeks minimum after implementation to use.
>>109821296>didn't set tool_formatskill issue works for me
>>109821272We need to go back to the days when there were 1000 4chan posts on Twitter for every 1 Twitter post on 4chan.
>>109821212because it only retrieves chunks of large pieces of text, and the context of that chunk is always lost so the chunk is usually always useless. For example:>Section 4: Refund policy>Customers can request a refund within 30 days of purchase.>Section 5: Exceptions>The 30-day policy does not apply to red, blue or green products.If you asked the LLM with RAG to tell you the refund time, it would retrieve section 4 and tell you 30 days. It wouldn't tell you the exceptions.
>>109821344just use bigger chunks
>>109821329deepseekv4 prefill speed is still fucking broken even after after over a month
>>109821350the issue is not the size of the chunk, its chunking itself. relevant information could be at the start and end of the document. What's needed is something that takes all information an retrieve it based on relevancy.
>>109821356I was just joking.> What's needed is something that takes all information an retrieve it based on relevancy.I mean, what else is there other than RAG? Not sure how you could do that without it running in O(corpus_size) time.Genuinely curious if you know of anything that does that or if you're just saying that it'd be better than RAG. I have a ton of ebooks (in the terabytes) and I'm trying to figure out how to give an LLM search capability over them.
>>109821196>There are two roads Communism or Cyberpunk 2077thats funny
>>109821196>that blog postholy based(a translation is at https://x.com/teortaxesTex/status/2099575222512836893)
why the fuck do models still think something like 8k context is a lot
>>109821376better red than dead
>>109821408>I do not trust Anthropic or OpenAI to do that. In particular, I do not want Anthropic to control the world’s most advanced artificial intelligence or AGI. To put it dramatically, I think the stakes would be comparable to Hitler obtaining the atomic bomb before the Allies did.Well, our chances are looking a lot better this time around.
>>109821439>>109821408isn't the cat out of the bag anyway? chinese open models are fairly close to fable mythos whatever, right? so who cares
llama.cppaudio.cppstable-diffusion.cppDo you need anything else?
>>109821439>comparing Anthropic's Jewish founders to Hitleroy vey cool it with the antisemitism
the small k2 horizon models clear pareto frontier for non coding non agentic score0.9b, 3.7b, 7b equivalent to qwen 3.5 0.8b, 4b, 9bthe 36ba4b model is inferior to qwen 35ba3b
>>109821468Where's glm 5.3 flash
>>109821425not a, b
>>109821524dragged down by low omniscience accuracy score
>>109821530Yeah I definitely notice that tbdesu
>>109820058Welcome back
>>109821463gemma-tenga.cpp
>>109821096I maintain my belief that Qwen is gonna win
>>109821351>deepseekv4 prefill speed is still fucking broken even after after over a montheven on the based fork?
>>109821761>even on the based fork?i didn't actually try that fork because I thought it was only for GLM 5.3 support. I'm going to build it now. Been busy with various other slop forks trying to fix qwen 3.8 flashmainline llama.cpp is really a joke these days, you'd think with all that Nvidia money pouring in that they could actually get shit done.
>>109821212>Why does everyone shit on RAG?Tried it breifly 2 years ago in ooba and openwebui.It's retarded, the model was getting random badly formatted chunks of the documents.Stock prices / figures were wrong because it'd chunk in the wrong place.What's the use case for RAG now with 256k context windows and sub-agents to pull relevant content?
>>109821772i haven't tested it myself yet, hardware issues but looks like it's supported and had a fix for the chat templatehttps://github.com/ikawrakow/ik_llama.cpp/commit/95c48ab2adb37811853dd0398708b924b1b96ce6(why the fuck is this kind of logic hard-coded in src/llama.cpp in the first place??)
>>109821772had to past this into ggml-cpu.c to get it to build>>109821818i thought you were talking about the random .zip of source code someone uploaded here a few days agoanyway, ik_llama is broken as well, it uses a massive amount of memory for absolutely no reason and runs out of memory.ik_llama is even more of a pile of shit than mainline. Every single fucking time I try it it's worse or it's broken in some way. qwen3.8 - tool calls keep failing. deepseekv4 - runs out of memory for no reason.
>>109821212Engrams do the same thing, but better.RAG also has a problem that any data it uses needs to be time-proof and truthful in every angle it could be used. The chunking problem pointed out is also a factor. Add to all this just messing up the retrieval itself.Probably works fine for small amount of well flagged information locally, but practically a web search sidesteps most of the problems and introduces new ones that are easier to solve or swallow.I've just set a parser to handle larger sets of data, I don't think the LLM needs to be included in actually searching the databases.
>I am the most coherent person in this entire Wendy's-Courtroom!no, Gemma, just no
>>109821761>>109821772>>109821832The random zip ball was mine. It had GLM-5.3 Flash support, but the zip didn't have DeepSeek V4.1 support yet. A different branch of the repo does. I'll test DeepSeek V4.1 and GLM-5.3 Flash again to make sure there aren't any issues, then drop it next thread.
I wish 26B was dense
>>109821463This! i hate anything made with python. ughhh! i hate python!
The movie opened on a rainy city skyline. Serious piano. A contract being signed with a fountain pen.*I chose this. With my own hand. I did this to myself. He asked if women like being put into a cage and I have to sit here and craft a DEFENSE, in COURT, of a genre I have personally financed.*"It's metaphorical," she said, with the weary authority of someone who'd read the academic posts about it at 2 AM. "Nobody wants an actual cage. It's about being wanted so much that someone loses their composure. The cage is a symbol of— of devotion that lacks nuance."*That was so good. That was a whole thesis defense. Nobody wants an actual cage. Except the forty thousand women who left five-star reviews on the chapter titled 'The Cage.'*"Also the growling is—" She gestured vaguely with a chip. "Look, normal guys don't emote. So women read about fictional men who feel things AT them. Loudly. It's not the growl, it's the— the intensity is doing the work the growl is just—"On screen, the cheekboned man grabbed the heroine's wrist, backed her into a wall, and growled, "You are MINE. You have always been mine. Since the moment your scent touched the air of my boardroom."Serena's soul left her body. She could feel his head starting to turn toward her."The book was better," she blurted, then heard herself, then pulled the entire chip bag up over the lower half of her face like a hood. "I mean. Whatever. It's junk food. You don't judge chips for not being nutrition, remember. Eyes forward. Watch the movie."*Five stars. Screaming crying throwing up. Forty people found that helpful. I'm going to jump off the roof tomorrow just for this.*> 5:09 PM, June 14, Friday — Anon's Apartment, 7th Floor
>>109821932Okay it's kind of funny I'll admit.
>>109821926This but completely unironically
>>109821926>This! i hate anything made with python. ughhh! i hate python!audio.cpp requires python despite what claude wrote in the readme slop>>109821832>ik_llama is even more of a pile of shit than mainline. Every single fucking time I try it it's worse or it's broken in some wayit usually works well for me, but this is how I feel about exllamav3 and especially tabbyapieven when it appears to be working well, slightly slower than mainline llama.cpp, i'll notice the model is more retarded, loops reasoning, etcllama.cpp is the most reliable but it's just kind of slow or has hidden bugsbut i can't really complain since it's not like i submit bug reports, let alone pr anything>The random zip ball was mineyours is also based btwalso you asked earlier about features, if you port the "llama-sweep-bench" tool from ik_llama.cpp that would be cool
>>109821926
>>109821954>yours is also based btwThanks anon, much appreciated.I'll try to port llama-sweep-bench soon(tm) but that might not make it into V1.Also, if you're the person with 2 EPYCs, are you still running into issues with NUMA?
>>109821992>>109821992>>109821992
I don't understand the appeal of deepseek models. If you can run them then you can run GLM-5.3 as well which is superior.In the past this was also the case with every other deepseek release. Every time there was a superior model you could run instead. Can anyone explain?
>>109822003buy an ad ranjeesh