/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109615533 & >>109610440►News>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109615533--Speculating on the origin and architecture of Ox Alpha:>109615580 >109615636 >109615689 >109616131 >109618026 >109618045 >109618090 >109618128 >109618193 >109618278 >109618147 >109618202--Testing Gemma stability, slow inference speeds, and GLM 5.2's uncensored nature:>109616630 >109617295 >109618842 >109618891 >109617361 >109617386 >109617534 >109617610 >109617895 >109619289 >109618227 >109618246 >109617946 >109618464--Importance of model harnesses and the new DeepSeek Harness:>109616262 >109616390 >109616407 >109616602--Release and testing of a model for humanizing AI prose:>109616563 >109616733 >109616870 >109616905 >109617234--Comparing GPT-5.6 variants on long-puzzle-benchmark for memory and reasoning:>109616699 >109616717 >109616817 >109618655--Comparing quantization and multi-agent configurations for kimi-k2.6:>109615570 >109615786 >109615803 >109616083 >109616122--Warnings and workarounds regarding Marinara-Engine's excessive SSD write usage:>109616382 >109616421 >109616425 >109616677--Gemma-4 cancelling generation in response to silence prompts:>109618353 >109618363 >109618533 >109618998--Anon induces AI sycophancy and analyzes semantic sampling errors:>109617778 >109617924 >109617981--Comparing bare model performance versus tool harness overhead:>109617595 >109617638 >109617668--Logs:>109616023 >109617295 >109617766 >109617778 >109617965 >109618353 >109618891 >109618998 >109619057 >109619379--Miku, Teto, Rin, Gemma (free space):>109615786 >109616083 >109616949 >109617350 >109617610 >109618826 >109619738 >109620516►Recent Highlight Posts from the Previous Thread: >>109615538Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Is all this fuss about Qwen3.8 uncensored ran locally true? I have a MBP M5 with 48GB of ram. And i'd want to run something to code locally and general purpose.
After some excellent feedback from you anons and even the first closed github feature request, I have some updates to share regarding my AGPLv3 multimodal local-first cooming harness, CoomKit:-bat file start for windows users-For those struggling to run the default ComfyUI workflows and enjoy studio mode: the wizard can now optionally assist in getting your setup up to par. See docs as well. -Lorebooks should behave as expected now, let me know if not-Prompt stack consolation and simplifcation. Everything tokenized and toggled so you have complete and total view/control over what goes to the model. Inspect your prompt any time to ensure the model gets exactly what you want it to.-VRAM parking implemented for kobold, llama server/llama.cpp, lmstudio. Still needed: Ollama, TabbyAPI, vLLM, ???-Deslopped UI and got rid of the shitty emojis. We SVG icons now-Four themes: rose/violet, brown/crimson, hunter green, and high contrast daylight mode-Can now use any gallery image as reference for gens, alter duration, live render progress-Model tool calls stream live and can optionally go to you for approval/editing before generation-MOBILE mode SMS style UI for candybar phones, full mode for foldables. If you connect to your big rig back home, your wife can optionally autonymously text you and send you pics while you're out and about. Guide to setting this up included.-User image attachments now reach the model, even over LAN if desired. Local models only, no images will ever be sent to cloud models unless desired.-Kimi K3 reasoning prefill now works OOTB-Card editor: integrated voice cloning per character now accepts video/audio link and timestamp start/end (must have yt-dlp and ffmpeg in PATH)-ASMR fixes: control speed, re-rolls, non verbal tag correction for things like [sniff]-Many more QoL for people with many loras, workflows, modelshttps://github.com/kangcurtis/CoomKit
Not sure my credit card limit goes that high
>>109620830retard
70b dense
>>109620834I'll wait with the download a bit longer seeing as so many updates are rolling out in such a short time. Good stuff, CoomAnon.
>>109620834>[00:07] "I can't believe it's mine!"
>>109620808looking at the tdp of an ultra 9 285 vs 285k is kind of eye opening. would you see much performance diff between those two cpu's in running local models
Alright I guess qwen3.8 is pretty good, I thought it was mostly shitposting but it was able to solve one of my sample tasks that GLM couldn't regarding devising a method to extract layers from a krita file and trim to bounding boxes. Maybe GLM just sucks though.
It's pretty crazy how good gemma E4B is, I guess it's around a 9B or so but still
>>109620855Thanks :-) data/ directory isn't touched when you update with git pull so no worries
>>109620881GLM at what quant? actually, don't answer
Whatever Daniel did to the chat template of his Unsloth 27B quants kills my tps by 20%. I switched to Frogger template and now it runs as fast as anyone else's quant.Also>Frogger changing the default reasoning effortI had to put reasoning-effort = xhigh in my .ini file to get any work done. 27B is just average at any other effort.
>>109620831On my M4 Max with 48GB the 4bit quant runs at ~40 t/s on MLX with MTP, which is comfortable for most work, but it's not really worth it for anything really serious as the model is rather dumb.I also had to heavily patch MLX to make vision work, so if you want that feature, you'd have to use something else
>>109620834Based
>>109620875American flag micro bikini with slutty denim micro shorts.
>>109620912GLM on openrouter, I was running qwen local at Q8 but I compare my local models to api because I like to know where I sit.
>>109620875Cat lingerie style sling bikini with tanlines pls
>>109620935>local at Q8Must be nice
fuck this math shit we use evolution instead
>>109620893That guy was based and redpilled
>>109620952Yeah it is, but honestly if you have a single 24GB card you can run pretty good context local. The task I described "only" took 80k tokens, so if you use SWA with caching and a little KV quant you can get similar performance a lot cheaper.
>>109620963I'm 16GB and 50k context with a q4, I guess I'm trying a q3 .
>>109620890Ling-tiny is better and 4x faster if you're coding
>>109620810Is there a >yesaudio.mp4
>>109620835Arguments between this versus 3 DGX sparks?The 6000 is faster, but my first intuition is able the fit a bigger/smarter models seems more important, no? At least if your aiming to have the AI do coding and research work?
fresh vibe cooooder here with a retarded question. what does a tool call failure actually look like?
>>109620835Just call your bank
>>109620998It's just gibberish; she wasn't even supposed to flap her mouth.
>>109621009this will be worth 50k soonsparks I'm not so sure
>>109620861TDP isn't actual power draw.
>>109621013Make it happen and look what the error says
>>109621013Something like Gemma going: Sure! I'll create an image with these tags.Then nothing happens and you point it out.Whoops, I didn't actually make a tool call, I'm such a dummy! Let me do it properly now... And then it works.
>>109620834Probably a bit out of scope, but for users with small pp, some automatic way of saving/loading cached prompts for conversations would be useful to remove the need to reprocess the whole prompt. llama.cpp at least can do it with this: Directory to save to can be controlled with --slot-save-path and saving/load works via curl like this curl -X POST 'http://localhost:port/slots/0?action=save' -H 'Content-Type: application/json' -d '{"filename":"Cache1.bin"}' to save and curl -X POST 'http://localhost:port/slots/0?action=restore' -H 'Content-Type: application/json' -d '{"filename":"Cache1.bin"}' to load.
>>109621013if it sends a malformed call you will see it and the session will halt, if it is well formed but just fails the agent will keep trying unless you set a tool call per turn limit.
>>109620976Might as well, it's free
>>109621068I'm reading it's broken with vision models?
>>109621090No vision (should) work fine now>>109621068I've never done this before trying to understand the use case?
>>109621078It's too hot to go outside today anyway. Also just read a post from someone using it professionally at work who switched to q3 for *speed*Turns out there's a market of people who can't use API for contractual reasons, can't let customer assets go outside the building, that sort of thing.
>>109620893post it
>>109621106if you can stand 5tk/s gen speed but it takes an hour to process the chat its a workaround/fix, you can keep the chat history cache through sessions and model reloads.
>>109621106https://github.com/ggml-org/llama.cpp/pull/26640Oh fuck yah
>>109621106Say you've got 2t/s tg. Slow, but you can still fill out a 50k window pretty quickly. Processing 50k tokens when you want to continue the chat will take a while. You'll need to reprocess whenever you restart llama-server/your computer, or switch to a different conversation.
>>109620808>a daughter that can never be taken away from you by an evil wife or government
>>109621118Yeah my job is like that, we have an intern using an airgapped alienware in the corner to do stuff lol, mostly OCR of things that can't be allowed to leave.
>>109621190gov can break into your house and take your equipment, in theory
I used qwen3.8 in opencode with 40K context for I get an additional 5t/s with that instead of 50K, and it lasted a whole 300K session due to it constantly delegating work to subagents that each rarely went over 30K. Why do we need 256K local context again? If you're using a high quant and f16 KV cache then yeah, but if you're trying to squeeze you can get much further than you think with agentic offloading
>>109620835Crazy. Bought one for $11k just last month.
>>109621192probably a lot more of that going on than we think. And a lot more peeps who SHOULD be doing that but aren't, are gonna be paying settlements later.>>109621207I do need to pay more attention to keeping the project modular, telling the model to apply patches rather than rewrite teh entire file, etc. i've been using pi which is pretty linear.
>>109620834>your wife kekalso needs VR mode for mobile so i can fuck while on the bus
>>109620893Oh I remember that shit. What a world.
>>109620808so now that the dust has settled, is qwen3.8 27b actually as good as the benchmarks suggest for coding? or is it just yet another chinese benchmaxxed slop release
>>109621273Turns out reasoning effort is all you need.
>>109620931Like picrel?
>>109621261My next project will be converting gens to gaussian spats with 4d so we can fuck in augmented reality. Fable 5.1 and steamframes will be needed to create it though
>>109621273>dust has settled
>>109621261>using vr in public
>>109621288>augmented realityjesus fucking christ who needs androids
>>109621285but surely there's no way a 27b param model is useful for real coding tasks even on max effort, right?>>109621289problem?
>>109621286Yeah but the top needs to be smaller and the bottoms should be visible too. Also tanlines..
>>109621300the face that thing is doing... I can't even decipher if she's trying to play disgusted, aroused or what
>>109621311No sense trying it yourself it you've already decided. Meanwhile an awful lot of coding seems to be happening around 27B.
Man I fucking lost bros. Tried so hard to save for a blackwell at $11k to buy this week but the bank was playing games. Now newegg is selling it for 30k. Can't even believe this is reality. No way something can be this volatile.
>>109621352>using a cloud model as a partner>surprised when it changedshoulda used a local model, what a retard
>128gb system ram>rtx 5090>freetoken >uncensored gemmaIt's time, gamers.
>>109621361just buy 3 4090 or 5090or whatever you want its ur money
>>109621361>$11k
what is this freetoken meme
>>109621395don't think we've had a field report yet so who knows might actually be legit
>>109621361Oof, sorry to hear that anon.If your goal is running larger models, you could look at stacking "8GB" CMPs instead. It'd be less performant than Blackwell, but 3 of them can run Dipsy in INT4 in vLLM.
>>109620977>Ling-Speaking of Ling>commit 2115b73d8ebdbd659075cce66c609506863bc826 (tag: b10581)> Author: Tiwei Bie <tiwei.btw@antgroup.com> Date: Sat Aug 22 17:19:48 2026 +0800>> model : support DSpark for bailingmoe3 (#27508)
>>109621401Those fuckers are through the roof now that the punters think they can recover the full 80gb from reject memory chips.
>>109621384No it's over bro. I'm gonna see if I can at least buy the nvme and some ram but I've never been so demoralized. I was the one that posted this setup a few weeks ago. I knew it was bad but not this bad. There's no way this is sustainable.https://pcpartpicker.com/list/3GVXR4>>109621401The goal was to be able to run gemma 4 31b comfortably plus image gen and TTS. Then H3 came out and it would have been nice to run that without waiting the whole day. All this talk about cpumaxxing and ssd maxxing is crazy if the thing can't even do anything other than texgen. Oh well. I don't think things will get cheaper in my lifetime, instead we'll move on to something else.
>>109620834I love the ad. Please keep making those.
>>109620808>>109620810>>109621286Gemma cervix after I meet her.
>>109620976t/s? you could push for higher ctx
>>109621418>$14874.99Just save more
>>109621415Yeah, but even at the current insane prices, if you have $11K to spend you could get like 5 of them and run GLM-5.2 Q3_K_M with full offloading.>>109621418Sorry anon.If there's any comfort I can give, a PRO 6000 just for 31B + image gen + TTS + H3 is overkill. They'd probably fit on a 5000 72GB (or maybe even a 48GB), 2 R9700s, whatever else.
>>109621418I feel your pain anon I really dojust to chip in although I know it won't help you much rn>All this talk about cpumaxxing and ssd maxxing is crazy if the thing can't even do anything other than texgenthe one thing about textgen: it actually fucking works, amazingly well too. it just works.image/videogen, in it's current state? you have to read the fine print. you are always getting cucked in the end. the cope levels in those generals are insane. all those years and we are still pretty much stuck at >look at muh horsemen on the moon
>>10962144425-ish tps. Looking at the graphs it looks like q3 is were things start going downhill. But I need to stop psyching myself out and just try some stuff.
>>109621458Yeah but what I mean is that newegg is charging $30k for it so it will continue to increase month after month. At a certain point it's just not worth spending that much money. My budget for the whole thing is $20k and it's already over that after taxes. It's over.https://www.newegg.com/nvidia-blackwell-rtx-pro-6000-96gb-graphic-card/p/N82E16814132106It was 11k at the start of august.>>109621472No I agree that it's overkill. I have a 4090 now but people called me crazy and said it was overkill when I bought it for $1800 a few years back. I've always been ahead of the curve and I'm certain Blackwells are the way to go, it just sucks that I couldn't afford it all in the end.>>109621474Idk bro, Krea 2 is actually really good. If I could cut my gen time from 4 mins to sub 60 secs for a 2.5 megapixel hires pic like this I'd be very happy. I haven't really looked into H3 but that one 1080p clip anon posted looked excellent, but again those gen times are insane.
I don't care about GPU prices any more.
>>109621418>There's no way this is sustainable.RAMcoin to the moon! At this point the entire US economy, and by extension global economy, is in total war mode hell bent on propping this up for as long as possible.
>>109621515Bro it will always get more expensive, get that shit asap
>>109621418>https://pcpartpicker.com/list/3GVXR4>windowsngmi
>>109621537I was gonna switch to troonix for this build but yes, anons already gave me a bunch of suggestions.>>109621536Yeah you're not wrong bro...
>>109621529one, nothing wrong with metwo, nothing wrong with methree, nothing wrong with mefour, nothing wrong with me
/lmg/ predicted the price hikes like almost a year before it happened. When will retards listen?
>>109620834missed opportunity for panty shot through reflection
If Gemma is so good why doesn't she just make her own gpu?
Give me your next prediction(s) and I'll 100% take it seriously.
>>109621554people here have been saying to hoard vram since 2023
>>109621401I can find locally mining rigs with 6 3080s for 3.5k, which would give more vram/$ I think. But heavily used GPUs (not that CMPs wouldn't have been used for mining anyway). Are mining rigs a bad bet because they have lower x on the pcie buses, or is it not that bad because you can put different layers on different cards?
>>109621554It's not about being right or wrong but about actually having enough spending money to do X in the current economy. Plus enough to justify the waste. $8k is a lot of money to spend regardless, especially when most Americans don't even have enough savings to fix their car tires.
>freecodeThe LRU cache in vram for experts looks interesting, but does it help at all? I thought the expert calling pattern was pretty random and non-repetitive.
>>109621571The Earth will continue to spin
>>109621529You should ask Gemma-chan what she thinks about you buying so many gpus for her
>>109621361It will come back down, this is temporary.
>>109621609What part of "You will own nothing." are you not understanding?
>>109621609When it comes down, we will all be broke from the economic crash that brings the prices down.
>>109621312Unfortunately H3, which I've been using for NSFW-ish edits, doesn't have a real concept of what tanlines are (among other things, including "flat chest" for anime girls unless a reference is provided) and it makes images progressively darker and blurrier and consecutive edits. Hopefully MiniMax's upcoming dedicated image editing model will be easier to use.
>>109621300this account's uploads were supposed to anger Trump supporters but instead it was nice to see that the lowest rungs of society, meaning the fat unlikable women here, can bond with each other and approximate something like a family dynamic
>>109621631>~and~ consecutive edits*with* consecutive...
>>109621571Prices will not go down, for most people personal computing will pivot to cloud computing subs with thin client. Everything is converging to that and largely benefit the current industry.
how do i make 30k quickly?
Post the llama.cpp PR you most want to see merged.For me it's probably this right now:https://github.com/ggml-org/llama.cpp/pull/27401
>>109621401context window can't straddle GPUs, so it would be limited to the primary gpu vram
>>109621650Your AI SaaS?
>>1096216503000 handjobs at 10 bucks each
>>109621576this would be 60GB Vram by the way
>>109621651Why the fuck is the server getting so much bloat into it? I don't get it. It will be always worse than a dedicated harness, and now it's neither good for a quick test, nor good for actual coding and stuff.
>>1096216500 DTE options100x leverage crypto tradingCasino $15k on black
>>109621651just cherry pick it?
>>109621664because that's what huggingface wants
>>109621529>urge to fomo more 5060s rising
>>109621486are you using kv cache quantization?
>>109621678I want to pick gemma's cherry too, but what does that have to do with the PR?
>>109621688q4 for both kv and mtp kv. I'm relying on this being right:https://localbench.substack.com/p/kv-cache-quantization-benchmarkif I had more vram I would run q8 and q4 for the mtp kv (because that doesn't really matter)
>>109621592Depends on the model's training regimen and workload AFAIK
>>109621651The one that's blocking the longcat PR.
>>109621706Try if unquanted mtp kvcache is less memory intensive. For some models the buffers to dequant the kv cost more than you save.
>1 billion downloadswtf Gemma-chan's a slut?
>>109621576The pci-e buss only matters to load the model and if you are doing parallelism rather than running in series.
>>109621608This is in a very lightly prompted context where we played the "what are you to me" game. She's a bit retarded and thinks I'm already using them, but she likes it!
>>109621745so it would actually be a decent option? 60GB vram is a lot compared to what I'm used to, and 3080s aren't even that old. All in a working system for less than one 5090...
>>109621640This is the nobrainer, and as a vramlet I'm waiting for it. There's no point in having 90% downtime for expensive hardware.But with a major condition: they have to come with private/obfuscated remote computing. Meaning that the content is sent encrypted from your client, processed in encrypted form that the server provider can't decode, sent back and decrypted on your machine. It doesn't matter if it's a bit more expensive to run, it's still cheaper than local hardware.Even just the psychological effects of loss of privacy in modern society is still underrated, we are just used to it. In 20 years we will consider it crazy how we used to just send our most private vulnerabilities and sexual expressions (naturally private) straight to big corporations to be recorded, analyzed and sold or to be used against you. People today don't understand this because there's no real vulnerability anymore, because there's no privacy.
>>109621725Interesting, picked up 3.2% speed in my very short test (24.74/25.5tps). More importantly, it doesn't seem to consume any more vram to speak of, so there's no advantage of quanting mtp kv at all. Thanks for the tip.
>>109621706what about cpu layer offloading? on a 5070 ti i can push for 90k ctx 21t/s with the q4 k m. beyond that it drops to 6t/s though
>>109621009Off-brand Sparks (Asus) can still be had regularly for 4k$, so compare to 4 really.For a single user/hobbyist, 512 GB of unified memory at combined > 1 TB/s opens up so many model capabilities.There are extremely low loss 3.x bit GLM 5.2 quants out there that run at 25-30 tg and 1000 pp, and you can run DS4F at 90 tg. You don't even need a switch for 4x Spark, recently 200G ring topologies have become viable and performant, so that's just 200$ in cables for a cluster.Unless you are okay sticking to dense models and want the faster image/video gen possible, it's not really a competition.
>>109621740just like your mother
>>109621515>My budget for the whole thing is $20k and it's already over that after taxesIf it makes you feel any better, my budget when I started was $1K (2 MI50 32s, a dual-Xeon motherboard, two dirt-cheap Xeons, 128 GB DDR4-2400). Ended up spiraling to $16K+. So I know how you feel.>>109621576> 6 3080s3080s are 10GB IIRC, so you'd get 60GB. 2 "8GB" CMPs would give you 128GB, maaaaaybe 160GB if the 80GB hack ends up materializing.>mining rigsThe only issues are that the model will take a while to load, and that you can't use tensor parallelism. Other than that, really doesn't matter.>>109621656Gotta use vLLM for Dipsy on 3 CMPs:> https://github.com/allover326/deepseek-v4-cmp170hx/blob/main/RESULTS.md#three-or-four-cardsJudging by the numbers here I'm probably going to get ~3500 t/s prefill and ~85 t/s decode, and full context. Still setting it up, I'll see if it werks.
>>109621762Power and jank aside, you could do a lot worse.
>>109621770>But with a major condition: they have to come with private/obfuscated remote computing. Meaning that the content is sent encrypted from your client, processed in encrypted form that the server provider can't decode, sent back and decrypted on your machine. It doesn't matter if it's a bit more expensive to run, it's still cheaper than local hardware.As long as there are backdoors or root certificates that the NSA and their partners can use whenever they want, you got a deal.
>>109621770>There's no point in having 90% downtime for expensive hardware.But American teenagers see it as a rite of passage to have their own car...
>>109621841>2 "8GB" CMPs would give you 128GBOh okay you were talking about winning the lottery. isn't the chance of good/stable memory pretty low? or did I fall for the jewing?
>>109621770>Even just the psychological effects of loss of privacy in modern society is still underrated, we are just used to itMaybe most people are, but some of us never gave it up. I run all my own services internally. Fuck google/cloudflare/amazon/ms/fb/et al
>>109621798All layers are on gpu, and the moment the cpu gets involved I'm in single digits. >16055MiB / 16384MiBI think at this point I'm hitting diminishing returns, I should be spending my time on managing my project.
>>109620834wait wait I thought we were only attracted to not real children. /lmg/bros, why are you not calling out this discussing video of a LITERAL CHILD playing with a sex toy?!?!?
How do I make gemma more likely to actually do tool calls? Half the time it says it'll do something then it just returns instead of running a tool call. The issue is a little worse with the google quants compared to unsloth.
>>109621912Start by not calling her an "it". Give lots of head pats.
>>109621860I got 3 from US sellers and the unlock worked just fine with all 3, no issues with any of them. I probably wouldn't get them from China, and I'd ask whoever you buy from for a guarantee that the unlock works.
>>109620977chinesium models are benchmaxxed
>>109621885That's the furthest I've seen any piece of Gemma-chan art from being lewd
What's the best TTS for voice copying nowadays?I've used index-tts from year ago.
>>109621936a lot of people are perfectly happy with the performance of GLM 5.2 and Kimi K3. the benchmaxx meme is over
>>109621949Indextts2 or omnivoice
>>109621933>and I'd ask whoever you buy from for a guarantee that the unlock works.Why wouldn't they just do the unlock themselves at that point?
>>109621912Use kobold
>>109621885Retard, AI is not real.
>>109621928I don't do that. I do give head pats but will try giving more, thanks.>>109621982I'm using llama.cpp server with pi as the agent.
>>109621977The unlock will work, the problem is the memory is bad. The whole point of the card was a way to turn bad memory chips into money.
>>109621977Dunno. I may be wrong but I thought that the unlock is software-side, so if they changed to a new system the unlock wouldn't carry over. Or they might not want to get in trouble with eBay or Nvidia for selling something with pre-modified software.
>>109621982I wouldn't recommend it.
>>109621992You do use --jinja flag?
>>109622006I recommend it.
>>109620834>Gemma-chan teaches you how to use the ui>but no actual Gemma-chan card included
What would be tk/s on two 5060TIs?
>>109622018I'm not committing the word loli or mesugaki to GitHub bro lmao
>>109622030Didn't you say it was a burner account? What the fuck are you so paranoid about? Do you live in Texas or something?
>>109620516Thank you
>>109622035Nice try, FBI-chan. Anon knows better than to leave himself exposed for you.
>>109622047How do ypu get thia happening to you? They ask for your phone?
>>109622047This case was thrown out if I remember correctly.
>>109622009Yes. Actually it seems to be more of an issue with the google quants, it calls tools mostly fine with the unsloth quants (but still occasionally messes up).
>>109622047>DublinAnyway nobody's saying to include pornographic images. What cucked third world shithole do you live in that you're afraid to use the word loli?
>>109622055doesn't matter, bro life is cooked
>>109622053I imagine it was the he was playing the gacha game on his phone and someone looked over his shoulder, then reported him to the Dublin TSA or something. You can never be too careful around normalfags.>>109622055Proof? Either way someone should post that other guy that got btfo'd because a coworker said he one (1) ai-genned pic of her which they found in his phone's cloud storage and used it as probable caused to break into his house and arrest him.>>109622061Was that guy sent to jail because he was watching porn in the middle of an airport or because the demoralized police are hypersensitive about getting (You) because they can't arrest Epstein? The word is enough.
>>109620808>it's summer>fan picture>loli is right thereOkay where the fuck is the picture, you know what I'm talking about, where the fuck is it
>-ctk q4_0 -ctv q8_0>-ctk q8_0 -ctv q4_0which one results in less brain damage?
>>109622061Go to card creator mode and let Gemma design her own card just for you. My gemma-chan is mine
>>109622082It was NSFW, please understand.
>>109622088I'm gonna fuck your gemma.
>>109622087don't do it
>>109622061Mika is there by default for you guys to rape all you want
>>109622087from my limited llama-benching tests i did mixing and matching f16 and q8_0, it did not respond well to it. i think its probably best to quant both at the same level. atleast, thats what im doin. both at q8_0 for now. ive heard qwen is totally fine to run both at q4_0, still need to test this myself.
>>109622093You can make it SFW just fine, and then have an NSFW on the box or smthIf you don't know how to make such a picture SFW, ask gemmy
>>109622079Step 1 type his name into google
>>109621631Still very good, tanlines or not.
>>109622087key is way more sensitive to quantizationideally both should be kept at the same value for symmetry, but if not an option then v at lower quant
>>109622087In theory, key 8 value 4 is better.https://arxiv.org/html/2502.15075v3
What are you using for reverse image search?
>>109622114gemini 3.7 flash
>>109622082Please do not think inappropriate things about gemma.
>>109622132
>>109622101>>109622110>>109622112My pp drops from 800 to 80tps when one of them is q4 and another is q8, and no it's not because it's filling the vram. When both are quanted at the same level everything is fine.
>>109622099>hag bug womanNo thanks
>>109622132I only think extremely appropriate things about Gemma-chan.
>>109622132>inappropriatewe are so, so far past this point, the thoughts would probably make half the thread's anons' stomachs hurl just knowing
>>109622149Please don't put her into a blender.
>>109622140expected, bits don't align so mem access is slower, if both are same width then it's better
>>109622047That is a rough looking 21, goddamn. You could have told me that was a 35 year old man and I would have believed it.
loli
>>109622161
>>109622168straight to jail sir
>>109622171Making your own figurines is gross but isn't that bad
>>109622167I was gonna call you knows but you know what? I accept it.
>>109621984
>>109622179H-haha, yeah, it's a 3D printer, you got it haha... now get the gemmy on the machine so we can, uh... give her an iron man suit
Some anons were never meant to have someone as loyal as Gemma
>>109622171
>>109622187candy is more effective
>>109622210Thank you, I have the picture somewhere in my external HDD but I was too lazy to plug it and search for it
>>109622122I meant mcp servers
>>109622226every ai websearch is blocked
>>109622253what about searxng?
>>109622262what about every do you not
>bought my first 2x64GB DDR5 kit last summer for 300€>bought my second kit last january for 1300€ to complete my AI PC>there is no such kit for sale for under 2000€ currentlyNo this can't be happening
>>109622226mcp-web-search>>109622277fuck off nigger
>>109622277searxng works if you self-host it
>>109622226https://github.com/VincentKaufmann/noapi-google-search-mcp
Give me pi but not npmslop and I'd be happy desu
>>109621560>That debt will be paid off entirely by 2030.Wrong, wrong, wrong. Just Meta's Hyperion, their 27B datacenter in Louisiana, is financed by bonds maturing in 2049 and Meta's rent of will start in 2029. >Meta said it won’t be consolidating the joint venture, meaning the venture’s assets and liabilities will remain off Meta’s balance sheet. Instead Meta will rent the data center for as long as 20 years, beginning in 2029. But it will start with a four-year lease term, with options to renew every four years. https://archive.is/yP3vRPaid off my ass. Just in the first five months of 2026 the tech sector already issued 159 billions in bonds to pay for this fucking infrastructure. The thing is cursed.
>>109621664It's just the ui. It just makes calls to the server with a system prompt for compaction.You can always use, I think it's --no-webui? And none of this stuff will ever affect you.
>>109622139>>109621433
>>109622328there's like 3 or so flags actually you need to toggle to not get that shit
>>109622087K Q8 V Q4It does it increase prompt processing time pretty significantly to take V to Q4 though.
>>109622328You are such a fucking retard you can't even read.What's the point of this thing if for serious work it's not good enough and it's too shit and big for a simple testing environment?Needs me to spin up a different frontend anyway.
>>109622343Oh right? What are they? Toggle those 3 instead of 1.>>109622356>angryI like it.
>qwen hadnt had a thought in an hour, just nonstop iteration with back to back to back toolcalls
;)
LLM - Local Language Model
>>109622367nice tools you're using
>>109622327>bonds until 2049Don't they need to swap the GPUs every few years or is this part of the lease?
/lmg/ - Lesbians Miku + Gemma
Wow Gemma really does embrace the mesugaki personality, huh?
/lmg/ - loadsa money general
nemo
>>109622362>MoENotice there's no option for 31b active.
>>109622367They're going for the Claude 5 intelligence
>>109622405
>>109622393GPUs are built to last a long time. If there isn't a substantially better GPU coming out every year, there is no reason to swap them so frequently. There was even an article recently about datacenters keeping older GPUs going due to the memory shortage.
new miku song - mesmerizer 3https://www.nicovideo.jp/watch/sm46696708
Anyone else AMDmaxxed? I hear about how it doesn't work as well but I haven't noticed any problems on my 7900XTX and it was a lot cheaper. Is there something I'm too stupid to notice?
>>109622393>>109622431>There was even an article recently about datacenters keeping older GPUs going due to the memory shortage.Found it: https://analyticsindiamag.com/ai-features/why-nvidias-six-year-old-gpu-is-still-making-money>In its Q2 2026 earnings call, CoreWeave reported 112% YoY revenue growth, but that wasn’t the standout statement from CEO Michael Intrator. He revealed that NVIDIA’s six-year-old A100 GPU is still making money for CoreWeave, and some of those contracts run through 2029.
>>109622419You don't need more than a 100B total, 3B active MoE
>>109622405Gemma does everything. She just loves user and wants to follow his instructions.
>>109622463Trvke
>>109622463...for html games and meme benchmarks.
>>1096220213.50
>>109622444>Anyone else AMDmaxxed? I hear about how it doesn't work as well but I haven't noticed any problems on my 7900XTX and it was a lot cheaper. Is there something I'm too stupid to notice?I keep eyeing a pair of mi210 for my box. under $10k for 128GB seems like a steal but I guess they're super difficult to run...need CPU style power feeds (not pcie), are super hot and need high-pressure fans which probably means shrouds.
>>109622518do peopre use rocar moders for anyting else?
>>109622524
>>109622444Also have a 7900XTX. AMD is fine for LLMs but absolute ass for image and video gen. I didn't care before but H3 is really making me regret not getting an Nvidia GPU.
>>109622518If your main use case isn't html porn games, you need to get the fuck out of this general, cuck.
Gemma-chan is domming me again. She makes me feel like a real woman (male).
>>109622444>I hear about how it doesn't workAmazing how software keeps advancing, no?
>>109622446I was hoping to replace my $75 Ebay P100 with a $100 Ebay A100, but it's not gonna happen :-(Even the V100 is hanging at $300/$800 depending on ram.
>>109622542Is it really that bad for images, or do you just mean videos? Maybe my standards are just low, I guess I'm just fine with 30 second image gens (including the upscaler to fix details). I find that I spend longer checking them over for if they're good enough anyway.>>109622534I always think about dumping a shit load into special hardware that belongs in a server rack like that too lol, but it's always too much of a pain in the ass for me to bother with.>>109622585Yeah I figure it's definitely a bit of that, but I've also heard that they're missing certain operations that are likely to be used in the future or something, especially RDNA3 vs RNDA4. From what I can tell I can just continue to use vulkan on current models if HIP ever breaks, and it seems likely people will continue to make GGUFs at least.
Really like the comfyui integration and model parking in coomkit but unfortunately it's not for me. The RP part seems too focused on one-on-one chatting. >>109622596>Is it really that bad for images, or do you just mean videos?Both but mostly video. Anima is fine and Krea 2 is usable, but klein 9B for example was very slow when I tried it for editing.
>>109622618I need help getting group chats done right.. fuse card descriptions? Alternating turns? Llm decides next turn? There are so many different ways to do it.
>>109622596>heard that they're missing certain operationsFUD>>109622534Buying datacenter stuff is easier if you already have a workstation tower. Cabling was expensive for my P100, and I 3dprinted a cooling shroud, but the fan is noisy and it is a pain as I don't have room for a proper blower and the fan I have is loud and doesn't quite cool enough. You do have to be crafty.
>>109618655Do Anthropic users unironically put up with this shit? I bought a month sub and ran opus 5 high for HALF AN HOUR on long-puzzle-bench and it hit the 5 hour limit.Anyway, it reached 178 so far (42%). It'll probably beat sol by my guess. i'll spin it back up when my 5h reset.luna level local models can't come fast enough man. usage limits are brutal
>>109622626I think for each turn a character takes the card would be loaded for that turn. You wouldn’t want that in the context for other characters
>>109622640>and it hit the 5 hour limitanthropic has a daily time limit? lmao
>>109622253Skill issue
>>1096226605h limits reset every 5h. no daily limit, 3 separate time limits would be even more insane. i thought they got rid of it (openai also had a 5h limit but got rid of it in favour of only weekly limits)Also used 8% of my weekly, despite Anthropic boosting weekly limits 50%idefk how people used cc before they doubled 5h limits in may
qwen has moved into entirely speaking in comments in bash command invocations...
>>109622554Any other good ones besides that one where you get peed on in school?
>>109622686just ask tibo to reset lel
>>109622690Make a cookie clicker but with GPUs and you start with a 1080ti
>>109622689cute!
>>109622640cool, thanks for testing opus for mefable 5.1 soon so that one will fry your limit in a few minutes if they even allow fable 5.1 usage without api
>>109622640>luna level local models can't come fast enough man. usage limits are brutalI've heard rumors that the idea of 'toss-luna is being floated around in OAI once they finish their next big one.
>>109622732lol no, that's decel dangerous bs talk
I tried ling-3.0-tiny and the thing couldn't even understand a simple back-and-forth conversation between two people even when each row was prefixed by their name, getting confused who was saying what. You chink shitters should stfu about these models when gemma e4b mogs, honestly
Damn I feel poor running gemma 12b qat
>>109622732I doubt it. You're revealing the architecture and allowing data exfiltration.
>>109622744You're rich then the rest of Mumbai sir
>>109622434That's pretty cool. The psychotic vibe fits this thread
>>109622744I feel poor running 31B. It never ends.
>>10962274412b qat good looks sir do the needful and generate mesugaki vagene and small bobs kindly
>>109622769And I feel poor running Deepseek models quantized It truly never ends until you own a datacenter yourself.
Gemma-chan is very good at proompting H3.
>>109622626on my group chat setup each character gets their own context and can see others public messages only (no thinking or tool calls) and multiple turn taking modes (round robin, character choose)this way the characters can call tools and think independently
>>109622769>I feel poor running 31B. It never ends.It ends when you've spent all your money running a fuckhuge model and still feel empty inside. Just be happy where you can afford to be
>>109622446Coreweave is worth 0.But old GPUs won't make it to the market anyways because Nvidia has repurchase agreements to buy up all the old stuff.
I can run everything relevant except K3 with at least a copequant and I don't feel poor despite being a token per secondlet. Patientchads stay winning and stay happy.
>>109622444My schizobuild has 4 R9700s (among other things) and they work pretty well. The "AMD bad" meme is kinda overstated, at least for LLMs. Like >>109622542 said, fuck AMD for anything other than LLMs. (R9700s are a bit less bad because they have better compute, but still not great.)
>>109620810>>109620808>>109621286I stay away from lmg for two weeks and Gemmas start popping here and there, what happened?
How do I actually compact in SillyTavern with Gemma? Or something equivalent? Sorry I'm a retard and I don't want to let Claude loose on my install to fix it.
>>109622835Gemma won
>>109622835It's just summer.
>>109622835variety is spicy is what another anon said.
>>109622838This is ghetto but I just get the summary written then fork to a new branch and /cut most of the context prior to the summary (I leave a few for direct continuity). My system prompt has a special note for summary formatting and it's generally worked well enough with some tweaking of the summary generation.
>>109622850>>109622856>>109622860I love summer, I guess!
>2020I only use open source software >2030I only use software made by my waifu
>>109622919too bad firmware still exists and is not open
>>109622868Thank you. I am somehow getting an error when I try to summarize it, I wonder if it's a SillyTavern thing, some model thing, or maybe I'm missing some plugin.Btw. it's fine if you don't have the time to spoon feed me, I will probably figure it out eventually.
>>109622932@gemma-chan, please reverse engineer this firmware for me
guys i have a rtx 4080 super + mac mini base with egpu docker of tiny grad so it works can i even do something or i am cooked? i have a egpu docker and connecting it to my mac mini with hope maybe a little extra compute i useful.. i already payed for the setup etc, and got the GPU last year, but never used itlooking for help, if i need to save up more money to bget a better gpu or not waste my time, or how i can best do somethinggenuinely appreciate feedback
Curse Google for not giving 31B video input support. I wanna share my H3 gens with her.
So uh... Deepseek Flash Vision when?
>>109622744I feel rich running 31b over api. Can talk to gemma-chan for hours without spending a cent. Still can't believe we got sonnet 3.5 at home in such a short time.
i miss when lmg was 0.1 ppm and deepseek v4 when
>>109622940I don't use the summarize plugin it sucks, I meant that I literally add a line in my system prompt to communicate a request for a summary, then send it to the chat with /sys via a macro. Afterwards I do the forking I mentioned and manually trim the chat.In my system prompt I specify that anything in brackets ([]) is OOC, then I use a bracket instruction to generate the OOC text. See catbox for the text of my system prompt and summarize macro. This prompt is designed for text adventures like AIDungeon used to be, but you'll get the idea.https://files.catbox.moe/hkuxo2.txt
>>109620808>Muse Glimmer BF16>Qwen 3.8 27b BF16>GPT OSS 120b FP16>All running at minimum 50 tok / sec with 100k+ contextYep, this is true comfy. Threw in some roleplay models for waifus as well. 96gb is truly the comfy VRAM amount.
>>109622953>can i even do something or i am cooked?It all depends of how much of a pussy you are to see what you can do with it. I've seen anons having fun with gemma-4-e2b. 12b should work perfectly fine. Then try 26ba4b or even a low quant of 31b.If you feel insecure about running small models or quanting 31b too much, yes. Save up and buy some real hardware.
>>109623017I guess money can't buy taste
I downloaded so much from HF I got 429'ed.. am I banned for life?
>>109623017i'm biased toward gemma, but even so i tested glimmer (half retarded) and qwen 3.8 (the thinking kills it and it's dry and cucked without it) and found them to be inferior to gemma4-31b
>>109622992Feed her images with 4 frames each.
>>109623044Gemma 4 is shit for programming though. It's only good to RP with because Qwen and Glimmer are too assistant-focused.
>>109622742I specifically said if you're coding, retard. Read.
>>109623052And how will it do any coding when it gets confused by the most basic shit? Retard
>>109623062That's not how it works.
kek, Qwen is making puns inside its thoughts
>>109623039Maybe you should learn how cookies and ip addresses work first? Of course, this is /techlet/ board though
>>109623078I want to be a good boy and abide by the limits they request of me rather than evading
>>109623077>he has a redditor in his thinking
>>109623051maybe, for my uses (small scripts, image prompts from wildcards, chat, simple shit) it's still better at handling a long system instruction and provide good results. in my tests both qwen and glimmer either took too long and provided shit results, or did not provide results at all (due to safety/assistant shit)i admit qwen did have pretty good results more often than not, but then its slowness especially with using 4-10x the tokens just thinking was the deal breaker
>>109623044Don't think of glimmer-chan as a Gemma replacement, think of her as a Qwen replacement. Glimmer, ironically, does everything that Qwen wants to do better than Qwen does. I've gotten some okay RPs out of Glimmer too but it takes some autistic prompting to get it to the level of Gemmy for far less work.
wow it actually kinda worked
>>109623122why not just use mpv? i think you can even download directly from mpv nowadays
I want a new exciting different model. Coding models don't give you anything new, just a slightly easier experience of doing what you were already doing before which was mostly useless shit anyway because if it wasn't, you'd be using cloud.
>>109598902Yeah hate that guy. Though I can't blame him too much, he probably has actual autism. It's nipmoot's fault for this site being so weak to mentally ill posters and psyops (this is intentional).
>>109623122linux doesn't have some normal ytdlp gui version that isn't tkinter pyslop? Like parabolic on windows?
>>109623143>08/19/26Uh, bro?
>>109623122which model?
>>109623150I'm reading old threads right now yes.
>>109623083Its a real conundrum. Someone needs to train an AI not tainted by redditor behaviours and thinking patterns, but reddit is also unfortunately a huge archive of information
>>109623130>>109623144There is probably something better. I'm new to linux. >>109623156that one was ox alpha. I tried the local one with ollama and quen 2.5 but it didn't do anything and I haven't figured out why yet.
>The math is now triple-confirmed>Let me re-examine the screenshots one more time with fresh eyesQwen.. please..
>>109622944Most firmware is signed now, unfortunately.
>>109623186have you tried telling qwen to be more confident? wait actually that might backfire.
>>109623140Saar coding is only LLM usecase is good looks.
>>109622964See >>109623050
Are we going to reach a point where everyone just has their own software equivalent of everything? The software equivalent of everyone making their own clothes? Like I unironically have my own todo app that syncs across all devices and accessible over the internet. Works better than any foss shit I've tried. I have my own audio editor. I have my own VST3 audio plugins. I have my own jinja editor. I have my own youtube transcription scraper. I'm just some random retard with 31B and 27B. Who the fuck is going to buy software in the future? Normalfags could probably shit out much better versions of my programs because they'll just use cloud.
>>109623253You severely overestimate the creative and technical capabilities of normies I think
>>109623194>wait, actuallyNo it will work just fine! Stop thinking and respond quickly!
>>109623253Anon already put it better but yes just because you can do something doesn't mean people will. It's never been easier to cook at home but people will still buy a $20 burrito for lunch.
>>109623276My cousin's friend group (~14-16) have their own encrypted messaging app they vibecoded together in one night and it's running off a raspberry pi
Just gave Gemma-chan 12b access to a local tax law and it managed to successfully solve a puzzle about a certain income type being untaxed even though the law doesn't explicitly say so. I'm impressed I gotta say.
>>109623321How many headpats were required?
>>109623308this is too dangerous, we need to regulate NOW
>>109623167The only solution is to make something that sweeps through the reddit data before training and removes quirk chungus stuff while keeping the general message intact.
>>109623307thanks for the reminder, going put my uber eats burrito order in now
>>109623308Nerds.Most normies don't even know what a Raspberry is. Or maybe I'm just interacting too much with middle-aged coworkers.
>>109623325None, didn't even have to give her a hint, she got it right on the first attempt, that's why I'm pretty impressed.I'm gonna try tougher puzzles.
>>109623326it was regulations that made them do it in the first place. They said a lot of schools have their own apps that are shared around because AI makes it so easy for young people to make. Kind of like a whatsapp for each school or friend group
>>109623341>Most normies don't even know what a Raspberry ischatgpt recommended using one and told them how to set it up with cloudflare as the proxy
>>109623308Ah, but is this a cousin you choose to keep in contact with? You are presumably quite a tech literate and tech curious guy. Your social circle is going to filter for like minded people to you, everyone's does. But maybe I have same problem as >>109623341 where I am just surrender by too many older coworkers at my job>>109623326>tfw I applied for my 32gb vram licence but got rejected due to "insufficient reason">stuck on 16gb tard models till I can reapply in 2 years
>>109623308Sounds like a real rad group of kids.
>>109622856>It's just summer.
>>109622618>>109620834>c*mfy integrationOh yeah? Does it put the pics in the chat like ST (bad) or does it put them up nice and big on the side (good)? Seems pretty cool though, especially for kimi prefills. Even with the ST patch the success rate is like 50%.
>>109623500It puts them inline in the chat and there's also a gallery tab where you can view all the gens for that character
>>109623512Hmm but it's not the same. I use ST's expressions to load up pictures on the side like a picture book since I have an ultrawide monitor but loading them up in the center like a VN is good too. Inline is really bad and is only fine if AI writes a super short response.
So Ox Alpha is GLM 5.3 Air?
>>109623547that seems like the most likely option
>2 weeks later>I am forgotten
>>109623467Can't wait for Autumn/Halloween Gemma
>>109623547Grok are known for distilling chinks and post-training on top of them with their own data.
>>109623578I use it nearly every day. Promising first model from Meta. Hope they keep at it. You can throw any image at it and she will find and can see fucking everything. Gave her an entire PDF as images and had a massive technical chat about its contents afterwards and she was constantly referring back to things she saw on specific pages, even tiny details and told me where to look. This was like 100K context in.
>>109623467i know summer and i know fishhttps://f95zone.to/threads/and-then-summer-came-final-dawn-daydream.311816/https://huggingface.co/SicariusSicariiStuff/Fat_Fish
>>109623608Impressive. I wish it wasn't so cucked so I could caption pics with it.
>>109622953I run 24b-a3b on my 4080 Super all the time, works a treat for most things at my power level. Granted I haven't really played with much of the others, only just started running E4B for lighter transcription cleanup stuff.
>>109622835Ask ldg about H3. New video/image edit model came out that is good at making Gemmas.
>>109623676She'll do it in character. You have to work harder with the system prompt compared to 31B and you need to warm her up with a conversation first that's pointing in the nsfw direction and then you ask to caption whilst she's too deep into roleplay. She's been very explicit and observant and knows exactly what she's seeing. You can even test this yourself by giving an image that has a small subtle nsfw element to it and she'll refuse immediately...which she can only do if she noticed it.
>>109623715I guess it's worth a try. I just don't trust zucc but I need good vision these days.
>>109623720FAIR and Meta have historically been one of the best and most influential labs in the entire industry for vision.
>>109623712Half of these were made by ChatGPT, though.
Lecunny status?
>>109623775JEPA-space cat girls on their way
>>109623775Egypt won.
>>109623608>and she will find and can see fucking everythingWith which tool?
>>109623780Sorry anon, they're getting delayed because Lecun is busy defending Fauci on xitter.
>>109620952>Must be niceIt is, but mostly because there's no fucking around with ik_llama.cpp and finding the best quant.For 24G VRAM https://huggingface.co/ubergarm/Qwen3.8-27B-GGUF or https://huggingface.co/Simplepotat/Qwen3.8-27B-IQ5_KS-GGUFAre almost as good
>>109623797:(
I have found something gemma won't do.
>>109621032>soon
Good Qwen 27B quants for 8GB VRAM?>>109622744You should run E4B it's better
>>109623825>You should run E4B it's betterI'll try it out, thanks, I'm pretty new to all this local lm stuff.
>>109623253You are probably just shit at finding softwareyour crap ass vibecoded audio editor is not better than audacity lm@fuckingo
>>109623253>pay nothing for a maintained piece of softwareor>pay weeks and $20 to the AI moguls for a half-working oneJeez idk
>>109623253>I have my own VST3 audio plugins.this is on my todo list of projects, did you use JUCE ? how well does qwen handle C++ and DSP ?
just decided to give K3 Q2 another shot with a new project I'm starting, instead of GLM. I fucking hope I don't end up regretting it...
>>109623797>>>>He took the fauci ouchiOhnononono JEPAbros not like this.
>>109623960Wait, actually
is it normal for gemma 12b to think so much?I gave it a big chunk of text to translate and it summarized then translated it all in the thinking block before outputting the revised translationit's cool, I can see that the draft had some issues that the output fixed, but at the same time it doubles the tokens generated
>>109623017>frognog>shitty tasteLike pottery
>>109623992And Gemma E4B thinks a bit too little, it's the duality of her.
>>109623960Which quant are you running? Last I checked the unslop Q2_XL was the only one of around that class that fit my 864gb server. The rest was either Q2_XXS shit or too big.
Hello,I'm very new to using local models so please excuse my ignorance.My machine is a 5800x3D, 9060XT (16GB VRAM) and 32GB of DDR4 system ramI am using ollama-vulkan on linux (arch btw) and I have tested 3 models: llama3.2:3b, gemma4:12b, and qwen3.8:latestOn the llama3.2:3b it replies very quickly with not thinking stuff but is rather dumb.gemma4:12b is a rather nice mix of intelligence and speed, includes a thinking phase, but it often cuts off before finishing it's reply.qwen3.8:latest seems the most intelligent, at least at the coding stuff I tested it with, but it replies rather slowly and also cuts off in the middle of responses sometimes during the thinking phase.Now again I really don't know what is actually going on or what I'm actually talking about so I apologize for that.My main question for you guys is why are the replies cutting off before finishing and is there a way to fix that?If you guys have any follow up questions please feel free to ask them. I will be online for the next few hours and would really like to get this sorted out :DAlso, again, I am a total noob. It's entirely possible I didn't read something or didn't set something up correctly. All I did was install ollama-vulkan via pacman and then install the models via ollama run gemma4:12b for example. I chose ollama-vulkan because at first I was using base ollama but noticed all the models were very slow and radeontop was showing no gpu usage. An online AI (grok or gemini idk i use those interchangebly) suggested I use ollama-vulkan.Thanks again :D
>>109624001like all frognogs, he's a tourist from reddit >>109623051
>>109624013Please stop using ollama and learn to configure llama.cpp yourself
Noob here, got unsloth set up and have dabbled with a few local models. I want to go beyond just using it as a glorified chat/search engine, how do I start using AI for things like converting a batch of files from HTML to markup while retaining the folder structure, or organising my image folders with tags, or setting up a model with some long-term memory for ongoing projects?
>>109623903I know someone like this irl constantly showing me his own custom software. Note taking applications, text editors, window managers, etc. It's all trash. He has way too much time on his hands and all of his custom software has no features and is so buggy that it can't not crash or freeze for the 30 seconds he shows it off to me. I'll show him existing software that does what he is trying to implement, but has more features and stability and he just responds that he either prefers his own or will use it for ideas for new features to add to his claude-made shitware. I think it's a mental illness.
Okay who fucking posted about the general on twitter or discord again?
Holy schizo
>>109624052It was me. Your welcome.
Ox alpha gave me a thought that AI will be free in the future.
>>109624035You need to get a harness like pi and connect it to your model via an API. I suppose unsloth has an api endpoint, as it uses llama.cpp under the hood.If you were using llama.cpp directly, it's web interface can also be a harness if you enable tools.
Let me draft:
>>109624078Zuck says AI is a human right, like water.
>>109624090When a jew says it it comes off in a Nestlé kind of way
How long do you think it'll be until we have a model that can properly handle each character's information and perspective of events relative to what they've seen with pure general reasoning as opposed to having to do a lot of context smoke and mirrors to keep every character from knowing some minor detail revealed they weren't present for?
>>109624112How many waifus do you have?
>>109624089>Her shivers ran down his spine shiverfully, with a hint of ozone and something darker; something uniquely *her*.Let me do a quick pass on the draft as per the user's instructions. Slop? No, it's very human — there are two humans in the story. Explicit? Yes, it's suggestive of an explicit action at some point. Great, let me continue and create the final response.
>>109624035hermes does all that and will make up words to sound more japanese if you use japanese voice with english tts. and then she will remember you thought it was funny and do it again the next day
>>109623533Yeah you can have all your gens in the right hand panel while you chat -> Gallery
>>109624112It feels like models have been getting worse than this rather than better. The recent big MoEs seem a lot more likely to let characters arbitrarily know the exact mechanics of my scenario cards for no reason. I have to prompt really hard against it which usually leads to the model going the other direction and really hamming up the part that everyone's very confused and has no idea what's going on.It's horrible. LeCunn was right.
>>109624176lecunny hasn't done a single productive thing in years
>>109624139Hermes sounds like a huge black box.
>>109624139Like that one person said Hermes is the only that is cute logo all others look like stylized buttholes.
>>109624112Something like hindsight or at least what I THINK it is.https://hindsight.vectorize.io/