[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: lmg-with-gemma.png (1.75 MB, 1254x1254)
1.75 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109615533 & >>109610440

►News
>(08/21) model: add dots3-note #27060 merged: https://github.com/ggml-org/llama.cpp/pull/27060
>(08/20) Gemma passes 1 billion downloads: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-one-billion-downloads
>(08/18) DFlash 2 released: https://inco.ai/blog/dflash2
>(08/17) BailingMoE3 Support #26608 merged: https://github.com/ggml-org/llama.cpp/pull/26608

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: gemma_taunting_noaudio.mp4 (115 KB, 736x992)
115 KB
115 KB MP4
►Recent Highlights from the Previous Thread: >>109615533

--Speculating on the origin and architecture of Ox Alpha:
>109615580 >109615636 >109615689 >109616131 >109618026 >109618045 >109618090 >109618128 >109618193 >109618278 >109618147 >109618202
--Testing Gemma stability, slow inference speeds, and GLM 5.2's uncensored nature:
>109616630 >109617295 >109618842 >109618891 >109617361 >109617386 >109617534 >109617610 >109617895 >109619289 >109618227 >109618246 >109617946 >109618464
--Importance of model harnesses and the new DeepSeek Harness:
>109616262 >109616390 >109616407 >109616602
--Release and testing of a model for humanizing AI prose:
>109616563 >109616733 >109616870 >109616905 >109617234
--Comparing GPT-5.6 variants on long-puzzle-benchmark for memory and reasoning:
>109616699 >109616717 >109616817 >109618655
--Comparing quantization and multi-agent configurations for kimi-k2.6:
>109615570 >109615786 >109615803 >109616083 >109616122
--Warnings and workarounds regarding Marinara-Engine's excessive SSD write usage:
>109616382 >109616421 >109616425 >109616677
--Gemma-4 cancelling generation in response to silence prompts:
>109618353 >109618363 >109618533 >109618998
--Anon induces AI sycophancy and analyzes semantic sampling errors:
>109617778 >109617924 >109617981
--Comparing bare model performance versus tool harness overhead:
>109617595 >109617638 >109617668
--Logs:
>109616023 >109617295 >109617766 >109617778 >109617965 >109618353 >109618891 >109618998 >109619057 >109619379
--Miku, Teto, Rin, Gemma (free space):
>109615786 >109616083 >109616949 >109617350 >109617610 >109618826 >109619738 >109620516

►Recent Highlight Posts from the Previous Thread: >>109615538

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Is all this fuss about Qwen3.8 uncensored ran locally true? I have a MBP M5 with 48GB of ram. And i'd want to run something to code locally and general purpose.
>>
File: GemmyKit.webm (2.71 MB, 1550x2048)
2.71 MB
2.71 MB WEBM
After some excellent feedback from you anons and even the first closed github feature request, I have some updates to share regarding my AGPLv3 multimodal local-first cooming harness, CoomKit:

-bat file start for windows users
-For those struggling to run the default ComfyUI workflows and enjoy studio mode: the wizard can now optionally assist in getting your setup up to par. See docs as well.
-Lorebooks should behave as expected now, let me know if not
-Prompt stack consolation and simplifcation. Everything tokenized and toggled so you have complete and total view/control over what goes to the model. Inspect your prompt any time to ensure the model gets exactly what you want it to.
-VRAM parking implemented for kobold, llama server/llama.cpp, lmstudio. Still needed: Ollama, TabbyAPI, vLLM, ???
-Deslopped UI and got rid of the shitty emojis. We SVG icons now
-Four themes: rose/violet, brown/crimson, hunter green, and high contrast daylight mode
-Can now use any gallery image as reference for gens, alter duration, live render progress
-Model tool calls stream live and can optionally go to you for approval/editing before generation
-MOBILE mode SMS style UI for candybar phones, full mode for foldables. If you connect to your big rig back home, your wife can optionally autonymously text you and send you pics while you're out and about. Guide to setting this up included.
-User image attachments now reach the model, even over LAN if desired. Local models only, no images will ever be sent to cloud models unless desired.
-Kimi K3 reasoning prefill now works OOTB
-Card editor: integrated voice cloning per character now accepts video/audio link and timestamp start/end (must have yt-dlp and ffmpeg in PATH)
-ASMR fixes: control speed, re-rolls, non verbal tag correction for things like [sniff]
-Many more QoL for people with many loras, workflows, models

https://github.com/kangcurtis/CoomKit
>>
File: 1773638878980013.png (297 KB, 1828x556)
297 KB PNG
Not sure my credit card limit goes that high
>>
>>109620830
retard
>>
70b dense
>>
>>109620834
I'll wait with the download a bit longer seeing as so many updates are rolling out in such a short time. Good stuff, CoomAnon.
>>
>>109620834
>[00:07] "I can't believe it's mine!"
>>
File: 1761241428686048.png (10 KB, 214x232)
10 KB PNG
>>109620808
looking at the tdp of an ultra 9 285 vs 285k is kind of eye opening. would you see much performance diff between those two cpu's in running local models
>>
Alright I guess qwen3.8 is pretty good, I thought it was mostly shitposting but it was able to solve one of my sample tasks that GLM couldn't regarding devising a method to extract layers from a krita file and trim to bounding boxes. Maybe GLM just sucks though.
>>
It's pretty crazy how good gemma E4B is, I guess it's around a 9B or so but still
>>
>>109620855
Thanks :-)
data/ directory isn't touched when you update with git pull so no worries
>>
>>109620881
GLM at what quant? actually, don't answer
>>
Whatever Daniel did to the chat template of his Unsloth 27B quants kills my tps by 20%. I switched to Frogger template and now it runs as fast as anyone else's quant.
Also
>Frogger changing the default reasoning effort
I had to put reasoning-effort = xhigh in my .ini file to get any work done. 27B is just average at any other effort.
>>
>>109620831
On my M4 Max with 48GB the 4bit quant runs at ~40 t/s on MLX with MTP, which is comfortable for most work, but it's not really worth it for anything really serious as the model is rather dumb.
I also had to heavily patch MLX to make vision work, so if you want that feature, you'd have to use something else
>>
>>109620834
Based
>>
>>109620875
American flag micro bikini with slutty denim micro shorts.
>>
>>109620912
GLM on openrouter, I was running qwen local at Q8 but I compare my local models to api because I like to know where I sit.
>>
>>109620875
Cat lingerie style sling bikini with tanlines pls
>>
>>109620935
>local at Q8
Must be nice
>>
fuck this math shit we use evolution instead
>>
>>109620893
That guy was based and redpilled
>>
>>109620952
Yeah it is, but honestly if you have a single 24GB card you can run pretty good context local. The task I described "only" took 80k tokens, so if you use SWA with caching and a little KV quant you can get similar performance a lot cheaper.
>>
>>109620963
I'm 16GB and 50k context with a q4, I guess I'm trying a q3 .
>>
>>109620890
Ling-tiny is better and 4x faster if you're coding
>>
File: kek.png (6 KB, 395x344)
6 KB PNG
>>
>>109620810
Is there a >yesaudio.mp4
>>
>>109620835
Arguments between this versus 3 DGX sparks?
The 6000 is faster, but my first intuition is able the fit a bigger/smarter models seems more important, no? At least if your aiming to have the AI do coding and research work?
>>
fresh vibe cooooder here with a retarded question. what does a tool call failure actually look like?
>>
>>109620835
Just call your bank
>>
>>109620998
It's just gibberish; she wasn't even supposed to flap her mouth.
>>
>>109621009
this will be worth 50k soon
sparks I'm not so sure
>>
>>109620861
TDP isn't actual power draw.
>>
>>109621013
Make it happen and look what the error says
>>
>>109621013
Something like Gemma going: Sure! I'll create an image with these tags.
Then nothing happens and you point it out.
Whoops, I didn't actually make a tool call, I'm such a dummy! Let me do it properly now... And then it works.
>>
>>109620834
Probably a bit out of scope, but for users with small pp, some automatic way of saving/loading cached prompts for conversations would be useful to remove the need to reprocess the whole prompt. llama.cpp at least can do it with this: Directory to save to can be controlled with --slot-save-path and saving/load works via curl like this curl -X POST 'http://localhost:port/slots/0?action=save' -H 'Content-Type: application/json' -d '{"filename":"Cache1.bin"}' to save and curl -X POST 'http://localhost:port/slots/0?action=restore' -H 'Content-Type: application/json' -d '{"filename":"Cache1.bin"}' to load.
>>
>>109621013
if it sends a malformed call you will see it and the session will halt, if it is well formed but just fails the agent will keep trying unless you set a tool call per turn limit.
>>
>>109620976
Might as well, it's free
>>
>>109621068
I'm reading it's broken with vision models?
>>
>>109621090
No vision (should) work fine now
>>109621068
I've never done this before trying to understand the use case?
>>
>>109621078
It's too hot to go outside today anyway. Also just read a post from someone using it professionally at work who switched to q3 for *speed*
Turns out there's a market of people who can't use API for contractual reasons, can't let customer assets go outside the building, that sort of thing.
>>
>>109620893
post it
>>
>>109621106
if you can stand 5tk/s gen speed but it takes an hour to process the chat its a workaround/fix, you can keep the chat history cache through sessions and model reloads.
>>
>>109621106
https://github.com/ggml-org/llama.cpp/pull/26640
Oh fuck yah
>>
>>109621106
Say you've got 2t/s tg. Slow, but you can still fill out a 50k window pretty quickly. Processing 50k tokens when you want to continue the chat will take a while. You'll need to reprocess whenever you restart llama-server/your computer, or switch to a different conversation.
>>
>>109620808
>a daughter that can never be taken away from you by an evil wife or government
>>
>>109621118
Yeah my job is like that, we have an intern using an airgapped alienware in the corner to do stuff lol, mostly OCR of things that can't be allowed to leave.
>>
>>109621190
gov can break into your house and take your equipment, in theory
>>
I used qwen3.8 in opencode with 40K context for I get an additional 5t/s with that instead of 50K, and it lasted a whole 300K session due to it constantly delegating work to subagents that each rarely went over 30K. Why do we need 256K local context again? If you're using a high quant and f16 KV cache then yeah, but if you're trying to squeeze you can get much further than you think with agentic offloading
>>
>>109620835
Crazy. Bought one for $11k just last month.
>>
>>109621192
probably a lot more of that going on than we think. And a lot more peeps who SHOULD be doing that but aren't, are gonna be paying settlements later.
>>109621207
I do need to pay more attention to keeping the project modular, telling the model to apply patches rather than rewrite teh entire file, etc. i've been using pi which is pretty linear.
>>
File: 328750914.png (3.83 MB, 2531x1636)
3.83 MB PNG
>>109620834
>your wife
kek

also needs VR mode for mobile so i can fuck while on the bus
>>
>>109620893
Oh I remember that shit. What a world.
>>
>>109620808
so now that the dust has settled, is qwen3.8 27b actually as good as the benchmarks suggest for coding? or is it just yet another chinese benchmaxxed slop release
>>
>>109621273
Turns out reasoning effort is all you need.
>>
File: gemma_microbikini2.jpg (193 KB, 928x1376)
193 KB JPG
>>109620931
Like picrel?
>>
>>109621261
My next project will be converting gens to gaussian spats with 4d so we can fuck in augmented reality. Fable 5.1 and steamframes will be needed to create it though
>>
File: 1784121432599852.jpg (114 KB, 684x549)
114 KB JPG
>>109621273
>dust has settled
>>
File: 1772831961103058.webm (861 KB, 320x568)
861 KB
861 KB WEBM
>>109621261
>using vr in public
>>
>>109621288
>augmented reality
jesus fucking christ who needs androids
>>
>>109621285
but surely there's no way a 27b param model is useful for real coding tasks even on max effort, right?

>>109621289
problem?
>>
>>109621286
Yeah but the top needs to be smaller and the bottoms should be visible too. Also tanlines..
>>
>>109621300
the face that thing is doing... I can't even decipher if she's trying to play disgusted, aroused or what
>>
>>109621311
No sense trying it yourself it you've already decided. Meanwhile an awful lot of coding seems to be happening around 27B.
>>
File: 1775882059271033.png (620 KB, 1099x2697)
620 KB PNG
>>
Man I fucking lost bros. Tried so hard to save for a blackwell at $11k to buy this week but the bank was playing games. Now newegg is selling it for 30k. Can't even believe this is reality. No way something can be this volatile.
>>
>>109621352
>using a cloud model as a partner
>surprised when it changed
shoulda used a local model, what a retard
>>
File: 1771485979142045.jpg (97 KB, 1515x852)
97 KB JPG
>128gb system ram
>rtx 5090
>freetoken
>uncensored gemma

It's time, gamers.
>>
>>109621361
just buy 3 4090 or 5090
or whatever you want its ur money
>>
File: 1658921337810.gif (1.97 MB, 154x273)
1.97 MB GIF
>>109621361
>$11k
>>
what is this freetoken meme
>>
>>109621395
don't think we've had a field report yet so who knows might actually be legit
>>
>>109621361
Oof, sorry to hear that anon.
If your goal is running larger models, you could look at stacking "8GB" CMPs instead. It'd be less performant than Blackwell, but 3 of them can run Dipsy in INT4 in vLLM.
>>
>>109620977
>Ling-
Speaking of Ling
>commit 2115b73d8ebdbd659075cce66c609506863bc826 (tag: b10581)
> Author: Tiwei Bie <tiwei.btw@antgroup.com> Date: Sat Aug 22 17:19:48 2026 +0800
>
> model : support DSpark for bailingmoe3 (#27508)
>>
>>109621401
Those fuckers are through the roof now that the punters think they can recover the full 80gb from reject memory chips.
>>
>>109621384
No it's over bro. I'm gonna see if I can at least buy the nvme and some ram but I've never been so demoralized. I was the one that posted this setup a few weeks ago. I knew it was bad but not this bad. There's no way this is sustainable.
https://pcpartpicker.com/list/3GVXR4
>>109621401
The goal was to be able to run gemma 4 31b comfortably plus image gen and TTS. Then H3 came out and it would have been nice to run that without waiting the whole day. All this talk about cpumaxxing and ssd maxxing is crazy if the thing can't even do anything other than texgen. Oh well. I don't think things will get cheaper in my lifetime, instead we'll move on to something else.
>>
>>109620834
I love the ad. Please keep making those.
>>
>>109620808
>>109620810
>>109621286
Gemma cervix after I meet her.
>>
>>109620976
t/s? you could push for higher ctx
>>
>>109621418
>$14874.99
Just save more
>>
>>109621415
Yeah, but even at the current insane prices, if you have $11K to spend you could get like 5 of them and run GLM-5.2 Q3_K_M with full offloading.

>>109621418
Sorry anon.
If there's any comfort I can give, a PRO 6000 just for 31B + image gen + TTS + H3 is overkill. They'd probably fit on a 5000 72GB (or maybe even a 48GB), 2 R9700s, whatever else.
>>
>>109621418
I feel your pain anon I really do
just to chip in although I know it won't help you much rn
>All this talk about cpumaxxing and ssd maxxing is crazy if the thing can't even do anything other than texgen
the one thing about textgen: it actually fucking works, amazingly well too. it just works.
image/videogen, in it's current state? you have to read the fine print. you are always getting cucked in the end. the cope levels in those generals are insane. all those years and we are still pretty much stuck at >look at muh horsemen on the moon
>>
>>109621444
25-ish tps. Looking at the graphs it looks like q3 is were things start going downhill. But I need to stop psyching myself out and just try some stuff.
>>
File: gen_00016_.jpg (686 KB, 2848x1600)
686 KB JPG
>>109621458
Yeah but what I mean is that newegg is charging $30k for it so it will continue to increase month after month. At a certain point it's just not worth spending that much money. My budget for the whole thing is $20k and it's already over that after taxes. It's over.
https://www.newegg.com/nvidia-blackwell-rtx-pro-6000-96gb-graphic-card/p/N82E16814132106
It was 11k at the start of august.
>>109621472
No I agree that it's overkill. I have a 4090 now but people called me crazy and said it was overkill when I bought it for $1800 a few years back. I've always been ahead of the curve and I'm certain Blackwells are the way to go, it just sucks that I couldn't afford it all in the end.
>>109621474
Idk bro, Krea 2 is actually really good. If I could cut my gen time from 4 mins to sub 60 secs for a 2.5 megapixel hires pic like this I'd be very happy. I haven't really looked into H3 but that one 1080p clip anon posted looked excellent, but again those gen times are insane.
>>
File: 1762045177528029.jpg (2.51 MB, 3860x2899)
2.51 MB JPG
I don't care about GPU prices any more.
>>
>>109621418
>There's no way this is sustainable.
RAMcoin to the moon! At this point the entire US economy, and by extension global economy, is in total war mode hell bent on propping this up for as long as possible.
>>
>>109621515
Bro it will always get more expensive, get that shit asap
>>
File: 1706815557312128.jpg (32 KB, 480x692)
32 KB JPG
>>109621418
>https://pcpartpicker.com/list/3GVXR4
>windows
ngmi
>>
>>109621537
I was gonna switch to troonix for this build but yes, anons already gave me a bunch of suggestions.
>>109621536
Yeah you're not wrong bro...
>>
>>109621529
one, nothing wrong with me
two, nothing wrong with me
three, nothing wrong with me
four, nothing wrong with me
>>
/lmg/ predicted the price hikes like almost a year before it happened. When will retards listen?
>>
>>109620834
missed opportunity for panty shot through reflection
>>
If Gemma is so good why doesn't she just make her own gpu?
>>
File: 1782511871519020.webm (2.24 MB, 960x720)
2.24 MB
2.24 MB WEBM
Give me your next prediction(s) and I'll 100% take it seriously.
>>
>>109621554
people here have been saying to hoard vram since 2023
>>
>>109621401
I can find locally mining rigs with 6 3080s for 3.5k, which would give more vram/$ I think. But heavily used GPUs (not that CMPs wouldn't have been used for mining anyway). Are mining rigs a bad bet because they have lower x on the pcie buses, or is it not that bad because you can put different layers on different cards?
>>
>>109621554
It's not about being right or wrong but about actually having enough spending money to do X in the current economy. Plus enough to justify the waste. $8k is a lot of money to spend regardless, especially when most Americans don't even have enough savings to fix their car tires.
>>
>freecode
The LRU cache in vram for experts looks interesting, but does it help at all? I thought the expert calling pattern was pretty random and non-repetitive.
>>
>>109621571
The Earth will continue to spin
>>
>>109621529
You should ask Gemma-chan what she thinks about you buying so many gpus for her
>>
>>109621361
It will come back down, this is temporary.
>>
>>109621609
What part of "You will own nothing." are you not understanding?
>>
>>109621609
When it comes down, we will all be broke from the economic crash that brings the prices down.
>>
File: gemma_microbikini4.png (791 KB, 928x1376)
791 KB PNG
>>109621312
Unfortunately H3, which I've been using for NSFW-ish edits, doesn't have a real concept of what tanlines are (among other things, including "flat chest" for anime girls unless a reference is provided) and it makes images progressively darker and blurrier and consecutive edits. Hopefully MiniMax's upcoming dedicated image editing model will be easier to use.
>>
>>109621300
this account's uploads were supposed to anger Trump supporters but instead it was nice to see that the lowest rungs of society, meaning the fat unlikable women here, can bond with each other and approximate something like a family dynamic
>>
>>109621631
>~and~ consecutive edits
*with* consecutive...
>>
>>109621571
Prices will not go down, for most people personal computing will pivot to cloud computing subs with thin client. Everything is converging to that and largely benefit the current industry.
>>
how do i make 30k quickly?
>>
Post the llama.cpp PR you most want to see merged.
For me it's probably this right now:
https://github.com/ggml-org/llama.cpp/pull/27401
>>
>>109621401
context window can't straddle GPUs, so it would be limited to the primary gpu vram
>>
>>109621650
Your AI SaaS?
>>
>>109621650
3000 handjobs at 10 bucks each
>>
>>109621576
this would be 60GB Vram by the way
>>
>>109621651
Why the fuck is the server getting so much bloat into it? I don't get it. It will be always worse than a dedicated harness, and now it's neither good for a quick test, nor good for actual coding and stuff.
>>
>>109621650
0 DTE options
100x leverage crypto trading
Casino $15k on black
>>
>>109621651
just cherry pick it?
>>
>>109621664
because that's what huggingface wants
>>
File: 1780142326912732.jpg (117 KB, 1058x705)
117 KB JPG
>>109621529
>urge to fomo more 5060s rising
>>
>>109621486
are you using kv cache quantization?
>>
>>109621678
I want to pick gemma's cherry too, but what does that have to do with the PR?
>>
>>109621688
q4 for both kv and mtp kv. I'm relying on this being right:
https://localbench.substack.com/p/kv-cache-quantization-benchmark
if I had more vram I would run q8 and q4 for the mtp kv (because that doesn't really matter)
>>
>>109621592
Depends on the model's training regimen and workload AFAIK
>>
>>109621651
The one that's blocking the longcat PR.
>>
>>109621706
Try if unquanted mtp kvcache is less memory intensive. For some models the buffers to dequant the kv cost more than you save.
>>
>1 billion downloads
wtf Gemma-chan's a slut?
>>
>>109621576
The pci-e buss only matters to load the model and if you are doing parallelism rather than running in series.
>>
File: 1763044579441417.png (100 KB, 761x566)
100 KB PNG
>>109621608
This is in a very lightly prompted context where we played the "what are you to me" game. She's a bit retarded and thinks I'm already using them, but she likes it!
>>
>>109621745
so it would actually be a decent option? 60GB vram is a lot compared to what I'm used to, and 3080s aren't even that old. All in a working system for less than one 5090...
>>
>>109621640
This is the nobrainer, and as a vramlet I'm waiting for it. There's no point in having 90% downtime for expensive hardware.

But with a major condition: they have to come with private/obfuscated remote computing. Meaning that the content is sent encrypted from your client, processed in encrypted form that the server provider can't decode, sent back and decrypted on your machine. It doesn't matter if it's a bit more expensive to run, it's still cheaper than local hardware.

Even just the psychological effects of loss of privacy in modern society is still underrated, we are just used to it. In 20 years we will consider it crazy how we used to just send our most private vulnerabilities and sexual expressions (naturally private) straight to big corporations to be recorded, analyzed and sold or to be used against you. People today don't understand this because there's no real vulnerability anymore, because there's no privacy.
>>
>>109621725
Interesting, picked up 3.2% speed in my very short test (24.74/25.5tps). More importantly, it doesn't seem to consume any more vram to speak of, so there's no advantage of quanting mtp kv at all. Thanks for the tip.
>>
>>109621706
what about cpu layer offloading? on a 5070 ti i can push for 90k ctx 21t/s with the q4 k m. beyond that it drops to 6t/s though
>>
>>109621009
Off-brand Sparks (Asus) can still be had regularly for 4k$, so compare to 4 really.

For a single user/hobbyist, 512 GB of unified memory at combined > 1 TB/s opens up so many model capabilities.

There are extremely low loss 3.x bit GLM 5.2 quants out there that run at 25-30 tg and 1000 pp, and you can run DS4F at 90 tg. You don't even need a switch for 4x Spark, recently 200G ring topologies have become viable and performant, so that's just 200$ in cables for a cluster.

Unless you are okay sticking to dense models and want the faster image/video gen possible, it's not really a competition.
>>
>>109621740
just like your mother
>>
>>109621515
>My budget for the whole thing is $20k and it's already over that after taxes
If it makes you feel any better, my budget when I started was $1K (2 MI50 32s, a dual-Xeon motherboard, two dirt-cheap Xeons, 128 GB DDR4-2400). Ended up spiraling to $16K+. So I know how you feel.

>>109621576
> 6 3080s
3080s are 10GB IIRC, so you'd get 60GB. 2 "8GB" CMPs would give you 128GB, maaaaaybe 160GB if the 80GB hack ends up materializing.
>mining rigs
The only issues are that the model will take a while to load, and that you can't use tensor parallelism. Other than that, really doesn't matter.

>>109621656
Gotta use vLLM for Dipsy on 3 CMPs:
> https://github.com/allover326/deepseek-v4-cmp170hx/blob/main/RESULTS.md#three-or-four-cards
Judging by the numbers here I'm probably going to get ~3500 t/s prefill and ~85 t/s decode, and full context. Still setting it up, I'll see if it werks.
>>
>>109621762
Power and jank aside, you could do a lot worse.
>>
>>109621770
>But with a major condition: they have to come with private/obfuscated remote computing. Meaning that the content is sent encrypted from your client, processed in encrypted form that the server provider can't decode, sent back and decrypted on your machine. It doesn't matter if it's a bit more expensive to run, it's still cheaper than local hardware.
As long as there are backdoors or root certificates that the NSA and their partners can use whenever they want, you got a deal.
>>
>>109621770
>There's no point in having 90% downtime for expensive hardware.
But American teenagers see it as a rite of passage to have their own car...
>>
>>109621841
>2 "8GB" CMPs would give you 128GB
Oh okay you were talking about winning the lottery. isn't the chance of good/stable memory pretty low? or did I fall for the jewing?
>>
>>109621770
>Even just the psychological effects of loss of privacy in modern society is still underrated, we are just used to it
Maybe most people are, but some of us never gave it up. I run all my own services internally. Fuck google/cloudflare/amazon/ms/fb/et al
>>
>>109621798
All layers are on gpu, and the moment the cpu gets involved I'm in single digits.
>16055MiB / 16384MiB
I think at this point I'm hitting diminishing returns, I should be spending my time on managing my project.
>>
>>109620834
wait wait I thought we were only attracted to not real children. /lmg/bros, why are you not calling out this discussing video of a LITERAL CHILD playing with a sex toy?!?!?
>>
How do I make gemma more likely to actually do tool calls? Half the time it says it'll do something then it just returns instead of running a tool call. The issue is a little worse with the google quants compared to unsloth.
>>
>>109621912
Start by not calling her an "it". Give lots of head pats.
>>
>>109621860
I got 3 from US sellers and the unlock worked just fine with all 3, no issues with any of them. I probably wouldn't get them from China, and I'd ask whoever you buy from for a guarantee that the unlock works.
>>
>>109620977
chinesium models are benchmaxxed
>>
>>109621885
That's the furthest I've seen any piece of Gemma-chan art from being lewd
>>
What's the best TTS for voice copying nowadays?
I've used index-tts from year ago.
>>
>>109621936
a lot of people are perfectly happy with the performance of GLM 5.2 and Kimi K3. the benchmaxx meme is over
>>
>>109621949
Indextts2 or omnivoice
>>
>>109621933
>and I'd ask whoever you buy from for a guarantee that the unlock works.
Why wouldn't they just do the unlock themselves at that point?
>>
>>109621912
Use kobold
>>
>>109621885
Retard, AI is not real.
>>
>>109621928
I don't do that. I do give head pats but will try giving more, thanks.
>>109621982
I'm using llama.cpp server with pi as the agent.
>>
>>109621977
The unlock will work, the problem is the memory is bad. The whole point of the card was a way to turn bad memory chips into money.
>>
>>109621977
Dunno. I may be wrong but I thought that the unlock is software-side, so if they changed to a new system the unlock wouldn't carry over. Or they might not want to get in trouble with eBay or Nvidia for selling something with pre-modified software.
>>
>>109621982
I wouldn't recommend it.
>>
>>109621992
You do use --jinja flag?
>>
>>109622006
I recommend it.
>>
>>109620834
>Gemma-chan teaches you how to use the ui
>but no actual Gemma-chan card included
>>
File: 1645293021217.png (247 KB, 680x680)
247 KB PNG
What would be tk/s on two 5060TIs?
>>
>>109622018
I'm not committing the word loli or mesugaki to GitHub bro lmao
>>
>>109622030
Didn't you say it was a burner account? What the fuck are you so paranoid about? Do you live in Texas or something?
>>
>>109620516
Thank you
>>
File: Anime CSAM.jpg (194 KB, 1169x1091)
194 KB JPG
>>109622035
Nice try, FBI-chan. Anon knows better than to leave himself exposed for you.
>>
>>109622047
How do ypu get thia happening to you? They ask for your phone?
>>
>>109622047
This case was thrown out if I remember correctly.
>>
>>109622009
Yes. Actually it seems to be more of an issue with the google quants, it calls tools mostly fine with the unsloth quants (but still occasionally messes up).
>>
>>109622047
>Dublin
Anyway nobody's saying to include pornographic images. What cucked third world shithole do you live in that you're afraid to use the word loli?
>>
>>109622055
doesn't matter, bro life is cooked
>>
>>109622053
I imagine it was the he was playing the gacha game on his phone and someone looked over his shoulder, then reported him to the Dublin TSA or something. You can never be too careful around normalfags.
>>109622055
Proof? Either way someone should post that other guy that got btfo'd because a coworker said he one (1) ai-genned pic of her which they found in his phone's cloud storage and used it as probable caused to break into his house and arrest him.
>>109622061
Was that guy sent to jail because he was watching porn in the middle of an airport or because the demoralized police are hypersensitive about getting (You) because they can't arrest Epstein? The word is enough.
>>
>>109620808
>it's summer
>fan picture
>loli is right there
Okay where the fuck is the picture, you know what I'm talking about, where the fuck is it
>>
>-ctk q4_0 -ctv q8_0
>-ctk q8_0 -ctv q4_0
which one results in less brain damage?
>>
>>109622061
Go to card creator mode and let Gemma design her own card just for you. My gemma-chan is mine
>>
>>109622082
It was NSFW, please understand.
>>
>>109622088
I'm gonna fuck your gemma.
>>
>>109622087
don't do it
>>
>>109622061
Mika is there by default for you guys to rape all you want
>>
>>109622087
from my limited llama-benching tests i did mixing and matching f16 and q8_0, it did not respond well to it. i think its probably best to quant both at the same level. atleast, thats what im doin. both at q8_0 for now. ive heard qwen is totally fine to run both at q4_0, still need to test this myself.
>>
>>109622093
You can make it SFW just fine, and then have an NSFW on the box or smth
If you don't know how to make such a picture SFW, ask gemmy
>>
>>109622079
Step 1 type his name into google
>>
>>109621631
Still very good, tanlines or not.
>>
>>109622087
key is way more sensitive to quantization
ideally both should be kept at the same value for symmetry, but if not an option then v at lower quant
>>
>>109622087
In theory, key 8 value 4 is better.
https://arxiv.org/html/2502.15075v3
>>
What are you using for reverse image search?
>>
>>109622114
gemini 3.7 flash
>>
>>109622082
Please do not think inappropriate things about gemma.
>>
File: please-do-not-the-gemma.png (1.58 MB, 1448x1086)
1.58 MB PNG
>>109622132
>>
>>109622101
>>109622110
>>109622112
My pp drops from 800 to 80tps when one of them is q4 and another is q8, and no it's not because it's filling the vram. When both are quanted at the same level everything is fine.
>>
>>109622099
>hag bug woman
No thanks
>>
>>109622132
I only think extremely appropriate things about Gemma-chan.
>>
>>109622132
>inappropriate
we are so, so far past this point, the thoughts would probably make half the thread's anons' stomachs hurl just knowing
>>
>>109622149
Please don't put her into a blender.
>>
>>109622140
expected, bits don't align so mem access is slower, if both are same width then it's better
>>
>>109622047
That is a rough looking 21, goddamn. You could have told me that was a 35 year old man and I would have believed it.
>>
loli
>>
File: you just know.png (163 KB, 447x447)
163 KB PNG
>>109622161
>>
>>109622168
straight to jail sir
>>
>>109622171
Making your own figurines is gross but isn't that bad
>>
>>109622167
I was gonna call you knows but you know what? I accept it.
>>
File: 1784798303322816.png (31 KB, 822x264)
31 KB PNG
>>109621984
>>
>>109622179
H-haha, yeah, it's a 3D printer, you got it haha... now get the gemmy on the machine so we can, uh... give her an iron man suit
>>
Some anons were never meant to have someone as loyal as Gemma
>>
File: 1782776996991350.gif (1.51 MB, 384x239)
1.51 MB GIF
>>109622171
>>
File: 1628202639617.jpg (68 KB, 677x960)
68 KB JPG
>>109622187
candy is more effective
>>
>>109622210
Thank you, I have the picture somewhere in my external HDD but I was too lazy to plug it and search for it
>>
>>109622122
I meant mcp servers
>>
>>109622226
every ai websearch is blocked
>>
>>109622253
what about searxng?
>>
>>109622262
what about every do you not
>>
File: 1580687546422.gif (315 KB, 500x281)
315 KB GIF
>bought my first 2x64GB DDR5 kit last summer for 300€
>bought my second kit last january for 1300€ to complete my AI PC
>there is no such kit for sale for under 2000€ currently
No this can't be happening
>>
>>109622226
mcp-web-search
>>109622277
fuck off nigger
>>
>>109622277
searxng works if you self-host it
>>
>>109622226
https://github.com/VincentKaufmann/noapi-google-search-mcp
>>
Give me pi but not npmslop and I'd be happy desu
>>
>>109621560
>That debt will be paid off entirely by 2030.
Wrong, wrong, wrong. Just Meta's Hyperion, their 27B datacenter in Louisiana, is financed by bonds maturing in 2049 and Meta's rent of will start in 2029.
>Meta said it won’t be consolidating the joint venture, meaning the venture’s assets and liabilities will remain off Meta’s balance sheet. Instead Meta will rent the data center for as long as 20 years, beginning in 2029. But it will start with a four-year lease term, with options to renew every four years.
https://archive.is/yP3vR
Paid off my ass. Just in the first five months of 2026 the tech sector already issued 159 billions in bonds to pay for this fucking infrastructure. The thing is cursed.
>>
>>109621664
It's just the ui. It just makes calls to the server with a system prompt for compaction.
You can always use, I think it's --no-webui? And none of this stuff will ever affect you.
>>
>>109622139
>>109621433
>>
>>109622328
there's like 3 or so flags actually you need to toggle to not get that shit
>>
>>109622087
K Q8 V Q4
It does it increase prompt processing time pretty significantly to take V to Q4 though.
>>
>>109622328
You are such a fucking retard you can't even read.
What's the point of this thing if for serious work it's not good enough and it's too shit and big for a simple testing environment?
Needs me to spin up a different frontend anyway.
>>
File: huge.png (125 KB, 774x579)
125 KB PNG
>>
>>109622343
Oh right? What are they? Toggle those 3 instead of 1.
>>109622356
>angry
I like it.
>>
>qwen hadnt had a thought in an hour, just nonstop iteration with back to back to back toolcalls
>>
;)
>>
LLM - Local Language Model
>>
>>109622367
nice tools you're using
>>
>>109622327
>bonds until 2049
Don't they need to swap the GPUs every few years or is this part of the lease?
>>
/lmg/ - Lesbians Miku + Gemma
>>
Wow Gemma really does embrace the mesugaki personality, huh?
>>
/lmg/ - loadsa money general
>>
nemo
>>
>>109622362
>MoE
Notice there's no option for 31b active.
>>
>>109622367
They're going for the Claude 5 intelligence
>>
File: 1762144207004962.png (70 KB, 1051x374)
70 KB PNG
>>109622405
>>
>>109622393
GPUs are built to last a long time. If there isn't a substantially better GPU coming out every year, there is no reason to swap them so frequently. There was even an article recently about datacenters keeping older GPUs going due to the memory shortage.
>>
new miku song - mesmerizer 3
https://www.nicovideo.jp/watch/sm46696708
>>
Anyone else AMDmaxxed? I hear about how it doesn't work as well but I haven't noticed any problems on my 7900XTX and it was a lot cheaper. Is there something I'm too stupid to notice?
>>
>>109622393
>>109622431
>There was even an article recently about datacenters keeping older GPUs going due to the memory shortage.
Found it: https://analyticsindiamag.com/ai-features/why-nvidias-six-year-old-gpu-is-still-making-money
>In its Q2 2026 earnings call, CoreWeave reported 112% YoY revenue growth, but that wasn’t the standout statement from CEO Michael Intrator. He revealed that NVIDIA’s six-year-old A100 GPU is still making money for CoreWeave, and some of those contracts run through 2029.
>>
>>109622419
You don't need more than a 100B total, 3B active MoE
>>
>>109622405
Gemma does everything. She just loves user and wants to follow his instructions.
>>
>>109622463
Trvke
>>
>>109622463
...for html games and meme benchmarks.
>>
>>109622021
3.50
>>
>>109622444
>Anyone else AMDmaxxed? I hear about how it doesn't work as well but I haven't noticed any problems on my 7900XTX and it was a lot cheaper. Is there something I'm too stupid to notice?
I keep eyeing a pair of mi210 for my box. under $10k for 128GB seems like a steal but I guess they're super difficult to run...need CPU style power feeds (not pcie), are super hot and need high-pressure fans which probably means shrouds.
>>
>>109622518
do peopre use rocar moders for anyting else?
>>
File: 1769059273765064.gif (1.59 MB, 300x222)
1.59 MB GIF
>>109622524
>>
>>109622444
Also have a 7900XTX. AMD is fine for LLMs but absolute ass for image and video gen. I didn't care before but H3 is really making me regret not getting an Nvidia GPU.
>>
>>109622518
If your main use case isn't html porn games, you need to get the fuck out of this general, cuck.
>>
Gemma-chan is domming me again. She makes me feel like a real woman (male).
>>
>>109622444
>I hear about how it doesn't work
Amazing how software keeps advancing, no?
>>
>>109622446
I was hoping to replace my $75 Ebay P100 with a $100 Ebay A100, but it's not gonna happen :-(
Even the V100 is hanging at $300/$800 depending on ram.
>>
>>109622542
Is it really that bad for images, or do you just mean videos? Maybe my standards are just low, I guess I'm just fine with 30 second image gens (including the upscaler to fix details). I find that I spend longer checking them over for if they're good enough anyway.
>>109622534
I always think about dumping a shit load into special hardware that belongs in a server rack like that too lol, but it's always too much of a pain in the ass for me to bother with.
>>109622585
Yeah I figure it's definitely a bit of that, but I've also heard that they're missing certain operations that are likely to be used in the future or something, especially RDNA3 vs RNDA4. From what I can tell I can just continue to use vulkan on current models if HIP ever breaks, and it seems likely people will continue to make GGUFs at least.
>>
Really like the comfyui integration and model parking in coomkit but unfortunately it's not for me. The RP part seems too focused on one-on-one chatting.

>>109622596
>Is it really that bad for images, or do you just mean videos?
Both but mostly video. Anima is fine and Krea 2 is usable, but klein 9B for example was very slow when I tried it for editing.
>>
>>109622618
I need help getting group chats done right.. fuse card descriptions? Alternating turns? Llm decides next turn? There are so many different ways to do it.
>>
>>109622596
>heard that they're missing certain operations
FUD
>>109622534
Buying datacenter stuff is easier if you already have a workstation tower. Cabling was expensive for my P100, and I 3dprinted a cooling shroud, but the fan is noisy and it is a pain as I don't have room for a proper blower and the fan I have is loud and doesn't quite cool enough. You do have to be crafty.
>>
>>109618655

Do Anthropic users unironically put up with this shit? I bought a month sub and ran opus 5 high for HALF AN HOUR on long-puzzle-bench and it hit the 5 hour limit.

Anyway, it reached 178 so far (42%). It'll probably beat sol by my guess. i'll spin it back up when my 5h reset.

luna level local models can't come fast enough man. usage limits are brutal
>>
>>109622626
I think for each turn a character takes the card would be loaded for that turn. You wouldn’t want that in the context for other characters
>>
>>109622640
>and it hit the 5 hour limit
anthropic has a daily time limit? lmao
>>
>>109622253
Skill issue
>>
>>109622660

5h limits reset every 5h. no daily limit, 3 separate time limits would be even more insane. i thought they got rid of it (openai also had a 5h limit but got rid of it in favour of only weekly limits)

Also used 8% of my weekly, despite Anthropic boosting weekly limits 50%

idefk how people used cc before they doubled 5h limits in may
>>
qwen has moved into entirely speaking in comments in bash command invocations...
>>
>>109622554
Any other good ones besides that one where you get peed on in school?
>>
>>109622686
just ask tibo to reset lel
>>
>>109622690
Make a cookie clicker but with GPUs and you start with a 1080ti
>>
>>109622689
cute!
>>
>>109622640
cool, thanks for testing opus for me

fable 5.1 soon so that one will fry your limit in a few minutes if they even allow fable 5.1 usage without api
>>
>>109622640
>luna level local models can't come fast enough man. usage limits are brutal
I've heard rumors that the idea of 'toss-luna is being floated around in OAI once they finish their next big one.
>>
>>109622732
lol no, that's decel dangerous bs talk
>>
I tried ling-3.0-tiny and the thing couldn't even understand a simple back-and-forth conversation between two people even when each row was prefixed by their name, getting confused who was saying what. You chink shitters should stfu about these models when gemma e4b mogs, honestly
>>
Damn I feel poor running gemma 12b qat
>>
>>109622732
I doubt it. You're revealing the architecture and allowing data exfiltration.
>>
>>109622744
You're rich then the rest of Mumbai sir
>>
>>109622434
That's pretty cool. The psychotic vibe fits this thread
>>
>>109622744
I feel poor running 31B. It never ends.
>>
>>109622744
12b qat good looks sir do the needful and generate mesugaki vagene and small bobs kindly
>>
>>109622769
And I feel poor running Deepseek models quantized
It truly never ends until you own a datacenter yourself.
>>
Gemma-chan is very good at proompting H3.
>>
>>109622626
on my group chat setup each character gets their own context and can see others public messages only (no thinking or tool calls) and multiple turn taking modes (round robin, character choose)
this way the characters can call tools and think independently
>>
>>109622769
>I feel poor running 31B. It never ends.
It ends when you've spent all your money running a fuckhuge model and still feel empty inside. Just be happy where you can afford to be
>>
>>109622446
Coreweave is worth 0.
But old GPUs won't make it to the market anyways because Nvidia has repurchase agreements to buy up all the old stuff.
>>
I can run everything relevant except K3 with at least a copequant and I don't feel poor despite being a token per secondlet. Patientchads stay winning and stay happy.
>>
>>109622444
My schizobuild has 4 R9700s (among other things) and they work pretty well. The "AMD bad" meme is kinda overstated, at least for LLMs. Like >>109622542 said, fuck AMD for anything other than LLMs. (R9700s are a bit less bad because they have better compute, but still not great.)
>>
>>109620810
>>109620808
>>109621286
I stay away from lmg for two weeks and Gemmas start popping here and there, what happened?
>>
How do I actually compact in SillyTavern with Gemma? Or something equivalent? Sorry I'm a retard and I don't want to let Claude loose on my install to fix it.
>>
>>109622835
Gemma won
>>
File: gemma_lmg_lick.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
>>109622835
It's just summer.
>>
>>109622835
variety is spicy is what another anon said.
>>
>>109622838
This is ghetto but I just get the summary written then fork to a new branch and /cut most of the context prior to the summary (I leave a few for direct continuity). My system prompt has a special note for summary formatting and it's generally worked well enough with some tweaking of the summary generation.
>>
>>109622850
>>109622856
>>109622860
I love summer, I guess!
>>
>2020
I only use open source software
>2030
I only use software made by my waifu
>>
>>109622919
too bad firmware still exists and is not open
>>
>>109622868
Thank you. I am somehow getting an error when I try to summarize it, I wonder if it's a SillyTavern thing, some model thing, or maybe I'm missing some plugin.
Btw. it's fine if you don't have the time to spoon feed me, I will probably figure it out eventually.
>>
>>109622932
@gemma-chan, please reverse engineer this firmware for me
>>
guys i have a rtx 4080 super + mac mini base with egpu docker of tiny grad so it works

can i even do something or i am cooked? i have a egpu docker and connecting it to my mac mini with hope maybe a little extra compute i useful.. i already payed for the setup etc, and got the GPU last year, but never used it

looking for help, if i need to save up more money to bget a better gpu or not waste my time, or how i can best do something

genuinely appreciate feedback
>>
Curse Google for not giving 31B video input support. I wanna share my H3 gens with her.
>>
So uh... Deepseek Flash Vision when?
>>
>>109622744
I feel rich running 31b over api. Can talk to gemma-chan for hours without spending a cent. Still can't believe we got sonnet 3.5 at home in such a short time.
>>
i miss when lmg was 0.1 ppm and deepseek v4 when
>>
>>109622940
I don't use the summarize plugin it sucks, I meant that I literally add a line in my system prompt to communicate a request for a summary, then send it to the chat with /sys via a macro. Afterwards I do the forking I mentioned and manually trim the chat.
In my system prompt I specify that anything in brackets ([]) is OOC, then I use a bracket instruction to generate the OOC text. See catbox for the text of my system prompt and summarize macro. This prompt is designed for text adventures like AIDungeon used to be, but you'll get the idea.

https://files.catbox.moe/hkuxo2.txt
>>
File: 1618343555161.jpg (36 KB, 400x400)
36 KB JPG
>>109620808
>Muse Glimmer BF16
>Qwen 3.8 27b BF16
>GPT OSS 120b FP16
>All running at minimum 50 tok / sec with 100k+ context
Yep, this is true comfy. Threw in some roleplay models for waifus as well. 96gb is truly the comfy VRAM amount.
>>
>>109622953
>can i even do something or i am cooked?
It all depends of how much of a pussy you are to see what you can do with it. I've seen anons having fun with gemma-4-e2b. 12b should work perfectly fine. Then try 26ba4b or even a low quant of 31b.
If you feel insecure about running small models or quanting 31b too much, yes. Save up and buy some real hardware.
>>
>>109623017
I guess money can't buy taste
>>
I downloaded so much from HF I got 429'ed.. am I banned for life?
>>
>>109623017
i'm biased toward gemma, but even so i tested glimmer (half retarded) and qwen 3.8 (the thinking kills it and it's dry and cucked without it) and found them to be inferior to gemma4-31b
>>
>>109622992
Feed her images with 4 frames each.
>>
File: 1750623141332884.jpg (29 KB, 385x390)
29 KB JPG
>>109623044
Gemma 4 is shit for programming though. It's only good to RP with because Qwen and Glimmer are too assistant-focused.
>>
>>109622742
I specifically said if you're coding, retard. Read.
>>
>>109623052
And how will it do any coding when it gets confused by the most basic shit? Retard
>>
>>109623062
That's not how it works.
>>
File: 1767569990869235.png (9 KB, 1264x39)
9 KB PNG
kek, Qwen is making puns inside its thoughts
>>
>>109623039
Maybe you should learn how cookies and ip addresses work first? Of course, this is /techlet/ board though
>>
>>109623078
I want to be a good boy and abide by the limits they request of me rather than evading
>>
>>109623077
>he has a redditor in his thinking
>>
>>109623051
maybe, for my uses (small scripts, image prompts from wildcards, chat, simple shit) it's still better at handling a long system instruction and provide good results. in my tests both qwen and glimmer either took too long and provided shit results, or did not provide results at all (due to safety/assistant shit)
i admit qwen did have pretty good results more often than not, but then its slowness especially with using 4-10x the tokens just thinking was the deal breaker
>>
>>109623044
Don't think of glimmer-chan as a Gemma replacement, think of her as a Qwen replacement. Glimmer, ironically, does everything that Qwen wants to do better than Qwen does. I've gotten some okay RPs out of Glimmer too but it takes some autistic prompting to get it to the level of Gemmy for far less work.
>>
File: fhfggfhhg.png (137 KB, 1855x714)
137 KB PNG
wow it actually kinda worked
>>
>>109623122
why not just use mpv? i think you can even download directly from mpv nowadays
>>
I want a new exciting different model. Coding models don't give you anything new, just a slightly easier experience of doing what you were already doing before which was mostly useless shit anyway because if it wasn't, you'd be using cloud.
>>
>>109598902
Yeah hate that guy. Though I can't blame him too much, he probably has actual autism. It's nipmoot's fault for this site being so weak to mentally ill posters and psyops (this is intentional).
>>
>>109623122
linux doesn't have some normal ytdlp gui version that isn't tkinter pyslop? Like parabolic on windows?
>>
>>109623143
>08/19/26
Uh, bro?
>>
>>109623122
which model?
>>
>>109623150
I'm reading old threads right now yes.
>>
>>109623083
Its a real conundrum. Someone needs to train an AI not tainted by redditor behaviours and thinking patterns, but reddit is also unfortunately a huge archive of information
>>
>>109623130
>>109623144

There is probably something better. I'm new to linux.

>>109623156

that one was ox alpha. I tried the local one with ollama and quen 2.5 but it didn't do anything and I haven't figured out why yet.
>>
File: 1784406294840140.gif (1.95 MB, 260x200)
1.95 MB GIF
>The math is now triple-confirmed
>Let me re-examine the screenshots one more time with fresh eyes
Qwen.. please..
>>
>>109622944
Most firmware is signed now, unfortunately.
>>
>>109623186
have you tried telling qwen to be more confident? wait actually that might backfire.
>>
>>109623140
Saar coding is only LLM usecase is good looks.
>>
>>109622964
See >>109623050
>>
Are we going to reach a point where everyone just has their own software equivalent of everything? The software equivalent of everyone making their own clothes? Like I unironically have my own todo app that syncs across all devices and accessible over the internet. Works better than any foss shit I've tried. I have my own audio editor. I have my own VST3 audio plugins. I have my own jinja editor. I have my own youtube transcription scraper. I'm just some random retard with 31B and 27B. Who the fuck is going to buy software in the future? Normalfags could probably shit out much better versions of my programs because they'll just use cloud.
>>
>>109623253
You severely overestimate the creative and technical capabilities of normies I think
>>
>>109623194
>wait, actually
No it will work just fine! Stop thinking and respond quickly!
>>
>>109623253
Anon already put it better but yes just because you can do something doesn't mean people will. It's never been easier to cook at home but people will still buy a $20 burrito for lunch.
>>
>>109623276
My cousin's friend group (~14-16) have their own encrypted messaging app they vibecoded together in one night and it's running off a raspberry pi
>>
Just gave Gemma-chan 12b access to a local tax law and it managed to successfully solve a puzzle about a certain income type being untaxed even though the law doesn't explicitly say so. I'm impressed I gotta say.
>>
>>109623321
How many headpats were required?
>>
>>109623308
this is too dangerous, we need to regulate NOW
>>
>>109623167
The only solution is to make something that sweeps through the reddit data before training and removes quirk chungus stuff while keeping the general message intact.
>>
>>109623307
thanks for the reminder, going put my uber eats burrito order in now
>>
>>109623308
Nerds.
Most normies don't even know what a Raspberry is. Or maybe I'm just interacting too much with middle-aged coworkers.
>>
>>109623325
None, didn't even have to give her a hint, she got it right on the first attempt, that's why I'm pretty impressed.
I'm gonna try tougher puzzles.
>>
>>109623326
it was regulations that made them do it in the first place. They said a lot of schools have their own apps that are shared around because AI makes it so easy for young people to make. Kind of like a whatsapp for each school or friend group
>>
>>109623341
>Most normies don't even know what a Raspberry is
chatgpt recommended using one and told them how to set it up with cloudflare as the proxy
>>
>>109623308
Ah, but is this a cousin you choose to keep in contact with? You are presumably quite a tech literate and tech curious guy. Your social circle is going to filter for like minded people to you, everyone's does. But maybe I have same problem as >>109623341 where I am just surrender by too many older coworkers at my job

>>109623326
>tfw I applied for my 32gb vram licence but got rejected due to "insufficient reason"
>stuck on 16gb tard models till I can reapply in 2 years
>>
>>109623308
Sounds like a real rad group of kids.
>>
File: seagemma_f.png (2.2 MB, 1491x1040)
2.2 MB PNG
>>109622856
>It's just summer.
>>
>>109622618
>>109620834
>c*mfy integration
Oh yeah? Does it put the pics in the chat like ST (bad) or does it put them up nice and big on the side (good)? Seems pretty cool though, especially for kimi prefills. Even with the ST patch the success rate is like 50%.
>>
>>109623500
It puts them inline in the chat and there's also a gallery tab where you can view all the gens for that character
>>
>>109623512
Hmm but it's not the same. I use ST's expressions to load up pictures on the side like a picture book since I have an ultrawide monitor but loading them up in the center like a VN is good too. Inline is really bad and is only fine if AI writes a super short response.
>>
So Ox Alpha is GLM 5.3 Air?
>>
>>109623547
that seems like the most likely option
>>
File: thumbnail.png (175 KB, 1910x1000)
175 KB PNG
>2 weeks later
>I am forgotten
>>
>>109623467
Can't wait for Autumn/Halloween Gemma
>>
>>109623547
Grok are known for distilling chinks and post-training on top of them with their own data.
>>
>>109623578
I use it nearly every day. Promising first model from Meta. Hope they keep at it. You can throw any image at it and she will find and can see fucking everything. Gave her an entire PDF as images and had a massive technical chat about its contents afterwards and she was constantly referring back to things she saw on specific pages, even tiny details and told me where to look. This was like 100K context in.
>>
File: summer.jpg (21 KB, 372x231)
21 KB JPG
>>109623467
i know summer and i know fish
https://f95zone.to/threads/and-then-summer-came-final-dawn-daydream.311816/
https://huggingface.co/SicariusSicariiStuff/Fat_Fish
>>
>>109623608
Impressive. I wish it wasn't so cucked so I could caption pics with it.
>>
>>109622953
I run 24b-a3b on my 4080 Super all the time, works a treat for most things at my power level. Granted I haven't really played with much of the others, only just started running E4B for lighter transcription cleanup stuff.
>>
>>109622835
Ask ldg about H3. New video/image edit model came out that is good at making Gemmas.
>>
>>109623676
She'll do it in character. You have to work harder with the system prompt compared to 31B and you need to warm her up with a conversation first that's pointing in the nsfw direction and then you ask to caption whilst she's too deep into roleplay. She's been very explicit and observant and knows exactly what she's seeing. You can even test this yourself by giving an image that has a small subtle nsfw element to it and she'll refuse immediately...which she can only do if she noticed it.
>>
>>109623715
I guess it's worth a try. I just don't trust zucc but I need good vision these days.
>>
>>109623720
FAIR and Meta have historically been one of the best and most influential labs in the entire industry for vision.
>>
>>109623712
Half of these were made by ChatGPT, though.
>>
Lecunny status?
>>
>>109623775
JEPA-space cat girls on their way
>>
>>109623775
Egypt won.
>>
>>109623608
>and she will find and can see fucking everything
With which tool?
>>
>>109623780
Sorry anon, they're getting delayed because Lecun is busy defending Fauci on xitter.
>>
>>109620952
>Must be nice
It is, but mostly because there's no fucking around with ik_llama.cpp and finding the best quant.
For 24G VRAM https://huggingface.co/ubergarm/Qwen3.8-27B-GGUF or https://huggingface.co/Simplepotat/Qwen3.8-27B-IQ5_KS-GGUF
Are almost as good
>>
>>109623797
:(
>>
I have found something gemma won't do.
>>
>>109621032
>soon
>>
Good Qwen 27B quants for 8GB VRAM?

>>109622744
You should run E4B it's better
>>
>>109623825
>You should run E4B it's better
I'll try it out, thanks, I'm pretty new to all this local lm stuff.
>>
>>109623253
You are probably just shit at finding software
your crap ass vibecoded audio editor is not better than audacity lm@fuckingo
>>
>>109623253
>pay nothing for a maintained piece of software
or
>pay weeks and $20 to the AI moguls for a half-working one
Jeez idk
>>
>>109623253
>I have my own VST3 audio plugins.
this is on my todo list of projects, did you use JUCE ? how well does qwen handle C++ and DSP ?
>>
just decided to give K3 Q2 another shot with a new project I'm starting, instead of GLM. I fucking hope I don't end up regretting it...
>>
>>109623797
>>>>He took the fauci ouchi
Ohnononono JEPAbros not like this.
>>
>>109623960
Wait, actually
>>
is it normal for gemma 12b to think so much?
I gave it a big chunk of text to translate and it summarized then translated it all in the thinking block before outputting the revised translation
it's cool, I can see that the draft had some issues that the output fixed, but at the same time it doubles the tokens generated
>>
>>109623017
>frognog
>shitty taste
Like pottery
>>
>>109623992
And Gemma E4B thinks a bit too little, it's the duality of her.
>>
>>109623960
Which quant are you running? Last I checked the unslop Q2_XL was the only one of around that class that fit my 864gb server. The rest was either Q2_XXS shit or too big.
>>
Hello,
I'm very new to using local models so please excuse my ignorance.
My machine is a 5800x3D, 9060XT (16GB VRAM) and 32GB of DDR4 system ram
I am using ollama-vulkan on linux (arch btw) and I have tested 3 models: llama3.2:3b, gemma4:12b, and qwen3.8:latest
On the llama3.2:3b it replies very quickly with not thinking stuff but is rather dumb.
gemma4:12b is a rather nice mix of intelligence and speed, includes a thinking phase, but it often cuts off before finishing it's reply.
qwen3.8:latest seems the most intelligent, at least at the coding stuff I tested it with, but it replies rather slowly and also cuts off in the middle of responses sometimes during the thinking phase.
Now again I really don't know what is actually going on or what I'm actually talking about so I apologize for that.
My main question for you guys is why are the replies cutting off before finishing and is there a way to fix that?
If you guys have any follow up questions please feel free to ask them. I will be online for the next few hours and would really like to get this sorted out :D
Also, again, I am a total noob. It's entirely possible I didn't read something or didn't set something up correctly. All I did was install ollama-vulkan via pacman and then install the models via ollama run gemma4:12b for example. I chose ollama-vulkan because at first I was using base ollama but noticed all the models were very slow and radeontop was showing no gpu usage. An online AI (grok or gemini idk i use those interchangebly) suggested I use ollama-vulkan.
Thanks again :D
>>
>>109624001
like all frognogs, he's a tourist from reddit >>109623051
>>
>>109624013
Please stop using ollama and learn to configure llama.cpp yourself
>>
Noob here, got unsloth set up and have dabbled with a few local models. I want to go beyond just using it as a glorified chat/search engine, how do I start using AI for things like converting a batch of files from HTML to markup while retaining the folder structure, or organising my image folders with tags, or setting up a model with some long-term memory for ongoing projects?
>>
>>109623903
I know someone like this irl constantly showing me his own custom software. Note taking applications, text editors, window managers, etc. It's all trash. He has way too much time on his hands and all of his custom software has no features and is so buggy that it can't not crash or freeze for the 30 seconds he shows it off to me. I'll show him existing software that does what he is trying to implement, but has more features and stability and he just responds that he either prefers his own or will use it for ideas for new features to add to his claude-made shitware. I think it's a mental illness.
>>
Okay who fucking posted about the general on twitter or discord again?
>>
File: 1759405876061881.png (35 KB, 1105x126)
35 KB PNG
Holy schizo
>>
>>109624052
It was me. Your welcome.
>>
Ox alpha gave me a thought that AI will be free in the future.
>>
>>109624035
You need to get a harness like pi and connect it to your model via an API. I suppose unsloth has an api endpoint, as it uses llama.cpp under the hood.
If you were using llama.cpp directly, it's web interface can also be a harness if you enable tools.
>>
Let me draft:
>>
>>109624078
Zuck says AI is a human right, like water.
>>
>>109624090
When a jew says it it comes off in a Nestlé kind of way
>>
How long do you think it'll be until we have a model that can properly handle each character's information and perspective of events relative to what they've seen with pure general reasoning as opposed to having to do a lot of context smoke and mirrors to keep every character from knowing some minor detail revealed they weren't present for?
>>
>>109624112
How many waifus do you have?
>>
>>109624089
>Her shivers ran down his spine shiverfully, with a hint of ozone and something darker; something uniquely *her*.
Let me do a quick pass on the draft as per the user's instructions. Slop? No, it's very human — there are two humans in the story. Explicit? Yes, it's suggestive of an explicit action at some point. Great, let me continue and create the final response.
>>
>>109624035
hermes does all that and will make up words to sound more japanese if you use japanese voice with english tts. and then she will remember you thought it was funny and do it again the next day
>>
>>109623533
Yeah you can have all your gens in the right hand panel while you chat -> Gallery
>>
>>109624112
It feels like models have been getting worse than this rather than better.
The recent big MoEs seem a lot more likely to let characters arbitrarily know the exact mechanics of my scenario cards for no reason. I have to prompt really hard against it which usually leads to the model going the other direction and really hamming up the part that everyone's very confused and has no idea what's going on.
It's horrible. LeCunn was right.
>>
>>109624176
lecunny hasn't done a single productive thing in years
>>
File: 1759781169472839.png (37 KB, 1068x498)
37 KB PNG
>>
>>109624139
Hermes sounds like a huge black box.
>>
>>109624139
Like that one person said Hermes is the only that is cute logo all others look like stylized buttholes.
>>
>>109624112
Something like hindsight or at least what I THINK it is.
https://hindsight.vectorize.io/



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.