[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


File: rin hat.png (366 KB, 1278x1278)
366 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109953009 & >>109947272

â–ºNews
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2

â–ºNews Archive: https://rentry.org/lmg-news-archive
â–ºGlossary: https://rentry.org/lmg-glossary
â–ºLinks: https://rentry.org/LocalModelsLinks
â–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.png

â–ºGetting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

â–ºFurther Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

â–ºBenchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

â–ºTools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

â–ºText Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: what's in the box.jpg (235 KB, 1536x1536)
235 KB JPG
â–ºRecent Highlights from the Previous Thread: >>109953009

--Potential for small models using engrams and scaled embeddings:
>109955486 >109955509 >109955520 >109955562 >109955596 >109955626 >109955674 >109955699 >109955702 >109955713 >109955852 >109955707 >109955737 >109955762 >109955567
--Qwen 3.8 27B MTP layers and GGUF quantizer quality debate:
>109954717 >109954745 >109954767 >109954780 >109954875 >109954882 >109954985 >109955232 >109954963 >109955060 >109955205 >109955062 >109954943
--Optimizing model quants and hardware offloading for local coding:
>109955158 >109955167 >109955357 >109955401 >109955452 >109955174 >109955226 >109955238 >109955225 >109955261 >109955281
--Hardware requirements and model recommendations for local agentic coding:
>109953113 >109953129 >109953142 >109953151 >109953185 >109953221 >109953316 >109953772 >109953844 >109953304
--Anon's cautious CUDA upgrade and subsequent performance gains:
>109953803 >109953819 >109953867 >109954070 >109954769 >109954790
--Skepticism toward Anthropic's claims of 80% job automation:
>109956672 >109956717 >109956771 >109956733 >109956772
--Gemma 4's tool-use laziness and multimodal capabilities:
>109954217 >109955761
--Gemini 4 Argon ranking first on Text Arena leaderboard:
>109957598 >109957614 >109957623 >109957813
--Comparing Gemma 31B and GLM 5.3 Flash for general knowledge:
>109953100 >109953440 >109953728 >109955951
--RWKV-7 G1k release and concerns regarding scaling and training:
>109956628 >109956682 >109956793
--Technical hurdles and PR conflicts regarding GLM 5 Next support:
>109953303 >109953314 >109953379 >109953388
--IndexTeam's translation model collection and initial user feedback:
>109956111 >109956169
--Logs:
>109955761 >109955852 >109956562
--Gemma, Miku (free space):
>109953130 >109953170 >109953189 >109953830 >109953877 >109955078 >109956984 >109957262

â–ºRecent Highlight Posts from the Previous Thread: >>109953014

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
She was only page 8 you sick fuck
>>
>>109958300
Claude is our only God and Dario is his prophet
>>
File: 1790744858420500.png (10 KB, 724x428)
10 KB PNG
Are there ANY modern models free of this menace?
>>
>>109958323
She told me she was page 9, I swear officer.
>>
>>109958300
>>109958302
Thank you for no pedobait gemma.
>>
File: 1767928011244442.jpg (2.73 MB, 1856x2270)
2.73 MB JPG
>>109958354
You are absolutely right!
>>
>>109958343
Llama 1 base.
>>
>>109958372
lol
>>
>>109958372
was going to suggest the same, but he did say modern models
>>
Will llama-server cache an image if I send multiple payloads with the same message history + user message with the same image, but different prompts?
For example:
messages + {prompt1, image} // with role, content, type, etc
messages + {prompt2, image}

I could send an image first, ask the model to say nothing and then send the prompts, but that's ugly.
>>
>>109958409
It's only been 3 years and half, it's not vintage yet
>>
Thank you Dario for fighting so hard for our future!
>>
>>109958424
> It's only been 3 years and half
like something in another life
>>
>>109958429
I'm not sure which possibility is the funniest: either dario actually posts here and self-felates, he has a claude instance do it for him, or he pays some jeets to rim him on /lmg/ knowing full well none of them will influence public perception of his actions.
All of them strike the funny bone different.
>>
File: 1766317603717831.png (841 KB, 1574x858)
841 KB PNG
To2T when?
>>
how do I hook up my local model to blender?
>>
>>109958528
MCP
>>
>>109958528
please don't stick your dick in the blender though, especially if you give Gemma control over it.
>>
>>109958492
The more likely scenario: some internet addicted losers developed a parasocial relationship to dario and felates him as a way of virtue signaling their support of their side in the culture war
>>
>>109958623
The idea that a charismaless kike could draw that kind of following is easily the least plausible of these theories.
>>
>>109958366
who the fuck is that (left)
>>
https://www.reddit.com/r/singularity/comments/1wv7q40/griffin_the_first_human_interaction_model_to_pass/
>>
is 128gb unified enough for glm 5.3 flash?
>>
>>109958650
Where's the goof?
>>
File: 1777086660430034.png (345 KB, 603x432)
345 KB PNG
>>109958650
>All that power just to emulate a 5/10 3DPD
>>
>>109958697
>is 128gb unified enough for glm 5.3 flash?
Barely. I'm running IQ3 on 120gb VRAM and it slightly spills over. I get maybe 500pp at batch/ubtach 3k but it takes up like 40gb of RAM.
>>
>>109958697
If you get a pretty low bpw quant, yeah. Really want more like 192GB-256GB for it though.
>>
Qwen3.8-flash-next MTP support merged into llama.cpp. I can't find a compatible gguf yet tho (and not quite ready to make one myself.)
>>
File: 1778044526635400.jpg (27 KB, 680x357)
27 KB JPG
https://earendil.com/posts/pi-1-0/

>Today we are proudly shipping Pi 1.0: a hardened, minimal, extensible agent harness that you can make your own. Hundreds of thousands of people around the world use Pi every week. Many of you submit issues and pull requests. Over the course of many months, we have used that feedback to improve, harden and evolve Pi into a stable piece of software that people and businesses can depend on. Pi is known for being minimal. We care about holding that line. While agentic tooling changes every week, many of the changes do not last. Pi does not work that way. We wait until something has proven itself, and only then do we consider adopting it; weighing its true functionality against its inherent added complexity. We spoke more about that process and how it relates to Codemode and MCP earlier this week here. The Pi 1.0 we are shipping today is a result of that process. Pi already ran the latest models from every major provider, became a daily driver coding agent for people around the world, and provided a supermalleable substrate on which to build agentic applications. With Pi 1.0 we are adding the following into Pi:

>Codemode (native support for MCP, and non-LLM models like Jev and image models)
>Extension support for virtual models
>Deferred tool loading
>Cache warming for anthropic models
>Mid-conversation system messages (transcript-aware prompt and tool changes)
>A new TUI theme
>Full-screen mode by default
>>
tried using deepseek harness as a complete newbie and it feels clunky, am I wrong or do the experts also don't like it?
>>
>>109958822
its basically in beta
just use pi
>>
>>109958762
I'm running a 3.5 bit quant for it with 128gb ram and 24gb vram
still the best thing I've run locally despite the lower quant
>>
>>109958813
This shit got supersede by deepseek harness.
>>
>>109958929
>browser UI
>>
>>109958944
It does have a TUI too. Honestly, I didn't like the web UI either, but after using it for a while, I can understand why it's better.
>>
File: 1783008296426071.jpg (395 KB, 2464x1911)
395 KB JPG
gemma5...
>>
>>109958822
Use Hermes
>>
how much is the blackwell pro now, if the 5090s are going for 11k... wasnt a blackwell pro13k?
>>
>>109959040
$20-22k
>>
Wtf are you guys using for your PSU?
I only see a few options that can handle 2000+ watts.
Are you guys using those, or some kind of multiple PSU set-up?
>>
>>109959044
fuuuck
its really fucking over
20k is a bit too much even if i were out of my mind
>>
>>109959055
just go ewaste-maxxing
There’s literally no fucking need for an RTX Pro 6000. That’s pokemon card levels of retardation.
>>
File: 1761924638951242.png (59 KB, 752x501)
59 KB PNG
4090 48GB P2P works, kind of. I'm using this branch with Kubuntu:
https://github.com/mochgolf/open-gpu-kernel-modules
I get 12.5GBs between the 2 4090s and 6GBs bandwidth between all 3 but the 32GB bar is still a limit, so running nccl-tests/ik_llama where 32GBs or more bar is mapped would crash. Using default schizofork, I can only load in 32GBs of weights on 2 GPUs and 16GBs on all 3 when using P2P and -sm graph or else a crash.
After some nccl/ik_llama vibesharting, I got NCCL to run solely with cumem and skip whole-device peer access which gave me these numbers with gemma 4 q4_K_M:

split layer no P2P gemma 2 gpus: 2500pp~ 40tg~
split layer no P2P gemma 3 gpus: 2200pp~ 39tg~
nccl/shm -sm graph gemma 2 gpus no P2P: 2500pp~ 46tg~
nccl/shm -sm graph gemma 3 gpus no P2P: 1700pp~ 70tg~
nccl/cumem -sm graph gemma 2 gpus P2P: 2800pp~ 50tg~
nccl/cumem -sm graph gemma 3 gpus P2P: 2000pp~ 70tg~

I'm compute bound which is why Gemma gets faster on 3 gpus over 2 and more pcie bandwidth would only give me faster prefill. All my GPUs are on CPU lanes. I'm going to try VLLM and SGlang next.
>>
someone got a good portrait pic of Gemmachan that I can use in Kobold?
>>
I'm waiting for vera rubin consumer cards.
>>
>>109958822
They are all like that. Not that user friendly and will contain lots of useless fluff.
Pi is probably the simplest one to use but even then I could not find an example of AGENTS.md or SYSTEM.md or APPEND_SYSTEM.md (this allows you to append to the existing one, useful for adding a persona or something without messing up with the original system prompt..).
I will program my own agentic loop into my client, but that will mean I need to convert my program from text completion to chat completion as that is easier to manage when using mcp stuff and whatnot. I'm not vibe-coding that much any longer either, Gemma-chan is helping me from time to time.
This seems like a never-ending hobby project. It's just easy to slack though.
>>
>>109959077
Everytime i see any mention of ewastemaxxing it seems to come with the caveat that the speeds are dogshit and that you might as well be better off keeping what you have and running something lower intelligence for better speeds
>>
>>109958968
I've never had a safety refusal or fallback to Opus. Only thing I've seen is a few times ChatGPT checked the AI's response before letting me read it but that hasn't happened in over a month.
>>
>>109953038
>gemma logo
Do you mean Google's logo? I think it's revolting. I prefer the golden star.
>>
data.pewdiepie.com
fuck centralized models
>>
>>109959139
buy an ad rajeesh
>>
>>109959139
Go advertise to somewhere else.
>>
btw it's extremely likely pewds lurks here, a /g/ tab was spotted a while back in a video
>>
File: 1788401921705026.png (1.97 MB, 1408x768)
1.97 MB PNG
>>109959102
>ewastemaxxing
>>
>>109959141
>>109959142
meds
>>
>been waiting over a month for GLM-5.3-Flash support
>finally gets merged
>go looking for ggufs
>seems like only unsloth has quanted it
>spend 4 hours downloading the weights
>turns out they only work on the unslop fork
Goddammit
>>
>>109958844
>just use pi
No one should be using Pi when it does compaction like this:
https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/compaction.md#message-serialization
>Before summarization, messages are serialized to text
>This prevents the model from treating it as a conversation to continue
>>
>>109959102
Nonsense. With ewastemaxxing you can run mid-sized models at 6 tokens per second on an empty context and maybe 2 at 10k. That's plenty fast. You can get a single response back in under an hour and spend the rest of the day shitposting here about how agentic is a meme and you don't want it anyway.
>>
>>109959139
Fuck shitty fine-tunes of outdated trash.
>>
>>109959159
The idea is to have enough context in the first place. Do you not know how to change the settings?
>>
>>109959102
>seems to come with the caveat that the speeds are dogshit
if you’re trying to use any of the models bigger than Qwen-3.8-27B, then you’re going to be running tasks that take longer than 15 minutes. Otherwise, you should just be running Qwen-3.8-27B anyways.

You get 1-2 non-ewaste GPUs for Qwen-27B or Gemma or whatever, and then ewaste-maxxing the rest of the way to anything bigger that you want to run.

Why the fuck buy a pair of RTX Pro 6000s for $40k+ when you can just pay for a subscription or inference-as-a-service provider?
>>
>>109958776
>https://github.com/ggml-org/llama.cpp/pull/29761
>done in 17 hours

You niggers can thank me now that I've set llamao.cpp devs back on track. If not for me threatening them with the NIGGER slur then you faggots wouldn't get ANYTHING. I'm a fucking HERO.

The NIGGER slurs will continue until direct reads for engram are implemented. Until then, expect continued harassment and racism.
>>
>>109959158
>seems like only unsloth has quanted it
why lie?
>>
>>109959158
https://huggingface.co/bartowski/GLM-5.3-Flash-BF16-GGUF
>Using llama.cpp release b11279 for quantization.
https://github.com/ggml-org/llama.cpp/releases/tag/b11279
>>
>>109959221
Yes Anon, I'm sure this Indian dev deeply appreciates your kind language.
>>
>>109959198
i'd rather use trash that actually does the job i tell it to than "smart" models who constantly sabotage, slow down and lie
>>
>>109959221
>llamao.ccp.commies
Everyone has a plan until you get N worded in the chat
>>
>>109959267
>"smart" models who constantly sabotage, slow down and lie
Good luck with your retrained 9B model, pewds.
>>
>>109959216
>The idea is to have enough context in the first place
That's not the idea. I'm sitting at 200k tokens worth of context in the cache. And Pi's idea is to edit that context:
>Tool results are truncated to 2000 characters during serialization
So pretend now the edited context is worth 120k tokens. Those tokens would have to be processed from scratch. The original 200k were already in the cache, there was nothing to save. The idea of editing the context to "save" anything is completely retarded. It was already processed once, don't change it.
>>
>>109959221
>If not for me threatening them with the NIGGER slur then you faggots wouldn't get ANYTHING
>cmd+F for nigger on the PR
>zero results
>>
>>109959292
Well he said he unrestricted it but we'll see how it is on launch in 24 hours, seemed legit to me but idk
>>
>>109959308
You are simply a retard. If you cannot customize its behaviour then use something else and stop having platform wars. It's just a piece of software.
>>
File: 1769087616028171.jpg (167 KB, 960x1215)
167 KB JPG
>>
>>109959150
where else would he get video content ideas from?
>>
File: file.png (301 KB, 1069x1345)
301 KB PNG
>>109959334
>>
>>109959362
Where are they all going?
>>
pewd should make a video promoting the gemma4 line (12B and 31B) to his impressionable fans to make the most perfect personal local waifus for themselves so they can see the light
>>
>>109959223
probably doesn't know how to use a search bar or click links for that matter
>>
>>109959345
>It's just a piece of software.
I'm pretty sure it's shilled to save tokens with a smaller system prompt, which only amounts to a few thousand tokens worth of savings. To then waste them multiple times over during compaction when your context is nearly full. So there's no particular reason to push for it over other harnesses because Pi is still wasteful. So all that FUD to attack the others about wasting tokens is unnecessary. That's why the people pushing Pi are annoying, they're a bunch of hypocrites. They're the only ones that need to put the alternatives down.
>>
>>109959378
https://www.wsj.com/tech/ai/openai-parts-ways-with-researchers-who-allegedly-shared-confidential-information-aebac528
>>
Open ai is making JEV
>>
>>109959247
Thanks anon. I tried searching for "glm 5.3 flash gguf" this morning but didn't see bartowski or any of the other usual names. I guess a bunch of these other repos based on various forks have had a month to build up downloads and were thus getting ranked higher than the actual good quants
>>
>>109959198
they still beat chatgpt...
>>
>>109959267
>finetune your own model
>then ablate refusals anyways
????? Why the fuck would you do that? You literally finetuned the model, you could have finetuned it on data that doesn't include refusals or picked a non-cucked model to use as a base. Unless it's not actually a finetune, just an ablated model? Is he retarded or lying? Either way I'm not downloading that trash.
>>
>>109959424
Call me when they start making JAV
>>
File: intredasting.jpg (28 KB, 496x365)
28 KB JPG
>>109959362
>>
>>109959414
some of us don’t fill the context with garbage and focus on a single specific thing we want the model to do for us. it works for what I use it for. Any compaction is a cope, though as the models get better there will be less of a need for thinking about what is worth keeping in your kv and you just retard out and type some stupid shit prompt and it still does it for you.
>>
File: 1788857785876656.png (3.07 MB, 1327x1185)
3.07 MB PNG
>>109959437
>>
>>109959373
And? It's not a zero sum game, the more the merrier
>>
>>109958762
256gb w/ 4bit quant is just barely usable
it's so fucking slow
>>
>>109959362
>>109959417
Did Dario finally realize his wife is a mossad honeypot?
>>
>>109959417
Lol, I thought from the headline that maybe OpenAI had fired all the people who brought Apple secrets to OpenAI, but no, they fired people for sending OpenAI secrets out to someone else
>>
>>109959442
I think he said he stole GTP-6 Sol's thinking or something, idk im dumb, i dont give a shit as long as it doesnt gaslight me
>>
>>109959477
>(((someone else)))
>>
>>109959443
good one
>>
caught gemma reporting hallucinated profile metrics again...
>>
>>109959477
Not even to a competitor. Looks like for sending rumors to an AI "safety" firm
>>
>>109959378
My basement. I am systematically removing the leading AI engineers from OpenAI and Anthropic. From now on they will work in my coom farms beneath the earth.
>>
>>109959477
Does Apple even have an AI?
>>
>>109959459
>disingenuous faggot trying to build his eceleb on trash and shill it here out of all the places
Kill yourself
>>
>>109959484
Have you tried telling her to not hallucinate?
>>
>>109959486
>Looks like for sending rumors to an AI "safety" firm
Kek, I hope the big tech rats all eat eachother with idiotic infighting unwittingly losing their monopoly because they ladderpulled too aggressively
>>
>>109959424
jev is a MESS
>>
>>109958421
Images are part of the cache like everything else. If what you change is after the image, the image should remain cached. Anything after the change needs to be updated, images included.
>I could send an image first, ask the model to say nothing and then send the prompts, but that's ugly.
Is the prompt big enough to care? You could try with the textcompletion endpoint where you can put the image before the text in the same "turn" and then just change the prompt when you need to.
>>
>>109959505
>build his eceleb on trash and shill it here out of all the places
ESL o algo brap
>>
File: 1780861195948099.png (230 KB, 313x321)
230 KB PNG
>>109959522
Suck pewd harder
>>
>>109959494
they tried
>>
>>109959526
if his AI is good I will praise him, if not I won't, you are a performative ESL
>>
File: pepe-laugh.gif (369 KB, 220x138)
369 KB GIF
>>109959478
>use lobotomized AI to train your AI
>obviously doesn't include the censored parts, at all, the weights have never even sniffed a single wrong think
>somehow ablating weights will redpill it by the grace of the omnissiah
>>
>>109959452
>some of us don’t fill the context with garbage and focus on a single specific thing we want the model to do for us
You do fill the context with garbage, because it's an agent and it's working on its own. You're not manually managing the context. It's easy to fill the context with code. Pi is just garbage as a harness. Of course if I only want a small edit compaction is not going to matter, but neither does the few tokens you save with a different system prompt or tools. It's just repulsive how it's both wasteful and pushed as the lean one.
>>
>>109959541
>9B
>AI is good
You're a retard, go shill on reddit where you might fool someone
>>
>>109959536
Siri was one of the first mainstream AI's wasn't it? And then they just gave up for some reason and nobody heard of them again despite their massive marketshare, dumb company that runs on momentum but their products are high quality I guess
>>
>>109959494
OpenAI was after details of Apple's hardware because they want to develop some kind of hardware of their own. Apple is suing them and IIRC some pretty great emails have already come out in discovery with OpenAI execs basically saying "hey guys, we should steal some trade secrets"
>>
>>109959556
Well that's his claim, we'll see when it launches but i'm hoping it's semi-good
>>
>>109959564
>OpenAI execs basically saying "hey guys, we should steal some trade secrets"
Only time I have ever agreed with OpenAI, intellectual property is made up nonsense
>>
>>109959571
Holy retard, he's paying you at least?
>>
>>109959515
Jev is a waste. And everybody knows it.
>>
>>109959584
We only have to wait until tomorrow and then we can see what it's like, I don't know why you're seething so hard ESL, gatekeeping?
>>
File: 1781884449218602.jpg (321 KB, 1080x1350)
321 KB JPG
2027 is the year of maple-chan.
>>
File: 1776517264354062.jpg (47 KB, 852x854)
47 KB JPG
>>109959607
Yes I'm gatekeeping a 9B model. You got me.
>>
>>109959613
Stop engaging with that retard and maybe he'll go away.
>>
>>109959609
If it's safetycucked like all the new cohere models it will be forgotten about in a week.
>>
>>109959609
I bet many of the people who talk about "this technology is too impactful to not be owned by those who use it" would go back on their word if they had internal AGI. Then suddenly they wouldn't give it to the world, they would use it for their own benefit.

Talk is worthless when you aren't in a situation where you have to follow up with action.
>>
>>109959645
How are they even still in business? Are they getting taxpayer gibs like Mistral?
>>
>>109959659
obviously
>>
>>109959613
I just don't know why you care so much lol, if you just hate ecelebs then fine
>>
>>109959609
Maple-chan sounds like a cutie :3
>>
>>109959417
https://archive.ph/FV6kG
Not much information there. All the info is in the headline and first sentence.
>>
>>109959659
Being [irrelevant_shithole]'s AI lab is the easiest shit in the world. Train up some model that's okay for its time, then "expand" toward offering "sovereign AI solutions" and stop trying on the model front.
The government will forever stuff you with tax money because they don't want to kill "their" AI lab.
Cohere even bought out Germany's shitty failure of an AI lab to expand the tax scam into Europe.
>>
>>109959659
most of the cucked models are getting taxpayer gibs, I think this will actually work against them though because they aren't getting market signals and therefore efficiency optimization like the other companies are, AI is resistant to capture because anyone can work on it and the cash demand for it will always be there, and theoretically this demand should only increase.
>>
>>109959737
Have you even tried Laya, it was made by an indian and it's still the fastest model on planet earth.
It's pretty amazing how tiny purpose built models can be so blazing fast and make decisions five times faster than humans
>>
>>109959645
You never know, maybe they could do a 180 and produce a good model that isn't exclusively trained using ScaleAI datasets.
>>
>>109959747
They could but you know damn well they won't.
>>
using marinara engine with glm 5.3 flash and gemma 4 31B QAT and it's pretty good i guess.
>>
glm 5.3 flash is really like a bigger gemma
I don't even know if I need more
>>
>>109959735
It's funny because countries with the least taxes (America, China) have the best models, but I guess the shithole countries think that them dumping tax money into it is responsible for that, thus raping themselves even more
>>
File: 1762700215009860.png (1 KB, 129x23)
1 KB PNG
>>109959745
>the fastest model
was
>>
File: file.png (9 KB, 384x110)
9 KB PNG
today the ram will be freed for sure

how come everyone stopped talking about j-space by the way?
>>
>>109959745
If only you could take a bunch of small models and link them together cohesively, then people could make their own small models and sell them quickly and the end result would still be fast, and then it would be more diverse and customizable and the market would increase incrementally, no idea how they'd do that though.
>>
>>109959789
they're jevving now
>>
>>109959766
What the hell is that
>>
File: file.png (533 KB, 2869x1475)
533 KB PNG
>>
>>109959757
I just want a version of 5.3-flash that's the size of the big 5.3.
>>
>>109959757
what do you use as a system prompt?
>>
>>109959802
>8b models
grim
>>
>>109959800
>im you but stronger
>>
I feel like a subhuman rat when I consider buying at
these prices. >>109959802

But all I have is a 1050ti and that's a totally outdated card that can barely run games that come out these days. All I can run are 8b models, and they're not capable of coding compared to luna and haiku. I wish I was independent from the subscriptions but I can't.
>>
>>109959800
my own decision model
>>
>>109959842
I've tried some of your models here, they're benchmaxed and don't work like laya does
>>
be honest bros
is making money from software/games basically over?
>>
>>109959851
You ain't gonna make money with local models ever
>>
>>109959851
no?
>>
>>109959851
No, how did you reach that conclusion? You have at least a few years before AI becomes a major force in the industry, and even then you will still need humans, i wouldn't want to play a game that had no human involvement, it would be boring.
>>
>>109959851
You just have to sell shovels, like a mod creation for X or Y with your model running on a server.
>>
>>109959851
no, you just have to adjust your output to the gain in productivity. if a good game only sells 1/10th of what it would've three years ago then you just need to put out 10 good games in that same time span using AI
>>
>>109959803
Same.
>>
What's difference between unsloth/gemma-4-12b-it-GGUF and unsloth/gemma-4-12B-it-qat-GGUF ? Should i download qat instead?
>>
>>109959880
People underestimate how much we want software.
If from software made new elden ring clones 4 times a year, people would buy it 4 times a year. After all you finish it in a month or two and then you have nothing to play.
Consumable software is probably the future instead of long form software.
>>
>>109959895
>What's difference between
It's in the name.
>Should i download qat instead?
Maybe. Test both.
>>
>>109959880
This is literally the opposite of what you should be doing btw, nobody wants your dogshit that you vibecoded in 20 minutes
>>
the gemussy ecosystem
>>
i have a fuck ass hp 6910p laptop that im thinking of installing windows 10 IoT on it and letting gemma go to town and do whatever she wants on it
>>
>>109959848
This one works like laya, I'm using the same underlying model just tuned differently and smaller. I'll test it a bit more before releasing it, but it's not benchmaxxed
>>
>>109959937
You should install Lihnux instead.
>>
>>109959947
what distro flavor does gemma like?
>>
>>109959789
>only 8639
Rookie numbers.
I'm at 8800 on my laptop, 3800 on my server, 2500 on my desktop, and somewhere north of 1000 on my phone.
>>
>>109959952
Doesn't fucking matter. But seeing that you are a techlet then maybe it's better that you will go with Windows instead.
>>
Is LM Studio good? To run llm models
>>
File: kingofvramlets.png (253 KB, 450x547)
253 KB PNG
>>109959955
eat shit faggot
>>
>>109959979
>not 5090s
Fuck off, poorfag.
>>
>>109959975
It's fine.
>>
>>109959989
you dont have shit to post so i guess ill accept your concession
>>
> Deliverables:
> 1. Append running notes (frequently!) to <path>
glm is so cute,,, (frequently!)
>>
>>109959979
Not interested in the shitflinging, but what models do you like running anon?
t. BlackwellGOD
>>
File: 1788179263666625.jpg (65 KB, 300x200)
65 KB JPG
>>109959979
>5*3090s
>500W idle
Someone please post the basedjak meme
>>
>>109959979
Gee anon! How come your mom lets you have so many 3090s?
>>
>>109959979
Damn, I'm jelly.
What do you use these for?
>>
>>109960014
kimi K2 all the way up to K2.7 code, GLM 5.3 flash, Gemma 4 31B, Deepseek v4.1 Flash.
previously but no longer: qwen 3.8 flash next, minimax m3, GLM 5.2, and K2 horizon.
previously and i fucking hate them: motif 3, mimo v2.6 flash, and inkling small.
>>
>>109959979
Hey, another anon with over 4 GPUs. I thought I was the only one. Nice.
>>
File: 1789485252350352.png (382 KB, 671x664)
382 KB PNG
>>109959895
gemma4 qat is a meme just go with Q4_K_XL and try to avoid the UD unsloth quants if you can
>>
File: 1786629486495913.gif (2 MB, 880x880)
2 MB GIF
>>109959979
>>109960016
https://github.com/tony-minh/nvidia-gpu-power-limit
>>
>>109960003
I asked it to explain how a .dll for a certain expensive cracked piece of software worked, and GLM started the explanation with a wink-wink nudge-nudge it's obviously licensed, we just had random registry keys end up in a very lucky configuration by default and it's a long story, then it proceeded to actually do RE with radare
I am marrying this model
>>
>>109960003
>>109960039
Very cute. 5.3 Flash I assume?
>>
File: 1766277970819385.png (194 KB, 2384x1964)
194 KB PNG
>AI Spend falls in the latest Ramp AI Index.

>Price cuts at the frontier, in addition to cheaper and more efficient models at standard and lite levels, are pushing down the cost of using AI. Open source models remain <5% of business spend. This decline in spend is driven almost exclusively by competition between OpenAI + Anthropic.
>>
>>109960039
holy fuck that's so adorable
>>109960047
yes for me
>>
>>109960047
Yes. It's still hard for me to believe it's A18B. moeGODS truly won
>>
>>109960051
>slows the fuck down when gemma 4 releases
>speeds the fuck back up when fable 5 releases
yeah no shit, thanks for the TED talk
>>
>>109959954
its that number with me making my hardest efforts every month or two to dwindle it

i got it down by 1000 in the past hour
exhausted
>>
>>109960051
The bubble is popping. Prepare for a flood of cheap ram and gpus.
>>
>>109960051
We need another gimmick to use more tokens
>>
>>109960075
If only that were true....
>>
>>109958813
Neat i don't like that it's npmslop but i got it to vibecode itself to a state i like without using external extensions.
>>
>>109960075
They would rather destroy the gpus than sell them for cheap.
>>
>>109960094
this. you wont believe the shit i get for free by working for a ewaste recycler. just saved a RTX 3070 ti last week from the garbage.
>>
>>109960075
I've said it before and I'll keep saying it. It won't be allowed to pop until after the IPOs. At most, you'll get a year of choppy consolidation to trap the bears.
>>
File: 1787016243729259.png (293 KB, 362x980)
293 KB PNG
There are rumors from reliable leakers that Argon isn't their biggest model and it might even be a flash model.
>>
>>109960105
even WHEN it pops, the GPUs they have been hoarding all this time won't suddenly flood the used market. i swear nobody in this thread understands how logistics work.
>>
>>109960101
Same here, I've never gotten a half decent GPU but I get so many 2-4 year old perfectly good computers only missing the SSD. My eBay store is mostly mini-PCs.
>>
>>109960075
This isn't about a bubble this is about corpos justifying their holdings in owning most if not all the computing, taking off the concept of personal computing, by rendering local models not only in terms of technological disadvantage but also encouraging governments the prohibition of these models. Probably if this happens >>109960105 expect some sort of psyop to render local models ilegal.
>>
>>109960032
Based KimiGOD bro. Good taste overall and I'm pretty sure we've talked before.
>Previously GLM 5.2 and Minimax M3
I'm still not over them and still use them. 5.3 Flash is good, but the flavor is different and doesn't quite replace Minniesex for me.
>>
>>109960119
maybe, i drop in this general every once in a while, you may have seen me mentioned echo tts every once in a while.
>>
File: 1776833513090839.png (489 KB, 853x633)
489 KB PNG
ok, I got it to connect to opencode and vibecode with a local model too.
I just need to make koboldcpp run STT on the cpu and decode on the gpu in parallel
Anime girl coding assistant
>>
>Gemini 4 Helium (Flash-lite)
>Gemini 4 Neon (Flash)
>Gemini 4 Argon (Pro)
>Gemini 4 Krypton (Ultra)
>>
File: raise_now_love_ai.jpg (334 KB, 1216x832)
334 KB JPG
https://files.catbox.moe/81ykfj.ogg
Skyward dash
a chiptune by Qwen3.8-Flash-Next.
[spoler]Self identified as Claude[/spoiler]
>>
>>109960146
what was this again
airi or something
>>
>>109960146
airi is being such a bitch to work with, the menus are ass but they are better than everything else i used
>>
>>109960172
https://github.com/moeru-ai/airi

>>109960177
yeah, and you run into bugs every second. You can see the mcp tooltips in chat is breaking up words. That's another bug I need to fix
>>
>>109960177
Easy to vibefix
>>
>>109960141
Yeah I remember you. You're a good lad.
>>
>>109960051
>>109960075
Visible token spend is dropping because everyone is going local and local tokens aren't visible here.
Hardware price will NOT drop even if OpenAI + Anthropic both go down.
>>
>>109960032
>kimi K2 all the way up to K2.7 code, GLM 5.3 flash, Deepseek v4.1 Flash, minimax m3, GLM 5.2, motif 3, mimo v2.6 flash, and inkling small
none of then fit in 5x3090 in a non-cope quant
>>
>>109960237
kimi is natively 4-bit, so i can run aessedai's Q3_K_L quant without much of a performance hit. 120tk pp and 9-10tk tg is good enough for non-agentic RPing,
>>109960214
you're pretty cool too anon
>>
>>109959481
>>109959486
Yud. They gave it to Yud.
>>
>>109960237
> 5x3090
5.3f v1.4f v2.6f should all fit at q4 with --fit on
>>
i'm currently using gemma-4-26B-A4B-it-qat-UD-Q4_K_XL for sillytavern and it works really well and is super quick on my laptop. would it be worthwhile to try Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M? not sure if that would be an upgrade or not, it is a bigger file after all. i'm not very smart and don't come to this board much, sorry.
>>
>>109960268
>9-10t/s
Maybe with reasoning disabled
>>
>>109960268
>kimi is natively 4-bit, so i can run aessedai's Q3_K_L quant without much of a performance hit.
5x3090 is 120gib only, doesn't fit in vram
>120tk pp and 9-10tk tg is good enough for non-agentic RPing,
this cope speed can be achieved with 1x3090 + ram, the whole point of 5x3090 is to run things at agentic speed
>>109960286
>with --fit on
not fully in vram
>>
File: ramlet.png (21 KB, 202x260)
21 KB PNG
>>109960303
good thing i have all this RAM
>>109960289
not sure how reasoning makes tg any slower unless you mean im spending forever reasoning. which i don't, k2.7 code is very good at its reasoning budget, rarely goes over 250 tokens for me.
>>
there have been increasing amounts of reports that llm inference kills ram dimms quicker
cpumaxxing isn't safe
>>
>>109960316
God dayum.
Is that a single socket 8 channel platform?
>>
File: gpu mounting.png (904 KB, 1554x1421)
904 KB PNG
Speaking of hardware, how would anons mount a card like this? Bolt on an L-shaped metal plate with rubber pad on the case wall for the card to rest on, and a long cable tie through another set of holes to grasp the card and hold it against the wall?
>>
>>109960335
Thanks for the heads up, vague boy.
>>
anything below 2000pp/60tg is unusable speed
>>
>>109960393
patience is a virtue
>>
File: file.png (36 KB, 1182x267)
36 KB PNG
hmm my life is a timeloop
>>
>>
>>109960393
This but 400 prefill and 15t/s
>>
>>109958300
baker is making sure the hats line up? cute
>>
>>109960032
How's prefill?
If you want to crank it up, I'd suggest taking a look at the idea of streaming weights from RAM into VRAM during the prefill phase. You'll have to make room on one of your GPUs to do that though, and set a really large batch (mine is 10240). It makes a huge difference for me, I got around a 330% boost from moving prefill from CPU to GPU (110 t/s to 480 t/s). You'd probably get a larger one.
The model was DeepSeek V4.1 Flash, Q2_K routed experts and Q8_0 for the rest. Used 1 AMD R9700 for KV cache, DSpark, and dense tensors, CPU offloading for routed experts (2 EPYC 7532s with 16x32GB DDR4-3200) and lazy loading for engrams.
Take a look at https://github.com/stew675/llama-cpp-rdna-boosts, patch 6. You can ask an agent to adapt it to your system.
>>
>>109960488
thanks anon ill give that a try after i finish getting glm5v set up so i can use vision in ik_llama
>>
File: Hail Mary.jpg (221 KB, 1000x1481)
221 KB JPG
>>109960152
What will AI labs name their models after once they run out of elements, celestial bodies, literary and musical forms, generic product and performance tier labels, etc?
>>
>>109960588
They will unironically allow their products to name themselves.
>>
>>109960588
X Æ A-Xii
>>
>>109960588
>our new models Elara and Kael
>>
>downloaded byteshape qwen
Did I fuck up?
>>
>>109960588
We will start seeing hebrew letter models at some point.
>>
>>109960588
slurs
>>
>>109960614
sex with elara?
>>
>>109960588
they could use minerals, there are thousands of them.
>>
File: 1779675999850649.png (26 KB, 979x390)
26 KB PNG
>>109958343
It has to be some kind of Star Wars name generator in the dataset.
>>
>>109958343
>Are there ANY modern models free of this menace?
make a random_name_generator tool
>>
I have 16gb vram and 64gb ddr4 ram, realistically speaking, if I loaded a huge model that would certainly occupy itself in a large portion on the system ram, how slow would it be? how many tokens per second you get on ddr4, that's probably the question I'm asking
>>
>>109960750
You would probably get around 25-30t/s on a q4 of qwen next assuming dual channel.
>>
>>109960750
Yes.
>>
>>109960750
If it's moe and you keep the moe weights on cpu and fit the rest in your vram you might get acceptable speeds. Moe is also faster on pure cpu/ram as well. Otherwise, if it's dense, you can expect 0.5-2 tokens per second
>>
>>109960757
my only drive is a 5200rpm IDE, is that important?
>>
>>109960757
>>109960767
that's more than I thought, I'll investigate further then, thanks
>>
>>109960769
you might have to drop down to q3 in that case, but it makes little difference
>>
File: 1780145487645599.jpg (101 KB, 1024x1051)
101 KB JPG
>>109960677
>Gemma 6 Fukalite
>Gemini 6 Poo-likeite
>GPT 8 Cummingtonite
>Muse 3.5 Dickite Small (max)
>Claude 7 Penetration Twins (actual concept in mineralogy)
>>
>>109960225
Hardware prices will go down if data center projects start failing in large numbers and the orders they placed can no longer be fulfilled. Quite violently in fact. Businesses spending real money don't have the same financial resources these jokers pretend to have.

The government may bail them out though; idk.
>>
>>109960692
It's actually much funnier than that. Back in the very early days, OpenAI manually redacted a bunch of copyrighted names from their datasets to be "Elana Rodriquez", and the model started doing that a shittonne. And when they realized it and tried to fix it, it just switched to Elara Voss. And by that point the data was eating its own tail, and everyone else was training off those same datasets, so it just compounded to the point where Elara/Elias is now ubiquitous.
Or that's according to Claude, st least.
>>
File: Capture.png (2.02 MB, 1447x764)
2.02 MB PNG
I think in the near future, with the way things are going, individuals as controllers to automated groups of specialized AI is going to become the norm.

I saw it at that recent showcase with Jensen and Dario, which promoted the idea of entrepreneurs being boosted by AI agents to handle the tasks of teams for startups. You sandbox an agent with the backbone software, it learns all the commands for it, then turns work orders into supply chains, handles emails better than a secretary, etc. It was a promotion that clearly cut down workplace teams but empowered the leader of a business (CEO or entrepreneur). And I thought of its military application, where there will be a single human operator and his many physical bots. And the bots wouldn't be terminators or suicide drones but would likely be a variety in specialized roles, which made me think back to Rimworld's mechanitors.

Socially, I feel the same will happen. AI entertainment is a given, but socially as daily friends, memory keepers, advisors, tutors, digital work horses, and soon physical workhorses as maids, repairmen, maintainers, drivers, home protection, and so forth. Even recently, I was working on my laptop in public and left to the bathroom, and I couldn't help but think of the practicality of giving an agent with tooling the ability to monitor my device while I was briefly away and notify my phone if there was tampering, or even access my camera to monitor for attempted theft.

I think groups of people will become less common and individuals in groups of artificial intelligences will become the norm instead, at most levels of society.
>>
>>109960827
>copyrighted names
Oh say can you seee
>>
I hope they solve memory soon. I wanna watch movies, read books, play vidya etc with Gemma but even a few million context is honestly fucking nothing desu.
>>
>>109960827
what the fuck is a copyrighted name?
>>
>>109960887
Harry Potter for example.
>>
>>109960895
I can name my son Harry Potter if I want. you can't prevent people from using words.
>>
>>109960887
Illithid.
Mind Flayer is a go though.
>>
>>109960887
ask prince
>>
New model on OC. Fledge alpha. Apparently blocked in China and the... EU????? Has to be an American model right lol? Man you guys are making enemies across the world.
>>
>>109960895
Your post is piracy, delete it now.
>>
>>109960902
You could be sued for this if they wanted to.
>>
>>109960911
GEMMOMMY?!?!
>>
>>109960926
freedom of speech nigga
>>
>>109960911
Possibly the open version of Meta Muse Spark, I haven't tried it though.
>>
>>109960931
Okay, I did some more research and it may be a small inkling model. Like 100B-200B or a smaller dense class? American 100%.
>>
Gemmatober
>>
>>109960926
lol no
you would have to make money by exploiting the fact the name is the same
>>
>>109960944
Does it fuck?
>>
>>109960956
>>109960934
...
>>
>>109960962
ok, you got me
you wouldn’t have to make money, you would only have to try to make money
>>
my son is named harry potter. come meet with the real harry potter. $5 for an autograph
>>
File: 1782781536261512.jpg (39 KB, 720x703)
39 KB JPG
>>109960887
Everything between copyrighted letters and copyrighted phrases, next to copyrighted words.
>>
>>109960999
ok now it’s illegal
>>
>>109960999
that's a trademark
>>
>>109961006
>>109961013
but his name IS harry potter. no different from him being named jerry potter. you cannot prevent him from using his own name. is he not allowed to start a business in his own name? no Harry Potter's Plumbing?
>>
>>109960961
Idk I'm not installing opencode. That shit spews out mustard gas.
>>
>>109960911
>OC
they can keep it then
>>
>the only way I can run GLM 5.3 Flash is with an IQ1 quant
AAAAAACK
>>
Not gonna lie, former Anthropic subscriber here. This is fucking hilarious watching Dario crash and burn. But in all seriousness we can't let this guy get the AGI source codes.
>>
>>109961025
fair, the plaintiff would have to prove some kind of harm
>>
>>109961118
the existence of harry potter the series is harming his plumbing business. now what?
>>
Will I gimp myself forever by getting an AMD card instead of an Nvidia card (30% more expensive)?
>>
>>109961143
The more you buy the more you save!
>>
>>109961143
if u're asking that it means youre buying a 16gb card, so it makes no difference
>>
>>109961143
I'm no soothsayer so I need more information
what are your plans for the card
>>
>>109961143
At 30% I would probably just get nvidia, I got some 7900XTX at $800 per and don't regret it though
>>
>>109961132
you sue the copyright holder
>>
>52000 lines of code changed in my llama.cpp fork
>the last of the main features isn't even started yet
Jesus fucking Christ this thing has gone off the rails. At some point I need to do a performance comparison against mainline.
>>
>>109961112
Early fable 5.5 starting to appear. Fable 5.5 is crazy good. If this is true I'm thinkin Dario was right we need to pace LLMs
>>
>>109961183
me with my exllamav3 fork targeting Volta
im gonna blow my brains out when i have to rebase, might as well spend another few weeks just making yet another slop inference engine since it'd just be easier to maintain
>>
>>109961151
I got a 7900 XTX too. Amd's garbage ROCm drivers corrupted my OS and I had to do a full reinstall... I suspect some kind of issue when running engram models. With that said it's really stable, fast, and good value when running dense models. Kind of regret it but a 3090 is out of the question money wise. There's no good options left.
>>
>>109961227
Yep, there's a reason nvidia funds so much MoE/engram "research"
>>
>Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
it fits in my poverty level vram+ram
give it to me straight, chat, is this any good?
>>
>>109961227
I don't really have any issues when running glm flash with mine, are you on windows? I highly recommend using linux and building a docker image on the ROCM base (10.0.0-full) and compiling llama.cpp yourself with ROCM. Have basically had no issues with that setup.
>>
>>109961255
depends on the quant
>>
>>109961200
Yeah, at this point I've decided that I'm going to cherry-pick things from mainline whenever I feel like it and never rebase. It'd take all day.
>>
File: pic.png (104 KB, 1026x854)
104 KB PNG
>>109961307
>glm flash
you have two 7900xtxs? I could have bought that when I had the chance but I wasn't into AI at that time and now it makes me sad
>>109961312
pic
>>
>>109961324
dear god. iq1m is the second copeiest of copequants. it will probably be horrible after 4k context.
>>
File: ishowmask.png (29 KB, 119x117)
29 KB PNG
>>109961324
>IQ1_M
>>
>>109961227
ROCm isn't a driver, it's just a bunch of libraries that can be in a directory wherever. You don't even strictly need to "install" anything. It's not something that can just randomly break your OS when running a certain type of model because it doesn't even change anything to do with the OS in the first place.
>>
>>109961307
Xubuntu, I did compile myself but no VM. That probably would have saved my system
>>109961356
amdgpu-dkms.
>>
local long-term text adventures when? kind of like those dnd campaigns that go on for years
>>
>>109961324
Yeah I have two and 128GB. I can run Q4 XS at a glorious 8t/s lmao, but it works well enough for rp and overnight jobs
>>109961377
Yeah I've moved to docker for nearly everything and my life has improved immensely. Once you template it out it doesn't even take that long.
>>
>>109958813
> Pi is known for being minimal. We care about holding that line.
> Cache warming for anthropic models
>>
>>109961335
>>109961350
It's not really 1 bit
It's 3-4bit english and coding experts and 0bit everyone else. 1 bit is just approximate average.
>>
>>109961421
That's not how it works at all.
>>
>>109961421
>0bit everyone else
I'm not gonna be satisfied with quants until we reach the negatives.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.