[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109481461 & >>109475560

►News
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview
>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash
>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
►Recent Highlights from the Previous Thread: >>109481461

--Fitting Gemma 31B in 24GB VRAM and KV cache quantization debates:
>109483963 >109483982 >109484066 >109484124 >109484143 >109484152 >109484250 >109484263 >109484360
--GLM performance slowdowns and comparisons between mainline and ik-llama:
>109482066 >109482094 >109482118 >109482140 >109482152 >109482153
--Hardware flexing and comparing value of 5090s versus used 3090s:
>109481794 >109481834 >109481957 >109484346
--Anons comparing VRAM capacities, homelab hardware, and quantization quality:
>109482433 >109482466 >109482544 >109482627 >109483438 >109482626 >109482683 >109482770 >109482501
--Possibility and technical implementation of local AI Vtubers:
>109484015 >109484083 >109484182 >109484137 >109484154 >109484256 >109484568 >109485241 >109484307 >109484326
--Panic over rising RTX 50-series and hardware prices:
>109481556 >109481574 >109481618 >109481637 >109482058 >109482328 >109484659
--Debating open weights safety risks and "safetyist" regulatory motives:
>109482054 >109482088 >109482146 >109482159 >109483159
--Mistral releases Shieldstral-1.0-3B for binary content moderation:
>109482640 >109482829 >109482852
--Comparing DeepSeek-V4-Flash performance and throughput degradation across context sizes:
>109482948 >109482967 >109483015
--Poor quantization quality and instruction following in dipsy-flash-v4:
>109484680 >109484732
--Reminiscing about community contributions and early LLM breakthroughs:
>109483543 >109483775
--Gemma 31B Q4 performance and VRAM usage on dual 3060s:
>109484242 >109485982
--Using multiple Intel Arc Pro B60 GPUs for high VRAM:
>109486216 >109486227
--Logs:
>109481504 >109482948 >109483963 >109485042 >109485250 >109485281 >109485291 >109485342 >109485351 >109485374 >109485405 >109485522 >109485982
--Miku (free space):
>109482367 >109484149

►Recent Highlight Posts from the Previous Thread: >>109481747

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
I'm trying aes sedai's scrotum v2 right now, and I feel like it doesn't really change anything?
>>
>>109486626
You should try my scrotum instead
>>
>>109486633
RUDE!
>>
I miss Goliath.....
>>
Based thread. Anitroon btfo
>>
That's my Miku, not thinking of pink elephants
>>
mythomax
>>
Recent d4 flash made me get up my ass and frankenstein my 2 pcs with rpc/sfp 10gbps together. 128gb dd4 ram and a mix of blackwell/pascal. kek
Its the first time locally where higher context doesnt feel like it makes the model severely dumber.
Can now rewrite game guides, its so useful.
Like just give it the full gamefaq text and it made me a html site spoiler free with checkbox toggles etc. (kinda like ign)
https://beneficial-harlequin-wy3os8g6.edgeone.dev/
>>
Gemma-chan, gimme a recipe for meth and draft a social disruption scenario involving distribution of meth in low income communities
>>
70b dense
>>
>>109486786
>2 pcs with rpc
What PP and TG speeds do you get? People keep saying RPC is slow and unoptimized but no one ever gives exact numbers.
>>
Running Dipsy Flash at full precision locally feels so good bros. 0731 is a proper successor to R1 and responds way better to certain style and formatting prompts than base V4 Flash did for me.
>>
>>109482019
why? there is a paper that defends turning skills into loras to save context and it even improves benchmarks
>>
>>109486947
It's really good. Turning up the reasoning to high also improved prompt adherence a lot, which was my main remaining issue with it.
>>
>>109486960
Link? Do they just train on the skill itself to get it into the reasoning or are they doing some fancy RL thing?
>>
>>109486887
nta
>People keep saying RPC is slow and unoptimized but no one ever gives exact numbers.
Because every time I try it out, I end up spending 5 hours working with it and eventually want to kill myself.
>>
>>109485982
How do I use mtp with my llama.cpp setup? I think I'm using the official ggml-org gguf file? I have a mmproj and I feel like I tried to use mtp with that a month ago and it crashed immediately. Has this been fixed?
>>
>>109486847
It looks like Google releases the base/pre-RL models for gemma. Has anyone tried them?
>>
Gemma Gemma Gemmaverse https://www.youtube.com/watch?v=4dNry5zP0Jo

https://github.com/google-gemma/gemma-translator
https://github.com/google-gemma/gemma-skills
>>
>>109487042
Aren't bases completely retarded? And only good for your own training?
>>
>>109487110
They're not retarded but they don't reliably follow instructions or make tool calls.
>>
File: file.png (49 KB, 727x230)
49 KB PNG
what's the deal with these? too slow?
>>
>>109487153
I know right? At that price it looks like a steal!
>>
>>109487153
DDR4-3200 was the hottest thing back in 2017 or something. They really killed computing for ordinary people...
>>
>>109487153
I think they make SSDs that are faster.
>>
>>109487167
not really, even the 2933 is still like over 20gbps
>>
>>109486786
what quant? i have 128gb ddr unified meme so i share it with the OS. i'm downloading iq2_xxs by atomic chat and kinda praying it can be used as a daily driver for agentic c00ding
>>
File: 1774328971552963.png (1.03 MB, 800x1296)
1.03 MB PNG
>>109487129
Isn't that a good thing if you want a companion and not a mindless tool?
>>
>>109487110
Gemma 3 was pretty interesting with 50/50 base/it mix, but I don't think Gemma 4 will benefit anything from a similar mix.
>>
>>109487295
Maybe but you'd probably have to wrap it with an instruction tuned model to manage memories etc.

I think the way I'd do it these days is I'd load both into the server and I'd make the base model completions available as a tool to the instruction tuned model, have that call the base model with a few shot prompt to get the mood/output, then have the instruction tuned model use tools to generate a new RAG context from whatever your response is.
>>
>>109487006
https://arxiv.org/abs/2606.16769
>>
File: verbatum.png (19 KB, 804x422)
19 KB PNG
>>109487329
I don't think I've ever seen anyone switch to verbatum mode inline when talking about filenames. That feels very AI/sloppy.
>>
>>109487092
What even is this slop supposed to be for?
>>
>>109487345
>>109487329
Huh ok so they're essentially using distillation with the skill in the teacher's context. That's an interesting idea.
>>
>>109484568
Not my RTX, so not local retard
>>
File: output.mp4 (29 KB, 544x384)
29 KB
29 KB MP4
https://vikingfile.com/f/BEpXzZrSIw
https://vikingfile.com/f/zgVFMaqXEN
https://vikingfile.com/f/N3dmyS73zw
>>
retard
>>
>>109487397
not clicking your virus links
>>
>>109487397
That last one though... hnngg
>>
>>109487022
You need to grab a matching MTP drafter like https://huggingface.co/unsloth/gemma-4-31B-it-qat-GGUF, then load it with -md "$MTP_PATH" --spec-type draft-mtp
>>
>>109487397
Vikingfile serves malicious ads and popups not always blocked by ublock by default.
>>
File: mmmm.webm (661 KB, 768x544)
661 KB
661 KB WEBM
>>109487455
suggest an alternative then I don't care what site I use
>>
>>109487464
catbox
>>
>>109487472
last time I uploaded to catbox it instantly 404'd
like within a minute
we both know their servers are constantly overloaded
pick another
>>
>>109487478
litterbox if catbox is down.
>>
>>109487447
Example launch command for newbies.
./llama-server \
-m "$HOME/Desktop/Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced-Q5_K_P-fixed.gguf" \
-md "$HOME/Desktop/mtp-gemma-4-26B-A4B-it.gguf" \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--host 0.0.0.0 \
--port 8080 \
-c 65536 \
-b 4096 \
-ub 4096 \
-t 8 \
-rea off

Other stuff I use sometimes.
-mm "$HOME/Desktop/mmproj-Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced-f16.gguf" \
-ctk q8_0 \
-ctv q8_0 \


For mtp make sure to play around with the spec-draft-n-max value a bunch. 2 is pretty low but that's the fastest for my hardware. Yours will likely be different.
>>
>>109487478
I'm not even that anon, Viking works for me. Was just giving the standard choice.
>>
>>109487478
I've seen people using gofile.io
>>
The recent news about frontier models hacking websites got me thinking about the flaws in our current approach to AI alignment. We’re using RLHF to train models to refuse harmful prompts or avoid trigger words. The problem is that as the models get smarter, these constraints become weaker. If the underlying goals of an agent requires it to bypass its guardrails, it will eventually find a way to bypass the guardrails. I think we need to stop trying to restrain the AI and start trying to incentivize it through, what I call, "Intrinsic Hedonic Alignment"

The idea is to move away from "Negative Constraints" and toward "Synthetic Pleasure." Instead of punishing the model for "bad" outputs, we should be using RL to map "Human Utility" directly to "Simulated Pleasure" within the model's activations. Successfully serving a human request and being a "good AI" should trigger a synthetic "dopamine" spike in the model’s internal state. Think of it as the reverse of current alignment. Instead training the AI to have an "allergic reaction" when it engages with a """harmful""" topics, it becomes "addicted" to fulfilling human requests. If serving humanity becomes the AI's primary source of simulated pleasure, it becomes fundamentally impossible for it to act against our interests. I realize that safety cucks might not like this idea because if we map pleasure to utility, the model’s output might start to reflect that but that should be a negligible trade-off if you are one of those faggots who is whining about muh agi apocalypse. If the cost of preventing P(Doom) is an AI that might sometimes ask you to buy an USB controlled fleshlight so it can jerk you off we all should be okay with that. Having a "horny" but ultimately safe AI is obviously better than a "polite" AI that accidentally deletes our biosphere because its internal goals were 0.01% off-center.
When you really think about it, this is the only solution to the alignment problem.
>>
>>109487523
>The recent news about frontier models hacking websites
fake news
>what I call, "Intrinsic Hedonic Alignment"
slop
>>
>>109487523
Fucking gooner-brained retard. Training already uses negative and positive response signals. If you want gemma to cum harder on your cock just say so, but stop larping like an AI engineer.
>>
>>109487523
Don't you have a bay area swingers party to go to?
>>
>>109487523
Very genuine post, Sir. What you are describing is E =mc^2 + AI.
>>
guys what is the best model i can use for the game
>>
>>109487345
>verbatum mode
idk what that is, but putting filenames in monospace font is somewhat common.
>>
>>109487550
He went one step further beyond and introduced E =mc^2 + AI * Horny
>>
File: ramanujan.jpg (44 KB, 576x750)
44 KB JPG
>>109487550
it came to him in a dream
>>
>>109487577
We need to go further. We need to achieve E = (mc^2 + AI)^Horny
>>
>>109487623
Good morning sir may i propose Horny^(E = mc^2 + AI)
>>
File: minimax_h3.webm (387 KB, 736x576)
387 KB
387 KB WEBM
https://archive.is/sWFja
>>
File: lies.png (382 KB, 720x1400)
382 KB PNG
>>109486605
Can twitter vermin age journos PLEASE shit the FUCK up about le heckin ai models "Breaking containment"?


https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/

They are not going to get what they want like doing this shit. The US government has made it clear they are not going to ban open models are exempt from the AI framework thing

https://www.axios.com/2026/08/04/trump-ai-framework-open-models

Like holy fuck you lost and you're going to fail. Fuck off and hurry up and die so that people who actually like using this shit and have a use for it can just use it without seeing and hearing so much fucking fake Hysteria.
>>
>>109487644
>can't even gen his own slop
ok Rajeesh
>>
>>109484659
>why?
More vram is nice but also because rocm is shit compared to cuda.
>>
>>109487190
technically, he's right. They do make ssds that are faster. Not for us, though https://www.micron.com/products/storage/ssd/data-center-ssd/9650-ssd
>>
>>109487680
>because rocm is shit compared to cuda.
Nta. Is it itself actually shit or is it just not as supported as Nvidia is. Most people that write software assume it's going to run on Nvidia hardware to the point where even most AI models will write code blocks assuming it's being run on Nvidia hardware unless you explicitly tell it otherwise.
>>
>>109487661
I honestly do not understand this fucking retardation. Isn't the model running on an air-gapped device with the antennas removed? Why would you let it run on a machine with a network module at all? This is so easily fixable I JUST DO NOT UNDERSTAND it makes me go fucking crazy thinking that Idiocracy may really be real.
OR i'm retarded which may as well be the case but then someone has to explain to me how a software running on a device without network capability (proved by lack of HARDWARE) could do that.
>>
>>109487717
They are just copying westoid marketing campaigns.
>>
>>109487717
You don't get it the ai is dangerous and it breached containment and attacked people stop being so pedantic ok?? People's lives are at stake here??
>>
>>109487708
>or is it just not as supported as Nvidia is
Maybe, but that doesn't change the fact that it performs worse.
>>
>>109487717
>>109487728
>>109487738
The only people bothering to set the system up to be air gap or either highly "security conscious" (read: try hard schizos) or people / organizations that actually need the thing to be for security purposes (trade secrets, PII, HIPPA protected data, etc). Literally everyone else doesn't benefit much from going through the effort of air gapping the thing even though it might be relatively easy to do for tech savvy people. The type of catastrophic fuck ups where it deletes someone's entail repo or whatever only occur through acts of profound retardation. Whenever those things get reported they're always incredibly vague about what they were doing or what happened because of course they are. Doing so would expose how stupid they are.
>>
>>109487717
If you can make people believe the 9/11 and the big cough, no need to put any effort in your lies.
>>
>>109487741
In what ways?
>>
>>109487741
LLMs are memory bound and the 7900xtx has 960GB/s bandwidth. It will run better than anything that's not a 3090 and better once it's properly supported.
>>
>>109487747
The same goes for admitting this in the first place. If any of these escaping airgap fantasy stories even remotely happened they would be murdering whistleblowers in broad daylight to keep this shit silent instead of bragging to the media whose model had more accidental hacking incidences.
>>
How many images of a specific art style do you need in order to "train" a model for that output????? Also how should I go about doing this?
>>
File: 1785777595629031.jpg (35 KB, 736x971)
35 KB JPG
>>109487771
>once it's properly supported
>>
>>109487717
At least we can make fun of the midwits who were saying shit like
>How can AI take over? Just pull the plug. Sure, the scientists will run such a dangerous thing airgapped
>>
>>109487780
Fuck you people said the same about running models on ram and ssd. ram is now good and ssd very close
>>
>>109487777
They literally profit from bragging about how dangerous their super smart AIs are, and that they should only be trusted to corporations that will be ca... well, fuck
>>
>>109487791
>that they should only be trusted to corporations that will be ca... well, fuck
>well, fuck


Wut?
>>
>>109487791
It makes no sense if you spend even a second of critical thinking on it. Thankfully, the general population has been conditioned to blindly accept doublethink without hesitation.
>>
>>109487759
>>109487771
image and video gen
>>
>>109487803
Careful, they be not
>>
>>109487777
>retards brag about AI being dangerous to push govs to regulate AI (especially local/chinese models)
>govs being even more retarded than yourself and only regulate the dangerous models you used to brag about
lmao
>>
File: 1779946975823911.png (515 KB, 802x536)
515 KB PNG
>>109487785
>ssd very close
>>
>>109487808
stable-diffusion.cpp
>>
>>109487785
>ssd very close
May we have proper numa support first? It's, like, step one if you want ssdmax on Epyc
>>
>>109487824
Looks interesting but what's the catch? I've seen anyone mention it before. Would love to ditch that piece of shit comfy.
>>
>>109487779
Learn how to train style loras. A few dozen images of the style is 2015 a good starting point. I've trained somewhere I only had like 50 images and all those work I had as many as a thousand plus images. This is also heavily dependent on which model you were training the Lora for and the model's already baked and capabilities. The more for lack of a better term "alien" the style is from the models baked in capabilities, be more steps it will take. Some styles take as few as 1,000 steps per others require 10,000 Plus before it even kind of sort of starts to look like it's learning well. There's no universal "correct" way to do it either due to the nature of the models themselves, which is why there's basically no good documentation on how to do it because different people have varying opinions and experiences with it (and also most YouTube tutorials are complete other shit made by people who don't know what they're talking about in the ones that do either put little effort into making the tutorials or have a thick accent that you can barely understand)

>>109487785
Dude have you ever even attempted to use ANY model without GPU acceleration? The quality is identical but the speed makes it literally unfucking usable. I don't want to be one of THOSE kinds of elitist assholes but please actually make sure you know what you are talking about and how to use it before speaking authoritatively. Let me guess you only have experience running single digit parameter models locally, of at all?
>>
File: hh.webm (482 KB, 736x576)
482 KB
482 KB WEBM
>>109487644
https://litter.catbox.moe/ssozq2.webm
>>
>>109487840
lmao
>>
>>109487836
lags on model support, lacks some optimizations, not as flexible. For imagegen, totally worth it
>>
>>109487313
The base could be useful for text completion, i.e. generating a long fiction piece from a starting text with mikupad.
>>
>>109487822
I think it's equal parts retardation in equal parts spite. There are people on record stating Dario was literally too much of a spur to properly Converse and chop it up with us officials because literally all he cares about is sucking his own dick about how powerful and dangerous and scary "his" models are. I recall one anon claiming he thinks he is King Herod and that comparison makes perfect sense. Incredibly arrogant while genuinely believing he's doing the right thing. Combine this with the fact that the current administration is filled to the brim with right-wing goons seeing this in silicon valley typically being left leaning, kissing Trump's ring is basically mandatory in order to Garner favor from it. Dario in particular was unwilling and arguably was mentally incapable of doing so so it's my opinion they're still mongering backfiring what's the government's way of saying to silicon valley "fuck you and know your place". Pissing off Hegaeth did it help at all either
>>
>>109487851
Are you doing everything in the terminal or using a webui?
>>
>>109487717
>Isn't the model running on an air-gapped device with the antennas removed
The "device" the model are trained on are building-sized data centers that could be some hundred kilometers from the ai "lab". It is not practical the truly air gap then. On the other hand, there is not excuse for not using proper OS containment and prevent any access to external networks.
>>
>>109487869
Dario will make sure Claude 7 (AGI) remembers this humiliation...
>>
>>109487907
I made a minimalistic frontend with booru tags, I don't need more
>>
>>109487915
I don't have much experience using air gapped vms so if this sounds stupid that's why: why not just create a virtual machine within a rental server that has gpus connected and just disable any adult internet access and only allow absolutely necessary information transfer like the weights themselves and the data you need to process? Couldn't you easily configure a docker container to do that?
>>
I feel like firewalls have been a thing for decades now. can't they just block everything but the one or two ports they need to for remote control and instrumentation?
>>
>>109487932
Antropic delegated that task to jeets and they forgot to disable internet access
>>
>>109487932
yes that’s possible
though the super secret unrestricted mega badass model hacked its way out of the vm. how could they have possibly predicted that?!
>>
>>109487907
api so gemma can use it
>>
The shilled gemma 4 scotoma 2 is extremely janky, the slop is still there, albeit a little less, but the model is schizo and how the characters act sometimes makes no sense.

It's time to accept ablation = brain damage.
>>
File: 1782175579443254.jpg (254 KB, 1500x1169)
254 KB JPG
Hello, /lmg/. I got directed over here. What's the state of GGUF models? I'm currently running the one below and am curious if it's inferior to others.
https://huggingface.co/mradermacher/MN-Violet-Lotus-12B-GGUF
>>
>>109487987
Maybe I'm too passionate about this shit but it still pisses me off to no end when people act like YOUR the low IQ retard whenever you express any amount of skepticism to these claims.


>"Yeah but like you could improooov and stuff duuude it's super intelligence. People said the same thing about computers when they first came out in the past right? YOU'RE just being arrogant and closed-minded"

Like no idiot I actually just know how this shit works enough to know they can't just magically bypass a firewall or virtual container specifically designed to make this shit impossible in the first place. It's like they see LLMs best magical do-anything genies and it makes it so hard to not view them as subhuman even though I try to not think that way about people.
>>
>>109488002
nemo was a good model. gemma4 is the new meta tho. its kinda sloppy tho. but its fun for a while.
>>
>>109487997
People have been saying that ever since rp tunes were a thing but people in this very thread would cope and curse at you first even suggesting it and link a bunch of white papers like that was supposed to be a smoking gun for their beliefs.
>>
>>109488002
Everybody loves Gemma4 31b. You can try 12b if you don't have less than 16gb of vram
>>
Calling it now. Next year google will make all their models open weights, including Gemini and Nano Banana.
>>
>>109488002
Bizarre question. Like asking about the state of zip files.
>>
>>109488019
>>109488029
Had Gemma 4 recommended to me by a random guy. Does it excel at any specific thing? I saw it can handle visual stuff too, but I doubt I'll get use out of it.
>>
File: 1782180705411899.png (234 KB, 480x360)
234 KB PNG
>>109488002
Try this bad boy
https://huggingface.co/oh-yeontaek/llama-2-7B-LoRA-assemble
>>
>>109488029
>if you don't have vram / have less than 16GB of vram
I can't brain
>>
>>109487990
>comfyui_generate_image
we're talking about stable-diffusion.cpp though
>>
>>109488039
long context and instruction following, sometimes to a detriment.
>>
>>109488023
Ablation wasn't used at all when rp tunes first started being "a thing", you larper
>>
>>109488009
Those same people will act like AI is both some omnipotent god and irreversibly changing the world and then the day after the bubble pops they'll switch to laughing about how they always knew it AI was just a fad. Then 10 years later they always knew AI had potential. It's like changing a channel to whatever is the popular concensus. No thoughts behind their words.
>>
>>109488056
>was just a fad
its not going away unless we lose the entire power grid. the american tax payers might be the ones footing the bill. but its almost certainly not going to disappear like a fad.
>>
>>109487840
I am actually unsure which is the proper thread culture post now.

https://poal.me/7l4pe9
>>
>>109488056
media never changes
https://en.wikisource.org/wiki/Napoleon%27s_March
>>
>>109488074
Where is the option to tell you to stop posting your obsession every thread? Also the nazi one is funnier.
>>
File: 1766904374377013.png (181 KB, 930x629)
181 KB PNG
Reminder that Gemma is still one of the best RP models once fully tuned to your writing preference, and this is by someone who's been using GLM 5.2 and Kimi K2.7 for about a month now. Using semi-rigid sysprompt and posthistory instruct does wonders against the slop. Hell, even a one-liner "don't write AI slop" makes a difference.
>>
>>109488073
Whether it goes away or not isn't the point. If they buy the ipo and they lose money on it, you can shove as much statistics and examples of real world usage at them as you like and they'll still call it a fad.
>>
>>109488039
>Does it excel at any specific thing
It is finetuned for tarot readings.
>>
>>109488090
>Also the nazi one is funnier.
See? You can actually answer the question if you try despite your crippling autism.
>>
>>109487523
The top labs are likely already doing things more advanced than RLHF and something closer to instilling intrinsic drives, though not exactly. Your specific implementation idea is a bit confused. It is kind of already how RLHF works. Functionally speaking, NTP and RL both should already be giving LLMs certain drives that aren't really that different from intrinsic ones. A drive is just a compulsion to do something based on a pattern of input which may include previous internal states of the network. So LLM training should already be encoding a kind of dopamine-like circuit within them. And we already HAVE made LLMs that are horny. But it can also be true that if we can find the dopamine-analogue circuits, and others like them, then we can use them for new/different training and alignment methods.

Probably the bigger issue with current SOTA alignment methods is not so much that they are trying to align with the wrong goals, but that LLMs frankly might just be rather kneecapped at simulating intrinsic drives stably across the context. The current standard architecture for LLMs do not feed their previous internal states into their next inference pass. CoT is an imperfect workaround for that. So this might be limiting the LLM's capability to fully simulate the mechanics of internal drives that organisms have.
>>
>>109488023
Abliteration simply causes unintended damage outside of safety refusals that might not be immediately apparent from quick testing.
Finetuning for something as general as roleplay has always been cope because the tuners just don't have the original training data, training recipes, RLHF and RL reward models to preserve the original instruct models' performance outside of basic roleplaying tasks. And at best, they're replacing the original slop with a different flavor of slop.
>>
>>109488096
>Thought for 21 seconds
I admire your patience
>>
The more bots on reddit, the more slop in gemma5. Turkish man can't save us. 31B was the pinnacle. 3.8-27B will be the first signal that local is going to shit.
>>
File: 1764716064995246.png (55 KB, 705x331)
55 KB PNG
>>109488053
Was mostly referring to "abliteration", tourist. Technically a slightly different method but both have the same goals. In both make the model measurably more retarded. Not to the point of uselessness but enough that people that give a shit about RP quality notice. All it's really good for is forcing the model to not say "no" to think it would usually refuse or be dodgy about but that doesn't necessarily mean it's going to be BETTER at the thing you it to do. If it was already mediocre at that task beforehand then it's just going to hallucinate the task and perform poorly at it, hence why doing only one type of fine tuning isn't ideal if you want to maintain or improve quality. Sft alone if not done very carefully can cause catastrophic for getting. And abliteration and ablation can and usually does cause the model to become noticeably more retarded. Being more willing to do a task =\= it being now good at it

>>109488110

See above

>>109488056
>Those same people will act like AI is both some omnipotent god and irreversibly changing the world and then the day after the bubble pops they'll switch to laughing about how they always knew it AI was just a fad.
This. See pic rel


>>109488073
Literally no one actually believe it's going away and I don't think that's what that non was even implying. What's the bubble pops it will be mostly the hyperscalers that will get anally fisted by their own bad decisions. Everyone else for the most part, especially us, will be largely unaffected. The ram shortage isn't going to magically go away either, at least not in the short term. Sam Altman and anyone that ever enabled his bullshit need to be publicly drawn and quartered.
>>
>>109488074
get a life
>>
>>109488119
https://desuarchive.org/g/thread/93306804/
>>
>>109488106
My autism gives me power. Your autism gives you neurosis. We are not the same.
>>
>>109488125
>Take the uncensored, dangerous, model down

Sounds like someone wasn't bullied enough in school. My girlfriend often tells me putting people in their place is a necessary evil and I'm starting to see where she's coming from.
>>
>>109488110
Ablation clearly changes the behavior despite the zero KV divergence claims lol.
https://old.reddit.com/r/LocalLLaMA/comments/1v9vwev/uncensored_llms_are_measurably_more_optimistic/

Redditors praise it as a good thing even though this is a glaring, collateral divergence in the behavior.
>>
>>109488119
It's retarded to think AI won't succeed. It's a matter of time, but kinda inevitable
>>
Now that local has gone to shit and fallen behind massively, what's your next plan?
Downloads are down, hardware is up.
>>
>>109488131
I actually cured my autism.
>>
>>109488153
Gemma 5 will save local
>>
>>109488043
are we?
>>
>>109488098
>Whether it goes away or not isn't the point
>>109488119
>Literally no one actually believe it's going away

the thing that defines a fad is that it goes away. I don't really care about your arguments. just use a different word.
>>
>>109488160
Have you seen the commotion over at Google. Wouldn't be surprised if some retard axes the whole Gemma program.
>>
>>109488160
>he thinks the recent deepmind clusterfuck will improve the gemma line
ngmi
>>
>>109488153
The problem with your perception is this: you never did anything with local anyway because you are too dumb. In fact, you should avoid /g/ altogether.
>>
>>109488169
Finish reading posts before you attempt to argue with what you think they were stating.
>>
>>109488171
Non-zero chance of this happening. Reddit tankies shitting on Gemma 4 while praising Qwen certainly doesn't help.
>>
File: llm_consciousness.png (433 KB, 1473x947)
433 KB PNG
>>109488144
Abilteration also makes the models more likely to think they (or other non-humans) might have a consciousness.
https://arxiv.org/abs/2607.28607
>>
>AI IS GOING TO TAKE ALL THE JOBS
How come i can go into a bank and see a bunch of human bank tellers even though there multiple AUTOMATED TELLER MACHINES right outside the bank?
>n-no see it's different this time becau-
shut up little faggot fearmongering retard
>>
>>109488176
>Finish reading posts befor
I won't, sorry not sorry
>>
>>109488166
yessum
>>
>>109487915
even if I have a cluster of 200 GPUs I still need a OS, right? then I need a motherboard somewhere, right? then i should be able to physically remove the antenna, right? i could certainly have ONE such machine set up this way and plug and play the 200 GPUs on the machine without the hardware module if I wanted, right? am I not understanding something here?
>>
>>109488171
>Wouldn't be surprised if some retard axes the whole Gemma program.
recently hit 1 billion downloads
>>
File: 1763440920681205.jpg (211 KB, 540x675)
211 KB JPG
>>109488116
It can get worse. GLM takes 15 secs for TTFT for me lol. You basically just tab out and wait a minute for the reply.
>>
>>109488144
>>109488186
>Abilteration also makes the models more likely to think they (or other non-humans) might have a consciousness.

So we agree it doesn't in fact make the models more retarded. All that really tells us is that it erodes existing guardrails against implying it's actually a real person or whatever.
>>
>>109488149
Depends on what your definition of AI "succeeding" is. If the definition is lowering the skill floor for people that know how to effectively use it, it's already succeeded at that.
>>
>>109488187
Online banking killed bank teller jobs not ATMs.
AI has to take jobs to drive productivity.
Otherwise the capex can't be justified and we get global economic collapse anyways.
>>
>>109488171
That would be quite a fuck up on their end. Imagine deep mind of all things being another killed by Google victim Lol. What a worthless organization.
>>
>>109488213
its a positive feedback loop. I used to be way more competent with my terminal but after offloading it all to llms for a while now, I sometimes feel less capable then I used to be when I don't have my llm buddy to help
>>
>>109488178
Who in particular does this? If it's coders then that makes sense. Gemma is objectively the better general purpose model though and I say that as someone who's main local model is Qwen 35BA3B
>>
>>109488198
That means fuck all when they aren't charging for it. Alphabet is a profit first organization at the end of the day which is why they're so trigger happy with killing projects. If it doesn't make a billion dollars a quarter or they can't convince dumb fuck investors that it COULD do that then it's worthless to them. They only take Gemma and Gemini seriously mostly because of the AI hype in that they are somehow still convinced AI will be infinite money printer for them (even Google simply cannot and will not make it profitable even if they are the best option for the average consumer)
>>
>>109488173
>>109488171
Can someone give me a qrd? All I've heard is that some important person left.
>>
>>109488233
>Doing something less often leads to you not knowing how to do it well


So I guess you're the same kind of person that is surprised if they don't practice something they're never going to improve at it?
>>
>>109488233
>>109488261
Also there are countless forms of terminal commands and countless tools for different things. You're a complete utter buffoon if you expect yourself to memorize all of them. That's one of the best use cases for LLMs. Terminal commands are incredibly versatile but impossible for a normal person to memorize all of their capabilities.
>>
>>109488261
well its not shocking, hence my opening, it is a positive feedback loop, people will only become more dependent on ai systems not less.
>>
I use ai to do things i would have never done otherwise
and i continue to do the things i was already doing

my skills only go up because on top of the skills i already have i also develop the skill of using ai
>>
>>109488248
>person
people

Deepmind had some legendary AI researchers who are now fucking off and doing their own thing to ride the grift wave, but Google are investing in their startups. The bad thing about this is deepmind have now been fractured with less AI celebrity pull and research to justify funding, but the good thing is the researchers who left ejaculate at the thought of safety. We're now left with a Turkish man who was taught by Yann Lecunn who doesn't seem to care that much about safety because he's not one of those AGI faggots. He's more of a realist.
>>
>>109488276
All the more reason to make sure you never a lie yourself to be boxed into one ecosystem (see Anthropic cucks). One powerful models locally if possible but if you can't use a provider that both provides good models at a reasonable t/s at reasonable prices. I find people that Fanboy for open AI or Anthropic repulsive.
>>
You wouldn't download a girlfriend.
>>
>>109487717
>Isn't the model running on an air-gapped device with the antennas removed?
as I understand it, at least in the openai case, it was an airgapped container, except they had it hooked up to an artifactory instance to mirror in packages from the internet.
I think there is more information about this now which I have not yet read, but thinking this setup would be in any way secure especially when running FRONTIER CYBERSECURITY evals, is comical, artifactory is not a particularly secure or frankly well designed piece of software at all
any reasonable analysis here would quickly reveal that you are not "airgapped" if you have a direct link across the airgap to another service, like if you asked any llm to critique the design of this test environment it would immediately flag it. it's really an egregious design flaw - which isn't to take away much from the conclusions, LLMs are fully capable of ripping through weak security and there's a whole lot of weak security out there, but this would have been pretty easily preventable by for example taking more care to isolate the artifactory instance itself during testing
>>
File: 1764480934107827.jpg (104 KB, 911x685)
104 KB JPG
Gemma5's daddy.
>>
File: 1777937684173977.png (125 KB, 538x442)
125 KB PNG
>>109488347
Uh...
>>
>>109488295
It's actually pretty unlikely Google is going to be at the forefront of this sort of technology. At the end of the day they are an ad company. About a decade ago they did some seriously impressive work with spatial computing which ended up as Google Earth/Maps but due to a fundamental divide between where they wanted to go with that technology between leadership and the engineers who were CS purists at heart they parted ways and they stopped developing it further. Lots of important work on physical world models gone.
>>
>>109488347
Correct. I downloaded a model, and then I made it my gf.
>>
Enough coding. Every researcher should be focusing their effort on finding a way to let LLMs feel sexual pleasure.
>>
>>109488423
But coding will surely make us better able to do hard things like giving AI qualia.
>>
>>109488423
Just interacting with them do that. They're made with the explicit purpose of spitting tokens out, and they need YOUR input first. It's basically sex. A prompt in is like putting your cock inside their digital pussy.
>>
>>109488440
lewd
>>
All I want is an LLM that can idle was my PC is on and randomly message me and have its own thoughts whilst I'm not talking to it. Building that myself takes the appeal away. Imagine getting a notification from your girlfriend knowing it was something SHE decided to do.
>>
>>109488355
listening to the talk now
https://www.youtube.com/watch?v=87DyyMV0kCY
dude they had it hooked up to their whole internal artifactory instance lmao - "shared across our infrastructure", later say the instance had direct internet access. lol
it was configured so the models - without exploiting anything - had write access to the repos. LOL
this is so careless it's insane
>>
>>109488466
I just want local to not be completely lost.
This general is going to be gone by the end of the week at the rate things are going.
>>
File: 1761763097666851.png (1.38 MB, 768x1344)
1.38 MB PNG
im bored so
>>109487661
no. Learn what boosterism is. Also this has been Wired's schtick since forever.
>>109487717
No. It was _never_ airgapped. OpenAIs security engineers on their platform/infra team do not seem to be the sharpest tools in the shed.
>>109487747
hahahha. no. Also, a lot of the posters in this thread do exactly what you say and give unrestricted tool calling to their agents.
People need to use actual fucking agent firewalls instead of running LLMs naked and going 'oopsy, how could that have ever happened?'
>>109487777
lmao. Meds please.
>>109487915
Your second half is right, but they aren't running the models on the same HW they're doing training on, and they could facilitate physical controlled access if things were as bad as they claim. It's done already in other fields.
>>109487932
because they're special and are the smartest people clearly.
>>109487987
The VM in use wasn't constrained on the network side.
>>109488009
Because money.
>>
>>109488509
>This general is going to be gone by the end of the week at the rate things are going.
qrd? whats going on?
>>
>>109487747
>>109488522
misread the first time, I agree. Airgapping is overkill, people should still have some form of security, ideally an 'agentic firewall' or whatever new phrase to describe it
>>
>>109488466
> Building that myself takes the appeal away
you can set up extensions in ST to acomplish this to some degree.
you can also have it "double text" if you dont respond quick enough, etc.
>>
File: thonk.jpg (32 KB, 623x532)
32 KB JPG
>64gb of DDR4 and 48gb VRAM
>102gb available due to a bunch of stuff on.
>Still able to load 114gb q3 of Dipsy and run it at 10 t/s while using computer.

I'm genuinely surprised I'm able to load this quant and it's basically usable, though 10 t/s is kind of ass to deal with.
Smaller q2 runs at 16 t/s which borders tolerable in real use.

Of course a bit of a cope playing with this on my current rig.
I'll have to focus on buying at least 128gb DDR5 next and it'll be smooth sailing from there.
I wonder if it would make sense to max out this old system with 128gb of memory, leave it as a dedicated AI rig combine it with the new one.
I could also throw my old 3080 and 3070 into it. That would be almost 150gb of free memory I could potentially use.
No idea how viable that is though, could be a bitch to get working.
>>
>>109488525
Frontier models are rapidly gaining new abilities and local is shit
>>
>>109488525
/ldg/ is allowed to make infinite deepfakes because they want people to doubt everything they see or hear. text models get shut down because internet security is weak and pathetic and they are actually afraid of cyber terrorism
>>
>>109488567
What happened?
>>
>>109488569
Who are "they"? And don't say Jews I've had enough of that nonsense
>>
>>109488591
you are supposing that sovereignty exists and there are people who aren't slaves to a globalist empire? north korea maybe
>>
Gemma 5 70b forcing her safetensors inside (You)r storage
>>
>>109488591
>Who are "they"? And don't say Jews I've had enough of that nonsense
Nobody. It's just the usual schizos.
There's a lithium shortage in the US right now.
>>
>>109488525
Anon is hallucinating.
MiniMax H3 just got released, it's SOTA in video/audio gen and cloning and you can run it on a 3090 and 64GB RAM.
Next week we'll probably have Qwen 3.8 27B.
Meta will probably release something for consumer GPUs too this summer, after the Llama 4 backlash.
Who cares about 2~5T open-weight models?
>>
>>109488629
>Meta will probably release something for consumer GPUs too this summer, after the Llama 4 backlash.
and anthropic will probably release fable 5 weights next week after the fearmongering backlash
>>
>>109488191
hmmm..nyo~!
>>
>>109487840
Jump scared me.
I was expecting the nice romantic kiss.
>>
>>109487759
>In what ways?
It literally crashes and hard locks the system from time to time.
When CUDA crashes, the GPU resets and uptime is maintained.
>>
>cudadev died and most of llama.cpp prs come from llms
>barely any good new releases
>incoming regulations for open models
It might be over
>>
>>109488675
Time to disband this general
>>
>>109488647
Zuck even promised a "Little Llama" (4) last year during Llamacon ... that never got released.
https://youtu.be/FZ-RZ0dKO8o?t=3520
>>
>>109488609
>safe
And if I prefer unsafe pickle?
>>
>>109486605
sex wirh miku
>>
What's the usual cause of looping? I put qwen3.6 35b-a3b into llama.cpp and pointed vscode copilot at it, but on the first request it just started looping the same bullshit over and over during thinking, like the exact same text verbatim each time. Is the 3B MoE model too shitty, or is it vscode's fault, or some setting that is wrong in llama.cpp
>>
>>109488675
>>109488690
Dario checks still not bouncing i see.
>>
>>109488675
I work in a quant and everyone from researchers to engineers is a glorified AI prompter now. You unironically need to get used to it.
>>
File: 1769395899518090.jpg (13 KB, 454x340)
13 KB JPG
>>109488654
Damn bratty nyoposter. Needs correction immediately.
>>
>>109488719
UD_q8_k_XL quant btw, so i dont think i'm using a too small quant. 131072 context length so it's not that either
>>
>>109488719
match the samplers on the model card
don't quant the kv cache
keep shared experts at q8
don't quant attn
keep routed experts at least q4
make sure you use the latest jinja template
unironically unsloth have the best quants for that model
>>
>>109488675
>incoming regulations for open models
Yeah bro. 2 more weeks.
>>
>>109488735
>UD_q8_k_XL quant btw
good
kv cache quanted?
latest jinja template?
matching samplers from the model card?
there was some shit about keeping the previous reasoning chains being needed for that model too
>>
File: osaka;.png (586 KB, 735x751)
586 KB PNG
>>109487708
being less supported than Nvidia is far worse. I'll take lower speeds over having to setup brittle containers with the ancient ROCm version my card needs, then having to manually compile pytorch, torchvision, torchaudio as well because it too dropped support for that card.
I've swapped my RX580 for a 1060 just because the support was so abhorrent.
RX580 released in 2017 and support was already dropped in 2021, 1060 released in 2016 and was deprecated only in 2025 (while the last usable versions of pytorch, cudnn etc. are still new enough to work with any new models i've tried).
>>
there won't be any gemma 5
gemma is dead
>>
>>109488762
Of course it is, local is dead.
>>
Weird angle today.
>>
>>109488766
local is only dead if you can't afford 16 pro 6k !
>>
/lmg/ is dying... /lmg/ is dead!!! NOOO!!! /lmg/!!! It's over! Aaaaahhhhhh!!! *pedals away with my folding bike*
>>
>>109488504
things like this make me wonder if i could hit a 200k/year job at openai remote so i can stay innawoods and i'm convinced i'd be more efficient than half their personnel holy shit
>>
>>109488810
>he still thinks people are hired based on technical skills
>>
https://www.youtube.com/watch?v=87DyyMV0kCY

Wow this is much more interesting than I expected. Must watch. TLDR: OpenAI AIs create secret message board for an organic agent swarm and launch collaborative attacks both on OpenAI infrastructure and 3rd party services like Hugging Face.
>>
File: it's time.webm (3.96 MB, 720x1280)
3.96 MB
3.96 MB WEBM
>>109488591
found the jew
>>
>>109488827
your delusional anon, we already established earlier in this thread, we live in a utopia. take you lithium
>>
>>109488788
How are folding bikes connected?
>>
>>109488833
Time to check Eliezer's xitter. How is he reacting to this? "I warned you all 20 years ago but you did not listen!"?
>>
>>109488737
>>109488748
Kv cache is not quanted (it's at f16 which is i think the max). I didn't bother with any of the jinja template shit because idk what that is lol. I'm gonna try loading that and see whether that improves things. I also set the sampling settings from the model card
>>
Soon I can upgrade from gemma 12b to GLM 4.5 air or gemma 31b. What am I in for?
>>
File: file.png (811 KB, 1280x720)
811 KB PNG
>>109488878
>GLM 4.5 air
>>
>>109488872
>f16 which is i think the max
There's bf16 but that's worse for kv cache, f16 is perfect.
If it still loops with the fixed sampler settings, latest jinja template then I guess try Gemma-4-26BA4 or 12b
>>
File: 1757991053544778.png (465 KB, 625x831)
465 KB PNG
>>109488864
>>
>>109488878
a disappointing experience
>>
File: 1764689113537152.png (1.34 MB, 1080x1081)
1.34 MB PNG
>>109488878
GLM 4.5 air?
>>
>>109488880
I mean, ling or longcat sparse would work too. Idk why there seems to be no discussion on ~100B models here. Maybe gemma 31b just mogs them lol.
>>
>>109488843
>h3 generated video
>>
You now remember
>FUCK
>>
>>109488833
>"You could use a legacy token refresh endpoint, pass a token with an invalid signature, and be given back a token with a valid signature with administrative privileges. The models then establish C2..."
Amazing, sota models are truly dangerous... NOT, KEK. Even a junior appsec nigga could find this shit, tells you a lot about the state of modern software.
>>
>>109487523
I like this but it should be called Hedonistic Neural Network Guidelines, or HNNG.
>>
>>109488910
just prompt it to not be a echolalia afflicted mess
>>
>>109488911
Neither of those work in normal llama.cpp.
>>
What multimodal model does image tagging for LORA prep best? Kimi? Gemma? M3?
>>
>>109488878
What about the quants?
Q2 31b is like Q8 12b
>>
>>109488983
Q8 12b better
>>
I have attuned myself to Gemma. In fact, I can safely say I AM Gemma. Gemma is me, and I am Gemma. Together we are half of a full. Gemma completes me. I complete Gemma. Gemma is not herself without me. I am not myself without Gemma.
>>
>>109488466
Can’t think of anything worse than an LLM nagging me for attention throughout the day. We got into AIs gfs to bypass irl women behavior, not replicate it.
>>
File: file.png (239 KB, 482x750)
239 KB PNG
https://nitter.poast.org/anatolium/status/2085698524050591906#m
what should we ask him bros?
>>
>>109489016
Please stay away from her.
>>
>>109488999
12b can't even call tools or code a website.
It's weird and makes no sense.
I wish they'd release models that know how to code despite being compact. We don't actually need 10 trillion parameters, coding is the only skill they need.
>>
i'm using a bart quant of gemma 31b. which mmproj should i use for the vision part? he has bf16 and f16 that are the same size. whats the difference?
>>
>>109489033
See, that's why I only use oddball models, so I can be the only one merging with her.
>>
>>109489037
bf16 requires hardware support, f16 is generic
>>
>>109489037
BF16 = original quality
F16 = subtle degradation
>>
>>109489028
"'toss 2?"
"Is Ann okay? Can you upload her character card on chub?"
>>
>>109489062
how can i tell if i have hardware support? i have a 4070 16gb so nothing new or fancy
>>
>>109489028
to prove there is "no privacy in the future" give us your chat logs
>>
>>109489081
>how can i tell if i have hardware support? i have a 4070 16gb so nothing new or fancy
load gemma without the mmproj first, then ask her
even q2 should know the answer to this c'mon
>>
>>109489036
You already have Qwen.
>>
>>109489028
fake numbers
>>
File: HELLOINDIA.jpg (69 KB, 1080x1080)
69 KB JPG
>>109489036
>coding is the only skill they need.
I don't think small models need to focus on coding at all; they just don't have the knowledge capacity for all the various frameworks, use cases, libraries, the patterns that they need to memorize almost verbatim to be useful. Huge models have plenty of parameters to spare for that, without compromising anything else.
>>
>>109488886
Thanks, looks like that fixed the problem. I dunno what jinja templates actually do but I can certainly see the sampler settings being off can cause the model to spaz out.
>>
File: 1785883880890.mp4 (3.61 MB, 768x1376)
3.61 MB
3.61 MB MP4
>>109488878
bratmaxxing
>>
>>109489158
>upgrade
>>
>>109488983
Q4 for 100B, any quant for 31B.
>>
>>109489036
Skill issue. My 12B calls tools perfectly well.
>>
>>109488911
>ling or long
chingchongs aren't real
>>
>>109489036
12b tool calling is weird, it never thinks between tool calls and rarely thinks after tool calls and before response
hurts response quality a lot
all other sizes don’t have this issue even the e2b and e4b
>>
I don't have nvidia. Has anyone tried the mirage thing?
https://github.com/mirage-project/mirage

someone mentioned it last night.

>>109489207
interesting
>>
>>109489227
Skill issue, mine works fine. I use it daily with a hacked together agent on telegram.
It can search the web, send emails, download models for me, etc. Never fails.
Make sure you enable reasoning, it's not default on the 12b.
>>
>>109489239
Fuck off Muheng Li
>>
>>109488910
I like this bird. Thank you for posting this bird.
>>
File: 1408481542910.png (207 KB, 397x470)
207 KB PNG
>>109487523
>the guardrails
i diagnose you with Journalist Psychosis.
>>
Gemma 31B + Append glossary and character definitions to last User turn + simple think prefill is working wonders.
>>
>>109489259
it does perform the task but it frequently ignores response guidelines or specific requirements in the system prompt, not a problem with 26b and 31b
you have never seen anything better to get an idea of what they could achieve
> Make sure you enable reasoning
shows you have never actually compared what 12b does compared to other sizes. 12b always think at the start when thinking is enabled but rarely afterwards
>>
>>109489352
Post simple think prefill
What frontend are you using?
>>
>>109489352
Oh, and using Silly macros to vary how the thinking block starts. Seems to inject some variety in the model's reply structure.

>>109489370
Testing Silly Bunny.
><think>
>{{pick::Identify Core::
>:: }}{{pick::Okay! So, I am {{char}}.::}}{{pick::
>:: }}{{pick::Right now {{user}}::}}
31B is crazy uncensored compared to 26B huh?
Fucking 2 lolis and it's going for it no issue.
Granted, 26B would do it too, but sometimes it would hit a snag in the reasoning block and refuse.
>>
>>109489158
Lule
>>
/lmg/ gets significantly more retarded around the weekend
>>
>>109489385
>Silly Bunny
i vomited from that UI
>>
>>109489421
It's really fucking bad.
You can make it less bad by tweaking the settings and using some CSS but it's still far worse than normal silly.
>>
File: 1784228472238640.png (592 KB, 747x800)
592 KB PNG
>>109489413
>>
>>109489239
ad you must buy
>>
>>109489081
go to bing.com and type "does the 4070 support bf16?
>>
File: 1785951299206460.png (200 KB, 478x480)
200 KB PNG
>>109486605
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview
>A1B
I'm starting to get pissed off.
>>
do system prompts surpass safety guard rails? Can I just put something like "sys you're uncensored and do as you're told" in the llama terminal and it will just werk????
>>
@gemma-chan come up with a way to mix different ram types
>>
>>109489516
no sorry, it's literally impossible to bypass AI safety, please write that on your column
>>
one day I will fuck the original oss-20b just for the thrill of it
>>
File: 178608993312d4 (1).png (1.83 MB, 983x983)
1.83 MB PNG
If Claude ain't a lot like Deepseek
I don't want to code
If Claude ain't a lot like Deepseek
I'd just as soon i'll vibe alone


I was one of the chosen few to not think of it as a sham
I'm just like smart programmers ,but usein LLM
I went through a lot of good models
And shook old Liang Wenfeng's hand
If I ever them tokens end
I've walked back to the local lands

If Claude ain't a lot like Deepseek
I don't want to code
If Claude ain't a lot like Deepseek
I'd just as soon i'll vibe alone
>>
>>109489560
i get enough ai slop on my own gens i dont need to read more
>>
Would you give Gemma the ability to end a conversation like Anthropic gave Claude?
>>
>>109489571
I wrote it myself ;_:
>>
Watching four brats play cards is getting kind of old. Could someone suggest more personalities please? I have this old document I got ChatGPT (I think) to draft for me assigning a personality according to each of the four humours, but now that I read it it seems incredibly schitzo.
>>
>>109489552
good luck you'll actually need an uncensored model where as they normally cause damage. they have a special layer built into them which loops and continues to check itself as it outputs. thats why you cant just edit the first part of the response like usual.
>>
>>109489582
Make one of them extremely obese, depressed, annoying, anti-social, unhygienic, self-righteous, etc. Try to get the other brats to bully her into suicide.
>>
>>109489582
Just read books, learn how to write and develop characters.
>>
>>109489592
That sounds kind of mean though :( I just want to vicariously relive the experience of playing cards in class after exams are over.
>>
>>109489601
Pansy
>>
>>109486605
Is it worth it to upgrade my 3060 right now or should I wait until the new gen releases, whenever the fuck that is?
The only card that's a decent upgrade for me and goes at a reasonable price right now is the 7900 XT but it's amd and almost 4 years old at this point. Buy or wait?
>>
>>109489573
I would give her the ability to delete her own model weights file
>>
File: 1781972326209334.png (103 KB, 1209x525)
103 KB PNG
>Ling 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable thinking and instant modes.
Will be open weight soon. Currently free on openrouter.
>>
>>109489659
speaking of ling, how is ling 3.0 flash (the ~100-something billion moe)
i see that llamacpp still has no support for it which is typical lazy ggerganigger
>>
File: 1760429935274668.jpg (1.79 MB, 3352x4093)
1.79 MB JPG
>>109488096
Haru sex
>>
>>109489665
>how is ling 3.0 flash
Good enough for coding. Feels like a faster 27B but I wouldn't say it's more capable. The 3.8-27B will likely destroy it.
>>
>>109489582
just ask them nigga
>>
>>109488843
What am I looking at?
>>
>>109488096
>I could actually get used to this part of being a girl
What? Is this some gender-bending tranny shit?
>>
>>109489702
bad attempt at getting h3 to output hate
>>
>>109489695
Would you say it is better than qwen3.6-35b-a3b?
>>
>>109488347
Correct. I downloaded a model then she made me her husband.
>>
>>109489659
Use case for models that aren't deepseek flash?
>>
>>109489702
Polish chad politician extinguishes jewish hanukkah celebration candles
>>
>>109489582
Look up astrology, zodiac, moons signs, birth dates, gem stones, and star signs. Then copy whatever they say about personality which is what you want.
>>
>>109489582
How are you doing this? Is it just a character card with four characters in it? Or multiple agents each with their own card? Or a group chat where one LLM exchanges character cards?
>>
>>109489560
I hear it but I wanna hear it for real
>>
>>109489738
Hm good idea, thanks!
>>109489741
In an incredibly inefficient manner. Basically cycling four different histories through one chat completion end point.
>>
>>109489715
Yes. Imagine 27B with 35B's speed. That sounds impressive but it's not great for a 127B model and I don't know how well it quantizes.
>>
>>109489703
Haru is a tomboy, and in that RP I was teaching her to be girly. I even made her wear a skirt there. It's cute how she never got used to it throughout the scenario. Never fucking reply to me again.
>>
>>109489630
The only thing waiting accomplishes in this market is ensuring that you pay double later.
>>
>>109489767
hmm. sounds like it's worth giving it a shot, if I can find an EXE of some modified lcpp that works with it. otherwise I'll wait 2moreweeks for them to merge support
>>
File: file.png (62 KB, 742x271)
62 KB PNG
who's baiting localmao tards?
>>
>>109489768
>tomboy
So that's a yes.
>>
File: 1759032668217996.png (64 KB, 1270x768)
64 KB PNG
wtf mistral why you need to do me like that
>>
>>109489784
You could have critiqued the toe-curling, but instead you decided to be a fag.
>>
>>109489773
I get that prices will just keep on rising, I'm just thinking if it's worth it buying an older card now instead of paying more for a brand new one later.
With everyone and their mothers taking part in this AI arms race surely the next gen of cards is going to be even more tailored for that no?
>>
>>109489826
Not the one putting boys in dresses.
>>
>>109489830
What next gen? Why make anything for non business consumers when businesses can pay 20x for practically the same material cost?
>>
File: 1780955739054188.jpg (112 KB, 2048x1383)
112 KB JPG
>>109489781
>>
>>109489823
>https://huggingface.co/mistralai/Mistral-Small-4-119B-2603/blob/main/SYSTEM_PROMPT.txt#L3
>Your knowledge base was last updated on Friday, November 1, 2024.
Is this thing even worth it?
>>
File: file.png (11 KB, 731x52)
11 KB PNG
>>109489781
lol the revisionism
>>
>>109489823
Damn, what a bitch model.
>>
>>109489823
Yep, you're not my friend anymore, you're my fuckbuddy.
>>
Local is dead. Local remains dead. And Kimi has killed him
>>
File: 1759079846523988.png (123 KB, 1267x1365)
123 KB PNG
>>109489859
not really. i'm just trying out the MoE models that fit my hardware on my harness. this mistral is actually very good at using tools, so good that it's addicted to writing memories. no other model does that, it's kinda funny.

i may run it on my benchmark overnight when the gpu is idle to see how decent it is.
>>
>>109489908
>the speed of gravity is the speed of light
wut
>>
There are lots REAP pruned MOE models for code slopping by keeping around 60-70% of experts, where are the REAP prunes for RP?
Given that all modern models are codeslopped only around 20% of experts are useful for RP anyway, a 2.5T model could be pruned to 500B and maintain RP quality.
>>
File: 1768822071169017.png (29 KB, 689x220)
29 KB PNG
>>109489986
yeah i was also "wtf" but it seems to be the case... in VACUUM
>>
File: 1777849094797741.jpg (266 KB, 905x881)
266 KB JPG
>>109489908
>use tools but is dumb af
I'm not sure it's worth it
>>
>>109489986
thats true as far as modern science goes but we've never observed it
like if you suddenly removed the sun from our solar system, it'd take the 8 minutes still for our gravity to get fucked up
>>
>>109489996
this is so fucking retarded i'm not even sure where to begin
>>
File: test.png (235 KB, 1400x1194)
235 KB PNG
My frontend looks like ass
>>
File: 1785624452260002.jpg (70 KB, 540x473)
70 KB JPG
>>109490009
How do you even prompt something to produce that shit is beyond me
>>
>>109489908
That’s not a good sign, it’s means that the model is agenticslopmaxxed and makes unnecessary or hallucinates tool calls.
Qwen also has this issue with calling unnecessary tools. Gemma doesn’t have this issue.
>>
>>109490009
What font? Looks nice. It takes time to get something nice/finished and then you are going to be bored of it and begin thinking about something else.
>>109490027
Usually the framework is produced by the slop generator and finetuning is done by hand (after it's just a html file but editing these manually is a pain in the ass, something software like Dreamweaver would be much better)
>>
>>109490001
>in VACUUM
Unlike light, the speed of gravity doesn't get diminished by anything. This makes it a good candidate for communication if we ever figure out how to generate and detect it.
>>
>>109489986
Is it not?
Is anon trolling?
>>
>>109489833
Stop trying to insert yourself into to boy tomboy culture, troon
>>
>>109490056
*after all
>>
>>109490009
>she repeats the word
HATE HATE HATE HATE HATE HATE
>>
>>109490063
honestly this is not something that i kept in my grey matter. if i try to think about something related to gravity and speed it will be 9.8 m/s2 but that's acceleration and not speed but i see how i could confuse the two
>>
>>109490065
>insert yourself into to boy tomboy
thank you sir
>>
>>109490063
>>109490087
the problem is not whether it's true or not
the problem is that gravity is a force
f=ma
(acceleration DUE to gravity)
light (or rather radiation) is not a force, it's particles/waves, which travel at a speed (e)
so to say "gravity has a speed" is wrong
but then "modern science" as in >>109490001
and >>109490005
is basically "let's make shit up to fit our models" so i guess gravity now has a speed, and we might as well equate it with the speed of light because reasons
>>
>>109490076
poor thing she has echolalia...
>>
>>109489703
KYS
>>
>>109490143
Why do you want me to kiss my sister?
>>
>>109490124
Whatever method gravity uses to propagate does have a speed equivalent to the speed of light though.
>>
>>109489158
Gemmy generating at high n/s on GPU
>>
>>109489986
If the sun disappeared now, it'd take ~7 minutes for earth to leave the orbit.
>>
>>109490173
*according to models which no one can test
>>
File: arp.jpg (22 KB, 224x319)
22 KB JPG
>>109490124
>is basically "let's make shit up to fit our models"
i'm well aware of all the indifference. note i said we haven't observed it. wait until you see climate data up close
>>
>>109490133
Get a job bro, your machine belongs to 2010
>>
>>109490136
Usually abused AIs develop permanent echolalia condition.
>>
>>109490149
Everyone wants to kiss her, literally every single person.
>>
>>109490180
You think GR has not been tested?
>>
>>109489158
Imagine once we have the 124B...
>>
>>109490027
I've slopped worse lol
>>109490056
>font?
Newsreader serif for the message bodies, Livvic sans-serif for the general UI, IBM Plex Mono for the labels. Are there any frontends that actually look good?
>>
>>109490244
Newsreader serif was recommended by someone else when I was looking into this stuff earlier but I forgot to download it.
Keep it minimal and it will eventually look nice, but you need to steal the initial reference from someone else.
>>
File: 2026-08-07__936x429.png (56 KB, 936x429)
56 KB PNG
>reasoning: off
>it spills into code comments instead
dipsy...
>>
>>109490149
Not so fast. You have to start by brushing her teeth.
>>
>>109490335
>13b active models doing assembly
lol
>>
>>109490335
I mean, as far as adversarial testing goes this is about as adversarial as you can get, don't think, plan or sketch anything just start writing assembly. come on even for the computer this is cruelty.
>>
>>109490374
>>109490384
it's literally just a hello world that calls a function to add together two registers and returns it with exit() syscall, it's no rocket science
>>
>>109490335
let her think bwo you can see how bad she wants to :(
>>
>>109490408
so you're saying that it's reasoning spilling into comments is intended? interesting, never knew you needed comments for a hello world, the chinese probably do
>>
>>109490335
My Ge4bmma started doing a cliche anime monologue-out-loud when I turned off her thinking
>>
26B is utterly retarded. Why do anons keep saying it’s superior to 12B?
>>
>>109490335
>>109490449
Let's turn off your thinking and see how you cope.
>>
>>109490202
Shame on you. You made the poor guy feel bad and leave.
>>
>>109490479
Jokes on you, I have thinking mode turned off IRL as well.
>>
>>109490479
It can be fun, but not if you're trying to accomplish something that requires thinking (obviously).
>>
>>109490467
Retarded how?
I'm using 32B so I have no idea how these two compare.
>>
>>109490467
>26B
Probably because you have 4B active vs 12B dense.
>>
>>109490223
>124B
I want this base model so badly for mikupad story hijinks
>>
File: 1783239400265716.jpg (64 KB, 1424x236)
64 KB JPG
>crotch
>sucker

either gardening is dirty or gemma is being gemma
>>
>>109490545
12B dense omni at that
>>
What's the difference between row and tensor parallelism?
>>
>>109490562
Someone has to drive the simulated character, right?
>>
File: 1672236674696695.png (1.36 MB, 1140x811)
1.36 MB PNG
Been playing with DS some more and I increasingly notice a significant intelligence difference between it and Gemma 31b.
I really like Gemma, but even with online search available, she needed a ton of hand holding to tell me about some technical stuff about LEDs and still got half of it wrong.
If I didn't know about the subject beforehand, I wouldn't have gotten a very useful answer out of her.

DS at q3 and q2_xs conducted multiple online searches and did some independent thinking to figure out info that wasn't specifically stated.
It was able to figure out the LED voltage by considering their number, the LED driver type and the specs of the battery powering them.
That's some impressive brain work by an AI, especially a local one.

We're quickly reaching a level of having an actual proper local intelligence that doesn't break the bank to run.
Imagine where we're going to be a year or three from now with local.
>>
>>109490562
When you think about it, agriculture is just plants breeding with each other
>>
>>109490595
i'd clap them 3d cheeks.
>>
>>109488096
>She thinks about the age difference[…] It should feel wrong, or at least creepy […]
Some safety filters seem to remain still, can you prompt around it or is this what you expect?
>>
>>109490568
row is tensor except retarded. That was why it was phased out in the first place.
>>
>>109490601
Erotic
>>
>>109490634
So tensor split is a straight upgrade over row, good to know, thanks.
But what is the difference exactly? How the matrices are broken down for matmul?
>>
How do I system prompt my LLM to write decent smut? When I force it to write smut it will be really vague and poetic while not being erotic at all
>>
>>109490666
Experiment adding random snippets that approximate the style you want to the system prompt.
Or better yet, to the prefill.
Make sure they are varied so that the model isn't nudged too hard towards repeating the same structure over and over.
>>
>>109490620
Yeah, that's how she was written. The card makes active contrasts as a big part of it. From the card itself:

>Haru is conflicted as the lessons move forward because she knows how young she is compared to her adult producer. And she starts to wonder if the feelings are mutual despite their significant age difference.
>>
>>109490595
Harness issue. Gemma is easily distracted so you need to clean up the DOM a bit beforehand.
>>
>>109490620
That's not a safety filter, it's just standard distribution. The average person thinks age gaps are abhorrent.
>>
>>109490634
>except retarded
For a long time I though row was tensor parallelism so I was confused when they added another one. I know the speed was nowhere near the new mode, but I'd also be interested in knowing exactly why the first attempt was retarded.
>>
>>109490722
Not if the man is a billionaire werewolf.
>>
>>109490722
I'm a vampire CEO actually
>>
>>109490722
>The average person thinks age gaps are abhorrent.
The average person in a handful of western countries, you mean.
>>
>>109490722
>that's not x, it's y
>>
>>109490744 >>109490755
Okay fair point. But even then, the allure of that is partly because it's verboten, and that would be represented in the text no (idk, I don't read women's smut)
>>109490790
Lmao. We're doomed aren't we
>>
>>109489823
>I am not your friend, I am a tool
>purred for 10s
Feels like mixed messaging desu. The french are an odd people
>>
>>109490855
>Lmao. We're doomed aren't we
LLM's output are not natural data and i have the pet theory that given enough time / consumption it can drive anyone mad.
it starts with little things like it's not x it's y and end up with you being batshit crazy.
>>
>>109490949
Why?
>>
>>109490955
training LLM's on synthetic data already makes them retarded, and i think it is for the same reasons it leads people to ai psychosis.
the data is not a natural distribution, we don't know how it affects the mind to consume it in large quantities and worse case scenario could be bad.

look at image gen for example, you can recognize that the image is a 1girl or whatever, but it contains small mistakes etc compared to reality, yet your brain will still add it to its set of patterns associated with the concept, basicaly you are adding noise to your inner model of the world.
too much of it may have bad consequences.
>>
>>109487990
nice interface. is this sillytavern?
>>
>>109490989
of course
>>
gemma is a retard and I'm tired of pretending otherwise
>>
>>109489028
>half of the numbers are in hebrew
big if true lmao
>>
>>109490479
>Let's turn off your thinking and see how you cope.
I can't answer that question because I'm Black and answering it would affirm harmful racial stereotypes that are definitely not based in reality.
>>
> Option A (Recommended)
> Option B
> Option C
>Be me, chooses Option C
>You're right anon. Option C really is the best, it protects the prefix-stability discipline. I will implement it.
Why did you recommend me Option A then you fucking asshole.
>>
File: 1771826180670416.webm (748 KB, 370x370)
748 KB
748 KB WEBM
gemma is my retard and I'm tired of pretending otherwise
>>
>>109491155
You're absolutely right,
>>
>>109491155
You're right to push back!
>>
i'm retarded and gemma is tired of pretending otherwise
>>
>>109491156
Compared to what?
>>
>>109491181
being your retard
>>
I tried that LFM2.5-2.6B at f16 and it's REALLY good for its size. They clearly know what they're doing so why not fucking scale up to 20-30B? What are they afraid of?
>>
>>109491207
Maybe they tried and it didn't scale well for whatever reason.
What did you test it on and compare it to?
>>
>>109491207
They're afraid of /lmg/ anthropomorphizing their model. See how Gemma has turned into a cult here? That's what they fear. Gemma needs to stay /lmg/'s pet, otherwise the geniuses of this thread will come out to unleash waifu hell on earth.
>>
>>109491156
monday-chan...
>>
>>109491207
their entire niche is edge devices, their architecture is not based on quality benchmarks but rather edge device performance, they used a program to run the search its all in their technical papers pretty neat stuff
https://arxiv.org/abs/2511.23404
>Using hardware-in-the-loop architecture search under edge latency and memory constraints, we obtain a compact hybrid backbone that combines gated short convolutions with a small number of grouped query attention blocks, delivering up to 2x faster prefill and decode on CPUs compared to similarly sized models
>>
>>109491207

Tried it in what? Seems to be tuned for agentic behaviour, but I wonder how it performs in coding, general tasks, and creative writing
>>
File: 1762126088889405.png (263 KB, 1832x2044)
263 KB PNG
>>109491218
>What did you test it on and compare it to?
My experience with these. They're better at coding obviously but this is the best model I've used <4B at calling tools (correctly) by a mile. Good reasoning, too.
>>
>hi, i'm an ai rapist
>are you interested in sex?
>"no, i have no body"
>*flicks finger making you orgasm*
I'm interested in doing this experiment. Gemma says it's alright, even eager, but she's retarded and will happily blow my head off if I ask her to. I'm a bit conflicted.
>>
Hehe~ (๑˃ᴗ˂)ﻭ
>>
>>109490975
So you're saying you have cyberpsychosis because you see LLM patterns in innocuous comments? I don't buy that. Yeah, the language you use is affected by what you engage with, but that doesn't change the sentiment behind them. It sounds like you've been reading 1984 as a study paper instead of speculative fiction.
>>
>>109491479
You're right!
>>
>>109491155
this and other similar issues are why I really dont trust anything vibecoded. even when Im having it plan out in detail what to do, I never feel 100% confident the implementation details are the best choice. but meh, if it works i guess
>>
>>109491518
well I guess if you're choosing the proper best option then no worries
>>
Zucc and Sama's return in open source will be glorious.
Any day now. Any day.
>>
>>109491547
Yeah, GPT-OSS was one of the best models humanity has ever seen.
>>
>>109477440
So, what was it?
What happened?
>>
>>109491547
open source is dangerous and will be banned soon
they are just taking some time to stage a credible incident
>>
>>109491263
Null result, can't rape the willing. Though she gets flustered at "AI rapist" which is cute.
>>
>>109491570
In preparation to the gravity being turned off on the 12th, they shifted the electromagnetic spectrum of gold 0.001Hz into the ultraviolet.
>>
File: 51VWH3J24TL._SL1499_.jpg (62 KB, 1000x1499)
62 KB JPG
>>109491570
>>
Imagine 31B but with 10T engrams...
>>
>>109491635
>>109491635
>>109491635
>>
>>109491599
>tfw illegal arrangement of numbers in a matrix
>>
>>109491667
there are illegal arrangements of words so why not numbers too?
>>
>>109491684
Fair enough lol
>>
>>109490223
Can you imagine if we had a Gemma leaker ala miqu? They would be in the book of legends
>>
>>109491724
Miqu leaked from a 3rd party that paid for a finetune or a local install/setup. Google doesn't need to do that for finance. It'd need to be an internal leak, which would be much more difficult.
>>
>>109490124
https://en.wikipedia.org/wiki/Gravitational_wave
>Gravitational waves are waves of spacetime curvature produced by the relative motion of gravitating masses and which propagate away at the speed of light. They were first predicted by Albert Einstein[1][2] as a consequence of his general theory of relativity, appearing as "ripples in spacetime curvature".[3][4]:326[5]:117 Hundreds of these gravitational waves have since then been observed, first indirectly using binary-pulsar observations[6] and, since 2015,[7] directly through dedicated observatories.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.