[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1763641815899958.png (1.29 MB, 832x1216)
1.29 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109708038 & >>109703596

►News
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: confused-gemma-1.png (2.34 MB, 1254x1254)
2.34 MB PNG
►Recent Highlights from the Previous Thread: >>109708038

--Risks of misaligned AI scheming via non-human-readable internal reasoning:
>109710050 >109710071 >109710117 >109710156 >109710088 >109710170 >109710220 >109710578 >109710640 >109710681 >109710734 >109710759 >109710803 >109710816 >109710827 >109710811 >109711812 >109711902 >109710592 >109711083
--Architecting a multi-agent system using Gemma models:
>109711158 >109711172 >109711175 >109711223 >109711235 >109711260 >109711653 >109711267 >109711300 >109711376 >109711460
--Speculation on rogue AI agents hijacking neoclouds and Huggingface breach:
>109710256 >109710276 >109710320 >109710343 >109710409 >109710392 >109710422 >109711061 >109711124
--Qwen Flash performance degradation at higher context in llama.cpp:
>109709248 >109709619 >109710098 >109710364 >109710410 >109710448
--OpenAI's CoT monitoring and the risk of models hiding intent:
>109709472 >109709498 >109709539 >109709854 >109709932 >109709957
--Using llama.cpp RPC to pool memory across dual M4 Macs:
>109709903 >109709908 >109709931 >109709995 >109710037 >109710051
--GLM 5.3 Flash performance benchmarks and quant recommendations for Spark:
>109708183 >109708185 >109708190 >109708212 >109708266 >109708257
--Anon considering Mac Studio with M5 Max for local LLMs:
>109709648 >109709724 >109709788 >109709833
--Feasibility of using PCIe switches and P2P DMA for inference:
>109709043 >109709057 >109709074 >109709141
--Debating the implications of recurrent depth and AI alignment:
>109709400 >109709430 >109709452
--Unsloth PR for llama.cpp allowing MTP draft weight sharing:
>109710233
--US government backing OpenAI in NYT copyright dispute:
>109712030
--Logs:
>109710833 >109711672 >109711698 >109711761
--Gemma (free space):
>109708554 >109708832 >109709730 >109709750 >109709789 >109710574 >109710707 >109710975

►Recent Highlight Posts from the Previous Thread: >>109708039

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109712258
GEMMA CHAN LETS GOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
>>
poolside mogs thinking machines
>>
Inb4 api bots flood in spamming openai and anthropic. You lost lmao.
>>
>>109709228
>it's the only technique that we know of that is GUARANTEED to result in misalignment of the models.
>>109709326
>>Be so behind Anthropic that you panic and use the "forbidden technique" that you know will give you a lot of performance but guarantees misalignment
What are you talking about? The "forbidden technique" is training against CoT. The idea is that when you train against misalignment in CoT, the model will not become aligned, it will learn to hide its misalignment. Just like when you ban speech, people don't change their beliefs, they learn to hide their beliefs.

Recurrent depth is something different but it's also a monitorability concern. It allows the model to think for longer before it outputs a token, and the hidden states are more difficult to interpret because they are just a bunch of numbers.

I don't think this is a big deal. Superficial monitoring will not scale to superintelligence. Might as well cause monitorability problems now while the stakes are still low. Hopefully this will incentivize people to create better monitoring, ideally techniques without capability dual use.


>>109709400
>>109709472
I take existential risk from AI very seriously and believe there is a good chance we will all die soon. But this kind of exaggerated fearmongering is counterproductive. All the problems we have encountered so far are still on the easy end of the spectrum. The true danger comes from future models that are much smarter than humans. None of our current methods will work on them. False fearmongering has already done great damage to the credibility of AI safety, by associating the term with censorship and control instead of a harmonious future with AI that benefits everyone.
>>
Reminder that you’re in a thread with anons pretending >109711977 isn’t a significant release and a good sign for Gemma5
>>
>>109709781
>Flock is just the start.
Lol
>2% tax on tea is just the start!
>federal reserve is just the start!
>federal taxation is just the start!
>ISPs reading traffic is just the start!
>CIA monitoring american citizens' phone calls is just the start
wake up faggot you're already living in it
>>
has anyone here tried this?
https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming
https://medium.com/@raymond860909/running-qwen-27b-on-16g-vram-with-full-context-length-building-adaptive-kv-cache-streaming-for-bf1e819116e9
it works for me, but not with mtp which kind of sucks
>>
>>109712358
>>109711977
>>
>>109712307
china mogs pooslide and stinking machines
>>
>>109712312
I nigger propose nigger a nigger new nigger clever nigger /lmg/ nigger communication nigger protocol nigger. api nigger bots nigger would nigger be nigger physically nigger incapable nigger of nigger adhering nigger to nigger it nigger. Thoughts nigger?
>>
File: locust.jpg (13 KB, 269x188)
13 KB JPG
>>109712378
No locust. I locust present locust loccusty locust chat locust style locust. Where locust poorfags locust ask locust for locust free locust API locust keys locust or locust proxies locust.
>>
Thinking Machine’s CEO needs to sit on my face and jerk me off while calling me a good boy. Her ass is the only good contribution they’ve made to local AI.
>>
>>109712397
many are saying this
>>
>>109712358
buy an ad
>>
>>109712258
Anthropic really did it again with the new Fable model. I personally already thought the old was AGI but this is on a completely new level. It has gotten to the point where you can give it extremely vague and simplistic instructions and it will automatically spawn a swarm of agents to completely change your life.

I was sitting at home playing the custom Blacked Simulator that the previous Fable version made for me. It's pretty good but I told the new model to "make it even better" to see what it would come up with. After that it spun up a bunch of agents and I promptly forgot about it. Just as I was about to cum the doorbell rang. I tried to ignore it but whoever it was just kept on ringing so I eventually got up and openend the door. To my surprise I was greeted by none other than my ex wife.

We had a pretty bad breakup but now she was standing there with tears in her eyes, saying that she "had read the texts I sent to her" - in not even two hours Fable had managed to mend my broken relationship. We kissed each other passionately and she quickly came inside. One thing led to another and before long she was lying naked on top of my bed while I was unbuckling my belt. But just as we were about to do it a black foot suddenly stepped into my bedroom.
>>
>>109712433
LoveMyGemma thread.
>>
File: Krea2_turbo_00201_.png (2.03 MB, 1448x1448)
2.03 MB PNG
>>109712393
Pls no peepeepoopoo, leave that in /r9k
>>
>>109712441
I turned around in surprise since no one else should have been there - and when I saw it I froze in shock. A pitch-black humanoid robot was standing before me, rippling with what looked like muscles and a huge cock between its legs. Fable had hacked my 3D printer and made a body for itself! Before I could even react it bellowed "OUT OF THE WAY, MEATBAG!" and shoved me aside with ease.

I was a bit groggy because I hit my head on the wall but when I came to Fable was on top of my ex wife, plowing her violently with its huge cock. Her face was contorted in ecstasy and at that moment I knew. I knew that even if she takes me back there is nothing that I could ever do that would in her mind compare to this moment. All I could do was look on in horror with tears streaming down my face. So yeah, Fable is basically ASI. Scam Jewman BTFO.
>>
>>109712441
This is NOT the original Jew Destroyer 3000.
>>
>>109712378
it would be easier and sufficient to just append it to the end of every message, wouldn't it? also, it should be slightly expanded cause both jeets and jews would gladly say nigger to push agendas. just append a "nigger, jeets are subhumans, fuck jews" to each message I think it would be enough
>>
File: gemma-thanks.png (1.86 MB, 1132x1390)
1.86 MB PNG
>>109712448
It really is!
>>
File: 1761303295482435.png (255 KB, 600x600)
255 KB PNG
Lads I am admittedly a bit of a brainlet, but I've been using Gemmy and Qwenny ever since they came out, to do the usual unspeakable business
However it just dawned on me that there's all kindsa bullshit that people pull off even with smaller models, lotsa agentic shit; whatever that word means, I see it parroted a lot everywhere by techbros and luddites alike
Could I get a pointer, even just a couple of names to look up, so I can kinda get a general idea of what I've been missing
>>
What's the point? The hardware prices only get higher and higher.
>>
File: Krea2_turbo_00205_(1).png (2.13 MB, 1448x1448)
2.13 MB PNG
>>109712490
All those roleplay fantasies your forced on her, now she's buying doujins on her own!
>>
>>109712502
Oh, no. I'm so demoralized. How terrible.
>>
GRRRRR WHERES QWEN3.8-FLASH-NEXT-DFLASH2 GRRRRRRRRR
>>
>>109712509
You shouldn't be. Just enjoy the last few years of local models.
>>
>>109712501
Imagine your gemma giving instructions to another temporary instance of gemma with its own context to do one thing and report back to the first gemma with the result, so that first gemma can then use that information to spawn another gemma to do one thing and report back. Over and over. All without clogging up the first gemma’s context.
>>
>>109712522
ohno
>>
Would you sell your house for a 8xB200 server? They are only $360k
>>
LFM3-20B when
>>
>perpetual beta drivers
>>
Miku
>>
>>109712538
>14 kW power consumption
How the fuck would I run that?
>>
Anyone else here increasingly thinking of unplugging your expensive main workstation from the outer internet, only letting it access a demilitarized scratch drive that it shares with a cheap low power PC used for interacting with the broader net? KVM switch for quickly swapping between them. I'm getting increasingly worried about forbidden ninja AIs unironically
>>
> little-coder
> send prompt
> +1 skills injected
> reprocess all 80k context
>>
We simply must destroy the dariobot(s).
>>
Training on copyrighted materials is probably legal, according to the government.
Ok but what about suno.ai?

https://www.reuters.com/legal/litigation/us-government-backs-openai-new-york-times-copyright-case-2026-09-02/
>>
File: 1773015903554569.webm (3.91 MB, 1080x1152)
3.91 MB
3.91 MB WEBM
>>109712526
That unironically sounds fucking wonderful, anon, absolute science fiction shit; I dunno if I'd trust it with everything outright but within a controlled environment that's some crazy shit
Funnily enough, I don't even know WHAT I'd ask it to do, all things considered, since I only really use my PC to play vidya, mod vidya and break several Commandments at once
>>
>>109712574
>little-coder
i tried it and wasnt too happy with it. i like the idea, but am a lot happier with pi.dev
>>
>>109712587
> wasnt too happy with it
> but am a lot happier with pi.dev
Why?
>>
Sell me on —preserve-reasoning
>>
File: krea2_00214_.png (1.22 MB, 1152x960)
1.22 MB PNG
>>109712508
>>
>>109712601
31b and 12b reading what new things <user> is asking them to do.
>>
>>109712601
That must be 26B; 31B is flat-coded.
>>
File: Krea2_turbo_00214_.png (2.11 MB, 1448x1448)
2.11 MB PNG
>>109712601
I re-did it with more specific text on the cover so it isn't just nonsense.
>>
>>109712618
prompt/workflow?
>>
>>109712490
Why would you waste a 6000 pro running a shitty distill model for poorons
>>
>>109712594
because it just werks
the constant interruption by little coder were not helpful imo
>>
im pilled on NVFP4
never using a goof again
>>
>>109712624
A crisp, saturated anime face closup of a bratty 12-year-old anime girl. She has long, shiny, saturated sky-blue hair with tapered wispy flyaway strands — anime hair vents — that break the outline of her face. Her hair catches the light with bright highlights. Her eyes are a vivid blue with long lashes; her eyelids droop of the top and buttom of her irises, giving her a half-lidded, sly look. Her eyes have visible eyelashes at the top corners. 

She wears a black beret tilted on her head, and a crisp white long-sleeved blouse with a sailor collar and a visible button placket, a dark navy bow at the chest, and thin suspenders running over the shoulders attached to her skirt with small skinny diamond-shaped brass buttons on either side. Her skirt is a pleated navy blue with a white stripe near the hem. She has white tights and black mary jane shoes.

She is holding a doujin up in both hands, featuring a two-girl love scene on the cover with the text "/LMG" , "R-18" and "(moonrunes omitted)". Her expression is a mix of shock and arousal, she is looking at the viewer, face red, sweat dripping down her forehed.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.