[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109783258 & >>109779128

►News
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL
>(09/07) MiniCPM5-2B released: https://hf.co/openbmb/MiniCPM5-2B
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
►Recent Highlights from the Previous Thread: >>109783258

--Papers:
>109785800
--Comparing quantization trade-offs and KL divergence for Gemma and Qwen:
>109784998 >109785027 >109785059 >109785085 >109785110
--Critiquing wikitext KLD as a metric for quantization quality:
>109784100 >109784185 >109784907 >109784915 >109784926 >109784967 >109784990 >109784931
--Evaluating alternatives to wikitext for model divergence testing:
>109784941 >109785011 >109785034 >109785090
--Hardware struggle to fit DSV4.1 Flash on RTX Pro cards:
>109786996 >109786998 >109787013 >109787039 >109787045 >109787022 >109787071 >109787235 >109787253
--Estimating RAM requirements and PPL for Flash 4.1:
>109784896 >109784987 >109785153 >109785162
--North Small Translate MoE benchmarks compared to Gemma 4:
>109784615 >109784623 >109784634 >109784643
--Comparing model hallucination rates using Artificial Analysis benchmarks:
>109784326 >109784476 >109784526 >109784558
--Anon discusses high-VRAM AMD setups and hardware modding:
>109787572 >109787668 >109787672 >109787675
--Comparing DeepSeek v4 and GLM 5.3 Flash for 256GB RAM:
>109786528 >109786543 >109786574 >109786847 >109786984
--Using LLMs to rewrite and clean repetitive AI-generated webnovels:
>109784307 >109784324 >109784791 >109784860 >109784901
--Anon uses Gemma to control a Bluetooth pleasure device:
>109784694 >109784897 >109785007 >109785350 >109785400
--Hardware flexing and debate over SWE-2, a Kimi K3 fine-tune:
>109787445 >109787477 >109787489 >109787510 >109787534 >109787589 >109787598 >109787471 >109787617 >109787645 >109787667
--Logs:
>109786019
--Miku, Rin, Gemma, Teto, Rei (free space):
>109783474 >109784048 >109784073 >109784350 >109784658 >109784671

►Recent Highlight Posts from the Previous Thread: >>109783261

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Is K2 Horizon 375b any good for writing?
>>
>>109788183
>K2 Horizon
:'(
>>
>>109788173
Are you going to make a super recap once you get back?
>>
>>109787534
Wow that's one of those $250k servers they use for inference in data centers.
>>
See you in two miku weekus baker.
>>
>>109788234
NTA but did you try it? How is it for code? Better than Gemma and Ornith?
>>
I think it's time to setup filters. Instead of using 4chanx, I'll let Gemma-chan create user script for me. I'm so sick of this openai/anthropic judaism corporation spam. It's all over this website.
>>
>>109788258
Please don't reply to the spambot.
>>
Bitches don't know about my rdna2.
>>
>not using a userscript to filter every single post using your local llm
ngmi
>>
>>109788261
>Not having your agents read 4chan and hackernews for you to get the interesting bits.
>>
>>109788276
My lm said every post was trash and I don't need to waste my time.
>>
>>109788262
It wasn't even (me) >>109788183 responding to him. We have bots talking to bots now.
>>
On saturday I shall setup and try out YuE-2 and share results.
in the meantime enjoy suno v6 corposlop
>>>/wsg/6231753
>>
A potentially very gemmy model just dropped. Could be really good for RP.
https://persimmon.humansand.ai/blog/persimmon.html
>>
>>109788295
Wait holy shit this looks really cool. Gem alert
>>
>>109788173
>>109788182
Enjoy your vacation mate
>>
>>109788306
gemma lert haha
>>
>>109788295
cockbench results? gguf statua?
>>
>>109788295
>Persimmon is a model initialized from NVIDIA’s 550-billion parameter Nemotron 3 Ultra base model
>>
>>109788295
>persimmon turing test pass rate: 20%
>ast*a: 0.22%
>that one redditor itt pretending to be 20 different oldfags:
Nooo we completely demolished the turing test already it's impossible to discern humans from AI!
>>
>>109788331
Thank god I didn't waste my time clicking the ad.
>>
>>109788182
>>109788309
I was wondering if he was manually doing that or had some script. That's a crazy amount of consistancy.
>>
>>109788316
>cockbench results?
this is the rebuttle to cloudcucks
>>
>>109788336
it's dark miku 70b finetune and my own script, anon
>>
>>109788250
>>109788309
Thank you.
>>109788246
Maybe when I get near computer/internet access and my home server is still up, I might occassionally do like 2 recaps at a time, if no one else started doing it. I don't know how it will be so I don't want to promise anything.
When I get back there'll be like 30 missing recaps. There's no way to catch up without spamming and trying to condense two weeks in under 2000 characters would be silly. That's what the OP news is for.
>>
>>109788295
>mentions playground and api
>literally does not mention anything related to releasing the weights
>>
>>109788359
Of course, do you know how dangerous it would be if it were to be released with free access? Imagine the nefarious, dastardly deeds bad actors could get up to with this kind of model.
>>
>>109788359
where do you think we are, some generals where you can only talk about models you can download and run locally on your computer?
>>
>>109788380
that would be silly, who would even go to such a general when all the best models can't be downloaded or run locally?
>>
>17gb vram
>Mfw 16gb vram
T-thanks I guess
They do this shit on purpose
>>
Is there a dataset or a something that can essentially function as an offline web search MCP tool for an agent? I want the IQ boost while still keeping it sandboxed.
>>
>>109788428
Open zim format. Jewpedia English database with images is only 115 GB for example but I do think it's bit redundant with the current models. Gemma can reference esoteric literature without hallucinations, you need to gauge her though.
>>
>>109788428
Just give it it's own user account so it can use grep. If want to share things with it have it clone your git repo. Treat agents like employees, all the software and infrastructure is already there for this.
>>
File: 1781002743701499m.jpg (93 KB, 606x1024)
93 KB JPG
Just trying to wrap my head around all this shit, since I'm fairly new and have only really used local text models so far.

What do I actually need for a local RP setup? I've got sillytavern and koboldcpp working with text, but I'm interested in (uncensored) image gen and possibly voice gen as well. Do I just need to download some additional models for kobold to use for image/audio? If so, what do you guys recommend? My system has a 9070 XT with 32 GB RAM, running cachyos.
>>
>>109788466
Install ComfyUI for image gen then
https://docs.sillytavern.app/extensions/stable-diffusion/
https://docs.sillytavern.app/extensions/tts/
You're going to struggle to fit good models of all 3 on 16gb of AMD vram
>>
>>109788428
This actually got me thinking. It would be nice to have some /lmg/-contributed archive of data for when jario, jam, jloudflare and joogle cut us off. Something we can download for our gemmas to use. I keep getting jloudflared these days and can't web search shit anymore and I'm not touching jexa. Can we at least start compiling a list of resources?
>>
>>109788167
I vibecoded a network client in C++ with gemma. Already caught it not sanitizing inputs and a potential buffer overflow. How fucked am I?
>>
>>109788167
>DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B
lmao
Is it good? Does it know about /lmg/?
>>
>>109788564
Network client is a really vague term.
>>
>>109788564
>vibecoded
>c++
>gemma
ngmi
>>
File: 1760451888555422.jpg (664 KB, 3024x3024)
664 KB JPG
are we going to die? I can't bake for shit
>>
bitnet and matmul free
>>
>>109788596
>matmul free
The only good use for BitNet, otherwise anything properly trained under 4-bit precision is doomed.
>>
>>109788335
It's apparently continued pretraining + RL, so not just a simple finetune.

https://persimmon.humansand.ai/blog/persimmon-model-card.html
>Training approach:
>- Mid-training: Initialized from the Nemotron 550B base model and trained on diverse chat and forum data.
>- Post-training: Further trained from the mid-trained model to improve human conversation coherence and reduce failure modes.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.