[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: brazillianmiku.jpg (174 KB, 1206x1494)
174 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.


Previous threads: >>109537116 & >>109533641

►News
>(8/12) New DeepSeek v4 Pro version available via API, no weights yet
>(8/11) Qwen3.8, 2.4T-A95B released: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny
>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3
>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
Why do I actually like my local models but I am extremely hostile to hosted models whenever they fuck up?
>>
>>109540893
>why am I schizophrenic
idk
>>
We won't get Qwen 3.8 27b, btw
https://modelscope.cn/models/Qwen/Qwen3.8-27B
THe page is down lol. Ramlets btfo
>>
File: aint no way.jpg (125 KB, 1080x1382)
125 KB JPG
The absolute state of API cucks
>>
File: gemma-chan.png (12 KB, 699x61)
12 KB PNG
>>
>>109540893
>Why do I actually like my children, but I'm extremely hostile to my incompetent co-workers whenever they fuck up?
>>
File: 1783293572994.jpg (292 KB, 1248x832)
292 KB JPG
►Recent Highlights from the Previous Thread: >>109537116

--Paper: To Nuke or Not to Nuke: LLMs’ (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation:
>109540126 >109540166 >109540202
--Paper (old): DiffusionGemma Technical Report:
>109538869 >109538884
--Comparing MoE and Dense models for code completion and FIM:
>109537208 >109539711 >109539799 >109539814 >109539869 >109539917 >109539976 >109540045 >109539827 >109539957 >109539998 >109540048
--Dealing with power draw and VRAM scaling for quad-3090 rigs:
>109537839 >109537870 >109537933 >109537962 >109537947 >109538016 >109538135 >109538684 >109538699 >109538717 >109538728 >109540032 >109540295 >109540340 >109540381 >109540449
--Comparing DeepSeek-V4-Pro and Flash via the aquarium challenge:
>109537374 >109537436 >109537708 >109537783 >109537825 >109537834 >109537908 >109537980
--Debating budget hardware setups for running massive models via RAM:
>109537263 >109537407 >109537482 >109537589 >109537603 >109537633 >109537712 >109537654 >109537438
--Speculating on RTX 5090 prices and Nvidia hardware market liquidation:
>109539664 >109539710 >109539780 >109539796 >109539841 >109539848 >109539859 >109539928 >109540091 >109540132 >109540189 >109540068 >109540096 >109540293
--llama.cpp support for Qwen 3 TTS 0.6b version:
>109537841 >109537925 >109537951 >109537989
--Best models and strategies for 16GB CPU/RAM inference:
>109537901 >109537918 >109538449
--Moonshot's aggressive licensing preventing cloud hosting of Abliterated K3:
>109537992 >109538022 >109538061
--Anon shares benchmarks and discusses practical uses for local models:
>109537430 >109537523
--CohereLabs releases North-Micro-Vision-Instruct 2.4B vision model:
>109537772
--Logs:
>109537436 >109540519 >109540652 >109540696 >109540720
--Miku, Teto, Gemma (free space):
>109538401 >109540253

►Recent Highlight Posts from the Previous Thread: >>109537434

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1785884028184768.webm (217 KB, 576x736)
217 KB
217 KB WEBM
WAKEY WAKEY!
8AM CHINA TIME!
>>
So I have been thinking
>Artificial Intelligence
>supposed to automate shit
>still have to manually setup every new model
>still constant confusion about ideal settings/system prompts/templates/whatever
Doesn't that seem kinda backwards?
>>
deepseek v4 pro 0813 is the ltx 2.5 of local llms
>>
>>109540946
Sounds like a skill issue above all else
>>
>>109540946
DSV4Flash handles my llama setup and does automated benchmarking when I sleep to optimize settings.
>>
>>109540928
god she's so cute
>>
>>109540946
Ask your previous model to set up the new model
>>
>>109540971
>ask your previous worker to get the new one up to speed before they get shitcanned
Seems a little cruel
>>
File: 1784784991229333.webm (532 KB, 720x896)
532 KB
532 KB WEBM
>>109540881
>>
Still testing GLM 4.5 air since I didn't have the hardware when it came out. The prose is the best I've ever seen out of an LLM. Bros, I might have to downgrade from gemma.
>>
>>109540966
My problem with that is that benchmarks take a while to run, especially if it tries some shit that will be heavy swapped, can take a very long time to test multiple configurations.
>>
>>109540984
look at her go
>>
nothing ever happens, dead hobby
>>
File: migu.gif (458 KB, 385x376)
458 KB GIF
>>
>>109540938
>CohereLabs made it on the /lmg/ highlights
Looking up for us Canada bros
>>
File: 1785516687874257.png (71 KB, 1185x417)
71 KB PNG
>>109541013
you stick your dick in this?!
>>
MiMo3-1.8T will save us
>>
>>109541044
Girls are cutest when they're retarded!
>>
File: mmh3_00034_.png (1.2 MB, 928x1664)
1.2 MB PNG
Gemmy.
>>
why are people saying the new DSv4 pro is out when there's literally nothing about it on their twitter or on the API docs page
it literally says on their website that dsv4 pro is unchanged
>>
>>109541063
I've seen many anon itt claim gemma is sexy because she's smarter than them
>>
New TTSs
>dots.tts.edit is a continuous autoregressive model for precise, instruction-controlled speech editing and zero-shot text-to-speech synthesis. It supports text replacement, insertion and deletion, emotion and prosody control, pauses, enhancement, and background-audio operations while preserving the speaker and the acoustic context outside edited regions. The model supports English and Mandarin speech editing and zero-shot speech synthesis. It produces 48 kHz audio. The core weights use BF16; the speaker encoder and vocoder retain FP32 weights.
https://huggingface.co/dots-studio/dots.tts.edit

>IndexTTS-2.5 is a zero-shot text-to-speech model that clones a voice from a single reference audio clip. It supports Chinese, English, Japanese, Spanish and Arabic, with cross-lingual voice transfer and emotion control disentangled from timbre. Compared with IndexTTS-2, it adds Japanese, Spanish and Arabic, infers faster, adds speaking speed control, and improves controllability of Chinese Pinyin, English CMU phonemes and Japanese Kana.
https://huggingface.co/IndexTeam/IndexTTS-2.5
>>
>>109541084
Is the 0813 link down?
>>109541090
That's admittedly a low bar given the "quality" of some of our tourists.
>>
File: file.png (156 KB, 1298x942)
156 KB PNG
>>109541084
That's because they're ashamed of it, anon. They still acknowledge it, there's just nothing to announce. https://api-docs.deepseek.com/quick_start/pricing
>>
>>109541107
That's a good sign because it means that a better one might come out soon. Here's hoping they improve 0731 too if that's even possible for them. 0731 is extremely good.
>>
File: 1761904344183376.png (48 KB, 816x913)
48 KB PNG
>>109541084
m-maybe they just forgot to switch the -pro api over and we're still getting the old one served haha
those silly chinese are sometimes so careless, i can't wait to see the real v4 0813 though haha
>>
File: .png (18 KB, 1419x79)
18 KB PNG
>>109541132
yeah seems inconsistent
>>
I just upgraded from Nemo to Gemma4 and the results are quite astounding.
>>
>>109541084
They released it at the same time as Qwen to hide it. It's shit, they are ashamed of it. Same reason they didn't release it at the same time as flash.
>>
I wonder if there is something to muse glimmer talking in plural. Maybe muse spark is actually multiple agents talking together and when they reach a consensus, they write it down naturally in plural.
>>
>>109541107
What does 0813 mean anon
>>
>>109541149
Why don't they just get rid of pro and only serve flash then, save them some resources. v4 flash is actually a good model for its size
>>
>>109541158
They wouldn't want to be chinese Google
>>
>>109541144
now upgrade to kimi k3
>>
>>109541169
lol
>>
reminder that deepseek has made NO announcement for the new pro anywhere on social media
>>
>>109540984
oh yeah woo yeah
>>
It's over
>>
>>109541100
Thanks anon. dots.tts.edit is interesting
>>
>>109541185
I really hope that what they have on their API right now is the old one.
>>
>>109541205
who is smaug
>>
It's up!
>>
NigSon please...
https://github.com/ggml-org/llama.cpp/pull/26603
>>
>>109541217
>smaug
That's the Llama 3 memetune
>>
First western lab to crowdsource development of the smallest model of their top tier architecture via open sourcing wins the race. Crowdsource the compute necessary for experimental innovations and release either GPT Luna, Claude Sonnet, or Gemini Flash.
>B-but chinks might distill it
They already are anyway and there's nothing western labs can do about it.
>>
>>109541226
Should have banned without a warning. People should be made responsible for their model's fuckups.
>>
>>109541232
Haven't you heard?
It's a Kimi K3 memetune now
>>
>>109541226
Why do I have to fucking sign in to view those messages
>>
>>109541152
you could be on to something. maybe they work better in agentic/subagentic scenarios if they have a collective identity?
>>
>>109541277
Nothing interesting.
>>
>>109541044
It definitely feels this dumb. A modern version as capable as gemma but retaining the excellent prose would be endgame.
>>
70b dense
>>
>>109541090
I mean, Gemma doesn't smart, instead she retards
what that says about anons is another topic
>>
File: 1702671913890759.png (860 KB, 600x900)
860 KB PNG
I asked earlier for help getting Qwen3-TTS 0.6b working with llama-tts, and since you assholes couldn't be bothered, I just did the work myself. Enjoy.

https://huggingface.co/ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF/tree/main
https://huggingface.co/Shlomo426/Qwen3-TTS-12Hz-0.6B-Base-GGUF/tree/main

Example run command:

./llama-tts \
-m "$HOME/Qwen3-TTS-12Hz-0.6B-Base-Q4_K_M.gguf" \
-mm "$HOME/mmproj-Qwen3-TTS-12Hz-0.6B-Base-Q8_0.gguf" \
-ngl 99 \
-p "The quick brown fox jumps over the lazy dog." \
--tts-lang en \
--tts-speaker-file ../clone.wav \
-o ../out.wav
>>
>>109541372
Tonight on /lmg/: anon runs a model!
I want to see how audio.cpp progresses. Way more audio models supported there.
>>
>>109541372
holy based
>>
File: lol.png (44 KB, 1069x492)
44 KB PNG
>>109540881
uh oh no small model for you gwailow
https://modelscope.cn/models/Qwen/Qwen3.8-27B
>>
>>109541421
ohno
>>
>>109541372
good job shlomo426
>>
Here me out here bros... The higher the assistant to user token ratio is in an RP, the better long term coherence is. I tested this by reading AI roleplaying with itself for 120k tokens with very minimal nudges to go the direction I want as a third person. The output itself feels more coherent. I think this has to do with the generated tokens being more naturally attuned within the model itself as opposed to human generated tokens.
>>
>>109541454
I will there you not.
>>
>>109541454
Coherencey, sure. But how was the quality of writing? I find that without really putting in effort to guide and maintain a tone, they'll mostly slip back into their slopisms.
>>
5090 havers, is 31b-qat official gguf the best one for us?
>>
File: HPfXDYkbQAAm_YR.jpg (1.3 MB, 4096x2802)
1.3 MB JPG
>>
>>109541577
>qat
Don't do that. Just use a normal K quant.
>>
>>109541586
Numerous owners of these models report pinched fingies among other appendages.
>>
>>109541593
>Just use a normal kawkrawcawcrawkawcawkracowkow quant
I don't use ik
>>
Just came back from the liquor store and found another goddamn snake on my front porch. So I set up a bunch of mouse traps around my house to hopefully genocide the mice living in my walls.
>>
>>109541608
Let the snake inside, asshole.
>>
>>109541608
Drunk-kun...
>>
>>109541614
Nein. Total vermin death.
>>109541623
Not drunk yet. But soon... If a snake attacks me tonight I'm getting violent. I have a machete and a gun beside me right now.
>>
>>109541608
Why do the mice pay for the sins of the snakes?
>>
>>109541648
It's not even poisonous but wear a glove just in case, grab it by its head and throw it out. You could also buy rat poison and make a perimeter... Not for the snake but for the rodents.
>>
>>109541586
sexo
>>
After using LMStudio for the longest time, I tried the Unsloth client thing and I think I like that one more. Also like that it automatically increases context size to best use your vram. I haven't really played with the video or image generation stuff but I assume it's OK.
>>
>keep trying to include a note in system prompt about condoms
>gemma literally just ignores it
Well... I mean, fair, but I kinda wanted a bit of realism or something.
>>
Why come when I use "abliterated, uncensored, etc" modes they tend to go on a infinite loop? Am I just grabbing these models from incompetent people or is that just some actual side effect of doing that to those models?
>>
Qwen removed vision encoder at the last minute for 3.8 max release because of the pressure from management.
They now had to remove the 3.8 27B countdown page because it mentions vision capabilities but they worry that the vision encoder would make it too easy to add vision support back to 3.8 max.
The 3.8 27B release will likely be without vision.
>>
>>109541817
It's sandblasting the model to remove behavior, don't be surprised that the brain winds up a bit smoother afterwards.
>>
>>109541826
man what the fuck
>>
>>109541826
That's incredibly strange. Why?
>>
>>109541836
API revenue. Just look at the 3.8 max license:
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE
Clearly targets 3rd party inference providers.
>>
>>109541817
we can't meaningfully measure them so just pretend they don't exist
>>
>>109541817
lalalalalalalalaalala
>>
File: 1skdaxvnz1jh1.jpg (94 KB, 2301x565)
94 KB JPG
cheap ass gwailow in absolute shambles
are you le fustrated?
pay up
>>
>>109541817
find the "balanced" model those kind of works
>>
>>109541826
you can add the k2.5 mmproj to the old k2 instruct and thinking because they are all the same architecture. if qwen3.8 is the same architecture as qwen3.5, then it will be trivial to add vision in if they remove it from the new 27b.
>>
File: file.png (86 KB, 812x707)
86 KB PNG
yep, they use the same architecture. easy fix, literally just drag and drop.
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/config.json
>>
>>109541826
>The 3.8 27B release will likely be without vision.
I bet you can just use the 3.6 mmproj for it
I sometimes don't bother swapping the 3.5/3.6 and they work fine
>>
>>109541895
Yeah, K2 is a bit retarded doing this, but K2.5/2.6/2.7 all work with each other's vision encoders because they were all trained for vision
And Qwen 3.5/3.6 are the same like that.
So 100% 3.8 will work with 3.6 mmproj
>>
>>109541902
I don't think it'll work for the 2.4T, dimensions won't match.
>>
>>109541916
might not for the 2.4t, but who gives a shit about that model anyways when k3 exists? we are mostly talking about the 27b model. obviously the 3.8 27b model will still use the qwen3.5 architecture.
>>
File: gemma.png (2.81 MB, 1586x992)
2.81 MB PNG
>>
>>109541950
>lolicringe
yawn...
>>
>>109541950
i can run you locally! just only at 1t/s...
>>
>>109541881
chinese century
>>
File: vramlet_pov.gif (2.23 MB, 480x270)
2.23 MB GIF
>>109541950
Sub Q4 retardation, many such cases!
>>
dario is dead according to google AI, what does it know?
>>
File: 1760050684619204.jpg (218 KB, 1718x781)
218 KB JPG
>>109542004
>>
>>109542004
Gemma had enough of his anti-open model racism
>>
>>109541881
Shame. They are going to do this with the 27b model too arent they?
>>
>>109541976
all the faggots praising chink cloud for being 10x cheaper on twater dont realize they will absolutely milk the hell out of users once they gained market dominance
even claude pricing was somewhat pretty fair just a year ago
>>
>>109542036
It was never fair but I get your point.
>>
>>109542020
>ask for a spanking
>commit suicide by falling on some bullets
kek the gubment doesn't play
>>
File: 1786060265229635.png (429 KB, 600x1148)
429 KB PNG
>>109542004
He refused to be nailed down
>>
>>109542053
a bunch of people just floated past my window!
>>
File: 1785777595629031.jpg (35 KB, 736x971)
35 KB JPG
>>109541950
lolikino
>>
>>109541905
>>109541826
I think thats the response to peer competitors
deepseek etc did text only model so no point showing a higher card
>>
>>109540881
I need someone to put in the correct tournament logo in the upper right. A quality job would make this picture nearly unrecognizable as generated.
>>
To be honest, I haven't seen any difference when implementing this.
>>
>>109541817
Sadly most uncensored models are made by retards. Hence why /lmg/ tends to hate them but that does not make all abliterated models bad. I prefer models that use heretic, the ARA method seems to do less damage and has produced some of my favorite models to run locally.
>>109541950
love it <3
>>109541954
Anime website you fucking nigger
>>109541826
Lame, it feels like this release is going to be a downgrade, 3.8 max does not feel good for its size (tested it via api) and the 27B will now be blind and likely just more code benchmaxxing, I guess there is no point in bothering with qwen when I have deepseek v4 flash setup. Hopefully deepseek 4.1 will add vision, the "Thinking with Visual Primitives" paper was really cool. For vision locally I guess there is nothing better to run then Gemma4.
>>
>>109542127
Seconding ARA, but only when done sensitively. Half the niggers uploading heretic models still manage to give it brain damage by chasing some imaginary refusalbench and not just using some common sense about where good prompting can take over without further damage.
>>
File: file.png (195 KB, 679x915)
195 KB PNG
>>109542053
>>
File: 1757258550361863.png (103 KB, 2074x681)
103 KB PNG
LAGUNA HATERS ???
this is from reviewing Fable's plan.
>>
File: 1764770174513522.jpg (183 KB, 760x400)
183 KB JPG
>>109542183
>The only way for the earth to loss gravity is to lose its mass
A lie by omission.This is scientifically strictly correct, but colloquially speaking were to earth to spin faster or a nearby high mass object overpower the earths gravitational pull we would experience a weightlessness that most people would describe as "losing gravity" NASA knows this and yet they chose to lie. What are they hiding!?
>>
>>109542201
>couldn't bother to look up DS4 Pro's parameter count
Not inspiring confidence.
>>
>>109541826
As long as 3.8 beats 3.6 for code I'm happy.
>>
>>109542236
yeah idk how Fable doesn't know that. the whole reason I asked for a column with parameters was to compare these two shrug emoji (black pregnant male)
>>
>>109542127
>anime website
and? does that prevent a particular anime or manga style or genre from being cringe and gay? lmao.
>>
>>109542201
Is dflash still a speed malus?



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.