[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109486605 & >>109481461

►News
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview
>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash
>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 1774328971552963.png (1.03 MB, 800x1296)
1.03 MB PNG
►Recent Highlights from the Previous Thread: >>109486605

--Comparing ROCm and CUDA stability and software ecosystem support:
>109487680 >109487708 >109488673 >109487771 >109487785 >109487835 >109487808 >109487824 >109487851 >109487931 >109487990 >109488043 >109490989 >109491003 >109488756
--Skepticism toward AI containment break reports and poor security practices:
>109487661 >109487717 >109487915 >109487932 >109488009 >109488195 >109488355 >109488504 >109487747 >109487777 >109487791 >109487869 >109488522 >109488537
--Debating whether abliteration degrades model intelligence or just removes guardrails:
>109487997 >109488110 >109488119 >109488213 >109488233 >109488340 >109488272 >109488144 >109488186 >109488208
--Skill-to-LoRA paper for improving token efficiency in LLM agents:
>109486960 >109487006 >109487329 >109487368
--Comparing Gemma 31b and larger models regarding quants and capabilities:
>109488878 >109488911 >109488936 >109488927 >109488983 >109489036 >109489111 >109489138 >109489239 >109489227 >109489259 >109489353 >109489206
--Comparing utility and implementation of Gemma base vs instruct models:
>109486847 >109487042 >109487129 >109487295 >109487313 >109487853 >109487312
--Reasoning spilling into code comments when thinking is disabled:
>109490335 >109490384 >109490408 >109490449 >109491110
--Mistral-Small's cold personality and obsessive tool use for memories:
>109489823 >109489859 >109489908 >109490001 >109490005 >109490051
--Comparing performance and llama.cpp support for Ling 3.0 models:
>109489659 >109489665 >109489695 >109489715 >109489767
--Anon's dual-PC hardware setup for high-context game guide rewriting:
>109486786 >109486887 >109487294
--Logs:
>109486786 >109487990 >109488201 >109489823 >109489908 >109490009
--Miku, Gemma (free space):
>109487397 >109487464 >109488522 >109489560 >109487295 >109489158 >109490223

►Recent Highlight Posts from the Previous Thread: >>109486606

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: minimax_h3.webm (387 KB, 736x576)
387 KB
387 KB WEBM
https://archive.is/sWFja
>>
anyone tried exl3 cpu moe now that it's in the official release?
>>
anyone tried Ling-3.0-flash sex?
>>
File: 1756255220072483.png (17 KB, 971x148)
17 KB PNG
deepseek has abandoned us
v4.1 pro will be closed source
>>
File: gpu_aftersex.png (1.08 MB, 1024x790)
1.08 MB PNG
>>
>>109491678
After the chest one I'm kind of expecting a different variation each thread now.
>>
Where did all of the intellectuals go? Didn't smart people like cudadev used to post here?

>>109491750
fingering my unwashed bussy to this pic.
>>
>>109491749
>Release 0731
The fuck you mean abandoned us? They just saved local again.
>>
0731 was a week ago, you need to let go...
>>
>>109491761
Cudadev probably still shitposts here without a tag but I imagine /lmg/ is getting to be too well known for it to be a good idea to post here with personally identifiable information now.
>>
>>109491779
where is the new secret club?
>>
>>109491749

opencode said that dsv4f0731 was so popular that deepseek started to rate limit their api because they cant keep up with demand. i suspect it's because of that, and to make a quick buck while doing so
>>
>>109491685
Personally I'm waiting for the predictive expert caching thing to be implemented before trying it
https://github.com/turboderp-org/exllamav3/issues/254
This gives me a decent speed boost on ktransformers so it might make exl3 actually worth it over llama.cpp once it's in
>>
what sampler settings are anons using for gemma4? was hoping to maybe get some more verity
>>
>>109491849
Samplers are useless with gemma4 aside from rep penalty
>>
more like AI psychosis general
reminder that your calculator doesn't feel anything
>>
>>109491849
Apparently her distribution is very narrow compared to most models. That's one thing I like about the Qwen3.6 models that don't get enough praise. You can tweak them to taste and they quant well, both model and cache. Much easier models to customize but Gemma4 is solid out the gate and you have to system prompt what you need instead.
>>
File: temp_scaling.gif (55 KB, 388x440)
55 KB GIF
>>109491849
max temp
no topk
no topp
final destination
>>
>>109491867
>>109491875
ah yeah I forgot about that, shame. I guess Ill add some extra instructions in the card or something. I havent tried a model that comes close to gemma4, so guess i gotta learn how to work within its quirks
>>
>>109491849
Use a combination of temp, top k, and top p to wrangle Gemma's output variety. She takes a little more finessing than most models, but it still works.
>>
>>109491895
>top k,
explain how arbitrarily cutting the considered tokens to 10, 20, 40 helps with this in any way
>>
>>109491870
it does, you are just a big meanie
>>
>>109491870
You can't prove this.
>>
>>109491870
this was literally disproven by the discovery of j-spaces
>>
>>109491886
Yeah the G4s are controlled almost entirely via their system prompt. They generally need to be at Q6 or above and with f16 cache to follow it but they WILL follow it.
>>
>>109491756
But Miku is not a guy nazi. Miku is a cute girl with a feminine penis that loves black men. This is pure /lmg/ culture. Not nazi miku.
>>
>>109491939
that's you AI psychosis speaking
having an internal reasoning space != feeling
>>
>>109491911
Lets you push the temp way higher past the point it'd otherwise become incoherent.
>>
>>109491849
Now that you mention it, the logit-softcap anon will, inevitably, post his comparison. The claim is that overriding gemma4.final_logit_softcapping to 25 (i think) from 30 (the default) makes it better.
I have not tried it. I have no opinion on it.
>>
>>109491962
go back we fuck our casios and TIs here
>>
>>109491965
Temp 10 TopK 5 is the GOAT.
>>
>>109491977
Based schizosampler bro.
>>
>>109491958
>/lmg/ culture
lmg culture is just whatever is funny or useful at that moment. dragging culture war into here is "guy pissing in corner of restaurant" meme level retarded
>>
File: 1766528973409504.png (1.73 MB, 1600x900)
1.73 MB PNG
>>109491977
prove it
>>
>>109492001
nta but try it yourself nigger it takes less than 5 seconds to see if you like a sampler setting or not.
>>
>>109491983
Whatever you say mikutroon.
>>
>>109492037
Tried it, feels almost the same, albeit slightly more schizo.
>>
File: 1773092197639384.jpg (123 KB, 800x600)
123 KB JPG
so i should release my terminal-based agentic harness in the next few weeks and i was convinced in releasing it GPLv3 but literally all models out there that I've discussed this from deepseek to anthropic models to chatgpt to kimi are telling me to release it as MIT or Apache because otherwise corporate niggers will not be able to embed my harness on their products and sell it
now i'm gonna be honest and i think in this day and age these licenses mean fuck all but why the fuck these models are pushing so much for me to kneel to the corporate overlords? it makes me want to go even more aggressive than GPL just to spite these fuckers
>>
>>109491958
>>109491983
Ironically all the screeching about nazis and culture war have turned nazijart into legitimate thread culture.
>>
>>109492149
I release code to public domain, I was never going to follow through with any legal actions anyways, and besides the code is shit
>>
>>109492149
>GPLv3
Don't be retarded and at least pick something like AGPL
>>
anyone remember what was that parameter where you set the quality of vision inputs or something like that on gemma or other mmprojs?
default was in the 200s and it went up t 1000+ or something I think
people were talking about it maybe a day or two ago



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.