[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Miku's Birthday Monday Edition #3

Previous threads: >>109695101 & >>109690289

►News
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
►Recent Highlights from the Previous Thread: >>109695101

--Optimizing 27B model inference speed using adaptive KV streaming and llama.cpp flags:
>109697228 >109697243 >109697357 >109697427
--Anon releases AI-powered TCG playtesting harness using LLM endpoints:
>109698227 >109698273 >109698286
--Speculating on looped models and hybrid architectures:
>109695371 >109695403 >109695455
--Comparing DeepSeek-V4-Flash-Vision stability and performance against Qwen and GLM:
>109695473 >109695501 >109695549 >109695702 >109695740 >109695748
--Speculation on new Gemma models appearing on Arena:
>109697335 >109697359 >109697444 >109697470 >109697589
--Anon considers learning GPU reballing for hardware rescue and flipping:
>109695351 >109695363 >109695436 >109695452 >109695489 >109695532 >109695623 >109695640
--LM Studio announces Bionic local agents for Linux:
>109696525 >109696571 >109697707 >109697719 >109697758 >109697709 >109697933
--Valuing second-hand hardware builds for running larger LLMs:
>109697866 >109697876 >109697895 >109697951 >109697998 >109698070
--Comparing Apple unified memory to Nvidia for local LLM hardware:
>109695404 >109695414 >109695505
--Comparing high-capacity server RAM versus smaller efficient models:
>109698125 >109698252
--Mistral employee confirms new model still in development:
>109696527 >109696587
--Performance penalties from mmap'ing Qwen 3.8 Flash Next embedding tables:
>109698531 >109698947
--Speculating on which GLM-5.3-Flash llama.cpp PR gets merged:
>109696833
--Logs:
>109699177
--Miku, Teto (free space):
>109695135 >109695163 >109695200 >109696360 >109696745 >109696958 >109696981 >109696984 >109697079 >109697446

►Recent Highlight Posts from the Previous Thread: >>109695103

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: Return_00161_.jpg (1.23 MB, 1776x2368)
1.23 MB JPG
>>
File: 1784121432599852.jpg (114 KB, 684x549)
114 KB JPG
JUST SHUT THE FUCK UP
SHUT THE FUCK UP
>>
File: 1778134258890543.png (83 KB, 2800x3000)
83 KB PNG
Hermes Agent v0.21.0: The Pantheon Release
>>
>>109699243
glad i use q8
>>
>>109699266
>Bot mode
I can use this to jack off!
>>
>>109699243
women like this exist
>>
>>109699266
all that... for what exactly
>>
>>109699287
>>109699281
>>
>Not having a slight interest in harem mode
>>
>>109699292
someone really should make an RP agentic harness already
>>
>>109699266
usecase?
>>
>the year is 2037
>mistral releases their new mistral tiny 1.3T
>"just because you cant run it doesnt its not local, h-heh"
>shekels up on openrouter
what will they mean by this?
>>
>>109693005
>this is what I would use if I had the money right now, but I would also need the 3 threadripper GPUs too. that's out of budget right now, but I'm hoping to have that within budget in a few months. Maybe build two of them.
I've got this exact board. Don't buy it for 3 GPUs. Only 2 PCIe slots are x16. The second slot is x4 and the bottom two are x8
>>
>>109699322
>buzzword buzzword
>>
>>109699378
If by 2037 you don't have at least 16Tb of VRAM then you might as well kys right now.
>>
File: HQ_GFhzXgAA9Pd1.jpg (253 KB, 1600x1318)
253 KB JPG
do I really need qwen3.8 37b uncensored? I asked it to hax legitimate app (against TOS) to see whats inside and it just does it
>>
>>109699514
uncensored/abliterated models are mostly for indians and children who can't fathom prompting a model beyond "gib child sex" on zero context
>>
is it almost the end of the two weeks yet?
>>
>Hmm — hmm, but wait: hmm,
actually, hmm — hold on. Let me reconsider
>>
>>109699266
Stage IV LLM psychosis. Once it reaches this point, it's fatal nearly 100% of the time.
>>
> Hmm — hmm, hmm, hmm. Hmm, wait, hmm: actually hold on. Hmm, let me reconsider ONE more time
this has to be an implementation error right?
>>
File: hmm1.png (10 KB, 429x143)
10 KB PNG
>>109699574
Hmm!
>>
>>109699292
The true thinker.
>>
>>109699572
Hmmm, nyo~
>>
They really released Ornith in this state. It's the most useless model out there right now.
>>
my model is super broken
>>
>>109699646
based
>>
>>109699583
I just updated it. Don't scare me bro
>>
>>109699583
I wait like 1-2 weeks before updating hermes and somehow the project has always gotten thousands of commits since the last update. These guys are sloppin at the speed of light
>>
File: prewriter.png (377 KB, 1848x924)
377 KB PNG
I finally integrated Prose Rewriter in Orb. Also trained the latest versions to rewrite more and keep less. It works pretty well, pic related is rewritten Gemma 4 prose. 4B Q4_KM runs perfectly well in CPU, DDR4-3200 is more than enough, run in GPU for instant edits.
>>
I added dreaming to public.swiley.net/agent.py
>>
File: ComfyUI_Krea_2_00434_.jpg (2.59 MB, 2048x2048)
2.59 MB JPG
Initial D LoRA anon here. Some of you may remember me from https://desuarchive.org/g/thread/109043554/#109043922

Local music has undergone quite a few improvements since then I'd like to share. Here's a few generations from a Yousei Teikoku LoRA I trained with even more optimized settings and improved VAE (only present in these gens I'm sharing):
https://files.catbox.moe/inbl3t.flac
https://files.catbox.moe/o5gsag.flac
https://files.catbox.moe/m42nvw.flac
https://files.catbox.moe/zczuto.flac

I wrote an inference guide here that explains in more details how I got these results- https://rentry.co/fetcad7s

The settings are based on a discovery I made a while back while playing with ACE-Step 1.5 XL merges, managed to create Base/Turbo 0.3 merge which is significantly better than the 0.5 merge I had discussed then.

I also briefly discuss on the guide why I'm still not using Minimax Music. While its sound quality is better than ACEStep XL, the devs have not released the RVQ encoder needed for LoRAs (though some are working on reverse engineering it, that will take a while). ACEStep XL with optimized settings still sounds very close to Minimax Music sound quality wise (especially depending on seed), just not fully perfect yet with vocals, but I'm sure future ACE-Step developments will bridge the gap as it's very close.
>>
Ass schizo at it again
>>
La la la la la la la
>>
>109687005
What? No one likes moral chatbot alignment more than Dario.

You RL them to follow instructions and write good agentic tool calls. RLing morality is a waste of time at *absolute* best.
>>
>CPU improvements never get reviewed
>Daniel's implementations of models suck
>ik quants never ever
>ik could do something to quadruple CPU performance and niggernov wouldn't even consider adding it
>>>ROCm
Who else is making their own llama fork?
>>
>>109699852
I just bought a mac mini instead.
>>
>>109699852
i'm having gemmachan add in V100-specific optimizations for llamacpp :)
>>
>>109699778
>https://rentry.co/fetcad7s
thanks
>>
>>109699778
How do I make one of these DnB channels: https://youtu.be/g2ZwpFe8G5Q
>>
>no one using qwen next
what went wrong?
>>
>>109700061
it's not implemented correctly right now
>>
>>109700061
I used it. I appreciate the ngrams and it's surprisingly smart for a 6b active.
Not much else to say about it.
>>
>>109700032
>DnB
That is pure Suno I bet, or it could also be using Treblo (which is free). You could train a LoRA on ACEStep and achieve decent results too, but those guys farming content aren't going to use their local hardware for it.
>>
*invalidates your checkpoints*
Heh... nothing personal, kid.
>>
>>109700061
I'm using it, it's pretty good, I like it :)
>>
>>109700061
wat r next?
>>
This is a pretty sad question but I dount many people know the answer so I wanted to ask my anons:
When buying things from the Nvidia store, I know they say limit one, but if I managed to buy one and they restock/the gpu was in stock next week does the limit still apply? Or how intense are the checks? Different shipping address enough or different billing details needed?
>>
File: 1782778560534285.png (2.33 MB, 1536x1024)
2.33 MB PNG
>>109699572
Maybe
>>
>>109699778
thanks for the guide anon. i tried MiniMax Music 3 a while back but didn't have too much fun, I'll try acestep and merge(s) and see how it goes!



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.