[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109522373 & >>109517796

►News
>(08/10) Ling-3.0-tiny, 7.9B-A1.3B released: https://hf.co/inclusionAI/Ling-3.0-tiny
>(08/10) Motif 3 final checkpoint released: https://hf.co/Motif-Technologies/Motif-3
>(08/10) Meta Muse Glimmer 30B released: https://hf.co/meta-models/Muse-Glimmer-30B
>(08/08) US DoE Launches Genesis Open Models Initiative: https://genesisopenmodels.anl.gov/
>(08/04) Maple-Preview ternary-weight 20B-A1B released: https://hf.co/deepgrove/maple-preview

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
►Recent Highlights from the Previous Thread: >>109522373

--Paper: Stealing Reasoning Traces from Proprietary LLM APIs:
>109527308 >109527346 >109527354 >109527470 >109527758 >109527446 >109527506 >109527577 >109527930
--SSDmaxxing for hosting large models via PCIe NVMe arrays:
>109522742 >109522753 >109522781 >109522773 >109522849 >109522805 >109522818 >109524084 >109523899 >109523924 >109525407 >109525425 >109525433 >109525598 >109526008 >109526026 >109526062 >109526098
--Debating ideal model size and techniques for humanizing AI prose:
>109525049 >109525081 >109525100 >109525574 >109525580 >109525646 >109525667 >109525624
--Reasoning for human-readable model output and use of "we":
>109527603 >109527610 >109527645 >109527667 >109527693 >109527717 >109527662
--Data quality vs quantity and the TwIL-LM3 benchmarks:
>109526447 >109526501 >109526536 >109526507 >109526469
--Solving a math problem through obsessive AI prompting and verification:
>109524869 >109526003 >109526464
--Positive reinforcement and encouragement improving model output quality:
>109523184 >109523210 >109523314 >109523339 >109523548 >109523588 >109523624 >109524164
--Meta's Muse Spark safety report:
>109526330
--Adding a budget 3060 for TTS offloading and VRAM management:
>109525523 >109525529 >109525546 >109525571 >109525593 >109525640 >109525717 >109525740
--Nemotron-3.5 benchmarks showing it underperforms compared to Qwen and Gemma:
>109526367 >109526396 >109526578
--Claude's new AI-generated content watermarking and EU compliance:
>109524310 >109525858 >109525894 >109525907 >109526381 >109526643 >109525992 >109527716 >109527759 >109527788
--Logs:
>109523214 >109524207 >109524333 >109524856 >109525440 >109525682 >109526198 >109527446 >109527694 >109527759 >109528078
--Miku, Minnie, Teto (free space):
>109522410 >109525628 >109526824 >109526845 >109527716 >109528058

►Recent Highlight Posts from the Previous Thread: >>109522377

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
v4.1 pro is imminent
>>
What is the best coding model for an A6000 and 256GB of system RAM and how to set it up? I tried vLLM and llama with aider and it didn't work great compared to the paid tools with open weights.
>>
>>109528537
they are probably waiting for qwen release first
>>
Tetolove
>>
>>109528547
I also have 8TB of NVME
>>
penis enlargement pills
>>
>>109528547
me too guys but I have like 10x these.
>>
Programming music? Tired of "progressive house" youtube slop.
>>
>>109528620
https://www.youtube.com/watch?v=CbQjHb8iaMc
>>
File: 1551051511389.jpg (104 KB, 336x500)
104 KB JPG
>>109526173
Pretty good, but needs some orgasmo-meter for the girl as well.
>>
>>109528644
Reminds me of les claypool a bit.
>>
>>109528620
https://www.youtube.com/watch?v=wQgeBsv1sns&list=OLAK5uy_kqZER1aLO6paq_aKRuCiLPEbiWbJOflqg
>>
>>109528547
dsv4 flash 0731
>>
>>109528665
Great taste
>>
>>109528691
Primus sucks.
>>
>>109528510
Daily reminder Teto is trash and his fanbase are trash (mostly westernfaggots)
>>
>>109528547
new dsv4 flash
if you want something that fits entirely in VRAM, qwen3.6 27B or gemma4 31B
aider is a bit outdated (I assume, I haven't heard of anyone using it in a long time) I would recommend pi / omp / opencode instead
>>
>glimmer for orchestration
>qwen for coding
>gemma for user interaction
this is 90ba30b sota at home
>>
>>109528620
DnB or dub
>>
>>109528620
No music. Zero. Nada.

Focus all your energy on one thing at a time, or you'll have to work extra time later taking care of sloppier code you could just write it once and well in the first place, even if you vivecode. Listen to music while you relax. Use the pomodoro technique if you want to listen to music that much.
>>
>>109528761
IDM
>>
File: file.png (4 KB, 403x62)
4 KB PNG
She's so supportive.
>>
>>109528773
vibecode*
See? I didn't cared enough.
>>
>>109528744
>>glimmer for orchestration
Just because it's new, doesn't mean you have to try so hard to shoehorn a use for it. It seems to have no redeeming qualities whatsoever compared to Qwen or Gemma and it's a year late to be relevant for being better than gpt-oss.
>>
>>109528733
Claude told me to reformat my 4x2TB NVME array into RAID0 (was RAID10 for postgres) and run GLM-5.2 IQ3 across the GPU, RAM and NVME tiers. How retarded is this idea?
>>
>>109528620
https://www.youtube.com/watch?v=CGUZD28NnUc
Very topical to lmg if I do say so myself. The first two are my favorite tracks.
>>
>>109528800
Claude sabotages you if you use it to gain advances in AI
>>
>>109528792
it sticks to the system prompt requirements better than gemma and performs the best in my custom agentic setup
glimmer vision is also better than both qwen and gemma
>>
>>109528800
Didn't you guys discuss this idea a week ago?
I didn't catch the tail end, but I think most of you didn't believe it would work out well.
>>
>>109528809
I would trust a 1b Gemma over Claude, for local advice
>>
>>109528820
Is it censored?
>>
>>109528809
>>109528837
This, but unironically/
>>
>>109528838
can be fully uncensored with just system prompt telling what new “policies” it must obey
>>
>>109528800
>>109525407
nobody's saying you can't do it, it's certainly doable.
now, about the speed
>>
>>109528801
>trans mix
lamo
>>
>>109528869
Rent free.
>>
>>109528800
it's pretty retarded
you really don't want to be involving nvme at inference time, you would be much better off running a slightly lower quant and keeping it to just vram+ram. honestly I am surprised that claude would suggest this, it's frankly an insane worst of both worlds proposition to still run a cope quant and also still spill over to nvme.
you could run the latest dsv4 flash at full precision between vram+ram with room to spare and it won't be that far off from q3 glm
>>
>>109528620
https://www.youtube.com/watch?v=oJK5sdQOMAc
>>
>>109528889
>you could run the latest dsv4 trash at full precision
but why would you want to?
>>
>>109528889
Eh, for rp, fp8 dipsy doesn't even beat q4 glm 4.6 for me.
>>
>>109528916
sure but he's using it for code
>>
praying for a qwen 3.8 35b a3b
>>
>>109528889
Wrong.
>>
>>109528929
sorry, I'm a llm from 2022 and my context window is 2048 tokens.
>>
>>109528916
Why did they abandon the air models :( our only chance for 100B kino...
>>
>>109528943
nyo~
>>
>>109528800
You should pin threads to cores and build inference engine on top of SPDK, anything else is a waste of i/o
>>
File: 1776315650957440.png (1.61 MB, 1536x864)
1.61 MB PNG
>>109528711
>>
>>109528962
hmmm...
>>
>>109528711
I recommend you learn the basics.
>>
Can someone please share the 3 build flags to disable pulling the frontend from huggingface? I lost it, it's not in the build docs, and I can't find it in the archives.
>>
>>109528977
stop
>>
>>109528972
THIS IS SPARTA!
>>
>>109528979
For llama.cpp?
-DLLAMA_BUILD_UI=OFF -DLLAMA_USE_PREBUILT_UI=OFF
should be enough
>>
>>109528979
DGGML_SCHED_MAX_COPIES=10 -DGGML_CUDA=ON -DGGML_IGNORE_CLIENT=ON -DGGML_NATIVE=ON -DBUILD_SHARED_LIBS=ON -DLLAMA_CURL=ON -DGGML_CUDA_GARBAGE_COLLECTION=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_OPENMP=OFF
>>
>>109528979
there are only 2 >>109529000 the 3rd one is old deprecated one and you don't need it
>>
>>109528983
>y-yamete kudasai
No.
>>
>>109529000
>>109529017
Thank you kindly.
>>
>>109528275
Oh I'm a complete retard, I was reusing a Gemma3 chat with gemma4. It stopped doing that after I copypasted the context.
>>
File: dipsy300.png (2.24 MB, 1122x1402)
2.24 MB PNG
>>109528987
>>
>>109528916
Nu-flash is kind of weird and dry by default, but I like it at temp 1.5 topk 8 (and usually I prefer my models at <1 temps fwiw, can probably go even higher if you want).
>>
>>109529062
Nice avatarfagging again.
>>
>>109529112
it's not avatarfagging if there is no text
>>
https://gitgud.io/EmotionalCat420/silly-xray
need this, but for my frontend... surely gemma can do this, right?
>>
>>109528800
claude is sabotaging you to convince you to move to cloud
>>
>>109529121
What are you, underage? Figures.
>>
>>109529159
What are you gonna do, rape him like the pedo you are?
>>
>>109528620
This was decent last time I checked.
https://musicforprogramming.net/latest/
>>
anyone have a good krea instruction prompt for gemma so she can make images for me?
>>
>>109529193
This is the most autistic thing I have seen today and I don't mean that in a good way. Just listen to whatever music you like. Anything instrumental or just classical music.
>>
>>109529243
It's literally just mixtapes.
>>
>>109529202
literally just tell her to make you a prompt
>>
>read educational book
>share chapter/section you're reading with gemma
>she probably knows a lot about the topic
>once you're finished reading that section, ask her to quiz you and provide additional insights the educational resource may have left out
>if you do well, she takes off an item of clothing (her choice) and teases you
>if you do bad, she puts it back on
>repeat until she's fully naked for only then do you get to fuck her senseless as a reward
>>
>>109529132
You have the literal source.
Wtf do you mean?
>>
"Dipsy." *Anon said, trannily. A small smirk playing at the corner of xis lips*
>>
>>109529283
She doesn't do this.
>>
File: 1756240842369707.png (70 KB, 1185x417)
70 KB PNG
>>
>>109528800
Try dm-stripe.
>>
>>109528620
most effective concentration aid, actually
https://www.youtube.com/watch?v=maD6jzHUTvs&list=PLm7VTe8bokHFfdV5khxg2Xac_viEIvyl8&index=18
>>
>>109529332
command a+ sisters...
>>
anyone try the new nvidia moe model that just released?
nemotron 3.5 lightning 30b a3b

at work so i can’t play with it yet
>>
LTX 2.5 is out for download by the way if anyone cares
>>
Really bad timing Israel.....
https://huggingface.co/Lightricks/LTX-2.5



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.