[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: RinHousekeeper.jpg (343 KB, 1536x1024)
343 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109470008 & >>109466178

►News
>(08/04) Ling-3.0-flash 124B-A5.1B released: https://hf.co/inclusionAI/Ling-3.0-flash
>(08/03) NemotronLabs VoiceChat 11B released: https://hf.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
>(08/02) DeepseekV4 MTP + DSpark support merged: https://github.com/ggml-org/llama.cpp/pull/25784
>(07/31) LongCat-Flash-Lite-Sparse 69B-A3B released: https://hf.co/meituan-longcat/LongCat-Flash-Lite-Sparse

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: what's in the box.jpg (235 KB, 1536x1536)
235 KB JPG
►Recent Highlights from the Previous Thread: >>109470008

--Anons reacting to Google DeepMind leadership changes and spinoffs:
>109470297 >109470334 >109470361 >109470375 >109471807 >109470342 >109470347 >109470468 >109470526 >109473139 >109470522
--Benchmarking model adherence to character identity and user addressing:
>109473295 >109473360 >109473343 >109473364 >109473410 >109473377 >109473681
--Optimizing llama.cpp thread count using physical cores to avoid cache thrashing:
>109474612 >109474626 >109474654 >109474688 >109474664
--Anon shares NEUROFORM, a simulated digital organism with emergent behavior:
>109470839 >109470913 >109470905 >109470961 >109471155 >109471018 >109471043 >109471071 >109471076 >109471122
--Comparing vector RAG and third-person text parsing for agent memory:
>109470258 >109470612 >109470677 >109470735 >109470751 >109472835
--DeepSeek Flash demonstrating autonomous agent capabilities for complex tool orchestration:
>109473235 >109473264 >109473277 >109474253 >109474385 >109474401
--Critiquing a paper's claims on BBP-type digit extraction formulas:
>109471919 >109471974 >109472057 >109472134 >109472244
--Exploring alternatives to RAG for integrating long-term memory into weights:
>109470159 >109470177 >109470200 >109470205
--Finetuning LFM models for agentic search and roleplay applications:
>109471420 >109471519 >109472072 >109472156
--DeepSeek V4 Flash performance and efficiency compared to major models:
>109471436 >109471502
--Trump AI framework allegedly exempts open-weight models from safety tests:
>109470792 >109470809
--Anthropic AI using fake profiles for cyber-attacks:
>109471483 >109471491
--Logs:
>109470961 >109471869 >109472814 >109473405 >109473815 >109474036 >109474213
--Miku, Gemma (free space):
>109470097 >109470140 >109470225 >109470325 >109473269 >109473321 >109473395 >109472253 >109474608

►Recent Highlight Posts from the Previous Thread: >>109470228

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: HO064y3WUAAZHZn.jpg (484 KB, 2400x2400)
484 KB JPG
isnt 1.5X KV div like mid Q3 territory? wtf?
i have like 179 gb vram/ram so cant go q8.
did people try the q4 quants?
>>
>>109475598
I'm using bart's Q4, haven't tried any others.
Looks like I might as well do bf16 since it's only 20GB more
>>
>>109475598
>IQ2_XXS 78.5 GB
bros i think i can run this
but for real is it worth it? why even release 2-bit or 1-bit quants? what is the use case?
>>
>>109475598
>>109475623(me)
Hang on, they didn't test the unsloth 4_XL did they?
In theory that should match BF16 perfectly since they left some bf16 in there for the 2d tensors
>>
Zucc and Sama's return in open source will be glorious.
Any day now. Any day.
>>
>>109475647
Show it to gemma and ask her if they tested it.
>>
>>109475623
can you tell me how much deepseek models use for context? like if i wanna do 64k context. would that take like 10gb or more?

>>109475647
Pic related. Im sceptical of the quants because I did see people complain a couple days ago already how dumb it is with longer reasoning compared to like openrouter api.
Thats never a good sign.
>>
>>109475646
that's going to be awful
>>
>>109475598
>wikitext-2
yawn
>>
>>109475666
>666
yes but what's the use case for these lobotomized quants? who use them and for what? i'm just curious
>>
File: no.png (59 KB, 710x353)
59 KB PNG
>>109475662
You're absolutely right to push back on this. Gemma says they didn't test it (picrel)
>>109475663
>can you tell me how much deepseek models use for context? like if i wanna do 64k context. would that take like 10gb or more?
I'll check, I only use f16 KV.
>>
>>109475690
>yes but what's the use case for these lobotomized quants?
I use them for regression testing when I make changes to llama.cpp
I've seen other devs do this as well.
And some models, like the original deepseek were okay at IQ2_XXS
>>
>>109475700
thanks anon, appreciated.
>>
>>109475560
uoh?
>>
File: ss1785993139.png (17 KB, 825x157)
17 KB PNG
>>109475700
Your gemma might be broken.
>>
>>109475663
128k:
32k:
llama_init_from_model: KV self size = 1376.00 MiB, K-only (f16): 1376.00 MiB; independent V-cache: not used
64k:
llama_init_from_model: KV self size = 2752.00 MiB, K-only (f16): 2752.00 MiB; independent V-cache: not used
128k:
llama_init_from_model: KV self size = 5536.25 MiB, K-only (f16): 5536.25 MiB; independent V-cache: not used
Also got this little fucker:
ensure_dsv4_cache_tensors: DSV4 cache: CSA K= 677.25 MiB (f16), HCA K= 25.00 MiB (f16), LID K= 169.31 MiB (f16), states= 11.64 MiB, total= 883.20 MiB, streams=1
>>
>>109475722
Stop.
>>
>>109475763
You, and your Gemma are broken.
Do either of you have reasoning enabled?
Look carefully at the teal
Wait actually, looking at the table, there is a UD-Q4_K_XL
But the graph only shows UD-IQ4_XL/IQ4_NL
I need to be careful here
>>
>>109475801
I trimmed out the thought block with spoilers on where to find the mysterious information. It saw the box, and also the line pointing to its spot on the graph.
>>
File: image.png (2.16 MB, 960x1072)
2.16 MB PNG
What are the current best models, methods and llamacpp args to run on a system with 12gb vram (intel arc) + 64gb ram for coding and general use? The last time I used local llms, it were qwen and gemma (both dense and moe), while convrot and speculations were just being introduced.

Please.
>>
Do you chat with 31B with reasoning? It takes sooo long
>>
>>109475778
wut, thats tiny AF. thanks for taking the time to look it up! might be able to run q8. (maybe)
>>
File: intelarc.png (149 KB, 1537x380)
149 KB PNG
>>109475897
>What are the current best models
Models you can haven't changed since then.
>methods
./bin/llama-server
>llamacpp args
You have to test shit on your own system. Stop being a pussy and try stuff yourself.
>convrot and speculations were just being introduced
Now they're more thoroughly introduced.
>>
>>109476007
Did you write a tool to detect the faggot or just remember him?
>>
>>109476131
It sounded familiar. I just searched the archive.
>>
>>109475951
>wut, thats tiny AF. thanks for taking the time to look it up!
No trouble, I'd just finished a 3 hour tweak+benchmark session with it anyway.
>>
File: 1759013501580378.jpg (205 KB, 819x800)
205 KB JPG
>>109475560
does cpu matter that much? like a ryzen 9 vs an ultra 9, with the same specs. would any disparity be noticeable with a multi gpu setup?
>>
>>109476158
Only for agentic
>>
File: kimijester.png (163 KB, 853x811)
163 KB PNG
>>
>>109476158
If you run most of the model on GPU, the CPU is mostly irrelevant. Processing power (mostly, instruction sets) helps with prompt processing and memory bandwidth (ddr4/ddr5, memory channels) helps with token generation.
At that range, I doubt it makes much of a difference, specially if you run mostly on GPU. Just beware of e-cores on intel if you're going to run part of the model on CPU. They're slow.
>>
>>109476209
can we get an analysis of this guy >>109474606 from the last thread? need to hear kimi rip him apart.
>>
>>109476209
doesn't really work when there's only a couple dozen posts so far



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.