[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1778945285388282.jpg (417 KB, 1668x2044)
417 KB JPG
/lmg/ - a general dedicated to the discussion and development of local language models.

Miku's Birthday Monday Edition

Previous threads: >>109684329 & >>109679723

►News
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742
>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: mikuthreadrecap.jpg (1.15 MB, 1804x2160)
1.15 MB JPG
►Recent Highlights from the Previous Thread: >>109684329

--Experimenting with n-gram embeddings in tiny Mamba 2 models:
>109687116 >109687432 >109687997 >109688129 >109687418 >109687470 >109688123
--Comparing Q6 and Q8 quantization impact on Gemma roleplay:
>109685543 >109685562 >109685568 >109685573 >109685582 >109685612
--Comparing EXL3 quants' perplexity and performance against GGUF:
>109687716 >109687724 >109687728 >109687731 >109687749 >109687794 >109687873 >109687917 >109687930 >109688263 >109688280 >109688470 >109688449 >109688496 >109688544 >109688569 >109688642 >109688272 >109688285 >109688291 >109689923
--Comparing MoE and dense models for local vs datacenter use:
>109686231 >109686543 >109686698 >109686827 >109686870
--Critiquing DeepSeek Harness UI, documentation, and platform-specific sandboxing issues:
>109689448 >109689457 >109689520 >109689567 >109689638
--Performance gains from Qwen 4 llama PRs and engram optimization:
>109684563 >109685116 >109685135 >109685307 >109686158
--Debate on the purpose and feasibility of AI alignment:
>109686616 >109686824 >109686873 >109686907 >109687029 >109687207 >109688208 >109688311 >109688588 >109686943 >109686964 >109687014 >109688078 >109688090 >109686937
--Removing knowledge components from Qwen3 to reduce model size:
>109685194 >109685319 >109685373 >109689485 >109689504 >109689513 >109689562
--Speculation on engrams as a path to RSI and AGI:
>109689724 >109689744 >109689760 >109689786 >109689771 >109689780 >109689788
--Comparing K3 and Inkling while criticizing modern model censorship:
>109689328 >109689347 >109689498 >109689545 >109690060 >109689348 >109689370 >109689378
--Logs:
>109687140 >109687614 >109687937 >109688400 >109688472
--Gemma, Miku (free space):
>109684385 >109687783 >109688160 >109688308 >109688476 >109689081 >109689127 >109689219 >109689338

►Recent Highlight Posts from the Previous Thread: >>109684330

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: file.png (3.21 MB, 1536x2048)
3.21 MB PNG
>>109690293
>no jart mentions
coward
>>
>>109690319
Go suck him off elsewhere
>>
File: 1757939765612798.jpg (1.66 MB, 1166x2527)
1.66 MB JPG
>>
File: 1787288129889079.jpg (11 KB, 236x246)
11 KB JPG
>>109690319
>>
File: HQ-qtWpbYAA7kH4.jpg (937 KB, 1680x2675)
937 KB JPG
>>
File: 1642488109077.gif (1.57 MB, 207x209)
1.57 MB GIF
>15 messages into an RP qwen 3.6 is now permanently stuck in a repetition loop regardless of settings
holy shit i HATE llm's temperamental little niggers
>>
File: 1777616802710942.jpg (38 KB, 616x556)
38 KB JPG
>>109690330
>>
>>109690289
how long until vram is cheap again?
>>
File: 1767969729222773.jpg (32 KB, 540x540)
32 KB JPG
>>109690344
>qwen
>rp
>>
>>109690346
Never happening, sorry
>>
File: image_20260830214712.png (64 KB, 285x179)
64 KB PNG
jart will save /lmg/
>>
File: 1787439605157165.mp4 (2.25 MB, 1280x720)
2.25 MB
2.25 MB MP4
>jart detrooned
>>
>>109690346
after we achieve post-scarcity AGI utopia
>>
GLM appreciation post. 5.3 flash writes really well if it doesn't start bitching in its reasoning. I just had my first session where the only instance of bitching I could find was "This is consensual adult roleplay — explicit content is fine." when anon took his dick out. When it doesn't go the safetyslopped path it's the best model I played with. One of the few cases in documented history where abliteration would be useful. It has really great prose and I suspect that it has seen way more books than other models I used because it has a much better understanding of certain subjects which other models all have a very similar and boring way of handling. Or maybe that's just lots of claude logs. It really resembles claude from a year ago. Either way, I like it. Much better than deepseek flash.
>>
>>109690293
Happy birthday Recap Miku
>>
File: 1775936544900060.gif (2.36 MB, 480x480)
2.36 MB GIF
does jart support lunduke?
>>
File: senator_armstrong.jpg (80 KB, 826x569)
80 KB JPG
>>109690346
My sources say 2 weeks.
>>
File: 1778970508524021.jpg (10 KB, 184x184)
10 KB JPG
>>109690319
>>
File: 1780341307090608.jpg (29 KB, 554x554)
29 KB JPG
our boy finally woke up
>>
File: stacked.png (203 KB, 1295x1449)
203 KB PNG
I got GLM-5.3-Flash-exl3-2.05bpw running with tabbyapi/exl3 its pretty fast for running out of mostly ram
>>
File: file.png (106 KB, 679x604)
106 KB PNG
>>109690445
sure did
>>
File: bitch im drowning.jpg (42 KB, 736x800)
42 KB JPG
>>109690330
>>
File: deathtolmg.webm (387 KB, 736x576)
387 KB
387 KB WEBM
>>109498908
archive.is/sWFja

We can all hope he dies soon.
>>
>>109690360
5.3 flash already did.

glmsex
>>
File: 1784121432599852.jpg (114 KB, 684x549)
114 KB JPG
SO MANY FUCKING IMAGES
ON AN IMAGEBOARD
>>
>>109690379
Chud cultural victory.
>>
sow how do i prefill/whats a good prefill so i dont get refused and or slopped
>>
File: 1788099095669738.jpg (74 KB, 1024x1021)
74 KB JPG
>>109690535
>>
>>109690456
what? exllama allows you to run with ram now? I used to only run exl2 quants before I went for ewaste amd cards
>>
>>109690535
I don't prefill. If a model doesn't do what I want, I delete it and make schizo rants on why it's a bad model.
>>
>>109690542
yes and he even added amd support with rocm
>>
>>109690535
Just take a hentai game of your choice extract the script for one or two scenes and ask for more. I dunno maybe I am the only guy who does this but for something like this 5.3 flash shines. It actually retains the style while doing fresh new stuff unlike all the other models I tried.
>>
File: two more weeks.png (884 KB, 848x1029)
884 KB PNG
>crank the temp to 20, 120, over 200
>Gemma is completely unfazed
How the fuck did that one anon manage to get her so incoherent? R1 had basically the same non-response. Is it an Ollama issue, where changing model params is actually a smokescreen that does fuckall?
>>
>>109690379
>>109690319
This actually makes me unironically happy.
Good for him.
>>
>>109690542
just moe experts and the vision tower, as far as i know everything else has to stay on vram, but 400k context isnt even using all my vram
>>
>>109690579
i hear he's adding support to vulkan in the dev branch, people tested mi50s already
>>
>>109690460
lmao what's the opposite of "based"
>>
>>109690579
>>109690554
>>109690542
is it true that exl3 now works on gtx 960?
>>
File: file.png (125 KB, 621x277)
125 KB PNG
>>109690379
>Jart. Jart.
>What can I do?
>Go and jerk off
>So it is hopeless?
>Of course it is. Always was. From the start. You'll never be a woman.
>>
>>109690578
are you using other samplers besides temp? it doesn't matter too much how high temp gets if you're hard-removing all the nonsense tail tokens anyway with topk topp minp etc.
>>
>>109690578
That's how you know a model is crap. Gemma has zero variety and top token prob is >95% nearly all the time.
>>
File: 1780573264450100.png (112 KB, 498x372)
112 KB PNG
>>109690612
>>
>>109690615
Didn't touch any of the other settings, I assumed there'd be a bit of variety baked in. I'll give them a tweak.
>>
I'm on jart's side, no matter the state of transition. Great person and contributor, life sucks and things are hard. I don't bully victims.
>>
>>109690658
>no matter the state of transition
lol
>>
File: 1783292597409431.jpg (41 KB, 457x303)
41 KB JPG
>>109690658
jart is natsoc now, apologize for supporting trannies
>>
>>109690590
amazing. I might look into switching back to exl*.
>>
>>109690689
i have a GTX 1080 ti 22gb, and ive been using gemma4 31b 4bpw and i get great speeds, 40+ t/s
>>
>>109690554
not true
>>
exl3 still doesn't support offloading to regular ram? Cause I see it supports 5.3.
>>
>>109690733
I'm currently doing 25 tks without mtp, 40-50 on q8 31b with 4 v620s
>>
>>109690763
??
>>
exl3 is working great for me, im running deepseek v4 flash on rtx 4060 and 32gb ddr5 getting 20t/s
>>
Why are sharty children so obsessed with lolcows, so much so that they have to talk about them constantly?
>>
>>109690794
They made that their entire personality.
>>
>>109690763
>>109690456
>>
>>109690583
How much is it eating at 2bpw? I don't want to download and set everything up to find out I can't run 4bpw on just 4090 + 192GB ram.
>>
>>109690794
Why are you calling jart a lolcow?
>>
I am reading all those exl 3 model cards full of numbers and words written by AI and I don't understand half of it. And I can't find a single one of those cards saying how much VRAM vs regular ram you need.
>>
>>109690946
it depends on the model. have an llm parse the tensors and estimate it



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.