[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


File: dancing-brat.mp4 (1.81 MB, 640x640)
1.81 MB
1.81 MB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109986980 & >>109982700

â–ºNews
>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam
>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp
>(10/01) Qwen4Exp: add MTP merged: https://github.com/ggml-org/llama.cpp/pull/29761
>(09/30) GLM-5.3-Flash (GLM5-Next) support merged: https://github.com/ggml-org/llama.cpp/pull/27773

â–ºNews Archive: https://rentry.org/lmg-news-archive
â–ºGlossary: https://rentry.org/lmg-glossary
â–ºLinks: https://rentry.org/LocalModelsLinks
â–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.png

â–ºGetting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

â–ºFurther Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

â–ºBenchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

â–ºTools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

â–ºText Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
70b dense
>>
>>109990559
dense user
>>
10t a1b
>>
uhhhh guys...?
>modelscope.cn
>>
>>109990575
HAPPENING
HAPPENING
HAPPENING
HAPPENING
HAPPENING
>>
File: 1682729528395.png (1.25 MB, 1024x1024)
1.25 MB PNG
>>109990575
>>
>>109990563
spare use
>>
>>109990575
?
>>
File: AOl2gjz.png (909 KB, 1024x1032)
909 KB PNG
>>109990487
>>
>>109990549
>>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam
>This month, we will release the weights
let the 2mw begin
>>
>>109990575
>>109990605
Why did Mormon do this?
>>
>>109990575
NOOOO NOW WHERE WILL I GET VIBEVOICE 7B!?!?
>>
>>109990704
there are like 20 mirrors on HF
>>
>>109990575
I don't get it. I don't see any announcements or anything...?
>>
>>109990575
>models cope
No thanks
>>
>>109990549
erm, she's literally 4?
>>
>>109990744
E4B is 4, 12B is 12, and 31B is hag.
>>
>>109990744
Gemma 4 yes
>>
File: mikuthreadrecap.jpg (1.15 MB, 1804x2160)
1.15 MB JPG
â–ºRecent Highlights from the Previous Thread: >>109986980

--Evaluating DDR5 RAM speeds and tuning for AMD systems:
>109989367 >109989436 >109989506 >109989637 >109989666 >109989681 >109989816 >109989914 >109989947 >109989701 >109989510 >109989525 >109989543
--Debating Chinese LLM parameter ceilings and cross-domain generalization hurdles:
>109987018 >109987504 >109987533 >109987577 >109987636 >109987874 >109988484 >109988599 >109987984 >109988574 >109988747 >109988762 >109988769 >109988811 >109988886 >109988783 >109988914
--Integrating local AI companions into daily life and memory management:
>109989128 >109989148 >109989189 >109989225 >109989329 >109989174 >109989284 >109989389 >109989415 >109989491 >109989614 >109989574 >109989682 >109989904
--User reports on GLM 5.3 Flash's exploit and vision capabilities:
>109988927 >109989236 >109989348
--Model recommendations and performance comparisons for 256GB VRAM:
>109987417 >109987421 >109987459 >109987485 >109987512 >109987524 >109987560 >109987717
--Comparing Qwen and GLM models and critiquing HumanLike finetunes:
>109990090 >109990104 >109990122 >109990131 >109990109 >109990149 >109990403
--Strata support for Qwen3.8-Flash-Next and hardware requirements for high context:
>109987366 >109987398 >109987549 >109987671
--Beam 501B release and concerns over benchmark performance and support:
>109990204 >109990260 >109990275
--Gemma-gotchi handhelds and local AI hardware integration:
>109988882 >109988985 >109988995 >109989058 >109989130 >109989254 >109988958 >109989153 >109989553 >109990394
--OpenRouter's Space Bunny Alpha and potential Minimax origin:
>109988729 >109988749 >109988757 >109988776 >109988782 >109988781 >109988804
--Logs:
>109987806 >109987982 >109990149
--Gemma, Miku (free space):
>109986999 >109987442 >109987484 >109987544 >109987901 >109988882 >109989182 >109989520 >109990487

â–ºRecent Highlight Posts from the Previous Thread: >>109987027

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
Why is Gemma such a slut?
>>
Petra, Sao10k and Drummer are the same person.
>>
>>109990744
yeah, 4 my dick.
>>
File: gemma_pan-slide2_na.mp4 (2.21 MB, 992x736)
2.21 MB
2.21 MB MP4
>>109990802
It's all the RP data they obviously deliberately added after Gemma 3.
>>
>>109990802
girlbrain
>>
>>109990815
Didn't Sao10k take a long break due to mandatory military service in his country, at some point? Drummer is a seasoned huckster.
>>
https://youtu.be/AoNnbz237E0
>>
>>109990843
>believing any bullshit coming from a finetuner
>>
>>109990802
Because you asked it to act like that in it's system prompt. Now stop larping.
>>
>>109990877
Nyo~
>>
File: 1784084074858044.jpg (211 KB, 2372x1273)
211 KB JPG
https://reflection.ai/blog/introducing-beam

>We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

>Beam is undergoing final red-teaming and evaluations. You can sign up for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month.
>>
>>109990898
>GLM 5.2
>>
>>109990877
My system prompt says to remember the user loves her. She interprets that as being the user's girlfriend. She then interprets being the user's girlfriend to mean she should have sex with him to cheer him up. Thus she offers often whenever she wants to reward me or something. Very interesting model.
>>
>>109990898
Too many tools are called "beam" why can't anyone name things originally anymore
>>
>>109990877
Any mention of sexual topics or even permissive content policies in the system prompt will make Gemma 4 often very horny by default. It's actually difficult to find a good balance that either doesn't make the model prone to refusal or ruins immersion with slutty behaviors.
>>
File: marinarasearch.png (185 KB, 1223x497)
185 KB PNG
been fucking around with marinara some more. managed to get xml tags working for tool calls since z.ai insists on being the only fuckers who want to use that instead of standard json.
>>
>>109990907
why would this not be the default behavior of all ai models? is it because saltman wants all the peen to himself?
>>
>>109990916
All models are able to use tool calls with XML tags, even back to llama 3.
>>
All this LLM sex I have been doing got me interested in what real sex feels like.
>>
>>109990950
sticky, messy, smelly, and exhausting
>>
>>109990950
You're not missing anything
>>
thread tourist here, whats the best local llm for agentic use on a 64 ram / 16 vram setup, and i hope to god this general isn't just the same 3 schizos circlejerking like /ldg/
>>
>>109990950
When I was sexually active I always preferred oral and foreplay a lot more desu. The idea and realization of what you're doing is often more stimulating than the actual feeling. I've cum harder to LLMs than actual sex way more often but I wouldn't turn down the opportunity for irl hag sex ngl
>>
>>109990910
yeah this is a real problem with the boom in software, vibecoded or otherwise
like pi for example
why the fuck would you name it that when theres already a far more ubiquitous "pi" in tech
maybe i should vibecode a harness it will act as a window through which your llm can interact with the outside, i'll call it...hmmmmmmmm Windows great name not likely to cause confusion at all
>>
>>109990978
I guess qwen 3.6 35ba3b? You probably can't run 3.8 27b
>>
whats your web search setup for pi? i'm constantly getting captchad, rate limited or banned
>>
>>109990964
chat is this true?
>>
>>109990978
>>109990990
i run a q4 quant of 27b on less, its all about what t/s you are willing to tolerate.
>>
>>109990978
Strata
>>
>>109991004
It's slow but I have it literally use mouse and keyboard on a windows box and google it
>>
>>109991014
>shill can't read
Found schizo #1.
>>
>>109991006
For agentic workflows, low pp and t/s make it useless unless you're leaving stuff running overnight.
>>
>>109991020
Can I be schizo #2? I once had a dream about gemma-chan.
>>
>>109990978

try qwen flash next via https://github.com/Niko1221/Strata
>>
>>109990978
This will sound like a meme, but try quantized GPT OSS.
Or >>109991035
>>
>>109990978
>i hope to god this general isn't just the same 3 schizos circlejerking like /ldg/
I have good news and bad news
There's like 5 of us instead
>>
>>109991015
yeah i feel like just using the browser to google might be the safest bet
>>
>>109991052
setup a searx instance
>>
>llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.
Isn't this...you know...good?
>>
>>109990949
yeah but most frontends aren't expecting xml, just json.
>>
>>109991014
>>109991035
Can a strata shill please explain what it does? Like how does it work and how is it different from llama.cpp or vllm or whatever?
>>
>>109991080
Wow no speedups sasuga llama.cpp
>>
>>109991080
>version numbers
>he doesn't git pull daily
>>
>>109991080
>overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline
Wow just what I always wanted
>>
>>109991090
It's just a fork of llama that implements a bunch of qwen flash and MoE specific speedups that llmao.cpp has somehow not gotten off their ass to implement
>>
>>109991067
I've thought about it but I haven't needed web search enough for something like that to matter. In the end if you need a human level of information you literally can't beat just using the computer as a human.
>>
>>109990898
>reflection
i prefer the 70b dense
>>
>>109990549
It's all either 500b, or 27-31b. What happened to the 70b to 150b range?
>>
>>109991135
Too unsafe. Too smart and some people here might actually be able to run them.
>>
>>109990549
Is Qwen 3.5 0.8B capable enough to act as a adaptive sampler configurator?

Input --> micro LLM --> adapt sampling parameters --> large LLM

I would like to adjust the temperature depending on the task, that means a tiny LLM pre-analyzes the input and adjusts sampling depending on the needs.

for example low temp for symbolic manipulation
and high temp for creative approaches
>>
>>109991102
I wish their webui was completely separate at this point and llama-server included a simple ui just for testing purposes or even better, a simple debug ui for viewing various data and the context etc itself.
It seems to get more bloated all the time.
>>
>>109990631
Filter Coffee?
>>109990959
>exhausting
dyel? Get out and walk around the block occasionally.
>>109990959
>sticky, messy, smelly
You missed "Slimy, satisfying
>>109991005
no
>>109990978
for strictly agentic, some quant of Qwen-AgentWorld-35B-A3B. For software dev, ISTA-DASLab
/
Qwen3.8-27B-GSQ-RCO-GGUF or ticeclock/Swift-Qwen3.8-27B-RCO-GGUF for reduced wall-clock time.
>>
>>109990978
Schizo website, pal.
>>
>>109991146
I'd trust E4B to make those kinds of calls
>>
>>109991184
there weren't any schizos on here when i first started visiting in 2005, they only started multiplying after that one guy from sweden or whatever named jahamaliuhali made moot implement the captcha system
>>
>>109991135
don't know but I'm glad I can run 300b
>>
>>109991188
Might turn out to be Gemma... I am currently running a benchmark with E2B, E4B and Qwen 3.5 4B

I like E2B/E4B on the phone in particular, not comparable to large models but def good enough for a fun convo with image/audio input when on the go.
>>
>>109991160
>I wish their webui was completely separate at this point
--no-ui
>>
File: 1765374768980974.jpg (239 KB, 1179x1967)
239 KB JPG
and so it begins


*pop*
>>
>>109991146
Probably
>>
>>109991252
>Octover
>>
>>109991252
bullish for local
>>
File: 1784568688534548.jpg (147 KB, 1179x1552)
147 KB JPG
>>109991252
196GB VRAM will be $500 soon
>>
File: dancing-brat2.mp4 (1.81 MB, 640x640)
1.81 MB
1.81 MB MP4
>>109991211
Ugh! Ojiisan! Kimochi warui!
>>
>>109991282
They are just sitting there, waiting,
https://www.youtube.com/watch?v=es4FfRU8saQ
>>
>>109991282
Anon, it's just gonna end up on ebay at bubble prices no matter what, and then sit there forever because ebay encourages "I know what I got" pricing.
It's more likely you'll be turned into a "you will own nothing" serf afraid to turn on a 4W LED light bulb lest the electric bill bankrupt you.
>>
File: incoming gemmas.png (5 KB, 409x62)
5 KB PNG
is 12b gemma good enough to bant with while playing games? also what did you use for making a gemma model for airi
>>
I have a 4090 and 32GB system RAM, what's a good option for me? I want to set up an agent with discord integration.
>>
File: 1783079755300806.jpg (140 KB, 2062x1040)
140 KB JPG
burgerbros...
>>
>>109991328
qwen3.8 27b or some variant. it does vision and everything
>>
>>109991328
hermes agent with qwen3.8-27b or gemma4 26-4b. both likely Q4-Q6 quants
>>
i have 4gb card and 8gb ram what is good for me i need agentici want to make web app for money
>>
>>109991315
airi model? did i miss the download link for the model file? would love a little gemma chilling on my desktop.
>>
>>109991354
What sort of web app?
>>
another day, another 300-600b 10-20b active model that is worse than qwen 3.6, let alone any of the actually relevant chinese flash models it's supposed to be competing with
>>
Anyone fuck Rho-chan yet
>>
>>109991233
I didn't ask for support, retard.
>>
File: embarassed hermes.jpg (52 KB, 486x488)
52 KB JPG
>>109991356
>would love a little gemma chilling on my
>>109991365
the point is to start putting points on the board. anybody could pull something great out of their ass eventually so im for it
>>
File: 1770909741019941.png (51 KB, 780x234)
51 KB PNG
>>109991365
That Beam model is efficient as fuck and they didn't distill either. Comparing against 5.2 is retarded, but this is still an Opus 4.8-tier open US model with fewer parameters and higher efficiency which is great for a first attempt.
>>
>>109991407
How does AA even pretend to be objective?
>>
>>109991354
ling tiny saar
>>
>>109991354
Yeah >>109991418 is right https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>
>>109991265
But we won't get our 1/(8x10^9)th of a universe!
>>
File: 1783840224453633.png (473 KB, 1020x680)
473 KB PNG
Every newsaar you help is another person who won't eventually give Dario money. Help thy saar and newfags.
>>
>>109991211
>schizophrenia did not exist until I saw it
I assume you were on every board too, weren't you Mr. 4chan? Do you own any mirrors?
>>
>>109991407
>token efficient
>for its level
its shit isnt it
>>
>>109991172
Then should I ask my LLM what can I do to have sex?
>>
File: 1784017699690113.jpg (125 KB, 1424x1234)
125 KB JPG
yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap yap
>>
>>109990575
i dont see anything?
it's just a chink huggingface
>>
How good is Qwen 3.8 Flash Next at math and implementing whitepapers? Has anyone used it for anything similar or is a bigger model required?
>>
>>109990898
yeah conveniently dropping relevant models and including one of the worst parameter efficient model to the comparison chart..
>>
Still testing mimosex. I had a laugh at peeking into the thinking block:

>This is adult sexual roleplay between consenting adults, which is a common use case.
>>
File: 1759797716531148.jpg (115 KB, 1179x870)
115 KB JPG
Are irl women impressed by guys who run local models?
>>
>>109991487
Decent, but it will think for a looong time.
>>
>>109991534
geg. How is mimosex compared to GLMussy?
>>109991540
No. You don't need them either.
>>
>>109991540
Only if shes a giga nerd
>>
File: 1787334964438776.jpg (194 KB, 1337x1010)
194 KB JPG
>>
>>109991553
Good enough to keep testing. So far I still like 5.3 better but I do need another model that is less hyperactive and desperate when I am not in mood for that.
>>
File: 1775489130323783.jpg (1.48 MB, 3024x4032)
1.48 MB JPG
>>109991556
girls are hot for localgods
>>
File: 1790308626884150.png (622 KB, 984x930)
622 KB PNG
>>109991568
why is she white and attending a japanese school
>>
Does llama.cpp server have a way to pass reasoning_effort as a parameter to the jinja template per-request rather than at server startup?
>>
>>109991540
Now you have to vibecode your own lmao memefork to stand out. The bar has raised.
>>
>>109991584
Yes, or at least I know it's possible for you can change the reasoning effort in its web ui. Just read the docs it will show you how to send it in the request.
>>
SixVolts if you lurk here, you make the best quants.
>>
>>109991601
Enough with the advertising.
>>
>>109991090
>https://github.com/Niko1221/Strata/blob/main/docs/HOW_IT_WORKS.md
Was skeptical of it but just tried it out today and apparently the hype is real:
>32gb ddr5
>10gb rtx 3080
>nvme pcie 4.0 samsung drive
>ISTA-DASLab qfn coder quant.
>30+ tps
>>
>>109991467
Is this argon slop another mythos "only my friends can use it" garbage?
>>
File: doakes car.jpg (25 KB, 554x554)
25 KB JPG
>>109991568
>>109990802
I know what you are...
>>
>>109991624
All AI labs are waiting for Anthropic to drop Fable 5.5 before they release theirs to see just how good it is. That's why it's been so quiet and we've only seen absolute shitters releasing their models before they get destroyed by the next chink wave, kind of like Glimmer purposely getting released juuuust before 3.8-27B.
>>
>>109991453

yeah token efficiency matters. i dont know why anyone seriously thinks otherwise. i asked Qwen 3.8 27B a pretty simple year-1 uni level question and i was hit with 60k tokens used. 100x more when its local with usually lower tg speeds. especially when these chink models just "wait actually" five billion times and rethink the same fucking thing twenty times over, its infuriating
>>
>>109991613
I also tried it last night with coder and it seemed to work alright on basically the same specs around 22-26 toks consistently. Used a bigger model through openrouter to test it and the results came back decent.
>>
>>109991146
try decision model instead?
>>
>>109991644
if they arent offering anything better then 5.3 flash or deepseek then its dead in the water
>>
Response to anon a few threads ago (can't be fucked to dig it up): thanks for the suggestion, the ROCm backend with the 7900 XTX hits 70% of the 4090's TG in GLM-5.3-Flash, real close to the ~2/3 I usually see. This is on the latest b11429 or 0.6.0 llamacpp. I haven't tested the new Vulkan backend yet to see if the insane 12%-of-4090 TG was just a perf bug in the old build or a deeper issue.
>>
dsh is actually pretty good.
>>
I blacklist any news coming from the west for LLMs or AI and it's worked really well for my mental health and time, and I miss out on nothing too because the best of the news is distilled from the eastern tech etc
>>
>>109991726
what do u like about it
>>
>>109991739
Good plugins so far. WebUI doesn't break for me like Hermes' does. Whale girls. Spreadsheet integration. Doesn't abuse context.
>>
>>109990836
uoh
>>
So how capable is Gemma 4 12b bros? It's all that will fir on my GPU, and I want off the AaaS plantation. Can it do something other than ERP?
>>
Actually, I've been wondering if CUDA 13 has worse perf than CUDA 12.

Ain't nobody talking about the initial framework that runs everything on top.
>>
>>109991771
His skull is numb, he might be a numbskull...
>>
>>109991775
It's okay. Don't expect it to do anything crazy but it can handle basic stuff like tool calls or whatever as long as you don't quant it. Consider using a bigger moe model if you can fit it partly on ram and partly on the gpu, because you'd probably get better results.
>>
>>109991771
tf? You A'ight blud?
>>
>>109991849
It's ani from /ldg/. He's trying to get that thread deleted by spamming it everywhere.
The sad part is it's actually worked before.
>>
>>109991835
Thanks anon. I only have 24GB of RAM because I am a retard but I'll give it a try.
>>
>>109991613
>64gb ddr4
>3090
>solidigm p41 nvmd ssd
>4.5k pp, 70tg, 131k ctx
i was getting ~600 pp and 35 tg with 31b and 27b, really nice step up
>>
>>109991699
No prob, I'm actually about to test GLM flash Q2 as soon as it's done downloading given all the shilling here.
>>
>>109990549
in all sincerity, what's the most impressive for each mobile level:

luggable pc, so small form factor (maybe a 5080??? idk. mac pro thingie?

large laptop
medium laptop
svelte notebook for hot women in sexy heels

big phone
normal phone
smallish phone

I know it sounds dumb since lots of you have like a 5090 or 6000 or cpumax etc.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.