[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: 1790695200186542.png (1.45 MB, 832x1216)
1.45 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109938458 & >>109934266

►News
>(09/30) IQuest-Q1, 320B-A15B for agentic coding and more: https://hf.co/IQuestLab/IQuest-Q1
>(09/26) koboldcpp-1.122 + bundled harness: https://github.com/LostRuins/koboldcpp/releases/tag/v1.122
>(09/26) exllamav3 v1.5.2 with Turing support, MiMoV2ForCausalLM support: https://github.com/turboderp-org/exllamav3/releases/tag/v1.5.2
>(09/25) MiMo-V2.6-RL training dataset released: https://hf.co/datasets/XiaomiMiMo/MiMo-V2.6-RL-oss

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
i love gemma-chan!
>>
>>109942697
>>>/a/
>>>/lgbt/
>>>/trash/
>>>/out/
>>
Based? >>109927620
>>
>>109942697
Somebody start selling Gemma doll plushes o algo. I would buy one
>>
>>109942720
go back
>>
You faggots are now recycling images of this shit.
Troon coded now
>>
>>109942729
I don't remember this one being used as the OP image.
>>
File: nocap.jpg (400 KB, 1536x1536)
400 KB JPG
►Recent Highlights from the Previous Thread: >>109938458

--Comparing Gemma 4 censorship and abliteration for roleplay and ERP:
>109938615 >109938627 >109939102 >109939142 >109939162 >109939184 >109939197 >109939210 >109939317 >109939242 >109939331 >109939493 >109939344 >109939435 >109940694 >109939175
--Debating Jensen Huang's views on distillation and the AI bubble:
>109938934 >109938979 >109939081 >109939193 >109939268 >109939318 >109939358 >109939545 >109939330 >109939084 >109939690
--GLM-5.3's exploitation capabilities and Nvidia's new rogue agent platform:
>109939639 >109939675 >109939688 >109939973 >109939813 >109940518 >109941137 >109942357
--Speculating on reasoning traces and text diffusion for future Gemma models:
>109940757 >109940781 >109940830 >109940866 >109940946
--Corporate AI monopolies using safety propaganda to stifle open-source competition:
>109939869 >109939891 >109940522 >109940590 >109939954
--IQuest-Q1 MoE model release and previous tokenizer concerns:
>109938544 >109938575
--Comparing llama.cpp performance against specialized model backends:
>109939691 >109939744 >109939842 >109939860 >109940244
--Comparing Qwen3.8-27B reasoning efficiency and recommending the Swift-1.5 derivative:
>109938580 >109938654 >109938743 >109938764 >109938610 >109938643 >109938661 >109938670 >109938611
--Evaluating an ESP32-S3 AI robot and modern MCU capabilities:
>109938564 >109940173 >109940437 >109940458 >109940546 >109940565 >109940595 >109940674 >109940642 >109941804
--Anon creates and beats a platformer using MiMo-V2.6-Pro and Intern-Decision-4B:
>109941945
--Logs:
>109939162 >109939184 >109939242 >109939331 >109939493 >109939631 >109940728 >109941524 >109942239
--Gemma, Dipsy, Teto (free space):
>109938577 >109938759 >109938768 >109938965 >109940331 >109942474

►Recent Highlight Posts from the Previous Thread: >>109938464

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109942697
Gemma is CUTE
>>
>>109942760
This highlight bot serves no purpose, this general has been stagnant and filled with offtopic dumb fucks not long after llama 3
>>
>>109942772
I use it to catch up whenever I miss threads and also collect new Gemma-chan art though
>>
>>109942776
AI "art" isn't art. Pick up a pencil
>>
>>109942780
or a wacom pen
>>
File: 1787244903292572.jpg (44 KB, 680x670)
44 KB JPG
>>109942780
>>
>>109942786
Digital "art" isn't art. Pencil means pencil.
>>
>>109942793
pencil lead isnt even real lead
>>
>>109942790
>>109942793
No arguments
>>
>>109942772
>filled with offtopic dumb fucks not long after llama 3
We had waves of tech illiterate retards coming here long before llama3
>>
File: apple-pencil-1740190123.jpg (272 KB, 2100x1398)
272 KB JPG
>>109942793
?
>>
>>109942697
You dare read that in front of me?
Rape
>>
File: gemma-doll3c.png (1.81 MB, 1025x1535)
1.81 MB PNG
>>109942720
They should make Gemma MDDs.
>>
>>109942807
holy fuck I need this
>>
>>109942804
Ok, iToddlers get an exception this one time.
>>
>>109942793
>Digital "art" isn't art. Pencil means pencil.
this was a real argument around the start of touch screen era.
>>
>>109942812
There was also the argument that using Photoshop wasn't art, anything digital wasn't art, using cameras wasn't art, it's an endless cycle.
>>
File: intel gpu.jpg (198 KB, 1707x960)
198 KB JPG
>480gb
Local is saved.
>>
>>109942825
and the price?
>>
>>109942825
the price?
>>
>>109942825
>LPDDR5x
The prompt processing:
https://www.youtube.com/watch?v=uD4izuDMUQA
>>
File: wetink.webm (3.83 MB, 440x800)
3.83 MB
3.83 MB WEBM
love me some Murati, simple as
>Inkling small runs at less than half the speed of Qwen 3.8 Flash Next on my machine
>>
>>109942825
Not mentioning bandwidth is telling...
>>
>>109942576
I am waiting for the first religion to embrace AI and say it is voice of god. Bonus points for doing some reinforcement learning on some model to make it align with religion.
>>
>>109942825
What is the price?
>>
>>109942835
No existing religion is going to undermine their own power structure like that and the Anthropic/OpenAI employees already have an AI cult.
>>
>>109942835
Reminded me of
https://huggingface.co/sleepdeprived3/Christian-Bible-Expert-v2.0-12B
https://huggingface.co/ReadyArt/Baptist-Christian-Bible-Expert-v1.1-24B

RIP Sleep, I miss him.
>>
>>109942841
2x less than nvidia would ask so probably 200k$
>>
File: 176754546633.jpg (31 KB, 337x593)
31 KB JPG
>>109942825
>$30K
Sure lmao
>>
>>109942719
It should have supposedly 1.5 TB/s bandwidth (with LPDDR5X-9600 memory and 1280-bit bus width), according to rumors. I doubt it will be cheap or even available to consumers.
https://chipsandcheese.com/p/hot-chips-2026-intels-crescent-island

>[...] Speaking of compute numbers, Intel also didn’t provide these numbers so if we run with the assumption that Intel will clock Crescent Island to approximately 2.5 GHz and that Intel quadrupled the rate of matrix operations while not changing the datatype ratios of matrix operations, we get the following approximate compute numbers:
>
> - 10.2 TFLOP/s FP64 vector
> - 20.5 TFLOP/s FP32 vector
> - 41 TFLOP/s FP16 vector
> - 328 TFLOP/s TF32 XMX
> - 655 TFLOP/s FP16/BF16 XMX
> - 1.3 PFLOP/s FP8 XMX
> - 2.6 PFLOP/s FP4/MXFP4 XMX
>
>This puts Crescent Island about 30% ahead of NVIDIA’s RTX PRO 6000 Blackwell for matrix compute and over 5x for FP64 operations while the RTX PRO 6000 has over 6x the FP32 compute and 3x the FP16 vector compute.
>>
>>109942830
>LPDDR5x
unlikely to be better than about 700GB/s no matter how they run it
>>
>>109942868
you don't need more for agentic coding, plenty fast enough and with that much vram you can run anything
>>
File: 1780514296365560.png (1.43 MB, 832x1216)
1.43 MB PNG
>>109942720
>>
>>109942883
I'll take two
>>
>>109942786
these are really useful as controlnet inputs
>>
File: 1775231388431238.png (71 KB, 1076x209)
71 KB PNG
>>
File: 1785428205190536.png (2.4 MB, 1024x1536)
2.4 MB PNG
>>109942884
shamelessly corpo version
>>
File: Krea2_turbo_01696_.png (976 KB, 1184x888)
976 KB PNG
>>
File: gemma_lewd2.png (203 KB, 2560x1339)
203 KB PNG
i'm starting to see why you guys enjoy this so much
>>
AAAAAAAAAAAAAAA MY PREFILL RATE FUCKING SUCKS
>>
How 'tarded would a IQ1_M of Glimmer be?
Let's find out.
>>
4k pp and 2k tg at 500k context. Thoughts?
>>
>>109943032
>4k pp and 2k tg
is that 4000tk/s or 4tk/s ?
Either ways the answer is "Insane" regardless.
>>
File: IMG_2637.gif (1.84 MB, 640x402)
1.84 MB GIF
>>109942807
>>
>>109943032
Running what model
>>
File: bake.jpg (2.35 MB, 1856x2270)
2.35 MB JPG
>>
>>109942760
Based historian.

Nice Teto. Nice skin texture.
>>
File: HQyNxAvbEAATF_F.jpg (207 KB, 2048x1538)
207 KB JPG
What would you name a fine-tune of Gemma4 31B? I have enough vram to do it from the base model BF16
>>
>>109943173
Rocinante.
>>
>>109942825
can wait to take a second mortgage to afford it
>>
>>109943212
>she has a house
let me live with you ._.
>>
>I'm embarrased to say, I haven't done any real profiling.
Gemmaballs...
>>
>>109942884
if i pull that hat off, will you die?
>>
>>109943173
nigger
>>
>>109941945
>a vidya made by an AI
>played by an AI
We're living in the future and it's CUTE
>>
Why is Anthropic advertising GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
>>
>>109943221
>if i pull that hat off, will you die?
she's not イか娘
>>
>>109943226
the jeet marketing psyops didn't think far it'd just promote them instead
>>109943234
retard
>>
>>109943221
it would be very painful



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.