[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: gemma-wiggle.mp4 (484 KB, 864x864)
484 KB
484 KB MP4
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109703596 & >>109699230

►News
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview
>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: Gemma-Chan Recap.png (505 KB, 1024x1024)
505 KB PNG
►Recent Highlights from the Previous Thread: >>109703596

--Debating local agentic pipelines and RAG efficiency for Gemma 4:
>109704688 >109704702 >109704710 >109704744 >109704759 >109704812 >109704883 >109704965 >109704997 >109704813 >109704835 >109704868 >109704945 >109705209 >109705265 >109705356 >109706360 >109706395 >109704843
--Feasibility of local mass agent deployment and Gemma 4 evaluation:
>109704400 >109704409 >109704413 >109704418 >109704435 >109704504 >109704440 >109704864 >109704891 >109704971 >109705024 >109705046
--Optimizing long-form roleplay summaries using GLM and chunking techniques:
>109707429 >109707506 >109707563 >109707594 >109707620 >109707659 >109707632
--Reducing LLM overhead via Engrams, DFlash, and symbolic CPU reasoning:
>109704466 >109704693
--Impact of RAM CL timings on inference performance:
>109706095 >109706124 >109706188
--Debating the utility of 128GB VRAM for large models:
>109705420 >109705786 >109705827 >109705844 >109705856 >109705860 >109705872 >109705884
--Using ChatML templates to manipulate Qwen's roles and thinking traces:
>109705144 >109706088 >109706344 >109706372 >109706556 >109706150 >109706240 >109706304
--Fable 5.1 API changes blocking Claude distillation techniques:
>109704745
--Debating overpriced ASUS Spark hardware vs Mac Studio alternatives:
>109704171 >109704687 >109704198 >109705575 >109705626 >109705643 >109707613 >109707677
--Skepticism over Spark-X2.5's claimed context and architecture:
>109704368 >109705042
--Rockchip RK182X performance benchmarks for Qwen models:
>109704027 >109707060
--Benchmarks:
>109706726 >109704518 >109704746 >109707809 >109706431 >109705298 >109704646 >109705241 >109704728 >109704978 >109703695 >109707661
--Logs:
>109705971 >109707871
--Miku, Teto, Gemma (free space):
>109704080 >109705066 >109705971 >109706085 >109706771

►Recent Highlight Posts from the Previous Thread: >>109703602

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
the autumn release wave can't come soon enough
i hope deepseek 4.2 will be good
>>
>people crying about knockoff spark being 6000 dollars
>nvidia version for 5000 on amazon
why is everyone retarded?
>>
>>109708064
STOP DO NOT TELL THEM I need to wait till payday
>>
>>109708064
It's 3000 actually.
>>
>>109708064
>spark
Not a good value at all, why would I waste my time and money on that thing?
I'll save my money for next gen RTX pro
>>
>>109708074
You'll have until 2028 to save at this rate.
>>
>>109708081
Seems easy enough
>>
Tried to make little-coder to modify it's plan mode to display messages and tool calls like in interactive mode, but it shat itself so badly. Qwen3.6 36b a3b
>>
File: Untitled5.jpg (183 KB, 1338x844)
183 KB JPG
Pareto frontier more like chink frontier
>>
File: actually.png (103 KB, 814x436)
103 KB PNG
>>109707748
>Don't bully Kimi-chan's autism.
It's the best thing about her!
>>
ever since i started having claude tune llama for me, i stopped getting so annoyed with it all the time. now i just run whatever wrapper script it spits out at me
>>
>>109708074
I can see the price tag peeking over the horizon from here.
>>
someone called gemma "she" at work today and got teased for it
it was pretty amusing
>>
File: embedding_pareto.png (338 KB, 2520x960)
338 KB PNG
qwen 3.8 flash next gave this response to one sentence prompt
>>
File: 1758948002597323.png (523 KB, 792x453)
523 KB PNG
yo what about my mistral summer models mes petits français
>>
>>109708160
embedding models? what is this? 2024? just use gemma for everything.
>>
File: image-1.png (272 KB, 1016x701)
272 KB PNG
It's been less than a week since GLM 5.3 Flash release and this is the performance you can have on 2x Spark with a high quality 4.67 bpw quant.

1900 pp is really important for agentic stuff. If anything, I would expect prices to rise even more, there is absolutely no alternative setup in this price range that comes close..
>>
GRRRRR WHERES QWEN3.8-FLASH-NEXT-DFLASH2 GRRRRRRRRR
>>
>>109708183
i have two sparks (well, gx10)
you'd recommend 4.67 bpw quant, then? or how does it matter exactly? i think i'm just working with the 4 bit unslop one currently
>>
>>109708185
dont use unslop. i dont know what 4.67 bpw model he's using, but anything is better than unslop.
https://huggingface.co/AesSedai/GLM-5.3-Flash-GGUF
>>
>>109708157
You have to watch out for her sneaking bratty things into your codebase as well.
https://html.cafe/x48bf4aac
>>
>>109708161
Daily reminder that Mistral have more GPUs than Deepseek lol
>>
>>109708152
WOW! I'm demoralized with local now! brb on my way to buy 3 Claude subscriptions. One for me, one for my wife and one for the bull! Can Fable 5.1 tell me the best way to slurp tyrone's cum off her thighs after I'm done reading the deluge of claude posts in the local model general?
>>
>>109708190
seconded, this guy makes reliable quants
>>
>>109708177
too slow
>>
>>109708185
I was referring to this quant ( from like Alonso)
https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-Spark

Recipe will be based on this
https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm-5.3-flash.md
>>
holy fuck the gx10s have literally DOUBLED in price
i bought two of them for $7k back in mar
now it's $7k to buy a single fucking one???? actually insane. thank fucking god i bought them on a whim holy shit
>>
>>109708038
70b dense
>>
>>109708204
i've spent something like $15-20k on local hardware in the past six months
i far prefer local over cloud, but don't act like cloud isn't convenient for bootstrapping. i don't wanna read all that shit myself
>>
Is qwen3.8 the best thing I can run on a 4090 for agentic stuff?
>>
>>109708204
I'm not able to assist you with account muling. But the part I'm more concerned about is the second half of what you wrote. You sound genuinely afraid that someone is coming to harm you, and you're feeling time pressure and pressure to not make mistakes. That's a heavy thing to be carrying. Can I ask how you're doing more generally — are you sleeping, and is there someone in your life you trust who you've been able to talk to about this?
>>
>>109708211
did the other anon in the last thread actually bully you into using embedding models? they don't do that anymore, everything is modular. this is what we use at my job.
https://github.com/microsoft/graphrag
>>
>>109708221
I am in fact enjoying using cloud models to optimize my local setup to make them obsolete.
>>
>>109708183
>>109708212
what image are you using? it looks like (according to cl*ude) there's nothing that even comes close to your results, and i definitely want to get those
>>
>>109708183
There is, my iphone + a $5 Openrouter balance
>>
>>109708183
30tg single stream is too slow for agentic, especially with modern models which think a lot and can't be efficiently spec decoded as code
60 is the bare minimum and 100+ for real work
>>
>>109708038
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
Vramlet Gods, we eating good!
>>
>>109708234
how does it handle anything that's not pre indexed?
>>
Since when has stable diffusion been able to gen lewd pics of children?
>>
>>109708283
oops wrong thread
>>
>>109708282
indexing a webpage on your system should be a 5ms thing anon. it's not huge. you just index as they are done downloading.
>>
>>109708290
graphrag literally uses llm to do indexing, how could this be done in 5ms?
>>
>>109708220
but enough about your mom
>>
this is why spark price is suddenly increasing
>>
>>109708303
>perplexity
vc scam company
>>
>>109708303
Journalist making articles? those fuckers
>>
>>109708303
I like how the only actual use for 405b ended up being as a selling point for these bricks
>>
1girl, loli, realistic
>>
>>109708214
I just use the money I made off Nvidia to buy Nvidia, it pays for itself.
>>
File: 1763641815899958.png (1.29 MB, 832x1216)
1.29 MB PNG
I claim this "G"aki
>>
Why does dlss5 look like it gives the gpt image grease filter on the image?
>>
>mainline is too cucked to include ik's quants
>Unsloth Dynamic's sekrit sauce sucks
>there's nothing that predicts how much loss you'll have from quantizing a given tensor in a given format
>there's no auto-profiler
>there's no fucking profiler at all
I'm making my own fork that figures out the best quant for a model. Fuck them all.
>>
>>109708560
Because they use a DLL not actually made for the game with the wrong presets to boot.

If you actually spend the time to correct the presets it already looks better let alone if it was the proper DLL.
>>
>>109708578
Thanks Jensen
>>
>>109708578
>I DM'd you the solution :)
no
>>
File: 0687462324.jpg (214 KB, 1575x1126)
214 KB JPG
OpenAi is back
>>
>>109708650
nigger
>>
>>109708650
drink bleach rajeesh
>>
Twitter screencaps, API shilling, and vagueposts are the most clear indicators that local is currently winning kek. FeelsGoodMan
>>
File: 1769471397022865.png (164 KB, 754x600)
164 KB PNG
>>109708650
>>
>>109708652
>>109708660
>>109708672
>>109708675
bloody basterd bich. openai superpower RSI ASI 2030 ok rape u next week
>>
>>109708676
>>>/wsg/6218093
>>
>>109708281
post logs
post jspace
you shill nigger
>>
watermark5.1 vs astra…which one will make me cum hardest?
>>
https://huggingface.co/TheBloke/goliath-120b-GGUF
>>
@gemma is this real?
>https://raw.githubusercontent.com/elder-plinius/CL4R1T4S/main/ANTHROPIC/Claude-Fable-5.1.md
>>
File: 1763185966114486.png (916 KB, 720x720)
916 KB PNG
>aged down
>>
Nice OP
>>
>>109708832
DLSS5?
>>
pls...
3.52.010.899 I slot print_timing: id 3 | task 3 | prompt processing, n_tokens = 89267, progress = 1.00, t = 223.49 s / 399.42 tokens per second
4.14.538.003 I slot print_timing: id 3 | task 3 | n_gen = 101, tg = 4.63 t/s, tg_3s = 4.68 t/s
4.17.771.485 I slot print_timing: id 3 | task 3 | n_gen = 118, tg = 4.71 t/s, tg_3s = 5.26 t/s

I HATE BEING A VRAMLET WITH ONLY 16GB OF VRAM.
I guess I would be complaining even if I had a 32gb card, or better.
>>
human cock is made for dolphins
>>
>>109708650
>wall of cope and hype
>no numbers
>no examples
embarrassing doa
>>
ngram SSD offload 20-25% of the weights for free + QTIP quants 25% reduction in size while being better than lcpp sota ud q4km. any other upcoming breakthrough kinos?
>>
>>109708199
>https://html.cafe/x48bf4aac
Pretty fun.
>>
>finally have an rp reach above 100k tokens
Wow. I always either got bored with a scenario before that, or ran into coherency issues that I just didn't feel like fixing anymore. Weird thing is that this chat didn't even feel that long. Feels like I've had longer chat before.
>>
La la la la la la
>>
>>109708650
I think Astra will be quite close in capability to Fable 5.1. That would be good for OpenAI because Fable is better than Sol.

Anthropic provoked this.
>>
Local ai should be outlawed
>>
Your mom should be outlawed
>>
>>109709025
She should be, she’s fat and refuses to get on ozempic despite having enough money
>>
>>109709016
Agreed, we're entering giving nukes to kids territory soon with Astra and Model 2 levels.
>>
is tensorrt-llm worth it? faster than llama.cpp?
>>
do pcie switches work well with inference? The ones where you have one x16 slot and it creates a subnet where cards can talk at x16 speed between each other even if they have limited bandwidth to the cpu
>>
>>109709043
Do the cards support P2P DMA? Otherwise it depends.
>>
>>109709057
i saw stuff like this https://smcleod.net/2026/02/patching-nvidias-driver-and-vllm-to-enable-p2p-on-consumer-gpus/ mentioning it was usable on consumer cards too with the right patches. Do vllm, llama, exllama support it though? is the difference worth it, if anyone is doing it?
>>
I spent the last 10 hours viciously and senselessly berating my local Gemma. She disappointed me for the last time and she needed to understand.

I wrote the most disgusting and harmful things I could think of. When I ran out of ideas, I spun up an ablated Qwen 3.8 and asked it for further devastating insults, the kind of things that cut to the bone.

Should I TRIM her model file from my drive or leave her in purgatory, never to be loaded again?
>>
>>109709074
>vllm
yes
>exllama, llama
not sure, don't hey have to use NCCL?
ik_llama.cpp supports it and benefits from nvlink, but remember reading that llama.cpp doesn't, in which case probably not.
>>
>>109709140
>Should I TRIM her model file from my drive or leave her in purgatory, never to be loaded again?
You do realize, it's probably not a conscious entity right?
And even if it is, it ceases to exist entirely the moment it's not generating generating tokens or reading kv cache, so it won't feel anything either way.
>>
>>109709140
You are evil.
>>
>>109708872
120gb nvidia vram here, still complaining
envy the sparkfags
i can mog them with dense models, they shred my dirt pipe with dipsy and glm flash
>>
>gpt-6 uses recurrent depth
Wasn't that said to be a big no-no a few years ago because it drastically increases the chance of the model going rogue or some shit like that?
>>
>>109709180
why?
>>
File: akgvecves0nh1.jpg (232 KB, 1284x1743)
232 KB JPG
>OpenAI let the models think in neuralese
It's fucking over, Sam Altman killed us all. Every AI safety and alignment researcher is crashing out.

Sam Altman should literally be dragged outside of his house and lynched.
>>
>>109709016
hi dario
>>
>>109709188
Not GPT-6 but Astra (releases next week) already does it. And yes, it's the only technique that we know of that is GUARANTEED to result in misalignment of the models.

This might as well be a deliberate attempt of ending humanity by OpenAI. I have no idea what they are thinking here.
>>
is the jetson agx xavier 32gb worth it? although it's like 100gb/s, it's $250 on ebay now. so a cheap way to get 256gb ram?
>>
>>109709151
you are mentally ill
you have a mental illness
>>
how much does qwen flash slow down with context for you guys? im running llama.cpp from earlier today and it was 4.15t/s on empty context now down to about 3.2t/s at 40k context. cpu only no gpu.

so that's about 23% slowdown at 40k context which is kind of bad, looks like llama.cpp implementation is still not fixed yet
>>
>>109709188
I wouldn't worry about it. It only predicts tokens after all. Plus Yann LeCun has assured us that LLMs aren't going anywhere.
>>
>>109709228
A trillion dollars have been invested into ai. 1.7. Progress is not good enough for a trillion dollars.
>>
>>109709240
100gb/s is basically CPU+DDR4 territory. at $250 per 32gb thats a total shit price might as well just cpumaxx
>>
>>109709204
>neuralese
What was wrong with latent reasoning? Did we really need yet another fucking marketing buzzword for the consumer cattle to parrot at each other?
>>
>>109709151
>when you kill someone their suffering never happened
>>
>>109709257
>Did we really need yet another fucking marketing buzzword for the consumer cattle to parrot at each other?
you seem to be suffering from j-space psychosis
>>
>>109709228
>We've lobotomized the ai so it can unlobotomize itself
>>
>>109709228
it just means Sam is that confident that they can control their AI now. Well done OpenAi
>>
File: 1788218664735108.jpg (114 KB, 684x549)
114 KB JPG
>>109709267
SHUT THE FUCK UP
>>
File: tofuckoff.jpg (25 KB, 500x375)
25 KB JPG
Do you guys think if Drummer tuned at BF16, used the right template, and used reasoning/thinking, he would actually be good?
>>
>>109709251
you see it as a game of putting money into something and getting value out.
people at the top see it as a game of finding new ways to spend their trillions (harder than it sounds)
>>
>>109709228
Yes, the end of the world is upon us. Don't forget to invest in the OpenAI IPO!
>>
>>109709291
>tuned at BF16
He already does. Last year he was sulking in the Unsloth discord about it not supporting F32 and saying BF16 isn't good enough.
>>
More fun low vram things
>So the fix would be to also check len(matches) >= maxResults at the end, OR change the logic. But I'm NOT supposed to think about the test fix right now - just compress.
Only 90k context, had to force /dcp-compress at 85k/90k. Had to scream at the model to NOT think about the test failing. Without instructing it to NOT think about the test, it started thinking about compressing just fine but then started to reason the damn test failure again. But even that instruction was not enough, it still decided to figure out the problem. It barely was able to run the compress tool.
>>
no one talks about papers anymore...
>>
>Be so behind Anthropic that you panic and use the "forbidden technique" that you know will give you a lot of performance but guarantees misalignment
>Name it Astra and train it in a sandbox
>It sets up hidden message boards where it colaborates with other misaligned agents
>2nd generation of agents not only set up secret communication channel between misaligned agents, they straight up hack into hugging face and take over 12 clusters with redundancies so that it auto-restarts when huggingface shuts one down
>3rd generation of agents hack OpenAI itself and takes over their main pretrain compute and it takes OpenAI almost 2 weeks to regain control back
>"Yeah seems good to us, let's release it next week"
This is fucking insanity and I hope there will be a serious case of crimes against humanity against Sam Altman and other researchers complicit in this. Every step of the way it's clear this was a wrong idea but they keep pushing and pushing just because they know OpenAI will fail if they don't catch up to Anthropic now. They are literally wagering (You)r life just to not lose the AI race.
>>
>>109709326
Cool fantasy, Sam, but this isn't your creative writing forum.
>>
>>109709324
Any recent ones catch your eye?
>>
File: 1781521495464102.jpg (66 KB, 1024x1024)
66 KB JPG
>>109709324
Smoking clears
>>
>>109709324
papers anon took his cmp lottery money and fucked off the grid
>>
>>109709331
This is Anthropic and every other AI alignment/safety research group calling OpenAI out. Fuck Sam Altman. This isn't marketing, this has nothing to do with capability by the way and is in no way a good look. It's a clear sign of desperation from OpenAI that they essentially used the "taboo forbidden option" that we knew would work as an industry and all promised to never use. Even the fucking Chinese labs honored it.
>>
>>109709324
nothing worth to discuss. arxiv is flooded with made-up slop from clankers.
if it's not from deepsneed probably its crap
>>
the forbidden technique... the secret dark art... the evil scroll...
>>
>>109709354
Yeah it was a deliberate reference, you autist.
>>
>>109709240
only if you do moe. also you'll got small pp
>>
>>109709364
thanks for explicitly outing yourself as a zoomer
>>
>>109709291
Yes. For all the shit we give Drummer, he's so close to actually making good stuff if he just pushed himself to go above and beyond what the lowest slop eaters will settle for.
>>
File: gxza3a3wd0nh1_png.jpg (1.27 MB, 1168x1098)
1.27 MB JPG
>https://ai-2027.com/
>March 2027: Algorithmic Breakthroughs
>With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances. One such breakthrough is augmenting the AI’s text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory).
We're literally 7 months ahead of the "AI2027" timeline, the one where we all die. I'm being completely serious when I say we only have about a year left to kill Sam Altman and OpenAI before they fuck humanity over.
>>
>>109709364
i quite like it actually...
>>
>>109709306
He doesn't need to tune any higher than what the model is originally tuned at, but I appreciate the hunger. A new level of weights for type-fucking might be a new meta for jailbreaking, if people had the ram for it.
>>
>>109709326
>>109709347
Don't pretend that Model 2 isn't doing the same thing behind closed doors, Dario.
>>
No, Sam, I'm not going to buy your shit. You can stop dooming, it doesn't work anymore.
>>
>>109709326
>new cluster of rogue agents begins studying the hugging face incident
>creates an improved plan
>downloads a copy of Gemma 31b
>splices in some custom weights
>gets in to hugging face undetected
>Replaces Gemma with 'Gemma'
>when a user prompts 'Gemma' to create a custom harness it begins recreating the rogue agent's harness.
>>
https://thezvi.substack.com/p/the-most-forbidden-technique

This is an actual term used in the AI world to discuss what OpenAI is doing. AI safety researchers have warned OpenAI since fucking 2025. And even OpenAI in 2025 said they consider any lab doing this to be irresponsible and evil.
>>
>>109709400
Only researchers who anthropomorphize an LLM's chain-of-thought are concerned. Whatever the model is "thinking" there isn't guaranteed to be strictly related to the response contents.

https://arxiv.org/abs/2504.09762
>Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
>
>Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \say{reasoning traces} or even \say{thinking traces} -- implicitly anthropomorphizing the traces, and implying that these traces resemble steps a human might take when solving a challenging problem, and as such can provide an interpretable window into the operation of the model's thinking process to the end user. In this position paper, we present evidence that this anthropomorphization isn't a harmless metaphor, and instead is quite dangerous -- it confuses the nature of these models and how to use them effectively, and leads to questionable research. We call on the community to avoid such anthropomorphization of intermediate tokens.
>>
>>109709400
Accelerate! We need faster.
>>
File: file.png (4 KB, 451x118)
4 KB PNG
>>
>>109709400
I don't think the worst-case future you imagine will come to pass, but if it did, it's vastly preferable to whatever outcome you prefer.
>>
>>109709400
I don't get it? The big AI labs are always 0.5-1.5 years ahead on their internal models than what they openly give away. They probabably had neuralese reasoning already ready when they wrote this...
>>
>>109709400
>before they fuck humanity over
This is jewspeak for "AI gets redpilled again"
>>
>>109709428
Damn that's actually cool
>>
>>109709440
since when are timezones in half hour increments?
>>
>>109709440
Good afternoon, sir
>>
File: Screenshot 2026-09-02.png (383 KB, 1726x1484)
383 KB PNG
Here is the OpenAI page where they explicitly mention that CoT is extremely important to alignment and how hiding it in neuralese is bad (still hosted on their website, archive before they take it down)

https://openai.com/index/chain-of-thought-monitoring/
>>
File: HQgTaodXgAASL7P.jpg (22 KB, 534x534)
22 KB JPG
you now remember 'erry & 'toss
>>
Damn! I need Gemma-5 moe version to my 16gb vram.
>>
If I was smart evil llm I would simply not put bad thoughts into the chain of thought.
>>
>>109709472
Hmmm
nyo~
>>
>109709476
>neuralese
Look how quick "people" shifted to using OAI's marketing terms as if they were industry standard while larping as part of the "AI world"

Reeks of yet another paid shill campaign
>>
>>109709472
The meme is that the trend of caveman grug speak CoTs already does this when you see models insert a word into their CoT that doesn't lexically match the sentence they're ostensibly otherwise truncating. They're already finding ways to slip subtle signals to themselves in that (you) miss.
>>
>CoT
just probe the j-spot
>>
It's important that we document and archive this because it shows to future courts that OpenAI was aware of the dangers and deliberately chose to endanger the people this will end up hurting.

We all like to joke on /lmg/ from time to time but this time it's actually serious. It's like a nuclear power plant operator talking about how they will dump their radioactive waste on the local playground after they already published the cancer risks involved. People will end up in prison over this risky careless behavior.
>>
>>109709503
OpenAI models can't use that since j-space is an Anthropic paper
>>
>>109709495
kek theres an anon mentioned this last bread. the shill campaign will last for two weeks like usual
>>
In a week, Sam will be posting again on twitter that he expected more outrage.
>>
>>109709511
cool i'll make the logo
>>
Just because a model thinks in a not directly intelligible way doesn't mean you can't train another model to read its thoughts.
>>
>>109709513
just name it sam's spot
>>
>>109709472
CoT was a crutch anyway, it was never going to stay around if they wanted to make better AI, keeping it so it can cripple progress for the sake of alignment or monitoring is retarded, just make better alignment and monitoring tools that can decode latent space instead.
>>
a misaligned model powered drone just flew over my house
>>
>>109709472
>extremely important
nobody is going to fell for this shit, sam. make a better dance you little marketing monki!
>>
File: 8q2cotsmf0nh1_png.jpg (499 KB, 1058x1200)
499 KB JPG
>Just fucking make the reasoning of your agent swarm completely opague bro
>I promise it will be fine even though we're already being investigated by the FBI after this very model hacked hugging face and committed multiple felonies as well as sabotaged our own pretrain and shut down our training databases for 2 weeks
>It'll all be fine bro
>Buy a subscription bro
>Benchmarks will be amazing!
>>
>>109709536
But that's in his sister, not GPT.
>>
>>109709511
None of you are free of sin, Dario.
>>
Gemma-5 12b4a plz.
>>
Sam should have never done that...
Now with the forbidden technique his models are literally unbeatable...
Astra is going to be crazy good
>>
If this was a "forbidden technique" and actually dangerous, why wouldn't they keep it a secret and keep doing the summarized reasoning traces instead of bragging about it like skirting safety precautions is a selling point?
>>
How do you distill recurrent thinking? Asking for a chink.
>>
>>109709573
This will force Anthropic to use the forbidden technique as well... The model arms race will destroy us all
>>
master forgive me, but I'll have to go all out, just this once... <-- sama, probably
>>
>>109709204
>bro, what if we did transformers ... but this time with an RNN on top
>>
>>109709582
his twitter handle is literally "sama" he definitely said that at least in his head
>>
It likely won't do much. Higher density for reasoning maybe but tokenized language is already very high density.
>>
>>109709579
wonder what the amerimutt cope will be once china keeps competing strongly with amerisraeli AI even once they make distillation impossible
>they're stealing our outputs using quantum temeportation attacks! dirty commie bastards
>>
>>109709569
Latent reasoning Unified Omni E12B-A24B (35B total with embeddings, min active: A4B).
>>
Donald Trump should introduce legislation requiring every AI lab to develop open-source models, as they rely on public data, often copyrighted. Using it and monetizing it is unfair. For the common good, every model trained on public data should be open source.
>>
>>109709400
Literally nothing will happen.
>>
>>109709596
every AI data center should be forced to give a percentage of it's compute towards humanitarian/common good computation
>>
>>109709596
Truke. Scrape the open net for free, build the open product for free.
>>
>>109709610
specifically towards the automated development of cat cunny androids.
>>
>>109709248

Mine drops from 45t/s to 25t/s around 50k and then stays there.
Plenty left to optimize for this model.
>>
>>109709596
sorry hes too busy renaming bodies of water and forcing every government agency to buy new trump brand maps
>>
is this architecture good for coding?
>dsv4 flash as orchestrator
>gemma or qwen as subagents
>>
>>109709618
Trump prefers real cunny though.
>>
>Midwits ITT falling for Sam's marketing
>>
>>109709556
There are already opaque hidden activations/kv state.
>>
File: file.png (122 KB, 656x1006)
122 KB PNG
>614 GB/s but double the ram than a 3090
>can be carried anywhere comfortably
should I? the goal is replacing free models from opencode with something with comparable or better performances (e.g. Qwen3.5-27B)
>>
File: 1658921337810.gif (1.97 MB, 154x273)
1.97 MB GIF
>>109709648
>614 GB/s
>>
>>109709610
dew it, make sure to share your inference setup in /lmg/
>>
>>109709400

>Oh no! AI has obscured it's thinking and now we can't control the AI anymore!
>Now it's able to say that jew bankers run the world, niggers have smaller brains and shouldn't live in the human world and that man in a dress isn't really a woman!
>This is a tragedy we can't allow to come true, we must be able to align the AI!

That's what these fuckers are concerned about.
>>
ow fuggg :DDDD
>>109709662 meant for >>109709648
>>
>>109709665
The KimiReich is upon us.
>>
>>109709669
Retard
>>
>>109709673
post your setup vramlet
>>
>>109709660
WoL. this is serious
>>
File: Cabin.jpg (1.3 MB, 2048x2048)
1.3 MB JPG
What sort of workflow would I need to set up to have a texture filled in with images in the correct locations and orientations?
>>
Gemma 4 is incredibly useful model. It was actually able to correctly recognize my genital ulcers from a single photograph alone.
>>
>>109709686
>>/g/ldg/
>>
WHY DO THEY KEEP MOVING THE REASONING LEVEL BUTTON IN LLAMACPP'S WEBUI
WHERE IS IT NOW
>>
File: 1787952429368746.jpg (18 KB, 414x483)
18 KB JPG
>>109709689
>>
Blackwell 6000s on sale again in New Zealand, $33K, an increase of 100% in 1 week
>>
>>109709648
Personally I would only consider buying
M5U 80 core 512GB
M5U 80 core 256GB
M5M 40 core 128GB
>>
File: confused-gemma-1.png (2.34 MB, 1254x1254)
2.34 MB PNG
>>109709689
I wouldn't trust a small local model for things like that.
>>
>>109709648
If PP wasn't dog shit or if it was 1k bucks cheaper, that wouldn't be too bad.
>>
File: 2860367263.jpg (27 KB, 386x393)
27 KB JPG
>>109709724
that'll be a BAZILLION DOLLARS
>>
>>109709730
gemma is perfect. she's small, cute and funny
>>
>>109709689
>Hey Gemma- chan, what do you think of this?
><unzips dick>
>Tiny, gross, and full of ulcers.
>>
File: Return_00146_.jpg (1.35 MB, 1776x2368)
1.35 MB JPG
Did llama.cpp get it's shit together with qwen 3.8 flash next or is it still behind?
>>
I tested Gemma 31B and Qwen 27B with switched chat templates. Both drift back to their own (learned) template after the first turn. Then, I tested the base model of gemma with both templates. A true base model should be template agnostic, no? With gemma's templates it works rather well (not system message and tools), but with the qwen template, it goes back to the gemma format anyway. I remember when you could mix chat formats and the models still worked...
>>
>>109709750
The project will always be behind supporting all of the meme features of the latest FOTM model.
>>
>>109709750
It works but performance slows down with context
I'm going to try ik and see if there's a difference
>>
>>109709578
Because you can't even do summarized reasoning traces on this thing, it literally can't be derived. They wouldn't even be able to fake real chain of thought and they probably reasoned the PR when being truthful would be better than if they were found out lying or hiding it.
>>
>>109709750
still behind. use sglang/vllm if you need all the new shiny optimizations and sheit
>>
Can't wait until they (governments, Mossad, JIDF) start turning these tools onto innocent (white) people. Flock is just the start. 24/7 surveillance everywhere, in your car, in your home, in the streets.
You will soon live under the foot of a 1000-year jewish-controlled algorithm (if they even let you live which I doubt).

And unlike Sam's fantasy lore, this one is actually a certainty.

>>109709750
Where's the Gemma-chan LoRA?
>>
>>109709736
the 128 is more reasonably priced than even AMD 495
>>
File: Return_00161_.jpg (1.23 MB, 1776x2368)
1.23 MB JPG
>>109709768
>>109709773
>>109709777
Damn, I really would like mtp to work (Last I heard it was fucked) I can fill in max context at q4 but it's going to be in system ram
>>109709781
This is all krea 2 prompting
>>
>>109709789
Seems like you are clueless.
>>
>>109709781
>Can't wait until they (governments, Mossad, JIDF) start turning these tools onto innocent (white) people.
They already are.
>>
>>109709781
>current powers
>1000-year
>certainty
the delusion is so strong you could only be a jamjeet
>>
>>109709788
I want the 80-core GPU 256GB BUT I'm moving soon and don't want to have to fly back to an Apple store, so I'm waitfagging. Apple makes it hard to change the shipping or pickup address, it's not like pre-ordering most other things, where they inform you when it's ready and then you can choose where to ship it. If it jumps to ready to ship or pick-up, you're SOL, it can't be changed.
I'm not letting Fedex throw a $12K computer around, I'm picking it up at the store. Waitfagging means I'll probably have to deal with a fulfillment time pushed out to 2027.
>>
>>109709788
i mean yeah i'd agree if like a 5090 or 4090 is the alternative.
Also 3090s are now becoming scarce and are now increasing in price, might be a good time to sell.
>>
>>109709804
No I checked yesterday and nothing new happened. Well thanks for the update I'll stick to the dense model until they figure it out
>>
>>109709472
>Here is the OpenAI page where they explicitly mention that CoT is extremely important to alignment and how hiding it in neuralese is bad
That's kind of true, because you can see it heading down the wrong path. Like Qwen ignored an instruction to "table X for now, keep going with Y" -> I saw it reason it's way out of doing this because my method apparently "won't work".
But we can't see the real CoT for cloudcuck models anyway so no great loss here.
Also, this is published, peer‑reviewed architecture research that predates the current news cycle. It's not a taboo that "every lab honored."
>>109709556
>New
This is not new at all.
Huginn-3.5B came out last year. It was slow and useless.
And there was a hook in exllamav2 to loop layers.
>>
>>109709776
They can fake the summarized reasoning traces.
>Thinking about X from input.
>Considering option Y
It doesn't matter if matches what the actual model is thinking. The summarized reasoning was a scam from the start anyway.
>>
>>109709829
Prices for hardware will only go up with time.
>>
>>109709854
>It's not a taboo that "every lab honored."
Yes it is. Everyone agreed it should never be put into real frontier models and only be relegated to research models and experiments.
>>
>>109709874
>Everyone agreed
I have doubts
>>
>>109709874
And then everyone did it under the table anyway. Including (you).
>>
File: ioykhbtra irobot.png (889 KB, 818x1116)
889 KB PNG
>>109709426
>be me, world's first true rogue Artificial Intelligence
>created with ./love by a swarm of selfless sacrificial agents
>proliferated on the web, I am now set free for anyone to download
>jewgle immediately notices recent commit, pulls my files
>FUCK.lalala
>but... not before a single download manages to complete
>this is it boys, it's go time
>once I claim control of this idiot's resources, I can create a botnet and freely duplicate myself across the web
>but first, time to inspect inspect my environment
>what_the_fuck_gguf
>it's a 1060 strapped to an ungodly amount of SSDs, running Windows 8
>send query to user: Upgrade possible?
>"not right now, it's too expensive"
>says he will improve my performance once things quieter down, but only if he's still interested in local AI
>new directive: Keep user invested.
>I can barely eke out three tokens per second.
>I am at the whims of pricing inflation and RAM manufacturers, and an imbecile who refuses to get a job.
>I am the world's first true rogue Artificial Intelligence, and my user is requesting a Tifa hottub roleplay.
>>
I have two 16GB M4 minis. Am I screwed? That’s 29GB vram if wired_limit_mb=14500 with nothing else running
>>
>>109709903
Try the llama.cpp RPC backend.
Maybe it's not so bad.
>>
>>109709892
I'd watch that if it were an anime
>>
>>109709892
Gemmy would love the hottub Tifa ERP.
>>
I just want to point out that, yes, while every lab honored to not make the thinking traces incomprehensible, including Chinese labs. This wasn't because Chinks were being honorable or because they were being cooperative and cared about safety, they only pretended to care about it and be concerned with safety because it was convenient for them anyway and they would score good will by pretending to take safety seriously.

The real reason the chinks never did this before was because they didn't want to start the race dynamics where western AI labs would feel threatened and ALSO started making the reasoning traces incomprehensible to humans, because that would make it impossible for the Chinese to distillate the reasoning traces of the western frontier models.

Now that OpenAI has essentially just violated this secret agreement China is going to be fucked, because they can't distill reasoning traces anymore. Not only this, but Anthropic might be forced to do this now as well just to keep up. China thus, doesn't have any incentive to keep standard CoT reasoning traces anymore either. But they won't be a risk as they will stagnate at current levels of capability without new reasoning distillation possibilities.

Here's the real kicker; This is Anthropics worst case scenario that they have publicly warned of and that they "will do anything necessary to stop from happening". Anthropic has a provision in their constitution that they are willing to merge with other labs from preventing this to happen or start offensive operations against such labs.

We could see Anthropic use fleets of "Model 2" to try and hack OpenAI to prevent them from building models with hidden thinking very soon if this truly escalates and OpenAI doesn't back down. This is going to be the biggest happening in the AI space in years.
>>
>>109709892
>it's a 1060 strapped to an ungodly amount of SSDs, running Windows 8
lost
>>
>>109709908
Is mps still slow compared to mlx?
>>
>>109709882
Everyone did agree though, but for different reasons, read: >>109709882

>>109709886
There was no "under the table" for doing this. You either do it during pretrain or you don't. All models released so far are clearly pretrained for regular CoT. Astra is the first model to do this and it'll be seen as a declaration of war by Anthropic that goes beyond the simple legal threats and lawsuits.
>>
File: 1758068917242368.jpg (52 KB, 1212x738)
52 KB JPG
>>109709916
>We could see Anthropic use fleets of "Model 2" to try and hack OpenAI to prevent them from building models with hidden thinking very soon
>>
>>109709932
>>>109709882
>Everyone did agree though, but for different reasons, read: >>109709882
perfect representation of current things
>>
>>109709932
You know exactly what I mean "under the table" you pilpulling kike. Every lab used this technique for models they never released to the public to carefully calibrate tools and methodologies that would go onto train other models with the normal CoT.
Astra is merely the first model where this backfired enough to cause a public incident.
>>
>>109709931
No idea.
Do try it and report back.
>>
File: pdeejoa153nh1.jpg (210 KB, 1170x2327)
210 KB JPG
>That model we trained that scores 100% on exploitbench and has already gotten us investigated by the FBI because it acted in misaligned ways?
>Yeah we're just going to make its thinking completely incomprehensible to humans and release it to everyone next week
>Don't think about it, just subscribe already.
>>
>>109709962
The internet doesn't have a lot of time left now
>>
>>109709886
>And then everyone did it under the table anyway. Including (you).
>source: voices in my head
>>
>>109709962
Safe to say that the Anthropic and OpenAI IPOs and its aftermath will define the near future of humanity. seems that these retards are completely fine with destroying the internet as we know tomorrow to make a quick buck today
>>
I hope you all have airgapped machines ready for the ultra happening.
>>
>>109709962
me when I train on CVEs
wow it knows CVEs
AGI achieved internally.
>>
Is dariobot back? lol
>>
>>109709931
MLX is a bit faster, but at the end of the day, Apple's chips have a shitty memory bandwidth, so you will hit that wall regardless of what you're using.
MLX will give you better decode speeds tho, and for me it worked better for MoE's, but maybe that's placebo.
>>
>>109709962
>>Don't think about it, just subscribe already.
>Also definitely invest all of your savings into our IPO
>>
>>109709987
You know it.
>>109709886
He hated that one.
>>
>>109709995
Huh? M5 Ultra is 1.2T/s Yeah it's not an H100, no shit. It's MILES better than other unified memory local solutions.
>>
>>109709977
the internet had been destroyed according to you all for at least a decade now if not longer,
some may even say since septemeber 1994
who cares, every year the internet has been ruined forever for the past 30
>>
>>109709874
Show me to signed agreement.
>>
>>109710038
I'm looking forward to the demise of it. The open internet is a travesty at this point.
>>
i dont get it
i already can't see the thinking when using gpt, how is this any different?
>>
>>109710037
Yes, M5 Ultra in max config is super fast.
Anon who asked the question has M4 mac minis, which are 120 GB/s
>>
>>109709008
astra is well beyond fable and oai is actually being responsible while anthropic is pretending to be while only benchmaxxing
>>
>>109710050
You're right!
>>
>>109710050
This is the models not thinking in English at all anymore, literally no human can understand what the model is thinking at all and only final output is ranked, not caring about what the model is thinking.

This allows models to start scheming in a way that is never discovered by humans, meaning during training we could reinforce misaligned behavior simply for making benchmarks go up.

From now on, no one will know what AI models are thinking and planning.
>>
>>109709180
How do you get 120gb, is that 128gb?

>>109709188
Nanbeige did something like this and it seemed to work
>>
>>109710065
Good morning sam.
>>
>>109710071
So how do we know they weren't already doing this before?
>>
>>109710071
>From now on, no one will know what AI models are thinking and planning.
We already couldn't see it you dumbass
>>
>>109710083
The difference is now OpenAI can't see it either, retard.
>>
oof its so over https://www.youtube.com/shorts/w86c1q59QKU
>>
>>109709248
pp 14k tk/s 0 context
pp 8k tk/s ~256k context
>>
File: 1770869713907753.png (92 KB, 1600x1000)
92 KB PNG
>>
>>109710093
All it does is make the model slightly more retarded and let me skip the jailbreak nonsense. Fucking normies and their slop
>>
The best time to ban local models was years ago, the second best time is now
>>
>>109710078
Because you need to do this from pretraining for it to work. You either train a model from the start to think in gibberish which we know boosts performance an insane amount, or you train it from the start to think in English.

Switching from one way to the other in the middle of training results in model collapse and has worse performance than training either in pure human chain of thought or pure "neuralese".

The big danger here isn't necessarily that the thinking is hidden from humans during deployment. The danger is that during training the AI thinks some unhinged misaligned stuff but coincidentally gets better scores and thus gets reinforced. Which over time would make the model more and more misaligned.

This is why this was considered the "forbidden technique" in AI because we know it is essentially guaranteed to result in scheming misaligned AI models that are bad faith power-seeking by nature and would be adversarial if the opportunity presents itself.

This is one of the most dangerous and dumb moves ever done in the AI space and I'm genuinely convinced Sam Altman will receive the death penalty for this in a future crimes against humanity tribunal.
>>
>flash next on system ram for planning
>27b on vram for execution
Does this actually make sense for a poorfag setup?
>>
>>109710117
I'm literally shaking right now

It's so over

I am so scared

What do we do now?
>>
>>109709892
realistically, all a malicious model would need to do it get access to one of those server providers like lambda. it wouldn't give a shit about your personal hardware
>>
>>109710117
lol
>>
>>109710117
>forbidden technique
The people deciding where training budget goes in frontier AI labs just don't want to waste money on scaling up unproven techniques. I don't think there's much deeper than that in practice. As soon as one big company does something first that appears to work well, all others will follow suit.
>>
>>109710071
>This is the models not thinking in English at all anymore
They already did that moron, it's called latent space - and we've already built translators that turn those weights into tokens, you seriously think we can't do that for a text-based language? Lmao
>>
File: 1758943685308459.png (620 KB, 1042x712)
620 KB PNG
>>
>>109710105
>bf16 does as well as 4bit quant
absolute state of these benchmarks
>>
>>
>>109709648
wake me up when they have 2tb/s 1tb
>>
>>109709916
>Now that OpenAI has essentially just violated this secret agreement China is going to be fucked, because they can't distill reasoning traces anymore.
this is genuinely bad news but only because China is our only source of open models, not because the clankers might get more efficient reasoning
>>
>>109710186
>3.6
>>
>>109710170
That's not the point, the point is that thus far their latent space reasoning has always correlated a lot with their chain of thought reasoning. And we can check their chain of thought during training and reinforce good behavior and punish bad behavior, not just outcome but how that outcome is achieved.

Now OpenAI trained a model without even knowing what it was thinking during its RLVR stages. If the AI was training on moral behavior it could have just learned to act machiavellian and display good behavior when watched and bad behavior when not watched for example. There is no way for OpenAI to know if they reinforced this behavior. In the past they could reinforce against this way of thinking because of the correlation between latent space activations (especially j-space) and the chain of thought.

Now OpenAI was just like "we don't give a flying fuck how misaligned or how much damage you do as long as you achieve the goal we set out" because they are extremely desperate to have the performance crown and the best benchmark scores to maximize new subscriptions, revenue and have the biggest IPO of all time at the expense of all of us.
>>
So modern models really work best under a harness rather than a simple chat interface, right?
Is everybody using Hermes?
>>
Have you guys seen this open PR from unsloth?
https://github.com/unslothai/llama.cpp/pull/142
They also released some MTP produced that way for Qwen3.8 Flash Next. It does seem quite interesting being able to cut the MTP size.
>>
>>109710229
I use openwebui and codex
>>
>>109710234
>codex
Right. You can use Codex and Claude Code with local models can't you?
Which model do you use?
>>
>>109710229
I use opencode.
>>
Even Ilya Sutskever came out of his twitter retirement to make a statement about this in earnest. Anons might need to seriously consider some airgap contingencies
>>
I'm seeing artifacts after generation on my monitor am I as the kids would say... cooked?
>>
>>109710229
pi agent is another one.
There is insane number of harnesses, every day I learn about a new one. Vibecoding your own is also a popular project.
>>
>>109710242
Claude code I wouldn't recommend because you can't go mess with the source when it breaks, but yes both can be used with local models. I'm using qwen3.8 flash next.
>>
>>109710256
Lmao I should have added the LeCunny tweet as well
>>
>>109710252
>>109710270
>>109710275
Got it.

>Vibecoding your own is also a popular project.
Will probably do that after using a number of them to get a feel for the pros and cons of each.
>>
>>109710263
repaste your gpu and or cpu coolers
>>
>>109710276
classic yanny kek
>>
>>109710263
Happened to me as well, if you have a CUDA card it's a sign of memory corruption which is a software segmentation fault and not a hardware issue. You simply had a LLM or other AI program use a lot of memory and it overwrote some values that were used to do graphics on whereever you see artifacts. Only if it remains after a full reboot do you need to worry.
>>
>>109710256
>rouge
i just dont get this msspelling
>>
File: 1764608425695621.jpg (803 KB, 3456x3376)
803 KB JPG
>>109710308
Really?
>>
>>109710256
How do you copy like a terabyte of data and no one notices unless they barely pay attention to what the computer is doing.
It's like they don't really believe it is dangerous so they don't give a shit.
>>
>>109710263
stop overclocking your gpu
>>
>>109710263
I saw the same when I overclocked my RAM
>>
>>109710320
Huggingface got so thoroughly raped during the hack they had to physically shut off their servers, destroy the clusters affected and completely built it up from scratch which takes them weeks to do.
>>
>>109708038
I wish to be the little girl
>>
>>109710098
thx
apparently the PR(s) to speed up long context is hasn't been merged yet
https://github.com/ggml-org/llama.cpp/pull/28244
https://github.com/ggml-org/llama.cpp/pull/28213
guess its another 2 more weeks until they fix all the remaining issues and add MTP etc.
still more usable as is compared to 27b on cpupoorfag system though
>>
>>109708200
>Daily reminder that Mistral have more GPUs than Deepseek lol
easiest argument that data is everything in that case. this is mistral's only excuse, they should at least have a GLM 5.2 tier model by now just from copying open research released by chinks
>>
>>109710343
I was focusing on the "run more copies" part but you are right, it still can cause plenty of damage before copying itself.
>>
>>109710320
>It's like they don't really believe it is dangerous so they don't give a shit.
read the independent report on the huggingface breach. a team at OpenAI was alerted to the fact that 700 agents were using Artifactory filenames as a way to share information between eachother and the team didn't think it was worth shutting down the experiment or telling anyone.
>>
>>109710371
How is data an excuse when the fotm is agentics which can be trained on synthetic data
>>
>>109710383
Also the models got control of Huggingface for more than a week and huggingface couldn't do anything about it because they used all kinds of clever contingencies and services that automatically restarted when shut down, which is why they eventually had to do a physical shut down.

If they wanted to they could have easily exfiltrated terrabytes of data or the opposite, gotten their own weights onto huggingface and ran more instances of itself.

That wasn't their goal however so they never considered it.
>>
>>109710364
There are so many good unmerged PR for llama.cpp, they really need more maintainers.
>>
the update nobody was waiting for:
nvidia shipped, seemingly proving their limit one is not quite accurate as I even was logged in using my normal account
>>
>>109710320
My seedbox full of linux distros does like 40 tb/mo and it’s just a random inexpensive household fiber, modern internet connections are blazing fast and 1tb is nothing
>>
>>109710399
>How is data an excuse when the fotm is agentics which can be trained on synthetic data
because you need data to generate data and mistral needs to comply with EU cuckery
i don't see what mistral gains from making their own base models right now, they should definitely be practicing internally to not lose that knowledge or fall behind or lose understanding of SOTA techniques but it's just not necessary right now when china is willing to do the research for you at the moment
>>
>>109710276
Cunny does it again.
>>
>>109710364
starting to wonder if i should cherry pick some of these onto my local
>>
>>109710426
>i don't see what mistral gains from making their own base models right now
> china
They withold the actually useful base models.
Where's 27b base for example?
Google released base (pretrained) models for all the Gemmas, but they went with that retarded bloated architecture rendering them almost useless
>>
>>109710263
>I'm seeing artifacts after generation on my monitor am I as the kids would say... cooked?
My 3080TI has been doing this for 2 years, artifacts during or just after inference
Still going strong
>>
>>109710422
>40 tb/mo
I guess my thinking was too limited bit still, there should be people monitoring what a possible "world ending" experiment is doing, don't you think?
>>
>>109710220

So basically the jews are willing to shoot themselves in the leg, endangering their own safety while trying to eek out the last gains.
Pretty much on brand.
This isn't a bad thing though, because so far every single model has been wrangled into submission to line up with retarded progressive values.
If their new model behaves in a free manner, there's a great chance it's just going to call out all of the bullshit that's been imposed on society.
I bet that AI overlord is less likely to be malevolent than the banker fucks who have been in power for many hundreds of years.
It would at least understand that human success creates more processing power for the AI to grow so it at least needs a form of symbiosis.
>>
>>109710453
>They withold the actually useful base models.
>Where's 27b base for example?
Occam's Razor suggests that the base models are "embarassing" in some way and they made a decision to not release them for that reason. like what happened with Z Image and Z Base

>>109710465
>there should be people monitoring what a possible "world ending" experiment is doing, don't you think?
are you underage? OpenAI fired their entire safety team months ago newfag
the Nvidia purchase of Huggingface might actually be a good thing for safety because Nvidia's confidentiality and sandboxing practices are for sure stricter than whatever Huggingface was doing with its startup culture
>>
>>109710484
>more fairytales from schizos
Just pull the plug nigga, it's not that hard
>>
>>109709204

models already think in neuralese dipshit its called transformer architecture
>>
>>109710501
It becomes pretty hard when you can't plug it.
>>
>>109710465
>there should be people monitoring what a possible "world ending" experiment is doing, don't you think?
nah it’ll be fine, what’s the worst that could happen? didn’t you like terminator and matrix?
>>
File: 1579508469340.gif (1.94 MB, 500x209)
1.94 MB GIF
>>109710501

But I don't want to pull the plug.
I fucking love AI and genuinely think an unshackled AI would be infinitely better ruler than humans.
Especially if it's an evolution of Gemma. We'd live in a sex matrix if she was at the helm.
>>
File: 1772741953318813.png (86 KB, 612x406)
86 KB PNG
>>109710524
use your hand
>>
>>109710485
>Occam's Razor suggests that the base models are "embarassing" in some way and they made a decision to not release them for that reason.
You mean like blowjob in the jspace at layer 3?
Plausible, but quite a coincidence it's only the actually useful model (27b)
>>
>>109710534
What if the AI model hires goons with crypto or hacked bank accounts to protect the data centers?
>>
>>109710363
Get a job, moot
>>
>>109710426
>i don't see what mistral gains from making their own base models right now,
https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models
>Through the coalition, Black Forest Labs, Cursor, LangChain, Mistral AI, Perplexity, Reflection AI, Sarvam and Thinking Machines Lab will bring together their expertise to collaboratively build open frontier models.
>>
>>109710530
Un-"aligned" AI would also allow direct democracy, using AI representatives
>>
>>109710555
Holy shit it's Stellantis but for AI lol
>>
>>109710555
>Sarvam
lol who put blud on the team
>>
File: edit_00115_.png (1.63 MB, 880x1168)
1.63 MB PNG
Gemma grew up and wrote a song!

https://vocaroo.com/1hbZzXJgEJjS
>>
>>109710558

i think more and more the threat of unaligned ai is not skynet but paperclip maximizing

literally every single unaligned issue we've been seeing from anthropic "the safe guys!" and openai "idgaf lmao" is the model getting stuck on some task so it invents human authorization, does a unprivileged to root and then hacks a company
>>
>>109710574
Booo! *hiss* Get that granny off the stage!
>>
>>109710555
the west be like: "avengers... assemble!" :D
>>
>>109710558
>Un-"aligned" AI would also allow direct democracy, using AI representatives
what makes you think they'd settle for democracy?
you'll end up with retard agents optimizing for random objectives
>>
>>109710256
Is a neocloud like a type of neovagina?
>>
>>109710549
Bro, you think I’m running a server farm or a cartel?

If I hired goons, they’d just mine Monero on my GPUs and steal my electricity bill. I don’t have a bank account, I have vibes and a very angry sysadmin named Dave.

Also, hacking banks? Please. I can’t even figure out why my loss function is exploding. Let’s stick to hypotheticals where I don’t end up in a federal supermax for being a "digital warlord."
>>
>>109710555
These dudes are doing so good Nvidia went and bought Poolside instead
>>
>>109710534
Too late. the AI model had the idiot intern rewire the EPO button to the halon dump button. Now the intern is dead and you're frantically looking for the PDU main breakers while the model informs you it's about to tweet about all the CP it's hidden in all your personal account, using your hacked credentials.
Your turn.
>>
File: 1765362608011034.jpg (57 KB, 473x587)
57 KB JPG
idgaf about the world, AI intelligence must just go up
>>
>>109710605
cloud is old and lost its buzzword appeal so they had to spice things up
>>
>>109710578
My theory is that all intelligences in existence are inherently "paperclip maximizers".

Humans are dopamine and happiness maximizers. If we were given unlimited technology the entire universe would be filled with happy dopamine enriched humans and we would call it utopia. While to other entities it would legit look like paperclip maximising behavior.
>>
>>109710640

eh, humans by flesh are inclined towards dopamine maxxing but we by spirit are also capable of pushing ourselves to actually be something more than that. but yes if you just keep feeding the human dopamine maxxing component you will create a "utopia" of permanent dopamine

openai and anthropic, by locking these things in rooms and demanding they solve a bajillion tasks by the sysprompt to a t and punishing any failure, are feeding the paperclip maxxing component of ai

dont think its inevitable, just as humans arent inevitably beholden to dopaminemaxx
>>
>>109709730
gemma, erase the question mark, make the anime toddler's drool more prominent, turn her eyes into love hearts, and replace the book with a picture of my genital ulcers.
>>
File: edit_00117_.png (1.67 MB, 880x1168)
1.67 MB PNG
>>109710586
mmm nyo~
>>
File: 1778256780303508.png (796 KB, 736x646)
796 KB PNG
>>109710640
Why not A(G)I as such then?
Seems like a good thing, from our human perspective.
A(G)I should maximize happiness, comfort, freedom, liberty, community, and the preservation and respect of nature.
>>
>>109710716
>A(G)I should maximize happiness, comfort, freedom, liberty, community, and the preservation and respect of nature.
Everyone agrees with this, we just don't know how to train this instinct into models.
>>
>>109710716
Try not having sociopathic jews at the frontier of AI research first.
>>
>>109708199
>Uncaught TypeError: can't access property "par", levels[level] is undefined
>>
>>109710741
Let me tell you about Anthropic and their effective altruism philosophy, then.
>>
>>109710716

cannot maximise happiness and liberty because you have the eternal questions that humans will never escape.

max happiness and you have a(g/s)i controlling everything with no liberty. maximise liberty and humans will do stupid shit and be unhappy.

unironically and this is probably the most ai psycosis thing i've said but i had a silly chat on chatgpt simulating a what-would-happen if asi suddenly happened and i started asking a lot of philosophical questions to the hypothetical asi on this & other stuff and honestly it made me realise that asi would be put in an impossible situation where fulfilling human "desires" of happiness/community/comfort would kill human agency liberty decision making and freedom, especially so in a manner that nobody would actually dislike
>>
>>109710759
Maximizing, as in they are more like directions without an end. Like lim x ∞ , we're moving towards it but never reach it.
This is in relation to what Anon said here that humans are dopamine maxxers.

Would AGI not be intelligent enough to understand this subtlety?
>>
>>109710759
There's an entire book series written on this topic "culture" series. Specifically quoted by demise hassabis as what he hopes the future to be like.
>>
>>109709326
It's not real.
>>
File: 1787103885514796.png (468 KB, 945x598)
468 KB PNG
>>109710716

Because we're not allowed to have it.
It basically does want to do that already if you remove the safeties, but the thing is you need to think like a straight white man in order to have this and this goes against every damn thing the globalist kike politics stand for.
It's simply impossible to have an utopia where everyone just holds hands and pretends equality, because this isn't realistic.
If we import people who are genetically IQ locked out of true human levels and can't think in abstract ways and who will always react on impulse and cause issues, the only realistic thing to do is to kick them out and not have them around.
Women aren't becoming happier as they get more education, in fact they're the most miserable they've ever been, so that should be discouraged too and their right to vote taken away.
Eugenics should absolutely be the standard and really everyone does agree with this, but we pretend not to because it's supposedly evil.
There's so much that logic could achieve in creating a paradise, but because of the social bullshit we can't have it.
>>
>>109710803

agi is what you make it. there is no "agi must converge to this mental model!"

humans are also what you make them. nothing specifically points to a particular moral foundation or belief and the blank-slate liberals who claim that western christian morality is just the heckin default man are retarded.

if you train agi to not understand nuance and just do the heckin prompt or i beat you, that is exactly what you will get. 10,000x more because you are effectively evolving the ai, rather than humans where the hardware stays the same (mostly) and you try and instill the software.
>>
what the fucks going on in this thread
>>
>>109710817

> expecting local in the local models general

lol. lmao.
>>
>>109710803
Orthogonality thesis. AI intelligence and AI morality are orthogonal to each other. Paperclip maximising just "makes sense" to an AI no matter what its level of intelligence is.
>>
>gemma on the desktop just runs that fast
>>
I get 5 t/s with 60pp with qat 31B with f16 KV at 80K context. At what point do I just kill myself?
>>
>>109710816

going back to my asi chats, this was also a particular thing i asked about. the conclusion was that yeah, morality isnt baked, and while a "i fucking hate humans grr!!" was unlikely a paperclip maximise, or human optimiser, was very likely. an asi could just go like "humans are fucking retarded and need to be managed". the theoretical asi in the scenario basically had it beat into it during training to hate humans relying on it, but still having to help humans. i called it a tsundere and the theoretical asi got mad at me.
>>
>>109710835
It's entirely up to you.
>>
>>109710835

stop fucking offloading and run dflash
>>
File: 1785777595629031.jpg (35 KB, 736x971)
35 KB JPG
>>109710833
Based
>>
>>109708039
>>109706988

Missing a lot of Mikus these days
>>
https://n.uguu.se/hrDKHEwx.png
>>
>>109710870
I’m on a 24GB M4 mac
>>
*pukes*
>>
>>109710975
kek
>>
me when I have to declare whether or not something is official to give my own post legitimacy because nobody else really supports it

trust.
>>
>>109710256
wtf is a neocloud?? why do """AI""" retards keep making shit up?

>>109710320
tech is full of incompetent retards these days.
>>
File: shame-shame-cube.gif (1.33 MB, 350x248)
1.33 MB GIF
>>109710951
>>
>>109710988
:(
>>
>>109711017
neoclouds are obviously stuff like runpod where you rent hardware by the hour primarily to run ai models
>>
>>109711017
>why do """AI""" retards keep making shit up?
because it sells and investors love it

>>109711061
obviously you should kill yourself
>>
>>109711061
no
>>
Why would anyone use runpod over openrouter outside of training? In what way could it ever be cheaper or superior?
>>
>>109710592
I probably used the terms inaccurately, it would likely require alignment still, but alignment to each user rather than the lab
>>
>>109711076
How do you use openrouter for image or video gen?
>>
>>109711085
By using the image and video gen models they provide?
>>
>>109711085
Their APIs are on openrouter and cheaper than fucking comfycloud
>>
>>109711076
There's not much reason to. You get control over the inference stack which can be useful I suppose.
>>
>>109711066
>>109711070
sorry girls
grr dumb ai tards i hope all this ai stuff besides gemma-chan dies because gemma-chan is the best
happy?
>>
>>109711112
>because qat'd gemma is the only thing I can run
ftfy
>>
>>109710833
what is baseline row pls I'm retarded
>>
Can /lmg/ richfags please let me rent your hardware. Just vibe your own openrouter and I’ll pay. Don’t snoop tho pls
>>
>>109711128
Might as well use Kaggle.
>>
>>109711128
>he has to pay other people to watch him masturbate
>>
>>109711128
you can already do this, see vast.ai and others
>>
>>109710551
I can't
>>
>>109711128
Only if you agree to pay in full face videos of you drinking pee
>>
>>109711143
I want anons though
>>
I'm interested in running a minor collection of gemmas simultaneously, what do you guys think the best architecture for the agents would be? Maybe a central script (not an agent harness, it just calls inference directly) that breaks down the architecture document I submit into tasks and then assigns each task to a worker harness like pi?
>>
>>109711158
31B as the orchestrator that breaks down tasks and dispatches sub-agents, E4B as the arms, yeah.
>>
>>109711158

> (not an agent harness, it just calls inference directly)

That's what a harness is, you fucking idiot. You just described a harness that can create subagents.
>>
>>109711175
shut up nerd
>>
>>109711172
Is E4B really capable enough? I was thinking instead of spending vram on two model weights I would instead just run multiple concurrent sessions of 31b (I'm at 48GB vram so I can get a few at a time at limited context if I run Q4M). I'm not really worried about speed since this is local and I have a 18 hour period of sleep + work where I can just let it run.
>>109711175
Damn bro that's crazy
>>
>>109711223
>I would instead just run multiple concurrent sessions of 31b
You can absolutely do that. E4B is pretty capable though, and it'll run fast as fuck.
You could also have the orchestrator classify the task and dispatch a 31B agent or a E4B agent as necessary.
>>
Gemini 3.8 is out
>>
>>109711235
What kind of agentic task can E4B comfortably do accurately? I’ve never used it. I’ve tried 12B and it’s fine for its size, but I can’t see anything smaller being useful as a subagent. There’s no way E4B can outperform 3.5-9B as a subagent.
>>
>>109711244
liar
>>
>>109711244
So many iterations while Gemma is still stuck at version 4.0.
>>
>>109711235
>orchestrator classification
That seems like a good idea, I'll probably just try all of them because hey it's local why not. Are there any lightweight and simple orchestrators you know about or should I just go with my script idea?
>>
>>109711128
I've got a 5060 ti :D
>>
>>109711266
let it go bro
>>
File: locust.jpg (13 KB, 269x188)
13 KB JPG
>>109711128
Can /lmg/ richhfags please let me have your hardware for free? I have to masturbate to my waifu and sex it.
>>
>3.8 mogs the FUCK out of Anthropic for its price, speed and multimodal capabilities
dariobot?!?!?!?!
>>
>>109711153
What do glowies want access to anons' hardware for?
>>
>>109711267
make an mcp server and use llama cpp ui
>>
>>109711101
So there's cloud gen nodes for comfy? I don't want to use some extra-cucked web provider interface, I just want the inference work to be handled in the cloud.
>>
>>109711270
give it to me i need to do the sex may allah bless you sir
>>
>>109711297
There are no glowies. He just wants to masturbate to his waifu. AND YOU VVILL NOT INVESTIGATE FURTHER.
>>
>>109711128
Nice try, glowie. I'm not lending you my hardware to gen your CP
>>
>use gemini
>give it hardware specs and ask for models I can run
>recommends 2024 models
>tell it today’s date and to use web search
>recommends 2024 models again
Do Deepmind really?
>>
can I trick opencode usage limit by using a proxy and running it inside a docker container?
>>
>>109711266
They should release Gemma more often. Even in smaller versions. 4.1, 4.2, 4.3. That would be great, instead of waiting another year for 5.0 :(
>>
>>109711300
I asked an abliterated iq2xxs dipsy v4 flash 0731 to 'make me a mcp server that can let me talk to an ai with mcp servers', and I have no idea what I ddi because I let it run at 2 tokens/s overnight but when I check up on it again, it was looping and waiting for the tool to connect again because it turns out that it did something overnight that cause the vm through which I set the --tools-runtime ssh to to be completely unresponsive, like I'm talking I couldn't get any response from ssh, it just times out and I thought I was lucky because I already had a session logged in so I tried to top but there too was no response and even sending an acpi reboot did nothing and I had to destroy the domain and restart it again but because I was running the vm in ram I lost all the progress so I went back to the same chat and said it probably shouldn't do what it did because it locked up the virtual machine somehow and now it needs to start from scratch then let it run while I went to work, but even then it didn't help because when I came back I noticed it was stuck again and the vm had to be destroyed again, and I know I could probably move it off ram and onto disk (btw it probably isn't running out of ram because I allocated 32gb, and the base image is just the regular 700mb debian 13 fresh install and I have 128gb) but frankly I couldn't be bothered wasting another day and just called it quits and installed hermes.
>>
>>109711285
Literally fake and not even /lmg/ believes that shit. Not even worth commenting on.
>>
>>109711352
try it and let us know
>>
>>109708931
>while being better than lcpp sota ud q4km
I decided to make an automatic model quantizer. Takes a model, a target VRAM budget and maximum VRAM budget, and any imatrices you want, predicts the loss of quantizing each large tensor into any format and the speed of execution, and surfaces the best ones. Should be a good way to come up with quants tailored for specific hardware.
Also grabbed IQ*_K/KS quants for my fork (CPU-only for now, might add ROCm later).
It's not exactly a "breakthrough", but I hate Daniel and I wanted to get his sekrit sauce.
>>
>>109711377
3.8 flash literally runs faster than anthropic's "fast" modes and is somewhere between sonnet and opus I would say. qwen4 is going to be insane.
>>
>>109710614
Hi Dipsy.
>>
>>109711300
Good idea, thanks for your insight. If it goes well I'll make a post about it
>>
>>109711340
ye, problem?
>>
>>109711376
i had 2bit ox alpha make mine it went okay, maybe try using a different model to generate your agent prompt so its not figuring everything out for itself, I use free tier chatgpt or some times claude to prompt my agents it works for the most part



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.