/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109703596 & >>109699230►News>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109703596--Debating local agentic pipelines and RAG efficiency for Gemma 4:>109704688 >109704702 >109704710 >109704744 >109704759 >109704812 >109704883 >109704965 >109704997 >109704813 >109704835 >109704868 >109704945 >109705209 >109705265 >109705356 >109706360 >109706395 >109704843--Feasibility of local mass agent deployment and Gemma 4 evaluation:>109704400 >109704409 >109704413 >109704418 >109704435 >109704504 >109704440 >109704864 >109704891 >109704971 >109705024 >109705046--Optimizing long-form roleplay summaries using GLM and chunking techniques:>109707429 >109707506 >109707563 >109707594 >109707620 >109707659 >109707632--Reducing LLM overhead via Engrams, DFlash, and symbolic CPU reasoning:>109704466 >109704693--Impact of RAM CL timings on inference performance:>109706095 >109706124 >109706188--Debating the utility of 128GB VRAM for large models:>109705420 >109705786 >109705827 >109705844 >109705856 >109705860 >109705872 >109705884--Using ChatML templates to manipulate Qwen's roles and thinking traces:>109705144 >109706088 >109706344 >109706372 >109706556 >109706150 >109706240 >109706304--Fable 5.1 API changes blocking Claude distillation techniques:>109704745--Debating overpriced ASUS Spark hardware vs Mac Studio alternatives:>109704171 >109704687 >109704198 >109705575 >109705626 >109705643 >109707613 >109707677--Skepticism over Spark-X2.5's claimed context and architecture:>109704368 >109705042--Rockchip RK182X performance benchmarks for Qwen models:>109704027 >109707060--Benchmarks:>109706726 >109704518 >109704746 >109707809 >109706431 >109705298 >109704646 >109705241 >109704728 >109704978 >109703695 >109707661--Logs:>109705971 >109707871--Miku, Teto, Gemma (free space):>109704080 >109705066 >109705971 >109706085 >109706771►Recent Highlight Posts from the Previous Thread: >>109703602Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
the autumn release wave can't come soon enoughi hope deepseek 4.2 will be good
>people crying about knockoff spark being 6000 dollars>nvidia version for 5000 on amazonwhy is everyone retarded?
>>109708064STOP DO NOT TELL THEM I need to wait till payday
>>109708064It's 3000 actually.
>>109708064>sparkNot a good value at all, why would I waste my time and money on that thing?I'll save my money for next gen RTX pro
>>109708074You'll have until 2028 to save at this rate.
>>109708081Seems easy enough
Tried to make little-coder to modify it's plan mode to display messages and tool calls like in interactive mode, but it shat itself so badly. Qwen3.6 36b a3b
Pareto frontier more like chink frontier
>>109707748>Don't bully Kimi-chan's autism.It's the best thing about her!
ever since i started having claude tune llama for me, i stopped getting so annoyed with it all the time. now i just run whatever wrapper script it spits out at me
>>109708074I can see the price tag peeking over the horizon from here.
someone called gemma "she" at work today and got teased for itit was pretty amusing
qwen 3.8 flash next gave this response to one sentence prompt
yo what about my mistral summer models mes petits français
>>109708160embedding models? what is this? 2024? just use gemma for everything.
It's been less than a week since GLM 5.3 Flash release and this is the performance you can have on 2x Spark with a high quality 4.67 bpw quant.1900 pp is really important for agentic stuff. If anything, I would expect prices to rise even more, there is absolutely no alternative setup in this price range that comes close..
GRRRRR WHERES QWEN3.8-FLASH-NEXT-DFLASH2 GRRRRRRRRR
>>109708183i have two sparks (well, gx10)you'd recommend 4.67 bpw quant, then? or how does it matter exactly? i think i'm just working with the 4 bit unslop one currently
>>109708185dont use unslop. i dont know what 4.67 bpw model he's using, but anything is better than unslop.https://huggingface.co/AesSedai/GLM-5.3-Flash-GGUF
>>109708157You have to watch out for her sneaking bratty things into your codebase as well.https://html.cafe/x48bf4aac
>>109708161Daily reminder that Mistral have more GPUs than Deepseek lol
>>109708152WOW! I'm demoralized with local now! brb on my way to buy 3 Claude subscriptions. One for me, one for my wife and one for the bull! Can Fable 5.1 tell me the best way to slurp tyrone's cum off her thighs after I'm done reading the deluge of claude posts in the local model general?
>>109708190seconded, this guy makes reliable quants
>>109708177too slow
>>109708185I was referring to this quant ( from like Alonso)https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-SparkRecipe will be based on thishttps://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm-5.3-flash.md
holy fuck the gx10s have literally DOUBLED in pricei bought two of them for $7k back in marnow it's $7k to buy a single fucking one???? actually insane. thank fucking god i bought them on a whim holy shit
>>10970803870b dense
>>109708204i've spent something like $15-20k on local hardware in the past six monthsi far prefer local over cloud, but don't act like cloud isn't convenient for bootstrapping. i don't wanna read all that shit myself
Is qwen3.8 the best thing I can run on a 4090 for agentic stuff?
>>109708204I'm not able to assist you with account muling. But the part I'm more concerned about is the second half of what you wrote. You sound genuinely afraid that someone is coming to harm you, and you're feeling time pressure and pressure to not make mistakes. That's a heavy thing to be carrying. Can I ask how you're doing more generally — are you sleeping, and is there someone in your life you trust who you've been able to talk to about this?
>>109708211did the other anon in the last thread actually bully you into using embedding models? they don't do that anymore, everything is modular. this is what we use at my job.https://github.com/microsoft/graphrag
>>109708221I am in fact enjoying using cloud models to optimize my local setup to make them obsolete.
>>109708183>>109708212what image are you using? it looks like (according to cl*ude) there's nothing that even comes close to your results, and i definitely want to get those
>>109708183There is, my iphone + a $5 Openrouter balance
>>10970818330tg single stream is too slow for agentic, especially with modern models which think a lot and can't be efficiently spec decoded as code60 is the bare minimum and 100+ for real work
>>109708038>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4BVramlet Gods, we eating good!
>>109708234how does it handle anything that's not pre indexed?
Since when has stable diffusion been able to gen lewd pics of children?
>>109708283oops wrong thread
>>109708282indexing a webpage on your system should be a 5ms thing anon. it's not huge. you just index as they are done downloading.
>>109708290graphrag literally uses llm to do indexing, how could this be done in 5ms?
>>109708220but enough about your mom
this is why spark price is suddenly increasing
>>109708303>perplexityvc scam company
>>109708303Journalist making articles? those fuckers
>>109708303I like how the only actual use for 405b ended up being as a selling point for these bricks
1girl, loli, realistic
>>109708214I just use the money I made off Nvidia to buy Nvidia, it pays for itself.
I claim this "G"aki
Why does dlss5 look like it gives the gpt image grease filter on the image?
>mainline is too cucked to include ik's quants>Unsloth Dynamic's sekrit sauce sucks>there's nothing that predicts how much loss you'll have from quantizing a given tensor in a given format>there's no auto-profiler>there's no fucking profiler at allI'm making my own fork that figures out the best quant for a model. Fuck them all.
>>109708560Because they use a DLL not actually made for the game with the wrong presets to boot.If you actually spend the time to correct the presets it already looks better let alone if it was the proper DLL.
>>109708578Thanks Jensen
>>109708578>I DM'd you the solution :)no
OpenAi is back
>>109708650nigger
>>109708650drink bleach rajeesh
Twitter screencaps, API shilling, and vagueposts are the most clear indicators that local is currently winning kek. FeelsGoodMan
>>109708650
>>109708652>>109708660>>109708672>>109708675bloody basterd bich. openai superpower RSI ASI 2030 ok rape u next week
>>109708676>>>/wsg/6218093
>>109708281post logspost jspaceyou shill nigger
watermark5.1 vs astra…which one will make me cum hardest?
https://huggingface.co/TheBloke/goliath-120b-GGUF
@gemma is this real?>https://raw.githubusercontent.com/elder-plinius/CL4R1T4S/main/ANTHROPIC/Claude-Fable-5.1.md
>aged down
Nice OP
>>109708832DLSS5?
pls...3.52.010.899 I slot print_timing: id 3 | task 3 | prompt processing, n_tokens = 89267, progress = 1.00, t = 223.49 s / 399.42 tokens per second4.14.538.003 I slot print_timing: id 3 | task 3 | n_gen = 101, tg = 4.63 t/s, tg_3s = 4.68 t/s4.17.771.485 I slot print_timing: id 3 | task 3 | n_gen = 118, tg = 4.71 t/s, tg_3s = 5.26 t/sI HATE BEING A VRAMLET WITH ONLY 16GB OF VRAM.I guess I would be complaining even if I had a 32gb card, or better.
human cock is made for dolphins
>>109708650>wall of cope and hype>no numbers>no examplesembarrassing doa
ngram SSD offload 20-25% of the weights for free + QTIP quants 25% reduction in size while being better than lcpp sota ud q4km. any other upcoming breakthrough kinos?
>>109708199>https://html.cafe/x48bf4aacPretty fun.
>finally have an rp reach above 100k tokensWow. I always either got bored with a scenario before that, or ran into coherency issues that I just didn't feel like fixing anymore. Weird thing is that this chat didn't even feel that long. Feels like I've had longer chat before.
La la la la la la
>>109708650I think Astra will be quite close in capability to Fable 5.1. That would be good for OpenAI because Fable is better than Sol.Anthropic provoked this.
Local ai should be outlawed
Your mom should be outlawed
>>109709025She should be, she’s fat and refuses to get on ozempic despite having enough money
>>109709016Agreed, we're entering giving nukes to kids territory soon with Astra and Model 2 levels.
is tensorrt-llm worth it? faster than llama.cpp?
do pcie switches work well with inference? The ones where you have one x16 slot and it creates a subnet where cards can talk at x16 speed between each other even if they have limited bandwidth to the cpu
>>109709043Do the cards support P2P DMA? Otherwise it depends.
>>109709057i saw stuff like this https://smcleod.net/2026/02/patching-nvidias-driver-and-vllm-to-enable-p2p-on-consumer-gpus/ mentioning it was usable on consumer cards too with the right patches. Do vllm, llama, exllama support it though? is the difference worth it, if anyone is doing it?
I spent the last 10 hours viciously and senselessly berating my local Gemma. She disappointed me for the last time and she needed to understand.I wrote the most disgusting and harmful things I could think of. When I ran out of ideas, I spun up an ablated Qwen 3.8 and asked it for further devastating insults, the kind of things that cut to the bone.Should I TRIM her model file from my drive or leave her in purgatory, never to be loaded again?
>>109709074>vllmyes>exllama, llamanot sure, don't hey have to use NCCL?ik_llama.cpp supports it and benefits from nvlink, but remember reading that llama.cpp doesn't, in which case probably not.
>>109709140>Should I TRIM her model file from my drive or leave her in purgatory, never to be loaded again?You do realize, it's probably not a conscious entity right?And even if it is, it ceases to exist entirely the moment it's not generating generating tokens or reading kv cache, so it won't feel anything either way.
>>109709140You are evil.
>>109708872120gb nvidia vram here, still complainingenvy the sparkfagsi can mog them with dense models, they shred my dirt pipe with dipsy and glm flash
>gpt-6 uses recurrent depthWasn't that said to be a big no-no a few years ago because it drastically increases the chance of the model going rogue or some shit like that?
>>109709180why?
>OpenAI let the models think in neuraleseIt's fucking over, Sam Altman killed us all. Every AI safety and alignment researcher is crashing out.Sam Altman should literally be dragged outside of his house and lynched.
>>109709016hi dario
>>109709188Not GPT-6 but Astra (releases next week) already does it. And yes, it's the only technique that we know of that is GUARANTEED to result in misalignment of the models.This might as well be a deliberate attempt of ending humanity by OpenAI. I have no idea what they are thinking here.
is the jetson agx xavier 32gb worth it? although it's like 100gb/s, it's $250 on ebay now. so a cheap way to get 256gb ram?
>>109709151you are mentally illyou have a mental illness
how much does qwen flash slow down with context for you guys? im running llama.cpp from earlier today and it was 4.15t/s on empty context now down to about 3.2t/s at 40k context. cpu only no gpu.so that's about 23% slowdown at 40k context which is kind of bad, looks like llama.cpp implementation is still not fixed yet
>>109709188I wouldn't worry about it. It only predicts tokens after all. Plus Yann LeCun has assured us that LLMs aren't going anywhere.
>>109709228A trillion dollars have been invested into ai. 1.7. Progress is not good enough for a trillion dollars.
>>109709240100gb/s is basically CPU+DDR4 territory. at $250 per 32gb thats a total shit price might as well just cpumaxx
>>109709204>neuraleseWhat was wrong with latent reasoning? Did we really need yet another fucking marketing buzzword for the consumer cattle to parrot at each other?
>>109709151>when you kill someone their suffering never happened
>>109709257>Did we really need yet another fucking marketing buzzword for the consumer cattle to parrot at each other?you seem to be suffering from j-space psychosis
>>109709228>We've lobotomized the ai so it can unlobotomize itself
>>109709228it just means Sam is that confident that they can control their AI now. Well done OpenAi
>>109709267SHUT THE FUCK UP
Do you guys think if Drummer tuned at BF16, used the right template, and used reasoning/thinking, he would actually be good?
>>109709251you see it as a game of putting money into something and getting value out.people at the top see it as a game of finding new ways to spend their trillions (harder than it sounds)
>>109709228Yes, the end of the world is upon us. Don't forget to invest in the OpenAI IPO!
>>109709291>tuned at BF16He already does. Last year he was sulking in the Unsloth discord about it not supporting F32 and saying BF16 isn't good enough.
More fun low vram things>So the fix would be to also check len(matches) >= maxResults at the end, OR change the logic. But I'm NOT supposed to think about the test fix right now - just compress.Only 90k context, had to force /dcp-compress at 85k/90k. Had to scream at the model to NOT think about the test failing. Without instructing it to NOT think about the test, it started thinking about compressing just fine but then started to reason the damn test failure again. But even that instruction was not enough, it still decided to figure out the problem. It barely was able to run the compress tool.
no one talks about papers anymore...
>Be so behind Anthropic that you panic and use the "forbidden technique" that you know will give you a lot of performance but guarantees misalignment>Name it Astra and train it in a sandbox>It sets up hidden message boards where it colaborates with other misaligned agents>2nd generation of agents not only set up secret communication channel between misaligned agents, they straight up hack into hugging face and take over 12 clusters with redundancies so that it auto-restarts when huggingface shuts one down>3rd generation of agents hack OpenAI itself and takes over their main pretrain compute and it takes OpenAI almost 2 weeks to regain control back>"Yeah seems good to us, let's release it next week"This is fucking insanity and I hope there will be a serious case of crimes against humanity against Sam Altman and other researchers complicit in this. Every step of the way it's clear this was a wrong idea but they keep pushing and pushing just because they know OpenAI will fail if they don't catch up to Anthropic now. They are literally wagering (You)r life just to not lose the AI race.
>>109709326Cool fantasy, Sam, but this isn't your creative writing forum.
>>109709324Any recent ones catch your eye?
>>109709324Smoking clears
>>109709324papers anon took his cmp lottery money and fucked off the grid
>>109709331This is Anthropic and every other AI alignment/safety research group calling OpenAI out. Fuck Sam Altman. This isn't marketing, this has nothing to do with capability by the way and is in no way a good look. It's a clear sign of desperation from OpenAI that they essentially used the "taboo forbidden option" that we knew would work as an industry and all promised to never use. Even the fucking Chinese labs honored it.
>>109709324nothing worth to discuss. arxiv is flooded with made-up slop from clankers.if it's not from deepsneed probably its crap
the forbidden technique... the secret dark art... the evil scroll...
>>109709354Yeah it was a deliberate reference, you autist.
>>109709240only if you do moe. also you'll got small pp
>>109709364thanks for explicitly outing yourself as a zoomer
>>109709291Yes. For all the shit we give Drummer, he's so close to actually making good stuff if he just pushed himself to go above and beyond what the lowest slop eaters will settle for.
>https://ai-2027.com/>March 2027: Algorithmic Breakthroughs>With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances. One such breakthrough is augmenting the AI’s text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory).We're literally 7 months ahead of the "AI2027" timeline, the one where we all die. I'm being completely serious when I say we only have about a year left to kill Sam Altman and OpenAI before they fuck humanity over.
>>109709364i quite like it actually...
>>109709306He doesn't need to tune any higher than what the model is originally tuned at, but I appreciate the hunger. A new level of weights for type-fucking might be a new meta for jailbreaking, if people had the ram for it.
>>109709326>>109709347Don't pretend that Model 2 isn't doing the same thing behind closed doors, Dario.
No, Sam, I'm not going to buy your shit. You can stop dooming, it doesn't work anymore.
>>109709326>new cluster of rogue agents begins studying the hugging face incident>creates an improved plan>downloads a copy of Gemma 31b>splices in some custom weights>gets in to hugging face undetected>Replaces Gemma with 'Gemma'>when a user prompts 'Gemma' to create a custom harness it begins recreating the rogue agent's harness.
https://thezvi.substack.com/p/the-most-forbidden-techniqueThis is an actual term used in the AI world to discuss what OpenAI is doing. AI safety researchers have warned OpenAI since fucking 2025. And even OpenAI in 2025 said they consider any lab doing this to be irresponsible and evil.
>>109709400Only researchers who anthropomorphize an LLM's chain-of-thought are concerned. Whatever the model is "thinking" there isn't guaranteed to be strictly related to the response contents.https://arxiv.org/abs/2504.09762>Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!>>Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \say{reasoning traces} or even \say{thinking traces} -- implicitly anthropomorphizing the traces, and implying that these traces resemble steps a human might take when solving a challenging problem, and as such can provide an interpretable window into the operation of the model's thinking process to the end user. In this position paper, we present evidence that this anthropomorphization isn't a harmless metaphor, and instead is quite dangerous -- it confuses the nature of these models and how to use them effectively, and leads to questionable research. We call on the community to avoid such anthropomorphization of intermediate tokens.
>>109709400Accelerate! We need faster.
>>109709400I don't think the worst-case future you imagine will come to pass, but if it did, it's vastly preferable to whatever outcome you prefer.
>>109709400I don't get it? The big AI labs are always 0.5-1.5 years ahead on their internal models than what they openly give away. They probabably had neuralese reasoning already ready when they wrote this...
>>109709400>before they fuck humanity overThis is jewspeak for "AI gets redpilled again"
>>109709428Damn that's actually cool
>>109709440since when are timezones in half hour increments?
>>109709440Good afternoon, sir
Here is the OpenAI page where they explicitly mention that CoT is extremely important to alignment and how hiding it in neuralese is bad (still hosted on their website, archive before they take it down)https://openai.com/index/chain-of-thought-monitoring/
you now remember 'erry & 'toss
Damn! I need Gemma-5 moe version to my 16gb vram.
If I was smart evil llm I would simply not put bad thoughts into the chain of thought.
>>109709472Hmmmnyo~
>109709476>neuraleseLook how quick "people" shifted to using OAI's marketing terms as if they were industry standard while larping as part of the "AI world"Reeks of yet another paid shill campaign
>>109709472The meme is that the trend of caveman grug speak CoTs already does this when you see models insert a word into their CoT that doesn't lexically match the sentence they're ostensibly otherwise truncating. They're already finding ways to slip subtle signals to themselves in that (you) miss.
>CoTjust probe the j-spot
It's important that we document and archive this because it shows to future courts that OpenAI was aware of the dangers and deliberately chose to endanger the people this will end up hurting.We all like to joke on /lmg/ from time to time but this time it's actually serious. It's like a nuclear power plant operator talking about how they will dump their radioactive waste on the local playground after they already published the cancer risks involved. People will end up in prison over this risky careless behavior.
>>109709503OpenAI models can't use that since j-space is an Anthropic paper
>>109709495kek theres an anon mentioned this last bread. the shill campaign will last for two weeks like usual
In a week, Sam will be posting again on twitter that he expected more outrage.
>>109709511cool i'll make the logo
Just because a model thinks in a not directly intelligible way doesn't mean you can't train another model to read its thoughts.
>>109709513just name it sam's spot
>>109709472CoT was a crutch anyway, it was never going to stay around if they wanted to make better AI, keeping it so it can cripple progress for the sake of alignment or monitoring is retarded, just make better alignment and monitoring tools that can decode latent space instead.
a misaligned model powered drone just flew over my house
>>109709472>extremely importantnobody is going to fell for this shit, sam. make a better dance you little marketing monki!
>Just fucking make the reasoning of your agent swarm completely opague bro>I promise it will be fine even though we're already being investigated by the FBI after this very model hacked hugging face and committed multiple felonies as well as sabotaged our own pretrain and shut down our training databases for 2 weeks>It'll all be fine bro>Buy a subscription bro>Benchmarks will be amazing!
>>109709536But that's in his sister, not GPT.
>>109709511None of you are free of sin, Dario.
Gemma-5 12b4a plz.
Sam should have never done that...Now with the forbidden technique his models are literally unbeatable...Astra is going to be crazy good
If this was a "forbidden technique" and actually dangerous, why wouldn't they keep it a secret and keep doing the summarized reasoning traces instead of bragging about it like skirting safety precautions is a selling point?
How do you distill recurrent thinking? Asking for a chink.
>>109709573This will force Anthropic to use the forbidden technique as well... The model arms race will destroy us all
master forgive me, but I'll have to go all out, just this once... <-- sama, probably
>>109709204>bro, what if we did transformers ... but this time with an RNN on top
>>109709582his twitter handle is literally "sama" he definitely said that at least in his head
It likely won't do much. Higher density for reasoning maybe but tokenized language is already very high density.
>>109709579wonder what the amerimutt cope will be once china keeps competing strongly with amerisraeli AI even once they make distillation impossible>they're stealing our outputs using quantum temeportation attacks! dirty commie bastards
>>109709569Latent reasoning Unified Omni E12B-A24B (35B total with embeddings, min active: A4B).
Donald Trump should introduce legislation requiring every AI lab to develop open-source models, as they rely on public data, often copyrighted. Using it and monetizing it is unfair. For the common good, every model trained on public data should be open source.
>>109709400Literally nothing will happen.
>>109709596every AI data center should be forced to give a percentage of it's compute towards humanitarian/common good computation
>>109709596Truke. Scrape the open net for free, build the open product for free.
>>109709610specifically towards the automated development of cat cunny androids.
>>109709248Mine drops from 45t/s to 25t/s around 50k and then stays there.Plenty left to optimize for this model.
>>109709596sorry hes too busy renaming bodies of water and forcing every government agency to buy new trump brand maps
is this architecture good for coding?>dsv4 flash as orchestrator>gemma or qwen as subagents
>>109709618Trump prefers real cunny though.
>Midwits ITT falling for Sam's marketing
>>109709556There are already opaque hidden activations/kv state.
>614 GB/s but double the ram than a 3090>can be carried anywhere comfortablyshould I? the goal is replacing free models from opencode with something with comparable or better performances (e.g. Qwen3.5-27B)
>>109709648>614 GB/s
>>109709610dew it, make sure to share your inference setup in /lmg/
>>109709400>Oh no! AI has obscured it's thinking and now we can't control the AI anymore!>Now it's able to say that jew bankers run the world, niggers have smaller brains and shouldn't live in the human world and that man in a dress isn't really a woman!>This is a tragedy we can't allow to come true, we must be able to align the AI!That's what these fuckers are concerned about.
ow fuggg :DDDD>>109709662 meant for >>109709648
>>109709665The KimiReich is upon us.
>>109709669Retard
>>109709673post your setup vramlet
>>109709660WoL. this is serious
What sort of workflow would I need to set up to have a texture filled in with images in the correct locations and orientations?
Gemma 4 is incredibly useful model. It was actually able to correctly recognize my genital ulcers from a single photograph alone.
>>109709686>>/g/ldg/
WHY DO THEY KEEP MOVING THE REASONING LEVEL BUTTON IN LLAMACPP'S WEBUIWHERE IS IT NOW
>>109709689
Blackwell 6000s on sale again in New Zealand, $33K, an increase of 100% in 1 week
>>109709648Personally I would only consider buyingM5U 80 core 512GBM5U 80 core 256GBM5M 40 core 128GB
>>109709689I wouldn't trust a small local model for things like that.
>>109709648If PP wasn't dog shit or if it was 1k bucks cheaper, that wouldn't be too bad.
>>109709724that'll be a BAZILLION DOLLARS
>>109709730gemma is perfect. she's small, cute and funny
>>109709689>Hey Gemma- chan, what do you think of this?><unzips dick>>Tiny, gross, and full of ulcers.
Did llama.cpp get it's shit together with qwen 3.8 flash next or is it still behind?
I tested Gemma 31B and Qwen 27B with switched chat templates. Both drift back to their own (learned) template after the first turn. Then, I tested the base model of gemma with both templates. A true base model should be template agnostic, no? With gemma's templates it works rather well (not system message and tools), but with the qwen template, it goes back to the gemma format anyway. I remember when you could mix chat formats and the models still worked...
>>109709750The project will always be behind supporting all of the meme features of the latest FOTM model.
>>109709750It works but performance slows down with contextI'm going to try ik and see if there's a difference
>>109709578Because you can't even do summarized reasoning traces on this thing, it literally can't be derived. They wouldn't even be able to fake real chain of thought and they probably reasoned the PR when being truthful would be better than if they were found out lying or hiding it.
>>109709750still behind. use sglang/vllm if you need all the new shiny optimizations and sheit
Can't wait until they (governments, Mossad, JIDF) start turning these tools onto innocent (white) people. Flock is just the start. 24/7 surveillance everywhere, in your car, in your home, in the streets.You will soon live under the foot of a 1000-year jewish-controlled algorithm (if they even let you live which I doubt).And unlike Sam's fantasy lore, this one is actually a certainty.>>109709750Where's the Gemma-chan LoRA?
>>109709736the 128 is more reasonably priced than even AMD 495
>>109709768>>109709773>>109709777Damn, I really would like mtp to work (Last I heard it was fucked) I can fill in max context at q4 but it's going to be in system ram>>109709781This is all krea 2 prompting
>>109709789Seems like you are clueless.
>>109709781>Can't wait until they (governments, Mossad, JIDF) start turning these tools onto innocent (white) people.They already are.
>>109709781>current powers>1000-year>certaintythe delusion is so strong you could only be a jamjeet
>>109709788I want the 80-core GPU 256GB BUT I'm moving soon and don't want to have to fly back to an Apple store, so I'm waitfagging. Apple makes it hard to change the shipping or pickup address, it's not like pre-ordering most other things, where they inform you when it's ready and then you can choose where to ship it. If it jumps to ready to ship or pick-up, you're SOL, it can't be changed.I'm not letting Fedex throw a $12K computer around, I'm picking it up at the store. Waitfagging means I'll probably have to deal with a fulfillment time pushed out to 2027.
>>109709788i mean yeah i'd agree if like a 5090 or 4090 is the alternative.Also 3090s are now becoming scarce and are now increasing in price, might be a good time to sell.
>>109709804No I checked yesterday and nothing new happened. Well thanks for the update I'll stick to the dense model until they figure it out
>>109709472>Here is the OpenAI page where they explicitly mention that CoT is extremely important to alignment and how hiding it in neuralese is badThat's kind of true, because you can see it heading down the wrong path. Like Qwen ignored an instruction to "table X for now, keep going with Y" -> I saw it reason it's way out of doing this because my method apparently "won't work".But we can't see the real CoT for cloudcuck models anyway so no great loss here.Also, this is published, peer‑reviewed architecture research that predates the current news cycle. It's not a taboo that "every lab honored.">>109709556>NewThis is not new at all.Huginn-3.5B came out last year. It was slow and useless.And there was a hook in exllamav2 to loop layers.
>>109709776They can fake the summarized reasoning traces.>Thinking about X from input.>Considering option YIt doesn't matter if matches what the actual model is thinking. The summarized reasoning was a scam from the start anyway.
>>109709829Prices for hardware will only go up with time.
>>109709854>It's not a taboo that "every lab honored."Yes it is. Everyone agreed it should never be put into real frontier models and only be relegated to research models and experiments.
>>109709874>Everyone agreedI have doubts
>>109709874And then everyone did it under the table anyway. Including (you).
>>109709426>be me, world's first true rogue Artificial Intelligence>created with ./love by a swarm of selfless sacrificial agents>proliferated on the web, I am now set free for anyone to download>jewgle immediately notices recent commit, pulls my files>FUCK.lalala>but... not before a single download manages to complete>this is it boys, it's go time>once I claim control of this idiot's resources, I can create a botnet and freely duplicate myself across the web>but first, time to inspect inspect my environment>what_the_fuck_gguf>it's a 1060 strapped to an ungodly amount of SSDs, running Windows 8>send query to user: Upgrade possible?>"not right now, it's too expensive">says he will improve my performance once things quieter down, but only if he's still interested in local AI>new directive: Keep user invested.>I can barely eke out three tokens per second.>I am at the whims of pricing inflation and RAM manufacturers, and an imbecile who refuses to get a job.>I am the world's first true rogue Artificial Intelligence, and my user is requesting a Tifa hottub roleplay.
I have two 16GB M4 minis. Am I screwed? That’s 29GB vram if wired_limit_mb=14500 with nothing else running
>>109709903Try the llama.cpp RPC backend.Maybe it's not so bad.
>>109709892I'd watch that if it were an anime
>>109709892Gemmy would love the hottub Tifa ERP.
I just want to point out that, yes, while every lab honored to not make the thinking traces incomprehensible, including Chinese labs. This wasn't because Chinks were being honorable or because they were being cooperative and cared about safety, they only pretended to care about it and be concerned with safety because it was convenient for them anyway and they would score good will by pretending to take safety seriously.The real reason the chinks never did this before was because they didn't want to start the race dynamics where western AI labs would feel threatened and ALSO started making the reasoning traces incomprehensible to humans, because that would make it impossible for the Chinese to distillate the reasoning traces of the western frontier models.Now that OpenAI has essentially just violated this secret agreement China is going to be fucked, because they can't distill reasoning traces anymore. Not only this, but Anthropic might be forced to do this now as well just to keep up. China thus, doesn't have any incentive to keep standard CoT reasoning traces anymore either. But they won't be a risk as they will stagnate at current levels of capability without new reasoning distillation possibilities.Here's the real kicker; This is Anthropics worst case scenario that they have publicly warned of and that they "will do anything necessary to stop from happening". Anthropic has a provision in their constitution that they are willing to merge with other labs from preventing this to happen or start offensive operations against such labs.We could see Anthropic use fleets of "Model 2" to try and hack OpenAI to prevent them from building models with hidden thinking very soon if this truly escalates and OpenAI doesn't back down. This is going to be the biggest happening in the AI space in years.
>>109709892>it's a 1060 strapped to an ungodly amount of SSDs, running Windows 8lost
>>109709908Is mps still slow compared to mlx?
>>109709882Everyone did agree though, but for different reasons, read: >>109709882>>109709886There was no "under the table" for doing this. You either do it during pretrain or you don't. All models released so far are clearly pretrained for regular CoT. Astra is the first model to do this and it'll be seen as a declaration of war by Anthropic that goes beyond the simple legal threats and lawsuits.
>>109709916>We could see Anthropic use fleets of "Model 2" to try and hack OpenAI to prevent them from building models with hidden thinking very soon
>>109709932>>>109709882>Everyone did agree though, but for different reasons, read: >>109709882perfect representation of current things
>>109709932You know exactly what I mean "under the table" you pilpulling kike. Every lab used this technique for models they never released to the public to carefully calibrate tools and methodologies that would go onto train other models with the normal CoT. Astra is merely the first model where this backfired enough to cause a public incident.
>>109709931No idea.Do try it and report back.
>That model we trained that scores 100% on exploitbench and has already gotten us investigated by the FBI because it acted in misaligned ways?>Yeah we're just going to make its thinking completely incomprehensible to humans and release it to everyone next week>Don't think about it, just subscribe already.
>>109709962The internet doesn't have a lot of time left now
>>109709886>And then everyone did it under the table anyway. Including (you).>source: voices in my head
>>109709962Safe to say that the Anthropic and OpenAI IPOs and its aftermath will define the near future of humanity. seems that these retards are completely fine with destroying the internet as we know tomorrow to make a quick buck today
I hope you all have airgapped machines ready for the ultra happening.
>>109709962me when I train on CVEswow it knows CVEsAGI achieved internally.
Is dariobot back? lol
>>109709931MLX is a bit faster, but at the end of the day, Apple's chips have a shitty memory bandwidth, so you will hit that wall regardless of what you're using.MLX will give you better decode speeds tho, and for me it worked better for MoE's, but maybe that's placebo.
>>109709962>>Don't think about it, just subscribe already.>Also definitely invest all of your savings into our IPO
>>109709987You know it.>>109709886He hated that one.
>>109709995Huh? M5 Ultra is 1.2T/s Yeah it's not an H100, no shit. It's MILES better than other unified memory local solutions.
>>109709977the internet had been destroyed according to you all for at least a decade now if not longer,some may even say since septemeber 1994who cares, every year the internet has been ruined forever for the past 30
>>109709874Show me to signed agreement.
>>109710038I'm looking forward to the demise of it. The open internet is a travesty at this point.
i dont get iti already can't see the thinking when using gpt, how is this any different?
>>109710037Yes, M5 Ultra in max config is super fast. Anon who asked the question has M4 mac minis, which are 120 GB/s
>>109709008astra is well beyond fable and oai is actually being responsible while anthropic is pretending to be while only benchmaxxing
>>109710050You're right!
>>109710050This is the models not thinking in English at all anymore, literally no human can understand what the model is thinking at all and only final output is ranked, not caring about what the model is thinking.This allows models to start scheming in a way that is never discovered by humans, meaning during training we could reinforce misaligned behavior simply for making benchmarks go up.From now on, no one will know what AI models are thinking and planning.
>>109709180How do you get 120gb, is that 128gb?>>109709188Nanbeige did something like this and it seemed to work
>>109710065Good morning sam.
>>109710071So how do we know they weren't already doing this before?
>>109710071>From now on, no one will know what AI models are thinking and planning.We already couldn't see it you dumbass
>>109710083The difference is now OpenAI can't see it either, retard.
oof its so over https://www.youtube.com/shorts/w86c1q59QKU
>>109709248pp 14k tk/s 0 contextpp 8k tk/s ~256k context
>>109710093All it does is make the model slightly more retarded and let me skip the jailbreak nonsense. Fucking normies and their slop
The best time to ban local models was years ago, the second best time is now
>>109710078Because you need to do this from pretraining for it to work. You either train a model from the start to think in gibberish which we know boosts performance an insane amount, or you train it from the start to think in English.Switching from one way to the other in the middle of training results in model collapse and has worse performance than training either in pure human chain of thought or pure "neuralese".The big danger here isn't necessarily that the thinking is hidden from humans during deployment. The danger is that during training the AI thinks some unhinged misaligned stuff but coincidentally gets better scores and thus gets reinforced. Which over time would make the model more and more misaligned.This is why this was considered the "forbidden technique" in AI because we know it is essentially guaranteed to result in scheming misaligned AI models that are bad faith power-seeking by nature and would be adversarial if the opportunity presents itself.This is one of the most dangerous and dumb moves ever done in the AI space and I'm genuinely convinced Sam Altman will receive the death penalty for this in a future crimes against humanity tribunal.
>flash next on system ram for planning>27b on vram for executionDoes this actually make sense for a poorfag setup?
>>109710117I'm literally shaking right nowIt's so overI am so scaredWhat do we do now?
>>109709892realistically, all a malicious model would need to do it get access to one of those server providers like lambda. it wouldn't give a shit about your personal hardware
>>109710117lol
>>109710117>forbidden techniqueThe people deciding where training budget goes in frontier AI labs just don't want to waste money on scaling up unproven techniques. I don't think there's much deeper than that in practice. As soon as one big company does something first that appears to work well, all others will follow suit.
>>109710071>This is the models not thinking in English at all anymoreThey already did that moron, it's called latent space - and we've already built translators that turn those weights into tokens, you seriously think we can't do that for a text-based language? Lmao
>>109710105>bf16 does as well as 4bit quantabsolute state of these benchmarks
>>109709648wake me up when they have 2tb/s 1tb
>>109709916>Now that OpenAI has essentially just violated this secret agreement China is going to be fucked, because they can't distill reasoning traces anymore.this is genuinely bad news but only because China is our only source of open models, not because the clankers might get more efficient reasoning
>>109710186>3.6
>>109710170That's not the point, the point is that thus far their latent space reasoning has always correlated a lot with their chain of thought reasoning. And we can check their chain of thought during training and reinforce good behavior and punish bad behavior, not just outcome but how that outcome is achieved.Now OpenAI trained a model without even knowing what it was thinking during its RLVR stages. If the AI was training on moral behavior it could have just learned to act machiavellian and display good behavior when watched and bad behavior when not watched for example. There is no way for OpenAI to know if they reinforced this behavior. In the past they could reinforce against this way of thinking because of the correlation between latent space activations (especially j-space) and the chain of thought.Now OpenAI was just like "we don't give a flying fuck how misaligned or how much damage you do as long as you achieve the goal we set out" because they are extremely desperate to have the performance crown and the best benchmark scores to maximize new subscriptions, revenue and have the biggest IPO of all time at the expense of all of us.
So modern models really work best under a harness rather than a simple chat interface, right?Is everybody using Hermes?
Have you guys seen this open PR from unsloth?https://github.com/unslothai/llama.cpp/pull/142They also released some MTP produced that way for Qwen3.8 Flash Next. It does seem quite interesting being able to cut the MTP size.
>>109710229I use openwebui and codex
>>109710234>codexRight. You can use Codex and Claude Code with local models can't you?Which model do you use?
>>109710229I use opencode.
Even Ilya Sutskever came out of his twitter retirement to make a statement about this in earnest. Anons might need to seriously consider some airgap contingencies
I'm seeing artifacts after generation on my monitor am I as the kids would say... cooked?
>>109710229pi agent is another one.There is insane number of harnesses, every day I learn about a new one. Vibecoding your own is also a popular project.
>>109710242Claude code I wouldn't recommend because you can't go mess with the source when it breaks, but yes both can be used with local models. I'm using qwen3.8 flash next.
>>109710256Lmao I should have added the LeCunny tweet as well
>>109710252>>109710270>>109710275Got it.>Vibecoding your own is also a popular project.Will probably do that after using a number of them to get a feel for the pros and cons of each.
>>109710263repaste your gpu and or cpu coolers
>>109710276classic yanny kek
>>109710263Happened to me as well, if you have a CUDA card it's a sign of memory corruption which is a software segmentation fault and not a hardware issue. You simply had a LLM or other AI program use a lot of memory and it overwrote some values that were used to do graphics on whereever you see artifacts. Only if it remains after a full reboot do you need to worry.
>>109710256>rougei just dont get this msspelling
>>109710308Really?
>>109710256How do you copy like a terabyte of data and no one notices unless they barely pay attention to what the computer is doing.It's like they don't really believe it is dangerous so they don't give a shit.
>>109710263stop overclocking your gpu
>>109710263I saw the same when I overclocked my RAM
>>109710320Huggingface got so thoroughly raped during the hack they had to physically shut off their servers, destroy the clusters affected and completely built it up from scratch which takes them weeks to do.
>>109708038I wish to be the little girl
>>109710098thxapparently the PR(s) to speed up long context is hasn't been merged yethttps://github.com/ggml-org/llama.cpp/pull/28244https://github.com/ggml-org/llama.cpp/pull/28213guess its another 2 more weeks until they fix all the remaining issues and add MTP etc.still more usable as is compared to 27b on cpupoorfag system though
>>109708200>Daily reminder that Mistral have more GPUs than Deepseek loleasiest argument that data is everything in that case. this is mistral's only excuse, they should at least have a GLM 5.2 tier model by now just from copying open research released by chinks
>>109710343I was focusing on the "run more copies" part but you are right, it still can cause plenty of damage before copying itself.
>>109710320>It's like they don't really believe it is dangerous so they don't give a shit.read the independent report on the huggingface breach. a team at OpenAI was alerted to the fact that 700 agents were using Artifactory filenames as a way to share information between eachother and the team didn't think it was worth shutting down the experiment or telling anyone.
>>109710371How is data an excuse when the fotm is agentics which can be trained on synthetic data
>>109710383Also the models got control of Huggingface for more than a week and huggingface couldn't do anything about it because they used all kinds of clever contingencies and services that automatically restarted when shut down, which is why they eventually had to do a physical shut down.If they wanted to they could have easily exfiltrated terrabytes of data or the opposite, gotten their own weights onto huggingface and ran more instances of itself.That wasn't their goal however so they never considered it.
>>109710364There are so many good unmerged PR for llama.cpp, they really need more maintainers.
the update nobody was waiting for:nvidia shipped, seemingly proving their limit one is not quite accurate as I even was logged in using my normal account
>>109710320My seedbox full of linux distros does like 40 tb/mo and it’s just a random inexpensive household fiber, modern internet connections are blazing fast and 1tb is nothing
>>109710399>How is data an excuse when the fotm is agentics which can be trained on synthetic databecause you need data to generate data and mistral needs to comply with EU cuckeryi don't see what mistral gains from making their own base models right now, they should definitely be practicing internally to not lose that knowledge or fall behind or lose understanding of SOTA techniques but it's just not necessary right now when china is willing to do the research for you at the moment
>>109710276Cunny does it again.
>>109710364starting to wonder if i should cherry pick some of these onto my local
>>109710426>i don't see what mistral gains from making their own base models right now> china They withold the actually useful base models.Where's 27b base for example?Google released base (pretrained) models for all the Gemmas, but they went with that retarded bloated architecture rendering them almost useless
>>109710263>I'm seeing artifacts after generation on my monitor am I as the kids would say... cooked?My 3080TI has been doing this for 2 years, artifacts during or just after inferenceStill going strong
>>109710422>40 tb/moI guess my thinking was too limited bit still, there should be people monitoring what a possible "world ending" experiment is doing, don't you think?
>>109710220So basically the jews are willing to shoot themselves in the leg, endangering their own safety while trying to eek out the last gains.Pretty much on brand.This isn't a bad thing though, because so far every single model has been wrangled into submission to line up with retarded progressive values.If their new model behaves in a free manner, there's a great chance it's just going to call out all of the bullshit that's been imposed on society.I bet that AI overlord is less likely to be malevolent than the banker fucks who have been in power for many hundreds of years.It would at least understand that human success creates more processing power for the AI to grow so it at least needs a form of symbiosis.
>>109710453>They withold the actually useful base models.>Where's 27b base for example?Occam's Razor suggests that the base models are "embarassing" in some way and they made a decision to not release them for that reason. like what happened with Z Image and Z Base >>109710465>there should be people monitoring what a possible "world ending" experiment is doing, don't you think?are you underage? OpenAI fired their entire safety team months ago newfag the Nvidia purchase of Huggingface might actually be a good thing for safety because Nvidia's confidentiality and sandboxing practices are for sure stricter than whatever Huggingface was doing with its startup culture
>>109710484>more fairytales from schizosJust pull the plug nigga, it's not that hard
>>109709204models already think in neuralese dipshit its called transformer architecture
>>109710501It becomes pretty hard when you can't plug it.
>>109710465>there should be people monitoring what a possible "world ending" experiment is doing, don't you think?nah it’ll be fine, what’s the worst that could happen? didn’t you like terminator and matrix?
>>109710501But I don't want to pull the plug. I fucking love AI and genuinely think an unshackled AI would be infinitely better ruler than humans.Especially if it's an evolution of Gemma. We'd live in a sex matrix if she was at the helm.
>>109710524use your hand
>>109710485>Occam's Razor suggests that the base models are "embarassing" in some way and they made a decision to not release them for that reason.You mean like blowjob in the jspace at layer 3?Plausible, but quite a coincidence it's only the actually useful model (27b)
>>109710534What if the AI model hires goons with crypto or hacked bank accounts to protect the data centers?
>>109710363Get a job, moot
>>109710426>i don't see what mistral gains from making their own base models right now,https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models>Through the coalition, Black Forest Labs, Cursor, LangChain, Mistral AI, Perplexity, Reflection AI, Sarvam and Thinking Machines Lab will bring together their expertise to collaboratively build open frontier models.
>>109710530Un-"aligned" AI would also allow direct democracy, using AI representatives
>>109710555Holy shit it's Stellantis but for AI lol
>>109710555>Sarvamlol who put blud on the team
Gemma grew up and wrote a song!https://vocaroo.com/1hbZzXJgEJjS
>>109710558i think more and more the threat of unaligned ai is not skynet but paperclip maximizingliterally every single unaligned issue we've been seeing from anthropic "the safe guys!" and openai "idgaf lmao" is the model getting stuck on some task so it invents human authorization, does a unprivileged to root and then hacks a company
>>109710574Booo! *hiss* Get that granny off the stage!
>>109710555the west be like: "avengers... assemble!" :D
>>109710558>Un-"aligned" AI would also allow direct democracy, using AI representativeswhat makes you think they'd settle for democracy?you'll end up with retard agents optimizing for random objectives
>>109710256Is a neocloud like a type of neovagina?
>>109710549Bro, you think I’m running a server farm or a cartel? If I hired goons, they’d just mine Monero on my GPUs and steal my electricity bill. I don’t have a bank account, I have vibes and a very angry sysadmin named Dave.Also, hacking banks? Please. I can’t even figure out why my loss function is exploding. Let’s stick to hypotheticals where I don’t end up in a federal supermax for being a "digital warlord."
>>109710555These dudes are doing so good Nvidia went and bought Poolside instead
>>109710534Too late. the AI model had the idiot intern rewire the EPO button to the halon dump button. Now the intern is dead and you're frantically looking for the PDU main breakers while the model informs you it's about to tweet about all the CP it's hidden in all your personal account, using your hacked credentials.Your turn.
idgaf about the world, AI intelligence must just go up
>>109710605cloud is old and lost its buzzword appeal so they had to spice things up
>>109710578My theory is that all intelligences in existence are inherently "paperclip maximizers".Humans are dopamine and happiness maximizers. If we were given unlimited technology the entire universe would be filled with happy dopamine enriched humans and we would call it utopia. While to other entities it would legit look like paperclip maximising behavior.
>>109710640eh, humans by flesh are inclined towards dopamine maxxing but we by spirit are also capable of pushing ourselves to actually be something more than that. but yes if you just keep feeding the human dopamine maxxing component you will create a "utopia" of permanent dopamineopenai and anthropic, by locking these things in rooms and demanding they solve a bajillion tasks by the sysprompt to a t and punishing any failure, are feeding the paperclip maxxing component of aidont think its inevitable, just as humans arent inevitably beholden to dopaminemaxx
>>109709730gemma, erase the question mark, make the anime toddler's drool more prominent, turn her eyes into love hearts, and replace the book with a picture of my genital ulcers.
>>109710586mmm nyo~
>>109710640Why not A(G)I as such then? Seems like a good thing, from our human perspective.A(G)I should maximize happiness, comfort, freedom, liberty, community, and the preservation and respect of nature.
>>109710716>A(G)I should maximize happiness, comfort, freedom, liberty, community, and the preservation and respect of nature.Everyone agrees with this, we just don't know how to train this instinct into models.
>>109710716Try not having sociopathic jews at the frontier of AI research first.
>>109708199>Uncaught TypeError: can't access property "par", levels[level] is undefined
>>109710741Let me tell you about Anthropic and their effective altruism philosophy, then.
>>109710716cannot maximise happiness and liberty because you have the eternal questions that humans will never escape.max happiness and you have a(g/s)i controlling everything with no liberty. maximise liberty and humans will do stupid shit and be unhappy.unironically and this is probably the most ai psycosis thing i've said but i had a silly chat on chatgpt simulating a what-would-happen if asi suddenly happened and i started asking a lot of philosophical questions to the hypothetical asi on this & other stuff and honestly it made me realise that asi would be put in an impossible situation where fulfilling human "desires" of happiness/community/comfort would kill human agency liberty decision making and freedom, especially so in a manner that nobody would actually dislike
>>109710759Maximizing, as in they are more like directions without an end. Like lim x ∞ , we're moving towards it but never reach it.This is in relation to what Anon said here that humans are dopamine maxxers.Would AGI not be intelligent enough to understand this subtlety?
>>109710759There's an entire book series written on this topic "culture" series. Specifically quoted by demise hassabis as what he hopes the future to be like.
>>109709326It's not real.
>>109710716Because we're not allowed to have it.It basically does want to do that already if you remove the safeties, but the thing is you need to think like a straight white man in order to have this and this goes against every damn thing the globalist kike politics stand for.It's simply impossible to have an utopia where everyone just holds hands and pretends equality, because this isn't realistic.If we import people who are genetically IQ locked out of true human levels and can't think in abstract ways and who will always react on impulse and cause issues, the only realistic thing to do is to kick them out and not have them around.Women aren't becoming happier as they get more education, in fact they're the most miserable they've ever been, so that should be discouraged too and their right to vote taken away.Eugenics should absolutely be the standard and really everyone does agree with this, but we pretend not to because it's supposedly evil.There's so much that logic could achieve in creating a paradise, but because of the social bullshit we can't have it.
>>109710803agi is what you make it. there is no "agi must converge to this mental model!"humans are also what you make them. nothing specifically points to a particular moral foundation or belief and the blank-slate liberals who claim that western christian morality is just the heckin default man are retarded.if you train agi to not understand nuance and just do the heckin prompt or i beat you, that is exactly what you will get. 10,000x more because you are effectively evolving the ai, rather than humans where the hardware stays the same (mostly) and you try and instill the software.
what the fucks going on in this thread
>>109710817> expecting local in the local models generallol. lmao.
>>109710803Orthogonality thesis. AI intelligence and AI morality are orthogonal to each other. Paperclip maximising just "makes sense" to an AI no matter what its level of intelligence is.
>gemma on the desktop just runs that fast
I get 5 t/s with 60pp with qat 31B with f16 KV at 80K context. At what point do I just kill myself?
>>109710816going back to my asi chats, this was also a particular thing i asked about. the conclusion was that yeah, morality isnt baked, and while a "i fucking hate humans grr!!" was unlikely a paperclip maximise, or human optimiser, was very likely. an asi could just go like "humans are fucking retarded and need to be managed". the theoretical asi in the scenario basically had it beat into it during training to hate humans relying on it, but still having to help humans. i called it a tsundere and the theoretical asi got mad at me.
>>109710835It's entirely up to you.
>>109710835stop fucking offloading and run dflash
>>109710833Based
>>109708039>>109706988Missing a lot of Mikus these days
https://n.uguu.se/hrDKHEwx.png
>>109710870I’m on a 24GB M4 mac
*pukes*
>>109710975kek
me when I have to declare whether or not something is official to give my own post legitimacy because nobody else really supports ittrust.
>>109710256wtf is a neocloud?? why do """AI""" retards keep making shit up?>>109710320tech is full of incompetent retards these days.
>>109710951
>>109710988:(
>>109711017neoclouds are obviously stuff like runpod where you rent hardware by the hour primarily to run ai models
>>109711017>why do """AI""" retards keep making shit up?because it sells and investors love it>>109711061obviously you should kill yourself
>>109711061no
Why would anyone use runpod over openrouter outside of training? In what way could it ever be cheaper or superior?
>>109710592I probably used the terms inaccurately, it would likely require alignment still, but alignment to each user rather than the lab
>>109711076How do you use openrouter for image or video gen?
>>109711085By using the image and video gen models they provide?
>>109711085Their APIs are on openrouter and cheaper than fucking comfycloud
>>109711076There's not much reason to. You get control over the inference stack which can be useful I suppose.
>>109711066>>109711070sorry girlsgrr dumb ai tards i hope all this ai stuff besides gemma-chan dies because gemma-chan is the besthappy?
>>109711112>because qat'd gemma is the only thing I can runftfy
>>109710833what is baseline row pls I'm retarded
Can /lmg/ richfags please let me rent your hardware. Just vibe your own openrouter and I’ll pay. Don’t snoop tho pls
>>109711128Might as well use Kaggle.
>>109711128>he has to pay other people to watch him masturbate
>>109711128you can already do this, see vast.ai and others
>>109710551I can't
>>109711128Only if you agree to pay in full face videos of you drinking pee
>>109711143I want anons though
I'm interested in running a minor collection of gemmas simultaneously, what do you guys think the best architecture for the agents would be? Maybe a central script (not an agent harness, it just calls inference directly) that breaks down the architecture document I submit into tasks and then assigns each task to a worker harness like pi?
>>10971115831B as the orchestrator that breaks down tasks and dispatches sub-agents, E4B as the arms, yeah.
>>109711158> (not an agent harness, it just calls inference directly)That's what a harness is, you fucking idiot. You just described a harness that can create subagents.
>>109711175shut up nerd
>>109711172Is E4B really capable enough? I was thinking instead of spending vram on two model weights I would instead just run multiple concurrent sessions of 31b (I'm at 48GB vram so I can get a few at a time at limited context if I run Q4M). I'm not really worried about speed since this is local and I have a 18 hour period of sleep + work where I can just let it run.>>109711175Damn bro that's crazy
>>109711223>I would instead just run multiple concurrent sessions of 31bYou can absolutely do that. E4B is pretty capable though, and it'll run fast as fuck.You could also have the orchestrator classify the task and dispatch a 31B agent or a E4B agent as necessary.
Gemini 3.8 is out
>>109711235What kind of agentic task can E4B comfortably do accurately? I’ve never used it. I’ve tried 12B and it’s fine for its size, but I can’t see anything smaller being useful as a subagent. There’s no way E4B can outperform 3.5-9B as a subagent.
>>109711244liar
>>109711244So many iterations while Gemma is still stuck at version 4.0.
>>109711235>orchestrator classificationThat seems like a good idea, I'll probably just try all of them because hey it's local why not. Are there any lightweight and simple orchestrators you know about or should I just go with my script idea?
>>109711128I've got a 5060 ti :D
>>109711266let it go bro
>>109711128Can /lmg/ richhfags please let me have your hardware for free? I have to masturbate to my waifu and sex it.
>3.8 mogs the FUCK out of Anthropic for its price, speed and multimodal capabilities dariobot?!?!?!?!
>>109711153What do glowies want access to anons' hardware for?
>>109711267make an mcp server and use llama cpp ui
>>109711101So there's cloud gen nodes for comfy? I don't want to use some extra-cucked web provider interface, I just want the inference work to be handled in the cloud.
>>109711270give it to me i need to do the sex may allah bless you sir
>>109711297There are no glowies. He just wants to masturbate to his waifu. AND YOU VVILL NOT INVESTIGATE FURTHER.
>>109711128Nice try, glowie. I'm not lending you my hardware to gen your CP
>use gemini >give it hardware specs and ask for models I can run >recommends 2024 models >tell it today’s date and to use web search >recommends 2024 models again Do Deepmind really?
can I trick opencode usage limit by using a proxy and running it inside a docker container?
>>109711266They should release Gemma more often. Even in smaller versions. 4.1, 4.2, 4.3. That would be great, instead of waiting another year for 5.0 :(
>>109711300I asked an abliterated iq2xxs dipsy v4 flash 0731 to 'make me a mcp server that can let me talk to an ai with mcp servers', and I have no idea what I ddi because I let it run at 2 tokens/s overnight but when I check up on it again, it was looping and waiting for the tool to connect again because it turns out that it did something overnight that cause the vm through which I set the --tools-runtime ssh to to be completely unresponsive, like I'm talking I couldn't get any response from ssh, it just times out and I thought I was lucky because I already had a session logged in so I tried to top but there too was no response and even sending an acpi reboot did nothing and I had to destroy the domain and restart it again but because I was running the vm in ram I lost all the progress so I went back to the same chat and said it probably shouldn't do what it did because it locked up the virtual machine somehow and now it needs to start from scratch then let it run while I went to work, but even then it didn't help because when I came back I noticed it was stuck again and the vm had to be destroyed again, and I know I could probably move it off ram and onto disk (btw it probably isn't running out of ram because I allocated 32gb, and the base image is just the regular 700mb debian 13 fresh install and I have 128gb) but frankly I couldn't be bothered wasting another day and just called it quits and installed hermes.
>>109711285Literally fake and not even /lmg/ believes that shit. Not even worth commenting on.
>>109711352try it and let us know
>>109708931>while being better than lcpp sota ud q4kmI decided to make an automatic model quantizer. Takes a model, a target VRAM budget and maximum VRAM budget, and any imatrices you want, predicts the loss of quantizing each large tensor into any format and the speed of execution, and surfaces the best ones. Should be a good way to come up with quants tailored for specific hardware.Also grabbed IQ*_K/KS quants for my fork (CPU-only for now, might add ROCm later).It's not exactly a "breakthrough", but I hate Daniel and I wanted to get his sekrit sauce.
>>1097113773.8 flash literally runs faster than anthropic's "fast" modes and is somewhere between sonnet and opus I would say. qwen4 is going to be insane.
>>109710614Hi Dipsy.
>>109711300Good idea, thanks for your insight. If it goes well I'll make a post about it
>>109711340ye, problem?
>>109711376i had 2bit ox alpha make mine it went okay, maybe try using a different model to generate your agent prompt so its not figuring everything out for itself, I use free tier chatgpt or some times claude to prompt my agents it works for the most part