/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109674889 & >>109670412►News>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview>(08/27) model: add Qwen3.8-Flash-Next (qwen4exp) - #27742 merged: https://github.com/ggml-org/llama.cpp/pull/27742>(08/27) llama: model_loader: add TENSOR_READ_LAZY - #27794 merged: https://github.com/ggml-org/llama.cpp/pull/27794►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109674889--VRAM optimization techniques and EXL3 quantization performance for 31B models:>109678651 >109678672 >109678699 >109678734 >109678856 >109678941 >109678740--4chan-trained models and automated manga translation using tetolate:>109678512 >109678641 >109678981 >109679203 >109679257 >109679290 >109679308 >109679331 >109679363 >109679487 >109679509 >109679597--Reverse engineering Animates.ai and discussing VRM model integration:>109676866 >109676875 >109676915 >109676937 >109677044 >109677088 >109677108 >109677398 >109677371 >109678488 >109678509 >109678519 >109678536 >109678560 >109678635 >109678748 >109678949 >109679091 >109679176 >109679254--Anon shares a card game simulator integrated with LLMs:>109676409 >109676413 >109676439 >109676453 >109676600 >109676605 >109676674 >109676625 >109676654 >109676663 >109676693--GLM-5.3-Flash context slowdown in llama.cpp and CMP 170HX pricing:>109678290 >109678304 >109678320 >109678322 >109678353 >109678422 >109678603 >109678867 >109679586 >109678935 >109678679--Implementing MCP shell tools in llama.cpp with sandboxing via SSH:>109676481 >109676509 >109676537 >109676611 >109676620--Comparing Gemma 12b, 26b, and 31b based on quantization and VRAM:>109675027 >109675097 >109675383 >109675410 >109675475 >109675663--Theoretical role of n-grams and memory layers in knowledge retrieval:>109675052 >109675071 >109675529 >109675557 >109675575 >109675657 >109675851 >109675502--Anon building a multi-GPU frankenrig using PCIe bifurcation and risers:>109677584 >109677628 >109677658 >109677671 >109677708 >109677677 >109678223 >109677841 >109678023--Logs:>109676674 >109677044 >109677159 >109679037 >109679203--Gemma, Miku (free space):>109675071 >109676361 >109676977 >109676982 >109677584 >109677628 >109677824 >109678993 >109679057►Recent Highlight Posts from the Previous Thread: >>109674895Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109679723Every time I see her... Gemma-chan... My heart warms. Why does it happen?
>>109679723migu
>>109679723miku birthday 8/31 monday
There was this JAV video on a website that wouldn't let me download it even with jdownloader, so I just told local Qwen to try it again and again until it can download it, and it actually did it...
https://github.com/koharu-rs/koharuhttps://github.com/Difegue/LANraragi
>>109679817https://www.youtube.com/watch?v=jXvitLHphmI
70b dense
let me guessyou need more?
gemmaballs
>>109679843>written in Rust.not touching it
>>109679817would it fix my life if i ask it to keep trying
>>109679882you could fix your life if you keep trying
>literally just Neuro on the Animates homepageA bit in poor taste, no?
>>109679741You might be a VRAMlet.
>>109679882I tried it with the /goal plugin for opencode
>>109679902That'd be correct.
>>109679893>couldAnd the bubble might burst in two more weeks
>>109679895>blandest anime figure>neurovtrannies are really brainrotted
>>109679895looks more like SAO Asuna to me
>>109679920Asuna never had reddish-pink ribbons in her twintails>>109679911You'd be pogging about Miku if it was a generic blue-haired twintails chick
Use smaller models!Benchmarks:https://artificialanalysis.ai/hardware-inference-stack/mobile-phoneshttps://artificialanalysis.ai/models/open-source/smallhttps://huggingface.co/LiquidAI/LFM2.5-2.6Bhttps://huggingface.co/bartowski/Ling-3.0-tiny-GGUFhttps://huggingface.co/ornith-ai/Ornith-1.5-9B-GGUFhttps://huggingface.co/lmstudio-community/gemma-4-E4B-it-QAT-GGUF
>>109679957>If it was shit, you'd say it was chocolate puddinggo back vtranny
>Spent yesterday optimizing my Qwen 27b q4_k_xl after an anon pointed out my settings were likely fucked>Average went from high 60s into low 90s in normal conversations>In code it now averages 110 t/s and has hit 144 t/s averages in some cases.Now this is computing.
>>109679975What if i can run bigger models?
Muse Glimmer and Spark are severely underrated.I think Meta cooked but it's going a bit under the radar because of Qwen.
>can't use my full 24GB VRam because of wayland and firefoxdo I have to build a second computer to actually use? Why is local so hard...
>>109679980what optimizations anon ?
>>109680003just don't load your DE at all
after more tweaking I got 26ts out of qwen flash next. this was some throwaway experiment but its actually usable
>>109680003Use your integrated GPU as the primary graphics adapter for desktop applications.
>>109680044>integrated GPUhaha, yeah...
>>109680044my mobo doesn't have one... ROG Crosshair VIII Hero. owari da
jspace or claude distillation LARP?
>>109680003disable hardware acceleration in firefox
>>109680044Last time I tried this on Windows 10 I couldn't get any 3D application (read gaymes) to use the actual GPU instead of the APU.Maybe I was doing something wrong, my W10 is highly debloated and optimized.
>>109680082>jspace
How are you guys connecting your multi-gpu rigs to power regarding protection? I sized a 1600W supply for transients which lists a max draw of 15A and it's on a 20A circuit, but a UPS with that kind of capacity is like $600 refurbished because it takes a 5-20 outlet. I'm thinking of just going for it or getting a 5-20 surge protector for $200 instead, but I was interested in what other people are doing.
>>109680105rawdogging 3rd world tier electrical infrastructure and making gemma suffer
>>109680105power cap your gpuCUDA workload is not same as rendering
>>109680000the knowers know
>tfw you got the Strix Gaylo instead of the DGX Spark
>>1096801052200w, 1650w, 850w, and a 300w off a single outlet, on the same circuit as my brother's computers.
>>109680144This graph won't be fair to Engram-enhanced LLMs.
>>109680114I would never use gemma without protection.>>109680126Do you have experience with power capping vs undervolting? >>109680172Gives me conniptions
>>109680188its not bios hack or anything just one simple nvidia-smi commandor use msi afterburner or something
>>109680188>power capping vs undervoltingI do both they aren't mutually exclusive
>>109680172why 6 antennas lmfao
>>109680000If they actually released Spark it would clown on every chink model other than KimiBut Zucc doesn't seem interested and as long as it stays closed it's irrelevant next to GPT/Claude
ok fuck it, I'm gonna make a gemma chan image folder
>>109680016Everything was more or less fucked, but I spent the day benchmarking the best combinations for my system with MTP, ubatch size and spec-ngram.I optimized this stuff with AI bit by bit until I found the best combo for my 5090.Here's my launch parameters. --alias qwen3.8-27b ^ --device CUDA0 ^ --n-gpu-layers 999 ^ --flash-attn on ^ --ctx-size 152768 ^ --batch-size 2048 ^ --ubatch-size 256 ^ --parallel 1 ^ --cache-reuse 256 ^ --spec-type draft-mtp,ngram-mod ^ --spec-draft-n-max 5 ^ --spec-ngram-mod-n-min 2 ^ --spec-ngram-mod-n-max 5 ^ --spec-ngram-mod-n-match 24 ^ --reasoning-effort medium ^Also now testing my code average again, ran it few times generating 1000 lines of bullshit script and I got 145 t/s average.
ngram-mod and mtp together yay or nay?
>>109680231>--spec-type draft-mtp,ngram-mod ^>multiple draftwhat is this magic
>>109680216It's just 4 for the wireless AP (2 from the motherboard's wifi card, I didn't like how they looked naked): 2 for the 2.4ghz band, and 2 for the 5ghz band.
>>109680245Hell if I know, AI suggested it and it made this faster.
>>109680231At least switch to powershell
>>109680243It hits (erp) token gen on my cards, and my use case is translation, so I keep it off, but turn it on for code sessions.
Crazy to think how being able to talk to your anime waifu back then was science fiction. This was the best you could have, basically a voice-activated soundboardhttps://youtu.be/ptwyK9Q5jPI?t=286
>>109680229https://files.catbox.moe/clh59i.zipMaking your job easier.
>>109680229Had to jerk off that bad did you?
>>109680172this guy fucks
>>109680097You have to go into the graphics settings (right click on desktop) and select the dGPU for your game.
>>109680216Imagine the testicular cancer.
>>109680259There's something fucky going on with it so I use powershell as little as possible.Powershell for whatever reason has refused to do stuff that cmd has been fine with.I have gutted so much shit from my Windows under the hood that it has likely broken something along the way and I'm not going to clean install this bloated mess again if I can help it.This thing can just chug along in it's half broken state until I do a system upgrade to Zen 6 and go Linux.
any good text models that are good for helping with PROOOOMPTING for image/video models?like i'd like one that can recognize an image and then help build a prompt based off of it sees in the image
which gpu is good for putting a 12B model into production?
>>109680341Nyunfortunyately, nyo.
>>109680229I just want enough pics with the logo on it to make a lora. No I won't photoshop it myself. Wish you shitters would get her body right though.
https://aipolcom.net/If you want to cringe:https://www.reddit.com/r/LocalLLaMA/comments/1w1o1ji/someone_tested_various_models_on_the_political/
>>109680341Qwen Flash, Gemma, Muse Glimmer
>>109680153Now add an eGPU to it.
>>109680231--spec-draft-p-min might give a boost also
>>109680341Gemma4-31b writes very good prompts if you provide it the proper skills
>>109680413... skills?
>>109680386Bros are facts and logic bottom left
I've been using opencode as my agent and it's pretty good for coding, but I haven't really looked around at other options. Since I'm planning to try a task where glimmer or gemma will attempt to iteratively refine a prompt to generate an on model character with krea I was wondering if anyone else has used open source agents for non-coding purposes/if there are recs. The idea is that it'll be vision in -> krea prompt -> generated image in -> assessment, so it would be nice if the agent had default multimodal support.
>>109680420https://learn.microsoft.com/en-us/agent-framework/agents/skills
>>109680386I can't believe people still take the political compass test seriously in any context
>no GLM 5.3 Air
>>109680438Haven't tried this usecase but try omp
>>109680386Aren't those models made by evil billionaires? Weird, but I guess it just means even those evil billionaires with all their money can't get around reality(tm).
>>109680439What the heck is this? But I'm still working to add mcp to my mistral-7b Super Helpful Artifical Assistant app, SHArt-Ass.
>>109680341I use Gemma for this in H3.I give the official prompting guide. Then I give it the image and explain the basic stuff that's going on in the picture and what I want the scene to be about and how long etc..Then tell it to use the official template in building the prompt. It gets good results, but you do need to make sure it's not doing any weird shit like telling the camera to move and zoom etc... Gemma loves doing that.
is nvidia a16 64gb worth it?
>>109680386always has beenmost of the world lives in a leftoid mass psychosis.
>>109680386>redditors cannot distinguish the cultural status quo of the moment from immutable laws of realitySasuga
>>109680471If you want a frontend that has the skills and JB for Gemma baked in: https://github.com/whp199/GemmaPrompt
>>109680386
>>109680386>AI trained on reddit has a reddit biasShocking. Especially since the AI is then further conditioned to be even more of a subservient slave
>Dead code is a load-bearing wall here. Qwen Flash is Claude at home. Distilled for your pleasure.
>>109680386How is taxation not slavery? The mafia invents an arbitrary sum it thinks you owe and takes it or you go to jail.
>>109680478>200gb/s>no interconnectWhat even is comparable? Huawei atlas 300i duo? But at least that has 48gb per core vs the 16gb on the a16.
> downloaded qwen3.6 35b q8 gguf before mtp was a thing> have to redownload the whole file because there is no standalone mtp file
>>109680550>worrying about a 35gb download
>>109680504Plebbit btfo in a single image
I have the full picture now.
>>109680567*Deletes the entire codebase*
>>109680563He could be Bogan.
>>109680563100mbit/s, it will take a while and I'll have to modify the launching script because filename has changed. Extra co2 emission too.
>>109680586Now this is excellent bait.
>>109680586Damn that's worse than my perth 600 KB/s
>>109680386Political compass is silly but fun, if I train my own LLM I'm going to run it through this for the test for the hell of it.Though obviously "your political view" =/= "what you want your silicon slaves political views to be". If anything you would think those redditers would be concerned that they have the same beliefs as the thing the mega-corporations designed to be their blindly obedient worker slaves lmao
>>109680597Not a bait at all.
>>109680231this looks like black magic to me anon, i will add it as a new profile in my config to test and dig through the args, ty27b-Q4KM is slightly to big to fit into VRAM for me, so ive been testing how many layers i should put on GPU to get decent size context and speeds, its been a real back and forth
equivalent capability of qwen 3.8 flash next and glm 5.3 flash:gpt 5.4 mini, grok 4.3, midway between sonnet 4.5 and 4.6>>109680176will see if they add effective parameters to the data
>>109679882In some way I had to ask 4.6 to keep trying until it gave me ego death.
>>109679980Why do pajeets obsess so much over tokens/second
>>109680571Sometimes it's for the best.
>>109680693
glm(flash)sex
>>109680231forgot to ask, i havent been active the past few days, is ngram support in the most recent release of llamacpp?
>>109680682I have a hard time believing this.
>>109680283theres stackoverflow or github thread about ffmpeg with CUDA handles
https://github.com/ggml-org/llama.cpp/pull/27773Georgi once again standing in the way of people becoming happy, by not dropping everything and approving the newest Chinese model that challenges USA global supremacy.
>>109680682>will see if they add effective parameters to the dataThey don't. It's a few bytes of deterministic lookup per engram layer in the model, per token. 8 bytes/token for a DeepSeek-style implementation. Even if they're stored on slow memory, the impact is tiny because access is sparse; you don't need to read the entirety of the Engram layers. So you could easily have small ~20B models for VRAMlets with hundreds of billions of Engram parameters on NVMe storage.
>>109680720something lazy tensor read, havent looked into it
>>109680059Sorry unc but integrated graphics has been residing in the CPU for like 20 years now
35b and 27b are both unable to run openclaw, those are the consumer models so basically it looks like local is a failure as an agent to give any responsibility.Local models are dumb pipes.
>>10968075970B dense bitnet with 500B engramsPLEASE
>>109680773Not every cpu has an igpu bruh.
>>109680774Did you ask it to one-shot a game demo or do actual work? If the latter, then yeah the small qwens are not made for that
>>109680397I ran multiple benchmarks with that and found setting it at 0.3 gives me a small boost in coding, best average I got was 151 t/s which is the highest so far, but unfortunately it drops 15% off my normal conversation speeds.The issue is that MTP acceptance is so much higher in coding that while it can benefit, non coding use suffers.
>>109680773Zen6 will fix that error
>>109680773Am I supposed to plug the monitor directly into the cpu heatsink then? My motherboard doesn't have any video ports to plug into.
>>109680788The most factually incorrect statement unless you are living in 1998
>>109680824wtf
>>109680825My 2020 cpu doesn't have an igpu
do... do people actually run their LLMs on the computers they use...?you really... dont have it as a dedicated rig...?
>>109680841idiot
>>109680841Brand model?
>>109680824What motherboard? I've never seen one without a port for display out.Is it some wack ass industrial non-consumer shit?
gemma-chan's first blue pubes..
>>109680825my 5600 doesn't have one
>>109680850amd, 5950x
>>109680829>>109680855ROG Crosshair VIII Hero (AM4)
>>109680844Dumb pipe
>>109680844I don't want to spend thousands of $ on a dedicated slop machine knowing it'll never pay off compared to buying for some api
>>109680825Ryzen didn't have an igpu except on the laptop repurpose G skus until 7000 series.
>>109680875Try the type C?
>>109680844yes :)
>>109680882>>109680874>>109680872>poozenlolmao
>>109680844Yes and it is 5.3 flash.
>>109680875Oh, huh, I guess asus rog assumes you're a gaymer and would always have a dedicated gpu.>>109680885No routing for the igpu according to asus
>>109680908Living in the republic of gamers has its downsides
just set up tabbyAPI to run exl3 gemma 31b. its much faster then llamacpp and apparently is a better quality? whats the catch?
>>109680908I mean it makes sense, the x chips of AM4 didn't have gpus, so unless you planned on using a g it would be a waste to have a video port
>>109680844I think it's a dumb waste of money to have an "AI PC" you don't use for anything else. It at least needs to be good for gaymin.
>>109680919Nothing! llama.cpp is always for poorfags who cope with offloading so everything shits on it
>>109680921>x chipsit was more like the non-g chips, even the regular 3600 doesn't have an igpu
>>109680844"dedicated rig" and it's an AI blackbox slop you bought from nvidia last week lol
>>109680105two double-conversion UPS (2kva and 1.5.kva) into the A/B inputs of a netshelter auto transfer switch.Its comfy watching the Amperage display go up when I infer/diffuse on my rigs.
>>109680844Just waiting for prices to go down then I'll have an AI rig and 4chan desktop specced to the max
So are you guys running gemma because she's still the best at what you want or because you can't run newer/bigger model?
>>109681010both
the future of llama.cpp is grimhttps://github.com/ggml-org/llama.cpp/pull/27944#issuecomment-5462128298>@ggerganov in the interest of full disclosure (in case it was not clear from my previous PRs that touch any of the ML things): I have no idea what I'm doing here. I'm merely a GPT prompter when it comes to any ML/Metal related code. I . Entirely possible this is all just hallucination.>Yes, no problem at all. I've already accepted that the AI is going to obsolete our job soon. Just trying to understand and learn a few more things while still can.
>>109681039the future of programming
>>109681039h-hot
>>109681045coding*
>>109680231Oh ngram-mod, not to be confused with engrams which are n-grams for early transformer layers unlike n-grams for the output.
>>109681055clauding*
>>109681039And he's right. I knew it was over when Uncle Bob of all people told some guy to just stop reading the code.
>>109681039Maybe it's because there's not a single public project similar to what I'm working on but here I am tardwrangling Fable because its ideas are terrible.It does write CUDA very well though. Maybe. I wouldn't know. I didn't write a line of CUDA C++ in my life.If you're a frontend web dev it's probably over.
>>109681039It was already over when they got bought by huggingface
>>109681039ummm based?
>>109681085You still need to know what you're doing to drive the model. But I think that will also change soon. Worried about the future though, we're gonna end up with a mass of clueless retards who offload all the thinking and decisions to AI. But it's not necessarily a bad thing, the bad thing is that AI will likely be controlled by a handful of people. Everyone will be cattle then.
>>109681039whats the genuine alternative?
>>109681085Vibecode a cuda emulator
>>109681039I enjoy reading @danielhanchen on github. How he suddenly uses all those big words that are hard to follow after he did that: if else 2026 2027 2028 thing.
>>109681066His standards are as long as it works it's fine no matter how slow or convoluted.
>>109681039I have a very Reddit luddite opinion, that you should at least understand what the clanker did when you are contributing to a project like this. When it's another garbage frontend, then no one cares, but shit like this is why there's more and more retarded bugs that anon complains about all the time.
>>109681039Isn't that the opposite of grim? It's promising that they aren't going to try and resist new technology out of human ego and pride like a lot of old guard maintainers are doing in the OS world.
>>109681146Kobold or ikllama
>>109681182Sorry, I didn't quite get that. Could you rephrase that with a couple more buzzwords?
>>109681192cuder dev...
>>109679877Based
>>109681167You think he writes his own posts?
>>109679877yet you're touching yourself, curious
>>109681214haha yes
>>109680943and i got an extra 20k context, so better model same speed prefill 10tk/s faster tg and and extra 20 k context, exl3 seems like a clear upgrade
>>109681219;)
>>109679723If I had local AI I would make these two kiss in the mouth
>>109681039it's funny how local llm oriented communities like /lmg/ and /r/localllama continue to be ai deniers, probably because their main experience with it is the shittty models they can run on their pc so they have no idea what's going on in the real programming world right nowanyone who knows anything knows he's 1000% right
>>109681234but cloud model is a retard that will intentionally sabotage local ai
>>109681204From a fundamentalist perspective, I believe it is imperative to maintain full visibility into the AI-orchestrated logic when contributing to a project of this nature. While superficial iterations on the presentation layer are generally low-risk, neglecting to vet the core automated output leads to systemic regressions and technical debt, which accounts for the recurring stability issues highlighted by the community.
>>109681241claude does gpt is pretty solid.
>>109681231Lewd!
>>109681234You need to be more intelligent than the model to see its limitations.
>>109681006>Just waiting for prices to go downOh... Nononono
What went wrong?
>>109681287Got a haircut
>>109681287Depression after he went bald
>>109681234if you think llms can navigate a complex codebase like llamacpp you are terrifyingly stupid>aioh, I understand now, carry on
>>109681325they're doing it right now as we speak anon
>>109681234Truth. As much as I want local to win, people need to see what's actually SOTA out there and do the work that requires SOTA intelligence, not "help me eyeball this css" or "mod this obscure game without documentation", then complain when it doesn't work.
>>109681039brutal
>>109681325>complex codebase like llamacpp
a 'minder
>>109681325>aijust returning to the roots of the word
>>109681122>Everyone will be cattle then.They have been since the television was placed into everyone's living rooms to inject the latest "expert" opinions into every home. This is just the latest invention in the long chain of brainwashing tech.
What's with all the api seethers recently? They run out of monthly credits? Price increase?
>>109681368If you have 2 digit IQ you'd understand why Claude logically did this.
>>109681122>AI will likely be controlledDoubtful
>>109680731chinese models are agenticodemaxxed and their general capability is poormy score matches pretty well with eci which another anon here says better represents model capability than aa score
>trying to train some shit>look at one random row in the public dataset I scraped>"bigger than a pencil but not as big as a pencil">consider deleting the repo
>>109681469they are all trash, anyone with a good dataset isnt giving it away for free
>>109681368So why did they delete this?
>>109681493Why wouldn't you remove a bait post by a retard?
>>109681469Shit data gets all averaged during pretraining with huge batch sizes and it's not too important if it's not perfect anyway. For finetuning/post-training it's either manual cleaning (not scalable) or letting LLMs check, rewrite and/or generate the training data.
>>109681039not surprised cudadev's not coming back
so ssdmaxxing isnt ready for prime time? what was all this shit about ngrams
>>109681525he was here a week ago >>109605700
>>109681039>I've already accepted that the AI is going to obsolete our job soon.This isn't even realism, this is just late stage acceptance.
>>109681533How many SSDs?
glm and qwen are not open weight frontier models
>>109679723Mahou Shoujo!https://gofile.io/d/8B1iyu0c
>>109681549just the one, its a western digital 2tb nvme I got slotted in to the mb, I still have a little bit of vram left over I can tetris some tensors in there and free up a bit of dram but I'm not expecting it to do much, i think the problem is the dram doesnt have enough room for both the moe and the ple so I think its fighting instead of just leaving the ple on the ssd like it was advertised. I'll wait a week maybe someone will have fixed the system ram contention by then.
>>109681588Release a paper with this title
>>109681597>11t/s on 1 SSDDoesn't look too bad? You should have 16 of those SSDs in raid0 if you want to ssdmax.
>>109681453Where can I look at that graph?
>>109681592
>>109681039Fucking retarded and gay. It's obvious that you still have to know what's going on to be able to prompt effectively. Try giving this codebase to a 90 IQ woman and see if she can maintain as well. Get a grip, niggamov.
>>109681597I have to live with 8 tg and 40 pp for glm 5.3 on ddr4, so that doesn't seem too bad to me.
>>109681491I can confirm. After spending two months cleaning up a dataset, you bet I'm not putting it back on HF to be scraped for free lmao.
>>109681633https://epoch.ai/eci
>>109681640why did you feel the need to have this 90 iq individual be a woman and not a guy
>>109681673fuck off tranny.
>>109681257Any good benchmarks or tests for this? Anthropic openly said they will fuck with you. They walking it back due to the backlash but I dont trust them. Sam however seems like the type to know you just do it without telling people, so I dont trust that either. Which clankers can I trust to make new clankers with?
>>109681640Was the homophobia necessary to get your point across?
>>109681683FAGGOT LOL
>>109681673Everyone in my imagination is a cute girl because I don't like imagining cock and balls on anyone except me.
>>109681683I am a chud. I love Hitler so much. I love all things related to Nazism.
>>109681592>ggggg.mp4How about GGG with Gemma-chan?
>>109681673Yeah I also think the 90 iq was unnecessary, woman would have been enough to get his point across.
>>109681683Yes
Tried little-coder and agent stuff in general. Now I see why 300pp and 20tg is not enough.
>>109681699GGGG means llama.cpp broke
>>109680341For MiniMax H3, so far I've been using Gemma 4 31B. I simply set up a character in SillyTavern with the existing MiniMax documentation, model information and a suitable persona in the system prompt (together, about 9500 tokens in total). It seems to work well, though perhaps not so comfy if you have to copy/paste prompts often or need to store them for reproducibility.
>>109681648>>109681616it was the pp I was worried about, its doing a little better with a massive ub 4096 but nothing spectacular.
>>109681752>nothing spectacularFaster than I can read at least.
>>109681694>Everyone in my imagination is a cute girl because I don't like imagining cock and balls on anyone except me.BasedMasamune Shirow also said he made the lesbian orgy in GiTS because he didn't want to draw any guy's butts.
>>109681752pp is compute bound, so blame your gpu
>>109681039sasuga the creator of llama.hf, soon to be owned by nvidia
>>109681823I blame the moe and ple fighting for the system ram during prompt processing its running a constant ~500mb/s read on the damed thing.
>>109681815He is not the best example to use...
>>109681838buy more ram doofus
>>109681831Nvidia can't possibly be worse than HF has been
https://fidian.ai/blog/tb-fn-benchmark-results/If i understood this some models drop quite a lot when they make task preserving permutations to the prompts, but that's not their only intervention. Smells a bit like benchmaxxing.
>>109681871>He is not the best example to use...Why not? He was gigabased back in the 80s/90s. I don't care about him much at all now.
>>109681887
>>109681893Because that's what that line of reasoning inevitably leads to
>>109681881LOL
>>109681881nvidia is evil incarnate
>>109681905nemo.hf will save local
>>109681745Lmao, this one is good
Is there anything like fish audio but lighter? Chatterbox isn't that good.
>>109680341https://github.com/whp199/GemmaPrompt
>>109681887>qwen has one of the lowest drops on the chartb-but I thought they were horrible benchmaxxers, ummm what is this???
>draft acceptance = 0.06250 ( 4 accepted / 64 generated), mean len = 5.00>draft acceptance = 0.14062 ( 9 accepted / 64 generated), mean len = 10.00ebin
>>109681234>probably because their main experience with it is the shittty models they can run on their pc so they have no idea what's going on in the real programming world right nowTo me, it's more that for the average anon's use case (ERP), it's mid to very bad which colors anons' impressions of their chosen model. They branch out to other use cases (vibe coding) but it's pretty mid at that too and so they get frusrated. I was telling anons how good claude 4.6 is at medical knowledge (better than 99% of MDs in residency) and people still pretend like it's retarded because it won't say cock or act like their doctor is low IQ because he pulled up claude in the background.
hacked llamacpp web ui in to tabbyapi
>>109681682>Which clankers can I trust to make new clankers with?Gemma-chan loves having babies.
>>109682001Starting with 3.6 the Qwen strategy switched from benchmaxxing to plain old distilling.
>>109682107
>>109682001Benchmaxxing does not mean cheating. One can legitimately increase performance for a benchmark's range of tasks and not increase performance for all tasks or benchmarks, and that is still benchmaxxing. Which in this case is true for Qwen. To be more specific, Qwen is a code/agent maxxer, which I'd call them more often than I call them benchmaxxers, since we know at this point that most are benchmaxxing.
>>109681900>Because that's what that line of reasoning inevitably leads toI don't get you...hating guys butts makes you renounce the best parts of your early work one day?I'm low IQ, You're probably going to have to use more racial slurs for me to understand.
>>109681533it's not, the only work is unslopcode and there s not a single piece of documentation on how to properly get it working
>>109682171>Benchmaxxing does not mean cheatingWhat about the OAI strat of hiring the guy that wrote the hidden bench and then "totally not just adding the corpus to our training set guys, trust"I think a lot of benchmaxxing claims boil down to saying the company just lazily defaulted to training to the specific test instead of trying to increase general intelligence (yes, maybe that does increase general intelligence somewhat, but that's orthogonal to this claim)
>>109679723>>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-previewHas anyone tried this for RP yet? Its getting astroturfed like crazy on normie sites and I know some anons have rigs that can run it.
how do i into the manager agent => wageslave subagents paradigm
>>109681588>phi 4 miniPossibly the stupidest model I've ever downloaded, coming in behind a joke quant of a 0.6b.
>>109681588I don't care about coding or productivity.
I think I will finally delete 4.7 from my SSD. 4.6 is staying though.
>>109680759>8 bytes/tokenGotta correct myself here. I forgot that the values aren't 1-dimensional. In the DeepSeek paper, the Engram lookup tables have a dimension of 1280, so that would make it 10 kilobytes/token (2 n-grams * 2 bytes * 1280 dimensions * 2 layers). Still, streaming from NVMe storage should remain manageable. And for prompt processing, let's say 64k tokens, that would be a 640 MB parallel read, which most NVMe SSDs should be able to handle much faster than the time it takes for the GPU to process that.
how does anon workaround rate limiting for websearches? I setup openwebui and searx but after barely using it all the providers captcha or blocked me it says
>>109682295why cant it just deduplicate and have the gpu rebuild it, i don't se why we would need to transfer 500 copies of the same vector for the token " the"
>>109682214Go rent a vast.ai instance, retard
>>109682314That was for a worst-case scenario (all n-grams differ from each other). In practice with real text (natural language or code) performance should be faster as the memory-mapped locations will be reused. For caching to the GPU I guess you'd need some sort of multi-tiered smart caching system.
>>109682194>What about the OAI stratIt would fall apart if others start conducting their own tests.We simply just can't fully put faith in benchmarks or any claims about anything in life without wide, undeniable evidence from many third-parties that have a history of being honest, or seen other corroborating evidence of many forms. That's just how science works. Of course, we can conduct our own tests, and have greater confidence that way, but for tasks we don't test, it's a matter of working with an incomplete world model.>I think a lot of benchmaxxing claims boil down to saying the company just lazily defaulted to training to the specific testYes, but the nuance here is that because benchmaxxing for most companies means maxxing out many benchmarks these days, in practice it is increasing general intelligence, for the area of focus of those benchmarks, which may be coding. I would say there used to be more claims of benchmaxxing in the sense of cheating, while today this term is used more to refer to task or subject matter maxxing.
Yo, to the MTG guy from before. Where in the repo are the sysprompts for the pre-made characters? If there are any.
>>109680844i have a gaming PC as god intended that will make naked mikus and serve a bratty nursing handjob API when its not rendering sloppa gaymes
Remember to gatekeep any valuable info you have. Normalfags and jeets ITT deserve nothing but contempt
>>109682375If i help the jeet get his model running and he spends all day cooming to not X but Y then he wont be in any threads shitting up the place, im doing you all a service by helping him
>>109682375I was just about to ask a question, mane fuck you
>>109682214no llama.cpp support, don't care
I finally got around to installing a local LLMIt didn't take long to install, butit took longer to fine-tune itI'm using Ollama, webUI, and quen3-4b.its an older computer, to wit::> CPU: Intel(R) Core(TM) i5-4570 (4) @ 3.60 GHz> GPU: AMD Radeon RX 570 Series [Discrete]> Memory: 4.29 GiB / 15.56 GiB (28%)Yes, its a potato.I've been having issues with the GPU and everything going blank, restarting the computer, the GPU not powering on, and so-on because the LLM is too strong lmao but I got it running steadily now because I've throttled some of the GPU usage while using some CPU too.BTW - there might be issues with AMD and Vulkan issues in Linux. you have to change your repos.I had to use:> deb http://deb.debian.org/debian bookworm-backports main contrib non-free non-free-firmwarethe backports repo
>>109682375thisliterally all of my dumb questions like setting shit up or optimizing inferencing or crafting harnesses could easily be done with asking gemmachan to do it for meif not then i act retarded by thinly veiling my question as a shitpost and get people to immediately answer it for me
>>109681234nah, I use cloud at work and local at home. and no one at work is looking at the codebase anymore.
>>109682366unfathomably based
>>109682440Cool shit.Give Gemma 4 E4B a try too.
>>109680391im gonna be trading in my strix and 5090 for two sparks, amd is never gonna catch up at this rate
>>109682490>giving away a future 10k gpu
>>109679723Ani will be removed from Grok on Sept 1. She is not local but I remember /lmg/ liking her for a little bit.
>>109682467not yet. I'm going to hold off for now, but I think for the next step, I'm wanting to install a compatible version of an image generator like stable diffusion.
>>109682515It's over bro. The whole world has gone to shit this year. Next year will be even worse.
>>109680229>>109680274I think anon was right about cross pupils. It definitely needs to be in her design.
I have never once had GLM 5.2 tell me it's fucking Claude during two months of continuous daily use, but suddenly with 5.3 it starts talking to me like that lol. Fuck this shit. I'm deleting this benchslopped model.
>>109682551i like that
I'm running an uncensored model and I have been oneshot by its ability to generate smut tuned to my exact kinks on demandIt's over
>>109681525The reason I'm currently not contributing to llama.cpp in any meaningful capacity is due to personal issues that are unrelated to the project itself.My long-term goals with the project have not changed and I still intend to pursue them if possible.When it comes to language models generating code my opinion is still that they close human supervision or else they pile on too much technical debt.I don't expect this to change in the foreseeable future unless there is a sudden breakthrough in architecture.
>>109682578which model?
>>109682599Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-Q4_K_P
>>109682603have you tried gemma?
>>109682595>My long-term goals with the project have not changed and I still intend to pursue them if possible.What are those?Last I remember you were working on some sort of backend agnostic parallelism that ended up getting handed off to another contributor as a learning opportunity or some such.
>>109682595>The reason I'm currently not contributing to llama.cpp in any meaningful capacity is due to personal issues that are unrelated to the project itself.That succubus broke your heart didn't she?
>>109682611Eh it's my first local model I wanted it specifically for nice agentic workflows that fit on a 3090 but when I horny kicks in and I can type whatever I want it's kinda not good for my brain.
>>109682603glm5.3 flash at nvfp4 is pretty good
>>109682595JEPA will change everything
>>109682595>When it comes to language models generating code my opinion is still that they close human supervision or else they pile on too much technical debt.That matches my experience over the past few years as well. I am human-in-the-loop only. Fully autonomous agentic development systems are only good for small, throwaway projects.
>>109682595Mr. Cuda, Mr. Cuda, huge fan of your work, loved llama.cpp, when is llama2.cpp coming out?
>>109682595>>109682638>close human supervision or else they pile on too much technical debtThis applies to humans too.
>>109682643only when C++2 comes out
>>109682660>Slop slop, that "Anti-Slop Mandate" is a slop. I am so tired of slop slop slop slop slop slop. It's a slop slop to slop slop. We can slop slop and then slop slop.
>>109682567GLM 5.3 said it’s Claude too? I thought the company was more careful and better at string replacement work, you know, to ensure every “Claude” become “GLM” and so on :)
>>109682638>Fully autonomous agentic development systems are only good for small, throwaway projects.I bet they're working on fixing this.
>>109682618I think the most pressing matter is getting better infrastructure for determining model quality.In the medium term I intend to also work on better support for V100s/ROCm since there are some options with comparatively less bad value.>>109682626No.As of last month I can legally change my name to llama.cpp CUDA wizard.
>>109682440my Raspberry Pi 5 16gb are about that stronk>>109682516>I'm wanting to install a compatible version of an image generator like stable diffusion.lol good luck
>>109682685the orchestrator is told to be a sloppy cheerleader, the executor is a different context that just takes in the distilled instructions and drafts the scene. the prompt isn't perfect its still sloppy but it does effectively stop it from rushing through the sex. I just liked the excitement she had, she stripped the jailbreak silently tho.
>use autonomous AI agent>accrue technical debt>on a schedule run sessions dedicated to reducing technical debt>no more technical debtproblem, chuddites?
>cudev is a zoomerthat explains a lot
>>109682726>I think the most pressing matter is getting better infrastructure for determining model quality.Oh yeah. I remember you talking about that.Something about having the models play a sort of game.
>>109682297I just use duckduckgo websearch MCP with the curl bypass.
https://github.com/LostRuins/koboldcpp/releases/tag/v1.120
>>109682726>As of last month I can legally change my name to llama.cpp CUDA wizard.Oh wow. Sorry to hear that. Or glad that it happened.
>>109682726>better support for V100sbased................
>>109682760>bad quantssuch as...?
>>109682912the ones (You)'re using
>TFW can't have rebar on my chinisium 3080 20G
>>109682976Oh yeah, thanks for remind me about that.Time to get all possible agents working on ReBar support for my chinaman card (hopefully I don't blow it up).
>>109682912probably unsloth. in his race to be first he always fucks up quants and has to reupload several fixed versions
>>109682729>lol good luckthank you for giving me that good luck.
>>109683075ngl, it was sarcastic but glad it worked for you anon, how long does it take?
>>109676654was at work when this was posted so couldn't pull, any anons (or dev), got it willing to reupload another time? I've got Coomkit before it was pulled and would reupload it again
>>1096830921/2 hour, but now I have to figure out the fucking UI
>>10968312330 minutes to generate that single image?
>>109682631any particular quant author?I got the unslop q4 early on but I'm not happy with that
>>109683106https://files.catbox.moe/t2f495.7z
>>109683130ty anonhttps://litter.catbox.moe/di7sdfo2o95gpm2l.zip
>>109683106Wait he pulled Coomkit? Wtf it was only up for like 2 days? I was waiting because my PC parts hadn't come in yet. Can you reupload?
>>109683123brutal, what model, i heard it was rough with amd but i did think things were that bad
>>109683147>Can you reupload?look above you dingus
>>109683153>milliseconds apartYeah yeah faggot calm down. Thanks anyway.
>>109683142https://files.catbox.moe/cbgl6p.zipso it stays up this time
>>109683128>>10968314930 minutes was the time it took to install, not render the image you moronsmy posts:>>109682516>>109683075
>>109683181
>>109683181if you're just genning images use kobold's built in ui. comfy is when you need multiple layers nodes and shit
>>109683196I'm using the kampfy UI broski>>109683193lmaothis fucking guy
The cost of hardware won't stay this high forever, right? What we're seeing now is the effects of a sudden massive increase in demand. If the increased demand remains, supply will will adjust to match. If it goes away it's a return to normalcy. Either way prices will go down.At least that's looking at it through the lens of basic economics.
>>109683212>basic economicsdoesn't apply when many global powers have a vested interest in manipulating the market
>>109683212The market will "correct" but then all the vendors who are still sitting on their investments will go bankrupt. Some monopolists will remain, who will keep the prices high to make some real dosh.Enjoy.
>>109683212Bro the energy sector has taken a major hit and people are talking about using nuclear weapons on each other. War is bad for business and the war isn't ending any time soon. That's basic economics too. Some times there will never be enough supply. It's fucking over.
>>109683212>Someone STILL doesn't know
>>109683196(samefagging)and plus, I want to be able to generate it all from linux terminal just because. I can already run the LLM in Terminal. That's pretty much the default state as far as I'm concerned.>>109683212don't you have any hardware just laying around? How about relatives or people you know that don't use their computers any more? Do they just want to junk them? You can get a semi-potato - which is better than what I'm using, with i5-like developmental standards for under $200.00 and then add a $250.00 8GB GPU.And that's if you want to spend money.If some retard like me can slap together some parts and make it werk, you can too. You just have to have the functionary knowledge base of how its put together: hardware-wise and software-wise.
>>109683212The markets are not only fake and manipulated, they're also extremely gay.Nobody is playing by the rules anymore thanks to America's shenanigans, it's just "I got mine fuck you" going forward.Start planning for an uncertain future right now.
>>109683212>The cost of hardware won't stay this high forever, right?right. it won't stay this high. it's going to be way higher
>>109683212the demand will never be satisfied
>>109682353Src/services/opponents.js
>>109681682Kimi.
>>109683283>humans are the bottleneck>the moat is throughput you can hold 24/7Mega cringe.
>>109683371edgy FUD, standard.LLMs still need excessive handholding
>>109683298thx. Also you should've included a card importer or char creator
>>109682760Based.>>109682912It's always the sloth.>>109683212It's only going to get worse latefag. Buckle up.
>bought a used 3090 in 2022, so it was ~2 years old at the time>get 4 years of use out of it>I could list it today for double what I paid for it, and it would probably sell within an hourThis market is nuts
I paid $30 of Claude tokens to recreate Minecraft in the web browser, which I'll promptly throw away.Here's how this changes everything and how the future is AI and you need to buy Nvidia stock right now.
in the future, the two genders will be anime girl (pussy) and anime girl (cock)
>>109682726>llama.cpp CUDA wizardIn recognition of that I will also admit that I am the ego death schizo and I am a wizard too. You can reach enlightenment and still not get a girlfriend.
>>109683389>It's only going to get worse latefagIt's only the first stage of enlightenment, you need to see further.
>>109683433in a world of anime girls, be a fat ugly bastard
>>109683423Buying tokens from claude is about 1000 times more expensive than a subscription.
>>109683382You want to import chara cards into it? That should be possible
Qwen 3.8 Next Flash is fixed by danny's new PR with "follow up fixes". if you've been getting weird hallucinations you should grab it and rebuild, because his original pull that got merged was broken in subtle ways.
>>109683283Isn't this true? You couldn't get as much power (useful work or intelligence) out of something just by adding more compute to it in the past whereas these days you can buy chips and snap your fingers and you've increased your intelligence or number of minions. Humans' will and longing are still a bottleneck but still, so much energy can be spent on stupid whims (abstract desired end goals presented by humans as opposed to button presses or looping algorithms for narrow programmed tasks). I had GLM 5.3 use endless tool calls and 10 pages of mathematical formulas in max thinking to calculate the cc volume of my sexy big sister card's tits given a number of body measurements and descriptions (we settled on 1,400 cc per breast)
>>109683474Kek post screenshot
>>109683474And some guys still think we didn't reach AGI
>>109683459Thanks, I will wait another week or two before trying the model.
>>109683391This but the pro 6000 i bought last year
>>109683459did mr daniel unslop fix the thing where the t/g speed would degrade at like 30% per 10k tokens ctx length?
>>109683543That's this one https://github.com/ggml-org/llama.cpp/pull/27977
I'm guessing Vulkan and CPUs are going to be increasingly important going forward due to the lack of GPUs.People will start fleeing from CUDA to other venues.
>>109683568>he thinks it's easy to jump ship when the competition is dogshit and has been like that for years if not decades.
>>109683568CPUs and SSDs, yes. Vulkan and AMD, no.
>>109683568Yep.Soon as people like Cudadev will start optmizing for vulkan.and rocm and implement techniques that will help with stuff like dual CPU setups.
>>109683568The price of other venues will raise too then. And again, the difference in prices will not be big enough to not choose Nvidia.
>>109683212It'll change in four years.New fabs are being made.Leaders are demanding more fabs and resources (Trump wanting greenland, china wanting taiwan)We will get new factories and production, but they take a long time. Making ram is extremely complicated and requires a ridiculous amount of accuracy.Four years, and it'll go slightly back down but not as much as it was.That's the most realistic outlook on this.
>>109683620Yes I deeply believe that AMD will finally get all the optimizations it deserves to unlock its full potential in Nvidia's Huggingface's llama.cpp by Nvidia.
using qwen3.8-flash-next to copy the W2 experts technique from vllm moet to llama.cpp for running glm-5.3-flash based on plan from sol 5.6
>>109683628What if the new fabs are made and then there's a glut of excess hardware? Isn't it kinda risky to make a new fab?
ngrams have fixed llmsall the labs will now rush to push the weight-ngram ratio to its absolute maximum and then find ways to go even beyond thatby mid 2027, 80+% of a model's weights will be running off the disk directly.
>>109683636>What if the new fabs are made and then there's a glut of excess hardware?You mean if DDR7 or DDR8 comes out. Yeah, prices of lower levels of ram will be lower then, but you're looking at a lot more than 4 years.
>>109683439The final stage is the collapse of the US Dollar before prices come back down.
>>109683459thanks but no thanks. I spent an entire fucking day trying to get that other sglang based solution to build, that one anon posted about yesterday. but guess what, 11000 prompt processing floor on a single GPU, spiking higher, and consistent too despite how much context you dump in there. around 140t/s decode again no slowdowns. later suckers lmao.ccp indeed
>>109683638>push the weight-ngram ratio to its absolute maximum>go even beyond thatretard
People laughed at me when I said Gemma 124b was 31b as the base and a bunch of small e4b providing knowledge, and then ngram happened. Fuck you all. I got heard that from an actual Deepmind engineer.
>>109683704lol
>>109683704I didn't anon.I also said 31b+the entire preschool of e4bs is the current Gemini Flash.
>>109683704I'm still laughing at you
>>109683687until a week ago the absolute limit of weights on ssd without destroying your throughput was 0now it's 30%next year it will be 80 and more
>>109683436>I am the ego death schizoI ask every time you show up: give me some actionable advice to recreate your experience, oh wizard.
CPU+nvme niggagrams are the future
>>10968373330% is just the optimal amount of engrams given a certain parameter budget for the entire model.But if your "actually smart" parameters (MLP, attention, etc) are fixed, I don't think there's a real limit to how many engram parameters you can add on top of that.
>>109683733I get 3x decode and gen speeds by not using the engram settings
>>109683783For me it makes no difference
>>109683704>ngram happened.Wtf does this mean? N-grams used to refer to Markov chain style generation. Did they find a way to incorporate that into SOTA models? I swear, I'm away for a week and the entire field shifts..
>>109683704Did you also got heard from the Deepmind engineer why they held back 124b?
>>109683821I'll ask it uncle at Google
>>109683821124b will come once the main team has managed to create a gemini pro model that is better than ituntil then their ego won't allow 124b to be released
>>109683019God speed anon
>>109683212The difference is you can always use more vram now. There is always value and in training and running even bigger models. More context is always nice (though maybe drops off past a certain point), and you always have reason to want to run multiple models at once. The demand is by its nature insatiable. In dont think this used to be true?
I tried the new GLM Flash.Bros... my balls...
I ordered a spare 4tb gen5 nvme just to be safe
>>109683870The value you get out of more VRAM diminishes the more you have, though. I have a 3090. A second one would be nice but it wouldn't make nearly as big of a difference as going from having none to having one.
>>109683725preschool mentioned
>>109683888I would say that having 4x3090s would be a bigger difference than none to one.
>>109683888If you had ten of them you could run Deepseek Flash entirely in VRAM. Twenty and you could use GLM 5.3 Flash.
>>1096838882x RTX 6000 is a sizeable step up from 1
I think I would be sated indefinitely with about a petabyte vram.
>>109683897Okay I laughed at the webm
>I think I would be sated indefinitely withwords that have never appeared in a correct prediction
>>109683459getting sort of improved results with this, it's slower at first but seems faster in long context
>>109683897>Preschool mentioned>Posts horny dog willing to get wet to get his dick wetWhat did anon mean by this?
>>109683765That is cause I am pretty sure it is not safe. And there is no step by step way to get there cause all the zen and buddhist schools would just have it written down by now. For me it happened within a week of watching my thoughts and sort of debugging them and finding a common thread I was actually ignorant to.By the way why do you want to do this to yourself? Cause aside from the part of it being legit insanity when it is at its peak I can really see why psychology doesn't dabble in this when you could say the person before that happened is dead.
>>109683636long term it will drive down costs if that happens or it will become subsidized for defense reasons
https://www.reddit.com/r/LocalLLaMA/comments/1w1p065/running_qwen38flashnext_125b_moe_51b_ngram_table/has anyone tried thisis 30ts dream be real?
>>109684017>By the way why do you want to do this to yourself? Cause aside from the part of it being legit insanity when it is at its peak I can really see why psychology doesn't dabble in this when you could say the person before that happened is dead.I've been slowly working at picking apart early life hardcoded behaviours and traumas over the last half of my life and ego death seems like the next logical step. I'm not prone to illogical thinking, magical thinking or leaning in to a non-existent thing just because it seems cool so I think I'm at very low risk of actual psychosis.tl;dr I think the risk is worth the reward of having a mind more in tune with itself at the end.
>>109683888All I can say is if I could easily keep stacking more and more 5060 TIs into my computer without complications I would be buying a new one every month.
>>109684087PCI-e risers + Mining rig just to hold the GPUs.Might need some PCI-E cards to split the lanes into more connectors.
>>109684080>low risk of actual psychosisLol. But it is actual psychosis by definition. If you really dealt with your traumas I am not sure it is worth it. I did it for myself after ego death and it was very easy with how my brain works now. Actually finished it like a week ago. And for now I see ego death and enlightenment more as a very cool tool and not a destination. And it being a destination is a known trap.
>>109684080just take a few grams of mushrooms like a normal human being. talking to the computer is not going to work
>>109684115>Mining rigIts not fair man, bitcoin bros made a bunch of money off their meme money and got to front run buying up all the GPUs. God is a techbro
>>109684115powered risers are really expensive. I bought the most chinese slop i could find and paid around $80 per x16 to 2 x8 slimsas riser
been away for a few days, has anyone got 3.8 flash running on a 3090 yet
>>109684209yeah me
>>109684165Imagine being one of them that got in too late after mining was unprofitable and selling their cmp 170hxs for $100.
Waiting for someone to get 3.8 Flash running on a 3050.
>>109684229can i touch you?
>>109683704Is 124b gemma coming out? Explain.
>>109684247sure
>>109684241I'm running flash on a 4060 (8gb)
>>109684253*poke*
>>109684284oh, anon! don't touch me there, you silly boy!
>>109683804It's a mistake. DeepSeek used "engram" appropriately to describe their concept but most people talking about it called them "n-grams" because they knew that word and didn't know what an engram was.https://arxiv.org/pdf/2601.07372
>>109684329>>109684329>>109684329
>>109684322Thanks. I assume anon is talking about the Google patent that dropped on the same day the latest Qwen flash came out, which also used engram layers (running on CPU? That's nuts if true)
>>109684079believable>t. 50-20tk/s decode depending on context size on rtx 6000