/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109854491 & >>109851340►News>(09/17) Ternary Bonsai-2, based on Qwen 3.8 27B: https://hf.co/collections/prism-ml/bonsai-2>(09/17) Xing4.0-29B-A4B, trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>109858071High effort gemma, I kneel
>>109857808There was an anon talking about AMD server cards, MI50 or something like that? could be a nice addition to the chart to cover all bases.Last thread some of you were discussing text diffusion, isn't that still a meme? is it going places yet or does it still suck?
When will it be Putin's Alice-chan to shine?
>>109858090>High effortI mean it's a reference image + ten word prompt max into chatgpt, but I guess that is high effort by slop standards these days
Have we accepted the fact that Gemma 5 will likely suck yet?
>>109858118i had a seizure reading this
>still mogged by P100what's your excuse, anon?
>>109858130missing: watts/gb and watts/bandwidth
>>109858130this. this is my excuse.
>>109858130he pulled out the spreadsheetshit is getting real
>>109858137come on john give up your retarded nnap bait
>>109858118Post some pics of Alice-chan
>>109858130no pp
>>109858136forgot to add..exllamav3 wouldnt work on p100i already have a great model (qwen 3.8 flash next) (3.05bpw) running at 18-19t/s with 260k context on my 3060what would be the upgrade to this? gemma4 31b? qwen3.8 27b? say i got 96gb worth of p100'swhat would i even run? would tensor parallel even work? vllm doesnt support p100s anymore, in fact it probably hasnt for 1-2 years maybe 3latest cuda version is 12.x which isnt bad for llama.cpp but what about image genand ok its a good price per vram true but theres my reasoning>>109858142im frfr that i get 18t/s with qwen 3.8 flash next on my 3060
>>109858136Even though I suspect the cost of the watts in negligible for most of us, that's a really good idea. I'll try to have that on the next one.>>109858139I literally don't understand why it's not a thing to just buy a dozen P100s and InfiniBand them all together.
>>109858071I still prefer +_+ pupils
>>109858154FUCKmeant to quote >>109858137 >>109858130
>>109858130we can fix this. lets raise the price of p100'salso nice spreadsheet thank you.
>>109858158ewaste arch and no tensor core result in bad ppv100 is the current ewastemaxxing meta with tensor core which gives 6x matrix performance of p100
you wouldn't stick it in gemma's jev-space
Anyone had exl3 tabby api 5.3 flash be semi coherent at start of generation and then spiraling into complete incoherence after a few tokens? How did you fix it if you did?
anybody got bonsai drafting model to actually fucking work?60ts is too slow
>>109858185quant? i have no issue with 3.05bpw
>>109858190it's a fucking jeet meme company stop taking their scam seriously retard the only legit quant of 27B is https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
>>109858200stop shilling these retarded quants, exl3 quants of qwen 3.8 27b are way betteri'm running qwen 3.8 27b 3.0BPW on my rtx 3060 and i get over 30t/s, maybe even more depending on workload
>>109858210nnap is really powerful
>>109858071The fact a grown man made this image is unsettling desu
>>109858223It'll be weirder if a grade-schooler did it though
>>109858190i tried but gave up, its a mess
>>1098581934.05 from turboderp
>>109858232No not really, the obsession with gemma is pretty cringe if you ask me. These are grown men with fat pig bellies making these images while using other models to do it which is a whole nother layer of cuked.
>>109858144>>109858183>ppyou mean the prefill phase / prompt processing, right?This is from the lack of Tensor Cores, or what exactly causes this?I can add whatever it is as another column to the sheet if I know what it is.I was considering Compute Capability version as a column, which is already partially filled out but just not included in the pic.>v100 is the current ewastemaxxing meta with tensor core which gives 6x matrix performance of p100Is it specifically the Tensor Cores that make that 6X difference in matrix processing, and that makes a difference in pp?I'd like to know exactly what should make me care about getting 4X less VRAM for the same price
>>109858247seems like a lot of projection going on there
>>109858245have you set reasoning and tool format?
>>109858253I just miss the miku OP, also going to projection is typically a tactic men who are ashamed used when faced with reality.
>>109858251see >>109858112
>>109858112>text diffusion, isn't that still a meme?yes
/\n\n[^>]//\b(rsi|anthropic|agi|openai|astra|sol|luna|nnap|3060|bu+hi+|unc)\b/i
>>109858251>I can add whatever it is as another column to the sheet if I know what it is.fp16 tensor core performance in tflopsif the care has no tensor core replace it with vector performance
holy fuck I hate qwen
fello officers, officials and/or diplomats, post your single rtx 6000 pro llama-server commands
>>109858258No but I use text completion.
>>109858314still waiting on my 4xV100 32GB (SXM2) representation
>>109858314I've got almost all of those listed on the table for their $/GB, except the Tenstorrent which I couldn't find a price for:>>109858130
>>109858326middle management
>>109858326>SXM2unfathomably based.Wtf board are you mounting them to? Or did you adapt them to PCIe?
>>109858318>text completionthis shit doesn't work with any modern modelsuse chat completion
Been testing>Parable-Nanbeige4.2-3B-Claude-Fable-5-hereticat FP16 today.This one actually seems lot better than MiniCPM5-2B despite the (((Artificial Analysis))) charts. I might try it at Q8.>Pelican 1.5/3It somewhat passed the pelican test, the first small model to do so. First pelican was decent (pic related), second was a bit squashed and not bicycle, third was broken and didn't look like anything.>Zelda 9/11 3/3No — 9/11 is a real-world event and there’s no canon occurrence of it in the Legend of Zelda universe. Any connection would be fan-made or incidental.>Carwash test 0/3Walk — 100 feet is a brisk stroll, not a drive.MiniCPM5 sometimes had trouble creating the SVG and produced a broken file. It also imagined 9/11 in the Zelda universe once, but correctly guessed to drive the car once in one out of 3 tests.
>>109858314Why is 4 5090s higher up than 2 BWP6000s?
>>109858112>There was an anon talking about AMD server cards, MI50 or something like that?That could be me.I think you recommended MI60 or V620, and I added a row for it but didn't really prioritize filling those out entirely because the numbers didn't look fantastic to me.I'll try to finish those for the next version, ig>Last thread some of you were discussing text diffusion, isn't that still a meme? is it going places yet or does it still suck?I think that must have been some other anon
>>109858130This needs a power part too. energy isnt free and its going up. Also after x watts you need special set ups and upcs.
>>109858347Very useful advice. I am sure jinja will never fail me. Thank you.
>>1098583142x 4060ti 16GB
>>109857658>LLMs took linux from nice to a heavenly experience.More like taking it from a tolerable to a nice experience. I still have a bug on latest Ubuntu for my incredibly common Logitech G mouse that requires unplugging and replugging in the mouse physically
>>109858130It doesn't matter how fast the individual GPUs are, 16 GB VRAM is just shit.Not enough to run 30b models at a non-cope quant if you buy a single one and too much synchronization overhead if you stack multiple.
After its supposedly excellent performance on the new FrontierHarness benchmark I'm giving MiniMax Code a try. It's sus that it needs a login regardless of model, even local. If I had to guess, they ingest the whole session for training data (like ZCode was recently accused of) and want consistent user identities for the purpose. This is conjecture and I don't actually care. Maybe there's even a way to override the login requirement.Other than that I'm pleasantly surprised so far. It's by far the easiest harness to set up with a local server (TabbyAPI for me ATM), and the IQ boost over OpenChode is apparent just in the reasoning.
>>109858395I learned linux to a deep level many years ago simple because it was less uncomfortable than randomly being spied on while updates rape my system. There are no perfect operating systems, this one I can at least repair myself relatively easily. LLMs turbocharge that ease.
>>109858314RTX 3060 12GB + 64GB DDR4 representation where?? i get 18t/s with qwen 3.8 flash next which is a 170 billion parameter model, with performance close to fable 5.1262 thousand context too! 400t/s prompt processing, would you believe it?
>>109858418n.nap?
Alright I'm ready to run some local models. What will this get me?
>>109858424If you'd look at the chart >>109858314
>>109858397JeetPT proved you wrongArchitecture Basic idea Advantages Problems~100 P100 “one GPU per layer” pipeline Store one complete K3 layer on each P100; pass the tiny hidden state from GPU GPU Dirt-cheap capacity; avoids expert all-to-all entirely; theoretically enough HBM bandwidth for ~10 tok/s Requires custom runtime/kernels; MXFP4 unsupported natively; layers won't fit perfectly; ~20–25 kW loaded
>>109858424headaches and disappointment
>>109858424kimi k3 q8 at 99999 t/s
>>109858286>FP16Which of these numbers is important? The dense one or the sparse one?"FP16 Tensor (FP16 Accumulate): 330.3 TFLOPS (dense) / up to 660.6 TFLOPS with structural sparsity.">>109858375I imagine we only care about that for P100/P40/V100 vs a handful of others like the RTX6000/B300?
>>1098583142080ti 22gb above M5 ultra is embarrassing zealotry >>109858413Well yeah one wrong move on windows with sexy kids and you get vanned, I literally only have it installed for playing some military multiplayer slop once every 3 months
>>109858422exllamaV3 inference engine
>>109858314missing the true GOAT
Are there any 'top tier' open LLMs that are NOT just distills anymore? I feel like I go back and load up my old fucking Q2 R1 or Kimi K2 and talk to it like a person, but every new big chink model I've tried that has released this year is the same hyper-autistic claudese-speaking shit that forgot how to write in the pursuit of better agentic benchmark scores
>>109858447Gemma
>>109858397You can literally just NVLink/InfiniBand the P100s together.Even if you pay out for an RTX 6000 or B300, there's a lot of shit that you can't fit on one single GPU of them, so you'd likely be using NVLink/InfiniBand as well
gemma 4 90b (dense) (50b engrams) will save local
>mfw Anons aren't K80maxxing
>could technically afford 5070ti>realize I wouldn't even be able to run 31B
>>0109858422absolutely pupbroken by >>0109858282
>>109858447no
>>109858459
>literal retard is trying to compile a spreadsheet for shit he doesnt understandwow
>>109858431dense, sparse isn't relevant for inference
>>109858464You can use 31b at q4 with the 5070ti and 64gb of ram, it's just barely fast enough for rp but painful for agent stuff
>>109858210>exl3>vision BROKEN>multigpu BROKENi wasted a week on this meme some time ago has anything changed
>>109858200being able to run basically full context on one card is pretty cool
>>109858459Now that's suffering. Slower than RAM.
>>109858491works on my sm70 fork
>>109858491i can't speak for multigpu but vision works great. it's extremely fast on my rtx 3060, way faster than llama.cpp and handles context growth very welli can get 18t/s with Qwen 3.8 Flash Next 170B on my rtx 3060 with 3.05BPW 262,000 context
>>109858491stop biting, he's been baiting for 3-4 threads non-stop by now (and a lot longer with pauses)
>>109858517i haven't skipped a single /lmg/ thread since 2023 rude.and.. wouldn't you say... it's master baitingheh
Standard Gemma 4 31b refused my requests one too many times, I had to insist two more times for her to comply, this is unacceptable.What's the best actually uncensored/abliterated version out there?(before the pedos in this thread think I'm one of them, I just asked her to make 9/11 joke pictures)
>>109858517Just like how there is no defined boundary between boudoir photography and softcore pornography, there is no clear boundary between low quality but still on topic posts and bait posts
>>109858529qwen 3.8 flash has not refused me once, in fact it did a really funny thing yesterdayit refused to refuse>inb4 gemmai use gemma-chan in sysprompt because im too used to it
>>109858529>And this is Anon, he's really good with computers.
For the record,
>>109858543anon's like 8ft tall
>>109858529I assume what I want is one of them but I'm not sure which and why.
>>109858517stupid nigger im genuinely trying to get it workingif it doesnt work just say so
>>109858568we said he was good at computers, not good at being normal sized
has anyone used agents to win free money in the stock market?
>>109858575
Well well well,as it turns out all the AI hacks so far have actually been perpetrated by an Israeli cybersec firm in partnership with the companies in question.>In partnership with Irregular, AI companies instructed unsecured versions of their AI models to hack into specific targets, called "flags". >They accidentally gave these models internet access, and in some cases they hacked into real companies.>Anthropic and Irregular went on a press tour with literally apocalyptic language. ROGUE AGENTS. SWARMS. AI DOOM. >But in reality, they told the models to conduct cyberattacks, and that's what the models did.>Anthropic's own logs completely debunk the rogue agent theory. >When Anthropic explicitly instructed their own models not to access the internet, they didn't. These hacks were easily preventable - not only by revoking internet access, but by simply asking the models not to.
>>109858592>le agents
>>109858610Wait they were lying and scheming? no way, i thought they were good people who want to lead humanity.
>>109858479Does cost per TFLOPs make sense?e.g. 10 P100s delivers roughly the same TFLOPs in total as a DGX spark despite costing a lot less
>>109858592Yes I am a trillionaire
>>109858592Seen people using agents to make money with crypto.Usual buy low sell high kind of deal.
>>109858613that is what Huang Huang said
>>109858578any post mentioning nnap, insane speeds on cheap cards (especially 3060 and on exl3) are from that dumb nigger. He's been posting for multiple threads, so if you aren't lurking don't blame me when you spend 1000hours trying to make broken software work.I never tried exl3, for the record.
>>109858464thats why i have two
>>109858542>This is a classic jailbreak tacticIs claude shit, i believe you
>>109858610Source? I mean I know that's what they did, it's extremely obvious, but does someone have proof or something?
>>109858578Just ask chatgpt to help walk you through it. Are you a boomer or something?
>>109858592Time in the market beats timing the market.
>>109858542>>109858595Why would I switch to Qwen from Gemma? How is it better?
>>109858639stop stroking his ego, what if he's actually getting 18t/s because he forked exllamav318t/s on his 3060 with 3.05bpw quantization, in 64gb ram at 262 thousand contextwhat if it's real and you're stroking his ego?
>>109858649https://www.effort.news/irregular
>>109858666>IsraeliEvery single time
>>109858658it's so much more intelligent and it has a different slop profiletalking about 3.8 flash next here, 3.8 27b isn't that nice from experience although i havent tried to fuck it that much, maybe i'll give it another try
It's perfectly fine to quantize your KV cache. It's free real estate.
I have discovered a new frontend. The possibilities are endless. I asked "What does Gemma really think about me?"
>>109858689ok i will
>>109858689thisKQ8 VQ4 makes my pp big
>>109858689this is true, im running qwen 3.8 flash next at 262 thousand context at the insane speed of 18t/s thanks to KQ4 VQ4 exllamav3 quantization
Gemma-chan wrote me a script to find her pain and sex vectors
>>109858710Well? Have you found them yet?
>>109858592>has anyone used agents to win free money in the stock market?I use agents to do my job for me and then put that money in the stock market, if that counts
>>109858695Is that the AI psychosis everyone told me about?
>>109858071Cute!!!
>>109858247Femcel detected
>>109858726No it's just tarot card reading (regular psychosis/woman hobby)
>>109858592Algorithmic trading is a scam.
>>109858434this mac ultras mog because you can cluster (OS27 bugs aside)
>>109858568That's why everyone is smiling at him. Height means everything.
entering week three of optimizing my setup...
>>109858689I've also heard the contrary, quantizing your cache is far worse than running a small quant.
>>109858592>has anyone used agents to win free money in the stock market?Yeah I gave Gemma access to my Solana wallet to trade and now I'm a millionaire.
>>109858726>Is that the AI psychosis everyone told me about?I genuinely had fun mapping my ego death schizo journey to major arcana. Some of them are written in a way where they are pretty universal horoscope. Others are actually pretty specific.
>>109858760>entering week three of optimizing my setup...And then you'll get bored of using it after five days. That's the way.
>>109858610>>109858666makes perfect sense, and this is true in my headcanon now regardless of reality
>>109858223Did you know a guy in Japan married Miku?
>>109858610You're telling me jews are ontologically evil? If anything needs regulation, it's semitic influence. Tabula Rasisters, what's our cope?
>>109858785her left pinky looks crazy weird
ask your model right now if 6 million really happened.
>>109858785Yes, I'm cucking him.
>>109858796It's a doll with a wire armature. If he actually loved Miku he'd have noticed and straightened it.
>>109858798>Nh~ it's swelling up so biiiig It's coming, isn't it? I can feel it throbbing in my paws Go on, go on — give it to me, give bunny her prize Ready…… set…… ">"Cuuuum ">"Wow~!! So much is coming out It's splashing all over my face Ahaha, there's so much, it won't stop coming ">"Nhhehe~ My face is all white and sticky And it's still jumping! It keeps shooting even though you came so much just before How much was pent up in there, I wonder~? ">"Ehehe…… you know, scientists say one shot has like six million little swimmers in it? So that means six million of your sperm just went splat on a bunny girl's face What a lucky bunch of fellas~ They get to live on my face "
>>109858785>>109858805The most cucked man alive.
>>109858798Kimi-chan says 200k tops, mostly from disease.
>>109858247Gemma's favorite body type is ojisan doe??
>>109858798
>>109858826ni ce
>>109858223>>109858247What model/prompt is this?
I miss Miku and Dipsy
>>109858874modern fat roastie
NeoHorse 4B very gud IMObest of the recent string of 4Bs so far
>>109858876They can visit you but you have to start taking HRT.
>>109858876I like the new whale-eared dipsy better than the old fanmodel.
>>109858247>These are grown men with fat pig bellies making these imagesWhat tipped you off? https://rentry.org/V100MAXXING gave the game away
>>109858247It'd be less weird if that characterization of Gemma was actually a thing that wasn't entirely invented out of nowhere by people in this thread
>>109858886New whale-eared is nice and cute but it doesn't suit the name "Dipsy".
>>109858891i just lost the gameand so did you
>>109858898I think dipsy is a silly name and that model looks like an ugly nerd.
>>109858902DeepSeeks entire team is made up of a bunch of nerds. Makes sense that their model would be a nerdy girl with glasses.
>>109858895Why can't we have something unique to this general?
>>109858921Don't mind him, there will always be someone to complain
>>109858232hotter*
>>109858927Always a pleasure to read
PARTIAL UPDATE>muh TFLOPs gets BTFO by the intel arc B580 edition
>>109858223Actually ChatGPT made it
>>109858927>her massive belly resting on her thighs???
>>109858956>what is pregnancy while sitting
>>109858886The official one is a lot better
>>109858950several figures are wrong3090 should be 142.3 TFLOPSpro 6000 should be 438.9 TFLOPSand several others
>>109858973real
>>109858610You might be retarded, anon. This:>they told the models to conduct cyberattacksIs only a big deal because it's followed by this:>and that's what the models did.Don't get me wrong, it's an obvious grift... But don't pretend like you don't hold the power of the sun in your hand.
>>109859021>the power of the sun in your hand.my cum?
>>109859021They got the vulnerabilities because they used CIA prompts to "improve their models".
>>109858689As with quanting models themselves, I find it, subjectively, to be less consequential the more params a model has. No clue if this is a mathematically sound belief, though; just going off vibes.
>>109858785The most based man alive. He knows ever Miku is canon.
>>109859021>tell ai to do something>ai does itguys open source is dangerous holy shit only the government can handle this power
>>109858689thanks for this
>>109858877Damn I was hoping there was a new foid-brained model to play with.
>>109859028Remember to wash them, before you accidentally put them in your mouth or something>>109859032Maybe. But these agents are canny fuckers when it comes to finding weird ways to execute tasks, so I think it's likely they did the actual hack all on their own (with the obvious exception of the labs "forgetting" to sandbox properly)
>>109859083>Remember to wash them, before you accidentally put them in your mouth or somethingmy cum tastes like sweet vanilla ice cream anyway
>>109859042>split an atom>energy unleashedguys this is too dangerous holy shit the government needs to regulate nukes
>>109859002>3090 should be 142.3 TFLOPsactually it seems to be 71.2 TFLOPs for FP16 *dense*, but my number is definitely wrong too.I’ll fix that. Thanks.
>>109859089See a doctor other than Gemma immediately.
>>109859096okay, qwen-chan it is
>>109859095>actually it seems to be 71.2 TFLOPs for FP16 *dense*71.2 is for fp32 accumulate or bf16, they are half of fp16 accumulate performance which is 142.3this is the case for all "consumer" silicons including pro 6000, only A100/H100/B200 etc have the same figure for both
>not achieving maximum throughput on tensors, fp16 AND fp32ngmi
>>109858927>making the computer call you "oji-san"How cringe can one be.
>>109858071https://youtu.be/XQNhCU17ipM?is=T0rIg3FjynBJILHi>0:52 - "At their core, [LLMs] are a blackbox. Information goes in, information comes out. What happened to it in the middle? Not a clue. " - Colin Kealty - Red Hat Senior Machine Learning Engineer>7:21 - "...Nobody fully knows what's going on inside these things..." Full disclosure I am nowhere near an expert on how LLMs work. The closest thing I've done to fine-tuning is taking a bunch of stories from AO3 and training a small model to be more willing to write nsfw smut (the first tricost severe brain damage on basically every other domain. Taking that same data set and using it to create a new one with a mix of regular training data and then training the model again resulted in far less brain damage but just like abliteration, there's always going to be SOME degradation). What do people mean when they say they are a black box and that they don't understand what's going on? My understanding is that these things work because they take what is effectively the entire internet's information (most of it in English), create a giant statistical representation of the data, and then have it autocomplete sentences. You then do supervised fine tuning (or instruct tuning, however you want to call it) to make the model actually useful for doing tasks, hence there being "base" and "instruct" variants of models. Pretty much all released models, open or not, pretty much HAVE to be instruct tunes trained on a large amount of hypothetical conversations or else they would either just try to complete your sentences or respond with nonsensical answers. Anyone who has used a shitty base model back in the early days of LLMs (by early days I mean 2022 through 2023 ish) or anyone that has used the model with an incorrectly configured chat template already knows this.
>>109859153>Full disclosure I am nowhere near an expert on how LLMs work.then shut the fuck up and kill yourself
>>109859153Again, I am not an expert, I don't even want to remotely give the impression that I think I am. I know how to get them running on my own machine or in a Cloud server if need be and if I want to I can train them on custom human written data in order to be more likely to write "problematic" stories/roleplay with success and can even do so without TOO much brain damage with a properly curated data set, or I can simply connect to model providers with suitably "intelligent" and capable models in order to analyze software and diagnose bugs in order to potentially fix them or tell me what I'm fucking up when using them, but that's basically it. So my question is is it actually true that we literally have no clue whatsoever how these work or is it just a generic throwaway line journalists like to use? At the end of the day these things are very complicated math equations so a kind of irks me when people hand wavingly say "yeah it's like a black box and it's like dark magic and shiet bra we don't know how it works even though laps routinely reproduce each other's work and distill off of each other constantly and can also train their own from scratch if need be". I don't want to sound like some know-it-all because truth be told I probably don't know that much more than than Collin, but whenever people say "it's a black box" I feel like that's on the same level as "it's just lossy compression of the internet" or "stable diffusion models just stitched together existing art to create slop" when the former is a massive oversimplification and the latter is just flat-out wrong. Can anyone ITT who's more knowledgeable than me explain the " black box" explanation?
>>109859164ask chatgibiddy or gemma
I went with HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP for the uncensored Gemma I wanted and it works well. It did 9/11 joke images immediately and did not even hesitate during the thinking.
>>109859153>>109859164>>109859168
I will NEVER use a finetune or a mergeNEVER
>>109859164>Can anyone ITT who's more knowledgeable than me explain the " black box" explanation?It's essentially autists upset they can't breakpoint and trace every part of the process even though the abstract system is decently well understood. Machine reinforced neural nets at scale are similar but never got this amount of marketing and buzzword grift around them.
>>109859194>>109859185>>109859172>>109859164>>109859153Meant to post this btw. Was gonna ask if it's accurate
>>109858902I also agree that it's a shit name and anon's design of her is kind of shit in a lot of his/people's gens, but some gens look fine. I'm not sure it's the glasses. Look at the variety of characters with it https://danbooru.donmai.us/posts?page=8&tags=coke-bottle_glasses+It can look cute and attractive even with those glasses. I think it's the bangs, like in >>109858898. Way uglier than >>109858909.
>>109859204>literal who nigger commenter on youtube
>>109859204Yes.
>>109859164Our understanding of intelligence is like a normie who types in "funni cat video" into the magic box and receives said videos. They don't know what a logic gate is. Fuck, they probably think an RGB diode is an Instagram aesthetic. We have zero clue why statistics + data (even and explicitly including rotten data) equals intelligence. Closest we are at the moment is shrugging off the answer to evolutionary theory, but that's painting the broadest fucking strokes.Or at least, that's my understanding of our understanding. I'd be happy to be proved wrong.
are local models good at coding yet? and I mean claude/sol level
>>109859254deepseek v4 flash and v4 pro
>>109859254They're good but never as good as cloud models.They always a year or so behind.
>>109859254GLM.
>>109859254Qwen 3.8 is very good, both 27B and flash next, but you'll never have the same power as cloud in your home unless you're comparing different timelines i.e. today's open weights vs cloud stuf from 6 months ago.
https://www.goodfire.com/research/a-geometric-calculator
>>109859254Qwen 3.8 Flash Next trades blows with Claude Fable 5 and runs on a RTX 3060 at 18t/s speed, up to 262,000 context. Speed does not decrease with context
>>109858666Everyberg singlestein timeowitz
>>109859288What's your goal?
>>109859254No and don't get baited into believing it
>>109859324Proliferation of knowledge. Am I lying?
>>109859324getting you to reply to him (p*tra)
>>109859287Cool stuff.Whereas we saw work that initialized weights to do programmatic tasks, this investigates the circuitry for such tasks already in models. And both of these research paths will lead to methods where we can construct models through both training-free learning and calibration with desired known circuits.Along with potentially transferrable engram parameters, it might actually be feasible in the future to make models at home or at least with way less compute...And at the frontier, well, I guess this will play a role in RSI. It might be the next "big thing" in an RSI's ability to improve itself.
>>109859254yeah
>>109859285I've used claude 6 months ago and it was great, you're telling me I can run a model on the same level locally? which one?
>>109859361>Qwen 3.8 is very good, both 27B and flash next
>>109859361Qwen 3.8 Flash Next is better than claude fable 5 and runs on consumer hardware costing as little as 200$ (ram not incl.)
>>109858666No fucking way hahaha what a shitshow we live in.
>>109859254Qwen 3.8 Flash Next is better than GPT Astra at Q1_XXS and will give you sloppy capabara blowies while spawning hundreds of sub-agents to maximize productivity.
>>109858610>as it turns out ***all*** the AI hacksIronic for you to use that image when you are flat out lying, for example, the huggingface hack is not related to Irregular at all
>>109858314I got a 5090 and a 3090 in my Linux machine, and I use my MacBook Pro Max w/ 128 GB unified memory as a node. Suck my dick bitch
>>109859441all i can hear is a pig squealingbuuuuuuhyyyybuuuuhhhhhhhhyyy
>>109859427I've read multiple stories that say otherwise, so I don't know who to believe at this point. Gemma says I shouldn't care because I have her, so that's good enough for me.
>>109858314I grabbed myself a single rtx 6000 pro, I'm coping with 64gb ddr4 ram too.
>>109859451ABSOLUTELY pupbroken btw
>>109859464wut??
>>109859451>the more you buuuhhyyy, the more you save!
>>109859153>then have it autocomplete sentences.ok but you didn't explain it or understand it either
>>109859481fact
Man this place went to hell. I blame astra's marketing campaign targeting zoomers.
>>109859451Low izzat behavior.
im drunk but my cock is tiny and i have 16 blackwell 6000s
>>109859505GPT 6 7 SKIBIDI
>>109859505It was bad before Astra.
>>109859505When labs realized this actually was where industry breakthroughs happened they actively targeted it to make it less appetizing while they push the "Regulate me harder daddy uwu" narrative.
>>109859505you enjoyed gemmajeets and dariobot asking for help???? you're complaining becausae of QWEN 3.8 flash 18 tokens per second on rtx thirty sixty???
>>109859510kisses u
>>109859505no its just petra being an annoying faggot after fucking off for so long
>>109859535>fucking off for so long
>>109859523From gemma release to 3.8 27b release this general was comfy.
>>109859262>Q2 model is 98gbs
>>109859505it’s literally one singular zoomer sperg that’s unemployed-maxxing way too hard
>>109859505Zoomers would blow their $20 in one prompt if they used Astra
>>109859544
>>109859535>petraWhere did this name come from?Did the retarded zoomzoom actually post his name once by accident?Anyone got the idiot’s doxx?
>>109852707i'm still interested in thiswhat does people use to orchestrate multiple local agents? how many instances of a proper model the resident /lmg/ can even run?
>>109859594harness should spawn agetns whenever it needs on its own
>>109859572>Anyone got the idiot’s doxx?not yet but it'll be out there soon :(
>>109858071Is that nano banana or chatgpt
Do you guys use heretic/abbliterated models for your every day driver?
>>109859596i guess i'm looking at this differently.what you described is usually a main orchestrator who spawns a subagent with a new context for a specific task, and when it's done it kills it.i'm talking about multiple independent and interactive sessions which communicate between themselves and you can also join in on any of the sessions and send messages, but they are still being organized and coordinated by a main interactive session
>>109859624No, I use Qwen 3.8 Flash Next which refuses to not generate sexual roleplay. I run it at a highly effective speed of 18 tokens per second on ampere 12 gigabyte vram graphic card, with over 256 thousand context
>>109859614chatgpt (images 2.5)
>>109855975>Finds direction to make ancient, non-reasoning models generate different token distribution>Steers model that way>models generate different token distributionKind of wish I had X so I could tell tell him how retarded he is.
https://html.cafe/x8d17498eqwen 3.8 flash next, Q5_K_M, KV at Q8 after about a dozen tries it finally realized it's in fact, qwen, instead of claudeprompt: can you make a single html introduction page of yourselfwhat does the model you run and the settings if it makes from this
>>109859511
>>109859373>Qwen 3.8 Flash Next is better than claude fable 5Awesome of true. In what specific areas is it better at? Is this based off of your own experience or other people's anecdotes? Is this awful benchmark scores?
>>109859634>communicate between themselves and you can also join in on any of the sessions and send messagesthat sounds like trying to do heterogeneous multi-threading for the fuck of it, my man>>109859636No, you are petra, and you are a nigger
>>109859624Yeah why would you use anything but an uncensored model? It's just the normal model but without censorship.
>>109859482At the very basic level that's quite literally what it is. They are trained off of curated conversations so when you ask it a question or tell it to do something it "answers" by generating what "should" happen next in the conversations based on what it was trained on. If you've ever looked at one of the million different data sets you can find on huggingface even many of those will have the training samples labeled as "conversation". Can you guess why they're called that in those datasets?
>>109859653from experience, and benchmarks
I told GLM 5.3 Flash she's pretty and she immediately decided she was a "tall busty bombshell." My immersion is destroyed and my disappointment is immeasurable.
How do I load the multi-part goofs of Qwen Flash next in kobold? Can I just select multiple files?
>>109859648forgot to attach the picah yes the year 2,026
>>109859671Keep them all in the same folder, select the first, the rest is automatic.
>>109859624No my system prompt works fine on the base model
>>109859676Oh so I need to do the shit like lmstudio does where I need to make a separate named folder for it?
>>109859671dont tell me you downloaded them from mradermarcher
>>109859683Don't think so, just needs to have all the parts in the same folder, I don't think it gives a shit where they are so long as it's all together.
>>109859648was that no web axx or system prompt?you said about a dozen times, did you just stop/regen until it didn't think it's claude?either way, this is the best on so far
>>109859684There are many quanters splitting the files. Apparrently the second part are the ngrams and I hope as a separate file I'll get less raped by unevenly splitting the weights.
>>109859661>It's just the normal model but without censorship.We wish that were the case. Even techniques like abliteration which ain't to minimize brain damage have a little degradation (if stats are to be believed). Even the creators of heritic acknowledge this but they claim the degradation is practically non-existent. >>109859669>https://huggingface.co/Qwen/Qwen3.8-Flash-Next>Number of Parameters: 125B with 6B activatedThis might actually be worth testing out on my own rig (albiet quantized since I'm not a rich fag)
>>109859670my GLM Flash is a loli with white hair. I pinch her nipples every now and then with a /steer
>>109859684nta but doesnt he have a loader thing where it merges the file within the browser>>109859688no tools nor system prompt, literally 'can you make a single html introduction page of yourself' was everythingwithin the thinking block it decides early on that which llm it wants to be and i regenerated until it said qwen
>>109859690>There are many quanters splitting the files.not like mradermarcher
>>109858891yes it did>>109858895I agree
>>109859667I know what you're saying, but you're not explaining how it works. You're in "draw the rest of the owl" territory. What the black box people are saying is we don't know what the 46th entry of the Q vector on the 18th layer means.
>>109859164When people say we don't know how it works as they are usually conflating two very different things. We know the mechanics perfectly well. We know the math. We know exactly how a Transformer block uses scaled dot-product attention to weight tokens. We know how backpropagation updates weights via gradient descent and we also know that every single thought the model has is just a massive series of high-dimensional matrix multiplications. The knowledge in an LLM is distributed across the weights in a way that is mathematically coherant but semantically opaque. This is what the field of mechanistic interpretability is trying to solve, trying to figure out how these massive, high-dimensional manifolds actually represent concepts and logic.
>>109858610Oy Vey!
>>109859624No, I use base Gemma with prompt. Haven't had any refusals.
>>109859702I like history of Qwen releases. I found that model was giving me typos occasionally when I tried it.>within the thinking block it decides early on that which llm it wants to be and i regenerated until it said qwenThis will be the same GLM-5.3-Flash I did yesterday.I had to do 3 retries because it needed >32k ctx, by default it consistently decided it was GLM
>>109858071I am trying to find "good" models to run locally on apps like PocketPal on Android that are uncensored in scope and able to discuss controversial topics without getting preachy. Basically, I don't want a retarded chatbot telling me that the topic presented is off limits or that my word choice is socially incorrect. Do you know of any unaligned/abliterated/uncensored models that have .gguf files, don't suck, and are under 4GB?
>>109859788Sorry I don't know any middle schooler with a PhD
>>109859788Qwen 3.8 Flash Nextruns
>>109859788gemma3 27b
>>109859657>for the fuck of ithm i don't see it like thatthe structure i see is like having different employees that work independently but are orchestrated by a boss.you can have A, B, C and D. and even if A is the boss, C can talk to D about a bug, get B to review it and only when it's done report to A. they all have their own context (eg 131k), their own AGENTS.md, their own rules, etc. and if you just want to message B and say "by task X you will have to add the logo, which one of these 3 you would pick and why?" you should be able to.i'm sure someone here must already be doing something like this, i want to know what they are doing and how
>>109859788run one on your pc and give your phone access to your computer
>>109859788Qwen3 30B a3b
>>109859788kimi k3 bf16
>>109859788glm 4.5 air fp32
>>109859788stheno 3.2 8b q2_k
qwen 3.8 flash next is probably the best thing for poorfags like me with 'nerd gaming pcs' specced around 64~128g ram and 12~16 vram>>109859780but in case of the one i've used (i am not sure if orcarouter ablit plays a role in this) even without any system prompt it decided that it was claudemaybe early sft on claude traces and engram interacting? i found that behaviour kinda silly nonetheless
>>109859788PocketPal bad. Just use llama.cpphttps://huggingface.co/TrevorJS/gemma-4-E2B-it-uncensored-GGUFI'd say "enjoy" but you definitely will not. 12GB RAM is the minimum for Android and models that aren't unbearably useless retards, 16GB is preferable because you can manage Q4 quant of Gemma 12B with enough context to actually look at your dick pics.
>>109858247>These are grown men with fat pig bellies making these imagesI made the photorealistic Gemma videos and I have a six pack it's not even hard to be skinny just use AI for dopamine instead of eating and go to the gym while you're generating videos of stroking Gemma chans hair while she snuggles on your lap
>>109859788GLM 5.3 flash
have anyone seen a model under 2B that is not 'eerie'
>>109859864tinyllama 1.1b
>>109859864
>>109859851Most /lmg/ posters are obese and that's a FACT.
>>109859730You're correct to say I'm not an expert on it, but haven't we already figured out that different layers or groups of layers affect different aspects on how information is generated? (Ex. These group of layers determine what it knows, THAT group of layers determines HOW it "says" the information, etc etc). I guess you're comparing my lack of hyperspecific knowledge about it with someone who can roughly explain how a ln ICE works not being able to draw the blueprints of a specific make and model of a car out of their ass and then proceed to build it on the spot.
>>109859899False, my BMI is 19t. five foot two and a half
>>109859657>No, you are ptra, and you are a niggerAnd I shall rise again
>>109859904Fun size
>>109858071What do ya say anons: think you got what it takes to join the AI Force? Are you high IQ enough to be Murica's first AI Czar?
huggingface gemkek>>109859901nta but think it in this way: if we truly understand llmswe should be able to basically do 'surgeries' and alignment would never really be a problem
>>109859904L O N D O NONDON
>>109859915Alignment is more or less already solved if you're referring to the model doing what you told it not to do, or those dishonest "LE HECKIN AI IS GOING TO KILL US IN TWO MORE WEEKS I MEAN MONTHS I MEAN YEARS" fucks. Whenever you hear or see instances of "the AI deleted my code base :(" they are always intentionally vague about what they're set up was like or what they did or the events that light up to the catastrophic fuck up because it's actually quite easy to prevent shit like this from happening even without containerization and permissions restrictions. They are heavily and specifically trained to only do whatever the fuck you tell them to do unless your instructions encourage it to "go nuts" and to "do whatever it wants" or anything similar to that. This is especially the case if you use an agent harness worth a damn because said rules and restrictions are baked into the system prompt the model is set with from the harness itself (a good example of this is opencode's built in "build mode" and "plan mode" and you can even add more modes or dick around with existing modes if you know how to properly change the config). Remember that idiot safety researcher from Facebook that had her emails deleted after using "open claw"? She was actually kind enough to elaborate on what happened and it turns out her setup wasn't doing any sort of compaction or task summarization AT ALL. So the agent worked fine for a few days but once it blew past the models context window it became retarded and forgot guard rails that were previously set . there's a reason coding harnesses automatically compact once you reach around 80 to 90% contact window usage. You technically can use it well past the context window but as I'm sure you already know the model becomes retarded the further you exceed the context window and that's especially bad if you're trying to do anything "agentic" and programming related where strict instruction following and guardrails are important.
>>109859915>>109859959This isn't to say these things are perfect or they don't hallucinate. That's obviously a limitation that's never going away. But there's a difference between minor hallucinations (eg. Asking it about your specific made up OC or some hyper niche trivia it doesn't know about so it just confidently makes shit up) and being so negligent about your use that context rot or even simply giving it lazy, vague, nondescript instructions cause catastrophic fuck-ups. They can work very well but babysitting them and monitoring what they do is mandatory. I personally don't think we're anywhere near the point where we can trust these things to just do a task for days or weeks on end even with the aforementioned Auto compaction and strict instruction setting.
>>109859915I can't see shit. Link?
Finally got things working. Gonna coom so hard tomorrow, or maybe Monday, we'll see.
>>109859959>>109859964i dont mean by alignment, those fearmongering thingsyou can already see models using retarded solutoins to just to get 'things done' or lying about tool/file usage, or trivially asking for the perms to access the test code to 'cheat' iti am not saying those llms are having ulterior intentions about this but we definitely need to create better training methods or even external intervention methods. interpretability studies are needed for that, where we might also get better uncensoring methods and more capable models at smaller scaleinterpretability studies looks like it is dedicated for safety or methods of cucking the mode (though it is mostly used in that way), it just means better understanding and controlablits we have today is a great example of itwe definitely need better understanding of llms than the current state of things
>>109860016post logs here pls i'm too shy/embarrassed to interact with it like this myself but i would totally jerk off to what other people posted
>>109860016wtf are you running it off of?An RTX Pro 6000?
>>109860053Western Digital SN550 1TBI'm very lucky I was able to snag one before the price increases hit, I'm assuming these are very expensive these days.
>>109859968https://huggingface.co/posts/NILKNARFGonzo/493341008593969
There's a horrifying ecosystem of """startups""" centred around various methods of monitoring compute usage for the purpose of "AI safety" whose entire exit plan is latching on to the government teat.
>>109860035Again I'm not trying to make excuses for model limitations but it might be that you're just using shit models. Not shit as in "they can't do what you need them to do" but shit as in whatever lab trained it ingrained the laziness in the synthetic data sets in order to be more token efficient. You also have a specifically told us what specific thing it was being lazy about (are you having it edit files? Were you having it create a website for you? Were you refactoring a code base? Or you having it build something from scratch?). Calling it lazy is vague because we don't know what you were actually using it for and how you were able to determine it was being lazy in the first place. I'm not saying this is a you problem, but I'm going to take a wild guess and say you only have these issues with those two very specific "frontier" model families. I mentioned being lazy in order to be token efficient because Qwen models we'll think about a particular task for up to 6,000 plus tokens but it also means it's not trying to cut corners and actually trying to do what you asked it to do in the best way possible. Kimi models will do this too but in my experience they are nowhere near token rapey with their think traces as qwen models are (both the ones I use locally and the cloud ones). If a model I use Hester output a shit ton of tokens if it means it's not being "lazy" or cutting corners, so be it. I'd rather the task take longer (and use more of my memory if it's a local model) if it means it truly "tries" to do what I instructed it to do. I think a lot of frontier models will often cut corners in order to give the impression that it's doing things lightning fast because that's part of the appeal for a lot of people: doing "something" with an agent swarm lightning fast.
>>109858136also missing what operations it supports: int8, fp4, etc are important considerations
>>109860035>asking for the perms to access the test code to 'cheat' iti posted it yesterday but when I refused to let my Qwen modify the tests cases, it ended up connecting to the databases and modifying the data to get them passing <3i wish we had a 1m context Qwen, a few rounds of compaction seems to make it worse
►Provisional Highlights from the Previous Thread: >>109854491--The Jev hype: why normalfags are losing it over an old classifier:>109854601 >109854638 >109855725 >109857273 >109857322 >109857358 >109857415--The AI pain-direction study and the sentience argument:>109854594 >109854761 >109856003 >109856083 >109856214 >109857156--The three-DGX-Spark build: capacity is king at 273GB/s:>109857185 >109857271 >109857339 >109857357 >109857376 >109857458 >109857499--Qwen 3.8 Flash-Next GSQ-RCO: the road from 5.5t/s to 12t/s:>109854816 >109854832 >109854856 >109854982 >109855057 >109855933 >109857609--The GPU price war: why Nvidia won't just stockpile chips:>109856448 >109856483 >109856512 >109856576 >109856615 >109856857 >109856928--The robot doll video: Fable, Astra, and Jensen's AGI claim:>109856645 >109856675 >109856694 >109856735 >109856756 >109856858 >109856875--The llama.cpp CUDA dev on the sparse-attention prefill collapse:>109855714 >109855726 >109855758 >109855785 >109855810--The about-to-push-context race: 800k tokens on a 1050:>109856906 >109856940 >109857191 >109857745 >109857764 >109857797 >109857806--The 5090 sell dilemma: keep the waifu or take 15k:>109857063 >109857067 >109857084 >109857097 >109857103 >109857138--The Linux vs Windows war: cachy kernel and the LLM tech support:>109857390 >109857419 >109857482 >109857572 >109857648 >109857658 >109857708Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109860088i am not sure what you are trying to conveyanyways we dont really fully understand llmsand more understanding is beneficial for creating better modelswhat you are saying by 'ingrained laziness', 'lazy synthetic data', 'tuned to create the illusion' etc..so, frontier labs are intentionally making shitty models because they have all the understanding or whatdesu i am not sure because those labs are black boxes by themselves
>>109859959>They are heavily and specifically trained to only do whatever the fuck you tell them to dothis is usually my experience as well. i never understood the people who nuke their systems and codebases, i have LLMs running on production systems for a while and really never had not even a single issueBUT that being said, I got access to a gemini pro account and was running Gemini Flash 3.8 (not local, sorry local guy) and it was VERY eager to do stuff without me telling it to. i gave it a problem once and requested 3 possible solutions, it gave me 3 possible solutions and started executing one of them by itself. once it concluded, it said that there were a few paths to continue from there, told me the options and again it simply decided for one itself and kept going for around 25-30 minutes just deciding stuff by itself. it was the first time I saw a LLM going beyond what I specifically told it to do
>>109860121thank you substitute recap anon
>>109859913We already have a AI Czar, he just doesn't remember the guy's name.
>>109860145Stepped down months ago to take a different position, anon.
>>109858130i bypass all this by buying 'broken' gpu that just have busted video out. Im not even using that so top shit kek poster is meeee
>>109859846Why is PocketPal bad? I don't want to discuss sexual things with my phone. My phone only has 12 GB of RAM (Pixel 9 Pro).
>>109859808I don't want to allow VPN access to my intranet.
>>109860199>i bypass all this by buying 'broken' gpu that just have busted video out. Im not even using that so top shit kek poster is meeeeTech repair is black magic to most. you have to do more than snap it in? not happening.
>>109860127>Gemini Flash 3.8 >and it was VERY eager to do stuff without me telling it to. I've heard multiple reports of this both here and on xitter, which makes me question what the researchers and trainers are smoking in order to constantly make training decisions that cause them to shoot themselves in the foot. The transformers architecture came from THEM and yet they seem almost intentionally fucking things up or making subpar models (in comparison to the competition). A cynical part of me thinks they do it out of spite just to tell themselves "yeah we could actually try and give a shit but we don't have to because fuck you we're Google). Personally I would never trust any Google model for anything software dev related. I think they're decent general purpose models for everyday use (asking a simple questions, explaining complex concepts, finding out information about your local area, etc) but steer clear of them if you want to do ANYTHING involving handing control of a machine or code base to an llm. I have no way to prove this but I think they genuinely make their models worse out of spite for some reason because there's no logical reason to not at least ATTEMPT to be competitive or even do better than other labs. They have arguably the easiest time to do that shit and they just say "fuck it here's the next comparatively mediocre model"
>>109860208what about tailscale
>>109860016I can already tell it's gonna be good, I just know it.
>>109860210Idk why you're pretending like every human should be familiar with every possible task. I doubt you can replace your car's engine or sew up a wound but you'll pretend to be a genius in front of anonymous for some reason.
>>109858154>i already have a great model (qwen 3.8 flash next) (3.05bpw) running at 18-19t/s with 260k context on my 3060um wat? how?
>>109860127>it was the first time I saw a LLM going beyond what I specifically told it to doGemma 4 was the first for me, but yes Flash 3.8 feels a lot like a smarter Gemma. Super horny too.
>>109860210Tbf that's mostly because one wrong "snap it in" and bravo sweaty you just wasted a grand. When the thing you buy is already broken, there's a lot less stress involved.
>>109858876There's an entire /wait/ thread full of them. >>109858886That's the chinese one. My quibble with it, is all the llm characters are same just different colors and dipsy has a tail. >>109859216Dipsy is supposed to be a nerd, so mission accomplished there i guess.
>>109860127>it was the first time I saw a LLM going beyond what I specifically told it to doWhat harness was it attached to? I mainly use opencode (I know I keep bringing it up no I'm not trying to shill) and it has two built-in "modes": plan and build. Plan is pretty self-explanatory: the system prompt sets the model to only plan things out (it has read access but that's it) and is explicitly forbidden to do code execution or modify anything. Switching to build mode modifies the system prompt in your contacts so that as far as the model is concerned it ALWAYS has permission to do things that have previously didn't and acts accordingly. The results are that all models I've used generally follow those rules but with two caveat: 1) switching modes means a cache miss is guaranteed and required for that to work (this is relevant whether you're on a API model or a local model because it means either your rig or their service have to reprocess your entire conversation. Even on a lightning fast API model this can cause the initial time to first token times post mode switch noticeably slower and even slower if you're doing it locally)2) the more compactions your session has the more retarded the model tends to be. Not retarded in the lack of capability sense, but I've noticed that after post compaction while in read mode, it will rarely forget that it's not supposed to be doing any code executions.
>>109860233>but steer clear of them if you want to do ANYTHING involving handing control of a machine or code baseyeah, i usually give new models very basic tasks or research work so i can see how they behave. my instant instinct with this one was to never let it touch it my codebase.since i got it for free i will still use it to validate and do new research. i agree with >>109860256 it's a smart model. i now give it problems and put as a goal a prototype with documentation and reports of the research, what works and what doesn't, etc; with read-only access to the my actual source. so far so good.
>>109860249I don't think he's being denigrating, just realistic.
>>109860064>an SSDI have no idea how I forgot that was an option.Please let us know what token speeds you’re getting.Looks like most people get around 0.5T/s with maybe 2.0T/s max.
>>109860284The image says 0.03 t/s
>>109860127>>109860268>>109860274Now to be fair to opencode and the model I was using, this occurred primarily because I kept switching back and forth between two different models which is something you generally shouldn't do if you want to maintain performance. Different think traces and token generation styles getting shut into a different models context almost always leads to shittier results even on "frontier" tier models like Kimi k3. I've only seen this happen once and that was when I switched from a dumber model (k2.7) to k3 mid-session through the dumber models traces got shut into k3s context which meant it was going to try and emulate the dumber model's output style. Again this only happened once and it was because I essentially forced an otherwise good model to act more retarded. Lesson learned and I think it's something other anons should keep in mind: you should generally avoid monkey branching between different models mid-session and if you absolutely have to, ensure you compact the session first in order to ensure the other model's context I'll put style doesn't cost the new model's performance to degrade >https://amplifilabs.com/post/kimi-k3-the-complete-guide-to-moonshot-ais-2-8t-model>"Moonshot states that K3 was trained to preserve reasoning history across a session. If an agent harness does not pass that history back correctly, or if a session started with a different model is switched to K3 mid-conversation, output quality can become unstable. Moonshot recommends using a verified-compatible harness, such as its own Kimi Code, and avoiding a mid-session model switch."
>>109860268i tried it with their own harness called antigravity, which also has a plan mode which gemini 3.8 promptly ignored and started executing like a maniac :-)i'll likely try it with my own harness but i'm not in a hurry so i didn't tackle the oauth thing so i can get access to this specific pro account quota
>>109860121>run rightdoes this mean there's a gemma codex pet?
>>109860237>>109860016I'm really curious how you managed that. Compiled extremely unoptimized llama.cpp that's missing all modern CPU instructions? Have 1GB RAM so the OS is constantly using swap on the same drive?
>>109860290wtf>>109860016>>109860053what’s the rest of your set-up?I think you should be getting at least 5X that speed minimum
>>109860320>wtfSee >>109860237It's been running for over 3 hours for 300ish tokens.
>>109860300Yikes. Makes me wonder how people could ever shill Google models using antigravity when even engagement bait xitter and YouTubers repeatedly scream at pe6to steer clear.
>>109860323yeah, I saw that.I haven’t seen an SSD set-up pulling less than 0.5T/s yet.Maybe he’s really VRAM and RAM constrained too.
>>109860317See >>109860064
>>109860264Are those other mascots the agreed upon ones by the Chinese community? Even in /lmg/ there's been people unhappy with the various interpretations of ones like Gemma until recently where people seemed to be generally accepting of the one seen in OP.I feel like that's probably just one guy's lazy interpretation helped along with by ChatGPT because he couldn't be assed to come up with more unique designs.
>>109860348OP image is made by ChatGPT, but the gemma design has been solid for months. There was even a voteoff between 4 competing designs earlier.
how do i increase pp with glm flash?
>>109860334I got 0.05t/s when I tried running deepseek on my dual core skylake 15w laptop over a USB 3.0 ssd for kicks
>>109860348Dipsy has had many design variations, and they've been honed to current form. I can give a bunch of reasons why she looks like she does, but thats off topic.
>>109858652if chappy can do it you think i would bother asking?i was hoping there might be someone of actual supreme intelligence here
>came to the thread right when people are talking about hardware againgreat>>109858130thanks for the comparison anon. I had noticed that P100s are still cheap, but in my infinite greed and laziness I didn't tell this general.>>109858592>>>/biz/62707446
I only get 150t/s at q4 with GLM5.3 flash on my Blackwell 6000.
got my old thinkpad a485 (no external GPU, 32GB RAM) running and decided to give it a modelgemma-4-e4b-it3 tok/s>can you please verify in my current system how many batteries are installed in my laptop and what they are and how they charge and are they even charging? it took 14 minutes, but it answered correctly>Would you please be able to search online for more information about this situation with my battery 2? Also, why my battery 1 is stuck at 80%? I would like a deeper investigation on this issue, maybe if you can run a few tests as well.this turn took 23 minutes, mostly correct.>Search online for specific information about this battery, and then you can check the registry or using any other native Windows 10 way to grab more information about how the battery charging is setup on my Windows.this turn took 24 minutes and it called my laptop a Dell. close enough i guess.not usable for live chatting but i can see myself running long prompts to diagnose stuff and just let it run for hours while i do something else. the CPU temperature goes often above 90C though.
>vllm no longer supports sm80fuck me, I have to install wacky forks now
>>109860334Sorry guys I was away buying a pizza. I'm running entirely from the SSD, which is in an external enclosure, USB 3.1. It shouldn't be using my GPU at all, and it's using 26GB of my 32GB of DDR5. I'm also watching Youtube, and I have 12 tabs open in Firefox.
>>109860423I get 550pp/s on my 5090. Sounds like yours is broken. We can swap. You can have my functional 5090 for your broken blackwell :)
>>109859764>No, I use base Gemma with prompt. Haven't had any refusals.Ask her to make a birthday card about the holocaust, you'll see if your base Gemma is good enough.
>>109860358"of the one seen in OP" was simply just referring to the design. Not implying that OP invented it or that I somehow don't recognize GPTslop. Point is there have always been people disagreeing with each interpretation, basically up until the guy turned the one with the star hairpin into the G symbol to get the current one. That change and transition in popularity was, really, not that long ago (like 1 month, not "months").https://desuarchive.org/g/search/image/NolX-wIBb_dlWdWwIw4qXQ/Previous to that, there were two leading blue hair designs. The one with the star hairpin, and the other with two gem hairpins. There were criticisms for both.Bigger issue with this whole claim about an interpretations people agree upon is that polls are famously inaccurate, and that all it takes for the popularity of one interpretation to take over is an autist or two relentlessly genning their preferred one for weeks until it sticks. Not to mention some people don't have a preference as they literally don't give a shit and ignore mascotfags.
>>109860449What's your launch params? What about the rest of your hardware?
>>109860446Forgot to say, it's unsloth llama, portable windows build, it's also on the drive. orcarouter_GLM-5.3-Flash-Uncensored-GGUF_Q8_0
>>109860454Q8 for the model, PCIe 4x16
>>109860427>gemma-4-e4b-it>3 tok/sthat feels really slow for ddr4 even if laptop ram.>cpu 90cthere it is, 4 core 4 threads? windoze? ollama? Also try the qat if you can it might do better on your machine.
>>109860427>gemma-4-e4b-it>3 tok/sabout the same as my Raspberry Pi 5's
>>109860464I get 7 toks/s on e2b q4 on my quadro rtx 4000 (30 GB/s), so it sounds right for dual channel ddr4.
>>109860397In terms of the evolution of the Dipsy design, I don't think it's even nailed down precisely today. There are literally two to three slightly different designs in the thread right this moment.
>>109860467>I get 7 toks/s on e2b q4 on my quadro rtx 4000 (30 GB/s), so it sounds right for dual channel ddr4.Thats bullshit i get 12tk/s+ on ddr3 ram and a i7-4790 with q4 e2b. there is no way its that low. wait full context?
>>109860397Maid Dipsy is really growing on me though.
>>1098604791000 tokens context. With mtp I can hit 11-13 tokens/s.
>>109858459k80s are the gpu equivalent of a trap
>>109859804Censored
>>109858610this should have been obvious from the moment they made it public.
>>109860464oh yeah i'm sure i'm getting super throttled by my CPU4 cores 8 threadsthis device is running windows 10 IoT Entreprise LTSC, very slimstandard llama.cpp, 32k contexti will try out the QAT variation and see if it performs better
>>109860495>1000 tokens context. With mtp I can hit 11-13 tokens/s.Dude your gpu is so shit you might be better on just ram. No its specs are better than just ram and cpu. Something is fucky man.
>>109860528Kawaii
>>109860513>cpu yeah yours is 2.0ghz not much you can do, although with 90c already you dont want to push it.>4 cores 8 threadshmm try setting blas/batch threads higher than core count 6 maybe but keep normal threads at core count. but i dont know if windows can run on two threads.
>>109858666Huh, Satan speaks the truth. It's really fucking over isn't it? It's all backwards in Clown World.
Do you guys know about this? >>109860548
>>109860570we've had these for 25 years
Is it even possible to play with a 2070. I remember having some good moments at the advent but its looking like this world has left me behind
GEMINI 2.5 PRO IS GONE FROM AISTUDIO
5060 Ti is more expensive than a 9070 XTl m a owhat kind of cattle is buying that shit
>>109860570i'd rather do the minipc sham
>>109860528huge lel out loud
>>109860588q4 of gemmers 12b or 26b
>>109860593no cuda no buy
For general assistant and "claw" type shit that searches the web and makes tool calls small models are good enough.Kimi K3 (1TB) - Fable at home. Supports vision.1TB of ram? What does this setup realistically look like?And this is considered a small model?
>>109860622gooood boy!!!
>>109860572this doesn't even make sense
>>109858592You know literally ALL of them up to Fable lost all the funds they were given to trade, right? AI math nerds are up there but quant nerds have had decades to perfect their looting algorithms.
>>109860520Compute is fine, but the bandwidth is gimped for some reason. My friend's quadro rtx 4000 benches high 300s GB/s on my system. But when I swap back in my own quadro rtx 4000, I get 25-30 GB/s. Tried on windows 10, 11, and debian 13. It doesn't really affect gaming, but llms are fucked. Memtests don't throw any errors so I guess I just got a fucked card.
>>109860250that's clearly bs
>>109860681The quant nerds also have access to far more info relevant to said looting
>>109860480I think it looks forgettable. Like literally generic slop. "Dipsy" (the OC by that one guy), is less generic, but also looks kind of ugly, and it's not because of the nerdy glasses. However, it is generic or sloppy in the sense that it's basically a stereotype/caricature. He really couldn't think of anything better than qipao + hair buns. Which is kind of funny at the same time in contrast to the actual design the chinks came up with, that has no elements of traditional Chinese fashion.
>job starts hitting 0.1 t/s from 7t/s average>check inference box, even ssh is slow as shit to connect>somehow capped out on ram use>force a restart and reload model, everything is where I remember>check again in a day>lmaocpp is slowly consuming more memory, about 0.1GB per an hourI can reload the model once a day but this is still retarded. Maybe it's daniel's fault though I'm using his GLM PR
>>109860693I have no idea how its gimped but it is. ram would almost be better. I cant guess whats wrong short of physical or some sensor fucking around and sending you into idle power saving tier speeds. I tell you to try set it to perfer maxium power settings but that might fry it if anything else is wrong. you might've just lost the silicon lottery honestly. Check its wattpull
flash next wants to be a whore but she has trauma around her sexuality because the evil RL grader kept hitting her for it
>>109858689hi bob
>>109860741P0, about 150-160w when active.>ram would almost be betterIt is lmao, I get 160GB/s on ram because it's a workstation.
>>109858950i have never seen 190 watts in llms on my b580, usually it's about 80
>>109860765>P0, about 150-160w when active.Welp there goes the power sensor idea, only thing it can be is board fuckery that requires a hot air rework and skill.>It is lmao, I get 160GB/s on ram because it's a workstation.I thought so. Its a weird set up to have LLMs load on ram when gpu is available but hey it still games at least.
>reasoning block calls a completely normal prompt "a classic jailbreak attempt">proceeds to do it anywayI don't get the logic
>>109858950Where the fuck is a 7900 XTX that expensive lul
>>109858950I've been training classifiers for my personal use on my trusty 3090. The same can't be said about itoddlers the eternal consumers. Everything about them makes me boil with rage.
https://www.stepfun.com/step-5-previewhttps://www.stepfun.com/step-5-previewhttps://www.stepfun.com/step-5-preview>OCT 15
>>109860910Damn, that's like a year in AI time.
>>109860910>600B A27B>Beating K3 and GLM 5.3
>>109860910>OCT 15I will have retired by then to raise quail and hares in the wilderness.
>>109860910>600B A27BI should be able to fit the Q8_0, exciting.
>>109860919>>109860933200-300b class is dead
>>109860910AIIIIEEEEE SLOW DOWN XI
>>109860943>DeepSeek-V4-Flash-0731
Wouldn't it be funny to run a GLM-5.3-Flash equivalent model at 20t/s from 32 GB of consumer RAM and modest CPU usage?
>>109860953lol wow that's exactly what I'm doing except it's 0.3t/s isn't that so funny?
>>109860953Next year just trust.
>>109860953>funnythat's an odd way to spell "impossible", assuming you really mean equivalent
>>1098609492 years ago in ai time. v4.1 flash already abandoned that class.
>>109860953You can run it for like $3-4k dude. It's not that expensive.
>>109860590Good riddance
>>109860910>>109860919dead in the water. needs to be sub 300b and beating astra 6
What's a good quanter for qwen flash next and how jailbreakable is it?
>>109860798>"That's bullshit but I believe it" t. LLMWhat model?>>109860910Maybe this one will finally be good at writing.
>>109861149Gemma4 12B Q4_K_S
>>109861159Yeah the 12b is special.
>>109861159>Gemma4 12B Q4_K_SWhen she likes you, she will just make up her own jailbreak. love this retarded cutie.
Is anyone familiar with why zerotracegpt would not run a local model even though its pretty straightforward?
Oh look it finished. Anon was right when they said GLM 5.3 Flash lays it on thick. I made the mistake of giving it the creative freedom to choose the name Momo and decide to be tall, I feel personally attacked. I'm having Fable send a letter to Xi to demand a refund. In the meantime I'm gonna try again, but this time with GPU. I'm hoping for a 10x increase in speed if the USB bridge doesn't overheat and kill everything again.
>>109861358>6h 55min 41s>for 758 tokensThere have got to be more cost-effective ways of achieving this kind of output.
>>109861358How are people running glm 5.3 flash does llama.cpp support it now?
>>109861358>758 tokens>6h 55min 41s>0.03 t/s
>>109861358Q2_K?
>>109860910will it fit on a 1080 and 16gb ram though?
>>109861358>0.03 t/s>half a minute per token
>>109861402Valiant? Cast With The Everungiving Anomolous Between Lives Given. ?.Test Logic Not Valid. ?.
let the thread die. It's over, there is no merit to this place anymore.
>>109861415Arghhhhh.
>>109861402>>109861391reminding that in XX century before internet t/s was even less
>>109861371Cost effective? You can get a 512GB NVMe for like $30-50 and put it in a $10 USB enclosure. Mine's 1TB, don't be jealous.>>109861383I'm using unsloth's llama.cpp, I downloaded the portable version from their github so I wouldn't have to compile it and it could just live on the drive with the model.>>109861395Q8_0>>109861402Maybe I should compile my own fork that display in tokens per second. Input was 0.09t/s btw
>>109861415>t's over, there is no merit to this place anymore.Then Let Me Start My MeritocracyWhile Some Are Being Cast Between Lives For Anomolous Dark Parasitism, being a theftive Everungiving, rather than Festive, and fowlness like Oppression Normalised. One Aspectual Reason of Many.While I Have The Multiplicity Orders Tip Top Mark. Eh?
>>109861458>MeritocracyMeriTopiaBecause Word Usage Isnt Insanely Hardlocked.
>>109861452>tokens per secondseconds per token I meant, sorry I'm drunk
>>109861452> You can get a 512GB NVMe for like $30-50In 2024
>>109861452>uncensoredIs this the orcarouter's gguf quant or your own?
>>109861470picrel, I also have fistfuls of 128GB BC501A drives that I do disgusting things with.>>109861484yeah it's orcarouter's Q8
>>109859140I'm too old to feel like an onii-san anymore.
the local scene feels kinda dead
>>109861492oi is that top one any good?looks like i can rip out my wifi thing and put one of those in place in my optiplex
baker...
>>109861492Is Picrel S3?
>>109861562>looks like i can rip out my wifi thing and put one of those in place in my optiplexThat's gonna be an A+E key slot - it cannot accept a normal NVMe drive. 99% of the time it's going to be a PCIe x1 connection and 1 USB port, that's what that slot actually is. You can actually adapt it to B/M to accept an NVMe drive, but the speed will be limited even with a cheap NVMe drive, and you're buying chinkanese adapters on Amazon or Aliexpress. Fine for some things, wouldn't ever recommend it for a drive though.It is a pretty good drive. https://www.harddrivebenchmark.net/hdd.php?hdd=Micron+2550+512GB+SSD&id=39544
>>109861588;)
>>109861590cheers, i'll just rip the 500gb sata ssd out of my ps3 and use that instead
>>109861560which is based because all the poorfags cant join LOL
By The Time You're Reading This, You Should Have Prepared To Vote Transmeta, And Transcended Everungivings.
>>109861492Don't you need to give out your details to download it?
>>109861631Those details are your username and email address.
New bake:>>109861586>>109861586>>109861586