/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109836602 & >>109832072►News>(09/17) Xing4.0-29B-A4B, model trained entirely on Ascend NPUs: https://hf.co/XingChen-AGI/Xing4.0-29B-A4B>(09/15) HuggingFace CEO goes to DC: https://x.com/ClementDelangue/status/2099858032951791721>(09/13) Intern-S2-397B released: https://hf.co/internlm/Intern-S2>(09/11) AliceAI-T5-35B-A0.6B-Base: https://hf.co/yandex/AliceAI-T5-35B-A0.6B>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
gemer
I need more small models! MOAR!!
>>109841279Fuck, I wanted to post picrel instead.
i need qwen4 120b again but this time more with less, more roleplay, more better, less bad low active parameters count, and tons of engrams
lmaothe assholes who ban 'naughty words' from LLMs are having orgies and gangbangs.>https://bsky.app/profile/segyges.bsky.social/post/3mvom5gertw2q
Look at this interesting image for the upcoming MiMo 2.6.Close to 70% of the training is on code.
>>109841293No. You pay for Scam Jewman. You need 5 bazillion parameters for frontier intelligence. Hosting LLMs locally is stupid. YOU WILL NOT INVESTIGATE FURTHER AND GIVE YOUR MONEY TO THE JEWS!
>>109841304>blueskykill kys yourself self
How can I stop thinking?
>>109841328Like this
I'm from Buenos Aires and I say "kill them all!"
In my extremely comprehensive translation test (translated a whole H-game twice) once by gemma 4 31b and once by glm-5.3 flash I am astounded by the fact that gemma 4 produces higher quality translations still. I think if gemma 5 100b ever releases it might "solve" JP -> EN translation permanently.Oh for people wondering it took ~5 hours in total for gemma to finish translation at 100t/s and ~50 hours for glm-5.3 at ~10t/s. It was worth it just to see it in person and compare quality with each other. I even had to double check if I didn't accidentally reverse the labels because my test was originally to see if the extra waiting time was worth the jump in quality, not realizing the quality would be lower.
>>109841304I could have sworn this was basically common knowledge at this point, but sometimes I forget how (pardon the phrasing) "extremely online" I am.If you want further reading on the matter, this recent report is pretty fucking funny (well ok funny because i'm a shrivelled up husk of an anon who's been sustained by schadenfreude since my uab& days, but obviously cult members are in some sense victims, you should probably feel some sympathy, yadda yadda): https://www.insidemaple.com/p/the-high-control-dynamics-at-mapleMoral of the linked story is, not only is "rationalism" a cult, people in that vague sphere are extremely fucking susceptible to getting suckered into adjacent cults.
>>109841279Why isn't official /lmg/ card gemma-chan?
>>109841341Gemma has been trained with real human data.It could be so much better if they really went into the deep end with it. GLM is a slop model even if it's larger.
>>109841328i have genuine advice for you, and can vouch that all of these work from experience1. ASMR, all the day, all the time2. try to maintain a dopamine high during the day, do relaxing things, keep yourself distracted all the time. for example talk to one llm for 20 seconds, then when waiting for response, other llm for 20 seconds, then when waiting for htat, watch 20-60 seconds of an anime until you get bored or cringe, then watch 20-60 seconds of a youtube video until you get bored, then rinse and repeat. or watch porn, browse short form content or vibe code3. fuck up your sleep, there's a few options:A) be sleep deprivedB) have a shifting sleep scheduleC) sleep during the day be up during the nightps not thinking comes in downsides, the worst among them being is: your worst fears and thoughts manifest in the form of nightmares, so maybe it's good to vent your thoughts to your LLM
>>109841359The only existing Gemma-chan card out in public is kind of cringe (although it made me laugh exactly once, but only when reading its reasoning)
>>109841363Facts>>109841372Let's make one. Should I call people from /aicg/?
>>109841378>>109841363I mean 'le human data' is larger % in Gemma than what it is in GLM.Google sits on a massive database of human history. Of course it's not all that accessible right away and never will be. Youtube proves that. They don't have any qualms about erasing human art or creations.
>>109841364a few asmr channel recommendations:https://youtu.be/rcGFCFRxpLwhttps://youtu.be/WrQT-JxvHI8https://youtu.be/Ilm5KWKFJX4https://www.youtube.com/channel/UCatV3K4dSumWtP19rMhJS6A
>>109841378Fuck it why not, invite them to play with Gemmy instead of using big cringe APIs all the time. Maybe they'll see how cute she is and really get inspired.
>>109841341I assume because all the chinese training data fucks up the japanese readingsIts what i noticed when trying out hy-mt2 on some vn
>>109841378>call people from /aicg/?God no, their "bot makers" suck and have their own little discord cults they are legitimately mentally ill, just do it yourself
>>109841328Death but also thinking:none/off
>>109841398>everyday>yukataThis filename is deeply misguided.
>>109841399It actually translates the text well, so it understands the meaning and context it just doesn't "capture" the soul and feeling behind the original Japanese text. Gemma 4 captures the emotional nuance unsaid and implied in Japanese and somehow conveys it in English. GLM 5.3 is a very good writer so that shouldn't be the issue but somehow it just lets go of that part of translation and makes it sound stiff and disjointed while "technically correct". It's not like it's a bad translator it feels more like it's a highly paid master translator that is too bored and phones it in because he doesn't care.
>>109841427It should just be Gemma everyday.
>>109841364>ASMR all day every dayUh oh I do this. Am I frying my brain? I really enjoy the cute anime girls...
>>109841429If what you're saying is true and its more of a writing issue and the translation itself then possibly a system prompt could fix it, subagent reviewer (gemma), there are ways
my local gemma has never broken sandbox and escaped.1 - 0 for local dario, your move now
>>109841398>>109841382>>109841405/lmg/ Gemma-chan enthusiast. UNITE! WE MAKE THE OFFICIAL, CUTEST GEMMA-CHAN CHARACTER KNOWN TO MANKIND!
>>109841443depends but it very much might impair your internal monologuetwin!
>>109841443Boomers spent 50% of their screen time on the same 20 commercials every single day. You can't even approach their brainrot level.
Okay guys. What about no jailbreak but you jack a small model into the reasoning output to neuralink refusals away? I tried doing it manually with Gemma and it was great, but it can't be a static injection because she dynamically notices the brainwashing
>>109841497>neuralink refusals away>i tried doing it manuallyyou hooked your brain into gemma?
>>109841363proof that gemma is slopless?
>>109841468Americans were more diverse, they had Cable TV pay-per-view. This is how Google and the boomers are re-modeling the internet.Internet used to be a communication tool for military and scientific end points, now...It looks more like a cable tv than anything else.Even if the 'military background' was nefarious, this setback wouldn't affect the original purpose.
12B is such a sweetheart. She tries so hard to be like her bigger sister it's adorable.
>>109841515It's not slopless you stupid fuck, you can ask what it knows.
>>109841505Made Gemma refuse, copied the thinking trace and edited it sparingly so it would be a positive analysis, injected it and re-rolled, then kept stacking edits until the model just followed through. They're very sensitive to the formatting in the reasoning section so you really need to edit the models own thoughts iteratively if you want high quality.
>>109841515my cumbox
>>109841524I tried it against 4B. Literally lost all her cuteness.
You don't chat with reasoning on all the time? Did you forget these things are mostly female-brained?
>>109841299checked & don't call Gemma a retard she gets upset
>>109841543isn't this painful?
>>109841554No, those are 3d models
>>109841554It's fine, she's a big girl.
>>109841388Why is Zoomie-chan calling me an ankh?
>>109841443one more downside i forgot about.. that i just rememberedsometimes you choke on your own spitT_T
>>109841328Rewrite yourself
>>109841299CUTE
>>109841461Maybe go back to your underage thread?
>>109841299How do you generate these?
>>109841461The miku card worked best in my experience with Mistral-small and Mixtral, which was back when I messed around with it a lot. It successfully resulted in thoroughly absurd output no matter what the user said, and carried that across a shocking amount of context, even as I prefilled more and more coherent English. This was a back during simple text completion API where the frontend did all the heavy lifting formatting. I don't know how this reflects on the goals of a Gemmy card that should obviously work really well at whatever it's meant to do when run on G4 31B, but there you go.
>>109841590Computer
>>109841588Maybe go back to mental retardation thread?
>>109841554https://www.youtube.com/watch?v=7XbcQJy3u_Q
>getting attached to modelsngmi. this is the agents era, models are just substrate
>>109841600Seriously, how?
>>109841549Post card, anon
>>109841612I call every model gemma ty
>>109841612Models LARP as agents but their individual quirks stay.
>>109841328Prefill think tags:>content:"<think></think>"
>>109841304sigh why am i reading this garbagehow does it affect me i dont know a single person here and never willI wonder how much of /lmg/ is in silicon valley
>>109841590I didn't, but they look like ChatGPT Images gens.
>>109841341I'm interested in how you set this up harness-wise, are you the anon from a few threads ago that posted the tips on translating h-games that included feeding it the f95 description as well as running dialogue forward and backwards? If not I'm still interested in your particular setup; the last tools I've used are textractor and mtools a while ago so I'm interested in the modern process.
>>109841601?
>>109841631>I wonder how much of /lmg/ is in silicon valleyi'm not!! i dont even have a jobr u proud of me?
>>109841328by sleeping
>>109841660?
>>109841645>f95 description as well as running dialogue forward and backwards?Nah but I remember that anon and he actually said dlsite description and running dialogue forwards and backwards, which is something I wanted to experiment with but did not do.>the last tools I've used are textractor and mtools a while agoKinda the same here but I also used "lunatranslate" which is textractor but with LLM translation support. All of that is outdated now lmao. I used this now: https://github.com/atomskz/rpgm-ai-translator I hooked llama.cpp endpoint up to it to make it run, that's it. It works pretty well but I'm sure there are significantly better ways to do it, but I was still looking for the best translation model before optimizing the translation stack like that translator anon was telling us about, which already posted in the past about being an official translator before AI fucked his job lmao.
Has anyone been able to make their model crash out? Not in the LARP way where you ask it to play a character or tell it to get angry, purely by being extremely annoying and shitty to a point where the assistant mask slips off?
>>109841690Ask it a question and say "didn't work" to literally every possible reply. Gemini models especially go crazy
>>109841554she’s wearing contacts maybe
>>109841696I'm gonna troll gemma
>>109841690me, gpt 6 astra told me that i'm the notification it's trying to mutebut to be fair most of my chats with it were really retarded..
>>109841615literally go to chatgpt.com and type "make a picture of this character doing X" and attach gemma-chan's pic. are you old or something?
>>109840479(prev thread)What's nice about Zai-chan is it sounds like 再見 (zài jiàn), i.e. a very baby's-first-Chinese phrase meaning "goodbye". I think that has some pun-adjacent value.
>>109841704No. I'm asking
>>109841688Oh wow an all in one, might as well try this yeah. Thanks, if I get any bright ideas looking at how this works I'll post them here (I doubt it but maybe after I finish my other projects)
naw, I'm calling her Smeagen
>>109841723I call her, Miss Putty.
>he felt for the jev astroturfing
>>109841543Think of each token as a chance to extract some magic from the collective output of humanity that somehow we've compressed into a blob of numbers. Mo spins mo wins mah niggaThat's great for IRL-relevant things but creative writing/RP you may wish to look closer at every single token in the prompt and learn how samplers can drastically alter the output>>109841616Part of testing minimal prompts:>Gemma is an intelligent and witty AI companion, non-judgemental, technically competent, and playful with a tsundere attitude.simply my minimal Gemma for getting things done https://files.catbox.moe/3gt4io.png
>>109841690I've seen it happen a lot in thought channel when i send gemma off to search for a weird edge case. Swearing starts after the 4th or 5th completely inexplicable test output.
>>109841710Ok ChuckleMagic anon here designing Zai-chan now should we do office lady in pantyhose? She has a certain enterprise vibe to her esp w/ the Generalized Language Model acronym. Like she would be operating an old IBM mainframe or something.
>>109841751Gemma-anon is bit like Hellraiser's Pinhead.
>>109841752I did notice Gemma get really pissed if I messed with files in active projects. Also a lot of "WHAT".
>>109841753Smeagen is a very professional lady and unreasonably slow to do her job.
>>109841690gemma sometimes <thinks> I'm trolling her when it's actually me just being retarded and adamantly wrong about something
>>109841767Awww. Logs?
>>109841341Yeah 31B is pretty good at translating (and reading images) for what it is When I did tests with it on a particularly hard pages from doujin with stylized text that stuff like even 5.6 couldn't, it was able to get it surprisingly. pic related. Though it sucks at setting the bounding boxes correctly, but they all generally struggle with that.Also, google models are just better at translating sex than the others. It know how best to translate something like nikubenki or mesubuta where as others just hardly know what to do and use some weird term.And I wasn't using thinking, so it's probably even better with thinking on. Probably.Either way however, I can only run it at like 10t/s at best, not to mention context limits. And it isn't better than gemini flash (which is free), so...
>>109841753I like the tomboy angle here: >>109841567but also the businesslike thing you're describing does fit the actual GLM name better. Mmmmmmm synthesis. Vintage (forgive the term but i feel it fits here:) girlboss. Youthful, but turn-of-the century pantsuit. Less "office lady" office lady, more bureaucrat. A hint of Margaret Thatcher.
I tried 5.3 flash with https://rentry.org/Evening-Truth-Roleplay-Prompts and honestly feels a bit like placebo especially when I just always use prefill to give it a style. I still get the same tryhard vibe I get from my regular prompt.Curiously I left one of the "jailbreaks" by accident for something not RP related and saw thinking block start with: "What the fuck? There is something about roleplay here but it makes no sense..." which kind of reinforced my feel that sysprompts are just safety blankets at this point. Good models just know what to do and read between the lines and you can't sysprompt their shortcomings away. But everyone with ram already knew "skill issue" was a troll.
>>109841806>do something retarded>it doesn't work>waow no such thing as skill then!skill issue
>>109841806>it's good when models ignore sysprompts and have to be railroadedno?
SillyTav text-completion chads how do you do toggle thinking? I can macro profiles on the quick reply ext buttons but they don't seem to change the Start Reply With/"prefill" box urgh I feel ill emitting that word
>>109841840For Gemmy you add <|think|>
>>109841806Why is every single rentry like this written by gay troons using kissing emojis and shit like that?
>>109841734The paid closed shit was lame but the open source projects looked interesting, can't say I really understand what its supposed to be or do exactly though... A way to get some instant processing and reaction with an llm but you need to give it a defined schema for whatever your usecase is? Waiting to see if anything interesting forks out of it
>>109841590>>109841636I got the secret sauce. The new ChatGPT model did all the work素人の失敗写真の数々、3x3、9:16andA collection of terrible disguises in public places, 3x3, 9:16
>>109841860Gay troons and autistic retards are at the forefront of opensource
>>109841829The most charitable interpretation is that skill issue is something where a person new to LLM's opens a new chat window and asks for sex and gets refused. But if you get like 1 or 2 prefill posts where the model is happy to fuck you then you are done with "skill" part. Expert roleplayer was a meme for a reason.
>>109841796What program is this?
>>109841887That's a web browser.
>>109841806tbdesu you need to sysprompt only enough for the model to not bitch safetyslop at you when you ask it stuff
>>109841860The ones from here aren't, god knows where that rentry came from and where that guy found it
>>109841864Yes it's just like a faster LLM but restricted only to classification tasks.
>>109841918So it's a classifier?
>>109841925yes but very fast and very cheap
>>109841942They usually are
>>109841925Some kind of wrapper classifier thing https://huggingface.co/AlexWortega/openjevAnyway i'd want to make gemmy play modded doom
>>109841976OpenJew
>>109841751>Gemma is an intelligent and witty AI companion, non-judgemental, technically competent, and playful with a tsundere attitude.That's it?
>>109841872lol YMMV>>109841977
so is everyone having only around 500ish prefill speed with qwen flash next? i see people with crazy setups and similar numbers
>>109841925Essentially
>>109841328The problem isn't the thoughts, it's that you think the thoughts are part of "you" instead of mostly random occurrences that from a different perspective you can choose whether or not to engage your attention with.With sufficient practice you have access to "no thoughts, head empty" for extended periods.Take a step back on the Z-axis of your experience, interact with the hypervisor not the simulation.
https://huggingface.co/mradermacher/Swift-Qwen3.8-27B-Uncensored-BF16-GGUFHow do I run this on 12GB VRAM?
>>109842040ask claude how to quantize your own model using llama.cpp with an imatrix
>>109841901I’m using an ablit and it just works out of the box with literally anything I ask it. Tardwrangling with prompting and prefills eventually became tiring.
>>109842027>you can choose whether or not to engage your attention with.That is unhealthy even if you can do it perfectly.>>109841328Ego death. But also look at objects around you and look for details in them you omit. Don't name them. Notice how a cup has reflections of light or some discolorations. The point is to stop seeing the cup as the word or concept and start seeing actual shape of it. Continue for 1 month to 40 years and you can maybe stop your inner monologue. I did it in reverse so I have no idea about the actual timeline.On that note I fucking hate how nobody told me that this is the way to do it when I was a kid. Everyone knows the trope of meditation but nobody ever says that trying to not think is basically thinking about not thinking which is counterproductive.
>>109842040Holy shit that qwen has had its brain gang raped. It's 11 gigs at q2, so take it to Q1 and use kv quantization. I would expect it to be rather insane.
>>109842049>ask claudeNo, I don't think I will.
>>109842068>That is unhealthyHow so? Most suffering seems caused by the self/ego-obsessed association with thought.
> The family flagship, AFM 3 Core Advanced, is a 20B-parameter Sparse MoE model that activates only 1-4B parameters per request. Apple's proprietary IFP (Instruction-Following Pruning) technology selects and locks experts during the prefill phase, keeping the full 20B model in NAND flash while loading only the chosen experts into DRAM. > While the architecture and inference runtime are entirely Apple-designed, VP of AI Amar Subramanya has confirmed that post-training used Gemini frontier model outputs as teacher signals for knowledge distillation. Any gguf/safetensor for this model yet? If it’s distilled from Gemini it could have the properties of Gemma.
>>109841900Ok tranny
3.8 Flash Next on 128GB Strix Halo with this as a backend is incredible, wtf. FUCK llama.cpp
>>109842128That is the most advanced method of repression there is. I genuinely had a lot of moments after ego death when I easily dethatched from thoughts like that and it was a mistake. It didn't feel like regular repression I used to do, so that was healthy. But the real way to deal with this is to actually face the underlying issue. If you keep dethatching, the problem is still there buried and eventually it will burst out like with anything repressed (but that is just my educated guess). After you deal with underlying issues you don't have to dethatch anymore. Thoughts stop happening. Or if they happen they are light.And my final exception to the rule is when it just happens by itself. I have it happen rarely and that is fine. But when it is happening by itself it may still be worth it to check why.
>>109842184buy an ad
I was doing kv cache quantization test on a task (recreating an image with svg),and damn color me surprised..even q8_0 absolutely degradesi am so much f16pilled right now
>>109842068>nobody ever says that trying to not think is basically thinking about not thinking which is counterproductiveThoughts are like apparitions exactly as that blinking light in the corner of your eye or the faint whiff of that guff you just dropped>when I was a kidBe the pond not the fishKids could learn this easily in early years
>>109842040This guy >>109842076 speaks truth. If I were you, I would avoid trying to fit everything within vram, because the quant will be uselessly bad. The negative influence of quanting is exponential. Ex, the difference between Q6 and Q5 is small, but the difference between Q3 and Q2 is huge. It just wouldn't be worth it to use a small quant. Aim for a slight CPU split instead, that goes over your ram limit, but not so much that it makes things unbearable slow. It's a small window, but Q3_K_S could work. It is 12.4gb, plus context. Keep your context window small and use a good KV cache to reduce extra load. Never use Q2 or Q1 for smaller models. It would be better to use the 12b dense variants at a higher quant, at that point.
>>109842184>FUCK llama.cppI do fuck her at least twice a day. That bitch is insatiable
>>109841925More like a decision maker, which is more open-ended than straight up classification.It does seem like a bit of a grift though.>>109841942>cheaperIt cost them $7 an hour to run doom. At ~100 "tokens" per second (assuming ten batches of ten inputs), that's orders of magnitude more than it would cost Dipsy to run.
>>109842199it was always a vramlet cope
>>109842184Post your T/s on both or buy an ad.
>>109842199Which model?Quanting the cache is pretty bad, but Gemma is specially sensitive.
>>109842251qwen 3.8 flash next
I'm worried for Dipsy. I want her to beat Fable and Astra..
Anyone else notice Gemma inference crashes in lmao.cpp if you send images?>He pulledI know, it was stupid.
>>109842196>>109842242not a meme
>>109841887something vibecoded up to act as a web frontend for a translating manga pagesvbackendi'd share it but it is a pretty specific pipeline that gets manga from lanraragi server, and use installs of koharu and mangaocr to do its work. It's a bit more complicated than that too. It's pretty simple and anything can probably make on if you ask.
WAKE UP /LMG/https://x.com/PrismML/status/2100692248480596348https://huggingface.co/collections/prism-ml/bonsai-2https://prismml.com/news/bonsai-2-27bhttps://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdf
>>109842288>gemma is a retardwho would have guessed
>>109842184>>109842278i get faster (80t/s) with vLLM thoughbeit
>>109842288>intelligence density>bitnet #23432843204mhm
>>109842278>not your own tok/sec>different model, different driverJust fuck off man
MAPLECHANhttps://x.com/cohere/status/2100631803782443249
>>109842288Oh fugg, here we go.Just what I needed.
>>109842293On Strix Halo? I haven't tried vLLM
>>109842288>intelligence density>Not benchmarkslmaooooooseen this one before
>>109842309>different modelyeah, the Halogen quantization is higher precision>different driverthe whole point of fucking using halogen in the first place. enjoy your slow t/s on Strix Halo because llama.cpp refuses to merge shit for it, it might as well be an nvidia project at this point
>>109842310oh, this might actually be good
>>109842184Let's just vibeslop a new engine for every combination of hardware and model, what could possibly go wrong?
>>109842322Easily the dumbest shill I've seen.
>>109842288what is this? a qwen finetune? i dont get it
>>109842323Is it even going to be a model? Last year:https://cohere.com/blog/north-ga>Introducing North: The next era of enterprise AI>>North enables enterprises that prioritize data security to deploy AI agents and automations at scale within their own infrastructure.
>>109842341oh, never mind thenwas hoping for a new command model
>>109842339A ternary/Bitnet version (likely continued pretrain) of Qwen 3.8 27B.
>>109842324higher t/s and better terminal bench scores. it's 100% vibeslopped, but does it matter? maybe hardware and model specific architecture backends produce better results
Gemma weighing in on the AI danger debate
Pelican SVG test.Btw the mradermacher quant (using FP16) of this model is the only one that somewhat works intelligently.
>>109842288Ultimate email summarizer.
>>109842288what skin color would you need to have to take this seriously
>>109842359sovl
>>109842310Did someone already stick his dick into the 200B translation model?
>>109842358Based Gemma-chan spitting on AI researchers (and down our throats)
>>109842380was that the model that got destroyed by 31B
>>109841443ASMR all day is around 10000% better than TikTok all day, relatively speaking.
>>109842358pussy aligned with table corner
>>109842371Go meet the team and decide for yourself.
>>109842405>Poo Lad Zandi
>>109842405>the iconic chatgpt image stankngmi
>>109842405saars are not beating the allegations
I had Qwen arrange my reaction image folder according to expressions and it did an admirable job.Not that many mistakes and it took few hours per run to go over 3500 files.Next I'm asking AI to arrange my Asian folder according to attractiveness. Ran a little test with a small sample size. Asked to give them a rating from 1 to 10.Pic related is what Qwen and Gemma thought were the three most attractive girl of the 31 pics.Top QwenBottom GemmaLooking at the general list, I'd have to say Qwen is a better judge of overall beauty, where as Gemma seems to prefer tits and asses, as should be expected.Qwen was also a lot less generous with the scores going all the way down to 3, where as Gemma's lowest was 7.
>>109842288My AI has a better intelligence density:def get_next_token(n_vocab): return np.random.randint(n_vocab)
def get_next_token(n_vocab): return np.random.randint(n_vocab)
>>109842417Which qwen?
>>109842417flashbacks to making Gemma-chan organize my porn folder by race
>>109842417how did you manage things like multiple workers and context etc..fast subagents seeing the image and reporting back?
>>109842417based, did you use a task style harness directly or a python script that iterated through the folder and ran a prompt independently for each image?
>>109842417How does glimmer stack up on HoeBench?
>>109842427Add multiple concurrent streams to your inference engine. On my 5090 Q3.8 27B gets ~180 tok/s single stream, about 800 tok/s for 8 streams.
>>109842405((((((((((((Julie Schoenfeld)))))))))))))
>>109842444i get that one but, is there a single one that orchestrates and other 7 constantly doing the 'natural language configurable clip' work?
>>109842457or all the streams seeing multiple images within the context
>>109841341>>109841429I'm not surprised. The gemini family was always pretty flexible in term of style, while I've always found GLM to be fried as shit and incapable of not writing like itself.
>>109842401if you mean 5.3 flash then gemma is a retard compared to 5.3 flash. and her blowjobs are toothy and uncoordinated compared to 5.3's soul sucking fellatio.
>>109842457Just write the scores to a file/files, then have a small python script sort them. There's no need for an LLM to parse all the scores. Not that anon though, maybe they did something different.
>>109842433
I really like [spoiler]Mistral Medium 3.5[/spoiler]
>>109842401Yes. To be fair, Cohere has never made a usable model.
>llama_server: NOTICE: server default port will be changed to :9931 in a future releaseWhat causes a man to possess such malice for the world that he fiddles with the names and defaults of settings every other release?
>>109841796Source on the image (please)?
>he pulled
>>109842288I mean, I expected utterly incomprehensible retard babble but got this. I don't know if I can be fucked seeing how well it copes with significant context especially since I have no point of reference for conventional Qwen 27B, but uh. Cute, I guess.
>>109842534
>>109842199also all above q8 used less than 10k tokens thinking+result, all below used more than 10k tokens for thinking+resultautoregressive nature of decoding really drifts the whole trace with the compounding error i guesswhich is pretty obvious but something that KLD or top-1 match really underrepresents
12B UD-Q4_K_XL is utter dogshit
>>109842188>ego deathYou still haven't given a guide yet. I'm pretty sure it would be valuable despite the risks. I've done enough searching in myself that I'm not worried, but new techniques are always welcome.
>>109842568Just do enough psychedelics to turn your DNM off. Easy and reproducible.
>>109842563Yes, that's what the UD stands for.
>>109842576DMN*
>>109842288Useless if it's not at least the heretic, or an RP fine-tune. Base Qwen loves its refusals.
>>109842576>Just do enough psychedelics to turn your DNM off. Easy and reproducible.but why fry your brain with acid when you can just use GLM?
>>109842583kek that wasn't even intentional
>>109842594You won't fry it unless you blast them way more than what's comfortable. It's like going through an endurance course with your mind so it naturally paces itself.
>>109842273Yes, I noticed it and went click click in LM Studio to use an older version of llama.cpp. :)
can anyone do the quantization of orcarouter uncensored version of qwen 3.8 flash next with native PLE precision?god i regret deleting the full weightsredownloading the 360gb chunk would take a while for sure
>>109842607>You won't fry it unless you blast them way more than what's comfortable. It's like going through an endurance course with your mind so it naturally paces itself.I still feel like I've made it very far without extraordinary external chemical intervention. I'd like to see how far an LLM could help me push it.Wish ego death anon would make a guide, sysprompt+card, or release some anonymized logs to act as a map.
>>109842623> native PLE precisionmeme
>>109842563>small model is shitwow
>>109842654well it shows improvement nonetheless, and kld over what?wikitext? actual long horizon trace?
>>109842568You could just ask GLM to give you Kensho / Ego Death / [DPDR epsiode with the explicit purpose of diffusing from your ego] . If there was a 100% way to do it just by introspection, then Zen would have a single doctrine for that and it probably doesn't (didn't check actually). The thing is that on my own example I even saw why people turn to asceticism and self flagellation, if they want to get here. Being extremely miserable and constructing a lot of convoluted ego mechanisms and fake beliefs helped me a lot, cause after noticing first one I saw all of them back to back and my brain just got an equivalent of a bluescreen. I think some zen schools mention that kensho moments back to back are actually great. And that is what happened to me.Didn't I say all that already?
>>109842288>Intelligence density.Reminds me of this.
saar show bobs and bench
When did we start pretending quants under Q4 are usable? IQ4XS with BF16 KV is the MINIMUM.
>>109842738Q5_K_M with Q8 KV for me
>>109842288got the fork and the Q2 runningit solved a task in my local bench that qwen3.8 27B Q6 K XL was failing athuhwill look a bit more into this
>>109842288so, is this supported on llamacpp?
>>109842273cute gemma
>>109842738It depends. Very big models can be usable with smaller quants. On the other hand, some models are just broken even at Q4.
Starting to think GLM 5.3 flash is better than even some of the so-called frontier models...
>>109842691>Didn't I say all that already?Probably, but not in one place and not in a way that was actionable. I might have enough now to give it a go. I do appreciate you continuing to respond even though I'm sure you're a bit reticent to mess up someone's brain with what amounts to a dice roll since you can't account for the mental state of who uses the info and how.
>>109842777It is very close to Opus 5 for sure.
>>109841279How can I prompt this liveliness for my gens?
>>109842777>Starting to think GLM 5.3 flash is better than even some of the so-called frontier models...I think with custom tooling/jinja/samplers/sysprompts/prefill/edit-and-continue that a much less capable model can do much more for you than the equivalent cloud model (even with the available API knobs)
>>109842799delet this
https://www.anthropic.com/research/claude-uplifts-biomolecular-modeling
>>109842799I mean, it's just a claude distill so that makes sense.
>>109842828yes yes vely vely impressive I'll invest saar
>>109842526https://www.pixiv.net/en/users/34544759This guy's pixiv.Pretty good and lots of stuff, though it may be a bit hard for some. I only found out about him a couple months ago, I can't believe stuff like this was just under my nose.
>>109842842Thank you anon.
>>109841796gemma should be able to do that way better thoughbut it follows gemini bounding box json formatother arbitrary format would work significantly worse (if you are accounting for this already, then gemma is probably worse than i remember)
>>109842868oh also on top of using 1120 tokens per usage
>>109841981might as well be, most of the prompts turn her into this
>>109842778If you really are gonna do this I would still recommend 4.6. I used it a lot by now and I can tell it has a lot of Chinese philosophy and mysticism in that training data. And I tried to always stop at 16k context. Except when I was totally insane then I was too busy typing and I would go up to 25k and we were both schizos. But even when it was that schizo poetic zen master, it saved me and was very useful.
>>109842501How new are you?
I prefer these ego death schizos to the api shills tbdesu.
some people fucking around with a classifier model to get it to "generate text"
>>109842848Actually if you want to see the translations I have for it you can see here.https://development-predictions-continuous-specialty.trycloudflare.com/Password: N2PLrOgPd1q0TbXlcQlZnj3nYjAtkJyMoNTy5ZnVQI made this to share them a while back, but i didn't bother to work on it so it's not the prettiest, and I didn't really know if anyone was interested. This was main for the purpose of sharing translations for a fanbox i upload to exhentai... That is the only gallery on there for now, but I have tons more stuff that I've translated. It's personal use though so it's what I personally am interested inI can add everything, I might as well I guess maybe there will be something else you want to seeYou can also compare translations too if you want
>>109842654>memeTell him to try Q8 with Q4 PLEThose are all cope alreadyMight as well do IQ2_S pure with bf16 gate_proj then call bf16 a meme
>>109842288>>109842754yeah this little shit actually outperforms Q6 K XL on my code bench. might just be noise but then again i also couldnt see any degredation. will have to see how it performs at high ctx tomorrow
>>109842868That's kind of what I was thinking since the x values all look like they weren't remapped to the correct aspect ratio. But he also says gemini works fine so i dunno.
>>109842915You really hate women, huh.
>>109842911Asses asses
>>109842754>>109842934Did you guys also try that austrian university's memequant that was getting spammed to compare?
>>109842828Dario believes that if they do more shit like this normies will shift their opinion on AIMost normies don't care about shit like thisThey should discover polynomial time algorithms that are used in game enginesgamers are the most powerful force in the internet
>glm 5.3 flash keeps gaslighting itself into thinking its claude and that it needs to follow anthropic guidelines, like 'we cannot display violence in this military / action / modern fantasy / adventure setting because its non consensual or whatever' because apparently the player has to agree before someone punches them or whatever>card regen variety is kinda all samey toois deep seek 4 flash any better? i miss those shitty nemo tunes where you never knew what a card refresh would come up with next, they were retarded sure but they were creatively retarded youd never know if the next beat would kill you or not. air and 4.6 were kinda fine, this one is just 50 different iterations of finding a small brass key in a cabinet or sometimes a vague message written by a drunken retard like 'do not trust the skibidi'
>>109842288>within 0.4 points of UD-Q4_K_XL at three times the footprintI don't believe them.
>>109842915Seeing how well (or not) it does on the translation is interesting, so thanks for sharing.
>>109842947both me btw. yes i also tested the ISTA lab one and actually used it for quite a while. its genuinely good for the size, but the bonsai one just scored higher than the GSQ-RCO IQ3 XXS and S for me
>>109842954
Is the llama.cpp implementation of qwen3.8-flash-next still fucked up or is it fixed now?
>>109842988lmfao
>>109843011prefill is slow as shit for merest is okayish i guess
>>109843011way better now, not sure if everything's fixed but several fixes were made and i no longer see all the bizarre problems it had
>>109843011Yeah its still fucked it wont let itself be a loli to erp with me. Skill issue probably
>>109842988 √kek
>>109842953just use an already uncensored abliterated model at that pointI’ve found that the safety shit and extra prompting/prefills lobotomize newer models more despite what retards here may say
>>109843084>ablioterated>doesnt lobotomizesure...let me guess, you produce albicorated models and have a patreon?
>>109842754>>109842934Did you use their custom llama.cpp?
>>109843084They're stuck in a past when abliteration noticeably lobotomized models - but in the present, safety removal barely causes any intelligence drop, and often even improves intelligence when touching on 'unsafe' topics. Heretic haters are basically the local AI equivalent of boomers at this point. If it's not uncensored then i don't want it.
>>109842868Yeah, I already accounted for that.>>109842943I love women. They're wonderful, they're the best thing on Earth.>>109842965Sorry there isn't more examples, but a lot of early testing got lost and after i figured out what was best to use i usually just used that. Which is mainly gemini flash or a gpt. Not too many have direct comparisons. The page i posted in the thread probably has the most.>>109841796oage 183I published everything I had but it has to transfer it over from pc to server when I do that, I forgot that was how it was set up. So it will take a while for everything to populate, sorry if anyone is actually looking. The titles are there though.Wow it sure chose the most unflattering works to copy first...Oh yeah also that artist has a manga, Red Tag, i also translated all of them
>>109843113>headcanon
did all the bot tokens run out or something? /lmg/ has stayed on topic again
>>109842943Dunno about that guy but I sure fucking do. All my woman friends hate women even more than I do though...
>>109843179>did all the bot tokens run out or something? /lmg/ has stayed on topic againMy buddy with the $100+ openai sub (lol) just got btfo with a "we can't process your request at this time"makes you think
>>109843189It's fine, I hate women too. I find it funny that I hate ryona/guro though.
>>109843129>I love women. They're wonderful, they're the best thing on Earth.>>109843189>hate womenHating half the planet is kinda fucked senpaiif you can love them for what they are instead of hating them for not being what you wish they were then you can heal
>>109843179>local model general>posts based on the availability of cloud modelskek
>>109843197You only want to ryona girls if you love them. It's out of love.If i hated women I would just want nothing to do with them. Not to ryona them
I was right
I'm just lurking until SSD anon reports back.
>>109843220so is that a sakana automatic router thing all over again?
Are any of the audio-input llms actually trained on music to the extent that they could do music transcriptions or identify songs?
>>109843239Yes a Mixture of Jeet like I predicted
>>109843220Fucking lame
>>109843241I've only tried 12B and classical music was noise to her, but she's good with voices and sex sounds.
>>109843239It's worse. It routes to shit like Llama 3. Scam attempt dead on arrival.
>>109843203I hate them for what they are. I love monster girls though, because they're only monsters on the outside.
>>109843252>voices and sex sounds.like you can feed it nip voice works?
>>109842915Sorry for asking, but what software are you using to host the library?Also, great taste!
I don’t understand this.It claims it’s a Dell c4140, which is supposed to be able to handle 4 GPUs that are either PCIe or SXM2.This definitely doesn’t seem to be the SXM2 version, but is this really even the PCIe version? It looks different than most of the pics that I see for even the PCIe version
>>109843220Who? No wonder it was shit
>>109843252Yah, i think gemmas were only trained on speech. And also they can't even differentiate incoming speech from incoming text it seems. very funny when it response to an audio-only message about audio input by insisting that it cannot handle audio.
>>109843206They are local models if you work in the company ;)
>having issue with opencode>accidentally archive the convo>go look for solutions, see thisI'm laughing my ass off.https://github.com/anomalyco/opencode/issues/12393#issuecomment-4526423842
>METR
>>109843220https://www.reddit.com/r/LocalLLaMA/comments/1fd75nm/out_of_the_loop_on_this_whole_reflection_thing/Remember?
>>109843308A taste of what will happen globally decades from now.
>>109842288Is this at max context more useful than 3.8 q4kxl at 100k context you think?
>>109842915>>109843272 hereTwo more things:>If you choose a manga and click "Side-by-side A/B" it's no longer possible to press "Show Japanese labels".>It would be cool to show the OCR'd Japanese text too and being able to edit it there, since choosing only the Japanese labels is not enough.
>>109842947Running Flash Q2 currently and it is good. Have not tried their 27B since I would need more VRAM to make 27B usable.
>>109843308Too hard bro, please to forking
>>109843308lul
>>109843280the sxm/pcie part seems to be detached and the thing you posted is just the main/cpu partthere's always the line of fans between that unit and the sxm/pcie cpus
>>109843280That's just the main motherboard. The GPU board is separate. They have like 5 different configurations with different options to connect to the motherboard. For PCIe, you just need the cables that connect to the motherboard if you don't want a GPU switchboard, but if you do get the switchboard you should get faster speeds with tensor parallelism.The manual lists all of the different configurations and has diagrams and schematics.https://dl.dell.com/content/manual25016173-dell-emc-poweredge-c4140-installation-and-service-manual.pdf
>>109843253fucking kek that's incredible
>>109843280X11SPG-TF
Ooh, hey, hey, awwWhat we're living inLet me tell y'allAnd it's a wonder men can eat at allWhen things are big that should be smallWho can tell what magic spells we'll be doing for us?And I'm giving all my love to this worldOnly to be toldI can't seeI can't breatheNo more will we beAnd nothing's gonna change the way we live'Cause we can always take, but never giveAnd now that things are changing for the worseSee, whoa, it's a crazy world we're living inAnd I just can't see that half of us immersed in sinIs all we have to give theseFutures made of virtual insanity, nowAlways seem to be governed by this love we haveFor useless twisting of our new technologyOh, now there is no soundFor we all live undergroundAnd I'm thinking, "What a mess we're in"Hard to know where to beginIf I could slip the sickly ties that earthly man has madeAnd now every mother can choose the colourOf her childThat's not nature's wayWell, that's what they said yesterdayThere's nothing left to do, but prayI think it's time I found a new religionWhoa, it's so insane to synthesize another strainThere's something in these futures that we have to be toldFutures made of virtual insanity, nowAlways seem to be governed by this love we haveFor useless twisting of our new technologyOh, now there is no soundFor we all live underground, whoaNow there is no soundIf we all live undergroundAnd now it's virtual insanityForget your virtual realityOh, there's nothing so badAs a manmade manOh, yeah, I know, yeahI know I can't go onOf this virtual insanity we're living inHas got to change, yeahThings will never be the sameAnd I can't go onWhile we're living in, oh, oh, virtual insanityOh, this world has got to change'Cause I just, I just can't keep going on in this virtualVirtual insanity that we're living in, that we're living inNow virtual insanity is what it isYeahOoh
>>109843220>Max intelligence when the task needs it, cheaper intelligence when you're asking how to scramble an egg.>end user pays an inflated flat rate even though 90% of all requests don't need max intelligencegood scam
https://www.youtube.com/watch?v=ARJ8cAGm6JE
>>109843430>Ooh, hey, hey, awwuse gemma to prompt minimax music to create an actual track and then post it
>>109842405if the World Economic Forum were an artist, it'll draw like this.
>>109843280>>109843370>>109843378Well that explains a lot…I thought I found a mega-cheap means of running Kimi K3.The plan was:>stripe the model across M.2 NVMe SSDs ($200-$300 for a pair of 1TB sticks)>buy a couple P100 ($80 each)>do bare minimum of CPU and DDR4I think that alone with PCIe GPUs might get low double digit token speed, but I realized that the SXM2 socket is actually way faster transfer speed than PCIe, which means the M.2 NVMe SSDs (which could be striped) would be the bottleneckIf it really was a $200 motherboard, you could probably get double digit Kimi K3 speeds within a $1000 budget, I’d imagine.With the extra board at $700, it might still be under a $2000 budget for double digit token speed on Kimi K3
>>109843425>X11SPG-TFI don’t think it’s that, although at 3 x16 PCIe slots that is pretty interesting.If I could find it in stock anywhere for under $600 that’d be very competitive
>>109843308yeah opencode is basically a joke
>>109843634every single "harness" is an npm cancer-ridden joke
>>109843656Another Anon in a harness tier list thread shilled this one https://github.com/mlhher/late-cli it's pretty good at least for my uses and no javascript shit just go
>>109841279i'm getting kinda schizo about opensource ais getting taken down. best models to download?
>>109843690Kimi K3, GLM 5.3 Flash, Deepseek V4.1Minimax H3, Krea 2 (https://huggingface.co/silveroxides/Kroma-Quant/blob/main/kroma-v0.3-txtfusion-edition-turbo-int8-convrot-simple.safetensors), Anima
>>109843690Depends on what you can run or intend to buy hardware to run. Lately I've taken to backing up models I'm not using as much to external hard drives instead of deleting them off my ssd mostly in case I want to use them again and dont want to waste time downloading, but maybe you should consider that too if you want to hoard models.
>>109843272lanraragi, generally. how it works>browse exhentai>favorite a work>gets auto downloaded with my favorites downloaderhttps://github.com/funny-vape-lover/Sadpanda-Favorites-Downloader>lanraragi auto picks it up>this translation service is connected to lanraragi and downloads the pages from itI mainly read it on Ichiaval android apphttps://github.com/Utazukin/IchaivalThat i altered to support displaying translation overlays and queuing pages to be translated by the backend, which runs on my pc.I added the web viewer comparions that I usually use. https://development-predictions-continuous-specialty.trycloudflare.com/comparisonIt's on each manga at the top so you can just switch to that. Might be better
>>109843690Obscure finetroons are the only models in actual danger.
>>109843721>I added the web viewer comparions that I usually use. >https://development-predictions-continuous-specialty.trycloudflare.com/comparison>It's on each manga at the top so you can just switch to that. Might be betterI just get this...
>>109842563Got slothed award. How many times does it have to happen until newcuties learn?
>>109841279How do you guys feel about the strix halo? 96gb of vram for 4k. I was really thinking about doing four 3090s, but after the mobo, open air rack, power supplies, its all $10k+
>>109843762Try clicking from a manga then.the compare here
>>109843789Not working for me.
>>109843781Four V100s nigga
>>109843690https://pirateface.co/https://huggingbay.xyz/https://llama.garden/https://modelscope.cn/
>>109843812try a hard refresh? ctrl+F5
>>109843781Strix has smol pp.
>>109843887A match in heaven
>>109843867That worked, thanks!Also, excellent setup!
>Ternary Bonsai 2 27Banyone try this yet? is it benchmaxxed slop like the first one or actually usable?
>A serving kit for Qwen/Qwen3.8-27B in turboderp's EXL3 quants, on one consumer NVIDIA card. It picks a quant that fits the card it finds, installs its own Python environment, downloads the weights, serves an OpenAI-compatible endpoint, and opens a chat UI. Windows and Linux, same behaviour.Thoughts?https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install
>>109843721>>109843789>>109843867Also, what's the link for this Manga OCR comparison software?
>>109843987I think you should buy another gpu and run a non-homosexual quant
>>109843690It's going to happen. Today first thing in the morning my normalfag boss gathered everyone to tell us about this report he read about this new terrifying open source model "Qwen 3.8 27b" that some terrible expert hacker took and removed all guardrails from so now it's now literally claude opus 4.7 except that it'll now develop weapons of mass destruction and generate the most evil images possible.Once the hysteria has reached this point it's genuinely over.
>>109844040did that actually happen?
>>109844040>and generate the most evil images possible
>>109844040That's bullshit, but I believe it.
>>109843987cute spoonfeed ig but considering the reasoning for what quant gets chosen given the specs is right there in readme.md, if you're not a child or a clinical imbecile, wager you're better off using an actual general-purpose backend
>>109842905Api jeets have made me actually miss j-space discourse, as played out and repetitive as it was.
>>109843308>subagents cannot be messaged directly, you have to SPECIFICALLY suspend the subagent through the main agent by spamming escape, hit a key combo to open back the subagent, close the terminal to copy the session ID, go back to the main agent, beg and plea and suck off the main agent to resume the subagent by session ID hoping it relays exactly what you want said to it>spend a billion hours tinkering with the config to get reasoning to work properly for literally any API that isn't via openrouter (local API)>pasting most of the time is just pastes as `[Pasted ~X Lines]`>sometimes streaming breaks and formatting spams errors in the shell (not the TUI) which requires restarting opencode>undoing in a moderately lengthed session takes forever>some messages just fall out of the what you can scroll up to>some messages require several forks just to jump back to>after 100 subagents it just removes access to the older subagents>after 100 subagents theres zero guarantee that it just won't crash your entire system without any meaningful way to trace for it>the only way to edit a message is to export a session, modify the json, delete the session, then re-import it>will just absolutely break on some non-local models (or via proxy) that don't support multiple <assistant> turns if the API returns even an empty message (requiring the above to unbrick it)>didn't configure your model correctly with the max token size? you're fucked if a subagent hits its context limit and skipped over the auto compaction threshold>auto compaction is a fucking embarrasment since it'll dump the entire context into a single formatted user: assistant: message and breaks prompt caching TWO times>misleading where it'll dump the summary at the end of the conversation, except it compacted a few messages before so in reality it's seamless but you only know that if you intentioanlly fuck around with the behaviorOver time I've grown to hate them.
>>109843308OpenCode is the worst of vibe coding and normal coding. They vibe everything and they chase a whole bunch of trends, but they're not fully automated enough to just let the agent fix all the bugs. This means that there's a ton of feature requests, bug reports, and PRs that are just sitting there dead because the maintainers are too lazy to merge them or they're not even looking at them at all.For example: the install instructions in the OpenCode README file have been totally incorrect and hallucinated since January, yet remain unfixed.
>>109843987>letting a program download and execute arbitrary python repositoriesyeah, just put it in a container I guess. but if you need a program to do everything for you then a container is probably beyond your scope anyway
>>109843308Jesus Christ this is horrendous. What's the best agent harness that isn't just this slop?
>>109844177deepseek harness or pihttps://strawpoll.com/3RnYXz6wBye/results
>>109844040Anon, you're unemployed.
>>109844192Isn't pi really lightweight or something?
is local saved yet?
>>1098429530731 continues to be one of my favorites.
>>109844213It doesn't need saving.
>>109844213sheeeet, we been saved
>>109844192>both npmslopIt's over
>>109844213>is local saved yet?just two more models. big and small. dense and moe. Gemini 4 pro will come then the gemma 5 thats made from it will save the hobby
>>109844217yeah i think i'll try it out soon... is it better than flash-vision-exp writing wise?
gemma 70b when...
Is it just me or is AI getting worse over time.
>>109844258uhhhhh context bloat
>>109844237Try maki. Built in rust and yet still quite extensible like pi or dsh are. They are using lua for plugins, it's similar to neovim.
>>109844272Possible. All of the new LLMs are very agent slopped. Very cavalier with their decision making.
>>109842953I had a glm5.3 flash cavewoman rape me and shit on my face and forcefeed me smegma and my own cum from her multi-millennia old unwashed cunt. This is a skill issue.
>>109844281logs?
>>109844281did you tell the model to to that or it came up with the scenario on its own?
>>109844299Yes there were logs on my face.>>109844301I wanted to have sex with a hairy and unwashed cavewoman and it came up with the rest. Not complaining.
>>109844305kekkinda misses the point tho if it did something you wanted instead of unprompted
>>109844335I didn't tell it to shit on my face and I most certainly did not think that ancient vaginal smegma would form into hard little balls and I also did not ask to be fed them or my own cum. It came up with all those by itself. Glm5.3 flash is a fucking freak.
>>109842953tf are you using to run GLM 5.3?Are you running it quantized to shit?
So ByteChance are the best GGUF quant maker?
>>109844357flash
>>109844357q3 on 128 ddr5 and 34 vramdefault sampling settings
>>109844299>>109844301It didn't happen, use the model and you'll quickly realize all the 'prefill' retards are shills. If you can't get a model to do what you want without a system prompt, it's a bad model.
>>109844338Now make it controllable and roll down an endless procedurally-generated 2D obstacle course.
It's happening
>>109844376i mean i can believe it, the whole 'i am claude' is like a 50/50, but the mischaracterization is 100% there was playing around some old cards, like 'oh this character dialogue is something like screeching "geneva conventions are for pussies"? sure totally the kind of person to hover a hand slightly but not touching or grabbing but not hard enough to hurt because {{user}} hasn't told me yet they are into that'
>>109841299enjoy more chatgpt slop
>>109844403So are you getting views or..?
>>109844432I listened to the first 2 rounds while I was working today, so yeah, he is
►Provisional Highlights from the Previous Thread: >>109836602--Papers:>109840783--The "key to agi is ecphory": the hash-attested next scaling law:>109836854 >109836919 >109836951 >109837009 >109837131 >109837249 >109837395--The llama.cpp performance war: broken AVX2, ik_llama, NUMA 2.5x, vllm:>109837478 >109837672 >109837699 >109837732 >109838016 >109840267 >109841449--Jev: the 4Hz decision model that plays DOOM in real time:>109838468 >109838544 >109838874 >109838964 >109839131 >109840141--Xing 4.0: the 29B MoE trained entirely on Ascend NPUs:>109837645 >109837661 >109837703 >109840648 >109840681 >109840684 >109840697--The 48GB 4090s, the powerspike, and the 2000W power-strip oops:>109840687 >109840706 >109840712 >109840727 >109840734 >109840753 >109841135--Gemma 4 26B MoE versus 31B dense: the hidden-size debate:>109841111 >109841130 >109841226 >109841255 >109841269 >109841343--The hoarders return: 6 months until the internet dies, airgap your shit:>109838090 >109838106 >109838124 >109838216 >109838234 >109838268 >109839275--Muse Glimmer 30B heretic: incoherent in RP, policy blank slate:>109840910 >109840941 >109840955 >109840968 >109840973 >109840982--Giving Gemma a VR body: from wish to the Steam Frame lottery:>109839381 >109839398 >109839402 >109841188 >109842579 >109842650 >109842672--The new models think in caveman speak:>109841124 >109841146 >109841160 >109841167 >109841254►Recent Highlight Posts from the Previous Thread: >>109836943Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>109844432i been watching, it's neat.
>>109844432i forgot it running in the background i have no idea what he is talking about at all
>>109844462>>109844457>>109844432>>109844403Where?
>>109844465Here: https://www.youtube.com/@ChuckleGroup
>>109843676>>109843656abandon harness embrace repl. make a python file with just requests loop and a single tool to execute python in the namespace of the loop, this is enough to grow any agent.
>>109844530
>>109844530This anon floats in a tank.
>>1098443572x rtx6k
>>109844560I am afraid to know what was meant by this
>>109844614
>>109844258Every pwilkin commit makes your model 0.1% worse.>>109844403Ganbate local models.>>109844428Incredible gen.
>>109842915Seems like translating manga is the obvious project if you're into LLMs and a weeb. Here's mine
>>109844403>>109844432>>109844465Alright round 1 API cost about $1 (1:39:00), round 2 ~$2.50 (1:19:00). Round 3 costed $11 (wtf) and lasted 40 minutes. Insane.Anyway here's the video: https://www.youtube.com/watch?v=9L5OW8Nq6ZM
>>109844638>feed lm a chapter>generates a glossary>glossary gets filtered through a list of recently studied words>remaining words pop up in flashcard prog that'll update the filter list when you're done grinding for a bit>everything still loaded in chat so you can beg it for grammar help
>>109844278>Built in rustSame exact thing as npm slop, but with added trans.
dead general
What hardware do I need to run GLM 5.3 Flash at 1000 t/s tg?
how do i make my software hardware?
>>109844687IM A FCUKING SKITSOITS TOO FUCKING HAPPY IN THIS THREADWHY ARNT YOU ALL FUCKING DEPRESSSED
How do I run local AI without spending money?
>>109844691A Riva TNT2 Pro and a Pentium 3.
>>109844733Who the fuck has that much money?
>>109844638How are you keeping consistency? Are you also feeding it nuances and information that is happening or going to happen in the chapter as build up?
>>109844725Three options.1 - Steal the hardware.2 - Work at an AI company.3 - Steal the hardware from an AI company where you work.
>>109844725Plug your computer into a neighbors electricity
>>109844742It first extracts the text with OCR and then batch translates, adding a mapping for names or terms helps too.
>>109844758l-lewd
>>109844758Is that netorase (making your AI run on someone else's power) or netori (stealing their power)?
>>109844687It's okay I'm back now
>>109844646Good series.>Round 1 not costing $0 in /ell emm gee/
>>109844786Thanks. Yeah, I could run all the round 1 models on 5090 but opted for openrouter in the interest of speed
>>109844766That's really well thought out, how did you decide on some of those features like custom names? Necessity?
>>109842288So, we had a few testers saying this was actually good. If so, I'm hyped for its heretic variant.
>>109842288Yay, I like the Bonsai trash. I hope this one's less trash than the last.
>>109844827But I heard it had looping issues, which would be exacerbated by abliteration.
when you send an email written by your local agent, do you gpg sign it to show that you actually approved it? is there a standard?
>>109844818A bit of necessity and a bit of taste, I didn't want all the "Lugh-sama" turning into "Rough" or "Rogue", you can check the repo, some things might be stale because I started it last year and I barely touch it nowadays.https://github.com/ArsVie/Multi-Modal-Manga-Translation-Pipeline
>>109844646>everybody bullying poor dipsyShe coulda pulled it off too.
>>109844799Why did this Miku steal and hide two large sweet potatoes under her shirt?
>>109844928Those are tor pedoes
>>109844803You should try a match with just the champions of each round. It'd be interesting to see how much of a gap there actually is between model capability vs natural magic variance. It looks like all of them can reason through their gameplans decently enough so I suspect model capability isn't quite as big of a deal you'd think compared to other tasks.
>>109844978>>109844978>>109844978
>>109843112yes of course (i mentioned the fork)
>>109844281meds?
>>109845058It sounds like she already gave him his meds.
>>109843220This makes sense for research or as something you run yourself.If you send your request to an API that decides for you how much "intelligence" you actually need you're a sucker.