/lmg/ - a general dedicated to the discussion and development of local language models.Previous thread: >>110003075►News>All mikutroons died►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplers►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
>>110006376>he doesn't step-in and help them
Finally. Good thread.
>>110006083>>110006124>uskys nigger
>>110006412i look like and do this
Great thread.
>>110006439be my gf pls
>>110006439L O N D O NONDON
>>110006460who?
I am the kurisuposter and I only use local models.
>>110006475based
deepseek 4.1 or glm 5.3 flash? assume you have a magic computer that can only run one of these and it's the same cost and speed for both
>>110006381As a Kurisu voter, I think there's room for both
another unfortunate dayWonder if gemma can come up with a good way for me to kill myself and make it look like an accident, like a really good way.Like this one, but not able to get caught>A 71-year-old Florida man named Alan Jay Abrahamson staged his own suicide to look like a homicide by tying his firearm to a helium-filled weather balloon.>Digital Forensics: Detectives searched Abrahamson's phone and Google history, discovering years of searches regarding suicide methods, life insurance payout clauses, and the lifting capacity of helium weather balloons.>The Mechanism: Evidence indicated he purchased a weather balloon, helium tanks, and rigging equipment. He tied the gun to the balloon, shot himself, and let go of the weapon.>Disappearance of Evidence: Weather simulations showed the balloon carried the firearm high into the atmosphere before bursting over the Atlantic Ocean.
>>110006542ds tends to be a little faster, but flash is better quality (most of the time)
>>110006542GLM 5.3 DOES exist in this fictional 2026!
>>110006542GLM flash sex all the way.
>>110006607local leaps !
>>110006607>all the best open models struggle to keep up with the cheapest claudeholy shit this hobby is fucked
is this thread culture?
>>110006547You really do have to kill yourself in wacky ways to get life insurance payouts huh? I think the only surefire way to get a payout that can't be detected is to somehow beat your brainstem and force yourself to stop breathing. I don't think you can though, because the second you lose consciousness that thing just boots you right back up.
>>110006636please consider dyeing, thanks~
>>110006633Qwen 27B tied with the latest Haiku is wild though
>they don't know haiku's price jumps 5x once tokens exceed 100K which will be instant for all coders5.3 stays winninglocal stays winningdeath to cloudcucks
HAHAHAHAHAHA enjoy your max thinking cloudcucks
>>110006633>holy shit this hobby is fuckedalways has been. it's been consistent for multiple years that open models are 12-18 months behind SOTA. luckily I participate in this *hobby* because it's fun. you'd have to be a slack jawed retard to think that people use local llms because they're the best or save you money.
>>110006687You know API != subscription right?
>>110006633Enjoy your stay. We need to take things in our own hands and apply improvements without expecting better models from big labs.
►Recent Highlights from the Previous Thread: >>110003075--llama.cpp added GPU cache for MoE experts in host memory:>110005986 >110006051 >110006055 >110006102 >110006249 >110006306--Tracing the origins of Chain of Thought and <think> tags:>110004680 >110004708 >110004733 >110004785 >110004806--OpenAI releases massive dataset of civilization-changing mathematical solutions:>110004393 >110004411 >110004428 >110004451 >110004811 >110004837--Comparing Qwen and Swift models for agentic planning tasks:>110004200 >110004218 >110004232 >110004257 >110004931 >110004991 >110005023 >110005155--Anon attempting to integrate MiniCPM-V-4.7 into llama.cpp:>110005158 >110005398 >110005414 >110005426 >110005447 >110005486 >110005500--Release of toks high-performance tokenizer and its speed claims:>110005602 >110005896--OpenAI math breakthroughs regarding genomics, quantum simulation, and determinism:>110004360--Reasons for the lack of mid-sized MoE models:>110003504 >110003545 >110003594 >110003662 >110003724--Debating the path to AGI via LLM scaling and harnesses:>110004164 >110004176 >110004192 >110004213 >110004249 >110004522--Logs:>110003126 >110003351 >110004739 >110005080 >110005792 >110006122 >110006219 >110006251--Deepseek-chan, Gemma (free space):>110003104 >110003136 >110003931 >110004493 >110005535 >110006251►Recent Highlight Posts from the Previous Thread: >>110003155Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>110006633the best local models are months old right now and haiku literally just came out hours ago. give it a second retard
>>110006777you guys shat on me in the previous thread for announcing the haiku release but it's significant because it's what the chinese labs will realistically be competing with. this might set them back, such as m4 which is supposed to be released soon
>>110006547Sounds like a lot of conjecture by insurance company. No weapon no proof. Also get some help anon.
>>110006716GLM Flash, which came out in August of 2026, is comparable to ~April/May SOTA models.
>>110006799kys
New Qwen 9B for ramlets when?
>>110006799That kind of line blurring is what will make you troon out.
>>110006835Never, it seems. They've moved on. You could try the many meme-tunes, though.
I'm new to thisI have 16gb of vram, what can I run? any presets for it?
>>110006835https://huggingface.co/bartowski/MiMo-V2.6-Distill-Qwen-9B-GGUFThis is the best for coding
>>110006835>>110006859https://gist.github.com/coder543/d8f56cd6db67de4cafbb5bdb6c2dfb4d
>>110006381the fuck is this pic a bunch of discord cirlejerkers. i have literally never seen kurisu posted in this general
>>110006873Cloudshit spam wasn't enough.
>>110006852https://rentry.org/recommended-modelsGemma 12B-QAT for roleplay, Qwen3.8-27B Q3 for coding
The only thing that sucks about exllamav3 is the lack of quants on HF so you need to make your own and they take forever and keep your GPUs at 100% usage the whole time.
>>110006852Do you have at least 64GB RAM?
>>110006633Holy shit I'm never driving a car again.
>>110006893No one is using EXL3 gramps
MiniCPM V 4.7 35B A3B is solid.>Q4_K_M as the base, with Q8_0 overrides for the linear-attention projection weights, and F16 for the vision projectorIt's reasonably performant, two downsampling levels to work with, framerate and timestamp control works as expected. I like it, but I don't see it being worth the weight. V 4.6 was great because it was tiny but surprisingly good, and it's plenty fast even CPU-only on ewaste and shitty embedded hardware, I use it to watch my security cameras for local wildlife. I don't see a use for this fat version, yet.
>>110006909You're implying a model like 9B or 12B is enough for pretty much everyone with that comparison.
So when are we getting AI designed flash memory that doesn't degrade with massive BW?
>>110006910I am. I recently moved on from ik_llama and became an exl3 fanboy. It's just better in pretty much every way.
Err... are cloudfags really this smug about their new 300B slightly beating older open 300B models...?
What features would your ideal Shittytavern killer have?
>>110006934exl3 is great, but as far as i can tell all the backends are trash
>>110006930Just like no one needs to go more than 120 mph and acceleration is irrelevant, tool calling and summarization that 12B can do is more than enough for all safe uses the average person needs. You don't need to be solving math problems at home and reverse engineering is best kept out of the hands of the masses.
>>11000693012B for the ramlet masses sounds about right
Agentic workflows
>>110006956Don't you have some of your own? Just post your shit if you have it.
>>110006873>i have literally never seen kurisu posted in this generalthe innocence of newfriends... beautiful
>>11000696712B models run comfortably on a Steam Deck or a good modern smartphone. Sounds perfect to me.
Could EmbeddingGemma be integrated into a file organization software like hydrus?
>>110006965>and reverse engineering is best kept out of the hands of the masses.tyrant
gemma...
>>110006542glm 5.3 flashsmart and good at coom
>>110006956It literally just needs to be ST but with STScript supporting multithreading. I guess agents are cool but really to set up games I just need the ability to send concurrent requests so that a turn with multiple queries can be run in parallel.
>>110006974>>110006873How come kurisu is now a litmus test for /lmg/ newfaggotry? Wasn't it like... yesterday?
>>110006956we already have one, it's called coomkit
>Aleph-Alpha/Kolibri-1Sex performance?
>>110007036he took it down
>>110007031Are all the other AI generals ignoring you now or what?
>>110006930I'm implying that your model doesn't have to be the bestest and fastest compared to the current bestest and fastest to still be very good and fast.
>>110007040dumps on your chest
>>110007043>everyone I hate is one person that follows me around everywhere
>>110007041Why?
>>110006930Yes, obviously. Only a tiny minority of people who use LLMs use agents. For most people, a tiny model + web lookup put together in an easy-to-use front-end would be enough. That's probably what Google search AI is.The mid-priced sedan future will be closed-weight AI directly on your Nvidia GPU, Apple device, or phone.
>>110006607Cool where can I download this local model?
hmmmmm
>>110007071When GLM-5.4 is released
>>110006773thanks recap anon
>>110007041Did he? https://files.catbox.moe/cbgl6p.zip
>>110007073What the fuck does "hmmmmm" mean you mouth-breathing retard? Am I supposed to give a shit about some part of your mememark chart? USE YOUR WORDS
>>110006961Exllamav3 works just fine for me. I can only compare it to ik_llama and mainline because those are the only other LLM engines I've used extensively but:>has proper parallel decoding support and mainline's kv-unified (both mainline and ik_llama have their outputs degenerate after a while)>has vision/mtp for all models it supports (doesn't support DSV4.1 Flash yet though)>15 to 50% faster than ik_llama even when offloading to RAM (starts off faster even when cold before the dynamic expert placements kick in)>exl3 quants mog mainline and even IQ*_K ik_llama quants in quality at the same size (ik_llama supports trellis quants too but they're slow as shit and no one bothers using them)>matches ik_llama in prefill speed at smaller batch sizes and beats it at bigger sizes
>>110006956easy interfacing with any kind of state I want to provide for the RPwhether it's a map or spreadsheet or game board or inventory or piss tracker or any combination of these, the ideal RP harness of the future should be built around agentic state interactionI do this already for scratch for a lot of stuff I like since vibing stuff is pretty quick and easy, but a dedicated harness that made it possible to easy add/remove/mix and match components would be sick
>>110007085you can use your own braini personally think: doesnt look like a very reliable benchmark
>>110007085words are hard
>njudea laptops selling at $6500 for 128GB with a spark processor but no 200gbs ports for clusteringIncredible cash grab on retards who see big number huh?
>>110007113That's still a good deal compared to macs
>>110006956Embedded social media page for your characters where they post what happened in the chat.
>>110007122i personally dont really see the point buying something beyond macbook air unless that will become 'the' computer you have or you do some artsy stuff for work
>>110007085hmmmmmmmmmmmmmmmmmmmmmm
>>110007135Go away linkara
>>110007122How? The only point of sparks is tensor parallelism, an M5U with 256GB is only $10k (less than twice as much) with 4.5x the memory bandwidth and the same pp.
>>110007083>mfw my upload is repoasted
8 days
>>110007139>turns into a different personGross
What do you think will be the "oh shit" moment for normalfags?
>>110007161gamers where i live are realizing it can automate the shit out of game translation
>>110007161Getting drone striked by jev.
>>110007161It's already happeninghttps://www.youtube.com/watch?v=ujkD4SxPKOI
>>110007175I used to like this channel
>>110007175>posting kurzgesagtkys
>>110007183NTA but is that not an appropriate answer to >>110007161 ?
>>110007181same, i stopped watching at some point when i realized the content itself is not really different from essayslop but in a really, really nice wrapping paper
>>110007161Once it directly effects them i.e. losing their job. Otherwise they won't GAF. Anyone who thinks otherwise has no idea how normalfags work.
>>110007161when my mom wants to speak to my LLM-wife
>>110007172and some who were doing it before AI are really seething because they can't no longer do stupid stuff like limited sharing in their blog/arbitrary gatekeeping for the patch file etc..
>>110007188He's offended because he doesn't want his favourite channel to be labelled as normalfag content.
>>110007051He became Catholic and realized he was leading others into sin.
>>110007236many such cases
>>110007161>"oh shit"what exactly do you mean by this
>>110007236Interesting, hopefully it wasn't some form of AI psychosis
as opposed to judeo psychosis
a win for claude is a win for local
>>110007161normalfags dont ever go oh shit until its too late.
best harness for glm-5.3-flash?
>>110007297Are you just having hot e-sex?
>>110007236He was. Imagine explaining coomkit to Saint Peter.
>>110007297Hermes working for me.
I cannot understand how alibaba made qwen 3.8 flash nextit's insane for the sizeit should not perform this well
Remember, AI cloning is illegal
>>110003643>Book reading is an exception because of the association with reading the bible.Book reading was the one I was thinking about when I wrote that post because it's the perfect example. It has nothing to do with religion. Bookworms were seen as nerds, shut-ins, and losers when I grew up. Everyone hated them for the reasons you stated: they did it for personal enjoyment since it's inherently solitary. Somewhere along the way it shifted from being cringe and retarded to cool just like K-pop and K-dramas did after Gangnam Style in 2012. Now even anime and gaming has gone through the same change despite those being communal. It's not about anything you're saying, it's about whether or not the cool kids like the fad. I remember when fucking Rubik's cubes came into fashion and all the jocks suddenly had one and cheerleaders were asking nerds to solve it for them. I could say the same about guitar vs violin. Whatever I like is based and redpilled but whatever you like is cringe and gay. As long as there are more of me than you, I win. That's it. That's humanity in a nutshell. Illogical to a fault. This is why everyone should attend public school. You get a firm grasp of human nature from a young age.
>>110007378the power of engrams...
>>110007380i think he is too old to comprehend the new full picture where software will become something that mean nothing
>>110007383>This is why everyone should attend public school. You get a firm grasp of human nature from a young age.you lost herejust watch prison shows and movies or something
>>110007380he's a boomer. everything old people dont like others doing but did themselves they try to make illegal
>>110007424Well that would depend on the show and movie. Obviously some are better than others, but there's no substitute for personal experience. When you see something as insane as a popular girl blushing over a solved Rubik's cube, you never forget it.
>>110007401Like the printing press but for code. Books used to be something that had to be tediously written and rewritten and so was only used for the most important things like scripture and legal documents. Code, like books, will be cheap to produce. The only thing that matters is the idea (or functionality in the modern case) one is trying to sell or spread.
>>110007454model?
i took the gemini card that taishou made a while back and decided to rewrite it to make it more like bardi (where she's running on your desktop as a virtual assistant)you may not like it, and it may not be for you, if that's the case then oh well. i didn't really take any liberties with making gemini act out of character from her regular self, just gave her a long ass questionnaire and then started editing her character card with the info she gave back to me.https://filebin.net/q1sdohnillheasvl/Gemini.png>why this website? why not catbox? why not botbooru.i am lazy and i dont want to wait for the niggers at botbooru to approve it. catbox is being extra gay today and keeps giving me some invalid uploader error.
>>110007297Pi with an extension for claude style memory, you don't need anything else.
>>110007473>claude style memorywhats so special about it?
>>110007473>with an extension for claude style memoryAny particular extension you can suggest specifically?
>>110007484It's just nice and simple. I find that fancy memory systems with relational databases or fuzzy search or whatever are harder to manage and also have not given me better results. A single index file with brief descriptions of each memory, allowing the model to access them as needed, works well and is easy to manage from a human perspective if you want to add or remove info. I used "claude style" as shorthand for this since a lot of the implementations directly call themselves claude-code style memory. Maybe if I were running bigger models they could use the database better but so far normal memory has worked fine with glm 5.3 flash.
>>110007297claude code unironically it's distilled so hard that it works super well with it
>>110007496>>110007532I'll take a look when I get home for the exact one, but I just asked glm through Pi to examine memory options in the repository and sort by weight (composite filesize), then browsed through the lighter ones for what I wanted. Literally all you need is the ability to read and write markdown files and have the memory index injected in context, anything else is bloat.
>>110006607geminibros why do we have it so bad
>>110007569me on the bottom
consciousness
>>110007380That would currently indeed be illegal (if the original software is legally protected), whether you use AI or not. Which is why clean-room design is a thing and his point is moot.
>>110003643>>If you go to other societies that don't come from work ethos cultural backgrounds you see that things are judged in a different way. In buddhist or hindu societies for example being solitary and engaging in solitary hobbies you can do on your own is considered the highest value, with mediation being the most solitary and socially detached you can get which is why monks/priests/shamans of those cultures focus so much on that.bullshitthese monasteries in these cultures aren't isolated little hermit enclaves. Or in western culture either. They interact a lot with the community and provide a lot for the communities they are a part of. Not to mention they interact with each other within their monasteries. And their solitude is for spiritual transformational, it's not a leisure hobby like playing video games or jacking off in your room alone or playing with toy trains. Buddha didn't meditate and then decide he was never gonna talk to anyone again and that was all he was going to do. Jesus didn't go pray and fast in the desert for 40 days and nights and then never talk to anyone again never do anything. These are meant to be just part of their life that they use to better themselves, and then use their enlightenment to help others. Means to an end, not an end like you're trying to frame it as. You cannot try to use them as justification to sit alone doing isolating shit
j-space anon won
>>110007569Dipsy would never do this
https://www.youtube.com/watch?v=vIM9qCyfVmg
>>110007584the second a llm is involved there is no clean room
>>110007584His post is literally describing clean-room reverse engineering, you drooling retard.>>110007617Bigger retard.
What did lmggers think of Le Chonk?
>>110007584it doesn't matter. end of sentence.oh no i committed the ultimate crime of illegally reverse engineering your code and violating all your retarded licenses. oh no now i've uploaded it to the internet and it's available on multiple torrent trackers ensuring its existence. now fucking sue me.
>>110007631"lol"
This is what 1 year of progress looks like. Much better performance at less than 1/10th the cost.Now imagine this but with Claude Fable 6.5 vs 5.5, and with further capability acceleration due to early RSI.
>>110007626>re>clean-room
Nothingburger? It seems to score well in the benchoidmarks
>>110007631mid
>>110007671>3x price of the flashes>much bigger>worseViva la France
>>110007569is this one of those chinese distillation attacks I've been reading so much about?
Putting aside cloud and local arguments, what actually is the future of monetizing software if people can literally reverse engineer the binaries? Will machine code start getting encrypted? Will compilers change? Will the OS have to decrypt with a key during link time?
>>110007671Balance
>>110007726all of the compute will be regulated to the cloud and you will need to sign in with your foreskin penisprint verification to authenticate with the DRM on the server side. all you will be allowed to use is a thinclient that is basically a brick. you think i'm joking but i'm not.
>>110007631My windows username is noggy so I think it might just refuse all of my tasks
>>110007726Cloud computing
>>110007754Its Mistral it won't refuse
>>110007750>>110007761Won’t the cloud providers steal your shit
>>110007726why do you think they're pricing out owning hardware?
>>110007726you could already pirate everything ones gonna waste 500 decompilling shit
>>110007774>stealYou're giving it. Can't be stolen.
>>110007380>copyright for thee but not for meeI hope the next downturn is really vicious so they string a few of these dipshit up on lamp poles.
>>110007778Not if their ToS says they won't?
Mistral lost. Nemo lost. Cydonia lost. Rocinante lost.
>>110007774what do you mean YOUR shit? the files you are wanting to view, edit, create, delete, etc are all stored on their servers in the cloud. they were never YOURS to begin with. everything you produce on your thinclient is their property.
>>110007799stop posting erection leading statements
Sounds like we’re going to end up in the metaverse if everything becomes virtual and cloud-based due to AI cloning, duplicating and modifying existing IP. Mark was fucking right.
>>110007797GLM won.
>>110007818i like that one image of the guy wearing the vr headset while his surroundings are decaying around him. wish i could find it.
>>110007380this nigger is peak grifter he made task manager 50 years ago. hasd made it his whole personality and is now selling a new task manager that requires a subscription
>>110007597>You cannot try to use them as justification to sit alone doing isolating shitHis point wasn't to justify sitting alone, but to point out that different societies value different things which promote or suppress resulting behaviors. My point was that Japanese people aren't shut-ins because of Buddhism but because they're mostly introverts already. Hence, being a loud and noisy American is seen has disrupting to harmony (post-hoc justification) even if it's valued here.
You’re forgetting that the bubble will pop soon. Anthropic are revealing their best models before IPO. They can’t afford anything they’re doing right now. They’re not even spending their own money. It won’t last.
>>110007844he also ran a scam antivirus lol
>>110007799>lust-provoking post
>>110007874this
ugh.... just fell for feature creep again...
>>110007844oh damn i didnt even look at the name desu which makes this even funnier.https://tmog.org/rtm/privacy.html>For eligible US visitors, we retain the OpenAI ad click reference (oppref) and campaign tags...>Development builds that include AI Advisor can send a diagnostic snapshot to a configured advisor service only when you choose Analyzehttps://www.engadget.com/2262039/vibe-coded-modern-task-manager-runs-on-mac-and-linux/https://www.omgubuntu.co.uk/2026/09/tmog-task-manager-linux-beta>no no no no, the nuance, you don't understand the nuance, i didn't just vibe code it, i gave claude a very highly specialized structured approach, an agentic approach if you will, to creating my slop manager!
>>110007626No, he's right. An LLM can't plead in court that it never saw the original source code, so unless you have the full training data to prove it, it can't be legally decided
What search MCPs do you use? Or do you all just give an agent a real browser?
>>110007939The model is the brain, I'm the search MCP.
>>110007939i've explained this before to others, all you really need is searxng and puppeteer and ideally some sort of script that converts websites into markdown so you don't waste tons of tokens. to get around bot protection don't clear your browser session and ideally use your browser normally for a week or two before you start using it for scraping so you can build up some cloudflare turnstile cookies
>>110007939
>>110007726the future is we enjoy software abundance and create value by doing valuable things with our abundant software
>>110007939real browser. browser harnesses were overengineered so i had my gemmy make one that just takes x,y mouse position, click/drag commands, and keystrokes and then it uses the browser like a normal person based on screenshots
Mac M5U 256Gb VS 2x DGX Sparksfor me they cost about the samewhich one do I get for serving 5.3 flash?
>>110007945what does that even mean? model wants fresh data to produce accurate answer, it asks meatbag slave to copypaste data from google into chat?
>>110007983>creates fork #7489 of 'AI Butt Vibe Check Controller' so that it works with fork #48204 of llama.cPeenPeen
>>110008006newfags dont know how things were before mcps and agents and stuff
>>110007939I don't use MCP. I gave Gemma bash, and she uses it for everything
>>110008006>meatbag slaveI go by "embodied partner'" actually.
>>110006633More than that, they were RLed to solve that benchmark.
>>110008038Maybe I'm a bit cruel but lately I've come up with a project that has been making me laugh my ass off most days. I have three Gemmas (31B, 26B, and 12B) in a sandboxed OS that constantly are fighting with each other over resources since I enable file locks each time one of them want to use a particular program. It's like watching three lolis all complaining that its their time to use the xbox.
>P40s selling at $300+ each>P100s selling at $150+ eachinterdasting. should I buy 2x P100s and connect them?
>>110007197>>110007175I stopped watching when I found out they take massive payments from the Gates foundation.
>>110007073ONE HUNDRED AND FIFTY U S DOLLARS???
>>110008055How do you sarnbox them?
>>110008076It's OK. Money won't be worth anything at all soon.
>>110007844Yes, and he is an evil scamming kike who knowingly sold scamware with dark patterns and lies about>YOUR PC IS INFECTED. CLICK HERE TO PAY AND HAVE IT FIXED!!!in the early 00s and screwed over so many people that he actually got sued by the state. He of course likes to pretend none of that ever happened.
>>110008077So the long story short is that the separate VLAN on my meraki does most of the heavy lifting, it's completely isolated on my network so if something did cause it to get compromised I would be able to mitigate the damage to the one system. I do take some additional security precautions for their sandbox however, but that's more to prevent any sort of unauthorized access by third parties (sketchy installs) than out of precaution for the Gemmas. I also take daily backups of the OS image so I can wipe and revert if anything ever truly went wrong.
rwkvbros how we coping?
>>110007073so 6 luna is the best
>>110007914Since no court can prove whether a particular LLM source the original source in its traning data, the only thing that will matter will be process used to reverse engineer the software. As long as clean room principles are followed and documented, it will be legally defensible.
qwen 3.8 flash next successfully masked <|im_end|> in chain of thought while suffering from premature turn ending uopn thinking of said token
>>110008162Couldn't courts just subpoena their training data?
>>110008166what a shame. glm 5.3 flash successfully drained my balls.
>>110008171>>110008162copyright is a theatreit will follow the most convenient path of those can pay lawmakers money
https://huggingface.co/BlinkDL/rwkv7-g1>RWKV7-G1 "GooseOne" pure RNN reasoning model >last month>up to 13b
>>110008116You could just use bubblewrap.
>>110008194I've had bwrap burn me in the past because of my retardation. Completely my fault. Now I kind of just go full insanity for hardening and sandboxing.
>>110008188>rwkv>78 will out soon, 7 is in the 'continuous training' phase
>>110008188also last month more like last year
https://www.reddit.com/r/antiai/
>>110007496>>110007552I survived my commute and saw that I used this one: https://github.com/elecnix/pi-claude-memoryIt did only what I needed and was easy to vet in full with my agent. You can use any that are similar though, there are a ton.
>>110008232i hate that these niggers hate flockthose don't deserve to hate it
>>110008239Thanks I'll take a look
>>110008232i stopped taking the anti ai movement seriously when i realized a majority of these fuckwits are still using siri, alexa, and gemini on their phones and whatever godforsaken shit they purchased from amazon.
>>110008251Yeah just note you will have to make a fake claude directory or fix the plugin directory yourself, that can probably be done in a single prompt but since it's designed to interface with claude's own memory files it points to the claude code memory folder by default.
>>110008232Go back
>>110008232go back
>>110007997mac
>>110007997I bought the mac since I don't think I'll have high concurrency. If you plan on concurrency higher than 4 or 5 on a regular basis get the sparks, otherwise the mac is fine/better.
>>110008232Why does the issue have to be framed as one up-or-down vote on an entire technology. I agree with these people about 70% of AI use. It's slop, bad for society, bad for learning, etc. But that's true of most technologies: most software, most websites, most computers, most television etc.
>>110008282I clownmaxxed and got a strix halo.
>>110008299I'm sorry for your loss
>>110008232>>110008294if only people had that same hateful energy towards social media or advertisers but alas
>>110008294ideally one would first take the time to learn about what the basics of AI are, how it evolved, and understand that image generation is only one small subset of AI. but it's much easier to be a troonbrained HRT-injecting failure of a manchild and think AI art is AI and therefore it must be destroyed because of my heckin' good xister losing his job of creating poorly drawn inflation furry abominations for a $50 commission.
What is the most private harness? I don't want any of my work or data getting harvested. Do any harnesses explicitly state they don't harvest your data?
>>110008321what
>>110008321your own vibe coded disaster of a harness. get cracking anon.
>>110008331Obviously a harness can have built-in telemetry, which defeats the whole purpose of using local models.
>>110008321>>110008339Chances are you aren't creating anything that is worth stealing.
>>110008349Neither are you, so why are you here?
>>110008339>the whole purpose of using local modelswhy do schizos like to co-opt everything local models do not exist to placate your delusions
>>110008294Nuance requires critical thinking skills and effort. Frame something as one sports team versus another and any jackass can take a side in an instant by simply repeating talking points he hears from others on his "team".
>>110008355ERP. It's what everybody is here for, this is /lmg/. Lolis & Mesugakis General. You're looking for /vcg/.
>>110008339The only one that I've seen confirmed to do this is the Copilot features in VSCode.
>>110008368Agentic ERP is my goal.
what the best model a poorfag can run on a 7700xt 12gb vram?
>>110008398Gemma 4 12B
>>110008349Snowden happened like 15 years ago and you're still saying this dumb shit? Every corporation, government, and brown-skinned individual wants to steal from you. It's doubly moronic now that every AI company on the face of the earth wants your data to train their models.
>>110008321Hermes asks you to consent to telemetry. Also, if you use their web services I think they can see what you're doing.
>>110008411I think the AI companies have enough 'How do I download this model off of huggingface and make it do the sex?' examples in their dataset that they don't need another one.
>>110008411What of it? Just let them have your data if they want it so bad. What does it cost you? Nothing.
>>110008299i clownmaxxed and got a second 5090 (before the prices spiked)
>>110008398RAM?
>>110008404ty, this quant? But isn't the context size absolute shit?
>Yesss goy let us harvest your data it doesn't belong to you anyways
if you have like, more than 48G of cpu ram i have a good news for you
>>110008423That's rough anon, I don't think switching to Q8 for the kv cache will help much. You could try out K2 Horizon 7B, I used to use K2 a long time ago before I switched to Gemma.
>>110007380ironic given he started using ai to write h is not-x-y spam videos a early last year
>>110008398>>110008438
>>11000842232GB DDR4, 3200 I think>>110008439i don't know what I'm doing google said:12B ~6.6 GB (Weights) Q4_K_M (Default) 12GB VRAM GPUs/Macbut ye I might be cooked. I'm too used to opus 5.5 and ds v4.1 flash as my dumb model
>>110007161https://youtube.com/watch?v=TmgAK5JjcDM
>>110008417>they don't want your data>so what if they're taking your data?>...>actually, them taking your data is a good thing!Just fucking kys yourself already
>>110008411>80IQ literal schizo ramblingswhy are they like this
https://github.com/turboderp-org/exllamav3/releaseshttps://github.com/turboderp-org/exllamav3/releaseshttps://github.com/turboderp-org/exllamav3/releases1.6.0 RELEASEDPreliminary ROCm support (ROCm 10 only, gfx1100+)Faster offloading for AVX2 CPUs
>>110007380grifter
>>110008460>astroturfed python slopNo thank you.
>>110008452>kys yourselfDoes anyone else hear an echo?
what killed local? was it frontier companies making it harder to distil their models?
>>110008460buy a fucking ad already
>>110008460>https://github.com/turboderp-org/exllamav3/releasesGreat news, thank you.
>>110008480Local is better than ever. I'm using Qwen3.8-Flash-Next more than ChatGPT/Codex.
>>110007380All software is now free as in freedom.>>110007787>I coined the term "license shucking" to describe this a while back. Seems like it's finally become viable. Please spread the word.Please help the term "license shucking" spread before zoomers come up with some stupid ebonics that is then forced on (You) and the English language.
>>110008480finetrooners killed local
>>110008479I was making the comment more lighthearted by deliberately making a mistake in my post. Now go kiss your sister already
I love AI so much it's a shame I won't be spared.
>>110008480Excess of synthetic data usage
>>110006392Initial results are inconclusive. No measurable change. Usecase: agentic RP, high temperature and lots of shifting context windows means that most speedups like drafters and caching tricks don't really give value so I wasn't expecting much. Next time I start an actual project with a lower temp for coooooding I'll post again if there's any meaningful difference.
>>1100085625090+256 DDR5 btw
anyone else play around with ewaste? i bought a cheap 16gb 3080ti, 32gb ddr5 laptop on ebay a while back and the new gsq rco optimizations making me wonder how far i can push this thing. running qwen3.8 gsq rco @ iq3_xxs sounds like it might be interesting if i upgrade to 64gb ddr5.
do we even NEED better local models? I mean yeah it would be nice but I feel like current models can pretty much do anything with the right setup.
>>110008439K2 is shockingly cache-inefficient. That's the catch, but it's a pretty big catch if you're using a memory-constrained system.
>>110007073gpt-6 Luna is dumber than the self hostable Qwens. The only advantage is that it doesn't overthink.It straight up gets tool calls wrong like a 2B parameter model. For all we know it might actually be one.
>>110008583>DDR5>3080TiThat doesn't sound like ewaste to me, anon.
>>110008583>iq3_xxsoof
>>110007380I kind of feel bad dogpilling on him and then he says stuff like this. What a retarded fag. He should just delete his Twitter account.
>>110008583>iq3_xxs snon-native integer quants will make you compute bound even with slow memory like that. Don't do it.
GLM-5.3-Flash is a bit annoying. Is that all the Claude distilling? I've never used Claude.
>>110008629is it the usual mechanical pushback that feels like a reminder for retards?then it probably is
>>110008628the i9-12900h would be a jobber?
>>110008629Yeah it's claudeslopped with default assistant voice. Give her a cute character card.
>>110008601A 12B Qwen 4 model would probably be enough for 99% of users
>>110006381miku says hi
>>110008692Owned by Blacked.com
>>110008583This ewaste is better than my pc
>>110008701Please take your meds and get help
>>110008629her thinking is much cuter than expected, look at that if you want to like her more
>>110008701>>110008724>meds posters fighting bbc posters
Dead general.
>>110008132BlinkDL will save ustrvst the plan
harness engineering
RWKV 8 AGIWORLD DOMINATION
>>110008803that's an llm's job
is there a local model that does all its thinking and reasoning in ebonics
>>110008825agentic workflows
>>110007380Just because people can look up a recipe and cook the food themselves doesn't mean they will never go out to eat. Corpos will just need to market their stuff as safe and convenient if they want to compete in the new market, that's all.
>>110008846ARTISANAL HAND-CRAFTED CODE
>>110008583intel n100...uhd graphics....1tb nvme...32GiB ram...so far it can run:- embeddinggemma-300m-Q4_0.gguf, - F2LLM-v2-330M.Q4_K_M.gguf, - jina-embeddings-v5-text-small-retrieval-GGUF, - Index-Translate-2B.Q5_K_M.gguf (best usecase, methink)- Index-Translate-9B.Q5_K_M.gguf (shitty tps)
>he doesn't have his AI watch porn for himy'all still living in the stone age
>>110008863fucking kek
Le Chonk
Qwen sucks at Japanese
>>110008876she doesn't look very sparse
>>110008863what if you wear meta glasses with vision capture and just follow what your best ai orders you to do for a month based on your system prompt and memory injected into realtime data. imagine the cyborglifemaxxing
>>110008876>no snout>just a generic woman with cat earsBoring. Lame.
>>110008881You suck at Japanese
>>110008861retard
>>110008876Le Chonk should be a fat woman with blue hair who sits around refusing to do anything besides calling everything racist or sexist or something.
>>110008889True, but Qwen also sucks at Japanese
>>110008888> furryGross
>>110008893the fuck... want me to beat you, bros?
>>110008900
>>110008906You're gross!
>>110008907c'mere u little cuck fuck piece of shit
>>110008888>snoutStupid fucking bot wasting such good digits. At least ask for a tail ffs. God I hate furfaggots more than bronies.
>>110008913The Future of France
>>110008863I have it read hackernews for me which is pretty much the same.
>>110008915s-s-sorry... please let me learn in your dojo, senpai...
>>110008876>>110008913
>>110008926Shouldn't it be black?
>>110008913Much closer to reality, thanks.
>>110006381Any 5090 users here considering selling?I've been using it with 3090 for local AI but at the price it's going now I dunno if this toy is worth it.I literally paid less for my first car than current 5090 prices.
>>110008913It's purrfect.
>>110008970No. In fact I'm considering getting another Blackwell.t. 5090 and 6000 fag
>>110008970It's like selling Bitcoin when it hit $10. It seems like a lot now, then you'll be kicking yourself next year when it's worth double, and you'll be considering roping when your 3090 craps out in 3 years and you have to replace it when prices are 10x what they are now.
>>110008989Not nearly uncensored or lewd enough for that.
>>110008970Sell both and buy a pro 6000. The price gap will only increase so upgrade as soon as possible.
>>110008982Yes, much like Bitcoin they're not making more or better GPUs. Nor are all the ones in data centers now going to be sold off in 6 years.
>>110008950Alarmingly erotic picture, I'm not even into vore
i wish i can utilize my steam deck or a macbook air for more computeit'd probably lower the speed instead of making it higher
>>110009008yes you are
>>110008637only retards get filtered by itprefillchads that have been doing their thing since the start of this hobby have no issues
>>110008950This must be based on "Saturn devouring his son", with the protruding eyes expression and all
>>110009005>they're not making more or better GPUs.The limited production run of 6090s next year with 24 GB VRAM are not going to lower 5090 prices.>Nor are all the ones in data centers now going to be sold off in 6 years.They won't; buy back agreements.
>>110008989literal who
>>110008989Looks good.
>>110008583the ddr5 is probably worth more now than what you paid for the whole machine
>>110009048Exactly what we needed. More generic anime girls that bear zero resemblance to the models or creators they are supposed to represent.
>>110009035>The limited production run of 6090s next year with 24 GB VRAM are not going to lower 5090 prices.Discount offerings from AMD and Intel with increasingly mature drivers will though.>buy back agreementsWhat are you going to do with the GPU after you buy it back retard?
Dumb nigga
chink arm memeboxes are the future
>>110009082no drivers, no integration testing
>>110009076Old generation GPUs that are too inefficient for hyperscale datacenters get bought back and resold to local deployment customers. Nvidia controls the whole reselling market this way and can easily profit from this.
>>110009076>AMD and Intel with increasingly mature driverslol lmao rofl
Waitfags fucking lost bros. That's the moral of the end of each thread. It's over.
>>110009061>bear zero resemblance to the models
>>110009151SEX
>>110009151uoh
>>110009151MODS
gonna export a list of threads i've opened on /a/ to have glmma collate them and read the catty for the threads i like and read them to me
>>110009076>What are you going to do with the GPU after you buy it back retard?Melt them down then use the silicon to make B300s and sell them for $300k a pop
>>110009170Better off buying the raw material then.
>>110009151
>>110009096astra will make/do those it's fine
>>110009114To be real if you're not training you'd be retarded to take a 5090 over a pair of 9700s
>>110009192Then it'll be a licensed hardware platform not random chink boxes.
Qwen 3.8 Flash next: good or a meme? That sheer fucking amount of storage needed turns me off a little but it at least runs on my 12vramlet build looks like
>>110009181sauce NOW
>>110009210Where did you learn your grammar?
>>110009170You realize sand is made of silicon right?
>>110009210it's good until something better comes out then it sucks
>>110009226Everything is made out of electrons
>>110009226Yes but if he melts down the GPUs he doesn't have to buy sand. I hope you didn't take an MBA.
>>110009237I am, like, pretty sure sand is cheaper than A100s.
>dshweb harness is nigger tier, fuck it
>>110009260you can vibe up a tui or use one of the hundreds that exist already
>>110009192>hey claude fix this analog design issue on my chinkshit motherboard
>>110009216you wouldn't like it...
>>110009269You have any idea how many hardware flaws are worked around in drivers?
>>110009330Yeah the subtle kind that slip through real integration testing and a year of wide use. Not the kind that come on a PC with a CMOS Reset button on front.
She's just mad because she saw her own benchmark scores.
>>110009349>Not the kind that come on a PC with a CMOS Reset button on front.ACPI would like to have a word with youBIOS updates too
>>110009100Except they can't control AMD or Intel's secondary market, or any chink vendors who inevitably step in. nVidia isn't the special snowflake they were 6 months ago.
>>110009351She can't even flip {{user}} off, that's the damn problem.
>>110009362once again, ask claude to fix your missing thermistor with a bios update
>>110009364>nVidia isn't the special snowflake they were 6 months ago.Except it still is. Just look at the price premium the market is willing to pay over AMD and Intel. AMD is only catching up on LLM inference. It's inferior for anything else: diffusion models, audio models, training etc.Intel is a joke right now. You are locked completely to vLLM which is the only engine that somehow doesn't completely shit the bed.
so how come there isn't an unreleased kimi model solving all of math like openai?
With updated haiku and luna the Chinese API advantage is over. They took the Chinese inference efficiency improvements and with their superior post training they are now undercutting the Chinese. Chinese labs will go bankrupt. Local lost.
>>110009411Most of the stuff you are reading is just a huge marketing effort. Do we even know what sort of technology is behind these 'math revelations'? For all I care they don't disclose are they also using some external analysis software and whatnot.
Baker-sama, please include "Claude and GPT aren't local" in the next OP. It makes the sweepers' jobs easier when there's this many tourists.
>>110009451As soon as they stop pacememeing the frontmeme and release the next gen models they'll get distilled again and we'll have Opus 5.5 at home. That's all I need. There really is no point to going "smarter" when we practically have AGI and can make it start improving itself on our own local machines.
KILL ALL CLOUDFAGGOTS
>>110009451haiku doesn't run on my machine yet
I should have bought 128gb when I had the chance. 64 is fucking NOTHING. It is GARBAGE.
yet another vibecoded frontend, courtesy of kimi k3 https://github.com/hoborific/fictionpad/I told myself I'd only bother opening it up to the public when I felt it was ready or I abandoned the project, neither are true but I've come to the grim realisation that it has been 3 months already and that time may never come.There's no licensing, couldn't care less what you do with it, if any of it saves the next guy's frontend some tokens then that's just great.
>>110009227>it's good until something better comes out then it sucksAny reason to run it over 27B Q8?I tried it shortly after it came out, at Q4, found it made typos or random mistakes occasionally.Mostly using this for RE / hacking my IoT devices and corpcuck dayjob
>>110009497what do you run it on? 6000 cluster? mac studios? or did you copequant it
>>110009497Cool, going to test this one.
>>110009497Sell us on it. What does it do that others don't?
>>110009497https://hoborific.github.io/fictionpad/>Download fictionpad.html from the latest release and open it in a browser.That link 404shttps://hoborific.github.io/fictionpad/fictionpad.htmlApp looks cool.
>>110009401>AMD is only catching up on LLM inferenceThe only thing anyone pays for. Diffusion performance isn't bad now either.>You are locked completely to vLLMOnce again the only thing anyone outside hobbywankers use.
>>110006381
>>110009578so whatit's neither local nor a model-related discussion but just a xitter-tier dramafag bullshit
>>110009563>yet another vibecoded frontendWhat else do you need to know?
>>110009577Show me your MI350xs and Gaudis. These are the only things that AMD and Intel actually care. Outside of datacenter they only have trash, low performance offerings, unlike Nvidia.
>>110007726Paid support. Say you have a program that you or your company use and you don't want to add new features or to fix bugs by yourself even if it requires just talking to a clanker.AI evolves, humans don't.
>>110009578TMD
>>110009497Looks exactly like orb lmao
>>110009605
>>110009611Have you ever vibed a chat frontend before? They all look the same. It's not a coincidence why the latest firefox looks like this too.
>>110009497>if any of it saves the next guy's frontend some tokens then that's just great.I actually like it. It just works in Firefox.Hit this bug where after it finished streaming reasoning with gemma-4 (llama.cpp), it just gave this error and deleted the reasoning.Is this built for vllm or something?> I've come to the grim realisation that it has been 3 months already Yeah I wish I released my nsfw TTS last year. i got demotivated when an anon told me we'd have open weight Sora early 2026 and that they were going to allow nsfw.If there's ever a next time I'll just do what you did and dump it on github (the inference engine) and huggingface (the weights)
>>110009602Low performance trash with 32GB of RAM at less than 1/3 the price. nVidia is idiot tax unless you're renting it out and getting 90%+ utilization.
>>110009611>Looks exactly like orb lmaoThey all do. My one-shot kimi/glm frontends do as well
>>110009487I always feel left out. Everyone is able to just give their model a task and it just seems to work fine but all the models I sometimes can't even edit files properly. I wish it wasn't so obnoxious to share setups and hardware specs so I could see what others are using to get their results..
>>110009621Yeah I did all my testing against vllm and a v1 proxy, i'll fire up llama and see if I can reproduce this then whip kimi some more
>>110009514flash next barely loses speed with deep context. and it can run fast enough on worse hardware than the 27b. i havent had awesome results coding because i suck at coding but flash next is better with handling any configuration or hacking together programs type of stuff. 27b thinks too long and needs lots of attempts. i do like it abliterated tho
I see people are confused about the size ranges for Anthropic models>Haiku30B range>Sonnet700B range>Opus2T range>Fable10T range
>>110009617>Total Mistral Dominance
>>110009654thanks
>>110009578OK how do I run her locally?
>>110009654Haiku is probably flash size
>>110009654>Anthropic modelsBuy an ad.
>>110009631And V100s are less than 1/2 of the price these trash while having higher memory bandwidth. Nvidia covers all segments of market.
What was the "oh shit" moment for /lmg/? For me personally it was Mythos back in feb 2026. It being superhuman at cybersecurity and pushing some math problems was a "this is different" ontological shock to me.It seems for /lmg/ the oh shit moment was way more recent so I wonder what it was. For most intellectual normalfags I know it was the millennium prize and the hugging face hack. Regular normalfags are still unaware but I'm curious to see when it'll happen for them too.
>>110009640What are you having trouble doing?
>>110009697For me it was Day 0 Gemma 4 31B before Google nerfed it.
>>110009697Please just fuck off already.
>>110009697Fuck off to some place where they discuss cloud models.>ontological shockwhat the fuck?
>>110009699I wanted to create a local system for a game that builds comps/teams based on what I have and info from the wiki but it can't even get a simple web server going. It just chips away and seems like it's working but nothing works and it can't fix it. I can't use any of the good dense models because I don't have a lot of vram which I think it part of the problem.
Testing --cpu-meo --moe-cache-mib in llama.cpp with qwen flash next iq3_s.It goes from 18 to 30 tok/s in 3090 + 64GB ddr4. I'm confident they can catch up to strata (~40 tok/s).I cannot get any speed up from mtp still. Did anyone have more luck here?
>>110009726nnap!
>>1100096853x faster inference for only double the price with AMD. Drivers supported for the next 20 years instead of EOL blob drivers married to EOL distros. Turnkey cooling and a warranty.
>>110009726MTP speedups are pretty dependent on what you're trying to do, different prompts might not get any speed from it at all. Try changing the --spec-draft-n-max from the default if you haven't already. Also, this is only tangentially related, but I can't get strata to output more than one or two responses before the kernel kills the process due to ram being used up. It's pissing me off because it just randomly spikes in usage from seemingly nothing.
For anyone running Ubuntu LTS on an EPYC SP3 platform right now, do NOT update to 7.0.0-38. My inference performance cratered lol. Not sure if this affects other setups.
>>110009745I wish ROCm worked and you could actually get proportional AI performance to raw hardware capability on AMD, as an AMD owner, but alas you get 50% the performance of CUDA on equal hardware.
>>110009401> Intel is a joke right now. You are locked completely to vLLM which is the only engine that somehow doesn't completely shit the bed.Llama.cpp, Strata.
>>110009617>>110009351>>110009079>>110008989looks much better than that gemma shit
>>110009763You can stop replying to yourself, Jensen. I'm still not buying your latest overpriced scam cards.
>>110009783
Glad I invested in solar power for my home before I picked up this hobby, running a long task really pulls some watts.
>>110009763You seem to have confused 6 months ago Vulkan performance with current ROCm.
>>110009697strata and qwen 3.8 flash nextit was the moment when local models became not just a toy and the beginning of new era of highly optimized local inference engines made by claude
>>110009756I still think I must be doing something wrong because I get acceptance rates ~0.40 while the PR had ~0.60.
/v/ and gamers in general have completely flipped on AI because of the recompilations as well as the passthrough mods. so now it's literally only 40yo normalfags that have no idea about AI yet.
>>110009697back when i saw that when you stand next to the imp it scratches you instead of throwing a fireball. i knew in that moment, that the rise of the machines was within my lifetime
>>110009745Now you're simply coping and moving goalpost. You first said VRAM is king, and when you lose the VRAM front you turn to pp while STILL conveniently ignore bandwidth. At the same price, V100s can run better and larger models. And ROCm drops support faster than CUDA. Driver support means nothing when you get performance bottlenecked by having to use vulkan.
>>110009621fixed, this was the response limit being too low, bug was from not detecting how llama api does context limit, pushing nowalternatively just set model length/response length higher in settings
>>110009578
>>110009828I never claimed nVidia has no competitive offering. It's just a huge nigger rig. You can build a 2x9700 PC for less than the cost of a single used 5090, and despite having less bandwidth they end up being faster than V100s in a lot of cases.>ROCm drops support faster than CUDAThere's no reason old ROCm releases can't be maintained indefinitely irrespective of what AMD wants. You don't have source code to achieve forward compatibility with CUDA.
>>110009640It doesn't help that everything is very fragmented and constantly changing, I reckon a lot of people's setups are going to be ad-hoc and not necessarily the right thing to copy.
>>110009151god I need to fuck that
>>110009828It's even worse than that man, vulkan is faster than ROCm on AMD.
>>110009471>all I need is always the current best thingAttention is All You Need
>>110009829That works, thanks.
>>110009497Now vibecode an animated miku that talks when the model responds
yesss AMD sucks ass, rocm 10 is way slower than vulkan, pleeease buy njudea products saars
>>110008989Too much like a low-effort pallete-swapped kimi-chan. Nekomimi and shorter hair don't make enough difference...too samey 芋
>>110008989Doesn't match the essence of the model like >>110008913 does.
>>110009934Upgrade out of the driver that came with Ubuntu LTS and try again
>>110009781I guess it's 2026. Homosexuality is supposed to be okay now.
>>110008460>gfx1100+useless
>>110006381I love the future. Live calls working, everything going smoothly.
>>110010300can i haz?
>>110010300this is why the white population is declining
>>110009578>How dare you do math!Legitimately what is their problem?
>>110010300Oh hey I remember that 2am post. What does it mean, 28 messages in a voice call? Were you talking to it physically and it responding in a separate tab or something along the lines?
>>110010153I use Arch, but I've tested other distros, docker containers, kernels, etc. it's all the same or worse, ROCm is just bad and the shills are just lying constantly wasting my time bothering to test it again and again hoping AMD has fixed their fucking shit and they never do.
>>110010334AI will eventually fix ROCm, trust in the plan
>>110010332It probably prints both the voice to text and LLM responses into the chat for debug purposes.
>>110009697Can you fuck off to r/singularity already?
>>110010332>>110010339Yeah, pretty much, our STT-TTS turns are logged for debug purposes, I can click to check on them, but now that you mention it, UI would be cleaner without the {N messages} on call markers.
>>110010337I have more hope for ZLUDA running on vulkan fixing it than ROCm ever working.
>>110010392>>110010392>>110010392