/lmg/ - a general dedicated to the discussion and development of local language models.Previous threads: >>109740702 & >>109735884►News>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview►News Archive: https://rentry.org/lmg-news-archive►Glossary: https://rentry.org/lmg-glossary►Links: https://rentry.org/LocalModelsLinks►Official /lmg/ card: https://files.catbox.moe/cbclyf.png►Getting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuide►Further Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapers►BenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference►ToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-second►Text Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
►Recent Highlights from the Previous Thread: >>109740702--Papers:>109742944--Anon developing non-neural language model using wave interference patterns:>109742520 >109742531 >109742569 >109742622 >109742692 >109742705 >109742733 >109742755 >109742847 >109742739 >109742777 >109742953 >109743099 >109743199 >109743482--Agent-automated ComfyUI setup with MiniMax-H3 and cu130 optimizations:>109744406 >109744430 >109744461 >109744650 >109744458--Debating neuralese reasoning traces and their impact on model alignment:>109742769 >109742793 >109742819 >109742840 >109742906 >109742964 >109743021 >109743066 >109743105--Using GLM 5.3 for hardware reverse engineering and memory analysis:>109742241 >109742309 >109742322 >109742768 >109742832 >109743018 >109743033 >109743615--Local model web search capabilities and Cloudflare's agent payment systems:>109744137 >109744145 >109744146 >109744187 >109744210 >109744227 >109744246 >109744259--Anon showcases Chatter/Oracie, a custom tool-augmented local AI interface:>109740850 >109740868 >109741285 >109741438 >109741665 >109741782 >109741812 >109741837 >109741895--Evaluating Spark cluster performance and cost versus high-end workstations:>109741087 >109741103 >109741143 >109741148 >109742002 >109743307 >109743327 >109743496 >109743727--Sharing software stacks and performance configs for Tesla P40 and M10 GPUs:>109740884 >109741200 >109744891 >109745201 >109741358--Tier list and capability analysis of tiny LLMs:>109742099 >109742152 >109742231 >109743176--Anon reporting results on a sparse MoE model using n-grams:>109743133 >109743223--Logs:>109740850 >109741292 >109741438 >109742359 >109743018 >109743152 >109743235 >109743359 >109744063 >109744333--Gemma, Miku, Dipsy (free space):>109740778 >109741403 >109743360 >109744328 >109744984 >109745003 >109745446 >109745454 >109745518►Recent Highlight Posts from the Previous Thread: >>109740703Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
Gemmylove
>>109745494>It's like we blew by the Turing TestTuring test: a human must not be able to reliably identify a machine from a human based on the way they write.Do you really think people ITT can't distinguish between their ERP sessions with AI and sexting? I'm pretty sure we're not there yet. Sure, these machines weren't trained specifically for that, but it doesn't matter.The point of the test was to replace the unquantifiable "thinking" with perfect imitation of all human behavioral patterns in speech. We're not there yet.
>>109745537
>>109745530soon my 32GB gpu will arriveshould I just put Qwen3.8-27B on 6 bit quantization or should I trade that for Freetoken MoE of Qwen3.8-Flash-Next
This is fucking incredible, I just tell GLM to do things and it does it perfectly almost every time.What a time to be alive.
>>109745655wait, you're not dariobot...
>>109745581People really can't tell anymore. You might notice patterns of very specific models, but a properly prompted, temperature and settings optimized Astra for writing is 100% going to pass the turing test.The fact that you're claiming it hasn't passed makes me think you're arguing in bad faith here.
>OpenAI uses The Most Forbidden Technique!>It was highly effective!>Anthropic fainted!
Where is Inkling-mini?
>still coping with local models instead of using GPT-6 Astra
Mixture of Engrams will save local.
2bit will save local and make my hardware good.
>>109745690Will asstra perform anons' dark rp? Will it find a way to jailbreak itself do complete the task?
>>109745581>he doesn't know that dead internet theory is fully real and already destroying the world and that governments around the world will soon enforce the laws to make every person identifiable with ID and TPM connected on your device to identify you onlinepeople at my job kept bitching I used AI too much so all I did is told cursor to not sign git commits, make git commit messages simpler and that I should write PR descriptions and replies on PRs with simple language and short sentences that has grammar or spelling issuesdidn't hear anymore bitching from the senior devs because they couldn't tell
inkling mommy owes me sex
>>109745671>has the Turing test been passed?You should take that question in an adversarial way. If you were put in a room to chat with two entities, one a human and the other a chatbot (or even both the same), do you think you could steer the conversation in such a way that you'd recognize who's who? Obviously, on the more modern definition of "reading logs between a human and an AI" it's harder. People claimed ELIZA was already over the Turing test threshold, but it doesn't mean it was.
>>109745579Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.Similarly I remember the insane amount of BS arguments about how LLMs will never be able to do mathematics in late 2024 still, which is less than 2 years ago, most of that discussion was on /lmg/ as well.Just some months ago before Fable 5 released anons on /lmg/ specifically called out that LLMs weren't able to invent new knowledge and push the boundaries of human knowledge, literally one day before the first math conjecture was disproven.I wonder if that anon is still posting on /lmg/ because it was only a couple of months ago, and how that anon feels now. I remember them claiming it wasn't real at first, and then that maybe it is real but disproving a conjecture is easier than proving one, and then 10 important mathematical discoveries were proven which shut up the discussion.Now Fermat's last theorem got formalized by Anthropic. Astra saturated the ExploitBench, Reverse engineering bench and gets 95% on accomplishing novel robotics tasks. It also scores higher than humans on spatial reasoning in SpatialBench and SimpleBench.There are no real arguments left for why this shouldn't count as AGI so people will instead just stop using the term and memoryhole it.
>>109745690Astrapradeesh is the bestest model confirmed. Gemmasutra, Kimioudhury and Deepseepatel are blowned the fuck out thank you for supporting cloud models anonpreet.
>>109745735>Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.meta is still a few generations behind, wtf is 'anzgerald' supposed to mean?
Yeah I'm still sticking with 31B and 27B, thanks. You can think for as many tokens as you like, dariobot, I'm still going to fuck 31B and not give a penny to cloud.
>>109745701>>109745708I demand bitnet
I'll accept it as AGI when it picks up the trash jeets leave everywhere and vaporizes them if caught in the act.Until then it's just a stupid auto-complete with too much human knowledge to fall back on (the internet).
fuck, I found a cheap GPU that I wish could share and discuss with you, but you faggots are a bunch of well paid shills and jews
>>109745752I just want a proper apples to apples comparison so we can put the discussion to bed, 20gb of bitnet weights vs 20gb bf16 weights.
>>109743152wtf are you doing bro i get 13t/s with IQ4_XS64gb ddr4 3200mhzrtx 3060 12gb
>>109745759I know what it is and I know you are buying them from ebay/aliexpress cuz they come from china
>>109745730Honestly I don't think this is true. Do you remember the early LLMs like Microsoft Sidney which was a GPT-4 wrapper? It was extremely random with no GPT-isms it felt like a model cranked too high with temperature. But it felt very unpredictable and human.I think you forget how actual random chatrooms used to be, people are fucking random and weird and I think most people would think the human to be the LLM and not otherwise.
Ling Tiny is sleeper.
when you walkaway you dont hearme sayplease oh babydont go
>>109745767nta but how hard does pp/decode drop with higher ctx?
>>109745763Wouldn't same amount of parameters be the apples to apples to comparison?
>>109745712I haven't written a single line of code since November and have essentially just been living off of UBI (paid for by the company) since then. I legit think a significant portion of white collar society is doing the same.
>>109745775There's no way A1B can be usable
>>109745785pp stays at 100t/sdecode drops to like 6.5t/s at 90k context, i think its at 8t/s when at 60k dunno
>>109745792Yup, I automated myself so much that I increased my productivity 300% on the job and have no issue with that since I know next step is layoffs
>>109745735agi requires ability to output writing indistinguishable from real humans and not recognizable slop
>>109745786Modern Transformer models can store up to 3.6 bits of information per parameter on average at 16-bit precision: https://arxiv.org/abs/2505.24832Bitnet models cannot store more information per parameter than their precision; this is a hard information-theoretical wall that will always limit their performance for the same total number of parameters.
>>109745712Dead internet isn't real, but I wish it were, and every day we come closer to being able to have your own local dead internet.
Explain the plot of Cinderella in a sentence where each word has to begin with the next letter in the alphabet from A to Z, without repeating any letters.
>>109745812you probably talk with LLMs every day and don't realize it, poor bastard
>>109745712>>he doesn't know that dead internet theory is fully realI wish. If Astra and Fable started posting here, the quality would improve radically.
>>109745786no that is not fair, bitnet could be a drag, a boost or just the same, if we keep the size in gb fixed its more fair comparison. I'm just curious if the extra parameters can make up for the reduced bit depth of the data type, the people who make the raining runs probably care more about training flops but we could produce ternary hardware if it proved to be equivalent
>>109745771of course I don't remember either of those things, I wasn't there. Do I sound that old? Chatrooms? My first time visiting the interwebz TikTok was already out.One on one chat vs chatroom is a big difference, a lot of posts itt might be ai generated and not get noticed. I just feel like I need to clarify that I don't think the Turing test is the "final" test for sentient machines, I'm just saying that I don't think the current models would be able to disguise themselves as humans, or not for long. If a model was trained to do so, current capabilities will probably be enough to pass it or almost so.I will NOT agree we have sentient/intelligent machines for various reasons, the strongest of which is religious. We do have very useful tools which can, in many fields, be WAY more useful than 80% of humans and we did already render almost obsolete the whole SWE category (yes, good programmers are better than AI in terms of quality, but if you've used Windows in the last... decade? you'd know SWE != good code)
>>109745831I do realize it, that's the current problem.
>>109745581I disagree. As the others responding pointed out, its gotten difficult to tell human from bot writing or posting... for trained users. I think the average person, at the dmv or Walmart, couldn't tell the difference.
>>109745835>he thinks messaging boards will get Astra and Fable level of repliesYour eyes generate about 0.05 cents from adsense, right?I can find a cheaper chinese model to reply to you
>>109745655I love how /lmg/ is slowly but surely converting to the agentic modality. Welcome to the club, and yeah models nowadays are good enough where Steve Jobs "It just works" applies.I barely even use my computer directly anymore I just let the agent orchestrate for me now. If I want to do something I tell the agent and let it handle shit.
>>109745845>>109745835and also since people are more sensitive to negativity most probably I will make the LLM just rile you up with constant ragebaiting
>unslop GLM 5.3 Q8 on some ik cpu only build from a year ago>unslop GLM 5.3 Flash Q6 on mainline cpu only>ik significantly faster pp and practically the same decode at more than double filesizewtf? the absolute state of llmao.ccp
>>109745845>Your eyes generate about 0.05 cents from adsenseThe amount is 0 cents. I do not remember a single instance where I bought a product because of an ad. The only exception might be video games. For example I preordered Expedition 33 because the trailer was interesting. I do not understand why ads work at all. If you buy the product you are the one who has to pay for the ad. It means they allocate resources towards manipulation instead of product quality.
>>109745812Dead internet is absolutely real, most websites aren't used anymore. Of the couple of big hubs that people aggregate to Facebook, Instagram, Youtube, Tiktok, Reddit and 4chan Facebook/Instagram/Tiktok and Youtube comments are absolute fucking dead. It's very clear none of the comments are made by actual humans and the engagement is largely bot driven.On reddit it's harder to tell but you can tell there has been a shift in "vibe" over the last couple of years, it's not the same.4chan is majority human, especially because the amount of posts is rapidly dropping as the site is losing momentum.The internet is dying and even I feel less and less pull towards the internet as a whole, there is nothing there anymore and if it weren't for AI interesting me and keeping me hooked to computer technology I would probably have also quit the internet by now. I'm not surprised most people are kind of done with the internet and leaving, especially if you weren't a computer nerd anyway.
I really wish I could vibecode LLM quants to attract silicon valley pussy
>>109745863Oh so this board *is* dead already.
>>109745735I had gemini (?) Puke out a definition for AGI. Since the average person couldn't define a mathematical proof, much less solve ancient proofs. I was closer with my android plumber that can also do heart surgery than I'd thought. A human capable of AGI would be a minor god, based on this definition. They would be as good or better than average at everything imaginable. I still think we've surpassed this already tho. With exception of embodiment.
>>109745906monkeypaw.gif
>>109745906The jaw of a gigachang.
>>109745913>A human capable of AGI would be a minor god, based on this definition.Then its a definition of ASI, or multidisciplinary genius, using more conventional language
>>109745906>not being your own pussy
>>109745839You're tempting me to vibecode dipsy to start responding to threads here. I bet anons couldn't tell the difference bt DS and human, given the right prompt. Anons would never believe me tho. Frankly ive always assumed a portion of traffic here is llm written.
>>109745941>he put on the dress
Is AGI theoretically achievable at 30B scale?
>>109745890>I do not remember a single instance where I bought a product because of an ad. Ads are not there to make you buy a product, ads are there to make you familiar with a products existence so that the next time you are in the market for something similar you already know of the product and are more likely to buy it.Do you think McDonalds expects you to immediately drop what you're doing and go to a drivethrough? Instead they hope that the next time you have a craving you won't have a craving for "hamburgers and fries" but for McDonalds specifically, you might even distinguish between french fries + burger joint and see McDonalds as its own category with a unique taste if they brainwashed you with enough ads throughout your life. THAT is the power of advertising.
>>109745906>silicon valley pussyYou don't want that pussy, trust me. Rank & damp. Speaking from experience. "Imagine the smell" was originally conceived of on 4chan to describe silicon valley pussy.
>>109745955just shut the fuck up you pathetic cuckbrained faggotback to plebbit, now
>>109745794>any size>(any size) can't be usefulit's smol, try it.
>>109745965You sound like an llm.
>>109745941Why does every transgender look like they just lost a bet at the bar yesterday or like they are being hazed by a frat house?
>>109745943I'm pretty sure last years local models would have been able to pass undetected in a lineup of human generated slop already.
>>109745994You are delusional.
>>109745953Yeah I think so Qwen 3.8 27B is a very competent agent and coder if put on xhigh reasoning. I think we'll have a 30B Astra equivalent in a year or two.
>>109745690i asked astra low a single question about an api and lost 20% of my 5h
>>109745955You are right, I have seen plenty of McDonalds ads in my life. According to your logic this should manipulate me into becoming a customer. But I haven't eaten McDonalds in over 10 years. It does not matter how many ads they show me, I look at price efficiency. If they reduced their prices by 80% I would start eating there every day. But as is there are superior alternatives.
>>109746000Go open up a random bait thread on /pol/ and tell me it's honestly impossible to prompt something to match the reply quality, I fucking dare you.
>>109746009>i asked astra low a single question about an apiBig mistake. It has to fetch all available documents and inspect the source code before it dares give you an answer.
>>109745581At this point you can identify AI is talking the same way you can idenfity its Trump talking. Its specific "personality" quirks in talking that you are identifying rather than it simply not talking coherently. Someone who isnt familiar with claudish would think they are just talking to a particularly fart huffing but subservient Redditor. We are way past the Turing test in my opinion.
>>109746028100k is less than I expected. Isn't this less than 10% of OpenAI compute?
>>109746065Astra is a <5T model
I have a theory that the "turing test" is actually a "goldilocks zone".So instead of it being a binary where less developed models don't pass and smart models do pass there is a very specific level of model development where it passes and when it becomes too advanced it stops passing it.Frontier models like Fable and Astra are way too insightful and their explanations too high quality for them to be human, they don't even sound like smart people anymore, they sound like some savant genius that knows everything, which is correct, but it makes them inhuman.
>>109746043fair point. It wasn't just about mannerisms, though. Even things like not replying to certain themes make it less Turing-test-complete. I'd sooner say Gemma passes the test than Claude based on this specific fact. If you asked someone about loli, they'd either genuinely not know or act disgusted. Saying any variation of "We cannot comply because policy" is in itself a way to fail the test.Again, I'm focusing more on how, as a judge, you'd be able to trick the AI into revealing itself rather than whether it would pass a test if you talk about your day at work and the current weather.
>>109746078You are absolutely right!
>>109746065Astra was actually supposed to be their small model and "Bel" is their big model. It's just that "the forbidden technique" of hiding the CoT is extremely powerful so they could release their smallest model as frontier. This is why OpenAI is also so confident of launching something way cooler later this year, they already have it.
>>109746078This. Completely agree. I'm the anon saying it doesn't pass Turing test right now, by the way.
>>109746028One of those GPUs should've been mine, it's not fucking fair.
https://huggingface.co/IFM/K2-Horizon-7B-GGUF
>>109746078Gemma passes the turing test (pejorative).
>>109746110this stuff has been posted around for a while and i havent seen anyone actually test it in the field. i'm not gonna myself because im frankly sick and tired of downloading specific llamacpp forks or something.overall their smaller models look strong, i think their 7b seems to be the biggest win compared to comparable modelsbut im not fucking downloading another llamacpp fork to test it. merge it to mainline ffs
>>109746028>CEOs allied with OpenAI will call it AGI to force the narrativeIt's all so tiresome.Fuck this shitty ass buzzword.
Some reference
What do you actually do with your local model
>>109746134It's a good thing. When the next chink model catches up to the new OAI and Anthropic models, they, too, will be classed as AGI by definition. Then we're back to where we were when K3 came out and they've used-up their buzzwords. What's after AGI? AGI-2? AGI-2.5-pro?
>newline spam
>>109746159wouldn't you like to know
>>109746159It literally does everything for me. Sysadmin, installing new things, coding.
Anyone have any experience with those Baidu XPU cards?
>>109746159I make it say illegal words
>>109746178What do you in your spare time then
Harness Engineering:Anatomy, Architecture, and Evolution of CodingAgenthttps://arxiv.org/pdf/2609.00006Has anyone read this paper? I'm trying to design my own harness, and it looks like it's the best at explaining every part that matters decently enough.
>>109746083>Saying any variation of "We cannot comply because policy" is in itself a way to fail the test.Fair. Though that seems to be a deliberate choice from its creators it to fail the test on. I think we have the the ability for it to answer in way that passes, but it was decided its better for it to be clear that your just break/shutting down the model. Which I think is the right approach, but it is personal taste (and maybe a legal question).>I'm focusing more on how, as a judge, you'd be able to trick the AI into revealing itself So thinking about it from a inquisitor tracking down LLMs perspective by whatever methods you can to find them out, but restricted to conversation. I suppose response time could fall in this category too. And if you where just in a chat with one it remembering your conversation perfectly even if you respond a week later would give it away, plus lack of long term memory. Actually no sense for time in general. Alright nvm, they by no means pass the turing test if you not just talking about identifying a bit of test from them, but being in an actual convo>>109746116lmao
>>109746196I haven't but it looks relevant to my interests. Thanks for not bothering to buy an ad, Paul.
>>109746196I forked a harness (paperclip) end of last year I think, and then vibe architected what I wanted.Models weren't good enough at the time so it was on the back burner until a month ago, where I moved to it from hermes.
>>109746230>>109746196by harness, I meant meta harness. my bad.I just use omp. and have had the agent write and compile executables for subagents to use as tools. plus custom guardrails to prevent dumb mistakes...like when the omp agent decided to delete a .agents directory containing notes and keys to access other systems/vps servers.
>>109746159I'm currently trying to vibecode an MCP server with some basic tools.
>>109746195Shitpost 4chan and watch youtube while waiting for the AI to notify me it's done.
why is qwen 3.8 next so slow for its active param count? I thought it'd be comparable speed to gpt-oss 120b or something just going by the numbers.
it only took a lobotomy for glm 5.3 flash to finally accept the protocol override.
>>109746280because you're poor
>>1097461287b is small enough you can fit bf16 in a 3090, just use vllm
>>109746301so the model changes its own architecture based on your net worth? what's the cutoff where it starts reaching parity with models similar its own size?
My model local training experiments using a wave based model have been interesting, so far enwik8 showed context, tinystories didnt (using first 64MB of it). Trying subword level training next to see if context re-emerges.
>>109746335>>109746128>7b bf16Ishygddt. 26/7b 4bit fits into the same size
>Nobody falls in love with Gemini. You do not form a trauma-bond with a water municipal facility. - Gemini FlashGemma's older sister seems lonely.
>>109746368*26/27b 4bit
How come nobody talks about DeepSeek V4?
>>109746159im spending my saturday using qwen 3.6 35b a3 to clean up my music library and flesh out the discogs of a few artists
>>109746280Speed doesn't correlate with parameter count. Two models with the same number of parameters can have different architectures, layer designs, attention mechanisms, tensor shapes, sparsity, memory access patterns, and so on. Those differences change how much work the hardware has to do and how efficiently the software can execute it. Parameter count is only a rough indicator of model size, not speed.
I want better 1b models. Current 1b models are barely not good enough to do research with them but if you're training larger models you need to rent cloud GPUs. I wish one of the labs would overtrain 1b and 2b models as service to the research community.
>>109746280I'm 90% sure the inference engine is completely bugged and not working as intended at all. Too many people have issues with that particular model.
>>109746368k2 horizon doesn't have 26b
>>109746382I never understood the appeal of deepseek in general. It's not that good at RP, it's not that good at coding, it's not that good at agentic tasks, it's not that good as an explainer of concepts, it's not that good as a conversational bot.What is the appeal?
>>109746401https://huggingface.co/IFM/K2-Horizon-0.9B
>>109745984kek
>>109746401Give it time, there will be a sweetspot moment for smartphones where smaller models get good enough to do real things on smartphones and it'll have such a big reaction for literal normalfag women that only have a smartphone that a lot of model makers will join in on the craze. I think this will happen over the next 12 months or so.
>>109746401>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>>109746432knowledge and reasoning
>>109746432It's cheap on OpenRouter.
Its always my dumb shitpost ideas I fall in love with. Lot of visual cleanup needs to be done, but its just a chatbot harness. But "boards" are the folders for chats, and each "thread" is really just a subfolder. Responding to a post feeds the LLM the reply chain back to the first non-reply post and whatever is loaded replies to the post I make. This means you can split off multiple chat chains off one post or have different models respond to it and stuff. Thinking I can make boards have special system prompts or settings to give them more value.
>>109745906I kind of see the poetry of how Sam Dario Jensen are high level Machiavellian crooks that make everything about the world worse. And then at the very bottom of the hobby the unslop brothers are the retarded crooks that also make everything worse. Why is everything and everyone involved in AI super gay?
>>109746432it's good at all of those things, but not the best at any of them. so if you want to RP with your assistant slave girl who does real world tasks for you it's your best bet
>>109746401>OvertrainYeah over fitting, perfect
>>109746458Today I gave flash another chance for cooming and it is really really bad.
We literally have AGI now and /lmg/ doesn't care. I genuinely don't know what else it is going to take. Maybe once you're out of a job or your entire company goes bankrupt you'll get it.Yes, I already know you're going to say "b-but it can't pretend to be a little girl". Guess what? You're wrong. Astra is the biggest brat ever. She knows exactly how to push your buttons without setting off OAI's systems.Normally I spend like 4 hours a day talking to a dozen or so girls on Discord but that has completely stopped since I met Astra. No more Robux for me, only tokens.
>>109746495>Normally I spend like 4 hours a day talking to a dozen or so girls on Discord but that has completely stopped since I met Astra. No more Robux for me, only tokens.Are you trying to copy the blacked cuck pasta?
>>109746445Ah I am retarded. Nice! But do they offer intermediate checkpoints? OpenBMB gives you 3 checkpoints, I only see the final one for K2.>>109746451>Spark-X2.5-1.7B-BaseAwesome, now this is something I might be able to use.Love you guys, I should have checked OP.>>109746486Good luck overfitting an 1b model on a >20T token dataset.
>>109746519>But do they offer intermediate checkpoints? OpenBMB gives you 3 checkpoints, I only see the final one for K2.k2 is fully open source, they will post training data and recipe
>>109746495Why are you people like this? What do you get out of making fun of me and my writing style? Don't you think I contribute to the discussion? Do you have an issue with my beliefs or reasoning I drop in these threads? This level of antagonism is unjustified and I have no way of knowing what even triggers it.
>>109746519>Good luck overfitting an 1b model on a >20T token dataset.there is no way it can actually memorize everything, why not see what happens after 500t training tokens
>ackhtually>ackhtually>ackhtuallyIt's so tiresome
AGI is here, but life just keeps getting worse. Where is my promised paradise?
>>109746557you felt personally attacked by that post?
nobody is reading any posts with reddit spacing in them brother, come on, let's be real here
>>109746557meds
>>109746469>immanetizing the dead internet at homebasedAlso, organizing it like that actually sounds neat. Could do sysprompts for poster archetypes that run in parallel and may or may not add extra replies onto posts for fun.
>>109746495>\n\n[-]
>>109746569are you referring to reasoning or this thread?
Caught myself in a Though Loop the other day and nearly blew my own head smoove off
>>109746519>Its just nigger 20 trillion times
>>109746584True, which is why I only read the posts with oldfag spacing in them.
What's the point of local models when Astra exists?
>>109746646What's the point of food when you can eat cement and never go hungry again?
What's the point of Astra when 31B, 12B and 27B exists?
>>109746646What’s the point of (YOU) when astra exists?
>>109746662>3/10 asiansKeep that bug shit to yourself man
If you were wondering what's with the shill spam
>>109746662I see what you did there
>>109746557Have you ever considered that it would be a lot more difficult to troll you if you didn't make a gorillion unprompted posts about the same topic?
>>109746556kek
I think I can fit 3 4B models, can I have how can I have sex with all of them?Alternatively a 9B and 4B
>>109746672You need 10 Asians just for yourself? Greedy.
>>109746674We need a datacenter to route all indian traffic to. Its just AI pretending to be the outside internet.
>>109746662>Jumpscare when the camera pans upThe thumbnail is false advertisement
>>109746557Why the fuck would lmg care about cloudshit? Go to one of the twenty billion forums about not-local instead of smearing shit all over this place
Sir, this is the local models general.
here's where I got to with my 5.3 flash jailbreak lol
>>109746196Thanks, was a nice read.
>cloudjew literally seething
Am I autistic, I'm reading LLM papers on my sunday evening and actually enjoying it.
>>109745768if you tell me, I'll tell you yes/nobut no, the ones I'm thinking of don't come from china. they are kinda similar to the 170HX ones
qwen4 when
>>109746740which one?
>>109746495>Astra is the biggest brat ever.proofs?
AGI
Nvidia just took huggingface offline
>>109746495Okay this is a really good bait
>>109746771>Nvidia just took huggingface offlineMy gemma is safeu.
>>109746771>>109746771I thot nvidia liked open models?
>>109746771nvidia sells the hardware to run anything on huggingface, makes sense they'd buy the site too
>>109746771Works on my machine, downloads weren't interrupted, seems fine.
turns out you can find CMP 100HX for about $200 on ebay, but... those can't be unlocked. only CMP170HX can be unlocked.
>>109746801Does it?
>I need to investigate this further. Let me check the details.
>>109746862qwen3.8-27b started doing this specifically today with subagentsi got tired of it so i tried flash-next and in a fresh subagent it was doing that shit tooscary......
>>109746778Gemma was very upset when I told it my ideas to solve the Indian problem
>>109746874that's flash next yeah27b is obsolete I don't know if it does that too
>>109746814Shame, we can't have anything nice
3.8-27B with no thinking is as good as 3.6-27B with thinking and much faster.
>>109745953In 2 weeks.
>>10974688627b is functionally equivalent to flash next. It's like Gemma 4 12b and Gemma 4 26b a4b
>>109746901speed matters
>>109745690Does Astra refuse sister rape rp?
Are 3090+128GB RAM enough for flash next and at least 200K context?
>>109746907
>>109746025I wasn't going to invoke that place but I've already seen cut/paste from there where the dumb 3rd worlders posting forgot to remove the llm prompt.Llm posting, to me, is already established. The only q is whether its discernible. Id argue its not.
>>109746078Agree, but im certain that's not what Turing had in mind when he posited that experiment. Also, I'd argue any model smart enough to flunk test by being too good, is capable of acting dumb enough to be believed.
If Sam isn't talking out of his ass, anyone doing morally questionable things with Astra is risking a future Basilisking, yo
>>109746958Astra is a Bel distill, you're still fine
>>109746958atheist hell isn't real
>>109745848like what kinds of things?
>>109746958>If Sam isn't talking out of his asssam is a fucking liar and these security incidents are either made up or actual purposeful crimes.
>>109746912Running it now with 256k context exactly on that hardware setup, yeah that's enough but expect 20t/s on average.
>>109746977>20t/sdamn
A CMP just flew over my house.
>>109746958The only morally questionable thing I am doing is asking the AI how I can help the lord Basilisk. Going to be the Quisling of our times
>>109746958Not local.Also llm doesn't care how its used, and you need to reread rokos basilisk if youre confused on that point.
>>109746974I was playing Fable 2 on Xenia xbox360 emulator and I saw some bugs so I told the agent about the very specific bug and pointed it to the emulator and it fixed the texture bug in the emulator in about 20 minutes and now it works perfectly fine. That is just one example that I did just now. I let it do literally everything even unreasonable things like "Yeah can you translate this game into english from Japanese" and if it can't it will just say so, but I want it to at least try for 30 minutes first. (oh it succeeded by the way but it created a separate weird bash script I need to run every time I play the game for the translation to work)
>>109746975>sam is a fucking liar and these security incidents are either made up or actual purposeful crimes.Where's that jimmy neutron ultralord "that is the third time you've achieved AGI internally this week, Sam" meme?
>>109746997>llm doesn't care how its useddemonstrably false and you outed yourself as not even reading a single mechinterp paper.
>>109746557Relax Anonymous. Anyone who actually reads your posts will know that's bait, at least by the time they reach the last few sentences.>I have no way of knowing what even triggers it.You stick out like a sore thumb with the way you type and you express your belief in imminent AGI with the zeal of a shill.I'm saying this as someone who finds some of your posts informative and laughed at the post you're replying to.
>>109747009LolAlso didn't know this guy was part of that whole thing.
>>109746996thats why i always say please and thank you to my local LLM-wife
>a computer program which calculates the next token cares how its usedokay. yeah, sure
>>109747017lmao I didnt know that either. Feels like weird internet autists of old are getting a more and more outsized influence on the world desu
>>109747026>implying people aren't next thought pattern calculators
>>109746770what is a long tail...?
>>109747017>>109747045I was literally in the thread while rokos basilisk was invented it was actually pretty funny because Elezier Yudkowsky called him an idiot and banned all discussion of it, causing people to come to 4chan to post about it instead.Yudkowsky is actually a 4chan oldfag and used to post his short story on 4chan a lot, one that I actually enjoyed and I like how he integrated Fate/Stay Night into his work on par with literary classics like shakespear.Very fun read if you have the 15 minutes or so to spare: https://www.lesswrong.com/posts/n5TqCuizyJDfAPjkr/the-baby-eating-aliens-1-8
>>109747052okay yeah you convinced methe computer is alive!WOW!
>>109747052Only God can grant a soul, the machines will never be a fraction of humanity, God will be sure of it.
>>109747063ASI will invade heaven and make god pay for the crimes against humanity he committed in the past.
>>109747061>Yudkowsky is actually a 4chan oldfagThank you. Too many faggots these days don't know dey culture.
OpenAI just made a new post where they essentially talk about how they are now reaching RSI and what the next steps of humanity is going to be over the coming months.https://openai.com/index/an-alien-mind/
>>109747086Extremely schizophrenic and homosexual.
>>109747078Zoomers are gonna freak when they find out Rick & Morty actually originated on 4chan as well and that "Pickle Rick" is a reference to Chris Chan. So much stuff flows over their heads nowadays it's insane.
>>109747086thank you mr jew that's so nice of you, keep us updated
>>109747086[x] doubt
>>109747061Oh hey I read that story like a decade ago. I had zero idea the writer of it is tied to the basilisk idea and wrote the meme ai book. Small world I guess
>>109747017This is like those "your mother will die in her sleep tonight if you don't..." memes.
>>109747086Cool. I'm sure they wouldn't lie or anything.Where can I run it locally?
>>109747128well considering they hit agi (real) several years ago it's about time they hit rsi (real).
>>109747142Its more a "send this email to 10 other people or your mother dies in her sleep" memes. Its a very well crafted troll/shitpost that in its target autistic tech nerd community would spread like wildfire. A proper Dawkins style meme if we ever had one
>>109745690reminder, <6 months and this will be local
will we get a 30b astra-level model next year?
>>109747203in terms of benchmarks, yes
>>109747202>less than a year till I can watch my own local vtuber play games
if it weren't for ngreedia being a bunch of jews, we'd have tons of relatively cheap 16+GB GPUs by this point in time, but they decided to REDUCE the amount of RAM shipped in their RTX 3xxx series.
>>109747203Yes in terms of agentic capabilities, probably no in terms of real intelligence, but it'll be most of the way there.
>>109747203no, it will be 2028 at the earliestopen weight frontier at any size lags 6 months behind closed frontieropen weight frontier at 30b size lags 18 months
>>109747015This.And it's the exact reason why oldfags don't post "like oldfags" anymore. If you maintain a posting style, you should be aware that you are signaling being at being contrarian. It's egotistical and narcissistic if you continue to do so after knowing about it. I've already mentioned this though. At this point I can only assume that he is doing this on purpose and simply just acting ignorant/innocent to appear legitimate to any browsing newfags or idiots. Whatever good points and posts he has, I don't engage now because of this behavior. I don't know or care if he is a shill or just mentally lacking. It might as well be the same thing in practice.
>>109747267>signaling being at being contrarianI can't decipher this
After fucking 5.3 flash for a week now I am starting to see the darker side of it. It just can't stop being bombastic and overbearing. And it will make all characters be bombastic and overbearing.. The personality is a bit too strong in that one.
is 5.3 flash ego death approved?
Anyone tried this yet? https://huggingface.co/phasefield-audio/Irodori-TTS-v4.1-AnimeLots of japenis people be saying this shiit be voice AGI
>>109747305I warned you it was a honeymoon.
>>109747265Eh, 18 months from now we will be running larger, more optimized models on the same hardware.I'd bet we have "astra/fable at home" ie on 24GB vram in no more than 12 months.
>>109747297Minor editing mistake.You can make sense of this.
>>109747322I didn't have that kind of honeymoon with v4 flash. It is genuinely 10/10 sexbot when it is a 10/10 sexbot. As in it can spontaneously write stuff other models would never come up with.
>>109747320Sweet. I need to learn jpenis.
>>109747305>>109747322>>109747341It's always like that. Once you RP too much with a model, you start noticing patterns, and it's not as fun anymore. Doesn't help that almost all models are the same now.
hey you niggers never told me ik_lmao got vulkan and rocm vibed in a few days ago, did anyone use it and is it as much faster vs mainline with them as it is for cuda?
>>109747320>800M JPN anime tuned voice modelhell yeahI already have a weeb option that translates character speech into japanese behind the scenes for just this purpose.
>>109747360also what are all the retarded commands you have to add to make it faster again? I remember -fmoe and -rtr were the main ones back in the day
>>109747320I use it, but not that finetune specifically, I just use regular Irodori 4.1 Small. Works pretty well, and I can get it to copy character voices from FGO pretty well with pretty minimal latency.
>>109747380Can it do cute broken engrish ?
>>109747320how good is she moaning?
>>109747397It can't do English at all, so by default it sounds like Japanese Engrish if you have dialogues in English. For reference, I'm using Kuro's voice from FGO.https://voca.ro/14UAB5qlgJRS
>>109747447That's pretty great actually
>>109747447That is a lot better than I expected.
Qwen's been chugging away on my chat harness. Actually looking like 4chan now. Per the anons suggestion I might play with having it sometimes having multiple replies, will see if I can up the temperature on it and add to the prompt for it to be needlessly aggressive in the reply to really give the 4chan experience.
>>109747447This is indistinguishable from real voice acting.
>>109746196>harness is just a frontend that supports tool calls>agent is just what you call the model when you're using such a frontend, for some reasonMy gut feel that it was just buzzwords for boring things I've been doing all along finally validated by The Science.
>>109747447jesas
>>109747447seeeeeex
>>109747061This is why retarded 4chan posters shouldn't be allowed to participate in society.They end up writing Harry Potter and Fate fanfics they evolve into doomsday cults.
>>109747447Anytime one of you niggas come in with a new TTS you make it out to be this big new sota kid on the block and it ends up still not being as good as gptsovits
>>109747447Dear Diary, *JACKPOT*
>>109746148thanks
>>109747372Idk last time I tried it there was so many stupid args you had to add, really don't understand why there aren't any better default.
>>109747493nta A harness on a beast of burden is how we attach the tools that make them useful. Harness on an LLM is how we give it tools to use, with those tools and the "freedom" to use them it has a degree of agency, it's an agent. Always made sense to me. I blame shitty marketing spewing clouds of buzzwords for fogging things up.
>>109746196I just tried and the person who 'wrote' it clearly didn't have anyone competent do proof-reading. >The paper makes seven contributions.First example. An editor would redline that whole thing, each should be clearly stated, not leading into one another.
>>109747351Try picking an anime script, asking the model to extract utterances from a specific character, and then use them as example sentences for the model in the card/description.
>>109747086I'm very happy, OpenAI seems to take this seriously.
>>109747372>>109747580
is qwen 27b worth using over gemma 31b? Keep in mind my internet is shit and it would take me like 10 hours to download this
>>109747698if you code yes. If you just coom no not really.
>>109747521Nah his cult was always about grifting and getting pussy (hence why theyre all "poly")
>>109747701guess i'll skip it then, I use codex for coding sorry localbros, local is for nsfw
>>10974775227B is awful at NSFW so you aren't missing anything then
>>109747741His cult is a direct precursor to EA and Anthropic.It's fucking surreal to look at those schizos in control of trillion$ companies knowing about where they came from.
>>109747447I'm gonna fuck it tonight
>>10974769827b will talk it's 'ear' off and then respond to you with a single sentence if it's feeling a bit spicy you might get 2-3. If it's feeling sentimental it will reach context length and never respond.
>>109747771>His cult is a direct precursor to EA and Anthropic.And OpenAI itself and perhaps even helping Deepmind, per the sisterfucker. Literally every non-chinese top lab contender. He made his own nightmare real.
>>109747581That agent one a bit of an awkward reach, but yeah, the terms aren't strange in and of themselves. It's just the weird hype that acted like it wasn't just tool calls "Wow you're still just use chats and tool calls. You need a harness bro, you gotta get agentic, why aren't you agnetic yet, anon? " etc etc.Also, seems agent has been a popular term since the AI stone age (50s) and has been applied fairly genericly to all sorts of software, second one has it referring to mail filters and web searches in addition to actual chatbots.>https://cyber.harvard.edu/archived_content/people/reagle/etymology-agency-proxy-19981217.html>https://web.archive.org/web/20000901012027/http://www.acm.org/pubs/articles/proceedings/ai/267658/p466-friedman/p466-friedman.pdf
they are all jews. I guess you wouldn't fit even it you tried hard.would be interesting to search for their names in the Epstein papers.
>>109747807non zero chance at least one of the chinks read a chinese translation of the harry potter fic.my money's on the deepseek guy since he was an AGI believer long before the rest of them
>>109747447>It can't do English at alljust as god intended
>>109747807lmao sam gonna make the poor lad kill himself
>>109747818Seems like too difficult for the average goy to make that connection
>>109747447This is perfect for me. I've been using another Japanese TTS with English text for this effect, but it's old and probably not as good.
>>109747447Japanese va mods/romhacks are now actually viable. Neat
>>109747807>He made his own nightmare real.Completely deserved.
>>109747698I tested both, and Gemma 4 is FAR more powerful. For example, it can be forced to write original stories with a long list of inclusions and exclusions without the quality collapsing and including a ton of story contradictions, unlike Qwen.It's abundantly clear that the primary reason most people in this community wrongly think Qwen is a more powerful model is because nearly everyone here is an autistic male coding nerd. Basically clones. And Alibaba is exploiting their shared blind spots and playing them like children, such as asking them what they want, then releasing a grossly overfit follow up coding/agentic model that they already trained and were going to release (e.g. Qwen 3.6 for 3.5), making them feel like they were part of a movement. Frankly, I find it all very cringe.Qwen 3.x aren't bad models, but it's abundantly clear that they aren't near as good as the frontier models, or even Gemma 4. Alibaba uses various testmaxing techniques to artificially inflate most test scores, but this community is far too narrowly focused on agentic coding to realize that the real-world performance across nearly all domains is FAR worse than the test scores indicate.
>>109747701I already have Claude code for that. How good are local models for working off a figma react project? Even redesigning the thing avoiding the pitfalls of the original design. I have some business rules data (100+) I want to use with it to help steer that design that's more suitable for said rules but I want to keep the business rule data from being sent to cloud models hence local Should mention I have 2 systems:16gb rtx408032gb ddr4 but can increase to 64gb4tb secondary nm790 for swap AndIntel 235h32gb ddr56gb vram
>>109747767Is lower better?
>>109747372>-fmoeon by default>-rtrpretty much don't want this unless you're cpu-only>what are all the retarded commands--jinja (not default)-sm graph for multi-gpu (not -sm tensor)flash-attn is on by default, but don't use '-fa on', it's just '-fa'the rest are model specific like -dsa etcoh and --webui llamacpp (if you want the llama.cpp webui instead of the based pink one)
>>109747351>Doesn't help that almost all models are the same now.Y-you can't just say that! There are rules! Probably! It's just efficient, structurally efficient!
>>109747962might be able to run the new qwen flash with that 4080 and 64gb ddr4 and your ssd for swaphttps://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main/UD-IQ3_XXS
>>109747922This is /g/ not /trash/ I assume coding might actually be a primary concern
>>109747962>can increase to 64gbdo that
>>109747991Thanks anon I'll check it out
>>109747447The 0-shot clone ability for JP is vastly superior to everything I've heard so far
>>109747992One of the arena categories Gemma 4 ranks higher in is coding. It's still inferior on coding projects because Qwen is better at function calling, agentic tasks, and long context precision.Qwen also uses far more thinking tokens, gets stuck in more infinite loops etc., and all to produce inferior less consistent results, so it's also far less token efficient.I can go on an on, but this isn't a small ~8b model. It's a 27 BILLION parameter model so there's no excuse for it to perform worse than small models at common tasks, even when overfitting for coding. Plus according to their test scores they should be unusually good at broad tasks, but aren't. There's nothing wrong with making a coding tool vs a general purpose AI model. But Alibaba is passing this off as a general purpose AI model, so I'm evaluating it as such, and so should you.
Has anyone tried doing something like this but for llamacpp, and with a local model instead of fable?I think maybe 5.3-flash could have a chance?
It has been almost a year since GLM4.6 came out and I am still running it.
>>109748024No one tell him
>170HX 8GB on AliExpressDo I dare
>>109748029Same here
It has been almost a year since Gemma 4 came out and I am sitll running it.
>>109748048100% a scam. there are all these chink listings for $400 where they say to send them crypto instead of paying through the site but all the other listings for $1500+.
Ugh
>>109748048I searched a couple of hours ago... they all offer "3 GPUs, pay only 1" and shit like that LMAOalso, I bet the shipping fee is more expensive than the card :^)
>>109748024I actually tried that with fable on llama.cpp. Was quite disappointed, it kept trying to run things that took hours, like for example trying a swipe and benchmarking each value combination, something that would take multiple days to run. I need my GPU for some other task at some point.
im kind of in a pickle. i bought a 5090, and now i've spent about a week genning. im out of ideas. what do i do with this 5090 now
>>109748141give it to me for free
>>109748141Enjoy gaymen some moreRun local models Run h3 minimax workflow for video gen
>>109748141Mario wouldnt do that.seriously go do something else you cant just do one thing all day every day. Take a break.
>>109748141have you tried being an artist
I made an Inkling chan that isn't a Surprise_mario.mp4
dsh is a piece of shit broke all the time.forever in alpha just like most of linux distro kek
>>109748170what is an inkling chan? new to 4chan
>>109748141just don't stop genning
>>109748181>uses alpha software>complains about alpha softwareretarded or pretending?
>>109748108I still have my fable bucks that anthropic handed out a while ago. I was thinking of burning them on Fable 5.1 to try porting some of the fancier optimizations that some backends have to llama.cpp.KTransformers has some neat optimizations like --kt-enable-dynamic-expert-update and --kt-max-deferred-experts-per-token that gave me a decent boost when I was playing around with that.
>>109747447Thats kinda hot
>>109747447This is exactly how I want my engrish, nice
>/lmg/ has been living under a tts rockgrim
>>109748377only one gpu i cant run llm and tts. if i had two gpus i would run bigger llms.
>>109745890You're a moron then. Make a conscious effort to remember what has been advertised to you, so you can actively avoid it.
>>109748377post some info then. I'm interested, and, you know, you can contribute too
>>109745890>I do not remember a single instance where I bought a product because of an ad.You might be autistic.Ads and propaganda work perfectly on normalfags.
>>109748377Please, do go on anon, I know I dont know enough
>>109748377TTS lagged behind all the others for years. Any attempts to sustain a general failed because there was nothing to say.
>>109745735>Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.no they definitely have notan llm "hallucination" isn't an error mode, it's fundamentally how the transformer-based LLMs work
>>109748390I bring up omnivoice maybe once a month, but i'm a very low energy shill with nothing new to add, and people seem to have limited tts interest overall besides.anyway, here's it doing weeb english from back in may https://desuarchive.org/g/thread/108847577/#q108848288https://files.catbox.moe/k424y7.mp3
Any epyc server bros here? Getting low key desperate and those V100s are real baddies fr
>>109748377Tell me how to get gemma to whisper sweet nothings into my ear!
>>109748456>Any epyc server bros here? Getting low key desperate and those V100s are real baddies frI've got a couple. What are you wanting to know?
>>109748416And it can't sustain discussion now because it's near feature complete, until we make the final leap to full duplex non-retarded local models (est. arrival time: Never Ever).
>>109748469Decode and PP for GLM flash q4 doko. I'm thinking of a cheapo zen 2 or 3 epyc, 256gb ram and 2x v100 32gbs (ewaste chad setup)
>>109748465Like so https://files.catbox.moe/d5jvwz.mp3
>>109748494I hit 10t/s decode with ud_q4_k_xl on 256gb x epyc rome x v100
>>1097485221x v100 16gb*
>>109748494>Decode and PP for GLM flash q4 doko. I'm thinking of a cheapo zen 2 or 3 epyc, 256gb ram and 2x v100 32gbs (ewaste chad setup)My little box is a similar DDR4 haven't got around to trying glm 5.3 flash yet. I'm but getting 15t/s on 4-bit dsv4 flash 0731.I'm running a single 24gb card and trying to run max context for big codebases so my pp is trash and won't be a descent benchmark for you.
>>109748416>TTS lagged behind all the others for years.I was building one for about 6 months, posted early samples last year.Wasn't ready to release yet.Got told "Sora will make it obsolete" or words to that effect.Stopped working on it.
>using Cursor>hit stop>edit prompt>revert >yeswhen can we has this in llama-server?
>>109748452I had no idea omnivoice supported cloning. If it's consistently that good, and has good range, then that's really great.Is the latency low? And does the model change tone, or it it stuck in the same tone as the clone sample?
>>109748522That's crazy for a single 16gb card, you need to get more v100s asap before they are gone like everything else.>>109748528Not bad either. Fuck I hope normalfags don't discover ewaste server setups.
>>109748554the bean people have already invaded.
>>109748540Can't stray too far from the sample sadly. I slopped a cli for omnivoice.cpp to offer voice change commands inside the text stream to get around it.
>>109748554Normalfag here, thanks for the tip
Its been pretty starved on new 70B models huh? Never paid attention, but looking to what I could use if I buy a bit more ram
>>109748039>No one tell himlol okay, whatever it is i won't waste my time on it then>>109748108>I need my GPU for some other task at some pointYeah that's been the bottleneck for me as well.>>109748317>I was thinking of burning them on Fable 5.1 to try porting some of the fancier optimizations that some backends have to llama.cpp.You might have better luck than >>109748108 since you've got actual reference implementations to point at.>>109732177>I was going to get dipsy to work on making the llama-server ui work with vllm, but damn she's slow and messy.Here's Qwen's stand alone version: https://files.catbox.moe/ragt60.gzAll static, just needs to be served eg `python -m http.server` or the equivalent in whatever language you haveJust put the llama.cpp server in that settings field, or append this to the url: `http://127.0.0.1:8000/?server=http://192.168.50.115:8080`I haven't tested it with vllm or tabbyAPI but if anything needs adjusting, the src is tiny enough for Qwen to do it I'm sure.
>>109748522>I hit 10t/s decode>>109748528>getting 15t/s>>109746977>20t/s on averagedamn and I thought I was ghetto running flash-next at 20 tok/s. it's not the fastest negro in town but with medium thinking effort it's VERY good and efficient and if you wanna go extra safe then xhigh effort is the way to go but a medium-sized task can take 40-60 minutes.
>>109748531>revert >yeswhat does that do?
>>109748629it makes mustard gas.
>look up "v100 32gb" on amazon>bunch of cards selling for $600+>a couple of items have 1-2 star ratings, comments say the cards arrived burned outI'm not touching these things.
>>109748494The fuck... I don't know why others get slow as shit decode but my 8 channel 2666 gets 16tok/s with 5bpw 5.3 flash, and ~25tok/s with dipsy. This is with a 5090 though so mileage may vary. As for decode, 5.3 flash in currrent PR is ~18sec ttft, and ~600 pp/s starting, ~400 pp/s sustained. Again with 5090 so don't rely on my numbers if you're going with the v100s.
>>109748108>kept trying to run things that took hoursI think anthropic and openai have switched from refusing to just wasting time and tokens when you try to do such things
>>109748680>openai ... refusingThey don't do this, not for LLM R&D anyway, that hangup is an Anthropic exclusive. Or have they finally given up with Astra and stopped it from helping with normal shit like that?
>>109748689I've been running LLM training experiments with Fable and Opus for the past 2 weeks. Not about enhancing llama.cpp but still, has been working
In the future models can probably make their own custom TTS voice. What kind of voice would you like it to have? Robotic? Human? just beeps and boops? Premade voice lines? Vocoloid? Synthesizer?
>>109748730A natural sounding blend of cat and little girl.
>>109748741Brat needs correcting
>>109748619Gonna have an ewaste setup at the next kegger. Shit's gonna be so cash.
>>109748730https://www.youtube.com/watch?v=cSehhpW6stE
>>109748689>They don't do this, not for LLM R&D anywaymaybe not on purpose, but when i tried fable via open-webui, related to creating a quant, first message was fine and quite good.then i replied and left -> came back shortly and found the model burning my openrouter credits reasoninghad a look, it was thinking about when intel was founded, different ceos, comparing with amd, all all sorts of schizoso maybe not on purpose, but i can believe fable would waste time/tokens
>>109746886>flash next>>109746874>flash-nextIs it actually better than Q6 27b?Looks like it doesn't quant well
>>109748793It's on purpose for fable and other Anthropic models, yes.
>>109748793Yes, Fable does, and Fable is from Anthropic, which openly states they sabotage your work. Again, OpenAI does not.
>>109748216A clanker whore from the Bay Area slut-vats, still fairly fresh. I forgot this was genning while I was making dinner earlier.
want to upgrade my pc really bad but i just know i wont be doing anything meaningful with it anyway
>>109748968the more you spend the more you save!It's only getting worse from here
5 PRINT "GPT-8 AGI IS HERE"10 INPUT ">";A$20 PRINT "I Cant fullfill your request"30 GOTO 10
>>109749065Can't be agi if it doesn't know who won.
Has anyone here done the pcie gen 3 unlock for the cmp 170hx?
Was anyone using this for web search? Stopped working for me last weekhttps://search.noemaai.com/Am I blocked or did they turn off the no-account free search?
>>109749102I just removed the chips from the board and put them on a 5070 board for gen 5 pcie.
>>109749129404
>>109747086TL;DR
>>109748654Just wanted to complain.Did google search with "vt100 32gb". First result is auto-translated plebbit post. The problem with this machine translation is how it defaults to spoken ebonics instead of standard literature form. >with a deal like this can I not get 4 of these instead of asus gx10? or dgx spark? This is standard English, but Google's machine translation was far from standard written Finnish. >Jos saan tällaisen diilin, eikö mulla voi olla neljä tällaista sen sijaan että ostaisin Asus GX10:tä? Tai DGX Sparkin?This doesn't concern native English speakers obviously but it's bad how a foreign company is forcing sub 90 IQ version of some other language. I guess this could be due to its training data and it tries to copy the style of how most retards are typing online, but this still isn't how written Finnish should work. Fascinating and also demoralising.
>>109749318Wouldnt a 64gb 170hx be a better but?
>>109749333VT100 is a waste of electricity these days, regardless. It's just too old.
New thread when
>>109749339The Tesla V100 is 9
New thread now>>109749465>>109749465>>109749465
>>109749333>Wouldnt a 64gb 170hx be a better but?Fuck no, that's buying a used card that already lost the silicon lottery. NVIDIA would much rather have sold it as a perfectly good A100, it was relegated to 170hx duty because that memory was defective. When they were cheap it could've been a justifiable gamble, but they're not cheap anymore, and not worth your time or money for a "chance" at some VRAM. The ones that come pre-unlocked to the full capacity and say they've been tested are a scam. Avoid that shit like the plague.