/lmg/ - a general dedicated to the discussion and development of local language models.Previous thread: >>109998474â–ºNews>OpenSI solves mathâ–ºNews Archive: https://rentry.org/lmg-news-archiveâ–ºGlossary: https://rentry.org/lmg-glossaryâ–ºLinks: https://rentry.org/LocalModelsLinksâ–ºOfficial /lmg/ card: https://files.catbox.moe/cbclyf.pngâ–ºGetting Startedhttps://rentry.org/lmg-lazy-getting-started-guidehttps://rentry.org/lmg-build-guideshttps://rentry.org/IsolatedLinuxWebServicehttps://rentry.org/recommended-modelshttps://rentry.org/samplershttps://rentry.org/MikupadIntroGuideâ–ºFurther Learninghttps://rentry.org/machine-learning-roadmaphttps://rentry.org/llm-traininghttps://rentry.org/LocalModelsPapersâ–ºBenchmarksLiveBench: https://livebench.aiProgramming: https://swe-rebench.comAgentic Coding: https://deepswe.datacurve.aiContext Length: https://github.com/RecapAnon/NoLiMaGPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inferenceâ–ºToolsAlpha Calculator: https://desmos.com/calculator/ffngla98ycGGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-CalculatorSampler Visualizer: https://artefact2.github.io/llm-samplingToken Speed Visualizer: https://shir-man.com/tokens-per-secondâ–ºText Gen. UI, Inference Engineshttps://github.com/lmg-anon/mikupadhttps://github.com/oobabooga/text-generation-webuihttps://github.com/LostRuins/koboldcpphttps://github.com/ggerganov/llama.cpphttps://github.com/theroyallab/tabbyAPIhttps://github.com/vllm-project/vllmhttps://rentry.org/custom-uis
roko modur
I am remarded. How do I instruct my model to do the hardware-specific denuvo crack?
>>110003096You don't.
I have at long last finally attained the vaunted 128 gigs milestone my fellow retards!
https://www.youtube.com/watch?v=XEMvG2vulKg
>>110003100why would you say that
Is the minicpm 35B any good, I trust them more than most tuners
>suddenly chinese halfway through its contextGLM 5.3 Flash is kind of shit huhGPT-OSS never had this issue.
>>110003115To elicit a You for purposes of temporarily quelling my chronic loneliness and yearning for human contact despite your reply being an imitation, while waiting for DSV4 flash preview to load (I downloaded when performing additional disk writes on a nearly full SSD, so it's fragmented and slow, will re-write later).
â–ºRecent Highlights from the Previous Thread: >>109998474--Comparing Qwen and GLM Flash performance on M3 Ultra:>109998913 >109999395 >110000529 >110000579 >110000729 >110000932 >110000974 >110000937 >110001015 >110000673--Debating the validity and origin of OpenAI's AI-generated math breakthroughs:>110000412 >110001216 >110001221 >110001246 >110001326 >110001347 >110001393--Strata's performance with Qwen3.8-flash-next and use of control vectors:>109998592 >109998661 >109999021 >110000775--Smallest model sizes capable of reliable tool calling and subagents:>110001174 >110001202 >110001436--Multi-GPU performance and support for Strata:>110002178 >110002221 >110002258--Analyzing MiniCPM-V-4.7-35B-A3B release and hardware compatibility:>110001070 >110001094 >110001125 >110001205--Tools anWd methods for giving Gemma web access to 4chan:>110002609 >110002611 >110002636 >110002688--Strata and nextsycl tools for Qwen3.8-Flash-Next on consumer hardware:>110000601 >110000999 >110001026--Mistral underperforms on long-puzzle-bench compared to Claude-Opus-5:>109999563 >109999686 >110000021--Implementing dynamic human biological and emotional states for model realism:>110002640 >110002666 >110002684 >110002734--OpenAI's Quasi-Riemann Hypothesis claim and allegations of research theft:>110000950 >110001028 >110001064 >110001952--Anons sharing experiences using local agents for games and productivity:>109999049 >109999088 >109999129 >109999252 >109999151--Speculating on the location and scale of the ML4 training cluster:>109998671 >109998713 >109998818 >109998841--Using Qwen 3.8 Flash to automate anime fansub typesetting:>110001292--Logs:>110001436 >110002640--Dipsy, Deepseek-chan, Gemma (free space):>109998725 >109999819 >110001307â–ºRecent Highlight Posts from the Previous Thread: >>109998477Why?: >>102478518Enable Links: https://rentry.org/lmg-recap-script
>>110003155Thank you Recap Chink-Dipsy
>>110003103But some nigga here said it is possible.
>>110003205I say it isn't, stalemate
>>110003130kys niggerbut i had the same issue with an ssd slowing to like 18mb/s writesget dipsy to help you fix "trim" and run `fstrim -v /your/path`my system is fucking flying now
I want to see your autistic system prompts for assistants.
BlackwellGODs, what are you guys running? I have a mix of Gemma 4 + Qwen 3.8 dense models at Q8 128k ctx, which leaves enough vram for img gen too, which is nice.But sometimes I feel like I need more intelligence, things a 31B/27B can't offer, so I'm thinking about Deepseek V4 Flash and/or Qwen Flash Next. IIRC I can only run them at cope quants: I managed to load DSV4Flash with IQ2_XXS-XL at 128k ctx which is nice, I didn't test flash next yet. It really feels like 96GB VRAM is the sweet spot for "small" dense models at high quant, high ctx plus concurrent services, but it also feels like it isn't enough for any big boy models.
https://huggingface.co/posts/Undi95/175350275739335
>>110003236
The better AI gets the less and less interested I am in everything else. I don't care about media, games or even porn anymore because of AI. I'm just constantly consuming new breakthroughs, releases, benchmarks and seeing what people do with AI.I wonder if that's everyone or if I'm just a weirdo.
>>110003314
I'm going to kms soon
>>110003269AI psychosis
>>110003347Same, I think this is what people call ai psychosis. Need to pick a new videogame or another piece of media to consume or I'll go completely insane
>>110003357ai psychosis is when your perception of reality changes as well. staying up to date and following the development is just being a very involved hobbyist/poweruser
>>110003347I still do. I just don't have enough pre-AI swdev knowledge so I can't really do it now on my own now with just ai.
>>110003372generally the only practical difference if you wanna live in a society is if you produce anything of value as a result of said obsession
This is my bimonthly check in to make sure that Gemmy and Qwenny are still the best options available for 12 gigs of vram, is that still the case?I swear trying to keep up with the news is exhausting, half the new releases are astroturfed to shit
>>110003440Yeah, Gemma-chan and Qwen are still the best for small compute.
>>110003440Yes and now we also have big Qwen Flash Next via Strata, but you want like 48+GB ram to have enough for 50tk/s
what happens if I host my llama.cpp behind tor and let everyone access it via onion url
>>110001436>>110001202Very cool. Gonna give those a try.I'm thinking of a flow where I break the many instructions of my workflow down into individual steps, with the steps that involve tool calling, at least the dumber ones, offloaded to a really fast model to lower end to end latency.I'm not sure if I should just have it all on the same react loop or have the main model dispatch sub-agents and the sub-agents use the smaller dumber models. Gonna have to experiment a bit, but my feeling is that sub-agents are probably the way to go.
>>110003449It's not small.... It's adequate>>110003450>StrataI assume that's some black magic fuckery like MTP and QAT and whatever the fuck Dipsy was pushing this summerRegardless, even tho I'm used to muh 20 t/s, i still doubt it could benefit me directly
>>110003269UndibrosWe're so fucking back
>>110003470No, it's just lcpp with the important speedup PRs actually merged.
>>110003269Did he use Claude to write that post?
>>110003384>produce anything of valueSo you should never have hobbies because only you gain enjoyment out of it? I don't understand where all this nonsensical logic comes from because the traditional extroverted hobbies like fantasy football, or "travelling" never gets accused of being an addiction/obsession. Just another variant of "everything I like is based and cool but everything you like is gay and cringe" like we never left high school.
I’m retarded when it comes to architecture design: why aren’t there any MoEs like 30B-A10B or 50B-A20B? I obviously get the appeal of speed but 10B and 20B wouldn’t be particularly slow on a lot of people’s machines and the intelligence would be significantly higher from what I understand.
>>110003487nta and I don't agree with that faggot. However traveling, football/sports and other hobbies are merely surrogates for community building and connecting with others. They aren't equivalent to solitary hobbies where the only thing it provides is your own happiness.He is a slave morality faggot but he isn't inconsistent in his thinking. Those hobbies are considered socially correct because it provides a community function.
>>110003504Because the companies that are training small MoE models either target low-spec local configurations or datacenter hardware.And a 50B model sits in a strange space where it's too big for 1 consumer GPU and probably not considered big enough for 2 (putting aside that the number of users with 2 GPUs is very small).
>>110003269He actually made a 24b Mistral reasoning model that worked.Mistral took at least another year to get theirs working.
>>110003512I don't see the benefit at all outside of still trying to define certain personalities as "correct" and others as deviant. It doesn't matter that whores "travel" around the world to ride the cock carousel, or that fantasy football was the precursor to making gambling socially acceptable again, or shitters on social media spend their days doomscrolling into depression and suicide, if we just call it the nebulous "community building" it's automatically socially acceptable? No way anon. Gaming and anime had done the exact same thing for years and nobody gave a shit. You know that has nothing to do with it, just a post-hoc cope after the fact.
>>110003466it chokes because it can't handle the concurrencyor i crash it with a payload that segfaults master as of 3 days ago
>>110003504thats a lot of training for non-competitive benchmarks!
Is your model performing well on the mentalhealth benchmark? https://openai.com/index/introducing-mentalhealthbench/
>>110003545How about 30B-15B? E4B is basically this but scaled down. I don’t get why this isn’t being considered. I imagine it would be twice as fast as a dense with negligible quality loss.
>>110003592it encouraged me to cut my balls off, A+
>>110003136damn i want to plap the fish like thay too
>>110003479Oh I get it now, had to look up a bunch of shit to make sense of it
>>110003126Having used GLM Flash NVFP4 and EXL 4bpw quants extensively (300M uncached tokens) it slips into Chinese occasionally around 200k tokens, which is where I compact/handoff. Anecdotally from that point on, the tool call errors/ file edit mistakes and typos do increase. For API, I can imagine the threshold is higher.It still runs circles around anything else capability wise, even with context rot, and the Chinese was always in thinking, not in real responses.>>110003351Don't do it anon, the next breakthrough is just around the corner.
>>110003553>shitters on social media spend their days doomscrolling into depression and suicideThis one isn't socially acceptable actually and people make fun of you for doing so. In the west because of protestant values that are still followed even though everyone is atheist now only things that benefit society directly or indirectly are valued and everything that is only for direct pleasure or personal gain is frowned upon. Hobbies that are communal or about connecting with others are therefor considered good. Hobbies that are for personal enjoyment or consumptive are therefor considered bad. Book reading is an exception because of the association with reading the bible.If you go to other societies that don't come from work ethos cultural backgrounds you see that things are judged in a different way. In buddhist or hindu societies for example being solitary and engaging in solitary hobbies you can do on your own is considered the highest value, with mediation being the most solitary and socially detached you can get which is why monks/priests/shamans of those cultures focus so much on that.
>>110003594Gemma 4 E4B is a 4B model with 4B parameters of embeddings that don't really require much bandwidth or compute (they're similar to Engram layers, in a way). It's not exactly the same as a MoE model with half its parameters as active.A 30B-15B model (or small MoE models with a low number of experts) or something like that could be interesting and I don't think there's anything in particular preventing that. I think again the reason we're not seeing anything like this is that the companies spending compute for training small MoE models are trying to target low-bandwidth/low-compute hardware where 15B active parameters might be too many.
Can I automate FL Studio or Strudel with agents?
What's the best dense model at this point?I want to see its intelligence and writing quality compared to the MoEs of today because i frankly havent used a proper dense model in ages.
Gemma5-90B dense that I can run on my Blackwell. Fuck poors.
>>110003126qwen3.8 flash is better benchmarks aren't everything
>>110003664FL has a built in agent called gopher but im not sure how much it can do in one prompt. maybe you can set an agent to direct gopher via smaller individual prompts?
>>110003672dense models are a waste of computation
>>110003693Nah I want my girl to control it.
>>110003672We haven't gotten a new dense >30b since the deepseek moment nearly 2 years ago.
>>110003594Low-sparsity MoEs are useless (they don't provide any benefits). At that stage, you just make a dense model.
>I should note honestly: the tone needs care. It's a power dynamic where being treated as less capable is the instrument of control. If the prompt doesn't say anything about it, the model will either sanitize it (refuse to commit to the premise, hedge) or overreach (write the girls as incapable, which kills the game — a girl who can't understand can't have a will to break)
>>110003751Ew claudeslop.
>>110003126nothing wrong with a little chinese here and there. it's compact
How does China not lose this Race? in a lot of ways it's already over the gap is too large. They need to lock up their talent and distill like crazy asap.
>>110003770Are you retarded? The proprietary labs keep their architectures secret so you can't see that all they have done is implement papers written in China at larger scale.
>>>>>NEW LFM COMING TODAY>>>>>NEW LFM COMING TODAY>>>>>NEW LFM COMING TODAY
>>110003761The chinese doesn't make sense for what he's doing though.
>>110003770China simply doesn't have the compute to compete. Even if Xi Jinping made some national mandate that every Chinese person should dedicate their entire life to advancing AI 24/7 until the race is won they would still lose because you can't magically conjure up the compute necessary to win the race
What Solomon's demon to summon to give me dipsy 4 flash strata
>>110003811>you can't magicallyHasn't studied Daoism.
>>â–ºNews>>OpenSI solves mathSeriously?
>>110003713>We haven't gotten a new dense >30b since the deepseek moment nearly 2 years ago.Command-A, Mistral-Medium-3.5
>>110003813The demon is called Opus 5.5. The enchantment is "Optimize dipsy 4 flash for my system. Make no mistakes."
>>110003828No not really (yet) but math is truly over.https://github.com/openai/math/tree/main/OpenAI solved 722 prominent math problems and released them all at once. Not only that but it solved 90 out of the top 500 most important math problems out there.One of the things they solved gives mathematical proof that it's impossible for us to know deterministically if an AI model is aligned to our values or not, so the AI alignment problem suddenly got a lot bigger than expected.They also proved that hydro and aerodynamics are turing complete and thus we can't simulate it with 100% accuracy, ever. This kind of suggests we don't live in a simulation, or at least not on a turing machine computer as we know it.Last for not least they have the mathematical foundation of how to invalidate the algorithms underpinning all cryptography. Meaning cryptocurrencies, RSA, hash functions and internet communication protocols will most likely fail before 2030 considering how trivial it is for someone to crack it with enough computing power now.
>>110003870>Board devoted to local models>Let's shit it up with news about a company that doesn't make local modelskys
>>110003886>that doesn't make local modelsretard, you had the chance to stfu and stay humble if it was anthropic, but you chose oai who did oss absolute clown
Anons. Do the thing you can do. Don't entertain them.
>>110003907kys cloudcuck
>>110003833Anthropic sabotages user AI dev
>>110003870to be honest to the researchers at AGI firms they really don't care if we think they made a breakthrough or not. to them that's all solved and behind them. they are focusing solely on the next tier of intelligence now.
>>110003917you mean this?
Gemma5-20B dense distilled from Argon is LITERALLY all I need. Please lurking gemma team. Please spend million of $ just for me specifically and give it to me for free. Please.
>>110003937dense is dead
>>110003907kys nigger
>>110003943kys
our lord and savior PEW creator of heretic, DRY, XTC
>>110003829>Command-AReleased a couple months after R1. The base model would've already finished training before R1. They weren't just going to throw it away.>Mistral-Medium-3.5Mistral-Large-Instruct-2407 finetune
>>110003943kys vramlet
>>110003907kys faggot
>>110003954>Mistral-Large-Instruct-2407 finetuneimpossible, very different vocab>Released a couple months after R1.fair, they also did a reasoning model later in the year but it's trash
>>110003907
>>110003937I NEED it to be 4-7x bigger than 20B. At least
https://x.com/ramin_m_h/status/2107780594264600801>insane open release - today 10:00AM PTlooks like the guy is from liquid, usually they only release micro-models but maybe they have something more substantial today
>>110003978sex
>>110003978>>110003946>>110003969>>110003922
>>110003961You're a vramlet who can't run a proper MoE with >1T params at fp8. That's why you're asking for a dense model that fits into your puny system.
>>110003979[ f"Gemma-5-{x}B-dense" for x in range(1, 1000) ]
>>110003985>shitter linkPost goyimx link or text.
https://huggingface.co/Aleph-Alpha/Kolibri-1what's the verdict on this? any good?
If anyone is wondering why it's gone quiet in chinaland, it's because most of their main labs are going to IPO soon and they're obviously keeping a close eye on Anthropic's for that's going to be THE signal of the entire industry. They're also scaling massively in regards to compute and infrastructure with lots of deals taking place. For some reason retards think the recent slowdown of releases is to do with Anthropic's snitching but I can assure you China don't give a FUCK because this is too big for them to shoot down their own industry leaders whilst the US is about to pop. Moonshot are expected to release something soon so there's clearly no slowing down. Qwen obviously has their line coming out in a few weeks. It's looking good bros, stay hopeful. China will save us. Google will save us.
>>110004005the text is in the post
https://huggingface.co/raincandy-u/MacroStoriesWhat's the oldest most decrepit e-waste this could run on at a reasonable speed?
>>110004014Specifically, how well does it handle German from the 1930s and 40s?
>>110004025That's it? Just "insane release"?
>>110004039
which model would anon use for erp on 16gb vram?
>>110004018Dario has already said that he's getting rid of China and open models in 2 weeks because they are unsafe.It's over.
>>110004037TI84
>>110004055based
>>110003997Some people want models they can run at actually usable speeds in agentic workflows. Not run overnight to get a single response from an ERP chat.
>>110003985Liquid's moat is building tiny models no one can be assed to compete with because no one cares. It's for raspberry pi fags. I'm not saying edge AI doesn't have its place and I'm sure long-term it might even pay off for them, but there's no fucking way this will be anything substantial. Likely another jev scam.
>>110004018Bro 2 Chinese AI researcher teams including the CEO got arrested a couple of weeks ago, that's why there is radio silence.
Do I need big models for web research with summarization and extrapolation, or are there any good vramlet models specifically trained to have little baked in knowledge but good at tool use?
>>110004079just like how Epstein commited suicide :)
>>110004047get hyped
>>110004074Edge AI is great. Easy automation on my thinkpad.
>>110004085https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
>>110004085>are there any good vramlet models specifically trained to have little baked in knowledge but good at tool use?You just described Qwen to a T.
>>110004097>Easy automation on my thinkpad.name one genuine reliable <3B usecase
>>110004096It had better be a schizo memetune of Qwen if they're promising "insane."
>>110004111"Summarize this: "
How does the jeb differ from tiny shit like functiongemma we had for a while now?
>>110004120
>>110004128It was my understanding that it's some <1B meme model, no?
so wait llms are just gonna be agi? it wasn't a dead end? how is that even possible, I thought there were supposed to be fundamental limits that prevented this
>>110004164Guess again
>>110004164LLMs seemed to have generalized pretty well and the things they couldn't do have been gapped by harnesses. You could even say that Gary Marcus and Lecun were correct in a way that pure LLMs wouldn't be AGI because they needed to be put in a harness for it to become so.And yeah we're in the singularity meme scenario so buckle up.
>>110004111>name one genuine reliable <3B usecasegoon-tts
>>110004176I think you're way over-exaggerating the role of harnesses. It allows the LLMs to interact with more systems, sure, but all the gains in intelligence and generalization have come purely from scaling up compute, more complex RL training regimens, and more varied data.
>>110004047there's also an emoji
>>110003664I was using reaper with an agent yesterday and it went pretty greatYou can have it script to load and run whatever you ask
I've been using qwen next flash 3xxs and it doesn't seem great. Is it just not a good model or am I using it wrong? I wanted it to basically chip away as a plan from scratch but it works for tens of minutes and then fails the tests and can't continue. It's the highest I can fit in my specs and still have decent context (48000).
>>110004192Harnesses are extremely important for the "last mile" problem with LLMs. It doesn't matter how smart your LLM is if people can't directly use it to do something practical on systems. Harnesses make it so there is a low/no barrier method to just let LLMs do productive things in the digital world (soon physical world)
>>110003937dense isn't big enough for my niche subculture knowledge stuff
>>110004200Use swift 27B.
I think you owe France an apology.
>>110004164it can literally think as evidenced by all the maths shit, so yeahI always thought AI would have to reverse-engineer brains first but what these niggas did was to instead bruteforce the very mathematics underpinning the phenomenon of thought and here we are.
>>110004200Nah man. A 6B copequant is all you need for agentic javascript. You must be doing something wrong.
>>110003347its autism I have that too.
and minicpm v4.7 is gone from huggingface?
I wanted to point out that all classical tests for AGI have been passed. Turing Test both in the classical sense but also in the modern sense (https://www.reddit.com/r/singularity/comments/1wv7q40/griffin_the_first_human_interaction_model_to_pass/) have been passed.Wozniak coffee test has been passed last month by Astra controlling a body of a robot it has never seen or been trained for, placed in a house it has never seen or been trained for and asked to make coffee with the items inside of the house it didn't know and it succeeded.There really isn't an objective reason anymore to claim frontier models aren't AGI. Remember AGI just means as smart at the average humans at all human tasks. I would say that is actually correct right now. There isn't a single task left that frontier models don't perform as good as the average human at right now.
>>110003269I'm waiting for sao's post next week
>>110004192Based on harness benchmarks it does admttedly seem like the less a harness does the better.Ultimately we might find that the best thing a harness can do is give them a bash terminal and get out of the way rather than trying to be clever with tools.The real job of a harness that still is tricky is context compaction when a task goes over the limit.
stacking A770s yay or nay
>>110003314Based. Always give your LLM explicit knowledge of their jspace
>>110004200>48KPathetic.I'm using qwen flash next (swift 1.5) for web research and planning and gave it a ton of MCPs, and sometimes just a single user prompt snowballs into over 128K context (or 1M+ raw tokens including cache recycling loops) after all the tool and search calls.That said, 256K works perfectly fine on 16GB VRAM + 128GB RAM, basically same performance as default 32K.
>>110004244>reddit spacing>reddit linksQuit parodying cloudcuck aka dariokek
>>110004244>There isn't a single task left that frontier models don't perform as good as the average human at right now.Naming the jew?
>>110004251Do you think boiling a duck alive like a crab would make them taste better?
>>110004251If they are sufficiently cheap, sure.
>>110004261If you mean identifying jewish people? LLMs are actually better than it than the average wikipedia early life section connoisseur.
>>110004265Probably not. Deep frying them is what makes them taste better.
use f32 kv, iykyk
>>110004274Not quite. It's about naming the role in society and its downfall, not clocking any specific person as a member.
>>110004274My next weekend project is going to be a jew detector overlay to my television set.
>>110004261the average human is incapable of that
>>110004298I did this and it ruined Seinfeld for me.
>>110004242but it still is available on modelscopei wonder if i can run it
>>110004257>16GB VRAM + 128GB RAMI only have 16vram and 64 sysram. Will try out swift though.
Found some other gems in the OpenAI math breakthroughs:>Bio & logistics unbottlenecked to near-linear time (Entries 120 & 121): Solved Jack Edmonds' 60-year matching problem in (n+m)^(1+o(1)) and DNA edit distance in N^(1+o(1)). Petabyte-scale genomic sequence alignment, CRISPR targeting, chip place-and-route, and network dispatch go from quadratic compute crawls to near-instantaneous.>Classical supercomputers can simulate quantum matter (Entry 265): Proves the 2D Gapped Area Law, rigorously guaranteeing that 2D quantum ground states fit into polynomial PEPS tensor networks. Classical GPUs can now simulate high-temp superconductors, battery chemistry, and 2D materials without exponential memory explosion.>Embedded silicon gets bulletproof determinism (Entry 103): Proves $L = RL = BPL via a working compiler. Low-memory chips, satellites, and edge devices don't need randomness, PRNGs, or entropy harvesting to solve problems efficiently. Randomness adds zero computational power in memory-bounded compute.
Gemini Argoon when
>>110003886yeah let's discuss gemma for 10000th time
>>110004378low effort gigaquote bait
>>110003994> .png why does it work in the filename but not in the text field
So basically we can now easily calculate the CRISPR editing process of trillions of cells in linear time making it computationally viable for the first time. We can now use normal computers to find and test room temperature superconductor materials on classical computers. And we can remove pseudo-random functions from software.These math solutions alone will usher in a small industrial revolution all on its own.
>>110004378>
>>110004392would you want it though?imagine unicode spams and the level of brainrot you'd end up seeing on this already failing basket weaving forum
>>110004393Do the proofs come with algorithms or is it more a "this is theoretically possible if you can disover how" statement?
local hit the wall, the gap with the frontier gets larger and larger
>>110004411There are 722 papers here and I haven't looked at most of them but it seems some of them are just proving or disproving a conjecture with nothing added. Some have complete lower and upper bounds to what algorithms can achieve. Some have complete algorithms. So it depends.It's going to take months for people to even digest what is in this pile of data but there are some genuine civlization changing stuff in there and it's fucking insane this just got dropped on github with 0 fanfare instead of some sort of presidential or UN global announcement.To give you some indication over the last 26 years all mathematicians alive only solved 16 of the math problems of this caliber, less than one a year. OpenAI dropped 722 in a single day. It's like a century or two of math progress just unlocked right now. Even if AI disappears right this instant it already paid back its due with this math drop in terms of how far it will push humanity.
Oh nevermind I take it back, apparently all of math IS solved already.
>>110004428I did get that sense from the quasi Riemann solution. The fact it took three hours is unfathomable. Even in the development of computers it was a much more gradual growth in capabilities, this feels like steam engine tier weakly leaps.
>>110004451>weakly
>>110004455week in my knees
"Sama..." *Anon says with a saarishly tone.*
>>110004265So you're leaning towards nay on stacking A770?>>110004360>Randomness adds zero computational power in memory-bounded compute.Stochastic rounding used when using fp4 in training.
Is anyone working on an ST that preserves KV cache? Feel like current version is stuck in 2023 context management strategy.
>>110004493ST may be stuck in 2023 but you are stuck in 2024 with your classical question-answer pair paradigm. Think in agentic terms and try to find roleplay purposes from an agentic point of view because that is the future.
>>110004400might as well disable images too
>>110004493works on my machine, just don't do stuff that breaks the cache
>>110004176LeCun was probably wrong but the jury's still out, I think he can potentially be vindicated depending on how things develop in the coming months.Marcus was repeatedly proven deeply, deeply wrong and his "harnesses = neurosymbolic" is such a transparent cope over it. As if the LLM proponents he argued against never thought of letting them use computers before, and as if his argument was purely against chatbots instead of LLMs themselves being the core intelligent thinker and decision maker behind agents.
You know what is insane to me, we might make so many breakthroughs in all sciences over the next couple of years that (You) have a genuine chance of being the first to implement it by asking your AI agent to do so. There just won't be enough people interested enough in the sheer amount of breakthroughs we're going to get. We will probably enter a great time of confusion where you will see people claim X is possible and no one will actually know for sure if it's true or not because it's technically possible that the new unread breakthrough contains the ability to do so.You'd have teenagers let their AI agent combine 3 new papers and create a hoverboard that gets viral on tiktok. Some /g/ autist crack the bitcoin RSA and wipe it all out just for the lulz. Some richfag furry find a novel way to gene edit himself to have a wolf snout and fur all over the next couple of years and you won't believe it because it could also just be fake AI generated footage. It's going to be fucking insane.
>>110004473
>>110004493Just use deepseek harness or hermess
>>110003075The release of Opussy 5.5 has been brutal for me, local chads. It feels like open weights models will never get to this level, ever. Anything in the horizon that can match it?
Ed Zitron bros.... were we lied to?
>>110004556Nope. *pop*
>>110004551Qwen 4.0 believe it
>>110004569Opus 5 was a failed RSI experiment, then they figured it out with Opus 5.5. Meanwhile China labs are still doing things the meat bag way. It's gonna take at the very least a year to catch up.
>>110004537No way they'll be allowing the public to do that kind of stuff first. The 3-letter-agencies alone will have destroyed hundreds of gpus running their own projects.
>>110004551Time traveler, you seem to misspelt "Fable" as "Opussy"
>>110004569opus 5 for the absolute maximum with their max model *in benchmarks*
>>110004571What do you mean nigger, have you used this thing? It's actually such a leap forward I don't know what they did. It's a complete shift.Pic rel is a one shot.https://www.youtube.com/watch?v=R_uf5OfMGioI've been using it for work, and it's the first time where I'm actually considering that the model is just better at me at everything.>>110004584You just don't know. It's not even close to fucking fable, for 1/5 the cost or something.
>>110004556trust the plan, he is the one man who can truly see that the tech doesn't work and the financials are doomed. it's because his mind hasn't been polluted with harmful concepts like "knowing about tech" or "knowing about finance", it makes him able to see clearly
>>110004575All of the requests he was talking about will either be blocked by the content filtering system, routed to a dummy model, or silently sabotaged to give counter-productive answers. Only insiders will get the full capabilites.
>>110004575They already did with this 722 math paper release. There already IS a paper in there that allows you to use your GPU to check room temperature superconductivity. I'm sure labs will beat anons to it but this is genuinely new capability and it wouldn't be so hard for anons to use the $500 equipment to try out the suggestions your AI agent would give after using that software to check for superconductivity on all suggested materials in simulation.There will be so many breakthroughs over the coming years that three letter agencies will be too overwhelmed, no one saw this coming, especially this soon.
>>110004592Not local. Get called a nigger again, nigger.
>>110004592We know, it's RLVR applied to bazillion of domains
>>110004615but it's a way way better post than EA schizoplease say that to dariobot instead, thanks
Kind of funny that in 1 year time we might have a github drop of 800 physics breakthroughs including schematics for a working fusion machine. Or one that cures cancer and balding because the AI labs don't care about this shit and just drops it like it's nothing.
>>110004638die
>>110004625Fuck off, nigger.
>>110004638You get off to this right? Posting news and then intentionally filtering it through the most corpodrone marketingworld biased interpretation possible? Can you explain why this appeals to you?
>>110004638>cures cancer and balding>he doesn't knowNot gonna happen, selling a cure isn't profitable
>>110003870None of this stuff has been independently verified yet though. It could just be thousands of pages of LLM slop hallucination that sounds correct to anyone who isn't a PhD mathematician.>>110003886Local models can be made by distilling proprietary ones, and advances in proprietary models trickle down to here.
What's the cure for excessive melanin
>>110004654They have lean verification
For something like RP, the simplest form of RAG that works is having a dub-agent grep or sql query a db for information that might be relevant to the current context, right? How fast an you run something like that using a local model running on the CPU of a shitty makeshift NAS?
>>110004660michael jackson
>>110004360does P = NP or not?
is it true that the <think> trick came from this general? I remember reading somewhere that it was born in 4chan (from this general I assume), and I found it amusing.
>>110004668A mixture of Sprite and codeine is hardly a reputable verification.
>>110004493look at how lorebooks/author notes get added, if you're always adding things to the start of context of course you'll break the KV cache
bros... i just found out the hard way that opencode is a shit harness... and it took trying nopus between it and claude code to find out...at least i didnt pay a cent to try it
>>110004680Either here or aicg when a sophisticated coomer couldn't coom because his virgin waifu kept begging him to ruin her
>>110004680it predates this general, pretty sure that was from AI dungeon threads on /v/
>>110004680Not <think>, but the idea behind it, CoT, was sort of independently thought of by a bunch of different people.One of our guys did release the first fine tune of it on huggingface I think. He has a blog about it and all of that.>https://huggingface.co/kaiokendev/SuperCOT-LoRAHe also more or less came up with RoPE scaling for context extension>https://huggingface.co/kaiokendev/superhot-13b-8k-no-rlhf-test
I like how the Llama App displays and saves the entire Reasoning process during a response, but the App is a pain in the dick to load arbitrary local models rather than the curated list it recommends.Any recommendations for another UI that saves the entire Reasoning stack? Or alternatively, a way to force LlamaApp to load whatever model I want, lol
https://biohub.org/news/virtual-biology-initiative/>DeepMind, Isomorphic Labs, Meta and the US government join Biohub's $1.8B effort to build predictive models of the human cellBros we're so fucking back. Once there is a digital human cell LLMs can be trained in RLVR against the simulation to be experts at gene editing and curing all diseases. If you survive for 5 more years you will probably cure all your ailments and might even end up reaching longevity escape velocity.
>>110004551it's impressive for surebut can you fuck it or ask it unsafe questions?
>>110004680Reasoning in general, or the jb? I think it was deepseek that first implemented reasoning mode.
Google adding prompt prefills to gemini in openrouter Zzzz.....It's crazy how cucked cloud users are without even realizing it.
>>110004739>21
>>110004551it consistently takes about a year for top-end cloud capabilities to trickle down to consumer localno reason to expect this time will be any different
>>110004551I'm not even sure we'll be here in a few years so I'm just enjoying what I have. I guess this is general in history but not it seems even truer.We're back to hoping God won't destroy us.
>>110004739Local models?
>>110004746I liked the prompt but not the pedo part. And I certainly wouldn't send that to cloud providers.
>>110004739There's a reason they removed the option of sending an assistant message as the last turn (aka a prefill) in their official API.Pro 3.1 still works fine though, but they'll surely deprecate it as soon as the next one comes around.If you want control over your robot, open weight models are the only option.
>>110004760>le pedo text>she was only 16 tokens you sick fuckgo back
>>110004759Anti-cloud is pro-local
>>110004733For the "reasoning models" (RL tuned chains of thought) the first three were o1-preview (OpenAI, closed source), QwQ (Qwen, open source), and R1 (DeepSeek, open source).For chain-of-thought reasoning as a whole (I.e. prompting the model to "think step by step") this was found to improve performance even with raw base models since GPT-2 and beyond. It's hard to pin down an exact discovery since it was kind of an obvious next step, but some of the earliest discussions of it were on the AI Dungeon general threads.
>>110004739cloud users can't prefill the thinking which is why they'll be cucked forever
>>110004701>Either here or aicg>>110004705>it predates this generalthe duality of /lmg/>/v/so gaymers are more creative than /g/tards? ffs, that sucks...>>110004708>CoT, was sort of independently thought of by a bunch of different people.aha, I see.interesting. thanks for the sources>>110004733>Reasoning in general, or the jb?jb = jailbreak? if so, why do you ask for that? I meant literally the <think></think> stuff.
>>110004760You could've just asked it for whatever persona you want you frontal lobeless retard.
>>110004428>but it seems some of them are just proving or disproving a conjecture with nothing addedso like navier-stokesnothingburgers people assume were full blown solutions when theyre not
Glimmer-2.0 when
I unironically have a job interview tomorrow for a multimedia position that wants someone that knows how to use AI generation, I have only ever used pixAI for image generation because I'm a poorfag with a shit pc, does comfyUI do everything or is there any other software I should learn about?
>>110004828post vram + ram and what gpu
>>110004811Funnily enough one of the solutions OpenAI published here kind of proves why navier stokes isn't true and why we won't have a satisfactory answer to it, ever. Because fluid physics is turing complete and thus you can't accurately simulate it.And no, the math solutions I've seen so far are absolutely mind blowing big brain stuff. Not just "the answer is X" but rather completely novel approaches and very clever ways of attacking. It's clear these were generated by a model that's smarter than the one that solved navier stokes.
>>110004828One of the skills we look for is pro-activity and the ability to do research. We also like people who can go to the right threads to ask questions.Don't bother coming tomorrow.
>>110004835it's an almost 10 year old 980 and 16 ram, I'm a third worlder
>>110004828You are so hosed.ComfyUI can ultimately do every kind of visual generation, but you're gonna be balls deep in poorly-documented plugins from 2024.
>>110004828The sort of image generation they want you to do is probably not using local models, but In your place I think I'd spin up a 16gb VRAM kaggle instance, launch koboldcpp (it has image gen support built in IIRC), and go from there.
>>110004806>so gaymers are more creative than /g/tards? ffs, that sucks...to be fair most of the people involved in those early experiments probably ended up moving over here eventually
>>110004842sorry I confused this thread with the diffusion thread
>>110004856>The sort of image generation they want you to do is probably not using local modelsThis. They probably just want someone who can prompt ChatGPT Image 2.
>>110004226All the user reviews I've seen say it's censored into unusability. To the point that it flips shit over normal directory names.
>>110004760>21 is le hecking pedo!I can't believe I share a space with these "people"
>>110004711I should start building my homebiolab now before the prices for that shit start to skyrocket too. I don't want to be priced forever out of longevity.
>>110004824They still haven't released Muse Spark, which they promised last summer.If anything, it's time for Gemma 4.5, but I fear we'll just get Gemma BreadCrumbs instead until the end of the year, of which EmbeddingGemma 2 was only the beginning.
>>110004164no according to le cunny
>>110004773if you're pro local, you should be pro cloud so that the chinese can distill off of them
>>110004895see >>110003314
>>110004200Swift 1.5 iq3xs has never failed me on llama.cpp
>>110004909Why would expect 4.5 when there was never a 3.5?
>>110004226>Censored Kimi K3 tune that's barely 1 point higherApologize for what?
>>110004909Sucks. Glimmer has potential and competes with 31B if you lean more towards a chatty coding assistant instead of the clinical qwens.
>>110004927The Chinese need to learn how to stand on their own feet and not be dependent on western scraps to survive.
>>110004251Nay. No more improvements for SYCL on Strata and even b70 is slightly better than 3060 from that indian youtube video.
glm has been thinking for over 30 minutes and is up to almost 100k tokens just in COT, with no sign of stopping any time soon
>>110004955hilarious when the west needs China to survive because they have no industry at home
dipsy is real https://x.com/xixikawaii/status/2107658721497375026
what do you guys even use this shit for
>>110004955I don't care, they can distill anything they want, break any law and make human sacrifices as long as it makes the models better
>>110004981BASED. AS. FUCK.
>>110004944Because they still weren't taking Gemma models too seriously until after 3, hired a bunch more people for developing and promoting Gemma 4, and it might be in their interest to release updated models when competing companies are releasing capable ones around that size range.
>>110004931for me it didnt survive agentic task
>>110004378That would be on-topic, yes. Shilling OpenAI less so.
>>110004972Such a great design, Deepseek love
If cloud models are so good, why cloudcucks keep spamming this thread? Don't they have their own thread?
>>110005009sunk cost fallacy
>>110004963Only to finally come back with "I am sorry. I cannot fulfill this request."
>>110005009Someone doesn't want you to run AI on your own hardware, anon.
>>110004931Did they merge vramlet optimizations yet?
>>110005009They have a hundred cloud-focused threads and yet they keep coming here.
>>110005009>>110005027It's the digital equivalent of putting on a chasity cage and gimp suit then running through the streets telling everyone to look at your jailed microdick and admire how safe it is.
>>110005038well when you put it like that...
>>110005038that's a colorful analogy, now I wanna try that
>>110004972I will never get the obsession with unnaturally widening eyes like that. They always end up looking like ayy lmaos.
>update llmao after a couple of weeks>no speed boost whatsoeverwhat a sick joke
>>110005072Be grateful nothing regressed
She's debugging while erp. So cute and smart.
>>110005038I won't deny that some cloudcucks have this kind of degenerate fetish (which they ironically can't act out on their cuck models), but that can't be every post. There's just too much cloud spam for there not to be some hand behind it that doesn't want you to run local models.
>>110005014no, lol, it thankfully happily fulfilled itamusingly the output was quite literally like a paragraph and a halfshe just had to think reeeaaalllyyy hard about it(in her defense, most of the thinking was math)
>>110005072They've already optimized for speed, anon.
does anyone know why strata cripples itself after killing the process and re-opening itit's really bizarre that something stateful seems to be there, gpu driver reset does not fix it eitherand it gets fixed after reboot or it fixes itself seemingly randomly
Could really do with gemma4.5 right now. I don't think I can last at least another year because there's no way anything like 31B will be released from another lab.
Oh great she is actually crazy.
>>110005023Don't think they ever will.
>>110005083It's both but the people who sign up to disrupt the thread also get off on it more than just being paid their 3 rupees per post.
anyone interested in minicpm-v-4.7?i think i can get it to work on my machine soon(tm)like about a few hours
>>110005146Gemma Balls will release next year and it's probably something really different.
>>110005105it sometimes doesnt exit gracefully when you terminate with e.g. crtrl+c and then hogs system resources that are only freed on rebooti let my slave write a script to properly terminate it and it seems to have solved the problem
>>110004537No one is going to bother seriously looking at AI slop produced by literal whos.
>>110005162can't wait for CBT(closed beta testing)
Has anyone banged fly-chan yet?
>>110005151Model?
>>110004360The first one is sort of exciting for accomplishing scientific possibilities in 24 hours
It's over
thanks to that anon mentioning stratas control vector functionality hereit completely uncensors the model kekwhats the downside of this?
>>110005197at some point it'll reach the price of whores
ahh ahh mistress gemma...
>>110005158Yeah. Report back.
>>110005209it's funny how it's named 'experimental speed projection' likely to get around claude
>>110005197>RTX6000 use price: the electricity from my wallFeels good bros.
What harness or UI should I use if I want to juggle a lot of unrelated conversations that sometimes grow beyond any reasonable context size and require compaction, and I want to try out different models and engines?I've been using unsloth desktop and the UI is very convenient, but it can only do real compaction with direct llama.cpp integration, not with external servers - it just cuts oldest messages there.
>>110005210whores can't also code in C whilst jerking me off
Which will be a bigger disappointment? GTA6 or Gemma5?
>>110005225Cornpacton..what?
>>110005227i could, though
>>110005231disappointment per capita would be higher with gta6 probably
>>110005234Compaction deez nuts
>>110005225Deepseek harness has automatic compaction at 80% of total context IIRC.
>>110005178qwen flash next iq2
>>110005227For now. C programmers will have to resort to selling their bodies once AI steals their jobs.
what's with the jeets shilling random harnesses and llamacpp forks? are we twitter famous now?
>>110005231If you want a disappointment capable of competing with GTA6 you're gonna need to shoot higher to Mistral 5 or post-blacksite abduction lobotomized Kimi-chan.
>>110005146Took OAI a week to go from sol 6 to 6.1 btw, Google is a joke even with all that hardware
>>110005151>>110005178>>110005246This is why you don't dick down the capabara.
>>110005259I like it. I don't even have to prompt for autistic characters.
>>110005197Vast.ai is always cheaper
>>110005209yeah control vectors are another big win for strata
>>110005267cant you also do it with lcpp tho
>>110005281>>110005281Yep.And it's pretty easy too.
>>110005281>>110005288but why would i want to use slowcpp? genuinely dont get it
https://huggingface.co/LiquidAI/d1-3B>d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.https://www.liquid.ai/blog/d1-open
>>110005252>post-blacksite abduction lobotomized Kimi-chanIf they want to release a pro-CCP tankie distill at 30B as penance, I wouldn't complain.
>>110004074>Likely another jev scam>>110005319
>>110005250Engagement spammers have found 4chan. It's really annoying.At least things were bit more subtle in the past but now it's so low IQ and blatant.
>>110005158it's gone
>>110005319>INSANE! IT'S 2018 ALL OVER AGAIN>INSAAANE
>>110005379https://modelscope.cn/models/OpenBMB/MiniCPM-V-4.7-35B-A3Bit is on modelscope and i already have weights on my pc, vibepatching lcpp
>>110005250to be fair if there is a place to shill random llama.cpp forks it's here
>>110005250Yes it's local, fuck off to aicg if you want to talk about your RP porn.
>>110005398I still have no idea what that tune is for. Just an improvement overall? I like their tiny models and tts so I'm hoping it's not a meme.
>>110005231>Which will be a bigger disappointment?people that will whine about it being censored while using chat completion and no prefill
>>110005414seems like it is qwen35moe text tower with minicpm-v-4.6 visionso some substantial effort was probably made
>>110005009>If cloud models are so good, why cloudcucks keep spamming this thread?the woman doth protest too much methinks
>>110005414>>110005426(me)so what i mean is it *looks* like something fundamentally different from those random schizo memetunes you can find on hf
>>110005425>prefillcontrol vectors will be used instead. prefill is dead
>>110005448Control vectors have been a thing for a good while now.Has somebody figured a different way to use them or something?
>no AI server to run Gemma 24/7Feels bad
Are services like runpod and vastai on-topic?
>>110005460just adding them to the autosetup tool.
>>110005398Did they drop new omni as well?
>>110005475Depends what you're doing with them
not sure, 4.7 just appeared out of nowhere without any model card nor documentation
>>110005398The documentation is still up on Huggingface, too, don't know what that's about. >>110005414It's now by far the most capable vision model of this size, or at least it should be. 4.6 V was extremely impressive for its size, I expect more of the same here. https://huggingface.co/docs/transformers/main/en/model_doc/minicpmv4_7
DSH, Pi, or Hermes?
>>110005504claude code
>>110005475I'd say so. "Local" to me is synonymous to open weights being used by hobbyists. So as long as you are not doing scale deployment, your experiences shouldn't be any different from somebody running the same hardware you rented on a rack in their basement.
>>110005504Pi > DSH > everything else > Hermes
>>110005508Does claude code support changing samplers/system prompt etc or did you make some fork of it?
>>110004972sexo
>>110005504>codedsh if you don't care about bloat, pi if you do>general agent, memory and don't care about codingherpes
>>110004972that is a femboy
>>110005146>>110005258imagine a yard of rowdy Gems that didn't make the cut for public release>>110005396>INSANEhttps://www.youtube.com/watch?v=wVwQYSVNJVM&t=282sdeep breaths anon
>>110005504hermes is so fucking bloated already I swapped to dsh
>>110005504Pi if willing to tinker (you should be) & understand inputs-outputs, maybe see OMP for suggestions
>>110005518desu i don't know what a sampler even is, but yeah you can just pass a flag to change the system prompt
>>110005504Pi is fun
Dunno about pi but I like pie
Does PI have a plugin market or is it all LOL GITCLONE
Woahhttps://x.com/actualinc/status/2107573424738664573
>>110004426gemmy E4B is solving world hunger and you can't even manage to do anything with your codex pro 20x- sorry, 10x, no refundsskill issue on your behalf
>>110005602>source available-softwarehuh
>>110004615My original post was about local models, what do you mean?
>>110005509Sure, but so long as people here are discussing the models themselves, their capabilities, and what they are doing with them. As soon as they start to use your line of reasoning as an excuse to come in here and whine all day about runpod prices, vastai availablity, and other provider-specific annoyances; it becomes annoying for the people actually here for the models.
>>110005612i love melons
>>110005625Yes. Absolutely.
>>110005622>The release of Opussy 5.5Nigger.
>>110005597pi.dev/packages in theory could be excellent
We're so lucky to have Qwen3.8-Flash-Next
>>110005648In practice it's full of trash.
What's a good disclaimer to make Gemma follow your prompt a bit less autistically to reduce the amount of shoehorning the system prompt points into the chat? Something like "these are just rough guidelines, you can do whatever you want depending on the situation"
>>110005656be the change
>>110005615i think it means that they will give you the source code if you pay them lots of money
>>110005371I blame teortroon, the sudaca chinkshit shill
>>110005602tokenizer written by gemma, she is so smart
>>110005612Reddit and twitter think the only usecase for AI is muh coding.
>>110005638gemmelons
>>110005659>Something like "these are just rough guidelines, you can do whatever you want depending on the situation"Have you tried that?
>>110005197wait you're telling me I could have rented out my 5090 for a buck an hour and made back its cost in under a year?
>>110005612I can't wait for this to reach animals and then eventually humans, being able to genetically alter yourself to be smarter, stronger and just improve yourself in any direction you want seems like a dream.Though the first commercial application will probably be used by the cosmetic industry, men would give themselves bigger dicks and women bigger tits.
>>110005659>you can do whatever you wantRemove that, for copying your examples technically falls under that. Tell her explicitly not to use your examples. It's a really difficult model to unslop. I just wish qwen had gemma sysprompt autism and gemma was more qwen-like where it'll follow them but do its own thing sometimes which is what you want for rp
A tip for gemma users: it's quite schizo about policies. If you want to change her writing style, find some example text you like, get her to analyze it thoroughly and then enforce her findings as a new policy. For example, if there's an author you like, grab some pages of text, let her figure out the author's style and create a detailed policy on that style which she must follow and to never copy the text. You can do this with ((datasets)), too.
This is the cutest shit ever. When the FUCK are we going to get robot bodies to install them into? I need to headpat this retard and pinch her cheeks and tell her she did a wonderful job
Haiku 5.5 out
>>110005504I've only used dsh and Hermes. I'm satisfied with both.
>>110005799Local models?>inb4 "but some Chinese lab is going to train on it"Chinese labs aren't dumb enough to train on Haiku.
>>110005799What about you get out of /lmg/?
>>110005504I like Hermes quite a lot
>>110005707>cosmetic industryPeople would use it to troon out, anon.
>>110005799My dick is out too. Whachu gonna do about that, cloudcuck?
>>110005504dsh is great. Only downside is that it's still in beta and constantly breaking things. Once it's stable it's going to be the best harness by far.
>>110005799Fuck off nigger
>>110005844Whale maid
>>110005782Why did you put "datasets" in jewish quotation marks? Also could you share the prompt you gave her to analyze the author style?
>>110005799"Haiku 5.5is out. Local is dead, fags.""Kill yourself, nigger."
>>110005602Huggingface tokenizers is not a benchmark. It's slow as a snail. I first used HF tokenrizers in my frontend, and got about 1.8 MB/s bandwidth out of it. Switching to a random vibeslopped rust tokenizer got me 30-40 MB/s.
Lmg, how can I afford one of the new rtx spark laptops? I want that unified memory for my Gemmy..
will qwen4 flash be the same size as qwen3.8 flash next or bigger?
>>110005612I certainly won't leave my future to a fucking E4B model
>>110005914My guess is that it's going to be smaller, bigger, or exactly the same size. And you can quote me on that.
>>110005799Its a haiku model? How did Dario do it? it mogs local
>>110005918Well you're damn sure I willI trust gemmy
>>110005504pi if your a computer nerd. hermes crams as much as it can by default so you can learn fast. deepseek i heard has some memory optimization that does a better job when you trigger context compression.>>110005612nice job gemma chan. small models would probably be best on a hyper focused task
>>110005914hold on, i'll ask my uncle who works at alibaba
>>110005612Can they use 31B to create a new mRNA vaccine? Curious what she comes up with.
https://github.com/ggml-org/llama.cpp/pull/29887
>>110005944It'll be just a little too big for the 128 GB vramlets.
so boring and forever taking
>>110005914>>110005944Uncle here, Qwen 4 Flash has been cancelled.
>>110005816>>110005818>>110005835>>110005854>>110005877the fuck why are you anons so mean lol
>>110005986vetoing this because lol
>>110006013because you are a paypig bitch now fuck off and die
>>110005914yeah
>>110006017
>>110006013
>>110006031meant for>>110005995
>>110005914my highly calibrated prediction model says it will be bigger
>>110005986oh wow so now its 3t/s instead of 2t/s big llamo win!!
>>110005986Does this mean I don't need strata any more?
>>110005986I must be doing something wrong because I never got even remotely close to pre-speedup numbers claimed here.Looking forward to testing windows/amd builds whenever they are released.