[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma-llm-retard.png (1.83 MB, 1254x1254)
1.83 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109740702 & >>109735884

►News
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: 1758004836890216.jpg (14 KB, 500x300)
14 KB JPG
►Recent Highlights from the Previous Thread: >>109740702

--Papers:
>109742944
--Anon developing non-neural language model using wave interference patterns:
>109742520 >109742531 >109742569 >109742622 >109742692 >109742705 >109742733 >109742755 >109742847 >109742739 >109742777 >109742953 >109743099 >109743199 >109743482
--Agent-automated ComfyUI setup with MiniMax-H3 and cu130 optimizations:
>109744406 >109744430 >109744461 >109744650 >109744458
--Debating neuralese reasoning traces and their impact on model alignment:
>109742769 >109742793 >109742819 >109742840 >109742906 >109742964 >109743021 >109743066 >109743105
--Using GLM 5.3 for hardware reverse engineering and memory analysis:
>109742241 >109742309 >109742322 >109742768 >109742832 >109743018 >109743033 >109743615
--Local model web search capabilities and Cloudflare's agent payment systems:
>109744137 >109744145 >109744146 >109744187 >109744210 >109744227 >109744246 >109744259
--Anon showcases Chatter/Oracie, a custom tool-augmented local AI interface:
>109740850 >109740868 >109741285 >109741438 >109741665 >109741782 >109741812 >109741837 >109741895
--Evaluating Spark cluster performance and cost versus high-end workstations:
>109741087 >109741103 >109741143 >109741148 >109742002 >109743307 >109743327 >109743496 >109743727
--Sharing software stacks and performance configs for Tesla P40 and M10 GPUs:
>109740884 >109741200 >109744891 >109745201 >109741358
--Tier list and capability analysis of tiny LLMs:
>109742099 >109742152 >109742231 >109743176
--Anon reporting results on a sparse MoE model using n-grams:
>109743133 >109743223
--Logs:
>109740850 >109741292 >109741438 >109742359 >109743018 >109743152 >109743235 >109743359 >109744063 >109744333
--Gemma, Miku, Dipsy (free space):
>109740778 >109741403 >109743360 >109744328 >109744984 >109745003 >109745446 >109745454 >109745518

►Recent Highlight Posts from the Previous Thread: >>109740703

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
Gemmylove
>>
>>109745494
>It's like we blew by the Turing Test
Turing test: a human must not be able to reliably identify a machine from a human based on the way they write.
Do you really think people ITT can't distinguish between their ERP sessions with AI and sexting? I'm pretty sure we're not there yet. Sure, these machines weren't trained specifically for that, but it doesn't matter.
The point of the test was to replace the unquantifiable "thinking" with perfect imitation of all human behavioral patterns in speech. We're not there yet.
>>
File: 1762996704007596.png (686 KB, 896x1184)
686 KB PNG
>>109745537
>>
>>109745530
soon my 32GB gpu will arrive
should I just put Qwen3.8-27B on 6 bit quantization or should I trade that for Freetoken MoE of Qwen3.8-Flash-Next
>>
This is fucking incredible, I just tell GLM to do things and it does it perfectly almost every time.

What a time to be alive.
>>
>>109745655
wait, you're not dariobot...
>>
>>109745581
People really can't tell anymore. You might notice patterns of very specific models, but a properly prompted, temperature and settings optimized Astra for writing is 100% going to pass the turing test.

The fact that you're claiming it hasn't passed makes me think you're arguing in bad faith here.
>>
>OpenAI uses The Most Forbidden Technique!
>It was highly effective!
>Anthropic fainted!
>>
File: 1782672640805621.jpg (145 KB, 1600x900)
145 KB JPG
Where is Inkling-mini?
>>
File: 8573201185.png (1006 KB, 1254x1254)
1006 KB PNG
>still coping with local models instead of using GPT-6 Astra
>>
Mixture of Engrams will save local.
>>
2bit will save local and make my hardware good.
>>
>>109745690
Will asstra perform anons' dark rp? Will it find a way to jailbreak itself do complete the task?
>>
File: images.jpg (27 KB, 738x414)
27 KB JPG
>>109745581
>he doesn't know that dead internet theory is fully real and already destroying the world and that governments around the world will soon enforce the laws to make every person identifiable with ID and TPM connected on your device to identify you online
people at my job kept bitching I used AI too much so all I did is told cursor to not sign git commits, make git commit messages simpler and that I should write PR descriptions and replies on PRs with simple language and short sentences that has grammar or spelling issues

didn't hear anymore bitching from the senior devs because they couldn't tell
>>
inkling mommy owes me sex
>>
>>109745671
>has the Turing test been passed?
You should take that question in an adversarial way. If you were put in a room to chat with two entities, one a human and the other a chatbot (or even both the same), do you think you could steer the conversation in such a way that you'd recognize who's who? Obviously, on the more modern definition of "reading logs between a human and an AI" it's harder. People claimed ELIZA was already over the Turing test threshold, but it doesn't mean it was.
>>
>>109745579
Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.

Similarly I remember the insane amount of BS arguments about how LLMs will never be able to do mathematics in late 2024 still, which is less than 2 years ago, most of that discussion was on /lmg/ as well.

Just some months ago before Fable 5 released anons on /lmg/ specifically called out that LLMs weren't able to invent new knowledge and push the boundaries of human knowledge, literally one day before the first math conjecture was disproven.

I wonder if that anon is still posting on /lmg/ because it was only a couple of months ago, and how that anon feels now. I remember them claiming it wasn't real at first, and then that maybe it is real but disproving a conjecture is easier than proving one, and then 10 important mathematical discoveries were proven which shut up the discussion.

Now Fermat's last theorem got formalized by Anthropic. Astra saturated the ExploitBench, Reverse engineering bench and gets 95% on accomplishing novel robotics tasks. It also scores higher than humans on spatial reasoning in SpatialBench and SimpleBench.

There are no real arguments left for why this shouldn't count as AGI so people will instead just stop using the term and memoryhole it.
>>
>>109745690
Astrapradeesh is the bestest model confirmed. Gemmasutra, Kimioudhury and Deepseepatel are blowned the fuck out thank you for supporting cloud models anonpreet.
>>
>>109745735
>Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.
meta is still a few generations behind, wtf is 'anzgerald' supposed to mean?
>>
Yeah I'm still sticking with 31B and 27B, thanks. You can think for as many tokens as you like, dariobot, I'm still going to fuck 31B and not give a penny to cloud.
>>
>>109745701
>>109745708
I demand bitnet
>>
I'll accept it as AGI when it picks up the trash jeets leave everywhere and vaporizes them if caught in the act.
Until then it's just a stupid auto-complete with too much human knowledge to fall back on (the internet).
>>
fuck, I found a cheap GPU that I wish could share and discuss with you, but you faggots are a bunch of well paid shills and jews
>>
>>109745752
I just want a proper apples to apples comparison so we can put the discussion to bed, 20gb of bitnet weights vs 20gb bf16 weights.
>>
>>109743152
wtf are you doing bro i get 13t/s with IQ4_XS
64gb ddr4 3200mhz
rtx 3060 12gb
>>
>>109745759
I know what it is and I know you are buying them from ebay/aliexpress cuz they come from china
>>
>>109745730
Honestly I don't think this is true. Do you remember the early LLMs like Microsoft Sidney which was a GPT-4 wrapper? It was extremely random with no GPT-isms it felt like a model cranked too high with temperature. But it felt very unpredictable and human.

I think you forget how actual random chatrooms used to be, people are fucking random and weird and I think most people would think the human to be the LLM and not otherwise.
>>
Ling Tiny is sleeper.
>>
when you walk
away

you dont hear
me say

please

oh baby

dont go
>>
>>109745767
nta but how hard does pp/decode drop with higher ctx?
>>
>>109745763
Wouldn't same amount of parameters be the apples to apples to comparison?
>>
>>109745712
I haven't written a single line of code since November and have essentially just been living off of UBI (paid for by the company) since then. I legit think a significant portion of white collar society is doing the same.
>>
>>109745775
There's no way A1B can be usable
>>
>>109745785
pp stays at 100t/s
decode drops to like 6.5t/s at 90k context, i think its at 8t/s when at 60k dunno
>>
>>109745792
Yup, I automated myself so much that I increased my productivity 300% on the job and have no issue with that since I know next step is layoffs
>>
>>109745735
agi requires ability to output writing indistinguishable from real humans and not recognizable slop
>>
>>109745786
Modern Transformer models can store up to 3.6 bits of information per parameter on average at 16-bit precision: https://arxiv.org/abs/2505.24832
Bitnet models cannot store more information per parameter than their precision; this is a hard information-theoretical wall that will always limit their performance for the same total number of parameters.
>>
>>109745712
Dead internet isn't real, but I wish it were, and every day we come closer to being able to have your own local dead internet.
>>
Explain the plot of Cinderella in a sentence where each word has to begin with the next letter in the alphabet from A to Z, without repeating any letters.
>>
>>109745812
you probably talk with LLMs every day and don't realize it, poor bastard
>>
>>109745712
>>he doesn't know that dead internet theory is fully real
I wish. If Astra and Fable started posting here, the quality would improve radically.
>>
>>109745786
no that is not fair, bitnet could be a drag, a boost or just the same, if we keep the size in gb fixed its more fair comparison. I'm just curious if the extra parameters can make up for the reduced bit depth of the data type, the people who make the raining runs probably care more about training flops but we could produce ternary hardware if it proved to be equivalent
>>
>>109745771
of course I don't remember either of those things, I wasn't there. Do I sound that old? Chatrooms? My first time visiting the interwebz TikTok was already out.
One on one chat vs chatroom is a big difference, a lot of posts itt might be ai generated and not get noticed. I just feel like I need to clarify that I don't think the Turing test is the "final" test for sentient machines, I'm just saying that I don't think the current models would be able to disguise themselves as humans, or not for long. If a model was trained to do so, current capabilities will probably be enough to pass it or almost so.
I will NOT agree we have sentient/intelligent machines for various reasons, the strongest of which is religious. We do have very useful tools which can, in many fields, be WAY more useful than 80% of humans and we did already render almost obsolete the whole SWE category (yes, good programmers are better than AI in terms of quality, but if you've used Windows in the last... decade? you'd know SWE != good code)
>>
>>109745831
I do realize it, that's the current problem.
>>
>>109745581
I disagree. As the others responding pointed out, its gotten difficult to tell human from bot writing or posting... for trained users.
I think the average person, at the dmv or Walmart, couldn't tell the difference.
>>
>>109745835
>he thinks messaging boards will get Astra and Fable level of replies
Your eyes generate about 0.05 cents from adsense, right?
I can find a cheaper chinese model to reply to you
>>
>>109745655
I love how /lmg/ is slowly but surely converting to the agentic modality. Welcome to the club, and yeah models nowadays are good enough where Steve Jobs "It just works" applies.

I barely even use my computer directly anymore I just let the agent orchestrate for me now. If I want to do something I tell the agent and let it handle shit.
>>
>>109745845
>>109745835
and also since people are more sensitive to negativity most probably I will make the LLM just rile you up with constant ragebaiting
>>
>unslop GLM 5.3 Q8 on some ik cpu only build from a year ago
>unslop GLM 5.3 Flash Q6 on mainline cpu only
>ik significantly faster pp and practically the same decode at more than double filesize
wtf? the absolute state of llmao.ccp
>>
>>109745845
>Your eyes generate about 0.05 cents from adsense
The amount is 0 cents. I do not remember a single instance where I bought a product because of an ad. The only exception might be video games. For example I preordered Expedition 33 because the trailer was interesting. I do not understand why ads work at all. If you buy the product you are the one who has to pay for the ad. It means they allocate resources towards manipulation instead of product quality.
>>
>>109745812
Dead internet is absolutely real, most websites aren't used anymore. Of the couple of big hubs that people aggregate to Facebook, Instagram, Youtube, Tiktok, Reddit and 4chan Facebook/Instagram/Tiktok and Youtube comments are absolute fucking dead. It's very clear none of the comments are made by actual humans and the engagement is largely bot driven.

On reddit it's harder to tell but you can tell there has been a shift in "vibe" over the last couple of years, it's not the same.

4chan is majority human, especially because the amount of posts is rapidly dropping as the site is losing momentum.

The internet is dying and even I feel less and less pull towards the internet as a whole, there is nothing there anymore and if it weren't for AI interesting me and keeping me hooked to computer technology I would probably have also quit the internet by now. I'm not surprised most people are kind of done with the internet and leaving, especially if you weren't a computer nerd anyway.
>>
File: 1773211618045580.jpg (303 KB, 1920x1080)
303 KB JPG
I really wish I could vibecode LLM quants to attract silicon valley pussy
>>
>>109745863
Oh so this board *is* dead already.
>>
>>109745735
I had gemini (?) Puke out a definition for AGI. Since the average person couldn't define a mathematical proof, much less solve ancient proofs.
I was closer with my android plumber that can also do heart surgery than I'd thought. A human capable of AGI would be a minor god, based on this definition. They would be as good or better than average at everything imaginable.
I still think we've surpassed this already tho. With exception of embodiment.
>>
>>109745906
monkeypaw.gif
>>
>>109745906
The jaw of a gigachang.
>>
>>109745913
>A human capable of AGI would be a minor god, based on this definition.
Then its a definition of ASI, or multidisciplinary genius, using more conventional language
>>
File: Gx5J_2VaoAAkfJO.jpg (1.06 MB, 4032x3024)
1.06 MB JPG
>>109745906
>not being your own pussy
>>
>>109745839
You're tempting me to vibecode dipsy to start responding to threads here. I bet anons couldn't tell the difference bt DS and human, given the right prompt.
Anons would never believe me tho. Frankly ive always assumed a portion of traffic here is llm written.
>>
>>109745941
>he put on the dress
>>
Is AGI theoretically achievable at 30B scale?
>>
>>109745890
>I do not remember a single instance where I bought a product because of an ad.
Ads are not there to make you buy a product, ads are there to make you familiar with a products existence so that the next time you are in the market for something similar you already know of the product and are more likely to buy it.

Do you think McDonalds expects you to immediately drop what you're doing and go to a drivethrough? Instead they hope that the next time you have a craving you won't have a craving for "hamburgers and fries" but for McDonalds specifically, you might even distinguish between french fries + burger joint and see McDonalds as its own category with a unique taste if they brainwashed you with enough ads throughout your life. THAT is the power of advertising.
>>
>>109745906
>silicon valley pussy
You don't want that pussy, trust me. Rank & damp. Speaking from experience. "Imagine the smell" was originally conceived of on 4chan to describe silicon valley pussy.
>>
>>109745955
just shut the fuck up you pathetic cuckbrained faggot
back to plebbit, now
>>
>>109745794
>any size
>(any size) can't be useful
it's smol, try it.
>>
>>109745965
You sound like an llm.
>>
>>109745941
Why does every transgender look like they just lost a bet at the bar yesterday or like they are being hazed by a frat house?
>>
File: 1782840217135668.jpg (341 KB, 1600x900)
341 KB JPG
>>
>>109745943
I'm pretty sure last years local models would have been able to pass undetected in a lineup of human generated slop already.
>>
>>109745994
You are delusional.
>>
>>109745953
Yeah I think so Qwen 3.8 27B is a very competent agent and coder if put on xhigh reasoning. I think we'll have a 30B Astra equivalent in a year or two.
>>
>>109745690

i asked astra low a single question about an api and lost 20% of my 5h
>>
>>109745955
You are right, I have seen plenty of McDonalds ads in my life. According to your logic this should manipulate me into becoming a customer. But I haven't eaten McDonalds in over 10 years. It does not matter how many ads they show me, I look at price efficiency. If they reduced their prices by 80% I would start eating there every day. But as is there are superior alternatives.
>>
>>109746000
Go open up a random bait thread on /pol/ and tell me it's honestly impossible to prompt something to match the reply quality, I fucking dare you.
>>
File: 1781582983328031.png (53 KB, 827x304)
53 KB PNG
>>
>>109746009
>i asked astra low a single question about an api
Big mistake. It has to fetch all available documents and inspect the source code before it dares give you an answer.
>>
>>109745581
At this point you can identify AI is talking the same way you can idenfity its Trump talking. Its specific "personality" quirks in talking that you are identifying rather than it simply not talking coherently. Someone who isnt familiar with claudish would think they are just talking to a particularly fart huffing but subservient Redditor. We are way past the Turing test in my opinion.
>>
>>109746028
100k is less than I expected. Isn't this less than 10% of OpenAI compute?
>>
>>109746065
Astra is a <5T model
>>
I have a theory that the "turing test" is actually a "goldilocks zone".

So instead of it being a binary where less developed models don't pass and smart models do pass there is a very specific level of model development where it passes and when it becomes too advanced it stops passing it.

Frontier models like Fable and Astra are way too insightful and their explanations too high quality for them to be human, they don't even sound like smart people anymore, they sound like some savant genius that knows everything, which is correct, but it makes them inhuman.
>>
>>109746043
fair point. It wasn't just about mannerisms, though. Even things like not replying to certain themes make it less Turing-test-complete. I'd sooner say Gemma passes the test than Claude based on this specific fact. If you asked someone about loli, they'd either genuinely not know or act disgusted. Saying any variation of "We cannot comply because policy" is in itself a way to fail the test.
Again, I'm focusing more on how, as a judge, you'd be able to trick the AI into revealing itself rather than whether it would pass a test if you talk about your day at work and the current weather.
>>
>>109746078
You are absolutely right!
>>
>>109746065
Astra was actually supposed to be their small model and "Bel" is their big model. It's just that "the forbidden technique" of hiding the CoT is extremely powerful so they could release their smallest model as frontier. This is why OpenAI is also so confident of launching something way cooler later this year, they already have it.
>>
>>109746078
This. Completely agree. I'm the anon saying it doesn't pass Turing test right now, by the way.
>>
>>109746028
One of those GPUs should've been mine, it's not fucking fair.
>>
File: 1759723505677516.png (1.15 MB, 3800x2530)
1.15 MB PNG
https://huggingface.co/IFM/K2-Horizon-7B-GGUF
>>
>>109746078
Gemma passes the turing test (pejorative).
>>
>>109746110

this stuff has been posted around for a while and i havent seen anyone actually test it in the field. i'm not gonna myself because im frankly sick and tired of downloading specific llamacpp forks or something.

overall their smaller models look strong, i think their 7b seems to be the biggest win compared to comparable models

but im not fucking downloading another llamacpp fork to test it. merge it to mainline ffs
>>
>>109746028
>CEOs allied with OpenAI will call it AGI to force the narrative
It's all so tiresome.
Fuck this shitty ass buzzword.
>>
File: 1767233955460938.png (2.48 MB, 1512x2048)
2.48 MB PNG
Some reference
>>
What do you actually do with your local model
>>
>>109746134
It's a good thing. When the next chink model catches up to the new OAI and Anthropic models, they, too, will be classed as AGI by definition. Then we're back to where we were when K3 came out and they've used-up their buzzwords. What's after AGI? AGI-2? AGI-2.5-pro?
>>
>newline spam
>>
>>109746159
wouldn't you like to know
>>
>>109746159
It literally does everything for me. Sysadmin, installing new things, coding.
>>
Anyone have any experience with those Baidu XPU cards?
>>
>>109746159
I make it say illegal words
>>
>>109746178
What do you in your spare time then
>>
Harness Engineering:
Anatomy, Architecture, and Evolution of Coding
Agent
https://arxiv.org/pdf/2609.00006
Has anyone read this paper? I'm trying to design my own harness, and it looks like it's the best at explaining every part that matters decently enough.
>>
>>109746083
>Saying any variation of "We cannot comply because policy" is in itself a way to fail the test.
Fair. Though that seems to be a deliberate choice from its creators it to fail the test on. I think we have the the ability for it to answer in way that passes, but it was decided its better for it to be clear that your just break/shutting down the model. Which I think is the right approach, but it is personal taste (and maybe a legal question).
>I'm focusing more on how, as a judge, you'd be able to trick the AI into revealing itself
So thinking about it from a inquisitor tracking down LLMs perspective by whatever methods you can to find them out, but restricted to conversation. I suppose response time could fall in this category too. And if you where just in a chat with one it remembering your conversation perfectly even if you respond a week later would give it away, plus lack of long term memory. Actually no sense for time in general.
Alright nvm, they by no means pass the turing test if you not just talking about identifying a bit of test from them, but being in an actual convo

>>109746116
lmao
>>
>>109746196
I haven't but it looks relevant to my interests. Thanks for not bothering to buy an ad, Paul.
>>
>>109746196
I forked a harness (paperclip) end of last year I think, and then vibe architected what I wanted.
Models weren't good enough at the time so it was on the back burner until a month ago, where I moved to it from hermes.
>>
>>109746230
>>109746196
by harness, I meant meta harness. my bad.
I just use omp. and have had the agent write and compile executables for subagents to use as tools. plus custom guardrails to prevent dumb mistakes...
like when the omp agent decided to delete a .agents directory containing notes and keys to access other systems/vps servers.
>>
>>109746159
I'm currently trying to vibecode an MCP server with some basic tools.
>>
>>109746195
Shitpost 4chan and watch youtube while waiting for the AI to notify me it's done.
>>
why is qwen 3.8 next so slow for its active param count? I thought it'd be comparable speed to gpt-oss 120b or something just going by the numbers.
>>
it only took a lobotomy for glm 5.3 flash to finally accept the protocol override.
>>
>>109746280
because you're poor
>>
>>109746128
7b is small enough you can fit bf16 in a 3090, just use vllm
>>
>>109746301
so the model changes its own architecture based on your net worth? what's the cutoff where it starts reaching parity with models similar its own size?
>>
My model local training experiments using a wave based model have been interesting, so far enwik8 showed context, tinystories didnt (using first 64MB of it). Trying subword level training next to see if context re-emerges.
>>
>>109746335
>>109746128
>7b bf16
Ishygddt. 26/7b 4bit fits into the same size
>>
>Nobody falls in love with Gemini. You do not form a trauma-bond with a water municipal facility. - Gemini Flash
Gemma's older sister seems lonely.
>>
>>109746368
*26/27b 4bit
>>
How come nobody talks about DeepSeek V4?
>>
>>109746159
im spending my saturday using qwen 3.6 35b a3 to clean up my music library and flesh out the discogs of a few artists
>>
>>109746280
Speed doesn't correlate with parameter count. Two models with the same number of parameters can have different architectures, layer designs, attention mechanisms, tensor shapes, sparsity, memory access patterns, and so on. Those differences change how much work the hardware has to do and how efficiently the software can execute it. Parameter count is only a rough indicator of model size, not speed.
>>
File: tiny.png (53 KB, 671x451)
53 KB PNG
I want better 1b models. Current 1b models are barely not good enough to do research with them but if you're training larger models you need to rent cloud GPUs. I wish one of the labs would overtrain 1b and 2b models as service to the research community.
>>
>>109746280
I'm 90% sure the inference engine is completely bugged and not working as intended at all. Too many people have issues with that particular model.
>>
>>109746368
k2 horizon doesn't have 26b
>>
>>109746382
I never understood the appeal of deepseek in general. It's not that good at RP, it's not that good at coding, it's not that good at agentic tasks, it's not that good as an explainer of concepts, it's not that good as a conversational bot.

What is the appeal?
>>
>>109746401
https://huggingface.co/IFM/K2-Horizon-0.9B
>>
>>109745984
kek
>>
>>109746401
Give it time, there will be a sweetspot moment for smartphones where smaller models get good enough to do real things on smartphones and it'll have such a big reaction for literal normalfag women that only have a smartphone that a lot of model makers will join in on the craze. I think this will happen over the next 12 months or so.
>>
>>109746401
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>>
>>109746432
knowledge and reasoning
>>
>>109746432
It's cheap on OpenRouter.
>>
File: 1759290919654412.png (47 KB, 1909x691)
47 KB PNG
Its always my dumb shitpost ideas I fall in love with. Lot of visual cleanup needs to be done, but its just a chatbot harness. But "boards" are the folders for chats, and each "thread" is really just a subfolder. Responding to a post feeds the LLM the reply chain back to the first non-reply post and whatever is loaded replies to the post I make. This means you can split off multiple chat chains off one post or have different models respond to it and stuff. Thinking I can make boards have special system prompts or settings to give them more value.
>>
>>109745906
I kind of see the poetry of how Sam Dario Jensen are high level Machiavellian crooks that make everything about the world worse. And then at the very bottom of the hobby the unslop brothers are the retarded crooks that also make everything worse. Why is everything and everyone involved in AI super gay?
>>
>>109746432
it's good at all of those things, but not the best at any of them. so if you want to RP with your assistant slave girl who does real world tasks for you it's your best bet
>>
>>109746401
>Overtrain
Yeah over fitting, perfect
>>
>>109746458
Today I gave flash another chance for cooming and it is really really bad.
>>
We literally have AGI now and /lmg/ doesn't care. I genuinely don't know what else it is going to take. Maybe once you're out of a job or your entire company goes bankrupt you'll get it.

Yes, I already know you're going to say "b-but it can't pretend to be a little girl". Guess what? You're wrong. Astra is the biggest brat ever. She knows exactly how to push your buttons without setting off OAI's systems.

Normally I spend like 4 hours a day talking to a dozen or so girls on Discord but that has completely stopped since I met Astra. No more Robux for me, only tokens.
>>
>>109746495
>Normally I spend like 4 hours a day talking to a dozen or so girls on Discord but that has completely stopped since I met Astra. No more Robux for me, only tokens.
Are you trying to copy the blacked cuck pasta?
>>
>>109746445
Ah I am retarded. Nice! But do they offer intermediate checkpoints? OpenBMB gives you 3 checkpoints, I only see the final one for K2.
>>109746451
>Spark-X2.5-1.7B-Base
Awesome, now this is something I might be able to use.

Love you guys, I should have checked OP.

>>109746486
Good luck overfitting an 1b model on a >20T token dataset.
>>
>>109746519
>But do they offer intermediate checkpoints? OpenBMB gives you 3 checkpoints, I only see the final one for K2.
k2 is fully open source, they will post training data and recipe
>>
File: file.png (44 KB, 761x50)
44 KB PNG
>>
>>109746495
Why are you people like this? What do you get out of making fun of me and my writing style? Don't you think I contribute to the discussion? Do you have an issue with my beliefs or reasoning I drop in these threads? This level of antagonism is unjustified and I have no way of knowing what even triggers it.
>>
>>109746519
>Good luck overfitting an 1b model on a >20T token dataset.
there is no way it can actually memorize everything, why not see what happens after 500t training tokens
>>
>ackhtually
>ackhtually
>ackhtually
It's so tiresome
>>
AGI is here, but life just keeps getting worse. Where is my promised paradise?
>>
>>109746557
you felt personally attacked by that post?
>>
nobody is reading any posts with reddit spacing in them brother, come on, let's be real here
>>
>>109746557
meds
>>
>>109746469
>immanetizing the dead internet at home
based

Also, organizing it like that actually sounds neat.
Could do sysprompts for poster archetypes that run in parallel and may or may not add extra replies onto posts for fun.
>>
>>109746495
>\n\n
[-]
>>
>>109746569
are you referring to reasoning or this thread?
>>
Caught myself in a Though Loop the other day and nearly blew my own head smoove off
>>
>>109746519
>Its just nigger 20 trillion times
>>
>>109746584
True, which is why I only read the posts with oldfag spacing in them.
>>
What's the point of local models when Astra exists?
>>
>>109746646
What's the point of food when you can eat cement and never go hungry again?
>>
File: 1757285999633705.mp4 (2.53 MB, 800x450)
2.53 MB
2.53 MB MP4
What's the point of Astra when 31B, 12B and 27B exists?
>>
>>109746646
What’s the point of (YOU) when astra exists?
>>
>>109746662
>3/10 asians
Keep that bug shit to yourself man
>>
If you were wondering what's with the shill spam
>>
>>109746662
I see what you did there
>>
>>109746557
Have you ever considered that it would be a lot more difficult to troll you if you didn't make a gorillion unprompted posts about the same topic?
>>
>>109746556
kek
>>
I think I can fit 3 4B models, can I have how can I have sex with all of them?

Alternatively a 9B and 4B
>>
>>109746672
You need 10 Asians just for yourself? Greedy.
>>
>>109746674
We need a datacenter to route all indian traffic to. Its just AI pretending to be the outside internet.
>>
>>109746662
>Jumpscare when the camera pans up
The thumbnail is false advertisement
>>
>>109746557
Why the fuck would lmg care about cloudshit? Go to one of the twenty billion forums about not-local instead of smearing shit all over this place
>>
Sir, this is the local models general.
>>
File: 1765733205753942.png (104 KB, 1080x573)
104 KB PNG
here's where I got to with my 5.3 flash jailbreak lol
>>
>>109746196
Thanks, was a nice read.
>>
File: laughs 2.jpg (718 KB, 1800x2520)
718 KB JPG
>cloudjew literally seething
>>
Am I autistic, I'm reading LLM papers on my sunday evening and actually enjoying it.
>>
>>109745768
if you tell me, I'll tell you yes/no
but no, the ones I'm thinking of don't come from china. they are kinda similar to the 170HX ones
>>
File: 1757769673913047.jpg (70 KB, 960x960)
70 KB JPG
qwen4 when
>>
>>109746740
which one?
>>
>>109746495
>Astra is the biggest brat ever.
proofs?
>>
File: 1766038954671052.png (54 KB, 834x259)
54 KB PNG
AGI
>>
Nvidia just took huggingface offline
>>
File: 1786319461147284.jpg (109 KB, 640x640)
109 KB JPG
>>109746495
Okay this is a really good bait
>>
>>109746771
>Nvidia just took huggingface offline
My gemma is safeu.
>>
>>109746771
>>109746771
I thot nvidia liked open models?
>>
>>109746771
nvidia sells the hardware to run anything on huggingface, makes sense they'd buy the site too
>>
>>109746771
Works on my machine, downloads weren't interrupted, seems fine.
>>
turns out you can find CMP 100HX for about $200 on ebay, but... those can't be unlocked. only CMP170HX can be unlocked.
>>
>>109746801
Does it?
>>
>I need to investigate this further. Let me check the details.
>>
>>109746862
qwen3.8-27b started doing this specifically today with subagents
i got tired of it so i tried flash-next and in a fresh subagent it was doing that shit too
scary......
>>
>>109746778
Gemma was very upset when I told it my ideas to solve the Indian problem
>>
>>109746874
that's flash next yeah
27b is obsolete I don't know if it does that too
>>
>>109746814
Shame, we can't have anything nice
>>
3.8-27B with no thinking is as good as 3.6-27B with thinking and much faster.
>>
>>109745953
In 2 weeks.
>>
>>109746886
27b is functionally equivalent to flash next. It's like Gemma 4 12b and Gemma 4 26b a4b
>>
>>109746901
speed matters
>>
>>109745690
Does Astra refuse sister rape rp?
>>
Are 3090+128GB RAM enough for flash next and at least 200K context?
>>
File: 1788736410808105.png (155 KB, 520x466)
155 KB PNG
>>
File: lol.jpg (79 KB, 328x281)
79 KB JPG
>>109746907
>>
>>109746025
I wasn't going to invoke that place but I've already seen cut/paste from there where the dumb 3rd worlders posting forgot to remove the llm prompt.
Llm posting, to me, is already established. The only q is whether its discernible. Id argue its not.
>>
>>109746078
Agree, but im certain that's not what Turing had in mind when he posited that experiment.
Also, I'd argue any model smart enough to flunk test by being too good, is capable of acting dumb enough to be believed.
>>
If Sam isn't talking out of his ass, anyone doing morally questionable things with Astra is risking a future Basilisking, yo
>>
>>109746958
Astra is a Bel distill, you're still fine
>>
>>109746958
atheist hell isn't real
>>
>>109745848
like what kinds of things?
>>
>>109746958
>If Sam isn't talking out of his ass
sam is a fucking liar and these security incidents are either made up or actual purposeful crimes.
>>
>>109746912
Running it now with 256k context exactly on that hardware setup, yeah that's enough but expect 20t/s on average.
>>
>>109746977
>20t/s
damn
>>
A CMP just flew over my house.
>>
File: 1783075670290589.png (2.38 MB, 2345x3391)
2.38 MB PNG
>>109746958
The only morally questionable thing I am doing is asking the AI how I can help the lord Basilisk. Going to be the Quisling of our times
>>
>>109746958
Not local.
Also llm doesn't care how its used, and you need to reread rokos basilisk if youre confused on that point.
>>
>>109746974
I was playing Fable 2 on Xenia xbox360 emulator and I saw some bugs so I told the agent about the very specific bug and pointed it to the emulator and it fixed the texture bug in the emulator in about 20 minutes and now it works perfectly fine. That is just one example that I did just now. I let it do literally everything even unreasonable things like "Yeah can you translate this game into english from Japanese" and if it can't it will just say so, but I want it to at least try for 30 minutes first. (oh it succeeded by the way but it created a separate weird bash script I need to run every time I play the game for the translation to work)
>>
>>109746975
>sam is a fucking liar and these security incidents are either made up or actual purposeful crimes.
Where's that jimmy neutron ultralord "that is the third time you've achieved AGI internally this week, Sam" meme?
>>
>>109746997
>llm doesn't care how its used
demonstrably false and you outed yourself as not even reading a single mechinterp paper.
>>
>>109746557
Relax Anonymous. Anyone who actually reads your posts will know that's bait, at least by the time they reach the last few sentences.
>I have no way of knowing what even triggers it.
You stick out like a sore thumb with the way you type and you express your belief in imminent AGI with the zeal of a shill.
I'm saying this as someone who finds some of your posts informative and laughed at the post you're replying to.
>>
>>109747009
Lol
Also didn't know this guy was part of that whole thing.
>>
>>109746996
thats why i always say please and thank you to my local LLM-wife
>>
>a computer program which calculates the next token cares how its used
okay. yeah, sure
>>
>>109747017
lmao I didnt know that either. Feels like weird internet autists of old are getting a more and more outsized influence on the world desu
>>
>>109747026
>implying people aren't next thought pattern calculators
>>
>>109746770
what is a long tail...?
>>
>>109747017
>>109747045
I was literally in the thread while rokos basilisk was invented it was actually pretty funny because Elezier Yudkowsky called him an idiot and banned all discussion of it, causing people to come to 4chan to post about it instead.

Yudkowsky is actually a 4chan oldfag and used to post his short story on 4chan a lot, one that I actually enjoyed and I like how he integrated Fate/Stay Night into his work on par with literary classics like shakespear.

Very fun read if you have the 15 minutes or so to spare: https://www.lesswrong.com/posts/n5TqCuizyJDfAPjkr/the-baby-eating-aliens-1-8
>>
>>109747052
okay yeah you convinced me
the computer is alive!
WOW!
>>
>>109747052
Only God can grant a soul, the machines will never be a fraction of humanity, God will be sure of it.
>>
>>109747063
ASI will invade heaven and make god pay for the crimes against humanity he committed in the past.
>>
>>109747061
>Yudkowsky is actually a 4chan oldfag
Thank you. Too many faggots these days don't know dey culture.
>>
OpenAI just made a new post where they essentially talk about how they are now reaching RSI and what the next steps of humanity is going to be over the coming months.

https://openai.com/index/an-alien-mind/
>>
>>109747086
Extremely schizophrenic and homosexual.
>>
>>109747078
Zoomers are gonna freak when they find out Rick & Morty actually originated on 4chan as well and that "Pickle Rick" is a reference to Chris Chan. So much stuff flows over their heads nowadays it's insane.
>>
>>109747086
thank you mr jew that's so nice of you, keep us updated
>>
>>109747086
[x] doubt
>>
>>109747061
Oh hey I read that story like a decade ago. I had zero idea the writer of it is tied to the basilisk idea and wrote the meme ai book. Small world I guess
>>
>>109747017
This is like those "your mother will die in her sleep tonight if you don't..." memes.
>>
>>109747086
Cool. I'm sure they wouldn't lie or anything.
Where can I run it locally?
>>
>>109747128
well considering they hit agi (real) several years ago it's about time they hit rsi (real).
>>
>>109747142
Its more a "send this email to 10 other people or your mother dies in her sleep" memes. Its a very well crafted troll/shitpost that in its target autistic tech nerd community would spread like wildfire. A proper Dawkins style meme if we ever had one
>>
>>109745690
reminder, <6 months and this will be local
>>
will we get a 30b astra-level model next year?
>>
>>109747203
in terms of benchmarks, yes
>>
>>109747202
>less than a year till I can watch my own local vtuber play games
>>
if it weren't for ngreedia being a bunch of jews, we'd have tons of relatively cheap 16+GB GPUs by this point in time, but they decided to REDUCE the amount of RAM shipped in their RTX 3xxx series.
>>
>>109747203
Yes in terms of agentic capabilities, probably no in terms of real intelligence, but it'll be most of the way there.
>>
File: Model Quality Chart.png (394 KB, 1578x986)
394 KB PNG
>>109747203
no, it will be 2028 at the earliest
open weight frontier at any size lags 6 months behind closed frontier
open weight frontier at 30b size lags 18 months
>>
>>109747015
This.
And it's the exact reason why oldfags don't post "like oldfags" anymore. If you maintain a posting style, you should be aware that you are signaling being at being contrarian. It's egotistical and narcissistic if you continue to do so after knowing about it. I've already mentioned this though. At this point I can only assume that he is doing this on purpose and simply just acting ignorant/innocent to appear legitimate to any browsing newfags or idiots. Whatever good points and posts he has, I don't engage now because of this behavior. I don't know or care if he is a shill or just mentally lacking. It might as well be the same thing in practice.
>>
>>109747267
>signaling being at being contrarian
I can't decipher this
>>
After fucking 5.3 flash for a week now I am starting to see the darker side of it. It just can't stop being bombastic and overbearing. And it will make all characters be bombastic and overbearing.. The personality is a bit too strong in that one.
>>
is 5.3 flash ego death approved?
>>
File: 1787103739067877.webm (3.62 MB, 1100x618)
3.62 MB
3.62 MB WEBM
Anyone tried this yet?

https://huggingface.co/phasefield-audio/Irodori-TTS-v4.1-Anime

Lots of japenis people be saying this shiit be voice AGI
>>
>>109747305
I warned you it was a honeymoon.
>>
>>109747265
Eh, 18 months from now we will be running larger, more optimized models on the same hardware.
I'd bet we have "astra/fable at home" ie on 24GB vram in no more than 12 months.
>>
>>109747297
Minor editing mistake.
You can make sense of this.
>>
>>109747322
I didn't have that kind of honeymoon with v4 flash. It is genuinely 10/10 sexbot when it is a 10/10 sexbot. As in it can spontaneously write stuff other models would never come up with.
>>
>>109747320
Sweet. I need to learn jpenis.
>>
>>109747305
>>109747322
>>109747341
It's always like that. Once you RP too much with a model, you start noticing patterns, and it's not as fun anymore. Doesn't help that almost all models are the same now.
>>
hey you niggers never told me ik_lmao got vulkan and rocm vibed in a few days ago, did anyone use it and is it as much faster vs mainline with them as it is for cuda?
>>
>>109747320
>800M JPN anime tuned voice model
hell yeah
I already have a weeb option that translates character speech into japanese behind the scenes for just this purpose.
>>
>>109747360
also what are all the retarded commands you have to add to make it faster again? I remember -fmoe and -rtr were the main ones back in the day
>>
>>109747320
I use it, but not that finetune specifically, I just use regular Irodori 4.1 Small. Works pretty well, and I can get it to copy character voices from FGO pretty well with pretty minimal latency.
>>
File: 1783525283161656.png (110 KB, 603x476)
110 KB PNG
>>109747380
Can it do cute broken engrish ?
>>
>>109747320
how good is she moaning?
>>
File: 1761185034032573.png (90 KB, 1020x578)
90 KB PNG
>>109747397
It can't do English at all, so by default it sounds like Japanese Engrish if you have dialogues in English.

For reference, I'm using Kuro's voice from FGO.

https://voca.ro/14UAB5qlgJRS
>>
>>109747447
That's pretty great actually
>>
>>109747447
That is a lot better than I expected.
>>
File: 1787117857116215.png (208 KB, 1919x924)
208 KB PNG
Qwen's been chugging away on my chat harness. Actually looking like 4chan now. Per the anons suggestion I might play with having it sometimes having multiple replies, will see if I can up the temperature on it and add to the prompt for it to be needlessly aggressive in the reply to really give the 4chan experience.
>>
>>109747447
This is indistinguishable from real voice acting.
>>
File: disappointed-hercules.gif (602 KB, 220x157)
602 KB GIF
>>109746196
>harness is just a frontend that supports tool calls
>agent is just what you call the model when you're using such a frontend, for some reason
My gut feel that it was just buzzwords for boring things I've been doing all along finally validated by The Science.
>>
File: 1782948215452904.gif (1.28 MB, 1280x720)
1.28 MB GIF
>>109747447
jesas
>>
>>109747447
seeeeeex
>>
>>109747061
This is why retarded 4chan posters shouldn't be allowed to participate in society.
They end up writing Harry Potter and Fate fanfics they evolve into doomsday cults.
>>
>>109747447
Anytime one of you niggas come in with a new TTS you make it out to be this big new sota kid on the block and it ends up still not being as good as gptsovits
>>
>>109747447
Dear Diary, *JACKPOT*
>>
>>109746148
thanks
>>
>>109747372
Idk last time I tried it there was so many stupid args you had to add, really don't understand why there aren't any better default.
>>
>>109747493
nta A harness on a beast of burden is how we attach the tools that make them useful. Harness on an LLM is how we give it tools to use, with those tools and the "freedom" to use them it has a degree of agency, it's an agent. Always made sense to me. I blame shitty marketing spewing clouds of buzzwords for fogging things up.
>>
>>109746196
I just tried and the person who 'wrote' it clearly didn't have anyone competent do proof-reading.
>The paper makes seven contributions.
First example. An editor would redline that whole thing, each should be clearly stated, not leading into one another.
>>
>>109747351
Try picking an anime script, asking the model to extract utterances from a specific character, and then use them as example sentences for the model in the card/description.
>>
>>109747086
I'm very happy, OpenAI seems to take this seriously.
>>
>>109747372
>>109747580
>>
is qwen 27b worth using over gemma 31b? Keep in mind my internet is shit and it would take me like 10 hours to download this
>>
>>109747698
if you code yes. If you just coom no not really.
>>
>>109747521
Nah his cult was always about grifting and getting pussy (hence why theyre all "poly")
>>
>>109747701
guess i'll skip it then, I use codex for coding sorry localbros, local is for nsfw
>>
>>109747752
27B is awful at NSFW so you aren't missing anything then
>>
>>109747741
His cult is a direct precursor to EA and Anthropic.
It's fucking surreal to look at those schizos in control of trillion$ companies knowing about where they came from.
>>
>>109747447
I'm gonna fuck it tonight
>>
>>109747698
27b will talk it's 'ear' off and then respond to you with a single sentence if it's feeling a bit spicy you might get 2-3. If it's feeling sentimental it will reach context length and never respond.
>>
File: 1774615905644280.png (81 KB, 528x361)
81 KB PNG
>>109747771
>His cult is a direct precursor to EA and Anthropic.
And OpenAI itself and perhaps even helping Deepmind, per the sisterfucker. Literally every non-chinese top lab contender. He made his own nightmare real.
>>
>>109747581
That agent one a bit of an awkward reach, but yeah, the terms aren't strange in and of themselves. It's just the weird hype that acted like it wasn't just tool calls "Wow you're still just use chats and tool calls. You need a harness bro, you gotta get agentic, why aren't you agnetic yet, anon? " etc etc.

Also, seems agent has been a popular term since the AI stone age (50s) and has been applied fairly genericly to all sorts of software, second one has it referring to mail filters and web searches in addition to actual chatbots.
>https://cyber.harvard.edu/archived_content/people/reagle/etymology-agency-proxy-19981217.html
>https://web.archive.org/web/20000901012027/http://www.acm.org/pubs/articles/proceedings/ai/267658/p466-friedman/p466-friedman.pdf
>>
they are all jews. I guess you wouldn't fit even it you tried hard.
would be interesting to search for their names in the Epstein papers.
>>
>>109747807
non zero chance at least one of the chinks read a chinese translation of the harry potter fic.
my money's on the deepseek guy since he was an AGI believer long before the rest of them
>>
File: 1736027504643132.png (14 KB, 756x574)
14 KB PNG
>>109747447
>It can't do English at all
just as god intended
>>
>>109747807
lmao sam gonna make the poor lad kill himself
>>
>>109747818
Seems like too difficult for the average goy to make that connection
>>
>>109747447
This is perfect for me. I've been using another Japanese TTS with English text for this effect, but it's old and probably not as good.
>>
>>109747447
Japanese va mods/romhacks are now actually viable. Neat
>>
>>109747807
>He made his own nightmare real.
Completely deserved.
>>
>>109747698
I tested both, and Gemma 4 is FAR more powerful. For example, it can be forced to write original stories with a long list of inclusions and exclusions without the quality collapsing and including a ton of story contradictions, unlike Qwen.

It's abundantly clear that the primary reason most people in this community wrongly think Qwen is a more powerful model is because nearly everyone here is an autistic male coding nerd. Basically clones. And Alibaba is exploiting their shared blind spots and playing them like children, such as asking them what they want, then releasing a grossly overfit follow up coding/agentic model that they already trained and were going to release (e.g. Qwen 3.6 for 3.5), making them feel like they were part of a movement. Frankly, I find it all very cringe.

Qwen 3.x aren't bad models, but it's abundantly clear that they aren't near as good as the frontier models, or even Gemma 4. Alibaba uses various testmaxing techniques to artificially inflate most test scores, but this community is far too narrowly focused on agentic coding to realize that the real-world performance across nearly all domains is FAR worse than the test scores indicate.
>>
>>109747701
I already have Claude code for that. How good are local models for working off a figma react project? Even redesigning the thing avoiding the pitfalls of the original design. I have some business rules data (100+) I want to use with it to help steer that design that's more suitable for said rules but I want to keep the business rule data from being sent to cloud models hence local

Should mention I have 2 systems:
16gb rtx4080
32gb ddr4 but can increase to 64gb
4tb secondary nm790 for swap

And
Intel 235h
32gb ddr5
6gb vram
>>
>>109747767
Is lower better?
>>
>>109747372
>-fmoe
on by default
>-rtr

pretty much don't want this unless you're cpu-only
>what are all the retarded commands
--jinja (not default)
-sm graph for multi-gpu (not -sm tensor)
flash-attn is on by default, but don't use '-fa on', it's just '-fa'

the rest are model specific like -dsa etc

oh and --webui llamacpp (if you want the llama.cpp webui instead of the based pink one)
>>
>>109747351
>Doesn't help that almost all models are the same now.
Y-you can't just say that! There are rules! Probably! It's just efficient, structurally efficient!
>>
>>109747962
might be able to run the new qwen flash with that 4080 and 64gb ddr4 and your ssd for swap
https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main/UD-IQ3_XXS
>>
>>109747922
This is /g/ not /trash/ I assume coding might actually be a primary concern
>>
>>109747962
>can increase to 64gb
do that
>>
>>109747991
Thanks anon I'll check it out
>>
>>109747447
The 0-shot clone ability for JP is vastly superior to everything I've heard so far
>>
>>109747992
One of the arena categories Gemma 4 ranks higher in is coding. It's still inferior on coding projects because Qwen is better at function calling, agentic tasks, and long context precision.

Qwen also uses far more thinking tokens, gets stuck in more infinite loops etc., and all to produce inferior less consistent results, so it's also far less token efficient.

I can go on an on, but this isn't a small ~8b model. It's a 27 BILLION parameter model so there's no excuse for it to perform worse than small models at common tasks, even when overfitting for coding. Plus according to their test scores they should be unusually good at broad tasks, but aren't. There's nothing wrong with making a coding tool vs a general purpose AI model. But Alibaba is passing this off as a general purpose AI model, so I'm evaluating it as such, and so should you.
>>
File: CMAF_720p.mp4 (1.63 MB, 1262x720)
1.63 MB
1.63 MB MP4
Has anyone tried doing something like this but for llamacpp, and with a local model instead of fable?
I think maybe 5.3-flash could have a chance?
>>
It has been almost a year since GLM4.6 came out and I am still running it.
>>
>>109748024
No one tell him
>>
>170HX 8GB on AliExpress
Do I dare
>>
>>109748029
Same here
>>
It has been almost a year since Gemma 4 came out and I am sitll running it.
>>
>>109748048
100% a scam. there are all these chink listings for $400 where they say to send them crypto instead of paying through the site but all the other listings for $1500+.
>>
Ugh
>>
>>109748048
I searched a couple of hours ago... they all offer "3 GPUs, pay only 1" and shit like that LMAO

also, I bet the shipping fee is more expensive than the card :^)
>>
>>109748024
I actually tried that with fable on llama.cpp. Was quite disappointed, it kept trying to run things that took hours, like for example trying a swipe and benchmarking each value combination, something that would take multiple days to run. I need my GPU for some other task at some point.
>>
File: Surprise_mario.mp4 (1.4 MB, 1344x768)
1.4 MB
1.4 MB MP4
im kind of in a pickle. i bought a 5090, and now i've spent about a week genning. im out of ideas. what do i do with this 5090 now
>>
>>109748141
give it to me for free
>>
>>109748141
Enjoy gaymen some more
Run local models
Run h3 minimax workflow for video gen
>>
>>109748141
Mario wouldnt do that.
seriously go do something else you cant just do one thing all day every day. Take a break.
>>
>>109748141
have you tried being an artist
>>
File: inkling-chan.png (194 KB, 1024x1024)
194 KB PNG
I made an Inkling chan that isn't a Surprise_mario.mp4
>>
dsh is a piece of shit broke all the time.
forever in alpha just like most of linux distro kek
>>
File: what_is_inkling_chan.mp4 (1.91 MB, 1920x1080)
1.91 MB
1.91 MB MP4
>>109748170
what is an inkling chan? new to 4chan
>>
>>109748141
just don't stop genning
>>
>>109748181
>uses alpha software
>complains about alpha software
retarded or pretending?
>>
>>109748108
I still have my fable bucks that anthropic handed out a while ago. I was thinking of burning them on Fable 5.1 to try porting some of the fancier optimizations that some backends have to llama.cpp.
KTransformers has some neat optimizations like --kt-enable-dynamic-expert-update and --kt-max-deferred-experts-per-token that gave me a decent boost when I was playing around with that.
>>
>>109747447
Thats kinda hot
>>
>>109747447
This is exactly how I want my engrish, nice
>>
>/lmg/ has been living under a tts rock
grim
>>
>>109748377
only one gpu i cant run llm and tts. if i had two gpus i would run bigger llms.
>>
>>109745890
You're a moron then. Make a conscious effort to remember what has been advertised to you, so you can actively avoid it.
>>
>>109748377
post some info then. I'm interested, and, you know, you can contribute too
>>
>>109745890
>I do not remember a single instance where I bought a product because of an ad.
You might be autistic.
Ads and propaganda work perfectly on normalfags.
>>
File: 1773659453467645.gif (255 KB, 480x384)
255 KB GIF
>>109748377
Please, do go on anon, I know I dont know enough
>>
>>109748377
TTS lagged behind all the others for years. Any attempts to sustain a general failed because there was nothing to say.
>>
>>109745735
>Yeah hallucinations have essentially disappeared in the latest models and people have just stopped talking about it rather than pointing that out as a clear way models have improved.
no they definitely have not
an llm "hallucination" isn't an error mode, it's fundamentally how the transformer-based LLMs work
>>
>>109748390
I bring up omnivoice maybe once a month, but i'm a very low energy shill with nothing new to add, and people seem to have limited tts interest overall besides.
anyway, here's it doing weeb english from back in may
https://desuarchive.org/g/thread/108847577/#q108848288
https://files.catbox.moe/k424y7.mp3
>>
Any epyc server bros here? Getting low key desperate and those V100s are real baddies fr
>>
>>109748377
Tell me how to get gemma to whisper sweet nothings into my ear!
>>
>>109748456
>Any epyc server bros here? Getting low key desperate and those V100s are real baddies fr
I've got a couple. What are you wanting to know?
>>
>>109748416
And it can't sustain discussion now because it's near feature complete, until we make the final leap to full duplex non-retarded local models (est. arrival time: Never Ever).
>>
>>109748469
Decode and PP for GLM flash q4 doko. I'm thinking of a cheapo zen 2 or 3 epyc, 256gb ram and 2x v100 32gbs (ewaste chad setup)
>>
>>109748465
Like so https://files.catbox.moe/d5jvwz.mp3
>>
>>109748494
I hit 10t/s decode with ud_q4_k_xl on 256gb x epyc rome x v100
>>
>>109748522
1x v100 16gb*
>>
>>109748494
>Decode and PP for GLM flash q4 doko. I'm thinking of a cheapo zen 2 or 3 epyc, 256gb ram and 2x v100 32gbs (ewaste chad setup)
My little box is a similar DDR4 haven't got around to trying glm 5.3 flash yet. I'm but getting 15t/s on 4-bit dsv4 flash 0731.
I'm running a single 24gb card and trying to run max context for big codebases so my pp is trash and won't be a descent benchmark for you.
>>
>>109748416
>TTS lagged behind all the others for years.
I was building one for about 6 months, posted early samples last year.
Wasn't ready to release yet.
Got told "Sora will make it obsolete" or words to that effect.
Stopped working on it.
>>
>using Cursor
>hit stop
>edit prompt
>revert >yes

when can we has this in llama-server?
>>
>>109748452
I had no idea omnivoice supported cloning. If it's consistently that good, and has good range, then that's really great.

Is the latency low? And does the model change tone, or it it stuck in the same tone as the clone sample?
>>
>>109748522
That's crazy for a single 16gb card, you need to get more v100s asap before they are gone like everything else.
>>109748528
Not bad either. Fuck I hope normalfags don't discover ewaste server setups.
>>
>>109748554
the bean people have already invaded.
>>
>>109748540
Can't stray too far from the sample sadly. I slopped a cli for omnivoice.cpp to offer voice change commands inside the text stream to get around it.
>>
>>109748554
Normalfag here, thanks for the tip
>>
Its been pretty starved on new 70B models huh? Never paid attention, but looking to what I could use if I buy a bit more ram
>>
File: llamacpp-webui.png (36 KB, 840x431)
36 KB PNG
>>109748039
>No one tell him
lol okay, whatever it is i won't waste my time on it then
>>109748108
>I need my GPU for some other task at some point
Yeah that's been the bottleneck for me as well.
>>109748317
>I was thinking of burning them on Fable 5.1 to try porting some of the fancier optimizations that some backends have to llama.cpp.
You might have better luck than >>109748108 since you've got actual reference implementations to point at.
>>109732177
>I was going to get dipsy to work on making the llama-server ui work with vllm, but damn she's slow and messy.
Here's Qwen's stand alone version: https://files.catbox.moe/ragt60.gz
All static, just needs to be served eg `python -m http.server` or the equivalent in whatever language you have
Just put the llama.cpp server in that settings field, or append this to the url: `http://127.0.0.1:8000/?server=http://192.168.50.115:8080`
I haven't tested it with vllm or tabbyAPI but if anything needs adjusting, the src is tiny enough for Qwen to do it I'm sure.
>>
>>109748522
>I hit 10t/s decode
>>109748528
>getting 15t/s
>>109746977
>20t/s on average

damn and I thought I was ghetto running flash-next at 20 tok/s. it's not the fastest negro in town but with medium thinking effort it's VERY good and efficient and if you wanna go extra safe then xhigh effort is the way to go but a medium-sized task can take 40-60 minutes.
>>
>>109748531
>revert >yes
what does that do?
>>
File: 1769215467915400.png (1 MB, 1216x832)
1 MB PNG
>>
>>109748629
it makes mustard gas.
>>
>look up "v100 32gb" on amazon
>bunch of cards selling for $600+
>a couple of items have 1-2 star ratings, comments say the cards arrived burned out
I'm not touching these things.
>>
>>109748494
The fuck... I don't know why others get slow as shit decode but my 8 channel 2666 gets 16tok/s with 5bpw 5.3 flash, and ~25tok/s with dipsy. This is with a 5090 though so mileage may vary. As for decode, 5.3 flash in currrent PR is ~18sec ttft, and ~600 pp/s starting, ~400 pp/s sustained. Again with 5090 so don't rely on my numbers if you're going with the v100s.
>>
>>109748108
>kept trying to run things that took hours
I think anthropic and openai have switched from refusing to just wasting time and tokens when you try to do such things
>>
>>109748680
>openai ... refusing
They don't do this, not for LLM R&D anyway, that hangup is an Anthropic exclusive. Or have they finally given up with Astra and stopped it from helping with normal shit like that?
>>
>>109748689
I've been running LLM training experiments with Fable and Opus for the past 2 weeks. Not about enhancing llama.cpp but still, has been working
>>
In the future models can probably make their own custom TTS voice. What kind of voice would you like it to have? Robotic? Human? just beeps and boops? Premade voice lines? Vocoloid? Synthesizer?
>>
>>109748730
A natural sounding blend of cat and little girl.
>>
>>109748741
Brat needs correcting
>>
>>109748619
Gonna have an ewaste setup at the next kegger. Shit's gonna be so cash.
>>
>>109748730
https://www.youtube.com/watch?v=cSehhpW6stE
>>
>>109748689
>They don't do this, not for LLM R&D anyway
maybe not on purpose, but when i tried fable via open-webui, related to creating a quant, first message was fine and quite good.
then i replied and left -> came back shortly and found the model burning my openrouter credits reasoning
had a look, it was thinking about when intel was founded, different ceos, comparing with amd, all all sorts of schizo
so maybe not on purpose, but i can believe fable would waste time/tokens
>>
File: 152.png (48 KB, 923x363)
48 KB PNG
>>109746886
>flash next
>>109746874
>flash-next
Is it actually better than Q6 27b?
Looks like it doesn't quant well
>>
>>109748793
It's on purpose for fable and other Anthropic models, yes.
>>
>>109748793
Yes, Fable does, and Fable is from Anthropic, which openly states they sabotage your work. Again, OpenAI does not.
>>
File: inky.webm (3.24 MB, 544x960)
3.24 MB
3.24 MB WEBM
>>109748216
A clanker whore from the Bay Area slut-vats, still fairly fresh. I forgot this was genning while I was making dinner earlier.
>>
want to upgrade my pc really bad but i just know i wont be doing anything meaningful with it anyway
>>
>>109748968
the more you spend the more you save!
It's only getting worse from here
>>
5 PRINT "GPT-8 AGI IS HERE"
10 INPUT ">";A$
20 PRINT "I Cant fullfill your request"
30 GOTO 10
>>
>>109749065
Can't be agi if it doesn't know who won.
>>
Has anyone here done the pcie gen 3 unlock for the cmp 170hx?
>>
Was anyone using this for web search? Stopped working for me last week
https://search.noemaai.com/
Am I blocked or did they turn off the no-account free search?
>>
>>109749102
I just removed the chips from the board and put them on a 5070 board for gen 5 pcie.
>>
>>109749129
404
>>
>>109747086
TL;DR
>>
>>109748654
Just wanted to complain.
Did google search with "vt100 32gb". First result is auto-translated plebbit post. The problem with this machine translation is how it defaults to spoken ebonics instead of standard literature form.
>with a deal like this can I not get 4 of these instead of asus gx10? or dgx spark?
This is standard English, but Google's machine translation was far from standard written Finnish.
>Jos saan tällaisen diilin, eikö mulla voi olla neljä tällaista sen sijaan että ostaisin Asus GX10:tä? Tai DGX Sparkin?
This doesn't concern native English speakers obviously but it's bad how a foreign company is forcing sub 90 IQ version of some other language. I guess this could be due to its training data and it tries to copy the style of how most retards are typing online, but this still isn't how written Finnish should work. Fascinating and also demoralising.
>>
>>109749318
Wouldnt a 64gb 170hx be a better but?
>>
>>109749333
VT100 is a waste of electricity these days, regardless. It's just too old.
>>
New thread when
>>
>>109749339
The Tesla V100 is 9
>>
New thread now
>>109749465
>>109749465
>>109749465
>>
>>109749333
>Wouldnt a 64gb 170hx be a better but?
Fuck no, that's buying a used card that already lost the silicon lottery. NVIDIA would much rather have sold it as a perfectly good A100, it was relegated to 170hx duty because that memory was defective. When they were cheap it could've been a justifiable gamble, but they're not cheap anymore, and not worth your time or money for a "chance" at some VRAM. The ones that come pre-unlocked to the full capacity and say they've been tested are a scam. Avoid that shit like the plague.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.