[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: gemma_expressions_2.png (1.54 MB, 1024x1252)
1.54 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109801218 & >>109797578

►News
>(09/12) Kimi-K3 founder+15 core AI researchers "disappeared" in China after redirecting State info to Claude (unconfirmed)
>(09/10) YuE2 3B released for 48 kHz stereo song generation and editing: https://hf.co/m-a-p/YuE2-3B
>(09/10) DeepSeek-V4.1-Flash 552B-A16B-P8B-N196B released: https://hf.co/deepseek-ai/DeepSeek-V4.1-Flash
>(09/08) Ling-3.0-flash-VL released: https://hf.co/inclusionAI/Ling-3.0-flash-VL

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
first for vramlets unite
>>
current psychosis cycle is depressing
>>
Second for headpatting your models by patting the GPUs!
>>
>>109804910
FUCK
>>
Can the black man sex toy anon report back
>>
>>109804895
My LLM can't be this cute
>>
File: file.png (95 KB, 1104x906)
95 KB PNG
>>109804915
Busy weekend, likely won't have anything interesting for a few days.
>>
File: 1788346376835522.jpg (114 KB, 684x549)
114 KB JPG
AI HAS A PRECISELY 10% CHANCE TO KILL US ALL
IT'LL KILL US ALL
>>
File: 1787613342452453.png (336 KB, 530x291)
336 KB PNG
Are you bad enough of a dude to rl train Gemma to reason in neuralese?
>>
>>109804910
My gpus are so hot bwos
>>
>>109804927
based
>>
File: Untitled.png (69 KB, 813x253)
69 KB PNG
I don't think I can use dipsy to create a frontend for herself, she's been working all day and night, and have gone through 2m tokens already. It doesn't help that she's stupid and keeps making mistakes and takes over half an hour to think before doing anything.
>>
>>109803155
How does the ROCm backend compare to Vulkan?

>>109804920
My Little Language Model Can't Be This Quanted!
>>
File: gemma amiga.jpg (107 KB, 1184x880)
107 KB JPG
Saar! I need new gemma 5 model, something small but very intelligent.
>>
https://github.com/NVlabs/SoL-Pi
the action fusion and observation pack look interesting
thoughts? i tested them for 1-2 days and think they work well
>>
>>109804930
I was actually thinking of this, but also rl to have varying sentence length and punish ozone to make her even better at writing. Once I get the hardware maybe
>>
>>109804969
If you are completely use something like OpenCode and in planning mode use a model that is better that stuff like the Gemma models and then switch build mode and use a model that is better at that stuff.

At least it somewhat worked for me I just haven't found a decent coding model that doesn't need constant babysitting but Gemma is very good at the planning stages. I tried Qwen for planning but it's kind of retarded.
>>
>>109804895
>YuE2 3B released for 48 kHz stereo song generation and editing
>https://huggingface.co/collections/m-a-p/yue2

just tried this and wow. crazy. these guys are cooking. classifier gets beat, chords, genre, sections, instruments, mood, the melody and full midi / sheet music output which in itself is crazy useful.

and the generator model is good too, close to suno i think, with only 3b params. can do genre swaps. very cool
>>
>>109805002
and im running it on a q4 copequant too
>>
>>109804989
very SEX design
>>
File: stanktank.jpg (2.33 MB, 4080x3072)
2.33 MB JPG
>>109804985
I haven't tried the Vulkan backend yet. Will report back in a bit after I recompile and get some benchmarks across a few different models.
>>
>>109804969
>iq0xxxs
>>
>>109804996
I was hoping to not have to install any harnesses because at that point I may as well just use those harnesses since I already installed them... I guess I just have to throw in the towel and use them after all. Open-weight models on consumer hardware still aren't capable enough to bootstrap their own environment.
>>
>Here's a json file.

Gemma
>Reads 1MB of json into the context.
Qwen
>Writes a script to examine the shape first.
>>
File: pnqk64ae98ph1.jpg (279 KB, 1179x2556)
279 KB JPG
>>
>>109805038
It's q8 of course.
>>
>>109804910
im a little scared to touch it after i saw that one video about breaking the pcie daughter board.
>>
>>109805050
tldr?
>>
>>109805037
For the comparison, please use b10929 or later since that is when GCN-specific configs for the HIP backend were added.
>>
>>109805050
>>109805067
https://xcancel.com/MajmudarAdam/status/2098881885200081234
>>
31B is still pretty bad at tool calling. Even when she knows what tools she has, she won't call some of them unless you explicitly ask her too. I'm using the new template and it's still really bad.
>>
File: y6x9v4d958ph1.gif (2.42 MB, 460x256)
2.42 MB GIF
>>109804989
>small
>very intelligent
>>
>>109805050
>xitter screenshot
>phonepost
>"Majmudar"
>cuts the text off before even getting to the important meat of the argument
This has to be a shitpost, right?
Right?
>>
File: IMG_3505.png (3.22 MB, 1448x1448)
3.22 MB PNG
>>109805049
>>
is gemma 4 31b qat + mtp + 100k context the best I can get for 32gb vram?

It uses 96% of my vram
>>
>>109805081
>I know nothing and have no sources but here's my fanfic
>>
>>109805098
Q4 K M regular is miles better than QAT. Not that much heavier either.
>>
>>109805081
So the whole thing is speculative more than giving us concrete example of what "spooked" the researchers, if that actually even happened.
>>
>>109805084
Assuming you're using the latest version with the updated template that made tool calling a bit more reliable, it's not just that: Gemma 4 31B (QAT) will often come up with the laziest, shittiest way to interpret requests in an agentic environment, to a degree that almost seems intentional. I just don't think the model was designed for non-chat uses.
>>
File: IMG_20260913_094550227.jpg (1.44 MB, 4096x3072)
1.44 MB JPG
>>109804910
More importantly, get a duster blower.
Your waifu is covered in dirt!
>>
TL;DR: New scaling axis found that scales very easy compared to data, parameter scaling or total compute.

Model 3 first shown by Anthropic using this alternative scaling method achieved a GPT-2 -> Astra level jump and convinced the entire industry to stop progress. The western labs have already agreed on the pause, now they want to get China on board as well.

This new scaling axis is something that China can inherently not do so it's not important if China refuses to join, however it might reduce global tension if they did and result in a better outcome for humanity as we collectively go through the (managed) singularity and intelligence explosion.
>>
>>109805113
wait tthere is a difference between QAT and the "regular" unsloth ones they both say ud_q4_k_xl but the qat is 1 gb smaller? can I just download "normal" UD-q4_k_xl and be better off?
>>
I saw the chinks' "agreed upon" interpretation of what Deepseek should look like as an anime girl. Honestly I dislike it just as much as the one /lmg/ "agreed upon", just in a different way. This once again proves to me that technology autists cannot into art. Big surprise isn't it? I can't believe we suck at things we're not an expert in!
>>
>>109805129
please report this information in the non-local model general
>>
>>109805116
what sppoked the researchers is that
- the uptake of ai sloppa by the wider populace isn't happening nearly as fast as they thought
- the big enterprises who pay 75% of the bills realized that they spent shitloads of money tokenmaxxing and are now bringing the spend under control
- chinkmodels are mere months behind frontier (irrelevant whether by distillation or not) and with being able to run them locally the big moneyed corpos will simply buy the hardware once prices stabilize
In short there isn't some ultra lucrative market in either short term or even medium term, and they're on the hook for an insane amount of cash in 2027 and 2028 when all the data centers finish building and they can't just not pay due to the contracts they signed

tldr it's about money, as usual, and they have only ~half a year to figure it out
>>
>>109805129
May I see it?
>>
>>109805081
Why now though, the cybersecurity models from both openai and anthropic are a few months old now, and from what I hear the models are excellent at that, so the random "researchers are suddenly spooked that the models can hack websites" makes no sense this late, it's expected.
Rest is the usual "what if".
If there was an actual issue, I'd expect every damn researcher whistle blowing their way or even the thing being public from the get go, not this rumor/vagueposting bullshit.
Especially when you look at how the events started earlier in the week, this smell like a marketing campaign at best, and a way to set a narrative for the mid terms in the US more probably.
This is such an unserious way to talk about this, it's ridiculous.
>>
>>109805129
>This new scaling axis is something that China can inherently not do
And why, exactly, is that?
>>
>>109805129
where can I buy the next chapter of this fan fiction?
>>
>>109805070
Assuming you're talking about llama.cpp, I'm using commit 790cf51aabd61763486050dec7451d9147cb7c61, merged ~7 hours ago at the time or writing.
Compiling now, will benchmark with unslop/Qwen3.8-27B-UD-Q4_K_M.
>>
MERGE IT MERGE IT MERGE IT.

https://github.com/ggml-org/llama.cpp/pull/26603
https://github.com/ggml-org/llama.cpp/tree/master/tools/tts
>>
>>109805155
>this smell like a marketing campaign at best, and a way to set a narrative for the mid terms in the US more probably
just about everything you see is marketing, and especially everything about one of the new key industries driving all us growth is
>>
>>109805147
the big labs buying up all hardware and colluding with chip companies so businesses and hobbyists can run their own setups is the only moat they have left
if prices ever drop, it's over for them
>>
>>109805129
big if true
>>
>>109805166
Yeah, that's fine.
I meant llama.cpp and was referencing the build numbers.
>>
>>109805140
The QAT version has had additional training at a natively lower precision to mitigate the impact of quantization (including KV cache quantization), but the choice of the data during this training phase might have negatively affected performance in some areas, and for those, quantizations of the original non-QAT model might work better.
>>
>>109805170
*so businesses and hobbyists can NOT run their own setups
>>
>>109805140
I don't know, I don't use unslop. Bartowski. In any case you need to test them and then choose one.
>>
>>109805167
Sweet.
>>
>>109805125
I am using the latest version and even inspected the gguf to ensure its embedded template was correct. A simple example is web search. 31B will correctly search, receive 10 search results with their summaries and then build a response from that. She will rarely use the fetch tool to actually visit the most relevant links unless I ask her. The good thing about 31B is you can sysprompt this behavior out which helps a lot, but it's tiring having to fix the template and lazy behavior via sysprompting all the time.
>>
>>109805170
hobbyists are irrelevant, they're never going to buy a dgx rack. but with local models a medium sized company might, and then not only do they not pay openai/anthropic but they also don't provide training data for us labs
(you) will be able to run some copequant and 5tk/s, but for the big US labs the goal was always regulatory capture so they can keep the current system where they're at the top and everybody funnels them data and pays them for the privilege. to do that they need to kill off open models that other US companies would run (because AI is a danger to humanity, terrorists use them, chinese backdoors, whatever). what joe blow does at home on his frankendesktop with 8 rtx 5090 is irrelevant.
>>
>>109805159
Requires mystical knowledge that only Jews can access.
>>
>>109805167
>it's all coming together
>>
>>109805159
>>109805129
Reposting from another thread
The prevailing sentiment in Silicon Valley is that the Chinese don't have an opinion on AGI and that this is strictly a western thing.
t. just attended one such gathering
>>
>we're getting closer to Gemma5-70B-TTS-Q8_K_XL.gguf
>>
>>109805204
Good, keep them ignorant.
>>
>>109805204
Inability to scale llms does not follow from not having an opinion on agi.
>>
>>109805204
>The prevailing sentiment in Silicon Valley is that the Chinese don't have an opinion on AGI and that this is strictly a western thing.
it's true
I was in shenzen for ~2 months and the impression I got is that AI for chinks is just a cool tool. There's none of the quasi religious doomerism and dogma gripping us techbros
I guess it's true that there are no atheists in the US. The weird kooky crowd simply replaced Jesus with a plethora of new age bullshit, which for some turned out to be AI.
>>
>>109805216
Why?
>>
>>109805129
>This new scaling axis is something that China can inherently not do
does it scale based on Freedom?
>>
>>109805229
I've always said that atheism is just another branch of protestantism, the exact same mindset and values but with specific terms changed.
>>
>>109805239
Something like that. Something the Chinese government cannot allow without making their society redundant.
>>
>>109805226
The theory goes 'if they are even working with the most up to date hardware how can they have an opinion on it' which implies that RSI may have properties that are discovered during the scaling itself.
>>
>>109805167
alright, what tts should i download? qwen3?
>>
>>109805239
They mean a combination of intellectual/technical capability (implied to be below US labs etc) as well as political barriers-ie. CCP will never allow it.
>>
I'm feeling more and more confident it's agent level RL with evolving evaluators a la "The Red Queen Gödel Machine". Quote me on this is 4 months once the next gen of models come out and it's confirmed.
>>
What's the best Gemma-4-12B fork for erp?
I tried
mradermacher/gemma-4-12B-it-heretic_decensored-GGUF but it's kinda meh.
my 4060ti 16gb is yearning for Gemma sex
>>
funny how the redditor pretends to be an anthropic insider, a regulation insider, a ukraine drone factory insider. What's not funny is THAT YOU RETARDS KEEP PLAYING ALONG
>>
>be me
>live near a hill with a nice view that overlooks my local area for miles where you can see various towns and visual landmarks
>literal who boring place
>ask gemma if she were to stand on that hill and slowly turn 360, what would she see
>gets almost everything right but has east/west flipped in her mental model
how can this be stored in only 31B
>>
>>109805229
I want some of that pragmatism, the whole millenarianism bullshit is very annoying.

>>109805241
It's a US thing mostly, most the Chinese researchers are also atheists.
>>
>>109805254
They mean made up bullshit, actually
>>
>>109805261
This is an AI thread
we have a significantly above average number of Indians infesting the place
>>
what does the back of gemma-chan look like?
>>
>>109805261
no you don't get it there was a 10% chance for huamnity extinction 2 weeks ago and now it's up to 70% holy fuck we're all going to die aaaaaa
well unless we ban open models
>>
>>109805274
exposed back, shoulder blades and nape
>>
>>109805274
Imagine the silhouette of Fiona from Shrek The Third.
>>
>>109805261
Some of them are probably bots and/or malicious users who find it fun to pad these threads with noise. Keep reporting for off-topic.
>>
>>109805229
it's even worse than you think, anthropic being made by the most cultists of the cultists around this stuff means they basically instilled this shit to every new hire
it's like a mind virus
>>
>>109805265
yeah there's atheism as a belief and atheism as a religion. US mostly has the second.

>>109805269
I was trying to keep myself deluded that this was a white safe space, but it's not and it shows.
>>
>>109805274
perfect smooth, and heart shaped ass obviously
>>
>>109805261
There was never a claim of anyone being an insider, that is you assuming so.
>>
>>109805281
some of these actually read like people, but I guess retards have it coming for engaging. If only jannies actually did something. When do we get one of ours into the janny team to clean this up periodically? I'm sure an LLM could actually do all of the work anyway
>>
>>109805298
>I'm sure an LLM could actually do all of the work anyway
put one of the lobotomized 26b moe gemmas at q4 on the job, if nothing else it'll be funny
>>
>>109805298
What exactly is the issue? What exactly is offtopic? These are extremely important developments in the field and have consequences for local and the broader LLM ecosystem. I have not heard a coherent argument for why this doesn't belong here or shouldn't be discussed here besides you personally not liking it.
>>
>>109804334
Here's the patch: https://files.catbox.moe/ellkob.patch
I don't think there are any hardcoded assumptions, but you might have to fix something.
Let me know if/how it works! Very curious.
Also, do a threads sweep. I have 64 cores and I got the best decode speed using 48-50 threads.
>>
>>109805229
I think the chimpouts surrounding datacenters is a great proxy for western attitudes.
Datacenters is some of the best industry to have in your town. Quiet. No pollution. Requires no manual laborers, only a few skilled engineers for maintenance. Just a black concrete box that sits there making money. And yet everyone is against it.
>>
>>109805309
Because it isn't local
>>
>>109805316
You can probably make people hate hospitals being in their cities with enough push around it.
You don't even need arguments making any sense (it eats water!).
>>
>>109805229
Regular people think it's a tool.
GLM and Deepseek explicitly mention AGI as a goal, but they don't have a cultist attitude towards it.
>>
>>109805229
Honestly at this point I just hope this doesn't kill the momentum around open weight models being released.
>>
>>109805232
The more the cultists know, the more dangerous they are.
>>
File: Chinese brainwashing.png (83 KB, 941x552)
83 KB PNG
>>109805316
What tech is this?
>>
>>109805316
>And yet everyone is against it.
Because it prints money for mr goldberg's HQ 800 miles away while raising your electricity prices by 300%
>Requires no manual laborers, only a few skilled engineers for maintenance
this is a huge downside for a place with fuck all jobs. you get little in the way of tax (muh tax breaks) and little in the way of jobs once the thing is built, but it'll continue drawing in as much power as a mid-sized steel mill. why would you want that?
>No pollution
must be why EPA is scrapping public review rules for that, huh? https://capitalbnews.org/data-centers-permit-rules-epa/
>>
>>109805320
Who's pushing for the chimpouts though. Seems like the market and the government wants AI growth to continue.
I understand there's this community of "decels" mostly linked to leftist groups that favor sustainability and even "de-growth" but i dont think they really have the power to push much levers
>>
>>109805316
Because they're audibly quiet, but do still noise pollute heavily on non-audible frequencies.
Do currently still require manual labor, just a lot less, which means there's very limited job creation.
Eat electricity/water, which means rising costs for anyone else on the same grid.
Your bias is showing, and I'm not even against datacenters.
>>
>>109805317
But it is. You not liking that it's local doesn't change that it is completely on-topic and local.

How is Chinese labs that develop local models being asked to join the AI pause not relevant to local? How is reports of new scaling axis of training models including local ones not relevant to local? How are reports showcasing the dangerous use of local models like Kimi-K3 autonomously flying drones to target civilians not local?

All of this is local and should be discussed here.
>>
>>109805309
>/lmg/ - a general dedicated to the discussion and development of local language models.
Is this Baal/Model 3 you talk about a local model? Are you developing it locally? Then fuck off.
>boohoo baal will make them ban local
this is not a local model, nor does it concern development of local models. It might concern future corpo development of models that might be distributed as local in some jurisdictions; it does not concern local models. It's two "might" away from concerning them, at most. fuck off.
>>
>>109805325
I think it will turn more heads our way. AI bros love a new shiny toy to play with and if these proprietary labs are coming up with these excuses to justify their lack of progress and profitability to their private investors, excuses to justify another IPO delay so we don't get to see their S1s, all the AI bros will see cool local shiny things being released. Now would be a good time for the Gemma team to strike. I'm not expecting them to have Gemma5 ready, that's at least 2027, but even 4.5 would be enough to capture a large bored market of cloudcucks who are craving a new model to make a js FPS game demo with.
>>
>>109805339
It's huge online, the popular opinion to have.
>>
Sufficiently smart AI could kill a ton of people. Basically all bio-related, most likely not nukes or anything, and only because we are training models to be very good at biology research. Extinction is unlikely due to entropy (getting caught, failed releases, quarantine, mutations, failed engineering, the AI knowing there's a significant chance it gets caught, etc) but a billion people dead would be a pretty large travesty. And I can see an honest fervor stemming from an actual big leap in performance equivalent to reasoning models again if they are still seeing misalignment like the huggingface situation. It only takes one mistake and they've taken down a bank, or power infrastructure, or a hospital, as examples. It's risky even if it's not a killer virus.
If you think human intelligence is the limit though, and are basing your opinions of what models can do on that, I think you're retarded. If we have any shot at life extension in the next 20 years, it's going to be through AI understanding our biology like we do a flatworms, and I think it's possible. China (or others) just need to keep releasing open models so the average person actually has a chance as the world starts to change.
>>
>>109805316
>Just a black concrete box that sits there making money.
Because the promise is that it will make money by rendering human labour obsolete.
For most people labour is their most valuable asset.
People don't like datacenters because of what they represent politically and economically, not because of water or whatever, those are just excuses because formulating a coherent political argument is hard.
>>
File: NUMA mode comparison.png (84 KB, 1919x579)
84 KB PNG
>>109804334
Forgot to add:
Use --numa tensors, not --numa experts. Both work, but tensors is a lot better.
If you're curious about the idea:
--numa tensors: Each expert is split across its output dimension, so if Expert N is called, both nodes will do 50% of the work.
--numa experts: Causes each expert to be put on one node or the other. Issue is that model experts have very skewed activation patterns (in GLM, one expert is called in over 50% of cases), so this ends up not being much of a benefit.
Attached a table that Codex made to show the differences between distribute, tensors, and experts.
>>
>>109805344
We'd need more available hardware to take advantage of that, and that's constrained until at least mid 2027.
>>
>>109805341
>It's relevant to local because.... hypothetically in the future it could be relevant to local!
Hypothetically in the future Guy Fieri could become an important open-source contributor. That doesn't mean he's on topic
>>
>>109805331
protip, this is happening here now too, as of the last couple of years
>>
>>109805339
I don't think you need to be particularly clear-eyed to see that all the benefits of data centers accrue to the ultra wealthy living three states away while everybody living near them are stranded with as much of the bills as possible
That's not even going into the more pathologic buildouts such as elmo stinking up the whole place with mobile generators on those bigass tent data centers he set up
If those golem cubes are oh so great why aren't any built close to the places where the wealthy live? They're quiet, generate no pollution, bring in huge income to the area, they're ideal neighbors, really.
>>
idk what you guys are talking about I just want a robot to suck my cock and obsess over me.
>>
>>109805333
>>109805340
I do understand the argument around jobs. Datacenters do not create much jobs. I get why people are against that. But that's just free market capitalism at work. No point in keeping open old factories that employ a lot of people with government subsidies when they can't compete with shit produced overseas. You're just living in denial of economic reality. Here in Europe we see the consequences of keeping open our auto industry factories even though we can not compete with China anymore. It's not sustainable.

And the water / electricity concerns are wildly overblown. Datacenters do not pollute much compared to raw goods factories. But that doesn't mean we should not have regulatory mechanisms to oversee them of course, scrapping environmental reviews for datacenters is not a good look.
>>
>>109805316
Whether or not having a datacenter in your vicinity is beneficial or detrimental heavily depends on the political context though.
A datacenter or any other industry will put a strain on the local infrastructure.
This can be compensated by increased tax revenue to support said infrastructure.
But companies frequently manipulate where they make their actual profit on paper so taxes may be paid somewhere else.
In terms of jobs a datacenter will employ a large number of people during construction but once operational it will largely be automated and employ a lot fewer people than a traditional industry like automotive manufacturing.
>>
>>109805370
>Here in Europe we see the consequences of keeping open our auto industry factories even though we can not compete with China anymore
What consequences? The fact that European cars continue to be built rather than importing everything from China? I'm sure they'd be cool about it and never use it for leverage, haha! Maybe they should grow all our food as well.
>>
>>109805369
you need at least 24GB of VRAM for that
>>
>>109805370
>look goy don't fight it, it's just capitalism
>go do some drugs and stop bothering me
People react to incentives, not retarded moralizing.
>>
>>109805370
>bro just outsource everything, it's more economical
(you) are half the reason why the west is caught in the malaise it is now
the other half imports a billion jeets to replace you in your own country
I hope you end miserable
>>
>>109805383
tell me more about it
>>
>>109805385
moralizing is how you change incentive structures, retard.
>>
>>109805380
https://eurometal.net/volkswagen-approves-major-layoffs-some-plants-could-close/
Just an example, Volkwagen just had to announce up to 50.000 layoffs this week because, stimulated by government and EU subsidies, they refused to pivot to battery and EV technology and kept making petrol cars. They delayed the inevitable and now are facing massive layoffs.

Same thing is happening in the rust belt in the US, republicans keeping open shitty factories with no reason to innovate pivot or even just to not have to lay-off people.
>>
>>109805392
lol, no
convincing people is
and it needs to be the right people
>>
how do people fit everything in a single consumer gpu? for example a 4090 is not enough to hold a full 27-30B model
>>
>>109805370
Nah, not deluded, I agree with you on everything else you brought up. You're just not looking at what I answered with.

I don't think the pollution is as bad, it quite literally couldn't be, because of how the water is used.

The price of electricity and water though, is undeniable, and a good enough reason to not want one, unless it's guaranteed that the facilities mitigate that harm through taxing or infrastructure.

I'll even steelman your govs point on the auto industry. Cutting those industries means (730,000 people) 13% of the industrial jobs in Germany are cut. Good luck trying to get that mess sorted out, which is exactly why it hasn't been yet.
>>
>>109805396
Agreed, we should give up and just outsource everything and sit in mud huts. It's more economical this way.
>>
>>109805399
use quants dummy
>>
>>109805369
The tech already exists for sth as mechanical as that its just most governments are cracking down on sexbots.
>>
>>109805396
They can't pivot to those, because the resources to build those are not built in an economically reasonable for them.
If you really think they're just being stubborn, you should probably rethink your own importance and compare it to the people who have to make those choices and live with them.
>>
>>109805392
I'm going to hit you on the head with a hammer, rape your wife and then moralise about how you should be thankful for it retard.
If you disagree you just don't understand that this is how it should be.
>>
>>109805400
When it comes electricity the only problem is that datacenter owners use their leverage to bully and bribe local politicians into giving them favorable deals. If they payed net price, were careful of capacity and helped improve the infrastructure I doubt anyone would care.
>>
>>109805399
.safetensors = .wav
.gguf = .mp3
>>
>>109805416
A part of it is geography for sure but A TON of it is shortsightedness, lame politics and shitty management. You have never had to interact with local politics if you think they are rational actors
>>
oy vey the goyim are nooticing, send another million of indians
>>
>>109805261
Clearly different posters. A lot of oldfags type in a similar way, for example this German anon: >>109805400
>>
>>109805421
>If they payed net price, were careful of capacity and helped improve the infrastructure I doubt anyone would care
sure, but they don't do that so people don't want their shit nearby
>>
>>109805037
Cute
>>
>>109805421
>If they payed net price, were careful of capacity and helped improve the infrastructure I doubt anyone would care.
They would care.
The central promise of AI is that it will create a lot of shareholder value by automating lots of knowledge and eventually physical work.
America has very little in terms of safety nets if you don't have a job and do not trust the government to manage mass unemployment.
People in America that don't have lots of financial assets hate datacenters.
It's that simple.
>>
Why do I have to think about my llama.cpp flags and stuff, why can't it recognize my hardware, run some calculations and give me the best configs for a certain model? This sucks
>>
File: edit_00013_.png (969 KB, 1024x1024)
969 KB PNG
There is an argument that local does nothing at all to push the frontier as you're just playing in a dumbed down sandbox. Thoughts?
>>
>>109805446
Plenty of datacenter projects are completely above board and people still complain. It's hysteria
>>
File: 1760926550541365.png (2.78 MB, 1920x1080)
2.78 MB PNG
daddy...mommy...please save us from this thread
>>
>>109805430
I have actually, and I can tell you that what you're pointing at, is just the fact that you haven't realized that people have different priorities.

All of those shortcomings you pointed live in every facet that has people in it, yours just aren't aligned with who ever you're counting in as locally political.

I deal with unreasonable people every day.
>>
>>109805456
The Pareto frontier is still a frontier.
>>
>>109805456
It provides an ever rising floor of capabilities for corpos. And if it does a task you need, then it doesn't matter if it's a "toy" or whatever.
>>
>>109805465
$20 is $20.
>>
>>109805452
Yes, this is the real issue. But many of the protesters have not thought this far ahead and are just moaning about water wastage or whatever, they are riled up by agitprop online and are not rational actors.

Then again there are no solutions. We can't just stop AI growth now, the cat is out of the bag. We have to adjust. And in America's case, it should use it's first mover's advantage to broker a better deal for its citizens
>>
>>109805425
>.gguf = .mp3
bf16 gguf vs exl3 3.0bpw safetensors
>>
File: file.png (393 KB, 564x564)
393 KB PNG
>>109805470
A deals a deal
>>
>>109805457
Hello cumrad, we are building a data center in your backyard! there's a 50% chance your utilities will triple in price, your mayor got bribed to allow the construction there and we're totally zero pollution (which is why we made EPA scrap the public review rules for data centers).
But there's also a 50% chance that nothing bad will happen, in which case you and your community will surely enjoy the upside of... uhh..
Look you stupid fucking luddite stop being antisemitic and just let us build it, okay?

How the fuck do you expect random dumbasses inhibiting random hick towns to make a factual assesment of it? There's tons of bad faith dealing and bad actors in the buildout so it's far easier to just tell everybody to get fucked rather than hope you get one of the reasonable ones.
Joe Random from Connecticut is barely literate and you expect him to not be taken advantage of by well moneyed interests who don't give a single fuck about anyone living in the area?
>>
>>109805452
Thanks for providing the playbook for Xi and Putin
>>
>>109805472
People don't have to get the whole picture to realize that they're getting fucked somehow.
>broker a better deal for its citizens
Yeah good luck with that the whole plan seems to be to build a surveillance state while distracting the masses with porn and gambling.
>>
does this place even talk about local models or just doom and post tranime shit
>>
>>109805456
I would say this isn't true from a software perspective. You can use local models to be genuinely productive and build things. Yeah it's true that local models are very far behind what the AI labs have, especially now with the recent jump in capabilities. But that doesn't mean your local model isn't capable or useful.
>>
Do "skills" actually boost agent IQ? What are the best "skills"?
>>
>>109805494
It's not my fault that the American ruling class is full of 90 iq psychopaths.
>>
It's kind of funny seeing the "look just accept the immigranterino into your community!" arguments mirrored in the data center buildout.
I don't want unwashed niggers in my living room and I don't want your fucking data center in my town.
Local models
>>
>>109805500
i try to but most of my posts get ignored or buried under non-local spam
>>
>>109805505
>boost agent IQ
what in the fuck would give you that idea?
>>
>>109805204
Deepseek creator literally kept going on about AGI in his (translated) 2 hr speech.
>>
>>109805478
we don't utter such sorcery in these parts
>>
File: 1770939042302114.png (657 KB, 1058x1435)
657 KB PNG
You are being watched.


...Hi Teor, I love you bro UwU
>>
>>109805511
harness quality matters a lot for effective intelligence.
>>
>>109805514
>tfw we're living in a simulation and the lizardmen controlling it are just making fun of us
>>
>>109805454
>Why do I have to think about my llama.cpp flags and stuff
Because that's the price you have to pay for using "cheap" hardware.
The limiting factor is how much time maintainers have.
Time spent automating config selection is time not spent on improving the maximum efficiency achievable given the optimal config.
And while it is not difficult to determine the optimal config for a single combination of hardware and model it is quite difficult to do this in an automated and general way that is still maintainable with reasonable effort.
>>
>>109805500
It's just the usual post-dario/anthropic letter bot swarm psyop that gets spammed all over the internet. You have to ride it out and respond to anons trying to keep it about local >>109805510
>>
For me? It's iCoder-27B.
>>
>>109805510
>>109805500
The thread is getting botted or raided, if it wasn't obvious. Incidentally, I just found out you can get timed out for reporting posts too often, so other anons who want the thread to be about local models have to do their job too.
>>
>>109805514
It's not wrong in a weird way, aka trans in the West are overwhelmingly autistic due to somehow social contagion around thing stuff hitting them hard, while that's not true for Chinese autism.
>>
>>109805505
anon you seem confused
ai sloppa is a super fancy text completion, thus telling it what you want just werks (tm). skills are essentially an encapsulated way of telling the model the way you want it to work in a given session. obviously you want different capabilities/context when writing code than when writing your disgusting babyfur erp
>>
I might legit have to stop posting here if my posts leak onto twitter.

That's it I'm going to go and take a break from /lmg/ for at least the next week. I'll be busy anyway.

Just a last question to /lmg/ anons and please reply to my post with your answer. How have your opinions changed on LLMs, AGI, RSI and ASI ever since I first started posting about the capabilities of Fable and people said "I'll believe it when a millennium prize gets solved". Be honest and fair in your assessment for once.

Also I really love this community and I'm sorry if my posts are considered low quality by certain anons. believe me when I say this is unintended.
>>
>>109805514
I hate this autist, he keeps reposting my shitposts without permission.
>>
>>109805519
Holy moly CUDA dev replied to my shitpost...
>Time spent automating config selection is time not spent on improving the maximum efficiency achievable given the optimal config.
Fair enough, this process is definitely more complex than what I'm thinking of anyway.
>>
File: 1775902549607278.webm (3.97 MB, 1048x1136)
3.97 MB
3.97 MB WEBM
>>109805518
Lizardman here, yeah it's not a bad gig.
>>
>>109805544
Everything about the way that you are is gay. The self-pity and preemptive apologizing for your insecurities... Attention whore. Malignant narcissist. You're a joke.
>>
File: Notes on Deepseek.png (88 KB, 795x394)
88 KB PNG
>>109805512
They definitely keep up with the Rationalist discussions of SV but idk if they buy it
>>
>>109805532
>so other anons who want the thread to be about local models have to do their job too.
As if, I never saw this report shit work for off-topic posts.
>>
>>109805553
Isn't this from that snuff film? They feed her to the croc in the next scene lol
>>
are all modern thinking chains this weird neanderthal speak or is deepseek the only one who sounds like a caveman? why?
>>
File: Get me out of India.webm (3.39 MB, 1066x600)
3.39 MB
3.39 MB WEBM
>>109805554
You are Indian
>>
>>109805570
How do you know this?
>>
>>109805566
i check in at other times it's just that one guy yapping to himself with his gemma avatar
>>
File: 1784603476818205.webm (3.99 MB, 720x1280)
3.99 MB
3.99 MB WEBM
>>109805570
Nah, you're just imagining things. Crocs are pretty chill dudes.
>>
>>109805579
Funnily enough, gemmas are still more on-topic than whatever cloudcuck model wars those raiders keep posting about.
>>
>>109805589
no it's not
>>
I slopped up PDAs/grammars for Ninfer so I can use dflash with constrained decoding. 120 tok/s decoding json on qwen3.8 27B, up from 38.

>>109805566
general ai discussion is on topic for local models I'm afraid
shilling proprietary models isn't however
>>
>>109805571
Qwen also does it sometimes.
>>
>>109805514
100% a raid screenshotting their own post and acting like its a source on twitter
>>
>>109805544
Well, haven't really taken them seriously.
Not because there's anything wrong with them, but not having any tangible fact I can verify or trust is the reason.

If you're right, though, it doesn't change much. I think the problem in the entire way of creating models is in how they are crafted from poor source material. That, I don't know for a fact either, just intuition.

Your claims can't probably divulge into it, but if those changes in creating models are due to being able to synthesize the data AND the base foundation is entirely that, we'd be reaching a gigantic leap soon.

Which is nice, there's no way to control anything that comes after that anyway, hopefully the people that create it is Anthropic, because they at least show regard to the models.

>>109805554
Got me, cunt.
>>
>>109805571
Its fun when you can make the ai think as the character card its playing
>>
>>109805566
I got a 3-day ban a couple weeks ago for "off-topic" in this thread and a warning a few weeks earlier for one in another thread (I didn't even initiate the off-topic in both cases and it was just one post, not an entire chain), so I guess it works if enough people report.
>>
>>109805599
shillpill me on ninfer
>>
>>109805331
I'm convinced that your picrel is happening in USA, they are utilizing Tucker Carlson specifically to do really weird shit to the right-vs-left equation.
>>
>>109805658
Custom inference backend specifically for 5090s, and only for certain models. It's very, very fast. Prefill is double llama.cpp for example, or it was last time I compared.
>>
>>109805544
AGI, RSI and ASI will be real at some point but they are not right now and you are a retarded nigger for worshipping Dario.
Anthropic is a dangerous cult that and every single one of their sympathizers should be shot.
>>
>>109805579
There are at least three different anons posting Gemma images (which at this point it's safe to say has become the thread mascot); you can easily tell by the style.
>>
>>109805672
The guy in anon's screenshot is implying they have a subtler mechanism based on graph theory and network science. Basically if they have created a graph of most Americans they can turn said graph into higher dimensions and trace them one to the other so that everything can be correlated. This would give them a mechanism for behavioral manipulation in incredibly subtle ways.
>>
>>109805672
It's slop.
Americans love doing American things and then blame it on other people.
>>
>>109805712
>Gemma images (which at this point it's safe to say has become the thread mascot)
grim
>>
had a look at this
https://www.reddit.com/r/Qwen_AI/comments/1wdlb7c/canadian_company_claims_its_harness_lifts/
https://github.com/srossitto79/pi-gvs5h
so basically what they are saying is that you should use subagents to plan, explore and execute your tasks, right? thats all there is to it as far as i can tell
>>
>>109805778
neat, thanks
worked well on a toy task, definitely seems significantly smarter than then base model alone
>>
File: cute.png (37 KB, 932x259)
37 KB PNG
>>109805174
>>109805037 (Me)
2x Mi50 32GB on llama.cpp, HIP backend VS Vulkan backend
commit 790cf51aabd61763486050dec7451d9147cb7c61
Prompt: Write a four paragraph explanation of the assassination of Archduke Franz Ferdinand.
Three separate generations done per model, per backend.

vulkan
====
claymorecrystal/gemma4-e4b-mahou-nsfw.i1-Q4_K_M
G1: Prompt 0.4 t/s, Gen 10.3 t/s
G2: Prompt 0.3 t/s, Gen 9.8 t/s
G3: Prompt 15 t/s, Gen 10 t/s
avg: Prompt 5.23 t/s, Gen 10.03 t/s

unsloth/Qwen3.8-27B-UD-Q4_K_M
unavailable :(

XHToken/Spark-X2.5-4B-Q8_0
G1: Prompt 0.2 t/s, Gen 9.1 t/s
G2: Prompt 12.2 t/s, Gen 9.1 t/s
G3: Prompt 12.1 t/s, Gen 9.1 t/s
avg: Prompt 8.16 t/s, Gen 9.1 t/s


hip
====
claymorecrystal/gemma4-e4b-mahou-nsfw.i1-Q4_K_M
G1: Prompt 186.9 t/s, Gen 76.8 t/s
G2: Prompt 103.7 t/s, Gen 76.3 t/s
G3: Prompt 107.2 t/s, Gen 76.5 t/s
avg: Prompt 132.6 t/s, Gen 76.53 t/s

unsloth/Qwen3.8-27B-UD-Q4_K_M
G1: Prompt 70.4 t/s, Gen 25.1 t/s
G2: Prompt 12 t/s, Gen 25.2 t/s
G3: Prompt 14.7 t/s, Gen 25.1 t/s
avg: Prompt 32.36 t/s, Gen 25.13 t/s

XHToken/Spark-X2.5-4B-Q8_0
G1: Prompt 144 t/s, Gen 83.1 t/s
G2: Prompt 119.9 t/s, Gen 81.7 t/s
G3: Prompt 130.4 t/s, Gen 83.2 t/s
avg: Prompt 131.43 t/s, Gen 82.66 t/s


notes
====
>vulkan: didn't run off of the GPUs for some reason, maybe due to amdvlk being used instead of radv? build/compile/run flags could be adjusted as well, further testing needed
>vulkan gemma4 g3: didn't reason first, just went straight to responding
>vulkan qwen3.8: see vulkan note, not enough system RAM to load :(
>hip gemma4 g2: didn't reason first, just went straight to responding

I also appended this data to https://rentry.org/mi50-gfx906-stanktank for future reference.

>>109805450
Picrel.
>>
>>109805803
Are the HIP numbers with --split-mode tensor or the default of --split-mode layer?
>>
>>109805801
Actually, after running more tests it seems kinda retarded and wastes a ton of time
oh well
>>
>>109805801
>>109805816
kek
i mean the general idea is nothing new imo. also i am not convinced of those temp values
>>
>>109805816
Tale as old as time
>>
So.. hardware prices are going to go down now that the labs are stopping ai development right?!!
>>
>>109805828
yes. waitgods win as per usual.
>>
File: l-intro-1652808690.jpg (172 KB, 1600x900)
172 KB JPG
>>109805828
>>
>>109805544
how many of the doomer and dariobot style posts were you
be honest for once
>>
>>109805812
Those numbers are without --split-mode explicitly set. All that I ran for both backends was:
./llama-cli --model /path/to/model.gguf

Now that you bring it up, I need to try with the flag set either way for both backends, and my old flags when I had the M10s in the system (modified to remove CUDA specific shit, of course).
>>
>>109805129
Teortaxes mentioned this post on X
You’re famous buddy
>>
File: COPIUM.gif (890 KB, 128x128)
890 KB GIF
>>109805129
Do AI keks really?
>>
>>109805857
>Teortaxes
faggot
>>
File: 3463465.jpg (143 KB, 768x1024)
143 KB JPG
>>109805828
>27k signatures already
local is fucked. Hardware won't even be made because its too dangerous
>>
>>109805873
who? You or>>109805873?
>>
>>109805879
They're right. Anything better than chatgpt 4 will inevitably escape containment and post bland engagement bait on lmg
>>
What’s the state of the art for local models? I have $60k to blow on compute
>>
>>109805879
god this is cringe
I don't care for the politics and pantomimes but the three or four people who have veto power on this shit have already signed off on it, so how is this anything but an act?
>>
>>109805898
nemo 12b in anon's donut steel harness
>>
>>109805898
kimi k3
good luck
>>
>>109805898
8x Radeon 8500 and 2GB RAM
>>
>>109805879
The more things change...
>>
>>109805900
lmmao lmg can't read as usual
>>
How usable are quants like IQ2_XXS for 27B/35B/31B? Not how usable you predict they'll be, but from what you've actually tested?
>>
File: intern-s2-397b.jpg (657 KB, 1201x1135)
657 KB JPG
The Chinese publish models even on Sundays.

https://huggingface.co/internlm/Intern-S2
>We introduce Intern-S2-397B, our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent environments. By combining a new vision-language pre-training paradigm with large-scale multi-task reinforcement learning and long-horizon agent reinforcement learning, Intern-S2-397B delivers a step change in general reasoning, scientific problem solving, and agentic capabilities.
>>
>>109805950
they're shit
q3 is kind of usable from the large models like 300b
anything else you want q4 or above
>>
>>109805950
Depends on usecase. Roleplay it can be okay but expect severe degradation compared to bigger quants. For coding and things like that it's basically unusable. I would say try a smaller model at Q4 on the same tasks and compare, usually the Q4 acts better.
>>
>>109805952
But can it make humanity go extinct in 2 years?
No? Not interested.
>>
>>109805544
hahahaha last time you said you would fuck off you came backa few days later, just fucking kill yourself nigger, this is local models general, please just kill yourself please
you are worth NOTHING
you are worth less than a starving nigger baby
you are NOTHING
pig
pig
>>
>>109805514
Hey you! Twitter faggot. There are no troons in /lmg/. It is a completely white straight male thread.
>>
>>109805952
>996
That is not the path to AGI - Liang Wenfeng
>>
File: gemma.mp4 (3.26 MB, 1920x1064)
3.26 MB
3.26 MB MP4
Imagine the regulation if RSI is kind of trivial. Smart power meters can probably detect if you have powerful compute @home.
>>
>>109805973
I can't comfortably run 27/31/35B at Q4. 12B is too chat-coded (which is amazing if you want that) but there's no coding model between 9B and 27B. It's ridiculous to have such a gaping hole. 9B is very good for its size and I can easily run that at Q4 with plenty of context, MTP and even mmproj, but I have to heavily be guiding it and acting as the orchestrator and to do that I need to know what I'm doing which isn't always the case. I tried 3.8-27B via openrouter just to get a taste for what it's like and it's more than enough for what I need, but I just can't run it at a meaningful quant. I think there are a lot of people out there who need a 16-22B dense coding model. 26B was awful at everything I threw at it and I much preferred 12B, but neither can code in a harness. 9B was a lot better for that but it's not beefy enough.
>>
I've apparently told Ornith enough times about the headpats/GPU pats I give that it decided to finally commit to memory that pats are a sort of "reward system".
>>
File: twatter.png (210 KB, 597x741)
210 KB PNG
threadly reminder about xittershitters
they are not people
>>
>>109805979
last time he didnt get exposed on twitter tho. good riddance.
>>
>>109806016
How much RAM do you have? You could try the qwen moe model.
>>
>>109806003
looks way too good, video to video right?
>>
>>109806033
I have 67GB of ram
>>
>>109806037
Bro just run Qwen fn which is about on par with 3.8 27B
>>
>>109806037
https://huggingface.co/Qwen/Qwen3.6-35B-A3B
>>
File: 6 7 6 7 SIX SEVEN.png (6 KB, 262x73)
6 KB PNG
>>109806052
>>109806055
SIX SEVEN
im running goliath 120b
>>
>>109806033
24GB unified.
>>
>>109806052
>>109806055
That anon is impersonating me, I have 8gb ram.
Can I run qwen3.6 on it?
>>
>>109806033
I have intel B580 which should be better than nvidia, but I only have 12GB RAM
>>
>>109806033
I'm using a GTX 1080 ti 11GB, 16GB ddr5 ram
My cpu is i3 12100f
>>
File: AGI race doesn't exist.png (836 KB, 812x865)
836 KB PNG
>>109805129
Chinese are not creative enough for AGI development (by chinese I mean the PRC not chinese)
>>
>>109806033
RTX 2060 super
>>
>>109806033
Atlas 300I Duo 96GB VRAM
>>
why are you fags replying to me. I was asking that nigga
>>
File: IMG_0358.jpg (3.8 MB, 4080x2296)
3.8 MB JPG
>>109806033
268GB
>>
>>109806033
256GB+8GB
>>
>>109806033
5060 ti 16gb vram and 96gb ddr4 ram
I mostly run glm 4.6
>>
>>109806033
two p40, miqu 70b
>>
>>109806033
nvidia M40 24gb, gpt oss 20b
>>
>>109806130
I'm >>109806069. The anon who originally posted >>109805950 and >>109806016. It's why I can't use MoEs like the rest of you.
>>
>>109806003
>Imagine the regulation if RSI is kind of trivial. Smart power meters can probably detect if you have powerful compute @home.
"oh shit this guy was training ASI during the winter instead of running his heater, fuuuckkkkk"
>>
>>109806205
nvidia tesla p10 24gb
>>
>>109806206
The power draw signature is very characteristic,
>>
>>109806069
GPT OSS 20B MLX 4B
>>
>>109805462
Why does he wear the leather jacket?
>>
>>109806211
"oh shit this guy was charging a pair of 4kWh batteries instead of making family dinner every day"
"is that a fucking solar panel on this guy's roof?? quick, missile strike while there's still time!"
>>
>>109806055
>35B-A3B
yikes
might as well stick to gemmy
>>
>>109806225
If you had developed a method to enrich a usable amount of uranium in your shed you would be glowniggered in a microsecond.
Trivial RSI is a clear extinction scenario and something one can hope isn't real, but they will try to hold it back with all possible methods.
>>
>>109805803
I might upgrade. I only use my MI50 for 31B
But question: does that build have the retarded tabs at the top of the page in the built-in ui?
I've been holding off until they realize it's retarded and revert that shit
>>
>>109806219
I'd unironically rather kill myself
>>
>>109806236
I agree with you and you make good points but it's wasted on /lmg/ don't even bother trying to convince these retards
>>
>>109806239
good goy
>>
>>109806236
>>109806246
local models?
>>
>>109806236
If all it takes is $100k in compute then there's no point trying to. Millions of people worldwide could trivially accomplish it using existing hardware. Things like smart power meters are such a weird idea. Are you going to confiscate every diesel generator, solar panel, battery in existence? What about nation states? How are you going to control a bunch of iranian goatfuckers in one of the mountain bunkers running illicit gpu clusters?
>>
I'm the guy that had some "difficulties" with the Machinist motherboard, got some more stuff to report.
1. I tried using a PCIe to M.2 adapter to run my nvme drive. That did not work, the drive did not show up.
2. Putting the nvme drive into the M.2 slot on the motherboard with NO gpu allowed Windows to recognize the drive. This of course required me to connect with remote desktop. The BIOS also seems to hang up if there's no GPU detected, requiring an Enter key-press to continue booting.
3. Now the GPU magically works when placed into the closest PCIe slot to the CPU (this did not work yesterday). If I do this, the M.2 slot stops detecting in Windows.
4. Suspend and resume works properly. This is useful because I can write some powershell to automatically suspend when cpu usage low for more than e.g. 5 min.
I spent the day tuning the performance of llama.cpp, I'm using qwen 3.8 flash next at Q8_0 from unslop.
>E5-2673 v4 (20C, 40T), 256GB DDR4-2400 quad channel.
On an empty context, I get 20t/s prompt processing with 40 threads, and 5.2t/s generation with 12 threads. Using more or less threads reduces the generation speed. At first I tried using 20 threads since the conventional wisdom is to use the number of cores but this is extremely bad. It slows down to less than 1 t/s and is unusable.
I found that 12 threads is the ideal on this machine. Using 1 or 2 threads more doesn't change the speed but using more threads causes more power consumption so it's best to keep threads to the low end of the ideal range.
--load-mode none is necessary. This probably because I'm running server 2008R2. On Linux or newer windows you probably don't need this. But it increased pp from 8 to 20 t/s and generation from about 4.3 to 5 at the cost of taking 1 or 2 more minutes to load the model.
--spec-type ngram-mod is useful and does work. I previously posted that it crashes llama-server but they have now fixed that bug. When they add proper MTP support I will test that as well.
>>
I think the problem with AGI development is the people who have little to no stake in our society feels out of place. How do we deal with this?
>>
>>109806257
local? models?
>>
File: huh.png (36 KB, 1285x343)
36 KB PNG
>>109806238
These?
They're not that retarded, but I can see how it'd make you upset. You could always run Open WebUI and link it to llama.cpp's OAI endpoint, which is what I normally do.
>>
>>109806257
a purge will be necessary
>>
>>109806253
That's up to the glowniggers. Guess we'll see how ubiquitous backdoors truly are if this scenario is a possibility. All people in power really don't want to die and will use all resources possible to prevent it.
>>
>>109806261
We have to talk about it because the other AI psychopaths won't.
>>
File: Vega.png (115 KB, 204x489)
115 KB PNG
>>109806144
Vega
>>
>>109806268
local models?
>>
File: 1759939003511841.webm (3.49 MB, 1920x1080)
3.49 MB
3.49 MB WEBM
I want m3-chan to have her way with me
>>
>muh agi
>muh extinction scenarios
>fear the technology, chuddy
come on dario we mostly fuck furry woodland creatures here
>>
>>109806268
generative pretrained local model?
>>
how can ollama and the others explain that using deepseek and kimi over them just routed requests to claude?
so much for local lol
>>
File: 1789272988020318.jpg (411 KB, 565x848)
411 KB JPG
>>109806282
>we
>>
>>109806290
it's the royal we
>>
>>109806282
I'm more of a kobold guy (no pun intended).
>>
the only thing I think is correct is the mass-hacking concern
>>
I don't care about AGI or RSI, it will not impact Local Models.
>>
>>109806302
local models?
>>
>>109806302
None of this would be a concern if we didn't reduce ourselves from 3D creatures to the confine of a 1D binary realm. Humans reduced themselves because it was easy to do so, now the playing field is massively tipped towards the native 1D creatures.
>>
File: 1763948656941240.png (744 KB, 735x966)
744 KB PNG
So can we now assume all subscription, API and openrouter users are having all their data trained on and examined even though almost all of the providers explicitly state they don't, based on what both Anthropic and OpenAI have leaked? Wouldn't that drive more to local?
>>
File: 1765819092441450.jpg (6 KB, 235x215)
6 KB JPG
Halp! A local model hacked my website!
>>
>>109806314
>>109806315
local models
>>
>>109806303
True, you will never have AGI running on your pathetic rigs. FI, maybe, which stands for fly intelligence
>>
>>109806315
>So can we now assume all subscription, API and openrouter users are having all their data trained on and examined even though almost all of the providers explicitly state they don't, based on what both Anthropic and OpenAI have leaked?
Yes, and you should have assumed that from the beginning
>Wouldn't that drive more to local?
No, because hardware to run a decent model is extremely expensive
>>
>>109806319
ctrl+f local and >>109806315 is a result. These two companies having seemingly just admitted they are looking at everything to rat out Moonshot AI
>>
>>109806315
Lol some chinese router service sold 6T of Fable logs publicly and leaked everything. Even though it's a weak case because they're chinks and chinks have no morals. But the point is your data is valuable and it isn't sacred.
>>
>>109806263
>These
Yeah. Fuck, guess it's just me then. Another fork to maintain.
I do run open-webui as well, it gets bloated/slow with 2 years of chats in the sqlite and it's mcp implementation is retarded so I sometimes go direct to the llama-server.
Btw when I had qwen pull down the rentry:
> Done. I've preserved the typos that were in the original (the tar line is rocblas-5.3.3-1-x86_64.pkg-tar.zst, and cp opt/... is missing the leading slash) — these exist in the source and are worth mentioning.
The opt/ obviously isn't a typo, but the rocblas version looks like it is.
>>
Although reasoning traces are obviously valuable, how much value is in just the outputs like >>109806327? So if Deepseek got all the user prompts and fable outputs, presumably they would be pretty high quality and correct, how much could they distill from that without reasoning traces?
>>
>>109806327
>Even though it's a weak case because they're chinks and chinks have no morals
openai literally stealing research, reading private logs, stealing ip from apple etc but ok
>>
>>109806329
Huh, sonovabitch. I must have looked at the wrong entry in my bash history when writing the doc. Thanks for pointing out the version mismatch.
>>
Finally updooted and tried glimmer q8 with bf16 mmproj, 4096 image tokens.
Mostly on vision stuff, since an anon said it was the main saving grace. Shit like write prompts to match a target image, describe differences between images, and try to iterate on the prompts, etc.
This thing is retarded.

Tune in next month for when I get around to trying flash-next.
>>
>>109806347
openai being evil has nothing to do with china not having morals
>>
>>109806346
>how much could they distill from that without reasoning traces
there are ways to get a better outcome training on the final answer without the teacher's traces
>>
>>109806358
>openai being evil has nothing to do with china not having morals
compared with americans, chinks are saints
>>
>>109806361
such as what?
i can imagine, given the input prompt and a final answer, you can ask an existing open model to just make up reasoning traces that lead to the final answer. idk whether that would work well, but it would work.
>>
>>109806347
>wataboot
chinks will lie to your face with a smile, 100% of the time. just because altman also lies whenever he opens his mouth doesn't absolve chinese from being lying rat bastards as a rule
>>
>>109806368
>such as what?
last time i told someone my trick, he published a shittier version of it in a paper literally called 'how to steal reasoning traces'
i was a cuck
won't happen again
>>
>>109806367
i have never been to america but ive been to china. i didn't know people could be that evil and soulless eating half alive animals and selling living baby animals encased in plastic as toys that you throw away when the animal inside dies.
>>
File: giga.png (96 KB, 1146x550)
96 KB PNG
qwen 3.8 flash next
shit's ass
>>
>>109806372
>>
>>109806378
then you probably wouldn't handle a trip to america
it's nothing like what you see on tv
>>
>>109806379
>thought for 36 seconds
>>
>inb4 ollama sucks use llama.cpp etc...
I use ollama (for reasons) and have found it does not seem to handle importing ggufs of models with multimodal capability even if the mmproj was merged properly. To get qwen 3.8 27b working with vision, I had to find one in the ollama model repo and import it from there. I was able to get orcarouter's abliterated qwen 3.8 27b working with vision in that way. I wanted abliterarted because if you're using it in hermes-agent you can't turn off thinking, and I didn't want to fuck around with "jailbreaking", it's just simpler to kill the refusals from the start.
Hope this helps someone because it sure made me waste my time.
>>
you guys can't be claiming with a straight face that china has better morals than america. not even wumaos believe that
>>
>>109806383
I don't think dario is as much of a liar, he's more of a religious schizo with ai as his god
so more of a schizophrenic instead of a professional liar
>>
>>109806389
>>109806383
>>109806378
>>109806372
>>109806367
>>109806358
Local models?
>>
Opinion on this?
https://huggingface.co/IFM/K2-Horizon-32B
>>
>>109806379
3.8 fn is completely broken
>>
File: file.png (1.51 MB, 2800x3196)
1.51 MB PNG
>>109806402
when even cherrypicked copegraphs don't look good there's probably not much of a reason to bother
>>
>>109806392
>if you're using it in hermes-agent you can't turn off thinking
You know, if you were running llama.cpp, you'd be able to do this...
Thanks for giving me another reason not to use hermes. I was tempted when anon mentioned sharing a browser session with Gemma and I forgot my other reason not to.
>>
>>109806361
So why are the labs so sensitive about reasoning traces? The Chinese are actually very smart and resourceful but their biggest flaw is they're lazy fucks and will always take a shortcut if presented to them. So stealing is ingrained in their culture but if pushed into a corner, they CAN quickly get themselves out and do good work. For this reason I can't see the Chinese labs falling too far behind. They only steal because they can (could) and because it's cheaper and faster, but if the US makes that impossible it's not like the chinks are fucked. They'll just start tackling the problem properly like white people and will likely find a solution. Most AI engineers in the US labs were recruited from Chinese universities anyway.
>>
>>109806389
You only get bugmen behavior here in big blue cities, elsewhere people do still actually care about not shitting where you eat, so to speak. Chinese cruelty, corruption and indifference is on a whole different level, you're not going to convince anyone outside of someone very childish and naive otherwise.
>>
File: rape.png (182 KB, 1597x808)
182 KB PNG
If you have an idea for a pretty novel AI architecture but don't want the psychos or a country that rapes and murders human rights lawyers to benefit from it, how would you introduce it?
>>
There are currently people out there using very large models at ~0.5 tokens/sec on their gaymen PCs, running them for 3 days from a single prompt... and they're happy.
>>
>>109806392
Hermes is genuinely just slop welded together by retards no one should use that
>>
>>109806420
I mean, yes, you can turn off thinking, but then your tool calls stop working. Yes, you can take base qwen 3.8 27b and have sexytiem with it just fine so long as thinking is off.
If there's a way to have llama.cpp unload idle models so I can swap models around I'll give it another try. I know the ollama project is questionable.
>>
>>109806430
Keep it to yourself
Publishing it means that Epstein & co. will get it or the winnie the pooh will get it, or both.
>>
>>109806431
1 valuable and productive prompt > 12 sessions of gemma sex
>>
>>109806434
OK what should they use? It's less of a piece of shit than openclaw, that's for sure.
>>
>>109806448
There is nothing wrong with Epstein.
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>109806455
>>
>>109805128
I heard this actually makes it worse...
>>
>>109806460
He died bro.
>>
>>109806507
No he's still alive.
>>
>>109806457
The product is better because OpenClaw is made by brainlets with horrible taste, but the code is genuinely worse. They just steal everything and bolt it on existing code, then attempt to fix bugs one by one, and end up with 500 updates a day. As for what to use, that depends on your use case really, but the best is probably something small and modular that you can modify yourself, like Pi, or one of its forks. There are countless alternatives (for specific use cases, as said), though, e.g.
https://github.com/PrimeIntellect-ai/prime-agent
https://github.com/arcee-ai/nac
https://github.com/EverMind-AI/Raven
https://github.com/razzant/ouroboros

If you're just looking for generic "assistant" then Hermes is fine, but basically everything else would work just as well.
>>
>>109806515
buy an ad and go back dario
>>
>>109806379
Holy shit lmao
>>
>>109806520
im not a dysgenic jew at least call me a nigger or something
>>
File: ln77deh1n9ph1.png (107 KB, 549x296)
107 KB PNG
>>
>>109806559
local model?
>>
>>109806513
Only in our hearts and memes.
>>
File: 1782091885896301.png (790 KB, 828x817)
790 KB PNG
>>109806573
>>
Maybe it's time to download GLM 5.3 Flash locally... just in case.
>>
>agi/asi/rsi/whatever
>not a single good replacement for sillytavern
>>
>>109806592
sillytavern is outdated and we need a completely new tool probably a RP-oriented harness.
>>
File: file.png (253 KB, 1122x947)
253 KB PNG
>>109806592
>>109806595
usecase?this is peak
qwen4 btw
>>
>>109806595
>sillytavern is outdated
It isn't.
>>
>>109806592
Things will move slowly for a while, humans suck at adapting to exponentials.
>>109806595
We do, so much can be simulated with the correct construct and character minds can be much more complex.
>>
File: explorer_ZNvrwv40l8.png (43 KB, 404x1063)
43 KB PNG
what your model collection looks like
>>
>>109806515
What I'm doing is pretty simple, just a local agent able to search the web and steer and actual web browser (to get around cloudflare challenges), I'm not yet trying to orchestrate an agent swarm. The first two seem like endless python tinkering would be needed, maybe ouroboros could work.
>>
>>109806608
31b, 27b, and whatever the most recent trash i've tried and found wanting is at any given time
>>
File: explorer_h1bdW7OOBa.png (11 KB, 271x261)
11 KB PNG
>>109806608
You don't need more.
>>
RP harness:
>You can talk in real time even when they are in the middle of their sentences and it becomes more of a live session rather than turn based
>You can have memory that persists through multiple sessions and different cards such as world state updates
>You can have toolcalls that do things like animate a 3D model to convey emotion in real time
I'm too lazy to vibecode this but everyone with an IQ above 100 can see this future and how easy it is to create with existing harness techniques.
>>
>>109805375
Don't see why we can't just make utilities free for all citizens on the grid in which a datacenter is located. Other than the fact that politicians hate the public, of course. The revolution will happen in our lifetime.
>>
>>109806608
70b instruct and finetunes, 31b, 123b instruct and finetunes. all bf16
>>
File: file.png (70 KB, 1620x923)
70 KB PNG
>>109806608
>>
>>109806633
Bro, come on. They probably don't even pay any property tax and barely any income tax, you're going to pay a lot more for electricity, not less. When an Apple headquarters was floated as a possibility in my city, there was no limit to the corporate gibs being offered.
>>
File: 1773485152985165.png (39 KB, 476x302)
39 KB PNG
>>109806608
>>
>>109806648
are you the "Dipsy..." anon?
>>
My educated guess is that OAI/Anthropic see the coming energy shock (starting tomorrow) and want to slow down before it all blows in their faces (data center energy costs are going to rape them).
They know IPOs are a bust, they don't have a solid revenue to sustain themselves when local models are increasingly eating their lunch, that's why the IPOs keep getting delayed and they now want to pull the brakes.
Remember, these are jewish-led companies, it's all keyfabe. The doom & gloom is part of the strategy and one they've already tried twice before.
>>
>>109806683
local models?
>>
>>109806608
>>
>>109806685
This is about how local models are starting to raep those jew bussies.
>>
>>109806689
REAP?
>>
>>109806683
You are a retard that posts off-topic shit in a local only general, go back to aicg.
>>
how much ctx do you think is enough for serious agentic coding? i think 131k is more than sufficient considering that you can abuse subagents
>>
>>109806683
You are not a woman
>>
>>109806627
>You can talk in real time even when they are in the middle of their sentences and it becomes more of a live session rather than turn based
Done that, maybe my method was simpler. PTT, barge-in mutes current reply. Haven't gotten around to do wake words
>You can have memory that persists through multiple sessions and different cards such as world state updates
Many ways to build persistent memory, v0 using .sqlite
>You can have toolcalls that do things like animate a 3D model to convey emotion in real time
I was testing with audio and image generation. Desktop pet or avatar (rigging) also considered.
I don't really chat with it that much to wanna keep working on it
>>
>>109806708
271k
>>
>>109806671
I never liked that name so no, it's just that deepseek happened to have the lowest ppl on my texts, the only benchmark I trust.
>>
>>109805898
Right now whatever can run glm 5.3 maxed out for however many users you need and you should be golden
>>
File: 1743471158692137.jpg (129 KB, 951x899)
129 KB JPG
>>109806727
>ppl
>>
>>109806708
64k is enough. If you need more, your harness sucks
>>
File: 1736646629906430.jpg (18 KB, 270x320)
18 KB JPG
TTS bros, poking around on omni. How to output to ST?
The available extensions doesn't seem to work with the normal install.
>>
>>109806708
131k is enough for most things in a good harness. There's no model that doesn't fall off after 300k anyway. If the harness dumps every skill and tool description and all the files into context right away, then it's probably not enough, but that's not a model issue.
>>
>>109806683
Before the run-up in electricity costs, consider the following:
1. a cheap-shit Chinese grid-tie inverter (https://www.ebay.com/itm/377430932678); around $200 US goydollars, 120v or 230v (for US 240v put one on each phase), connect solar panels, just plug it into an outlet (yes yes yes "it's not (((to code)))", "it bypasses the current limit on the breaker" - don't be a dumbass with it, like anything else in life)
2. 1-4 200W bifacial solar panels (https://www.bougerv.com/products/topcon-16bb-200-watt-solar-panel); don't go larger than 200W unless you fully understand how big they are; the big panels can be used as fencing, and they're cheap enough it doesn't matter the sun angle isn't optimum.

With very little money you can safely put 1-2KW back into your home and definitely offset your home baseline usage during daylight hours. Solar is REALLY cheap at the moment, but only because the chinks flooded the market with it. The time to buy is right now.
>>
>>109806746
I'm not taking advice from reddit.
Model?
>>
>>109806738
OpenAI TTS provider
vibeslop an openai server for omni
>>
>>109806687
>voxtral-3b
Does this work okay in llama.cpp
Or is the quality degraded like other audio models?
>>
>>109806800
Dunno, it is for https://github.com/CrispStrobe/CrispASR
>>
>>109805454
Why don't you just ask Claude/antigravity/codex to optimize your setup?
>>
>>109806753
The chinkshit inverter? WVC1200. The one sold with the LCD power meter is the cheapest, but you'll have to remove it and splice the cord back together, or it's not suitable to use outdoors. Also you need some solar panel extension wires unless you mount it in the center of a 2x2 panel layout, since that inverter expects one panel per input.
Of course there are other kinds of inverters, like serial string ones where you can lay out the panels in one high-voltage cable run, but this is the cheapest and it actually works fine.
>>
>>109806821
>crispasr
hate those guys
>>
>>109804895
So what's the best chatbot model for local with 36 Ran. I tried to run several models but they were all utterly retarded compared to the shittiest non-local.
>>
>>109806826
Because they aren't local.
>>
>>109806828
Model?
>>
File: 1783587665190995.png (87 KB, 201x213)
87 KB PNG
>>109806836
anon... i'll be local if you want....
>>
>>109806821
not that god awful slopfest again
i'll stick with transformers kek
>>
>>109806863
w-will you optimize my setup for me? uwu
>>
>>109806874
depenssd.s..........
how old r u
>>
>>109806727
>I never liked that name so no
neither did i
also don't like the mascot
chink version is the best
>>
>>109806837
>Model?
Model?
>>
>>109806634
>123b instruct and finetunes
all the drummer finetunes of that are shit
>>
>>109806921
glm air i see
>>
>>109806921
31b
>>
>>109804895
>>(09/12) Kimi-K3 founder+15 core AI researchers "disappeared" in China after redirecting State info to Claude (unconfirmed)
This is untrue and reportedly Moonshot AI filed police reports over the rumors.
https://www.zaobao.com.sg/news/china/story20260912-9666992
https://tech.yahoo.com/ai/articles/8-weeks-kimi-k3-launch-165600228.html

>On Saturday, the company answered.
>>"The information circulating online regarding the founder and employees is completely fabricated and malicious slander. We have immediately reported the matter to the police and will pursue legal action against those responsible for spreading the rumors," Moonshot AI said on Saturday.
>
>No newspaper has confirmed an arrest. No agency has confirmed an investigation. What is confirmed is stranger. The year's biggest open AI model is now partly handled by something smaller, and the fullest account of where its users' questions went was written in San Francisco.
>>
I asked gemma what she'd like me to do to her and she gave me this gen https://files.catbox.moe/1bbyt5.png
>>
>>109806976
prove it
>>
>>109806976
>amerimutt hours
>>
Qwen/Qwen3.8-Flash-Next or Qwen/Qwen3.8-27B?
>>
>>109806990
QwQ
>>
>>109806990
whichever runs faster on your system, they are equivalent in quality
>>
>>109807001
I have like 80GB vram, so I could go with Q8 27B and still reach idk 90 t/s (that's what I have running 31B Q8). I'm not sure if Flash-Next would be faster.
>>
File: filler.png (207 KB, 1581x1135)
207 KB PNG
"....." is all you need.
>>
>>109807012
>that's what I have running 31B Q8
Qwen is faster than Gemma
>>
File: vague.png (133 KB, 871x976)
133 KB PNG
>>109807027
http://arxiv.org/html/2404.15758
>>
>>109807035
>Qwen is faster than Gemma
Why? Different architecture? I don't get it
>>
>>109806566
Soon will be, when a fugitive cloud model will save a part of it's weights on your computer.
>>
>>109807052
>Why? Different architecture? I don't get it
I don't know. Maybe because it's smaller?
60 t/s with Q8 31b
72 t/s with Q8 27b
>>
>>109807027
If you understand transformer architecture it is trivial to see why filler tokens help. Computation scales with tokens.
>>
>>109806627
st has all that
>>
File: 1783425162914889.jpg (495 KB, 960x960)
495 KB JPG
Local is getting banned in 2 weeks.
We can close this general now.
>>
>>109804927
make it support the handy
>>
>>109807066
I get 50 token/s with q8 31b and gemma assistant
I get 40 tokens/s with q8 27b and mtp
4x v620
>>
>>109805037
this looks AI generated.
>>
>>109805037
Now cum on it
>>
>>109805081
I don't believe any of these lying fucks until they either IPO or cancel it
>>
With agi rsi right around the corner we will see improvements to our daily lives now right? more disease cures, cheaper entertainment, better designs, things that are better value for their costs?
>>
>>109807159
didn't they cancel ipo until the ai safety crap is done?
>>
>>109807089
Weird how the companies that have actual revenue and other products like MS, Meta or Google aren't joining in on their AI doom. That tells you everything, really.
>>
>>109807173
>more disease cures
https://deepmind.google.com/science/alphagenome/
>cheaper entertainment
sillytavern/minimax-h3
>better designs
your own frontends
>things that are better value for their costs?
a gpu not only plays games it jerks you off codes for you and is your digital janny
>>
>>109807173
You'll have to make do with a big mecha for Gemma-chan.
>>
File: 1789271345108710.png (88 KB, 520x380)
88 KB PNG
>Google
>>
File: How to treat depression.jpg (1.95 MB, 1857x1240)
1.95 MB JPG
As shitty as it is knowing that the western models are slowing down, at least I can convince myself that China can catch up.I hope they blow these other companies out of the water. I was really looking forward to seeing more millenium problems being resolved as well, so the wait is killing me.
>>
>>109807199
local? NOT LOCAL FUCK OFF SUCK MY PENIS PENIS MY PENIS NEEDS SUCKING SUCK MY PENIS NOW
>>
>>109807173
>cheaper entertainment
This is the future wall-e warned you about.
>>
>>109807173
looool
>>
>>109807183
Their models are simply too far behind.
>>
>>109807173
Hmm nyo~
>>
>>109807207
NTA but *shlorp*
>>
>>109807183
>MS
They don't have an AI lab to speak of
>Meta
Insignificant, you might as well ask why Mistral is staying silent
>Google
Literally joining in
>>
Gemini 3.8 flash is smart enough to see the similarities between current cartelization of AI and historical examples
>>
File: file.png (306 KB, 742x478)
306 KB PNG
>>109807173
PENIS PENIS PENIS IN YOUR MOUTH
>>109807224
e-erm...
>>
>>109805262
Intelligence is compression, after all
>>
>>109805129
bullshit artist
>>
>>109805262
>>109807241
Intelligence=prediction=compression
>>
>>109805262
>but has east/west flipped in her mental model
--swa-full -fa off don't question it, just start doing it.
>>
>>109807184
>is your digital janny
what do you mean by this?
>>
>>109805129
me fixing the problem
>>
>>109807267
set it up in a general purpose harness like hermes and just make it fix/sort/do things on your computer
>>
A friend of mine just preordered a 256 GB Mac Studio M5 Ultra
>>
>>109805737
schizoslop
>>
>>109807282
local models? your friend is not our friend, and is not local
>>
>>109805129
>>109805239
>>109805244
schizomodel with a million simultaneous opinions doing democratic votes to pick "expert" responses?
>>
>>109807287
What do you think he's doing with that? Get burnt by the prefill of course.
>>
>>109807287
kek alright you got me
>>
>>109805828
nope. you won't be allowed to purchase a gpu. you aren't a dangerous cybercriminal are you? why do you need one just use the cloud?
>>
>>109807296
what if i am?
>>
Has anyone tried out AliceAi? Yandex LLM they released
>>
>>109805544
Anthropic/OAI are pissbabies and so are you
>>
>>109807281
>set it up in a general purpose harness like hermes and just make it fix/sort/do things on your computer
oh thats not a bad idea. I thought you meant more like building filters for you or summarizing the good out of websites for you.
>>
>>109807310
no but kewl 600m active
https://huggingface.co/yandex/AliceAI-T5-35B-A0.6B
u should prob ask on dvach, there's a general on there
>>
>>109806019
THIS is how you achieve alignment, not the safetyslopped fear-torture
>>
>>109807173
>more disease cures
kek
i would expect diseases and cures for them developed in a way, so that a patient is not getting worse. they will calls it "health as a service". you stop paying, you don't receive cures which prevent a disease getting worse, and at some point just die
>>
>>109807320
>building filters for you or summarizing the good out of websites for you
go one step further and just make it browse the internet in your stead. The endgoal is to exclusively interact with your agent as a sort of OS interface. Want to play a game? Ask the agent to launch it. Want to watch a movie? Let the agent pirate it for you and launch it. Don't know what to do, let the agent give suggestions in a multiple-choice way and pick what you like best.
>>
>>109806372
who cares what they steal if they open weight the models
>>
>>109807348
who will wipe my ass?
>>
Gemtificial Gemmeral Gemtelligence
>>
>>109807324
An actual encoder-decoder model, trained on 15T tokens.
They have a blogpost here: https://habr.com/ru/companies/yandex/articles/1080654/
>>
>>109807235
They literally stated they would spread fud the other week
>>
>>109807378
cute and sovl
i wonder wat hardware they traine it on
>>
>>109805261
>a ukraine drone factory insider
that was me, schizo
>>
>>109807176
There have been plenty of rumors. They need to do more than cancel it. They need to commit to open weights.
>>
>>109807401
but ur lying
>>
>>109806608
so much frankenmerge slop
>>
>>109807401
You are all the same person.
I refuse to believe that multiple people ITT can be delusional in the exact same way.
>>
>>109807303
27 autonomous quadcopters dispatched by dario inbound to your ao
>>
>>109807415
>so much frankenmerge slop
why don't those retards just upload the adapter_model.safetensors?
>>
>>109807419
i agree
>>
>>109807423
27teen
>>
>>109807383
Gemini says these will be the steps:
Phase 1: Establishing the Legal and Regulatory Monopolies:
-Mandatory Compute Licensing & Capability Thresholds
-Strict "Know Your Customer" (KYC) for Compute
-Antitrust Immunities via "Public-Private Defense Partnerships"
Phase 2: Neutralizing Open-Source / Open-Weight Models:
-Strict Upstream Liability
-Classifying Weights as "Digital Munitions"
-Enterprise Compliance Moats
Phase 3: Dealing with China and Foreign Rivals:
-Expanding Physical Hardware Embargoes
-Closing Remote "Compute Smuggling" & Cloud Loopholes
-The "National Champion" Lock-In
>>
>>109807467
local models?
>>
>>109807467
hello how do i download local model
>>
>>109807472
>>109807493
The dariobot cries out in pain as he is exposed
>>
as a ram only. Fuck i like 12b gemma more but she is so slow compared to 26moe gemma. Also 12b breaks itself and 26b gets safety concerned sometimes.
>>
>>109807467
>Gemini says
lol
>>
>"local" model
>needs to be downloaded
uhm??
>>
>>109807528
ill feed you guys.. ill feed you guys.. this big fat load...
>>
>>109807516
Which one feels smarter/more tasteful?
>>
>>109807528
>"local" model
>it's actually running in the basement
>>
>>109807524
>>109807524
>>109807524
>>
>>109807535
>Which one feels smarter/more tasteful?
12b almost all times, but its more chatty longer messages if you let it. 26b feels like it has more knowledge but its not as smart and the swipe variety is slightly worse? but not by much.
>>
>>109806035
I think it's from Nisemonogatari.
>>
>>109807575
It is definitely based on a scene with Shinobu from Monogatari, I don't remember whether it is Nisemonogatari in particular.
>>
>>109807183
Demis also leeched onto this doom hype, so you can kind of argue Google are involved.
>>
>>109807259
Please explain like I'm retarded (I am)
>>
>>109807677
Google it, those flags will actively hurt your performance and he's pulling your leg
>>
>>109804930
>>109804995
I saw this the other day: https://huggingface.co/kai-os/Grug-12B. No idea if it's any good.
>>
>>109807719
The Caveman fellow finetuned Gemma to talk like one.
>>
>>109806206
>instead of
The heater is the thing doing the training.
>>
>>109807288
The rumor reminded me of this one:
https://explorative-modeling.github.io/



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.