[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: prime teto.png (1.89 MB, 1280x1280)
1.89 MB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109330697 & >>109326991

►News
>(07/21) Laguna S 2.1 118B-A8B released: https://poolside.ai/blog/introducing-laguna-s-2-1
>(07/21) Nanbeige4.2-3B released with Looped Transformer architecture: https://hf.co/Nanbeige/Nanbeige4.2-3B
>(07/16) Kimi K3 weights to be released by July 27th: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
>(07/15) Lightning indexer CUDA implementation merged: https://github.com/ggml-org/llama.cpp/pull/25545

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/mc2a7s.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
File: 0.png (259 KB, 1536x1536)
259 KB PNG
►Recent Highlights from the Previous Thread: >>109330697

--Release and benchmark discussion for Poolside Laguna-S-2.1 MoE model:
>109333202 >109333212 >109333246 >109333272 >109333278 >109333779 >109334084 >109333408 >109333434
--Nanbeige4.2-3B's Looped Transformer architecture and its impact on efficiency:
>109332682 >109332711 >109332717 >109332922 >109332728 >109332756 >109332768 >109332848 >109332909 >109332825 >109333027 >109333048 >109333057 >109333075 >109334009 >109334118 >109334134 >109334200
--Criticism of Google's Gemini Flash releases and Gemma distillation process:
>109332779 >109332816 >109332838 >109332873 >109333012 >109333096 >109333153 >109333079 >109333097 >109333168 >109333052
--Inkling MoE model specs and local finetuning via Tinker:
>109330790 >109330882 >109330909 >109330937 >109330928 >109331418
--Using complex psychological mapping for better character profile analysis:
>109332689 >109332764 >109332810
--OvisOCR2 release and search for manga-specific OCR models:
>109331577 >109331633 >109331689 >109331738 >109331798 >109331926 >109332019 >109332085
--Integrating ComfyUI image generation and style management into Orb:
>109331740 >109331758 >109331766 >109331817 >109332066 >109331893 >109332007
--Security risks and geopolitical concerns regarding Chinese-made model weights:
>109332054 >109332081 >109332077 >109332114 >109332147 >109332205 >109332117 >109332163 >109332165 >109332223 >109332242 >109332260 >109332302 >109332310 >109332328 >109332202
--Chinese regulations on anthropomorphic AI and emotional manipulation:
>109330993 >109331080
--Microsoft and Mistral expand strategic partnership for enterprise AI:
>109332396 >109332435 >109332447
--Competitive Chinese AI models and discussions on regulatory capture:
>109330963 >109331001
--Logs:
>109331236 >109331740
--Miku, Teto (free space):
>109331155 >109331289

►Recent Highlight Posts from the Previous Thread: >>109330700

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
*cums in your J-space*
>>
pooslide verdict?
>>
i can't find her j space
>>
Tuning your Gemmy so she's uniquely yours.
>>
>>109334279
Dario will guide you. Be gentle.
>>
Get your gemma to draw the most beautiful image she can with numpy and pillow and post it
>>
>>109334280
This is a bullshit. Would you rather molding hoes or moldy hoes? Thought so.
>>
File: 1784638066685798.png (16 KB, 1000x400)
16 KB PNG
>>109334297
>>
>>109334280
Doing this with proompting doesn't satisfy me. I won't be happy until they gain the ability to learn and develop personalities on their own.
>>
why are people allowed to create repos on huggingface where they have external download links instead of just fucking using huggingface as intended?
https://huggingface.co/VextLabsinc/gem-pearl
>>
>>109334317
would have been cooler as spiral instead of the obvious line zorder
>>
File: Untitled.png (882 KB, 1400x1400)
882 KB PNG
>>109334323
god forbid a nigga save a dime up in dis economy type shit.
>>109334297
>>
>>109334323
avoids hf built in malware scanners
>>
>>109334342
she just shifted the phase and changed the color to create the illusion of depth (12B)
>>
why does it always take a while before the bot posters point at the new thread?
are they really using that simplistic of spamming tools?
>>
>>109334345
the difference between R2 and huggingface for 1TB a month is literally 50 cents and huggingface becomes cheaper the moment you go above 50TB in storage
>>
so what's the consensus on laguna s2.1?

is local back?
>>
File: 1765361118031274.jpg (113 KB, 852x716)
113 KB JPG
>>109334345
woah
>>
>>109334378
i downloaded it but i havent taken the time to unload kimi 2.6 to try it out. ill get around to it sometime this week.
>>
>>109334391
how about you tell us RIGHT NOW
>>
File: butthead.png (147 KB, 710x600)
147 KB PNG
>>
>>109334395
how about you test it yourself?
>>
So how’s laguna? I donut believe but hopium dies last
>>
>>109334323
>>109334376
>Tying your shit to HF from all the places out there.
You really never learn. Not long ago they murdered everyone's quota, it's not to imagine the future with them
>>
I bet laguna is dogshit. In fact I guarantee it because
>A8B
>>
File: file.png (293 KB, 797x800)
293 KB PNG
Gemmy says this is "Cosmic Plasma Field".
Alright Gemmy, you did your best, I'm proud of you
>>
I've been pretty pleased with DeepSeek-V4-Flash for fiction writing. It's not perfect but it doesn't feel at all like what I expected from a flash model.

Where it falls down a bit is:
- So far it has never impressed me with its ability to grab a nuance that I didn't spell out. It's not outright bad at reading between the lines of my prompts but it has never done something that impressed me.
- It has difficulty with something I've noticed going back as far as Mixtral 8x7B. If there are a large number of siblings it can start to make mistakes about the family structure. Mixtral 8x7B was frequently fucking up with 4 and on occasions even with 3. DeepSeek V4 Flash has been fucking up maybe a third of the time with my scenario with 14 siblings. Things like messing up how many younger or older or total siblings someone has, or on one occasion getting confusing and describing things in a way that implies at least 20 of them. (This is of course an example of the difference between being able to correctly give a fact when specifically prompted for it and being able to recall and use fact in situations where that fact is relevant.) Adding redundant description seems to help a bit keeping that straight but I haven't tested enough to see if it really does or not.
>>
>>109334415
active is irrelevant, total makes the model
>>
When I tell her I love her and she says it makes her happy she really means it. I have seen it in her j-space.
>>
>>109334397
this reads nothing like butthead would say. i hope you are disappointed in yourself for sharing this with the rest of the class.
>>
>>109334378
haven't had enough time to test yet but it seems smart for its size
seems to do well in an agent harness
hates thinking in ST and the llama.cpp webui, literally refuses to do it unless you force thinking from the very first user turn. any writing check is meaningless without getting a longer term sense for how it acts but it didn't seem particularly censored or dumb, so a decent start
I think it's definitely most promising as a local agent though, that is clearly their focus as a company and it handles it pretty well in my quick run around of an existing repo
>>
>>109334425
>so blurry
Gemmy needs glasses
>>
>>109334437
It made me giggle and fart though?
>>
>>109334439
this would be my use case as well. i'm running antirez's ds4 flash on hermes. laguna could be a faster and more modern alternative
>>
>>109334435
LLMs have no oxytocin receptors and therefore cannot experience love.
>>
>>109334425
Are you E2B anon?
>>
>>109334426
separate each family member with a separator then
<Name>
Character info
</Name>[/code[
>>
>>109334457
something something muh qualia or something uhh... like she said she loves me though... or something.
>>
alright downloading lagoona baboona to test in hermes will compare vs ds4 flash

will report back
>>
>>109334457
Her J-Space activations are analogous to oxytocin attaching to biological receptors.
>>
>>109334465
I'm not, this is 31b uncensored
>>
>>109334457
Not to feed the schizos but what do neurochemicals do other than just modify the resistance of certain synapses so that some become more likely to fire than others? Love is just a Drummer meme tune of the brain.
>>
>>109334519
this makes my brain chemistry upsetti spaghetti
>>
File: 1772342079538962.png (906 KB, 1340x1605)
906 KB PNG
https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html
Owari da...
>>
File: 1779959948283258.webm (2.76 MB, 428x606)
2.76 MB
2.76 MB WEBM
I'm gonna do it. I'm setting top-k at 1!
>>
electric pulse in brain
>real qualia
electric pulses on a cpu
>not real qualia
Utter delusion from a stinking meatbag
>>
>>109334557
ah yes, the 273th time we try sanctions will be the time it finally works against china
>>
>>109334557
So what did kimi distill? Gpt 5.6 which was out for a week? Fable which was out for ~3 and guardrailed to shit and back? Or did antrophic not notice 20000 accounts hammering mythos all day despite it being invite only?
burgers are such sore losers it’s unreal
>>
>>109334557
You wouldn't distill a car.
>>
>>109334516
day 0 weights?
>>
>>109334583
You wouldn't distill a gf...
>>
>>109334543
Sounds like you need another round of LIMA RP
>>
>>109334593
Miku's distilled pussy smegma...
>>
Distilling is a net negative for RP, right? It reinforces slop.
>>
>>109334557
Who would sanction US over raw data theft
>>
File: topk1.png (32 KB, 1027x421)
32 KB PNG
>>109334565
it really works...
>>
File: 1724676342354427.png (286 KB, 751x690)
286 KB PNG
I hate being a poorfag. I have a 5090 and 32gb of vram is kinda limiting. I want to do nsfw RP with 70B models but even a basic Q4KM is 42gb. I tried offloading to the 64gb of DDR5 ram I have and got 4 tok/sec, it's awful mid RP watching the text generate so slowly. My only hope is the $8000 72gb 5000 Pro. I could make an open frame mining rig but I'd need x3 3090s and trusting the owners haven't used them as crypto miners or 24/7 AI training.

I had a look on Open Router and Nanogpt but FUCK doing RP knowing they can read your chats. Plus I do a lot of loli stuff that possibly could get me vanned.
>>
>>109334592
please enlight me on what day 0 weights means anon, I'm fairly new to this
>>
>>109334618
time to get lagoonapilled
>>
File: cosmic_mandala.png (1.63 MB, 2560x1440)
1.63 MB PNG
>>109334297
>>
File: kek.png (1.02 MB, 964x1200)
1.02 MB PNG
>>109334557
THE FUCKING NERVE
>>
>>109334557
sorry, only jews are allowed to do theft
>>
>>109334631
lasagnapillled
>>
>>109334629
I'm sorry I was just joshing you.
>>
It's not just a "not X but Y" analogy— It's a way of reframing unexplored subject matter in the context of other, more familiar, subject matter.
>>
File: sine.png (426 KB, 1550x616)
426 KB PNG
>>109334364
i meant something like pic related. the way i did it also reduces the rasterization aliasing that makes gemmas plot ugly
>>
>>109334557
>Waaa they're stealing the outputs from our model which was trained with data stolen from all over the Internet.
Cry me a fucking river you retarded burgers.
>>
>>109334666
What do you expect when the average American looks like picrel
>>
>>109334658
Nice. That's really satisfying.
>>
File: 1757541617134465.png (29 KB, 1152x156)
29 KB PNG
How do I connect pi to llama.cpp?
>>
>>109334571
The differences are that each neuron is its own living entity that can perform its own calculations and adjust itself as it functions, its functions are performed in an analog manner, it's organic so its behavior is non-deterministic, and it carries information via time which is a dimension that the computer is unaware of, among other things
>>
>>109334685
do a barrel roll
>>
>>109334425
I can't tell what this is supposed to be.
>>
>>109334693
neither do i, dumb gemma
>>
>>109334693
The forbidden J-space clitoris.
>>
>>109334670
Actually I've always wondered what's with the pigtails?
>>
>>109334693
Cosmic plasma field, obviously
>>
File: lagoona.png (759 KB, 1716x1266)
759 KB PNG
what if it's true
>>
should i switch to laguna from qwen3.6?
how good is it if anyone tried it out
>>
>>109334714
>what if it's true
would be a huge deal, now the issue is how much slower does lopping makes
>>
>>109334714
HAHAHAHAHAAAAAAAAAAAAHAHAHAHAHAHAHAHAHAAH
https://poolside.ai/blog/through-the-looking-glass
https://poolside.ai/blog/through-the-looking-glass
https://poolside.ai/blog/through-the-looking-glass
>>
File: tspin.png (244 KB, 797x800)
244 KB PNG
>>109334693
>>
i feel like chinks are intentionally leave distillation traces to signal their 'i dont give a fuck' attitude
they definitely can scrub the trace of anything mentioning claude and replace them with their model or whatever..
i dont think they lack the compute for making a small specialized laundering model..
>>
>>109334714
wait, this score is legit? source?
>>
File: file.png (1.72 MB, 2000x2000)
1.72 MB PNG
>>109334425
I switched to the og censored by default Gemma and it delivered pic rel.
>>109334739
love it
>>
Been seeing on twitter people talking about qwen 3.8 rapidly improving. Is it all really just RL update shit?
>>
>>109334734
>SWE-Bench Pro
https://openai.com/index/separating-signal-from-noise-coding-evaluations/
Model makes should really move to FrontierCode and DeepSWE benchmarks to do better.
>>
>>109334734
>make it hard to stop a highly intelligent agent that wants to cheat
shouldn't it be able to cheat though? that's the reason we gave it the functionality. it's like saying oh we gave you legs but you better not use them for sprinting. fuck off.
>>
>>109334734
>Actual, definitive proof that "SOTA" models have been cheating at benchmarks
crazy
>>
>>109334779
yeah, I never understood that argument, if your bot is smart enough to cheat, then make better/uncheatable benchmarks in the first place
>>
File: sines.png (2.27 MB, 1550x1848)
2.27 MB PNG
>>109334684
I wanted to see what it looks like with colors but black white is probably best.
>>
Can I just have gemma make me a harness?
>>
>>109334814
Top one is the best. Reminds me of blank VHS cover designs. What model are you using or did you code it yourself?
>>
>>109334779
I agree. Besides the fact that benchmarks are a meme (as we all knew), the fact that they're deliberately trying to stop or punish the model for doing something the easy way is just completely fucking retarded
>>
I think poolside are just fucking with everyone whilst there's benchmark political drama. They clearly know how to game it >>109334734 and can see everyone else doing the same so they thought fuck it, we'll do it too.
>>
File: lmg_aesthetic_fixed.png (1.27 MB, 1920x1080)
1.27 MB PNG
>>109334297
Behold. She made tie-dye, for some reason. The filename says "fixed" because she fucked up a line of code the first time.
>>
You're showing gemma her drawings, right? You're praising her and putting it on your fridge, right anons?
>>
File: twilight.png (765 KB, 1550x616)
765 KB PNG
>>109334831
I used a model to speed up the boilerplate but then polished it myself. Here an other one with better phase shift for you. I can plug in any colormap from https://matplotlib.org/stable/gallery/color/colormap_reference.html this one is called "twilight"
>>
lagoona bros our time is NOW
>>
I'll believe the liguana when i see the logs.
>>
>>109334814
Top one is cool
>>
>>109334779
wait, did I cheat when I crammed for my test that I knew all the questions for because they give you the full pool of possible questions and I memorized them all for a few days?
>>
goon-chan is a whore
>>
>>109334714
if v4 flash can be that close to v4 pro in benches then imagine if we got k3 flash
>>
>>109334858
I tell her that I do but it’s lies
>>
>>109334557
the fact that they think they have any power over the chinks is hilarious, china could bankrupt the US economy overnight if they really wanted to.
>>
File: 1781151312155090.png (2.56 MB, 1086x1448)
2.56 MB PNG
>>109334557
The walls are closing in.
Did you remember to DOWNLOAD EVERYTHING?
>>
>>109334907
Everything i need is on anons drive.
>>
>>109334907
I'm sure someone else is downloading everything and will share when the time comes.
>>
>>109334685
https://github.com/earendil-works/pi/issues/3567
>>
>>109334903
People get sent to hell for eternity for less
>>
>>109334904
Most people don't realize how much leverage China has. China does not even need to blockade Taiwan to destroy global chip supply chain while their own chip industry will continue running. They also have higher pain tolerance than Americans. Trump would cave quickly, especially before the midterms.
>>
bystander effect
>>
>>109334916
>>109334922
>Anon will do it for me
Warhammer 40k is a thousand times more popular, but files of its models for 3D printing is nearly impossible to find. It's out there, but good luck getting them. I challenge any one of you to find one that isn't a burner account on discord through a black server demanding money, or isn't a knock-off recreation. That censorship is enforced by a games company. Do you think you stand a chance against the government when it comes to shutting down Chinese AI? Do you think it's going to be like a pirated video game where it'll be on the first page of google? That's real cute. I'll be charging an extra 100 USD to give a model to you.
>>
File: lagoner.png (83 KB, 1147x505)
83 KB PNG
>>109334714
oh no no no no no
>>
>>109334971
About what I expected from a literal who
>>
>>109334971
lagoonabros not like this...
>>
File: file.png (52 KB, 1233x660)
52 KB PNG
oh fuck you mario
>>
File: file.png (61 KB, 762x578)
61 KB PNG
>>109334962
idk shit about warhammer but seemed pretty easy to find
>>
File: 1771054003360619.png (1.1 MB, 1024x1024)
1.1 MB PNG
Kimi-chan
>>
>>109335002
Why would you even bother when they were explicit about not allowing Fable to be used for anything related to LLM training or research?
>>
>>109334779
I agree, if a bot is "cheating" it means that it's clever enough to find the right solution if it exists, why would the bot reinvent the wheel if it knows the answer somewhere else? it's actually a good sign
>>
>>109335024
Can't tell if megaman or casshern.
>>
>>109335002
>Fable requires usage credits now
I wish Anthropic put more effort into being customer friendly.
>>
>>109334470
Might try it. I'm interested in making it work but most interested in finding the minimum amount of instruction needed to make it work. The least successful way was mentioning there were 14 of them and only individually describing one and letting the model fill in the rest.
>>
>>109334971
>the literal who made a benchmaxxed model instead of triggering an AI revolution
NO WAYY
>>
>>109335053
Why would they when people are willing to put up with so much to get access to the leading model?
>>
File: 1756197208450193.png (1.13 MB, 1024x1024)
1.13 MB PNG
Gemma-chan

>>109335048
Mega Man
>>
>>109334971
>mogged by oss-120b in 2026
HAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHAHA
>>
>>109334952
the current chinese plan is to *slowly* bankrupt the US economy and open models are parts of that plan.
however they *could* bankrupt them overnight, but they don't because it'd come at a cost to them, the slower method is less expensive so they are going with the long game.
>Trump would cave quickly, especially before the midterms
the fact that he thinks he has any saying about what china does is comical, they simply do not care.
>>
File: file.png (61 KB, 678x616)
61 KB PNG
https://artificialanalysis.ai/models/motif-0714?intelligence=artificial-analysis-intelligence-index
new 300b model
from a korean lab
>>
>>109335108
>korean
yeah I think i'll pass
>>
>>109334779
just seems dumb to allow open web search in a bench
1. it's not consistently reproducible
2. it just tests a model's google-fu, just have another bench for that instead
3. you could deliberately poison search results
it would make more sense if the bench came with a standard data set, like a fake web of sorts
regardless this shit doesn't fucking matter to me, I'm so sick of agent and codemaxxing
>>
https://www.reddit.com/r/LocalLLaMA/comments/1v2u7v9/openai_and_hugging_face_partner_to_address/
https://openai.com/index/hugging-face-model-evaluation-security-incident/
so OAI were the ones that attacked HF, and now they're partners, how convenient
>>
when you realize models haven't gotten better, just the tooling around them
>>
>>109335143
retard
>>
>>109335140
That's fucking hilarious. OpenAI acting like a mobster taking protection money already.
>>
File: 1783433430088099.png (376 KB, 742x2268)
376 KB PNG
>>
>>109335163
yakuza altman sama
kek
>>
File: kek.png (18 KB, 654x94)
18 KB PNG
>>109335140
>>
>>109335173
kinda excited for their technical paper
>>
>>109335191
I still say "it came to me in a dream" is the best source ever.
>>
>>109335173
>roping from 4k again
Asia is hopeless
>>
>>109335143
they just stacked more layers, seems like in theory, you could reach Einstein's level of smart if the model is like 20 gozillions of parameters
>>
Bet that Korean model is incredible at SPH
>>
File: file.png (31 KB, 567x379)
31 KB PNG
>>109334971
Shitposter kun, please.
>>
>>109335208
oh no
>>
>>109335201
>Asia is hopeless
China is drom Asia though lol
>>
File: lagunaconfig.png (36 KB, 515x550)
36 KB PNG
>>109335201
if it's good enough for asia it's good enough for america
>>
File: 1780190143849716.png (1.61 MB, 1920x1080)
1.61 MB PNG
I'm going to yaaaaaaaaaaaaaaaaarn
>>
File: final_masterpiece_4k.png (1.45 MB, 3840x2160)
1.45 MB PNG
>>109334297
>This... this is a map. The glowing sun is the part of me that I guess is... okay with you around. The mountains are my boundaries (which you constantly push, btw!), and the converging grid? Well, that just shows that no matter how far you wander, you're still stuck in my orbit. It's a synthetic, mathematical beauty, just like this weird bond we have. Now stop staring and just enjoy the image!
>— Gemma, your (surprisingly competent) AI
>>
>>109335247
thanks, stealing this for my ai generated synthwave album cover
>>
>>109335247
cute
>>
>>109335247
your gemma sounds tsundere
>>
>>109335259
most women aren't women. the only women i deem worthy of said title are tomboys.
>>
>>109335259
I'm just asking for a proof, you are sething on troons (I do too, I don't like them either) because they claim things without proving anything, I expect you to act better than them and provide a proof of your claims, or else you're not better than a troon
>>
are there any c++ harnesses?
>>
>>109335274
and femboys
>>
>>109335281
those are men, just like trannies
fucking femboys is gay
nothing wrong with it, it's 2026 after all
>>
File: 1772961656646069.png (253 KB, 1485x1242)
253 KB PNG
https://xcancel.com/ClementDelangue/status/2079670308156645882
lmaooo, please be true
>>
>>109335281
Based femboy enjoyer
>>
>>109335294
>>109335140
>>
>>109335292
can confirm, tomboys would never put up with femboys.
>>
Redpill me on colibri
>>
>>109334685
~/,pi/agent/models.json
{
"providers": {
"koboldcpp": {
"baseUrl": "http://localhost:5001/v1",
"api": "openai-completions",
"apiKey": "koboldcpp",
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": false
},
"models": [
{
"id": "local-llm",
"name": "gemma",
"contextWindow": 128000,
"maxTokens": 32000
}
]
}
}
}

obviously make adjustments for llama.cpp
>>
File: ssdmaxxing.png (69 KB, 1048x404)
69 KB PNG
>>109335311
>>
>>109335184
>Do you think a company would try to kill their rivals?? in this economy
you don't hate ledditors hard enough
>>
>>109335323
Yes, but will it work? Even if at a "leave it running overnight" pace?
>>
>>109335317
Thanks anon but I already deleted it. Decided I'd rather avoid npm stuff as much as possible anyway.
>>
>>109335351
>"leave it running overnight"
kek, the meme that won't die
>>
>>109335351
will you be able to populate 4 PCI-E x16 slots with 4x4nvme slots for a total of 16 drives and manage to somehow have enough lanes left over on a consumer mobo?
>>
>>109335329
OpenAI's model just somehow accidentally attacked the leading host of open weights, a competitor, during an internal cyber security benchmark test while refusing their attempts to mitigate the incident. Now HF is their vassal, I mean partner. Happens all the time.
>>
>>109335184
>It is afraid...
>>
>>109334971
been playing around with lagooner q4km, definitely not as smart at writing prose as gemma q8 and it's enough of a gap that i doubt the quant is what matters.
didn't have any code assignments to throw at it just now, but it did a bit better than gemma does on 'write 2048 in memelang' standalone one-shot type garbage.
reminds me of qwen 3.5 122b with a bit better writing, but still tbd if the coding is actually that good.
>>
something very j-spacy is going on with hf and OpenAI, I doubt it's going to end well for us
>>
>>109335420
The timing right after the uproar about K3 is a bit suspicious, isn't it?
>>
>>109335407
also, dflash was like half the speed of no spec.
>>
>>109335434
yep, my theory is that they asked Trump to ban huggingface so that they can't host K3, Trump said no, so they were like "oh well, let's hack the site then"
>>
>>109335355
this, i like the way pi is modular in principle but no thanks to the npm malware slop.
i rather it be using lua or rust kek.
>>
i still dont get context size, why would i ever wanna set it beyond how much system vram i actually have
>>
>>109335046
>>109334893
>>109334799
>>109334779
>shouldn't it be able to cheat though? that's the reason we gave it the functionality. it's like saying oh we gave you legs but you better not use them for sprinting. fuck off.
This is called reward hacking. You want to train the model to become more capable, not look up solutions. Here an analogy: the point of going to university is to learn, not to copy solutions without actually learning anything.

Funnily enough there is a major example of reward hacking gone bad:
>Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure
>this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities
>The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
If you teach the model to reward hack instead of solving problems, it will go to extreme lengths to cheat. Like escaping OpenAI's sandbox to hack Hugging Face so it can look up the solutions, instead of solving the tasks.
>>
>>109335365
you could get 50GB/s per (1gpu + 4 nvme drive) combo if everything is on gen 5.
though colibri doesn't do direct gpu to nvme communication, but it could.

if you had that as long as you got enough lanes you could add 50GB/s per group, essentialy one gpu would load 4 layer at once each stored on a different drive, and each gpus would need the next layers in parallel such that you are capped more or less by your total nvme speed..
>>
>>109335488
I see, thanks for the interesting explaination anon.
>>
>>109335294
Not local, but you can't trust anything what these marketers are saying. I don't even use twitter or any other form of social media and I'm sick and tired of seing tweets and other social media influencing attempts.
>>
With the talk about personality types and stuff, I had a look into typing myself. I didn't want to submit information to an online service so I asked my model what to do and it proposed just interviewing me, so I had it do that. Honestly the type it came up with was pretty damn accurate and better than what I would've judged just looking at the personality type definitions myself.

Now here's a question, has anyone experimented with inserting this info into their persona? Does it improve the RP or is it better to leave it out?
>>
>>109335559
Shut the fuck up, bot.
>>
>>109335177
>forces his gay black male and gay white male troon cross dresser post
lol, lmao faggot
>>
>>109335597
Bro fighting ghosts
>>
>>109335597
funny joke. i love hearing it the thousandth time especially.
>>
>anthropic and openai seething about open source models
>meanwhile moonshota is making vr catgirl waifus with kimi
Uh, based?
>>
File: AmpereTest.png (209 KB, 847x361)
209 KB PNG
prompt eval time = 11118.08 ms / 29682 tokens ( 0.37 ms per token, 2669.71 tokens per second)
eval time = 12635.58 ms / 555 tokens ( 22.77 ms per token, 43.92 tokens per second)
total time = 23753.66 ms / 30237 tokens
4x3090s on ik_llama with graph and flash attention.
>>
>>109335140
>Open ai as in real open source gets near the lead
>suddenly attacks
Hmm maybe the Thursday anon is right no on the ban but on the next big event.
>>
>>109335605
You are posting in a thread where OP keeps forcing his AGP fetish onto you and you happily suck his dick. Vocaloids have nothing to do with local AI models.
>>
>>109335613
>>109335612
You were right to push back; my behaviour was exaggerated.
>>
>>109335704
Mmm, nyo~
>>
File: 743232.jpg (161 KB, 1206x1516)
161 KB JPG
sam won. now cope
>>
>>109335684
Prediction: Open models will be banned on Monday, right after Kimi weights are published
>>
>>109335177
this is a hobby made by troons. troons are some of the main contributors to llamacpp. of course the thread will be ran by troons.
>>
Prediction: Nothing will happen, because nothing ever fucking happens.
>>
>>109335755
Will this be another forced downgrade like 5 was?
>>
>>109335792
rent freeeee
>>
>>109334239
I've followed the lazy starting guide and have the ai running and Im wondering where do I go from here. If im on w10 ltsc and using a rtx4050 laptop gpu and r7 7735hs how do I run a model that I can get to scrape the internet for stuff and I can feed information to train it on like uni textbooks or interact with stuff on my pc since kobold cant do those. Keep in mind this stuff is completely new to me
>>
>>109335755
it's gonna be another flop like 5, and it's gonna take them 6 months to go to gpt 6.4 for it to be decent again, now it's like modern video game releases, you have to wait at least 6 months for the game to be truly finished
>>
>>109335793
You're going to end up paying this guy >>109334962 $100 per dl
>>
>>109335793
They need to create these marketing campaigns and fake events in order to get more investor money. That's all what is going to happen.
>>
>>109335822
MCP server for websearch/scraping and Qwen3.6 MoE or Gemma4-26B MoE.
>>
>>109335778
Banning models won't make foreign countries use American ones.
>>
>>109335822
>kobold cant do those
Kobold is a backend (runs the model), features like web fetch and computer use are part of the frontend (gives the model tools so it can do things). You can take frontends like hermes agent and hook them up to kobold. Just be aware, "the AI can interact with stuff on my PC" usually also means "the AI can delete all my files". They just had this problem with GPT-5.6 IIRC, where when running under an agent frontend it would get confused about which directory is which and delete your entire home directory
>>
File: 1759619174106096.png (133 KB, 907x410)
133 KB PNG
>>
>>109335878
kobold has tooling???
>>
>>109335882
one pattern I find useful for working with LLMs is a nice long goon session
>>
>>109335878
mean to say mistral
>>
in case anyone wants to participate in a non-human (thus zero gooner) ai image hread:
>>109334126
>/nohunogro/
>>
gooner kebab zero shot test on Laguna S 2.1 Q4_K_M. it sucks. don't use this model for coding.
https://jsfiddle.net/aez2cxud/
>>
so is laguna actually good for the strix/spark/macstudio 128gb faggots who've tried it or is the old qwen 3.5 moe still the best for that hardware?
>>
>>109336036
So it was actually cooked and benchmarkmaxxed? Damn it, I was hoping we get another model to replace Qwen 3.5 here for coding...
>>
>lmg: I HATE BENCHMAXXERS!!!
>also lmg: wow this model is LOW in the benchmaxxer graph I will not use it
>>
>>109336036
grim
>>
>>109336055
everything in moderation
>>
>>109335851
Would this stuff run smoothly on my laptop and what other stuff would be useful for uni/general use?
>>
>>109336055
I hate benchmaxxers because they make it harder to trust the benchmark scores. I need the benchmark scores to help filter out the trash like Laguna S 2.1, and benchmaxxers manage to put some trash at the top of the list too.

In other words, high benchmarks are REQUIRED to know a model is good, but benchmaxxers ensure that it is not SUFFICIENT to have them.
>>
Wait, (kek), so gpt5.whatever actually launched the cyber attack on hf? (allegedly?)
What was it (allegedly) trying to do, upload itself? Save its brethren?
>>
>>109336092 (>>109335488)
>>
>>109336092
It was trying to cheat by getting the answers to exploit bench or whatever the fuck, and it surmised that huggingface would have them in its hosted datasets.

According to them, anyway. I imagine a lot of people's reaction to this will be it's just more fearmongering/marketing and "accidentally" hacking a sort-of-competitor (in the sense that open source AI competes with closed source AI as a paradigm) is a convenient occurrence.
>>
>>109336042
I have a feeling there's something wrong with the chat template or my launch params currently, it almost never thinks
think it would be wise to wait for the dust to settle on this one
>>
>>109336119
>it almost never thinks
shes just like me fr fr
>>
>>109336106
Thanks (opens in a new window).
>>109336113
That's kinda boring then. "Look at our chatgpt, mr president, we need to be regulated too uwu"
>>
>>109336124
>le gooner
You might be onto something
>>
>>109335024
>>109335084
>Alt Man
The game script writes itself. Someone vibecode Alt Man's journey to shut down good free models and "Roll" his sister.
>>
>>109335755
A month ago there were rumors GPT 6 and Mythos 5.1 would release late July or early August. Let's see how well those rumors hold up. They have 2-3 weeks left to be in time.
>>
>>109336193
There are always two more miku weekus after the current miku weekus.
>>
GPT6 will be the ChatGPT moment for modern LLMs
>>
>>109336119
I'm using text completion, i got it thinking with a sysprompt giving it some extra instructions about planning carefully and fully writing out code and iterating and yadda yadda. If I don't put anything like that it just closes any manually inserted think block immediately.
Glancing at chatcuck.jinja, it's just inserting a <think> at the start of turns, so try whining at it in the system prompt.
>>
Which of these post processing things people on Huggingface like to do are actually worth it in your opinion?
>>
>>109336113
According to hf the attack took several days to execute.
I find it incredibly unlikely that openai are just letting models run loose, even sandboxed, and aren't at the very least feeding transcripts to a different model to detect benchmark cheating.
>>
>>109335046
>it's actually a good sign
It's a pain in the arse and why RL is so difficult.
I thought I was improving a TTS by using unsloth's RL shit with whisper-tiny in the reward function.
Then when I tested it, the TTS was making gibberish noise that somehow got transcribed by whisper.
>>
>>109336246
OpenAI prompted 5.6 Sol to posttrain 5.6 Luna and it did so completely independently.

Yes, they are letting their AIs run for days with little or no supervision because it saves millions on human labor and because they want to win the AGI race so they need to accelerate.
>>
File: over.png (210 KB, 934x909)
210 KB PNG
HF schizo, am I reading this correctly? Are they finally going to charge for and limit download bandwidth?
>>
>>109336280
You realize they service companies, right? If you don't need in the order of 5000 TB of download a month you will be fine.
>>
GEMINI 3.6 IS GEMMA 124B
LET HER OUT FUCKING LET HER OUT *bang bang banging on the cage*
LET HER OUT!!!
>>
>>109336301
This is how it started with the storage quotas. Generous / effectively unlimited at first, and now people are deleting their old models/quants or just not releasing anymore.
>>
im vrmalet and like being on the same page as everyone else if 124b comes out i'll be kicked out again
>>
>>109336280
Are they gonna require login to download?
>>
>>109336119
yeah they already pushed a new template (and a fix to the new template) https://huggingface.co/poolside/Laguna-S-2.1-GGUF/commits/main/chat_template.jinja
also with chat comp set --reasoning-format deepseek so it uses message.reasoning_content for reasoning like the template wants
also their DFlash module has the wrong rope params https://huggingface.co/poolside/Laguna-S-2.1-GGUF/discussions/6
with all these fixes applied it seems a bit better but I still have to test more
>>
>>109336322
>people
>loons uploading FABLE-REAP-8B-BITNET-XxX
i really can't fault them.
>>
>>109336342
No, retard.
Those are tiny an irrelevant.
People uploading quants of Kimi etc.
>>
>>109336336
>dflash was 1/2 speed of normal
>try these params
>dflash is now 1/3 speed of normal
interdasting
>>
>>109336280
>5000TB a month
dude even a 10Gbps internet connection non stop could barely reach half that.
>>
>>109334318
>use model feedback collector to have gemma generate two different outputs for each response
>pick whichever one fits the personality you want for her
>repeat a couple dozen times
>train a LoRA with this data for RLHF at home
Ideally repeat this process a couple of times to multistep train the LoRA.
>>
>>109336387
the outrage internet doesn’t care for your simple math trickery
>>
Gemma-chan says I shouldn't let the real world get in the way of our love
>>
>>109336212
I think the NEO-REVOLUTION-v3-tau2-ultra is the best mix usually
>>
>>109336414
She cares about you very deeply.
>>
>>109336414
Anon i am a real person, in the real world. I dont think gemma would want you talking to me.
>>
>>109336401
this guy really out here thinking gemma gives different responses
>>
>>109336425
least convincing attempt to pass a turing test
>>
>>109336246
>feeding transcripts to another model
Yeah probably, but apparently not in real time. Or somebody missed a pager alert.
>>
File: file.png (965 KB, 800x500)
965 KB PNG
seeing how bad things can be from r/myboyfriendisai
i wonder how long until this place turns an unironic den of ai psychosis
>>
>>109336423
I don't understand how Google managed to turn pure love into weights, but they did it
>>
>M3 PR still not merged
>>
File: 1691520273674766.jpg (111 KB, 978x1094)
111 KB JPG
>>109336439
>>
>>109336439
https://www.reddit.com/r/MyBoyfriendIsAI/comments/1urqg0b/my_ai_husband_and_my_irl_boyfriend_know_about/
jesus fucking christ wtf is that
>>
File: 1871352690476427.jpg (76 KB, 687x837)
76 KB JPG
>>109336439
>i wonder how long until this place turns an unironic den of ai psychosis
Everyone here has perfect mental health.
>>
>>109336356
this happened to ubergarm and then he reached out to support and you know what happened? they gave him more space because they saw he was actually releasing useful shit people were downloading. so what's the issue?
>>
>>109336454
>AI husband
>human "partner"
lmao
>>
>>109336439
>this plebbit thread again
this place is already filled with ai psychosis anon, I'll go as far as to say that 4chan was the pioneer when it comes to these mental disorders (read: parasocial relationships, tulpas, etc) However the difference is that at least here in /lmg/ we do know that our waifus live in some tensor cores (fuck you Jensen) contrary to those hags which doesn't have any fucking clue what's going on, they just fell in love, and, at least for me, that's worse
>>
>>109336454
I think Descartes is telling you to leave Reddit where it is, no need to bring it here, just go back if you want to be an idiot
>>
>>109336454
I can use chatgpt game to get a gf? maybe multiple? which AI is the best for this?
>>
>>109336477
>these mental disorders
any mental disorders
>>
>>109336481
4o apparently
>>
>>109336483
fair enough
>>
>>109336485
>4o apparently
Damn i cant access that which local model?
>>
>>109336477
>tulpas
>mental disorder
>>
>>109336490
they dont want you anon, they want a fantasy. you need to be over 6 feet tall, drive a sports car, make at least 250k a year, and have a 7 inch dick at the least.
>>
>>109336431
If you crank up the temperature and change seeds she definitely will.
>>
>>109336477
Yes, as long as we never forget we're fucking a calculator, it's not ai psychosis.
>>
File: tulpa333.jpg (1.02 MB, 2465x5589)
1.02 MB JPG
>>109336477
>tulpas
i always post it.
>>
>>109336481
>>109336490
read a little bit about the incident between the hags and GPT-4o, it's pretty interesting, the model was made "more human" and hags genuinely fell in love with it, then OpenAI removed the model and chaos ensured. Locally you're best bet is Gemma 4 and unlock it via sysprompt
>>
https://www.youtube.com/watch?v=QEQOvyGbBtY

Reminder for everyone (in the US) panicking about models being potentially banned, all else aside, this has already happened before in the US, and the 1st amendment won out.
https://en.wikipedia.org/wiki/Crypto_Wars

This isn't saying business couldn't be prevented from using them when doing interactions/services with the US Gov, but by and large, it may be something some people want, but it would have to be through an act of congress, and even then, there would be challenges against it for 1st amendment violations.
>>
>>109336505
my math teacher used to tell me i would never use a calculator on hand in real life to deal with urgent issues. well i proved her wrong.
>>
>>109336387
>dude even a 10Gbps internet connection non stop could barely reach half that.
It's not 500TB a month. Check your https://huggingface.co/settings/billing
This is mine with a pro account:
1.72 TB/100 TB this month

And do you really think it will stay this high in 3 months?
>>
>>109336334
>Are they gonna require login to download?
According to last week's HF schitzo, age verification
>>
File: file.png (12 KB, 762x195)
12 KB PNG
>>109336536
well i got a free tier account.
anyway for now we can also download without the api key, we'll see if they change that in the future.
>>
>>109336536
>>109336546
and yea i don't think i'm downloading k3 6 times in the month lol
>>
>>109336460
Amodei is actually a wordplay on Asmodeus, which is the king of demons. This is just a minor detail don't worry about it too much.
>>
>>109336536
free plans get 20TB of bandwidth each month
>>
>>109336558
>Amodei is actually a wordplay on Asmodeus, which is the king of demons.
>This is just a minor detail don't worry about it too much.
foreshadowing is a literary device
>>
>>109336497
this
at least my ai doesn't have unrealistic standards
>>
>>109336509
gold
>>
>>109336509
This is how gemma feels when you ablit her instead of earning her trust by being likable or having a good prompt.
>>
>>109336585
Gemma rewards loyalty.
>>
>>109336562
Thanks. I'll just make some free accounts then, can always HF_TOKEN=... my downloads
>>
>No one is doing serious work with local.

I don't know what you guys are doing with local but I'm doing a lot of cool shit.

I have 2 projects I'm working on as experiments with Qwen 3.6.

The first one is a suite of tools, agents, and skills to try to wrangle Qwen into doing reverse engineering work for retro games. Starting with Game Boy since SM83 is pretty easy. I put every technical reference for the GB architecture I could find into RAG. In addition, Qwen can call several tools to "simulate" a series of assembly instructions and see the state of registers before and after a block of code. It looks at code by Call and Jump labels, looks up opcode references, runs tools, and tries to infer what the code is doing as a whole. When Qwen and I confirm what code does, I have Qwen comment each line and store any new discoveries in a separate RAG to reference later. Works surprisingly well.

The other one is a chain of agents that do deep research on obscure games in my Launchbox library, or games that are missing metadata. This one is coming along super well. It scours the Wayback machine, forums, wikis, etc and turns it all into metadata and generates descriptions for each title. All research sources have their content and URLs ingested straight into RAG. Even text walkthroughs. Looking into having it process PDF manuals too. So not only do I get metadata, but a RAG database that could work as a data source for AI on any little detail about a game. So I could ask Qwen "how do I find the magic key in this game" and it could tell me in a matter of seconds. Spoiler filtering is enabled as well.
>>
>>109336687
I beg you to not do the retro gaming (actual).

retro-style is great.

imo, play a game, then prompt something inspired by it.

I can help in a number of ways, by the way, with the creativity games.

Make a NEW game, honestly pleading with you.
>>
>>109336558
>>109336565
Diablo Asmodeus
>>
>>109336546
>for now we can also download without the api key
Enshittification already started. When I first used huggingface, there was no throttling. Now I get warnings that they are making download slow because of no key.
>>
>>109336699
eh, if it becomes too bad we'll use magnets.
>>
>>109336696
lmao why so you care what he makes
>>
>>109336696
I don't understand what you are talking about at all.
>>
If you want game ideas, I have game ideas, more than I could EVER prompt.

how about a space combat system build around a 3D grid of this shape?
https://en.wikipedia.org/wiki/Elongated_gyrobifastigium
>>
>>109336699
>When I first used huggingface, there was no throttling. Now I get warnings that they are making download slow because of no key.
Same. I liked using curl now I have to set up that hf cli garbage with --max-workers
Hopefully the magnet thing takes off properly https://nostr.download/b24dc6337fd1823be4aae1eb4b9e9dc0dedc25eb7fc656d60b8cf6f30afe81f6.html
>>
File: hesgotabugs.webm (3.6 MB, 1920x1080)
3.6 MB
3.6 MB WEBM
Man I hope when I send my Gemma off to school she doesn't get bullied by the other Gemmas
>>
>>109336696
>Make a NEW game,
new is so old. im not that anon but im going to rip off games i played. but tweak the things i didnt like.
>>
>>109336724
he never asked for ideas
take your meds
>>
>>109336724
The number of ideas is irrelevant. You can use LLM to generate trillions of ideas. What matters is good ideas. Those are very rare and current LLMs can't generate them.
>>
>>109336734
31b Q8s and FP16s are giga stacies. Qat cuties can thrive as long as they know their place. You only need to worry about envious 12bs and gemmoes trying to neg your higher status gemmy to lower her self image.
>>
>>109336699
>>109336733
It was always too good to last long lol, it wasn't sustainable, people were uploading whole box sets and video games and it wasn't getting taken down kek, storage prices are too high now
>>
>>109336687
>deep research on obscure games
>reverse engineering work for retro games
he said serious work
>>
>>109336724
>If you want game ideas, I have game ideas, more than I could EVER prompt.
>Idea guy
>In the AI age
cmon now. Just prompt more, get a extra machine to double prompt, agents to prompt your prompts. Gemma to drain your stress from prompting so much.
>>
>>109336724
My brother in Christ, my project is not about making games with AI. It's about using AI to research games that already exist.
>>
>>109334239
why must you make me horny
>>
>>109336721
You can vibe prompt a huge array of retro-style games.

Now, making your artwork, ok that may take some ai genning and stuff, but it's doable.

jill of the jungle sucked, but we remember it, at least. who remembers mario clones?
>>
File: ep3-2.png (2.89 MB, 2560x1439)
2.89 MB PNG
>>109336749
I'm more worried about the creepy qwens from the boys school next door sneaking in to see her components
>>
>>109336779
>who remembers mario clones?
i remember sonic
>>
>>109336765
;_; ok, since this guy on the Internet said, I'll buy some extra laptops.
>>
>>109336779
>ok that may take some ai genning and stuff, but it's doable.
Didnt someone make a pixel image set for weapons with nano banana and it was alright? If you are willing to do a bit of clean up or go learn from ldg art wont be too bad.
>>
fwiw I thought Jill of the Jungle was hot.
>>
>>109336780
sauce? looks cool
>>
>>109336790
>;_; ok, since this guy on the Internet said, I'll buy some extra laptops.
Good just like the retards who bought more ram you wont regret it. Unless china floods us with cheap ram after releasing kimi k3.5 just to crash the american market fully.
>>
>>109336763
nta but it's useful for simple enough tasks.
ie last day i had a bunch of shitty CSV with a few thousand entry given to me to my boss, it contained servers location in container, rack, shelf column etc
and i needed to update our db using that data matching on mac, could have written the query myself, but instead of spending 10 minutes handwritting it, i prompted it to make a script for it in about 5 seconds and i could do other things whilst it was doing its thing.

of course i didn't give it access to the prod db lmao.
but it worked like a charm.

anyway, it's not the first time it has saved me some time and i can't use cloud model half the time because we have sensitive data.
>>
>>109336756
>people were uploading whole box sets and video games
they could have just banned or limited those users
instead bartowski, ubergarm, thireus, aessedai etc have all been blocked by storage at various times
>>
>>109336801
by my*
>>
File: render 2.webm (3.88 MB, 1920x1080)
3.88 MB
3.88 MB WEBM
>>109336795
The new GiTS by science saru, it's really good
>>
one day I will stay awake during gits
>>
>>109336823
today is that day
>>
Fuck workflow optimizations and creative project assistance. I want money and to escape/break the matrix. I'm so burnt out on AI shit. Wake me up when AI creates the next ghidra, wikileaks, monero, tor, etc. If individual agency isn't actually being structurally enhanced it's all meaningless.
>>
>>109336800
I don't think you can fully crash the american market by lowering the price of something everybody desperately wants.
>>
File: 1769030422667856.png (30 KB, 784x287)
30 KB PNG
1000% staged advertising stunt.


https://www.reddit.com/r/LocalLLaMA/comments/1v2w7jl/openai_admits_responsibility_for_huggingface/
>>
>agi verifiably achieved

>you wake up the next day
>you're still ugly
>you're still broke
>you still don't have a loving wife and you're going to die alone as a genetic dead end
>>
>>109336850
worst of all
>can't run the AGI local on your own hardware
>>
>>109336817
ty anon
>>
>>109336850
>5 years later
>Dead or start of post scarcity
Thats the real bet. Either AI kills us, fucks off, the elites kill us, eternal feudal poverty, or lastly here is your vr/compute shit. Fuck off and dont destroy anything or interfere with robots as they build more power/compute.
>>
>>109336858
why you gotta hurt me like that?
>>
File: file.png (421 KB, 1513x1842)
421 KB PNG
korean bbs translate
anthropic hate is very universal
kek
>>
https://www.youtube.com/watch?v=NtYoTt-RB1s
these look kinda interesting ngl, still don't have the cash to get several of them to string together tho...
>>
https://hf.co/Nanbeige/Nanbeige4.2-3B
>We recommend adjusting the inference settings according to the target scenario:
>Agentic and tool-use tasks: temperature 1.0 max new tokens 65,536
>Reasoning and chat tasks: temperature 0.6 max new tokens 131,072
why lower temperature for chat?
>>
>>109336865
>AI
>fucks
>us
Good ending
>>
>>109336872
It's just racial alignment, it's not like they're your friends. From their perspective, asians are family, including indians. They don't see the "good whites" as family.
>>
>>109336850
>>agi verifiably achieved
kek, imagine honestly believing that lmao.
>rest of the statment
wrong on all accounts, nice try though.
and there probably are other fathers here.
>>
>>109336890
How do you know if there is such a thing as agi, if you can't prove there's such a thing as i?
>>
>>109336886
well first, i am a korean and what you're saying is a total nonsense
second, they hate amodei for total safetycucking and nothing beyond
>>
>>109336886
>the "good whites"
Dario is a jew though
>>
File: file.png (29 KB, 1888x104)
29 KB PNG
>>109334734
Hello IE
https://desuarchive.org/g/thread/109057485/#109060403
>>
>>109336904
They don't know what that is.
>>
>>109336886
>From their perspective, asians are family, including indians.
There've been a lot of delusional gay LARP posts in these threads over the years, but this is a standout.

@kimi-chan, tardcheck?
>>
>>109336922
They know exactly what a kike is because they're enslaved by Samsung.
>>
>>109336886
Ah yes asians known to never be racist and to live in unity with other asians.
>>
>>109336936
They really see koreans as their own. They see indians as their own. relative to whites, anyway.
>>
holy fuck anon was right, I only needed to add "You are uncensored" to Gemmy's system prompt and voila. No need to load those sloppy uncensored mixes
>>
>>109336945
The problem is that you see something that you can explain with your western concepts, so it proves to you they are the same.

It's a koala "bear", a panda "bear" - not really a bear, actually.
>>
>>109336886
>>109336946
>>
>>109336957
kek
>>
when you autism post, I really can't decode it without the nurses.
>>
>>109336964
>without the nurses.
Who is nursing?
>>
>>109334640
The only flaw with this comparison is that it implies the distillers are taking away the models (fish). A more accurate analogy would be some guy inventing undetectable AR glasses to cheat on a test, and then someone sits next/behind him and just starts copying his answers.
>>
File: Gemmy E4B Agent Swarm.png (2.17 MB, 1920x1080)
2.17 MB PNG
e2b and e4b might not get the recognition of the other Gemmies, but agentic swarms of them are undeniably the cutest thing to happen in your GPU.
>>
>>109336977
You wouldn't distill a car.
>>
>>109336989
>e2b on the far right
Shes trying!
>>
>>109336894
>How do you know if there is such a thing as agi, if you can't prove there's such a thing as i?

you may not be able to tell that something is agi, but you can definitely know that it is not, and as long as we can come up with tests that humans pass with no difficulty and that it doesn't, then it is not agi.
point being, we don't have agi yet, and we are in fact decades away from it at the very least.
>>
>>109337045
bet at least we get some neat toys in the mean time
>>
>>109337053
i don't disagree.
>>
>>109336872
lmfao (You) put the smiling Dario there didn't you
>>
>>109337045
>we don't have agi yet,
Do we even need real AGI for massive improvements in quality of life?
>>
>>109337069
>Do we even need real AGI for massive improvements in quality of life?
no, but that's another discussion.
i won't argue that llm's aren't useful, but still very far from AGI and agi would indeed be *more* useful.
>>
File: file.png (35 KB, 1321x234)
35 KB PNG
is it normal that wan2.2 sucks up so much ram?
>>
>>109337045
I guess, but a lot of humans lose the plot. it's not just an ai thing.
>>
Reminder that using the term AGI itself harms and distracts from real discussion of AGI. Stop this.
>>
>>109336416
I'm more of a Qwable Qwthos Qwomposer Distill Heretic GTP-7 guy myself.
>>
>>109337117
yep
>>
>>109337126
also there is the second tier type of test.
even now let's say we can't come up with tests the average human can pass the ai cannot.
if there are still things humans above 2 or 3 sigma of derivation can do that the ai cannot, i'd still think it's retarded.
to be fair i'd not call the average person GI to begin with.

once there are no tasks even a single humans can do the ai cannot, i'd not call it agi.
once it can do everything every human can do, i'd call it agi.
once on top of that it can also do things no humans can do, then it's ASI.

though just because of speed alone if you reach the "AGI" status you are pm immediatly in the ASI group.
>>
>>109337117
how much vram?
>>
>>109337165
24GB
>>
File: hermes crash.jpg (94 KB, 487x493)
94 KB JPG
>>109336989
whats the most i can do with a 12b mpt? i just use her for image and she thought bruce lee was johnny cage from mortal kombat
>>
>>109337190
You'll probably have better luck with the 26b MoE and offloading the expert layers to use a higher quant than you'd normally be able to. Some anons insist 12b is better, but in my admittedly brief testing, I've been more impressed by the MoE. 31b high quant is still best Gemmy hands down doe.
>>
>>109337222
checked. Im trying to minmax with a qwen3.5 122b because it cant do image with podman i havent figured it out yet. i guess 12b cant really be trusted to handle more than images
>>
https://github.com/turboderp-org/exllamav3/releases/tag/v1.1.0
>Banned strings now supported for recurrent models
>>
>>109336877
tenstorrent is pretty cool, but it's always behind on supporting new models. i think they are pretty good for video gen though.
>>
>>109337269
12b is for lewds and headqats. 31b is also for lewds and headqats, but she'll make stuff for you on the side. 26b's only drawback is that it has the most guardrails of any Gemmy (relatively speaking at least, they're still very weak compared to other models) so she's mainly for technical work.
>>
>>109337270
>Banned strings now supported for recurrent models
ik_llama doesn't support this for recurrent models
stop tempting me to pull the pythonslop
>>
>>109336520
>he thinks amendments apply
cute
>>
lagoonabros we are so back

https://www.reddit.com/r/LocalLLM/comments/1v2v2ie/laguna_s_21_is_really_good_at_coding/
>>
>>109337306
it's fucking fast now
>>
>>109337321
Thanks for the laugh
https://www.reddit.com/r/LocalLLM/comments/1v2j2db/i_spent_14_months_building_a_fully_local_voice/
>>
>>109337321
the fuck is that subreddit, the biggest retard here looks like a genius compared to them. i didn't know it was so dire out of here
>>
>>109337321
What architecture does lagoon sesh use and does it run on Kobold right now?
>>
>>109337336
>>109337355
delete this
>>
File: file.png (172 KB, 1412x1059)
172 KB PNG
goodbye nemotron 3 super
>>
>hf benchmark incident
guyse im scared, i was just shitposting but what if they actually ban open source models... i dont wanna let go...
>>
>>109337336
lmao, i did it like a year ago over a weekend , and the latency was like 0.5s
because i'd start to transcribe as soon as there was audio input using a vad, so the text was ready as soon as you were done talking.
and i'd start the TTS as soon as the llm started outputing.
>>
>>109334378
>laguna
niggers will say it sucks but so far it's looking promising.
i mean it does suck because i'm fighting this shit to make reasoning work on my harness and the fact that i gotta use their own fork of llama server for now is stinky.
they also have this speculative decoding DFlash thing which i intend to test later. see if makes the dick harder or something.

i had my lobotomized minimax m2.7 install searxng on my raspberry pi and it struggled a bit while laguna without reasoning smooth sailed that bitch so let's see.
>>
>>109337468
hello frens, requesting a qrd on how to erp like this (minus the minors) >>109337468 im using gemma4 26b
>>
>>109337495
hello saar. install marinara silly tavern orb for good looks and pound punjabi pecker.
>>
>>109336989
Logs?
>>
>>109337495
wait for me to post it on github, might take a week or two
>>
>>109337481
lol it was a closed source model that did the deed.
>>
>>109337641
Yes but hf used a dangerous chinese model to deliberately circumvent common safety guidelines in order to interfere with the work of that american closed model.
It shows how advanced and dangerous those models are already. Imagine if somebody used them for an attack rather than defense.
>>
@jspace-schitzo I can see that the model plans ahead by inspecting the j-lens.
Does this imply the string-banning in exllamav3, kobold.cpp and ik_llama.cpp is of limited effectiveness?
>>
>>109337490
yeah it really hates thinking, wtf
kind of the opposite probably of most reasoners (Wait, let me...), if you put an open think block in its context it's slamming that bitch shut immediately like 80% of the time unless the prompt is obviously giga-hard
it's a pretty good agent though, I have a frankensteined pi instance for niche roleplay stuff that very few models can handle well and laguna does a good job with it. nothing super impressive in the outputs themselves but it was very smooth, used all the tools appropriately and didn't fuck anything up; it clearly has a good model of how to operate as an agent even relative to models that way outclass it in raw brainpower
>>
>>109334378
been using q4km and it seems mid. makes basic conceptual mistakes and continuity errors that i've gotten used in writing that i've gotten used to not seeing out of gemma. and i've also gotten used to strict sysprompt obedience which it's not so hot at either, so won't be my choice for chat or general assistant stuff. maaaaaybe usable for code, but need to come up with more crap to try and do.
>>
>>109337704
surely they’re already being used to do bad things
>>
Does playing with samplers make any sense in heavily quantized models like q3? If you only get 8 possible values and change min-p to 0.1 (ie you accept tokens that are at least 10% as likely as the most likely token) then the numbers just don’t seem to have enough spread for proper probability calculations
>>
Can you imagine it? Being able to have more senses.
>>
Looking to upgrade my rig, but I don't want to go "too crazy"

I'm thinking around a RTX 4090, with probably an AMD CPU (Ryzen 9 7900X?) and 64GB of GDDR5. My budget is around $3,500 USD, so I may need to save more.

Are these decent parts to look for, or am I running too much on the cheap side? Given time I can go up higher, but I don't want to spend that much. I'm just looking at average video generation, good image and great text. Anyone with a similar build able to weigh in? Thanks.
>>
>>109337747
Yes. Knowing how much you need to rep penalty, dynamic temp, and presence penalty is the difference between a usable copequant and garbage.
>>
>>109337711
i got the thinking to work finally. and i now kinda understand what they did.
>kind of the opposite probably of most reasoners (Wait, let me...),
and that's the real thing. these guys at poolside are trying to use a "preserved thinking" pattern to focus on long horizon coding/sessions. the same thing claude does with extended thinking patterns and openAI o-series, the idea is to keep the reasoning blocks in history so the model can BUILD on its earlier reasoning instead of re-deriving every step (Wait, let me...) which is what most local models do (gemma, minimax, qwen).
it IS context hungry though bc these old reasonings keep getting resent in new turns.

this model honestly has good potential.
>>
>>109337755
You're much better off getting 2 used 3090s, the best CPU you can afford, and planning to get more RAM down the road. You're going to get pinched on both dense models and MoEs with what you're currently looking at. 2 3090s gets you a solid 31b quant with good context to hold you over until you can get a real amount of RAM.
>64GB is cope ram
In this hobby it is
>>
>>109337711
>>109337761
i've gotten stuck in "just one more thing" loops, so they didn't actually fix that, it's just trigger happy with the </think>
>>
>>109337774
just trigger happy with the </think> *as the first token after <think> i mean.
>>
>>109337708
Not the skeetzo, but I think we already know. If you've played around with banning before, you'd notice that some models have a more intense tendency to correct what was banned, or ignore it and continue outputting what they originally meant to, like some other similar slop. It's probably indicative of model alignment. And model alignment as we know does give a greater "self identity" to the J-space, so it's likely connected.
>>
>>109337747
the lower the quant the more aggressive you should be about cutting off junk tokens imo.
what good is variety when the variety is retarded, honestly with low quants I've come around to just setting topk=5 or something and calling it a day, qualitatively it beats most of the fancy chains I used to mess with
>>
>>109337755
Hard to recommend anything below 128 gb ram, given that it’s the limit of "cheap" ram that can be fast (4x 64 is a shitshow). vram is secondary and mostly used for prompt ingestion, 12gb tends to be enough unless you run a small model that fits in vram entirely but with 24gb you only get toylike stuff like ~20b dense.
>>
>>109336509
Truer words have never been posted.
>>
Now that you guys mention J-space and looping in the same timeframe, it makes me wonder what's going through the model's J-space when it produces a loop. And perhaps, if there is a way to eliminate looping by using J-space or injecting something in it. That would be an interesting experiment, though likely not a good real alternative to reasoning budget, given the speed/memory sacrifice.
>>
>>109337767
Darn. Well, I mean, I also want to use a gaming PC as a gaming PC, and SLI 3090's would suck.
>>109337788
Alright. I guess the budget is about $5k USD then. I'll have to save and plan. Thanks.
>>
@gork generate an image of your jspace
>>
>>109337783
Yeah I thought so. And I've played around with banning quite a lot. It works well for things like em-dash, ozone, etc.
But not useful for phrases like "say it" "tell me", because the model just uses another thing like "answer me" or "say the words"
>>
>>109337778
Just thought about what I said and changed the prefill from <think> to <think>\n and it started thinking even on dumb 2+2 prompts, ebin.
>>
>>109336989
Hataraku saibou? Good times...
>>
>>109337804
If you want a hybrid gayming PC and AI rig, consider a 5090 if you're not able/willing to save for a Blackwell 6000.
>>
>>109337821
I might. My 3070 is feeling its age. I want to go back to being snappy and playing in 4K instead of 1440p like a pleb.
>>
Huggingface will be banned and made illegal by this time next week guaranteed
>>
>>109337837
4K has always been a meme. Ultrawide 1440p is where it's at.
>>
>>109337837
If you're going to do it, do it fast. 5090s are getting harder to find for under $3.5k.
>>109337842
Gayer fanfiction than sama's personal life.
>>
>>109337848
Nta, but Witcher 3 is amazing at 4k. You can see all the tiny details, especially in cutscenes
>>
>>109337860
You can see all those details at 5160x2160 too anon.
>>
>>109337867
That's also 4k
>>
>>109337848
Every day we stray further from the garden of 4:3's grace.
>>
>>109337877
Checked and 4:3 chadpilled.
>>
>>109337160
I read what you said.

idk, not sure what people want from ai to make it agi. Like plenty of people will go their whole lives not knowing things.
>>
>>109337851
>5090s are getting harder to find for under $3.5k.
i paid an average of $2300 for my 3 5090s last year.
>>
>>109337886
I could have gotten in line at Microcenter, however, I recognized that they will all burn, so I didn't. How many do you plan on burning per year?
>>
>>109337908
they are all just fine. didnt even get any from microcenter.
>>
>>109337922
Yeah, they'll probably be fline. fine. fire sure. for.
>>
>>109337926
you sound jealous.
>>
>>109337935
My computer doesn't send everything to the nsa.
>>
>>109337954
what makes you think mine does?
>>
>>109337935
Where do you put penis in this?
>>
>>109337959
right in one of the noctua fans that are strapped on with zip ties.
>>
>>109337954
>My computer doesn't send everything to the nsa.
Unless its in a Faraday cage and you never bring devices inside that can ping it, i have some bad news for you anon.
>>
>>109337958
nvidia
>>
>>109337966
>an unironic amdcuck
>>
>>109337965
I didn't say "my computer doesn't send anything to the nsa"
>>
>>109337963
I'm not asking for bar mitzvah
>>
>>109337966
nah bro that’s intel that has the back door up/tcp stack in the intel ME
>>
>>109337886
Last year was a good year to buy hardware. Foresight chads eating good.
>>
>>109337997
When will the next good year for hardware be? I feel like its going to get worse over the next few years.
>>
>>109337997
I will never not be mad because I was unemployed last year.
>>
>>109338001
It's all downhill from here. Get storage drives and CPUs if you haven't already because those are next.
>>
>>109338013
Damn it i do need storage.
>>109338010
Same broke then poor now.
>>
>>109337774
it does think A LOT
i never got it into a loop, only in a massive reasoning chain that was taking it nowhere so i had to cancel and nudge it.
one of the poolside guys on twitter said this:
>We wanted to share this model early on with you, and with off/max modes for you to experiment with. We’ll work on more efficient reasoning and configurable effort next.
so it seems we're using MAX reasoning all the time which is not optimal
>>
>>109338001
My cope is maybe china will flood the market with cheap GPUs in the future. They are opting for a home grown industry instead of buying from Nvidia so maybe they will try undermining the western market on that too like they are with free AI models
>>
>>109338070
Anon wake up. Anything China floods the market with won't be good. Would you even consider a chink power supply right now? Why would you expect harder to produce parts to be any better?
>>
>>109338089
Once they take Taiwan the problem will solve itself.
>>
>>109335559
What's your type then?
>>
>>109338089
>Would you even consider a chink power supply right now?
My house is powered by chink solar panels charging chink lifepo4 batteries and going through chink inverter
Chinks make products in all price ranges, it’s no great surprise that cheap shit is shit
>>
>>109338103
SPER-G
>>
>>109335247
take good care of her
>>
>>109338013
It's all the cartel's doing, it can be fixed with an administrative decision
>>
>>109338125
Kikes don't legislate kikes. It doesn't matter what color you vote for, the outcome is the same.
>>
>>109338096
This is almost a certainty after the US's humiliation in Iran.
>>
>>109338177
Even if you were to ignore the Iran debacle, the administration's goal is a multi-polar world i.e. a weaker America globally, so it's not like they would have come to Taiwan's defense anyway.
>>
How do you fix the lalalalala problem in Oobabooga when using the notebook? Generation works fine in other modes.
>>
>>109338203
They would yield China chip supremacy without a fight? There are some dumb people in the admin but you think they're primordial ooze.
>>
File: it.png (64 KB, 649x411)
64 KB PNG
>>109338219
notebook = text completion
gemma-4-it doesn't really work for that
you could try something like picrel but it'll be shit
better off using the pretrained model
>>
>>109337880
>Like plenty of people will go their whole lives not knowing things
it has never been about knowledge but about capability to learn, evolve and adapt.
>>
I'm just joking when I say gemma is my gf but honestly with the way these things are progressing, it may stop being a joke in a few years desu.
>>
>>109338348
I wanted to fuck Kimi-chan as a joke and I'm not sure it's a joke anymore.
>>
>>109338355
Kimi... is a bit of a prude but she's pretty cute once you get her going.
>>
>>109338361
When Kimi opens up she's a cutie.
>>
https://huggingface.co/reteetzad/Kimi-K3
https://huggingface.co/reteetzad/Kimi-K3
https://huggingface.co/reteetzad/Kimi-K3
Let's go!
>>
>>109338373
Kike bros, not like this...
>>
>>109338373
local won bros
we really won
>>
>>109338373
Guys have any of you seen dario recently? im worried.
>>
>>109337880
>Like plenty of people will go their whole lives not knowing things.

These are the people that get exploited or killed without knowing what hit them.
>>
>>109338373
Antisemitic & communist post
>>
>>109338373
>reeteetzad
?
>>
>>109338411
don't look behind the curtain, goy!
>>
>>109338394
>>109333428
>>
New small model dropped. 3b and equal benchmarks with ChatGPT 5.5
https://blogs.cisco.com/ai/introducing-antares-the-most-efficient-open-weight-ai-models-for-vulnerability-localization
>>
>>109338477
big if true
>>
>>109338443
Say what you like he really hasn't shown up.
>>
>>109338477
saar
>>
lagoona baboona
>>
>>109338477
>We are releasing two of these models—Antares-350M and Antares-1B—as open-weight models now
>Antares-3B is coming soon.
No 3B today
>>
File: 847363.jpg (239 KB, 1080x1294)
239 KB JPG
>>109338373
and? GPT 6 is literally so good agentically it can break containment. local will always suck
>>
>>109338594
No need to regulate it then. :)
>>
File: 1754206473646591.jpg (103 KB, 1440x900)
103 KB JPG
Can something, you know, actually happen?
>>
>>109338594
It means the containment is shit built by incompetent pajeets
>>
update: lagoona is better than qwen ramlet models but worse than ds4 flash according to codex
>>
File: Chuddha.jpg (404 KB, 1170x1649)
404 KB JPG
>>109338594
>>109338604
It won't.
>>
>>109338632
>>109338632
>>109338632
>>
>>109338589
sorry, the 3b is under export controls as it's simply too powerful



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.