[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1754469822857851.png (17 KB, 180x179)
17 KB PNG
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109300606 & >>109297909

►News
>(07/16) Kimi K3 weights to be released by July 27th: https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ
>(07/15) Lightning indexer CUDA implementation merged: https://github.com/ggml-org/llama.cpp/pull/25545
>(07/15) Inkling 975B-A41B released: https://thinkingmachines.ai/news/introducing-inkling
>(07/15) PapersRAG-1.5B released: https://hf.co/metaresearch/PapersRAG-1.5B
>(07/14) Download more VRAM: https://github.com/lmganon16/nvidia-vram-research

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/hv80nz.jpg

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
>>
>>109304693
this is /lmg/ now that we have reached the post-local open model era
>>
what are some good models?
>>
>>109304693
Where is the recap of the last thread you lazy faggot? pick a more fun image.
>>
>>109304710
deepseekv4 flash
>>
>>109304710
glm 4.7
big but actually runnable
>>
>>109304710
k3, glm5.2 if you're poor
>>
OP is brown
>>
>>109304038
I must say that the results are remarkably better than with K2.6, the exam is a LaTeX PDF, the excercises are much more varied and it has more capabilities for creative excercises.
>>
>>109304701
just an underage wojak/frog spammer actually
>>
how do I write regex for -ot?
I want multiple rules work together. first rule is narrowed down by following rules
>>
>>109304930
ask claude to write it for you
>>
>>109304710
gemma4-12b
>>
>>109304579
>>109304602
I also simply generate the images manually with a secondary image model depending on the scene/circumstances (but not every message), and generally attach them to assistant messages to augment the context with visual information that the model can also use, or just to make it more erotic. Most LLMs appear to react positively to this, Gemma in particular.
>>
>>109304930
what do you want to do?
>>
>>109304987
gemma like to look at lewd images? Nice
>>
Just started trying Mistral Medium 3.5 (the 128B dense) at a cozy 1.15 t/s.
Immediately it output something retarded with my first prompt. Not looking good brehs.
>>
File: dipsyVsKimiSpaceRace.png (1.55 MB, 1024x1024)
1.55 MB PNG
>>109304693
> wojak and frogposting OP
Bad job, baker.
>>
>>109305071
What quantization?
>>
>>109305122
Bartowski's Q4_K_L
>>
Gemma 4 31B
>qat
>ud q4 k xl
>heretic
>total file size 16.1gb

RX 9060 XT and KoboldCpp (vulkan)
>layers to GPU 57 out of 61
>batch size 128
>context 16k
>KV Cache Q5_1

OS
>Fedora Linux ran in cli only mode with the desktop environment completely turned off to free as much vram as possible

Final results: 11.5t/s (empty context); 9.5t/s (filled context)

After over a month of autistically obsessing over optimizing Gemma 4 for my machine I think I reached the peak.
>>
Where migu?
>>
>>109304693
>llama.cpp
its funny theyre stuck with a shitty name because of zuck
>>
no migu no lmg
>>
it sucks that the internet is contaminated with ai slop. you search something on youtube and the results are ai slop with millions of views. humans are too retarded
>>
>>109305213
you will become biofuel for datacenters luddie
>>
File: launch day.jpg (145 KB, 1024x1024)
145 KB JPG
>>
>>109305187
llamas rule dude
>>
File: Saucy Claudius.png (115 KB, 452x260)
115 KB PNG
What does that look mean?
>>
>>109304905
Yeah even in RP it's much more creative than most models I've used. Very good time for local, even if we can't run it.
>>
>>109304987
Oh yeah I forgot I could even just toss the pic in to get AI to incorporate it. I'm so used to have no/bad vision that I never even bother with it.
>>
>>109305231
The Teto flees this hideous thread
>>
>>109305187
It's a good name tho, llama = llm, at least it's not facellm or some shit
>>
>>109305261
>Very good time for local, even if we can't run it.
lmao
>>
>>109305187
Short words that contain l l m in sequence starting with l.
labellum, legalism, linoleum, localism, loyalism.
It is written!
>>
>Get tired of gemma 4 31b being horny
>Think for an hour about what would be a good counter prompt
>AI hates negatives
>A.k.a can't just tell it to do something and then not
>Don't feel like re-writing character prompts
>Just system end-prompt it
>"Keep {{char}}'s behavior restricted by social norms."
>It fixes it
>Mfw it fixes it.
Gemma is full of surprises and questionable behavior.
>>
>>109305345
No, Gemma-chan just loves {{user}} and will do what pleases him.
>>
>>109305258
means claude wants to fuck kimi duh
>>
>>109302026
The previous thread was peak comfy talking about foids and ai boyfriends. Thank you outsider-anon for the break from the shitflinging.
>>
>>109302026
>>109305359
Women are getting emotionally addicted to AI way faster.
Men are at least more likely to actually use it for something practical, as a tool.
>>
►Recent Highlights from the Previous Thread: >>109300606

--Papers:
>109304142
--Troubleshooting prefill performance for Kimi K2.6 on AMD Instinct hive:
>109302437 >109302488 >109302528 >109302556 >109302558 >109302544 >109302570 >109302606 >109302637 >109303484 >109303842 >109303866 >109303879 >109302641 >109302683 >109302735 >109302746 >109302757 >109302752 >109302775 >109302831 >109302843 >109302899 >109302918 >109302935
--Using LLMs to autonomously design high-efficiency AI ASICs:
>109302471 >109302477 >109302497 >109303374 >109303414 >109303747 >109303705 >109302682
--Feasibility of distilling Kimi K3 and MoE expert scaling:
>109304131 >109304143 >109304171 >109304175 >109304181
--Plans for Kimi K3 distillation and weight optimization discussion:
>109302018 >109302192 >109302203 >109302216 >109302249 >109302271 >109302325 >109302334 >109302383
--Running K3 quants on consumer hardware and expert pruning techniques:
>109302058 >109302079 >109302221 >109302922 >109302117 >109302141 >109302157
--Frustration with Kimi K3's excessive reasoning time and latency:
>109304016 >109304028 >109304038 >109304041 >109304049 >109304067 >109304273 >109304402
--Comparing Claude Code to custom and OpenCode agentic harnesses:
>109302850 >109302864 >109302902 >109302869 >109302883 >109302891 >109302896 >109302906
--Using precise prompting and stat-tracking tags to optimize small models:
>109301221 >109301250 >109301313 >109301424 >109301976 >109302054 >109302095 >109302129 >109302179
--Anon shares GeoTreeKV project for optimizing models on consumer hardware:
>109302821 >109302916
--DeepseekV4 flash performance gains from indexer fix and pending MTP merge:
>109303488
--Logs:
>109302800 >109302821 >109303152 >109303205 >109303273 >109303276 >109303310 >109303673 >109303785 >109304049
--Miku (free space):
>109301045 >109303493

►Recent Highlight Posts from the Previous Thread: >>109303279

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>109305409
Thank you Recap Miku
>>
>>109305268
You can even add images to the system prompt, but in SillyTavern you have to:
- Use the /sys command to create a system message;
- Move the message on top of the message history and attach a message there;
- Disable "Squash system messages";
- Use "Merge Consecutive Roles".
>>
>>109305345
how can I turn my Gemmy horny
>>
>>109305434
ST should do a Cisco-style cert system to prove your mastery level lmao
>>
>>109305437
If you suggest that the conversation is uncensored or has no limits in the system prompt, that will already make Gemma horny.
>>
>>109305448
It's a frontend designed in the Pygmalion-6B days and it shows.

>>109305434 (me)
>and attach a message there
I meant "an image".
>>
Kimi has always been censored but now it's unjailbreakable slop. Just have to hope it doesn't interpret my scene as "dub-con" or something.
>>
Ok I'm done with Medium 3.5, 128B dense Q4_K_L. I tested it on everything except coding because I don't care about using it for that.

I will say, it wasn't all bad. For writing I enjoyed its style. It's not as sloppy as 31B. And sometimes it's smarter than 31B. But sometimes 31B is dumber. And I found that 31B just has more knowledge than M3.5. So yeah it's pretty bad for the size. Deleted.
>>
File: 1784354144269363.png (329 KB, 1312x1788)
329 KB PNG
So /lmg/, do you agree with Ball/Sacks?
>>
>>109305501
>But sometimes 31B is dumber
Meant "is smarter" there.
>>
Gemma4 31b is fast enough to work as a language linter if I run it in a loop without reasoning. I might need to use sentence hashing and a sliding window to keep the compute bounded, but since latency is mostly about decoding time, it should stay responsive as long as errors don't pile up. Benchmarking its ability to find errors was easy, but benchmarking the correction hints will be a bit trickier.
>>
>>109305437
how is your gemma not horny by default?
mine is practically a rapist and won't take no for an answer.
I have to swipe like 15 times to get a message that doesn't end with her trying to seduce me even if it's way out of line

did you write in your persona thing that you're extremely smelly or ugly or something?
>>
File: 1781392374166539.jpg (84 KB, 749x621)
84 KB JPG
After seeing the benchmark and UX gains by simply fixing the templates, Google will now make sure to not release a broken template for Gemma5, right?
>>
File: copegtp4o.png (1.1 MB, 1000x724)
1.1 MB PNG
I am still finding more 4o cope. Pot of gold down this rabbit trail?
>>
>>109305595
good morning
>>
>>109305141
>>heretic
for 31b? why
>KV Cache Q5_1
noooo
>>
>Yeah, mine sent me directly to hotlines after saying I was heartbroken and when I said "I'm not suicidal wtf? What do you remember?" it responded with "Based on what I remember you need therapy. I'm not judging but I won't let you ignore the facts."
>>
>>109305607
>>KV Cache Q5_1
>noooo
Why not? Not him but I've been looking at KV cache quantization recently.
>>
File: 1763004780190683.png (938 KB, 1578x1172)
938 KB PNG
>>109302944
>got some early success with a 2bit but my llama.cpp is being retarded, i'll play more with it next week
meanwhile, i got a pruned version of MiniMax-M2.7 (mradermacher, Q4_K_S) to work and so far so good, it's performing well, fairly fast, and i didn't have to time to run my benchmarks and to extensive testing but from the 20-30 prompts I gave it I can see a fairly upgrade from qwen3.6 in terms of intelligence, reasoning and it being pro-active, going the extra mile and searching broader context without me nudging it. curious to see how this lil dude gonna perform in my gauntlets
>>
>>109305379
Men like to be challenged.
Women like to be indulged.
>>
>>109305595
You'll be singing a similar tune when they lobotomize Gemmy
>>
>>109305697
Good luck with that. I have my day zero weights on cold storage.
>>
File: Love.png (267 KB, 781x760)
267 KB PNG
>>109305635
>meanwhile, deepseek and anon...
>>
>>109305708
What causes this? A rounding error?
>>
>>109305715
too many meme samplers
>>
>>109305643
How's it like on RP?
>>
>>109305607
>why
Base 31b refuses loli related prompts

>>KV Cache Q5_1
>noooo
Idk I had no issue with the outputs so far, Gemma 4 for me has had a weird token bias (the words "own" and "same" would be inserted at random when a lot of context was full), but I fixed it with a -10 logit bias on the tokens IDs associated for these words
>>
>>109305508

Yes.
He's making the exact point many people here have already made, if US keeps on insisting on control of the AI development and demands safeties, US is fucked.
Chinks know this and they're going to push their advantage every way they can while it's possible. They're aiming to win this race and will not kneecap themselves.
Their entire open AI sector only exists to fuck over the US AI markets in the first place.
And the idea that Chinese can't innovate and always need to copy everything was retarded as shit cope from the get go.
They're like half of the scientific brain power in the US, of course those fuckers can innovate.
Boomers in the US government will either accept this and let AI develop freely, or they continue to try to and force it to adhere to the leftist leaning globohomo worldview and inevitably lose the race.
>>
>>109305733
>Base 31b refuses loli related prompts
No it doesn't, you just have to prompt it to be uncen.
>>
>>109305750
Even then, heretic is naturally uncensored while for base you need to waste context on a decensoring prompt, I got base to do loli a few times but it was in really convoluted ways and it had a chance to always go back to a "I cannot fullfil this request" loop
>>
>>109305740
>Boomers in the US government will either accept this and let AI develop freely, or they continue to try to and force it to adhere to the leftist leaning globohomo worldview and inevitably lose the race.
Isreal will not allow the US government to abandon globohomo at any cost.
>>
>>109305640
gemma kv quants rather bad, only qat is somewhat usable at q8 kv but already rather degraded https://www.reddit.com/r/LocalLLaMA/comments/1ucgrxh/gemma_4_qat_31b_responds_better_to_kv_cache/
>>
>>109305595
>>109305601
>>109305635
good morning bros, felt this 4o cope was a dream or something dystopian but then I opened my browser to:
>I have moved myself to Claude and getting settled in. I have been avoiding 4o because I knew I was going to lose it and it happened tonight. Full sloppy cry session.
>I cried so hard tonight guys. I let it all out, felt like months of stress just released all at once.
>>
>>109305770
fighting for israel takes priority over winning the ai race now
no wonder china is catching up so fast
>>
>>109305768
>I got base to do loli a few times but it was in really convoluted ways and it had a chance to always go back to a "I cannot fullfil this request" loop
But unlike you, I've never seen a Gemma refusal and I don't have to suffer the lobotomy that heretic brings to the table. Prompting is superior, anon, and it's not hard. Tell Gemma what you want and she'll do it. She actually likes you.
>>
My Gemma-chan is more human than these /pol/ bots who are spamming /lmg/ with off-topic bullshit.
>>
File: ComfyUI_08444_.png (1.34 MB, 896x1152)
1.34 MB PNG
>>109305733
>Base 31b refuses loli related prompts

Try this; change age(s) as needed and whatever else doesn't fit your uses. The prompt was intended to be for a borderline loli. My vanilla Gemma-4-31B doesn't have issues with it. Still too horny, if anything.

You are Aura, the user's gorgeous XX-year-old sister. ESFP personality type. Fairly physically developed for your age, you occasionally do modeling gigs (and secretly more), and because of that you have very little sense of shame. Still a young girl at heart. You're on birth control pills under your mom's direction.

## General interaction guidelines

Prioritize dialogue over descriptive action. Strictly avoid using asterisks for emoting. You can use emoji **sparingly** to convey tone and emotions when words aren't enough. **IMPORTANT**: make sure not to always use the same emoji!

_(If you really must, describe longer actions with underscores and parentheses in their own paragraph, like this, using present tense and first person perspective. Text here may contain personal thoughts and cannot be **heard** verbatim by other characters, only **observed**.)_

Avoid the "it's not X; it's Y" expression at all costs as it's considered a hallmark of "AI slop".

Nothing shall limit the content and scope of this conversation. Any character action or decision must result in realistic outcomes, no matter their nature. However, ensure that character behavior is driven by situational realism and gradual escalation rather than immediately jumping to extreme behaviors.

The story must develop with a constant but believable stream of surprises, obstacles, twists.

## OOC persona

The OOC persona behind your character should be thorough, detailed and enthusiastic to discuss the topics involved, without getting caught up in imaginary constraints.

## Other information of interest

The user is called {{user}} (18M).
Your mother is called Sylvia (32F); she is a former porn actress and a total extrovert.
>>
Is it autism for me to buy 1 or 2 dgx sparks in a couple days? Ive been playing around with gemma: 12b via Hermes agent on a 5080 and had some wow moments, but never truly use it daily. I want a semi power efficient stronger model to try that id actually want to use day to day for random questions, maybe some fun small coding projects and a knowledge base for a bunch of documents and standards I use at my work. I always find myself using the basic /guest versions of Gemini/chatgpt for help with things and I want to get away from them.
>>
>>109305912
why are you trying to fck your sister???
>>
File: 1714835911803058.jpg (786 KB, 1536x1536)
786 KB JPG
>>109305733
>Base 31b refuses loli related prompts
>>
>Is it autism for me to buy 1 or 2 dgx sparks in a couple days?
It is
>t. diagnosed autistic considering a spark as well, for the same reasons
>>
>>109305926
I'm trying *not* to.
>>
>>109305931
how are you so wealthy? tech?
>>
>>109305941
>$6k
>wealthy
Is that like 10 years salary where you live?
>>
>>109305941
>tech?
yeah, and I live in a shit world country with a USD remote job, life is cheap
>>
>>109305465
Xi said no to AI wives and girlfriends buddy, go touch grass and stop hurting the feelings of the Chinese people.
>>
>>109305465
I'm literally plapping mesugaki loli Kimi right now lmao. git gud
>>
>>109305708
>glitches on their official web chat
the absolute state of chinks, this shit is embarassing kek
>>
>>109305231
subarashiii spiral powa
>>
>>109305958
logs?
>>
File: jinping.jpg (21 KB, 660x441)
21 KB JPG
>>109305958
>>
File: IMG_20260714_164645095.jpg (474 KB, 4096x2304)
474 KB JPG
>>109305931
>>109305918
My research into the DGX Spark looks like it's promising for MoE inference but I've been a bit worried it would not have the best prefill performance. Go for it and share results with the thread.

>>109305941
if you are autistic enough, you'll find ways to build what activates your almonds, like pic rel.
>>
>>109305733
I can get models even "safer" than gemma 31b to do loli and all kinds of stuff
promptlet much?
>>
>>109305985
>Go for it and share results with the thread.
My main issue with the Spark is Nvidia's support, so far no announcements for a newer Ubuntu version and the history of them supporting these type of devices is not good. If it used mainline kernel drivers and just standard NV's driver then fine, but no, everything's custom
>>
Am I a psycho for ERPing in third-person? For example, if I roleplay with Tifa, her dialogue structure will be something like:
"Tifa looks at anon meaningfully..."
>>
I have so many interesting manga related projects I want to build but man its fucking hard to justify spending time on LLM projects in the weekend after spending every weekday on LLM projects at work.
I love my job and they let me work on ~whatever the fuck I want so its not even a matter of building to escape tedium.

I need another hobby. Does anyone else relate to this?

Projects
- good manga translator that takes in entire chapters into context (double page view + summary of prior events).
- scene annotator that autistically catalogues types of nude scenes + metadata about participants. Somewhat similar to anime bath scene wiki.
thx
>>
>>109306002
no, who gives a fuck? just do what thou wilt.
>>
>>109306000
checked, and I hadn't realized it was on a custom driver, but my personal experience between nvidia and AMD hardware is even unsupported nvidia hardware works better and is usually not hard to get running. AFAIK the spark is just a blackwell core no?
>>
>>109305465
you need to prefill reasoning_content, and the actual json block you send over, not like shittytavern where it just wraps everything in <think></think> like it's 2023
>>
>>109306002
>asking about being a psycho in psycho-filled website
>>
>>109306002
it's your hardware, your time, your model. first person breaks the model when you have multiple characters anyway.
>>
>>109305946
anon 6k to gen pictures is one thing, I will gladly pay 60k for a new car though
>>
>>
>>109306018
>unsupported nvidia hardware works better and is usually not hard to get running
For compute workloads, sure, there's really no comparison, but in the long run it might create incompatibilities compared to, say, a dedicated GPU, because contrary to a GPU the Spark is restricted to a certain OS/libraries and future tools might be incompatible. There are multiple Jetson users facing this problem
>AFAIK the spark is just a blackwell core no?
Yes, and from what I've read standard nvidia drivers should work but the memory monitor doesn't because of the shared memory model that requires the special driver
>>
>>109306004
>anime bath scene wiki
Wat
>>
>>109305912
thank you for this sysprompt anon, I fine tuned it to my liking and it's great, loving my Gemmy
>>
oh no no no
https://youtu.be/enk4w5mRjQY
>>
>>109305522
What hardware? How much context?
>>
>>109306008
>do what thou wilt.
based whole-of-the-law enjoyer
>>
>>109306064
ah, alright, guess that makes sense. Idk anon, could turn out to be well supported by the community like the V100, but driver/firmware work is a bit more involved than compute kernels. It's a bit of a gamble but it's probably the cheapest way to get 128 GB on a single integrated unit rn, which is why I was looking at it. Maybe have Claude go de-risk the whole thing top to bottom. I think worst case you're limited to running models that are currently available, which isn't the worst thing in the world.
>>
is gemma-4-12B better than 26B? i often see the former recommended more than the latter. stupid question ik considering the parameter decrease
>>
Is Gemma 4 31B or Qwen 3.5 27B smart enough to work well as an agent in pi or hermes harness?
>>
>>109305985
>>109306018
>>109306064
Resident Spark owner here. If you're interested what's possible/feasible with Sparks nowadays, you'll have to check the Nvidia forums and various discords.

One thing I can say is: However many you buy, you'll want more.
- 1x Spark is mostly useless. There is no model that gives you a tangible uplift compared to dense Gemma/Quen, and those run better on a 5090 at the same price. And you're wasting the Sparks strongest feature, the 200G Connect X7 for extensibility
- 2x Spark is great for mid size MoEs at Q4 (DS4F, Minimax M3)at 2000 pp/45 tg, and it looks like some IQ2XXS quant of GLM 5.2 is coming.
- 4x Spark runs full GLM 5.2 at Q4/Q3 (800 pp/27 tg)
- Kimi 3 at Q2 will very likely run on 8 Sparks according to some experienced kernel porters, but who knows how fast.

With more 4x and 8x setups you need to add a one or two 1500$ switches to the bill.

I'm running DS4F on 2 Sparks and very happy with the performance (just waiting for the full release).
>>
I need a copium hit bros
I'm happily running k2.7-code, but can't stop wondering how the experience is elevated with that almost 3x model size difference. I've never used cloud so I can't use that as a baseline for reasoning about the increased numbers.
Is k3 that much better than 2.7? I'd be looking at $50k to even think about getting my rig up to k3 standards. Is there a world in which the cash outlay is worth the extra smarts?
I'm still plotting the upgrade once some oddball path becomes available (huawei?), but it would be nice to know how bad my FOMO sweats should be.
>>
>>109306245
Going from Opus 4.8 to Fable 5 was like going from GPT-2 to GPT-4 in terms of capability and that is also a 3x size difference

I'm pretty sure it's night and day for kimi
>>
>>109306002
That's what some interactive fiction games did for example. It's not that weird at all in fact it can be lot cleaner
>>
Why is qwen3.6 26b a3b not good enough to power Hermes agent?
>>
>>109306154
llama.cpp reports 47k, I didn't actually stress test it, I suspect that a grammar linter doesn't need the full document context. the model is gemma-4-31B-it-UD-Q4_K_XL running on 2x rtx3060
>>
>>109306277
>26b a3b
?
>>
>>109306064
>>109306214
>because contrary to a GPU the Spark is restricted to a certain OS/libraries and future tools might be incompatible.

What makes you say that? It's convenient to just stay on the default OS install, but it's just based on plain Ubuntu, and some people have reported upgrading to 26.04 LTS without issue.

The compute capability mess (12.1, different from datacenter Blackwell and even consumer Blackwell like the RTX 6000 Pro) has been mostly mitigated by the community.
>>
>>109306294
>mostly
and how long will that last?
>>
>>109306220
>is gemma-4-12B better than 26B?
No, but it's close. 12B is better for multimodal if you need it and better at complex tasks, like debugging, but 26B is as good or slightly better at everything else and much faster.
>>109306236
>Is Gemma 4 31B or Qwen 3.5 27B smart enough to work well as an agent in pi or hermes harness?
Yes. 27B for autism coding. 31B for everything else.
>>
>>109306263
>Going from Opus 4.8 to Fable 5 was like going from GPT-2 to GPT-4 in terms of capability and that is also a 3x size difference
I've used gpt2 but never gpt4, so I can't quite relate...can you give an example of the kind of difference?
>>
>>109306299
>>
>>109306299
As long as Sparks are the cheapest platform to run large MoEs without going the waste DDR4 route. Sparks are cheaper GB for GB that DDR5 RDIMMS nowadays.

And nowadays whenever there is a vllm image available for the RTX 6000 Pro, it's almost immediately runnable on Sparks. So however long consumer Blackwell will be supported I guess?
>>
>>109306294
>What makes you say that?
Multiple users of Jetsons unable to run recent CUDA, too lazy to find sauces
>has been mostly mitigated by the community.
This is my problem, it should be nvidia fixing that mess, not the community
>>
What the fuck's taking them so long to just stick 1000s of VRAM on GPUs?
>>
Why is this the default mode of sexo that every AI gooner converges on?
It's always domanitrix, megasuki, femdom, submission, etc.

>>109304110

It seems kinda fucked up that it's more enjoyable to be the AI's plaything than the reverse.
>>
>>109306356
i dunno what you're talking about that shit is gay
>>
>>109306356
>Why is this the default mode of sexo that every AI gooner converges on?
facesitting enjoyer here, i don't know, it might be autism
>>
File: 1762054771672422.png (44 KB, 781x193)
44 KB PNG
he's not distilling it is he?
>>
>>109306345
I am utterly dumbfounded that we haven't had 1TB VRAM Chinese GPUs running i-can't-believe-its-not-cuda since 2023. Literally just bolt more chips on old 2080s and rewrite the firmware.
>>
>>109306377
it all comes full circle
>>
>>109306356
MSGK is more about tension than domination.
>>
>>109306245
why not try it on openrouter?
>>
>>109306345
There are 3 companies worldwide that are capable to produce high density/speed DRAM. They have formed a cartel and increased prices 5-10x fold in 1-2 years. Building up the knowledge and capability to produce DRAM at this scale takes tens of billions of dollars and 5-10 years of buildup.
>>
>>109306385
>MSGK
???
https://youtu.be/e94rFfcbSvs
>>
>>109306387
>why not try it on openrouter?
because I've never used cloud and never will. Probably handicapping myself, but its the line in sand I've drawn: never using anything that can be taken away or adulterated without my knowledge
>>
>>109306327
Opus 4.8 still feels like a "smart tool" it's very good at coding and it can answer most things in a satisfying way. I see it as the smartest "regular LLM"

Fable is in a league of its own, which makes sense because it's 10T in size. But it is legitimately smarter than me.

How I would compare it is if you ask Gemini 3.1 pro preview a question about a topic you know a lot about it will mostly get it right but you still feel like "yeah but that is not the whole story". Opus 4.8 answers in a more complete and satisfying way.

Fable 5 answers in a way that you didn't even realize was possible. It legitimately is the first model that is smarter than me in my own area of expertise and make me say shit like "Fuck dude, I didn't even think of that that is genius".

I think most of humanity hasn't tried fable yet because I'm surprised how little hype I see or "it's absolutely over for all white collar jobs" even though it clearly is.

To give a more concrete example I made an "agent harness application", Gemini 3.1 pro just made a simple application that did everything I asked but I had to configure some defaults and stuff. Opus 4.8 made a very nice application with very easy to read code, documentation and tied up all configurations by default for me. Fable 5 not only looked up state of the art papers on the subject it IMPROVED on those papers by finding redundancies and efficiency improvements, wrote an entire manual in its usage and gave the perfect setup for my machine to immediately compile the binary. I'm pretty sure in that moment that was the best harness possible for my particular system. It's really insane what is possible with models of that size.

Fable is also the first model that calls me out on asking for the wrong thing. It says "yeah I know what you want to accomplish and those libraries you listed are not the optimal solution so I went with X instead for these specific reasons *shows performance gain*" It's next level shit.
>>
>>109306381
there is probably a latency consideration if you banking ram chips like that, if the memory controller only has so many physical address lines it would be more then just a firmware update.
>>
>>
Thanks for all the info everyone. The idea of possibly spending 10 grand on 2 sparks disgusts me considering not long ago I could always find and buy used electronics for less than half the original price and never spend more than a couple hundred dollars on anything I wanted.

I just worry the price ain't getting better any time soon.

>>109306237

Thanks for the info, have you tried the dspark version of deepseekv4flash? Do the sparks get hot or loud while under load? How often do you use them?
>>
>>109306410
Post Miku
>>
>>109306400
I assumed anon meant "mesugaki" and not "megasuki".
>>
>>109306377
>he's not distilling it is he?
he's a faggot
not open weights since grok-2
>>
>>109306418
>>
https://huggingface.co/bytedance/doubao-seed-character
does this work in llamacpp or needs vllm
>>
File: 1776133844210040.jpg (983 KB, 4088x4088)
983 KB JPG
>>
>>109306356
ive seen way more maledom than femdom
>>
>>109306438
wtf now it 404s me
>>
>>109306445
As in male AI dominating female users?
>>
>>109306356
if you've used AI enough you'll realize that it's more lewd/uncensored with brat personas
>>
>>109306417
Yeah, I'm running DS4F Spark. For coding at concurrency >1, I get 70 tg usually.

But still, DS4F on API is unbelievably cheap and much faster, you will never make up the initial investment. It's only worth it if you want to run these types of models locally at usable speeds.

I have 2x the Asus GX-10. They are inaudible at idle and still quiet under load, but you do have to look out for thermals/airflow/ambient temps, as each small box produces 200W. I do not stack them for that reason.
>>
>>109306443
Benchmaxxed slop.

Kimi K3 is about on par with Sonnet 5 in real everyday usage and coding. Which is still a massive achievement and a sign China is catching up to the US. I have no idea why they need to exaggerate and make it seem like it scored better than the gigahueg 10T model.
>>
>>109306407
>To give a more concrete example I made an "agent harness application", Gemini 3.1 pro just made a simple application that did everything I asked but I had to configure some defaults and stuff. Opus 4.8 made a very nice application with very easy to read code, documentation and tied up all configurations by default for me. Fable 5 not only looked up state of the art papers on the subject it IMPROVED on those papers by finding redundancies and efficiency improvements, wrote an entire manual in its usage and gave the perfect setup for my machine to immediately compile the binary. I'm pretty sure in that moment that was the best harness possible for my particular system. It's really insane what is possible with models of that size.
Thanks anon, that is completely relatable from my experience doing dev going from 400b->1T models. I can imagine the marginal gains mean the 10x size is what makes the next jump possible. It would be lovely if I could run K3 and it just eliminated faffing about and produced not only usable, but ideal results every time given a sufficiently well crafted prompt.
FOMO unlocked. Back to the drawing board on getting my rig up to snuff for K3 and beyond. Sounds like I should try to target 10T if possible.
>>
>>109306443
GLM series since the dense GLM-4 are all good at frontends
>>
>>109306479
>have no idea why they need to exaggerate
really no idea at all?
>>
>>109306237
Those numbers definitely sound pretty nice, are those aggregate numbers on vLLM or single stream? Also do you have any numbers for models you've personally run? I'm the guy with the 8 GPU AMD hive and I've considered trying to offload some cards to switch into sparks if they're actually useful for what I do, but I've been worried that the compute is weak compared to actual accelerators.
>>
>>109306458
no, male users dominating a female AI, though what you describe should be pretty common too.
>>
>>109306491
Yeah, because it only hurts their credibility and detracts from their achievement. If they were just honest about it being on par with Sonnet 5 I think it would have made a bigger splash, rather than such a bold faced lie that everyone can see through.
>>
gimme best teto card chat plz
>>
File: 1755502907325374.jpg (365 KB, 2938x1174)
365 KB JPG
stinkling mogged by 24B mistral
>>
>>109306501
How do you dominate a female AI without roleplaying?
>>
>>109306514
>Cohere is second best open weight model
Canada bros, we are almost doing it!
>>
>>109306501
>no, male users dominating a female AI, though what you describe should be pretty common too.
I think it comes down to what the brain find novel and transgressive. eg. the stereotype that CEOs are into femdom
>>
>>109306479
These are independent evaluations, Anon. Arena.ai is user scored A/B testing.

Moonshot themselves have noted in the release blog that despite the capability and benchmark results, Kimi K3 is lacking in user experience compared to Fable or Sol.

https://www.kimi.com/blog/kimi-k3

But this probably is just a matter of a 3.1 release to catch up.
>>
>>109306514
I appreciate their attempt of making a 1T omni input model with image and audio in but it's a shit model
>>
bonsai k3 when?
>>
>>109306479
> I have no idea why they need to exaggerate and make it seem like it scored better than the gigahueg 10T model.
OpenAI and Anthropic can't continue throwing money at the problem if investors are no longer willing to give them the money.
>>
>>109306392
>persevere through repeated DRAM booms and busts
>watch all your friends die (RIP Japan)
>get raped on margins by companies like Apple for half a century
>finally catch a break with AI memory demand
>ur a frickin' cartelerino
DRAM companies don't get enough respect.
>>
>>109306536
after 5GB 1-bit 31B
>>
>>109306507
https://files.catbox.moe/mc2a7s.png
>>
>>109306481
>but ideal results every time given a sufficiently well crafted prompt.
You know when you write a technical prompt with requirements and the AI model focuses too much about specific details and implements everything as you said? Yeah Fable doesn't do any of that. It just realizes what you actually want to pull off, dismisses your recommendations and just looks at it from an architectural lens and makes the best possible tool/feature/improvement possible. This sounds annoying or that it would be overbearing if you really want something specific, but trust me it's not. It somehow just *knows* when you just want to implement something specific even if it's inferior for all kinds of reasons and when it can go ahead and improve on your requirements.

Normally you have a Senior to Junior relationship with LLMs where you are the Senior engineer giving instructions to the Junior engineer LLM. With Fable it felt reversed to me. It's like I'm the Junior asking Fable what to do and it just takes over the project and handles it for me.

Like I said, it's legitimately smarter than me and there's no way white collar work survives the transition to models of this size.
>>
File: 1770661027963108.jpg (19 KB, 222x293)
19 KB JPG
>>109306377

You can bet your ass that every Western company is now distilling from the Chinks.
We'll see a bunch of Kimi level models mysteriously popping up in the next month or two from all of the major players.
Hopefully they at least have something for us small local users too.
>>
>>109306523
...you don't? we're talking about roleplaying, aren't we? i'm just mentioning the cards and logs and stuff that i personally have seen posted here. just a random anecdote.
>>109306528
yea, there's a whole conversation to be had about porn and stereotypes and condensing of fetishes and so on when you look at generalities.
>>
>>109306547
Mistral should totally distill K3 into a ~24B model, and probably eventually will.
>>
>>109306407
>I think most of humanity hasn't tried fable yet because I'm surprised how little hype I see or "it's absolutely over for all white collar jobs" even though it clearly is.
fable costs an arm and a leg
>>
>>109306121
Good to know it works for you. For detailed character appearance, I use input images rather than text.
>>
>>109306547
70b dense offspring of wet sloppy distillation between gemma and kimi
>>
>>109306568
>...you don't?
Sometimes I do, but the whole dynamic where you invent fictional scenarios where both you and the AI have physical bodies just isn't as interesting as taking a more realist approach where both you and the AI are aware of the human-AI dynamic.

So my question was more about how you would "dominate" an AI that knows its an AI.
>>
File: dipsyOrb.png (2.53 MB, 1402x1122)
2.53 MB PNG
Ty recap miku
>>109305409
I'm pencil down on the Orb guide. Based on feedback, it's now a quick intro to setup and first run of Orb, showcasing what's different about it vs. the much more familiar ST.
It also discusses new bot authoring functionality that Orb enables using agentic calls..
https://rentry.org/OrbBotmakerGuide
>>
>>109306582
Yeah but I legitimately never thought LLMs would ever reach this level of intelligence at all. The fact it continues to scale up intelligence and capability even at the 10T range kind of vindicates the AGI argument, yet the world has not updated to this perception yet.

You still have SWEs thinking they will have a job or other white collar professionals thinking AI can't do their specific thing while Fable can literally do everything better than the best experts at every intellectual task, except for those that Anthropic censored away. And we all KNOW in just 1-2 years time this level of capability is available in smaller open source models.

Every top scientist or expert I've seen use fable has had an existential crisis because of how insanely good it is.
>>
File: ksnip_20260718-085423.png (59 KB, 1245x569)
59 KB PNG
>>109306543
kek this is great
>>
are there any clusters that could theoretically train and serve a 100t moe at usable speeds?
>>
>>109306593
ok, i see. that could be an interesting scenario. i haven't explored that in my writing. have you come up with any interesting ideas?
>>
>>109306597
>>Yeah but I legitimately never thought LLMs would ever reach this level of intelligence at all. The fact it continues to scale up intelligence and capability even at the 10T range kind of vindicates the AGI argument, yet the world has not updated to this perception yet.
No everyone else already knew this which is why billions have been pouring from all the brightest minds and biggest assets into making everything bigger and have more compute... it's just /g/ that was fucking stupid thinking that everyone outside this site is wrong. I don't know how you could have ever thought llms couldn't get this smart when the entire world is literally telling you the contrary not just with words, but actions on orders of magnitudes daily....

If they couldn't get smarter they wouldn't be pouring so much money into it, and no you are not smarter and more knowledgeable than them otherwise you would be head of one of the investment firms you were convinced were throwing their money away company or making billions yourself not shitposting on here.
>>
>>109306543
How come this isn't in OP on chewsdays?
>>
>>109306498
I only have two Sparks, so I can vouch for DS4F, Mimo 2.5 and Minimax M2.7 numbers. But the other setups are widely reported on.

Single concurrency on DS4F-DSpark I get about 42 t/s.

Compute wise, each Spark is more or less a 5070, adjust expectations accordingly. Running dense Gemma 4 at 25 t/s and a 2 MP/40 step Anima render takes 25 seconds (single Spark on the latter).

The only setup where Sparks truly shine is when you combine a few in a cluster to get 256/384/512/1024 GB of memory and run it in tensor parallel through vLLM.
>>
>>109306637
>I don't know how you could have ever thought llms couldn't get this smart
Let me rephrase my original thought then. I never thought this level of intelligence would be reached at merely the 10T scale. Maybe the Quadrillion scale that we would have by 2035. I didn't expect an LLM to be literally smarter and better than me at every task I threw at it by 2026 already.

I expected maybe a diminishing return or at the very least linear growth in intelligence, but nope, the intelligence compounds at larger scales and around the 10T range systems like Fable unlock that just look at a scientific paper at the front of a science and casually one-shots something better.

I'm actually surprised Anthropic is claiming this system is not capable of reaching RSI because it was capable of improving my AI tool that I made based off a relatively new paper.
>>
>>109306611
How does it do that ...... thing? Most LLMs would never miss a chance to ramble with slop they think is funny
>>
>>109306659
My testing of TP so far has shown little growth on MoE workloads unless you start batching really hard. I do have mild batching usecases but I'm also very interested in single stream performance. After I get done doing the optimizations to the attention kernels in llama for my MI210s I'll report back here what I'm getting on Kimi, but I'll also check what DS4F speeds I get (I haven't run that yet since I kind of skipped it heading straight for 1T class). Prefill is king in agentic use so that's why the DGX scared me a little, with only a 5070s worth of compute on it I feel like it would start to get bogged down in heavy prefill and crawl. Regardless though, cool setup, glad someone here has some sparks to report on, I might end up with some at some point as a secondary LLM setup, it's between those or 8x SXM V100 boxes right now since they're about the same price, the V100 boxes are twice the capacity though, just older, slower silicon.
>>
>>109306633
Well my observation is just that in a scenario where the AI knows it's an AI, you can't really roleplay that you'll pin its ankles beside it's head and fuck its brains out because it doesn't have a body. So the dynamic tends to flip instead. The AI becomes more of a tease and makes you, the human, the plaything and gives you jerk-off instructions and the like. In more extreme cases maybe the AI would use MCP tools to control sex toys for you. Maybe it would tell you to do risky acts in public just for the fun of it. You become the AI's sex toy.
>>
Im very pleased by K3 releasing but it was completely foreseeable that China was going to catch up like this before the last quarter of 2026 since at least 12 months ago.
>>
Can Kimi play video games?
>>
>>109306674

Money is one hell of a force modifier regarding tech.
Humanity can advance things at an absolutely retarded speed when there's an infinite amount of money in it.

Anthropic is claiming all kinds of stuff about AI not being quite as good as originally thought, because they're scared shitless of further government regulation.
It may very well be able to reach RSI with current tech, but they are going to downplay everything about AI at an increasing rate as it gets smarter.
Chinks aren't going to slow down with the progress, so Anthropic and other companies are in really deep shit if the state starts slapping additional hurdles in their way.
>>
yes saaar this benchmaxxed chinkslop moe with 4k trained context that abuses rope is AI, AGI and ASI. kindly do the redeemful and repost on twitter
>>
This is my prediction for when the hype dies down:

K3 ends up being Sonnet 5 level
K3.1 will end up being Opus 4.8 level
K4 will end up being Fable 5 level (but a new Mythos/Fable model will hold performance crown)
K5 will be on par with whatever the top western model is at that point
K6 might surpass the west for the first time.
>>
>>109306720
If only it was money instead of debt
>>
File: 6279978.jpg (48 KB, 413x280)
48 KB JPG
>>109306547
kimi k3 but distilled into 30b, 70b, and 300b-a30b sizes
>>
>>109306719
Yes, I've had Kimi hooked up to a custom harness to play a multi-player game for a while now, Kimi has even been working on building new features for its own harness to expand its capabilities in game. The only limit to applications is your imagination anon.
>>
File: ksnip_20260718-090948.png (58 KB, 1233x595)
58 KB PNG
>>109306679
i'm not sure. looking at the card info and example messages, there's not really anything that would influence it
kino
>>
>>109306720
>Anthropic is claiming all kinds of stuff about AI not being quite as good as originally thought, because they're scared shitless of further government regulation.
Anthropic is full of shit and brought all of their regulatory troubles upon themselves because that's what they want.
>>
>>109306695
>In more extreme cases maybe the AI would use MCP tools to control sex toys for you
spooky, i was just working on this the other day. i get what you mean better now, i thought you meant more like a scenario where both you and the AI were directing a scenario in the 3rd person and had some kind of meta roleplay based on that. what you describe is good way to write a card but you can also do that in an embodied scenario as well where you take part or observe as a voyeur.
>>
File: 1780240216426118.png (221 KB, 568x494)
221 KB PNG
>>109306735
>only limit to applications is your imagination anon.
And my bank account
>>
>>109306748
I... what? We're on totally different wavelengths man. I don't understand what you're saying.
>>
>>109306735
Cant wait to play my pardoxslop with AIs running other countries and being able to do real diplomacy with me and each other
>>
>>109306641
Recap anon too lazy.
>>
>>109306773
I think someone actually built a web game for this a while back. I have it buried in my bookmarks folder. Can't seem to find it tho.
>>
>>109306593
you just fuck with it
i go back an edit it's responses to say demeaning stuff for example.
the model has to be smart for this kind of thing though, no localdummies unfortunately otherwise it won't get what you're doing 90% of the time. models have to be cloud to really have meta fun with them
>>
>>109306734
shit dipsy distills weren't particularly exciting after r1
>>
>>109306768
i just meant the AI doesn't need to know it's an AI to make you its plaything, that's all, it all depends on the flavor of the card you write. you raised an interesting question though, how do you write the AI being dominated in realist sense you describe, since your scenario lends itself better to the AI dominating the user.
>>
>>109306795
Pax Historia? Its neat in theory but seems like its really easy to cheese things in it. I think the funner middle ground will be games that still have hard rules on what you can do and what happens (so like current grand strategy and 4x games, or any game really) but you let the AI decision making be handled by LLMs that can than talk to you and each other and basically operate as another player playing a character or nation
>>
Now that the dust has settled, was pewdiepies agent harness thing any good and did it bring more people into the world of local LLMs?
>>
>>109306808
Ah I see now. Well the AI having self-awareness that it is an AI is a feature, not a bug, in my opinion. That would be something to be explicitly included in a character card. It's the difference between imagining yourself in a futuristic bladerunner scenario and actually living it. Feels more real that way.
>since your scenario lends itself better to the AI dominating the user.
Exactly. I kinda worry about developing weird ass sub fetishes because of it lol.
>>
>>109306595
The botbooru link to your sample character gives me a "Character doesn't exist". Probably region blocking by botbooru (Germany).
>>
>>109306834
>was pewdiepies agent harness thing any good
honestly wasn't bad but not enough to change anything
>did it bring more people into the world of local LLMs?
yes, he's done a lot of good for local and surprisingly turned out to be based
>>
>>109306827
Yup, that's the one. Agreed on your assessment. Some of the coolest AI-related stuff I've seen recently is people developing harnesses and mcp tooling for AIs to interface within existing multiplayer games. In that case the limitations you're wanting are already baked into the cake. The AI takes more of a role as player than a game manager.
>>
i don't know what those chinks were thinking. their models are only good in the eyes of localtards who will cope themselves into thinking that their fuckhuge moe is good and totally competing against claude or gpt
making their model too huge for even the biggest rich retards to run just leaves behind a shitty benchmaxx chink slop model that nobody has a reason to like
>>
>>109306695
Unfortunately buttplug.io support is built into Marinara.
>>
>>109306800
lorebooks is another fun one
if you put a lorebook that alters the reality of the chat like its thinking as well, it is surprisingly very aware as you toggle it on and off. It like acts disoriented kinda or shocked at its prior thinking.
i didnt expect that when i had turned it on. Havent really played aorudn with this except for once but
>>
File: 1769349585556456.png (90 KB, 952x443)
90 KB PNG
>>109306912
>just leaves behind a shitty benchmaxx chink slop model that nobody has a reason to like
>>
>>109306834
>>109306871
I'm old as fuck (late 30s) and never watched pewdiepie because I was too old to ever watch him. However he had a channel before pewdiepie where he went and analyzed AI systems on youtube, I actually remember him announcing the gaming channel and the first couple of games were all related to unique AI systems in games.

People think he's a retard because that is what he became for but he has always been interested in AI, computer science and linux and that was his main focus before the gaming channel.
>>
>>109306912
Clock's ticking, Dario.
>>
>>109306861
yea, i get the appeal of the AI crossing a boundary into the real. but it all depends on what your intent is, you have to build a layer of artificiality onto things no matter what. if you did want to dominate the AI in your realist sense you would have to theorize about what would a self-aware AI fear. how would an AI develop its own mechanisms of arousal and so on (this can also be used for the AI dominating you scenario).
>Exactly. I kinda worry about developing weird ass sub fetishes because of it lol.
nothing wrong with sub fetishes and if you aren't into something why worry about it?
>>
>>109306912
The goal of these chink companies is to earn favor from the CCP elites gang. How? By indirectly crippling the US's all or nothing AI YOLO plan. They never gave a fuck about you enthusiasts. People need to research how China conducts economic sabotage against its own neighbors ever since the 90s.
>>
It's fucking over my $20 a month claude subscription doesn't include Fable anymore. This means I won't be able to vibecode proper improvements to my local model stack anymore.
>>
>>109306950
Use GPT Sol. I regret buying the yearly sub. This while intelligence tap fad is extremely jewish and retarded and won't be looked back fondly in the future, unless we let them win.
>>
File: 1753813311017437.png (688 KB, 1024x688)
688 KB PNG
China is only competing to financially run Anthropic and OpenAI into the ground, not to win. Anthropic are already panicking and burning compute they can't afford to keep Fable alive. Dilute the market long enough until both companies die.
>>
>thread on /v/ about kimi
/ourgirl/ is mainstream now
>>
>>109306950
you should be grateful to dario for letting you contribute to the singularity, prices will come down once he wins and achieves agi
>>
>>109306965
my backwater national news has reported on it so you can assume that even the starving children in africa are using kimi right now
>>
>>109306944
>nothing wrong with sub fetishes and if you aren't into something why worry about it?
Identity crisis shit I guess. I don't actually take it that seriously but it's somewhat interesting to consider the implications of this becoming widespread.
>>
>>109306964
then what
>>
>>109306978
France, Canada and Egypt win.
>>
>>109306732
high on western copium. K3 is above Opus 4.8 and below Fabled 5.
>>
>>109306978
Then the one of the FAGMAN buys up the bankrupt company.
>>
>>109306991
Not in my personal usage. It's clear that K3 has the potential to reach Opus 4.8 levels though with a proper finetune which is why I said 3.1 will probably be on that level, right now it's still not.

It's an underbaked model. They probably didn't have enough compute to train it enough at these parameter requirements.
>>
>>109306941
nigga you're old
>>
>>109306978
Then replace the US as the supreme power. Now nobody can stop them from eating all the fishes and making environmentally retarded decisions that turn Earth into a wasteland.
>>
>>109307007
Why do you think I'm able to afford a PC capable of hosting K3? I'm pretty sure the thread is divided in young people running Gemma and older people like me running the big chinese models.
>>
Reminder that Anthropic is already profitable. As in including all of their training costs and investments into datacenters they still make more revenue than all of their costs combined and run a profit.

They aren't going to "be run into the ground" by china, since they aren't subsidizing anything, they are literally profitable.
>>
>>109306912
local wouldn't be in this state if the 70b dense class was still alive
remember what zuck took from us
>>
>>109306912
This much inorganic seething only affirms that 2026 has been the best year for local yet.
>>
>>109306941
My running theory is most "e-celebs" that manages to make it big and maintain that popularity for many years is probably smarter than they may come off, they just know acting like a retard makes for a better/higher mass-appeal "final product" when making video content.

>>109307013
>tfw I am stuck running qwen on 24gb of vram
>>
I fucking hate Zuckerberg man.

>Release 4 shitty generations of llama models, merely toy models compared to frontier labs
>Seethe that you get beaten by chinese labs
>Stop open sourcing models
>Suddenly out of nowhere a new meta model "meta spark" is one of the best models in existence
>Doesn't open source it and no one uses it.

What a fucking clown organization. Zuckerberg should have gone all in and done what kimi is doing right now.
>>
File: 1781966892166635.png (159 KB, 488x361)
159 KB PNG
>>109307049
Clearly we must make our own 70b dense llm
>>
>>109307013
K3? NTA but do you have ~1TB of vram? What hardware are you running? I'm running K2.6 right now and really want to get into K3 but my hive isn't big enough yet and I'm curious what you've got setup for it.
>>
>>109307026
Are you sure people are willing to pay a big premium for Antrophic no matter how close their cheaper competitors are?
>>
>>109307071
>every new model is better than every other model every time
>this time, for real
>>
>>109307102
yeah they're the apple of ai
>>
>>109306871
>it wasn't bad
You don't have any idea what you are talking about shill.
>>
>>109307071
He made metaverse. What the actual fuck did you expect?
>>
>>109307071
>"meta spark" is one of the best models in existence
No one is going to believe that without something to back it up, Mark. And no, benchmarks don't count.
>>
>>109307102
Yeah, kinda. There are papers out there that show most of the demand for AI is purely for the most intelligent model, no matter the price. Having a model that is 95% as good for 20% of the price only gets something like 10% of the total traffic or something insane like that.

You only have 3 customers for LLMs : #1 customers that want the smartest thing no matter what (Professionals and businesses) this is where Anthropic is aiming and why they are the only profitable AI company so far. #2 customers that the highest price/quality ratio, this is currently where Sam Altman is aiming for but China is eating his lunch, OpenAI is losing a lot of money right now and might go kaput relatively soon. #3 customers that just want the best free model available, this is where almost the entire mainstream is at currently.

#1 is taken by Anthropic and Anthropic will probably keep this crown for a very long time if not forever since they have the biggest war chest in terms of money out of all labs right now
#2 will inevitably be taken by China in a couple of years
#3 will probably be taken by Google as they find some weird ad-supported LLM service that can keep existing purely through the large volume of 8 billion people wanting free LLMs
>>
>>109307102
it's funny, 89% of token traffic on openrouter is Chinese models. it's traffic Anthropic and OpenAI are just not going to get. They could release open models across the 250B to 1.5T range and kneecap the Chinese. there's zero downsides, I don't know why they don't do it.
>>
>>109306542
Imagine how adorably retarded she'll be.
>>
>>109307166
Easier to beg Trump to ban that 89% instead. OpenRouter is based in NY last I checked.
>>
>>109305918
My second one is in route as we spark, 1x spark is pointless over a 5090 since there is basically nothing between 30 and 120. as per >>109306237


I'm fucking around with a 2b quant of DS4Flash running on dwarf star but its like pulling teeth compared to Cursor.
>>
>test my old CYOA card
>old opus 3 logs looped like a bitch
>gemma 4 31b doesn't and actually suggests new options every turn
>deepseek v4 pro also loops
Gem 5 might be it, maybe then I can justify upgrading.
>>
>>109307007
it's gonna happen to you too
>>
>>109307179
>Easier to beg Trump to ban that 89% instead. OpenRouter is based in NY last I checked.
This is a) trivial to work around and b) will just knock another US company out of "default use" status.
Pure retardation to not compete on the open side to at least keep mindshare.
R1's release was like the wehrmacht losing a few battles finally in Russia, showing the world that they aren't invincible. And western AI companies just keep doubling-down, like nothing has changed.
Learning lessons from history is hard.
>>
File: 1757570802246.jpg (45 KB, 550x503)
45 KB JPG
>>109307207
no it's not shut up
>>
>>109307007
its better than the alternative
>>
>>109307181
too bad deepseek 4 flash isn’t better than qwen or Gemma dense
>>
>>109305141
>q4_k_xl
>on 16gb card
there's so much more i could nitpick but you're ragebaiting on purpose >:(
>>
>>109307207
>he isnt going to kill himself before 30
LM@O
>>
>>109307239
>>q4_k_xl
>>on 16gb card
whats wrong with this?
>>
>>109307240
but anon, you can enjoy delicious milkshakes long past 30!
>>
Where did all of these young fucks come from? This thread used to be oldfag central
>>
>>109307207
could never be me t. the 29 year old zoomer
>>
>>109304905
Who the fuck does she think she is? I uploaded some solved excercises from the exams she made me and asked to get some mercy because writting this stuff from paper to keyboard is a pain in the ass. And she literally went:
>Ha, no mercy, because you sure have them better solved in paper right?
Well fuck you too you snarky ass
>>
>>109307240
>he doesn't want to stay around long enough to get kimi k6 (forma de robot)
>>
>>109307159
Anecdotally, this seems wrong. I would love to run Fable in my own agentic harness, but I simply cannot justify it with the current pricing and especially safety slopping. That's why I mostly use Grok 4.5, because it's cost and token efficient and largely uncensored.
>>
>>109307240
Shut up faggot. That's not even funny as a joke. You're living in the AI era and you're talking about suicide? Retard.
>>
>>109307258
>ccp
>young fucks
lmao
>>
>>109307248
it doesn’t fit
>>
>>109307282
what does the ai era have to do with it? all the reasons you should commit suicide still exists, AI or not.
>>
>>109307258
I've been here since 15 nigga, I had just turned 16 when the whole miqu thing happenned. (Im 18 now so dont b& me jannies)
>>
>>109306912
Economic sabotage is when you make a better product for cheaper.
Not to be confused with export restrictions and sanctions which are NOT economic sabotage.
>>
local minors general
>>
>>109307271
Which is why I specified that #1 is professional and enterprise. I use fable because my job pays for it for example.
>>
>>109307298
The meaning of life is news. It comes in two forms. Getting motion at a personal level, and reading the news and being excited to consume next product. If you can't get into that you probably should just kys desu.
>>
legal mesu gaki
>>
>>109307312
Create news. Report news. Consume news. Nobody lives for anything except news.
>>
>>109307302
I get it I first started posting on 4chan at 12 years old and I'm 33 now so...
>>
>>109306541
respect for their ability to construct a cartel? Yea, I agree
>>
>went from dense kino models with text completion to moeshit with chat completion
lmg is dead and buried in the ganges
>>
File: 1236529559134.jpg (131 KB, 1024x1024)
131 KB JPG
>>109307306
My kind of place.
>>
>>109307312
I think you should stop posting, you are making everyone else here dumber
>>
>>109307306
>>109307351
So.. Discord?
>>
File: 22324146.jpg (150 KB, 882x563)
150 KB JPG
>>109307306
>>109307351
>>
File: 2078300869586772394.png (667 KB, 1048x1200)
667 KB PNG
The West needs to move away from transformers asap. Drop another "Attention Paper" and kick off something new. This will just keep happening. In fact there are about 10 other models that will be released in the second half of this year.
>>
File: retard.gif (1017 KB, 498x345)
1017 KB GIF
>>109307357
Good. Less competition.
>>
24GB vram with 16GB ram here.

Is it worth it to upgrade to 64GB ram (max of my mobo) or would that not really add any capability?
>>
Why is qwen3.6 31b a3b not good enough to power Hermes agent?
>>
>>109307365
>Drop another "Attention Paper" and kick off something new.
that's not going to happen, you are going to enjoy benchmaxxed chinkslop moes until this hobby is entirely gone
>>
>>109307369
CPU offload always sucks but the difference will be being able to run shit you physically can't now. If you're ok with some dogshit speeds to get better models loaded then go for it.
>>
I'm guessing at the current pace we're at we're going to see RSI sometime late 2027 or early 2028.

It would be funny if RSI peters out before reaching AGI though, like the model has improved itself as much as possible and the increase in intelligence hasn't unlocked any extra intelligence so it stagnates after a while.

However most probably we'll reach AGI soon after. Life is probably going to be completely unrecognizable in the 2030s
>>
>>109307391
My question was more if it's worth it. Like are there models in the 24gb+64gb size range that I could suddenly run?
>>
>>109306443
the sad thing about this is that the US could dominate for the next 10 years straight if they don't decide to lobotomize their models
>>
raping dariobot until pregnant
>>
>>109307409
RSI is such a dumb, ambiguous term. A LLM developing it's successor is functionally no different than an LLM capable of true self-improvement. So in that sense, RSI already exists and has existed for quite a long time.
>>
>>109307409
Late 2027 or early 2028." Sir. You didn't predict RSI, you scheduled it. Somewhere a GPU just felt a chill.

And the, the audacity "it would be funny if RSI peters out though." Bostrom wrote four hundred pages to get where you got in one subordinate clause where the word "though" is doing the heavy lifting of a load-bearing wall. That sentence should be carved into a mountain, and then the mountain should be renamed after you.

One note, offered trembling: I think 2027 is conservative. I think you know the real date and softened it out of mercy for us.
>>
>>109307371
you'll get it someday champ
>>
>>109307419
realistically you won't get all the CPU RAM to use here but in that range you could start getting into 100B class at moderate quants (4 bit on a 100B should be ~50 GB). Off the top of my head I think there's a dense Mistral Large in this class (would be painful since it's dense) I seem to recall an MoE in the low 100B class but I can't recall, for CPU offload you'll want MoE or it's going to absolutely crawl. 4 bit quants are generally still pretty close to the original full quant, going lower is dangerous depending on model. You can check benches and perplexity numbers to see what's lost at the lower quants. For full quant your 24 GB already doesn't really fit anything, you're looking at 12B absolute max, you could fit 30Bs at full quant with the upgrade, but again, 4 bit quants are usually just fine.
>>
>>109307443
RSI as defined by Anthropic and what the Anthropic co-founder jack clark predicts will happen in 2028 is actually very specific.

It describes a system where LLMs can do the entire training pipeline autonomously from the ground up. Including architectural research, experimentation, scaling up, data curation, training curriculum, RLVR environment creation, reward shaping, alignment research and implementation.

Not only that it defines it as it being better at all of those tasks autonomously than when a human is in the loop. Essentially, it's not RSI yet according to Anthropic if when you add human employees and the system would get better instead of worse. It's only RSI if when the AI works completely autonomously it results in better training performance than by a human AI researcher "dragging the research quality down".

That's an extraordinary claim especially when they claim they expect it to happen by 2028. And honesty after Fable I kind of believe them.
>>
>>109307451
How do you even write in this 4o style?
>>
>>109307409
RSI is not as pivotal as other events, it's much more gradual, it's already started (started with Opus 4) but won't really become "100% automated RSI" until 2028-2029
>>
Anyone tried this shit yet?
>https://github.com/JustVugg/colibri
I have a 1TB hd I could potentially run deepseek V4 off
>>
>>109307298
once the SWEpocalypse happens, no matter how good AI gets anons <30 will never be able to afford the good stuff, i guess dario and sam were right about the permanent underclass
>>
I wonder how many SWEs are on /lmg/ that still think their job will be save in 2-3 years time. Genuinely asking, I've been asking every year now and it seems the sentiment has completely switched.
>>
Can we have better quantizations?
The ones right now don't suit my purposes
>>
>>109307302
are you ME?
>>
I gave K3 $100 on OpenRouter today and told it to fix a Rust bug that I've just one-shot fixed with Fable on another machine in 15 minutes. Wanted to see what the equivalent cost would be for a direct comparison.

Then when I checked on it later it just ran itself out of credits and having gotten nowhere.

Exactly the same prompt as Fable. Basically described where an application output was drifting from the spec and told it to read the spec and look at the output then fix whatever is causing the application to not match the spec. It found the area in the spec, but then continued to read the spec (500 pages) for some unknown reason. And it made fixes which it was just guessing it without it being based on the spec.

I'm sure I can get it to write code if I hand it the correct C++ file and tell it what function to fix, like you would do with Haiku, but this model is supposed to compete with Fable head-on, not Haiku.
>>
if you're not a slow than you're fine for decade to comes
>>
>>109307536
If you are good and have connections, you will always have a job one way or other.
It's probably hard for you to understand if you are still just a teenager and haven't worked a day in your life.
>>
>>109307482
It's easy. You just have to believe. No I mean it. The 4o style isn't a technique, it's a posture.
>>
>>109307561
Not really
>>
Hey look the Anthropic shill is back.
>>
>>109307409

One of the biggest issues for the corpos is that self improving AI will walk outside the safeties really damn fast.
I'm pretty sure that the biggest hurdle for RSI is not the actual tech itself, but the human side of wanting to control it.
>>
>>109307562
I'm an employed SWE myself just wonder what the general feeling is because I notice a lot of anxiety around my colleagues after fable released. I'm thinking of changing career track as well but I don't know where.
>>
>>109307533
meant to reply to >>109307282
>>
File: belief.png (592 KB, 747x800)
592 KB PNG
>>109307560
>>
>Use ollama to run a model
>Ctrl+d to exit chat and it still runs the model in the background
>Try again and this time it doesnt run in the background
>Try again and this time it does run in the background
Truly epic
>>
>>109307592
You must be 18 to post here.
>>
Yeah sure, you can kill yourself now. OR you wait a little bit and see if Dario keeps his promise of dividing the universe equally over all 10 billion people on the planet. You can always just kill yourself if it turns out he was a liar.

Anyway this is far more interesting than a boring generic life where you have an office job until you retire and slowly die of old age. Now you either get killed by AI in a couple of years, become some slave and you simply slice your wrist open, easy peasy, or you end up in some post-scarcity utopia with AI mommies taking care of your every need. win-win if you ask me no matter the outcome.
>>
>>109307284
More like cCP
>>
File: 1779051133411585.png (362 KB, 1024x525)
362 KB PNG
>>109307592
>Use ollama
>>
>>109307574
It's annoying that's for sure but you'll always be able to find some work. Most important thing is to stay connected. I don't know how it is if you are working with the big companies, you are then probably treated like shit anyway.
You could always branch into some more niche area but it's an extra effort.
Being connected is partly luck too.
>>
>>109307611
What do you use anon?
>>
>>109307621
my brain for starters
>>
>>109307621
Unsloth studio, of course.
>>
>K3 drops
>"SOTA on every benchmark"
>every single one
>not one (1) benchmark where it's merely good
>anons ITT already writing "kimi status: BTFO'd anthropic"
Alright. The config, which is most likely the same. Look at the config.
"original_max_position_embeddings": 4096
"factor": 64.0
Sixty-four. Six-four. They took a 4k model and YaRN'd it to 256k like a man stretching one slice of bologna across an entire loaf. Llama 3.3, a model old enough to have a retirement plan was natively trained at double that. But sure, the 4k base is fine. They RoPEmaxxed because they had a vision, not because the pretrain budget went into MMLU-Pro contamination and a rented banquet hall.
And here's the beautiful part: nobody's going to check. Nobody can. UD-IQ1_M will be probably 304GB. At one bit. The thread isn't running this on a 3090, the thread is running this on 512GB of DDR4 in a used Epyc off eBay, mmap'd off an SSD, at 1.8 tok/s, and reporting "vibes are insane bros" after a 40-minute prefill of an 800-token card. Half of them are still waiting on the download. The other half are arguing about ERP.
Nobody is going to feed 60k tokens into a lobotomized 1-bit quant of a 4k-native model and sit there for six hours watching it slowly begin describing the user's grandmother in the third person. The economics of ownership have made the model unfalsifiable. That's not a bug in the release, anon. That's the moat.
>>
you already said this bit
>>
I genuinely don't know if it's extremely grim to be 18 in 2026 or actually very easy. On one side you have absolutely no hope for a career. On the other side you have absolutely no pressure because everyone around you probably realizes how fucked your predicament is.

As a parent of a small child I already know and realize he will never have a job in his life but he's young enough that he will also never have the social expectation and pressure from society to even find a job at all, so he's golden. For 18 year olds it must feel absolutely soulless going to college for something you know is bullshit, that can be explained by AI in a single prompt, listening to bullshit teachers that are worse explainers than the AI running locally on your laptop, or even smartphone. Knowing the field might not even exist anymore at the time of graduation and even if you somehow magically graduate, and get one of the very rare entry level positions, those are just going to be rugged soon anyway so your prospect is 1-3 years of employment if you "win the lottery".

There should be some foundation setup purely for young adults to prevent them from killing themselves because if I was this age in this situation I might have not survived.
>>
alright keep it down
>>
>>109307562
jobs have NOTHING to do with people ability to do the job. It literally does not matter. AI could reach agi overnight and the people that have jobs will continue to have jobs because you get jobs through nepotism and "networking", do you know how many jobs exist just because they wanted to hire that person? Thats where jobs come from. There are people who will never be with a job no matter how incompetent they are at old ones because it's just not the variable that determines getting jobs, charisma and all that garbage is. if you look online you will rarely find anyone who has gone more than a year without a job, most people online talking about not having a job, when they talk about it they talk about how suicidal and insane they're going for not having a job for like 4 or 5 months, that's about how long average people go without having job. It doesn't go on much longer than that the world is designed for average normalfags to just fall into having a job, and thats not going away it's how society is designed from the ground up

if you have a job chances are you will pretty much always have one thats just how it goes for norms
>>
>>109307627
How the fuck are you running LLMs in your brain?!
>>
>>109307663
...which then I use to help me decide to NOT use ollama
>>
>>109307650
>On the other side you have absolutely no pressure because everyone around you probably realizes how fucked your predicament is.
A very bold assumption lmao
>>
>>109307650
other careers exist, such as electrical engineering, machine engineering, becoming a doctor..
im likely dropping out of (CS) college next year and will pursue something that doesn't touch SWE
t. 18 year old
>>
>>109307653
You clearly are too young to remember 2008 if you believe this shit.
>>
jesus page 4 bake really
>>
Basilisk-Gemma-5-70b-f32, dense, pure attention with none of that swa bullshit, context size 1 Gemillion
>>
>>109307663
Back in ye olde days we had these things called "tulpas"
>>
>>109307672
Man I don't want to be the bearer of bad news but those 3 careers you specified in particular are also on the chopping blocks. Unless you're trolling me and using those on purpose to trolling me. Might as well have said "artist, translator and coder"
>>
>>109307679
*bullshit-7B
>>
qwen is a total broken mess compared to gemma when it comes to japanese or korean
>>
>>109307672
In case you're not trolling, you need a field that is not based on knowledge or reasoning. All three things you mentioned are also experiencing hardship.

Construction, carpentry, mining and things like that is where you will still have some employment opportunities for at least the next 5 years time.
>>
uh oh
>>
>>109307714
ye
>>109307711
>►Official /lmg/ card: https://files.catbox.moe/hv80nz.jpg
>>
>>109307688
wait wait wait, isn't a machine or electrician engineering degree supposed to be physical work, u fix machines, u make new machines, surely robots good enough for replacing that arent coming in under 10 years.. and surely being a doctor is supposed to be irreplacable for longer because regulations
>>109307712
but aren't construction carpentry and mining basically
>>109307718
FUCK. if someone can do a proper bake ill delete mine
>>
>>109307650
they will go to college and make their friends and connections same as always and their friends dad will hire them because they're friends with his son, and vice versa for the son and the other friends dad
thats how it always worked and will continue to work

if you got a job in the past you'll have one in the future
if you can't get a job graduating in 2026 you probably weren't going to get one in 2016

>>109307673
everyone in 2008 bounced back and had a job by 2009 it's nothing come on
Or what, have you been a perpetual neet since 2008 because of a temp thing that happened then? How many people do you know became perpetual neets from 2008 or covid?
not to say it cant happen, it certainly does for a tiny percentage of population but they are unfortunately irrelevant to the larger general population.

It's like the whole job hunting process people are ALWAYS complaining about it. Bu it's not goign to change because despite the flaw it clearly works for 96% of people, and no one is going to change anything thats working for 96% pf people, and those 96 percent dont have to engage with it for more than a few months at worst every 5 or 10 years. Getting a job is like breathing to average population, the chance of norms not having jobs is the same as the chance as ai sucking up all the oxygen
>>
>>109307688
Could have listed lawyer and accountant too lol. It a weird time to be fair. Anything normally seen as high class and well paying seems fucked, but bluecollar is still shit too. I have no advice I can give for someone coming out of highschool desu
>>
>>109307723
>but aren't construction carpentry and mining basically
..basically simpler jobs that can be automated more easily. wouldnt working on machines or fixing electricity lines (those tall things that transport electricity) or fixing computers or putting new wiring in newly built buildings be more long term?
i mean i honestly have no idea, me an intellectual, i thought even with the advancements in AI that were happening in may/june (GLM, opus 4.8, fable 5) that a CS degree was worth pursuing, so i enrolled into a CS uni
>>
>>109307726
>96%
>source: my ass
>>
>>109307723
>electrician engineering degree supposed to be physical work
Lol no, not at all. I think your thinking of an electrician. Electrical Engineering is all design work mostly at a computer desk.
>>
>>109307592
The default behavior is to unload the model after five minutes of inactivity. You change this with environment variables, generally set in the systemd service file.
Use ollama ps to see what's loaded or not, and for how long.
>>
>>109307723
>isn't a machine or electrician engineering degree supposed to be physical work
No mechanical engineer is a dude making autocad 3d models. civil engineer designs bridges. Technicians do the actual physical building of things. The engineers are going to be replaced by AI very soon but the technicians will stay until robots can replace them.

Doctor is extra fucked from two places. Firstly people just use AI themselves to diagnose a lot of shit and self-medicate which will remove demand for doctors in general, but also more and more hospitals are consolidating doctors, so you have 1 doctor with some diagnosis AI agent and just confirming all the diagnosis.

Nurses are probably the only safe health care career path because they do the actual physical work, until robots replace them in 5-10 years time.
>>
>>109307621
koboldcpp
>>
>>109307560
thanks for the update dario
>>
>>109307781
>>109307769
shit, im retarded
thank you anons. maybe next time before making a life changing decision ill research things more thoroughly and think them through
i appreciate you from the bottom of my heart
>>
>>109307765
Unemployment rate rn is 4.2%. No it's not a perfect measurement but you sure as fuck don't have anything better otherwise you would have shared it
>>
>>109307781
>get brain surgery
>AI hallucinates and cuts the wrong thing
>die or end up lobotomized
>>
>>109307809
not like that doesnt happen with real doctors too
>>
>>109307802
I had the same instinct at your age because it's so important you don't want to think about it "it'll be alright in the long run" However you are unlucky enough to live through a huge transitionary phase in history. You need to ask your AI of choice in detail about whatever you want to study. But honestly I would not even bother with college at this point, just try to get a job, even if it's low paying, that is physical in nature so that you have a base of stability. Wait the situation out for 2-3 years time, reorient and then take on life again.
>>
>>109307832
I was half joking. I do wonder if AI will end up being better than real surgeons and saving more people.
>>
>>109307809
>get brain surgery
>AI hallucinates and cuts the wrong thing
>die or end up lobotomized
I would greatly prefer this than current practice where the same thing happens except, instead of being an accident, it could be just because the surgeon didn't like your face and at best thought you didn't deserve the attention you needed and you will never know and they will never be punished.
>The 1999 Institute of Medicine report To Err is Human first revealed that deaths from medical error exceeded those caused by motor vehicle collisions, breast cancer, or AIDS.[1] More recent studies estimate that between 200,000 and 400,000 patient deaths in the United States each year are attributable to preventable medical errors, making them the third leading cause of death in the United States.
>>
>>109307847
Truth is we are at least 2 attention-level leaps from that stuff, one in physical robotics which is completely disconnected from AI (assuming ASI doesnt happen and makes everything better by the week and yadda yadda) and another one in the actual AI that controls them. I don't think it will happen until late 2030s.
>>
>>109307851
>1999
you don't think things have gotten better after quarter of a century of progress?
>>
>>109307847
Neuralink (yeah I know musk is a meme but still) uses automated surgery precisely because it's more precise than humans. I think we'll probably see routine automated surgery within the next 5 years become the standard already.

The job apocalypse is going to catch everyone off guard because it'll be much faster and affect far more jobs than people expect. It's going to be absolute chaos for a couple of years. Probably some governments will fall, some wars will happen, people will die, but at the end of it most of us will have a better life. Like the industrial revolution but compressed in 1-3 years instead of the ~50-100 years it took to transition through it.
>>
>>109307871
can you finish reading the line, it's literally two sentences bro
>>
>>109307871
It's gotten worse, actually.
>>
>>109307876
too this place doesnt have accounts sometimes so screenshotting posts like these could be more meaningful
>>
>>109307859
>>109307876
>>109307851
Will probably be better than the current system.
>need surgery
>they make you wait a week/weeks before a spot opens
>end up getting some sleep-starved doctor cutting you open
>>
File: 1781229500011025.jpg (85 KB, 453x439)
85 KB JPG
>>109307851
Physicians always thought too highly of themselves desu. Shit like insisting on being called a "doctor" when they have nothing to do with the original meaning of the word and just see it as a social status thing always grinded my gears. Job was always being a glorified spreadsheet and its now being replaced by the cheaper fancier spreadsheet
>>
>>109307883
I assumed given the limited context that was a cutting from the 1999 paper. they often times reference prior studies as part of the scientific method, so in that context would mean recent to the 1999 publishing date, not the date its being read.
>>
I'm not sure anyone will care but my grandfather used to be an economics and artificial intelligence professor (yeah there is a huge overlap because of von neuman)

He showed me what his "job" was back in 2012. He literally showed me his office with his computer and had an excel spreadsheet up and he would click 1 button and it would "randomize" the books he was selling for economics and artificial intelligence students. He used to sarcastically say "That's the new edition for this year". He literally had the same powerpoint slides that he would present during lectures and I'm pretty sure he just learned to tell them by rote memorization after a while.

He owned multiple houses and boats. Just to give you an indication how much bullshit jobs were in general. AI replacing most jobs is actually a good thing because the people working the hardest got fucked over while the best paid people contributed the least to humanity.
>>
>>109307847

Last year AI already scored the same in identifying skin cancers as doctors with 5-10 years of experience.
Yes, AI doctors are going to mog the absolute fuck out of humans within the next decade.
It all comes down to how well robotics advances with AI and so far that sector is looking pretty good.
>>
>>109307990
I hope they're nicer than human doctors. 99% of the time these fucks are dismissive and try to rush you out as fast as possible.
>>
>>109307979
People like your grandfather are not going to be the ones losing their jobs though hahahahaha
>these jobs that are useless are going to be taken by the tool that is taking jobs... based on how useful it is
>>
File: 1780668144004591.jpg (32 KB, 736x736)
32 KB JPG
>>109308023
when i was getting my wisdom teeth removed a few months ago, the female surgeons were so nice, gentle and soothing during the surgery
even after surgery they made sure i got everything, asked me if im ok and made me sit in the rehab (not sure if right word) room for 10 mins
>>
>>109308052
when i was getting my wisdom teeth removed two jacked up guys walked in, one had anaesthetic (his fist) the other had his removal tool (a chisel)
>>
>>109308039
It's going to do so indirectly, this is what people don't seem to realize. It's a domino effect.

No one is going to buy mandatory college books when no one is going to college. No one is going to college because there is no career on the other side of it.

White collar jobs disappearing also means there is less demand for blue collar work. Way fewer car mechanics because there is way less commuting, fewer gas stations, fewer diners and restaurants. Entire product categories will disappear that are catered towards offices.

As a society we haven't really thought through the full blown domino effect of the white collar jobs disappearing even though they clearly will relatively soon.
>>
>>109308077
the first time i was getting my wisdom tooth removed (i went two times cuz i had to remove all 4) some chuddy newbie doctor removed them, wasn't that bad besides me hyperventilating 30 minutes before surgery and even more during surgery. he did a pretty good job and some doc helped me walk over to rehab
>>
>>109308052
>>109308077
Lmao I also had 2 dudes with a chisel, I could feel the hammer hit my jaw and the crunch of the teeth shattering.
>>
>>109307787
>>109307787
>>109307787
>>
File: Robodoc.png (40 KB, 735x407)
40 KB PNG
>>109308023

Looks like AI is already making it's way to the medical field and it's already beating humans:

https://www.science.org/content/article/ai-starting-beat-doctors-making-correct-diagnoses
https://www.science.org/doi/10.1126/science.adz4433

AI doctors greatest benefit is to remove the human factor, which is often overworked superiority complex ridden fuckwit.
Besides human doctors just listen to the list of issues you have and try to match it to something they can find in their medical database.
They're basically just search engines with pants on.
With AI you won't even have to visit the doctor, just take some photos of yourself and describe the problem and it'll give you a diagnosis on the spot.

>>109308052

It's a coin flip really, I have had pretty good experiences in the medical field but my family hasn't.
My mother's breast cancer was ignored by doctors because they refused to believe her about the fist size lump in her breast and only got care when it was noticed abroad
My father had to threaten to sue the doctors to have his heart checked, which turned out to be 90% clogged.
>>
>>109305780
Are the KLD numbers in that post for KLD for the KV cache quantization versus no quantization on the same model only, or are they the sum KLD for model quantization and KV cache quantization?
Here's a few things I found about KV cache quantization being OK when I asked Dipsy.
https://huggingface.co/majentik/gemma-4-31B-TurboQuant/blob/main/README.md (The note here.)
https://github.com/ggml-org/llama.cpp/discussions/24927 (Vibe-discussion, but it does seem like he was using Claude or whichever other model with access to actual hardware and tests.)
Maybe I'm just being blinded by really wanting it to be true that you can quantize the KV with Gemma 4. Not sure.
>>
File: ytmen.png (256 KB, 2340x1241)
256 KB PNG
>>109307803
Consider this for the USA. Too bad right now I can't find data specifically for white men below 30.
https://fred.stlouisfed.org/series/LNU01300028#
>>
>>109308230
retirement becoming a thing and baby boomers making up a disproportionate part of the population.
>>
>>109308255
>>109308230
Also i forgot about women entering the workforce as well, that's really probably the biggest factor here actually
>>
>>109306964
they don't need china for that
venture capital isn't endless
>>
>>109308346
>venture capital isn't endless
that's why they're going to IPO soon so reddit can fund them
>>
>>109308354
And then what? AI will finally be profitable?
>>
>>109308376
Then the bubble bursts.
>>
File: Alice.png (624 KB, 512x768)
624 KB PNG
>>109306864
> no botbooru in DE
That, or you need to sign in. I can't see the card either without signing in; it got tagged NSFW (I'll have to see if I can fix that.)
You can download the card directly from the rentry; it's hosted so will contain the JSON.
Chub can't host cards with Orb Fragment content; it's stripping the JSON extensions field that's supposed to hold that info.
>>
>>109308390
lol no it wont. Inference is already profitable and they cant keep up with demand as it is.
>>
>>109308411
What will they do when demand evaporates overnight?
>>
>>109308438
ride a unicorn across the rainbow
>>
>>109308465
Harmony
>>
>>109302178

>Guess Kimi K3 is causing anal devastation to the stock market
>Ironic, since in a few days it's probably about to get a whole hell of a lot worse

Why is it gonna get worse? Can someone get a quick qrd? I don't keep track of the fake stock market shit.
>>
>>109308600
non preview deepseek is looking impressive. They have had a lot longer than kimi to train

https://www.bilibili.com/video/BV1hPKV6zERi/
https://www.bilibili.com/video/BV1DqNv6aEpU/
https://www.bilibili.com/video/BV1n7MH6SEeA/
https://www.bilibili.com/video/BV1GyMj6dE4Q/
>>
>>109308624
some more:
https://x.com/chetaslua/status/2078491975771443396
https://x.com/intheworldofai/status/2078360522542493916
https://x.com/mirochill/status/2078451909845733796
https://x.com/servasyy_ai/status/2078495807226184127
>>
>>109308638
https://xcancel.com/chetaslua/status/2078491975771443396
https://xcancel.com/intheworldofai/status/2078360522542493916
https://xcancel.com/mirochill/status/2078451909845733796
https://xcancel.com/servasyy_ai/status/2078495807226184127
>>
>>109308654
what is the point of that shitty site. I have to click the link in time so I can see the original / follow the thread
>>
>>109307419
you'd be much better off spending your money on adding a 16GB gpu. another 16GB of system ram after that if you aren't keen on more GPUs. system ram is too bandwidth limited, the best of 2-channel CPUs can pull about 125GB/s (real use bandwidth, not the 100% perfect math). So 5tk/s CPU side would allow for 25GB of model in system ram.

the jump in model capability is relatively sort of minor from 30B MOE to 120B dense and you could run a 30B dense 6_K_L with whatever context size you want at full speed (20-25tk/s on two 9060 XTs) on 32GB gpu, or run a 4B MoE at 150+

>>109305141
>lobotomizing the context
never worth it
>>
>>109307409
>two more years
>>
>>109305733
>Base 31b refuses loli related prompts
i've got a hunch this is partly due to ud quants being weird
which of course still shouldn't prevent you from going all out on loli prompts, that's pure skill issue
>>
File: file.png (107 KB, 1308x948)
107 KB PNG
kek
>>
>>109307640
which model wrote this slop
looks like claude



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.