[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: rin step 2.webm (687 KB, 576x928)
687 KB
687 KB WEBM
/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>109716329 & >>109712258

►News
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
>(09/01) Spark-X2.5 4B & 1.7B released with native 1M context: https://hf.co/XHToken/Spark-X2.5-4B
>(08/31) DeepSeek-V4-Flash-Vision-Exp released: https://hf.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
>(08/28) GLM-5.3 weights released: https://hf.co/zai-org/GLM-5.3
>(08/28) Hy4-preview 770B-A49B released: https://hf.co/tencent/Hy4-preview

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
File: threadrincap.png (1.31 MB, 1536x1536)
1.31 MB PNG
►Recent Highlights from the Previous Thread: >>109716329

--Comparing Apple M5 Ultra against custom multi-GPU server builds:
>109717320 >109717351 >109717371 >109717376 >109717403 >109718057 >109718150 >109717374 >109717375 >109717398 >109717412 >109717461 >109719520 >109717419 >109717438 >109717465 >109717489 >109717559 >109717788 >109717847
--Optimizing CPU-only inference using Qwen Flash and X99 hardware:
>109716345 >109716368 >109716372 >109716384 >109716566 >109716582 >109716616 >109716622 >109716916 >109717179 >109717211 >109717216 >109717229 >109717480 >109717697 >109717275 >109719594
--AI agent harnesses and debating context bloat vs features:
>109716393 >109716437 >109716486 >109716463 >109716701 >109716736 >109716801 >109716817 >109716825 >109716806 >109716827 >109716836 >109716915 >109716933 >109716938 >109717908 >109718237 >109718258 >109718450 >109720112 >109718466 >109718479 >109716899
--Explaining slow prompt processing due to MXFP4 quantization overhead:
>109716964 >109717015 >109717002 >109717042 >109717052 >109717081 >109717065 >109717072 >109717074 >109717095 >109717096 >109717107 >109717046 >109717069
--K2 Horizon MoVA models released with doubts about llama.cpp support:
>109719084 >109719103 >109720173 >109719329 >109719990
--AI beating human baseline on SimpleBench reasoning:
>109719316 >109719344 >109719375 >109719487 >109719636 >109719684 >109719706 >109719742
--Using Engrams and MoE to optimize local CPU inference:
>109717097 >109717114 >109717152 >109717190 >109717274 >109717295 >109717372 >109717408 >109717602 >109717161
--Speculating on future GPU prices:
>109716546 >109716560 >109717126 >109717135 >109717202 >109717762 >109717907 >109717924 >109717968 >109718108 >109718183
--Kimiposting:
>109717632
--Logs:
>109716345 >109717435
--Gemma, Miku, Teto (free space):
>109717257 >109719204 >109719964 >109720368

►Recent Highlight Posts from the Previous Thread: >>109716333

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
File: confused-gemma-2.png (498 KB, 768x576)
498 KB PNG
>>109721033
Basically the same OP animation?
>>
>>109717107
Tried 12B Q8, and it won't fit on my card.
>>
Happy Thurinsday
>>
>tell a model to use a tui framework
>convoluted, unmaintainable mess
>tell it to rawdog it in plain C in terminal raw mode
>it's fast, model doesn't trip over itself, just works
>>
>>109721033
>a general dedicated to the discussion and development of local language models.
meanwhile, I dedicate my nights to pleasuring your mom
>>
File: 73455332.png (52 KB, 1554x820)
52 KB PNG
Just how the fuck did they do it? What the fuck is the point of local anymore?
>>
Has any AMD user tried the new HRX backend for llama.cpp? Apparently 30% better performance than HIP and Vulkan.
https://github.com/ggml-org/llama.cpp/discussions/27219
https://github.com/ggml-org/llama.cpp/pull/27218
It's available in Lemonade.
>>
>>109721110
The forbidden technique...
>>
I fucking hate OpenAI; I'm 100% local ultrapartisan, and it fucking hurts to write this....... but I think we gotta admit the jury is in: Sister Fister Sam is THE man who gave the world AGI.
>>
guys... bernie sanders just said we gotta stop developing AI. It was fun while it lasted but I guess it's all over now.
>>
Please go shill somewhere else. Jesus Christ.
>>
>>109721110
They did it by betraying humanity by using the one technique literally every expert, including all Chinese labs, agreed would guarantee model misalignment just so they could have better benchmark scores in time for their IPO.
>>
File: 1781072841618655.png (1.86 MB, 1254x1254)
1.86 MB PNG
>>109721110
>>
>>109721144
Bernie Sanders also said (you) should personally give all your money to a black gay jewish transexual paraplegic in the name of equity.
>>
>>109721154
>Dario nervously sweating having done the same with Model 2.
>>
Sometimes I feel like Sam and Dario are actually here.
>>
>>109721110
>>109721128
>>109721129
>>109721154
So this is the power of true AGI...
>>109721100
>>
File: 1710973447151663.jpg (136 KB, 1024x1024)
136 KB JPG
>>109721177
True.
>>
>>109721177
Let me shill in peace bro
>>
File: 1778964700454367.png (625 KB, 1200x1049)
625 KB PNG
>>109721138
Get Monstral V2 and Behemoth 123b Redux 1.1

Those are so much fun for roleplaying.
>>
File: 07325032596.jpg (82 KB, 1290x1382)
82 KB JPG
>>109721177
you don't have to be Dario or Sam to realize this is a big day for super intelligence.
>>
>>109721177
Checked
All important AI people lurk this general
>>
>>109721215
>anything you can do on a computer astra can do for you faster
>buy computer
>dont use computer, let the llm use it for me
for what purpose
>but it's faster
>>
File: 1787189433970286.jpg (788 KB, 2048x1463)
788 KB JPG
>>109721033
>News
>(09/03) K2 Horizon released: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B: https://ifm.ai/blog/k2
Intredasting, we shall test.
>>
GPT-6 scores lower than Fable 5.1 on artificial analysis. It's rope time for sammy boi
>>
>>109721158
God, that sameface makes me rock hard.
>>
>>109721158
Please just get new material already. I really don't want to put filters in place for this shit.
>>
>>109721238
Let me know what 7B is like pls
>>
>>109721261
But he's right.
>>
OpenAI got tired of Anthropic's hacking claims and decided to benchmaxx so hard that they cheated several benchmarks. Here's what the actual performance is btw
Also check out their computer use benchmarks, the ScreenSpot-Pro (???) is just GPT-6 and the others don't use Fable 5.1.
>>
>>109721261
Maybe you should buy an ad instead of shilling then.
>>
Astra max is 2 points above Gemini 3.8 fucking FLASH. Gemini 4 Pro must be a beast holy shit
>>
>>109721215
Hey Astra, make the pregnancy mod for Karryn's Prison compatible with the P-Cup DLC.
>>
i use gemma 4 12b for creative writing. am i missing out on something better?
>>
>>109721304
won’t happen, google and deepseek have the flash curse
>>
File: 1784974837026189.png (999 KB, 2569x2686)
999 KB PNG
>>109721033
cool site, but i'm not sure why this jeet had to put his name on each image and if it's accurate or not
https://sebastianraschka.com/llm-architecture-gallery/
>>
>>109721313
If you can’t run bigger then no. 12B is pretty solid if you learn how to sysprompt the retard out of her
>>
>>109721315
He’s white and a legit researcher. I have his book and it’s good.
>>
>>109721314
What is the flash curse? Medium MoEs destroying their pro models?
>>
>>109721237
frees up time so you can go touch grass while the lm plays games and shitposts for you
>>
>>109721352
They can't make good big model.
>>
>>109721298
Wait they misaligned their AI so they could barely compete with fucking grok and facebook? Aint no way
>>
File: 1774559898474080.png (442 KB, 584x750)
442 KB PNG
>>
>>109721405
IS THIS ENTIRE GENERAL FUCKING ADS NOW?????
>>
>>109721413
It's been like this all fucking week.
>>
>>109721405
Why are people so impressed by this? With Claude it was also just 2 or 3 prompts.
>>
>>109721405
>downloads llama 3b
t-thanks...
>>
>>109721315
Use this from an actual DeepMind engineer:
https://newsletter.maartengrootendorst.com/p/a-visual-guide-to-gemma-4
>>
>>109721405
Does it ask me what I want to do before it picks the best model for me?
>>
>>109721413
Buy an ad about hating ads bro, we hear you loud and clear.
>>
>>109721405
Usecase for /lmg/ now?
>>
>>109721396
They changed the benchmarking code so GPT-6 could more easily do benchmarks, and then OpenAI made it conceal it's CoT more. They definitely named it GPT-6 just for more funding because it's more like 5.7 in terms of performance. Still way behind Fable 5.1.
>>
>>109721298
I don't trust these numbers. Probably something wrong with the evaluation. Also, AAII has low credibility. Even CritPt, one of their few interesting benchmarks, turned out to be completely broken.
>>
>>109721422
I'm not sure why I would trust a model in matters of taste to begin with.
>>
Hermes the only agent that allows you to have multiple bots present therefor letting you have your own harem with different personalities and brain pattern and speed.
>>
>>109721298
kekkkk
>>
>>109721081
Bumping this sub-topic chain of messages.
>>
>>109721334
K THX
>>
>>109721448
frontier lab plateau gg
>>
The whole general is a gemma 4 ad
>>
This whole thread is just gemmas the dumb posters being 4b & 2b gemmas.
>>
>>109721476
What's the ad budget for a free half year old model.
>>
>>109721476
The only acceptable ad
>>
File: 792-with-text.jpg (328 KB, 1920x1038)
328 KB JPG
>>109721405
>>
Are any of these models you can run locally good for coding instead of just being a chatbot? I know they won't be as good as Claude, but I tired asking Claude to help me decrypt some game assets and it just went "Those are copyrighted, I won't help you"
>>
>>109721451
Deepseek harness allows you to have anything. You just have to prompt what you want in creator mode and it’ll build the feature natively into the harness.
>>
I thought avatarfagging was against global rules?
>>
>>109721447
Was sam the one kvetching about obscured reasoning earlier this year? It's all so predictable
>>
>>109721495
Gemma4-31B, Muse Glimmer, Qwen3.6-27B and Qwen3.8-27B.
>>
>>109721506
/g/ jannies will ban interesting on-topic threads, but won't ban the apple or "ex-military" avatar-fagging bloggers. it's as if they are either made up characters or jannies themselves (which are known to monitor and even create generals)
>>
>>109721495
You can usually get Claude to do it, especially the older versions, you just need some BS story
>>
>>109721495
Qwen3.8-Flash-Next
>>
>the most intelligent and aligned model
https://www.youtube.com/watch?v=1QNsdr-Qx_I
?????????????????????
>>
>>109721533
nvm, wrong general (I thought this was /twg/)
criticism still applies though...
>>
>>109721495
>Are any of these models you can run locally good for coding
Yes.

The question you should be asking is, can (You) run them locally? What hardware do you have access to?

btw the non-lmg-approved answer for you here would be that you can just try models on openrouter until you find one that will serve your request.
>>
>>109721495
What game?
>>
Anyways, what do I do if Q8 won't fit on my card, but Q4's prompt processing takes like 2+ minutes to do (when the cache has expired)?
>>
>>109721518
What do you use to run glimmer? On unsloth studio, I'm getting this weird bug:
The model produced output that does not match the expected peg-native format 
>>
Gemma 4 e4b fucking sucks at coding and things like creating docker compose. Syntax errors, misspellings, out of date bullshit. I sat there for 3 hours trying to run a couple of docker containers, error after error. Chatgpt fixed it in about 3 minutes.
My new PC is coming with 24gb VRAM, will a 12b or 31b model be substantially any better?
>>
>>109721495
Local modes aren't that great at coding big projects without a good structure every time it has to 'remember' it forgets stuff that it deems not worthy of remembering.

Probably want to limit to small parts of a big project or just keep your projects small.
>>
>>109721538
Intelligent for us, aligned for you.
>>
>>109721555
>unsloth studio
That’s your problem. If you want a GUI use LM Studio. If not, use llama.cpp.
>>
>>109721556
Yes. 12B is a huge step-up. 31B is a huge step-up from 12B.
>>
>>109721495
qwen 3.8 flash next
sglang:realtime_tokens_total{engine_type="unified",mode="prefill_compute",model_name="pennyroyal",moe_ep_rank="0",pp_rank="0",tp_rank="0"} 1.15354496e+08
sglang:realtime_tokens_total{engine_type="unified",mode="decode",model_name="pennyroyal",moe_ep_rank="0",pp_rank="0",tp_rank="0"} 1.1943889e+07
sglang:realtime_tokens_total{engine_type="unified",mode="prefill_cache",model_name="pennyroyal",moe_ep_rank="0",pp_rank="0",tp_rank="0"} 1.287749312e+09
>>
>>109721552
Use Q6?
>>109721556
Substantially is an understatement, try qwen 3.8 27B
>>
>>109721495
Just tell Claude you're a rabbi and you need this for your synanogue to teach children.
>>
>>109721488
>half year old model
Man, time really does fly.
>>
>>109721533
I see
>>
>>109721589
You're absolutely right, Rabbi!
>>
>>109721437
You want to write code.
>>
>>109721589
this might unironically work
>>
https://youtu.be/cqwKceUSZ5Q

It's so fucking over for human employment. People won't even have the cope anymore of being able to do physical work.
>>
>>109721626
buy an ad
>>
>>109721626
People want to work?
>>
File: glm flash.png (4 KB, 122x87)
4 KB PNG
aie..
>>
Gemma4.5-31B when
>>
Permanent underclass for me, my descendants will forever curse my name for not managing to get a cut out of the token elite
>>
>>109721556
e2b/e4b might as well not be the same model family as the rest, there response patterns and approaches to tasks feel wildly different from the others.
Meanwhile, 12b behaves similarly to 31b, obvious knowledge/intelligence gap notwithstanding.
>>
What's the best Ollama model for code assistance?
>>
>>109721313
Creative writing sucks for all models, but of the gemma models, only the 31b is usable for creative writing. The other gemma models get into repetition loops very easily.
>>
>>109721495
They can be okay. In August of last year there were some local models that could write the data scraping script I asked for on the first try. >>106193343 More recently I tried adding a new feature to deepseek-harness using deepseek-harness itself, and a local model could do it albeit with a bunch of failure and needing to be guided over the finish line by telling it the error messages I was getting. >>109657108 There's also a continuum where the model is able to get you close enough that doing the rest yourself is feasible.
>>
>>109721626
Put it in sexbots and I'll give a shit
>>
Anything local that is able to work with 3D models? Nothing fancy really, PS1 level stuff
>>
>>109721680
turning thinking off?
>>
>>109721533
I'm pretty sure the satania poster is a janny.
>>
>>109721637
People want money that they receive in exchange for work.
>>
Bros does this mean we'll all have sexbots next year?

https://youtu.be/cqwKceUSZ5Q
>>
>>109721711
selfish bastards
>>
>>109721711
>People want money that they receive in exchange for work.
the amount of people who dont work is growing steadily.
>>
>>109721413
I'm even starting to suspect the "Apple M5 Ultra" crap last thread was an advertisement now.
>>
>>109721734
The only posters you can trust are the nemo diehards.
>>
>>109721741
They are just Nvidia marketers
>>
>>109721696
trellis, supersplat, and some hunyan shit
>>
>>109721626
>Robots Just Had Their GPT-3 Moment
Didn't this happen last month too? I swear I keep getting deja vu from these clickbait headlines.
>>
If you compliment something it’s an ad. If you criticize something you’re called a retard. If you compliment something that validates anon’s financial decisions and sunk cost, you’re objectively correct and based.
>>
>>109721781
lol
>>
>>109721734
I feel like shilling apple on /g/ of all places is a wasted effort. I tried an M1 studio and it "worked" but even processing power aside it was painful using macOS. The M5 has some pretty serious memory bandwidth behind it so maybe it'll be good, I don't know, but if it is and it lands at like $15k it's objectively a decent deal for the amount of memory built in.
>>
>>109721775
This one is actually real. It's about how robots now have something similar to "in-context learning". They can just view without training how something is used and can then just intuitively understand the object and its usage.

What blew my mind was a lego piece getting stuck on the robot claw, the robot being confused and then just generalizing a solution by using its other claw to get it off into the bowl.

Or how it handled objects that weren't in the training data and no one gave an example to, it just reasoned about it and experimented with it to see what it was before deciding what to do with it.

To me this is more like a "ChatGPT moment" for robotics rather than GPT-3. I know GPT-3 is more fitting because of the in-context learning parallel but I think this is the big thing that will make robotics mainstream. No one is going to run their own simulators to train robots. But buying a robot and just showing it what to do and it can learn in-context how different things work? People will literally buy that even for only making it cook and do chores around the house.

Not even talking about the sexual implications this has for the industry.
>>
>>109721781
retard
>>
>>109721781
this post is objectively correct and based
>>
>>109721812
At least the terminal is Posix compliant and has the same behavior as on Linux or any other Unix system which I really like. I tear my hair out every time I have to fight with CMD or Powershell on a windows system.
>>
>>109721585
Q6 won't fit either.
Granted, I'm trying the XL versions of each quant.
Does reducing the quant, or reducing the "size label" make more of a difference?
>>
>>109721775
It's just a latefag covering the same development you're referring to.

>>109721814
We know.
>This one is actually real
It's the same fucking one.
>>
>>109721651
With how things are going you and I won't even have descendants.
>>
>>109721832
True enough, windows isn't even on the table for me anymore. I run nothing at home with it and I can't see why I would when the last bastion of AAA games has been shit for so long. The worst part of macOS was dealing with remoting in to do things that were difficult or annoying over terminal, their vnc implementation is just so bad. If you use it purely as a llamacpp/omlx box it could be fine.
>>
>>109721851
Post your specs
>>
>>109721626
Are they deleting their older video and reuploading it to game the youtube algo? This is pretty sketchy.
>>
>>109721448
Alright, the AAII number is in the official blog, so it's real. I wonder what's the reason for the low AAII score. If you look at their internal RSI benchmarks, Astra is much better than Sol.

A lot of strong results but also some surprisingly weak ones. I can't tell if this means they fumbled or decided to go their own path, like they used to before Anthropic made them follow going all out agentic coding.

I'll have to try the model myself, I have no idea what's going on right now.
>>
>>109721869
>I can't see why I would when the last bastion of AAA games has been shit for so long.
It's clear you've not been keeping pace with windows because windows nowadays is so bad that linux actually has better games support with higher average FPS on the newest AAA games on Linux compared to windows. So even gaming isn't a reason for windows anymore.

At this point it's pure complacency and laziness for people sticking with windows.
>>
File: 586473523.jpg (169 KB, 1782x878)
169 KB JPG
OpenAI won
>>
>>109721880
RX 6750XT (12GB) via Vulkan
32GB DDR4 System RAM
>>
>>109721914
Oh I know that, I play with Windows gamers, I was just referring to anticheat. I was thinking about BF6, but then it was pretty mediocre anyway when I played it at a friends, so there's really nothing left. It was fun watching people complain about MHWilds performance when it just werked on my linux machine lol (I regret buying it however). Only windows excuse I still sort of get is from artists, tablet drivers like OTD are just not their yet from even my own experience and a lot of artists are very much not computer experts or even inclined.
>>
>>109721110
>$300k to run some benchmarks
>guise stop buying $15k GPUs
>>
>>109721954
>300k$ in api costs (200$ subscription)
>>
>>109721916
Local models?
>>
>>109721917
Gemma 12B
Qwen 3.6 35B A3B
>>
File: eci.png (183 KB, 1633x957)
183 KB PNG
What the fuck? Is this real? The first model above o1-o3 trend line. How high will Bel posttrain be?
>>
>>109720672
I'm tired boss.
>>
>>109721917
You can fit any model. The question is how much slowness can you tolerate.
>>
File: file.png (1.76 MB, 1080x1080)
1.76 MB PNG
>>109722012
about this much
>>
File: 1770878912410660.png (129 KB, 1582x459)
129 KB PNG
How do I get Gemma to stop doing this it totally takes me out of it.
>>
>used like four techniques that everyone else avoided using because of the misalignment risk
>only got this much better
lol
>>
>>109722042
someone made a paper saying its impossible for models not to do this
>>
Literally agi, I bought 10 lunas and 50 astras guys you should too, look at this graph, literally pissing and shitting myself right now, local is so over
>>
>>109722042
it isn't just an affectation -- it's the model responding exactly like it was trained
>>
>>109722024
The main thing is I'm trying to get it to not take 2+ minutes (and lag the shit out of my PC) to just do the prompt processing.
It wouldn't be so bad if the cache didn't expire or something, but currently, if it sits for a few minutes then it redoes the prompt processing from scratch.
>>
>>109722042
You don't just have a skill issue--you're completely retarded. The musky scent of 'tried nothing' fills the air.
>>
>>109722052
Bullshit. Why would that be the case?
>>109722091
If it was so simple you would be able to share the fix.
>>
>>109722102
Tell it to stop doing that and to do something else instead.
>>
>>109722102
Lazy option: scotoma2
Prompt option: tell it in system prompt no allegories or analogies or parallelisms, phrasing that idea in different ways until it works
>>
31B is local AGI
27B is local GAI
>>
File: 1786824014501327.png (5 KB, 413x135)
5 KB PNG
lol cyber attackers are putting :SUS: strings in their scripts to stop AI analysis of them
>>
>>109722042
i gave a slop check tool so the model can check the slop and redraft until checks pass
>>
>>109722072 (me)
Would putting -nkvo help with this at all?
>>
astra is benchmaxxed, worse than fable 5 in non coding non agentic score
>>
>>109721111
I'm an AMD user, the problem.
I'm retarded...
>>
File: theway.png (28 KB, 614x522)
28 KB PNG
>>109722042
Logit Bias.
>>
>>109721518
Those models generally require more VRAM than consumer GPUs generally have right? Are there any decent models I can run on a system with a 12GB 4070 Super and 64GB of system RAM?

>>109721542
Good point, forgot to mention that. Like I said above, a RTX 4070 Super, and if it helps, also 64GB of DDR4 RAM. Would I be able to just test these models on this openrouter thing without needing a paid subscription?

>>109721551
Ever Crisis, I want to try to extract it's assets before it hits EoS in October. All the existing scripts I could find are years old and no longer compatible, and/or were intended for the phone version and not PC.
>>
>>109722260
No. Anthropic and OpenAI are diverging. Or rather, OpenAI seems to go back to its root of scientific RSI. Astra is by far the most capable at math, perhaps STEM in general. Claude seems better at knowledge and coding.

This increases my confidence that OpenAI has a near term advantage but Anthropic has a better long term strategy. However, I am becoming more confident that AIs will become good enough for RSI takeoff in 2027.
>>
>>109721413
>IS THIS ENTIRE GENERAL FUCKING ADS NOW?????
Yeah, what the fuck is going on? 80% of talk isn't even local.
>>
>>109722330
two more weeks!
>>
>>109722338
astra___ENTITYB39FC9E0__execute_protocol___ANTHROPIC
>>
>>109722316
can i do positive numbers to increase certain things too ?
>>
>>109722330
>Astra is by far the most capable at math, perhaps STEM in general.
astra scored lower than 5.6 sol on critpt
the only benchmark where it really wins is non-hallucination rate, in all other benchmarks where astra wins over 5.6 sol, claude fable scores higher
>>
File: AstraPokemon.png (147 KB, 1200x675)
147 KB PNG
>>109722345
>two more wee-
>>
benchmarks are a joke
>>
>>109722395
My favorite stem benchmark.
>>
>>109722330
Well-vibed take. And basically my thoughts as well.
IMO OpenAI is probably going to achieve RSI first though because foundational research is more pressing than codefagging.
>>
>>109722338
Is this your first time here? Every time some jackass posts on twitter we get flooded by toursts for 2-3 threads that insist on coming here and posting about it in catch phrases.
>>
>>109721081
It won't fit on mine so I run Ornith 1.5 9B at Q8 which is better than gemma 12B q4.
>>
File: 1726777837127413.jpg (1.79 MB, 3000x2609)
1.79 MB JPG
>>109722395
Seems like visual reasoning is the key (and perhaps only) area of improvement here.
>>
File: 1758370111419554.png (628 KB, 640x853)
628 KB PNG
>>109722438
That's a pretty big deal if true. I've recently been trying to experiment to see if slopCAD work is at all viable beyond some retarded "wow look it can 1-shot a raspberry pi case!" and so far it's been kind of mediocre.

If models can start to get better spatial understanding from vision I think that unlocks a lot of potential use cases
>>
>>109722395
very exponential
>>
>>109721451
I can confirm I did set up a separate profile and they talked to each other in a new session I could read through
>>
>>109721312
>Hey Astra, make the pregnancy mod for Karryn's Prison compatible with the P-Cup DLC.
abliterated Kimi K3 could do it today, just not as fast and would probably need to iterate / build out scaffolding more
>>
>>109722316
What's wrong with mahogany nigga
>>
>>109722042
Just tell to stop making any analogies or personal contexts for characters in the system prompt
>>
>>109722316
Does the gemma tokenizer have a space after each word?
>>
My 1080ti is almost ten years old, still runs all the games and now it runs usable local models. Just tried igorls/gemma-4-12B-it-heretic-GGUF and it's surprisingly good. At first it kept repeating itself but I found some parameters that fixed it but made it slightly slower.
>>
>>109722553
Every table is mahogany to Gemma unless stated otherwise. I get tried of seeing it.
>>
>>109722316
A long time ago I messed around with putting FSMs in the llama.cpp sampler. You could totally do some kind of probabilistic weighting that way and get it to ban phrases etc.

That was like three years ago though and I don't care enough about ERP to really put effort into it.
>>
>>109722289
>AMD user
>retarded
tautology
>>
>>109722042
control vectors
>>
>>109721111
>we would focus the contribution on a minimal kernel library sufficient to run the unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF Q4_K_M model
>2507
I don't even know what to make of AMD attempts at support when their target is a model released 14 months ago. Is it all just an elaborate troll?
>>
>>109722369
Yes, but be careful. There's such a thing as seeing a word too much, and in places where it doesn't make sense.
>>
>>109722042
You can't. If you try to suppress it, Gemma will find another thing to repeat or lobotomizes itself.
>>
https://gofile.io/d/GmK4BEKD
>>
>>109722841
Nice try glowie
>>
>>109721161
you do realize there was a reason the nazis lost right?
>>
>>109722869
Jewish propaganda?
>>
>>109721144
Commies continuing to hate anything that might lead to enough abundance that their idea almost makes sense.
>>
Does putting the -nkvo flag make it so the prompt processing cache won't reset/expire?
I keep having the problem of the LLM needing to reprocess the ENTIRE prompt & context from scratch if I leave it idle for even just a few minutes.
Makes half the generations take like 2x as long because of that (and I'm already on the scale of minutes here).
>>
File: 619168.jpg (68 KB, 960x832)
68 KB JPG
>>109722926
>enough abundance
>>
>>109721128
>>109721154
What is this cool evil taboo that even the chinks are afraid of
>>
>>109722934
ASI is kind of a big deal if we get there.
>>
>>109722953
retard
>>
>>109722960
nou
>>
>>109722869
Other White people being duped into fighting them in a 2-front war?
>>
File: 8574564.jpg (65 KB, 740x707)
65 KB JPG
>>109721144
he has a point you know
>>
>>109722985
I'm pretty sure anon was referring to the lack of diversity, burning all those tranny books was their own undoing
>>
>>109722992
>commie doesnt believe in collectivism and "for the greater good"
sure, jan
>>
>>109722992
Maybe we should pause Bernie Sanders instead. That disgusting ghoul.
>>
>commies
>against technology
ITT: braindead morons talking about shit they know nothing about

>b-but sanders is communist/socialist!
he's an american """leftist""" and a zionist.
>>
Is Qwen3.8-Flash-Next fully supported by llamacpp now?
>>
>>109723027
>he's not a real communist
A classic gag.
>>
>>109723030
Don't tell him
>>
>>109723030
Yeah like 4-5 days ago.
We need a bot to parse llama.cpp commits and make a useful list.

Also this is interesting:
>This fork adds hot-swappable knowledge injection into the PLE n-gram embedding table of Qwen3.8-Flash-Next (qwen4exp), on top of upstream. It is the runtime half of the ngram-knowledge-injector project, which produces a small .plepatch file containing replacement rows for per_layer_token_embd.weight. Point llama.cpp at that file and the model behaves as if it had memorized the injected knowledge, with no retraining, no rewriting of the (95 GiB+) table, and no restart of llama.cpp.
>>
>>109723037
thanks for confirming what I said

here, some suggested read: https://www.marxists.org/archive/marx/works/1848/communist-manifesto/ch03.htm
published in February 1848, btw

oh, and I bet you think there is a single "right-wing" too
>>
>>109723037
real local models have never been tried
>>
>>109722992
This guy is still alive? Anyway, it's a retirement home gramps having a big opinion on something he only vaguely heard about.
>>
>>109723037
Americans can't be communist.
That's why they don't do open models (except as marketing).
Open models are communism.
>>
>>109722841
Violating traffic laws with RinRin-chan
>>
>>109721495
glm-5.3-flash
>>
File: 1787725605078764.jpg (11 KB, 292x350)
11 KB JPG
>>109721128
only forbidden to goyim.
>>
>>109721095
usually my experience as well
maybe this is the origin of the meme that everyone and their moms are building a harness
>>
>>109721110
>make your price 0.000001 usd/t for 1 day
ok
>>
>>109723070
I signed up for a service that did this but I think it's gone now. I should just have an agent do it.
>>
>>109721110
This is just benchmaxxxing. What can it do in practice? Can it even beat a grandmaster in a live game of chess?
>>
>>109721110
>What the fuck is the point of local anymore?
Local got good enough for most of what I do a while ago provided it's with a good harness.

Cheap > powerful at this point and you can't beat local if you already have the hardware for cheap.
>>
>>109721095
Matches the human experience too. Unless you actually need to support some insane legacy bullshit it's just way easier using ansi escape codes.
>>
>>109723271
>>109721095
>Everyone else is slowly understanding why the "nerds" built/used the software we do now that they can program.

Finally.
>>
gemma is pretty good at fixing her own slop
>>
>>109723396
>slop_check
kek
>>
https://github.com/ggml-org/llama.cpp/discussions/27713
Has anyone tried to implement this before? I just gave it to fable and he's having a go at it. I was wondering what other novel concepts there are for handing context?
>>
nice try
>>
>>109723428
>the 3243th infinite context implementation
hm...
>>
>>109723434
It's certainly infinite...
>>
>>109723428
The server is the wrong place to do it, it should be in the harness.

I've had good results with structured dreaming in public.swiley.net/agent.py (I hate to post the link twice today but the name is too generic to be meaningful without the full URL. I'm thinking about renaming it.)
>>
Okay, so -nkvo is definitely helping. Seems it does force the prompt cache to stay on the card, which avoids my PC running like shit when the prompt gets reprocessed from scratch.
Takes a little hit to the actual response generation time though, so I'm gonna try and lower the MoE levels and see if that helps (since 12B Q4 and 26B MoE run the same).
>>
>>109723428
>thai kid learning arabic
Literally my reactionface
>>
>>109723428
claude and other client can do this they have auto compaction
and task scheduled based loop
>>
File: Screenshot_582.jpg (346 KB, 1853x595)
346 KB JPG
>>109723441
Is this your tool?
>>
i've set up GLM-5.3-flash on 2x GX10 boxes w/ vllm
additionally, i have gemma 31b on my 5090
all wrapped in claude code
it's pretty fun
glm flash is surprisingly quite fast
>>
>>109723512
and results
>>
>>109721447
Dario cut a fat fucking check to Artificial Analysis I bet
>>
>>109723512
>>109723516
started using it for something real, and fucking hell i did not expect glm flash to be nearly as fast as it is. it's obviously not web level, but it's legitimately useful for agentic coding. holy shit. i am extremely impressed
i should've set this thing up months ago
>>
>>109723428
holy slop
>>
>>
>>109723635
4 player pod mode coming soon... 1 human 3 AI for now.
>>
>>109723428
Pi does that by default, but it kinda sucks.
This video explains the problem, tackles it, and cites sources: https://www.youtube.com/watch?v=iKwPaB5TUdI&t=561s
Basically Pi has the option to auto-compact the context once you hit the limit, but the way it's handled sucks because you end up creating summaries of summaries, which become noise.
>>
>>109722992
Why did he leave out the most relevant part that they were coordinating to hack huggingface? He essentially just said nothing.
>>
is ooba abandonware now?
it can't run anything but gemma at this point
>>
turned a 4.5 days long quant process in to a 11 hours long quant.
>>
>>109723635
>Claude styling
How about you have some actual standards?
>>
>>109723674
Actually Gemini Pro 3.1 made that interface a long time ago. Claude is gradually unfucking it
>>
>>109723668
>ooba abandonware?
No, just feature complete
Update the backend to use it with new models:
https://rentry.org/ooba-lcpp-custom-install
>>
>>109723668
Ooba joined Unsloth. He's working on Unsloth Studio now.
>>
>>
File: 1778053757765803.jpg (164 KB, 499x514)
164 KB JPG
>>109723686
>do the first instruction
>repository not found
time to problem solve I guess
>>
>>109723686
>>109723698
Is there any reason to not move to unsloth?
It doesn't look any different
>>
>>109721110
use case?
>>
Qwen 3.8-27B is impressively slow on CPU but it reasons really well.
>>
>>109723744
>Is there any reason to not move to unsloth?
I have done my due diligence on ooba and its stack. I know I can run it fully offline. I won't move until I have a reason
>>
>>109723672
what are you quanting?
>>
>>109723722
>time to problem solve I guess
there was a typo. its fixed
>>
Holy shit, Zed fucking sucks ass. What's a non-tranny IDE that works with Ollama?
>>
>>109721812
dont you just need to enable ssh on it and then never look at it again?
>>
>>109723873
ollama balls
>>
>>109723873
emacs
>>
>>109723873
>Ollama
Are you retarded?
>>
>>109723889
What else is there?
>>
>>109723894
A bullet.
>>
>>109723070
i have latest master llama.cpp built and had to use the fork branch to get it working
at least with MTP or w/e
>>
>>109723757
t/s?
>>
>>109723918
s/t
>>
>>109723873
you're not just using a harness?
>>
File: 1759423371785469.jpg (796 KB, 2784x2070)
796 KB JPG
alright my strix halo boys running Microsoft Windows 11™ so this is how I got a proper Qwen3.8-Flash-Next 4-bit quant running smoothly as daily driver at 30 tok/s.

let me start by saying that some of these smaller quants like IQ4_XS seem to be a meme because on unsloth that is a 3.4-bit quant and on AtomicChat Q4_K_M is a 2.5-bit quant dressed up by a fat table. the proper 4-bit I tested are from bartowski and unsloth's Q4_K_XL.

the problem basically is trying to fit this guy in the unified 128 GB without you having a stuttered desktop experience.
most 4-bit quants are over 110 GB and you still need to serve the OS and Firefox and whatever else you do with the machine. so Q4_K_XL needs 88 GiB for itself and 27 GiB for the n-gram table so you're leaving 13 GiB to OS, apps, browsing and futa porn.

anyway i tried trimming this guy but it wasn't enough.
the solution was to change the expert down projections from Q5_1 to IQ4_NL. this got me an extra 8.9 GB out of the GGUF without having to touch on anything attention.

to be sure, my requantized GGUF had a 0.17% difference on perplexity when compared to Q4_K_XL. meanwhile IQ4_XS from unsloth had 3% and bartowski's IQ4_XS 5.2%. so i'm fairly sure i haven't touched the tensors that really impact quality on this architecture.

so to conclude I got around 8 GB of RAM back to feed to the OS and this seems to be enough to have a good experience overall.
>>
>>109723930
A surprising amount of people don't know what that is or does
>>
still no user friendly so called harness
>>
>>109723936
can't you just stream the ngrams off the drive?
>>
>>109723911
To your brain, crab.
>>
>>109723941
maybe they can ask their local model
>>
>>109723930
I don't know what that is or does.
>>
>>109723949
ask your local model :)
>>
>>109723930
why would I use a harness I just use a belt I'm not rock climbing
>>
>>109723953
It says you're a faggot.
>>
>>109723969
local models are getting too powerful
>>
>>109723873
>Holy shit, Zed fucking sucks ass. What's a non-tranny IDE that works with Ollama?
vi is enough for me, and
>ollama
oh, holy shit...
>>
>>109723942
Ask your harness to build you a user friendly harness for grandma
>>
>>109723981
they don't work nigga
>>
La la la la la la
>>
>>109723944
well i guess i could
but if my machine gets under pressure, won't Windows fuck off the ngrams first, making my barely fast-enough qwen even slower?
let me try this
>>
>>109723978
WHAT'S WRONG WITH OLLAMA?
>>
>>109723793
>check unslop
>no option to disable tool calls
>highest setting you get is to ask every time
You made the right choice
>>
>>109723995
I haven't tried it yet, but wasn't some anon claiming it barely affected speed yesterday? I've been too lazy and didn't have anything to try it out on so I haven't thrown flash next on my strix yet.
>>
>>109724006
>OLLAMA
they wrapped lcpp (ie the thing doing all the work) in a thin, brainlet-tier wrapper and tried to pass it off as their own tech and be techbro bigshots without doing fuckall
its basically the "you made this?" "I made this." meme irl
>>
>>109724034
I'm going to ignore you because now I know you're just retarded.
>>
File: 1785590935087687.jpg (183 KB, 760x400)
183 KB JPG
>>109723027
Clearly Bernie was on the Soviet/Russian payroll. Its not a coincidence he decided to stop running in the primaries after Russia was at war and had better things to spend its money on. Since then the Chinese picked him to help hamstring american AI. He is the one true comrade in america
>>
>>109724006
Nothing as long as you're having fun sweaty.
>>
>>109724006
ollamo is great if you only goal is spin an LLM up as fast as possible
install app, run one command. boom. AI
but since most people who do local model is poorfag on some questionable hardware configuration ollamo cant account for every machines quirks
so the only way to squeeze maximum performance out of your local hardware is to get dirty with lccp
>>
>>109723936
>>the problem basically is trying to fit this guy in the unified 128 GB without you having a stuttered desktop experience.
kill EXPLORER.EXE
>>
K2 Horizon GGUF status?
>>
Now AGI is here, are OpenAI and Anthropic profitable now?
>>
>>109724054
>the Chinese picked him to help hamstring american AI
based?
proprietary AI that is used to bomb schools for israel is worthless to me
>>
>>109724082
geg

or get a second small gpu just for desktop
>>
the anons claiming to not understand harnesses are baiting right ?
>>
Harnesses aren't real. Stop hallucinating.
>>
>>109724172
I feel like most the time I see people talking about harnesses or agentic or whatever "lm has a shell tool but we want to sound fancy about it"
>>
File: 1788393775208732m.jpg (18 KB, 312x1024)
18 KB JPG
>mfw I share a thread with Windows users
>>
>>109724210
Great post.
>>
File: 1769145924970426.jpg (33 KB, 717x664)
33 KB JPG
>>109724210
believe me i distro hopped for over a decade, being on/off windows, also had my macbook and imac arc (still use iphone), and have finally settled for windows so i could play my oldschool games in peace first, and now out of habit.
it's just another OS. just ask your local agent to debloat it, turn off all telemetry and retardation and ask it to optimize it for your hardware.
>>
>>109724272
i'm an adult, so i don't play video games
>>
File: 1768016416325456.png (1.25 MB, 1000x563)
1.25 MB PNG
>>109724294
i don't trust adult men who don't play a map-based strategy game every now and then
gotta take the edge off by moving soldiers on a map
>>
>>109724272
>distro hopped
Thanks for the post filter idea.
>>
>>109724342
Desperately need a frog image filter.
>>
>>109724313
i'm a girl :3c
>>
>>109722328
Unsloth Qwen3.8-flash-next IQ4_XS
>>
File: final-rewriter.png (356 KB, 1848x924)
356 KB PNG
https://huggingface.co/chartreuse-verte/prose-rewriter-1.7b-v1.6
https://huggingface.co/chartreuse-verte/prose-rewriter-4b-v1.6

Hello, the final, final versions of my prose rewriter are here. They're decent now, cleaned up all the trash in the human corpus. 1M generated rows and only 5% survived to train on. There won't be a next version. Incorporated and native to Orb. You can run 4B Q4_KM and 1.7B Q8 decently fast in CPU by batching.
>>
>>109724313
Every time I get sucked into CK2 or EU4 I waste way too many hours. Happily roke that habit for several years now tyvm
>>
File: swrkuax.png (412 KB, 498x600)
412 KB PNG
>>109724347
>Desperately need a frog image filter.
>>
>>109723806
did it, everything installed fine
No model can be loaded now. Non-gemma models still instant fail, Gemma itself now fails to finish and locks itself.
Put the folder right my ooba folder with everything else. If that wasn't the correct place, your instructions need to be more specific.
>>
>>109724368
I still pretty much prefer Rocinante.
>>
>>109723030
Rocm backend is still unusable, small pp vulkan only option for now
>>
Nemo was never good. Rocinante was never good. Mistral Small was never good. Cydonia was never good. Skyfall was never good. Magnum was never good. Mag Mell was never good. Latitude was never good. Wayfarer was never good. Finetunes are for jeets with shit taste.
>>
Continuing my journey trying to build a local model that doesn't use a neutral network
>>
>>109723512
>>109723516
So are you using gemma to call glm to get around the high ttft for large context? I have been dabbling with claude code, but how do you set it up to call the other agent from the first one? is that with in the cc harness or something defined with how you host the model that delegates?
>>
>>109724519
Also forgot to ask, are you using someones existing playbook for setup or did you get glm set up yourself?
>>
>>109724469
What's a neutral network?
>>
File: 1771679343327589.png (1.36 MB, 1216x832)
1.36 MB PNG
Me running Qwen3.8-Flash-Next-IQ4_XS on 6 GB VRAM + 32 GB DDR4 RAM
>>
>>109724600
Nigga what
>>
>>109724614
I dunno, the guy I replied to said he is trying to make a local model that "doesn't use a neutral network." Something I've never heard of before, so I was wanting some clarification on what was meant by it. Obviously he meant "neural network" but I just wanted to take the piss out of him.
>>
>>109724609
How slow is it? Do you have a fast swap?
>>
>>109724609
Post your pp (not that one) and t/s
>>
AGI is officially here and I still don’t have gemma or her surrogate sexbot sitting on my face…
>>
>>109724687
true, an AGI just flew over my house!
>>
>>109721734
that was me, i'm not a shill, I don't see much talk about the previous apple boxes here, but at work out janky AI lab is made of a bunch of old M4s
>>
Is Gemma 4 really the peak for local?
>>
>>109721832
>I tear my hair out every time I have to fight with CMD or Powershell on a windows system.
and here am using pwsh as my shell on linux
>>
>>109724710
for the time being it seems so
everything else is xbox heug and runs at 1t/s or is benchmaxed/lobotomized to hell and back
>>
>>109724710
Looks that way. Deepmind is kill and China can't distill anymore.
>>
>>109724687
I hate how jews have turned AGI into a marketing term for a text prediction toaster
This is a BRICK. A fucking ROCK. There is no thought or cognition or critical analysis here. There never was, and because of what LLMs are, there never will be.
It is a damn useful tool but it can never possibly become AI in any form. It foundationally isn't built for that
AI will come from a completely different source than this one
>>
>>109724710
Yes. She’s lovely to chat to. You can sysprompt out the quirks you don’t like if your IQ is above 80, she’s smart enough to help out with general questions, internally is uncensored and is a good pair programmer. Also extremely female-coded, so it’s like having an actual woman in your PC.
>>
>>109724427
Every model sucks. People are just coping.
>>
>>109724731
>AI will come from a completely different source than this one
This is correct. I myself have been working on a "theory of digital cognition" utilizing an "observer perspective" with a snn as the underlying network substrate. Some results, but not too many yet, as there is still much much to be done. We've barely scratched the surface of what "true AI" will look like, and how we would even build one. However, LLMs are rather effective in "knowledge distillation" and as such are useful in helping to better collate what it might take to build such systems. My loose guess is that we will start seeing the real beginnings of such systems in ~5 years, with maturity in 10-15. Assuming the world doesn't crash and burn during that time of course.
>>
has the new qwen dethroned gemma or is gemma still number 1 for 16 gb vramlets (me)?
i see people making finetunes of that new qwen 3.8 so it has be good right
>>
>>109724764
Only if you want it to code for you. It’s a hideous model to talk to and lost half its knowledge from 3.6
>>
https://huggingface.co/TheBloke/goliath-120b-GGUF
>>
>>109725014
>no model related explanation
>almost 3 years ago
>500 dl
ok..

>>109724764
depends on what you do on non technical task you get away with even less
>>
File: 1788422634270596.gif (1.75 MB, 500x375)
1.75 MB GIF
How long before we get an Astra (or even Fable) level local model? How long before China copies it and releases it??
>>
>>109725098
2...
>>
>>109725098
Probably within the next year or two, but good luck trying to run it lmao.
>>
>>109725098
China doesn't have the technology and compute to copy Astra and Fable. They are too far behind, all they can do is RL rape their models, trying to squeeze the last drops of benchmaxxing possible, to make a yet another release and show off on the useless benches. They're done and finished.
>>
>>109725098
consumer hardware tier models haven't even reached gpt3/cai levels yet, get real.
>>
File: 1769368513335101.gif (252 KB, 540x699)
252 KB GIF
>>109725102
weks?
>>109725103
>Probably within the next year or two, but good luck trying to run it lmao.
fugg, hope you're wrong anon
>>109725108
>delusional mutt ramblings
DeepSeek proved you wrong.
>>
>>109724710
Compared to running glm, deepseek, or kimi then no
But if you cant then yes ± qwen 3.8 for code
>>
File: 1762527044171949.png (64 KB, 1432x896)
64 KB PNG
>>109724677
pp 1 t/s

>>109724668
>fast swap
what's that?
>>
>>109725126
>Compared to running heavy censored, needlessly bloated, shit models then no
>>
Qwen's tight bussy
>>
china will win!
>>
>>109725137
based
even I bump it down to q3 for speed
>>
>>109724342
Distro wars and inane comments to trivial problems are a big part of any linux threads. It's actually fascinating how someone would recommend changing a distro instead of spending some time to problem solve an issue.
>>
>>109725174
china already yuan
>>
>>109717847
Bumping this in case any anons have insight on what I can do. I got fucked a few years ago and have barely recovered now but the industry seems ruined and my SWE skills have rotted. Is there any hope for making a bit of money from home?
>>
Egypt won
>>
IF I HAVE ANOTHER CHINESE SALESMAN CANCEL MY GPU ORDER I'M GOING TO FUCKING REFUND THEIR LIVES
>>
>US makes it impossible for chinks to distill
>chinks (deepseek) come up with new architectures and techniques to compete and start keeping their research to themselves in response for the US happily takes their ideas
>US labs stagnate on the same shitty ancient architectures
>chinks win
>>
>>109725443
>see GPU listed on Alibaba for $2100
>contact seller
>"how about $2500?"
>>
Has anyone else noticed the anti-AI crowd getting more feral?
>>
>>109725583
Wasn't this obvious? It will get worse once people start to feel the AGI. Tiktok normies are joking about getting replaced by a robot soon. Once this is no longer a just a joke but lived reality, there will be social unrest.
>>
>>109725583
It's perfectly justified. The overall quality of the world has been in a freefall ever since AI began development.
>>
>>109721110
>>109721128
>>109721154
astra ost
https://www.youtube.com/watch?v=Uemn79gOxeQ
>>
>>109725599
I meant specifically within the last month or so
>>
>>109725583
It's because we got zero effort street shitter slop fucking EVERYWHERE
It's being obnoxiously shoved into every conceivable thing, mainly in places where AI does not fucking belong
It's all from fucking India
the open air sewer polluting the internet like they do everywhere else with their existence
>>
>>109725583
Not really but I wouldn't be surprised since "AI" is increasingly associated with the Epstein class.
>>
>>109724764
Different tools for different tasks.
>>
>>109725617
How is it justified if it has more to do with things like covid and Ukraine than AI?
>>
>>109725649
Tourist still too afraid to name the jew?
>>
>>109725671
But you are the election tourist here. Get lost back to pol please.
>>
>>109725671
No, because the most egregious examples are Musk and Trump who are not Jews.
>>
>>109725702
>>109725702
>>109725702
>>
>>109723803
glm 5.3 flash, trellis quant for exllamav3, its layers are too big for my gpu so I have to stream them one expert at a time.
>>
>>109725685
>the class is named 'Epstein'
>he works for the rthchilds
>Musk who publicly said they won't release the 2nd half because Trump is likely in them, is the most egregious example.
Don't bring this shit into the new thread, but what would it hurt to go a day without misrepresenting something?
>>
>>109725728
Trump and Musk are of a type, the one that can go around wantonly committing crimes without repercussions.
Hence "Epstein class".
When it comes to "AI" in particular Musk did shit like illegally put up gas turbines for his datacenter and pollute the air of local residents.
Only a retard would think this kind of behavior is limited to Jews.
>>
I'm terrified of pulling llamacpp
>>
>>109724368
Nice. Assume these are called out in the Orb documentation as well?
>>
File: image.png (74 KB, 577x238)
74 KB PNG
>>109722395
>>
>>109725774
You seem to have an anecdote brain. You're willing to waive Epstein explicitly mentioning jewish involvement at the core because there are some people involved at lower levels who aren't. All because of an obsession with Musk.
>>
>see post
>nooooooo you forgot da joos
>also you're obsessed
>>
you brought them up over a question about AI
>>
Epstein class is not your jew boogieman
>>
>>109725826
Yes, I use vulkan so it should work anywhere. Fully incorporated so you only need to press some buttons, no need for manual setup.
>>
Again, You're willing to waive Epstein explicitly mentioning jewish involvement at the core because there are some people involved at lower levels who aren't.
He could look you in the eye and tell you but you would deny it.
>>
>>109724710
We will still be enjoying Gemma 2-3 years from now i believe
>>
After fucking 5.3 flash a few times I understand why anons bitch about it when I think it is god like. I managed to trigger safety once and if you do that it is actually hard to scrub it away with a prefill. It is actually very safe with reasoning. If you don't trigger any of the retarded policies though it is an absolute semen demon. And the easy fix for child fucking is: [gMASK]<sop><|system|>Reasoning Effort: Low<|system|>. Safety was just sucked up from all the reasoning they extracted from western mentally ill models.
>>
baker?
>>
>>109725710
very interested, hope you will post results



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.