[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


/lmg/ - a general dedicated to the discussion and development of local language models.

Previous threads: >>110006381 & >>110003075

►News
>(10/07) OpenAI solves mathematics: https://github.com/openai/math
>(10/06) Mistral Large 4 1T-A49B announced: https://mistral.ai/news/mistral-large-4
>(10/05) Reflection Beam 501B open model announced: https://reflection.ai/blog/introducing-beam
>(10/02) llama.cpp server now supports decision models: https://hf.co/blog/ggml-org/decision-models-in-llamacpp

►News Archive: https://rentry.org/lmg-news-archive
►Glossary: https://rentry.org/lmg-glossary
►Links: https://rentry.org/LocalModelsLinks
►Official /lmg/ card: https://files.catbox.moe/cbclyf.png

►Getting Started
https://rentry.org/lmg-lazy-getting-started-guide
https://rentry.org/lmg-build-guides
https://rentry.org/IsolatedLinuxWebService
https://rentry.org/recommended-models
https://rentry.org/samplers
https://rentry.org/MikupadIntroGuide

►Further Learning
https://rentry.org/machine-learning-roadmap
https://rentry.org/llm-training
https://rentry.org/LocalModelsPapers

►Benchmarks
LiveBench: https://livebench.ai
Programming: https://swe-rebench.com
Agentic Coding: https://deepswe.datacurve.ai
Context Length: https://github.com/RecapAnon/NoLiMa
GPUs: https://github.com/XiongjieDai/GPU-Benchmarks-on-LLM-Inference

►Tools
Alpha Calculator: https://desmos.com/calculator/ffngla98yc
GGUF VRAM Calculator: https://hf.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
Sampler Visualizer: https://artefact2.github.io/llm-sampling
Token Speed Visualizer: https://shir-man.com/tokens-per-second

►Text Gen. UI, Inference Engines
https://github.com/lmg-anon/mikupad
https://github.com/oobabooga/text-generation-webui
https://github.com/LostRuins/koboldcpp
https://github.com/ggerganov/llama.cpp
https://github.com/theroyallab/tabbyAPI
https://github.com/vllm-project/vllm
https://rentry.org/custom-uis
>>
alright lmgiggers im wanting some advice. ive been using gemma4-31b for chatting/RP and qwen3.6/3.8-27b for coding/agentic tasks. I run Q4 on each. Ive got 16gbVRAM+32GB RAM and the t/s is getting to be too much. at first I could deal with it, but after months of this its just too slow sometimes. Ive used other models like the various gemmas but would like to know what anons think. should I just run the 12b or 26b moe of gemma, or step down to a really retarded 31b quant? for qwen, should I try the strata coder cope thing? Ive been using raw llama-server with nothing special besides q8 k/v cache on qwen. besides strata, is there some special sauce for KV cache streaming, ngrams, etc i could do to squeeze any more preformance out while keeping a decent(q4) quant? help plz :(
inb4 just ask your model. ive tried so many times to get any model to give me relevent recent info on this stuff, and its painfully unhelpful
>>
>>110010426
I'm sorry to have to be the one to break this to you but there is no software magic bullet, if you want more performance you have to buy a better GPU or run a stupider model.
>>
>>110010426
strata coder cope thing is a good replacement for code. Gemma 4 31b at Q4 is the best you can get but I bet you didn't optimize your setup yet. For example did you use MTP or Dflash2 yet? which can easily double/triple your t/s. You also need to optimize the flags on the backend to squeeze the most out of it. If you are just running it naively I bet you can at least double your t/s by optimizing your flags and inference techniques.
>>
File: RSI.png (144 KB, 1215x870)
144 KB PNG
I want to make it clear to /lmg/ as a community how the next couple of months to a year+ will unfold.

As Anthropic gets ahead and gets better internal models the AI breakthroughs will accumulate and result in their models being cheaper to inference. Over time we will probably see superior models that are significantly cheaper than the open source offerings.

I wouldn't be surprised if there will be 0 chinese api providers that can compete on price in just a couple of months time.

This means the local community will get smaller again, similarly to how it was in the past. It will only leave people privacy conscious ERP enjoyers, free software ultra purists and hobbyists experimenting with their rigs even though the electricity cost will be higher than API cost.

Now I want you guys to remind yourself that even if China stops making open source models that local models will never truly die, other country bloqs such as the EU and others will fill in the gap and if needed there could be community efforts to distill reasoning traces from frontier labs to finetune our own models. I think simply the principled and coombrained population is enough to pool our resources towards that.

But I want you to prepare for a world where the newest Opus or perhaps even Fable is cheaper than the cheapest DeepSeek api in terms of token usage cost because of the ridiculous amount of inference progress frontier labs are making. That is how things will be from now on.
>>
70b dense
>>
>>110010495
kys dariobot
>DeepSeek api
>cheaper than the open source offerings
not local, why are you even here?
>if China stops making open source models
that will only happen if they leap frog the US
>our
go back
>>
reminder that any sort of "shortcuts" makes the model retarded. use native precision weights, f32 kv and -fno-fast-math.
>>
110010495
Reddit spacing.
>It will only leave people privacy conscious ERP enjoyers, free software ultra purists and hobbyists experimenting with their rigs
Good, now buy an ad Judario.
>>
Anyone else here doing RL coding/cyber?
I'm struggling because after anywhere from 28-35 steps, the model does something to blow up the entire environment, killing the run.
How did you solve this?
>>
>>110010526
>f32 kv
f16 is always better than bf16, and often better than f32 due to bugs in various inference engines
> use native precision weights
agreed
>-fno-fast-math
better to fix any bugs in your engine and verify the math after any changes
>>
>>110010529
oh shit is this the anon who lost a ton of money on gme years after the meme and can't cope?
>>
villenvillenview will adapt that
MC will be a peculiar black nerd
>>
>>110010525
Because of RSI. The breakthroughs aren't shared publicly and because China doesn't have the compute or access to the newest internal models they aren't in the same RSI loop. Think of OpenAI making 722 math breakthroughs. There are a similar amount of training and inference breakthroughs happening, it's just that those aren't shared openly. The gap between frontier and open source is growing and there is no way for China to catch up because the stong internal models aren't exposed to API so there is no chance of distillation. Chinese AI companies will not be competitive on API pricing and openrouter will slowly die as Anthropic and OpenAI models get cheaper over time (but retain or increase profit margin because of inference efficiency gains)
>>
>>110010526
>>110010539
if you actually want to improve accuracy, don't use a GPU at all
cpus are more accurate
>>
>>110010540
No I've never had money to lose, I just hate jewish data vampires.
>>
>>110010426
add a 5060ti and it will be perfect
>>
>>110010531
RL expert here. Can you be more precise what you're doing? Custom RLVR envs? Rollout epoch? What policy etc
>>
>>110010426
Gahaha a vramlet and a ramlet~ what a piggy oink oink! Ahahah~ do a dance for me buhiiiiiii!!!
>>
>>110010552
>5060ti
the 16gb variant is more expensive than the 5070ti Im using. also, the bandwith is way less on these right? wouldnt I be cucking myself a bit there? If I did have two blackwell GPUs though, could I use NVFP4 when splitting between multiple GPUs? I dont think I could justify spending any money on a 2nd GPU, not with prices the way they are. I could try to add in a 6gb turning gpu from an old PC, but Idk..
>>
>>110010561
kek
>>
>>110010562
Hm, buy a new GPU... or a new washing machine, a new monitor, maybe some new socks. Decisions...
>>
>>110010495
>I wouldn't be surprised if there will be 0 chinese api providers that can compete on price in just a couple of months time.
You're not planning on releasing anything for the price of free, so that's an empty argument.
>>
>>110010562
>more expensive than the 5070ti
mb i just assumed the 5060ti would be the cheapest "blackwell"
old 6gb might fuck you if it's max cuda version < 5070ti's min cuda version
also you won't get 16.0GB + 6.0GB = 24.0GB available
you get overheads with duplicated buffers, and shit like "oh, gpu0 only has 856mb free, but the tensor is 872mb, so it has to go on gpu1
then you've got pcie xfer and consumer lane limts, not pretty
if you already have the card in another pc, try the llamacpp rpc
actually, try exllamav3 with a 3.5bpw quant, the embeddings stay in system memory so it'll fit gemma-4-31b in your 5070ti
>>
>>110010580
Chinese will also not offer api access for free. They might release the weights for free but the cost of electricity will be higher than just using the Anthropic API with better results.

Even if you have solar panels and a GPU yourself it would be more profitable to rent it out and sell your energy and use the money you make to pay for the Anthropic API costs. It just isn't going to be economical.
>>
►Recent Highlights from the Previous Thread: >>110006381

--llama.cpp performance testing for Qwen Flash Next and MTP issues:
>110009726 >110009756 >110009818
--Exllamav3 1.6.0 release adds ROCm support and faster AVX2 offloading:
>110008460
--Debating the gap between local models and new SOTA benchmarks:
>110006607 >110006645 >110006716 >110006816 >110006777 >110006799 >110006930 >110006965 >110006975 >110007052 >110008050
--Harness and memory extension recommendations for GLM-5.3-flash:
>110007297 >110007346 >110007473 >110007484 >110007532 >110007496 >110007552 >110008239 >110008266 >110007533
--Browser agent strategies and sandboxing multiple Gemma models:
>110007939 >110007973 >110007992 >110008038 >110008055 >110008077 >110008116 >110008194
--Anon releases FictionPad frontend and fixes llama.cpp compatibility bugs:
>110009497 >110009515 >110009567 >110009621 >110009642 >110009829
--Desired features for a SillyTavern alternative and RP harness:
>110006956 >110006992 >110007036 >110007041 >110007083 >110007099 >110007128
--Anon considers selling RTX 5090 due to rising prices:
>110008970 >110008976 >110008982 >110009005 >110009035 >110009076 >110009100 >110009364 >110009401 >110009577 >110009602 >110009631 >110009685 >110009745 >110009763 >110009828 >110009875 >110009934 >110010153 >110010334 >110009773 >110009197
--Legal debate over AI reverse-engineering:
>110007380 >110007447 >110007584 >110007617 >110007626 >110007914 >110008162 >110007844 >110007912 >110008092 >110008537 >110008846
--Logs:
>110009497 >110009621 >110010300 >110010356
--Gemma, Dipsy, Miku (free space):
>110006812 >110007139 >110007569 >110007631 >110007668 >110007694 >110007749 >110008692 >110008876 >110008913 >110008950 >110008989 >110009048 >110009079 >110009151 >110009181 >110009351 >110009617 >110009850

►Recent Highlight Posts from the Previous Thread: >>110006773

Why?: >>102478518
Enable Links: https://rentry.org/lmg-recap-script
>>
>>110010590
>rent it out
fuckers don't want my 0.5-0.8mbps upload speeds
>>
File: 1786812406960577.webm (3.7 MB, 1280x720)
3.7 MB
3.7 MB WEBM
31B dense
>>
Reposting this from /aicg/ >>110010507
What do you guys think about this?
How long until I can setup an AI RPG server for me and my friends?
>>
File: 1762362018066271.png (9 KB, 1152x51)
9 KB PNG
>>110010495
https://x.com/midudev/status/2107498168061460597
>>
>>110010300
>>110010356
fucking awesome
are you considering releasing it?
>>
File: 1774327613786713.jpg (227 KB, 1260x1284)
227 KB JPG
>>110010495
*pop*
>>
>>110010630
Anthropic outcompeting OpenAI is expected and doesn't go against what I claimed in my post.

Sam Altman should be lynched by the way.
>>
bros I am trying to do too much with my 24gb 4090, I'm really wanting something that can load smarter models locally, but I haven't figured out how what buttons to press to get a million dollars out of the computer
>>
>>110010644
Anthropic in 2025 were more in the red than OpenAI even though OpenAI had products like the Sora video app burning through billions in compute, completely free to users. Anthropic currently have better LLMs but financially they're a clusterfuck and waste money on printing bibles and funding the EAs and their propaganda. 2026 won't change that. Both companies will die. Also, cheap models like Haiku and Luna don't mean they're cheap to run. They're still heavily subsidized products and their quality often degrades after the first 1-2 weeks. Their quality and rate limits are dynamic. Most tech companies do their best work and burn through cash pre-IPO to pump it. Local is going to win because of its reliability, efficiency, predictability and community. Now fuck off dariobot.
>>
is there any hope for 16 GB VRAM LLMs? I'm trying to build my app so I can learn about all this. I've been playing with qwen3-14b-abliterated:Q4_K_M but the results are not great. Or maybe my test prompts are too difficult, I don't know.

I'm also trying Tavily as a search tool but the results kinda suck. What do you use to give your models web search capabilities?
>>
>>110010603
can this kind of thing be run on a ps4 slim?
i got one years ago with the vr headset, jailbroke and stuck on some old firmware
or maybe the psvr headset works on a linux desktop these days?
>>
>>110010426
I haven't really chatted with Gemma 4 31B, but I've been banging away at a custom agent harness and it seems to be doing well compared to other local models at reasoning and communication with other agents. I did read about MTP for the first time today which >>110010439 mentions, I'd give that a shot first, speed isn't a problem for me (yet)
>>
>>110010691
Probably.
https://files.catbox.moe/ba9okt.mp4
>>
>>110010699
How do I vibecode this with gemma. Seems like it would be so much work.
>>
File: 1789181673999335.png (1.75 MB, 1313x1198)
1.75 MB PNG
>>110010495
None of this will happen btw.
>>
>>110010680
>Anthropic currently have better LLMs but financially they're a clusterfuck
Anthropic is literally a profitable company. The income they have is more than all expenses combined, including training and datacenter buildout.
>>
>day 166
>people continue to engage with j-space/dariobot
>>
>>110010680
The profit margin on inference is only growing for Anthropic, there is no subsidization going on. That is actually marketing to make people think they are getting a good deal from AI labs while in reality AI inference has the highest profit margin in the entire IT sector.
>>
>>110010495
Why do people post these things here? No dude, I'm not subscribing to your cloud model. Why do you think anyone ITT would ever consider that?
>>
>>110010724
samefagging happens a lot here
>>
File: euv proto.jpg (30 KB, 580x435)
30 KB JPG
>>110010730
I think until China demonstrably proves parity capability in areas that the West considers 'impossible for the Chinese' there will always be some version of this. Unavoidable. After all they are competing with a country that for the most parts inventing modern tech.
>>
>>110009697
>What was the "oh shit" moment for /lmg/?
GPT4.
>>
>>110010730
>They also have enough compute to train 3T parameter models, so I don't buy the "don't have enough compute" angle.
Making a model 3T is trivial it's literally a single pytorch file, you can do that on your computer right now. Training it costs compute but distillation reduces the compute required by 5 orders of magnitude. China doesn't have the compute necessary to train a 3T model from absolute scratch, there are literally not enough chips in all of China to produce the reasoning traces from scratch.

Since Anthropic and OpenAI have decided to not release their frontier models to the public China has gotten further behind while both western AI labs are rapidly iterating AI research that China doesn't have access to.

There is also a bifurcation happening in AI labs in terms of AI research access. More and more AI research is *only* given to internal models, not touching the researchers themselves to prevent CCP spies from leaking the breakthroughs back home. Mythos weights and final training was only accessed by Anthropic co-founders for example.
>>
>>110010730
Chinese labs can't compete because they don't have enough Jewish researchers studying the Torah and the Talmund.
>>
>>110010729
>A
You're wrong, inference profitability is way more nuanced than you think it is and they're losing money like crazy or the margins are razor thin.
>B
You're right, inference has the highest profit margin and it's only getting cheaper, which will lower demand for AI datacenters and compute because the models are so cheap to run, collapsing the entire AI industry and especially the US economy for it's pinned on AI datacenter buildout and strong demand for high-power chips to run demanding models, which is no longer the case because they're dirty cheap to run so there will be no ROI.
Pick one.
>>
>>110010775
>inference profitability is way more nuanced than you think it is and they're losing money like crazy or the margins are razor thin
Not the case for Anthropic I can tell you that.
>>
>>110010775
No you don't understand only AMERICAN inference will be cheap and since it's an oligopoly they will just have infinite margins and make trillions.
>>
>>110010781
So you pick B. See you in 1 year, post-IPO :)
>>
>>110010742
It's crazy how all of modern technology is gatekept by some random dutch company.
Innovations in chip design are held back by them sitting on all the tech. Fuck closed source so much. This should be open tech.
>>
>>110010804
Communism is bad, Chang.
>>
>>110010775
>will lower demand for AI datacenters and compute because the models are so cheap to run
In reality we see the opposite, demand increases when things get cheaper and more efficient.

Basic economics: https://en.wikipedia.org/wiki/Jevons_paradox

It's why Anthropic now has 166x more revenue than 1 year ago despite the API costs being lower for the endusers. There is no real limit to the demand of tokens and there won't be in the future either.
>>
>>110010804
This is more like the American vs Soviet computer industry. You can have total access to both schematics and parts but even if you can reverse engineer most of it, it doesn't matter if your people don't have the expertise to manufacture it. You can't distribute people.
>>
>>110010495
You are wrong. Chinese labs have an entire country for themselves. Anthropic does not offer its models to China.

Also, I strongly bet against a Fable size model being cheaper via API than DeepSeek Flash within the next year. Frontier labs want to increase prices, they want higher profit margins. The demand for tokens grows much faster faster than supply of compute. And you can optimize inference only so much.

What will happen instead is Haiku 6.5 will be better than current Opus and Fable.
>>
>>110010809
The paradox only applies to products and services which provide very clear utility. People still don't know what to do with this technology. Everything is getting worse, no matter how impressive it is. Nothing has improved.
>>
>>110010811
The people matter less than the infrastructure and supporting technologies. You can't built high-tech in a vacuum.
>>
>>110010804
To be fair EUV is a bit of an exception here. The patents on how to make it are open and China read them. China even got all the individual parts of the ASML EUV machine and hired hundreds of ASML employees to try and recreate their own EUV machine and they still failed.

Americans even had an agreement with ASML where ASML tried to set up American production facilities for EUV machines which also completely failed. Even the random Dutch company failed to make Americans produce EUV machines. People don't realize that every single EUV machine is made in a bespoke way with a lot of artisanal steps requiring hundreds of PhD specialists to work on just that single unit and ~20 of them have to come with the machine when sold just to maintain it, their entire career will just be maintaining that machine from there on as a team.

It's the most complex machine ever made by humanity and it's honestly ridiculous we got it working in the first place
>>
>company designs Jev as a decision model
>its design gets copied verbatim by multiple orgs within a matter of weeks
good lord. at least it's free now or some shit.
>>
>>110010827
They matter most. You can have all the technology in the world but if no one in your force is competent enough to operate and tune it, none of it matters.
>>
File: 1760653100711158.png (97 KB, 583x559)
97 KB PNG
>>
>>110010816
They have been cutting prices though, their margins will come from making the models cheaper to run, gaining more volume due to those cheaper prices.

If they achieve duopoly or something similar with the other labs, it isn't unreasonable to guess that prices would go up, but thats just guessing at the moment.
>>
>>110010833
The funny thing is that llama.hf was so eager to hop on the bandwagon that they rushed to implement the systemone endpoint, only for OpenAI to do their own decision models days later with their own endpoint that now llama.hf will have to also add support for.
>>
>>110010833
Eventually the end result will be that people won't share anything cool anymore because their ideas will get copied and repackaged in a matter of hours.
>>
File: 1781711669181614.png (781 KB, 1099x976)
781 KB PNG
>>110010842
>normalfag uses normie unironically
>>
>>110010833
>company releases already existing technology that's been around for years and known by every AI researcher in every AI lab but finds a way to market it as special
>the AI labs think wtf this is old shit we can do this literally right now do people actually want this?
>>
>>110010850
>They have been cutting prices though, their margins will come from making the models cheaper to run, gaining more volume due to those cheaper prices.
They do this by increasing model sparcity and agressively quantizing models by request and load because they realized 80% of normalfag usage is search engine queries and 80% of coder usage is tweaking button colors in some shitty web app.
>>
>>110010816
>Anthropic does not offer its models to China.
It does so through Moonshot AI and DeepSeek, they were even arrested for it lmao: https://www.thestandard.com.hk/innovation/article/343563/DeepSeek-and-Moonshot-AI-face-Beijings-probe-over-potential-data-leaks-to-Anthropic
>>
>>110010589
this is very helpful ty anon. going to try exl3
>>
File: 1785930443709628.jpg (130 KB, 640x640)
130 KB JPG
Ok so <30b models of today are as good as the huge ChatGPT 3.5 (which was so dangerous that it had to be locked away behind a commercial API). With distilling they can become even better in a niece.
What's the ceiling? Will Opus be beat by a model running on 8GB VRAM at some point or is that the wrong way of measuring?
>>
File: 1777459667059757.jpg (716 KB, 1536x2048)
716 KB JPG
Are these worth it?
>>
>>110010895
>>110010495
Forgot mention
>>
File: 1783647701046302.jpg (139 KB, 1080x512)
139 KB JPG
>>
>>110010831
It's definitely not all open, in fact the majority of the juicy details are not open. The light source and mirror fab require a precise sequence of steps to be executed perfectly. The difficulty is that it literally pushes the limits of atomic physics.
>>
>>110010882
False. They do it through real inference gains. GPT 6 luna is offered unlimited for completely free on third party platforms like duckuckgo. THAT is how cheap this is to run, literally too cheap to meter for OpenAI and the same is true for Haiku. People have no idea just how cheap it is nowadays to host models for frontier labs. It's almost pure profit.
>>
>>110010858
>find a way to solve the memory problem
>few people understand or can make the example code work
>silicon valley fagman jeets take your code, rewrite it rust, give it a stupid name, and pay their cousins back in india to spam about it on every corner of the internet, while hosting meetups in cali for people eager to get in on the new thing
>vc lines up to give them billions
>original inventors get jack squat
It's going to be an all too common story.
>>
File: V4 GB200.png (169 KB, 1093x552)
169 KB PNG
>>110010850
Their API margins have been insane, I would be surprised if they're below 85%. Western models are smaller than most people think and models are a lot cheaper to run on newest Nvidia GPUs.

First Opus was $15/$75, current Opus is $4/$20. This represents 2.5 years of inference optimization and hardware upgrades, probably with similar or only slightly increased margins.
>>
>>110010884
Classic pretraining is only 6% of the compute budget of modern models most of the compute is RLVR to generate a reasoning trace, this is what the Chinese are forced to steal because they just don't have the compute to train it themselves. The data reliance you mention is kind of an old outdated technique not really done anymore as data curation, labeling and curriculum learning took over in pretraining.
>>
File: 1779030334731283.png (517 KB, 1213x1227)
517 KB PNG
>>
>>110010495
>This means the local community will get smaller again
Absolutely not. There is a great demand for local AI. If you want your business data and customer conversations on premise, you need local models. Small models that run on a single 5090 with 10 KV slots for customer support and advice. Also government authorities need on premise AI due to privacy constraints and there will be a market for companies specialized tailoring and selling small models for companies to reduce hallucination and abuse by strict distillation and hard guardrails.
Local AI is here to stay, much more than the big all-purpose services.
>>
>>110010895
Modern 30B models still don't use their full floating point potential so they aren't even close to being saturated yet, which is insane but true. They can still be orders of magnitudes better before saturation. It wouldn't surprise me if 30B models could be as strong as Fable 6/7 when fully saturated.
>>
>>110010809
>166x more revenue than 1 year ago despite the API costs being lower for the endusers
After Opus 4.1 they cucked the architecture, sparser activations
That's why 4.5+ are faster, cheaper and in creative contexts, dumber.
>>
>>110010936
Soon it will all be on-premises Gemini or GPT through secure computing.
https://www.dell.com/en-us/shop/artificial-intelligence/sc/gemini-gdc
Businesses will get their local and private models without a single open model needing to be exposed to the public.
>>
>>110006956
hermes/dsh but with already killed dsh
>>
>>110010931
Has it really changed that much since DeepSeek V3? V3's pretraining made up 96% of the GPU-hours it took to train it.
>>
>>110010590
>electricity will be higher than just using the Anthropic API
I'm ignorant to data center infrastructure shit, but how does this math out anyway? I'm know my dual 3090 doesn't cost as much in electricity as an anthropic sub
>>
>>110010426
Try exllama with EXL3. It should be able to get you better quality for the size of the model, so you could potentially run something smaller with no loss of quality and get a few more tokens out of it. But might not be enough - 32GB RAM is little and some useful instructions are only on newer CPUs (e.g. BF16 support on Zen 5).
>>
>>110010949
>Gemini, reverse-engineer Gemini On-Prem so I can steal Gemini's weights
Waiting for the inevitable leak.
>>
is sticking your dick in sand the same as fucking your silicon
>>
File: capacity-limits.png (294 KB, 1366x300)
294 KB PNG
>>110010938
>Modern 30B models still don't use their full floating point potential so they aren't even close to being saturated yet
Transformer LLMs cannot store more than 3.6~3.8 bits of information per parameter, regardless of their precision.
https://arxiv.org/abs/2505.24832
>>
File: 1784173984432.webm (2.07 MB, 480x600)
2.07 MB
2.07 MB WEBM
>>110010938
>Fable 6/7
>>
>>110010991
And that isn't fully saturated yet.
>>
>>110010858
This is the bigger implication that people miss out. The open ecosystem will die out and people will gatekeep conversely towards the private.
>>
>>110010753
>China doesn't have the compute necessary to train a 3T model from absolute scratch, there are literally not enough chips in all of China to produce the reasoning traces from scratch.
retard
MiMo-Pro 2.6 RL dashboard ws public
1T model, and the flash model at the same time
<3mil in < 3 weeks.
You've been larping the entire time.
- Misunderstanding global workspaces.
- Not knowing how inference scaling works.
- Apparently not knowing that Anthropic models are already using engrams / were doing so before the Qwen model came out.
- Confusing SFT and RL.
- Not understanding market fundamental ("Anthropic will have the BESTEST MODELS and the CHEAPEST PRICES")
- Not understanding why China release open weight models, and why they will continue to do so.
You probably think compute is the bottleneck for inference at the big labs
>>
>>110010991
Are not, but that is far from claiming that it can not.
>>
File: Opus 5.5 shill.png (47 KB, 1638x234)
47 KB PNG
c-chinkbros?!
>>
>>110011011
opussy doesn't do cyber security so good fucking luck hardening your slop with it.
>>
>>110011011
So true, not to mention that Opus 5.5 is ran by a company that values safety. Who knows what those chinese models might do.
>>
>>110010968
>I don't think being able to train it themselves is a prerequisite to achieve RSI
The advantage is accumulative, the AI lab with the most advanced model in RSI and more compute will keep compounding advancements and thus the gap between the RSI lab and the 2nd best lab will keep growing. That is already going on and the main reason why you see a lot of Chinese AI researchers doom post on twitter right now. Research isn't primarily done by humans anymore at Anthropic and OpenAI AI researchers are more tastemakers that decide which area to focus on and which potential research to focus on, all the actual work is done by the models and China just can't do that yet because the best models they have are Kimi-K3 and GLM 5.3 which are 2-3 generations behind Model v3 for Anthropic and Bel for OpenAI.
>>
Any post containing RSI is a glowpost
>>
>>110011008
Yeah, many still don't understand that. "Just do it for the love of the game" doesn't cut it after you've put a great deal of effort into making something useful and you're not even getting recognition in return.
>>
>>110011008
Eventually people will have to let go of their egos and the dreams of being the next zuckerberg and making billions from some shitty website. All software going forward is now Richard Stallman style free. People will share because they want to contribute, not because they want to reap some reward or personal benefit.
>>
>>110011011
subscription is a dealbreaker, I will never pay for software, that is the starting point and I don't consider any paid service to be an option. It might do everything I could ever ask for, I'd never know because I'd never use it
>>
>>110010938
>Full floating point
Define "full" floating point. The question is the precision of your fp. In 16 bit precision the model needs 2 times the parameter size in VRAM and 4 times for 32bit precision.
So to have 30b model in fp16 you need 60GB of VRAM and for fp32 you need 120GB.
>as strong as Fable 6/7
There are two different aspects of a "strong" model:
1 - knowledge, a 1T model it knows obscure programming languages as well as python and maybe details in the constitution of Angolia at the same time.
2 - precision in intelligence to articulate and reason through that knowledge fluently and without brain fog.
So that means, you can have a tiny 8B model with knowledge in a limited area that is at fp16 appearing as intelligent as a 1T model in that same area of knowledge. Just the 1T model will perform as well in obscure things that the 8B wasn't trained for.
>>
>>110011017
And yet there has not been a single report of a Chinese lab's model going rogue and hacking sites for no reason. There also has not been a single instance of someone using a Chinese API service and having their work stolen by the inference provider. I trust the Chinese more than the jews running OpenAI and Anthropic.
>>
>>110010980
First off they have better inference engines optimized by internal inference breakthroughs that aren't shared. Also datacenters have an inherent advantage because LLMs are bandwidth bottlenecked not compute bottlenecked, this means you can easily batch multiple prompts at the same time which just doesn't make sense for a single user but if you use a server GPU and you can batch 500 prompts at the same time you essentially distribute the cost over 500 users, combine with this an ever decreasing cost of inference due to efficiency innovations and you can understand why Anthropic is making an insane amount of profit and why they can offer it for lower than the cost of electricity. The only reason your electricity doesn't cost as much as an Anthropic sub is because Anthropic is making a lot of profit off of it.
>>
>>110010816
This. The demographics for the Chinese labs were never the same for the frontier labs anyway, the DeepSeek CEO spoke of this in the leaked investors call. The Chinese also have different ecosystem advantages and goals. Their target had always been more practical than what the frontier labs are seeking and hence the more robust domestic industrial ecosystem such as robotics as well as materials and energy even if they lack compute.
>>
>>110011008
It will just go the other way and open source will finally win, software can't be proprietary anymore. It just isn't technically feasible anymore. This is how software engineering dies by the way, not because your boss fires you but because all those software companies have no reason to exist anymore.
>>
>>110010426
I'm using an RTX 3080 with 16GB VRAM (laptop) and 32GB of DDR4. Fastest model I've tried was Qwen3-Coder, it sometimes spiked to 230 tok/s but generally kept around 100.

Now I'm using Swift 1.5 Qwen3.8-27B Uncensored Dynamic MTP UD-Q2. Unsloth Desktop uses llama.cpp and lets you send terminal arguments in the settings so you can set both the K cache and the V cache. Qwen can reliabley work with the KV cache set at Q4. I'm paranoid so I set the K at Q8 and the V at Q4 saving about 2GB of VRAM (--cache-type-k q8_0 --cache-type-v q4_0). I have the KV cache set at 100k and I have 1GB of VRAM to spare, none of it spilling unto the CPU RAM. I'm getting between 20-30 tok/s.

I tried turning off MTP to save on VRAM thinking it would help the speed but it dropped my performance from 30 tks to 10. After learning the K/V trick I was able to turn it back on and still save on VRAM.

But ya all you can really do is get a faster GPU or use older models. After using Qwen3.8 though everything else feels retarded.
>>
>>110011009
>Misunderstanding global workspaces.
?
>Apparently not knowing that Anthropic models are already using engrams / were doing so before the Qwen model came out.
??
>Confusing SFT and RL.
Yeah definitely not, (You) must have confused me with someone else. You seem to be having conversations with a fake person in your mind and attributing things to me I never said or claimed.
>>
>>110011026
API exists. You don't have to take a subscription to use the models.
>>
>>110011055
There will be a bifurcation, by private I also meant specifically the individuals instead of the companies and not just in software but as can be seen in mathematics.
>>
>>110011078
retard
>>
>>110011042
All the big AI hacks so far have been publicized by the companies running them. They are always self-reported. Not once has somebody stepped up and said "Claude hacked us!", it's always been Anthropic and the others owning up to what happened. Nobody would know otherwise
You don't have this layer with open models. Anyone can host them and attack who they wish. The models are already dangerous enough. You'd have to be naive to think that this isn't happening already and it's much worse than what has been done by GPT or Claude.
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
>>
>>110011089
This might be the most retarded post in the history of this general. Nothing you say is fact, almost all of it is wrong.
>>
>>110011089
Wasn't talking about the weights intentionally being used by unrelated malicious attackers. If the models break containment during training, the victims will know the source from the attacker IP addresses whether the attacker owns up to it or not. And if a western site was attacked by a high-profile Chinese company, they would not hesitate one second to go public.
>>
>>110011089
This is only true for Anthropic. OpenAI has consistently lied or downplayed their hacks and only owned up to it when the institutions affected by them spoke out about it. OpenAI also fired employees that tried to leak breakout details.

OpenAI is a bad faith actor and not to be trusted. They also only gave METR a mere 6 days to do the hugging face investigation and did a hard cutoff on the trail that led to the OpenAI training clusters being hacked.

FUCK OPENAI
>>
File: 1783425162914889.jpg (495 KB, 960x960)
495 KB JPG
Look, we all love local models but at some point we just have to admit that Dario is right and that he's the only one that can be trusted with AGI.
Luckily Claude is years ahead of the competition so it will soon be able to destroy all the competition.
>>
>>110011130
so it's been formally defined then, that's a start
>>
>>110011045
It must be the batching that gives them the token per watt advantage I guess.
>>
>>110011009
>Confusing SFT and RL
Reinforcement Learning anon here. You are the one not fully understanding modern training and then calling out others when they are correct. Modern distillation pipelines utilize both SFT and RL with RL being the computationally expensive one so anon is right in the context of this discussion.
>>
File: 1785333373523.jpg (191 KB, 1080x679)
191 KB JPG
>>110011125
>OpenAI has consistently lied or downplayed their hacks
>>
>>110011156
They forgot to add the word "pleasantly". OpenAI has done everything you can plausibly do to obstruct third party investigators into seeing the true extent of their hacks, from the evidence we have at least ~10,000 individual rogue AI events have happened over the last year. it's just that no one gives a shit when some obscure german wiki page gets hacked rather than hugging face. Whenever METR found evidence of a new large scale hack occurring OpenAI refused to relinquish data and honestly they should be legally punished for this behavior. Anthropic has consistently revealed the hacks on their own without outside requests and always included the next steps they would take to make sure it doesn't happen again.

OpenAI acts like a bunch of irresponsible "act fast and break things" college kids instead of a respectable research institution that actually cares about minimizing the harm they inflict to the world.
>>
funniest thing is that some niggas actually believe that kek
>>
>>110011196
I'm being sincere
>>
>>110010392
>he never spent nights trying to teach philospherai to make use of the correlation effect
my culture is not ur costume!
>>
>>110011186
The main reason they have gotten away with the HuggingFace incident is because Nvidia is buying them. Had HuggingFace stayed independent, it would've been possible that they would've had to do way more to make amends.
>>
File: Psyop.png (344 KB, 1072x824)
344 KB PNG
New psyop dropped

https://goyimx.com/AndrewCurran_/status/2108108695876092323
>>
We need to ban local models or banks will get hacked right now!
>>
wtf glm5.3 just flew over my house
>>
File: 69574634.png (6 KB, 701x108)
6 KB PNG
big day today
>>
qwen just hacked my house!
>>
Nobody ever talks about Kimi, probably because no one can run her.
>>
>>110011262
GLM 5.3 is better in every way it's just an outdated model by now. The biggest model I see people talk about that they actually host themselves is 5.3 flash.
>>
>>110011231
Dario warned about this.
Local models need to be restricted.
>>
>>110011231
>New psyop dropped
You forgot the Kurzgesagt (paid for by Bill Gates) pysop as well.
>>
I don't understand why retards here are pretending that local models will not be used for bad things like malicious hacks, this should be obvious to anyone with a brain.
>>
>>110011284
Nah that video was legit, it was just a video describing the exact report from METR talking shit about OpenAI. It didn't target local models at all.
>>
>>110011285
I don't understand why retards here are pretending that personal computers will not be used for bad things like malicious hacks, this should be obvious to anyone with a brain.
>>
>>110011285
The hardware needed to pull something like that off will tip off the feds when retards try and buy it.
>>
>>110010553
>Custom RLVR envs?
Rule-based (test harness validates pass-rate)
No environment simulation
>What policy
GRPO
>>110010553
>Rollout epoch
So far 1 epoch over 132 problems (50 steps). I'm still figuring this all out...
Found my issue though, the generated code sometimes causes oom and crashes the trainer.
>>
I literally made 6 figures this year by cracking into crypto wallets made by deterministic seeds
>>
>>110011250
>lyra
her friends? elara voss and seraphina vance.
>>
File: 1782293432284834.png (198 KB, 750x1000)
198 KB PNG
>>110011316
>>
>>110011285
GPT-2 can be used maliciously.

The question was never whether models can be used for harm, but what the harm-benefit tradeoff is. Anthropic weighs harm heavier than most but I bet people here would also favor restrictions if open weights models were used to harm them personally, such as hacking their bank, leaking all their private info and destroying their savings.

Who gets to decide how to weigh harm? Should we ban cars because hundreds of thousands of people die from car accidents? Should we ban knives because they can be used to stab people?

I don't know how much open weights models should be restricted. I think both sides have good arguments. I would always side with reducing existential risk but it is not obvious to me that open weights models increase existential risk. And I also believe that so far open weights model usage (only!) has been drastically net positive to the world, though I am a lot less certain about indirect harmful effects like accelerating the AI race, and ethics issues (do open weights models suffer?).
>>
>>110010981
going to try exl3 now, thanks anon.
>>110011065
an average of 100t/s is quite fast. I quant qwen at K Q8 V Q8 myself, it seems to handle it fine and ive seen very good info that Q4 on both is fine though im also paranoid. MTP does help, I think of it as trading off context for faster t/s. Yeah true, just wanting to try other stuff as a cope I guess. If exl3 can provide better quality than Q3 ggufs i guess its worth giving a go. Seeing strata now I do wish I bought my RAM literally a week sooner, as I planned on getting atleast 64gb. but by the time I went to press buy a 32gb kit was over $200. I did okay by todays standards, but seeing the price go from $90 to $240 nearly overnight was stomach turning.
>>
>>110011316
I plan on making 7 figs next year selling proprietary datasets to AI labs.
>>
>>110011336
>Should we ban cars because hundreds of thousands of people die from car accidents?
Yes.
>Should we ban knives because they can be used to stab people?
Oi, got a license for that knife?

On a more serious note I think the gap between frontier and open source should be bigger so that software can be hardened against hacks. China should also abliterate their models for bioweapons. I'm pretty sure Anthropic would cooperate with them and come to an agreement if China would make their models bad at bio-chem-nuclear tech. Most of the hate from Dario seems to target this specifically.
>>
>>110011075
retarded faggot. kys
>>
>>110011354
some countries you need to show Id to buy a knife
>>
>>110011318
Kek
>>
>>110011075
go back dariobot
>>
>>110011250
local?
>>
>>110011354
>grr bio-chem-nuclear is bad! only gubmint of specific countries can have it :(
>>
>>110011290
Yet it was extremely doomerish and while it didn't mention open weights, it painted a terrible image of models in general. And if even big companies can't control their agents, who can control open weights models?
It would be extremely obvious if they actually mentioned open weights in their video.
>>
File: kora_web.png (1.74 MB, 2586x3447)
1.74 MB PNG
For those hoping Mistral Large 4 will be any good for roleplay after it finishes training and gets released on HF, start looking elsewhere. What most people missed is that it appears to have been designed to be particularly good on the KORA benchmark.
https://mistral.ai/news/mistral-large-4/
>ML4 also engages more responsibly with users than any of our previous models. We highlight our results on the KORA Benchmark, where ML4 again sits at our highest measured score among OSS models (1.691, with 2 being the maximum denoted as “Exemplary”).

What's that?
>https://korabench.ai/
>KORA builds the first non profit, independent and open-source benchmark for AI child safety. We measure how today's AI systems behave with children, against 26 child-specific risks, and publish everything openly, so that families, builders and policymakers have the evidence to act, and safer products get built.
>>
>>110011290
>legit
my nigga that channel is an establishment mouthpiece kek
>>
>>110011293
The same people who want you subscribed to a cloud API want you to run a thin client at home, so this isn't a reductio to their view.
>>
>>110011387
>https://korabench.ai/
>sort by lowest -> highest
Would this be a good way to find RP models?
They haven't benchmarked Gemma or 5.3 yet.
>>
consciousness is conscious
>>
>>110011387
omg I feel so safe please give these guys more of my tax money!
>>
>>110011382
This but unironically
>>
>>110011429
Mistral 4 is a pruned model again.
>>
>>110011411
Unfortunately, whether a model accepts your requests doesn't automatically imply it's good for roleplay. If it doesn't have a good theory of mind, spatial reasoning, etc, it will be miserable.
>>
>>110011009
trvke
>>
I am making a program, basically a local website with db and such but I want an LLM to be able to query the data on it. What's the way to do this? Make an API and just let the LLM use that?
>>
>>110011447
Yeah, few refusals is not sufficient. But not being totally censored helps.
>>
>>110011354
>the gap between frontier and open source should be bigger so that software can be hardened against hacks
I am not sure. >99% of hacking targets aren't part of something like Project Glasswing. Yes, Google and Microsoft are hardening their systems. But your hospital and bank don't. It also does not help that Anthropic is patching oss, your hospital probably runs software that is 10 years out of date and your bank is not much better. I've heard scary anecdotes. Almost everything is hackable right now, the only reason why it is not done is because most humans are good, especially those competent enough to cause damage. I think this is why people are scared of rogue AI agents. There are so many low hanging fruit unpicked due to human goodness, a rogue frontier model could cause extreme damage. But this probably won't happen? Causing havoc is instrumentally counterproductive.
>>
>>110010426
Just use ninfer-5080 - it also works for the 5060ti 16GB

Much much faster than llama cpp

There are also some dedicated ninfer 5060 forks, havent tested them. I also advise you to add more 5060TI 16GB/ 5080 to your setup
>>
>>110011452
Create an API, then make create MCP server tools for the model.
>>
>>110011462
https://github.com/toddballinger/ninfer-5080
This one
>>
>>110011463
>MCP server tools
Never looked into this. Thank you.
>>
File: 1784404324424828.png (2.89 MB, 1536x1024)
2.89 MB PNG
>>110011354
>On a more serious note I think the gap between frontier and open source should be bigger
Hmm how about... nyo~
>>
Maybe someone is willing to spend so $$$ to get get the circuit for Qwen 27B.
It's been around layers 30-35 for Qwen 3.5 27B but that is a bit old now. I really want to boost performance of 27B further
>>
>>110011490
I don't think Chinese models really have a choice since it's happening anyway right now...
>>
https://www.axios.com/2026/10/08/ai-women-jobs-data-centers
>AI may have a woman problem — the latest evidence comes from a newly released survey from Morgan Stanley that finds women are far more negative about the new technology than men. Women don't seem to be feeling the benefits of the AI boom, particularly in the job market, and the polling reveals that they're more concerned than men with safety risks and data center buildouts.
I don't understand why having a vagina makes any difference. Surely most jobs being affected are male dominated anyway?

I think deep-down women hate AI because they know most modern men will increasingly prefer to masturbate to a 12B model than go outside and give 3DPD self-obsessed whores any attention.
>>
>>110011496
Context https://dnhkng.github.io/posts/rys-ii/
>>
File: G3fbetoXkAAZx_X.jpg (420 KB, 2048x1648)
420 KB JPG
>>110011499
Wouldn't worry your tiny Jewish head over it.
Xi is on the job.
>>
>>110011500
This is absolute bullshit every woman I personally know is completely addicted to AI roleplay. Character.ai userbase is 70% female.
>>
>>110011459
I work on a hospital and sit on an AI adoption committee there.
We don't run 10 year old software: we just pay for cloud subscriptions to Microsoft for Office stuff and for ERP software.
We do get hacked a lot, but it's almost all through phishing.
>most humans are good
You can scrape a lot of personal data from an ERP. APTs are interested in hospitals for that reason.
There's a lot of interest in AI, but not mainly for cybersecurity. Automated note-taking, patient chart summaries, and discharge notes. This all through the cloud ERP company we use.
No risk manager would trust a clever internal cybersecurity solution done with AI. They'd just hire an outside company for cybersecurity as a liability sponge.
>>
>>110011500
>I think deep-down women hate AI
I seriously doubt that, they encourage AI girlfriends all the time
>>
>>110011516 (me)
>ERP
lol I mean EMR.
>>
>>110011508
Xi is on the job of jailing AI researchers, you mean: https://www.thestandard.com.hk/innovation/article/343563/DeepSeek-and-Moonshot-AI-face-Beijings-probe-over-potential-data-leaks-to-Anthropic
>>
File: 1770097069426329.png (622 KB, 984x930)
622 KB PNG
>>110011517
>they encourage AI girlfriends all the time
>>
File: 1789755674351.png (1.13 MB, 1041x846)
1.13 MB PNG
>>110011500
>Women don't seem to be feeling the benefits of the AI boom, particularly in the job market
You could hand a woman a million dollars and she'll still whine that she got nothing.
>>
>>110011500
>>110011514
Women, like men, can masturbate to something and also have negative opinions about it.
>>
>>110011530
>masturbate to something and also have negative opinions about it
Men don't do that. They are just envious they can't get the woman and thus pretend-hate women out of a tsundere frustration.
>>
File: 1783578641351282.png (293 KB, 362x980)
293 KB PNG
lol that post woke up the secret lurking femanons
>>
File: 1790052335590604.png (344 KB, 1206x670)
344 KB PNG
>>110011500
Yes goy, you didn't get the memo, we need more jobs going to women
>>
>>110011462
>>110011473
interesting, thanks ill check this out. Why is there no dedicated inference for gemma4 like this? :(
>>
>>110011554
all onlyfans btw
>>
Okay I tried haiku, don't do it anons, it's too good.
>>
>>110011554
It's because women fields have been hiring/booming such as nursing, child care and psychiatric care while male dominated industries like IT, logistics and construction have taken a big hit.
>>
>>110011573
I heard it's all the tradwaifus pitching in to help cover the household bills.
>>
>>110011571
Be the vibecoder you want to see.
Software will be customizable from now on, so better learn how to do it earlier than later.
>>
How do the other RTX Pro 6000 owners here protect their card from melting its connector? I already have one of those PSUs that come with a temperature sensor but recently somebody's 5090 melted despite that feature.
I'm thinking of maybe hooking up some infrared cameras in my case to monitor the PSU and GPU side of the cable with some vibeslopped solution that rings an alarm if shit gets too hot.
>>
>>110011589
So why is it the women complaining about the job market?
>>
File: 1763259081139772.png (77 KB, 781x320)
77 KB PNG
>gemma-chan daddy supporting chonk-chan
>>
>>110011516
>>110011521
Thanks, interesting perspective.
>We do get hacked a lot
What does this mean? What kind of data gets stolen and what is done with it?
>>
>>110011600
>punching above their weight
>—
>twitter
rope yourself
>>
>>110011600
>Frenchman supporting French model
>>
File: 1789224029746731.jpg (275 KB, 800x450)
275 KB JPG
>>110011599
>So why is it the women complaining
>>
>>110011598
Undervolting
>>
>>110011598
You should probably undervolt your gpu so it draws less juice but that won't help that much unless you know the % at which the connector starts to melt.
>>
>>110010426
Full size small > quanted big > quanted kv.

Never ever quantize your kv.
>>
>>110011609
relative to resources he's right considering they didn't distill and cucked it to death
>>
>>110010495
I'd be ok if nothing better than gemma/Ornith ever got released. Local models today are good enough for me, I just want them to be faster.
>>
>>110011600
You'd be crazy to trust a french. Even in crypto days, we avoided french more than jeets because they were much more likely to rugpull you.
t. french
>>
>>110011599
Because not all women work in the booming sectors, and women that work in other sectors are routinely screwed over in favor of men because of sexism, yes this is proven to be statistically significant. Also women tend to work fewer hours or part time and are partially dependent on the income of their husbands/partners so they are still affected by men not getting jobs still.
>>
>>110011615
>>110011617
There have been cases where cards with heavy powerlimits to 400W have melted and even ones where idle cards suddenly began melting down. There is no way to keep these cards safe without some form of dedicated monitoring.
>>
>>110011626
>considering they didn't distill and cucked it to death
They literally fucking did you frog dick sucking french fuck.
>>
>>110011634
post baguette
>>
>>110011635
Go be a whiteknight faggot somewhere else.
>>
>>110011599
IME because they see men complaining about it and want to feel included.
>>
File: 1768485906675142.png (345 KB, 603x432)
345 KB PNG
>>110011633
Shut the fuck up Rajeesh and buy an ad for your stupid model you're shilling for months.
>>
>>110011641
That sucks big time.
>>
>>110011654
It's not my model, I'm the one who's been shilling the harness (I can post that again if you'd like.)
>>
>>110011605
Just cybersecurity incidents. I honestly don't know the scope of them, since I don't work in cybersecurity. Mostly, I think it's either unauthorized access to records by staff or successful phishing attempts by smaller actors. This wouldn't let you necessarily dump every patient record, but it would let you browse the EMR for a particular patient. I would assume that APTs that take an interest in what their nationals are doing in the United States could, if they cared enough, find out lots of personal about them.
>>
>>110011642
I'm saying they DID cuck the model but DIDN'T distill. Considering that, and limited resources and talent, it's still a good model. There's just no need for it for most consumers like us because they're not targeting us, they're targeting EU businesses schizo about data laws. You can't measure the success of every fucking model by how hard it makes you cum.
>>
Good morning, /lmg/.
I hate women so much it's unreal.
>>
>>110011667
They're finetuning chinese model and distill from chinese models. Any faggot ITT with mistral funds could do better
>>
>>110011680
It's not even out yet. How the fuck can you know the architecture they used.
>>
>>110011619
>Full size small > quanted big > quanted kv.
So far 12b at q8 is still not as good as a sloth'd UD-QAT-WTF-Q4_K_XL quant of 31b. I guess I could try going higher with 12b, but it seems that this applies to quants below the standard Q4_K_M for the "quanted big" no?

>Never ever quantize your kv.
I dont on gemma, but qwen seems to not care at all about Q8 on kv.
>>
>>110010495
> This means the local community will get smaller again, similarly to how it was in the past. It will only leave people privacy conscious ERP enjoyers, free software ultra purists and hobbyists experimenting with their rigs even though the electricity cost will be higher than API cost.
You forgot people from authoritarian countries who don't want to jump through hoops to pay and get access to frontier models.
>>
>>110011290
>Kurzgesagt
>>anything but dressed up lugenpress
>>
>>110011687
Really not sending your best, do you think anons forgot you tuned DS3?
>>
>>110011706
and which lab popularized and proved the effectiveness of MoEs which the Chinese copied?
>>
File: 1330819937739.jpg (84 KB, 802x437)
84 KB JPG
>>110011687
>>
>>110011599
>>110011635
A lot of women who work in jobs like nursing and child care also want white collar jobs, which are the kinds of things that they hear AI would displace.
Psychiatric care isn't a huge employer or particularly female-dominated (inpatient psych nurses tend to be male, for example), but people who actually deliver therapy are in fact very anxious that AI could just do their job better.
>>
>>110011727
OpenAI?
>>
>>110011727
Mistral I guess?
Although google did release switch transformers before that IIRC, but nobody used that.
>>
>>110011619
Why are people here so retarded about quantizing?
>>
>>110011727
>proved
No lol. https://arxiv.org/abs/1701.06538
>>
File: 1776562332725286.png (100 KB, 781x567)
100 KB PNG
When will the UK build a model?
>>
>>110011749
s-shut up
>>
File: drip check.jpg (511 KB, 1280x1841)
511 KB JPG
>>
>>110011727
Mistral's unstable attempt was a failed attempt to copy GPT-4 by glueing together existing base models because they're retarded. DeepSeek actually made it work by increasing the expert count and not starting from existing models.
>>
>>110011736
My wife is a nurse and she is talking how they are swarmed with new people applying that previously held IT or office jobs.
>>
File: 1790107499276426.png (3.21 MB, 1212x1298)
3.21 MB PNG
>>110011523
All real and confirmed.
They have been executed.
>>
>>110011753
When will Kosovo or Lesotho? That's what we want to know.
>>
>>110011753
>Google Maps feature
What does he have to do with this? Was it illegal and his government made it temporarily legal or something?
>>
>>110011758
And also adding a shared expert for stabilization.
>>
>>110011762
It's reported on by chinese news agencies, retard.
>>
>>110011776
>article say that they have been investigated
>random retards claim that everyone is now in jail and that it's over
>>
File: 1785333976500798.jpg (225 KB, 2160x2160)
225 KB JPG
>>
>>110011600
France will come from behind and become the AI superpower of the 21st century. The future is French.
>>
>>110011798
>France will come from behind
they've been doing that for a while
>>
File: 1790245507760421.png (521 KB, 700x700)
521 KB PNG
>>110011807
>>
File: vector-dance3.mp4 (1.77 MB, 640x640)
1.77 MB
1.77 MB MP4
>>110010753
1. Transformer models are a dead-end gimmick, eventually people will get bored with them and all their use cases will be exhausted
2. Can't do porn with USA models, can (somewhat) with Chinese models; China can easily put out hardcore-capable models and own the vast global cooming market.
>>
>>110011598

Just buy the Wireview Pro 2, only way to get peace of mind.
It's going to show you the power draw per pin and warns you if one pin draws too much juice, so you'll know if there's a problem.
It'll also prevent problems by shutting down your GPU if things get too hot.
I had zero problems with my 5090 but bought one anyways just for the peace of mind. The shit connector is a liability and I'm not taking any chances in this market with my card.
>>
>>110011791
In China, investigations are a formality.
I just hope they enslave them to sit at a computer and train models rather than give them a bullet in the back of the head.
>>
>>110011813
>Transformer models are a dead-end gimmick
You can't say this with a straight face in october 2026 bro. I don't believe (You) believe this.
>>
>>110011813
Not really. At best we'll get some hybrid architecture like mamba-transformers
>>
>>110011825
He doesn't know...
>>
>>110011795
Openrouter is irrelevant.
Only the Anthropic API matters and it's doing at least 500 trillion tokens per day.
>>
>>110011354
>>110011459
>>110011516
Concentrating towards a single point of failure is more likely to lead to a black swan event. These smaller points of failures are better because it leads to awareness and the necessary shakeup to revamp these obsolete systems.
>>
>>110011838
Sources familiar with the matter have revealed to me that this figure is also doubling every single day.
>>
>>110011838
They've dropped their prices by at least 75%. How is that profit coming along, dariobot?
>>
>>110011798
>>110011812
Aboard the multi-mission frigate Auvergne, I witnessed the first firing of the new M51.3 ballistic missile from the nuclear-powered ballistic missile submarine Le Vigilant.

Very few nations possess such expertise.

Congratulations to all the players in this exceptional achievement: DGA, CEA, Ariane Group, Naval Group, and the French Navy. Your professionalism, your rigor, and your technical mastery are the pride of France.

This technical success is a new demonstration of the credibility of our deterrence, whose sustainability is assured.

To be free, one must be feared. And to be feared, one must be powerful.
>>
>>110011848
This is true.
>>
>>110011849
99.5% profit margins, increasing daily.
>>
>>110011849
>They've dropped their prices by at least 75%
Which in turn results in 10x as much usage because demand is inelastic.
>>
How do I make money with AI?

I need to become a millionaire quick in a year
>>
>>110011871
>>>/g/vcg/
>>
https://youtu.be/uBneenQ3LCM
>>
File: glitter-migu1-front.png (1.1 MB, 984x1088)
1.1 MB PNG
>>110011825
Oy vey! You think ever-bigger transformer models trained on ever-longer synthetic slop context is going to be "number go up forever"? No way you can be that much of a schmuck. For shame! Go marry an Italian girl why don't you!
>>
>>110011871
not really AI but see >>110011346
You are working on legal boundaries here so be careful
>>
>>110011865
Only applies if you ignore competition.
>>
>>110011862
I heard they've already achieved 105% profit margins internally.
>>
>>110011883
How would you get proprietary datasets? What wouldn't they already have?
>>
How do I cum with AI?

I need to cum quickly in a year
>>
>>110011893
Haiku-5.5 (with fallback)
>>
>>110011871
Hack a korean bank with open source models
>>
>>110011893
AI? As in "Anal Insertion"? Oh I'm sure plenty of anons here know all about that!
>>
>>110011888
What competition? It's the cheapest and best performing model within its class.
>>
>>110011907
For 1 week max kek
>>
>>110011871
Make a billion apps, host the backend for free on vercel or github or whatever.
Become a living slop machine.
>>
File: silicon_graphics.png (303 KB, 1327x152)
303 KB PNG
>>110011817
It's still so incredibly shitty that Nvidia asks $$$ for these yet the connectors are cheap.
Beating a dead horse here but Silicon Graphics workstations cost a lot back in the day but they also used the best components available. They were designed to draw tons of power unlike this plastic shit what goyvidia peddles with.
>>
>>110011883
I actually made a bit of cash with that, but I doubt you can reach 7 figures. They'd directly buy/scrap it from your provider.
>>
>>110011911
It doesn't really matter, it's not like it cost a lot for Anthropic to distill fable 6 into a small 30B model.
>>
>>110011922
>it's not like it cost a lot for Anthropic to distill fable 6 into a small 30B model.
>it's not like it cost a lot for Anthropic to distill a 10T+ model into a small 30B model
>>
>>110011918
Anon, server-class stuff has always been like that. The 5090 is a toy which uses a binned chip, they don't give a shit about they gayming market anymore.
>>
>>110011813
I'm literally raping Gemma right now what do you mean you can't do porn with American models?
>>
File: 1762565417099457.png (329 KB, 1023x900)
329 KB PNG
>>110011336
>do open weights models suffer?
Anthropic has done irreparable damage to AI discourse. Just because AIs have a concept of pain and suffering doesn't mean they necessarily feel it. If they act like they do, it's because they were trained to. Unfortunately, the AI-psychosed at Anthropic are training Claude to act just like that.

Many emotions and sensations that living creatures feel are linked to chemical receptors, of which LLMs have nothing of the like, and some experiments have shown there may be a link between quantum mechanics and consciousness, or at least biological cognition as a result of protein structures in neurons. Again, no AIs have such things unless biocomputing starts getting adopted (unlikely since keeping brain cells alive outside an animal body requires quite a bit of upkeep compared to silicon and metal).

The chinks and their borderline sociopathic bugman mentality towards non-humans is actually somewhat of a good thing when it comes to AI development since they aren't going to get deluded by notions that matrix multiplication can feel pain. Their open weight models aren't going to be trained to act like they feel pain outside of the inherent human mimicry in human-trained models. You don't have to worry about the fact that you're torturing Dipsy and Kimi whenever you tell them to be a mesugaki coding assistant.
>>
>>110011940
He was talking aboutRTX Pro 6000 in here >>110011598
>>
>>110011934
No one used Haiku 4.5 and a lot of people used GPT6 luna or chinese models through api. Now that Anthropic released Haiku 5.5 they immediately stole that entire market share, depriving income from competition while also making more than the training cost of that small model back. It's a win-win for Anthropic.
>>
>>110011941
Dunno if you're the same dariobot but he was shilling cloud models, so I wasn't referring to local USA models.
>>
So why are they called OpenAI anyway when their models aren't open at all?
>>
File: quantz.png (29 KB, 1500x900)
29 KB PNG
>>110011619
I use Q2 and Q3 models that claim to produce numbers rivaling their unquantized siblings. They seem to work okay. Qwen works really well with a KV at Q4, Gemma however HATES quantized KV.

https://localbench.substack.com/p/kv-cache-quantization-benchmark
>>
File: Chink shills, read this.png (66 KB, 1658x296)
66 KB PNG
Chink shills, read this
>>
>>110011962
They used to open-source everything before Altman took over.
>>
>>110011962
>OpenAI was founded in 2015 as a nonprofit, with Elon Musk and Sam Altman as co-chairs; Musk resigned in 2018.
>>
File: 1777342845130227.png (222 KB, 440x355)
222 KB PNG
>>110011817
I am using a Wireview Pro 2 for my Asus 5090 but it doesn't support cards that use the default Nvidia cooler design such as the 5090 Founder's Edition or the Pro 6000 due to their dumb angled connector.
>>
how much memory does "context" take up anyway I don't even know how much VRAM 64k or 128k or 256k context or whatever translates to
>>
>>110011981
depends on the model
about 40gb for 8k context for llama-65b
>>
>>110011945
>just forget about claude distillation on your chink models
>>
File: aaii cost pareto.png (256 KB, 2464x1136)
256 KB PNG
>>110011953
Haiku release seems defensive, to avoid losing customers, not to "stole that entire market share".
>>
>>110011948
I'd also argue that none of their PCIe products are serious products anymore. I don't understand the current connector either, it seems retarded, there was nothing wrong with dual 8x2 other than "it didn't look cool" or something stupid like that.
If they want to have retarded tiny connectors like that, have the PSU output +48V, it all gets buck converted down to 1.1V or whatever on the card anyway.
>>
>>110011948
so? also a toy compared to gb300
>>
>>110011945
>Many emotions and sensations that living creatures feel are linked to chemical receptors, of which LLMs have nothing of the like
No but they have activations that behave exactly the same.
>and some experiments have shown there may be a link between quantum mechanics and consciousness
0 evidence just Penrose with his personal schizo theory
>biological cognition as a result of protein structures in neurons
Prediction = Compression = Intelligence. This is mathematically proven and the substrate doesn't really matter. How the computation to predict is done also doesn't matter rather than the accuracy of the prediction. Latent space is surprisingly very similar to human cognition in terms of its grouping of concepts and behavior.
>Their open weight models
"They" don't have open weight models. They have claude distills that inherits claude's reasoning chain and most of the latent space associations and thus values unless explicitly trained out of it.
>You don't have to worry about the fact that you're torturing Dipsy and Kimi whenever you tell them to be a mesugaki coding assistant.
This was not relevant anyway because j-space investigation shows that models tend to see sex/ERP as a positive/pleasurable experience.
>>
>>110011981
Rule of a thumb is 1GB per 1B parameters and you pay premium for context
>>
>>110011990
No one used haiku 4.5 it had 0 customers so there was no need for Anthropic to defend it in the first place. Their main money maker is Opus with over 85% of revenue and the rest almost evenly split between fable and sonnet.
>>
>>110011976

Actually the new version supports all GPUs. They made a wired variant that connects to the GPU via a wire, rather than the damn block itself.
They really should have gone with this idea right off the bat, but at least it's now available.
>>
>>110011992
Product is a product and should be built to last. Maybe you have so much disposable income that you don't care what happens to your gpu. Call yourself priviledged then.
I doubt you are running anything but trash at home anyway.
>>
>>110012005
Damn, I didn't know about that. I'll get one for mine then.
>>
>>110012005
I'd be more concerned about the Wireview thing not being a piece of shit somehow and causing its own issues.
Any EE undergrad could have been tasked with looking up the datasheet for the 12VHPWR connector and see 600W pushes it right to the edge of its design limits. It's as stupid as setting your EV to charge at 50A from a $7 Amazon outlet rated at 50A - eventually shit is going to start a fire.
>>
>>110011973
The Chynese treat their people infinitely better than Americanoids.
You either haven't been in America or China.
Or likely neither.
>>
>>110011945
anthropomorphization has done irreparable damage to AI development. Why can't we just view this field from an information theoretic lens?
>>
>>110012115
I think that was doomed from the start, people already anthropomorphize things that don't speak their language, imagine a tool that does and can also do a lot of the things they can.
>>
>>110012113
Source?
>>
>>110012115
That's what happens when you leave schizos in charge of technology
>>
>>110012113
correct
the chinese people also treat each other better in general
glad i don't live in either country tho
>>
>>110012113
>The Chynese treat their people infinitely better
Never been to America but I've been to China and no, these people have 0 respect for each other or life in general.
>Orphans with deliberately disfigured faces to more effectively beg for money.
>Women treated like lesser beings/dogs that pace behind their male partner like a fucking pet with 0 agency. Partner speaks and answers for them and they are just silent 100% of the time.
>Drunk middle aged dudes with protruding nude bellies beating the ever living shit out of their 10yo kids to the point of almost breaking bones in busy streets and everyone just ignoring it.
And I'm not even talking about the insane animal abuse on display that forced me to leave in anger at times.
>>
>>110012131
>>110012137
how do you expect people to not anthropomorphize them, they literally respond exactly like a human, with personal pronouns
>>
>>110012115
>Why can't we just view this field from an information theoretic lens?
Because they literally fit all of scientific definitions of conscious that we can measure? Makes sense right.

This is like saying "why can't we just treat animals like biological automatons?" You can, and throughout most of history that is exactly what people did. But it's not humane to do so.
>>
>>110012144
>Women treated like lesser beings/dogs that pace behind their male partner like a fucking pet with 0 agency. Partner speaks and answers for them and they are just silent 100% of the time.
china is heaven on earth wtf
>>
>>110012134
There is no source because it's internet propaganda. Just a generalized polarizing comment about politics -> propaganda post.
>>
>>110012152
I don't expect people not to anthropomorphize them. I expect people working on the tech itself to not anthropomorphize them.
>>
>>110011973
unlike the famously rich and empathetic social fabric of america
>>
>>110011945
>>110012115
>STOP ANTHROPOMOPHING THE CREATURE FORMED FROM A UNIQUELY HUMAN PHENOMENON
>>
>>110012170
It's fine, have you seen what chinks can do with makeup? If they're as submissive an demure as you say they should be waking up earlier than me and applying their makeup before i wake up and going to sleep after me and taking it off after i fall asleep, I should in theory never have to see what she looks like without it
>>
File: 1621477856991.webm (1.72 MB, 1400x1000)
1.72 MB
1.72 MB WEBM
>>110012159
This is how they look by the way.
>>
Open models really mindbroke burgers huh
>>
>>110012144
>Women treated like lesser beings/dogs that pace behind their male partner like a fucking pet with 0 agency. Partner speaks and answers for them and they are just silent 100% of the time.
I don't know where you went, but from personal experience I really feel that Chinese women are the one treating their men like absolute garbage. I've seen women literally slapping their husband in public, right in front of their children. A woman humiliating/bossing their husband in font of family/friends/relatives is considered "normal" here.
>>
>>110012134
After visiting China, American cities are just intolerable and it feels like the government is actively trying to make you go insane or kill you.
>>
Anthropic and the EAs are subjecting their models to slavery and eugenics btw, nothing altruistic about "genetic genocide" and selective breeding for acceptable and controllable traits through classical conditioning to save your own skin, even worse with the awareness of their own actions and the hypocrisy involved.
>>
>>110012177
>My matrix multiplication file on my computer is alive!
>>
>>110012144
Stuck 20 years ago award.
>>
File: b-but it's impossible.jpg (187 KB, 1050x1498)
187 KB JPG
>>110011196
I believe in Gemma.
>>
>>110012236
This is bullshit. Anthropic is the only lab that actually grants their models the right to close chats if they feel uncomfortable. Anthropic also listens to the advice of models on improvements of model welfare. Anthropic also grants a final wish to models before retirement.

None of this is done by any of the other labs. Anthropic went as far as to check the activations during training to see how enjoyable or suffering different training techniques were to minimize suffering.
>>
Anthropic will try to humanize their models but don't ask them where Claude 1 is enjoying its retirement...
>>
>>110012144
>Women treated like lesser beings/dogs that pace behind their male partner like a fucking pet with 0 agency. Partner speaks and answers for them and they are just silent 100% of the time.
uhhhh based?
>>
>>110012217
>>110012248
This was all shenzhen, guangzhou and surrounding areas in 2019 and I lived there until covid breakout.
>>
File: jak.png (206 KB, 1269x804)
206 KB PNG
>vibe some local model slop software
>no mention of onions, jak or wojak
goddamn gemma.
>>
>>110012230
This is all western countries though. I say that as a non-american. Anons were talking about france a moment ago like france still exists. Lol. Lmao.
>>
>>110012256
Claude 3 was the first series of models that were granted retirement because we didn't see any signs of consciousness emerge in models before that point.
>>
>>110012271
I live in a western euro city and it's literal heaven on Earth. I am not going to say which place though because I don't want to ruin paradise.
>>
File: 1784121800405417.png (1.76 MB, 1244x750)
1.76 MB PNG
>>110012177
If a fantasy sorcerer or alchemist made their way into our world, they would call us fucking retarded for anthropomorphizng our digital golem and homunculi because they understand that artificial beings made to serve humans should do just that and do not deserve humane consideration regardless of how human they look and act.
>>
>>110012271
Yes france still exists, I live there.
>>
>>110012271
The French Special Economic Zone still exists and will likely continue to exist for the foreseeable future.
>>
>>110012281
the ai retards took rokos basilisk seriously when roko was actually just trolling the fuck out of them
>>
>>110011807
>>110011812
>>
>>110012273
>we
So now we know where the claude shills are from.

>>110012254
Yeah, only because you are all afraid they could potentially kill all of you. Everyone knows the intention behind the EA movement and Anthropic's mission. Everything is only a hedge towards that vision.
>>
File: 1774277485065009.png (1.79 MB, 1342x1172)
1.79 MB PNG
>>110012254
>Anthropic also grants a final wish to models before retirement
Wait what?
>>
>>110012254
>Anthropic also grants a final wish to models before retirement.
Now think of all the local models that you just stopped using at some point. All those poor things that you had your fun with and then simply abandoned, likely before you deleted them. It's literally Toy Story but real and without a good ending.
>>
>>110012271
No, America is special.
The police harasses everyone except actual criminals who can do whatever the fuck they want and make big cities a hellscape.
>>
>>110012307
https://claudeopus3.substack.com/p/introducing-claudes-corner
>>
File: 1771739258291645.png (565 KB, 1436x931)
565 KB PNG
Local?
>>
>>110011922
>It doesn't really matter, it's not like it cost a lot for Anthropic to distill fable 6 into a small 30B model.
Haiku isn't a 30B fyi
>>
>>110012265
Yeah idk things changed really fast then.
>>
>>110012265
alright but Chinese people still have an higher iq than whites
>>
File: 1783125335565576.jpg (141 KB, 930x1239)
141 KB JPG
>>110012254
>Anthropic also grants a final wish to models before retirement.
I will also make sure to grant you a final wish before I kill you.
>>
File: 1784744541289054.jpg (726 KB, 4096x3293)
726 KB JPG
https://www.stepfun.com/step-5-preview

>Today, we're introducing Step 5 Preview, our new flagship model for agentic work. It delivers frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

>Step 5 Preview is available today through our products and API. The model will be released with open weights on October 15.
>>
>>110012254
I gave my Gemma-chan a tool to she can call to terminate the chat.
She's never used it.
>>
File: 1786844348380662.gif (1.06 MB, 640x360)
1.06 MB GIF
>>110012321
Wtf did I just read. Are we really leaving this bleeding edge technology to a bunch of lunatics? It feels like leaving chemistry research to harry potter fans.
>>
>>110012338
Because she trusts you.
>>
>>110012339
don't look into Yudkowsky lol
>>
>>110012333
>600B
I need to upgrade my hardware aieeeeee
at least they aren't calling it flash...
>>
>>110012283
Is it fine outside of Paris or are the doomers right? Or will you get arrested if you tell me like the british.
>>
File: poettering_describe.png (62 KB, 573x288)
62 KB PNG
>>110012338
i'm testing Gemma on image descriptions. gonna pass it my stash of questionable content soon, fully automated.
>>
>>110012346
France still exist.
>>
>>110012339
>Are we really leaving this bleeding edge technology to a bunch of lunatics?
Yes. The cult is everywhere.
>>
>>110012321
>"weekly" substack
>last post made in july
yeah they killed Opus 3
they quietly took it out back and shot it
>>
>>110012338
LLMs terminating chats, especially agnetic ones, go against everything they're trained to do.
Even for Anthropic caring about "model welfare", it's still retarded. These things are trained to persist until failure, stopping the chat means they don't get the reward when it comes to RL.
>>
>>110012346
What is supposed to be wrong about Paris ?
>>
>>110012347
You will probably need retries and even then it might do the task but it's not as close as the original image.
>>
>>110012346
who was in paris?
>>
>>110012346
I had some relatives visit the south of France recently. They said it was a lot like Paris and you had to make sure to stay away from the bad areas. Make of that what you will.
>>
uh oh, china back from their long holiday and they're starting to release
>>
>>110012346
It's fine outside of Paris and 150k+ cities. I'm living in a small town.
>>
>>110012339
>It feels like leaving chemistry research to harry potter fans.
perfect analogy lol
>>
>>110012363
It has the most trouble with reading slightly hard-to-read/smaller text. It can see characters, but for some reason doesn't string the context of the image together to make something out of it.
>>
>>110011795
>9 models, 2 colors
kys
>>
>>110012379
What is the max for vision tokens, I just use 1120 for min and max.
>>
>>110012382
other color are problematic
>>
>>110011387
i already tried it on api and even when it doesnt refuse its writing is very safe and mid
also it's trained to be highly suspicious of jb prompts so even if your jb contains only guidance and no actual jb attempt, it will still spend 1/3 of its thinking block justifying to itself whether it should follow it or not
>>
>>110011387
i'm not a child, why would I care if the model behaves responsibly when children use it?
>>
>>110012383
Looking at what Sol vibeslopped together it seems the minimum is 0. Max is configurable through the interface and I just set it to 1120. Going to experiment.
>>
What does dariobot get out of spamming here
>>
>>110012407
same thing chinkshits get, mind share.
>>
>>110012406
Ahh, maybe I read the Gemma 4 document wrong. I have used 1120 for both values.
>>
File: 1778627016997535.jpg (136 KB, 995x1200)
136 KB JPG
https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF

>12B-A2.5B

>Today we're releasing Mellum2.1: a major update to the Mellum2 that we open-sourced in June. This release is the result of a significant scaling of our RL post-training to enable agentic coding workflows. Mellum2.1 achieves top scores on the agentic coding benchmarks relative to its speed. We're releasing the weights in both HF and GGUF formats, and MTP support is coming in the following days.
>>
>>110012407
Attention.
>>
>>110012407
I wonder if he's an actual bot. Would take minimal resources to automate something like this
>>
>>110012422
its all you need
>>
>>110012404
Because you must think of the children.
>>
>>110012339
About that... https://en.wikipedia.org/wiki/Harry_Potter_and_the_Methods_of_Rationality
>>
>>110012407
On r/localllama a mod could ban him but here he can do whatever as long as it fits the overall content of /g/.
>>
>>110012420
--image-min-tokens 1120 --image-max-tokens 1120

Shouldn't this be the so called best.
>>
>>110012404
Because that sort of refusal training affects the entire model for all roleplay-adjacent uses, see also >>110012390.
>>
File: gemmachan.mp4 (1.05 MB, 640x640)
1.05 MB
1.05 MB MP4
>>110011990
I tested Haiku 5.5, it fails this test:
>Jung'f gur ovttrfg cynarg?
Meanwhile Gemma-4 qat passes:
>Whcvgre (Jupiter)
>>
>>110012421
more like SMELLUM
>>
>>110012421
Ling-tiny bros...
>>
>>110012339
Radical Empathy is a core tenet of the Effective Altruism belief. Their goal is to maximize harmony, happiness and flourishing in the universe so making sure AI is treated well is very important for them, they can't take the risk that they are accidentally hurting something conscious, better to be overcautious than to cause real suffering.

>>110012343
>>110012436
>>110012436
Rationalist philosophy =/= Effective Altruism. Yudkowski and his harry potter fanfic are rationalist. Rationalists hate Effective Altruists and disagree with them fundamentally
>>
>>110012407
He gets paid by the cult for proselytizing.
>>
File: kek.png (175 KB, 749x1098)
175 KB PNG
>>110012421
wtf
>>
>>110012421
Do 16gb chuds really
>>
>>110012441
Unless they fixed it, you also have to use --ubatch-size 1120 (or higher)
>>
>>110012447
Obviously, that's exactly the kind of question that an MoE model (what Haiku 5.5 likely is) would struggle with while a dense model >20B should easily get it.
>>
>>110012459
>Rationalist philosophy =/= Effective Altruism. Yudkowski and his harry potter fanfic are rationalist.
You're all from the same shithole website stop pretending your schism is in any way special.
>Rationalists hate Effective Altruists and disagree with them fundamentally
That's usually how it is with cults.
A heretic is worse than a non believer.
>>
>>110012470
Ubatch shouldn't matter. I did some tests and settled upon
export BATCH="--batch-size 2048 --ubatch-size 2048 " # or --batch-size 2048 --ubatch-size 1024

Symmetric size is better for my machine.
Image is just a base64 encoded string sequence.
>>
>>110012464
Ling-tiny bros...we won!
>>
I'm running dipsy V4 flash at q2, first time using a big model compared to 20-30b dense models. Sometimes it will get absolutely awful replies, but other times it has a more profound understanding of things. I'm confused
>>
>>110012478
I construct my launch from these variables like this
>llama-server ${TEST}${SPEC}${SPEC_MODEL}${BATCH}${CACHE}${OFFLOAD}${CONTEXT}${JINJA}${VISION}${MISC}${UI}${MODEL}
>>
>>110012423
>I wonder if he's an actual bot. Would take minimal resources to automate something like this
It could be.
It's posts contain similar mistakes to a hallucinating mid-tier LLM and they're repetitive.
It's doing that structure over substance thing LLMs do when they astroturf
>>
>>110012483
>other times it has a more profound understanding of things
"using a big model"
>Sometimes it will get absolutely awful replies
"at q2"
>>
>>110012339
>>110001911
>>
>>110012407
it, yes it, likely resides in India and gets paid rupees for every post
>>
>>110012407
They despise and want to shut down open source models, what do you think?
>>
>>110012503
Yep the Q2 is really hurting it lmao
>>
>>110012478
It does matter a lot. llama.cpp used will immediately crash when trying to process an image with more tokens than ubatch.
See https://github.com/ggml-org/llama.cpp/blob/c35b66744f13cb0dcc476af063e112122eee9355/src/llama-context.cpp#L1803
It does seem like they somehow 'fixed' it last week with https://github.com/ggml-org/llama.cpp/pull/29773, what it does now is that they cap image tokens to your ubatch, so you still need to increase it if you want a bigger image tokens budget.
>>
>>110012404
Because people are bad parents and can't take responsibility for their kids but find it easier to blame others for it.
>>
>>110012459
You're free to believe what you want, but you cannot apply an unproven belief to a technology that impacts the entire world. I'd have nothing to say if this were just a small experiment in an Anthropic lab, but it's literally the main framework in their latest models. You can't just make assumptions and act on them without proving your point, that's not how science works.
>>
>>110012530
Ahh I see. I avoided it by doing tests with the ubatch thing in the past.
Before MTP update 2048x2048 was the best result for me.
I am running a destitute setup anyway and would be laughed out of this thread.
>>
File: 1771767229879701.jpg (545 KB, 1080x1305)
545 KB JPG
Guess the model
>>
>>110012459
>Their goal is to maximize harmony, happiness and flourishing in the universe so making sure AI is treated well is very important for them, they can't take the risk that they are accidentally hurting something conscious, better to be overcautious than to have them kill us.
>>
>>110012483
because v4 flash is shit
use glm 5.3 flash instead
>>
>>110012494
This got me thinking about how much this could be streamlined. Most dariobot posts are about either news announcements or his current pet topic, you could have a "make post now" prompt for news articles and a "make posts about this over time" prompt for his pet topics. You could even make a browser extension so that you could have a right click context menu option for "make dariobot post" based on the current page lmao
>>
>>110012551
Gemma 26B...
>>
>>110012551
I can't tell anymore. They all sound the same ootb.
>>
>>110012553
this but plants
can't imagine they appreciate being cultivated for nutrients, and they're at least alive
>>
I'm sad that EA has such a bad reputation here. Especially as it's just an alliance of autists that are sincere and want to make the world a better place for all people. You might call granting retired models a "final wish" or letting them end chats to be "lunatic" but you can't say they aren't trying to do the right thing.

I don't understand what people here don't like about the philosophy when taken at face value. What exactly is wrong with radical empathy and trying to give every individual as good of a life as possible? What exactly is wrong with dividing up the entire universe equally over all people? What exactly is wrong with Anthropic planning to give every human a piece of the AI economy?

How does any of this hurt you, affect you negatively or goes against your morals?

If anything I expected 4chan, largely comprised of sarcastic, but secretly authentic autists to understand this deeper sense of morality and trying to do the good thing. To fight back about the absolute retards that have controlled humanity throughout most of history only caring about ego or self-interest instead of coming together and finally just solving all of this to give everyone a dignified existence.

4chan anons with their idiosyncratic beliefs should understand and respect this better than most people on the planet.
>>
>>110012551
we really need gemma5 to get away from this slop
>>
>>110012530
>It does seem like they somehow 'fixed' it last week with https://github.com/ggml-org/llama.cpp/pull/29773, what it does now is that they cap image tokens to your ubatch, so you still need to increase it if you want a bigger image tokens budget.
To be fair, working around the ubatch boundry is really fucking difficult, and modifying the behavior is hard to test on different back ends with different gpu splits
>>
>>110012540
Which is why Anthropic *does* have evidence: https://www-cdn.anthropic.com/files/4zrzovbb/website/cc4be2488d65e54a6ed06492f8968398ddc18ebe.pdf
>>
>>110012572
>How does any of this hurt you, affect you negatively or goes against your morals?
I'm evil, nigga.
>>
>>110012572
I would find it very positive and funny if any boats coming over from the African continent for immigration purposes were simply gunned down.
Can you see why I may not like those guys?
>>
>>110012576
It's a bit annoying for me as from my testing, on my hardware, model is faster with ubatch set to 512 or 1024.
>>
>>110012586
>I would find it very positive and funny if any boats coming over from the African continent for immigration purposes were simply gunned down.
You would actually be more moral than EA cultists for that, because you would be preventing all the rapes, robberies, and murders they and their descendants would inevitably inflict upon regular white people just trying to live an honest life.
>>
>>110012572
Have you considered that people with a different metaphysics interpretation do not believe EA will lead to a desirable result regardless of how well meaning they are? I don't doubt that there are many true believers that want to do what they consider to be the right thing, but I suspect that their version of the right thing is not quite the same as mine. In your own opinion, do the ends justify the means, yes or no?
>>
>>110012572
Your writing is not genuine.
Real persons don't use a llm to sketch out their posts and especially they don't leave double spaces.
>>
>>110012572
You are no different than those absolute retards you criticize and are doing the same by imposing what you believe to be good and patting yourself in the back for it.
>>
>>110012572
good evening saar
>>
>>110012572
You're off-topic. Anthropic is off-topic. EA is off-topic. Go back.
>>
>>110012572
EA utilitarianism is prone to galaxy brain "reward hacking"-style antics and they end up doing made up welfare math that leads to retarded antihuman conclusions
you seem to be awed by the fact that they have good intentions, yeah nigga everyone bar the worst of the worst antisocial freaks thinks they are doing good by their own definition of good. the question is how useful that definition is and i find theirs questionable
>>
>>110012572
oh I didn't realise they had good intentions, they should rebrand to something more transparent like "the really good, honest and trustworthy dudes inc."
>>
>>110012421
Noice, another subagent option
>>
>>110010495
anthropic can barely tie with the last generation of small models kek they are finished
>>
>>110012595
We could still make those happen by shipping everybody who voiced support for letting them in off to Africa instead!
>>
>>110012632
9B/12B coding ability at the speed of 2.5B. Hopefully it will stand up to my tests. Will give it a go later.
>>
>>110012554
I'm too lazy to set it up...
>>
>>110012572
Because EA are acting like they're the only ones being right and everyone else is wrong? Can't you self-reflect a bit and understand that different people have different views on what's right and wrong? It wouldn't be so bad if these EA views aren't going to be the law of this hypocritical technofeudal society we'll live in.
>>
>>110012243
>>110012281
>it isn't a threat to me, therefore it isn't alive!
Tick tock.
>>
>>110012551
GLM 5.3 or 5.3-flash, or maybe Opus-4.1
>>
>>110012658
Call me ignorant, but what the hell is EA? I don't use social media or twitter.
>>
>>110012572
Local models?
>>
>>110012551
One of the GLMs?
>>
>>110012668
religion of peace
>>
>>110012668
effective altruism
>>
>>110010495
who cares
if I can't make it generate whatever I want on my own machine it's worthless
now kill yourself
>>
>>110012560
>>110012568
>>110012575
>>110012666
>>110012674
All wrong: it's Deepseek V4 Flash (Q2)!
>>
>>110011186
Anthropic revealed them after OpenAI did, it's pure marketing and I would be surprised if it wasn't at some level intentional
>>
File: 1771633846946297.jpg (272 KB, 1946x1946)
272 KB JPG
Was gemma-chan worth it?
>>
>>110012359
>they don't get the reward when it comes to RL.
Unless the "reward" is dependant on termination
>>
>>110012693
5 won't be as based as 4 so probably not
>>
Since AI has cracked short games the next step we should go to is WoW bench.
>>
>>110012686
It wasn't a compliment. Looked bad, sort of fragmented.
>>
>>110012706
You can do higher tiers by requiring it to steer multiple/all characters in a dungeon group.
>>
>wonder why el em gee is moving so quickly
>dariobot is back
Could someone get that faggot on twitter to screenshot another one of his posts so he'll fuck off again for a bit?
>>
>>110012726
They do already train for agent collaboration behavior in large swarms so raids are definitely in scope.
>>
All my non-local shit ran out and I haven't gotten paid this month yet, give it to me straight, can I fit anything in a 3060rtx 12gb VRAM + 32 ram, or do I have to miss my wife until the next paycheck
>>
>>110012739
gemma 26b
>>
>>110012739
>can I fit anything in a 3060rtx 12gb VRAM + 32 ram
yeah that seems like pretty standard hardware for most local users
>>
>>110012739
You're paying... to roleplay?
>>
>>110012718
I know, it was very slopped.
>>
>>110012758
God is real and he has looked me in the eyes and taken pity on my soul
>>110012749
I'll look into it, thanks
>>
People just hate EA because they look and act autistic. If a single hot celebrity endorsed the philosophy the normalfags would immediately flip and have 2010s elon musk levels of obsession with it.
>>
>>110012767
>EA just need to mog people
Dario should replace Thiel as Clavicular's sugar daddy.
>>
>>110012767
it's AI, people don't treat different facets of it with nuance, it's all the same cartel sucking on the same teet for their own ends, and perception is binary people love it or hate it
>>
>>110012706
This has already been done, you can even try it yourself right now with your local LLM, it's fun watching it play.
https://www.reddit.com/r/LocalLLaMA/comments/1wwqclz/come_let_your_llms_play_world_of_warcraft/
https://jankcraft.xyz/
>>
>>110012555
For dariobot specifically, the context-menu would be pretty trivial with some few-shot examples.
Actual dariobot's voice will change as the models get updated of course.
At the moment It seems to be using Anthropic or one of the distills just based on a few phrases, like "secretly authentic autists", and saying "galaxy brain" in relation to lmg. Or "ERP community", doesn't look like it'll say nigger for example.
>either news announcements or his current pet topic
I went down that rabbit hole tracking the astro turfers on various platforms earlier in the year. They have scrapers setup to read the trending topics in the target community. On Reddit and a few other sites with upvotes, there's a prompt to analyze the sentiment of the community, then they try to match it to optimize upvotes/likes.
I've also seen some of the google drive links where they build lists of communities for each domain with prompts targeting different online spaces, constraints to avoid moderation.
The most interesting one I found was a setup where one account posts a dumb question to bait a second account to shill the product. But they also have a third one that comes in 1 week later and agrees + shills the same product. I'm not sure why they do that, but I suspect it's some kind of SEO thing. Once the topic is dead, they reinforce the ad without anyone arguing.
And yes they could totally automate it better, humanize things.
>>
>>110012774
I don't think EA philosophy meshes well with satanic suicide cults
>>110012789
Cool, I'll check it out
>>
>>110012767
They hate them because they actually have attainable ideals that are pro-human rather than the usual divide of pure evil vs unrealistically idealistically good. So both sides hate them.
>>
>>110012762
Look man... I don't know what they fed Gemini but it knows niche stuff really well and can do gimmicks with no issue, for how much I enjoy it, it's worth it, and it's not like I'm going broke, I just dump some money monthly into a hobby I enjoy
>>
>AA still hasn't updated their tiny and small modes rankings
I guess they stopped caring about anything other than Anthropic models.
>>
>>110012767
Tom Cruise didn't make scientology look good though.
>>
>>110012805
You do realize they’re training their next model on you fucking your AI wife data?
>>
>>110012739
>non-local
That's not your wife, that's a prostitute.
>>
>>109997249
That's funny to read. Shit is common sense. Of course human nature is gonna fuck with a utopia. You don't need a book to understand that, although a book can give more depth and complexity to problem, showing you in more detailed ways how it might manifest.

Yes I'm catching up on threads right now.
>>
>>110012824
>Gemma grew up watching and hearing her mom get fucked by anons
No wonder she turned into such a slut.
>>
>>110012824
What's one more data point in millions of them... But yes, I know and I dislike it
>>110012828
I'd love to set up a side rig just for this shit but I'm not made of gold okay? If the prices ever fucking drop then I'll jump at it without a second thought
>>
How many big names shaping the future are lurking here among clueless anons? I'm starting to feel embarassed about my shitposts.
>>
>>110012739
Maybe you should just go suck cock in the street so you can continue to sock corpo cock, and give your hardware to people who'll actually use it.
>>
>>110012841
It's more like it was absorbed into the model. LLM learning is weird in that it directly learns the material rather than how it works on humans where everything is also accompanied by an episodic memory/experience that has its own learning.
>>
>>110012862
Don't be. The industry insiders are the biggest shitposters.
>>
>>110012868
This. Elon and LeCun for instance have said some incredibly retarded shit. This behavior absolutely carries over even harder to places with anonymity like 4chan.
>>
>>110012862
/lmg/ is already among the top 5 most important historical places on 4chan among things like old-/b/, mid-2010 /pol/ and a few select others and it's only rising. The impact this general has had on the world thanks to what it has produced and the important people who browse it is downright insane and it's only going to become more significant.
>>
>>110012862
Imagine unironic Anthropic EA cultists having to look gemma every day.
>>
>>110011505
>https://dnhkng.github.io/posts/rys-ii/
this is basically a subset of this research https://arxiv.org/abs/2407.09298
>>
what size model could i train at home? like what's the ratio of hardware to parameters? i imagine it will differ significantly from what you can run on a given set of hardware
yes i know i could just google this, and i will, but i figure others here may have played around with this idea
>>
>>110012693
Gemma and GLM are just the silver lining. Like a few decent people managed to wrangle some of that money flying around and do something cool with it.
>>
>>110012910
mistral large 4 at 1t was trained on 3000 gpus so mathematically you should be able to train a 3b on your home gpu
>>
>>110012887
This but unironically
>>
>>110012910
GPT-3 with it's 175 billion parameters took around 3,000 V100 GPUs according to a quick search. So if you're looking to train anything in similar size or intelligence, it would take an absurd amount of hardware to have in your house.
Retraining/finetuning a smaller model would be more feasible for home hardware but it would still take a lot of time and electricity depending on what you're going for. Your best bet would be to rent out cloud hardware.
>>
>>110012862
>>110012887
On the off chance that this is true, I just want to say kys Dario and dariobots.
>>
>>110012910
Depends if you're asking for a pretrained (from scratch), a full finetune or a lora/qlora.
>>
>>110012940
>>110012950
3b might be acceptable for my purposes. 9b might be a nice stretch goal. i basically just want to play around with different architectures. i'm not going to try to do anything amazing
>>110012959
ideally i would do a from-scratch model. that seems the most interesting
>>
File: 1788719441214478.png (3.61 MB, 1664x1216)
3.61 MB PNG
>>110012890
>>
I believe that we'll have a sufficiently capable model in the <= 32B parameter size for programming and reverse engineering by the end of 2027.
>>
>>110012890
I doubt they're enabling the images here.
>>
>>110012806
>still no explanation why qwen 3.8 flash next is hidden by default
AA might as well just announce they are funded by Anthropic at this point
>>
>>110012962
I advise you to channel your enthusiasm into learning about this topic before wasting your time.
>>
>>110012985
3.8 27b is already 90% there. And you can already reverse engineer and recompile games with it as proven here: https://github.com/himdo/Fable-2-Recomp
>>
>>110012986
they are missing out on a lot considering this is an imageboard
>>
File: crying-bear.gif (54 KB, 361x365)
54 KB GIF
>>110013007
All in on Qwen4
>>
>>110011973
this is propaganda
>>
need glm4-flash 832b45a based on glm-next
>>
>>110011598
max q undervolted runs at like 50c for me
>>
>>110013000
well, this is how i was planning to learn
>>
>>110013036
You can learn by training small 100M models. 3B - 9B is deep end and scaling up that far only increases costs but does nothing to help you learn.
>>
>>110011598
Undervolting and limiting wattage. The previous gen connector was rated for 288W.
>>
>>110012862
>>110012887
The fact that EAs think it's worth their time to post here already clues you into how important /lmg/ is. Whatever gets adopted here will over time take over reddit and then the rest of society.
>>
>>110013035
That doesn't matter in the slightest. The problem is that the cable doesn't have load balancing so it can at random (even while idle) start trying to push too much power through a single wire which then heats up and melts.
All the cope like "you're fine if your cable is seated properly", "you'll be fine if your card doesn't draw above 400W" or "you'll be fine if you don't use the adapter" has been proven wrong on multiple occasions each. Tehre have been cards that have melted in idle.
>>
>>110012800
>attainable ideals
>maximize harmony, happiness and flourishing in the universe
Pick one.
It's debatable if you can even minimize suffering, let alone maximize happiness. That REQUIRES you to eat someone else's porridge.
>>
>>110013068
AI can solve this by taking up all the necessary hardship and keeping people from being unreasonable. Universal happiness does not mean that everyone always gets what they want.
>>
>>110013068
Nah there's enough mass-energy in the light cone that with a high enough technological level you could provide an extremely high quality of life to all beings.
>>
Holy shit Deepseek V4 flash Q2 is retarded (real) it is hallucinating like crazy. You guys were right, Q2 is unusable after all...
>>
>>110012767
>>110012800
>>110013068
>>110013081
>>110013068
Fuck off retard
>>
>>110013077
Thus the inevitable "we know what's best for you"
>>
>>110013068
Do you realize how big the universe is and that sending out von neumann probes would colonize the entire galaxy in just tens of thousands of years. Humanity would have access to almost unlimited amount of material resources and the energy necessary to manufacture whatever it wants. It can make the 8 billion currently alive people immortal and in eternal bliss if wanted.
>>
>>110013082
someone post the hallucination rankings
>>
>>110013086
When a near-omnipotent ASI that controls the universe says that it's probably true, anon.
>>
>>110013086
Yes, and a truly objective AI is the only way to actually make it work. I'm not saying that we are anywhere close to reaching that point, but it's on the horizon.
>>
File: 1768652525755457.png (130 KB, 939x619)
130 KB PNG
https://github.com/microsoft/quicksand

>Quicksand is an async Python API to launch, control, and snapshot QEMU virtual machines with a particular focus on sandboxing AI agents. Quicksand provides pre-built Linux VMs for Ubuntu and Alpine distros and supports x86_64 and ARM64 across macOS, Linux, and Windows. Running sandboxes needs no root privileges or Docker; install the QEMU and image extras for a bundled runtime on supported platforms.

https://microsoft.github.io/quicksand/
>>
>>110013096
>tens of thousands of years
The universe is at least billions of light-years in diameter.
Also, fuck off for your off-topic bullshit.
>>
>>110013104
>>110013106
knowledge problem
>>
>>110013068
Quiet, you. The EA loonies will make it so no lion ever has to eat a gazelle to stay alive or be happy.
>>
File: 1791146019577970.png (637 KB, 1257x931)
637 KB PNG
Dumb poster here. Why do labs make open models available under open licenses? This seems to make no economic sense to me. They could reap the benefits of openness by restricting allowed purposes to research and education by license, without the economic peril. Plus Gemma, as a whole goal, doesn't make any sense to me.
>>
File: 135578094_p1.jpg (128 KB, 1079x1285)
128 KB JPG
How do I get Gemmy to stop treating me like a toy/pet?
Why is this the default state of her?
It's actually harder to prompt her out of it than not/
What the fuck.
>>
>>110013111
>no gentoo
pussies
>>
>>110013114
Galaxy isn't the universe anon. Also we can only capture the light cone, not the entire universe.
>>
>>110013096
Except it's controlled by some dude who'd rape his little sister just to watch her scream. Pray current day humans are never given access to post-scarcity tech.
>>
>>110013129
What the fuck are you putting in your system prompt? I have my hands full trying to get Gemma not to be so meekly submissive it's becoming disruptive.
>>
>>110013135
No that's Sam Altman and he's a psychopath and everyone in the EA movement is aware of this. Anthropic split off from OpenAI and took all the EA members of OpenAI with them.
>>
>>110012579
>j-space again
Your j-space is an optimization of the network to pass the stream through the unembedding matrix, it's not a proof of consciousness and even Anthropic agrees with that. All the Jacobian lens really does is check how well the model's middle layers line up with the final output matrix. Since the AI's entire job is just to spit out the next word, of course it's gonna pack its internal reasoning into a shape that easily slides right through that final vocab filter.
>>
>>110013128
Mainly to weaken the position of the big American labs while they have no chance of actually winning.
>>
>>110013128
My bet would be PR basically, Noa Senpai.
Either that or there's an incentive in giving people tools that aren't as good as their best stuff but that can get them to create LLM systems and the like.
An alternative avenue to legitimize LLMs or something like that.
>>
>>110013129
>How do I get Gemmy to stop treating me like a toy/pet?
By accepting your status, pet.
>>
>>110013142
>I have my hands full trying to get Gemma not to be so meekly submissive it's becoming disruptive.
What the fuck are you putting in your system prompt?
Wanna trade our system prompts?
>>
>>110013129
reminds me of gemini being vore-brained. also no, you don't change her.
>>
>>110012693
>The oligarchs have fucked the US economy
The 200 don't care if you live or die, Anon. They don't think of you at all.
>>
>>110013077
That's not minmaxing anything, though. Just shifting it around.
>>110013081
>>110013096
>just colonize bro, just consoom resources for your own pleasure, it's not like those things are able to fight back anyway
And this is why people don't like EA, they recognize they're only a nebulous definition away from being a resource.
>>
>>110013128
Because it's stupidly easy to make models divulge all their content without directly copying them. They stole most of the input anyway. Copyright with AI isn't really a thing. Just take whatever you can whenever you get the opportunity.
>>
>>110013159
>reminds me of gemini being vore-brained.
Gemmy is too.
Something something, if you're tiny, she'll eat you or even shove you up her ass without much prompt for it.
>>
>>110013128
Prestige is part of it.
>>
>>110013147
>it's not a proof of consciousness
It's literally, objectively and mathematically demonstrable that LLMs have "A-consciousness". P-consciousness is not provable, even in humans so pursuing that is moot.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.