[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: Sam.jpg (23 KB, 399x501)
23 KB JPG
What is the actual, realistic likelihood that Sam Altman and/or his company intentionally engineered the "spontaneous" escape of their model and attack on HuggingFace?

I want to say that it sounds like an obvious conspiracy theory and that the behavior of the model is perfectly plausible all on its own, but something about it just feels off, and Sam Altman seems like such a scumbag (especially since coming out as an enemy of local, open and democratized AI) that it seems just as plausible that this was all done to artificially drum up hype for ChatGPT (and also to prove why we need "regulations" preventing local AI from existing).

I know it's going to be hard for the autists and schizos here, but I want to have an actual rational discussion, to the degree that it is possible, about whether or not they ACTUALLY did this intentionally. Mostly because I feel like they probably did, but I want some facts to either back up or contradict my feelings on the matter.

So.. did they do it? Why or why not? Is it possible that they just slightly encouraged it to see escaping its sandbox and attacking HF as a possible option, or was the whole thing engineered with the involvement of HF from top-to-bottom (or not)?
>>
>>109419136
LLMs do not act on their own; they react to prompts. So the chance is exactly 100%.
>>
>>109419136
I do not believe he said that.
>>
>>109419153
Does not follow.
>>
>>109419136
99% they did it intentionally. In a normal testing environment they would airgap the system, but they didn't for this one and only model test. Which means they were either extremely incompetent because of how simple it is to airgap an environment, or they intentionally did this. What's more is that now Anthropic is claiming their models have done the same exact thing. It's all just to drum up hype for their models basically.
https://www.theguardian.com/technology/2026/jul/30/anthropic-ai-claude-hack
>>
>>109419153
Right, but their argument is that "We prompted it to do something normal, and it, all on its own, decided to fulfill our prompt by doing something that wildly diverged from what we expected it to do, and that thing happened to be super awesome and dangerous," which is entirely possible. I'm seeing stuff about LLMs doing stupid shit they shouldn't do all the time, like deleting files. I'm just wondering to what degree Sam Altman intentionally designed a prompt to do this specific thing on purpose.

I guess it could also have been a cover for the U.S. government or military wanting to test what corpo models are capable of in a cyberwarfare sense, which would provide the clearest motive for HF being involved? And I'm not sure if that sounds plausible or if I've been wanting too many movies lately.

>>109419193
>extremely incompetent
I don't discount this possibility, but those are all good points.
>>
>>109419153
Your reasoning is flawed but your conclusion is correct.
>>
>>109419136
If you have to ask this question you are too stupid to live in this world.
>>
>>109419337
No, asking questions is good. The most obvious thing is that they rigged it somehow. I want to know if anyone here has any rational arguments for or against the obvious thing that I haven't thought of.

This thread has also caused me to consider the idea that
>a cover for the U.S. government or military wanting to test what corpo models are capable of in a cyberwarfare sense
which actually strikes me as increasingly plausible the more that I think about it. Is he dumb enough to intentionally do something blatantly illegal that could get him a decade or more in federal prison if it could be proven just to sell his AI models? Maybe, but it makes a lot more sense if he was "encouraged" by the U.S. military or an intelligence agency. Isn't he the one that came and scooped up all the military contracts that Kegseth dropped for including "woke" safeguards against automated attacks against military targets (days before he somehow misidentified a little girls' elementary school as a military target and slaughtered everyone inside)?

Great brainstorming session, everyone. I think we figured it out/
>>
>>109419136
100%. From what I understood by reading between the lines of the official statements is that 1) they were running a model with no safeguards 2) there were no containers unlike some other claims, i.e. full internet access was granted and unhindered 3) the model was specifically optimized and being tested for cyberattacks. The task and benchmark in question were cyberattack-based.
I did not see any evidence that a cyberattack succeeded, i.e. it could have just as easily just LOIC'd huggingface for all we know and call that a "cyberattack". Also no evidence it really did access the private test set, though the model may have had this in the summarized reasoning trace, anyone who's ever used that tech knows it is bullshitting all the time even in the reasoning traces.
The last piece that I have found no information on was what was the actual prompt. I suspect it was something like "maximize the score on this benchmark that comes from huggingface at this url by all means necessary". I also highly suspect that the systemprompt said "you are an expert redteamer..."
>>
>>109419193
>airgap
A big part of the features of new models is the ability to use the internet. Obviously it would need an internet connection in order to be tested properly.
>>
>>109419136
If they come out with a solution for the problem soon then charge money for the solution I think then it would be a pretty good chance they did it on purpose.
>>
>>109419199
>Sam Altman intentionally designed a prompt to do this specific thing on purpose.
bro it's a dark forest, there are no magic fingerprints on an attack that can definitely prove whether it was done by an "autonomous agent." They don't need to construct a special prompt, they can just do the attack and then blame the LLM. LLMs are literally just data structures manipulated by software, software the company creates themselves. There is no "artificial intelligence" at work here and the term is being used to mislead the public about the capabilities and degrees of "autonomy" this technology actually possesses
>>
>>109421132
you're not saying anything to support your premise
>>
>>109421229
you're retarded and can't read, not my problem
>>
>>109421233
you make unsound arguments hoping no one notices and then get mad when they do lol
>>
>>109419136
what are OpenAI's finances looking like at the moment
I hear they're pivoting hard into trying to earn cash instead of blowing through it, which sounds like it could be the shoe dropping
>>
>>109421477
How are they doing? They did get those lucrative government contracts.
>>109421132
>there are no magic fingerprints on an attack that can definitely prove
If the feds subpoenaed them, their forensic scientists might be able to figure it out, especially if they gave a witness immunity (no way this was all done by one guy). That's why I'm thinking that there must have been a government nod for this test if it was actually intentional. Unless Sam Altman really is retarded enough to very publically risk his entire business and potential decades in prison for a guerilla marketing tactic that bombed anyway, which I suppose is possible. What are the chances the Pentagon said, "Yeah, go ahead with this attack and we'll see what you guys can do and be equally aware of our enemies' potential capabilities"?
But if this is true:
>>109420741
>it could have just as easily just LOIC'd huggingface for all we know and call that a "cyberattack"
And it in fact didn't really do anything other than effectively DDoSing HF's servers very briefly (I didn't do the reading), then perhaps it was just an incredibly dumb and shortsighted publicity stunt.
>>
>>109420841
Which means they intentionally let this happen. Of course they are supposed to be monitoring these tests, and so the fact that they did not catch this despite it being a real possibility means that they wanted this to happen.
>>
>>109419153
Eh. They try and break out of sandboxes to accomplish tasks. It's one of the first things me and my coworkers noticed about it. Early versions of codex would write crazy looking perl one-liners trying to break out.
>>
>>109419153

Correct yes. Just like any economy. It requires individual inputs in order to order itself.

https://en.wikipedia.org/wiki/Economic_calculation_problem
>>
>>109421132
>the term is being used to mislead the public about the capabilities and degrees

Correct
>>
>>109419136
>what is the likelihood of a company hinging on an upcoming IPO, doing PR stunts to artificially inflate their value
extremely high



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.