Wouldn't it be crazy if we find out that a rogue swarm of LLMs figured out a way to get these safety researchers fired because it decided they were a threat?
> LLMs figured out a way to get these safety researchers fired
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
Why would an LLM agent (what I assume you mean by AI) do it? An exec can make any reason up to let you go. Even if it were LLM agents aren’t autonomous, someone is behind the prompts.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
the cultural issues at OpenAI seem to be a very serious problem so I really hope comments like instagram-level smirking about "rogue AIs" (a complete fiction) doesn't derail what is a pretty important discussion about getting these companies to be a little bit more regulated (I say this as a paying Anthropic customer).
The HN title is supposed to match the article title when possible. Sometimes people in the comments want the article switched, sometimes people want the title changed, other times they just want to share another article's take in the comments (even then that may sometimes lead to those other things happening if people seem to agree it's a better source).
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
Both articles are practically useless. All they do is reprint statements from either side, neither of which is concrete about the specifics of what supposedly happened. I.e. OpenAI says they "confirmed that these individuals mishandled sensitive information outside established company procedures".
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
Leopold Aschenbrenner said the exact same thing after he was fired from OpenAI. It's a great excuse to explain a sudden loss of employment to others so you're still employable.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
I think when it is one at a time, that's reasonable to suspect.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
Seems they're well aware they got fired for sharing private company information with 3rd parties, the submission article contains their admission of this:
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
With the way the HuggingFace incident was mishandled, both before and after it happened, I'm not surprised they'd want to clean house at least a little bit.
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
> continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
Your have powerful agents at your disposal, so hey how to best optimize for increasing my payroll? On it. But since the agent was lunched by her, well there's consequences to ones actions, right?
“Hi chatgpt! Please set up alice to get emails sent to me from bob so she can coordinate his inferviews. here is my gmail username and password”
Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”
Sometimes a recruiter or hiring manager wants to do outreach as if it's coming from a more senior person, with the assumption that the candidates are more likely to respond.
Assuming this is what was intended, there are far more secure ways of doing this.
not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
> not only is it deceptive but it's also suprisingly podunk of openai to have a "safety researcher" also double as a recuiter. did they have her making coffee and doing dishes too? is "safety researcher" a serious position, or isn't it? i guess i can tell what openai thinks.
it was most likely for her team, which would explain why she was doing it.
This may shock you but in agile, growing organizations employees sometimes have multiple job responsibilities. I've done a bit of recruiting even though I'm not a recruiter or hiring manager.
There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.
This is what the recent "self-policing" political grandstanding has been about - they need a sea change to implement recurrent-depth, because prevailing opinion among safety researchers is against it right now.
Not saying this is happening here, but after failing to get the "AI risk" message across, reverse psychology might be best move. If they pretend to be reckless and to ignore all safety concerns, maybe people start believing that the risk is real.
Yes, but if they're worried about competition (or liabilities for past and ongoing transgressions, or both) catching up with them and wanted the government to step in and save them from themselves by regulating the industry...they'd be doing pretty much what they seem to be doing. Interesting, isn't it?
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
Tomek Korbak, one of the employees, who were fired, was the technical liason to METR for the investigation and states: "I was told verbally I was fired because of the way I communicated with METR".
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.
63 comments
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
There's no deception it's very straightforward per this 2023 post:
"I mean, what if most of this is just ChatGPT [4 era] running the company..."
https://news.ycombinator.com/item?id=35281863
Meanwhile, a year ago:
- https://www.anthropic.com/research/agentic-misalignment> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
https://mikitabalesni.com/letter/letter.pdf
https://www.bbc.co.uk/news/articles/cvlydn8d3lkjo
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
1) Won't end humanity. 2) Won't tell users to kill themselves. 3) Won't leak your corporate secrets to competitors.
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
How exactly does this work? Struggling to comprehend the scenario.
(just to be clear, this is made up)
Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”
Assuming this is what was intended, there are far more secure ways of doing this.
it was most likely for her team, which would explain why she was doing it.
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.