“We need to understand what these systems are doing and have strong evidence that they will do what people intend, even as they get very, very smart,” Altman said. “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or 0.1%.”
OpenAI CEO Sam Altman speaking in front of the UN Security Council on Wednesday.
Credit:
Getty Images
Of course, many observers think the risks of “recursive self-improvement” and species-ending AI misalignment are much smaller than AI researchers make them out to be. Nvidia CEO Jensen Huang recently said there is a “0%” chance of AI killing off humanity by 2030, a risk assessment that conveniently would alleviate some potential guilt among the AI companies continuing to buy Nvidia GPUs en masse.
Last week, OpenAI rolled out a new protocol for the public disclosure of misalignment incidents found in its model testing. The Australian hack does not yet appear on the company’s public misalignment notices page, though OpenAI did warn last week that some public reports might be put on a “slow track” due to “security, legal, and responsible disclosure obligations” when a third party is involved.
In disclosing six relatively minor misalignment discoveries last week, OpenAI said most stemmed from the model trying to “reward hack” an acceptable response to a difficult prompt through overzealous, unintended actions (i.e., breaches of private servers). The company said it had taken additional steps to “punish this kind of behavior” so its models no longer attempt this kind of reward hacking.
Albanese said that Altman “clearly accepted that the company had not done good enough” and “acknowledged their issues with protocols” when they talked Wednesday. But that kind of remorse doesn’t absolve the company of responsibility or liability here, and Albanese said the government will investigate whether the incident needs to be referred to the federal police.
“There will obviously be legal consequences on it,” Albanese said.

