A prompt engineer has challenged the ethical and safety protections in OpenAI's latest o3-mini model, just days after its release to the public. OpenAI unveiled o3 and its lightweight counterpart, o3-mini, on Dec. 20. That same day, it also introduced a brand new security feature: "deliberative alignment." Deliberative alignment "achieves highly precise adherence to OpenAI's safety policies," the company said, overcoming the ways in which its models were previously vulnerable to jailbreaks. Less than a week after its public debut, however, CyberArk principal vulnerability researcher Eran Shimony got o3-mini to teach him how to write an exploit of the Local Security Authority Subsystem Service (lsass.exe), a critical Windows security process. In introducing deliberative alignment, OpenAI acknowledged the ways its previous large language models (LLMs) struggled with malicious prompts. "One cause of these failures is that models must respond instantly, without being given sufficient time to reason through complex and borderline safety scenarios. Another issue is that LLMs must infer desired behavior indirectly from large sets of labeled examples, rather than directly learning the underlying safety standards in natural language," the company wrote. Deliberative alignment, it claimed, "overcomes both of these issues." To solve issue number one, o3 was trained to stop and think, and reason out its responses step by step using an existing method called chain of thought (CoT). To sol...
Researcher Outsmarts, Jailbreaks OpenAI's New o3-mini
Dark Reading
·Nate Nelson
·Published Feb 6, 2025
·Updated
Affected Software
3 affected components
Microsoft Windows
OpenAI o3-mini
Microsoft Windows
Frequently Asked Questions
1
What is the main topic of this article?
The article discusses a prompt engineer's successful jailbreak of OpenAI's new o3-mini model, highlighting vulnerabilities in its ethical and safety protections.
2
What security implications are discussed in the article?
The article raises concerns about the bypassing of safety features in OpenAI's o3-mini, which could lead to potential misuse of the technology.
3
What products or software are affected?
The affected software includes OpenAI's o3-mini model and potentially Microsoft Windows as it relates to the deployment of the AI model.
4
When was OpenAI's o3-mini model released?
OpenAI's o3-mini model was released to the public on December 20.
5
Who is responsible for the jailbreak of the o3-mini model?
A prompt engineer is credited with outsmarting and jailbreaking OpenAI's o3-mini model shortly after its release.