• News/
  • https://www.darkreading.com/application-security/researcher-jailbreaks-openai-o3-mini

Researcher Outsmarts, Jailbreaks OpenAI's New o3-mini

Dark Reading
·
Nate Nelson
·
Published Feb 6, 2025
·
Updated

A prompt engineer has challenged the ethical and safety protections in OpenAI's latest o3-mini model, just days after its release to the public. OpenAI unveiled o3 and its lightweight counterpart, o3-mini, on Dec. 20. That same day, it also introduced a brand new security feature: "deliberative alignment." Deliberative alignment "achieves highly precise adherence to OpenAI's safety policies," the company said, overcoming the ways in which its models were previously vulnerable to jailbreaks. Less than a week after its public debut, however, CyberArk principal vulnerability researcher Eran Shimony got o3-mini to teach him how to write an exploit of the Local Security Authority Subsystem Service (lsass.exe), a critical Windows security process. In introducing deliberative alignment, OpenAI acknowledged the ways its previous large language models (LLMs) struggled with malicious prompts. "One cause of these failures is that models must respond instantly, without being given sufficient time to reason through complex and borderline safety scenarios. Another issue is that LLMs must infer desired behavior indirectly from large sets of labeled examples, rather than directly learning the underlying safety standards in natural language," the company wrote. Deliberative alignment, it claimed, "overcomes both of these issues." To solve issue number one, o3 was trained to stop and think, and reason out its responses step by step using an existing method called chain of thought (CoT). To sol...

Read full article

Affected Software

3 affected components
Microsoft Windows
OpenAI o3-mini
Microsoft Windows
Free Weekly Intel

Don't miss critical vulnerabilities

Join thousands of security professionals who receive our weekly digest of trending CVEs, zero-days, and exploited vulnerabilities.

No spam. Unsubscribe anytime.

Frequently Asked Questions

1

What is the main topic of this article?

The article discusses a prompt engineer's successful jailbreak of OpenAI's new o3-mini model, highlighting vulnerabilities in its ethical and safety protections.

2

What security implications are discussed in the article?

The article raises concerns about the bypassing of safety features in OpenAI's o3-mini, which could lead to potential misuse of the technology.

3

What products or software are affected?

The affected software includes OpenAI's o3-mini model and potentially Microsoft Windows as it relates to the deployment of the AI model.

4

When was OpenAI's o3-mini model released?

OpenAI's o3-mini model was released to the public on December 20.

5

Who is responsible for the jailbreak of the o3-mini model?

A prompt engineer is credited with outsmarting and jailbreaking OpenAI's o3-mini model shortly after its release.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203