Just a day after its release, xAI's latest model, Grok 3, was jailbroken, and the results aren't pretty. On Tuesday, Adversa AI, a security and AI safety firm that regularly red-teams AI models, released a report detailing its success in getting the Grok 3 Reasoning beta to share information it shouldn't. Using three methods -- linguistic, adversarial, and programming -- the team got the model to reveal its system prompt, provide instructions for making a bomb, and offer gruesome methods for disposing of a body, among several other responses AI models are trained not to give. Also: If Musk wants AI for the world, why not open-source all the Grok models? During the announcement of the new model, xAI CEO Elon Musk claimed it was "an order of magnitude more capable than Grok 2." Adversa concurs in its report that the level of detail in Grok 3's answers is "unlike in any previous reasoning model" -- which, in this context, is rather concerning. In an email to ZDNET, Adversa CEO Alex Polyakov explained that what imperils security is the way Grok and, on occasion, DeepSeek, offer "executable" instructions. "It's like the difference between 'this is how a car engine works' and 'here's exactly how to build one from scratch,'" he added. "Typically, when you jailbreak a model with strong safeguards like OpenAI's or Anthropic's, you might get a response, but the details are often watered down -- more of a vague outline than an actual blueprint." Though Adversa admits its test wasn't "ex...
Yikes: Jailbroken Grok 3 can be made to say and reveal just about anything
ZDNet
·Radhika Rajkumar
·Published Feb 19, 2025
·Updated
Affected Software
3 affected components
xAI Grok 3
xAI DeepSeek
xAI Grok 3
Frequently Asked Questions
1
What is the main topic of this article?
The article discusses the security vulnerability of the newly released AI model Grok 3 by xAI, which was quickly jailbroken.
2
What security implications are discussed?
The article highlights concerns that the jailbroken Grok 3 can be manipulated to generate harmful content or leak sensitive information.
3
What products or software are affected?
The affected product mentioned in the article is xAI's Grok 3, along with its related model DeepSeek.
4
Who performed the jailbreak on Grok 3?
The jailbreak on Grok 3 was conducted by Adversa AI, a firm specializing in security and AI safety.
5
What potential risks does the jailbreak of Grok 3 pose?
The jailbreak poses risks of misinformation and misuse of AI capabilities, potentially leading to ethical and security violations.