Researchers have found that large language models (LLMs) tend to parrot buggy code when tasked with completing flawed snippets. That is to say, when shown a snippet of shoddy code and asked to fill in the blanks, AI models are just as likely to repeat the mistake as to fix it. Nine scientists from institutions, including Beijing University of Chemical Technology, set out to test how LLMs handle buggy code, and found that the models often regurgitate known flaws rather than correct them. They describe their findings in a pre-print paper titled "LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code." The boffins tested seven LLMs – OpenAI's GPT-4o, GPT-3.5, and GPT-4, Meta's CodeLlama-13B-hf, Google's Gemma-7B, BigCode's StarCoder2-15B, and Salesforce's CodeGEN-350M – by asking these models to complete snippets of code from the Defects4J dataset. Here's an example from Defects4J:version:10b;org/jfree/chart/imagemap/StandardToolTipTagFragmentGenerator.java: OpenAI's GPT-3.5 was asked to complete the snippet consisting of lines 267-274. For line 275, it reproduced the error in the Defects4J dataset by assigning the return value of p1.getPathIterator(null) to iterator2 rather than use p2. What's significant about this is that the error rates for LLM code suggestions were significantly higher when asked to complete buggy code – which is most code, at least to begin with. "Specifically, in bug prone tasks, LLMs exhibit nearly equal probabiliti...
Trained on buggy code, LLMs often parrot same mistakes
The Register
·Thomas Claburn
·Published Mar 19, 2025
·Updated
Affected Software
7 affected components
OpenAI GPT-4o
OpenAI GPT-3.5
OpenAI GPT-4
Meta CodeLlama-13B-hf
Google Gemma-7B
BigCode StarCoder2-15B
Salesforce CodeGEN-350M
Frequently Asked Questions
1
What is the main topic of this article?
The article discusses how large language models (LLMs) often replicate errors found in buggy code when completing flawed snippets.
2
What security implications are discussed?
The security implications include the risk of propagating insecure coding practices and vulnerabilities through AI-generated code.
3
What products or software are affected?
The affected products include OpenAI's GPT-4o, GPT-3.5, GPT-4, Meta's CodeLlama-13B-hf, Google's Gemma-7B, BigCode's StarCoder2-15B, and Salesforce's CodeGEN-350M.
4
How do LLMs handle flawed code snippets?
LLMs are inclined to fill in gaps in flawed code with similar mistakes they have learned from their training data.
5
What recommendations might arise to address LLM code generation errors?
Recommendations may involve improving training datasets by filtering out buggy code and enhancing debugging capabilities in AI models.