CVE-2025-49847: llama.cpp Vulnerable to Buffer Overflow via Malicious GGUF Model
llama.cpp is an inference of several LLM models in C/C++. Prior to version b5662, an attacker‐supplied GGUF model vocabulary can trigger a buffer overflow in llama.cpp’s vocabulary‐loading code. Specifically, the helper trycopy in llama.cpp/src/vocab.cpp: llamavocab::impl::tokentopiece() casts a very large sizet token length into an int32t, causing the length check (if (length < (int32t)size)) to be bypassed. As a result, memcpy is still called with that oversized size, letting a malicious model overwrite memory beyond the intended buffer. This can lead to arbitrary memory corruption and potential code execution. This issue has been patched in version b5662.
Affected Software
Remediation
Event History
Frequently Asked Questions
What is the severity of CVE-2025-49847?
CVE-2025-49847 has been classified as a high severity vulnerability due to the potential for buffer overflow exploitation.
How do I fix CVE-2025-49847?
To mitigate CVE-2025-49847, update to llama.cpp version b5662 or later as it resolves the buffer overflow issue.
What type of vulnerability is CVE-2025-49847?
CVE-2025-49847 is a buffer overflow vulnerability occurring during the loading of attacker-supplied GGUF model vocabulary.
What software is affected by CVE-2025-49847?
CVE-2025-49847 affects llama.cpp versions prior to b5662.
Can CVE-2025-49847 be exploited remotely?
Yes, CVE-2025-49847 can be exploited remotely if an attacker supplies a malicious GGUF model vocabulary.