CVE-2026-88047: Tesseract: ReadNormProtos stack buffer overflow
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, Classify::ReadNormProtos in src/classify/normmatch.cpp parses the NORMPROTO component of a .traineddata file and uses std::istream::operator>>(char) to extract a whitespace-delimited token into a fixed 61-byte stack buffer without setting a stream width. The 100-byte line buffer can carry a token of up to 99 characters, so a token longer than 60 characters writes up to 39 attacker-controlled bytes past the buffer during TessBaseAPI::Init of the legacy engine, causing stack corruption, denial of service, and potentially control-flow hijacking on affected standard-library implementations. Builds using Apple's libc++ C++20 bounded array overload are incidentally protected, while typical libstdc++ builds remain affected. No fixed release is available as of this review.
Affected Software
Remediation
Recommended actions to resolve this vulnerability, in priority order.
- Configuration
Ensure the NORMPROTO component tokens in the .traineddata files do not exceed 60 characters to prevent overflowing the 61-byte stack buffer in Classify::ReadNormProtos (src/classify/normmatch.cpp) when extracting into a fixed char[61] buffer during TessBaseAPI::Init.
Tesseract (legacy engine) NORMPROTO token length = ≤ 60 characters - Compensating control
Avoid using typical libstdc++ C++20 builds for Tesseract where the std::istream::operator>>(char*) extraction occurs without a stream width; rely on Apple libc++ C++20 bounded array overload builds as stated (“Builds using Apple's libc++ C++20 bounded array overload are incidentally protected”).
- Operational
Assess and restart/restore service for systems that may have been running affected Tesseract versions (5.5.3 and earlier) with untrusted .traineddata, due to potential stack corruption (DoS and potential control-flow hijacking).
Event History
Frequently Asked Questions
What must an attacker control to trigger this issue?
An attacker needs to supply or cause the application to load a malicious .traineddata file whose NORMPROTO component contains a whitespace-delimited token longer than 60 characters. The overflow occurs during TessBaseAPI::Init when using the legacy engine.
Are all Tesseract builds affected?
No. Versions 5.5.3 and earlier are affected in typical libstdc++ builds, while builds using Apple's libc++ C++20 bounded array overload are incidentally protected. The issue is tied to parsing legacy-engine traineddata files.
What can be done if no fixed release is available?
Do not load untrusted .traineddata files, and restrict the sources and write access for traineddata assets. Where feasible, use an environment built with Apple's libc++ C++20 bounded array overload, which is described as incidentally protected.
What impact should responders expect from exploitation?
A crafted token can overwrite up to 39 attacker-controlled bytes past the 61-byte stack buffer. This can cause stack corruption and denial of service, with potential control-flow hijacking on affected standard-library implementations.