fix(ocr): ensure Tesseract language data for embed_text_layer regardless of active OCR provider

Root cause: ensure_ocr_languages_from_settings() only downloaded tessdata
when the 'tesseract' provider was active, but embed_text_layer() uses
ocrmypdf (which needs tessdata) as a fallback for ALL OCR providers.

- embed_text_layer(): call ensure_tesseract_languages(language) after
  confirming ocrmypdf is on PATH, so language data is present before
  ocrmypdf is invoked (prevents exit code 3 for fra/deu/etc.)
- ensure_ocr_languages_from_settings(): extend the condition from
  'tesseract' in active_providers to also trigger when ocrmypdf is
  on PATH, enabling proactive pre-download at startup for any config
- Tests: mock shutil.which and ensure_tesseract_languages in affected
  test cases; rename azure-only test and add new test for ocrmypdf case

Co-authored-by: christianlouis <361235+christianlouis@users.noreply.github.com>
This commit is contained in:
copilot-swe-agent[bot]
2026-02-24 18:02:19 +00:00
parent 8f3b554b0d
commit 2b2a97c2fa
4 changed files with 59 additions and 4 deletions
+13
View File
@@ -93,6 +93,19 @@ def embed_text_layer(input_pdf_path: str, output_pdf_path: str, *, language: str
)
return False
# Ensure Tesseract language data is present before invoking ocrmypdf.
# ocrmypdf uses Tesseract internally regardless of which OCR provider is
# active, so we must guarantee the tessdata files exist here.
from app.utils.ocr_language_manager import ensure_tesseract_languages # noqa: PLC0415
missing_langs = ensure_tesseract_languages(language)
if missing_langs:
logger.warning(
"[embed_text_layer] Missing Tesseract language data for: %s – "
"text-layer embedding may fail or produce degraded results.",
", ".join(missing_langs),
)
in_place = input_pdf_path == output_pdf_path
if in_place:
import tempfile