fix: Fehlerzaehlung, Upload-Fehler und ocr.timeout scharf (v0.4.0)

- [ocr].timeout wird als tesseract_timeout (pro Seite) an ocrmypdf
  durchgereicht; Default 1800 -> 300, 0 = ocrmypdf-Default
- Exceptions nach dem OCR zaehlen als Fehler, Datei wird nach error/ gerettet
- Fehlgeschlagene Uploads zaehlen als Fehler und loesen Fehler-Mail aus
- name_mode wird im Preflight geprueft, nicht erst pro Datei
- Fehlende [paths]-Sektion -> ConfigError mit klarer Meldung statt KeyError
- Stabilitaets-Timeout zaehlt als Fehler (--once liefert Exit 1)
- upload_folder nutzt shutil.copyfile statt read_bytes/write_bytes
- OcrConfig.pdfa_level Default "2" -> "" (Ghostscript-Bug, Issue #3)
- 35 neue Tests (92 gesamt), pytest.ini
- AI_AGENT_BRIEFING.md auf Stand 0.4.0 gebracht

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-22 21:01:38 +02:00
parent cbdc9d6664
commit 578472872e
17 changed files with 890 additions and 76 deletions
+3 -3
View File
@@ -83,10 +83,10 @@ Vollständiges Beispiel: [`config.example.toml`](config.example.toml). Wichtigst
languages = "deu+eng" # Tesseract-Sprachen
jobs = 4 # Threads pro PDF
skip_text = true # bereits OCR-haltige Seiten überspringen
pdfa_level = "2" # "1", "2", "3" oder "" für reines PDF
pdfa_level = "" # "1", "2", "3" oder "" für reines PDF (Default "" wegen Ghostscript-Bug, s.u.)
deskew = true
max_workers = 2 # parallele PDFs
timeout = 1800
timeout = 300 # max. Sekunden pro SEITE (Tesseract), 0 = ocrmypdf-Default
```
### `[output]`
@@ -239,5 +239,5 @@ MIT — © Sonith UG
---
**Version:** 0.3.1
**Version:** 0.4.0
**Repo:** https://gitea.sonith.de/sonith_ug/pdf-ocr-hotfolder