fix: Fehlerzaehlung, Upload-Fehler und ocr.timeout scharf (v0.4.0)
- [ocr].timeout wird als tesseract_timeout (pro Seite) an ocrmypdf durchgereicht; Default 1800 -> 300, 0 = ocrmypdf-Default - Exceptions nach dem OCR zaehlen als Fehler, Datei wird nach error/ gerettet - Fehlgeschlagene Uploads zaehlen als Fehler und loesen Fehler-Mail aus - name_mode wird im Preflight geprueft, nicht erst pro Datei - Fehlende [paths]-Sektion -> ConfigError mit klarer Meldung statt KeyError - Stabilitaets-Timeout zaehlt als Fehler (--once liefert Exit 1) - upload_folder nutzt shutil.copyfile statt read_bytes/write_bytes - OcrConfig.pdfa_level Default "2" -> "" (Ghostscript-Bug, Issue #3) - 35 neue Tests (92 gesamt), pytest.ini - AI_AGENT_BRIEFING.md auf Stand 0.4.0 gebracht Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -83,10 +83,10 @@ Vollständiges Beispiel: [`config.example.toml`](config.example.toml). Wichtigst
|
||||
languages = "deu+eng" # Tesseract-Sprachen
|
||||
jobs = 4 # Threads pro PDF
|
||||
skip_text = true # bereits OCR-haltige Seiten überspringen
|
||||
pdfa_level = "2" # "1", "2", "3" oder "" für reines PDF
|
||||
pdfa_level = "" # "1", "2", "3" oder "" für reines PDF (Default "" wegen Ghostscript-Bug, s.u.)
|
||||
deskew = true
|
||||
max_workers = 2 # parallele PDFs
|
||||
timeout = 1800
|
||||
timeout = 300 # max. Sekunden pro SEITE (Tesseract), 0 = ocrmypdf-Default
|
||||
```
|
||||
|
||||
### `[output]`
|
||||
@@ -239,5 +239,5 @@ MIT — © Sonith UG
|
||||
|
||||
---
|
||||
|
||||
**Version:** 0.3.1
|
||||
**Version:** 0.4.0
|
||||
**Repo:** https://gitea.sonith.de/sonith_ug/pdf-ocr-hotfolder
|
||||
|
||||
Reference in New Issue
Block a user