Compare commits
8 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 2062476252 | |||
| 8da0b7da1c | |||
| 578472872e | |||
| cbdc9d6664 | |||
| a23a3968ef | |||
| 9cdc9ae443 | |||
| 6f7cadfc63 | |||
| 985a33d3f9 |
@@ -4,6 +4,7 @@ __pycache__/
|
||||
venv/
|
||||
env/
|
||||
.venv/
|
||||
.pytest_cache/
|
||||
*.egg-info/
|
||||
build/
|
||||
dist/
|
||||
|
||||
+176
-34
@@ -1,8 +1,8 @@
|
||||
# AI Agent Briefing — PDF OCR Hotfolder
|
||||
|
||||
**Zuletzt aktualisiert:** 2026-04-08
|
||||
**Version:** 0.1.0
|
||||
**Status:** Initiale Implementation, nicht produktiv getestet
|
||||
**Zuletzt aktualisiert:** 2026-09-22
|
||||
**Version:** 0.5.0
|
||||
**Status:** Multi-Instanz-Betrieb, Preflight-Checks und Fehlerzählung vorhanden, Test-Suite grün (95 pytest-Tests). Ein Produktiv-Einsatz ist im Repo (README/CHANGELOG) nicht dokumentiert — die bisherigen Fixes stammen aus Issues #1–#6, nicht aus einem belegten Dauerbetrieb.
|
||||
|
||||
## 🎯 Projektziel
|
||||
|
||||
@@ -13,16 +13,28 @@ Eingehende gescannte PDFs werden automatisch durch OCR (ocrmypdf + Tesseract) in
|
||||
```
|
||||
pdf-ocr-hotfolder/
|
||||
├── pdf_ocr_hotfolder/
|
||||
│ ├── __init__.py # Versionsstring
|
||||
│ ├── __main__.py # CLI-Entrypoint (argparse, --once, --config)
|
||||
│ ├── config.py # TOML-Loader, Dataclasses
|
||||
│ ├── service.py # Hauptservice (watchdog + ThreadPool)
|
||||
│ ├── processor.py # ocrmypdf + veraPDF
|
||||
│ └── uploaders.py # folder, nextcloud (WebDAV), sftp, email
|
||||
│ ├── __init__.py # Versionsstring (__version__)
|
||||
│ ├── __main__.py # CLI (argparse: --config, --once, --version); Exit 0/1/2
|
||||
│ ├── config.py # TOML-Loader, Dataclasses, ConfigError
|
||||
│ ├── service.py # HotfolderService (watchdog + ThreadPool), Preflight, Zähler
|
||||
│ ├── processor.py # ocrmypdf-Call, veraPDF, Ausgabename, Original-Entsorgung
|
||||
│ └── uploaders.py # folder, nextcloud (WebDAV), sftp, E-Mail-Notify
|
||||
├── tests/ # pytest-Suite (95 Tests, ocrmypdf wird gemockt)
|
||||
│ ├── conftest.py # Fixtures tmp_config / dummy_pdf
|
||||
│ ├── test_config_errors.py
|
||||
│ ├── test_error_counting.py
|
||||
│ ├── test_ghostscript_version.py
|
||||
│ ├── test_ocr_timeout.py
|
||||
│ ├── test_once_exit_code.py
|
||||
│ ├── test_output_naming.py
|
||||
│ ├── test_preflight.py
|
||||
│ └── test_upload_folder.py
|
||||
├── systemd/
|
||||
│ └── pdf-ocr-hotfolder.service # Template (Platzhalter __SERVICE_USER__/__SERVICE_GROUP__)
|
||||
│ ├── pdf-ocr-hotfolder@.service # Template-Unit (Instanz = %i)
|
||||
│ └── lxc-compat.conf # Drop-in-Vorlage: Hardening für LXC abschalten
|
||||
├── pytest.ini # testpaths = tests
|
||||
├── config.example.toml
|
||||
├── install.sh # Interaktiver Installer
|
||||
├── install.sh # Interaktiver Installer + Instanz-Manager
|
||||
├── update.sh # Update aus Repo
|
||||
├── requirements.txt
|
||||
├── VERSION
|
||||
@@ -35,46 +47,159 @@ pdf-ocr-hotfolder/
|
||||
| Komponente | Technologie |
|
||||
|------------|-------------|
|
||||
| Sprache | Python 3.11+ (für `tomllib` aus stdlib) |
|
||||
| OCR | `ocrmypdf` (als Library, nicht via Subprozess) |
|
||||
| OCR | `ocrmypdf` (als Library, nicht via Subprozess; Import ist lazy) |
|
||||
| Engine | Tesseract |
|
||||
| Watcher | `watchdog` |
|
||||
| HTTP | `requests` (Nextcloud WebDAV) |
|
||||
| SFTP | `paramiko` |
|
||||
| Email | `smtplib` (stdlib) |
|
||||
| Service | systemd |
|
||||
| Tests | `pytest` |
|
||||
| Service | systemd (Template-Unit) |
|
||||
|
||||
## 🖥️ Installations-Layout
|
||||
## 🖥️ Installations-Layout (Multi-Instanz)
|
||||
|
||||
| Pfad | Inhalt |
|
||||
|------|--------|
|
||||
| `/opt/pdf-ocr-hotfolder/` | Code + venv (`venv/bin/python`) |
|
||||
| `/etc/pdf-ocr-hotfolder/config.toml` | Konfiguration (mode 640, root:<service-group>) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/{incoming,working,outgoing,error}/` | Datenverzeichnisse |
|
||||
| `/var/log/pdf-ocr-hotfolder/` | Logs (zusätzlich zu journald) |
|
||||
| `/etc/systemd/system/pdf-ocr-hotfolder.service` | systemd-Unit |
|
||||
| `/opt/pdf-ocr-hotfolder/` | Code + venv (für alle Instanzen gemeinsam) |
|
||||
| `/opt/pdf-ocr-hotfolder/.repo_path` | Pfad zum Repo, aus dem installiert wurde (nutzt `update.sh`) |
|
||||
| `/etc/pdf-ocr-hotfolder/<instanz>.toml` | Config pro Instanz (mode 640, root:<service-group>) |
|
||||
| `/etc/systemd/system/pdf-ocr-hotfolder@.service` | Template-Unit |
|
||||
| `/etc/systemd/system/pdf-ocr-hotfolder@.service.d/lxc-compat.conf` | Drop-in für Container (optional) |
|
||||
| `/etc/systemd/system/pdf-ocr-hotfolder@<instanz>.service.d/user.conf` | Drop-in für abweichenden User (optional) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/<instanz>/{incoming,working,outgoing,error}/` | Daten pro Instanz |
|
||||
| `/var/backups/pdf-ocr-hotfolder/` | Update-Backups |
|
||||
|
||||
Ein eigenes Logverzeichnis gibt es **nicht** (seit 0.4.1 auch nicht mehr vom
|
||||
Installer angelegt): `_setup_logging()` nutzt `logging.basicConfig()` ohne
|
||||
FileHandler, alles geht nach stdout → journald.
|
||||
|
||||
```bash
|
||||
journalctl -u pdf-ocr-hotfolder@<instanz> -f # eine Instanz mitlesen
|
||||
journalctl -u 'pdf-ocr-hotfolder@*' --since today # alle Instanzen, heute
|
||||
```
|
||||
|
||||
## 👤 Service-User
|
||||
|
||||
Der Installer fragt interaktiv:
|
||||
1. Username (default `pdfocr`)
|
||||
2. Falls User existiert (lokal oder AD via SSSD/Winbind): wird übernommen, primäre Gruppe automatisch erkannt
|
||||
3. Falls nicht: Frage nach lokaler Anlage als System-User
|
||||
- Basis-Install legt Default-User `pdfocr` an (als System-User, falls nicht schon vorhanden)
|
||||
- Beim Anlegen einer Instanz fragt der Installer nach dem Service-User (default `pdfocr`)
|
||||
- Wird ein **abweichender** User gewählt, wird ein systemd-Drop-in erstellt (`pdf-ocr-hotfolder@<instanz>.service.d/user.conf`) mit `User=/Group=` Override
|
||||
- Existierende User (lokal oder AD via SSSD/Winbind) werden übernommen, primäre Gruppe via `id -gn` ermittelt
|
||||
- Bei AD-Usern mit lokaler UID werden Datei-Berechtigungen über die UID gesetzt — transparent
|
||||
|
||||
**Wichtig:** Bei AD-Usern mit lokaler UID werden Datei-Berechtigungen über die UID gesetzt — funktioniert transparent.
|
||||
## 🗂️ Instanz-Management
|
||||
|
||||
`install.sh` ist gleichzeitig **Installer und Instanz-Manager**:
|
||||
|
||||
- Erster Lauf: Basis-Install + erste Instanz anlegen (Pflicht)
|
||||
- Folgender Lauf: Basis-Install wird übersprungen (erkannt an `venv` + Template-Unit), bestehende Instanzen werden gelistet, weitere Instanzen können ergänzt werden
|
||||
- Eingaben pro Instanz (seit 0.5.0 fünf statt drei):
|
||||
1. Name (`[a-z0-9][a-z0-9-]*`)
|
||||
2. Basis-Pfad (default `/var/lib/pdf-ocr-hotfolder/<name>`)
|
||||
3. Service-User (default `pdfocr`)
|
||||
4. **OCR-Sprachen** (default `deu+eng`) — Format `^[a-z]{3}(_[A-Za-z]+)?(\+…)*$`,
|
||||
bei Unsinn wird erneut gefragt. Jeder Code wird gegen `tesseract --list-langs`
|
||||
geprüft; fehlt einer, bietet der Installer `tesseract-ocr-<code>` an
|
||||
(Unterstrich → Bindestrich, `chi_sim` → `tesseract-ocr-chi-sim`). Ablehnung
|
||||
oder fehlgeschlagene Installation → Warnung, dass OCR mit dieser Sprache
|
||||
**pro Datei** scheitert, und die Sprach-Abfrage beginnt von vorn (kein
|
||||
harter Abbruch). Ist `tesseract` nicht aufrufbar, wird die Prüfung
|
||||
übersprungen und die Eingabe unverändert übernommen.
|
||||
5. **Original nach erfolgreichem OCR archivieren?** (default **nein** →
|
||||
`original_on_success = "delete"`). Bei ja wird der Archiv-Pfad abgefragt
|
||||
(Vorschlag `$BASE/archive`, absoluter Pfad Pflicht), angelegt und auf
|
||||
`$SVC_USER:$SVC_GROUP` gechownt — innerhalb von `$BASE` erledigt das
|
||||
bestehende `chown -R` das schon, nur ein Archiv **außerhalb** bekommt ein
|
||||
eigenes `chown -R`.
|
||||
- **Sprachen sind bewusst instanz-lokal**, nicht global: ein Hotfolder
|
||||
`buchhaltung` läuft mit `deu`, ein Hotfolder `export` mit `deu+eng+fra`.
|
||||
`LANGS`/`ORIG_MODE`/`ARCHIVE_DIR` sind `local` in `create_instance()` — jeder
|
||||
Durchlauf fragt neu, `deu+eng` ist nur der vorgeschlagene Default. Die Liste
|
||||
gehört eng gehalten: jede zusätzliche Sprache kostet Laufzeit **und**
|
||||
Erkennungsqualität.
|
||||
- Basis-Install prüft zusätzlich die Ghostscript-Version und bietet auf Debian 12 bookworm-backports an; erkennt Container (`systemd-detect-virt --container`) und bietet das LXC-Drop-in an
|
||||
- `<instanz>.toml` wird aus `config.example.toml` per `sed` generiert. Substituiert
|
||||
werden die vier `[paths]`-Zeilen **sowie** (seit 0.5.0) `[ocr].languages`,
|
||||
`[output].original_on_success` und `[output].archive_dir`. Die Ausdrücke sind
|
||||
am Zeilenanfang verankert (`^key[[:space:]]*=`), damit die deutschen
|
||||
Kommentarzeilen über den Keys nicht getroffen werden (im Beispiel steht z.B.
|
||||
`"archive" : Original wird in archive_dir verschoben` als Kommentar);
|
||||
Pfad-Variablen laufen vorher durch `sed_escape_repl()` (maskiert `\`, `&`, `|`).
|
||||
Nach dem sed-Lauf liest `config_value()` die drei Keys zurück und vergleicht
|
||||
sie mit der Eingabe; erst wenn das passt, nennt die Zusammenfassung Sprachen
|
||||
und Archiv-Verzeichnis.
|
||||
- Instanz wird sofort `enable --now` gestartet
|
||||
|
||||
Manuelles Löschen einer Instanz:
|
||||
```bash
|
||||
systemctl disable --now pdf-ocr-hotfolder@<name>
|
||||
rm /etc/pdf-ocr-hotfolder/<name>.toml
|
||||
rm -rf /etc/systemd/system/pdf-ocr-hotfolder@<name>.service.d
|
||||
systemctl daemon-reload
|
||||
# Datenverzeichnis /var/lib/pdf-ocr-hotfolder/<name> manuell aufräumen
|
||||
```
|
||||
|
||||
## 🔄 Update-Verhalten
|
||||
|
||||
`update.sh`:
|
||||
1. Findet das Repo (eigenes Verzeichnis oder `/opt/pdf-ocr-hotfolder/.repo_path`)
|
||||
2. Ermittelt alle **aktiven** `pdf-ocr-hotfolder@*.service` Units und stoppt sie
|
||||
3. Backup nach `/var/backups/pdf-ocr-hotfolder/` (tar.gz, ohne venv/`__pycache__`)
|
||||
4. Kopiert Code + requirements + VERSION + config.example aus dem Repo
|
||||
5. `pip install --upgrade` im venv
|
||||
6. Aktualisiert Template-Unit + `daemon-reload`
|
||||
7. Setzt den Code-Eigentümer auf den User, dem `venv` gehört (default `pdfocr`)
|
||||
8. Startet alle zuvor aktiven Instanzen wieder, Exit 1 wenn eine nicht mehr hochkommt
|
||||
|
||||
Config-Dateien werden **nie** überschrieben. Das Repo muss erhalten bleiben — `update.sh` kopiert daraus.
|
||||
|
||||
## ⚙️ Konfiguration (Überblick)
|
||||
|
||||
Vollständiges Beispiel mit Kommentaren: `config.example.toml`. Sektionen:
|
||||
|
||||
| Sektion | Zweck |
|
||||
|---------|-------|
|
||||
| `[paths]` | `incoming`, `outgoing`, `working`, `error` — **Pflicht**, fehlt einer → `ConfigError` + Exit 2 |
|
||||
| `[ocr]` | `languages`, `jobs`, `skip_text`, `oversample`, `pdfa_level`, `deskew`, `clean`, `max_workers`, `timeout` (Sekunden **pro Seite**) |
|
||||
| `[output]` | `name_mode` (`prefix`/`suffix`/`none`), `name_tag`, `original_on_success` (`delete`/`archive`), `archive_dir` |
|
||||
| `[verapdf]` | `enabled`, `binary`, `flavour` — optionale PDF/A-Validierung per CLI |
|
||||
| `[upload.folder]` | `enabled`, `target` (leer = `[paths].outgoing`, dann No-op) |
|
||||
| `[upload.nextcloud]` | `enabled`, `url`, `username`, `password`, `remote_path`, `verify_ssl` |
|
||||
| `[upload.sftp]` | `enabled`, `host`, `port`, `username`, `key_file`, `password`, `remote_path` |
|
||||
| `[notify.email]` | `enabled`, SMTP-Daten, `from_addr`, `to_addrs`, `on` = `always`/`errors`/`never` |
|
||||
| `[logging]` | `level` = DEBUG/INFO/WARNING/ERROR |
|
||||
|
||||
Unbekannte Keys in einer Sektion werden beim Laden **still verworfen** (`config.py` filtert gegen die Dataclass-Annotationen) — Tippfehler in Key-Namen fallen also nicht auf.
|
||||
|
||||
## 🔄 Verarbeitungs-Flow
|
||||
|
||||
1. `watchdog` triggert auf Datei-Event in `incoming/`
|
||||
2. `_wait_until_stable()` wartet, bis Datei nicht mehr wächst (Scanner schreibt mehrmals)
|
||||
3. Move nach `working/`
|
||||
4. `ocrmypdf.ocr()` als **Library-Call** (kein Subprozess-Start pro PDF — schneller)
|
||||
5. Optional: veraPDF-Validierung (CLI-Subprozess)
|
||||
6. Move nach `outgoing/` als `OCR_<originalname>.pdf`
|
||||
7. Aktive Upload-Targets ausführen (folder/nextcloud/sftp)
|
||||
8. Optional E-Mail-Notify
|
||||
**Beim Start (`run()` wie `run_once()`), vor allem anderen:**
|
||||
1. `check_preflight()` — `tesseract` und `gs` müssen im PATH sein; ist `pdfa_level` gesetzt, wird zusätzlich die Ghostscript-Version gegen den 10.0.0–10.02.0-Bug geprüft
|
||||
2. `check_output_config()` — validiert `original_on_success`, `archive_dir` (Pflicht bei `archive`) und `name_mode`
|
||||
3. Scheitert eines davon → `PreflightError`, CLI beendet sich mit **Exit-Code 2** (ebenso bei kaputter/fehlender Config)
|
||||
|
||||
Fehler → Move nach `error/`, Service läuft weiter (kein `exit 1` wie im alten Bash-Tool).
|
||||
**Pro Datei:**
|
||||
1. `watchdog` triggert auf `created`/`moved`/`closed` in `incoming/` (beim Start greift `_scan_existing()` bereits liegende PDFs auf)
|
||||
2. `_wait_until_stable()` wartet, bis die Datei nicht mehr wächst (max. ~60s)
|
||||
3. Move nach `working/`
|
||||
4. `ocrmypdf.ocr()` als **Library-Call** (kein Subprozess-Start pro PDF)
|
||||
5. Optional: veraPDF-Validierung (CLI-Subprozess) — bei FAIL geht das OCR-Ergebnis nach `error/`, das Original folgt `original_on_success` (wird also bei `archive` **nicht** gelöscht)
|
||||
6. Move nach `outgoing/` unter dem laut `[output]` gebauten Namen (`build_output_name()`: `prefix`/`suffix`/`none` + `name_tag` — das harte `OCR_`-Präfix aus 0.1.0 ist nur noch der Default)
|
||||
7. Original in `working/` wird laut `original_on_success` **gelöscht** oder nach `archive_dir` **archiviert** (Kollision → Timestamp-Suffix)
|
||||
8. Aktive Upload-Targets ausführen (folder/nextcloud/sftp)
|
||||
9. E-Mail-Notify je nach `[notify.email].on`
|
||||
|
||||
**Fehlerbehandlung (Stand 0.4.1):**
|
||||
|
||||
| Fehlerfall | Zählt als Fehler | Wo liegt die Datei danach |
|
||||
|------------|------------------|----------------------------|
|
||||
| Stabilitäts-Check läuft in den Timeout | ja | bleibt in `incoming/`, wird beim nächsten Lauf erneut versucht |
|
||||
| Datei verschwindet vor der Verarbeitung | nein | — |
|
||||
| OCR wirft (ocrmypdf) | ja | `error/` |
|
||||
| veraPDF FAIL | ja | OCR-Ergebnis nach `error/`, Original laut `original_on_success` (`delete` → weg, `archive` → `archive_dir`; seit 0.4.1) |
|
||||
| Beliebige Exception aus `process_pdf()` (z.B. `shutil.move` nach `outgoing/`) | ja | `_rescue_to_error()` sucht in `incoming/` und `working/` und verschiebt nach `error/` |
|
||||
| Mindestens ein Upload-Ziel schlägt fehl | ja | PDF bleibt **bewusst in `outgoing/`** (das OCR war ja erfolgreich), Fehler-Mail nennt die Ziele |
|
||||
|
||||
Der Service läuft in allen Fällen weiter (kein `exit 1` wie im alten Bash-Tool). Im `--once`-Modus liefert die CLI **Exit-Code 1**, sobald `error_count > 0` ist, sonst 0.
|
||||
|
||||
## 🧠 Performance-Entscheidungen
|
||||
|
||||
@@ -83,8 +208,17 @@ Fehler → Move nach `error/`, Service läuft weiter (kein `exit 1` wie im alten
|
||||
- **`--jobs` an ocrmypdf**: Tesseract parallelisiert Seiten innerhalb eines PDFs
|
||||
- **`skip_text=True`**: bereits OCR-haltige Seiten werden nicht neu verarbeitet
|
||||
- **Stabilitäts-Check** statt magic-file `new` (alte Bash-Krücke)
|
||||
- **`upload_folder()` nutzt `shutil.copyfile()`** statt `read_bytes()`/`write_bytes()` — große PDFs landen nicht komplett im RAM
|
||||
- veraPDF nur wenn `enabled=true` (JVM-Start ist teuer)
|
||||
|
||||
## ⚠️ Fallstricke
|
||||
|
||||
- **Ghostscript 10.0.0–10.02.0 zerschießt OCR.** Das ist der Debian-12-Default. In Kombination aus `[ocr].pdfa_level` + `skip_text = true` blockiert ocrmypdf komplett (Issue #3). Deshalb ist `pdfa_level = ""` der sichere Default, und der Preflight bricht mit Exit 2 ab, wenn `pdfa_level` gesetzt **und** die GS-Version betroffen ist. Abhilfe: Ghostscript ≥ 10.02.1 aus bookworm-backports (der Installer bietet das an).
|
||||
- **`[ocr].timeout` ist ein Timeout PRO SEITE**, kein Gesamt-Timeout pro PDF. Der Wert geht als `tesseract_timeout` (ocrmypdf-Option `--tesseract-timeout`) durch; ocrmypdf kennt kein Dokument-Timeout. Wer noch den alten Default `1800` in einer Config stehen hat, gibt Tesseract 30 Minuten **je Seite** — Richtwert ist 300. Ein durchgereichtes `0` würde ocrmypdf dazu bringen, OCR **still zu überspringen**, deshalb wird bei `0` (oder negativ) gar nichts übergeben und der ocrmypdf-Default greift.
|
||||
- **systemd-Hardening bricht in LXC-Containern** (`Error 226/NAMESPACE` durch `PrivateTmp`, `ProtectSystem` usw., Issue #4). Gegenmittel ist das Drop-in `systemd/lxc-compat.conf` nach `/etc/systemd/system/pdf-ocr-hotfolder@.service.d/`; der Installer erkennt Container via `systemd-detect-virt --container` und bietet es an.
|
||||
- **Das Paket wird nicht pip-installiert, sondern nach `/opt/pdf-ocr-hotfolder` kopiert.** Gestartet wird per `python -m pdf_ocr_hotfolder`, gefunden wird das Modul nur über das Arbeitsverzeichnis — `WorkingDirectory=/opt/pdf-ocr-hotfolder` in der Unit ist daher Pflicht, nicht Kosmetik (Issue #5).
|
||||
- **Klartext-Passwörter in der Instanz-Config**: SMTP-, Nextcloud- und SFTP-Zugangsdaten stehen unverschlüsselt in `/etc/pdf-ocr-hotfolder/<instanz>.toml`. Deshalb `chmod 640` und `chown root:<service-gruppe>`, und `/etc/pdf-ocr-hotfolder` selbst `750 root:pdfocr`. Beim Debuggen nicht versehentlich in ein Ticket oder Log kopieren.
|
||||
|
||||
## 🛠️ Entwicklung
|
||||
|
||||
Lokaler Test ohne Installation:
|
||||
@@ -97,9 +231,17 @@ cp config.example.toml /tmp/config.toml
|
||||
python -m pdf_ocr_hotfolder --config /tmp/config.toml
|
||||
```
|
||||
|
||||
Tests (aus dem Repo-Root, `pytest.ini` setzt `testpaths = tests`):
|
||||
```bash
|
||||
pytest # aktuell 95 Tests
|
||||
```
|
||||
|
||||
`ocrmypdf` muss dafür **nicht** installiert sein: der Import in `processor.py` ist lazy, und `tests/test_ocr_timeout.py` schiebt ein Dummy-Modul in `sys.modules`. Die übrigen Tests mocken `process_pdf` bzw. arbeiten nur auf Config-Ebene.
|
||||
|
||||
## 📋 Roadmap / TODO
|
||||
|
||||
- [ ] Tests (`pytest`) für `processor` und `uploaders`
|
||||
- [x] Tests (`pytest`) für `processor` und `uploaders` — 95 Tests
|
||||
- [ ] Test-Lücken schließen: der watchdog-Eventpfad (`_Handler`/`Observer`) wird nirgends getestet, `run_verapdf()` ebenso wenig (der FAIL-*Pfad* in `process_pdf()` ist getestet, die veraPDF-CLI-Anbindung selbst nicht), und `run_ocr()` nur gegen ein gemocktes ocrmypdf — es gibt keinen Test mit einer echten PDF-Datei. Auch `upload_nextcloud()` und `upload_sftp()` sind ungetestet (nur `upload_folder()`).
|
||||
- [ ] Prometheus-Metriken (verarbeitete PDFs, Fehlerquote, Laufzeit)
|
||||
- [ ] CLI-Subkommandos: `pdf-ocr-hotfolder reprocess <error-file>`
|
||||
- [ ] Optional: S3/MinIO Upload-Target
|
||||
|
||||
+193
@@ -1,5 +1,198 @@
|
||||
# Changelog
|
||||
|
||||
## [0.5.0] - 2026-09-22
|
||||
|
||||
### Added
|
||||
- Der Installer weist einen Archiv-Pfad ab, der auf `incoming/`, `outgoing/`,
|
||||
`working/` oder `error/` der Instanz zeigt — im Eingang wuerde das Original
|
||||
sonst endlos neu aufgegriffen.
|
||||
- **`install.sh` fragt beim Anlegen einer Instanz die OCR-Sprachen ab**
|
||||
(`Tesseract-Sprachen [deu+eng]:`). Die Wahl gilt bewusst **pro Instanz** —
|
||||
ein Hotfolder `buchhaltung` kann mit `deu` laufen, ein Hotfolder `export` mit
|
||||
`deu+eng+fra`. Der Installer weist vorher darauf hin, dass jede zusaetzliche
|
||||
Sprache Laufzeit **und** Erkennungsqualitaet kostet, die Liste also eng
|
||||
gehalten werden sollte. Das Eingabeformat wird geprueft (Sprachcodes mit `+`
|
||||
verbunden, `chi_sim` & Co. erlaubt); bei Unsinn wird erneut gefragt statt
|
||||
abzubrechen.
|
||||
- **Sprachpakete werden nachinstalliert.** Jeder eingegebene Code wird gegen
|
||||
`tesseract --list-langs` geprueft. Fehlt eine Sprachdatei, bietet der
|
||||
Installer das passende apt-Paket an (`tesseract-ocr-<code>`, Unterstrich wird
|
||||
zum Bindestrich: `chi_sim` → `tesseract-ocr-chi-sim`). Lehnt der User ab oder
|
||||
laesst sich das Paket nicht installieren, warnt der Installer, dass OCR mit
|
||||
dieser Sprache **bei jeder Datei** scheitern wuerde, und fragt die Sprachen
|
||||
erneut ab — so kann die Sprache einfach wieder rausgeworfen werden. Ist
|
||||
`tesseract` nicht aufrufbar, wird die Pruefung uebersprungen und die Eingabe
|
||||
unveraendert uebernommen.
|
||||
- **Abfrage `Original nach erfolgreichem OCR archivieren? [j/N]:`** — Default
|
||||
nein, also weiterhin `original_on_success = "delete"`. Bei ja wird der
|
||||
Archiv-Pfad abgefragt (Vorschlag `<basis>/archive`), angelegt und auf den
|
||||
Service-User gechownt; ein Archiv ausserhalb des Instanz-Basis-Pfads bekommt
|
||||
ein eigenes `chown -R`.
|
||||
|
||||
### Changed
|
||||
- Die Instanz-Config wird weiterhin per `sed` aus `config.example.toml`
|
||||
erzeugt, substituiert jetzt aber zusaetzlich `[ocr].languages`,
|
||||
`[output].original_on_success` und `[output].archive_dir` — bisher waren das
|
||||
die Beispiel-Defaults, `archive_dir` musste von Hand nachgetragen werden.
|
||||
Die Ausdruecke sind am Zeilenanfang verankert (`^key[[:space:]]*=`), damit die
|
||||
deutschen Kommentarzeilen ueber den Keys unangetastet bleiben, und
|
||||
Pfad-Variablen laufen durch `sed_escape_repl()` (maskiert `\`, `&`, `|`) —
|
||||
Pfade mit Sonderzeichen landen damit korrekt in der Config.
|
||||
- Nach dem sed-Lauf liest der Installer die drei Keys aus der erzeugten Config
|
||||
zurueck und vergleicht sie mit der Eingabe. Erst wenn das passt, nennt die
|
||||
Abschluss-Zusammenfassung zusaetzlich die gewaehlten **Sprachen** und (bei
|
||||
Archivierung) das **Archiv-Verzeichnis**; sonst gibt es eine Warnung.
|
||||
|
||||
## [0.4.1] - 2026-09-22
|
||||
|
||||
### Fixed
|
||||
- **veraPDF-FAIL hat das Original immer gelöscht.** Schlug die PDF/A-Validierung
|
||||
fehl, wanderte das OCR-Ergebnis nach `error/` und das Original wurde per
|
||||
`unlink()` entfernt — unabhängig von `[output].original_on_success`. Wer
|
||||
`archive` konfiguriert hatte, verlor die Datei also ausgerechnet im
|
||||
Fehlerfall. Der FAIL-Pfad nutzt jetzt dieselbe `_dispose_original()`-Logik
|
||||
wie der Erfolgsfall: `archive` legt das Original samt
|
||||
Timestamp-Kollisionsschutz im `archive_dir` ab, `delete` verhält sich wie
|
||||
bisher. Die Log-Meldung nennt jetzt beides — wohin das OCR-Ergebnis ging und
|
||||
was mit dem Original passiert ist.
|
||||
|
||||
### Removed
|
||||
- Das nie benutzte Logverzeichnis `/var/log/pdf-ocr-hotfolder/` wird nicht mehr
|
||||
vom Installer angelegt und ist aus README und Briefing entfernt. Es hat nie
|
||||
ein Logfile enthalten: `_setup_logging()` nutzt `logging.basicConfig()` ohne
|
||||
FileHandler, der Dienst loggt nach stdout → journald. **journald ist damit die
|
||||
einzige Log-Quelle** (`journalctl -u pdf-ocr-hotfolder@<instanz> -f`).
|
||||
Weder Installer noch Updater fassen das Verzeichnis an: ein vorhandenes,
|
||||
leeres `/var/log/pdf-ocr-hotfolder/` kann auf bestehenden Installationen
|
||||
gefahrlos von Hand entfernt werden (`sudo rmdir /var/log/pdf-ocr-hotfolder`).
|
||||
|
||||
### Added
|
||||
- 3 neue Tests für den veraPDF-FAIL-Pfad (`delete`, `archive`,
|
||||
Archiv-Namenskollision); veraPDF wird dabei gemockt. Suite jetzt 95 Tests.
|
||||
|
||||
## [0.4.0] - 2026-09-22
|
||||
|
||||
### Added
|
||||
- `[ocr].timeout` ist jetzt wirksam: der Wert wird als `tesseract_timeout`
|
||||
(Sekunden pro Seite) an ocrmypdf durchgereicht. Bisher war der Key zwar
|
||||
dokumentiert, wurde aber nirgends gelesen.
|
||||
- `check_output_config()` validiert zusätzlich `[output].name_mode`. Ein Tippfehler
|
||||
führt jetzt beim Start zum Abbruch mit Exit-Code 2, statt erst pro Datei
|
||||
zuzuschlagen — und zwar bisher **nach** dem Verschieben nach `working/`,
|
||||
wo die Datei dann liegen blieb.
|
||||
- Neue Exception `ConfigError` in `pdf_ocr_hotfolder.config` — fehlende
|
||||
`[paths]`-Sektion oder ein fehlender Pfad-Eintrag liefern eine deutsche
|
||||
Fehlermeldung mit Datei- und Key-Nennung statt eines nackten `KeyError`-Tracebacks.
|
||||
Die CLI bricht damit sauber mit Exit-Code 2 ab.
|
||||
- `pytest.ini` mit `testpaths = tests`, damit `pytest` aus dem Repo-Root läuft.
|
||||
- 35 neue Tests: Fehlerzählung (Exception, Upload, Stabilitäts-Timeout),
|
||||
Config-Fehlermeldungen, `tesseract_timeout`-Durchreichung (ocrmypdf gemockt)
|
||||
und `upload_folder()`.
|
||||
|
||||
### Changed
|
||||
- **`[ocr].timeout` hat eine neue Bedeutung — für bestehende Installationen relevant!**
|
||||
Der Wert ist kein (nie implementiertes) Gesamt-Timeout pro PDF mehr, sondern
|
||||
das Limit **pro Seite** für Tesseract. Der Default sinkt entsprechend von
|
||||
`1800` auf `300`. Wer den alten Wert `1800` in seiner `config.toml` stehen hat,
|
||||
gibt Tesseract damit 30 Minuten **je Seite** — bitte auf einen Seiten-Wert
|
||||
anpassen (Richtwert 300).
|
||||
`0` bedeutet "kein eigenes Limit": der Wert wird dann gar nicht erst
|
||||
durchgereicht, weil ocrmypdf `tesseract_timeout=0` als "OCR komplett
|
||||
überspringen" interpretiert.
|
||||
- `_dispatch_uploads()` liefert jetzt die Namen der fehlgeschlagenen Upload-Ziele
|
||||
zurück; die doppelte `enabled`-Prüfung (Service + Uploader) ist entfallen —
|
||||
die Uploader prüfen das selbst.
|
||||
- `upload_folder()` kopiert mit `shutil.copyfile()` statt
|
||||
`read_bytes()`/`write_bytes()` — große PDFs landen nicht mehr komplett im
|
||||
Speicher. Die Selbst-Ziel-Erkennung bleibt unverändert.
|
||||
|
||||
### Fixed
|
||||
- `OcrConfig.pdfa_level` hatte im Code noch den Default `"2"`, obwohl
|
||||
`config.example.toml` seit 0.2.2 bewusst `""` setzt (Ghostscript-Bug, Issue #3).
|
||||
Eine Config ohne `[ocr]`-Sektion bzw. ohne den Key lief damit ungewollt in
|
||||
PDF/A. Default im Code jetzt ebenfalls `""`.
|
||||
- Eine Exception **nach** dem OCR (z.B. ein fehlgeschlagener
|
||||
`shutil.move()` nach `outgoing/`) wurde nur im Worker-Callback geloggt.
|
||||
`error_count` blieb 0 und `--once` lieferte trotz Fehlschlag Exit-Code 0.
|
||||
Jede Exception aus `process_pdf()` zählt jetzt als Fehler, wird geloggt,
|
||||
löst eine Fehler-Mail aus und die Datei wandert — soweit noch auffindbar
|
||||
(`incoming/` oder `working/`) — nach `error/`.
|
||||
- Fehlgeschlagene Uploads waren folgenlos: die Rückgabewerte der Uploader wurden
|
||||
verworfen, es ging sogar eine Erfolgs-Mail raus. Jetzt zählt mindestens ein
|
||||
fehlgeschlagenes Ziel als Fehler und die E-Mail geht als **FEHLER** raus, mit
|
||||
Nennung der betroffenen Ziele. Das OCR-PDF bleibt bewusst in `outgoing/`
|
||||
liegen (das OCR selbst war ja erfolgreich) — das steht so auch im Log.
|
||||
- Lief der Stabilitäts-Check einer Datei in den 60-Sekunden-Timeout, gab es nur
|
||||
ein `log.warning`; `--once` meldete Exit-Code 0. Jetzt `log.error` +
|
||||
`error_count`. Die Datei bleibt bewusst in `incoming/` liegen und wird beim
|
||||
nächsten Lauf erneut versucht. Eine zwischenzeitlich *verschwundene* Datei
|
||||
wird davon unterschieden und zählt weiterhin nicht als Fehler.
|
||||
|
||||
## [0.3.1] - 2026-04-10
|
||||
|
||||
### Fixed
|
||||
- **Issue #4**: LXC/Container-Kompatibilität — systemd-Hardening (`PrivateTmp`, `ProtectSystem`, etc.)
|
||||
verursacht Error 226/NAMESPACE in LXC-Containern. Installer erkennt Container-Umgebung automatisch
|
||||
und bietet ein Drop-in an. Zusätzlich liegt `systemd/lxc-compat.conf` als Vorlage im Repo.
|
||||
- **Issue #5**: `WorkingDirectory=/opt/pdf-ocr-hotfolder` in der systemd Template-Unit ergänzt —
|
||||
ohne diesen Eintrag konnte das Python-Modul nicht gefunden werden.
|
||||
- **Issue #6**: Auf Debian 12 bietet der Installer bei betroffenen Ghostscript-Versionen (10.0.0–10.02.0)
|
||||
jetzt automatisch an, bookworm-backports zu aktivieren und GS zu upgraden (statt nur zu warnen).
|
||||
|
||||
## [0.3.0] - 2026-04-09
|
||||
|
||||
### Added
|
||||
- Neue Config-Sektion `[output]` mit:
|
||||
- `name_mode` — Platzierung des Tags im Dateinamen: `"prefix"`, `"suffix"` (vor Extension), `"none"`
|
||||
- `name_tag` — verbatim einzufügender String, z.B. `"OCR_"` oder `"_OCR"`
|
||||
- `original_on_success` — `"delete"` (alter Default) oder `"archive"`
|
||||
- `archive_dir` — Zielverzeichnis für `"archive"`, mit Kollisions-Schutz (Timestamp-Suffix)
|
||||
- Runtime-Validierung der Output-Config in `check_output_config()`
|
||||
- 20 neue Tests für `build_output_name()`, `check_output_config()` und `process_pdf()`
|
||||
mit allen Kombinationen aus Modus + Original-Behandlung
|
||||
|
||||
### Changed
|
||||
- `process_pdf()` nimmt jetzt `output_cfg: OutputConfig` als Pflicht-Argument
|
||||
|
||||
## [0.2.2] - 2026-04-09
|
||||
|
||||
### Fixed
|
||||
- **Issue #3**: Ghostscript 10.0.0–10.02.0 (Debian 12 default) zerschießen OCR mit PDF/A + `skip_text=true`.
|
||||
- `config.example.toml`: `pdfa_level = ""` als sicherer Default
|
||||
- Runtime-Preflight: Prüft `gs --version` wenn `pdfa_level` gesetzt ist, bricht mit klarer Fehlermeldung ab
|
||||
- `install.sh`: warnt bei betroffenen GS-Versionen mit Upgrade-Hinweis auf bookworm-backports
|
||||
|
||||
### Added
|
||||
- `is_ghostscript_broken()` / `detect_ghostscript_version()` in `pdf_ocr_hotfolder.service`
|
||||
- 19 weitere pytest-Tests für GS-Versions-Detection (parametrisiert) und Preflight-Kombinationen
|
||||
|
||||
## [0.2.1] - 2026-04-09
|
||||
|
||||
### Fixed
|
||||
- **Issue #1**: Preflight-Check beim Start prüft jetzt `tesseract` und `gs` (Ghostscript). Fehlt eine Abhängigkeit, beendet sich der Service sofort mit Exit-Code 2 und klarer Fehlermeldung statt erst bei der ersten Datei.
|
||||
- **Issue #2**: `--once`-Modus liefert jetzt Exit-Code `1`, sobald **mindestens ein** PDF fehlgeschlagen ist. Exit-Code `0` nur bei vollständigem Erfolg (inkl. "keine Dateien vorhanden"). Exit-Code `2` bei Preflight-Fehler.
|
||||
|
||||
### Added
|
||||
- Public API: `HotfolderService.run_once()`, `.success_count`, `.error_count`, `.ensure_dirs()`
|
||||
- `check_preflight()` / `PreflightError` in `pdf_ocr_hotfolder.service`
|
||||
- pytest-Test-Suite (`tests/`) mit 11 Tests — deckt alle Szenarien aus Issue #1 und #2 ab
|
||||
- `ocrmypdf`-Import in `processor.py` ist jetzt lazy (Tests ohne ocrmypdf-Installation möglich)
|
||||
|
||||
## [0.2.0] - 2026-04-08
|
||||
|
||||
### Added
|
||||
- **Multi-Instanz-Support** via systemd Template-Unit `pdf-ocr-hotfolder@<name>.service`
|
||||
- Pro Instanz: eigene Config (`/etc/pdf-ocr-hotfolder/<name>.toml`), eigene Datenverzeichnisse (`/var/lib/pdf-ocr-hotfolder/<name>/…`), optional eigener Service-User via Drop-in
|
||||
- **Instanz-Manager in `install.sh`**: erkennt bestehende Instanzen bei Re-Run, fragt nach weiteren, listet Namen + Status
|
||||
- `update.sh` stoppt/startet automatisch **alle** laufenden Instanzen
|
||||
|
||||
### Changed
|
||||
- Single-Unit `pdf-ocr-hotfolder.service` durch Template-Unit `pdf-ocr-hotfolder@.service` ersetzt
|
||||
- Installer fragt nicht mehr einmalig nach Service-User, sondern **pro Instanz**
|
||||
|
||||
### Removed
|
||||
- Alte Single-Config unter `/etc/pdf-ocr-hotfolder/config.toml` — wird nicht mehr erzeugt
|
||||
|
||||
## [0.1.0] - 2026-04-08
|
||||
|
||||
### Added
|
||||
|
||||
@@ -23,49 +23,118 @@ cd pdf-ocr-hotfolder
|
||||
sudo ./install.sh
|
||||
```
|
||||
|
||||
Der Installer fragt nach dem Service-User. Standardmäßig wird ein lokaler System-User `pdfocr` angelegt. Wenn der User bereits existiert (z.B. AD via SSSD), wird er einfach übernommen.
|
||||
Der Installer:
|
||||
1. Installiert einmalig Code + venv + systemd-Template-Unit
|
||||
2. Fragt **pro Instanz** ab:
|
||||
- Instanz-Name
|
||||
- Basis-Pfad für die Daten
|
||||
- Service-User
|
||||
- **OCR-Sprachen** (Tesseract, Default `deu+eng`) — fehlende Sprachpakete
|
||||
(`tesseract-ocr-<code>`) werden erkannt und auf Wunsch nachinstalliert
|
||||
- **Original nach erfolgreichem OCR archivieren?** (Default nein = löschen;
|
||||
bei ja zusätzlich der Archiv-Pfad, vorgeschlagen `<basis>/archive`)
|
||||
3. Legt so viele Hotfolder-Instanzen an, wie du willst (`Weitere Instanz anlegen? [j/N]`)
|
||||
|
||||
Danach Konfiguration anpassen:
|
||||
|
||||
```bash
|
||||
sudo nano /etc/pdf-ocr-hotfolder/config.toml
|
||||
sudo systemctl restart pdf-ocr-hotfolder
|
||||
```
|
||||
Bei jedem erneuten Aufruf erkennt der Installer bestehende Instanzen und fragt nur nach neuen.
|
||||
|
||||
Test:
|
||||
|
||||
```bash
|
||||
cp irgendein-scan.pdf /var/lib/pdf-ocr-hotfolder/incoming/
|
||||
journalctl -u pdf-ocr-hotfolder -f
|
||||
cp irgendein-scan.pdf /var/lib/pdf-ocr-hotfolder/<instanz>/incoming/
|
||||
journalctl -u pdf-ocr-hotfolder@<instanz> -f
|
||||
```
|
||||
|
||||
Nach wenigen Sekunden liegt das OCR-PDF unter `/var/lib/pdf-ocr-hotfolder/outgoing/OCR_irgendein-scan.pdf`.
|
||||
Nach wenigen Sekunden liegt das OCR-PDF im `outgoing/`-Ordner der Instanz.
|
||||
|
||||
## Multi-Instanz-Betrieb
|
||||
|
||||
Das Tool arbeitet komplett **instanzbasiert** über eine systemd Template-Unit `pdf-ocr-hotfolder@<name>.service`. Jede Instanz hat:
|
||||
|
||||
- eigene Config-Datei: `/etc/pdf-ocr-hotfolder/<name>.toml`
|
||||
- eigene Datenverzeichnisse: `/var/lib/pdf-ocr-hotfolder/<name>/{incoming,working,outgoing,error}/`
|
||||
- eigene systemd-Unit: `pdf-ocr-hotfolder@<name>.service`
|
||||
- optional eigenen Service-User (via Drop-in `/etc/systemd/system/pdf-ocr-hotfolder@<name>.service.d/user.conf`)
|
||||
- **eigene OCR-Sprachen und eigene Original-Behandlung** (löschen oder archivieren)
|
||||
|
||||
### Sprachen pro Instanz
|
||||
|
||||
Die Tesseract-Sprachen werden bewusst **je Instanz** abgefragt, nicht global:
|
||||
Hotfolder haben unterschiedliche Post. Ein Buchhaltungs-Hotfolder sieht nur
|
||||
deutsche Belege, ein Export-Hotfolder internationale Korrespondenz:
|
||||
|
||||
```toml
|
||||
# /etc/pdf-ocr-hotfolder/buchhaltung.toml
|
||||
languages = "deu"
|
||||
|
||||
# /etc/pdf-ocr-hotfolder/export.toml
|
||||
languages = "deu+eng+fra"
|
||||
```
|
||||
|
||||
**Die Liste so eng wie möglich halten.** Jede zusätzliche Sprache kostet
|
||||
Laufzeit *und* Erkennungsqualität: Tesseract muss mehr Modelle gegeneinander
|
||||
abwägen und verwechselt dabei Wörter, die in der einen Sprache eindeutig wären.
|
||||
`deu+eng+fra` auf reinen Deutsch-Scans ist also kein Sicherheitsnetz, sondern
|
||||
ein Rückschritt.
|
||||
|
||||
Beispiel für 3 Hotfolder:
|
||||
|
||||
```bash
|
||||
sudo ./install.sh
|
||||
# → legt z.B. kunde-a, kunde-b, buchhaltung an
|
||||
|
||||
systemctl status 'pdf-ocr-hotfolder@*'
|
||||
journalctl -u pdf-ocr-hotfolder@kunde-a -f
|
||||
```
|
||||
|
||||
Manuell eine weitere Instanz anlegen geht auch — einfach `install.sh` erneut starten, er fragt wieder nach.
|
||||
|
||||
## Verzeichnisse
|
||||
|
||||
| Pfad | Zweck |
|
||||
|------|-------|
|
||||
| `/etc/pdf-ocr-hotfolder/config.toml` | Konfiguration |
|
||||
| `/var/lib/pdf-ocr-hotfolder/incoming` | Eingang (Scanner schreibt hier rein) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/working` | Arbeitsverzeichnis während OCR |
|
||||
| `/var/lib/pdf-ocr-hotfolder/outgoing` | Ausgang (fertige PDFs) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/error` | PDFs, die nicht verarbeitet werden konnten |
|
||||
| `/opt/pdf-ocr-hotfolder/` | Code + venv |
|
||||
| `/var/log/pdf-ocr-hotfolder/` | Logs (zusätzlich zu journald) |
|
||||
| `/opt/pdf-ocr-hotfolder/` | Code + venv (für alle Instanzen gemeinsam) |
|
||||
| `/etc/pdf-ocr-hotfolder/<instanz>.toml` | Config pro Instanz |
|
||||
| `/etc/systemd/system/pdf-ocr-hotfolder@.service` | systemd Template-Unit |
|
||||
| `/var/lib/pdf-ocr-hotfolder/<instanz>/incoming` | Eingang (Scanner schreibt hier rein) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/<instanz>/working` | Arbeitsverzeichnis während OCR |
|
||||
| `/var/lib/pdf-ocr-hotfolder/<instanz>/outgoing` | Ausgang (fertige PDFs) |
|
||||
| `/var/lib/pdf-ocr-hotfolder/<instanz>/error` | Fehlgeschlagene PDFs |
|
||||
| `/var/backups/pdf-ocr-hotfolder/` | Update-Backups |
|
||||
|
||||
## Konfiguration
|
||||
|
||||
Vollständiges Beispiel: [`config.example.toml`](config.example.toml). Wichtigste Sektionen:
|
||||
|
||||
Der Installer fragt `[ocr].languages`, `[output].original_on_success` und
|
||||
`[output].archive_dir` pro Instanz ab und schreibt sie direkt in die
|
||||
Instanz-Config — die Werte unten sind nur die Beispiel-Defaults.
|
||||
|
||||
### `[ocr]`
|
||||
```toml
|
||||
languages = "deu+eng" # Tesseract-Sprachen
|
||||
languages = "deu+eng" # Tesseract-Sprachen (Installer fragt pro Instanz)
|
||||
jobs = 4 # Threads pro PDF
|
||||
skip_text = true # bereits OCR-haltige Seiten überspringen
|
||||
pdfa_level = "2" # "1", "2", "3" oder "" für reines PDF
|
||||
pdfa_level = "" # "1", "2", "3" oder "" für reines PDF (Default "" wegen Ghostscript-Bug, s.u.)
|
||||
deskew = true
|
||||
max_workers = 2 # parallele PDFs
|
||||
timeout = 1800
|
||||
timeout = 300 # max. Sekunden pro SEITE (Tesseract), 0 = ocrmypdf-Default
|
||||
```
|
||||
|
||||
### `[output]`
|
||||
```toml
|
||||
# Dateiname im outgoing/:
|
||||
# "prefix" → OCR_scan.pdf
|
||||
# "suffix" → scan_OCR.pdf (vor der Extension)
|
||||
# "none" → scan.pdf (unverändert)
|
||||
name_mode = "prefix"
|
||||
name_tag = "OCR_"
|
||||
|
||||
# Nach erfolgreichem OCR mit dem Original:
|
||||
# "delete" → löschen
|
||||
# "archive" → in archive_dir verschieben
|
||||
# Beides fragt der Installer beim Anlegen der Instanz ab:
|
||||
original_on_success = "delete"
|
||||
archive_dir = "" # absoluter Pfad, Pflicht bei "archive"
|
||||
```
|
||||
|
||||
### `[upload.nextcloud]`
|
||||
@@ -101,9 +170,24 @@ on = "errors" # always | errors | never
|
||||
## Service-Verwaltung
|
||||
|
||||
```bash
|
||||
sudo systemctl status pdf-ocr-hotfolder
|
||||
sudo systemctl restart pdf-ocr-hotfolder
|
||||
journalctl -u pdf-ocr-hotfolder -f
|
||||
# Eine bestimmte Instanz
|
||||
sudo systemctl status pdf-ocr-hotfolder@kunde-a
|
||||
sudo systemctl restart pdf-ocr-hotfolder@kunde-a
|
||||
journalctl -u pdf-ocr-hotfolder@kunde-a -f
|
||||
|
||||
# Alle Instanzen
|
||||
sudo systemctl status 'pdf-ocr-hotfolder@*'
|
||||
sudo systemctl restart 'pdf-ocr-hotfolder@*'
|
||||
```
|
||||
|
||||
### Logs
|
||||
|
||||
Der Dienst schreibt **kein eigenes Logfile** — alles geht nach stdout und damit
|
||||
ins journal:
|
||||
|
||||
```bash
|
||||
journalctl -u pdf-ocr-hotfolder@<instanz> -f # eine Instanz mitlesen
|
||||
journalctl -u 'pdf-ocr-hotfolder@*' --since today # alle Instanzen, heute
|
||||
```
|
||||
|
||||
## Update
|
||||
@@ -114,15 +198,22 @@ git pull
|
||||
sudo ./update.sh
|
||||
```
|
||||
|
||||
`update.sh`:
|
||||
1. Stoppt alle laufenden Instanzen
|
||||
2. Sichert den alten Code nach `/var/backups/pdf-ocr-hotfolder/`
|
||||
3. Aktualisiert Code + venv + systemd-Template-Unit in `/opt/pdf-ocr-hotfolder/`
|
||||
4. Startet alle zuvor laufenden Instanzen neu
|
||||
|
||||
Config-Dateien unter `/etc/pdf-ocr-hotfolder/` werden **nie** überschrieben.
|
||||
Das Repo muss bestehen bleiben — `update.sh` kopiert daraus.
|
||||
|
||||
## Manueller Lauf (One-Shot)
|
||||
|
||||
Bestehende PDFs im Eingang einmalig verarbeiten und beenden:
|
||||
Bestehende PDFs einer Instanz einmalig verarbeiten und beenden:
|
||||
|
||||
```bash
|
||||
sudo -u pdfocr /opt/pdf-ocr-hotfolder/venv/bin/python -m pdf_ocr_hotfolder \
|
||||
--config /etc/pdf-ocr-hotfolder/config.toml --once
|
||||
--config /etc/pdf-ocr-hotfolder/kunde-a.toml --once
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
@@ -141,6 +232,24 @@ Service-User braucht **rw** auf alle vier Verzeichnisse unter `/var/lib/pdf-ocr-
|
||||
sudo chown -R DOMAIN\\scanuser:DOMAIN\\scangroup /var/lib/pdf-ocr-hotfolder
|
||||
```
|
||||
|
||||
### LXC/Container: Error 226/NAMESPACE
|
||||
In LXC-Containern schlagen systemd-Hardening-Optionen fehl. Der Installer erkennt Container automatisch und bietet ein Drop-in an. Manuell:
|
||||
```bash
|
||||
sudo mkdir -p /etc/systemd/system/pdf-ocr-hotfolder@.service.d/
|
||||
sudo cp /opt/pdf-ocr-hotfolder/systemd/lxc-compat.conf \
|
||||
/etc/systemd/system/pdf-ocr-hotfolder@.service.d/
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl restart 'pdf-ocr-hotfolder@*'
|
||||
```
|
||||
|
||||
### Ghostscript PDF/A-Bug auf Debian 12
|
||||
GS 10.00.0–10.02.0 (Debian 12 Default) zerstört OCR bei `pdfa_level` + `skip_text=true`. Der Installer bietet automatisch bookworm-backports an. Manuell:
|
||||
```bash
|
||||
echo 'deb http://deb.debian.org/debian bookworm-backports main' | \
|
||||
sudo tee /etc/apt/sources.list.d/bookworm-backports.list
|
||||
sudo apt update && sudo apt install -t bookworm-backports ghostscript
|
||||
```
|
||||
|
||||
### veraPDF-Validierung schlägt immer fehl
|
||||
veraPDF binary prüfen (`[verapdf].binary`). Wenn nicht zwingend gebraucht: `enabled = false`.
|
||||
|
||||
@@ -172,5 +281,5 @@ MIT — © Sonith UG
|
||||
|
||||
---
|
||||
|
||||
**Version:** 0.1.0
|
||||
**Version:** 0.5.0
|
||||
**Repo:** https://gitea.sonith.de/sonith_ug/pdf-ocr-hotfolder
|
||||
|
||||
+27
-3
@@ -21,15 +21,39 @@ skip_text = true
|
||||
# Auflösung für gerasterte Seiten
|
||||
oversample = 300
|
||||
# PDF/A-Konformitätsstufe ("1", "2", "3" oder leer für keinen PDF/A-Output)
|
||||
pdfa_level = "2"
|
||||
# ACHTUNG: Ghostscript 10.0.0 bis 10.02.0 (Debian 12 default!) haben einen Bug,
|
||||
# der mit pdfa_level + skip_text=true ocrmypdf komplett blockiert.
|
||||
# Sicherer Default ist "" — nur auf "1"/"2"/"3" setzen, wenn gs >= 10.02.1 installiert ist.
|
||||
pdfa_level = ""
|
||||
# Schiefe Scans automatisch begradigen
|
||||
deskew = true
|
||||
# Hintergrund säubern
|
||||
clean = false
|
||||
# Maximale parallele PDFs (Hauptsystem hat selten mehr als 1-2 gleichzeitig)
|
||||
max_workers = 2
|
||||
# Timeout pro PDF in Sekunden
|
||||
timeout = 1800
|
||||
# Max. Sekunden, die Tesseract pro SEITE laufen darf (0 = kein Limit).
|
||||
# ocrmypdf kennt kein Gesamt-Timeout pro Dokument, nur dieses Seiten-Limit
|
||||
# (ocrmypdf-Option --tesseract-timeout). Läuft eine Seite in den Timeout,
|
||||
# wird sie ohne Textebene ins Ergebnis übernommen; die Verarbeitung der
|
||||
# restlichen Seiten läuft weiter.
|
||||
# 0 = wir geben kein Limit vor und überlassen es dem ocrmypdf-Default.
|
||||
timeout = 300
|
||||
|
||||
[output]
|
||||
# Wie soll die Ziel-Datei im outgoing/-Ordner benannt werden?
|
||||
# "prefix" : name_tag wird vor den Dateinamen gestellt (OCR_scan.pdf)
|
||||
# "suffix" : name_tag wird vor die Extension gestellt (scan_OCR.pdf)
|
||||
# "none" : Dateiname bleibt wie das Original
|
||||
name_mode = "prefix"
|
||||
# Verbatim einzufügender String. Leerer String = kein Tag (wie mode="none").
|
||||
# Beispiele: "OCR_", "[OCR]_", "_OCR", "_searchable"
|
||||
name_tag = "OCR_"
|
||||
# Was passiert mit dem Original, wenn OCR erfolgreich war?
|
||||
# "delete" : Original wird gelöscht (alter Standard)
|
||||
# "archive" : Original wird in archive_dir verschoben
|
||||
original_on_success = "delete"
|
||||
# Absoluter Pfad; nur relevant wenn original_on_success = "archive"
|
||||
archive_dir = ""
|
||||
|
||||
[verapdf]
|
||||
# PDF/A-Validierung (optional)
|
||||
|
||||
+383
-87
@@ -1,11 +1,13 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# PDF OCR Hotfolder — Installer für Debian 12/13
|
||||
# PDF OCR Hotfolder — Installer / Instanz-Manager für Debian 12/13
|
||||
#
|
||||
# Fragt interaktiv nach dem Service-User. Unterstützt:
|
||||
# - Lokal anlegen (neuer System-User)
|
||||
# - Bereits existierender lokaler User
|
||||
# - AD-User mit lokaler UID (z.B. via SSSD/Winbind)
|
||||
# Basis-Installation erfolgt einmalig (Code, venv, systemd-Template-Unit).
|
||||
# Danach werden Hotfolder-Instanzen verwaltet:
|
||||
# - Beim Erstlauf: mindestens eine Instanz wird angelegt
|
||||
# - Beim Folgelauf: bestehende Instanzen werden erkannt; neue können ergänzt werden
|
||||
#
|
||||
# Unterstützt lokale System-User und AD-User mit lokaler UID (SSSD/Winbind).
|
||||
#
|
||||
|
||||
set -euo pipefail
|
||||
@@ -14,7 +16,7 @@ RED='\033[0;31m'; GREEN='\033[0;32m'; YELLOW='\033[1;33m'; BLUE='\033[0;34m'; NC
|
||||
log_info() { echo -e "${GREEN}[INFO]${NC} $*"; }
|
||||
log_warn() { echo -e "${YELLOW}[WARN]${NC} $*"; }
|
||||
log_error() { echo -e "${RED}[ERROR]${NC} $*"; }
|
||||
log_step() { echo -e "${BLUE}==>${NC} $*"; }
|
||||
log_step() { echo -e "\n${BLUE}==>${NC} $*"; }
|
||||
|
||||
if [ "${EUID}" -ne 0 ]; then
|
||||
log_error "Bitte als root ausführen: sudo ./install.sh"
|
||||
@@ -23,9 +25,9 @@ fi
|
||||
|
||||
INSTALL_DIR="/opt/pdf-ocr-hotfolder"
|
||||
CONFIG_DIR="/etc/pdf-ocr-hotfolder"
|
||||
DATA_DIR="/var/lib/pdf-ocr-hotfolder"
|
||||
LOG_DIR="/var/log/pdf-ocr-hotfolder"
|
||||
SERVICE_NAME="pdf-ocr-hotfolder"
|
||||
DATA_ROOT="/var/lib/pdf-ocr-hotfolder"
|
||||
SERVICE_TEMPLATE="pdf-ocr-hotfolder@.service"
|
||||
DEFAULT_USER="pdfocr"
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
REPO_DIR="$SCRIPT_DIR"
|
||||
@@ -35,123 +37,417 @@ if [ ! -f "$REPO_DIR/pdf_ocr_hotfolder/__init__.py" ]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo
|
||||
echo "=========================================="
|
||||
echo " PDF OCR Hotfolder — Installation"
|
||||
echo "=========================================="
|
||||
echo
|
||||
|
||||
# ============ 1. System-Dependencies ============
|
||||
log_step "Installiere System-Pakete"
|
||||
# ============================================================
|
||||
# Basis-Installation (idempotent)
|
||||
# ============================================================
|
||||
|
||||
install_base() {
|
||||
log_step "System-Pakete installieren"
|
||||
apt-get update -qq
|
||||
apt-get install -y --no-install-recommends \
|
||||
python3 python3-venv python3-pip \
|
||||
tesseract-ocr tesseract-ocr-deu tesseract-ocr-eng \
|
||||
ghostscript qpdf unpaper pngquant \
|
||||
icc-profiles-free \
|
||||
ca-certificates curl
|
||||
icc-profiles-free ca-certificates curl
|
||||
log_info "System-Pakete ok ✓"
|
||||
|
||||
log_info "System-Pakete installiert ✓"
|
||||
|
||||
# ============ 2. Service-User ============
|
||||
log_step "Service-User konfigurieren"
|
||||
|
||||
read -r -p "Service-User-Name [pdfocr]: " SERVICE_USER
|
||||
SERVICE_USER="${SERVICE_USER:-pdfocr}"
|
||||
|
||||
if id "$SERVICE_USER" &>/dev/null; then
|
||||
log_info "User '$SERVICE_USER' existiert bereits (lokal oder via AD)."
|
||||
SERVICE_GROUP="$(id -gn "$SERVICE_USER")"
|
||||
log_info "Verwende bestehende primäre Gruppe: $SERVICE_GROUP"
|
||||
# Ghostscript-Versions-Check (Issue #3 + Issue #6)
|
||||
if command -v gs >/dev/null 2>&1; then
|
||||
GS_VER="$(gs --version 2>/dev/null || echo 0.0)"
|
||||
log_info "Ghostscript: $GS_VER"
|
||||
case "$GS_VER" in
|
||||
10.0.0|10.00.0|10.01.*|10.02.0)
|
||||
echo
|
||||
log_warn "═══════════════════════════════════════════════════════════════"
|
||||
log_warn "Ghostscript $GS_VER ist vom PDF/A-Bug betroffen (10.0.0–10.02.0)."
|
||||
log_warn "Mit pdfa_level + skip_text=true kann ocrmypdf KEINE PDFs verarbeiten."
|
||||
log_warn "═══════════════════════════════════════════════════════════════"
|
||||
echo
|
||||
# Prüfe ob Debian bookworm (12) — Backports anbieten
|
||||
if grep -q 'bookworm' /etc/os-release 2>/dev/null; then
|
||||
read -r -p "Ghostscript via bookworm-backports upgraden? [J/n]: " UPGRADE_GS
|
||||
UPGRADE_GS="${UPGRADE_GS:-J}"
|
||||
if [[ "$UPGRADE_GS" =~ ^[JjYy]$ ]]; then
|
||||
log_info "Aktiviere bookworm-backports..."
|
||||
if ! grep -q 'bookworm-backports' /etc/apt/sources.list /etc/apt/sources.list.d/*.list 2>/dev/null; then
|
||||
echo 'deb http://deb.debian.org/debian bookworm-backports main' \
|
||||
> /etc/apt/sources.list.d/bookworm-backports.list
|
||||
apt-get update -qq
|
||||
fi
|
||||
apt-get install -y -t bookworm-backports ghostscript
|
||||
GS_VER_NEW="$(gs --version 2>/dev/null || echo '?')"
|
||||
log_info "Ghostscript aktualisiert: $GS_VER → $GS_VER_NEW ✓"
|
||||
else
|
||||
log_warn "User '$SERVICE_USER' existiert nicht."
|
||||
read -r -p "Lokal als System-User anlegen? [J/n]: " CREATE_USER
|
||||
CREATE_USER="${CREATE_USER:-J}"
|
||||
if [[ "$CREATE_USER" =~ ^[JjYy]$ ]]; then
|
||||
adduser --system --group --home "$DATA_DIR" --shell /usr/sbin/nologin "$SERVICE_USER"
|
||||
SERVICE_GROUP="$SERVICE_USER"
|
||||
log_info "Lokaler System-User '$SERVICE_USER' angelegt ✓"
|
||||
log_warn "Workaround: In der Config [ocr].pdfa_level = \"\" setzen (Default ab v0.2.2)"
|
||||
fi
|
||||
else
|
||||
log_error "User '$SERVICE_USER' muss vor der Installation existieren (z.B. via AD/SSSD)."
|
||||
log_error "Lege ihn an oder wähle einen existierenden Namen."
|
||||
exit 1
|
||||
log_warn "Kein Debian bookworm erkannt — manuelles Upgrade nötig."
|
||||
log_warn "Workaround: In der Config [ocr].pdfa_level = \"\" setzen (Default ab v0.2.2)"
|
||||
fi
|
||||
echo
|
||||
;;
|
||||
esac
|
||||
fi
|
||||
|
||||
# LXC/Container-Erkennung (Issue #4)
|
||||
if systemd-detect-virt --container -q 2>/dev/null; then
|
||||
VIRT_TYPE="$(systemd-detect-virt --container 2>/dev/null || echo 'container')"
|
||||
log_warn "Container-Umgebung erkannt ($VIRT_TYPE)."
|
||||
log_warn "systemd-Hardening kann in Containern fehlschlagen (Error 226/NAMESPACE)."
|
||||
read -r -p "LXC-Kompatibilitäts-Drop-in installieren? [J/n]: " LXC_FIX
|
||||
LXC_FIX="${LXC_FIX:-J}"
|
||||
if [[ "$LXC_FIX" =~ ^[JjYy]$ ]]; then
|
||||
local LXC_DROPIN_DIR="/etc/systemd/system/pdf-ocr-hotfolder@.service.d"
|
||||
mkdir -p "$LXC_DROPIN_DIR"
|
||||
cp "$REPO_DIR/systemd/lxc-compat.conf" "$LXC_DROPIN_DIR/lxc-compat.conf"
|
||||
systemctl daemon-reload
|
||||
log_info "LXC-Kompatibilitäts-Drop-in installiert ✓"
|
||||
fi
|
||||
fi
|
||||
|
||||
# ============ 3. Verzeichnisse ============
|
||||
log_step "Verzeichnisse erstellen"
|
||||
log_step "Default-User '$DEFAULT_USER' prüfen"
|
||||
if id "$DEFAULT_USER" &>/dev/null; then
|
||||
log_info "'$DEFAULT_USER' existiert bereits"
|
||||
else
|
||||
adduser --system --group --home "$DATA_ROOT" --shell /usr/sbin/nologin "$DEFAULT_USER"
|
||||
log_info "System-User '$DEFAULT_USER' angelegt ✓"
|
||||
fi
|
||||
|
||||
mkdir -p "$INSTALL_DIR" "$CONFIG_DIR" "$LOG_DIR"
|
||||
mkdir -p "$DATA_DIR"/{incoming,outgoing,working,error}
|
||||
log_step "Verzeichnisse anlegen"
|
||||
mkdir -p "$INSTALL_DIR" "$CONFIG_DIR" "$DATA_ROOT"
|
||||
chown root:"$DEFAULT_USER" "$CONFIG_DIR"
|
||||
chmod 750 "$CONFIG_DIR"
|
||||
|
||||
log_step "Code kopieren"
|
||||
rm -rf "$INSTALL_DIR/pdf_ocr_hotfolder"
|
||||
cp -r "$REPO_DIR/pdf_ocr_hotfolder" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/requirements.txt" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/VERSION" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/config.example.toml" "$INSTALL_DIR/"
|
||||
echo "$REPO_DIR" > "$INSTALL_DIR/.repo_path"
|
||||
|
||||
if [ ! -f "$CONFIG_DIR/config.toml" ]; then
|
||||
cp "$REPO_DIR/config.example.toml" "$CONFIG_DIR/config.toml"
|
||||
log_info "Beispiel-Konfig nach $CONFIG_DIR/config.toml kopiert"
|
||||
else
|
||||
log_info "Bestehende Konfig $CONFIG_DIR/config.toml bleibt unverändert"
|
||||
fi
|
||||
|
||||
log_info "Verzeichnisse erstellt ✓"
|
||||
|
||||
# ============ 4. Python venv ============
|
||||
log_step "Python venv anlegen"
|
||||
|
||||
log_step "Python venv"
|
||||
if [ ! -d "$INSTALL_DIR/venv" ]; then
|
||||
python3 -m venv "$INSTALL_DIR/venv"
|
||||
fi
|
||||
"$INSTALL_DIR/venv/bin/pip" install --upgrade pip -q
|
||||
"$INSTALL_DIR/venv/bin/pip" install -r "$INSTALL_DIR/requirements.txt" -q
|
||||
log_info "venv ok ✓"
|
||||
|
||||
log_info "venv bereit ✓"
|
||||
log_step "systemd Template-Unit installieren"
|
||||
cp "$REPO_DIR/systemd/$SERVICE_TEMPLATE" "/etc/systemd/system/$SERVICE_TEMPLATE"
|
||||
systemctl daemon-reload
|
||||
log_info "Template-Unit installiert ✓"
|
||||
|
||||
# ============ 5. Berechtigungen ============
|
||||
log_step "Berechtigungen setzen"
|
||||
chown -R "$DEFAULT_USER":"$DEFAULT_USER" "$INSTALL_DIR"
|
||||
}
|
||||
|
||||
chown -R "$SERVICE_USER:$SERVICE_GROUP" "$INSTALL_DIR" "$DATA_DIR" "$LOG_DIR"
|
||||
chown root:"$SERVICE_GROUP" "$CONFIG_DIR"
|
||||
chmod 750 "$CONFIG_DIR"
|
||||
if [ -f "$CONFIG_DIR/config.toml" ]; then
|
||||
chown root:"$SERVICE_GROUP" "$CONFIG_DIR/config.toml"
|
||||
chmod 640 "$CONFIG_DIR/config.toml"
|
||||
# ============================================================
|
||||
# Instanz-Verwaltung
|
||||
# ============================================================
|
||||
|
||||
list_instances() {
|
||||
find "$CONFIG_DIR" -maxdepth 1 -name '*.toml' -type f 2>/dev/null \
|
||||
| sed 's|.*/||; s|\.toml$||' \
|
||||
| sort
|
||||
}
|
||||
|
||||
show_existing_instances() {
|
||||
local instances
|
||||
mapfile -t instances < <(list_instances)
|
||||
if [ "${#instances[@]}" -eq 0 ]; then
|
||||
log_info "Keine bestehenden Instanzen gefunden."
|
||||
return
|
||||
fi
|
||||
echo
|
||||
log_info "Bestehende Instanzen:"
|
||||
for name in "${instances[@]}"; do
|
||||
local active
|
||||
active=$(systemctl is-active "pdf-ocr-hotfolder@${name}.service" 2>/dev/null || echo inactive)
|
||||
printf " • %-30s [%s]\n" "$name" "$active"
|
||||
done
|
||||
echo
|
||||
}
|
||||
|
||||
# Liest den Wert eines Keys (erste Zuweisung am Zeilenanfang) aus einer Config
|
||||
config_value() {
|
||||
local file="$1" key="$2"
|
||||
sed -n "s|^${key}[[:space:]]*=[[:space:]]*\"\(.*\)\"[[:space:]]*$|\1|p" "$file" | head -n1
|
||||
}
|
||||
|
||||
# Maskiert Sonderzeichen, damit ein Pfad gefahrlos in eine sed-Ersetzung darf
|
||||
# (Trennzeichen '|', Rueckverweis '&', Backslash).
|
||||
sed_escape_repl() {
|
||||
printf '%s' "$1" | sed -e 's/[\\&|]/\\&/g'
|
||||
}
|
||||
|
||||
# Prueft jeden Tesseract-Sprachcode gegen die installierten Sprachdateien und
|
||||
# bietet fehlende Pakete zur Installation an.
|
||||
# Rueckgabe: 0 = alle Sprachen verfuegbar (oder Pruefung nicht moeglich),
|
||||
# 1 = mindestens eine Sprache fehlt weiterhin.
|
||||
ensure_tesseract_langs() {
|
||||
local langs="$1"
|
||||
local raw installed code pkg answer rc=0
|
||||
local -a codes
|
||||
|
||||
if ! command -v tesseract >/dev/null 2>&1; then
|
||||
log_warn "tesseract ist nicht aufrufbar — Sprachpruefung wird uebersprungen."
|
||||
log_warn "Eingabe '$langs' wird unveraendert uebernommen."
|
||||
return 0
|
||||
fi
|
||||
if ! raw="$(tesseract --list-langs 2>/dev/null)"; then
|
||||
log_warn "'tesseract --list-langs' schlug fehl — Sprachpruefung wird uebersprungen."
|
||||
log_warn "Eingabe '$langs' wird unveraendert uebernommen."
|
||||
return 0
|
||||
fi
|
||||
installed="$(printf '%s\n' "$raw" | grep -vi '^List of available' || true)"
|
||||
|
||||
IFS='+' read -r -a codes <<< "$langs"
|
||||
for code in "${codes[@]}"; do
|
||||
[ -n "$code" ] || continue
|
||||
if printf '%s\n' "$installed" | grep -qxF "$code"; then
|
||||
log_info "Sprache '$code' ist installiert ✓"
|
||||
continue
|
||||
fi
|
||||
pkg="tesseract-ocr-${code//_/-}"
|
||||
log_warn "Sprache '$code' ist nicht installiert (Paket: $pkg)."
|
||||
read -r -p "Paket '$pkg' jetzt installieren? [J/n]: " answer
|
||||
answer="${answer:-J}"
|
||||
if [[ "$answer" =~ ^[JjYy]$ ]]; then
|
||||
if ! apt-get install -y --no-install-recommends "$pkg"; then
|
||||
log_error "Paket '$pkg' liess sich nicht installieren."
|
||||
elif tesseract --list-langs 2>/dev/null | grep -qxF "$code"; then
|
||||
log_info "Paket '$pkg' installiert ✓"
|
||||
continue
|
||||
else
|
||||
log_error "Paket '$pkg' ist da, aber tesseract kennt '$code' weiterhin nicht."
|
||||
fi
|
||||
fi
|
||||
log_warn "Ohne die Sprachdatei '$code' scheitert das OCR bei JEDER Datei."
|
||||
rc=1
|
||||
done
|
||||
return $rc
|
||||
}
|
||||
|
||||
create_instance() {
|
||||
echo
|
||||
read -r -p "Instanz-Name (nur a-z, 0-9, -): " INST
|
||||
if [[ ! "$INST" =~ ^[a-z0-9][a-z0-9-]*$ ]]; then
|
||||
log_error "Ungültiger Name. Abbruch."
|
||||
return 1
|
||||
fi
|
||||
if [ -f "$CONFIG_DIR/$INST.toml" ]; then
|
||||
log_error "Instanz '$INST' existiert bereits. Abbruch."
|
||||
return 1
|
||||
fi
|
||||
|
||||
log_info "Berechtigungen gesetzt ✓"
|
||||
local default_base="$DATA_ROOT/$INST"
|
||||
read -r -p "Basis-Pfad für Daten [$default_base]: " BASE
|
||||
BASE="${BASE:-$default_base}"
|
||||
|
||||
# ============ 6. systemd-Unit ============
|
||||
log_step "systemd-Unit installieren"
|
||||
read -r -p "Service-User [$DEFAULT_USER]: " SVC_USER
|
||||
SVC_USER="${SVC_USER:-$DEFAULT_USER}"
|
||||
|
||||
sed -e "s|__SERVICE_USER__|$SERVICE_USER|g" \
|
||||
-e "s|__SERVICE_GROUP__|$SERVICE_GROUP|g" \
|
||||
"$REPO_DIR/systemd/pdf-ocr-hotfolder.service" \
|
||||
> "/etc/systemd/system/${SERVICE_NAME}.service"
|
||||
local SVC_GROUP
|
||||
if id "$SVC_USER" &>/dev/null; then
|
||||
SVC_GROUP="$(id -gn "$SVC_USER")"
|
||||
log_info "User '$SVC_USER' existiert (Gruppe: $SVC_GROUP)"
|
||||
else
|
||||
log_warn "User '$SVC_USER' existiert nicht."
|
||||
read -r -p "Lokal als System-User anlegen? [J/n]: " CREATE_USER
|
||||
CREATE_USER="${CREATE_USER:-J}"
|
||||
if [[ "$CREATE_USER" =~ ^[JjYy]$ ]]; then
|
||||
adduser --system --group --home "$BASE" --shell /usr/sbin/nologin "$SVC_USER"
|
||||
SVC_GROUP="$SVC_USER"
|
||||
log_info "User '$SVC_USER' angelegt ✓"
|
||||
else
|
||||
log_error "User muss existieren (z.B. via AD/SSSD). Abbruch."
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# --- OCR-Sprachen ---
|
||||
echo
|
||||
log_info "Tesseract-Sprachen — gelten NUR fuer diese Instanz '$INST'."
|
||||
log_info "Jede zusaetzliche Sprache kostet Laufzeit und verschlechtert zugleich"
|
||||
log_info "die Erkennung — also so eng wie moeglich waehlen (z.B. nur 'deu')."
|
||||
local LANGS
|
||||
while true; do
|
||||
read -r -p "Tesseract-Sprachen [deu+eng]: " LANGS
|
||||
LANGS="${LANGS:-deu+eng}"
|
||||
if [[ ! "$LANGS" =~ ^[a-z]{3}(_[A-Za-z]+)?(\+[a-z]{3}(_[A-Za-z]+)?)*$ ]]; then
|
||||
log_error "Ungueltiges Format. Erwartet: Sprachcodes mit '+' verbunden,"
|
||||
log_error "z.B. 'deu', 'deu+eng' oder 'chi_sim+eng'."
|
||||
continue
|
||||
fi
|
||||
if ensure_tesseract_langs "$LANGS"; then
|
||||
break
|
||||
fi
|
||||
log_warn "Bitte Sprachen erneut angeben (fehlende Sprache einfach weglassen)."
|
||||
echo
|
||||
done
|
||||
|
||||
# --- Original archivieren? ---
|
||||
echo
|
||||
local ORIG_MODE="delete"
|
||||
local ARCHIVE_DIR=""
|
||||
local ARCHIVE_ANS
|
||||
read -r -p "Original nach erfolgreichem OCR archivieren? [j/N]: " ARCHIVE_ANS
|
||||
ARCHIVE_ANS="${ARCHIVE_ANS:-N}"
|
||||
if [[ "$ARCHIVE_ANS" =~ ^[JjYy]$ ]]; then
|
||||
ORIG_MODE="archive"
|
||||
local default_archive="$BASE/archive"
|
||||
while true; do
|
||||
read -r -p "Archiv-Verzeichnis [$default_archive]: " ARCHIVE_DIR
|
||||
ARCHIVE_DIR="${ARCHIVE_DIR:-$default_archive}"
|
||||
if [[ "$ARCHIVE_DIR" != /* ]]; then
|
||||
log_error "Bitte einen absoluten Pfad angeben (beginnt mit '/')."
|
||||
continue
|
||||
fi
|
||||
# Das Archiv darf keines der Arbeitsverzeichnisse sein: im Eingang
|
||||
# wuerde das Original endlos neu aufgegriffen, in den uebrigen
|
||||
# kollidiert es mit der Verarbeitung.
|
||||
case "${ARCHIVE_DIR%/}" in
|
||||
"$BASE/incoming"|"$BASE/outgoing"|"$BASE/working"|"$BASE/error")
|
||||
log_error "Das Archiv darf nicht incoming/outgoing/working/error sein."
|
||||
continue
|
||||
;;
|
||||
esac
|
||||
break
|
||||
done
|
||||
else
|
||||
log_info "Original wird nach erfolgreichem OCR geloescht (original_on_success = \"delete\")."
|
||||
fi
|
||||
|
||||
log_info "Lege Datenverzeichnisse unter $BASE an..."
|
||||
mkdir -p "$BASE"/{incoming,outgoing,working,error}
|
||||
if [ -n "$ARCHIVE_DIR" ]; then
|
||||
mkdir -p "$ARCHIVE_DIR"
|
||||
fi
|
||||
chown -R "$SVC_USER":"$SVC_GROUP" "$BASE"
|
||||
# Innerhalb von $BASE erledigt das chown -R oben schon alles; nur ein Archiv
|
||||
# ausserhalb braucht eigenes mkdir/chown.
|
||||
if [ -n "$ARCHIVE_DIR" ] && [[ "$ARCHIVE_DIR" != "$BASE"/* ]] && [ "$ARCHIVE_DIR" != "$BASE" ]; then
|
||||
chown -R "$SVC_USER":"$SVC_GROUP" "$ARCHIVE_DIR"
|
||||
log_info "Archiv-Verzeichnis $ARCHIVE_DIR angelegt (liegt ausserhalb von $BASE)"
|
||||
fi
|
||||
|
||||
log_info "Erstelle Config $CONFIG_DIR/$INST.toml..."
|
||||
# Verankerte Ausdruecke (Zeilenanfang + Key + '='), damit die deutschen
|
||||
# Kommentarzeilen ueber den Keys unangetastet bleiben.
|
||||
local ESC_BASE ESC_ARCHIVE ESC_LANGS
|
||||
ESC_BASE="$(sed_escape_repl "$BASE")"
|
||||
ESC_ARCHIVE="$(sed_escape_repl "$ARCHIVE_DIR")"
|
||||
ESC_LANGS="$(sed_escape_repl "$LANGS")"
|
||||
sed \
|
||||
-e "s|^incoming[[:space:]]*=.*|incoming = \"$ESC_BASE/incoming\"|" \
|
||||
-e "s|^outgoing[[:space:]]*=.*|outgoing = \"$ESC_BASE/outgoing\"|" \
|
||||
-e "s|^working[[:space:]]*=.*|working = \"$ESC_BASE/working\"|" \
|
||||
-e "s|^error[[:space:]]*=.*|error = \"$ESC_BASE/error\"|" \
|
||||
-e "s|^languages[[:space:]]*=.*|languages = \"$ESC_LANGS\"|" \
|
||||
-e "s|^original_on_success[[:space:]]*=.*|original_on_success = \"$ORIG_MODE\"|" \
|
||||
-e "s|^archive_dir[[:space:]]*=.*|archive_dir = \"$ESC_ARCHIVE\"|" \
|
||||
"$INSTALL_DIR/config.example.toml" > "$CONFIG_DIR/$INST.toml"
|
||||
chown root:"$SVC_GROUP" "$CONFIG_DIR/$INST.toml"
|
||||
chmod 640 "$CONFIG_DIR/$INST.toml"
|
||||
|
||||
# Erzeugte Config gegenpruefen: tragen die drei Keys wirklich die Auswahl?
|
||||
local CFG_OK=1 got key want
|
||||
for key in languages original_on_success archive_dir; do
|
||||
case "$key" in
|
||||
languages) want="$LANGS" ;;
|
||||
original_on_success) want="$ORIG_MODE" ;;
|
||||
archive_dir) want="$ARCHIVE_DIR" ;;
|
||||
esac
|
||||
got="$(config_value "$CONFIG_DIR/$INST.toml" "$key")"
|
||||
if [ "$got" != "$want" ]; then
|
||||
log_error "Config-Pruefung: $key ist \"$got\", erwartet \"$want\""
|
||||
CFG_OK=0
|
||||
fi
|
||||
done
|
||||
if [ "$CFG_OK" -eq 1 ]; then
|
||||
log_info "Config-Pruefung ok ✓ (languages / original_on_success / archive_dir)"
|
||||
else
|
||||
log_warn "Bitte $CONFIG_DIR/$INST.toml von Hand nachziehen."
|
||||
fi
|
||||
|
||||
# Drop-in für abweichenden Service-User
|
||||
if [ "$SVC_USER" != "$DEFAULT_USER" ]; then
|
||||
local DROPIN_DIR="/etc/systemd/system/pdf-ocr-hotfolder@${INST}.service.d"
|
||||
mkdir -p "$DROPIN_DIR"
|
||||
cat > "$DROPIN_DIR/user.conf" <<EOF
|
||||
[Service]
|
||||
User=$SVC_USER
|
||||
Group=$SVC_GROUP
|
||||
EOF
|
||||
log_info "Drop-in für User '$SVC_USER' erstellt"
|
||||
fi
|
||||
|
||||
systemctl daemon-reload
|
||||
systemctl enable "${SERVICE_NAME}.service"
|
||||
systemctl enable --now "pdf-ocr-hotfolder@${INST}.service"
|
||||
sleep 1
|
||||
if systemctl is-active --quiet "pdf-ocr-hotfolder@${INST}.service"; then
|
||||
log_info "✅ Instanz '$INST' läuft"
|
||||
else
|
||||
log_warn "Instanz '$INST' läuft nicht. Logs: journalctl -u pdf-ocr-hotfolder@${INST}"
|
||||
fi
|
||||
|
||||
log_info "systemd-Unit installiert & enabled ✓"
|
||||
echo
|
||||
echo " Config: $CONFIG_DIR/$INST.toml"
|
||||
echo " Eingang: $BASE/incoming"
|
||||
echo " Ausgang: $BASE/outgoing"
|
||||
echo " User: $SVC_USER ($SVC_GROUP)"
|
||||
if [ "$CFG_OK" -eq 1 ]; then
|
||||
echo " Sprachen: $LANGS"
|
||||
if [ "$ORIG_MODE" = "archive" ]; then
|
||||
echo " Archiv: $ARCHIVE_DIR"
|
||||
fi
|
||||
fi
|
||||
echo
|
||||
}
|
||||
|
||||
# ============ 7. Start ============
|
||||
log_step "Service starten"
|
||||
systemctl restart "${SERVICE_NAME}.service"
|
||||
sleep 2
|
||||
systemctl --no-pager --lines=10 status "${SERVICE_NAME}.service" || true
|
||||
# ============================================================
|
||||
# Main
|
||||
# ============================================================
|
||||
|
||||
echo
|
||||
echo "=========================================="
|
||||
echo " Installation abgeschlossen"
|
||||
echo " PDF OCR Hotfolder — Installer"
|
||||
echo "=========================================="
|
||||
|
||||
if [ ! -d "$INSTALL_DIR/venv" ] || [ ! -f "/etc/systemd/system/$SERVICE_TEMPLATE" ]; then
|
||||
log_step "Basis-Installation"
|
||||
install_base
|
||||
else
|
||||
log_info "Basis-Installation bereits vorhanden ($INSTALL_DIR)"
|
||||
log_info "Überspringe Basis-Setup (nutze update.sh für Code-Updates)"
|
||||
fi
|
||||
|
||||
show_existing_instances
|
||||
|
||||
# Erste Instanz ist Pflicht, wenn noch keine vorhanden
|
||||
mapfile -t existing < <(list_instances)
|
||||
if [ "${#existing[@]}" -eq 0 ]; then
|
||||
log_info "Lege erste Hotfolder-Instanz an."
|
||||
create_instance || true
|
||||
fi
|
||||
|
||||
while true; do
|
||||
read -r -p "Weitere Instanz anlegen? [j/N]: " MORE
|
||||
MORE="${MORE:-N}"
|
||||
if [[ "$MORE" =~ ^[JjYy]$ ]]; then
|
||||
create_instance || true
|
||||
else
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
echo
|
||||
echo " Konfiguration: $CONFIG_DIR/config.toml"
|
||||
echo " Eingang: $DATA_DIR/incoming"
|
||||
echo " Ausgang: $DATA_DIR/outgoing"
|
||||
echo " Service-User: $SERVICE_USER ($SERVICE_GROUP)"
|
||||
echo
|
||||
echo " Logs: journalctl -u $SERVICE_NAME -f"
|
||||
echo "=========================================="
|
||||
echo " Fertig"
|
||||
echo "=========================================="
|
||||
show_existing_instances
|
||||
echo " Logs: journalctl -u pdf-ocr-hotfolder@<instanz> -f"
|
||||
echo " Neustart: systemctl restart pdf-ocr-hotfolder@<instanz>"
|
||||
echo " Update: sudo ./update.sh"
|
||||
echo
|
||||
|
||||
@@ -1,3 +1,3 @@
|
||||
"""PDF OCR Hotfolder — Scanner-PDFs automatisch durchsuchbar machen."""
|
||||
|
||||
__version__ = "0.1.0"
|
||||
__version__ = "0.5.0"
|
||||
|
||||
@@ -7,8 +7,8 @@ import sys
|
||||
from pathlib import Path
|
||||
|
||||
from . import __version__
|
||||
from .config import load_config
|
||||
from .service import HotfolderService
|
||||
from .config import ConfigError, load_config
|
||||
from .service import HotfolderService, PreflightError
|
||||
|
||||
|
||||
def _setup_logging(level: str) -> None:
|
||||
@@ -36,18 +36,28 @@ def main() -> int:
|
||||
print(f"Config nicht gefunden: {cfg_path}", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
try:
|
||||
cfg = load_config(cfg_path)
|
||||
except ConfigError as e:
|
||||
print(f"FEHLER: {e}", file=sys.stderr)
|
||||
return 2
|
||||
_setup_logging(cfg.log_level)
|
||||
|
||||
service = HotfolderService(cfg)
|
||||
|
||||
if args.once:
|
||||
service._ensure_dirs() # noqa: SLF001
|
||||
service._scan_existing() # noqa: SLF001
|
||||
service._executor.shutdown(wait=True) # noqa: SLF001
|
||||
return 0
|
||||
try:
|
||||
errors = service.run_once()
|
||||
except PreflightError as e:
|
||||
print(f"FEHLER: {e}", file=sys.stderr)
|
||||
return 2
|
||||
return 1 if errors > 0 else 0
|
||||
|
||||
try:
|
||||
service.run()
|
||||
except PreflightError as e:
|
||||
print(f"FEHLER: {e}", file=sys.stderr)
|
||||
return 2
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
return 0
|
||||
|
||||
@@ -7,6 +7,10 @@ from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
class ConfigError(RuntimeError):
|
||||
"""Konfigurationsdatei ist unvollständig oder fehlerhaft."""
|
||||
|
||||
|
||||
@dataclass
|
||||
class Paths:
|
||||
incoming: Path
|
||||
@@ -21,11 +25,26 @@ class OcrConfig:
|
||||
jobs: int = 4
|
||||
skip_text: bool = True
|
||||
oversample: int = 300
|
||||
pdfa_level: str = "2"
|
||||
# Default bewusst leer: pdfa_level + skip_text zerschießt OCR mit
|
||||
# Ghostscript 10.0.0-10.02.0 (Debian-12-Default), siehe Issue #3
|
||||
pdfa_level: str = ""
|
||||
deskew: bool = True
|
||||
clean: bool = False
|
||||
max_workers: int = 2
|
||||
timeout: int = 1800
|
||||
# Max. Sekunden, die Tesseract pro Seite laufen darf (0 = kein eigenes Limit)
|
||||
timeout: int = 300
|
||||
|
||||
|
||||
@dataclass
|
||||
class OutputConfig:
|
||||
# "prefix" | "suffix" | "none"
|
||||
name_mode: str = "prefix"
|
||||
# Tag-String, verbatim eingefügt (Leerstring = kein Tag)
|
||||
name_tag: str = "OCR_"
|
||||
# "delete" | "archive"
|
||||
original_on_success: str = "delete"
|
||||
# Absoluter Pfad; Pflicht wenn original_on_success == "archive"
|
||||
archive_dir: str = ""
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -79,6 +98,7 @@ class EmailNotify:
|
||||
class Config:
|
||||
paths: Paths
|
||||
ocr: OcrConfig
|
||||
output: OutputConfig
|
||||
verapdf: VeraPdfConfig
|
||||
folder: FolderUpload
|
||||
nextcloud: NextcloudUpload
|
||||
@@ -94,21 +114,46 @@ def _section(data: dict[str, Any], *keys: str) -> dict[str, Any]:
|
||||
return cur if isinstance(cur, dict) else {}
|
||||
|
||||
|
||||
def _require_path(p: dict[str, Any], key: str, cfg_path: Path) -> Path:
|
||||
"""Holt einen Pflicht-Pfad aus der [paths]-Sektion.
|
||||
|
||||
Wirft ConfigError mit klarer Meldung statt eines nackten KeyError.
|
||||
"""
|
||||
value = p.get(key)
|
||||
if value is None or (isinstance(value, str) and not value.strip()):
|
||||
raise ConfigError(
|
||||
f"{cfg_path}: In der Sektion [paths] fehlt der Eintrag '{key}' "
|
||||
f"(oder er ist leer). Bitte ergänzen, z.B. "
|
||||
f'{key} = "/var/lib/pdf-ocr-hotfolder/{key}" '
|
||||
f"— siehe config.example.toml."
|
||||
)
|
||||
return Path(str(value))
|
||||
|
||||
|
||||
def load_config(path: str | Path) -> Config:
|
||||
path = Path(path)
|
||||
with path.open("rb") as f:
|
||||
data = tomllib.load(f)
|
||||
|
||||
if not isinstance(data.get("paths"), dict):
|
||||
raise ConfigError(
|
||||
f"{path}: Die Sektion [paths] fehlt (oder ist keine Tabelle). "
|
||||
"Sie muss die Einträge incoming, outgoing, working und error "
|
||||
"enthalten — siehe config.example.toml."
|
||||
)
|
||||
|
||||
p = _section(data, "paths")
|
||||
paths = Paths(
|
||||
incoming=Path(p["incoming"]),
|
||||
outgoing=Path(p["outgoing"]),
|
||||
working=Path(p["working"]),
|
||||
error=Path(p["error"]),
|
||||
incoming=_require_path(p, "incoming", path),
|
||||
outgoing=_require_path(p, "outgoing", path),
|
||||
working=_require_path(p, "working", path),
|
||||
error=_require_path(p, "error", path),
|
||||
)
|
||||
|
||||
ocr = OcrConfig(**{k: v for k, v in _section(data, "ocr").items()
|
||||
if k in OcrConfig.__annotations__})
|
||||
output = OutputConfig(**{k: v for k, v in _section(data, "output").items()
|
||||
if k in OutputConfig.__annotations__})
|
||||
verapdf = VeraPdfConfig(**{k: v for k, v in _section(data, "verapdf").items()
|
||||
if k in VeraPdfConfig.__annotations__})
|
||||
folder = FolderUpload(**{k: v for k, v in _section(data, "upload", "folder").items()
|
||||
@@ -123,7 +168,7 @@ def load_config(path: str | Path) -> Config:
|
||||
log_level = _section(data, "logging").get("level", "INFO")
|
||||
|
||||
return Config(
|
||||
paths=paths, ocr=ocr, verapdf=verapdf,
|
||||
paths=paths, ocr=ocr, output=output, verapdf=verapdf,
|
||||
folder=folder, nextcloud=nextcloud, sftp=sftp, email=email,
|
||||
log_level=log_level,
|
||||
)
|
||||
|
||||
@@ -7,12 +7,39 @@ import subprocess
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
import ocrmypdf
|
||||
|
||||
from .config import OcrConfig, VeraPdfConfig
|
||||
from .config import OcrConfig, OutputConfig, VeraPdfConfig
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
# Erlaubte Werte für [output].name_mode — wird auch vom Preflight geprüft
|
||||
VALID_NAME_MODES = ("prefix", "suffix", "none")
|
||||
|
||||
|
||||
def build_output_name(src_name: str, mode: str, tag: str) -> str:
|
||||
"""Erzeugt den Ziel-Dateinamen für ein OCR-PDF.
|
||||
|
||||
Args:
|
||||
src_name: Original-Dateiname (z.B. "scan.pdf")
|
||||
mode: "prefix" | "suffix" | "none"
|
||||
tag: Einzufügender String (verbatim, leer = kein Tag)
|
||||
|
||||
Beispiele:
|
||||
prefix "OCR_": "scan.pdf" -> "OCR_scan.pdf"
|
||||
suffix "_OCR": "scan.pdf" -> "scan_OCR.pdf"
|
||||
suffix "_OCR": "scan.tar.gz.pdf" -> "scan.tar.gz_OCR.pdf"
|
||||
none: "scan.pdf" -> "scan.pdf"
|
||||
"""
|
||||
if mode == "none" or not tag:
|
||||
return src_name
|
||||
if mode == "prefix":
|
||||
return f"{tag}{src_name}"
|
||||
if mode == "suffix":
|
||||
# Nur die letzte Extension abspalten, sonst "foo.bar.pdf" kaputt gemacht
|
||||
p = Path(src_name)
|
||||
stem, ext = p.stem, p.suffix
|
||||
return f"{stem}{tag}{ext}"
|
||||
raise ValueError(f"Unbekannter name_mode: {mode!r}")
|
||||
|
||||
|
||||
@dataclass
|
||||
class ProcessResult:
|
||||
@@ -25,6 +52,8 @@ class ProcessResult:
|
||||
|
||||
def run_ocr(src: Path, dst: Path, cfg: OcrConfig) -> None:
|
||||
"""Führt ocrmypdf als Library-Call aus (kein Subprozess-Overhead)."""
|
||||
import ocrmypdf # lazy, damit Tests ohne ocrmypdf laufen
|
||||
|
||||
kwargs: dict = {
|
||||
"language": cfg.languages,
|
||||
"jobs": cfg.jobs,
|
||||
@@ -39,6 +68,15 @@ def run_ocr(src: Path, dst: Path, cfg: OcrConfig) -> None:
|
||||
else:
|
||||
kwargs["output_type"] = "pdf"
|
||||
|
||||
# [ocr].timeout = max. Sekunden, die Tesseract pro Seite laufen darf.
|
||||
# ocrmypdf kennt kein Gesamt-Timeout für ein Dokument, nur `tesseract_timeout`
|
||||
# (pro Seite). ACHTUNG: ocrmypdf interpretiert tesseract_timeout=0 als
|
||||
# "OCR komplett überspringen" — deshalb wird 0 bei uns als "kein eigenes
|
||||
# Limit" behandelt und gar nicht erst durchgereicht (dann gilt der
|
||||
# ocrmypdf-Default).
|
||||
if cfg.timeout and cfg.timeout > 0:
|
||||
kwargs["tesseract_timeout"] = float(cfg.timeout)
|
||||
|
||||
log.info("OCR start: %s", src.name)
|
||||
ocrmypdf.ocr(str(src), str(dst), **kwargs)
|
||||
log.info("OCR done: %s", dst.name)
|
||||
@@ -71,11 +109,13 @@ def process_pdf(
|
||||
error_dir: Path,
|
||||
ocr_cfg: OcrConfig,
|
||||
vera_cfg: VeraPdfConfig,
|
||||
output_cfg: OutputConfig,
|
||||
) -> ProcessResult:
|
||||
"""Verarbeitet eine einzelne PDF: move→OCR→validate→outgoing/error."""
|
||||
out_name = build_output_name(src.name, output_cfg.name_mode, output_cfg.name_tag)
|
||||
work_src = working_dir / src.name
|
||||
work_out = working_dir / f"OCR_{src.name}"
|
||||
final_out = outgoing_dir / f"OCR_{src.name}"
|
||||
work_out = working_dir / f"__ocr_{out_name}" # Temp-Name, damit er != src.name ist
|
||||
final_out = outgoing_dir / out_name
|
||||
|
||||
try:
|
||||
shutil.move(str(src), str(work_src))
|
||||
@@ -93,17 +133,58 @@ def process_pdf(
|
||||
if vera_cfg.enabled:
|
||||
vera_ok = run_verapdf(work_out, vera_cfg)
|
||||
if not vera_ok:
|
||||
# Das OCR-Ergebnis ist unbrauchbar und wandert nach error/. Das
|
||||
# Original wird aber NICHT bedingungslos gelöscht: es folgt derselben
|
||||
# [output].original_on_success-Regel wie im Erfolgsfall, sonst
|
||||
# verliert man es ausgerechnet im Fehlerfall (archive!).
|
||||
_move_to_error(work_out, error_dir)
|
||||
work_src.unlink(missing_ok=True)
|
||||
_dispose_original(work_src, src.name, output_cfg)
|
||||
log.error(
|
||||
"veraPDF FAIL: %s — OCR-Ergebnis nach %s verschoben, Original %s",
|
||||
src.name, error_dir,
|
||||
"archiviert" if output_cfg.original_on_success == "archive" else "gelöscht",
|
||||
)
|
||||
return ProcessResult(src, final_out, False,
|
||||
"verapdf validation failed", verapdf_passed=False)
|
||||
|
||||
outgoing_dir.mkdir(parents=True, exist_ok=True)
|
||||
shutil.move(str(work_out), str(final_out))
|
||||
work_src.unlink(missing_ok=True)
|
||||
_dispose_original(work_src, src.name, output_cfg)
|
||||
return ProcessResult(src, final_out, True, verapdf_passed=vera_ok)
|
||||
|
||||
|
||||
def _dispose_original(work_src: Path, original_name: str, cfg: OutputConfig) -> None:
|
||||
"""Entsorgt das Original laut [output].original_on_success — löschen oder archivieren.
|
||||
|
||||
Wird nach erfolgreichem OCR aufgerufen und ebenso, wenn veraPDF die
|
||||
Validierung ablehnt: auch dann soll `archive` das Original erhalten.
|
||||
"""
|
||||
if not work_src.exists():
|
||||
return
|
||||
mode = cfg.original_on_success
|
||||
if mode == "delete":
|
||||
work_src.unlink(missing_ok=True)
|
||||
return
|
||||
if mode == "archive":
|
||||
if not cfg.archive_dir:
|
||||
log.error("original_on_success=archive aber archive_dir ist leer — lösche stattdessen")
|
||||
work_src.unlink(missing_ok=True)
|
||||
return
|
||||
archive = Path(cfg.archive_dir)
|
||||
archive.mkdir(parents=True, exist_ok=True)
|
||||
dest = archive / original_name
|
||||
# Bei Namens-Kollision mit Timestamp umbenennen
|
||||
if dest.exists():
|
||||
from datetime import datetime
|
||||
ts = datetime.now().strftime("%Y%m%d-%H%M%S")
|
||||
dest = archive / f"{dest.stem}_{ts}{dest.suffix}"
|
||||
shutil.move(str(work_src), str(dest))
|
||||
log.info("Original archiviert: %s", dest)
|
||||
return
|
||||
log.warning("Unbekannter original_on_success=%r — lösche stattdessen", mode)
|
||||
work_src.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def _move_to_error(p: Path, error_dir: Path) -> None:
|
||||
error_dir.mkdir(parents=True, exist_ok=True)
|
||||
try:
|
||||
|
||||
+235
-12
@@ -2,7 +2,10 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import re
|
||||
import shutil
|
||||
import signal
|
||||
import subprocess
|
||||
import threading
|
||||
import time
|
||||
from concurrent.futures import Future, ThreadPoolExecutor
|
||||
@@ -12,12 +15,112 @@ from watchdog.events import FileSystemEvent, FileSystemEventHandler
|
||||
from watchdog.observers import Observer
|
||||
|
||||
from .config import Config
|
||||
from .processor import ProcessResult, process_pdf
|
||||
from .processor import VALID_NAME_MODES, ProcessResult, _move_to_error, process_pdf
|
||||
from .uploaders import notify_email, upload_folder, upload_nextcloud, upload_sftp
|
||||
|
||||
log = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class PreflightError(RuntimeError):
|
||||
"""Erforderliche externe Binaries fehlen."""
|
||||
|
||||
|
||||
# Pflicht-Binaries für ocrmypdf
|
||||
_REQUIRED_BINARIES = ("tesseract", "gs")
|
||||
|
||||
# Ghostscript-Versionen mit bekanntem PDF/A+skip_text Bug (Issue #3):
|
||||
# 10.0.0 .. 10.02.0 (inklusive). Ab 10.02.1 wieder nutzbar.
|
||||
_GS_BROKEN_MIN = (10, 0, 0)
|
||||
_GS_BROKEN_MAX = (10, 2, 0)
|
||||
|
||||
|
||||
def _parse_version(text: str) -> tuple[int, ...] | None:
|
||||
"""Extrahiert die erste X.Y[.Z] Version aus einem String."""
|
||||
m = re.search(r"(\d+)\.(\d+)(?:\.(\d+))?", text)
|
||||
if not m:
|
||||
return None
|
||||
return tuple(int(x) if x is not None else 0 for x in m.groups())
|
||||
|
||||
|
||||
def is_ghostscript_broken(version: str | None) -> bool:
|
||||
"""Prüft, ob eine Ghostscript-Version vom PDF/A+skip_text Bug betroffen ist.
|
||||
|
||||
Betrifft 10.0.0 bis einschließlich 10.02.0. Ab 10.02.1 wieder sicher.
|
||||
"""
|
||||
if not version:
|
||||
return False
|
||||
parsed = _parse_version(version)
|
||||
if parsed is None:
|
||||
return False
|
||||
# Auf 3-Tupel normalisieren
|
||||
while len(parsed) < 3:
|
||||
parsed = parsed + (0,)
|
||||
parsed = parsed[:3]
|
||||
return _GS_BROKEN_MIN <= parsed <= _GS_BROKEN_MAX
|
||||
|
||||
|
||||
def detect_ghostscript_version() -> str | None:
|
||||
"""Ruft `gs --version` auf und gibt den Versionsstring zurück (oder None)."""
|
||||
gs = shutil.which("gs")
|
||||
if gs is None:
|
||||
return None
|
||||
try:
|
||||
result = subprocess.run([gs, "--version"], capture_output=True,
|
||||
text=True, timeout=5)
|
||||
except (OSError, subprocess.TimeoutExpired):
|
||||
return None
|
||||
return result.stdout.strip() or None
|
||||
|
||||
|
||||
def check_output_config(mode: str, archive_dir: str,
|
||||
name_mode: str = "prefix") -> None:
|
||||
"""Validiert die [output]-Section. Wirft PreflightError bei Problemen."""
|
||||
valid_modes = {"delete", "archive"}
|
||||
if mode not in valid_modes:
|
||||
raise PreflightError(
|
||||
f"[output].original_on_success={mode!r} ungültig. "
|
||||
f"Erlaubt: {sorted(valid_modes)}"
|
||||
)
|
||||
if mode == "archive" and not archive_dir:
|
||||
raise PreflightError(
|
||||
"[output].original_on_success='archive' erfordert [output].archive_dir"
|
||||
)
|
||||
# Früh prüfen: sonst schlägt ein Tippfehler erst pro Datei zu — und zwar
|
||||
# NACH dem Move nach working/, wo die Datei dann liegen bleibt.
|
||||
if name_mode not in VALID_NAME_MODES:
|
||||
raise PreflightError(
|
||||
f"[output].name_mode={name_mode!r} ungültig. "
|
||||
f"Erlaubt: {sorted(VALID_NAME_MODES)}"
|
||||
)
|
||||
|
||||
|
||||
def check_preflight(pdfa_level: str = "") -> None:
|
||||
"""Prüft externe Abhängigkeiten.
|
||||
|
||||
- Tesseract und Ghostscript müssen im PATH sein
|
||||
- Bei gesetztem pdfa_level wird die Ghostscript-Version gegen den
|
||||
bekannten 10.0.0–10.02.0 Bug geprüft
|
||||
|
||||
Wirft PreflightError bei fehlenden Binaries oder unsicherem Ghostscript.
|
||||
"""
|
||||
missing = [b for b in _REQUIRED_BINARIES if shutil.which(b) is None]
|
||||
if missing:
|
||||
raise PreflightError(
|
||||
"Fehlende Abhängigkeiten: " + ", ".join(missing)
|
||||
+ ". Bitte installieren: sudo apt install tesseract-ocr ghostscript"
|
||||
)
|
||||
|
||||
if pdfa_level:
|
||||
gs_version = detect_ghostscript_version()
|
||||
if is_ghostscript_broken(gs_version):
|
||||
raise PreflightError(
|
||||
f"Ghostscript {gs_version} ist mit pdfa_level='{pdfa_level}' nicht "
|
||||
"kompatibel (bekannter Bug in 10.0.0–10.02.0). "
|
||||
"Entweder ghostscript auf >=10.02.1 upgraden (z.B. via bookworm-backports) "
|
||||
"oder in der Config [ocr].pdfa_level = \"\" setzen."
|
||||
)
|
||||
|
||||
|
||||
def _is_pdf(path: Path) -> bool:
|
||||
return path.suffix.lower() == ".pdf" and path.is_file()
|
||||
|
||||
@@ -70,10 +173,20 @@ class HotfolderService:
|
||||
self._stop = threading.Event()
|
||||
self._inflight: set[str] = set()
|
||||
self._lock = threading.Lock()
|
||||
self._success_count = 0
|
||||
self._error_count = 0
|
||||
|
||||
@property
|
||||
def success_count(self) -> int:
|
||||
return self._success_count
|
||||
|
||||
@property
|
||||
def error_count(self) -> int:
|
||||
return self._error_count
|
||||
|
||||
# ---- Setup ----
|
||||
|
||||
def _ensure_dirs(self) -> None:
|
||||
def ensure_dirs(self) -> None:
|
||||
for p in (self.cfg.paths.incoming, self.cfg.paths.outgoing,
|
||||
self.cfg.paths.working, self.cfg.paths.error):
|
||||
p.mkdir(parents=True, exist_ok=True)
|
||||
@@ -81,7 +194,11 @@ class HotfolderService:
|
||||
# ---- Lifecycle ----
|
||||
|
||||
def run(self) -> None:
|
||||
self._ensure_dirs()
|
||||
check_preflight(self.cfg.ocr.pdfa_level)
|
||||
check_output_config(self.cfg.output.original_on_success,
|
||||
self.cfg.output.archive_dir,
|
||||
self.cfg.output.name_mode)
|
||||
self.ensure_dirs()
|
||||
self._scan_existing()
|
||||
|
||||
self._observer = Observer()
|
||||
@@ -98,6 +215,23 @@ class HotfolderService:
|
||||
finally:
|
||||
self.shutdown()
|
||||
|
||||
def run_once(self) -> int:
|
||||
"""Verarbeitet alle bereits im incoming-Ordner liegenden PDFs und beendet sich.
|
||||
|
||||
Returns:
|
||||
Anzahl fehlgeschlagener PDFs (0 = alles ok).
|
||||
"""
|
||||
check_preflight(self.cfg.ocr.pdfa_level)
|
||||
check_output_config(self.cfg.output.original_on_success,
|
||||
self.cfg.output.archive_dir,
|
||||
self.cfg.output.name_mode)
|
||||
self.ensure_dirs()
|
||||
self._scan_existing()
|
||||
self._executor.shutdown(wait=True)
|
||||
log.info("One-shot fertig: %d ok, %d Fehler",
|
||||
self._success_count, self._error_count)
|
||||
return self._error_count
|
||||
|
||||
def shutdown(self) -> None:
|
||||
log.info("Shutdown läuft...")
|
||||
if self._observer:
|
||||
@@ -134,13 +268,33 @@ class HotfolderService:
|
||||
|
||||
# ---- Processing ----
|
||||
|
||||
def _count_success(self) -> None:
|
||||
with self._lock:
|
||||
self._success_count += 1
|
||||
|
||||
def _count_error(self) -> None:
|
||||
with self._lock:
|
||||
self._error_count += 1
|
||||
|
||||
def _process(self, path: Path) -> None:
|
||||
if not _wait_until_stable(path):
|
||||
log.warning("Datei nicht stabilisiert, überspringe: %s", path)
|
||||
if not path.exists():
|
||||
# Datei wurde währenddessen entfernt — kein Fehlerfall
|
||||
log.info("Datei vor der Verarbeitung verschwunden: %s", path)
|
||||
return
|
||||
# Bewusst als Fehler zählen: sonst liefert --once trotz liegen
|
||||
# gebliebener Datei Exit 0.
|
||||
log.error(
|
||||
"Datei hat sich nicht stabilisiert (Timeout): %s — bleibt in %s "
|
||||
"liegen und wird beim nächsten Lauf erneut versucht",
|
||||
path, self.cfg.paths.incoming,
|
||||
)
|
||||
self._count_error()
|
||||
return
|
||||
if not path.exists():
|
||||
return
|
||||
|
||||
try:
|
||||
result: ProcessResult = process_pdf(
|
||||
src=path,
|
||||
working_dir=self.cfg.paths.working,
|
||||
@@ -148,18 +302,87 @@ class HotfolderService:
|
||||
error_dir=self.cfg.paths.error,
|
||||
ocr_cfg=self.cfg.ocr,
|
||||
vera_cfg=self.cfg.verapdf,
|
||||
output_cfg=self.cfg.output,
|
||||
)
|
||||
except Exception as e: # noqa: BLE001 - kein Fehler darf die Zählung umgehen
|
||||
log.exception("Unerwarteter Fehler bei der Verarbeitung von %s", path.name)
|
||||
self._count_error()
|
||||
self._rescue_to_error(path)
|
||||
self._notify(ProcessResult(
|
||||
path, self.cfg.paths.outgoing / path.name, False,
|
||||
f"unerwarteter Fehler: {e}",
|
||||
))
|
||||
return
|
||||
|
||||
if result.success:
|
||||
self._dispatch_uploads(result.output)
|
||||
if not result.success:
|
||||
self._count_error()
|
||||
self._notify(result)
|
||||
return
|
||||
|
||||
failed = self._dispatch_uploads(result.output)
|
||||
if failed:
|
||||
log.error(
|
||||
"Upload fehlgeschlagen (%s) für %s — das OCR selbst war "
|
||||
"erfolgreich, die Datei bleibt daher in %s liegen und wird "
|
||||
"NICHT nach error/ verschoben",
|
||||
", ".join(failed), result.output.name, result.output.parent,
|
||||
)
|
||||
self._count_error()
|
||||
self._notify_upload_failure(result, failed)
|
||||
return
|
||||
|
||||
self._count_success()
|
||||
self._notify(result)
|
||||
|
||||
def _dispatch_uploads(self, pdf: Path) -> None:
|
||||
upload_folder(pdf, self.cfg.folder, self.cfg.paths.outgoing)
|
||||
if self.cfg.nextcloud.enabled:
|
||||
upload_nextcloud(pdf, self.cfg.nextcloud)
|
||||
if self.cfg.sftp.enabled:
|
||||
upload_sftp(pdf, self.cfg.sftp)
|
||||
def _rescue_to_error(self, src: Path) -> None:
|
||||
"""Bringt eine Datei nach einer unerwarteten Exception ins error-Verzeichnis.
|
||||
|
||||
Die Datei kann je nach Abbruchzeitpunkt noch in incoming/ oder schon in
|
||||
working/ liegen. Der erste Treffer wird verschoben (keine Doppel-Moves),
|
||||
Fehler beim Verschieben werden nur geloggt.
|
||||
"""
|
||||
error_dir = self.cfg.paths.error
|
||||
for candidate in (src, self.cfg.paths.working / src.name):
|
||||
try:
|
||||
if not candidate.is_file():
|
||||
continue
|
||||
if candidate.parent.resolve() == error_dir.resolve():
|
||||
return # liegt bereits im error-Verzeichnis
|
||||
except OSError:
|
||||
continue
|
||||
_move_to_error(candidate, error_dir)
|
||||
return
|
||||
log.warning("Datei %s nach Fehler nicht mehr auffindbar — "
|
||||
"kein Verschieben nach error/ möglich", src.name)
|
||||
|
||||
def _dispatch_uploads(self, pdf: Path) -> list[str]:
|
||||
"""Schiebt das fertige PDF an alle Upload-Ziele.
|
||||
|
||||
Die uploader prüfen `cfg.enabled` jeweils selbst und liefern für
|
||||
deaktivierte Ziele True.
|
||||
|
||||
Returns:
|
||||
Namen der fehlgeschlagenen Ziele — leere Liste = alle erfolgreich.
|
||||
"""
|
||||
failed: list[str] = []
|
||||
if not upload_folder(pdf, self.cfg.folder, self.cfg.paths.outgoing):
|
||||
failed.append("folder")
|
||||
if not upload_nextcloud(pdf, self.cfg.nextcloud):
|
||||
failed.append("nextcloud")
|
||||
if not upload_sftp(pdf, self.cfg.sftp):
|
||||
failed.append("sftp")
|
||||
return failed
|
||||
|
||||
def _notify_upload_failure(self, result: ProcessResult, failed: list[str]) -> None:
|
||||
"""Fehler-Mail, wenn das OCR lief, aber mindestens ein Upload scheiterte."""
|
||||
subject = f"[pdf-ocr] FEHLER Upload: {result.source.name}"
|
||||
body = (
|
||||
f"OCR erfolgreich: {result.output}\n\n"
|
||||
f"Fehlgeschlagene Upload-Ziele: {', '.join(failed)}\n\n"
|
||||
f"Das OCR-PDF bleibt in {result.output.parent} liegen und wurde "
|
||||
"NICHT nach error/ verschoben. Details siehe Log.\n"
|
||||
)
|
||||
notify_email(self.cfg.email, subject, body, False)
|
||||
|
||||
def _notify(self, result: ProcessResult) -> None:
|
||||
if result.success:
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import shutil
|
||||
import smtplib
|
||||
import ssl
|
||||
from email.message import EmailMessage
|
||||
@@ -25,7 +26,9 @@ def upload_folder(pdf: Path, cfg: FolderUpload, default_target: Path) -> bool:
|
||||
try:
|
||||
if pdf.resolve() == dest.resolve():
|
||||
return True
|
||||
dest.write_bytes(pdf.read_bytes())
|
||||
# copyfile statt read_bytes/write_bytes: große PDFs nicht komplett
|
||||
# in den Speicher laden
|
||||
shutil.copyfile(pdf, dest)
|
||||
log.info("Folder upload OK: %s", dest)
|
||||
return True
|
||||
except OSError as e:
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
[pytest]
|
||||
testpaths = tests
|
||||
@@ -0,0 +1,10 @@
|
||||
# Drop-in für LXC/Container-Betrieb
|
||||
# Kopieren nach: /etc/systemd/system/pdf-ocr-hotfolder@.service.d/lxc-compat.conf
|
||||
# Danach: systemctl daemon-reload && systemctl restart 'pdf-ocr-hotfolder@*'
|
||||
|
||||
[Service]
|
||||
PrivateTmp=false
|
||||
ProtectSystem=false
|
||||
ProtectKernelTunables=false
|
||||
ProtectKernelModules=false
|
||||
ProtectControlGroups=false
|
||||
@@ -1,13 +1,14 @@
|
||||
[Unit]
|
||||
Description=PDF OCR Hotfolder
|
||||
Description=PDF OCR Hotfolder (Instance: %i)
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
User=__SERVICE_USER__
|
||||
Group=__SERVICE_GROUP__
|
||||
ExecStart=/opt/pdf-ocr-hotfolder/venv/bin/python -m pdf_ocr_hotfolder --config /etc/pdf-ocr-hotfolder/config.toml
|
||||
User=pdfocr
|
||||
Group=pdfocr
|
||||
WorkingDirectory=/opt/pdf-ocr-hotfolder
|
||||
ExecStart=/opt/pdf-ocr-hotfolder/venv/bin/python -m pdf_ocr_hotfolder --config /etc/pdf-ocr-hotfolder/%i.toml
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
KillMode=mixed
|
||||
@@ -0,0 +1,54 @@
|
||||
"""Gemeinsame pytest-Fixtures."""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.config import (
|
||||
Config,
|
||||
EmailNotify,
|
||||
FolderUpload,
|
||||
NextcloudUpload,
|
||||
OcrConfig,
|
||||
OutputConfig,
|
||||
Paths,
|
||||
SftpUpload,
|
||||
VeraPdfConfig,
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def tmp_config(tmp_path: Path) -> Config:
|
||||
"""Minimal-Config mit tmp_path-Verzeichnissen, alle Uploads deaktiviert."""
|
||||
paths = Paths(
|
||||
incoming=tmp_path / "incoming",
|
||||
outgoing=tmp_path / "outgoing",
|
||||
working=tmp_path / "working",
|
||||
error=tmp_path / "error",
|
||||
)
|
||||
for p in (paths.incoming, paths.outgoing, paths.working, paths.error):
|
||||
p.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
return Config(
|
||||
paths=paths,
|
||||
ocr=OcrConfig(max_workers=1),
|
||||
output=OutputConfig(),
|
||||
verapdf=VeraPdfConfig(enabled=False),
|
||||
folder=FolderUpload(enabled=False),
|
||||
nextcloud=NextcloudUpload(enabled=False),
|
||||
sftp=SftpUpload(enabled=False),
|
||||
email=EmailNotify(enabled=False),
|
||||
log_level="DEBUG",
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def dummy_pdf(tmp_config: Config) -> Path:
|
||||
"""Legt eine Datei mit .pdf-Extension im incoming-Ordner ab.
|
||||
|
||||
Achtung: kein echtes PDF. Für Tests wird `process_pdf` gemockt.
|
||||
"""
|
||||
pdf = tmp_config.paths.incoming / "test.pdf"
|
||||
pdf.write_bytes(b"%PDF-1.4 fake\n")
|
||||
return pdf
|
||||
@@ -0,0 +1,79 @@
|
||||
"""Tests für verständliche Fehlermeldungen beim Laden der Config."""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.config import ConfigError, load_config
|
||||
|
||||
_FULL_PATHS = """
|
||||
[paths]
|
||||
incoming = "/tmp/in"
|
||||
outgoing = "/tmp/out"
|
||||
working = "/tmp/work"
|
||||
error = "/tmp/err"
|
||||
"""
|
||||
|
||||
|
||||
def _write(tmp_path: Path, content: str) -> Path:
|
||||
cfg = tmp_path / "config.toml"
|
||||
cfg.write_text(content)
|
||||
return cfg
|
||||
|
||||
|
||||
def test_missing_paths_section(tmp_path: Path) -> None:
|
||||
"""Fehlt [paths] komplett → ConfigError statt KeyError."""
|
||||
cfg = _write(tmp_path, '[ocr]\nlanguages = "deu"\n')
|
||||
with pytest.raises(ConfigError) as exc:
|
||||
load_config(cfg)
|
||||
msg = str(exc.value)
|
||||
assert "[paths]" in msg
|
||||
assert str(cfg) in msg
|
||||
|
||||
|
||||
def test_empty_config_file(tmp_path: Path) -> None:
|
||||
cfg = _write(tmp_path, "")
|
||||
with pytest.raises(ConfigError, match=r"\[paths\]"):
|
||||
load_config(cfg)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("missing", ["incoming", "outgoing", "working", "error"])
|
||||
def test_missing_single_path_key(tmp_path: Path, missing: str) -> None:
|
||||
"""Fehlt ein einzelner Key, wird genau dieser genannt."""
|
||||
lines = [line for line in _FULL_PATHS.strip().splitlines()
|
||||
if not line.startswith(missing)]
|
||||
cfg = _write(tmp_path, "\n".join(lines) + "\n")
|
||||
with pytest.raises(ConfigError) as exc:
|
||||
load_config(cfg)
|
||||
msg = str(exc.value)
|
||||
assert missing in msg
|
||||
assert str(cfg) in msg
|
||||
|
||||
|
||||
def test_empty_path_value_is_rejected(tmp_path: Path) -> None:
|
||||
"""Ein leerer Pfad ist genauso falsch wie ein fehlender."""
|
||||
cfg = _write(tmp_path, _FULL_PATHS.replace('working = "/tmp/work"',
|
||||
'working = ""'))
|
||||
with pytest.raises(ConfigError, match="working"):
|
||||
load_config(cfg)
|
||||
|
||||
|
||||
def test_complete_paths_section_loads(tmp_path: Path) -> None:
|
||||
cfg = _write(tmp_path, _FULL_PATHS)
|
||||
loaded = load_config(cfg)
|
||||
assert loaded.paths.incoming == Path("/tmp/in")
|
||||
assert loaded.paths.error == Path("/tmp/err")
|
||||
|
||||
|
||||
def test_main_returns_2_on_broken_config(tmp_path: Path, monkeypatch, capsys) -> None:
|
||||
"""CLI bricht sauber mit Exit-Code 2 ab — ohne Traceback."""
|
||||
cfg = _write(tmp_path, '[ocr]\nlanguages = "deu"\n')
|
||||
monkeypatch.setattr(sys, "argv",
|
||||
["pdf-ocr-hotfolder", "--config", str(cfg), "--once"])
|
||||
from pdf_ocr_hotfolder.__main__ import main
|
||||
assert main() == 2
|
||||
err = capsys.readouterr().err
|
||||
assert "FEHLER" in err
|
||||
assert "[paths]" in err
|
||||
@@ -0,0 +1,257 @@
|
||||
"""Tests für die Fehlerzählung im Service.
|
||||
|
||||
Deckt drei bisher stumme Fehlerpfade ab:
|
||||
- Exception aus `process_pdf()` (z.B. fehlgeschlagener Move nach outgoing/)
|
||||
- fehlgeschlagene Uploads
|
||||
- Timeout im Stabilitäts-Check
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from pdf_ocr_hotfolder.processor import ProcessResult
|
||||
from pdf_ocr_hotfolder.service import HotfolderService
|
||||
|
||||
|
||||
def _run_once(tmp_config, **patches):
|
||||
"""Führt run_once() mit gemocktem Preflight aus und gibt den Service zurück."""
|
||||
stack = [
|
||||
patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None),
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable",
|
||||
return_value=patches.pop("stable", True)),
|
||||
]
|
||||
for target, kwargs in patches.items():
|
||||
stack.append(patch(f"pdf_ocr_hotfolder.service.{target}", **kwargs))
|
||||
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
for p in stack:
|
||||
p.start()
|
||||
service.run_once()
|
||||
finally:
|
||||
for p in reversed(stack):
|
||||
p.stop()
|
||||
service._executor.shutdown(wait=False)
|
||||
return service
|
||||
|
||||
|
||||
# ---------------- Exception aus process_pdf ----------------
|
||||
|
||||
def test_exception_from_process_pdf_counts_as_error(tmp_config) -> None:
|
||||
"""Eine Exception aus process_pdf() darf die Zählung nicht umgehen."""
|
||||
(tmp_config.paths.incoming / "boom.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def explode(src, **kwargs):
|
||||
raise OSError("move to outgoing failed")
|
||||
|
||||
service = _run_once(tmp_config, process_pdf={"side_effect": explode})
|
||||
|
||||
assert service.error_count == 1
|
||||
assert service.success_count == 0
|
||||
|
||||
|
||||
def test_exception_moves_file_to_error_dir(tmp_config) -> None:
|
||||
"""Die Datei landet nach einer Exception im error-Verzeichnis."""
|
||||
(tmp_config.paths.incoming / "boom.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def explode(src, **kwargs):
|
||||
raise RuntimeError("kaputt")
|
||||
|
||||
_run_once(tmp_config, process_pdf={"side_effect": explode})
|
||||
|
||||
assert (tmp_config.paths.error / "boom.pdf").exists()
|
||||
assert not (tmp_config.paths.incoming / "boom.pdf").exists()
|
||||
|
||||
|
||||
def test_exception_after_move_to_working_rescues_from_working(tmp_config) -> None:
|
||||
"""Realistischer Fall: process_pdf hat schon nach working/ verschoben."""
|
||||
src = tmp_config.paths.incoming / "boom.pdf"
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def explode(src: Path, working_dir: Path, **kwargs):
|
||||
# process_pdf verschiebt zuerst nach working/, dann knallt der Move
|
||||
# nach outgoing/
|
||||
src.rename(working_dir / src.name)
|
||||
raise OSError("move to outgoing failed")
|
||||
|
||||
service = _run_once(tmp_config, process_pdf={"side_effect": explode})
|
||||
|
||||
assert service.error_count == 1
|
||||
assert (tmp_config.paths.error / "boom.pdf").exists()
|
||||
assert not (tmp_config.paths.working / "boom.pdf").exists()
|
||||
|
||||
|
||||
def test_exception_with_vanished_file_does_not_raise(tmp_config) -> None:
|
||||
"""Ist die Datei nicht mehr auffindbar, wird nur geloggt — kein Crash."""
|
||||
(tmp_config.paths.incoming / "boom.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def explode(src: Path, **kwargs):
|
||||
src.unlink(missing_ok=True)
|
||||
raise RuntimeError("kaputt")
|
||||
|
||||
service = _run_once(tmp_config, process_pdf={"side_effect": explode})
|
||||
|
||||
assert service.error_count == 1
|
||||
assert not (tmp_config.paths.error / "boom.pdf").exists()
|
||||
|
||||
|
||||
def test_exception_triggers_error_notification(tmp_config) -> None:
|
||||
"""Auch bei einer Exception geht eine Fehler-Mail raus (success=False)."""
|
||||
(tmp_config.paths.incoming / "boom.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def explode(src, **kwargs):
|
||||
raise RuntimeError("kaputt")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.notify_email") as mail:
|
||||
_run_once(tmp_config, process_pdf={"side_effect": explode})
|
||||
|
||||
assert mail.call_count == 1
|
||||
args = mail.call_args[0]
|
||||
assert "FEHLER" in args[1]
|
||||
assert args[3] is False # success-Flag
|
||||
|
||||
|
||||
# ---------------- Upload-Fehler ----------------
|
||||
|
||||
def _fake_success(src: Path, working_dir, outgoing_dir, error_dir, **kwargs):
|
||||
out = outgoing_dir / f"OCR_{src.name}"
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_bytes(b"%PDF-1.4 ocr\n")
|
||||
src.unlink(missing_ok=True)
|
||||
return ProcessResult(src, out, True)
|
||||
|
||||
|
||||
def test_failed_upload_counts_as_error(tmp_config) -> None:
|
||||
"""Ein fehlgeschlagener Upload zählt als Fehler, nicht als Erfolg."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
service = _run_once(
|
||||
tmp_config,
|
||||
process_pdf={"side_effect": _fake_success},
|
||||
upload_nextcloud={"return_value": False},
|
||||
)
|
||||
|
||||
assert service.error_count == 1
|
||||
assert service.success_count == 0
|
||||
|
||||
|
||||
def test_failed_upload_sends_error_mail_naming_targets(tmp_config) -> None:
|
||||
"""Die Fehler-Mail nennt die fehlgeschlagenen Ziele."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.notify_email") as mail:
|
||||
_run_once(
|
||||
tmp_config,
|
||||
process_pdf={"side_effect": _fake_success},
|
||||
upload_nextcloud={"return_value": False},
|
||||
upload_sftp={"return_value": False},
|
||||
)
|
||||
|
||||
assert mail.call_count == 1
|
||||
_cfg, subject, body, success = mail.call_args[0]
|
||||
assert "FEHLER" in subject
|
||||
assert success is False
|
||||
assert "nextcloud" in body
|
||||
assert "sftp" in body
|
||||
|
||||
|
||||
def test_failed_upload_keeps_pdf_in_outgoing(tmp_config) -> None:
|
||||
"""Das OCR war erfolgreich — die Datei bleibt in outgoing/, nicht error/."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
_run_once(
|
||||
tmp_config,
|
||||
process_pdf={"side_effect": _fake_success},
|
||||
upload_folder={"return_value": False},
|
||||
)
|
||||
|
||||
assert (tmp_config.paths.outgoing / "OCR_a.pdf").exists()
|
||||
assert not (tmp_config.paths.error / "OCR_a.pdf").exists()
|
||||
|
||||
|
||||
def test_successful_uploads_count_as_success(tmp_config) -> None:
|
||||
"""Gegenprobe: wenn alle Uploads durchgehen, zählt es als Erfolg."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
service = _run_once(tmp_config, process_pdf={"side_effect": _fake_success})
|
||||
|
||||
assert service.success_count == 1
|
||||
assert service.error_count == 0
|
||||
|
||||
|
||||
def test_dispatch_uploads_reports_failed_targets(tmp_config) -> None:
|
||||
"""_dispatch_uploads() liefert die Namen der fehlgeschlagenen Ziele."""
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
pdf = tmp_config.paths.outgoing / "x.pdf"
|
||||
pdf.write_bytes(b"%PDF-1.4\n")
|
||||
with patch("pdf_ocr_hotfolder.service.upload_folder", return_value=True), \
|
||||
patch("pdf_ocr_hotfolder.service.upload_nextcloud", return_value=False), \
|
||||
patch("pdf_ocr_hotfolder.service.upload_sftp", return_value=True):
|
||||
assert service._dispatch_uploads(pdf) == ["nextcloud"]
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.upload_folder", return_value=True), \
|
||||
patch("pdf_ocr_hotfolder.service.upload_nextcloud", return_value=True), \
|
||||
patch("pdf_ocr_hotfolder.service.upload_sftp", return_value=True):
|
||||
assert service._dispatch_uploads(pdf) == []
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
|
||||
# ---------------- Stabilitäts-Timeout ----------------
|
||||
|
||||
def test_unstable_file_counts_as_error(tmp_config) -> None:
|
||||
"""Stabilisiert sich eine Datei nicht, ist das ein Fehler (Exit 1)."""
|
||||
(tmp_config.paths.incoming / "slow.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None), \
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable", return_value=False), \
|
||||
patch("pdf_ocr_hotfolder.service.process_pdf") as proc:
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
errors = service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
assert errors == 1
|
||||
assert service.error_count == 1
|
||||
proc.assert_not_called()
|
||||
|
||||
|
||||
def test_unstable_file_stays_in_incoming(tmp_config) -> None:
|
||||
"""Die instabile Datei bleibt bewusst in incoming/ liegen."""
|
||||
(tmp_config.paths.incoming / "slow.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None), \
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable", return_value=False):
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
assert (tmp_config.paths.incoming / "slow.pdf").exists()
|
||||
assert not (tmp_config.paths.error / "slow.pdf").exists()
|
||||
|
||||
|
||||
def test_vanished_file_is_not_an_error(tmp_config) -> None:
|
||||
"""Verschwundene Datei ist kein Fehler — _wait_until_stable liefert dafür
|
||||
ebenfalls False."""
|
||||
pdf = tmp_config.paths.incoming / "weg.pdf"
|
||||
pdf.write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def vanish(path: Path, **kwargs) -> bool:
|
||||
path.unlink(missing_ok=True)
|
||||
return False
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None), \
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable", side_effect=vanish):
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
errors = service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
assert errors == 0
|
||||
assert service.error_count == 0
|
||||
@@ -0,0 +1,72 @@
|
||||
"""Tests für Issue #3: Ghostscript 10.0.0–10.02.0 PDF/A-Bug-Erkennung."""
|
||||
from __future__ import annotations
|
||||
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.service import (
|
||||
PreflightError,
|
||||
check_preflight,
|
||||
is_ghostscript_broken,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("version,expected", [
|
||||
# Betroffene Versionen
|
||||
("10.0.0", True),
|
||||
("10.00.0", True),
|
||||
("10.01.0", True),
|
||||
("10.01.1", True),
|
||||
("10.01.2", True),
|
||||
("10.02.0", True),
|
||||
# Sichere Versionen
|
||||
("10.02.1", False),
|
||||
("10.03.0", False),
|
||||
("10.04.0", False),
|
||||
("11.0.0", False),
|
||||
("9.56.1", False), # Debian 11 / Ubuntu 22.04
|
||||
("9.55.0", False),
|
||||
# Edge cases
|
||||
("", False),
|
||||
(None, False),
|
||||
("garbage", False),
|
||||
])
|
||||
def test_is_ghostscript_broken(version, expected) -> None:
|
||||
assert is_ghostscript_broken(version) is expected
|
||||
|
||||
|
||||
def test_check_preflight_without_pdfa_passes_with_broken_gs() -> None:
|
||||
"""Ohne pdfa_level darf der betroffene GS verwendet werden."""
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value="/usr/bin/fake"), \
|
||||
patch("pdf_ocr_hotfolder.service.detect_ghostscript_version",
|
||||
return_value="10.0.0"):
|
||||
check_preflight(pdfa_level="") # darf nicht werfen
|
||||
|
||||
|
||||
def test_check_preflight_with_pdfa_fails_on_broken_gs() -> None:
|
||||
"""Mit pdfa_level + kaputtem GS → PreflightError mit hilfreicher Meldung."""
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value="/usr/bin/fake"), \
|
||||
patch("pdf_ocr_hotfolder.service.detect_ghostscript_version",
|
||||
return_value="10.0.0"):
|
||||
with pytest.raises(PreflightError, match="Ghostscript 10.0.0"):
|
||||
check_preflight(pdfa_level="2")
|
||||
|
||||
|
||||
def test_check_preflight_with_pdfa_passes_on_fixed_gs() -> None:
|
||||
"""Mit pdfa_level + gefixtem GS → ok."""
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value="/usr/bin/fake"), \
|
||||
patch("pdf_ocr_hotfolder.service.detect_ghostscript_version",
|
||||
return_value="10.02.1"):
|
||||
check_preflight(pdfa_level="2") # darf nicht werfen
|
||||
|
||||
|
||||
def test_default_config_pdfa_level_is_empty() -> None:
|
||||
"""Default-Config der Beispiel-Datei soll pdfa_level='' enthalten (Issue #3)."""
|
||||
from pathlib import Path
|
||||
import tomllib
|
||||
cfg_path = Path(__file__).parent.parent / "config.example.toml"
|
||||
with cfg_path.open("rb") as f:
|
||||
data = tomllib.load(f)
|
||||
assert data["ocr"]["pdfa_level"] == "", \
|
||||
"config.example.toml muss pdfa_level='' als sicheren Default haben"
|
||||
@@ -0,0 +1,81 @@
|
||||
"""Tests für [ocr].timeout → ocrmypdf `tesseract_timeout`.
|
||||
|
||||
ocrmypdf wird hier komplett gemockt (per sys.modules), es läuft also nie
|
||||
wirklich — die Tests laufen auch ohne installiertes ocrmypdf.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
import tomllib
|
||||
from pathlib import Path
|
||||
from types import ModuleType
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.config import OcrConfig
|
||||
from pdf_ocr_hotfolder.processor import run_ocr
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fake_ocrmypdf(monkeypatch) -> ModuleType:
|
||||
"""Schiebt ein Dummy-ocrmypdf in sys.modules und merkt sich die kwargs."""
|
||||
mod = ModuleType("ocrmypdf")
|
||||
mod.calls = [] # type: ignore[attr-defined]
|
||||
|
||||
def ocr(src, dst, **kwargs):
|
||||
mod.calls.append({"src": src, "dst": dst, "kwargs": kwargs}) # type: ignore[attr-defined]
|
||||
Path(dst).write_bytes(b"%PDF-1.4 ocr\n")
|
||||
|
||||
mod.ocr = ocr # type: ignore[attr-defined]
|
||||
monkeypatch.setitem(sys.modules, "ocrmypdf", mod)
|
||||
return mod
|
||||
|
||||
|
||||
def _run(fake, tmp_path: Path, cfg: OcrConfig) -> dict:
|
||||
src = tmp_path / "in.pdf"
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
run_ocr(src, tmp_path / "out.pdf", cfg)
|
||||
assert len(fake.calls) == 1
|
||||
return fake.calls[0]["kwargs"]
|
||||
|
||||
|
||||
def test_timeout_is_passed_as_tesseract_timeout(fake_ocrmypdf, tmp_path: Path) -> None:
|
||||
kwargs = _run(fake_ocrmypdf, tmp_path, OcrConfig(timeout=120))
|
||||
assert kwargs["tesseract_timeout"] == 120.0
|
||||
|
||||
|
||||
def test_timeout_zero_means_no_limit(fake_ocrmypdf, tmp_path: Path) -> None:
|
||||
"""0 = kein Limit → der Key darf NICHT durchgereicht werden.
|
||||
|
||||
ocrmypdf würde tesseract_timeout=0 als 'OCR überspringen' auslegen.
|
||||
"""
|
||||
kwargs = _run(fake_ocrmypdf, tmp_path, OcrConfig(timeout=0))
|
||||
assert "tesseract_timeout" not in kwargs
|
||||
|
||||
|
||||
def test_negative_timeout_is_ignored(fake_ocrmypdf, tmp_path: Path) -> None:
|
||||
kwargs = _run(fake_ocrmypdf, tmp_path, OcrConfig(timeout=-5))
|
||||
assert "tesseract_timeout" not in kwargs
|
||||
|
||||
|
||||
def test_default_timeout_is_passed(fake_ocrmypdf, tmp_path: Path) -> None:
|
||||
kwargs = _run(fake_ocrmypdf, tmp_path, OcrConfig())
|
||||
assert kwargs["tesseract_timeout"] == 300.0
|
||||
|
||||
|
||||
def test_other_kwargs_still_present(fake_ocrmypdf, tmp_path: Path) -> None:
|
||||
"""Der neue Key ersetzt nichts Bestehendes."""
|
||||
kwargs = _run(fake_ocrmypdf, tmp_path,
|
||||
OcrConfig(languages="deu", jobs=2, pdfa_level=""))
|
||||
assert kwargs["language"] == "deu"
|
||||
assert kwargs["jobs"] == 2
|
||||
assert kwargs["output_type"] == "pdf"
|
||||
assert kwargs["skip_text"] is True
|
||||
|
||||
|
||||
def test_config_default_matches_example(tmp_path: Path) -> None:
|
||||
"""Dataclass-Default und config.example.toml dürfen nicht auseinanderlaufen."""
|
||||
cfg_path = Path(__file__).parent.parent / "config.example.toml"
|
||||
with cfg_path.open("rb") as f:
|
||||
data = tomllib.load(f)
|
||||
assert data["ocr"]["timeout"] == OcrConfig().timeout == 300
|
||||
@@ -0,0 +1,96 @@
|
||||
"""Tests für Issue #2: --once Modus muss Exit-Code != 0 bei Fehlern liefern."""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from pdf_ocr_hotfolder.processor import ProcessResult
|
||||
from pdf_ocr_hotfolder.service import HotfolderService
|
||||
|
||||
|
||||
def _fake_success(src: Path, working_dir, outgoing_dir, error_dir, **kwargs):
|
||||
out = outgoing_dir / f"OCR_{src.name}"
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_bytes(b"%PDF-1.4 ocr\n")
|
||||
src.unlink(missing_ok=True)
|
||||
return ProcessResult(src, out, True)
|
||||
|
||||
|
||||
def _fake_failure(src: Path, working_dir, outgoing_dir, error_dir, **kwargs):
|
||||
error_dir.mkdir(parents=True, exist_ok=True)
|
||||
dest = error_dir / src.name
|
||||
src.rename(dest)
|
||||
return ProcessResult(src, outgoing_dir / f"OCR_{src.name}", False,
|
||||
error="fake ocr failure")
|
||||
|
||||
|
||||
def _run(tmp_config, fake_process):
|
||||
"""Helper: führt run_once() mit gemocktem process_pdf und preflight aus."""
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None), \
|
||||
patch("pdf_ocr_hotfolder.service.process_pdf", side_effect=fake_process), \
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable", return_value=True):
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
return service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
|
||||
def test_once_exit_0_when_no_files(tmp_config) -> None:
|
||||
"""Szenario: Keine PDFs vorhanden → Exit 0."""
|
||||
errors = _run(tmp_config, _fake_success)
|
||||
assert errors == 0
|
||||
|
||||
|
||||
def test_once_exit_0_when_all_success(tmp_config) -> None:
|
||||
"""Szenario: Alle PDFs erfolgreich → Exit 0."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
(tmp_config.paths.incoming / "b.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
errors = _run(tmp_config, _fake_success)
|
||||
assert errors == 0
|
||||
|
||||
|
||||
def test_once_exit_nonzero_when_all_fail(tmp_config) -> None:
|
||||
"""Szenario: Alle PDFs fehlgeschlagen → Exit != 0 (Issue #2)."""
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
(tmp_config.paths.incoming / "b.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
errors = _run(tmp_config, _fake_failure)
|
||||
assert errors == 2
|
||||
|
||||
|
||||
def test_once_exit_nonzero_when_some_fail(tmp_config) -> None:
|
||||
"""Szenario: Teilweise fehlgeschlagen → Exit != 0."""
|
||||
(tmp_config.paths.incoming / "ok.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
(tmp_config.paths.incoming / "bad.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def mixed(src, *args, **kwargs):
|
||||
if "bad" in src.name:
|
||||
return _fake_failure(src, *args, **kwargs)
|
||||
return _fake_success(src, *args, **kwargs)
|
||||
|
||||
errors = _run(tmp_config, mixed)
|
||||
assert errors == 1
|
||||
|
||||
|
||||
def test_counters_track_success_and_failure(tmp_config) -> None:
|
||||
"""success_count und error_count sollen korrekt mitzählen."""
|
||||
(tmp_config.paths.incoming / "ok.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
(tmp_config.paths.incoming / "bad.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
def mixed(src, *args, **kwargs):
|
||||
if "bad" in src.name:
|
||||
return _fake_failure(src, *args, **kwargs)
|
||||
return _fake_success(src, *args, **kwargs)
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None), \
|
||||
patch("pdf_ocr_hotfolder.service.process_pdf", side_effect=mixed), \
|
||||
patch("pdf_ocr_hotfolder.service._wait_until_stable", return_value=True):
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
service.run_once()
|
||||
assert service.success_count == 1
|
||||
assert service.error_count == 1
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
@@ -0,0 +1,315 @@
|
||||
"""Tests für Feature: konfigurierbare Dateinamen und Original-Behandlung."""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.config import OcrConfig, OutputConfig, VeraPdfConfig
|
||||
from pdf_ocr_hotfolder.processor import build_output_name, process_pdf
|
||||
from pdf_ocr_hotfolder.service import PreflightError, check_output_config
|
||||
|
||||
|
||||
# ---------------- build_output_name ----------------
|
||||
|
||||
@pytest.mark.parametrize("src,mode,tag,expected", [
|
||||
# prefix
|
||||
("scan.pdf", "prefix", "OCR_", "OCR_scan.pdf"),
|
||||
("scan.pdf", "prefix", "[OCR] ", "[OCR] scan.pdf"),
|
||||
# suffix (Tag vor Extension)
|
||||
("scan.pdf", "suffix", "_OCR", "scan_OCR.pdf"),
|
||||
("scan.pdf", "suffix", "-ocr", "scan-ocr.pdf"),
|
||||
# none
|
||||
("scan.pdf", "none", "OCR_", "scan.pdf"),
|
||||
# leerer Tag = none
|
||||
("scan.pdf", "prefix", "", "scan.pdf"),
|
||||
("scan.pdf", "suffix", "", "scan.pdf"),
|
||||
# Mehrfach-Punkte im Namen: nur letzte Extension zählt
|
||||
("rechnung.2026.pdf", "suffix", "_OCR", "rechnung.2026_OCR.pdf"),
|
||||
("rechnung.2026.pdf", "prefix", "OCR_", "OCR_rechnung.2026.pdf"),
|
||||
# Name ohne Extension
|
||||
("NO_EXT", "suffix", "_OCR", "NO_EXT_OCR"),
|
||||
])
|
||||
def test_build_output_name(src, mode, tag, expected) -> None:
|
||||
assert build_output_name(src, mode, tag) == expected
|
||||
|
||||
|
||||
def test_build_output_name_invalid_mode() -> None:
|
||||
with pytest.raises(ValueError, match="name_mode"):
|
||||
build_output_name("x.pdf", "bogus", "OCR_")
|
||||
|
||||
|
||||
# ---------------- check_output_config ----------------
|
||||
|
||||
def test_check_output_config_delete_ok() -> None:
|
||||
check_output_config("delete", "") # ok
|
||||
|
||||
|
||||
def test_check_output_config_archive_requires_dir() -> None:
|
||||
with pytest.raises(PreflightError, match="archive_dir"):
|
||||
check_output_config("archive", "")
|
||||
|
||||
|
||||
def test_check_output_config_archive_with_dir_ok() -> None:
|
||||
check_output_config("archive", "/var/archive") # ok
|
||||
|
||||
|
||||
def test_check_output_config_invalid_mode() -> None:
|
||||
with pytest.raises(PreflightError, match="ungültig"):
|
||||
check_output_config("trash", "")
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name_mode", ["prefix", "suffix", "none"])
|
||||
def test_check_output_config_accepts_valid_name_modes(name_mode) -> None:
|
||||
check_output_config("delete", "", name_mode) # ok
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name_mode", ["prefixx", "Prefix", "", "postfix"])
|
||||
def test_check_output_config_invalid_name_mode(name_mode) -> None:
|
||||
"""Tippfehler in name_mode muss schon im Preflight auffallen."""
|
||||
with pytest.raises(PreflightError, match="name_mode"):
|
||||
check_output_config("delete", "", name_mode)
|
||||
|
||||
|
||||
def test_run_once_aborts_on_invalid_name_mode(tmp_config) -> None:
|
||||
"""Der Dienst bricht beim Start ab, bevor eine Datei angefasst wird."""
|
||||
from unittest.mock import patch
|
||||
|
||||
from pdf_ocr_hotfolder.service import HotfolderService
|
||||
|
||||
tmp_config.output.name_mode = "bogus"
|
||||
(tmp_config.paths.incoming / "a.pdf").write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
with patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None):
|
||||
with pytest.raises(PreflightError, match="name_mode"):
|
||||
service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
# Datei wurde nicht angefasst
|
||||
assert (tmp_config.paths.incoming / "a.pdf").exists()
|
||||
|
||||
|
||||
def test_main_returns_2_on_invalid_name_mode(tmp_path: Path, monkeypatch) -> None:
|
||||
"""CLI liefert Exit-Code 2 — gleicher Mechanismus wie die übrigen Preflights."""
|
||||
import sys
|
||||
from unittest.mock import patch as _patch
|
||||
|
||||
cfg_file = tmp_path / "cfg.toml"
|
||||
cfg_file.write_text(f"""
|
||||
[paths]
|
||||
incoming = "{tmp_path / 'in'}"
|
||||
outgoing = "{tmp_path / 'out'}"
|
||||
working = "{tmp_path / 'work'}"
|
||||
error = "{tmp_path / 'err'}"
|
||||
|
||||
[output]
|
||||
name_mode = "bogus"
|
||||
""")
|
||||
monkeypatch.setattr(sys, "argv",
|
||||
["pdf-ocr-hotfolder", "--config", str(cfg_file), "--once"])
|
||||
with _patch("pdf_ocr_hotfolder.service.check_preflight", return_value=None):
|
||||
from pdf_ocr_hotfolder.__main__ import main
|
||||
assert main() == 2
|
||||
|
||||
|
||||
# ---------------- process_pdf mit Original-Behandlung ----------------
|
||||
|
||||
def _fake_ocr(src: Path, dst: Path, cfg: OcrConfig) -> None:
|
||||
"""Simuliert ocrmypdf: kopiert Inhalt, erzeugt Zieldatei."""
|
||||
dst.write_bytes(b"%PDF-1.4 OCRed\n" + src.read_bytes())
|
||||
|
||||
|
||||
def _prepare(tmp_path: Path) -> dict:
|
||||
dirs = {
|
||||
"working": tmp_path / "working",
|
||||
"outgoing": tmp_path / "outgoing",
|
||||
"error": tmp_path / "error",
|
||||
"archive": tmp_path / "archive",
|
||||
"incoming": tmp_path / "incoming",
|
||||
}
|
||||
for d in dirs.values():
|
||||
d.mkdir(parents=True, exist_ok=True)
|
||||
src = dirs["incoming"] / "scan.pdf"
|
||||
src.write_bytes(b"%PDF-1.4 original\n")
|
||||
return {"src": src, **dirs}
|
||||
|
||||
|
||||
def test_process_pdf_prefix_delete(tmp_path: Path) -> None:
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="delete")
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr):
|
||||
result = process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=False),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
assert result.success
|
||||
assert (env["outgoing"] / "OCR_scan.pdf").exists()
|
||||
# Original ist weg, weder in incoming noch in working
|
||||
assert not env["src"].exists()
|
||||
assert not (env["working"] / "scan.pdf").exists()
|
||||
|
||||
|
||||
def test_process_pdf_suffix_delete(tmp_path: Path) -> None:
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="suffix", name_tag="_OCR",
|
||||
original_on_success="delete")
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr):
|
||||
result = process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=False),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
assert result.success
|
||||
assert (env["outgoing"] / "scan_OCR.pdf").exists()
|
||||
|
||||
|
||||
def test_process_pdf_none_mode(tmp_path: Path) -> None:
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="none", name_tag="OCR_",
|
||||
original_on_success="delete")
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr):
|
||||
result = process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=False),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
assert result.success
|
||||
# Ausgang hat GLEICHEN Namen wie Original
|
||||
assert (env["outgoing"] / "scan.pdf").exists()
|
||||
|
||||
|
||||
def test_process_pdf_archive_original(tmp_path: Path) -> None:
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="archive",
|
||||
archive_dir=str(env["archive"]))
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr):
|
||||
result = process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=False),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
assert result.success
|
||||
assert (env["outgoing"] / "OCR_scan.pdf").exists()
|
||||
# Original liegt jetzt im Archiv
|
||||
archived = env["archive"] / "scan.pdf"
|
||||
assert archived.exists()
|
||||
assert archived.read_bytes() == b"%PDF-1.4 original\n"
|
||||
|
||||
|
||||
def test_process_pdf_archive_name_collision(tmp_path: Path) -> None:
|
||||
"""Bei Namens-Kollision im Archiv wird Timestamp angehängt."""
|
||||
env = _prepare(tmp_path)
|
||||
# Vorhandene Kollisions-Datei
|
||||
(env["archive"] / "scan.pdf").write_bytes(b"old")
|
||||
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="archive",
|
||||
archive_dir=str(env["archive"]))
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr):
|
||||
process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=False),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
# Alte Datei unverändert
|
||||
assert (env["archive"] / "scan.pdf").read_bytes() == b"old"
|
||||
# Neue Datei mit Timestamp-Suffix
|
||||
archived = list(env["archive"].glob("scan_*.pdf"))
|
||||
assert len(archived) == 1
|
||||
assert archived[0].read_bytes() == b"%PDF-1.4 original\n"
|
||||
|
||||
|
||||
# ---------------- veraPDF FAIL: Original folgt original_on_success ----------------
|
||||
|
||||
def _run_with_vera_fail(env: dict, out_cfg: OutputConfig):
|
||||
"""process_pdf mit gemocktem OCR und einem veraPDF, das FAIL meldet."""
|
||||
with patch("pdf_ocr_hotfolder.processor.run_ocr", side_effect=_fake_ocr), \
|
||||
patch("pdf_ocr_hotfolder.processor.run_verapdf", return_value=False):
|
||||
return process_pdf(
|
||||
src=env["src"],
|
||||
working_dir=env["working"],
|
||||
outgoing_dir=env["outgoing"],
|
||||
error_dir=env["error"],
|
||||
ocr_cfg=OcrConfig(),
|
||||
vera_cfg=VeraPdfConfig(enabled=True),
|
||||
output_cfg=out_cfg,
|
||||
)
|
||||
|
||||
|
||||
def test_process_pdf_verapdf_fail_delete_removes_original(tmp_path: Path) -> None:
|
||||
"""delete: Verhalten wie bisher — OCR-Ergebnis nach error/, Original weg."""
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="delete")
|
||||
result = _run_with_vera_fail(env, out_cfg)
|
||||
|
||||
assert not result.success
|
||||
assert result.verapdf_passed is False
|
||||
# OCR-Ergebnis liegt in error/
|
||||
assert (env["error"] / "__ocr_OCR_scan.pdf").exists()
|
||||
# Original ist weg
|
||||
assert not env["src"].exists()
|
||||
assert not (env["working"] / "scan.pdf").exists()
|
||||
assert not (env["outgoing"] / "OCR_scan.pdf").exists()
|
||||
|
||||
|
||||
def test_process_pdf_verapdf_fail_archive_keeps_original(tmp_path: Path) -> None:
|
||||
"""archive: das Original darf im Fehlerfall NICHT verloren gehen."""
|
||||
env = _prepare(tmp_path)
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="archive",
|
||||
archive_dir=str(env["archive"]))
|
||||
result = _run_with_vera_fail(env, out_cfg)
|
||||
|
||||
assert not result.success
|
||||
assert result.verapdf_passed is False
|
||||
# OCR-Ergebnis liegt in error/
|
||||
assert (env["error"] / "__ocr_OCR_scan.pdf").exists()
|
||||
# Original liegt unversehrt im Archiv
|
||||
archived = env["archive"] / "scan.pdf"
|
||||
assert archived.exists()
|
||||
assert archived.read_bytes() == b"%PDF-1.4 original\n"
|
||||
assert not (env["working"] / "scan.pdf").exists()
|
||||
assert not (env["outgoing"] / "OCR_scan.pdf").exists()
|
||||
|
||||
|
||||
def test_process_pdf_verapdf_fail_archive_name_collision(tmp_path: Path) -> None:
|
||||
"""Auch im veraPDF-FAIL-Pfad greift der Timestamp-Kollisionsschutz."""
|
||||
env = _prepare(tmp_path)
|
||||
(env["archive"] / "scan.pdf").write_bytes(b"old")
|
||||
out_cfg = OutputConfig(name_mode="prefix", name_tag="OCR_",
|
||||
original_on_success="archive",
|
||||
archive_dir=str(env["archive"]))
|
||||
_run_with_vera_fail(env, out_cfg)
|
||||
|
||||
assert (env["archive"] / "scan.pdf").read_bytes() == b"old"
|
||||
archived = list(env["archive"].glob("scan_*.pdf"))
|
||||
assert len(archived) == 1
|
||||
assert archived[0].read_bytes() == b"%PDF-1.4 original\n"
|
||||
@@ -0,0 +1,75 @@
|
||||
"""Tests für Issue #1: Preflight-Check bei fehlendem Tesseract."""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from unittest.mock import patch
|
||||
|
||||
import pytest
|
||||
|
||||
from pdf_ocr_hotfolder.service import (
|
||||
HotfolderService,
|
||||
PreflightError,
|
||||
check_preflight,
|
||||
)
|
||||
|
||||
|
||||
def test_preflight_passes_when_all_binaries_present() -> None:
|
||||
"""Wenn tesseract + gs im PATH sind, darf kein Fehler fliegen."""
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value="/usr/bin/fake"):
|
||||
check_preflight() # darf nicht werfen
|
||||
|
||||
|
||||
def test_preflight_fails_when_tesseract_missing() -> None:
|
||||
"""Fehlendes tesseract → PreflightError mit passender Meldung."""
|
||||
def fake_which(name: str) -> str | None:
|
||||
return None if name == "tesseract" else "/usr/bin/fake"
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", side_effect=fake_which):
|
||||
with pytest.raises(PreflightError, match="tesseract"):
|
||||
check_preflight()
|
||||
|
||||
|
||||
def test_preflight_fails_when_ghostscript_missing() -> None:
|
||||
def fake_which(name: str) -> str | None:
|
||||
return None if name == "gs" else "/usr/bin/fake"
|
||||
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", side_effect=fake_which):
|
||||
with pytest.raises(PreflightError, match="gs"):
|
||||
check_preflight()
|
||||
|
||||
|
||||
def test_preflight_lists_all_missing_binaries() -> None:
|
||||
"""Bei mehreren fehlenden Binaries werden alle genannt."""
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value=None):
|
||||
with pytest.raises(PreflightError) as exc_info:
|
||||
check_preflight()
|
||||
msg = str(exc_info.value)
|
||||
assert "tesseract" in msg
|
||||
assert "gs" in msg
|
||||
|
||||
|
||||
def test_run_once_raises_preflight_error(tmp_config) -> None:
|
||||
"""HotfolderService.run_once() wirft PreflightError, wenn tesseract fehlt."""
|
||||
service = HotfolderService(tmp_config)
|
||||
try:
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value=None):
|
||||
with pytest.raises(PreflightError):
|
||||
service.run_once()
|
||||
finally:
|
||||
service._executor.shutdown(wait=False)
|
||||
|
||||
|
||||
def test_main_returns_2_on_preflight_error(tmp_config, tmp_path, monkeypatch) -> None:
|
||||
"""CLI liefert Exit-Code 2 bei Preflight-Fehler (Issue #1 Szenario)."""
|
||||
cfg_file = tmp_path / "cfg.toml"
|
||||
cfg_file.write_text(f"""
|
||||
[paths]
|
||||
incoming = "{tmp_config.paths.incoming}"
|
||||
outgoing = "{tmp_config.paths.outgoing}"
|
||||
working = "{tmp_config.paths.working}"
|
||||
error = "{tmp_config.paths.error}"
|
||||
""")
|
||||
monkeypatch.setattr(sys, "argv", ["pdf-ocr-hotfolder", "--config", str(cfg_file), "--once"])
|
||||
with patch("pdf_ocr_hotfolder.service.shutil.which", return_value=None):
|
||||
from pdf_ocr_hotfolder.__main__ import main
|
||||
assert main() == 2
|
||||
@@ -0,0 +1,66 @@
|
||||
"""Tests für upload_folder() — Kopie per shutil.copyfile statt read_bytes()."""
|
||||
from __future__ import annotations
|
||||
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from pdf_ocr_hotfolder.config import FolderUpload
|
||||
from pdf_ocr_hotfolder.uploaders import upload_folder
|
||||
|
||||
|
||||
def test_upload_folder_copies_file(tmp_path: Path) -> None:
|
||||
src = tmp_path / "out" / "OCR_scan.pdf"
|
||||
src.parent.mkdir()
|
||||
src.write_bytes(b"%PDF-1.4 inhalt\n")
|
||||
target = tmp_path / "ziel"
|
||||
|
||||
assert upload_folder(src, FolderUpload(enabled=True, target=str(target)),
|
||||
tmp_path / "out") is True
|
||||
assert (target / "OCR_scan.pdf").read_bytes() == b"%PDF-1.4 inhalt\n"
|
||||
# Quelle bleibt liegen (Kopie, kein Move)
|
||||
assert src.exists()
|
||||
|
||||
|
||||
def test_upload_folder_uses_copyfile_not_read_bytes(tmp_path: Path) -> None:
|
||||
"""Große PDFs dürfen nicht komplett in den Speicher gelesen werden."""
|
||||
src = tmp_path / "out" / "OCR_scan.pdf"
|
||||
src.parent.mkdir()
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
target = tmp_path / "ziel"
|
||||
|
||||
with patch("pdf_ocr_hotfolder.uploaders.shutil.copyfile") as copyfile:
|
||||
upload_folder(src, FolderUpload(enabled=True, target=str(target)),
|
||||
tmp_path / "out")
|
||||
copyfile.assert_called_once()
|
||||
|
||||
|
||||
def test_upload_folder_skips_self_target(tmp_path: Path) -> None:
|
||||
"""Ist das Ziel = outgoing, wird nicht auf sich selbst kopiert."""
|
||||
out = tmp_path / "out"
|
||||
out.mkdir()
|
||||
src = out / "OCR_scan.pdf"
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.uploaders.shutil.copyfile") as copyfile:
|
||||
assert upload_folder(src, FolderUpload(enabled=True, target=""), out) is True
|
||||
copyfile.assert_not_called()
|
||||
assert src.read_bytes() == b"%PDF-1.4\n"
|
||||
|
||||
|
||||
def test_upload_folder_disabled_returns_true(tmp_path: Path) -> None:
|
||||
src = tmp_path / "OCR_scan.pdf"
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
assert upload_folder(src, FolderUpload(enabled=False), tmp_path) is True
|
||||
|
||||
|
||||
def test_upload_folder_reports_failure(tmp_path: Path) -> None:
|
||||
"""OSError beim Kopieren → False (wird vom Service als Fehler gezählt)."""
|
||||
src = tmp_path / "out" / "OCR_scan.pdf"
|
||||
src.parent.mkdir()
|
||||
src.write_bytes(b"%PDF-1.4\n")
|
||||
|
||||
with patch("pdf_ocr_hotfolder.uploaders.shutil.copyfile",
|
||||
side_effect=OSError("disk full")):
|
||||
assert upload_folder(src, FolderUpload(enabled=True,
|
||||
target=str(tmp_path / "ziel")),
|
||||
tmp_path / "out") is False
|
||||
@@ -2,6 +2,10 @@
|
||||
#
|
||||
# PDF OCR Hotfolder — Update-Script
|
||||
#
|
||||
# Aktualisiert Code und venv unter /opt/pdf-ocr-hotfolder/ sowie die
|
||||
# systemd Template-Unit. Danach werden alle laufenden Instanzen neu gestartet.
|
||||
# Config-Dateien unter /etc/pdf-ocr-hotfolder/ bleiben unverändert.
|
||||
#
|
||||
set -euo pipefail
|
||||
|
||||
RED='\033[0;31m'; GREEN='\033[0;32m'; YELLOW='\033[1;33m'; NC='\033[0m'
|
||||
@@ -15,7 +19,8 @@ if [ "${EUID}" -ne 0 ]; then
|
||||
fi
|
||||
|
||||
INSTALL_DIR="/opt/pdf-ocr-hotfolder"
|
||||
SERVICE_NAME="pdf-ocr-hotfolder"
|
||||
CONFIG_DIR="/etc/pdf-ocr-hotfolder"
|
||||
SERVICE_TEMPLATE="pdf-ocr-hotfolder@.service"
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
if [ -f "$SCRIPT_DIR/pdf_ocr_hotfolder/__init__.py" ]; then
|
||||
@@ -42,12 +47,18 @@ log_info "Install: $INSTALL_DIR"
|
||||
log_info "Version: $OLD_VERSION → $NEW_VERSION"
|
||||
echo
|
||||
|
||||
# Service-User aus systemd-Unit lesen
|
||||
SERVICE_USER="$(awk -F= '/^User=/{print $2}' /etc/systemd/system/${SERVICE_NAME}.service 2>/dev/null || echo pdfocr)"
|
||||
SERVICE_GROUP="$(awk -F= '/^Group=/{print $2}' /etc/systemd/system/${SERVICE_NAME}.service 2>/dev/null || echo pdfocr)"
|
||||
# Laufende Instanzen ermitteln
|
||||
mapfile -t RUNNING < <(systemctl list-units --no-legend --state=active 'pdf-ocr-hotfolder@*.service' 2>/dev/null | awk '{print $1}')
|
||||
if [ "${#RUNNING[@]}" -gt 0 ]; then
|
||||
log_info "Laufende Instanzen: ${RUNNING[*]}"
|
||||
else
|
||||
log_info "Keine laufenden Instanzen."
|
||||
fi
|
||||
|
||||
log_info "Stoppe Service..."
|
||||
systemctl stop "${SERVICE_NAME}.service" 2>/dev/null || true
|
||||
log_info "Stoppe laufende Instanzen..."
|
||||
for unit in "${RUNNING[@]}"; do
|
||||
systemctl stop "$unit" || true
|
||||
done
|
||||
|
||||
log_info "Backup erstellen..."
|
||||
BACKUP_DIR="/var/backups/pdf-ocr-hotfolder"
|
||||
@@ -60,29 +71,40 @@ rm -rf "$INSTALL_DIR/pdf_ocr_hotfolder"
|
||||
cp -r "$REPO_DIR/pdf_ocr_hotfolder" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/requirements.txt" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/VERSION" "$INSTALL_DIR/"
|
||||
cp "$REPO_DIR/config.example.toml" "$INSTALL_DIR/"
|
||||
echo "$REPO_DIR" > "$INSTALL_DIR/.repo_path"
|
||||
|
||||
log_info "Dependencies aktualisieren..."
|
||||
"$INSTALL_DIR/venv/bin/pip" install --upgrade pip -q
|
||||
"$INSTALL_DIR/venv/bin/pip" install --upgrade -r "$INSTALL_DIR/requirements.txt" -q
|
||||
|
||||
log_info "systemd-Unit aktualisieren..."
|
||||
sed -e "s|__SERVICE_USER__|$SERVICE_USER|g" \
|
||||
-e "s|__SERVICE_GROUP__|$SERVICE_GROUP|g" \
|
||||
"$REPO_DIR/systemd/pdf-ocr-hotfolder.service" \
|
||||
> "/etc/systemd/system/${SERVICE_NAME}.service"
|
||||
log_info "systemd Template-Unit aktualisieren..."
|
||||
cp "$REPO_DIR/systemd/$SERVICE_TEMPLATE" "/etc/systemd/system/$SERVICE_TEMPLATE"
|
||||
systemctl daemon-reload
|
||||
|
||||
log_info "Berechtigungen setzen..."
|
||||
chown -R "$SERVICE_USER:$SERVICE_GROUP" "$INSTALL_DIR"
|
||||
# Eigentümer des Codes bleibt der primäre User (pdfocr); Instanzen laufen
|
||||
# ggf. als anderer User, lesen aber nur den Code.
|
||||
PRIMARY_USER="$(stat -c '%U' "$INSTALL_DIR/venv" 2>/dev/null || echo pdfocr)"
|
||||
chown -R "$PRIMARY_USER":"$PRIMARY_USER" "$INSTALL_DIR"
|
||||
|
||||
log_info "Service starten..."
|
||||
systemctl start "${SERVICE_NAME}.service"
|
||||
sleep 2
|
||||
|
||||
if systemctl is-active --quiet "${SERVICE_NAME}.service"; then
|
||||
log_info "✅ Service läuft (Version $NEW_VERSION)"
|
||||
log_info "Starte Instanzen wieder..."
|
||||
FAIL=0
|
||||
for unit in "${RUNNING[@]}"; do
|
||||
systemctl start "$unit" || true
|
||||
sleep 1
|
||||
if systemctl is-active --quiet "$unit"; then
|
||||
log_info " ✅ $unit"
|
||||
else
|
||||
log_error "Service läuft nicht. journalctl -u $SERVICE_NAME -n 30"
|
||||
log_error " ❌ $unit — journalctl -u $unit -n 30"
|
||||
FAIL=1
|
||||
fi
|
||||
done
|
||||
|
||||
echo
|
||||
if [ "$FAIL" -eq 0 ]; then
|
||||
log_info "Update auf $NEW_VERSION abgeschlossen ✓"
|
||||
else
|
||||
log_warn "Update abgeschlossen, aber mindestens eine Instanz läuft nicht."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
Reference in New Issue
Block a user