Hi Boris,
I got an error in Paperless. Here is the analysis (helped by KI):
Paperless-ngx: `train_classifier` fails every hour – NLTK `stopwords` corpus is missing from the snap
Description
On a Syncloud appliance running the `paperless` snap, the scheduled `train_classifier` task fails on every run because the NLTK `stopwords` corpus is not shipped with the snap and `/usr/share/nltk_data` does not exist. As a result, the document classifier is never trained, so automatic matching of tags, document types, correspondents and storage paths does not work.
Environment
- Syncloud appliance on Raspberry Pi (aarch64), Debian-based OS
- `paperless` snap revision 207 (previously 204), Paperless-ngx 3.1.3
- Python 3.14 (`/snap/paperless/207/paperless/usr/local/lib/python3.14`)
Observed behaviour
- `GET /api/status/` reports `tasks.classifier_status: “ERROR”` permanently.
- `train_classifier` runs hourly and fails every time. In the last 30 days: 249 failures of this task (254 failed tasks in total, 1794 successes).
- Journal (`snap.paperless.celery`):
File ".../nltk/corpus/util.py", line 129, in __getattr__
...
File ".../nltk/data.py", line 877, in find
raise LookupError(resource_not_found)
LookupError:
Resource 'stopwords' not found.
Please use the NLTK Downloader to obtain the resource:
>>> nltk.download('stopwords')
Attempted to load 'corpora/stopwords'
Searched in:
- PosixPath('/usr/share/nltk_data')
- The directory does not exist on the device: `ls /usr/share/nltk_data` → `No such file or directory`. A filesystem search finds no `nltk_data` or `stopwords*` anywhere.
Expected behaviour
The classifier trains successfully; `classifier_status` is `OK`.
Cause
The snap bundles the `nltk` Python package but not the corpus data it needs (`stopwords`, and likely also `punkt`/`punkt_tab` and `snowball_data`, which Paperless-ngx’s Docker image downloads at build time). The snap’s filesystem is read-only, so users cannot fix this without modifying the system, and manual changes would not survive a snap refresh.
Suggested fix
Download the NLTK data at snap build time (as the official Paperless-ngx Docker image does) and include it in the snap. Point Paperless to it with `PAPERLESS_NLTK_DIR` in the snap’s configuration (the default is `/usr/share/nltk_data`).
Impact
Not a data-loss issue, but every installation of this snap appears to be affected. The failing hourly task also fills the task history with errors and keeps health monitoring permanently red.
How to reproduce
1. Install the `paperless` snap on Syncloud.
2. Wait for the hourly `train_classifier` task, or check `/api/status/` → `tasks.classifier_status`.