# Paperless-ngx: \`train\_classifier\` fails every hour – NLTK \`stopwords\` corpus is missing from the snap

**URL:** <https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687>\
**Category:** Existing Apps\
**Tags:** paperless\
**Created:** [1 October 2026 13:35 UTC](https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687 "2026-10-01T13:35:42Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ralfb](https://yyz2.discourse-cdn.com/free1/user_avatar/syncloud.discourse.group/ralfb/32/288_2.png) [@ralfb](https://syncloud.discourse.group/u/ralfb)\
**Post date:** [1 October 2026 13:35 UTC](https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687/1 "2026-10-01T13:35:42Z")

</div>

Hi Boris,

I got an error in Paperless. Here is the analysis (helped by KI):

**Paperless-ngx: `train_classifier` fails every hour – NLTK `stopwords` corpus is missing from the snap**

**Description**

On a Syncloud appliance running the `paperless` snap, the scheduled `train_classifier` task fails on every run because the NLTK `stopwords` corpus is not shipped with the snap and `/usr/share/nltk_data` does not exist. As a result, the document classifier is never trained, so automatic matching of tags, document types, correspondents and storage paths does not work.

**Environment**

- Syncloud appliance on Raspberry Pi (aarch64), Debian-based OS

- `paperless` snap revision 207 (previously 204), Paperless-ngx 3.1.3

- Python 3.14 (`/snap/paperless/207/paperless/usr/local/lib/python3.14`)

**Observed behaviour**

- `GET /api/status/` reports `tasks.classifier_status: “ERROR”` permanently.

- `train_classifier` runs hourly and fails every time. In the last 30 days: 249 failures of this task (254 failed tasks in total, 1794 successes).

- Journal (`snap.paperless.celery`):

```auto

File ".../nltk/corpus/util.py", line 129, in __getattr__

  ...

File ".../nltk/data.py", line 877, in find

    raise LookupError(resource_not_found)

LookupError:

  Resource 'stopwords' not found.

  Please use the NLTK Downloader to obtain the resource:

  >>> nltk.download('stopwords')

  Attempted to load 'corpora/stopwords'

  Searched in:

    - PosixPath('/usr/share/nltk_data')

```

- The directory does not exist on the device: `ls /usr/share/nltk_data` → `No such file or directory`. A filesystem search finds no `nltk_data` or `stopwords*` anywhere.

**Expected behaviour**

The classifier trains successfully; `classifier_status` is `OK`.

**Cause**

The snap bundles the `nltk` Python package but not the corpus data it needs (`stopwords`, and likely also `punkt`/`punkt_tab` and `snowball_data`, which Paperless-ngx’s Docker image downloads at build time). The snap’s filesystem is read-only, so users cannot fix this without modifying the system, and manual changes would not survive a snap refresh.

**Suggested fix**

Download the NLTK data at snap build time (as the official Paperless-ngx Docker image does) and include it in the snap. Point Paperless to it with `PAPERLESS_NLTK_DIR` in the snap’s configuration (the default is `/usr/share/nltk_data`).

**Impact**

Not a data-loss issue, but every installation of this snap appears to be affected. The failing hourly task also fills the task history with errors and keeps health monitoring permanently red.

**How to reproduce**

1. Install the `paperless` snap on Syncloud.

2. Wait for the hourly `train_classifier` task, or check `/api/status/` → `tasks.classifier_status`.

---

<div class="post-metadata">

**Author:** ![boris](https://avatars.discourse-cdn.com/v4/letter/b/74df32/32.png) [@boris](https://syncloud.discourse.group/u/boris)\
**Post date:** [1 October 2026 14:41 UTC](https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687/2 "2026-10-01T14:41:48Z")

</div>

Let me check, the fix will be coming soon

---

<div class="post-metadata">

**Author:** ![boris](https://avatars.discourse-cdn.com/v4/letter/b/74df32/32.png) [@boris](https://syncloud.discourse.group/u/boris)\
**Post date:** [2 October 2026 05:57 UTC](https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687/3 "2026-10-02T05:57:52Z")

</div>

Fix is released, could you check please?

---

<div class="post-metadata">

**Author:** ![ralfb](https://yyz2.discourse-cdn.com/free1/user_avatar/syncloud.discourse.group/ralfb/32/288_2.png) [@ralfb](https://syncloud.discourse.group/u/ralfb)\
**Post date:** [2 October 2026 07:21 UTC](https://syncloud.discourse.group/t/paperless-ngx-train-classifier-fails-every-hour-nltk-stopwords-corpus-is-missing-from-the-snap/687/4 "2026-10-02T07:21:21Z")

</div>

Hi Boris,  
fix works, problem is solved. Thanks a lot!  
Regards,  
Ralf
