✨(backend) add custom metrics when monitoring is enabled

We added custom metrics when the monitoring is enabled, we will measure
the cost of calling outside services like yhub and the converters, the
database pool and celery queue.
This commit is contained in:
Manuel Raynaud
2026-09-21 14:48:43 +02:00
parent 68655603cc
commit 18b9315ed5
11 changed files with 698 additions and 47 deletions
+1
View File
@@ -134,6 +134,7 @@ These are the environment variables you can set for the `impress-backend` contai
| OIDC_USE_NONCE | Use nonce for OIDC | true |
| POSTHOG_KEY | Posthog key for analytics | |
| PROMETHEUS_API_KEY | Bearer token required on `/metrics`. Required when the metrics are enabled (the application refuses to start without it). Can be given as a file with PROMETHEUS_API_KEY_FILE | |
| PROMETHEUS_CELERY_QUEUE_METRICS_ENABLED | When the metrics are enabled, report the length of the default Celery queue, asked to the broker at every scrape | True |
| PROMETHEUS_DB_METRICS_ENABLED | When the metrics are enabled, also count and time the SQL queries (swaps the postgresql engine for the instrumented one of django-prometheus) | True |
| PROMETHEUS_METRICS_ENABLED | Instrument the application with django-prometheus and serve `/metrics`. OFF by default. See documentation/metrics.md | False |
| PROMETHEUS_MULTIPROC_DIR | Directory the uvicorn workers share their metrics through. Local to the host or pod, owned by the user running the application | `<tmp>/impress-prometheus-<uid>` |
+33
View File
@@ -49,6 +49,34 @@ scrape_configs:
With `DB_PSYCOPG_POOL_ENABLED`, `django_db_new_connections_total` counts the
connections taken from the pool, not the connections opened to Postgres.
- Calls to the other services, in
`docs_outgoing_request_duration_seconds{service,operation,method,status}` and
`docs_outgoing_requests_inflight{service,operation,method}`. `service` is
`yhub`, `y-provider` or `docspec`; `operation` is the endpoint (`ydoc`,
`create-ydoc`, `reset-connections`, `convert`, ...), never the url; `status`
is the http status, `timeout` when the call was given up on, `error` when it
never got an answer. These calls are made inside requests (`duplicate`,
`formatted-content`, document creation from a file) and from the Celery
tasks: the in-flight gauge is what piles up when the collaboration server
slows down.
- The psycopg pool, when `DB_PSYCOPG_POOL_ENABLED` is on:
`docs_db_pool_size`, `docs_db_pool_available` and
`docs_db_pool_requests_waiting` (what the pool of each worker looked like at
the end of its last request, added up), and the exact counters of the pool:
`docs_db_pool_requests_total`, `docs_db_pool_requests_queued_total`,
`docs_db_pool_requests_wait_seconds_total`,
`docs_db_pool_requests_errors_total`, `docs_db_pool_connections_total`,
`docs_db_pool_connections_errors_total`. A rising
`rate(docs_db_pool_requests_queued_total)` with
`rate(docs_db_pool_requests_wait_seconds_total)` is the application waiting
for connections, before Postgres shows anything.
- `docs_celery_queue_length{queue}`: tasks waiting on the default Celery queue,
asked to the broker when the metrics are scraped (Redis/Valkey brokers only;
turn it off with `PROMETHEUS_CELERY_QUEUE_METRICS_ENABLED=False`). It is one
queue for the whole deployment, so every replica reports the same number:
read it with `max`, never `sum`. The Celery workers have no endpoint of
their own, which is why the backend reports it.
No label ever carries a path, a user or a document identifier. A request that
matches no route is counted under `<unnamed view>`.
@@ -82,6 +110,11 @@ application refuses a directory owned by another user or a symbolic link. Set
`PROMETHEUS_MULTIPROC_DIR` to put it elsewhere — it must be writable, and local
to the host or the pod: it is **not** shared between replicas.
The gauges (in-flight calls, pool state) are kept per worker process and added
up over the living ones. uvicorn has no hook telling when a worker is gone, so
whichever worker answers a scrape first drops the gauge files of the processes
that no longer exist.
Known limit: uvicorn recycles its workers (`--limit-max-requests`) and has no
hook to tell when one is gone. Every new worker writes two new 64 KiB files,
and the files of the workers that are gone must stay — their counters are part