Commit Graph
755 Commits
Author SHA1 Message Date
Manuel Raynaud 8f67c256c1 ♻️(backend) index the document content from updated_content endpoint
the search indexer reads it with `YHubService`, and the indexation
of an edited document is triggered by the `content-updated` call
the collaboration server makes — nothing else sees the content change
anymore.
It is queued as a celery task, throttled like the other updates, so
no indexation ever runs in the process serving the request.
A document whose content cannot be read is left out of the batch
rather than indexed empty, which would have erased it from the search
backend
2026-09-21 14:46:12 +02:00
Manuel Raynaud b2be3bd2cb (collaboration) notify the backend when the worker persists new content
notify the backend when the worker persists new content for
a document, so the lists ordered by `updated_at` follow the edits made on the
collaboration server. The backend serves it on
`POST /api/v1.0/documents/{id}/content-updated/`, authenticated with a short
lived RS256 JWT the collaboration server signs (`aud: "docs-backend"`) and
the backend verifies against the JWKS the collaboration server publishes on
`/collaboration/jwks/v1` — the mirror of the admin token the backend signs to
call it, so no long lived secret is shared and either side can roll its key
on its own
2026-09-21 14:46:11 +02:00
Manuel Raynaud 6f77d73e0e (backend) serve documents/{id}/formatted-content/ from yhub
The formatted-content endpoint was using the `document.content` to fetch
the ydoc from s3, we want to move from this usage to using yhub to
retrieve the content, so yhub is becoming our source of thruth.
2026-09-21 14:46:11 +02:00
Manuel Raynaud 216758437b 💥(backend) remove the documents/{id}/content/ endpoint
Both its PATCH and its GET: the content of a document is saved and
served by the collaboration server. The `content_patch` and
`content_retrieve` abilities go with it.
2026-09-21 14:46:11 +02:00
Manuel Raynaud d92dd27aed (backend) duplicate a document through the collaboration server
its stateis fetched from yhub and seeded into the copy instead
of being copied from the content stored by Django
2026-09-21 14:46:10 +02:00
Manuel Raynaud afe41106aa (collaboration) add a get-ydoc endpoint on yhub
`GET /collaboration/get-ydoc/v1/docs/{id}` answers the current Yjs state of a
document as a raw binary update, the read counterpart of create-ydoc, and
204 when the document has no content yet
2026-09-21 14:46:10 +02:00
Manuel Raynaud f33d078a75 (backend) call YHubService to seed initial document content
When a new Docs is created and a file is sent, as before we convert it
first and we need to use the raw content to seed it by calling the
create-ydoc api in the YHub service.
2026-09-21 14:46:10 +02:00
Manuel Raynaud 207201f1e2 ️(backend) reintroduce the reset connection mechanism
When an access change or is deleted or a link configuration changes, we
call the yhub server to reset connections and remove them if needed. The
YHubService is used for this.
2026-09-21 14:46:09 +02:00
Manuel Raynaud e2ee11b764 (backend) implement reset-connections and create-ydoc in YHubService
The reset-connections and create-ydoc are the first action we want to
implement in the YHubService. They will be used in next commits.
2026-09-21 14:46:09 +02:00
Manuel Raynaud 391ee77b7e ♻️(backend) audience is an enum to be used by the JWTService
To ease the use of the audience with the JWTService, we choose to create
an enum holding all the possible values and then use them in the Yhub
and Y-converter services.
2026-09-21 14:46:09 +02:00
Manuel Raynaud e76da04f31 (backend) add a service to call the yhub REST API
The backend application will have to call the yhub REST API for some
operations. We want to use a dedicated service to do that. This first
commit introduces the shape of this service, it only does the
configuration for now, calling actions will be implemented later.
2026-09-21 14:46:08 +02:00
Kevin JahnsandManuel Raynaud 85587422cb (collaboration) test the legacy migrations against a real yhub
Cover both paths off the legacy Django store end to end: the lazy seed on
first access, and the migrate endpoint replaying every S3 version. The tests
need no database — the admin JWT short-circuits document authorization, so a
fixture is an S3 object on a random uuid — and read the timeline through
yhub 0.5.0's `Accept: application/json`, which spares python a lib0 decoder.
CI grows a valkey service and starts a collaboration server alongside the
backend test job; the tests skip themselves when nothing answers on the new
COLLABORATION_API_URL setting, so `make test` without the dev stack still
passes.

Writing them turned up three things worth fixing in the server.

Backend reads now seed too. getAccessType short-circuited on the admin token
before reaching the legacy store, so a server-side read of an unmigrated
document answered with an empty one, and a create-ydoc against it would have
written a second lineage beside the content the first user access was about
to seed in.

Seeding no longer decides access; the backend's answer alone does. A legacy
object that cannot be migrated — it does not decode, or it exceeds the size
we load — opens as a new document instead of denying, since no retry can fix
it and refusing would leave the document unopenable by anyone. The cause is
logged once per attempt with the bucket, key and stack, and every later access
logs that it admitted a caller without migrating.

That made the failure classifier dangerous, so it is inverted. It was an
allowlist of retryable errors — eight socket errnos — which left every way S3
can refuse (AccessDenied on a rotated key, NoSuchBucket, a region redirect)
counting as "this object is unusable". Denying, that was survivable; opening
empty, one misscoped credential would fork every document touched during the
window. Now only a failure raised while interpreting bytes we already hold is
permanent, marked at the throw site, and everything else answers a retryable
503. Guessing wrong that way costs a retry; the other way costs the document.

The admin seed is also fenced to the org and to main, like the user path
above it. The legacy store is branchless — {docid}/file is main — and the
bookkeeping is per document, so seeding ?branch=draft would have written
main's content into an orphan room and left the real one permanently empty.

Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-21 14:46:07 +02:00
Anthony LCandManuel Raynaud 50cc5a5dd6 🛂(backend) add audience to jwt
Add audience to the jwt, scoping the token to it
prevents an admin JWT issued for another backend
service from being replayed against y-provider.
2026-09-21 14:42:14 +02:00
Anthony LCandManuel Raynaud 6e0d5a1494 🛂(django) use jwt token for converter services
The Y_PROVIDER_API_KEY shared secret is replaced by a
signed admin JWT when Django calls the y-provider
conversion endpoint.
2026-09-21 14:42:13 +02:00
Manuel Raynaud 065048dbec 🔥(backend) remove CollaborationService and can-edit endpoint
The CollaborationService was doing nothing since we started the
migration to yhub, all the code using it is now removed. Also the
`can-edit` endpoint and all the safeguard mechanism relying on the
presence of other users connected to the websocket will not be used
anymore, it will be possible to replace all of this with yhub, so all
this code is also removed.
2026-09-21 14:42:13 +02:00
Manuel Raynaud aa9a17deb9 (backend) add a method to create a dedicated admin token
For now the only token we will need is ont with the admin claim set to
True. To not repeat the creation of this token again and again, we
created a dedicated method to issue this token in the JWTService class.
2026-09-21 14:42:11 +02:00
Manuel Raynaud d74f239f3e (backend) publish the JWT public key on a JWKS endpoint
The yhub service will need our public key in order to validate the jwt
token we will used. We choose to expose a jwks endpoint as it is a
standard wat to do this.
2026-09-21 14:42:11 +02:00
Manuel Raynaud 5617765969 (backend) add a service generating cached RS256 JWT tokens
We want to generate jwt token using the RS256 algotrithm. This token
will be used for internal call with the yhub service.
2026-09-21 14:42:11 +02:00
Kevin JahnsandManuel Raynaud 76066d3123 ♻️(collaboration) switch collaboration server from hocuspocus to yhub
Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-21 14:42:10 +02:00
Anthony LC d01372fd90 ♻️(backend) return the full document in the duplicate response
The duplicate endpoint used to respond with only `{"id": ...}`. It now
returns the complete duplicated document representation, consistent
with the other document detail endpoints, so the frontend doesn't have
to make a follow-up request to get the new document's data.

This required setting `is_favorite` explicitly on the duplicated
document before serializing it: it is normally set by the
`annotate_is_favorite` queryset method, which the newly created
document never goes through. Being a read-only serializer field, it
was silently dropped from the response instead of raising an error. A
document can't be a favorite right after being created, so it is set
to `False` directly.
2026-09-18 16:17:31 +02:00
risk-altandAnthony LC 5b661d7224 🥅(frontend) warn before uploading a file over the size limit
Dropping a file larger than the allowed size showed a bare "unknown
error" in the editor. The proxy in front of the API cuts the request
and answers a 413 with an HTML body, so errorCauses threw while
parsing it as JSON and no cause ever reached the error panel.

The size limit the backend already enforces is now exposed by the
config endpoint, and the editor checks the file against it before
sending anything, with the same toast wording the document import
uses. errorCauses no longer throws on a body it cannot parse, and a
413 without a usable cause falls back to an explicit message, which
covers the instances whose proxy limit is lower than the application
one.

The size formatting duplicated in the import hook moved to a shared
util.

Signed-off-by: risk-alt <aldu6974@gmail.com>
2026-09-17 16:04:26 +02:00
Anthony LC 147bf68dda 🔖(release) minor 5.7.0
Added:
- 🔧(backend) fine tune redis cache options
- (frontend) make the full last-update date available
- 💄(frontend) redesign email confirmation standalone page

Changed:
- ⬆️(backend) upgrade celery to version 5.6.3
- ️(backend) stop using LEFT(value, LENGTH(path)) in sql queries
- 🚚(project) switch docspec image to ghcr.io/docspec/api
- 🚚(global) move favorite documents API endpoint
  to `/documents/favorites/`

Fixed:
- 🐛(backend) skip session creation for the liveness probe
- 🐛(frontend) preserve page titles when adding an emoji
- 🐛(frontend) scroll to the linked block in read-only documents
- 🐛(frontend) hide the selection highlight on presenter images
- 🐛(y-provider) prevent process crash on malformed websocket frames
- 🐛(frontend) keep commented text sharp when printing to PDF
- 🐛(docker) pull minio images from quay.io
- ️(frontend) restore presenter focus trapping after share links
- 🐛(frontend) export any raster image supported by the browser to a PDF
2026-09-15 16:19:39 +02:00
AntoLCandAnthony LC da4f409907 🌐(i18n) update translated strings
Update translated files with new translations
2026-09-15 14:00:47 +02:00
Manuel RaynaudandGitHub d596df9512 (backend) allow configuring trace sampling
Add a configuration knob for the trace sampling rate, so we can
enable tracing on middleware and cache spans when debugging slow
requests in production.
    
Sampling is set to 0 by default, so tracing stays fully off unless
explicitly enabled.
Copied from suitenumerique/meet#1690
2026-09-14 18:59:01 +00:00
Julien MaupetitandGitHub 31cff890b8 🚚(global) move favorite documents API endpoint to /documents/favorites/
To respect the globally used pattern, we can safely switch to a simpler
path
2026-09-14 10:01:36 +00:00
Manuel Raynaud 00cc95aa05 🔧(backend) move the DockerflowMiddleware higher in the middleware list
We decided to move the DockerflowMiddleware higher in the middleware
list to prevent future access to the database or redis in other
middleware that can have an impact on the liveness probe.
2026-09-11 12:55:20 +02:00
Manuel Raynaud 451499016e ️(backend) increase nb_accesses cache TTL
The nb_accesses cache TTL was very short, 30 seconds. That mean that the
user will hit the cache for a very short period and the cache is
probably not be hit. This is what we can see in the slow queries from
the pg_stat_statements table. The query to compute the nb_accesses is
executed a little bit less than the number of queries to list or
retrieve documents, meaning the cache is not used.
2026-09-11 12:55:20 +02:00
Manuel Raynaud 6baf20aaeb ️(backend) improve DocumentViewset.get_queryset
The filtering made in the DocumentViewset.get_queryset method is not
optimal and lead to a full scan of the Document table. The heavy part is
on the filtering on what the user can access between the accesses and
the link traces. To have better performance we make an union operation
of both document_id list and the filter the id on this list. Postgresql
will use the index on the id column.
2026-09-11 12:55:20 +02:00
Manuel Raynaud fbc3ef83ba ️(backend) stop using LEFT(value, LENGTH(path)) in sql queries
Comparing path with LEFT(value, LENGTH(path)) makes a sequential scan on
all the Document table, the more this table grow, the more the query
using it will be slow. We dediced instead to lookup on the path
extracting all ancestors path for a given document and then make a path
IN statement to use the index existing on the path column.
2026-09-11 12:55:20 +02:00
Manuel RaynaudandGitHub 673a670dd9 🔧(backend) configure request.summary logger
The request.summary logger is used by dockerflow. The INFO level is
always empty and is used everytime the liveness or readiness endpoint
are fetch.
2026-09-09 14:46:19 +00:00
Manuel Raynaud 137cecc0e1 🔧(backend) fine tune redis cache options
We want to configure other options on the redis cache. By default there
is no timeout on the connection to socket and no timeout for read/write
operations. We set default values in all caches used in production. The
settings IGNORE_EXCEPTIONS differ between the default and the session
cache. Activating it behaves like a missed cache. Enabling it for the
session should lead to unwanted side effects, by returning falsy on the
session creation, a retry mechanism of 10000 attempts is made in the
SessionStore.create method, the request can stay in this loop for a long
time.
2026-09-08 16:26:15 +02:00
Manuel Raynaud 0b8808f9a7 🐛(backend) skip session creation for the readiness probe
The readiness probe should also not create a new session. A new session
will live in redis and increase the number of keys inside it for
nothing. The readiness path is isgnored in the ForceSessionMiddleware
2026-09-08 14:44:57 +02:00
Manuel Raynaud 36a890a119 ⬆️(backend) upgrade celery to version 5.6.3
Celery version 5.6 has several fixes we want : two significant memory
leaks have been resolved and also a fix allowing a better use of psycopg
pool.
2026-09-08 14:44:56 +02:00
Manuel Raynaud 1e61b4a789 🐛(backend) skip session creation for the liveness probe
The ForceSessionMiddleware force the session creation, we want to
ignore it when the request is the liveness probe. The liveness probe
must not check if redis is available, this is the readiness probe job
2026-09-08 08:56:42 +02:00
Manuel Raynaud 3c1275c88d 🔖(release) patch 5.6.1
Added

- (frontend) export presenter slides as PDF #2487

Fixed

- 🐛(frontend) hide Leave in the doc menu when not logged in #2626
- 🐛(backend) allow to configure settings DATA_UPLOAD_MAX_MEMORY_SIZE
2026-09-04 18:19:06 +02:00
renovate[bot]andGitHub 36bf78558a ⬆️(dependencies) update django to v5.2.16 [SECURITY] 2026-09-04 16:05:54 +00:00
Manuel Raynaud efdac7444d (backend) add servestatic dependency
We removed previously whitenoise because it was not working with asgi
application. By removing it we also removed the way to serve the static
files in the application. There is an existing fork of whitenoise,
servestatic, that manage async application and we can use it to serve
static files.
2026-09-04 16:28:16 +02:00
Manuel Raynaud 3bede0d9a0 🐛(backend) allow to configure settings DATA_UPLOAD_MAX_MEMORY_SIZE
Release 3.17.2 of DRF now takes care of DATA_UPLOAD_MAX_MEMORY_SIZE
and is checked when the body request is parsed. Before that, DRF wasn't
using it at all and we were only looking for custom settings linked to
the media and conversion file upload. We must now also configure this
setting.
2026-09-04 10:21:43 +02:00
Anthony LC 059f1d004a 🔖(release) minor 5.6.0
Added:
- (frontend) Add "Copy link to block" feature
- (frontend) add word count to doc header toolbox
- (frontend) add find and replace feature to the editor

Changed:
- ️(frontend) use anchor links for interlinking sub-documents
- (frontend) reset side panel state between documents
- ️(frontend) announce search loading state for screen readers
- ♻️(frontend) change favorite to star
- 🚚(frontend) add doc move to doc options
- ♻️(frontend) unified menu
- (frontend) hide decorative emojis in document titles from SR
- ♻️(frontend) save the doc with a keepalive request when
  leaving the page

Fixed:
- 🐛(frontend) fix clipped formatting toolbar in new comment
  composer
- 🐛(backend) fix duplicating a document that has no content
- 📄(frontend) allowed partially export when MIT
- 🐛(backend) manage async support for Docs custom middleware

Removed:
- 🔥(backend) remove whitenoise package
2026-09-03 17:00:36 +02:00
AntoLCandAnthony LC 8b9fce27fd 🌐(i18n) update translated strings
Update translated files with new translations
2026-09-03 17:00:35 +02:00
Manuel Raynaud c69e790703 🔥(backend) remove whitenoise package
whitenoise middleware is failing a lot with a cancelled exception from
asyncio. Using whitenoise is not needed in our case, we are just serving
an API with django and DRF. We decided to completely remove it.
2026-09-03 08:25:15 +02:00
Manuel Raynaud 53bf783447 🐛(backend) manage async support for Docs custom middleware
Docs have 2 custom middlewares, both are only managing sync
requests. With Python 3.13 we didn't have any errors, but
since we upgraded to Python 3.14, we have a CancelledError
exception. We decided to use the MiddlewareMixin from Django
that is sync and async capable and will be responsible for
executing both middleware in the good mode.
2026-09-03 08:25:03 +02:00
Amine BOUKERFAandGitHub 681f9a8c40 🐛(backend) fix duplicating a document that has no conten
Document.content reads from object storage and returns None when nothing
was ever written there. That None, raised "content should be a string.",
so the duplicate endpoint answered a 500. Default to an empty string instead.
    
Signed-off-by: BOUKERFA Mohamed El Amine <boukerfa.ma@gmail.com>
2026-09-02 06:46:37 +00:00
renovate[bot]andGitHub 4633cc7690 ⬆️(dependencies) update djangorestframework to v3.17.2 [SECURITY] 2026-09-02 02:42:52 +00:00
Anthony LC 4e6d28e259 ♻️(frontend) update ui logo Docs
The logo Docs seems to have again changed in the
design system. We update the logo part accordingly
to the new design system logo.
2026-08-26 10:32:03 +02:00
Anthony LC 02195154a0 🔖(release) minor 5.5.0
Added:
- ️(frontend) restore skip to content link after header redesign
- 🌐(i18n) rename cn_CN to zh_CN, add eo_PL and zh_TW locales
- (backend) conditional email notification in server to server api
- (backend) profile api using django-silk

Changed:
- ️(frontend) use semantic `<dl>` structure in document info card
- ️(frontend) replace onboarding assets with webm and webp
- 💄(frontend) use the same highlight color for cells and moves
- ️(backend) optimize media_auth endpoint
- 🚸(frontend) print from document options menu

Fixed:
- 🐛(frontend) refresh pins after document deletion and restoration
- 🐛(frontend) redirect homepage to login when homepage feat
  is disabled
- 🐛(backend) ignore CSPs for API docs in development
- 🐛(frontend) export images embedded with a relative url
- 🐛(y-provider) fix sentry init
- 🐛(backend) handle object storage metadata keys case-insensitively
- 🐛(keycloak) fix database env variables in the self-hosting example
- 🐛(helm) show the database error while jobs wait for it to be ready
2026-08-24 21:40:18 +02:00
AntoLCandAnthony LC c8ed9ec349 🌐(i18n) update translated strings
Update translated files with new translations
2026-08-24 14:51:36 +02:00
Anthony LC 1df56199c6 🌐(i18n) add Polish language to django system
A new language has been added to the Django system,
allowing for Polish translations and localization
support.
We need to initialize the Polish language files before
being able to download the translations from
Crowdin. This commit includes the initial setup for
the Polish language, including the necessary configuration
files and directory structure.
2026-08-24 14:37:36 +02:00
Manuel Raynaud 4111e4e5ed ️(backend) optimize media_auth cpu usage
Once the sql queries improved we have still a bottleneck on large
concurrent requests on this endpoint. We notive in the profiles generated
that lot of time was spent in creating a new s3 client instance on each
request. django_storage use a thread local cache for signed and unsigned
connection, but using uvicorn we have a new thread for each request, so
on each request a new s3 client is generated and it appears to be an
expensive operation. To fix this issue, we cache the client and share it
accross all the thread and requests.
2026-08-20 17:28:31 +02:00
Manuel Raynaud 7372c4610f ️(backend) optimize media_auth sql queries
On the media_auth endpoint the first bottleneck we have is with
postgresql. We are looking for too much data and no index is used on the
attachments colum. When the lookup filter on the attachement columns, a
full scan is made on all the document table looking for each element in
the array, this operation is really expensive. To fix this we created a
GIN index on the attachments column. Also the readable_per_se lookup was
selecting too much data combined with the filter_descendants function.
We remove the usage of the filter_descendants, we choose to first fetch
all the paths where the attachment is found, this operation is fast
thanks to the new index, split all the paths in candidate paths and then
filter readable_per_se queryset with these paths. All these
modifications make the endpoint faster.
2026-08-20 16:28:31 +02:00