Commit Graph
767 Commits
Author SHA1 Message Date
Manuel Raynaud 789f79e853 ♻️(collab) retrieve history from user access
Previously we added to retrieve document response a new property
user_access_since that contains the date from when the user started to
have access to the document. The way t was made added an other
annotation to the Document queryset making the sql query more and more
complex. We decided to lighten the queryset and expose the access the
user has on the document instead and read the history from the
created_at property.
2026-09-23 12:06:05 +02:00
Kevin JahnsandManuel Raynaud a290748c0f ✨(collaboration) build the version history from the activity api
The version history has been dead since the migration. It listed S3 object
versions of the legacy `{pk}/file` key, and nothing writes that key any more,
so every document's list has been frozen at its migration date; restoring one
was a stub that closed the modal and did nothing, while still promising that
the document would be replaced.

It now reads the collaboration server, which is what keeps the history: the
list comes from `activity`, a selected version is previewed from `changeset`
as the document stood at that moment, and restoring one is a `rollback`.

A version is a minute of editing — changes less than a minute apart become
one, and none spans more than a minute. The collaboration server groups only
changes by the same author, so the browser merges what is left across authors:
a version is a moment in the document, not a moment in one person's editing.
Both are needed, and both use the same rule.

This grants `history.rollback` to editors, which is the first time a browser
may change the past rather than read it, and publishes the rollback route.
A reader is refused it twice over — the collaboration server treats it as a
dead grant without document write access, and the endpoint is withheld as well.
Mutations refuse where reads clamp, so a rollback reaching further back than
the history a user was granted is rejected rather than trimmed: nobody can
undo work that predates their own access, and a rollback with no bound at all
is refused outright. `prune`, which erases, stays granted to nobody. Restoring
is not destructive: it appends a change that undoes another, so what it
replaced stays in the history and can be restored again.

The backend's version endpoints are untouched and now have no caller. They are
marked deprecated with the condition for removing them, since until a document
has been replayed by `migrate_documents` they hold the only record of what it
looked like before it moved.

Also fixes the e2e helper that waited for the removed content endpoint, so it
never returned, and the three version tests that hung behind it.

Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-23 12:06:05 +02:00
Kevin JahnsandManuel Raynaud af55589a4f ✨(collaboration) show a user the history since they got access
The collaboration server's activity and changeset routes are opened to
the browser, so a document's editing history can be read from where it
actually lives now. What a user may see of it is bounded to the moment
they were given access to the document: joining a document that has been
written for a year does not hand them the year.

That rule is not new. It is the one the version endpoints have always
applied - "only those created after the user got access to the document"
- and the date is the same one: the earliest access the user holds on
the document or on any of its ancestors, so sharing a folder shares its
subtree from that moment. It was computed twice in the backend,
differently, and exposed nowhere. It is now a single annotation,
user_access_since, that the version endpoints and the collaboration
server both read, the latter through the document detail response it
already fetches to authorize a connection.

The bound is applied server-side and silently: a client asks for
whatever range it likes and receives only its own share, so there is no
bound for it to get wrong and none it can widen. It is a stored date
rather than a wall-clock-relative one, which is what keeps it stable
across a websocket re-check, and it is never zero - the one value that
would also unlock a full-history connection.

A reader who reaches a document through its link alone holds no access
and so has no date to bound a history with. They get none, which is why
the backend has always refused them their versions. rollback and prune
stay refused to everyone: restoring a version is a separate decision.

Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-23 12:06:05 +02:00
Manuel Raynaud ed5bd54820 ✨(backend) allow too migrate a specific document
The management command migrating document to yhub didn't allow to target
a specific document. This can be usefull for debugging purpose but also
to replay the migration of a specific document.
2026-09-23 12:06:05 +02:00
Manuel Raynaud e263001033 ✨(collaboration) erase content in yhub from clean_document command
The clean_document command makes a reset of a document deleting its
content and all the attachments linked to this subdocument and its
children. The hard delete api in yhub make the room, so the document id,
not usable at all and this is not not what we want. We added a new
custom api in yhub to manage this case, the document is hard deleted and
then the Tombstone to make the room reusable again.
2026-09-23 12:06:05 +02:00
Manuel Raynaud b5978acb91 ✨(backend) wired soft deletion with yhub server
yhub is the source of truth, when a user delete a document, it should
also be deleted in the yhub server. We call the yhub server in the
perform_destroy action but also the restore endpoint of yhub when a
document is restored.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 5a93cc86b3 ✨(backend) add a migrate_documents command
command replaying the legacy content of
the documents into the collaboration server, one call to its migrate endpoint
per document. Resumable and safe to re-run: what became of every document is
recorded (`impress_document_migration`), a server that is unwell is retried
with a backoff and a document it refuses is left for a later run
(`--retry-failed`). Bounded by `--concurrency`, `--rate` and `--limit`, most
recently edited documents first
2026-09-23 12:06:05 +02:00
Manuel Raynaud 2b1267f3d4 ✅(backend) correctly reload urls in tests
After removing most of the usage of S3 in the tests, these ones are
faster and make some flakyness more relevant. For example, in tests
related to the external api we have to reload the urls based on the
settings. We now have some race conditions where tests collapsed and
urls are not correctly reloaded.
2026-09-23 12:06:05 +02:00
Manuel Raynaud e530c0b6ec ♻️(backend) remove usage of s3 for document.content in tests
Tehe DocumentFactory was always creating a content and this content was
saved on S3. This leads to the creation of huge amount of content in the
S3 storage but not necesseraly used in the tests. In order to keep the
refactor to remove the usage of content from document.content but from
Yhub service, this content is no more generated. It is kept for part of
the code not yet refactor like the versionning feature.
2026-09-23 12:06:05 +02:00
Manuel Raynaud b13a1a4d16 ♻️(backend) seed the content of the demo documents using yhub
seed the content of the demo documents in the collaboration
server: `create_demo` no longer writes it to the object storage, which
nothing reads anymore, and fails with an explicit message when the
collaboration server is not running rather than building a corpus of
documents that would open empty
2026-09-23 12:06:05 +02:00
Manuel Raynaud 10d01455ed ♻️(backend) take adavantage of yhub 0.5.0 json encoding returns
The version 0.5.0 can manage response format by using accept and
content-type headers. In python we can't use for now the lib0 decoder so
we have to use the json format. When the lib0 decoder will be available
in pycrdt we will use it. So we can now use directly the /ydoc api to
fetch a document content instead the custom api made for this.
2026-09-23 12:06:05 +02:00
Manuel Raynaud ec84b9dbae ♻️(backend) duplicate the onboarding sandbox using YHub service
Duplicate the onboarding sandbox document through the
collaboration server: its content is read from there and copied under the
identity of the user the sandbox is created for. A collaboration server that
cannot be reached skips the sandbox, as a missing template already did, and
never fails the signup
2026-09-23 12:06:05 +02:00
Manuel Raynaud d22fac15b0 ♻️(backend) index the document content from updated_content endpoint
the search indexer reads it with `YHubService`, and the indexation
of an edited document is triggered by the `content-updated` call
the collaboration server makes — nothing else sees the content change
anymore.
It is queued as a celery task, throttled like the other updates, so
no indexation ever runs in the process serving the request.
A document whose content cannot be read is left out of the batch
rather than indexed empty, which would have erased it from the search
backend
2026-09-23 12:06:05 +02:00
Manuel Raynaud 5011c7eeb7 ✨(collaboration) notify the backend when the worker persists new content
notify the backend when the worker persists new content for
a document, so the lists ordered by `updated_at` follow the edits made on the
collaboration server. The backend serves it on
`POST /api/v1.0/documents/{id}/content-updated/`, authenticated with a short
lived RS256 JWT the collaboration server signs (`aud: "docs-backend"`) and
the backend verifies against the JWKS the collaboration server publishes on
`/collaboration/jwks/v1` — the mirror of the admin token the backend signs to
call it, so no long lived secret is shared and either side can roll its key
on its own
2026-09-23 12:06:05 +02:00
Manuel Raynaud 6d43f93645 ✨(backend) serve documents/{id}/formatted-content/ from yhub
The formatted-content endpoint was using the `document.content` to fetch
the ydoc from s3, we want to move from this usage to using yhub to
retrieve the content, so yhub is becoming our source of thruth.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 56d2a0bde6 💥(backend) remove the documents/{id}/content/ endpoint
Both its PATCH and its GET: the content of a document is saved and
served by the collaboration server. The `content_patch` and
`content_retrieve` abilities go with it.
2026-09-23 12:06:05 +02:00
Manuel Raynaud a2fbec6c4c ✨(backend) duplicate a document through the collaboration server
its stateis fetched from yhub and seeded into the copy instead
of being copied from the content stored by Django
2026-09-23 12:06:05 +02:00
Manuel Raynaud 4240bda2f0 ✨(collaboration) add a get-ydoc endpoint on yhub
`GET /collaboration/get-ydoc/v1/docs/{id}` answers the current Yjs state of a
document as a raw binary update, the read counterpart of create-ydoc, and
204 when the document has no content yet
2026-09-23 12:06:05 +02:00
Manuel Raynaud 9b1a3d1099 ✨(backend) call YHubService to seed initial document content
When a new Docs is created and a file is sent, as before we convert it
first and we need to use the raw content to seed it by calling the
create-ydoc api in the YHub service.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 6d326c88c9 ⏪️(backend) reintroduce the reset connection mechanism
When an access change or is deleted or a link configuration changes, we
call the yhub server to reset connections and remove them if needed. The
YHubService is used for this.
2026-09-23 12:06:05 +02:00
Manuel Raynaud edf9dee9a4 ✨(backend) implement reset-connections and create-ydoc in YHubService
The reset-connections and create-ydoc are the first action we want to
implement in the YHubService. They will be used in next commits.
2026-09-23 12:06:05 +02:00
Manuel Raynaud dd0599dce4 ♻️(backend) audience is an enum to be used by the JWTService
To ease the use of the audience with the JWTService, we choose to create
an enum holding all the possible values and then use them in the Yhub
and Y-converter services.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 3ae723d7b1 ✨(backend) add a service to call the yhub REST API
The backend application will have to call the yhub REST API for some
operations. We want to use a dedicated service to do that. This first
commit introduces the shape of this service, it only does the
configuration for now, calling actions will be implemented later.
2026-09-23 12:06:05 +02:00
Kevin JahnsandManuel Raynaud 2f7ac19ec9 ✅(collaboration) test the legacy migrations against a real yhub
Cover both paths off the legacy Django store end to end: the lazy seed on
first access, and the migrate endpoint replaying every S3 version. The tests
need no database — the admin JWT short-circuits document authorization, so a
fixture is an S3 object on a random uuid — and read the timeline through
yhub 0.5.0's `Accept: application/json`, which spares python a lib0 decoder.
CI grows a valkey service and starts a collaboration server alongside the
backend test job; the tests skip themselves when nothing answers on the new
COLLABORATION_API_URL setting, so `make test` without the dev stack still
passes.

Writing them turned up three things worth fixing in the server.

Backend reads now seed too. getAccessType short-circuited on the admin token
before reaching the legacy store, so a server-side read of an unmigrated
document answered with an empty one, and a create-ydoc against it would have
written a second lineage beside the content the first user access was about
to seed in.

Seeding no longer decides access; the backend's answer alone does. A legacy
object that cannot be migrated — it does not decode, or it exceeds the size
we load — opens as a new document instead of denying, since no retry can fix
it and refusing would leave the document unopenable by anyone. The cause is
logged once per attempt with the bucket, key and stack, and every later access
logs that it admitted a caller without migrating.

That made the failure classifier dangerous, so it is inverted. It was an
allowlist of retryable errors — eight socket errnos — which left every way S3
can refuse (AccessDenied on a rotated key, NoSuchBucket, a region redirect)
counting as "this object is unusable". Denying, that was survivable; opening
empty, one misscoped credential would fork every document touched during the
window. Now only a failure raised while interpreting bytes we already hold is
permanent, marked at the throw site, and everything else answers a retryable
503. Guessing wrong that way costs a retry; the other way costs the document.

The admin seed is also fenced to the org and to main, like the user path
above it. The legacy store is branchless — {docid}/file is main — and the
bookkeeping is per document, so seeding ?branch=draft would have written
main's content into an orphan room and left the real one permanently empty.

Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-23 12:06:05 +02:00
Anthony LCandManuel Raynaud 27752a17bd 🛂(backend) add audience to jwt
Add audience to the jwt, scoping the token to it
prevents an admin JWT issued for another backend
service from being replayed against y-provider.
2026-09-23 12:06:05 +02:00
Anthony LCandManuel Raynaud c1944dca90 🛂(django) use jwt token for converter services
The Y_PROVIDER_API_KEY shared secret is replaced by a
signed admin JWT when Django calls the y-provider
conversion endpoint.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 68f5f5cdf7 🔥(backend) remove CollaborationService and can-edit endpoint
The CollaborationService was doing nothing since we started the
migration to yhub, all the code using it is now removed. Also the
`can-edit` endpoint and all the safeguard mechanism relying on the
presence of other users connected to the websocket will not be used
anymore, it will be possible to replace all of this with yhub, so all
this code is also removed.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 14cd463c1b ✨(backend) add a method to create a dedicated admin token
For now the only token we will need is ont with the admin claim set to
True. To not repeat the creation of this token again and again, we
created a dedicated method to issue this token in the JWTService class.
2026-09-23 12:06:05 +02:00
Manuel Raynaud e8421824cc ✨(backend) publish the JWT public key on a JWKS endpoint
The yhub service will need our public key in order to validate the jwt
token we will used. We choose to expose a jwks endpoint as it is a
standard wat to do this.
2026-09-23 12:06:05 +02:00
Manuel Raynaud 3addd04755 ✨(backend) add a service generating cached RS256 JWT tokens
We want to generate jwt token using the RS256 algotrithm. This token
will be used for internal call with the yhub service.
2026-09-23 12:06:05 +02:00
Kevin JahnsandManuel Raynaud b18f0c7bc8 ♻️(collaboration) switch collaboration server from hocuspocus to yhub
Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
2026-09-23 12:06:05 +02:00
Anthony LC d01372fd90 ♻️(backend) return the full document in the duplicate response
The duplicate endpoint used to respond with only `{"id": ...}`. It now
returns the complete duplicated document representation, consistent
with the other document detail endpoints, so the frontend doesn't have
to make a follow-up request to get the new document's data.

This required setting `is_favorite` explicitly on the duplicated
document before serializing it: it is normally set by the
`annotate_is_favorite` queryset method, which the newly created
document never goes through. Being a read-only serializer field, it
was silently dropped from the response instead of raising an error. A
document can't be a favorite right after being created, so it is set
to `False` directly.
2026-09-18 16:17:31 +02:00
risk-altandAnthony LC 5b661d7224 🥅(frontend) warn before uploading a file over the size limit
Dropping a file larger than the allowed size showed a bare "unknown
error" in the editor. The proxy in front of the API cuts the request
and answers a 413 with an HTML body, so errorCauses threw while
parsing it as JSON and no cause ever reached the error panel.

The size limit the backend already enforces is now exposed by the
config endpoint, and the editor checks the file against it before
sending anything, with the same toast wording the document import
uses. errorCauses no longer throws on a body it cannot parse, and a
413 without a usable cause falls back to an explicit message, which
covers the instances whose proxy limit is lower than the application
one.

The size formatting duplicated in the import hook moved to a shared
util.

Signed-off-by: risk-alt <aldu6974@gmail.com>
2026-09-17 16:04:26 +02:00
Anthony LC 147bf68dda 🔖(release) minor 5.7.0
Added:
- 🔧(backend) fine tune redis cache options
- ✨(frontend) make the full last-update date available
- 💄(frontend) redesign email confirmation standalone page

Changed:
- ⬆️(backend) upgrade celery to version 5.6.3
- ⚡️(backend) stop using LEFT(value, LENGTH(path)) in sql queries
- 🚚(project) switch docspec image to ghcr.io/docspec/api
- 🚚(global) move favorite documents API endpoint
  to `/documents/favorites/`

Fixed:
- 🐛(backend) skip session creation for the liveness probe
- 🐛(frontend) preserve page titles when adding an emoji
- 🐛(frontend) scroll to the linked block in read-only documents
- 🐛(frontend) hide the selection highlight on presenter images
- 🐛(y-provider) prevent process crash on malformed websocket frames
- 🐛(frontend) keep commented text sharp when printing to PDF
- 🐛(docker) pull minio images from quay.io
- ♿️(frontend) restore presenter focus trapping after share links
- 🐛(frontend) export any raster image supported by the browser to a PDF
2026-09-15 16:19:39 +02:00
AntoLCandAnthony LC da4f409907 🌐(i18n) update translated strings
Update translated files with new translations
2026-09-15 14:00:47 +02:00
Manuel RaynaudandGitHub d596df9512 ✨(backend) allow configuring trace sampling
Add a configuration knob for the trace sampling rate, so we can
enable tracing on middleware and cache spans when debugging slow
requests in production.
    
Sampling is set to 0 by default, so tracing stays fully off unless
explicitly enabled.
Copied from suitenumerique/meet#1690
2026-09-14 18:59:01 +00:00
Julien MaupetitandGitHub 31cff890b8 🚚(global) move favorite documents API endpoint to /documents/favorites/
To respect the globally used pattern, we can safely switch to a simpler
path
2026-09-14 10:01:36 +00:00
Manuel Raynaud 00cc95aa05 🔧(backend) move the DockerflowMiddleware higher in the middleware list
We decided to move the DockerflowMiddleware higher in the middleware
list to prevent future access to the database or redis in other
middleware that can have an impact on the liveness probe.
2026-09-11 12:55:20 +02:00
Manuel Raynaud 451499016e ⚡️(backend) increase nb_accesses cache TTL
The nb_accesses cache TTL was very short, 30 seconds. That mean that the
user will hit the cache for a very short period and the cache is
probably not be hit. This is what we can see in the slow queries from
the pg_stat_statements table. The query to compute the nb_accesses is
executed a little bit less than the number of queries to list or
retrieve documents, meaning the cache is not used.
2026-09-11 12:55:20 +02:00
Manuel Raynaud 6baf20aaeb ⚡️(backend) improve DocumentViewset.get_queryset
The filtering made in the DocumentViewset.get_queryset method is not
optimal and lead to a full scan of the Document table. The heavy part is
on the filtering on what the user can access between the accesses and
the link traces. To have better performance we make an union operation
of both document_id list and the filter the id on this list. Postgresql
will use the index on the id column.
2026-09-11 12:55:20 +02:00
Manuel Raynaud fbc3ef83ba ⚡️(backend) stop using LEFT(value, LENGTH(path)) in sql queries
Comparing path with LEFT(value, LENGTH(path)) makes a sequential scan on
all the Document table, the more this table grow, the more the query
using it will be slow. We dediced instead to lookup on the path
extracting all ancestors path for a given document and then make a path
IN statement to use the index existing on the path column.
2026-09-11 12:55:20 +02:00
Manuel RaynaudandGitHub 673a670dd9 🔧(backend) configure request.summary logger
The request.summary logger is used by dockerflow. The INFO level is
always empty and is used everytime the liveness or readiness endpoint
are fetch.
2026-09-09 14:46:19 +00:00
Manuel Raynaud 137cecc0e1 🔧(backend) fine tune redis cache options
We want to configure other options on the redis cache. By default there
is no timeout on the connection to socket and no timeout for read/write
operations. We set default values in all caches used in production. The
settings IGNORE_EXCEPTIONS differ between the default and the session
cache. Activating it behaves like a missed cache. Enabling it for the
session should lead to unwanted side effects, by returning falsy on the
session creation, a retry mechanism of 10000 attempts is made in the
SessionStore.create method, the request can stay in this loop for a long
time.
2026-09-08 16:26:15 +02:00
Manuel Raynaud 0b8808f9a7 🐛(backend) skip session creation for the readiness probe
The readiness probe should also not create a new session. A new session
will live in redis and increase the number of keys inside it for
nothing. The readiness path is isgnored in the ForceSessionMiddleware
2026-09-08 14:44:57 +02:00
Manuel Raynaud 36a890a119 ⬆️(backend) upgrade celery to version 5.6.3
Celery version 5.6 has several fixes we want : two significant memory
leaks have been resolved and also a fix allowing a better use of psycopg
pool.
2026-09-08 14:44:56 +02:00
Manuel Raynaud 1e61b4a789 🐛(backend) skip session creation for the liveness probe
The ForceSessionMiddleware force the session creation, we want to
ignore it when the request is the liveness probe. The liveness probe
must not check if redis is available, this is the readiness probe job
2026-09-08 08:56:42 +02:00
Manuel Raynaud 3c1275c88d 🔖(release) patch 5.6.1
Added

- ✨(frontend) export presenter slides as PDF #2487

Fixed

- 🐛(frontend) hide Leave in the doc menu when not logged in #2626
- 🐛(backend) allow to configure settings DATA_UPLOAD_MAX_MEMORY_SIZE
2026-09-04 18:19:06 +02:00
renovate[bot]andGitHub 36bf78558a ⬆️(dependencies) update django to v5.2.16 [SECURITY] 2026-09-04 16:05:54 +00:00
Manuel Raynaud efdac7444d ➕(backend) add servestatic dependency
We removed previously whitenoise because it was not working with asgi
application. By removing it we also removed the way to serve the static
files in the application. There is an existing fork of whitenoise,
servestatic, that manage async application and we can use it to serve
static files.
2026-09-04 16:28:16 +02:00
Manuel Raynaud 3bede0d9a0 🐛(backend) allow to configure settings DATA_UPLOAD_MAX_MEMORY_SIZE
Release 3.17.2 of DRF now takes care of DATA_UPLOAD_MAX_MEMORY_SIZE
and is checked when the body request is parsed. Before that, DRF wasn't
using it at all and we were only looking for custom settings linked to
the media and conversion file upload. We must now also configure this
setting.
2026-09-04 10:21:43 +02:00