Some networks refuse a websocket upgrade - corporate proxies, captive portals - and a browser is told nothing more than "the connection closed", so those users could not edit at all. The editor now runs a second transport next to the socket, polling the collaboration server's REST api on the same room, with the same session cookie and the same authorization, and only while the socket is down. Local changes go out about a second after the last keystroke and remote ones arrive within ten seconds, so editing works with visibly more latency rather than not at all. The socket keeps being retried underneath, so a client that fell back during an outage returns to it on its own, and nothing is lost in either direction - both transports publish from the same document. This makes /collaboration/ydoc/ a route browsers call, so COLLABORATION_SERVER_ORIGIN is now handed to yhub as its cors configuration and gates the http routes as well as the websocket. Signed-off-by: Kevin Jahns <kevin.jahns@protonmail.com>
6.0 KiB
Collaboration
By default with Docs, collaboration is enabled. To allow the collaboration between users, a connection to a websocket server is made (the yhub service), you only have to configure the Django backend URL and the allowed origin in your yhub service:
COLLABORATION_BACKEND_BASE_URL: https://{yourdocsdomain.tld}
COLLABORATION_SERVER_ORIGIN: https://{yourdocsdomain.tld}
The collaboration server keeps the live state of a document in Redis and persists it to a PostgreSQL database of its own, so it needs both:
REDIS: redis://{redis-host}:6379/0
POSTGRES: postgres://{user}:{password}@{postgres-host}:5432/yhub
Nothing creates that schema at startup: the server never runs DDL. Run the script yhub ships (npm run init-db, which the helm chart runs as a job) once before starting it, and again after every upgrade that adds a table. It creates the database when it is missing, it is idempotent, and until it has run every document read fails with relation "..." does not exist.
The Django backend reads and writes document content there too, so point it at the service:
YHUB_API_BASE_URL: http://{yhub-service}:443
Prefer the internal service url: the routes the backend calls are not meant to be reachable from the outside. Route /collaboration/ws/ to the service publicly — that is the one the browsers open — plus the document routes (/collaboration/ydoc/, rollback, prune, changeset, activity) and /collaboration/jwks/, which carries public keys and nothing else. Keep reset-connections, migrate, restore-ydoc, reset-ydoc and create-ydoc in-cluster.
Both directions are authenticated with short-lived RS256 JWTs rather than a shared secret, and each side verifies the other against the JWKS it publishes — so both need a signing key of their own, and neither needs a copy of the other's:
# Django
JWT_PRIVATE_KEY_FILE: /path/to/backend-private.pem
# yhub
YHUB_JWT_PRIVATE_KEY_FILE: /path/to/yhub-private.pem
They are ordinary PKCS#8 RSA keys (openssl genpkey -algorithm RSA -pkeyopt rsa_keygen_bits:2048), and rolling one needs no change on the other side. Without them the documents still open and edit, but the backend cannot create, delete or restore a document's content, and yhub cannot tell it that a document changed — its updated_at stops following the edits.
Generating them on the cluster
The helm chart generates both for you, so that no key has to be created by hand, put in a values file or in a secret:
jwtKeys:
enabled: true
A job then creates the two keys, once, in a secret every service mounts read-only, and points the backend and yhub at them. It generates them with openssl in a pod-local volume and hands them to kubectl create secret, so they never touch a disk, a manifest or a values file. The secret is left alone when it is already there, so the job is safe to re-run — it runs on every sync — and rolling the keys is deleting the secret and letting the next run create it again. Both sides follow: they pick the verification key by its kid and fetch the set again when they meet one they do not know.
The job is the only thing allowed near that secret: the chart gives it a service account whose role can create a secret and read whether that one exists, nothing more. The services never call the kubernetes API — they read a mounted file. The secret is not part of the release either, so uninstalling keeps the same identities; delete the secret to start over.
Deployments already holding their keys in a secret of their own point the chart at it instead, and the job and its rights are not created at all:
jwtKeys:
enabled: true
existingSecret: my-jwt-keys # holding private.pem and yhub-private.pem
Setting JWT_PRIVATE_KEY_FILE or YHUB_JWT_PRIVATE_KEY_FILE yourself keeps priority over what the job provides, so a deployment holding its keys in a secret of its own can leave jwtKeys disabled and mount them where it wants.
Several replicas can serve the same document: they exchange updates through Redis, so no sticky routing is needed on the websocket ingress.
What happens when connection to the websocket is not allowed?
Some networks refuse a websocket upgrade — corporate proxies, captive portals — and a browser is told nothing more than "the connection closed". For those clients the editor falls back to polling the collaboration server over plain http, on the same room, with the same session cookie and the same authorization. Nothing has to be configured: the fallback is installed next to the websocket and only ever sends a request while the socket is down.
That means /collaboration/ydoc/ has to be routed publicly, not only in-cluster — the browsers of
those users call it directly. And the origins a browser may reach the server from are the ones in
COLLABORATION_SERVER_ORIGIN, which now gate the http routes as well as the websocket:
COLLABORATION_SERVER_ORIGIN: https://{yourdocsdomain.tld}
A comma-separated list is allowed, and each entry is a bare origin — https://host[:port], no path
and no trailing slash. A deployment serving the frontend from another origin than the collaboration
server has to list it here or the fallback is refused, the same way the websocket already is.
What the fallback does not do is hide the difference. It publishes local changes about a second after the last keystroke, and it retrieves the document every ten seconds, so someone else's edits arrive with up to that much delay and remote cursors move at poll resolution. Each round transfers the whole document, so a large document polled by many clients is real egress. It is a way to keep editing, not a replacement for the socket — and the socket keeps being retried underneath, so a client that fell back during an outage returns to it on its own.
Documents are never in conflict either way: both transports publish from the same Yjs document, and Yjs merges. Before the fallback existed, users who could not open a websocket edited a document that was saved wholesale and erased each other's modifications; that is what this removes.