## Summary The sqlite delta history silently drops a parent checkpoint whose id sorts above its child's, losing that parent's stored value and its pending writes. The channel hydrates short with no error. Fixes #8550 ## Problem Stage 1 walked ancestors with: ```sql WHERE thread_id = ? AND checkpoint_ns = ? AND checkpoint_id <= ? ORDER BY checkpoint_id DESC ``` Ancestry is defined by the `parent_checkpoint_id` column. These two predicates add a second requirement: that every child's id sorts above its parent's. The contract promises monotonic ids, but that only holds within one process, so ids from processes with different clocks can break it. When the requirement is violated the parent is excluded from the stream and its seed and writes go with it. Dropping the range filter alone does not fix it: in `checkpoint_id DESC` order that parent arrives *before* the target, so the walk streams past it before it has started. ## Fix A recursive CTE anchored at the target, following `parent_checkpoint_id`: ```sql WITH RECURSIVE ancestors(checkpoint_id, parent_checkpoint_id, type, checkpoint) AS ( SELECT ... FROM checkpoints WHERE thread_id = ? AND checkpoint_ns = ? AND checkpoint_id = ? UNION ALL SELECT c.... FROM ancestors a CROSS JOIN checkpoints c ON c.checkpoint_id = a.parent_checkpoint_id WHERE c.thread_id = ? AND c.checkpoint_ns = ? ) SELECT checkpoint_id, type, checkpoint FROM ancestors ``` Rows now arrive in walk order (target, parent, grandparent, ...), so `step_walk_with_row` no longer needs its off-path skip or its `parent_cid` tracking; both are removed. The query reads only true ancestors, where the old one read every row at or below the target including sibling branches. `CROSS JOIN` pins the join order. The saver never runs `ANALYZE`, and with a plain `JOIN` sqlite put `checkpoints` as the outer loop, scanning the whole thread on every recursion step. With `ancestors` outside, each step is one primary key lookup. Through `get_delta_channel_history`: | chain length | plain `JOIN` | `CROSS JOIN` | | -- | -- | -- | | 1000 | 0.032s | 0.001s | | 2000 | 0.124s | 0.003s | | 4000 | 0.475s | 0.006s | ## Cycle guard Following pointers can loop where a bounded id scan could not, and a loop is reachable through `put` alone: `put` writes with `INSERT OR REPLACE`, so re-putting an existing checkpoint id under a descendant's config repoints that checkpoint at its own descendant. The walk stops on a repeated `checkpoint_id` (one set insert per row, no depth ceiling that could truncate a long migrated thread). sqlite yields recursive rows lazily, so abandoning the cursor ends the recursion. `test_walk_terminates_when_put_makes_the_parent_chain_cycle` fails by hanging, not by asserting, if the guard regresses (confirmed by deleting the guard locally). The package has no `pytest-timeout`, so the CI job timeout is the backstop. ## Postgres No equivalent change needed. It pages the whole thread with no id bound and follows parent pointers in Python, and its upsert never rewrites `parent_checkpoint_id`, so it can neither miss this parent nor form the loop. `BaseCheckpointSaver` and `InMemorySaver` also walk parent pointers. ## Test plan New `libs/checkpoint-sqlite/tests/test_delta_parent_walk.py`: - [x] Sync and async, parametrised over both id orders; the sync case also asserts equality with `BaseCheckpointSaver` on the same rows. `parent_id_sorts_above_child` is the bug, `parent_id_sorts_below_child` the control. - [x] `test_walk_reaches_root_of_long_chain_with_descending_ids`: 40 checkpoints, only stored value at the root. - [x] `test_walk_terminates_when_put_makes_the_parent_chain_cycle`. - [x] `test_walk_step_looks_up_the_parent_by_primary_key`: asserts the recursive step's `EXPLAIN QUERY PLAN` is a key lookup, so a plain `JOIN` can't come back. Fails with it. - [x] On `main`: 3 of the 6 walk tests fail (both `parent_id_sorts_above_child` cases and the long chain). The cycle test passes on `main` too, since the old bounded scan could not loop; it guards the new path. - [x] #8550's repro returns `{'writes': [('task', 'ch', 'write-root')], 'seed': 'seed'}` sync and async (was `{'writes': []}` on `main`). - [x] `libs/checkpoint-sqlite`: `make format`, `make lint` clean; full suite 125 passed, 2 skipped. - [x] `libs/langgraph`: `-k "delta or sqlite"` 739 passed, 1 skipped. Thanks to @lylelllll for the report, the minimal repro, the base-saver comparison that isolated it to the fast path, and for suggesting the recursive CTE. Co-authored-by: lylelllll <59271327+lylelllll@users.noreply.github.com>
LangGraph SQLite Checkpoint
To help you ship LangGraph apps to production faster, check out LangSmith. LangSmith is a unified developer platform for building, testing, and monitoring LLM applications.
Quick Install
uv add langgraph-checkpoint-sqlite
🤔 What is this?
This library provides a SQLite implementation of LangGraph's checkpoint saver, with both sync and async support via aiosqlite. Use it when you want LangGraph state persistence backed by SQLite for local development, testing, or lightweight deployments.
📖 Documentation
For full documentation, see the API reference. For conceptual guides on persistence and memory, see the LangGraph Docs.
Security
Important
Set
LANGGRAPH_STRICT_MSGPACK=trueor pass an explicitallowed_msgpack_moduleslist when creating your checkpointer. This restricts checkpoint deserialization to known-safe types, preventing code execution if the database is compromised. See the langgraph-checkpoint README for details.
Usage
from langgraph.checkpoint.sqlite import SqliteSaver
write_config = {"configurable": {"thread_id": "1", "checkpoint_ns": ""}}
read_config = {"configurable": {"thread_id": "1"}}
with SqliteSaver.from_conn_string(":memory:") as checkpointer:
checkpoint = {
"v": 4,
"ts": "2024-07-31T20:14:19.804150+00:00",
"id": "1ef4f797-8335-6428-8001-8a1503f9b875",
"channel_values": {
"my_key": "meow",
"node": "node"
},
"channel_versions": {
"__start__": 2,
"my_key": 3,
"start:node": 3,
"node": 3
},
"versions_seen": {
"__input__": {},
"__start__": {
"__start__": 1
},
"node": {
"start:node": 2
}
},
}
# store checkpoint
checkpointer.put(write_config, checkpoint, {}, {})
# load checkpoint
checkpointer.get(read_config)
# list checkpoints
list(checkpointer.list(read_config))
Async
from langgraph.checkpoint.sqlite.aio import AsyncSqliteSaver
async with AsyncSqliteSaver.from_conn_string(":memory:") as checkpointer:
checkpoint = {
"v": 4,
"ts": "2024-07-31T20:14:19.804150+00:00",
"id": "1ef4f797-8335-6428-8001-8a1503f9b875",
"channel_values": {
"my_key": "meow",
"node": "node"
},
"channel_versions": {
"__start__": 2,
"my_key": 3,
"start:node": 3,
"node": 3
},
"versions_seen": {
"__input__": {},
"__start__": {
"__start__": 1
},
"node": {
"start:node": 2
}
},
}
# store checkpoint
await checkpointer.aput(write_config, checkpoint, {}, {})
# load checkpoint
await checkpointer.aget(read_config)
# list checkpoints
[c async for c in checkpointer.alist(read_config)]
📕 Releases & Versioning
See our Releases and Versioning policies.
💁 Contributing
As an open-source project in a rapidly developing field, we are extremely open to contributions, whether it be in the form of a new feature, improved infrastructure, or better documentation.
For detailed information on how to contribute, see the Contributing Guide.