This allows us to keep the full history of semantically important events, while not spamming the history and avoiding excessive use of disk space. See issues #370 and #423. This is a substantial change. Highlights follow: * Added a developer documentation section * No longer using events for manifest/crl generation (#370) * No longer using events for publication deltas (#423) * Removed pre 0.6.0 migration code - people will have to upgrade to at least 0.6.0 first * Added migration code for 0.6.0-0.8.1 to this * Migrate repository by doing a keyroll. (#370) * Remove archiving code for commands (no longer applicable) Minor other fixes: * Use a swap file when writing (avoid corrupt json if disk is full) (#370) * Make removing publisher content idempotent for publishers already removed. Co-authored-by: Ximon Eighteen <3304436+ximon18@users.noreply.github.com> Co-authored-by: Jasper den Hertog <jasper@plainspace.com>
3.1 KiB
Event Sourcing Overview
The Audit is the Truth
Event Sourcing is a technique based on the concept that the current state of a thing is based on a full record of past events - which can be replayed - rather than saving / loading the state using something like object-relational-mapping or serialization.
A commonly used example to explain this concept is bookkeeping. In bookkeeping one tracks all financial transactions, like money coming in and going out of an account. The current state of an account is determined by the culmination of all these changes over time. Using this approach, and assuming that no one altered the historical records, we can be 100% certain that the current state reflects reality. Conveniently, we also have a full audit trail which can be shown.
CQRS and Domain Driven Design
Event Sourcing is often combined with CQRS: Command Query Responsibility Segregation. The idea here is that there is a separation between intended changes (commands) sent to your entity, the internal state of the entity, and its external representation.
Separating these concerns we can also borrow heavily from "Domain Driven Design". In DDD structures are organized into so-called of "aggregates" of related structures. At their root they are joined by an "aggregate root".
Commands, Events and Data
This separation means that the aggregate is a bit like a black box to the outside world, but in a positive sense. Users of the code just need to know what messages of intent (commands) they can send to it. This interface, like an API, should be fairly stable.
The first thing to do when a command is sent is that an aggregate is retrieved from storage, so that it can receive a command. The state of the aggregate is the result of all past events - but, because this can get slow, aggregate snapshots are often used for efficiency. If a snapshot does not include the latest events, then they are simply re-applied to it.
When an aggregate receives a command it can see if a change can be applied. It is fully in charge of its own consistency. If it isn't then the aggregate root is usually at the wrong level. The result of the command can either be an error - the command is rejected - or a number of events which represent state changes are applied to the aggregate.
When events are applied they are also saved, so that they can be replayed later. Furthermore, implementations often use message queues where events are posted as well. This allows other components in the code to be triggered by changes in an aggregate.
In its purest form (hint: we don't do this), these messages can then be used to (re-)generate one or more data representations that users can query. E.g. you could populate data tables in a SQL database if that floats your boat.
Credits / Read More..
This combination of techniques has been championed by various people, most notably Greg Young and Martin Fowler. You can do your own internet search find out much more about how this can work, and how it is done in other projects.