When Is an Order Safely Recorded?
Once Shopend confirms a purchase, it must preserve the order. With one database server, a purchase saved durably on disk can survive a server crash. Assuming the disk remains intact, we wait for the server to restart and recover the purchase.
Failover lets us resume work without waiting for that server. But even the replica furthest along may not have received the last purchase.
Confirmed, but missing after failover
Suppose the following happens:
| Step | What happens |
|---|---|
| 1 | The primary records a purchase. |
| 2 | Shopend tells the shopper it succeeded. |
| 3 | The primary fails before passing the changes to a replica. |
| 4 | A replica takes over as the new primary. |
The new primary does not have the purchase. The shopper received a successful response, but the database now serving Shopend has no record of that order. The old primary’s disk might still hold it, but that does not make it part of the new primary’s data.
Saving the purchase on the primary was enough to recover it after restarting that server. It was not enough to preserve it when another server took over. We need to decide what must happen before Shopend confirms success.
Asynchronous replication
With asynchronous replication, the primary reports success without waiting for the other replicas to record the changes:
sequenceDiagram
participant A as Shopend API
participant P as Primary
participant R as Replica
A->>P: Write purchase
P->>P: Durably record purchase
P-->>A: Success
P->>R: Replicate changes
R->>R: Durably record changes
R-->>P: Acknowledgment
Durably recording the changes means saving them so they can be recovered after a restart. The replica sends an acknowledgment when it has done so. In this example, the primary has already reported success by then.
Replication can start before the primary reports success. What makes it asynchronous is that success does not wait for the replica’s acknowledgment. This avoids that wait, but a confirmed purchase may still exist only on the primary when it fails.
Synchronous replication
With synchronous replication, the primary waits for the required replica acknowledgments before reporting success. Here is the same purchase with one replica required to acknowledge it:
sequenceDiagram
participant A as Shopend API
participant P as Primary
participant R as Replica
A->>P: Write purchase
P->>P: Durably record purchase
P->>R: Replicate changes
R->>R: Durably record changes
R-->>P: Acknowledgment
P-->>A: Success
Now the purchase is recorded on both servers before Shopend confirms it. Losing the primary still leaves a copy of the purchase. The failover policy must preserve those acknowledged changes when choosing a replacement; it cannot select a replica that lacks them and discard the copy that has them.
Waiting has a cost. The response takes longer, and if too few of the required replicas respond, the primary cannot confirm success under this policy. Keeping additional copies protects the purchase against losing a server, but the copies must actually record it before we rely on that protection.
When Shopend confirms a write
For a purchase, we wait for a replica to durably record the transaction’s changes before confirming success. That includes the order, the stock reductions, and the cart update. We accept the wait because losing an accepted purchase would break our promise to the shopper.
For an ordinary cart edit, we choose to confirm after the primary records it. If that change is lost during failover, the shopper may have to add an item again or change how many they want to buy. For this version of Shopend, we accept that possibility in exchange for avoiding the replica wait on those edits.
The acknowledgment choice can be made for each operation or transaction. It does not have to be the same for every write in the database.