Practice — Distributed Transactions, Consensus & Coordination (6 questions)
Diagnosing a 2PC Blocking Incident Permalink →
A payments platform uses XA-style two-phase commit across an orders
database and a ledger database, coordinated by a single transaction
manager. During a deploy, the transaction manager process is killed
after collecting YES votes from both participants but before it sends
the COMMIT decision. On-call notices orders rows locked and new
transactions queuing behind them.
- Explain precisely why the participants cannot resolve this on their own.
- What are the two possible correct resolutions once the coordinator comes back, and what determines which one is correct?
- Propose two concrete changes to the deployment/architecture that would prevent this class of incident going forward.
Share this question
Why Byzantine Fault Tolerance Needs an Extra Node Permalink →
Your Raft-based configuration store tolerates 1 crashed node using a cluster of 3 (majority quorum, N = 2f+1 for f=1). A separate team is building a multi-party settlement ledger where the participating nodes are run by different companies — one of them might not just crash, but could actively send conflicting or false messages to different peers (a Byzantine fault, not a crash fault).
To tolerate f=1 Byzantine (actively malicious or arbitrarily faulty) node using a classic BFT protocol like PBFT, how many total nodes are required?
Share this question