Built
Process Pair
Jim Gray's 1985 process pair rebuilt on one laptop and killed 420 times to measure the checkpoint tradeoff he described and never tested. Checkpointing every request costs 25% of throughput, lost work is N/2, and the 1974 formula for the best interval lands 20x too high because it assumes lost work gets recomputed.
Why: Built as the first stop in rederiving fault tolerance from the original papers. Two things I did not expect: Gray's own fix, client replay, closes the lost-work hole for free, while a durable log costs 99% of throughput because fsync() on macOS does not wait for the disk. And my first takeover result was wrong: macOS AirPlay listens on port 5000 and was accepting the client's reconnects, which is the gray-failure problem in miniature.