Papershelf
I read papers around storage engines, distributed systems and vector databases. The list is grouped by topic so I can follow a thread instead of working through one very long queue.
✓ means read. ✗ means still to read.
Suggested reading paths
- Storage: The Google File System → LSM-Tree → Bigtable → WiscKey → Spanner
- Consensus: Time and clocks → distributed snapshots → Raft → Paxos → ZooKeeper
- Replication: Epidemic algorithms → CRDTs → Dynamo → TAO → consistency trade-offs
- Distributed processing: MapReduce → Borg → Quincy → Dataflow → Dapper
Distributed-systems foundations
- ✓ The Google File System
- ✗ Time, Clocks, and the Ordering of Events in a Distributed System
- ✗ Timestamps in Message-Passing Systems
- ✗ Logical Physical Clocks and Consistent Snapshots
- ✗ Distributed Snapshots: Determining Global States
- ✗ Linearizability: A Correctness Condition for Concurrent Objects
- ✗ Impossibility of Distributed Consensus with One Faulty Process
- ✗ Brewer’s Conjecture and the Feasibility of Consistent, Available, Partition-Tolerant Web Services
- ✗ A Critique of the CAP Theorem by Martin Kleppmann
Consensus and coordination
- ✗ In Search of an Understandable Consensus Algorithm
- ✗ The Part-Time Parliament
- ✗ Paxos Made Simple
- ✗ Viewstamped Replication Revisited
- ✗ Practical Byzantine Fault Tolerance
- ✗ Elections in a Distributed Computing System
- ✗ ZooKeeper: Wait-free Coordination for Internet-Scale Systems
- ✗ The Chubby Lock Service for Loosely-Coupled Systems
Replication and consistency
- ✗ Epidemic Algorithms for Replicated Database Maintenance
- ✗ Conflict-Free Replicated Data Types
- ✗ A Comprehensive Study of Convergent and Commutative Replicated Data Types
- ✗ Dynamo: Amazon’s Highly Available Key-value Store
- ✗ TAO: Facebook’s Distributed Data Store for the Social Graph
- ✗ Consistency Tradeoffs in Modern Distributed Database System Design
Storage engines and indexing
- ✗ Ceph: A Scalable, High-Performance Distributed File System
- ✗ Organization and Maintenance of Large Ordered Indexes by Rudolf Bayer and Edward M. McCreight (1972)
- ✗ The Log-Structured Merge-Tree (LSM-Tree) by Patrick O’Neil, Edward Cheng, Dieter Gawlick and Elizabeth O’Neil (1996)
- ✗ ARIES: A Transaction Recovery Method Supporting Fine-Granularity Locking and Partial Rollbacks
- ✗ Bigtable: A Distributed Storage System for Structured Data
- ✗ WiscKey: Separating Keys from Values in SSD-Conscious Storage
Distributed databases and transactions
Distributed processing and infrastructure
- ✗ MapReduce: Simplified Data Processing on Large Clusters
- ✗ The Dataflow Model
- ✗ Resilient Distributed Datasets
- ✗ Kafka: A Distributed Messaging System for Log Processing
- ✗ Large-Scale Cluster Management at Google with Borg
- ✗ Quincy: Fair Scheduling for Distributed Computing Clusters
- ✗ Dapper, a Large-Scale Distributed Systems Tracing Infrastructure
Hashing, lookup and load balancing
- ✗ Consistent Hashing and Random Trees
- ✗ Chord: A Scalable Peer-to-Peer Lookup Service
- ✗ Kademlia: A Peer-to-peer Information System Based on XOR Metric
- ✗ The Power of Two Choices in Randomized Load Balancing
- ✗ The Anatomy of a Large-Scale Hypertextual Web Search Engine