Database — NoSQL & SQL vs NoSQL¶
What is NoSQL?¶
NoSQL is a collection of data items represented as a key-value store, document store, wide column store, or graph database.
- Data is denormalized; joins are done in application code.
- Most NoSQL stores lack true ACID transactions and favor eventual consistency.
BASE (contrast with ACID)¶
BASE describes NoSQL properties and corresponds to choosing availability over consistency in the CAP theorem:
| Letter | Meaning |
|---|---|
| Basically available | The system guarantees availability. |
| Soft state | State may change over time, even without input. |
| Eventual consistency | The system becomes consistent over time (given no new input). |
Types of NoSQL databases¶
Key-Value Store¶
- Abstraction: hash table.
- O(1) reads/writes, often backed by memory or SSD.
- Keys can be kept in lexicographic order for efficient range retrieval.
-
High performance; used for simple data models or rapidly-changing data (e.g., in-memory cache layer).
-
Only a limited set of operations — extra complexity shifts to the app.
- Basis for document stores (and some graph databases).
- Examples: Redis, Memcached.
Document Store¶
- Abstraction: key-value store with documents (XML, JSON, binary) as values.
-
A document stores all info for a given object; query by the document's internal structure.
-
Documents can have completely different fields from one another.
- Examples: MongoDB, CouchDB; DynamoDB supports both KV and documents.
- High flexibility; good for occasionally-changing data.
Wide Column Store¶
-
Abstraction: nested map
ColumnFamily<RowKey, Columns<ColKey, Value, Timestamp>>. -
Basic unit is a column (name/value pair); columns group into column families (like a SQL table).
-
Access each column independently by row key; same row key = a row.
- Each value has a timestamp for versioning/conflict resolution.
-
Examples: Bigtable (Google, the first), HBase (open source, Hadoop ecosystem), Cassandra (Facebook).
-
High availability & scalability; good for very large datasets.
Graph Database¶
- Abstraction: graph.
- Each node is a record; each arc is a relationship between nodes.
- Optimized for complex relationships (many foreign keys / many-to-many).
- Example: social networks.
- Relatively new, fewer tools/resources; many are only accessible via REST APIs.
- Examples: Neo4j, FlockDB (Twitter).
Summary table¶
| Type | Abstraction | Examples | Best for |
|---|---|---|---|
| Key-value | Hash table | Redis, Memcached | Caches, simple/fast data |
| Document | KV + documents | MongoDB, CouchDB, DynamoDB | Flexible, semi-structured data |
| Wide column | Nested map | Bigtable, HBase, Cassandra | Very large datasets |
| Graph | Graph | Neo4j, FlockDB | Complex relationships |
SQL or NoSQL?¶
Choose SQL when¶
- Structured data
- Strict schema
- Relational data
- Need complex joins
- Need transactions (ACID)
- Clear patterns for scaling
- More established (developers, community, tools)
- Fast lookups by index
Choose NoSQL when¶
- Semi-structured data
- Dynamic/flexible schema
- Non-relational data
- No need for complex joins
- Storing many TB/PB of data
- Very data-intensive workloads
- Very high throughput for IOPS
Sample data well-suited to NoSQL¶
- Rapid ingest of clickstream/log data
- Leaderboard or scoring data
- Temporary data (e.g., shopping cart)
- Frequently accessed ("hot") tables
- Metadata / lookup tables
Key takeaways¶
-
ACID (SQL) prioritizes correctness/transactions; BASE (NoSQL) prioritizes availability/scale.
-
Pick the right NoSQL type for the data model, not just "NoSQL" generically.
- When in doubt: SQL for transactions and relations, NoSQL for scale and flexibility.
Common Interview Questions¶
Design a key-value store like Redis
Why it's asked here: a KV store is the foundational NoSQL type.
Key points to discuss:
- Hash-table abstraction, O(1) reads/writes, in-memory with optional persistence.
- Data structures (lists, sets, sorted sets) beyond plain strings.
- Eviction (LRU) and TTL for bounded memory.
- Consistency model (BASE/eventual) and why ACID isn't required.
flowchart LR
C[Client] --> K[KV store]
K -->|hash| N1[Node 1]
K -->|hash| N2[Node 2]
K -->|hash| N3[Node 3]
Design a social network's data structures
Why it's asked here: social graphs map naturally to a graph database.
Key points to discuss:
- Nodes = users, edges = relationships (follow, friend).
- Why a graph DB (Neo4j/FlockDB) beats relational joins for deep traversal.
- Traversal queries (friends-of-friends) and their performance.
- When to use adjacency lists vs a graph store.
flowchart LR
U[User node] -->|follows| F[Followed node]
U -->|friend| FR[Friend node]
G[Graph DB] --> U
G --> F
G --> FR
Design a recommendation system
Why it's asked here: recommendation data (user-item interactions) fits wide-column/ document stores.
Key points to discuss:
- Store user-item interaction history (wide-column, e.g., Cassandra).
- Document store for item metadata; KV for counters/lookups.
- Offline batch (MapReduce/Spark) vs online serving paths.
- BASE/eventual consistency for precomputed recommendations.
flowchart LR
E[User-item events] --> W[(Wide-column store)]
W -->|batch| B[MapReduce / Spark]
B --> R[(Recs store)]
R --> U[User]