# Why is Redis so Fast despite being (mostly) single-threaded?

# What is Redis?

Redis is one of the most popular and versatile data-stores today, favored for its speed and simplicity. Redis offers features that resemble common data structures, such as strings, lists, hashes, sets, sorted sets, streams, and even geospatial indexes. Redis works by storing key-value pairs, such that a key names one value, and the value can be one of the aforementioned structures. Think of it like a directory, where you use a name (the key) to locate information about the person (the value).

For example, you could use a Redis hash to store user data by using a `user_id` key, then the value would be an object with hash fields mapped to values:

```json
"user_123" : {
    "name": "Bob", 
    "age": 40, 
    "email": "bob@gmail.com"
  }
```

# What makes Redis so Fast?

## Leveraging Physical Attributes

While a variety of factors contribute to Redis’s speedy performance, the first one to understand is in-memory storage. All data in Redis lives in RAM; disk is used only for persisting data, never to serve a read. Think of RAM as your desk: limited space, but everything on it is within arm's reach. Meanwhile, disk storage is the filing room down the hall – far more capacity, but retrieving an item means getting up and walking.

To help quantify this advantage, DRAM access is ~100 ns, NVMe SSD is ~20-100 μs, and spinning disk (HDD) 5-10 ms. That makes DRAM ~200-1000x faster vs. NVMe and ~50,000-100,000x compared to HDD.

While tempting to conclude that Redis outperforms disk-backed databases like Postgres simply because it reads from RAM while the latter reads from disk, this doesn’t reflect how modern databases operate. Indeed, databases contain their own buffer pools (RAM managed directly by the DB) that store recently-accessed pages. Similarly, the [kernel](https://en.wikipedia.org/wiki/Kernel_\(operating_system\)) itself retains recently-read file blocks in RAM automatically, which the database can utilize. Taken together, even disk-backed databases do not primarily read from disk.

The real difference is the overhead required to *support* a disk representation, which is costly even when the disk is never read. A conventional database must parse and plan the query, consult a buffer manager to locate and pin the page, evaluate [MVCC visibility](https://www.geeksforgeeks.org/dbms/what-is-multi-version-concurrency-control-mvcc-in-dbms/), and deserialize the tuple from its on-disk format. [Harizopoulos et al](https://dl.acm.org/doi/10.1145/1376616.1376713). (2008) benchmarked a conventional database with its data already fully cached in memory, and found that buffer management, latching, locking, and recovery consumed the large majority of instructions, leaving only a small fraction as useful query work. Redis has no on-disk format to translate from, so the structure it queries is the structure in memory.

Thus, our earlier analogy needs an amendment. A conventional database *does* keep frequently accessed pages on the desk; however, those pages are copies from the filing room, and the office enforces rules about them. Each time you want one, someone has to work out which file you actually need, look up where its copy is being held, and confirm that the version you have is the one you're supposed to read. By contrast, Redis's data is not a copy of anything. There is no filing room, no index, and no checkout procedure. What you see on the desk is the original, and only one.

![](https://cdn.hashnode.com/uploads/covers/68c19865376514b2a5927314/6696048e-ea00-4c4f-9521-d3b4851c6c08.png align="center")

While operating purely from memory contributes to Redis's speed, that alone does not explain what sets it apart; Redis also makes the most of low-level data structures.

## Efficient Data Storage

Redis [stores data incredibly efficiently](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/memory-optimization/), with small values being encoded in compact memory formats: listpack (which replaced ziplist in Redis 7.0), intset, and quicklist, that improve CPU cache locality. Hashes and sets are automatically converted to hash tables, lists become quicklists, and sorted sets become skiplists when exceeding the configured max size ([Object Encoding docs](https://redis.io/docs/latest/commands/object-encoding/)).

To understand the advantage of these compact memory formats, we first must explore some computer architecture. Random access memory (RAM) is a massive array of bytes. Each byte has a number representing its position in the array, which also serves as the address. When the CPU runs a load instruction, it requests a small number of bytes at a specific address (eight, in the case of a pointer). However, the smallest unit that can be transferred from memory to the CPU is a 64-byte cache line. Memory is pre-divided into 64-byte blocks at fixed boundaries (0-63, 64-127, 128-191, …), so whichever block the address falls into is the one you get, along with everything else inside it.

Between CPU and RAM there are several levels of smaller and faster cache: L1 on the core, L2 beside it, and L3 shared across the chip. Each level is larger and slower than the last – ~1 ns for L1, a few nanoseconds for L2, ~15 ns for L3, and ~100 ns for main RAM. Indeed, a single trip to RAM amounts to about a hundred L1 hits, meaning that even once data is in memory where it sits can influence performance.

![](https://cdn.hashnode.com/uploads/covers/68c19865376514b2a5927314/cacca702-678d-4be4-a7aa-847869676991.png align="center")

Now consider what happens when Redis reads a field from each structure. With a compact memory format such as a listpack, data is stored contiguously, meaning each field exists immediately after the previous. The first access costs a full trip to RAM, returning 64 bytes (most of the structure for a small collection). Since the scan moves forward through memory predictably, the CPU’s prefetcher proactively pulls in subsequent cache lines. Once the bytes are in L1, comparing each field costs ~1 ns, so twenty comparisons add ~20 ns.

On the other hand, a hash table stores its bucket array, entries, its key strings, and its value as separate allocations scattered across the heap. Each hop must be completed before the next address is available, making the four resulting trips to RAM dependent and non-overlapping. Thus, at small sizes, the O(N) scan outperforms the O(1) lookup.

However, the four trips involved in the hash table fetch are constant – they do not grow with the number of fields whereas the listpack’s scan does. After a certain size, the scan cost exceeds the cost of the RAM trips, and the structure stops fitting in the cache, making the comparisons themselves more expensive. Redis resolves this by automatically converting the compact structures into hash tables once they grow past the configured threshold.

![](https://cdn.hashnode.com/uploads/covers/68c19865376514b2a5927314/5bfa07c2-0bcd-4069-bb52-08441a76443f.png align="center")

In addition to clever utilization of data structures, architectural decisions further contribute to Redis's notable performance.

## Single Threaded Architecture

Redis employs single-threaded command execution, avoiding the complexity and overhead associated with shared-state multithreaded systems: context switching, thread scheduling, lock contention, and even cache lines moving between cores. This design is viable because CPU is rarely the bottleneck for Redis, usually it is memory or network ([FAQ](https://redis.io/docs/latest/develop/get-started/faq/)). Since typical Redis commands are very small, adding the cost of coordination (which is roughly the same regardless of the operation size) becomes a bad tradeoff.

Of note, Redis is not strictly a single-threaded process – modern versions use threads for a variety of supporting functions, such as closing files, cleaning memory, or reading/writing to client sockets. Threaded I/O, introduced in Redis 6, parallelizes socket reads, writes and protocol parsing ([`redis.conf`](https://github.com/redis/redis/blob/unstable/redis.conf), THREADED I/O). However, *only* the main thread executes commands that touch the global keyspace, such as lookups, mutations, and triggering expiry/eviction to name a few. This architecture has the added benefit of preserving atomicity, since only a single thread is modifying Redis’s in-memory data at a given time.

What enables Redis to handle thousands of clients with a single thread is [multiplexing and non-blocking I/O.](https://redis.io/docs/latest/develop/reference/clients/) Every client connection is a socket, and the kernel holds a receive buffer for each one.

Redis sets these sockets to non-blocking, so read and write operations return immediately if nothing is ready rather than blocking the thread. It then registers all of them with an event notification interface such as [`epoll`](https://en.wikipedia.org/wiki/Epoll), asking to be notified when one of them has something ready for it, then sleeps. A socket is *ready* when a `read` or `write` wouldn't halt execution: readable means that at least one byte has arrived, writeable means there's room in socket's send buffer. Because the kernel is the one putting arriving bytes into the buffers, it can add the socket to a ready list at that moment. When data arrives for one of those sockets, the kernel wakes Redis, which checks the short list of ready descriptors, runs the handler registered for each, and goes back to sleep. As a result, the thread never blocks on any individual client – it only pauses when no client has anything ready.

A helpful analogy is a restaurant with forty tables (client connections) and a single waiter (the Redis thread). If the waiter were to approach the table and wait until everyone was ready to order, this prevents the waiter from being able to attend other tables while they deliberate. Non-blocking here translates to the waiter asking first if a table is ready, and if not, he leaves. The issue is that the waiter now has to continuously walk the floor to determine which tables are ready, checking the whole room each lap, regardless of anyone actually is ready (this corresponds to [`select`](https://en.wikipedia.org/wiki/Select_\(Unix\))/[`poll`](https://en.wikipedia.org/wiki/Poll_\(Unix\))). Meanwhile, `epoll` is akin to giving each table a call button that allows them to notify the waiter when they are ready. This allows the waiter to attend tables that are guaranteed to be ready without constantly checking all the tables.

![](https://cdn.hashnode.com/uploads/covers/68c19865376514b2a5927314/31e6dd92-60c8-4b35-9e1a-3b59c0ea4791.png align="center")

It's worth noting an important drawback that this design introduces: a single long-running command will block all clients. The documentation highlights commands that involve many elements, such as `SORT`, `LREM`, `SUNION`. To troubleshoot this, the [Diagnosing latency issues](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/latency/) documentation provides a comprehensive guide.

Multiplexing and non-blocking I/O are part of the networking story, but two more factors facilitate Redis's network I/O efficiency.

## Optimizations for Network I/O

Redis is further optimized for handling network I/O by using a custom protocol called [Redis Serialization Protocol](https://redis.io/docs/latest/develop/reference/protocol-spec/) (RESP) and [pipelining](https://redis.io/docs/latest/operate/oss_and_stack/management/optimization/benchmarks/). RESP offers a simple implementation, fast parsing, and human readable commands, enabling Redis to quickly parse commands with minimal CPU cycles. This performance benefit is achieved by using prefixed lengths, thereby removing the need to scan for special characters or quote/escape the payload. Meanwhile, pipelining enables multiple commands to be sent with a single `write` operation by the client. The client can skip reading replies and continue to send commands to the query buffer. Redis then drains the query buffer and executes each command as it’s parsed, appending each reply to the client’s output buffer for the replies to go out together before Redis sleeps. This process reduces the latency by decreasing the total number of network round trips, and increases throughput by minimizing socket I/O, condensing multiple `read()` and `write()` syscalls into fewer (note that a read caps at 16KB, so a large pipeline may still require multiple syscalls).

![](https://cdn.hashnode.com/uploads/covers/68c19865376514b2a5927314/03a5e25e-a1bf-43c3-bd34-5cec1b1222b1.png align="center")

# Conclusion

There is no one thing that can take credit for Redis's performance. In reality it is the culmination of a deep understanding for how computers store and access memory combined with architectural decisions that maximize resource utilization based on Redis's use case. The fact that Redis operates entirely in-memory facilitates incredibly fast access and avoids costly operations involved in maintaining multiple data representations. Memory is further optimized by dynamically switching between compact memory formats and hash tables, which guarantees the best possible performance when working with smaller data structures. Moreover, as Redis is not bound by CPU and the operations are typically small, keeping all command execution on a single thread avoids unnecessary costs associated with shared-state, multi-threaded applications. Finally, networking optimizations such as a lightweight, custom communication protocol and pipelining minimize wasted CPU cycles without impacting the functionality.
