Introduction
Memory-mapped I/O (mmap) allows a process to map a file directly into its virtual address space, making file contents accessible through ordinary pointer dereferences.
The operating system's virtual memory subsystem handles page faults, eviction, and writeback transparently.
On the surface, this seems ideal for database storage engines: the kernel manages a buffer pool for free, the code is simpler, and the system can leverage decades of OS virtual memory engineering.
In practice, the tradeoffs are more complex, and most high-performance database systems either avoid mmap entirely or use it in carefully constrained ways.
How mmap Works
When a process calls mmap() on a file, the kernel creates a mapping between a range of virtual addresses and the file's contents on disk.
No data is loaded immediately.
Instead, the first access to a mapped page triggers a page fault.
The kernel's page fault handler reads the corresponding file block into a physical page frame, updates the page table, and resumes the faulting instruction.
Subsequent accesses to the same page hit physical memory directly, with no system call overhead.
The kernel manages mapped pages using the same LRU-approximate eviction policies it applies to the page cache.
Dirty pages (modified via stores to the mapped region) are written back to disk asynchronously by the kernel's writeback threads (historically pdflush, replaced in Linux 2.6.32 by per-backing-device writeback threads, and later managed via kworker threads) or by kswapd under memory pressure, or synchronously via msync().
Key System Calls
mmap(addr, length, prot, flags, fd, offset): Establishes the mapping.munmap(addr, length): Removes the mapping.msync(addr, length, flags): Forces dirty pages to disk.madvise(addr, length, advice): Provides hints to the kernel (e.g.,MADV_SEQUENTIAL,MADV_RANDOM,MADV_DONTNEED).
Why Database Systems Consider mmap
The appeal of mmap for database engines rests on several practical advantages:
Reduced code complexity. A buffer pool manager is one of the most intricate components of a database system.
It must handle page pinning, reference counting, eviction policies, dirty page tracking, and concurrency control on its own metadata structures.
Using mmap delegates much of this work to the OS.
Zero-copy access. Data is accessed in place.
There is no need to copy pages from a kernel buffer into a user-space buffer pool.
This eliminates one copy per page access and can reduce memory pressure.
Large virtual address spaces. On 64-bit systems, the virtual address space is large enough to map entire databases.
On conventional x86-64 hardware with 4-level paging, user space typically has 47 usable address bits (128 TiB).
Newer hardware and kernels supporting 5-level paging (e.g., Linux with CONFIG_X86_5LEVEL) extend this to 56 usable bits (64 PiB).
ARM64 configurations vary similarly.
In all cases, the virtual address space is sufficient to map very large database files, allowing the engine to treat the whole file as an in-memory array and simplifying pointer arithmetic for page access.
Transparent prefetching. The kernel can detect sequential access patterns and issue readahead, reducing I/O latency for scan-heavy workloads without explicit prefetch logic.
The Problems with mmap for Databases
Despite its appeal, mmap introduces several serious problems that have led most production database systems to implement their own buffer managers.
1. Lack of Control over Eviction
The OS page cache uses a global eviction policy (typically a clock or LRU variant) that is unaware of database-level access patterns.
A database system knows which pages are hot (e.g., the root of a B-tree), which are about to be accessed (through query plan analysis), and which can be safely evicted (e.g., pages from a completed sequential scan).
The OS has none of this information.
Under memory pressure, the kernel may evict a critical index page to make room for a page that will never be accessed again.
2. Uncontrollable I/O Stalls
Page faults are synchronous from the perspective of the faulting thread.
When a thread dereferences a pointer to an unmapped page, it blocks until the kernel completes the disk read.
The database engine has no way to schedule this I/O, batch it with other reads, or handle it asynchronously.
This creates unpredictable latency spikes, which is particularly problematic for OLTP workloads with tight latency requirements.
3. Error Handling
When a buffer pool manager issues a pread(), it receives an error code on failure.
When a page fault fails (due to I/O error, for example), the kernel delivers a SIGBUS signal to the process.
Handling SIGBUS reliably is difficult, and most recovery strategies involve terminating the process or the faulting transaction.
This makes robust error handling significantly harder compared to explicit I/O.
4. Dirty Page Writeback
The kernel decides when to flush dirty pages to disk.
This interacts poorly with write-ahead logging (WAL) protocols.
For correctness, a database must ensure that the WAL record for a modification is durable before the modified data page reaches disk.
With mmap, the kernel can write back a dirty data page at any time, potentially violating the WAL protocol.
Preventing this requires either mprotect() tricks (which are expensive due to TLB shootdowns) or careful use of msync() to coordinate flushes.
5. TLB Shootdowns and Scalability
When a mapping is modified or removed (e.g., via munmap() or mprotect()), the kernel must invalidate the corresponding TLB entries on all cores that may have cached them.
This inter-processor interrupt (IPI) is expensive and does not scale well with core count.
Database systems that use mmap with frequent mapping changes can become bottlenecked on TLB shootdown overhead.
6. Non-Trivial Concurrency Issues
Multiple threads accessing the same mapped region can cause contention on the kernel's internal page table locks and the address space semaphore.
In older Linux kernels this was mmap_sem; since Linux 5.8, it has been renamed mmap_lock to better reflect its semantics, though its role is the same.
This read-write lock is acquired on every page fault and on every call to mmap, munmap, madvise, or mprotect.
Under high concurrency, this lock becomes a severe bottleneck.
Walkthrough
The following walkthrough illustrates the lifecycle of a page access in an mmap-based storage engine versus a traditional buffer pool manager.
mmap-Based Access
1. Engine opens database file and calls mmap() to map the entire file.
2. To read page N, engine computes: ptr = base_address + (N * PAGE_SIZE)
3. Engine dereferences ptr.
4. If page is resident in physical memory:
- Hardware resolves the virtual address via TLB/page table.
- Data is returned. No kernel involvement.
5. If page is NOT resident (page fault):
a. CPU traps into kernel.
b. Kernel identifies the fault address and the associated file mapping.
c. Kernel allocates a physical page frame.
d. Kernel issues I/O to read the page from disk (synchronous to faulting thread).
e. Kernel updates the page table entry.
f. Kernel returns to user space; the original instruction resumes.
6. To modify the page, engine writes through the pointer.
7. Kernel marks the page dirty in the page cache.
8. At some future time (kernel-determined), the dirty page is written back to disk.
9. Engine has NO direct control over step 8 timing.
Buffer Pool Manager Access
1. Engine maintains a hash table mapping (file_id, page_id) -> buffer frame.
2. To read page N:
a. Engine looks up (file_id, N) in the hash table.
b. If found (cache hit):
- Pin the page (increment reference count).
- Return pointer to the buffer frame.
c. If not found (cache miss):
- Select a victim frame using engine's eviction policy (LRU, CLOCK, etc.).
- If victim is dirty, write it to disk (under WAL constraints) via pwrite().
- Issue pread(fd, frame_ptr, PAGE_SIZE, N * PAGE_SIZE).
- Check return value for errors.
- Insert new mapping into hash table.
- Pin the page and return pointer.
3. To modify the page:
a. Log the modification to the WAL first.
b. Modify the page in the buffer frame.
c. Mark the frame as dirty.
4. Engine flushes dirty pages to disk at controlled checkpoints,
always ensuring WAL records are durable first.
The buffer pool approach requires more code, but every I/O decision is explicit and controllable.
Systems That Use mmap
Despite the drawbacks, several notable systems use mmap:
- LMDB (Lightning Memory-Mapped Database) uses a single read-only mmap for its entire database file. Writes go through a separate path. The read-only mapping avoids the dirty page writeback problem.
- MongoDB (WiredTiger) historically used mmap for its original storage engine (MMAPv1), but replaced it with WiredTiger, which uses a custom buffer pool. The switch was motivated by multiple factors, including the limitations of mmap-based writeback control, as well as the need for document-level concurrency and compression support — capabilities that MMAPv1's architecture could not cleanly provide.
- SQLite can optionally use mmap for read access via its
PRAGMA mmap_sizesetting. - RavenDB uses mmap and has documented strategies for working around its limitations.
Crotty, Leis, and Pavlo published a thorough study at CIDR 2022 that benchmarked mmap-based approaches against buffer pool managers and found that mmap consistently underperformed for database workloads due to the problems outlined above.
When mmap Can Work
mmap is most suitable when the database workload has specific characteristics: the dataset fits in memory (eliminating most page faults), the workload is read-heavy (avoiding dirty page writeback issues), concurrency is moderate (avoiding mmap_lock contention), and latency predictability is not critical.
Embedded databases with these properties can benefit from the code simplicity that mmap provides.
Key Points
- mmap delegates buffer management to the OS kernel, reducing code complexity but sacrificing control over eviction, I/O scheduling, and dirty page writeback.
- Synchronous page faults create unpredictable latency stalls that are invisible to the database engine, making mmap unsuitable for latency-sensitive OLTP workloads.
- The kernel may flush dirty mapped pages to disk at any time, potentially violating write-ahead logging protocols required for crash recovery.
- I/O errors during mmap access surface as SIGBUS signals rather than return codes, making robust error handling substantially harder.
- Contention on the kernel's mmap_lock (formerly mmap_sem, renamed in Linux 5.8), and TLB shootdown costs degrade performance under high concurrency on multi-core hardware.
- Most production database systems (PostgreSQL, MySQL/InnoDB, Oracle) use custom buffer pool managers for the control they provide over page lifecycle.
- mmap remains viable for read-heavy, moderate-concurrency workloads where the dataset largely fits in memory and code simplicity is valued.
References
Crotty, A., Leis, V., and Pavlo, A. "Are You Sure You Want to Use MMAP in Your Database Management System?" Proceedings of CIDR 2022 (12th Annual Conference on Innovative Data Systems Research), 2022.
Graefe, G. "A Survey of B-Tree Locking Techniques." ACM Transactions on Database Systems, Vol. 35, No. 3, 2010.
Stonebraker, M. "Operating System Support for Database Management." Communications of the ACM, Vol. 24, No. 7, 1981.
Silberschatz, A., Korth, H. F., and Sudarshan, S. "Database System Concepts." 7th Edition, McGraw-Hill, 2019.