### What is RAID? RAID is a storage virtualization technology that combines multiple physical hard drives or SSDs into one logical storage unit. It was originally designed to improve: * **Performance** (faster data access) * **Reliability** (protecting against drive failures) * **Availability** (keeping systems running even when a drive fails) Originally, RAID stood for **Redundant Array of Inexpensive Disks**, but it is now commonly referred to as **Redundant Array of Independent Disks**. Data is distributed across drives using different techniques called **RAID levels**, each balancing speed, capacity, and fault tolerance differently. ([Harvey Mudd College][1]) --- ## History RAID was introduced in **1987** by researchers at the University of California, Berkeley. At the time, many small, inexpensive disks could be combined to match or outperform a single large, expensive disk (known as a **SLED – Single Large Expensive Disk**). Over time, RAID became a standard feature in servers, enterprise storage systems, and some desktop computers. ([Harvey Mudd College][1]) --- ## How RAID Works RAID organizes data using three main techniques: ### 1. Striping * Data is divided into blocks. * Blocks are spread across multiple drives. * Increases read and write speed because several drives work simultaneously. Example: | Disk 1 | Disk 2 | | ------- | ------- | | Block 1 | Block 2 | | Block 3 | Block 4 | | Block 5 | Block 6 | --- ### 2. Mirroring * Every piece of data is copied onto another drive. * If one drive fails, the copy remains available. Example: | Disk 1 | Disk 2 | | ------ | ------ | | File A | File A | | File B | File B | --- ### 3. Parity Parity stores extra information that allows lost data to be reconstructed if a drive fails. Instead of storing complete copies like mirroring, RAID calculates parity values from the data blocks. Example: | Disk 1 | Disk 2 | Disk 3 | | ------ | ------ | ------ | | Data A | Data B | Parity | If Disk 1 fails, Data A can be rebuilt using Data B and the parity information. Parity is much more storage-efficient than mirroring but requires additional calculations. --- # Standard RAID Levels ## RAID 0 – Striping **Minimum disks:** 2 Features: * Fastest RAID level * No redundancy * Uses 100% of storage capacity Advantages: * Excellent read/write performance * Full disk capacity available Disadvantages: * If one disk fails, all data is lost Typical use: * Gaming * Video editing * Temporary data --- ## RAID 1 – Mirroring **Minimum disks:** 2 Features: * Every disk has an identical copy. * Very reliable. Advantages: * Survives one disk failure. * Easy recovery. Disadvantages: * Only 50% of storage is usable. Example: Two 2 TB drives become **2 TB usable**, not 4 TB. --- ## RAID 2 Uses bit-level striping with error-correcting codes. It is rarely used today because modern drives already perform internal error correction. --- ## RAID 3 Uses: * Byte-level striping * One dedicated parity disk Advantages: * High sequential throughput Disadvantages: * Dedicated parity disk becomes a bottleneck. Rarely used today. --- ## RAID 4 Uses: * Block-level striping * One dedicated parity disk Improves random reads but still suffers from the parity-disk bottleneck. --- ## RAID 5 **Minimum disks:** 3 Uses: * Striping * Distributed parity Parity blocks are spread across all drives instead of using one dedicated parity disk. Advantages: * Good performance * Efficient storage * Can survive one drive failure Storage formula: **(Number of drives − 1) × Drive size** Example: Four 4 TB drives → **12 TB usable** Disadvantages: * Slow rebuild after a failure * A second drive failure during rebuild causes complete data loss --- ## RAID 6 Similar to RAID 5 but stores **two parity blocks**. Advantages: * Can survive two simultaneous disk failures. * Better suited for very large storage arrays. Disadvantages: * More storage overhead * Slightly slower writes due to additional parity calculations Storage formula: **(Number of drives − 2) × Drive size** --- ## RAID 10 (1+0) Combines: * RAID 1 (mirroring) * RAID 0 (striping) Requires at least four disks. Advantages: * Excellent performance * High fault tolerance * Fast rebuilds Disadvantages: * Only 50% of storage is usable Widely used for: * Databases * Virtual machines * Enterprise servers --- # Nested RAID Levels Some RAID systems combine multiple RAID levels, including: * RAID 01 * RAID 10 * RAID 50 * RAID 60 These combinations aim to balance speed, storage efficiency, and fault tolerance for specific workloads. ([Harvey Mudd College][1]) --- # RAID Implementations ### Hardware RAID A dedicated RAID controller manages the array. Advantages: * Better performance * Less CPU usage * Often includes battery-backed cache Disadvantages: * More expensive * Hardware failures may require a compatible controller --- ### Software RAID The operating system manages the RAID. Advantages: * No special hardware required * Lower cost * Flexible configuration Disadvantages: * Uses CPU resources * Performance depends on the operating system Examples include Linux `mdadm`, Windows Storage Spaces, and ZFS/Btrfs (which provide RAID-like capabilities). --- ### Firmware ("Fake RAID") Implemented in motherboard firmware with operating system drivers. Advantages: * Cheaper than dedicated hardware RAID Disadvantages: * Often offers little performance benefit over software RAID * May have compatibility issues --- # Reliability and Weaknesses Although RAID improves reliability, it has limitations: ### RAID is **not a backup** RAID protects against hardware failures but does **not** protect against: * Accidental deletion * Malware or ransomware * File corruption * Fire, flood, or theft * Human error Regular backups are still essential. --- ### Rebuild Time When a failed drive is replaced, RAID rebuilds the lost data. Modern large-capacity drives (e.g., 20 TB or more) can take many hours or even days to rebuild, during which the array is more vulnerable to additional failures. ([Harvey Mudd College][1]) --- ### Unrecoverable Read Errors (UREs) During a rebuild, if another disk has an unreadable sector, the rebuild may fail—especially in RAID 5. This is one reason RAID 6 is often preferred for large arrays. ([Harvey Mudd College][1]) --- ## Choosing the Right RAID Level | RAID Level | Speed | Fault Tolerance | Storage Efficiency | Typical Use | | ---------- | --------: | --------------: | -----------------: | ---------------------------------- | | RAID 0 | Excellent | None | 100% | High-performance workloads | | RAID 1 | Good | High | 50% | Critical personal or business data | | RAID 5 | Good | One drive | High | General-purpose servers | | RAID 6 | Good | Two drives | Moderate | Large storage systems | | RAID 10 | Excellent | High | 50% | Databases and virtualization | ### Key Takeaways * RAID combines multiple drives to improve **performance**, **availability**, and/or **fault tolerance**. * Different RAID levels use **striping**, **mirroring**, and **parity** in different combinations. * RAID 0 maximizes speed but provides no protection. * RAID 1 focuses on data redundancy through mirroring. * RAID 5 and RAID 6 use parity to balance capacity and reliability. * RAID 10 delivers both high performance and strong redundancy but requires more disks and sacrifices half the raw storage capacity. * RAID helps protect against disk failures, but it is **not a substitute for regular backups**.