RAID emergencies · triage

RAID won’t rebuild? Stop before it costs the array.

A rebuild that fails, hangs at a percentage, or drops a second drive mid-run is the most dangerous moment in an array’s life — because every instinct says retry, and retrying is how one-drive problems become total losses. Here’s why rebuilds fail, the three moves that make things worse, and the discipline that keeps a wounded array recoverable.

Members imaged first
Free 48-hour diagnostic
No fix, no fee · most jobs
// the paradox

The rebuild is the most dangerous thing an array ever does.

Rebuilding reads every sector of every surviving drive at sustained maximum intensity — applied to disks the same age and wear as the one that just died. It’s a stress test scheduled at the worst possible time.

Never
Force a member online
Never
Re-initialise the array
Never
Retry, retry, retry
Always
Label · image · copies
// why rebuilds fail

Three reasons a rebuild dies mid-run.

A read error on a survivor. Rebuilding RAID 5 means reading everything from every remaining drive; one unreadable sector on one survivor and many controllers abort — or worse, drop that drive too. On today’s multi-terabyte disks, the odds of hitting at least one such sector during a full-array read are uncomfortably real, which is why big rebuilds fail so much more often than small ones ever did. The batch-twin problem. Arrays get built from identical drives bought the same day, aged by identical workloads — when one dies of wear, its siblings are the next most likely drives in the building to die, and the rebuild hands them the heaviest week of their lives. The wrong-drive accident. Under pressure, a healthy member gets pulled instead of the failed one, or a hot spare kicks off an automatic rebuild against a faltering survivor — human-factor failures that account for a startling share of the arrays on our bench. Different causes, one shared property: after the failure, the data still exists on the members — unless the next move destroys it.

// the three killers

The buttons that turn degraded into destroyed.

Force online. The controller dropped a member for a reason and at a moment; forcing it back merges stale data into a live array, and the ‘repair’ that follows calculates garbage into every stripe. Re-initialise / recreate the array. Offered by every controller as the fresh-start option — it writes new metadata (and sometimes new parity) across all members, over the top of your volume. The array becomes healthy; its contents become the recovery job. Serial retries. Rebuild fails, reboot, rebuild again — each pass re-runs the full-intensity read over drives that just demonstrated they can’t survive it. If one retry didn’t work, the second one is spending your data’s remaining odds. The moment any of these tempts you is the moment to power down instead.

// the discipline

What keeps a wounded array recoverable.

The protocol fits on a label-maker. Power down the box — a degraded array at rest is stable; the same array under load is deteriorating. Photograph and label every drive with its bay position before anything moves; member order is reconstruction gold. Change nothing in the controller — no config saves, no ‘repairs’, no firmware updates mid-crisis. Then the professional sequence: every member is imaged read-only — including the ‘failed’ one, which is often 95% readable on gentler hardware and holds the stripes the survivors lack — and the array is reassembled virtually from those images, where parameters can be tested and mistakes cost nothing. Your data comes off the virtual volume; the original drives are never gambled with. That’s the substance of our RAID recovery work — rack servers and desktop NAS boxes alike — and the deeper why-arrays-fail arithmetic lives in the RAID 10 guide.

// questions

Asked before you ask, answered.

If the array holds data you can’t lose and you have no verified backup: no — pause it if the controller allows, and stop adding load either way. A crawling rebuild usually means a surviving drive is struggling through read errors, and hours more of maximum-intensity reading is the exact stress that kills second drives. A paused, degraded array is recoverable; a rebuild that killed drive two often isn’t — not fully.

Forcing a member online is the single most destructive button in RAID. The controller dropped that drive at a moment in time; the array has changed since. Force it back and the controller may treat stale data as current — corrupting the very parity it then uses to ‘repair’ everything else. If the array matters, the honest sequence is power down and image every member; reassembly happens virtually, where mistakes cost nothing.

Same bench, same method. A four-bay Synology or QNAP under a desk is architecturally a small RAID server — mdadm or a vendor variant on standard disks. Send the labelled drives only (no enclosure needed): every member is imaged, the array logic is reconstructed from the images, and the volume comes back file by file. Details on the NAS recovery page.

// rebuild failing right now?

Power down. Then we take it from copies.

Free 48-hour diagnostic on multi-disk arrays at the Bristol lab — every member imaged before any array logic runs, written quote before any work.

Call us — 0117 332 1137
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →