Open, and work is coming in · weekdays, 9am–5:30pm Faster by phone: 0800 6890668
PDR Plymouth Data Recovery 0800 6890668 Get a price
PDR / Casebook / Two Failed Bays of Four

NAS & RAID · job record · PLY-2025-0341

One Bay Failed, Then Another Nine Months On.

On a shelf above the grinding bench in a Saltash boatyard workshop sat a four-bay Synology, with the fan pulling dust off the bench into it year after year. Bay 2 had reported bad sectors from before Christmas onward, and somebody had turned the warning emails off instead of ringing anyone. The day an order file would not open, the unit was powered down and started again; it came up reading volume crashed. On the lid, in marker pen: stick a new drive in and let it sort itself out. That instruction was the one thing we did not do.

The customer checked and agreed No names published

Sounds like yours? Talk it through.
0800 6890668

Why it happens.

A RAID 5 set carries one spare disk and not a fraction more. Parity is thin — a little in every stripe — enough to reconstruct one missing member, no help at all once two have gone. Two absences leave the arithmetic nothing to solve for. Powering the unit up for a look would have cost twice over: two tired disks pressed back into service, and a crashed volume able to write fresh metadata across the very structures a recovery depends on. So it stayed unplugged. Our refusal of the marker-pen instruction went in writing, with the reason: pushing a failing member into a live set only loads more work onto the disks still doing theirs. Nothing was switched on. Four disks were copied, and everything after that ran on the copies.

The kit this job called for.

How the job runs →
The kitWhat it did hereWhy we keep it
DeepSpar Disk Imager 4Both failed members were given new head stacks and read, the newer failure firstScores every head before it starts, works one surface at a time, and holds resets, timeouts and power in check
Atola TaskForce 2Copied the two working disks in parallel while the failed pair waited their turnCopies the whole array in one run, so no member waits its turn on the bench
UFS Explorer RAID RecoveryRead the mdadm superblocks off, then rebuilt the volume across the four imagesRebuilds the set, then works back through whatever volume layers the NAS added on top

How it ran.

01

The two working disks were tired as well

Months of carrying a degraded array leaves a mark, and by this point both of the working members had pending sectors of their own. The TaskForce took the pair together while the two failed disks waited, which saved several days and, more to the point, meant no disk in the set was declared healthy without being read end to end.

02

What the event log actually said

Both failed disks accepted a donor head stack and began to read. The event log then reshaped the job. That noisy bay had actually dropped out of the array the previous January, leaving its contents nine months behind the rest — a view of the volume from long before the crash. Everything current sat on the disk that had gone most recently, where the unreadable areas formed bands it was possible to work around.

03

Build the volume from three images and set January aside

None of the geometry was guessed. The array's own metadata gave up the member order, the width of the stripes, and which way parity rotates. Assembly then used the three current images alone; the January disk played no part, because blocks written months apart, blended together, yield corruption neat enough to survive a consistency check and still be wrong. Whatever the most recent disk declined to return lay in unallocated space.

How it ended up.

It mounted first time. Drawings, job cards and the yard's own order books were checked against its numbering before replacement disks went into the post. The Synology has moved into the office now, away from the bench, and the alert emails are back on.

In short: Parity covers a single failed disk, and not one beyond that. Ignore an amber warning for a whole winter and the next failure has already been paid for.

Seeing the same thing?

Leave it switched off and post it in. Nothing happens until the diagnosis, which lists what can still be read and what cannot.

0800 6890668