Bench record · NAS & RAID · PLY-2025-0642
The Power Cut Was Free. The Rebuild Was Not.
An engineering subcontractor working into the docks at Falmouth found a six-bay shelf two members short when the office opened after the bank holiday; the mains had been off since the Saturday. One of them had been showing amber since the Friday
, and nobody had followed it up. By the Tuesday a second disk had gone. Their IT firm dropped a spare into the empty bay and set a rebuild running
, and by the Wednesday it had stalled somewhere near sixty per cent
. The power cut itself took nothing from them. The rebuild took a fortnight.
Sounds like yours? Talk it through.
0800 6890668
Why it happens.
Two failed disks are workable. A rebuild laid on top of them is not. RAID 5 carries one member in hand, and that margin went the morning the first bay turned amber; with the second disk gone the controller found itself two values short in every stripe, and no route to either. A rebuild begun in that condition returns nothing and does real damage, laying fresh parity through the very stripes a recovery must find exactly as they were left. Server shelves add their own complications. SAS disks are commonly formatted with sector sizes an ordinary desktop cannot reach, and the vendor's own structures live in places consumer software will either overwrite or step straight past. Beneath the lot sat an amber warning nobody answered, and parity carrying the whole set by itself.
The kit this job called for.
How the job runs →| The kit | What it did here | Why we keep it |
|---|---|---|
| PC-3000 SAS/SCSI | Put all six server disks on a native SAS interface to read them | An ordinary desktop cannot address a SAS or SCSI disk taken out of a server |
| Atola TaskForce 2 | Imaged the six together instead of in sequence, which saved several days | Copies the whole array in one run, so no member waits its turn on the bench |
| UFS Explorer RAID Recovery | Read the geometry off the metadata and mounted a virtual volume over the images | Rebuilds the set, then works back through whatever volume layers the NAS added on top |
How it ran.
Image all six, the working ones too
The shelf stayed switched off. Each of the six disks was taken out and labelled with the bay it had sat in, then put on an imager — the four sound ones to begin with, so that neither failure was asked for anything before it had been properly sized up. Get the interface or the sector size wrong at this stage and every image made afterwards is quietly worthless.
Reading what the two failures still held
One was losing a head; the other grew fresh bad sectors as it was being read. Neither was anywhere close to unreadable. Both came off in short passes with long rests between them, and most of both surfaces went into images. Taking each of them twice earned its keep: where a sector read differently on the two passes, the cleaner version was kept, parity being in no position to settle it.
Let the metadata set the geometry
None of it was assumed from how that make of controller usually behaves. Member order, stripe width, the rotation of parity and the delay on it were all lifted out of the array's own structures. With those four settled, a virtual volume went together over all six images, the filesystem was read out of it, and the physical disks stayed switched off from start to finish.
How it ended up.
The rebuilt volume mounted, and its contents were checked against the directory tree it held before anything went back on new media, return postage paid here. One qualification is on the file: inside the band Wednesday's rebuild had already written across, a live project folder returned incomplete. A copy of it from a fortnight before was still on the engineer's own laptop.
Pages people open after this.
Other jobs on RAID & NAS.
Seeing the same thing?
Leave it switched off and post it in. Nothing happens until the diagnosis, which lists what can still be read and what cannot.