Taking work now — the first look is freeServer and RAID drives posted in from anywhere in the UK, or handed in at ten drop-off pointsQuicker still, give us a ring:0800 6890668
RDDRAID Drive Data Recovery 0800 6890668 Price my job
RDD / What is inside a RAID member drive

SAS, SED, 520 bytes, TLER, helium

What is actually inside a RAID member drive. And why the controller gave up on it before you did.

A drive out of a server slot is a desktop drive's cousin with a different education. It speaks SAS or SCSI rather than SATA, or SATA with firmware that reports a bad sector in seconds so a RAID controller can act on it; it may carry 520 bytes in every sector, a key that only its controller holds, helium sealed inside it, or a vendor's firmware with a clock in it. This page says what each piece is, plainly, for anyone whose controller has just said Failed and who wants to understand what the bench is about to do to the drive, and why the rebuild button is not the answer tonight.

Rather talk it through? An engineer answers the bench line
0800 6890668

The interfaces, and why a desktop cannot see most of them.

SATA is the desktop interface: one port, the ATA command set, and a controller in every PC. SAS is the server interface: two ports for redundancy, the SCSI command set with its sense keys and log pages, and a connector that accepts a SATA drive but not the other way round. Parallel SCSI came before it, on 68-pin and 80-pin connectors with terminated buses, and Fibre Channel put the same command set on a loop for the SANs of the 2000s. NVMe is the newest, speaking directly to PCIe from U.2 and U.3 carriers, E1.S rulers and M.2 sticks. A SAS, SCSI, FC or NVMe drive needs a host adaptor of its own kind to be seen at all, and the bench has imagers for each, which is the first reason a drive from a server should not be tried in a PC.

Enterprise firmware, and time-limited error recovery.

A desktop drive that meets a bad sector retries for as long as it takes, sometimes for minutes, because there is nobody else to ask. A drive built for RAID has time-limited error recovery: it gives up after about seven seconds and reports the sector unreadable, because a controller that waited longer would assume the drive had died and drop it anyway. The result is that a RAID controller sees errors quickly and acts on them, marking the drive predictive when its counts cross a threshold and failed when a read fails outright. The drive the controller dropped is usually still readable on a bench that turns the time limit off and reads the sector the slow way, which is the second reason a dropped drive is not a dead one.

Sector formats: 512, 4Kn, 520 and 528.

Drives store data in sectors, and for decades every sector held 512 bytes. Advanced Format drives hold 4,096, either emulating 512-byte sectors to the host (512e) or presenting 4,096-byte ones natively (4Kn); a 4Kn drive in a host that expects 512 reports the wrong capacity or none. Storage arrays go further: NetApp, Sun and IBM format their drives with 520 bytes per sector and EMC with 528, the extra bytes holding a checksum or T10 Protection Information the array checks on every read. A standard host refuses such a drive as 0 GB or incompatible, and the forum's fix, reformatting it to 512 bytes, writes every sector and erases it. The bench images the drive at its native size and strips the extra bytes from the image.

Self-encrypting drives, and the PSID.

Most SAS drives and many others sold in the last decade are self-encrypting. The drive generates a media encryption key at manufacture and encrypts everything it stores with it, whether or not anyone turned encryption on; locking wraps that key with an authentication key the host supplies, a PERC passphrase, an HPE encryption key, a KMIP server's credential. Unlocked, the drive is transparent; locked and separated from its controller, it returns ciphertext. The 32-character PSID printed on the label resets a locked drive by regenerating the media key, which makes the old data noise forever. It is the right command for a drive being decommissioned and the wrong one for any drive whose data is wanted; the bench unlocks a SED on the imager with the key you supply and cannot help without it.

Helium, and shingled recording.

Helium drives, from HGST's He6 through every Exos X, Ultrastar HC5xx and Toshiba MG07 onwards, are laser-welded shut with helium inside so that more platters can spin closer together. Their heads fly on helium, and a head swap means opening the case once in clean air and imaging immediately, because those heads fly badly in air and not for long. Shingled drives, WD's Red EFAX and Ultrastar HC6xx among others, overlap their tracks like roof tiles and keep a map of where each block really is; they stall on the random writes of a rebuild, and when the map fails the drive reads as blank with the data still on the tracks. Both are recoverable, on different benches with different rules.

Why the controller drops a drive, and what predictive failure means.

A controller watches every drive's counters: SMART attributes on SATA, the grown defect list and error log pages on SAS. When one crosses the controller's threshold the drive is flagged predictive failure while it still reads; when a read fails outright the drive is dropped and the set is degraded. On a healthy set, predictive is a replacement under warranty and the rebuild reads clean survivors. On a degraded set, the same replacement asks the rebuild to read the predictive drive from end to end, including the sectors it has just said it cannot promise, and the first one it hits is where the rebuild stops. Pending sectors, SMART attribute 197, are the number that stops rebuilds; every one is a read the drive cannot answer and the rebuild will ask for.

The firmware clocks.

Some firmware has a counter in it that the code cannot handle past a certain value. HPE's SAS SSDs on firmware before HPD8 failed at 32,768 power-on hours, which is two to the power fifteen; another line before HPD7 at 40,000; Dell's LT-series at 40,000 before D417, reporting 0 GB afterwards; HPE's NVMe drives on one line at 4,700. Drives fitted together reach the hour together, and a RAID 5 of six can lose four before lunch. The fix prevents the failure and does not undo it; some failed drives answer at controller level on the bench and some never will, and the free look is honest about which.

Imaging on a SAS imager, and head swaps on 10K and 15K drives.

The bench reads a drive on equipment that controls how long a read may take and how many times it is retried, on the interface the drive speaks, at the sector size it was formatted with. A drive that clicks goes to the clean bench first, where its head stack is replaced with a matched donor and its service area, the drive's own firmware data on the platters, is checked and repaired; a 15K 2.5-inch drive has more heads flying closer than any desktop drive's, and its donors are scarcer every year. The drive is then imaged head by head with the weakest last, and from then on the original is not touched.

How the image feeds the rebuild.

Where one member failed and the set is otherwise healthy, the image goes back on a fresh drive of the same interface and format, sector-identical, and the controller rebuilds from a drive that answers every read. Where the set itself has failed, every member is imaged and the array reassembled in software from the images: the order, stripe size and parity rotation read from the controller's metadata on the drives, or recovered from the parity arithmetic where the metadata was cleared, with the first-dropped drive's stripes filling any survivor's gaps. The file system is repaired on the virtual volume and the files copied out. That second half is our RAID array site's subject, and it happens on the same bench.

Where that leaves your drive.

Nearly every member drive that reaches us was dropped by a controller that could not wait, or failed at the heads or in the firmware with its data intact behind the fault, and nearly every one is imaged and its image sent home. A smaller number arrive after a rebuild into a second failure, a cleared configuration, an initialise, a reformat or a PSID revert, and those recover less or nothing. The free look tells you which yours is, and what it will cost, before anything chargeable happens.

The questions that come up first.

Why does the controller say failed when the drive spins?

Because enterprise firmware reports a bad sector within seconds and the controller acts on the first serious one. On a bench that waits, the drive usually reads.

Why can't I read the drive in a PC?

A SAS, SCSI, FC or NVMe drive needs a host adaptor of its own kind; a 520-byte drive cannot be addressed by a standard host at all; a locked SED returns noise. And a PC that can see a SATA member will offer to initialise it.

Is a helium drive recoverable if the heads have failed?

Yes, once: opened in clean air, a matched head stack fitted, and imaged immediately. It is not reassembled for use.

What is a PSID revert?

The reset code on a SED's label. It erases the drive by destroying its key. It is not a recovery step.

Does any of this cost me anything to find out?

No. The free look identifies the drive, the format and the fault, and one figure follows in writing. £300 + VAT for one drive; £500 + VAT upwards for a set.

Now you know what the bench is about to do.

Send the form with the model strings, the controller and what it says, and the first look tells you which of these your drive needs, and what it would cost.

0800 6890668