Turns out running ZFS on a USB-attached JBOD is a bad idea
How bad could it be?
In the beginning, there was a second-hand OptiPlex with a consumer-grade 8TB drive: my home server. This setup was good enough for a media server, and maybe one or two cron jobs, but the lack of redundancy made me queasy when thinking of relying on it for any of my personal data. This server was my daily driver and performed admirably through my university years, but once I got a full-time job, and thus a bit more disposable income, I decided to replace it with something more robust and maybe learn a thing or two along the way.
At the time, I was living in a charming but compact 1-bedroom apartment with no space for a server rack, and the noise of a standard server would likely not get me many brownie points with my partner. Thus, I went with a mini PC, a Minisforum MS-01, and decided to connect HDDs to it through a USB JBOD (Just A Bunch Of Disks) enclosure. I made sure to check and double-check my choices, paying special attention to get an enclosure that would pass the drives' SMART data through to the PC, and that the connection was fast enough not to bottleneck four drives. For the drives, I went for three recertified 12TB Seagate Exos server-grade drives. I was aware that the USB dock wasn't the ideal deployment setup but had seen others online do similar things with moderate levels of success. I thought to myself, "How bad could it be?"
I installed Proxmox on the host and created the zpool. Everything went well! I didn't run into any problems creating or managing the pool, and a quick reboot ensured that the drive IDs were stable. Success! I kept going and slowly started to move my entire life into the server. Nearly a year passed, and one fateful Sunday morning, I started getting strange errors in my Nextcloud instance. After some troubleshooting, I decided to check the zpool and was greeted like this:
pool: bulk
state: FAULTED
status: The pool metadata is corrupted and the pool cannot be opened.
action: Destroy and re-create the pool from a backup source.
see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-72
config:
NAME STATE READ WRITE CKSUM
bulk FAULTED 0 0 1 corrupted data
raidz1-0 UNAVAIL 0 0 0 insufficient replicas
usb-ST12000N_M000J-2TY103-0:0 ONLINE 0 0 0
usb-ST12000N_M000J-2TY103-0:1 UNAVAIL 0 0 0 corrupted data
usb-ST12000N_M000J-2TY103-0:2 UNAVAIL 0 0 0 cannot open
Oh no. Something had gone very wrong overnight. Alright then, just restore from a backup right? Well, backups are expensive, and while I had a bit more money to invest into my own infrastructure I had been prioritizing it for other, more important things in my life, and had inadvertently let myself rely on my server. I didn't have a backup. RAID-Z1 is fault tolerant, right? I'd assumed that if a drive failed, I could just shut down my pool and buy another drive to replace it. Well, two drives had failed at once.
I started thinking about what it could be: how rare was a correlated failure like this? I'd bought all three at once. Had they come from the same batch, with a shared defect that made them fail together? I just couldn't believe it. I had all of my documents and photos in that pool. So I began trying to see if I could recover anything.
ZFS is copy-on-write, so it never overwrites data in place. Every time it commits changes, it writes a new version of the pool's state, and each disk keeps pointers to the most recent ones (called uberblocks). My thinking was that if the current state was broken, maybe I could import an older one from before whatever had happened. So I went hunting:
$ zdb -ul /dev/sdb1 | grep -E 'txg|timestamp'
txg = 4817263
timestamp = 1776549312 UTC = Sat Apr 18 21:55:12 2026
txg = 4817262
timestamp = 1776549307 UTC = Sat Apr 18 21:55:07 2026
...
$ zpool import -o readonly=on -T 4817240 bulk
cannot import 'bulk': one or more devices is currently unavailable
I kept desperately working my way back through older and older transaction groups, and got the same answer every time. It took me a long evening to realize why this was never going to work: every one of those older versions still lived on the same three disks, and ZFS still couldn't see two of them. Going back in time doesn't help much when two thirds of your pool is missing.
I decided to look elsewhere. Maybe it wasn't a ZFS problem? I started looking at what the kernel could see of the disks:
$ lsblk -o NAME,SIZE,TRAN,MODEL
NAME SIZE TRAN MODEL
sdb 10.9T usb ST12000NM000J-2TY103
sdc 2T usb ST12000NM000J-2TY103
One drive was missing entirely, and the other was reporting as 2T, which was definitely not what I had paid for. After some digging, I found out that 2T isn't a random number: it's exactly 2³² sectors × 512 bytes, the most a disk can report through the old 10-byte SCSI capacity command. Bigger disks need the 16-byte version, and if the USB bridge messes that up, the kernel falls back to the old one and gets a capped answer. This also explained ZFS's "corrupted data". ZFS keeps copies of its labels at both the start and the end of each disk, so a disk that suddenly ends at 2TB has lost half of them.
This opened up a new possibility. Maybe it wasn't the drives? I decided to test this theory by swapping the two failing drives' bays. The 2T drive was sitting in bay 2, and the failed drive in bay 3. After the swap, the problem followed the bay: the drive now in bay 3 vanished, and the one that had been missing showed up in bay 2, truncated to 2TB. The fault was in the bay, not the drive!
Later, when I went through the kernel logs, I found the exact moment it happened:
Apr 18 21:55:14 pve kernel: scsi host6: uas_eh_device_reset_handler start
Apr 18 21:55:14 pve kernel: usb 4-1: reset SuperSpeedPlus Gen 2x1 USB device number 2 using xhci_hcd
Apr 18 21:55:15 pve kernel: scsi host6: uas_eh_device_reset_handler success
Apr 18 21:55:15 pve kernel: sd 6:0:0:1: [sdc] Very big device. Trying to use READ CAPACITY(16).
Apr 18 21:55:15 pve kernel: sd 6:0:0:1: [sdc] Read Capacity(16) failed: Result: hostbyte=DID_ERROR driverbyte=DRIVER_OK
Apr 18 21:55:15 pve kernel: sd 6:0:0:1: [sdc] Using 0xffffffff as device size
Apr 18 21:55:15 pve kernel: sd 6:0:0:1: [sdc] 4294967296 512-byte logical blocks: (2.20 TB/2.00 TiB)
Apr 18 21:55:15 pve kernel: sdc: detected capacity change from 23437770752 to 4294967296
Apr 18 21:55:16 pve kernel: WARNING: Pool 'bulk' has encountered an uncorrectable I/O failure and has been suspended.
I scrambled to order an HBA card, and moved the drives over to proper SATA connections as soon as it arrived. I opened a terminal and checked the drives:
$ for d in sdb sdc sdd; do smartctl -H -A /dev/$d | grep -E 'result|Reallocated_Sector_Ct|Power_On_Hours'; done
SMART overall-health self-assessment test result: PASSED
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 0
9 Power_On_Hours 0x0032 069 069 000 Old_age Always - 27514
SMART overall-health self-assessment test result: PASSED
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 0
9 Power_On_Hours 0x0032 065 065 000 Old_age Always - 31208
SMART overall-health self-assessment test result: PASSED
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 0
9 Power_On_Hours 0x0032 070 070 000 Old_age Always - 26877
All three passed, with zero reallocated sectors and around three years of previous service each, which is about what you'd expect from recertified drives. So the drives really were fine. Then came the real test:
$ zpool import bulk
It imported! No rollbacks, no recovery flags, nothing. The pool I had spent a whole evening trying to bring back from the dead had been alive all along. Well, mostly. The first zpool status showed 247 data errors, and ZFS immediately started a resilver on its own:
scan: resilvered 214G in 01:10:42 with 18 errors on Fri Apr 24 23:41:07 2026
The same logs had another surprise: the bay 3 drive had actually dropped off the bus on March 14th, five weeks before the pool died, and bulk had been running degraded on the other two drives ever since without me noticing.
$ journalctl -k --since 2026-03-14 --until 2026-03-15 | grep -E 'sdd|host6|usb 4-1|zio'
Mar 14 03:12:41 pve kernel: sd 6:0:0:2: [sdd] tag#18 uas_eh_abort_handler 0 uas-tag 4 inflight: CMD OUT
Mar 14 03:12:41 pve kernel: sd 6:0:0:2: [sdd] tag#18 CDB: Write(16) 8a 00 00 00 00 02 4a 5c 81 80 00 00 01 00 00 00
Mar 14 03:12:41 pve kernel: scsi host6: uas_eh_device_reset_handler start
Mar 14 03:12:41 pve kernel: usb 4-1: reset SuperSpeedPlus Gen 2x1 USB device number 2 using xhci_hcd
Mar 14 03:12:42 pve kernel: scsi host6: uas_eh_device_reset_handler success
...
Mar 14 03:13:43 pve kernel: sd 6:0:0:2: Device offlined - not ready after error recovery
Mar 14 03:13:43 pve kernel: I/O error, dev sdd, sector 9837511040 op 0x1:(WRITE) flags 0x700 phys_seg 32 prio class 0
Mar 14 03:13:43 pve kernel: zio pool=bulk vdev=/dev/disk/by-id/usb-ST12000N_M000J-2TY103-0:2-part1 error=5 type=2 offset=5036808798208 size=131072 flags=180880
Mar 14 03:13:44 pve zed[2231]: eid=86 class=statechange pool='bulk' vdev=usb-ST12000N_M000J-2TY103-0:2-part1 vdev_state=FAULTED
Once the bay 3 drive was back, the resilver caught it up on five weeks of writes, and with a third drive to check against, ZFS could rebuild most of the blocks the JBOD had damaged on its way down.
But not everything. Eighteen errors were still left after the resilver, so I ran a scrub. Then I ran another one, just in case, and got the same result:
scan: scrub repaired 0B in 03:47:14 with 18 errors on Sun May 10 04:11:15 2026
config:
NAME STATE READ WRITE CKSUM
bulk ONLINE 0 0 0
raidz1-0 ONLINE 0 0 0
ata-ST12000NM000J-2TY103_WV70DEB1 ONLINE 0 0 3.61K
ata-ST12000NM000J-2TY103_ZRT0CQHB ONLINE 0 0 3.61K
ata-ST12000NM000J-2TY103_ZRT0CKV1 ONLINE 0 0 3.61K
errors: Permanent errors have been detected in the following files:
/bulk/vm-drives/images/101/vm-101-disk-0.raw
/bulk/vm-drives/images/102/vm-102-disk-0.qcow2
"Permanent errors" is not a phrase you want to see next to the disk image with your only copy of years of photos. READ and WRITE were zero on every drive, so none of them was having trouble talking to the HBA anymore. The only non-zero column was CKSUM, and it was the same on all three disks. I later learned this is what RAID-Z does when it can't rebuild a block: it can't tell which disk had the bad piece, so it blames all of them equally.
RAID-Z1 can rebuild a block from parity if one piece of the stripe is bad, which is how the resilver fixed most of those 247 errors. For these 18 blocks, though, the JBOD had mangled more than one piece of the same stripe, and there wasn't enough left to rebuild from. The scrub had found these blocks and given up on fixing them. While this seemed dire, the VMs were booting fine. No system-critical files seemed to have been corrupted, and as far as I could tell, I was able to access all of my photos and documents again. The errors wouldn't go away, but they were a drop in the bucket compared to what I had started with. I could live with 18 errors.
The worst part was that ZFS did notice. ZED, the little daemon that watches for ZFS events, saw the drive fall off and dutifully sent an email about it, like Proxmox sets it up to do by default. The problem was where it sent it. When I installed Proxmox, I clicked right past the email field and left it on the default, mail@example.invalid. I never set up proper alerting, so for five weeks my server was trying to tell me something was wrong and sending every warning into the void. Now ZED sends alerts to ntfy on my phone, and I tested it by pulling a drive. Lesson learned. Now, about those backups...