On a trip through Syracuse, we decided to stop by the switch room and do a little maintenance on puck. One of the new 400GB drives was reporting read errors during SMART tests, so it needed to be replaced before it started failing for real. This took 90 minutes but seemed to go OK and puck was running when we left. By the time we got back to Pittsburgh, though, puck was nonresponsive on the network and on the console port. Cycling the power didn't increase the load on the PDU, which suggested a serious hardware failure. So, yesterday we drove back to Syracuse. It turned out to be that a HDD power cable was smashed under the case lid when it was reassembled yesterday, and it took a little while for the rack vibration to wear through the insulation and finally short the +5VDC line to the chassis. Some creative use of cable ties fixed the problem and puck is now back online.
To cap things off, today when I was upgrading atd, I accidentally triggered the checkfs script which failed and automatically rebooted the system.
So, sorry about all the unexpected downtime this weekend. Soon everyone will have their own huge hunk of space on /data to make up for it.
In the near future, I will be making an email announce list to notify everyone of outages in advance, and hopefully that will bring more coordination to maintenance scheduling.