Thursday, August 10, 2006
Power outage
Apparently there was a 30-minute power outage at the co-lo this morning which was enough to kill the old dying battery in the UPS. Everything shut down nicely, though, and now the power has been restored and we're back up.
Thursday, June 08, 2006
Brief outage on Sunday, July 11 at noon
I hate to do this so soon after the last major outage, but we will be taking puck down briefly on Sunday to add another 512MB of memory and upgrade the kernel. This should only take 10-20 minutes, barring any unexpected hardware problems.
Update 6/13: Of course, I meant June 11, not July 11 as the title says. Hope nobody was confused. The upgrade was completed without trouble.
Update 6/13: Of course, I meant June 11, not July 11 as the title says. Hope nobody was confused. The upgrade was completed without trouble.
Monday, May 29, 2006
Memorial Day Weekend outages
On a trip through Syracuse, we decided to stop by the switch room and do a little maintenance on puck. One of the new 400GB drives was reporting read errors during SMART tests, so it needed to be replaced before it started failing for real. This took 90 minutes but seemed to go OK and puck was running when we left. By the time we got back to Pittsburgh, though, puck was nonresponsive on the network and on the console port. Cycling the power didn't increase the load on the PDU, which suggested a serious hardware failure. So, yesterday we drove back to Syracuse. It turned out to be that a HDD power cable was smashed under the case lid when it was reassembled yesterday, and it took a little while for the rack vibration to wear through the insulation and finally short the +5VDC line to the chassis. Some creative use of cable ties fixed the problem and puck is now back online.
To cap things off, today when I was upgrading atd, I accidentally triggered the checkfs script which failed and automatically rebooted the system.
So, sorry about all the unexpected downtime this weekend. Soon everyone will have their own huge hunk of space on /data to make up for it.
In the near future, I will be making an email announce list to notify everyone of outages in advance, and hopefully that will bring more coordination to maintenance scheduling.
To cap things off, today when I was upgrading atd, I accidentally triggered the checkfs script which failed and automatically rebooted the system.
So, sorry about all the unexpected downtime this weekend. Soon everyone will have their own huge hunk of space on /data to make up for it.
In the near future, I will be making an email announce list to notify everyone of outages in advance, and hopefully that will bring more coordination to maintenance scheduling.
Bookmark this site so when any Litech services are down, you can find out what's going on.
Subscribe to:
Posts (Atom)