edin Our cable op had some kind of week long fiber outage in the entire system early april, even different hubs and customers hundred+ miles away. It might have had something to do with provisioning. ONT’s were not getting IP addresses, and at first it was random. Initially the common denominator was subscribers with the Plume router, so at first they thought a bad firmware update went out. A message on the website said a bad firmware outage causes a massive spike in traffic causing a systemwide outage. But then there were reports of people going out of service even with their own routers.
So then the thoughts were something to do with provisioning the ONT’s. I’m not sure if its a Nokia based server since the chassis and ONT’s are all Nokia based, or if theres some other third party software that runs on it. I know at least the Voice/Video/DOCSIS system is on the CSG billing platform. Its almost like the fiber system didn’t recognize paying customers so it wasn’t allowing DHCP to complete and dish out IP addresses. Something about they were doing an upgrade on a server in one site at our headend and it wasn’t correctly power protected. There was a power outage and in mid-upgrade the server rebooted and corrupted the update. Then when it came online it replicated its corruption to the backup server up at their backup headend 100 miles away and so now they had two faulty servers.
Engineers were rotating sleeping on site, didn’t even get to leave to go home for days. Its almost if they had to rebuild the entire system and manually enter every customer or something. It was a huge deal and when done a releif for those who spent 80+ hour work week working with vendor and getting customers online.
Meanwhile anyone on DOCSIS was completely unaffected. So whatever happened did not cross over. Therefore I do not think SECV uses DPoE on their Nokia platform. I believe its a completely separated system. Both DOCSIS and PON can rent out Plumes, so the initial thought of some kind of bad Plume “broadcast storm” is false. I’m not sure what the IP schemes are on the PON system until I get it myself. I do think DOCSIS dishes out a bunch of /22 networks in the 10.x.x.x range thats then routed to upstream interfaces.
They never did put out an official RCA / Press release with a lessons learned statement. So while I’m excited for fiber and I’m definitely going to jump on it when I can… I’m rightfully concerned that they had an entire 5 day long outage ONLY on the fiber system, and it didn’t matter if you lived in Hazelton, Bloomsburg, Birdsboro, Earl township, Amity, etc… Sites 100 miles apart, served from different hubs all experienced the same issue, while the back end internet was fine on DOCSIS. BIZARRE. I’d love to be a fly on the wall to know exactly all the details and what they are doing to ensure it doesn’t happen again. I’m almost nervous that “what if” it does happen again when I move to Fiber? I think the extremely bad social media publicity (that they since scrubbed on their Facebook posts) and a week long outage with multiple engineers working 80+ hours would bring some very valuable lessons learned and mitigations put in place to not let an event like that happen again.