I was under the impression that gitlab use gitlab.com for their work. Surely someone would have noticed within seconds that it was down?
Why have the misleading "updated a few seconds" ago text if it doesn't update on complete failure? :)
I was under the impression that gitlab use gitlab.com for their work. Surely someone would have noticed within seconds that it was down?
Why have the misleading "updated a few seconds" ago text if it doesn't update on complete failure? :)
The delay in updating status is a result of our Incident Management process [0]. We have a Communications Manager on Call (CMOC) who leads communication throughout an incident. One of their responsibilities includes updating the status page. The slight delay between noticing the issue and updating the status page is a result of the time it takes for the CMOC to get alerted, assess the situation, and write the communication that is shared on the status page.
I'm not sure how the "updated a few seconds ago" messages are generated but I'll try to find out once the incident has been resolved.
0 - https://about.gitlab.com/handbook/engineering/infrastructure...
At first glance it looks like everything is operational with no issues.
"Active Incident" remains because our team is still working towards full recovery.
"System Wide Outage" is the description of the incident at its onset.
Also, most alerting systems like check multiple times before declaring a public outage, many times 2 to 3 failures some seconds apart are needed.
1. External engineers will start to automate recovery/mitigation processes around your status page if it has real time status.
2. You now need to bug test your status page thoroughly because of #1. It basically becomes an actual API.