Chargetrip status page
  • Status
  • Events
  • Monitors
Get in touch
Chargetrip
Back
Feb 14, 2023
4 years ago
Service Outage
StationRoute (legacy)TileChargetrip GO
Resolved · February 17 at 11:26 AM (UTC) (in 3 days)

Service outage on Tuesday, February 14 2023

Summary:

In the early morning of February 14, Chargetrip received a high volume of requests that started to back up our exchange broker queue. This resulted in our exchange broker seizing, a service slowdown, and, ultimately, a service outage. All customers were affected either with slower calculation times or a complete service outage. Chargetrip recovered all systems within three hours.

Impact:

All customers were affected by this outage by slow calculation times or denial of service. Affected services were: the routing engine, station database, vehicle database, tile service and go.chargetrip.com.

Timeline:

  • Around 07:30 CET, we started receiving large volumes of messages that would ultimately seize our exchange broker.
  • Around 09:00 CET, our internal warning system began sending alerts to our DevOps team.
  • 09:28 We noticed our exchange broker was using an incredible amount of memory (1.4M messages in sync stations queue).
  • Around 09:30 CET, our resources were increased to unblock our queue.
  • 09:32 We scaled our sync daemon to 40 replicas.
  • 09:33 We updated our status page.
  • 09:34 We added more memory to our sync daemon.
  • 09:40 We Scaled our sync daemon to 80 replicas to further distribute the load.
  • 09:55 We increased the memory for our exchange broker and changed our replicas to 5 to further unblock queueing.
  • At 09:55 CET, all systems were back online.

Contributing factors:

  • Our Prometheus alerting system needed updated configurations and was slow to alert our engineers.
  • Our monitoring system monitored calculation times and time deviations but failed to alert us of calculation errors.

Action items:

Prometheus and our monitor have been re-configured to detect memory outages and irregular volume. As a result, our recovery time for a similar incident should be below 20 minutes.

Resolved · · February 14 at 9:09 AM (UTC) (3 days earlier)

This incident has been fully resolved. All services are back up. We will continue to monitor the services.

Monitoring · · February 14 at 9:07 AM (UTC) (2 minutes earlier)

Station Database is back up. The routing engine is still showing degraded performance which might result in slower routes.

Identified · · February 14 at 9:00 AM (UTC) (7 minutes earlier)

Chargetrip GO is fully operational again.

Identified · · February 14 at 8:59 AM (UTC) (37 seconds earlier)

Our tileservice has been restored.

Identified · · February 14 at 8:38 AM (UTC) (21 minutes earlier)

The issue has been identified and a fix is being implemented.

Investigating · · February 14 at 8:34 AM (UTC) (4 minutes earlier)

We are continuing to investigate this issue.

Investigating · · February 14 at 8:33 AM (UTC) (1 minute earlier)

We are currently investigating this issue.

powered by openstatus.dev