---
title: "Service Outage"
description: "Resolved · affects Station, Route (legacy), Tile, Chargetrip GO · Feb 14, 2023"
base_url: "https://chargetrip.openstatus.dev"
canonical: "https://chargetrip.openstatus.dev/events/report/3487"
homepage_url: "https://www.chargetrip.com"
contact_url: "https://www.chargetrip.com/contact/support"
---

# Service Outage

[Status](/.md) › [Events](/events.md) › Service Outage

Feb 14, 2023 · 4 years ago · affects: Station, Route (legacy), Tile, Chargetrip GO · 3 days

## Updates

### + Resolved — Feb 17, 11:26 AM

# Service outage on Tuesday, February 14 2023

## **Summary:**

In the early morning of February 14, Chargetrip received a high volume of requests that started to back up our exchange broker queue. This resulted in our exchange broker seizing, a service slowdown, and, ultimately, a service outage. All customers were affected either with slower calculation times or a complete service outage. Chargetrip recovered all systems within three hours.

## **Impact:**

All customers were affected by this outage by slow calculation times or denial of service. Affected services were: the routing engine, station database, vehicle database, tile service and go.chargetrip.com.

## **Timeline:**

* Around 07:30 CET, we started receiving large volumes of messages that would ultimately seize our exchange broker.
* Around 09:00 CET, our internal warning system began sending alerts to our DevOps team.
* 09:28 We noticed our exchange broker was using an incredible amount of memory \(1.4M messages in sync stations queue\).
* Around 09:30 CET, our resources were increased to unblock our queue.
* 09:32 We scaled our sync daemon to 40 replicas.
* 09:33 We updated our status page.
* 09:34 We added more memory to our sync daemon.
* 09:40 We Scaled our sync daemon to 80 replicas to further distribute the load.
* 09:55 We increased the memory for our exchange broker and changed our replicas to 5 to further unblock queueing.
* At 09:55 CET, all systems were back online.

## **Contributing factors:**

* Our Prometheus alerting system needed updated configurations and was slow to alert our engineers.
* Our monitoring system monitored calculation times and time deviations but failed to alert us of calculation errors.

## **Action items:**

Prometheus and our monitor have been re-configured to detect memory outages and irregular volume. As a result, our recovery time for a similar incident should be below 20 minutes.

### + Resolved — Feb 14, 9:09 AM

affects: Station (operational), Route (legacy) (operational), Tile (operational), Chargetrip GO (operational)

This incident has been fully resolved. All services are back up. We will continue to monitor the services.

### x Monitoring — Feb 14, 9:07 AM

affects: Station (operational), Route (legacy) (degraded performance), Tile (operational), Chargetrip GO (operational)

Station Database is back up. The routing engine is still showing degraded performance which might result in slower routes.

### x Identified — Feb 14, 9:00 AM

affects: Station (degraded performance), Route (legacy) (partial outage), Tile (operational), Chargetrip GO (operational)

Chargetrip GO is fully operational again.

### x Identified — Feb 14, 8:59 AM

affects: Station (degraded performance), Route (legacy) (partial outage), Tile (operational), Chargetrip GO (partial outage)

Our tileservice has been restored.

### x Identified — Feb 14, 8:38 AM

affects: Station (degraded performance), Route (legacy) (partial outage), Tile (partial outage), Chargetrip GO (partial outage)

The issue has been identified and a fix is being implemented.

### x Investigating — Feb 14, 8:34 AM

affects: Station (degraded performance), Route (legacy) (partial outage), Tile (partial outage), Chargetrip GO (partial outage)

We are continuing to investigate this issue.

### x Investigating — Feb 14, 8:33 AM

affects: Station (degraded performance), Route (legacy) (partial outage), Tile (partial outage)

We are currently investigating this issue.

---

_Powered by [openstatus.dev](https://openstatus.dev)_
