Experiencing an issue that is not shown here? Please contact our support team. Account-specific sending or deliverability issues may not be displayed on this page.
On 23 September, Lettermint stopped accepting new message work during two periods: approximately 18:51–18:57 and 19:02–19:04 CEST. A RabbitMQ disk alarm caused both interruptions. We then carried out emergency maintenance. Some Sending API requests failed during that work, while other requests succeeded.
We are very sorry for the disruption.
During the two intake interruptions, API and SMTP submissions could fail or time out. Some tracking requests failed, and webhook delivery was delayed. Our logs also show temporary errors on three incoming SMTP attempts.
During the later maintenance, some Sending API requests returned errors until about 20:04 CEST. This was intermittent. We retried queued jobs after recovery.
The public incident remained open until 22:01 CEST while we monitored the services. It was not a continuous outage for that full period.
Two message-processing queues had jobs that remained unacknowledged. Their queue logs grew while older log data remained in use. Each RabbitMQ broker had a volume and was configured to keep at least some GB free. When free space crossed that limit, RabbitMQ protected the queued data by stopping new message publishing across the cluster.
Our disk warnings did not give us advance notice. The volume warning required less than 15% free space, with a forecast that the volume would fill. At the first RabbitMQ alarm, the volumes were only about 47% used. We also had an alert for an active RabbitMQ resource alarm, but that alert could fire only after RabbitMQ had stopped publishing. We did not alert on the remaining space above RabbitMQ's own limit.
The RabbitMQ version we ran could also retain old queue log data when messages remained unacknowledged. RabbitMQ 4.3 improves log cleanup in this case.
We expanded the volumes to a larger storage size, one broker at a time. We reduced RabbitMQ's free disk limit to a lower, still safe, amount and upgraded the brokers from to 4.3.6.
All three brokers now run 4.3.6, all three volumes are healthy, and no broker disk alarm is active.
All times are CEST.
18:51: The first disk alarm stopped new message intake.
18:52: We reported the incident.
18:57: The alarm cleared, and intake resumed.
19:02: A second disk alarm stopped intake.
19:04: The second alarm cleared, and intake resumed.
19:33: We announced emergency maintenance.
About 19:39–19:57: We expanded the broker volumes, one at a time. Some Sending API requests failed during maintenance.
About 20:03–20:10: We upgraded the brokers to RabbitMQ 4.3.6. The main period of maintenance-related API errors ended at about 20:04.
20:23: We moved the public incident to monitoring.
22:01: We resolved the incident.
We have increased broker storage, changed the free disk limit and upgraded RabbitMQ.
We will add an early warning based on the space remaining above RabbitMQ’s configured limit, not only the percentage of a volume that is used. We will also monitor sustained queue log growth and old unacknowledged jobs, and update the procedure for storage expansion that requires a broker to be offline.
We are sorry that our warning came after message intake had already stopped and are always aiming to do better.