# Alerting when a RabbitMQ queue has no consumers

*Originally published at [warrenops.io](https://warrenops.io/blog/alert-rabbitmq-queue-no-consumers.html).*

A queue with messages and zero consumers is the quietest outage there is. Nothing errors. Producers publish happily. The dashboard shows the broker green. Then the TTL runs out, or the queue hits its limit, and an hour of orders is in a dead-letter queue. Here is how to get paged before that.

## What to alert on

"No consumers" alone is not the alert. Plenty of queues have no consumers by design: retry-wait queues, parking-lot DLQs, queues of a batch job that runs at night. The signal is the combination:

- `consumers == 0`
- and `messages_ready > 0` (something is waiting)
- and it has been like that for a few minutes (a deploy restarts consumers; do not page on a 20-second gap)
- and the queue is not on your list of intentionally consumer-less queues.

Two neighbours of this alert are worth setting up at the same time, because they catch the cases "consumers exist but do nothing":

- **Stuck consumers:** `consumers > 0`, `messages_ready > 0`, and the ack rate is zero for five minutes. The consumer process is alive, holds the connection, and is deadlocked or stuck on a downstream call.
- **Growth:** `messages_ready` grew by more than N in the last M minutes. Consumers are there but cannot keep up.

## Where the numbers come from

### Management HTTP API

The management plugin exposes everything you need per queue. Ask for only the columns you want, otherwise the response for a broker with a thousand queues is huge:

```
curl -s -u monitor:secret \
  'http://rabbit:15672/api/queues?columns=vhost,name,consumers,messages_ready,messages_unacknowledged,message_stats.ack_details.rate'
```

Relevant fields:

| Field | Meaning |
|---|---|
| `consumers` | Number of consumers currently attached to the queue. |
| `messages_ready` | Messages waiting to be delivered. This is the "backlog". |
| `messages_unacknowledged` | Delivered but not yet acked. High and flat means consumers hold messages without finishing them. |
| `message_stats.ack_details.rate` | Acks per second, averaged over the last sampling interval. Zero with a backlog and consumers present is the stuck-consumer signal. |
| `consumer_utilisation` | Fraction of time consumers could take new messages. Low with a backlog means consumers are the bottleneck. |
| `idle_since` | Set when the queue had no activity at all since that time. Useful to ignore truly dead queues. |

The monitoring user needs the `monitoring` tag, nothing more. Do not use an administrator account for a poller.

### Prometheus plugin

Since RabbitMQ 3.8, `rabbitmq_prometheus` ships with the broker. By default `/metrics` is *aggregated*: you get totals per node, not per queue, which is useless for this alert. Per-queue values come from either:

- `/metrics/per-object`, every metric for every object. Fine for a few hundred queues, heavy for thousands.
- `/metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count` (3.9+), only the families you ask for. Metrics from this endpoint are prefixed `rabbitmq_detailed_`.
- Or set `prometheus.return_per_object_metrics = true` in `rabbitmq.conf` to make `/metrics` per-object.

```
rabbitmq-plugins enable rabbitmq_prometheus
curl -s 'http://rabbit:15692/metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count' | grep orders
```

## Prometheus alert rules

With the detailed endpoint scraped as its own job, three rules cover the three signals above. Adjust the prefix (drop `detailed_`) if you use per-object metrics.

```
groups:
- name: rabbitmq-queues
  rules:
  - alert: RabbitMQQueueNoConsumers
    expr: |
      rabbitmq_detailed_queue_consumers == 0
      and on (vhost, queue) rabbitmq_detailed_queue_messages_ready{queue!~".*(\\.wait|\\.dlq|\\.parking)$"} > 0
    for: 5m
    labels: {severity: page}
    annotations:
      summary: "{{ $labels.vhost }}/{{ $labels.queue }} has {{ $value }} messages and no consumer"

  - alert: RabbitMQQueueConsumersStuck
    expr: |
      rabbitmq_detailed_queue_consumers > 0
      and on (vhost, queue) rabbitmq_detailed_queue_messages_ready > 0
      and on (vhost, queue) rate(rabbitmq_detailed_queue_messages_acked_total[5m]) == 0
    for: 5m
    labels: {severity: page}

  - alert: RabbitMQQueueGrowing
    expr: |
      delta(rabbitmq_detailed_queue_messages_ready[15m]) > 1000
    for: 0m
    labels: {severity: warn}
```

The regex on `queue` excludes the queues that have no consumers on purpose; more on that below. Because the exclusion sits on the `messages_ready` side of the `and`, those queues never produce a result, whatever their consumer count.

Route `severity: page` to your on-call and `warn` to a channel. A queue growing at 2 a.m. may be a batch job; a queue with no consumer at 2 a.m. is not.

## No Prometheus? Cron and curl

A twenty-line script on a box that can reach the management API does the job for a small setup. This one posts to a Slack incoming webhook when a queue has a backlog and no consumer, and keeps a marker file so it does not repeat itself every minute:

```
#!/usr/bin/env bash
set -euo pipefail
API=http://rabbit:15672; AUTH=monitor:secret
HOOK=https://hooks.slack.com/services/T000/B000/xxxx
IGNORE='\.(wait|dlq|parking)$'
STATE=/var/tmp/rabbit-noconsumer; mkdir -p "$STATE"

# "vhost/queue<TAB>ready" for every queue that is in trouble right now
ALERTING=$(curl -sf -u "$AUTH" "$API/api/queues?columns=vhost,name,consumers,messages_ready" \
  | jq -r --arg ign "$IGNORE" \
      '.[] | select(.consumers == 0 and .messages_ready > 0 and (.name | test($ign) | not))
           | "\(.vhost)/\(.name)\t\(.messages_ready)"')

# notify once per incident: a marker file per queue
while IFS=$'\t' read -r q ready; do
  [ -n "$q" ] || continue
  key="$STATE/$(printf '%s' "$q" | md5sum | cut -c1-32)"
  [ -e "$key" ] && continue
  printf '%s\n' "$q" > "$key"
  curl -sf -X POST -H 'content-type: application/json' "$HOOK" \
    -d "$(jq -nc --arg t ":rotating_light: $q has $ready messages and no consumer" '{text: $t}')"
done <<< "$ALERTING"

# recovered: drop markers for queues that are no longer alerting
for f in "$STATE"/*; do
  [ -e "$f" ] || continue
  grep -qF "$(cat "$f")" <<< "$ALERTING" || rm -f "$f"
done
```

Run it every minute from cron. The "for 5 minutes" part is missing here; add it by requiring the marker to be older than five minutes before posting, or accept the occasional deploy-time notification. Teams users swap the Slack payload for an Adaptive Card or a plain `{"text": ...}` to an incoming webhook.

## Deciding which queues to exclude

Do this once and write it down, otherwise the alert gets muted the first week:

- **Retry-wait queues** (TTL, dead-letter back to the work queue): never have consumers. Exclude by name pattern.
- **Dead-letter and parking-lot queues**: no consumers, and messages in them are a *different* alert (see below). Exclude here, alert separately.
- **Batch queues** consumed on a schedule: exclude, or alert only outside the schedule.
- **Everything else** with a backlog and no consumer is an incident.

Name your queues so a regex can tell these apart. `orders.process`, `orders.retry.wait`, `orders.dlq` is a convention that makes every rule on this page a one-liner.

## The companion alert: something landed in a DLQ

The no-consumer alert is the early warning. The late warning is `messages_ready > 0` on a queue matching `\.dlq$`, because by then messages have already failed. Give that one a lower threshold than you think: one dead letter in a queue that is normally empty is news, three hundred is a postmortem. If you want fewer false alarms, alert on the rate of arrivals into the DLQ (`message_stats.publish_details.rate` on the DLQ, or `rate(rabbitmq_detailed_queue_messages_published_total[5m])`) rather than on the count.

## Keep the history

Whatever alerts you build, keep the per-queue time series for at least a week. The question after a no-consumer page is always "since when", and the management UI only shows the last hour at any useful resolution. Prometheus gives you this for free. The cron script does not; if you go that route, at least append the numbers to a file per run.

---

## Where Warren fits

If you do not run Prometheus, Warren's Team edition does this out of the box: alert rules match queues by regex on one or all clusters, with conditions for messages above a threshold, no consumers while messages wait, and growth within a window, each with a sustained-for time. Notifications go to Slack, Teams or a JSON webhook, repeat after a configurable interval and resolve when the queue recovers. Every queue is sampled every 30 seconds and kept for seven days, so "since when" has an answer.

[Try Warren in a minute](https://warrenops.io/#deploy) (one compose file, demo broker with real dead letters included) · [Documentation](https://warrenops.io/docs/alert-rules.html)
