Rate limits
Plan send caps, the test-mode cap, and the per-caller request rate.
Two limits apply before a send is accepted. Plan caps count recipients. A separate request rate limits authenticated HTTP calls. After acceptance, worker send admission still paces delivery.
Plan caps
| Plan | Live recipients per UTC day | Live recipients per UTC month |
|---|---|---|
| Free | 100 | 3 000 |
| Pro | No daily cap | 100 000 |
| Scale | No daily cap | 1 000 000 |
Tenant overrides replace the plan defaults. Zero disables sends for that period. Test keys share a separate tenant cap of 1 000 recipients per UTC day and do not consume live caps. Each recipient counts, including recipients in batch and scheduled requests. A rejected request creates no email row and no usage debit. UTC days reset at midnight. Months reset at midnight on the first day.
GET /usage, with a full or read_only key, reports the counts, the effective limits (null means uncapped) and the reset timestamps. A read_only key can call only that route.
Request rate
Every API key and every dashboard user has a request rate of 10 requests per second, with burst capacity 20. The limit is shared across that caller's authenticated routes. The limiter key is the API key id or the dashboard user id. Replays and rejected authenticated requests count. GET /health and GET /ready are unauthenticated and do not spend this limit.
Wait for the integer Retry-After seconds. The value is at least 1. For a send cap it is the time until the UTC reset. The error hint names the cap and reset, or the request rate. A 429 is not stored as an idempotent response. Retry with the same Idempotency-Key after the wait.
Every authenticated response, including errors and replays, carries:
X-RateLimit-Limit: request rate per second, 10.X-RateLimit-Remaining: remaining requests in the burst capacity.X-RateLimit-Reset: Unix seconds, rounded up, when the burst fully replenishes.
These headers describe request admission. GET /usage describes send caps. If Valkey cannot admit the request, admission fails open and logs a warning. The headers then report capacity 20 and a reset two seconds ahead, as an estimate. Postgres send caps still enforce the hard stop. The worker send limiter stays fail closed and does not spend these request tokens.
Readiness
GET /health reports liveness. GET /ready checks configured dependencies within 500 ms and returns 503 when one is unavailable.
Worker send admission
The default is 50 sends per second per tenant, with a burst equal to the rate. Operators set SENDTIER_TENANT_SEND_RATE. Test-mode jobs spend tenant tokens, so test traffic shares this admission limit. Live sends also share a regional limit of 80% of SES SendQuota.MaxSendRate, with a burst of one. When admission is refused, the worker waits and keeps the queued job. This is separate from the HTTP request rate and the plan caps.
Sources: rate_limited (docs/errors/rate_limited.md), ADR-0016 (docs/adr/0016-send-rate-limiting.md), plan caps (internal/domain/plans.go), request rate (internal/limiter/request.go), usage (internal/api/usage.go), readiness (internal/api/system.go), request admission (internal/api/request_limits.go), configuration (internal/config/config.go), worker limiter (internal/limiter/limiter.go), API middleware (internal/api/server.go).