Rate Limits
Understanding Flameup API rate limits and quotas
Overview
Rate limits are applied in scopes. Which scope a request falls into depends on the route and how the request is authenticated — not on a single global pool.
The two that matter for an API-key integration are ingestion at 12,000/min and reads at 1,200/min, both counted per workspace. Writes are deliberately the generous one: reads are a tenth of that, so a polling loop will hit its ceiling long before an event firehose does.
Deletes are the exception to that split: DELETE /people/{id} is metered as a
read, not a write, so a bulk erasure job draws down the 1,200/min budget
rather than the generous ingestion one.
Limits
| Scope | Limit | Burst | Counted per | Applies to |
|---|---|---|---|---|
| Ingestion | 12,000/min | 2,000 | Workspace | /track, /track/batch, /identify, people writes, push sends (/push/send, /push/send/batch, token validation) |
| Device registration | 6,000/min | 1,000 | Workspace | /people/{id}/devices register and unregister |
| Reads | 1,200/min | 300 | Workspace | People, device and event reads, deletes, campaign triggers, push reads |
| Dashboard | 1,200/min | 300 | User | Firebase-authenticated dashboard sessions |
| Unauthenticated | 120/min | 60 | IP | Public routes, e.g. invitation lookup |
| Provider webhooks | 600/min | 300 | IP | Inbound email (Resend) and SMS (Twilio) delivery callbacks |
| Edge backstop | 6,000/min | 600 | IP | Every request, as a floor under the above |
Limits are per workspace, not per key: three keys in one workspace share one ingestion budget.
Burst
Each scope has a burst allowance on top of its per-minute rate. Burst is what a short spike can consume before the sustained rate applies — 2,000 ingestion requests can arrive at once, then the 12,000/min rate governs.
Rate Limit Headers
Rate limit headers are returned only on 429 Too Many Requests responses — successful responses do not include them:
X-RateLimit-Limit: 12000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1705320660
X-RateLimit-Scope: ws-ingest
Retry-After: 37
| Header | Description |
|---|---|
Retry-After | Seconds to wait. This is the one to honour — the X-RateLimit-* trio is convention, Retry-After is the standard clients and proxies already understand. Capped at 3600. |
X-RateLimit-Limit | The per-minute limit of the scope you hit |
X-RateLimit-Remaining | Always 0 — the header is only sent on a 429 |
X-RateLimit-Reset | Unix timestamp when the window resets |
X-RateLimit-Scope | Which budget was exhausted. A ws-* scope means your workspace's budget; edge means the per-IP backstop under everything. |
429, you cannot use them to pace yourself in advance. Back off when you see one rather than trying to track a budget.Handling Rate Limits
When you exceed the rate limit, you'll receive a 429 Too Many Requests response:
{
"error": "Rate limit exceeded",
"message": "Too many requests. Please try again later.",
"scope": "ws-ingest",
"gate": "local"
}
scope names the budget you exhausted, matching X-RateLimit-Scope — that is
the field that tells you whose budget it was (ws-ingest is your workspace;
edge is the per-IP backstop). gate says which enforcement layer of that
same budget rejected you: local is the replica's own token bucket, shared
is the cross-replica counter that coordinates the identical budget between
replicas. Both gates mean the same thing for you — back off. Quote both fields
if you contact support.
Implementing Backoff
The headers only appear on a 429, so there is nothing to pace against in
advance. React to the rejection instead:
const MAX_RETRIES = 5; // retries after the first attempt, so 6 requests at worst
async function request(url, options = {}, attempt = 0) {
const response = await fetch(url, {
...options,
headers: { Authorization: `Bearer ${API_KEY}`, ...options.headers }
});
if (response.status === 429 && attempt < MAX_RETRIES) {
// Retry-After is in seconds and is the header to trust.
const retryAfter = parseInt(response.headers.get('Retry-After') || '0');
const wait = (retryAfter || 2 ** attempt) * 1000;
await new Promise(r => setTimeout(r, wait));
return request(url, options, attempt + 1);
}
return response;
}
import time, requests
MAX_RETRIES = 5 # retries after the first attempt, so 6 requests at worst
def request(method, url, attempt=0, **kwargs):
headers = {'Authorization': f'Bearer {API_KEY}', **kwargs.get('headers', {})}
r = requests.request(method, url, **{**kwargs, 'headers': headers})
if r.status_code == 429 and attempt < MAX_RETRIES:
retry_after = int(r.headers.get('Retry-After', 0))
time.sleep(retry_after or 2 ** attempt)
return request(method, url, attempt + 1, **kwargs)
return r
Best Practices
Instead of making individual requests, batch when possible:
// Bad: 100 individual requests
for (const user of users) {
await flare.createPerson(user);
}
// Good: 1 batch request
await flare.batchUpsertPeople(users);
Batch endpoints let you do more work per request, which helps you stay within your workspace's per-minute budget.
Requesting Higher Limits
If you need higher rate limits:
- Optimize your integration - Use the batch endpoints and cache reads
- Contact support - Explain your use case for a limit increase