ResourcesRate Limits

Rate Limits

Understanding Flameup API rate limits and quotas

Overview

Rate limits are applied in scopes. Which scope a request falls into depends on the route and how the request is authenticated — not on a single global pool.

The two that matter for an API-key integration are ingestion at 12,000/min and reads at 1,200/min, both counted per workspace. Writes are deliberately the generous one: reads are a tenth of that, so a polling loop will hit its ceiling long before an event firehose does.

Deletes are the exception to that split: DELETE /people/{id} is metered as a read, not a write, so a bulk erasure job draws down the 1,200/min budget rather than the generous ingestion one.

Limits

ScopeLimitBurstCounted perApplies to
Ingestion12,000/min2,000Workspace/track, /track/batch, /identify, people writes, push sends (/push/send, /push/send/batch, token validation)
Device registration6,000/min1,000Workspace/people/{id}/devices register and unregister
Reads1,200/min300WorkspacePeople, device and event reads, deletes, campaign triggers, push reads
Dashboard1,200/min300UserFirebase-authenticated dashboard sessions
Unauthenticated120/min60IPPublic routes, e.g. invitation lookup
Provider webhooks600/min300IPInbound email (Resend) and SMS (Twilio) delivery callbacks
Edge backstop6,000/min600IPEvery request, as a floor under the above
Ingestion and device registration are metered separately, on purpose. A burst of events cannot use up the budget that registers push tokens — a token that fails to register means every later notification silently goes nowhere.

Limits are per workspace, not per key: three keys in one workspace share one ingestion budget.

Burst

Each scope has a burst allowance on top of its per-minute rate. Burst is what a short spike can consume before the sustained rate applies — 2,000 ingestion requests can arrive at once, then the 12,000/min rate governs.

Rate Limit Headers

Rate limit headers are returned only on 429 Too Many Requests responses — successful responses do not include them:

X-RateLimit-Limit: 12000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1705320660
X-RateLimit-Scope: ws-ingest
Retry-After: 37
HeaderDescription
Retry-AfterSeconds to wait. This is the one to honour — the X-RateLimit-* trio is convention, Retry-After is the standard clients and proxies already understand. Capped at 3600.
X-RateLimit-LimitThe per-minute limit of the scope you hit
X-RateLimit-RemainingAlways 0 — the header is only sent on a 429
X-RateLimit-ResetUnix timestamp when the window resets
X-RateLimit-ScopeWhich budget was exhausted. A ws-* scope means your workspace's budget; edge means the per-IP backstop under everything.
Because the headers only appear on a 429, you cannot use them to pace yourself in advance. Back off when you see one rather than trying to track a budget.

Handling Rate Limits

When you exceed the rate limit, you'll receive a 429 Too Many Requests response:

{
  "error": "Rate limit exceeded",
  "message": "Too many requests. Please try again later.",
  "scope": "ws-ingest",
  "gate": "local"
}

scope names the budget you exhausted, matching X-RateLimit-Scope — that is the field that tells you whose budget it was (ws-ingest is your workspace; edge is the per-IP backstop). gate says which enforcement layer of that same budget rejected you: local is the replica's own token bucket, shared is the cross-replica counter that coordinates the identical budget between replicas. Both gates mean the same thing for you — back off. Quote both fields if you contact support.

Implementing Backoff

The headers only appear on a 429, so there is nothing to pace against in advance. React to the rejection instead:

const MAX_RETRIES = 5; // retries after the first attempt, so 6 requests at worst

async function request(url, options = {}, attempt = 0) {
  const response = await fetch(url, {
    ...options,
    headers: { Authorization: `Bearer ${API_KEY}`, ...options.headers }
  });

  if (response.status === 429 && attempt < MAX_RETRIES) {
    // Retry-After is in seconds and is the header to trust.
    const retryAfter = parseInt(response.headers.get('Retry-After') || '0');
    const wait = (retryAfter || 2 ** attempt) * 1000;
    await new Promise(r => setTimeout(r, wait));
    return request(url, options, attempt + 1);
  }

  return response;
}

Best Practices

Instead of making individual requests, batch when possible:

// Bad: 100 individual requests
for (const user of users) {
  await flare.createPerson(user);
}

// Good: 1 batch request
await flare.batchUpsertPeople(users);

Batch endpoints let you do more work per request, which helps you stay within your workspace's per-minute budget.

Requesting Higher Limits

If you need higher rate limits:

  1. Optimize your integration - Use the batch endpoints and cache reads
  2. Contact support - Explain your use case for a limit increase