Rate Limits
Request budgets per plan, burst behaviour, the headers to read and how to back off.
Every API key has a sustained rate in requests per minute and a burst bucket that refills continuously. Reads against cached endpoints such as quotes and screener results are cheap; bar history is weighted by the number of bars returned, so one request for five thousand bars costs considerably more than five requests for a hundred. Streaming connections are counted separately, per key, as concurrent subscriptions.
Three response headers tell you where you stand: x-algobeam-limit is the ceiling for the window, x-algobeam-remaining is what is left, and x-algobeam-reset is the unix second at which the window rolls. Read them on every response rather than counting requests yourself — retries, redirects and cached hits all affect the count in ways a local counter will get wrong.
A 429 response includes a retry-after header in seconds. Honour it, then apply exponential backoff with jitter on top for any further failures. Hammering a limited key extends the penalty window, and a fleet of workers that all retry on the same schedule will synchronise into a thundering herd that keeps itself limited indefinitely.
Most limit problems are cache problems in disguise. Indicator values change only when a bar closes, so a time-to-live matched to the timeframe removes almost all repeat traffic. Batch endpoints accept up to two hundred symbols per call and count as a single weighted request, which is dramatically cheaper than a loop. If you genuinely need more sustained throughput, plan limits are listed on the pricing page and higher tiers are available for teams.