rate_limiter
langroid/language_models/rate_limiter.py
Pro-active rate limiting for OpenAI-compatible chat-completion calls.
This module is standalone: it knows nothing about Langroid agents, and it does
not touch the retry-with-exponential-backoff logic in
langroid.language_models.utils, which stays as the reactive fallback.
The limiter discovers the account's actual limits instead of asking the user to configure them. OpenAI (and most OpenAI-compatible gateways) report the live budget on every response::
x-ratelimit-limit-requests / x-ratelimit-remaining-requests
x-ratelimit-limit-tokens / x-ratelimit-remaining-tokens
x-ratelimit-reset-requests / x-ratelimit-reset-tokens
reset is the time until the bucket refills to limit, so the sustained
refill rate implied by a single response is::
rate = (limit - remaining) / reset
which is exactly the account's limit for that model, in requests (or tokens)
per second. The limiter paces sends at rate * (1 - headroom) so a request is
held briefly rather than rejected.
For providers that send no such headers (local servers, Groq/Cerebras, litellm) the limiter falls back to an additive-free AIMD scheme: multiply the inter-send interval up on a 429, decay it down on every success. No prior knowledge of any limit is needed.
Limiters are shared process-wide by key (see get_rate_limiter), so the many
cloned agents of a run_batch_tasks job pace against one budget.
Known limitation: until the first response arrives there are no headers to learn from, so an initial burst of concurrent calls is sent unpaced; pacing engages from the first observed response onwards.
RateLimitSnapshot
¶
Bases: BaseModel
One provider-reported view of the remaining rate-limit budget.
has_rate_limit_info
property
¶
Did the provider report any rate-limit budget at all?
request_refill_rate
property
¶
Implied sustained limit, in requests/second (None if unknowable).
token_refill_rate
property
¶
Implied sustained limit, in tokens/second (None if unknowable).
from_headers(headers)
classmethod
¶
Build a snapshot from HTTP response headers (case-insensitive).
Source code in langroid/language_models/rate_limiter.py
RateLimitConfig
¶
Bases: BaseSettings
Settings for the pro-active rate limiter.
Every field can be overridden by an env var with the
LANGROID_RATE_LIMIT_ prefix, e.g. LANGROID_RATE_LIMIT_ENABLED=1.
RateLimiter(config=None, name='')
¶
Paces sends so provider rate limits are approached, not exceeded.
Thread-safe and asyncio-safe: state is guarded by a threading.Lock held
only for the (non-blocking) bookkeeping, while the wait itself happens
outside the lock via time.sleep or asyncio.sleep.
Slots are handed out as reservations off a shared monotonic clock, so N concurrent callers queue rather than all retrying at once.
Source code in langroid/language_models/rate_limiter.py
acquire()
¶
Block until it is this caller's turn to send. Returns seconds waited.
acquire_async()
async
¶
Async variant of acquire. Returns seconds waited.
observe_response(headers=None, tokens_used=None)
¶
Take in what a response reveals about the budget.
Information only: callers may call this more than once per request (the
headers arrive with the response, a streaming request's token usage
only later), so it must NOT be where the AIMD recovery is applied --
see observe_success, which is called exactly once per request.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
headers
|
Optional[Mapping[str, Any]]
|
Response headers, if the provider/transport exposes them. |
None
|
tokens_used
|
Optional[int]
|
Total tokens billed for this request, if reported. Used as an EWMA estimate of the per-request token cost, which is what makes token-budget pacing possible. It is an estimate, not a guarantee of provider quota compliance. |
None
|
Source code in langroid/language_models/rate_limiter.py
observe_success(tokens_used=None)
¶
Record that one request was accepted; recover the send rate.
Call this EXACTLY once per request that the provider accepted. It is the only place the header-free AIMD interval decays, so calling it twice for one request would halve the backoff twice over.
Source code in langroid/language_models/rate_limiter.py
observe_rate_limit_error(headers=None)
¶
Learn from a 429: back off, and stall until the budget recovers.
Source code in langroid/language_models/rate_limiter.py
stats()
¶
Counters and discovered state, for tests and diagnostics.
sends is the "did it run at all" signal: zero means the limiter was
never consulted, which is different from "it was consulted and never
needed to wait" (sends > 0, waits == 0).
Source code in langroid/language_models/rate_limiter.py
parse_reset_duration(value)
¶
Parse an OpenAI x-ratelimit-reset-* duration into seconds.
The header is a Go-style duration, e.g. "120ms", "1s", "6m0s",
"1h2m3s", "7.66s". A bare number is read as seconds, which is what
some OpenAI-compatible gateways send.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
value
|
Optional[str]
|
Raw header value, or None. |
required |
Returns:
| Type | Description |
|---|---|
Optional[float]
|
Duration in seconds, or None if |
Source code in langroid/language_models/rate_limiter.py
get_rate_limiter(key, config=None)
¶
Get (creating if needed) the process-wide limiter for key.
Sharing by key is what lets the cloned agents of a batch job pace against one budget. The config of the first caller for a given key wins; later callers get the existing limiter unchanged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key
|
str
|
Sharing key, e.g. |
required |
config
|
Optional[RateLimitConfig]
|
Settings to use if the limiter does not exist yet. |
None
|
Returns:
| Type | Description |
|---|---|
RateLimiter
|
The shared |
Source code in langroid/language_models/rate_limiter.py
reset_rate_limiters()
¶
rate_limit_error_headers(exc)
¶
Classify an exception as a rate-limit error and return its headers.
Returns:
| Type | Description |
|---|---|
Optional[Mapping[str, Any]]
|
None if |
Optional[Mapping[str, Any]]
|
headers, which may be an empty mapping if the provider sent none. |
Optional[Mapping[str, Any]]
|
An empty mapping is therefore meaningfully different from None. |