baldur.llm — LLM client wrap
One line around an LLM SDK client: every call it makes shares the provider's
waits across workers, retries on Baldur's coordinated ladder instead of the
SDK's own, moves to the next endpoint when waiting cannot save it, and raises
LLMUnavailableError when no endpoint answers. Covers the OpenAI Python SDK
(and every OpenAI-compatible server it reaches), the Anthropic SDK and the
Google Gen AI SDK (google-genai).
baldur.llm is loaded on first access, so import baldur does not import it.
How the provider's answers are classified, and what happens to a job no
endpoint could answer, is on the LLM API rate limits
page.
wrap
wrap(
client: Endpoint,
*,
fallbacks: Sequence[Any] = (),
name: str | None = None,
timeout: Any = _UNSET
) -> Any
wrap(
client: ClientT,
*,
fallbacks: Sequence[Any] = (),
name: str | None = None,
timeout: Any = _UNSET
) -> ClientT
wrap(
client: Any,
*,
fallbacks: Sequence[Any] = (),
name: str | None = None,
timeout: Any = _UNSET
) -> Any
Wrap an LLM SDK client so every call it makes runs under Baldur.
The call sites do not change: wrapped.chat.completions.create(...) is
the SDK call, sent to the first endpoint that answers. Each endpoint gets
its own shared wait (a provider's rate limit or overload is waited out by
every worker at once, for at least as long as the provider asked), its own
retry ladder and its own circuit breaker. A call that waiting cannot save
moves to the next endpoint; a request the provider rejected as invalid is
neither retried nor moved; when no endpoint answers the call raises
LLMUnavailableError.
Supported SDKs: the OpenAI Python SDK (which also reaches Azure OpenAI and
every OpenAI-compatible server), the Anthropic SDK and the Google Gen AI
SDK (google-genai). Each OpenAI or Anthropic client is copied with the
SDK's own retries switched off (with_options(max_retries=0)), so
Baldur's coordinated retry is the only retry loop; the client you pass in
is not changed. Plain attributes and client-level methods
(wrapped.close()) are the primary client's own.
Only a call with a model= keyword moves between endpoints; other calls
run on the primary alone, still waited, retried and broken. A helper that
sends its request after it returns (stream, with_streaming_response)
is the primary client's own, as you passed it, SDK retries included.
The wrapped object is not an instance of the SDK client's class: a library that type-checks its client argument needs the raw client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Any
|
The primary SDK client, or an :class: |
required |
fallbacks
|
Sequence[Any]
|
Further endpoints, tried in order — clients of the same SDK,
each either bare or in an :class: |
()
|
name
|
str | None
|
The primary endpoint's name — see :class: |
None
|
timeout
|
Any
|
A request timeout applied to every OpenAI or Anthropic
endpoint, as the SDK's own |
_UNSET
|
Returns:
| Type | Description |
|---|---|
Any
|
An object that answers like |
Raises:
| Type | Description |
|---|---|
ValueError
|
An endpoint uses a different SDK, or a sync client is mixed
with an async one, or two endpoints would always share one name —
sharing, the fallback would wait out the primary's wait and stand
behind its breaker, and never answer where the primary did not.
Give one of them |
Endpoint
Endpoint(
client: Any,
*,
model: str | None = None,
name: str | None = None
)
One place a wrapped call can be sent.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Any
|
An SDK client — the same SDK as the wrap's other endpoints. |
required |
model
|
str | None
|
The model this endpoint serves. A call sent here has its
|
None
|
name
|
str | None
|
The name this endpoint's breaker, shared wait and metrics are
kept under. Unset, it is derived from the client's host and the
model: |
None
|