Skip to content

baldur.llm — LLM client wrap

One line around an LLM SDK client: every call it makes shares the provider's waits across workers, retries on Baldur's coordinated ladder instead of the SDK's own, moves to the next endpoint when waiting cannot save it, and raises LLMUnavailableError when no endpoint answers. Covers the OpenAI Python SDK (and every OpenAI-compatible server it reaches), the Anthropic SDK and the Google Gen AI SDK (google-genai).

baldur.llm is loaded on first access, so import baldur does not import it. How the provider's answers are classified, and what happens to a job no endpoint could answer, is on the LLM API rate limits page.

wrap

wrap(
    client: Endpoint,
    *,
    fallbacks: Sequence[Any] = (),
    name: str | None = None,
    timeout: Any = _UNSET
) -> Any
wrap(
    client: ClientT,
    *,
    fallbacks: Sequence[Any] = (),
    name: str | None = None,
    timeout: Any = _UNSET
) -> ClientT
wrap(
    client: Any,
    *,
    fallbacks: Sequence[Any] = (),
    name: str | None = None,
    timeout: Any = _UNSET
) -> Any

Wrap an LLM SDK client so every call it makes runs under Baldur.

The call sites do not change: wrapped.chat.completions.create(...) is the SDK call, sent to the first endpoint that answers. Each endpoint gets its own shared wait (a provider's rate limit or overload is waited out by every worker at once, for at least as long as the provider asked), its own retry ladder and its own circuit breaker. A call that waiting cannot save moves to the next endpoint; a request the provider rejected as invalid is neither retried nor moved; when no endpoint answers the call raises LLMUnavailableError.

Supported SDKs: the OpenAI Python SDK (which also reaches Azure OpenAI and every OpenAI-compatible server), the Anthropic SDK and the Google Gen AI SDK (google-genai). Each OpenAI or Anthropic client is copied with the SDK's own retries switched off (with_options(max_retries=0)), so Baldur's coordinated retry is the only retry loop; the client you pass in is not changed. Plain attributes and client-level methods (wrapped.close()) are the primary client's own.

Only a call with a model= keyword moves between endpoints; other calls run on the primary alone, still waited, retried and broken. A helper that sends its request after it returns (stream, with_streaming_response) is the primary client's own, as you passed it, SDK retries included.

The wrapped object is not an instance of the SDK client's class: a library that type-checks its client argument needs the raw client.

Parameters:

Name Type Description Default
client Any

The primary SDK client, or an :class:Endpoint around one.

required
fallbacks Sequence[Any]

Further endpoints, tried in order — clients of the same SDK, each either bare or in an :class:Endpoint (to pin the model it serves, or name it).

()
name str | None

The primary endpoint's name — see :class:Endpoint.

None
timeout Any

A request timeout applied to every OpenAI or Anthropic endpoint, as the SDK's own timeout option. Unset, the SDK's default stands.

_UNSET

Returns:

Type Description
Any

An object that answers like client.

Raises:

Type Description
ValueError

An endpoint uses a different SDK, or a sync client is mixed with an async one, or two endpoints would always share one name — sharing, the fallback would wait out the primary's wait and stand behind its breaker, and never answer where the primary did not. Give one of them Endpoint(name=...) or a different model.

Endpoint

Endpoint(
    client: Any,
    *,
    model: str | None = None,
    name: str | None = None
)

One place a wrapped call can be sent.

Parameters:

Name Type Description Default
client Any

An SDK client — the same SDK as the wrap's other endpoints.

required
model str | None

The model this endpoint serves. A call sent here has its model= replaced with it, so a fallback can name the model its own provider calls the same job by.

None
name str | None

The name this endpoint's breaker, shared wait and metrics are kept under. Unset, it is derived from the client's host and the model: llm.<host>.<model>.

None