Skip to content

Celery Quickstart

Protect Celery tasks with Baldur, and let Baldur run its own scheduled maintenance on the beat you already have.

Supports Python 3.11–3.13 and Celery 5.3+ (5.4 in CI). Assumes you have a working Celery app.

Baldur meets Celery in two places: the same @baldur.protected facade wraps a task body, and an app-wide signal integration records every task's health for you without touching each task.

1. Install

pip install baldur-framework[celery]

2. Protect a task body

@baldur.protected composes into a Celery task exactly as into any function. Declare it just below @app.task so Celery registers the wrapped function:

import baldur
from celery import Celery

app = Celery("myproject")

@app.task
@baldur.protected("charge-customer", retry=True)
def charge_customer(order_id):
    return payment_gateway.charge(order_id)

The call now travels through a circuit breaker and retry. Add fallback= for a safe default, idempotency_key="order_id" to dedup a re-delivered task, or dlq=True to set a final failure aside for replay. See Composing with @baldur.protected.

3. Observe every task automatically

Rather than decorate each task, connect Baldur to Celery's task signals once — in the module that builds your Celery app. Task failures, retries, and successes then feed Baldur's circuit breaker and metrics on their own, trace/actor context is carried across the enqueue → execute hop, and each worker process initializes Baldur at startup — reading your storage configuration and starting Baldur's own background maintenance, exactly as the Django, FastAPI, and Flask adapters do:

from celery import Celery
from baldur.adapters.celery import setup_baldur_signals

app = Celery("myproject")

setup_baldur_signals(
    app=app,
    task_domain_mapping={
        "myproject.tasks.charge_customer": "payment",
        "myproject.tasks.sync_inventory": "inventory",
    },
)

task_domain_mapping groups tasks under a shared circuit-breaker / metric domain (every payment task trips one breaker). An unmapped task is matched against built-in keyword patterns first (a name containing pay, order, stock, email, … lands in that domain), then falls under the first meaningful segment of its dotted name — myproject.tasks.sync_products becomes domain myproject, not the full task name — so map explicitly any task whose grouping matters. The circuit-breaker, dead-letter, metrics, and forensic-capture hooks each toggle independently (cb_enabled / dlq_enabled / metrics_enabled / forensics_enabled), all on by default. Failure capture and replay ship in the OSS core — the queue lives in process memory by default and in Redis once one is configured; PRO adds the operate-at-scale surface on top (batch replay from the console, adaptive pacing, archive/purge).

4. Run Baldur's scheduled maintenance on your beat

Baldur relies on a handful of background jobs to heal itself — circuit-breaker recovery probes, dead-letter archival and cleanup, expired-override cleanup, metric collection. On a single host it elects itself with a local lock and runs them out of the box. Across multiple hosts, hand them to your existing Celery beat instead — one call injects Baldur's schedule (and its queues and routes) into your app:

from baldur.adapters.celery import configure_baldur_celery

configure_baldur_celery(app)

This arms the same worker startup as step 3, so it is a complete setup on its own if you skipped that step.

Then run beat and a worker as usual:

celery -A myproject beat -l info
celery -A myproject worker -l info

Each job lane is opt-out through an include_* flag, and multi-service deployments can isolate queues with queue_prefix=. One lane composes itself conditionally: the canary watchdog appears only when the PRO distribution is installed and entitled, and its two mutating actions — automatic stage promotion and automatic rollback — stay off until you opt in (see Canary watchdog). The multi-worker coherence runbook walks through the single-host-lock vs. distributed-beat decision.

See Baldur's events

The worker behaves as a plain Celery worker: its -l level, handlers and stdout redirection apply, and Baldur's events go through Celery's handler in Celery's format. Set BALDUR_LOG_LEVEL=INFO to watch circuit breaker and retry events as your tasks run:

export BALDUR_LOG_LEVEL=INFO   # circuit opened/closed, retries, ...

A few JSON lines from baldur.init() precede Celery's logging setup at boot — Baldur initialises at worker_init, before Celery configures the worker's logging. To keep Baldur's JSON handler on the worker instead, set worker_hijack_root_logger = False on the app or connect your own setup_logging receiver.

Going to production

The in-memory fallback is single-worker only

Celery almost always runs more than one worker, so this matters from the first deploy. The zero-config path keeps circuit breaker state, idempotency keys, and counters in a per-process store — across workers they diverge silently, which breaks correctness, not just scale. This is a hazard, not a tuning knob: give Baldur a shared backend before you run a second worker.

Point Baldur at Redis so every worker shares state. With step 3 or step 4 in place, no further code changes — set one environment variable before starting the workers:

pip install baldur-framework[celery,redis]
export BALDUR_REDIS_URL=redis://localhost:6379/0
export BALDUR_ENVIRONMENT=production

Reading those variables is baldur.init()'s job, and steps 3 and 4 each run it in every worker for you. If you use neither and rely on @baldur.protected alone, the FAQ has the one line that does it — without it Baldur keeps protecting calls, on per-process state.

Declaring the environment is what turns the hazard above into a rule Baldur enforces: with BALDUR_ENVIRONMENT=production set and BALDUR_REDIS_URL missing, baldur.init() refuses to start rather than let a shared guarantee degrade to per-worker memory. For a deliberate single-worker deployment on in-memory state, BALDUR_TEST_MODE=true opts out of that check and of every other production check (the write-ahead log directory, backend construction, the PRO requirements below), and keeps every store in per-process memory even if you set a backend later.

A worker container that cannot write /var/log/baldur keeps Baldur's write-ahead log in a writable fallback directory and says so in a warning at startup; one where nothing is writable (a read-only root filesystem with no writable mount) refuses to start. Set BALDUR_RESILIENT_STORAGE_WAL_DIR to point the log at a volume. With PRO active, production additionally needs BALDUR_SECRETS_AUDIT_SIGNING_KEY and BALDUR_SQL_DSN (or, in a Django project, its DATABASES) — see Environment Variables.

See also