ADR-030: Operational and Infrastructure Task Queues

Status

Accepted

Context

Smarter’s Celery tasks are of two kinds:

  • Operational tasks record what happens on the platform: a prompt’s history, tool calls and plugin usage, charges, budgets, and the aggregation of these records. They are quick, there are many of them, and users depend on them: a conversation whose history is not recorded loses its context, and a charge that is not recorded escapes its budget.

  • Infrastructure tasks deploy, verify and destroy cloud, Kubernetes and DNS resources: an LLMClient’s DNS record, ingress and TLS certificate, a custom domain, an LLMHost’s node group, a vectorstore’s database. A certificate takes minutes to be issued, and a DNS record or a domain’s name servers take up to a day to propagate.

Both kinds shared one queue and one pool of workers, and the infrastructure tasks waited by sleeping: deploy_default_api slept 10 minutes and then polled its certificate for 30 more, and verify_domain and verify_custom_domain slept between checks for up to 4 and 24 hours. Deploying the built-in LLMClients occupied every worker for half an hour, until Celery’s task time limit killed them, while the operational tasks waited behind them. Conversations lost their history, and failed on their next prompt.

Decision

  1. Separate queues and workers. Operational tasks run in smarter_settings.llmclient_tasks_celery_task_queue (default_celery_task_queue), served by the smarter-worker. Infrastructure tasks run in smarter_settings.infrastructure_tasks_celery_task_queue (infrastructure_celery_task_queue), served by its own smarter-worker-infrastructure. A deployment can never block an operational task.

  2. Never wait in a task. A task that waits for something outside Smarter checks once, and if it is not ready, schedules itself to check again later, with apply_async(countdown=...), passing the number of attempts made. It does not time.sleep(), so it holds a worker only while it works.

  3. Each task declares its queue in its @app.task(queue=...), and each Celery Beat entry names the same queue in its options, because Beat sends a task by name, without importing it. A queue that no worker consumes is never processed.

Alternatives Considered

  • More workers in one queue. It postpones the problem: enough deployments still occupy every worker.

  • Celery task priorities. They reorder waiting tasks, but cannot free a worker that a task holds.

  • Separate queues only. Infrastructure tasks would still hold their own workers while they sleep, and run into Celery’s task time limit, so a slow DNS propagation would fail a deployment.

Consequences

  • Positive: - Operational tasks run promptly, however many deployments are in progress. - A deployment that waits hours for DNS holds no worker, and is not killed by the task time limit. - The two kinds of task can be scaled, monitored and restarted independently.

  • Negative: - One more worker Deployment (templates/deployment-worker-infrastructure.yaml) and docker-compose service. - A new task must be assigned to the right queue, and a task that waits must be written as a sequence

    of checks rather than a loop.