In this post, I share how we built the workflow execution engine and handled complex async state machines using BullMQ, Redis, and LangChain.
Building an autonomous workflow automation platform requires robust queue management and state machine recovery. In Sprint 3 of AutoOps AI, we focused on replacing inline execution with resilient background worker queues.
### 1. The Core Architecture Challenge
When users build multi-step workflow pipelines, individual agent tasks can take anywhere from 500ms to 30 seconds depending on LLM response times and external API calls. Running these synchronously inside HTTP endpoints caused severe gateway timeouts.
We needed a decoupled worker model that could pause workflows mid-step, log execution telemetry, and retry failed API calls automatically without blocking the main event loop.
> **Architectural Principle**: Never process long-running agentic LLM calls directly inside web request handlers. Always offload them to isolated background workers.
### 2. Redis & BullMQ Workflow Queue
We integrated BullMQ with Redis to manage step execution jobs. Each node in the workflow graph maps to a discrete job with snapshot metadata.
```typescript
import { Queue, Worker } from 'bullmq';
import { redisConnection } from '../config/redis';
export const workflowQueue = new Queue('workflow-execution', {
connection: redisConnection,
defaultJobOptions: {
attempts: 3,
backoff: { type: 'exponential', delay: 1000 },
},
});
export const worker = new Worker('workflow-execution', async (job) => {
const { workflowId, stepId, payload } = job.data;
console.log(`[Worker] Executing step ${stepId} for workflow ${workflowId}`);
// Agentic execution runner logic...
}, { connection: redisConnection });
```
### 3. Agent Orchestration Engine
LangChain agents are wrapped in custom schema validators. If an agent produces invalid JSON structured output, a feedback loop sends the schema violation back to the LLM to self-correct.
With Sprint 3 complete, AutoOps AI can execute resilient multi-step workflows with 99.9% queue reliability. Next up in Sprint 4 is real-time WebSocket telemetry and frontend canvas node monitoring!
#Next.js
#LangChain
#Redis
#TypeScript
#Docker