> ## Documentation Index
> Fetch the complete documentation index at: https://docs.crewai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Conversational Flows

> Build multi-turn chat apps with handle_turn per turn, message history, intent routing, tracing, and structured streaming.

## Overview

Conversational apps treat each user line as a **new flow run** with the **same session id**. CrewAI adds helpers for message history, optional intent routing, deferred tracing, structured turn streaming, and a local `flow.chat()` REPL.

| Concept | Implementation |
| - | - |
| Session id | `handle_turn(..., session_id=...)` → `kickoff(inputs={"id": ...})` → `state.id` |
| User line | `handle_turn(message)` appends to `state.messages` before the graph runs |
| Turn complete | `conversation_turn_completed`; with default trace deferral, `FlowFinished` waits for `finalize_session_traces()` |
| Full-session trace | `ConversationConfig(defer_trace_finalization=True)` + `finalize_session_traces()` |

## Turn APIs

Use **`flow.handle_turn(message, session_id=...)`** for every user message from REST, WebSocket, tests, and custom UIs. Use **`flow.chat()`** when you want a local terminal chat loop for a conversational `Flow`.

`Flow.kickoff()` does **not** accept `user_message=` or `session_id=` keyword arguments. For conversational flows, `handle_turn()` stores the pending message and calls `kickoff(inputs={"id": session_id})` internally after resetting per-turn execution state.

| API | Use for |
| - | - |
| `handle_turn(message, session_id=...)` | Ergonomic one-turn wrapper for conversational `Flow` |
| `stream_turn(message, session_id=...)` | Stream one conversational turn as ordered runtime frames |
| `chat()` | Local terminal REPL for conversational `Flow` |
| `kickoff(inputs={...})` | Advanced flow execution without conversational turn handling |
| `ask()` | Blocking prompt **inside** one step (wizard, clarification) |
| `@human_feedback` | Approve/reject **a step output** — not the next chat line |

`handle_turn()`, `stream_turn()`, and `chat()` raise `ValueError` unless conversational mode is enabled. Applying `@ConversationConfig(...)` enables it automatically; otherwise set `conversational = True`.

## Quick start

```python theme={null}
from uuid import uuid4

from crewai import Flow
from crewai.flow import listen
from crewai.flow import (
    ConversationConfig,
    ConversationState,
)


@ConversationConfig(defer_trace_finalization=True)
class SupportFlow(Flow[ConversationState]):
    def route_turn(self, context):
        message = (self.state.current_user_message or "").lower()
        if "order" in message:
            return "order"
        if "bye" in message or "goodbye" in message:
            return "goodbye"
        return "help"

    @listen("order")
    def handle_order(self):
        reply = "Your order is on the way."
        self.append_assistant_message(reply)
        return reply

    @listen("help")
    def handle_help(self):
        reply = "How can I help?"
        self.append_assistant_message(reply)
        return reply

    @listen("goodbye")
    def handle_goodbye(self):
        reply = "Goodbye!"
        self.append_assistant_message(reply)
        return reply


session_id = str(uuid4())
flow = SupportFlow()

try:
    flow.handle_turn("Where is my order?", session_id=session_id)
    flow.handle_turn("What about returns?", session_id=session_id)
finally:
    flow.finalize_session_traces()  # one trace link for the whole chat
```

## Streaming a turn

Use `stream_turn()` when a UI or runtime needs structured events for one chat turn. It returns a stream session with ordered frames for Flow routing, LLM chunks, tool activity, and conversation messages.

```python theme={null}
stream = flow.stream_turn("Where is my order?", session_id=session_id)

with stream:
    for frame in stream.events:
        if frame.channel == "llm" and frame.type == "llm_stream_chunk":
            print(frame.content, end="", flush=True)

result = stream.result
```

For the full frame contract and channel list, see [Streaming Runtime Contract](/edge/en/learn/streaming-runtime-contract).

## Turn lifecycle

Each `handle_turn` runs this pipeline:

1. **Turn setup** — stores the pending user message, resolves the session id, resets per-turn execution tracking, and calls `kickoff(inputs={"id": session_id})`.
2. **State restore** — if `inputs["id"]` exists and `@persist` is configured, loads the latest snapshot.
3. **`FlowStarted`** — emitted on the first deferred session turn only.
4. **Pending turn hydration** — appends the user message to `state.messages`, sets `current_user_message` / `last_user_message`, and optionally classifies when `intents` / `default_intents` + `intent_llm` are set.
5. **Graph execution** — user-defined `@start` methods (if any) → `route_conversation` (the built-in start/router) → the selected `@listen` handler. `route_conversation` also calls the overridable `conversation_start()` helper.
6. **End of run** — per-turn `flow_finished` and trace finalization are **skipped** when deferral is enabled; nested `Agent.kickoff()` / crews do not close the parent batch either.

Handlers should call **`append_assistant_message(reply)`** when the visible reply is not the return value, or when you trim history. A public string return is also recorded as assistant and included in the `@persist` snapshot, so a fresh Flow instance restores it. The user line is already stored by `handle_turn` — do not append it again in handlers.

## Configuration overview

Decorating a `Flow` subclass with `ConversationConfig` both attaches the chat defaults and enables conversational mode. See the [full field reference](#conversationconfig) below. Override pre-classification per turn with `handle_turn(..., intents=..., intent_llm=...)`.

## Lower-level `ChatState` helpers

`ChatState`, the legacy `ConversationalConfig`, and `crewai.flow.conversation` helpers are still importable for advanced orchestration, tests, or custom wrappers. They are separate from the `ConversationState` / `ConversationConfig` API and do not add `user_message=` or `session_id=` keyword arguments to `Flow.kickoff()`.

```python theme={null}
from crewai.flow import ChatState


class MyChatState(ChatState):
    # Inherited: id, messages, last_user_message, last_intent, session_ready
    research_turn_count: int = 0
    custom_flag: bool = False
```

| Field | Role |
| - | - |
| `id` | Session UUID (same as `inputs["id"]`) |
| `messages` | `list` of `{role, content}` for LLM history |
| `last_user_message` | Latest user line for this turn |
| `last_intent` | Route label after classification (if used) |
| `session_ready` | One-time bootstrap flag (permissions, caches, etc.) |

`ConversationalInputs` is a `TypedDict` for conventional `kickoff(inputs={...})` keys: `id`, `user_message`, `last_intent`.

`ConversationState` stores `messages` as `ConversationMessage` objects and additionally provides `current_user_message`, `ended`, `events`, and `agent_threads`. Use `conversation_messages` when passing its canonical history to an LLM.

## `Flow` conversational API

### `handle_turn` parameters

| Parameter | Purpose |
| - | - |
| `message` | This turn’s text |
| `session_id` | Conversation UUID → `inputs["id"]` / `state.id` |
| `intents` | Outcome labels for pre-kickoff `classify_intent` |
| `intent_llm` | LLM for classification (required with `intents`) |
| `**kickoff_kwargs` | Forwarded to `kickoff()` for options like `input_files`, `from_checkpoint`, and `restore_from_state_id` |

### `kickoff` parameters

`Flow.kickoff()` accepts `inputs`, `input_files`, `from_checkpoint`, and `restore_from_state_id`. Pass `inputs={"id": session_id}` when you need raw flow execution, but use `handle_turn()` when the call represents a chat message.

### Instance attributes

| Attribute | Purpose |
| - | - |
| `conversational` | Set to `True` to enable the conversational graph and `handle_turn()` |
| `defer_trace_finalization` | Optional instance override. Otherwise `_should_defer_trace_finalization()` reads `ConversationConfig.defer_trace_finalization`. |
| `suppress_flow_events` | Hides console flow panels and suppresses method execution events; flow start/finish events still emit |
| `stream` | Generic Flow streaming flag. For conversational turns, use `stream_turn()` instead of combining this flag with `handle_turn()`. |

### Methods and properties

| Name | Description |
| - | - |
| `append_assistant_message(content)` | Append a user-visible assistant reply to `state.messages` |
| `append_message(role, content, **extra)` | Lower-level append to `state.messages` |
| `conversation_messages` | Read-only history for LLM calls |
| `classify_intent(text, outcomes, *, llm, context=None)` | Map text to one outcome (same collapse logic as `@human_feedback`) |
| `receive_user_message(text, *, outcomes=None, llm=None)` | Append user message; optionally set `last_intent` |
| `finalize_session_traces()` | Emit deferred `flow_finished` and finalize the session trace batch |
| `_should_defer_trace_finalization()` | Advanced/internal hook that resolves whether per-turn trace finalization is deferred |
| `input_history` | Audit trail of `ask()` prompts and responses |

### Module helpers (`crewai.flow.conversation`)

Importable from `crewai.flow.conversation` for tests or custom orchestration. These helpers use the legacy `ConversationalConfig` shape; `prepare_conversational_turn()` also clears `last_intent`, unlike `handle_turn()`, which preserves it as router context.

| Function | Description |
| - | - |
| `normalize_kickoff_inputs(inputs, user_message=..., session_id=...)` | Merge conversational kwargs into `inputs` |
| `get_conversation_messages(flow)` | Read messages from state or internal buffer |
| `append_message(flow, role, content, **extra)` | Same as instance method |
| `prepare_conversational_turn(flow, user_message=..., intents=..., intent_llm=..., config=...)` | Lower-level turn hydration for custom wrappers |
| `receive_user_message(flow, text, ...)` | Same as instance method |
| `set_state_field(flow, name, value)` | Set a field on dict or Pydantic state |
| `get_conversational_config(flow)` | Read class `conversational_config` |
| `input_history_to_messages(entries)` | Convert `input_history` to LLM message format |

## Intent routing patterns

### A. Pre-classify via `ConversationConfig` (simplest)

Set `default_intents` and `intent_llm`. Each `handle_turn()` pre-classifies the current message. A non-empty result returned by a custom `route_turn()` takes precedence; otherwise `route_conversation` uses the current turn's classified intent.

### B. Classify inside `route_turn` (richer prompts)

Set `default_intents=None` so `handle_turn()` only appends the user message. In `route_turn()`, call `classify_intent` with a custom prompt or descriptions:

```python theme={null}
def route_turn(self, context):
    intent = self.classify_intent(
        self._routing_prompt(self.state.current_user_message),
        ("GREETING", "ORDER", "RESEARCH", "GOODBYE"),
        llm="gpt-4o-mini",
    )
    self.state.last_intent = intent
    return intent
```

Use **`@listen("RESEARCH")`** (or similar) for steps that run `Agent.kickoff()` with tools — not bare `LLM.call()` — when you need web research or multi-step tool use.

## When the flow finishes but the user keeps chatting

Each `handle_turn()` completes one graph run, and the conversation continues with another `handle_turn()` using the same `session_id`. With the default deferred trace lifecycle, that run emits `conversation_turn_completed`, while `FlowFinished` is emitted once when `finalize_session_traces()` closes the session. `@persist` restores `messages`, flags, and context.

**Persist pattern:** prefer `@persist` on a **single terminal step** (for example `finalize`) rather than on the whole `Flow` class. Class-level persist saves after every method; `load_state` uses the latest row, which may be a mid-run snapshot (for example right after `bootstrap`) and miss handler updates from the same turn.

Do **not** use `@human_feedback` for follow-up chat lines unless a human must approve a specific step output before it is shown.

## Conversational `Flow`

Opt into the conversational chat graph by setting `conversational = True` on a `Flow` subclass or applying `@ConversationConfig(...)`. The base `Flow` then supplies `route_conversation` as the built-in start/router plus the `converse_turn` and `end_conversation` listeners. The deprecated `answer_from_history_turn` listener remains available for compatibility. The framework manages `state.messages`, can drive a router LLM, and keeps the trace batch open across turns. You write the **custom routes**; the framework owns the rest.

Use this when you want a multi-turn chat with a router and per-route handlers without wiring the lifecycle yourself. Use `Flow[ChatState]` (the lower-level pattern above) when you need full control.

### Quick example

```python theme={null}
from crewai import Flow
from crewai.flow import listen
from crewai.flow import (
    ConversationConfig,
    ConversationState,
)


@ConversationConfig(defer_trace_finalization=True)
class SupportFlow(Flow[ConversationState]):
    def route_turn(self, context: dict) -> str | None:
        message = (self.state.current_user_message or "").lower()
        if "search" in message or "news" in message:
            return "INTERNET_SEARCH"
        if "docs" in message or "crewai" in message:
            return "CREWAI_DOCS"
        return "converse"

    @listen("INTERNET_SEARCH")
    def handle_internet_search(self) -> str:
        """Fresh web research, current news, real-time lookups."""
        reply = "I would run the web research route here."
        self.append_assistant_message(reply)
        return reply

    @listen("CREWAI_DOCS")
    def handle_crewai_docs(self) -> str:
        """Look up the CrewAI documentation for framework/API questions."""
        reply = "I would look up the CrewAI docs here."
        self.append_assistant_message(reply)
        return reply


flow = SupportFlow()
try:
    flow.handle_turn("What can you do?")              # routes to converse
    flow.handle_turn("Search the web for AI news.")   # routes to INTERNET_SEARCH
    flow.handle_turn("Check the CrewAI docs.")         # routes to CREWAI_DOCS
finally:
    flow.finalize_session_traces()
```

For a local terminal chat, use `chat()`:

```python theme={null}
def kickoff() -> None:
    SupportFlow().chat()
```

`chat()` wraps `handle_turn()` in a REPL, exits on `exit` / `quit`, skips blank lines by default, and calls `finalize_session_traces()` when the session ends.

### `ConversationConfig`

Class decorator that attaches per-class chat defaults.

| Field | Default | Purpose |
| - | - | - |
| `system_prompt` | `slices.conversational_system_prompt` from i18n | System message used by the built-in `converse_turn`. Pass `""` to opt out entirely. |
| `llm` | `None` | Conversation LLM (used by `converse_turn` and as router fallback). |
| `router` | `None` | Optional `RouterConfig` overrides. With custom listeners and a resolvable LLM, routing auto-enables even when this is omitted. |
| `answer_from_history_prompt` | Framework default | **Deprecated.** Use the `converse` system prompt or override `converse_turn()`. |
| `answer_from_history_llm` | `None` | **Deprecated.** Use `llm`; `converse` already receives canonical history. |
| `intent_llm` | `None` | LLM for legacy `intents=`/`default_intents` pre-classification. |
| `default_intents` | `None` | Outcome labels for legacy pre-classification. |
| `visible_agent_outputs` | `None` | `"all"`, or a list of agent names whose `append_agent_result()` calls should be promoted to public assistant messages. |
| `defer_trace_finalization` | `True` | Keep one trace batch open across `handle_turn()` calls. |

<Warning>
  `answer_from_history_prompt`, `answer_from_history_llm`, and the
  `answer_from_history` route are deprecated and will be removed in a future
  release. They duplicate `converse`, add an eligibility LLM call, and are
  bypassed when the normal auto-router returns a route. Existing configurations
  continue to work and emit `DeprecationWarning`.
</Warning>

With no custom routes, turns fall through to `converse`. With custom routes and a conversation/router LLM, the framework synthesizes a default `RouterConfig`; provide one explicitly only to customize its prompt, route list, descriptions, or fallback behavior. Setting `default_intents` uses the legacy pre-classification path instead.

If no conversation LLM is configured, the built-in `converse_turn` returns a configuration placeholder rather than generating an answer.

### `RouterConfig` and the auto-built route catalog

```python theme={null}
from typing import Literal

from pydantic import BaseModel

from crewai import LLM
from crewai.flow import RouterConfig


class MyRoute(BaseModel):
    intent: Literal["INTERNET_SEARCH", "CREWAI_DOCS", "converse"]


ROUTER_LLM = LLM(model="gpt-4o-mini")

router_config = RouterConfig(
    prompt="Optional domain framing (policy, voice, persona).",
    response_format=MyRoute,        # optional; auto-generated otherwise
    llm=ROUTER_LLM,                  # falls back to ConversationConfig.llm
    routes=["INTERNET_SEARCH", "CREWAI_DOCS"],   # optional; inferred from listeners
    route_descriptions={
        "INTERNET_SEARCH": "Override the docstring for this one route.",
    },
    default_intent="converse",       # used when LLM call fails or no LLM available
    fallback_intent="converse",      # used when LLM returns an invalid route
    intent_field="intent",
)
```

The router prompt that gets sent to the LLM is built automatically. For each route the framework picks a description with this precedence:

1. `RouterConfig.route_descriptions[label]` — explicit override.
2. `Flow.builtin_route_descriptions[label]` — framework-canned text for `converse`, `end`, and the deprecated `answer_from_history` compatibility route (phrased for the router LLM).
3. The method's declared `description` (used by declarative flows and DSL projections).
4. First non-empty line of the `@listen(label)` handler's docstring.
5. Empty (the route is listed without a description).

So in practice, **adding a new route is `@listen("X")` + a one-line docstring**:

```python theme={null}
from crewai.flow import listen


@listen("INTERNET_SEARCH")
def handle_internet_search(self) -> str:
    """Fresh web research, current news, real-time lookups."""
    ...
```

### Naming handlers

The string in `@listen("…")` is a **router route label** (an event name), not the Python method name. Route labels and method completion events share one trigger namespace, so naming a handler the same as its route causes the handler to re-trigger itself in a loop.

Use a different method name — the docs examples use a `handle_*` prefix:

```python theme={null}
@listen("create_video")
def handle_create_video(self) -> str:
    """User wants a new video."""
    ...
```

Do **not** mirror the route label on the method:

```python theme={null}
@listen("create_video")
def create_video(self) -> str:  # rejected at flow instantiation
    ...
```

…and the router LLM sees:

```
Routes:
- CREWAI_DOCS: Look up the CrewAI documentation for framework/API questions.
- INTERNET_SEARCH: Fresh web research, current news, real-time lookups.
- converse: Ordinary chat, follow-ups, summaries, clarifications…
- end: User signals the conversation is finished (goodbye, exit, done).
```

`RouterConfig.prompt` is for **domain framing** (assistant persona, business rules, voice). The route catalog is auto-built — don't list routes in `prompt`; they'll drift the moment you add a handler.

### Built-in routes

| Route | Handler | Purpose |
| - | - | - |
| `converse` | `converse_turn` | Default chat handler. Calls `ConversationConfig.llm` with the system prompt + canonical message history. |
| `end` | `end_conversation` | Sets `state.ended = True` and emits a terminator reply. |
| `answer_from_history` | `answer_from_history_turn` | **Deprecated compatibility route.** Use `converse`, which already receives canonical history. |

You can override any of these by defining a same-named handler in your subclass.

### `handle_turn()` semantics

`flow.handle_turn(message)` runs one turn:

1. Resets per-execution tracking (`_completed_methods`, `_method_outputs`) so the graph re-runs — without this, repeated `kickoff` calls on the same flow instance would short-circuit on turn 2+ because `Flow.kickoff_async` treats `inputs={"id": ...}` as a checkpoint restore.
2. Appends the user message to `state.messages`, sets `current_user_message` / `last_user_message`. `last_intent` is **preserved from the prior turn** so the router LLM can use it as a signal.
3. Runs user-defined `@start` methods (if any), then `route_conversation` as the built-in start/router, then the chosen `@listen` handler. `route_conversation` invokes the overridable `conversation_start()` helper.
4. The router stores its decision in `state.last_intent` (visible to the next turn's router context).
5. If your handler returned a string and didn't already call `append_assistant_message`, `handle_turn` appends it for you and persists the updated `state.messages` so `@persist` restore includes the assistant turn.

Call `handle_turn()` for chat messages. Calling `kickoff(inputs={"id": ...})` directly runs the flow graph without applying the conversational turn wrapper.

### `chat()` for local REPLs

`flow.chat()` is the batteries-included terminal wrapper around `handle_turn()`:

```python theme={null}
flow = SupportFlow()
flow.chat()
```

It handles the common local loop:

1. Prompts for a user message.
2. Stops on `exit` / `quit`, `EOFError`, or `KeyboardInterrupt`.
3. Calls `handle_turn(message, session_id=...)`.
4. Prints the assistant result.
5. Finalizes deferred session traces in a `finally` block.

`chat(defer_trace_finalization=True)` temporarily enables the instance deferral flag for the REPL and restores its prior value on exit.

Customize the terminal behavior with injectable I/O:

```python theme={null}
flow.chat(
    session_id="demo-session",
    prompt="You: ",
    assistant_prefix="Assistant: ",
    exit_commands=("exit", "quit", "bye"),
)
```

For web apps, background workers, tests, and custom transports, keep using `handle_turn()` directly.

### Custom router behavior

To run side effects (event bus setup, telemetry) on every routing decision, override `route_turn`:

```python theme={null}
from typing import Any

from crewai import Flow
from crewai.flow import ConversationState


class SupportFlow(Flow[ConversationState]):
    conversational = True

    def route_turn(self, context: dict[str, Any]) -> str | None:
        self.event_bus = MyBus(self)
        return super().route_turn(context)
```

To bypass the LLM router entirely and pick a route programmatically, return a non-empty string from `route_turn`. A falsy return does **not** invoke `_route_with_config()` from your override; routing falls through to this turn's pre-classified intent, then the deprecated `answer_from_history` compatibility path when configured, and finally `converse`. A previous turn's `last_intent` is available in router context but is never replayed as a fallback.

### `append_assistant_message` and `append_agent_result`

Inside a `@listen(label)` handler, choose:

* `self.append_assistant_message(text)` — adds a user-visible assistant turn to `state.messages`. The next turn's `converse_turn` sees it.
* `self.append_agent_result(agent_name, result, visibility="private")` — records a structured event in `state.events` and a thread in `state.agent_threads[agent_name]`. Public visibility also calls `append_assistant_message` for you. Use private results for scratch work that shouldn't pollute the canonical history.

`ConversationConfig.visible_agent_outputs` can promote specific agents' private results to public globally (`"all"`, or a list of agent names).

## Declaring a conversational flow in JSON/YAML

A [declarative Flow](/edge/en/concepts/cli) can be conversational too. Add a top-level `conversational` block and declare your own routes as methods that `listen` to a route label:

```yaml theme={null}
schema: crewai.flow/v1
name: SupportFlow

conversational:
  system_prompt: You are a terse support assistant.
  llm: gpt-4o-mini
  router:
    llm: gpt-4o-mini

methods:
  handle_order:
    description: Order status, shipping and delivery questions.
    listen: order
    do:
      call: agent
      with:
        role: Support specialist
        goal: Answer order questions accurately
        backstory: Knows the fulfilment pipeline.
        input: "${state.current_user_message}"
```

Declaring the block is the opt-in — `enabled` defaults to `true`. Set `enabled: false` to keep the configuration while turning chat off. This also disables built-in method synthesis, so the declaration must provide a normal non-conversational graph.

Three things are supplied for you:

| Supplied | Detail |
| - | - |
| The built-in graph | `route_conversation`, `converse_turn`, and `end_conversation` are added automatically. Deprecated `answer_from_history_turn` is retained for compatibility. Declare a method under one of those names to override it. |
| Conversation state | `ConversationState` is used when there is no `state` block. A Pydantic `ref` or `json_schema` state is automatically composed with the conversational fields; it does not need to extend `ConversationState`. |
| The route catalog | Inferred from non-router methods with `listen` labels, excluding internal routes. Descriptions follow the precedence above, and explicit `router.routes` can limit the choices. |

Declarative `llm`, `router.llm`, and `intent_llm` fields accept either a model id or a configuration mapping such as `{model: openai/gpt-4o-mini, max_tokens: 512}`. The `conversational` block also supports `default_intents`, `visible_agent_outputs`, `defer_trace_finalization`, and the `RouterConfig` fields shown above. Deprecated `answer_from_history_prompt` / `answer_from_history_llm` declarations remain accepted for compatibility.

Run it from Python with the same turn APIs as a class-based conversational Flow:

```python theme={null}
from crewai.flow import Flow

flow = Flow.from_declaration(path="flow.yaml")

try:
    flow.handle_turn("Where is my order?", session_id="session-1")
finally:
    flow.finalize_session_traces()
```

### Naming routes

Route labels and method names share one trigger namespace, so a handler must not be named after the route it listens to — `create_video` listening to `create_video` is rejected when the flow is built. Use a `handle_*` prefix.

### What a declaration cannot express

| Not expressible | Use instead |
| - | - |
| A live `LLM` instance or a custom `BaseLLM` | A model id string or static configuration mapping |
| `router.response_format` as a live model class | Name the class with a python ref: `response_format: {python: my_project.schemas.ConversationRoute}`. Omit it and the framework synthesizes one |
| A `route_turn()` override | Author the Flow in Python, or replace the declarative `route_conversation` method with a `call: code` / expression action |
| A `can_answer_from_history()` override | Deprecated. Use `converse` or override `converse_turn()` in Python. |

`crewai run` opens the chat TUI for a declarative conversational flow — the same one a Python conversational Flow gets. A chat loop needs a terminal, so a headless run exits non-zero with guidance instead of running a single turn; drive it from Python there with `handle_turn()` or `stream_turn()`. A declarative method with a `human_feedback:` block (Python: `@human_feedback`) runs on a terminal REPL, because the runtime collects feedback with a blocking prompt the TUI cannot service. `--inputs` is not accepted for a conversational flow — each turn's input is the message you type — and resuming a session by id is not wired into the CLI yet; use `flow.handle_turn(message, session_id=...)` from Python for that.

## Tracing across turns

With `defer_trace_finalization=True` (default in `ConversationConfig`):

* **One trace batch** for the whole chat session.
* **`flow_started`** on the first turn only; **`flow_finished`** once in `finalize_session_traces()`.
* **Per-turn** `kickoff` does not print “Trace batch finalized”.
* **Nested work** (`Agent.kickoff()`, crews, Exa tools) appends to the **parent** batch; inner `AgentExecutor` flows do not close the session batch early.

```python theme={null}
flow.chat(session_id=session_id)
```

`flow.chat()` calls `finalize_session_traces()` for you. When you own the loop
with `handle_turn()`, call `finalize_session_traces()` when
the session ends.

`suppress_flow_events=True` hides Rich console panels and suppresses method execution events. Flow start/finish events still emit, so the outer Flow lifecycle remains traceable, but individual method spans are omitted.

### Conversational `Flow` trace lifecycle

The [conversational `Flow`](#conversational-flow) uses the same tracing lifecycle: `defer_trace_finalization` defaults to `True`, so each `handle_turn()` keeps the session trace open. Deferred turns also suppress per-turn `flow_failed`; on a turn error or session abort, finalize the session explicitly. This closes the batch with the session-level `FlowFinished` event rather than a per-turn `FlowFailed` event. Always wrap your REPL/loop in `try/finally` and call `flow.finalize_session_traces()` on exit. Without it, the trace batch stays open and the final conversation may never export.

## Streaming

For conversational UIs, use `stream_turn()` and iterate its ordered `StreamFrame` objects:

```python theme={null}
stream = flow.stream_turn("Where is my order?", session_id=session_id)

with stream:
    for frame in stream.events:
        if frame.channel == "llm" and frame.type == "llm_stream_chunk":
            print(frame.content, end="", flush=True)

reply = stream.result
```

For a non-conversational Flow, setting `stream = True` makes `kickoff()` return a `StreamSession`. Do not set `flow.stream = True` when using `handle_turn()`; `stream_turn()` owns the conversational streaming lifecycle.

## Imports

```python theme={null}
from crewai.flow import (
    ChatState,
    ConversationalConfig,
    ConversationalInputs,
    Flow,
    listen,
    persist,
    router,
    start,
)
from crewai.flow.conversation import prepare_conversational_turn
from crewai.flow import (
    ConversationConfig,
    ConversationState,
    RouterConfig,
)
```

## See also

* [Mastering Flow State Management](/en/guides/flows/mastering-flow-state) — persistence, Pydantic state, `@persist`
* [Build Your First Flow](/en/guides/flows/first-flow) — flow basics

## Experimental jobs spanning conversational turns

This API is experimental and may change. Import it explicitly from `crewai.experimental.flow_jobs`; it is not part of the stable `crewai.flow` surface.

`JobRecord`, `JobState`, `JobUpdate`, `JobWorkState`, `JobWorkFlow`, and `JobRunner` provide an opt-in, in-memory job contract. Existing conversational Flows keep their behavior. A job has its own identity, originating turn, revision, execution attempt, status, current stage, completed stages, and error. Applications add typed inputs and outputs; no separate artifact class is required.

Declare writable result fields in `output_fields`. Worker updates cannot change input, ownership, or lifecycle fields through an output payload. The complete candidate record is validated before a commit; malformed, foreign, stale, duplicate, and post-terminal updates leave accepted state unchanged.

```python theme={null}
import asyncio
from typing import ClassVar
from pydantic import Field
from crewai.flow import ConversationState, start
from crewai.experimental.flow_jobs import (
    JobRecord, JobRunner, JobState, JobWorkFlow, JobWorkState, add_job,
)

class ReportJob(JobRecord):
    output_fields: ClassVar[frozenset[str]] = frozenset({"notes"})
    question: str
    notes: list[str] = Field(default_factory=list)
    stage: str = "collect"

class WorkChatState(ConversationState, JobState[ReportJob]):
    pass

class ReportWork(JobWorkFlow[JobWorkState[ReportJob]]):
    @start()
    def collect(self):
        self.check_open()
        self.publish_update("started")
        self.publish_update(
            "stage_completed", stage="collect", outputs={"notes": ["Example fact"]},
        )
        self.publish_update("completed")

async def run_job():
    state = WorkChatState()
    changes = asyncio.Queue()

    def make_work(job, inputs, publish):
        return ReportWork(
            initial_state=JobWorkState[ReportJob](job=job),
            publish=publish, suppress_flow_events=True,
        )

    runner = JobRunner(state, make_work, on_update=changes.put)
    try:
        async with runner.state_lock:
            job = ReportJob(session_id=state.id, question="Example report")
            if not add_job(state, job):
                raise RuntimeError("Job registration rejected")
            runner.submit(job, {})
        while True:
            snapshot = await changes.get()
            if snapshot["jobs"][0]["status"] in {"completed", "failed"}:
                return snapshot
    finally:
        await runner.aclose()

# Run in a script: asyncio.run(run_job())

```

Construct `JobRunner(state, make_work, on_update=publish_snapshot)` on a running asyncio loop. `make_work(job, inputs, publish)` returns a non-streaming `JobWorkFlow` with its own typed state, copied job/inputs, and `publish=publish`. `on_update` is an async callback receiving JSON-compatible snapshots (`session_id`, `seq`, `jobs`); adapt them to your existing transport. `format_error` can redact provider details, and `on_worker_event` observes worker start/settlement.

Hold `runner.state_lock` for **every** foreground turn, job admission, and parent-state mutation. Register with `add_job(state, job)` and then call `runner.submit(job, inputs)`. Run synchronous turns in a worker thread while holding that lock; never run two turns concurrently on one Flow. Worker `publish_update()` also runs off the runner's event loop and waits for an accepted commit before the next stage. A `stage_started` update requires the previous stage to be committed; `stage_completed` retains typed outputs before successful completion. A later failure preserves earlier outputs.

Project relevant job status and outputs through `build_router_context()` and `build_agent_context()`. Snapshots are not automatically public chat messages or spoken replies. Domain interpretation and admission capacity remain application policy.

At session teardown, call `runner.request_close()` on its owning loop before draining foreground execution, then `await runner.aclose()`. This closes admission/publication, signals each worker, releases waiting publishers, and settles owned tasks. Work methods must call `check_open()` at stage boundaries. An in-flight provider operation is not aborted. Subscriber failure closes the runner rather than stranding publishers. Stopping speech alone should not close the runner.

The initial lifecycle is `queued → running → completed/failed`. Pause/resume, replacement, automatic reply delivery, and distributed recovery are not supplied. Records can be serialized with existing state facilities, but restoring them does not recreate live workers or guarantee exactly-once external effects.

## Experimental turn and public reply identities

Import `TurnRecord`, `ReplyRecord`, `TurnState`, `ReplyEvent`, and `TurnRunner` explicitly from `crewai.experimental.flow_turns`. This opt-in API may change. It wraps an existing conversational Flow; the stable `handle_turn()`/`stream_turn()` APIs, return values, and experimental job contract are unchanged.

`TurnState` extends `ConversationState` with serializable `turns` and `replies`. One invocation of `TurnRunner.stream_turn()` tracks one committed input and its primary public reply. Every accepted input has a session-scoped, single-use `turn_id`; a new utterance or correction gets a new turn ID. `input_revision` is a positive integer (default `1`), not automatic provisional transcript replacement. `delivery_id` is the sole reply identity; there is no second `reply_id`. Reply kinds are `acknowledgment`, `answer`, and `progress`.

```python theme={null}
from crewai.flow import Flow, listen
from crewai.experimental.flow_turns import TurnRunner, TurnState

class Chat(Flow[TurnState]):
    conversational = True

    def route_turn(self, context):
        return "acknowledge"

    @listen("acknowledge")
    def handle_acknowledge(self):
        return "Request received."

flow = Chat(suppress_flow_events=True)
runner = TurnRunner(flow)
events = runner.stream_turn("Hello", turn_id="turn-1", kind="acknowledgment")
try:
    for event in events:
        if runner.accept_event(event):
            print(event.model_dump_json())
finally:
    events.close()

snapshot = runner.snapshot()
reply = snapshot["replies"][0]
runner.interrupt(
    session_id=reply["session_id"], turn_id=reply["turn_id"],
    input_revision=reply["input_revision"], delivery_id=reply["delivery_id"],
)
```

Hardcoded replies and final-only providers emit the same public `text`/`completed` contract as streaming responders. The existing Flow remains responsible for canonical history. `seq` orders events within a delivery; `segment_id` numbers text deltas starting at one. A text delta is not necessarily a complete spoken sentence. Route and generation observation events include the same identity. `completed` means text generation finished, **not** that audio was played or heard.

Only dedicated public responder methods (`converse_turn` and `answer_from_history_turn` by default), or explicitly selected public agent IDs, allow live model text. Configure `public_methods` and the `public_agent_ids` callback deliberately; do not designate a method that mixes private model work with public response generation. Router streams, reasoning, tool-call chunks, tool-enabled calls, and unrelated agents are excluded. Private agent results do not become replies. Providers without observable deltas still publish their canonical final message; no model timing is invented.

`runner.accept_event(event)` claims one runner-issued event before transport handoff. It rejects foreign identity, duplicate/out-of-order sequence, and output invalidated by interruption or failure. Invoke it after any application queue/delay, immediately before handoff. It is not a validator for untrusted client playback receipts. Already handed-off output needs adapter-side identity fencing and audio interruption.

`runner.interrupt(session_id=..., turn_id=..., input_revision=..., delivery_id=...)` targets exactly one reply and is idempotent for valid repeated requests. During generation it marks the turn/reply interrupted; after generation it preserves completed text and marks the reply interrupted without changing job success. `on_interrupt` may signal existing cooperative stage checks; `should_interrupt` can observe an application's existing stop signal. These callbacks do not abort arbitrary provider calls. The stream must settle before another tracked turn can execute on the same Flow. Closing the iterator signals interruption and drains native execution; it can wait for an in-flight provider.

To compose with background work, declare `class WorkChatState(TurnState, JobState[MyJob]): ...`. Hold `JobRunner.state_lock` around the **entire** foreground iteration and all parent job mutations. TurnRunner separately protects its records and interruption controls. Do not run native turn APIs concurrently with the wrapper. After successfully admitting a job, `runner.associate_job(job)` links its session/originating turn and revision/attempt to the active reply; it never admits, submits, or cancels work. A new conversational turn does not invalidate independent research.

`elapsed_ms` uses a server monotonic clock from tracked-turn startup. `model_first_text_ms`, when available, is the difference between provider request-start and first public text **event timestamps**; it includes provider/SDK/network work and is not isolated inference-engine TTFT. Never subtract browser and server timestamps. TTS first chunk, cached audio, physical audible onset, and playback completion remain adapter measurements.

This increment does not supply background reply queues, speaking-floor scheduling, playback receipts, exact word alignment, speculative generation, or job pause/resume. Snapshots contain records, not locks, cancellation signals or running executions. They do not resume work or replay delivery. Restore state before creating a new tracked live turn; `from_checkpoint` and `restore_from_state_id` are deliberately rejected by this wrapper. Durable reconciliation and persistence of terminal tracking records remain separate work. Continue using the stable Flow APIs for existing checkpoint workflows.

During tracked execution, the current restored state remains authoritative: automatic session reload is suppressed so it cannot discard admitted turn/reply records or accepted background job updates. Native Flow turn APIs retain their existing persistence reload behavior. Closing after a terminal event does not interrupt the completed reply. Failure remains an acceptable terminal lifecycle signal if interruption arrives after failure.

## Experimental queued background replies

`crewai.experimental.flow_replies` adds `ReplyQueueState`, `ReplyQueue`, `DeliveryRecord`, `DeliveryFloor`, `CoveredUpdate`, and `ClientActivity` on top of the turn/reply contract. Declare state that combines `ReplyQueueState` and `JobState[MyJob]`, then create **one** `TurnRunner` and **one** `ReplyQueue` for that live Flow. Existing Flows remain opt-in; the job executor and stable conversational APIs do not change.

A successful job is immediately available in state, even when its announcement waits. `enqueue(job)` accepts only a current, committed `completed` job with an existing originating turn. The pending delivery stores that job's revision, attempt and accepted update sequence; duplicate notifications coalesce to the same `delivery_id`. This ID is also the eventual public `ReplyRecord` identity. Each job completion gets one attributable announcement; intermediate stage progress does not create a speaking queue.

```python theme={null}
from crewai.experimental.flow_jobs import JobState
from crewai.experimental.flow_turns import TurnRunner
from crewai.experimental.flow_replies import (
    ClientActivity, CoveredUpdate, ReplyQueue, ReplyQueueState,
)

class WorkChatState(ReplyQueueState, JobState[MyJob]):
    pass

# Your conversational Flow uses WorkChatState; jobs is its JobRunner.
turns = TurnRunner(flow)
replies = ReplyQueue(turns, is_relevant=lambda job: job.job_id not in cancelled_job_ids)

async def on_job_snapshot(snapshot):
    async with jobs.state_lock:
        for job in flow.state.jobs.values():
            if job.status == "completed":
                replies.enqueue(job)
    # Publish snapshots and wake your application-owned delivery task.

async def try_delivery():
    async with jobs.state_lock:
        delivery = replies.claim()
        if delivery is None:
            return
        events = replies.prepare(
            delivery.delivery_id,
            lambda current_jobs: "\n\n".join(job.answer for job in current_jobs),
        )
    for event in events:
        async with jobs.state_lock:
            if not replies.accepts_output(delivery.delivery_id):
                replies.settle(delivery.delivery_id, "skipped")
                return
            if turns.accept_event(event):
                await send_public_event(event)
    # Generation has completed. The speaking floor is still leased.
    # Validate whole-message adapter feedback before calling:
    # replies.settle(delivery.delivery_id, "completed")
```

Serialize enqueue, claim, preparation, coverage changes, and output handoff with the same `JobRunner.state_lock` used by foreground turns and job commits. `ReplyQueue` also protects its own controls with a lock so two claims cannot acquire the floor. `prepare()` reads copied **current accepted artifacts**, rather than cached text captured at enqueue, and emits the shared `started`/`text`/`completed` public reply events with origin and job identity. It does not execute an LLM, append history, or schedule TTS. The application chooses intentional public text, history insertion, delivery tasks and transport.

`set_foreground(True)` holds the gate through foreground generation and draining, including pending user requests; release it when that work has settled. `observe_client(ClientActivity(...))` accepts only the owning session and increasing client sequence, with strict boolean `recording`, `playback`, and `muted` observations. Identified playback must belong to the Flow; a stale stop for a different reply cannot clear newer playback. Null `delivery_id` accommodates a local cached acknowledgment. Recording, playback, mute, a running turn, session closure or an already active delivery blocks `claim()`. If recording/mute or a new user request arrives during an announcement, the adapter must interrupt and stop that active delivery, then update activity and wake scheduling.

`claim()` atomically changes one eligible delivery from `pending` to `scheduled`. Before preparation and **each** handoff, recheck relevance with `accepts_output(delivery_id)` and claim text events with `turns.accept_event(event)`. Identity/version changes, missing jobs, non-successful work, ended sessions, already-covered updates and application `is_relevant(job)` policy invalidate announcements. The callback must be side-effect free; use it for application cancellation/supersession beyond the existing job contract. Wake the delivery task when activity or relevance changes. Relevance skips never remove retained artifacts.

After a requested summary has actually been handed off, `mark_covered(CoveredUpdate.from_job(job))` records the accepted version and skips matching pending announcements. Merely requesting a summary is not delivery. A later job version remains eligible. Do not mark all context jobs covered by a status response or unrelated conversation.

`settle(delivery_id, status, reason=...)` releases only the active delivery, once. Its terminal status is `completed`, `interrupted`, `skipped`, or `failed`; generation completion alone does **not** settle delivery. Validate feedback session, turn, input revision and delivery identity at the adapter boundary, and accept completion only after output generation has settled. Whole-message adapter completion is a scheduling observation, not proof of audible words. On interruption, failure or invalidation, stop local playback and in-flight adapter delivery, fence remaining public output, and leave the result available. Interrupted/failed announcements are not automatically replayed; users can request a fresh summary. TTS failure does not change job success.

Call `close()` before cancelling and draining your application-owned delivery tasks. It fences pending/active deliveries while retaining job outputs. Snapshot data includes floor observations and delivery records, not a timer, media connection or resumable lease. Restoration does not restart delivery; reconcile active records through application recovery policy. Segment-level receipts, exact spoken context, automatic job cancellation/replacement, distributed scheduling and playback timeouts remain separate features. A client that never reports completion conservatively retains the floor until interrupted or closed; applications can add an explicit timeout policy.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.