Skip to main content

Overview

Conversational apps treat each user line as a new flow run with the same session id. CrewAI adds helpers for message history, optional intent routing, deferred tracing, structured turn streaming, and a local flow.chat() REPL.

Turn APIs

Use flow.handle_turn(message, session_id=...) for every user message from REST, WebSocket, tests, and custom UIs. Use flow.chat() when you want a local terminal chat loop for a conversational Flow. Flow.kickoff() does not accept user_message= or session_id= keyword arguments. For conversational flows, handle_turn() stores the pending message and calls kickoff(inputs={"id": session_id}) internally after resetting per-turn execution state. handle_turn(), stream_turn(), and chat() raise ValueError unless conversational mode is enabled. Applying @ConversationConfig(...) enables it automatically; otherwise set conversational = True.

Quick start

Streaming a turn

Use stream_turn() when a UI or runtime needs structured events for one chat turn. It returns a stream session with ordered frames for Flow routing, LLM chunks, tool activity, and conversation messages.
For the full frame contract and channel list, see Streaming Runtime Contract.

Turn lifecycle

Each handle_turn runs this pipeline:
  1. Turn setup — stores the pending user message, resolves the session id, resets per-turn execution tracking, and calls kickoff(inputs={"id": session_id}).
  2. State restore — if inputs["id"] exists and @persist is configured, loads the latest snapshot.
  3. FlowStarted — emitted on the first deferred session turn only.
  4. Pending turn hydration — appends the user message to state.messages, sets current_user_message / last_user_message, and optionally classifies when intents / default_intents + intent_llm are set.
  5. Graph execution — user-defined @start methods (if any) → route_conversation (the built-in start/router) → the selected @listen handler. route_conversation also calls the overridable conversation_start() helper.
  6. End of run — per-turn flow_finished and trace finalization are skipped when deferral is enabled; nested Agent.kickoff() / crews do not close the parent batch either.
Handlers should call append_assistant_message(reply) when the visible reply is not the return value, or when you trim history. A public string return is also recorded as assistant and included in the @persist snapshot, so a fresh Flow instance restores it. The user line is already stored by handle_turn — do not append it again in handlers.

Configuration overview

Decorating a Flow subclass with ConversationConfig both attaches the chat defaults and enables conversational mode. See the full field reference below. Override pre-classification per turn with handle_turn(..., intents=..., intent_llm=...).

Lower-level ChatState helpers

ChatState, the legacy ConversationalConfig, and crewai.flow.conversation helpers are still importable for advanced orchestration, tests, or custom wrappers. They are separate from the ConversationState / ConversationConfig API and do not add user_message= or session_id= keyword arguments to Flow.kickoff().
ConversationalInputs is a TypedDict for conventional kickoff(inputs={...}) keys: id, user_message, last_intent. ConversationState stores messages as ConversationMessage objects and additionally provides current_user_message, ended, events, and agent_threads. Use conversation_messages when passing its canonical history to an LLM.

Flow conversational API

handle_turn parameters

kickoff parameters

Flow.kickoff() accepts inputs, input_files, from_checkpoint, and restore_from_state_id. Pass inputs={"id": session_id} when you need raw flow execution, but use handle_turn() when the call represents a chat message.

Instance attributes

Methods and properties

Module helpers (crewai.flow.conversation)

Importable from crewai.flow.conversation for tests or custom orchestration. These helpers use the legacy ConversationalConfig shape; prepare_conversational_turn() also clears last_intent, unlike handle_turn(), which preserves it as router context.

Intent routing patterns

A. Pre-classify via ConversationConfig (simplest)

Set default_intents and intent_llm. Each handle_turn() pre-classifies the current message. A non-empty result returned by a custom route_turn() takes precedence; otherwise route_conversation uses the current turn’s classified intent.

B. Classify inside route_turn (richer prompts)

Set default_intents=None so handle_turn() only appends the user message. In route_turn(), call classify_intent with a custom prompt or descriptions:
Use @listen("RESEARCH") (or similar) for steps that run Agent.kickoff() with tools — not bare LLM.call() — when you need web research or multi-step tool use.

When the flow finishes but the user keeps chatting

Each handle_turn() completes one graph run, and the conversation continues with another handle_turn() using the same session_id. With the default deferred trace lifecycle, that run emits conversation_turn_completed, while FlowFinished is emitted once when finalize_session_traces() closes the session. @persist restores messages, flags, and context. Persist pattern: prefer @persist on a single terminal step (for example finalize) rather than on the whole Flow class. Class-level persist saves after every method; load_state uses the latest row, which may be a mid-run snapshot (for example right after bootstrap) and miss handler updates from the same turn. Do not use @human_feedback for follow-up chat lines unless a human must approve a specific step output before it is shown.

Conversational Flow

Opt into the conversational chat graph by setting conversational = True on a Flow subclass or applying @ConversationConfig(...). The base Flow then supplies route_conversation as the built-in start/router plus the converse_turn and end_conversation listeners. The deprecated answer_from_history_turn listener remains available for compatibility. The framework manages state.messages, can drive a router LLM, and keeps the trace batch open across turns. You write the custom routes; the framework owns the rest. Use this when you want a multi-turn chat with a router and per-route handlers without wiring the lifecycle yourself. Use Flow[ChatState] (the lower-level pattern above) when you need full control.

Quick example

For a local terminal chat, use chat():
chat() wraps handle_turn() in a REPL, exits on exit / quit, skips blank lines by default, and calls finalize_session_traces() when the session ends.

ConversationConfig

Class decorator that attaches per-class chat defaults.
answer_from_history_prompt, answer_from_history_llm, and the answer_from_history route are deprecated and will be removed in a future release. They duplicate converse, add an eligibility LLM call, and are bypassed when the normal auto-router returns a route. Existing configurations continue to work and emit DeprecationWarning.
With no custom routes, turns fall through to converse. With custom routes and a conversation/router LLM, the framework synthesizes a default RouterConfig; provide one explicitly only to customize its prompt, route list, descriptions, or fallback behavior. Setting default_intents uses the legacy pre-classification path instead. If no conversation LLM is configured, the built-in converse_turn returns a configuration placeholder rather than generating an answer.

RouterConfig and the auto-built route catalog

The router prompt that gets sent to the LLM is built automatically. For each route the framework picks a description with this precedence:
  1. RouterConfig.route_descriptions[label] — explicit override.
  2. Flow.builtin_route_descriptions[label] — framework-canned text for converse, end, and the deprecated answer_from_history compatibility route (phrased for the router LLM).
  3. The method’s declared description (used by declarative flows and DSL projections).
  4. First non-empty line of the @listen(label) handler’s docstring.
  5. Empty (the route is listed without a description).
So in practice, adding a new route is @listen("X") + a one-line docstring:

Naming handlers

The string in @listen("…") is a router route label (an event name), not the Python method name. Route labels and method completion events share one trigger namespace, so naming a handler the same as its route causes the handler to re-trigger itself in a loop. Use a different method name — the docs examples use a handle_* prefix:
Do not mirror the route label on the method:
…and the router LLM sees:
RouterConfig.prompt is for domain framing (assistant persona, business rules, voice). The route catalog is auto-built — don’t list routes in prompt; they’ll drift the moment you add a handler.

Built-in routes

You can override any of these by defining a same-named handler in your subclass.

handle_turn() semantics

flow.handle_turn(message) runs one turn:
  1. Resets per-execution tracking (_completed_methods, _method_outputs) so the graph re-runs — without this, repeated kickoff calls on the same flow instance would short-circuit on turn 2+ because Flow.kickoff_async treats inputs={"id": ...} as a checkpoint restore.
  2. Appends the user message to state.messages, sets current_user_message / last_user_message. last_intent is preserved from the prior turn so the router LLM can use it as a signal.
  3. Runs user-defined @start methods (if any), then route_conversation as the built-in start/router, then the chosen @listen handler. route_conversation invokes the overridable conversation_start() helper.
  4. The router stores its decision in state.last_intent (visible to the next turn’s router context).
  5. If your handler returned a string and didn’t already call append_assistant_message, handle_turn appends it for you and persists the updated state.messages so @persist restore includes the assistant turn.
Call handle_turn() for chat messages. Calling kickoff(inputs={"id": ...}) directly runs the flow graph without applying the conversational turn wrapper.

chat() for local REPLs

flow.chat() is the batteries-included terminal wrapper around handle_turn():
It handles the common local loop:
  1. Prompts for a user message.
  2. Stops on exit / quit, EOFError, or KeyboardInterrupt.
  3. Calls handle_turn(message, session_id=...).
  4. Prints the assistant result.
  5. Finalizes deferred session traces in a finally block.
chat(defer_trace_finalization=True) temporarily enables the instance deferral flag for the REPL and restores its prior value on exit. Customize the terminal behavior with injectable I/O:
For web apps, background workers, tests, and custom transports, keep using handle_turn() directly.

Custom router behavior

To run side effects (event bus setup, telemetry) on every routing decision, override route_turn:
To bypass the LLM router entirely and pick a route programmatically, return a non-empty string from route_turn. A falsy return does not invoke _route_with_config() from your override; routing falls through to this turn’s pre-classified intent, then the deprecated answer_from_history compatibility path when configured, and finally converse. A previous turn’s last_intent is available in router context but is never replayed as a fallback.

append_assistant_message and append_agent_result

Inside a @listen(label) handler, choose:
  • self.append_assistant_message(text) — adds a user-visible assistant turn to state.messages. The next turn’s converse_turn sees it.
  • self.append_agent_result(agent_name, result, visibility="private") — records a structured event in state.events and a thread in state.agent_threads[agent_name]. Public visibility also calls append_assistant_message for you. Use private results for scratch work that shouldn’t pollute the canonical history.
ConversationConfig.visible_agent_outputs can promote specific agents’ private results to public globally ("all", or a list of agent names).

Declaring a conversational flow in JSON/YAML

A declarative Flow can be conversational too. Add a top-level conversational block and declare your own routes as methods that listen to a route label:
Declaring the block is the opt-in — enabled defaults to true. Set enabled: false to keep the configuration while turning chat off. This also disables built-in method synthesis, so the declaration must provide a normal non-conversational graph. Three things are supplied for you: Declarative llm, router.llm, and intent_llm fields accept either a model id or a configuration mapping such as {model: openai/gpt-4o-mini, max_tokens: 512}. The conversational block also supports default_intents, visible_agent_outputs, defer_trace_finalization, and the RouterConfig fields shown above. Deprecated answer_from_history_prompt / answer_from_history_llm declarations remain accepted for compatibility. Run it from Python with the same turn APIs as a class-based conversational Flow:

Naming routes

Route labels and method names share one trigger namespace, so a handler must not be named after the route it listens to — create_video listening to create_video is rejected when the flow is built. Use a handle_* prefix.

What a declaration cannot express

crewai run opens the chat TUI for a declarative conversational flow — the same one a Python conversational Flow gets. A chat loop needs a terminal, so a headless run exits non-zero with guidance instead of running a single turn; drive it from Python there with handle_turn() or stream_turn(). A declarative method with a human_feedback: block (Python: @human_feedback) runs on a terminal REPL, because the runtime collects feedback with a blocking prompt the TUI cannot service. --inputs is not accepted for a conversational flow — each turn’s input is the message you type — and resuming a session by id is not wired into the CLI yet; use flow.handle_turn(message, session_id=...) from Python for that.

Tracing across turns

With defer_trace_finalization=True (default in ConversationConfig):
  • One trace batch for the whole chat session.
  • flow_started on the first turn only; flow_finished once in finalize_session_traces().
  • Per-turn kickoff does not print “Trace batch finalized”.
  • Nested work (Agent.kickoff(), crews, Exa tools) appends to the parent batch; inner AgentExecutor flows do not close the session batch early.
flow.chat() calls finalize_session_traces() for you. When you own the loop with handle_turn(), call finalize_session_traces() when the session ends. suppress_flow_events=True hides Rich console panels and suppresses method execution events. Flow start/finish events still emit, so the outer Flow lifecycle remains traceable, but individual method spans are omitted.

Conversational Flow trace lifecycle

The conversational Flow uses the same tracing lifecycle: defer_trace_finalization defaults to True, so each handle_turn() keeps the session trace open. Deferred turns also suppress per-turn flow_failed; on a turn error or session abort, finalize the session explicitly. This closes the batch with the session-level FlowFinished event rather than a per-turn FlowFailed event. Always wrap your REPL/loop in try/finally and call flow.finalize_session_traces() on exit. Without it, the trace batch stays open and the final conversation may never export.

Streaming

For conversational UIs, use stream_turn() and iterate its ordered StreamFrame objects:
For a non-conversational Flow, setting stream = True makes kickoff() return a StreamSession. Do not set flow.stream = True when using handle_turn(); stream_turn() owns the conversational streaming lifecycle.

Imports

See also

Experimental jobs spanning conversational turns

This API is experimental and may change. Import it explicitly from crewai.experimental.flow_jobs; it is not part of the stable crewai.flow surface. JobRecord, JobState, JobUpdate, JobWorkState, JobWorkFlow, and JobRunner provide an opt-in, in-memory job contract. Existing conversational Flows keep their behavior. A job has its own identity, originating turn, revision, execution attempt, status, current stage, completed stages, and error. Applications add typed inputs and outputs; no separate artifact class is required. Declare writable result fields in output_fields. Worker updates cannot change input, ownership, or lifecycle fields through an output payload. The complete candidate record is validated before a commit; malformed, foreign, stale, duplicate, and post-terminal updates leave accepted state unchanged.
Construct JobRunner(state, make_work, on_update=publish_snapshot) on a running asyncio loop. make_work(job, inputs, publish) returns a non-streaming JobWorkFlow with its own typed state, copied job/inputs, and publish=publish. on_update is an async callback receiving JSON-compatible snapshots (session_id, seq, jobs); adapt them to your existing transport. format_error can redact provider details, and on_worker_event observes worker start/settlement. Hold runner.state_lock for every foreground turn, job admission, and parent-state mutation. Register with add_job(state, job) and then call runner.submit(job, inputs). Run synchronous turns in a worker thread while holding that lock; never run two turns concurrently on one Flow. Worker publish_update() also runs off the runner’s event loop and waits for an accepted commit before the next stage. A stage_started update requires the previous stage to be committed; stage_completed retains typed outputs before successful completion. A later failure preserves earlier outputs. Project relevant job status and outputs through build_router_context() and build_agent_context(). Snapshots are not automatically public chat messages or spoken replies. Domain interpretation and admission capacity remain application policy. At session teardown, call runner.request_close() on its owning loop before draining foreground execution, then await runner.aclose(). This closes admission/publication, signals each worker, releases waiting publishers, and settles owned tasks. Work methods must call check_open() at stage boundaries. An in-flight provider operation is not aborted. Subscriber failure closes the runner rather than stranding publishers. Stopping speech alone should not close the runner. The initial lifecycle is queued → running → completed/failed. Pause/resume, replacement, automatic reply delivery, and distributed recovery are not supplied. Records can be serialized with existing state facilities, but restoring them does not recreate live workers or guarantee exactly-once external effects.

Experimental turn and public reply identities

Import TurnRecord, ReplyRecord, TurnState, ReplyEvent, and TurnRunner explicitly from crewai.experimental.flow_turns. This opt-in API may change. It wraps an existing conversational Flow; the stable handle_turn()/stream_turn() APIs, return values, and experimental job contract are unchanged. TurnState extends ConversationState with serializable turns and replies. One invocation of TurnRunner.stream_turn() tracks one committed input and its primary public reply. Every accepted input has a session-scoped, single-use turn_id; a new utterance or correction gets a new turn ID. input_revision is a positive integer (default 1), not automatic provisional transcript replacement. delivery_id is the sole reply identity; there is no second reply_id. Reply kinds are acknowledgment, answer, and progress.
Hardcoded replies and final-only providers emit the same public text/completed contract as streaming responders. The existing Flow remains responsible for canonical history. seq orders events within a delivery; segment_id numbers text deltas starting at one. A text delta is not necessarily a complete spoken sentence. Route and generation observation events include the same identity. completed means text generation finished, not that audio was played or heard. Only dedicated public responder methods (converse_turn and answer_from_history_turn by default), or explicitly selected public agent IDs, allow live model text. Configure public_methods and the public_agent_ids callback deliberately; do not designate a method that mixes private model work with public response generation. Router streams, reasoning, tool-call chunks, tool-enabled calls, and unrelated agents are excluded. Private agent results do not become replies. Providers without observable deltas still publish their canonical final message; no model timing is invented. runner.accept_event(event) claims one runner-issued event before transport handoff. It rejects foreign identity, duplicate/out-of-order sequence, and output invalidated by interruption or failure. Invoke it after any application queue/delay, immediately before handoff. It is not a validator for untrusted client playback receipts. Already handed-off output needs adapter-side identity fencing and audio interruption. runner.interrupt(session_id=..., turn_id=..., input_revision=..., delivery_id=...) targets exactly one reply and is idempotent for valid repeated requests. During generation it marks the turn/reply interrupted; after generation it preserves completed text and marks the reply interrupted without changing job success. on_interrupt may signal existing cooperative stage checks; should_interrupt can observe an application’s existing stop signal. These callbacks do not abort arbitrary provider calls. The stream must settle before another tracked turn can execute on the same Flow. Closing the iterator signals interruption and drains native execution; it can wait for an in-flight provider. To compose with background work, declare class WorkChatState(TurnState, JobState[MyJob]): .... Hold JobRunner.state_lock around the entire foreground iteration and all parent job mutations. TurnRunner separately protects its records and interruption controls. Do not run native turn APIs concurrently with the wrapper. After successfully admitting a job, runner.associate_job(job) links its session/originating turn and revision/attempt to the active reply; it never admits, submits, or cancels work. A new conversational turn does not invalidate independent research. elapsed_ms uses a server monotonic clock from tracked-turn startup. model_first_text_ms, when available, is the difference between provider request-start and first public text event timestamps; it includes provider/SDK/network work and is not isolated inference-engine TTFT. Never subtract browser and server timestamps. TTS first chunk, cached audio, physical audible onset, and playback completion remain adapter measurements. This increment does not supply background reply queues, speaking-floor scheduling, playback receipts, exact word alignment, speculative generation, or job pause/resume. Snapshots contain records, not locks, cancellation signals or running executions. They do not resume work or replay delivery. Restore state before creating a new tracked live turn; from_checkpoint and restore_from_state_id are deliberately rejected by this wrapper. Durable reconciliation and persistence of terminal tracking records remain separate work. Continue using the stable Flow APIs for existing checkpoint workflows. During tracked execution, the current restored state remains authoritative: automatic session reload is suppressed so it cannot discard admitted turn/reply records or accepted background job updates. Native Flow turn APIs retain their existing persistence reload behavior. Closing after a terminal event does not interrupt the completed reply. Failure remains an acceptable terminal lifecycle signal if interruption arrives after failure.

Experimental queued background replies

crewai.experimental.flow_replies adds ReplyQueueState, ReplyQueue, DeliveryRecord, DeliveryFloor, CoveredUpdate, and ClientActivity on top of the turn/reply contract. Declare state that combines ReplyQueueState and JobState[MyJob], then create one TurnRunner and one ReplyQueue for that live Flow. Existing Flows remain opt-in; the job executor and stable conversational APIs do not change. A successful job is immediately available in state, even when its announcement waits. enqueue(job) accepts only a current, committed completed job with an existing originating turn. The pending delivery stores that job’s revision, attempt and accepted update sequence; duplicate notifications coalesce to the same delivery_id. This ID is also the eventual public ReplyRecord identity. Each job completion gets one attributable announcement; intermediate stage progress does not create a speaking queue.
Serialize enqueue, claim, preparation, coverage changes, and output handoff with the same JobRunner.state_lock used by foreground turns and job commits. ReplyQueue also protects its own controls with a lock so two claims cannot acquire the floor. prepare() reads copied current accepted artifacts, rather than cached text captured at enqueue, and emits the shared started/text/completed public reply events with origin and job identity. It does not execute an LLM, append history, or schedule TTS. The application chooses intentional public text, history insertion, delivery tasks and transport. set_foreground(True) holds the gate through foreground generation and draining, including pending user requests; release it when that work has settled. observe_client(ClientActivity(...)) accepts only the owning session and increasing client sequence, with strict boolean recording, playback, and muted observations. Identified playback must belong to the Flow; a stale stop for a different reply cannot clear newer playback. Null delivery_id accommodates a local cached acknowledgment. Recording, playback, mute, a running turn, session closure or an already active delivery blocks claim(). If recording/mute or a new user request arrives during an announcement, the adapter must interrupt and stop that active delivery, then update activity and wake scheduling. claim() atomically changes one eligible delivery from pending to scheduled. Before preparation and each handoff, recheck relevance with accepts_output(delivery_id) and claim text events with turns.accept_event(event). Identity/version changes, missing jobs, non-successful work, ended sessions, already-covered updates and application is_relevant(job) policy invalidate announcements. The callback must be side-effect free; use it for application cancellation/supersession beyond the existing job contract. Wake the delivery task when activity or relevance changes. Relevance skips never remove retained artifacts. After a requested summary has actually been handed off, mark_covered(CoveredUpdate.from_job(job)) records the accepted version and skips matching pending announcements. Merely requesting a summary is not delivery. A later job version remains eligible. Do not mark all context jobs covered by a status response or unrelated conversation. settle(delivery_id, status, reason=...) releases only the active delivery, once. Its terminal status is completed, interrupted, skipped, or failed; generation completion alone does not settle delivery. Validate feedback session, turn, input revision and delivery identity at the adapter boundary, and accept completion only after output generation has settled. Whole-message adapter completion is a scheduling observation, not proof of audible words. On interruption, failure or invalidation, stop local playback and in-flight adapter delivery, fence remaining public output, and leave the result available. Interrupted/failed announcements are not automatically replayed; users can request a fresh summary. TTS failure does not change job success. Call close() before cancelling and draining your application-owned delivery tasks. It fences pending/active deliveries while retaining job outputs. Snapshot data includes floor observations and delivery records, not a timer, media connection or resumable lease. Restoration does not restart delivery; reconcile active records through application recovery policy. Segment-level receipts, exact spoken context, automatic job cancellation/replacement, distributed scheduling and playback timeouts remain separate features. A client that never reports completion conservatively retains the floor until interrupted or closed; applications can add an explicit timeout policy.