模块 3 · 铁律与流式

两条铁律 +
prepare_customer_stream

客户线能跑通,靠文件头写死的两条 铁律。 流式 SSE 看起来是「边生成边推」,实现上是先跑完整张 LangGraph,再切块推送。

铁律(指挥 AI 改代码时别碰)

铁律 1 · 数值不经过 LLM 编造

持仓、流水、风评等查询结果 100% 来自 Core 只读查询 Tool(core_ro_tool.py)的 fact_text。LLM 在 interpret / generate 里只做解读与组织语言,不能自己算数。

铁律 2 · customer_id 只来自 JWT

四 Agent 对话 HTTP 入口(chat.py)在进图之前已从令牌解析客户号并做归属校验。图内 state["customer_id"] 不信用户消息里的「帮我查 CUST-xxx」。

和 advisor 线的差异: 理财师走顾问通用编排(agent_service.py)的 tool→llm→guard;客户走独立 14 节点图。别混用对话 Tool 编排(tool_service.py)关键词意图那套来改客户 RAG 分支。

流式真相:图跑完再切块

方案 C 下 prepare_customer_stream 先调用 run_customer_chat 跑完整图,再用 _chunk_reply_text 切成 16 份左右推 SSE——Tool/RAG 仍是同步完成,不是 token 级流式。

CODE · prepare_customer_stream
def prepare_customer_stream(
    ctx, message, session_id, customer_id, end_session=False
) -> dict[str, Any]:
    reply, has_disclaimer, intent, transfer = run_customer_chat(
        ctx, message, session_id, customer_id, end_session
    )
    return {
        "reply": reply,
        "has_disclaimer": has_disclaimer,
        "intent": intent,
        "transfer_to_human": transfer,
        "chunks": _chunk_reply_text(reply),
    }
白话

入参里的 customer_id 是四 Agent 对话 HTTP 入口(chat.py)验完 JWT 后塞进来的,函数内部不再解析用户文本。

先拿到完整 reply(含 intent、是否转人工、是否附免责),再切片给 SSE 层逐块写。

改「真流式」要先动 LangGraph 执行模型,不是只改前端 ChatPanel。

相关文件(改流式 / 铁律时打开这些)

app/service/
登录客户 14 节点 LangGraph 编排(customer_service.py)— prepare_customer_stream
游客试聊 9 节点 LangGraph 编排(visitor_service.py)— 无 Tool 查持仓
RAG 检索层(rag_service.py)— fin_faq / fin_product / fin_policy
app/api/
四 Agent 对话 HTTP 入口(chat.py)— customer 分支分流 + SSE 写帧
宿主 AuthContext → 模块 AuthContext 适配(auth_adapter.py)— host_auth_for_customer_service
app/tool/
客户 Agent Core 只读查询工具(core_ro_tool.py)— query_holdings / query_trades 等

客户 SSE 流式回复,LangGraph 什么时候跑完?