Files
group_fqcd_jr/app/core/knowledge_schema.py
T
张胜宇 5d0becb67d 客服 Agent 重构收口:五出口决策链 + 知识库档位隔离 + 前端入参边界(答辩演示版本)
一、客服 Agent 智能增强(正面回应"不智能、动不动就转人工")
- 决策链由 2 个出口扩到 5 个:E1 澄清 / E2 计算型 / E3 知识直返 / E4 证据约束生成 / E5 分级回退
- 转人工从"默认动作"降为最后一档 E5c,只保留 4 类白名单:
  P0 反诈 / P1 账户与个人数据 / P2 写操作与争议 / 用户明确要求人工
- 46 条金标实测(修复前 → 修复后):
  转人工率 43.5% → 10.9%;出口准确率 45.7% → 100%;事实正确率 69.6% → 100%
  禁忌违反 1 → 0;档位越权 / 无出处数字 / 误拒 四项零容忍全 0
- 安全不变量 INV-1~INV-5;零容忍规则未删,改的是挂载点
  (输出侧字面黑名单 → 检索层档位隔离 + 判定层合规词表 + 输出守护)

二、知识库:档位单点化与物理隔离
- 新增 app/core/knowledge_tier.py 作为档位规则唯一落点(G-03),
  knowledge_contracts.py 原定义块改为显式再导出(X as X,非副本)
- 档位过滤由 bool 默认值(fail-open)改为 tiers 必填集合(缺参即 TypeError)
- Milvus 侧四集合按 visibility 分区键物理隔离;双 schema 收敛为一套
- 新增 app/core/actor.py:访客三元组与匿名判定的唯一构造/判定点(G-01/G-01b)
- 新增 app/core/fund_fee_rules.py:费率计算纯函数

三、前端入参边界对齐(本轮 W11 新修,4 处"校验宽于存储")
- message 加 max_length=8000(与浮窗 widget.js 的 maxlength 一致)
- session_id 加 1—64;idempotency_key 上限 128 → 64(对齐列宽 String(64))
- feedback_type 加 max_length=32(对齐列宽 String(32))
- 8 条路径参数补 min_length=1 + max_length=64 + 字符集正则
  ({session_id} / {run_id} / {handover_id})
- 改前超限值会落到 MySQL 才失败(500);改后一律 422 AGENT_INPUT_INVALID + 字段级定位
- 新增 tests/unit/api/test_frontend_boundaries.py(33 例),含"端点表 ↔ OpenAPI 全量对照"

四、投顾模块整体清除(D4.4 / D4.5)
- 删除投顾相关 controller / schema / model / repository / service 及门户页面
- tools/portal_api_check.py 同步作废 AD003/AD005/AD011/A047 四条用例与 advisor_t 登录
  (端点与账号均已不存在,此前稳定报 3 条假红)

五、验证(提交前实测)
- pytest -q:1856 passed / 2 skipped / 0 failed
- ruff check app tools tests:19(= 基线);mypy app:2(= 基线)
- 前端接口契约体检 portal_api_check.py:38 项,通过 34,失败 0,跳过 4
- 全链路冒烟 e2e_smoke_test.py --read-only:31/31
- HTTP 全链路探针 http_probe.py:11/11 succeeded
- 跨文档一致性 _consistency.py:GATE PASS
- 真机边界复验 12 条:12/12 符合预期

六、纪律与文档
- 可改文件白名单 A-09(docs/46)与底座会签申请单 A-10(docs/47,组 1—组 4 全部受理)
- 零 DDL:未新增/修改任何表结构,89 张业务表与基线一致
- 证据留痕:docs/evidence/**(含 46 条金标 score、快照、清除与重建记录)
- 未提交(刻意排除,见提交说明):仓库内 客服agent/ 与 开发文档/ 是 2026-09-16 前的
  过期副本(Todolist 440 行 vs 权威 D2.1 1167 行),权威正本在仓库外;
  _chunks_report.txt 是 tools/build_knowledge_chunks.py 生成的本地产物
2026-09-20 14:33:30 +08:00

223 lines
9.1 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""知识集合的**运行时字段探测**:把"逻辑字段名"解析成该集合真实的物理字段名。
## 为什么需要它
同一批集合名(`fin_faq_collection` 等)在不同环境里可能是**两套完全不同的 schema**,
这是实测确认的事实(2026-09-11,两个开发环境各自 `describe_collection`):
| 逻辑字段 | 环境甲(本机) | 环境乙(架构师机) |
|---|---|---|
| 文档标识 | `knowledge_id` | `doc_id` |
| 正文 | `snippet` | `content` |
| 可见性 | **无** | `visibility` |
| 来源文件 | **无** | `source_file` |
| 章节 | **无** | `chapter` / `section` / `doc_no` |
| 行数 | 106 / 177 / 73 | 125 / 297 / 214 |
**硬编码任何一套都会打挂另一套**:Milvus 对不存在的字段直接报错
(`field doc_id not exist`)→ 三个集合全失败 → `degraded=True` → 客服一律"引导人工"。
把字段名换成探测之后,两套 schema 都能跑,**没有需要改回去的东西**,也不需要迁移或重灌。
## 设计要点
1. **只探测一次**:`detect_schema()` 带缓存,`describe_collection` 是纯元数据调用、不查数据;
探测失败不抛异常,返回"该集合不可用",由调用方走降级路径(与检索服务一贯口径一致)。
2. **缺失字段不报错、只记录**:`chapter`/`section`/`visibility` 缺失时,依赖它们的增强逻辑
(父子块选择、内部资料过滤)自然退化为"不启用",而不是让整条检索失败。
3. **关键字段缺失才失败**:既没有文档标识字段、又没有正文字段的集合,**必须明确报出来**
(`missing_required`)——那种集合检索不出任何有意义的结果,静默零召回比报错更难定位。
4. **不 import 检索服务**:本模块只依赖 `collections.abc`/`dataclasses`/`typing`,
避免与 `knowledge_search_service` 形成循环依赖。
"""
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from typing import Any, Protocol
#: 逻辑字段名 → 该字段在各环境里**可能**的物理名(按优先级排列)。
#:
#: 顺序有意义:第一个命中的即采用。例:某环境同时有 `doc_id` 与 `knowledge_id` 时优先 `doc_id`
#: (那是灌库脚本的正式设计名),保持与架构师那套一致的语义。
FIELD_CANDIDATES: Mapping[str, tuple[str, ...]] = {
"doc_id": ("doc_id", "knowledge_id"),
"content": ("content", "snippet"),
"title": ("title",),
"tags": ("tags",),
"version": ("version",),
"intent": ("intent",),
# 以下为**可选**增强字段:缺了就退化为"不启用",不影响检索可用性
"visibility": ("visibility",),
"source_file": ("source_file",),
"chapter": ("chapter",),
"section": ("section",),
"doc_no": ("doc_no",),
# v1.4(2026-09-18)新增:`D2.4` 附录A 的「同族合并」与「计算型参数位」两条能力
# 需要它们。缺了就退化为"不启用",不影响检索可用性(与其余可选增强字段一致)。
"family_id": ("family_id",),
"param_class": ("param_class",),
}
#: 缺了就无法检索的字段(既无标识、又无正文的集合没有可用结果)
REQUIRED_LOGICAL_FIELDS: tuple[str, ...] = ("doc_id", "content")
class SchemaProbe(Protocol):
"""只依赖 `describe_collection`(pymilvus 的 `MilvusClient`/`AsyncMilvusClient` 都有)。"""
def describe_collection(self, **kwargs: Any) -> Any: ...
@dataclass(frozen=True)
class CollectionSchema:
"""一个集合的字段解析结果。
`fields` 是「逻辑名 → 物理名」的映射,**只含实际存在的字段**。
`missing_optional` 记录缺失的可选增强字段(用于日志与诊断,不影响检索)。
`missing_required` 非空表示这个集合不可用(调用方应记 `degraded`)。
"""
collection: str
fields: Mapping[str, str]
physical_names: frozenset[str]
missing_optional: tuple[str, ...] = ()
missing_required: tuple[str, ...] = ()
error: str | None = None
#: 是否具备检索的最低字段条件
@property
def usable(self) -> bool:
return not self.missing_required and self.error is None
def resolve(self, logical: str) -> str | None:
"""逻辑名 → 物理名;该集合没有这个字段时返回 None(调用方据此跳过)。"""
return self.fields.get(logical)
@property
def output_fields(self) -> tuple[str, ...]:
"""本次检索应从 Milvus 取回的物理字段(只取存在的,避免"字段不存在"报错)。"""
return tuple(self.fields.values())
def has(self, logical: str) -> bool:
return logical in self.fields
def _physical_names(description: Any) -> frozenset[str]:
"""从 `describe_collection` 的返回里取出字段名集合。
pymilvus 不同版本返回形状略有差异(`{"fields": [{"name": ...}]}` 或带 `schema` 包装),
这里做兼容解析;解析不出任何字段名时返回空集合 → 调用方会判定为不可用。
"""
if not isinstance(description, Mapping):
return frozenset()
fields = description.get("fields")
if fields is None:
schema = description.get("schema")
fields = schema.get("fields") if isinstance(schema, Mapping) else None
if not isinstance(fields, Sequence):
return frozenset()
names: list[str] = []
for item in fields:
if isinstance(item, Mapping):
name = item.get("name")
if isinstance(name, str) and name:
names.append(name)
return frozenset(names)
def resolve_schema(collection: str, description: Any) -> CollectionSchema:
"""把一次 `describe_collection` 结果解析成 `CollectionSchema`(纯函数,不抛异常)。
单独抽出来是为了让单测能直接喂两套 schema 的替身描述,不必连 Milvus。
"""
names = _physical_names(description)
if not names:
return CollectionSchema(
collection=collection, fields={}, physical_names=frozenset(),
missing_required=REQUIRED_LOGICAL_FIELDS,
error="describe_collection 未返回字段信息",
)
resolved: dict[str, str] = {}
missing_optional: list[str] = []
missing_required: list[str] = []
for logical, candidates in FIELD_CANDIDATES.items():
found = next((c for c in candidates if c in names), None)
if found is not None:
resolved[logical] = found
continue
if logical in REQUIRED_LOGICAL_FIELDS:
missing_required.append(logical)
else:
missing_optional.append(logical)
return CollectionSchema(
collection=collection,
fields=resolved,
physical_names=names,
missing_optional=tuple(missing_optional),
missing_required=tuple(missing_required),
)
class SchemaCache:
"""按集合名缓存探测结果(进程内,一次探测)。
缓存的是**元数据**:集合重建(换 schema)后需要重启进程或调用 `invalidate()`。
这是刻意的取舍——检索是热路径,不能每次调用都打一次 describe_collection。
"""
def __init__(self) -> None:
self._cache: dict[str, CollectionSchema] = {}
def get(self, collection: str) -> CollectionSchema | None:
return self._cache.get(collection)
def put(self, schema: CollectionSchema) -> None:
self._cache[schema.collection] = schema
def invalidate(self, collection: str | None = None) -> None:
if collection is None:
self._cache.clear()
else:
self._cache.pop(collection, None)
def detect_schema(
client: Any, collection: str, *, cache: SchemaCache | None = None
) -> CollectionSchema:
"""探测一个集合的字段映射。**不抛异常**:任何失败都转成不可用的 `CollectionSchema`。
`client` 需要提供 `describe_collection(collection_name=...)`(同步或异步都可——
这里只调用同步形式;pymilvus 的 `MilvusClient` 是同步的,检索服务用的就是它)。
"""
if cache is not None:
cached = cache.get(collection)
if cached is not None:
return cached
probe = getattr(client, "describe_collection", None)
if probe is None:
schema = CollectionSchema(
collection=collection, fields={}, physical_names=frozenset(),
missing_required=REQUIRED_LOGICAL_FIELDS,
error="客户端不支持 describe_collection",
)
if cache is not None:
cache.put(schema)
return schema
try:
description = probe(collection_name=collection)
except Exception as exc: # 探测失败=该集合不可用,由调用方降级
schema = CollectionSchema(
collection=collection, fields={}, physical_names=frozenset(),
missing_required=REQUIRED_LOGICAL_FIELDS, error=f"{type(exc).__name__}: {exc}",
)
if cache is not None:
cache.put(schema)
return schema
schema = resolve_schema(collection, description)
if cache is not None:
cache.put(schema)
return schema