chore(sync): zsy_developcc 全量同步至 qyqy_develop(W26 口径)
- 分支内容对齐 qyqy_develop b6ec3aa,树完全一致(同步后 git diff 为空) - 覆盖本轮全部交付:客服 Agent 重构(安全路由 / 五出口 / 记忆与画像 / RAG 全链路) + 开发文档 62 份编号体系(D1.1 v1.17 索引) + 新增 D2.10-客服Agent端到端答辩文档-2026-09-21.html - 基线:e239eb7(2026-09-17 品牌口径统一快照),本提交为其直接后继
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
"""幂等创建知识检索用的三个 Milvus 集合(Task 6)。
|
||||
"""幂等创建知识检索用的四个 Milvus 集合(Task 6)。
|
||||
|
||||
用法::
|
||||
|
||||
@@ -9,11 +9,26 @@
|
||||
1. **幂等**:集合已存在则**直接跳过**,不重建、不覆盖、不清数据;
|
||||
2. **不覆盖不同结构**:同名集合已存在但字段/维度与本脚本定义不同时,**停下来报告**
|
||||
并以非零退出码结束 —— 宁可人工确认,也不动别人的数据;
|
||||
3. 只处理 `ALLOWED_COLLECTIONS` 里的三个集合,集合名**不接受外部参数**。
|
||||
3. 只处理 `ALLOWED_COLLECTIONS` 里的四个集合,集合名**不接受外部参数**。
|
||||
|
||||
schema 三个集合逐字相同(见实施计划 Task 6):`knowledge_id`(VARCHAR64, 主键) /
|
||||
`title`(256) / `snippet`(4000) / `tags`(512) / `version`(16) / `intent`(32) /
|
||||
`embedding`(FLOAT_VECTOR dim=1024),索引 `AUTOINDEX` + `metric_type="COSINE"`。
|
||||
schema 四个集合逐字相同,且**本模块是唯一权威定义**(`H-05` ④「两套建表脚本收敛为一套」):
|
||||
|
||||
| 字段 | 类型 / 长度 | 角色 |
|
||||
|---|---|---|
|
||||
| `doc_id` | VARCHAR(64) | **主键** |
|
||||
| `title` / `content` | VARCHAR(1024) / VARCHAR(16384) | 检索与输出 |
|
||||
| `chapter` / `section` / `tags` | VARCHAR(512) | 定位与分组 |
|
||||
| `doc_no` / `version` / `effective_date` / `expire_date` | VARCHAR(64/32/32/32) | 溯源与时效 |
|
||||
| `source_url` / `reviewer` / `source_file` | VARCHAR(512/64/128) | 溯源与合规留痕 |
|
||||
| `family_id` / `param_class` / `intent` | VARCHAR(64/16/32) | 同族 / 参数类型 / 意图(v1.4) |
|
||||
| `visibility` | VARCHAR(16) | **分区键(档位隔离)** |
|
||||
| `embedding` | FLOAT_VECTOR(dim=1024) | 向量 |
|
||||
|
||||
索引 `AUTOINDEX` + `metric_type="COSINE"`。**为什么收敛到这一套**:`app/core/knowledge_schema.py`
|
||||
的 `FIELD_CANDIDATES` 优先取 `doc_id` / `content`(注释即写明"那是灌库脚本的正式设计名"),
|
||||
`app/infrastructure/milvus_knowledge_writer.py` 的逻辑主键/正文名也是 `doc_id` / `content`,
|
||||
`knowledge/_chunks.jsonl` 的产物字段同样是这一套 —— 旧脚本里的 `knowledge_id` / `snippet` /
|
||||
`intent` 属**另一套环境的历史命名**,保留它等于让"同名集合两套字段"继续活着。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -28,37 +43,82 @@ sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
from app.core.config import get_settings # noqa: E402
|
||||
from app.core.knowledge_contracts import ALLOWED_COLLECTIONS, VECTOR_DIM # noqa: E402
|
||||
|
||||
PRIMARY_FIELD = "knowledge_id"
|
||||
PRIMARY_FIELD = "doc_id"
|
||||
VECTOR_FIELD = "embedding"
|
||||
INDEX_NAME = "knowledge_autoindex"
|
||||
|
||||
#: 档位字段 = **分区键**。集合内按档位物理分桶,检索按该字段过滤时由引擎做分区裁剪,
|
||||
#: 不可见档位不进候选集 —— 这是"检索层硬隔离",不依赖上层自觉。
|
||||
#:
|
||||
#: ⚠️ 两条实测结论(2026-09-18,Milvus v2.5.3):
|
||||
#: 1. 分区键模式下 **Milvus 禁止手工 `create_partition`**(报
|
||||
#: `disable create partition if partition key mode is used`)⇒ 新增档位值
|
||||
#: **不需要任何运维动作**,由引擎按哈希自动路由。设计文档里"档位值变更 = 建分区"
|
||||
#: 的表述据此修正为"档位值变更 = 无需动作,引擎自动路由"。
|
||||
#: 2. `upsert` / `insert` / 按分区键过滤的 `search` 在该模式下**均正常**
|
||||
#: (已用抛废集合实测),所以写入侧(走 `upsert`)不受影响。
|
||||
PARTITION_KEY_FIELD = "visibility"
|
||||
|
||||
#: 分区桶数。**创建后不可改**,所以宁可一次给足:档位枚举预计会增长,16 桶可保证
|
||||
#: 不同档位值大概率落在不同桶、裁剪真正生效;600+ 行规模下多桶开销可忽略。
|
||||
NUM_PARTITIONS = 16
|
||||
|
||||
#: (字段名, VARCHAR 最大长度);顺序与写入侧口径一致。
|
||||
#:
|
||||
#: 这份字段表是**唯一权威**:灌库脚本(`tools/load_knowledge_milvus.py`)从本模块
|
||||
#: import `VARCHAR_FIELDS` / `FIELD_LIMITS` 做截断,**不再自带第二份定义**(`H-05` ④)。
|
||||
#: 长度按**实际数据最坏情况**定:`content` 取 16384(实测最长块 2828 字符,留足余量),
|
||||
#: `title` 取 1024(含完整章节路径)。
|
||||
VARCHAR_FIELDS: tuple[tuple[str, int], ...] = (
|
||||
(PRIMARY_FIELD, 64),
|
||||
("title", 256),
|
||||
("snippet", 4000),
|
||||
("title", 1024),
|
||||
("content", 16384),
|
||||
("chapter", 512),
|
||||
("section", 512),
|
||||
("tags", 512),
|
||||
("version", 16),
|
||||
("doc_no", 64),
|
||||
("version", 32),
|
||||
("effective_date", 32),
|
||||
("expire_date", 32),
|
||||
("source_url", 512),
|
||||
("reviewer", 64),
|
||||
("source_file", 128),
|
||||
# v1.4(2026-09-18):`D2.4` 附录F 的「同族合并 / 计算型参数位 / 意图标签」三条能力
|
||||
# 需要它们。`family_id` 64 足够(父块编号最长 `POL-AST-012` 这类);`param_class`
|
||||
# 是固定枚举 `none/rate/threshold/scale/count`;`intent` 是五类业务意图。
|
||||
("family_id", 64),
|
||||
("param_class", 16),
|
||||
("intent", 32),
|
||||
(PARTITION_KEY_FIELD, 16),
|
||||
)
|
||||
|
||||
#: 允许缺省的字段。`intent` 是**稀疏标签**:知识契约里 `None` 表示"无显式标签、由集合名推断"
|
||||
#: (见 `app/core/knowledge_contracts.py` 的 `intent_for_qa_id` docstring),
|
||||
#: 因此写入侧对普通知识**省略该字段**。若集合把它定义为 non-nullable 且无默认值,
|
||||
#: `upsert` 会抛 `Insert missed an field 'intent'`(已实测:106 条里 86 条普通知识全部写不进)。
|
||||
NULLABLE_FIELDS = frozenset({"intent"})
|
||||
#: 字段名 → 最大长度。写入侧据此截断,**与集合定义同源**,避免"脚本截到 500、
|
||||
#: 集合只给 256"这类只在写入时才暴露的错配。
|
||||
FIELD_LIMITS: dict[str, int] = dict(VARCHAR_FIELDS)
|
||||
|
||||
#: 允许缺省的字段。**空集 = 全部必填**:VARCHAR 由写入侧补空串
|
||||
#: (`MilvusKnowledgeWriter` 的既有行为),分区键字段则要求**显式声明**——
|
||||
#: `visibility` 缺省不是"写空串"而是"不计入任何档位",属档位越权风险,
|
||||
#: 必须在入库门禁(`B-01`)就挡住,不能靠集合默认值兜底。
|
||||
NULLABLE_FIELDS: frozenset[str] = frozenset()
|
||||
|
||||
|
||||
def expected_fields() -> dict[str, dict[str, Any]]:
|
||||
"""本脚本期望的字段结构:`{字段名: {type, is_primary?, max_length?, dim?}}`。"""
|
||||
"""本脚本期望的字段结构:`{字段名: {type, is_primary, is_partition_key, max_length?, dim?}}`。"""
|
||||
fields: dict[str, dict[str, Any]] = {}
|
||||
for name, max_length in VARCHAR_FIELDS:
|
||||
fields[name] = {
|
||||
"type": "VARCHAR",
|
||||
"max_length": max_length,
|
||||
"is_primary": name == PRIMARY_FIELD,
|
||||
"is_partition_key": name == PARTITION_KEY_FIELD,
|
||||
}
|
||||
fields[VECTOR_FIELD] = {"type": "FLOAT_VECTOR", "dim": VECTOR_DIM, "is_primary": False}
|
||||
fields[VECTOR_FIELD] = {
|
||||
"type": "FLOAT_VECTOR",
|
||||
"dim": VECTOR_DIM,
|
||||
"is_primary": False,
|
||||
"is_partition_key": False,
|
||||
}
|
||||
return fields
|
||||
|
||||
|
||||
@@ -80,9 +140,17 @@ def describe_mismatch(described: Any) -> list[str]:
|
||||
"type": raw.get("type"),
|
||||
"params": params,
|
||||
"is_primary": bool(raw.get("is_primary")),
|
||||
"is_partition_key": bool(raw.get("is_partition_key")),
|
||||
}
|
||||
|
||||
problems: list[str] = []
|
||||
# 分区桶数也是结构的一部分,而且**创建后不可改** —— 不一致必须报出来,
|
||||
# 否则"看起来建好了"的集合其实没有按预期分桶、分区裁剪名存实亡。
|
||||
actual_partitions = described.get("num_partitions")
|
||||
if isinstance(actual_partitions, int) and actual_partitions != NUM_PARTITIONS:
|
||||
problems.append(
|
||||
f"num_partitions 不一致(实际 {actual_partitions},期望 {NUM_PARTITIONS})"
|
||||
)
|
||||
for name, spec in expected_fields().items():
|
||||
found = actual.get(name)
|
||||
if found is None:
|
||||
@@ -90,6 +158,11 @@ def describe_mismatch(described: Any) -> list[str]:
|
||||
continue
|
||||
if found["is_primary"] != spec["is_primary"]:
|
||||
problems.append(f"字段 {name} 主键标记不一致(实际 {found['is_primary']})")
|
||||
if found["is_partition_key"] != spec["is_partition_key"]:
|
||||
problems.append(
|
||||
f"字段 {name} 分区键标记不一致"
|
||||
f"(实际 {found['is_partition_key']},期望 {spec['is_partition_key']})"
|
||||
)
|
||||
if name == VECTOR_FIELD:
|
||||
dim = found["params"].get("dim")
|
||||
if not isinstance(dim, int) or dim != VECTOR_DIM:
|
||||
@@ -101,7 +174,10 @@ def describe_mismatch(described: Any) -> list[str]:
|
||||
return problems
|
||||
|
||||
|
||||
def _build_schema() -> Any:
|
||||
def build_schema() -> Any:
|
||||
"""构造四个集合的统一 schema。**公开**:灌库脚本直接 import 本函数,
|
||||
以保证"建表的字段定义"与"写库的字段填充"是同一份声明(`H-05` ④)。
|
||||
"""
|
||||
from pymilvus import DataType, MilvusClient # type: ignore[import-untyped]
|
||||
|
||||
schema = MilvusClient.create_schema(auto_id=False, enable_dynamic_field=False)
|
||||
@@ -111,14 +187,17 @@ def _build_schema() -> Any:
|
||||
datatype=DataType.VARCHAR,
|
||||
max_length=max_length,
|
||||
is_primary=(name == PRIMARY_FIELD),
|
||||
# 稀疏标签字段必须可空:写入侧对普通知识**省略**该字段(不是写空串)。
|
||||
# 档位字段是**分区键**(每集合只能有一个,且不可空)。
|
||||
is_partition_key=(name == PARTITION_KEY_FIELD),
|
||||
# 只有被显式列为可缺省的字段允许为空;分区键字段永远不可空。
|
||||
nullable=name in NULLABLE_FIELDS,
|
||||
)
|
||||
schema.add_field(field_name=VECTOR_FIELD, datatype=DataType.FLOAT_VECTOR, dim=VECTOR_DIM)
|
||||
return schema
|
||||
|
||||
|
||||
def _build_index_params() -> Any:
|
||||
def build_index_params() -> Any:
|
||||
"""构造向量索引参数(`AUTOINDEX` + `COSINE`)。**公开**,理由同 `build_schema`。"""
|
||||
from pymilvus import MilvusClient # type: ignore[import-untyped]
|
||||
|
||||
index_params = MilvusClient.prepare_index_params()
|
||||
@@ -132,7 +211,7 @@ def _build_index_params() -> Any:
|
||||
|
||||
|
||||
async def ensure_collections(uri: str, token: str = "") -> tuple[list[str], list[str], list[str]]:
|
||||
"""确保三集合存在。返回 `(创建, 已存在, 结构冲突)`;结构冲突时不覆盖。"""
|
||||
"""确保四个集合存在。返回 `(创建, 已存在, 结构冲突)`;结构冲突时不覆盖。"""
|
||||
from pymilvus import AsyncMilvusClient # type: ignore[import-untyped]
|
||||
|
||||
names = sorted(ALLOWED_COLLECTIONS)
|
||||
@@ -156,8 +235,9 @@ async def ensure_collections(uri: str, token: str = "") -> tuple[list[str], list
|
||||
|
||||
await client.create_collection(
|
||||
collection_name=name,
|
||||
schema=_build_schema(),
|
||||
index_params=_build_index_params(),
|
||||
schema=build_schema(),
|
||||
index_params=build_index_params(),
|
||||
num_partitions=NUM_PARTITIONS,
|
||||
)
|
||||
await client.load_collection(name)
|
||||
created.append(name)
|
||||
|
||||
Reference in New Issue
Block a user