feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -1,81 +1,88 @@
{
"caseId": "chat-diagnosis-flow",
"query": "What is the standard troubleshooting flow for an application incident?",
"retrievedAt": "2026-07-06T00:00:00Z",
"lookupResult": {
"found": true,
"evidenceBlocks": [
{
"source": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"retrievalLayer": "L1",
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
"score": 0.82,
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"source": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"retrievalLayer": "L1",
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
"score": 0.55,
"hitReasons": []
}
],
"contextPack": {
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
"strategy": "top_evidence_blocks",
"charBudget": 3500,
"usedChars": 192,
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
"omittedSources": []
"caseId" : "chat-diagnosis-flow",
"query" : "What is the standard troubleshooting flow for an application incident?",
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
"searchMode" : "hybrid",
"kbScope" : "rag-eval",
"lookupResult" : {
"found" : true,
"evidenceBlocks" : [ {
"docId" : "incident-diagnosis-flow",
"chunkIndex" : 1,
"evidenceKey" : "incident-diagnosis-flow#chunk-1",
"source" : "incident-diagnosis-flow",
"title" : "Diagnosis Flow",
"breadcrumb" : "AIOps > Diagnosis Flow",
"retrievalLayer" : "L1",
"content" : "## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.",
"score" : 0.032786883413791656,
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"docId" : "incident-diagnosis-flow",
"chunkIndex" : 0,
"evidenceKey" : "incident-diagnosis-flow#chunk-0",
"source" : "incident-diagnosis-flow",
"title" : "AIOps",
"breadcrumb" : "AIOps > Diagnosis Flow",
"retrievalLayer" : "L1",
"content" : "# AIOps",
"score" : 0.016129031777381897,
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
} ],
"contextPack" : {
"packedText" : "[Evidence 1]\nsource: incident-diagnosis-flow\ntitle: Diagnosis Flow\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.\n\n[Evidence 2]\nsource: incident-diagnosis-flow\ntitle: AIOps\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
"strategy" : "ranked_evidence_char_budget",
"charBudget" : 4000,
"usedChars" : 1032,
"includedSources" : [ "incident-diagnosis-flow", "incident-diagnosis-flow" ],
"omittedSources" : [ ]
},
"retrievalTrace": {
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
"categoryFilter": "AIOps",
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"queryHints": {
"domains": ["AIOps"],
"matched_keywords": ["collect evidence", "verify", "remediation"],
"entities": ["application incident"],
"l0_titles": ["Incident Diagnosis Flow"],
"l0_match_count": 1
"retrievalTrace" : {
"originalQuery" : "What is the standard troubleshooting flow for an application incident?",
"rewrittenQuery" : "What is the standard troubleshooting flow for an application incident?",
"categoryFilter" : "ops",
"selectedAttempt" : "FILTERED_VECTOR",
"fallbackReason" : null,
"evidenceStatus" : "supported",
"queryHints" : {
"domains" : [ "ops" ],
"matched_keywords" : [ "standard troubleshooting flow", "application incident" ],
"entities" : [ "standard troubleshooting flow", "application incident" ],
"l0_titles" : [ "Incident Diagnosis Flow" ],
"l0_match_count" : 1
},
"attempts": [
{
"name": "FILTERED_VECTOR",
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
"categoryFilter": "AIOps",
"candidateCount": 2,
"usable": true,
"durationMs": 10,
"topScore": 0.82,
"topSimilarity": 0.82
}
]
"attempts" : [ {
"name" : "FILTERED_VECTOR",
"query" : "What is the standard troubleshooting flow for an application incident?",
"categoryFilter" : "ops",
"candidateCount" : 2,
"usable" : true,
"errorMessage" : null,
"durationMs" : 1540,
"topScore" : 0.032786883413791656,
"topSimilarity" : 0.6828859150409698
} ]
},
"rerankTrace": {
"items": [
{
"finalRank": 1,
"source": "incident-diagnosis-flow",
"baseScore": 0.82,
"finalScore": 1.07,
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "rag-chunk-context-reconstruction",
"baseScore": 0.55,
"finalScore": 0.55,
"boostReasons": []
}
]
}
"rerankTrace" : {
"items" : [ {
"finalRank" : 1,
"source" : "incident-diagnosis-flow",
"baseScore" : 0.6828859150409698,
"finalScore" : 0.6828859150409698,
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
}, {
"finalRank" : 2,
"source" : "incident-diagnosis-flow",
"baseScore" : 0.3166210651397705,
"finalScore" : 0.3166210651397705,
"boostReasons" : [ "l0_domain_overlap" ]
} ]
},
"evidenceCandidateCount" : 2,
"evidenceBlockCount" : 2,
"relevanceLevel" : "REFERENCE",
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
"retrievedDomainsThisSession" : null,
"message" : null
}
}
}