awoooi

Author	SHA1	Message	Date
Your Name	3059897318	feat(governance): auto-deprecate low-trust unused playbooks (>30d) Some checks failed Code Review / ai-code-review (push) Successful in 41s Details CD Pipeline / tests (push) Successful in 3m29s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details trust_drift previously fired alerts forever for playbooks stuck below the 0.2 threshold. With user authorization for governance-class auto-fixes, check_trust_drift now retires playbooks that have been unused for 30+ days (or never used and created 30+ days ago) by flipping status to 'deprecated' before alerting. Alerts now report drifted_count, auto_deprecated_count, and the kept playbook_ids that still need human review (those in their 30d trial window). Existing alert noise from the four currently-drifted playbooks should drop to whatever fraction is genuinely in trial. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-02 12:31:37 +08:00
Your Name	607358c4dd	fix(approval): route SSH actions through SSHProvider on manual approve parse_operation_from_action only knew kubectl and Chinese restart phrases, so any "ssh host '...'" action approved via Telegram fell through to "Could not parse operation type" and reported a fake failure even though the LLM had proposed a valid host repair. Adds OperationType.SSH_HOST, makes the parser detect ssh prefixes (with optional flags / user@host) before kubectl patterns, and routes the SSH_HOST branch in approval_execution.execute_in_background through SSHProvider with the same tool keywords decision_manager uses (ssh_docker_prune / ssh_docker_restart / ssh_systemctl_restart / ssh_diagnose). Unroutable SSH actions now fail loudly with a descriptive error instead of silently breaking. Trigger: 2026-05-02 incidents INC-20260502-D6D0B7 / E12EE4 / 557055 were approved by the user but executor reported "Could not parse" and left the alerts pending. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-02 12:31:37 +08:00
Your Name	3156ff1c69	feat(aiops): add ssh_docker_prune to auto-repair flywheel for disk-full alerts Adds Group B SSH MCP tool ssh_docker_prune (image+volume+builder prune with ≥75% disk usage gate) and routes "docker prune" actions through it. Flips HostDiskUsageHigh from auto_repair=false to true with mcp_provider routing labels so the flywheel can self-heal next disk-full event without hitting the emergency_channel Telegram path. Trigger: 2026-05-01 → 05-02 Telegram alert storm (peak 53/hr) caused by empty ssh-mcp-key/known_hosts secret rejecting all SSH and forcing every disk-full alert through "Host key is not trusted → escalate" loop. known_hosts patched live; this commit closes the playbook gap so the next occurrence resolves without manual intervention. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-02 12:31:37 +08:00
Your Name	8cf559215c	docs(awooop): add Phase 1 Isolation Foundation implementation plan (ADR-106 P1)	2026-05-02 12:28:33 +08:00
Your Name	443947ffa1	fix(ci): avoid code review sigpipe on large diffs [skip ci]	2026-05-01 20:59:14 +08:00
Your Name	7795f027d2	fix(aiops): persist emergency intervention traces Some checks failed CD Pipeline / tests (push) Successful in 2m56s Details Code Review / ai-code-review (push) Failing after 39s Details CD Pipeline / build-and-deploy (push) Successful in 12m54s Details CD Pipeline / post-deploy-checks (push) Successful in 4m40s Details	2026-05-01 20:34:33 +08:00
Your Name	8e49f2ea88	fix(ci): preserve ssh mcp known hosts [skip ci]	2026-05-01 17:18:32 +08:00
Your Name	433f7b068e	fix(aiops): close ssh and telegram remediation gaps All checks were successful CD Pipeline / tests (push) Successful in 2m7s Details Code Review / ai-code-review (push) Successful in 42s Details CD Pipeline / build-and-deploy (push) Successful in 13m14s Details CD Pipeline / post-deploy-checks (push) Successful in 4m29s Details	2026-05-01 16:53:02 +08:00
Your Name	3650fc727a	docs(ci): record runner user service takeover state All checks were successful Code Review / ai-code-review (push) Successful in 45s Details	2026-05-01 16:30:54 +08:00
Your Name	bc295eaec2	fix(ci): allow user service for gitea host runner Some checks failed Code Review / ai-code-review (push) Has been cancelled Details	2026-05-01 16:24:45 +08:00
Your Name	cb5ab900c4	fix(ci): preserve gitea runner jobs on shutdown All checks were successful Code Review / ai-code-review (push) Successful in 46s Details	2026-05-01 16:16:27 +08:00
Your Name	b0da6da1e9	feat(aiops): structure agent loop shadow output Some checks failed CD Pipeline / tests (push) Successful in 2m50s Details Code Review / ai-code-review (push) Successful in 33s Details CD Pipeline / build-and-deploy (push) Failing after 25m48s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details	2026-05-01 15:09:57 +08:00
Your Name	f8e44971c1	feat(aiops): enable read-only agent loop canary All checks were successful CD Pipeline / tests (push) Successful in 1m43s Details Code Review / ai-code-review (push) Successful in 31s Details CD Pipeline / build-and-deploy (push) Successful in 10m22s Details CD Pipeline / post-deploy-checks (push) Successful in 4m3s Details	2026-05-01 14:20:16 +08:00
Your Name	6ec3f116fd	fix(ci): normalize migration database url for psql All checks were successful CD Pipeline / tests (push) Successful in 1m30s Details Code Review / ai-code-review (push) Successful in 27s Details CD Pipeline / build-and-deploy (push) Successful in 13m20s Details CD Pipeline / post-deploy-checks (push) Successful in 3m36s Details	2026-05-01 13:30:32 +08:00
Your Name	7e4d995e4b	feat(aiops): add mcp agent loop foundation Some checks failed CD Pipeline / tests (push) Successful in 1m59s Details Code Review / ai-code-review (push) Successful in 28s Details run-migration / migrate (push) Failing after 24s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-05-01 13:21:19 +08:00
Your Name	9db87f177e	fix(aiops): suppress repeated llm alert loops Some checks failed CD Pipeline / tests (push) Successful in 1m37s Details Code Review / ai-code-review (push) Successful in 28s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-05-01 13:02:07 +08:00
Your Name	11673d80ea	fix(aiops): route backup decisions through ssh Some checks failed CD Pipeline / tests (push) Successful in 1m35s Details Code Review / ai-code-review (push) Successful in 34s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-05-01 12:50:01 +08:00
Your Name	3a6acae408	fix(km): add phase25 knowledge enum labels Some checks failed CD Pipeline / tests (push) Successful in 2m14s Details Code Review / ai-code-review (push) Successful in 26s Details run-migration / migrate (push) Failing after 24s Details CD Pipeline / build-and-deploy (push) Has been cancelled Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details	2026-05-01 11:03:03 +08:00
Your Name	2c12bce135	fix(aiops): use existing escalation event type Some checks failed CD Pipeline / tests (push) Successful in 1m54s Details Code Review / ai-code-review (push) Successful in 29s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-05-01 10:56:59 +08:00
Your Name	97be5dedd7	fix(aiops): escalate failed host verification Some checks failed CD Pipeline / tests (push) Successful in 1m27s Details Code Review / ai-code-review (push) Successful in 29s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-05-01 10:47:42 +08:00
Your Name	e4aef6ac4e	fix(aiops): block k8s playbooks for host repair All checks were successful CD Pipeline / tests (push) Successful in 1m27s Details Code Review / ai-code-review (push) Successful in 26s Details CD Pipeline / build-and-deploy (push) Successful in 8m6s Details CD Pipeline / post-deploy-checks (push) Successful in 3m31s Details	2026-05-01 10:33:52 +08:00
Your Name	ca22ec2fd2	fix(aiops): route backup failures rule-first All checks were successful CD Pipeline / tests (push) Successful in 1m51s Details Code Review / ai-code-review (push) Successful in 30s Details Deploy Alert Rules / Deploy Prometheus Alert Rules (push) Successful in 42s Details CD Pipeline / build-and-deploy (push) Successful in 8m21s Details CD Pipeline / post-deploy-checks (push) Successful in 4m18s Details	2026-05-01 10:11:10 +08:00
Your Name	f154ac022e	feat(playbook): version generated playbooks All checks were successful CD Pipeline / tests (push) Successful in 1m34s Details Code Review / ai-code-review (push) Successful in 28s Details Type Sync Check / check-type-sync (push) Successful in 1m10s Details CD Pipeline / build-and-deploy (push) Successful in 10m19s Details CD Pipeline / post-deploy-checks (push) Successful in 3m1s Details	2026-04-30 23:59:39 +08:00
Your Name	f0d14ab6c4	fix(aiops): escalate blocked auto repair Some checks failed CD Pipeline / tests (push) Successful in 1m33s Details Code Review / ai-code-review (push) Successful in 28s Details Deploy Alert Rules / Deploy Prometheus Alert Rules (push) Successful in 40s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-04-30 23:49:17 +08:00
Your Name	6e04fe9c8a	feat(playbook): generate drafts with local llm Some checks failed CD Pipeline / tests (push) Successful in 1m28s Details Code Review / ai-code-review (push) Successful in 29s Details Type Sync Check / check-type-sync (push) Failing after 2m41s Details CD Pipeline / build-and-deploy (push) Successful in 8m40s Details CD Pipeline / post-deploy-checks (push) Successful in 3m10s Details	2026-04-30 23:04:58 +08:00
Your Name	95110971f3	fix(telegram): close remaining DM alert routes Some checks failed CD Pipeline / tests (push) Successful in 1m27s Details Code Review / ai-code-review (push) Successful in 29s Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details	2026-04-30 23:02:17 +08:00
Your Name	712d3e5a77	fix(ci): send workflow alerts to SRE group All checks were successful CD Pipeline / tests (push) Successful in 1m30s Details Code Review / ai-code-review (push) Successful in 26s Details CD Pipeline / build-and-deploy (push) Successful in 7m48s Details CD Pipeline / post-deploy-checks (push) Successful in 2m58s Details	2026-04-30 15:05:16 +08:00
Your Name	61f5a6a419	fix(telegram): route alerts to SRE war room Some checks failed CD Pipeline / tests (push) Has been cancelled Details CD Pipeline / build-and-deploy (push) Has been cancelled Details CD Pipeline / post-deploy-checks (push) Has been cancelled Details Code Review / ai-code-review (push) Has been cancelled Details	2026-04-30 15:01:23 +08:00
Your Name	ed2a4838f2	fix(auto): use action parser for repair gates Some checks failed CD Pipeline / tests (push) Failing after 1m2s Details CD Pipeline / build-and-deploy (push) Has been skipped Details CD Pipeline / post-deploy-checks (push) Has been skipped Details Code Review / ai-code-review (push) Successful in 24s Details	2026-04-30 14:06:09 +08:00
Your Name	4723499955	fix(cd): install playwright system deps for smoke All checks were successful CD Pipeline / tests (push) Successful in 1m34s Details Code Review / ai-code-review (push) Successful in 24s Details CD Pipeline / build-and-deploy (push) Successful in 6m58s Details CD Pipeline / post-deploy-checks (push) Successful in 3m7s Details	2026-04-30 11:02:12 +08:00
Your Name	e27b462bef	fix(ops): keep disabled gitea runner stopped All checks were successful Code Review / ai-code-review (push) Successful in 27s Details	2026-04-30 10:59:46 +08:00
Your Name	0f7e9d3467	fix(cd): run docker builds on host runner All checks were successful CD Pipeline / tests (push) Successful in 1m33s Details Code Review / ai-code-review (push) Successful in 25s Details CD Pipeline / build-and-deploy (push) Successful in 9m20s Details CD Pipeline / post-deploy-checks (push) Successful in 1m33s Details	2026-04-30 10:43:33 +08:00
Your Name	7cc10b2599	fix(cd): serialize gitea docker builds Some checks failed CD Pipeline / build-and-deploy (push) Failing after 40s Details Code Review / ai-code-review (push) Successful in 24s Details	2026-04-30 10:11:50 +08:00
Your Name	e91db52858	docs(logbook): record `639bb64` prod deployment [skip ci]	2026-04-30 09:45:48 +08:00
Your Name	639bb64788	feat(flywheel): surface ai automation and code review Some checks failed Code Review / ai-code-review (push) Successful in 31s Details CD Pipeline / build-and-deploy (push) Failing after 5m23s Details	2026-04-30 00:09:25 +08:00
Your Name	4a57c2d04f	feat(flywheel): expose incident processing timeline All checks were successful CD Pipeline / build-and-deploy (push) Successful in 10m56s Details	2026-04-29 23:38:30 +08:00
Your Name	f5f41543c9	docs: ADR-105 推翻 A2 + LOGBOOK 2026-04-29 LLM 飛輪復活戰 ADR-105 完整記錄推翻 A2 鐵律的決策： - Context: A2 歷史背景 + 2 個月後事實基礎變化（GPU + qwen2.5:7b） - Decision: 4 處修改（IntentType.DIAGNOSE override / chain / openclaw.py task_type / 6 regression test） - Consequences: 正面（飛輪復活）+ 負面（Ollama 單點）+ 已知債（ADR-106-109 後續） - Validation: 部署前 1635 tests 全綠，部署後 5 項驗證指標 - Rollback: env 切換 / git revert LOGBOOK 加 2026-04-29 條目： - 真根因：4 provider 全死 + A2 鐵律排除 Ollama - CD 連環血淚：5 個 commit 全 failure（setup_test_schema.sql 缺欄） - 已落地（不依賴 CD）：Prometheus 17 條 rule + Gemini sanitize - Memory 索引同步更新（指向 project_revert_a2_ollama_primary.md）注意：docs/ 不在 cd.yaml paths trigger，此 commit 不影響 CD。 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-29 20:59:53 +08:00
Your Name	6eb33594c2	docs(logbook): T0 12-Agent 全景驗證紀錄承接前段 session wave2 (commit `143c15f0`) + DB cleanup + Gitea HMAC + ArgoCD/Sentry MCP，派四位專家並行驗證（critic / db-expert / debugger / tool-expert）。詳情：B1/B2 鬼魂按鈕 + KM 早期吞例外 + M1-M4 中度問題 + G1-G3 環境治理 gap。此 commit 主要為 LOGBOOK 索引補齊，本次 P0/P1 修復內容詳見前 2 個 commit。 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-29 10:44:39 +08:00
Your Name	025a493f06	feat(p3.2+adr-100): Model Version Tracker + SLO 自治 + KB rot cleaner Some checks failed run-migration / migrate (push) Failing after 12s Details CD Pipeline / build-and-deploy (push) Has been cancelled Details Wave 8 P3.2 模型版本追蹤 + ADR-100 SLO 自我治理 + 配套： P3.2 — Model Version Tracking: - model_version_probe.py (268 行) — 探測 Ollama / OpenRouter 等 provider 的 model version - model_version_tracker.py (101 行) — 對齊 PG provider_version_history 表 - migrations/p3_2_provider_version_history.sql + rollback — 25 行 schema - db/models.py +32 行 — ProviderVersionHistory ORM ADR-100 — AI 自主化 SLO: - docs/adr/ADR-100-ai-autonomous-slo.md (167 行) — 飛輪 SLO 設計與閾值 - ops/monitoring/slo-rules.yml (254 行) — Prometheus SLO recording rules + alerts - ops/monitoring/tests/test_slo_rules.yaml (242 行) — promtool unit tests 整合修改: - main.py +72 行 — Lifespan 啟動 model_version_probe + KB rot cleaner schedule - gitea_webhook.py +45 行 — webhook 接收 model 版本變化通知 - ci_auto_repair.py / evidence_snapshot.py / pre_decision_investigator.py — 配合接線新測試: - test_kb_rot_cleaner_schedule.py (120 行) — 9 tests pass - test_slo_rules.yaml — promtool 驗收 Tests: 9 passed (test_kb_rot_cleaner_schedule) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: Multiple Engineers (P3.2 + ADR-100) <noreply@anthropic.com>	2026-04-27 14:54:19 +08:00
Your Name	fb130c9a28	feat(p3.1-t2): DiagnosisAggregator stub tests + sanitization 補強 + metrics 補欄 Some checks failed CD Pipeline / build-and-deploy (push) Failing after 2m16s Details Wave 8 P3.1-T2 後續補測 + 配套：新增測試: - test_diagnosis_aggregator_stub.py (238 行) — 15 tests · stub fixture 驗證 _collect_diagnosis_aggregator 接線 · feature flag default off 不呼叫 · timeout 邊界 / exception fail-soft 修改: - core/metrics.py +23 — 新增 DiagnosisAggregator 相關 Prometheus 指標 - sanitization_service.py +24 — 補強 prompt sanitize 邊界（vuln #4 配套） - RUNBOOK-AGENT-STEP-LATENCY.md / agent_step_latency_rules.yaml — 微調 Tests: 15 passed Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-27 08:30:26 +08:00
Your Name	fefe4c21cd	fix(inc-20260425): A1+A2 後續 — Solver/Critic timeout + auto_repair 接線 + Runbook + Grafana Some checks failed CD Pipeline / build-and-deploy (push) Has been cancelled Details 延續 `595629c0` INC-20260425 修復，補三段 Agent + 全鏈路觀測： A1 後續 — Solver/Critic 三段 timeout 接線: - solver_agent.py: AGENT_SOLVER_TIMEOUT_SEC=20.0（env override） - critic_agent.py: AGENT_CRITIC_TIMEOUT_SEC=15.0（env override） - protocol.py: 三 Agent 共用 observe_agent_step() 包裹呼叫 · success/timeout/error outcome label · histogram 寫入 aiops_agent_step_duration_seconds A2 後續 — auto_repair_service 改用 _diagnose_fallback_chain: - auto_repair_service.py +46 行 — 切換 DIAGNOSE 路由到新 chain（NEMO→GEMINI→CLAUDE） - 完全避開 Ollama CPU 238s 二次 timeout 新增 metrics: - core/metrics.py +59 行 — 配合 observe_agent_step 的 histogram bucket + label cardinality 新增測試 (862 行): - test_agent_step_timeouts.py (475) — 三 Agent 各 timeout 邊界 + outcome label - test_ai_router_diagnose_fallback.py (387) — _diagnose_fallback_chain 正確序新增配套: - docs/runbooks/RUNBOOK-AGENT-STEP-LATENCY.md (350) — INC 故障排查 + 觀測指引 - ops/monitoring/grafana/agent_step_latency_rules.yaml (160) · 三 Agent histogram alert rules（p99 > timeout 80% → warning）驗收: 33 tests pass (test_agent_step_timeouts 22 + test_ai_router_diagnose_fallback 11) INC-20260425 雙修總工作量（595629c0 + 此 commit）: · 5 個 service/agent 檔修改 · 1 個新 observability 模組 · 4 個新測試/配套檔 · 1372+187 = 1559 行新增 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Sonnet 4.6 (INC-20260425 後續) <noreply@anthropic.com>	2026-04-27 08:15:53 +08:00
Your Name	1ab6786ce3	feat(ops): Ollama 容災 Runbook + Grafana 儀表板 + Consensus K8s ConfigMap patch Some checks failed run-migration / migrate (push) Failing after 13s Details CD Pipeline / build-and-deploy (push) Failing after 2m1s Details Wave 6 P2.3 ops 配套 + tool-expert 部署文件：新增: - docs/runbooks/RUNBOOK-OLLAMA-FAILOVER.md (240 行) · 三大鐵律驗證步驟（自動切 Gemini / 自動切回 / quota 熔斷） · failover/recovery 完整 SOP · 故障排查清單（Ollama 111/188 不通、Gemini quota 超發等） - ops/monitoring/grafana/dashboards/ollama_failover.json (295 行) · 4 panel：current primary / failover events / quota usage / health status · 對應 P2.3 metrics: OLLAMA_FAILOVER_TRIGGERED_TOTAL / GEMINI_DAILY_CALL_COUNT - k8s/awoooi-prod/04-configmap.yaml.patch-consensus · ENABLE_12AGENT_CONSENSUS / ENABLE_AIOPS_P2_FUSION feature flag patch Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: tool-expert agent (Wave 6) <noreply@anthropic.com>	2026-04-27 08:11:40 +08:00
Your Name	cc547736ab	feat(wave6-8): P2.1 fusion + P2.2 governance + P2.4 consensus + Wave 7/8 BLOCKER 修復承接 Wave 6/7/8 多 engineer 在 agent 限額前完成的代碼，補 commit 解 production HEAD 隱性 import error（decision_fusion 已被 decision_manager 引用但檔案 untracked）。新增（後端核心）: - decision_fusion.py (562 行) — P2.1 方法 III（OpenClaw + Hermes + Elephant 三 LLM 融合） - aiops_timeline.py + aiops_timeline_service.py — critic B4 修復 /api/v1/aiops/timeline endpoint，DB 存取抽到 service 層遵守 leWOOOgo 積木化 - migrations/p2_decision_fusion_columns.sql + rollback — approval_records fusion 欄位修改（後端整合）: - decision_manager.py — fusion 三斷鏈修補（critic B1+B2+B3）： · B1: 寫 _evidence_snapshot_ref 到 token.proposal_data · B2: fusion 前計算 complexity_score 並寫 token · B3: fusion composite 寫 token.proposal_data["decision_fusion"] - auto_approve.py — fusion + consensus 認識（critic B3+B5）： · composite > 0.7 → auto_execute_eligible bypass min_confidence · source=consensus_engine + score>=0.6 → 規則可信路徑 - consensus_engine.py — db-fix _save_consensus 重用 agent_sessions - governance_agent.py — db-fix _alert PG 寫入 ai_governance_events - approval_db.py — fusion 3 欄位 + 2 partial index + CheckConstraint - db/models.py — schema 對齊 migration - core/config.py — vuln #1 修復：OLLAMA_URL/_FALLBACK_URL field_validator 拒絕公網 IP + 外部域名，僅允許私網/loopback/K8s SVC 白名單 - core/feature_flags.py — P2 fusion + consensus flags - main.py — governance_agent lifespan 啟動 - failover_alerter.py — Wave8-X2: in-memory dedup fallback（Redis 拒絕後不 fail-open） - ollama_*.py — metrics 整合 + recovery 改善 - auto_repair_service.py — verifier 接線新增（測試 2438 行）: - test_decision_fusion.py / test_governance_agent.py / test_consensus_integration.py - test_p2_db_fixes.py / test_wave8_fusion_fixes.py - test_config_url_validation.py（vuln #1 12 tests） - test_failover_alerter.py +Wave8-X2 in-memory dedup 補測驗收: 116 tests pass (decision_fusion + wave8_fusion + config_url + consensus + governance + p2_db_fixes + failover_alerter) Conflict resolution: - 3 檔（config.py + auto_approve.py + decision_manager.py）git stash pop 衝突保留 stashed (engineer 最終版)，補回 ValueError 「公網 IP」字樣對齊 test Note: 此 commit 解 production HEAD 隱性 import error 仍未修: vuln #4 prompt injection / debugger B14 quota fail-closed / B25-B26 drain_pending_tasks / B8 governance fail alert Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-Authored-By: Multiple Engineers (Wave 6/7/8) <noreply@anthropic.com>	2026-04-27 08:11:40 +08:00
Your Name	7cd53c0228	fix(monitoring): 記憶體告警改用 working_set，停止 page cache 假告警 - alerts-unified.yml: - SentryClickHouseMemoryPressure: usage_bytes → working_set_bytes，0.8 → 0.85 - GiteaMemoryPressure: 同步修正（同樣 page cache 虛高根因） - ops/monitoring/tests/clickhouse_memory_test.yml: promtool 4 cases - 04-awoooi-devops-commander.md v2.8: Prometheus 指標選擇規範 + Gitea HMAC Webhook 規範 - LOGBOOK: 記錄 T0 五大並行任務（A 按鈕 / B ClickHouse / C Gitea webhook / D ElephantAlpha / F Code review）鐵證: 2026-04-23 23:13 sentry-clickhouse usage_bytes=88.5% vs working_set=7.8% 根因: container_memory_usage_bytes 含 OS page cache，OOM killer 不視為壓力修法: 改用 K8s/cadvisor 認可的 working_set_bytes (RSS + active cache)，閾值 0.85 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-26 20:16:12 +08:00
Your Name	689839cd83	docs(logbook): 記錄 2026-04-25 自動化飛輪四修 + Hermes + qwen3 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 09:49:50 +08:00
Your Name	d467cac709	fix(hermes): 改用 anthropic Python SDK 直呼，棄用需要 claude CLI 的 claude-agent-sdk Some checks failed CD Pipeline / build-and-deploy (push) Has been cancelled Details 根因：claude-agent-sdk 需要 spawn claude CLI，prod pod 沒有 CLI 所以 SDK 回空。修法：改用 anthropic.AsyncAnthropic().messages.create() 直呼 API。 model: claude-haiku-4-5-20251001（快速低成本，適合 Telegram QA） Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 03:08:51 +08:00
Your Name	7d1c85eb86	fix(hermes): ANTHROPIC_API_KEY 注入 + solver 信心度修法 A + 12-Agent 治理文件 Some checks failed CD Pipeline / build-and-deploy (push) Has been cancelled Details - nl_gateway.py: ClaudeAgentOptions 透過 env= 注入 ANTHROPIC_API_KEY（CLAUDE_API_KEY alias），修復 SDK 找不到 API key 的問題（SDK 讀 ANTHROPIC_API_KEY，K8s secret 名稱是 CLAUDE_API_KEY） - solver_agent.py: 修法 A — kubectl_command 欄位優先路徑，OpenClaw Nemo 回傳完整指令時不再被語意合成壓縮 confidence（0.9 → min(0.5) 的 bug），9 tests pass - AGENTS.md: Codex CLI 對應版 CLAUDE.md（Codex Session 啟動用） - docs/12-agent-game-rules.md: 12-Agent 任務判型 + 主責/協作派工 + 9 skills 對照（v1.0） - .agents/skills/06-awoooi-monorepo-master.md: v1.6，新增 12-agent 協作治理章節 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 02:33:43 +08:00
Your Name	86ee013cdf	feat(hermes-complete): Hermes NL 三項補強 + ConsensusEngine + ADR 收尾 All checks were successful CD Pipeline / build-and-deploy (push) Successful in 9m32s Details ## Hermes NL 補強（nl_gateway.py） - T1 hermes_dispatch_log DB 寫入（asyncio.create_task 非阻擋） - T2 Redis 速率限制：per-chat_id 20 req/min，fail-open - T3 Multi-turn session：hermes:session:{chat_id}:{user_id} TTL=300s，最近 3 輪 ## ConsensusEngine（ADR-095 宣告式設計） - consensus_engine.py: CONSENSUS_WEIGHTS class 屬性 security=0.4 鎖定，9 個 Claude Code agent 分配 0.6 - config.py: ENABLE_12AGENT_CONSENSUS=False feature flag ## ADR 狀態 - ADR-093/094/095: Proposed → 🟡 批准實作中 - 各 ADR 加 v1.1 變更紀錄 ## K8s ConfigMap - prod 04-configmap.yaml: 加 3 個 feature flags（均 false） - dev 02-configmap.yaml: 同步加入 ## LOGBOOK - 記錄 WS0–WS6 + 補強完成，feature flags 啟用指引 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 02:22:40 +08:00
Your Name	5675e7c3b0	fix(phase2+aiops): Phase 2 Agent timeout + AI Router intent hint + signoz incident_id ## Phase 2 Agent timeout（防止單步 LLM 拖垮整場辯證） - critic_agent.py: asyncio.wait_for + PHASE2_STEP_TIMEOUT_SEC=20s - diagnostician_agent.py: 同等超時保護 - solver_agent.py: 同等超時保護 ## AI Router 優化 - ai_router.py: _resolve_intent_from_context() Phase 2 agents 傳 intent_hint → Router 快路徑，不重跑 intent LLM ## SignOz Webhook 修復 - signoz_webhook.py: incident_id 補傳 send_approval_card()（移除 TODO 2026-04-05） ## Alert 處理流程修復 - webhooks.py: _should_bypass_alertmanager_llm() Host 類 NO_ACTION 告警直接走人工排查卡片，不再誤觸 LLM Agent Debate - incident_repository.py: update_incident_status 加 resolved_at 參數 - incident_service.py / proposal_service.py / incident_approval_service.py: 小修 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 02:10:06 +08:00
Your Name	054d0ae422	docs(ws0): Hermes × 12-Agent Telegram 整合治理文件（ADR-093/094/095） ## 新增 - ADR-093: Telegram 告警全面遷移至 SRE 戰情室群組 - 混合策略 allowlist 模式（TYPE-3/4/4D/8M → 群組 + user_id binding） - nonce 新格式 apr:{short_id}:{action}:{user_id_hash} + Redis 後端映射 - Feature flag TG_GROUP_CUTOVER 灰階切流 - ADR-094: Hermes 自然語言介面（@mention 對話） - Option C：單 bot + Claude Agent SDK 虛擬分派 - Webhook secret_token + allowed_updates = [message, callback_query, chat_member] - Prompt Injection 防護：query/describe/summarize only，mutate 走 ApprovalRecord - Redis session TTL=300s + turn>=5 壓縮 - ADR-095: 12-Agent Claude SDK 整合 × Telegram 視覺分派 - 12 位 agent 完整 emoji/hashtag/handle 表格 - ConsensusEngine weights 擴充（security=0.4 鎖定） - display_names.py 命名隔離（.claude/agents/ vs src/agents/） ## 更新 - ADR-009: 加 v0.3 變更紀錄指向 ADR-095 - ADR-075: 加更新引用表（ADR-093 D4 allowlist 子條款、ADR-094/095） - docs/design/hermes-telegram-flows/hermes-flows.html: F1-F7 完整流程圖 ## Pre-Flight 確認 - approval_records 表尚不存在 → 將用 BIGINT 全新建立 - docker-compose.yml:78 明碼 token 🔴 P0 待 WS1 修復 - awoooi_migrator 角色尚未建立 → WS2 建立 - claude-agent-sdk 升至 0.1.66（最新） Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-25 02:10:06 +08:00

1 2 3 4 5 ...

389 Commits