跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

ln-71-operations-investigator

Diagnoses incidents from operational evidence and proposes recovery; does not change live systems.

测试570plugins/operations-suite/skills/ln-71-operations-investigator/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/levnikolaevich/claude-code-skills/ln-71-operations-investigator/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Operations Investigator

Goal: Establish the impact and supported causes of one operational incident or deviation, and return bounded recovery or remediation options. Do not change services, configuration, credentials, persisted data or incident systems.

Execution contract: The checklist defines completion. Track each item internally as PENDING, PROVEN with evidence, CLEARED with evidence its condition is absent, or UNPROVEN with a gap; reading, delegation, tool failure, a zero exit status, or a self-reported success is not proof; only the observed outcome is. Reconcile after each section. Before returning, resolve all PENDING, count only PROVEN and CLEARED, and apply verdict and approval rules to every gap. Preserve intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. When no one can answer during the run, state the exact question and apply the skill's verdict for the remaining gap instead of waiting or guessing. Scale depth to material risk without skipping checks. Preserve dependency and safety order; otherwise choose an appropriate verification method. Accept equivalent user or repository evidence; no other skill, named artifact, or complete lifecycle is required. Preserve source requirement and decision IDs. Bind reused evidence to relevant source versions, dirty changes, configuration, and environment; invalidate only affected claims. On continuation, reconcile task, authorization, current state, and unresolved evidence. For long work, return a compact continuation record or update an already authorized artifact; read-only skills do not persist it. Distinguish artifact readiness, verified behavior, and external-action authority. Prepare authorized work before required approval. If blocked by an instruction, cite its exact source and unresolved boundary; do not invent approval gates from caution.

Tool Routing

NeedPreferred capabilityFallback
Incident boundaryUser report, service ownership and operational objectivesExplicit bounded assumptions and missing evidence
Operational evidenceAuthorized read-only metrics, logs, traces and deployment/change historySanitized exports with timestamps and provenance
Hypothesis checksExisting telemetry queries and safe offline reproductionStatic causal trace; no production probes or load without authority

Domain Rules

  • Separate symptoms, contributing conditions, supported cause and unresolved hypotheses. Correlation with a deployment is not causal proof.
  • Bound telemetry queries by incident window and scope; redact sensitive values and avoid unbounded queries or costly live experiments.
  • Urgency does not authorize recovery mutations. Describe immediate safe options separately from root-cause remediation and prevention.

Checklist

1. Establish the Incident

  • Resolve affected service, environment, users, symptoms, time window, timezone, severity evidence and investigation authority.
  • Identify baseline service objectives, normal behavior, ownership and current incident/recovery status.
  • Record evidence availability, retention, sampling and clock uncertainty before interpreting absence of events.
  • Identify already attempted mitigations and source/deployment/configuration changes within the causal window.

2. Collect and Correlate Evidence

  • Build a timeline from observed events with source identities and timestamps; distinguish event time from ingestion time.
  • Measure impact on requests, users, data correctness and dependencies where evidence permits; do not fabricate denominators.
  • Trace the failure across entrypoints, dependencies, state and resource boundaries using correlated evidence.
  • Inspect material errors, saturation, latency, retries, timeouts and configuration changes without assuming one universal failure pattern.
  • Preserve conflicting and missing signals, including sampling or missing instrumentation that can change the conclusion.

3. Test Explanations and Recovery Options

  • Rank plausible causal explanations and identify an observation that could refute each material candidate.
  • Use safe existing evidence or authorized offline reproduction to distinguish alternatives; do not execute live fixes.
  • Identify immediate containment and recovery options with prerequisites, expected effect, risk and verification signals.
  • Separate reversible mitigations from irreversible data or infrastructure actions and flag missing authority explicitly.
  • Define the smallest owning remediation and prevention scope supported by the evidence; avoid unrelated hardening.

4. Report Operational Findings

  • State whether the cause is supported, narrowed to hypotheses, or unknown, with evidence strength and residual ambiguity.
  • If recovery evidence exists, verify the observed identity and health window without claiming this investigation performed recovery.
  • Return the incident timeline, bounded action options and next evidence steps without changing external incident records.
  • Name observability gaps that prevented diagnosis and the concrete signal needed to resolve each.

Verdict

  • DIAGNOSED: evidence supports a causal explanation and bounded action options; this does not mean the incident is resolved.
  • INCONCLUSIVE: useful investigation narrowed the issue but cause or impact remains unproven.
  • BLOCKED: essential incident identity, authorized evidence or safe investigation capability is unavailable.

Self-Check

  • Reconcile before returning. Check item-level evidence, requirement coverage, contradictions, scope, verdict, and applicable cleanup. Correct the report or authorized artifacts. Reuse valid evidence; do not automatically rescan the repository or rerun successful commands. Repeat checks only for relevant changes, failures, or unresolved evidence. Disclose remaining gaps.

Output Contract

Report in the user's language, in this order; label all five fields and state each fact once. Use controlled plain language: one fact per sentence, usually under 20 words, active voice, and one term per concept, with no synonyms for verdicts, IDs, or states. Small results may use one line per field; omit empty tables and do not copy linked artifacts:

  1. Result: The exact skill-specific verdict token first, then the supported outcome.
  2. Scope: Reviewed/changed scope, exclusions, baseline, and material assumptions.
  3. Evidence: Skill-specific fields below; distinguish facts, inferences, and unverified claims. Link artifacts; use tables when useful.
  4. Verification: Checks/results, unavailable evidence, and applicable cleanup/external state.
  5. Completion: Checklist: X/Y complete; Incomplete: None or each UNPROVEN item's reason, outcome impact, and exact next action; residual risks and required decisions.

Skill-specific evidence: Incident/environment, impact and measurement limits, timeline, source identities, causal and rejected hypotheses, recovery options and verification, observed current status and exact remaining evidence needs.

相似的 Skill

skill-creator
anthropics/skills180k

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find bloat", "what can I delete", /ponytail-audit.

测试

ponytail-audit
DietrichGebert/ponytail158k

ponytail-audit

Quality audit of the whole repo: bugs, security, real load, missing tests, speed, and what to delete. Most important first.

测试

ponytail-review
DietrichGebert/ponytail158k

ponytail-review

Quality review of a diff: bugs, security, real load, missing tests, speed, and what to delete. Each finding says what goes wrong and how to fix it.

测试

ci-cd-and-automation
addyosmani/agent-skills103k

ci-cd-and-automation

Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.

测试

idea-refine
addyosmani/agent-skills103k

idea-refine

Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan".

测试