跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

pmstudio-dr

Generate a Disaster Recovery plan with RTO/RPO targets. Use when someone asks to "create a DR plan", "disaster recovery", "business continuity", or "RTO RPO". Do NOT use for step-by-step restoration runbooks — use /pmstudio-recovery instead. Designed for SaaS platform products where Coco Inc is the customer — focuses on service continuity, data recovery, and vendor dependency management rather than infrastructure rebuild.

运维与云511skills/dr-plan/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/coco-research/coco/dr-plan/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

DR Plan — Disaster Recovery Plan

Purpose

Generates a Disaster Recovery plan scoped to a specific product. For SaaS products (where the vendor owns infrastructure), this plan focuses on what Coco Inc controls: data exports, integration failover, access recovery, communication, and business continuity.

Process

Step 1: Read Context

Read all that exist:

  • CLAUDE.local.md — architecture, integrations, vendor info, data strategy
  • PRD/*.html or PRD/*.md — NFRs, technical considerations, integrations, data architecture
  • Architecture/ — system diagrams, data flows
  • Operational/IRP-*.html — incident response plan (if exists, reference for communication)

Step 2: Determine Product Type

From context, classify the product:

TypeDR FocusExample
SaaS (customer)Vendor dependency, data portability, alternative workflows(third-party SaaS)
Self-hostedInfrastructure recovery, backup/restore, failoverCustom app on EKS
HybridBoth vendor and self-hosted componentsSaaS + custom middleware

Adjust the plan scope accordingly. For SaaS, skip infrastructure sections (that's the vendor's problem). Add vendor SLA and data portability sections.

Step 3: Ask Discovery Questions

Only ask what can't be inferred:

  1. Business criticality tier? (Tier 1 critical / Tier 2 important / Tier 3 standard)
  2. What's the maximum tolerable downtime? (This becomes RTO)
  3. What's the maximum tolerable data loss? (This becomes RPO)
  4. Are there manual workarounds if the product is down? (e.g., "we can use Excel for 48 hours")
  5. Vendor SLA? (uptime commitment, support response times, data export capabilities)
  6. Backup strategy? (Does data export to Snowflake/SharePoint? How often?)

Step 4: Generate DR Plan

Output: Operational/DR-Plan-{ProductName}-{Date}.html

Self-contained HTML with clean typography and print CSS.

11 sections:

1. Purpose & Scope

  • Product name, ProdID, instances covered
  • What this plan covers and doesn't cover
  • Relationship to vendor's own DR plan

2. Service Classification

  • Business criticality tier with justification
  • Data classification (Purple/Red/Yellow/Green)
  • Regulatory or compliance requirements affecting recovery
  • Business impact of downtime (per hour/day)

3. RTO/RPO Targets

Per-component target table:

ComponentRTORPOJustification
Core platform4 hours24 hoursVendor SLA: 99.9% uptime
Data in Snowflake2 hours0 (real-time sync)Analytics feeds downstream
Integrations (SSO)1 hourN/AUsers locked out
Integrations (ServiceNow)8 hoursN/ATicket creation manual fallback

4. Disaster Scenarios

Ranked by likelihood x impact:

  1. Vendor platform outage (most likely) — vendor is down, product inaccessible
  2. Data corruption — bad data entered or sync error corrupts records
  3. Access loss — SSO failure, license revocation, permission misconfiguration
  4. Integration failure — upstream/downstream system breaks connection
  5. Security breach — unauthorized access, data exfiltration
  6. Vendor business failure — vendor acquired, shut down, or ends product

For each scenario: description, likelihood, impact, detection method.

5. Recovery Procedures

Per-scenario step-by-step. See references/saas-dr-patterns.md for SaaS-specific recovery patterns.

Structure per scenario:

  • Trigger condition (how do we know this is happening)
  • Immediate actions (first 15 minutes)
  • Short-term recovery (first 4 hours)
  • Full recovery (to meet RTO)
  • Verification checklist

6. Communication Plan

  • Who to notify per scenario and severity
  • Templates (reference IRP if it exists, or provide standalone)
  • Vendor communication (support ticket + account manager)
  • Stakeholder updates cadence during outage

7. Dependencies

  • External systems required for recovery
  • Vendor contacts (support, account manager, escalation)
  • SLA commitments and how to invoke them
  • Third-party services (SSO provider, email, etc.)

8. Data Backup & Restoration

  • What data is backed up (by vendor, by Coco Inc)
  • Backup frequency and retention
  • Where backups are stored
  • Restoration procedure (step-by-step)
  • Data validation after restoration

9. Manual Workarounds

  • What business processes can continue without the product
  • Excel/SharePoint fallback procedures
  • Duration the workaround is sustainable
  • Data reconciliation procedure when product is restored

10. Testing Schedule

  • DR test cadence (annually minimum, semi-annually recommended)
  • Test scenarios (tabletop, partial recovery, full recovery)
  • Last test date and results
  • Next scheduled test
  • Test success criteria

11. Review & Maintenance

  • Review cadence: semi-annually or after any DR event
  • Approval chain (PM, Security POC, Sponsor)
  • Version history
  • Distribution list

Step 5: Present for Review

Show the plan outline with key decisions highlighted:

  • RTO/RPO targets
  • Scenario prioritization
  • Manual workaround availability

Ask for approval before writing.

Critical Rules

  1. RTO/RPO must be numbers, not words. "As fast as possible" is not an RTO. Push for specific hours.
  2. Vendor DR is not your DR. The vendor's uptime SLA is an input to your plan, not a substitute for it. Your plan covers what happens on Coco Inc's side.
  3. Manual workarounds are critical. For every scenario, document what the business does while the product is down. If there's no workaround, that's a risk to flag.
  4. Data portability is DR. If you can't get data out of the product, you can't recover from vendor failure. Document export capabilities.
  5. Test the plan. A plan that's never been tested is a wish list. Include a testing schedule and success criteria.

相似的 Skill

internal-comms
anthropics/skills180k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

运维与云

shipping-and-launch
addyosmani/agent-skills104k

shipping-and-launch

Prepares production launches. Use when preparing to deploy to production, or when asking what needs to be in place before shipping. Use when you need a pre-launch checklist, when setting up monitoring, when planning a staged rollout, or when you need a rollback strategy.

运维与云

observability-and-instrumentation
addyosmani/agent-skills104k

observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the available data.

运维与云

internal-comms
ComposioHQ/awesome-claude-skills77k

internal-comms

A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use. Claude should use this skill whenever asked to write some sort of internal communications (status reports, leadership updates, 3P updates, company newsletters, FAQs, incident reports, project updates, etc.).

运维与云

vercel-react-best-practices
CherryHQ/cherry-studio52k

vercel-react-best-practices

React and Next.js performance optimization guidelines from Vercel Engineering. This skill should be used when writing, reviewing, or refactoring React/Next.js code to ensure optimal performance patterns. Triggers on tasks involving React components, Next.js pages, data fetching, bundle optimization, or performance improvements.

运维与云

acceptance-orchestrator
sickn33/agentic-awesome-skills47k

acceptance-orchestrator

Use when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and acceptance verification with minimal human re-intervention.

运维与云