跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

data-governance

Establishes ownership, definitions, quality, access, and lineage for the organization's data. Use this when metrics disagree between teams, when nobody knows which dataset is authoritative, when setting up data ownership or access policy, when data quality is unreliable, or before opening a dataset to a wider audience.

数据库与数据2kplugins/data-analytics/skills/data-governance/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/cbrock84/headcount/data-governance/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Data governance

Governance has a reputation for bureaucracy because it is usually implemented as approval queues. Done properly it is the opposite: it makes data usable without asking anyone.

Start with definitions, not policy

The highest-value governance artifact is a metric dictionary. For each business metric:

  • The plain-language definition — what it counts, and what it deliberately excludes.
  • The computation, unambiguously: source table, filters, time grain, timezone.
  • The owner — a person who decides when it is disputed.
  • Known caveats — when it is misleading, and what changed historically.

Most metric disputes dissolve once both parties read the same definition and discover they were measuring different things. Almost none require a policy.

Watch the ones that look obvious. "Active customer," "revenue," and "signup" each have half a dozen defensible definitions, and the ambiguity surfaces at the worst moment.

Ownership

Every dataset has a named owner accountable for its quality and access — a person, not a team. Unowned datasets decay, and nobody notices until a decision is made on stale data.

The owner should sit with the business meaning, not with the pipeline. The team that generates the data understands what it means; the platform team understands how it moves.

Quality, measured rather than asserted

Test data like code, continuously, and alert on failures:

  • Freshness — did it arrive when expected?
  • Volume — is the row count within its normal range? A silent drop to zero is the classic failure.
  • Uniqueness and nullity on key fields.
  • Referential integrity across joins.
  • Distribution — has the shape shifted in a way nothing explains?

The point is finding breakage before a decision is made on it. A pipeline that fails loudly is better than one that silently produces yesterday's numbers.

Access

Default to open for internal, non-personal data. Restrictive-by-default drives the shadow spreadsheet layer, which is genuinely less safe than a governed warehouse.

Personal, financial, and regulated data are the exception: least privilege, purpose stated, reviewed periodically, with Legal & Risk involved on anything with a lawful-basis question.

Lineage

Know where a number came from and what feeds it. Without lineage, you cannot answer the two questions that matter during an incident: what broke upstream, and what downstream is now wrong.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Let two systems each claim to be the source of truth for the same fact.
  • Fix a data-quality issue in a dashboard. Fix it upstream or it recurs in every other consumer.
  • Retire a dataset because it looks unused — you cannot see every consumer. Deprecate, announce, then remove.

相似的 Skill

xlsx
anthropics/skills180k

xlsx

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like "the xlsx in my downloads") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

数据库与数据

deprecation-and-migration
addyosmani/agent-skills103k

deprecation-and-migration

Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to maintain or sunset existing code.

数据库与数据

host-observer
thedotmack/claude-mem98k

host-observer

Use this when fulfilling claude-mem observer jobs on Grok Bot: reply only skip_summary or one full observation XML, never prose.

数据库与数据

babysit
thedotmack/claude-mem98k

babysit

Watch a pull request or review cycle until it is ready to merge. Use when asked to babysit, monitor, or keep checking PR comments, reviews, and CI until all actionable issues are resolved.

数据库与数据

mem-search
thedotmack/claude-mem98k

mem-search

Search claude-mem's persistent cross-session memory database. Use when user asks "did we already solve this?", "how did we do X last time?", or needs work from previous sessions.

数据库与数据

Agent Cost Report
thedotmack/claude-mem98k

Agent Cost Report

Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.

数据库与数据