跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

dataset-quality-audit

Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and actionable suggestions. Triggered when users ask to check data quality, find missing or duplicate values, detect outliers, validate formats, profile data, or clean data.

数据库与数据4.7kskills/dataset-quality-audit/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/zebbern/claude-code-guide/dataset-quality-audit/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

dataset-quality-audit

A data quality auditing tool that runs 12-dimension quality checks on tabular data, producing per-dimension scores (0–100), an overall grade, and actionable fix suggestions.

Capabilities

DimensionDescription
Missing ValuesCount and percentage of null/NaN values per column
Duplicate RowsNumber and percentage of fully duplicated rows
Type ConsistencyMixed types within a single column (e.g., numbers mixed with text)
Value Range / OutliersOutlier detection using the IQR method
Format ComplianceConsistency of date, email, phone number, and other formatted fields
Uniqueness ConstraintsWhether ID-type columns contain duplicates
Whitespace IssuesLeading/trailing spaces, empty strings, whitespace-only values
Constant ColumnsColumns with only a single unique value (zero information)
Distribution SkewnessWhether numeric columns have excessive skewness
Column NamingSpaces, special characters, or inconsistent casing in column names
Cardinality AnomaliesUnusually high or low number of unique values
Cross-Column ConsistencyLogical checks across columns (e.g., start date before end date)

Quick Start

# Basic quality check
python3 scripts/data_quality_checker.py data.csv

# Save report as JSON
python3 scripts/data_quality_checker.py data.csv --output report.json

# Specify ID columns (for uniqueness checks)
python3 scripts/data_quality_checker.py users.csv --id-columns "user_id,email"

# Specify date columns (for format checks)
python3 scripts/data_quality_checker.py orders.csv --date-columns "created_at,updated_at"

Detailed Usage

Basic Invocation

python3 scripts/data_quality_checker.py <data-file> [options]

Parameters

ParameterShortRequiredDefaultDescription
input—Yes—Path to input file (CSV/TSV/Excel/JSON)
--output-oNostdoutPath for the JSON report output
--id-columns-idNoAuto-detectComma-separated column names that should be unique
--date-columns-dcNoAuto-detectComma-separated column names containing dates
--sample-sNoAll rowsNumber of rows to sample (useful for large files)
--encoding-eNoutf-8File encoding

Output Format (JSON)

{
  "file": "data.csv",
  "rows": 10000,
  "columns": 15,
  "overall_score": 78.5,
  "grade": "B",
  "dimensions": {
    "missing_values": {
      "score": 85.0,
      "issues": [
        {"column": "age", "missing_count": 150, "missing_pct": 1.5, "suggestion": "Fill with median or mode"}
      ]
    },
    "duplicates": {
      "score": 95.0,
      "issues": [...]
    }
  },
  "top_suggestions": [
    "Column 'age' has 1.5% missing values — consider filling with the median",
    "Found 200 fully duplicated rows — consider deduplication"
  ]
}

Grading Scale

GradeScore RangeMeaning
A+95–100Excellent quality — ready for use as-is
A90–95Good quality — minor issues only
B80–90Moderate quality — recommended to fix before use
C60–80Poor quality — significant cleaning required
D40–60Very poor quality — many issues need attention
F0–40Essentially unusable — requires re-collection or major cleanup

Dependencies

  • Python 3.8+
  • pandas
  • numpy
pip install pandas numpy

相似的 Skill

xlsx
anthropics/skills180k

xlsx

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like "the xlsx in my downloads") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

数据库与数据

deprecation-and-migration
addyosmani/agent-skills103k

deprecation-and-migration

Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to maintain or sunset existing code.

数据库与数据

claude-mem-install
thedotmack/claude-mem99k

claude-mem-install

Use this when setting up claude-mem on Grok Bot: local worker plus CMEM Pro observer (default), optional host-login observer, or remote cmem.ai. No Cursor required.

数据库与数据

do
thedotmack/claude-mem99k

do

暂无描述

数据库与数据

Agent Cost Report
thedotmack/claude-mem99k

Agent Cost Report

Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.

数据库与数据

Agent Cost Report
thedotmack/claude-mem99k

Agent Cost Report

Believable agent cost report for any period, default the last 7 full days PT, not counting today. Measured tokens from Claude Code transcripts priced at OpenRouter list prices (ESTIMATED), measured provider spend when a sanctioned source exists, note-taker cost separate, Timing-style HTML/PDF plus report.json, line-items.csv, evidence.json.

数据库与数据