Data Governance Guide

This guide walks through a complete data governance workflow using metapod’s quality, profiling, compliance, and cost probes together.

The probes are designed to chain: the output of one probe feeds into the next, building a complete picture of your data estate.

The governance pipeline

Step 1: Profile   (ydata)      What data do I have?
Step 2: Quality   (soda)       Is the data correct?
Step 3: PII scan  (presidio)   Does it contain personal data?
Step 4: Classify  (presidio)   What sensitivity tier?
Step 5: Contract  (schema)     Does the schema match expectations?
Step 6: Access    (audit)      Who can access sensitive data?
Step 7: Retention (retention)  Is data kept within legal limits?
Step 8: Cost      (cost)       How much does storage cost?

Each step produces a JSON file in ./output/ that the next step can consume. You can run all steps or just the ones you need.

Prerequisites

# Install metapod with all governance extras
pip install 'metapod-probes[soda,ydata,presidio]'

# Or from source
pip install -e '.[soda,ydata,presidio]'

# Download spaCy model for PII detection (once)
python -m spacy download en_core_web_lg

Step 1: Profile your data

YData Profiling generates per-column statistics: missing values, distinct counts, distributions, correlations.

# config_ydata.yaml
source:
  type: ydata
  host: csv                        # csv, parquet, sql, or blob
  schema_name: ./data/transactions.csv
  port: 5000                       # sample 5000 rows
metapod run "ydata.profiling.profile" -c config_ydata.yaml

Output (per column):

{
  "type": "column",
  "name": "amount",
  "dtype": "Numeric",
  "missing_count": 0,
  "missing_pct": 0.0,
  "distinct_count": 4523,
  "mean": 245.67,
  "std": 189.34,
  "min": 0.01,
  "max": 9999.99,
  "p25": 89.50,
  "p50": 198.00,
  "p75": 350.25
}

Step 2: Detect anomalies

Compare the current profile against a baseline (previous run) to flag unexpected changes.

# First run: save as baseline
cp output/ydata/profiling/profile.json baseline_profile.json

# Later runs: compare against baseline
metapod run "ydata.profiling.anomalies" \
  -s output/ydata/profiling/profile.json \
  --config config_ydata_anomalies.yaml
# config_ydata_anomalies.yaml
source:
  type: ydata
  host: csv
  schema_name: output/ydata/profiling/profile.json
  service_name: baseline_profile.json   # baseline path

Output:

{
  "column": "amount",
  "check": "mean_shift",
  "severity": "alert",
  "value": 512.30,
  "baseline": 245.67,
  "z_score": 3.42,
  "message": "Mean shifted from 245.67 to 512.30 (z=3.4)"
}

Step 3: Run quality checks

Soda Core validates business rules using the SodaCL language.

# checks/transactions.yml (SodaCL format)
checks for TRANSACTIONS:
  - row_count > 0
  - missing_count(amount) = 0
  - duplicate_count(transaction_id) = 0
  - avg(amount) between 10 and 10000
  - freshness(created_at) < 2d
# config_soda.yaml
source:
  type: soda
  host: my_postgres            # data source name in soda config
  schema_name: checks/transactions.yml
metapod run "soda.quality.checks" -c config_soda.yaml

If Soda runs as a separate step (CI/CD, Databricks), import results:

# Soda runs independently
soda scan -d my_postgres -c soda/ checks/ --output-file results.json

# metapod imports
metapod run "soda.quality.scan_results" -s results.json

Step 4: Detect PII

Presidio scans text columns for personal data. ~75 built-in entity types cover Financial Services, Healthcare, Telco, and Public Admin.

# config_presidio.yaml
source:
  type: presidio
  host: it                     # language: it, en, de, es, fr
  schema_name: ./data/transactions.csv
  port: 1000                   # sample rows
metapod run "presidio.compliance.pii_scan" -c config_presidio.yaml

Output (per column):

{
  "column": "customer_name",
  "pii_detected": true,
  "rows_with_pii": 987,
  "pii_pct": 98.7,
  "entity_types": [
    {"entity": "PERSON", "count": 987},
    {"entity": "IT_FISCAL_CODE", "count": 42}
  ],
  "risk_level": "high"
}

Custom PII patterns

Define your own patterns without modifying code:

# custom_patterns.yml
patterns:
  - entity: INTERNAL_CUSTOMER_ID
    name: "Internal Customer ID"
    regex: '\bCUST\d{8}\b'
    score: 0.8
    context: ["customer", "cliente", "id"]
    sensitivity: CONFIDENTIAL
# config_presidio.yaml
source:
  type: presidio
  host: it
  schema_name: ./data/transactions.csv
  service_name: custom_patterns.yml   # path to custom patterns

Step 5: Classify sensitivity

Builds on the PII scan to assign each column a sensitivity tier.

metapod run "presidio.compliance.sensitivity_report" \
  -s output/presidio/compliance/pii_scan.json

Tiers: PUBLIC < INTERNAL < CONFIDENTIAL < RESTRICTED

{
  "column": "customer_name",
  "sensitivity": "CONFIDENTIAL",
  "entities": ["PERSON"],
  "pii_pct": 98.7
}

Step 6: Validate schema contract

Define what your schema SHOULD look like:

# contract.yml
contracts:
  - dataset: TRANSACTIONS
    columns:
      - name: transaction_id
        type: integer
        nullable: false
      - name: amount
        type: decimal
        nullable: false
      - name: created_at
        type: timestamp
    rules:
      min_columns: 3
      max_columns: 50
metapod run "schema_contract.quality.validate" -s contract.yml

Output:

{"column": "transaction_id", "status": "ok"}
{"column": "amount", "status": "ok"}
{"column": "new_column", "status": "extra", "message": "Not in contract"}

Detect schema drift between runs:

metapod run "schema_contract.quality.drift" -s current_schema.json

Step 7: Audit access control

Cross-reference PII findings with RBAC role assignments to identify who has access to sensitive data.

# config_access.yaml
source:
  type: access_audit
  schema_name: output/presidio/compliance/pii_scan.json
  service_name: output/azure_rbac/governance/role_assignments.json
metapod run "access_audit.compliance.pii_access" -c config_access.yaml

Output:

{
  "column": "customer_pan",
  "sensitivity": "RESTRICTED",
  "users_with_access": 12,
  "risk": "excessive_access",
  "message": "12 users have access to RESTRICTED data (threshold: 5)"
}

Step 8: Check retention compliance

Verify that data is not kept beyond legal retention limits.

# retention_policy.yml
policies:
  - dataset: "*"
    sensitivity: RESTRICTED
    max_age_days: 2555        # 7 years (GDPR financial data)
  - dataset: "logs_*"
    max_age_days: 365         # 1 year
  - dataset: "*"
    sensitivity: CONFIDENTIAL
    max_age_days: 3650        # 10 years
metapod run "retention.compliance.check_retention" -c config_retention.yaml

Output:

{
  "dataset": "transactions_2018",
  "sensitivity": "RESTRICTED",
  "oldest_record": "2018-01-15",
  "age_days": 2997,
  "max_age_days": 2555,
  "status": "expired",
  "message": "Dataset exceeds retention by 442 days"
}

Step 9: Estimate storage costs

metapod run "cost.cost.storage_cost" \
  -s output/adls_gen2/metadata/paths.json

Output:

{
  "directory": "raw/transactions",
  "total_bytes": 15234567890,
  "total_gb": 14.19,
  "tier": "hot",
  "estimated_monthly_eur": 0.26
}

Running all steps together

# Full governance pipeline
metapod run "ydata.profiling.*" -c config_ydata.yaml
metapod run "soda.quality.*" -c config_soda.yaml
metapod run "presidio.compliance.*" -c config_presidio.yaml
metapod run "schema_contract.quality.*" -s contract.yml
metapod run "access_audit.compliance.*" -c config_access.yaml
metapod run "retention.compliance.*" -c config_retention.yaml
metapod run "cost.cost.*" -s output/adls_gen2/metadata/paths.json

# Or use parallel execution
metapod run "presidio.compliance.*" -c config_presidio.yaml --parallel

All results are saved in ./output/ as JSON files. Push them to a catalog platform or process them with your own scripts.

Integration with azure-monitoring

The governance probes feed into the azure-monitoring system’s check engine. When a probe detects an anomaly (quality check failed, PII in unexpected column, schema drift, expired retention), it produces a CheckResult that flows through the standard alerting pipeline:

metapod probe -> CheckResult -> WorkItemTracker -> Azure DevOps Board
                                                -> Email notification
                                                -> JSON for blob storage

See the azure-monitoring wiki for the full data observability architecture.