data privacy / Python

piiscope

Find, score, and remediate personal data in files and databases.

piiscope scans tabular data and free text for direct identifiers, then measures privacy risk across quasi-identifiers. It combines practical detection with k-anonymity, l-diversity, t-closeness, and column-level remediation workflows.

Installation

pip install piiscope

Platform requirements and configuration are documented in the repository instructions.

Capabilities

  • Scan CSV, JSON, JSONL, Parquet, text files, directories, and supported database sources
  • Map findings to GDPR, CCPA, KVKK, and LGPD categories
  • Measure k-anonymity, l-diversity, and t-closeness on quasi-identifiers
  • Generate machine-readable reports and apply column-level remediation

Use cases

  • Checking datasets before sharing, analysis, or model training
  • Adding a repeatable personal-data check to local and CI workflows
  • Investigating privacy risk beyond simple regex matches

Evidence

published benchmark and repository fixture

All 512 structured benchmark cases pass before a clean remediation rescan.

The public synthetic benchmark covers 32 multilingual positive and hard-negative families across 512 rows. The committed customers fixture separately exercises the full remediation round trip.

benchmark 512 / 512 exact
precision 1.000
recall    1.000
rescan    0 / low

GitHub Actions

Run a privacy gate and emit SARIF without matched values.

- uses: barissozudogru/[email protected]

Limitations

  • Detection is heuristic and should be reviewed against the context of the dataset
  • A clean scan is not proof of legal or regulatory compliance
  • Optional Parquet and NLP features require their matching extras