Data inventory & classification
You can’t govern what you can’t see. We help you build a living inventory of the personal and sensitive data you hold, and a classification scheme (public, internal, PII, special category, AI-generated) with clear criteria, so every record carries a class and that class drives how it’s protected.
Data flow mapping & lineage
We map how data actually moves through your systems, databases, APIs, logs, third-party tools and backups, and assess how far you can trace it: table-level (“a system touched this”) versus field-level (“this column carries this obligation and propagates here”), so you can answer an auditor and a deletion request with confidence.
Roles, ownership & accountability
We help you define who is accountable for each data domain, how decisions get made and escalated, and how governance connects to onboarding new tools and features, so data ownership is a named responsibility rather than everyone’s problem and no one’s job.
Retention & deletion by policy
We help you design retention tied to class: each category gets a defined lifespan, after which data is deleted or archived automatically, including across replicas and backups. This closes minimisation and the right to erasure at the process level and shrinks the blast radius of any future breach.
Access governance & least privilege
We review who can reach which data and systems, and assess it for excess against least-privilege: standing privilege, orphaned accounts, vendor access that outlived its contract, over-broad access to regulated data. You get a map of over-permissioned access with priorities and a path to least privilege.
Data quality & policy framework
Good governance needs a backbone of documentation. We help you build the policy framework (data governance policy, classification standard, retention schedule, access rules) as one coherent system, each document mapped to a real requirement, with an owner, a version and evidence it’s actually applied.
Shadow data & proliferation control
Copies of personal data sprawl into shadow data, caches, replicas, backups and orphaned resources that live outside your policies: unclassified, uncovered by retention, invisible to a data subject request. We help you find where data proliferates and bring those blind spots back under governance.
Training data provenance & registry
We trace the origin of every dataset in your training pipeline: where it came from, what rights apply, whether consent or a licence exists, and build a dataset-labelling scheme and source registry as an auditable chain of custody, in the format the transparency duties expect for GPAI providers.
Synthetic data pipelines & validation
We help you design a synthetic-data pipeline so teams work without touching real PII, then validate the output from both sides: that it doesn’t lead back to real people (privacy) and doesn’t reproduce protected content (IP). Generating synthetic data from real data is itself processing, so we build it privacy-by-design.