Data Classification That Drives Controls

Four labels sit at the top of the policy. Public, internal, confidential, restricted. Every system has one, the spreadsheet is up to date, and the access rules on those systems are identical. That programme has a taxonomy rather than a control.

A data classification scheme earns its cost at the point where a label changes what happens to a dataset. Until then it is documentation, and expensive documentation at that.

Data classification by impact rather than sensitivity

The standard worth borrowing from comes from the United States federal system. FIPS 199 categorises information and information systems by the potential adverse impact of losing confidentiality, integrity or availability.

Each objective gets rated separately as low, moderate or high, and the security category takes the form of a triplet. A payroll extract might carry high confidentiality, moderate integrity and low availability. A public status page might carry no applicable confidentiality rating, moderate integrity and high availability.

One data classification label cannot hold that. Calling both datasets confidential tells an engineer nothing about which one has to survive an outage.

The high-water mark, and the trap inside it

FIPS 199 then aggregates. Where a system holds several information types, the value assigned to each objective is the highest among those determined for the resident types. One high rating anywhere pulls the whole system up.

That rule is deliberate and it creates a predictable problem. Mixing one high-impact dataset into a moderate system raises the categorisation of everything in it. That is an argument for segregating data, not for lowering the rating. Programmes that quietly downgrade to avoid the cost have inverted the control.

Making data classification decide something

Under Article 32 of the GDPR, a controller implements measures appropriate to the risk. It weighs the state of the art, the cost of implementation, and the nature, scope, context and purposes of the processing. Appropriate to the risk is the phrase doing the work. Data classification is how an organisation makes that judgement repeatable instead of case by case.

So each level needs a named consequence. Who may grant access, and who approves it. Whether encryption at rest is required or optional. What the retention period is, and what destruction method closes it. Which environments the data may reach, including test systems. Whether the level triggers an assessment before a new use.

Classify the information, then the system

Order matters here and programmes get it backwards. Rate the information types first, because impact lives there. Derive the system rating afterwards, from what the system holds. Starting with systems produces labels that stop being true the moment a new feed arrives.

That also makes the scheme maintainable. A new dataset gets rated once. The system rating then recalculates from the types it now contains.

Data classification in the CIPM exam

The Body of Knowledge is the IAPP’s published outline of what each certification exam tests. Protecting personal data is Domain IV. It carries data classification, control purposes and limitations, and access controls in one competency. That grouping is the hint.

Scenario questions hand you a dataset and ask for the data classification. Or they hand you a classification and ask which control fits. Read for proportionality. The answer applying maximum protection to everything is usually wrong, because it fails the appropriateness test as surely as under-protection does. A second trap offers a classification decided by data type alone, ignoring the context that changes the impact. Our piece on privacy metrics covers how this work gets reported. The piece on privacy governance models covers who owns the scheme.

Take your own scheme and write the consequence beside each level. Any level with no consequence beside it is a label. The CIPM trial exam at €65 gives you a timed read on whether the domain is solid.

Similar Posts