DataNinja

Removing customer names before you send a spreadsheet

You want to paste a customer list into an AI, or send it to an analyst outside the business. Four common ways of handling that, three of which fail quietly.

Everything else in a spreadsheet fails loudly. A wrong total shows up in the reconciliation; a bad account code gets the import refused. This one fails silently, the person harmed isn't you, and by the time anybody finds out the file has already been sent.

Hashing the names doesn't work

Running names or phone numbers through MD5 or SHA-256 feels like the technical answer. It isn't, because the set of possible values is tiny.

There are roughly a hundred million possible Australian mobile numbers. Hashing all of them and looking yours up in the list takes seconds on a laptop — no key needed, no weakness in the algorithm, just an exhaustive list of every answer. First names are far worse: there are only a few thousand common ones.

A hash of a value drawn from a small set is a lookup, not a disguise.

Substituting realistic fake names is worse than it looks

Replacing Priya Raman with Sarah Chen produces a file that reads naturally, which is exactly the problem. Nothing downstream knows the names are invented, so somebody acts on one — quotes it in a report, emails it, follows it up.

And generate enough of them and one will be a real person, now attached to somebody else's transactions. Sequential labels — Customer 001 — are ugly and can't be mistaken for a person.

Deleting the column breaks the file

It's the one approach that's genuinely safe for that column, and it costs you the analysis. Without a consistent identifier you can't group by customer, so you can't ask which customers are worth anything, which is usually why the file was going anywhere.

It also misses the names everywhere else. A name in a Notes column identifies somebody just as well as one in Customer, and free-text fields are full of them — "rang Priya about the invoice".

Find-and-replace in Excel wrecks the file

Excel's Replace All has no concept of a word. Replace a customer called Ali and every ali inside another word goes with it: a real 611-row file came back with Australian Capital Territory reading Austr[label]an Capital [label]tory — 234 states and 162 localities destroyed.

The damage is invisible, too. Nothing checks whether a word you destroyed was ever a name, so nothing tells you it happened. You'd find it by reading the file.

What actually works

Sequential labels Customer 001. Not a hash, not a fake person. Nothing to reverse and nothing to mistake for real.
The same label every time One person gets one label across the whole file, so grouping and totals still work and the file is still worth analysing.
Replaced everywhere, not just in its own column Every name you replace has to be hunted down in every other column, and so do its parts. Replacing Priya Raman and leaving "rang Priya" is the whole failure.
Matched as whole words Which is what keeps Ali out of the middle of Australia. The cost is that a name run together with the word beside it isn't found.
You keep the original It's your key. When an answer comes back about Customer 001 you match it up yourself, against the file you still have. Nobody else needs to hold that mapping.

What nobody can promise you

No tool can tell you a file is anonymous, safe, or "PII compliant". There is no such standard to be compliant with, and de-identified data becomes personal information again the moment somebody can work out who it's about. A file is never compliant — how an organisation handles it might be. Treat any tool that tells you otherwise as the risk.

Two specific limits worth knowing, whatever you use:

A person mentioned only in free text can't be found. If they appear nowhere in a column anything recognised as holding people, there's no way to know the word is a person. Read the file before you send it.

Dates and amounts identify people too, in combination. A date of birth beside a postcode and a suburb is the textbook re-identifier, and it doesn't need a name attached. Those columns are usually the reason the file is being sent anywhere, so removing them isn't an option — but knowing it is.

What we do about all this

The De-identify tab does the list above: sequential labels, consistent across the file, swept through every other column including free text, matched as whole words. Before it hands the file back it searches the whole output for every value it replaced and every part of one, and one survivor stops the run rather than producing a file that looks handled.

You get an inventory of every column, including the ones left untouched, so you can see what was done rather than take it on trust — and you can overrule it either way, because you know Guarantor is a person and Product name isn't, and no word list does.

It's free, nothing is stored, and the file never leaves Australia. Take the people out of a spreadsheet.