Skip to content
Back to blog
Benchmarks13 Aug 20264 min readRanjan Singh

Omit vs Google Cloud DLP: The Fast Tier

Google Cloud DLP recognises 213 kinds of sensitive text. Our free-with-every-install Fast tier recognises 191 of them on your own machine, with nothing uploaded, and matches or beats Google on every comparable type. Here is the measurement, including where we are weaker.

In short

Omit's Fast tier matched or beat Google Cloud DLP on 21 of the 21 entity types where both tools compete, at an average strict F2 of 0.8400 against Google's 0.4465. The corpus was 56 scored cells across 34 structured entity types, majority German and English. Strict F2 is recall-weighted and requires both an exact span and an exact type match, a harder bar than most published comparisons use. Of Google's 213 infoTypes, 171 map one to one onto an Omit detector, 34 map at a coarser label, and 8 are caught only by a generalized catch-all.

If you are evaluating PII detection tools, you will hit Google Cloud DLP, now sold as Sensitive Data Protection. It is the closest thing this field has to a reference catalogue, and the thing people ask us about most. This is part one of a two-part answer, and it measures our Fast tier: the lighter of our two detection tiers, the one bundled with every install, and the one that runs inside Protect, the clipboard guard we give away free.

The result first. Google recognises 213 kinds of sensitive text, which it calls infoTypes. A default Omit install has a live detector for 191 of them, on your own machine, with nothing uploaded anywhere. On accuracy, measured with a strict scoring standard explained below, our Fast tier matched or beat Google's best configuration on 21 of the 21 entity types where both tools actually compete, at an average strict F2 of 0.8400 against Google's 0.4465. Part two covers the second tier and a corpus of ordinary prose, where the gap is narrower and more interesting.

191
of Google's 213

Google Cloud DLP recognises 213 kinds of sensitive text. A default Omit install has a live detector for 191 of them, on your own machine, with nothing uploaded. The remaining 22 need the broader coverage mode switched on.

171

Mapped one to one

34

Found at a coarser label

8

Only via a generalized catch-all

0

With no Omit detector at all

What's in the coverage number

Coverage is the easier half of the question, and it is not uniform. Of the 213 infoTypes, 171 map one to one: Google's US_SOCIAL_SECURITY_NUMBER is our US_SSN, different name, same thing. 34 map at a coarser label: Google separates FIRST_NAME, LAST_NAME and gendered names where we detect all of them as PERSON, which finds and redacts the text identically but loses the fine label in an audit report. 8 are caught only through a generalized catch-all, the weakest form of coverage on the list. None are left with no detector at all. And 22 of the 213, the 8 catch-alls plus 14 demographic types backed by wordlists, only fire once the broader coverage mode is switched on, which it is not by default, so 191 is the honest default number, not 213.

Detection is only half the job

Detection is also only half of a redaction tool. Omit's PDF redaction does not draw a black rectangle over text and save it, which is the failure mode behind most recovered-redaction news stories. The text itself is removed and the output is checked before it is written; if the value is still recoverable, the page is escalated to a stronger treatment. Excel and CSV redaction rewrites the cell rather than the pixel. The original file is never modified in any format: Omit writes a new file and leaves yours untouched.

What we measured

What we measured. The corpus is 56 scored cells across 34 structured entity types, synthetic, majority German and English with single cells in four more languages. The metric is strict F2, recall-weighted and requiring an exact span and an exact type match, which is a harder bar than most published comparisons use. Google was scored twice on 22 July 2026, at its POSSIBLE likelihood threshold: once on its default infoType configuration, and once handed the exact infoType for every entity in the corpus, close to a best case for it. Our own numbers were re-measured on 17 August 2026 against those same frozen Google predictions.

Question one: detector quality, on the 21 types both tools attempt

21 of 21

Matched or beat Google's default configuration

21 of 21

Matched or beat Google configured at its best

Every type Google ships a text detector for, with no exclusions needed. Average strict F2 0.8400, against 0.4465 for Google at its best and 0.1143 on its defaults. On the July build the same figure was 0.6603.

Question two: does the PII come out redacted, across all 34 types

Omit, Fast tier30 of 34
4 missed, every one of them ours
Google, best configuration16 of 34
18 missed, 13 of them ship no detector
Google, default configuration4 of 34
30 missed, 13 of them ship no detector

What's still weak, and what we don't claim

What our Fast tier is still weak at, and it is a short list. UK_BANK_ACCOUNT, IMSI and US_ABA_ROUTING are all below 0.78. Our two genuine misses on the full 34-type set are NATIONAL_ID and TAX_ID, both open-ended catch-all types with no fixed shape for a pattern to match, and both are exactly what the Accuracy tier in part two is for: it catches both, at 0.8664 and 0.7688. The full per-entity table, and the two corrections we have made to our own published numbers along the way, are on the benchmark data page.

What we are not claiming. This is our harness, our corpus, our run, and nobody independent has audited it. Synthetic data is kinder to pattern-based detectors than real documents, and 56 cells is a small benchmark. We are also not claiming Omit replaces Google Cloud DLP: if you need to scan cloud storage across an organisation with central dashboards and case management, that is what Google's product is for and we have none of it. We wrote about that boundary on the Omit versus cloud DLP page.

What we do claim is narrower, and I think it is worth something. For sensitive text on a desktop, the lighter of our two tiers, the one that ships in every install and powers a clipboard guard we charge nothing for, holds its own against a metered cloud service configured by an expert, without the document ever leaving the machine. For a lot of documents, being inspected in someone else's cloud is itself the disclosure you were trying to avoid, and no accuracy score fixes that.

Get the next one by email

New benchmarks and release notes as they go up. Nothing else, and unsubscribe by replying.

New posts onlyunsubscribe by replying