Omit vs Google Cloud DLP: The Full Benchmark
The complete entity-by-entity data behind our Google Cloud DLP comparison: every structured identifier and prose entity type, both of our detection tiers, both of Google's configurations, methodology, and the corrections we have made in public.
This is the reference page behind our Google Cloud DLP comparison: every scored cell, both corpora, both tiers, both Google configurations, in one place. The blog posts tell the story; this page is the working. Nothing here is narrative, so start with the numbers.
Structured identifiers, 34 types
VAT numbers, national IDs, passports, bank codes. Fixed shape, often a checksum.
21 of 21 comparable entity types matched or beaten, either configuration.
Ordinary prose, 14 comparable types
Emails, notes, records. Names, dates and places in running sentences.
12 of 14 comparable entity types beaten outright. URL and VIN are Google's, both within 1 to 2 points of a tie.
Method
Two corpora. Benchmark v1 is synthetic, 34 structured identifier types with a fixed shape and often a checksum: VAT numbers, national IDs, passports, bank codes. The ai4privacy corpus is 1,500 documents of ordinary prose in English, German, French and Italian: emails, notes, records, the kind of text a person actually pastes into something. The metric is strict F2 on both: recall-weighted, and a detection only counts when the span and the type both match exactly. Google Cloud DLP v2 was run at its POSSIBLE likelihood threshold. On the structured corpus it was scored twice, once on its default infoType configuration and once handed the exact infoType for every entity in the corpus, which is close to a best case for it. On the prose corpus it was scored once, with an explicit infoType list rather than its unrestricted default, because an unrestricted scan on ordinary text is likely to return infoTypes our label map never anticipated. We have not measured Google's unrestricted default on either corpus, and defaults are usually worse, not better.
Structured identifiers, all 34 types
| Entity | Fast F2 | Accuracy F2 | Google, default | Google, expert |
|---|---|---|---|---|
| ADDRESS | 0.9762 | 1.0000 | 0.0000 | 0.8548 |
| DE_HEALTH_INSURANCE | 1.0000 | 1.0000 | n/a | n/a |
| DE_ID_CARD | 1.0000 | 1.0000 | 0.0000 | 0.5000 |
| DE_LICENSE_PLATE | 0.9184 | 1.0000 | n/a | n/a |
| DE_PASSPORT | 0.9762 | 0.9902 | 0.0000 | 0.4505 |
| DE_SVNR | 1.0000 | 1.0000 | n/a | n/a |
| DE_TAX_ID | 1.0000 | 1.0000 | 0.0000 | 0.0000 |
| DE_VAT | 1.0000 | 1.0000 | 0.0000 | 0.0000 |
| DRIVER_LICENSE | 0.8539 | 0.9762 | 0.0000 | 0.5000 |
| ES_NIF | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| EU_VAT | 0.8656 | 0.0000 | 0.0000 | 0.2381 |
| FI_PERSONAL_IDENTITY_CODE | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| ICCID | 1.0000 | 1.0000 | 0.0000 | 0.5309 |
| IMSI | 0.6989 | 0.0000 | 0.0000 | 0.0000 |
| IN_AADHAAR | 1.0000 | 1.0000 | 0.0000 | 0.5000 |
| IN_AADHAAR_VID | 1.0000 | 1.0000 | n/a | n/a |
| IN_GSTIN | 1.0000 | 1.0000 | 0.0000 | 0.5000 |
| IN_IFSC | 1.0000 | 1.0000 | n/a | n/a |
| IN_PAN | 1.0000 | 1.0000 | 0.0000 | 0.0610 |
| IN_PASSPORT | 1.0000 | 0.0000 | 0.0000 | 1.0000 |
| IN_UPI_VPA | 1.0000 | 0.6744 | n/a | n/a |
| IN_VEHICLE_REGISTRATION | 1.0000 | 0.7671 | n/a | n/a |
| IN_VOTER | 1.0000 | 0.6989 | n/a | n/a |
| IT_FISCAL_CODE | 1.0000 | 1.0000 | 0.0000 | 1.0000 |
| MRZ | 1.0000 | 1.0000 | n/a | n/a |
| NATIONAL_ID | 0.0000 | 0.8664 | n/a | n/a |
| PASSPORT | 0.8333 | 0.8333 | 0.0000 | 0.7980 |
| PL_PESEL | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| POSTAL_CODE | 0.0000 | 0.0000 | 0.0000 | 0.0000 |
| STREET_ADDRESS | 0.0000 | 0.0000 | 0.0000 | 0.0000 |
| TAX_ID | 0.0000 | 0.7688 | n/a | n/a |
| UK_BANK_ACCOUNT | 0.6044 | 1.0000 | n/a | n/a |
| UK_SORT_CODE | 1.0000 | 0.8333 | n/a | n/a |
| US_ABA_ROUTING | 0.7774 | 0.7774 | 0.5000 | 0.5000 |
Ordinary prose, all 14 comparable types
| Entity | Fast F2 | Accuracy F2 | Google F2 | Winner |
|---|---|---|---|---|
| BIC_SWIFT | 0.9955 | 0.9955 | 0.4169 | Omit |
| CREDIT_CARD | 0.0475 | 0.8639 | 0.0558 | Omit |
| DATE_TIME | 0.4640 | 0.7186 | 0.4281 | Omit |
| EMAIL_ADDRESS | 0.9931 | 0.9664 | 0.9861 | Omit |
| IBAN_CODE | 0.9449 | 0.9844 | 0.9652 | Omit |
| IMEI | 1.0000 | 0.8748 | 0.2906 | Omit |
| IP_ADDRESS | 1.0000 | 0.9535 | 0.9962 | Omit |
| LOCATION | 0.2928 | 0.5250 | 0.4124 | Omit |
| MAC_ADDRESS | 0.9437 | 1.0000 | 0.6464 | Omit |
| PERSON | 0.2453 | 0.5178 | 0.4108 | Omit |
| PHONE_NUMBER | 0.6188 | 0.8617 | 0.3519 | Omit |
| URL | 0.9582 | 0.9846 | 0.9976 | |
| USERNAME | 0.0000 | 0.6364 | 0.0000 | Omit |
| VIN | 0.9904 | 0.7572 | 1.0000 |
Why the Accuracy tier's structured-ID score reads lower
On the structured corpus the Accuracy tier covers more of the text than the Fast tier, 96.80% of gold spans against 92.73%, but scores a lower strict F2. That is the metric, not the redaction. Strict F2 needs an exact type match, and the transformer often returns a broader-but-true label instead of the specific one our gold data expects.
Same characters, same span boundaries, higher confidence, different name
VAT number: NL557993085B24
Broader, and true. A VAT number is a tax ID.
Passport H5876797 issued.
Broader, and true. Less specific than our gold label.
IMSI 579825746804539
Simply wrong. Fifteen digits reads as a card number.
What we found and fixed on the way here
The first prose measurement against Google, on 13 Aug 2026, found Google beating both our tiers outright on nine of fourteen entity types. Chasing that down turned up real defects, not benchmark tuning: a URL match was swallowing the sentence's closing period, which strict scoring counts as a total miss; IMEI numbers printed with dashes were not recognised at all; a validated VIN was losing an overlap fight it should always win, because a confident but wrong general-purpose guess was allowed to outscore an exact-format match on the same span, a bug URLs shared. Fixing those closed most of the gap. Two entities were still behind on 17 Aug: BIC_SWIFT, because we deliberately kept its confidence low to avoid flagging an ordinary capitalised word as a bank code, and a separate registration bug that left four languages (Spanish, Italian, Polish, Finnish) with no financial-identifier detection at all, in the shipped app, not only the benchmark. Both are fixed now: BIC_SWIFT moved from 0.2470 to 0.9955, and the fourth language, Italian, moved from 0.0 to 1.0 once the detector was actually wired up for it.
Corrections to our own published numbers
We have corrected our own published aggregate twice. The first time, our scoring code counted the 13 entity types Google ships no detector for as Google losses instead of coverage gaps, which overstated our lead. The second time, we published a corrected but stale figure, average strict F2 0.6603, measured against a three- week-old build; re-running the identical corpus against the current build moved it to 0.8400. Both corrections are logged with dates in the source documents linked below.
Cost
| Corpus | Tier | Per document | Peak memory |
|---|---|---|---|
| Structured (1,763 docs) | Fast | 13.9 ms | 2,166 MB |
| Structured (1,763 docs) | Accuracy | 107.8 ms | 3,426 MB |
| Prose (1,500 docs) | Fast | 19.1 ms | 2.5 GB |
| Prose (1,500 docs) | Accuracy | 182.0 ms | 3.9 GB |
What this does not measure
Both corpora are ours, both benchmarks are small, and nobody independent has audited either. The structured corpus's Google numbers are several weeks older than ours, because rerunning them costs money and the predictions are stable, but the service may have moved since. The prose corpus covers four languages, not the seven we ship. And we are not publishing the detector inventory, the models, the weights or the thresholds behind any of this: a published description of how detection works is a published description of how to evade it. The numbers are real; the mechanism stays private.
Every file behind this page is in the product repository: the raw scored cells for both tiers on both corpora, the frozen Google cells, and the scripts that produced all of it.