Skip to content
Back to the comparison
Benchmark dataUpdated 17 Aug 2026Ranjan SinghGitHubLinkedIn

Omit vs Google Cloud DLP: The Full Benchmark

The complete entity-by-entity data behind our Google Cloud DLP comparison: every structured identifier and prose entity type, both of our detection tiers, both of Google's configurations, methodology, and the corrections we have made in public.

This is the reference page behind our Google Cloud DLP comparison: every scored cell, both corpora, both tiers, both Google configurations, in one place. The blog posts tell the story; this page is the working. Nothing here is narrative, so start with the numbers.

Structured identifiers, 34 types

VAT numbers, national IDs, passports, bank codes. Fixed shape, often a checksum.

Omit, Fast tier0.8400
Omit, Accuracy tier0.7758
Google, expert configuration0.4465
Google, default configuration0.1143

21 of 21 comparable entity types matched or beaten, either configuration.

Ordinary prose, 14 comparable types

Emails, notes, records. Names, dates and places in running sentences.

Omit, Accuracy tier0.8314
Omit, Fast tier0.6782
Google, expert configuration0.5684

12 of 14 comparable entity types beaten outright. URL and VIN are Google's, both within 1 to 2 points of a tie.

Method

Two corpora. Benchmark v1 is synthetic, 34 structured identifier types with a fixed shape and often a checksum: VAT numbers, national IDs, passports, bank codes. The ai4privacy corpus is 1,500 documents of ordinary prose in English, German, French and Italian: emails, notes, records, the kind of text a person actually pastes into something. The metric is strict F2 on both: recall-weighted, and a detection only counts when the span and the type both match exactly. Google Cloud DLP v2 was run at its POSSIBLE likelihood threshold. On the structured corpus it was scored twice, once on its default infoType configuration and once handed the exact infoType for every entity in the corpus, which is close to a best case for it. On the prose corpus it was scored once, with an explicit infoType list rather than its unrestricted default, because an unrestricted scan on ordinary text is likely to return infoTypes our label map never anticipated. We have not measured Google's unrestricted default on either corpus, and defaults are usually worse, not better.

Structured identifiers, all 34 types

EntityFast F2Accuracy F2Google, defaultGoogle, expert
ADDRESS0.97621.00000.00000.8548
DE_HEALTH_INSURANCE1.00001.0000n/an/a
DE_ID_CARD1.00001.00000.00000.5000
DE_LICENSE_PLATE0.91841.0000n/an/a
DE_PASSPORT0.97620.99020.00000.4505
DE_SVNR1.00001.0000n/an/a
DE_TAX_ID1.00001.00000.00000.0000
DE_VAT1.00001.00000.00000.0000
DRIVER_LICENSE0.85390.97620.00000.5000
ES_NIF1.00001.00001.00001.0000
EU_VAT0.86560.00000.00000.2381
FI_PERSONAL_IDENTITY_CODE1.00001.00001.00001.0000
ICCID1.00001.00000.00000.5309
IMSI0.69890.00000.00000.0000
IN_AADHAAR1.00001.00000.00000.5000
IN_AADHAAR_VID1.00001.0000n/an/a
IN_GSTIN1.00001.00000.00000.5000
IN_IFSC1.00001.0000n/an/a
IN_PAN1.00001.00000.00000.0610
IN_PASSPORT1.00000.00000.00001.0000
IN_UPI_VPA1.00000.6744n/an/a
IN_VEHICLE_REGISTRATION1.00000.7671n/an/a
IN_VOTER1.00000.6989n/an/a
IT_FISCAL_CODE1.00001.00000.00001.0000
MRZ1.00001.0000n/an/a
NATIONAL_ID0.00000.8664n/an/a
PASSPORT0.83330.83330.00000.7980
PL_PESEL1.00001.00001.00001.0000
POSTAL_CODE0.00000.00000.00000.0000
STREET_ADDRESS0.00000.00000.00000.0000
TAX_ID0.00000.7688n/an/a
UK_BANK_ACCOUNT0.60441.0000n/an/a
UK_SORT_CODE1.00000.8333n/an/a
US_ABA_ROUTING0.77740.77740.50000.5000
Strict F2, benchmark v1 (synthetic, 34 structured identifier types, 1,763 documents). n/a means Google ships no text detector for that type at all, scored as a coverage gap, never as a Google miss. POSTAL_CODE and STREET_ADDRESS are zero for every engine here, a defect in how our benchmark splits ADDRESS, not a result for anyone.Verified 17 Aug 2026.

Ordinary prose, all 14 comparable types

EntityFast F2Accuracy F2Google F2Winner
BIC_SWIFT0.99550.99550.4169Omit
CREDIT_CARD0.04750.86390.0558Omit
DATE_TIME0.46400.71860.4281Omit
EMAIL_ADDRESS0.99310.96640.9861Omit
IBAN_CODE0.94490.98440.9652Omit
IMEI1.00000.87480.2906Omit
IP_ADDRESS1.00000.95350.9962Omit
LOCATION0.29280.52500.4124Omit
MAC_ADDRESS0.94371.00000.6464Omit
PERSON0.24530.51780.4108Omit
PHONE_NUMBER0.61880.86170.3519Omit
URL0.95820.98460.9976Google
USERNAME0.00000.63640.0000Omit
VIN0.99040.75721.0000Google
Strict F2, ai4privacy prose corpus (1,500 documents, en/de/fr/it), the 14 entity types Google ships a text detector for. Google Cloud DLP v2, explicit infoTypes matching these 14 entities, POSSIBLE likelihood threshold, not Google's unrestricted default, which is untested here and usually scores lower. Winner is the better of our two tiers against Google's single configuration. Google still wins outright on URL and VIN, both within 1 to 2 points of a tie.Verified 17 Aug 2026, both sides, same day.

Why the Accuracy tier's structured-ID score reads lower

On the structured corpus the Accuracy tier covers more of the text than the Fast tier, 96.80% of gold spans against 92.73%, but scores a lower strict F2. That is the metric, not the redaction. Strict F2 needs an exact type match, and the transformer often returns a broader-but-true label instead of the specific one our gold data expects.

Same characters, same span boundaries, higher confidence, different name

VAT number: NL557993085B24

Fast tierEU_VAT0.65
Accuracy tierTAX_ID0.96

Broader, and true. A VAT number is a tax ID.

Passport H5876797 issued.

Fast tierIN_PASSPORT0.90
Accuracy tierPASSPORT0.98

Broader, and true. Less specific than our gold label.

IMSI 579825746804539

Fast tierIMSI0.40
Accuracy tierCREDIT_CARD0.97

Simply wrong. Fifteen digits reads as a card number.

What we found and fixed on the way here

The first prose measurement against Google, on 13 Aug 2026, found Google beating both our tiers outright on nine of fourteen entity types. Chasing that down turned up real defects, not benchmark tuning: a URL match was swallowing the sentence's closing period, which strict scoring counts as a total miss; IMEI numbers printed with dashes were not recognised at all; a validated VIN was losing an overlap fight it should always win, because a confident but wrong general-purpose guess was allowed to outscore an exact-format match on the same span, a bug URLs shared. Fixing those closed most of the gap. Two entities were still behind on 17 Aug: BIC_SWIFT, because we deliberately kept its confidence low to avoid flagging an ordinary capitalised word as a bank code, and a separate registration bug that left four languages (Spanish, Italian, Polish, Finnish) with no financial-identifier detection at all, in the shipped app, not only the benchmark. Both are fixed now: BIC_SWIFT moved from 0.2470 to 0.9955, and the fourth language, Italian, moved from 0.0 to 1.0 once the detector was actually wired up for it.

Corrections to our own published numbers

We have corrected our own published aggregate twice. The first time, our scoring code counted the 13 entity types Google ships no detector for as Google losses instead of coverage gaps, which overstated our lead. The second time, we published a corrected but stale figure, average strict F2 0.6603, measured against a three- week-old build; re-running the identical corpus against the current build moved it to 0.8400. Both corrections are logged with dates in the source documents linked below.

Cost

CorpusTierPer documentPeak memory
Structured (1,763 docs)Fast13.9 ms2,166 MB
Structured (1,763 docs)Accuracy107.8 ms3,426 MB
Prose (1,500 docs)Fast19.1 ms2.5 GB
Prose (1,500 docs)Accuracy182.0 ms3.9 GB
CPU only, no GPU numbers exist for the Accuracy tier. The Accuracy tier costs roughly eight to ten times the Fast tier's latency and about 1.3 GB more resident memory.

What this does not measure

Both corpora are ours, both benchmarks are small, and nobody independent has audited either. The structured corpus's Google numbers are several weeks older than ours, because rerunning them costs money and the predictions are stable, but the service may have moved since. The prose corpus covers four languages, not the seven we ship. And we are not publishing the detector inventory, the models, the weights or the thresholds behind any of this: a published description of how detection works is a published description of how to evade it. The numbers are real; the mechanism stays private.

Every file behind this page is in the product repository: the raw scored cells for both tiers on both corpora, the frozen Google cells, and the scripts that produced all of it.

Get the next one by email

New benchmarks and release notes as they go up. Nothing else, and unsubscribe by replying.

New posts onlyunsubscribe by replying