Honest limits

What Sunda does — and what it doesn't.

A recruiter needs to know precisely what is detected, what isn't, and what risk remains after the swap. This page is that list, kept in sync with the extension. If something here changes, the product changed.

In one paragraph

Sunda Privacy Guard reads Word (.docx), text files, and PDFs you can select text in — not scans, photos or anything else that is really an image — and replaces structured identifiers (emails, phone numbers, links, card and bank numbers, Israeli ID numbers, numeric dates) with placeholders using always-on rules, everywhere. Names and postal addresses are handled by an optional on-device AI model that runs only in the popup flow, on machines with WebGPU. A redacted PDF is rebuilt from its text, so it keeps the words and page breaks but not the layout — and PDFs are the popup's job only, never the upload guard's. Employer names, titles and schools are never removed. Detection is good, not perfect. It reduces exposure; it makes nobody compliant with anything.

Formats: Word, text, and PDFs you can select text in

Most CVs arrive as PDFs, and Sunda now reads them — in the popup, and only when the file has a real text layer. It extracts that text, redacts it, and writes you a new PDF. That rebuild is the point: a PDF stores the same string in several places at once (the page content, the embedded font subsets, the document metadata, annotations, and any edit history appended to the file), so deleting the visible copy would leave the rest behind. A file that looks redacted while still carrying the candidate's name is the one outcome this product cannot ship. The cost of rebuilding is that you get the words and the page breaks, not the original layout, fonts or images.

Two PDFs still get refused outright, and refused loudly — nothing is downloaded and the popup tells you why. A scan or photo has no text to extract (there is no OCR), and a PDF written in Hebrew, Arabic or CJK uses characters the built-in PDF font cannot write, which would mean dropping them. And on Protected sites, a PDF still uploads untouched — like an image or a zip — because the upload guard only acts on the formats it can clean instantly. Hiding any of that would cost you a leak and us a one-star review, so here it is in the first section.

Supported

  • Word (.docx) — layout, headers, footers and footnotes preserved in the output.
  • Plain text (.txt) and Markdown (.md).
  • PDFs with a text layer, in the popup — rebuilt as a new PDF, so the words and page breaks survive and the layout does not.
  • Files up to 20 MiB on Protected sites; larger supported files are blocked rather than leaked.

Not supported

  • Scanned documents, photos, screenshots — any text inside an image. No OCR: refused, never passed through.
  • PDFs in Hebrew, Arabic or CJK — the built-in PDF font can't write those scripts; use .docx.
  • PDFs on Protected-site uploads — they go through untouched; PDF is the popup flow only.
  • The original PDF's layout — the output is rebuilt from the text, not edited in place.
  • Old .doc, Apple Pages, Google Docs edited in the browser — download as .docx first.
  • Spreadsheets, presentations, archives.

What to do with a PDF CV

  1. Try it. Open the PDF, and if you can select the text with your cursor, Sunda can read it — protect it straight from the popup.
  2. If Sunda says the PDF has no text in it, it is a scan or a photo. Ask the candidate, or your ATS export, for a .docx or a PDF you can select text in. There is no shortcut here and no tool that offers one without OCR.
  3. If you need the original layout preserved — sending it on to a client rather than to an AI — convert to .docx first (open the PDF in Word, Save As .docx) and protect that: the Word path keeps the formatting.
  4. Same for a CV in Hebrew, Arabic or CJK: protect it as .docx.

Detection: exactly what is caught, by which engine, in which flow

Two engines do the work. The always-on rules are pattern matches with checksums where one exists; they run instantly in every flow. The optional on-device AI model finds the things patterns can't — names and addresses — and runs only when you protect a file by hand in the popup, on a machine with WebGPU (Chrome 138+), after a one-time ~900 MB download from Hugging Face. Site uploads are rules-only on purpose: an upload has to be instant, and the model can take up to ~45 s to load cold.

IdentifierPlaceholderEngineWhere it runsNotes
Email addresses[EMAIL_n]RulesPopup + Protected sitesReliable.
Phone numbers[PHONE_n]RulesPopup + Protected sitesIsraeli (05x-…, 0x-…), international (+…), US-grouped. Not every national format.
Links[URL_n]RulesPopup + Protected sitesOnly with http(s):// or www.. A bare linkedin.com/in/name is not caught.
Credit/debit cards[CREDIT_CARD_n]RulesPopup + Protected sitesLuhn-checked, 13–19 digits.
IBANs[IBAN_n]RulesPopup + Protected sitesMod-97 checksum.
Israeli ID (Teudat Zehut)[ISRAELI_ID_n]RulesPopup + Protected sites9 digits, checksum-verified — the check rejects about nine in ten arbitrary 9-digit numbers, so an unrelated 9-digit reference is occasionally redacted.
Numeric dates[DATE_n]RulesPopup + Protected sites2026-08-19, 19/08/2026, 19.8.26.
People's names[PERSON_n]ModelPopup only, WebGPU machinesGood, not perfect. Misses happen.
Postal addresses[ADDRESS_n]ModelPopup only, WebGPU machinesGood, not perfect.
Written-out dates[DATE_n]ModelPopup only, WebGPU machines“March 2019”-style dates.
Account numbers, secrets[ACCOUNT_n], [SECRET_n]ModelPopup only, WebGPU machinesRare on CVs.

Within one document the same value always gets the same placeholder, so a CV stays readable and two candidates stay distinguishable. When a file is protected you get a count report (“Replaced 3 items — 1 email, 1 phone number, 1 Israeli ID”) and an honest note about which engine did the checking.

What is never removed

Consequence: a redacted CV is pseudonymised, not anonymous. A distinctive career path can still point at one person. That is normal and expected — see pseudonymisation vs anonymisation — but don't call the output anonymous.

Detection is not perfect. Human review is still required.

Rules are reliable for structured data; the model is good for names and addresses and will still miss some — a name split across a line break, an unusual format, a nickname. Before you share a protected file: read the report, open the output, scan the header block and the footer. Thirty seconds of review is the difference between “reduced” and “assumed”.

Two flows, two different guarantees

  1. Protect a file (popup). Rules + model (where available), and the only flow that handles PDFs. Output goes to your Downloads as name.redacted.docx — or .redacted.pdf, rebuilt from the text. Nothing is uploaded. The first protect after a browser start may wait for the model to load.
  2. Protected sites (upload guard). Rules only, instant, on the sites you added and nowhere else. Supported files that can't be processed or exceed 20 MiB are blocked (fail-closed); unsupported types pass through unchanged. One deliberate exception: if your 14-day trial has ended and no Pro license is active, the file is not protected and is sent as-is — you get a visible “Free trial ended — this upload was NOT protected” notice instead of a block, because a licensing verdict should never break your upload. With an empty list the extension runs on no web page at all.

Trial, license, network

Sunda reduces exposure. It does not make you compliant with anything. No GDPR badge, no Amendment 13 seal, no certification. Using Sunda does not satisfy a legal obligation; it removes identifiers from files before they leave your device. A lawful basis, a candidate privacy notice, a processor agreement, retention and deletion are still yours to handle — ask whoever does data protection for your business, and read the regulator's guidance.

Questions recruiters ask

Does Sunda Privacy Guard work on PDF CVs?

Yes, if you can select the text in it. In the popup, Sunda extracts that text, redacts it and writes you a new PDF — which keeps the words and the page breaks but not the layout, fonts or images, because it is rebuilt rather than edited. A scan or photo has no text to read and is refused outright (there is no OCR), as is a PDF in Hebrew, Arabic or CJK, whose characters the built-in PDF font cannot write; protect those as .docx. On a Protected site a PDF still uploads untouched — PDF redaction is the popup flow only.

Will it remove the candidate's name?

Names and postal addresses are detected by the optional on-device AI model, which runs in the popup “Protect a file” flow on machines with WebGPU (Chrome 138 or later) after a one-time download. On Protected-site uploads, and on machines without WebGPU, only the always-on rules run, and they do not detect names. Check the report and the output before you share.

Does it remove employer names, universities or job titles?

No, deliberately. A CV with no companies, schools or titles is useless for screening. Be aware that a distinctive career history can still identify a person even after the name, phone, email and address are gone.

Does using Sunda satisfy the GDPR or Amendment 13 for me?

No. Sunda reduces what you disclose by removing identifiers from a file before it leaves your device. It does not create a lawful basis, write your candidate privacy notice, sign a processor agreement, or satisfy any legal obligation. No tool does. Ask whoever handles data protection for your business and read the regulator's own guidance.

Is detection guaranteed?

No. Structured identifiers — emails, phone numbers in common formats, links with http(s):// or www., card numbers, IBANs, checksum-valid Israeli ID numbers, numeric dates — are caught reliably by rules. Names and addresses depend on the optional model and will miss some. The report tells you what was replaced; human review of the output is still required.