RedactCA
Redaction & anonymization for Indian financial and personal data
1Add files
2Review & tune
3Apply & download
.pdf.xlsx.csv.docx.json.txt
Every file is processed locally in your browser. There is no server and no API call carrying your data.
What does touch the network
Library code — and, on first OCR use, the Tesseract language model — is fetched from a public CDN.
That traffic contains no document content. For fully offline use, save the libraries alongside this file.
This is a supplementary tool. Detection is probabilistic and will miss or over-flag items —
review every output before sharing it.
Anonymized output is pseudonymized, not anonymous. While the vault exists, every token can be
reversed to the person it came from — the output is still personal data under DPDP/GDPR. Keep the vault
and the anonymized files in separate places: together they are the original.
Project vaultclosed
Files
Drop files here
or click to browse · multiple files supported
.txt .csv .json .pdf .xlsx .xls .docx
Redaction mode
Pseudonyms are consistent: the same value always maps to the same placeholder, across every file in this session.
Per-type overrides
Detection patterns
Account number by length
Bank account and CIF numbers have no reliable fixed format. Prefer column/key targeting for
structured files. This length rule is a blunt fallback and will over-match.
Custom patterns & term lists
PDF & OCR options
Both are true redaction — the text is gone from the file, not merely covered.
Content-stream editing is verified after writing: any page where the text survives is
automatically flattened instead. Stream editing can shift the remainder of an edited line
slightly; flattening cannot, but drops the text layer and enlarges the file.
Review defaults
Preview & review
Add a file to begin.
Restore report
Detected entities
0
Selected
0
Found
No file selected.
Output
The vault is never bundled with the anonymized files
and never downloads automatically — vault + output together reconstruct the originals.
Originals are never modified — every download is a new file. Source data is wiped
from memory once all outputs are downloaded.