Find personal information in your text and replace it with placeholders. Ai4Privacy gives you the synthetic datasets, the browser-based chat, the REST API and the local SDKs to do it.
Draft a reply to Maya Chen[FULLNAME_1] at maya.chen@example.org[EMAIL_1] about order EU-84921[ORDERNUMBER_1].
Original text: Draft a reply to Maya Chen at maya.chen@example.org about order EU-84921. Protected text: Draft a reply to [FULLNAME_1] at [EMAIL_1] about order [ORDERNUMBER_1].
3 entities ready
Cited by
Ai4Privacy's work is cited in research and industry publications from these institutions. Each logo links to its source.
Training a model, working with prompts, shipping a product or running it locally. Compare all products, then choose the path that matches the work in front of you.
Train and evaluate with synthetic PII data
Synthetic and multilingual, with no real personal data. Each release has its own page.
Start with the decisions that shape a training corpus, then carry the same rules through annotation and model evaluation.
PII Masking for AI Training Data: A Practitioner’s Guide
PII masking for AI training data is the process of finding personal data in a text corpus and replacing it, so that less identifiable information is carried into a model’s weights.
DefineInventoryChoose the personal-data types that matter.
FindDetectLocate each personal span in context.
ProtectReplaceKeep useful language around the placeholder.
Explore the complete guide set
Dr. Rina PatelFULLNAME joined North LabORGANIZATION.
PII Annotation Guidelines: Labeling Personal Data for NER Training
A PII annotation guideline is the written standard a labeler, or annotator, works from. It says which spans of text to mark, where each span starts and stops, and what to do when a case is ambiguous.
How to Evaluate a PII Detection Model: Precision, Recall and What They Miss
Evaluating a personally identifiable information detection model means measuring how many personal data spans it finds, how many things it flags that are not personal data, and what the aggregate figure is hiding.
How to Train a PII Detection Model with Named Entity Recognition
Build a PII detection model around explicit task framing, aligned labels, leak-free data splits and a clear decision about when training a new model is justified.
Data Minimization for AI Training Data Under the GDPR
Map data-minimization principles to practical decisions about what a training corpus contains, what gets masked and what should not be collected at all.
Research citations show relevance. These figures show the work being distributed and expanded in practice by Artificial Intelligence Suisse SA and its partners.
Short answers to the practical questions teams ask when comparing datasets, local packages and hosted access.
What does Ai4Privacy do?
Ai4Privacy helps teams find personal information in text and replace it with placeholders. The same workflow is available through synthetic datasets, a browser-based chat, a REST API and local Python and JavaScript SDKs.
Do the training datasets contain real personal data?
The datasets presented here use synthetic PII rather than real personal data. Each release has its own coverage, access and license details, so those facts should be checked on the individual release page.
Can PII detection run inside my own application?
Yes. The Python and JavaScript packages are designed for local execution inside your own application or workflow. The REST API is the separate hosted integration path.
Start with our open-source work.
The datasets are on Hugging Face, the projects are on GitHub, and the packages install from PyPI and npm.
If you are evaluating Ai4Privacy for your company, talk to us about access and integration.