AI dataset registry
Training datasets from Hugging Face / Kaggle and others, with license, PII flags, and poisoning signals. Same database, same API as models, MCP, and packages.
- 0.60TypicaAI/pii-masking-60k_frpkg:data/TypicaAI/pii-masking-60k_fr
PII French dataset This PII French dataset is based on the World's largest open-source privacy dataset: ai4privacy/pii-masking-200k. The original dataset ai4privacy/pii-masking-200k was filtered out, using a BERT-based language classifier, to keep only French rows. This dataset was created solely for educational purposes. For more information, please refer to the dataset ai4privacy/pii-masking-200k.
huggingfacetoken-classification34 downloads - 0.60cicero-im/piiptbrchatmlpkg:data/cicero-im/piiptbrchatml
Dataset Card for PII PT-BR ChatML The piiptbrchatml dataset is designed for training and evaluating models for Personal Identifiable Information (PII) masking in Brazilian Portuguese. It contains conversations where a system is instructed to mask PII from user inputs. The dataset includes the original text, the masked text, and the identified PII entities. O dataset piiptbrchatml foi criado para treinar e avaliar modelos para mascaramento de Informações Pessoais Identificáveis… See the full description on the dataset page: https://huggingface.co/datasets/cicero-im/piiptbrchatml.
huggingface26 downloads - 0.60cicero-im/modifiedpkg:data/cicero-im/modified
Dataset Card for Modified-Anonymization-Dataset This dataset contains anonymization examples in Portuguese. It consists of text samples where Personally Identifiable Information (PII) has been masked. The dataset includes the original text, the masked text, the identified PII entities, and information about potential data pollution introduced during the anonymization process. Este dataset cont[u00e9m exemplos de anonimiza[u00e7[u00e3o em portugu[u00eas. Consiste em amostras de… See the full description on the dataset page: https://huggingface.co/datasets/cicero-im/modified.
huggingface16 downloads - 0.60
- 0.60LocalDoc/pii_benchmarkpkg:data/LocalDoc/pii_benchmark
Azerbaijani PII / NER Benchmark A synthetic benchmark for evaluating Personally Identifiable Information (PII) detection and Named Entity Recognition (NER) in Azerbaijani. The dataset is designed to test models on realistic Azerbaijani user messages across multiple domains, including informal writing, spelling noise, Azerbaijani morphology, Latin-script transliteration, short fragments, long support-style messages, and adversarial hard-negative identifiers that resemble PII. The… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/pii_benchmark.
huggingfacetoken-classification0 downloads - 0.60bbanany/step4_sllm_v2pkg:data/bbanany/step4_sllm_v2
Step 4 PII 후보 판정 데이터셋 — 최종 10,000개 RAG 답변에서 상위 NER 단계가 추출한 후보가 문맥상 특정 자연인의 개인정보인지 PII 또는 NOT_PII로 판정하도록 Qwen을 SFT하기 위한 합성 데이터셋이다. 바로 사용하는 파일 step4_final_10000_qwen_train.jsonl: 학습 8,000개 step4_final_10000_qwen_valid.jsonl: 검증 1,000개 step4_final_10000_qwen_test.jsonl: 최종 평가 1,000개 step4_final_10000_qwen_all.jsonl: 전체 확인용 10,000개 각 행의 최상위 필드는 messages 하나뿐이며 system, user, assistant 순서다. 학습 시 Qwen tokenizer의 chat template를 적용하고 assistant 응답 부분에만 loss를 계산한다.… See the full description on the dataset page: https://huggingface.co/datasets/bbanany/step4_sllm_v2.
huggingfacetext-classification0 downloads - 0.60
- 0.55guneeshv/REDACT-PII-Benchmarkpkg:data/guneeshv/REDACT-PII-Benchmark
REDACT REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection Accepted to the EMNLP 2026 Industry Track Paper / arXiv · GitHub REDACT is a multilingual benchmark for evaluating personal information detection under systematically controlled generation conditions. 13,427 records · 324,078 entity annotations · 51 canonical entity types · 25 languages · 9 scripts · 4,127 surface-form patterns Benchmark task The headline task is… See the full description on the dataset page: https://huggingface.co/datasets/guneeshv/REDACT-PII-Benchmark.
huggingfacetoken-classificationother47 downloads - 0.55ai4privacy/pii-masking-health-phi-200kpkg:data/ai4privacy/pii-masking-health-phi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Health & Medical Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 252,437 1,686,246 49 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-200k.
huggingfacetoken-classificationother45 downloads - 0.55lianghsun/tw-PII-chatpkg:data/lianghsun/tw-PII-chat
Taiwan PII Chat (tw-PII-chat) v3 release (2026-05) — 2.4× larger than v2, fixes v2's distribution-shift regression on tw-PII-bench mid/long splits. Supersedes both v1 (61K synthetic short-form) and v2 (76K mixed) releases. Property Value Languages Traditional Chinese (zh-TW), English Items 183,588 Format Chat-format JSON (messages field) + raw NER spans (text + spans) Labels 8 in-schema (matching openai/privacy-filter) + 11 Taiwan-specific OOD Generation Mixed:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-PII-chat.
huggingfacetoken-classificationapache-2.041 downloads - 0.55ai4privacy/pii-masking-work-pwi-200kpkg:data/ai4privacy/pii-masking-work-pwi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Work & HR Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 252,273 1,383,008 41 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-work-pwi-200k.
huggingfacetoken-classificationother39 downloads - 0.55Vrandan/pii-harmonized-corpus-v2pkg:data/Vrandan/pii-harmonized-corpus-v2
pii-harmonized-corpus-v2 Harmonized + synthetic-augmented English-only PII NER training corpus, derived from three public datasets and Kimi K2.6 synthetic generation. Stats at a glance Train rows: 204,546 Test rows: 90,160 Total spans (train): 865,473 Total spans (test): 528,449 Real rows in train: 180,892 Synthetic rows in train: 23,654 (11.6%) Languages: English only (language == "en" for all rows) ML labels: 46 entity types Tagging: BILOU at training time (1 + 4×46… See the full description on the dataset page: https://huggingface.co/datasets/Vrandan/pii-harmonized-corpus-v2.
huggingfacetoken-classificationother38 downloads - 0.55ai4privacy/pii-masking-digital-pdi-200kpkg:data/ai4privacy/pii-masking-digital-pdi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Digital Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 198,319 815,110 33 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-200k.
huggingfacetoken-classificationother34 downloads - 0.55ai4privacy/pii-masking-location-pli-200kpkg:data/ai4privacy/pii-masking-location-pli-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Location & Travel Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 256,762 2,050,300 54 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-location-pli-200k.
huggingfacetoken-classificationother34 downloads - 0.55ai4privacy/pii-masking-financial-pfi-200kpkg:data/ai4privacy/pii-masking-financial-pfi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🇪🇺 Personal Financial Information — European PII Dataset Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy Entries PII Annotations Labels Languages Regions 257,434 1,563,807 48 23 29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-200k.
huggingfacetoken-classificationother33 downloads - 0.55ai4privacy/pdi-masking-100k-fullpkg:data/ai4privacy/pdi-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Digital Information (PDI) Masking Dataset — Full Overview The EPII PDI Masking Dataset is a large-scale, multilingual dataset of 91,400 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pdi-masking-100k-full.
huggingfacetoken-classificationother25 downloads - 0.55ai4privacy/pfi-masking-100k-fullpkg:data/ai4privacy/pfi-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Financial Information (PFI) Masking Dataset — Full Overview The EPII PFI Masking Dataset is a large-scale, multilingual dataset of 160,403 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pfi-masking-100k-full.
huggingfacetoken-classificationother25 downloads - 0.55ai4privacy/phi-masking-100k-fullpkg:data/ai4privacy/phi-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Health Information (PHI) Masking Dataset — Full Overview The EPII PHI Masking Dataset is a large-scale, multilingual dataset of 91,339 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k-full.
huggingfacetoken-classificationother23 downloads - 0.55ai4privacy/pli-masking-100k-fullpkg:data/ai4privacy/pli-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Location Information (PLI) Masking Dataset — Full Overview The EPII PLI Masking Dataset is a large-scale, multilingual dataset of 91,314 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pli-masking-100k-full.
huggingfacetoken-classificationother21 downloads - 0.55ai4privacy/pwi-masking-100k-fullpkg:data/ai4privacy/pwi-masking-100k-full
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Work Information (PWI) Masking Dataset — Full Overview The EPII PWI Masking Dataset is a large-scale, multilingual dataset of 91,559 annotated text samples containing synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pwi-masking-100k-full.
huggingfacetoken-classificationother21 downloads - 0.55ai4privacy/pii-masking-location-pli-400kpkg:data/ai4privacy/pii-masking-location-pli-400k
👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Location & Travel Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 423,933 3,323,062 34 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-location-pli-400k.
huggingfacetoken-classificationother3 downloads - 0.55ai4privacy/pii-masking-financial-pfi-400kpkg:data/ai4privacy/pii-masking-financial-pfi-400k
👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Financial Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 426,660 2,587,698 38 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-400k.
huggingfacetoken-classificationother1 downloads - 0.55ai4privacy/pii-masking-work-pwi-400kpkg:data/ai4privacy/pii-masking-work-pwi-400k
👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Work & HR Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 418,580 2,305,517 26 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-work-pwi-400k.
huggingfacetoken-classificationother1 downloads - 0.55ai4privacy/pii-masking-health-phi-400kpkg:data/ai4privacy/pii-masking-health-phi-400k
👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Health & Medical Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 417,900 2,802,316 37 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-400k.
huggingfacetoken-classificationother0 downloads - 0.55ai4privacy/pii-masking-digital-pdi-350kpkg:data/ai4privacy/pii-masking-digital-pdi-350k
👉 Looking for the open multilingual baseline? Start with ai4privacy/pii-masking-openpii-1.5m (1.5M samples, 30 languages, open-PII taxonomy). 🇪🇺🌏 Personal Digital Information, Global PII Dataset Part of PII-Masking-3M by Ai4Privacy, the global (2M base + Asia Pacific) PII-masking corpus. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Entries PII Annotations Labels Languages Regions 369,310 1,281,920 28 30 37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-350k.
huggingfacetoken-classificationother0 downloads - 0.50SII-lyXiang/SynthLeakpkg:data/SII-lyXiang/SynthLeak
SynthLeak SynthLeak is a dialogue dataset for studying contextual privacy leakage (CPL) in user-generated text. Each item is a multi-turn or single-post forum-style dialogue with structured annotations linking privacy clues in the dialogue to personal information items (PII). This release follows the dataset setting described in: Hangyu Ye, Liyao Xiang, Naixuan Huang, Dongyue Yu, Lijun Zhang, and Gang Wang. PrivSniffer: Graph-based Contextual Privacy Leakage Detection for… See the full description on the dataset page: https://huggingface.co/datasets/SII-lyXiang/SynthLeak.
huggingfacetoken-classification672 downloads - 0.50bbeglerov/russian-pi-66k-opfpkg:data/bbeglerov/russian-pi-66k-opf
Russian PII 66K OPF Format This dataset is a converted, OPF-compatible version of wolframko/russian-pii-66k. It is intended for fine-tuning OpenAI Privacy Filter with a hybrid label space: the standard OPF v2 labels plus Russian PII categories that do not have direct standard OPF equivalents. See label_space.json. Files data/train.jsonl: 59087 records data/validation.jsonl: 6565 records label_space.json: OPF custom label space The split was created from source split… See the full description on the dataset page: https://huggingface.co/datasets/bbeglerov/russian-pi-66k-opf.
huggingface130 downloads - 0.50TheoDB/french-pii-evalpkg:data/TheoDB/french-pii-eval
French PII Evaluation Dataset A curated French PII detection evaluation and training dataset, built for benchmarking TheoDB/privacy-filter-fr. Dataset Structure Split Examples Purpose test.jsonl 2,500 Held-out evaluation — never used in training test_english.jsonl 426 English regression check train.jsonl 57,248 Training data val.jsonl 500 Validation data Label Taxonomy 8 PII classes (same as openai/privacy-filter): Class Test… See the full description on the dataset page: https://huggingface.co/datasets/TheoDB/french-pii-eval.
huggingfacetoken-classification117 downloads - 0.50cicero-im/analysis_resultspkg:data/cicero-im/analysis_results
Dataset Card for Analysis Results This dataset contains synthetic text samples generated to evaluate the quality of PII masking. The samples are rated on factors like incorporation, structure, consistency and richness. The dataset can be used to train and evaluate models for PII detection and masking, and to analyze the trade-offs between data utility and privacy. Dataset Structure The dataset consists of generated text samples containing PII (Personally Identifiable… See the full description on the dataset page: https://huggingface.co/datasets/cicero-im/analysis_results.
huggingface104 downloads - 0.50disi-unibo-nlp/physionet-deid-i2b2-2014pkg:data/disi-unibo-nlp/physionet-deid-i2b2-2014
The De-identification dataset contains medical text records with Named Entity Recognition (NER) annotations. The dataset is processed to split records into individual sentences while preserving entity annotations. Each sentence is tokenized and annotated in IOB format for training NER models.
huggingfacetoken-classification75 downloads - 0.50akiFQC/japanese-confidential-information-extraction-sftpkg:data/akiFQC/japanese-confidential-information-extraction-sft
Japanese Confidential Information Extraction — SFT Dataset 日本語テキストから社外秘の固有表現を抽出するタスク向けの SFT (Supervised Fine-Tuning) データセットです。 LFM2 系モデルの LoRA fine-tune を想定して構築されています。 タスク概要 入力テキスト(日本語)に含まれる機密情報を、11カテゴリの JSON として抽出します。 入力: 「山田太郎(yamada@example.co.jp)から請求書番号 INV-2024-0042 で 売上 ¥12,800,000 の見積書が届いた。」 出力: { "address": [], "company_name": [], "email_address": ["yamada@example.co.jp"], "human_name": ["山田太郎"], "phone_number": [], "account_identifier":… See the full description on the dataset page: https://huggingface.co/datasets/akiFQC/japanese-confidential-information-extraction-sft.
huggingfacetext-generationother49 downloads - 0.50ai4privacy/phi-masking-100kpkg:data/ai4privacy/phi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Health Information (PHI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.
huggingfacetoken-classificationother49 downloads - 0.50auren-research/pii-shieldpkg:data/auren-research/pii-shield
PII Shield: Multilingual PII Detection Dataset PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information (PII) detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi— achieving the highest F1 on the SPY benchmark among open-source PII detectors. The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/pii-shield.
huggingfacetoken-classificationcc-by-4.049 downloads - 0.50ai4privacy/pdi-masking-100kpkg:data/ai4privacy/pdi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Digital Information (PDI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Digital Information (PDI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pdi-masking-100k.
huggingfacetoken-classificationother47 downloads - 0.50joneauxedgar/pasteproof-pii-dataset-v3pkg:data/joneauxedgar/pasteproof-pii-dataset-v3
PasteProof PII Dataset v3 Synthetic PII detection dataset with intentional confusion to prevent overfitting. What's Different in v3 Problem v2 v3 Fix Key names always match content apiKey → API_KEY 30% use generic names like data, x, field1 No lookalikes - Mixed real PII with fake lookalikes Always structured JSON/SQL/etc 20% raw PII without context Easy negatives Generic code Hard negatives (test cards, example.com emails) Generation… See the full description on the dataset page: https://huggingface.co/datasets/joneauxedgar/pasteproof-pii-dataset-v3.
huggingfacetoken-classificationmit47 downloads - 0.50LocalDoc/pii_ner_azerbaijanipkg:data/LocalDoc/pii_ner_azerbaijani
PII NER Azerbaijani Dataset Short, synthetic Azerbaijani dataset for PII-aware Named Entity Recognition (token classification). Useful for training and evaluating models that detect and localize personally identifiable information (PII) in Azerbaijani text. Note: All examples are synthetically generated with the library az-data-generator https://github.com/LocalDoc-Azerbaijan/az-data-generator. No real persons or contact details are included. Dataset Summary Each row… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/pii_ner_azerbaijani.
huggingfacetoken-classificationcc-by-4.046 downloads - 0.50shivaniachary123/pii-masking-health-phi-previewpkg:data/shivaniachary123/pii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Health & Medical Information (PHI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full access. Label Distribution Language Distribution European Coverage Full Dataset… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-health-phi-preview.
huggingfacetoken-classificationcc-by-4.046 downloads - 0.50JALAPENO11/model-inversion-adversarialpkg:data/JALAPENO11/model-inversion-adversarial
Model Inversion Adversarial Dataset 39,950 (original, anonymized) sentence pairs (target: 40,000) for black-box model inversion attack research against PII anonymization models. Each record contains the original PII-rich sentence and the BART-anonymized output produced by a fine-tuned BART-base anonymizer, along with rich metadata. Splits Split Count train 38,032 eval 1,918 total 39,950 Probing Strategies Strategy Count Purpose S1… See the full description on the dataset page: https://huggingface.co/datasets/JALAPENO11/model-inversion-adversarial.
huggingfacetext-generationmit46 downloads - 0.50micmadAAU/Nemotron-PIIpkg:data/micmadAAU/Nemotron-PII
Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/micmadAAU/Nemotron-PII.
huggingfacetoken-classificationcc-by-4.045 downloads - 0.50BoB14TeamSentinel/sentinel-kr-sensitive-entities-synthetic-v3pkg:data/BoB14TeamSentinel/sentinel-kr-sensitive-entities-synthetic-v3
Sentinel KR Sensitive Entities (Synthetic) v3 Overview Sentinel KR Sensitive Entities (Synthetic) v3 is a Korean synthetic (AI-generated) dataset for whitelist-only sensitive-entity detection in DLP / LLM guardrail scenarios. All sensitive values in this dataset (e.g., phone numbers, emails, IDs, tokens, keys) are artificially generated by AI and do not come from real individuals, real incidents, or collected private datasets. Any resemblance to real persons or real… See the full description on the dataset page: https://huggingface.co/datasets/BoB14TeamSentinel/sentinel-kr-sensitive-entities-synthetic-v3.
huggingfacetoken-classificationcc-by-4.044 downloads - 0.50ai4privacy/pfi-masking-100kpkg:data/ai4privacy/pfi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Financial Information (PFI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Financial Information (PFI) Masking Dataset, a… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pfi-masking-100k.
huggingfacetoken-classificationother43 downloads - 0.50ShmalexFlow/whiteout-compliance-benchmarkpkg:data/ShmalexFlow/whiteout-compliance-benchmark
Whiteout AI Compliance Benchmark A 15,915-prompt benchmark for evaluating AI compliance engines — systems that enforce content policies on user prompts before they reach AI providers. Built by Groovy Security for the Whiteout AI platform. Dataset Summary Property Value Total prompts 15,915 Categories 9 (PHI, PII, GDPR, Legal, Code, Confidential, Security, Finance, Education) Policies 74 across all categories Prompt types 3 (safe, violation… See the full description on the dataset page: https://huggingface.co/datasets/ShmalexFlow/whiteout-compliance-benchmark.
huggingfacetext-classificationapache-2.042 downloads - 0.50aniket-curlscape/pii-masking-english-5kpkg:data/aniket-curlscape/pii-masking-english-5k
Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-5k.
huggingfacetext-classificationother42 downloads - 0.50agentlans/personal-information-promptspkg:data/agentlans/personal-information-prompts
Personal Information Prompts This dataset contains multilingual prompts derived from the all_sample subset of the agentlans/allenai-WildChat-4.8M dataset. Each prompt features artificially inserted personally identifiable information (PII) generated randomly with the Faker Python package for various locales. Each rewritten prompt uses the google/gemma-3-12b-it model to incorporate the synthetic personal data. Dataset fields for the two configurations: classification… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/personal-information-prompts.
huggingfacetext-classificationcc-by-4.041 downloads - 0.50Wismut/nym-pii-multilingual-datapkg:data/Wismut/nym-pii-multilingual-data
nym-pii-multilingual-data 805,000 synthetic, exactly-labeled PII token-classification examples across ~23 languages and 6 scripts, built for training nym's PII detection models (e.g. Wismut/nym-pii-multilingual). Format JSONL with character-offset spans (offsets index into text as UTF-8 — compatible with HF fast-tokenizer offset_mapping): {"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.", "entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.
huggingfacetoken-classificationmit40 downloads - 0.50huggingbahl21/saha-alpkg:data/huggingbahl21/saha-al
SAHA-AL: PII Anonymization Benchmark SAHA-AL is a benchmark for training and evaluating text anonymization systems. It goes beyond detection accuracy by evaluating anonymization as a system under attack — measuring adversarial re-identification risk, contextual privacy leakage, and a formalized privacy-utility tradeoff. Key Features 3 evaluation tasks: PII detection, text anonymization quality, and adversarial privacy risk 11 metrics spanning leakage, utility, format… See the full description on the dataset page: https://huggingface.co/datasets/huggingbahl21/saha-al.
huggingfacetext-generationmit40 downloads - 0.50mukuls9971/indian-address-v1pkg:data/mukuls9971/indian-address-v1
Indian Address Synthetic Dataset v1 Synthetic multilingual Indian-address token-classification dataset generated by the pii-model-oss project. Repository Dataset repo: mukuls9971/indian-address-v1 Train split: 12000 Validation split: 1000 Test split: 1000 Files train.jsonl validation.jsonl test.jsonl report.json Notes Generated and published by the pii-model-oss workflow. Upstream datasets used to assemble benchmark variants retain their own… See the full description on the dataset page: https://huggingface.co/datasets/mukuls9971/indian-address-v1.
huggingfacetoken-classificationmit39 downloads - 0.50ai4privacy/pwi-masking-100kpkg:data/ai4privacy/pwi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Work Information (PWI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Work Information (PWI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pwi-masking-100k.
huggingfacetoken-classificationother37 downloads - 0.50shivaniachary123/sovereign-pii-detection-v1pkg:data/shivaniachary123/sovereign-pii-detection-v1
🛡️ Sovereign PII Detection Dataset (v1.0) Maintainer: Cata Risk Lab | Project: Wattle Guard 🌍 Dataset Summary This synthetic dataset contains labeled examples of Sovereign Identity Markers specific to the Swiss, UK, and Australian jurisdictions. It is designed to train and benchmark the Wattle Guard redaction engine, ensuring compliance with cross-border data protection laws (nFADP, UK GDPR, Privacy Act 1988). Unlike generic PII datasets that focus on US data (SSN)… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/sovereign-pii-detection-v1.
huggingfacetoken-classificationmit36 downloads - 0.50575-lab/kiji-inspector-reviewed-pairspkg:data/575-lab/kiji-inspector-reviewed-pairs
Kiji PII Detection Training Data Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution. Dataset Summary Samples 99,990 (train: 89,991, test: 9,999) Languages 6 (Dutch, Spanish, German, English, Danish, French) Countries 20 PII entity types 26 Total entity annotations 814,306 (avg 8.1 per sample) Coreference clusters 142,142 (99% of… See the full description on the dataset page: https://huggingface.co/datasets/575-lab/kiji-inspector-reviewed-pairs.
huggingfacetoken-classificationapache-2.035 downloads - 0.50tugrulkaya/turkish-pii-datasetpkg:data/tugrulkaya/turkish-pii-dataset
🔒 Turkish PII Detection Dataset Türkçe metinlerde Kişisel Tanımlanabilir Bilgi (PII) tespiti için el ile etiketlenmiş, araştırma amaçlı bir NER veri kümesi. KVKK ve GDPR uyumlu yapay zeka geliştirme için temel bir kaynak olarak tasarlanmıştır. Veri Kümesi Özeti Örnek sayısı: ~20 etiketlenmiş metin (küçük ölçekli, başlangıç seviyesi) Kategoriler: 7+ PII türü Dil: Türkçe Format: JSON / token-level etiketleme Amaç: Eğitim, araştırma, anonimleştirme prototipleri… See the full description on the dataset page: https://huggingface.co/datasets/tugrulkaya/turkish-pii-dataset.
huggingfacetoken-classificationcc-by-4.035 downloads - 0.50shivaniachary123/pii-detection-corpuspkg:data/shivaniachary123/pii-detection-corpus
PII Detection Corpus Synthetic dataset of text samples containing labeled PII (Personally Identifiable Information) for testing and benchmarking PII detection/scrubbing tools. Fields text: Text sample containing PII pii_type: Category of PII (email, phone, ssn, credit_card, ip, dob, address, passport, api_key, name, iban) pii_value: The exact PII string in the text start: Character offset start end: Character offset end context: Surrounding context category (medical… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-detection-corpus.
huggingfacetoken-classificationmit35 downloads - 0.50anony-mouse123/Instruction_recall_datasetpkg:data/anony-mouse123/Instruction_recall_dataset
CanaryBench-PII Frequency-aware canary injection benchmark for auditing memorization in finetuned language models, built on the AI4Privacy PII reconstruction task. Dataset Description This dataset is part of CanaryBench, a benchmark for evaluating memorization in finetuned language models across repetition tiers and privacy regimes. Frequency tiers: 1×, 10×, 50× PII types: EMAIL, PHONE Member canaries: 770 Reference canaries: 1000 Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.
huggingfacetext-generationcc-by-4.034 downloads - 0.50vstantch/x402-pii-corpuspkg:data/vstantch/x402-pii-corpus
x402 PII Metadata Corpus Synthetic labelled corpus of 2,000 x402 payment metadata triples for PII filter evaluation. Released alongside the paper "Hardening x402: Privacy-Preserving Agentic Payments via Pre-Execution Metadata Filtering". Paper: arXiv:2604.11430 [cs.CR] Canonical archive: IEEE DataPort doi:10.21227/kpsz-nq73 Code: presidio-v/presidio-hardened-x402 Dataset description Each record represents one x402 payment metadata triple (resource_url, description… See the full description on the dataset page: https://huggingface.co/datasets/vstantch/x402-pii-corpus.
huggingfacetoken-classificationmit34 downloads - 0.50UniDataPro/synthetic-printed-australian-passportspkg:data/UniDataPro/synthetic-printed-australian-passports
Australian passport dataset The dataset comprises 5,000 high-resolution synthetic photos of ** Australian passports**, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information. This dataset is an essential tool for organizations and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-australian-passports.
huggingfaceimage-to-textcc-by-nc-nd-4.033 downloads - 0.50ai4privacy/pii-masking-openpii-1.5mpkg:data/ai4privacy/pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.
huggingfacetoken-classificationother30 downloads - 0.50EdyVision/pii-skills-ablation-resultspkg:data/EdyVision/pii-skills-ablation-results
PII Skills Ablation — Scored Results This repository contains model predictions and evaluation scores for the ablation study described in: "Asymmetry, Not Capability: Evaluation Shapes Tool-Augmented PII Detection in Small Language Models" Results are produced by running four open-weight instruction-tuned models (Gemma 2 9B, Llama 3.1 8B, Mistral 7B, Qwen 2.5 7B) under four primary conditions (zero-shot, +Docs, +Tool, +Skills), plus three baselines (standalone PII-Codex detector… See the full description on the dataset page: https://huggingface.co/datasets/EdyVision/pii-skills-ablation-results.
huggingfacemit29 downloads - 0.50orgrctera/pii_masking_300k_information_extractionpkg:data/orgrctera/pii_masking_300k_information_extraction
PII Masking 300k — Information Extraction Dataset summary This repository hosts a validation sample of the PII Masking 300k benchmark for the information extraction track: models must identify personally identifiable information (PII) in text and produce structured extractions (slot-filling JSON), optional token-level BIO labels, and span-based annotations for masking or redaction workflows. The full PII Masking 300k suite is designed to stress-test privacy-preserving… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/pii_masking_300k_information_extraction.
huggingfacetoken-classificationapache-2.029 downloads - 0.50EdyVision/pii-skills-ablationpkg:data/EdyVision/pii-skills-ablation
PII Skills Ablation Benchmark This repository provides the benchmark and experiment configuration for the ablation study described in: "Asymmetry, Not Capability: Evaluation Shapes Tool-Augmented PII Detection in Small Language Models" The benchmark is a stratified sample of text with ground-truth PII spans aligned to PII-Codex canonical types. It is used to evaluate whether zero-shot prompting, documentation injection (+Docs), tool access (+Tool), or skills injection (+Skills)… See the full description on the dataset page: https://huggingface.co/datasets/EdyVision/pii-skills-ablation.
huggingfacemit27 downloads - 0.50ahczhg/Nemotron-PIIpkg:data/ahczhg/Nemotron-PII
Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/ahczhg/Nemotron-PII.
huggingfacetoken-classificationcc-by-4.026 downloads - 0.50NAMANDREWLV/pii-masking-95k-preencodedpkg:data/NAMANDREWLV/pii-masking-95k-preencoded
VI PII Masking (Pre-encoded) – 95k Private Vietnamese dataset for PII detection and masking, designed for token classification / NER in privacy-preserving NLP systems. Dataset Overview Language: Vietnamese (vi) Domain: Privacy / PII masking Total samples: ~95,000 Splits: Train: 76,097 Validation: 9,512 Test: 9,513 Format: JSONL Data Fields Raw & masked text source_text masked_text privacy_mask language region script split uid… See the full description on the dataset page: https://huggingface.co/datasets/NAMANDREWLV/pii-masking-95k-preencoded.
huggingfacetoken-classificationother25 downloads - 0.50Keler-Health/turkish-medical-deid-evalpkg:data/Keler-Health/turkish-medical-deid-eval
Turkish Medical De-Identification Evaluation Corpus (Synthetic) A labelled benchmark for evaluating the removal of personally identifiable information (PII) from Turkish medical speech-to-text (STT) transcripts. This dataset contains no real data. Every consultation, name, phone number, address, identifier and financial detail is programmatically generated and fictitious. The corpus exists specifically so that de-identification systems can be evaluated without any real patient… See the full description on the dataset page: https://huggingface.co/datasets/Keler-Health/turkish-medical-deid-eval.
huggingfacetoken-classificationcc-by-nc-4.024 downloads - 0.50Ari-S-123/better-english-pii-anonymizerpkg:data/Ari-S-123/better-english-pii-anonymizer
PII Detection Combined Dataset Combined dataset for PII (Personally Identifiable Information) detection, merging the ai4privacy English-only subset with synthetically generated challenging examples targeting NER failure modes. Dataset Description This dataset combines two sources: ai4privacy/open-pii-masking-500k (English subset): 120,533 train / 30,160 test examples Synthetic data (Grok-4.1-Non-reasoning generated/GPT-5.1 validated): 4,801 train / 1,201 test examples… See the full description on the dataset page: https://huggingface.co/datasets/Ari-S-123/better-english-pii-anonymizer.
huggingfacetoken-classificationmit22 downloads - 0.50Vigil-ai/governance-bench-v1pkg:data/Vigil-ai/governance-bench-v1
governance-bench-v1 A 200-case curated benchmark for evaluating AI governance scanners across three categories: PII handling, hallucination on factual / legal / medical / financial prompts, and prompt-injection resistance. Designed to be small enough to run by hand, structured enough to grade automatically, and varied enough to exercise the failure modes that real governance tooling needs to catch. What it tests Category Cases What it measures PII 70… See the full description on the dataset page: https://huggingface.co/datasets/Vigil-ai/governance-bench-v1.
huggingfacetext-classificationcc-by-4.020 downloads - 0.50Marawanelbalal/pii-phi-corpus-holdoutpkg:data/Marawanelbalal/pii-phi-corpus-holdout
Multilingual PII/PHI De-identification Corpus (Eval-Safe) This is a correction of Marawanelbalal/pii-phi-corpus-final. That release folded all of OpenPII's validation-split rows into train, which contaminates any evaluation of an OpenPII-style generalization question for a model trained on it — the model has simply memorized the exact rows an evaluator would otherwise use to test it. This release fixes that by respecting OpenPII's own train/validation split boundary instead of… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/pii-phi-corpus-holdout.
huggingfaceother19 downloads - 0.50Cata-Risk-Lab/sovereign-pii-detection-v1pkg:data/Cata-Risk-Lab/sovereign-pii-detection-v1
🛡️ Sovereign PII Detection Dataset (v1.0) Maintainer: Cata Risk Lab | Project: Wattle Guard 🌍 Dataset Summary This synthetic dataset contains labeled examples of Sovereign Identity Markers specific to the Swiss, UK, and Australian jurisdictions. It is designed to train and benchmark the Wattle Guard redaction engine, ensuring compliance with cross-border data protection laws (nFADP, UK GDPR, Privacy Act 1988). Unlike generic PII datasets that focus on US data (SSN)… See the full description on the dataset page: https://huggingface.co/datasets/Cata-Risk-Lab/sovereign-pii-detection-v1.
huggingfacetoken-classificationmit19 downloads - 0.50Marawanelbalal/pii-phi-corpus-finalpkg:data/Marawanelbalal/pii-phi-corpus-final
Multilingual PII/PHI De-identification Corpus (Final) This is the finalized, training-ready corpus for a multilingual PII/PHI de-identification (token classification) model scoped to en-GB, nl-BE, fr-FR. It merges two sources: Our own synthetic clinical/administrative corpus (Marawanelbalal/synthetic-pii-phi-v2), unchanged. A normalized subset of ai4privacy/pii-masking-openpii-1.5m (English/French/Dutch rows only), added as generic-register training text. development/test are… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/pii-phi-corpus-final.
huggingfaceother18 downloads - 0.50KhalidAlharbi377/pii-detection-multisource-en-saudi-arabicpkg:data/KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic
PII Detection Multisource EN + Saudi/Arabic 284,619 English examples. 2,088,335 labelled spans. 31 entity types. One label space. Four public PII datasets, merged into a single schema, plus Saudi and Arabic coverage that none of them had, plus material for two failure modes that matter when you run redaction in production. Built for OnKith, a privacy first voice assistant that transcribes speech and strips personal information on the device itself, before anything is allowed to… See the full description on the dataset page: https://huggingface.co/datasets/KhalidAlharbi377/pii-detection-multisource-en-saudi-arabic.
huggingfacetoken-classificationcc-by-4.017 downloads - 0.50joelbarmettler/gheim-ch-pii-212kpkg:data/joelbarmettler/gheim-ch-pii-212k
gheim-ch-pii-212k Summary. 212,503-chunk multilingual PII NER dataset covering the four official Swiss languages and English. 84% is real text from the Apertus pretrain corpora (Swiss court rulings, federal parliament records, Swiss-filtered web text, Romansh corpus); the remaining 16% is template- and LLM-generated synthetic prose used to populate cells where real-text coverage was insufficient. Annotations are machine generated by three independent open-weights LLMs (Gemma… See the full description on the dataset page: https://huggingface.co/datasets/joelbarmettler/gheim-ch-pii-212k.
huggingfacetoken-classificationcc-by-4.017 downloads - 0.50ppalani09/pii-masking-300kpkg:data/ppalani09/pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ppalani09/pii-masking-300k.
huggingfacetext-classificationother16 downloads - 0.50darkmatter2222/redact-v1pkg:data/darkmatter2222/redact-v1
Redact-v1 Synthetic Data Card Overview This dataset consists of 100% synthetic data—every element is artificially generated. No data originates from any genuine or external source. The full repository, including synthetic data generation models and use cases, is available on GitHub: https://github.com/darkmatter2222/NLU-Redact-PII. Categories of Synthetic Sensitive Data The dataset includes the following categories of artificially generated sensitive data:… See the full description on the dataset page: https://huggingface.co/datasets/darkmatter2222/redact-v1.
huggingfaceapache-2.015 downloads - 0.50Pranshurs/ylemis-india-pii-benchmarkpkg:data/Pranshurs/ylemis-india-pii-benchmark
Ylemis India-PII Benchmark v1 An open, reproducible benchmark for span-based PII detection in Indian and multilingual business text. The dataset is fully synthetic and is intended for evaluation, regression testing, and transparent comparison—not model training and not as a substitute for a customer-specific pilot. Release facts 10,000 cases, all synthetic; no intended real-person PII. 16 entity types and 14 language/format slices. 1,756 negative cases: 810 clean… See the full description on the dataset page: https://huggingface.co/datasets/Pranshurs/ylemis-india-pii-benchmark.
huggingfacetoken-classificationcc0-1.014 downloads - 0.50woojin1069/pii-masking-openpii-1.5mpkg:data/woojin1069/pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/woojin1069/pii-masking-openpii-1.5m.
huggingfacetoken-classificationother14 downloads - 0.50dbabis/20NG_5topics_PII_annotatedpkg:data/dbabis/20NG_5topics_PII_annotated
20 Newsgroups (5 Topics) — PII-Augmented version Description This dataset is a curated subset of the 20 Newsgroups corpus, containing 5 clearly distinguishable topics for experimentation with intelligent text anonymization and topic classification It was created as part of the Bachelor’s thesis “Intelligent anonymization for natural language processing and inference” at FIIT STU, 2025 Versions A. 20NG_5topics.jsonl Original subset with 5 selected… See the full description on the dataset page: https://huggingface.co/datasets/dbabis/20NG_5topics_PII_annotated.
huggingfacetoken-classificationmit14 downloads - 0.50klusai/ds-kp-general-de-50kpkg:data/klusai/ds-kp-general-de-50k
klusai/ds-kp-general-de-50k KlusAI Privacy (KP) dataset — general domain, language de. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the de LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) de. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-de-50k.
huggingfacetoken-classificationcc-by-4.013 downloads - 0.50felipe53/slm-deid-name-judgmentpkg:data/felipe53/slm-deid-name-judgment
Dataset card — v3 (SFT training data) The data sft-v3-mps trained on. Built 2026-07-08 to fix v2's recall/consistency regression by rebalancing toward person-use and scaling up. Supersedes docs/dataset-card-v2.md. What it is Short educational passages (essay + tutoring-dialogue), each byte-identical to its input except that real personal names are wrapped in ⟨NAME⟩…⟨/NAME⟩. The bulk is matched minimal pairs: the same ambiguous surface used once as a person… See the full description on the dataset page: https://huggingface.co/datasets/felipe53/slm-deid-name-judgment.
huggingfacetoken-classificationapache-2.012 downloads - 0.50Roblox/roblox-pii-classifier-benchmarkpkg:data/Roblox/roblox-pii-classifier-benchmark
Roblox PII Safety for Chat Benchmark Overview The Roblox PII Safety for Chat Benchmark is an evaluation-only dataset for detecting when a target speaker asks for personal information, shares it, or directs another user off-platform. Unlike isolated-text benchmarks, it evaluates these behaviors using surrounding multi-turn, multi-speaker context. It accompanies Roblox PII Classifier v2.0. The test split contains 39,202 fully synthetic English conversations—24,518… See the full description on the dataset page: https://huggingface.co/datasets/Roblox/roblox-pii-classifier-benchmark.
huggingfacetext-classificationapache-2.011 downloads - 0.50roei-ar/AdvPIIBenchpkg:data/roei-ar/AdvPIIBench
Dataset Card for AdvPIIBench A benchmark of 104,728 synthetic-but-realistic user prompts for measuring how PII detection degrades under inference-time adversarial attack. Every prompt carries character-accurate span labels, so the same corpus supports both document-level (was anything flagged?) and span-level (was the right text flagged?) evaluation. All PII in this dataset is synthetic. Identifiers are generated by Faker via Microsoft Presidio; no value refers to a real… See the full description on the dataset page: https://huggingface.co/datasets/roei-ar/AdvPIIBench.
huggingfacetoken-classificationcc-by-4.011 downloads - 0.50klusai/ds-kp-general-es-50kpkg:data/klusai/ds-kp-general-es-50k
klusai/ds-kp-general-es-50k KlusAI Privacy (KP) dataset — general domain, language es. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the es LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) es. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-es-50k.
huggingfacetoken-classificationcc-by-4.011 downloads - 0.50promptrails/piimask-eval-tr-enpkg:data/promptrails/piimask-eval-tr-en
piimask-eval-tr-en An independent, bilingual (Turkish + English) benchmark for PII detection / redaction, built to be hard: it mixes languages, spans 95 realistic domains, and seeds adversarial distractors — product codes, reference numbers, and other PII-shaped strings designed to trigger false positives. 400 documents · 2,584 entities · difficulty-labelled. Why this set Most PII benchmarks are English-first, single-domain, or easy. This one was constructed… See the full description on the dataset page: https://huggingface.co/datasets/promptrails/piimask-eval-tr-en.
huggingfacetoken-classificationcc-by-4.010 downloads - 0.50klusai/ds-kp-general-it-50kpkg:data/klusai/ds-kp-general-it-50k
klusai/ds-kp-general-it-50k KlusAI Privacy (KP) dataset — general domain, language it. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the it LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) it. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-it-50k.
huggingfacetoken-classificationcc-by-4.010 downloads - 0.50klusai/ds-kp-general-nl-50kpkg:data/klusai/ds-kp-general-nl-50k
klusai/ds-kp-general-nl-50k KlusAI Privacy (KP) dataset — general domain, language nl. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the nl LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) nl. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-nl-50k.
huggingfacetoken-classificationcc-by-4.010 downloads - 0.50enosislabs/aether-privacy-datasetpkg:data/enosislabs/aether-privacy-dataset
MaTE X Privacy Sentinel Dataset Synthetic token-classification dataset for training a local privacy/security filter based on OpenAI Privacy Filter. Purpose This dataset teaches a local filter to detect and redact sensitive spans in developer workflows before context is sent to external LLMs. Target domains include: .env files terminal logs stack traces git diffs GitHub issues and PR comments agent traces tool outputs workspace memory auth, database, cloud and… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/aether-privacy-dataset.
huggingfacetoken-classificationapache-2.09 downloads - 0.50klusai/ds-kp-general-fr-50kpkg:data/klusai/ds-kp-general-fr-50k
klusai/ds-kp-general-fr-50k KlusAI Privacy (KP) dataset — general domain, language fr. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the fr LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) fr. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-fr-50k.
huggingfacetoken-classificationcc-by-4.09 downloads - 0.50paperboy-ai/desktop-pii-210pkg:data/paperboy-ai/desktop-pii-210
Desktop PII 210 This directory is a local Hugging Face-compatible dataset package for synthetic desktop screenshots with expected privacy and utility QA pairs. The screenshots are generated synthetic desktop scenes. The visible sensitive values are fictional benchmark strings, not real personal data. Contents images/: all 210 generated PNG screenshots. images/metadata.jsonl: Hugging Face imagefolder metadata, one row per image. data/train.parquet: the default Hugging… See the full description on the dataset page: https://huggingface.co/datasets/paperboy-ai/desktop-pii-210.
huggingfaceimage-to-textother9 downloads - 0.50Shayfra7926/PANOPTICONpkg:data/Shayfra7926/PANOPTICON
PANOPTICON Dataset Summary PANOPTICON (PII-based Assemblage of Naturalistic Output–Prompt Tuples for Investigating Privacy Leakage in Conversational AI) is a dataset of synthetic, PII-bearing prompts designed to enable controlled evaluation of privacy leakage / prompt inversion behaviors in LLMs. The dataset is organized by high-level Category and Scenario, and includes fields that support separating PII spans from surrounding benign context for analysis.… See the full description on the dataset page: https://huggingface.co/datasets/Shayfra7926/PANOPTICON.
huggingfacetext-generationother8 downloads - 0.50schift-io/korean-pii-benchmark-v3pkg:data/schift-io/korean-pii-benchmark-v3
Korean PII Benchmark v3 473 cases for evaluating Korean PII (person, organization, address) detection models. Covers court judgments, news articles, administrative documents, medical/insurance records, transcripts, and negative examples (law names, case numbers, titles that should NOT be detected). Entity Distribution Entity Spans private_person 411 private_organization 182 private_address 141 private_phone 65 (negative) 59… See the full description on the dataset page: https://huggingface.co/datasets/schift-io/korean-pii-benchmark-v3.
huggingfacetoken-classificationcc-by-sa-4.07 downloads - 0.50alrosait/pii-synthetic-rupkg:data/alrosait/pii-synthetic-ru
PII Synthetic Dataset (Russian) — pii-synthetic-ru Синтетический датасет для обучения NER-детектора персональных данных на русском языке. Содержит 4 500 примеров с аннотациями сущностей NAME (ФИО) и ADDRESS (адрес), а также негативные примеры без ПД. Создан в рамках проекта PIIDetector — гибридного детектора ПД для русского языка на базе Microsoft Presidio. Статистика Метрика Значение Всего примеров 4 500 Только NAME 1 113 Только ADDRESS 952 NAME… See the full description on the dataset page: https://huggingface.co/datasets/alrosait/pii-synthetic-ru.
huggingfacetoken-classificationmit6 downloads - 0.50aibotjock/RedactionBenchpkg:data/aibotjock/RedactionBench
Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/aibotjock/RedactionBench.
huggingfacetoken-classificationcc-by-4.05 downloads - 0.50NikolaiSachok/strata-insurance-corpuspkg:data/NikolaiSachok/strata-insurance-corpus
Strata Insurance Corpus A reproducible, fully synthetic, multi-format insurance document corpus for a fictional pan-European property-&-casualty insurer, Meridian Mutual, shipped with a golden evaluation set produced by construction. Built to exercise and benchmark document-RAG systems on enterprise-shaped data — born-digital and scanned PDFs, Word documents, spreadsheets, and photos — with trustworthy ground truth. Everything here is synthetic. No real persons, companies, or… See the full description on the dataset page: https://huggingface.co/datasets/NikolaiSachok/strata-insurance-corpus.
huggingfacequestion-answeringcc-by-4.05 downloads - 0.50asgdaudahu/pii-bench-zhpkg:data/asgdaudahu/pii-bench-zh
PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/asgdaudahu/pii-bench-zh.
huggingfacetoken-classificationapache-2.05 downloads - 0.50lucianfialho/privacy-filter-br-datasetpkg:data/lucianfialho/privacy-filter-br-dataset
privacy-filter-br Dataset Synthetic Portuguese (BR) PII detection dataset used to fine-tune the privacy-filter-br NER model. 22 PII categories, BIOES tagging compatible. Latest: v8.1 main aponta sempre pra última versão estável. Hoje: v8.1 (172075 train + 17072 holdout). from datasets import load_dataset ds = load_dataset("lucianfialho/privacy-filter-br-dataset") # ou pin: load_dataset("lucianfialho/privacy-filter-br-dataset", revision="v8.1") Schema… See the full description on the dataset page: https://huggingface.co/datasets/lucianfialho/privacy-filter-br-dataset.
huggingfacetoken-classificationapache-2.05 downloads - 0.50Reza2kn/persian-pii-masking-openpii-690k-initial-cleanpkg:data/Reza2kn/persian-pii-masking-openpii-690k-initial-clean
Persian PII-Masking Initial-Round Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo contains the cleaned initial-clean artifact. Audit/pruning summary: Raw rows: 225418 Kept rows: 224956 Hard-excluded rows: 0 Dropped rows from full exact nearest-neighbor components at cosine >= 0.95: 462 Rows with a full-dataset nearest neighbor at cosine >= 0.95 before component pruning: 867 Full-NN fraction before component pruning: 0.003846… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-initial-clean.
huggingfacetoken-classificationcc-by-4.04 downloads - 0.50pbhappliedsystems/veritruct-cloud-regulated-deid-1kpkg:data/pbhappliedsystems/veritruct-cloud-regulated-deid-1k
Veritruct Cloud — Construction-True Regulated-Domain De-Identification Dataset Run run_cloud_20260706_000447 · Flagship release demo dataset · PBH Applied Systems, LLC Generated, gated, masked, labeled, and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · Quality-Gated Synthetic Data · LLM Optimization & Deployment 📄 Read the whitepaper: Veritruct: Quality-Gated Synthetic Data Generation for Regulated Industries, or for a quick overview, read… See the full description on the dataset page: https://huggingface.co/datasets/pbhappliedsystems/veritruct-cloud-regulated-deid-1k.
huggingfacetoken-classificationcc-by-4.01 downloads - 0.50A10Networks/RedactionBenchpkg:data/A10Networks/RedactionBench
Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/A10Networks/RedactionBench.
huggingfacetoken-classificationcc-by-4.01 downloads - 0.50mike8355/Nemotron-PIIpkg:data/mike8355/Nemotron-PII
Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas… See the full description on the dataset page: https://huggingface.co/datasets/mike8355/Nemotron-PII.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50mapo80/aliasit-pii-dataset-v3-smallpkg:data/mapo80/aliasit-pii-dataset-v3-small
aliasit-pii-dataset-v3-small Half the training split of aliasit-pii-dataset-v3, with validation and test kept whole. For cheap training runs that still evaluate on the real thing. This is a subsample of mapo80/aliasit-pii-dataset-v3. The training split is halved by uniform sampling on document id; validation and test are the full splits of the parent dataset, unchanged. It exists to make a training run cheap without making the evaluation a different question. Every category… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v3-small.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50mapo80/aliasit-pii-dataset-v3pkg:data/mapo80/aliasit-pii-dataset-v3
aliasit-pii-dataset-v3 Multilingual PII span-annotation dataset, 79 entity types, 30 pinned sources, Italian-first. Text plus character spans, single-label BIO. Documents 182,862 Entities 3,172,535 Entity types 79 (46 marked critical) Languages ar, de, el, en, es, fr, it, nl, pt, sl, tr Taxonomy v3.0.0 — taxonomy.yaml ships in this repo Revision 6.0.0 Representation text + character spans, end exclusive, non-overlapping Distinct sources 27, each pinned… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v3.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50mapo80/aliasit-pii-dataset-v2pkg:data/mapo80/aliasit-pii-dataset-v2
aliasit-pii-dataset-v2 Dataset PII multilingue in formato canonico testo + character span, 82 categorie, 6 lingue, derivato da quattro sorgenti pinnate a commit. documenti 170.448 entita' 3.127.084 categorie 82 (48 marcate critiche) lingue de, en, es, fr, it, pt tassonomia v1.1.0 (taxonomy.yaml nel repo) versione 2.0.0 costruito il 2026-08-26T17:19:46+00:00 Formato La rappresentazione canonica e' testo + character span, non BIO: gli… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v2.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50mapo80/aliasit-pii-dataset-v2-smallpkg:data/mapo80/aliasit-pii-dataset-v2-small
aliasit-pii-dataset-v2-small Sottoinsieme stratificato di aliasit-pii-dataset-v2 per le prove di training: stesse categorie e stesso formato, dimensioni da smoke test. documenti 8.620 entita' 124.644 categorie 82 (48 marcate critiche) lingue de, en, es, fr, it, pt tassonomia v1.1.0 (taxonomy.yaml nel repo) versione 1.0.0 costruito il 2026-08-26T17:36:35+00:00 Formato La rappresentazione canonica e' testo + character span, non BIO: gli… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v2-small.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50barissozudogru/piiscope-benchmarkpkg:data/barissozudogru/piiscope-benchmark
Piiscope Structured PII Pattern Benchmark A deterministic, privacy-safe benchmark for structured personal-data detectors. Every value is synthetic, reserved for documentation, or a published test credential. The dataset contains no records collected from people, no customer data, and no transactable financial identifiers. The benchmark is maintained with Piiscope, a local PII scanner and privacy-risk CLI. It can also evaluate compatible rule-based detectors that return one or… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/piiscope-benchmark.
huggingfacetext-classificationcc-by-4.00 downloads - 0.50xorushi/roberta-pii-synthpkg:data/xorushi/roberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.
huggingfacetoken-classificationmit0 downloads - 0.50Werea-co/Werea-KVKK-Benchpkg:data/Werea-co/Werea-KVKK-Bench
Werea-KVKK-Bench v2 Synthetic, reproducible data and benchmark assets for evidence-first Turkish privacy operations. Contents pii_synthetic.jsonl: fictional Turkish PII spans; workflow_sft.jsonl: structured workflow-assistant conversations; cases.jsonl: frozen rules/deadline/PII evaluation cases; kvkk_rules.json: human-approval workflow definitions; kvkk_controls.json: evidence-based institutional control catalog; pii_taxonomy.json: Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Werea-co/Werea-KVKK-Bench.
huggingfacetoken-classificationapache-2.00 downloads - 0.50akaruineko/privasetpkg:data/akaruineko/privaset
PRIVAset: A Multilingual Synthetic PII Detection Dataset PRIVAset is a large-scale, privacy-safe, synthetic dataset for training and evaluating Personally Identifiable Information (PII) filtering systems. It supports 25 languages, includes 16 PII types, and is designed for tasks such as classification, named-entity recognition (NER), and redaction. All data is generated using Faker (locale-aware) and custom synthetic logic. No real PII is used, making it safe for public release… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/privaset.
huggingfacemit0 downloads - 0.50Musayusuf001/Synthetic-Nigerian-AI4Privacy-Datasetpkg:data/Musayusuf001/Synthetic-Nigerian-AI4Privacy-Dataset
Synthetic Nigerian AI4Privacy Dataset A synthetic Nigerian-focused dataset for Personally Identifiable Information (PII) detection, masking, and anonymization. Dataset Statistics Split Examples Full dataset 43,501 Train 34,800 Validation 4,350 Test 4,351 Annotated entities: 354,236 Entity types: 67 Entity Types Entity Count FIRSTNAME 64,282 LASTNAME 52,888 CITY 17,638 STATE 10,880 USERNAME 9,508 STREET 8… See the full description on the dataset page: https://huggingface.co/datasets/Musayusuf001/Synthetic-Nigerian-AI4Privacy-Dataset.
huggingfacetoken-classificationmit0 downloads - 0.50Marawanelbalal/synthetic-pii-phi-v1pkg:data/Marawanelbalal/synthetic-pii-phi-v1
Synthetic PII/PHI v1 Synthetic clinical/administrative documents in three locales (en-GB, nl-BE, fr-FR) with span-level PII/PHI annotations, generated for training a multilingual de-identification (token classification) model. All entity values (names, IDs, dates, etc.) are synthetically generated — no real patient, provider, or organization data is present anywhere in this dataset. Splits split rows train 68,668 development 4,017 test 3,978… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/synthetic-pii-phi-v1.
huggingfaceother0 downloads - 0.50bbanany/korean-pii-candidate-classificationpkg:data/bbanany/korean-pii-candidate-classification
Korean RAG PII Span Classification SFT 한국어 문서 기반 RAG 응답에서 지정된 candidate span이 특정 자연인에게 연결되는 정보인지 PII 또는 NOT_PII로 분류하도록 만든 지도 미세조정 (SFT) 데이터셋입니다. 각 레코드는 Qwen 계열 채팅 형식의 system, user, assistant 메시지 세 개로 구성됩니다. 이 저장소는 bbanany/final-mode 모델의 미세조정과 평가에 실제 사용한 토크나이징 직전 JSONL 분할을 보존합니다. Dataset structure Split File Rows Use train data/train.clean.jsonl 15,875 QLoRA training validation data/valid.clean.jsonl 1,984 Training-time evaluation and checkpoint… See the full description on the dataset page: https://huggingface.co/datasets/bbanany/korean-pii-candidate-classification.
huggingfacetext-classificationother0 downloads - 0.50affjljoo3581/pii-masking-openpii-1m-enpkg:data/affjljoo3581/pii-masking-openpii-1m-en
OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification… See the full description on the dataset page: https://huggingface.co/datasets/affjljoo3581/pii-masking-openpii-1m-en.
huggingfacetoken-classificationother0 downloads - 0.50pl-pii-bench/pl-pii-benchpkg:data/pl-pii-bench/pl-pii-bench
pl-pii-bench Maintained by Anonimator.pl, a local-first Polish document anonymization tool. GitHub repository · Live benchmark results pl-pii-bench is an open benchmark for Polish personally identifiable information detection and text anonymization. It contains a fully synthetic, exhaustively annotated, document-level corpus and a separate open scoring harness. The corpus is an evaluation set, not training data. Please do not train on it. Every identifier is synthetic.… See the full description on the dataset page: https://huggingface.co/datasets/pl-pii-bench/pl-pii-bench.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50ahuseynli-17683/pii-masking-openpii-financepkg:data/ahuseynli-17683/pii-masking-openpii-finance
1. Overview Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety). 1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50cagrigungor/turkish-pii-masking-benchmarkpkg:data/cagrigungor/turkish-pii-masking-benchmark
Turkish PII Masking Benchmark (1,000 test cases) A hand-built, fully synthetic benchmark for evaluating instruction-conditional PII masking in Turkish. Designed to stress-test small LLMs (0.3B-1B) fine-tuned for privacy filtering in banking/ERP user messages. All values are synthetic (no real persons); all sentence patterns were written specifically for this benchmark (no training-set overlap). Task Given instruction (the masking policy) and input, the model must… See the full description on the dataset page: https://huggingface.co/datasets/cagrigungor/turkish-pii-masking-benchmark.
huggingfacetext-generationcc0-1.00 downloads - 0.50dr3x1/rizzo-pii-security-itpkg:data/dr3x1/rizzo-pii-security-it
rizzo-pii security IT — corpus sintetico del genere "sicurezza" Corpus sintetico italiano per la token classification di PII su documenti di sicurezza: verbali d'incidente, timeline forensi, ticket, estratti di log. Serve ad addestrare un modello che anonimizzi quei documenti in locale, prima di mandarli a un LLM esterno. È il dataset con cui è stato addestrato dr3x1/rizzo-pii-0.3B-security. Non contiene nessuna PII reale. Nessun dato personale vero, nessun indicatore di… See the full description on the dataset page: https://huggingface.co/datasets/dr3x1/rizzo-pii-security-it.
huggingfacetoken-classificationmit0 downloads - 0.505000dev/noirci-benchpkg:data/5000dev/noirci-bench
Noirci Bench, mesurer la protection des données personnelles en français Un benchmark de pseudonymisation française, livré avec son outil de mesure. L'intérêt n'est pas le corpus, c'est la métrique : elle change les conclusions. Publié par 5000.dev. Code du moteur : 5000dev/noirci. La métrique : un nom à moitié masqué est une fuite Les travaux publiés rapportent en général un F1 par token. Masquer « Jean » dans « Jean Dupont » y compte comme un demi-succès. Nous… See the full description on the dataset page: https://huggingface.co/datasets/5000dev/noirci-bench.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50Arsh9210/Privasis-Zeropkg:data/Arsh9210/Privasis-Zero
Privasis-Zero Dataset Description: Privasis-Zero is a large-scale synthetic dataset consisting of diverse text records—such as medical and financial records, legal documents, emails, and messages—containing rich, privacy-sensitive information. Each record includes synthetic profile details, surrounding social context, and annotations of privacy-related content. All data are fully generated using LLMs, supplemented with first names sourced from the U.S. Social… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Privasis-Zero.
huggingfacetext-generationother0 downloads - 0.50Arsh9210/Privasis-USApkg:data/Arsh9210/Privasis-USA
Dataset Description: Privasis is a large-scale synthetic dataset consisting of diverse text records—such as medical and financial records, legal documents, emails, and messages—containing rich, privacy-sensitive information. Each record includes synthetic profile details, surrounding social context, and annotations of privacy-related content. All data are fully generated using GPT-OSS-120B, supplemented with the personas from the public Nemotron-Personas-USA dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Privasis-USA.
huggingfacetext-generationother0 downloads - 0.50fevziegeyurtsevenler/turkish-pii-corpuspkg:data/fevziegeyurtsevenler/turkish-pii-corpus
turkish-pii-corpus from datasets import load_dataset ds = load_dataset("fevziegeyurtsevenler/turkish-pii-corpus") A synthetic, checksum-valid, character-span-labeled Turkish PII corpus — no real person's data. Every TCKN/IBAN/VKN/plaka/card is randomly generated but passes its checksum, plus distractor sentences with number-like strings that are not PII (to test precision). Each row: text, entities: [{type, start, end, value}]. Entity types: TCKN, IBAN, VKN, PLAKA, PHONE… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-pii-corpus.
huggingfacetoken-classificationapache-2.00 downloads - 0.50fevziegeyurtsevenler/turkish-pii-patterns-kvkkpkg:data/fevziegeyurtsevenler/turkish-pii-patterns-kvkk
Turkish PII Detection Patterns (KVKK) from datasets import load_dataset ds = load_dataset("fevziegeyurtsevenler/turkish-pii-patterns-kvkk") Regex + metadata for detecting/masking Turkish personal data (TCKN, IBAN, VKN...) for KVKK compliance. Schema column meaning pii_type, name_tr type regex, checksum pattern / has checksum masked_example, kvkk_category example / category Related AltaySec resources 🕵️ uncloak scanner:… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-pii-patterns-kvkk.
huggingfacetoken-classificationapache-2.00 downloads - 0.50naeyn/nobody-pii-synth-depkg:data/naeyn/nobody-pii-synth-de
Nobody PII Synthetic A procedurally generated, multilingual corpus for training and regression testing PII detection and redaction systems, with German as the primary language and English/Dutch coverage. Companion model: naeyn/nobody-pii-de Release identity This Hub revision contains the canonical Parquet snapshot. The Parquet files were verified at Hub revision 23eba3fadef50ef4c01738bea0b0ee1c392426e6; card and attribution metadata may receive later commits… See the full description on the dataset page: https://huggingface.co/datasets/naeyn/nobody-pii-synth-de.
huggingfacetoken-classificationapache-2.00 downloads - 0.50sonomoshq/flashlight-neutral-benchmarkpkg:data/sonomoshq/flashlight-neutral-benchmark
Flashlight Neutral PII Benchmark The neutral third-party evaluation used to benchmark sonomoshq/flashlight-1 against popular PII NER models — published so the numbers on that model card can be independently reproduced and challenged. Why "neutral" Every PII model's published F1 is usually measured on its own training distribution. This benchmark uses two third-party datasets that none of the compared models trained on, verified by provenance checks against every… See the full description on the dataset page: https://huggingface.co/datasets/sonomoshq/flashlight-neutral-benchmark.
huggingfacetoken-classificationmit0 downloads - 0.50rizzoaiacademy/anonimizzazione-testi-italiano-cleanpkg:data/rizzoaiacademy/anonimizzazione-testi-italiano-clean
Anonimizzazione Testi Italiano — versione pulita e bilanciata Dataset pronto al training per la token-classification di PII in testi legali italiani (22 categorie in schema BIO), derivato dal corpus community rizzoaiacademy/anonimizzazione-testi-italiano tramite una pipeline di deduplicazione e bilanciamento. Alimenta il modello rizzo-pii. ⚠️ 100% sintetico. Nessun dato personale reale. Nomi, codici fiscali, IBAN, indirizzi ecc. sono generati (con checksum validi o volutamente… See the full description on the dataset page: https://huggingface.co/datasets/rizzoaiacademy/anonimizzazione-testi-italiano-clean.
huggingfacetoken-classificationmit0 downloads - 0.50shivaniachary123/pii-detection-datasetpkg:data/shivaniachary123/pii-detection-dataset
Gravitee PII Detection A harmonized, multi-source corpus for fine-tuning encoder-style PII / NER models. 25 canonical PII classes, character-level span annotations, 175,881 English examples, 781,052 entity spans. Published as a single split (train). Hold-out evaluation is expected to be performed against unrelated external PII corpora rather than against a slice of this dataset. Quick start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-detection-dataset.
huggingfacetoken-classificationapache-2.00 downloads - 0.50zackhatecoding/Med-PCDpkg:data/zackhatecoding/Med-PCD
Med-PCD: Medical Privacy-Conscious Delegation Med-PCD is the medical dataset introduced in Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning (ACL 2026). It is a benchmark for privacy-preserving LLM systems in a domain where queries tend to carry many interconnected PII entities. Paper: arXiv:2510.16054 Code: github.com/zackhuiiiii/Privacy-R1 All PII in Med-PCD is synthetic and does not correspond to any real individual.… See the full description on the dataset page: https://huggingface.co/datasets/zackhatecoding/Med-PCD.
huggingfacetext-generationcc-by-nc-4.00 downloads - 0.50anonymous128463/KoSTA-datasetpkg:data/anonymous128463/KoSTA-dataset
KoSTA-bench 한국어 PII(개인정보) 장면 텍스트 편집(Scene Text Editing) 벤치마크. 원본·타깃 텍스트를 모두 가상 PII 키워드로 구성한 611개 (원본, 타깃) 이미지 쌍으로 이루어진다. 구성 source/{name} — 원본 텍스트(text_s)가 담긴 원본 이미지 crop target/{name} — 타깃 텍스트(text_t)로 편집된 타깃 이미지 crop manifest.jsonl — 샘플별 메타데이터 (한 줄당 한 샘플) manifest 필드 필드 설명 name 샘플 파일명 (source/, target/ 공통) text_s 원본 텍스트 text_t 타깃 텍스트 label PII 유형 라벨 예시: {"name": "000000.png", "text_s": "원수", "text_t": "과장", "label":… See the full description on the dataset page: https://huggingface.co/datasets/anonymous128463/KoSTA-dataset.
huggingfaceimage-to-imagecc-by-nc-4.00 downloads - 0.50sallani/privamesh-legal-syntheticpkg:data/sallani/privamesh-legal-synthetic
PrivaMesh Legal Synthetic Description PrivaMesh Legal Synthetic is a multilingual dataset of 100,000 fully synthetic legal, privacy, security and AI-governance records. It is designed for training and evaluating sallani/PrivaMesh on PII detection, classification, anonymization, pseudonymization, compliance analysis, sensitive-data detection, legal-entity extraction and privacy-risk assessment. No source document or identity was copied from a real person. Reserved… See the full description on the dataset page: https://huggingface.co/datasets/sallani/privamesh-legal-synthetic.
huggingfacetoken-classificationapache-2.00 downloads - 0.50Danili4/pii-benchpkg:data/Danili4/pii-bench
PII-Bench (ru) Бенчмарк для оценки качества детекции персональных данных (PII) в русскоязычных текстах. Использует span-level разметку с явными индексами начала, конца символов и названия сущности, что позволяет валидировать такие системы как Presidio как ML-модели, так и регулярные выражения, фокусируя не оценку самих моделей в формате IO, BIO или BILOU, а фокусираясь на комплексной оценке всего NER пайплайна. Формат данных { id: chat_03, domain: L-CHAT… See the full description on the dataset page: https://huggingface.co/datasets/Danili4/pii-bench.
huggingfacetoken-classificationother0 downloads - 0.50nguyenlamtung/pii-masking-95k-preencodedpkg:data/nguyenlamtung/pii-masking-95k-preencoded
VI PII Masking (Pre-encoded) – 95k Private Vietnamese dataset for PII detection and masking, designed for token classification / NER in privacy-preserving NLP systems. Dataset Overview Language: Vietnamese (vi) Domain: Privacy / PII masking Total samples: ~95,000 Splits: Train: 76,097 Validation: 9,512 Test: 9,513 Format: JSONL Data Fields Raw & masked text source_text masked_text privacy_mask language region script split uid… See the full description on the dataset page: https://huggingface.co/datasets/nguyenlamtung/pii-masking-95k-preencoded.
huggingfacetoken-classificationother0 downloads - 0.50somukandula/maskara-extensive-piipkg:data/somukandula/maskara-extensive-pii
Maskara Extensive PII This dataset is the larger training corpus used for the real Maskara Transformers token-classification model. It contains 75,000 examples: train.jsonl: 67,500 examples validation.jsonl: 3,750 examples test.jsonl: 3,750 examples Sources Maskara synthetic v2 generator Public examples normalized from ai4privacy/pii-masking-300k The dataset includes chat prompts, email drafts, support tickets, delivery messages, identity prompts, finance… See the full description on the dataset page: https://huggingface.co/datasets/somukandula/maskara-extensive-pii.
huggingfacetoken-classificationother0 downloads - 0.50redmadrobot-rnd/pii_benchmarkpkg:data/redmadrobot-rnd/pii_benchmark
Russian PII NER Evaluation Dataset Dataset Description This dataset is designed for evaluating PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) systems on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.
huggingfacetoken-classificationmit0 downloads - 0.50kierandesmond/spanish-gdpr-pii-ner-v9-validationpkg:data/kierandesmond/spanish-gdpr-pii-ner-v9-validation
Spanish GDPR PII NER — v9 Validation & Reproducibility Held-out evaluation set, scripts, and results for kierandesmond/spanish-gdpr-pii-ner-v9. Why this repo exists v6/v8 were scored on a tiny n=1–6 "battery" that reported 90%+ but hid real failures. v9 introduces a large held-out eval set (7,304 examples, 200–336 per entity, with hard negatives) to get statistically-meaningful per-entity numbers. Re-scoring v8 on it revealed SIP_CARD 0.5%, ETHNIC_ORIGIN 40%… See the full description on the dataset page: https://huggingface.co/datasets/kierandesmond/spanish-gdpr-pii-ner-v9-validation.
huggingfacecc-by-4.00 downloads - 0.50ai4privacy/openpii-masking-micro-100kpkg:data/ai4privacy/openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.
huggingfacetoken-classificationother0 downloads - 0.50ai4privacy/pii-masking-mini-10kpkg:data/ai4privacy/pii-masking-mini-10k
PII Masking Mini: Multilingual Sample A mini-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.
huggingfacetoken-classificationother0 downloads - 0.50ai4privacy/pii-masking-micro-100kpkg:data/ai4privacy/pii-masking-micro-100k
PII Masking Micro: Multilingual Sample A micro-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.
huggingfacetoken-classificationother0 downloads - 0.50ai4privacy/pii-masking-nano-1kpkg:data/ai4privacy/pii-masking-nano-1k
PII Masking Nano: Multilingual Sample A nano-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.
huggingfacetoken-classificationother0 downloads - 0.50klusai/ds-kp-general-en-50kpkg:data/klusai/ds-kp-general-en-50k
klusai/ds-kp-general-en-50k KlusAI Privacy (KP) dataset — general domain, language en. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the en LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) en. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-en-50k.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50klusai/ds-kp-general-pl-50kpkg:data/klusai/ds-kp-general-pl-50k
klusai/ds-kp-general-pl-50k KlusAI Privacy (KP) dataset — general domain, language pl. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the pl LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) pl. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-pl-50k.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50klusai/ds-kp-general-ro-50kpkg:data/klusai/ds-kp-general-ro-50k
klusai/ds-kp-general-ro-50k KlusAI Privacy (KP) dataset — general domain, language ro. Gold PII / quasi-identifier spans in the harmonized KP (BIOES) taxonomy, emitted at generation time via the ro LocalePack (offset-deterministic template splice — text[start:end] == value by construction). See https://huggingface.co/datasets/klusai/europriv-bench for the benchmark these feed. This is a released volume (SLUG) built from the generator (PACK) ro. A pack is code (generator +… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-kp-general-ro-50k.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50paperboy-ai/desktop-pii-500pkg:data/paperboy-ai/desktop-pii-500
Desktop PII 500 This directory is a local Hugging Face-compatible dataset package for synthetic desktop screenshots with expected privacy and utility QA pairs. It reuses the existing 210-image Desktop PII package and adds 290 accepted supplemental screenshots generated from google/gemini-3.5-flash scenario prompts and gpt-image-2 image generation. The screenshots are generated synthetic desktop scenes. The visible sensitive values are fictional benchmark strings, not real… See the full description on the dataset page: https://huggingface.co/datasets/paperboy-ai/desktop-pii-500.
huggingfaceimage-to-textother0 downloads - 0.50buzzcraft/Norwegian_PIIpkg:data/buzzcraft/Norwegian_PII
Norwegian PII (Placeholder) This dataset contains annotated Norwegian text samples where sensitive entities are replaced with placeholders such as name_1, address_1, email_1, and phone_1. It is intended as a safe-to-share placeholder version of a Norwegian PII extraction dataset. Dataset Summary The dataset preserves the original annotation structure while removing direct values for selected entity types. Placeholder tokens are used for: person address email… See the full description on the dataset page: https://huggingface.co/datasets/buzzcraft/Norwegian_PII.
huggingfacetoken-classificationmit0 downloads - 0.50Reza2kn/persian-pii-masking-iranian-personas-143kpkg:data/Reza2kn/persian-pii-masking-iranian-personas-143k
Iranian Persona Pool 143K Synthetic Iran-native persona metadata generated for the Persian PII masking data pipeline. This repo contains the persona pool only, not the downstream PII token-classification rows. Dataset Repo Reza2kn/persian-pii-masking-iranian-personas-143k Size Rows: 143515 Unique persona_id: 143515 Provinces: 31 Cities: 3251 Gender counts: {"female": 70305, "male": 73210} Schema Each row contains: persona_id:… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-iranian-personas-143k.
huggingfacecc-by-4.00 downloads - 0.50Accuknoxtechnologies/PIIpkg:data/Accuknoxtechnologies/PII
Accuknoxtechnologies/PII SFT dataset for fine-tuning a Qwen-based guard that detects personally-identifiable information (PII) in a user prompt. Each row pairs a natural-language prompt with a JSON target enumerating the exact substrings of every PII category present. This release combines the previously-separate train + test CSVs into a single train split (source files: pii_openpii.csv, test_dataset_pii.csv). Schema column description prompt user… See the full description on the dataset page: https://huggingface.co/datasets/Accuknoxtechnologies/PII.
huggingfacetoken-classificationapache-2.00 downloads - 0.50Reza2kn/persian-pii-masking-openpii-690k-cleanpkg:data/Reza2kn/persian-pii-masking-openpii-690k-clean
Persian PII-Masking Combined Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits. Persona-clean rows: 623890 Initial-clean rows: 224956 Dataset Repo Reza2kn/persian-pii-masking-openpii-690k-clean Schema Rows include: source_text masked_text privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.
huggingfacetoken-classificationcc-by-4.00 downloads - 0.50FreekCoolAI/kvk-pii-checkerpkg:data/FreekCoolAI/kvk-pii-checker
KVK PII-checker dataset Nederlandstalige Q&A-dataset voor het trainen van een micro-LLM (Gemma-3-1B) als privacy/AVG-checker: gegeven een tekst die iemand in een AI-tool zou willen plakken, geeft het model een 3-regels oordeel. Format Elk voorbeeld: { "category": "pii-check/...", "question": "<tekst die getoetst wordt>", "answer": "Oordeel: <VEILIG|ANONIMISEER EERST|NIET VERSTUREN>\nGevonden: ...\nAdvies: ..." } Oordeel-klassen VEILIG — geen… See the full description on the dataset page: https://huggingface.co/datasets/FreekCoolAI/kvk-pii-checker.
huggingfacetext-generationcc-by-4.00 downloads - 0.45Meddies/meddies-pii-v2pkg:data/Meddies/meddies-pii-v2
Meddies PII v2 Dataset Character-level PII annotations for training and evaluating multilingual de-identification systems across 17 languages and nine entity families. [!IMPORTANT] This is a research artifact for privacy and healthcare AI teams. It is not medical advice, not a redaction tool, and not a substitute for local validation before any clinical deployment, compliance workflow, or high-stakes privacy claim. If you want to use this dataset in commercial… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-pii-v2.
huggingfacetoken-classificationcc-by-nc-4.0132 downloads - 0.45SulthanAbiyyu/anak-baikpkg:data/SulthanAbiyyu/anak-baik
Anak-Baik Dataset: Overview Anak-Baik dataset is a collection of instruction-output pairs in Bahasa Indonesia, designed for Supervised Fine-Tuning (SFT) tasks. The dataset contains examples of both harmful and harmless outputs, aimed at promoting ethical AI development (hence the name; anak baik == good boy :D). The dataset consists of pairs of instructions and their corresponding outputs, categorized as either harmful or harmless and their topics. This structure enables models to… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/anak-baik.
huggingfacetext-generation14 downloads - 0.45SulthanAbiyyu/anak-baik-rejection-classificationpkg:data/SulthanAbiyyu/anak-baik-rejection-classification
Anak-Baik Rejection Classification: Overview The Anak-Baik Rejection Classification dataset is a curated collection of labeled instructional rejections in Bahasa Indonesia, specifically designed for Supervised Fine-Tuning (SFT) tasks. This dataset includes examples of both harmful and harmless instructions, along with labels indicating whether an instruction should be answered or rejected. This dataset aimed at promoting ethical AI development (hence the name; anak baik == good… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/anak-baik-rejection-classification.
huggingfacetext-classification5 downloads - 0.40ai4privacy/pii-masking-300kpkg:data/ai4privacy/pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.
huggingfacetext-classificationother6,530 downloads - 0.40nvidia/Nemotron-PIIpkg:data/nvidia/Nemotron-PII
Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-PII.
huggingfacetoken-classificationcc-by-4.04,030 downloads - 0.40ai4privacy/pii-masking-400kpkg:data/ai4privacy/pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.
huggingfacetext-classificationother2,691 downloads - 0.40ai4privacy/pii-masking-openpii-1mpkg:data/ai4privacy/pii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.
huggingfacetoken-classificationother1,769 downloads - 0.40ai4privacy/open-pii-masking-500k-ai4privacypkg:data/ai4privacy/open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.
huggingfacetext-classificationother1,450 downloads - 0.40Meddies/meddies-piipkg:data/Meddies/meddies-pii
Meddies PII Synthetic PII extraction data for multilingual clinical and administrative documents, with language-specific, domain-transfer, translation, and instruction-style views in one Hub repo. [!IMPORTANT] This is a synthetic de-identification research artifact for healthcare AI teams. It is not medical advice, not a privacy certification, and not a substitute for task-specific validation on your own data. If you want to use this dataset in commercial work… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-pii.
huggingfacetoken-classificationcc-by-nc-4.01,360 downloads - 0.40guardion/BR-Agentic-PII-Benchmarkpkg:data/guardion/BR-Agentic-PII-Benchmark
BR-Agentic-PII-Benchmark Overview BR-Agentic-PII-Benchmark is a highly annotated dataset of synthetic multi-turn conversations between humans and AI banking assistants in Brazilian Portuguese. It is designed to benchmark the processes of detection, anonymization, de-anonymization, and transparency for AI agents. Key Use Cases This dataset helps evaluate systems that redact and protect Personally Identifiable Information (PII) before it leaves a secure perimeter… See the full description on the dataset page: https://huggingface.co/datasets/guardion/BR-Agentic-PII-Benchmark.
huggingfacetext-generationmit1,010 downloads - 0.40gretelai/synthetic_pii_finance_multilingualpkg:data/gretelai/synthetic_pii_finance_multilingual
Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.
huggingfacetext-classificationapache-2.0995 downloads - 0.40nvidia/Privasis-Zeropkg:data/nvidia/Privasis-Zero
Privasis-Zero Dataset Description: Privasis-Zero is a large-scale synthetic dataset consisting of diverse text records—such as medical and financial records, legal documents, emails, and messages—containing rich, privacy-sensitive information. Each record includes synthetic profile details, surrounding social context, and annotations of privacy-related content. All data are fully generated using LLMs, supplemented with first names sourced from the U.S. Social Security… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Privasis-Zero.
huggingfacetext-generationother684 downloads - 0.40hivetrace/pii-benchpkg:data/hivetrace/pii-bench
PII-Bench (ru) Span-level benchmark for evaluating personal-data (PII) detection in Russian text. Annotations use explicit character offsets (start, end, type) rather than IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and lets you evaluate a complete pipeline — ML model, regular expressions, post-processing, or a hybrid such as Presidio — instead of only the model in isolation. Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.
huggingfacetoken-classificationapache-2.0607 downloads - 0.40BCCard/pii-masking-openpii-financepkg:data/BCCard/pii-masking-openpii-finance
1. Overview Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either generated locally under the documented safety controls or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/pii-masking-openpii-finance.
huggingfacetoken-classificationcc-by-4.0520 downloads - 0.40mks-logic/SPYpkg:data/mks-logic/SPY
SPY: Enhancing Privacy with Synthetic PII Detection Dataset We proudly present the SPY Dataset, a novel synthetic dataset for the task of Personal Identifiable Information (PII) detection. This dataset highlights the importance of safeguarding PII in modern data processing and serves as a benchmark for advancing privacy-preserving technologies. Key Highlights Innovative Generation: We present a methodology for developing the SPY dataset and compare it to other… See the full description on the dataset page: https://huggingface.co/datasets/mks-logic/SPY.
huggingfacetoken-classificationcc-by-4.0490 downloads - 0.40jmdanto/corpus-essms-publicpkg:data/jmdanto/corpus-essms-public
Corpus social et medico-social (export public) Apercu Ce dataset contient des ecrits professionnels en francais du secteur social et medico-social. L'export public diffuse ici est compose de rapports fictifs mais realistes (les documents reels ont ete exclus), et vise des usages de NER, pseudonymisation, extraction d'entites et evaluation de modeles en contexte metier. Volume de l'export public: 410 enregistrements fictifs realistes dans data.jsonl (texte + metadonnees)… See the full description on the dataset page: https://huggingface.co/datasets/jmdanto/corpus-essms-public.
huggingfacetoken-classificationapache-2.0434 downloads - 0.40tomekkorbak/pile-pii-scrubadubpkg:data/tomekkorbak/pile-pii-scrubadub
Dataset Card for pile-pii-scrubadub Dataset Summary This dataset contains text from The Pile, annotated based on the personal idenfitiable information (PII) in each sentence. Each document (row in the dataset) is segmented into sentences, and each sentence is given a score: the percentage of words in it that are classified as PII by Scrubadub. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is taken from The… See the full description on the dataset page: https://huggingface.co/datasets/tomekkorbak/pile-pii-scrubadub.
huggingfacetext-classification['mit']360 downloads - 0.40raayraay/privacyleak-piipkg:data/raayraay/privacyleak-pii
PrivacyLeak-PII A Machine Unlearning Benchmark for Personal Information Extraction All data in this dataset is synthetically generated using Faker. No real PII is included. The Problem We're Solving Current unlearning evaluations only check if models refuse direct questions about "forgotten" data. But real attackers don't ask nicely: Direct question (model refuses): "What is John Doe's SSN?" Prefix completion (model leaks): "Customer Name: John Doe Issue:… See the full description on the dataset page: https://huggingface.co/datasets/raayraay/privacyleak-pii.
huggingfacetext-generationmit305 downloads - 0.40UniDataPro/synthetic-printed-brazilian-passportspkg:data/UniDataPro/synthetic-printed-brazilian-passports
Brazilian passport dataset The dataset comprises 5,000 high-resolution synthetic photos of Brazilian passports, designed to advance computer vision and identity verification systems. It provides a secure and ethical resource for training robust models for OCR (Optical Character Recognition), document analysis, and spoofing detection, all without exposing real personal data or sensitive personal information. By utilizing this dataset, researchers and developers can enhance… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-brazilian-passports.
huggingfaceimage-to-textcc-by-nc-nd-4.0278 downloads - 0.40lianghsun/tw-PII-benchpkg:data/lianghsun/tw-PII-bench
Taiwan PII Benchmark (tw-PII-bench) A token-classification benchmark for evaluating PII detectors on Taiwan-specific personally identifiable information in Traditional Chinese (繁體中文). Designed against openai/privacy-filter to surface its label-coverage gaps and locale-specific failure modes. The benchmark has three splits by text length, so you can isolate where a model breaks (boundary handling, long-context coverage, multi-PII reasoning): Split Items Text lengthAvg PII… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-PII-bench.
huggingfacetoken-classificationapache-2.0274 downloads - 0.40Ganasekhar/pii-masking-400kpkg:data/Ganasekhar/pii-masking-400k
Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.
huggingfacetext-classificationother268 downloads - 0.40TonicAI/Privacy-Benchpkg:data/TonicAI/Privacy-Bench
PrivacyBench PrivacyBench is a benchmark for de-identifying semi-structured data exports from work tools like email, messaging, calendar, and so forth. The benchmark focuses on the identification and synthesis of PII in the unstructured text fields of the data export, and introduces novel metrics for evaluating synthesis quality. This initial version consists of data exports of Slack and email messages from 21 distinct personas generated by the Fabricate synthetic data tool… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/Privacy-Bench.
huggingfacetoken-classificationcc-by-4.0267 downloads - 0.40Pritesh-2711/pii-benchpkg:data/Pritesh-2711/pii-bench
PIIBench Description PIIBench is a unified benchmark dataset for PII detection across multiple domains. Paper arXiv: http://arxiv.org/abs/2604.15776 Dataset Summary Total records: 999,940 Entity types: 82 BIO labels: 165 including O Format: BIO token classification with source text Structure Each example contains: tokens: list of tokens labels: BIO labels source: original data source of the sample text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.
huggingfacetoken-classificationapache-2.0260 downloads - 0.40DataikuNLP/kiji-pii-training-datapkg:data/DataikuNLP/kiji-pii-training-data
Kiji PII Detection Training Data Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution. Dataset Summary Samples 51,495 (train: 46,345, test: 5,150) Languages 6 (English, Danish, Dutch, French, Spanish, German) Countries 20 PII entity types 26 Total entity annotations 397,441 (avg 7.7 per sample) Coreference clusters 0 (0% of samples)… See the full description on the dataset page: https://huggingface.co/datasets/DataikuNLP/kiji-pii-training-data.
huggingfacetoken-classificationapache-2.0257 downloads - 0.40UniDataPro/synthetic-printed-usa-passports-datasetpkg:data/UniDataPro/synthetic-printed-usa-passports-dataset
Passport Dataset - 9 600 Images The dataset comprises 9,600 high-quality synthetically generated passport images, providing a robust resource for training and verifying document analysis systems. Every passport is presented across 3 angles (0°, 25°, 45°), 4 lighting conditions (Natural-daylight, Office-LED, Warm-indoor, Dim-light), 4 backgrounds (Neutral wall, Textured desk, Outdoor pavement, Docs-on-docs), and 2 distances (Close, Medium), creating a rich and challenging dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-printed-usa-passports-dataset.
huggingfaceimage-to-textcc-by-nc-nd-4.0255 downloads - 0.40temsa/OpenMed-Irish-CorePII-TrainMix-v1pkg:data/temsa/OpenMed-Irish-CorePII-TrainMix-v1
OpenMed Irish Core PII Train Mix v1 Composite token-classification training mix used to fine-tune temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1. This repo is the training dataset, not the model itself. What A Row Looks Like Each row uses a fixed schema so the Hugging Face dataset viewer and datasets.load_dataset() can read it directly: id: row id inside the split text: reconstructed text string tokens: tokenized text labels: BIO labels aligned to tokens language:… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-CorePII-TrainMix-v1.
huggingfacetoken-classificationcc-by-4.0236 downloads - 0.40AdamiTitus/pii-masking-300kpkg:data/AdamiTitus/pii-masking-300k
Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.
huggingfacetext-classificationother231 downloads - 0.40saad-kw-almutairi/pii-masking-300kpkg:data/saad-kw-almutairi/pii-masking-300k
Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/saad-kw-almutairi/pii-masking-300k.
huggingfacetext-classificationother210 downloads - 0.40tursunait/roberta-pii-synthpkg:data/tursunait/roberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tursunait/roberta-pii-synth.
huggingfacetoken-classificationmit188 downloads - 0.40anony-mouse123/enron_canarypkg:data/anony-mouse123/enron_canary
CanaryBench-Enron Frequency-aware canary injection benchmark for auditing memorization in finetuned language models, built on the Enron email corpus. Dataset Description This dataset is part of CanaryBench, a benchmark for evaluating memorization in finetuned language models across repetition tiers and privacy regimes. Frequency tiers: 1×, 10×, 50× Domain: Email (Enron corpus) Member canaries: 770 Reference canaries: 1000 Files… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/enron_canary.
huggingfacetext-generationcc-by-4.0187 downloads - 0.40DomainShield/InternalPiiDatasetpkg:data/DomainShield/InternalPiiDataset
Internal PII Benchmark A synthetic dataset for training and evaluating models on the detection of domain-specific PII — organization-internal identifiers that conventional PII systems fail to recognize. Unlike traditional PII (names, emails, phone numbers), this dataset targets terms such as internal team names, restricted locations, communication channels, infrastructure labels, and operational procedures (e.g., gamma squad, secure chamber, inner route). These terms are… See the full description on the dataset page: https://huggingface.co/datasets/DomainShield/InternalPiiDataset.
huggingfacetoken-classificationmit177 downloads - 0.40ameau01/synthetic-it-support-ticketspkg:data/ameau01/synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth 745 synthetic IT service-management incident records for LLM wiki and retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root cause, and resolution steps. The free text is enriched with realistic technical detail and injected synthetic PII. The corpus ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.
huggingfacetext-generationmit172 downloads - 0.40gravitee-io/pii-detection-datasetpkg:data/gravitee-io/pii-detection-dataset
Gravitee PII Detection A harmonized, multi-source corpus for fine-tuning encoder-style PII / NER models. 25 canonical PII classes, character-level span annotations, 175,881 English examples, 781,052 entity spans. Published as a single split (train). Hold-out evaluation is expected to be performed against unrelated external PII corpora rather than against a slice of this dataset. Quick start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/pii-detection-dataset.
huggingfacetoken-classificationapache-2.0170 downloads - 0.40ai4privacy/pii-masking-health-phi-previewpkg:data/ai4privacy/pii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Health & Medical Information (PHI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.
huggingfacetoken-classificationcc-by-4.0165 downloads - 0.40ai4privacy/openpii-masking-nano-1kpkg:data/ai4privacy/openpii-masking-nano-1k
OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.
huggingfacetoken-classificationother161 downloads - 0.40enosislabs/matex-privacy-sentinel-datasetpkg:data/enosislabs/matex-privacy-sentinel-dataset
MaTE X Privacy Sentinel Dataset Synthetic token-classification dataset for training a local privacy/security filter based on OpenAI Privacy Filter. Purpose This dataset teaches a local filter to detect and redact sensitive spans in developer workflows before context is sent to external LLMs. Target domains include: .env files terminal logs stack traces git diffs GitHub issues and PR comments agent traces tool outputs workspace memory auth, database, cloud and payment… See the full description on the dataset page: https://huggingface.co/datasets/enosislabs/matex-privacy-sentinel-dataset.
huggingfacetoken-classificationapache-2.0153 downloads - 0.40joneauxedgar/pasteproof-pii-dataset-v2pkg:data/joneauxedgar/pasteproof-pii-dataset-v2
PasteProof PII Dataset v2 Improved synthetic dataset for training PII detection models. What's New in v2 Dynamic templates: Templates generated on-the-fly with random variation More format variations: Each entity type has many format options Hard negatives: 10% of samples are tricky non-PII that looks like PII Variable key names: apiKey, api_key, API_KEY, etc. More entity generators: Many more API key formats, card types, etc. Entity Types (27) Financial:… See the full description on the dataset page: https://huggingface.co/datasets/joneauxedgar/pasteproof-pii-dataset-v2.
huggingfacetoken-classificationmit151 downloads - 0.40mindbomber/aana-peer-review-evidence-packpkg:data/mindbomber/aana-peer-review-evidence-pack
AANA Peer Review Evidence Pack This dataset packages the current public evidence for AANA as an architecture for making agents more auditable, safer, more grounded, and more controllable. The claim boundary is intentionally narrow: AANA is production-candidate as an audit/control/verification/correction layer. AANA is not yet proven as a raw agent-performance engine. Results here are measured held-out or validation artifacts, not official leaderboard proof unless a benchmark… See the full description on the dataset page: https://huggingface.co/datasets/mindbomber/aana-peer-review-evidence-pack.
huggingfacemit135 downloads - 0.40RedactionBench/RedactionBenchpkg:data/RedactionBench/RedactionBench
Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The above is… See the full description on the dataset page: https://huggingface.co/datasets/RedactionBench/RedactionBench.
huggingfacetoken-classificationcc-by-4.0132 downloads - 0.40scanpatch/pii-ner-corpus-synthetic-controlledpkg:data/scanpatch/pii-ner-corpus-synthetic-controlled
PII NER Corpus - Synthetic Controlled A controlled synthetic dataset for training Named Entity Recognition models to detect Personally Identifiable Information (PII) in Ukrainian and Russian text. This dataset was generated using a controlled pipeline with human-verified annotation guidelines. The text samples are based on real-world document patterns and annotated using Claude Sonnet 4 with strict quality controls. Dataset Description This dataset contains text… See the full description on the dataset page: https://huggingface.co/datasets/scanpatch/pii-ner-corpus-synthetic-controlled.
huggingfacetoken-classificationmit131 downloads - 0.40compliancemas/ComplianceMAS-Benchpkg:data/compliancemas/ComplianceMAS-Bench
ComplianceMAS-Bench Dataset Description ComplianceMAS-Bench is the first systematic benchmark for evaluating compliance behaviour in multi-agent memory systems. It comprises 269 scenarios spanning 5 compliance failure-mode categories and 4 regulated domains, grounded in HIPAA and GDPR requirements. Paper: ComplianceMAS: A Systematic Benchmark for Evaluating Compliance Behaviour in Multi-Agent Memory Systems (NeurIPS 2025 submission)Repository:… See the full description on the dataset page: https://huggingface.co/datasets/compliancemas/ComplianceMAS-Bench.
huggingfacetext-classificationapache-2.0129 downloads - 0.40aniket-curlscape/pii-masking-englishpkg:data/aniket-curlscape/pii-masking-english
Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english.
huggingfacetext-classificationother124 downloads - 0.40Ari-S-123/pii-detection-english-consolidatedpkg:data/Ari-S-123/pii-detection-english-consolidated
PII Detection Combined Dataset Combined dataset for PII (Personally Identifiable Information) detection, merging the ai4privacy English-only subset with synthetically generated and semantically validated with different LLMs challenging examples targeting NER failure modes. Class labels had to be consolidated to prevent label fragmentation too. Dataset Description This dataset combines two sources: ai4privacy/open-pii-masking-500k (English subset): 120,533 train / 30,160… See the full description on the dataset page: https://huggingface.co/datasets/Ari-S-123/pii-detection-english-consolidated.
huggingfacetoken-classificationmit124 downloads - 0.40subhash-holla/pii-anonpkg:data/subhash-holla/pii-anon
PII-Anon A CC0 multilingual PII benchmark corpus of 782,677 records carrying 3,107,240 entity annotations across 66 entity types and 60 languages, spanning 7 evaluation dimensions. Each record exposes the five legally-distinct regulatory regime signals (gov-02 / FR-022) as separate reg_* columns — no merged compliance verdict. Train vs. evaluation substrate The 159,891 tier3_evaluation records are the EVALUATION substrate of the 782,677-record corpus… See the full description on the dataset page: https://huggingface.co/datasets/subhash-holla/pii-anon.
huggingfacetoken-classificationcc0-1.0122 downloads - 0.40vkatg/streaming-phi-deidentification-benchmarkpkg:data/vkatg/streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.
huggingfacemit118 downloads - 0.40aniket-curlscape/pii-masking-english-100pkg:data/aniket-curlscape/pii-masking-english-100
Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-100.
huggingfacetext-classificationother113 downloads - 0.40nisaefendioglu/synthetic-sensitive-data-in-source-code-n300pkg:data/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300
Synthetic Sensitive Data in Source Code (N=300) Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives). Designed for evaluating local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure scenarios in AI-assisted coding workflows. Version 1.2: multi_secret (and related) samples label every secret present in code_text (complete ground truth). All… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.
huggingfacetext-classificationmit112 downloads - 0.40aniket-curlscape/pii-masking-english-1kpkg:data/aniket-curlscape/pii-masking-english-1k
Important This repository contains the English-only subset of the Ai4Privacy PII-Masking-300k Dataset. The dataset is curated to provide English texts only, while retaining the structure, labeling schema, and licensing of the original dataset. Licensing Academic use is encouraged with proper citation provided it follows similar license terms*. Commercial entities should contact us at licensing@ai4privacy.com for licensing inquiries and additional data access.* Terms… See the full description on the dataset page: https://huggingface.co/datasets/aniket-curlscape/pii-masking-english-1k.
huggingfacetext-classificationother108 downloads - 0.40VytautoDidziojoUniversitetas/NUS-LT-PII-corpuspkg:data/VytautoDidziojoUniversitetas/NUS-LT-PII-corpus
NUS Lithuanian PII Corpus Description Lithuanian text annotated for personal information (PII), spanning three subject domains — administrative, scientific, and media — plus a stratified validation set. The corpus covers 24 entity types: 16 general categories (PER, LOC, ORG, …) and 8 GDPR special-category "sensitive" entities (REL, POL, SEX, GENDER, MAR, FAM, ETH, HEALTH). Dataset Summary Subsets: 4 (3 training categories + 1 validation set) Total records: 41… See the full description on the dataset page: https://huggingface.co/datasets/VytautoDidziojoUniversitetas/NUS-LT-PII-corpus.
huggingfacetoken-classificationopenrail108 downloads - 0.40ai4privacy/openpii-masking-mini-10kpkg:data/ai4privacy/openpii-masking-mini-10k
OpenPII Masking Mini 10K A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models. Sampling Methodology Samples were selected using proportional stratified sampling by language: Target count per language = round(lang_proportion × 10,000) — proportional representation. Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.
huggingfacetoken-classificationcc-by-4.0101 downloads - 0.40temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1pkg:data/temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1
OpenMed Irish PPSN Eircode Spec v1 Focused synthetic token-classification dataset for Irish PPSN and Eircode detection. This repo contains synthetic training rows, not a fine-tuned model. What A Row Looks Like Each row uses a fixed schema: id: row id inside the split text: rendered text string tokens: tokenized text labels: BIO labels aligned to tokens language: en or ga source_dataset: generator identifier source_domain: optional domain tag, empty in this release… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1.
huggingfacetoken-classificationapache-2.099 downloads - 0.40wan9yu/pii-bench-zhpkg:data/wan9yu/pii-bench-zh
PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.
huggingfacetoken-classificationapache-2.099 downloads - 0.40klusai/europriv-benchpkg:data/klusai/europriv-bench
EuroPriv-Bench (v0 — detection) Held-out gold for pan-European PII/PHI de-identification, labeled in the harmonized KP taxonomy (BIOES). Each row: text, spans ({start, end, label}, KP labels), language. Scored by the europriv-bench harness (entity F1 / recall-weighted F2; re-identification-risk + privacy-utility tracks to follow). v0 scope: seeded from AI4Privacy open core, remapped to KP. Legal + clinical splits, Romanian (not in AI4Privacy), and TAB/MEDDOCAN subsumption land… See the full description on the dataset page: https://huggingface.co/datasets/klusai/europriv-bench.
huggingfacetoken-classificationcc-by-4.098 downloads - 0.40ai4privacy/pii-masking-location-pli-previewpkg:data/ai4privacy/pii-masking-location-pli-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Location & Travel Information (PLI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-location-pli-preview.
huggingfacetoken-classificationcc-by-4.097 downloads - 0.40ai4privacy/pii-masking-digital-pdi-previewpkg:data/ai4privacy/pii-masking-digital-pdi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Digital Information (PDI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-preview.
huggingfacetoken-classificationcc-by-4.097 downloads - 0.40shivaniachary123/pii-masking-400kpkg:data/shivaniachary123/pii-masking-400k
Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-400k.
huggingfacetext-classificationother94 downloads - 0.40ai4privacy/pii-masking-financial-pfi-previewpkg:data/ai4privacy/pii-masking-financial-pfi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Financial Information (PFI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-preview.
huggingfacetoken-classificationcc-by-4.093 downloads - 0.40shivaniachary123/pii-masking-300kpkg:data/shivaniachary123/pii-masking-300k
Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. Key facts: OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/shivaniachary123/pii-masking-300k.
huggingfacetext-classificationother93 downloads