Turkish NLP · v02 · 270M parameters
Turkish PII detection
and masking.
A small model by halilneed that transforms Turkish text according to a masking instruction. Use it to redact personal-data expressions before text reaches an LLM, log store or third party.
It generates masked text, rather than NER spans or entity offsets. The policy can request full masking, selected fields or everything except named fields. Local inference is possible; outputs still need validation.
- Size
- 270M
- Label schema
- 53 labels
- Reported v02 exact match
- 0.944 / 1,000 synthetic rows
- Model terms
- Gemma
01 / Task
Choose what to mask.
The model reads a Turkish instruction and an input text, then generates a transformed version. This fictional example illustrates the task contract; it is not a newly recorded prediction.
Instruction: Metindeki e-posta adreslerini maskele;
diğer içeriği koru.
Input: E-posta demo@example.com; durum açık.
Target: E-posta [EMAIL]; durum açık.Names, national identifiers and contact information are among the personal-data expressions described by the model. Consult the current model card for the masking policy and known failures before designing strict output checks.
02 / Usage
Run the Python example locally.
The companion repository provides a command-line example with the model's prompt format and deterministic decoding. The following command selects the v02 snapshot recorded on 5 October 2026.
git clone https://github.com/halilneed/turkish-pii-detection.git
cd turkish-pii-detection
python -m venv .venv
source .venv/bin/activate
python -m pip install torch transformers accelerate
python examples/mask.py \
--revision 28644718923ae38b0105c9f3d2be57312ad0ced3 \
--text 'E-posta demo@example.com; durum açık.' \
--instruction 'Metindeki e-posta adreslerini maskele; diğer içeriği koru.'Without --revision, the companion script retains its historical pinned v01 default. The Hugging Face model's unpinned main revision serves v02. The commands above use macOS/Linux shell syntax and were not executed as an inference benchmark for this page.
The published card reports about 600 ms per short sentence and 1.6 GB RAM on an Intel i5-12400F using fp32, eight CPU threads and batch size one. These are publisher-reported conditions, not measurements performed here; speed depends on input and hardware.
03 / Evidence
Read the score with its boundary.
The current model card reports whole-row exact match on a synthetic Turkish PII Masking Benchmark. The results below were not independently rerun while preparing this page.
| Metric / slice | v01 rerun | v02 |
|---|---|---|
| Exact match · 1,000 rows | 0.880 | 0.944 |
| Schema-neutral · 903 rows | 0.900 | 0.951 |
| B partition · 531 rows | 0.885 | 0.945 |
| Long text | 0.656 | 0.844 |
| Multiple people | 0.600 | 0.750 |
| Suffixed personal data | 0.758 | 0.803 |
The card says aggregate benchmark feedback influenced development. The interpolation ratio was selected using the A partition (469 rows), a development set and internal probes; the B partition (531 rows) was not used for selection. Benchmark row contents were not used for training or templates, according to the publisher. This is a reported boundary, not an independent contamination audit.
The historical v01 card reported 0.882 / 0.902; the current card attributes its slightly lower rerun to bf16 batched inference differences. Exact match measures the complete output string. It does not measure entity recall or establish that 94.4% of real documents are safe to share.
Release and evaluation notes ↗04 / Limits
Validate the transformed text.
- Evaluation is synthetic. Measure the model on your own representative documents before relying on it.
- Known weaker slices include several people in one text, Turkish suffixes and uppercase inputs.
- Very short, context-free fields may remain unmasked. Supply the relevant context and inspect the output.
- Generation can miss personal data, change unrelated content or produce unexpected labels. Empty or truncated output is also a failure.
- Restricted instructions intentionally leave some fields visible. Masking alone does not establish anonymization or KVKK/GDPR compliance.
Model weights are subject to Gemma Terms of Use. The companion repository's MIT license covers its original code and documentation; it does not relicense model weights.
05 / Questions
Before integrating the model.
Is this a Turkish NER model?
Its output is generated, masked text. It does not return token labels or entity offsets. Use a span-based detector if your application needs structured positions.
Can it run without sending text to an API?
Yes, local inference is supported. Model files and Python dependencies must be downloaded first. The card includes CPU settings; this page has not independently verified inference speed.
How is masking different from KVKK classification?
Masking transforms the text. The separate Turkish KVKK classifier predicts data categories, which can inform a masking policy. Category labels do not determine legal permission to process data.
Which version will I download?
The canonical Hugging Face repository serves v02 on main. Use revision="v01" for the older release, or a recorded commit SHA for reproducibility. The companion script's default remains pinned to v01; pass the v02 revision in the quick start.
06 / Türkçe
Türkçe kişisel veri maskeleme modeli.
turkish-pii-detection, halilneed tarafından geliştirilen, verilen talimata göre Türkçe metindeki kişisel verileri maskeleyen 270M parametreli bir modeldir. Tüm kişisel verileri, yalnızca seçilen alanları veya belirtilen alanlar dışındakileri maskelemek için kullanılabilir.
Çıktısı maskelenmiş metindir; NER konumları veya KVKK sınıflandırma etiketleri üretmez. Güncel v02 kartında 1.000 sentetik örnekte 0.944 tam eşleşme raporlanır. Model seçiminde kullanılmadığı belirtilen 531 örneklik B bölümündeki skor 0.945'tir. Bu sayfa için yeniden ölçüm yapılmadı.
Yerel kullanım örnekleri GitHub'da, model ağırlıkları ve teknik açıklamalar Hugging Face'tedir. Gerçek belgelerde kullanmadan önce kendi veriniz üzerinde ölçün; özellikle çok kişili, ek almış ve bağlamsız kısa metinlerde çıktıyı doğrulayın.
Code, weights and the write-up.
- Current model card and weights
- Companion repository and evaluation utilities
- Medium: Daha büyük model eğitmedim, zayıf dilimleri eğittim
- Other Turkish privacy models by halilneed
Technical source: current model card, revision 28644718923ae38b0105c9f3d2be57312ad0ced3, read 5 October 2026.