00 halilneed · Istanbul
Hi, I’m
Halil.
I measure things and publish the numbers.
I’m a Boğaziçi University graduate working where business and artificial intelligence meet. Most of what I do comes down to finding out what actually happens — inside a model, in a coding agent’s session log, in a team’s own data — and writing it down.
What I’m working on
- Turkish privacy models Small models that mask personal data and classify KVKK categories in Turkish text. They run on a CPU.
- Tools for coding agents Six open-source plugins that show what an agent actually did, read from logs already on your disk.
- AI adoption in corporate teams Starting from a team’s own data instead of a tool. One team’s daily incoming requests went from 100 to 35.
- Writing it up Long-form on Medium about what the logs and benchmarks show, including where things got worse.
Latest writing — Medium
-
2Models on Hugging Face
-
1Open benchmark
-
6Agent plugins
-
5Dispatches
01 Work
What I have built
and shipped.
Models, repositories and writing. Each line carries the number its model card, README or session log reports.
Models & datasets — Hugging Face
A 270M model that masks personal data in Turkish text according to a one-sentence policy. 0.885 → 0.945 exact match on the held-out half of the benchmark; 1,359 downloads in its first 14 days.
Tells which of 24 KVKK personal-data categories a Turkish text contains, in 22 ms on a CPU. 0.874 micro-F1 on 987 held-out texts.
The evaluation set behind the classifier: 987 gold texts, 24 labels, CC0 — published so other models can be measured on it.
Repositories — GitHub
agent-blackbox, skillbench, scar, gardener, painradar and devpersona: six Claude Code plugins that read the session logs already on your disk and show what the agent actually did.
A local-first memory intelligence layer for coding agents: inspects what your agents know and moves context between sessions, machines and agents.
Pulls live data from Reddit, GitHub Trending and Hacker News, then surfaces the tools that matter to solo developers from inside your assistant.
An agent skill for the HyperFrames composition contract, a terminal-style prompt quiz in three languages, and the marketplace manifest.
Writing — Medium
02 Models
Small Turkish models
for personal data.
Published on Hugging Face. Each one runs on a CPU, so the text never has to leave the machine — and each card says where the model still fails.
Replaces names, national IDs, IBANs and phone numbers in Turkish text with tags before it reaches an LLM, a log store or a vendor. The masking policy is a sentence: mask everything, only these fields, or everything except those. v02 went from 0.885 to 0.945 exact match on the half of the benchmark never used for model selection. Weakest slice: text with several people in it, 0.750.
Says which personal-data categories a Turkish text contains: 14 VERBİS categories plus the 10 special categories of KVKK Art. 6. It sits in front of the masker — classify, pick a policy, mask. 0.874 micro-F1 on 987 held-out texts; “special category present?” gets 0.960 precision and 0.884 recall, so about one in nine still slips through.
The evaluation set behind the classifier, published so other models can be measured on it: 987 gold texts, 24 labels. Texts and both blind label passes are LLM-generated; gold is the 87.7% where the two passes agreed exactly. It was not verified line by line by a human.
A local inference example and a dependency-free evaluator for the masking model, with notes on how it was trained and where its evaluation is limited.
03 Repositories
Open source,
on GitHub.
The agent tools have their own page: the six plugins, how to install them and the four rules they are built on. Everything else public is below.
agent-blackbox, skillbench, scar, gardener, painradar and devpersona. Each one reads something that already exists on your disk and cites where every claim came from. Five of the six make no network calls at all.
A local-first memory intelligence layer for coding agents. It reads what your tools already write — session transcripts, auto-memory, project instructions — to surface context risks and conflicts, and to move context between sessions, machines and agents.
The manifest that makes the six plugins installable side by side. One repo,
one marketplace.json, added once.
Pulls live data from Reddit, GitHub Trending and Hacker News, then surfaces the tools that matter to solo developers from inside your assistant. Generic mode works out of the box; role-based mode tunes the sources to your job.
Teaches coding agents the HyperFrames composition contract so a render is correct on the first attempt: 19 rule files covering timing attributes, GSAP timelines, transitions, Lottie, 3D and the render pipeline.
A terminal-style prompt quiz. Five scenarios scored on clarity, creativity, technical depth and strategy, ending in a shareable Prompt IQ card. Turkish, English and Spanish.
04 Dispatches
Notes from building these.
Long-form on Medium, indexed here. What the logs actually show, how the models were trained and where they got worse, and what the agent ecosystem still gets wrong.
05 About
Boğaziçi graduate.
Business and AI.
Boğaziçi University graduate, based in Istanbul. I work across business and artificial intelligence — data analysis, product management and software — and spend the rest of the time on what is listed above: Turkish-language privacy models, tooling for coding agents, and data-loss-prevention research.
The common thread is measurement. A model card says where the model still fails; a tool cites the log line a finding came from; an article publishes the numbers, including the ones that did not flatter me.
06 Contact
Open to interesting problems.
If you are working on agent observability, developer tooling, Turkish-language privacy tooling, or anything where a system has to prove what it did rather than claim it — I want to hear about it.