Can Jev redact PII?
Jev scores 1.43× GPT-6 Luna’s F2 on PII masking.
PII masking F2
Loading measured scores…
Mean across four sources. Higher is better. F2 = 5PR / (4P + R), where P is precision and R is recall.
Background
Redacting personal information means catching names, emails, IDs, and other sensitive spans without blanking out the whole document. We tested whether Jev’s yes/no judgments could do that job better than an LLM, a dedicated PII model, or a few regexes.
Methodology
We sampled 400 short English sentences from PIIMB: 100 each from Ai4Privacy, Nemotron-PII, Gretel, and Privy. Jev marked whitespace spans at probability ≥0.5, GPT-6 Luna returned character offsets, and GLiNER Multi PII ran on CPU at threshold 0.3. Regex and mask-all supplied cheap baselines. All used the same text and gold masks; gold never entered inference. We scored character-level F2, which weights recall more than precision, then averaged the four source scores.
Results
Jev reached 78.1% mean F2, ahead of GLiNER at 70.3% and Luna at 54.6%. It led on three of four sources. GLiNER led on Privy, 58.6% versus Jev’s 51.1%. The full Jev run cost about $0.015 in API usage; Luna cost $0.017. GLiNER had no API fee and ran on an existing CPU VM.
Discussion
The average hides a sharp weak spot: Jev’s precision on Privy was 24.4%, so it masked plenty of text it should have left alone. Luna produced 19 invalid spans or JSON outputs, scored as misses. These are different workflows, not a clean test of model weights alone. The sample is short, English, and public, with possible training overlap. A good F2 score is not a guarantee that private data stays private.
Conclusion
Jev is worth testing as a cheap PII detector, especially if your inputs resemble the three sources it led on. Do not ship unattended redaction from this score alone. Check misses and over-masking on your own documents before trusting any of these methods.
Data and run details
Results
Masking F2 (5 × precision × recall / (4 × precision + recall)) by source. Higher is better.
| Source | Jev | GPT-6 Luna | GLiNER PII | Regex | Mask-all |
|---|
Partial cells show n/100; compare only when all reach 100.
Precision and recall
| Source | Jev P | Jev R | Luna P | Luna R | GLiNER P | GLiNER R |
|---|
P = precision; R = recall.
Method
PIIMB (CC BY-NC 4.0): 100 short English sentences each from Ai4Privacy, Nemotron-PII, Gretel, and Privy. This is a pilot, not the full benchmark.
Jev chooses whitespace spans at probability ≥0.5. Luna returns character offsets in JSON. GLiNER Multi PII v1 runs on CPU at threshold 0.3. All three use the same text and gold masks, but different workflows. Gold never enters inference.
Regex masks emails, URLs, and long numeric strings. Mask-all shows the whitespace tokenizer's maximum recall, not an F2 ceiling. We score character-level, label-agnostic F2.
Sampling, models, and cost
We hash sentence IDs with seed pii-jev-v1, skip duplicate IDs, and keep nonempty sentences ≤240 characters and ≤32 whitespace tokens. Dataset revision 7d797b9fc8dc1942cef60fbe532e7d1a0e31b655; sample SHA-256 9230a9d372021759ed89b237f1cba9994cdee59c393e416b0c8c9332c76801ef.
Jev 1.13 asks one yes/no question per candidate in a single call. Luna uses the 22 Sept 2026 revision through OpenRouter Azure US, reasoning off, with PIIMB's published field names. Two broad-Azure setup calls were excluded before provider selection. GLiNER is Apache-2.0 and uses the same published labels. Jev, regex, and mask-all retain their original audited scores.
Jev: 364,706 input tokens, estimated $0.0153. Luna: $0.0171 provider-reported API usage; failed calls may add unobserved cost. GLiNER: no API fee, CPU time on an existing VM.
Limits
Short English sentences cannot establish safe redaction on real documents. Labels include dates, places, and demographics. Luna returned 19 invalid spans or JSON outputs, scored as misses. Public test data may overlap training. No sampled PIIMB sentences or their per-case predictions are published.
Reproducibility
Illustrative example
Synthetic sentence, not part of the evaluation. We sent it to Jev 1.13 with the Ai4Privacy scope.
Ava Chen emailed ava.chen@example.com.
Request construction:
const text = "Ava Chen emailed ava.chen@example.com.";
const scope = "Mask every annotated personal data field: names, title, age, gender, dates, city, address, phone, email, IDs, tax, passport, payment numbers.";
const spans = tokenize(text); // whitespace tokens, edge punctuation trimmed
const request = {
model: "jev-1.13.0",
state: { text },
questions: Object.fromEntries(spans.map(({ start, end }, i) => [
`s${i}`, { type: "noul",
instructions: `In \`text\`, should characters ${start}-${end} (${JSON.stringify(text.slice(start, end))}) be masked? ${scope} Answer yes only if this span is part of a field to mask.`
}
]))
};
Response from that call:
{
"model": "jev-1.13.0",
"answers": {
"s0": { "type": "noul", "noul": 0.88 }, // Ava
"s1": { "type": "noul", "noul": 0.92 }, // Chen
"s2": { "type": "noul", "noul": 0.02 }, // emailed
"s3": { "type": "noul", "noul": 0.98 } // ava.chen@example.com
},
"usage": { "input_tokens": 568, "output_tokens": 72 }
}