Jev-Urdu:
Urdu-first AI,
measured in decisions.
A multilingual mmBERT-base encoder. An Urdu-first focus.
My work on decision-making across Urdu script, Roman Urdu, and mixed Urdu-English workflows.

Accuracy over 5,010 decisions
95% paired interval +11.2 to +14.6 pp
Decisions per second · batch size 64
I built Jev-Urdu around a practical question: can a small decision model make useful choices when the conversation happens in Urdu?
That conversation rarely stays in one neat format. A customer writes in Urdu script, follows up in Roman Urdu, then drops an English product name into the same sentence. For a support system, the useful next step may be an issue category, a request for a human, or a check that a claim is supported by the supplied text.
I wanted a focused interface for those choices. Jev-Urdu takes text and typed questions, then scores the available answers. A choice question selects from options; a yes_no question returns a probability of yes. It does not write a free-form reply. That gives an application a clear response shape, while leaving the responsibility for correctness with the evaluation and the surrounding workflow.
The mmBERT-base
architecture.
Jev-Urdu uses the multilingual mmBERT-base encoder with a two-layer decision head. It has approximately 321.9 million parameters and is released under Apache-2.0. I built it from scratch, including the Urdu training, task interfaces, and benchmark engineering around the release.
The attraction for me is control: explicit options, a consistent Python interface, and batch inference on local hardware. Those are useful engineering properties. They do not, by themselves, make the answers reliable.
Model size matters.
Jev-Urdu has the same size as the multilingual comparison checkpoint and roughly 24% fewer parameters than Laya English. The routed result selects a checkpoint for each request; it is not a separate model.
| Model | Parameters | Weight file |
|---|---|---|
| Jev-Urdu | 321.9M | 643.8 MB |
| Laya multilingual | 321.9M | 643.8 MB |
| Laya English | 421.3M | 842.6 MB |
| Laya routed | 321.9M or 421.3M per request | 1.49 GB for both checkpoints |
| TypeSafe Jev 1.13 | Not publicly disclosed | Not available for download |
Disk sizes use decimal MB/GB and cover weights only, excluding tokenizers and runtime GPU memory. The routed total covers the two checkpoints used in this benchmark; one is selected per request. TypeSafe's public model reference does not disclose Jev 1.13's parameter count or downloadable weights.
Measure the claim.
Keep the context.
The figures here come from my saved 1 October 2026 benchmark run using Laya 0.3.22 on a Tesla T4. I checked the 32 published local accuracy rows against the saved prediction records. These are reported results from that run, not a live rerun each time this page loads.
This article covers three independent text benchmarks: Urdu natural-language inference with XNLI, Urdu domain classification with MASSIVE ur-PK (18 scenario classes), and Roman Urdu sentiment. Every local model received the same requests for a given condition. I kept the narrow internal workflow tests separate, and compared the TypeSafe API on its own matched sample.
Accuracy is only one view. Macro F1 gives each class equal weight. Expected calibration error (ECE) measures the gap between confidence and observed correctness. Log loss penalizes confident mistakes. The explorer lets you switch between them. Accuracy-change intervals use 2,000 paired bootstrap resamples of decisions, preserving the pairing between model answers; they describe this evaluation, not every future domain.
Explore the evidence.
Four complete suites. Four local models. The wins, the regressions, and the probability quality — in the same view.
The share of decisions that match the reference answer. Higher is better. Prompt language changes the questions and options; passages stay Urdu or Roman Urdu.
Urdu prompts · Accuracy · Higher is better ↑ · common scale 0–100%
XNLI · Urdu inference
5,010 decisionsAccuracy change vs Laya multilingual: +12.9 pp95% paired interval +11.2 to +14.6 pp
MASSIVE · Urdu domains
2,974 decisionsAccuracy change vs Laya multilingual: −5.1 pp95% paired interval −6.6 to −3.7 pp
Roman Urdu · sentiment
2,048 decisionsAccuracy change vs Laya multilingual: +3.2 pp95% paired interval +0.7 to +5.6 pp
Read the figures as a table
| Benchmark | Model | Decisions | Accuracy |
|---|---|---|---|
| XNLI · Urdu inference | Laya English | 5,010 | 34.6% |
| XNLI · Urdu inference | Laya multilingual | 5,010 | 46.1% |
| XNLI · Urdu inference | Laya routed | 5,010 | 46.1% |
| XNLI · Urdu inference | Jev-Urdu | 5,010 | 59.0% |
| MASSIVE · Urdu domains | Laya English | 2,974 | 21.7% |
| MASSIVE · Urdu domains | Laya multilingual | 2,974 | 48.4% |
| MASSIVE · Urdu domains | Laya routed | 2,974 | 48.4% |
| MASSIVE · Urdu domains | Jev-Urdu | 2,974 | 43.3% |
| Roman Urdu · sentiment | Laya English | 2,048 | 41.0% |
| Roman Urdu · sentiment | Laya multilingual | 2,048 | 44.0% |
| Roman Urdu · sentiment | Laya routed | 2,048 | 42.8% |
| Roman Urdu · sentiment | Jev-Urdu | 2,048 | 47.2% |
All paired accuracy comparisons and intervals
| Benchmark | Comparison model | Change | 95% paired interval |
|---|---|---|---|
| XNLI · Urdu inference | Laya multilingual | +12.9 pp | +11.2 to +14.6 pp |
| XNLI · Urdu inference | Laya routed | +12.9 pp | +11.2 to +14.6 pp |
| XNLI · Urdu inference | Laya English | +24.4 pp | +22.4 to +26.2 pp |
| MASSIVE · Urdu domains | Laya multilingual | −5.1 pp | −6.6 to −3.7 pp |
| MASSIVE · Urdu domains | Laya routed | −5.1 pp | −6.6 to −3.7 pp |
| MASSIVE · Urdu domains | Laya English | +21.6 pp | +19.7 to +23.5 pp |
| Roman Urdu · sentiment | Laya multilingual | +3.2 pp | +0.7 to +5.6 pp |
| Roman Urdu · sentiment | Laya routed | +4.4 pp | +2.0 to +6.9 pp |
| Roman Urdu · sentiment | Laya English | +6.3 pp | +3.8 to +8.7 pp |
Where it helps.
Where it gives ground.
The clearest gain is Urdu-prompt XNLI: Jev-Urdu reaches 59.0% accuracy, compared with 46.1% for Laya multilingual. The change is +12.9 percentage points, with a paired 95% interval from +11.2 to +14.6. That supports a specific improvement in these inference decisions.
Roman Urdu sentiment improves in accuracy, from 44.0% to 47.2%, but macro F1 falls from 0.404 to 0.384 and ECE becomes worse. I would not translate that into a blanket sentiment-quality claim. On Urdu-prompt MASSIVE, accuracy drops from 48.4% to 43.3%. Specialization has a cost, and those results belong beside the headline.
Prompt wording also matters. With English questions and answer options, Urdu XNLI is essentially tied: 63.2% for Jev-Urdu and 63.1% for Laya multilingual. Its paired interval includes zero. There is no state-of-the-art claim here; the value is seeing exactly where this version improves and where a fallback may be the better choice.
Urdu as it is
actually written.
I describe Jev-Urdu as an Urdu-first decision model using the multilingual mmBERT-base encoder. The evaluated release focuses on Urdu script, Roman Urdu, and mixed Urdu-English workflows. The independent suites measure Urdu and Roman Urdu; mixed writing appears in the internal workflow harness. Further languages are an evaluation roadmap, not a measured release claim.
The English-prompt toggle changes the language of questions and options. The passage remains Urdu or Roman Urdu. It is not an English-text benchmark and does not establish broad multilingual performance.
URDU SCRIPTمیرا آرڈر ابھی تک نہیں پہنچا۔
ROMAN URDUMera order abhi tak nahi pohancha.
MIXED URDU–ENGLISHمیرا order ابھی تک deliver نہیں ہوا۔
The internal workflow check
Alongside the independent tests, I used a template-based harness covering seven task families: support triage, sentiment, grounding and inference, consent, instruction boundaries, tool-result verification, and extraction. It contains 751 decisions from 320 source inputs: 301 Urdu-script, 133 Roman Urdu, and 317 mixed-script decisions.
Jev-Urdu answers all 751 correctly, and all 601 in the unseen-template filter. These are narrow internal workflow results. Related and overlapping task templates can make this environment easier than real messages; filtering unseen templates does not remove that limitation. I treat 100% as a check that the intended interfaces work within this harness, never as broad real-world accuracy.
Same items.
A different comparison.
The TypeSafe Jev 1.13 API comparison uses 100 identical items per external suite and prompt condition. The internal sample has 100 inputs and 232 decisions. All models below are scored on those matched samples.
100 identical items per external suite and prompt condition. These small-sample scores are separate from the full-suite results above. TypeSafe Jev 1.13 was accessed through its API.
XNLI · Urdu inference
Urdu prompts · 100 decisionsJev-Urdu minus TypeSafe Jev: −8.0 pp95% paired interval −19.0 to +3.0 pp · includes zero
Read matched scores and every API comparison
| Model | Decisions | Accuracy |
|---|---|---|
| Jev-Urdu | 100 | 0.620 |
| TypeSafe Jev 1.13 | 100 | 0.700 |
| Laya multilingual | 100 | 0.420 |
| Laya routed | 100 | 0.420 |
| Laya English | 100 | 0.300 |
| Majority baseline | 100 | 0.260 |
| Benchmark | Change | 95% paired interval |
|---|---|---|
| Internal workflow suite | +17.7 pp | +12.9 to +22.8 pp |
| XNLI · Urdu inference | −8.0 pp | −19.0 to +3.0 pp |
| MASSIVE · Urdu domains | −48.0 pp | −59.0 to −38.0 pp |
| Roman Urdu · sentiment | −16.0 pp | −30.0 to −3.0 pp |
The API is stronger on several general tasks. With Urdu prompts, its matched-sample accuracy is 70% on XNLI versus Jev-Urdu's 62% and 79% on MASSIVE versus 31%. XNLI's interval includes zero; the MASSIVE gap does not. The smaller sample makes uncertainty especially important.
These figures should be compared within this explorer, not against the complete-suite numbers above. The API scores are transcribed at the saved report's three-decimal precision. Jev-Urdu is my separate model project and has no affiliation with TypeSafe's Jev API.
Useful speed.
Honest confidence.
Batch inference reaches 279.272 decisions per second on the Tesla T4, close to Laya multilingual's 279.840. Each timing covers 22,775 decisions at batch size 64. This suggests the specialization preserves the local batch profile; it does not tell us single-request latency, concurrent service behavior, or how a hosted API would compare.
Tesla T4 · batch 64 · 22,775 decisions per model · no API timing
The larger production concern is confidence. On Urdu-prompt XNLI, Jev-Urdu's average confidence is 98.3% while accuracy is 59.0%. ECE is 0.393, versus 0.248 for Laya multilingual. The exported probabilities are overconfident on these independent tasks. A high score should not become an automatic trust threshold without validation and recalibration.
For a support workflow, I would start with routing suggestions and human review. For grounding, I would surface a decision alongside the passage and keep a fallback for ambiguous cases. Consent and tool-result checks should support explicit application rules, with review before consequential actions. I would measure each deployment's class balance, script mix, and failure costs before deciding which model gets the final say.
A Python library.
Ready to install.

I have published jev-urdu 0.1.0 on PyPI, a Python library for using the model in real applications. Developed through LughaatNLP, it brings loading, built-in Urdu tasks, custom questions, and batching into one interface. Inference runs directly with PyTorch, Transformers, Safetensors, and Hugging Face Hub.
Install, load, ask
The library requires Python 3.10 or newer. Install the published release:
python -m pip install jev-urdu==0.1.0The package name is jev-urdu; the Python import is jev_urdu. Loading uses the published Jev-Urdu checkpoint on Hugging Face. This example prints actual predictions:
import jev_urdu
model = jev_urdu.load()
print(model.triage(
"میرا آرڈر دو ہفتے سے نہیں پہنچا۔ کسی نمائندے سے بات کرنی ہے۔"
))
print(model.sentiment(
"Yeh phone bohat acha hai, battery bhi zabardast hai."
))Importing the library does not download the model. The first load() downloads and caches it, then loads it for inference. A GPU is optional: the loader chooses available CUDA, then Apple MPS, then CPU. Use jev_urdu.load(device="cpu") to select CPU explicitly.
One interface, six built-in tasks
triage classifies a support issue, priority, and human-agent request. sentiment evaluates sentiment and dissatisfaction. check_claim compares a claim with supplied context. consent, detect_injection, and verify_tool_result expose permission, instruction-boundary, and tool-receipt checks. These are model judgments for the surrounding application to validate.
Define your own decisions
Use choice for a defined set of options and yes_no for a binary question:
import jev_urdu
from jev_urdu import choice, yes_no
jev = jev_urdu.load()
answers = jev.ask(
"کل سے انٹرنیٹ بند ہے، کام رکا ہوا ہے۔",
{
"topic": choice("مسئلہ کس شعبے سے متعلق ہے؟",
["بلنگ", "تکنیکی", "دیگر"]),
"urgent": yes_no("کیا یہ فوری مسئلہ ہے؟"),
},
)
print(answers)ask() returns a dictionary keyed by your question IDs, with an answer and probability for each. Choice results include every option's probability. For yes/no questions, probability always means P(yes), even when the selected answer is false.
Batch messages, keep control
Apply the same questions to several messages with ask_many. Results keep the input order:
results = jev.ask_many(
["میرا بل غلط آیا ہے۔", "Internet kal se band hai."],
{"urgent": yes_no("کیا یہ فوری مسئلہ ہے؟")},
batch_size=8,
)
print(results)For more control, predict and predict_batch expose typed results and a truncation flag. The loader also accepts a local checkpoint directory, local_files_only=True, and a model commit SHA as revision. Pinning the package version and the model revision are separate choices.
The library release packages the inference interface; the benchmark figures above remain results from the recorded model evaluation. Batch size, hardware, and input length affect deployment performance. Validate the returned probabilities for your use case.
A release.
And a starting point.
With the model and Python library published, my next priorities are broader independent language evaluation, recalibration on representative inputs, and testing with real users. I also want to understand which prompt formulations preserve the inference gains without sacrificing domain and sentiment quality. Those are open engineering questions, not capabilities I can claim for this release.
Jev-Urdu is original work in how I have specialized, packaged, and measured an Urdu decision model. The strongest part of sharing it is making those choices inspectable: the useful gain, the uncomfortable regressions, and the work still ahead.