← ALL FIELD NOTES
01 / MODEL ENGINEERINGURDU-FIRST. EVIDENCE-LED.

Jev-Urdu:
Urdu-first AI,
measured in decisions.

A multilingual mmBERT-base encoder. An Urdu-first focus.
My work on decision-making across Urdu script, Roman Urdu, and mixed Urdu-English workflows.

A sculptural silver ribbon connects cobalt and orange spheres, suggesting two reading directions
Two directions. One decision. Original conceptual launch artwork.
URDU-PROMPT XNLI59.0%

Accuracy over 5,010 decisions

VS LAYA MULTILINGUAL+12.9pp

95% paired interval +11.2 to +14.6 pp

BATCH INFERENCE / TESLA T4279/s

Decisions per second · batch size 64

I built Jev-Urdu around a practical question: can a small decision model make useful choices when the conversation happens in Urdu?

That conversation rarely stays in one neat format. A customer writes in Urdu script, follows up in Roman Urdu, then drops an English product name into the same sentence. For a support system, the useful next step may be an issue category, a request for a human, or a check that a claim is supported by the supplied text.

I wanted a focused interface for those choices. Jev-Urdu takes text and typed questions, then scores the available answers. A choice question selects from options; a yes_no question returns a probability of yes. It does not write a free-form reply. That gives an application a clear response shape, while leaving the responsibility for correctness with the evaluation and the surrounding workflow.

The mmBERT-base
architecture.

Jev-Urdu uses the multilingual mmBERT-base encoder with a two-layer decision head. It has approximately 321.9 million parameters and is released under Apache-2.0. I built it from scratch, including the Urdu training, task interfaces, and benchmark engineering around the release.

The attraction for me is control: explicit options, a consistent Python interface, and batch inference on local hardware. Those are useful engineering properties. They do not, by themselves, make the answers reliable.

Model size matters.

Jev-Urdu has the same size as the multilingual comparison checkpoint and roughly 24% fewer parameters than Laya English. The routed result selects a checkpoint for each request; it is not a separate model.

Model sizes · published FP16 weight files
ModelParametersWeight file
Jev-Urdu321.9M643.8 MB
Laya multilingual321.9M643.8 MB
Laya English421.3M842.6 MB
Laya routed321.9M or 421.3M per request1.49 GB for both checkpoints
TypeSafe Jev 1.13Not publicly disclosedNot available for download

Disk sizes use decimal MB/GB and cover weights only, excluding tokenizers and runtime GPU memory. The routed total covers the two checkpoints used in this benchmark; one is selected per request. TypeSafe's public model reference does not disclose Jev 1.13's parameter count or downloadable weights.

02 / EVALUATION

Measure the claim.
Keep the context.

The figures here come from my saved 1 October 2026 benchmark run using Laya 0.3.22 on a Tesla T4. I checked the 32 published local accuracy rows against the saved prediction records. These are reported results from that run, not a live rerun each time this page loads.

This article covers three independent text benchmarks: Urdu natural-language inference with XNLI, Urdu domain classification with MASSIVE ur-PK (18 scenario classes), and Roman Urdu sentiment. Every local model received the same requests for a given condition. I kept the narrow internal workflow tests separate, and compared the TypeSafe API on its own matched sample.

Accuracy is only one view. Macro F1 gives each class equal weight. Expected calibration error (ECE) measures the gap between confidence and observed correctness. Log loss penalizes confident mistakes. The explorer lets you switch between them. Accuracy-change intervals use 2,000 paired bootstrap resamples of decisions, preserving the pairing between model answers; they describe this evaluation, not every future domain.

03 / INDEPENDENT BENCHMARKS

Explore the evidence.

Four complete suites. Four local models. The wins, the regressions, and the probability quality — in the same view.

EXPLORER 01 / FULL SUITESDownload metrics CSV
Question & option language
Measure

The share of decisions that match the reference answer. Higher is better. Prompt language changes the questions and options; passages stay Urdu or Roman Urdu.

Show models

Urdu prompts · Accuracy · Higher is better ↑ · common scale 0–100%

XNLI · Urdu inference

5,010 decisions
Jev-Urdu59.0%
Laya multilingual46.1%
Laya routed46.1%
Laya English34.6%

Accuracy change vs Laya multilingual: +12.9 pp95% paired interval +11.2 to +14.6 pp

MASSIVE · Urdu domains

2,974 decisions
Jev-Urdu43.3%
Laya multilingual48.4%
Laya routed48.4%
Laya English21.7%

Accuracy change vs Laya multilingual: −5.1 pp95% paired interval −6.6 to −3.7 pp

Roman Urdu · sentiment

2,048 decisions
Jev-Urdu47.2%
Laya multilingual44.0%
Laya routed42.8%
Laya English41.0%

Accuracy change vs Laya multilingual: +3.2 pp95% paired interval +0.7 to +5.6 pp

Read the figures as a table
Independent full-suite results · Urdu prompts · Accuracy
BenchmarkModelDecisionsAccuracy
XNLI · Urdu inferenceLaya English5,01034.6%
XNLI · Urdu inferenceLaya multilingual5,01046.1%
XNLI · Urdu inferenceLaya routed5,01046.1%
XNLI · Urdu inferenceJev-Urdu5,01059.0%
MASSIVE · Urdu domainsLaya English2,97421.7%
MASSIVE · Urdu domainsLaya multilingual2,97448.4%
MASSIVE · Urdu domainsLaya routed2,97448.4%
MASSIVE · Urdu domainsJev-Urdu2,97443.3%
Roman Urdu · sentimentLaya English2,04841.0%
Roman Urdu · sentimentLaya multilingual2,04844.0%
Roman Urdu · sentimentLaya routed2,04842.8%
Roman Urdu · sentimentJev-Urdu2,04847.2%
All paired accuracy comparisons and intervals
Jev-Urdu minus each local comparison model · Urdu prompts · percentage points
BenchmarkComparison modelChange95% paired interval
XNLI · Urdu inferenceLaya multilingual+12.9 pp+11.2 to +14.6 pp
XNLI · Urdu inferenceLaya routed+12.9 pp+11.2 to +14.6 pp
XNLI · Urdu inferenceLaya English+24.4 pp+22.4 to +26.2 pp
MASSIVE · Urdu domainsLaya multilingual−5.1 pp−6.6 to −3.7 pp
MASSIVE · Urdu domainsLaya routed−5.1 pp−6.6 to −3.7 pp
MASSIVE · Urdu domainsLaya English+21.6 pp+19.7 to +23.5 pp
Roman Urdu · sentimentLaya multilingual+3.2 pp+0.7 to +5.6 pp
Roman Urdu · sentimentLaya routed+4.4 pp+2.0 to +6.9 pp
Roman Urdu · sentimentLaya English+6.3 pp+3.8 to +8.7 pp

Where it helps.
Where it gives ground.

The clearest gain is Urdu-prompt XNLI: Jev-Urdu reaches 59.0% accuracy, compared with 46.1% for Laya multilingual. The change is +12.9 percentage points, with a paired 95% interval from +11.2 to +14.6. That supports a specific improvement in these inference decisions.

Roman Urdu sentiment improves in accuracy, from 44.0% to 47.2%, but macro F1 falls from 0.404 to 0.384 and ECE becomes worse. I would not translate that into a blanket sentiment-quality claim. On Urdu-prompt MASSIVE, accuracy drops from 48.4% to 43.3%. Specialization has a cost, and those results belong beside the headline.

Prompt wording also matters. With English questions and answer options, Urdu XNLI is essentially tied: 63.2% for Jev-Urdu and 63.1% for Laya multilingual. Its paired interval includes zero. There is no state-of-the-art claim here; the value is seeing exactly where this version improves and where a fallback may be the better choice.

04 / MULTILINGUAL FOUNDATION

Urdu as it is
actually written.

I describe Jev-Urdu as an Urdu-first decision model using the multilingual mmBERT-base encoder. The evaluated release focuses on Urdu script, Roman Urdu, and mixed Urdu-English workflows. The independent suites measure Urdu and Roman Urdu; mixed writing appears in the internal workflow harness. Further languages are an evaluation roadmap, not a measured release claim.

The English-prompt toggle changes the language of questions and options. The passage remains Urdu or Roman Urdu. It is not an English-text benchmark and does not establish broad multilingual performance.

ILLUSTRATIVE INPUTS / NO MODEL OUTPUT SHOWN
URDU SCRIPT

میرا آرڈر ابھی تک نہیں پہنچا۔

ROMAN URDU

Mera order abhi tak nahi pohancha.

MIXED URDU–ENGLISH

میرا order ابھی تک deliver نہیں ہوا۔

The internal workflow check

Alongside the independent tests, I used a template-based harness covering seven task families: support triage, sentiment, grounding and inference, consent, instruction boundaries, tool-result verification, and extraction. It contains 751 decisions from 320 source inputs: 301 Urdu-script, 133 Roman Urdu, and 317 mixed-script decisions.

Jev-Urdu answers all 751 correctly, and all 601 in the unseen-template filter. These are narrow internal workflow results. Related and overlapping task templates can make this environment easier than real messages; filtering unseen templates does not remove that limitation. I treat 100% as a check that the intended interfaces work within this harness, never as broad real-world accuracy.

05 / A SEPARATE COMPARISON

Same items.
A different comparison.

The TypeSafe Jev 1.13 API comparison uses 100 identical items per external suite and prompt condition. The internal sample has 100 inputs and 232 decisions. All models below are scored on those matched samples.

EXPLORER 02 / MATCHED SAMPLESACCURACY · HIGHER IS BETTER
Question & option language

100 identical items per external suite and prompt condition. These small-sample scores are separate from the full-suite results above. TypeSafe Jev 1.13 was accessed through its API.

XNLI · Urdu inference

Urdu prompts · 100 decisions
Jev-Urdu62.0%
TypeSafe Jev 1.1370.0%
Laya multilingual42.0%
Laya routed42.0%
Laya English30.0%
Majority baseline26.0%

Jev-Urdu minus TypeSafe Jev: −8.0 pp95% paired interval −19.0 to +3.0 pp · includes zero

Read matched scores and every API comparison
Matched sample accuracy · Urdu prompts · XNLI · Urdu inference · reported to three decimals
ModelDecisionsAccuracy
Jev-Urdu1000.620
TypeSafe Jev 1.131000.700
Laya multilingual1000.420
Laya routed1000.420
Laya English1000.300
Majority baseline1000.260
Jev-Urdu minus TypeSafe Jev · all Urdu-prompt samples
BenchmarkChange95% paired interval
Internal workflow suite+17.7 pp+12.9 to +22.8 pp
XNLI · Urdu inference−8.0 pp−19.0 to +3.0 pp
MASSIVE · Urdu domains−48.0 pp−59.0 to −38.0 pp
Roman Urdu · sentiment−16.0 pp−30.0 to −3.0 pp

The API is stronger on several general tasks. With Urdu prompts, its matched-sample accuracy is 70% on XNLI versus Jev-Urdu's 62% and 79% on MASSIVE versus 31%. XNLI's interval includes zero; the MASSIVE gap does not. The smaller sample makes uncertainty especially important.

These figures should be compared within this explorer, not against the complete-suite numbers above. The API scores are transcribed at the saved report's three-decimal precision. Jev-Urdu is my separate model project and has no affiliation with TypeSafe's Jev API.

06 / ENGINEERING REALITY

Useful speed.
Honest confidence.

Batch inference reaches 279.272 decisions per second on the Tesla T4, close to Laya multilingual's 279.840. Each timing covers 22,775 decisions at batch size 64. This suggests the specialization preserves the local batch profile; it does not tell us single-request latency, concurrent service behavior, or how a hosted API would compare.

LOCAL BATCH THROUGHPUT / DECISIONS PER SECOND ↑
Laya English85.792
Laya multilingual279.840
Jev-Urdu279.272

Tesla T4 · batch 64 · 22,775 decisions per model · no API timing

The larger production concern is confidence. On Urdu-prompt XNLI, Jev-Urdu's average confidence is 98.3% while accuracy is 59.0%. ECE is 0.393, versus 0.248 for Laya multilingual. The exported probabilities are overconfident on these independent tasks. A high score should not become an automatic trust threshold without validation and recalibration.

For a support workflow, I would start with routing suggestions and human review. For grounding, I would surface a decision alongside the passage and keep a fallback for ambiguous cases. Consent and tool-result checks should support explicit application rules, with review before consequential actions. I would measure each deployment's class balance, script mix, and failure costs before deciding which model gets the final say.

07 / NOW ON PYPI

A Python library.
Ready to install.

I have published jev-urdu 0.1.0 on PyPI, a Python library for using the model in real applications. Developed through LughaatNLP, it brings loading, built-in Urdu tasks, custom questions, and batching into one interface. Inference runs directly with PyTorch, Transformers, Safetensors, and Hugging Face Hub.

Install, load, ask

The library requires Python 3.10 or newer. Install the published release:

python -m pip install jev-urdu==0.1.0

The package name is jev-urdu; the Python import is jev_urdu. Loading uses the published Jev-Urdu checkpoint on Hugging Face. This example prints actual predictions:

import jev_urdu

model = jev_urdu.load()

print(model.triage(
    "میرا آرڈر دو ہفتے سے نہیں پہنچا۔ کسی نمائندے سے بات کرنی ہے۔"
))
print(model.sentiment(
    "Yeh phone bohat acha hai, battery bhi zabardast hai."
))

Importing the library does not download the model. The first load() downloads and caches it, then loads it for inference. A GPU is optional: the loader chooses available CUDA, then Apple MPS, then CPU. Use jev_urdu.load(device="cpu") to select CPU explicitly.

One interface, six built-in tasks

triage classifies a support issue, priority, and human-agent request. sentiment evaluates sentiment and dissatisfaction. check_claim compares a claim with supplied context. consent, detect_injection, and verify_tool_result expose permission, instruction-boundary, and tool-receipt checks. These are model judgments for the surrounding application to validate.

Define your own decisions

Use choice for a defined set of options and yes_no for a binary question:

import jev_urdu
from jev_urdu import choice, yes_no

jev = jev_urdu.load()
answers = jev.ask(
    "کل سے انٹرنیٹ بند ہے، کام رکا ہوا ہے۔",
    {
        "topic": choice("مسئلہ کس شعبے سے متعلق ہے؟",
                        ["بلنگ", "تکنیکی", "دیگر"]),
        "urgent": yes_no("کیا یہ فوری مسئلہ ہے؟"),
    },
)
print(answers)

ask() returns a dictionary keyed by your question IDs, with an answer and probability for each. Choice results include every option's probability. For yes/no questions, probability always means P(yes), even when the selected answer is false.

Batch messages, keep control

Apply the same questions to several messages with ask_many. Results keep the input order:

results = jev.ask_many(
    ["میرا بل غلط آیا ہے۔", "Internet kal se band hai."],
    {"urgent": yes_no("کیا یہ فوری مسئلہ ہے؟")},
    batch_size=8,
)
print(results)

For more control, predict and predict_batch expose typed results and a truncation flag. The loader also accepts a local checkpoint directory, local_files_only=True, and a model commit SHA as revision. Pinning the package version and the model revision are separate choices.

The library release packages the inference interface; the benchmark figures above remain results from the recorded model evaluation. Batch size, hardware, and input length affect deployment performance. Validate the returned probabilities for your use case.

08 / THE NEXT ITERATION

A release.
And a starting point.

With the model and Python library published, my next priorities are broader independent language evaluation, recalibration on representative inputs, and testing with real users. I also want to understand which prompt formulations preserve the inference gains without sacrificing domain and sentiment quality. Those are open engineering questions, not capabilities I can claim for this release.

Jev-Urdu is original work in how I have specialized, packaged, and measured an Urdu decision model. The strongest part of sharing it is making those choices inspectable: the useful gain, the uncomfortable regressions, and the work still ahead.