Trainable text classifiers
Train it on your data in minutes. More accurate than the generalists, answers in under 20 ms of server time (measured p50 12.7 ms) -- in over 50 languages. Your examples never train anyone else's model.
“Ich komme nicht mehr in mein Konto, das Passwort funktioniert nicht.”
account_access99.9%bug0.1%billing0.0%Jex trains a small model on top of a shared multilingual encoder. 10,000 examples train in about a minute on CPU.
Measured p50 12.7 ms on the serving CPU. One encoder pass and a tiny head. Send up to 50 texts per call; large batches under heavy concurrent load queue behind each other -- see the docs' limits.
Train in English, classify German, Spanish, Japanese or Chinese. Or train in any of them.
Accuracy, per-label precision and recall, a confusion matrix and a confidence cut, all on examples the model never saw.
Examples are used for your classifier only and discarded after training unless you ask us to keep them. Export or delete at any time.
Paste a CSV on this page, or POST examples from your app. The same classifier either way.
During the preview, the demo below gives you a key that lasts a day. Account keys arrive with launch.
A list of {text, label}, or {text, labels} for multi-label. Add a group (a conversation or document id) so near-copies never land on both sides of the test.
curl -X POST https://getjex.dev/v1/classifiers \
-H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
-d '{"name": "tickets", "examples": [
{"text": "I was charged twice", "label": "billing"},
{"text": "Export crashes on Safari", "label": "bug"}]}'
# 202 {"id": "c_…", "status": "queued"}
import requests
r = requests.post("https://getjex.dev/v1/classifiers",
headers={"Authorization": f"Bearer {JEX_KEY}"},
json={"name": "tickets", "examples": examples}) # [{"text": ..., "label": ...}]
classifier_id = r.json()["id"]
const r = await fetch("https://getjex.dev/v1/classifiers", {
method: "POST",
headers: { Authorization: `Bearer ${JEX_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ name: "tickets", examples }),
})
const { id } = await r.json()GET /v1/classifiers/{id} returns the status and, once ready, the report: accuracy and macro-F1 on a held-out test split, per-label precision and recall, the confusion matrix, and the confidence at which it is right 90% of the time.
curl -X POST https://getjex.dev/v1/classifiers/c_…/classify \
-H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
-d '{"texts": ["Refund my annual plan"]}'
# {"results": [{"label": "billing", "confidence": 0.94, "confident": true, "scores": {…}}]}
r = requests.post(f"https://getjex.dev/v1/classifiers/{classifier_id}/classify",
headers={"Authorization": f"Bearer {JEX_KEY}"},
json={"texts": ["Refund my annual plan"]})
print(r.json()["results"][0]["label"])
const res = await fetch(`https://getjex.dev/v1/classifiers/${id}/classify`, {
method: "POST",
headers: { Authorization: `Bearer ${JEX_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ texts: ["Refund my annual plan"] }),
})
console.log((await res.json()).results[0].label)The report lists the test examples it got most wrong. Add examples for the weak labels and train again; each training returns a fresh report so you can compare.
DELETE /v1/classifiers/{id} removes the model, its report and any kept examples, and returns a deletion receipt.
Send just your label names (and, better, a one-line description of each). POST /v1/classify answers immediately -- every call, including the very first one on a brand-new label set. The first few calls on a label set you have never sent before come back as mode: "zero_shot_warming" (an instant answer from label-embedding similarity, no training yet); once Jex has generated and fit a small model for that exact label set in the background (seconds, cached from then on), every later call is mode: "instant_head" and more confident. Common label shapes (spam/ham, sentiment, support intents, topics) are pre-warmed, so they never show the cold-start mode at all.
curl -X POST https://getjex.dev/v1/classify \
-H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
-d '{"labels": {"billing": "a payment, invoice, or charge issue",
"bug": "something in the product is broken",
"feature_request": "a request for a new capability"},
"texts": ["I was charged twice this month for one seat"]}'
# first call on this label set (zero-shot, warming a model in the background):
# {"results": [{"label": "billing", "confidence": 0.4092,
# "scores": {"billing": 0.4092, "bug": 0.3302, "feature_request": 0.2605}}],
# "mode": "zero_shot_warming", "ms": 51.0}
# a later call, same label set, once it has warmed:
# {"results": [{"label": "billing", "confidence": 0.804,
# "scores": {"billing": 0.804, "bug": 0.1375, "feature_request": 0.0585}}],
# "mode": "instant_head", "ms": 11.8}
import requests
labels = {"billing": "a payment, invoice, or charge issue",
"bug": "something in the product is broken",
"feature_request": "a request for a new capability"}
r = requests.post("https://getjex.dev/v1/classify",
headers={"Authorization": f"Bearer {JEX_KEY}"},
json={"labels": labels, "texts": ["I was charged twice this month for one seat"]})
j = r.json()
print(j["results"][0]["label"], j["mode"], j["ms"]) # 'mode' is zero_shot_warming at first, instant_head once warm
const labels = { billing: "a payment, invoice, or charge issue",
bug: "something in the product is broken",
feature_request: "a request for a new capability" }
const res = await fetch("https://getjex.dev/v1/classify", {
method: "POST",
headers: { Authorization: `Bearer ${JEX_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ labels, texts: ["I was charged twice this month for one seat"] }),
})
const j = await res.json()
console.log(j.results[0].label, j.mode, j.ms)
2-200 labels, 25+ word-or-two label descriptions recommended for the best cold-start answer. No training examples are billed or stored; generation happens on our side. See the docs for errors, limits and the trained-classifier mode above.
label, textThe demo allows 2 classifiers of up to 2,000 examples, kept for 1 day. The report needs about 25 examples per label.
Train the classifier to see accuracy and macro-F1 on a held-out test split, the most-common-label baseline, and precision and recall per label.
Jex vs Jev (by TypeSafe) on 13 public datasets, STORED numbers from a 2026-09-28 run -- no new calls to Jev. Every system saw the same 400 rows from each official test split. Jex's settings were chosen on separate development data before the test was read.
Both instant rows use the same 13 datasets as the table below, measured the same way, 2026-10-01. Second opinion sends only the specific low-confidence texts it checks to an open-model provider (DeepSeek, via our inference gateway) -- off by default, opt-in only. Docs.
| Dataset | 0 | 8 | 32 | 128 | full |
|---|---|---|---|---|---|
| AG News (topics, 4) · research licence, pending legal | 81.8Jev 86.8 | 81.2Jev 86.8 | 86.0Jev 86.8 | 88.8Jev 86.8 | 95.0Jev 86.8 |
| Emotion (6) · research licence, pending legal | 48.0Jev 62.3 | 39.0Jev 62.3 | 50.2Jev 62.3 | 57.5Jev 62.3 | 94.0Jev 62.3 |
| GoEmotions (28, multi-label) | 23.4Jev 24.9 | — | — | — | 56.4Jev 24.9 |
| Banking77 (intents, 77) | 72.5Jev 88.2 | 83.8Jev 88.2 | 91.5Jev 88.2 | 93.0Jev 88.2 | 93.5Jev 88.2 |
| CLINC150 (intents, 151) | 65.0Jev 71.0 | 77.5Jev 71.0 | 83.8Jev 71.0 | 87.8Jev 71.0 | 88.8Jev 71.0 |
| TREC (question type, 6) | 36.8Jev 87.8 | 66.2Jev 87.8 | 83.5Jev 87.8 | 89.5Jev 87.8 | 96.5Jev 87.8 |
| SST-2 (sentiment, 2) | 85.8Jev 97.0 | 85.5Jev 97.0 | 88.2Jev 97.0 | 88.2Jev 97.0 | 93.5Jev 97.0 |
| Yahoo Answers (topics, 10) · research licence, pending legal | 49.0Jev 67.5 | 50.7Jev 67.5 | 62.7Jev 67.5 | 63.0Jev 67.5 | 72.5Jev 67.5 |
| MASSIVE English (60) | 61.3Jev 83.0 | 71.5Jev 83.0 | 81.8Jev 83.0 | 85.8Jev 83.0 | 89.5Jev 83.0 |
| MASSIVE German (60) | 58.5Jev 79.0 | 63.5Jev 79.0 | 75.0Jev 79.0 | 82.8Jev 79.0 | 81.5Jev 79.0 |
| MASSIVE Spanish (60) | 57.8Jev 80.0 | 68.0Jev 80.0 | 78.8Jev 80.0 | 81.5Jev 80.0 | 82.0Jev 80.0 |
| MASSIVE Japanese (60) | 61.3Jev 82.8 | 69.2Jev 82.8 | 77.5Jev 82.8 | 81.2Jev 82.8 | 83.8Jev 82.8 |
| MASSIVE Chinese (60) | 55.8Jev 77.5 | 67.2Jev 77.5 | 72.8Jev 77.5 | 80.2Jev 77.5 | 82.5Jev 77.5 |
Honest picture: Jev leads at no examples or only a handful per label on almost every set, and leads at every size on SST-2 sentiment specifically. Jex catches up around 32-128 examples per label on most sets and leads on most sets with full training data. Instant mode (above) trails this table's "0 examples" column on average (53.8/62.2 vs 58.2) -- it's a different technique (open-weight synthetic generation vs. the original closed-model read) and not directly comparable cell-for-cell; shown honestly as its own line rather than merged into the table.
Latency: Jex p50 12.7 ms server time per classification (measured on the serving CPU, trained/ready-made head, single text); Jev's stored per-item latency on these runs was ~2-230 ms depending on the call shape, not independently re-measured here. Price: Jex $0.005-0.01 per 1,000 texts depending on mode (instant vs. ready-made/custom-trained; second opinion, opt-in, +$0.01/1,000 texts it actually checks); Jev (TypeSafe)'s published list price is $0.042 per million input tokens with output free, which works out to roughly $0.002-0.008 per 1,000 calls depending on text length -- cited from TypeSafe's own pricing, not a new call. Cells show Jex's score over Jev's (accuracy; micro-F1 for GoEmotions). Columns are examples per label; "0" means label names only. Measured 2026-09-28, stored numbers, no new Jev calls. Methodology
Select by name with {"classifier": "jex/..."} -- no training, no warm-up. Each trained once on licensed public data; numbers measured 2026-10-01.
| Model | Trained on | Held-out accuracy |
|---|---|---|
jex/spam | SMS Spam Collection (CC BY 4.0) + synthetic emails | 98.1% own test; 81.2% / 66.9% spam recall on Enron-Spam -- a domain-tuned eval, not held-out (the synthetic styles were chosen by looking at Enron's misses) |
jex/intent | CLINC150 (CC BY 3.0), full official split, 151 labels incl. out-of-scope | 98.3% accuracy, held-out test |
jex/intent-banking | Banking77 (PolyAI, CC BY 4.0), full official split, 77 labels | 94.3% accuracy, held-out test |
jex/sentiment | Synthetic (open-weight model, 8 review domains) -- no commercially-licensed sentiment corpus found | 85.75% on SST-2, eval only, never trained on |
jex/topic | Synthetic (open-weight model, 4 news categories) -- same reasoning as sentiment | 82.0% on AG News, eval only, never trained on |
Full numbers, model cards and licence notes: docs.
second_opinion sends the specific low-confidence texts it applies to (not your full batch) to DeepSeek, an open-weight model, through our inference gateway -- off unless you explicitly set second_opinion: true.Two per label is the minimum. For a trustworthy report card, aim for 25 or more per label; accuracy keeps improving into the hundreds.
Not yet on its own. Today Jex needs labelled examples. A cold-start mode for label names only is in progress.
The encoder covers more than 50 languages. You can train in one language and classify text in another, though accuracy is best when your examples include the languages you expect.
It depends on your labels and examples, which is why every model comes with a report measured on examples it never saw during training, plus the confidence level at which it is right 90% of the time.
Yes. Send labels: [...] instead of label and Jex trains a multi-label classifier with its own threshold.
See Privacy above. In short: used for your classifier only, discarded after training by default, deletable with a receipt.