MEASURED, NOT PROMISED

Evidence before
the headline.

Results for the launch reader, which is still the Standard reader in version 1.6. First published 30 September 2026; updated 3 October 2026.

How we measured

Just Another Classifier is a taught classifier: you give it a few examples of each option. For each public benchmark, we taught the launch reader with a fixed number of examples per intent from that dataset's training split, then tested on the full official test split. The test sets were Banking77 (3,080 messages), CLINC150 (4,500 in-scope messages) and MASSIVE en-US (2,974 messages).

Taught accuracy (%)

Top-1 accuracy on official test splits
Benchmark2 examples5 examples10 examplesAll examples
Banking7772.678.382.189.8
CLINC15082.786.789.193.4
MASSIVE (en-US)73.073.875.482.5

Our previous internal engine, measured the same way, scored 67.9 / 75.6 / 80.2 / 90.8 on Banking77; 76.3 / 83.0 / 86.6 / 92.8 on CLINC150; and 64.1 / 66.9 / 69.9 / 80.6 on MASSIVE, in the 2 / 5 / 10 / All order. The launch reader improves most with few examples. With all Banking77 training examples, its 89.8% is one point below the earlier engine's 90.8%.

Saying no

A classifier should refuse a message that fits none of your options instead of guessing. On CLINC150's out-of-scope set, taught with every training example, the launch reader reached 87.8% balanced accuracy: 86.8% of in-scope messages were answered correctly and 91.9% of out-of-scope messages were refused. With 10 examples per intent, those figures were 82.9%, 81.0% and 91.5%. Refusal is a normal result with a receipt, not an error.

Macro-F1, the format of public intent tables

Many published intent-classification tables report macro-F1 rather than top-1 accuracy, and count CLINC150's out-of-scope messages as a 151st class. On 3 October 2026 we measured the same reader that way, once, with the method written down before the run. The test sets were the full official splits of Banking77 (3,080 messages) and CLINC150 including its out-of-scope messages (5,500 messages).

Macro-F1 (%) on official test splits
Benchmark10 examplesAll examples
Banking7781.890.7
CLINC150 with out-of-scope (151 classes)85.590.5

The Banking77 row uses the reader's top answer. In the product, a message the reader is unsure of is refused instead. Counting those refusals as wrong gives 81.5 and 90.3, with 3.3% and 2.0% of messages refused. On CLINC150, 81.7% and 88.0% of in-scope messages were answered correctly, and 91.4% and 90.7% of out-of-scope messages were refused. As with the results above, the reader's training data includes banking and payments domains written for training; no benchmark test message was used.

Speed and memory

Measured on a Windows PC with a 12th-gen Intel Core i9 (i9-12900K). Times are medians for the launch reader.

Reader latency by CPU allocation
MessageOne coreFour cores
Typical message15 ms8.5 ms
Long email, about 185 tokens92 ms34 ms

With Reading set to Standard, the engine stayed under 1 GB of memory, peaking at about 0.8 GB while reading long emails. Close Read and Second Opinion use more memory while in use; see the system requirements. Teaching a 40-option Flow with 10 examples each, including tuning, took about 3 seconds. Numbers questions do not load the text reader. Hardware and tasks differ, so these observations are a guide, not a speed promise.

Cold and warm reads

Our 3 October benchmark graphic gives a speed of 1.8 to 5.5 ms per message. Those times come from the macro-F1 runs above: one test message at a time, against a taught Banking77 (77 options) or CLINC150 (150 options) Flow, on the same PC with four cores.

The launch table above was measured in an earlier, separate run, and its times are the more conservative.

Where it is weaker

We publish losses beside wins.

About the training data

The launch reader was trained on a wide range of customer-service domains written for training, including banking and payments. None of the benchmark test messages were used. These are measurements of the reader in the product you install, not a claim that every task will score the same.

Always improving

We first published these results on 30 September 2026 and added the macro-F1 results on 3 October 2026. Every release will be measured the same way and published here, including where it falls short. We are working on the next improvements now, and we are grateful to every customer who teaches it, rates its answers and tells us what to fix. Thank you.

Back to Just Another Classifier