MEASURED, NOT PROMISED
Evidence before
the headline.
Results for the launch reader, which is still the Standard reader in version 1.6. First published 30 September 2026; updated 3 October 2026.
How we measured
Just Another Classifier is a taught classifier: you give it a few examples of each option. For each public benchmark, we taught the launch reader with a fixed number of examples per intent from that dataset's training split, then tested on the full official test split. The test sets were Banking77 (3,080 messages), CLINC150 (4,500 in-scope messages) and MASSIVE en-US (2,974 messages).
- The score is top-1 accuracy: the share of messages whose first answer is the right intent.
- For 2, 5 and 10 examples, results are the mean of five random draws. “All” uses every training example once.
- Data loaders were pinned and checksum-verified. No test message was used for training, teaching or tuning.
Taught accuracy (%)
| Benchmark | 2 examples | 5 examples | 10 examples | All examples |
|---|---|---|---|---|
| Banking77 | 72.6 | 78.3 | 82.1 | 89.8 |
| CLINC150 | 82.7 | 86.7 | 89.1 | 93.4 |
| MASSIVE (en-US) | 73.0 | 73.8 | 75.4 | 82.5 |
Our previous internal engine, measured the same way, scored 67.9 / 75.6 / 80.2 / 90.8 on Banking77; 76.3 / 83.0 / 86.6 / 92.8 on CLINC150; and 64.1 / 66.9 / 69.9 / 80.6 on MASSIVE, in the 2 / 5 / 10 / All order. The launch reader improves most with few examples. With all Banking77 training examples, its 89.8% is one point below the earlier engine's 90.8%.
Saying no
A classifier should refuse a message that fits none of your options instead of guessing. On CLINC150's out-of-scope set, taught with every training example, the launch reader reached 87.8% balanced accuracy: 86.8% of in-scope messages were answered correctly and 91.9% of out-of-scope messages were refused. With 10 examples per intent, those figures were 82.9%, 81.0% and 91.5%. Refusal is a normal result with a receipt, not an error.
Macro-F1, the format of public intent tables
Many published intent-classification tables report macro-F1 rather than top-1 accuracy, and count CLINC150's out-of-scope messages as a 151st class. On 3 October 2026 we measured the same reader that way, once, with the method written down before the run. The test sets were the full official splits of Banking77 (3,080 messages) and CLINC150 including its out-of-scope messages (5,500 messages).
- Macro-F1 averages the F1 score of every class equally, so rare intents count as much as common ones.
- CLINC150's out-of-scope class was taught as decline examples: the 100 out-of-scope messages in its training split, in both settings. A refusal counts as the out-of-scope answer.
- For 10 examples, results are the mean of three random draws. “All” uses every training example once. No test message was used for training, teaching or tuning.
| Benchmark | 10 examples | All examples |
|---|---|---|
| Banking77 | 81.8 | 90.7 |
| CLINC150 with out-of-scope (151 classes) | 85.5 | 90.5 |
The Banking77 row uses the reader's top answer. In the product, a message the reader is unsure of is refused instead. Counting those refusals as wrong gives 81.5 and 90.3, with 3.3% and 2.0% of messages refused. On CLINC150, 81.7% and 88.0% of in-scope messages were answered correctly, and 91.4% and 90.7% of out-of-scope messages were refused. As with the results above, the reader's training data includes banking and payments domains written for training; no benchmark test message was used.
Speed and memory
Measured on a Windows PC with a 12th-gen Intel Core i9 (i9-12900K). Times are medians for the launch reader.
| Message | One core | Four cores |
|---|---|---|
| Typical message | 15 ms | 8.5 ms |
| Long email, about 185 tokens | 92 ms | 34 ms |
With Reading set to Standard, the engine stayed under 1 GB of memory, peaking at about 0.8 GB while reading long emails. Close Read and Second Opinion use more memory while in use; see the system requirements. Teaching a 40-option Flow with 10 examples each, including tuning, took about 3 seconds. Numbers questions do not load the text reader. Hardware and tasks differ, so these observations are a guide, not a speed promise.
Cold and warm reads
Our 3 October benchmark graphic gives a speed of 1.8 to 5.5 ms per message. Those times come from the macro-F1 runs above: one test message at a time, against a taught Banking77 (77 options) or CLINC150 (150 options) Flow, on the same PC with four cores.
- Cold: the reader sees the message for the first time. Most new messages are read cold. Measured again with every message cold, the median was 4.9 to 6.8 ms on four cores and 10.4 to 15.1 ms on one core. The graphic's 5.5 ms is a cold figure.
- Warm: the same message was read earlier while the engine was running, so the reader reuses its earlier reading. The median was 1.8 to 2.9 ms on four cores. The graphic's 1.8 ms is a warm figure. The engine keeps up to 8,192 recent messages this way, in memory only, and clears them when it closes.
The launch table above was measured in an earlier, separate run, and its times are the more conservative.
Where it is weaker
We publish losses beside wins.
- Very few examples: Banking77 accuracy was 72.6% with only 2 examples per option. Start with about 10.
- Near-identical options: on our hardest internal detail-sensitive test, top-1 accuracy was about 59% with 10 examples. Examples that show the difference help most.
- Unfamiliar writing styles: our internal cross-writer test scored about 64% with 10 examples. Rating real answers can add your customers' phrasing.
About the training data
The launch reader was trained on a wide range of customer-service domains written for training, including banking and payments. None of the benchmark test messages were used. These are measurements of the reader in the product you install, not a claim that every task will score the same.
Always improving
We first published these results on 30 September 2026 and added the macro-F1 results on 3 October 2026. Every release will be measured the same way and published here, including where it falls short. We are working on the next improvements now, and we are grateful to every customer who teaches it, rates its answers and tells us what to fix. Thank you.