Most models that classify text are trained on whatever labelled datasets happen to exist: sentiment here, topics there, a few hundred tasks collected from benchmarks. The result is a model shaped by the accidents of what people chose to label. It can be excellent at the questions those datasets ask, and you have little idea which questions it has never been taught.
When we started Cox, a small model that answers typed questions about text with calibrated probabilities, we asked a different question first: given a piece of text, what kinds of question can be asked about it at all?
Two short lists
Every question Cox answers is pinned down by two things, each drawn from a short list.
The answer form. A decision comes out in one of three shapes, which we call primes:
- yes / no: a probability that the answer is yes;
- choice: a probability for each of a set of named options;
- score: a probability for each level of an ordered scale, so the answer can sit between two levels.
Everything else a caller might want, such as several yes/no features at once or a score per aspect, is built from these three.
What the question is about. We call this the target, and ten of them cover what we have seen so far: a property of the text (topic, tone, urgency); the presence of something in it (a date, an amount); a relation between two texts; compliance with a rule carried in the question; selection of the best item from a set; grading against a rubric; counting; change between two versions; attribution (who said or did what); and consistency (whether the parts of a text cohere).
Cross the two lists and you get a grid. Each cell is a skill. Hover over a cell to see an example of the question it stands for.
| what it's about ↓ answer form → | Yes / no a probability of yes | Choice a probability for each option | Score a probability for each level of an ordered scale |
|---|---|---|---|
| Property an attribute of the text: its topic, tone, urgency, intent | |||
| Presence whether something is in the text: a date, an amount, a name, a claim | |||
| Relation how two texts relate: support, contradiction, paraphrase, similarity | |||
| Compliance whether the text satisfies a rule carried in the question | |||
| Selection which item of a set best meets a criterion | |||
| Grading quality against a rubric or a reference | |||
| Counting how many of something, in buckets | |||
| Change what differs between two versions, and how much | |||
| Attribution who said or did what, or what caused what | |||
| Consistency whether the parts of a text cohere |
Hover over, focus or tap a cell to see an example. N/A cells are not gaps: each is a remapping of a cell that exists.
What the grid showed us
The first time we placed our training data in the grid, the picture was lopsided. Almost everything sat in the top row: properties of a text, asked as yes/no or as a choice. Of the 176 instruction-style tasks in our mix, 135 were binary. Whole rows were empty. Nothing taught the model to grade against a rubric, to count, to compare two versions, to say who agreed to what, or to pick the best item from a list. Graded relations (how similar are these two texts?) had no training data at all.
None of that was visible from a list of datasets. It was obvious from the grid.
The target decides where the truth comes from
The second thing the grid gave us was a plan. Each target implies where a correct label can come from:
- For presence, counting, change, compliance and consistency, code can decide the right answer first: insert a date or don’t, put four action items in the notes, edit one clause, apply the rule, break one sentence. A language model then writes natural text around that decision, and a second model checks it blind. The label is right by construction, not by someone’s opinion.
- For property and grading, the label is a judgement, so it needs people or a strong teacher model, and it carries their uncertainty.
- Relations sit in between: some are structural, some are judgements.
So the empty cells became a generation agenda. We built families of examples for counting, change, compliance, selection, grading, attribution, record matching and number comparison, each targeting cells the existing data left empty.
Four cells that are not gaps
Four cells are marked N/A: presence, compliance, selection and attribution, each asked as a score. At first they look like gaps. They are not questions of their own.
“How present is a date, on a 0-3 scale?” is the yes/no presence probability mapped onto the caller’s own scale. “Rate this claim’s compliance from 1 to 5” either rescales the yes/no answer or hides several rules inside one number. Because the probabilities are calibrated (an answer given at 80% is right about 80% of the time), callers can draw those lines themselves, and the model needs nothing new to learn.
Measuring by cell
The grid also changed how we measure. “How good is the model?” has no single answer. “Can it count? Can it grade against a rubric it has never seen? Can it tell which of two people agreed?” do. We score every question in our evaluation under its cell, keep real-world test data apart from generated data, and look at the grid rather than a single average. A strong average can hide an empty row.
Here is what that looks like for our current research build (9 October 2026), on the generated families built to fill the empty cells. New texts are unseen examples from the same domains the model trained on; new domains are whole domains held out of training (a counting family trained on lists of fruit and cities, tested on tools and rivers, for instance). Skill is 0 for always giving the most common answer and 1 for always being right; calibration error is the gap between stated confidence and observed accuracy.
| Family (target) | New texts: skill | New domains: skill | New domains: calibration error |
|---|---|---|---|
| Change between versions | +1.00 | +0.99 | 0.006 |
| Counting | +1.00 | +0.99 | 0.006 |
| Compliance with a rule | +0.98 | +0.74 | 0.061 |
| Selection of the best item | +0.97 | +0.79 | 0.094 |
| Graded levels (urgency, severity, formality) | +0.97 | +0.93 | 0.070 |
| Relations between two texts | +0.94 | +0.96 | 0.009 |
| Number comparison | +0.80 | +0.88 | 0.059 |
| Record matching | +0.59 | +0.54 | 0.098 |
Two things stand out. Where the label is decided by code, the model learns the skill almost perfectly and carries it to new domains (counting, change, relations). Where a skill depends on careful reading of every field (record matching) or on applying an unfamiliar rule in an unfamiliar domain (compliance, selection), new domains are markedly harder, and those are the rows we are now generating more data for. These are our own generated tests, written in the format the model was trained on; real-world datasets are scored separately, in our comparison of decision models.
What the grid is not
It is not a taxonomy of everything language models do. Cox reads and decides; it doesn’t write, so generation and free-text extraction are outside it by design. Two further axes refine a cell: the shape of the input (one text, a pair, a list, a record, a conversation, code), and the source of truth above. And ten targets are where we are today, not a claim that the list is complete.
But if you are building a model that makes decisions about text, or a benchmark that tests one, placing each question in the grid is a cheap exercise. It tells you what you are measuring, and, more usefully, what you are not.