Software that reads text keeps needing small decisions. Which team should handle this ticket? Is this alert urgent?
Do these two payment records match? Does this expense break a policy? A large language model can answer, but it
replies in prose you have to parse, it takes seconds, and its “I’m fairly sure” is a sentence rather than a number you
can act on.
Cox is our research into a better fit for those decisions.
Why we built it
Cox started earlier this year as a classifier for Vitund Sandbox’s events. A sandbox that records everything an agent does
produces a stream of events, and many of them call for a quick judgement: is this request sending out data the agent
read earlier, is this file sensitive, should this action wait for a person? Those decisions have to be cheap enough to
make on every event, and honest enough about their uncertainty to drive a threshold, so we set out to build a small
model that answers typed questions with calibrated probabilities. When the first general-purpose decision models
appeared, they showed how broadly the same approach applies, and we widened Cox from security events to
classification in general. Bringing it back into the sandbox, fine-tuned on sandbox event data, is on our roadmap.
What it does
You give Cox a text and a set of typed questions: yes/no, pick one of these options, or rate on this scale. It
returns a probability for every option of every question, from a single pass through a small model (4 billion
parameters, running on one consumer GPU). It can’t answer outside the options you gave, so there is nothing to parse
and nothing to go off-script.
The probabilities are calibrated: answers given at 80% confidence are right about 80% of the time. That is what
makes them usable in code. Automate above a threshold, send the rest to a person, weigh asymmetric costs, or notice
when the text supports two answers.
How it works
- Read, don’t write. We start from an open language model and train it to answer by reading. The text and all
the questions go in together, and each answer is read from the model’s internal state at the end of its question.
No text is generated.
- Each question is answered as if asked alone. Asking ten questions gives the same answers as asking them one at a
time, but in a single call.
- Map the question space first. Rather than starting from whichever datasets exist, we asked what kinds of
question can be asked about a text: an answer form (yes/no, choice, score) crossed with what it’s about (a property,
presence, a relation, compliance with a rule, selection, grading, counting, change, attribution, consistency). That
grid shows which skills are untrained and becomes the data plan.
- Truth by construction. Where no dataset exists, code decides the right answer first, an open model writes
natural text around that decision, and a second model checks it blind. The labels are correct by design.
How we measure
A fixed suite of 20 real-world tasks is never trained on: their texts are fingerprinted and refused if they ever
appear in training data. One half of the suite uses question shapes the model has trained on, in new settings. The
other half uses question shapes it has never seen. For each task we report accuracy, skill (improvement over always
guessing the most common answer) and calibration error.
We also check the test labels. Where a model and the dataset disagree, a person reviews the item, and corrections are
published as a separate errata list with a fixed rule for what changes. The original data is never edited. Where a
label records a fact, such as the star rating a reviewer actually gave, it stays, even when keeping it lowers our own
score.
Where it stands
As of October 2026, our best model (an average of several 4B training runs) reaches:
- skill +0.56 on familiar question shapes in new settings, and +0.61 on question shapes it never trained on;
- 90% accuracy on items the text clearly answers;
- calibration error of about 0.01 on long documents it never saw;
- the same scores at 4-bit (within 0.005 skill), at roughly 2.3 GB.
Not everything worked. Longer training fitted familiar question shapes harder and lost skill on unseen ones. Several
plausible ideas did nothing measurable. We record those results alongside the ones that did work.
What’s next
- New question families: grading against a rubric, attribution (who said or caused what), record matching, and number
comparison.
- A release build trained only on data that permits commercial use, with open weights.
- Write-ups of the method and results in Articles.
The Python package for running Cox, coxlm, is open source, with
worked examples of real inputs and outputs. The weights aren’t public yet; to
hear when they are, subscribe or email info@vitund.ai.