Examples: questions in, probabilities out
Each example below is the exact code that ran and the output it produced, recorded on 6 October 2026 from Cox 4B (our current best model) through a coxlm server. Nothing is edited or selected after the fact; where the model is wrong or unsure, the note says so.
Every example starts with import coxlm,
from coxlm import Questions, questions, choice, score, yesno, multilabel, multiscore, order_items
and model = coxlm.connect("http://localhost:8000").
Routing support tickets
Three questions of different types, written as a class and answered together for each ticket.
class Ticket(Questions):
team = choice(["billing", "support", "sales"], instructions="Which team handles this?")
urgency = score(["low", "medium", "high"], instructions="How urgent is it?")
refund = yesno("Is the customer asking for a refund?")
tickets = model.decide_batch([
"My card was charged twice!! I want my money back today.",
"Hi, could you send me a quote for 50 seats on the annual plan?",
"The export button does nothing since yesterday's update. Not urgent, we can wait.",
], Ticket) My card was charged twice!! I want my money back today.
team · choice
- billing 84%
- support 11%
- sales 5%
urgency · score, expected level 2.82
- low 5%
- medium 9%
- high 86%
refund · yesno
- no 3%
- yes 97%
Hi, could you send me a quote for 50 seats on the annual plan?
team · choice
- billing 18%
- support 9%
- sales 73%
urgency · score, expected level 1.32
- low 76%
- medium 17%
- high 7%
refund · yesno
- no 96%
- yes 4%
The export button does nothing since yesterday's update. Not urgent, we can wait.
team · choice
- billing 10%
- support 42%
- sales 48%
urgency · score, expected level 1.13
- low 91%
- medium 7%
- high 3%
refund · yesno
- no 95%
- yes 5%
The first two tickets are routed confidently. The third is a bug report, which belongs with support; the model leans to sales, but only at 48% against 42% for support. A split that close is the signal to hand the ticket to a person instead of routing it automatically.
Matching records
The state can be a record as well as text: here an invoice and a bank-statement line.
class Payment(Questions):
match = yesno("Do the invoice and the bank line describe the same payment?")
matches = model.decide_batch([
{"invoice": "ACME Corp, $1,240.00, 2026-09-14",
"bank_line": "ACME CORPORATION 1240.00 USD 15SEP"},
{"invoice": "ACME Corp, $1,240.00, 2026-09-14",
"bank_line": "ACME HARDWARE 124.00 USD 15SEP"},
], Payment) invoice: ACME Corp, $1,240.00, 2026-09-14 bank_line: ACME CORPORATION 1240.00 USD 15SEP
match · yesno
- no 30%
- yes 70%
invoice: ACME Corp, $1,240.00, 2026-09-14 bank_line: ACME HARDWARE 124.00 USD 15SEP
match · yesno
- no 91%
- yes 9%
The first pair is the same payment written two ways, and the model says yes at a cautious 70%. The second pair differs in merchant and amount: no, at 91%.
Several yes/no features at once
multilabel asks one yes/no question per feature; each is answered independently.
schema = multilabel(["mentions a deadline", "mentions money",
"is a complaint", "threatens to cancel"])
ans = model.decide("We were billed $400 twice and need it fixed by Friday or we cancel.", schema) We were billed $400 twice and need it fixed by Friday or we cancel.
mentions a deadline · yesno
- no 3%
- yes 97%
mentions money · yesno
- no 3%
- yes 97%
is a complaint · yesno
- no 4%
- yes 96%
threatens to cancel · yesno
- no 3%
- yes 97%
All four features are present, and each is answered on its own, so asking for more features does not change the answers to the others.
Rating several aspects on one scale
multiscore rates each aspect on a shared scale. The score is the expected level (1 = poor, 4 = excellent).
schema = multiscore(["cleanliness", "staff", "wifi", "breakfast"],
["poor", "fair", "good", "excellent"])
ans = model.decide(
"The room was spotless and the staff were lovely, "
"but the wifi kept dropping and breakfast was cold.",
schema,
) The room was spotless and the staff were lovely, but the wifi kept dropping and breakfast was cold.
cleanliness · score, expected level 3.65
- poor 2%
- fair 5%
- good 18%
- excellent 75%
staff · score, expected level 3.19
- poor 3%
- fair 7%
- good 60%
- excellent 31%
wifi · score, expected level 1.21
- poor 89%
- fair 4%
- good 3%
- excellent 3%
breakfast · score, expected level 1.25
- poor 86%
- fair 6%
- good 5%
- excellent 3%
One review, four different ratings: high for cleanliness and staff, low for wifi and breakfast.
Putting events in order
order_items puts a set of items in order. In pairwise mode it asks, for every pair, whether one comes before the other (21 yes/no questions for 7 events, one call), then assembles the answers into an order.
events = [
"checkout latency alarms fire",
"the on-call engineer is paged",
"the configuration change is deployed",
"the change is rolled back",
"the incident report is written",
"checkout recovers",
"the database is found refusing connections",
]
review = (
"Post-incident review, checkout outage, 14 September. At 09:12 a configuration "
"change raised the database connection limit; it was deployed without a canary. "
"Twenty minutes later checkout latency alarms fired. The on-call engineer was "
"paged, found the database refusing connections, and rolled the change back. "
"Checkout recovered at 10:05, and the team wrote the incident report the next day."
)
result = order_items(model, events, mode="pairwise", context=review,
instructions="In what order did these events happen?") Order returned
- 1the configuration change is deployed
- 2checkout latency alarms fire
- 3the on-call engineer is paged
- 4the database is found refusing connections
- 5the change is rolled back
- 6checkout recovers
- 7the incident report is written
Probability of this exact order: 71% · self-consistency of the pairwise answers: 100% · orders compatible with the confident answers: 1
Every event is in its correct place. All 21 answers agree with one another (no contradictory cycles), only one order is consistent with them, and that order carries 71% of the probability.
The same events, one question per event
Score mode asks each event its position (7 questions instead of 21). It is cheaper and gets the broad shape, but here it misplaces several events that pairwise mode orders correctly.
# events and review as in the previous example
result = order_items(model, events, mode="score", context=review,
instructions="In what order did these events happen?") Order returned
- 1the configuration change is deployed
- 2the on-call engineer is paged
- 3the database is found refusing connections
- 4checkout latency alarms fire
- 5checkout recovers
- 6the change is rolled back
- 7the incident report is written
Expected position of each event (1 = first): the configuration change is deployed 2.3the on-call engineer is paged 3.0the database is found refusing connections 3.4checkout latency alarms fire 3.8checkout recovers 4.5the change is rolled back 5.3the incident report is written 5.4
Score mode gets the beginning and the end right but swaps the middle: it puts the page and the diagnosis before the alarm, and the recovery before the rollback. When the order matters, pairwise mode is worth the extra questions.
What can happen at the same time
Not every set of steps has one correct order. Pairwise mode returns the dependencies the model is confident about, the steps that can run in parallel, and the pairs it leaves free.
steps = [
"preheat the oven (5 mins)",
"chop the vegetables (15 mins)",
"set the table (5 mins)",
"roast the chicken (90 minutes)",
"make the salad (15 minutes)",
"serve the cooked dinner (5 minutes)",
]
result = order_items(
model, steps, mode="pairwise",
context="Cooking a roast-chicken dinner for guests tonight in the most efficient time possible.",
instructions="Which of these tasks must be finished before which others can start?",
) Order returned
- 1preheat the oven (5 mins)
- 2set the table (5 mins)
- 3chop the vegetables (15 mins)
- 4make the salad (15 minutes)
- 5roast the chicken (90 minutes)
- 6serve the cooked dinner (5 minutes)
Stages (steps in a stage can run at the same time)
- Stage 1 preheat the oven (5 mins) · set the table (5 mins) · chop the vegetables (15 mins)
- Stage 2 make the salad (15 minutes)
- Stage 3 roast the chicken (90 minutes)
- Stage 4 serve the cooked dinner (5 minutes)
Confident dependencies (must finish → before), probability
- preheat the oven (5 mins) → roast the chicken (90 minutes) 0.96
- chop the vegetables (15 mins) → make the salad (15 minutes) 0.92
- set the table (5 mins) → serve the cooked dinner (5 minutes) 0.94
- roast the chicken (90 minutes) → serve the cooked dinner (5 minutes) 0.96
- make the salad (15 minutes) → roast the chicken (90 minutes) 0.87
The confident dependencies are mostly what a cook would say: preheat before roasting, roast before serving, set the table before serving, chop the vegetables before making the salad. Preheating, setting the table and chopping can all start at once. One dependency has no basis in the task: the model says the salad must be finished before the chicken goes in (0.87).
The probabilities are calibrated on our test suite: answers given at 80% are right about 80% of the time on average. That is what makes a 48% answer a reason to ask a person. See how Cox compares with other decision models and the coxlm repository.