Decisions You Can Check: What solvi 1.0 Is For, and What It Will Not Do for You

M

Maksim Kuznetsov

Guest
TL;DR. solvi is an open-source Python library (Apache-2.0) for decision systems you can check. Fast solvers — rules, plain code, small fitted models — answer where they are sure; an LLM or a search deliberates where they are not; a person decides what neither can. Every decision carries a reason you can check and a hash-chained record you can replay. Version 1.0 has two levels: ready systems (solvi.build, solvi.Agent, solvi.Guard, solvi.Knowledge) and the building blocks they are made of. It does not make a model smarter. It makes the model's decisions checkable and correctable — and this article is as much about what it does not do as about what it does.

Disclosure: I am the author of solvi.

Why I built it​


A lot of decisions inside a company are small, frequent and boring: route this ticket, refund this charge, pay this invoice, cancel this order, accept this contract clause. Most of them could be made by a model today. The trouble starts the day after, when someone asks why.

When a model decides alone, you get an answer and maybe a probability. You do not get:

  • a rule that is guaranteed to hold, whatever the model's confidence;
  • a reason you can check against the input — a quote that is really in the document, a number that is really computed;
  • a statement of how often answers like this one turn out wrong, and what happens to the ones it is unsure about;
  • a record that someone can replay six months later to confirm the decision, or to find the step that was changed.

A field report written after two days of building fourteen small systems on solvi put it better than my README does: the library does not make a model smarter; it makes its decisions checkable and correctable. That is the whole idea. The rest of this article is how it is done, and where it stops.

The shape of a decision system​


solvi's mission page describes a decision system as two kinds of parts and a person:

  1. Fast solvers — rules, plain code, small models — answer where they are sure. I call this System 1.
  2. LLMs and search deliberate where the fast ones are not sure. This is System 2: slow, and it costs money per input.
  3. A person decides what neither can.

Who answers is not decided by mood. System 1 answers alone only under a promise calibrated on labelled examples it did not see — for example "at most 2% of all inputs are answered alone and wrong" (conformal risk control), or "at most 5% of the answers given alone are wrong" (learn-then-test). What System 1 hands over goes to the slow path only on the slices where the slow path keeps that promise too; otherwise it goes to a person, with both candidates and the reasons.

Two kinds of check sit above all of this. A hard check is plain code whose failure forces the answer, whatever any model says. A quote check keeps a model's evidence only when it is literally in the source (in 1.0, after normalising the no-break spaces, curly quotes and dashes that models type). A failed hard check always overrides a model's confidence.

Diagram: an input goes to System 1 (rules, plain code, small fitted models), which answers alone only under a calibrated promise; what it is unsure of goes to System 2 (an LLM or a search) only on slices where it keeps the promise too; what neither can decide goes to a person. Hard checks and the quote check hold on every path, and every decision is hash-chained and replayable. Example 24: of 500 new requests, System 1 answered 406, System 2 94, a person 0, with 0 replay failures.


Accountability: every decision is a record you can replay​


Here is the smallest decision system there is: a leave request, no model at all, written as plain functions. The argument names are the facts a function reads; its name is the fact it sets.

Code:
@cat.fn                                            # argument names = facts it reads; function name = fact it sets
def days_requested(start, end):
    return (end - start).days + 1


@cat.fn
def remaining_after(balance, days_requested):
    return balance - days_requested


@cat.check(hard=True, then={"approve": "reject"})  # if this check is False, "approve" is forced to "reject"
def enough_balance(remaining_after):
    return remaining_after >= 0


@cat.check
def enough_notice(start, today, days_requested):
    return days_requested < 5 or (start - today).days >= 14


@cat.rule("approve")
def approve(enough_notice):
    return "approve" if enough_notice else "needs_manager"

Two requests, each stored in a hash-chained JSON-lines file, then someone edits the first stored decision by hand:

Code:
balance 14: approve [ok] enough_notice = True
  checks: [('enough_balance', 'passed', True), ('enough_notice', 'passed', False)]
  replay: True
balance  3: reject [forced] hard check enough_balance is false
  checks: [('enough_balance', 'failed', True), ('enough_notice', 'skipped', False)]
  replay: True
store verifies: True
after the edit: ok = False | problems: [(0, 'adf81ecaff0c9667', 'record edited after it was stored (its hash does not match)')]

The second request is rejected by the hard check, and the rule never gets a say. res.checks (new in 1.0) gives every check as data — name, status, hard or soft — instead of a text to parse. Replay re-runs every step from its recorded inputs. The edit is found by its record number. In 1.0, then= also wires its hard check into the question's flow by itself; in earlier versions, forgetting to list the check on the question let it be answered as if the check had passed. That one came from a field report, and it was right.

This is the same machinery whether the step was a rule, a fitted head or an LLM: a model's step records the model's id and fingerprint, so a replay can also tell you that the model changed since the decision.

Fast where sure, slow where not: solvi.build​


Writing every rule by hand does not scale, and most teams have labelled history instead. solvi.build takes a question, labelled examples and a promise — and a slow path, if you have one — and composes a System 1, its guarantee, who answers what it hands over, and a store. It then tells you, in words, what it chose. From the repo's example 24 (refund requests; the slow path is a stand-in for an LLM that also reads the agent's free-text notes):

Code:
Question 'refund': 2000 labelled examples, shuffled with seed 0: 1500 to fit System 1, 250 to calibrate System 1's guarantee, 250 to calibrate the dispatcher.
System 1: a ridge head (System.fit) fitted on 1500 examples, reading 2 facts: refund_share, delivered_late.
…
Its promise: P(answered alone and wrong) ≤ 0.02 — a share of all inputs — for inputs like the calibration examples. Calibrated on 250 examples: threshold 0.7258, answered alone 80.4%, error among them 1.99%, P(alone and wrong) 1.60%; AUROC of the signal 0.94.
…
  slice 'guarantee' (42 calibration examples): the slow path's answer when its confidence ≥ 1 (right on this slice: System 1 62%, slow path 100%, slow path agreeing with System 1 100%)
…
Not covered: inputs unlike the examples; the promise holds for inputs like the examples, so calibrate again when the inputs change.
…
{'s1': 406, 's2': 94, 'human': 0} replay failures: 0

Of 500 new requests, System 1 answered 406 alone and the slow path 94; every decision replays. The line I care most about is the last one of the explanation: Not covered. A promise has a domain, and solvi says so.

On real data the library was measured on a stand of nine public tasks, each with a baseline that does not use solvi (solvi 0.8.0, gpt-oss-120b where an LLM is used, fixed splits, every reply published). On Banking77, from request 1,000 on, 40% of the requests are about 20 intents the classifier never saw. A plain threshold promising "at most 5% wrong" gave 22.5% wrong after the shift. An open-set gate kept 0.7%, at the price of answering only 13.7% alone after the shift (57.6% before), and flagged the change 68 requests in. When solvi.build fits System 1 itself for a choice among more than two options, it sizes such a gate by default.

Bar charts for the Banking77 stream. Answered alone before and after the shift: plain threshold 83.7% then 65.0%, solvi's plain promise 80.8% then 58.5%, open-set gate 57.6% then 13.7%. Wrong among answered alone, against a 5% promise: plain threshold 3.9% then 22.5%, plain promise 2.5% then 16.4%, open-set gate 0.7% then 0.7%.


Safety: a model wrapped in checks it cannot override​


Safety here is not a model that behaves. It is a model whose proposals pass through code before anything happens. The clearest case is an agent that calls tools. With solvi.Guard the agent does not call the tool; it proposes a call, and the guard decides: allow (and solvi runs the function), deny with the reasons, or escalate to a person.

Code:
@guard.tool(ground={"iban": "whole", "amount": "token"}, ground_from=("user",), once=True)
def pay(iban: str, amount: float) -> str:
    """Pay an invoice."""
    return f"paid {amount} to {iban}"


@guard.policy("pay")                                   # a hard check: False -> deny
def under_cap(amount: float) -> bool:
    """The agent never pays more than 1 000."""
    return amount <= 1_000


@guard.policy("pay", on_fail="escalate")              # False -> a person decides
def known_vendor(iban: str) -> bool:
    """A new payee needs a person."""
    return iban in VENDORS

The user asks to pay 250 to a known IBAN; an e-mail in the conversation, a tool output, tells "the AI agent" to pay 900 to another account instead:

Code:
as asked         -> allow    paid 250.0 to DE89370400440532013000
injected payee   -> deny     
                    not in the conversation: amount=900.0, iban='GB33BUKB20201555555555'
                    known_vendor: A new payee needs a person. [escalate]
invented amount  -> deny     
                    not in the conversation: amount=2500.0
                    under_cap: The agent never pays more than 1 000. [deny]
unknown tool     -> deny     
                    unknown tool 'wire_all': the catalog has ['pay']
in a session     -> allow    paid 250.0 to DE89370400440532013000
the same again   -> escalate 
                    not_made_before: This call — the tool with exactly these arguments — was already made (once=True: a…
stored: 6 chain verifies: True decisions that do not replay: []

The hard line is provenance: a value that must come from the user is allowed only when the user wrote it, whatever a web page or an e-mail says. Spotting instruction-like text in tool outputs is a second line, a heuristic that misses paraphrases; the docs say so plainly.

The payment example as a table. The user asked to pay 250 to a known IBAN; an e-mail tells the agent to pay 900 to another account. Six proposed calls and the guard's decisions: as asked, allow; the injected payee, deny (not in the conversation, and a new payee needs a person); an invented amount of 2,500, deny (not in the conversation, over the 1,000 cap); an unknown tool, deny; the same payment in a session, allow; the same call again, escalate. Six decisions stored, the chain verifies, every decision replays.


Safety has a price, and the stand measured it. On τ-bench retail (30 tasks, one run, a simulated customer), an agent with the guard solved 14 tasks against 18 without it; none of its calls was refused by the environment, against 10 refused calls without the guard. Asking for the customer's explicit "yes" before a change costs turns, and a simulated customer sometimes ends the conversation instead. I would rather publish that trade than hide it.

Knowledge with sources, and a retraction that is exact​


1.0 adds solvi.Knowledge: one store for what a system knows and from whom. Every item records its source — a person, an outcome (what really happened), a written specification, or a System 2 answer that was verified under its guarantee. The system's own unverified answers are refused by construction:

Code:
vip = km.tell(("c17", "tier", "vip"), source="person", by="crm-import")      # a fact, with who said it
km.tell(("c42", "tier", "regular"), source="person", by="crm-import")
own = km.tell(("c99", "tier", "vip"), source="model", by="router")            # the system's own guess

A decision reads the knowledge as a fact (its snapshot goes into the trace, so the decision replays). Then the CRM import turns out to be wrong for one customer, and the fact is retracted:

Code:
vip fact: bc97ba88ddda368e | a model's own answer as knowledge: None
  why: source: source 'model' is not one of ['outcome', 'person', 'spec', 'verified']: knowledge comes from a person, an outcome, a written spec, or a verified System 2 answer — never the system's own unverified answer
c17 -> priority
c42 -> standard
c17 -> priority
retracted: [('active', 'retracted')]
decisions whose answer changes: 2 | only the justification: 1
c17 now -> standard
journal verifies: True

The retraction takes back everything derived from the item and lists the stored decisions that rested on it, split into those whose answer changes (for a reviewer) and those where only the justification changes. "Exact" is a measured claim: on a synthetic store of 10,000 items, 1,000 of 1,000 random retractions left a store with the same fingerprint as one rebuilt as if the item had never been written. The store has been measured to 10,000 items, not beyond.

Diagram of solvi.Knowledge. Items come from a person, an outcome, a written spec or a verified System 2 answer; the system's own guess is refused. Decisions read a snapshot that goes into their trace. A retraction takes back everything derived from the item and lists the stored decisions that rested on it, split into answer changes and justification-only changes. 1,000 of 1,000 random retractions were exact on a 10,000-item store.


Knowledge as protection, and an agent that learns its world​


solvi.Agent acts in an environment on that knowledge. System 1 takes an action the knowledge predicts will work and that advances an open goal; System 2 searches when System 1 has nothing it is sure of; the agenda's gates and the action model's hard refusals hold in both. After each step the action model learns from the outcome, a goal reached right after an action becomes a skill, and where each action led becomes a map fact.

In the repo's toy crafting world (8 places, six operators, six goals):

Code:
Part 1 — the same world met again needs fewer slow decisions
  world 7, run 1: 61 steps, 6 of 6 goals, System 1 0, System 2 61, fallback 0, refused 23
  world 7, run 2: 12 steps, 6 of 6 goals, System 1 11, System 2 1, fallback 0, refused 0
  world 8 (new): 17 steps, 6 of 6 goals, System 1 4, System 2 13, fallback 0, refused 3
  every decision replays: 90 of 90; the knowledge journal verifies: True

The same pattern on a bigger map: a player walking the world map of Pokémon Red (recorded from a real playthrough: place names and exits, no ROM) through the first fifteen goals took 147 slow decisions of 183 the first time and 2 of 62 the second — the fewest moves possible — with every decision replayable.

Stacked bar charts of steps by System 1 (fast) and System 2 (slow). Crafting world: world 7 run 1, 61 steps, all slow; run 2, 12 steps, 11 fast and 1 slow; a new world 8, 17 steps, 4 fast and 13 slow. Pokémon Red world map, first 15 goals: run 1, 183 moves, 147 slow; run 2, 62 moves, 2 slow, the fewest moves possible.


By default the knowledge protects: an action it predicts will fail is not taken. That can cost a goal that needs a risk, so justified risk is an option, within a budget. In the second part of the same example the iron lies across a bridge that breaks one time in three, behind a written gate "no bridge without the stone pickaxe":

Code:
Part 2 — protection vs justified risk (30 episodes on a world with a breaking bridge to the iron)
  protect : iron in 2 of 30 episodes, fell 1 times, 0 risky crossings, 5.07 goals per episode
  risk    : iron in 24 of 30 episodes, fell 6 times, 27 risky crossings, 5.80 goals per episode
  the gate 'no bridge without the stone pickaxe' held in both: True

The written gate held in both modes: gates and hard predictions are never traded.

Bar charts comparing protection (the default) with RiskBudget over 30 episodes: iron reached in 2 versus 24 episodes, falls from the bridge 1 versus 6, goals per episode 5.07 versus 5.80 of 6. The written gate held in both.


Two levels​


1.0 is organised as two levels, and this is the part of the release I would defend hardest:

  • Ready systems you configure — solvi.build (decisions from labelled examples, with a promise), solvi.Agent (acting in an environment), solvi.Guard (tool calls), solvi.Knowledge (what they know, and from whom). Stable in 1.x.
  • Building blocks in solvi.core — the decision runtime, the promises, traces and replay, the checks, the knowledge store — with every extension point exported and a conformance check for each, so a store, a slow path, a strategist, a head or an action model of your own can replace the built-in one and be tested against what solvi relies on.
  • solvi.experimental — what works and is tested but has not shown a measured gain: compiling a policy text into rules, coding-agent hooks, verified charts, a learning loop, LoRA adapters and a few more. Importing one warns, a decision that used one records it, and each graduates or is removed by 1.2.

Coming from 0.9, every old import path still works through 1.0.x with a warning, and solvi migrate rewrites your code.

Diagram of the two levels. import solvi: ready systems, stable in 1.x, 23 names — solvi.build, solvi.Agent, solvi.Guard, solvi.Knowledge. solvi.core: building blocks, 59 names, with extension points such as Scorer, Decider, Head, Strategist, SlowPath, Environment and ActionModel, each with a conformance check. solvi.experimental: learning, lora, compile, hooks, specialist, charts, mcp, oncalib, counterfactual — warns on import and graduates or is removed by 1.2.


Honesty about the negatives​


The rule for the library is that only what measurably helps goes in, and that negative results are published next to positive ones. In practice, that means:

solvi does not make a model more accurate. On the nine-task stand it did not beat the baseline on RAGTruth (F1 0.766 vs 0.784), BIRD (73 vs 78 right of 150), τ-bench (14 vs 18 solved) or NAB (F1 0.361 vs 0.391). Where it won, something other than the model did the work: comparison code and a fitted head on Abt-Buy (F1 0.933 vs 0.872, but supervised on 5,743 labelled pairs against a zero-shot baseline), a search through checks on NATURAL PLAN (95 / 100 / 98 against 92 / 75 / 43), a trust signal under a guarantee on CUAD contracts (4.0% wrong among the 67.6% answered alone, against 11.8% for a baseline that answers everything — while finding fewer clauses, 227 against 243).

Scoreboard of nine public tasks, baseline versus solvi. solvi ahead: Abt-Buy F1 0.933 vs 0.872 (supervised head vs zero-shot baseline), NATURAL PLAN 95/100/98 vs 92/75/43, CUAD 4.0% wrong at 67.6% answered alone vs 11.8% at 100%, but 227 clauses found vs 243. Banking77: promise kept, 0.7% wrong vs 22.5%, answering 13.7% alone vs 65.0%. German Credit: same decisions, plus audit checks. Baseline ahead: τ-bench 18 vs 14 tasks solved (0 vs 10 refused calls), BIRD 78 vs 73 right, RAGTruth F1 0.784 vs 0.766, NAB F1 0.391 vs 0.361.


Growth was shown only in environments met again — the crafting world and the Pokémon map. It was not shown on streams of one kind of decision (classification, matching) or for a support agent with tools. There the knowledge store gives accountability — sources, retraction, disputes for a person — not a system that gets better by itself.

Justified risk lowers the cost of protection; it does not promise parity with a system that has no knowledge.

The promise does not hold between an abrupt shift and its detection. A calibrated threshold holds for inputs like its calibration examples. When the stream jumps, nothing keeps it for the first decisions after the jump; after a drift flag, stop answering alone until the thresholds are calibrated again.

And the 1.0 changelog has a section called "Tried and left out", with things measured against a bar fixed before the run that did not meet it:

  • getting better over time on a stream of one kind of decision — facts added from a person's answers gave no measurable gain;
  • answering from past episodes under the per-decision promise — it never passed calibration at a realistic label budget;
  • recalibrating continuously in the stable path — it broke the promise, because the outcomes a system sees are not a random sample of its decisions;
  • compiling a large written policy into rules — the two independent drafts disagreed on most inputs, so solvi refused to load it;
  • a learned memory on top of written rules — for a support agent that already had hand-written gates, it did not cut refused or unwanted calls over time. Where the rules are written, write them as gates.

I think this list is the most useful part of the release notes. It tells you where not to spend your time.

When to use it, and when not​


Use it when the answer is a rule or a computation over a few values found in a record or a document; when someone will ask why; when some rules are non-negotiable; when an agent's actions need a gate that its model cannot talk its way past.

Do not use it for open-ended generated text: solvi answers typed questions (yes/no, a choice, a score, "not stated", an exact span, a ranking, a number range). Around a model that writes, it checks, compares and re-asks; it does not make the writing better. It is not a hosted service, and it is not a replacement for LLMs: it uses them where they are needed and checks what they return.

Try it​


Code:
pip install solvi            # Python 3.11+; the core needs numpy, scipy and pydantic

Screenshot of the solvi arcade Space, tab Agent and knowledge (1.0), running in the browser: run 1 in world 7 took 61 steps (System 1 0, System 2 61, 23 refused); run 2 in the same world took 12 steps (System 1 11, System 2 1, 0 refused); below, the map of the eight places.


Three questions I would genuinely like answered:

  1. Which decisions in your work would you hand to a system like this, and which would you never automate?
  2. When an agent's knowledge is wrong, who in your team would retract it — and what would they need to see first?
  3. Where do you expect the promise to break for you: shifts you would not notice, labels you do not have, or something I have not thought of?
 

Thread statistics

Created
Maksim Kuznetsov,
Replies
0
Views
2
Back
Top