Tarski: tune how diffusion models make decisions

Many business workflows need a small, explicit decision: which team should receive a request, whether a condition is met, or which rating best fits a document. You can write down the possible answers and collect examples of what a good decision looks like.

Tarski helps you use those examples to tune how a diffusion language model reads the questions. It compares ways of presenting the same task, selects a read policy, and adjusts the confidence scores. The model's weights stay fixed. The result is a JSON configuration you can evaluate and use for local inference.

What a read policy changes

A diffusion decision reader places answer slots alongside a description of the situation and the questions about it. How that input is arranged, and how the answer slots are initialized, can affect the resulting label probabilities. Tarski turns those choices into settings you can compare on your own labeled examples.

SettingWhat changes
Prompt layoutWhere the situation and questions appear in the prompt, including stock and state-first formats
Question orderWhether questions keep their given order, reverse, or rotate through a different starting point
Answer-slot initializationWhether the reader uses its native slot initialization or a vocabulary-mean embedding
Confidence temperatureHow concentrated the output probabilities are, fitted separately for each question type

The first three settings affect the read itself. Temperature calibration changes the confidence distribution without changing the highest-ranked label. A better confidence estimate and a more accurate decision are separate things to measure.

Tarski selects one global combination of layout, order, and slot setting from the candidates you supply. It does not invent a different policy for every incoming record. The fitted file records the chosen policy, model identity, seed, and data fingerprints so you can keep track of what was evaluated.

A concrete example: routing a support request

Consider the message, “I was charged twice.” You might ask two questions about it: which team should handle it, and whether the service is unavailable. A human-labeled example could assign billing to the first question and no to the second.

Tarski accepts records with a stable ID, a state, a dictionary of questions, and optional gold labels. The state can be text, an object, or an array. It supports three question types:

  • Choice: select one of 2–255 named criteria, such as billing or technical support.
  • Yes/no: answer a binary question; the interface calls this type noul.
  • Score: choose among 2–10 described rating levels, using string labels such as "0" and "1".

Supply gold labels to select and evaluate a policy. Incoming records do not need those labels for prediction. The output includes a selected label, probabilities for the available labels, and a confidence score for each question.

Your application decides what happens next. For example, it could route some results automatically and send uncertain cases to a person, after testing that rule on representative data. Tarski provides the decision output; it does not connect itself to your help desk or implement that review workflow.

Try the file workflow before loading a model

The current source package requires Python 3.12 or newer. Start in a virtual environment. Cloning the repository gives you the example files as well as the CLI:

git clone https://github.com/looskis/tarski.git
cd tarski
python3 -m venv .venv
source .venv/bin/activate
python -m pip install .
tarski schema

You can then run the supplied offline example in a new working directory:

mkdir -p work
tarski init --backend mlx --out work/base.json
tarski validate --data examples/calibration.reads.jsonl --kind reads
tarski tune --config work/base.json \
  --reads examples/calibration.reads.jsonl --out work/fitted.json
tarski eval --config work/fitted.json \
  --reads examples/evaluation.reads.jsonl

These fixtures contain synthetic, already-collected probabilities. They demonstrate validation, policy selection, configuration output, and evaluation. They do not measure the quality of a model. These commands use the Python standard library and do not load MLX, PyTorch, or model weights. Existing output files are protected from overwrite unless you explicitly request it.

If you only need the installed CLI, the equivalent repository install is pip install git+https://github.com/looskis/tarski.git. That does not provide a working checkout of the example files.

Use your own labeled examples

For actual model reads, follow the backend installation instructions. The adapters use a pinned OpenJev prompt/canvas implementation plus the selected backend: MLX for DiffusionGemma on Apple silicon, or PyTorch for DiffusionGemma or LLaDA on CUDA. The first collection or prediction may download model weights. The documented 4-bit DiffusionGemma checkpoint is about 14 GB, before runtime memory requirements.

With the backend installed and two separate labeled files prepared, the core loop is:

tarski doctor --config work/base.json
tarski validate --data calibration.jsonl
tarski collect --config work/base.json --data calibration.jsonl \
  --out work/calibration.reads.jsonl \
  --formats stock state_first --orders given reversed rotate:1 \
  --slots native mean
tarski tune --config work/base.json \
  --reads work/calibration.reads.jsonl --out work/production.json
tarski collect --config work/production.json --data heldout.jsonl \
  --out work/heldout.reads.jsonl
tarski eval --config work/production.json --reads work/heldout.reads.jsonl

The example sweep runs 12 candidate reads per state: two layouts, three orders, and two slot settings. Start small. doctor checks dependencies, not available memory or whether weights are cached.

The tuning report describes performance on the calibration examples used to select the settings. Use a separate held-out file to assess whether those choices help on other examples. eval rejects IDs already used in calibration and checks state hashes when available. Stable IDs still matter: older experiment files without state hashes cannot detect the same state under a new ID.

Evaluation reports accuracy and probability-quality metrics, including Brier score, negative log likelihood, and expected calibration error. Compare both raw and calibrated probabilities. Calibration can fail to improve a held-out sample, especially when the labeled set is small or unrepresentative.

Once the evaluation supports your use case, predict on new records:

tarski predict --config work/production.json \
  --data incoming.jsonl --out predictions.jsonl

Successful commands return JSON; predictions are JSONL, with one record per line. Progress and errors go to stderr, and tarski schema describes the interface for scripts and agents.

What the published measurements establish

The repository's experiment log reports the following comparisons for DiffusionGemma 26B-A4B with unchanged model weights:

ComparisonReported resultScope
Decision accuracy65.9% for the default order/random-slot single read; 71.2% with a selected order and vocabulary-mean slot400 typed-decisions states containing 2,000 decisions
End-to-end latencyAbout 995 ms per state for four decoder passes; about 560 ms for oneA separate 20-state timing sample on the same Apple-silicon laptop, using 4-bit weights and including prefill

The accuracy and timing figures come from separate comparisons. They are project-reported research results, not measurements of your workflow or a promise that every model benefits. In the repository's LLaDA experiments, vocabulary-mean slots performed worse than the native mask slot. Settings need to be selected for the reader and data you actually use.

Where to start, and what to watch

Choose one bounded decision with clear labels and enough representative examples to keep calibration and evaluation separate. Tarski can help compare read settings and confidence estimates without updating the model weights, but it still needs labeled data for tuning. Changing the checkpoint or fitted policy calls for renewed evaluation and calibration.

The current CLI processes records and results in memory and writes output when the command completes. It has no streaming or resume support. Each record's questions must fit in one answer canvas; oversized question sets fail explicitly instead of being split automatically. These limits favor small initial experiments before larger batch runs.

Collection and inference use local model backends. Plan for the initial model download and decide how your own application stores inputs, predictions, and logs. A local model does not by itself define the data boundary of every system you connect to it.

Get Tarski on GitHub, read the CLI reference, or explore the other Looskis tools.

Frequently asked questions

Does Tarski fine-tune the model weights?

The current read-policy CLI keeps model weights fixed. It selects input layout, question order, and answer-slot settings, then fits confidence temperatures from labeled reads. This is distinct from training the model weights.

Do I need labeled data?

Yes for tuning and evaluation. Gold labels tell Tarski which candidate policies work on your examples and allow it to fit confidence temperatures. Prediction on incoming records does not require gold labels.

Does confidence calibration change the chosen answer?

Temperature calibration changes the distribution of probabilities but preserves the highest-ranked label. Evaluate accuracy and confidence quality separately on held-out examples.

Can I try Tarski without downloading a model?

Yes. The repository includes synthetic reads for the init, validate, tune, and eval workflow. Those offline commands use the Python standard library. Actual collection and prediction require a backend, the documented OpenJev dependency, and model weights.

Will the reported speed and accuracy gains apply to my workflow?

They are measurements from specific project experiments, with separate accuracy and timing comparisons. Results depend on the model, input, hardware, and read policy. Evaluate on representative held-out data before relying on a setting.

Sources & further reading

Talk with us about your workflow →