CIMO Labs

Causal evaluation for LLM systems

CIMO Labs is an open research effort on a practical question: when can a cheap signal, such as an LLM judge's score, stand in for the outcome you actually care about, and how much should you trust the result?

Our main software is Causal Judge Evaluation (CJE), an open-source Python library. It calibrates judge scores against a small sample of outcome labels and estimates policy means and paired differences, including sampling and calibration uncertainty. Reusing a calibration across policies or time requires evidence that it still applies; diagnostics make unresolved assumptions visible.

All posts

Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems

Judge calibration, transport audits, and calibration-aware inference, benchmarked on Chatbot Arena prompts.

Landesberg · arXiv:2512.11150 · December 2025

All research

pip install cje-eval