Causal evaluation for LLM systems
CIMO Labs is an open research effort on a practical question: when can a cheap signal, such as an LLM judge's score, stand in for the outcome you actually care about, and how much should you trust the result?
Our main software is Causal Judge Evaluation (CJE), an open-source Python library. It calibrates judge scores against a small sample of outcome labels and estimates policy means and paired differences, including sampling and calibration uncertainty. Reusing a calibration across policies or time requires evidence that it still applies; diagnostics make unresolved assumptions visible.
Start here
- Your AI Metrics Are Lying to You
Why the 'You're absolutely right!' meme reveals a deep flaw in AI evaluation, and how to fix it with calibrated surrogates.
November 2025 · 30 min read
- High Agreement, Wrong Decisions
Why iterating a judge prompt until it agrees with your labels is training on the test set.
February 2026 · 12 min read
- Calibrating LLM Judges to Business Value
Calibration, a holdout, and uncertainty for the outcome you actually care about.
March 2026 · 15 min read
- Belief Cartography: Turning LLM Product Intuitions into an Experiment Portfolio
Auditing LLM opportunity maps and settling them with experiments, with results on 500 real A/B tests.
July 2026 · 14 min read
Paper
Judge calibration, transport audits, and calibration-aware inference, benchmarked on Chatbot Arena prompts.
Landesberg · arXiv:2512.11150 · December 2025
Software
pip install cje-eval
