Blauweiss Teleprinter — Mod. BW‑26 Ready PrintLN 1207

Measuring the machines: how RIGOR scores a reservoir‑engineering agent

Ian Matejka · 14 Aug 2026 · 12 min

Agents are arriving in reservoir engineering faster than the field can agree on how to judge them. RIGOR is our attempt at a yardstick: hand-authored OPM Flow tasks, verifier-grade scoring, and a baseline that keeps everyone honest.

Reservoir simulation has always had a trust problem, and agents inherit it with interest. When a person writes a simulator deck, you can ask them why they chose a relative-permeability model or where a well constraint came from. When an agent writes one, the deck either has to carry its own evidence — or you need a common, public way to measure whether the agent can be trusted with the job at all.

RIGOR is that measurement. It puts an agent harness through hand-authored OPM Flow tasks spanning the actual work of the discipline: writing simulator decks from scratch, modifying and debugging existing ones, and analyzing results. Every task is hand-authored and expert-reviewed. Gold references stay verifier-only, so nothing in the benchmark can leak into a training set and quietly inflate a score.

◆

Scored on

01Deck validity — is the generated deck well-formed and runnable? 02Semantic requirements — does it capture what the case asked for? 03Successful runs — does the simulation complete? 04Output match — do results line up with the gold reference? 05Analysis correctness — is the engineering read-back right?

A benchmark that only measures a fully-tooled harness measures the harness, not the agent. So every RIGOR run is paired with a naked baseline: the same model, stripped of guidance, skills, and tools. The gap between the two is the honest number — the real lift your scaffolding provides.

A score you can’t interrogate is just a bigger black box. The point of RIGOR is that every number decomposes back into a task, a deck, and a run you can read.

Because a yardstick only works if everyone is allowed to hold it, we’re open-sourcing the task harness and the scoring criteria — so the whole field can measure progress on the same footing, including against CLARISSA. When our own agent has a bad criterion, that’s not a marketing problem; that’s the roadmap.

The repository lands soon. If you run agents against simulators — or you’re the engineer who would have to trust one — we’d like RIGOR to be the first thing you point at it.

· · · End of Printout · · ·