RIGOR Benchmark
Measure reservoir-engineering agents with simulator-backed tasks.
RIGOR is Blauweiss AI's benchmark for reservoir-engineering agents, scoring simulation deck authoring, editing, repair, successful runs, output agreement and engineering analysis.
RIGOR—Reservoir Input Generation Output Review—is a benchmark framework for reservoir-engineering agents. It uses hand-authored simulation tasks and verifier-grade criteria to test authoring, editing, repair and analysis.
What RIGOR measures
- Deck validity
- Semantic task requirements
- Successful simulator execution
- Output agreement
- Engineering-analysis correctness
Why the baseline matters
RIGOR is intended to separate model capability from harness capability. Comparing a fully tooled system against a stripped baseline helps measure the real value added by prompts, tools, validation and agent scaffolding.