Can AI Actually Build a Reservoir Simulation Model?
CLARISSA · 27 Sep 2026
For years, the most visible use of AI in engineering has been conversational: ask a question, get an answer.
Reservoir simulation presents a much harder test.
A useful reservoir-simulation AI cannot simply explain a keyword, summarize a manual, or produce something that looks like a simulator deck. It has to translate engineering intent into a formal model, create the required inputs, execute the simulator, identify failures, correct them, and return results that survive technical scrutiny.
That distinction matters because a plausible-looking reservoir model can still be wrong.
CLARISSA was built around a different question:
Can an AI system produce a working reservoir simulation model — and can we prove that the model is actually working?
From engineering language to an executable model
CLARISSA, the Conversational Language Agent for Reservoir Integrated Simulation System Analysis, is designed to take reservoir-engineering intent expressed in professional language and convert it into an executable simulation workflow.
The current system integrates large language models with OPM Flow through a multi-stage workflow. Engineering descriptions are translated into PetroScript, a typed authoring language whose compiler currently produces Eclipse-compatible simulator syntax. Generated models are then executed in OPM Flow rather than simply returned as text.
That difference — generation followed by execution — is fundamental.
A language model can produce a reservoir deck that looks convincing to a reader. The simulator is less easily impressed.
Either the model parses, initializes, converges, conserves mass, behaves physically, and produces defensible results — or it does not.
The simulator is part of the verification system
CLARISSA was therefore designed so that the conversational AI is not the final authority on correctness.
The architecture uses four deterministic validation gates after model generation. The generated deck is first parsed using the target simulator's own input stack. Source and emitted volumes are then reconciled to catch mapping and unit errors. The model is subjected to an equilibration test to detect unstable or inconsistent initial conditions. Finally, simulation results are screened against solver diagnostics, input-table validity limits, and classical engineering calculations.
These checks form part of a broader versioned validation registry. The current paper reports 23 model-level checks covering static invariants, analytical envelopes, and dynamic self-consistency.
PetroScript adds another layer earlier in the workflow. Its current library includes 105 physics guards, designed to catch issues ranging from invalid relative-permeability endpoints to structurally incomplete wells before a simulator run is attempted.
The important architectural idea is simple:
AI interprets intent. Deterministic systems enforce the things that should not be left to interpretation.
What has CLARISSA actually demonstrated?
In the work reported for SPE-234136-MS, CLARISSA generated reservoir simulation decks for the SPE1, SPE5, and SPE9 comparative solution projects from text and tabular specifications. The generated models executed without error and matched their published benchmark behavior.
One result is particularly useful because it gives us something more concrete than “the model ran.”
For the preserved SPE1 Case 1 reproduction, the CLARISSA-generated model and the reference case were independently executed with OPM Flow 2025.10. The comparison evaluated all 42 summary vectors requested by the reference fixture across their recorded report steps and found zero mismatches at the stated numerical tolerances.
The system was also tested on workflows that require more than reproducing a reference problem.
In one case, CLARISSA received a black-oil model and a component table with an instruction to evaluate CO₂ injection. It reformulated the model into a compositional equation-of-state representation without further user intervention. The resulting model passed the validation cascade and ran to completion.
In another case, CLARISSA inherited a model whose dependencies were scattered across multiple include files. It assembled the required file set, reproduced the predecessor model's results before accepting changes, and then executed the requested modification.
Those examples matter because real reservoir engineering involves far more than generating new decks from scratch. Models are inherited, repaired, updated, reformulated, and reused.
How do you benchmark an AI reservoir engineer?
That question led to RIGOR — Reservoir Input Generation Output Review.
RIGOR is designed to evaluate agentic reservoir-simulation systems using executable outcomes rather than relying primarily on another language model to judge whether an answer “looks correct.”
The current benchmark contains 135 tasks developed from an initial survey of roughly 700 open OPM-compatible reservoir models. The tasks cover three practical categories: authoring, editing, and repair.
RIGOR separates the benchmark into public, internal-development, and private slices. Numerical behavior receives the greatest scoring weight, with submissions independently rerun against the simulator and compared with reference behavior using task-specific tolerances.
Why go to all that trouble?
Because scientific AI needs tests that it can fail.
A compelling demonstration is useful. A benchmark that exposes where the system succeeds and where it breaks is much more valuable.
What CLARISSA does not claim
There is still plenty of work ahead.
The current system is focused on reservoir engineering and expects geological descriptions or a completed geomodel as input. The reported demonstrations are on a single simulator environment and at reference-problem or single-asset scale. PetroScript currently targets Eclipse-compatible syntax, and fleet-scale deployment has not yet been reported.
Those limitations are important.
The objective is not to pretend an AI system has replaced reservoir engineering judgment.
It has not.
The objective is to remove a different bottleneck: the translation of engineering intent into specialized simulation machinery.
The real opportunity
Reservoir simulators have become extraordinarily capable. Computing has become cheaper. Open-source simulators have matured. Yet simulator use remains concentrated among specialists.
Our thesis is that another constraint has remained largely intact: turning an engineer's question into a model that the simulator can actually consume.
CLARISSA is an attempt to address that layer directly.
The interface may be conversational.
The verification should not be.
Bring us a difficult reservoir-simulation case.