ENGINEERING
Building OmniRecon: governed AI for reconciliation
The deterministic system settles clear cases. In the implemented prototype, a model helps triage only the residue, while independent policy and later verification control what can happen next.
2026 · 16 min read
OmniRecon is a reconciliation system for organisations whose systems disagree about the same event. One system says a customer used a service. Another says what they were billed. The useful question is not only whether the numbers differ. It is what should happen next, and who is allowed to make it happen.
This is an account of the implemented core, the reasoning behind it, and the work deliberately still ahead. It is not a claim of production deployment or measured customer outcomes. The repository uses synthetic data and fixture-backed connectors.
The problem
Operational systems rarely agree perfectly. Usage metering, billing, rate cards, orders, fulfilment and settlement systems all describe parts of the same reality from different viewpoints. Small disagreements accumulate into two different risks:
- Financial impact: delivered value that was not billed.
- Exposure impact: a customer billed for value that was not delivered.
Those risks are asymmetric. A queue ordered only by dollar value can treat a possible overcharge as less urgent than a larger missed charge. OmniRecon keeps both measures visible so a reviewer can make the right decision for the case.
The design question was therefore straightforward: how can software resolve the clear cases quickly without giving an AI system unchecked authority over the ambiguous ones?
The core idea
The system makes deterministic software the engine and AI the exception handler. Rules and tolerances settle straightforward correspondences. Only the residue becomes a case for investigation and review.
For a non-technical reader, the important point is simple: a model does not decide whether it may charge, credit or alter a record. In the implemented prototype, it helps triage unresolved cases. Investigation and corrective proposals are future work. A separate policy and, when needed, a person decide whether anything happens.
Start with boundaries, not a prompt
The project began with constraints rather than an agent prompt. The core is domain-neutral: it understands facts, correspondences, cases, evidence, policy and actions. A domain pack supplies the specific fields, matching rules, tolerances, policy defaults and synthetic fixtures.
The first pack is SaaS revenue reconciliation. Usage metering, billing and a rate card are compared by account, period and usage type. A deliberately thin order-fulfilment pack tests whether the core boundary survives a different problem rather than merely claiming it is generic. The same shape also fits supplier invoices against goods received, or payments against settlement records.
This arrangement makes a useful distinction. Domain knowledge belongs in an explicit, inspectable pack. The system that decides what is safe to do remains separate from both the pack and any model.
Preserve evidence before drawing conclusions
Each imported assertion is an immutable, provenance-stamped fact. If a source record changes, a new version is recorded and the prior version remains available for audit. The matcher uses the current version, while a case can pin the exact facts and configuration it used as evidence.
That solves a practical investigation problem: a reviewer should be able to answer "what did the system know at that time?" without reconstructing it from mutable source records.
The reviewer UI is consequently designed as an evidence-first ledger rather than a generic AI control room. The case queue and detail view show the reconciliation ledger, pinned evidence, action payload, revision, audit events and linked cases before any generated explanation.
Put ordinary arithmetic ahead of AI
The matching engine normalises records into a canonical correspondence, applies a pack-defined tolerance and settles cases that are safely within it. It opens or reuses one active case when a mismatch remains.
Routine arithmetic, identity checks and tolerance handling are therefore deterministic, testable and cheap. The current agent workflow performs triage. The broader investigation, adjudication and execution workflow remains a design direction, not a claim of completed autonomous operation.
| Component | Responsibility | Model involved |
|---|---|---|
| Connectors | Narrow typed reads and separately gated writes | No |
| Fact store | Immutable, provenance-stamped ingestion | No |
| Matcher | Correspondence rules and tolerances | No |
| Case and event log | Replayable audit record | No |
| Agent workflow | Current triage; future investigation and proposals | Yes |
| Policy engine | Decide auto, approval or deny | No |
| Executor | Apply policy-checked writes | No |
Most of the system is deterministic. That is the point.
Treat model output as untrusted input
Source records can contain text written by an external party. That text can be wrong, misleading or adversarial. Prompting a model to ignore it is not a sufficient defence.
OmniRecon separates trusted instructions, validated domain declarations and untrusted source data. Facts such as amounts, dates, quantities and identifiers stay in the deterministic matcher wherever possible. Text that must reach a model is fenced and labelled as data, constrained by field-specific limits and handled through scoped tools.
A proposed action is not an authorised action. Every proposal is evaluated through a pure policy function outside the agent framework and outside the connectors. The outcome is one of three states:
- Automatic for narrowly bounded actions.
- Human approval required for corrective documents, amendments and fuzzy merges.
- Denied for actions that must never run autonomously.
The policy is fail-closed. Known low-risk actions can proceed, corrective proposals require approval and unknown actions are denied. Reviewer approval is bound to the displayed action revision and payload hash, avoiding approval of a quietly changed request.
Design closure as verification, not optimism
An executed correction is not automatically considered resolved. It remains awaiting verification until a later ingestion pass finds the expected correction and reconciles the original correspondence. If that correction is absent or ineffective, the system records the failure and opens a linked follow-up case.
This prevents a common automation failure mode: treating "we attempted a write" as proof that the business state is now correct. Every correction becomes another claim to reconcile rather than an assumption that disappears after a button is pressed.
Build the reviewer surface around evidence
The reviewer queue ranks unresolved cases, then the case detail view anchors decisions in the evidence ledger and exact proposed payload. The reviewer can approve, revise or reject the displayed revision. A changed proposal is not silently resubmitted.
Make generated interfaces auditable
Domain packs can produce a compact, read-only UI artifact. The default artifact is deterministic and generated offline. An optional model-assisted mode generates JSON UI constrained to an allowlisted component catalog.
The artifact records model selection, prompt version and SHA-256 input and output hashes. The application falls back to the deterministic template if no generated artifact is available. This allows useful generation without accepting arbitrary model-generated code into the product surface.
Build sequence
The implementation followed the risk order of the problem rather than the visual order of a product demo:
- Architecture and safety: chose a single durable store, an independent policy boundary, immutable evidence and domain packs before adding model behaviour.
- Reconciliation core: added facts, currentness, deterministic matching, case lifecycle, action revisions and correction verification with synthetic fixtures.
- Review and operations: built the reviewer queue and evidence ledger, then assurance, connector, cost and automation views so decisions can be inspected.
- Safe demonstrations: added local authentication, repeatable seeded demos and browser-local onboarding that never presents simulated connection details as a real integration.
- Constrained generated interfaces: added optional pack UI generation with a fixed component catalog, validation, provenance hashes and a deterministic fallback.
- Explainability in the interface: added a six-month automation chart that illustrates a growing workload, automation catching up and an enduring ambiguous queue. It is not a reported production metric.
The sequence matters. It makes it possible to demonstrate an agentic product without making the demo more autonomous than the safety and evidence model can support.
Decisions and trade-offs
| Decision | Why it was chosen | What it gives up |
|---|---|---|
| Deterministic matching before agents | Clear cases should be predictable, cheap and reproducible. | Less apparent AI coverage. The hard residue remains hard. |
| Independent policy boundary | An agent must not authorise itself. | More explicit states and approval design work. |
| Immutable facts and event log | Historical decisions need replayable evidence. | More storage and version-handling complexity. |
| Verify corrections on later ingestion | A write is not proof of a correct outcome. | Resolution is intentionally delayed. |
| Fixture-backed connectors and synthetic data | Safe, reproducible demos and tests. | No claim of live production integrations. |
| Evidence-first review UI | Humans can judge data and exact proposed effect. | Less focus on conversational AI theatrics. |
| Constrained generated UI artifact | Keep generation inspectable and bounded. | Less expressive than free-form generated frontends. |
What exists today
The repository includes immutable fact storage and a current-version projection; deterministic SaaS revenue matching and settlement handling; active-case creation, state transitions, event records, action revisions and fixture-backed correction verification.
It also includes a compiled policy foundation, reviewer API, authenticated review UI and operational screens for automation, assurance, integrations and cost. The system has a triage workflow, synthetic fixtures, a local demo flow and an in-repository test suite.
The reviewer UI includes a browser-local onboarding and recruiter walkthrough. It simulates source and destination setup, MCP URLs, authentication values, notifications and a sample workflow. It creates no account, connects no system and sends no email. The distinction is explicit so a demo never implies that a production integration occurred.
What is deliberately not claimed
- No production customer data or live customer connector is used here.
- No real external billing write-back is claimed.
- No benchmark accuracy, safety rate, cost saving or latency result is published until a reproducible evaluation exists.
- Multi-tenant isolation enforcement, real connector admission and the full multi-stage agent workflow remain future work.
The intended evaluation is explicit: deterministic settlement, agent residue, policy outcomes, false auto-execution, injection resistance, accuracy by difficulty band, cost per case and traceable workflow behaviour. Until those runs exist, value claims remain potential rather than observed.
Where it goes next
The next meaningful step is not making the agent more autonomous everywhere. It is earning autonomy narrowly. Repeated, well-evidenced reviewer decisions can be proposed as deterministic configuration changes, evaluated against labelled fixtures and historical cases, and approved by a person before promotion.
That path keeps the original principle intact: the more a pattern becomes known, the less it should need a model.
