Evidence-backed investigation system for enterprise transaction exceptions.
Enterprise transaction exceptions—such as invoice rate variances, unapproved service fees, and billing disputes—are notoriously time-consuming and expensive to investigate. In typical enterprises, accounts payable and finance teams spend hours manually cross-referencing records because supporting evidence is fragmented across disconnected systems:
- Master Services Agreements (MSAs) defining baseline terms, fee structures, and discount tiers
- Contract amendments modifying specific clauses, rates, or validity periods
- Statements of Work (SOWs) detailing deliverable milestones and scope caps
- Operational approvals granting variance authorizations via emails, tickets, or formal sign-offs
- Invoices & billing records containing disputed line items, quantities, and dates
- Renewal notices & ancillary records establishing effective dates and customer entity bindings
When an exception occurs, a human reviewer must reconstruct the multi-hop relationship path that links the invoice line item back to governing contractual authority. Disconnected relational queries often pull extraneous terms across multiple contracts for the same customer, while pure LLM solutions hallucinate terms, miss amendment precedence, or fabricate authority.
ExceptionLineage investigates enterprise transaction exceptions by tracing relationship lineage across contracts, amendments, SOWs, approvals, and invoices to produce evidence-backed, auditable determinations.
Rather than relying on unconstrained LLM text generation or rigid static SQL scripts, ExceptionLineage implements a strict architectural boundary:
"AI investigates. Deterministic code verifies."
The agent dynamically plans and navigates graph traversal paths to gather necessary documentary evidence. Once evidence is gathered, a deterministic validation engine verifies dates, rate thresholds, product scopes, and approval requirements—ensuring contractual authority is always code-governed, fail-closed, and audit-grade.
Enterprise exception investigation does not follow a uniform, linear checklist:
- Some invoices match base contract fee schedules immediately, rendering amendment or SOW lookups wasteful.
- Other invoices involve unanchored entities (e.g., missing contracts), where further queries for approvals or deliverables are pointless.
- Complex exceptions branch: an invoice might require checking multiple amendments, discovering an executive approval waiver, or pivoting to an SOW deliverable schedule.
The agent's specific role is adaptive investigation:
- It evaluates intermediate findings after each step.
- It dynamically chooses which available tool to execute next.
- It recognizes when sufficient evidence exists to stop early, avoiding unnecessary queries.
- It pivots investigation paths when initial branches hit dead ends.
Crucially, the agent never decides the business outcome. The agent determines what evidence to retrieve next; deterministic code determines whether the transaction is legally and contractually authorized.
An investigation follows a controlled, evidence-driven loop:
Invoice Exception Flagged
│
▼
Target Contract Identified (Governing MSA / Terms)
│
▼
Amendments Traversed (Active rate revisions & clauses)
│
▼
Statements of Work Inspected (Deliverables, caps & scopes)
│
▼
Approvals & Sign-offs Retrieved (Formal variance authorizations)
│
▼
Related Evidence Gathered (Verbatim clause excerpts & validity windows)
│
▼
Deterministic Validation (Rule checks: PASS | FAIL | UNKNOWN)
│
▼
Evidence-Backed Finding (VERIFIED | NOT_VERIFIED | INSUFFICIENT_EVIDENCE | NEEDS_REVIEW)
The system coordinates client requests, agent planning, graph retrieval, rule evaluation, and immutable audit logs:
USER / CLIENT
│
▼
INPUT (POST /api/investigations)
│
▼
INVESTIGATION SERVICE (Lifecycle & State Machine: QUEUED → INVESTIGATING → VALIDATING)
│
▼
INVESTIGATION AGENT (Bounded reasoning loop; max 10 steps)
│
├──> TOOL REGISTRY (Typed, schema-validated queries)
│ │
│ ▼
│ RELATIONSHIP-AWARE RETRIEVAL (Neo4j Property Graph / Multi-hop traversal)
│ │
│ ▼
│ INVESTIGATION CONTEXT (Assembled evidence, clauses, entities)
│
▼
DETERMINISTIC VALIDATION ENGINE (Pure rule evaluation: PASS | FAIL | UNKNOWN)
│
▼
EVIDENCE & STATE MACHINE TRANSITION (Authoritative terminal status emitted)
│
▼
IMMUTABLE AUDIT TRACE & FINDING (30-event chronological trace + cited evidence)
ExceptionLineage enforces a non-negotiable architectural separation between probabilistic reasoning and deterministic authority:
| Responsibility Domain | System Component | Specific Tasks Handled |
|---|---|---|
| AI / Agent | InvestigationAgentAgentModel |
• Dynamic query planning • Ambiguity resolution & clause interpretation • Next-step selection based on intermediate evidence • Branching path navigation • Deciding when sufficient evidence is gathered |
| Deterministic Code | ValidationEngineInvestigationStateMachine |
• Legal effective date validation • Product scope & SKU matching • Approval presence & signer authority checks • Financial calculations (exact Decimal math)• Amendment conflict & precedence resolution • State machine transitions & final verification outcome |
The system never allows an LLM to state: "I verified this invoice." Verification is an objective calculation performed by code against gathered evidence.
Enterprise contracts form a dense, directed property graph:
- A
Customeris bound by multipleContracts. - Each
Contractis amended by specificAmendmentsand executed viaSOWs. - Specific
Approvalsauthorize specificExceptions. - Evidence nodes link to discrete clauses and validity windows.
In our controlled retrieval experiments, explicit relationship traversal proved load-bearing:
- Contract Boundary Isolation: In flat relational queries, querying amendments by customer ID pulled amendments belonging to unrelated contracts of the same customer, contaminating the validation context and creating false rate conflicts.
- Directional Traversal: Traversing
(Contract)-[:AMENDED_BY]->(Amendment)ensures that only amendments explicitly linked to that specific contract are considered. - Multi-Hop Lineage: Graph queries maintain end-to-end evidence chains from the invoice line item back to the governing contract with 100% provenance completeness.
Limitation Note: This experiment demonstrates the value of explicit relationship-aware retrieval for the tested workload. It does not establish that Neo4j is universally superior to a well-designed relational schema.
Every investigation automatically produces a machine-readable, chronological evidence trace accessible via GET /api/investigations/{id}/trace:
INPUT
↓
AGENT_DECISION (tool planning rationale & step context)
↓
TOOL_CALL (controlled invocation with schema-validated arguments)
↓
GRAPH_RETRIEVAL (Cypher traversal, relationship path, cited evidence IDs)
↓
VALIDATION (deterministic check execution: PASS | FAIL | UNKNOWN)
↓
OUTCOME (authoritative final determination & state transition)
- Complete Secret Redaction: All API keys, authorization tokens, and credentials in tool arguments or metadata are automatically sanitized to
[REDACTED]. - Immutable Audit Log: Every state transition and agent action is recorded as an immutable event with UTC timestamps.
- Verifiable Citations: Every finding explicitly cites the supporting evidence IDs (
EV-001,EV-002, etc.) containing verbatim source excerpts, document locators, and validity windows.
ExceptionLineage includes a fully reproducible quantitative evaluation harness (evaluation/):
| Evaluation Area | System / Method Tested | Baseline / Alternative | Measured Result |
|---|---|---|---|
| Claim A: Adaptive Investigation Value | Adaptive Agent (state-dependent tool selection) | Heuristic Baseline (fixed query sequence) | • 100.0% accuracy vs 80.0% • 25.0% tool call reduction (24 vs 32 calls) • 0 unnecessary tool calls (avoided 7 wasted queries) • 3 dynamic early stops • 60.0% branching path recovery rate (3/5 cases in |
| Claim B: Relationship-Aware Retrieval | Graph Traversal (Neo4jLineageRepository) |
Flat Relational Mock (FlatRetrievalAdapter) |
• 100.0% validation accuracy vs 87.5% (flat lookups caused false amendment conflicts) • 0 irrelevant records retrieved vs 37 extraneous records • 100.0% multi-hop provenance vs 20.0% • 1.0 query/case vs 8.25 operations/case |
| Claim C: End-to-End Evidence Chain | Chronological Audit Trace Generator | Schema & Redaction Verifier | • 30 total chronological events verified end-to-end (1 input, 7 agent decisions, 7 tool calls, 6 graph retrievals, 8 validation checks, 1 outcome) • 8 cited evidence items across 8 validation checks • secrets_redacted = true• authority_boundary_preserved = true
|
| Live LLM Evaluation Status |
LLMDecisionModel (live OpenAI API) |
Provider Quota / Availability | Provider quota was exhausted (HTTP 429). In accordance with evaluation standards, live LLM accuracy was marked UNMEASURABLE (null) rather than 0% to prevent conflating provider availability with model reasoning ability. The system safely failed closed with 0 hallucinations. |
ExceptionLineage treats failure and uncertainty as first-class states, never as unhandled exceptions:
- Missing Evidence (
INSUFFICIENT_EVIDENCE): When a required document (e.g., executive approval for a rate variance) is absent from the graph, the validation engine marks the check asUNKNOWN. Crucially,UNKNOWNevidence does not become a confident conclusion. The investigation terminates withINSUFFICIENT_EVIDENCE. - Contractual Non-Compliance (
NOT_VERIFIED): When an invoice charge violates active terms (e.g., expired amendment date or unauthorized rate hike), the check fails deterministically and recordsNOT_VERIFIED. - Conflicting Authority (
NEEDS_REVIEW): When multiple concurrent amendments specify competing terms without clear precedence, the engine escalates the transaction for human review viaNEEDS_REVIEW. - Infrastructure / System Failure (
FAILED): If an invalid identifier is supplied or graph queries fail, the state machine transitions toFAILEDwith an explicit notice: "No determination was made." Infrastructure failure never converts into a business outcome.
All demonstrations, seed records, contracts, invoices, and evaluation cases in this repository are simulated, synthetic data created specifically for testing and benchmark evaluation.
No real-world enterprise production data, proprietary customer records, or live ERP systems are used in this repository.
To maintain strict technical honesty, reviewers should note the following current limitations:
- Synthetic Evaluation Dataset: Benchmarks evaluate controlled synthetic enterprise scenarios (
benchmark-v1,adaptive-v1). The system has not yet been validated on multi-million record production enterprise ERPs. - In-Memory Persistence: Current investigation records and events persist in application memory during runtime. Future production milestones will integrate persistent relational storage (PostgreSQL).
- External LLM Provider Availability: Live LLM evaluation depends on active provider API quotas; during testing, provider quota exhaustion prevented measuring live LLM token efficiency, though deterministic mock and adaptive models execute 100% offline.
- No Production Accuracy Claim: We claim verified accuracy on the tested controlled scenarios, not universal or production enterprise accuracy.
- Event: Open Agent Hackathon 2026
- Selected Track: Track 03: The Agent That Can Explain Why (also strongly aligned with Track 02: Autonomous Agent)
- Why ExceptionLineage Fits This Track: The "Agent That Can Explain Why" track specifically challenges builders to demonstrate a compelling use of relationships and evidence to produce explainable conclusions. ExceptionLineage is purpose-built to answer why enterprise transaction exceptions occur by traversing relational lineage across contracts, amendments, SOWs, approvals, and invoices, producing an auditable chain of cited evidence and deterministic validation checks rather than black-box assertions.
- Repository: https://github.com/Anekenonso/ExceptionLineage
- Video Walkthrough:
[TODO: Link to 3-5 minute demo video] - Sponsor Technology Actually Used:
- Neo4j: Knowledge graph storage, directional relationship traversal, Cypher query execution, multi-hop evidence lineage.
- Sponsor Technologies Not Used:
- Meterless: Not integrated into current persistence tier.
- Zetaris: Not integrated into current data virtualization tier. (We do not claim sponsor integrations that are not load-bearing in the codebase.)
- Node.js 18+
- Python 3.11+
- npm
cd apps/api
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # macOS/Linux
pip install -r requirements.txt
uvicorn app.main:app --port 8000 --reloadThe API will be available at http://localhost:8000. Test health at http://localhost:8000/health.
cd apps/web
npm install
npm run devThe application will be available at http://localhost:3000.
# Run backend test suite from repository root (296 passed, 2 skipped)
.\apps\api\venv\Scripts\python.exe -m pytest apps/api/tests/ -v
# Run frontend production build & TypeScript validation
cd apps/web
npm run buildExceptionLineage is production-hardened for a decoupled cloud architecture:
INTERNET
│
▼
Vercel / Next.js
apps/web
│
HTTPS
│
▼
Render / FastAPI
apps/api
│
Bolt + TLS (neo4j+s://)
│
▼
Neo4j Aura
- Create a free or professional Neo4j AuraDB instance at console.neo4j.io.
- Save your instance connection URI (e.g.
neo4j+s://<dbid>.databases.neo4j.io), username (neo4j), and generated password. - Zero-step bootstrapping: On first connection, ExceptionLineage detects if the database is unpopulated and automatically initializes all unique constraints, indexes, and loads the benchmark seed dataset.
Deploy apps/api using the included render.yaml blueprint or manual setup:
-
Environment: Python 3.11+ (or Docker via
apps/api/Dockerfile) -
Build Command:
pip install -r apps/api/requirements.txt -
Start Command:
PYTHONPATH=apps/api uvicorn app.main:app --host 0.0.0.0 --port $PORT -
Health Check Path:
/health -
Environment Variables:
-
PORT: Automatically set by Render -
CORS_ORIGINS: Comma-separated list including your Vercel URL (e.g.https://exceptionlineage.vercel.app,http://localhost:3000) -
CORS_ORIGIN_REGEX:^https://.*\.vercel\.app$(supports preview branch deployments) -
NEO4J_URI:neo4j+s://<dbid>.databases.neo4j.io -
NEO4J_USERNAME:neo4j -
NEO4J_PASSWORD:<your-neo4j-password> -
NEO4J_DATABASE:neo4j -
AGENT_MODEL:heuristic(default deterministic) orllm -
AGENT_LLM_PROVIDER:geminioropenai(optional) -
AGENT_LLM_API_KEY:<api-key>(optional)
-
Deploy apps/web to Vercel:
- Import repository on vercel.com/new.
- Set Root Directory to
apps/web. - Set Framework Preset to
Next.js. - Configure Environment Variables:
NEXT_PUBLIC_API_URL:https://<your-render-service>.onrender.com(trailing slashes are automatically sanitized).
- Click Deploy. The site will build cleanly and connect securely to the Render API.
ExceptionLineage provides a secure, sandboxed testing path allowing judges and evaluators to test custom invoice exceptions and contractual lineages without altering permanent storage or benchmark records.
- Strict Format: Exclusively
.jsonfiles under 1 MB. - Single Shared Engine: Runs through the identical
InvestigationServiceandValidationEnginepipeline. - Zero Persistence: Evaluated strictly in request-scoped ephemeral memory. Never written to Neo4j, relational databases, disk, or browser storage.
- Click "Review an invoice" in the top navigation.
- Select the "Test Your Own Case" tab.
- Click "Download Template" to get a fully populated starter schema, or drag and drop your own
.jsoncase. - Click "Run Test Case" to view the real-time deterministic validation and end-to-end evidence chain.
# 1. Fetch starter template
curl -s http://localhost:8000/api/investigations/test-case/template > my_case.json
# 2. Upload and evaluate in ephemeral sandbox
curl -s -X POST http://localhost:8000/api/investigations/test-case \
-F "file=@my_case.json;type=application/json" | jq .Run the reproducible quantitative evaluation suite from the repository root:
# Complete Architectural Evaluation (Claims A, B, and C):
python evaluation/runner.py --stage-18-5
# Standard Evaluation Suite (Deterministic Baseline A and Heuristic Baseline B):
python evaluation/runner.py
# Offline Mock LLM Evaluation:
python evaluation/runner.py --mock-llm
# Live LLM Evaluation (requires active API key):
export AGENT_LLM_API_KEY="your-api-key"
python evaluation/runner.py --with-llm
# Specific Baseline Evaluation:
python evaluation/runner.py --baseline deterministic
python evaluation/runner.py --baseline heuristic
python evaluation/runner.py --baseline adaptiveGenerated reports:
evaluation/reports/stage-18-5-latest.json(Architectural claims data)evaluation/reports/stage-18-5-latest.md(Human-readable claims markdown)evaluation/reports/latest.json(Standard baseline report)evaluation/reports/latest.md(Standard baseline markdown)
ExceptionLineage/
├── apps/
│ ├── web/ # Next.js 16 frontend (TypeScript, Tailwind CSS, React 19)
│ └── api/ # FastAPI backend (Python 3.11+, Pydantic v2, Neo4j Driver)
├── data/ # Synthetic seed records, benchmark cases, and ground truth
│ ├── cases/ # Investigation case definitions
│ ├── ground_truth/ # Immutable ground truth determinations
│ └── seed/ # Initial contracts, invoices, amendments, approvals
├── evaluation/ # Quantitative evaluation harness & reporting
│ ├── adapters/ # Baseline adapters (Deterministic, Heuristic, Adaptive, Flat)
│ ├── reports/ # Machine-readable JSON and human-readable Markdown reports
│ ├── dataset.py # Benchmark and branching scenario loader
│ ├── metrics.py # Transparent accuracy, recall, tool, and failure formulas
│ ├── runner.py # CLI evaluation runner
│ └── schemas.py # Pydantic evaluation schemas
├── neo4j/ # Graph schema and Cypher constraints
├── docs/ # Technical architecture, decisions, and evaluation documentation
├── .env.example # Environment variable template
├── .gitignore # Git ignore rules
└── README.md # Public-facing repository documentation
- System Architecture
- Architectural Decisions
- Evaluation Framework
- Controlled Failure Modes
- Demo Walkthrough Script
- Demo Scenarios Matrix
- Data Contracts & Schemas
- Architectural Traps Avoided
- Project State
ExceptionLineage features an editorial, calm, and audit-grade interface designed with classical typography, architectural linework, and restrained visual indicators:
Searchable directory with real-time status filter pills, multi-parameter sorting, and dense ledger view.

Side-by-side contract vs billed discrepancy terms, deterministic check matrix, and interactive governance graph.

- Built for the Open Agent Hackathon 2026
- Graph Database: Neo4j
- Backend Framework: FastAPI & Pydantic
- Frontend Framework: Next.js & React
© 2026 ExceptionLineage. All rights reserved.
