Skip to content
pradneshfernandezPublic

About

Developer-facing benchmark for structured geospatial LLM output. Implements VOYGR's Phase 2 evaluation (Section 7, Q1 2026 report) measuring how Claude, OpenAI, and Gemini handle geocoding, GeoJSON validity, and batch POI validation, and quantifying VOYGR's verification layer in dollars per caught error.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DevBench

A developer-facing benchmark for structured geospatial LLM output.

DevBench implements VOYGR's announced Phase 2 evaluation (Section 7, Q1 2026 report). It measures how accurately Claude, OpenAI, and Gemini produce structured geospatial data — geocoding accuracy, GeoJSON validity, and batch POI validation — and quantifies what VOYGR's verification layer is worth in dollars per caught error.

It is open, reproducible, and runs end-to-end on a laptop in a few commands.


Why this exists

VOYGR's own report graded models using Google Search + Maps grounding, then flagged in Section 8 that this introduces evaluator bias — Gemini is effectively graded by its own tool stack. DevBench removes that bias by verifying against an open, neutral source:

PostGIS + OpenStreetMap, not Google Maps. This divergence is deliberate and documented — see docs/architecture.md.

For a full technical reference covering the data model, API, scoring methodology, and current state of the benchmark, see docs/documentation.md.


The pipeline

Prompt set → [EXTRACT] → [VERIFY] → [SCORE] → DynamoDB → Dashboard
Category What it tests Verifier Needs VOYGR key
A Geocoding accuracy PostGIS ST_DWithin @ 50 m vs OSM No
B GeoJSON validity PostGIS ST_IsValid + schema No
C Batch POI validation VOYGR /v1/business-status (Redis-cached) Yes

Scoring uses VOYGR's published R0–R4 rubric (out of 100) with a 2× hallucination penalty: a fatally flawed entity counts twice in the denominator. See docs/architecture.md.


Quick start

Requirements: Docker + Docker Compose, Python 3.13, osm2pgsql, and at least one LLM API key.

git clone <repo> && cd devbench
cp .env.example .env          # fill in at least one of ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt

make setup                    # start PostGIS + DynamoDB Local + Redis, load OSM data (~30 MB download)

Run the baseline (no VOYGR key needed)

make run-baseline             # Categories A + B, verified locally against OSM

You need at least one LLM provider key — not all three. Any provider whose key is absent is logged and skipped, so a Gemini-only setup works. Restrict the set explicitly with the PROVIDERS env var:

PROVIDERS=gemini make run-baseline          # only Gemini
PROVIDERS=gemini,openai make run-baseline   # two of three

Run the full benchmark (Category C needs a VOYGR key)

# add VOYGR_API_KEY to .env, then:
make run-integrated           # A + B + C. Without the key, C is skipped gracefully.

View results

make serve                    # FastAPI dashboard on http://localhost:8080

Open http://localhost:8080/docs for the interactive OpenAPI explorer.


API endpoints

Endpoint Returns
GET /health Liveness probe
GET /results/summary Latest run, per-provider × per-category score breakdown
GET /results/longitudinal?provider=&category= Score / latency / cost over time
GET /results/voygr-impact The headline: fatal-flaw catch rate and cost-per-caught-error (Category C)

/results/voygr-impact is the proof-of-work metric: it compares the raw fatal-flaw count in each model's output against how many VOYGR caught, and divides total cost by errors caught to answer "what does one prevented bad place cost?"


Running everything in Docker

The devbench service builds the app image and joins the same network as the data services:

docker compose up -d --build           # dashboard on :8080
docker compose run --rm devbench python scripts/run_baseline.py

Make targets

Target Action
make setup Start services + load OSM data
make run-baseline Categories A + B (no VOYGR key)
make run-integrated All categories (C needs VOYGR_API_KEY)
make serve Start the results dashboard on :8080
make test Run the test suite
make logs / make down / make clean Tail logs / stop / tear down + wipe volumes

Project layout

prompts/          Benchmark prompt dataset — v1.json is source of truth
spatial/          PostGIS + OSM verifier (Categories A + B)
runner/           Extract → Verify → Score pipeline
  providers/      One async client per LLM provider
  pipeline.py     End-to-end orchestration (used by the run scripts)
scoring/          R0–R4 rubric + 2× hallucination penalty
storage/          DynamoDB write/read path
instrumentation/  CloudWatch metrics + cost tracking (stdout fallback)
api/              FastAPI results dashboard
infra/            CloudWatch dashboard definition
scripts/          run_baseline.py, run_integrated.py
tests/            Unit + integration tests (offline fixtures)
docs/             architecture.md, errors.md, handoff.md

Testing

One command runs everything:

make test            # == pytest tests/

The suite self-skips by tier depending on what's running, so it's always safe to run as a single command:

  • No services up — unit tests pass; integration tests skip.
  • docker compose up -d — storage + GeoJSON-validity tests run; 4 OSM-dependent coordinate tests skip.
  • make setup (services + OSM loaded) — all 120 run.

Unit/logic tests mock every external dependency (LLM APIs, VOYGR, Redis, DynamoDB, CloudWatch), so they need no API keys. CI (.github/workflows/tests.yml) runs the single pytest tests/ with PostGIS, Redis, and DynamoDB Local service containers up.


Configuration

All configuration is via .env (copied from .env.example, which documents every variable). Notable knobs:

  • PROVIDERS — comma-separated subset to run (e.g. gemini). Unset = every provider whose key is present. Providers with no key are skipped gracefully.
  • VOYGR_API_KEY — blank disables Category C gracefully.
  • GEOCODING_THRESHOLD_METERS — Category A pass radius (default 50 m).
  • AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY — blank = metrics log to stdout; set = ship to CloudWatch. (Local DynamoDB needs no real AWS credentials.)

Design rules (do not break)

  1. PostGIS + OSM is the verifier — never Google Maps (removes evaluator bias).
  2. The 2× hallucination penalty matches VOYGR's published methodology exactly.
  3. Category C degrades gracefully without VOYGR_API_KEY.
  4. All LLM calls are async (asyncio + httpx).
  5. Data crossing a module boundary is always a Pydantic model.
  6. Category C checks Redis before calling VOYGR (sha256(name|address), 7-day TTL).

See CLAUDE.md and docs/architecture.md for the full rationale.


Status

First live end-to-end run completed 2026-06-20 with Gemini only (Categories A + B):

4 results · avg score 30.0/100 · total cost $0.0004

Low scores are expected — prompts run without search grounding, so Gemini relies on training-data coordinates which PostGIS checks at 50 m precision against real OSM data. That gap is the hallucination rate the benchmark is designed to surface.

Full test suite: 120/120 passing. See docs/errors.md for all issues encountered and fixed during development and the live run.

About

Developer-facing benchmark for structured geospatial LLM output. Implements VOYGR's Phase 2 evaluation (Section 7, Q1 2026 report) measuring how Claude, OpenAI, and Gemini handle geocoding, GeoJSON validity, and batch POI validation, and quantifying VOYGR's verification layer in dollars per caught error.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages