A developer-facing benchmark for structured geospatial LLM output.
DevBench implements VOYGR's announced Phase 2 evaluation (Section 7, Q1 2026 report). It measures how accurately Claude, OpenAI, and Gemini produce structured geospatial data — geocoding accuracy, GeoJSON validity, and batch POI validation — and quantifies what VOYGR's verification layer is worth in dollars per caught error.
It is open, reproducible, and runs end-to-end on a laptop in a few commands.
VOYGR's own report graded models using Google Search + Maps grounding, then flagged in Section 8 that this introduces evaluator bias — Gemini is effectively graded by its own tool stack. DevBench removes that bias by verifying against an open, neutral source:
PostGIS + OpenStreetMap, not Google Maps. This divergence is deliberate and documented — see
docs/architecture.md.
For a full technical reference covering the data model, API, scoring methodology, and
current state of the benchmark, see docs/documentation.md.
Prompt set → [EXTRACT] → [VERIFY] → [SCORE] → DynamoDB → Dashboard
| Category | What it tests | Verifier | Needs VOYGR key |
|---|---|---|---|
| A | Geocoding accuracy | PostGIS ST_DWithin @ 50 m vs OSM |
No |
| B | GeoJSON validity | PostGIS ST_IsValid + schema |
No |
| C | Batch POI validation | VOYGR /v1/business-status (Redis-cached) |
Yes |
Scoring uses VOYGR's published R0–R4 rubric (out of 100) with a 2× hallucination
penalty: a fatally flawed entity counts twice in the denominator. See
docs/architecture.md.
Requirements: Docker + Docker Compose, Python 3.13, osm2pgsql, and at least one LLM API key.
git clone <repo> && cd devbench
cp .env.example .env # fill in at least one of ANTHROPIC_API_KEY / OPENAI_API_KEY / GEMINI_API_KEY
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
make setup # start PostGIS + DynamoDB Local + Redis, load OSM data (~30 MB download)make run-baseline # Categories A + B, verified locally against OSMYou need at least one LLM provider key — not all three. Any provider whose key is
absent is logged and skipped, so a Gemini-only setup works. Restrict the set explicitly
with the PROVIDERS env var:
PROVIDERS=gemini make run-baseline # only Gemini
PROVIDERS=gemini,openai make run-baseline # two of three# add VOYGR_API_KEY to .env, then:
make run-integrated # A + B + C. Without the key, C is skipped gracefully.make serve # FastAPI dashboard on http://localhost:8080Open http://localhost:8080/docs for the interactive OpenAPI explorer.
| Endpoint | Returns |
|---|---|
GET /health |
Liveness probe |
GET /results/summary |
Latest run, per-provider × per-category score breakdown |
GET /results/longitudinal?provider=&category= |
Score / latency / cost over time |
GET /results/voygr-impact |
The headline: fatal-flaw catch rate and cost-per-caught-error (Category C) |
/results/voygr-impact is the proof-of-work metric: it compares the raw fatal-flaw count
in each model's output against how many VOYGR caught, and divides total cost by errors
caught to answer "what does one prevented bad place cost?"
The devbench service builds the app image and joins the same network as the data
services:
docker compose up -d --build # dashboard on :8080
docker compose run --rm devbench python scripts/run_baseline.py| Target | Action |
|---|---|
make setup |
Start services + load OSM data |
make run-baseline |
Categories A + B (no VOYGR key) |
make run-integrated |
All categories (C needs VOYGR_API_KEY) |
make serve |
Start the results dashboard on :8080 |
make test |
Run the test suite |
make logs / make down / make clean |
Tail logs / stop / tear down + wipe volumes |
prompts/ Benchmark prompt dataset — v1.json is source of truth
spatial/ PostGIS + OSM verifier (Categories A + B)
runner/ Extract → Verify → Score pipeline
providers/ One async client per LLM provider
pipeline.py End-to-end orchestration (used by the run scripts)
scoring/ R0–R4 rubric + 2× hallucination penalty
storage/ DynamoDB write/read path
instrumentation/ CloudWatch metrics + cost tracking (stdout fallback)
api/ FastAPI results dashboard
infra/ CloudWatch dashboard definition
scripts/ run_baseline.py, run_integrated.py
tests/ Unit + integration tests (offline fixtures)
docs/ architecture.md, errors.md, handoff.md
One command runs everything:
make test # == pytest tests/The suite self-skips by tier depending on what's running, so it's always safe to run as a single command:
- No services up — unit tests pass; integration tests skip.
docker compose up -d— storage + GeoJSON-validity tests run; 4 OSM-dependent coordinate tests skip.make setup(services + OSM loaded) — all 120 run.
Unit/logic tests mock every external dependency (LLM APIs, VOYGR, Redis, DynamoDB,
CloudWatch), so they need no API keys. CI (.github/workflows/tests.yml) runs the
single pytest tests/ with PostGIS, Redis, and DynamoDB Local service containers up.
All configuration is via .env (copied from .env.example, which documents every
variable). Notable knobs:
PROVIDERS— comma-separated subset to run (e.g.gemini). Unset = every provider whose key is present. Providers with no key are skipped gracefully.VOYGR_API_KEY— blank disables Category C gracefully.GEOCODING_THRESHOLD_METERS— Category A pass radius (default 50 m).AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY— blank = metrics log to stdout; set = ship to CloudWatch. (Local DynamoDB needs no real AWS credentials.)
- PostGIS + OSM is the verifier — never Google Maps (removes evaluator bias).
- The 2× hallucination penalty matches VOYGR's published methodology exactly.
- Category C degrades gracefully without
VOYGR_API_KEY. - All LLM calls are async (
asyncio+httpx). - Data crossing a module boundary is always a Pydantic model.
- Category C checks Redis before calling VOYGR (
sha256(name|address), 7-day TTL).
See CLAUDE.md and docs/architecture.md for the
full rationale.
First live end-to-end run completed 2026-06-20 with Gemini only (Categories A + B):
4 results · avg score 30.0/100 · total cost $0.0004
Low scores are expected — prompts run without search grounding, so Gemini relies on training-data coordinates which PostGIS checks at 50 m precision against real OSM data. That gap is the hallucination rate the benchmark is designed to surface.
Full test suite: 120/120 passing. See docs/errors.md for all issues encountered
and fixed during development and the live run.