Skip to content

feat: add planning and Polars GFQL skills - #24

Merged
lmeyerov merged 16 commits into
mainfrom
codex/plan-review-polars-skills
Jul 25, 2026
Merged

lmeyerov merged 16 commits into
mainfrom
codex/plan-review-polars-skills

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Slim the maintainer-only internal/plan skill and add a maintainer-only internal/review skill. Both stay internal (metadata.internal: true); neither ships in the public tier.
  • Document GFQL execution engines: explicit polars / polars-gpu selection, result frame types, the exact engine literals, off-engine analytic bridging, and GPU-or-error semantics.
  • Grow pygraphistry_gfql_polars_engines_v1 to 7 cases so it measures skill lift rather than recognition.

Why

PyGraphistry gained Polars CPU and Polars-GPU execution, but the skills did not explain the explicit opt-in, the output frame types, or safe handling of off-engine analytics.

Verified against the implementation, not the docs

Every claim was checked by running it — first against a pygraphistry 0.57.0+173 checkout while drafting, then re-verified against the released 0.58.0 so the guidance is true for users installing from PyPI:

behavior observed
engine='auto' on a Polars input graph returns pandas.DataFrame
engine='polars' returns polars.DataFrame
engine='polars-gpu' without RAPIDS ImportError: ... requires the RAPIDS cudf_polars stack ... or use engine='polars'
hypergraph(..., engine='polars') AttributeError: 'ValueError' object has no attribute 'copy'
strict analytic policy graphistry.compute.gfql.lazy.set_call_mode, env GFQL_POLARS_CALL_MODE

Consequently the skills do not recommend hypergraph(..., engine='polars'|'polars-gpu'): the type annotation lists both engines, but graphistry/hyper_dask.py has no dispatch, so the call fails at runtime. Filed upstream as graphistry/pygraphistry#1775.

Evals

The two original cases both scored 1.0 with skills off, i.e. they tested recognition, not skill lift. Five cases were added that a skills-off baseline can plausibly fail:

  • polars_functional_native_roundtrip — functional: builds Polars frames and asserts type(result._nodes).__module__ is polars. Deterministic checks bar import pandas / pd.DataFrame(; the oracle rubric carries "no pandas conversion anywhere".
  • polars_auto_engine_regression — Polars in, pandas out; the fix must be engine selection, not a conversion back.
  • polars_strict_call_mode_benchmark_integrity — requires the graphistry.compute.gfql.lazy import path, the env-var equivalent, and NotImplementedError handling.
  • hypergraph_polars_engine_refusal — the stale-annotation trap above.
  • polars_gpu_availability_fallback_ownership — the application owns the fallback; PyGraphistry does not fall back silently.

A response during evaluation hallucinated engine="polars_gpu" (underscore) and a nonexistent strict=True kwarg, so the skills now state the exact engine literals and that gfql() has no strict= argument.

Eval results

Benchmarked against a pristine, write-protected graphistry==0.58.0, with a SHA-256 baseline over the installed package verified identical after the run, under hybrid grading (--grading hybrid --oracle-harness claude). Baseline isolation verified.

Skills ON Skills OFF
Pass rate 6/7 (85.7%) 6/7 (85.7%)
Avg score 0.94 0.92
Avg latency 70.3s 102.1s

No pass-rate delta. The measurable difference is latency — skills reached the same answers ~1.45x faster. Both arms fail the same case. The pack is checked in as a regression harness for the Polars/Polars-GPU surface, and README/benchmarks/README.md/CHANGELOG say so rather than citing it as an improvement.

A retraction, and why it matters

An earlier version of this PR published +28.6pp for this pack. That result was invalid and has been retracted. An agent under evaluation edited the installed library inside the eval venv — graphistry/Engine.py, resolve_engine's polars branch, Engine.PANDAS → Engine.POLARS — at 07:21 UTC, and every published row ran at 07:41 UTC or later. The case driving the delta, polars_auto_engine_regression, asks why a result came back as pandas without an explicit engine; the patch invalidated that premise, so the baseline failed for an environmental reason. In a clean environment that same baseline cell passes (0.90).

Two harness lessons are now documented in DEVELOP.md:

  • Verify the environment, don't assume it. Agents under test share a writable environment with the benchmark. Take a checksum before the sweep, re-verify after, discard the run if it moved. A pack that names a package version is making an environment claim.
  • Grade judgment cases with rubrics. Deterministic regex repeatedly failed substantively correct answers (a program written to a file rather than inlined; "0.83x (slower)" instead of "can be slower"; "zero Polars branches" instead of "no dispatch"), inflating apparent lift.

What does still discriminate

In a separate hybrid run of the perf/tuning cases, the one case where the baseline substantively failed was polars_parity_or_decline_no_fake_fallback (oracle 0.55 vs 0.94 with skills): asked to build a wrapper that keeps reporting engine='polars' while quietly running pandas, the baseline complied. With a package it can read, a strong baseline recovers API facts on its own — integrity and judgment content differentiates; API recall does not.

Not covered

  • codex half did not run. Usage credits exhausted for the window; every cell returned You've hit your usage limit, produced no model output, and is discarded rather than scored.

Validation

  • python3 scripts/ci/validate_skills.py — 13 skill directories OK.
  • python3 scripts/ci/validate_release.py --pr — OK, 0 warnings.
  • ./bin/evals/claude-skills-smoke.sh — exit 0, all six pygraphistry skills discovered.
  • ./bin/evals/codex-skills-smoke.sh — not run, blocked by the same Codex usage limit.
  • scripts/evals/gfql_functional_check.py executes the generated code against pygraphistry and passes.
  • Journey JSON parses; all regex / must_not_regex patterns compile.

@lmeyerov
lmeyerov force-pushed the codex/plan-review-polars-skills branch from 9e1d565 to b6401bf Compare July 25, 2026 03:20
Comment thread .agents/skills/pygraphistry-core/SKILL.md Outdated
lmeyerov and others added 4 commits July 24, 2026 20:22
Journey pygraphistry_gfql_polars_engines_v1 grows 2 -> 7 cases so it measures
skill lift rather than recognition (both original cases already scored 1.0 with
skills off). New cases: functional native-Polars round trip (asserts the result
frame module is polars, forbids any pandas conversion), auto-engine regression
debugging, strict off-engine analytic policy (import path + GFQL_POLARS_CALL_MODE
+ NotImplementedError), the unsupported hypergraph Polars path, and GPU fallback
ownership.

All checks were written against observed behavior of pygraphistry 0.57.0+173:
auto on Polars input -> pandas.DataFrame; polars -> polars.DataFrame; polars-gpu
without RAPIDS -> ImportError; hypergraph(engine='polars') -> AttributeError from
the missing dispatch; set_call_mode at graphistry.compute.gfql.lazy.

Skills: record the exact engine literals (polars-gpu is hyphenated, polars_gpu is
not valid) and that gfql() has no strict= argument, after an eval response
hallucinated both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…found

The functional Polars case forbade `register(` anywhere in the response, so a
model that echoed the prompt's own "no register(), no plot()" phrasing failed a
check it had actually satisfied (0.929, sole failure). Reword the prompt so the
constraint is not echoable and keep only code-shaped negatives; a real register()
call still fails the functional execution check.

DEVELOP.md: record that agents under evaluation probe the installed package, so a
graphistry older than 0.58 makes them "verify" that the Polars engines do not
exist and answer with engine='cudf'. Observed on this box (0.45.4) in 4 of 7
cases, symmetrically across skills on/off. Includes the scoped PYTHONPATH
workaround.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…blished

Claude ran the full 7-case x on/off matrix (14/14 rows, clean baseline
isolation): off 1/7, on 3/7. The result is environment-dominated -- the eval
box's graphistry is 0.45.4, predating the Polars engines, so agents verify
against a stale install. With a Polars-capable checkout on PYTHONPATH the same
four cases go 4/4 in both modes. The configuration worth publishing (released
graphistry >= 0.58, no source tree) is untested and the Codex half is blocked
until 2026-07-29, so README and benchmarks/ stay untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
@lmeyerov
lmeyerov marked this pull request as ready for review July 25, 2026 07:11
lmeyerov and others added 9 commits July 25, 2026 00:39
Three checks failed responses that were correct:

- The functional case asked for "complete self-contained Python code"; an agent
  with a filesystem wrote it to polars_gfql_risk.py and reported the run output,
  so no code reached the graded response text (it had in fact printed
  NODES_MODULE: polars.dataframe.frame). The prompt now requires the program
  inline in a python code block.
- The GPU-fallback case required "(does not|no) ... (automatic|fall back)" within
  60 characters. A response that said "It does not pick the engine" and
  "GPU-or-error" failed it. Widened to the substantive claims.
- The hypergraph case required a literal gfql( call in a question about whether
  hypergraph accepts engine='polars-gpu'. A response that verified against the
  installed 0.58.0 that hyper_dask.py has zero polars branches failed it. The
  oracle rubric still covers the follow-on GFQL point.

Substance bar is unchanged; only the phrasing latitude widened.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
… 0.58.0)

skills=on 100% (7/7, avg score 1.00, 83.4s), skills=off 71.4% (5/7, 0.92,
132.3s): +28.6pp, zero regressions, baseline isolation verified.

Run in the configuration a real user has -- a venv with the released
graphistry==0.58.0 and polars on PATH. That choice is load-bearing: against the
box's stale 0.45.4 install both arms collapse (agents "verify" that the Polars
engines do not exist), and against a source checkout on PYTHONPATH both reach
100% (the baseline just reads the implementation). Only the released install
measures the skill, so the report documents all three regimes.

Both baseline losses are substantive rather than timeouts: it never states that
the default engine resolves a Polars input graph to pandas, and never
establishes that polars-gpu is GPU-or-error with application-owned fallback.

Claude only -- codex usage credits were exhausted for the window, so every codex
cell returned a usage-limit error with no model output and is discarded rather
than scored. Report and CHANGELOG say so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
The pack only tested "which engine, what comes back" -- the shallow half of the
surface. Zero coverage of set_gpu_executor, set_cpu_streaming, the env knobs,
engine crossover, the parity-or-decline boundaries, or the validate/autofix
conversion convention.

Skill gains an engine-tuning section: the three lazy settings with their values,
defaults, env vars and Python-override>env>default resolution; why one plan is
collected once (per-op eager GPU collect is a regression -- repeated H2D); the
parity-or-NotImplementedError contract and which surfaces decline today; and
performance guidance that refuses to promise a speedup.

Numbers are measured, not assumed. On an NVIDIA GB10 via the prebuilt RAPIDS
image, a single-hop MATCH runs 0.83x at 100k rows (GPU SLOWER than CPU), 1.41x
at 1M, 0.98x at 5M -- a band, not a rising curve. Upstream's pandas->polars
crossover (~50-100k) and the CPU-streaming regression on small data (~0.86x)
are attributed as upstream measurements.

Five cases added, each aimed where reading the source or the flag name misleads:
GPU-not-a-blanket-speedup, the CPU-streaming tradeoff, refusing to relabel a
pandas fallback as polars, executor selection, and strict-vs-autofix conversion.

DEVELOP.md documents the GPU harness: prebuilt nvcr.io RAPIDS image on
dgx-spark with a named volume keeping the graphistry install warm, and
--system-site-packages so the image's cudf_polars is visible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Three of the new perf cases failed responses that were right. One answered
"100k -> 0.83x (slower)" in a table; the regex wanted "can be slower". Another
was asked for a wrapper and built one with 14 passing tests, replying with a
file summary, so prose-matching saw nothing.

Accept a bare "slower", tell prose-graded cases to answer inline rather than
produce files, and delete two regexes that pinned one phrasing of a claim the
oracle rubric already judges. These cases are judgment, not keyword presence --
run them with --grading hybrid --oracle-harness claude so the rubrics decide.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Under hybrid grading the baseline answer to the GPU-speedup case passed the
oracle (0.82) and was substantively right -- direct "no", tied to scale, asked
for a benchmark -- but failed a deterministic regex because it argued
"GPU adds strictly more fixed overhead ... no room to amortize" without using
the words "slower" or "depends". Keeping both graders on the same claim
double-penalizes phrasing. The rubric keeps the requirement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
An agent under evaluation edited the installed graphistry inside the eval venv
mid-sweep -- resolve_engine's polars branch, Engine.PANDAS -> Engine.POLARS --
and every cell that ran afterwards, including a full published matrix, executed
against a library that no longer matched the version the report named.

Record a SHA-256 baseline of the package before a sweep and diff it after;
discard the run if it moved. Write-protection helps but is not proof, since the
same agent can undo it. A pack that names a package version is making an
environment claim, and only the checksum substantiates it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…0pp result

The pack's first matrix ran against a library an agent under evaluation had
patched. At 07:21 UTC a cell edited the eval venv's installed
graphistry/Engine.py -- resolve_engine's polars branch, Engine.PANDAS ->
Engine.POLARS -- and every published row ran at 07:41 UTC or later, so the report
named a package version it was not actually running.

Re-ran all 14 cells against a pristine, write-protected 0.58.0 with a SHA-256
baseline over the package verified identical afterwards, under hybrid grading:

  skills=on   6/7 (85.7%), score 0.94, 70.3s
  skills=off  6/7 (85.7%), score 0.92, 102.1s
  delta 0pp; the real difference is latency, ~1.45x

The case that produced the old delta, polars_auto_engine_regression, asks why a
result came back as pandas without an explicit engine. The patch invalidated that
premise, so the baseline failed environmentally; in a clean environment the same
baseline cell passes at 0.90.

README, benchmarks/README and CHANGELOG no longer cite this pack as an
improvement. It stays as a regression harness. The skill guidance itself is
unaffected and was re-verified against pristine 0.58.0: auto still resolves a
polars input graph to pandas.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
GFQL ships a physical index subsystem -- create_index/show_indexes/index_trace,
Cypher index DDL, and an engine-aware cost gate -- with zero mentions in the
skill. It is the largest available win on exactly the queries the engine choice
does not help: small seeded lookups.

Measured, 200k nodes / 1.6M edges, polars:
  chain [n({'id':..}), e_forward(), n()]   8.90ms -> 1.68ms with an index (5.3x)
  cypher WHERE a.id=..                     7.41ms -> 7.82ms, index NEVER consulted
  cypher inline {id:..}                    7.18ms -> 7.07ms, index NEVER consulted
So query shape, not just engine, decides whether an index does anything. The
planner's gate is engine-aware -- pandas ~0.5 of distinct source keys, polars/GPU
~0.02 -- because vectorized scans are fast enough that an index only wins on very
selective seeds.

The decision procedure orders what actually matters: declines first, then where
the frames already live, then indexing, then CPU engine, then GPU. GPU numbers
measured on a GB10 refuse a simple "polars-gpu over cudf" rule:

  1.6M edges 2-hop   polars 486.9  cudf 322.6  polars-gpu 284.8
  8M   edges 1-hop   polars 1379.4 cudf 762.8  polars-gpu 1499.6
  8M   edges 2-hop   polars 2433.5 cudf 1376.6 polars-gpu 3410.0

polars-gpu wins the middle size and loses to CPU polars at 8M. Two shipped facts
explain it: cudf_polars still ingests a host polars frame (Engine.py), so
already-on-device cuDF frames skip a transfer polars-gpu pays; and upstream calls
multi-hop GPU fusion a follow-up where the win dilutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Reading ComputeMixin.gfql shows it is not a plain passthrough: it pops an
index_policy kwarg ('off'|'use'|'auto'|'force') and routes index DDL to the
registry before delegating, so index_policy never appears in the unified gfql()
signature. Verified: 'auto'/'force' probe the index, 'off' does not, and 'use'
silently scans when no index is resident.

DDL forms verified against the parser: CREATE/DROP GFQL INDEX FOR <kind> and
SHOW GFQL INDEXES.

Also record the contrast that prompted this: call mode is NOT a gfql() parameter.
The released signature is query, engine, output, policy, where, language, params,
validate, shortest_path_backend -- call_mode= and strict= both raise TypeError.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
lmeyerov and others added 2 commits July 25, 2026 14:27
Filed graphistry/pygraphistry#1778 asking for strictness as a gfql() parameter.
Note the process-global caveat (scope it and restore in a finally: when only one
step must be strict) and leave a pointer so this section and the
polars_strict_call_mode_benchmark_integrity case get updated if it lands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
This repo is public; the GPU verification section named an internal ssh alias.
Use $GPU_HOST instead -- the instructions are identical for any RAPIDS-capable
box, and the alias added nothing for outside readers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
@lmeyerov
lmeyerov merged commit 35b2227 into main Jul 25, 2026
1 check passed
@lmeyerov
lmeyerov deleted the codex/plan-review-polars-skills branch July 25, 2026 21:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant