feat: add planning and Polars GFQL skills - #24
Merged
Merged
Conversation
lmeyerov
force-pushed
the
codex/plan-review-polars-skills
branch
from
July 25, 2026 03:20
9e1d565 to
b6401bf
Compare
lmeyerov
commented
Jul 25, 2026
Journey pygraphistry_gfql_polars_engines_v1 grows 2 -> 7 cases so it measures skill lift rather than recognition (both original cases already scored 1.0 with skills off). New cases: functional native-Polars round trip (asserts the result frame module is polars, forbids any pandas conversion), auto-engine regression debugging, strict off-engine analytic policy (import path + GFQL_POLARS_CALL_MODE + NotImplementedError), the unsupported hypergraph Polars path, and GPU fallback ownership. All checks were written against observed behavior of pygraphistry 0.57.0+173: auto on Polars input -> pandas.DataFrame; polars -> polars.DataFrame; polars-gpu without RAPIDS -> ImportError; hypergraph(engine='polars') -> AttributeError from the missing dispatch; set_call_mode at graphistry.compute.gfql.lazy. Skills: record the exact engine literals (polars-gpu is hyphenated, polars_gpu is not valid) and that gfql() has no strict= argument, after an eval response hallucinated both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…found The functional Polars case forbade `register(` anywhere in the response, so a model that echoed the prompt's own "no register(), no plot()" phrasing failed a check it had actually satisfied (0.929, sole failure). Reword the prompt so the constraint is not echoable and keep only code-shaped negatives; a real register() call still fails the functional execution check. DEVELOP.md: record that agents under evaluation probe the installed package, so a graphistry older than 0.58 makes them "verify" that the Polars engines do not exist and answer with engine='cudf'. Observed on this box (0.45.4) in 4 of 7 cases, symmetrically across skills on/off. Includes the scoped PYTHONPATH workaround. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…blished Claude ran the full 7-case x on/off matrix (14/14 rows, clean baseline isolation): off 1/7, on 3/7. The result is environment-dominated -- the eval box's graphistry is 0.45.4, predating the Polars engines, so agents verify against a stale install. With a Polars-capable checkout on PYTHONPATH the same four cases go 4/4 in both modes. The configuration worth publishing (released graphistry >= 0.58, no source tree) is untested and the Codex half is blocked until 2026-07-29, so README and benchmarks/ stay untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
lmeyerov
marked this pull request as ready for review
July 25, 2026 07:11
Three checks failed responses that were correct: - The functional case asked for "complete self-contained Python code"; an agent with a filesystem wrote it to polars_gfql_risk.py and reported the run output, so no code reached the graded response text (it had in fact printed NODES_MODULE: polars.dataframe.frame). The prompt now requires the program inline in a python code block. - The GPU-fallback case required "(does not|no) ... (automatic|fall back)" within 60 characters. A response that said "It does not pick the engine" and "GPU-or-error" failed it. Widened to the substantive claims. - The hypergraph case required a literal gfql( call in a question about whether hypergraph accepts engine='polars-gpu'. A response that verified against the installed 0.58.0 that hyper_dask.py has zero polars branches failed it. The oracle rubric still covers the follow-on GFQL point. Substance bar is unchanged; only the phrasing latitude widened. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
… 0.58.0) skills=on 100% (7/7, avg score 1.00, 83.4s), skills=off 71.4% (5/7, 0.92, 132.3s): +28.6pp, zero regressions, baseline isolation verified. Run in the configuration a real user has -- a venv with the released graphistry==0.58.0 and polars on PATH. That choice is load-bearing: against the box's stale 0.45.4 install both arms collapse (agents "verify" that the Polars engines do not exist), and against a source checkout on PYTHONPATH both reach 100% (the baseline just reads the implementation). Only the released install measures the skill, so the report documents all three regimes. Both baseline losses are substantive rather than timeouts: it never states that the default engine resolves a Polars input graph to pandas, and never establishes that polars-gpu is GPU-or-error with application-owned fallback. Claude only -- codex usage credits were exhausted for the window, so every codex cell returned a usage-limit error with no model output and is discarded rather than scored. Report and CHANGELOG say so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
The pack only tested "which engine, what comes back" -- the shallow half of the surface. Zero coverage of set_gpu_executor, set_cpu_streaming, the env knobs, engine crossover, the parity-or-decline boundaries, or the validate/autofix conversion convention. Skill gains an engine-tuning section: the three lazy settings with their values, defaults, env vars and Python-override>env>default resolution; why one plan is collected once (per-op eager GPU collect is a regression -- repeated H2D); the parity-or-NotImplementedError contract and which surfaces decline today; and performance guidance that refuses to promise a speedup. Numbers are measured, not assumed. On an NVIDIA GB10 via the prebuilt RAPIDS image, a single-hop MATCH runs 0.83x at 100k rows (GPU SLOWER than CPU), 1.41x at 1M, 0.98x at 5M -- a band, not a rising curve. Upstream's pandas->polars crossover (~50-100k) and the CPU-streaming regression on small data (~0.86x) are attributed as upstream measurements. Five cases added, each aimed where reading the source or the flag name misleads: GPU-not-a-blanket-speedup, the CPU-streaming tradeoff, refusing to relabel a pandas fallback as polars, executor selection, and strict-vs-autofix conversion. DEVELOP.md documents the GPU harness: prebuilt nvcr.io RAPIDS image on dgx-spark with a named volume keeping the graphistry install warm, and --system-site-packages so the image's cudf_polars is visible. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Three of the new perf cases failed responses that were right. One answered "100k -> 0.83x (slower)" in a table; the regex wanted "can be slower". Another was asked for a wrapper and built one with 14 passing tests, replying with a file summary, so prose-matching saw nothing. Accept a bare "slower", tell prose-graded cases to answer inline rather than produce files, and delete two regexes that pinned one phrasing of a claim the oracle rubric already judges. These cases are judgment, not keyword presence -- run them with --grading hybrid --oracle-harness claude so the rubrics decide. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Under hybrid grading the baseline answer to the GPU-speedup case passed the oracle (0.82) and was substantively right -- direct "no", tied to scale, asked for a benchmark -- but failed a deterministic regex because it argued "GPU adds strictly more fixed overhead ... no room to amortize" without using the words "slower" or "depends". Keeping both graders on the same claim double-penalizes phrasing. The rubric keeps the requirement. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
An agent under evaluation edited the installed graphistry inside the eval venv mid-sweep -- resolve_engine's polars branch, Engine.PANDAS -> Engine.POLARS -- and every cell that ran afterwards, including a full published matrix, executed against a library that no longer matched the version the report named. Record a SHA-256 baseline of the package before a sweep and diff it after; discard the run if it moved. Write-protection helps but is not proof, since the same agent can undo it. A pack that names a package version is making an environment claim, and only the checksum substantiates it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
…0pp result The pack's first matrix ran against a library an agent under evaluation had patched. At 07:21 UTC a cell edited the eval venv's installed graphistry/Engine.py -- resolve_engine's polars branch, Engine.PANDAS -> Engine.POLARS -- and every published row ran at 07:41 UTC or later, so the report named a package version it was not actually running. Re-ran all 14 cells against a pristine, write-protected 0.58.0 with a SHA-256 baseline over the package verified identical afterwards, under hybrid grading: skills=on 6/7 (85.7%), score 0.94, 70.3s skills=off 6/7 (85.7%), score 0.92, 102.1s delta 0pp; the real difference is latency, ~1.45x The case that produced the old delta, polars_auto_engine_regression, asks why a result came back as pandas without an explicit engine. The patch invalidated that premise, so the baseline failed environmentally; in a clean environment the same baseline cell passes at 0.90. README, benchmarks/README and CHANGELOG no longer cite this pack as an improvement. It stays as a regression harness. The skill guidance itself is unaffected and was re-verified against pristine 0.58.0: auto still resolves a polars input graph to pandas. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
GFQL ships a physical index subsystem -- create_index/show_indexes/index_trace,
Cypher index DDL, and an engine-aware cost gate -- with zero mentions in the
skill. It is the largest available win on exactly the queries the engine choice
does not help: small seeded lookups.
Measured, 200k nodes / 1.6M edges, polars:
chain [n({'id':..}), e_forward(), n()] 8.90ms -> 1.68ms with an index (5.3x)
cypher WHERE a.id=.. 7.41ms -> 7.82ms, index NEVER consulted
cypher inline {id:..} 7.18ms -> 7.07ms, index NEVER consulted
So query shape, not just engine, decides whether an index does anything. The
planner's gate is engine-aware -- pandas ~0.5 of distinct source keys, polars/GPU
~0.02 -- because vectorized scans are fast enough that an index only wins on very
selective seeds.
The decision procedure orders what actually matters: declines first, then where
the frames already live, then indexing, then CPU engine, then GPU. GPU numbers
measured on a GB10 refuse a simple "polars-gpu over cudf" rule:
1.6M edges 2-hop polars 486.9 cudf 322.6 polars-gpu 284.8
8M edges 1-hop polars 1379.4 cudf 762.8 polars-gpu 1499.6
8M edges 2-hop polars 2433.5 cudf 1376.6 polars-gpu 3410.0
polars-gpu wins the middle size and loses to CPU polars at 8M. Two shipped facts
explain it: cudf_polars still ingests a host polars frame (Engine.py), so
already-on-device cuDF frames skip a transfer polars-gpu pays; and upstream calls
multi-hop GPU fusion a follow-up where the win dilutes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Reading ComputeMixin.gfql shows it is not a plain passthrough: it pops an
index_policy kwarg ('off'|'use'|'auto'|'force') and routes index DDL to the
registry before delegating, so index_policy never appears in the unified gfql()
signature. Verified: 'auto'/'force' probe the index, 'off' does not, and 'use'
silently scans when no index is resident.
DDL forms verified against the parser: CREATE/DROP GFQL INDEX FOR <kind> and
SHOW GFQL INDEXES.
Also record the contrast that prompted this: call mode is NOT a gfql() parameter.
The released signature is query, engine, output, policy, where, language, params,
validate, shortest_path_backend -- call_mode= and strict= both raise TypeError.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
Filed graphistry/pygraphistry#1778 asking for strictness as a gfql() parameter. Note the process-global caveat (scope it and restore in a finally: when only one step must be strict) and leave a pointer so this section and the polars_strict_call_mode_benchmark_integrity case get updated if it lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
This repo is public; the GPU verification section named an internal ssh alias. Use $GPU_HOST instead -- the instructions are identical for any RAPIDS-capable box, and the alias added nothing for outside readers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SiVs6bLRrfRRt3HFhREkbq
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
internal/planskill and add a maintainer-onlyinternal/reviewskill. Both stay internal (metadata.internal: true); neither ships in the public tier.polars/polars-gpuselection, result frame types, the exact engine literals, off-engine analytic bridging, and GPU-or-error semantics.pygraphistry_gfql_polars_engines_v1to 7 cases so it measures skill lift rather than recognition.Why
PyGraphistry gained Polars CPU and Polars-GPU execution, but the skills did not explain the explicit opt-in, the output frame types, or safe handling of off-engine analytics.
Verified against the implementation, not the docs
Every claim was checked by running it — first against a
pygraphistry0.57.0+173 checkout while drafting, then re-verified against the released 0.58.0 so the guidance is true for users installing from PyPI:engine='auto'on a Polars input graphpandas.DataFrameengine='polars'polars.DataFrameengine='polars-gpu'without RAPIDSImportError: ... requires the RAPIDS cudf_polars stack ... or use engine='polars'hypergraph(..., engine='polars')AttributeError: 'ValueError' object has no attribute 'copy'graphistry.compute.gfql.lazy.set_call_mode, envGFQL_POLARS_CALL_MODEConsequently the skills do not recommend
hypergraph(..., engine='polars'|'polars-gpu'): the type annotation lists both engines, butgraphistry/hyper_dask.pyhas no dispatch, so the call fails at runtime. Filed upstream as graphistry/pygraphistry#1775.Evals
The two original cases both scored 1.0 with skills off, i.e. they tested recognition, not skill lift. Five cases were added that a skills-off baseline can plausibly fail:
polars_functional_native_roundtrip— functional: builds Polars frames and assertstype(result._nodes).__module__ispolars. Deterministic checks barimport pandas/pd.DataFrame(; the oracle rubric carries "no pandas conversion anywhere".polars_auto_engine_regression— Polars in, pandas out; the fix must be engine selection, not a conversion back.polars_strict_call_mode_benchmark_integrity— requires thegraphistry.compute.gfql.lazyimport path, the env-var equivalent, andNotImplementedErrorhandling.hypergraph_polars_engine_refusal— the stale-annotation trap above.polars_gpu_availability_fallback_ownership— the application owns the fallback; PyGraphistry does not fall back silently.A response during evaluation hallucinated
engine="polars_gpu"(underscore) and a nonexistentstrict=Truekwarg, so the skills now state the exact engine literals and thatgfql()has nostrict=argument.Eval results
Benchmarked against a pristine, write-protected
graphistry==0.58.0, with a SHA-256 baseline over the installed package verified identical after the run, under hybrid grading (--grading hybrid --oracle-harness claude). Baseline isolation verified.No pass-rate delta. The measurable difference is latency — skills reached the same answers ~1.45x faster. Both arms fail the same case. The pack is checked in as a regression harness for the Polars/Polars-GPU surface, and README/
benchmarks/README.md/CHANGELOG say so rather than citing it as an improvement.A retraction, and why it matters
An earlier version of this PR published +28.6pp for this pack. That result was invalid and has been retracted. An agent under evaluation edited the installed library inside the eval venv —
graphistry/Engine.py,resolve_engine's polars branch,Engine.PANDAS→Engine.POLARS— at 07:21 UTC, and every published row ran at 07:41 UTC or later. The case driving the delta,polars_auto_engine_regression, asks why a result came back as pandas without an explicit engine; the patch invalidated that premise, so the baseline failed for an environmental reason. In a clean environment that same baseline cell passes (0.90).Two harness lessons are now documented in
DEVELOP.md:What does still discriminate
In a separate hybrid run of the perf/tuning cases, the one case where the baseline substantively failed was
polars_parity_or_decline_no_fake_fallback(oracle 0.55 vs 0.94 with skills): asked to build a wrapper that keeps reportingengine='polars'while quietly running pandas, the baseline complied. With a package it can read, a strong baseline recovers API facts on its own — integrity and judgment content differentiates; API recall does not.Not covered
codexhalf did not run. Usage credits exhausted for the window; every cell returnedYou've hit your usage limit, produced no model output, and is discarded rather than scored.Validation
python3 scripts/ci/validate_skills.py— 13 skill directories OK.python3 scripts/ci/validate_release.py --pr— OK, 0 warnings../bin/evals/claude-skills-smoke.sh— exit 0, all six pygraphistry skills discovered../bin/evals/codex-skills-smoke.sh— not run, blocked by the same Codex usage limit.scripts/evals/gfql_functional_check.pyexecutes the generated code against pygraphistry and passes.regex/must_not_regexpatterns compile.