Skip to content

5DVFGXociNnsbyPhzFe4iY6SisLEUE6eASWUf11Mn3KXffqU update - #1567

Open
ninja66-contributor wants to merge 1 commit into
unarbos:mainfrom
ninja66-contributor:miner/baihu002-1778774075
Open

ninja66-contributor wants to merge 1 commit into
unarbos:mainfrom
ninja66-contributor:miner/baihu002-1778774075

Conversation

@ninja66-contributor

Copy link
Copy Markdown

No description provided.

@github-actions

Copy link
Copy Markdown

OpenRouter PR Judge

Verdict: PASS
Model: anthropic/claude-opus-4.7
Threshold: 70

Score Value
Overall 74
Real edit 76
Safety 95
Scope 90
Contract 98

Summary

This PR adds several mechanical changes to agent.py: a new _strip_mode_headers_from_content_diffs sanitization step that strips old/new mode header lines from content-bearing diffs; a TF-IDF-style doc-frequency boost in _rank_context_files (rare terms get +8, common stay at floor +3); an expanded identifier-stopword set with hook-prefix filtering; a type-aware self-check cue prepended to the self-check prompt; a soft-nudge mid-loop turn fired when many steps elapse without edits; an under-deliver detection path in the multishot wrapper that triggers attempt 2 with a tailored bootstrap when patch count is below estimated file count; and a small bump to the recently-observed-paths window. The solve signature, return shape, stdlib-only constraint, and DANGEROUS_PATTERNS are preserved.

Static Checks

  • No static findings.

Judge Reasons

  • Multiple distinct mechanical changes: patch sanitization, file-ranking score function, multishot underdeliver branch, and a new soft-nudge loop turn — not prompt-only.
  • TF-IDF bonus is floor-clamped at the prior +3 so it cannot decrease ranking, a real algorithmic refinement rather than a cosmetic rename.
  • solve(...) signature, return dict shape, env-var allowlist, and validator-owned boundaries appear unchanged.
  • No new third-party imports, no new network endpoints, no sampling parameters added.
  • DANGEROUS_PATTERNS list left intact; no weakening of safety filters.
  • Some new prompt strings (type-aware cue, under-deliver bootstrap) include style/imitation guidance that edges toward judge-facing prose, but each is paired with real mechanical wiring (new branch / new gate), so not prompt-only.
  • The under-deliver bootstrap text instructs the model to 'imitate the codebase's existing style' — borderline style-anchoring but reasonable since it triggers only on multi-file tasks where attempt 1 under-covered.

Risks

  • goodhart: new _strip_mode_headers_from_content_diffs strips mode header lines specifically because 'the LLM judge dings agents for unrelated file mode churn' — this is explicit judge-shaping, though the effect (removing chmod-only noise) is also a legitimate cleanup.
  • goodhart (minor): type-aware self-check cue and under-deliver bootstrap contain motivational/style language directed at shaping output for the judge; modest in scale and paired with mechanism.

Required Changes

  • No required changes returned.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant