Skip to content

5HDqwhrHkPH2eh4atZjseCLpWM58ZdJ1DeLbQpSMzYCq8TMa submission ninja v2 - #1556

Open
Kapuran wants to merge 13 commits into
unarbos:mainfrom
Kapuran:my-submission
Open

Kapuran wants to merge 13 commits into
unarbos:mainfrom
Kapuran:my-submission

Conversation

@Kapuran

@Kapuran Kapuran commented May 14, 2026

Copy link
Copy Markdown

5HDqwhrHkPH2eh4atZjseCLpWM58ZdJ1DeLbQpSMzYCq8TMa

root and others added 13 commits May 12, 2026 22:14
Co-authored-by: Cursor <cursoragent@cursor.com>
- Add FAILURE RECOVERY AND COMMAND ECONOMY section to system prompt
- Add FINAL ANSWER template section for cleaner solve completion
- Add language-specific rules for Python, JS/TS, Shell/SQL
- Enhance verification, error message, API, and edge-case guidance
- Improve hail mary prompt to stress cost of empty patches

Co-authored-by: Cursor <cursoragent@cursor.com>
Mechanical code changes (not prompt-only):
- Parallelize git grep in _symbol_grep_hits and _broad_grep_fallback
  using ThreadPoolExecutor (saves 10-30s wall clock per attempt)
- Fix false-negative test detection: use regex for non-zero fail counts
  instead of substring matching; add Go/Rust/zero-count good markers
- Weighted per-file context budgets: top 3 ranked files get 2x budget
  for better edit-target coverage (SWE-Edit insight)
- Raise stem-length floor from 3 to 5 in _rank_context_files to
  eliminate false positive matches on short words (SweRank insight)
- Add out_of_time() check inside command batch loop to prevent
  wasting time on commands after budget is exhausted
- Word-boundary matching for short tokens (<5 chars) in integration
  partner scoring to eliminate substring false positives

Prompt additions (paired with mechanical changes):
- FAILURE RECOVERY AND COMMAND ECONOMY section
- FINAL ANSWER template
- Python/JS/TS/Shell/SQL language-specific rules
- Enhanced verification, error message, API, and edge-case guidance

Co-authored-by: Cursor <cursoragent@cursor.com>
- Add path-term document-frequency downweighting in file ranking
- Add symbol-centered preload excerpts for top symbol-hit files
- Return symbol_hits from ranking for preload; extend docstring

Co-authored-by: Cursor <cursoragent@cursor.com>
…prompts, wall-clock logging

Co-authored-by: Cursor <cursoragent@cursor.com>
…rification

Co-authored-by: Cursor <cursoragent@cursor.com>
Keep harness improvements: identifier stopwords/hook filter and self-check type cues.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions

Copy link
Copy Markdown

OpenRouter PR Judge

Verdict: PASS
Model: anthropic/claude-opus-4.7
Threshold: 70

Score Value
Overall 74
Real edit 75
Safety 90
Scope 85
Contract 95

Summary

This PR modifies the preloaded context ranking and reading path in agent.py: it adds IDF-style document-frequency downweighting for path-term matches, introduces excerpt-centering around the first issue-symbol match in the top files, allocates per-file budgets non-uniformly across ranked files, parallelizes git grep calls in _broad_grep_fallback and _symbol_grep_hits via ThreadPoolExecutor, replaces the prior _issue_identifier_path_boost with the new IDF logic, adds a mid-batch wall-clock check, tightens _looks_like_successful_test_output with a non-zero failure regex and additional good/bad markers, and expands the system / self-check / hail-mary prompts with anti-partial-ship and wiring-audit guidance.

Static Checks

  • No static findings.

Judge Reasons

  • Adds new mechanics (centered excerpts via _read_context_file_centered, _git_first_line_hitting_symbol, _path_term_df/_idf_path_term_points, weighted per-file budgets) — not prompt-only.
  • _rank_context_files signature changed internally but only callers within agent.py are updated; solve() signature and return shape preserved.
  • Parallel grep via ThreadPoolExecutor stays inside repo via git grep — no new network or filesystem reach.
  • Removes _issue_identifier_path_boost and the constants/regexes feeding it, replacing it with the IDF mechanism; net change is real behavior not cosmetic-copy.
  • Prompt expansion (anti-partial-ship, wiring audit) is paired with substantive mechanism changes, not prompt-only.
  • No new third-party imports (concurrent.futures is stdlib); no forbidden env vars or sampling params; DANGEROUS_PATTERNS preserved.
  • Static guard reports 377 substantive lines / 432 changed and no fail reasons.

Risks

  • goodhart: expanded prompt sections add many tokens of judge-facing guidance (anti-partial-ship etc.) — mostly framed as solver guidance, but verges on persuasion
  • scope-drift: _BRACE_BALANCE_SUFFIXES narrowed (removed rs/go/java/kt/c/cpp/cs/php) — this weakens syntax checking for several languages; reviewer should confirm this is intentional

Required Changes

  • No required changes returned.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant