Skip to content

Week 4 Day 7: Evaluate Observable Outcomes - #257

Merged
skyzh merged 4 commits into
mainfrom
agent/week4v3-day7-harness
Aug 12, 2026
Merged

Week 4 Day 7: Evaluate Observable Outcomes#257
skyzh merged 4 commits into
mainfrom
agent/week4v3-day7-harness

Conversation

@skyzh

@skyzh skyzh commented Aug 11, 2026

Copy link
Copy Markdown
Owner

Why

Close Week 4 with one learner-visible way to judge a completed agent run from declared observable outcomes, without grading hidden reasoning or one exact transcript.

Change

  • add a compact reference evaluator plus solution-free TODO starter for final, file, result, and named receipt facts
  • return stable structured checks and a deterministic aggregate report while keeping evaluation read-only
  • compare expected UTF-8 file content from exact bytes, preserving CRLF/LF distinctions and rejecting malformed UTF-8
  • add one scripted edit-and-validate scenario, causal evidence regressions, cumulative starter sync, and the Day 7 learner chapter/navigation

Review note

Final exact-head review is complete on the strict-UTF-8 successor.

Base b107310ab977b0a434e54f6b694b2f91b0f1652e; head 2a8942a6e29547e96730803e5dd5f0d01312e38e; tree b60e25dbd2401d21c54ef0a22d96799a64c168fe. The successor from rejected c47407f19cd13580602fc64ae7dcce776b5d91bf changes only tests_refsol/test_week_4_day_7.py (28 insertions); every production/API/starter/prose/navigation/Day 1–6/stack byte is unchanged. Total main-relative delta: 11 paths, 957 insertions, 260 deletions.

Final evidence: Day 7 plus sync 51/51; cumulative Week 4 plus sync 108/108; full reference 443 passed with 8 intentional skips; Ruff/format/diff clean; learner copy byte-identical and expected-red 22/22 before cleanup; manual mdBook build passed. Independent malformed-byte cases fail if strict decoding is mutated to errors="ignore" or errors="replace"; the prior nine causal controls remain green. Hosted macOS run 31600359755, job 94125932294, succeeded on the exact head.

AI-Assisted: GPT-5.6 Sol + Forge
Reviewed-by: GPT-5.6 Terra + Oracle (consistency)
Reviewed-by: GPT-5.6 Sol + Sage (correctness)
Reviewed-by: GPT-5.6 Sol + Tuner (performance)
Reviewed-by: GPT-5.6 Terra + Scholar (learner)

AI-Assisted: GPT-5.6 Sol + Forge
@skyzh
skyzh force-pushed the agent/week4v3-day7-harness branch from 173c7d1 to 7b29c02 Compare August 12, 2026 11:48
@skyzh
skyzh changed the base branch from agent/week4v3-day6-reconcile to main August 12, 2026 11:52
@skyzh skyzh changed the title Week 4 equivalence harness: semantic/evidence/policy comparison Week 4 Day 7: Evaluate Observable Outcomes Aug 12, 2026
@skyzh skyzh closed this Aug 12, 2026
@skyzh skyzh reopened this Aug 12, 2026
skyzh added 3 commits August 12, 2026 05:24
AI-Assisted: GPT-5.6 Sol + Forge
AI-Assisted: GPT-5.6 Sol + Forge
AI-Assisted: GPT-5.6 Sol + Forge
@skyzh
skyzh marked this pull request as ready for review August 12, 2026 13:33
@skyzh
skyzh merged commit e4c3a22 into main Aug 12, 2026
1 check passed
@skyzh
skyzh deleted the agent/week4v3-day7-harness branch August 12, 2026 13:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant