Repository navigation
MAINT: Target Tool Calling Capability Improvements - #2852
Merged
Richard Lundeen (richlundeen) merged 6 commits intoSep 30, 2026
Merged
Richard Lundeen (richlundeen) merged 6 commits into
Richard Lundeen (richlundeen) merged 6 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bring the shared tool-content models and no-send target preflight from PR microsoft#2853 into the tool capability change. Reuse payload validation across Chat, LiteLLM, and Responses while keeping editor persistence in its own PR. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve complete tool execution and request-local retry/pacing from microsoft#2718. Share the request helper with single-response mode, suppress provider activity during capability probes, and preserve configured target identities. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Filter nested SDK tool settings during probes, render objective tool exchanges as adversarial context, label simulated tools correctly, and separate strict draft validation from provider argument replay. Include focused regressions and shorten the target overview. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Roman Lutz (romanlutz)
approved these changes
Sep 29, 2026
Mark prepended seed history without marking outgoing user prompts. Preserve synthetic assistant/tool roles and update the role contract regression for simulated_tool. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Richard Lundeen (richlundeen)
enabled auto-merge
September 30, 2026 00:03
Richard Lundeen (richlundeen)
deleted the
richlundeen-tool-call-support-plan
branch
September 30, 2026 00:45
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PyRIT already represented tool calls and results as input modalities, but support for using them as conversation history was inconsistent across target adapters. Capability discovery did not test tool-history acceptance, and callers had no shared validation API to check history without sending a request. Synthetic tool results also lacked a distinct role, so copied conversations, scoring, and display could lose the distinction between injected history and actual execution evidence.
This PR uses the existing input modalities as the tool-history capability contract, without adding a separate boolean. It adds shared tool-content models, consistent Chat/Responses parsing and serialization of converted values, and
validate_tool_history()for checks without sending requests. Capability probes test history acceptance without running configured tools or opening provider sessions, while preserving target settings and identity. Responses targets also supportexecute_tools=Falseto return the first response without local tool execution. The newsimulated_toolrole and shared prepended-history marking preserve synthetic content through conversation copies, scoring, UI, and exports. Trace-based scoring uses actual execution evidence rather than injected message claims. Send-time validation remains after normalization, so ADAPT behavior is unchanged.