Evaluation report: Needle 2 as a tool-call router (4-phase experiment: retrieval, fine-tuning, argument accuracy, confidence gating)
Sharing results from evaluating Needle 2 as a candidate to replace an existing tool-routing layer in a small local voice-assistant project. Not a bug report (we independently isolated the same bug down to a 2-example toy dataset before finding it had already been reported and fixed as #53, thanks for the quick turnaround). Just a data point that might be useful for others sizing up Needle for a similar use case, and a question for the maintainers on whether our results match expectations.
Setup
- Catalogue: 16 tools across 6 categories (1-5 tools per category), declared with realistic natural-language descriptions and typed parameters via
@tool.
- Prompts and tool descriptions in a non-English language (Spanish). Flagging this since Needle 2's training data mix may skew English-heavy; it's a real confound we can't fully separate from raw model capability.
- Semantic overlap: several tools within the same category share similar verbs/shapes (e.g. multiple "create an X" actions with overlapping parameter names, multiple mail-lifecycle actions like search, read, draft, send that take similar-looking string arguments). Probably close to a worst case for a 45M-param retrieval+extraction model, and likely a big part of the gap below.
- Baseline comparison: an existing production 2-stage router (lightweight LLM classification into a category, then a ~30B local LLM constrained to that category's tools), measured at 3/3 and 3/3 on the two test scenarios below.
- Used only
Needle.complete() throughout (decision only); tool execution is always gated by application code, never by the model itself, so we never called .run().
Phase 1: baseline accuracy (no fine-tuning)
| Scenario |
Tool selection |
Argument accuracy (on correct-tool picks) |
| Full 16-tool catalogue, single target tool, 3 phrasings |
2/3 |
0/3 (arguments consistently disconnected from the query, e.g. a "task description" field kept getting filled with an unrelated noun from the prompt rather than the actual task) |
| 5-tool category subset, 5 distinct queries |
1/5 |
1/1 (only correct pick had correct args) |
Phase 2: fine-tuning
~300 hand-templated examples (not LLM-generated) across the 4 busiest categories, positive + negative + ambiguous examples, query/tools/answers/reasoning format. Hit the NaN-at-step-2 bug here, isolated it down to a 2-example toy dataset before finding the issue was already reported. Installed from main post-fix and training completed cleanly (final loss 1.40, LoRA rank 16, defaults otherwise, 3 epochs).
Phase 3: post-fine-tuning re-test (same two scenarios)
| Scenario |
Tool selection |
Argument accuracy (hand-audited, not substring-matched) |
| Full 16-tool catalogue |
2/3 (unchanged) |
0/3 (unchanged) |
| 5-tool category subset |
3/5 (up from 1/5) |
2/5 fully correct end-to-end (one pick had the right tool + one right field but a fabricated date the query never provided) |
Headline finding: fine-tuning on a small clean dataset measurably improved retrieval (which tool) but did not proportionally improve argument grounding (correct field values). Invented/hallucinated arguments persisted even on some correct tool picks, post-fine-tune.
Phase 4: confidence-gated hybrid (fall back to the existing router below a threshold)
Using the 8 combined test cases with Needle's own reported confidence score:
| Threshold |
% handled by Needle alone |
Precision on handled cases |
Hybrid avg. latency* |
| 0.001 |
87.5% |
28.6% |
0.95s |
| 0.05 |
50.0% |
25.0% |
3.01s |
| 0.1 |
37.5% |
33.3% |
3.70s |
| 0.2 |
25.0% |
50.0% |
4.39s |
| 0.35 |
12.5% |
100.0% |
5.08s |
| 0.4 |
0% |
N/A |
5.76s (about equal to baseline) |
* baseline (existing router alone) is about 5.5s avg per turn in our production logs; hybrid latency = Needle's own decision time on handled turns, plus Needle's probe time and baseline latency on fallback turns.
The only threshold with perfect precision only covers 1 of our 8 cases (n=8 is tiny, please don't read the 100% as a real number, it's a single data point), for an ~8% average latency improvement in the best case. Below that threshold, precision degrades quickly (down to 20-50%), which we judged too risky for an assistant that sometimes proposes irreversible actions.
Takeaway (for us)
Not adopting Needle for this router in its current form. The combination of "argument grounding doesn't fine-tune away as easily as retrieval does" and "confidence doesn't cleanly separate good/bad decisions at usable coverage" was the deciding factor, not raw speed (which is excellent: decisions consistently landed in 0.1-0.8s regardless of correctness).
Genuinely curious whether the argument-grounding gap is a known characteristic at this parameter count / with this training approach, or something we could address differently (larger fine-tuning set? different confidence calibration? separating "which tool" from "which arguments" into two passes?). Happy to share more methodology detail if useful.
Related: #53 (NaN during fine-tuning, independently reproduced during this evaluation).
Evaluation report: Needle 2 as a tool-call router (4-phase experiment: retrieval, fine-tuning, argument accuracy, confidence gating)
Sharing results from evaluating Needle 2 as a candidate to replace an existing tool-routing layer in a small local voice-assistant project. Not a bug report (we independently isolated the same bug down to a 2-example toy dataset before finding it had already been reported and fixed as #53, thanks for the quick turnaround). Just a data point that might be useful for others sizing up Needle for a similar use case, and a question for the maintainers on whether our results match expectations.
Setup
@tool.Needle.complete()throughout (decision only); tool execution is always gated by application code, never by the model itself, so we never called.run().Phase 1: baseline accuracy (no fine-tuning)
Phase 2: fine-tuning
~300 hand-templated examples (not LLM-generated) across the 4 busiest categories, positive + negative + ambiguous examples,
query/tools/answers/reasoningformat. Hit the NaN-at-step-2 bug here, isolated it down to a 2-example toy dataset before finding the issue was already reported. Installed frommainpost-fix and training completed cleanly (final loss 1.40, LoRA rank 16, defaults otherwise, 3 epochs).Phase 3: post-fine-tuning re-test (same two scenarios)
Headline finding: fine-tuning on a small clean dataset measurably improved retrieval (which tool) but did not proportionally improve argument grounding (correct field values). Invented/hallucinated arguments persisted even on some correct tool picks, post-fine-tune.
Phase 4: confidence-gated hybrid (fall back to the existing router below a threshold)
Using the 8 combined test cases with Needle's own reported confidence score:
* baseline (existing router alone) is about 5.5s avg per turn in our production logs; hybrid latency = Needle's own decision time on handled turns, plus Needle's probe time and baseline latency on fallback turns.
The only threshold with perfect precision only covers 1 of our 8 cases (n=8 is tiny, please don't read the 100% as a real number, it's a single data point), for an ~8% average latency improvement in the best case. Below that threshold, precision degrades quickly (down to 20-50%), which we judged too risky for an assistant that sometimes proposes irreversible actions.
Takeaway (for us)
Not adopting Needle for this router in its current form. The combination of "argument grounding doesn't fine-tune away as easily as retrieval does" and "confidence doesn't cleanly separate good/bad decisions at usable coverage" was the deciding factor, not raw speed (which is excellent: decisions consistently landed in 0.1-0.8s regardless of correctness).
Genuinely curious whether the argument-grounding gap is a known characteristic at this parameter count / with this training approach, or something we could address differently (larger fine-tuning set? different confidence calibration? separating "which tool" from "which arguments" into two passes?). Happy to share more methodology detail if useful.
Related: #53 (NaN during fine-tuning, independently reproduced during this evaluation).