Internalized Tool Calling: How a NeoSmith SLM Beats GPT-5.2 at Multi-Turn Customer Support
By NeoSmith AI Research Team · March 19, 2026 · 22 min read
TL;DR
We trained Neosmith AI, a NeoSmith SLM, to handle complex customer support scenarios requiring multi-turn tool calling, database queries, policy lookups, and booking modifications. In a blind benchmark across 20 real Swiss Airlines support scenarios, Neosmith V5 won 16 out of 20 evaluations against GPT-5.2, scoring 6.0/10 vs 0.5/10, with Claude Opus 4 as the neutral judge. The breakthrough: internalized tool calling combined with a 5-signal reward function and Reasoning Cache for multi-turn refinement.
Key Results
- 16/20 wins vs GPT-5.2 (80% win rate)
- 6.0 vs 0.5 average quality score
- 94% cost reduction ($0.02 vs $0.35 per session)
- Dominates on hard multi-step scenarios (4-0 with 2 ties)
What Is Internalized Tool Calling?
Traditional LLM tool calling works through an external API layer. Internalized tool calling means the model generates tool calls as regular tokens, flowing through the attention mechanism as a continuous sequence. This enables multi-step planning, error recovery, and policy compliance as learned behaviors rather than engineered rules.
The 5-Signal Reward Function
- Tool Execution (35%) - Did tool calls succeed with valid arguments?
- Correctness (25%) - Is the database in the right state?
- Policy Compliance (15%) - Proper operational procedures followed?
- LLM Quality (10%) - Claude Opus holistic evaluation
- Response Quality (15%) - Natural language quality around tool calls
Training Journey
From 4-15 loss (V3, 59 samples) to 10-10 tie (V4, 648 samples) to 16-1 victory (V5, 806 samples with hard multi-step traces). The key lever: adding 158 hard multi-step traces covering 4-7 tool chains.
Key Takeaways
- Hard examples matter more than volume for tool calling
- Internalized tool calls create emergent multi-step planning
- Real database execution is the most impactful training signal
- A neutral judge (Claude Opus) produces more honest evaluations
- Multi-turn Reasoning Cache dramatically improves complex scenarios