How NeoSmith SLM Trained on a Single Repo Beats Claude Opus 4 at Code Review
By NeoSmith AI Research Team · March 14, 2026 · 18 min read
TL;DR
We trained Neosmith AI, a NeoSmith SLM, using reinforcement learning (GDPO) to perform automated PR code reviews. Trained on a single complex enterprise repository with 1,000+ files and multiple interlinked sub-repos, in a blind benchmark across 100 real pull requests, Neosmith V2 beat Claude Opus 4 in 65 out of 100 evaluations, scoring an average of 34.5/50 vs 29.4/50 on quality dimensions, while costing 93% less per review.
The Challenge
Automated code review is one of the highest-value applications of LLMs in software engineering. But frontier models have three fundamental problems for code review at scale: cost ($15/1M input tokens for Claude Opus 4), generic knowledge (no codebase-specific context), and latency/control limitations.
The Architecture
We built the NeoSmith SLM with full control over fine-tuning, RL training, and inference — no API vendor lock-in, dramatically lower costs ($1.25/$5.00 per 1M tokens vs $15/$75 for Opus), and deep codebase knowledge embedded in model weights.
Phase 1: Training Data from a Complex Enterprise Repository
The entire training dataset was sourced from a single production-grade repository: a large-scale multi-agent system spanning 1,000+ files. We generated 156 PR review samples and 749 codebase comprehension samples — 905 total training samples. Quality and domain relevance matter far more than quantity.
Phase 3: RL Training with GDPO
We used Group Distilled Policy Optimization (GDPO), our enhanced variant of GRPO with Dr-GRPO tricks + 15% SFT replay. The V1→V2 reward rebalancing shifted from 55% YAML format compliance to 50% LLM quality grading — transforming the model from "great at formatting" to "great at finding bugs."
Benchmark Results: Neosmith V2 vs Claude Opus 4
100 real PRs, blind evaluation, position randomization, Claude Opus 4 as neutral judge. Neosmith V2 won 65/100 evaluations with average quality score 34.5/50 vs 29.4/50. Cost per review: $0.01 vs $0.24 — 93% savings.
Key Takeaways
- Domain-specialized SLMs trained on a single repo can outperform frontier models at specialized tasks
- Reward design matters more than model scale — the V1→V2 breakthrough came from rebalancing reward weights
- 93% cost reduction ($8,280/year saved at 100 PRs/day) while delivering better review quality
About the author: NeoSmith AI Research Team builds automated distillation and reinforcement learning systems for production AI agents.