Introducing Recursion: The RL platform for enterprise specialist agents
Introducing Recursion: The RL platform for enterprise specialist agents
AI agents are moving from isolated experiments into real business workflows. They connect to enterprise tools, operate across internal systems, and execute increasingly complex multi-step tasks across the organization.
But there is a fundamental limitation that will not disappear with better foundation models: general intelligence alone does not create specialized outcomes.
As base models become more accessible, the differentiator will shift from who has access to intelligence to who can continuously improve it. The organizations that win in the AI era will not simply deploy AI. They will build learning systems that turn their expertise, workflows, and decisions into compounding intelligence.
Today, we are introducing Recursion, a unified platform for developing, evaluating, and deploying specialist AI models. Recursion is built on a simple idea: in the age of AI, the ability to learn and improve continuously matters more than any single model. The learning loop is your most durable competitive advantage.
The specialist always beats the generalist
Your business has accumulated something no foundation model has: years of domain expertise, decision patterns, and workflow knowledge embedded in your people and processes. General LLMs approximate this expensively, inconsistently, and at full inference cost with every prompt.
That same expertise, encoded into a specialist model through reinforcement learning, consistently outperforms general LLMs on your specific workflows at a fraction of the token cost. The model that knows your workflows beats the model that knows everything.
A smaller specialist model beats a larger generalist model
Comparison to other systems on the same held-out agentic finance tasks.
| System | Mean reward | pass@1 | pass@3 | pass@5 |
|---|---|---|---|---|
| #### Qwen3.6 35B (tuned) ours |
0.23 | 0.37 | 0.60 | 0.67 |
| #### Qwen3.5-397B | 0.25 | 0.20 | 0.41 | 0.50 |
In an agentic finance workflow, a Recursion-tuned Qwen3.5-35B open-source model reached 0.37 pass@1, 0.60 pass@3, and 0.67 pass@5, ahead of Qwen3.5-397B on each pass@k metric. This comparison was measured on unseen tasks, which matters: the model learned more robust agentic behaviors rather than overfitting to specific financial procedures. In practice, this translates to a finance agent that reliably reaches correct outcomes across multi-step reasoning, tool use, and structured decision-making without iterative retries.
Customer Support
A customer service agent developed with Recursion achieves these results: more tickets resolved, fewer hallucinations, faster responses, and lower inference cost.
| Metric | Finetuned OS model | GPT-5.5 | Claude Opus 4.8 |
|---|---|---|---|
| Resolution rate | 84% | 76% | 73% |
| Reduction in hallucinations | 72% | 46% | 41% |
| Cost per 1,000 support tickets | $32 | $158 | $176 |
| Time to first token | 0.42s | 1.10s | 1.28s |
How Recursion works
Recursion connects environments, evaluation, and training into a unified reinforcement learning loop. Rather than treating deployment as the endpoint, production becomes the training surface. Every outcome becomes an opportunity to improve.
RL environments that reflect real work
Recursion turns workflows, tools, policies, and edge cases into executable environments for RL training and evaluation. Powered by WorldSim, these environments recreate the full enterprise software stack — with configurable world effects that generate diverse, realistic scenarios at scale.
Evaluation systems that measure real execution
The difference between agents that stagnate and agents that improve is measurement. Recursion builds evaluation systems that score intelligence and skill at every level — final outcomes, intermediate decisions, and execution quality — so every run generates the signal your models need to get better.
A training loop that compounds from real work
Every rollout produces graded trajectories that can feed fine-tuning and reinforcement learning. The result is a specialist model that improves from enterprise execution signals instead of synthetic benchmarks alone.
| Metric | Value |
|---|---|
| Source | 4,300 financial analysis tasks |
| Task family | DCF, LBO, acquisition, projection |
| Base model | GLM 5.1 |
| Training method | GRPO |
| Compute | GKE, H100 cluster |
| Endpoints | Baseline and tuned |
Evaluations
The enterprise advantage is a learning system
Historically, software helped organizations scale execution. AI introduces something fundamentally different: the ability to scale learning itself.
Every organization possesses a unique set of workflows, judgments, decision patterns, and domain expertise embedded within its people. The question is whether that expertise remains trapped inside individuals or becomes a durable organizational asset.
As foundation models improve and become accessible to everyone, sustainable advantage shifts away from the model itself and toward an organization's ability to continuously capture, evaluate, and improve its own expertise.
From workflow to specialist model — continuously
Every cycle captures additional expertise. Every execution generates a new signal. Every improvement compounds your proprietary edge.
Built for enterprise scale
AI will make intelligence broadly accessible. The enduring advantage will come from what organizations do with it.
The firms that thrive will be those that continuously transform expertise into systems, systems into learning, and learning into proprietary advantage. Recursion helps enterprises build company intelligence that outlasts any individual model.
Privacy and security for enterprise agents
Recursion is built for high-stakes workflows where prompts, traces, reward signals, and training data need enterprise-grade governance. Labelbox applies the same privacy, security, and compliance posture across the systems that power specialist agents.
Recursion supports a growing set of enterprise domains: finance, legal, security, insurance, manufacturing, and operations, with expansion into additional high-value workflows underway.
As agents take on more responsibility across the enterprise, the need for systems that can evaluate, learn, and improve from real execution will only grow. Reliable agents are not discovered. They are engineered through a compounding learning loop.
That is why we built Recursion. Each cycle evaluates execution, turns the result into training signal, and feeds that signal back into the next generation of specialist models agents. The system continually gets better because the learning loop repeats with better context, better measurements, and better outcomes.