Introducing Recursion: The RL platform for enterprise specialist agents

Introducing Recursion: The RL platform for enterprise specialist agents

AI agents are moving from isolated experiments into real business workflows. They connect to enterprise tools, operate across internal systems, and execute increasingly complex multi-step tasks across the organization.

But there is a fundamental limitation that will not disappear with better foundation models: general intelligence alone does not create specialized outcomes.

As base models become more accessible, the differentiator will shift from who has access to intelligence to who can continuously improve it. The organizations that win in the AI era will not simply deploy AI. They will build learning systems that turn their expertise, workflows, and decisions into compounding intelligence.

Today, we are introducing Recursion, a unified platform for developing, evaluating, and deploying specialist AI models. Recursion is built on a simple idea: in the age of AI, the ability to learn and improve continuously matters more than any single model. The learning loop is your most durable competitive advantage.

The specialist always beats the generalist

Your business has accumulated something no foundation model has: years of domain expertise, decision patterns, and workflow knowledge embedded in your people and processes. General LLMs approximate this expensively, inconsistently, and at full inference cost with every prompt.

That same expertise, encoded into a specialist model through reinforcement learning, consistently outperforms general LLMs on your specific workflows at a fraction of the token cost. The model that knows your workflows beats the model that knows everything.

A smaller specialist model beats a larger generalist model

Comparison to other systems on the same held-out agentic finance tasks.

System Mean reward pass@1 pass@3 pass@5
#### Qwen3.6 35B (tuned)
ours
0.23 0.37 0.60 0.67
#### Qwen3.5-397B 0.25 0.20 0.41 0.50

In an agentic finance workflow, a Recursion-tuned Qwen3.5-35B open-source model reached 0.37 pass@1, 0.60 pass@3, and 0.67 pass@5, ahead of Qwen3.5-397B on each pass@k metric. This comparison was measured on unseen tasks, which matters: the model learned more robust agentic behaviors rather than overfitting to specific financial procedures. In practice, this translates to a finance agent that reliably reaches correct outcomes across multi-step reasoning, tool use, and structured decision-making without iterative retries.

Customer Support

A customer service agent developed with Recursion achieves these results: more tickets resolved, fewer hallucinations, faster responses, and lower inference cost.

Metric Finetuned OS model GPT-5.5 Claude Opus 4.8
Resolution rate 84% 76% 73%
Reduction in hallucinations 72% 46% 41%
Cost per 1,000 support tickets $32 $158 $176
Time to first token 0.42s 1.10s 1.28s

How Recursion works

Recursion connects environments, evaluation, and training into a unified reinforcement learning loop. Rather than treating deployment as the endpoint, production becomes the training surface. Every outcome becomes an opportunity to improve.

RL environments that reflect real work

Recursion turns workflows, tools, policies, and edge cases into executable environments for RL training and evaluation. Powered by WorldSim, these environments recreate the full enterprise software stack — with configurable world effects that generate diverse, realistic scenarios at scale.

Evaluation systems that measure real execution

The difference between agents that stagnate and agents that improve is measurement. Recursion builds evaluation systems that score intelligence and skill at every level — final outcomes, intermediate decisions, and execution quality — so every run generates the signal your models need to get better.

A training loop that compounds from real work

Every rollout produces graded trajectories that can feed fine-tuning and reinforcement learning. The result is a specialist model that improves from enterprise execution signals instead of synthetic benchmarks alone.

Metric Value
Source 4,300 financial analysis tasks
Task family DCF, LBO, acquisition, projection
Base model GLM 5.1
Training method GRPO
Compute GKE, H100 cluster
Endpoints Baseline and tuned

Evaluations

The enterprise advantage is a learning system

Historically, software helped organizations scale execution. AI introduces something fundamentally different: the ability to scale learning itself.

Every organization possesses a unique set of workflows, judgments, decision patterns, and domain expertise embedded within its people. The question is whether that expertise remains trapped inside individuals or becomes a durable organizational asset.

As foundation models improve and become accessible to everyone, sustainable advantage shifts away from the model itself and toward an organization's ability to continuously capture, evaluate, and improve its own expertise.

From workflow to specialist model — continuously

Every cycle captures additional expertise. Every execution generates a new signal. Every improvement compounds your proprietary edge.

Built for enterprise scale

AI will make intelligence broadly accessible. The enduring advantage will come from what organizations do with it.

The firms that thrive will be those that continuously transform expertise into systems, systems into learning, and learning into proprietary advantage. Recursion helps enterprises build company intelligence that outlasts any individual model.

Privacy and security for enterprise agents

Recursion is built for high-stakes workflows where prompts, traces, reward signals, and training data need enterprise-grade governance. Labelbox applies the same privacy, security, and compliance posture across the systems that power specialist agents.

Recursion supports a growing set of enterprise domains: finance, legal, security, insurance, manufacturing, and operations, with expansion into additional high-value workflows underway.

As agents take on more responsibility across the enterprise, the need for systems that can evaluate, learn, and improve from real execution will only grow. Reliable agents are not discovered. They are engineered through a compounding learning loop.

That is why we built Recursion. Each cycle evaluates execution, turns the result into training signal, and feeds that signal back into the next generation of specialist models agents. The system continually gets better because the learning loop repeats with better context, better measurements, and better outcomes.