Recursion | Labelbox

The specialist always beats the generalist

Recursion is the RL platform for developing, evaluating, and deploying specialist AI models that improve from real enterprise execution.

A customer service agent resolved more tickets with fewer hallucinations, faster responses, and lower inference cost

Resolution rate

Model Resolution Rate
Finetuned OS model 84%
GPT-5.5 76%
Claude Opus 4.8 73%

Reduction in hallucinations

Model Reduction Rate
Finetuned OS model 72%
GPT-5.5 46%
Claude Opus 4.8 41%

Cost per 1,000 support tickets

Model Cost
Finetuned OS model $32
GPT-5.5 $158
Claude Opus 4.8 $176

Time to first token

Model Time
Finetuned OS model 0.42s
GPT-5.5 1.10s
Claude Opus 4.8 1.28s

General agents produce general outcomes

General-purpose agents can appear strong in isolation. They break in consistent ways once deployed into real enterprise environments: struggling to maintain reliability across multi-step workflows, adapt to edge cases, and preserve quality as the environment changes underneath them.

The core issue isn't model capability. It's that general models are built to be broadly useful, not tuned to the specific structure, constraints, and decision patterns that define how your business actually works.

Without a tight loop between execution, measurement, and training, every run is just another task completed, not another signal the system gets to learn from.

A unified RL platform for specialist models

Recursion connects environments, evaluation, and training into a closed-loop reinforcement learning system. Rather than treating deployment as the endpoint, production becomes the training surface.

RL environments that reflect real work

Recursion turns workflows, tools, policies, and edge cases into executable environments for RL training and evaluation. Powered by WorldSim, these environments recreate the full enterprise software stack — with configurable world effects that generate diverse, realistic scenarios at scale.

0554-crude-pipeline-lbo-waterfall

Context

You are an associate at a private equity fund preparing an investment committee update for a crude oil pipeline project acquired through an LBO. The task is to build an analyst-ready workbook that connects operating performance, the pre-computed debt schedule, exit valuation, and sponsor-management waterfall.

Workbook to build

  1. Financing Assumptions
  2. Operating Assumptions
  3. Debt Schedule
  4. Model
  5. Returns

Evaluation systems that measure real execution

The difference between agents that stagnate and agents that improve is measurement. Recursion builds evaluation systems that score intelligence and skill at every level — final outcomes, intermediate decisions, and execution quality — so every run generates the signal your models need to get better.

Model Accuracy Token Efficiency Latency
GLM 5.1 FTv2 86% 84% 1.1s
GPT 5.5 82% 38% 2.6s
Claude 4.8 79% 42% 2.9s
GLM 5.1 FTv1 78% 76% 1.7s
GLM 5.1 base 56% 54% 2.4s

A training loop that compounds from real work

Every rollout produces graded trajectories that can feed fine-tuning and reinforcement learning. The result is a specialist model that improves from enterprise execution signals instead of synthetic benchmarks alone.

Metric Value
Source 4,300 financial analysis tasks
Task family DCF, LBO, acquisition, projection
Base model GLM 5.1
Training method GRPO
Compute GKE, H100 cluster
Endpoints Baseline and tuned

From workflow to specialist model — continuously

Define

Identify the knowledge work tasks: the workflows, decisions, and domain-specific judgments agents need to perform.

Connect

Labelbox agents connect to enterprise data sources and convert them into structured RL data representations.

Evaluate

Design the evaluation system: rubrics, success criteria, and scoring logic that define what good performance looks like.

Generate

Create task distributions that capture long-tail edge cases and operational variability.

Train

Run RL training on graded rollouts to produce specialist models that improve with every cycle.

Serve

Recursion offers managed agents: pre-built, configurable agent harnesses that run in managed infrastructure for long-running tasks and asynchronous work.

Privacy and security for enterprise agents

Your advantage is the learning loop, not the model

As foundation models improve and become accessible to everyone, sustainable advantage shifts away from the model itself. The organizations that win won't simply deploy intelligence; they'll own the learning loops that transform their expertise, workflows, and decisions into durable competitive advantage.

Every organization has unique workflows, judgments, and domain expertise embedded in its people. The question is whether that expertise remains trapped inside individuals or becomes a compounding organizational asset.

Recursion is designed to close that gap.