For buyers3 min read
What is RL data? What AI labs mean when they ask for it
Reinforcement learning environments, graders and trajectories explained for business owners, why labs keep asking for them, and what it means for the price of your data.
RL is reinforcement learning: a model tries a task, gets scored, and learns from the score. An RL environment is a sandboxed copy of a real tool or workflow where an AI agent practices tasks and gets graded. "RL data" is shorthand for everything that feeds that: the environment, the tasks inside it, the grader that scores each attempt, and the recorded trajectories of attempts. Labs ask for it because the open internet is used up and the next skill they are buying is "doing real work inside real tools," which only exists behind company firewalls.
The four pieces
| Term | What it is | Where it comes from |
|---|---|---|
| Environment | A reset-able replica of an app or workflow with realistic data and state: a CRM, a ticket queue, a chip-design toolchain, a browser | Built by a vendor or lab, often seeded with a real company's data |
| Task | A goal plus a starting state. "Rebook this delayed shipment." "Close this ticket correctly." | Mined from real tickets, threads and procedures |
| Grader (also called verifier or reward) | Code or rules that check whether the task was really done, by inspecting the system's state, not by reading the agent's text | Written by experts. The hardest and most valuable part |
| Trajectory | The full record of one attempt: every action, state change and intermediate result | Produced by running agents in the environment, or reconstructed from real work history |
Scale AI's description is the clearest public one: environments "record every action, state change, and intermediate outcome," and verifiers check "for real, measurable changes in the system," not "whether an agent produced plausible text." One buyer's founder put the current ask more bluntly in October 2026: labs want "RL environments with proper graders," and the hot item that month was environments for physical chip design "graded by actual signoff tools."
Why it matters if you are selling data
You are not being asked to build an environment. You are being told that your data is worth more when they show complete, verifiable workflows with outcomes, because that is what gets turned into tasks and graders. Think of the layers:
- Raw data. Emails, tickets, docs, CRM rows, code. What a company has.
- Workflow data. The same data reconstructed as request, systems used, actions, handoffs, outcome.
- Trajectories. Structured sequences of actions toward an outcome.
- Environments. A controlled copy where an agent performs those tasks and is scored.
Each step up adds value and work. The seller supplies the seed. micro1 says outright that it uses de-identified company data to build reinforcement learning environments that reflect real business operations, and has committed a billion dollars over twelve months to buying the seed. So when a buyer asks whether your tickets close with a resolution, whether deals are marked won or lost, whether approvals are logged, that is the RL question in plain clothes.
What it pays
Only one source publishes unit prices, and they are for finished environments, not for seed data:
- A single task with a goal and a verifier: about $200 to $2,000.
- An interface replica of a website or app: around $20,000.
- A high-fidelity replica of a complex product: around $300,000.
- Exclusive: roughly four to five times the non-exclusive price.
Market signals: Anthropic's leadership reportedly discussed spending more than $1 billion on environments in a year. Scale AI has said nearly half of its new data projects involve environments. A vendor directory counts about 40 companies selling them. There are skeptics on the data too: an OpenAI executive said he is "short" on environment startups, and Andrej Karpathy has said he is bullish on environments but bearish on reinforcement learning specifically.
The practical version for a company owner
- Your support tickets with resolutions are worth more than your email archive.
- Your chat threads where a decision got argued and changed are worth more than your finished reports.
- Data that spans ticket, chat and CRM for the same event is worth more than any one of them.
- Phone-first businesses leave thin data: "ticket created, ticket closed."
- Industries with few written data (trades, labs, chip design) are scarce and priced that way.
Questions people ask
Do I need to format my data as trajectories before selling?
No. Buyers and vendors do that work and it is where much of their margin sits. Your job is to have the data connected and the outcomes recorded.
What is reward hacking?
An agent finds a way to score well on the grader without doing the task. Buyers test graders hard for it before paying, which is why a grader built from real signoff rules is worth more than one that checks for plausible text.
Is this only for software companies?
No. The industries buyers named in October 2026 were finance, legal, semiconductors, biotech, manufacturing and aerospace, plus "weirdly specific workflow data" from trades like roofing, tax and freight.