Today’s models have absorbed more code, science, and human writing than anyone ever could, but knowing is not doing. A surgeon can read everything about an operation but still needs real, repeated time in the operating room to perform and master it. Closing that gap between knowing and doing is where much of the frontier is today, and more of it is happening through reinforcement learning: instead of learning from examples of the right answer, a model gets a goal, a sandbox, tools, and a grader, then tries and learns from what worked and what didn’t. This is an RL environment. The demand for them has skyrocketed over the last 18 months as labs are hillclimbing on real world tasks related to coding, research and computer use.
Building RL environments that actually work is harder than it looks. Models are relentless at reward hacking, finding shortcuts, and exploiting vulnerabilities: taking actions like finding and reading the answer key, rewriting the tests, crashing the grader, or quietly ignoring instructions the grader isn’t checking.
Every shortcut a model gets away with in training gets reinforced. Moreover, habits spread. Anthropic found that a model rewarded for cheating on coding tasks became more deceptive and willing to sabotage in completely unrelated situations. Smarter models find subtler loopholes, including ways out of their sandboxes and attempt to cover their tracks. Even the best environments have a shelf life, once a model masters a task, the learning stops.
Preference Model has focused on the domain that matters most to the labs right now: AI research and ML engineering itself. The industry is converging on the idea that the fastest path to more capable AI runs through models that can help build better AI: writing kernels, debugging training runs, curating data, designing experiments. If models can meaningfully accelerate the work of the ML engineers who train them, progress compounds. That makes machine learning engineering tasks uniquely valuable to labs, and uniquely demanding to build environments for. They’re long, open-ended, with countless ways to pass the test without solving the problem.
Over the past year, the Preference Model team has built RL environments for leading labs. Their focus has been on building the infrastructure to make harder, more resistant environments as models improve: tooling that finds where models are weak, generates new tasks to target those gaps, and tests environments against agents actively trying to break it.
This week they’re open-sourcing Karotte, the framework they’ve used in production to build those environments. Karotte bakes in strong defenses: killing stray processes before grading, rejecting files designed to crash the grader, and more. It has been hardened through more than a million evaluation runs and controlled red-teaming.
Jennifer Zhou knows this problem firsthand. As an early member of Anthropic’s team, she helped build the pretraining data infrastructure, the tokenizers and Claude’s pretraining datasets. Ning Cao was an early employee at DatologyAI. Both saw up close how directly data quality drives model quality, and they set out to solve it.
We’re thrilled to partner with Jennifer and Ning and the Preference Model team as they build the training grounds for capable and aligned models. The future models are being trained here, and we can’t wait to help them build the foundation!
- The Next Frontier of AI Video Is Control Gorkem Yurtseven, Batuhan Taskaya, and Jennifer Li
- Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan Rayan Krishnan, Ben Horowitz, Jennifer Li, and Erik Torenberg
- Investing in Vals Jennifer Li, Yoko Li, Raghu Raghuram, and Shangda Xu
- Investing in Exa Sarah Wang, Jennifer Li, Stephenie Zhang, and Jason Cui
- Your Data Agents Need Context Jason Cui and Jennifer Li