Access
The public set is the proof, not the product.
Every one of the 1,000 public tasks is browsable, downloadable and gradeable, which is exactly why it cannot be the thing you score a model on for a claim you intend to defend. Anything published is a candidate for the next pretraining corpus. These are not published.
Private holdout
Access upon requestA never-published evaluation set for scoring a model you intend to make claims about.
Labs and agent teams that need a score they can defend publicly.
What is includedRL environments
Access upon requestExecutable environments with a program as the reward signal, for training rather than only measuring.
Post-training teams doing RL on code, who need a reward that is not a model opinion.
What is includedCustom environments
Access upon requestEnvironments authored against a specific weakness in a specific model, chosen from measurement rather than guesswork.
Teams that already know their model has a gap and want volume on exactly that gap.
What is includedNot sure which applies? Start with the public data and the research writeups. If the method holds up, the rest is a conversation.