Muence
BenchmarkDataResearchAccess

Access

The public set is the proof, not the product.

Every one of the 1,000 public tasks is browsable, downloadable and gradeable, which is exactly why it cannot be the thing you score a model on for a claim you intend to defend. Anything published is a candidate for the next pretraining corpus. These are not published.

Private holdout

Access upon request

A never-published evaluation set for scoring a model you intend to make claims about.

Labs and agent teams that need a score they can defend publicly.

What is included

RL environments

Access upon request

Executable environments with a program as the reward signal, for training rather than only measuring.

Post-training teams doing RL on code, who need a reward that is not a model opinion.

What is included

Custom environments

Access upon request

Environments authored against a specific weakness in a specific model, chosen from measurement rather than guesswork.

Teams that already know their model has a gap and want volume on exactly that gap.

What is included

Not sure which applies? Start with the public data and the research writeups. If the method holds up, the rest is a conversation.