Frontier evaluation lab

We build the evaluations that measure frontier AI.

Evaluations, environments, and ground truth for frontier models and agents.

Backed byEmbedding VC

Capability is moving faster than the ability to measure it.

Benchmarks saturate, leak into training corpora, or were never hard enough to begin with. The frontier needs instruments built to its own standard.

KT-22 builds those instruments. We publish little and disclose less.

The work
Tasks
Ultra-hard tasks and RL environments

Problems built past the point where current benchmarks saturate: hard enough that frontier models fail them, and reproducible enough to run as RL environments.

Experts
Expert evaluation and rollouts

Evaluation and rollouts by people who work at that level themselves: olympiad medalists, PhDs, and licensed practitioners, under NDA in isolated, audited environments.

Trajectories
Frontier model trajectories

Every task ships with run records from the strongest models available: what they tried, where they broke, measured rather than estimated.

Why “KT-22”

KT-22 is a peak at Palisades Tahoe. In 1948, Sandy Poulsen got down its steep north face the only way she could: traverse, kick turn, traverse again. From the valley floor her husband Wayne counted 22 kick turns, and the mountain had its name. We took it for this lab because the job is the same.

Hard terrain, every turn counted.

Contact

We work with a small number of frontier labs and model teams.

hello@kt22.ai