Increasing the overlap between human and AI values.

We study how to scalably steer advanced AI systems to act on human values—so progress in capability is matched by progress in safety.

Our approach

Safety research should connect rigorous theory to the systems people actually build and use.

01

Scalability

02

Generality (architecture-agnosticism)

03

Low interpretability requirement

04

Low capabilities externalities

Work with us

Bring more perspectives into the overlap.

We welcome conversations with researchers, engineers, funders, and policy teams working toward safer advanced AI.