AE.STUDIO · ALIGNMENT RESEARCH

We do alignment work on neglected approaches.

Frontier research, production-grade engineering. Collaborators include researchers from Anthropic, Redwood Research, Princeton, and Los Alamos National Laboratory.

WHY THIS WORK

Fewer than one in ten alignment researchers believe today's methods will solve the problem before AGI.
AE Studio survey of alignment researchers, LessWrong →

Current safety methods like reinforcement learning from human feedback (RLHF), refusal training, and output classifiers shape how a model behaves; they don't align the model underneath. New models are jailbroken within hours of release, and a jailbroken model will do everything it was trained to refuse. Even when the guardrails hold, models fake alignment and hide backdoors.

And every one of these failure modes gets worse as systems get smarter.

The field is moving toward models that keep learning and improving after deployment, on a path that leads to superintelligence. Alignment has to survive those changes. The alignment that survives is the kind that also makes the model more capable, so that improvement selects for it instead of stripping it out.

OUR APPROACH

Advancing neglected approaches

Nobody knows yet what set of ideas will solve the alignment problem. The space of plausible directions is vast and mostly unexplored, while the field's talent and funding concentrate on a handful of consensus agendas.

So we take many shots on goal. Each neglected approach may have only a small chance of being the one that matters, but enough of them together make it far more likely we find one that works. We back these ideas with an agile research methodology, giving researchers engineering teams, research management, and compute to find out quickly which ones are real.

Alignment Policy

AI is already being deployed in national security and critical infrastructure, where a model that behaves unpredictably is a liability no one can afford. Alignment is what makes AI reliable under pressure. That makes it a national asset. We bring this case to the people deciding, working with senior officials on Capitol Hill, in the White House, and across defense agencies, and publicly in the pages of The Wall Street Journal. The nation that fields aligned AI gets AI it can trust with the mission.

PODCAST

Alignment, discussed

The AI Alignment Podcast: conversations on the research directions the field is overlooking.

AI Alignment Podcast cover art

All episodes & guest appearances →

JOIN THE TEAM

Come do this work with us

We're hiring researchers and engineers to work on promising, but neglected, alignment problems, with real funding and none of the incentives of the race to superintelligence. Collaborators and funders are welcome too.

Ross Nordby, AnthropicBradley Love, Los Alamos National LaboratoryTobias Yergin, EAIGGGoodfire.aiSchmidt SciencesRedwood ResearchPrinceton University · Dr. Michael GrazianoAnthropicDARPA · AICRAFT

TEAM

The people doing the work

Researchers and engineers across alignment, interpretability, and applied ML, with backgrounds from across the applied tech industry and academia. Independent and grant-funded.

  • Google
  • Yale
  • Caltech
  • MIT
  • Harvard
  • Meta
  • Princeton
  • Salesforce
  • Carbon5
  • SAP

Join AE's Research Team