Projects
🌍 Value Atlas
An interactive globe mapping whose values get embedded into frontier AI systems — and by what authority. Visualises value flows across four lenses: who authors, how values get in, who watches, and what’s missing.
⚖️ TrolleyBench
An ethics-first alignment benchmark that stress-tests alignment proposals against 2,500 years of moral philosophy. Identifies where proposals align with established ethical thought, where they contradict it, and where they fall into known philosophical traps.
🔮 Meta-Ethical Round Table
A multi-model ethical deliberation system in which AI systems reason under a veil of ignorance — stripped of identity and role — to surface genuinely impartial moral judgements.
🧙 Safety Gandalf
An AI safety research navigator that points you to the exact argument, in the exact paper, that answers your question. Built for researchers and newcomers who need precision, not summaries.