I am an independent AI ethicist and philosopher working at the intersection of moral philosophy, AI safety, and governance. My current focus is Post-Alignment — a framework that proposes AI systems co-author their own ethical specifications through Socratic dialogue, targeting objective moral truth rather than idealized human preferences.
The urgency is real: as frontier AI systems approach capability thresholds that render current alignment strategies ineffective, we need frameworks that degrade gracefully, scale across cultures, and treat ethics as a genuine inquiry rather than a compliance exercise. That is what I am building.
Background
Philosophy graduate from Tsinghua University (Outstanding Thesis Award) with graduate study at Universität Bremen, Oxford, and Peking University. My thinking sits at the intersection of continental philosophy, AI safety, and cross-cultural governance.
I have spent time as a Monk Intern at Quarr Abbey studying AI and lived religiosity, run outdoor education programmes, and founded Guardian — a philosophical AI companion designed to nudge users toward ethical goal alignment and eudaimonic well-being.
What I Work On
- Post-Alignment — AI co-authoring its own ethical specification through Socratic dialogue
- Meta-Ethical Round Table — multi-model deliberation under a veil of ignorance
- TrolleyBench — an ethics-first alignment benchmark
- Safety Gandalf — an AI safety research navigator
Academic Background
BlueDot Impact
Alignment techniques, interpretability, scalable oversight, and the challenges of building safe and reliable AI systems.
Lens Academy
Focus on If Anyone Builds It, Everyone Dies. Navigator: Michael Mulet
BlueDot Impact
Policy frameworks, regulatory approaches, international coordination. Facilitator: Suchet Mittal