Building safer AI systems through interpretability research and alignment work
I'm a PhD candidate at Università degli Studi di Milano, working on AI safety, interpretability, and natural language processing. My research focuses on how models work internally and how they structure their knowledge — spanning manifold geometry, refusal behaviour, instruction-following mechanisms, training dynamics, and the encoding of aesthetic concepts.