Applied AI Scientist & Engineer
I build AI systems people can rely on — turning research on large language models, evaluation, and human-AI decision-making into products that ship. PhD cum laude from the University of Trento; today I lead AI research and engineering at an early-stage startup while continuing my scientific work with clinical and academic collaborators.
I work at the intersection of large language models and human decision-making — building AI systems reliable enough to trust, and understanding precisely when they shouldn't be.
I'm an applied AI scientist and engineer. My PhD at the University of Trento was on machine learning for human-AI decision-making: when to trust a model, when to reject its answer, and how to measure what a model is really worth. As a postdoctoral researcher with the Structured Machine Learning Group I brought that lens to large language models in medicine and law. Today I lead AI research and engineering at a stealth AI startup, where I spend most of my time designing, building and running LLM products in production and the rest leading a small team of engineers. Alongside that, I continue my scientific work independently with clinical and academic collaborators.
I lead AI research and engineering at an early-stage startup: designing, building and running LLM products in production, and leading a small team of engineers. I combine research rigour (evaluation, statistics, guardrails) with production engineering on Python, AWS and Kubernetes.
I continue my scientific work alongside my industry role, with clinical and academic collaborators. Current topics: agentic AI, AI for healthcare, and the evaluation of multimodal LLMs.
Researched hybrid human-LLM reasoning and decision-making within the EU Horizon TANGO project: diagnostic dialogue systems evaluated with hospital physicians, reliable retrieval for large legal corpora, and how to evaluate LLMs in clinical settings. Supervised MSc and PhD students.
Thesis: “Towards Reliable Hybrid Human-Machine Classifiers.” Value-aware evaluation metrics and active-learning methods that account for the real-world cost of errors, so that model choice reflects application value, not accuracy alone.
Research on social network analysis and privacy. Taught and supported courses including Operating Systems, Data Structures, Data Mining, Probability & Statistics, and Cryptography.
A controlled study with hospital physicians: dialogue with an LLM improves diagnostic accuracy, with the largest gains for junior doctors on hard cases.
Benchmarking frontier multimodal LLMs against dental students on radiograph landmark localisation.
A failure mode of RAG on large legal corpora — retrieving from the wrong document — and a chunking method that substantially reduces it.
Value-aware evaluation: choosing models by what their predictions are worth in the application, not by accuracy alone.
When and how LLMs can usefully challenge expert decisions — designing safe human-AI interaction for high-stakes medicine.
Why learning to reject model predictions is central to ML — and what role humans play. Blue Sky Best Paper, AAAI HCOMP 2021.
For “The Science of Rejection: A Research Area for Human Computation,” at the 9th AAAI Conference on Human Computation and Crowdsourcing.
For “On the State of Reporting in Crowdsourcing Experiments and a Checklist to Aid Current Practices.”
For “A Novel Approach to Information Spreading Models for Social Networks,” at the 6th International Conference on Data Analytics.
Open to conversations about applied AI, research collaborations, and advisory roles.