Burcu Sayin Günel
Trento, Italy
applied AI · research × production

Burcu Sayin Günel

Applied AI Scientist & Engineer

I build AI systems people can rely on — turning research on large language models, evaluation, and human-AI decision-making into products that ship. PhD cum laude from the University of Trento; today I lead AI research and engineering at an early-stage startup while continuing my scientific work with clinical and academic collaborators.

LLMs Generative AI AI Agents Evaluation & Statistics Human-AI Decision-Making Production ML
290+
Citations
10
h-index
26
Publications
2×
Best-Paper Awards
About

Research depth, shipped at industry pace

I work at the intersection of large language models and human decision-making — building AI systems reliable enough to trust, and understanding precisely when they shouldn't be.

I'm an applied AI scientist and engineer. My PhD at the University of Trento was on machine learning for human-AI decision-making: when to trust a model, when to reject its answer, and how to measure what a model is really worth. As a postdoctoral researcher with the Structured Machine Learning Group I brought that lens to large language models in medicine and law. Today I lead AI research and engineering at a stealth AI startup, where I spend most of my time designing, building and running LLM products in production and the rest leading a small team of engineers. Alongside that, I continue my scientific work independently with clinical and academic collaborators.

Experience

Where I've worked

Mar 2026 — Present

Lead AI Engineer & Head of AI Research

Stealth AI Startup · Italy

I lead AI research and engineering at an early-stage startup: designing, building and running LLM products in production, and leading a small team of engineers. I combine research rigour (evaluation, statistics, guardrails) with production engineering on Python, AWS and Kubernetes.

Mar 2026 — Present

Independent Researcher

Self-directed research · Trento, Italy

I continue my scientific work alongside my industry role, with clinical and academic collaborators. Current topics: agentic AI, AI for healthcare, and the evaluation of multimodal LLMs.

Nov 2022 — Feb 2026

Postdoctoral Researcher

Structured Machine Learning Group · University of Trento

Researched hybrid human-LLM reasoning and decision-making within the EU Horizon TANGO project: diagnostic dialogue systems evaluated with hospital physicians, reliable retrieval for large legal corpora, and how to evaluate LLMs in clinical settings. Supervised MSc and PhD students.

Nov 2018 — Sep 2022

Ph.D. in Information & Communication Technology (cum laude)

University of Trento · research internship at TU Delft (2021)

Thesis: “Towards Reliable Hybrid Human-Machine Classifiers.” Value-aware evaluation metrics and active-learning methods that account for the real-world cost of errors, so that model choice reflects application value, not accuracy alone.

Jan 2016 — Sep 2018

Research & Teaching Assistant

Izmir Institute of Technology · Dept. of Computer Engineering

Research on social network analysis and privacy. Taught and supported courses including Operating Systems, Data Structures, Data Mining, Probability & Statistics, and Cryptography.

Research

Selected publications

arXiv · 2026 · under review

Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care

A controlled study with hospital physicians: dialogue with an LLM improves diagnostic accuracy, with the largest gains for junior doctors on hard cases.

Dentomaxillofacial Radiology · 2026

Multimodal Large Language Models and Dental Students in Radiographic Landmark Identification: A Novel Grid-Based Assessment

Benchmarking frontier multimodal LLMs against dental students on radiograph landmark localisation.

NLLP @ EMNLP · 2025

Towards Reliable Retrieval in RAG Systems for Large Legal Datasets

A failure mode of RAG on large legal corpora — retrieving from the wrong document — and a chunking method that substantially reduces it.

Artificial Intelligence Review · 2025

Rethinking and Recomputing the Value of Machine Learning Models

Value-aware evaluation: choosing models by what their predictions are worth in the application, not by accuracy alone.

Clinical NLP @ NAACL · 2024

Can LLMs Correct Physicians, Yet? Investigating Effective Interaction Methods in the Medical Domain

When and how LLMs can usefully challenge expert decisions — designing safe human-AI interaction for high-stakes medicine.

AAAI HCOMP · 2021 · Best Paper

The Science of Rejection: A Research Area for Human Computation

Why learning to reject model predictions is central to ML — and what role humans play. Blue Sky Best Paper, AAAI HCOMP 2021.

Looking for the full list? See my Google Scholar profile.
Recognition

Awards

🏆

Blue Sky Best Paper Award

AAAI HCOMP 2021

For “The Science of Rejection: A Research Area for Human Computation,” at the 9th AAAI Conference on Human Computation and Crowdsourcing.

🎖️

Methods Recognition

ACM CSCW 2021

For “On the State of Reporting in Crowdsourcing Experiments and a Checklist to Aid Current Practices.”

🥇

Best Paper Award

Data Analytics 2017

For “A Novel Approach to Information Spreading Models for Social Networks,” at the 6th International Conference on Data Analytics.

Contact

Let's build something together

Open to conversations about applied AI, research collaborations, and advisory roles.