My research lies at the intersection of statistics, machine learning and natural language processing, with an emphasis on reliable inference, low-resource language technologies and responsible evaluation.
Statistical Inference for AI
- Statistically valid inference from model-generated evidence
- Calibration and uncertainty quantification
- Risk control
- Conformal prediction
- Semi-supervised and prediction-powered inference
Natural Language Processing & Legal AI
- Turkish and low-resource NLP
- Legal document understanding and classification
- Extractive and abstractive summarisation
- Dense and hybrid retrieval
- Retrieval augmented generation
- Benchmark construction and evaluation methodology
Responsible and Efficient AI
- Green AI and energy-aware evaluation
- Efficiency accuracy trade-offs in NLP systems
- Benchmarking transformer and non-transformer architectures
Statistical Modelling & Data Science
- Statistical modelling and inference
- Computational statistics
- Time series and state-space models
- Stochastic processes
- Resampling and bootstrap methods
- Survey and sampling design
- Experimental design
Crisis Data and Data Governance
I contributed to the quantitative and qualitative analysis of 146 survey responses in CODATA’s Data Policy Committee work on data policies for times of crisis. The work included close analysis of open-ended responses concerning training, standards, data governance and collaboration between researchers and practitioners.