Project
OCPSG-Benchmarking project
Benchmarking LLMs and Fine-Tuned Models for Multilingual Policy Agenda Annotation in European Parliamentary Speeches
About the project
**Principal Investigador:** Bastián González-Bustamante **Research Team:** Tom Bellens, Christopher Klamm, and Marta Koch This project benchmarks zero-shot, few-shot and reasoning LLMs against fine-tuned Transformer baselines for multilingual policy agenda annotation using European parliamentary speeches as the core corpus. We will construct a reproducible evaluation pipeline comparing open and closed LLMs with fine-tuned BERT-family models across multiple European languages, quantifying trade-offs in performance, robustness, computational costs, and carbon footprint. Models will be assessed on a fixed, stratified test split with uniform prompts and controlled prompt variations. Decoding will be controlled through deterministic settings and contrasted with stochastic runs. Human-in-the-loop validation will be triggered by uncertainty and model disagreement to adjudicate complex cases and iteratively improve label quality. Outputs include a peer-reviewed article, fine-tuned model releases, and a working paper with practical guidelines.

Research outputs
Conference presentations
| Date | Event | Presentation | Location | Authors |
|---|---|---|---|---|
| 7 Aug 2026 | FAE-UDP KeynoteFaculty of Administration and Economics, Universidad Diego Portales | The Politics of Attention: Using AI to Map the Public Policy Agenda | Santiago, Chile | Bastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch |
| 6 Aug 2026 | CICS-UDD WorkshopResearch Centre for Social Complexity Workshop, Universidad del Desarrollo | From Classifier Consensus to Benchmark Data: An Uncertainty-Aware Silver Standard for Multilingual Policy Agenda Annotation | Santiago, Chile | Bastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch |
| 1 Aug 2026 | MPP-UDP KeynoteMaster in Public Policy, Universidad Diego Portales | The Politics of Attention: Using AI to Map the Public Policy Agenda | Santiago, Chile | Bastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch |
| 21–24 Jul 2026 | ALACIP 2026XIII Latin American Congress of Political Science | Benchmarking LLMs for Comparative Political Science: Lessons from Automated Legislative Agenda Annotation | Buenos Aires, Argentina | Bastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch |
| 20 May 2026 | RadiUnce MeetingRadiUnce Project Meeting, Utrecht University | From Benchmark to Lessons: Guidelines from Computational Analysis of Parliamentary Speech | Utrecht, Netherlands | Bastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch |
| 15–16 Apr 2026 | APSA Workshop 2026APSA Virtual Research Meeting 2026: Virtual Research Group on Aligning Computational Tools for the Political Science Research Lifecycle | Evaluating Large Language Models for Multilingual Policy Coding: A Transparent Benchmarking and Reporting Framework | Virtual | Christopher Klamm, Bastián González-Bustamante, Marta Koch, Tom Bellens |
| 27 Mar 2026 | Oxford Research CafeOCPSG Research Cafe Research Cafe, University of Oxford | Benchmarking LLMs and Fine-Tuned Models for Multilingual Policy Agenda Annotation | Virtual | Bastián González-Bustamante, Tom Bellens, Christopher Klamm, Marta Koch |
Keynote presentation
These are the project's research outputs in which I am involved; the project may have additional outputs.
Funding
DPIR-OGPSGThis project is part of the Oxford Computational Political Science Group (OCPSG).