Project

OCPSG-Benchmarking project

Benchmarking LLMs and Fine-Tuned Models for Multilingual Policy Agenda Annotation in European Parliamentary Speeches

About the project

**Principal Investigador:** Bastián González-Bustamante **Research Team:** Tom Bellens, Christopher Klamm, and Marta Koch This project benchmarks zero-shot, few-shot and reasoning LLMs against fine-tuned Transformer baselines for multilingual policy agenda annotation using European parliamentary speeches as the core corpus. We will construct a reproducible evaluation pipeline comparing open and closed LLMs with fine-tuned BERT-family models across multiple European languages, quantifying trade-offs in performance, robustness, computational costs, and carbon footprint. Models will be assessed on a fixed, stratified test split with uniform prompts and controlled prompt variations. Decoding will be controlled through deterministic settings and contrasted with stochastic runs. Human-in-the-loop validation will be triggered by uncertainty and model disagreement to adjudicate complex cases and iteratively improve label quality. Outputs include a peer-reviewed article, fine-tuned model releases, and a working paper with practical guidelines.

OCPSG-Benchmarking project project visual

Research outputs

Conference presentations

Conference presentations
DateEventPresentationLocationAuthors
7 Aug 2026FAE-UDP KeynoteFaculty of Administration and Economics, Universidad Diego PortalesThe Politics of Attention: Using AI to Map the Public Policy AgendaSantiago, ChileBastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch
6 Aug 2026CICS-UDD WorkshopResearch Centre for Social Complexity Workshop, Universidad del DesarrolloFrom Classifier Consensus to Benchmark Data: An Uncertainty-Aware Silver Standard for Multilingual Policy Agenda AnnotationSantiago, ChileBastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch
1 Aug 2026MPP-UDP KeynoteMaster in Public Policy, Universidad Diego PortalesThe Politics of Attention: Using AI to Map the Public Policy AgendaSantiago, ChileBastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch
21–24 Jul 2026ALACIP 2026XIII Latin American Congress of Political ScienceBenchmarking LLMs for Comparative Political Science: Lessons from Automated Legislative Agenda AnnotationBuenos Aires, ArgentinaBastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch
20 May 2026RadiUnce MeetingRadiUnce Project Meeting, Utrecht UniversityFrom Benchmark to Lessons: Guidelines from Computational Analysis of Parliamentary SpeechUtrecht, NetherlandsBastián González-Bustamante, Christopher Klamm, Tom Bellens, Marta Koch
15–16 Apr 2026APSA Workshop 2026APSA Virtual Research Meeting 2026: Virtual Research Group on Aligning Computational Tools for the Political Science Research LifecycleEvaluating Large Language Models for Multilingual Policy Coding: A Transparent Benchmarking and Reporting FrameworkVirtualChristopher Klamm, Bastián González-Bustamante, Marta Koch, Tom Bellens
27 Mar 2026Oxford Research CafeOCPSG Research Cafe Research Cafe, University of OxfordBenchmarking LLMs and Fine-Tuned Models for Multilingual Policy Agenda AnnotationVirtualBastián González-Bustamante, Tom Bellens, Christopher Klamm, Marta Koch

Keynote presentation

These are the project's research outputs in which I am involved; the project may have additional outputs.

Funding

DPIR-OGPSG logoDPIR-OGPSG

This project is part of the Oxford Computational Political Science Group (OCPSG).

Project website