Conceptual

Benchmarking LLMs on Journalistic Sourcing Annotation

A benchmark scenario for evaluating how well large language models identify and annotate sourcing in news stories: a five-category schema (adapted from journalism studies) covering sourced statements, source names and types, and source justifications; a hand-built ground-truth dataset of annotated articles; and an accuracy-scoring method that matches model output to ground truth via text embeddings and cosine-similarity thresholds. Findings show current LLMs struggle most at recovering every sourced statement and at spotting source justifications.