Usage and Evaluation of LLM in Media and Organizations [pdf] (tech.ebu.ch)

🤖 AI Summary
The EBU’s Technical Report TR 083 (Jan 2025) surveys how public broadcasters are experimenting with Large Language Models (LLMs) and generative AI across newsroom and production workflows. Members including BBC, RAI, VRT, NHK, SRG and Radio France document pilots for content tagging, transcript chaptering, summarization, title generation, translation, subtitles and writing assistance. The report stresses practical significance: LLMs can automate routine editorial tasks and speed production, but their safe, scalable deployment depends on robust, mixed-method evaluation and careful platform selection. Technically, the report highlights evaluation challenges—automated metrics (ROUGE, BLEU, BERTScore, BLEURT, BARTScore) are useful but insufficient for assessing summaries, chapters or world knowledge; human-in-the-loop and task-specific benchmarks remain essential. It catalogues tools and models in use (GPT-4/4o, Gemini, Claude, Llama3.1, Mistral, Kosmos-2, multimodal LLaVA/Kosmos, plus toolkits like LangChain, DeepEval and LM Eval Harness) and recommends cost-effective strategies such as parameter-efficient fine-tuning (LoRA) and using smaller models for classification/tagging at scale. Annex material (VRT’s Smart News Assistant) exemplifies a pragmatic, privacy-aware approach to benchmarking and deployment. Overall the report argues for adaptive evaluation pipelines combining automated metrics, curated datasets and human review to manage quality, bias and scalability in media-grade LLM adoption.
Loading comments...
loading comments...