Interdisciplinary Research into French News Media Ecosystems

Datasets, benchmarks and open models for studying the French press.

Non-profit association under French law (loi 1901) · Paris · RNA W751279600

Supported by

  • Project 4beta by Oxylabs
  • Amazon Web Services
  • Microsoftfor Nonprofits
  • Googlefor Nonprofits
  • ElevenLabs
  • GENCI
  • IDRIS · Jean Zay

Le French News Lab explores how French news shapes political life. Our work looks at the stories, frames, and language that structure public debate, combining computational methods with media and political analysis.

We develop datasets, methods, and public-facing research that make the French media ecosystem easier to study, question, and understand.

News

Grants, papers and data releases from the lab, newest first.

All news

Announcement

Le French News Lab wins access to France’s Jean Zay supercomputer

GENCI has awarded Le French News Lab computing time on Jean Zay, bringing France’s national supercomputing infrastructure to our research on French news.

Read the announcement

Two papers accepted for oral presentation at the International Conference on Natural Language and Speech Processing 2026

FrIdéo and Framing by Wording, Framing by Selection, both accepted for oral presentation.

Le French News Lab introduces FrenchNews-7

87,637 articles from 13 publishers, harmonised into 7 shared editorial desk categories.

What the lab has built

Scales, corpora and framing audits for the French press — each one documented enough that a disagreement about the media can become a disagreement about data.

Ideology scale

Thirty outlets, one continuous score

FrIdéo places French national news outlets on a single ideology scale by fusing nine independent evidence families under empirical Bayes shrinkage. Each placement decomposes into the evidence that produced it; leave-one-out ablation holds at ρ ≥ 0.992.

Accepted for oral presentation at the 9th International Conference on Natural Language and Speech Processing, Trento, September 2026.

30
Outlets
9
Evidence families
Read the paper

Headline framing

902,111 headlines, two dimensions of framing

To our knowledge the largest framing audit of the French press: wording and selection measured separately across 902,111 headlines from 25 outlets, 2022–2025. A released 10,000-headline supervision set carries per-item annotator agreement rather than a collapsed consensus.

Accepted for oral presentation at the 9th International Conference on Natural Language and Speech Processing, Trento, September 2026.

902,111
Headlines audited
25
Outlets
10,000
Supervision set
2
Dimensions
Read the paper

Role framing

LFI as conflict, RN as strategy

In 28,592 headlines from 25 outlets, La France insoumise is cast as a conflict actor and Rassemblement National as a strategic competitor. The asymmetry holds after outlet and year controls; at Le Journal du Dimanche it inverted after the 2023 editorial takeover.

28,592
Headlines
25
Outlets
2022–2025
Coverage window
Read the paper

Editorial desk classification

FrenchNews-7: a shared label space across publishers

87,637 articles from 13 publishers, mapped onto seven editorial desks. Where a URL slug is unambiguous the label is deterministic (72.2%); the ambiguous remainder is LLM-labelled and audited against a blinded human sample at κ = 0.806. The CamemBERT reference model reaches macro-F1 0.847 on the held-out split.

87,637
Articles
13
Publishers
7
Desks
0.847
Macro-F1
κ 0.806
Human audit
Read the paper

Our approach

We build the measurement infrastructure that makes the French news ecosystem possible to study: annotated corpora, reproducible benchmarks, open models, and peer-reviewed methods for questions that are usually argued about without evidence.

Measurement before argument

Claims about bias, framing and slant in the French press are usually made without a number anyone can check. We build scales and benchmarks that are documented, versioned and auditable, so that a disagreement about the media can become a disagreement about data.

Methods you can check

Papers are published with the datasets behind them. Results are reported as means across multiple random seeds rather than a single best run, because a number that only appears once is not a result.

Released in the open

Our datasets and models are published on Hugging Face, data.gouv.fr and Kaggle under open licences, free to use. Nothing on this site is behind a paywall or a signup form.

Read the research

Open data and models

Everything we build for a paper is released alongside it. These are the current public datasets and models — free to download, with data dictionaries and full metadata.

FrenchNews-7 corpus

A cross-publisher corpus of French news articles annotated by editorial desk, built so that topic classifiers can be tested on publishers they were never trained on — the case where reported accuracy usually falls apart. Released with the train and test splits used in the paper.

From the paper

CamemBERT FrenchNews-7 model

A CamemBERT-base classifier fine-tuned on FrenchNews-7 for editorial desk classification, published with its model card and the per-seed evaluation scores reported in the paper.

Available fromHugging Face
From the paper

FrIdéo ideology scores

Continuous ideology scores for thirty French national news outlets, fusing nine independent families of evidence rather than resting on any single one. Published as CSV and JSON with a data dictionary and full metadata, under CC BY 4.0.

Available fromdata.gouv.frKaggle
From the paper

Work with us

We collaborate with researchers, journalists, civil-society organisations and academic institutions working on the French information ecosystem. If you want to use our data, replicate a result, or propose a study, get in touch — we answer every message.

© 2026 LE FRENCH NEWS LAB. All rights reserved.

Non-profit association (loi 1901) · W751279600SIRET 108 676 511 00018Paris 20e, France