Language

Dark Mode
Accepted paper · ICNLSP 2026·Peer-reviewed

Framing by wording, framing by selection

An outlet frames a story twice: once when it decides to run it, and again when it decides how to word it. This audit measures the two separately across 902,111 French headlines from 25 outlets, 2022–2025.

By Amr Sobhy

Oral presentation · Trento, September 2026

Most attempts to score a newsroom produce one number, which mixes those two decisions together. Separating them changes the picture for about half this panel — and shows that French headlines got measurably hotter over four years without the mix of stories behind them changing much.

Year
2026
Venue
ICNLSP 2026
Oral presentation · Trento, September 2026
Research area
Media framing, Political communication, French news media, Computational journalism
Read the paperThe mapData and code
902,111
headlines, 25 French national outlets, 2022–2025
2
dimensions measured separately: wording and story selection
31.8% → 37.4%
headlines carrying a wording device, 2022 to 2025, story mix held constant
40%
of the difference between outlets on one dimension the other does not explain

Two decisions, not one

Every story involves two choices. Whether to run it — an outlet that leads on crime three days in five has framed the country before writing a single adjective. And how to word it: « Les migrants envahissent la Manche » and « Des migrants traversent la Manche » report the same crossing.

Both are framing. Both usually get collapsed into one score, which cannot tell them apart. An outlet with sober language on a concentrated crime agenda and one with a broad agenda and inflammatory language land in the same place from opposite directions.

Le Point and Le Parisien word their headlines identically hard — a device appears in 40.2% and 40.3%. They are nowhere near each other on the other axis: 43.8% of Le Parisien’s headlines are crime, conflict, scandal or crisis stories, against 29.3% of Le Point’s.
The mirror case: Les Echos words 24.2% of its headlines against BFMTV’s 29.9%, so BFMTV words slightly harder. But 40.8% of BFMTV’s headlines are high-charge story types against 10.4% of Les Echos’s — four times as many.
0%18%35%53%71%89%0%16%31%47%63%78%Headlines carrying a wording device (%)Headlines that are high-charge story forms (%)Le ParisienOuest-FranceBFMTVLes EchosLe PointCauseur
Distinctive on bothDistinctive wordingDistinctive agendaClose to the panel averagePoint size: headlines behind the estimate

Each dot is one outlet. Horizontal: share of headlines carrying a wording device. Vertical: share that are high-charge story forms. The diagonal is where the two match.

Across all 25 outlets the two rates are correlated: an outlet that words hard tends to select hard. But the gap between the two runs from −11 points at BFMTV to +29 at Causeur. One number per outlet means landing somewhere in that range without saying where.

What gets counted

The wording measure looks for three devices, each defined so a neutral rewrite of the same event would drop it. The selection measure classifies what a headline reports happening, into ten categories.

Loaded vocabulary

Charged wording whose neutral paraphrase keeps the event intact

« Covid : comment le vaccin d’AstraZeneca est devenu le mal aimé »

Les Echos

Blame attribution

Wording that assigns responsibility for a bad outcome to a named actor

« Martinique : un cadre de Groupama sera jugé pour insultes racistes »

Le Figaro

Threat framing

Wording that casts an event as danger, invasion, crisis or security risk

« Les stations de ski des Pyrénées sont en danger, pour la Cour des comptes »

Le HuffPost

None of these three headlines is about a contested political issue. These are wording patterns, not a politics detector.

Two supplementary fields

Interrogative headlines. The released data calls this field rhetorical_question, but the annotation guide instructs coders to mark any headline containing a question mark, explicitly telling them not to judge whether the question is rhetorical. It is close to a punctuation count. A high rate means an outlet writes questions, not that it insinuates.

Us-versus-them. Explicit ingroup/outgroup contrast. The raters agree on it least of the five fields, and it falls below the precision floor set for it in advance.

Both appear in the released data and in outlet profiles below. No finding on this page rests on either.

Story form

Selection is measured by what a headline reports happening, not what it is about: POLICY, PARLIAMENT, CONFLICT, CRIME, SCANDAL, ELITE, SOCIAL, LAW, ELECTIONS, OTHER.

« Immigration : vers plus de contrôle des mariages des personnes étrangères en situation irrégulière » is POLICY. « Naufrage de migrants dans la Manche : quatre personnes mises en examen » is CRIME. Same subject, different journalistic action — the difference a topic model misses.

A headline is measured on its own, which is how most readers meet it. That also means the measurement cannot see context that would change the reading. Quotation, irony, and blame voiced by a source rather than the outlet are often unrecoverable. Charged language inside quotation marks counts, because from the headline it is not decidable whose language it is.

Where the labels come from

A 10,000-headline training set was labelled by three language models from different families, with each label set by two-out-of-three agreement. Where all three disagreed on story form — 642 headlines — a human decided. The models are not treated as ground truth: they share pretraining, so they can share a mistake, and majority vote would hide it.

The check is external. Two annotators with no connection to the project re-labelled a held-out sample of 499 headlines from a written guide, without seeing any model output.

On 232 of those 499 headlines — 46.5% — the two annotators disagreed about at least one device. That is not a failure of the annotators; it is the size of the judgement being asked for. Both read « Violences urbaines : 243 établissements scolaires dégradés » as threat framing, and split on whether violences urbaines is itself loaded. Both read « Le Pen revendique un parti "professionnalisé" » as a factual report; one heard the quotation marks as ironic and one did not.
FieldBetween the two humans (κ)Humans vs. models (κ)
Threat framing.613.668
Blame attribution.601.718
Loaded vocabulary.542.608
Interrogativesupplementary.894.869
Us-vs-themsupplementary.510.596

The classifier is set to catch devices rather than be sure about them: recall is high, precision lower. Every rate on this page is therefore an upper bound. Two fields — loaded vocabulary and us-vs-them — fall below the precision floors set for them before results were seen.

Two of the three models were released during or after the period they labelled. A period-stratified comparison found no sign of contamination: agreement between models is flat or slightly lower in the later period, the opposite of what contamination produces. That test cannot clear specific event clusters — Israel/Gaza from October 2023, the June 2024 elections. Anyone needing contamination-free labels there should use the human-annotated 499-headline sample, released in full.

Four years

The audit covers 2022 to 2025, and the four years are not the same.

In 2022, 31.8% of headlines in this corpus carried at least one wording device. In 2025, 37.4% did, holding the mix of outlets, sections and stories constant. That is a rise of about a fifth across 902,111 headlines.

The obvious explanation is that the news got worse — more crime, more conflict, more of the story types that attract charged language. Because the two dimensions are measured separately, that can be tested directly, and it accounts for little of it.

Before holding the mix constant the raw rise is 6.3 points, and it splits: about 1.2 points come from the mix of stories changing and about 4.9 from the same kinds of stories being worded more strongly. Crime coverage did grow, from 12.5% of headlines to 15.7%. But the larger movement is on the other axis, and it shows up inside almost every story category.

0%10%20%30%40%2022202320242025

2022: 31.8% → 2025: 37.4% standardised. Of the 6.3-point raw rise, 4.9 points come from wording and 1.2 from the mix of stories.

Monthly share of headlines carrying a wording device, and the share that are high-charge story forms. A steady climb, not a spike: removing the October–December 2023 and June–July 2024 windows changes nothing.

Twenty-three of the twenty-five outlets rose. It is not a story about the fringe: Fdesouche, already the most charged outlet at 71.5%, moved 1.3 points because there was little room. The large moves are in the middle of the panel — Le Figaro +8.9 points, Le Parisien +10.3, CNews +11.9, Le Point +12.5.

The same classifier scores all four years, so it cannot drift. It also cannot tell the difference between newsrooms changing their house style and newsrooms covering four years that warranted stronger language. Holding the story mix constant controls for the category of an event, not its intensity — a war and a summit are both CONFLICT.

Every rate is an upper bound, so the level is less trustworthy than the change; both years are estimated the same way, which is what makes the comparison meaningful. Four years is a short series, and this is a headline audit — whether article bodies moved the same way is not something it can see.

The map

Twenty-five outlets, placed by how far their wording departs from the corpus average and how far their story mix does.

The centre of this map is the French press average, not neutral journalism. Both axes measure distance from what these 25 outlets did on average, in three sections, between 2022 and 2025. An outlet near the origin is typical of this panel — not unframed. An outlet far from it is unusual relative to these 25 — not biased. Change the panel and every position moves.
-0.0180.0190.0550.0920.1290.165-0.0060.0070.0190.0320.0440.057Wording divergence from the corpus averageStory-selection divergence from the corpus averageLe FigaroOuest-FranceTF1 INFOJDDFdesoucheMarianneSlate.frCauseur
Distinctive on bothDistinctive wordingDistinctive agendaClose to the panel averagePoint size: headlines behind the estimate

Point size is the number of headlines behind the estimate. Quadrants are median splits and a reading aid; the boundaries are soft and several outlets sit near them. The two axes are computed over different numbers of categories and are not comparable in magnitude — read them as rankings within an axis.

Four outlets that make the point

Most outlets sit near the diagonal. The ones that do not are why measuring both is worth the trouble.

Marianne: charged language, ordinary rhetoric

Marianne carries charged wording more often than almost any outlet — loaded vocabulary in 41.9% of headlines against a corpus 17.0%, some device in 61.2%. Its wording divergence is nevertheless below the panel median. The measure compares an outlet’s mix of devices to the corpus mix, and Marianne’s mix is ordinary: it does what the French press does, more often. Volume of charged language and a distinctive rhetorical signature are different properties.

Fdesouche: distinctive on both

The most extreme outlet on both axes. 72.6% of headlines carry a device and 82.1% are high-charge; blame attribution appears in 54.0% against a corpus 13.6%, threat framing in 31.8% against 8.5%. Its selection is equally concentrated: 58% of its headlines are CRIME, against 14.4% across the corpus. The two operations reinforce each other rather than substitute. Every finding here was recomputed without it; the correlation between the axes falls from 0.74 to 0.69 and nothing else changes.

Ouest-France: the low end, and what produces it

A device appears in 7.2% of its headlines against a corpus 34.6%, and 4.6% are high-charge — both the lowest on the panel by a wide margin. This is where the measurement needs explaining rather than celebrating. Ouest-France publishes at high volume in a compressed regional-brief register, and 40% of its headlines fall into OTHER, the residual category. Some of the gap is editorial style; some is a short headline format meeting a classifier trained mostly on national-desk headlines. Read it as “these headlines look unlike the rest of the panel”, not as a finding about restraint.

Slate.fr and Causeur: charged wording, uncharged agenda

Both are distinctive on wording and run mild agendas relative to how they write. Causeur uses evaluative vocabulary in 45.6% of headlines against a corpus 17.0% — the highest on the panel — while 30.4% of its headlines are high-charge. Slate.fr words 49.6% and selects high-charge in 23.1%; its blame-attribution rate, 7.3%, is about half the corpus average. Strong wording on a mild agenda is the opposite of Fdesouche, and invisible to any single score. Both are small: 3,056 and 3,907 headlines.

Framing intensity and political position

These 25 outlets also appear on this lab’s ideology scale, which places French outlets on a left–right axis from independent evidence. Putting the two together answers a question this page will otherwise be asked: does one side frame harder?

It does not. An outlet’s position on the left–right axis explains almost none of how often its headlines carry a wording device — the correlation is −0.03. Its distance from the centre explains a good deal, at 0.67. The outlets that frame hardest sit at both ends: L’Humanité, Mediapart and Blast on one side, Fdesouche, Causeur and Valeurs actuelles on the other. Those closest to the centre of the ideology scale are among the least distinctive on this one.

Two measurements, 25 outlets, a descriptive correlation — not a claim about cause. The two scales share one kind of input: the ideology scale reads article vocabulary and this audit reads headline wording. Removing that shared component from the ideology scale leaves the result unchanged, at 0.67.

-0.03
correlation with left–right position
+0.67
correlation with distance from the centre
0.015.731.447.162.878.5-2.9-1.9-0.80.31.32.4Ideology score (left ← → right)Headlines carrying a wording device (%)Le FigaroOuest-FranceL'HumanitéFdesoucheValeurs actuellesMediapart
Distinctive on bothDistinctive wordingDistinctive agendaClose to the panel averagePoint size: headlines behind the estimate

Horizontal: ideology score. Vertical: share of headlines carrying a wording device.

Outlet profiles

Every outlet on the panel, with the rates behind its position on the map.

Per-outlet framing rates, 25 French national outlets, 2022 to 2025
Dominant story form
Corpus average902 11134.631.317.013.68.55.65.3
Le Figaro108 30733.231.216.214.08.44.04.9OTHER 25%Close to the panel average
Le Parisien93 87740.343.820.219.08.94.64.5CRIME 25%Close to the panel average
Ouest-France90 4237.24.63.91.21.51.60.6OTHER 40%Distinctive on both
Franceinfo83 57431.728.413.810.48.96.14.7SOCIAL 25%Close to the panel average
20 Minutes62 83239.142.612.819.58.97.93.7CRIME 27%Distinctive wording
BFMTV57 61729.940.810.116.07.72.93.8CRIME 29%Distinctive wording
Les Echos55 69424.210.415.93.45.94.21.9OTHER 32%Distinctive on both
TF1 INFO44 01438.329.615.98.510.013.72.7SOCIAL 29%Distinctive wording
Le Monde39 01737.228.320.013.39.13.37.1SOCIAL 24%Close to the panel average
La Croix35 98029.323.014.48.57.34.74.4OTHER 23%Close to the panel average
Libération33 92145.734.526.716.18.87.29.3SOCIAL 20%Close to the panel average
Le Point32 52240.229.323.510.98.78.95.4SOCIAL 18%Close to the panel average
CNews27 25238.543.116.319.212.04.77.5CRIME 25%Close to the panel average
L'Humanité17 91149.933.230.917.29.28.915.2SOCIAL 26%Distinctive agenda
Le HuffPost17 73737.431.216.417.26.77.25.6POLICY 16%Close to the panel average
JDD17 59841.935.924.218.011.05.89.8POLICY 17%Distinctive agenda
Le Nouvel Obs17 51742.234.521.815.48.68.17.7SOCIAL 22%Close to the panel average
Fdesouche16 12472.682.133.354.031.81.421.4CRIME 58%Distinctive on both
Valeurs actuelles14 28652.054.625.928.415.43.511.2CRIME 29%Distinctive on both
L'Express11 79545.127.529.19.911.310.86.2SOCIAL 23%Distinctive on both
Marianne9 41261.241.941.920.810.614.214.3SOCIAL 19%Distinctive agenda
Mediapart7 09258.848.737.326.910.53.614.5SOCIAL 18%Distinctive on both
Slate.fr3 90749.623.123.37.310.922.84.1SOCIAL 48%Distinctive on both
Causeur3 05659.130.445.69.310.414.811.7ELITE 35%Distinctive on both
Blastcase profile64667.357.956.231.717.02.621.8SCANDAL 21%Distinctive on both

* Supplementary fields. The interrogative field is close to a punctuation count and us-vs-them falls below its precision floor; no finding on this page rests on either. Rates are classifier estimates over headlines only, from the Politics, Economy and Society sections, 2022–2025, and are upper bounds. They describe an outlet’s headline output relative to the other 24. They say nothing about the accuracy, rigour or good faith of its journalism.

How to misread this

Six sentences this audit does not support.

Le Parisien uses more framing than Le Monde.

43.8% of Le Parisien’s headlines are high-charge story forms, against 28.3% for Le Monde.

“More framing” is not a quantity this measures. There are two dimensions and they can point in different directions for the same pair. Name which one.

Outlets near the centre of the map are neutral.

Outlets near the centre are typical of this 25-outlet panel.

The centre is the panel average, and a device appears in more than a third of all headlines in it. There is no neutral point on this map.

The right frames more than the left.

Outlets far from the centre — in either direction — frame more than outlets near it.

Which side an outlet is on is uncorrelated with how often its headlines carry a device. Distance from the centre is what tracks.

Blast is the most charged outlet in France.

Nothing. Blast contributes 646 headlines and is a case profile.

Its interval on any single rate is several points wide, and it is two orders of magnitude smaller than the largest outlets here. Small outlets are the wrong place to draw comparisons.

34.6% of French headlines contain loaded language.

Up to 34.6% of headlines in this corpus carry at least one of the measured devices.

Three corrections: it is an upper bound, loaded language is one device rather than all of them, and the corpus is 25 outlets in three sections, not the French press.

Headlines mentioning Muslims are more hostile.

Headlines mentioning Muslims more often appear in crime and security story forms, and carry devices at a higher rate than the corpus average.

The measurement has no hostility term and no target. Coverage of an attack on a group and coverage attacking a group score the same.

This is a headline audit. It does not read article bodies, and it does not know whether a headline matches the piece beneath it. It also cannot see the stories an outlet chose not to cover — it observes what was published, which is narrower than an agenda.

Questions

Is this a media bias score?

No. It measures two properties of headline output — which story forms an outlet publishes, and which wording devices appear in them — against the average of the other outlets on the panel. Nothing in it tracks accuracy, and nothing tracks a left–right position. An outlet can be far from the panel average in either direction and be entirely accurate.

Why headlines rather than whole articles?

Because headlines are where most people meet a story, and a headline is a complete editorial artefact that can be measured consistently at this scale. The cost is real: the audit cannot see article-body framing, and cannot recover the context that would tell you whether charged wording is the outlet’s or a quoted source’s.

Why isn’t a given outlet here?

The panel is 25 national outlets chosen for reach, format diversity and ideological range, restricted to Politics, Economy and Society. Adding one is not appending a row: both axes are distances from the panel average, so a new outlet changes the baseline and every position moves.

Isn’t the vocabulary list itself a political choice?

The device definitions are, and so is the group lexicon. Both are published in full so the choices can be argued with rather than guessed at. Two independent annotators applying those definitions disagreed on 46.5% of headlines, which is the most direct available evidence of how contestable they are.

Can I use this data?

The derived tables are CC BY 4.0. Two things to check first: the classifier is recall-oriented, so corpus rates are upper bounds rather than point estimates, and the loaded-vocabulary and us-vs-them heads fall below the precision floors set for them. The paper reports per-device estimates.

Citation

@inproceedings{sobhy2026framing,
  title     = {Framing by Wording, Framing by Selection: A Large-Scale
               Two-Dimensional Audit of French News Headlines, 2022--2025},
  author    = {Sobhy, Amr},
  booktitle = {Proceedings of the 9th International Conference on Natural
               Language and Speech Processing (ICNLSP 2026)},
  year      = {2026},
  address   = {Trento, Italy},
  note      = {Accepted for oral presentation. To appear.},
}

Accepted for oral presentation at ICNLSP 2026, Trento. The citation will be updated when the camera-ready is final.

This work was supported by

AWS

Contact

For research questions, collaborations, or media inquiries.

Amr Sobhy