espectroquem cobre o quê, na imprensa portuguesa

Methodology

espectro shows, for every significant Portuguese news story, which outlets covered it — and which ignored it. That requires two things: grouping articles about the same event, and knowing each outlet's editorial position. Both are done as follows.

Sourcing

We read the public RSS feeds outlets publish themselves — no page scraping. We store only the headline, a lead of up to 200 characters, and the link to the original article. Article bodies, and the images outlets attach to them, are never reproduced — the portrait that sometimes heads a story is a different thing, and “Where the pictures come from” below says what. Any outlet can request title-only display or removal, honoured within 48 hours.

The same article sometimes arrives twice. RTP files one piece under more than one section path — /economia/… and /lusa/… — so the address differs while the article does not. The real identifier sits at the end of the address, and that is what we count: two articles from the same outlet carrying the same identifier are one, and the first one collected is the one kept. Four cases in 3899 articles. The rule reads the address and never the text, so it has no way to hide two genuinely different stories on a resemblance.

It has a cost, and it is better said out loud: in two of the four the second address carried the headline the newsroom had already corrected, so keeping the first keeps the earlier wording — in one the difference is a typo, in the other it changes what the sentence means. That is the visible edge of a wider gap that holds for every outlet: when a newsroom corrects a headline without changing the address, we never find out, because we only ever look at new articles.

Story grouping

A multilingual text-similarity model groups articles describing the same event within a 72-hour window, tuned to prefer not merging when uncertain: a wrong merge corrupts a coverage bar, a miss merely duplicates a story. Manual corrections are recorded in a public corrections log — every entry, with the reason the editor wrote. That page, like the rest of the site, is in Portuguese.

Two articles can be grouped in either of two ways, and neither is an average of the other. Either the text is very similar and the two were published within six hours; or the text is somewhat less similar, they were published within eighteen hours, and they share enough concrete detail — quantities with units, proper names, acronyms, two- and three-word phrases almost nobody else uses. Each detail is weighted by how rare it is: a word appearing in thousands of articles proves nothing, one appearing in five proves a lot.

The second route exists because the first misses exactly the case this site is built to show. On 4 August 2026 Renascença ran “bank profits fall by 70 million to June”, RTP ran “profits at the main banks down 3.4%”, and Esquerda.net ran “banks make 14 million a day off their customers”. Same set of results, inverted framing, and by vocabulary the headlines barely touch — so they landed in separate stories, which is another way of saying the bar that showed the difference in framing never existed.

What reaches the front page

The main list carries only stories covered by two or more outlets: a site built on the question “who covered this?” cannot lead with something that has nothing to compare. It is also a volume rule — a single high-output newsroom can account for more than two thirds of everything collected in a week, and without the gate the front page would be one publisher's wire with everyone else's reporting buried under it.

Nothing is discarded. Single-outlet stories are published in a separate section below, capped at two per outlet. Within a story, articles are grouped by outlet — latest visible, earlier ones collapsed — so the page reflects what the bar counts: outlets, not article volume. Ordering is outlet count decayed over the day since the last article; no editorial hand.

Where the subject comes from

Every row carries a subject — politics, economy, society, world, regions. It is not ours: it is the label the newsroom itself put on the article, and it comes from two places, both published by the outlets. The first is the feed the article arrived on, since a newsroom publishing an economy-only feed has already said what is in it. The second is the category many newsrooms attach to each item, and that is what settles the common case: most articles arrive on a paper’s main feed, which is its own front page and not a subject at all.

Only labels on a fixed list are translated, and the whole label has to match — never a word found inside it. Opinião Pública is Portuguese for public opinion, i.e. polling, so a system hunting for the word would have hidden that story as commentary. Anything not on the list is left out with nothing put in its place: we do not open the article to find it a subject, nor infer one from the headline. That would be classifying content, which is precisely what this site does not do.

So some stories carry no subject, and they stay that way. There is also no sports or culture section here, and those are the two commonest labels discarded — the known gap in this rule, and better than filing news under headings the press did not use. Place and country names are not subjects either: Algarve says where, not what about.

The same label enforces the rule at the foot of this page. An article the outlet marks as opinion or commentary is stored — so it is not reconsidered on every poll, and so the count can be audited — but never becomes a story and never enters a bar. Until now that was achieved by not collecting opinion feeds, which missed the column published in the middle of the main one, where most of it runs.

Related stories

A story may carry a short list of others at the foot of the page. The link is vocabulary, not meaning: article titles are reduced to stems by Postgres's Portuguese dictionary and the two sets compared. Terms under four characters are dropped, as are terms carried by more than a tenth of recent stories — “government” and “millions” relate everything to everything.

The score is a proportion rather than a count: shared terms over the geometric mean of both vocabularies. A raw count would rank by how many articles a story attracted, the same volume artifact the two-outlet rule exists to correct. The thresholds are four shared terms and a score above 0.18, computed over each story's twenty most recent articles — a story does not become more related by being updated more often.

This measures words, not sense. Nobody read the two stories to confirm they belong together, and the method cannot tell two reports of one fire from two reports that happen to share wording. It is deliberately conservative: below those thresholds nothing is shown, and a story having no related stories at all is the normal case.

Search

Search runs over article titles, not article bodies, which we do not store. Accents are ignored on both sides, so habitação and habitacao return the same thing, as do incêndio and incendio. Typing without diacritics is ordinary on a phone, and an empty page would read as “nobody covered this” when the only problem was a missing tilde.

Results are stories, not articles, and the ordering asks how much of the story is about the term: matching articles over the story's articles, multiplied by the log of its outlet count. The first half is relevance, the second is weight, and it counts outlets rather than articles for the usual reason. A real example from this database: searching habitação (housing) also returns the Valpaços fire, because one headline mentions eight first- or second-home houses. That is a legitimate hit, but it ranks second, behind the story that is actually about housing. Every result shows the title the term was found in, so the reason it appears is visible.

What the bar means

The bar counts outlets per editorial position covering a story. It never rates individual articles, and nothing is labelled true or false.

The bar itself always draws all five positions. The legend beside it varies with the space available: in lists — front page, sections, search, blindspots — the five positions are grouped into three sides, centre-left counting as left and centre-right as right, with unrated outlets kept separate. Six spelled-out positions do not fit on a card line and would push the bar itself onto a second one. The story page lists all five separately, each with its own count. Same tally, different grain: a list never shows a number the story page contradicts.

One masthead is not one newsroom

The bar counts outlets, and in Portugal that leaves a question unanswered: how many different newsrooms actually wrote the story. A large part of what looks like independent coverage is one wire dispatch going out under different logos — the same sentence, in the same minute, under eleven mastheads. 11 outlets is true and misleading at once. So when the two numbers disagree, the counts line shows both: 11 outlets · 10 newsrooms. When they agree there is no second number to show, and none is shown.

Both side by side, and the bar stays as it is. Collapsing it would make every number on the site depend on a detector whose precision has not yet been measured against properly clustered stories, and would silently change what the bar means: a reader opening the story would find eleven headlines and a bar saying ten. It is reversible on purpose. The day the detector is measured above 95% precision on real data, promoting newsrooms to the headline number is one line of code — and that is a right to earn rather than assume.

Lusa, the national agency, is not ingested: its feeds are a paid subscription product, so it has zero feeds and zero articles here, and there is no original to compare against. Everything below is outlet against outlet — two mastheads publishing the same text — and the system never asserts whose text it was.

The first signal is the title. Normalised titles are compared — lower case, accents and punctuation stripped — and one containing the other counts as a match, because outlets trim and extend the front of the same headline: “Preventiva” against “Prisão preventiva”, “Montenegro” against “Luís Montenegro”. The short label a newsroom puts in front of its own title (“Médio Oriente:”, “Incêndios.”) is removed before comparing. The floor is 40 characters: below that the rule stops matching dispatches and starts matching stock phrases.

The second signal is the lead, the first 52 characters after the same normalisation. Only 22% of articles carry a lead — 867 of 3899 — so it is structurally thin, and it stays because it catches what the title cannot: the case where the newsroom rewrote the headline and published the agency’s lead untouched. The 52 was swept from 20 to 150: the count does not move between 49 and 54; at 55 a true pair is lost, a regional newsroom having swapped “alertou hoje” for “alertou esta sexta-feira”; and at 48 the first false ones appear, because “O Presidente da República, António José Seguro,” is 45 characters of formal title before a single word of news.

Two conditions close the rule. The two publications must be less than 24 hours apart. Of the 88 pairs the title signal finds with no window at all, all but four are under six hours apart; those four were read, and two of them — at 8.7 hours and 18.6 hours — really are the same dispatch, while the other two are 51 hours apart and are a quotation reused as a headline two days later. Cutting at six hours would kill two true pairs to remove two false ones; at 24 hours both stay and both go, with 32 hours of empty gap for the line to sit in. And an outlet never counts as a copy of itself — a paper republishing its own text is a problem of ours, not one newsroom fewer. The final count is connected components across outlets: if A and B share text and so do B and C, the three count as one newsroom. That is the charitable reading, and it is deliberate — it errs towards fewer independent newsrooms than there are, never more.

What was checked by hand: the 91 pairs the detector found in this corpus were read one by one — all 91, not a sample — and all 91 are the same text. 84 matched on the title alone, 5 on the lead alone and 2 on both. The weakest span that ever matched is 45 characters.

Two caveats, because neither shows up in those numbers. The hand-check judges “same dispatch” largely from the same textual evidence the detector uses, so it is partly circular; what it establishes independently is that no pair joins two different events and that no matched span is a house formula. And the second number is a floor: copying is only visible between the outlets we follow, and a newsroom that rewrote the dispatch in its own words counts as a newsroom of its own. That is right for the label — it did write it — and it means this number never claims to measure originality.

Two signals were tried and are not here, written down so they are not proposed again as an oversight. The “(Lusa)” credit line is unambiguous when it appears, and it appears once in 3899 articles. And the byline, which intuition says agency copy should be missing: RTP’s feed carries no author field on any item, so the signal is not weak, it is absent.

Outlet ratings

Every rating requires at least two independent public sources, weighted in this order: Reuters Institute / OberCom audience self-placement data for Portuguese brands; peer-reviewed academic characterisations; international raters (e.g. Media Bias/Fact Check) where they cover the outlet; documented editorial positions and ownership. Outlets without sufficient sources are shown, honestly, as unrated. Each rating publishes its sources and full change history; ratings change only in quarterly reviews or through documented corrections.

Between two sources and none there is a middle case, and it needs stating. Where a position has been established but cannot yet take effect, it is published as provisional: written on the outlet’s own page and in the outlet register, marked as such in both, and counted nowhere. In story bars, in the per-side percentages, in blindspot detection and in the register’s own coverage counts, an outlet holding a provisional position is counted as unrated. Withholding what has already been established would not be more honest, only quieter; counting it would mean the two-source rule bends.

Provisional for one of two reasons, and the outlet’s page always says which. Either the public sources fall short of the minimum, in which case what is missing is evidence — it may never turn up, and until it does the position stays where it is. Or the dossier already holds the two sources, written out with the citations in view, and nobody has signed it. Signing is the step where whoever answers for this site reads the whole dossier, accepts the conclusion and fixes the date from which it holds — which is what the “in force since” line on an outlet’s page means. Without that step a rating would be the output of having collected sources, and collecting sources is not deciding.

Any outlet may contest its rating. While the review runs, the rating is marked under review — again in two places, in the register beside the position and on the outlet’s page, the latter with the date it opened and what is being re-examined, written by the person who opened it rather than by the system. It is deliberately absent from story bars: the bar counts outlets, and an outlet under review still counts exactly as it did before. That is the part that matters. The rating stays in force until the review concludes, because withdrawing it while waiting for the outcome would be changing it without having reviewed it — precisely what the previous paragraph forbids. Opening and closing are both recorded in the public corrections log, like any other editorial decision.

Factuality and access

Editorial position says which side an outlet writes from. It says nothing about how rigorous it is, and those are different questions: a paper from any part of the spectrum can be careful or careless. Factuality is therefore a separate axis with three levels only — high, mixed, low. Three rather than seven because that is as far as citable evidence goes; a finer scale would claim a precision the sources do not have. It obeys the same two-independent-sources rule, carries the same history, and stays unassessed where no dossier exists — never high by default. It applies to the outlet over time, never to the article on screen.

Today it is unassessed on all 57, and that is a measurement rather than a backlog. On 07/08/2026 we went looking for citable factuality ratings of the Portuguese press. NewsGuard does not count Portugal among the markets it covers. Media Bias/Fact Check has entries for some national titles — we found Público, Observador and Expresso — and its country profile sits behind a subscription, so not even the full list can be read. For the regional press, which is 27 of the 57 outlets in this register, the public searches we ran returned nothing at all.

Rating three outlets and leaving 54 blank would not be information. It would be one American organisation’s verdict on three national papers and silence about everyone else — and silence, sitting next to three ratings, reads as a failing grade. So the axis stays unassessed everywhere until a source covers the country, and this paragraph is the answer, rather than a rating invented to fill the column.

Each outlet also carries an access flag — free, partial, or subscribers only — so a reader knows before clicking. That is a verifiable fact rather than a judgement, so it needs no sources. Where nobody has checked it reads unverified: the absence of a badge never means the article is free.

Who owns what

Editorial position and factuality are judgements with public sources behind them. Ownership is not a judgement; it is a matter of record, and it is kept separate for that reason. Every story adds up the outlets covering it a second time, by owner group: eleven newsrooms can be five owners, and it is the second count that says anything about plurality.

The group name is the key that count is made on, not a label. Written two ways, one owner counts as two and the line hides the concentration it exists to show — so names are normalised, and the variants (the registered company, the commercial brand, the local operating firm) live in the outlet’s ownership note rather than in the key.

That note follows the ratings rule: who controls the capital, since when, and on what public sources. A page that prints sources for this rating beside the editorial position and then asserts ownership citing nothing would be one page held to two standards. Where the sources are not yet gathered the profile says the note is missing instead of writing one anyway, and the register counts how many outlets are in each state.

The part most easily read backwards is what happens to what we do not know. Not established is not independent. Most outlets in the register have no group recorded, nearly all of them regional, and calling those independent would publish the strongest claim available about a newsroom — that it answers to nobody — on the strength of not having looked. They therefore count as a group nowhere: in a story’s ownership breakdown they are stated separately, outside the list, and never at the head of it by weight of numbers.

There is nevertheless a column in the register that reads Independent, and it means the opposite of the absence above. Each outlet is characterised as one of three kinds — general-interest press, party press, independent — and the last two are claims about who a newsroom answers to, so neither is written without public sources cited on the outlet’s own profile, exactly as the editorial position is. Independent means we looked and established that it belongs to no media group and no party structure. An outlet with no group recorded and no such check does not become independent: it stays uncharacterised, and the register says how many of those there are.

Party press is the strongest of the three and it removes nothing from anyone. A party paper counts in the coverage bar like any other, there is no switch that hides it, and the party is recorded as the owner group — which is what it is. What the label does is say so in the open, on the profile and in the register, so that the bar stays a count rather than a selection of ours.

Knowing who owns a paper is not a verdict on it. A large group does not make a report false, and an owner nobody has established does not make a newsroom free. Documented ownership is one input among others to a rating; everywhere else on the site it is what it is — a matter of record, printed next to the story, for the reader to draw conclusions from.

Reader personalisation, without an account

One page describes the reader rather than the press: it shows which side of the spectrum their own reading has come from. It works with no account, no cookie and no server — the tally is written to that browser’s local storage and read back by the same browser, and no code path sends it anywhere, which anyone can confirm in a network panel in about ten seconds. What is stored is an outlet’s short name and how many times it was opened; never the headline, never the timestamp, because a list of those is a reading history and a reading history on a shared computer is a liability created for no gain. Editorial positions are not stored either — they are applied at render time, so a re-rating moves the reader’s own history with it. Percentages appear only after ten articles from rated outlets. The front page is identical for everyone who opens it, and stays that way.

Blindspots

When a story with at least five rated outlets is covered 60%+ by one side while the other side barely touches it, the system proposes a blindspot. Publication always requires human review, and both sides of the spectrum get flagged over time.

That last sentence is a claim, so the blindspots page is built so that it can be checked without counting anything: two sections, one per side, both always present, each carrying its own count in the heading. If a week’s flags all go one way, the other section reads zero and says so. Headings name the press rather than the reader — rarely covered by the left-leaning press — because rated outlets present on the story is what the system actually counts. A story can be flagged for both audiences, in which case it appears in both sections.

A flagged story also says so on its own page — below the numbers, above the timeline — carrying the note the reviewer left, in their words rather than the system’s. It does not restate the percentages sitting two lines above it; it adds what the bar cannot say. When the link is shared, the preview text leads with the finding, because it is the one thing on that page no other site is saying.

Where the pictures come from

A story may open with a photograph, and it is worth being exact about what it is. It is not the picture the outlet ran with its article — that one is never shown — nor a photograph of the event. It is a portrait of who the story is about: a person or a party. A library portrait, telling you who is being discussed, not what happened.

It comes from Wikimedia Commons, Wikipedia’s free-media repository, and only from there. It arrives the same way as everything else here — a request to the public interface, no scraping — and is used only if the licence is free: public domain, CC0, or Creative Commons with attribution (CC BY, CC BY-SA). Anything requiring non-commercial use or forbidding changes is left out, even where it exists. We keep our own copy and serve it ourselves, so the page fetches nothing from third-party servers.

The caption under the portrait is not decoration: it is what the licence requires. It names the author, names the licence with a link to its deed, and links back to the original file on Commons, so anyone can confirm the provenance and the terms. Where the Commons record names no author, the caption says so rather than inventing one.

The subject is a name drawn from the story’s own headlines, and it is used only if it matches a person or party with a Portuguese Wikipedia page — the same two-proof demand made everywhere else, applied here to identity. In any doubt we show no face at all: the wrong one would be worse than none. So most stories carry no photograph, which is the expected case, not a failure.

What we never do