A River of News

Methodology

A River of News is a list. Headlines go in at the top as they arrive and the list is sorted by time and nothing else. This page says exactly how that works.

Where the headlines come from

Publishers expose RSS and Atom feeds. A separate service called NewsMarkets (newsmarkets.org) polls several hundred of those feeds and stores each item in a database. This site reads that database and shows the items in a different way.

For each item we keep five things: the headline, the publisher's name, a timestamp, the link to the original, and the byline when the feed gives one. We do not store or show article text, summaries or images. Every row links to the publisher and names it.

How items are ordered

Items are sorted by the time the publisher says the item was published. Two cases break that rule:

The first-seen time is kept next to the published time for every item, so the rule can change later without losing data.

NewsMarkets stores one time per item and replaces missing or future times with its fetch time before we see them. That means we cannot always tell a real publisher time from a substituted one. Treat times as accurate to a few minutes, not to the second.

Time zones

The server writes every time in UTC. If your browser allows scripts, they rewrite the times and the date headers into your own time zone. With scripts off you see UTC, and the page works the same.

Archive pages (a day, or a category on a day) always use UTC days, so a link means the same thing for everyone.

Categories

A category comes from the feed an item arrived in, not from the item itself. A feed filed under Sports puts everything it carries under Sports. That is simple and predictable, and it is sometimes wrong: a general feed can carry a story that belongs elsewhere. Items from feeds with no category appear only in the unfiltered list.

The categories are Politics, World, Business, Tech, AI, Crypto, Sports, Science, Weather, Culture, Music, Startups and Health.

Duplicates

An item that arrives twice, through two feeds of one publisher or under two forms of one link, is shown once. So is the same headline from the same publisher within the same hour. The same story from different publishers is not a duplicate: each publisher gets its own row. That is the point of a river.

Authors

A byline comes from the feed. Feeds write them every way, so we tidy them: we drop placeholders such as "author", email addresses and very long values, take the job title off "Jane Doe, senior correspondent", and split "A and B" into two authors. A desk or an agency (such as "TOI WORLD DESK" or "FRANCE24") counts as an author because that is the byline the publisher chose. About half of the items have one, and some publishers send none. Two spellings of one name can still appear as two authors when they differ by more than case, accents, punctuation or a space before a number.

Each author has a page with a summary and every article we hold. The summary's publisher and category counts cover the newest 500 articles. An author who writes for several outlets appears under all of them. We do not merge people who share a name.

Headlines and language

Some feeds double-encode special characters. We decode them, so an ampersand shows as an ampersand. We do not rewrite headlines in any other way.

Language is detected only when it is clear. A headline with several Spanish, German, French, Portuguese, Italian or Dutch function words and few English ones gets a small tag such as ES. Everything else is treated as English, which is not always right: a short headline in another language, or one made mostly of names, goes unlabeled. The search filter works on the same labels.

Keywords

A keyword is a word or a two-word phrase from a headline, with filler words removed. It ranks by how many different publishers used it that day, so one outlet repeating itself does not push a word up. A phrase replaces its words when most publishers who used the word used the phrase. "Rising today" lists terms used by at least one and a half times as many publishers in the last 24 hours as on a usual day over the previous week. A day with no bar on a term's chart means it was below the top list that day, not that nobody used it.

First to report

Over the last 30 days we look at stories covered by three or more publications. A publication leads a story when its earliest article came at least a minute before every other and no more than six hours before the next. It measures who published first in the feeds we read, not who had the story first, and it depends on the times feeds give us.

Audience lean

On the perspective spectrum and on the story pages for politics, world and business, publications are sorted by the political lean of their audience. We take it from a published study: Robertson, Jiang, Joseph, Friedland, Lazer and Wilson, "Auditing Partisan Audience Bias within Google Search", Proceedings of the ACM on Human-Computer Interaction 2(CSCW), 2018. The authors scored web domains by the share of registered Democrats and of registered Republicans on Twitter who shared their links. We use their data with their permission.

What this is, and is not:

Freshness

New items are collected about every five minutes. An item can arrive late, with a publisher time hours in the past, so the most recent six hours are rebuilt on every run. Older hours are final unless we correct them by hand.

Search

Search covers the headline, the publisher name and the author, over every item we hold. Words you type must all appear, and the last word also matches as a prefix, so "ukra" finds "Ukraine". Put words in quotes to match a phrase. Results are newest first, never ranked by relevance. Newest here means first seen by us, which differs from the published time only for late-arriving items. The date filter uses the same first-seen order, so an item can sit a few places either side of a day boundary.

The "story" link

When an item belongs to a story that NewsMarkets has grouped from two or more sources, the row shows a small link to that story's page on newsmarkets.org. That is the only link that leaves for a site we run.

What this site does not do

It does not rank, cluster, score, summarize, personalize or recommend. It has no accounts, comments, search or saved items.