sources/sources.yaml
# Source registry for the unverified lane. Feeds are taken at their offered depth
# only (teaser/headline/summary); no full-text scraping, no robots, paywall, or
# bot-gate bypass, no third-party relays whose terms forbid non-personal use.
# Each entry records the basis on which we use it; the site renders this list
# in its About pane so the public can audit it.
# id is recorded in every article and essence record as provenance.
- id: aljazeera
name: Al Jazeera
via: rss
url: https://www.aljazeera.com/xml/rss/all.xml
depth: summary
site: https://www.aljazeera.com
takes: Article titles and summaries as published in the newsroom's public RSS feed.
basis: Public syndication feed offered by the newsroom. We take only feed-depth content, attribute every record, and link to the original article. Nonprofit use.
- id: bbc-world
name: BBC News (World)
via: rss
url: https://feeds.bbci.co.uk/news/world/rss.xml
depth: summary
site: https://www.bbc.com/news
takes: Article titles and summaries as published in the BBC's public RSS feed.
basis: Public feed offered by the newsroom. We take only feed-depth content, attribute every record, and link to the original article. Nonprofit use.
- id: guardian-world
name: The Guardian (World)
via: rss
url: https://www.theguardian.com/world/rss
depth: summary
site: https://www.theguardian.com/world
takes: Article titles and summaries as published in The Guardian's open RSS feed.
basis: Open RSS feed offered for syndication. We take only feed-depth content, attribute every record, and link to the original article. Nonprofit use.
- id: reuters
name: Reuters
via: deferred
url: none
depth: headline
site: https://www.reuters.com
takes: Nothing at present.
basis: No legitimate route exists today: Reuters retired its public RSS, its news sitemap and site are bot-gated, and third-party relays (e.g. Google News RSS) state terms that forbid non-personal use. Deferred until a partnership or approved API.