- services/feeds.py parses RSS 2.0/1.0 and Atom into (title, date) entries - prepare_headlines auto-detects feeds by content and parses them directly - strip feeds./rss./moxie. subdomains so feeds collapse to the outlet domain - add 7 verified RSS feeds (BBC, Guardian, Al Jazeera, WaPo, The Hill, Vox, CNBC) - get_sources skips blank lines and # comments - broaden excluded-phrase list (skip links, share/follow, newsletter, utility)master
| @@ -10,8 +10,9 @@ question: *which stories are multiple independent sources reporting right now?* | |||||
| ## How it works | ## How it works | ||||
| 1. **Load** source URLs and stopwords from `resources/`. | 1. **Load** source URLs and stopwords from `resources/`. | ||||
| 2. **Fetch** each source and extract candidate headline text from `<a>`/`<span>` | |||||
| tags, associating each headline with the nearest `<time>` publication timestamp. | |||||
| 2. **Fetch** each source. RSS/Atom feeds are parsed as structured XML (titles and | |||||
| publication timestamps); HTML pages are scraped for `<a>`/`<span>` text, | |||||
| associating each headline with the nearest `<time>` publication timestamp. | |||||
| 3. **Normalize** each headline (lowercase, strip punctuation, remove stopwords). | 3. **Normalize** each headline (lowercase, strip punctuation, remove stopwords). | ||||
| 4. **Cluster** headlines into stories using pairwise cosine similarity, linking | 4. **Cluster** headlines into stories using pairwise cosine similarity, linking | ||||
| matches transitively so a chain of near-duplicates collapses into one story. | matches transitively so a chain of near-duplicates collapses into one story. | ||||
| @@ -39,7 +40,8 @@ Run from the project root — the resource paths are relative (`./resources/...` | |||||
| ## Configuration | ## Configuration | ||||
| - `resources/sources.txt` — one news source URL per line. | |||||
| - `resources/sources.txt` — one source URL per line (HTML pages **or** RSS/Atom | |||||
| feeds; `#` comments and blank lines are ignored). | |||||
| - `resources/stopwords.txt` — one stopword per line. | - `resources/stopwords.txt` — one stopword per line. | ||||
| - `resources/excluded_phrases.txt` — boilerplate link text to ignore (see below). | - `resources/excluded_phrases.txt` — boilerplate link text to ignore (see below). | ||||
| - Defaults live at the top of `main.py` (`SIMILARITY_THRESHOLD = 0.75`, | - Defaults live at the top of `main.py` (`SIMILARITY_THRESHOLD = 0.75`, | ||||
| @@ -89,7 +91,13 @@ News stories covered by at least 2 distinct sources (published on or after 2026- | |||||
| - **A "source" is a distinct domain**, not a distinct URL. `www.cnn.com/us` and | - **A "source" is a distinct domain**, not a distinct URL. `www.cnn.com/us` and | ||||
| `cnn.com/politics` both normalize to `cnn.com`, so one outlet counts once even | `cnn.com/politics` both normalize to `cnn.com`, so one outlet counts once even | ||||
| when listed under multiple sections. | |||||
| when listed under multiple sections. Feed/redirect subdomains (`feeds.`, `rss.`, | |||||
| `moxie.`) are also stripped, so a feed and an HTML page from the same outlet | |||||
| still collapse to one source. | |||||
| - **RSS/Atom sources are preferred when available.** They're structured and far | |||||
| more reliable than scraping JavaScript-heavy or paywalled pages, and they carry | |||||
| publication timestamps directly. Feed type is auto-detected from the content, so | |||||
| HTML and feed URLs can live side by side in `sources.txt`. | |||||
| - **Date windowing drops undated headlines.** When `--since` or `--until` is set, | - **Date windowing drops undated headlines.** When `--since` or `--until` is set, | ||||
| any headline whose page carries no parseable timestamp is excluded because its | any headline whose page carries no parseable timestamp is excluded because its | ||||
| recency can't be established (the count is logged). Without a date flag, | recency can't be established (the count is logged). Without a date flag, | ||||
| @@ -117,7 +125,8 @@ anya/ | |||||
| ├── main.py # CLI entry point and orchestration | ├── main.py # CLI entry point and orchestration | ||||
| ├── services/ | ├── services/ | ||||
| │ ├── headlines.py # fetch + parse headlines (and timestamps) | │ ├── headlines.py # fetch + parse headlines (and timestamps) | ||||
| │ ├── normalization.py # stopword loading and headline normalization | |||||
| │ ├── feeds.py # RSS/Atom feed detection and parsing | |||||
| │ ├── normalization.py # stopword/phrase loading and headline normalization | |||||
| │ ├── similarity.py # cosine similarity over token lists | │ ├── similarity.py # cosine similarity over token lists | ||||
| │ ├── sources.py # load source URLs | │ ├── sources.py # load source URLs | ||||
| │ ├── stories.py # cluster headlines into stories | │ ├── stories.py # cluster headlines into stories | ||||
| @@ -3,38 +3,71 @@ | |||||
| # phrase matches whenever it appears as a contiguous run of words in a link's | # phrase matches whenever it appears as a contiguous run of words in a link's | ||||
| # text (so "privacy choices" also matches "Your Privacy Choices"). | # text (so "privacy choices" also matches "Your Privacy Choices"). | ||||
| # Lines starting with '#' are ignored. | # Lines starting with '#' are ignored. | ||||
| # | |||||
| # Note: phrases shorter than three words are largely redundant here because the | |||||
| # headline filter already rejects text under three words — they are kept only | |||||
| # as defensive documentation. Focus additions on 3+ word phrases that would | |||||
| # otherwise slip through. | |||||
| # Skip / navigation | # Skip / navigation | ||||
| skip to content | skip to content | ||||
| skip to main content | skip to main content | ||||
| skip to navigation | |||||
| jump to content | jump to content | ||||
| jump to main content | jump to main content | ||||
| main navigation | main navigation | ||||
| open menu | |||||
| close menu | |||||
| main menu | |||||
| back to top | |||||
| return to top | |||||
| # Privacy / consent | # Privacy / consent | ||||
| your privacy choices | |||||
| privacy choices | privacy choices | ||||
| privacy policy | privacy policy | ||||
| cookie settings | cookie settings | ||||
| cookie preferences | cookie preferences | ||||
| manage cookies | |||||
| manage consent | manage consent | ||||
| do not sell my personal information | do not sell my personal information | ||||
| do not sell or share my personal information | do not sell or share my personal information | ||||
| do not sell or share my personal data | do not sell or share my personal data | ||||
| do not sell my info | |||||
| # Legal | # Legal | ||||
| terms of service | terms of service | ||||
| terms and conditions | terms and conditions | ||||
| terms of use | terms of use | ||||
| conditions of use | conditions of use | ||||
| terms of sale | |||||
| accessibility statement | accessibility statement | ||||
| copyright notice | copyright notice | ||||
| reprints and permissions | |||||
| # Advertising / corporate | |||||
| advertise with us | advertise with us | ||||
| advertise here | |||||
| about our ads | |||||
| work with us | |||||
| investor relations | |||||
| media kit | |||||
| # Apps / downloads | |||||
| download the app | download the app | ||||
| terms of sale | |||||
| reprints and permissions | |||||
| download our app | |||||
| get the app | |||||
| # Newsletter / subscription | |||||
| sign up for our newsletter | |||||
| subscribe to our newsletter | |||||
| newsletter sign up | |||||
| # Social / sharing | |||||
| share this article | |||||
| share this story | |||||
| share on facebook | |||||
| share on twitter | |||||
| share on linkedin | |||||
| follow us on | |||||
| # Search / account / utility | |||||
| search this site | |||||
| search the site | |||||
| manage your account | |||||
| reset your password | |||||
| @@ -17,4 +17,12 @@ https://www.nytimes.com/section/us | |||||
| https://www.nytimes.com/section/politics | https://www.nytimes.com/section/politics | ||||
| https://www.usnews.com/news | https://www.usnews.com/news | ||||
| https://time.com/ | https://time.com/ | ||||
| https://www.axios.com/politics-policy | |||||
| https://www.axios.com/politics-policy | |||||
| # RSS feeds (parsed as structured XML rather than scraped HTML) | |||||
| https://feeds.bbci.co.uk/news/world/rss.xml | |||||
| https://www.theguardian.com/us-news/rss | |||||
| https://www.aljazeera.com/xml/rss/all.xml | |||||
| https://feeds.washingtonpost.com/rss/world | |||||
| https://thehill.com/feed/ | |||||
| https://www.vox.com/rss/index.xml | |||||
| https://www.cnbc.com/id/100003114/device/rss/rss.html | |||||
| @@ -0,0 +1,65 @@ | |||||
| import logging | |||||
| import xml.etree.ElementTree as ET | |||||
| from services.dates import parse_datetime | |||||
| logger = logging.getLogger(__name__) | |||||
| # Distinctive root tags used to sniff feed (XML) content apart from HTML. | |||||
| FEED_MARKERS = ('<rss', '<feed', '<rdf:rdf') | |||||
| def looks_like_feed(content): | |||||
| """Best-effort sniff for RSS/Atom XML content (vs. HTML pages).""" | |||||
| if not content or not isinstance(content, str): | |||||
| return False | |||||
| head = content[:4000].lower() | |||||
| return any(marker in head for marker in FEED_MARKERS) | |||||
| def _local(tag): | |||||
| """Return an element's local name, ignoring any XML namespace prefix.""" | |||||
| return tag.rsplit('}', 1)[-1] if '}' in tag else tag | |||||
| def _child_text(elem, names): | |||||
| """First direct child whose local name is in ``names``, with text content.""" | |||||
| for child in elem: | |||||
| if _local(child.tag) in names: | |||||
| text = (child.text or '').strip() | |||||
| if text: | |||||
| return text | |||||
| return None | |||||
| def parse_feed(content): | |||||
| """Extract ``(title, published_at)`` tuples from an RSS or Atom feed. | |||||
| Understands RSS 2.0 (``<item>``), RSS 1.0 (``<rdf:RDF>`` items) and Atom | |||||
| (``<entry>``). Publication timestamps are taken from ``pubDate``, | |||||
| ``published``, ``updated`` or ``dc:date`` (``None`` when absent or | |||||
| unparseable). | |||||
| """ | |||||
| items = [] | |||||
| try: | |||||
| root = ET.fromstring(content) | |||||
| except ET.ParseError as e: | |||||
| logger.error("Failed to parse feed XML: %s", e) | |||||
| return items | |||||
| except Exception as e: | |||||
| logger.error("Unexpected error parsing feed: %s", e, exc_info=True) | |||||
| return items | |||||
| # Work over every <item>/<entry> element regardless of nesting, which covers | |||||
| # RSS 2.0, RSS 1.0 (RDF) and Atom in one pass. | |||||
| entries = [el for el in root.iter() if _local(el.tag) in ('item', 'entry')] | |||||
| for entry in entries: | |||||
| title = _child_text(entry, ('title',)) | |||||
| if not title: | |||||
| continue | |||||
| published = _child_text(entry, ('pubDate', 'published', 'updated', 'date')) | |||||
| items.append((title, parse_datetime(published))) | |||||
| logger.info("Parsed %d entries from feed", len(items)) | |||||
| return items | |||||
| @@ -4,6 +4,7 @@ import re | |||||
| import requests | import requests | ||||
| from services.dates import parse_datetime | from services.dates import parse_datetime | ||||
| from services.feeds import looks_like_feed, parse_feed | |||||
| from services.normalization import is_excluded, normalize_headline | from services.normalization import is_excluded, normalize_headline | ||||
| from structs.headline import Headline | from structs.headline import Headline | ||||
| @@ -86,6 +87,21 @@ def _extract_candidates(source_content): | |||||
| return candidates | return candidates | ||||
| def _collect_feed_headlines(source_content, source_url, stopwords, excluded_phrases): | |||||
| """Turn parsed feed entries into headline objects, applying normal filters.""" | |||||
| collected = [] | |||||
| for title, published_at in parse_feed(source_content): | |||||
| cleaned_text = title.strip() | |||||
| if not cleaned_text or not is_headline(cleaned_text, stopwords, excluded_phrases): | |||||
| continue | |||||
| try: | |||||
| normalized = normalize_headline(cleaned_text, stopwords) | |||||
| collected.append(Headline(cleaned_text, normalized, source_url, published_at)) | |||||
| except Exception as e: | |||||
| logger.error("Failed to normalize feed title %r from '%s': %s", cleaned_text, source_url, e, exc_info=True) | |||||
| return collected | |||||
| def is_headline(text, stopwords=None, excluded_phrases=None): | def is_headline(text, stopwords=None, excluded_phrases=None): | ||||
| if not text or not isinstance(text, str): | if not text or not isinstance(text, str): | ||||
| return False | return False | ||||
| @@ -142,6 +158,17 @@ def prepare_headlines(sources, stopwords, timeout=DEFAULT_TIMEOUT, headers=None, | |||||
| logger.error("Unexpected error fetching source '%s': %s", source_url, e, exc_info=True) | logger.error("Unexpected error fetching source '%s': %s", source_url, e, exc_info=True) | ||||
| continue | continue | ||||
| if looks_like_feed(source_content): | |||||
| logger.info("Detected RSS/Atom feed for source '%s'", source_url) | |||||
| try: | |||||
| feed_headlines = _collect_feed_headlines(source_content, source_url, stopwords, excluded_phrases) | |||||
| except Exception as e: | |||||
| logger.error("Failed to parse feed from '%s': %s", source_url, e, exc_info=True) | |||||
| continue | |||||
| headlines.extend(feed_headlines) | |||||
| logger.info("Successfully extracted %d headlines from source '%s'", len(feed_headlines), source_url) | |||||
| continue | |||||
| logger.debug("Parsing HTML content from '%s' for headline candidates", source_url) | logger.debug("Parsing HTML content from '%s' for headline candidates", source_url) | ||||
| try: | try: | ||||
| candidates = _extract_candidates(source_content) | candidates = _extract_candidates(source_content) | ||||
| @@ -4,20 +4,15 @@ logger = logging.getLogger(__name__) | |||||
| def get_sources(path, delimiter='\n'): | def get_sources(path, delimiter='\n'): | ||||
| """Load source URLs from a file, one per line. | |||||
| Blank lines and lines starting with ``#`` are ignored, so the file can be | |||||
| annotated with comments (e.g. grouping HTML vs. RSS feeds). | |||||
| """ | |||||
| logger.info("Attempting to load sources from file: '%s' with delimiter: %r", path, delimiter) | logger.info("Attempting to load sources from file: '%s' with delimiter: %r", path, delimiter) | ||||
| sources = None | |||||
| try: | try: | ||||
| with open(path, 'r', encoding='utf-8') as source_file: | |||||
| logger.debug("Successfully opened file '%s' for reading", path) | |||||
| with open(path, 'r', encoding='utf-8-sig') as source_file: | |||||
| content = source_file.read() | content = source_file.read() | ||||
| logger.debug("Read %d bytes/characters from '%s'", len(content), path) | |||||
| sources = content.split(delimiter) | |||||
| logger.info("Successfully read and split %d source entries from '%s'", len(sources), path) | |||||
| for idx, src in enumerate(sources): | |||||
| if not src.strip(): | |||||
| logger.warning("Source at index %d is empty or whitespace-only: %r", idx, src) | |||||
| else: | |||||
| logger.debug("Source [%d]: %s", idx, src) | |||||
| except FileNotFoundError: | except FileNotFoundError: | ||||
| logger.error("Source file not found at path: '%s'", path, exc_info=True) | logger.error("Source file not found at path: '%s'", path, exc_info=True) | ||||
| raise | raise | ||||
| @@ -28,4 +23,12 @@ def get_sources(path, delimiter='\n'): | |||||
| logger.error("Failed to read sources from '%s': %s", path, e, exc_info=True) | logger.error("Failed to read sources from '%s': %s", path, e, exc_info=True) | ||||
| raise | raise | ||||
| sources = [] | |||||
| for line in content.split(delimiter): | |||||
| stripped = line.strip() | |||||
| if not stripped or stripped.startswith('#'): | |||||
| continue | |||||
| sources.append(stripped) | |||||
| logger.info("Successfully loaded %d source entries from '%s'", len(sources), path) | |||||
| return sources | return sources | ||||
| @@ -28,17 +28,21 @@ class Headline: | |||||
| def _extract_domain(url: str): | def _extract_domain(url: str): | ||||
| """Return a normalized outlet domain for a source URL, or None if absent. | """Return a normalized outlet domain for a source URL, or None if absent. | ||||
| Strips scheme, surrounding whitespace and a leading ``www.`` so that | |||||
| ``https://www.cnn.com/us`` and ``https://cnn.com/politics`` both resolve | |||||
| to ``cnn.com`` — i.e. one outlet counts once, regardless of section URL. | |||||
| Strips scheme, whitespace and a leading ``www.``/feed/redirect subdomain | |||||
| (``www.``, ``feeds.``, ``rss.``, ``moxie.``) so that ``https://www.cnn.com/us`` | |||||
| and ``https://feeds.npr.org/1001/rss.xml`` both resolve to their outlet | |||||
| (``cnn.com``, ``npr.org``) — one outlet counts once regardless of section | |||||
| URL or feed host. | |||||
| """ | """ | ||||
| if not url or not isinstance(url, str) or not url.strip(): | if not url or not isinstance(url, str) or not url.strip(): | ||||
| return None | return None | ||||
| host = urlparse(url.strip()).netloc.lower() | host = urlparse(url.strip()).netloc.lower() | ||||
| if not host: | if not host: | ||||
| return None | return None | ||||
| if host.startswith("www."): | |||||
| host = host[4:] | |||||
| for prefix in ("www.", "feeds.", "rss.", "moxie."): | |||||
| if host.startswith(prefix): | |||||
| host = host[len(prefix):] | |||||
| break | |||||
| return host | return host | ||||
| def compare_headlines(self, other_headline): | def compare_headlines(self, other_headline): | ||||
| @@ -0,0 +1,103 @@ | |||||
| import os | |||||
| import tempfile | |||||
| import unittest | |||||
| from datetime import datetime, timezone | |||||
| from unittest.mock import patch, MagicMock | |||||
| from services.feeds import looks_like_feed, parse_feed | |||||
| from services.sources import get_sources | |||||
| from services.headlines import prepare_headlines | |||||
| from structs.headline import Headline | |||||
| RSS_2_0 = """<?xml version="1.0" encoding="UTF-8"?> | |||||
| <rss version="2.0"><channel> | |||||
| <title>Test feed</title> | |||||
| <item><title>First news story</title><link>https://x/1</link><pubDate>Mon, 14 Sep 2026 12:00:00 GMT</pubDate></item> | |||||
| <item><title>Second news story</title><link>https://x/2</link></item> | |||||
| </channel></rss> | |||||
| """ | |||||
| ATOM = """<?xml version="1.0" encoding="utf-8"?> | |||||
| <feed xmlns="http://www.w3.org/2005/Atom"> | |||||
| <entry><title>Atom story one</title><published>2026-09-14T14:30:00Z</published></entry> | |||||
| <entry><title>Atom story two</title><updated>2026-09-13T10:00:00Z</updated></entry> | |||||
| </feed> | |||||
| """ | |||||
| class TestLooksLikeFeed(unittest.TestCase): | |||||
| def test_rss_and_atom_detected(self): | |||||
| self.assertTrue(looks_like_feed(RSS_2_0)) | |||||
| self.assertTrue(looks_like_feed(ATOM)) | |||||
| def test_html_not_detected(self): | |||||
| self.assertFalse(looks_like_feed("<html><body><a>Story</a></body></html>")) | |||||
| def test_empty_not_detected(self): | |||||
| self.assertFalse(looks_like_feed("")) | |||||
| self.assertFalse(looks_like_feed(None)) | |||||
| class TestParseFeed(unittest.TestCase): | |||||
| def test_rss_2_0(self): | |||||
| items = parse_feed(RSS_2_0) | |||||
| titles = [t for t, _ in items] | |||||
| self.assertEqual(titles, ["First news story", "Second news story"]) | |||||
| self.assertEqual(items[0][1].year, 2026) | |||||
| self.assertIsNone(items[1][1]) | |||||
| def test_atom(self): | |||||
| items = parse_feed(ATOM) | |||||
| titles = [t for t, _ in items] | |||||
| self.assertEqual(titles, ["Atom story one", "Atom story two"]) | |||||
| self.assertIsNotNone(items[0][1]) | |||||
| self.assertIsNotNone(items[1][1]) | |||||
| def test_malformed_returns_empty(self): | |||||
| self.assertEqual(parse_feed("not xml at all <<<"), []) | |||||
| class TestGetSources(unittest.TestCase): | |||||
| def test_skips_comments_and_blanks(self): | |||||
| path = None | |||||
| try: | |||||
| with tempfile.NamedTemporaryFile("w", suffix=".txt", delete=False) as f: | |||||
| f.write("# a comment\n\nhttps://a.com\n\nhttps://b.com\n# another\n") | |||||
| path = f.name | |||||
| self.assertEqual(get_sources(path), ["https://a.com", "https://b.com"]) | |||||
| finally: | |||||
| if path: | |||||
| os.unlink(path) | |||||
| class TestFeedDomainNormalization(unittest.TestCase): | |||||
| def test_feed_subdomain_stripped(self): | |||||
| self.assertEqual(Headline("a", ["a"], "https://feeds.npr.org/1001/rss.xml").domain, "npr.org") | |||||
| self.assertEqual(Headline("a", ["a"], "https://rss.nytimes.com/x").domain, "nytimes.com") | |||||
| self.assertEqual(Headline("a", ["a"], "https://www.vox.com/rss/index.xml").domain, "vox.com") | |||||
| class TestPrepareHeadlinesFeed(unittest.TestCase): | |||||
| def setUp(self): | |||||
| self.stopwords = {"the", "a", "an", "in", "on", "and", "of", "to"} | |||||
| @patch("services.headlines.requests.get") | |||||
| def test_feed_source_is_parsed(self, mock_get): | |||||
| mock_response = MagicMock() | |||||
| mock_response.status_code = 200 | |||||
| mock_response.content = RSS_2_0.encode("utf-8") | |||||
| mock_response.text = RSS_2_0 | |||||
| mock_get.return_value = mock_response | |||||
| headlines = prepare_headlines(["https://feeds.bbci.co.uk/news/world/rss.xml"], self.stopwords) | |||||
| titles = [h.display_text for h in headlines] | |||||
| self.assertEqual(titles, ["First news story", "Second news story"]) | |||||
| # Domain comes from the feed host, normalized to the outlet. | |||||
| self.assertEqual(headlines[0].domain, "bbci.co.uk") | |||||
| self.assertIsNotNone(headlines[0].published_at) | |||||
| if __name__ == "__main__": | |||||
| unittest.main() | |||||