services/ssr.py walks JSON-LD and Next.js __NEXT_DATA__ for article-like
objects (headline, or title+url, or name with article signals), recovering
JS-rendered sites without a browser. prepare_headlines favors embedded JSON
over regex link scraping. Forbes is re-added as a live proof.
- services/feeds.py parses RSS 2.0/1.0 and Atom into (title, date) entries
- prepare_headlines auto-detects feeds by content and parses them directly
- strip feeds./rss./moxie. subdomains so feeds collapse to the outlet domain
- add 7 verified RSS feeds (BBC, Guardian, Al Jazeera, WaPo, The Hill, Vox, CNBC)
- get_sources skips blank lines and # comments
- broaden excluded-phrase list (skip links, share/follow, newsletter, utility)
Add resources/excluded_phrases.txt (skip links, privacy/consent, legal
boilerplate) and match it case- and punctuation-insensitively as a
contiguous run of words. Phrases live in a data file so they can be
extended without touching code.