You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
Jared Bell c765b53169 Expand excluded phrases list to filter additional boilerplate 3 天之前
.idea first commit 3 天之前
resources Expand excluded phrases list to filter additional boilerplate 3 天之前
services Filter boilerplate link text via configurable excluded phrases 3 天之前
structs Group headlines into stories and filter by source count and date 3 天之前
tests Filter boilerplate link text via configurable excluded phrases 3 天之前
.gitignore Group headlines into stories and filter by source count and date 3 天之前
README.md Filter boilerplate link text via configurable excluded phrases 3 天之前
anya-updates.bundle Add browser-like request headers for resilience 3 天之前
main.py Filter boilerplate link text via configurable excluded phrases 3 天之前

README.md

Anya

A news headline aggregator that clusters near-duplicate headlines into stories and reports only the stories covered by a minimum number of distinct outlets.

Instead of printing every pairwise “duplicate found” alert, Anya groups matching headlines transitively (union-find over cosine similarity) and answers the useful question: which stories are multiple independent sources reporting right now?

How it works

  1. Load source URLs and stopwords from resources/.
  2. Fetch each source and extract candidate headline text from <a>/<span> tags, associating each headline with the nearest <time> publication timestamp.
  3. Normalize each headline (lowercase, strip punctuation, remove stopwords).
  4. Cluster headlines into stories using pairwise cosine similarity, linking matches transitively so a chain of near-duplicates collapses into one story.
  5. Filter & report stories with at least --min-sources distinct outlets, optionally restricted to a date window.

Requirements

  • Python 3.11+ (developed and tested on 3.14)
  • requests
pip install requests

Setup

git clone https://ikibani.com/jbell730/anya.git
cd anya
pip install requests

Run from the project root — the resource paths are relative (./resources/...).

Configuration

  • resources/sources.txt — one news source URL per line.
  • resources/stopwords.txt — one stopword per line.
  • resources/excluded_phrases.txt — boilerplate link text to ignore (see below).
  • Defaults live at the top of main.py (SIMILARITY_THRESHOLD = 0.75, MIN_SOURCES = 2) and can be overridden on the command line.

Usage

python main.py [options]
Option Description Default
--min-sources N Only output stories reported by at least N distinct sources 2
--threshold T Cosine similarity used to consider two headlines the same story 0.75
--since YYYY-MM-DD Only consider headlines published on or after this date (none)
--until YYYY-MM-DD Only consider headlines published on or before this date (none)
--verbose Enable debug logging off

Examples

# Stories reported by 3+ distinct outlets
python main.py --min-sources 3

# Stories from the last week, need 2+ outlets
python main.py --since 2026-09-07

# A specific range with a stricter similarity threshold
python main.py --since 2026-09-01 --until 2026-09-14 --threshold 0.8

Sample output:

News stories covered by at least 2 distinct sources (published on or after 2026-09-13)
============================================================

[3 source(s)] Federal Reserve holds interest rates steady (2026-09-14)
    cnn.com, foxnews.com, reuters.com

[2 source(s)] Senate passes major infrastructure bill (2026-09-14)
    npr.org, nbcnews.com

Design notes

  • A “source” is a distinct domain, not a distinct URL. www.cnn.com/us and cnn.com/politics both normalize to cnn.com, so one outlet counts once even when listed under multiple sections.
  • Date windowing drops undated headlines. When --since or --until is set, any headline whose page carries no parseable timestamp is excluded because its recency can’t be established (the count is logged). Without a date flag, everything is included.
  • Publication dates are drawn from <time> elements (their datetime/title attributes or inner text) and associated with headlines in document order; the association resets at each <article>/<li> boundary so undated headlines don’t inherit a neighboring story’s date.
  • Boilerplate link text is filtered. Navigation/footer/legal links like “skip to content” or “your privacy choices” are dropped by matching against excluded_phrases.txt (case- and punctuation-insensitive; an entry matches when it appears anywhere in the link text as a run of words, so privacy choices also catches “Your Privacy Choices”). Extend the file rather than editing code.

Tests

python -m unittest discover -s tests -p 'test_*.py'

Project structure

anya/
├── main.py                 # CLI entry point and orchestration
├── services/
│   ├── headlines.py        # fetch + parse headlines (and timestamps)
│   ├── normalization.py    # stopword loading and headline normalization
│   ├── similarity.py       # cosine similarity over token lists
│   ├── sources.py          # load source URLs
│   ├── stories.py          # cluster headlines into stories
│   └── dates.py            # date parsing and window filtering
├── structs/
│   ├── headline.py         # Headline model (text, domain, published_at)
│   └── story.py            # Story model (sources, representative, latest_date)
├── resources/
│   ├── sources.txt            # one source URL per line
│   ├── stopwords.txt          # one stopword per line
│   └── excluded_phrases.txt   # boilerplate link text to ignore
└── tests/                  # unit tests