| @@ -0,0 +1,126 @@ | |||
| # Anya | |||
| A news headline aggregator that clusters near-duplicate headlines into **stories** | |||
| and reports only the stories covered by a minimum number of distinct outlets. | |||
| Instead of printing every pairwise "duplicate found" alert, Anya groups matching | |||
| headlines transitively (union-find over cosine similarity) and answers the useful | |||
| question: *which stories are multiple independent sources reporting right now?* | |||
| ## How it works | |||
| 1. **Load** source URLs and stopwords from `resources/`. | |||
| 2. **Fetch** each source and extract candidate headline text from `<a>`/`<span>` | |||
| tags, associating each headline with the nearest `<time>` publication timestamp. | |||
| 3. **Normalize** each headline (lowercase, strip punctuation, remove stopwords). | |||
| 4. **Cluster** headlines into stories using pairwise cosine similarity, linking | |||
| matches transitively so a chain of near-duplicates collapses into one story. | |||
| 5. **Filter & report** stories with at least `--min-sources` distinct outlets, | |||
| optionally restricted to a date window. | |||
| ## Requirements | |||
| - Python 3.11+ (developed and tested on 3.14) | |||
| - `requests` | |||
| ```bash | |||
| pip install requests | |||
| ``` | |||
| ## Setup | |||
| ```bash | |||
| git clone https://ikibani.com/jbell730/anya.git | |||
| cd anya | |||
| pip install requests | |||
| ``` | |||
| Run from the project root — the resource paths are relative (`./resources/...`). | |||
| ## Configuration | |||
| - `resources/sources.txt` — one news source URL per line. | |||
| - `resources/stopwords.txt` — one stopword per line. | |||
| - Defaults live at the top of `main.py` (`SIMILARITY_THRESHOLD = 0.75`, | |||
| `MIN_SOURCES = 2`) and can be overridden on the command line. | |||
| ## Usage | |||
| ```bash | |||
| python main.py [options] | |||
| ``` | |||
| | Option | Description | Default | | |||
| | --- | --- | --- | | |||
| | `--min-sources N` | Only output stories reported by at least N distinct sources | `2` | | |||
| | `--threshold T` | Cosine similarity used to consider two headlines the same story | `0.75` | | |||
| | `--since YYYY-MM-DD` | Only consider headlines published on or after this date | *(none)* | | |||
| | `--until YYYY-MM-DD` | Only consider headlines published on or before this date | *(none)* | | |||
| | `--verbose` | Enable debug logging | off | | |||
| ### Examples | |||
| ```bash | |||
| # Stories reported by 3+ distinct outlets | |||
| python main.py --min-sources 3 | |||
| # Stories from the last week, need 2+ outlets | |||
| python main.py --since 2026-09-07 | |||
| # A specific range with a stricter similarity threshold | |||
| python main.py --since 2026-09-01 --until 2026-09-14 --threshold 0.8 | |||
| ``` | |||
| Sample output: | |||
| ``` | |||
| News stories covered by at least 2 distinct sources (published on or after 2026-09-13) | |||
| ============================================================ | |||
| [3 source(s)] Federal Reserve holds interest rates steady (2026-09-14) | |||
| cnn.com, foxnews.com, reuters.com | |||
| [2 source(s)] Senate passes major infrastructure bill (2026-09-14) | |||
| npr.org, nbcnews.com | |||
| ``` | |||
| ## Design notes | |||
| - **A "source" is a distinct domain**, not a distinct URL. `www.cnn.com/us` and | |||
| `cnn.com/politics` both normalize to `cnn.com`, so one outlet counts once even | |||
| when listed under multiple sections. | |||
| - **Date windowing drops undated headlines.** When `--since` or `--until` is set, | |||
| any headline whose page carries no parseable timestamp is excluded because its | |||
| recency can't be established (the count is logged). Without a date flag, | |||
| everything is included. | |||
| - **Publication dates** are drawn from `<time>` elements (their `datetime`/`title` | |||
| attributes or inner text) and associated with headlines in document order; the | |||
| association resets at each `<article>`/`<li>` boundary so undated headlines don't | |||
| inherit a neighboring story's date. | |||
| ## Tests | |||
| ```bash | |||
| python -m unittest discover -s tests -p 'test_*.py' | |||
| ``` | |||
| ## Project structure | |||
| ``` | |||
| anya/ | |||
| ├── main.py # CLI entry point and orchestration | |||
| ├── services/ | |||
| │ ├── headlines.py # fetch + parse headlines (and timestamps) | |||
| │ ├── normalization.py # stopword loading and headline normalization | |||
| │ ├── similarity.py # cosine similarity over token lists | |||
| │ ├── sources.py # load source URLs | |||
| │ ├── stories.py # cluster headlines into stories | |||
| │ └── dates.py # date parsing and window filtering | |||
| ├── structs/ | |||
| │ ├── headline.py # Headline model (text, domain, published_at) | |||
| │ └── story.py # Story model (sources, representative, latest_date) | |||
| ├── resources/ | |||
| │ ├── sources.txt # one source URL per line | |||
| │ └── stopwords.txt # one stopword per line | |||
| └── tests/ # unit tests | |||
| ``` | |||