|
|
hace 1 día | |
|---|---|---|
| .idea | hace 3 días | |
| resources | hace 1 día | |
| services | hace 1 día | |
| structs | hace 3 días | |
| tests | hace 1 día | |
| .gitignore | hace 3 días | |
| README.md | hace 1 día | |
| anya-updates.bundle | hace 3 días | |
| main.py | hace 1 día | |
A news headline aggregator that clusters near-duplicate headlines into stories and reports only the stories covered by a minimum number of distinct outlets.
Instead of printing every pairwise “duplicate found” alert, Anya groups matching headlines transitively (union-find over cosine similarity) and answers the useful question: which stories are multiple independent sources reporting right now?
resources/.<a>/<span> text, associating each headline with
the nearest <time> publication timestamp.--min-sources distinct outlets,
optionally restricted to a date window.requestspip install requests
git clone https://ikibani.com/jbell730/anya.git
cd anya
pip install requests
Run from the project root — the resource paths are relative (./resources/...).
resources/sources.txt — one source URL per line (HTML pages or RSS/Atom
feeds; # comments and blank lines are ignored).resources/stopwords.txt — one stopword per line.resources/excluded_phrases.txt — boilerplate link text to ignore (see below).resources/aliases.toml — full-name → short-form aliases (see below).main.py (SIMILARITY_THRESHOLD = 0.64,
MIN_SOURCES = 3) and can be overridden on the command line.python main.py [options]
| Option | Description | Default |
|---|---|---|
--min-sources N |
Only output stories reported by at least N distinct sources | 3 |
--threshold T |
Cosine similarity used to consider two headlines the same story | 0.64 |
--since YYYY-MM-DD |
Only consider headlines published on or after this date | (none) |
--until YYYY-MM-DD |
Only consider headlines published on or before this date | (none) |
--verbose |
Enable debug logging | off |
# Stories reported by 3+ distinct outlets
python main.py --min-sources 3
# Stories from the last week, need 2+ outlets
python main.py --since 2026-09-07
# A specific range with a stricter similarity threshold
python main.py --since 2026-09-01 --until 2026-09-14 --threshold 0.8
Sample output:
News stories covered by at least 3 distinct sources (published on or after 2026-09-13)
============================================================
[3 source(s)] Federal Reserve holds interest rates steady (2026-09-14)
cnn.com, foxnews.com, reuters.com
[2 source(s)] Senate passes major infrastructure bill (2026-09-14)
npr.org, nbcnews.com
www.cnn.com/us and
cnn.com/politics both normalize to cnn.com, so one outlet counts once even
when listed under multiple sections. Feed/redirect subdomains (feeds., rss.,
moxie.) are also stripped, so a feed and an HTML page from the same outlet
still collapse to one source.sources.txt.__NEXT_DATA__ SSR state —
Anya parses that directly and never executes JavaScript. This is far lighter than
a headless browser, at the cost of some per-site variation in the JSON shape.--since or --until is set,
any headline whose page carries no parseable timestamp is excluded because its
recency can’t be established (the count is logged). Without a date flag,
everything is included.<time> elements (their datetime/title
attributes or inner text) and associated with headlines in document order; the
association resets at each <article>/<li> boundary so undated headlines don’t
inherit a neighboring story’s date.excluded_phrases.txt (case- and punctuation-insensitive; an entry matches when
it appears anywhere in the link text as a run of words, so privacy choices
also catches “Your Privacy Choices”). Extend the file rather than editing code.aliases.toml maps full names to a shared
short form so “Federal Reserve” and “the Fed” normalize to the same token and
cluster. Matching is whole-word only (“united states” won’t match inside
“united statesman”) and longest-phrase-first (“president of the united states”
becomes “potus” before “united states” can fire). The canonical form must not be
a stopword — e.g. “united nations” uses unitednations, not un, because un
is stripped during normalization. Extend the file rather than editing code.python -m unittest discover -s tests -p 'test_*.py'
anya/
├── main.py # CLI entry point and orchestration
├── services/
│ ├── headlines.py # fetch + parse headlines (and timestamps)
│ ├── feeds.py # RSS/Atom feed detection and parsing
│ ├── ssr.py # JSON-LD / Next.js embedded-JSON extraction
│ ├── normalization.py # stopword/alias/phrase loading and headline normalization
│ ├── similarity.py # cosine similarity over token lists
│ ├── sources.py # load source URLs
│ ├── stories.py # cluster headlines into stories
│ └── dates.py # date parsing and window filtering
├── structs/
│ ├── headline.py # Headline model (text, domain, published_at)
│ └── story.py # Story model (sources, representative, latest_date)
├── resources/
│ ├── sources.txt # one source URL per line
│ ├── stopwords.txt # one stopword per line
│ ├── aliases.toml # full-name → short-form aliases
│ └── excluded_phrases.txt # boilerplate link text to ignore
└── tests/ # unit tests