選択できるのは25トピックまでです。 トピックは、先頭が英数字で、英数字とダッシュ('-')を使用した35文字以内のものにしてください。
Jared Bell c765b53169 Expand excluded phrases list to filter additional boilerplate 3日前
.idea first commit 3日前
resources Expand excluded phrases list to filter additional boilerplate 3日前
services Filter boilerplate link text via configurable excluded phrases 3日前
structs Group headlines into stories and filter by source count and date 3日前
tests Filter boilerplate link text via configurable excluded phrases 3日前
.gitignore Group headlines into stories and filter by source count and date 3日前
README.md Filter boilerplate link text via configurable excluded phrases 3日前
anya-updates.bundle Add browser-like request headers for resilience 3日前
main.py Filter boilerplate link text via configurable excluded phrases 3日前

README.md

Anya

A news headline aggregator that clusters near-duplicate headlines into stories and reports only the stories covered by a minimum number of distinct outlets.

Instead of printing every pairwise “duplicate found” alert, Anya groups matching headlines transitively (union-find over cosine similarity) and answers the useful question: which stories are multiple independent sources reporting right now?

How it works

  1. Load source URLs and stopwords from resources/.
  2. Fetch each source and extract candidate headline text from <a>/<span> tags, associating each headline with the nearest <time> publication timestamp.
  3. Normalize each headline (lowercase, strip punctuation, remove stopwords).
  4. Cluster headlines into stories using pairwise cosine similarity, linking matches transitively so a chain of near-duplicates collapses into one story.
  5. Filter & report stories with at least --min-sources distinct outlets, optionally restricted to a date window.

Requirements

  • Python 3.11+ (developed and tested on 3.14)
  • requests
pip install requests

Setup

git clone https://ikibani.com/jbell730/anya.git
cd anya
pip install requests

Run from the project root — the resource paths are relative (./resources/...).

Configuration

  • resources/sources.txt — one news source URL per line.
  • resources/stopwords.txt — one stopword per line.
  • resources/excluded_phrases.txt — boilerplate link text to ignore (see below).
  • Defaults live at the top of main.py (SIMILARITY_THRESHOLD = 0.75, MIN_SOURCES = 2) and can be overridden on the command line.

Usage

python main.py [options]
Option Description Default
--min-sources N Only output stories reported by at least N distinct sources 2
--threshold T Cosine similarity used to consider two headlines the same story 0.75
--since YYYY-MM-DD Only consider headlines published on or after this date (none)
--until YYYY-MM-DD Only consider headlines published on or before this date (none)
--verbose Enable debug logging off

Examples

# Stories reported by 3+ distinct outlets
python main.py --min-sources 3

# Stories from the last week, need 2+ outlets
python main.py --since 2026-09-07

# A specific range with a stricter similarity threshold
python main.py --since 2026-09-01 --until 2026-09-14 --threshold 0.8

Sample output:

News stories covered by at least 2 distinct sources (published on or after 2026-09-13)
============================================================

[3 source(s)] Federal Reserve holds interest rates steady (2026-09-14)
    cnn.com, foxnews.com, reuters.com

[2 source(s)] Senate passes major infrastructure bill (2026-09-14)
    npr.org, nbcnews.com

Design notes

  • A “source” is a distinct domain, not a distinct URL. www.cnn.com/us and cnn.com/politics both normalize to cnn.com, so one outlet counts once even when listed under multiple sections.
  • Date windowing drops undated headlines. When --since or --until is set, any headline whose page carries no parseable timestamp is excluded because its recency can’t be established (the count is logged). Without a date flag, everything is included.
  • Publication dates are drawn from <time> elements (their datetime/title attributes or inner text) and associated with headlines in document order; the association resets at each <article>/<li> boundary so undated headlines don’t inherit a neighboring story’s date.
  • Boilerplate link text is filtered. Navigation/footer/legal links like “skip to content” or “your privacy choices” are dropped by matching against excluded_phrases.txt (case- and punctuation-insensitive; an entry matches when it appears anywhere in the link text as a run of words, so privacy choices also catches “Your Privacy Choices”). Extend the file rather than editing code.

Tests

python -m unittest discover -s tests -p 'test_*.py'

Project structure

anya/
├── main.py                 # CLI entry point and orchestration
├── services/
│   ├── headlines.py        # fetch + parse headlines (and timestamps)
│   ├── normalization.py    # stopword loading and headline normalization
│   ├── similarity.py       # cosine similarity over token lists
│   ├── sources.py          # load source URLs
│   ├── stories.py          # cluster headlines into stories
│   └── dates.py            # date parsing and window filtering
├── structs/
│   ├── headline.py         # Headline model (text, domain, published_at)
│   └── story.py            # Story model (sources, representative, latest_date)
├── resources/
│   ├── sources.txt            # one source URL per line
│   ├── stopwords.txt          # one stopword per line
│   └── excluded_phrases.txt   # boilerplate link text to ignore
└── tests/                  # unit tests