ソースを参照

Add README with setup, usage, and design notes

master
Jared Bell 3日前
コミット
975bc68a61
1個のファイルの変更126行の追加0行の削除
  1. +126
    -0
      README.md

+ 126
- 0
README.md ファイルの表示

@@ -0,0 +1,126 @@
# Anya

A news headline aggregator that clusters near-duplicate headlines into **stories**
and reports only the stories covered by a minimum number of distinct outlets.

Instead of printing every pairwise "duplicate found" alert, Anya groups matching
headlines transitively (union-find over cosine similarity) and answers the useful
question: *which stories are multiple independent sources reporting right now?*

## How it works

1. **Load** source URLs and stopwords from `resources/`.
2. **Fetch** each source and extract candidate headline text from `<a>`/`<span>`
tags, associating each headline with the nearest `<time>` publication timestamp.
3. **Normalize** each headline (lowercase, strip punctuation, remove stopwords).
4. **Cluster** headlines into stories using pairwise cosine similarity, linking
matches transitively so a chain of near-duplicates collapses into one story.
5. **Filter & report** stories with at least `--min-sources` distinct outlets,
optionally restricted to a date window.

## Requirements

- Python 3.11+ (developed and tested on 3.14)
- `requests`

```bash
pip install requests
```

## Setup

```bash
git clone https://ikibani.com/jbell730/anya.git
cd anya
pip install requests
```

Run from the project root — the resource paths are relative (`./resources/...`).

## Configuration

- `resources/sources.txt` — one news source URL per line.
- `resources/stopwords.txt` — one stopword per line.
- Defaults live at the top of `main.py` (`SIMILARITY_THRESHOLD = 0.75`,
`MIN_SOURCES = 2`) and can be overridden on the command line.

## Usage

```bash
python main.py [options]
```

| Option | Description | Default |
| --- | --- | --- |
| `--min-sources N` | Only output stories reported by at least N distinct sources | `2` |
| `--threshold T` | Cosine similarity used to consider two headlines the same story | `0.75` |
| `--since YYYY-MM-DD` | Only consider headlines published on or after this date | *(none)* |
| `--until YYYY-MM-DD` | Only consider headlines published on or before this date | *(none)* |
| `--verbose` | Enable debug logging | off |

### Examples

```bash
# Stories reported by 3+ distinct outlets
python main.py --min-sources 3

# Stories from the last week, need 2+ outlets
python main.py --since 2026-09-07

# A specific range with a stricter similarity threshold
python main.py --since 2026-09-01 --until 2026-09-14 --threshold 0.8
```

Sample output:

```
News stories covered by at least 2 distinct sources (published on or after 2026-09-13)
============================================================

[3 source(s)] Federal Reserve holds interest rates steady (2026-09-14)
cnn.com, foxnews.com, reuters.com

[2 source(s)] Senate passes major infrastructure bill (2026-09-14)
npr.org, nbcnews.com
```

## Design notes

- **A "source" is a distinct domain**, not a distinct URL. `www.cnn.com/us` and
`cnn.com/politics` both normalize to `cnn.com`, so one outlet counts once even
when listed under multiple sections.
- **Date windowing drops undated headlines.** When `--since` or `--until` is set,
any headline whose page carries no parseable timestamp is excluded because its
recency can't be established (the count is logged). Without a date flag,
everything is included.
- **Publication dates** are drawn from `<time>` elements (their `datetime`/`title`
attributes or inner text) and associated with headlines in document order; the
association resets at each `<article>`/`<li>` boundary so undated headlines don't
inherit a neighboring story's date.

## Tests

```bash
python -m unittest discover -s tests -p 'test_*.py'
```

## Project structure

```
anya/
├── main.py # CLI entry point and orchestration
├── services/
│ ├── headlines.py # fetch + parse headlines (and timestamps)
│ ├── normalization.py # stopword loading and headline normalization
│ ├── similarity.py # cosine similarity over token lists
│ ├── sources.py # load source URLs
│ ├── stories.py # cluster headlines into stories
│ └── dates.py # date parsing and window filtering
├── structs/
│ ├── headline.py # Headline model (text, domain, published_at)
│ └── story.py # Story model (sources, representative, latest_date)
├── resources/
│ ├── sources.txt # one source URL per line
│ └── stopwords.txt # one stopword per line
└── tests/ # unit tests
```

読み込み中…
キャンセル
保存