Move the hard-coded ALIASES dict into resources/aliases.toml (loaded with
stdlib tomllib), expand it with agencies, international bodies, and country
names, and fix two correctness bugs in the naive replace():
- Whole-word matching so 'united states' no longer collapses inside
'united statesman', and 'inflation rate' no longer mangles 'inflation rates'.
- Longest-phrase-first application so 'president of the united states'
resolves to 'potus' before 'united states' fires.
- Correct a stopword collision: 'united nations' -> 'un' was silently dropped
because 'un' is a stopword; it now maps to 'unitednations'.
Aliases are threaded through normalize_headline -> is_headline ->
_build_headlines -> prepare_headlines and loaded in main.py. Includes unit
tests and README coverage.
Add resources/excluded_phrases.txt (skip links, privacy/consent, legal
boilerplate) and match it case- and punctuation-insensitively as a
contiguous run of words. Phrases live in a data file so they can be
extended without touching code.