I spent the last year gathering and normalizing as many resources for
Usenet messages and built a search engine over public Usenet
discussions going back to 1981. Posting here because this seemed like
the one group that would actually care.
https://www.usenet-rewind.com/
Since Google Groups stopped indexing new Usenet content a
while back, and search over its historical archive has always been
rough- I wanted something that treats this as a real research
archive -- full-text search, date-range/filtering,etc and most
importantly as much Usenet data as possible going back as far as I could >find.
The index currently holds (to date) roughly 980 million messages and is
still growing. Sources are archival backups, data donations, and ongoing >crawls of several thousand still-active public NNTP servers. Binary and >yEnc-encoded content is stripped where possible to keep the index
focused on text.
The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu >22.04.5 LTS & custom Python scripts, all on nvme disk.
This is public archival content, same legal basis as DejaNews and
Google Groups before it. There's a removal process for anyone who
wants their own posts taken down, and author contact info is masked
by default.
Over in alt.folklore.computers
Craig Stadler wrote:
I spent the last year gathering and normalizing as many resources for
Usenet messages and built a search engine over public Usenet
discussions going back to 1981. Posting here because this seemed like
the one group that would actually care.
https://www.usenet-rewind.com/
Since Google Groups stopped indexing new Usenet content a
while back, and search over its historical archive has always been
rough- I wanted something that treats this as a real research
archive -- full-text search, date-range/filtering,etc and most
importantly as much Usenet data as possible going back as far as I could
find.
The index currently holds (to date) roughly 980 million messages and is
still growing. Sources are archival backups, data donations, and ongoing
crawls of several thousand still-active public NNTP servers. Binary and
yEnc-encoded content is stripped where possible to keep the index
focused on text.
The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu
22.04.5 LTS & custom Python scripts, all on nvme disk.
This is public archival content, same legal basis as DejaNews and
Google Groups before it. There's a removal process for anyone who
wants their own posts taken down, and author contact info is masked
by default.
I thought that some might find it interesting.
| Sysop: | Jacob Catayoc |
|---|---|
| Location: | Pasay City, Metro Manila, Philippines |
| Users: | 4 |
| Nodes: | 4 (0 / 4) |
| Uptime: | 497101:01:44 |
| Calls: | 182 |
| Files: | 744 |
| D/L today: |
41 files (6,007K bytes) |
| Messages: | 73,594 |