← All Licenses
Custom Research & Evaluation License
Added 2026-05-12
DataCustomHuggingFaceProprietary
Full Text
USENET CORPUS 1980–2013
Custom Research & Evaluation License
Copyright (c) 2024. All rights reserved.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SCOPE
This license governs use of the sample files made available in
this repository. The sample files consist of approximately
65,000 posts drawn from the full Usenet Corpus 1980–2013.
The full corpus (103.1 billion tokens, 408 million posts) is
governed by separate written licensing agreements available
upon request.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
PERMITTED USES — SAMPLE FILES
The sample files in this repository are freely available and
permitted for the following uses without requiring a separate
agreement:
1. Evaluation and inspection of dataset quality, schema, and
content for the purpose of assessing suitability for
research or commercial use.
2. Non-commercial academic research and publication,
provided that appropriate attribution is included
(see Citation section in README.md).
3. Personal, non-commercial experimentation and exploration,
including non-commercial fine-tuning and model training
experiments using the sample files only.
4. Educational use in academic or research contexts.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
PROHIBITED USES — WITHOUT A SEPARATE WRITTEN AGREEMENT
The following uses are expressly prohibited without a separate
written licensing agreement:
1. Incorporation of the full corpus or any substantial
portion thereof into any AI model training pipeline,
whether commercial or non-commercial.
2. Commercial AI model training using the sample files.
3. Redistribution, resale, sublicensing, or transfer of
the full corpus or any derivative thereof to any
third party.
4. Use of the full corpus in any commercial product,
service, or offering.
5. Bulk downloading, scraping, or systematic extraction
of content beyond the sample files provided in this
repository.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
FULL CORPUS ACCESS
The full Usenet Corpus 1980–2013 (103.1 billion tokens,
408 million posts) is available for licensing under separate
terms. Licensing arrangements may include:
- Data partnership agreements with AI laboratories
- Research institution licensing for academic use
- Commercial licensing for inclusion in training pipelines
- API-based query access arrangements
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
DERIVATIVE WORK NOTICE
This dataset is a derivative work produced through substantial
original effort applied to public Usenet archives, including:
- Conversion from MBOX to structured JSONL format
- Deduplication at scale
- Removal of binary-encoded content via MIME header
inspection and pattern detection
- Sanitization and sensitive content removal
- Email address and PII redaction
- Message-ID anonymization via SHA-256 hashing
- Language detection at record level
- Compression and indexing
The underlying posts are public Usenet communications. The
compiled, processed corpus represents original work protected
separately from the underlying source material, consistent
with standard database and compilation rights frameworks.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
DISCLAIMER
THIS DATASET IS PROVIDED "AS IS" WITHOUT WARRANTY OF ANY KIND,
EXPRESS OR IMPLIED. THE LICENSOR MAKES NO REPRESENTATIONS
REGARDING THE ACCURACY, COMPLETENESS, OR FITNESS FOR ANY
PARTICULAR PURPOSE OF THE DATASET OR ITS CONTENTS.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
CONTACT
To inquire about licensing, data partnerships, or access
arrangements for the full corpus:
usenetoverlord@gmail.com
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━