All Licenses

Custom Research & Evaluation License

Added 2026-05-12
DataCustomHuggingFaceProprietary

Full Text

USENET CORPUS 1980–2013 Custom Research & Evaluation License Copyright (c) 2024. All rights reserved. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ SCOPE This license governs use of the sample files made available in this repository. The sample files consist of approximately 65,000 posts drawn from the full Usenet Corpus 1980–2013. The full corpus (103.1 billion tokens, 408 million posts) is governed by separate written licensing agreements available upon request. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ PERMITTED USES — SAMPLE FILES The sample files in this repository are freely available and permitted for the following uses without requiring a separate agreement: 1. Evaluation and inspection of dataset quality, schema, and content for the purpose of assessing suitability for research or commercial use. 2. Non-commercial academic research and publication, provided that appropriate attribution is included (see Citation section in README.md). 3. Personal, non-commercial experimentation and exploration, including non-commercial fine-tuning and model training experiments using the sample files only. 4. Educational use in academic or research contexts. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ PROHIBITED USES — WITHOUT A SEPARATE WRITTEN AGREEMENT The following uses are expressly prohibited without a separate written licensing agreement: 1. Incorporation of the full corpus or any substantial portion thereof into any AI model training pipeline, whether commercial or non-commercial. 2. Commercial AI model training using the sample files. 3. Redistribution, resale, sublicensing, or transfer of the full corpus or any derivative thereof to any third party. 4. Use of the full corpus in any commercial product, service, or offering. 5. Bulk downloading, scraping, or systematic extraction of content beyond the sample files provided in this repository. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ FULL CORPUS ACCESS The full Usenet Corpus 1980–2013 (103.1 billion tokens, 408 million posts) is available for licensing under separate terms. Licensing arrangements may include: - Data partnership agreements with AI laboratories - Research institution licensing for academic use - Commercial licensing for inclusion in training pipelines - API-based query access arrangements ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ DERIVATIVE WORK NOTICE This dataset is a derivative work produced through substantial original effort applied to public Usenet archives, including: - Conversion from MBOX to structured JSONL format - Deduplication at scale - Removal of binary-encoded content via MIME header inspection and pattern detection - Sanitization and sensitive content removal - Email address and PII redaction - Message-ID anonymization via SHA-256 hashing - Language detection at record level - Compression and indexing The underlying posts are public Usenet communications. The compiled, processed corpus represents original work protected separately from the underlying source material, consistent with standard database and compilation rights frameworks. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ DISCLAIMER THIS DATASET IS PROVIDED "AS IS" WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED. THE LICENSOR MAKES NO REPRESENTATIONS REGARDING THE ACCURACY, COMPLETENESS, OR FITNESS FOR ANY PARTICULAR PURPOSE OF THE DATASET OR ITS CONTENTS. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ CONTACT To inquire about licensing, data partnerships, or access arrangements for the full corpus: usenetoverlord@gmail.com ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━