43B Tokens Released
Key point
A 590GB SEC filings EDGAR dataset has been released on Hugging Face.
Details
Datamule, Teraflop AI, and Eventual have released the SEC-EDGAR dataset on Hugging Face.
The dataset is 590GB in size, containing 8 million samples and 43 billion tokens, broadly covering major filings from SEC EDGAR. It targets corporate filing documents such as 10-Q and 10-K, and the key point is that it provides free, open access to data for which some paid unofficial APIs charge hundreds of dollars per month.
The collection was carried out using datamule-python and the official datamule API. The authors explain that due to EDGAR's rate limit of 10 requests per second and network overhead, large-scale crawling at the 8-million-item scale takes over 10 days, emphasizing the need for open-source tools for large-scale collection and processing.
- Distribution location: TeraflopAI/SEC-EDGAR on Hugging Face
- Data source: SEC's EDGAR filing database
- Point: Providing free open data instead of expensive commercial access
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.