A developer operating under the alias hashfunction has published a dataset containing metadata for approximately 5.6 billion public TikTok videos on the open-source AI repository Hugging Face. The data, hosted under the datasocial account, spans from July 2014 through October 2026 and is available for free download under a non-commercial license. The dataset consists of 460 GB of Parquet files organized by month, including fields such as captions, hashtags, sound IDs, on-screen text, TikTok Shop product identifiers, and engagement metrics like views, likes, comments, shares, saves, and downloads. It also incorporates TikTok’s internal labels regarding whether content is flagged as AI-generated or restricted from the For You page.

The acquisition method conflicts with TikTok’s terms of service, specifically Section 3.4 of its U.S. agreement, which prohibits automated data extraction without written approval. According to DataSocial’s documentation, the data was collected via TikTok’s private mobile API using generated device identities that mimic Android phones, reverse-engineered request signatures, and spoofed TLS handshakes, all without requiring user login. While the Hugging Face version is gated only by a non-commercial license, commercial access, creator profiles, and daily updates are routed through datasocial.ai, where the scraper’s source code is sold for $1,699. This release occurs amid ongoing legal disputes over data scraping, including Reddit’s lawsuit against Perplexity and other firms alleging industrial-scale harvesting of content for AI training.