A developer operating under the alias hashfunction has published a dataset containing metadata for approximately 5.6 billion public TikTok videos on the open-source AI repository Hugging Face. The data, hosted under the datasocial account, spans from July 2014 through October 2026 and is available for free download under a non-commercial license. The dataset consists of 460 GB of Parquet files organized by month, including fields such as captions, hashtags, sound IDs, on-screen text, TikTok Shop product identifiers, and engagement metrics like views, likes, comments, shares, saves, and downloads. It also incorporates TikTok’s internal labels regarding whether content is flagged as AI-generated or restricted from the For You page.
The acquisition method conflicts with TikTok’s terms of service, specifically Section 3.4 of its U.S. agreement, which prohibits automated data extraction without written approval. According to DataSocial’s documentation, the data was collected via TikTok’s private mobile API using generated device identities that mimic Android phones, reverse-engineered request signatures, and spoofed TLS handshakes, all without requiring user login. While the Hugging Face version is gated only by a non-commercial license, commercial access, creator profiles, and daily updates are routed through datasocial.ai, where the scraper’s source code is sold for $1,699. This release occurs amid ongoing legal disputes over data scraping, including Reddit’s lawsuit against Perplexity and other firms alleging industrial-scale harvesting of content for AI training.
The publication of this massive metadata set highlights the growing tension between open-source AI development needs and proprietary platform controls. By providing raw material for models predicting virality and commerce trends, the dataset circumvents TikTok’s official Research Tools, which are limited to approved academic institutions in specific jurisdictions. This unauthorized access undermines the platform’s ability to control how its ecosystem data is utilized, potentially exposing sensitive patterns in user behavior and content moderation algorithms to external developers who have not undergone compliance vetting.
From a regulatory and operational risk perspective, this incident underscores the fragility of current anti-scraping enforcement mechanisms. The use of sophisticated techniques to bypass mobile API protections suggests that traditional terms-of-service violations may be insufficient deterrents against well-resourced actors. Furthermore, the dual-track distribution model—free non-commercial data versus paid commercial access—creates a gray area for liability, as it incentivizes widespread dissemination while monetizing deeper insights. Legal precedents involving similar scraping cases remain unresolved, leaving platforms and developers navigating an uncertain landscape regarding intellectual property rights and data ownership.


