_id large_stringlengths 24 24 | id large_stringlengths 5 123 | author large_stringlengths 2 42 | cardData large_stringlengths 2 1.09M ⌀ | disabled bool 1
class | gated large_stringclasses 3
values | lastModified timestamp[us]date 2021-02-05 16:03:35 2026-07-25 13:22:28 | likes int64 0 9.77k | trendingScore float64 0 123 | private bool 1
class | sha large_stringlengths 40 40 | description large_stringlengths 0 6.67k ⌀ | downloads int64 0 8.55M | downloadsAllTime int64 0 143M | mainSize float64 0 306,846B ⌀ | tags listlengths 1 7.92k | createdAt timestamp[us]date 2022-03-02 23:29:22 2026-07-25 13:20:19 | paperswithcode_id large_stringclasses 712
values | citation large_stringlengths 0 10.7k ⌀ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
6a615c95fb10b1093e0ea9ed | HuggingFaceCode/stack-v3-train | HuggingFaceCode | {"thumbnail": "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png", "annotations_creators": [], "language_creators": ["crowdsourced", "expert-generated"], "language": ["code"], "license": ["odc-by"], "multilinguality": ["multilingual"], "size_categories": ["100M<n<1B"], "sourc... | false | False | 2026-07-24T18:36:04 | 128 | 123 | false | de81e3ca7151fc8b8769dbc2dfc0af5ac92b6d1e |
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled dir... | 35,376 | 35,376 | 4,711,370,231,977 | [
"task_categories:text-generation",
"language_creators:crowdsourced",
"language_creators:expert-generated",
"multilinguality:multilingual",
"language:code",
"license:odc-by",
"size_categories:100M<n<1B",
"format:parquet",
"modality:tabular",
"modality:text",
"library:datasets",
"library:dask",
... | 2026-07-23T00:13:09 | null | null |
6a4e1fe2df56b09d5f449aa8 | SupraLabs/reasoning-corpus-4K-5M-v1 | SupraLabs | {"license": "apache-2.0", "task_categories": ["text-generation"], "language": ["en"], "tags": ["reasoning", "CoT", "code", "agentic", "thinking", "think", "deepseek-v4", "qwen3", "qwen3next"], "pretty_name": "Reasoning Corpus 5M", "size_categories": ["1M<n<10M"]} | false | False | 2026-07-21T02:31:33 | 107 | 79 | false | 1fc50907cffff48186f61e52a750dabec2c09435 | Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other reposito... | 2,167 | 2,167 | 68,664,454,256 | [
"task_categories:text-generation",
"language:en",
"license:apache-2.0",
"size_categories:1M<n<10M",
"format:json",
"modality:text",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissant",
"region:us",
"reasoning",
"CoT",
"code",
"agentic",
"thinking",
"think",
... | 2026-07-08T10:01:06 | null | null |
6a4cc0ac90ce9cc602189d11 | FlyRank/internship-warehouse | FlyRank | {"license": "other", "language": ["en"], "tags": ["seo", "content-performance", "data-warehouse", "tabular", "education", "flyrank-internship"], "pretty_name": "FlyRank Internship \u2014 Warehouse Star Schema (Pseudonymized, Gated)", "size_categories": ["10M<n<100M"], "extra_gated_prompt": "By requesting access you agr... | false | auto | 2026-07-07T10:02:21 | 347 | 47 | false | 50cbf7c3909d07be4d1b5906b4d09e882e5acbf2 |
FlyRank Internship — Pseudonymized Warehouse Release (v20260703)
The open-ended, warehouse-shaped dataset (~81.8M rows; daily fact
78,835,655 rows) for advanced capstone work. Star schema with salted, namespaced,
fingerprinted hash keys. Built from warehouse v2 full history (frozen snapshot,
export date ... | 5,195 | 5,195 | 1,168,719,310 | [
"language:en",
"license:other",
"size_categories:10M<n<100M",
"modality:tabular",
"modality:text",
"region:us",
"seo",
"content-performance",
"data-warehouse",
"tabular",
"education",
"flyrank-internship"
] | 2026-07-07T09:02:36 | null | null |
6a4509196c643209b19b2fc7 | Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset | Manusagents | {"license": "mit", "language": ["en", "multilingual"], "task_categories": ["text-generation", "other"], "tags": ["distillation", "instruction-tuning", "sft", "reasoning", "coding", "code-repositories", "cybersecurity", "attack", "defense", "exploit", "penetration-testing", "red-team", "blue-team", "open-source", "colle... | false | False | 2026-07-18T18:01:17 | 68 | 43 | false | f0aa1d8326d7ca5c4a01982ca8299a783bc59faf |
📖 The Open Distillation Codex
🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌
Where 73 open-source minds converge into one unified stream of intelligence
18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~7... | 7,265 | 7,265 | 76,526,135,473 | [
"task_categories:text-generation",
"task_categories:other",
"language:en",
"language:multilingual",
"license:mit",
"size_categories:10M<n<100M",
"format:json",
"modality:text",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissant",
"region:us",
"distillation",
"in... | 2026-07-01T12:33:29 | null | null |
6a292cbbe1b5c7903e6fbe30 | openbmb/UltraX-Preview | openbmb | {"language": ["en"], "license": "apache-2.0", "size_categories": ["10B<n<100B"], "task_categories": ["text-generation"], "pretty_name": "UltraX", "tags": ["llm", "pretraining", "web-corpus", "data-refinement", "programmatic-editing", "function-calling"], "configs": [{"config_name": "UltraX-FineWeb", "data_files": [{"sp... | false | False | 2026-07-17T03:02:12 | 264 | 42 | false | a88527587389fd4ab352e9ad1273f4c0a234d8df |
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
📜 Paper |
💻 Code |
🤖 Models |
📦 UltraData Collection
English |
中文
📚 Introduction
UltraX is a function-calling refinement framework for large-scale pre-training data that ad... | 6,347 | 6,353 | 486,915,481,612 | [
"task_categories:text-generation",
"language:en",
"license:apache-2.0",
"size_categories:100M<n<1B",
"format:parquet",
"modality:text",
"library:datasets",
"library:dask",
"library:polars",
"library:mlcroissant",
"arxiv:2607.08646",
"region:us",
"llm",
"pretraining",
"web-corpus",
"dat... | 2026-06-10T09:22:03 | null | null |
6a5ae883a5f7ad08ccdbda43 | greghavens/kimi-k3-coding-and-debugging-traces | greghavens | {"pretty_name": "Kimi K3 Coding, Tool Use & Instruction Following Traces", "license": "cc-by-4.0", "language": ["en"], "annotations_creators": ["machine-generated"], "task_categories": ["text-generation"], "size_categories": ["1K<n<10K"], "tags": ["traces", "code", "agentic", "tool-use", "coding-agent", "coding-agents"... | false | False | 2026-07-25T01:18:22 | 37 | 36 | false | b594ab95ab7e3964012d8c072e22e71ed81d0b5f |
Kimi K3 Coding, Tool Use & Instruction Following Traces
702 TRAJECTORIES · 4,928 TRAINING ROWS · 3 MB PARQUET · 90 MB JSONL
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Behavior-preserving instruction-following, tool... | 2,442 | 2,442 | 5,018,189 | [
"task_categories:text-generation",
"annotations_creators:machine-generated",
"language:en",
"license:cc-by-4.0",
"size_categories:1K<n<10K",
"format:parquet",
"modality:tabular",
"modality:text",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissant",
"region:us",
"t... | 2026-07-18T02:44:19 | null | null |
6a2cd0828137fb18cecbcc06 | Glint-Research/Fable-5-traces | Glint-Research | {"license": "agpl-3.0", "pretty_name": "Fable 5 Pi Agent Traces", "annotations_creators": ["machine-generated"], "language": ["en"], "size_categories": ["1K<n<10K"], "task_categories": ["text-generation"], "tags": ["agent-traces", "pi-agent", "claude-code", "fable-5", "chain-of-thought", "tool-use", "coding-agents", "s... | false | False | 2026-06-29T15:10:20 | 666 | 35 | false | e05c417852fc59fd8da758e68b352732423ca0cb |
Glint Research Dataset Card
Fable 5 Pi Agent Traces
A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation.
... | 61,491 | 90,678 | 187,507,989 | [
"task_categories:text-generation",
"annotations_creators:machine-generated",
"language:en",
"license:agpl-3.0",
"size_categories:1K<n<10K",
"format:json",
"format:agent-traces",
"modality:tabular",
"modality:text",
"library:datasets",
"library:dask",
"library:polars",
"library:mlcroissant",
... | 2026-06-13T03:37:38 | null | null |
6a437ed52e089285573dcfd3 | markov-ai/gaming-500-hours | markov-ai | {"configs": [{"config_name": "default", "data_files": [{"split": "train", "path": "metadata.jsonl"}]}]} | false | False | 2026-06-30T11:56:39 | 214 | 30 | false | 5af703f2810306e7d75eb4394ae59591f1f6e8a2 |
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and... | 41,011 | 41,011 | 1,598,371,626,719 | [
"size_categories:n<1K",
"format:json",
"modality:tabular",
"modality:text",
"modality:video",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissant",
"region:us"
] | 2026-06-30T08:31:17 | null | null |
6a60a044d3559d7ff7b5590d | r0b0tlab/qwen3.8-max-distillation-50k | r0b0tlab | {"license": "other", "task_categories": ["text-generation", "question-answering"], "language": ["en"], "tags": ["distillation", "knowledge-distillation", "reasoning", "chain-of-thought", "supervised-fine-tuning", "math", "code", "instruction-following", "tool-use", "qwen"], "size_categories": ["10K<n<100K"], "pretty_na... | false | False | 2026-07-22T11:27:58 | 29 | 29 | false | ab9f8b289423c249fc0054507f045a12efb54b1b |
Qwen3.8-Max Distillation 50K
A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation.
The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, tho... | 247 | 247 | 70,765,792 | [
"task_categories:text-generation",
"task_categories:question-answering",
"language:en",
"license:other",
"size_categories:10K<n<100K",
"format:parquet",
"format:optimized-parquet",
"modality:tabular",
"modality:text",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissa... | 2026-07-22T10:49:40 | null | null |
621ffdd236468d709f184284 | wikimedia/wikipedia | wikimedia | {"language": ["ab", "ace", "ady", "af", "alt", "am", "ami", "an", "ang", "anp", "ar", "arc", "ary", "arz", "as", "ast", "atj", "av", "avk", "awa", "ay", "az", "azb", "ba", "ban", "bar", "bbc", "bcl", "be", "bg", "bh", "bi", "bjn", "blk", "bm", "bn", "bo", "bpy", "br", "bs", "bug", "bxr", "ca", "cbk", "cdo", "ce", "ceb"... | false | False | 2024-01-09T09:40:51 | 1,330 | 24 | false | b04c8d1ceb2f5cd4588862100d08de323dccfbaa |
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the co... | 250,729 | 2,682,246 | 71,792,022,791 | [
"task_categories:text-generation",
"task_categories:fill-mask",
"task_ids:language-modeling",
"task_ids:masked-language-modeling",
"language:ab",
"language:ace",
"language:ady",
"language:af",
"language:alt",
"language:am",
"language:ami",
"language:an",
"language:ang",
"language:anp",
"... | 2022-03-02T23:29:22 | null | null |
69d185a53c023c2c9072697a | netflix/Vera-Layered-Video-Dataset | netflix | {"license": "apache-2.0", "task_categories": ["text-to-video"], "tags": ["diffusion", "layered-diffusion", "video", "layered-video-dataset", "video-editing", "video-generation"]} | false | False | 2026-07-17T01:33:49 | 51 | 18 | false | 8e0b98ee9bce66fdae345e75aa766ef7c0a04d4e |
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing
Hongkai Zheng¹²* ·
Ta-Ying Cheng² ·
Benjamin Klein² ·
Yisong Yue¹ ·
Zhuoning Yuan²†
¹California Institute of Technology ²Netflix, Inc.
*Work done... | 14,756 | 24,169 | 319,639,945,834 | [
"task_categories:text-to-video",
"license:apache-2.0",
"size_categories:10K<n<100K",
"modality:video",
"arxiv:2606.23610",
"region:us",
"diffusion",
"layered-diffusion",
"video",
"layered-video-dataset",
"video-editing",
"video-generation"
] | 2026-04-04T21:41:57 | null | null |
645e8da96320b0efe40ade7a | roneneldan/TinyStories | roneneldan | {"license": "cdla-sharing-1.0", "task_categories": ["text-generation"], "language": ["en"]} | false | False | 2024-08-12T13:27:26 | 1,094 | 17 | false | f54c09fd23315a6f9c86f9dc80f725de7d8f9c64 | Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation los... | 91,128 | 1,585,597 | 7,621,978,240 | [
"task_categories:text-generation",
"language:en",
"license:cdla-sharing-1.0",
"size_categories:1M<n<10M",
"format:parquet",
"modality:text",
"library:datasets",
"library:dask",
"library:polars",
"library:mlcroissant",
"arxiv:2305.07759",
"region:us"
] | 2023-05-12T19:04:09 | null | null |
6a34e9d01b6b6e116d313e13 | Crownelius/Complete-FABLE.5-traces-2M | Crownelius | {"license": "mit", "pretty_name": "Complete FABLE.5 Traces 2M", "annotations_creators": ["machine-generated"], "language": ["en"], "language_creators": ["found", "machine-generated"], "multilinguality": ["monolingual"], "size_categories": ["10K<n<100K"], "task_categories": ["text-generation"], "task_ids": ["language-mo... | false | False | 2026-07-16T18:20:04 | 127 | 17 | false | f4530f12b1a1f46531f26051d66b62f2ad2de63c |
Complete FABLE.5 Traces 2M
Provenance-cleaned FABLE.5 / Claude corpus — trimmed to content-verified traces only.
Dataset Viewer | Parquet
This dataset is a post-closure compilation of FABLE.5 / Claude trace datasets found on Hugging Face after the closure of Fable and Mythos. It is deduplic... | 13,967 | 15,256 | 497,799,384 | [
"task_categories:text-generation",
"task_ids:language-modeling",
"annotations_creators:machine-generated",
"language_creators:found",
"language_creators:machine-generated",
"multilinguality:monolingual",
"language:en",
"license:mit",
"size_categories:10K<n<100K",
"format:parquet",
"modality:tabu... | 2026-06-19T07:03:44 | null | null |
6a54a57f36d31ec6cbee47d6 | ianncity/GLM-5.2-Conversation | ianncity | {"license": "apache-2.0", "task_categories": ["text-generation", "question-answering"], "language": ["en"], "tags": ["reasoning", "chain-of-thought", "science", "physics", "chemistry", "biology", "distillation", "sft", "glm-5.2", "math", "programming"], "size_categories": ["10K<n<100K"], "pretty_name": "GLM-5.2 Convers... | false | False | 2026-07-16T11:30:15 | 21 | 17 | false | c831fbec04d34c982bbf1ef73d07b723c77a4fa8 |
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Cre... | 1,136 | 1,136 | 485,528,881 | [
"task_categories:text-generation",
"task_categories:question-answering",
"language:en",
"license:apache-2.0",
"size_categories:10K<n<100K",
"format:json",
"modality:text",
"library:datasets",
"library:pandas",
"library:polars",
"library:mlcroissant",
"region:us",
"reasoning",
"chain-of-tho... | 2026-07-13T08:44:47 | null | null |
End of preview. Expand in Data Studio
Changelog
NEW Changes March 11th 2026
- Added new split:
arxiv_papers, sourced from the Hugging Face/api/papersendpoint paperscontinues to point todaily_papers.parquet, which is the Daily Papers feed
NEW Changes July 25th
- added
baseModelsfield to models which shows the models that the user tagged as base models for that model
Example:
{
"models": [
{
"_id": "687de260234339fed21e768a",
"id": "Qwen/Qwen3-235B-A22B-Instruct-2507"
}
],
"relation": "quantized"
}
NEW Changes July 9th
- Fixed issue with
ggufcolumn with integer overflow causing import pipeline to be broken over a few weeks ✅
NEW Changes Feb 27th
Added new fields on the
modelssplit:downloadsAllTime,safetensors,ggufAdded new field on the
datasetssplit:downloadsAllTimeAdded new split:
paperswhich is all of the Daily Papers
Updated Daily
- Downloads last month
- 6,393