Dataset Viewer
Auto-converted to Parquet Duplicate
_id
large_stringlengths
24
24
id
large_stringlengths
5
123
author
large_stringlengths
2
42
cardData
large_stringlengths
2
1.09M
disabled
bool
1 class
gated
large_stringclasses
3 values
lastModified
timestamp[us]date
2021-02-05 16:03:35
2026-07-25 13:22:28
likes
int64
0
9.77k
trendingScore
float64
0
123
private
bool
1 class
sha
large_stringlengths
40
40
description
large_stringlengths
0
6.67k
downloads
int64
0
8.55M
downloadsAllTime
int64
0
143M
mainSize
float64
0
306,846B
tags
listlengths
1
7.92k
createdAt
timestamp[us]date
2022-03-02 23:29:22
2026-07-25 13:20:19
paperswithcode_id
large_stringclasses
712 values
citation
large_stringlengths
0
10.7k
6a615c95fb10b1093e0ea9ed
HuggingFaceCode/stack-v3-train
HuggingFaceCode
{"thumbnail": "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png", "annotations_creators": [], "language_creators": ["crowdsourced", "expert-generated"], "language": ["code"], "license": ["odc-by"], "multilinguality": ["multilingual"], "size_categories": ["100M<n<1B"], "sourc...
false
False
2026-07-24T18:36:04
128
123
false
de81e3ca7151fc8b8769dbc2dfc0af5ac92b6d1e
🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled dir...
35,376
35,376
4,711,370,231,977
[ "task_categories:text-generation", "language_creators:crowdsourced", "language_creators:expert-generated", "multilinguality:multilingual", "language:code", "license:odc-by", "size_categories:100M<n<1B", "format:parquet", "modality:tabular", "modality:text", "library:datasets", "library:dask", ...
2026-07-23T00:13:09
null
null
6a4e1fe2df56b09d5f449aa8
SupraLabs/reasoning-corpus-4K-5M-v1
SupraLabs
{"license": "apache-2.0", "task_categories": ["text-generation"], "language": ["en"], "tags": ["reasoning", "CoT", "code", "agentic", "thinking", "think", "deepseek-v4", "qwen3", "qwen3next"], "pretty_name": "Reasoning Corpus 5M", "size_categories": ["1M<n<10M"]}
false
False
2026-07-21T02:31:33
107
79
false
1fc50907cffff48186f61e52a750dabec2c09435
Reasoning Corpus 5M · Within 5k sequence length About Dataset This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other reposito...
2,167
2,167
68,664,454,256
[ "task_categories:text-generation", "language:en", "license:apache-2.0", "size_categories:1M<n<10M", "format:json", "modality:text", "library:datasets", "library:pandas", "library:polars", "library:mlcroissant", "region:us", "reasoning", "CoT", "code", "agentic", "thinking", "think", ...
2026-07-08T10:01:06
null
null
6a4cc0ac90ce9cc602189d11
FlyRank/internship-warehouse
FlyRank
{"license": "other", "language": ["en"], "tags": ["seo", "content-performance", "data-warehouse", "tabular", "education", "flyrank-internship"], "pretty_name": "FlyRank Internship \u2014 Warehouse Star Schema (Pseudonymized, Gated)", "size_categories": ["10M<n<100M"], "extra_gated_prompt": "By requesting access you agr...
false
auto
2026-07-07T10:02:21
347
47
false
50cbf7c3909d07be4d1b5906b4d09e882e5acbf2
FlyRank Internship — Pseudonymized Warehouse Release (v20260703) The open-ended, warehouse-shaped dataset (~81.8M rows; daily fact 78,835,655 rows) for advanced capstone work. Star schema with salted, namespaced, fingerprinted hash keys. Built from warehouse v2 full history (frozen snapshot, export date ...
5,195
5,195
1,168,719,310
[ "language:en", "license:other", "size_categories:10M<n<100M", "modality:tabular", "modality:text", "region:us", "seo", "content-performance", "data-warehouse", "tabular", "education", "flyrank-internship" ]
2026-07-07T09:02:36
null
null
6a4509196c643209b19b2fc7
Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset
Manusagents
{"license": "mit", "language": ["en", "multilingual"], "task_categories": ["text-generation", "other"], "tags": ["distillation", "instruction-tuning", "sft", "reasoning", "coding", "code-repositories", "cybersecurity", "attack", "defense", "exploit", "penetration-testing", "red-team", "blue-team", "open-source", "colle...
false
False
2026-07-18T18:01:17
68
43
false
f0aa1d8326d7ca5c4a01982ca8299a783bc59faf
📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~7...
7,265
7,265
76,526,135,473
[ "task_categories:text-generation", "task_categories:other", "language:en", "language:multilingual", "license:mit", "size_categories:10M<n<100M", "format:json", "modality:text", "library:datasets", "library:pandas", "library:polars", "library:mlcroissant", "region:us", "distillation", "in...
2026-07-01T12:33:29
null
null
6a292cbbe1b5c7903e6fbe30
openbmb/UltraX-Preview
openbmb
{"language": ["en"], "license": "apache-2.0", "size_categories": ["10B<n<100B"], "task_categories": ["text-generation"], "pretty_name": "UltraX", "tags": ["llm", "pretraining", "web-corpus", "data-refinement", "programmatic-editing", "function-calling"], "configs": [{"config_name": "UltraX-FineWeb", "data_files": [{"sp...
false
False
2026-07-17T03:02:12
264
42
false
a88527587389fd4ab352e9ad1273f4c0a234d8df
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing 📜 Paper | 💻 Code | 🤖 Models | 📦 UltraData Collection English | 中文 📚 Introduction UltraX is a function-calling refinement framework for large-scale pre-training data that ad...
6,347
6,353
486,915,481,612
[ "task_categories:text-generation", "language:en", "license:apache-2.0", "size_categories:100M<n<1B", "format:parquet", "modality:text", "library:datasets", "library:dask", "library:polars", "library:mlcroissant", "arxiv:2607.08646", "region:us", "llm", "pretraining", "web-corpus", "dat...
2026-06-10T09:22:03
null
null
6a5ae883a5f7ad08ccdbda43
greghavens/kimi-k3-coding-and-debugging-traces
greghavens
{"pretty_name": "Kimi K3 Coding, Tool Use & Instruction Following Traces", "license": "cc-by-4.0", "language": ["en"], "annotations_creators": ["machine-generated"], "task_categories": ["text-generation"], "size_categories": ["1K<n<10K"], "tags": ["traces", "code", "agentic", "tool-use", "coding-agent", "coding-agents"...
false
False
2026-07-25T01:18:22
37
36
false
b594ab95ab7e3964012d8c072e22e71ed81d0b5f
Kimi K3 Coding, Tool Use & Instruction Following Traces 702 TRAJECTORIES · 4,928 TRAINING ROWS · 3 MB PARQUET · 90 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool...
2,442
2,442
5,018,189
[ "task_categories:text-generation", "annotations_creators:machine-generated", "language:en", "license:cc-by-4.0", "size_categories:1K<n<10K", "format:parquet", "modality:tabular", "modality:text", "library:datasets", "library:pandas", "library:polars", "library:mlcroissant", "region:us", "t...
2026-07-18T02:44:19
null
null
6a2cd0828137fb18cecbcc06
Glint-Research/Fable-5-traces
Glint-Research
{"license": "agpl-3.0", "pretty_name": "Fable 5 Pi Agent Traces", "annotations_creators": ["machine-generated"], "language": ["en"], "size_categories": ["1K<n<10K"], "task_categories": ["text-generation"], "tags": ["agent-traces", "pi-agent", "claude-code", "fable-5", "chain-of-thought", "tool-use", "coding-agents", "s...
false
False
2026-06-29T15:10:20
666
35
false
e05c417852fc59fd8da758e68b352732423ca0cb
Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. ...
61,491
90,678
187,507,989
[ "task_categories:text-generation", "annotations_creators:machine-generated", "language:en", "license:agpl-3.0", "size_categories:1K<n<10K", "format:json", "format:agent-traces", "modality:tabular", "modality:text", "library:datasets", "library:dask", "library:polars", "library:mlcroissant", ...
2026-06-13T03:37:38
null
null
6a437ed52e089285573dcfd3
markov-ai/gaming-500-hours
markov-ai
{"configs": [{"config_name": "default", "data_files": [{"split": "train", "path": "metadata.jsonl"}]}]}
false
False
2026-06-30T11:56:39
214
30
false
5af703f2810306e7d75eb4394ae59591f1f6e8a2
Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play session, trimmed to pure gameplay — login screens, launchers, desktop, collection-app references, and any watching/streaming are removed. In-game menus, lobbies, loading, and...
41,011
41,011
1,598,371,626,719
[ "size_categories:n<1K", "format:json", "modality:tabular", "modality:text", "modality:video", "library:datasets", "library:pandas", "library:polars", "library:mlcroissant", "region:us" ]
2026-06-30T08:31:17
null
null
6a60a044d3559d7ff7b5590d
r0b0tlab/qwen3.8-max-distillation-50k
r0b0tlab
{"license": "other", "task_categories": ["text-generation", "question-answering"], "language": ["en"], "tags": ["distillation", "knowledge-distillation", "reasoning", "chain-of-thought", "supervised-fine-tuning", "math", "code", "instruction-following", "tool-use", "qwen"], "size_categories": ["10K<n<100K"], "pretty_na...
false
False
2026-07-22T11:27:58
29
29
false
ab9f8b289423c249fc0054507f045a12efb54b1b
Qwen3.8-Max Distillation 50K A curated dataset of 49,772 teacher-generated traces from qwen3.8-max-preview, prepared for supervised fine-tuning and off-policy knowledge distillation. The teacher responses are preserved as returned by the API. Where the model emitted visible <think>...</think> blocks, tho...
247
247
70,765,792
[ "task_categories:text-generation", "task_categories:question-answering", "language:en", "license:other", "size_categories:10K<n<100K", "format:parquet", "format:optimized-parquet", "modality:tabular", "modality:text", "library:datasets", "library:pandas", "library:polars", "library:mlcroissa...
2026-07-22T10:49:40
null
null
621ffdd236468d709f184284
wikimedia/wikipedia
wikimedia
{"language": ["ab", "ace", "ady", "af", "alt", "am", "ami", "an", "ang", "anp", "ar", "arc", "ary", "arz", "as", "ast", "atj", "av", "avk", "awa", "ay", "az", "azb", "ba", "ban", "bar", "bbc", "bcl", "be", "bg", "bh", "bi", "bjn", "blk", "bm", "bn", "bo", "bpy", "br", "bs", "bug", "bxr", "ca", "cbk", "cdo", "ce", "ceb"...
false
False
2024-01-09T09:40:51
1,330
24
false
b04c8d1ceb2f5cd4588862100d08de323dccfbaa
Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the co...
250,729
2,682,246
71,792,022,791
[ "task_categories:text-generation", "task_categories:fill-mask", "task_ids:language-modeling", "task_ids:masked-language-modeling", "language:ab", "language:ace", "language:ady", "language:af", "language:alt", "language:am", "language:ami", "language:an", "language:ang", "language:anp", "...
2022-03-02T23:29:22
null
null
69d185a53c023c2c9072697a
netflix/Vera-Layered-Video-Dataset
netflix
{"license": "apache-2.0", "task_categories": ["text-to-video"], "tags": ["diffusion", "layered-diffusion", "video", "layered-video-dataset", "video-editing", "video-generation"]}
false
False
2026-07-17T01:33:49
51
18
false
8e0b98ee9bce66fdae345e75aa766ef7c0a04d4e
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing Hongkai Zheng¹²* &nbsp;·&nbsp; Ta-Ying Cheng² &nbsp;·&nbsp; Benjamin Klein² &nbsp;·&nbsp; Yisong Yue¹ &nbsp;·&nbsp; Zhuoning Yuan²† ¹California Institute of Technology &nbsp;&nbsp; ²Netflix, Inc. *Work done...
14,756
24,169
319,639,945,834
[ "task_categories:text-to-video", "license:apache-2.0", "size_categories:10K<n<100K", "modality:video", "arxiv:2606.23610", "region:us", "diffusion", "layered-diffusion", "video", "layered-video-dataset", "video-editing", "video-generation" ]
2026-04-04T21:41:57
null
null
645e8da96320b0efe40ade7a
roneneldan/TinyStories
roneneldan
{"license": "cdla-sharing-1.0", "task_categories": ["text-generation"], "language": ["en"]}
false
False
2024-08-12T13:27:26
1,094
17
false
f54c09fd23315a6f9c86f9dc80f725de7d8f9c64
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation los...
91,128
1,585,597
7,621,978,240
[ "task_categories:text-generation", "language:en", "license:cdla-sharing-1.0", "size_categories:1M<n<10M", "format:parquet", "modality:text", "library:datasets", "library:dask", "library:polars", "library:mlcroissant", "arxiv:2305.07759", "region:us" ]
2023-05-12T19:04:09
null
null
6a34e9d01b6b6e116d313e13
Crownelius/Complete-FABLE.5-traces-2M
Crownelius
{"license": "mit", "pretty_name": "Complete FABLE.5 Traces 2M", "annotations_creators": ["machine-generated"], "language": ["en"], "language_creators": ["found", "machine-generated"], "multilinguality": ["monolingual"], "size_categories": ["10K<n<100K"], "task_categories": ["text-generation"], "task_ids": ["language-mo...
false
False
2026-07-16T18:20:04
127
17
false
f4530f12b1a1f46531f26051d66b62f2ad2de63c
Complete FABLE.5 Traces 2M Provenance-cleaned FABLE.5 / Claude corpus — trimmed to content-verified traces only. Dataset Viewer | Parquet This dataset is a post-closure compilation of FABLE.5 / Claude trace datasets found on Hugging Face after the closure of Fable and Mythos. It is deduplic...
13,967
15,256
497,799,384
[ "task_categories:text-generation", "task_ids:language-modeling", "annotations_creators:machine-generated", "language_creators:found", "language_creators:machine-generated", "multilinguality:monolingual", "language:en", "license:mit", "size_categories:10K<n<100K", "format:parquet", "modality:tabu...
2026-06-19T07:03:44
null
null
6a54a57f36d31ec6cbee47d6
ianncity/GLM-5.2-Conversation
ianncity
{"license": "apache-2.0", "task_categories": ["text-generation", "question-answering"], "language": ["en"], "tags": ["reasoning", "chain-of-thought", "science", "physics", "chemistry", "biology", "distillation", "sft", "glm-5.2", "math", "programming"], "size_categories": ["10K<n<100K"], "pretty_name": "GLM-5.2 Convers...
false
False
2026-07-16T11:30:15
21
17
false
c831fbec04d34c982bbf1ef73d07b723c77a4fa8
GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Cre...
1,136
1,136
485,528,881
[ "task_categories:text-generation", "task_categories:question-answering", "language:en", "license:apache-2.0", "size_categories:10K<n<100K", "format:json", "modality:text", "library:datasets", "library:pandas", "library:polars", "library:mlcroissant", "region:us", "reasoning", "chain-of-tho...
2026-07-13T08:44:47
null
null
End of preview. Expand in Data Studio

Changelog

NEW Changes March 11th 2026

  • Added new split: arxiv_papers, sourced from the Hugging Face /api/papers endpoint
  • papers continues to point to daily_papers.parquet, which is the Daily Papers feed

NEW Changes July 25th

  • added baseModels field to models which shows the models that the user tagged as base models for that model

Example:

{
  "models": [
    {
      "_id": "687de260234339fed21e768a",
      "id": "Qwen/Qwen3-235B-A22B-Instruct-2507"
    }
  ],
  "relation": "quantized"
}

NEW Changes July 9th

  • Fixed issue with gguf column with integer overflow causing import pipeline to be broken over a few weeks ✅

NEW Changes Feb 27th

  • Added new fields on the models split: downloadsAllTime, safetensors, gguf

  • Added new field on the datasets split: downloadsAllTime

  • Added new split: papers which is all of the Daily Papers

Updated Daily

Downloads last month
6,393

Spaces using cfahlgren1/hub-stats 18