South Korea to Open Source 1.5 Trillion Tokens for Foundation Model Training
The Ministry of Science and ICT of South Korea, acting through the National Information Society Agency (NIA), will publish on AI Hub the training corpora assembled by five teams during the first…
Shane Barrett·updated August 31, 2026

The Ministry of Science and ICT of South Korea, acting through the National Information Society Agency (NIA), will publish on AI Hub the training corpora assembled by five teams during the first stage of the government's independent AI foundation model project. The release — 29 data types totaling approximately 1.56 trillion tokens across roughly 35.44 million items — is positioned as sufficient material to train a model in the 70–80 billion parameter range. The participating teams are Naver Cloud, Upstage, SK Telecom, NC AI and LG AI Research.
Composition and Resource Allocation
Each team constructed its corpus according to an internally defined strategy aligned with its specific model development objectives, rather than pooling into a single shared schema. Total government expenditure on data acquisition and processing during the prior fiscal year reached 15 billion won. The aggregated material covers pre-training text, multimodal video and audio assets, post-training corpora targeted at advanced reasoning and agentic systems, and red-teaming datasets intended for safety verification. In aggregate scale terms, the release falls just above the empirical threshold often cited for training models in the 50–70B parameter regime from scratch.
Team-Level Dataset Profile
Upstage contributed approximately 1 trillion pre-training tokens and roughly 500,000 post-training items, supplying both from-scratch training material and downstream agent-oriented fine-tuning data. Naver Cloud focused on video understanding and generative modeling, producing 2.34 million public video items, 520,000 broadcast video items, 15.5 million video-clip-derived textual captions and 12 million voice Q&A pairs. SK Telecom assembled multi-step reasoning corpora across mathematics, science and legal domains, accompanied by voice and image post-training items and approximately 10,000 Korean-context red-teaming samples calibrated against domestic legal norms. NC AI constructed seven capability-specific datasets sourced from industrial environments, including manufacturing documentation and complaint-counseling audio, targeting long-context understanding, step-by-step reasoning, multi-turn dialogue, question-answer and multimodal integration. LG AI Research generated more than 170,000 video clips capturing 50 distinct household activities across 50 in-home filming setups, annotated with object, segmentation, pose and contextual metadata for vision-language model and physical AI training.
Empirical Boundaries and What Remains Unverified
All 29 datasets will be released at no cost on AI Hub; the Ministry has additionally indicated that data accumulated during the project's second-stage evaluation is also scheduled for publication, though specific timelines were not disclosed. Independent verification of token counts, decontamination procedures, licensing terms, dataset-card completeness and format heterogeneity across constituent dataset types remains outstanding. Rigorous downstream evaluation — including ablation studies and head-to-head benchmark comparisons against established open-source corpora — will determine whether the reported improvements in reasoning, video understanding and physical AI capabilities are reproducible outside the participating teams' native training pipelines. The broader structural pattern is consistent with adjacent initiatives in domain-specific training data for game development curricula, where corpora are likewise engineered for narrow institutional objectives rather than harvested through indiscriminate web scraping.