News & Notice

Press Releases

Phase 2 Evaluation Results Unveiled for the Sovereign AI Foundation Model Project

담당부서
작성자
연락처

- All four elite teams demonstrated achievements in developing world-class sovereign AI models.

- Three elite teams (Upstage, SK Telecom, and LG AI Research Institute) advanced to the next round.


【Relevant National Tasks】

21. Becoming the world’s most AI-literate country



The Ministry of Science and ICT (MSIT, Deputy Prime Minister and Minister: Bae Kyung-hoon) and the National IT Industry Promotion Agency (NIPA, President: Park Yun-kyu) announced the Phase 2 evaluation results of the “Sovereign AI Foundation Model” project.


[Achievements of Phase 2 of the Sovereign AI Foundation Model Project (H1 of 2026)]


During the first half of 2026, the elite teams participating in Phase 2 of the “Sovereign AI Foundation Model” project developed globally competitive sovereign AI foundation models despite having relatively limited resources compared to global big tech companies, and released them as open source, thereby demonstrating Korea’s AI competitiveness and its capacity to build an open ecosystem.


Company

Model (Parameter Scale)

Key Features

Motif Technologies

Motif 3

(314B-A13B)

- An agent-specialized AI model that recorded the highest performance among AI models from countries other than the US and China (based on AAII)

-  Integrates specialized AI capabilities in areas such as coding and agents into a single model

Upstage

Solar Open2

(250B-A15B)

- An AI-model optimized for agents

- Capable of processing hundreds of pages of documents or more within a single prompt and running on just two NVIDIA H200 chips, making it particularly cost-efficient

SK Telecom

A.X K2

(688B-A33B)

- Recorded top-tier performance among global open-source models on Korean-language and mathematics benchmarks

- Its applications have expanded to various industries including manufacturing, defense, and biotechnology, and the company has begun providing full-scale support to enable AI transformation throughout society

LG AI Research Institute

K-EXAONE 2.0

(750B-A37B)

- Korea’s largest model with 750 billion parameters

- Released under the Apache 2.0 license and supporting 10    languages

- Achieved a 10% improvement in performance over its predecessor K-EXAONE 1.0 (based on combined benchmarks)

Following the five AI models unveiled in Phase 1 at the end of last year, the four AI models recently revealed by the four elite teams participating in Phase 2 were all listed as “Notable AI Models” by Epoch AI, a US-based nonprofit AI research organization. Currently, only three countries (the United States, China, and Korea) have models recognized as Notable AI Models by Epoch AI in 2026. 

 

In addition, the four elite teams all scored 31 or higher on the Artificial Analysis Intelligence Index (AAII), a holistic benchmark of intelligence and performance of global AI models measured and released by Artificial Analysis (AA), a platform specializing in AI model evaluation. The AI models of the elite teams exceeded major AI models from Europe, Canada, and others (e.g., Mistral Medium 3.5 etc.), and one Korean model was included among those scoring above 40, following the US and China.

By narrowing the gap in AAII scores* with top-tier AI models and ranking 6th** on the AAII global ranking of open-weight models, Korea has strengthened its position as one of the top three AI powers following the US and China.

* The gap with the top-tier models narrowed compared to Phase 1: 19-point gap in Jan 2026 (Score of GPT-5.2: 51, The highest score achieved by a Korean elite team: 32) → 16-point gap in Aug 2026 (Claude Opus 5: 63 vs The highest score achieved by a Korean elite team: 47) ※ Changes in AAII’s evaluation criteria should also be factored in
** 1st place: Kimi K3 (Score: 60) → 2nd place: Qwen3.8 2.4T A95B (Score: 58) → 3rd place: GLM 5.2 (Score: 53) → ... → 6th place: The highest score earned by a Korean elite team (Score: 47)

Korea also came in third following the US and China in the comparison of the performance of frontier language models by each country (based on AAII).



Meanwhile, based on AAII rankings of the highest-scoring model from each AI company worldwide, “Motif 3,” the sovereign AI model developed by Motif Technologies, one of the elite teams, also ranked among the global Top 10.

“Solar Open2,” the sovereign AI model developed by Upstage, has a context window — the amount of text it can remember and process at once — of up to 1 million (1M) tokens, equivalent to hundreds of pages of documents or more, securing information processing capacity on par with global top-tier AI models.

※ Currently, global top-tier AI models such as Claude Opus 5 and GPT-5.6 Sol all have context windows of 1 million tokens.

“A.X K2,” the sovereign AI model developed by SK Telecom, scored at the gold medal level in an evaluation based on the 2026 International Mathematical Olympiad (IMO) problems. The model also achieved a tied top score on the MathArena AIME 2026* leaderboard with 97.1% accuracy, proving its world-class mathematical reasoning capability.

* An evaluation based on the problems of the American Invitational Mathematics Examination (AIME) for US high school students

“K-EXAONE 2.0,” the sovereign AI model developed by LG AI Research Institute, ranked 9th worldwide on the AA-Omniscience Non-Hallucination Rate — an AAII benchmark metric that measures the minimization of AI hallucination* — demonstrating highly reliable AI model performance.

* A metric that measures whether an AI model honestly acknowledges not knowing an answer, rather than fabricating a plausible-sounding lie (hallucination), when faced with unfamiliar information. This metric imposes a heavier penalty on the tendency to guess or pretend to know the answer to unfamiliar questions and get it wrong.

In addition, Hugging Face, the world’s largest open-source platform, has recognized Korea’s Sovereign AI Foundation Model project as an exemplary case of government-led open-source policy, and proposed extensive cooperation including joint promotional activities with Korea as an “alternative” AI power beyond the US and China. The two sides are currently discussing specific plans for implementation.

As such, this is significant in that, through the AI model development competition and evaluations conducted in the first half of this year, Korea has demonstrated both domestically and internationally that its sovereign AI models have secured competitiveness comparable to global top-tier AI models.

In addition to the four elite teams, various Korean AI companies including KT, Kakao, Naver, and NC AI continue to deliver meaningful results in their respective fields*. These K-AI companies’ technologies are being widely applied not only to industrial sectors such as semiconductor processes, legal precedent analysis, and new materials development, but also to public sectors including government administrative networks, R&D budget review, and the AI Challenge for All**, demonstrating the capabilities of Korea’s AI industry and broader ecosystem.

* (KT) Developed Mi:dm K2.0 Omni Basic, its first omni AI model, in July 2026
(Kakao) Unveiled four lightweight on-device foundation models (including Kanana-2-1.3B-base) in July 2026
(Naver) Unveiled “HyperCLOVA X SEED 4B,” an omni-modal AI optimized for the military, in June 2026
(NC AI) Unveiled “VARCO 3D 2.0,” a 3D generation model in June 2026, and began strengthening its foundation for physical AI 
**Referred to the series of ten promotional press releases (issued from May to July 2026) on real-world applications of K-AI technologies

This shows that the Sovereign AI Foundation Model project has enhanced Korea’s AI model development capabilities and competitiveness, and that these achievements are extending beyond the four elite teams to drive the development and adoption of AI models across Korea’s AI ecosystem as a whole, ultimately strengthening, growing, and expanding the capabilities of Korea’s entire AI ecosystem.

[Overview of the Phase 2 Evaluation] 

This Phase 2 evaluation was conducted as a multi-dimensional assessment combining (1) benchmark evaluation (40 points), (2) expert evaluation (35 points), and (3) user evaluation (25 points).



Compared to the Phase 1 evaluation (Jan 2026), the Phase 2 evaluation has been further refined and enhanced by (1) introducing highly challenging global benchmarks recognized for their credibility (benchmark evaluation), (2) expanding the assessment to include the application and expansion of AX (AI Transformation) in real-world settings, going beyond AI model development (expert evaluation), and (3) newly introducing a public-participation usability evaluation (user evaluation).

The significance of Phase 2 evaluation lies not simply in selecting a smaller number of teams through competition among the elite teams, but in the government’s active design and continuous refinement of the evaluation criteria to encourage the elite teams to: (1) strengthen their capabilities to develop world-class AI models, (2) facilitate the adoption and expansion of their sovereign AI models across real-world applications in Korea’s AI ecosystem, and (3) develop AI models that meet the needs and expectations of the public.

This is especially significant because the evaluation criteria were established through extensive consultation with the four elite teams, which are composed of leading AI experts in Korea.

[1. The Benchmark Evaluation and Its Outcomes]

The benchmark evaluation consisted of the AAII benchmark evaluation (25 points) and the National Information Society Agency (NIA) benchmark evaluation (15 points).

The AAII benchmark evaluation was conducted using nine highly challenging benchmarks across four core areas: agents, coding, general, and scientific reasoning. In collaboration with AA, AA itself conducted the benchmark evaluation, enhancing the reliability and fairness of the results.

* Benchmark by AAII Area 
· Agents: GDPval-AA v2, τ3-Banking
· Coding: Terminal-Bench v2.1, Scicode
· General: AA-LCR, AA-Omniscience
· Scientific Reasoning: Humanity’s Last Exam, GPQA Diamond, and CritPt

AAII is known for its highly credible benchmarks, which even leading global big tech companies find difficult to score high on. Its benchmarks are updated irregularly to enhance the index’s ability to distinguish between models, while some test questions and answers are kept confidential to minimize benchmark contamination. In addition, AAII incorporates challenging benchmarks that are difficult even for top-tier AI models to achieve high scores on. 

MSIT introduced this AAII benchmark evaluation to drive the elite teams’ AI model development capabilities toward the global frontier level, while ensuring the credibility of the benchmark evaluation both domestically and internationally.

Meanwhile, the NIA benchmark evaluation comprises benchmarks covering not only mathematics, knowledge, long-text comprehension, instruction following, and Korean language, but also safety and reliability. It is designed to provide a comprehensive assessment of overall performance.

※ Compared to the NIA benchmark used in the Phase 1 evaluation, the instruction-following and Korean language evaluation areas were added to the NIA benchmark for Phase 2.

Not all questions and answers in the NIA benchmark are publicly disclosed, significantly reducing the risk of benchmark contamination. While the AAII benchmark focuses on evaluating the general-purpose performance of AI models for global use, the NIA benchmark comprehensively assesses AI model performance in the Korean language and Korean context, as well as safety and reliability. This makes it a meaningful measure of the actual competitiveness of domestic AI models, which can be difficult to assess using only global general-purpose benchmarks.



The four elite teams recorded an average score of 22.5 points on the benchmark evaluation (40 points), with a 4.0-point gap between the 1st- and 4th-ranked teams.

 [2. The Expert Evaluation and Its Outcomes ]

The expert evaluation was conducted by an expert committee composed of external AI specialists, which conducted the evaluation over approximately one week based on the materials submitted by the elite teams.

※ To facilitate the evaluation, Q&A sessions were held between the expert committee (evaluators) and the elite teams (evaluated teams) alongside the evaluation process.

The committee conducted an in-depth evaluation of the elite teams centered on (1) development strategies and technologies (10 points), (2) development outcomes and future plans (10 points), and (3) ripple effects and contribution plans (15 points).

The expert committee was composed of 10 members with expertise in areas such as AI algorithms, services, and data, and with no conflicts of interest.

Going beyond simply assessing whether sovereignty had been secured, the expert evaluation incorporated “sovereignty and the degree of performance improvement achieved based on it” as an evaluation criterion, allowing teams whose sovereign technologies lead to actual improvements in AI model performance to receive higher scores.

In addition, by increasing the weight given to the ripple effects and contribution plans for the domestic and international AI ecosystem, the evaluation went beyond simply assessing the development of high-performance AI models, with a focus on promoting the broader adoption and expansion of AX in real-world settings across Korea.



The four elite teams recorded an average score of 28.8 points on the expert evaluation (35 points), with a 2.4-point gap between the 1st- and 4th-ranked teams.


[3. The User Evaluation and Its Outcomes]

The user evaluation consisted of two components: a professional AI user evaluation (15 points) and a general citizen evaluation (10 points). Together, these two components were used to assess AI models’ usability, among other factors.

In the professional AI user evaluation, 49 professional AI users including CEOs of AI startups took part. In the general citizen evaluation, 200 citizens were selected, of whom 185 ultimately participated in the evaluation.

The user evaluation was designed to look beyond AI model performance and capture something performance metrics alone can’t show, including how usable and effective real users found each model to be in actual use. This matters because it encourages the elite teams to develop not only high-performing AI models, but also the ability to turn those models into AI services that work well from an end user’s point of view.

Building on this goal, and in contrast to the Phase 1 evaluation, the Phase 2 evaluation focused on establishing a more multi-faceted evaluation framework that assesses AI usability not only from the perspective of AI professionals but also from the general public.



Prompt guidelines were provided to each evaluation group to support in-depth AI usability evaluation, and the evaluation groups’ assessments were conducted through absolute evaluation, rather than comparative evaluation among the elite teams.

The four elite teams recorded an average score of 17.6 points on the user evaluation (25 points), with a 5.0-point gap between the 1st- and 4th-ranked teams.


[Phase 2 Evaluation Results]

As a result of the Phase 2 evaluation, which combined the results of (1) the benchmark evaluation (40 points), (2) the expert evaluation (35 points), and (3) the user evaluation (25 points), three of the four elite teams — Upstage, SK Telecom, and LG AI Research (in Korean alphabetical order) — advanced to the next round by a narrow margin.

According to the expert committee’s evaluation, Upstage was praised for “pursuing integration with the ‘Daum’ portal and the ‘Timely’ platform to enable the public to experience firsthand the outcomes of its development,” and “collaborating with FuriosaAI on NPUs to reduce reliance on foreign hardware (lock-in) and demonstrating the technology through a proof of concept (PoC) in the actual ‘Daum’ service environment, setting a leading example of bringing together Korea’s AI software and hardware ecosystems.”

SK Telecom was praised for “securing top-tier performance among the models in the comparison group in mathematical reasoning and the Korean language, and demonstrating a clear competitive edge in usability and practicality with its model actually deployed in large-scale commercial services,” and “demonstrating the applicability and practicality of its model across a range of industries by providing the A.X K1 model for the defense sector, validating a manufacturing-specialized agent, and showcasing applications in the legal and tax fields.”

LG AI Research Institute received meaningful comments for “establishing collaboration strategies with international organizations and other global partners to increase its global impact, as well as demonstrating distinctive capabilities in Agentic AI,” and “clearly demonstrating its efforts and related activities toward developing safe and reliable AI, while also taking sustainability into consideration through its focus on securing model reliability and managing risk factors.”

Motif Technologies was recognized for its deep technical expertise with comments such as “It developed everything in-house, from the architecture and tokenizer to the optimizer and kernel, eliminating external dependencies,” “Its architectural improvements and training efficiency techniques were successfully applied to the Motif 3 model, achieving strong results in global evaluations,” and “The technical limitations and trial-and-error data generated in the process of applying its sovereign technology stack are also highly valuable as national assets.”

In particular, Motif Technologies recorded the world’s best performance on AAII among AI models from countries other than the US and China, demonstrating performance comparable to that of leading global big tech companies in challenging benchmark areas such as agents and coding. Given these meaningful achievements in the development of sovereign AI models, the team is expected to contribute to Korea’s domestic AI ecosystem in various ways going forward.

[Takeaways and Plans]

The Sovereign AI Foundation Model project not only drove unprecedented, large-scale collaboration among domestic industry, academia, and research institutes, but also recorded meaningful achievements by pooling their capabilities to strengthen Korea’s underlying strength in AI model development and showcasing this both domestically and internationally. Major countries are also assessed to be showing deep interest in Korea’s policy approach.

MSIT plans to expand its support of high-performance GPUs (B200) for the elite teams that advanced to the next round (H1 2026: approximately 768 B200 units → H2 2026: approximately 1,000 B200 units), and through this, the elite teams are expected to focus more intensively on developing higher-performing sovereign AI models.

※ GPU support trend (B200 basis): H2 2025: about 500 → H1 2026: about 768 → H2 2026: about 1,000

Meanwhile, as global competition surrounding AI models intensifies and AI models make notable advances in performance across a wide range of fields, major countries are increasingly moving to position AI models as strategic national security assets, with such efforts becoming increasingly evident and taking concrete shape.

Against this backdrop, MSIT also recognizes the need to develop top-tier sovereign AI models and is currently discussing with relevant ministries a new support framework that goes beyond existing approaches. Specific details will be prepared and announced at a later date.


For further information, please contact the Public Relations Division (Phone: +82-44-202-4034, E-mail: msitmedia@korea.kr) of the Ministry of Science and ICT.

Please refer to the attached PDF.
KOGL Korea Open Government License, BY Type 1 : Source Indication The works of the Ministry of Science and ICT can be used under the terms of "KOGL Type 1".
TOP