The significance of Phase 2 evaluation lies not simply in selecting a smaller number of teams through competition among the elite teams, but in the government’s active design and continuous refinement of the evaluation criteria to encourage the elite teams to: (1) strengthen their capabilities to develop world-class AI models, (2) facilitate the adoption and expansion of their sovereign AI models across real-world applications in Korea’s AI ecosystem, and (3) develop AI models that meet the needs and expectations of the public.
This is especially significant because the evaluation criteria were established through extensive consultation with the four elite teams, which are composed of leading AI experts in Korea.
[1. The Benchmark Evaluation and Its Outcomes]
The benchmark evaluation consisted of the AAII benchmark evaluation (25 points) and the National Information Society Agency (NIA) benchmark evaluation (15 points).
The AAII benchmark evaluation was conducted using nine highly challenging benchmarks across four core areas: agents, coding, general, and scientific reasoning. In collaboration with AA, AA itself conducted the benchmark evaluation, enhancing the reliability and fairness of the results.
* Benchmark by AAII Area
· Agents: GDPval-AA v2, τ3-Banking
· Coding: Terminal-Bench v2.1, Scicode
· General: AA-LCR, AA-Omniscience
· Scientific Reasoning: Humanity’s Last Exam, GPQA Diamond, and CritPt
AAII is known for its highly credible benchmarks, which even leading global big tech companies find difficult to score high on. Its benchmarks are updated irregularly to enhance the index’s ability to distinguish between models, while some test questions and answers are kept confidential to minimize benchmark contamination. In addition, AAII incorporates challenging benchmarks that are difficult even for top-tier AI models to achieve high scores on.
MSIT introduced this AAII benchmark evaluation to drive the elite teams’ AI model development capabilities toward the global frontier level, while ensuring the credibility of the benchmark evaluation both domestically and internationally.
Meanwhile, the NIA benchmark evaluation comprises benchmarks covering not only mathematics, knowledge, long-text comprehension, instruction following, and Korean language, but also safety and reliability. It is designed to provide a comprehensive assessment of overall performance.
※ Compared to the NIA benchmark used in the Phase 1 evaluation, the instruction-following and Korean language evaluation areas were added to the NIA benchmark for Phase 2.
Not all questions and answers in the NIA benchmark are publicly disclosed, significantly reducing the risk of benchmark contamination. While the AAII benchmark focuses on evaluating the general-purpose performance of AI models for global use, the NIA benchmark comprehensively assesses AI model performance in the Korean language and Korean context, as well as safety and reliability. This makes it a meaningful measure of the actual competitiveness of domestic AI models, which can be difficult to assess using only global general-purpose benchmarks.

The four elite teams recorded an average score of 22.5 points on the benchmark evaluation (40 points), with a 4.0-point gap between the 1st- and 4th-ranked teams.
[2. The Expert Evaluation and Its Outcomes ]
The expert evaluation was conducted by an expert committee composed of external AI specialists, which conducted the evaluation over approximately one week based on the materials submitted by the elite teams.
※ To facilitate the evaluation, Q&A sessions were held between the expert committee (evaluators) and the elite teams (evaluated teams) alongside the evaluation process.
The committee conducted an in-depth evaluation of the elite teams centered on (1) development strategies and technologies (10 points), (2) development outcomes and future plans (10 points), and (3) ripple effects and contribution plans (15 points).
The expert committee was composed of 10 members with expertise in areas such as AI algorithms, services, and data, and with no conflicts of interest.
Going beyond simply assessing whether sovereignty had been secured, the expert evaluation incorporated “sovereignty and the degree of performance improvement achieved based on it” as an evaluation criterion, allowing teams whose sovereign technologies lead to actual improvements in AI model performance to receive higher scores.
In addition, by increasing the weight given to the ripple effects and contribution plans for the domestic and international AI ecosystem, the evaluation went beyond simply assessing the development of high-performance AI models, with a focus on promoting the broader adoption and expansion of AX in real-world settings across Korea.

The four elite teams recorded an average score of 28.8 points on the expert evaluation (35 points), with a 2.4-point gap between the 1st- and 4th-ranked teams.
[3. The User Evaluation and Its Outcomes]
The user evaluation consisted of two components: a professional AI user evaluation (15 points) and a general citizen evaluation (10 points). Together, these two components were used to assess AI models’ usability, among other factors.
In the professional AI user evaluation, 49 professional AI users including CEOs of AI startups took part. In the general citizen evaluation, 200 citizens were selected, of whom 185 ultimately participated in the evaluation.
The user evaluation was designed to look beyond AI model performance and capture something performance metrics alone can’t show, including how usable and effective real users found each model to be in actual use. This matters because it encourages the elite teams to develop not only high-performing AI models, but also the ability to turn those models into AI services that work well from an end user’s point of view.
Building on this goal, and in contrast to the Phase 1 evaluation, the Phase 2 evaluation focused on establishing a more multi-faceted evaluation framework that assesses AI usability not only from the perspective of AI professionals but also from the general public.

Prompt guidelines were provided to each evaluation group to support in-depth AI usability evaluation, and the evaluation groups’ assessments were conducted through absolute evaluation, rather than comparative evaluation among the elite teams.
The four elite teams recorded an average score of 17.6 points on the user evaluation (25 points), with a 5.0-point gap between the 1st- and 4th-ranked teams.
[Phase 2 Evaluation Results]
As a result of the Phase 2 evaluation, which combined the results of (1) the benchmark evaluation (40 points), (2) the expert evaluation (35 points), and (3) the user evaluation (25 points), three of the four elite teams — Upstage, SK Telecom, and LG AI Research (in Korean alphabetical order) — advanced to the next round by a narrow margin.
According to the expert committee’s evaluation, Upstage was praised for “pursuing integration with the ‘Daum’ portal and the ‘Timely’ platform to enable the public to experience firsthand the outcomes of its development,” and “collaborating with FuriosaAI on NPUs to reduce reliance on foreign hardware (lock-in) and demonstrating the technology through a proof of concept (PoC) in the actual ‘Daum’ service environment, setting a leading example of bringing together Korea’s AI software and hardware ecosystems.”
SK Telecom was praised for “securing top-tier performance among the models in the comparison group in mathematical reasoning and the Korean language, and demonstrating a clear competitive edge in usability and practicality with its model actually deployed in large-scale commercial services,” and “demonstrating the applicability and practicality of its model across a range of industries by providing the A.X K1 model for the defense sector, validating a manufacturing-specialized agent, and showcasing applications in the legal and tax fields.”
LG AI Research Institute received meaningful comments for “establishing collaboration strategies with international organizations and other global partners to increase its global impact, as well as demonstrating distinctive capabilities in Agentic AI,” and “clearly demonstrating its efforts and related activities toward developing safe and reliable AI, while also taking sustainability into consideration through its focus on securing model reliability and managing risk factors.”
Motif Technologies was recognized for its deep technical expertise with comments such as “It developed everything in-house, from the architecture and tokenizer to the optimizer and kernel, eliminating external dependencies,” “Its architectural improvements and training efficiency techniques were successfully applied to the Motif 3 model, achieving strong results in global evaluations,” and “The technical limitations and trial-and-error data generated in the process of applying its sovereign technology stack are also highly valuable as national assets.”
In particular, Motif Technologies recorded the world’s best performance on AAII among AI models from countries other than the US and China, demonstrating performance comparable to that of leading global big tech companies in challenging benchmark areas such as agents and coding. Given these meaningful achievements in the development of sovereign AI models, the team is expected to contribute to Korea’s domestic AI ecosystem in various ways going forward.
[Takeaways and Plans]
The Sovereign AI Foundation Model project not only drove unprecedented, large-scale collaboration among domestic industry, academia, and research institutes, but also recorded meaningful achievements by pooling their capabilities to strengthen Korea’s underlying strength in AI model development and showcasing this both domestically and internationally. Major countries are also assessed to be showing deep interest in Korea’s policy approach.
MSIT plans to expand its support of high-performance GPUs (B200) for the elite teams that advanced to the next round (H1 2026: approximately 768 B200 units → H2 2026: approximately 1,000 B200 units), and through this, the elite teams are expected to focus more intensively on developing higher-performing sovereign AI models.
※ GPU support trend (B200 basis): H2 2025: about 500 → H1 2026: about 768 → H2 2026: about 1,000
Meanwhile, as global competition surrounding AI models intensifies and AI models make notable advances in performance across a wide range of fields, major countries are increasingly moving to position AI models as strategic national security assets, with such efforts becoming increasingly evident and taking concrete shape.
Against this backdrop, MSIT also recognizes the need to develop top-tier sovereign AI models and is currently discussing with relevant ministries a new support framework that goes beyond existing approaches. Specific details will be prepared and announced at a later date.
For further information, please contact the Public Relations Division (Phone: +82-44-202-4034, E-mail: msitmedia@korea.kr) of the Ministry of Science and ICT.
Please refer to the attached PDF.