# How a Small Korean Team Hit a Top 5 AA II Score in open model | Nemotron Labs

## Executive summary

Motif, a small Korean AI team, achieved a top 5 AA II score by employing a highly resource-efficient and engineering-intensive strategy. The process involved iterative model scaling, rigorous data curation, and a focus on post-training refinement, particularly enhancing agentic capabilities. The success highlights that world-class AI systems can be built by small teams by prioritizing system-level optimization and deep technical expertise over sheer compute scale.

## Key takeaways

- Iterative Scaling and Resource Management: Motif started with a small model (2.6 billion parameters) to validate hypotheses before scaling up to 12.7 billion parameters. This approach allowed them to achieve significant results with limited resources (approximately 30 people and less than a thousand GPUs) [Timestamp: 03:20].
- Engineering Focus Over Scale: The team emphasized that the engineering aspect—including data curation, monitoring anomalies, and maximizing GPU utilization—is crucial for training Large Representation Models (LRMs), comparing it to the meticulous process of molecular cuisine [Timestamp: 03:40].
- Failure-Driven Development: Post-training focused on creating a feedback loop by analyzing model weaknesses (e.g., failure to conclude or recover errors in agentic scenarios). This led to failure-driven Supervised Fine-Tuning (SFT) [Timestamp: 05:20].
- Importance of Code Integrity: Due to limited resources, the team stressed the necessity of rigorous code review, manually hunting bugs in AI-generated PRs, and fixing unintended bugs found even in widely used frameworks like Torch Titan [Timestamp: 04:10].

## Technical details

- Model Architecture & Attention: Motif utilized Differential Attention, which was selected as the winner of an ablation study, to reduce attention background noise. They plan to continue improving this series, including Group Differential Latent Attention [Timestamp: 08:40].
- Training Methodology: The post-training process involved using Nemo RL for Reinforcement Learning (RL) and applying MOPD (Method of Partial Distillation) to combine multiple specialized teachers, a concept conceptually linked to Mixture of Experts (MoE) [Timestamp: 05:20].
- Compute Requirements: For pre-training, the team utilized an infrastructure of approximately 700 B200 GPUs over a period of two to three months [Timestamp: 09:50].

## Practical implications

- Small teams can compete with large tech companies by focusing on engineering rigor, system optimization, and iterative, failure-driven development cycles.
- The use of specialized, proprietary datasets (internal/publicly crawled) combined with foundational datasets is critical for model performance.
- Advanced techniques like Differential Attention and MOPD can be leveraged to enhance model capabilities without requiring massive, monolithic training runs.

## Topics

AI Model Development, Large Language Models (LLMs), Reinforcement Learning (RL), System Engineering, Korean AI Initiative (KI), NVIDIA Nemotron Labs, Nemotron free training dataset, Nemo RL, Torch Titan, Motif Chat

Source: https://www.youtube.com/watch?v=I-Bjwi0qLfo
