Features Core Innovations Results
Kuaishou GameMind Lab / Technical Report · Aug 2026

Kuairp1.0 Role-Play Models Technical Report

Yipeng Wang· Ziwei Zhang· Jiahui Zhang· Qi Gan· Kai Sheng Kuaishou GameMind Lab

A complete technical recipe for the Kuairp series of role-play models.

Model Features

Role-play tailored to game scenarios

A small dedicated model embeds role-play ability and domain world knowledge while preserving general tool-calling, to support diverse in-game play.

01

Small & Efficient

Goal 1 desc.

02

Prompt-Engineering Friendly

Goal 2 desc.

03

Built-in World Knowledge

Goal 3 desc.

04

General Tool-Calling Ability

Goal 4 desc.

Core Innovations

Three Core Innovations

From the SFT data pipeline and reinforcement learning to two-stage OPD on-policy distillation — with CDD at its core.

Instruction-injected user behavior
Feat 1 desc.
Reverse Profile Filtering (RPF)
Feat 2 desc.
Turn-count stratified sampling
Feat 3 desc.
PHASE 1 Domain Expert Training
General Model SFT RL Domain Expert
PHASE 2 OPD Alignment & Fusion
General Model Domain Expert
OPD-stage1 OPD-stage2 Expert with
General Ability
OPD Stage 1: inject format & style into the general model
Stage 1 desc.
OPD Stage 2: inject world knowledge into the general model
Stage 2 desc.
Dive into CDD

Cumulative-Divergence Decay (CDD)

The mechanism
Student
T1 T2 T3 T4 T5 T6
Teacher
d1
d2
d3
d4
d5
d6
c1
c2
c3
c4
c5
c6
w1
w2
w3
w4
w5
w6
Σ
l1
l2
l3
l4
l5
l6

di: the student-teacher divergence for token i
ci: the student-teacher divergence for token i
wi: the loss weight for token i
li: the weighted loss for token i
All li are summed to update the Student.

Evaluation Results

Domain ability peaks, general ability preserved

qwen3-8b (baseline) Kuairp 1.0 M2-HER
85.14
92.08
89.22
60.00
72.50
60.00
61.48
97.44
98.89
97.57
99.12
99.62
93.75
100.00
96.40
1.90
32.86
n/a
24.83
24.78
n/a
Char-Consist.
Memory
Diversity
LangQuality
Length
World Knowledge
Tool Calling*‡

BFCL footnote.

Kuairp footnote.

M2-HER footnote.

More comparison data in our report

Domain role-play ability greatly improved
Point 1 desc.
Domain world knowledge internalized
Point 2 desc.
General tool-calling ability preserved
Point 3 desc.