Model Features
Role-play tailored to game scenarios
A small dedicated model embeds role-play ability and domain world knowledge while preserving general tool-calling, to support diverse in-game play.
01
Small & Efficient
Goal 1 desc.
02
Prompt-Engineering Friendly
Goal 2 desc.
03
Built-in World Knowledge
Goal 3 desc.
04
General Tool-Calling Ability
Goal 4 desc.
Core Innovations
Three Core Innovations
From the SFT data pipeline and reinforcement learning to two-stage OPD on-policy distillation — with CDD at its core.
Instruction-injected user behavior
Feat 1 desc.
Reverse Profile Filtering (RPF)
Feat 2 desc.
Turn-count stratified sampling
Feat 3 desc.
PHASE 1
Domain Expert Training
General Model
SFT
RL
Domain Expert
PHASE 2
OPD Alignment & Fusion
General Model
Domain Expert
OPD-stage1
OPD-stage2
Expert withGeneral Ability
OPD Stage 1: inject format & style into the general model
Stage 1 desc.
OPD Stage 2: inject world knowledge into the general model
Stage 2 desc.
Cumulative-Divergence Decay (CDD)
The mechanism
Student
T1
T2
T3
T4
T5
T6
Teacher
d1
d2
d3
d4
d5
d6
c1
c2
c3
c4
c5
c6
w1
w2
w3
w4
w5
w6
Σ
l1
l2
l3
l4
l5
l6
di: the student-teacher divergence for token ici: the student-teacher divergence for token iwi: the loss weight for token ili: the weighted loss for token i
All li are summed to update the Student.
Evaluation Results
Domain ability peaks, general ability preserved
qwen3-8b (baseline)
Kuairp 1.0†
M2-HER
Char-Consist.
Memory
Diversity
LangQuality
Length
World Knowledge‡
Tool Calling*‡
BFCL footnote.
Kuairp footnote.
M2-HER footnote.
More comparison data in our report
Domain role-play ability greatly improved
Point 1 desc.
Domain world knowledge internalized
Point 2 desc.
General tool-calling ability preserved
Point 3 desc.