That’s a fair point. At this scale, instruction tuning is definitely a capacity tradeoff rather than a free capability layer, so preserving the base model’s strengths will be one of the main acceptance criteria.
Regarding the post-merge result, we have done additional validation runs and the performance appears directionally consistent, although I would prefer to publish the full repeated-evaluation results once the methodology and comparison conditions are finalized. At 90M, even small evaluation details can meaningfully affect the reported ranking.
The instruct version will be treated as a separate checkpoint rather than a replacement for the base model.