Qwen3.8-Flash-Next는 비용 효율을 높인 멀티모달 MoE 모델 아키텍처를 소개합니다.
Qwen3.8-Flash-Next는 새로운 멀티모달 MoE 모델을 도입하여 비용 효율성을 높입니다. 이 모델은 125B 본체와 51B N-gram 임베딩을 사용하며, 토큰당 6B의 매개변수만 활성화됩니다. GDN과 QSA를 통해 과거 정보를 압축하고 중요한 문맥을 효율적으로 검색하는 성능 개선도 이루어졌습니다.
Qwen3.8-Flash-Next introduces a cost-efficient multimodal MoE model architecture.
Qwen3.8-Flash-Next presents a new cost-efficient multimodal MoE model architecture. It utilizes a 125B backbone and a 51B N-gram embedding, activating only 6B parameters per token. The integration of GDN and QSA allows for the compression of historical information and efficient retrieval of important context.