GPT-5.6 Sol의 Ultrafast Mode가 초당 750토큰 출력을 제공.
Cerebras와 OpenAI가 OpenAI API의 새 서비스 계층인 Ultrafast Mode를 공개했습니다. 이 모드는 GPT-5.6 Sol에서 품질 저하 없이 초당 최대 750개 출력 토큰을 제공합니다. 웨이퍼 크기 칩의 44GB SRAM을 활용해 모델 가중치의 이동 병목 현상을 해결하고 있습니다.
Ultrafast Mode for GPT-5.6 Sol offers up to 750 tokens per second.
Cerebras and OpenAI have unveiled the new service layer of the OpenAI API called Ultrafast Mode. This mode provides up to 750 output tokens per second for GPT-5.6 Sol without quality degradation. It addresses the model weight transfer bottleneck using 44GB SRAM on wafer-scale chips.