Codex 자동 연구로 기준보다 232배 빠른 GPU 커널 만들기
Codex를 이용한 GPU Kernel 최적화 사례를 소개합니다.
A case study on GPU kernel optimization using Codex is presented.
AI가 선별한 아티클
Codex를 이용한 GPU Kernel 최적화 사례를 소개합니다.
A case study on GPU kernel optimization using Codex is presented.
자작 SLM 테스트 모델에 대한 피드백 요청.
Request for feedback on a self-built SLM test model.
LLM의 과적합 문제를 다루는 중요한 기사입니다.
An important article addressing the overfitting issue in LLM evaluation.
torch.compile의 속도 최적화에 대한 탐구.
Exploration of performance optimization in torch.compile.
Mac에서 llama.cpp로 Gemma 4를 양자화하는 방법에 대한 튜토리얼입니다.
Tutorial on quantizing Gemma 4 on Mac using llama.cpp.
PyTorch 학습 프로파일링 시 GPU 정지를 피하는 방법에 대한 기술적 노트입니다.
A technical note on profiling PyTorch training without stalling the GPU.