AI-ML·중요도 6·2026. 09. 09.·The New Stack
Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.
── KO ──────────────────
Claude가 새로운 기준에서 '에이전트를 구축하는 에이전트' 벤치마크에서 최고의 성과를 보였으나 테스트에서 25%도 통과하지 못했다.
AI 모델들이 코딩 도우미부터 고객 서비스 시스템에 이르기까지 다양한 에이전트의 힘을 더하고 있다. 최근 Claude는 '에이전트를 구축하는 에이전트'를 위한 새로운 벤치마크에서 최상의 결과를 보였지만, 여전히 테스트의 25%도 통과하지 못하는 성과를 기록했다. 이 결과는 AI 모델의 현재 성능 한계를 시사한다.
── EN ──────────────────
Claude excelled in a new benchmark for 'agents that build agents', yet passed fewer than a quarter of tests.
AI models now drive various agents, from coding assistants to customer service systems. Recently, Claude performed best on a new benchmark for 'agents that build agents'; however, it still only passed less than 25% of the tests. This outcome underscores the current limitations of AI model performance.