MEMORY INDUSTRY INTELLIGENCE
훈련에서 프로덕션까지, NVIDIA와 CoreWeave가 에이전트 AI의 루프를 완성하다
한국어 번역·요약·분석
원문 제목: From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI
원문 문장이 일치하지 않은 주장 1건은 근거 등록에서 제외했습니다.
핵심 요약
CoreWeave는 NVIDIA Vera Rubin NVL72 시스템과 Spectrum-X 102.4T 이더넷 네트워킹의 가용성을 발표했으며, Cognition이 첫 프로덕션 고객이다. CoreWeave는 AI 에이전트용 첫 CPU인 NVIDIA Vera도 제공하고, 훈련·평가·개선을 위한 통합 환경 CoreWeave Forge를 출시했다. Cognition의 초기 테스트에서 Vera Rubin NVL72는 GB200 NVL72 대비 SWE-2 추론 워크로드에서 최대 4.8배 토큰 처리량 향상을 보였다. NVIDIA Vera CPU는 단일 랙에 128개 CPU와 11,264개 코어를 제공하며, 테스트에서 에이전트 샌드박스 시작 시간이 3배 이상 빨라졌다. CoreWeave Forge는 Weights & Biases, OpenPipe, marimo를 통합하고 ARIA, Agent Lens, Sandboxes 등의 기능을 제공한다.
메모리 산업 영향 분석
이 문서는 NVIDIA와 CoreWeave의 에이전트 AI 인프라 협력 확대를 다루며, 메모리 산업에 직접적인 영향은 제한적입니다. Vera Rubin NVL72는 고성능 AI 가속기로, HBM 등 고대역폭 메모리 수요를 견인할 수 있으나 문서에는 구체적인 메모리 사양이나 채택 규모가 명시되지 않았습니다. Cognition의 4.8배 토큰 처리량 향상은 추론 효율 개선을 의미하며, 이는 메모리 대역폭 요구를 높일 수 있는 계층 이동 효과로 해석될 수 있으나 확인되지 않은 가설입니다. CoreWeave Forge의 사후 훈련 및 RL 기능은 훈련 인프라 수요를 변화시킬 수 있으나 메모리 산업과의 직접 연결 근거는 부족합니다. 반대 근거로는 에이전트 AI가 메모리 집약적이지 않을 가능성, 또는 기존 GPU의 메모리 용량으로 충분할 수 있다는 점이 있습니다. 확인할 지표로는 Vera Rubin의 HBM 탑재량, CoreWeave의 GPU 배포 규모, 그리고 메모리 공급업체와의 협력 여부가 있습니다.
한국어 번역 읽기
수집된 원문 v1의 전체 본문 기준 · 8174자
거의 10년에 걸친 공동 엔지니어링을 바탕으로 CoreWeave는 NVIDIA 컴퓨팅, 네트워킹 및 소프트웨어를 AI를 위해 특별히 구축된 클라우드에 내장했으며, 이는 여러 세대의 배포에 걸쳐 여전히 투자 수익을 내고 있습니다. 이제 CoreWeave는 차세대 NVIDIA 인프라를 프로덕션에 도입하고 있습니다.
이번 주 샌프란시스코에서 열리는 CoreWeave Fully Connected에서 CoreWeave는 Spectrum-X 102.4T 이더넷 네트워킹을 갖춘 NVIDIA Vera Rubin NVL72 시스템의 가용성을 발표했습니다. Devin AI 소프트웨어 엔지니어를 만든 응용 AI 연구소인 Cognition이 Vera Rubin에서 프로덕션 워크로드를 실행하는 첫 번째 고객입니다.
CoreWeave는 또한 AI 에이전트를 위해 만들어진 첫 번째 CPU인 NVIDIA Vera를 제공할 예정입니다. 또한 CoreWeave는 NVIDIA 가속 컴퓨팅에서 모델과 에이전트를 훈련, 평가 및 개선하기 위한 연결된 환경인 CoreWeave Forge를 출시했습니다.
NVIDIA의 하이퍼스케일 및 고성능 컴퓨팅 부사장 Ian Buck은 “NVIDIA 가속 컴퓨팅은 세대를 넘어 가치를 제공합니다.”라고 말했습니다. “CoreWeave의 NVIDIA V100 GPU는 Volta 출시 후 거의 10년이 지난 지금도 고객 워크로드를 실행하고 있으며, 동시에 CoreWeave는 Vera Rubin NVL72를 프로덕션에 도입하고 있습니다. 이것이 NVIDIA 플랫폼의 강점입니다. 수년간 수익을 계속 창출하는 인프라와 올바른 워크로드에 올바른 GPU를 배치할 수 있는 유연성입니다.”
Cognition, Vera Rubin NVL72에서 4.8배 높은 토큰 처리량으로 실행
Cognition은 CoreWeave에서 Devin의 훈련, 강화 학습 및 프로덕션 추론을 실행합니다. 이 회사는 9개월 만에 CoreWeave에서 수천 개의 GPU로 확장하여 Cognition 추론 워크로드를 지원했습니다.
이번 달 초 CoreWeave는 첫 Vera Rubin NVL72 프로덕션 랙을 받았습니다. 얼마 후 Cognition은 실제 소프트웨어 엔지니어링 워크로드를 사용하여 GB200 NVL72 기준선에 대해 Vera Rubin의 추론 성능을 벤치마킹했습니다. 이 워크로드를 생성하기 위해 FrontierCode에서 작업의 하위 집합을 샘플링하고 AI 에이전트를 배포하여 해결했습니다.
초기 테스트에서 Cognition은 Vera Rubin NVL72가 GB200 NVL72 대비 SWE-2 추론 워크로드에 대해 총 토큰 처리량에서 최대 4.8배 증가를 제공하는 것을 확인했습니다. Devin에게 이러한 이득은 더 빠른 실시간 코드 생성과 더 반응적인 다단계 추론을 의미합니다.
Cognition 창립 팀의 Silas Alberti는 “에이전트 코딩은 복잡한 워크로드입니다. 긴 컨텍스트, 높은 동시성 및 토큰 볼륨으로 인해 토큰당 비용이 우리가 출시할 수 있는 것을 결정합니다.”라고 말했습니다. “NVIDIA와 CoreWeave 엔지니어들이 우리와 함께 어려운 문제를 해결하면서 모든 것을 하나의 플랫폼에서 처리하는 것이 어떤 단일 사양보다 우리에게 더 중요합니다.”
CoreWeave, CoreWeave Cloud에서 NVIDIA Vera Rubin 가용성 발표
CoreWeave는 CoreWeave Cloud에서 NVIDIA Vera Rubin NVL72의 가용성을 발표하여 고객에게 플랫폼을 제공하는 최초의 클라우드 제공업체 중 하나가 되었습니다.
얼리 액세스 고객은 CoreWeave Cloud에서 NVIDIA의 풀스택 AI 팩토리 플랫폼의 성능을 신속하게 활용할 수 있습니다. 며칠 만에 CoreWeave는 Cognition을 위한 프로덕션 Vera Rubin 클러스터를 구축했으며, 이는 인프라부터 제공되는 토큰에 이르기까지 스택 전반에 걸친 NVIDIA와 CoreWeave 간의 공동 설계 및 협업의 결과로 달성되었습니다.
용량은 CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes 및 CoreWeave Inference를 통해 운영할 수 있습니다.
NVIDIA Vera CPU, CoreWeave Cloud에 출시 예정, 테스트에서 에이전트 샌드박스 시작이 3배 이상 빠름
에이전트 AI는 두 방향에서 인프라에 압력을 가합니다. 에이전트를 제공하려면 대규모로 저지연 컴퓨팅이 필요하고, 사후 훈련을 통해 개선하려면 동시에 실행되는 수천 개의 격리된 환경이 필요합니다.
NVIDIA Vera CPU는 에이전트 워크로드를 위해 특별히 제작되었습니다. 에이전트 AI의 경우 핵심 성능 지표는 동시에 실행할 수 있는 격리된 에이전트 환경의 수와 그 수가 증가함에 따라 각 환경이 얼마나 일관되고 성능을 유지하는지입니다.
CoreWeave의 Vera 배포는 단일 랙에 128개의 CPU와 11,264개의 코어를 제공하며, 각각 하나의 코어로 11,000개 이상의 동시 환경을 실행할 수 있습니다. CoreWeave Sandboxes를 사용하면 이러한 환경은 하드웨어로 격리되고 지원하는 훈련 작업과 함께 실행되며, Spectrum-X 이더넷 스위치와 BlueField-4 DPU가 안전하고 고성능의 안전한 에이전트 통신을 저지연으로 보장합니다.
테스트에서 CoreWeave는 NVIDIA Vera CPU에서 에이전트 샌드박스 시작 시간이 3배 이상 빨라져 샌드박스를 가속화하고 확장했습니다. 이는 AI 팀이 CoreWeave에서 격리된 환경에서 코드를 실행할 수 있게 하는 강화 학습(RL), 에이전트 도구 사용 및 모델 평가를 위한 실행 계층입니다.
Terminal-Bench에서 CoreWeave는 Vera CPU에서 모든 통과 작업에 걸쳐 1.7배 성능 향상을 확인했습니다.
CoreWeave Forge: 프로덕션에서 훈련으로 돌아가는 AI 루프 완성
모델과 에이전트는 루프를 실행하여 개선됩니다. 프로덕션 동작이 다음 훈련 실행에 정보를 제공하고, 각 평가가 다음 버전을 개선합니다. 이 루프는 역사적으로 여러 공급업체의 도구에 걸쳐 분할되어 각 핸드오프에서 신호가 손실되었습니다.
CoreWeave Forge는 Weights & Biases, OpenPipe의 사후 훈련 전문 지식 및 오픈 소스 marimo 노트북 프로젝트를 지속적인 모델 및 에이전트 개선을 위해 구축된 하나의 연결된 환경으로 통합합니다. 모델, 프레임워크 및 클라우드 전반에 걸쳐 개방형을 유지합니다.
사용 가능한 새롭고 확장된 기능은 다음과 같습니다.
CoreWeave ARIA — 이제 일반 제공 — 사용자가 AI 루프 전반에서 학습, 연구, 코딩 및 반복할 수 있도록 지원하며, 실행 분석, 실험 제안, 코드 변경 제안 및 GitHub에 저장, 실험 데이터를 분석하고 변경을 주도한 요인을 표면화하며 실행할 다음 실험을 제안하는 실행 가능한 증거를 제공합니다.
CoreWeave Agent Lens — 새로운 서비스 — 프로덕션 에이전트 관찰 가능성을 지속적인 개선과 이해 가능한 통찰력으로 전환합니다. 실패 감지를 20% 개선하고 절반의 비용으로 문제를 수정하여 수천만 개의 프로덕션 에이전트 추적을 수정을 주도하는 통찰력으로 전환합니다.
CoreWeave Sandboxes — 이제 일반 제공 — 사용자가 격리된 CPU 또는 GPU 실행 환경에서 에이전트, 도구 호출, RL 및 평가를 실행할 수 있게 하며, 서버리스 인프라 또는 이미 훈련 중인 인프라에서 실행됩니다. 모든 에이전트 도구 호출, RL 실행 또는 평가를 위해 새롭고 격리된 환경을 제공합니다.
사후 훈련은 훈련 클러스터 없이 사용자 자체 프로덕션 신호를 활용하여 모델 품질을 개선하고 지연 시간과 비용을 절감합니다.
서버리스 지도 미세 조정 및 서버리스 RL을 통해 사용자는 자체 훈련 레시피를 실험할 수 있습니다. 서버리스 RL은 자체 관리 설정보다 40% 낮은 비용으로 1.4배 빠르게 훈련합니다.
NVIDIA Dynamo — AI 팩토리를 위한 오픈 소스 추론 프레임워크 — CoreWeave의 관리형 추론 서비스와 현재 비공개 미리 보기 중인 RL Rollouts를 지원합니다. RL Rollouts는 실행 중인 라이브 배포에 새 체크포인트를 로드하여 재배포 없이 강화 학습이 계속되고 사후 훈련이 프로덕션과 동일한 추론 효율성을 얻습니다.
Canva, Capital One 및 MasterClass는 Forge를 기반으로 구축하는 첫 번째 회사 중 하나입니다.
NVIDIA Nemotron 오픈 모델은 Forge의 팀에게 에이전트 워크플로를 위한 추론 및 멀티모달 모델을 사용자 정의하고 배포할 수 있는 직접적인 경로를 제공합니다.
스타트업에서 글로벌 기업까지 입증된 영향
AI 연구소, AI 네이티브 및 글로벌 기업은 공동 엔지니어링된 NVIDIA 및 CoreWeave 플랫폼을 사용하여 프로토타입에서 프로덕션으로 더 빠르게 이동하고 있습니다.
의료 분야에서 15개 주에 걸쳐 약 50,000명의 고위험 Medicare 환자에게 서비스를 제공하는 재택 의료 제공업체인 Ennoble Care는 임상 AI 추론을 실행하기 위해 CoreWeave를 선택했습니다. 이 회사는 CoreWeave Kubernetes Service에서 예약된 NVIDIA RTX PRO 6000 GPU 용량을 사용하여 임상 문서화, 의사 결정 지원 및 백오피스 자동화를 위한 AI 에이전트를 확장할 예정입니다.
CoreWeave는 모든 훈련 및 추론 라운드에서 기록적인 MLPerf 결과를 제공했습니다. SemiAnalysis ClusterMAX 1.0, 2.0 및 3.0에서 플래티넘 등급을 보유한 유일한 클라우드 제공업체이며, 10개 주요 AI 연구소 중 9개에 서비스를 제공합니다.
NVIDIA와 CoreWeave는 함께 고객에게 실험적 에이전트를 소프트웨어를 작성하고 임상의를 지원하며 실제 세계에서 유용한 작업을 수행하는 프로덕션 시스템으로 전환할 수 있는 플랫폼을 제공하고 있습니다.
CoreWeave Fully Connected에서 NVIDIA 세션, 데모 및 워크숍에 참석하여 자세히 알아보세요.
이번 주 샌프란시스코에서 열리는 CoreWeave Fully Connected에서 CoreWeave는 Spectrum-X 102.4T 이더넷 네트워킹을 갖춘 NVIDIA Vera Rubin NVL72 시스템의 가용성을 발표했습니다. Devin AI 소프트웨어 엔지니어를 만든 응용 AI 연구소인 Cognition이 Vera Rubin에서 프로덕션 워크로드를 실행하는 첫 번째 고객입니다.
CoreWeave는 또한 AI 에이전트를 위해 만들어진 첫 번째 CPU인 NVIDIA Vera를 제공할 예정입니다. 또한 CoreWeave는 NVIDIA 가속 컴퓨팅에서 모델과 에이전트를 훈련, 평가 및 개선하기 위한 연결된 환경인 CoreWeave Forge를 출시했습니다.
NVIDIA의 하이퍼스케일 및 고성능 컴퓨팅 부사장 Ian Buck은 “NVIDIA 가속 컴퓨팅은 세대를 넘어 가치를 제공합니다.”라고 말했습니다. “CoreWeave의 NVIDIA V100 GPU는 Volta 출시 후 거의 10년이 지난 지금도 고객 워크로드를 실행하고 있으며, 동시에 CoreWeave는 Vera Rubin NVL72를 프로덕션에 도입하고 있습니다. 이것이 NVIDIA 플랫폼의 강점입니다. 수년간 수익을 계속 창출하는 인프라와 올바른 워크로드에 올바른 GPU를 배치할 수 있는 유연성입니다.”
Cognition, Vera Rubin NVL72에서 4.8배 높은 토큰 처리량으로 실행
Cognition은 CoreWeave에서 Devin의 훈련, 강화 학습 및 프로덕션 추론을 실행합니다. 이 회사는 9개월 만에 CoreWeave에서 수천 개의 GPU로 확장하여 Cognition 추론 워크로드를 지원했습니다.
이번 달 초 CoreWeave는 첫 Vera Rubin NVL72 프로덕션 랙을 받았습니다. 얼마 후 Cognition은 실제 소프트웨어 엔지니어링 워크로드를 사용하여 GB200 NVL72 기준선에 대해 Vera Rubin의 추론 성능을 벤치마킹했습니다. 이 워크로드를 생성하기 위해 FrontierCode에서 작업의 하위 집합을 샘플링하고 AI 에이전트를 배포하여 해결했습니다.
초기 테스트에서 Cognition은 Vera Rubin NVL72가 GB200 NVL72 대비 SWE-2 추론 워크로드에 대해 총 토큰 처리량에서 최대 4.8배 증가를 제공하는 것을 확인했습니다. Devin에게 이러한 이득은 더 빠른 실시간 코드 생성과 더 반응적인 다단계 추론을 의미합니다.
Cognition 창립 팀의 Silas Alberti는 “에이전트 코딩은 복잡한 워크로드입니다. 긴 컨텍스트, 높은 동시성 및 토큰 볼륨으로 인해 토큰당 비용이 우리가 출시할 수 있는 것을 결정합니다.”라고 말했습니다. “NVIDIA와 CoreWeave 엔지니어들이 우리와 함께 어려운 문제를 해결하면서 모든 것을 하나의 플랫폼에서 처리하는 것이 어떤 단일 사양보다 우리에게 더 중요합니다.”
CoreWeave, CoreWeave Cloud에서 NVIDIA Vera Rubin 가용성 발표
CoreWeave는 CoreWeave Cloud에서 NVIDIA Vera Rubin NVL72의 가용성을 발표하여 고객에게 플랫폼을 제공하는 최초의 클라우드 제공업체 중 하나가 되었습니다.
얼리 액세스 고객은 CoreWeave Cloud에서 NVIDIA의 풀스택 AI 팩토리 플랫폼의 성능을 신속하게 활용할 수 있습니다. 며칠 만에 CoreWeave는 Cognition을 위한 프로덕션 Vera Rubin 클러스터를 구축했으며, 이는 인프라부터 제공되는 토큰에 이르기까지 스택 전반에 걸친 NVIDIA와 CoreWeave 간의 공동 설계 및 협업의 결과로 달성되었습니다.
용량은 CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes 및 CoreWeave Inference를 통해 운영할 수 있습니다.
NVIDIA Vera CPU, CoreWeave Cloud에 출시 예정, 테스트에서 에이전트 샌드박스 시작이 3배 이상 빠름
에이전트 AI는 두 방향에서 인프라에 압력을 가합니다. 에이전트를 제공하려면 대규모로 저지연 컴퓨팅이 필요하고, 사후 훈련을 통해 개선하려면 동시에 실행되는 수천 개의 격리된 환경이 필요합니다.
NVIDIA Vera CPU는 에이전트 워크로드를 위해 특별히 제작되었습니다. 에이전트 AI의 경우 핵심 성능 지표는 동시에 실행할 수 있는 격리된 에이전트 환경의 수와 그 수가 증가함에 따라 각 환경이 얼마나 일관되고 성능을 유지하는지입니다.
CoreWeave의 Vera 배포는 단일 랙에 128개의 CPU와 11,264개의 코어를 제공하며, 각각 하나의 코어로 11,000개 이상의 동시 환경을 실행할 수 있습니다. CoreWeave Sandboxes를 사용하면 이러한 환경은 하드웨어로 격리되고 지원하는 훈련 작업과 함께 실행되며, Spectrum-X 이더넷 스위치와 BlueField-4 DPU가 안전하고 고성능의 안전한 에이전트 통신을 저지연으로 보장합니다.
테스트에서 CoreWeave는 NVIDIA Vera CPU에서 에이전트 샌드박스 시작 시간이 3배 이상 빨라져 샌드박스를 가속화하고 확장했습니다. 이는 AI 팀이 CoreWeave에서 격리된 환경에서 코드를 실행할 수 있게 하는 강화 학습(RL), 에이전트 도구 사용 및 모델 평가를 위한 실행 계층입니다.
Terminal-Bench에서 CoreWeave는 Vera CPU에서 모든 통과 작업에 걸쳐 1.7배 성능 향상을 확인했습니다.
CoreWeave Forge: 프로덕션에서 훈련으로 돌아가는 AI 루프 완성
모델과 에이전트는 루프를 실행하여 개선됩니다. 프로덕션 동작이 다음 훈련 실행에 정보를 제공하고, 각 평가가 다음 버전을 개선합니다. 이 루프는 역사적으로 여러 공급업체의 도구에 걸쳐 분할되어 각 핸드오프에서 신호가 손실되었습니다.
CoreWeave Forge는 Weights & Biases, OpenPipe의 사후 훈련 전문 지식 및 오픈 소스 marimo 노트북 프로젝트를 지속적인 모델 및 에이전트 개선을 위해 구축된 하나의 연결된 환경으로 통합합니다. 모델, 프레임워크 및 클라우드 전반에 걸쳐 개방형을 유지합니다.
사용 가능한 새롭고 확장된 기능은 다음과 같습니다.
CoreWeave ARIA — 이제 일반 제공 — 사용자가 AI 루프 전반에서 학습, 연구, 코딩 및 반복할 수 있도록 지원하며, 실행 분석, 실험 제안, 코드 변경 제안 및 GitHub에 저장, 실험 데이터를 분석하고 변경을 주도한 요인을 표면화하며 실행할 다음 실험을 제안하는 실행 가능한 증거를 제공합니다.
CoreWeave Agent Lens — 새로운 서비스 — 프로덕션 에이전트 관찰 가능성을 지속적인 개선과 이해 가능한 통찰력으로 전환합니다. 실패 감지를 20% 개선하고 절반의 비용으로 문제를 수정하여 수천만 개의 프로덕션 에이전트 추적을 수정을 주도하는 통찰력으로 전환합니다.
CoreWeave Sandboxes — 이제 일반 제공 — 사용자가 격리된 CPU 또는 GPU 실행 환경에서 에이전트, 도구 호출, RL 및 평가를 실행할 수 있게 하며, 서버리스 인프라 또는 이미 훈련 중인 인프라에서 실행됩니다. 모든 에이전트 도구 호출, RL 실행 또는 평가를 위해 새롭고 격리된 환경을 제공합니다.
사후 훈련은 훈련 클러스터 없이 사용자 자체 프로덕션 신호를 활용하여 모델 품질을 개선하고 지연 시간과 비용을 절감합니다.
서버리스 지도 미세 조정 및 서버리스 RL을 통해 사용자는 자체 훈련 레시피를 실험할 수 있습니다. 서버리스 RL은 자체 관리 설정보다 40% 낮은 비용으로 1.4배 빠르게 훈련합니다.
NVIDIA Dynamo — AI 팩토리를 위한 오픈 소스 추론 프레임워크 — CoreWeave의 관리형 추론 서비스와 현재 비공개 미리 보기 중인 RL Rollouts를 지원합니다. RL Rollouts는 실행 중인 라이브 배포에 새 체크포인트를 로드하여 재배포 없이 강화 학습이 계속되고 사후 훈련이 프로덕션과 동일한 추론 효율성을 얻습니다.
Canva, Capital One 및 MasterClass는 Forge를 기반으로 구축하는 첫 번째 회사 중 하나입니다.
NVIDIA Nemotron 오픈 모델은 Forge의 팀에게 에이전트 워크플로를 위한 추론 및 멀티모달 모델을 사용자 정의하고 배포할 수 있는 직접적인 경로를 제공합니다.
스타트업에서 글로벌 기업까지 입증된 영향
AI 연구소, AI 네이티브 및 글로벌 기업은 공동 엔지니어링된 NVIDIA 및 CoreWeave 플랫폼을 사용하여 프로토타입에서 프로덕션으로 더 빠르게 이동하고 있습니다.
의료 분야에서 15개 주에 걸쳐 약 50,000명의 고위험 Medicare 환자에게 서비스를 제공하는 재택 의료 제공업체인 Ennoble Care는 임상 AI 추론을 실행하기 위해 CoreWeave를 선택했습니다. 이 회사는 CoreWeave Kubernetes Service에서 예약된 NVIDIA RTX PRO 6000 GPU 용량을 사용하여 임상 문서화, 의사 결정 지원 및 백오피스 자동화를 위한 AI 에이전트를 확장할 예정입니다.
CoreWeave는 모든 훈련 및 추론 라운드에서 기록적인 MLPerf 결과를 제공했습니다. SemiAnalysis ClusterMAX 1.0, 2.0 및 3.0에서 플래티넘 등급을 보유한 유일한 클라우드 제공업체이며, 10개 주요 AI 연구소 중 9개에 서비스를 제공합니다.
NVIDIA와 CoreWeave는 함께 고객에게 실험적 에이전트를 소프트웨어를 작성하고 임상의를 지원하며 실제 세계에서 유용한 작업을 수행하는 프로덕션 시스템으로 전환할 수 있는 플랫폼을 제공하고 있습니다.
CoreWeave Fully Connected에서 NVIDIA 세션, 데모 및 워크숍에 참석하여 자세히 알아보세요.
브리프용 요약 초안
NVIDIA와 CoreWeave가 Vera Rubin NVL72와 Vera CPU를 프로덕션에 도입하며 에이전트 AI 인프라를 강화하고 있다. Cognition의 초기 테스트에서 4.8배 토큰 처리량 향상이 보고되었으나, 메모리 산업에 대한 직접적 영향은 문서에서 확인되지 않는다. HBM 수요 증가 가능성은 있으나 구체적 사양이나 채택 규모가 명시되지 않아 추가 확인이 필요하다.
원문 텍스트
원문 열기 ↗Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production.
At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced availability of
NVIDIA Vera Rubin NVL72
systems with Spectrum-X 102.4T Ethernet networking. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.
CoreWeave will also offer
NVIDIA Vera
, the first CPU built for AI agents. In addition, CoreWeave launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.
“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production. That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.
”
Cognition Runs on Vera Rubin NVL72 With 4.8x Higher Token Throughput
Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled to thousands of GPUs on CoreWeave in nine months, powering Cognition inference workloads.
Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks. Shortly after, Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using a real-world software engineering workload. To generate this workload, it sampled a subset of tasks from FrontierCode and deployed AI agents to solve them.
In its early tests, Cognition saw Vera Rubin NVL72 deliver up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72. For Devin, those gains mean faster real-time code generation and more responsive multistep reasoning.
“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” said Silas Alberti, founding team at Cognition. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec.”
CoreWeave Announces NVIDIA Vera Rubin Availability on CoreWeave Cloud
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform in customers’ hands.
Early-access customers can put the performance of NVIDIA’s full-stack AI factory platform to work quickly on CoreWeave Cloud. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served.
Capacity can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference.
NVIDIA Vera CPU to Come to CoreWeave Cloud, Tests Show More Than 3x Faster Agentic Sandbox Startups
Agentic AI puts pressure on infrastructure from two directions: serving agents demands low-latency compute at scale, while improving them through post-training requires thousands of isolated environments running at once.
NVIDIA Vera CPU is purpose-built for agentic workloads. For agentic AI, a key performance measure is how many isolated agent environments can run at once and how consistent and performant each one stays as that number grows.
CoreWeave’s deployment of Vera puts 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each. With CoreWeave Sandboxes, these environments are hardware-isolated and run alongside the training jobs they support, with Spectrum-X Ethernet switches and BlueField-4 DPUs ensuring secure, high-performance, secure agent communication at low latency.
In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs, accelerating and scaling its sandboxes, an execution layer for reinforcement learning (RL), agent tool use and model evaluation that let AI teams run code in isolated environments on CoreWeave.
On Terminal-Bench, CoreWeave saw a 1.7x performance gain on Vera CPU across all passing tasks.
CoreWeave Forge: Closing the AI Loop From Production Back to Training
Models and agents improve by running a loop: production behavior informs the next training run, and each evaluation sharpens the next version. This loop has historically been split across tools from different vendors, with signal lost at every handoff.
CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project in one connected environment built for continuous model and agent improvement. It stays open across models, frameworks and clouds.
New and expanded capabilities available include:
CoreWeave ARIA
— now generally available — helps users learn, research, code and iterate across the AI loop, analyzing runs, proposing experiments, recommending code changes and storing them in GitHub, and bringing back actionable evidence that analyzes experiment data, surfaces what drove a change and proposes the next experiments to run.
CoreWeave Agent Lens
— a new service — turns production agent observability into continuous improvement and understandable insights. It improves failure detection by 20% and fixes issues at half of the cost, which turns tens of millions of production agent traces into insights that drive fixes.
CoreWeave Sandboxes
— now generally available — let users run agents, tool calls, RL and evaluations in isolated CPU or GPU execution environments, on serverless infrastructure or on the infrastructure they already train on, providing a fresh, isolated environment for every agent tool call, RL run or evaluation.
Post-training improves model quality and cuts latency and costs
harnessing users’ own production signals, with no training cluster required.
Serverless supervised fine-tuning
and
serverless RL
let users experiment with their own training recipes. Serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup.
NVIDIA Dynamo
, an open source inference framework for AI factories, powers CoreWeave’s managed inference service as well as RL Rollouts, now in private preview. RL Rollouts loads new checkpoints into a live deployment while it’s running, so reinforcement learning continues without redeploys, and post-training gets the same inference efficiency as production.
Canva, Capital One and MasterClass are among the first companies building on Forge.
NVIDIA Nemotron
open models give teams on Forge a direct path to customizing and deploying reasoning and multimodal models for agentic workflows.
Proven Impact From Startups to Global Enterprises
AI labs, AI-natives and global enterprises are using the co-engineered NVIDIA and CoreWeave platform to move from prototype to production faster.
In healthcare, Ennoble Care, a home-based care provider serving about 50,000 high-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back-office automation.
CoreWeave has delivered record
MLPerf results
in every round of training and inference. It’s the only cloud provider who holds the Platinum ranking in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0, and serves nine of the 10 leading AI labs.
Together, NVIDIA and CoreWeave are giving customers a platform to turn experimental agents into production systems that write software, support clinicians and do useful work in the real world.
Learn more by attending
NVIDIA sessions, demos and workshops at CoreWeave Fully Connected
.
At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced availability of
NVIDIA Vera Rubin NVL72
systems with Spectrum-X 102.4T Ethernet networking. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.
CoreWeave will also offer
NVIDIA Vera
, the first CPU built for AI agents. In addition, CoreWeave launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.
“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production. That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.
”
Cognition Runs on Vera Rubin NVL72 With 4.8x Higher Token Throughput
Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled to thousands of GPUs on CoreWeave in nine months, powering Cognition inference workloads.
Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks. Shortly after, Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using a real-world software engineering workload. To generate this workload, it sampled a subset of tasks from FrontierCode and deployed AI agents to solve them.
In its early tests, Cognition saw Vera Rubin NVL72 deliver up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72. For Devin, those gains mean faster real-time code generation and more responsive multistep reasoning.
“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” said Silas Alberti, founding team at Cognition. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec.”
CoreWeave Announces NVIDIA Vera Rubin Availability on CoreWeave Cloud
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform in customers’ hands.
Early-access customers can put the performance of NVIDIA’s full-stack AI factory platform to work quickly on CoreWeave Cloud. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served.
Capacity can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference.
NVIDIA Vera CPU to Come to CoreWeave Cloud, Tests Show More Than 3x Faster Agentic Sandbox Startups
Agentic AI puts pressure on infrastructure from two directions: serving agents demands low-latency compute at scale, while improving them through post-training requires thousands of isolated environments running at once.
NVIDIA Vera CPU is purpose-built for agentic workloads. For agentic AI, a key performance measure is how many isolated agent environments can run at once and how consistent and performant each one stays as that number grows.
CoreWeave’s deployment of Vera puts 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each. With CoreWeave Sandboxes, these environments are hardware-isolated and run alongside the training jobs they support, with Spectrum-X Ethernet switches and BlueField-4 DPUs ensuring secure, high-performance, secure agent communication at low latency.
In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs, accelerating and scaling its sandboxes, an execution layer for reinforcement learning (RL), agent tool use and model evaluation that let AI teams run code in isolated environments on CoreWeave.
On Terminal-Bench, CoreWeave saw a 1.7x performance gain on Vera CPU across all passing tasks.
CoreWeave Forge: Closing the AI Loop From Production Back to Training
Models and agents improve by running a loop: production behavior informs the next training run, and each evaluation sharpens the next version. This loop has historically been split across tools from different vendors, with signal lost at every handoff.
CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project in one connected environment built for continuous model and agent improvement. It stays open across models, frameworks and clouds.
New and expanded capabilities available include:
CoreWeave ARIA
— now generally available — helps users learn, research, code and iterate across the AI loop, analyzing runs, proposing experiments, recommending code changes and storing them in GitHub, and bringing back actionable evidence that analyzes experiment data, surfaces what drove a change and proposes the next experiments to run.
CoreWeave Agent Lens
— a new service — turns production agent observability into continuous improvement and understandable insights. It improves failure detection by 20% and fixes issues at half of the cost, which turns tens of millions of production agent traces into insights that drive fixes.
CoreWeave Sandboxes
— now generally available — let users run agents, tool calls, RL and evaluations in isolated CPU or GPU execution environments, on serverless infrastructure or on the infrastructure they already train on, providing a fresh, isolated environment for every agent tool call, RL run or evaluation.
Post-training improves model quality and cuts latency and costs
harnessing users’ own production signals, with no training cluster required.
Serverless supervised fine-tuning
and
serverless RL
let users experiment with their own training recipes. Serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup.
NVIDIA Dynamo
, an open source inference framework for AI factories, powers CoreWeave’s managed inference service as well as RL Rollouts, now in private preview. RL Rollouts loads new checkpoints into a live deployment while it’s running, so reinforcement learning continues without redeploys, and post-training gets the same inference efficiency as production.
Canva, Capital One and MasterClass are among the first companies building on Forge.
NVIDIA Nemotron
open models give teams on Forge a direct path to customizing and deploying reasoning and multimodal models for agentic workflows.
Proven Impact From Startups to Global Enterprises
AI labs, AI-natives and global enterprises are using the co-engineered NVIDIA and CoreWeave platform to move from prototype to production faster.
In healthcare, Ennoble Care, a home-based care provider serving about 50,000 high-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back-office automation.
CoreWeave has delivered record
MLPerf results
in every round of training and inference. It’s the only cloud provider who holds the Platinum ranking in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0, and serves nine of the 10 leading AI labs.
Together, NVIDIA and CoreWeave are giving customers a platform to turn experimental agents into production systems that write software, support clinicians and do useful work in the real world.
Learn more by attending
NVIDIA sessions, demos and workshops at CoreWeave Fully Connected
.