MEMORY INDUSTRY INTELLIGENCE

훈련에서 프로덕션까지, NVIDIA와 CoreWeave가 에이전틱 AI의 루프를 완성하다

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v2 · 10.09 23:18사용자 검토 전 초안

원문 제목: From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI

핵심 요약

CoreWeave는 NVIDIA Vera Rubin NVL72 시스템과 Spectrum-X 102.4T 이더넷 네트워킹의 CoreWeave Cloud 제공을 발표했으며, Cognition이 Vera Rubin에서 프로덕션 워크로드를 실행하는 첫 고객이다. Cognition은 GB200 NVL72 대비 SWE-2 추론 워크로드에서 최대 4.8배 토큰 처리량 향상을 조기 테스트에서 확인했다고 밝혔다. CoreWeave는 에이전트용으로 설계된 NVIDIA Vera CPU도 제공 예정이며, 단일 랙에 128개 CPU·11,264코어를 탑재해 11,000개 이상의 동시 환경을 지원한다고 설명했다. 또한 훈련·평가·개선을 연결하는 CoreWeave Forge를 출시하고 Weights & Biases, OpenPipe, marimo를 통합한다고 발표했다. NVIDIA Dynamo가 관리형 추론 서비스와 RL Rollouts를 구동하며, Canva·Capital One·MasterClass 등이 Forge를 초기 도입한다.

메모리 산업 영향 분석

이 문서는 NVIDIA Vera Rubin NVL72, Vera CPU, Spectrum-X 이더넷, BlueField-4 DPU, RTX PRO 6000 등 가속기·네트워킹·CPU 중심의 CoreWeave 클라우드 프로덕션 도입 소식으로, 메모리 산업과의 직접 연결 근거는 본문에 명시되어 있지 않습니다. Vera Rubin NVL72나 GB200 NVL72의 HBM 탑재량·스택 수·용량, 서버 DRAM·LPDDR·SOCAMM 구성, eSSD·NAND 채택에 대한 언급이 없으므로 HBM·서버 DRAM·eSSD 수요로 직접 연결하는 것은 분석가 가설에 해당합니다. 다만 에이전틱 AI의 긴 컨텍스트·높은 동시성·토큰 볼륨 증가와 11,000개 이상 동시 격리 환경, 수천 개 GPU 확장은 향후 가속기당 메모리 탑재량과 서버 메모리·스토리지 사용량을 늘릴 수 있는 방향성 신호로 해석될 수 있으나, 구체적 비트 수요·용량·ASP는 미확인입니다. NAND와 eSSD를 합산하지 않았고, 서버 SOCAMM/LPDDR 소식이 아니므로 모바일 LPDDR 수요로 오인하지 않았습니다. 확인할 지표는 Vera Rubin NVL72의 HBM 세대·스택·용량, 서버당 DRAM/LPDDR 구성, CoreWeave의 eSSD 채택 여부, Cognition 워크로드의 메모리 집약도입니다.
한국어 번역 읽기

수집된 원문 v2의 전체 본문 기준 · 8837자

거의 10년에 걸친 공동 엔지니어링을 바탕으로, CoreWeave는 NVIDIA 컴퓨팅, 네트워킹 및 소프트웨어를 AI를 위해 특별히 구축된 클라우드에 내장했으며, 이는 여러 세대의 배포에 걸쳐 여전히 투자 수익을 내고 있습니다. 이제 CoreWeave는 차세대 NVIDIA 인프라를 프로덕션에 도입하고 있습니다.

이번 주 샌프란시스코에서 열리는 CoreWeave Fully Connected에서 CoreWeave는 Spectrum-X 102.4T 이더넷 네트워킹과 함께 [NVIDIA Vera Rubin NVL72](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/) 시스템의 제공을 발표했습니다. Devin AI 소프트웨어 엔지니어를 만든 응용 AI 연구소인 Cognition이 Vera Rubin에서 프로덕션 워크로드를 실행하는 첫 고객입니다.

CoreWeave는 또한 AI 에이전트를 위해 만들어진 최초의 CPU인 [NVIDIA Vera](https://www.nvidia.com/en-us/data-center/vera-cpu/)도 제공할 예정입니다. 또한 CoreWeave는 NVIDIA 가속 컴퓨팅에서 모델과 에이전트를 훈련, 평가 및 개선하기 위한 연결된 환경인 CoreWeave Forge를 출시했습니다.

NVIDIA의 하이퍼스케일 및 고성능 컴퓨팅 부문 부사장 Ian Buck은 “NVIDIA 가속 컴퓨팅은 세대를 넘어 가치를 제공합니다.”라고 말했습니다. “CoreWeave의 NVIDIA V100 GPU는 Volta 출시 후 거의 10년이 지난 지금도 고객 워크로드를 실행하고 있으며, 동시에 CoreWeave는 Vera Rubin NVL72를 프로덕션에 도입하고 있습니다. 이것이 NVIDIA 플랫폼의 강점입니다. 수년간 수익을 계속 창출하는 인프라와 올바른 워크로드에 올바른 GPU를 배치할 수 있는 유연성입니다.”

## **Cognition, 4.8배 높은 토큰 처리량으로 Vera Rubin NVL72에서 실행**

Cognition은 CoreWeave에서 Devin의 훈련, 강화 학습 및 프로덕션 추론을 실행합니다. 이 회사는 9개월 만에 CoreWeave에서 수천 개의 GPU로 확장하여 Cognition 추론 워크로드를 지원했습니다.

이달 초 CoreWeave는 첫 Vera Rubin NVL72 프로덕션 랙을 받았습니다. 얼마 후 Cognition은 실제 소프트웨어 엔지니어링 워크로드를 사용하여 Vera Rubin의 추론 성능을 GB200 NVL72 기준과 벤치마킹했습니다. 이 워크로드를 생성하기 위해 FrontierCode에서 작업의 하위 집합을 샘플링하고 AI 에이전트를 배포하여 해결했습니다.

초기 테스트에서 Cognition은 Vera Rubin NVL72가 GB200 NVL72 대비 SWE-2 추론 워크로드에 대해 총 토큰 처리량에서 최대 4.8배 증가를 제공하는 것을 확인했습니다. Devin에게 이러한 이득은 더 빠른 실시간 코드 생성과 더 반응성이 좋은 다단계 추론을 의미합니다.

Cognition 창립 팀의 Silas Alberti는 “에이전틱 코딩은 복잡한 워크로드입니다. 긴 컨텍스트, 높은 동시성, 그리고 토큰당 비용이 우리가 출시할 수 있는 것을 결정하는 토큰 볼륨입니다.”라고 말했습니다. “NVIDIA와 CoreWeave 엔지니어들이 어려운 문제를 우리와 함께 해결하면서 모든 것을 하나의 플랫폼에서 처리하는 것이 어떤 단일 사양보다 우리에게 더 중요합니다.”

## **CoreWeave, CoreWeave Cloud에서 NVIDIA Vera Rubin 제공 발표**

CoreWeave는 CoreWeave Cloud에서 NVIDIA Vera Rubin NVL72의 제공을 발표하여 고객 손에 플랫폼을 제공하는 최초의 클라우드 제공업체 중 하나가 되었습니다.

얼리 액세스 고객은 CoreWeave Cloud에서 NVIDIA의 풀스택 AI 팩토리 플랫폼의 성능을 신속하게 활용할 수 있습니다. 며칠 만에 CoreWeave는 Cognition을 위한 프로덕션 Vera Rubin 클러스터를 구축했으며, 이는 인프라부터 제공되는 토큰에 이르기까지 스택 전반에 걸친 NVIDIA와 CoreWeave 간의 공동 설계 및 협업의 결과입니다.

용량은 CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes 및 CoreWeave Inference를 통해 운영할 수 있습니다.

## **NVIDIA Vera CPU, CoreWeave Cloud에 출시 예정, 테스트에서 3배 이상 빠른 에이전틱 샌드박스 시작 확인**

에이전틱 AI는 두 방향에서 인프라에 압력을 가합니다. 에이전트를 서빙하려면 대규모로 저지연 컴퓨팅이 필요하고, 사후 훈련을 통해 개선하려면 동시에 실행되는 수천 개의 격리된 환경이 필요합니다.

NVIDIA Vera CPU는 에이전틱 워크로드를 위해 특별히 제작되었습니다. 에이전틱 AI의 경우 핵심 성능 지표는 동시에 실행할 수 있는 격리된 에이전트 환경의 수와 그 수가 증가함에 따라 각 환경이 얼마나 일관되고 성능을 유지하는지입니다.

CoreWeave의 Vera 배포는 단일 랙에 128개의 CPU와 11,264개의 코어를 배치하여 각각 하나의 코어로 11,000개 이상의 동시 환경을 지원할 수 있습니다. CoreWeave Sandboxes를 사용하면 이러한 환경은 하드웨어 격리되며 지원하는 훈련 작업과 함께 실행되며, Spectrum-X 이더넷 스위치와 BlueField-4 DPU가 안전하고 고성능의 안전한 에이전트 통신을 저지연으로 보장합니다.

테스트에서 CoreWeave는 NVIDIA Vera CPU에서 3배 이상 빠른 에이전트 샌드박스 시작 시간을 달성하여 샌드박스를 가속화하고 확장했습니다. 샌드박스는 강화 학습(RL), 에이전트 도구 사용 및 모델 평가를 위한 실행 계층으로, AI 팀이 CoreWeave에서 격리된 환경에서 코드를 실행할 수 있게 합니다. Terminal-Bench에서 CoreWeave는 Vera CPU에서 모든 통과 작업에 걸쳐 1.7배 성능 향상을 확인했습니다.

## **CoreWeave Forge: 프로덕션에서 훈련으로 돌아가는 AI 루프 완성**

모델과 에이전트는 루프를 실행하여 개선됩니다. 프로덕션 동작이 다음 훈련 실행에 정보를 제공하고, 각 평가가 다음 버전을 정교하게 만듭니다. 이 루프는 역사적으로 서로 다른 공급업체의 도구에 걸쳐 분할되어 있었고, 각 핸드오프마다 신호가 손실되었습니다.

CoreWeave Forge는 Weights & Biases, OpenPipe의 사후 훈련 전문 지식 및 오픈 소스 marimo 노트북 프로젝트를 지속적인 모델 및 에이전트 개선을 위해 구축된 하나의 연결된 환경으로 통합합니다. 모델, 프레임워크 및 클라우드 전반에 걸쳐 개방적으로 유지됩니다.

사용 가능한 신규 및 확장 기능은 다음과 같습니다.

* **CoreWeave ARIA**— 이제 일반 제공 — 사용자가 AI 루프 전반에서 학습, 연구, 코딩 및 반복할 수 있도록 돕고, 실행을 분석하고, 실험을 제안하고, 코드 변경을 권장하고 GitHub에 저장하며, 실험 데이터를 분석하고 변경을 주도한 요인을 표면화하며 다음에 실행할 실험을 제안하는 실행 가능한 증거를 제공합니다.
* **CoreWeave Agent Lens**— 새로운 서비스 — 프로덕션 에이전트 관찰 가능성을 지속적인 개선과 이해 가능한 통찰력으로 전환합니다. 실패 감지를 20% 개선하고 절반의 비용으로 문제를 수정하여 수천만 개의 프로덕션 에이전트 추적을 수정을 주도하는 통찰력으로 전환합니다.
* **CoreWeave Sandboxes** — 이제 일반 제공 — 사용자가 에이전트, 도구 호출, RL 및 평가를 격리된 CPU 또는 GPU 실행 환경에서, 서버리스 인프라 또는 이미 훈련 중인 인프라에서 실행할 수 있게 하여 모든 에이전트 도구 호출, RL 실행 또는 평가를 위한 새롭고 격리된 환경을 제공합니다.
* **사후 훈련은 모델 품질을 개선하고 지연 시간과 비용을 절감**하며 사용자 자체 프로덕션 신호를 활용하고 훈련 클러스터가 필요하지 않습니다.[서버리스 지도 미세 조정](https://www.coreweave.com/products/coreweave-forge/serverless-sft) 및[서버리스 RL](https://www.coreweave.com/products/coreweave-forge/serverless-rl)을 통해 사용자는 자체 훈련 레시피를 실험할 수 있습니다. 서버리스 RL은 자체 관리 설정보다 1.4배 빠르게 훈련하고 40% 낮은 비용으로 훈련합니다.

[NVIDIA Dynamo](https://www.nvidia.com/en-us/ai/dynamo/)는 AI 팩토리를 위한 오픈 소스 추론 프레임워크로, CoreWeave의 관리형 추론 서비스와 현재 비공개 미리보기 중인 RL Rollouts를 구동합니다. RL Rollouts는 실행 중인 라이브 배포에 새 체크포인트를 로드하여 재배포 없이 강화 학습이 계속되고, 사후 훈련이 프로덕션과 동일한 추론 효율성을 얻습니다.

Canva, Capital One 및 MasterClass는 Forge에서 구축하는 최초의 회사 중 하나입니다.

[NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/) 오픈 모델은 Forge의 팀에게 에이전틱 워크플로를 위한 추론 및 멀티모달 모델을 사용자 정의하고 배포할 수 있는 직접적인 경로를 제공합니다.

## **스타트업에서 글로벌 기업까지 입증된 영향**

AI 연구소, AI 네이티브 및 글로벌 기업은 공동 엔지니어링된 NVIDIA 및 CoreWeave 플랫폼을 사용하여 프로토타입에서 프로덕션으로 더 빠르게 이동하고 있습니다.

헬스케어에서 15개 주에 걸쳐 약 50,000명의 고위험 Medicare 환자에게 서비스를 제공하는 재택 기반 치료 제공업체인 Ennoble Care는 임상 AI 추론을 실행하기 위해 CoreWeave를 선택했습니다. 이 회사는 CoreWeave Kubernetes Service에서 예약된 NVIDIA RTX PRO 6000 GPU 용량을 사용하여 임상 문서화, 의사 결정 지원 및 백오피스 자동화를 위한 AI 에이전트를 확장할 예정입니다.

CoreWeave는 모든 훈련 및 추론 라운드에서 기록적인 [MLPerf 결과](https://coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra)를 제공했습니다. SemiAnalysis ClusterMAX 1.0, 2.0 및 3.0에서 플래티넘 등급을 보유한 유일한 클라우드 제공업체이며, 10개 주요 AI 연구소 중 9개에 서비스를 제공합니다.

NVIDIA와 CoreWeave는 함께 고객에게 실험적 에이전트를 소프트웨어를 작성하고, 임상의를 지원하고, 현실 세계에서 유용한 작업을 수행하는 프로덕션 시스템으로 전환할 수 있는 플랫폼을 제공하고 있습니다.

_자세한 내용은_[_CoreWeave Fully Connected의 NVIDIA 세션, 데모 및 워크숍_](https://www.nvidia.com/en-us/events/coreweave-fully-connected/) _참석을 통해 알아보세요._
브리프용 요약 초안
CoreWeave가 NVIDIA Vera Rubin NVL72와 Spectrum-X 102.4T 이더넷을 CoreWeave Cloud에서 제공하고 Cognition이 첫 프로덕션 고객으로 GB200 NVL72 대비 최대 4.8배 토큰 처리량을 보고했습니다. NVIDIA Vera CPU와 CoreWeave Forge, Agent Lens, Sandboxes 등 에이전틱 AI 실행·관찰 계층도 확장됩니다. 다만 본문에는 HBM·서버 DRAM·eSSD 등 메모리 제품의 용량·채택 근거가 없어 메모리 수요 직접 연결은 미확인입니다.

원문 텍스트

원문 열기 ↗
Building on nearly a decade of co-engineering, CoreWeave has built NVIDIA compute, networking and software into a cloud purpose-built for AI that’s still returning on investment across multiple generations of deployment. Now, CoreWeave is bringing the next generation of NVIDIA infrastructure to production.

At CoreWeave Fully Connected, running this week in San Francisco, CoreWeave announced availability of [NVIDIA Vera Rubin NVL72](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/) systems with Spectrum-X 102.4T Ethernet networking. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.

CoreWeave will also offer [NVIDIA Vera](https://www.nvidia.com/en-us/data-center/vera-cpu/), the first CPU built for AI agents. In addition, CoreWeave launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.

“NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. “CoreWeave’s NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production. That’s the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.”

## **Cognition Runs on Vera Rubin NVL72 With 4.8x Higher Token Throughput**

Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled to thousands of GPUs on CoreWeave in nine months, powering Cognition inference workloads.

Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks. Shortly after, Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using a real-world software engineering workload. To generate this workload, it sampled a subset of tasks from FrontierCode and deployed AI agents to solve them.

In its early tests, Cognition saw Vera Rubin NVL72 deliver up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72. For Devin, those gains mean faster real-time code generation and more responsive multistep reasoning.

“Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” said Silas Alberti, founding team at Cognition. “Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec.”

## **CoreWeave Announces NVIDIA Vera Rubin Availability on CoreWeave Cloud**

CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform in customers’ hands.

Early-access customers can put the performance of NVIDIA’s full-stack AI factory platform to work quickly on CoreWeave Cloud. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served.

Capacity can be operated through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference.

## **NVIDIA Vera CPU to Come to CoreWeave Cloud, Tests Show More Than 3x Faster Agentic Sandbox Startups**

Agentic AI puts pressure on infrastructure from two directions: serving agents demands low-latency compute at scale, while improving them through post-training requires thousands of isolated environments running at once.

NVIDIA Vera CPU is purpose-built for agentic workloads. For agentic AI, a key performance measure is how many isolated agent environments can run at once and how consistent and performant each one stays as that number grows.

CoreWeave’s deployment of Vera puts 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each. With CoreWeave Sandboxes, these environments are hardware-isolated and run alongside the training jobs they support, with Spectrum-X Ethernet switches and BlueField-4 DPUs ensuring secure, high-performance, secure agent communication at low latency.

In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on NVIDIA Vera CPUs, accelerating and scaling its sandboxes, an execution layer for reinforcement learning (RL), agent tool use and model evaluation that let AI teams run code in isolated environments on CoreWeave. On Terminal-Bench, CoreWeave saw a 1.7x performance gain on Vera CPU across all passing tasks.

## **CoreWeave Forge: Closing the AI Loop From Production Back to Training**

Models and agents improve by running a loop: production behavior informs the next training run, and each evaluation sharpens the next version. This loop has historically been split across tools from different vendors, with signal lost at every handoff.

CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project in one connected environment built for continuous model and agent improvement. It stays open across models, frameworks and clouds.

New and expanded capabilities available include:

* **CoreWeave ARIA**— now generally available — helps users learn, research, code and iterate across the AI loop, analyzing runs, proposing experiments, recommending code changes and storing them in GitHub, and bringing back actionable evidence that analyzes experiment data, surfaces what drove a change and proposes the next experiments to run.
* **CoreWeave Agent Lens**— a new service — turns production agent observability into continuous improvement and understandable insights. It improves failure detection by 20% and fixes issues at half of the cost, which turns tens of millions of production agent traces into insights that drive fixes.
* **CoreWeave Sandboxes** — now generally available — let users run agents, tool calls, RL and evaluations in isolated CPU or GPU execution environments, on serverless infrastructure or on the infrastructure they already train on, providing a fresh, isolated environment for every agent tool call, RL run or evaluation.
* **Post-training improves model quality and cuts latency and costs** harnessing users’ own production signals, with no training cluster required.[Serverless supervised fine-tuning](https://www.coreweave.com/products/coreweave-forge/serverless-sft) and[serverless RL](https://www.coreweave.com/products/coreweave-forge/serverless-rl) let users experiment with their own training recipes. Serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup.

[NVIDIA Dynamo](https://www.nvidia.com/en-us/ai/dynamo/), an open source inference framework for AI factories, powers CoreWeave’s managed inference service as well as RL Rollouts, now in private preview. RL Rollouts loads new checkpoints into a live deployment while it’s running, so reinforcement learning continues without redeploys, and post-training gets the same inference efficiency as production.

Canva, Capital One and MasterClass are among the first companies building on Forge.

[NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/) open models give teams on Forge a direct path to customizing and deploying reasoning and multimodal models for agentic workflows.

## **Proven Impact From Startups to Global Enterprises**

AI labs, AI-natives and global enterprises are using the co-engineered NVIDIA and CoreWeave platform to move from prototype to production faster.

In healthcare, Ennoble Care, a home-based care provider serving about 50,000 high-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back-office automation.

CoreWeave has delivered record [MLPerf results](https://coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in every round of training and inference. It’s the only cloud provider who holds the Platinum ranking in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0, and serves nine of the 10 leading AI labs.

Together, NVIDIA and CoreWeave are giving customers a platform to turn experimental agents into production systems that write software, support clinicians and do useful work in the real world.

_Learn more by attending_[_NVIDIA sessions, demos and workshops at CoreWeave Fully Connected_](https://www.nvidia.com/en-us/events/coreweave-fully-connected/)_._