MEMORY INDUSTRY INTELLIGENCE
Crusoe의 NVIDIA GB200 NVL72 기반 MLPerf Inference v6.1 결과
한국어 번역·요약·분석
원문 제목: Crusoe's MLPerf Inference v6.1 results on NVIDIA GB200 NVL72
핵심 요약
Crusoe는 MLPerf Inference v6.1에 처음 제출한 결과로 NVIDIA GB200 NVL72에서 gpt-oss-120b와 Qwen3-VL-235B-A22B를 벤치마크했다. Qwen3-VL-235B-A22B에서 4x GB200 NVL72 노드 기준 Offline 58.40 samples/s, Server 49.70 queries/s를 기록해 NVIDIA의 참고치(약 56, 46)를 상회했고, gpt-oss-120b는 GPU당 약 12,700 tokens/s로 closed division 최고 수준이라고 밝혔다. 1개 노드에서 4개 노드로의 확장은 사실상 선형이었으며, prefill:decode GPU 비율을 75:25에서 50:50으로 재조정해 Interactive 처리량이 15.85에서 39.85 queries/s로 약 2.5배 향상됐다. 모든 결과는 Crusoe Cloud의 Managed Kubernetes 및 Slurm 환경에서 closed division 규칙을 통과했고, 정확도 기준도 충족했다. 이 글은 결과, 방법론, 멀티노드 prefill/decode 분리 토폴로지 구현을 다루며 AMD MI355X 제출 결과는 별도 글로 다룬다.
메모리 산업 영향 분석
이 문서는 NVIDIA GB200 NVL72에서의 추론 성능 벤치마크 결과로, 메모리 산업과의 직접 연결 근거는 제한적이다. 원문은 HBM, DDR, LPDDR, GDDR 등 특정 메모리 제품 규격이나 용량, 대역폭, 공급업체를 명시하지 않는다. GB200 NVL72는 HBM을 탑재한 GPU 가속기 시스템이지만, 이 글은 소프트웨어 스택, 토폴로지 튜닝, 처리량 수치에 초점을 맞추고 있어 메모리 수요나 계층 이동을 직접 확인할 수 없다. 다만 prefill/decode 분리와 KV 캐시 전송(NIXL)이 언급되어, 대규모 추론 서빙에서 메모리 대역폭과 용량이 성능에 영향을 줄 수 있다는 분석가 가설은 가능하다. 그러나 이는 원문에서 직접 확인된 사실이 아니며, HBM 수요 증가 또는 감소를 단정할 수 없다. 관련 제품·고객·재무 지표에 대한 구체적 정보가 없어 메모리 산업 영향 분석은 미확인으로 남는다.
한국어 번역 읽기
수집된 원문 v1의 전체 본문 기준 · 7860자
Cloud
Engineering
2026년 9월 16일
Crusoe의 NVIDIA GB200 NVL72 기반 MLPerf Inference v6.1 결과
Crusoe의 첫 MLPerf Inference v6.1 제출은 Crusoe Cloud의 NVIDIA GB200 NVL72에서 gpt-oss-120b와 Qwen3-235B를 벤치마크했습니다. 이들은 고객이 사용하는 것과 동일한 Managed Kubernetes 및 Slurm 인프라에서 실행됩니다.
Young Jeong
Staff Solutions Engineer
Alex Akesson
Principal Product Manager
2026년 9월 16일
목차
This is some text inside of a div block.
공유:
오늘 MLCommons가 MLPerf® Inference v6.1 결과를 발표했으며, 여기에는 Crusoe의 첫 MLPerf 제출도 포함됩니다. 우리는 Crusoe Cloud의 NVIDIA GB200 NVL72에서 실행되는 gpt-oss-120b와 Qwen3-VL-235B-A22B에 대한 결과를 제출했으며, 이는 NVIDIA CUDA® 및 NVIDIA TensorRT-LLM부터 NVIDIA Dynamo, NVIDIA NVLink™, NVIDIA Quantum InfiniBand에 이르는 NVIDIA Accelerated Computing Platform이 고객이 실제로 소비하는 방식 그대로 작동함을 검증합니다.
세 가지 결과가 원시 수치보다 더 중요합니다:
우리는 Qwen3-VL-235B-A22B에 대해 NVIDIA가 공개한 단일 노드 참조를 초과했습니다.
NVIDIA의 4x GB200 NVL72 노드에 대한 가이던스는 대략 Offline 56 samples/s, Server 46 queries/s였습니다. 우리 제출은 58.40과 49.70을 측정했습니다. 우리의 gpt-oss-120b 제출 역시 closed division에서 GPU당 최고 GB200 NVL72 결과로, GPU당 초당 약 12,700 토큰입니다.
1개 노드에서 4개 노드로의 확장은 사실상 선형이었습니다.
16 GPU에서 Qwen3-VL Server 처리량은 4 GPU 결과의 4.00배였고, Offline은 선형 외삽과 1% 미만 차이로 근접했습니다. 랙 스케일 NVLink와 NVIDIA Quantum InfiniBand는 이 워크로드 형태에서 패브릭을 비요인으로 만듭니다.
토폴로지 튜닝 패스는 동일 하드웨어에서 약 2.5배의 Interactive 처리량 향상을 만들어냈습니다.
분리형 NVIDIA Dynamo 배포에서 prefill:decode GPU 분할을 75:25에서 50:50으로 재조정함으로써, 16 GPU에서 Qwen3-VL Interactive 처리량이 15.85에서 39.85 queries/s로 증가했습니다. 동일한 GPU, 동일한 패브릭, 극적으로 더 많은 서빙 용량입니다.
이 글은 결과, 그 뒤의 방법론, 그리고 멀티노드 prefill/decode 분리 서빙 토폴로지를 MLPerf의 closed-division 규칙을 통과시키기까지의 엔지니어링을 다룹니다. 동반 글은 우리의 AMD MI355X 제출을 다룹니다.
NVIDIA GB200 NVL72 기반 MLPerf Inference v6.1 결과
모든 결과는 MLPerf Inference v6.1, closed division, Crusoe Cloud의 NVIDIA GB200 NVL72(gb200-186gb-nvl-ib.4x 인스턴스, 노드당 4x GB200, GPU당 1200W TGP)입니다.
모델
시스템
시나리오
결과
gpt-oss-120b
4x GB200 NVL72
Offline
50,912.9 tok/s
Server
50,613.5 tok/s
Qwen3-VL-235B-A22B
4x GB200 NVL72
Offline
58.40 samples/s
Server
49.70 queries/s
Interactive
3.99 queries/s
16x GB200 NVL72
Offline
230.09 samples/s
Server
198.94 queries/s
Interactive
39.85 queries/s
gpt-oss-120b는 두 시나리오 모두에서 GPU당 초당 약 12,700 토큰으로 계산됩니다. MLPerf는 Server 시나리오에서 지연 시간에 제한된 처리량을 측정하므로, 흥미로운 숫자는 비율입니다: 우리의 Server 결과는 Offline의 0.6% 이내에 들어옵니다. 이 플랫폼에서 지연 SLA 비용은 거의 없습니다. 최대 배치 처리량을 내는 동일 노드가 v6.1 한도 내에서 p99 time-to-first-token과 time-per-output-token을 유지하며, 이는 프로덕션 배포 규모 산정 시 중요한 속성입니다.
모든 실행은 closed-division 정확도를 통과했습니다: gpt-oss-120b는 82.30% 하한 대비 83.69% exact-match(Offline) 및 83.64%(Server), Qwen3-VL은 78.24% 하한 대비 78.56%에서 78.82% 사이의 hierarchical F1을 기록했습니다. gpt-oss-120b는 컴플라이언스 감사 TEST07 및 TEST09를 통과했습니다.
공개된 결과는 MLCommons MLPerf Inference v6.1 closed-division 결과에서 확인할 수 있습니다.
소프트웨어 구성
구성 요소
gpt-oss-120b
Qwen3-VL-235B-A22B
추론 스택
TensorRT-LLM (feat/1.2-mlpinf) via trtllm-serve
vLLM (CentML mlperf-inf-mm-q3vl) + NVIDIA Dynamo (mlperf-v6.0-dynamo-v0.8.0)
CUDA 스택
TensorRT 10.14, CUDA 13.1, NVIDIA cuDNN 9.17
NVIDIA Dynamo with NIXL KV transfer
노드 유형
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace™ CPU, aarch64
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace CPU, aarch64
패브릭
NVL72 랙 내 NVLink
NVL72 랙 내 NVLink
오케스트레이션
Crusoe Managed Kubernetes (단일 노드)
NVIDIA sflow를 통한 Crusoe Managed Slurm 클러스터(멀티노드), VAST의 공유 NFS
Prefill/decode 분리: 토폴로지만으로 2.5배 이득
16-GPU Qwen3-VL 배포는 우리가 제출한 것 중 아키텍처적으로 가장 흥미로운 시스템입니다. 4개의 독립적인 노드 크기 복제본을 실행하는 대신, 우리는 NVIDIA Dynamo를 prefill/decode 분리 토폴로지로 배포했습니다: prefill 워커(TP2)는 프롬프트 처리를 담당하고, decode 워커(TP1)는 토큰 생성을 담당하며, KV 캐시는 NVIDIA Inference Xfer Library(NIXL)를 통해 이들 사이를 이동합니다.
분리가 중요한 이유는 prefill과 decode가 반대의 하드웨어 요구를 갖기 때문입니다. Prefill은 컴퓨트 바운드이며 긴 프롬프트에 걸쳐 병렬화됩니다; decode는 메모리 대역폭 바운드이며 최대 배치 상주를 원합니다. 두 단계를 동일 GPU에서 서빙하면 타협이 강제됩니다; 분리하면 각 풀이 자연스러운 동작점에서 실행될 수 있습니다. 이는 대규모 프로덕션 추론 시스템 뒤의 동일한 아키텍처이며, 타이트한 p99 지연 한도를 가진 MLPerf의 Interactive 시나리오는 정확히 이점이 나타나는 곳입니다.
우리의 첫 Interactive 실행은 75:25 prefill:decode GPU 분할을 사용했고 15.85 queries/s를 측정했습니다. 프로파일링 결과 decode가 병목이었습니다: prefill GPU가 p99 한도 아래에서 decode GPU가 토큰을 스트리밍할 수 있는 것보다 더 빠르게 프롬프트를 완료했습니다. 50:50으로 재조정하고 지속 가능한 쿼리 속도를 이진 탐색하여 제출을 39.85 queries/s로 끌어올렸으며, 이는 하드웨어 변경 없이 약 2.5배 개선입니다. 최종 토폴로지는 3개 노드에 걸친 6개의 prefill 워커와 네 번째 노드의 4개의 decode 워커였습니다. 현재까지 이 결과는 GB200 NVL72 시스템에서 해당 라운드의 최고 Qwen3-VL Interactive 점수입니다.
분리형 서빙을 배포하는 모든 사람을 위한 두 가지 시사점:
prefill:decode 비율은 워크로드 의존적이며 스윕할 가치가 있습니다.
지연 바운드 Interactive 워크로드에 대한 최적 분할은 처리량 지향 기본값과 매우 달랐습니다.
확장은 선형이었습니다.
4개 노드에서 Offline 처리량(230.09 samples/s)은 3개 노드 실행에서의 선형 외삽(172.35가 229.8을 예측)과 일치했습니다. Server는 설정 수정 후 첫 시도에서 목표 200 대비 198.94를 기록했으며, 이는 구성된 속도의 99.5%입니다.
프로덕션이 실행되는 방식으로 MLPerf 실행
우리는 오늘날 고객 워크로드를 구동하는 오케스트레이션 서비스에서 의도적으로 이 벤치마크를 실행했습니다.
단일 노드 gpt-oss-120b 실행은 Crusoe Managed Kubernetes에서 실행되었으며, 이는 우리의 다른 MLPerf 작업과 동일한 플랫폼이고 Crusoe Managed Metrics를 통한 동일한 관측성을 갖습니다. 멀티노드 Qwen3-VL 실행은 NVIDIA의 sflow 런처를 사용하는 Crusoe Managed Slurm 클러스터에서 실행되었으며, 모델 저장소는 VAST 기반 공유 NFS에 있고 타이밍 실행 중 스토리지 핫 패스가 네트워크를 건너지 않도록 노드별 NVMe로 스테이징되었습니다.
다음 단계
이 제출은 Crusoe에서 NVIDIA Accelerated Computing Platform을 두 가지 규모로 보여줍니다: 공개된 참조 처리량을 초과하는 단일 GB200 노드, 그리고 NVLink와 NVIDIA Quantum InfiniBand를 가로질러 선형으로 확장되는 멀티노드 분리 토폴로지. 모든 Crusoe 클러스터는 고객이 만지기 전에 브링업 중 엔드투엔드로 검증되며, MLPerf는 그 과정의 한 도구입니다.
우리는 GB200 NVL72, GB300 NVL72 및 그 이후에 걸쳐 성능과 소프트웨어 공동 설계를 위해 NVIDIA와 계속 협력하고 있으며, Inference와 Training 워크로드 모두에 대한 결과를 계속 게시할 것입니다.
이 규모로 서빙할 준비가 되셨나요?
Crusoe Cloud의 NVIDIA GB200 NVL72에서 추론 실행에 대해 우리 팀과 상담하세요.
최신 글
2026년 10월 9일
Crusoe의 추론 엔진은 얼마나 친환경적인가?
2026년 10월 9일
Lone Star State에서 커뮤니티 구축
2026년 10월 7일
Triton을 사용한 GPU 프로그래밍, 파트 1
멋진 것을 만들 준비가 되셨나요?
문의하기
Engineering
2026년 9월 16일
Crusoe의 NVIDIA GB200 NVL72 기반 MLPerf Inference v6.1 결과
Crusoe의 첫 MLPerf Inference v6.1 제출은 Crusoe Cloud의 NVIDIA GB200 NVL72에서 gpt-oss-120b와 Qwen3-235B를 벤치마크했습니다. 이들은 고객이 사용하는 것과 동일한 Managed Kubernetes 및 Slurm 인프라에서 실행됩니다.
Young Jeong
Staff Solutions Engineer
Alex Akesson
Principal Product Manager
2026년 9월 16일
목차
This is some text inside of a div block.
공유:
오늘 MLCommons가 MLPerf® Inference v6.1 결과를 발표했으며, 여기에는 Crusoe의 첫 MLPerf 제출도 포함됩니다. 우리는 Crusoe Cloud의 NVIDIA GB200 NVL72에서 실행되는 gpt-oss-120b와 Qwen3-VL-235B-A22B에 대한 결과를 제출했으며, 이는 NVIDIA CUDA® 및 NVIDIA TensorRT-LLM부터 NVIDIA Dynamo, NVIDIA NVLink™, NVIDIA Quantum InfiniBand에 이르는 NVIDIA Accelerated Computing Platform이 고객이 실제로 소비하는 방식 그대로 작동함을 검증합니다.
세 가지 결과가 원시 수치보다 더 중요합니다:
우리는 Qwen3-VL-235B-A22B에 대해 NVIDIA가 공개한 단일 노드 참조를 초과했습니다.
NVIDIA의 4x GB200 NVL72 노드에 대한 가이던스는 대략 Offline 56 samples/s, Server 46 queries/s였습니다. 우리 제출은 58.40과 49.70을 측정했습니다. 우리의 gpt-oss-120b 제출 역시 closed division에서 GPU당 최고 GB200 NVL72 결과로, GPU당 초당 약 12,700 토큰입니다.
1개 노드에서 4개 노드로의 확장은 사실상 선형이었습니다.
16 GPU에서 Qwen3-VL Server 처리량은 4 GPU 결과의 4.00배였고, Offline은 선형 외삽과 1% 미만 차이로 근접했습니다. 랙 스케일 NVLink와 NVIDIA Quantum InfiniBand는 이 워크로드 형태에서 패브릭을 비요인으로 만듭니다.
토폴로지 튜닝 패스는 동일 하드웨어에서 약 2.5배의 Interactive 처리량 향상을 만들어냈습니다.
분리형 NVIDIA Dynamo 배포에서 prefill:decode GPU 분할을 75:25에서 50:50으로 재조정함으로써, 16 GPU에서 Qwen3-VL Interactive 처리량이 15.85에서 39.85 queries/s로 증가했습니다. 동일한 GPU, 동일한 패브릭, 극적으로 더 많은 서빙 용량입니다.
이 글은 결과, 그 뒤의 방법론, 그리고 멀티노드 prefill/decode 분리 서빙 토폴로지를 MLPerf의 closed-division 규칙을 통과시키기까지의 엔지니어링을 다룹니다. 동반 글은 우리의 AMD MI355X 제출을 다룹니다.
NVIDIA GB200 NVL72 기반 MLPerf Inference v6.1 결과
모든 결과는 MLPerf Inference v6.1, closed division, Crusoe Cloud의 NVIDIA GB200 NVL72(gb200-186gb-nvl-ib.4x 인스턴스, 노드당 4x GB200, GPU당 1200W TGP)입니다.
모델
시스템
시나리오
결과
gpt-oss-120b
4x GB200 NVL72
Offline
50,912.9 tok/s
Server
50,613.5 tok/s
Qwen3-VL-235B-A22B
4x GB200 NVL72
Offline
58.40 samples/s
Server
49.70 queries/s
Interactive
3.99 queries/s
16x GB200 NVL72
Offline
230.09 samples/s
Server
198.94 queries/s
Interactive
39.85 queries/s
gpt-oss-120b는 두 시나리오 모두에서 GPU당 초당 약 12,700 토큰으로 계산됩니다. MLPerf는 Server 시나리오에서 지연 시간에 제한된 처리량을 측정하므로, 흥미로운 숫자는 비율입니다: 우리의 Server 결과는 Offline의 0.6% 이내에 들어옵니다. 이 플랫폼에서 지연 SLA 비용은 거의 없습니다. 최대 배치 처리량을 내는 동일 노드가 v6.1 한도 내에서 p99 time-to-first-token과 time-per-output-token을 유지하며, 이는 프로덕션 배포 규모 산정 시 중요한 속성입니다.
모든 실행은 closed-division 정확도를 통과했습니다: gpt-oss-120b는 82.30% 하한 대비 83.69% exact-match(Offline) 및 83.64%(Server), Qwen3-VL은 78.24% 하한 대비 78.56%에서 78.82% 사이의 hierarchical F1을 기록했습니다. gpt-oss-120b는 컴플라이언스 감사 TEST07 및 TEST09를 통과했습니다.
공개된 결과는 MLCommons MLPerf Inference v6.1 closed-division 결과에서 확인할 수 있습니다.
소프트웨어 구성
구성 요소
gpt-oss-120b
Qwen3-VL-235B-A22B
추론 스택
TensorRT-LLM (feat/1.2-mlpinf) via trtllm-serve
vLLM (CentML mlperf-inf-mm-q3vl) + NVIDIA Dynamo (mlperf-v6.0-dynamo-v0.8.0)
CUDA 스택
TensorRT 10.14, CUDA 13.1, NVIDIA cuDNN 9.17
NVIDIA Dynamo with NIXL KV transfer
노드 유형
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace™ CPU, aarch64
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace CPU, aarch64
패브릭
NVL72 랙 내 NVLink
NVL72 랙 내 NVLink
오케스트레이션
Crusoe Managed Kubernetes (단일 노드)
NVIDIA sflow를 통한 Crusoe Managed Slurm 클러스터(멀티노드), VAST의 공유 NFS
Prefill/decode 분리: 토폴로지만으로 2.5배 이득
16-GPU Qwen3-VL 배포는 우리가 제출한 것 중 아키텍처적으로 가장 흥미로운 시스템입니다. 4개의 독립적인 노드 크기 복제본을 실행하는 대신, 우리는 NVIDIA Dynamo를 prefill/decode 분리 토폴로지로 배포했습니다: prefill 워커(TP2)는 프롬프트 처리를 담당하고, decode 워커(TP1)는 토큰 생성을 담당하며, KV 캐시는 NVIDIA Inference Xfer Library(NIXL)를 통해 이들 사이를 이동합니다.
분리가 중요한 이유는 prefill과 decode가 반대의 하드웨어 요구를 갖기 때문입니다. Prefill은 컴퓨트 바운드이며 긴 프롬프트에 걸쳐 병렬화됩니다; decode는 메모리 대역폭 바운드이며 최대 배치 상주를 원합니다. 두 단계를 동일 GPU에서 서빙하면 타협이 강제됩니다; 분리하면 각 풀이 자연스러운 동작점에서 실행될 수 있습니다. 이는 대규모 프로덕션 추론 시스템 뒤의 동일한 아키텍처이며, 타이트한 p99 지연 한도를 가진 MLPerf의 Interactive 시나리오는 정확히 이점이 나타나는 곳입니다.
우리의 첫 Interactive 실행은 75:25 prefill:decode GPU 분할을 사용했고 15.85 queries/s를 측정했습니다. 프로파일링 결과 decode가 병목이었습니다: prefill GPU가 p99 한도 아래에서 decode GPU가 토큰을 스트리밍할 수 있는 것보다 더 빠르게 프롬프트를 완료했습니다. 50:50으로 재조정하고 지속 가능한 쿼리 속도를 이진 탐색하여 제출을 39.85 queries/s로 끌어올렸으며, 이는 하드웨어 변경 없이 약 2.5배 개선입니다. 최종 토폴로지는 3개 노드에 걸친 6개의 prefill 워커와 네 번째 노드의 4개의 decode 워커였습니다. 현재까지 이 결과는 GB200 NVL72 시스템에서 해당 라운드의 최고 Qwen3-VL Interactive 점수입니다.
분리형 서빙을 배포하는 모든 사람을 위한 두 가지 시사점:
prefill:decode 비율은 워크로드 의존적이며 스윕할 가치가 있습니다.
지연 바운드 Interactive 워크로드에 대한 최적 분할은 처리량 지향 기본값과 매우 달랐습니다.
확장은 선형이었습니다.
4개 노드에서 Offline 처리량(230.09 samples/s)은 3개 노드 실행에서의 선형 외삽(172.35가 229.8을 예측)과 일치했습니다. Server는 설정 수정 후 첫 시도에서 목표 200 대비 198.94를 기록했으며, 이는 구성된 속도의 99.5%입니다.
프로덕션이 실행되는 방식으로 MLPerf 실행
우리는 오늘날 고객 워크로드를 구동하는 오케스트레이션 서비스에서 의도적으로 이 벤치마크를 실행했습니다.
단일 노드 gpt-oss-120b 실행은 Crusoe Managed Kubernetes에서 실행되었으며, 이는 우리의 다른 MLPerf 작업과 동일한 플랫폼이고 Crusoe Managed Metrics를 통한 동일한 관측성을 갖습니다. 멀티노드 Qwen3-VL 실행은 NVIDIA의 sflow 런처를 사용하는 Crusoe Managed Slurm 클러스터에서 실행되었으며, 모델 저장소는 VAST 기반 공유 NFS에 있고 타이밍 실행 중 스토리지 핫 패스가 네트워크를 건너지 않도록 노드별 NVMe로 스테이징되었습니다.
다음 단계
이 제출은 Crusoe에서 NVIDIA Accelerated Computing Platform을 두 가지 규모로 보여줍니다: 공개된 참조 처리량을 초과하는 단일 GB200 노드, 그리고 NVLink와 NVIDIA Quantum InfiniBand를 가로질러 선형으로 확장되는 멀티노드 분리 토폴로지. 모든 Crusoe 클러스터는 고객이 만지기 전에 브링업 중 엔드투엔드로 검증되며, MLPerf는 그 과정의 한 도구입니다.
우리는 GB200 NVL72, GB300 NVL72 및 그 이후에 걸쳐 성능과 소프트웨어 공동 설계를 위해 NVIDIA와 계속 협력하고 있으며, Inference와 Training 워크로드 모두에 대한 결과를 계속 게시할 것입니다.
이 규모로 서빙할 준비가 되셨나요?
Crusoe Cloud의 NVIDIA GB200 NVL72에서 추론 실행에 대해 우리 팀과 상담하세요.
최신 글
2026년 10월 9일
Crusoe의 추론 엔진은 얼마나 친환경적인가?
2026년 10월 9일
Lone Star State에서 커뮤니티 구축
2026년 10월 7일
Triton을 사용한 GPU 프로그래밍, 파트 1
멋진 것을 만들 준비가 되셨나요?
문의하기
브리프용 요약 초안
Crusoe가 NVIDIA GB200 NVL72에서 MLPerf Inference v6.1 결과를 발표했으며, Qwen3-VL-235B-A22B에서 NVIDIA 참조치를 상회하고 4노드 확장이 선형임을 보였다. prefill:decode 비율을 50:50으로 조정해 Interactive 처리량이 약 2.5배 향상됐다. 메모리 제품·용량·공급업체에 대한 직접 언급은 없어 메모리 산업 영향은 확인되지 않는다.
원문 텍스트
원문 열기 ↗Cloud
Engineering
September 16, 2026
Crusoe's MLPerf Inference v6.1 results on NVIDIA GB200 NVL72
Crusoe's debut MLPerf Inference v6.1 submission benchmarked gpt-oss-120b and Qwen3-235B on NVIDIA GB200 NVL72 on Crusoe Cloud. They run on the same Managed Kubernetes and Slurm infrastructure customers use.
Young Jeong
Staff Solutions Engineer
Alex Akesson
Principal Product Manager
September 16, 2026
Table of contents
This is some text inside of a div block.
Share:
Today, MLCommons published the MLPerf® Inference v6.1 results, and with them Crusoe's first-ever MLPerf submission. We submitted results for gpt-oss-120b and Qwen3-VL-235B-A22B running on NVIDIA GB200 NVL72 on Crusoe Cloud, validating that the NVIDIA Accelerated Computing Platform, from NVIDIA CUDA® and NVIDIA TensorRT-LLM through NVIDIA Dynamo, NVIDIA NVLink™, and
NVIDIA Quantum InfiniBand
, runs exactly as our customers consume it.
Three results matter more than the raw numbers:
We exceeded NVIDIA's published single-node reference for Qwen3-VL-235B-A22B.
NVIDIA's guidance for a 4x GB200 NVL72 node was roughly 56 samples/s Offline and 46 queries/s Server. Our submission measured 58.40 and 49.70. Our gpt-oss-120b submission is likewise the highest per-GPU GB200 NVL72 result in the closed division, at roughly 12,700 tokens per second per GPU.
Scaling from one node to four was effectively linear.
Qwen3-VL Server throughput at 16 GPUs came in at 4.00x the 4-GPU result, and Offline landed within a fraction of a percent of the linear extrapolation. Rack-scale NVLink and NVIDIA Quantum InfiniBand make the fabric a non-factor for this workload shape.
A topology tuning pass produced a ~2.5x Interactive throughput gain on identical hardware.
By rebalancing the prefill:decode GPU split in a disaggregated NVIDIA Dynamo deployment from 75:25 to 50:50, Qwen3-VL Interactive throughput on 16 GPUs went from 15.85 to 39.85 queries/s. Same GPUs, same fabric, dramatically more serving capacity.
This post covers the results, the methodology behind them, and the engineering it took to get a multi-node prefill/decode-disaggregated serving topology through MLPerf's closed-division rules. A companion post covers our AMD MI355X submission.
MLPerf Inference v6.1 results on NVIDIA GB200 NVL72
All results are MLPerf Inference v6.1, closed division, on NVIDIA GB200 NVL72 on Crusoe Cloud (gb200-186gb-nvl-ib.4x instances, 4x GB200 per node, 1200W TGP per GPU).
Model
System
Scenario
Result
gpt-oss-120b
4x GB200 NVL72
Offline
50,912.9 tok/s
Server
50,613.5 tok/s
Qwen3-VL-235B-A22B
4x GB200 NVL72
Offline
58.40 samples/s
Server
49.70 queries/s
Interactive
3.99 queries/s
16x GB200 NVL72
Offline
230.09 samples/s
Server
198.94 queries/s
Interactive
39.85 queries/s
gpt-oss-120b works out to roughly 12,700 tokens per second per GPU in both scenarios. MLPerf measures latency-bounded throughput in the Server scenario, so the interesting number is the ratio: our Server result lands within 0.6% of Offline. The latency SLA costs almost nothing on this platform. The same node that maxes batch throughput also holds p99 time-to-first-token and time-per-output-token under the v6.1 limits, which is the property that matters when sizing a production deployment.
All runs cleared closed-division accuracy: gpt-oss-120b at 83.69% exact-match (Offline) and 83.64% (Server) against an 82.30% floor, and Qwen3-VL between 78.56% and 78.82% hierarchical F1 against a 78.24% floor. gpt-oss-120b passed compliance audits TEST07 and TEST09.
Published results are available in the
MLCommons MLPerf Inference v6.1 closed-division results.
Software configuration
Component
gpt-oss-120b
Qwen3-VL-235B-A22B
Inference stack
TensorRT-LLM (feat/1.2-mlpinf) via trtllm-serve
vLLM (CentML mlperf-inf-mm-q3vl) + NVIDIA Dynamo (mlperf-v6.0-dynamo-v0.8.0)
CUDA stack
TensorRT 10.14, CUDA 13.1, NVIDIA cuDNN 9.17
NVIDIA Dynamo with NIXL KV transfer
Node type
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace™ CPU, aarch64
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace CPU, aarch64
Fabric
NVLink within the NVL72 rack
NVLink within the NVL72 rack
Orchestration
Crusoe Managed Kubernetes (single node)
Crusoe Managed Slurm cluster via NVIDIA sflow (multi-node), shared NFS on VAST
Prefill/decode disaggregation: a 2.5x gain from topology alone
The 16-GPU Qwen3-VL deployment is the most architecturally interesting system we submitted. Rather than running 4 independent node-sized replicas, we deployed NVIDIA Dynamo in a prefill/decode-disaggregated topology: prefill workers (TP2) handle prompt processing, decode workers (TP1) handle token generation, and KV cache moves between them over NVIDIA Inference Xfer Library (NIXL).
Disaggregation matters because prefill and decode have opposite hardware appetites. Prefill is compute-bound and parallelizes across long prompts; decode is memory-bandwidth-bound and wants maximum batch residency. Serving both phases on the same GPUs forces a compromise; splitting them lets each pool run at its natural operating point. This is the same architecture behind large-scale production inference systems, and MLPerf's Interactive scenario, with its tight p99 latency limits, is precisely where it pays off.
Our first Interactive run used a 75:25 prefill:decode GPU split and measured 15.85 queries/s. Profiling showed decode was the bottleneck: prefill GPUs finished prompts faster than decode GPUs could stream tokens under the p99 limit. Rebalancing to 50:50 and binary-searching the sustainable query rate brought the submission to 39.85 queries/s, a roughly 2.5x improvement with zero hardware changes. The final topology was 6 prefill workers across 3 nodes and 4 decode workers on the fourth. To date, that result is the highest Qwen3-VL Interactive score in the round on the GB200 NVL72 system.
Two takeaways for anyone deploying disaggregated serving:
The prefill:decode ratio is workload-dependent and worth sweeping.
The optimal split for a latency-bound Interactive workload was very different from the throughput-oriented defaults.
Scaling was linear.
Offline throughput at 4 nodes (230.09 samples/s) matched the linear extrapolation from the 3-node run (172.35 predicted 229.8). Server hit 198.94 against a target of 200 on the first attempt after setup fixes, 99.5% of the configured rate.
Running MLPerf the way production runs
We deliberately ran these benchmarks on the orchestration services that power our customers’ workloads today.
The single-node gpt-oss-120b runs executed on
Crusoe Managed Kubernetes
, the same platform as our other MLPerf work, with the same observability through Crusoe Managed Metrics. The multi-node Qwen3-VL runs executed on a
Crusoe Managed Slurm
cluster using NVIDIA's sflow launcher, with the model repository on VAST-backed shared NFS and staged to per-node NVMe so the storage hot path never crossed the network during timed runs.
What comes next
This submission demonstrates the NVIDIA Accelerated Computing Platform on Crusoe at two scales: a single GB200 node exceeding published reference throughput, and a multi-node disaggregated topology scaling linearly across NVLink and NVIDIA Quantum InfiniBand. Every Crusoe cluster is validated end to end during bring-up before a customer touches it, and MLPerf is one instrument in that process.
We are continuing to work with NVIDIA on performance and software co-design across GB200 NVL72, GB300 NVL72, and beyond, and we will keep publishing results for both Inference and Training workloads.
Ready to serve at this scale?
Talk to our team
about running inference on NVIDIA GB200 NVL72 on
Crusoe Cloud
.
Latest articles
October 9, 2026
How green is Crusoe's inference engine?
October 9, 2026
Building community in the Lone Star State
October 7, 2026
GPU programming with Triton, part 1
Are you ready to build something amazing?
Contact Us
Engineering
September 16, 2026
Crusoe's MLPerf Inference v6.1 results on NVIDIA GB200 NVL72
Crusoe's debut MLPerf Inference v6.1 submission benchmarked gpt-oss-120b and Qwen3-235B on NVIDIA GB200 NVL72 on Crusoe Cloud. They run on the same Managed Kubernetes and Slurm infrastructure customers use.
Young Jeong
Staff Solutions Engineer
Alex Akesson
Principal Product Manager
September 16, 2026
Table of contents
This is some text inside of a div block.
Share:
Today, MLCommons published the MLPerf® Inference v6.1 results, and with them Crusoe's first-ever MLPerf submission. We submitted results for gpt-oss-120b and Qwen3-VL-235B-A22B running on NVIDIA GB200 NVL72 on Crusoe Cloud, validating that the NVIDIA Accelerated Computing Platform, from NVIDIA CUDA® and NVIDIA TensorRT-LLM through NVIDIA Dynamo, NVIDIA NVLink™, and
NVIDIA Quantum InfiniBand
, runs exactly as our customers consume it.
Three results matter more than the raw numbers:
We exceeded NVIDIA's published single-node reference for Qwen3-VL-235B-A22B.
NVIDIA's guidance for a 4x GB200 NVL72 node was roughly 56 samples/s Offline and 46 queries/s Server. Our submission measured 58.40 and 49.70. Our gpt-oss-120b submission is likewise the highest per-GPU GB200 NVL72 result in the closed division, at roughly 12,700 tokens per second per GPU.
Scaling from one node to four was effectively linear.
Qwen3-VL Server throughput at 16 GPUs came in at 4.00x the 4-GPU result, and Offline landed within a fraction of a percent of the linear extrapolation. Rack-scale NVLink and NVIDIA Quantum InfiniBand make the fabric a non-factor for this workload shape.
A topology tuning pass produced a ~2.5x Interactive throughput gain on identical hardware.
By rebalancing the prefill:decode GPU split in a disaggregated NVIDIA Dynamo deployment from 75:25 to 50:50, Qwen3-VL Interactive throughput on 16 GPUs went from 15.85 to 39.85 queries/s. Same GPUs, same fabric, dramatically more serving capacity.
This post covers the results, the methodology behind them, and the engineering it took to get a multi-node prefill/decode-disaggregated serving topology through MLPerf's closed-division rules. A companion post covers our AMD MI355X submission.
MLPerf Inference v6.1 results on NVIDIA GB200 NVL72
All results are MLPerf Inference v6.1, closed division, on NVIDIA GB200 NVL72 on Crusoe Cloud (gb200-186gb-nvl-ib.4x instances, 4x GB200 per node, 1200W TGP per GPU).
Model
System
Scenario
Result
gpt-oss-120b
4x GB200 NVL72
Offline
50,912.9 tok/s
Server
50,613.5 tok/s
Qwen3-VL-235B-A22B
4x GB200 NVL72
Offline
58.40 samples/s
Server
49.70 queries/s
Interactive
3.99 queries/s
16x GB200 NVL72
Offline
230.09 samples/s
Server
198.94 queries/s
Interactive
39.85 queries/s
gpt-oss-120b works out to roughly 12,700 tokens per second per GPU in both scenarios. MLPerf measures latency-bounded throughput in the Server scenario, so the interesting number is the ratio: our Server result lands within 0.6% of Offline. The latency SLA costs almost nothing on this platform. The same node that maxes batch throughput also holds p99 time-to-first-token and time-per-output-token under the v6.1 limits, which is the property that matters when sizing a production deployment.
All runs cleared closed-division accuracy: gpt-oss-120b at 83.69% exact-match (Offline) and 83.64% (Server) against an 82.30% floor, and Qwen3-VL between 78.56% and 78.82% hierarchical F1 against a 78.24% floor. gpt-oss-120b passed compliance audits TEST07 and TEST09.
Published results are available in the
MLCommons MLPerf Inference v6.1 closed-division results.
Software configuration
Component
gpt-oss-120b
Qwen3-VL-235B-A22B
Inference stack
TensorRT-LLM (feat/1.2-mlpinf) via trtllm-serve
vLLM (CentML mlperf-inf-mm-q3vl) + NVIDIA Dynamo (mlperf-v6.0-dynamo-v0.8.0)
CUDA stack
TensorRT 10.14, CUDA 13.1, NVIDIA cuDNN 9.17
NVIDIA Dynamo with NIXL KV transfer
Node type
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace™ CPU, aarch64
gb200-186gb-nvl-ib.4x: 4x NVIDIA GB200 (186 GB), 2x NVIDIA Grace CPU, aarch64
Fabric
NVLink within the NVL72 rack
NVLink within the NVL72 rack
Orchestration
Crusoe Managed Kubernetes (single node)
Crusoe Managed Slurm cluster via NVIDIA sflow (multi-node), shared NFS on VAST
Prefill/decode disaggregation: a 2.5x gain from topology alone
The 16-GPU Qwen3-VL deployment is the most architecturally interesting system we submitted. Rather than running 4 independent node-sized replicas, we deployed NVIDIA Dynamo in a prefill/decode-disaggregated topology: prefill workers (TP2) handle prompt processing, decode workers (TP1) handle token generation, and KV cache moves between them over NVIDIA Inference Xfer Library (NIXL).
Disaggregation matters because prefill and decode have opposite hardware appetites. Prefill is compute-bound and parallelizes across long prompts; decode is memory-bandwidth-bound and wants maximum batch residency. Serving both phases on the same GPUs forces a compromise; splitting them lets each pool run at its natural operating point. This is the same architecture behind large-scale production inference systems, and MLPerf's Interactive scenario, with its tight p99 latency limits, is precisely where it pays off.
Our first Interactive run used a 75:25 prefill:decode GPU split and measured 15.85 queries/s. Profiling showed decode was the bottleneck: prefill GPUs finished prompts faster than decode GPUs could stream tokens under the p99 limit. Rebalancing to 50:50 and binary-searching the sustainable query rate brought the submission to 39.85 queries/s, a roughly 2.5x improvement with zero hardware changes. The final topology was 6 prefill workers across 3 nodes and 4 decode workers on the fourth. To date, that result is the highest Qwen3-VL Interactive score in the round on the GB200 NVL72 system.
Two takeaways for anyone deploying disaggregated serving:
The prefill:decode ratio is workload-dependent and worth sweeping.
The optimal split for a latency-bound Interactive workload was very different from the throughput-oriented defaults.
Scaling was linear.
Offline throughput at 4 nodes (230.09 samples/s) matched the linear extrapolation from the 3-node run (172.35 predicted 229.8). Server hit 198.94 against a target of 200 on the first attempt after setup fixes, 99.5% of the configured rate.
Running MLPerf the way production runs
We deliberately ran these benchmarks on the orchestration services that power our customers’ workloads today.
The single-node gpt-oss-120b runs executed on
Crusoe Managed Kubernetes
, the same platform as our other MLPerf work, with the same observability through Crusoe Managed Metrics. The multi-node Qwen3-VL runs executed on a
Crusoe Managed Slurm
cluster using NVIDIA's sflow launcher, with the model repository on VAST-backed shared NFS and staged to per-node NVMe so the storage hot path never crossed the network during timed runs.
What comes next
This submission demonstrates the NVIDIA Accelerated Computing Platform on Crusoe at two scales: a single GB200 node exceeding published reference throughput, and a multi-node disaggregated topology scaling linearly across NVLink and NVIDIA Quantum InfiniBand. Every Crusoe cluster is validated end to end during bring-up before a customer touches it, and MLPerf is one instrument in that process.
We are continuing to work with NVIDIA on performance and software co-design across GB200 NVL72, GB300 NVL72, and beyond, and we will keep publishing results for both Inference and Training workloads.
Ready to serve at this scale?
Talk to our team
about running inference on NVIDIA GB200 NVL72 on
Crusoe Cloud
.
Latest articles
October 9, 2026
How green is Crusoe's inference engine?
October 9, 2026
Building community in the Lone Star State
October 7, 2026
GPU programming with Triton, part 1
Are you ready to build something amazing?
Contact Us