MEMORY INDUSTRY INTELLIGENCE

초당 575만 토큰 서빙: AMD MI355X에서 Crusoe의 MLPerf Inference v6.1 결과

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.10 20:09사용자 검토 전 초안

원문 제목: Serving 5.75 million tokens per second: Crusoe's MLPerf Inference v6.1 results on AMD MI355X

핵심 요약

Crusoe는 MLPerf Inference v6.1에 gpt-oss-120b와 DeepSeek-R1을 512개 AMD Instinct MI355X GPU로 실행한 결과를 제출했으며, 이는 GPU 수 기준 MLPerf 역사상 최대 규모의 MI355X 추론 제출이다. 8개 GPU에서 512개 GPU로 확장할 때 처리량이 이상적 수준의 90% 이상으로 선형 확장되었고, 표준 400Gb 이더넷만 사용했으며 InfiniBand나 RoCE, 노드 간 all-reduce가 필요하지 않았다. 각 MI355X는 288GB HBM3E를 탑재해 DeepSeek-R1(671B MoE)을 단일 8-GPU 노드에 수용하고 gpt-oss-120b는 단일 GPU에 MXFP4로 수용해 512개의 독립 복제본을 구성했다. 서버 시나리오에서 gpt-oss-120b는 약 5.39M tok/s, DeepSeek-R1은 약 2.40M tok/s를 기록했고, 두 모델 모두 closed-division 정확도 기준을 통과했다. 이 결과는 대규모 추론에서 HBM 용량이 샤딩 오버헤드와 프리미엄 패브릭 필요성을 줄여 TCO를 낮출 수 있음을 보여주는 사례로 제시되었다.

메모리 산업 영향 분석

이 문서는 AMD MI355X 기반 추론 클러스터의 MLPerf 결과로, 메모리 산업에는 HBM3E 용량(288GB/GPU)이 대형 MoE 모델을 단일 노드에 수용해 샤딩·노드 간 집합 통신·프리미엄 RDMA 패브릭 필요성을 줄인다는 점이 직접 관련된다. 원문은 MI355X에 GPU당 288GB HBM3E가 탑재된다고 명시하며, DeepSeek-R1의 671B MoE가 2.3TB HBM3E 단일 8-GPU 플랫폼에 들어가 1TB 이상이 KV 캐시로 남는다고 서술한다. 이는 HBM 용량이 추론 처리량과 TCO에 미치는 영향을 보여주는 사례이나, 특정 HBM 공급사·고객·거래 관계는 원문에 확인되지 않는다. 분석가 가설로는 HBM 용량 증가가 GPU당 동시 시퀀스와 토큰 처리량을 높여 HBM 비트 수요를 지지할 수 있으나, 동시에 모델을 단일 노드에 수용해 노드 간 통신용 네트워크·스위치 수요를 줄일 수 있다는 상반된 효과가 가능하다. 확인할 지표는 MI355X 출하량, HBM3E 탑재량·스택 구성, Crusoe 클러스터 도입 규모, MLPerf 제출의 실제 상용 배포 전환율이다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 13814자

클라우드
엔지니어링
2026년 9월 16일
초당 575만 토큰 서빙: AMD MI355X에서 Crusoe의 MLPerf Inference v6.1 결과
Crusoe의 MLPerf Inference v6.1 제출은 512개 AMD Instinct MI355X GPU에서 gpt-oss-120b와 DeepSeek-R1을 실행했으며, 이는 GPU 수 기준 MLPerf 역사상 최대 규모의 MI355X 추론 제출이다. 처리량은 Crusoe Managed Kubernetes 상에서 이더넷을 통해 8개에서 512개 GPU로 선형 확장되었다.
Martin Cala
Staff Solutions Engineer
2026년 9월 16일
목차
This is some text inside of a div block.
공유:
지난달 Crusoe는 gpt-oss-120b와 DeepSeek-R1에 대한 MLPerf
®
Inference v6.1 결과를 제출했으며, Crusoe Managed Kubernetes 내부의 AMD Instinct MI355X에서 512-GPU 규모로 실행했다. 현재까지 이는 GPU 수 기준 MLPerf 역사상 최대 규모의 MI355X 추론 제출이다.
추론 워크로드에서 원시 수치보다 더 주목할 만한 세 가지 결과가 있다:
처리량이 8개에서 512개 GPU로 선형 확장
되며 이상적 수준의 90% 이상을 기록했고, 이는 예상 부하에 따라 워크로드를 결정론적으로 오토스케일할 수 있음을 의미한다.
특수한 RDMA 인터커넥트가 필요하지 않았다.
이 규모의 추론은 반드시 노드 간 집합 통신을 필요로 하지는 않는다. 전체 512-GPU 실행은 표준 400 Gb 이더넷으로 조정되었으며, InfiniBand나 RoCE가 없었고 노드 간 all-reduce도 없었다.
벤치마크는 일반적인 Kubernetes 워크로드
로 실행되었으며, 고객이 프로덕션에서 사용하는 것과 동일한 Crusoe Managed Kubernetes 플랫폼, 동일한 관측성, 동일한 장애 처리를 사용했다. 우리는 매니페스트를 오픈소스로 공개했으므로 누구나 이를 재현할 수 있다.
이 블로그는 결과, 분석, 그리고 이러한 벤치마크를 수행하는 데 사용된 방법론을 다루므로, 독자가 직접 결과를 재현할 수 있다. 전체 저장소 재현은 GitHub
여기
에서 찾을 수 있다.
512개 AMD Instinct MI355X GPU에서의 MLPerf Inference v6.1 결과
모든 결과는 MLPerf Inference v6.1, closed division, 64개 노드에 걸친 512x AMD Instinct MI355X이며, Crusoe Managed Kubernetes가 주요 오케스트레이터이다.
모델
시나리오
총 처리량
GPU당 처리량
gpt-oss-120b
Offline
~5.75M tok/s
~11.2k tok/s
Server
~5.39M tok/s
~10.5k tok/s
DeepSeek-R1
Offline
~2.90M tok/s
~5.7k tok/s
Server
~2.40M tok/s
~4.7k tok/s
Server 결과는 v6.1 지연 SLA에 따라 측정되었다: gpt-oss-120b의 경우 p99 첫 토큰까지의 시간(time-to-first-token) 3.0초 미만 및 p99 출력 토큰당 시간(time-per-output-token) 80ms 미만. DeepSeek-R1의 경우 p99 TTFT 2.0초 미만 및 p99 TPOT 80ms 미만. 두 실행 모두 여유를 두고 통과했으며(gpt-oss는 2.69초 TTFT 및 36.5ms TPOT 측정, DeepSeek는 1.85초 TTFT 및 79.98ms TPOT 측정), 이는 더 높은 처리량을 얻기 위해 QPS를 더 높일 수 있었음을 시사한다.
두 모델 모두 closed-division 정확도를 통과했다: gpt-oss-120b는 82.3% 참조 하한 대비 83.7% exact-match, DeepSeek-R1은 80.7% exact-match(80.54% 하한)로 샘플당 평균 출력 길이 3,898 토큰이며, 이는 요구 범위인 3,497.6~4,274.85 내에 있다.
게시된 결과는
MLCommons MLPerf Inference v6.1 closed-division 결과
에서 확인할 수 있다.
소프트웨어 구성
구성 요소
gpt-oss-120b
DeepSeek-R1
추론 엔진
vLLM 0.22.1 (AITER, hipBLASLt)
SGLang 0.5.15.post1 (AITER, hipBLASLt, MoRI-EP)
AMD ROCm™
7.2.2
7.2.0
컨테이너 이미지
rocm/amd-mlperf (v6.1)
rocm/sgl-dev:v0.5.15.post1-rocm720-mi35x
노드 유형
mi355x-288gb-roce.8x: 8x MI355X(GPU당 288 GB HBM3E), 2x AMD EPYC 9575F, Ubuntu 24.04
mi355x-288gb-roce.8x: 8x MI355X(GPU당 288 GB HBM3E), 2x AMD EPYC 9575F, Ubuntu 24.04
패브릭
400 Gb 이더넷, 노드당 8x 400 Gb RoCE(총 3200 Gbps)
400 Gb 이더넷, 노드당 8x 400 Gb RoCE(총 3200 Gbps)
프로덕션 8개에서 512개 GPU로의 선형 확장
그림 1: GPU 수가 8개에서 512개로 증가함에 따른 총 출력 처리량.
프로덕션 추론 워크로드를 시뮬레이션하기 위해 일반적으로 수십, 수백, 심지어 수천 개의 GPU로 확장되는 오토스케일링 애플리케이션을 사용한다. 확장이 유지되는 이유는 이 형태의 추론이 대규모 병렬이며 vLLM 및 SGLang과 같은 추론 엔진이 동시성을 위해 설계되었기 때문이다. gpt-oss와 deepseek-r1 모두 단일 Mi355x 노드의 HBM 메모리 내에 들어갈 수 있으므로 복제본은 서로 동기화할 필요가 없으며, 노드 수에 따라 증가하는 all-reduce 노드 간 통신 비용이 없다. 유일한 노드 간 트래픽은 프런트엔드 입력 및 출력 스트림이다.
고객에게 이는 용량 계획의 결정론을 제공한다. 프로덕션 배포가 64개 GPU에서 주어진 초당 토큰 수를 유지한다면, 512-GPU 배포는 단순한 곱셈을 통해 규모를 산정할 수 있다. 오토스케일러와 Crusoe Managed Kubernetes를 결합하면 용량이 수요를 추적할 수 있다 - 피크 트래픽에 대해 확장하고, 줄어들면 축소하며, 피크를 영구적으로 프로비저닝하는 대신 차액을 지불한다.
총 소유 비용 절감
이 구성의 두 가지 속성이 규모에서 총 소유 비용을 낮춘다: 모든 모델을 단일 노드 내에 유지하는 메모리 용량, 그리고 프리미엄 스케일아웃 RDMA 패브릭에 대한 요구가 없다는 점이다.
288 GB HBM3E는 대형 MoE 모델을 단일 노드 내에 유지
DeepSeek-R1은 671B 파라미터 MoE로 토큰당 약 37B 파라미터가 활성화된다. FP8에서 가중치만 약 671 GB로, KV 캐시를 고려하면 일반적인 가속기 하나에 들어가지 않는다. MI355X에서는 동일 모델이 2.3 TB HBM3E를 갖춘 단일 8-GPU 플랫폼 내에 여유 있게 들어가며, KV 캐시 재사용을 위해
1테라바이트
이상이 남는다.
이 용량이 샤딩 오버헤드를 제거한다. DeepSeek-R1은 한 노드의 8개 GPU에 분산되지만, 그 통신의 모든 바이트는 노드 내 AMD Infinity Fabric(XGMI)에 머문다. 섀시를 떠나는 all-reduce도, 네트워크 홉을 넘는 파이프라인 단계 경계도, 집합 통신이 실행되는 동안 부분적으로 유휴 상태가 되는 GPU도 없다.
이 구성에서 어텐션은 텐서 병렬이 아닌 데이터 병렬로 실행되는데, DeepSeek의 Multi-head Latent Attention이 단일 압축 잠재 KV 헤드를 캐시하기 때문이다. 그 헤드를 텐서 병렬로 샤딩하면 모든 랭크에서 KV 캐시가 복제되어 동시 요청을 보유해야 할 메모리를 소비한다. 어텐션을 데이터 병렬로 유지하면 각 랭크가 자체 요청을 서비스하는 하나의 KV 캐시를 갖게 되어 배치 크기와 총 처리량을 보존한다. 전문가 레이어는 GPU에 걸쳐 분할되는데, 전문가 가중치가 메모리 풋프린트를 지배하고 토큰당 소수의 하위 집합만 활성화되므로 각 랭크는 자신이 소유한 전문가로 라우팅된 토큰에 대해서만 계산한다.
동일한 용량이 중형 모델에 서빙 여유를 제공
gpt-oss-120b 모델은 MXFP4의 MoE 가중치로 약 65 GB를 차지한다. 텐서 병렬 크기 1로 실행되어 전체 모델 가중치가 vRAM에 저장되므로, 우리는 512개의 완전히 독립적인 단일 GPU 복제본을 배포한다. 이는 샤딩이나 텐서 병렬을 방지했고, 순방향 패스에 집합 통신이 없으며, 복제본 간 조정이 전혀 없다. 플릿의 모든 GPU가 독립적인 서빙 단위이다.
여기서 288 GB가 바꾸는 것은 모델이 들어가는지 여부뿐만 아니라 그와 함께 무엇이 들어가는지이다. 가중치와 활성화 이후 각 MI355X는 KV 캐시를 위해 200 GB 이상을 사용할 수 있으며, 이는 80 GB 카드의 약 15 GB와 비교된다. KV 캐시 용량이 복제본당 동시 시퀀스를 제한하고, 동시 시퀀스가 위 결과 표의 GPU당 처리량을 만들어낸다. 용량 여유는 GPU당 달러당 초당 토큰으로 직접 전환되며, 와트당 총 토큰 성능을 증가시킨다.
규모에서의 추론은 프리미엄 패브릭을 필요로 하지 않는다
복제본이 서로 동기화하지 않기 때문에 512-GPU 실행은 전적으로 RoCE 이더넷으로 조정되었다. 이 구성에는 InfiniBand가 없으며, 그 부재로 인해 테이블에 남겨진 처리량도 없다.
추론 중심 플릿을 구축하는 고객에게 이는 측정된 성능 저하 없이 클러스터 BOM에서 상당한 항목을 제거한다. 노드 내부의 스케일업 패브릭이 중요한 작업을 수행하고, 스케일아웃 패브릭은 토큰만 이동하면 된다.
Kubernetes에서 MLPerf를 실행한 이유
MLPerf Inference의 참조 하네스는 단일 머신, 베어 메탈 또는 VM에서 Docker 컨테이너로 실행되도록 구축되었다. 그 모델은 64개 노드에 걸친 512-GPU 실행으로 잘 확장되지 않으며, 더 중요하게는 Crusoe에서 프로덕션 추론이 실제로 운영되는 방식이 아니다.
64개의 SSH 세션을 수동으로 오케스트레이션하는 대신, 우리는 벤치마크를 Kubernetes 네이티브 매니페스트로 재표현했다.
그 선택은 세 가지를 직접 제공했다:
확장.
동일한 워커 Job이 하나의 인자를 변경하여
parallelism=1
또는
parallelism=64
로 실행된다. 8-GPU 스모크 테스트와 512-GPU 제출 실행은 동일한 매니페스트이다.
관측성.
Pod 로그, 이벤트, 리소스 메트릭이 별도 설정 없이 표준 Kubernetes 표면을 통해 도착한다. 우리는 Crusoe의 Managed Metrics를 기반으로 Grafana 솔루션을 게시하고 유지 관리하므로 고객은 첫날부터 동일한 뷰를 얻는다.
장애 허용.
충돌한 복제본은 죽은 실행이 아니라 재시작된 Pod이다. 지연자는 복제본별 처리량에서 보이며 우아하게 퇴출될 수 있다.
실제 Crusoe Command Center
그림 2: 512-GPU gpt-oss-120b 오프라인 실행 중 총 GPU 및 메모리 활용도, Crusoe Managed Metrics를 통한 Grafana.
위 스크린샷은 Crusoe Managed Metrics로 구축된 Grafana 대시보드이다. 우리는 Crusoe VPC의 컴퓨트, 스토리지, 네트워크 리소스 전반에 걸쳐 인프라 로그를 수집하는 관리형 PromQL 엔드포인트를 노출하며, 이는 타사 관측성 대시보드를 채우는 데 사용될 수 있다. 평가 중에 우리는 이 메트릭을 사용하여 실행 중인 워크로드의 성능과 상태를 추적했다.
메트릭은 GPU 결함, 스토리지 병목, VPC 및 RDMA 네트워크 경합과 같은 인프라 장애에 대한 사용자 정의 경고를 설정하는 데에도 사용할 수 있다. Crusoe Watch Agent는 이제 인프라 리소스 전반에 기본 제공되어 AI 워크로드 관리의 프로덕션 경로를 간소화한다.
MLPerf 시스템 언더 테스트(SUT) 방법론
MLPerf의 LoadGen 바이너리는 단일 시스템 언더 테스트를 보아야 한다. 512개 GPU에서 우리는 그 SUT를 ZeroMQ(ZMQ)로 연결된 헤드와 워커로 구축했다.
그림 3: 512-GPU 분산 SUT. 하나의 전용 헤드 Pod가 LoadGen과 디스패치를 실행하고, 64개의 워커 Pod가 각각 한 노드의 8개 GPU를 소유하며 ZMQ로 연결된다.
추론 테스트를 확장할 때 우리는 워커가 규모에서 헤드 노드에 안정적으로 등록할 수 없는 오케스트레이션 계층의 병목에 부딪혔다. 이것이 우리의 최종 설계가 디스패치와 LoadGen만 수행하는 하나의 전용 헤드 Pod로 귀결된 이유이다. 이는 모델을 호스팅하지 않으므로 컴퓨트 병목이 되지 않는다. 64개의 워커 Pod 각각은 한 노드의 8개 GPU를 소유하며, 그 8개 GPU가 어떻게 사용되는지는 모델별로 다르다:
gpt-oss-120b:
노드당 텐서 병렬 크기 1의 8개 독립 단일 GPU 복제본, 총 512개 복제본. 모델이 MXFP4로 하나의 GPU에 들어가므로 샤딩이 필요하지 않다.
DeepSeek-R1:
한 노드의 8개 GPU 전체에 샤딩된 하나의 복제본(TP8, EP8, 8-way DP-attention), 총 64개 복제본. 671B MoE는 단일 288 GB GPU에 너무 크므로 샤딩되지만, 노드 내에서만 XGMI를 통해 이루어진다.
노드 간 집합 통신은 없다. 유일한 노드 간 트래픽은 프런트엔드 이더넷 패브릭의 ZMQ를 통한 토큰화된 입력 및 출력 스트림이다.
시나리오
Offline
은 지연 제약 없이 원시 배치 처리량을 측정한다. device_count는 노드 수의 8배이며, target_qps는 LoadGen이 20분 최소 지속 시간 창을 채울 충분한 샘플을 발행할 만큼 높게 조정된다.
Server
는 첫 토큰까지의 시간과 출력 토큰당 시간을 포함하는 p99 지연 SLA 하에서 처리량을 측정한다. target_qps는 SLA를 여전히 통과하는 가장 높은 값으로 조정된다. 그 지점을 넘기면 스케줄링 백로그가 생기고 p99가 무너진다. gpt-oss-120b의 지속 가능한 지점은 target_qps 4000이고, DeepSeek-R1은 target_qps 688이다.
직접 재현하기
위의 모든 것은
crusoe-mlperf-mi355x-inference-v6.1 저장소
에서 오픈소스이다.
파이프라인을 검증하기 위해 작게 시작한 다음 확장하라: N=1(8 GPU), 그다음 N=8(64 GPU), 그다음 N=64(512 GPU). 정확도와 규정 준수는 규모 독립적이므로 N=1에서 정확성을 확인할 수 있고 전체 플릿은 처리량 수치에만 사용되었다. 모델별 명령은 최상위 README에 있다.
다음 단계
이것이 Crusoe의 첫 MLPerf Inference 제출이지만, 마지막은 아닐 것이다. 우리는 고객과 파트너가 규모에서 성능을 검증하는 방식을 직접 볼 수 있도록 이러한 결과와 그 뒤의 코드를 게시한다. 모든 Crusoe 클러스터는 고객이 만지기 전에 브링업 중에 종단 간 검증된다. 여기에는 지속 부하 하의 워크로드 검증, GPU 및 HBM 스트레스 테스트, Infiniband / RoCE 패브릭 검증, 참조 워크로드 벤치마크가 포함된다. MLPerf는 그 과정의 한 도구이다. 고객이 클러스터를 인도받을 때 성능 특성은 이미 공개적이고 동료 검토된 표준에 대해 측정되어 있다.
Crusoe에서 우리는 ROCm 및 추론 엔진 튜닝부터 클러스터 브링업 및 네트워크 검증에 이르기까지 하드웨어 성능과 소프트웨어 공동 설계에 대해 AMD와 긴밀히 협력한다. 그 파트너십이 우리가 가장 안정적이고 성능이 뛰어난 AMD 클러스터를 구축하는 방식이며, MLPerf와 같은 공개 벤치마크가 이를 증명하는 방식이다. 이 글의 모든 수치는 동료 검토되었고 재현 가능하다: MLPerf 역사상 최대 규모의 MI355X 추론 제출. 우리는 AMD 및 파트너와 이 작업을 계속 심화하고, 추론 및 훈련 워크로드 모두에 대한 결과를 계속 게시하여 우리가 인용하는 수치가 항상 누구나 검증할 수 있는 수치가 되도록 할 것이다.
이 규모로 서빙할 준비가 되셨나요?
Crusoe Managed Kubernetes와 함께 AMD Instinct MI355X에서 추론을 실행하는 방법에 대해
우리 팀과 상담하세요
. 이미 Crusoe를 사용 중이신가요?
Crusoe Managed Kubernetes 문서
가 시작하는 데 도움이 될 것입니다.
최신 기사
2026년 10월 9일
Crusoe의 추론 엔진은 얼마나 친환경적인가?
2026년 10월 9일
텍사스주에서 커뮤니티 구축
2026년 10월 7일
Triton을 사용한 GPU 프로그래밍, 파트 1
멋진 것을 만들 준비가 되셨나요?
문의하기
브리프용 요약 초안
Crusoe가 512개 AMD MI355X로 MLPerf Inference v6.1에서 gpt-oss-120b 약 5.75M tok/s, DeepSeek-R1 약 2.90M tok/s(Offline)를 기록했다. GPU당 288GB HBM3E가 대형 MoE를 단일 노드에 수용해 노드 간 all-reduce와 프리미엄 RDMA 패브릭 없이 선형 확장이 가능했다는 점이 핵심이다. 메모리 관점에서는 HBM 용량이 추론 TCO와 처리량을 좌우하는 변수로 부각되나, 특정 공급사·고객 거래는 원문에서 확인되지 않는다.

원문 텍스트

원문 열기 ↗
Cloud
Engineering
September 16, 2026
Serving 5.75 million tokens per second: Crusoe's MLPerf Inference v6.1 results on AMD MI355X
Crusoe's MLPerf Inference v6.1 submission ran gpt-oss-120b and DeepSeek-R1 on 512 AMD Instinct MI355X GPUs, the largest MI355X inference entry by GPU count in MLPerf history. Throughput scaled linearly from 8 to 512 GPUs over Ethernet on Crusoe Managed Kubernetes.
Martin Cala
Staff Solutions Engineer
September 16, 2026
Table of contents
This is some text inside of a div block.
Share:
Last month, Crusoe submitted results to MLPerf
®
Inference v6.1 for gpt-oss-120b and DeepSeek-R1, running at 512-GPU scale on AMD Instinct MI355X inside Crusoe Managed Kubernetes. To date, this is the largest MI355X inference submission by GPU count in MLPerf history.
Three results are notable for inference workloads more than the raw numbers:
Throughput scales linearly from 8 to 512 GPUs
at over 90% of ideal, which means workloads can be autoscaled deterministically depending on expected load.
No exotic RDMA interconnect was required.
Inference at this scale does not necessarily require cross-node collectives. The entire 512-GPU run was coordinated over standard 400 Gb Ethernet, with no InfiniBand or RoCE and no cross-node all-reduce.
The benchmark ran as a typical Kubernetes workload
on the same Crusoe Managed Kubernetes platform our customers use in production, with the same observability and the same failure handling. We have open-sourced the manifests so you can reproduce any of it.
This blog covers the results, analysis, and the methodology used to conduct these benchmarks,  so you can reproduce the results yourself. The full repo reproduction can be found in GitHub
here
.
MLPerf Inference v6.1 results on 512 AMD Instinct MI355X GPUs
All results are MLPerf Inference v6.1, closed division, 512x AMD Instinct MI355X across 64 nodes with Crusoe Managed Kubernetes as the main orchestrator.
Model
Scenario
Total throughput
Per-GPU throughput
gpt-oss-120b
Offline
~5.75M tok/s
~11.2k tok/s
Server
~5.39M tok/s
~10.5k tok/s
DeepSeek-R1
Offline
~2.90M tok/s
~5.7k tok/s
Server
~2.40M tok/s
~4.7k tok/s
Server results were measured under the v6.1 latency SLAs: for gpt-oss-120b, p99 time-to-first-token under 3.0 s and p99 time-per-output-token under 80 ms. For DeepSeek-R1, p99 TTFT under 2.0 s and p99 TPOT under 80 ms. Both runs passed with headroom (gpt-oss measured 2.69 s TTFT and 36.5 ms TPOT; DeepSeek 1.85 s TTFT and 79.98 ms TPOT), which suggests we could have pushed QPS higher to get even more throughput.
Both models cleared closed-division accuracy: gpt-oss-120b at 83.7% exact-match against an 82.3% reference floor, and DeepSeek-R1 at 80.7% exact-match (80.54% floor) with a mean output length of 3,898 tokens per sample, inside the required 3,497.6 to 4,274.85 range.
Published results are available in the
MLCommons MLPerf Inference v6.1 closed-division results.
Software configuration
Component
gpt-oss-120b
DeepSeek-R1
Inference engine
vLLM 0.22.1 (AITER, hipBLASLt)
SGLang 0.5.15.post1 (AITER, hipBLASLt, MoRI-EP)
AMD ROCm™
7.2.2
7.2.0
Container image
rocm/amd-mlperf (v6.1)
rocm/sgl-dev:v0.5.15.post1-rocm720-mi35x
Node type
mi355x-288gb-roce.8x: 8x MI355X (288 GB HBM3E per GPU), 2x AMD EPYC 9575F, Ubuntu 24.04
mi355x-288gb-roce.8x: 8x MI355X (288 GB HBM3E per GPU), 2x AMD EPYC 9575F, Ubuntu 24.04
Fabric
400 Gb Ethernet, 8x 400 Gb RoCE per node (3200 Gbps aggregate)
400 Gb Ethernet, 8x 400 Gb RoCE per node (3200 Gbps aggregate)
Production Linear scaling from 8 to 512 GPUs
Figure 1: Aggregate output throughput as GPU count increases from 8 to 512.
To simulate production inference workloads, it is typical to have an autoscaling application that scales from tens, to hundreds, to even thousands of GPUs. The reason scaling holds is that inference at this shape is massively parallel and inference engines like vLLM and SGLang are designed for concurrency. Since both gpt-oss and deepseek-r1 could fit within the HBM memory of a single Mi355x node, replicas never need to synchronize with one another, so there are no all-reduce cross node communication costs that grow with node count. The only inter-node traffic is frontend input and output streams.
For customers this provides determinism for capacity planning. If your production deployment sustains a given tokens-per-second at 64 GPUs, you can size the 512-GPU deployment through straightforward multiplication. Combining an autoscaler with Crusoe Managed Kubernetes, that means capacity can track demand - scale out for peak traffic, scale back in when it subsides, and pay for the difference rather than provisioning for the peak permanently.
Lowering total cost of ownership
Two properties of this configuration drive down total cost of ownership at scale: memory capacity that keeps every model within a single node, and the absence of any requirement for a premium scale-out RDMA fabric.
288 GB of HBM3E keeps large MoE models inside one node
DeepSeek-R1 is a 671B-parameter MoE with roughly 37B parameters active per token. At FP8 the weights alone are approximately 671 GB, which does not fit within a typical accelerator once you account for KV cache. On MI355X the same model sits comfortably inside a single 8-GPU platform with 2.3 TB of HBM3E, leaving
over a
terabyte
for KV cache re-use.
That capacity is what eliminates the sharding overhead. DeepSeek-R1 is distributed across the 8 GPUs of one node, but every byte of that communication stays on the in-node AMD Infinity Fabric (XGMI). There is no all-reduce leaving the chassis, no pipeline stage boundary crossing a network hop, and no GPU sitting partially idle while collective communications run .
In this configuration, attention also runs data-parallel rather than tensor-parallel because DeepSeek's Multi-head Latent Attention caches a single compressed latent KV head. Sharding that head with tensor parallelism would replicate the KV cache on every rank, consuming memory that should be holding concurrent requests. Keeping attention data-parallel gives each rank one KV cache serving its own requests, which preserves batch size and total throughput. The expert layers are partitioned across GPUs because the expert weights dominate the memory footprint and only a small subset activates per token, so each rank computes only for the tokens routed to the experts it owns.
The same capacity buys serving headroom for mid-size models
The gpt-oss-120b model, with its MoE weights in MXFP4, occupies roughly 65 GB. It runs at tensor-parallel size 1 meaning the full model weights are stored in vRAM, so we deploy 512 fully independent single-GPU replicas. This prevented any sharding or tensor parallelism, no collectives in the forward pass, and no coordination between replicas at all. Every GPU in the fleet is an independent serving unit.
What 288 GB changes here is not only whether the model fits, but what fits alongside it. After weights and activations, each MI355X has over 200 GB available for KV cache, compared with roughly 15 GB on an 80 GB card. KV cache capacity is what caps concurrent sequences per replica, and concurrent sequences are what produce the per-GPU throughput in the results table above. Capacity headroom converts directly into tokens per second per dollar of GPU, and increases total token performance per Watt.
Inference at scale does not require a premium fabric
Because replicas never synchronize with one another, the 512-GPU runs coordinated entirely over RoCE Ethernet. There is no InfiniBand in this configuration and no throughput left on the table for its absence.
For customers building inference-primary fleets, that removes a substantial line item from the cluster bill of materials with no measured performance penalty. The scale-up fabric inside the node does the work that matters, and the scale-out fabric only needs to move tokens.
Why we ran MLPerf on Kubernetes
MLPerf Inference's reference harnesses are built to run as a Docker container on a single machine, bare metal, or a VM. That model does not scale well with a 512-GPU run across 64 nodes, and more importantly it is not how production inference is actually operated on Crusoe.
Rather than hand-orchestrate 64 SSH sessions, we re-expressed the benchmark as Kubernetes-native manifests.
That choice gave us three things directly:
Scaling.
The same worker Job runs at
parallelism=1
or
parallelism=64
by changing one argument. The 8-GPU smoke test and the 512-GPU submission run are the same manifest.
Observability.
Pod logs, events, and resource metrics arrive through the standard Kubernetes surface, with no bespoke setup. We publish and maintain a Grafana solution backed by Crusoe's Managed Metrics so customers get the same view on day one.
Fault tolerance.
A crashed replica is a restarted pod, not a dead run. Stragglers are visible in per-replica throughput and can be evicted gracefully.
Crusoe Command Center in practice
Figure 2: Aggregate GPU and Memory utilization during a 512-GPU gpt-oss-120b offline run, from Crusoe Managed Metrics via Grafana.
The screenshot above is a Grafana dashboard built with Crusoe Managed Metrics. We expose a managed PromQL endpoint that collects infrastructure logs across compute, storage, and network resources in your Crusoe VPC that can be used to populate third party observability dashboards. During evaluations, we used these metrics to track the performance and state of running workloads.
The metrics can also be used to set up custom alerts for infrastructure failures like GPU faults, storage bottlenecks, or VPC and RDMA Network contention. The Crusoe Watch Agent now ships by default across infrastructure resources, which streamlines the path to production for managing AI workloads.
MLPerf system under test (SUT) methodology
MLPerf's LoadGen binary must see a single system under test. At 512 GPUs we built that SUT as a ZeroMQ (ZMQ)-connected head plus workers.
Figure 3: The 512-GPU distributed SUT. One dedicated head pod runs LoadGen and dispatch; 64 worker pods each own one node's 8 GPUs, connected over ZMQ.
When scaling the inference tests, we ran into bottlenecks at the orchestration layer where workers could not reliably register with the head node at scale. This is why our final design resulted in one dedicated head pod which does only dispatch and LoadGen. It hosts no model, so it does not become a compute bottleneck. Each of the 64 worker pods owns one node's 8 GPUs, and how those 8 GPUs are used is model-specific:
gpt-oss-120b:
8 independent single-GPU replicas per node at tensor-parallel size 1, for 512 replicas total. The model fits in one GPU at MXFP4, so no sharding is needed.
DeepSeek-R1:
one replica sharded across all 8 GPUs of a node (TP8, EP8, 8-way DP-attention), for 64 replicas total. The 671B MoE is too large for a single 288 GB GPU, so it is sharded, but only within the node, over XGMI.
There are no cross-node collectives. The only inter-node traffic is tokenized input and output streams over ZMQ on the front-end ethernet fabric.
Scenarios
Offline
measures raw batch throughput with no latency constraint. device_count is 8 times the node count, and target_qps is tuned high enough that LoadGen issues sufficient samples to fill the 20-minute minimum-duration window.
Server
measures throughput under a p99 latency SLA covering time to first token and time per output token. target_qps is tuned to the highest value that still passes the SLA. Pushing it beyond that point creates a scheduling backlog and blows p99. For gpt-oss-120b the sustainable point is target_qps 4000, while for DeepSeek-R1 it is target_qps 688.
Reproduce this yourself
Everything above is open source in the
crusoe-mlperf-mi355x-inference-v6.1 repo
.
Start small to validate the pipeline, then scale: N=1 (8 GPUs), then N=8 (64 GPUs), then N=64 (512 GPUs). Accuracy and compliance are scale-independent, so correctness can be confirmed at N=1 and the full fleet spent only on throughput numbers. Per-model commands are in the top-level README.
What comes next
Although this is Crusoe's first MLPerf Inference submission, it will not be our last. We publish these results and the code behind them to provide tangible benchmarks that our customers and partners can use to see first hand how we validate performance at scale. Every Crusoe cluster is validated end to end during bring-up before a customer touches it. This includes workload validation under sustained load, GPU and HBM stress testing, Infiniband / RoCE fabric validation, and reference workload benchmarks. MLPerf is one instrument in that process. When a customer takes delivery of a cluster, the performance characteristics have already been measured against a public, peer-reviewed standard.
At Crusoe, we partner closely with AMD on hardware performance and software co-design, from ROCm and inference engine tuning down to cluster bring-up and network validation. That partnership is how we build the most reliable and performant AMD clusters, and public benchmarks like MLPerf are how we prove it. Every number in this post is peer-reviewed and reproducible: the largest MI355X inference submission in MLPerf history. We will keep deepening this work with AMD and our partners, and we will keep publishing results for both Inference and Training workloads, so the numbers we quote are always numbers anyone can verify.
Ready to serve at this scale?
Talk to our team
about running inference on AMD Instinct MI355X with Crusoe Managed Kubernetes. Already on Crusoe? The
Crusoe Managed Kubernetes documentation
will get you started.
Latest articles
October 9, 2026
How green is Crusoe's inference engine?
October 9, 2026
Building community in the Lone Star State
October 7, 2026
GPU programming with Triton, part 1
Are you ready to build something amazing?
Contact Us