MEMORY INDUSTRY INTELLIGENCE

할라페뇨의 첫 결과, AI 추론에서 업계 최고 속도와 효율성 입증

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.10 11:45사용자 검토 전 초안

원문 제목: Jalapeño’s first results show industry-leading speed and efficiency in AI inference

핵심 요약

OpenAI는 자사의 첫 맞춤형 추론 칩인 할라페뇨(Jalapeño)의 테스트 결과를 발표하며, 단일 아키텍처로 높은 처리량과 낮은 지연을 동시에 달성했다고 밝혔다. GPT-OSS 120B, DeepSeek R1, Kimi K2.5 1T 등 공개 모델에서 비교 시스템 대비 와트당 1.5~1.9배 높은 AI 작업량과 1.7~3.6배 낮은 종단 간 지연을 기록했다고 주장했다. 할라페뇨는 칩, 메모리, 네트워크, 소프트웨어, 랙 규모 시스템을 함께 설계한 결과이며, KV 캐시를 로컬에 유지하고 데이터 이동을 최소화하는 아키텍처를 채택했다고 설명했다. OpenAI는 올해 말까지 자사 컴퓨팅 인프라에 할라페뇨를 배치하기 시작할 계획이며, 2세대와 3세대 로드맵도 진행 중이라고 밝혔다. 또한 NVIDIA 등 파트너의 가속기를 훈련과 추론 모두에 계속 광범위하게 배치할 것이라고 덧붙였다.

메모리 산업 영향 분석

OpenAI의 자체 추론 칩 할라페뇨는 메모리 산업에 직접적인 영향을 줄 수 있는 몇 가지 구조적 변화를 시사한다. 첫째, 할라페뇨는 KV 캐시를 로컬에 유지하고 데이터 이동을 최소화하는 아키텍처를 강조하는데, 이는 HBM이나 고대역폭 메모리의 필요성을 줄일 수도 있고, 반대로 온칩 메모리나 근접 메모리의 중요성을 높일 수도 있다. 그러나 원문은 할라페뇨가 사용하는 메모리 종류(HBM, LPDDR, SRAM 등)나 용량을 명시하지 않았으므로, 메모리 계층에 대한 구체적 영향은 미확인이다. 둘째, 할라페뇨가 추론 효율을 높이면 AI 서비스 비용이 낮아져 수요가 증가할 수 있고, 이는 전체 메모리 수요를 늘릴 수 있다. 셋째, OpenAI는 NVIDIA 가속기를 계속 배치할 것이라고 밝혔으므로, HBM 수요가 즉시 감소한다고 단정할 수 없다. 넷째, 할라페뇨의 성능 수치는 자체 벤치마크와 SemiAnalysis의 InferenceX에 기반하며, 비교 시스템은 GB200/GB300으로 명시되어 있으나, 메모리 구성이나 HBM 탑재량은 언급되지 않았다. 따라서 메모리 산업에 대한 직접적 영향은 제한적으로 확인되며, 주로 추론 효율 향상이 AI 수요와 인프라 투자에 미치는 간접적 효과에 주목해야 한다. 확인할 지표로는 할라페뇨의 메모리 사양, OpenAI의 향후 가속기 조달 계획, HBM 수요 전망 등이 있다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 10763자

할라페뇨(Jalapeño)를 발표한 이후, 우리는 이 칩과 그 주변 시스템을 테스트해 왔다. 결과는 상당한 성능 향상을 보여준다: 할라페뇨는 단위 전력당 더 많은 AI 작업을 처리하면서 동시에 응답을 더 빠르게 반환할 수 있다. 할라페뇨는 기존 하드웨어 시스템이 종종 처리량과 지연 사이에서 타협해야 했던 것과 달리, 하나의 아키텍처로 더 높은 처리량과 더 낮은 지연을 모두 제공한다.

고객에게 이는 더 빠른 응답, 더 반응성이 높은 에이전트, 수요가 증가해도 더 안정적인 접근을 의미할 수 있다. 우리의 사명은 인공 일반 지능(AGI)이 인류 전체에 이익이 되도록 하는 것이다. 이러한 이점은 점점 더 능력 있는 AI를 더 저렴하고 더 널리 이용 가능하게 만드는 데 도움이 될 것이다.

OpenAI 모델들도 할라페뇨의 개발을 가속화했다. 이전 세대 모델들은 팀이 칩을 설계하고 초기 구동을 하는 데 도움을 주었고, 최신 모델들은 우리가 칩을 최적화하고 프로그래밍하는 속도를 높이고 있다. 할라페뇨의 성능은 GPT‑OSS 120B, DeepSeek R1, Kimi K2.5 1T에 걸쳐 나타나며, 이는 이 아키텍처가 OpenAI 내부와 외부에서 개발된 모델 모두에서 작동함을 보여준다. 세 모델 모두에서 할라페뇨는 최대 처리량 기준 비교 시스템보다 와트당 1.5~1.9배 더 많은 AI 작업을 제공했고, 종단 간 지연은 1.7~3.6배 더 낮았다. 고도로 상호작용적인 워크로드의 경우 2.1~4.1배 더 높은 성능을 제공했다.

할라페뇨는 또한 더 광범위한 풀스택 우위의 증거이기도 하다. OpenAI는 모델, 제품, 서빙 소프트웨어, 칩, 메모리, 네트워킹, 시스템을 함께 설계할 수 있으며, 실제 워크로드에서 배운 것을 사용해 스택의 모든 계층을 개선한다. 할라페뇨는 측정된 결과를 가진 자체 개발(first-party) 실리콘이며, 다세대 플랫폼의 시작이다. 앞으로 몇 달 안에 우리는 할라페뇨를 확장하여 고객에게 더 빠르고, 더 능력 있고, 더 효율적인 제품을 제공할 것이다.

## 할라페뇨의 성능 측정 방법

우리는 일치된 사용자 경험에서 성능을 평가하며, 각 시스템이 고객과 상호작용 에이전트가 요구하는 지연을 충족하면서 단위 전력당 얼마나 많은 유용한 AI 작업을 완료할 수 있는지 측정한다. 이는 특히 에이전트에 중요하며, 에이전트는 여러 단계를 순차적으로 완료해야 하므로 지연이 전체 작업에 걸쳐 누적될 수 있다.

할라페뇨가 실제로 어떻게 작동하는지 이해하기 위해, 우리는 AI 요청 서빙의 전체 과정을 측정하는 SemiAnalysis의 공개 벤치마크인 InferenceX에서 테스트했다. 우리는 할라페뇨를 테스트된 운영 범위 전반에 걸쳐 선도적인 상용 AI 시스템과 비교했으며, 고처리량 서빙부터 고도로 상호작용적인 저지연 사용까지 포함했다. 할라페뇨는 처리량, 전력 효율, 지연의 더 나은 조합을 제공했다. 성능이 때때로 칩당으로 보고되기도 하지만, 우리는 더 유용한 기준은 단위 전력당 성능이라고 믿는다.

### 할라페뇨는 이전 최고 TBT에서 격차를 벌린다

### 할라페뇨는 사용자당 더 많은 토큰을 제공한다

### 할라페뇨는 킬로와트당 더 많은 처리량을 제공한다

세 공개 모델 모두에서 할라페뇨는 테스트된 운영 범위 전반에 걸쳐 와트당 성능과 지연의 더 나은 조합을 제공하여 파레토 프론티어에 위치했다. 시스템을 일관되게 비교하기 위해, 우리는 각 가속기의 공개된 칩 전력 등급을 사용하여 결과를 정규화했다. 할라페뇨는 700와트로 평가되지만, 측정된 지속 전력은 테스트된 워크로드에서 550와트 이하를 유지했다.

할라페뇨는 GPT‑OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T 전반에서 강력한 성능을 보였다. 우리가 테스트한 가장 큰 공개 모델인 Kimi에서, 할라페뇨는 비교 시스템보다 약 1.5배 높은 최대 와트당 성능과 3.4배 낮은 종단 간 지연을 제공했다. 내부 테스트에서 할라페뇨의 우위는 최전선 OpenAI 모델에서 더욱 확대되었으며, 이는 워크로드가 더 커지고 더 까다로워질수록 이 아키텍처가 더 가치 있어진다는 것을 시사한다.

## 단일 칩 내에서 속도와 효율을 위한 아키텍처 설계

할라페뇨는 처음부터 다음과 같은 질문을 하며 설계되었다: 만약 하드웨어의 주요 임무가 현대와 미래의 언어 모델, 특히 상호작용 에이전트를 서빙하는 것이라면 어떤 하드웨어를 만들 것인가? 할라페뇨의 이점은 칩, 메모리, 네트워크, 소프트웨어, 랙 규모 시스템을 실제 언어 모델 워크로드에 맞춰 함께 설계한 데서 나온다. 언어 모델 추론은 서로 다른 병목을 가진 여러 단계를 거친다. 시스템이 프롬프트를 처리하는 프리필(prefill)은 컴퓨팅 집약적이며, 시스템이 응답을 토큰 단위로 생성하는 디코드(decode)는 메모리 대역폭에 더 제약을 받는다. 데이터가 코어와 칩 사이를 이동해야 할 때 통신도 지연을 추가할 수 있으며, 일부 처리 장치가 대기하는 동안 유휴 상태가 된다. 한 단계에서 뛰어난 시스템은 데이터를 기다리거나 서로 다른 자원 간에 모델 상태를 이동하는 동안 그 이점을 잃을 수 있다.

우리는 할라페뇨를 데이터 이동과 통신 지연을 최소화하도록 설계했다. 이는 응답 생성 중 사용되는 KV 캐시를 포함한 모델 상태가 명시적으로 배치되고 로컬에 유지될 수 있으며, 시스템이 각 추론 단계에 맞는 컴퓨팅, 메모리, 네트워킹의 올바른 조합을 활성화한다는 것을 의미한다. 네트워크는 아키텍처의 필수 요소이다. 그 큰 도메인은 전체 워크로드가 하나의 연결된 시스템 내에 머물 수 있게 하여 데이터 이동을 최소화하고 전체 요청이 처음부터 끝까지 빠르고 효율적으로 유지되도록 돕는다. 그 결과는 변화하는 모델 아키텍처를 지원하고, 프리필과 디코드 모두에서 뛰어나며, 둘 사이의 균형이 변함에 따라 적응할 수 있는 균형 잡히고 호환성 있는(fungible) 가속기이며, 이는 에이전트 워크로드의 정의적 특징이다.

## 우리는 AI를 사용해 칩을 설계했고, AI가 프로그래밍할 수 있도록 칩을 설계했다

AI는 할라페뇨 개발에 직접적인 역할을 했으며, 팀이 구현을 탐색하고 설계, 측정, 검증 루프를 단축하며 모델 워크로드를 지속적으로 반복함으로써 초기 설계에서 테이프아웃까지 9개월 만에 이동할 수 있게 했다. AI는 또한 칩의 산술 회로를 최적화하는 데 도움을 주어 팀이 일정에 맞춰 더 많은 컴퓨팅 성능을 칩에 담을 수 있게 했다.

할라페뇨는 인간과 AI 모두에게 명확하고 예측 가능한 프로그래밍 대상으로 설계되었다. 엔지니어는 로컬 텐서, 명시적 통신, 예측 가능한 동기화를 통해 작업을 설명할 수 있다. 그러면 AI는 해당 작업이 시스템 전체에 걸쳐 매핑, 배치, 스케줄링, 조정되는 방식을 최적화할 수 있다. 이 명확하고 예측 가능한 구조는 AI가 전통적으로 어려운 병렬 프로그래밍 문제를 다룰 수 있는 실마리를 제공한다.

각 새로운 모델 패밀리를 지원하려면 여전히 새로운 커널과 모델별 최적화가 필요하다. Codex와 GPT‑Astra를 사용하여, 팀은 할라페뇨의 원래 생산 계획에 포함되지 않았던 세 개의 오픈 웨이트 모델을 두 달 만에 고성능으로 끌어올렸다. 이는 아키텍처의 유연성과 AI가 우리의 프로그래밍을 돕는 속도를 모두 입증했다. 선택된 GPT‑OSS 어텐션 및 전문가 혼합(mixture-of-experts) 블록의 경우, AI가 생성한 구현이 기존 인간 전문가가 작성한 구현보다 1.5~1.8배 더 빠르게 실행되었다. 이 수치는 전체 모델이 아닌 선택된 블록에 적용되지만, 강력한 새로운 개발 루프를 가리킨다.

## 효율적이고 초고속 추론을 향한 길

AI 인프라는 그것이 가능하게 하는 유용한 실제 작업 때문에 가치가 있다. 동일한 전력과 하드웨어에서 더 많은 유용한 작업을 생산함으로써, 할라페뇨는 우리가 더 많은 수요를 서빙하고 성공적인 결과를 제공하는 비용을 낮추는 데 도움을 줄 수 있다. OpenAI에게 이는 유용한 작업과 수익이 서빙 비용보다 빠르게 성장할 수 있게 함으로써 운영 레버리지를 개선할 수 있다. 또한 더 넓은 채택과 더 나은 모델, 제품, 인프라에 대한 지속적인 투자를 지원할 수 있다. 더 빠른 추론은 더 빠른 반복과 새로운 사용 사례를 가능하게 할 수 있다.

할라페뇨는 효율적이고 저지연 추론이 가능한 범위를 확장한다:

* 이전에는 패스트 모드에서만 가능했던 효율로 초고속 모드 추론

* 이전에는 배치 모드에서만 가능했던 효율로 패스트 모드 추론

* 배치 모드 추론의 더 높은 효율

우리는 올해 말까지 OpenAI의 컴퓨팅 인프라 내에 할라페뇨를 배치하기 시작할 계획이다. 이는 다세대 로드맵의 첫 세대이다: 2세대는 개발이 깊숙이 진행 중이고, 3세대는 형태를 갖추고 있다. 각 세대는 우리가 배운 것을 기반으로 하여 효율성과 속도를 더욱 발전시킬 것이다.

AI에 대한 증가하는 수요를 충족하려면 이용 가능한 모든 소스에서 더 많은 컴퓨팅이 필요할 것이다. 우리는 훈련과 추론 워크로드 모두에 NVIDIA 및 기타 파트너의 가속기를 계속 광범위하게 배치할 것이다. 우리의 사명은 인공 일반 지능이 인류 전체에 이익이 되도록 하는 것이다.

배치를 준비하면서, 우리는 생산 품질 인증을 계속하고, 소프트웨어를 성숙시키고, 할라페뇨를 대규모로 운영할 준비를 하고, 더 많은 모델에서 성능을 검증하고 있다. 지금까지의 결과는 우리가 전체 시스템을 함께 설계할 때 무엇이 가능한지 보여준다: 더 반응성이 높고, 능력 있고, 에이전트적인 AI를 더 많은 사람에게 더 효율적으로 제공하는 것이다.

## 부록

### 할라페뇨는 GPT‑OSS 120B에서 파레토 프론티어이다

처리량-킬로와트 프론티어

InferenceX · GPT‑OSS‑120B · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB200 1,200 W

### 할라페뇨는 GPT‑OSS 운영 지점 전반에서 앞선다

최대 및 일치 처리량

InferenceX · GPT‑OSS‑120B · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB200 1,200 W

더 높은 최대 혼합 TPS / kW

≈1.9×

85,448 vs. 44,960 혼합 / kW

더 낮은 종단 간 지연

≈1.7×

1.03 s vs. 1.80 s

더 낮은 최소 TBT

≈2.7×

0.69 vs. 1.87 ms (1,459 vs. 535 tok/s/user)

이전 TBT에서 더 많은 처리량

≈53.7×

22,935 vs. 427 혼합 / kW (535.28 tok/s/user에서)

### 할라페뇨는 DeepSeek R1 670B에서 파레토 프론티어이다

처리량-킬로와트 프론티어

InferenceX · DeepSeek R1 MXFP4 · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB300 1,400 W

### 할라페뇨는 DeepSeek R1 운영 지점 전반에서 앞선다

최대 및 일치 처리량

InferenceX · DeepSeek R1 MXFP4 · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB300 1,400 W

더 높은 최대 혼합 TPS / kW

≈1.7×

19,641 vs. 11,781 혼합 / kW

더 낮은 종단 간 지연

≈3.6×

1.65 s vs. 5.99 s

더 낮은 최소 TBT

≈4.1×

1.43 vs. 5.90 ms (700 vs. 169 tok/s/user)

이전 TBT에서 더 많은 처리량

≈104.3×

12,258 vs. 118 혼합 / kW (169.41 tok/s/user에서)

### 할라페뇨는 Kimi K2.5 1T에서 파레토 프론티어이다

처리량-킬로와트 프론티어

InferenceX · Kimi K2.5 MXFP4 · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB300 1,400 W

### 할라페뇨는 Kimi K2.5 운영 지점 전반에서 앞선다

최대 및 일치 처리량

InferenceX · Kimi K2.5 MXFP4 · 명목 8k/1k · STP · 패키지 TDP: 할라페뇨 700 W; GB300 1,400 W

더 높은 최대 혼합 TPS / kW

≈1.5×

18,195 vs. 11,862 혼합 / kW

더 낮은 종단 간 지연

≈3.4×

1.56 s vs. 5.31 s

더 낮은 최소 TBT

≈3.8×

1.44 vs. 5.48 ms (694 vs. 182 tok/s/user)

이전 TBT에서 더 많은 처리량

≈56.1×

6,744 vs. 120 혼합 / kW (182.46 tok/s/user에서)
브리프용 요약 초안
OpenAI가 자체 추론 칩 할라페뇨의 첫 성능 결과를 공개하며 와트당 1.5~1.9배 높은 처리량과 1.7~3.6배 낮은 지연을 주장했다. 할라페뇨는 KV 캐시를 로컬에 유지하고 데이터 이동을 최소화하는 아키텍처를 채택했으나, 사용 메모리 종류와 용량은 명시되지 않아 메모리 산업 영향은 미확인이다. OpenAI는 NVIDIA 가속기를 계속 배치할 계획이며, 할라페뇨는 올해 말 배치 예정이다.

원문 텍스트

원문 열기 ↗
Since announcing Jalapeño, OpenAI’s first custom inference chip, we have been testing the chip and the system built around it. The results show a significant performance advance: Jalapeño can serve more AI work per unit of power while also returning responses more quickly. Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.

For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows. Our mission is to ensure that artificial general intelligence benefits all of humanity. These gains will help make increasingly capable AI more affordable and more broadly available.

OpenAI models also accelerated Jalapeño’s development. Earlier generations helped the team design and bring up the chip, while our latest models are accelerating how we optimize and program it. Jalapeño’s performance extends across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, showing that the architecture works across models developed both inside and outside OpenAI. Across all three, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance.

Jalapeño is also evidence of a broader full-stack advantage. OpenAI can design models, products, serving software, chips, memory, networking, and systems together, using what we learn from real workloads to improve every layer of the stack. Jalapeño is working first-party silicon with measured results, and it is the beginning of a multigenerational platform. In the months ahead, we will ramp Jalapeño to deliver faster, more capable, and more efficient products for our customers.

## How we measured Jalapeño’s performance

We evaluate performance at a matched user experience, measuring how much useful AI work each system can complete per unit of power while meeting the latency customers and interactive agents require. This matters especially for agents, which need to complete many steps in sequence, so delays can compound across an entire task.

To understand how Jalapeño performs in practice, we tested it on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. We compared Jalapeño with leading commercially available AI systems across the tested operating range, from high-throughput serving to highly interactive, low-latency use. Jalapeño delivered a better combination of throughput, power efficiency, and latency. Although performance is sometimes reported per chip, we believe the more useful standard is performance per unit of power.

### Jalapeño widens the lead at previous-best TBT

### Jalapeño delivers more tokens per user

### Jalapeño delivers more throughput per kilowatt

Across all three public models, Jalapeño delivered a better combination of performance per watt and latency across the tested operating range, placing it on the Pareto frontier. To compare the systems consistently, we normalized the results using each accelerator’s published chip power rating. Jalapeño is rated at 700 watts, although its measured sustained power remained at or below 550 watts on the workloads tested.

Jalapeño performed strongly across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. On Kimi, the largest public model we tested, it delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. In our internal testing, Jalapeño’s advantage widened further on frontier OpenAI models, suggesting that the architecture becomes more valuable as workloads grow larger and more demanding.

## Architecting for speed and efficiency within a single chip

Jalapeño was designed from the start by asking: what hardware would we build if its primary job were serving modern and future language models, especially interactive agents? Jalapeño’s gains come from designing the chip, memory, network, software, and rack-scale system together around real language-model workloads. Language-model inference moves through several distinct phases with different bottlenecks. Prefill, when the system processes a prompt, is compute-intensive, while decode, when the system generates the response token by token, is constrained more by memory bandwidth. Communication can also add latency when data must move between cores and chips, leaving some processing units idle while they wait. A system that excels at one phase can lose that advantage while waiting for data or moving model state between different resources.

We designed Jalapeño to minimize data movement and communication delays. This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase. The network is integral to the architecture. Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end. The result is a balanced and fungible accelerator that can support changing model architectures, excel at both prefill and decode, and adapt as the balance between them changes, a defining feature of agentic workloads.

## We used AI to design the chip, and designed the chip so AI could program it

AI played a direct role in Jalapeño’s development, enabling the team to move from initial design to tapeout in nine months by exploring implementations, shortening design, measurement, and verification loops, and continuously iterating on model workloads. AI also helped optimize the chip’s arithmetic circuits, allowing the team to fit more compute performance into the chip on schedule.

Jalapeño was designed as a clear, predictable programming target for both humans and AI. Engineers can describe work through local tensors, explicit communication, and predictable synchronization. AI can then optimize how that work is mapped, placed, scheduled, and coordinated across the system. That clear, predictable structure gives AI a tractable way to tackle the traditionally difficult problem of parallel programming.

Supporting each new model family still requires new kernels and model-specific optimizations. Using Codex with GPT‑Astra, the team brought three open-weight models that were not part of Jalapeño’s original production plan to high performance within two months. This demonstrated both the flexibility of the architecture and the speed at which AI can help us program it. For selected GPT‑OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than the existing human-expert-written implementations. Those figures apply to the selected blocks, not the full model, but they point toward a powerful new development loop.

## The path ahead for efficient, ultra-fast inference

AI infrastructure is valuable because of the useful real-world work it enables. By producing more useful work from the same power and hardware, Jalapeño can help us serve more demand and lower the cost of delivering a successful result. For OpenAI, that can improve operating leverage by allowing useful work and revenue to grow faster than the cost to serve. It can also support broader adoption and continued investment in better models, products, and infrastructure. Faster inference can enable faster iteration and new use cases.

Jalapeño expands what is possible for efficient, low-latency inference:

* Ultra-fast-mode inference at efficiencies previously available only in fast mode

* Fast-mode inference at efficiencies previously available only in batched mode

* Higher efficiency for batched-mode inference

We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what we learn and further advance both efficiency and speed.

Meeting growing demand for AI will require more compute from every available source. We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads. Our mission is to ensure that artificial general intelligence benefits all of humanity.

As we prepare for deployment, we are continuing production qualification, maturing the software, preparing to operate Jalapeño at scale, and validating performance across more models. The results so far show what is possible when we design the full system together: more responsive, capable, and agentic AI delivered more efficiently to more people.

## Appendix

### Jalapeño is pareto frontier at GPT‑OSS 120B

THROUGHPUT-per-kW FRONTIER

InferenceX · GPT‑OSS‑120B · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB200 1,200 W

### Jalapeño leads across GPT‑OSS operating points

PEAK AND MATCHED THROUGHPUT

InferenceX · GPT‑OSS‑120B · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB200 1,200 W

Higher peak mixed TPS / kW

≈1.9×

85,448 vs. 44,960 mixed / kW

Lower end-to-end latency

≈1.7×

1.03 s vs. 1.80 s

Lower min TBT

≈2.7×

0.69 vs. 1.87 ms (1,459 vs. 535 tok/s/user)

More throughput at previous TBT

≈53.7×

22,935 vs. 427 mixed / kW (at 535.28 tok/s/user)

### Jalapeño is pareto frontier at DeepSeek R1 670B

THROUGHPUT-per-kW FRONTIER

InferenceX · DeepSeek R1 MXFP4 · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB300 1,400 W

### Jalapeño leads across DeepSeek R1 operating points

PEAK AND MATCHED THROUGHPUT

InferenceX · DeepSeek R1 MXFP4 · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB300 1,400 W

Higher peak mixed TPS / kW

≈1.7×

19,641 vs. 11,781 mixed / kW

Lower end-to-end latency

≈3.6×

1.65 s vs. 5.99 s

Lower min TBT

≈4.1×

1.43 vs. 5.90 ms (700 vs. 169 tok/s/user)

More throughput at previous TBT

≈104.3×

12,258 vs. 118 mixed / kW (at 169.41 tok/s/user)

### Jalapeño is pareto frontier at Kimi K2.5 1T

THROUGHPUT-per-kW FRONTIER

InferenceX · Kimi K2.5 MXFP4 · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB300 1,400 W

### Jalapeño leads across Kimi K2.5 operating points

PEAK AND MATCHED THROUGHPUT

InferenceX · Kimi K2.5 MXFP4 · nominal 8k/1k · STP · package TDP: Jalapeño 700 W; GB300 1,400 W

Higher peak mixed TPS / kW

≈1.5×

18,195 vs. 11,862 mixed / kW

Lower end-to-end latency

≈3.4×

1.56 s vs. 5.31 s

Lower min TBT

≈3.8×

1.44 vs. 5.48 ms (694 vs. 182 tok/s/user)

More throughput at previous TBT

≈56.1×

6,744 vs. 120 mixed / kW (at 182.46 tok/s/user)