MEMORY INDUSTRY INTELLIGENCE

OpenAI와 Broadcom, LLM에 최적화된 추론 칩 공개

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.10 11:44사용자 검토 전 초안

원문 제목: OpenAI and Broadcom unveil LLM-optimized inference chip

핵심 요약

OpenAI와 Broadcom이 OpenAI의 첫 Intelligence Processor인 Jalapeño를 공개했다. 이 칩은 LLM 추론을 위해 처음부터 설계된 가속기로, 9개월 만에 설계부터 테이프아웃까지 완료되었으며, 초기 테스트에서 현재 최첨단 제품보다 현저히 우수한 전력당 성능을 보일 것으로 나타났다. Broadcom의 실리콘 구현과 네트워킹 기술, Celestica의 보드·랙·시스템 전문성을 결합해 2026년 말까지 초기 배포를 목표로 하며, Microsoft 등 파트너와 함께 기가와트 규모 데이터센터에 다세대에 걸쳐 배포될 예정이다. OpenAI는 칩 아키텍처, 커널, 메모리 시스템, 네트워킹, 스케줄링, 배포 시스템, 제품 경험을 아우르는 풀스택 전략을 강화하고 있다. 최종 성능은 아직 측정 중이며, 상세 기술 보고서는 향후 몇 달 내 발표될 예정이다.

메모리 산업 영향 분석

OpenAI와 Broadcom이 공개한 Jalapeño는 LLM 추론에 특화된 맞춤형 ASIC으로, 2026년 말 초기 배포를 목표로 한다. 이 칩은 데이터 이동을 줄이고 컴퓨트·메모리·네트워킹 자원의 균형을 맞추는 아키텍처를 채택했으며, Broadcom의 Tomahawk 네트워킹 실리콘과 결합된다. 메모리 산업 관점에서 이 칩의 메모리 계층 구성(HBM, 서버 DRAM, LPDDR 등)은 원문에 명시되지 않아 직접적인 영향은 미확인이다. 다만 추론 효율 개선이 단위 메모리 사용량을 줄일 가능성과, 동시에 AI 사용량 증가로 전체 메모리 수요가 늘어날 가능성이 공존한다. OpenAI의 풀스택 전략은 자체 칩을 통한 컴퓨트 조달 다변화를 의미하며, 이는 NVIDIA GPU 의존도 변화와 HBM 수요 구조에 간접적 영향을 줄 수 있으나, 구체적 메모리 사양·채택 규모는 확인되지 않았다. Microsoft가 배포 파트너로 언급되었으나, 이 칩이 Microsoft 데이터센터에 실제로 채택되는지 여부는 원문에서 확정되지 않았다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 7194자

* 초기 테스트에 따르면 1세대 가속기는 현재 최첨단 제품보다 전력당 성능이 현저히 우수할 것입니다.

* 업계 전반의 현재 및 미래 LLM을 위해 처음부터 구축되었습니다.

* OpenAI의 모델로 가속화되어 설계부터 생산까지 9개월 만에 개발되었습니다.

* 제품에서 모델, 이제 칩에 이르기까지 OpenAI의 풀스택 플랫폼을 확장합니다.

* 데이터센터 파트너와 함께 여러 세대에 걸쳐 기가와트 규모로 배포될 예정입니다.

OpenAI와 Broadcom(NASDAQ: AVGO)은 오늘 OpenAI의 첫 Intelligence Processor인 Jalapeño를 공개했습니다. 이는 OpenAI의 미래 LLM 추론 비전을 중심으로 설계된 가속기이며, 두 회사가 함께 구축 중인 다세대 컴퓨트 플랫폼의 첫 AI 가속기로, 고급 AI를 더 빠르고, 더 안정적이며, 더 많은 사람이 접근할 수 있도록 만들기 위한 것입니다.



Jalapeño는 Broadcom의 회장 겸 CEO Hock Tan과 사장 Charlie Kawwas에 의해 OpenAI CEO Sam Altman과 사장 Greg Brockman에게 전달되었으며, 이는 OpenAI가 모델과 제품 뒤의 풀스택을 구축하려는 전략에서 중요한 진전입니다.

OpenAI는 LLM의 기본 원리에 대한 깊은 이해를 바탕으로 모델, 커널, 서빙 시스템, 제품 요구사항 로드맵에 따라 칩을 처음부터 설계했으며, 파트너인 Broadcom과 Celestica가 칩 구현, 보드, 랙 시스템 통합, 고성능 네트워킹, 확장 가능한 생산 시스템을 통해 플랫폼을 산업화하는 데 도움을 주었습니다. Jalapeño는 업계 전반의 현재 및 미래 AI 모델의 추론 요구에 대한 OpenAI의 통찰을 바탕으로 모든 LLM과 호환될 수 있는 유연성을 갖도록 설계되었습니다. Jalapeño 칩의 엔지니어링 샘플은 생산 목표 주파수와 전력으로 연구소에서 ML 워크로드를 실행 중이며, 여기에는 GPT‑5.3‑Codex‑Spark가 포함됩니다.

OpenAI는 아직 최종 성능을 측정 중이지만, 초기 테스트에 따르면 Jalapeño는 현재 최첨단 제품보다 전력당 성능이 현저히 우수할 것입니다. 성능에 대한 상세 기술 보고서는 향후 몇 달 내 발표될 예정입니다. 이 아키텍처는 데이터 이동을 줄이고 컴퓨트, 메모리, 네트워킹 자원의 균형을 맞추어 이론적 최대 성능에 훨씬 더 가까운 실현 활용률을 달성합니다. Broadcom의 실리콘 구현 및 네트워킹 기술(예: Tomahawk 네트워킹 실리콘)은 플랫폼을 대규모 생산으로 이끄는 데 기여합니다.

“세계는 컴퓨트 기반 경제로 이동하고 있습니다.”라고 OpenAI의 사장 겸 공동 창립자인 Greg Brockman은 말했습니다. “Jalapeño는 컴퓨트를 더 풍부하게 만들어 AI를 더 빠르고, 더 안정적이며, 사람과 기업에게 더 저렴하게 하고, 더 중요한 문제를 해결하는 데 사용될 수 있도록 하는 장기 풀스택 인프라 전략의 일부입니다. 스택을 더 많이 직접 설계함으로써 우리는 더 큰 효율로 더 많은 지능을 제공하고 고급 AI를 더 넓은 접근으로 밀어붙일 수 있습니다.”

“Jalapeño는 OpenAI 연구원들과의 긴밀한 협업에서 얻은 상세한 통찰을 사용하여 LLM 추론을 위해 처음부터 설계되었습니다.”라고 OpenAI의 하드웨어 프로그램을 이끄는 Richard Ho는 말했습니다. “우리는 프론티어 AI 모델에 가장 중요한 커널, 메모리 이동, 네트워킹, 서빙 패턴을 중심으로 아키텍처를 최적화했습니다. 초기 테스트에 따르면 Jalapeño는 하드웨어의 이론적 한계에 가깝게 우리의 가장 중요한 워크로드를 효율적으로 실행할 것입니다.”

“OpenAI와의 협력은 향후 10년의 AI에 필요한 물리적 인프라를 확장하겠다는 근본적인 약속을 나타냅니다.”라고 Broadcom의 사장 겸 CEO Hock Tan은 말했습니다. “이것은 다세대 로드맵의 시작에 불과합니다. 업계를 선도하는 우리 실리콘을 OpenAI와 직접 공동 개발함으로써 우리는 2026년부터 Microsoft 및 기타 파트너와 함께 기가와트 규모 데이터센터의 배포를 가능하게 하고 있습니다.”

## LLM을 위한 최고의 추론 플랫폼으로 설계

Jalapeño는 이전 AI 워크로드에서 적응된 범용 가속기가 아니라 현대 LLM 추론을 위한 백지 설계입니다. 이는 OpenAI가 ChatGPT, Codex, API, 미래 에이전트 제품 전반에서 매일 실행하는 시스템에 기반하며, 동시에 업계 전반의 현재 및 미래 LLM을 위해 설계되었습니다. 목표는 오늘날 선도적인 AI 가속기의 성능과 처리량을 가장 빠른 특화 추론 시스템에 가까운 지연 시간과 결합하여 Jalapeño를 대규모 대화형 LLM 제품에 매우 적합하게 만드는 것입니다.

이것이 풀스택 이점입니다. OpenAI는 프론티어 모델을 개발하거나 그 위에 제품을 구축하는 것뿐만 아니라 그 아래의 인프라(칩 아키텍처, 커널, 메모리 시스템, 네트워킹, 스케줄링, 배포 시스템, 제품 경험)를 설계하고 있습니다. OpenAI가 스택 전반에 걸쳐 운영되기 때문에 각 계층은 동일한 목표, 즉 모델을 사용자에게 더 빠르고, 더 안정적이며, 더 저렴하게 만드는 것을 중심으로 최적화될 수 있습니다.

Jalapeño는 OpenAI의 진보 뒤에 있는 플라이휠을 강화합니다. 더 나은 인프라는 컴퓨트 효율을 높입니다. 더 큰 컴퓨트 효율은 더 나은 훈련과 서빙을 가능하게 하여 궁극적으로 더 유능한 AI 모델을 구동합니다. 더 나은 모델은 사람, 개발자, 기업에게 더 나은 제품이 됩니다. 더 나은 제품은 더 많은 사용, 더 많은 고객, 더 많은 수익을 이끌어 OpenAI가 차세대 인프라에 재투자할 수 있게 합니다. 시간이 지나면서 이 순환은 지능을 모든 사람에게 더 유능하고, 더 안정적이며, 더 저렴하게 만드는 데 기여합니다.

## OpenAI 모델로 가속화된 9개월 테이프아웃

Jalapeño는 초기 설계부터 제조 테이프아웃까지 단 9개월 만에 공동 개발되었으며, 이 맞춤형 AI 가속기 프로그램은 고성능 첨단 반도체에서 달성된 가장 빠른 ASIC 개발 주기라고 믿습니다. 이러한 속도는 OpenAI 엔지니어링 팀과의 깊은 소프트웨어-하드웨어 공동 개발, Broadcom의 실리콘 구현 전문성, 그리고 설계 및 최적화 과정의 일부를 가속화하기 위한 OpenAI 모델 사용을 반영합니다.

사용자에게 제공되는 동일한 모델이 미래 모델을 실행하는 데 사용되는 인프라를 개선하는 데 도움을 주고 있습니다. AI가 엔지니어가 더 나은 칩을 더 빠르게 설계하도록 도울 수 있다면, 업계 전반의 컴퓨트 비용을 낮추고 고급 AI에 대한 접근을 민주화하는 데 도움이 될 수 있습니다.

## 파트너와 함께 다세대 플랫폼 구축

Jalapeño는 2026년 말까지 초기 배포를 위해 설계된 다세대 컴퓨트 플랫폼의 첫 단계이며, 향후 몇 년간 확장될 예정입니다. 이 플랫폼은 OpenAI가 설계한 가속기와 Broadcom의 실리콘 구현, 네트워킹, 연결 기술, 그리고 Celestica의 보드, 랙, 시스템 전문성을 결합합니다.

## 고급 AI를 더 널리 사용 가능하게 만들기

이 작업의 요점은 간단합니다. 추론은 AI가 사람에게 도달하는 곳입니다. 비용, 속도, 안정성의 모든 개선은 더 빠른 ChatGPT 답변, 더 적은 대기로 더 많은 단계를 수행할 수 있는 Codex 작업, 구축 비용이 더 저렴한 API 제품, 또는 수요가 높을 때 더 신뢰할 수 있는 접근으로 나타날 수 있습니다.

AI 민주화는 고급 모델을 더 많은 사람이 매일 사용할 수 있을 만큼 사용 가능하고, 신뢰할 수 있으며, 저렴하게 만드는 것을 의미합니다. Jalapeño는 OpenAI가 더 많은 인프라를 학생, 개발자, 소규모 기업, 연구원, 기업, 그리고 배우고, 창조하고, 어려운 문제를 해결하려는 모든 사람을 위한 유용한 지능으로 전환하는 데 도움을 줍니다.
브리프용 요약 초안
OpenAI와 Broadcom이 LLM 추론 특화 ASIC 'Jalapeño'를 공개했다. 9개월 만에 테이프아웃되었고 2026년 말부터 기가와트 규모 배포를 목표로 한다. 메모리 사양은 공개되지 않아 HBM·서버 DRAM 수요에 대한 직접적 영향은 아직 미확인이다.

원문 텍스트

원문 열기 ↗
* Early testing shows that the first-generation accelerator will deliver performance per watt substantially better than current state-of-the-art

* Built from the ground up for current and future LLMs across the industry

* Developed from design to production in nine months, accelerated by OpenAI’s models

* Expands OpenAI’s full-stack platform, from products to models and now to chips

* To be deployed at gigawatt scale with data center partners, over multiple generations

OpenAI and Broadcom (NASDAQ: AVGO) today unveiled Jalapeño, OpenAI’s first Intelligence Processor: an accelerator architected around OpenAI’s vision for the future of LLM inference, and the first AI accelerator in a multi-generation compute platform the companies are building together to make advanced AI faster, more reliable, and more accessible to more people.



Jalapeño was delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom President and CEO Hock Tan and President Charlie Kawwas, marking an important step in OpenAI’s strategy to build the full stack behind its models and products.

OpenAI designed the chip from scratch around its deep understanding of LLM fundamentals, informed by its roadmap of models, kernels, serving systems, and product needs, with partners Broadcom and Celestica, helping industrialize the platform through chip implementation, board, rack system integration, high-performance networking, and scalable production systems. Jalapeño is designed with flexibility to work with all LLMs guided by OpenAI’s insights into the inference needs of current and future AI models across the industry. Engineering samples of the Jalapeño chip are running ML workloads in the lab at production target frequency and power, including GPT‑5.3‑Codex‑Spark.

While OpenAI is still measuring final performance, early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art. A detailed technical report on performance will be presented in the coming months. The architecture reduces data movement and balances compute, memory, and networking resources to achieve realized utilization much closer to theoretical peak performance. Broadcom’s silicon implementation and networking technologies, including Tomahawk networking silicon, help bring the platform to large-scale production.

“The world is moving to a compute-powered economy,” said Greg Brockman, President and Co-Founder of OpenAI. “Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant, resulting in AI which is faster, more reliable, more affordable for people and businesses, and can be used to solve more important problems. By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access.”

“Jalapeño was designed from the ground up for LLM inference using detailed insights from our close collaboration with OpenAI researchers,” said Richard Ho, who leads OpenAI’s hardware program. “We optimized the architecture around the kernels, memory movement, networking, and serving patterns that matter most for frontier AI models. Based on early testing, Jalapeño will efficiently execute our most important workloads close to the hardware’s theoretical limits.”

“Our collaboration with OpenAI represents a fundamental commitment to scaling the physical infrastructure required for the next decade of AI,” said Hock Tan, President and CEO, Broadcom. “This is just the beginning of a multi-generation roadmap. By co-developing our industry-leading silicon directly with OpenAI, we are enabling the deployment of gigawatt scale data centers with Microsoft and other partners beginning in 2026.”

## Designed to be the best inference platform for LLMs

Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads. It is informed by the systems OpenAI runs every day across ChatGPT, Codex, the API, and future agentic products, while also being designed for current and future LLMs across the industry. The goal is to combine the power and throughput of today’s leading AI accelerators with latency closer to the fastest specialized inference systems, making Jalapeño well suited for interactive LLM products at scale.

That is the full-stack advantage. OpenAI is not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience. Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable for users.

Jalapeño strengthens the flywheel behind OpenAI’s progress. Better infrastructure drives compute efficiency. Greater compute efficiency enables better training and serving, ultimately powering more capable AI models. Better models become better products for people, developers, and businesses. Better products drive more usage, more customers, and more revenue, which lets OpenAI reinvest in the next generation of infrastructure. Over time, that cycle helps make intelligence more capable, more reliable, and less expensive for everyone.

## Nine-month tape-out, accelerated by OpenAI models

Jalapeño was co-developed from initial design to manufacturing tape-out in just nine months, and the custom AI accelerator program represents what we believe to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. That speed reflects deep software-hardware co-development with OpenAI’s engineering teams, Broadcom’s silicon implementation expertise, and the use of OpenAI models to accelerate parts of the design and optimization process.

The same models served to users are helping improve the infrastructure used to run future models. If AI can help engineers design better chips faster, it can lower the cost of compute across the industry and help democratize access to advanced AI.

## Building a multi-generation platform with partners

Jalapeño is the first step in a multi-generation compute platform designed for initial deployment by the end of 2026 and expanding in the years ahead, combining OpenAI-designed accelerators with Broadcom silicon implementation, networking, and connectivity technologies; and Celestica’s board, rack, and system expertise.

## Making advanced AI more broadly available

The point of this work is simple: inference is where AI reaches people. Every improvement in cost, speed, and reliability can show up as a faster ChatGPT answer, a Codex task that can take more steps with less waiting, an API product that is cheaper to build, or more dependable access when demand is high.

Democratizing AI means making advanced models available, dependable, and affordable enough for more people to use every day. Jalapeño helps OpenAI turn more of its infrastructure into useful intelligence for students, developers, small businesses, researchers, enterprises, and anyone trying to learn, create, or solve hard problems.