MEMORY INDUSTRY INTELLIGENCE

[AI 인프라 인사이트] 더 빠른 GPU만으로는 AI 성능을 구현할 수 없는 이유 | SK하이닉스 뉴스룸

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.11 19:12사용자 검토 전 초안

원문 제목: [AI Infrastructure Insight] Why faster GPUs alone can’t deliver AI performance | SK hynix Newsroom

핵심 요약

SK하이닉스 뉴스룸은 AI 성능이 컴퓨트 속도만으로 결정되지 않으며, 메모리·스토리지·네트워크·전력·냉각이 함께 최적화되어야 한다고 주장한다. UC Berkeley, ICSI, LBNL 연구진의 'AI and Memory Wall' 논문을 인용해 지난 20년간 서버 하드웨어의 최대 FLOPS는 약 2년마다 3배 증가한 반면 DRAM 대역폭은 1.6배, 인터커넥트 대역폭은 1.4배만 증가했다고 밝혔다. 추론 서비스 확대와 대규모 학습의 체크포인팅 과정에서 스토리지와 네트워크가 병목이 될 수 있으며, MLCommons의 MLPerf Storage 벤치마크가 이를 별도로 측정한다고 설명한다. 메모리 기업의 역할이 제품 공급을 넘어 고객 시스템의 데이터 흐름 공동 설계와 시스템 수준 메모리 아키텍처 결정으로 확대된다고 전망한다. 이 글은 특정 제품·고객·거래·수치를 새로 공시하지 않으며, 산업 구조 변화에 대한 해석과 전망을 제시한다.

메모리 산업 영향 분석

이 원문은 SK하이닉스 뉴스룸의 산업 해석·전망 글로, 특정 메모리 제품의 공급 배분·고객 인증·제품 믹스·투자 일정 변경을 직접 공시하지 않는다. 원문이 제시하는 변화는 AI 시스템 성능이 컴퓨트 단독이 아니라 메모리·스토리지·네트워크·전력·냉각의 데이터 경로 최적화에 달려 있다는 시스템 수준 논지이며, 이는 메모리 기업이 제품 공급을 넘어 고객 시스템의 데이터 흐름 공동 설계와 시스템 수준 메모리 아키텍처 결정 역량을 갖춰야 한다는 역할 확대 전망으로 이어진다. 메모리 산업에 대한 직접적 함의는 HBM·서버 DDR·서버 LPDDR·eSSD 등 데이터 경로상 계층별 제품의 상대적 중요성이 시스템 설계 선택에 따라 달라질 수 있다는 점이나, 원문은 특정 제품·고객·용량·세대·인증·거래를 명시하지 않으므로 제품별 수요 증감을 단정할 근거는 부족하다. 특히 원문이 인용한 'AI and Memory Wall' 논문의 FLOPS·DRAM 대역폭·인터커넥트 대역폭 증가율은 연구 결과(research)로 분류되며, SK하이닉스의 제품 로드맵이나 고객 채택을 확인하는 근거가 아니다. 원문의 ChatGPT 주간 메시지 180억 건, Gemini 월간 사용자 10억 명 수치는 각각 OpenAI와 Google이 공개한 사용량 지표로, 메모리 비트 수요나 제품 믹스로 직접 환산할 수 없다. 따라서 이 원문만으로는 공급 배분·고객 인증·제품 믹스·투자 일정 중 어떤 선택이 바뀌는지 확정할 수 없으며, 메모리 산업과의 직접 연결 근거가 부족하다. 다만 원문이 강조하는 데이터 이동 최소화·전력 효율·시스템 공동 설계 방향은 향후 메모리 계층 구조와 모듈 형태 선택에 영향을 줄 수 있는 가설적 조건으로, 실제 제품 채택 여부는 별도 공시·인증 자료로 확인해야 한다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 13475자

AI는 더 이상 단일 모델이나 칩으로 정의되지 않는다. AI가 실제 서비스와 산업 응용 분야에서 효과적으로 작동하려면 더 빠른 컴퓨트, 더 넓은 메모리 대역폭, 더 높은 성능의 네트워크, 더 효율적인 스토리지, 그리고 더 안정적인 전력과 냉각이 매끄럽게 함께 작동해야 한다. 이것이 AI 경쟁이 개별 기술을 넘어 전체 인프라의 설계와 운영으로 이동하고 있는 이유다.

**AI Infrastructure Insight** 시리즈에서는 컴퓨트, 메모리, 스토리지, 네트워크, 전력 및 냉각에서부터 이들이 하나의 통합 시스템으로 결합하는 방식에 이르기까지 AI 인프라 뒤의 아키텍처가 어떻게 진화하고 있는지, 그리고 이것이 업계에 무엇을 의미하는지 탐구한다.

**[시리즈 개요]**
① AI 데이터 센터 내부에서 무엇이 변하고 있는가?
**② 더 빠른 GPU만으로는 AI 성능을 구현할 수 없는 이유**
③ 전력과 냉각이 AI 데이터 센터의 다음 과제가 된 이유
④ 미래를 위해 AI 인프라가 어떻게 설계될 것인가

더 빠른 GPU*를 사용하면 AI 연산이 더 빨라지는가? GPU와 AI 가속기*는 대규모 모델의 학습과 복잡한 추론 요청 처리를 가능하게 한 핵심 기술이다. 그러나 실제 서비스 환경에서 AI 성능은 컴퓨트 속도만으로 결정되지 않는다.

* Graphics processing unit(GPU): 원래 그래픽 연산을 위해 개발된 병렬 처리 장치. 대량의 데이터를 동시에 처리하는 능력 덕분에 AI 학습과 추론에 널리 사용되게 되었다
* AI accelerator: AI 학습과 추론에 필요한 대규모 연산을 신속하게 수행하도록 설계된 반도체 또는 컴퓨팅 장치. 예로는 GPU, NPU, TPU가 있다

프로세서가 계산을 수행할 준비가 되어 있더라도 데이터가 제때 공급되지 않으면 값비싼 가속기가 대기 상태에 놓여 전체 시스템 효율이 떨어질 수 있다.
따라서 AI 시스템은 가속기 컴퓨트 능력, 데이터 전달 속도, 데이터 이동 경로가 효과적으로 함께 작동할 때에만 최대 성능을 발휘할 수 있다.

연구는 이 점을 뒷받침한다.
UC Berkeley, ICSI, LBNL의 연구진이 작성한 논문 “AI and Memory Wall”*은 지난 20년 동안 서버 하드웨어의 최대 FLOPS*가 약 2년마다 3배 증가한 반면, DRAM 대역폭*은 1.6배, 인터커넥트* 대역폭은 1.4배만 증가했다고 밝혔다.
컴퓨트 성능은 빠르게 발전한 반면, 데이터를 공급하고 이동하는 능력은 상대적으로 느리게 발전했다. 이 격차가 커질수록 컴퓨트 성능뿐 아니라 메모리와 데이터 이동 경로의 병목*이 AI 시스템에서 더 두드러질 수 있다.

* AI and Memory Wall: UC Berkeley, ICSI, LBNL의 연구진이 AI 시스템에서 컴퓨트 성능의 성장 속도와 메모리 및 인터커넥트 대역폭의 성장 사이의 격차를 분석한 연구
* Peak floating point operations per second(FLOPS): 컴퓨팅 시스템의 이론적 최대 부동소수점 컴퓨트 성능. FLOPS는 초당 수행할 수 있는 부동소수점 연산 수를 의미한다
* Bandwidth: 주어진 시간 동안 전송할 수 있는 데이터 양의 척도. 일반적으로 메모리와 네트워크의 데이터 전송 성능을 설명하는 데 사용된다
* Interconnect: 서버, 칩, 메모리, 가속기 등의 구성 요소를 연결하여 이들 사이에서 데이터가 이동할 수 있게 하는 아키텍처 또는 기술
* Bottleneck: 전체 시스템 성능을 제한하는 지점. AI 시스템에서 병목은 프로세서뿐 아니라 데이터 전달, 메모리, 스토리지, 네트워크에서도 발생할 수 있다



AI 시스템에서 데이터는 한 곳에 정체되어 있지 않다. 학습 데이터, 모델 파라미터*, 사용자 요청, 컨텍스트 정보, 검색 결과, 연산 결과는 처리 단계에 따라 서로 다른 위치 사이를 이동한다. 일부 데이터는 대용량 스토리지에서 검색되고, 다른 데이터는 가속기에 더 가까운 메모리에 도달하기 전에 시스템 메모리를 통과한다. 처리가 완료되면 결과는 다시 저장되거나 네트워크를 통해 다른 서버와 서비스로 전송된다.

* Model parameter: AI 모델이 학습하면서 조정되는 내부 값으로, 모델이 입력을 해석하고 출력을 생성하는 방식을 결정한다

> **☑︎ AI가 답하기 전에 데이터는 어디를 이동하는가?**
>
>
> 사용자가 챗봇이나 AI 검색 엔진에 질문을 입력하면, 데이터 센터는 사용자의 요청만 처리하는 것이 아니다. 이전 대화 컨텍스트와 검색 결과, 관련 문서, 모델이 필요로 할 수 있는 기타 정보가 필요에 따라 수집된다. 이 데이터는 가속기에 더 가까운 계층에 도달하기 전에 스토리지와 메모리를 통과한다. 연산이 완료되면 결과는 다시 저장되거나 네트워크를 통해 사용자에게 반환된다.
>
>
> AI 검색과 챗봇의 사용이 증가함에 따라 이러한 추론 요청은 더 자주, 더 큰 규모로 발생하고 있다. 궁극적으로 사용자가 경험하는 응답 속도는 모델 성능뿐 아니라 필요한 데이터를 얼마나 빠르게 수집하고 이동할 수 있는지에 달려 있다.
>
>
> 2025년 7월 기준으로 OpenAI에 따르면 ChatGPT에서 매주 180억 건의 메시지가 교환되었으며, Gemini는 2026년 8월 이후 월간 사용자 10억 명을 넘어섰다.
>
>
>
>
>
> ▲ Stargate Data Center 출처 : OpenAI, 「Stargate Advances with Partnership with Oracle」(2025)

이 흐름 안에서 각 계층은 서로 다른 역할을 한다. 스토리지는 대량의 정보를 보존하고, 시스템 메모리는 전체 서버를 위한 작업 공간을 제공한다. 가속기에 가까이 위치한 메모리는 프로세서가 필요로 하는 정보를 신속하게 공급하며, 네트워크는 여러 서버와 가속기를 하나의 시스템으로 연결한다.

따라서 AI 성능은 필요한 데이터가 존재하기만 하면 보장되지 않는다. 동일한 GPU를 사용하더라도 데이터가 얼마나 효율적으로 공급되고 이동되는지에 따라 시스템은 서로 다른 실제 성능을 낼 수 있다.

AI 시스템의 병목은 한 곳에서만 발생하지 않는다. 프로세서가 아무리 빠르더라도 메모리 대역폭이 충분하지 않으면 시스템이 필요할 때 충분한 데이터를 공급하지 못할 수 있다. 이 경우 가속기는 연산할 준비가 되어 있지만 여전히 데이터를 기다리는 시간을 보내게 되어 실제 활용률이 떨어진다.

스토리지도 병목* 지점이 될 수 있다. AI 서비스가 대량의 문서, 이미지, 로그, 사용자 이력, 검색 결과를 검색해야 할 때 필요한 데이터를 충분히 빠르게 읽을 수 없으면 응답 속도가 느려진다.
추론 서비스 요청 수가 증가함에 따라 이러한 지연은 서비스 품질을 직접 저하시킬 수 있다.

대규모 AI 학습 중에도 유사한 문제가 발생한다. AI 모델을 학습하려면 모델의 중간 학습 상태를 주기적으로 저장하는 체크포인팅*이 필요하다. 스토리지가 이 대량의 데이터를 충분히 빠르게 저장하거나 검색하지 못하면 전체 학습 클러스터에 지연이 발생할 수 있다. 이것이 MLCommons*가 MLPerf Storage* 벤치마크를 통해 AI 및 머신러닝 학습 워크로드의 스토리지 성능을 별도로 측정하는 이유다. 데이터 관리는 AI 학습과 추론의 핵심 요소이며, 스토리지 성능은 AI 인프라에서 점점 더 중요한 부분이 되고 있다.

* Checkpoint: AI 모델 학습의 현재 상태를 주기적으로 기록하여 중단이나 실패 후 저장된 지점부터 학습을 재개할 수 있게 하는 프로세스 또는 저장된 데이터 세트
* MLCommons: AI 시스템의 성능과 효율을 측정하기 위한 산업 표준 벤치마크를 개발하는 글로벌 AI 엔지니어링 컨소시엄
* MLPerf Storage: AI 및 머신러닝 워크로드를 지원하는 스토리지 시스템의 성능을 평가하는 MLCommons 벤치마크 모음. 데이터 읽기와 전달, AI 학습 중 체크포인트 저장 및 복구가 시스템 효율에 어떻게 영향을 미치는지 측정한다

네트워크도 잠재적 병목 원인이다. 대규모 AI 학습이나 고성능 추론에서는 여러 서버와 가속기가 단일 작업을 분할 처리하기 위해 동시에 데이터를 교환한다. 서버 간 데이터 전송이 지연되거나 네트워크 대역폭이 불충분하면 시스템의 한 부분에서의 지연이 전체 처리 속도를 저하시킬 수 있다. 개별 장치가 빠른 것만으로는 충분하지 않다. 여러 장치가 데이터를 효율적으로 교환하고 하나의 통합 시스템으로 작동해야 한다.

전력과 냉각도 성능을 제약할 수 있다. 고성능 장비는 최고 성능을 유지하기 위해 안정적인 전력 공급과 효과적인 열 관리가 필요하다. 이 시리즈의 3부에서 이러한 문제를 더 자세히 다룰 것이지만, 핵심은 분명하다. 병목과 성능 제한은 메모리, 스토리지, 네트워크를 포함한 데이터 이동 경로에서부터 전력 및 냉각 인프라에 이르기까지 시스템 전반에서 발생할 수 있다.



AI 인프라의 핵심은 더 많은 가속기를 확보하는 것뿐 아니라 그 가속기들이 지속적으로 연산에 참여하도록 시스템을 설계하는 데 있다. 데이터가 너무 느리게 공급되면 가속기 활용률*이 떨어져 투자 대비 성능이 낮아진다. 이것이 고성능 하드웨어를 갖춘 시스템조차 최대 잠재 성능을 발휘하지 못할 수 있는 이유다.

* Accelerator utilization: GPU 또는 AI 가속기가 연산에 적극적으로 사용되는 시간의 비율. 데이터 전달 지연은 가속기를 대기하게 하고 활용률을 떨어뜨릴 수 있다

이를 방지하려면 데이터가 필요할 때 필요한 곳에 공급될 수 있도록 데이터 경로를 설계해야 한다. 자주 사용되는 데이터는 프로세서 가까이에 배치하고, 대량의 데이터는 스토리지와 네트워크를 통해 효율적으로 검색해야 한다. 이는 모든 정보를 가장 빠른 위치에 배치하는 것이 불가능하기 때문이다. 궁극적으로 데이터의 위치와 이동 경로는 속도, 용량, 비용, 전력 효율을 고려하여 세심하게 설계되어야 한다.

소프트웨어는 실제 운영을 위해 이 아키텍처를 조정한다. 어떤 작업을 먼저 처리하는지, 언제 데이터를 검색하는지, 그리고 이들이 여러 장치에 어떻게 분산되는지에 따라 가속기 활용률과 전체 성능이 달라진다. 빠른 하드웨어를 갖추고도 비효율적인 스케줄링과 배치는 새로운 병목을 만들 수 있다.

데이터 이동을 최소화하도록 설계하면 전력 효율도 개선할 수 있다. Intel은 시스템 내에서 데이터를 이동하는 과정이 에너지를 소비하지만 연산 자체에는 직접 기여하지 않는다고 설명한다. 이것이 데이터가 더 짧은 거리를 이동하고 필요한 곳에 더 가까이 배치되어야 하는 이유다.

이 아키텍처 안에서 메모리는 필요할 때 필요한 곳에 데이터를 공급하여 가속기가 지속적으로 연산에 참여하도록 돕는 핵심 계층 역할을 한다. 따라서 컴퓨트와 메모리가 얼마나 효율적으로 연결되는지가 실제 시스템 성능을 결정할 수 있다.



데이터 경로를 중심으로 시스템을 설계하는 이러한 접근 방식은 AI 인프라 산업의 경쟁 구도를 변화시키고 있다. 빠른 GPU, 높은 메모리 대역폭, 대용량 스토리지, 고속 네트워크는 개별적으로 모두 중요하지만, 고립되어 작동하면 충분하지 않다. AI 시스템은 데이터가 스토리지에서 프로세서로 원활하게 이동하고 결과가 필요한 곳에 전달될 때에만 실제 성능을 발휘할 수 있다.

결과적으로 전체 산업 생태계에 걸친 협업 방식도 변화하고 있다. 반도체, 서버, 네트워킹, 클라우드, 스토리지 기업이 각자 자신의 구성 요소만 독립적으로 최적화하면 전체 시스템 성능을 극대화하는 데 한계가 있다. AI 워크로드가 더 커지고 복잡해짐에 따라 전체 데이터 흐름에 기반한 협업이 중요해진다. 이는 고객이 어떤 모델을 사용하고, 어떤 데이터를 처리하며, 어느 수준의 응답 속도가 필요한지에 따라 필요한 인프라 아키텍처가 달라지기 때문이다.

메모리 기업의 역할도 확대되고 있다. 단순히 제품을 공급하는 것을 넘어, 고객 시스템 내의 데이터 흐름을 공동 설계하고 프로세서, 스토리지, 네트워크를 포괄하는 시스템 수준 관점에서 어떤 메모리 아키텍처가 필요한지 결정하는 역량이 점점 더 요구된다.

컴퓨트는 여전히 필수적이다. 더 빠른 GPU와 AI 가속기는 계속해서 AI 인프라의 핵심 구성 요소가 될 것이다. 그러나 그들의 능력을 실제 성능으로 전환하려면 시스템 전반에 걸친 데이터 흐름 최적화도 필요하다.

앞으로 AI 인프라를 이해하려면 프로세서의 속도뿐 아니라 데이터가 시스템을 통해 이동하는 방식을 살펴야 한다. 메모리, 스토리지, 네트워크, 전력 및 냉각은 독립적인 구성 요소가 아니라 데이터가 중단 없이 흐르도록 하는 하나의 시스템의 유기적으로 상호 연결된 부분으로 간주되어야 한다.

더 빠른 GPU에 대한 필요성은 그 어느 때보다 필수적이다. 그러나 이제 질문은 한 단계 더 진화했다. 데이터가 GPU에 얼마나 빠르고 효율적으로 도달하여 처리되는가?

AI 성능을 결정하는 다음 병목은 그 질문에서 시작된다.

**<References****>**

* Amir Gholami et al., “[AI and Memory Wall](https://arxiv.org/abs/2403.14123),” IEEE Micro, Vol. 44, No. 3, 2024.
* OpenAI, “[How People Use ChatGPT](https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf),” 2025.
* Google, “[Google’s Gemini app hits 1 billion monthly active users](https://blog.google/innovation-and-ai/products/gemini-app/one-billion-monthly-users/),”
Google Blog, Aug. 11, 2026.
브리프용 요약 초안
SK하이닉스 뉴스룸은 AI 성능이 GPU 컴퓨트만으로 결정되지 않으며 메모리·스토리지·네트워크·전력·냉각의 데이터 경로 최적화가 필요하다고 주장했다. UC Berkeley 등의 'AI and Memory Wall' 연구를 인용해 컴퓨트와 메모리·인터커넥트 대역폭 성장 격차를 제시했으나, 특정 메모리 제품·고객·거래·수치는 새로 공시하지 않았다. 메모리 기업의 역할이 시스템 수준 데이터 흐름 공동 설계로 확대된다는 전망이 핵심이다.

원문 텍스트

원문 열기 ↗
AI is no longer defined by a single model or chip. For AI to operate effectively in real-world services and industrial applications, it takes faster compute, wider memory bandwidth, higher-performance networks, more efficient storage, and more stable power and cooling working seamlessly together. This is why competition in AI is shifting beyond individual technologies to the design and operation of the entire infrastructure.

In the**AI Infrastructure Insight** series, we explore how the architecture behind AI infrastructure is evolving — from compute, memory, storage, network, and power and cooling to the way they come together as a unified system — and what this means for the industry.

**[Series overview]**
① What is changing inside AI data centers?
**② Why faster GPUs alone can’t deliver AI performance**
③ Why power and cooling have become the next challenge for AI Data Centers
④ How AI infrastructure will be designed for the future

Does using a faster GPU* result in faster AI operations? GPUs and AI accelerators* are core technologies that have enabled the training of large-scale models and the processing of complex inference requests. In real-world service environments, however, AI performance is not determined by compute speed alone.

* Graphics processing unit(GPU): A parallel processing device originally developed for graphics computation. Its ability to process large volumes of data simultaneously has led to its widespread use in AI training and inference
* AI accelerator: A semiconductor or computing device designed to rapidly perform the large-scale computations required for AI training and inference. Examples include GPUs, NPUs, and TPUs

Even when a processor is ready to perform calculations, an expensive accelerator can be left waiting if data is not supplied in time, reducing overall system efficiency.
An AI system can therefore deliver its full performance only when accelerator compute capability, data delivery speed, and data movement paths work together effectively.

The research supports this point.
The paper “AI and Memory Wall,”* by researchers from UC Berkeley, ICSI, and LBNL, found that over the past 20 years, peak FLOPS* in server hardware increased by approximately threefold every two years, while DRAM bandwidth* increased by only 1.6 times and interconnect* bandwidth by 1.4 times.
While compute performance has advanced rapidly, the ability to supply and move data has progressed at a comparatively slower pace. As this gap widens, bottlenecks* in memory and data movement paths — not just compute performance — may become more pronounced in AI systems.

* AI and Memory Wall: A study by researchers from UC Berkeley, ICSI, and LBNL analyzing the gap between the rate of growth in compute performance and the growth of memory and interconnect bandwidth in AI systems
* Peak floating point operations per second(FLOPS): The theoretical maximum floating-point compute performance of a computing system. FLOPS refers to the number of floating-point operations that can be performed per second
* Bandwidth: A measure of the amount of data that can be transferred over a given period of time. It is commonly used to describe data transfer performance in memory and networks.
* Interconnect: An architecture or technology that connects components such as servers, chips, memory, and accelerators, enabling data to move between them
* Bottleneck: A point that limits overall system performance. In AI systems, bottlenecks can occur not only in processors but also in data delivery, memory, storage, and networks



In an AI system, data does not remain stagnant in one place. Training data, model parameters*, user requests, contextual information, search results, and compute results move between different locations depending on the stage of processing. Some data is retrieved from high-capacity storage, while other data passes through system memory before reaching memory closer to the accelerator. Once processing is complete, the results are stored again or transmitted over the network to other servers and services.

* Model parameter: An internal value adjusted as an AI model learns, determining how the model interprets inputs and generates outputs

> **☑︎ Where does data travel before AI answers?**
>
>
> When a user enters a question into a chatbot or AI search engine, the data center processes more than just the user’s request. Previous conversational context, along with search results, relevant documents, and other information the model may need, is gathered as required. This data moves through storage and memory before reaching layers closer to the accelerator. Once computation is complete, the result is stored again or returned to the user over the network.
>
>
> As the use of AI search and chatbots grows, these inference requests are occurring more frequently and at greater scale. Ultimately, the response speed experienced by users depends not only on model performance but also on how quickly the necessary data can be gathered and moved.
>
>
> As of July 2025, 18 billion messages were exchanged on ChatGPT each week, according to OpenAI, while Gemini surpassed 1 billion monthly users after August 2026.
>
>
>
>
>
> ▲ Stargate Data Center Source : OpenAI, 「Stargate Advances with Partnership with Oracle」(2025)

Within this flow, each layer plays a different role. Storage retains large volumes of information, while system memory provides a workspace for the entire server. Memory located close to the accelerator rapidly supplies the information needed by the processor, while networks connect multiple servers and accelerators into a single system.

AI performance, therefore, is not guaranteed simply by having the necessary data available. Even when using the same GPU, systems can deliver different real-world performance depending on how efficiently data is supplied and moved.

Bottlenecks in AI systems do not occur in a single location. No matter how fast a processor is, insufficient memory bandwidth can prevent the system from supplying enough data when it is needed. In this case, an accelerator may be ready to compute but still spend time waiting for data, reducing actual utilization rates.

Storage can also become a bottleneck* point. When an AI service needs to retrieve large volumes of documents, images, logs, user histories, or search results, response speeds slow down if the required data cannot be read quickly enough.
As the number of requests for inference services grows, such delays can directly degrade service quality.

Similar challenges arise during large-scale AI training. Training an AI model requires checkpointing*, which periodically saves the model’s intermediate training state. If storage fails to save or retrieve these large volumes of data quickly enough, latency can occur across the entire training cluster. This is why MLCommons* separately measures storage performance for AI and machine learning training workloads through its MLPerf Storage* benchmark. Data management is a core element of AI training and inference, making storage performance an increasingly important part of AI infrastructure.

* Checkpoint: A process or set of saved data that periodically records the current state of AI model training, allowing training to resume from the saved point following an interruption or failure
* MLCommons: A global AI engineering consortium that develops industry-standard benchmarks for measuring the performance and efficiency of AI systems
* MLPerf Storage: A suite of MLCommons benchmarks that evaluates the performance of storage systems supporting AI and machine learning workloads. It measures how data reads and delivery, as well as checkpoint saving and recovery during AI training, affect system efficiency

Networks are another potential source of bottlenecks. In large-scale AI training or high-performance inference, multiple servers and accelerators exchange data simultaneously to divide and process a single task. If data transmission between servers are delayed or network bandwidth is insufficient, slowdowns in one part of the system can reduce overall processing speed. It is not enough for an individual device to be fast. Multiple devices need to exchange data efficiently and operate as a unified system.

Power and cooling can also constrain performance. High-performance equipment requires a stable power supply and effective thermal management to sustain peak performance. Part 3 of this series will explore these issues in greater detail, but the key point is clear: Bottlenecks and performance limitations can arise throughout the system — from data movement paths involving memory, storage, and networks to power and cooling infrastructure.



The key to AI infrastructure lies not only in securing more accelerators but also in designing systems that keep those accelerators continuously engaged in computation. When data is supplied too slowly, accelerator utilization* falls, reducing performance relative to the investment made. This is why even systems equipped with high-performance hardware may fail to perform at their full potential.

* Accelerator utilization: The proportion of time a GPU or AI accelerator is actively used for computation. Delays in data delivery can leave accelerators waiting and reduce utilization

To prevent this, data paths need to be designed so that data can be supplied where and when it is needed. Frequently used data should be placed close to processors, while large volumes of data should be retrieved efficiently through storage and networks. This is due to the fact that it is impossible to place all information in the fastest location. Ultimately, the location and movement paths of data must be meticulously designed with speed, capacity, cost, and power efficiency in mind.

Software coordinates this architecture for real-world operations. Accelerator utilization and overall performance vary depending on which tasks are processed first, when data is retrieved, and how they are both distributed across multiple devices. Even with fast hardware, inefficient scheduling and placement can create new bottlenecks.

Designing to minimize data movement can also improve power efficiency. Intel explains that while the process of moving data within a system consumes energy, it does not directly contribute to computation itself. This is why data should travel shorter distances and be placed closer to where it is needed.

Within this architecture, memory serves as a critical layer that supplies data where and when it is needed, helping accelerators remain continuously engaged in computation. How efficiently compute and memory are connected can therefore determine real-world system performance.



This approach of designing systems around data paths is bringing about changes in the competitive landscape of the AI infrastructure industry. While fast GPUs, high memory bandwidth, high-capacity storage, and high-speed networks are all important individually, they are insufficient if they operate in isolation. AI systems can deliver real-world performance only when data moves smoothly from storage to the processor, and results are delivered where they are needed.

Consequently, collaboration methods across the entire industry ecosystem are also shifting. There are limitations to maximizing overall system performance if semiconductor, server, networking, cloud, and storage companies each optimize only their own components independently. As AI workloads grow larger and more complex, collaboration based on the entire data flow becomes crucial. This is due to the fact that required infrastructure architecture varies depending on which models customers use, what data they process, and what level of response speed is required.

The role of memory companies is expanding as well. Beyond merely supplying products, they increasingly need the capabilities to jointly design data flows within customer systems and determine what memory architectures are required from a system-level perspective encompassing processors, storage, and networks.

Compute remains essential. Faster GPUs and AI accelerators will continue to be core components of AI infrastructure. However, translating their capabilities into real-world performance requires optimizing the flow of data across the system as well.

Looking ahead, understanding AI infrastructure requires examining not only the speed of processors but also the ways in which data moves through the system. Memory, storage, networks, and power and cooling should not be considered as independent components but as organically interconnected parts of a single system that keep data flowing without interruption.

The need for faster GPUs remains as essential as ever. But the question has now evolved one step further: How quickly and efficiently does the data reach the GPU for processing?

The next bottleneck determining AI performance begins with that question.

**<References****>**

* Amir Gholami et al., “[AI and Memory Wall](https://arxiv.org/abs/2403.14123),” IEEE Micro, Vol. 44, No. 3, 2024.
* OpenAI, “[How People Use ChatGPT](https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf),” 2025.
* Google, “[Google’s Gemini app hits 1 billion monthly active users](https://blog.google/innovation-and-ai/products/gemini-app/one-billion-monthly-users/),”
Google Blog, Aug. 11, 2026.