MEMORY INDUSTRY INTELLIGENCE

[AI 생태계] 인프라 재설계: 아키텍처가 성능을 결정하는 이유 | SK하이닉스 뉴스룸

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.11 19:18사용자 검토 전 초안

원문 제목: [AI Ecosystem] Redesigning infrastructure: Why architecture determines performance | SK hynix Newsroom

핵심 요약

이 글은 AI 워크로드가 추론·에이전트 AI로 진화하면서 시스템 성능의 기준이 컴퓨팅 파워에서 데이터 저장 위치·이동 속도·처리 효율로 이동하고 있다고 주장한다. LLM 추론은 모델 가중치·컨텍스트·중간 결과를 반복적으로 접근하며, KV 캐시와 어텐션 연산으로 메모리 용량·대역폭 요구가 급증한다고 설명한다. DRAM 접근 에너지가 단순 연산보다 150~2,000배 클 수 있고, 대규모 ML 모델에서 시스템 에너지의 90% 이상이 메모리 접근·데이터 이동에 소비될 수 있다는 연구를 인용한다. 해법으로 HBM, CXL, NVLink/NVSwitch, 근접 메모리 가속, 메모리 중심 컴퓨팅, 분산 근접 메모리 가속 시스템(Tesseract 연구)을 제시한다. 다음 글에서 HBM·첨단 패키징·PIM을 다룰 예정이라고 밝힌다.

메모리 산업 영향 분석

이 원문은 SK하이닉스 뉴스룸에 게재된 게스트 기고(ETH 취리히 Onur Mutlu 교수)로, AI 인프라 병목이 컴퓨트에서 데이터 이동·메모리로 이동하고 있다는 아키텍처 논의다. 메모리 산업에 대한 직접적 함의는 HBM이 GPU/NPU/TPU 옆에서 대역폭을 공급하는 핵심 계층으로 명시되고, CXL이 메모리 확장·공유·풀링을 가능하게 하는 인터커넥트로, 근접 메모리 가속·PIM이 다음 글의 주제로 예고된다는 점이다. 다만 이는 기술·아키텍처 담론이며 특정 제품의 공급 배분, 고객 인증, 제품 믹스, 투자 일정에 대한 회사 차원의 결정이나 수치를 담고 있지 않다. 따라서 공급 배분·고객 인증·제품 믹스·투자 일정 중 어느 선택에 영향을 준다고 단정할 근거는 원문에 없고, '메모리 산업과의 직접 연결 근거 부족'으로 표시한다. 분석가 가설로는, 만약 메모리 중심·근접 메모리 가속 아키텍처가 실제 채택되면 HBM·고용량 DRAM의 계층 내 위치가 강화되고 CXL 기반 메모리 풀링이 서버 DDR 수요 구조를 바꿀 수 있다는 방향성이 있으나, 이는 원문이 제시한 개념·연구 수준이며 채택·양산·인증 근거는 없다. 반대 근거로는 원문 스스로 '최선의 접근은 모델 아키텍처·컨텍스트 길이·요청 수·네트워크 구성에 따라 달라진다'고 밝혀 단일 제품·규격으로의 수렴을 부정한다. 판단 변경 조건은 다음 글에서 HBM·첨단 패키징·PIM의 구체적 개발·인증·상용화 단계가 제시되는지, 그리고 CXL·NVLink 채택이 특정 고객·플랫폼에서 확인되는지다. research_topics의 HBM 세대·스택·인증, 서버 DDR·LPDDR·SOCAMM 규격, NAND·eSSD·HBF 개발 단계, SK하이닉스/Solidigm 신제품·고객 인증 질문은 이 원문만으로 답할 수 없어 미확인으로 남긴다.
한국어 번역 읽기

수집된 원문 v1의 분석에 제공된 본문 기준 · 15792자

AI 워크로드가 추론과 에이전트 AI로 진화함에 따라 시스템 성능의 기준도 변하고 있다. 컴퓨팅 파워만으로는 더 이상 충분하지 않다. 데이터가 어디에 저장되고, 얼마나 빠르게 이동하며, 얼마나 효율적으로 처리되는지가 이제 AI 시스템의 성능과 효율에 직접적인 영향을 미친다.

이 기사 시리즈는 소프트웨어, 데이터센터 인프라, 반도체, 메모리 기술이 함께 작동하는 AI 생태계의 변화를 탐구한다. 또한 메모리 반도체가 왜 AI 시대의 기반 기술이 되었는지 살펴보고, 차세대 AI 메모리 생태계를 구축하기 위한 SK하이닉스의 비전을 제시한다.

① AI 컴퓨팅의 패러다임 전환
② 진짜 병목: 컴퓨트가 아닌 데이터
**③ 인프라 재설계: 아키텍처가 성능을 결정하는 이유**
**– Onur Mutlu 교수, ETH 취리히**
④ 메모리로 이동하는 반도체 패러다임
⑤ AI의 완성: 피지컬 인텔리전스와 메모리의 역할

ChatGPT, Gemini, Claude와 같은 대규모 언어 모델(LLM) 서비스가 하나의 질문에 답하기 위해서는 모델이 방대한 양의 모델 가중치, 이전 컨텍스트, 중간 결과에 반복적으로 접근해야 한다. 응답이 길어지고 더 많은 추론 단계를 포함할수록 읽어야 하는 데이터의 양도 증가한다. 따라서 AI 시스템의 핵심 과제 중 하나는 컴퓨트 장치에 대규모 데이터를 얼마나 빠르고 안정적으로 공급할 수 있는가였다.

이 과제는 AI의 무게 중심이 모델 개발에서 대규모 서비스 운영으로 이동하면서 점점 더 중요해졌다. 사전 학습은 막대한 초기 투자가 필요하지만, 추론은 사용자가 질의를 제출할 때마다 반복적으로 발생한다. 생성형 AI 서비스, 엔터프라이즈 코파일럿, 에이전트 AI가 확산되면서 요청 수가 늘어나고, 각 요청이 유지해야 하는 컨텍스트도 길어진다. 그 결과 AI 인프라 경쟁은 단순히 더 빠른 GPU/NPU/TPU와 더 큰 클러스터를 배치하는 것을 넘어서고 있다. 초점은 GPU/NPU/TPU, 메모리, 스토리지, 네트워크를 통합 시스템으로 얼마나 효율적으로 공동 설계하고 연결하며, 그 시스템을 통해 데이터가 얼마나 효율적으로 흐를 수 있는가, 더 적절하게는 시스템이 구성 요소 간에 데이터를 얼마나 최소한으로 이동시키는가에 점점 더 맞춰지고 있다.

오늘날의 컴퓨팅 시스템은 AI·머신러닝, 유전체 분석, 대규모 데이터 분석과 같이 점점 더 중요한 워크로드를 가능하게 하는 매우 강력한 인프라를 제공한다. 그러나 그 근본 아키텍처는 여전히 프로세서 중심이다. 데이터는 처리되기 전후에 센서, 스토리지, 메모리, 캐시에서 컴퓨트 장치까지 그리고 다시 돌아오는 메모리 계층 전반을 끊임없이 이동해야 한다.

이 아키텍처는 범용 컴퓨팅의 성능을 향상시키는 데 오랫동안 효과적이었다. 그러나 AI 워크로드가 데이터 양과 접근 빈도 모두에서 극적으로 증가하면서, 데이터 접근과 데이터 이동이 시스템 효율의 주요 병목이 되었다. 모델은 방대한 양의 가중치와 중간 결과에 지속적으로 접근하며, 추론은 이전 컨텍스트와 새로 생성된 토큰을 반복적으로 읽는다. 데이터가 더 멀리, 더 많이 이동할수록 지연, 전력 소비, 하드웨어 비용이 커진다.

이러한 데이터 접근 및 이동 병목을 완화하기 위해 시스템 설계자들은 컴퓨터 설계를 크게 복잡하게 만들었다. 현대 컴퓨팅 시스템은 캐시 계층*, 필요하기 전에 데이터를 미리 가져오는 프리페칭, 여러 작업을 병렬로 실행하는 멀티스레딩, 비순차 실행*과 같은 복잡하고 값비싼 메커니즘을 많이 사용한다. 이러한 기술은 컴퓨트 장치의 지연을 줄이는 데 제한적인 효과를 보였지만, 시스템에 상당한 복잡성, 에너지 비효율, 하드웨어 비용을 초래했다.

* 캐시 계층: 최근 또는 자주 사용되는 데이터를 컴퓨트 장치 가까이에 임시 저장해 메모리 접근 시간을 줄이는 계층 구조. 일반적으로 L1, L2, L3 캐시와 같은 여러 수준으로 구성된다
* 비순차 실행: 프로그램에 명시된 순서와 다르더라도 실행 준비가 된 명령을 먼저 실행해 컴퓨트 장치 유휴 시간을 줄이는 기법

오늘날의 컴퓨팅 시스템에서 연산은 상대적으로 저렴하지만, 메모리 접근과 데이터 이동은 성능과 에너지 모두에서 비싸다. DRAM에서 데이터를 한 번 가져오는 데 필요한 에너지는 단순 산술 연산에 필요한 에너지보다 150~2,000배 클 수 있다. 텐서 처리 장치(TPU) 시스템에서 실행되는 신경망 모델에 관한 연구는 이러한 격차가 에너지와 성능에 미치는 막대한 영향을 입증했다*. 대규모 머신러닝 모델의 경우 시스템 에너지의 90% 이상이 연산이 아니라 메모리 접근과 데이터 이동에 소비될 수 있다*.

컴퓨트-메모리 격차는 대규모 언어 모델(LLM)에서 더욱 두드러진다. LLM이 토큰을 생성할 때마다 이전 컨텍스트를 참조하고 KV 캐시*와 어텐션 연산을 통해 방대한 중간 데이터를 처리한다. 모델이 더 오래 추론하고 더 긴 컨텍스트를 유지할수록 메모리 용량과 대역폭*에 대한 요구가 급격히 증가한다. 이러한 메모리 수요가 증가할수록 메모리 병목은 더 커진다. 컴퓨트 장치가 아무리 빨라져도 시스템이 데이터가 컴퓨트 유닛에 도착하기를 기다리는 데 대부분의 시간을 쓰기 때문에 시스템 효율은 떨어진다.

AI 워크로드가 확장됨에 따라 고대역폭 메모리(High Bandwidth Memory)는 GPU에 더 많은 양의 데이터를 더 높은 속도로 공급하는 데 점점 더 중요해진다. GPU/NPU/TPU 옆에 위치한 HBM은 초고속으로 대량의 데이터를 전달해 컴퓨트 장치 지연을 줄이는 데 도움을 준다. 그러나 HBM 성능을 완전히 활용하려면 GPU/NPU/TPU와 HBM뿐만 아니라 스토리지, 네트워크, 인터커넥트*를 포함한 전체 데이터 경로가 통합 시스템으로 설계되어야 한다.

* 키-값 캐시(KV 캐시): LLM이 이후 토큰을 생성할 때 사용하기 위해 이전 토큰의 어텐션 연산 결과를 저장하는 메모리 영역. 컨텍스트 길이가 증가하면 KV 캐시가 커져 메모리 용량과 대역폭 요구가 증가한다
* 메모리 대역폭: 주어진 기간 내에 메모리에서 전송할 수 있는 데이터의 양. 메모리 대역폭이 높을수록 컴퓨트 장치가 한 번에 더 많은 데이터를 받아 처리할 수 있다
* 인터커넥트: 프로세서, 메모리, 가속기, 스토리지와 같은 시스템 구성 요소를 연결하는 통신 경로. AI 인프라에서 인터커넥트는 데이터 이동의 속도와 효율을 결정하는 핵심 요소다

시스템 설계의 근본 가정도 바뀌어야 한다. 프로세서 중심 아키텍처는 대체로 데이터를 컴퓨트 장치로 이동시켜 처리하는 방식으로 작동한다. 데이터 양이 늘고 접근 빈도가 증가하면서 모든 데이터를 중앙 컴퓨트 장치로 보내는 아키텍처는 빠르게 한계에 도달한다. 이제 시스템은 데이터가 어디에 있는지와 처리되기까지의 경로를 중심으로 재설계되어야 한다.

이 과제를 해결하기 위해 개발된 개념 중 하나가 메모리 중심 컴퓨팅*이다. 메모리 병목을 근본적으로 더 잘 해결하려면 데이터 이동을 줄이고 메모리에 더 가까이에서 연산을 가능하게 해야 한다. 이것이 메모리 중심 컴퓨팅의 기본 아이디어다. 메모리 중심 컴퓨팅은 데이터 위치와 이동 비용을 시스템 설계의 핵심 고려 사항으로 삼아 메모리 배치와 컴퓨트 아키텍처를 함께 최적화한다*. 일부 데이터는 GPU/NPU/TPU 옆의 HBM에 상주할 수 있고, 다른 데이터는 공유 메모리 풀*에서 관리되며, 특정 연산은 메모리에 더 가까이에서 수행된다. 최선의 접근 방식은 모델 아키텍처, 컨텍스트 길이, 사용자 요청 수, 네트워크 구성과 같은 요인에 따라 달라진다.

좋은 소식은 이러한 메모리 중심 시스템이 이미 학계뿐 아니라 산업계에서도 설계되고 평가되고 있다는 점이다. 데이터 이동 병목을 제거하려는 움직임이 커지고 있으며, 이는 가능한 한 데이터를 이동하지 않는 방식으로 가장 잘 달성된다.

* 메모리 중심 컴퓨팅: 데이터를 컴퓨트 장치로 이동시키는 기존 구조에서 벗어나 메모리에 더 가까이에서 데이터를 처리하도록 설계된 컴퓨팅 패러다임. 데이터 이동을 줄여 성능과 전력 효율을 모두 개선하는 데 초점을 맞춘다
* Onur Mutlu, et al., “_[Memory-Centric Computing: Solving Computing’s Memory Problem](https://arxiv.org/abs/2505.00458)_,” 17th IEEE International Memory Workshop(IMW) (2025)
* 메모리 풀: 여러 장치가 접근할 수 있는 공유 메모리 자원. 각 장치가 필요에 따라 메모리 용량을 사용할 수 있게 함으로써 메모리 풀링은 전체 시스템 자원 활용도를 향상시킬 수 있다

프로세서, 메모리, 가속기는 오랫동안 다양한 인터페이스를 통해 연결되어 왔다. 그러나 AI 워크로드가 처리하는 데이터 양이 급격히 증가하면서 기존 연결 아키텍처는 요구되는 통신 대역폭과 메모리 접근 효율을 제공하는 데 어려움을 겪고 있다. 결과적으로 장치 간에 더 빠르고 유연한 데이터 경로를 가능하게 하는 고속 인터커넥트의 중요성이 커지고 있다. 오늘날 CXL*과 NVLink/NVSwitch*는 서로 다른 영역에서 이러한 전환을 이끄는 두 가지 대표 기술이며, 구성 요소를 긴밀하게 통합하려는 다양한 학술 연구도 존재한다.

* Compute Express Link(CXL): CPU, GPU/NPU/TPU, 메모리, 가속기 간에 고속 연결을 제공하는 차세대 인터커넥트 표준. CXL은 메모리 확장과 공유를 가능하게 하는 핵심 기술이다
* NVLink와 NVSwitch: NVIDIA가 개발한 고속 인터커넥트 기술. NVLink는 GPU 간 또는 GPU와 다른 장치 간에 고대역폭 연결을 제공하며, NVSwitch는 이 연결성을 확장해 더 많은 수의 GPU 간 고속 통신을 가능하게 한다. 이러한 기술은 대규모 AI 클러스터에서 GPU 간 데이터 병목을 줄이는 데 사용된다

CXL은 메모리 확장과 공유를 가능하게 하는 핵심 인터커넥트 기술이다. 전통적인 서버 아키텍처에서 메모리는 대체로 특정 CPU나 GPU/NPU/TPU에 부착된 전용 자원으로 기능했다. 한 장치에 여유 메모리 용량이 있고 다른 장치가 부족하더라도 그 용량을 유연하게 공유하기 어려웠다. CXL은 프로세서, 가속기, 메모리 장치 간 연결성을 확장해 유연한 메모리 확장, 공유, 풀링의 기반을 제공한다. CXL을 사용하면 메모리 근처에서 처리를 수행하는 장치나 컨트롤러로 연산을 오프로드할 수도 있으며, CXL 장치는 유연한 인터페이스를 통해 서로 통신할 수 있다.

NVLink와 NVSwitch는 GPU 클러스터 내에서 고대역폭 연결을 제공한다. 대규모 AI 모델이 여러 GPU와 노드에 걸쳐 실행될 때 모델 파라미터, 활성화, KV 캐시와 같은 중간 데이터는 GPU 간에 이동해야 한다. GPU 간 연결이 느리거나 유연하지 않으면 각 개별 GPU가 아무리 빨라도 전체 처리 성능이 저하될 수 있다

[원문 후반 발췌]

그리고 공유를 가능하게 하며, NVLink/NVSwitch는 고대역폭 GPU 연결을 제공한다. 그러나 둘 다 필요한 데이터를 필요한 곳에 빠르게 전달함으로써 AI 인프라의 데이터 병목을 줄인다는 더 넓은 목표를 지원한다. 이러한 연결 아키텍처가 구축되면 메모리와 가속기는 고정된 구성 요소가 아니라 워크로드 요구 사항에 맞춰 구성할 수 있는 시스템 자원이 된다.

고속 인터커넥트가 가능하게 하는 유연한 연결성은 컴퓨트 장치와 메모리가 배치되는 방식도 바꾸고 있다. 장치 간에 더 많은 데이터 경로가 사용 가능해지면서 컴퓨트와 메모리 자원은 워크로드 요구 사항에 따라 배치될 수 있다. 이러한 변화는 단순히 더 많은 데이터를 더 먼 거리로 이동시키는 것이 아니라, 데이터가 있는 곳에 더 가까이에서 필요한 처리를 수행하는 아키텍처를 향한다.

데이터 이동 병목을 줄이는 것은 단순히 연결 속도를 높이는 것 이상으로 나아갈 수 있다. CXL과 NVLink/NVSwitch가 서로 다른 방식으로 메모리와 GPU/NPU/TPU 간 데이터 경로를 개선하는 반면, 근접 메모리 가속은 연산 자체를 데이터에 더 가까이 배치함으로써 이동해야 하는 데이터의 양을 줄인다.

중앙 GPU/NPU/TPU가 모든 데이터를 가져와 처리하는 대신, 메모리 내부나 근처에 위치한 가속기들이

[원문 후반 발췌]

가속기는 고속 인터커넥트를 통해 긴밀하게 결합되어, 전체 데이터셋을 중앙 GPU/NPU/TPU로 보내는 대신 데이터를 먼저 메모리 가까이에서 처리하고 필요한 결과만 다른 장치와 교환할 수 있다. HBM과 같은 3D 적층 고대역폭 메모리* 또는 기타 고용량 메모리 근처에서 연산이 이루어지면 앞뒤로 이동해야 하는 데이터가 훨씬 줄어든다. 이는 지연과 전력 소비를 줄이면서 여러 노드로 확장하기도 더 쉽게 만들 수 있다. 이 아키텍처는 긴밀하게 연결된 근접 메모리 가속기의 분산 시스템*으로 알려져 있다.

이 접근 방식의 잠재력은 Tesseract*에서 볼 수 있다. Tesseract는 병렬 그래프 처리에 분산 근접 메모리 가속을 적용한 연구 아키텍처(2015년 제42회 ACM/IEEE 국제 컴퓨터 아키텍처 심포지엄에서 발표)다. 메모리 가까이에 위치한 여러 가속기가 할당된 데이터를 독립적으로 처리하고 필요할 때만 서로 통신한다. 다양한 병렬 그래프 처리 워크로드에서 이 연구는 기존 프로세서 중심 설계와 비교해 약 10배 수준의 성능 향상과 이에 상응하는 에너지 절감을 입증했다.

이러한 결과의 의의는 숫자 자체를 넘어선다. 처리를

[원문 후반 발췌]

기존 설계보다 확장성과 에너지 효율이 더 우수하다. 더 많고 다양한 가속기가 추가될수록 메모리 용량, 메모리 대역폭*, 연산 능력이 모두 비례적으로 확장된다. 즉, 핵심은 단순한 분산 처리가 아니라, 여러 근접 메모리 가속기가 긴밀하게 연결되어 최소한의 데이터 이동으로 협력 연산을 수행하는 아키텍처다.

* 3D 적층 고대역폭 메모리: 여러 메모리 다이를 수직으로 적층해 고대역폭을 제공하는 메모리 아키텍처. HBM이 대표적인 예
* 긴밀하게 연결된 근접 메모리 가속기의 분산 시스템: 메모리 가까이에 위치한 여러 가속기를 고속 인터커넥트로 연결해 각 가속기가 할당된 데이터를 처리하고 필요한 결과만 교환할 수 있게 하는 아키텍처
* Tesseract: 메모리 가까이에 위치한 여러 가속기가 고속 인터커넥트를 통해 긴밀하게 연결되는 근접 메모리 가속용 연구 아키텍처. 각 가속기는 할당된 데이터를 로컬에서 처리하고 필요할 때만 다른 가속기와 통신한다
* 원 아키텍처는 2015년 제42회 ACM/IEEE 국제 컴퓨터 아키텍처 심포지엄에서 “_[A Scalable Processing-in-Memory Accel

[원문 후반 발췌]

프로세서 중심 패러다임으로 인해 발생한다. 위에서 살펴본 고속 인터커넥트, 근접 메모리 가속기 아키텍처, 분리형(disaggregated) 아키텍처는 모두 데이터 이동의 큰 부담을 줄이기 위한 시스템 수준 접근 방식이다.

이제 질문은 반도체와 메모리 자체의 아키텍처로 향한다. 단순히 데이터를 컴퓨트 장치에 더 가까이 배치하는 것을 넘어, 일부 연산이 메모리 자체 내부나 주변에서 수행된다면 데이터 이동을 얼마나 더 줄일 수 있을까?

다음 기사에서는 HBM, 첨단 패키징, 메모리 내 연산(PIM)에 초점을 맞춰 AI 반도체의 무게 중심이 왜 메모리로 이동하고 있는지 탐구할 것이다.



_**면책 고지:** 이 기사에 표현된 의견은 전적으로 저자의 것이며 SK하이닉스의 공식 입장을 반드시 반영하지는 않습니다._

본문이 길어 24063자 중 15792자만 처리했습니다.

브리프용 요약 초안
SK하이닉스 뉴스룸 기고에서 ETH 취리히 Onur Mutlu 교수는 AI 병목이 컴퓨트에서 데이터 이동·메모리로 이동했으며, HBM·CXL·NVLink/NVSwitch·근접 메모리 가속·메모리 중심 컴퓨팅이 해법이라고 주장했다. DRAM 접근 에너지가 단순 연산의 150~2,000배, 대규모 ML 모델 시스템 에너지의 90% 이상이 메모리 접근·데이터 이동에 소비될 수 있다는 연구를 인용했다. 다음 글에서 HBM·첨단 패키징·PIM을 다룰 예정이라고 예고했으나, 구체적 제품·고객·수치·일정은 제시되지 않았다.

원문 텍스트

원문 열기 ↗
As AI workloads evolve toward inference and agentic AI, the criteria for system performance are also changing. Computing power alone is no longer enough. Where data is stored, how quickly it moves, and how efficiently it is processed now have a direct impact on the performance and efficiency of AI systems.

This article series explores the transformation of the AI ecosystem, where software, data center infrastructure, semiconductors, and memory technologies work together. It also examines why memory semiconductors have become a foundational technology for the AI era and outlines SK hynix’s vision for building the next-generation AI memory ecosystem.

① The paradigm shift in AI computing
② The real bottleneck: Data, not compute
**③ Redesigning infrastructure: Why architecture determines performance**
**– Professor Onur Mutlu, ETH Zurich**
④ The semiconductor paradigm shifts toward memory
⑤ Completing AI: Physical intelligence and the role of memory

For a large language model(LLM) service such as ChatGPT, Gemini, or Claude to answer a single question, the model must repeatedly access vast amounts of model weights, previous context, and intermediate results. As responses become longer and involve more reasoning steps, the amount of data that must be read also increases. Therefore, one of the key challenges for AI systems has been how quickly and reliably they can supply massive volumes of data to compute devices.

This challenge has become increasingly important as the center of gravity in AI shifts from model development to large-scale service operations. Pre-training requires a massive upfront investment, while inference occurs repeatedly every time a user submits a query. As generative AI services, enterprise copilots, and agentic AI proliferate, the number of requests grows, and the context that each request must maintain becomes longer. As a result, AI infrastructure competition is moving beyond simply deploying faster GPUs/NPUs/TPUs and larger clusters. The focus is increasingly on how efficiently GPUs/NPUs/TPUs, memory, storage, and networks can be co-designed and connected as an integrated system and how efficiently data can flow through that system. Or more appropriately, how minimally the system moves data across components.

Today’s computing systems provide a very powerful infrastructure that has enabled increasingly important workloads such as AI and machine learning, genomic analysis, and large-scale data analytics. Yet their fundamental architecture remains processor-centric. Data must continuously move across the memory hierarchy — from sensors, storage, memory, and caches all the way to the compute device(s) and back again — before and after being processed.

This architecture has long been effective at improving the performance of general-purpose computations. As AI workloads have grown dramatically in both data volume and access frequency, however, data access and data movement have become major bottlenecks in system efficiency. Models continuously access vast amounts of weights and intermediate results, while inference repeatedly reads previous context and newly generated tokens. The farther and more data must travel, the greater the latency, power consumption, and hardware cost.

To mitigate these data access and data movement bottlenecks, system designers have complicated the design of computers greatly. Modern computing systems employ many complex and expensive mechanisms, such as cache hierarchies*, prefetching data before it is needed, multithreading to execute multiple tasks in parallel, and out-of-order execution*. These technologies have had limited effectiveness at reducing the latency of compute devices, but they have come at significant complexity, energy inefficiency, and hardware cost in the system.

* Cache hierarchies: A hierarchical structure that temporarily stores recently or frequently used data close to compute devices to reduce memory access time. It typically consists of multiple levels, such as L1, L2, and L3 caches
* Out-of-order execution: A technique that reduces compute device idle time by executing instructions that are ready to execute ahead of others, even when this differs from the order specified in the program

In today’s computing systems, computation is relatively cheap, yet memory access and data movement are expensive — in both performance and energy. The energy required to retrieve data from DRAM once can be 150 to 2,000 times greater than that required for a simple arithmetic operation. Research involving neural network models running on Tensor Processing Unit(TPU) systems has demonstrated the huge energy and performance impact of this disparity*. For large machine learning models, more than 90% of system energy can be consumed by memory access and data movement rather than computation*.

The compute-memory disparity becomes even more pronounced with large language models(LLMs). Each time an LLM generates a token, it refers to the previous context and handles vast amounts of intermediate data through the KV cache* and attention operations. As models reason for longer periods and maintain longer contexts, their requirements for memory capacity and bandwidth* drastically increase. The memory bottleneck becomes larger and larger as these memory demands increase. No matter how fast the compute device becomes, system efficiency declines because the system spends most of its time waiting for data to arrive at the compute units.

As AI workloads scale, High Bandwidth Memory becomes increasingly important for supplying GPUs with larger volumes of data at higher speeds. Positioned next to the GPU/NPU/TPU, HBM delivers large amounts of data at ultra-high speeds, helping reduce compute device latency. To fully utilize HBM performance, however, the entire data path must be designed as an integrated system — not only the GPU/NPU/TPU and HBM, but also storage, networks, and interconnects*.

* Key-value cache(KV cache): A memory area where an LLM stores attention computation results from previous tokens for use when generating subsequent tokens. As context length increases, the KV cache grows, increasing memory capacity and bandwidth requirements
* Memory Bandwidth: The amount of data that can be transferred from memory within a given period of time. Higher memory bandwidth allows compute devices to receive and process more data at once
* Interconnect: A communication pathway connecting system components such as processors, memory, accelerators, and storage. In AI infrastructure, interconnects are a key factor in determining the speed and efficiency of data movement

The fundamental assumptions behind system design also need to change. Processor-centric architectures largely operate by moving data to compute devices for processing. As data volumes grow and access frequency increases, architectures that send all data to a central compute device quickly reach their limits. Now, the system must instead be redesigned around where data resides and the path it takes to be processed.

One concept developed to address this challenge is memory-centric computing*. Solving the memory bottleneck in a fundamentally better way requires reducing data movement and enabling computation closer to memory. This is the basic idea behind memory-centric computing. Memory-centric computing takes data location and movement costs as key considerations in system design, optimizing memory placement and compute architecture together*. Some data may reside in HBM next to the GPU/NPU/TPU, while other data is managed in a shared memory pool*, with certain computations performed closer to memory. The best approach depends on factors such as model architecture, context length, the number of user requests, and network configuration.

The good news is that such memory-centric systems are already being designed and evaluated not only in academia but also in the industry. There is an increasing push toward eliminating data movement bottlenecks, and this is best done by not moving data as much as possible.

* Memory-centric computing: A computing paradigm designed to process data closer to memory, moving away from the existing structure that moves data to/from the compute device. It focuses on reducing data movement to improve both performance and power efficiency
* Onur Mutlu, et al., “_[Memory-Centric Computing: Solving Computing’s Memory Problem](https://arxiv.org/abs/2505.00458)_,” 17th IEEE International Memory Workshop(IMW) (2025)
* Memory pool: A shared memory resource that can be accessed by multiple devices. By allowing each device to use memory capacity as needed, memory pooling can improve overall system resource utilization

Processors, memory, and accelerators have long been connected through a variety of interfaces. As the volume of data processed by AI workloads grows rapidly, however, conventional connectivity architectures are struggling to provide the required communication bandwidth and memory access efficiency. Consequently, the importance of high-speed interconnects that enable faster and more flexible data paths between devices is increasing. Today, CXL* and NVLink/NVSwitch* are two representative technologies driving this shift in different areas and various academic works to tightly integrate components exist.

* Compute Express Link(CXL): A next-generation interconnect standard that provides high-speed connectivity among CPUs, GPUs/NPUs/TPUs, memory, and accelerators. CXL is a key technology for enabling memory expansion and sharing
* NVLink and NVSwitch: High-speed interconnect technologies developed by NVIDIA. NVLink provides high-bandwidth connections between GPUs or between GPUs and other devices, while NVSwitch extends this connectivity to enable high-speed communication among larger numbers of GPUs. These technologies are used to reduce GPU-to-GPU data bottlenecks in large-scale AI clusters

CXL is a key interconnect technology that enables memory expansion and sharing. In traditional server architectures, memory has largely functioned as a dedicated resource attached to a specific CPU or GPU/NPU/TPU. Even when one device had spare memory capacity while another faced a shortage, it was difficult to share that capacity flexibly. CXL expands connectivity among processors, accelerators, and memory devices, providing a foundation for flexible memory expansion, sharing, and pooling. With CXL, computation can also be offloaded to devices or controllers that perform processing near memory, while CXL devices can communicate with one another through the flexible interface.

NVLink and NVSwitch provide high-bandwidth connectivity within GPU clusters. When large-scale AI models run across multiple GPUs and nodes, intermediate data such as model parameters, activations, and the KV cache must move between GPUs. If GPU-to-GPU connectivity is slow or inflexible, overall processing performance can suffer regardless of how fast each individual GPU is.



CXL and NVLink/NVSwitch serve different roles: CXL focuses on memory expansion and sharing, while NVLink/NVSwitch provide high-bandwidth GPU connectivity. Both, however, support the same broader goal of reducing data bottlenecks in AI infrastructure by enabling the delivery of the required data quickly to where it is needed. When such a connection architecture is established, memory and accelerators cease to become fixed components and become system resources that can be configured around workload requirements.

The flexible connectivity enabled by high-speed interconnects is also changing how compute devices and memory are arranged. As more data paths become available between devices, compute and memory resources can be positioned according to workload requirements. This shift points toward an architecture that does not simply move more data over longer distances but instead performs the required processing closer to where the data resides.

Reducing data movement bottlenecks can go a step beyond simply increasing connection speeds. While CXL and NVLink/NVSwitch improve data paths between memory and GPUs/NPUs/TPUs in different ways, near-memory acceleration reduces the amount of data that must be moved by placing computation itself closer to the data.

Instead of having a central GPU/NPU/TPU retrieve and process all data, accelerators located in or near memory devices divide the workload among themselves. These accelerators are tightly coupled through high-speed interconnects, allowing data to be processed close to memory first, with only the necessary results exchanged with other devices, rather than sending the entire dataset to a central GPU/NPU/TPU. When computation takes place near 3D-stacked high-bandwidth memory* such as HBM or other high-capacity memory, much less data needs to travel back and forth. This can reduce latency and power consumption while making it easier to scale across multiple nodes. This architecture is known as a distributed system of tightly connected near-memory accelerators*.

The potential of this approach can be seen in Tesseract*. Tesseract is a research architecture(published at the 42nd ACM/IEEE International Symposium on Computer Architecture in 2015) that applies distributed near-memory acceleration to parallel graph processing. Multiple accelerators located close to memory process their assigned data independently and communicate with one another only when necessary. For various parallel graph-processing workloads, the research demonstrated an order-of-magnitude performance improvement along with a similar energy reduction compared with conventional processor-centric designs.

The significance of these results extends beyond the numbers themselves. By processing each node’s assigned data locally and reducing communication volume, the architecture can achieve greater scalability and energy efficiency than conventional designs. As more and diverse accelerators are added, memory capacity, memory bandwidth*, and computation capability all scale proportionally. In other words, the key is not simply distributed processing, but an architecture in which multiple near-memory accelerators are tightly connected and perform computation cooperatively with minimal data movement.

* 3D-stacked high-bandwidth memory: A memory architecture that vertically stacks multiple memory dies to provide high bandwidth. HBM is a representative example
* Distributed system of tightly connected near-memory accelerators: An architecture that connects multiple accelerators located close to memory via high-speed interconnects, enabling each accelerator to process its assigned data and exchange only the necessary results
* Tesseract: A research architecture for near-memory acceleration in which multiple accelerators located close to memory are tightly connected through high-speed interconnects. Each accelerator processes its assigned data locally and communicates with others only when necessary
* The original architecture is published at the 42nd ACM/IEEE International Symposium on Computer Architecture in 2015, with a paper entitled “_[A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing](https://ieeexplore.ieee.org/document/7284059)_”.



More recently, research has also targeted LLM inference. For example, a CXL-based LLM inference study presented at ASPLOS* 2025 demonstrated how memory expansion architectures and processing units located close to memory can help to greatly alleviate capacity and bandwidth constraints in large-scale inference infrastructure(called the CENT design)*. As memory capacity and bandwidth requirements increase, architectures that concentrate all computations on a central GPU/NPU/TPU alone face growing challenges in improving efficiency. Such distributed near-memory architectures like the CENT design that minimize data movement using near-memory processing and efficient & scalable interconnects in intelligent and novel ways can improve all metrics(performance, energy, cost, hardware area) at the same time.

* ASPLOS: Short for Architectural Support for Programming Languages and Operating Systems, a major international conference covering computer architecture, programming languages, and operating systems
* Yufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang, Xavier Servot, Onur Mutlu, Ravi Iyer, and Reetuparna Das,“_[PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference](https://arxiv.org/abs/2502.07578)_,” ASPLOS 2025

Disaggregated architecture* starts from the same fundamental challenge but takes a somewhat different approach. Unlike near-memory accelerator architectures, which move computation closer to data, disaggregated architectures organize CPU, GPU/NPU/TPU, memory, storage, and network resources into a system-wide shared pool. In this structure, the necessary resources can then be flexibly combined according to workload requirements. Workloads that require large memory capacity can use an expanded memory pool, while workloads whose performance depends on bandwidth can be handled by near-memory accelerators. In essence, resources are deployed according to the workload’s requirements rather than being fitted into a fixed server configuration. This is very much in-line with the “Asymmetry Everywhere” vision outlined in 2010 to enable customized, flexible, and reconfigurable processing to maximize energy efficiency and performance*.

* Disaggregated architecture: An architecture that separates resources such as CPUs, GPUs, memory, and storage at the system level rather than fixing them within individual servers, allowing resources to be combined as needed
* Onur Mutlu, “_[Asymmetry Everywhere(with Automatic Resource Management](https://people.inf.ethz.ch/omutlu/pub/mutlu\_asymmetry-everywhere-talk\_nsf-acar10.pdf))_,” CRA Workshop on Advancing Computer Architecture Research, 2010



Data that requires substantial memory capacity, such as the KV cache, and memory-intensive operations, such as the attention layer*, are representative use cases for this type of architecture. With high-speed, flexible interconnects and flexible, compute-capable memory configurations, these workloads can be assigned to separate memory resources or near-memory acceleration domains. This reduces the burden on GPUs/NPUs/TPUs to retrieve and process all data directly, while allowing devices to locally and efficiently perform computations and to exchange results more efficiently.

The effectiveness of a disaggregated architecture depends on balancing resource allocation with communication costs. How to minimize communication and how to perform efficient data flow within and across accelerators continue to be important questions because excessive communication can actually increase latency and energy consumption. It is therefore critical to determine which resources and workloads should be disaggregated, which should remain close to memory, and where exactly what computation should be performed.

* Attention layer: A layer in an AI model that determines which information within a given context should receive greater attention. It plays a key role in the Transformer architecture underlying LLMs, with memory usage and data access requirements increasing as context length grows.

The competitiveness of AI infrastructure depends on how fast and efficient a path it can provide to the data required for computation. High-speed interconnects such as CXL and NVLink, near-memory accelerator architectures such as Tesseract, and disaggregated architectures that flexibly and efficiently combine computation and memory resources according to workload requirements, including CXL-based memory pooling and customized near-memory processing, are all moving in this direction. This is also why major Big Tech and cloud companies continue to invest heavily in AI infrastructure architecture. The economics of AI competition now extend beyond acquiring more compute devices to how efficiently one can design AI systems across the stack with memory as a first-class citizen: enabling efficient use of memory, power, networks, silicon, and minimizing data movement to greatly improve inference performance and efficiency.

This shift is transforming AI infrastructure from static infrastructure into a dynamic system. As models grow and evolve, contexts lengthen, and inference calls increase, infrastructure must become more flexible and efficient. It must be able to select the appropriate processing path for each workload and model.

As system architectures rapidly evolve, so does the role of memory. Traditionally, memory has primarily been treated as a device for storing and supplying data. With AI workloads, however, the distance between where data resides and where computation takes place directly affects performance and power efficiency. Accordingly, memory is quickly evolving to become an active first-class citizen to perform computation and improve overall system efficiency.

Recent research results are particularly fascinating, showing that some computation is already possible within existing DRAM chips themselves*. By changing how the memory controller* accesses DRAM and activating multiple rows simultaneously, systems can perform large-scale(TeraOps per second* level) bit-level operations or data copying & initialization purely inside modern DRAM chips. These capabilities arise from the fundamental operational principles of DRAM circuitry itself. We call this approach processing using DRAM, since one “uses” (existing) DRAM chips to perform computation. This recent discovery on the extensive computational capabilities of real DRAM chips offers only a glimpse, yet an important one, of how memory could take on a more active role in future computing architectures and provides a foundation for building new processing-in-DRAM mechanisms into future DRAM chips and standards.

AI performance cannot be fully explained simply by adding up the specifications of individual components. Future competitiveness will increasingly depend on how precisely systems are designed and how computation is performed around where data resides, which devices perform computation, and which paths are used to exchange results.

* Ismail Emir Yuksel, et al., “_[Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis](https://arxiv.org/abs/2402.18736)_,” 30th International Symposium on High-Performance Computer Architecture(HPCA) (2024)
* Memory controller: A device that controls data read and write commands and timing between the processor and memory
* Tera Operations Per Second(TOPS): A throughput metric that measures the number of operations performed per second in trillions. One TOPS represents one trillion operations per second; here, it refers to the scale of bit-level operations performed within DRAM

Bottlenecks in AI infrastructure arise not only from individual components such as GPUs or memory, but especially from how computation is distributed and communication is performed, i.e., from how GPUs/NPUs/TPUs, HBM, storage, networks, and interconnects exchange data. The major bottleneck we face today is the huge amounts of data movement caused by the processor-centric paradigm. High-speed interconnects, near-memory accelerator architectures, and disaggregated architectures explored above are all system-level approaches to reducing the large burden of data movement.

Now, the question turns to the architecture of semiconductors and memory themselves. Beyond simply placing data closer to compute devices, how much more data movement could be reduced if some computations were performed within or around memory itself?

The next article will explore why the center of gravity in AI semiconductors is shifting toward memory, with a focus on HBM, advanced packaging, and processing-in-memory(PIM).



_**Disclaimer:**The opinions expressed in this article are solely those of the author and do not necessarily reflect the official position of SK hynix._