MEMORY INDUSTRY INTELLIGENCE
[AI 생태계] 진짜 병목은 컴퓨트가 아니라 데이터 | SK하이닉스 뉴스룸
한국어 번역·요약·분석
원문 제목: [AI Ecosystem] The real bottleneck: Data, not compute | SK hynix Newsroom
핵심 요약
이 글은 AI 워크로드가 사전학습에서 추론·에이전트 AI로 이동하면서 시스템 병목이 GPU 컴퓨트에서 데이터의 위치와 이동 경로로 옮겨가고 있다고 주장한다. 1994년 Wulf·McKee의 메모리 월 경고를 인용하며, 프로세서 성능은 연 약 60%, DRAM 접근 속도는 연 약 7% 개선에 그친 격차가 현재 대규모 모델 실시간 서비스에서 현실화되고 있다고 설명한다. KV 캐시·긴 컨텍스트·체인오브소트·에이전트 AI로 토큰 수와 메모리 요구가 커지며, 700억 파라미터 모델 가중치 약 140GB를 넘어설 수 있는 KV 캐시 사례를 든다. FlashAttention·PagedAttention·추측 디코딩·양자화 등 소프트웨어 최적화가 확산되고, Google이 ICLR 2026에서 발표했다고 언급된 TurboQuant는 16비트 KV 캐시를 약 3~4비트로 줄여 메모리 사용량을 약 5분의 1로 낮춘다고 소개된다. 글은 소프트웨어 압축이 이론적 한계에 가까워지면서 HBM, HBF, CXL, PIM 등 메모리와 컴퓨트를 가깝게 하는 아키텍처 재설계가 필요하다고 전망한다.
메모리 산업 영향 분석
이 원문은 SK하이닉스 뉴스룸의 AI 생태계 시리즈 2편으로, AI 병목이 컴퓨트에서 데이터 이동·메모리 아키텍처로 이동한다는 분석과 함께 HBM·HBF·CXL·PIM을 메모리-컴퓨트 근접화 방향의 기술로 소개한다. 메모리 산업에 대한 직접적 함의는 HBM이 GPU 가까이에서 데이터 대기 시간을 줄이는 제품으로 재조명되고, HBF는 NAND 기반 차세대 메모리로, CXL은 메모리 공유·분리, PIM은 메모리 내 연산으로 제시된다는 점이다. 다만 이 글은 제품 출하량·고객 인증·양산 일정·투자 계획을 제시하지 않으므로, 공급 배분·고객 인증·제품 믹스·투자 일정에 대한 구체적 선택을 원문만으로 확정할 수 없다. 분석가 가설로는 소프트웨어 압축이 이론적 한계에 가까워질수록 동일 메모리 아키텍처 내 개선 여지가 줄어 HBM·HBF·CXL·PIM 같은 아키텍처 변화의 상대적 필요성이 커질 수 있으나, 이는 원문의 전망이며 실제 채택·물량은 미확인이다. 반대 근거로는 TurboQuant 같은 소프트웨어 최적화가 KV 캐시 메모리 사용량을 약 5분의 1로 낮춰 단위 메모리 수요를 줄일 수 있다는 점이 원문에 명시되어 있어, 소프트웨어 효율 개선이 메모리 수요를 억제하는 상충 효과가 존재한다. 판단 변경 조건은 HBM·HBF·CXL·PIM의 실제 고객 인증·양산·가동이 공식 발표로 확인되는지, 그리고 소프트웨어 압축의 추가 한계 도달 여부다. research_topics의 HBM 세대·스택·인증·패키징 제약, 서버 DDR·LPDDR·SOCAMM 실제 규격, NAND·eSSD·HBF 개발·인증 단계, SK하이닉스/Solidigm 신제품·고객 인증 관련 질문은 이 원문만으로 답할 수 없어 미확인으로 남긴다.
한국어 번역 읽기
수집된 원문 v1의 분석에 제공된 본문 기준 · 15759자
AI 워크로드가 추론과 에이전트 AI로 진화하면서 시스템 성능의 기준도 바뀌고 있다. 컴퓨팅 파워만으로는 더 이상 충분하지 않다. 데이터가 어디에 저장되고, 얼마나 빠르게 이동하며, 얼마나 효율적으로 처리되는지가 이제 AI 시스템의 성능과 효율에 직접 영향을 미친다.
이 기사 시리즈는 소프트웨어, 데이터센터 인프라, 반도체, 메모리 기술이 함께 작동하는 AI 생태계의 변화를 탐구한다. 또한 메모리 반도체가 왜 AI 시대의 기반 기술이 되었는지 살펴보고, 차세대 AI 메모리 생태계를 구축하기 위한 SK하이닉스의 비전을 제시한다.
① AI 컴퓨팅의 패러다임 전환
**② 진짜 병목: 컴퓨트가 아니라 데이터 – KAIST 유회준 교수**
③ 인프라 재설계: 아키텍처가 성능을 결정하는 이유
④ 반도체 패러다임이 메모리로 이동
⑤ AI 완성: 피지컬 인텔리전스와 메모리의 역할
AI의 무게 중심이 사전학습에서 추론으로, 일회성 응답에서 에이전트 AI로 이동하면서 시스템 병목의 위치도 바뀌었다. 핵심 문제는 점점 컴퓨트 자체의 양보다 데이터의 흐름에 있다. 모델이 더 긴 컨텍스트를 유지하고 중간 결과를 반복적으로 참조하면서 메모리와 컴퓨트 장치 사이의 데이터 교환 빈도가 증가한다.
따라서 질문은 더 이상 GPU 컴퓨트 능력에 국한되지 않는다. AI 성능은 정말 GPU가 얼마나 계산할 수 있는지에 의해 제약되는가, 아니면 메모리 아키텍처가 필요한 데이터를 검색해 필요할 때 컴퓨트 장치에 전달하는 능력에 의해 제약되는가? AI 시대의 핵심 병목은 컴퓨트 자체가 아니라 데이터가 어디에 있고 시스템을 통해 어떻게 이동하는지에 있다.
지난 10년 동안 AI 성능을 향상시키는 공식은 명확해 보였다. 더 많은 GPU, 더 큰 모델, 더 긴 학습을 우선시하는 컴퓨트 중심 스케일링*이다. 이 공식은 대규모 사전학습이 AI 성능 향상의 주요 경로였던 시대에는 유효했다. 더 많은 GPU를 배치하면 더 큰 모델을 더 오래 학습할 수 있었고, 이는 모델 정확도와 범용성을 모두 향상시켰다.
* Scaling: GPU, 모델 크기, 학습 데이터 양 등 컴퓨트 자원을 늘려 AI 성능을 향상시키는 접근법
그 결과 GPU는 AI 시대 성능의 상징이 되었다. 기업들은 얼마나 많은 GPU를 확보했는지, 클러스터 규모가 얼마나 큰지, 초당 얼마나 많은 연산을 수행할 수 있는지를 강조하기 위해 경쟁했다. 이러한 지표는 분명 중요하다. GPU의 병렬 컴퓨트 능력 없이는 오늘날의 대규모 AI 모델을 달성하기 어렵다.
그러나 AI 시스템의 실제 성능은 GPU의 이론적 컴퓨트 능력만으로 결정되지 않는다. 시스템에 컴퓨트 장치가 아무리 많아도 필요한 데이터가 제때 도착하지 않으면 최대 성능을 발휘할 수 없다. 즉, 중요한 것은 GPU가 얼마나 빨리 계산할 수 있는지만이 아니라, 필요한 데이터가 컴퓨트 장치 가까이에 위치해 지연 없이 접근되어야 GPU 성능이 완전히 실현된다는 점이다.
이 구분은 추론 중심 AI와 에이전트 AI의 확산으로 더욱 뚜렷해졌다. 대규모 학습에서는 많은 데이터를 배치*로 함께 처리해 컴퓨트 장치를 비교적 잘 활용할 수 있었다. 반면 실제 추론 환경은 낮은 지연, 긴 컨텍스트, 반복 호출을 요구한다. 모델은 필요한 정보를 반복적으로 참조하고, 중간 결과를 저장하며, 다음 토큰*을 생성하기 위해 데이터를 계속 검색한다.
* Batch: AI 모델이 여러 데이터를 함께 처리하는 단위. 큰 배치는 대규모 학습에서 GPU를 효율적으로 활용하는 데 도움이 되지만, 낮은 지연 요구와 개별 요청 처리 필요 때문에 실제 서비스 추론에는 사용하기 더 어렵다
* Token: AI 모델이 처리하는 텍스트의 기본 단위. LLM은 이전 토큰을 참조해 각 새 토큰을 순차적으로 생성하며, 토큰 수가 늘어날수록 추론 시간과 메모리 사용량도 증가한다.
결과적으로 병목의 초점은 GPU 컴퓨트 능력에서 데이터의 위치와 이동 경로로 이동하고 있다. 핵심 과제는 메모리 대역폭과 접근 지연의 한계, 즉 메모리 월*을 어떻게 극복할 것인가이다.
* Memory wall: 메모리 데이터 전달 속도 향상이 프로세서 성능 발전을 따라가지 못할 때 발생하는 병목
1994년 미국 컴퓨터 과학자 William A. Wulf와 Sally A. McKee는 짧은 논문*에서 프로세서 성능과 메모리 접근 속도 사이의 격차가 커지고 있음을 지적했다. 그들은 프로세서 성능이 매년 약 60%씩 개선되는 반면 DRAM 접근 속도는 약 7%만 개선되고 있다고 경고했다. 이 격차가 계속 벌어지면 시스템 성능이 궁극적으로 메모리 속도에 의해 제약될 수 있었다. 이 병목은 메모리 월로 알려지게 되었다.
* Wm. A. Wulf and Sally A. McKee, “Hitting the Memory Wall: Implications of the Obvious.” ACM SIGARCH Computer Architecture News, Vol. 23, No. 1 (1995): 20-24
30년 후, 그 경고는 수십억에서 수천억 개의 파라미터를 가진 대규모 모델이 실시간으로 작동하는 AI 서비스 환경에서 현실이 되고 있다. 이러한 모델을 처리하는 최신 AI 가속기의 컴퓨트 성능은 페타플롭스* 범위에 도달했지만, 모델 가중치와 활성화를 메모리에서 검색할 수 있는 속도는 컴퓨트 발전을 따라가지 못했다. 그 결과 대형 칩에 통합된 많은 컴퓨트 유닛은 데이터를 기다리는 데 더 많은 시간을 보내며 실제 활용도가 낮아진다.
이 격차를 좁히는 두 가지 주요 접근법이 있다. 메모리 대역폭 자체를 늘리거나, 메모리를 프로세서 가까이에 배치해 데이터가 이동해야 하는 거리를 줄이는 것이다.
* Petaflops(PFLOPS): 초당 1천조 회의 부동소수점 연산에 해당하는 컴퓨트 성능 단위
대규모 언어 모델*과 추론 집약적 워크로드가 확산되면서 메모리 병목은 더 복잡해졌다. 이전의 일회성 추론 작업과 달리 오늘날 모델은 단일 쿼리에 응답해 수천, 때로는 수만 개의 토큰을 순차적으로 생성한다. 체인오브소트*와 에이전트 AI 같은 단계별 추론 접근법은 토큰 수를 더 높일 수 있다.
* Large language model(LLM): 대량의 텍스트 데이터로 학습해 맥락을 이해하고 응답을 생성하는 AI 모델. 시퀀스가 길어질수록 모델은 이전 토큰의 정보를 계속 참조해야 하므로 추론 중 상당한 메모리와 데이터 이동이 필요하다.
* Chain-of-thought(CoT): 모델이 답을 내기 전에 중간 추론 단계를 거치는 프롬프팅 및 생성 기법으로, 추론 정확도를 높이는 데 도움이 된다
이 과정에서 가장 큰 과제 중 하나는 KV 캐시*이다. Transformer* 기반 모델이 토큰을 생성할 때마다 모든 선행 토큰의 핵심 정보를 key와 value 벡터로 메모리에 유지한다. 시퀀스가 길어질수록 캐시도 커진다. 일부 긴 컨텍스트 추론 환경에서는 KV 캐시에 필요한 메모리가 700억 파라미터 모델의 가중치를 담는 데 필요한 약 140GB를 초과할 수도 있다.
* KV cache: 이전에 생성된 key와 value 벡터를 저장하고 재사용해 동일한 계산을 반복하는 비효율을 피하는 기법
* Transformer: Google이 2017년에 도입한 신경망 아키텍처. 자기 주의를 사용해 맥락을 이해하며 거의 모든 현대 LLM의 기반을 이룬다.
다시 말해, 사용자가 ChatGPT 같은 LLM 서비스에 쿼리를 보내면 GPU 작업의 절반 이상이 계산이 아니라 데이터 이동에 관련될 수 있다. 이는 단순히 GPU를 더 추가하는 것으로 해결할 수 없는 구조적 과제다.
이러한 병목을 완화하기 위해 업계는 다양한 소프트웨어 기반 솔루션을 채택했다. 가장 두드러진 것 중 하나는 FlashAttention*이다. 이 기법은 LLM의 핵심 연산인 attention*을 더 작은 블록으로 나누고 가능한 한 GPU 내부 SRAM*에서 처리하도록 재구성한다. 이는 HBM과 SRAM 사이의 중간 데이터 반복 읽기·쓰기를 줄여 동일 하드웨어에서 추론 효율을 향상시킨다. 메모리 사용 패턴과 토큰 생성을 최적화하는 PagedAttention*과 추측 디코딩* 같은 기법도 빠르게 주목받고 있다.
* FlashAttention: attention 연산을 더 작은 블록으로 나누어 GPU의 빠른 온칩 SRAM에서 처리함으로써 외부 메모리와의 데이터 전송을 줄이는 알고리즘
* Attention: AI 모델이 다음 단어를 생성할 때 앞선 맥락을 참조하는 과정
* Static random-access memory(SRAM): CPU와 GPU 내부 캐시에 사용되는 고속 메모리. DRAM보다 빠르지만 더 많은 회로 면적이 필요해 대용량 저장에는 적합하지 않다. AI 컴퓨트에서 SRAM은 자주 접근하는 중간 데이터를 프로세서 가까이에 유지하는 데 사용된다
* PagedAttention: attention 연산 중 GPU 메모리를 가상 메모리처럼 페이지 단위로 관리하는 기법. 불필요한 메모리 낭비를 줄여 더 많은 사용자 요청을 동시에 처리하고 추론 효율을 높인다
* Speculative decoding: 더 작은 보조 모델이 여러 토큰을 미리 예측하고 주 모델이 이를 함께 검증하는 추론 가속 기법. 생성 품질을 유지하면서 응답 속도를 높이는 데 사용된다.
또 다른 주요 트렌드는 양자화*이다. 이 방법은 모델 가중치, 활성화 값, KV 캐시를 16비트 대신 8비트나 4비트 같은 더 적은 비트로 표현하고 저장한다. 표현할 데이터가 줄어들면 메모리 사용량과 데이터 이동이 모두 감소한다. 최근에는 유사한 값을 함께 묶어 표현을 개선하는 데이터 형식, 자주 발생하는 패턴을 더 짧은 코드로 대체하는 벡터 양자화* 등 데이터 표현 효율을 높이는 다양한 압축 기법이 등장했다.
* Vector quantization(VQ): 모델 가중치, 활성화, KV 캐시 등의 데이터를 더 적은 비트로 표현해 메모리 사용량과 데이터 이동을 줄이는 기법. 예를 들어 16비트 데이터를 8비트나 4비트로 표현하면 저장 요구량과 데이터 전송 오버헤드를 모두 줄일 수 있다
최근 이 분야에서 주목받는 발전 중 하나는 KV 캐시를 더 적은 비트로 표현해 메모리 사용량을 줄이는 양자화 기법인 TurboQuant*이다. 기존 압축 기법은 입력 데이터 분포가 바뀔 때마다 압축 규칙을 재보정해야 했다. 이 기법은 정확도가 높지만 처리 시간도 증가한다.
* TurboQuant: Google이 ICLR 2026에서 발표한 KV 캐시 압축 기법. 회전 변환을 사용해 입력 데이터를 표준화된 분포로 매핑한 후 미리 정의된 압축 규칙을 적용해 16비트
[원문 후반 발췌]
미리 정의된 규칙 세트로. 이 아이디어는 고객마다 새 옷을 맞추는 대신 고객의 체형을 표준화한 후 같은 사이즈의 옷을 입히는 것과 유사하다. 연구에 따르면 TurboQuant는 16비트로 저장된 KV 캐시를 약 3~4비트로 줄여 응답 품질을 크게 훼손하지 않으면서 메모리 사용량을 약 5분의 1로 낮춘다.
이 TurboQuant 기법은 압축 효율, 즉 동일한 데이터를 더 적은 비트로 표현하는 것을 이론적 한계에 가깝게 밀어붙인다. 다시 말해 정보 손실 없이 데이터를 더 압축할 여지는 제한적일 수 있다.
여기에는 중요한 함의가 있다. 소프트웨어 압축이 이론적 한계에 접근하면서 동일한 메모리 아키텍처 내에서 추가 개선의 가능성도 줄어든다. 소프트웨어 최적화는 분명 효과적이지만, 컴퓨트 장치가 더 멀리 있는 메모리에서 데이터를 계속 가져와야 한다면 근본적 한계가 남는다.
그렇다면 다음은 무엇인가? 출발점은 메모리와 컴퓨트를 더 가깝게 배치해 필요한 데이터를 더 짧은 경로로 처리하는 아키텍처를 만드는 것이다. 이는 단순히 더 나은 알고리즘을 찾는 문제가 아니다. 시스템과 반도체 자체의 아키텍처를 재설계해야 하는 문제다.
여러 가지가
[원문 후반 발췌]
예는 High Bandwidth Memory*이다. HBM은 여러 DRAM 다이를 수직으로 적층하고 GPU 가까이에 배치해 데이터가 컴퓨트 장치에 도달하는 거리를 단축하면서 한 번에 전송할 수 있는 데이터 양을 늘린다. HBM은 단순히 메모리 속도를 높이기 때문이 아니라 컴퓨트 장치가 데이터를 기다리는 시간을 줄여주기 때문에 주목받고 있다.
* High Bandwidth Memory: 여러 DRAM 다이를 적층하고 through-silicon via(TSV) 기술로 연결해 GPU와 AI 가속기에 매우 높은 대역폭을 제공하는 메모리.
이와 같은 비판적 관점은 다른 메모리와 시스템 기술도 이끌고 있다. High Bandwidth Flash*는 더 큰 용량과 더 높은 대역폭을 모두 확보하려 하고, Compute Express Link*는 메모리를 특정 GPU 전용 자원에서 시스템 수준의 공유 자원으로 전환하려는 움직임을 반영한다. Processing-in-memor*는 한 걸음 더 나아가 일부 컴퓨트를 데이터가 있는 메모리 가까이에서 처리한다. 이 세 기술은 접근 방식은 다르지만 모두 같은 질문에서 출발한다. 데이터가 처리를 위해 계속 먼 거리를 이동해야 하는가, 아니면 더 많은 작업을 데이터가 있는 곳 가까이에서 처리해야 하는가?
* High Bandwidth Flash(HBF): 차세대 NAND 기반 메모리 기술
[원문 후반 발췌]
과 SSD. HBM보다 훨씬 큰 용량을 제공하면서 비슷한 대역폭과 더 큰 비용 효율을 목표로 한다.
* Compute Express Link(CXL): 일관성 캐시 프로토콜을 통해 CPU, GPU, 메모리, 가속기를 연결해 메모리를 시스템 수준에서 분리하고 공유할 수 있게 하는 차세대 인터커넥트 표준
* Processing-in-memory(PIM): 메모리에 처리 능력을 추가해 메모리와 프로세서 사이의 데이터 병목을 줄이고 성능을 크게 향상시키는 데 도움이 되는 차세대 메모리 기술
궁극적으로 이러한 접근법은 같은 방향을 가리킨다. 컴퓨트 장치 가까이에 둘 수 있는 데이터는 가까이 있어야 하고, 이동해야 하는 데이터에는 더 넓고 효율적인 경로가 필요하다. AI 병목은 더 이상 단일 칩 수준에서 해결할 수 있는 문제가 아니다. 모델 아키텍처, GPU, 메모리, 인터커넥트, 패키징, 데이터센터 아키텍처 사이의 복잡한 연결을 포함하는 시스템 수준의 과제다.
이는 컴퓨트 장치와 메모리를 분리하는 기존 폰 노이만 아키텍처*가 곧 사라진다는 뜻은 아니다. 그러나 수십 년 동안 당연하게 여겨졌던 접근법, 즉 프로세서와 메모리를 분리하고 버스로 연결하는 방식은
[원문 후반 발췌]
메모리와 컴퓨트 장치. 결과적으로 AI 성능은 개별 GPU의 컴퓨트 능력보다 데이터가 어디에 저장되고, 시스템을 통해 어떻게 이동하며, 얼마나 효율적으로 처리되는지에 점점 더 의존한다.
이 과제는 개별 칩이나 메모리 기술만으로 해결하기 어렵다. GPU, 메모리, 스토리지, 네트워크, 데이터센터가 하나의 시스템으로 함께 작동해야 한다. 다음 편에서는 이 과제를 AI 인프라 관점에서 살펴본다. 궁극적으로 AI 성능을 결정하는 것은 더 빠른 구성 요소인가, 아니면 컴퓨트와 데이터가 흐르는 전체 아키텍처인가?
_**면책 조항:** 이 기사에 표현된 의견은 전적으로 저자의 것이며 SK하이닉스의 공식 입장을 반드시 반영하지는 않습니다._
이 기사 시리즈는 소프트웨어, 데이터센터 인프라, 반도체, 메모리 기술이 함께 작동하는 AI 생태계의 변화를 탐구한다. 또한 메모리 반도체가 왜 AI 시대의 기반 기술이 되었는지 살펴보고, 차세대 AI 메모리 생태계를 구축하기 위한 SK하이닉스의 비전을 제시한다.
① AI 컴퓨팅의 패러다임 전환
**② 진짜 병목: 컴퓨트가 아니라 데이터 – KAIST 유회준 교수**
③ 인프라 재설계: 아키텍처가 성능을 결정하는 이유
④ 반도체 패러다임이 메모리로 이동
⑤ AI 완성: 피지컬 인텔리전스와 메모리의 역할
AI의 무게 중심이 사전학습에서 추론으로, 일회성 응답에서 에이전트 AI로 이동하면서 시스템 병목의 위치도 바뀌었다. 핵심 문제는 점점 컴퓨트 자체의 양보다 데이터의 흐름에 있다. 모델이 더 긴 컨텍스트를 유지하고 중간 결과를 반복적으로 참조하면서 메모리와 컴퓨트 장치 사이의 데이터 교환 빈도가 증가한다.
따라서 질문은 더 이상 GPU 컴퓨트 능력에 국한되지 않는다. AI 성능은 정말 GPU가 얼마나 계산할 수 있는지에 의해 제약되는가, 아니면 메모리 아키텍처가 필요한 데이터를 검색해 필요할 때 컴퓨트 장치에 전달하는 능력에 의해 제약되는가? AI 시대의 핵심 병목은 컴퓨트 자체가 아니라 데이터가 어디에 있고 시스템을 통해 어떻게 이동하는지에 있다.
지난 10년 동안 AI 성능을 향상시키는 공식은 명확해 보였다. 더 많은 GPU, 더 큰 모델, 더 긴 학습을 우선시하는 컴퓨트 중심 스케일링*이다. 이 공식은 대규모 사전학습이 AI 성능 향상의 주요 경로였던 시대에는 유효했다. 더 많은 GPU를 배치하면 더 큰 모델을 더 오래 학습할 수 있었고, 이는 모델 정확도와 범용성을 모두 향상시켰다.
* Scaling: GPU, 모델 크기, 학습 데이터 양 등 컴퓨트 자원을 늘려 AI 성능을 향상시키는 접근법
그 결과 GPU는 AI 시대 성능의 상징이 되었다. 기업들은 얼마나 많은 GPU를 확보했는지, 클러스터 규모가 얼마나 큰지, 초당 얼마나 많은 연산을 수행할 수 있는지를 강조하기 위해 경쟁했다. 이러한 지표는 분명 중요하다. GPU의 병렬 컴퓨트 능력 없이는 오늘날의 대규모 AI 모델을 달성하기 어렵다.
그러나 AI 시스템의 실제 성능은 GPU의 이론적 컴퓨트 능력만으로 결정되지 않는다. 시스템에 컴퓨트 장치가 아무리 많아도 필요한 데이터가 제때 도착하지 않으면 최대 성능을 발휘할 수 없다. 즉, 중요한 것은 GPU가 얼마나 빨리 계산할 수 있는지만이 아니라, 필요한 데이터가 컴퓨트 장치 가까이에 위치해 지연 없이 접근되어야 GPU 성능이 완전히 실현된다는 점이다.
이 구분은 추론 중심 AI와 에이전트 AI의 확산으로 더욱 뚜렷해졌다. 대규모 학습에서는 많은 데이터를 배치*로 함께 처리해 컴퓨트 장치를 비교적 잘 활용할 수 있었다. 반면 실제 추론 환경은 낮은 지연, 긴 컨텍스트, 반복 호출을 요구한다. 모델은 필요한 정보를 반복적으로 참조하고, 중간 결과를 저장하며, 다음 토큰*을 생성하기 위해 데이터를 계속 검색한다.
* Batch: AI 모델이 여러 데이터를 함께 처리하는 단위. 큰 배치는 대규모 학습에서 GPU를 효율적으로 활용하는 데 도움이 되지만, 낮은 지연 요구와 개별 요청 처리 필요 때문에 실제 서비스 추론에는 사용하기 더 어렵다
* Token: AI 모델이 처리하는 텍스트의 기본 단위. LLM은 이전 토큰을 참조해 각 새 토큰을 순차적으로 생성하며, 토큰 수가 늘어날수록 추론 시간과 메모리 사용량도 증가한다.
결과적으로 병목의 초점은 GPU 컴퓨트 능력에서 데이터의 위치와 이동 경로로 이동하고 있다. 핵심 과제는 메모리 대역폭과 접근 지연의 한계, 즉 메모리 월*을 어떻게 극복할 것인가이다.
* Memory wall: 메모리 데이터 전달 속도 향상이 프로세서 성능 발전을 따라가지 못할 때 발생하는 병목
1994년 미국 컴퓨터 과학자 William A. Wulf와 Sally A. McKee는 짧은 논문*에서 프로세서 성능과 메모리 접근 속도 사이의 격차가 커지고 있음을 지적했다. 그들은 프로세서 성능이 매년 약 60%씩 개선되는 반면 DRAM 접근 속도는 약 7%만 개선되고 있다고 경고했다. 이 격차가 계속 벌어지면 시스템 성능이 궁극적으로 메모리 속도에 의해 제약될 수 있었다. 이 병목은 메모리 월로 알려지게 되었다.
* Wm. A. Wulf and Sally A. McKee, “Hitting the Memory Wall: Implications of the Obvious.” ACM SIGARCH Computer Architecture News, Vol. 23, No. 1 (1995): 20-24
30년 후, 그 경고는 수십억에서 수천억 개의 파라미터를 가진 대규모 모델이 실시간으로 작동하는 AI 서비스 환경에서 현실이 되고 있다. 이러한 모델을 처리하는 최신 AI 가속기의 컴퓨트 성능은 페타플롭스* 범위에 도달했지만, 모델 가중치와 활성화를 메모리에서 검색할 수 있는 속도는 컴퓨트 발전을 따라가지 못했다. 그 결과 대형 칩에 통합된 많은 컴퓨트 유닛은 데이터를 기다리는 데 더 많은 시간을 보내며 실제 활용도가 낮아진다.
이 격차를 좁히는 두 가지 주요 접근법이 있다. 메모리 대역폭 자체를 늘리거나, 메모리를 프로세서 가까이에 배치해 데이터가 이동해야 하는 거리를 줄이는 것이다.
* Petaflops(PFLOPS): 초당 1천조 회의 부동소수점 연산에 해당하는 컴퓨트 성능 단위
대규모 언어 모델*과 추론 집약적 워크로드가 확산되면서 메모리 병목은 더 복잡해졌다. 이전의 일회성 추론 작업과 달리 오늘날 모델은 단일 쿼리에 응답해 수천, 때로는 수만 개의 토큰을 순차적으로 생성한다. 체인오브소트*와 에이전트 AI 같은 단계별 추론 접근법은 토큰 수를 더 높일 수 있다.
* Large language model(LLM): 대량의 텍스트 데이터로 학습해 맥락을 이해하고 응답을 생성하는 AI 모델. 시퀀스가 길어질수록 모델은 이전 토큰의 정보를 계속 참조해야 하므로 추론 중 상당한 메모리와 데이터 이동이 필요하다.
* Chain-of-thought(CoT): 모델이 답을 내기 전에 중간 추론 단계를 거치는 프롬프팅 및 생성 기법으로, 추론 정확도를 높이는 데 도움이 된다
이 과정에서 가장 큰 과제 중 하나는 KV 캐시*이다. Transformer* 기반 모델이 토큰을 생성할 때마다 모든 선행 토큰의 핵심 정보를 key와 value 벡터로 메모리에 유지한다. 시퀀스가 길어질수록 캐시도 커진다. 일부 긴 컨텍스트 추론 환경에서는 KV 캐시에 필요한 메모리가 700억 파라미터 모델의 가중치를 담는 데 필요한 약 140GB를 초과할 수도 있다.
* KV cache: 이전에 생성된 key와 value 벡터를 저장하고 재사용해 동일한 계산을 반복하는 비효율을 피하는 기법
* Transformer: Google이 2017년에 도입한 신경망 아키텍처. 자기 주의를 사용해 맥락을 이해하며 거의 모든 현대 LLM의 기반을 이룬다.
다시 말해, 사용자가 ChatGPT 같은 LLM 서비스에 쿼리를 보내면 GPU 작업의 절반 이상이 계산이 아니라 데이터 이동에 관련될 수 있다. 이는 단순히 GPU를 더 추가하는 것으로 해결할 수 없는 구조적 과제다.
이러한 병목을 완화하기 위해 업계는 다양한 소프트웨어 기반 솔루션을 채택했다. 가장 두드러진 것 중 하나는 FlashAttention*이다. 이 기법은 LLM의 핵심 연산인 attention*을 더 작은 블록으로 나누고 가능한 한 GPU 내부 SRAM*에서 처리하도록 재구성한다. 이는 HBM과 SRAM 사이의 중간 데이터 반복 읽기·쓰기를 줄여 동일 하드웨어에서 추론 효율을 향상시킨다. 메모리 사용 패턴과 토큰 생성을 최적화하는 PagedAttention*과 추측 디코딩* 같은 기법도 빠르게 주목받고 있다.
* FlashAttention: attention 연산을 더 작은 블록으로 나누어 GPU의 빠른 온칩 SRAM에서 처리함으로써 외부 메모리와의 데이터 전송을 줄이는 알고리즘
* Attention: AI 모델이 다음 단어를 생성할 때 앞선 맥락을 참조하는 과정
* Static random-access memory(SRAM): CPU와 GPU 내부 캐시에 사용되는 고속 메모리. DRAM보다 빠르지만 더 많은 회로 면적이 필요해 대용량 저장에는 적합하지 않다. AI 컴퓨트에서 SRAM은 자주 접근하는 중간 데이터를 프로세서 가까이에 유지하는 데 사용된다
* PagedAttention: attention 연산 중 GPU 메모리를 가상 메모리처럼 페이지 단위로 관리하는 기법. 불필요한 메모리 낭비를 줄여 더 많은 사용자 요청을 동시에 처리하고 추론 효율을 높인다
* Speculative decoding: 더 작은 보조 모델이 여러 토큰을 미리 예측하고 주 모델이 이를 함께 검증하는 추론 가속 기법. 생성 품질을 유지하면서 응답 속도를 높이는 데 사용된다.
또 다른 주요 트렌드는 양자화*이다. 이 방법은 모델 가중치, 활성화 값, KV 캐시를 16비트 대신 8비트나 4비트 같은 더 적은 비트로 표현하고 저장한다. 표현할 데이터가 줄어들면 메모리 사용량과 데이터 이동이 모두 감소한다. 최근에는 유사한 값을 함께 묶어 표현을 개선하는 데이터 형식, 자주 발생하는 패턴을 더 짧은 코드로 대체하는 벡터 양자화* 등 데이터 표현 효율을 높이는 다양한 압축 기법이 등장했다.
* Vector quantization(VQ): 모델 가중치, 활성화, KV 캐시 등의 데이터를 더 적은 비트로 표현해 메모리 사용량과 데이터 이동을 줄이는 기법. 예를 들어 16비트 데이터를 8비트나 4비트로 표현하면 저장 요구량과 데이터 전송 오버헤드를 모두 줄일 수 있다
최근 이 분야에서 주목받는 발전 중 하나는 KV 캐시를 더 적은 비트로 표현해 메모리 사용량을 줄이는 양자화 기법인 TurboQuant*이다. 기존 압축 기법은 입력 데이터 분포가 바뀔 때마다 압축 규칙을 재보정해야 했다. 이 기법은 정확도가 높지만 처리 시간도 증가한다.
* TurboQuant: Google이 ICLR 2026에서 발표한 KV 캐시 압축 기법. 회전 변환을 사용해 입력 데이터를 표준화된 분포로 매핑한 후 미리 정의된 압축 규칙을 적용해 16비트
[원문 후반 발췌]
미리 정의된 규칙 세트로. 이 아이디어는 고객마다 새 옷을 맞추는 대신 고객의 체형을 표준화한 후 같은 사이즈의 옷을 입히는 것과 유사하다. 연구에 따르면 TurboQuant는 16비트로 저장된 KV 캐시를 약 3~4비트로 줄여 응답 품질을 크게 훼손하지 않으면서 메모리 사용량을 약 5분의 1로 낮춘다.
이 TurboQuant 기법은 압축 효율, 즉 동일한 데이터를 더 적은 비트로 표현하는 것을 이론적 한계에 가깝게 밀어붙인다. 다시 말해 정보 손실 없이 데이터를 더 압축할 여지는 제한적일 수 있다.
여기에는 중요한 함의가 있다. 소프트웨어 압축이 이론적 한계에 접근하면서 동일한 메모리 아키텍처 내에서 추가 개선의 가능성도 줄어든다. 소프트웨어 최적화는 분명 효과적이지만, 컴퓨트 장치가 더 멀리 있는 메모리에서 데이터를 계속 가져와야 한다면 근본적 한계가 남는다.
그렇다면 다음은 무엇인가? 출발점은 메모리와 컴퓨트를 더 가깝게 배치해 필요한 데이터를 더 짧은 경로로 처리하는 아키텍처를 만드는 것이다. 이는 단순히 더 나은 알고리즘을 찾는 문제가 아니다. 시스템과 반도체 자체의 아키텍처를 재설계해야 하는 문제다.
여러 가지가
[원문 후반 발췌]
예는 High Bandwidth Memory*이다. HBM은 여러 DRAM 다이를 수직으로 적층하고 GPU 가까이에 배치해 데이터가 컴퓨트 장치에 도달하는 거리를 단축하면서 한 번에 전송할 수 있는 데이터 양을 늘린다. HBM은 단순히 메모리 속도를 높이기 때문이 아니라 컴퓨트 장치가 데이터를 기다리는 시간을 줄여주기 때문에 주목받고 있다.
* High Bandwidth Memory: 여러 DRAM 다이를 적층하고 through-silicon via(TSV) 기술로 연결해 GPU와 AI 가속기에 매우 높은 대역폭을 제공하는 메모리.
이와 같은 비판적 관점은 다른 메모리와 시스템 기술도 이끌고 있다. High Bandwidth Flash*는 더 큰 용량과 더 높은 대역폭을 모두 확보하려 하고, Compute Express Link*는 메모리를 특정 GPU 전용 자원에서 시스템 수준의 공유 자원으로 전환하려는 움직임을 반영한다. Processing-in-memor*는 한 걸음 더 나아가 일부 컴퓨트를 데이터가 있는 메모리 가까이에서 처리한다. 이 세 기술은 접근 방식은 다르지만 모두 같은 질문에서 출발한다. 데이터가 처리를 위해 계속 먼 거리를 이동해야 하는가, 아니면 더 많은 작업을 데이터가 있는 곳 가까이에서 처리해야 하는가?
* High Bandwidth Flash(HBF): 차세대 NAND 기반 메모리 기술
[원문 후반 발췌]
과 SSD. HBM보다 훨씬 큰 용량을 제공하면서 비슷한 대역폭과 더 큰 비용 효율을 목표로 한다.
* Compute Express Link(CXL): 일관성 캐시 프로토콜을 통해 CPU, GPU, 메모리, 가속기를 연결해 메모리를 시스템 수준에서 분리하고 공유할 수 있게 하는 차세대 인터커넥트 표준
* Processing-in-memory(PIM): 메모리에 처리 능력을 추가해 메모리와 프로세서 사이의 데이터 병목을 줄이고 성능을 크게 향상시키는 데 도움이 되는 차세대 메모리 기술
궁극적으로 이러한 접근법은 같은 방향을 가리킨다. 컴퓨트 장치 가까이에 둘 수 있는 데이터는 가까이 있어야 하고, 이동해야 하는 데이터에는 더 넓고 효율적인 경로가 필요하다. AI 병목은 더 이상 단일 칩 수준에서 해결할 수 있는 문제가 아니다. 모델 아키텍처, GPU, 메모리, 인터커넥트, 패키징, 데이터센터 아키텍처 사이의 복잡한 연결을 포함하는 시스템 수준의 과제다.
이는 컴퓨트 장치와 메모리를 분리하는 기존 폰 노이만 아키텍처*가 곧 사라진다는 뜻은 아니다. 그러나 수십 년 동안 당연하게 여겨졌던 접근법, 즉 프로세서와 메모리를 분리하고 버스로 연결하는 방식은
[원문 후반 발췌]
메모리와 컴퓨트 장치. 결과적으로 AI 성능은 개별 GPU의 컴퓨트 능력보다 데이터가 어디에 저장되고, 시스템을 통해 어떻게 이동하며, 얼마나 효율적으로 처리되는지에 점점 더 의존한다.
이 과제는 개별 칩이나 메모리 기술만으로 해결하기 어렵다. GPU, 메모리, 스토리지, 네트워크, 데이터센터가 하나의 시스템으로 함께 작동해야 한다. 다음 편에서는 이 과제를 AI 인프라 관점에서 살펴본다. 궁극적으로 AI 성능을 결정하는 것은 더 빠른 구성 요소인가, 아니면 컴퓨트와 데이터가 흐르는 전체 아키텍처인가?
_**면책 조항:** 이 기사에 표현된 의견은 전적으로 저자의 것이며 SK하이닉스의 공식 입장을 반드시 반영하지는 않습니다._
본문이 길어 18075자 중 15759자만 처리했습니다.
브리프용 요약 초안
SK하이닉스 뉴스룸은 AI 병목이 GPU 컴퓨트에서 데이터 위치·이동 경로로 이동하고 있다고 분석하며, 소프트웨어 압축이 이론적 한계에 가까워지면서 HBM·HBF·CXL·PIM 등 메모리-컴퓨트 근접화 아키텍처가 필요하다고 전망했다. 다만 이 글은 제품 출하·고객 인증·양산 일정·투자 계획을 제시하지 않아 메모리 공급 배분·제품 믹스·투자 일정에 대한 구체적 선택은 미확인이다.
원문 텍스트
원문 열기 ↗As AI workloads evolve toward inference and agentic AI, the criteria for system performance are also changing. Computing power alone is no longer enough. Where data is stored, how quickly it moves, and how efficiently it is processed now have a direct impact on the performance and efficiency of AI systems.
This article series explores the transformation of the AI ecosystem, where software, data center infrastructure, semiconductors, and memory technologies work together. It also examines why memory semiconductors have become a foundational technology for the AI era and outlines SK hynix’s vision for building the next-generation AI memory ecosystem.
① The paradigm shift in AI computing
**② The real bottleneck: Data, not compute – Professor Hoi-Jun Yoo, KAIST**
③ Redesigning infrastructure: Why architecture determines performance
④ The semiconductor paradigm shifts toward memory
⑤ Completing AI: Physical intelligence and the role of memory
As the center of gravity of AI shifts from pretraining to inference and from one-off responses to agentic AI, the location of system bottlenecks has also changed. The key issue increasingly lies in the flow of data rather than the volume of compute itself. As models maintain longer contexts and repeatedly reference intermediate results, the frequency of data exchanges between memory and compute devices increases.
The question, therefore, is no longer limited to GPU compute capability. Is AI performance really constrained by how much a GPU can compute or by the memory architecture’s ability to retrieve the required data and deliver it to compute devices when needed? In the AI era, the core bottleneck lies not in compute itself, but in where data resides and how it moves through the system.
For the past decade, the formula for improving AI performance seemed clear: compute-centric scaling* that prioritizes more GPUs, larger models, and longer training runs. This formula held true during the era when large-scale pretraining was the primary path to improving AI performance. Deploying more GPUs made it possible to train larger models for longer periods, which in turn improved both model accuracy and versatility.
* Scaling: An approach to improving AI performance by increasing compute resources such as GPUs, model size, and the amount of training data
As a result, GPUs became a symbol of performance in the AI era. Companies competed to highlight how many GPUs they had secured, the size of their clusters, and how many operations they could perform per second. These metrics clearly matter. Without the parallel compute capabilities of GPUs, today’s large-scale AI models would be difficult to achieve.
However, the real-world performance of an AI system is not determined solely by a GPU’s theoretical compute capability. No matter how many compute devices a system has, it cannot deliver its full performance if the required data does not arrive in time. In other words, what matters is not only how quickly a GPU can compute — the required data must also be located close to the compute device and accessed without delay for the GPU’s performance to be fully realized.
This distinction has become more pronounced with the proliferation of inference-centric AI and agentic AI. During large-scale training, compute devices could be kept relatively well utilized by processing large amounts of data together in batches*. In contrast, real-world inference environments require low latency, long contexts, and repeated calls. Models repeatedly reference required information, store intermediate results, and continuously retrieve data to generate the next token*.
* Batch: A unit in which an AI model processes multiple pieces of data together. Large batches can help utilize GPUs efficiently during large-scale training but are more difficult to use for inference in real-world services because of low-latency requirements and the need to process individual requests
* Token: The basic unit of text processed by an AI model. An LLM generates each new token sequentially by referring to previous tokens; as the token count grows, so does inference time and memory usage.
As a result, the focus of the bottleneck is shifting from GPU compute capability to the location and movement path of data. The key challenge is how to overcome the limits of memory bandwidth and access latency — the memory wall*.
* Memory wall: A bottleneck that occurs when improvements in memory data delivery speed fail to keep pace with advances in processor performance
In 1994, U.S. computer scientists William A. Wulf and Sally A. McKee highlighted the growing gap between processor performance and memory access speed in a short paper*. They warned that while processor performance was improving by around 60% each year, DRAM access speed was improving by only about 7%. If that gap continued to widen, system performance could ultimately become constrained by memory speed. This bottleneck became known as the memory wall.
* Wm. A. Wulf and Sally A. McKee, “Hitting the Memory Wall: Implications of the Obvious.” ACM SIGARCH Computer Architecture News, Vol. 23, No. 1 (1995): 20-24
Thirty years later, that warning is becoming a reality in AI service environments where large models with tens or hundreds of billions of parameters operate in real time. While the compute performance of the latest AI accelerators processing these models has reached the petaflops* range, the speed at which model weights and activations can be retrieved from memory has not kept pace with advances in compute. As a result, many of the compute units integrated into large chips spend more time waiting for data, reducing their actual utilization.
There are two primary approaches to narrowing this gap: increasing memory bandwidth itself or placing memory closer to the processor to reduce the distance data must travel.
* Petaflops(PFLOPS): A unit of compute performance equal to one quadrillion floating-point operations per second
As large language models* and reasoning-intensive workloads become widespread, memory bottlenecks have become more complex. Unlike earlier, one-off inference tasks, today’s models generate thousands — sometimes tens of thousands — of tokens sequentially in response to a single query. Step-by-step reasoning approaches such as chain-of-thought*and agentic AI can push token counts even higher.
* Large language model(LLM): An AI model trained on large volumes of text data to understand context and generate responses. As sequences grow longer, the model must continually refer back to information from previous tokens, requiring substantial memory and data movement during inference.
* Chain-of-thought(CoT): A prompting and generation technique in which a model works through intermediate reasoning steps before producing an answer, helping improve reasoning accuracy
One of the biggest challenges in this process is the KV cache*. Each time a Transformer*-based model generates a token, it retains key information from all preceding tokens in memory as key and value vectors. As the sequence grows, so does the cache. In some long-context inference environments, the memory required of the KV cache can even exceed the roughly 140 GB required to hold the weight of a 70-billion-parameter model.
* KV cache: A technique that stores and reuses previously generated key and value vectors, avoiding the inefficiency of repeatedly performing the same computations
* Transformer: A neural network architecture introduced by Google in 2017. It uses self-attention to understand context and forms the foundation of nearly all modern LLMs.
Put another way, when a user sends a query to an LLM service such as ChatGPT, more than half of the GPU’s work can involve moving data rather than computing it. This is a structural challenge that cannot be solved simply by adding more GPUs.
To ease these bottlenecks, the industry has adopted a range of software-based solutions. One of the most prominent is FlashAttention*. This technique restructures attention*, a core operation in LLMs, by dividing it into smaller blocks and processing as much as possible within the GPU’s internal SRAM*. This reduces repeated reads and writes of intermediate data between HBM and SRAM, improving inference efficiency on the same hardware. Techniques such as PagedAttention* and speculative decoding*, which optimize memory usage patterns and token generation, have also quickly gained traction.
* FlashAttention: An algorithm that divides attention operations into smaller blocks and processes them in the GPU’s fast on-chip SRAM, reducing data transfers to and from external memory
* Attention: The process by which an AI model refers to preceding context when generating the next word
* Static random-access memory(SRAM): High-speed memory used for caches inside CPUs and GPUs. It is faster than DRAM but requires more circuit area, making it unsuitable for high-capacity storage. In AI compute, SRAM is used to keep frequently accessed intermediate data close to the processor
* PagedAttention: A technique that manages GPU memory in pages, similar to virtual memory, during attention operations. By reducing unnecessary memory waste, it enables more user requests to be processed simultaneously and improves inference efficiency
* Speculative decoding: An inference acceleration technique in which a smaller auxiliary model predicts multiple tokens in advance and the main model verifies them together. It is used to improve response speed while maintaining generation quality.
Another major trend is quantization*. This method represents and stores model weights, activation values, or the KV cache using fewer bits, such as 8 or 4 bits instead of 16. With less data to represent, both memory usage and data movement decrease. Recently, various compression techniques have emerged to improve data representation efficiency, such as data formats that group similar values together to enhance representation, and vector quantization*, which replaces frequently occurring patterns with shorter codes.
* Vector quantization(VQ): A technique that represents model weights, activations, KV caches, and other data using fewer bits to reduce memory usage and data movement. For example, representing 16-bit data with 8 or 4 bits can reduce both storage requirements and data transfer overhead
One recent development attracting attention in the field is TurboQuant*, a quantization technique that represents the KV cache with fewer bits to reduce its memory footprint. Conventional compression techniques have had to recalibrate their compression rules whenever the distribution of input data changes. While this technique offers high accuracy, it also increases processing time.
* TurboQuant: A KV cache compression technique presented by Google at ICLR 2026. It uses a rotation transformation to map input data to a standardized distribution before applying predefined compression rules, reducing 16-bit data to around 3-4 bits without additional training while largely preserving model accuracy.
TurboQuant takes a different approach. It first applies a mathematical operation known as a rotation to transform different inputs into similar distributions, then applies a set of predefined rules. The idea is similar to standardizing customers’ body shapes before fitting them with the same size of clothing, rather than tailoring a new outfit for every customer. According to the research, TurboQuant reduces a KV cache stored at 16 bits to around 3-4 bits, lowering memory usage to roughly one-fifth without significantly compromising response quality.
This TurboQuant technique pushes compression efficiency — representing the same data with fewer bits — close to its theoretical limit. In other words, there may be limited room to compress the data much further without information loss.
There is a key implication here. As software compression approaches its theoretical limits, the potential for further improvement within the same memory architecture also diminishes. Software optimization is clearly effective, but fundamental limitations remain if compute devices must continue fetching data from memory located farther away.
So what comes next? The starting point is to bring memory and compute closer together, creating architectures that process the data they need over shorter paths. This is not simply a matter of finding better algorithms. It is a problem that requires redesigning the architecture of systems and semiconductors themselves.
Several solutions have already emerged. One prominent example is High Bandwidth Memory*. HBM stacks multiple DRAM dies vertically and places them close to the GPU, shortening the distance data must travel to reach the compute device while increasing the amount of data that can be transferred at once. HBM is drawing attention not simply because it increases memory speed, but also because it reduces the time compute devices spend waiting for data.
* High Bandwidth Memory: Memory that stacks multiple DRAM dies and connects them using through-silicon via(TSV) technology to provide very high bandwidth to GPUs and AI accelerators.
This same critical perspective is driving other memory and system technologies. High Bandwidth Flash* seeks to secure both greater capacity and higher bandwidth, while Compute Express Link* reflects a shift toward turning memory from a dedicated resource for a specific GPU into a shared resource at the system level. Processing-in-memor* goes a step further, processing some compute closer to the memory where data resides. Although these three technologies differ in their approaches, they all stem from the same question: Should data continue traveling long distances for processing, or should more work be processed closer to where the data is located?
* High Bandwidth Flash(HBF): A next-generation NAND-based memory technology being explored as a new memory layer between HBM and SSDs. It aims to provide significantly greater capacity than HBM while targeting comparable bandwidth and greater cost efficiency.
* Compute Express Link(CXL): A next-generation interconnect standard that connects CPUs, GPUs, memory, and accelerators through a coherent cache protocol, enabling memory to be disaggregated and shared at the system level
* Processing-in-memory(PIM): A next-generation memory technology that adds processing capabilities to memory, helping reduce data bottlenecks between memory and processors and significantly improve performance
Ultimately, these approaches point in the same direction. Data that can be kept close to the compute device should stay close, while data that must travel requires wider, more efficient paths. The AI bottleneck is no longer a problem that can be addressed at the level of a single chip. It is a system-level challenge involving the complex connections between model architecture, GPUs, memory, interconnects, packaging, and data center architecture.
This does not mean that conventional von Neumann architecture*, which separates compute devices from memory, is about to disappear. It does mean, however, that an approach taken for granted for decades — keeping processors and memory separate and connecting them through a bus — must now be reexamined. AI system performance is determined not only by the compute device itself, but also by where memory is located and how data is read.
* Von Neumann architecture: A stored-program computer architecture in which programs and data are held in the same memory and processed sequentially by the CPU. It forms the foundation of modern computer architecture
Let us return to the initial question: What is the real bottleneck in AI? It lies in a place that may not be immediately visible — in an architecture where data constantly moves back and forth between memory and compute devices.
The software-based compression techniques discussed earlier can alleviate some of this bottleneck, but a fundamental question remains. Why must information travel such a long path every time? Can more processing be done where the data resides? Can the gap between memory and compute be narrowed?
How effectively this bottleneck is addressed will determine the competitiveness of next-generation AI infrastructure. This is not a challenge limited to any single country or company. It is a challenge for the global AI ecosystem, requiring AI accelerators, memory, advanced packaging, interconnects, and software to work together. Within this context, memory technology capabilities have become more critical than ever.
This is ultimately where the next phase of competition in AI infrastructure will be decided. Simply securing more compute devices will not be enough. The advantage will go to those that can co-design memory layouts and compute structures to process data faster and more efficiently. Understanding this shift and translating it into new architectures is one of the most important tasks facing semiconductor researchers today.
The bottleneck examined in Part 2 is not simply a matter of memory speed. As inference and agentic AI become more widespread, models must maintain long contexts, repeatedly reference intermediate results, and continuously exchange the necessary data between memory and compute devices. As a result, AI performance increasingly depends less on the compute capability of an individual GPU and more on where data is stored, how it moves through the system, and how efficiently it is processed.
This challenge is difficult to solve with an individual chip or memory technology alone. GPUs, memory, storage, networks, and data centers must work together as a single system. In the next installment, we examine this challenge from an AI infrastructure perspective. What ultimately determines AI performance: faster components or the overall architecture through which compute and data flow?
_**Disclaimer:**The opinions expressed in this article are solely those of the author and do not necessarily reflect the official position of SK hynix._
This article series explores the transformation of the AI ecosystem, where software, data center infrastructure, semiconductors, and memory technologies work together. It also examines why memory semiconductors have become a foundational technology for the AI era and outlines SK hynix’s vision for building the next-generation AI memory ecosystem.
① The paradigm shift in AI computing
**② The real bottleneck: Data, not compute – Professor Hoi-Jun Yoo, KAIST**
③ Redesigning infrastructure: Why architecture determines performance
④ The semiconductor paradigm shifts toward memory
⑤ Completing AI: Physical intelligence and the role of memory
As the center of gravity of AI shifts from pretraining to inference and from one-off responses to agentic AI, the location of system bottlenecks has also changed. The key issue increasingly lies in the flow of data rather than the volume of compute itself. As models maintain longer contexts and repeatedly reference intermediate results, the frequency of data exchanges between memory and compute devices increases.
The question, therefore, is no longer limited to GPU compute capability. Is AI performance really constrained by how much a GPU can compute or by the memory architecture’s ability to retrieve the required data and deliver it to compute devices when needed? In the AI era, the core bottleneck lies not in compute itself, but in where data resides and how it moves through the system.
For the past decade, the formula for improving AI performance seemed clear: compute-centric scaling* that prioritizes more GPUs, larger models, and longer training runs. This formula held true during the era when large-scale pretraining was the primary path to improving AI performance. Deploying more GPUs made it possible to train larger models for longer periods, which in turn improved both model accuracy and versatility.
* Scaling: An approach to improving AI performance by increasing compute resources such as GPUs, model size, and the amount of training data
As a result, GPUs became a symbol of performance in the AI era. Companies competed to highlight how many GPUs they had secured, the size of their clusters, and how many operations they could perform per second. These metrics clearly matter. Without the parallel compute capabilities of GPUs, today’s large-scale AI models would be difficult to achieve.
However, the real-world performance of an AI system is not determined solely by a GPU’s theoretical compute capability. No matter how many compute devices a system has, it cannot deliver its full performance if the required data does not arrive in time. In other words, what matters is not only how quickly a GPU can compute — the required data must also be located close to the compute device and accessed without delay for the GPU’s performance to be fully realized.
This distinction has become more pronounced with the proliferation of inference-centric AI and agentic AI. During large-scale training, compute devices could be kept relatively well utilized by processing large amounts of data together in batches*. In contrast, real-world inference environments require low latency, long contexts, and repeated calls. Models repeatedly reference required information, store intermediate results, and continuously retrieve data to generate the next token*.
* Batch: A unit in which an AI model processes multiple pieces of data together. Large batches can help utilize GPUs efficiently during large-scale training but are more difficult to use for inference in real-world services because of low-latency requirements and the need to process individual requests
* Token: The basic unit of text processed by an AI model. An LLM generates each new token sequentially by referring to previous tokens; as the token count grows, so does inference time and memory usage.
As a result, the focus of the bottleneck is shifting from GPU compute capability to the location and movement path of data. The key challenge is how to overcome the limits of memory bandwidth and access latency — the memory wall*.
* Memory wall: A bottleneck that occurs when improvements in memory data delivery speed fail to keep pace with advances in processor performance
In 1994, U.S. computer scientists William A. Wulf and Sally A. McKee highlighted the growing gap between processor performance and memory access speed in a short paper*. They warned that while processor performance was improving by around 60% each year, DRAM access speed was improving by only about 7%. If that gap continued to widen, system performance could ultimately become constrained by memory speed. This bottleneck became known as the memory wall.
* Wm. A. Wulf and Sally A. McKee, “Hitting the Memory Wall: Implications of the Obvious.” ACM SIGARCH Computer Architecture News, Vol. 23, No. 1 (1995): 20-24
Thirty years later, that warning is becoming a reality in AI service environments where large models with tens or hundreds of billions of parameters operate in real time. While the compute performance of the latest AI accelerators processing these models has reached the petaflops* range, the speed at which model weights and activations can be retrieved from memory has not kept pace with advances in compute. As a result, many of the compute units integrated into large chips spend more time waiting for data, reducing their actual utilization.
There are two primary approaches to narrowing this gap: increasing memory bandwidth itself or placing memory closer to the processor to reduce the distance data must travel.
* Petaflops(PFLOPS): A unit of compute performance equal to one quadrillion floating-point operations per second
As large language models* and reasoning-intensive workloads become widespread, memory bottlenecks have become more complex. Unlike earlier, one-off inference tasks, today’s models generate thousands — sometimes tens of thousands — of tokens sequentially in response to a single query. Step-by-step reasoning approaches such as chain-of-thought*and agentic AI can push token counts even higher.
* Large language model(LLM): An AI model trained on large volumes of text data to understand context and generate responses. As sequences grow longer, the model must continually refer back to information from previous tokens, requiring substantial memory and data movement during inference.
* Chain-of-thought(CoT): A prompting and generation technique in which a model works through intermediate reasoning steps before producing an answer, helping improve reasoning accuracy
One of the biggest challenges in this process is the KV cache*. Each time a Transformer*-based model generates a token, it retains key information from all preceding tokens in memory as key and value vectors. As the sequence grows, so does the cache. In some long-context inference environments, the memory required of the KV cache can even exceed the roughly 140 GB required to hold the weight of a 70-billion-parameter model.
* KV cache: A technique that stores and reuses previously generated key and value vectors, avoiding the inefficiency of repeatedly performing the same computations
* Transformer: A neural network architecture introduced by Google in 2017. It uses self-attention to understand context and forms the foundation of nearly all modern LLMs.
Put another way, when a user sends a query to an LLM service such as ChatGPT, more than half of the GPU’s work can involve moving data rather than computing it. This is a structural challenge that cannot be solved simply by adding more GPUs.
To ease these bottlenecks, the industry has adopted a range of software-based solutions. One of the most prominent is FlashAttention*. This technique restructures attention*, a core operation in LLMs, by dividing it into smaller blocks and processing as much as possible within the GPU’s internal SRAM*. This reduces repeated reads and writes of intermediate data between HBM and SRAM, improving inference efficiency on the same hardware. Techniques such as PagedAttention* and speculative decoding*, which optimize memory usage patterns and token generation, have also quickly gained traction.
* FlashAttention: An algorithm that divides attention operations into smaller blocks and processes them in the GPU’s fast on-chip SRAM, reducing data transfers to and from external memory
* Attention: The process by which an AI model refers to preceding context when generating the next word
* Static random-access memory(SRAM): High-speed memory used for caches inside CPUs and GPUs. It is faster than DRAM but requires more circuit area, making it unsuitable for high-capacity storage. In AI compute, SRAM is used to keep frequently accessed intermediate data close to the processor
* PagedAttention: A technique that manages GPU memory in pages, similar to virtual memory, during attention operations. By reducing unnecessary memory waste, it enables more user requests to be processed simultaneously and improves inference efficiency
* Speculative decoding: An inference acceleration technique in which a smaller auxiliary model predicts multiple tokens in advance and the main model verifies them together. It is used to improve response speed while maintaining generation quality.
Another major trend is quantization*. This method represents and stores model weights, activation values, or the KV cache using fewer bits, such as 8 or 4 bits instead of 16. With less data to represent, both memory usage and data movement decrease. Recently, various compression techniques have emerged to improve data representation efficiency, such as data formats that group similar values together to enhance representation, and vector quantization*, which replaces frequently occurring patterns with shorter codes.
* Vector quantization(VQ): A technique that represents model weights, activations, KV caches, and other data using fewer bits to reduce memory usage and data movement. For example, representing 16-bit data with 8 or 4 bits can reduce both storage requirements and data transfer overhead
One recent development attracting attention in the field is TurboQuant*, a quantization technique that represents the KV cache with fewer bits to reduce its memory footprint. Conventional compression techniques have had to recalibrate their compression rules whenever the distribution of input data changes. While this technique offers high accuracy, it also increases processing time.
* TurboQuant: A KV cache compression technique presented by Google at ICLR 2026. It uses a rotation transformation to map input data to a standardized distribution before applying predefined compression rules, reducing 16-bit data to around 3-4 bits without additional training while largely preserving model accuracy.
TurboQuant takes a different approach. It first applies a mathematical operation known as a rotation to transform different inputs into similar distributions, then applies a set of predefined rules. The idea is similar to standardizing customers’ body shapes before fitting them with the same size of clothing, rather than tailoring a new outfit for every customer. According to the research, TurboQuant reduces a KV cache stored at 16 bits to around 3-4 bits, lowering memory usage to roughly one-fifth without significantly compromising response quality.
This TurboQuant technique pushes compression efficiency — representing the same data with fewer bits — close to its theoretical limit. In other words, there may be limited room to compress the data much further without information loss.
There is a key implication here. As software compression approaches its theoretical limits, the potential for further improvement within the same memory architecture also diminishes. Software optimization is clearly effective, but fundamental limitations remain if compute devices must continue fetching data from memory located farther away.
So what comes next? The starting point is to bring memory and compute closer together, creating architectures that process the data they need over shorter paths. This is not simply a matter of finding better algorithms. It is a problem that requires redesigning the architecture of systems and semiconductors themselves.
Several solutions have already emerged. One prominent example is High Bandwidth Memory*. HBM stacks multiple DRAM dies vertically and places them close to the GPU, shortening the distance data must travel to reach the compute device while increasing the amount of data that can be transferred at once. HBM is drawing attention not simply because it increases memory speed, but also because it reduces the time compute devices spend waiting for data.
* High Bandwidth Memory: Memory that stacks multiple DRAM dies and connects them using through-silicon via(TSV) technology to provide very high bandwidth to GPUs and AI accelerators.
This same critical perspective is driving other memory and system technologies. High Bandwidth Flash* seeks to secure both greater capacity and higher bandwidth, while Compute Express Link* reflects a shift toward turning memory from a dedicated resource for a specific GPU into a shared resource at the system level. Processing-in-memor* goes a step further, processing some compute closer to the memory where data resides. Although these three technologies differ in their approaches, they all stem from the same question: Should data continue traveling long distances for processing, or should more work be processed closer to where the data is located?
* High Bandwidth Flash(HBF): A next-generation NAND-based memory technology being explored as a new memory layer between HBM and SSDs. It aims to provide significantly greater capacity than HBM while targeting comparable bandwidth and greater cost efficiency.
* Compute Express Link(CXL): A next-generation interconnect standard that connects CPUs, GPUs, memory, and accelerators through a coherent cache protocol, enabling memory to be disaggregated and shared at the system level
* Processing-in-memory(PIM): A next-generation memory technology that adds processing capabilities to memory, helping reduce data bottlenecks between memory and processors and significantly improve performance
Ultimately, these approaches point in the same direction. Data that can be kept close to the compute device should stay close, while data that must travel requires wider, more efficient paths. The AI bottleneck is no longer a problem that can be addressed at the level of a single chip. It is a system-level challenge involving the complex connections between model architecture, GPUs, memory, interconnects, packaging, and data center architecture.
This does not mean that conventional von Neumann architecture*, which separates compute devices from memory, is about to disappear. It does mean, however, that an approach taken for granted for decades — keeping processors and memory separate and connecting them through a bus — must now be reexamined. AI system performance is determined not only by the compute device itself, but also by where memory is located and how data is read.
* Von Neumann architecture: A stored-program computer architecture in which programs and data are held in the same memory and processed sequentially by the CPU. It forms the foundation of modern computer architecture
Let us return to the initial question: What is the real bottleneck in AI? It lies in a place that may not be immediately visible — in an architecture where data constantly moves back and forth between memory and compute devices.
The software-based compression techniques discussed earlier can alleviate some of this bottleneck, but a fundamental question remains. Why must information travel such a long path every time? Can more processing be done where the data resides? Can the gap between memory and compute be narrowed?
How effectively this bottleneck is addressed will determine the competitiveness of next-generation AI infrastructure. This is not a challenge limited to any single country or company. It is a challenge for the global AI ecosystem, requiring AI accelerators, memory, advanced packaging, interconnects, and software to work together. Within this context, memory technology capabilities have become more critical than ever.
This is ultimately where the next phase of competition in AI infrastructure will be decided. Simply securing more compute devices will not be enough. The advantage will go to those that can co-design memory layouts and compute structures to process data faster and more efficiently. Understanding this shift and translating it into new architectures is one of the most important tasks facing semiconductor researchers today.
The bottleneck examined in Part 2 is not simply a matter of memory speed. As inference and agentic AI become more widespread, models must maintain long contexts, repeatedly reference intermediate results, and continuously exchange the necessary data between memory and compute devices. As a result, AI performance increasingly depends less on the compute capability of an individual GPU and more on where data is stored, how it moves through the system, and how efficiently it is processed.
This challenge is difficult to solve with an individual chip or memory technology alone. GPUs, memory, storage, networks, and data centers must work together as a single system. In the next installment, we examine this challenge from an AI infrastructure perspective. What ultimately determines AI performance: faster components or the overall architecture through which compute and data flow?
_**Disclaimer:**The opinions expressed in this article are solely those of the author and do not necessarily reflect the official position of SK hynix._