MEMORY INDUSTRY INTELLIGENCE

훈련에서 추론으로: 메모리 계층 구조의 패러다임 전환

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.11 17:27사용자 검토 전 초안

원문 제목: From Training to Inference: A Paradigm Shift in the Memory Hierarchy

핵심 요약

조사기관 전망·해석 문서로, AI 산업의 초점이 훈련에서 추론으로 이동하면서 메모리 계층 구조가 구조적 변화를 겪고 있다고 분석한다. 2026년 상반기 NVIDIA, Google 등 주요 AI 기업이 추론 전용 칩을 발표했고, Cerebras·SambaNova·Etched 등 AI ASIC 스타트업도 신제품과 전략적 파트너십을 발표했다. 훈련 단계는 HBM이 지배했으나, 추론 단계에서는 KV Cache 용량 병목을 해결하기 위해 HBF, SSD POD, SRAM 등 다양한 접근이 실험되고 있다. SanDisk와 SK하이닉스는 2025년 8월 HBF 공동 개발 파트너십을 발표했으며, HBF는 아직 양산 단계에 도달하지 못했다. NVIDIA는 2026년 SSD POD 개념을 도입했고, Groq·Cerebras 등은 온칩 SRAM을 활용해 디코드 속도를 높이고 있다.

메모리 산업 영향 분석

이 문서는 조사기관의 전망·해석으로, AI 산업이 훈련에서 추론으로 이동하면서 메모리 계층 구조가 HBM 중심에서 HBF·SSD POD·SRAM 등으로 다양화될 것이라는 분석을 제시한다. 메모리 산업에 대한 직접적 영향은 추론용 KV Cache 용량 병목이 HBM 단독으로 해결되지 않아 HBF(NAND 기반)와 SSD POD(로컬 SSD와 공유 스토리지 사이 계층) 같은 새로운 메모리 계층이 등장한다는 점이다. 이는 HBM 수요 증가와 동시에 NAND 기반 HBF·eSSD 수요를 촉진할 수 있으나, HBF가 아직 양산 전이고 SSD POD는 개념 단계이므로 실제 공급 배분·고객 인증·제품 믹스·투자 일정에 미치는 영향은 미확인이다. 추론 ASIC 스타트업의 온칩 SRAM 사용은 HBM 수요를 감소시킬 수 있으나, 동시에 추론 수요 증가로 인한 HBM 사용량 증가 효과와 상충할 수 있다. NVIDIA의 Groq LPX 통합, Google TPU v8i의 HBM3e 72 GB 추가 등은 HBM 수요에 긍정적이나, 이는 원문에서 확인된 계획·발표이며 실제 양산·출하량은 미확인이다. 반대 근거로는 HBF의 양산 지연, SSD POD의 실제 채택 불확실성, SRAM의 용량 한계 등이 있다. 확인할 지표로는 HBF 양산 일정, SSD POD 채택 고객, 추론 ASIC의 HBM 탑재량, NAND 공급 동향 등이 있다. 공급 배분 측면에서는 HBM과 NAND 생산 능력 배분 선택이, 고객 인증 측면에서는 HBF·SSD POD의 고객 채택 여부가, 제품 믹스 측면에서는 HBM·HBF·eSSD 비중 변화가, 투자 일정 측면에서는 HBF 개발 및 양산 투자 시점이 영향을 받을 수 있다. 다만 원문은 조사기관 전망·해석이므로 실제 기업의 공식 확인·납품·가동과 구분해야 한다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 8556자

테스트 시점 스케일링(test-time scaling)이 자리 잡으면서 AI 산업의 초점은 빠르게 훈련에서 추론으로 이동하고 있다. NVIDIA, Google 및 기타 주요 AI 플레이어들은 2026년 상반기에 추론 전용 칩을 출시했으며, Cerebras, SambaNova, Etched와 같은 AI ASIC 스타트업들도 신제품과 전략적 파트너십을 발표했다. 이러한 전환과 함께 메모리 계층 구조도 구조적 변화를 겪고 있으며, 벤더들은 기존 HBM의 용량 및 대역폭 한계를 넘어서기 위해 HBF, SSD POD, SRAM으로 눈을 돌리고 있다.

이 글은 AI 메모리 요구사항에 대한 명확하고 기초적인 이해를 구축하려는 독자를 위해 작성되었다.

**목차:**

1. AI 훈련을 위한 메모리 요구사항

2. AI 추론을 위한 메모리 요구사항

3. 산업이 훈련에서 추론으로 이동함에 따라 메모리 구조가 어떻게 변하는가

AI 훈련은 훈련 데이터셋을 사용하여 AI 모델에 새로운 능력을 가르치는 과정을 말한다. AI 추론은 이미 훈련된 모델을 새로운 데이터에 적용하여 응답을 생성하는 과정을 말한다. 이 두 단계는 매우 다른 유형의 메모리를 필요로 하기 때문에, 최근 몇 년간 메모리 아키텍처가 빠르게 다양화되었다.

[](https://substackcdn.com/image/fetch/$s_!paa_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364581a8-9dbf-46ad-af2d-4f3031ad08de_1012x451.png)

AI 훈련 및 AI 추론 워크플로 다이어그램. 출처: NVIDIA

오늘날 지배적인 Transformer 아키텍처를 예로 들면, 훈련 과정은 세 가지 주요 단계로 구성된다:

1. **순방향 패스(Forward Pass):** 각 레이어의 가중치가 메모리에서 컴퓨트 다이로 읽혀지고, 대규모 행렬 곱셈이 예측 출력을 생성한다. 그런 다음 손실 함수가 이 출력을 정답과 비교하여 오차(손실)를 측정한다.

2. **역방향 패스(Backward Pass):** 연쇄 법칙에 따라, 모델은 순방향 패스 중 임시로 저장된 활성화 값을 레이어별로 읽고, 가중치 크기에 맞는 기울기를 계산한다.

3. **옵티마이저 업데이트(Optimizer Update):** 기울기를 읽은 후, 모델은 가중치 자체와 가중치와 동일한 크기의 두 옵티마이저 상태 텐서(1차 모멘트 m 및 2차 모멘트 v)를 업데이트하고 다시 쓴다.

모델이 손실을 최소화하고 수렴하기 위해 이 세 단계를 반복해야 하므로, 메모리 접근 속도가 훈련 효율을 결정하는 핵심 요소가 된다.

[](https://substackcdn.com/image/fetch/$s_!cAfc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a085fbe-e57c-4618-877b-893eeaad8f7c_1080x560.jpeg)

이러한 맥락에서 HBM(High Bandwidth Memory)이 등장했다: 여러 층의 DRAM을 수직으로 적층하고 TSV로 연결한 것이다. 상대적으로 느린 PCIe(PCIe 6.0 x16은 단방향 128 GB/s 대역폭 제공)로 연결되는 기존 DRAM과 달리, HBM은 컴퓨트 다이와 나란히 동일 인터포저에 패키징되어 데이터 경로를 단축한다. 단일 HBM 스택은 2 TB/s 이상의 대역폭을 제공할 수 있으며(HBM4는 2,048핀에 걸쳐 단방향으로 핀당 1 GB/s 이상 제공), 이는 AI의 산업화를 크게 가속화했다.

[](https://substackcdn.com/image/fetch/$s_!iEN4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe555eeb7-8ae3-4aa8-a6cd-e8441389977d_1080x560.jpeg)

메모리 계층 구조는 훈련 시대에 HBM이라는 새로운 계층을 얻었으며, 온칩 SRAM과 기존 DRAM 사이의 격차를 메웠다.

> _**관련 보고서: [3Q26 HBM 데이터시트](https://www.trendforce.com/research/download/RP240710GF?utm\_source=tf\_substack&utm\_medium=post)**_

[](https://substackcdn.com/image/fetch/$s_!YQ5h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0a1bd87-cdd6-4914-a709-57ea06017fa7_1080x560.jpeg)

AI 추론은 주로 두 단계로 구성된다:

1. **프리필(Prefill)**: 입력 프롬프트가 단일 패스에서 병렬로 처리된다. 각 레이어의 Key 및 Value 벡터가 계산되어 메모리(KV Cache)에 저장되고, 동시에 첫 번째 출력 토큰이 생성된다.

2. **디코드(Decode)**: Key 및 Value 벡터가 KV Cache에서 읽혀져 토큰을 하나씩 생성하는 동시에, 새로 계산된 Key 및 Value 벡터가 응답이 완료될 때까지 지속적으로 KV Cache에 다시 쓰여진다.

**디코드** 단계는 KV Cache에 대한 토큰별 반복 읽기 및 쓰기를 필요로 하고, KV Cache 크기가 대화 턴 수, 컨텍스트 윈도우 길이, 배치 크기에 따라 계속 증가하기 때문에, HBM만으로는 증가하는 추론 수요를 따라잡기 점점 어려워지고 있다. 메모리 용량이 대역폭과 함께 두 번째 핵심 병목이 되었다.

[](https://substackcdn.com/image/fetch/$s_!9wQd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F194332e3-8ff9-4af5-87ca-2027e2a47561_1080x560.jpeg)

AI 훈련을 위해 구축된 상대적으로 성숙한 메모리 아키텍처와 달리, AI 추론 메모리 아키텍처는 이제 막 테스트 시점 스케일링과 함께 빠르게 발전하고 있다. 벤더들은 여전히 다양한 메모리 접근 방식을 적극적으로 실험하고 있으며, 메모리 계층 구조는 패러다임 전환의 중간에 있다.

KV Cache의 증가하는 용량 요구를 해결하기 위해, 메모리 제조사 SanDisk와 SK하이닉스는 2025년 8월에 HBF(High Bandwidth Flash)를 개발하기 위한 파트너십을 발표했으며, 이는 여러 층의 NAND를 수직으로 적층하고 TSV로 연결한다. 예상 대역폭 1.6 TB/s에서 단일 HBF 스택은 약 512 GB의 용량을 제공할 것으로 예상되며, 이는 HBM(16-Hi HBM4 스택은 48 GB 제공)의 약 10배 이상으로 HBM과 기존 DRAM 사이의 격차를 메운다.

HBF는 아직 양산 단계에 도달하지 못했지만, 많은 벤더들이 동시에 KV Cache 오프로딩을 사용하여 덜 자주 사용되는 KV Cache를 하위 메모리 계층으로 이동시켜 대량의 KV Cache 저장 부담을 완화하고 있다. 이러한 맥락에서 NVIDIA는 2026년에 SSD POD 개념을 도입하여 로컬 SSD와 공유 스토리지 사이의 격차를 메웠다.

특히, HBF와 SSD POD 외에도 일부 추론 ASIC 스타트업들은 HBM의 속도 한계를 넘어서기 위해 고속 SRAM을 도입했다. 예를 들어 Groq(추론 기술 라이선스와 핵심 팀이 2025년 12월 NVIDIA에 확보됨)와 Cerebras는 대면적 온칩 SRAM을 사용하여 디코드 속도를 크게 높인다.

> _**관련 보고서: [2026 트렌드: 새로운 AI 추론 수요를 위한 메모리](https://www.trendforce.com/research/download/RP260612EF?utm\_source=tf\_substack&utm\_medium=post)**_

[](https://substackcdn.com/image/fetch/$s_!gg74!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b302e8b-8793-4e14-ab18-1d4486f333bc_1080x560.jpeg)

2026년 상반기, 추론 수요가 계속 급증하면서 NVIDIA는 Groq LPX 칩을 Rubin 시리즈에 통합할 것이라고 발표했으며, 올해 말 출시될 예정이고, CMX(Context Memory Storage Platform)와 Vera CPU Rack도 도입할 예정이다.

Google도 추론 전용으로 설계된 TPU v8i를 공개했으며, 이는 훈련용으로 설계된 TPU v8t보다 72 GB 더 많은 HBM3e와 384 MB 더 많은 SRAM을 탑재한다.

이 두 주요 플레이어 외에도, 추론 전용 칩을 만드는 여러 AI ASIC 스타트업들이 2026년 상반기에 활발히 활동했다: SambaNova는 2026년 2월에 5세대 칩 SN50을 출시했고; Cerebras는 2026년 4월에 IPO를 신청하고 OpenAI 및 AWS와의 파트너십을 발표했으며; Etched는 2026년 6월에 Sohu 칩이 TSMC에서 성공적으로 테이프아웃되었다고 발표했다.

훈련 시대가 거의 전적으로 HBM에 의해 지배된 것과 비교하여, 오늘날 훨씬 더 다양한 추론 환경을 보면, 메모리 계층 구조의 설계 선택이 앞으로도 계속 다양화될 것으로 예상할 수 있다.

[](https://substackcdn.com/image/fetch/$s_!_Upw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e470b7-e07d-4007-ae9b-023c976c7b4e_1080x560.jpeg)

[](https://substackcdn.com/image/fetch/$s_!DxEt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e2bda23-c931-4f49-a5f3-f5bfabcd3420_1080x560.jpeg)

> _다음 글에서는 업계 리더들이 메모리 병목을 해결하기 위해 CXL 확장과 KV Cache 압축으로 눈을 돌리는 방법과 이것이 향후 메모리 시장에 미치는 영향을 자세히 살펴본다. 구독하여 놓치지 마세요._
브리프용 요약 초안
AI 산업이 훈련에서 추론으로 이동하면서 메모리 계층 구조가 HBM 중심에서 HBF·SSD POD·SRAM 등으로 다양화되고 있다. SanDisk와 SK하이닉스는 2025년 8월 HBF 공동 개발을 발표했으나 아직 양산 전이며, NVIDIA는 2026년 SSD POD 개념을 도입했다. 추론 ASIC 스타트업들의 온칩 SRAM 사용은 HBM 수요에 상충 효과를 줄 수 있어 실제 채택과 양산 일정 확인이 필요하다.

원문 텍스트

원문 열기 ↗
As test-time scaling takes hold, the focus of the AI industry is quickly shifting from training to inference. NVIDIA, Google, and other major AI players rolled out inference-specific chips in the first half of 2026, while AI ASIC startups such as Cerebras, SambaNova, and Etched have announced new products and strategic partnerships. Alongside this shift, the memory hierarchy is undergoing structural change, with vendors turning to HBF, SSD POD, and SRAM to push past the capacity and bandwidth limits of conventional HBM.

This piece is written for readers looking to build a clear, foundational understanding of AI memory requirements.

**Contents:**

1. Memory requirements for AI training

2. Memory requirements for AI inference

3. How the memory structure changes as the industry moves from training to inference

AI training refers to the process of using a training dataset to teach an AI model new capabilities. AI inference refers to the process of applying an already-trained model to new data to generate a response. Because these two stages require very different types of memory, memory architecture has diversified rapidly in recent years.

[](https://substackcdn.com/image/fetch/$s_!paa_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364581a8-9dbf-46ad-af2d-4f3031ad08de_1012x451.png)

Diagram of AI Training and AI Inference Workflow. Source: NVIDIA

Taking today’s dominant Transformer architecture as an example, the training process consists of three main stages:

1. **Forward Pass:** The weights of each layer are read from memory into the compute die, where large-scale matrix multiplication produces a predicted output. A loss function then compares this output against the correct answer to measure the error (loss).

2. **Backward Pass:** Following the chain rule, the model reads the activations temporarily stored during the forward pass, layer by layer, and computes gradients that match the size of the weights.

3. **Optimizer Update**: After reading the gradients, the model updates and writes back the weights themselves, along with two optimizer state tensors of the same size as the weights (first-moment m and second-moment v).

Because the model must repeat these three stages over and over to minimize loss and converge, memory access speed becomes the key factor determining training efficiency.

[](https://substackcdn.com/image/fetch/$s_!cAfc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a085fbe-e57c-4618-877b-893eeaad8f7c_1080x560.jpeg)

This is the context in which HBM (High Bandwidth Memory)emerged: multiple layers of DRAM stacked vertically and connected via TSV. Unlike conventional DRAM, which connects over relatively slow PCIe (PCIe 6.0 x16 offers 128 GB/s of unidirectional bandwidth), HBM is packaged side by side with the compute die on the same interposer, shortening the data path. A single HBM stack can deliver more than 2 TB/s of bandwidth (HBM4 offers over 1 GB/s per pin in one direction across 2,048 pins), which has significantly accelerated the industrialization of AI.

[](https://substackcdn.com/image/fetch/$s_!iEN4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe555eeb7-8ae3-4aa8-a6cd-e8441389977d_1080x560.jpeg)

The memory hierarchy gained a new tier, HBM, during the training era, filling the gap between on-chip SRAM and conventional DRAM.

> _**Related report: [3Q26 HBM Datasheet](https://www.trendforce.com/research/download/RP240710GF?utm\_source=tf\_substack&utm\_medium=post)**_

[](https://substackcdn.com/image/fetch/$s_!YQ5h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff0a1bd87-cdd6-4914-a709-57ea06017fa7_1080x560.jpeg)

AI inference consists mainly of two stages:

1. **Prefill**: The input prompt is processed in parallel in a single pass. The Key and Value vectors for each layer are computed and stored in memory (the KV Cache), and the first output token is generated at the same time.

2. **Decode**: Key and Value vectors are read from the KV Cache to generate tokens one at a time, while newly computed Key and Value vectors are continuously written back into it, until the response is complete.

Because the **Decode** stage requires repeated token-by-token reads and writes to the KV Cache, and because KV Cache size keeps growing with the number of conversation turns, context window length, and batch size, relying on HBM alone is increasingly insufficient to keep up with growing inference demand. Memory capacity has become a second critical bottleneck, alongside bandwidth.

[](https://substackcdn.com/image/fetch/$s_!9wQd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F194332e3-8ff9-4af5-87ca-2027e2a47561_1080x560.jpeg)

Unlike the relatively mature memory architecture built for AI training, AI inference memory architecture is only now developing rapidly alongside test-time scaling. Vendors are still actively experimenting with different memory approaches, and the memory hierarchy is in the middle of a paradigm shift.

To address the growing capacity needs of KV Cache, memory makers SanDisk and SK hynix announced a partnership in August 2025 to develop HBF (High Bandwidth Flash), which stacks multiple layers of NAND vertically and connects them via TSV. At an expected bandwidth of 1.6 TB/s, a single HBF stack is projected to offer around 512 GB of capacity, roughly ten times or more that of HBM (a 16-Hi HBM4 stack offers 48 GB), filling the gap between HBM and conventional DRAM.

While HBF has yet to reach mass production, many vendors are concurrently using KV Cache offloading to move less frequently used KV Cache down to lower tiers of memory, easing the burden of storing large volumes of KV Cache. In this context, NVIDIA introduced the concept of SSD POD in 2026, filling the gap between local SSD and shared storage.

Notably, beyond HBF and SSD POD, some inference ASIC startups have introduced high-speed SRAM to break past HBM’s speed limits. Groq (whose inference technology license and core team were secured by NVIDIA in December 2025) and Cerebras, for example, use large areas of on-chip SRAM to significantly increase decode speed.

> _**Related report: [2026 Trends: Memory for New AI Inference Demand](https://www.trendforce.com/research/download/RP260612EF?utm\_source=tf\_substack&utm\_medium=post)**_

[](https://substackcdn.com/image/fetch/$s_!gg74!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b302e8b-8793-4e14-ab18-1d4486f333bc_1080x560.jpeg)

In the first half of 2026, as inference demand continued to surge, NVIDIA announced it will incorporate the Groq LPX chip into its Rubin series, expected to be available later this year, alongside the introduction of CMX (Context Memory Storage Platform) and the Vera CPU Rack.

Google also unveiled the TPU v8i, designed specifically for inference, which carries 72 GB more HBM3e and 384 MB more SRAM than the TPU v8t, which is designed for training.

Beyond these two major players, a number of AI ASIC startups building inference-specific chips have been active in the first half of 2026: SambaNova launched its fifth-generation chip, the SN50, in February 2026; Cerebras filed for an IPO in April 2026 and announced partnerships with OpenAI and AWS; and Etched announced in June 2026 that its Sohu chip had successfully taped out at TSMC.

Looking at how the training era was dominated almost entirely by HBM, compared with the far more varied inference landscape today, the design choices for the memory hierarchy can be expected to keep diversifying going forward.

[](https://substackcdn.com/image/fetch/$s_!_Upw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e470b7-e07d-4007-ae9b-023c976c7b4e_1080x560.jpeg)

[](https://substackcdn.com/image/fetch/$s_!DxEt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e2bda23-c931-4f49-a5f3-f5bfabcd3420_1080x560.jpeg)

> _In our next piece, we take a closer look at how industry leaders are turning to CXL expansion and KV Cache compression to address the memory bottleneck, along with what this means for the memory market going forward. Subscribe so you don’t miss it._