MEMORY INDUSTRY INTELLIGENCE

AI 메모리 장벽 극복: CXL 메모리 풀링이 확장 가능한 AI 컴퓨팅의 차세대 도약을 가능하게 하는 방법 - Compute Express Link

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.12 02:47사용자 검토 전 초안

원문 제목: Overcoming the AI Memory Wall: How CXL Memory Pooling Powers the Next Leap in Scalable AI Computing - Compute Express Link

원문 문장이 일치하지 않은 주장 1건은 근거 등록에서 제외했습니다.

핵심 요약

XConn Technologies는 SC25의 CXL 파빌리온(부스 #817)에서 CXL 메모리 풀링 라이브 데모를 선보일 계획이다. 이 데모는 각각 NVIDIA H100 GPU(80GB)를 탑재한 두 서버에서 OPT-6.7B 모델을 64개 프롬프트, 프롬프트당 512토큰으로 실행하며, 프리필과 디코드 단계를 분리해 공유 CXL 메모리가 AI 추론을 가속하는 방식을 보여준다. XConn은 RDMA 기반 공유 대비 200G RDMA에서 3.8배, 100G RDMA에서 6.5배 속도 향상을 달성했다고 주장하며, TTFT 감소와 대역폭 효율 개선을 강조한다. 또한 PCIe GPU가 CXL을 네이티브 지원하지 않아도 XConn의 하이브리드 스위치와 Ultra IO Transformer를 통해 CXL 메모리 풀에 직접 접근할 수 있다고 설명한다. 실제 적용 사례로 AI 추론·KV 캐시 확장, PNNL Crete와 같은 과학·HPC 워크로드, 클라우드 데이터베이스를 제시한다.

메모리 산업 영향 분석

이 원문은 XConn Technologies가 SC25에서 CXL 메모리 풀링 데모를 선보인다는 계획과 그 기술적 배경을 설명하는 공급사 자료입니다. 메모리 산업 관점에서 CXL은 연결 프로토콜로서, GPU의 HBM 용량 한계를 CXL로 연결한 DRAM 풀로 보완하려는 시도입니다. 원문은 CXL 메모리 풀이 GPU VRAM을 보강해 KV 캐시를 저장하고 TTFT를 줄인다고 주장하며, 이는 HBM 수요를 대체하기보다는 HBM 탑재 GPU 옆에 별도 DRAM 계층을 추가하는 방향입니다. 다만 원문은 CXL 풀에 어떤 DRAM 규격(DDR5, LPDDR 등)이 사용되는지, 어느 메모리 공급사의 제품인지 명시하지 않아 미확인입니다. 또한 3.8배·6.5배 속도 향상은 XConn의 데모 조건(OPT-6.7B, H100 80GB 2대, 64 프롬프트, 512토큰)에 한정된 수치로, 실제 양산 환경의 성능·채택 규모는 확인되지 않습니다. 공급 배분 측면에서는 CXL 풀용 서버 DDR 수요가 HBM과 별개로 증가할 수 있으나, 고객 인증·제품 믹스·투자 일정에 대한 구체적 영향은 원문에 없습니다. NVIDIA Dynamo·KV Block Manager와의 통합 언급은 소프트웨어 생태계 연결을 시사하나, NVIDIA가 이 데모에 참여했는지 여부는 원문에서 확인되지 않습니다. 반대 근거로는 CXL의 200~500ns 지연이 HBM의 온패키지 대역폭에 크게 못 미친다는 점, PCIe GPU의 CXL 미지원을 하이브리드 스위치로 우회한다는 점이 확장성·비용의 불확실성으로 남습니다. 확인할 지표는 CXL 풀에 탑재되는 DRAM 규격·용량, 실제 고객 채택 사례, 클러스터당 100TiB 확장의 실현 조건, 그리고 이 데모가 HBM 수요에 미치는 순효과입니다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 4527자

작성자: XConn Technologies

대규모 언어 모델(LLM)과 생성형 AI 워크로드가 계속 복잡해짐에 따라 메모리는 빠르게 새로운 병목으로 떠오르고 있습니다. GPU는 병렬 컴퓨팅 성능에서 비할 데 없지만, 온보드 메모리 용량은 여전히 제한적입니다. 현대 AI 워크로드, 특히 KV 캐시 사용량이 많은 LLM 추론은 GPU당 80~120GB를 일상적으로 초과하여 높은 지연 시간과 시스템 전반의 비용이 큰 데이터 이동을 초래합니다.

슈퍼컴퓨팅 2025(SC25) 기간 중 [CXL 파빌리온(부스 #817)](https://computeexpresslink.org/event/supercomputing-2025/)에서 XConn Technologies는 CXL 메모리 풀링이 이 병목을 깨고 확장 가능한 메모리 중심 AI 아키텍처의 새로운 부류를 열 수 있음을 보여주는 라이브 데모를 선보일 예정입니다.

현재 다중 GPU 추론 환경에서는 데이터가 길고 비효율적인 경로를 거쳐야 합니다:

GPU → DRAM → NIC → 스토리지 서버 → NIC → DRAM → GPU

각 홉은 오버헤드, 지연 시간, 에너지 비용을 추가합니다. OPT-6.7B나 GPT 변형과 같은 대규모 LLM을 서빙할 때, 프리필/디코드 KV 캐시 관리의 작은 비효율조차도 수 초의 지연과 낭비되는 컴퓨팅 사이클로 증폭됩니다. 그 결과:

* 첫 토큰까지의 시간(TTFT) 증가
* 낮은 GPU 활용률
* 높은 데이터 이동 에너지 비용

CXL은 200~500ns 범위의 지연 시간으로 메모리 시맨틱 접근을 제공하며, 이는 NVMe 기술의 약 100μs, 스토리지 기반 메모리 공유의 10ms 초과와 비교됩니다. 이러한 지연 시간 개선은 컴퓨팅 노드 전반에 걸쳐 메모리 자원을 진정으로 동적이고 세밀하게 공유할 수 있게 합니다.

XConn의 CXL 스위치는 GPU 메모리의 직접 접근, 저지연 확장으로 작동하는 공유 CXL 메모리 풀을 도입합니다. 네트워크 인터페이스와 스토리지 서버를 통해 데이터를 라우팅하는 대신, GPU(또는 호스트 CPU)가 CUDA 호환 시맨틱으로 CXL 메모리 풀에 직접 읽기/쓰기를 수행하여 중복 복사와 두꺼운 소프트웨어 스택을 제거할 수 있습니다.

**데모 개요**

이 데모는 각각 NVIDIA H100 GPU(80GB)를 탑재한 두 서버에서 OPT-6.7B 모델을 요청당 64개 프롬프트, 프롬프트당 512토큰으로 실행하여 CXL 메모리 풀링의 성능 이점을 보여줍니다. 워크로드를 프리필과 디코드 단계로 분리함으로써, 이 구성은 공유 CXL 메모리가 AI 추론을 가속할 수 있는 방법을 보여줍니다.

RDMA 기반 공유와 비교하여, CXL 메모리 풀은 200G RDMA 대비 3.8배, 100G RDMA 대비 6.5배의 속도 향상을 달성했으며, TTFT의 극적인 감소와 대역폭 효율 개선을 이루었습니다. 이를 통해 이 데모는 AI 워크로드가 클러스터당 최대 100TiB까지의 동적이고 확장 가능한 메모리 확장, 비용 효율적인 자원 활용, 최소한의 CPU 개입으로 에너지 효율적인 데이터 접근, 그리고 NVIDIA Dynamo 및 KV Block Manager와 같은 AI 프레임워크와의 통합으로부터 어떻게 이점을 얻는지 강조합니다.

이 아키텍처의 핵심 구현 요소는 XConn의 차세대 기술인 Ultra IO Transformer로, 대부분 CXL을 네이티브로 지원하지 않는 PCIe GPU가 XConn의 하이브리드 스위치를 통해 CXL 메모리 풀에 직접 접근할 수 있게 하여 저지연, 고대역폭 통신을 유지합니다.

**실제 응용 사례**

CXL 메모리 풀링은 실험실을 넘어 실제 데이터센터와 AI 환경 전반에 걸쳐 실질적인 이점을 제공하고 있습니다. 대규모 메모리 풀에 대한 유연하고 공유된 접근을 가능하게 함으로써, 조직이 데이터 집약적 워크로드를 가속하고, 자원 활용을 개선하며, 총 소유 비용(TCO)을 절감하도록 돕습니다. CXL 메모리 풀링의 실제 응용 사례는 다음과 같습니다:

* **AI 추론 및 KV 캐시 확장:** CXL 메모리는 KV 캐시 저장을 위해 GPU VRAM을 보강하여 토큰 디코딩을 가속하고 LLM 서빙의 TTFT를 감소시킵니다.
* **과학 및 HPC 워크로드:** PNNL Crete와 같은 프로젝트는 컴퓨팅 노드 전반에 걸친 고처리량 메모리 공유를 위해 CXL 풀을 사용합니다.
* **클라우드 데이터베이스:** 대규모 인메모리 데이터베이스는 CXL 메모리 풀을 통합하여 유연한 확장이 가능한 고성능 데이터베이스 버퍼 풀을 구현합니다.

메모리 장벽은 오랫동안 AI 확장성의 제한 요인이었습니다. CXL 메모리 풀을 배포하면 고속, 분리형 메모리의 새로운 계층이 생성되어 AI 인프라를 구축하고 배포하는 방식을 재편합니다.

XConn의 데모는 CXL 아키텍처가 단지 이론에 그치지 않음을 증명합니다. CXL 시스템은 오늘날 실제 LLM 추론을 지원할 준비가 되어 있으며, 더 높은 성능, 더 낮은 지연 시간, 확장 가능한 메모리 용량을 제공하고 모든 규모의 AI에 대해 더 낮은 TCO를 제공합니다.

11월 18일부터 20일까지 SC25의 [CXL 파빌리온(부스 #817)](https://computeexpresslink.org/event/supercomputing-2025/)에서 데모를 확인하세요. 여러분을 만나기를 기대합니다!
브리프용 요약 초안
XConn Technologies가 SC25에서 CXL 메모리 풀링 라이브 데모를 선보일 계획이다. H100 GPU 2대에서 OPT-6.7B를 실행해 RDMA 대비 3.8~6.5배 속도 향상을 주장하며, CXL을 GPU 메모리 확장 계층으로 제시한다. 다만 CXL 풀에 쓰이는 DRAM 규격·공급사·실제 채택 규모는 원문에서 확인되지 않는다.

원문 텍스트

원문 열기 ↗
By: XConn Technologies

As large language models (LLMs) and generative AI workloads continue to grow in complexity, memory is quickly becoming the new bottleneck. While GPUs are unmatched in parallel compute performance, their onboard memory capacity remains limited. Modern AI workloads, especially LLM inference with heavy KV cache usage, routinely exceed 80 to 120 GB per GPU, leading to high latency and costly data movement across systems.

At the [CXL Pavilion (Booth #817)](https://computeexpresslink.org/event/supercomputing-2025/) during Supercomputing 2025 (SC25), XConn Technologies will showcase a live demo illustrating how CXL memory pooling can shatter this bottleneck and unlock a new class of scalable, memory-centric AI architectures.

Currently, in a multi-GPU inference setup, data must traverse a long and inefficient path:

GPU → DRAM → NIC → Storage Server → NIC → DRAM → GPU

Each hop adding overhead, latency, and energy costs. When serving large-scale LLMs such as OPT-6.7B or GPT variants, even small inefficiencies in prefill/decode KV cache management multiply into seconds of delay and wasted compute cycles. This results in:

* Longer Time to First Token (TTFT)
* Low GPU utilization
* High data-movement energy cost

CXL provides memory-semantic access with latency in the 200–500 ns range, compared to ~100 μs for NVMe technology and >10 ms for storage-based memory sharing. This latency improvement enables truly dynamic, fine-grained sharing of memory resources across compute nodes.

XConn’s CXL switch introduces a shared CXL memory pool that acts as a direct-access, low-latency extension of GPU memory. Instead of routing data through network interfaces and storage servers, GPUs (or their host CPUs) can perform direct reads/writes to the CXL memory pool with CUDA-compatible semantics, eliminating redundant copies and thick software stacks.

**Demo Overview**

The demo will show two servers, each with an NVIDIA H100 GPU (80GB), run the OPT-6.7B model with 64 prompts per request and 512 tokens per prompt to showcase the performance benefits of CXL memory pooling. By disaggregating the workload between the prefill and decode stages, the setup demonstrates how shared CXL memory can accelerate AI inferencing.

Compared to RDMA-based sharing, the CXL memory pool achieved 3.8 times speedup compared with 200G RDMA, 6.5 times speedup compared with 100G RDMA, and a dramatic reduction in TTFT and improved bandwidth efficiency. With this, the demonstration highlights how AI workloads benefit from dynamic, scalable memory expansion up to 100 TiB per cluster, cost-effective resource utilization, energy-efficient data access with minimal CPU involvement, and integration with AI frameworks like NVIDIA Dynamo and KV Block Manager

A key enabler of this architecture is XConn’s next-generation technology, the Ultra IO Transformer, which allows PCIe GPUs, most of which don’t natively support CXL, to directly access the CXL memory pool through XConn’s hybrid switch, maintaining low-latency, high-bandwidth communication.

**Real World Applications**

CXL memory pooling is moving beyond the lab to deliver tangible benefits across real-world data center and AI environments. By enabling flexible, shared access to large pools of memory, it helps organizations accelerate data-intensive workloads, improve resource utilization, and reduce total cost of ownership. Real world applications of CXL memory pooling include:

* **AI Inference & KV Cache Scaling:** CXL memory augments GPU VRAM for KV cache storage, accelerating token decoding and reducing TTFT for LLM serving.
* **Scientific & HPC Workloads:** Projects like PNNL Crete use CXL pools for high-throughput memory sharing across compute nodes.
* **Cloud Databases:** Large in-memory databases integrate CXL memory pools to enable a high-performance database buffer pool with flexible scaling.

The memory wall has long been the limiting factor for AI scalability. Deploying a CXL memory pool creates a new tier of high-speed, disaggregated memory, reshaping how we build and deploy AI infrastructure.

XConn’s demo proves that CXL architectures are not just theoretical. CXL systems are ready to power real-world LLM inference today, bringing higher performance, lower latency, and scalable memory capacity, with lower TCO for AI at any scale.

Check out the demo at the [CXL Pavilion (Booth #817)](https://computeexpresslink.org/event/supercomputing-2025/) at SC25 from November 18-20. We hope to see you there!