MEMORY INDUSTRY INTELLIGENCE

CXL로 토큰당 비용 개선: FMS 2026 요약 - Compute Express Link

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v1 · 10.10 23:58사용자 검토 전 초안

원문 제목: Increasing Tokens-per-Dollar with CXL: FMS 2026 Recap - Compute Express Link

핵심 요약

CXL 컨소시엄은 2026년 8월 4~6일 캘리포니아 산타클라라에서 열린 FMS(Future of Memory and Storage) 2026에 참가해 CXL 기술이 메모리 확장, 스토리지 병목 완화, 메모리 자원 활용 효율화를 통해 AI 인프라의 토큰당 비용(tokens-per-dollar)을 개선할 수 있음을 홍보했다. 오픈 스탠더드 파빌리온의 키오스크에서는 AI 추론, 메모리 확장, 영구 메모리, 분리형 메모리 관련 회원사 영상 데모가 전시됐고, CXL 4.0과 CXL 메모리 풀링이 FMS 2026 Best of Show 어워드의 Connectivity Award를 수상했다. CXL 컨소시엄 후원 패널에서는 KV 캐시 데이터 증가에 대응해 CXL 메모리 풀링·공유, 메모리 계층화, KV 캐시 오프로딩이 스토리지 병목을 줄이고 컴퓨트 자원 활용을 개선할 수 있다는 논의가 이뤄졌다. 전시장에서는 CXL 3.x 기반 대규모 메모리 확장(테라바이트급 DDR5 용량 추가), 멀티호스트 메모리 구성과 동적 할당, CXL 3.2 컨트롤러와 64 GT/s 연결, 검증 도구, 영구 메모리, 메모리 근접 처리, 컴포저블 인프라 등의 진전이 소개됐다. 2026년 전망으로 AI 추론과 에이전트형 AI가 더 큰 메모리 풋프린트를 요구함에 따라 CXL이 CPU·GPU·가속기 간 메모리 확장·풀링·공유 옵션을 제공해 토큰당 비용을 개선할 수 있다고 밝혔다.

메모리 산업 영향 분석

이 원문은 CXL 컨소시엄의 FMS 2026 참가 홍보 자료로, CXL 기반 메모리 확장·풀링·분리와 KV 캐시 오프로딩이 AI 추론의 토큰당 비용을 개선할 수 있다는 생태계 차원의 주장을 담고 있다. 메모리 산업 관점에서는 CXL이 CPU에 직접 연결된 DRAM과 별개로 테라바이트급 DDR5 용량을 추가할 수 있다는 데모가 언급되어 서버 DDR 수요의 추가 경로가 될 수 있으나, 이는 CXL로 연결한 DDR 용량을 서버 DDR 총량에 별도로 더하지 않는다는 전제에서 관찰해야 한다. CXL 메모리 풀링·계층화는 HBM·DRAM·SSD·HBF·CXL 사이 데이터/워크로드 위치 이동(tier_shift)을 만들 수 있는데, KV 캐시를 HBM에서 CXL 확장 메모리로 오프로드하면 HBM 사용량이 줄고 CXL 연결 DDR 사용량이 늘어나는 방향과, 반대로 더 큰 메모리 풋프린트가 전체 메모리 사용량을 늘리는 방향이 동시에 존재할 수 있어 순효과는 미확인이다. 원문은 CXL 4.0과 CXL 메모리 풀링의 수상, CXL 3.2 컨트롤러와 64 GT/s 연결, 검증 도구, 영구 메모리, 메모리 근접 처리, 컴포저블 인프라 진전을 언급하지만 실제 제품 양산·출하량·고객 채택·주문량은 확인되지 않는다. Intel은 패널 사회자로, Micron은 패널 참여자로 언급될 뿐 이 원문에서 특정 제품 공급·거래 관계는 확인되지 않는다. 확인할 지표로는 CXL 연결 DDR5의 실제 출하·채택, KV 캐시 오프로딩 도입 시 HBM·서버 DDR·eSSD 수요 변화, CXL 3.2 컨트롤러 양산 시점, 표준 발표와 실제 제품 양산의 구분이 있다. research_topics의 표준·밸류체인 질문 중 새 표준·설계·인터페이스 언급은 있으나 실제 채택 사례·양산·일정·주문량은 이 원문만으로 확인되지 않아 미확인으로 남긴다.
한국어 번역 읽기

수집된 원문 v1의 전체 본문 기준 · 6027자

CXL 컨소시엄은 2026년 8월 4~6일 캘리포니아 산타클라라에서 열린 [FMS(Future of Memory and Storage) 2026](https://computeexpresslink.org/event/future-of-memory-and-storage-fms-2026/)에 다시 참가해, CXL 기술이 사용 가능한 메모리를 확장하고 스토리지 병목을 줄이며 메모리 자원을 더 효율적으로 사용하도록 함으로써 AI 인프라의 토큰당 비용(tokens-per-dollar)을 높이는 데 어떻게 기여할 수 있는지 선보였다.

CXL 컨소시엄 키오스크의 영상 데모와 AI 추론에 관한 전문가 패널부터 새로운 CXL 3.x 솔루션과 전시장 곳곳의 회원사 데모에 이르기까지, FMS는 CXL 기술에 대한 증가하는 모멘텀을 부각했다.

**오픈 스탠더드 파빌리온의 CXL 컨소시엄 키오스크**

오픈 스탠더드 파빌리온 내 CXL 컨소시엄 키오스크에서 참석자들은 컨소시엄 관계자들과 만나 CXL이 더 효율적인 메모리 아키텍처, 개선된 자원 활용, 확장 가능한 AI 인프라를 통해 조직이 토큰당 비용을 높일 수 있게 하는 방식을 살펴봤다.

키오스크에서는 CXL 컨소시엄 회원사들의 영상 데모가 상영되어 AI 추론, 메모리 확장, 영구 메모리, 분리형 메모리 전반에서 CXL 솔루션이 실제로 작동하는 모습을 보여줬다. 영상 데모는 CXL 생태계가 확장·풀링·분리형 메모리를 어떻게 실천에 옮기고 있는지 보여줬다. 사용 가능한 메모리를 늘리고 더 많은 워크로드 데이터를 컴퓨트에 더 가깝게 유지함으로써, CXL은 더 느린 스토리지 계층에 대한 의존을 줄이고 자원 활용을 개선하며 조직이 AI 인프라에서 더 많은 것을 얻도록 도울 수 있다.

[FMS 2026 CXL 영상 데모 보기](https://computeexpresslink.org/fms-2026-cxl-demos/)

**FMS에서 인정받은 CXL 혁신**

CXL 생태계 혁신은 FMS 2026 Best of Show 어워드에서도 인정받았으며, CXL 4.0과 CXL 메모리 풀링이 Connectivity Award를 수상했다.

이 수상은 점점 더 데이터 집약적인 워크로드에 필요한 연결성, 메모리, 인프라 기술을 기업들이 개발함에 따라 CXL 생태계 전반에서 혁신이 계속되고 있음을 반영한다.



**CXL 전문가들이 AI 인프라 효율성 향상 방안을 논의하다**

CXL 컨소시엄이 후원한 패널 “메모리 장벽 너머: CXL 메모리 풀링과 공유가 AI 추론을 어떻게 바꾸고 있는가”는 메모리 아키텍처와 AI 인프라 전문가들을 모아 CXL이 AI 추론 배포가 직면한 가장 시급한 과제 일부를 어떻게 해결할 수 있는지 검토했다.

CXL 컨소시엄 마케팅 워킹 그룹 의장인 아닐 고드볼(Intel)이 사회를 맡았고, 패널에는 산딥 다타프라사드(Astera Labs), 수밋 푸리(Liqid), 루이스 안카하스(Micron), 로넨 하얏(UnifabriX)이 참여했다.



AI 모델과 컨텍스트 윈도가 계속 커지면서 추론 중 저장·접근해야 하는 KV 캐시 데이터의 양도 늘어난다. 패널은 CXL 메모리 풀링과 공유가 컴퓨트에 더 가까운 추가 메모리 용량을 제공할 수 있는 방식과, 더 느린 스토리지에 대한 의존을 줄이는 메모리 계층화 및 KV 캐시 오프로딩 전략을 탐구했다.

KV 캐시 워크로드의 경우 스토리지 병목을 줄이면 추론에 필요한 데이터 접근을 가속화하는 동시에 값진 컴퓨트 자원의 활용도를 개선할 수 있다. CXL 메모리 풀링은 또한 CPU, GPU, 가속기가 이기종 인프라 전반에서 메모리 자원에 더 효율적으로 접근하고 공유할 기회를 만든다.



**전시장에서 드러난 CXL 생태계 모멘텀**

CXL에 대한 모멘텀은 FMS 전시장 전체로 이어졌으며, 컨소시엄 회원사들은 CXL 3.x 기반 기술을 선보이고 생태계가 메모리 장치, 컨트롤러, 풀링 아키텍처, 소프트웨어 전반에서 계속 성숙하고 있음을 시연했다.

여러 데모는 CXL을 통해 테라바이트급 DDR5 용량을 추가할 수 있는 솔루션을 포함해 대규모 메모리 확장을 부각했다. 이러한 사례는 데이터센터가 CPU에 연결된 DRAM과 별개로 사용 가능한 메모리를 늘려 AI, 분석, 인메모리 데이터베이스 및 기타 메모리 집약적 애플리케이션을 지원할 유연성을 높일 수 있음을 보여줬다.

다른 데모는 멀티호스트 메모리 구성과 동적 메모리 할당을 포함한 CXL 3.x 메모리 풀링 및 분리에 초점을 맞췄다. 이러한 기능은 인프라 운영자가 메모리를 개별 프로세서에 정적으로 붙이는 대신 더 효율적으로 프로비저닝하고 필요한 곳에 용량을 제공할 수 있음을 보여준다.

AI 추론도 또 다른 주요 주제였다. 생태계 전반의 데모는 CXL 풀링·확장 메모리가 KV 캐시 저장, 전송, 재사용을 지원해 점점 더 까다로워지는 AI 워크로드를 위한 더 큰 메모리 계층을 만드는 방식을 보여줬다. 서로 다른 메모리와 스토리지 기술을 결합한 접근 방식은 CXL이 용량, 성능, 비용의 균형을 맞추는 계층형 아키텍처를 가능하게 하는 방식도 보여줬다.

생태계는 또한 CXL 3.2 컨트롤러와 64 GT/s 연결, 검증 도구, 영구 메모리, 메모리 근접 처리, 컴포저블 인프라에서의 지속적인 진전을 선보였다.

종합적으로 이러한 발전은 CXL이 실리콘과 컨트롤러부터 메모리 장치, 소프트웨어, 완전한 시스템 아키텍처에 이르기까지 전체 생태계에서 진전하고 있음을 보여준다.

**2026년 전망**

AI 추론과 새롭게 부상하는 에이전트형 AI 애플리케이션이 더 큰 메모리 풋프린트를 요구함에 따라, 제한된 로컬 메모리와 더 느린 스토리지 사이에서 데이터를 이동하는 것은 값진 컴퓨트 자원을 제약할 수 있다. CXL은 CPU, GPU, 가속기 전반에서 메모리를 확장·풀링·공유할 새로운 옵션을 제공해 이러한 까다로운 워크로드를 위한 더 유연한 메모리 아키텍처를 만든다.

사용 가능한 메모리를 확장하고 스토리지 병목을 줄이며 메모리 활용도를 개선하고 KV 캐시 계층화 및 오프로드에 대한 새로운 접근을 가능하게 함으로써, CXL은 조직이 컴퓨트 인프라에서 더 많은 가치를 얻고 토큰당 비용을 높이도록 도울 수 있다.

FMS 2026에 함께해 CXL 생태계의 지속적인 모멘텀을 보여준 모든 회원사, 발표자, 참석자에게 감사한다. 9월 15~17일 산타클라라에서 열리는 [AI Infra Summit](https://computeexpresslink.org/event/ai-infra-summit-2026/)에서 많은 분을 만나기를 기대한다.
브리프용 요약 초안
CXL 컨소시엄은 FMS 2026에서 CXL 기반 메모리 확장·풀링·분리와 KV 캐시 오프로딩이 AI 추론의 토큰당 비용을 개선할 수 있다고 홍보했고, CXL 4.0과 CXL 메모리 풀링이 Connectivity Award를 수상했다. 전시장에서는 테라바이트급 DDR5 용량 추가, 멀티호스트 메모리 구성, CXL 3.2 컨트롤러와 64 GT/s 연결 등이 소개됐으나 실제 제품 양산·출하·고객 채택은 확인되지 않았다.

원문 텍스트

원문 열기 ↗
The CXL Consortium returned to [Future of Memory and Storage (FMS) 2026](https://computeexpresslink.org/event/future-of-memory-and-storage-fms-2026/), August 4–6 in Santa Clara, California, to showcase how CXL technology can help AI infrastructure increase tokens-per-dollar by expanding available memory, reducing storage bottlenecks and enabling more efficient use of memory resources.

From video demos at the CXL Consortium kiosk and an expert panel on AI inference to new CXL 3.x solutions and demonstrations from members across the show floor, FMS highlighted the growing momentum behind CXL technology.

**CXL Consortium Kiosk at the Open Standards Pavilion**

At the CXL Consortium kiosk in the Open Standards Pavilion, attendees connected with Consortium representatives and explored how CXL enables organizations to increase tokens-per-dollar through more efficient memory architectures, improved resource utilization and scalable AI infrastructure.

The kiosk featured video demonstrations from CXL Consortium members, highlighting CXL solutions in action across AI inference, memory expansion, persistent memory and disaggregated memory. The video demonstrations showed how the CXL ecosystem is putting expanded, pooled and disaggregated memory into practice. By increasing available memory and keeping more workload data closer to compute, CXL can reduce dependence on slower storage tiers, improve resource utilization and help organizations get more from their AI infrastructure.

[View the FMS 2026 CXL Video Demos](https://computeexpresslink.org/fms-2026-cxl-demos/)

**CXL Innovation Recognized at FMS**

CXL ecosystem innovation was also recognized through the FMS 2026 Best of Show Awards, with CXL 4.0 and CXL Memory Pooling receiving the Connectivity Award.

The recognition reflects the continued innovation across the CXL ecosystem as companies develop the connectivity, memory and infrastructure technologies needed for increasingly data-intensive workloads.



**CXL Experts Explore How to Increase AI Infrastructure Efficiency**

The CXL Consortium-sponsored panel, “Beyond the Memory Wall: How CXL Memory Pooling and Sharing Are Transforming AI Inference,” brought together experts in memory architecture and AI infrastructure to examine how CXL can address some of the most pressing challenges facing AI inference deployments.

Moderated by Anil Godbole (Intel), CXL Consortium Marketing Working Group Chair, the panel featured Sandeep Dattaprasad (Astera Labs), Sumit Puri (Liqid), Luis Ancajas (Micron) and Ronen Hyatt (UnifabriX).



As AI models and context windows continue to grow, so does the amount of KV Cache data that needs to be stored and accessed during inference. The panel explored how CXL memory pooling and sharing can provide additional memory capacity closer to compute, along with memory tiering and KV Cache offloading strategies that reduce reliance on slower storage.

For KV Cache workloads, reducing the storage bottleneck can accelerate access to the data needed for inference while improving the utilization of valuable compute resources. CXL memory pooling also creates opportunities for CPUs, GPUs, and accelerators to access and share memory resources more efficiently across heterogeneous infrastructure.



**CXL Ecosystem Momentum on Display**

The momentum behind CXL extended throughout the FMS show floor, where Consortium members showcased technologies based on CXL 3.x and demonstrated how the ecosystem continues to mature across memory devices, controllers, pooling architectures and software.

Several demonstrations highlighted large-scale memory expansion, including solutions capable of adding terabytes of DDR5 capacity through CXL. These examples showed how data centers can increase available memory independently of CPU-attached DRAM, creating greater flexibility to support AI, analytics, in-memory databases and other memory-intensive applications.

Other demonstrations focused on CXL 3.x memory pooling and disaggregation, including multi-host memory configurations and dynamic memory allocation. These capabilities show how CXL can help infrastructure operators provision memory more efficiently and make capacity available where it is needed, rather than statically attaching it to individual processors.

AI inference was another major theme. Demonstrations from across the ecosystem showed how CXL pooled and expanded memory can support KV Cache storage, transfer and reuse, creating larger memory tiers for increasingly demanding AI workloads. Approaches combining different memory and storage technologies also demonstrated how CXL can enable tiered architectures that balance capacity, performance and cost.

The ecosystem also showcased continued progress in CXL 3.2 controllers and 64 GT/s connectivity, validation tools, persistent memory, processing near memory and composable infrastructure.

Collectively, these developments show that CXL is progressing across the full ecosystem, from silicon and controllers to memory devices, software and complete system architectures.

**Looking Ahead for 2026**

As AI inference and emerging agentic AI applications require larger memory footprints, moving data between limited local memory and slower storage can constrain valuable compute resources. CXL provides new options to expand, pool and share memory across CPUs, GPUs and accelerators, creating more flexible memory architectures for these demanding workloads.

By expanding available memory, reducing storage bottlenecks, improving memory utilization and enabling new approaches to KV Cache tiering and offload, CXL can help organizations get more value from their compute infrastructure and increase tokens-per-dollar.

Thank you to all the members, speakers, and attendees who joined us at FMS 2026 and helped demonstrate the CXL ecosystem’s continued momentum. We look forward to seeing many of you at the [AI Infra Summit](https://computeexpresslink.org/event/ai-infra-summit-2026/) event in Santa Clara, from September 15 – 17.