MEMORY INDUSTRY INTELLIGENCE
CXL®로 구현하는 DRAM 자원 확장성 - Compute Express Link
한국어 번역·요약·분석
원문 제목: DRAM Resource Scalability Enabled by CXL® - Compute Express Link
핵심 요약
이 문서는 서버 아키텍처가 AI·ML 등으로 인해 더 많은 메모리 용량과 대역폭을 요구하게 되면서, CXL(Compute Express Link)이 메모리 의미론(memory semantic) 프로토콜로서 확장·공유 가능한 인터페이스를 제공한다고 설명한다. CXL은 PCIe 전기 인터페이스를 활용해 외부 메모리 컨트롤러를 연결함으로써 프로세서에 추가 메모리 계층을 붙이는 방식이며, CXL 2.0 스위치를 통해 여러 프로세서와 메모리 장치를 재할당(re-provisioning)할 수 있다고 주장한다. 문서는 CXL 기반 메모리 컨트롤러가 16레인 기준 최대 4 GB/s 대역폭을 제공하고, DDR4와 DDR5를 혼용해 비용 효율을 높일 수 있으며, DIMM과 EDSFF 폼팩터를 활용할 가능성이 높다고 전망한다. 또한 CXL이 직접 부착 메모리 대비 핀 수와 전기적 도달 거리 문제를 완화하고, 메모리 자원 활용률을 최적화하며 열·전력 관리 부담을 줄인다고 주장한다. 이 문서는 CXL 컨소시엄의 기술 백서 성격으로, 특정 제품 양산·고객 채택·거래에 대한 확인은 포함하지 않는다.
메모리 산업 영향 분석
이 문서는 CXL 컨소시엄의 기술 백서로, CXL이 DRAM 확장을 위한 연결 프로토콜로서 서버 메모리 계층에 추가 계층을 형성할 수 있음을 설명한다. 메모리 산업 관점에서 직접적인 영향은 서버 DDR(제품 ID 2)의 공급 배분과 제품 믹스 선택에 있다. CXL 연결 DDR은 서버 DDR의 총량을 늘리는 것이 아니라, 기존 서버 DDR 모듈을 CXL 스위치·컨트롤러를 통해 재할당·공유하는 방식이므로, 서버 DDR 비트 수요에 미치는 효과는 불확실하다. 다만 DDR4와 DDR5를 혼용할 수 있다는 점은 DDR4 수요를 일부 연장할 수 있고, DDR5 전환 속도를 늦출 수 있는 조건부 효과가 있다. 고객 인증 측면에서는 CXL 2.0 스위치와 외부 메모리 컨트롤러의 상호 운용성 확보가 관건이며, 문서는 DIMM과 EDSFF 폼팩터 채택 가능성을 전망하지만 실제 양산·채택 사례는 확인되지 않는다. 투자 일정 측면에서 CXL은 다중 소켓 CPU·메모리 솔루션을 단순화해 CPU 코어 수를 줄일 수 있다고 주장하므로, 서버 CPU 코어 수 감소가 서버 DDR 탑재량에 미치는 영향은 상충 효과가 있다. HBM·LPDDR·GDDR·NAND·eSSD와의 직접 연결 근거는 이 문서에 없다. CXL 메모리의 실제 DRAM 규격·사용처와 HBF 개발·인증 상태는 미확인이다. CXL이 서버 DDR의 계층 이동(tier_shift)을 일으키는지 여부는 CXL 연결 DDR이 네이티브 DDR을 대체하는지 보완하는지에 달려 있으며, 문서는 보완적 확장을 강조하지만 실제 배분 비율은 확인되지 않는다. 판단 변경 조건은 CXL 2.0 스위치·컨트롤러의 실제 서버 채택, DDR4/DDR5 혼용 구성의 고객 인증, CXL 연결 DDR의 지연 시간이 HPC 애플리케이션 요구를 충족하는지에 대한 실측이다.
한국어 번역 읽기
수집된 원문 v1의 전체 본문 기준 · 7113자
## 서론
컴퓨터 서버 아키텍처는 분석 애플리케이션에 제공되는 더 큰 데이터 세트로 더 빠른 분석을 지원하기 위해 지속적으로 진화하고 있다. 중앙처리장치(CPU)와 그래픽처리장치(GPU)의 컴퓨팅 성능은 인공지능(AI)과 머신러닝(ML) 같은 현대 애플리케이션을 가능하게 하기 위해 증가하고 있다. 프로세서의 코어 수가 증가함에 따라 처리되는 데이터가 더 많아지고, 더 많은 메모리 용량과 대역폭이 필요해진다. 애플리케이션의 다양성은 메모리 활용도에 따라 CPU 코어와 메모리 요구 사항이 크게 달라진다는 것을 의미한다. 시스템 공급업체와 관리자는 모든 애플리케이션을 지원할 충분한 메모리를 확보하면서도 메모리 과잉 프로비저닝과 저활용을 피하도록 시스템 자원을 최적화하는 과제를 안고 있다.
CXL®(Compute Express Link®)은 최고 성능의 PCIe® 전기 인터페이스 규격을 활용하여 차세대 플랫폼을 위한 구성 가능하고 확장 가능하며 공유 가능한 인터페이스를 제공하는 메모리 의미론(memory semantic) 프로토콜이다. CXL 인터페이스의 도입은 메모리와 처리 할당의 유연성을 높여 고립된 자원을 제거하고 애플리케이션 요구 사항에 맞게 재프로비저닝할 수 있도록 함으로써 확장과 자원 공유를 가능하게 하는 중요한 진전이다.
모든 유형의 프로세서(CPU, GPU, 가속기 장치)는 일반적으로 DRAM 메모리에 최고 성능을 제공하는 최신 세대 DDR 메모리 인터페이스를 구현한다. CXL은 외부 메모리 컨트롤러 장치를 연결하여 처리 엔진이 사용할 수 있는 메모리 자원을 확장할 수 있게 한다. 즉, 추가 메모리 계층을 부착하는 것이다. 데이터는 CXL을 통해 외부 메모리 컨트롤러로 전달되며, 이 컨트롤러가 메모리 매체(일반적으로 DRAM)의 읽기/쓰기 작업을 관리한다. CXL 인터페이스는 매우 낮은 지연 시간을 갖도록 설계되어, 성능은 네이티브 DDR 인터페이스보다 낮지만 그에 필적하는 성능을 제공한다.
CXL 연결 DDR은 필요에 따라 자원을 추가하거나 제거할 수 있는 주요 애플리케이션이다. CXL 2.0 스위치는 변화하는 애플리케이션 요구에 대응하여 여러 처리 유형과 메모리 장치를 상호 연결하고 재프로비저닝(재할당)할 수 있게 하여, 애플리케이션 변경이 필요할 때 저활용될 위험 없이 더 큰 메모리 자원 배포를 가능하게 한다.
## 외부 CXL 메모리 컨트롤러의 장점
지연 시간 요구는 각각 3200MT/s와 4800MT/s를 지원할 수 있는 DDR4/DDR5 같은 최신 메모리 기술을 통해 충족될 수 있지만, 직접 부착 DRAM 인터페이스 수를 늘리는 것은 고성능 컴퓨팅(HPC) 애플리케이션에 필요한 코어당 성능 향상을 달성한다. 이는 몇 가지 시스템 설계 과제를 제시한다. DRAM 인터페이스는 고속 신호를 위해 많은 장치 핀(288개)을 필요로 하며, 전기적 도달 거리(장치 간 거리) 문제는 프로세서의 근접 배치와 관련된 레이아웃 및 실면적 문제를 야기하여 국소적인 전력 및 냉각 요구를 악화시킬 수 있다. 그리고 직접 연결 메모리는 최고 성능을 보장하지만 유연성이 없고 사용되지 않은 메모리 용량을 시스템의 다른 프로세서와 공유할 수 없다.
## CXL 기반 메모리 확장
CXL 인터페이스는 각각 32 GT/s의 16레인에 최적화되어 있어 메모리 확장을 지원하는 핀 수 오버헤드를 크게 줄이고 DDR 확장에 필요한 데이터 대역폭을 쉽게 지원한다. 감소된 전체 대역폭으로 x8 및 x4 채널 폭을 지원하는 옵션도 있으며, CXL 전기 인터페이스는 PCIe Gen 5 및 Gen 6에 사용되는 것과 동일하여 시스템 설계에 대한 신뢰성과 견고한 채널 도달 거리를 보장한다.
CXL에는 특정 유형의 장치를 대상으로 하는 세 가지 프로토콜이 있다. CXL.mem과 CXL.cache는 메모리 접근 성능을 위해 최저 지연 시간을 보장하는 로드/스토어 메모리 의미론 프로토콜이다. CXL.io는 장치 관리, 오류 및 상태 보고에 사용되며 PCIe 트랜잭션을 기반으로 한다. CXL은 서로 다른 특성을 가진 세 가지 유형의 장치를 정의한다.
* 유형 1: 호스트 프로세서
* 유형 2: 처리, 캐싱 및 메모리 공유 기능을 가진 장치
* 유형 3: 메모리 장치
DDR 메모리 컨트롤러는 가장 일반적으로 유형 3 기능으로, 고성능, 용량 확장, 배포 확장성이라는 세 가지 주요 기능을 제공한다.
CPU가 CXL 인터페이스를 통해 컨트롤러와 연결되고 DDR 인터페이스를 통해 DRAM 메모리와 연결되는 CXL 기반 메모리 컨트롤러의 사용은 16레인 CXL 인터페이스를 통해 최대 4 GB/s를 달성하여 대역폭 요구 사항을 충족한다.
CXL은 프로세서 측에 필요한 핀 수를 줄여 DDR5 속도에서 직접 부착 메모리의 신호 무결성 문제와 실면적 문제를 완화하고, 동시에 HPC 애플리케이션의 지연 시간 요구를 충족하면서 더 많은 DPC를 지원한다.
## CXL 구현의 비용 효율성
CXL은 메모리 확장과 공유를 통해 이기종 처리 시스템을 가능하게 하여 메모리 자원 활용을 최적화하는 데 도움을 준다. 비용이 많이 들고 복잡한 다중 소켓 CPU 및 메모리 솔루션은 CXL 메모리 확장을 통해 단순화되며 전체 비용이 절감되고 애플리케이션 성능이 향상된다.
CXL 기반 메모리 컨트롤러 솔루션은 비용을 관리하고 이전 기술을 재사용할 수 있도록 다양한 메모리 기술을 지원하는 유연성을 제공한다. 예를 들어, DDR4 메모리는 여전히 시장에서 주를 이루고 있으며 비트당 비용이 DDR5보다 낮다. 시스템은 직접 부착 프로세서 메모리에는 고성능 DDR5를 사용하고 확장 CXL 부착 메모리에는 DDR4를 사용하도록 구성할 수 있으며, 나중에 업그레이드할 수 있는 옵션도 있다.
CXL 메모리는 프로세서와 메모리 장치 사이에 더 큰 물리적 거리를 허용하여 전력 활용 최적화와 더 저렴한 열 냉각 솔루션에 도움을 줄 수 있다. 더 큰 규모의 메모리 확장, 즉 분리(disaggregation)는 메모리 자원 풀을 공유하고 필요에 따라 재할당할 수 있게 한다.
메모리 애플리케이션의 경우, CXL은 가장 효율적인 밀도와 호환성을 제공하므로 DIMM(Dual In-Line Memory Module) 및 EDSFF(Enterprise and Datacenter Standard Form Factor) 폼팩터를 활용할 가능성이 높다. 공통 폼팩터 채택은 공급업체 간 상호 운용성을 보장하고 시스템을 쉽게 업그레이드하거나 재구성할 수 있게 한다.
## 결론
CXL 기반 메모리 솔루션은 클라우드 컴퓨팅, AI/ML, 클러스터 네트워크, HPC와 같은 처리 집약적 애플리케이션의 요구에 가장 적합한 기술을 제공한다. 이는 다음을 가능하게 한다.
* 다중 코어 프로세서의 메모리 대역폭, 용량 및 지연 시간 요구 증가
* 저활용 자원을 줄이기 위한 메모리 공유 및 재할당
* 혁신적인 AI, ML 및 신경 아키텍처를 위한 애플리케이션을 제공하는 이기종 시스템 가능
* 채택 용이성과 CPU 코어 수 감소를 통한 비용 효율적인 솔루션 제공
* 열 및 전력 요구의 관리 용이성 증가
미래 시스템 설계에서 비용 효율적인 메모리 확장을 가능하게 하려면 오늘날 CXL 기반 메모리 솔루션의 구현을 고려하라.
컴퓨터 서버 아키텍처는 분석 애플리케이션에 제공되는 더 큰 데이터 세트로 더 빠른 분석을 지원하기 위해 지속적으로 진화하고 있다. 중앙처리장치(CPU)와 그래픽처리장치(GPU)의 컴퓨팅 성능은 인공지능(AI)과 머신러닝(ML) 같은 현대 애플리케이션을 가능하게 하기 위해 증가하고 있다. 프로세서의 코어 수가 증가함에 따라 처리되는 데이터가 더 많아지고, 더 많은 메모리 용량과 대역폭이 필요해진다. 애플리케이션의 다양성은 메모리 활용도에 따라 CPU 코어와 메모리 요구 사항이 크게 달라진다는 것을 의미한다. 시스템 공급업체와 관리자는 모든 애플리케이션을 지원할 충분한 메모리를 확보하면서도 메모리 과잉 프로비저닝과 저활용을 피하도록 시스템 자원을 최적화하는 과제를 안고 있다.
CXL®(Compute Express Link®)은 최고 성능의 PCIe® 전기 인터페이스 규격을 활용하여 차세대 플랫폼을 위한 구성 가능하고 확장 가능하며 공유 가능한 인터페이스를 제공하는 메모리 의미론(memory semantic) 프로토콜이다. CXL 인터페이스의 도입은 메모리와 처리 할당의 유연성을 높여 고립된 자원을 제거하고 애플리케이션 요구 사항에 맞게 재프로비저닝할 수 있도록 함으로써 확장과 자원 공유를 가능하게 하는 중요한 진전이다.
모든 유형의 프로세서(CPU, GPU, 가속기 장치)는 일반적으로 DRAM 메모리에 최고 성능을 제공하는 최신 세대 DDR 메모리 인터페이스를 구현한다. CXL은 외부 메모리 컨트롤러 장치를 연결하여 처리 엔진이 사용할 수 있는 메모리 자원을 확장할 수 있게 한다. 즉, 추가 메모리 계층을 부착하는 것이다. 데이터는 CXL을 통해 외부 메모리 컨트롤러로 전달되며, 이 컨트롤러가 메모리 매체(일반적으로 DRAM)의 읽기/쓰기 작업을 관리한다. CXL 인터페이스는 매우 낮은 지연 시간을 갖도록 설계되어, 성능은 네이티브 DDR 인터페이스보다 낮지만 그에 필적하는 성능을 제공한다.
CXL 연결 DDR은 필요에 따라 자원을 추가하거나 제거할 수 있는 주요 애플리케이션이다. CXL 2.0 스위치는 변화하는 애플리케이션 요구에 대응하여 여러 처리 유형과 메모리 장치를 상호 연결하고 재프로비저닝(재할당)할 수 있게 하여, 애플리케이션 변경이 필요할 때 저활용될 위험 없이 더 큰 메모리 자원 배포를 가능하게 한다.
## 외부 CXL 메모리 컨트롤러의 장점
지연 시간 요구는 각각 3200MT/s와 4800MT/s를 지원할 수 있는 DDR4/DDR5 같은 최신 메모리 기술을 통해 충족될 수 있지만, 직접 부착 DRAM 인터페이스 수를 늘리는 것은 고성능 컴퓨팅(HPC) 애플리케이션에 필요한 코어당 성능 향상을 달성한다. 이는 몇 가지 시스템 설계 과제를 제시한다. DRAM 인터페이스는 고속 신호를 위해 많은 장치 핀(288개)을 필요로 하며, 전기적 도달 거리(장치 간 거리) 문제는 프로세서의 근접 배치와 관련된 레이아웃 및 실면적 문제를 야기하여 국소적인 전력 및 냉각 요구를 악화시킬 수 있다. 그리고 직접 연결 메모리는 최고 성능을 보장하지만 유연성이 없고 사용되지 않은 메모리 용량을 시스템의 다른 프로세서와 공유할 수 없다.
## CXL 기반 메모리 확장
CXL 인터페이스는 각각 32 GT/s의 16레인에 최적화되어 있어 메모리 확장을 지원하는 핀 수 오버헤드를 크게 줄이고 DDR 확장에 필요한 데이터 대역폭을 쉽게 지원한다. 감소된 전체 대역폭으로 x8 및 x4 채널 폭을 지원하는 옵션도 있으며, CXL 전기 인터페이스는 PCIe Gen 5 및 Gen 6에 사용되는 것과 동일하여 시스템 설계에 대한 신뢰성과 견고한 채널 도달 거리를 보장한다.
CXL에는 특정 유형의 장치를 대상으로 하는 세 가지 프로토콜이 있다. CXL.mem과 CXL.cache는 메모리 접근 성능을 위해 최저 지연 시간을 보장하는 로드/스토어 메모리 의미론 프로토콜이다. CXL.io는 장치 관리, 오류 및 상태 보고에 사용되며 PCIe 트랜잭션을 기반으로 한다. CXL은 서로 다른 특성을 가진 세 가지 유형의 장치를 정의한다.
* 유형 1: 호스트 프로세서
* 유형 2: 처리, 캐싱 및 메모리 공유 기능을 가진 장치
* 유형 3: 메모리 장치
DDR 메모리 컨트롤러는 가장 일반적으로 유형 3 기능으로, 고성능, 용량 확장, 배포 확장성이라는 세 가지 주요 기능을 제공한다.
CPU가 CXL 인터페이스를 통해 컨트롤러와 연결되고 DDR 인터페이스를 통해 DRAM 메모리와 연결되는 CXL 기반 메모리 컨트롤러의 사용은 16레인 CXL 인터페이스를 통해 최대 4 GB/s를 달성하여 대역폭 요구 사항을 충족한다.
CXL은 프로세서 측에 필요한 핀 수를 줄여 DDR5 속도에서 직접 부착 메모리의 신호 무결성 문제와 실면적 문제를 완화하고, 동시에 HPC 애플리케이션의 지연 시간 요구를 충족하면서 더 많은 DPC를 지원한다.
## CXL 구현의 비용 효율성
CXL은 메모리 확장과 공유를 통해 이기종 처리 시스템을 가능하게 하여 메모리 자원 활용을 최적화하는 데 도움을 준다. 비용이 많이 들고 복잡한 다중 소켓 CPU 및 메모리 솔루션은 CXL 메모리 확장을 통해 단순화되며 전체 비용이 절감되고 애플리케이션 성능이 향상된다.
CXL 기반 메모리 컨트롤러 솔루션은 비용을 관리하고 이전 기술을 재사용할 수 있도록 다양한 메모리 기술을 지원하는 유연성을 제공한다. 예를 들어, DDR4 메모리는 여전히 시장에서 주를 이루고 있으며 비트당 비용이 DDR5보다 낮다. 시스템은 직접 부착 프로세서 메모리에는 고성능 DDR5를 사용하고 확장 CXL 부착 메모리에는 DDR4를 사용하도록 구성할 수 있으며, 나중에 업그레이드할 수 있는 옵션도 있다.
CXL 메모리는 프로세서와 메모리 장치 사이에 더 큰 물리적 거리를 허용하여 전력 활용 최적화와 더 저렴한 열 냉각 솔루션에 도움을 줄 수 있다. 더 큰 규모의 메모리 확장, 즉 분리(disaggregation)는 메모리 자원 풀을 공유하고 필요에 따라 재할당할 수 있게 한다.
메모리 애플리케이션의 경우, CXL은 가장 효율적인 밀도와 호환성을 제공하므로 DIMM(Dual In-Line Memory Module) 및 EDSFF(Enterprise and Datacenter Standard Form Factor) 폼팩터를 활용할 가능성이 높다. 공통 폼팩터 채택은 공급업체 간 상호 운용성을 보장하고 시스템을 쉽게 업그레이드하거나 재구성할 수 있게 한다.
## 결론
CXL 기반 메모리 솔루션은 클라우드 컴퓨팅, AI/ML, 클러스터 네트워크, HPC와 같은 처리 집약적 애플리케이션의 요구에 가장 적합한 기술을 제공한다. 이는 다음을 가능하게 한다.
* 다중 코어 프로세서의 메모리 대역폭, 용량 및 지연 시간 요구 증가
* 저활용 자원을 줄이기 위한 메모리 공유 및 재할당
* 혁신적인 AI, ML 및 신경 아키텍처를 위한 애플리케이션을 제공하는 이기종 시스템 가능
* 채택 용이성과 CPU 코어 수 감소를 통한 비용 효율적인 솔루션 제공
* 열 및 전력 요구의 관리 용이성 증가
미래 시스템 설계에서 비용 효율적인 메모리 확장을 가능하게 하려면 오늘날 CXL 기반 메모리 솔루션의 구현을 고려하라.
브리프용 요약 초안
CXL 컨소시엄 백서는 CXL을 DRAM 확장·공유를 위한 메모리 의미론 프로토콜로 설명하며, CXL 2.0 스위치를 통한 재할당과 DDR4/DDR5 혼용으로 비용 효율을 높일 수 있다고 전망한다. 서버 DDR 공급 배분과 제품 믹스에 영향을 줄 수 있으나, 실제 양산·고객 채택·CXL 연결 DDR의 지연 시간 실측은 확인되지 않아 미확인으로 남긴다.
원문 텍스트
원문 열기 ↗## Introduction
Computer server architectures continually evolve to support faster analytics with larger data sets delivered for analysis applications. Computing capabilities of Central Processing Units (CPUs), and Graphics Processing Units (GPUs) increase to enable modern applications such as Artificial Intelligence (AI) and Machine Learning (ML). As core counts for processors increase, there is more data being processed requiring more memory capacity and bandwidth. The diversity of applications means that the CPU core and memory requirements vary widely based on their utilization of memory. System vendors and managers are challenged to optimize system resources to ensure there is enough memory to support all applications while avoiding memory overprovisioning and underutilization.
CXL® (Compute Express Link®) is a memory semantic protocol that leverages the highest performance PCIe® electrical interface specifications to deliver a configurable, scalable, sharable interface for a new generation of platforms. The introduction of the CXL interface is a major step toward enabling expansion and resource sharing by increasing the flexibility of memory and processing allocation to eliminate isolated resources and allow re-provisioning to suit application requirements.
Processors of all types (CPUs, GPUs, and accelerator devices) typically implement the latest generation of DDR memory interfaces delivering the highest performance for DRAM memory. CXL enables the connection of an external memory controller device to expand the memory resources available to the processing engines – effectively attaching an extra layer of memory. Data is passed over CXL to an external memory controller that manages the read/write operations for the memory media (usually DRAM). The CXL interface is designed with very low latency, so although the performance is lower than the native DDR interface, it still delivers comparable performance.
CXL connected DDR is a primary application that allows resources to be added or removed as required. A CXL 2.0 switch enables the interconnection and re-provisioning (re-allocation) of multiple processing types and memory devices in response to changing application needs – allowing for larger memory resource deployment without the risk of being underutilized when application changes are needed.
## Advantages of External CXL Memory Controllers
While the latency needs can be met through newer memory technologies – such as DDR4/DDR5 – capable of supporting 3200MT/s and 4800 MT/s respectively, increasing the number of direct-attached DRAM interfaces achieves the increased performance per core needed for high performance computing (HPC) applications. This presents several system design challenges: DRAM interfaces require a lot of device pins for high-speed signals (288), and the issue of electrical reach (distance between devices) presents layout and real estate concerns regarding the close proximity of the processors which can exacerbate localized power and cooling requirements. And while directly connected memory does ensure the highest performance, it is inflexible and does not allow any unused memory capacity to be shared with other processors in the system.
## CXL-based Memory Expansion
The CXL interface is optimized for 16 lanes each at 32 GT/s which substantially reduces the pin count overhead to support memory expansion, easily supporting the necessary data bandwidth for DDR expansion. There are options to support x8 and x4 channel widths at reduced overall bandwidth and the CXL electrical interface is the same as used for PCIe Gen 5 and Gen 6 which ensures reliability and robust channel reach for system designs.
CXL has three protocols that are targeted for specific types of devices. CXL.mem and CXL.cache are load/store memory semantic protocols that ensure the lowest latency for memory access performance. CXL.io is used for device management, error, and status reporting, and is based on PCIe transactions. CXL defines three types of devices with different characteristics:
* Type 1: Host processors
* Type 2: Devices that have processing, caching, and memory sharing functions
* Type 3: Memory devices
DDR Memory controllers are most commonly type-3 functions providing three major features: high performance, capacity expansion, and scalability of deployment.
The use of a CXL-based memory controller – where the CPU would connect with the controller through the CXL interface and with DRAM memory through the DDR interface – meets bandwidth requirements by achieving up to 4 GB/s through a 16 lane CXL interface.
CXL reduces the signal integrity challenges and real-estate footprint issues of direct-attach memory at DDR5 rates due to the reduced pin count needed on the processor side, and also supports more DPCs while at the same time meeting the latency needs of HPC applications.
## Cost Efficiencies of CXL Implementations
CXL enables heterogeneous processing systems with memory expansion and sharing that helps to optimize memory resource utilization. Expensive and complex multi-socket CPU and memory solutions are simplified through CXL memory expansion while overall cost is reduced and application performance is improved.
A CXL-based memory controller solution provides the flexibility of supporting different memory technologies to manage costs and reuse older technologies. For example, DDR4 memories are still predominant in the market and the cost per bit is lower than DDR5. A system could be configured to use high performance DDR5 for direct attached processor memory and DDR4 for the expansion CXL attached memory, with the option to upgrade later.
CXL memory allows for greater physical distance between the processor and the memory devices that can help in the optimization of power utilization and less expensive thermal cooling solutions. Larger scale memory expansion – or disaggregation – enables pools of memory resources to be shared and reallocated on demand.
For memory applications, CXL will likely utilize DIMM (Dual In-Line Memory Module) and EDSFF (Enterprise and Datacenter Standard Form Factor) form factors, as this provides the most efficient density and compatibility. Common form factor adoption ensures interoperability between vendors and allows systems to be easily upgraded or reconfigured.
## Conclusion
CXL based memory solutions provide the best technology for the needs of processing intensive applications such as cloud computing, AI/ML, cluster networks, and HPC as it facilitates:
* Increasing memory bandwidth, capacity, and latency needs for multi-core processors
* Sharing and reallocation of memory to reduce underutilized resources
* Enables heterogeneous systems to serve the applications for innovative AI, ML, and Neural architectures
* Provides a cost effective solution through ease of adoption and reduction in CPU core count
* Increases manageability of thermal and power demands
Consider the implementation of CXL based memory solutions today to enable cost-efficient memory expansion in your future system designs.
Computer server architectures continually evolve to support faster analytics with larger data sets delivered for analysis applications. Computing capabilities of Central Processing Units (CPUs), and Graphics Processing Units (GPUs) increase to enable modern applications such as Artificial Intelligence (AI) and Machine Learning (ML). As core counts for processors increase, there is more data being processed requiring more memory capacity and bandwidth. The diversity of applications means that the CPU core and memory requirements vary widely based on their utilization of memory. System vendors and managers are challenged to optimize system resources to ensure there is enough memory to support all applications while avoiding memory overprovisioning and underutilization.
CXL® (Compute Express Link®) is a memory semantic protocol that leverages the highest performance PCIe® electrical interface specifications to deliver a configurable, scalable, sharable interface for a new generation of platforms. The introduction of the CXL interface is a major step toward enabling expansion and resource sharing by increasing the flexibility of memory and processing allocation to eliminate isolated resources and allow re-provisioning to suit application requirements.
Processors of all types (CPUs, GPUs, and accelerator devices) typically implement the latest generation of DDR memory interfaces delivering the highest performance for DRAM memory. CXL enables the connection of an external memory controller device to expand the memory resources available to the processing engines – effectively attaching an extra layer of memory. Data is passed over CXL to an external memory controller that manages the read/write operations for the memory media (usually DRAM). The CXL interface is designed with very low latency, so although the performance is lower than the native DDR interface, it still delivers comparable performance.
CXL connected DDR is a primary application that allows resources to be added or removed as required. A CXL 2.0 switch enables the interconnection and re-provisioning (re-allocation) of multiple processing types and memory devices in response to changing application needs – allowing for larger memory resource deployment without the risk of being underutilized when application changes are needed.
## Advantages of External CXL Memory Controllers
While the latency needs can be met through newer memory technologies – such as DDR4/DDR5 – capable of supporting 3200MT/s and 4800 MT/s respectively, increasing the number of direct-attached DRAM interfaces achieves the increased performance per core needed for high performance computing (HPC) applications. This presents several system design challenges: DRAM interfaces require a lot of device pins for high-speed signals (288), and the issue of electrical reach (distance between devices) presents layout and real estate concerns regarding the close proximity of the processors which can exacerbate localized power and cooling requirements. And while directly connected memory does ensure the highest performance, it is inflexible and does not allow any unused memory capacity to be shared with other processors in the system.
## CXL-based Memory Expansion
The CXL interface is optimized for 16 lanes each at 32 GT/s which substantially reduces the pin count overhead to support memory expansion, easily supporting the necessary data bandwidth for DDR expansion. There are options to support x8 and x4 channel widths at reduced overall bandwidth and the CXL electrical interface is the same as used for PCIe Gen 5 and Gen 6 which ensures reliability and robust channel reach for system designs.
CXL has three protocols that are targeted for specific types of devices. CXL.mem and CXL.cache are load/store memory semantic protocols that ensure the lowest latency for memory access performance. CXL.io is used for device management, error, and status reporting, and is based on PCIe transactions. CXL defines three types of devices with different characteristics:
* Type 1: Host processors
* Type 2: Devices that have processing, caching, and memory sharing functions
* Type 3: Memory devices
DDR Memory controllers are most commonly type-3 functions providing three major features: high performance, capacity expansion, and scalability of deployment.
The use of a CXL-based memory controller – where the CPU would connect with the controller through the CXL interface and with DRAM memory through the DDR interface – meets bandwidth requirements by achieving up to 4 GB/s through a 16 lane CXL interface.
CXL reduces the signal integrity challenges and real-estate footprint issues of direct-attach memory at DDR5 rates due to the reduced pin count needed on the processor side, and also supports more DPCs while at the same time meeting the latency needs of HPC applications.
## Cost Efficiencies of CXL Implementations
CXL enables heterogeneous processing systems with memory expansion and sharing that helps to optimize memory resource utilization. Expensive and complex multi-socket CPU and memory solutions are simplified through CXL memory expansion while overall cost is reduced and application performance is improved.
A CXL-based memory controller solution provides the flexibility of supporting different memory technologies to manage costs and reuse older technologies. For example, DDR4 memories are still predominant in the market and the cost per bit is lower than DDR5. A system could be configured to use high performance DDR5 for direct attached processor memory and DDR4 for the expansion CXL attached memory, with the option to upgrade later.
CXL memory allows for greater physical distance between the processor and the memory devices that can help in the optimization of power utilization and less expensive thermal cooling solutions. Larger scale memory expansion – or disaggregation – enables pools of memory resources to be shared and reallocated on demand.
For memory applications, CXL will likely utilize DIMM (Dual In-Line Memory Module) and EDSFF (Enterprise and Datacenter Standard Form Factor) form factors, as this provides the most efficient density and compatibility. Common form factor adoption ensures interoperability between vendors and allows systems to be easily upgraded or reconfigured.
## Conclusion
CXL based memory solutions provide the best technology for the needs of processing intensive applications such as cloud computing, AI/ML, cluster networks, and HPC as it facilitates:
* Increasing memory bandwidth, capacity, and latency needs for multi-core processors
* Sharing and reallocation of memory to reduce underutilized resources
* Enables heterogeneous systems to serve the applications for innovative AI, ML, and Neural architectures
* Provides a cost effective solution through ease of adoption and reduction in CPU core count
* Increases manageability of thermal and power demands
Consider the implementation of CXL based memory solutions today to enable cost-efficient memory expansion in your future system designs.