MEMORY INDUSTRY INTELLIGENCE
CXL Type 2: 액티브 메모리 계층화 및 근접 메모리 가속기 사용 사례 - Compute Express Link
한국어 번역·요약·분석
원문 제목: CXL Type 2: Use Cases for Active Memory Tiering and Near Memory Accelerators - Compute Express Link
핵심 요약
Altera의 기술 리더 Divya Vijayaraghavan이 CXL 생태계의 두 주요 사용 사례인 액티브 메모리 계층화와 근접 메모리 컴퓨트 가속을 설명한다. 액티브 메모리 계층화는 핫/콜드 페이지를 로컬·원격 메모리 계층 간 마이그레이션하는 기법으로, 핫 페이지 분류 정확도와 마이그레이션 속도, 하드웨어/소프트웨어 분담이 과제로 남아 있으며 업계는 핫 페이지 감지를 하드웨어로 오프로드하는 추세다. 근접 메모리 가속은 CPU 기능을 CXL 디바이스의 EMIF 컨트롤러 근처에서 처리하는 것으로, 어떤 워크로드를 가속할지, SDK 구현, 컴퓨테이셔널 메모리의 높은 비용이 과제이며 FPGA를 활용해 실험 유연성을 확보하고 TCO를 낮출 수 있다고 주장한다. Altera Agilex 7 FPGA에서 구현된 계층화 솔루션에서 호스트 DRAM이 CPU 수요의 약 50%만 충족하던 상황이 페이지 분류 후 40%의 호스트 DRAM을 확보해 개선되었고, KSM 기능을 CXL Type 2 디바이스로 오프로드한 사례에서 테일 레이턴시가 83% 감소했으며, CXL.cache 레이턴시가 PCIe 대비 68% 낮게 측정되었다고 제시한다. 이는 CXL 프로토콜이 캐시 가능 읽기/쓰기를 통해 호스트-디바이스 간 통신 레이턴시를 낮추고 커널 호출·DMA 설정을 회피한다는 주장이다.
메모리 산업 영향 분석
이 원문은 CXL Type 2의 액티브 메모리 계층화와 근접 메모리 가속 사용 사례를 다루며, 메모리 산업에는 간접적·조건부 영향이 있다. 원문에서 직접 확인되는 것은 CXL 프로토콜이 캐시 가능 읽기/쓰기로 호스트-디바이스 간 레이턴시를 낮추고 커널 호출·DMA 설정을 회피한다는 주장, Altera Agilex 7 FPGA에서 호스트 DRAM의 약 50%만 CPU를 서비스하던 상황이 페이지 분류 후 40%의 호스트 DRAM을 확보했다는 예시, KSM 오프로드 시 테일 레이턴시 83% 감소, CXL.cache 레이턴시가 PCIe 대비 68% 낮다는 측정 예시다. 분석가 가설로는 계층화가 확산되면 서버 DDR의 계층 구성(고가·저가 DRAM 혼합)과 CXL 확장 메모리 모듈의 채택 조건이 바뀔 수 있으나, 원문은 특정 DRAM 규격·모듈 형태·고객 인증·양산 일정을 명시하지 않으므로 공급 배분·고객 인증·제품 믹스·투자 일정 중 어느 선택이 실제로 바뀌는지는 미확인이다. 대안으로 CXL 계층화가 서버 DDR 수요를 늘리는 방향과, 페이지 분류·오프로드로 기존 DRAM 사용 효율이 높아져 단위 서버당 DRAM 비트 수요 증가가 둔화되는 상충 효과가 모두 가능하다. 판단 변경 조건은 CXL Type 2 디바이스의 실제 양산·고객 채택·표준 인증이 확인되는 경우이며, 그 전까지는 메모리 산업과의 직접 연결 근거가 부족하다. research_topics의 표준·밸류체인 질문에 대해 원문은 CXL IP·FPGA 설계 예제 수준의 공급사 기술 자료만 제공하고 실제 제품 양산·채택 사례는 확인되지 않아 미확인으로 남긴다.
한국어 번역 읽기
수집된 원문 v1의 전체 본문 기준 · 5363자
작성자: Divya Vijayaraghavan, Altera 기술 리더
CXL 생태계가 발전함에 따라 여러 사용 사례가 등장했으며, 그중 두 가지 두드러진 선두 주자는 액티브 메모리 계층화와 근접 메모리 컴퓨트 가속입니다.
1. 액티브 메모리 계층화는 시스템 내 로컬 메모리 계층과 원격 메모리 계층 사이에서 핫 페이지와 콜드 페이지를 마이그레이션하는 기법을 활용합니다. 핫이라는 용어는 자주 접근됨과 동의어입니다.
2. 근접 메모리 컴퓨트 가속은 컴퓨테이셔널 메모리라고도 하며, 시스템 내 원격 메모리 근처에서 수행되는 처리 또는 가속을 설명하는 데 사용됩니다.
업계에서 액티브 메모리 계층화를 구현하는 접근 방식은 소프트웨어 주도, 하드웨어 주도, 그리고 두 가지의 조합 등 여러 가지가 있지만, 보편적인 과제는 남아 있습니다. 핫 페이지 분류는 얼마나 정확한가, 시스템이 분류에 따라 얼마나 빠르게 페이지를 마이그레이션할 수 있는가, 어떤 워크로드가 계층화의 이점을 얻는가, 어떤 기능이 하드웨어와 소프트웨어 중 어디에 구현되는가? 일반적으로 시스템은 핫 페이지 감지를 하드웨어로 더 많이 오프로드하는 추세로 가고 있으며, 그에 따라 cxl.mem HDM(Host Managed Device Memory) 인터페이스에서 향상된 메모리 접근 보고 기능이 제공됩니다.
근접 메모리 컴퓨트 가속 사용 사례에서는 일반적으로 CPU에서 기능이 오프로드되어 CXL 디바이스의 EMIF(External Memory Interface) 컨트롤러 근처에서 처리 또는 가속됩니다. 과제로는 어떤 워크로드 또는 기능을 가속할지 결정하는 것, 가속 기능을 위한 SDK(Software Development Kit) 구현, 컴퓨테이셔널 메모리의 높은 비용 등이 있습니다. 이러한 과제를 완화하는 한 가지 방법은 FPGA를 활용해 디바이스 연결 메모리 근처에서 기능을 가속함으로써, 어떤 기능이 성능 개선에 가장 적합한지 워크로드를 실험할 수 있는 유연성을 제공하는 것입니다. 메모리 계층화와 컴퓨테이셔널 메모리를 결합하면 시스템 TCO(총소유비용)를 낮출 수 있습니다.
아래 그림은 Altera의 Agilex 7 FPGA에 구현된 메모리 계층화 솔루션을 보여줍니다. 왼쪽에서는 호스트 DRAM이 CPU의 약 50%만 서비스할 수 있고 나머지 CPU는 메모리 부족 상태입니다! 오른쪽에서는 페이지가 분류되었습니다. 자주 사용되지 않는 페이지는 더 저렴한 메모리 계층에 할당되어 호스트 DRAM의 40%를 확보하며, 이는 이전에 메모리 부족 상태였던 CPU 부분을 서비스하는 데 사용될 수 있습니다.
CXL Type 2 근접 메모리 가속이 기존 PCIe® 대비 갖는 레이턴시 이점을 보여주는 공개된 성능 지표가 등장했습니다. 아래 표는 커널 동일 페이지 병합(KSM) 기능을 CPU에서 CXL Type 2 디바이스로 오프로드한 결과 테일 레이턴시가 83% 낮아졌음을 보여주며, 여기서 테일 레이턴시는 시스템 레이턴시 분포의 99번째 백분위수를 의미합니다.
아래 표는 측정된 CXL.cache 레이턴시가 PCIe보다 68% 낮은 예를 보여줍니다. CXL.cache 측정은 CPU의 최종 레벨 캐시(LLC)에서 적중을 가정하고, PCIe 측정은 읽기 데이터가 호스트 연결 DRAM에서 얻어진다고 가정합니다.
위 데이터 포인트에서 보듯이 CXL 프로토콜은 호스트, 디바이스 연결 메모리, 가속기 간에 캐시 가능한 읽기와 쓰기를 도입함으로써 호스트와 가속기 디바이스 간의 더 낮은 레이턴시 통신을 가능하게 합니다. 낮은 레이턴시 외에도 CXL은 PCIe에서 필요한 커널 호출과 DMA 설정을 회피함으로써 통신을 단순화합니다.
요약하면, CXL 기반 메모리 계층화와 근접 메모리 가속은 시스템 TCO 절감, CPU에서의 워크로드 처리 오프로드, 특정 워크로드의 처리 레이턴시 감소와 같은 이점을 제공합니다.
일부 과제는 FPGA 기반 솔루션으로 완화할 수 있습니다. Altera FPGA는 구성 가능한 사전 구축 가속기 기능과 함께 CXL IP 및 설계 예제를 제공합니다. FPGA의 동적 재구성 기능은 워크로드 분할과 CPU로부터의 오프로드를 실험할 수 있게 합니다. Altera FPGA가 제공하는 추가 가치 기능으로는 하드 프로세서 시스템, 그리고 PCIe 5.0/6.0 연결로 CXL을 랙 너머로 확장할 수 있는 100-800 GbE 기능이 있습니다.
더 알아보려면 Altera FPGA 팀([https://www.altera.com/contact.html](https://www.altera.com/contact.html))에 문의하십시오.
**참고문헌**
(1) Meta, UMichigan – TPP: CXL 지원 계층형 메모리를 위한 투명 페이지 배치: [https://arxiv.org/abs/2206.02878](https://arxiv.org/abs/2206.02878)
(2) Google – [https://doi.org/10.1145/3582016.3582031](https://doi.org/10.1145/3582016.3582031)
(3) SKHynix – [https://www.youtube.com/watch?v=pbnTlY41h08](https://www.youtube.com/watch?v=pbnTlY41h08)
(4) SmartModular – https://www.youtube.com/watch?v=A_PML20fk-Y
(5) VMWare – https://www.youtube.com/watch?v=4tX8wJ-BJj8&list=PL0pU5hg9yniY-WFJ-uvAXKWVVaXXHLR7O&index=12
(6) Unifabrix – [https://www.unifabrix.com/](https://www.unifabrix.com/)
**작성자**
Divya Vijayaraghavan은 San Jose에 있는 Altera의 기술 리더입니다. 그녀는 경력 동안 UPI 및 CXL 고객 지원과 확산을 위한 버티컬 챔피언 및 주제 전문가, FPGA 가속 솔루션 기술 리드, 여러 주요 고객 및 생태계 파트너를 위한 기술 매니저 등 다양한 역할을 맡았습니다. 그녀는 26개의 등록 특허를 보유하고 있으며 PCI-SIG와 같은 산업 표준 위원회에서 Altera를 대표했습니다.
CXL 생태계가 발전함에 따라 여러 사용 사례가 등장했으며, 그중 두 가지 두드러진 선두 주자는 액티브 메모리 계층화와 근접 메모리 컴퓨트 가속입니다.
1. 액티브 메모리 계층화는 시스템 내 로컬 메모리 계층과 원격 메모리 계층 사이에서 핫 페이지와 콜드 페이지를 마이그레이션하는 기법을 활용합니다. 핫이라는 용어는 자주 접근됨과 동의어입니다.
2. 근접 메모리 컴퓨트 가속은 컴퓨테이셔널 메모리라고도 하며, 시스템 내 원격 메모리 근처에서 수행되는 처리 또는 가속을 설명하는 데 사용됩니다.
업계에서 액티브 메모리 계층화를 구현하는 접근 방식은 소프트웨어 주도, 하드웨어 주도, 그리고 두 가지의 조합 등 여러 가지가 있지만, 보편적인 과제는 남아 있습니다. 핫 페이지 분류는 얼마나 정확한가, 시스템이 분류에 따라 얼마나 빠르게 페이지를 마이그레이션할 수 있는가, 어떤 워크로드가 계층화의 이점을 얻는가, 어떤 기능이 하드웨어와 소프트웨어 중 어디에 구현되는가? 일반적으로 시스템은 핫 페이지 감지를 하드웨어로 더 많이 오프로드하는 추세로 가고 있으며, 그에 따라 cxl.mem HDM(Host Managed Device Memory) 인터페이스에서 향상된 메모리 접근 보고 기능이 제공됩니다.
근접 메모리 컴퓨트 가속 사용 사례에서는 일반적으로 CPU에서 기능이 오프로드되어 CXL 디바이스의 EMIF(External Memory Interface) 컨트롤러 근처에서 처리 또는 가속됩니다. 과제로는 어떤 워크로드 또는 기능을 가속할지 결정하는 것, 가속 기능을 위한 SDK(Software Development Kit) 구현, 컴퓨테이셔널 메모리의 높은 비용 등이 있습니다. 이러한 과제를 완화하는 한 가지 방법은 FPGA를 활용해 디바이스 연결 메모리 근처에서 기능을 가속함으로써, 어떤 기능이 성능 개선에 가장 적합한지 워크로드를 실험할 수 있는 유연성을 제공하는 것입니다. 메모리 계층화와 컴퓨테이셔널 메모리를 결합하면 시스템 TCO(총소유비용)를 낮출 수 있습니다.
아래 그림은 Altera의 Agilex 7 FPGA에 구현된 메모리 계층화 솔루션을 보여줍니다. 왼쪽에서는 호스트 DRAM이 CPU의 약 50%만 서비스할 수 있고 나머지 CPU는 메모리 부족 상태입니다! 오른쪽에서는 페이지가 분류되었습니다. 자주 사용되지 않는 페이지는 더 저렴한 메모리 계층에 할당되어 호스트 DRAM의 40%를 확보하며, 이는 이전에 메모리 부족 상태였던 CPU 부분을 서비스하는 데 사용될 수 있습니다.
CXL Type 2 근접 메모리 가속이 기존 PCIe® 대비 갖는 레이턴시 이점을 보여주는 공개된 성능 지표가 등장했습니다. 아래 표는 커널 동일 페이지 병합(KSM) 기능을 CPU에서 CXL Type 2 디바이스로 오프로드한 결과 테일 레이턴시가 83% 낮아졌음을 보여주며, 여기서 테일 레이턴시는 시스템 레이턴시 분포의 99번째 백분위수를 의미합니다.
아래 표는 측정된 CXL.cache 레이턴시가 PCIe보다 68% 낮은 예를 보여줍니다. CXL.cache 측정은 CPU의 최종 레벨 캐시(LLC)에서 적중을 가정하고, PCIe 측정은 읽기 데이터가 호스트 연결 DRAM에서 얻어진다고 가정합니다.
위 데이터 포인트에서 보듯이 CXL 프로토콜은 호스트, 디바이스 연결 메모리, 가속기 간에 캐시 가능한 읽기와 쓰기를 도입함으로써 호스트와 가속기 디바이스 간의 더 낮은 레이턴시 통신을 가능하게 합니다. 낮은 레이턴시 외에도 CXL은 PCIe에서 필요한 커널 호출과 DMA 설정을 회피함으로써 통신을 단순화합니다.
요약하면, CXL 기반 메모리 계층화와 근접 메모리 가속은 시스템 TCO 절감, CPU에서의 워크로드 처리 오프로드, 특정 워크로드의 처리 레이턴시 감소와 같은 이점을 제공합니다.
일부 과제는 FPGA 기반 솔루션으로 완화할 수 있습니다. Altera FPGA는 구성 가능한 사전 구축 가속기 기능과 함께 CXL IP 및 설계 예제를 제공합니다. FPGA의 동적 재구성 기능은 워크로드 분할과 CPU로부터의 오프로드를 실험할 수 있게 합니다. Altera FPGA가 제공하는 추가 가치 기능으로는 하드 프로세서 시스템, 그리고 PCIe 5.0/6.0 연결로 CXL을 랙 너머로 확장할 수 있는 100-800 GbE 기능이 있습니다.
더 알아보려면 Altera FPGA 팀([https://www.altera.com/contact.html](https://www.altera.com/contact.html))에 문의하십시오.
**참고문헌**
(1) Meta, UMichigan – TPP: CXL 지원 계층형 메모리를 위한 투명 페이지 배치: [https://arxiv.org/abs/2206.02878](https://arxiv.org/abs/2206.02878)
(2) Google – [https://doi.org/10.1145/3582016.3582031](https://doi.org/10.1145/3582016.3582031)
(3) SKHynix – [https://www.youtube.com/watch?v=pbnTlY41h08](https://www.youtube.com/watch?v=pbnTlY41h08)
(4) SmartModular – https://www.youtube.com/watch?v=A_PML20fk-Y
(5) VMWare – https://www.youtube.com/watch?v=4tX8wJ-BJj8&list=PL0pU5hg9yniY-WFJ-uvAXKWVVaXXHLR7O&index=12
(6) Unifabrix – [https://www.unifabrix.com/](https://www.unifabrix.com/)
**작성자**
Divya Vijayaraghavan은 San Jose에 있는 Altera의 기술 리더입니다. 그녀는 경력 동안 UPI 및 CXL 고객 지원과 확산을 위한 버티컬 챔피언 및 주제 전문가, FPGA 가속 솔루션 기술 리드, 여러 주요 고객 및 생태계 파트너를 위한 기술 매니저 등 다양한 역할을 맡았습니다. 그녀는 26개의 등록 특허를 보유하고 있으며 PCI-SIG와 같은 산업 표준 위원회에서 Altera를 대표했습니다.
브리프용 요약 초안
Altera가 CXL Type 2의 액티브 메모리 계층화와 근접 메모리 가속 사용 사례를 정리하며, Agilex 7 FPGA에서 호스트 DRAM 40% 확보, KSM 오프로드 시 테일 레이턴시 83% 감소, CXL.cache 레이턴시 PCIe 대비 68% 감소 예시를 제시했다. 다만 특정 DRAM 규격·모듈·고객 채택·양산 일정은 확인되지 않아 메모리 산업 직접 영향은 조건부다.
원문 텍스트
원문 열기 ↗By: Divya Vijayaraghavan, Technical Leader, Altera
As the CXL ecosystem evolves, multiple use cases have emerged with two prominent frontrunners – active memory tiering and near memory compute acceleration.
1. Active memory tiering utilizes schemes to migrate hot and cold pages between local and remote memory tiers in a system. The term hot is synonymous with frequently accessed.
2. Near memory compute acceleration, also called computational memory, is used to describe processing or acceleration performed near the remote memory in a system.
There are multiple approaches to implementing active memory tiering in the industry – software-driven, hardware-driven, and a combination of the two, but universal challenges remain. How accurate is the hot-page classification, how quickly can a system act on the classification and migrate pages, which workloads benefit from tiering, and which functions are implemented in hardware vs. software? In general, systems have been trending towards an increased offload of hot-page detection to hardware and enhanced memory access reporting capabilities are consequently provided on the cxl.mem HDM (Host Managed Device Memory) interface.
In the near memory compute acceleration use case, a function is typically offloaded from the CPU and processed or accelerated close to the EMIF (External Memory Interface) controller on the CXL device. Challenges include determining which workload or function to accelerate, the implementation of a Software Development Kit (SDK) for the accelerated function, and the prohibitive cost of computational memory. One way to mitigate these challenges is to utilize FPGAs to accelerate functions near the device-attached memory, providing the flexibility to experiment with workloads to determine which functions are most conducive to performance improvements. System TCO (Total Cost of Ownership) could be reduced by combining memory tiering and computational memory.
The figures below show a memory tiering solution implemented on Altera’s Agilex 7 FPGA. On the left, the host DRAM is able to service only about 50% of the CPU, with the rest of the CPU starved for memory! On the right, pages have been classified. Infrequently used pages are assigned to a cheaper memory tier, freeing up 40% of the host DRAM, which can be used to service the portion of the CPU that was previously starved for memory.
Publicly disclosed performance metrics have emerged, illustrating the latency advantage of CXL Type 2 near memory acceleration vs. traditional PCIe®. The table below shows the offloading of the Kernel Same-Page Merging (KSM) feature from the CPU to a CXL Type 2 device, resulting in an 83% lower tail latency, where the term tail latency refers to the 99th percentile of the latency distribution of a system.
The table below shows an example in which the measured CXL.cache latency is 68% lower than PCIe. The CXL.cache measurement assumes a hit in the CPU’s Last Level Cache (LLC), and the PCIe measurement assumes that read data is obtained from the host-attached DRAM.
As illustrated by the data points above, the CXL protocol allows for lower latency communication between the host and the accelerator device by introducing cacheable reads and writes between the host, device-attached memory, and accelerator. In addition to lower latency, CXL also simplifies communication by avoiding the kernel calls and DMA setups that are required by PCIe.
In summary, CXL-based memory tiering and near memory acceleration provide advantages such as reducing system TCO, offloading workload processing from the CPU, and decreasing processing latency for specific workloads.
Some challenges can be mitigated by FPGA-based solutions. Altera FPGAs provide CXL IP and design examples with configurable pre-built accelerator functions. The FPGA’s dynamic reconfigurability capability enables experimentation with workload partitioning and offloading from the CPU. Additional value-added features provided by Altera FPGAs include a Hard Processor System, and 100-800 GbE capability to expand CXL beyond the rack, with PCIe 5.0/6.0 connectivity.
To learn more, contact the Altera FPGA team at [https://www.altera.com/contact.html](https://www.altera.com/contact.html).
**References**
(1) Meta, UMichigan – TPP: Transparent Page Placement for CXL-enabled Tiered-Memory: [https://arxiv.org/abs/2206.02878](https://arxiv.org/abs/2206.02878)
(2) Google – [https://doi.org/10.1145/3582016.3582031](https://doi.org/10.1145/3582016.3582031)
(3) SKHynix – [https://www.youtube.com/watch?v=pbnTlY41h08](https://www.youtube.com/watch?v=pbnTlY41h08)
(4) SmartModular – https://www.youtube.com/watch?v=A_PML20fk-Y
(5) VMWare – https://www.youtube.com/watch?v=4tX8wJ-BJj8&list=PL0pU5hg9yniY-WFJ-uvAXKWVVaXXHLR7O&index=12
(6) Unifabrix – [https://www.unifabrix.com/](https://www.unifabrix.com/)
**Author**
Divya Vijayaraghavan is a Technical Leader at Altera in San Jose. She has worn many hats in her career ranging from being a vertical champion and subject matter expert for UPI and CXL customer enablement and proliferation to a technical lead for FPGA acceleration solutions to a technical manager for several key customers and ecosystem partners. She has 26 granted patents and has represented Altera on industry standards committees such as the PCI-SIG.
As the CXL ecosystem evolves, multiple use cases have emerged with two prominent frontrunners – active memory tiering and near memory compute acceleration.
1. Active memory tiering utilizes schemes to migrate hot and cold pages between local and remote memory tiers in a system. The term hot is synonymous with frequently accessed.
2. Near memory compute acceleration, also called computational memory, is used to describe processing or acceleration performed near the remote memory in a system.
There are multiple approaches to implementing active memory tiering in the industry – software-driven, hardware-driven, and a combination of the two, but universal challenges remain. How accurate is the hot-page classification, how quickly can a system act on the classification and migrate pages, which workloads benefit from tiering, and which functions are implemented in hardware vs. software? In general, systems have been trending towards an increased offload of hot-page detection to hardware and enhanced memory access reporting capabilities are consequently provided on the cxl.mem HDM (Host Managed Device Memory) interface.
In the near memory compute acceleration use case, a function is typically offloaded from the CPU and processed or accelerated close to the EMIF (External Memory Interface) controller on the CXL device. Challenges include determining which workload or function to accelerate, the implementation of a Software Development Kit (SDK) for the accelerated function, and the prohibitive cost of computational memory. One way to mitigate these challenges is to utilize FPGAs to accelerate functions near the device-attached memory, providing the flexibility to experiment with workloads to determine which functions are most conducive to performance improvements. System TCO (Total Cost of Ownership) could be reduced by combining memory tiering and computational memory.
The figures below show a memory tiering solution implemented on Altera’s Agilex 7 FPGA. On the left, the host DRAM is able to service only about 50% of the CPU, with the rest of the CPU starved for memory! On the right, pages have been classified. Infrequently used pages are assigned to a cheaper memory tier, freeing up 40% of the host DRAM, which can be used to service the portion of the CPU that was previously starved for memory.
Publicly disclosed performance metrics have emerged, illustrating the latency advantage of CXL Type 2 near memory acceleration vs. traditional PCIe®. The table below shows the offloading of the Kernel Same-Page Merging (KSM) feature from the CPU to a CXL Type 2 device, resulting in an 83% lower tail latency, where the term tail latency refers to the 99th percentile of the latency distribution of a system.
The table below shows an example in which the measured CXL.cache latency is 68% lower than PCIe. The CXL.cache measurement assumes a hit in the CPU’s Last Level Cache (LLC), and the PCIe measurement assumes that read data is obtained from the host-attached DRAM.
As illustrated by the data points above, the CXL protocol allows for lower latency communication between the host and the accelerator device by introducing cacheable reads and writes between the host, device-attached memory, and accelerator. In addition to lower latency, CXL also simplifies communication by avoiding the kernel calls and DMA setups that are required by PCIe.
In summary, CXL-based memory tiering and near memory acceleration provide advantages such as reducing system TCO, offloading workload processing from the CPU, and decreasing processing latency for specific workloads.
Some challenges can be mitigated by FPGA-based solutions. Altera FPGAs provide CXL IP and design examples with configurable pre-built accelerator functions. The FPGA’s dynamic reconfigurability capability enables experimentation with workload partitioning and offloading from the CPU. Additional value-added features provided by Altera FPGAs include a Hard Processor System, and 100-800 GbE capability to expand CXL beyond the rack, with PCIe 5.0/6.0 connectivity.
To learn more, contact the Altera FPGA team at [https://www.altera.com/contact.html](https://www.altera.com/contact.html).
**References**
(1) Meta, UMichigan – TPP: Transparent Page Placement for CXL-enabled Tiered-Memory: [https://arxiv.org/abs/2206.02878](https://arxiv.org/abs/2206.02878)
(2) Google – [https://doi.org/10.1145/3582016.3582031](https://doi.org/10.1145/3582016.3582031)
(3) SKHynix – [https://www.youtube.com/watch?v=pbnTlY41h08](https://www.youtube.com/watch?v=pbnTlY41h08)
(4) SmartModular – https://www.youtube.com/watch?v=A_PML20fk-Y
(5) VMWare – https://www.youtube.com/watch?v=4tX8wJ-BJj8&list=PL0pU5hg9yniY-WFJ-uvAXKWVVaXXHLR7O&index=12
(6) Unifabrix – [https://www.unifabrix.com/](https://www.unifabrix.com/)
**Author**
Divya Vijayaraghavan is a Technical Leader at Altera in San Jose. She has worn many hats in her career ranging from being a vertical champion and subject matter expert for UPI and CXL customer enablement and proliferation to a technical lead for FPGA acceleration solutions to a technical manager for several key customers and ecosystem partners. She has 26 granted patents and has represented Altera on industry standards committees such as the PCI-SIG.