MEMORY INDUSTRY INTELLIGENCE

생산적이고 내구성 있으며 범용적인: NVIDIA AI 팩토리가 투자 수익을 극대화하는 방법

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v2 · 10.09 23:17사용자 검토 전 초안

원문 제목: Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment

핵심 요약

NVIDIA는 AI 팩토리 투자 수익이 수익 창출 능력, 유효 수명, 수요라는 세 가지 요소에 의해 결정된다고 주장한다. 회사는 자사 플랫폼이 메가와트당 토큰 처리량을 극대화하고(생산적), 기존 하드웨어가 수년간 수익을 창출하며(내구성), 모든 유형의 AI 및 비AI 워크로드를 실행할 수 있다(범용적)고 설명한다. 근거로 SemiAnalysis AgentX의 Vera Rubin NVL72 대 GB300 NVL72 성능 비교, A100의 6년 후 상업적 사용, CoreWeave의 2029년까지 예약 연장, 여러 운영사의 감가상각 일정 연장, Barkr·Silicon Data·Ornn Data의 중고 가치 및 임대료 데이터를 제시한다. Lilly, Pinterest, Revolut, Runway, Texas A&M, Cosm, Dassault Systèmes, Unilever 등의 고객 사례를 통해 범용성을 입증한다. 이 문서는 NVIDIA의 마케팅 성격 글이다.

메모리 산업 영향 분석

이 문서는 NVIDIA의 AI 팩토리 마케팅 글로, 메모리 산업에 대한 직접적 언급은 '컴퓨팅, 네트워킹 및 메모리'라는 한 구절에 불과하다. AI 팩토리 확대는 HBM, 서버 DRAM, eSSD 등 데이터센터 메모리 수요와 간접적으로 연결될 수 있으나, 본문은 특정 메모리 제품·용량·고객·거래를 명시하지 않는다. Vera Rubin NVL72의 처리량·비용 개선은 가속기 세대 변화가 메모리 계층 이동(HBM→HBF 등)을 유발한다는 근거로 쓰기 어렵다. A100의 장수명과 감가상각 연장은 기존 서버 DRAM·HBM 수요의 교체 주기 연장 가능성을 시사하나, 이는 분석가 가설이며 원문은 메모리 수요에 대해 직접 논하지 않는다. 확인할 지표: NVIDIA 데이터센터 매출, HBM 공급사 인증·물량, 서버 DRAM 탑재량, eSSD 채택률. 반대 근거: 토큰 비용 하락이 컴퓨팅 수요를 늘린다는 주장은 검증되지 않았으며, 메모리 산업과의 직접 연결 근거는 부족하다.
한국어 번역 읽기

수집된 원문 v2의 전체 본문 기준 · 9588자

AI 팩토리는 메가와트 단위로, 심지어 기가와트 단위로 건설됩니다. 각 메가와트 팩토리는 약 6천만 달러가 들며, AI 팩토리 운영자는 투자 수익에 대한 명확한 전망이 있어야만 그러한 규모의 자본을 투입합니다. AI 팩토리 수익을 결정하는 세 가지 핵심 요소는 다음과 같습니다:

1. **수익 창출 능력:** 팩토리가 생산할 수 있는 모든 토큰을 판매한다면 1년에 벌 수 있는 금액.
2. **유효 수명**: AI 하드웨어가 계속 수익을 창출하는 기간.
3. **수요**: 해당 토큰에 대한 수요의 크기.

강점이 다른 약점을 완전히 상쇄할 수는 없습니다. 높은 수익 창출 능력은 팩토리가 생산할 수 있는 것의 일부만 판매한다면 거의 의미가 없습니다. 높은 수요도 1년 만에 최대 용량으로 생산을 중단하면 별로 중요하지 않습니다. 또한 이 세 가지는 독립적이지 않습니다. 더 많은 종류의 워크로드를 실행할 수 있는 팩토리는 더 많은 수요를 찾아 매년 수익을 창출합니다.

NVIDIA AI 팩토리는 이 세 가지를 모두 극대화하도록 설계되었습니다. 그것들은:

* **생산적**: 메가와트당 최고 처리량과 토큰당 최저 비용을 제공하여 **수익 창출 능력을 극대화**합니다.
* **내구성**: NVIDIA GPU와 시스템은 출하 후 수년간 계속 수익을 창출하여 **유효 수명을 연장**합니다.
* **범용적**: 모든 유형의 AI를 — 모든 단계와 모든 장소에서 — 실행할 뿐만 아니라 AI와 관련 없는 많은 워크로드도 실행하여 **서비스할 수 있는 수요를 심화하고 확장**합니다.



NVIDIA 플랫폼은 생산적이고 내구성 있으며 범용적이어서 AI 팩토리 수익을 극대화합니다.

전체 스택에 걸친 엔지니어링 코드설계는 AI 팩토리 처리량을 극대화하고, 지속적인 소프트웨어 최적화는 설치된 하드웨어가 출하 후 수년간 생산성을 유지하도록 합니다. [NVIDIA CUDA-X](https://www.nvidia.com/en-us/technologies/cuda-x/) 라이브러리를 통해 팩토리는 모든 가속 워크로드를 실행할 수 있습니다. 표준화된 아키텍처는 이 모든 것을 모든 운영자가 접근할 수 있게 하며, 검증된 참조 설계에서 배포할 수 있습니다.

## **생산적: 메가와트당 최고 토큰 수와 최저 토큰 비용**

전력은 AI 팩토리의 구속 조건입니다. 이로 인해 메가와트당 초당 토큰 수가 수익 창출 능력을 좌우하는 숫자가 됩니다. 고정된 전력 범위 내에서 더 많은 토큰은 더 많은 수익을 의미합니다. 토큰당 비용이 낮을수록 그에 대한 마진이 더 커집니다.

[SemiAnalysis AgentX](https://newsletter.semianalysis.com/p/vera-rubin-nvl72-agentic-inference) 데이터에 따르면 NVIDIA Vera Rubin NVL72 시스템은 NVIDIA GB300 NVL72보다 메가와트당 처리량이 30배 이상 높고, DeepSeek V4 Pro 모델에서 백만 토큰당 비용이 최대 45배 낮습니다. 이러한 규모의 이득은 전체 스택에 걸친 [극단적 코드설계](https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/)에서 비롯됩니다: 모델과 워크로드에서 소프트웨어를 거쳐 컴퓨팅, 네트워킹, 메모리에 이르기까지 모두 함께 최적화됩니다.



GLM 5.3에 대한 NVIDIA GB300 NVL72 성능에 관한 SemiAnalysis AgentX 분석.

두 가지 질문이 따릅니다. 모든 세대가 토큰을 극적으로 저렴하게 만든다면 컴퓨팅 수요가 줄어들까요?

아니요, 확장됩니다.

더 저렴한 토큰은 더 많은 사용 사례를 경제적으로 만들고, 그 사용 사례는 절약된 효율성보다 더 많은 토큰을 소비합니다.

두 번째 질문은 내구성에 관한 것입니다. 각 세대가 이전 세대보다 훨씬 뛰어나다면 이전 세대는 어떻게 될까요?

## **내구성: 설치 기반이 계속 수익을 창출**

모든 워크로드에 최신 시스템이 필요한 것은 아닙니다. 적합한 선택은 워크로드의 복잡성과 형태에 따라 달라지므로, 다음 세대가 도착한 후에도 이전 세대가 계속 수익을 창출합니다.

NVIDIA A100 GPU는 2020년에 출하되어 6년 후인 지금도 상업적으로 사용되며 지속적인 경제적 가치를 입증합니다. CoreWeave는 최근 2020년에 처음 출시된 장치의 예약을 2029년까지 연장했습니다. 수년에 걸쳐 모든 주요 운영자는 서버의 감가상각 일정을 연장했는데, 이는 하드웨어가 언제 수익 창출을 멈추는지에 대한 추측이며, 계속 늘어나고 있습니다. 2026년 9월 Sprout 분석, "[데이터 센터 GPU의 생산 수명](https://www.sproutup.com/resources/blog/how-long-does-a-data-center-gpu-actually-last)"은 모든 주요 운영자에서 이 일정이 어떻게 변화했는지 추적합니다.



모든 주요 운영자가 서버 수명을 연장했습니다. 출처: Sprout, "데이터 센터 GPU의 생산 수명", 2026년 9월. Sprout이 수집한 회사 공시 및 언론 보도를 기반으로 한 데이터. 회계 수명은 물리적 수명의 보수적 대리 지표입니다 — Microsoft의 NVIDIA V100 플릿은 6년 장부 수명에 대해 8.4년을 운영했습니다.

[Barkr](https://barkr.ai/market-report#the-resale-standard-forward-looking-gpu-valuations)는 8-GPU H100 시스템의 유효 수명을 5~6년, GB300 NVL72의 경우 9~10년으로 추정하며, 이는 해당 시스템의 재판매 가격을 기반으로 합니다. [Silicon Data](https://x.com/Silicon_Data/status/2100302896646512643)는 6년 된 A100 GPU가 여전히 원래 비용의 4분의 1의 가치가 있다고 보여주는데, 5년 감가상각 일정에서는 1년 전에 이미 0이었습니다. [Ornn Data](https://data.ornn.com/the-economics-of-open-weight-inference.pdf)는 시장이 A100 GPU를 5년 계약으로 임대할 때 1개월 계약의 80%만큼 지불하고 있다고 발견했습니다.

NVIDIA GPU를 프로그래밍하는 소프트웨어 플랫폼인 [CUDA](https://developer.nvidia.com/cuda)는 세대에 걸쳐 실행되므로, 새로운 아키텍처가 도착해도 운영자가 이미 소유한 것은 고립되지 않습니다. 지속적인 소프트웨어 및 커널 최적화는 기존 하드웨어가 할 수 있는 일을 계속 개선합니다.

동일한 플랫폼이 머신 러닝, 딥 러닝, 생성형 AI, 추론, 에이전트 AI 및 물리적 AI를 실행합니다. 각각의 새로운 작업 유형은 이미 설치된 하드웨어에 도착했습니다.

그것이 범용성입니다. 시스템이 더 많은 종류의 작업을 받아들일수록 더 오래 작업을 찾습니다.

## **범용적: 모든 유형의 AI, 모든 단계, 모든 장소**

한 가지 종류의 작업을 위해 구축된 팩토리는 그 작업이 유지된다는 베팅입니다. 모든 것을 실행하는 팩토리는 작업이 바뀌어도 유용하고 수익을 창출합니다.

NVIDIA AI 팩토리는 모든 유형의 AI 모델을 — 오픈 및 독점 — 언어, 비전, 생물학, 물리학 및 로봇 공학 전반에 걸쳐 실행합니다. 데이터 처리부터 사전 학습, 사후 학습 및 추론에 이르기까지 모든 단계를 실행합니다. 그리고 하이퍼스케일 및 AI 클라우드에서 주권 프로그램, 기업 데이터 센터 및 엣지에 이르기까지 모든 장소에서 실행합니다.



NVIDIA 플랫폼은 범용적이며 모든 유형의 AI 워크로드를 모든 단계와 장소에서 실행합니다.

전부가 AI 모델을 구축하고 실행하는 것은 아닙니다. 동일한 인프라가 데이터 처리, 과학 컴퓨팅, 시뮬레이션, 그래픽 등을 실행합니다.

AI와 비AI를 막론한 모든 워크로드는 동일한 병렬 수학으로 귀결되며, NVIDIA GPU는 수천 개의 코어에서 동시에 정확히 그것을 실행하도록 구축되었습니다. CUDA는 하나의 칩이 빛을 시뮬레이션하고, 단백질을 접고, 다음 토큰을 예측할 수 있는 이유입니다. 1,000개 이상의 기성 CUDA-X 라이브러리와 모델이 그 위에 있으며, 심층 신경망, 계산 리소그래피 및 양자 회로 시뮬레이션에서 벡터 검색 및 기후 모델링에 이르기까지 모든 것을 포괄하며, 1,000만 명 이상의 개발자가 그 위에 구축하고 있습니다.

이것이 NVIDIA GPU를 하나의 워크로드를 위해 구축된 맞춤형 ASIC이 아닌 범용 가속 컴퓨팅으로 만드는 것입니다. 범용이라는 것이 일반적이라는 것을 의미하지는 않습니다: Tensor Core와 Transformer Engine은 프로그래밍 가능한 아키텍처 내부에 AI에 최적화된 하드웨어를 배치하여 하나의 칩에서 전문화와 유연성을 모두 제공합니다. 이 모든 것을 실행하는 하나의 아키텍처가 높은 활용도를 유지하고 그에 따른 수익을 유지합니다.

이러한 다재다능함은 고객의 실제 운영에서 나타납니다:

* [**Lilly**](https://blogs.nvidia.com/blog/lilly-ai-factory-live/)**:** 1,016-GPU 온프레미스 클러스터에서 단백질, 소분자 및 유전체 모델을 구축 및 실행하고, 자체 팀을 위한 챗봇 및 에이전트 워크플로우도 운영합니다.
* [**Pinterest**](https://www.nvidia.com/en-us/case-studies/pinterest/)**:** NVIDIA Blackwell, Hopper 및 이전 아키텍처에 걸친 14,000개의 GPU를 사용하여 하이퍼스케일 클라우드에서 비전 언어 모델을 사후 학습하고 배포합니다.
* [**Revolut**](https://www.nvidia.com/en-us/case-studies/revolut/)**:** NVIDIA cuDF로 수십억 건의 거래 기록 데이터를 처리한 후 AI 클라우드에서 파운데이션 모델을 학습하고 배포합니다.
* [**Runway**](https://runway.com/research/introducing-gwm-worlds-2)**:** NVIDIA Hopper에서 월드 모델을 학습하고 클라우드 인프라를 사용하여 NVIDIA Blackwell 플랫폼에서 서비스합니다.
* [**Texas A&M University**](https://www.nvidia.com/en-us/case-studies/texas-a-m-university/)**:** 슈퍼컴퓨터에서 분자 시뮬레이션 및 AI 신약 발견을 실행하며, 26개 프로젝트와 7개 기관에서 95-98%의 활용도를 보입니다.

AI를 넘어서도 마찬가지입니다. [Cosm](https://www.nvidia.com/en-us/case-studies/cosm-immersive-live-sports-platform/)은 고해상도 비디오 재생, 스트리밍, 실시간 그래픽 및 동기화된 디스플레이 전반에 걸쳐 유연한 컴퓨팅 리소스를 제공하는 분산 NVIDIA 기반 데이터 센터를 운영합니다. [Dassault Systèmes](https://www.nvidia.com/en-us/case-studies/dassault-systemes/)는 Wichita State의 항공기 인증과 Lucid Motors의 차량 설계 뒤에 있는 가상 트윈 시뮬레이션을 지원합니다. 그리고 [Unilever](https://www.nvidia.com/en-us/case-studies/unilever/)는 사진 촬영 대신 디지털 트윈에서 제품 이미지를 만들어 생산 비용을 절반으로 줄입니다.

여기의 예시 목록은 완전할 수 없습니다 — 그것이 핵심입니다.

NVIDIA AI 팩토리는 생산적이고 내구성 있으며 범용적으로 설계되었습니다: 더 수익성 있는 토큰, 더 긴 유효 수명, 더 깊은 수요. 그것이 수익을 극대화하는 방법입니다.

_NVIDIA AI 팩토리에 대해 더 알아보려면 NVIDIA 창립자 겸 CEO Jensen Huang의_[_GTC Berlin 기조연설_](https://www.nvidia.com/en-eu/gtc/keynote/)_을 10월 21일 수요일 오전 11시 CEST에 시청하세요._
브리프용 요약 초안
NVIDIA가 AI 팩토리의 투자 수익을 생산성·내구성·범용성으로 설명하는 마케팅 글을 게시했다. Vera Rubin NVL72가 GB300 NVL72 대비 메가와트당 처리량 30배, 토큰 비용 45배 개선을 주장하나 메모리 제품·용량·고객은 명시하지 않았다. 메모리 산업 직접 영향은 확인되지 않으며, AI 팩토리 확대에 따른 간접 수요는 미확인이다.

원문 텍스트

원문 열기 ↗
AI factories are built by the megawatt, even by the gigawatt. Each megawatt factory costs roughly $60 million, and AI factory operators will only commit capital on that scale with a clear view of the return on investment. Three key things shape AI factory returns:

1. **Earning capacity:** What the factory could earn in a year if it sold every token it can produce.
2. **Useful life**: How long its AI hardware keeps earning.
3. **Demand**: How much demand there is for those tokens.

Strength cannot fully offset weakness in another. High earning capacity counts for little if the factory sells only part of what it can produce. High demand matters little if it stops producing at full capacity in just a year. Nor are the three independent. A factory that can run more kinds of workloads finds more demand, keeping it earning year after year.

NVIDIA AI factories are engineered to maximize all three. They’re:

* **Productive**: Delivering the highest throughput per megawatt and the lowest cost per token, which **maximizes their earning capacity**.
* **Durable**: NVIDIA GPUs and systems keep earning years after they ship,**extending useful life**.
* **Fungible**: They run every type of AI — in every phase and every place — as well as many workloads that don’t involve AI at all, which**deepens and broadens the demand** they can serve.



NVIDIA platform is productive, durable and fungible, which maximizes AI factory returns.

Engineering codesign across the full stack maximizes AI factory throughput, and continuous software optimization keeps installed hardware productive years after it ships. [NVIDIA CUDA-X](https://www.nvidia.com/en-us/technologies/cuda-x/) libraries let a factory run any accelerated workload. A standardized architecture then puts all of it within reach of any operator, deployable from a validated reference design.

## **Productive: Highest Tokens Per Megawatt and Lowest Token Cost**

Power is the binding constraint on an AI factory. This makes tokens per second per megawatt the number that governs earning capacity. More tokens inside a fixed power envelope means more revenue. Lower cost per token means more margin on it.

[SemiAnalysis AgentX](https://newsletter.semianalysis.com/p/vera-rubin-nvl72-agentic-inference) data shows NVIDIA Vera Rubin NVL72 systems deliver over 30x higher throughput per megawatt than NVIDIA GB300 NVL72, and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model. Gains of that size come from[extreme codesign](https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/) across the full stack: from models and workloads down through software to compute, networking and memory, all optimized together.



SemiAnalysis AgentX analysis about NVIDIA GB300 NVL72 performance for GLM 5.3.

Two questions follow. If every generation makes tokens dramatically cheaper, does demand for compute shrink?

No, it expands.

Cheaper tokens make more use cases economical, and those use cases consume more tokens than the efficiency saved.

The second question is about durability. If each generation is so much better than the last, what happens to the older generations?

## **Durable: The Installed Base Keeps Earning**

Not every workload needs the newest system. The right fit depends on a workload’s complexity and shape, which is why the last generation keeps earning after the next one arrives.

The NVIDIA A100 GPU shipped in 2020 and is still in commercial service six years later, demonstrating its continued economic value; CoreWeave recently extended bookings for units first introduced in 2020 through 2029. Over the years, every major operator has extended the depreciation schedule on its servers, which is a guess about when hardware stops earning, and one that keeps moving out. A September 2026 Sprout analysis, “[The Productive Life of a Data Center GPU](https://www.sproutup.com/resources/blog/how-long-does-a-data-center-gpu-actually-last),” tracks how that schedule has shifted across every major operator.



Every major operator has extended server life. Source: Sprout, “The Productive Life of a Data Center GPU,” September 2026. Data based on company disclosures and press reporting compiled by Sprout. Accounting life is a conservative proxy for physical life — Microsoft’s NVIDIA V100 fleet ran 8.4 years against a six-year book life.

[Barkr](https://barkr.ai/market-report#the-resale-standard-forward-looking-gpu-valuations) puts useful life at five to six years for an eight-GPU H100 system and nine to 10 years for GB300 NVL72, based on what those systems resell for. [Silicon Data](https://x.com/Silicon_Data/status/2100302896646512643)shows a six-year-old A100 GPU is still worth a quarter of what it cost, where a five-year depreciation schedule had it at zero more than a year ago.[Ornn Data](https://data.ornn.com/the-economics-of-open-weight-inference.pdf) finds the market paying 80% as much to rent an A100 GPU on a five-year contract as on a one-month contract.

[CUDA](https://developer.nvidia.com/cuda), the software platform with which NVIDIA GPUs are programmed, runs across generations, so nothing an operator already owns is stranded when a new architecture arrives. Continuous software and kernel optimization keeps improving what existing hardware can do.

The same platform runs machine learning, deep learning, generative AI, reasoning, agentic AI and physical AI. Each new kind of work arrived on hardware that was already installed.

That’s fungibility. The more kinds of work a system can take, the longer it keeps finding work.

## **Fungible: Every Type of AI, Every Phase, Every Place**

A factory built for one kind of work is a bet that the work stays. A factory that runs everything stays useful and revenue-generating even when the work changes.

NVIDIA AI factories run every type of AI model — open and proprietary — across language, vision, biology, physics and robotics. They run every phase, from data processing through pretraining, post-training and inference. And they run in every place, from hyperscale and AI clouds to sovereign programs, enterprise data centers and the edge.



The NVIDIA platform is fungible and runs every type of AI workload, in every phase and place.

Not all of it is building and running AI models. The same infrastructure runs data processing, scientific computing, simulation, graphics and more.

All of those workloads, AI and non-AI alike, reduce to the same parallel math, and NVIDIA GPUs are built to run exactly that across thousands of cores at once. CUDA is why one chip can simulate light, fold a protein and predict the next token. More than 1,000 ready-made CUDA-X libraries and models sit on top, covering everything from deep neural networks, computational lithography and quantum circuit simulation to vector search and climate modeling, with more than 10 million developers building on them.

That’s what makes NVIDIA GPUs general-purpose accelerated computing rather than a custom ASIC built for one workload. Being general purpose does not mean being generic: Tensor Cores and the Transformer Engine put AI-optimized hardware inside a programmable architecture, delivering specialization and flexibility in one chip. One architecture running all of it is what keeps utilization high, and the return with it.

This versatility shows up in production across customers:

* [**Lilly**](https://blogs.nvidia.com/blog/lilly-ai-factory-live/)**:** Building and running protein, small-molecule and genomics models on a 1,016-GPU, on-premises cluster, plus chatbots and agentic workflows for its own teams.
* [**Pinterest**](https://www.nvidia.com/en-us/case-studies/pinterest/)**:** Post-training and deploying a vision language model on a hyperscale cloud, across 14,000 GPUs spanning NVIDIA Blackwell, Hopper and earlier architectures.
* [**Revolut**](https://www.nvidia.com/en-us/case-studies/revolut/)**:**Processing data for billions of transaction records with NVIDIA cuDF, then training and deploying a foundation model on an AI cloud.
* [**Runway**](https://runway.com/research/introducing-gwm-worlds-2)**:** Training a world model on NVIDIA Hopper and serving on the NVIDIA Blackwell platform using cloud infrastructure.
* [**Texas A&M University**](https://www.nvidia.com/en-us/case-studies/texas-a-m-university/)**:** Running molecular simulation and AI drug discovery on its supercomputer, at 95-98% utilization across 26 projects and seven institutions.

The same is true beyond AI. [Cosm](https://www.nvidia.com/en-us/case-studies/cosm-immersive-live-sports-platform/) operates a distributed NVIDIA-powered data center providing flexible compute resources across high-resolution video playback, streaming, real-time graphics and synchronized display. [Dassault Systèmes](https://www.nvidia.com/en-us/case-studies/dassault-systemes/)powers the virtual twin simulation behind aircraft certification at Wichita State and vehicle design at Lucid Motors. And [Unilever](https://www.nvidia.com/en-us/case-studies/unilever/) builds product imagery from digital twins rather than photo shoots, cutting production costs in half.

No list of examples here would be complete — that’s the point.

NVIDIA AI factories are engineered to be productive, durable and fungible: more profitable tokens, longer useful life and deeper demand. That’s what maximizes their return.

_Learn more about NVIDIA AI factories by tuning in to NVIDIA founder and CEO Jensen Huang’s_[_GTC Berlin keynote_](https://www.nvidia.com/en-eu/gtc/keynote/)_on Wednesday, Oct. 21, at 11 a.m. CEST._