MEMORY INDUSTRY INTELLIGENCE
Crusoe Cloud에서 페스티벌 규모의 AI 영상 생성
한국어 번역·요약·분석
원문 제목: AI video generation at festival scale on Crusoe Cloud
원문 문장이 일치하지 않은 주장 5건은 근거 등록에서 제외했습니다.
핵심 요약
Monks는 Boomtown 페스티벌 참석자 약 77,000명에게 48시간 내 애프터무비를 제공하려 했고, Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200 시스템과 Quantum-2 InfiniBand로 두 개의 오픈소스 파이프라인을 구축했다. Pipeline A는 실제 프레임 사이에 AI 생성 '오버그로스' 전환을 만들고, Pipeline B는 Qwen-Image-Edit로 이름을 깃발에 렌더링한 뒤 LTX-2.3으로 5초·121프레임 영상을 생성했다. 최종적으로 10,195개의 개인화 깃발 영상을 포함해 약 180,000개 클립을 렌더링하고 77,000개 영상을 전달했으며, 표준 이름 기준 성공률 100%, 총 약 226 GPU-시간을 기록했다. 전통 방식 대비 인력 3~5 FTE에서 1~2 FTE, 비용은 영상당 약 $2.94에서 약 $1.96로, 기간은 3주 이상에서 1~2주로 개선되었다고 원문은 밝힌다. 1% 미만의 특수 문자 이름에서 이미지 플레이트 호환성 문제로 정지 깃발이 생성된 사례가 있었고, 하드웨어 관련 실패는 없었다고 서술한다.
메모리 산업 영향 분석
이 문서는 Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200 GPU 16개로 페스티벌 규모의 AI 영상 생성 파이프라인을 운영한 사례다. 메모리 산업 관점에서 직접 확인되는 연결은 H200 GPU에 탑재된 141GB HBM3e와 클러스터 전체 2,256GB HBM3e 용량, 그리고 Intel Xeon Sapphire Rapids CPU가 사용되었다는 점이다. 이는 HBM 수요와 서버 DDR 수요의 존재를 보여주는 응용 사례이지만, 특정 메모리 공급사·제품 구매량·거래 관계는 원문에 없어 미확인이다. 분석가 가설로는 생성형 AI 영상 워크로드가 확산될 경우 HBM 탑재 가속기 수요와 서버 호스트 DDR 수요가 늘어날 수 있으나, 이 문서는 단일 프로젝트의 226 GPU-시간 규모로 메모리 산업 전체에 미치는 직접 효과는 제한적이다. 계층 이동은 확인되지 않으며, HBM·서버 DDR·NAND 등 제품별 사용량 효과는 원문 수치로 분리할 수 없다. 반대 근거로는 오픈소스 모델과 기존 H200 인프라를 재사용하므로 추가 메모리 구매가 수반되지 않을 수 있다는 점, 그리고 원문이 하드웨어 실패가 없었다고만 밝혀 메모리 수요 증가를 직접 입증하지 않는다는 점을 들 수 있다. 확인할 지표는 H200 출하·HBM 탑재량, Crusoe Cloud의 GPU 인스턴스 증설, 해당 리전의 서버 DDR 구성, 그리고 향후 유사 워크로드의 확장 여부다.
한국어 번역 읽기
수집된 원문 v1의 전체 본문 기준 · 14370자
Cloud
Engineering
2026년 9월 28일
한 주말에 10,195개의 AI 클립, 청정 에너지로 구동
글로벌 마케팅 및 기술 서비스 회사인 Monks는 Boomtown 참석자들에게 자신의 이름이 들어간 영상을 제공하려 했습니다. 그들은 Crusoe Cloud의 아이슬란드 리전에서 NVIDIA H200 GPU에 두 개의 오픈소스 파이프라인을 구축하고 한 번의 프로덕션 윈도우에서 10,195개의 5초 클립을 전달했습니다. 방법은 다음과 같습니다.
Rohit Kalmankar
Senior Staff Solutions Engineer
2026년 9월 28일
목차
This is some text inside of a div block.
공유:
Boomtown은 전형적인 음악 페스티벌이 아닙니다. 정교한 세계관 구축, 지속가능성 약속, 몰입형 스토리텔링으로 유명하며, 매년 수만 명의 참석자를 며칠 동안만 존재하는 임시 도시로 끌어들입니다.
2026년, 글로벌 크리에이티브 기술 회사 Monks는 그 야망에 걸맞은 새로운 것을 시도했습니다. 페스티벌 후 48시간 이내에 모든 방문객(총 약 77,000명)을 위한 애프터무비를 공유하여 각자의 페스티벌 여정에 대한 기억을 제공하는 것입니다. 영상의 특정 부분에서 그들은 NVIDIA HGX™ H200 시스템과
NVIDIA Quantum-2 InfiniBand
네트워킹을 Crusoe Cloud 아이슬란드에서 연결한 새로운 개인화 콘텐츠 제작 방식을 탐구했습니다.
결과: 10,195개의 고유하게 개인화된 깃발 클립과 수천 개의 AI 생성 시네마틱 전환이 몇 주가 아닌 며칠 만에, 전통적으로 필요한 수동 작업의 일부만으로 제작되었습니다.
그들이 어떻게 했는지 소개합니다.
솔루션 아키텍처
과제: 페스티벌 규모의 영상 제작
라이브 이벤트를 위한 개인화 영상 콘텐츠 제작은 항상 느리고, 비싸고, 확장이 불가능했습니다. 10,000개의 개인화된 5초 클립에 대한 전통적 접근 방식은 다음과 같습니다.
인력:
3~5명의 전임 크리에이티브
상업적 가치:
>$30,000
툴링:
Adobe After Effects 및 라이선스 크리에이티브 소프트웨어
소요 시간:
3주 이상
주말 동안 지속되는 페스티벌에는 그 일정이 맞지 않습니다. 그리고 모든 참석자에게 자신만의 영상을 제공하는 개인화는 합리적인 비용으로는 사실상 불가능했습니다.
Monks는 더 나은 방법이 있다고 믿었습니다. 하지만 이를 증명하려면 로컬 하드웨어에서는 실행할 수 없는 무거운 오픈소스 AI 모델을 실험할 심각한 GPU 컴퓨트가 필요했습니다.
왜 Crusoe이고, 왜 아이슬란드인가
단 하나의 영상이 생성되기 전에 Monks는 다른 클라우드 제공자가 줄 수 없는 것을 제공할 인프라 파트너가 필요했습니다. 즉, 다른 곳에서 겪었던 프로비저닝 장벽 없이 대규모 모델 R&D와 프로덕션 워크로드를 실행할 충분한 NVIDIA HGX H200 노드에 대한 접근입니다.
참여 초기부터 Crusoe의 솔루션 엔지니어링 팀은 파이프라인이 설계되기도 전에 이 워크로드에 적합한 인프라로 Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200을 추천했습니다. 그 추천은 세 가지 요인에 근거했습니다. NVIDIA H200 GPU의 141GB HBM3e 메모리가 대형 오픈소스 생성 모델을 전체 정밀도로 실행하는 데 이상적이라는 점, 아이슬란드가 영국 페스티벌에 지리적으로 가까워 데이터 전송이 더 빠르다는 점, Crusoe가 다른 플랫폼에서 겪었던 대기 시간 없이 전용 격리 GPU 노드를 대규모로 프로비저닝할 수 있다는 점입니다.
"Crusoe는 우리에게 더 많은 GPU에 대한 접근을 제공했고, 이 R&D에 필요한 GPU를 프로비저닝하려 할 때 겪었던 것과 같은 장벽이 없었습니다." - Monks 팀
그 초기 인프라 결정은 옳았습니다. NVIDIA H200 GPU 노드는 며칠 동안 지속적인 혼합 이미지·영상 생성을 단 한 번의 하드웨어 관련 실패 없이 처리했습니다.
아이슬란드 선택은 지연 시간 외에도 의도적이었습니다. Boomtown은 지속가능성을 최우선으로 하는 페스티벌입니다. 관객들은 환경 영향에 깊이 관심이 있으며, 화석 연료로 구동되는 AI 생성 콘텐츠는 페스티벌이 추구하는 모든 것과 상충했을 것입니다. Crusoe의 아이슬란드 인프라는 지열 및 수력 발전의 청정 재생 에너지로 운영됩니다.
NVIDIA의 GPU 가속 컴퓨트 스택, 특히 141GB HBM3e 메모리와 NVIDIA Quantum-2 InfiniBand 네트워킹을 갖춘 H200은 두 파이프라인의 핵심인 오픈소스 모델을 실행하는 데 필요한 원시 성능을 제공했습니다.
인프라
Crusoe는 이 프로덕션 실행을 위해 특별히 예약된 공유 테넌시가 없는 프라이빗 VPC인
eu-iceland1-a
에 두 개의 전용 NVIDIA HGX H200 인스턴스를 프로비저닝했습니다.
클러스터:
2x Crusoe 인스턴스, 각 8x NVIDIA H200 GPU(총 16 GPU)
GPU 메모리:
GPU당 141GB HBM3e, 인스턴스당 1,128GB, 클러스터 전체 2,256GB
CPU / RAM:
인스턴스당 176 vCPU(Intel Xeon Sapphire Rapids), 2,000GB RAM
리전:
eu-iceland1-a
, 지열 및 수력 발전, 영국에 가장 가까운 Crusoe 리전
네트워킹:
NVIDIA Quantum-2 InfiniBand, 3,200 Gbps, 클러스터 내 지연 시간 <1ms, VPC 175 Gbps
스토리지:
인스턴스당 8x 1.92TB NVMe
데이터 전송:
~398 MB/s 병렬 인바운드(아이슬란드에서 프랑크푸르트)
모델:
오픈소스, 독점 소프트웨어 라이선스 없음, 벤더 종속 없음
오케스트레이션:
두 노드에 걸친 결정적 샤딩을 갖춘 재개 안전 작업 큐
인스턴스 간 NVIDIA Quantum-2 InfiniBand 네트워킹은 두 노드가 하나의 조정된 클러스터로 작동할 수 있게 했으며, 이는 10,195개 이름 워크로드를 병렬 GPU 큐에 중복이나 노드 장애 시 데이터 손실 없이 분할하는 데 중요했습니다.
Pipeline A: 오버그로스 클립 전환 시스템
첫 번째 과제는 창작적 연속성이었습니다. 페스티벌 영화는 여러 무대, 날짜, 순간에 걸쳐 촬영된 수백 개의 클립으로 구성됩니다. 하드 컷이나 단순 디졸브로 이들을 자르면 마법이 사라집니다. Monks 팀은 더 살아 있는 느낌의 전환을 원했습니다.
해결책: 클립 사이에 자연을 자라게 하기.
두 페스티벌 클립 사이에서 AI는 아이비, 버섯, 꽃, 나비로 이루어진 짧은 "오버그로스" 브리지를 생성하여 한 클립을 다음 클립으로 모핑한 뒤 실제 영상으로 하드 컷합니다. 전체 전환은 8~9초가 걸리며 실제 프레임에서 전적으로 AI 생성되며, Crusoe Cloud의 NVIDIA HGX H200 노드에서 배치 작업으로 실행됩니다.
작동 방식
전환을 말로 설명하는 대신, 모델에는 네 개의 실제 프레임이 앵커로 주어집니다. 클립 A 끝에서 두 개, 클립 B 시작에서 두 개입니다. 모델은 A 쌍을 "이전"으로, B 쌍을 "이후"로 취급한 다음 그 사이의 모든 프레임을 하나의 연속 시퀀스로 생성합니다.
핵심 통찰: 한쪽당 두 프레임, 하나가 아닙니다.
한쪽당 하나의 고정 프레임은 전환 중간에 군중과 카메라가 정지하여 영상이 일시정지한 것처럼 보입니다. 0.1초 차이의 두 프레임은 속도 정보를 담고 있어 군중이 오버그로스를 통해 자연스럽게 계속 움직입니다.
생성된 전환이 실제 카메라 움직임과 일치하도록 각 클립에 걸쳐 수백 개의 추적점이 매핑됩니다. 거의 모든 궤적이 같은 방향으로 쓸면 그 공유 움직임이 카메라 무브입니다. 이는 모델에 어떤 프레임을 가이드로 사용할지 알려줍니다.
착지 지점에서 AI는 드리프트하는 경향이 있습니다. 모델의 마지막 프레임을 신뢰하는 대신, 파이프라인은 실제 클립 B와 픽셀 매칭되는 생성 프레임을 자동으로 찾아 거기서 자릅니다. 디졸브도 크로스페이드도 없습니다. 움직임이 그대로 이어집니다.
세 가지 순환 모티프가 긴 전환 시퀀스를 신선하게 유지합니다. 아이비와 이끼, 주황 버섯, 꽃과 나비입니다. 모두 NVIDIA H200 GPU 노드에서 배치 작업으로 실행되며 한 번에 수천 개의 클립을 처리합니다.
Pipeline B: 대규모 개인화 깃발 영상
R&D 단계에서 Crusoe의 NVIDIA H200 GPU 노드에서 오픈소스 모델을 테스트하고 각 작업에 가장 잘 수행하는 모델을 식별한 것이 확장 프로덕션 파이프라인 설계에 반영되었습니다. 가설: 실제 깃발 천에 개별 참석자의 이름을 렌더링한 수천 개의 고유하게 개인화된 깃발 영상을 16x NVIDIA H200 GPU에 걸쳐 병렬로 실행되는 2단계 파이프라인으로 대량 생산하는 것입니다.
모델
Qwen-Image-Edit
는 각 참석자의 이름을 주름, 조명, 원근을 존중하며 깃발 천에 렌더링합니다.
LTX-2.3
은 정적 깃발 플레이트를 5초, 121프레임의 펄럭이는 영상으로 애니메이션화합니다.
둘 다 Crusoe 인스턴스에 직접 다운로드된 오픈소스 모델입니다. 라이선스 비용 없음. API 호출 없음. NVIDIA H200 GPU의 141GB HBM3e 메모리는 두 모델을 전체 정밀도로 실행할 수 있게 했고, 10,195개의 고유 이름 렌더링 전체에서 일관된 품질을 보장했습니다.
파이프라인
1단계, 이미지 생성(Qwen-Image-Edit)
Qwen은 사진 같은 깃발 플레이트를 받아 수신자의 이름을 천의 주름, 조명, 원근을 존중하며 천에 직접 렌더링합니다. 결과는 이름이 거기에 인쇄된 것처럼 보입니다.
2단계, 영상 생성(LTX-2.3)
LTX-2.3은 정적 깃발 플레이트를 바람에 펄럭이는 5초, 121프레임 영상으로 애니메이션화합니다. 최종 경량 크롭 단계가 결과물을 패키징합니다.
모든 것은 빨강, 초록, 보라의 세 가지 색상 변형에 걸쳐 실행되며, 단계별 고정 시드로 같은 이름이 항상 같은 픽셀로 재실행되어 대규모에서 샘플링 아티팩트가 없습니다.
Crusoe에서 실행
이름 목록은 결정적 샤딩으로 두 개의 8x NVIDIA H200 GPU 인스턴스에 분할되어 어떤 노드가 충돌해도 작업을 중복하거나 누락하지 않고 재개할 수 있었습니다. 모든 GPU는 재개 안전 큐에서 작업을 가져와 노드별 원장에 결과를 기록하여 단계별 실시간 가시성(완료, 재시도, 대기, 성공률, 누적 컴퓨트 시간)을 제공했습니다.
팀의 접근 방식: 항상 배치 테스트 먼저. 10개 이미지, 그다음 50개, 그다음 100개, 그다음 품질 확인, 그다음 10,000개. 그런 다음 영상에 대해 반복. 프로덕션 윈도우 전에 테스트용 전용 Crusoe 노드를 사용할 수 있었기에 가능했던 이 단계적 접근은 라이브 실행을 원활하고 예측 가능하게 만들었습니다.
텍스트 인식(OCR) 스크립트가 이미지 생성 후 인스턴스에서 실행되어 영상 단계 전에 모든 이름이 올바르게 렌더링되었는지 확인했습니다. 후반 크롭은 Crusoe 인스턴스에서 직접 스크립트로 실행되어 별도의 후처리 단계를 완전히 제거했습니다.
최종 전달: 완료된 영상은 Crusoe 인스턴스에서 직접 더 넓은 클라이언트 전달 파이프라인으로 가져와 파일명으로 각 참석자와 매칭되었습니다.
숫자로 보기
지표
결과
최종 프로덕션의 고유 이름
10,195개(빨강 3,776 / 초록 3,210 / 보라 3,209)
생성된 총 영상
테스트 배치와 시드 스윕 포함 10,000개 이상
성공률
표준 이름 100%, 모든 단계에서 중단된 작업 0개
중간 이미지 생성 시간
개인화 이미지당 약 33초(NVIDIA H200 GPU)
중간 영상 생성 시간
121프레임 영상당 약 54초(NVIDIA H200 GPU)
중간 크롭 시간
크롭당 약 1.8초
총 컴퓨트
클러스터 전체 약 226 GPU-시간
클러스터
2× Crusoe 인스턴스 · 각 8× NVIDIA H200 GPU · 16 GPU 병렬
이전 vs 이후
전통(비AI)
Crusoe의 NVIDIA H200 GPU에서 AI
필요 인력
3~5 FTE
1~2 FTE
상업적 가치
>$30,000
~$20,000
소요 시간
3주 이상
1~2주
툴링
Adobe After Effects(라이선스)
오픈소스(Qwen-Image-Edit + LTX-2.3)
개인화
규모에서 불가능
10,195개 고유 영상
영상당 비용
~$2.94
~$1.96
언뜻 보면 비용 차이는 크지 않아 보입니다. 하지만 규모가 커지면 경제성이 빠르게 복리로 작용합니다. 영상당 약 33% 절감(전통 ~$2.94 vs Crusoe ~$1.96)으로, 연간 여러 페스티벌에 같은 파이프라인을 실행하면 연간 $100,000 이상의 절감을 의미하며, 전통적 접근으로는 달성할 수 없었던 규모의 개인화를 가능하게 합니다.
중요한 점은 프로덕션 실행 후에도 클러스터에 여유가 남아 있었다는 것입니다. 즉, 같은 인프라가 같은 윈도우 내에서 훨씬 더 많은 영상을 처리할 수 있었고, 영상당 비용을 더 낮출 수 있었습니다. 파이프라인 설정 비용은 대부분 고정되어 있습니다. 더 많은 볼륨을 밀어 넣을수록 단위 경제성이 좋아집니다.
실패한 것과 배운 것
이름의 작은 하위 집합, 1% 미만, 대부분 악센트 문자와 고유 철자가 대체 이미지 생성 도구의 이미지 플레이트를 필요로 했습니다. 그 플레이트는 영상 단계에서 정지하여 움직임 대신 정적 깃발을 생성했습니다. 근본 원인: 이미지 소스와 영상 모델 간의 플레이트 호환성이며, 파이프라인이나 하드웨어의 결함이 아닙니다.
NVIDIA H200 GPU 클러스터는 며칠 동안 지속적인 혼합 이미지·영상 생성을 단 한 번의 하드웨어 관련 실패 없이 실행했습니다. 프로덕션 실행 중 디버깅된 모든 사고는 소프트웨어 측이었습니다. 인프라가 예측 가능하게 작동하자 팀은 창작 및 소프트웨어 문제에 집중할 수 있었습니다.
두 프로덕션 실행이 함께 증명한 것
두 파이프라인 모두 같은 방식으로 확장됩니다. Crusoe 노드를 추가하고 큐를 재샤딩합니다. 여기서 구축된 툴링은 Boomtown 2027 및 그 이후, 그리고 관객에게 개인적인 것을 집에 가져가게 하고 싶은 모든 라이브 이벤트를 위한 미래 생성 영상 제품의 재사용 가능한 템플릿입니다.
"노드는 우리가 필요로 할 때 정확히 무거운 생성 부하를 감당했고, 그것이 프로젝트에 실질적인 차이를 만들었습니다." - Monks 팀, Boomtown 2026
10,195개의 개인화 깃발 영상은 훨씬 더 큰 Boomtown 프로덕션의 일부였습니다. 두 파이프라인에 걸쳐 Monks는 총 약 180,000개의 클립을 렌더링하여 참석자에게 77,000개의 영상을 전달했으며, 그중 10,195개가 각 개인의 이름으로 고유하게 개인화되었습니다.
이 참여를 시작한 인프라 추천, 즉 아이슬란드의 H200은 파이프라인 코드 한 줄이 작성되기 전에 이루어졌으며, 이후 모든 것의 올바른 기반임이 입증되었습니다.
오픈소스. 청정 에너지. 하드웨어 실패 제로.
독점 소프트웨어 라이선스 없음. 벤더 종속 없음. 두 개의 오픈소스 모델(Qwen-Image-Edit 및 LTX-2.3), Crusoe의 재생 에너지 아이슬란드 인프라의 16 NVIDIA H200 GPU.
그것이 10,195명의 페스티벌 참가자에게 결코 잊지 못할 무언가를 주기 위해 필요했던 전부입니다.
결론
Boomtown 2026은 페스티벌 규모의 AI 생성 개인화 영상 콘텐츠가 더 이상 개념이 아님을 증명했습니다. 그것은 반복 가능하고 비용 효율적인 프로덕션 파이프라인입니다. 오픈소스 모델과 Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200 시스템 위에 구축된 여기서 개발된 파이프라인은 다음 페스티벌과 다음 크리에이티브 브리프를 위해 다시 실행할 준비가 되어 있습니다.
파이프라인 작업을 해준 Monks 기술 팀에 감사드립니다.
자체 생성 영상 파이프라인을 구축하시나요?
Crusoe Cloud를 살펴보거나
우리 팀과 이야기하세요
.
최신 기사
2026년 10월 9일
Crusoe의 추론 엔진은 얼마나 친환경적인가?
2026년 10월 9일
Lone Star State에서 커뮤니티 구축
2026년 10월 7일
Triton을 이용한 GPU 프로그래밍, 파트 1
멋진 것을 만들 준비가 되셨나요?
문의하기
Engineering
2026년 9월 28일
한 주말에 10,195개의 AI 클립, 청정 에너지로 구동
글로벌 마케팅 및 기술 서비스 회사인 Monks는 Boomtown 참석자들에게 자신의 이름이 들어간 영상을 제공하려 했습니다. 그들은 Crusoe Cloud의 아이슬란드 리전에서 NVIDIA H200 GPU에 두 개의 오픈소스 파이프라인을 구축하고 한 번의 프로덕션 윈도우에서 10,195개의 5초 클립을 전달했습니다. 방법은 다음과 같습니다.
Rohit Kalmankar
Senior Staff Solutions Engineer
2026년 9월 28일
목차
This is some text inside of a div block.
공유:
Boomtown은 전형적인 음악 페스티벌이 아닙니다. 정교한 세계관 구축, 지속가능성 약속, 몰입형 스토리텔링으로 유명하며, 매년 수만 명의 참석자를 며칠 동안만 존재하는 임시 도시로 끌어들입니다.
2026년, 글로벌 크리에이티브 기술 회사 Monks는 그 야망에 걸맞은 새로운 것을 시도했습니다. 페스티벌 후 48시간 이내에 모든 방문객(총 약 77,000명)을 위한 애프터무비를 공유하여 각자의 페스티벌 여정에 대한 기억을 제공하는 것입니다. 영상의 특정 부분에서 그들은 NVIDIA HGX™ H200 시스템과
NVIDIA Quantum-2 InfiniBand
네트워킹을 Crusoe Cloud 아이슬란드에서 연결한 새로운 개인화 콘텐츠 제작 방식을 탐구했습니다.
결과: 10,195개의 고유하게 개인화된 깃발 클립과 수천 개의 AI 생성 시네마틱 전환이 몇 주가 아닌 며칠 만에, 전통적으로 필요한 수동 작업의 일부만으로 제작되었습니다.
그들이 어떻게 했는지 소개합니다.
솔루션 아키텍처
과제: 페스티벌 규모의 영상 제작
라이브 이벤트를 위한 개인화 영상 콘텐츠 제작은 항상 느리고, 비싸고, 확장이 불가능했습니다. 10,000개의 개인화된 5초 클립에 대한 전통적 접근 방식은 다음과 같습니다.
인력:
3~5명의 전임 크리에이티브
상업적 가치:
>$30,000
툴링:
Adobe After Effects 및 라이선스 크리에이티브 소프트웨어
소요 시간:
3주 이상
주말 동안 지속되는 페스티벌에는 그 일정이 맞지 않습니다. 그리고 모든 참석자에게 자신만의 영상을 제공하는 개인화는 합리적인 비용으로는 사실상 불가능했습니다.
Monks는 더 나은 방법이 있다고 믿었습니다. 하지만 이를 증명하려면 로컬 하드웨어에서는 실행할 수 없는 무거운 오픈소스 AI 모델을 실험할 심각한 GPU 컴퓨트가 필요했습니다.
왜 Crusoe이고, 왜 아이슬란드인가
단 하나의 영상이 생성되기 전에 Monks는 다른 클라우드 제공자가 줄 수 없는 것을 제공할 인프라 파트너가 필요했습니다. 즉, 다른 곳에서 겪었던 프로비저닝 장벽 없이 대규모 모델 R&D와 프로덕션 워크로드를 실행할 충분한 NVIDIA HGX H200 노드에 대한 접근입니다.
참여 초기부터 Crusoe의 솔루션 엔지니어링 팀은 파이프라인이 설계되기도 전에 이 워크로드에 적합한 인프라로 Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200을 추천했습니다. 그 추천은 세 가지 요인에 근거했습니다. NVIDIA H200 GPU의 141GB HBM3e 메모리가 대형 오픈소스 생성 모델을 전체 정밀도로 실행하는 데 이상적이라는 점, 아이슬란드가 영국 페스티벌에 지리적으로 가까워 데이터 전송이 더 빠르다는 점, Crusoe가 다른 플랫폼에서 겪었던 대기 시간 없이 전용 격리 GPU 노드를 대규모로 프로비저닝할 수 있다는 점입니다.
"Crusoe는 우리에게 더 많은 GPU에 대한 접근을 제공했고, 이 R&D에 필요한 GPU를 프로비저닝하려 할 때 겪었던 것과 같은 장벽이 없었습니다." - Monks 팀
그 초기 인프라 결정은 옳았습니다. NVIDIA H200 GPU 노드는 며칠 동안 지속적인 혼합 이미지·영상 생성을 단 한 번의 하드웨어 관련 실패 없이 처리했습니다.
아이슬란드 선택은 지연 시간 외에도 의도적이었습니다. Boomtown은 지속가능성을 최우선으로 하는 페스티벌입니다. 관객들은 환경 영향에 깊이 관심이 있으며, 화석 연료로 구동되는 AI 생성 콘텐츠는 페스티벌이 추구하는 모든 것과 상충했을 것입니다. Crusoe의 아이슬란드 인프라는 지열 및 수력 발전의 청정 재생 에너지로 운영됩니다.
NVIDIA의 GPU 가속 컴퓨트 스택, 특히 141GB HBM3e 메모리와 NVIDIA Quantum-2 InfiniBand 네트워킹을 갖춘 H200은 두 파이프라인의 핵심인 오픈소스 모델을 실행하는 데 필요한 원시 성능을 제공했습니다.
인프라
Crusoe는 이 프로덕션 실행을 위해 특별히 예약된 공유 테넌시가 없는 프라이빗 VPC인
eu-iceland1-a
에 두 개의 전용 NVIDIA HGX H200 인스턴스를 프로비저닝했습니다.
클러스터:
2x Crusoe 인스턴스, 각 8x NVIDIA H200 GPU(총 16 GPU)
GPU 메모리:
GPU당 141GB HBM3e, 인스턴스당 1,128GB, 클러스터 전체 2,256GB
CPU / RAM:
인스턴스당 176 vCPU(Intel Xeon Sapphire Rapids), 2,000GB RAM
리전:
eu-iceland1-a
, 지열 및 수력 발전, 영국에 가장 가까운 Crusoe 리전
네트워킹:
NVIDIA Quantum-2 InfiniBand, 3,200 Gbps, 클러스터 내 지연 시간 <1ms, VPC 175 Gbps
스토리지:
인스턴스당 8x 1.92TB NVMe
데이터 전송:
~398 MB/s 병렬 인바운드(아이슬란드에서 프랑크푸르트)
모델:
오픈소스, 독점 소프트웨어 라이선스 없음, 벤더 종속 없음
오케스트레이션:
두 노드에 걸친 결정적 샤딩을 갖춘 재개 안전 작업 큐
인스턴스 간 NVIDIA Quantum-2 InfiniBand 네트워킹은 두 노드가 하나의 조정된 클러스터로 작동할 수 있게 했으며, 이는 10,195개 이름 워크로드를 병렬 GPU 큐에 중복이나 노드 장애 시 데이터 손실 없이 분할하는 데 중요했습니다.
Pipeline A: 오버그로스 클립 전환 시스템
첫 번째 과제는 창작적 연속성이었습니다. 페스티벌 영화는 여러 무대, 날짜, 순간에 걸쳐 촬영된 수백 개의 클립으로 구성됩니다. 하드 컷이나 단순 디졸브로 이들을 자르면 마법이 사라집니다. Monks 팀은 더 살아 있는 느낌의 전환을 원했습니다.
해결책: 클립 사이에 자연을 자라게 하기.
두 페스티벌 클립 사이에서 AI는 아이비, 버섯, 꽃, 나비로 이루어진 짧은 "오버그로스" 브리지를 생성하여 한 클립을 다음 클립으로 모핑한 뒤 실제 영상으로 하드 컷합니다. 전체 전환은 8~9초가 걸리며 실제 프레임에서 전적으로 AI 생성되며, Crusoe Cloud의 NVIDIA HGX H200 노드에서 배치 작업으로 실행됩니다.
작동 방식
전환을 말로 설명하는 대신, 모델에는 네 개의 실제 프레임이 앵커로 주어집니다. 클립 A 끝에서 두 개, 클립 B 시작에서 두 개입니다. 모델은 A 쌍을 "이전"으로, B 쌍을 "이후"로 취급한 다음 그 사이의 모든 프레임을 하나의 연속 시퀀스로 생성합니다.
핵심 통찰: 한쪽당 두 프레임, 하나가 아닙니다.
한쪽당 하나의 고정 프레임은 전환 중간에 군중과 카메라가 정지하여 영상이 일시정지한 것처럼 보입니다. 0.1초 차이의 두 프레임은 속도 정보를 담고 있어 군중이 오버그로스를 통해 자연스럽게 계속 움직입니다.
생성된 전환이 실제 카메라 움직임과 일치하도록 각 클립에 걸쳐 수백 개의 추적점이 매핑됩니다. 거의 모든 궤적이 같은 방향으로 쓸면 그 공유 움직임이 카메라 무브입니다. 이는 모델에 어떤 프레임을 가이드로 사용할지 알려줍니다.
착지 지점에서 AI는 드리프트하는 경향이 있습니다. 모델의 마지막 프레임을 신뢰하는 대신, 파이프라인은 실제 클립 B와 픽셀 매칭되는 생성 프레임을 자동으로 찾아 거기서 자릅니다. 디졸브도 크로스페이드도 없습니다. 움직임이 그대로 이어집니다.
세 가지 순환 모티프가 긴 전환 시퀀스를 신선하게 유지합니다. 아이비와 이끼, 주황 버섯, 꽃과 나비입니다. 모두 NVIDIA H200 GPU 노드에서 배치 작업으로 실행되며 한 번에 수천 개의 클립을 처리합니다.
Pipeline B: 대규모 개인화 깃발 영상
R&D 단계에서 Crusoe의 NVIDIA H200 GPU 노드에서 오픈소스 모델을 테스트하고 각 작업에 가장 잘 수행하는 모델을 식별한 것이 확장 프로덕션 파이프라인 설계에 반영되었습니다. 가설: 실제 깃발 천에 개별 참석자의 이름을 렌더링한 수천 개의 고유하게 개인화된 깃발 영상을 16x NVIDIA H200 GPU에 걸쳐 병렬로 실행되는 2단계 파이프라인으로 대량 생산하는 것입니다.
모델
Qwen-Image-Edit
는 각 참석자의 이름을 주름, 조명, 원근을 존중하며 깃발 천에 렌더링합니다.
LTX-2.3
은 정적 깃발 플레이트를 5초, 121프레임의 펄럭이는 영상으로 애니메이션화합니다.
둘 다 Crusoe 인스턴스에 직접 다운로드된 오픈소스 모델입니다. 라이선스 비용 없음. API 호출 없음. NVIDIA H200 GPU의 141GB HBM3e 메모리는 두 모델을 전체 정밀도로 실행할 수 있게 했고, 10,195개의 고유 이름 렌더링 전체에서 일관된 품질을 보장했습니다.
파이프라인
1단계, 이미지 생성(Qwen-Image-Edit)
Qwen은 사진 같은 깃발 플레이트를 받아 수신자의 이름을 천의 주름, 조명, 원근을 존중하며 천에 직접 렌더링합니다. 결과는 이름이 거기에 인쇄된 것처럼 보입니다.
2단계, 영상 생성(LTX-2.3)
LTX-2.3은 정적 깃발 플레이트를 바람에 펄럭이는 5초, 121프레임 영상으로 애니메이션화합니다. 최종 경량 크롭 단계가 결과물을 패키징합니다.
모든 것은 빨강, 초록, 보라의 세 가지 색상 변형에 걸쳐 실행되며, 단계별 고정 시드로 같은 이름이 항상 같은 픽셀로 재실행되어 대규모에서 샘플링 아티팩트가 없습니다.
Crusoe에서 실행
이름 목록은 결정적 샤딩으로 두 개의 8x NVIDIA H200 GPU 인스턴스에 분할되어 어떤 노드가 충돌해도 작업을 중복하거나 누락하지 않고 재개할 수 있었습니다. 모든 GPU는 재개 안전 큐에서 작업을 가져와 노드별 원장에 결과를 기록하여 단계별 실시간 가시성(완료, 재시도, 대기, 성공률, 누적 컴퓨트 시간)을 제공했습니다.
팀의 접근 방식: 항상 배치 테스트 먼저. 10개 이미지, 그다음 50개, 그다음 100개, 그다음 품질 확인, 그다음 10,000개. 그런 다음 영상에 대해 반복. 프로덕션 윈도우 전에 테스트용 전용 Crusoe 노드를 사용할 수 있었기에 가능했던 이 단계적 접근은 라이브 실행을 원활하고 예측 가능하게 만들었습니다.
텍스트 인식(OCR) 스크립트가 이미지 생성 후 인스턴스에서 실행되어 영상 단계 전에 모든 이름이 올바르게 렌더링되었는지 확인했습니다. 후반 크롭은 Crusoe 인스턴스에서 직접 스크립트로 실행되어 별도의 후처리 단계를 완전히 제거했습니다.
최종 전달: 완료된 영상은 Crusoe 인스턴스에서 직접 더 넓은 클라이언트 전달 파이프라인으로 가져와 파일명으로 각 참석자와 매칭되었습니다.
숫자로 보기
지표
결과
최종 프로덕션의 고유 이름
10,195개(빨강 3,776 / 초록 3,210 / 보라 3,209)
생성된 총 영상
테스트 배치와 시드 스윕 포함 10,000개 이상
성공률
표준 이름 100%, 모든 단계에서 중단된 작업 0개
중간 이미지 생성 시간
개인화 이미지당 약 33초(NVIDIA H200 GPU)
중간 영상 생성 시간
121프레임 영상당 약 54초(NVIDIA H200 GPU)
중간 크롭 시간
크롭당 약 1.8초
총 컴퓨트
클러스터 전체 약 226 GPU-시간
클러스터
2× Crusoe 인스턴스 · 각 8× NVIDIA H200 GPU · 16 GPU 병렬
이전 vs 이후
전통(비AI)
Crusoe의 NVIDIA H200 GPU에서 AI
필요 인력
3~5 FTE
1~2 FTE
상업적 가치
>$30,000
~$20,000
소요 시간
3주 이상
1~2주
툴링
Adobe After Effects(라이선스)
오픈소스(Qwen-Image-Edit + LTX-2.3)
개인화
규모에서 불가능
10,195개 고유 영상
영상당 비용
~$2.94
~$1.96
언뜻 보면 비용 차이는 크지 않아 보입니다. 하지만 규모가 커지면 경제성이 빠르게 복리로 작용합니다. 영상당 약 33% 절감(전통 ~$2.94 vs Crusoe ~$1.96)으로, 연간 여러 페스티벌에 같은 파이프라인을 실행하면 연간 $100,000 이상의 절감을 의미하며, 전통적 접근으로는 달성할 수 없었던 규모의 개인화를 가능하게 합니다.
중요한 점은 프로덕션 실행 후에도 클러스터에 여유가 남아 있었다는 것입니다. 즉, 같은 인프라가 같은 윈도우 내에서 훨씬 더 많은 영상을 처리할 수 있었고, 영상당 비용을 더 낮출 수 있었습니다. 파이프라인 설정 비용은 대부분 고정되어 있습니다. 더 많은 볼륨을 밀어 넣을수록 단위 경제성이 좋아집니다.
실패한 것과 배운 것
이름의 작은 하위 집합, 1% 미만, 대부분 악센트 문자와 고유 철자가 대체 이미지 생성 도구의 이미지 플레이트를 필요로 했습니다. 그 플레이트는 영상 단계에서 정지하여 움직임 대신 정적 깃발을 생성했습니다. 근본 원인: 이미지 소스와 영상 모델 간의 플레이트 호환성이며, 파이프라인이나 하드웨어의 결함이 아닙니다.
NVIDIA H200 GPU 클러스터는 며칠 동안 지속적인 혼합 이미지·영상 생성을 단 한 번의 하드웨어 관련 실패 없이 실행했습니다. 프로덕션 실행 중 디버깅된 모든 사고는 소프트웨어 측이었습니다. 인프라가 예측 가능하게 작동하자 팀은 창작 및 소프트웨어 문제에 집중할 수 있었습니다.
두 프로덕션 실행이 함께 증명한 것
두 파이프라인 모두 같은 방식으로 확장됩니다. Crusoe 노드를 추가하고 큐를 재샤딩합니다. 여기서 구축된 툴링은 Boomtown 2027 및 그 이후, 그리고 관객에게 개인적인 것을 집에 가져가게 하고 싶은 모든 라이브 이벤트를 위한 미래 생성 영상 제품의 재사용 가능한 템플릿입니다.
"노드는 우리가 필요로 할 때 정확히 무거운 생성 부하를 감당했고, 그것이 프로젝트에 실질적인 차이를 만들었습니다." - Monks 팀, Boomtown 2026
10,195개의 개인화 깃발 영상은 훨씬 더 큰 Boomtown 프로덕션의 일부였습니다. 두 파이프라인에 걸쳐 Monks는 총 약 180,000개의 클립을 렌더링하여 참석자에게 77,000개의 영상을 전달했으며, 그중 10,195개가 각 개인의 이름으로 고유하게 개인화되었습니다.
이 참여를 시작한 인프라 추천, 즉 아이슬란드의 H200은 파이프라인 코드 한 줄이 작성되기 전에 이루어졌으며, 이후 모든 것의 올바른 기반임이 입증되었습니다.
오픈소스. 청정 에너지. 하드웨어 실패 제로.
독점 소프트웨어 라이선스 없음. 벤더 종속 없음. 두 개의 오픈소스 모델(Qwen-Image-Edit 및 LTX-2.3), Crusoe의 재생 에너지 아이슬란드 인프라의 16 NVIDIA H200 GPU.
그것이 10,195명의 페스티벌 참가자에게 결코 잊지 못할 무언가를 주기 위해 필요했던 전부입니다.
결론
Boomtown 2026은 페스티벌 규모의 AI 생성 개인화 영상 콘텐츠가 더 이상 개념이 아님을 증명했습니다. 그것은 반복 가능하고 비용 효율적인 프로덕션 파이프라인입니다. 오픈소스 모델과 Crusoe Cloud 아이슬란드 리전의 NVIDIA HGX H200 시스템 위에 구축된 여기서 개발된 파이프라인은 다음 페스티벌과 다음 크리에이티브 브리프를 위해 다시 실행할 준비가 되어 있습니다.
파이프라인 작업을 해준 Monks 기술 팀에 감사드립니다.
자체 생성 영상 파이프라인을 구축하시나요?
Crusoe Cloud를 살펴보거나
우리 팀과 이야기하세요
.
최신 기사
2026년 10월 9일
Crusoe의 추론 엔진은 얼마나 친환경적인가?
2026년 10월 9일
Lone Star State에서 커뮤니티 구축
2026년 10월 7일
Triton을 이용한 GPU 프로그래밍, 파트 1
멋진 것을 만들 준비가 되셨나요?
문의하기
브리프용 요약 초안
Crusoe Cloud 아이슬란드 리전에서 NVIDIA HGX H200 16개로 10,195개의 개인화 영상을 포함한 약 180,000개 클립을 생성한 사례가 공개됐다. H200의 141GB HBM3e와 Intel Xeon Sapphire Rapids가 사용됐지만 특정 메모리 공급사나 구매량은 확인되지 않는다. 단일 프로젝트 규모로 메모리 산업 직접 영향은 제한적이며, 향후 생성형 영상 워크로드 확산 시 HBM·서버 DDR 수요 연결 가능성을 관찰할 필요가 있다.
원문 텍스트
원문 열기 ↗Cloud
Engineering
September 28, 2026
10,195 AI clips in one weekend, powered by clean energy
Monks, a global marketing and technology services company, set out to give Boomtown attendees a video with their own name on it. They built two open source pipelines on NVIDIA H200 GPUs in Crusoe Cloud's Iceland region and shipped 10,195 5-second clips in one production window. Here is how.
Rohit Kalmankar
Senior Staff Solutions Engineer
September 28, 2026
Table of contents
This is some text inside of a div block.
Share:
Boomtown is not a typical music festival. Known for its intricate world-building, sustainability commitments, and immersive storytelling, it draws tens of thousands of attendees every year to a temporary city that exists for just a few days.
In 2026, Monks, a global creative technology company, set out to match that ambition with something new: sharing an aftermovie for every visitor (~77,000 in total) within 48 hours after the festival, providing a memory about everyone's individual festival journey. For a specific part of the videos they explored new ways to create personalized content, powered by NVIDIA HGX™ H200 systems connected with
NVIDIA Quantum-2 InfiniBand
networking on Crusoe Cloud in Iceland.
The result: 10,195 uniquely personalized flag clips and thousands of AI-generated cinematic transitions, produced in days, not weeks, with a fraction of the manual effort traditionally required.
This is how they did it.
Solution architecture
The challenge: video production at festival scale
Producing personalized video content for a live event has always been slow, expensive, and impossible to scale. The traditional approach to 10,000 personalized five-second clips would look something like this:
Human effort:
3 to 5 full-time creatives
Commercial value:
>$30,000
Tooling:
Adobe After Effects and licensed creative software
Turnaround:
3+ weeks
For a festival that lasts a weekend, that timeline simply doesn't work. And personalization, giving every single attendee their own video, was effectively impossible at any reasonable cost.
Monks believed there was a better way. But to prove it, they needed serious GPU compute to experiment with heavyweight open-source AI models that simply couldn't run on local hardware.
Why Crusoe, and why Iceland
Before a single video was generated, Monks needed an infrastructure partner that could give them something other cloud providers couldn't: access to enough NVIDIA HGX H200 nodes to run large-scale model R&D and production workloads, without the provisioning barriers they had experienced elsewhere.
From the start of the engagement, Crusoe's solutions engineering team recommended the NVIDIA HGX H200 in Crusoe Cloud's Iceland region as the right infrastructure for this workload, before the pipeline was even designed. That recommendation was based on three factors: the NVIDIA H200 GPU's 141GB HBM3e memory making it ideal for running large open-source generative models at full precision; Iceland's geographic proximity to the UK festival for faster data transfer; and Crusoe's ability to provision dedicated, isolated GPU nodes at scale without the queue times the team had faced on other platforms.
"Crusoe gave us access to more GPUs, we didn't have the same barrier we had trying to provision the ones we needed for this R&D." - Monks team
That early infrastructure decision turned out to be the right one. The NVIDIA H200 GPU nodes handled sustained, mixed image-and-video generation for days without a single hardware-related failure.
The choice of Iceland was equally deliberate beyond latency. Boomtown is a sustainability-first festival. Its audiences care deeply about environmental impact, and AI-generated content powered by fossil fuels would have been at odds with everything the festival stands for. Crusoe's Iceland infrastructure runs on clean, renewable energy from geothermal and hydroelectric power.
NVIDIA's GPU-accelerated compute stack, specifically the H200 with its 141GB HBM3e memory and NVIDIA Quantum-2 InfiniBand networking, provided the raw power needed to run the open-source models at the heart of both pipelines.
The infrastructure
Crusoe provisioned two dedicated NVIDIA HGX H200 instances in
eu-iceland1-a
, a private VPC with no shared tenancy, reserved specifically for this production run.
Cluster:
2x Crusoe instances, 8x NVIDIA H200 GPUs each (16 GPUs total)
GPU memory:
141GB HBM3e per GPU, 1,128GB per instance, 2,256GB total across the cluster
CPU / RAM:
176 vCPUs (Intel Xeon Sapphire Rapids), 2,000 GB RAM per instance
Region:
eu-iceland1-a
, geothermal and hydroelectric power, closest Crusoe region to the UK
Networking:
NVIDIA Quantum-2 InfiniBand, 3,200 Gbps, <1ms intra-cluster latency, 175 Gbps VPC
Storage:
8x 1.92TB NVMe per instance
Data transfer:
~398 MB/s parallel inbound (Iceland to Frankfurt)
Models:
Open source, no proprietary software licenses, no vendor lock-in
Orchestration:
Resume-safe job queues with deterministic sharding across both nodes
The NVIDIA Quantum-2 InfiniBand networking between instances meant both nodes could operate as a single coordinated cluster, which mattered for splitting the 10,195-name workload across parallel GPU queues with no duplication or data loss on any node failure.
Pipeline A: the overgrowth clip-transition system
The first challenge was creative continuity. A festival film is built from hundreds of clips shot across multiple stages, days, and moments. Cutting between them with a hard cut or a simple dissolve loses the magic. The Monks team wanted something more, a transition that felt alive.
The solution: grow nature between clips.
Between any two festival clips, the AI generates a short "overgrowth" bridge of ivy, mushrooms, flowers and butterflies that morphs one clip into the next and then hard-cuts back into real footage. The whole transition takes 8 to 9 seconds and is entirely AI-generated from real frames, running as a batch job across the NVIDIA HGX H200 nodes on Crusoe Cloud.
How it works
Rather than describing the transition in words, the model is given four real frames as anchors: two from the end of clip A and two from the start of clip B. The model treats the A pair as "before" and the B pair as "after," then generates every frame in between as one continuous sequence.
The key insight: two frames per side, not one.
A single frozen frame per side causes the crowd and camera to slow to a standstill mid-transition, so it looks like the video hits pause. Two frames a fraction of a second apart carry velocity information, so the crowd keeps moving naturally through the overgrowth.
To ensure the generated transition matches the real camera movement, hundreds of tracking points are mapped across each clip. When almost all trails sweep the same direction, that shared motion is the camera move. This tells the model which frames to use as guides.
At the landing point, the AI tends to drift. Rather than trusting the model's last frame, the pipeline automatically finds the generated frame that pixel-matches real clip B, and cuts there. No dissolve and no crossfade. The motion simply continues.
Three rotating motifs keep a long sequence of transitions feeling fresh: ivy and moss, orange mushrooms, and flowers with butterflies. It all runs as a batch job on the NVIDIA H200 GPU nodes, thousands of clips at a time.
Pipeline B: personalized flag videos at scale
The R&D phase, testing open-source models on Crusoe's NVIDIA H200 GPU nodes and identifying which performed best at each task, informed the design of the scaled production pipeline. The hypothesis: mass-produce thousands of uniquely personalized flag videos, each carrying an individual attendee's name rendered onto real flag fabric, using a two-stage pipeline running in parallel across 16x NVIDIA H200 GPUs.
The models
Qwen-Image-Edit
, renders each attendee's name onto flag fabric respecting folds, lighting, and perspective
LTX-2.3
, animates the static flag plate into a 5-second, 121-frame waving video
Both are open-source models downloaded directly onto the Crusoe instances. No licensing fees. No API calls. The NVIDIA H200 GPU's 141GB HBM3e memory allowed both models to run at full precision, ensuring consistent quality across all 10,195 unique name renders.
The pipeline
Stage 1, image generation (Qwen-Image-Edit)
Qwen takes a photographic flag plate and renders the recipient's name directly onto the fabric, respecting the folds, lighting, and perspective of the cloth. The result looks like the name was printed there.
Stage 2, video generation (LTX-2.3)
LTX-2.3 animates the static flag plate into a 5-second, 121-frame video of the flag waving in the wind. A final lightweight crop stage packages the deliverable.
Everything runs across three color variants, red, green, and purple, with fixed seeds per stage so the same name always reruns to the same pixels, with no sampling artifacts at scale.
Running it on Crusoe
The name list was split across the two 8x NVIDIA H200 GPU instances by deterministic sharding, so any node could crash and resume without duplicating or dropping work. Every GPU pulled jobs from a resume-safe queue and wrote results to a per-node ledger, giving live visibility per stage: completed, retries, pending, success rate, and cumulative compute time.
The team's approach: batch testing first, always. 10 images, then 50, then 100, then a quality check, then 10,000. Then repeat for video. This phased approach, made possible by having dedicated Crusoe nodes available for testing before the production window, meant the live run was smooth and predictable.
A text recognition (OCR) script ran on the instances after image generation to verify every name was correctly rendered before the video stage. Post-production cropping ran as a script directly on the Crusoe instances, eliminating a separate post-processing step entirely.
Final delivery: completed videos were pulled directly from the Crusoe instances into the broader client delivery pipeline, matched by filename to each attendee.
By the numbers
Metric
Result
Unique names in final production
10,195 (red 3,776 / green 3,210 / purple 3,209)
Total videos generated
10,000+ including test batches and seed sweeps
Success rate
100% on standard names, zero abandoned jobs across all stages
Median image generation time
~33 seconds per personalized image (NVIDIA H200 GPU)
Median video generation time
~54 seconds per 121-frame video (NVIDIA H200 GPU)
Median crop time
~1.8s per crop
Total compute
~226 GPU-hours across the cluster
Cluster
2× Crusoe instances · 8× NVIDIA H200 GPU each · 16 GPUs in parallel
Before vs. after
Traditional (no AI)
AI on Crusoe’s NVIDIA H200 GPU
People required
3–5 FTE
1–2 FTE
Commercial value
>$30,000
~$20,000
Turnaround
3+ weeks
1–2 weeks
Tooling
Adobe After Effects (licensed)
Open source (Qwen-Image-Edit + LTX-2.3)
Personalization
Not feasible at scale
10,195 unique videos
Cost per video
~$2.94
~$1.96
At first glance the cost difference appears modest. But the economics compound quickly at scale. At roughly 33% savings per video (~$2.94 traditional vs ~$1.96 on Crusoe), running the same pipeline across multiple festivals per year translates to over $100,000 in annual savings, while unlocking personalization at a scale that was simply not achievable with the traditional approach.
Importantly, the cluster had headroom remaining after the production run. That means the same infrastructure could have processed significantly more videos within the same window, further driving down the cost per video. The pipeline setup cost is largely fixed. The more volume you push through it, the better the unit economics get.
What failed and what we learned
A small subset of names, under 1%, mostly accented characters and unique spellings, needed image plates from an alternative image generation tool. Those plates froze in the video stage, producing a static flag instead of motion. Root cause: plate compatibility between image sources and the video model, not a flaw in the pipeline or the hardware.
The NVIDIA H200 GPU cluster ran sustained, mixed image-and-video generation for days without a single hardware-related failure. Every incident debugged during the production run was on the software side. With the infrastructure behaving predictably, the team could focus on the creative and software problems.
What the two production runs proved together
Both pipelines scale the same way: add Crusoe nodes, reshard the queue. The tooling built here is a reusable template for future generative video products, for Boomtown 2027 and beyond, and for any live event that wants to give its audience something personal to take home.
"The nodes carried a heavy generation load exactly when we needed them, and that made a real difference to the project." - Monks team, Boomtown 2026
The 10,195 personalized flag videos were one part of a much larger Boomtown production. Across both pipelines, Monks rendered approximately 180,000 clips in total, delivering 77,000 videos to attendees, 10,195 of which were uniquely personalized with each individual's name.
The infrastructure recommendation that started this engagement, H200s in Iceland, made before a single line of pipeline code was written, proved to be the right foundation for everything that followed.
Open source. Clean energy. Zero hardware failures.
No proprietary software licenses. No vendor lock-in. Two open-source models (Qwen-Image-Edit and LTX-2.3), 16 NVIDIA H200 GPUs on Crusoe's renewably-powered Iceland infrastructure.
That's what it took to give 10,195 festival-goers something they'll never forget.
Conclusion
Boomtown 2026 proved that AI-generated, personalized video content at festival scale is no longer a concept. It's a repeatable, cost-effective production pipeline. Built on open-source models and NVIDIA HGX H200 systems in Crusoe Cloud's Iceland region, the pipeline developed here is ready to go again, for the next festival and the next creative brief.
Thanks to the Monks technical team for their work on the pipelines.
Building a generative video pipeline of your own?
Explore Crusoe Cloud
, or
talk to our team
.
Latest articles
October 9, 2026
How green is Crusoe's inference engine?
October 9, 2026
Building community in the Lone Star State
October 7, 2026
GPU programming with Triton, part 1
Are you ready to build something amazing?
Contact Us
Engineering
September 28, 2026
10,195 AI clips in one weekend, powered by clean energy
Monks, a global marketing and technology services company, set out to give Boomtown attendees a video with their own name on it. They built two open source pipelines on NVIDIA H200 GPUs in Crusoe Cloud's Iceland region and shipped 10,195 5-second clips in one production window. Here is how.
Rohit Kalmankar
Senior Staff Solutions Engineer
September 28, 2026
Table of contents
This is some text inside of a div block.
Share:
Boomtown is not a typical music festival. Known for its intricate world-building, sustainability commitments, and immersive storytelling, it draws tens of thousands of attendees every year to a temporary city that exists for just a few days.
In 2026, Monks, a global creative technology company, set out to match that ambition with something new: sharing an aftermovie for every visitor (~77,000 in total) within 48 hours after the festival, providing a memory about everyone's individual festival journey. For a specific part of the videos they explored new ways to create personalized content, powered by NVIDIA HGX™ H200 systems connected with
NVIDIA Quantum-2 InfiniBand
networking on Crusoe Cloud in Iceland.
The result: 10,195 uniquely personalized flag clips and thousands of AI-generated cinematic transitions, produced in days, not weeks, with a fraction of the manual effort traditionally required.
This is how they did it.
Solution architecture
The challenge: video production at festival scale
Producing personalized video content for a live event has always been slow, expensive, and impossible to scale. The traditional approach to 10,000 personalized five-second clips would look something like this:
Human effort:
3 to 5 full-time creatives
Commercial value:
>$30,000
Tooling:
Adobe After Effects and licensed creative software
Turnaround:
3+ weeks
For a festival that lasts a weekend, that timeline simply doesn't work. And personalization, giving every single attendee their own video, was effectively impossible at any reasonable cost.
Monks believed there was a better way. But to prove it, they needed serious GPU compute to experiment with heavyweight open-source AI models that simply couldn't run on local hardware.
Why Crusoe, and why Iceland
Before a single video was generated, Monks needed an infrastructure partner that could give them something other cloud providers couldn't: access to enough NVIDIA HGX H200 nodes to run large-scale model R&D and production workloads, without the provisioning barriers they had experienced elsewhere.
From the start of the engagement, Crusoe's solutions engineering team recommended the NVIDIA HGX H200 in Crusoe Cloud's Iceland region as the right infrastructure for this workload, before the pipeline was even designed. That recommendation was based on three factors: the NVIDIA H200 GPU's 141GB HBM3e memory making it ideal for running large open-source generative models at full precision; Iceland's geographic proximity to the UK festival for faster data transfer; and Crusoe's ability to provision dedicated, isolated GPU nodes at scale without the queue times the team had faced on other platforms.
"Crusoe gave us access to more GPUs, we didn't have the same barrier we had trying to provision the ones we needed for this R&D." - Monks team
That early infrastructure decision turned out to be the right one. The NVIDIA H200 GPU nodes handled sustained, mixed image-and-video generation for days without a single hardware-related failure.
The choice of Iceland was equally deliberate beyond latency. Boomtown is a sustainability-first festival. Its audiences care deeply about environmental impact, and AI-generated content powered by fossil fuels would have been at odds with everything the festival stands for. Crusoe's Iceland infrastructure runs on clean, renewable energy from geothermal and hydroelectric power.
NVIDIA's GPU-accelerated compute stack, specifically the H200 with its 141GB HBM3e memory and NVIDIA Quantum-2 InfiniBand networking, provided the raw power needed to run the open-source models at the heart of both pipelines.
The infrastructure
Crusoe provisioned two dedicated NVIDIA HGX H200 instances in
eu-iceland1-a
, a private VPC with no shared tenancy, reserved specifically for this production run.
Cluster:
2x Crusoe instances, 8x NVIDIA H200 GPUs each (16 GPUs total)
GPU memory:
141GB HBM3e per GPU, 1,128GB per instance, 2,256GB total across the cluster
CPU / RAM:
176 vCPUs (Intel Xeon Sapphire Rapids), 2,000 GB RAM per instance
Region:
eu-iceland1-a
, geothermal and hydroelectric power, closest Crusoe region to the UK
Networking:
NVIDIA Quantum-2 InfiniBand, 3,200 Gbps, <1ms intra-cluster latency, 175 Gbps VPC
Storage:
8x 1.92TB NVMe per instance
Data transfer:
~398 MB/s parallel inbound (Iceland to Frankfurt)
Models:
Open source, no proprietary software licenses, no vendor lock-in
Orchestration:
Resume-safe job queues with deterministic sharding across both nodes
The NVIDIA Quantum-2 InfiniBand networking between instances meant both nodes could operate as a single coordinated cluster, which mattered for splitting the 10,195-name workload across parallel GPU queues with no duplication or data loss on any node failure.
Pipeline A: the overgrowth clip-transition system
The first challenge was creative continuity. A festival film is built from hundreds of clips shot across multiple stages, days, and moments. Cutting between them with a hard cut or a simple dissolve loses the magic. The Monks team wanted something more, a transition that felt alive.
The solution: grow nature between clips.
Between any two festival clips, the AI generates a short "overgrowth" bridge of ivy, mushrooms, flowers and butterflies that morphs one clip into the next and then hard-cuts back into real footage. The whole transition takes 8 to 9 seconds and is entirely AI-generated from real frames, running as a batch job across the NVIDIA HGX H200 nodes on Crusoe Cloud.
How it works
Rather than describing the transition in words, the model is given four real frames as anchors: two from the end of clip A and two from the start of clip B. The model treats the A pair as "before" and the B pair as "after," then generates every frame in between as one continuous sequence.
The key insight: two frames per side, not one.
A single frozen frame per side causes the crowd and camera to slow to a standstill mid-transition, so it looks like the video hits pause. Two frames a fraction of a second apart carry velocity information, so the crowd keeps moving naturally through the overgrowth.
To ensure the generated transition matches the real camera movement, hundreds of tracking points are mapped across each clip. When almost all trails sweep the same direction, that shared motion is the camera move. This tells the model which frames to use as guides.
At the landing point, the AI tends to drift. Rather than trusting the model's last frame, the pipeline automatically finds the generated frame that pixel-matches real clip B, and cuts there. No dissolve and no crossfade. The motion simply continues.
Three rotating motifs keep a long sequence of transitions feeling fresh: ivy and moss, orange mushrooms, and flowers with butterflies. It all runs as a batch job on the NVIDIA H200 GPU nodes, thousands of clips at a time.
Pipeline B: personalized flag videos at scale
The R&D phase, testing open-source models on Crusoe's NVIDIA H200 GPU nodes and identifying which performed best at each task, informed the design of the scaled production pipeline. The hypothesis: mass-produce thousands of uniquely personalized flag videos, each carrying an individual attendee's name rendered onto real flag fabric, using a two-stage pipeline running in parallel across 16x NVIDIA H200 GPUs.
The models
Qwen-Image-Edit
, renders each attendee's name onto flag fabric respecting folds, lighting, and perspective
LTX-2.3
, animates the static flag plate into a 5-second, 121-frame waving video
Both are open-source models downloaded directly onto the Crusoe instances. No licensing fees. No API calls. The NVIDIA H200 GPU's 141GB HBM3e memory allowed both models to run at full precision, ensuring consistent quality across all 10,195 unique name renders.
The pipeline
Stage 1, image generation (Qwen-Image-Edit)
Qwen takes a photographic flag plate and renders the recipient's name directly onto the fabric, respecting the folds, lighting, and perspective of the cloth. The result looks like the name was printed there.
Stage 2, video generation (LTX-2.3)
LTX-2.3 animates the static flag plate into a 5-second, 121-frame video of the flag waving in the wind. A final lightweight crop stage packages the deliverable.
Everything runs across three color variants, red, green, and purple, with fixed seeds per stage so the same name always reruns to the same pixels, with no sampling artifacts at scale.
Running it on Crusoe
The name list was split across the two 8x NVIDIA H200 GPU instances by deterministic sharding, so any node could crash and resume without duplicating or dropping work. Every GPU pulled jobs from a resume-safe queue and wrote results to a per-node ledger, giving live visibility per stage: completed, retries, pending, success rate, and cumulative compute time.
The team's approach: batch testing first, always. 10 images, then 50, then 100, then a quality check, then 10,000. Then repeat for video. This phased approach, made possible by having dedicated Crusoe nodes available for testing before the production window, meant the live run was smooth and predictable.
A text recognition (OCR) script ran on the instances after image generation to verify every name was correctly rendered before the video stage. Post-production cropping ran as a script directly on the Crusoe instances, eliminating a separate post-processing step entirely.
Final delivery: completed videos were pulled directly from the Crusoe instances into the broader client delivery pipeline, matched by filename to each attendee.
By the numbers
Metric
Result
Unique names in final production
10,195 (red 3,776 / green 3,210 / purple 3,209)
Total videos generated
10,000+ including test batches and seed sweeps
Success rate
100% on standard names, zero abandoned jobs across all stages
Median image generation time
~33 seconds per personalized image (NVIDIA H200 GPU)
Median video generation time
~54 seconds per 121-frame video (NVIDIA H200 GPU)
Median crop time
~1.8s per crop
Total compute
~226 GPU-hours across the cluster
Cluster
2× Crusoe instances · 8× NVIDIA H200 GPU each · 16 GPUs in parallel
Before vs. after
Traditional (no AI)
AI on Crusoe’s NVIDIA H200 GPU
People required
3–5 FTE
1–2 FTE
Commercial value
>$30,000
~$20,000
Turnaround
3+ weeks
1–2 weeks
Tooling
Adobe After Effects (licensed)
Open source (Qwen-Image-Edit + LTX-2.3)
Personalization
Not feasible at scale
10,195 unique videos
Cost per video
~$2.94
~$1.96
At first glance the cost difference appears modest. But the economics compound quickly at scale. At roughly 33% savings per video (~$2.94 traditional vs ~$1.96 on Crusoe), running the same pipeline across multiple festivals per year translates to over $100,000 in annual savings, while unlocking personalization at a scale that was simply not achievable with the traditional approach.
Importantly, the cluster had headroom remaining after the production run. That means the same infrastructure could have processed significantly more videos within the same window, further driving down the cost per video. The pipeline setup cost is largely fixed. The more volume you push through it, the better the unit economics get.
What failed and what we learned
A small subset of names, under 1%, mostly accented characters and unique spellings, needed image plates from an alternative image generation tool. Those plates froze in the video stage, producing a static flag instead of motion. Root cause: plate compatibility between image sources and the video model, not a flaw in the pipeline or the hardware.
The NVIDIA H200 GPU cluster ran sustained, mixed image-and-video generation for days without a single hardware-related failure. Every incident debugged during the production run was on the software side. With the infrastructure behaving predictably, the team could focus on the creative and software problems.
What the two production runs proved together
Both pipelines scale the same way: add Crusoe nodes, reshard the queue. The tooling built here is a reusable template for future generative video products, for Boomtown 2027 and beyond, and for any live event that wants to give its audience something personal to take home.
"The nodes carried a heavy generation load exactly when we needed them, and that made a real difference to the project." - Monks team, Boomtown 2026
The 10,195 personalized flag videos were one part of a much larger Boomtown production. Across both pipelines, Monks rendered approximately 180,000 clips in total, delivering 77,000 videos to attendees, 10,195 of which were uniquely personalized with each individual's name.
The infrastructure recommendation that started this engagement, H200s in Iceland, made before a single line of pipeline code was written, proved to be the right foundation for everything that followed.
Open source. Clean energy. Zero hardware failures.
No proprietary software licenses. No vendor lock-in. Two open-source models (Qwen-Image-Edit and LTX-2.3), 16 NVIDIA H200 GPUs on Crusoe's renewably-powered Iceland infrastructure.
That's what it took to give 10,195 festival-goers something they'll never forget.
Conclusion
Boomtown 2026 proved that AI-generated, personalized video content at festival scale is no longer a concept. It's a repeatable, cost-effective production pipeline. Built on open-source models and NVIDIA HGX H200 systems in Crusoe Cloud's Iceland region, the pipeline developed here is ready to go again, for the next festival and the next creative brief.
Thanks to the Monks technical team for their work on the pipelines.
Building a generative video pipeline of your own?
Explore Crusoe Cloud
, or
talk to our team
.
Latest articles
October 9, 2026
How green is Crusoe's inference engine?
October 9, 2026
Building community in the Lone Star State
October 7, 2026
GPU programming with Triton, part 1
Are you ready to build something amazing?
Contact Us