MEMORY INDUSTRY INTELLIGENCE

오픈 사이언스가 연구자들의 차기 팬데믹 대비를 어떻게 도울 수 있는가

한국어 번역·요약·분석

처리 완료Alibaba · deepseek-v4.1-flash · 원문 v2 · 10.09 23:26사용자 검토 전 초안

원문 제목: How Open Science Can Help Researchers Prepare for the Next Pandemic

원문 문장이 일치하지 않은 주장 3건은 근거 등록에서 제외했습니다.

핵심 요약

엔비디아는 구글 딥마인드, EMBL-EBI 등과 협력해 2,800종 이상 바이러스의 단백질 복합체 예측 3D 구조를 AlphaFold 데이터베이스를 통해 공개했다. 이 데이터셋은 AlphaFold2와 NVIDIA BioNeMo Inference Runtime 최적화로 생성되었으며, GPU 가속 워크플로인 BioNeMo Structure Prediction Pipeline도 오픈소스로 공개된다. 새로 추가된 단백질 상호작용의 약 30%는 과학계에 완전히 새로운 것이다. 이는 백신·치료제 개발을 위한 기초 생물학 지식 비축을 목표로 하며, 2050년까지 COVID-19급 팬데믹 발생 확률을 약 50%로 추정하는 분석에 대응한다. 이 프로젝트는 메모리 산업과 직접적 관련성이 확인되지 않는다.

메모리 산업 영향 분석

이 문서는 엔비디아가 주도하는 오픈 사이언스 및 AI 기반 단백질 구조 예측 프로젝트로, 메모리 산업과의 직접적 연결 근거는 확인되지 않는다. GPU 가속 워크플로(BioNeMo)가 언급되지만 이는 AI 추론 소프트웨어이며, HBM·DRAM·NAND 등 메모리 제품의 수요·공급·가격·기술 로드맵에 대한 구체적 언급이나 수치가 없다. 따라서 메모리 산업 영향은 미확인으로 분류한다. 관련 제품·고객·계층 이동·사용량 효과에 대한 분석은 원문 근거가 부족하여 작성하지 않는다.
한국어 번역 읽기

수집된 원문 v2의 전체 본문 기준 · 6205자

COVID-19가 출현했을 때 과학자들에게는 결정적인 이점이 있었다. 코로나바이러스에 대한 수십 년간의 선행 연구 덕분에 바이러스의 핵심 단백질을 충분히 이해하여 기록적인 시간 내에 백신을 설계할 수 있었다. 차기 팬데믹은 같은 출발선을 제공하지 않을 수 있다.

확률을 높이기 위해 엔비디아는 구글 딥마인드와 유럽 분자생물학 연구소의 유럽 생물정보학 연구소(EMBL-EBI)를 포함한 글로벌 연구 기관 연합에 합류하여 2,800종 이상의 바이러스 단백질 복합체에 대한 예측 3D 구조를 공개했다. 이는 전 세계 어느 과학자든 AlphaFold 데이터베이스를 통해 열람할 수 있다.

새로 공개된 데이터셋의 구조는 단백질이 3D 형태로 접히는 방식을 예측하는 구글 딥마인드의 AI 모델인 AlphaFold2를 사용하여 추론되었으며, NVIDIA BioNeMo Inference Runtime의 최적화가 적용되었다. 이를 통해 팀은 추론을 수천 개의 바이러스 프로테옴으로 확장하여 각 바이러스에 암호화된 복합체, 즉 상호작용하는 단백질 그룹을 예측할 수 있었다.

구글 딥마인드의 생명과학 파트너십 매니저인 리샤 파텔은 "AlphaFold 데이터베이스에 대한 우리의 야망은 항상 기초 생물학에 대한 접근을 대규모로 민주화하는 것이었다"며 "수천 개의 바이러스 복합체를 데이터베이스에 추가하는 이 협력은 전 세계 과학자들에게 향후 발병에 대비하는 데 필요한 통찰력을 제공할 것"이라고 말했다.

엔비디아는 또한 데이터셋 생성에 사용된 GPU 가속 워크플로인 BioNeMo Structure Prediction Pipeline을 공개적으로 배포하여 연구자들이 자체 표적에 대해 단백질 서열에서 예측 3D 구조로 이동할 수 있도록 한다.

차기 팬데믹 대비는 지금 시작해야 한다. 글로벌 개발 센터(Center for Global Development)의 분석에 따르면 2050년까지 세계가 COVID-19만큼 심각한 팬데믹에 직면할 확률은 약 50%로 추정된다.

글래스고 대학 의학연구위원회 바이러스 연구 센터의 분자 바이러스학 교수이자 이 프로젝트 협력자인 조 그로브는 "차기 팬데믹이 발생하면 예상치 못한 것이 나올 수 있고, 우리는 COVID에 대해 가졌던 지식이 부족할 것"이라며 "우리가 하려는 것은 그 지식의 일부를 미리 비축하는 것"이라고 말했다.

데이터베이스에 추가되는 단백질 상호작용의 약 30%는 과학계에 완전히 새로운 것으로, 실험적으로 결정된 단백질 구조의 주요 저장소인 단백질 데이터 뱅크(Protein Data Bank)에 한 번도 문서화되지 않은 상호작용 형태를 보여준다. 이는 생물학계가 새로운 지식을 생성하기 위해 탐구하고 활용할 수 있는 새로운 통찰력으로 이어진다.

엔비디아 디지털 생물학 응용 연구 과학 팀 리드인 크리스 달라고는 "이 데이터베이스는 가설 생성을 위한 엔진"이라며 "우리는 생물학자와 AI 커뮤니티가 단백질 상호작용을 단일 분자가 아닌 복합체로 조사할 수 있도록 하여 전체 분야가 발전할 수 있게 한다"고 말했다.

## **복합 단백질 구조 예측**

대부분의 단백질은 단독으로 작동하지 않는다. 여러 분자의 복합체로 모여 정교한 기능을 수행한다. 이러한 구조는 종종 백신이나 약물이 바이러스 기능을 방해하기 위해 표적으로 삼아야 하는 대상이다.

예를 들어 COVID-19 바이러스 스파이크 단백질의 3D 구조를 이해하는 것은 백신 설계의 기초가 되었다. 수천 개의 다른 바이러스에 대해서는 오늘날 그러한 구조적 지식이 존재하지 않는다. 이 데이터셋이 그 격차를 메우기 시작한다.

단백질 구조를 결정하는 전통적인 방법인 단백질 결정화 및 X선 조사는 구조당 수년이 걸리고 수천 달러의 비용이 들 수 있다. 엔비디아 GPU에서 효율적으로 실행되도록 NVIDIA BioNeMo로 최적화된 AlphaFold2는 몇 분 만에 구조를 예측하고 대량으로 실행할 수 있다. 과학자들은 이후 실험적 방법을 통해 높은 신뢰도의 예측을 검증할 수 있다.

이 프로젝트에서 팀은 일반 감기 바이러스부터 엠폭스(Mpox)와 같은 신종 위협에 이르기까지 인간을 감염시키는 것으로 알려진 바이러스 과의 단백질 구조를 체계적으로 작업했다.

## **글로벌 협력과 글로벌 접근**

이 협력은 전염병 대비 혁신 연합(CEPI), EMBL-EBI, 구글 딥마인드, 엔비디아, 서울대학교, 성균관대학교, 스위스 생물정보학 연구소, 글래스고 대학에 걸쳐 있다.

이번 데이터셋 공개는 이번 주 뉴욕에서 세계경제포럼이 소집한 유엔 총회 팬데믹 예방·대비·대응 회의와 동시에 이루어지며, AlphaFold 데이터베이스에 기여한다. 이 데이터베이스는 현재 과학계에 알려진 거의 모든 단백질을 포괄하는 2억 6천만 개 이상의 단백질 및 단백질 복합체 예측을 보유하고 있다.

EMBL-EBI의 임시 소장인 조 맥엔타이어는 "이 데이터를 공개하는 것은 바이러스 진단을 이해하고 치료제와 백신을 개발하는 데 중요하다"며 "이 데이터셋은 덜 연구된 바이러스도 다루며 발병에 직접 맞서는 저자원 환경의 과학자들의 장벽을 낮춘다"고 말했다.

공개 데이터셋의 예측은 신뢰도로 표시된다. 구조는 바이러스 복합체가 어떻게 보일 수 있는지, 개별 단백질이 바이러스 프로테옴 내에서 어떻게 상호작용할 수 있는지를 보여준다.

전반적으로 새로운 데이터는 디지털 생물학 및 질병 연구 전반의 과학자들이 이용할 수 있는 정보에 대한 주요 기여를 나타낸다.

그로브는 "박사 과정 때 우리가 조사하던 단백질 중 어떤 것에 대한 구조도 없었다. 어둠 속에서 작업하는 것과 같았고, 무슨 일이 일어나는지 추측해야 했다"며 "이 데이터셋은 지금 박사 과정을 밟는 모든 연구자에게 강력한 도구이며, 고품질 구조 데이터를 제공하여 기초 과학을 가속화할 것"이라고 말했다.

_AlphaFold 데이터베이스 팬데믹 대비 포털에서 바이러스 단백질 복합체 데이터셋을 탐색하고, BioNeMo Structure Prediction Pipeline으로 단백질 표적의 구조를 예측하며, NVIDIA BioNeMo에 대해 더 알아보세요._
브리프용 요약 초안
엔비디아가 구글 딥마인드, EMBL-EBI 등과 협력해 2,800종 이상 바이러스의 단백질 복합체 예측 3D 구조를 AlphaFold 데이터베이스에 공개했다. BioNeMo Inference Runtime으로 최적화된 AlphaFold2를 사용했으며, 약 30%는 과학계에 새로운 상호작용이다. 메모리 산업과의 직접적 관련성은 확인되지 않는다.

원문 텍스트

원문 열기 ↗
When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start.

To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI), to release predicted 3D structures for the protein complexes of more than 2,800 viruses — openly available to any scientist, anywhere, through the AlphaFold Database.

The structures in the newly released dataset were inferred using AlphaFold2 — Google DeepMind’s AI model for predicting how proteins fold into 3D shapes — with optimization from [NVIDIA BioNeMo Inference Runtime](https://docs.nvidia.com/bionemo/inference-runtime/overview). This allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus.

“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks.”

NVIDIA is also openly releasing the [BioNeMo Structure Prediction Pipeline](https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline), the GPU-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets.

Preparation for the next pandemic must begin now. An analysis by the Center for Global Development estimates a [roughly 50% chance](https://www.cgdev.org/blog/the-next-pandemic-could-come-soon-and-be-deadlier) of the world facing a pandemic as severe as COVID-19 by 2050.

“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is stockpile some of that knowledge ahead of time.”

About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures. This translates to new insights for the biological community to explore and harness to generate new knowledge.

“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward.”

## **Predicting Complex Protein Structures**

Most proteins don’t work alone — they come together in complexes of multiple molecules to perform sophisticated functions. Those structures are often what a vaccine or drug must target to disrupt viral function.

Understanding the 3D structure of the COVID-19 virus’ spike protein, for example, proved foundational to vaccine design. For thousands of other viruses, no such structural knowledge exists today. This dataset begins to fill that gap.

Traditional methods for determining protein structures — crystallizing proteins and shooting X-rays at them — can take years and cost thousands of dollars per structure. AlphaFold2, which was optimized with [NVIDIA BioNeMo](https://github.com/NVIDIA-BioNeMo) to efficiently run on NVIDIA GPUs, predicts a structure in minutes and can be run in bulk. Scientists can then verify high-confidence predictions through experimental methods.

For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common-cold viruses to emerging threats like Mpox.

## **A Global Collaboration With Global Access**

The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.

The dataset release — coinciding with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City — contributes to the AlphaFold Database, which now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science.

“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI. “The dataset also covers lesser-studied viruses and lowers the barriers for scientists in low-resource settings who are confronting outbreaks firsthand.”

Predictions in the open dataset are labeled by confidence. The structures show what viral complexes may look like and how individual proteins might interact within a viral proteome.

Overall, the new data represents a major contribution to the information available for scientists across digital biology and disease research.

“When I did my Ph.D., there were no structures for any of the proteins we were investigating. It was like working in the dark — we had to guess what was going on,” said Grove. “This dataset is a powerful tool for all the researchers doing their Ph.D.s now, giving them high-quality structural data that’s going to accelerate fundamental science.”

_Explore the viral protein complex dataset on the_[_AlphaFold Database Pandemic Preparedness Portal_](https://alphafold.ebi.ac.uk/)_, predict structures for protein targets with the_[_BioNeMo Structure Prediction Pipeline_](https://github.com/NVIDIA-BioNeMo/BioNeMo-Structure-Prediction-Pipeline)_, and learn more about_[_NVIDIA BioNeMo_](https://github.com/NVIDIA-BioNeMo)_._