MEMORY INDUSTRY INTELLIGENCE
추론 병목 우회: Retrieve-for-Train으로 복잡한 AI 검색 가속화
한국어 번역·요약·분석
원문 제목: Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
핵심 요약
이 글은 검색·추천 시스템이 단일 최적 결과가 아닌 상호 보완적인 결과 집합을 반환해야 한다는 요구에서 출발해, 제로샷 LLM의 쿼리 분해가 패러프레이즈 붕괴와 자기회귀 지연 병목을 겪는다고 주장한다. 저자들은 오프라인 강화학습으로 집합 수준 속성(근거성·다양성·정렬)에 부합하는 팬아웃을 탐색하고 이를 감독 데이터로 컴파일한 뒤, 53.9M 파라미터 확산 검색기로 증류하는 Retrieve-for-Train 프레임워크를 제안한다. 실험은 패션 데이터셋(CLIP 기반)과 독점 음악 플레이리스트 데이터셋(MuLan 기반)에서 Gemma3-4B·Qwen3-4B를 Soft-GRPO로 학습해 수행되었으며, 각 프롬프트당 정확히 10개의 하위 쿼리를 생성하도록 했다. 결과로 자기회귀 대비 12~20배 속도 향상과 대규모 컨텍스트 배치에서 자기회귀 팬아웃 지연이 약 50초까지 선형 확장되는 반면 확산 모델은 1초 미만에서 수 초 사이를 유지한다고 보고한다. 이는 연구·프레임워크 제안이며, 메모리 산업과의 직접적 연결은 원문에 명시되지 않았다.
메모리 산업 영향 분석
이 문서는 검색·추천 시스템의 쿼리 팬아웃을 위한 오프라인 RL + 확산 증류 프레임워크 연구로, 메모리 산업과의 직접 연결 근거는 원문에 전혀 없다. 추론 지연을 12~20배 줄이고 대규모 배치에서 자기회귀 지연이 약 50초까지 확장되는 반면 확산 모델은 1초 미만~수 초를 유지한다고 보고하지만, 이는 알고리즘·모델 아키텍처 변화이며 HBM·DRAM·SSD·HBF·CXL 등 메모리 계층 간 데이터/워크로드 위치 이동을 시사하지 않는다. 53.9M 파라미터 확산 검색기와 4B 팬아웃 모델은 GPU/가속기에서 실행될 것이나, 원문은 특정 하드웨어·메모리 용량·대역폭·고객·조달을 명시하지 않으므로 메모리 수요·계층 이동·사용량 효과를 추정할 수 없다. 분석가 가설로는 경량 추론 모델의 확산이 추론 인프라의 컴퓨팅·메모리 요구를 낮출 가능성이 있으나, 이는 원문 근거가 없는 미확인 항목이다. 확인할 지표: 배치 크기·컨텍스트 길이별 메모리 사용량, 임베딩 차원·데이터베이스 규모, 실제 배치 하드웨어.
한국어 번역 읽기
수집된 원문 v2의 전체 본문 기준 · 13587자
현대의 검색 또는 추천 애플리케이션은 점점 더 단일 최적 일치 항목이 아니라 일관성 있는 결과 집합을 반환할 것이 기대된다. 예를 들어 사용자가 "캠핑 장비"를 검색할 때, 그들은 4인용 텐트의 약간씩 다른 열 가지 변형을 원하지 않는다. 그들은 텐트, 침낭, 휴대용 스토브, 헤드램프 같은 필수 캠핑 장비를 포함하는 일관성 있고 상호 보완적인 슬레이트를 원한다.
이를 위해 시스템은 하나의 광범위한 프롬프트를 잠재적 사용자 관심사를 포괄하는 여러 관련 하위 쿼리로 분해하는 [쿼리 팬아웃](https://blog.google/products-and-platforms/products/search/ai-mode-search/) 기법을 사용한다. 그러나 LLM이 데이터베이스 인식 [쿼리 분해](https://www.emergentmind.com/topics/query-decomposition)를 동적으로 수행하도록 가르치는 것은 막대한 사고 예산을 소모한다. 설계상 [제로샷](https://www.promptingguide.ai/techniques/zeroshot) LLM은 범용 [자기회귀](https://aws.amazon.com/what-is/autoregressive-models/) 텍스트 예측기이며, 대상 코퍼스의 특정한 [기하학적 다양체](https://medium.com/@adnan.mazraeh1993/manifold-learning-and-geometry-based-approaches-a-comprehensive-explanation-7bc33d29cc04)를 탐색하도록 최적화되어 있지 않다. 결과적으로 이들은 고차 집합 수준 속성(예: 다양성, 포괄성, 상호 보완성, 일관성)을 최적화하면서 고정된 데이터베이스에 대해 근거를 유지하는 결과 모음을 반환하기 위해 확장된 테스트 시점 연산을 필요로 한다.
우리의 [ICML 2026](https://icml.cc/virtual/2026/poster/66354) 논문 "[Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion](https://arxiv.org/abs/2603.06397)"에서 우리는 보상-대-데이터 컴파일 프레임워크를 통해 이 분해 병목을 다룬다. 추론 시점에 모델이 큰 사고 예산을 소모하도록 강제하는 대신, 우리의 Retrieve-for-Train 프레임워크는 오프라인 강화학습(RL)을 사용해 보상에 정렬된 팬아웃을 발견하고 이를 감독 데이터로 컴파일한다. 이러한 최적화된 탐색 행동을 경량 확산 검색기로 증류함으로써, 우리는 추론 시점에 매우 효율적인 단일 패스 쿼리 팬아웃을 가능하게 한다. 이는 테스트 시점 사고 토큰의 오버헤드 없이 수학적으로 정식화된 집합 수준 속성을 달성한다.
## 일상적인 AI가 검색 전문가가 아닌 이유
복잡한 검색어 그룹을 브레인스토밍하는 작업이 주어졌을 때, 추론 시점에 표준 기성 LLM을 단순히 배치해 그 일을 처리하도록 하는 것은 유혹적이다. 그러나 데이터베이스 인식 쿼리 분해를 위해 범용 모델에 의존하는 것은 두 가지 중요한 문제를 야기한다:
1. _패러프레이즈 붕괴_**_:_** 데이터베이스 인식 최적화가 없으면 제로샷 LLM은 종종 [패러프레이즈 붕괴](https://arxiv.org/html/2605.04665v2)를 겪는다. 주제의 상호 보완적인 측면을 탐색하는 대신, 이들은 중복되고 거의 동의어인 쿼리를 생성하는 경향이 있다. 예를 들어 "보헤미안 페스티벌 스타일"이라는 광범위한 프롬프트가 주어지면, 신중한 프롬프트 엔지니어링이 없는 표준 LLM은 "보헤미안 페스티벌 패션"과 "보헤미안 페스티벌 의류"를 게으르게 생성할 수 있다. 이러한 의미적 반복은 동질적인 결과 슬레이트를 만들어내며, 패션 전문가가 식별할 프린지 재킷, 크로셰 드레스, 스웨이드 부츠 같은 뚜렷하고 유용한 의미적 방향을 완전히 놓친다.
2. _자기회귀 지연 병목_**_:_** 표준 LLM은 근본적으로 순차적 자기회귀 생성에 제약된다. 복잡한 쿼리를 상호 보완적인 측면으로 성공적으로 분해하기 위해, 현대 모델은 일반적으로 상당한 사고 예산을 필요로 하며, 실제 검색어를 출력하기 전에 확장을 계획하기 위해 수백 개의 중간 [사고 연쇄](https://blog.bluedot.org/p/faithful-chain-of-thought?utm_source=google&utm_medium=pmax&utm_campaign=FoAI_&utm_term=&utm_content=&gad_source=1&gad_campaignid=22554833691&gbraid=0AAAAA_kXnYMdypNLRSDdkci909e4FZdzE&gclid=Cj0KCQjwnbrUBhDOARIsAKKhPpe_x5PmJtrYhjcdmCSdnrM6xrWWnbjdI_8RAn6WxkMZ55NPyB399VMaAv0WEALw_wcB) (CoT) 추론 토큰(즉, AI 모델이 복잡한 질문에 답하기 전에 생성하는 중간 단계 또는 내부 처리 단위)을 생성한다. 이러한 의도적 추론은 대화형 AI에는 허용되지만, 집합값 검색(예: 위에서 언급한 프린지 재킷이나 크로셰 드레스 같은 상호 보완적인 결과 슬레이트 검색)에는 심각한 구조적 병목을 초래한다. 시스템이 대규모 하위 쿼리 슬레이트를 동시에 브레인스토밍해야 할 때, 연속적 컨텍스트 처리와 확장된 추론 토큰 생성의 결합 오버헤드는 확장성이 나쁘다. 고급 서빙 최적화가 있더라도, 이 토큰별 아키텍처는 프로덕션 검색창이 요구하는 1초 미만 응답 시간과 근본적으로 상충하는 지연 하한을 만든다.
## Retrieve-for-Train 프레임워크
Retrieve-for-Train은 AI의 학습을 사용자가 기다리는 동안 현장에서 치러야 하는 시험이 아니라 오프라인 연습 세션처럼 취급한다. AI가 좋은 검색의 규칙을 천천히 파악하고 누군가 쿼리를 입력할 때마다 막대한 처리 예산을 소모하도록 강제하는 대신, Retrieve-for-Train은 오프라인 RL 학습 프로그램을 한 번 실행한다.
이 프로그램은 엄격한 보상 시스템을 사용해 "결과가 다양하고 실제로 재고가 있도록 보장"과 같은 추상적 목표를 정확한 단계별 지침 매뉴얼로 바꾼다. 그 매뉴얼이 구축되면, AI는 실제 검색 중에 지연 없이 즉시 실행할 수 있다.
파이프라인은 세 가지 뚜렷한 단계로 작동한다:
* _팬아웃 언어 모델 학습:_ RL은 집합 수준 속성 검사 보상으로 점수가 매겨진 속성 정렬 하위 쿼리를 생성하도록 팬아웃 언어 모델을 학습시킨다. 이는 각 결과를 개별적으로 점수화하는 대신 전체 결과 그룹을 하나로 평가한다.
* _감독 데이터 합성:_ 동결된 팬아웃 언어 모델이 지도 학습을 위해 (쿼리 → 대상 집합) 쌍을 전적으로 오프라인으로 합성하며, 인간 레이블이 필요하지 않다.
* _확산 검색기 학습:_ 소형 53.9M 파라미터 확산 모델이 쿼리 임베딩을 하나의 비자기회귀 패스로 완전한 대상 임베딩 집합에 직접 매핑하는 법을 학습하여, 텍스트 기반 CoT 추론 토큰의 필요성을 공식적으로 우회한다.
### 집합을 위한 설계: 복합 보상의 힘
Retrieve-for-Train 프레임워크의 성공은 전적으로 "좋은" 검색 행동을 어떻게 정의하느냐에 달려 있다. 전통적인 지도 학습은 [랭킹 학습을 통한 점별 관련성](https://en.wikipedia.org/wiki/Learning_to_rank)을 평가하여 각 검색 항목을 개별적으로 점수화한다. 그러나 진정한 전문가 검색 슬레이트는 분해 불가능한 집합 수준 속성으로 정의된다. 단일 항목의 다양성이나 상호 보완성을 측정할 수 없다; 이러한 속성은 검색된 결과의 전체 모음을 평가할 때만 수학적으로 존재한다.
이러한 팬아웃 속성을 강제하기 위해 모호한 자연어 지시에 의존하는 대신, Retrieve-for-Train은 엄격한 수학적 복합 보상을 사용한 강화학습을 통해 4B 오픈소스 언어 모델([Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) 및 [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B))을 미세 조정한다. 우리의 개방형 추상 검색 작업에서 이 복합 보상은 세 가지 경쟁 축의 가중 균형이다:
* _근거성:_ 데이터베이스 다양체까지의 거리를 벌점화하여, 생성된 모든 하위 쿼리가 데이터베이스에서 실제로 검색 가능한 항목에 대응하도록 보장한다.
* _다양성:_ 전체 하위 쿼리 집합에 대해 [Vendi Score](https://arxiv.org/abs/2210.02410)로 측정되며, 모델이 넓은 의미적 폭을 탐색하도록 강제한다.
* _정렬:_ 후보 하위 쿼리를 원래의 광범위한 프롬프트에 고정시켜 의미적 표류를 방지한다.
### 상호 대항 앵커와 소프트-GRPO 학습
학습 중에 우리는 소프트 [근접 정책 최적화](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88%5C) (PPO)와 함께 [그룹 상대 정책 최적화](https://arxiv.org/abs/2505.22257) (GRPO)를 사용하여 이러한 기하학적 현실에 대해 팬아웃 언어 모델을 최적화한다.
이 특정한 세 가지 보상 조합은 이들이 상호 대항 앵커로 작용하기 때문에 중요하다. 모델이 순전히 근거성에 대해서만 최적화되면, 우연히 특정 데이터베이스 좌표에 수학적으로 매핑되는 퇴화되고 무의미한 문자열을 생성하여 시스템을 보상 해킹할 것이다. 무의미함을 고치기 위해 정렬을 추가하면, 정책은 단순히 사용자 프롬프트의 반복적인 패러프레이즈로 붕괴하여 속임수를 쓴다.
Vendi Score를 대항 앵커로 주입함으로써, Retrieve-for-Train은 이러한 지름길 해법을 효과적으로 차단한다. 높은 보상 상태를 달성하기 위해, 정책은 원래 의도의 유효하고 엄격하게 근거를 둔, 그러나 의미적으로 구별되는 변형을 발견해야 하는 임베딩 공간의 균형 잡힌 영역으로 강제된다.
## 실험
Retrieve-for-Train 프레임워크를 평가하기 위해, 우리는 동결된 데이터셋별 멀티모달 임베딩 백본과 쿼리 확장에 최적화된 [오픈소스 언어 모델](https://deepmind.google/models/gemma/gemma-3/)의 조합을 사용했다. 우리는 이 설정을 두 가지 뚜렷한 집합값 검색 체제에 걸쳐 평가했다:
* _개방형 추상 검색:_ 고유한 정답이 존재하지 않고 품질이 전적으로 다양성, 쿼리 정렬, 데이터베이스 근거성을 포함한 집합 수준 속성으로 측정되는 설정.
* _약지도 조성 검색:_ 쿼리가 쿼리 의도의 하나의 그럴듯한 실현으로서 역할을 하는 약한 참조 집합과 쌍을 이루는 설정.
멀티모달 임베딩 백본의 경우, 우리는 두 도메인에 걸쳐 실험을 수행했다: 텍스트-이미지 실험에 사용된 사용자 큐레이션 의상의 대규모 패션 데이터셋([CLIP](https://openai.com/index/clip/) 기반 검색기로 평가), 그리고 텍스트-음악 평가에 사용된 전문가 생성 음악 플레이리스트의 독점 산업 데이터셋([MuLan](https://research.google/pubs/mulan-a-joint-embedding-of-music-audio-and-natural-language/)으로 평가).
언어 모델의 경우, 쿼리 팬아웃 과정은 4B 오픈소스 모델, 구체적으로 [Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it)와 [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)에 의해 구동되었으며, 이들은 처리하는 모든 단일 주요 검색 프롬프트에 대해 정확히 10개의 하위 쿼리를 생성하는 임무를 맡았다. 우리는 이러한 팬아웃 모델에 대한 RL 학습을 [Soft-GRPO](https://arxiv.org/abs/2511.06411)를 통해 구현했으며, 이는 [소프트 PPO 정규화](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88)와 함께 [그룹 상대 정책 최적화](https://cameronrwolfe.substack.com/p/grpo)를 사용하는 접근법이다.
## 결과
### 검색 품질과 정확도
두 검색 작업 모두에서 Retrieve-for-Train은 전통적인 단일 쿼리 검색, 제로샷 확장, 심지어 고도로 최적화된 [Best-of-N 기준선](https://openai.com/index/measuring-goodharts-law/)을 능가했다.
정성적으로, 제로샷 LLM 기준선은 거의 동의어인 패러프레이즈(예: "보헤미안 페스티벌 스타일" 대 "보헤미안 페스티벌 패션")를 생성하는 경향이 있어 중복 결과를 초래했다. Retrieve-for-Train은 데이터베이스 다양체 내에 엄격하게 근거를 두면서도 매우 다양하고 뚜렷한 하위 쿼리(예: "부츠" 또는 "레이스"로 분기)를 생성했다.
### 압도적으로 빠른 추론
RL로 튜닝된 언어 모델을 직접 배치하면 뛰어난 검색 품질을 얻을 수 있었지만, 이는 표준 자기회귀 지연 제약을 물려받았고 높은 계산 사고 예산을 요구했다.
그 학습된 행동을 53.9M 파라미터 Retrieve-for-Train 확산 모델로 증류함으로써, 우리는 지연 병목을 성공적으로 깨뜨렸다. 확산 모델은 연속 임베딩 공간에서 단일 비자기회귀 병렬 패스로 모든 대상 방향을 동시에 생성하기 때문에, 자기회귀 접근법 대비 12~20배의 막대한 속도 향상을 제공한다.
규모에서, 자기회귀 팬아웃 지연이 대규모 컨텍스트 배치에서 거의 50초까지 선형적으로 확장되는 반면, Retrieve-for-Train-Diffusion은 1초 미만에서 수 초 사이를 유지하며, 계산 비용의 일부로 프로덕션 준비가 된 전문가 수준 검색을 제공한다.
### 안티 해킹 앵커 (절제 연구 통찰)
우리의 보상 최적화 과정에서 우리는 검색을 위한 팬아웃 언어 모델 학습에 관한 근본적인 무언가를 발견했다. 다양성 항이 없으면, 모델은 데이터베이스의 벡터 좌표를 수학적으로 악용하기 위해 퇴화되고 무의미한 문자열(예: _"line ending line ending"_)을 생성하는 것으로 빠르게 붕괴한다. 기하학적 다양성 지표(Vendi Score)를 주입하는 것은 중요한 대항 앵커로 작용하여, 모델이 진정한 검색 전문가처럼 행동함으로써만 보상을 극대화할 수 있는 임베딩 공간의 안정적인 영역으로 강제한다.
## 결론
우리는 RL이 온라인 추론 엔진이 아니라 일회성 "목표 변환기"로 사용될 때 매우 효과적일 수 있음을 입증했다. 보상 기반 행동 탐색의 무거운 계산을 최종 배치 모델로부터 분리함으로써, 우리의 프레임워크는 온라인 LLM 배치의 전형적인 가파른 추론 지연과 높은 계산 오버헤드를 성공적으로 우회한다.
이러한 복잡한 집합 수준 행동을 경량 확산 사전으로 증류하면, 프로덕션 검색 시스템이 다양성과 정렬 같은 고차 속성을 효과적으로 최적화할 수 있다. 궁극적으로 Retrieve-for-Train은 인간 레이블이 붙은 속성 정렬 학습 쌍이 부족하거나 비용이 많이 드는 특수 또는 멀티모달 도메인에서 집합 검색을 위한 고도로 확장 가능하고 데이터 효율적인 파이프라인을 확립한다. 자세한 내용은 [논문](https://arxiv.org/abs/2603.06397)을 참조하라.
이를 위해 시스템은 하나의 광범위한 프롬프트를 잠재적 사용자 관심사를 포괄하는 여러 관련 하위 쿼리로 분해하는 [쿼리 팬아웃](https://blog.google/products-and-platforms/products/search/ai-mode-search/) 기법을 사용한다. 그러나 LLM이 데이터베이스 인식 [쿼리 분해](https://www.emergentmind.com/topics/query-decomposition)를 동적으로 수행하도록 가르치는 것은 막대한 사고 예산을 소모한다. 설계상 [제로샷](https://www.promptingguide.ai/techniques/zeroshot) LLM은 범용 [자기회귀](https://aws.amazon.com/what-is/autoregressive-models/) 텍스트 예측기이며, 대상 코퍼스의 특정한 [기하학적 다양체](https://medium.com/@adnan.mazraeh1993/manifold-learning-and-geometry-based-approaches-a-comprehensive-explanation-7bc33d29cc04)를 탐색하도록 최적화되어 있지 않다. 결과적으로 이들은 고차 집합 수준 속성(예: 다양성, 포괄성, 상호 보완성, 일관성)을 최적화하면서 고정된 데이터베이스에 대해 근거를 유지하는 결과 모음을 반환하기 위해 확장된 테스트 시점 연산을 필요로 한다.
우리의 [ICML 2026](https://icml.cc/virtual/2026/poster/66354) 논문 "[Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion](https://arxiv.org/abs/2603.06397)"에서 우리는 보상-대-데이터 컴파일 프레임워크를 통해 이 분해 병목을 다룬다. 추론 시점에 모델이 큰 사고 예산을 소모하도록 강제하는 대신, 우리의 Retrieve-for-Train 프레임워크는 오프라인 강화학습(RL)을 사용해 보상에 정렬된 팬아웃을 발견하고 이를 감독 데이터로 컴파일한다. 이러한 최적화된 탐색 행동을 경량 확산 검색기로 증류함으로써, 우리는 추론 시점에 매우 효율적인 단일 패스 쿼리 팬아웃을 가능하게 한다. 이는 테스트 시점 사고 토큰의 오버헤드 없이 수학적으로 정식화된 집합 수준 속성을 달성한다.
## 일상적인 AI가 검색 전문가가 아닌 이유
복잡한 검색어 그룹을 브레인스토밍하는 작업이 주어졌을 때, 추론 시점에 표준 기성 LLM을 단순히 배치해 그 일을 처리하도록 하는 것은 유혹적이다. 그러나 데이터베이스 인식 쿼리 분해를 위해 범용 모델에 의존하는 것은 두 가지 중요한 문제를 야기한다:
1. _패러프레이즈 붕괴_**_:_** 데이터베이스 인식 최적화가 없으면 제로샷 LLM은 종종 [패러프레이즈 붕괴](https://arxiv.org/html/2605.04665v2)를 겪는다. 주제의 상호 보완적인 측면을 탐색하는 대신, 이들은 중복되고 거의 동의어인 쿼리를 생성하는 경향이 있다. 예를 들어 "보헤미안 페스티벌 스타일"이라는 광범위한 프롬프트가 주어지면, 신중한 프롬프트 엔지니어링이 없는 표준 LLM은 "보헤미안 페스티벌 패션"과 "보헤미안 페스티벌 의류"를 게으르게 생성할 수 있다. 이러한 의미적 반복은 동질적인 결과 슬레이트를 만들어내며, 패션 전문가가 식별할 프린지 재킷, 크로셰 드레스, 스웨이드 부츠 같은 뚜렷하고 유용한 의미적 방향을 완전히 놓친다.
2. _자기회귀 지연 병목_**_:_** 표준 LLM은 근본적으로 순차적 자기회귀 생성에 제약된다. 복잡한 쿼리를 상호 보완적인 측면으로 성공적으로 분해하기 위해, 현대 모델은 일반적으로 상당한 사고 예산을 필요로 하며, 실제 검색어를 출력하기 전에 확장을 계획하기 위해 수백 개의 중간 [사고 연쇄](https://blog.bluedot.org/p/faithful-chain-of-thought?utm_source=google&utm_medium=pmax&utm_campaign=FoAI_&utm_term=&utm_content=&gad_source=1&gad_campaignid=22554833691&gbraid=0AAAAA_kXnYMdypNLRSDdkci909e4FZdzE&gclid=Cj0KCQjwnbrUBhDOARIsAKKhPpe_x5PmJtrYhjcdmCSdnrM6xrWWnbjdI_8RAn6WxkMZ55NPyB399VMaAv0WEALw_wcB) (CoT) 추론 토큰(즉, AI 모델이 복잡한 질문에 답하기 전에 생성하는 중간 단계 또는 내부 처리 단위)을 생성한다. 이러한 의도적 추론은 대화형 AI에는 허용되지만, 집합값 검색(예: 위에서 언급한 프린지 재킷이나 크로셰 드레스 같은 상호 보완적인 결과 슬레이트 검색)에는 심각한 구조적 병목을 초래한다. 시스템이 대규모 하위 쿼리 슬레이트를 동시에 브레인스토밍해야 할 때, 연속적 컨텍스트 처리와 확장된 추론 토큰 생성의 결합 오버헤드는 확장성이 나쁘다. 고급 서빙 최적화가 있더라도, 이 토큰별 아키텍처는 프로덕션 검색창이 요구하는 1초 미만 응답 시간과 근본적으로 상충하는 지연 하한을 만든다.
## Retrieve-for-Train 프레임워크
Retrieve-for-Train은 AI의 학습을 사용자가 기다리는 동안 현장에서 치러야 하는 시험이 아니라 오프라인 연습 세션처럼 취급한다. AI가 좋은 검색의 규칙을 천천히 파악하고 누군가 쿼리를 입력할 때마다 막대한 처리 예산을 소모하도록 강제하는 대신, Retrieve-for-Train은 오프라인 RL 학습 프로그램을 한 번 실행한다.
이 프로그램은 엄격한 보상 시스템을 사용해 "결과가 다양하고 실제로 재고가 있도록 보장"과 같은 추상적 목표를 정확한 단계별 지침 매뉴얼로 바꾼다. 그 매뉴얼이 구축되면, AI는 실제 검색 중에 지연 없이 즉시 실행할 수 있다.
파이프라인은 세 가지 뚜렷한 단계로 작동한다:
* _팬아웃 언어 모델 학습:_ RL은 집합 수준 속성 검사 보상으로 점수가 매겨진 속성 정렬 하위 쿼리를 생성하도록 팬아웃 언어 모델을 학습시킨다. 이는 각 결과를 개별적으로 점수화하는 대신 전체 결과 그룹을 하나로 평가한다.
* _감독 데이터 합성:_ 동결된 팬아웃 언어 모델이 지도 학습을 위해 (쿼리 → 대상 집합) 쌍을 전적으로 오프라인으로 합성하며, 인간 레이블이 필요하지 않다.
* _확산 검색기 학습:_ 소형 53.9M 파라미터 확산 모델이 쿼리 임베딩을 하나의 비자기회귀 패스로 완전한 대상 임베딩 집합에 직접 매핑하는 법을 학습하여, 텍스트 기반 CoT 추론 토큰의 필요성을 공식적으로 우회한다.
### 집합을 위한 설계: 복합 보상의 힘
Retrieve-for-Train 프레임워크의 성공은 전적으로 "좋은" 검색 행동을 어떻게 정의하느냐에 달려 있다. 전통적인 지도 학습은 [랭킹 학습을 통한 점별 관련성](https://en.wikipedia.org/wiki/Learning_to_rank)을 평가하여 각 검색 항목을 개별적으로 점수화한다. 그러나 진정한 전문가 검색 슬레이트는 분해 불가능한 집합 수준 속성으로 정의된다. 단일 항목의 다양성이나 상호 보완성을 측정할 수 없다; 이러한 속성은 검색된 결과의 전체 모음을 평가할 때만 수학적으로 존재한다.
이러한 팬아웃 속성을 강제하기 위해 모호한 자연어 지시에 의존하는 대신, Retrieve-for-Train은 엄격한 수학적 복합 보상을 사용한 강화학습을 통해 4B 오픈소스 언어 모델([Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) 및 [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B))을 미세 조정한다. 우리의 개방형 추상 검색 작업에서 이 복합 보상은 세 가지 경쟁 축의 가중 균형이다:
* _근거성:_ 데이터베이스 다양체까지의 거리를 벌점화하여, 생성된 모든 하위 쿼리가 데이터베이스에서 실제로 검색 가능한 항목에 대응하도록 보장한다.
* _다양성:_ 전체 하위 쿼리 집합에 대해 [Vendi Score](https://arxiv.org/abs/2210.02410)로 측정되며, 모델이 넓은 의미적 폭을 탐색하도록 강제한다.
* _정렬:_ 후보 하위 쿼리를 원래의 광범위한 프롬프트에 고정시켜 의미적 표류를 방지한다.
### 상호 대항 앵커와 소프트-GRPO 학습
학습 중에 우리는 소프트 [근접 정책 최적화](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88%5C) (PPO)와 함께 [그룹 상대 정책 최적화](https://arxiv.org/abs/2505.22257) (GRPO)를 사용하여 이러한 기하학적 현실에 대해 팬아웃 언어 모델을 최적화한다.
이 특정한 세 가지 보상 조합은 이들이 상호 대항 앵커로 작용하기 때문에 중요하다. 모델이 순전히 근거성에 대해서만 최적화되면, 우연히 특정 데이터베이스 좌표에 수학적으로 매핑되는 퇴화되고 무의미한 문자열을 생성하여 시스템을 보상 해킹할 것이다. 무의미함을 고치기 위해 정렬을 추가하면, 정책은 단순히 사용자 프롬프트의 반복적인 패러프레이즈로 붕괴하여 속임수를 쓴다.
Vendi Score를 대항 앵커로 주입함으로써, Retrieve-for-Train은 이러한 지름길 해법을 효과적으로 차단한다. 높은 보상 상태를 달성하기 위해, 정책은 원래 의도의 유효하고 엄격하게 근거를 둔, 그러나 의미적으로 구별되는 변형을 발견해야 하는 임베딩 공간의 균형 잡힌 영역으로 강제된다.
## 실험
Retrieve-for-Train 프레임워크를 평가하기 위해, 우리는 동결된 데이터셋별 멀티모달 임베딩 백본과 쿼리 확장에 최적화된 [오픈소스 언어 모델](https://deepmind.google/models/gemma/gemma-3/)의 조합을 사용했다. 우리는 이 설정을 두 가지 뚜렷한 집합값 검색 체제에 걸쳐 평가했다:
* _개방형 추상 검색:_ 고유한 정답이 존재하지 않고 품질이 전적으로 다양성, 쿼리 정렬, 데이터베이스 근거성을 포함한 집합 수준 속성으로 측정되는 설정.
* _약지도 조성 검색:_ 쿼리가 쿼리 의도의 하나의 그럴듯한 실현으로서 역할을 하는 약한 참조 집합과 쌍을 이루는 설정.
멀티모달 임베딩 백본의 경우, 우리는 두 도메인에 걸쳐 실험을 수행했다: 텍스트-이미지 실험에 사용된 사용자 큐레이션 의상의 대규모 패션 데이터셋([CLIP](https://openai.com/index/clip/) 기반 검색기로 평가), 그리고 텍스트-음악 평가에 사용된 전문가 생성 음악 플레이리스트의 독점 산업 데이터셋([MuLan](https://research.google/pubs/mulan-a-joint-embedding-of-music-audio-and-natural-language/)으로 평가).
언어 모델의 경우, 쿼리 팬아웃 과정은 4B 오픈소스 모델, 구체적으로 [Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it)와 [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)에 의해 구동되었으며, 이들은 처리하는 모든 단일 주요 검색 프롬프트에 대해 정확히 10개의 하위 쿼리를 생성하는 임무를 맡았다. 우리는 이러한 팬아웃 모델에 대한 RL 학습을 [Soft-GRPO](https://arxiv.org/abs/2511.06411)를 통해 구현했으며, 이는 [소프트 PPO 정규화](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88)와 함께 [그룹 상대 정책 최적화](https://cameronrwolfe.substack.com/p/grpo)를 사용하는 접근법이다.
## 결과
### 검색 품질과 정확도
두 검색 작업 모두에서 Retrieve-for-Train은 전통적인 단일 쿼리 검색, 제로샷 확장, 심지어 고도로 최적화된 [Best-of-N 기준선](https://openai.com/index/measuring-goodharts-law/)을 능가했다.
정성적으로, 제로샷 LLM 기준선은 거의 동의어인 패러프레이즈(예: "보헤미안 페스티벌 스타일" 대 "보헤미안 페스티벌 패션")를 생성하는 경향이 있어 중복 결과를 초래했다. Retrieve-for-Train은 데이터베이스 다양체 내에 엄격하게 근거를 두면서도 매우 다양하고 뚜렷한 하위 쿼리(예: "부츠" 또는 "레이스"로 분기)를 생성했다.
### 압도적으로 빠른 추론
RL로 튜닝된 언어 모델을 직접 배치하면 뛰어난 검색 품질을 얻을 수 있었지만, 이는 표준 자기회귀 지연 제약을 물려받았고 높은 계산 사고 예산을 요구했다.
그 학습된 행동을 53.9M 파라미터 Retrieve-for-Train 확산 모델로 증류함으로써, 우리는 지연 병목을 성공적으로 깨뜨렸다. 확산 모델은 연속 임베딩 공간에서 단일 비자기회귀 병렬 패스로 모든 대상 방향을 동시에 생성하기 때문에, 자기회귀 접근법 대비 12~20배의 막대한 속도 향상을 제공한다.
규모에서, 자기회귀 팬아웃 지연이 대규모 컨텍스트 배치에서 거의 50초까지 선형적으로 확장되는 반면, Retrieve-for-Train-Diffusion은 1초 미만에서 수 초 사이를 유지하며, 계산 비용의 일부로 프로덕션 준비가 된 전문가 수준 검색을 제공한다.
### 안티 해킹 앵커 (절제 연구 통찰)
우리의 보상 최적화 과정에서 우리는 검색을 위한 팬아웃 언어 모델 학습에 관한 근본적인 무언가를 발견했다. 다양성 항이 없으면, 모델은 데이터베이스의 벡터 좌표를 수학적으로 악용하기 위해 퇴화되고 무의미한 문자열(예: _"line ending line ending"_)을 생성하는 것으로 빠르게 붕괴한다. 기하학적 다양성 지표(Vendi Score)를 주입하는 것은 중요한 대항 앵커로 작용하여, 모델이 진정한 검색 전문가처럼 행동함으로써만 보상을 극대화할 수 있는 임베딩 공간의 안정적인 영역으로 강제한다.
## 결론
우리는 RL이 온라인 추론 엔진이 아니라 일회성 "목표 변환기"로 사용될 때 매우 효과적일 수 있음을 입증했다. 보상 기반 행동 탐색의 무거운 계산을 최종 배치 모델로부터 분리함으로써, 우리의 프레임워크는 온라인 LLM 배치의 전형적인 가파른 추론 지연과 높은 계산 오버헤드를 성공적으로 우회한다.
이러한 복잡한 집합 수준 행동을 경량 확산 사전으로 증류하면, 프로덕션 검색 시스템이 다양성과 정렬 같은 고차 속성을 효과적으로 최적화할 수 있다. 궁극적으로 Retrieve-for-Train은 인간 레이블이 붙은 속성 정렬 학습 쌍이 부족하거나 비용이 많이 드는 특수 또는 멀티모달 도메인에서 집합 검색을 위한 고도로 확장 가능하고 데이터 효율적인 파이프라인을 확립한다. 자세한 내용은 [논문](https://arxiv.org/abs/2603.06397)을 참조하라.
브리프용 요약 초안
ICML 2026 논문은 오프라인 RL로 팬아웃을 컴파일하고 53.9M 확산 검색기로 증류해 자기회귀 대비 12~20배 빠른 집합 검색을 보고했다. 다만 원문은 특정 하드웨어·메모리 용량·고객을 명시하지 않아 메모리 산업 영향은 직접 근거가 없다.
원문 텍스트
원문 열기 ↗Modern search or recommendation applications are increasingly expected to return a coherent set of results rather than a single best match. For example, when a user searches for "camping gear", they don’t want ten slight variations of four-person tents. They want a coherent, complementary slate that includes essential camping gear, such as a tent, sleeping bag, portable stove, and headlamp.
To do this, systems use a [query fan-out](https://blog.google/products-and-platforms/products/search/ai-mode-search/) technique that breaks a single broad prompt into several related sub-queries to cover potential user interests. However, teaching an LLM to perform database-aware [query decomposition](https://www.emergentmind.com/topics/query-decomposition) dynamically drains a massive thinking budget. By design, [zero-shot](https://www.promptingguide.ai/techniques/zeroshot) LLMs are general [autoregressive](https://aws.amazon.com/what-is/autoregressive-models/) text predictors; they aren’t optimized to navigate the specific, [geometric manifold](https://medium.com/@adnan.mazraeh1993/manifold-learning-and-geometry-based-approaches-a-comprehensive-explanation-7bc33d29cc04) of a target corpus. Consequently, they need extended test-time computation to return a collection of results that optimizes higher-order set-level properties (e.g., diversity, coverage, complementarity, coherence) while remaining grounded with respect to a fixed database.
In our [ICML 2026](https://icml.cc/virtual/2026/poster/66354) paper, “[Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion](https://arxiv.org/abs/2603.06397)”, we address this decomposition bottleneck via a reward-to-data compilation framework. Instead of forcing the model to expend a large thinking budget at inference, our Retrieve-for-Train framework uses offline reinforcement learning (RL) to discover reward-aligned fan-outs and compile them into supervision. By distilling these optimized exploration behaviors into a lightweight diffusion retriever, we enable highly efficient, single-pass query fan-out at inference time. This achieves mathematically formulated, set-level properties without the overhead of test-time thinking tokens.
## Why everyday AI isn't a search expert
When tasked with brainstorming a complex group of search terms, it’s tempting to simply deploy a standard, off-the-shelf LLM at inference time to handle the job. However, relying on generic models for database-aware query decomposition introduces two critical challenges:
1. _Paraphrastic collapse_**_:_** Without database-aware optimization, zero-shot LLMs frequently suffer from [paraphrastic collapse](https://arxiv.org/html/2605.04665v2). Rather than exploring complementary facets of a topic, they tend to generate redundant, near-synonymous queries. For example, given the broad prompt "Bohemian festival style”, a standard LLM without careful prompt engineering might lazily generate "bohemian festival fashion" and "bohemian festival clothes”. This semantic looping produces a homogeneous slate of results, entirely missing the distinct, helpful semantic directions a fashion expert would identify, such as fringe jackets, crochet dresses, or suede boots.
2. _Autoregressive latency bottlenecks_**_:_** Standard LLMs are fundamentally constrained by sequential, autoregressive generation. To successfully decompose a complex query into complementary facets, modern models typically require a substantial thinking budget, generating hundreds of intermediate [chain-of-thought](https://blog.bluedot.org/p/faithful-chain-of-thought?utm_source=google&utm_medium=pmax&utm_campaign=FoAI_&utm_term=&utm_content=&gad_source=1&gad_campaignid=22554833691&gbraid=0AAAAA_kXnYMdypNLRSDdkci909e4FZdzE&gclid=Cj0KCQjwnbrUBhDOARIsAKKhPpe_x5PmJtrYhjcdmCSdnrM6xrWWnbjdI_8RAn6WxkMZ55NPyB399VMaAv0WEALw_wcB) (CoT) reasoning tokens (i,e., the intermediate steps or internal processing units an AI model generates before answering a complex question) to plan their expansion before outputting the actual search terms. While this deliberate reasoning is acceptable for conversational AI, it introduces a severe structural bottleneck for set-valued search (e.g., retrieving a complementary slate of results, such as fringe jackets or crochet dresses mentioned above). When a system must brainstorm a large slate of sub-queries simultaneously, the combined overhead of continuous context processing and generating extended reasoning tokens scales poorly. Even with advanced serving optimizations, this token-by-token architecture creates a latency floor that is fundamentally at odds with the sub-second response times required by a production search bar.
## The Retrieve-for-Train framework
The Retrieve-for-Train treats the AI's training like an offline practice session rather than a test it has to take on the spot while a user is waiting. Instead of forcing the AI to slowly figure out the rules of a good search and drain a massive processing budget every single time someone types a query, Retrieve-for-Train runs an offline RL training program once.
This program uses a rigorous reward system to turn abstract goals like "ensure the results are diverse and actually in stock" into an exact step-by-step instruction manual. Once that manual is built, the AI can execute it instantly during a real search without delay.
The pipeline operates in three distinct steps:
* _Fan-out language model training:_ RL trains a fan-out language model to emit property-aligned sub-queries scored by a set-level property-check reward. This evaluates the entire group of results as a whole, rather than scoring each result in isolation.
* _Supervision synthesis:_ The frozen fan-out language model synthesizes (query → target-set) pairs entirely offline for supervised learning, requiring no human labels.
* _Diffusive retriever training:_ A compact, 53.9M-parameter diffusion model learns to map a query embedding directly to a complete set of target embeddings in one non-autoregressive pass, officially bypassing the need for text-based CoT reasoning tokens.
### Designing for the set: The power of composite rewards
The success of the Retrieve-for-Train framework hinges entirely on how we define "good" search behavior. Traditional supervised training evaluates [pointwise relevance via learning to rank](https://en.wikipedia.org/wiki/Learning_to_rank), scoring each retrieved item in isolation. However, a truly expert search slate is defined by non-decomposable, set-level properties. You can’t measure the diversity or complementarity of a single item; these properties only exist mathematically when evaluating the entire collection of retrieved results.
Rather than relying on ambiguous natural language instructions to enforce these fan-out properties, Retrieve-for-Train fine-tunes the 4B open-source language models ([Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) and [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)) via reinforcement learning using a strict mathematical composite reward. For our open-ended abstract retrieval tasks, this composite reward is a weighted balance of three competing pillars:
* _Groundedness:_ Penalizes distance to the database manifold, ensuring every generated sub-query corresponds to a real, retrievable item in the database.
* _Diversity:_ Measured using the [Vendi Score](https://arxiv.org/abs/2210.02410) over the entire set of sub-queries, forcing the model to explore broad semantic breadth.
* _Alignment:_ Anchors candidate sub-queries to the original broad prompt to prevent semantic drift.
### Mutual counter-anchors and soft-GRPO training
During training, we optimize the fan-out language model against these geometric realities using [group relative policy optimization](https://arxiv.org/abs/2505.22257) (GRPO) with soft [proximal policy optimization](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88%5C) (PPO).
This specific triad of rewards is critical because they act as mutual counter-anchors. If a model is optimized purely for groundedness, it will reward-hack the system by generating degenerate, nonsensical strings that happen to mathematically map to a specific database coordinate. If alignment is added to fix the nonsense, the policy simply cheats by collapsing into repetitive paraphrases of the user's prompt.
By injecting the Vendi Score as a counter-anchor, Retrieve-for-Train effectively closes off these shortcut solutions. To achieve a high-reward state, the policy is forced into a balanced region of the embedding space where it must discover valid, strictly grounded, yet semantically distinct variations of the original intent.
## Experiments
To evaluate the Retrieve-for-Train framework, we used a combination of frozen, dataset-specific multimodal embedding backbones and [open-source language models](https://deepmind.google/models/gemma/gemma-3/) optimized for query expansion. We evaluated this setup across two distinct set-valued retrieval regimes:
* _Open-ended abstract retrieval:_ A setting where no unique ground truth exists and quality is exclusively measured by set-level properties, including diversity, query alignment, and database groundedness.
* _Weakly supervised compositional retrieval:_ A setting where queries are paired with a weak reference set that serves as just one plausible realization of the query intent.
For the multimodal embedding backbones, we conducted experiments across two domains: A large-scale fashion dataset of user-curated outfits used for text-to-image experiments (evaluated using a [CLIP](https://openai.com/index/clip/)-based retriever), and a proprietary industrial dataset of expert-generated music playlists used for text-to-music evaluations (evaluated using [MuLan](https://research.google/pubs/mulan-a-joint-embedding-of-music-audio-and-natural-language/)).
For the language models, the query fan-out process was driven by 4B open-source models, specifically [Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) and [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B), which were tasked with generating exactly 10 sub-queries for every single main search prompt they processed. We implemented the RL training for these fan-out models via [Soft-GRPO](https://arxiv.org/abs/2511.06411), an approach that uses [group relative policy optimization](https://cameronrwolfe.substack.com/p/grpo) with [soft PPO regularization](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88).
## Results
### Retrieval quality and accuracy
Across both retrieval tasks, Retrieve-for-Train outperformed traditional single-query search, zero-shot expansion, and even the heavily optimized [Best-of-N baseline](https://openai.com/index/measuring-goodharts-law/).
Qualitatively, zero-shot LLM baselines tended to generate near-synonymous paraphrases (e.g., "bohemian festival style" vs. "bohemian festival fashion"), causing redundant results. Retrieve-for-Train generated highly diverse, distinct sub-queries (e.g., branching into "boots" or "lace") that remained strictly grounded within the database manifold.
### Order-of-magnitude faster inference
Directly deploying our RL-tuned language model yielded exceptional search quality, but it inherited standard autoregressive latency constraints and demanded a high computational thinking budget.
By distilling that learned behavior into the 53.9M-parameter Retrieve-for-Train diffusion model, we successfully smashed the latency bottleneck. Because the diffusion model generates all target directions simultaneously in a single, non-autoregressive parallel pass in continuous embedding space, it delivers a massive 12 to 20 speedup over autoregressive approaches.
At scale, while autoregressive fan-out latency expands linearly to nearly 50 seconds under large context batches, Retrieve-for-Train-Diffusion stays between sub-second to a few seconds, delivering production-ready, expert-level search at a fraction of the computational cost.
### The anti-hacking anchor (ablation insights)
During our reward optimization process, we discovered something fundamental about training a fan-out language model for search. Without a diversity term, the model quickly collapses into generating degenerate, nonsensical strings (like _"line ending line ending"_) to mathematically exploit the vector coordinates of the database. Injecting a geometric diversity metric (the Vendi Score) acts as a vital counter-anchor, forcing the model into a stable region of the embedding space where it can only maximize its reward by acting like a true search expert.
## Conclusion
We demonstrated that RL can be highly effective when used as a one-time "objective transducer" rather than an online inference engine. By decoupling the heavy computation of reward-driven behavior exploration from the final deployed model, our framework successfully bypasses the steep inference latency and high computational overhead typical of online LLM deployment.
Distilling these complex, set-level behaviors into a lightweight diffusion prior allows production retrieval systems to optimize for higher-order properties like diversity and alignment effectively. Ultimately, Retrieve-for-Train establishes a highly scalable, data-efficient pipeline for set retrieval in specialized or multimodal domains where human-labeled, property-aligned training pairs are otherwise scarce or costly to obtain. See the [paper](https://arxiv.org/abs/2603.06397) for more details.
To do this, systems use a [query fan-out](https://blog.google/products-and-platforms/products/search/ai-mode-search/) technique that breaks a single broad prompt into several related sub-queries to cover potential user interests. However, teaching an LLM to perform database-aware [query decomposition](https://www.emergentmind.com/topics/query-decomposition) dynamically drains a massive thinking budget. By design, [zero-shot](https://www.promptingguide.ai/techniques/zeroshot) LLMs are general [autoregressive](https://aws.amazon.com/what-is/autoregressive-models/) text predictors; they aren’t optimized to navigate the specific, [geometric manifold](https://medium.com/@adnan.mazraeh1993/manifold-learning-and-geometry-based-approaches-a-comprehensive-explanation-7bc33d29cc04) of a target corpus. Consequently, they need extended test-time computation to return a collection of results that optimizes higher-order set-level properties (e.g., diversity, coverage, complementarity, coherence) while remaining grounded with respect to a fixed database.
In our [ICML 2026](https://icml.cc/virtual/2026/poster/66354) paper, “[Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion](https://arxiv.org/abs/2603.06397)”, we address this decomposition bottleneck via a reward-to-data compilation framework. Instead of forcing the model to expend a large thinking budget at inference, our Retrieve-for-Train framework uses offline reinforcement learning (RL) to discover reward-aligned fan-outs and compile them into supervision. By distilling these optimized exploration behaviors into a lightweight diffusion retriever, we enable highly efficient, single-pass query fan-out at inference time. This achieves mathematically formulated, set-level properties without the overhead of test-time thinking tokens.
## Why everyday AI isn't a search expert
When tasked with brainstorming a complex group of search terms, it’s tempting to simply deploy a standard, off-the-shelf LLM at inference time to handle the job. However, relying on generic models for database-aware query decomposition introduces two critical challenges:
1. _Paraphrastic collapse_**_:_** Without database-aware optimization, zero-shot LLMs frequently suffer from [paraphrastic collapse](https://arxiv.org/html/2605.04665v2). Rather than exploring complementary facets of a topic, they tend to generate redundant, near-synonymous queries. For example, given the broad prompt "Bohemian festival style”, a standard LLM without careful prompt engineering might lazily generate "bohemian festival fashion" and "bohemian festival clothes”. This semantic looping produces a homogeneous slate of results, entirely missing the distinct, helpful semantic directions a fashion expert would identify, such as fringe jackets, crochet dresses, or suede boots.
2. _Autoregressive latency bottlenecks_**_:_** Standard LLMs are fundamentally constrained by sequential, autoregressive generation. To successfully decompose a complex query into complementary facets, modern models typically require a substantial thinking budget, generating hundreds of intermediate [chain-of-thought](https://blog.bluedot.org/p/faithful-chain-of-thought?utm_source=google&utm_medium=pmax&utm_campaign=FoAI_&utm_term=&utm_content=&gad_source=1&gad_campaignid=22554833691&gbraid=0AAAAA_kXnYMdypNLRSDdkci909e4FZdzE&gclid=Cj0KCQjwnbrUBhDOARIsAKKhPpe_x5PmJtrYhjcdmCSdnrM6xrWWnbjdI_8RAn6WxkMZ55NPyB399VMaAv0WEALw_wcB) (CoT) reasoning tokens (i,e., the intermediate steps or internal processing units an AI model generates before answering a complex question) to plan their expansion before outputting the actual search terms. While this deliberate reasoning is acceptable for conversational AI, it introduces a severe structural bottleneck for set-valued search (e.g., retrieving a complementary slate of results, such as fringe jackets or crochet dresses mentioned above). When a system must brainstorm a large slate of sub-queries simultaneously, the combined overhead of continuous context processing and generating extended reasoning tokens scales poorly. Even with advanced serving optimizations, this token-by-token architecture creates a latency floor that is fundamentally at odds with the sub-second response times required by a production search bar.
## The Retrieve-for-Train framework
The Retrieve-for-Train treats the AI's training like an offline practice session rather than a test it has to take on the spot while a user is waiting. Instead of forcing the AI to slowly figure out the rules of a good search and drain a massive processing budget every single time someone types a query, Retrieve-for-Train runs an offline RL training program once.
This program uses a rigorous reward system to turn abstract goals like "ensure the results are diverse and actually in stock" into an exact step-by-step instruction manual. Once that manual is built, the AI can execute it instantly during a real search without delay.
The pipeline operates in three distinct steps:
* _Fan-out language model training:_ RL trains a fan-out language model to emit property-aligned sub-queries scored by a set-level property-check reward. This evaluates the entire group of results as a whole, rather than scoring each result in isolation.
* _Supervision synthesis:_ The frozen fan-out language model synthesizes (query → target-set) pairs entirely offline for supervised learning, requiring no human labels.
* _Diffusive retriever training:_ A compact, 53.9M-parameter diffusion model learns to map a query embedding directly to a complete set of target embeddings in one non-autoregressive pass, officially bypassing the need for text-based CoT reasoning tokens.
### Designing for the set: The power of composite rewards
The success of the Retrieve-for-Train framework hinges entirely on how we define "good" search behavior. Traditional supervised training evaluates [pointwise relevance via learning to rank](https://en.wikipedia.org/wiki/Learning_to_rank), scoring each retrieved item in isolation. However, a truly expert search slate is defined by non-decomposable, set-level properties. You can’t measure the diversity or complementarity of a single item; these properties only exist mathematically when evaluating the entire collection of retrieved results.
Rather than relying on ambiguous natural language instructions to enforce these fan-out properties, Retrieve-for-Train fine-tunes the 4B open-source language models ([Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) and [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)) via reinforcement learning using a strict mathematical composite reward. For our open-ended abstract retrieval tasks, this composite reward is a weighted balance of three competing pillars:
* _Groundedness:_ Penalizes distance to the database manifold, ensuring every generated sub-query corresponds to a real, retrievable item in the database.
* _Diversity:_ Measured using the [Vendi Score](https://arxiv.org/abs/2210.02410) over the entire set of sub-queries, forcing the model to explore broad semantic breadth.
* _Alignment:_ Anchors candidate sub-queries to the original broad prompt to prevent semantic drift.
### Mutual counter-anchors and soft-GRPO training
During training, we optimize the fan-out language model against these geometric realities using [group relative policy optimization](https://arxiv.org/abs/2505.22257) (GRPO) with soft [proximal policy optimization](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88%5C) (PPO).
This specific triad of rewards is critical because they act as mutual counter-anchors. If a model is optimized purely for groundedness, it will reward-hack the system by generating degenerate, nonsensical strings that happen to mathematically map to a specific database coordinate. If alignment is added to fix the nonsense, the policy simply cheats by collapsing into repetitive paraphrases of the user's prompt.
By injecting the Vendi Score as a counter-anchor, Retrieve-for-Train effectively closes off these shortcut solutions. To achieve a high-reward state, the policy is forced into a balanced region of the embedding space where it must discover valid, strictly grounded, yet semantically distinct variations of the original intent.
## Experiments
To evaluate the Retrieve-for-Train framework, we used a combination of frozen, dataset-specific multimodal embedding backbones and [open-source language models](https://deepmind.google/models/gemma/gemma-3/) optimized for query expansion. We evaluated this setup across two distinct set-valued retrieval regimes:
* _Open-ended abstract retrieval:_ A setting where no unique ground truth exists and quality is exclusively measured by set-level properties, including diversity, query alignment, and database groundedness.
* _Weakly supervised compositional retrieval:_ A setting where queries are paired with a weak reference set that serves as just one plausible realization of the query intent.
For the multimodal embedding backbones, we conducted experiments across two domains: A large-scale fashion dataset of user-curated outfits used for text-to-image experiments (evaluated using a [CLIP](https://openai.com/index/clip/)-based retriever), and a proprietary industrial dataset of expert-generated music playlists used for text-to-music evaluations (evaluated using [MuLan](https://research.google/pubs/mulan-a-joint-embedding-of-music-audio-and-natural-language/)).
For the language models, the query fan-out process was driven by 4B open-source models, specifically [Gemma3-4B](https://huggingface.co/google/gemma-3-4b-it) and [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B), which were tasked with generating exactly 10 sub-queries for every single main search prompt they processed. We implemented the RL training for these fan-out models via [Soft-GRPO](https://arxiv.org/abs/2511.06411), an approach that uses [group relative policy optimization](https://cameronrwolfe.substack.com/p/grpo) with [soft PPO regularization](https://medium.com/@kdk199604/ppo-efficient-stable-and-scalable-policy-optimization-15b5b9c74a88).
## Results
### Retrieval quality and accuracy
Across both retrieval tasks, Retrieve-for-Train outperformed traditional single-query search, zero-shot expansion, and even the heavily optimized [Best-of-N baseline](https://openai.com/index/measuring-goodharts-law/).
Qualitatively, zero-shot LLM baselines tended to generate near-synonymous paraphrases (e.g., "bohemian festival style" vs. "bohemian festival fashion"), causing redundant results. Retrieve-for-Train generated highly diverse, distinct sub-queries (e.g., branching into "boots" or "lace") that remained strictly grounded within the database manifold.
### Order-of-magnitude faster inference
Directly deploying our RL-tuned language model yielded exceptional search quality, but it inherited standard autoregressive latency constraints and demanded a high computational thinking budget.
By distilling that learned behavior into the 53.9M-parameter Retrieve-for-Train diffusion model, we successfully smashed the latency bottleneck. Because the diffusion model generates all target directions simultaneously in a single, non-autoregressive parallel pass in continuous embedding space, it delivers a massive 12 to 20 speedup over autoregressive approaches.
At scale, while autoregressive fan-out latency expands linearly to nearly 50 seconds under large context batches, Retrieve-for-Train-Diffusion stays between sub-second to a few seconds, delivering production-ready, expert-level search at a fraction of the computational cost.
### The anti-hacking anchor (ablation insights)
During our reward optimization process, we discovered something fundamental about training a fan-out language model for search. Without a diversity term, the model quickly collapses into generating degenerate, nonsensical strings (like _"line ending line ending"_) to mathematically exploit the vector coordinates of the database. Injecting a geometric diversity metric (the Vendi Score) acts as a vital counter-anchor, forcing the model into a stable region of the embedding space where it can only maximize its reward by acting like a true search expert.
## Conclusion
We demonstrated that RL can be highly effective when used as a one-time "objective transducer" rather than an online inference engine. By decoupling the heavy computation of reward-driven behavior exploration from the final deployed model, our framework successfully bypasses the steep inference latency and high computational overhead typical of online LLM deployment.
Distilling these complex, set-level behaviors into a lightweight diffusion prior allows production retrieval systems to optimize for higher-order properties like diversity and alignment effectively. Ultimately, Retrieve-for-Train establishes a highly scalable, data-efficient pipeline for set retrieval in specialized or multimodal domains where human-labeled, property-aligned training pairs are otherwise scarce or costly to obtain. See the [paper](https://arxiv.org/abs/2603.06397) for more details.