SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
Ensembles of generative large language models (LLMs) can integrate the strengths of different LLMs to compensate for the limitations of individual models.
However, recent work has focused on training an additional fusion model to combine complete responses from multiple LLMs, failing to tap into their collaborative potential to generate higher-quality responses.
Moreover, as the additional fusion model is trained on a specialized dataset, these methods struggle with generalizing to open-domain queries from online users.
In this paper, we propose SpecFuse, a novel ensemble framework that outputs the fused result by iteratively producing the next segment through collaboration among LLMs.
This is achieved through cyclic execution of its inference and verification components.
In each round, the inference component invokes each base LLM to generate candidate segments in parallel, and the verify component calls these LLMs again to predict the ranking of the segments.
The top-ranked segment is then broadcast to all LLMs, encouraging them to generate higher-quality segments in the next round.
This approach also allows the base LLMs to be plug-and-play, without any training or adaptation, avoiding generalization limitations.
Furthermore, to conserve computational resources, we propose a model exit mechanism that dynamically excludes models exhibiting poor performance in previous rounds during each query response.
In this way, it effectively reduces the number of model calls while maintaining overall performance.
We conduct extensive experiments using ensembles of five LLMs with different architectures across six benchmarks, covering instruction-response, reasoning, commonsense, and instruction-following tasks. The experimental results demonstrate that SpecFuse consistently enhances performance across all benchmarks, with RougeL scores improving by +3.13.1+3.1+ 3.1 on the Chinese and +3.03.0+3.0+ 3.0 on the English human-computer interaction benchmarks. Furthermore, the model exit mechanism reduces the average models invoked per round from 5555 to 2.42.42.42.4, with only a slight reduction in performance.
https://arxiv.org/html/2412.07380v1
Reinforcement Learning on Pre-Training Data
The growing disparity between the exponential scaling of computational resources and the finite growth of high-quality text data now constrains conventional scaling approaches for large language...
https://arxiv.org/abs/2509.19249


Seonglae Cho