Qwen-AgentWorld: Language World Models for General Agents
摘要
论文围绕 “world model(世界模型)” 用于 Agent 推理与规划展开,探索用语言模型构建环境动态预测能力。作者提出两个模型版本:Qwen-AgentWorld-35B-A3B 与 Qwen-AgentWorld-397B-A17B,用于在 7 个领域中进行 Agent 环境模拟,并支持长链式推理。训练使用超过 1000 万条真实环境交互轨迹,并采用三阶段流程:CPT 注入通用世界建模能力(基于状态转移与增强语料),SFT 激活下一状态预测推理能力,RL 通过混合规则与评分标准提升模拟精度。同时提出评测基准 AgentWorldBench,基于 5 个前沿模型在 9 个现有 benchmark 中的真实交互构建。实验结果表明,该方法整体优于现有前沿模型。除此之外,论文还提出两种世界模型增强 Agent 的路径:作为可解耦环境模拟器,用于大规模可控环境生成并提升 Agent RL 效果;以及作为统一 Agent 基础模型,在预训练阶段作为 warm-up,提升 7 个 Agent 基准任务表现。
荐读理由
论文给出的三阶段训练流程(CPT注入世界建模、SFT激活next-state推理、RL用混合奖励精炼)和两个范式(解耦模拟器支持大规模Agent RL、统一基础模型warm-up提升7基准),基于10M+真实轨迹覆盖7领域,这为你判断Agent项目是否采用语言世界建模架构提供了可迁移的新方法与依据。
原文
Computer Science > Computation and Language
[Submitted on 23 Jun 2026]
Title:Qwen-AgentWorld: Language World Models for General Agents
Authors:Yuxin Zuo, Zikai Xiao, Li Sheng, Fei Huang, Jianhong Tu, Yuxuan Liu, Tianyi Tang, Xiaomeng Hu, Yang Su, Qingfeng Lan, Yantao Liu, Qin Zhu, Yinger Zhang, Bowen Yu, Haiquan Zhao, Haiyang Xu, Jianxin Yang, Jiayang Cheng, Junyang Wang, Lianghao Deng, Mingfeng Xue, Tianyi Bai, Yang Fan, Yubo Ma, Yucheng Li, Zeyu Cui, Zhihai Wang, Zhihui Xie, Zhuorui Ye, An Yang, Dayiheng Liu, Jingren Zhou, Ning Ding
View a PDF of the paper titled Qwen-AgentWorld: Language World Models for General Agents, by Yuxin Zuo and 32 other authors
Abstract:A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of general agents. (i) We first focus on building foundation models for agentic environment simulation. We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning. Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the state transition dynamics and augmented professional corpora, SFT activates next-state-prediction reasoning, and RL sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. To evaluate language world models, we present AgentWorldBench, a comprehensive benchmark constructed from real-world interactions of 5 frontier models on 9 established benchmarks. Empirical results demonstrate that Qwen-AgentWorld significantly outperforms existing frontier models. (ii) Beyond foundation models, we further investigate two complementary paradigms through which world modeling enhances general agents. First, as a decoupled environment simulator, Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone. Second, as a unified agent foundation model, world-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks. Code: this https URL
https://doi.org/10.48550/arXiv.2606.24597
arXiv-issued DOI via DataCite (pending registration)
| Subjects: | Computation and Language (cs.CL) |
|---|---|
| Cite as: | arXiv:2606.24597 [cs.CL] |
| (or arXiv:2606.24597v1 [cs.CL] for this version) | |
Submission history
From: Fei Huang [view email] [v1] Tue, 23 Jun 2026 13:53:55 UTC (3,883 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Qwen-AgentWorld: Language World Models for General Agents, by Yuxin Zuo and 32 other authors
Current browse context:
cs.CL
Change to browse by:
References & Citations
Loading...
BibTeX formatted citation
Data provided by:
Bookmark
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Connected Papers (What is Connected Papers?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
alphaXiv (What is alphaXiv?)
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Hugging Face (What is Huggingface?)
ScienceCast (What is ScienceCast?)
Demos
Replicate (What is Replicate?)
Hugging Face Spaces (What is Spaces?)
TXYZ.AI (What is TXYZ.AI?)
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
这条对你有帮助吗?