How Can Reinforcement Learning Achieve Expert-Level [Chip] Placement?
摘要
标题《How Can Reinforcement Learning Achieve Expert-level Placement?》来自 arXiv 2604.25191,作者 Ruo-Tong Chen 等。摘要指出 RL 方法训练重心在线长优化,常难达专家水平;核心原因是奖励设计缺失,论文绕过复杂过程,直接从最终专家布局推断 step-by-step 专家轨迹,使用这些轨迹作为演示或偏好训练模型,捕捉专家结果的隐式奖励。实验表明该框架能高效学习,甚至单次设计即可良好泛化至未见案例。
荐读理由
EIM框架从单个专家最终布局直接推导出步骤轨迹(演示或偏好),训练隐含奖励模型;实验证明对已见芯片能让RL策略达到专家PPA水平,对未见芯片也能有意义进步
原文
Computer Science > Hardware Architecture
[Submitted on 28 Apr 2026 (v1), last revised 1 Jun 2026 (this version, v2)]
Title:How Can Reinforcement Learning Achieve Expert-level Placement?
Authors:Ruo-Tong Chen, Ke Xue, Chengrui Gao, Yunqi Shi, Tian Xu, Peng Xie, Siyuan Xu, Mingxuan Yuan, Chao Qian, Zhi-Hua Zhou
View a PDF of the paper titled How Can Reinforcement Learning Achieve Expert-level Placement?, by Ruo-Tong Chen and 9 other authors
Abstract:Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore often fail to achieve expert-quality layouts. We identify the reward design as the primary cause for the performance gap with experts, and instead of formalizing intricate processes, we circumvent this by directly learning from expert layouts to derive a reward model. Our approach starts from the final expert layouts to infer step-by-step expert trajectories. Using these trajectories as demonstrations or preferences, we train a model that captures the latent implicit rewards in expert results. Experiments show that our framework can efficiently learn from even a single design and generalize well to unseen cases.
https://doi.org/10.48550/arXiv.2604.25191
arXiv-issued DOI via DataCite
| Comments: | |
|---|---|
| Subjects: | Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.25191 [cs.AR] |
| (or arXiv:2604.25191v2 [cs.AR] for this version) | |
Submission history
From: Chao Qian [view email] [v1] Tue, 28 Apr 2026 03:55:03 UTC (358 KB) [v2] Mon, 1 Jun 2026 14:43:02 UTC (358 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled How Can Reinforcement Learning Achieve Expert-level Placement?, by Ruo-Tong Chen and 9 other authors
Current browse context:
cs.AR
Change to browse by:
References & Citations
Loading...
BibTeX formatted citation
Data provided by:
Bookmark
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Connected Papers (What is Connected Papers?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
alphaXiv (What is alphaXiv?)
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Hugging Face (What is Huggingface?)
ScienceCast (What is ScienceCast?)
Demos
Replicate (What is Replicate?)
Hugging Face Spaces (What is Spaces?)
TXYZ.AI (What is TXYZ.AI?)
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
这条对你有帮助吗?