Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
摘要
本文是 Neon 与 Castform 的联合技术博客,介绍 Castform 如何利用 Neon 的 Lakebase Search 和 Postgres 基础设施,对开源模型进行强化学习后训练,使其在代理式检索任务上以约 100 倍更低的成本超越 GPT-5.6 Sol 等前沿模型。文章对比了 2022 年嵌入检索到 2025 年代理式检索的演进,指出多跳搜索循环调用前沿模型导致高成本高延迟,而 RL 后训练可缩小开源模型差距。Castform 将企业现有语料自动转化为训练任务和奖励函数,管理整个 RL 循环,并提供训练可观测性。Neon 的动态计算扩缩容和分支功能支撑训练时的高突发负载及有状态代理的隔离环境。文末附有代码示例和训练运行链接。
荐读理由
正文给出了具体技术路径:用Neon的Lakebase Search作为工具环境,将企业语料自动生成训练任务和奖励函数,对开源模型做RL后训练,并展示了成本对比(gpt-5.6-sol多轮检索>10秒、约$0.03,开源模型100倍便宜)。这为独立开发者提供了一条可参考的模型定制思路,但文章本质是两家公司的产品宣传,缺少独立基准或第三方验证,实际效果需自行测试。
原文

“Most teams' best training data is just sitting in their databases. The problem is that turning raw data into something usable is hard, and letting agents read, search, and mutate data cheaply at scale requires advanced infra. Pointing Castform at Neon skips both.”
Ying Hang Seah, cofounder, Castform
A "good agent" needs to be strong in 2 areas:
Context: can we provide the tools to find the right data?
Model: can the model decide what to search for?
Neon (Lakebase Postgres) and their new Search extensions solve the first; Castform solves the second.
Evolution of agentic search
In ~2022, the industry was going all in on embedding search. Every database provider added one, and pgvector was Neon's most downloaded extension. To provide context to LLMs, engineers handcrafted RAG pipelines, which in essence, is some form of embedding similarity search.
In ~2025, agents started to gain more traction. Developers started creating multi-hop search workflows, decomposing big problems into smaller ones. Retrieval has shifted from the one-shot search systems to agentic retrieval. Instead of issuing a single query, models plan and search multiple times in a loop. Every loop iteration meant another call to the frontier model, increasing the overall cost and latency per user request.

Concretely, a typical multi-turn search request with gpt-5.6-sol takes >10s and costs ~$0.03 end-to-end, making it prohibitively slow and expensive.
Meanwhile, small open-weights models are 100x cheaper. But, out of the box, their capabilities lag behind closed api models. RL post-training helps bridge this gap. On specific tasks like search, post-trained open-source models can match & beat frontier models while costing orders of magnitude less per request.
That is why we built Castform: to enable developers to RL post-train models without having to deal with machine learning & gpu internals. The goal's to make post-training as approachable as prompt engineering.
How does Castform use Neon?
Castform's pipeline runs against Neon via Lakebase Search:
| Stage | Neon + Lakebase Search |
|---|---|
| Corpus storage | Raw documents live in Postgres on Neon |
| Synthetic data generation | Castform training pipeline uses lakebase_text and lakebase_vector to write training tasks |
| RL Training | Every rollout's search tool call uses Lakebase Search on Neon |
| Production Inference | The final model uses the same search tool call during inference |
Your best training data already exists
To perform RL post-training effectively, you need a task (e.g. answer a user's question), the environment for the agent to run in (e.g. a search tool for your corpus) and a reward function (e.g. is the answer correct?).
With all 3 pieces in place, the RL post-training is a loop of trial and error: the model attempts the task given the tools, the reward function scores the attempt, and the feedback signal guides the model on how to hill-climb its way to optimal performance.
Yet, most companies do not have a clean dataset of tasks and reward functions ready for post-training.
Enterprises do have a large set of proprietary data:
internal documentation
product records
support articles
customer interactions
wikis
operational databases
This data contains the knowledge an agent needs, but turning it into an effective training dataset normally requires substantial data engineering and manual labeling.
That leads many teams to dismiss post-training for one of two reasons:
"We don't have the training data."
"Fine-tuning is too difficult and requires infrastructure we don't have."
Castform addresses both. It turns an existing corpus into training tasks, then manages the RL loop needed to teach an open-source model how to use that data effectively.
Using Castform
With Castform, you can turn your company knowledge base into a model:
Document (from your data): Trains booked through Navan will be paid by GitLab travel card. Train rides must be standard cabin class with 14 day booking lead time
Ground truth (inferred from your data): Train rides must be standard cabin class with a 14 day booking lead time.
Question (synthetically generated): When booking a rail trip in Navan, what are the rules for how early I need to reserve it and which seating level I'm expected to choose?
With the generated question-answer dataset, Castform lets you scaffold the training run by specifying the tools the agent has access to and a reward function.
The reward function specifies what you want your model to get good at. In our case, we want it to retrieve the correct chunks, cite the right sources along with providing the right final answer.
def run_tool(tool, tool_args):
"""Single tool: hybrid search over Lakebase."""
if tool == "search":
query = tool_args["query"]
bm25 = neon.lakebase_text(query, k)
vector = neon.lakebase_vector(query, k)
return rrf_merge(bm25, vector, k)
def reward(trace, ground_truth):
""""""
answer = parse_trace(trace)
retrieval = ... # did it retrieve the right source
citation = ... # did it cite the right chunk
correctness = ... # did it land on the right answer
return retrieval + citation + correctness
See a comprehensive code example here.
Observability: Watch the model learn
Castform gives you full observability into your RL run. You can monitor your reward climb with each step, but more importantly you can drop into individual tasks/prompts to watch how the model performs qualitatively, allowing you to debug problems such as broken tools or reward hacking.
For more details on how to monitor your training runs, you can check out the Castform blog here. You can also check out our example training run here.

Why Neon 'just works'
During training, the agent repeatedly calls Lakebase Search until it has enough context to answer. Across thousands of parallel rollouts, each potentially making dozens of calls, this creates a highly bursty workload.

Neon's dynamic compute scaling absorbs these peaks without requiring Castform to provision for maximum capacity around the clock. Training runs get low-latency search when demand spikes, while compute scales down during idle periods.
This infrastructure becomes even more valuable as agents move beyond search and begin modifying data. Training stateful agents requires isolated environments that can be created and reset cheaply, preventing one rollout's actions from affecting another or touching production.
Neon branching can give each rollout an isolated database state, while time-travel queries make it possible to reconstruct and inspect the state an agent encountered. Combined with autoscaling and scale-to-zero, this creates a path toward training thousands of stateful agent rollouts without maintaining thousands of continuously running environments.
Castform makes it easy for any developer to post-train open-source models to be cheaper, faster, better than the frontier. Post-train your first model today at castform.com.
这条对你有帮助吗?