Please don't discontinue Gemini 2.5 Flash
摘要
帖主感谢 Gemini 团队提供访问权限,并表示其内部工作流严重依赖 Gemini 2.5 Flash;对比显示 Gemini 3 Flash 与 3.1 Flash Lite 性能不足且存在问题,2.5 Flash 是最佳低延迟选择;强调其仅在澳大利亚部署,完成时间 300-400ms,适合语音代理;3.5 Flash 延迟 600-800ms、成本约 3 倍且未部署澳大利亚,无法满足需求;呼吁 Google 团队延长保留 2.5 Flash 以保护用户流量与竞争力。
荐读理由
我们依赖的内部基准显示 Gemini 2.5 Flash 在延迟+性能上无竞争,3.1 flash lite 甚至有 thoughts leaking 问题,成本从 2.5 Flash 到 3.5 Flash 会涨三倍;建议找 open weight 模型替代
原文
Please don't discontinue Gemini 2.5 Flash
Firstly, I want to give my thanks to the Gemini team for providing access to such great models. It has been so helpful.
We have some very specific workflows that rely on Gemini 2.5 Flash. Our internal benchmarks show that Gemini 3 flash does not perform as well (even after attempting to tweak prompting following the new prompting guidelines and other changes).
I am sure there are others with the same experience as well, and there is no easy switch.
It would be extremely appreciated if 2.5 flash was not discontinued.
Yes, even our own benchmarks show that the closest model in latency + performance, which is 3.1 flash lite, doesn’t even come close to 2.5 Flash. Seeing issues with thoughts leaking out. I should say, 2.5 flash has been the best model we’ve seen for all round usage and works pretty well with most of tasks.
I am pretty sure, a huge chunk of their traffic and usage comes from this model. Would really appreciate it if the Google team can retain this model for much longer time.
Preach. Retiring 2.5 Flash will be such a massive loss. It’s the only low latency model that is deployed in Australia. I’m able to get 300-400ms completions from 2.5 Flash, making it suitable for voice agents.
3.5 Flash offers 600-700ms completions and doesn’t even get an Australian deployment, so the actual latency is closer to 700-800ms, completely breaking the voice agent use case.
There is honestly no other model deployed in this part of the world that comes close to the quality that 2.5 Flash offers for low latency applications.
I am more concerned about the cost step up from Gemini 2.5 Flash to 3.5 Flash, with the latter being roughly 3x more expensive. I thought the intention of the Flash models was to be relatively low-latency and more affordable compared to Pro, but the newer Flash models aren’t being priced as such. Then again, the era of cheap and plentiful AI might be coming to an end…
Seconded, the alternative for us is not upgrading to flash-3 but rather finding an appropriate open weight model
+1 I have some critical workflows that no other model is good at for the same price/intelligence! It would be a huge hit to have this model discontinued - we’d likely switch to an open source model if this happened but the latency of 2.5 flash is something hard to beat.
这条对你有帮助吗?