Hacker News Viewer

Mercury 2.5 LLM hits 770 tokens per second

by Retro_Dev on 9/23/2026, 10:16:19 PM

https://artificialanalysis.ai/models/mercury-2-5

Comments

by: bearjaws

If you care about speed Cerebras gpt-oss-120b is 1400tk&#x2F;s and &quot;just as smart&quot; in ranking.<p>I&#x27;ve used it on a few for fun projects and its decent but the speed is crazy to watch.

9/23/2026, 11:22:16 PM


by: walrus01

Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don&#x27;t see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.

9/23/2026, 11:03:53 PM


by: rvz

The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.

9/23/2026, 10:24:59 PM