Mercury 2.5 LLM hits 770 tokens per second
by Retro_Dev on 9/23/2026, 10:16:19 PM
https://artificialanalysis.ai/models/mercury-2-5
Comments
by: bearjaws
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.<p>I've used it on a few for fun projects and its decent but the speed is crazy to watch.
9/23/2026, 11:22:16 PM
by: walrus01
Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
9/23/2026, 11:03:53 PM
by: rvz
The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
9/23/2026, 10:24:59 PM