LLM & Text Generation2 min reading time

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

Apple Research Blog
Read full post
Researchers from UC Berkeley and affiliated institutions developed ARBITRAGE, a step-level speculative decoding method that dynamically routes generation between draft and target models to improve reasoning efficiency in large language models. ARBITRAGE reduces inference latency by up to 2x while maintaining accuracy on mathematical reasoning benchmarks, outperforming previous methods.

More in LLM & Text Generation

DeepSeek V4.1 Flash now available on AI Gateway

Covered by 2 sources

OpenAI’s GPT-Live-1 Arrives in the API at $0.05 Per Minute

Unite.AI