LLM & Text Generation2 min reading time

Arbitrage: Efficient Reasoning via Advantage-Aware Speculation

Apple Research Blog
Read full post
Researchers from UC Berkeley and affiliated institutions developed ARBITRAGE, a step-level speculative decoding method that dynamically routes generation between draft and target models to improve reasoning efficiency in large language models. ARBITRAGE reduces inference latency by up to 2x while maintaining accuracy on mathematical reasoning benchmarks, outperforming previous methods.

More in LLM & Text Generation

Peter Thiel-Backed AI Startup Cognition Raises Funds at $48 Billion Valuation

Covered by 2 sources

Build more natural voice experiences with GPT‑Live‑1 in the API

Covered by 2 sources