ARC-Decode: Accelerated Decoding with Risk-Bounded Acceptance

ICML 2026
Ying Li1, Zhaode Wang2, Zhiwen Chen2, Chengfei Lv2, Huan Wang1
1Westlake University  ยท  2Alibaba Group
Code Paper Apache 2.0

Abstract

As larger language models deliver stronger capabilities, their autoregressive inference becomes increasingly expensive. Speculative decoding accelerates generation by letting a fast draft propose tokens that the target model verifies in parallel. Yet under sampling (T > 0), observed speedups consistently lag behind those under greedy decoding, as the classical lossless verification rule tends to over-reject low-risk drafts, leading to lower acceptance rates and limited acceleration. To address this gap, we propose ARC-Decode (Acceptance with Risk Control), a training-free method that augments speculative decoding without extra forward passes. ARC-Decode enables relaxed acceptance by identifying drafts whose acceptance is expected to induce limited local distributional deviation under a calibrated risk criterion based on Jensen–Shannon divergence. It combines confidence-based pre-verification filtering with a risk-bounded acceptance criterion derived from an analytic upper bound on the potential distributional deviation. Integrated into the state-of-the-art EAGLE-3 pipeline, ARC-Decode increases accept length per cycle and reduces verification compute, achieving up to 1.6× end-to-end speedup over EAGLE-3 under sampling with comparable generation quality across the evaluated benchmarks.

Problem

Problem
Why does speculative decoding slow down under sampling?

Key Problem. EAGLE-3 exhibits significant degradation in decoding efficiency under sampling, with performance further deteriorating as temperature increases, leading to reduced acceptance length in speculative decoding.

Motivation

Are the rejected tokens under sampling actually harmful?

Motivation 1
Many rejected drafts yield highly consistent continuations under multiple surrogate semantic metrics.

Observation 1. Many rejected drafts yield highly consistent continuations under multiple surrogate semantic metrics, suggesting that token-level rejection does not necessarily imply a downstream change.

Motivation 2
Counterfactual rollouts reveal substantial recoverable acceptance space.

Observation 2. Counterfactual rollouts reveal substantial recoverable acceptance space: many full-rejection events contain rejected draft candidates that still preserve the correct final answer.

Key insight. Under sampling, posterior rejection sampling can reject draft candidates that would still preserve the reasoning trajectory and final answer, revealing additional acceptance space that can be exploited under bounded risk.

Method

ARC-Decode method overview
Overview of ARC-Decode: entropy-guided pruning, local shift estimation, and risk-bounded acceptance.
  1. Entropy-guided pruning reduces verification cost. Draft nodes are ranked by path probability, target entropy, and depth to form a compact prefix-closed subtree.
  2. Local shift estimation finds low-risk relaxation opportunities. For rejected tokens, ARC-Decode combines target logit margin and whitened embedding distance to estimate local distribution shift.
  3. Risk-bounded acceptance extends accepted sequences. A calibrated Local Tolerance Score selectively relaxes low-shift rejections while retaining standard posterior acceptance as the floor.
Summary. Training-free, plug-and-play, and free of extra target forward passes.

Results

Full-data results on NVIDIA A100 GPUs, temperature=1.0, total_tokens=60, depth=7. Each benchmark reports accept length, throughput (tok/s), and speedup (method / Base). Best in bold.

Llama 8B
methodMTB accMTB tpMTB spHE accHE tpHE spGS accGS tpGS spAlp accAlp tpAlp sp
Base1.00033.0931.00×1.00031.5501.00×1.00031.9641.00×1.00032.0781.00×
Eagle-3 base3.12166.0992.00×4.52990.7122.88×5.184108.8283.40×5.054105.5463.29×
ARC (Ours)6.324131.0623.96×5.11593.5812.97×6.648125.7133.93×6.223125.6673.92×
Qwen 8B
methodMTB accMTB tpMTB spHE accHE tpHE spGS accGS tpGS spAlp accAlp tpAlp sp
Base1.00027.5971.00×1.00027.2061.00×1.00027.6141.00×1.00027.9371.00×
Eagle-3 base4.04678.6102.85×4.17071.7412.64×5.36792.6323.35×4.44481.0252.90×
ARC (Ours)4.97895.2463.45×4.23276.0582.80×5.78296.8503.51×5.00892.1963.30×
Vicuna 13B
methodMTB accMTB tpMTB spHE accHE tpHE spGS accGS tpGS spAlp accAlp tpAlp sp
Base1.00029.1551.00×1.00028.3631.00×1.00030.1091.00×1.00030.1271.00×
Eagle-3 base5.354106.4583.65×5.656102.3633.61×5.731117.0323.89×5.22899.4243.30×
ARC (Ours)6.880134.5964.62×6.905131.0574.62×6.796126.5424.20×6.582131.5064.37×

Quality: HumanEval = pass@10, GSM8K = exact match. Best in bold.

llama
methodhumanevalgsm8k
Base0.60370.7125
Eagle-3 base0.55490.6375
ARC (Ours)0.74390.7074
qwen
methodhumanevalgsm8k
Base0.85370.7875
Eagle-3 base0.85370.8125
ARC (Ours)0.85370.8211
vicuna
methodhumanevalgsm8k
Base0.12200.1625
Eagle-3 base0.11590.1625
ARC (Ours)0.27440.1971
A6000 results
Additional efficiency results on a single NVIDIA A6000 GPU (total_tokens=32, depth=6).

BibTeX

@inproceedings{li2026arcdecode,
  title     = {ARC-Decode: Accelerated Decoding with Risk-Bounded Acceptance},
  author    = {Li, Ying and Wang, Zhaode and Chen, Zhiwen and Lv, Chengfei and Wang, Huan},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning (ICML)},
  year      = {2026}
}