Reliable Parallel Decoding
in Masked Diffusion Language Models

Department of Computer Science, University of Virginia

Commit stable predictions together.
Keep upstream uncertainty in check.

Up to 6.1×throughput vs. default decoding
No retraininguses the current forward pass
Full canvasparallelism without fixed blocks

Real decoding trajectories

Loading example…

Loading recorded trajectories…

Download this replay as a GIF ↓

MaskedCommitted this stepCommitted earlier

Abstract

Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct.

Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.

The idea

Reliable predictions need ready context

Two complementary signals decide what to commit in parallel.

RPD overview: persistent layerwise predictions with limited confidence drop become candidates; an upstream entropy budget allows A and B to commit together while C waits.
Layerwise stability identifies candidates. Cumulative upstream uncertainty determines which can advance together. Figure 3 from the paper.
01 · Select candidates

Look beyond final confidence

Check whether the predicted token persists across the final layers and whether its probability drops. Very high-confidence predictions bypass the stability test.

02 · Commit together

Account for unresolved context

Scan candidates from left to right and track the entropy of earlier positions that will remain masked. Commit candidates within the budget, wherever they lie on the canvas.

RPD-block uses the same stability-based selection within fixed 32-token blocks, without the cumulative entropy budget. We report it separately from full-canvas RPD.

Speed with answer quality in view

LLaDA-8B-Instruct · MBPP
Methodpass@1 (%) ↑NFE ↓Tokens/s ↑Throughput vs. default
Default39.60256.005.931.00×
Fast-dLLM40.2044.7333.135.59×
EB-Sampler39.2053.2626.964.55×
LoPA39.4023.9421.583.64×
DAPD37.8047.1628.624.83×
RPD-block38.6040.6133.915.72×
RPD40.6038.1635.926.06×

On LLaDA MBPP, full RPD reaches 6.06× the throughput of default decoding, with pass@1 increasing from 39.60% to 40.60%.

Citation

@misc{he2026reliableparallel,
  title={Reliable Parallel Decoding in Masked Diffusion Language Models},
  author={Zhenghao He and Bohan Liu and Guangzhi Xiong and Aidong Zhang},
  year={2026},
  eprint={2609.36452},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2609.36452}
}