Paper brief

SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval

SPLADE v2 improves a learned sparse retriever through max pooling, a document-only option, and distillation—but its strongest scores and efficiency claims belong to different configurations.

Why teach an index new words?

A first-stage retriever has to search a large collection before a slower reranker can inspect promising passages. Traditional inverted indexes make this cheap by matching terms, but a relevant passage can use different words from the query. Dense retrievers can bridge that vocabulary mismatch, at the cost of searching a different kind of index.

The original SPLADE paper, published at SIGIR 2021, proposed another path: learn which vocabulary terms a query or passage should carry, including useful terms absent from its literal text, while keeping the resulting vectors sparse enough for inverted-index retrieval. The author preprint reports strong first-stage passage ranking, but it leaves room to improve the balance between quality and scoring work. SPLADE v2 is a separate arXiv follow-up, not a new version of that SIGIR proceedings article.

A sparse score with room to expand

Both generations turn text into a vector over BERT’s WordPiece vocabulary. A masked-language-model head predicts a weight for each vocabulary term at each input position. Nonnegative, log-saturated weights are pooled across positions, and a query–passage pair is scored by the dot product of their sparse vectors. A passage can therefore acquire a weighted term it did not literally contain; zero-weight dimensions need not be stored in its posting lists.

The original SPLADE sums the contributions from input positions. V2 changes that pooling operation to a maximum: for a given vocabulary term, the strongest input-position contribution survives. It retains separate sparsity penalties for query and passage representations. These penalties discourage expensive, widely activated term dimensions while letting the model learn useful expansions. The point is not that max pooling turns a sparse retriever into a dense one. It changes how a sparse term’s weight is formed.

Three configurations, three different questions

The paper’s SPLADE-max asks whether the changed pooling rule improves a two-sided, expansion-aware retriever. Query encoding still happens at search time; passages can be encoded and indexed ahead of time.

SPLADE-doc asks a different question: what if the model expands and weights passages but leaves queries unexpanded and unweighted? At search time, query terms can look up learned passage weights without running a neural query encoder. This removes one online operation, but it also gives up learned query-side expansion. “Cheaper query path” is the design claim; the paper does not provide a measured end-to-end production latency comparison for this variant.

DistilSPLADE-max changes training as well as pooling. The authors use harder negative passages and scores from a cross-encoder teacher, then train the sparse retriever with a margin-based distillation loss. Its strongest result is therefore a result for the combined trained configuration, not an experiment that isolates max pooling by itself. V2’s new models also use DistilBERT-base initialization, whereas the original paper used BERT-base. Cross-paper differences should not be assigned to one switch.

Read the result row before the headline

The main comparison is first-stage retrieval over MS MARCO passages, not reranking with a BERT cross-encoder. Table 1 reports MRR@10 and Recall@1000 on the MS MARCO development set, plus NDCG@10 and Recall@1000 on TREC DL 2019. The latter has 43 assessed queries; the MS MARCO development set has 6,980. All four metrics reward higher values.

On MS MARCO dev MRR@10, the original SPLADE row is 0.322, SPLADE-max 0.340, SPLADE-doc 0.322, and DistilSPLADE-max 0.368. On TREC DL 2019 NDCG@10, the same rows are 0.665, 0.684, 0.667, and 0.729. The last comparison to the original is an absolute gain of 0.064 NDCG@10, or about 9.6% relative. It bundles training and model differences; it is not a measured 9.6% gain from max pooling alone.

The document-only row makes the trade-off visible. Its TREC Recall@1000 is 0.747, against 0.813 for original SPLADE and 0.851 for SPLADE-max. Removing query expansion may simplify serving, but this table does not say that it is free in retrieval quality.

Scatterplot of MS MARCO development MRR at 10 versus estimated scoring FLOPS. Blue DistilSPLADE-max circles generally reach higher MRR than black SPLADE-max and red original-SPLADE triangles at comparable plotted costs; an orange doc2query-T5 square and brown SparTerm lexical pentagon mark fixed baselines. One blue point near zero FLOPS has much lower MRR.
Formal, Lassance, Piwowarski, and Clinchant, SPLADE v2, Figure 1 (unchanged PNG from the versioned arXiv source; CC BY-NC-SA 4.0). The SPLADE series vary regularization strength; doc2query-T5 and SparTerm lexical are fixed comparison baselines. The vertical axis displays MRR@10 as a percentage, while the prose uses proportions; the horizontal axis estimates matching FLOPS, not elapsed query time. Open the full-resolution figure.

The chart shows why the authors vary sparsity rather than select solely for accuracy. At one reported distilled setting, MRR@10 is about 0.368 around 4 estimated FLOPS; a more regularized setting reaches about 0.35 around 0.3 FLOPS. These are estimated query–passage scoring operations, not an index-build bill, memory measurement, wall-clock latency, or a guarantee that two implementations would serve at the same speed.

Beyond one passage benchmark

The authors also evaluate a subset of BEIR datasets available to them. Table 2 reports average NDCG@10 of 0.500 across the listed datasets for DistilSPLADE-max, versus 0.460 for SPLADE-max and 0.455 for the cited ColBERT baseline. On the zero-shot subset, the corresponding averages are 0.506, 0.464, and 0.457. These are paper-reported comparisons, not a fresh independent rerun of every baseline.

Averages hide misses. On Natural Questions, distilled SPLADE scores 0.521, just below ColBERT’s 0.524. On Touché-2020 (v1), its 0.364 trails the table’s BM25 0.614. Five BEIR datasets were unavailable in this study. “Strong transfer on the evaluated subset” is a fair reading; “best on every dataset” is not.

Before using the result

  • Choose the configuration first. A max-pooled encoder, a document-only index, and a distilled model answer different quality–cost questions.
  • Measure the real system. FLOPS in this paper is a scoring proxy; include query encoding, postings traversal, memory, indexing, and end-to-end latency in your own test.
  • Keep the protocol attached to the number. MS MARCO dev MRR@10, TREC NDCG@10, and BEIR subset averages are not interchangeable.
  • Test out-of-domain failures. The BEIR average is informative, but the Touché result is a reason to examine individual query sets before deployment.