Daily Signal card · August 21, 2026

Every Citation Was Real

A bibliography can contain valid references while repeatedly excluding relevant branches of the available literature.

A researcher studies a small selected stack while a much broader archive remains visible around the table.
Citation coverageSelection trailResearch verificationPreprint limits

Every reference in a bibliography can exist while the shelf still gets narrower.

What changed

A new preprint tested whether language models keep choosing the same real papers after retrieval rankings and visible prestige cues are removed.

That setup used 120 knowledge-distillation papers published from 2015 through 2022. Eleven models from three vendors received random panels of thirty papers and could cite up to ten. The titles and abstracts were authentic. Author names were fabricated, years were reassigned, and venues and citation counts were hidden.

The result was concentration across all eleven models beyond the study’s indifferent-selection baseline. The top tenth of papers received 23.3 to 30.2 percent of model citations, compared with 15.6 percent under the matched null. One component explained 68 to 73 percent of the variation across model preference maps. Even the best cross-fitted mixture retained 55 percent of the measured excess.

The models came from different vendors, yet much of the pull pointed toward the same subset.

By contrast, eight domain experts completed 53 matched prompts. Their pooled choices showed no comparable shared pattern. That comparison needs restraint: one expert completed 41 prompts, and the human sample was much smaller than the model study.

Why it matters

Citation verification answers one question: does this source exist? It does not show whether the bibliography represents the available shelf.

That distinction matters as models move deeper into literature search and synthesis. A review can pass a link check while repeatedly excluding valid work that approaches the subject from another direction. Mixing model vendors may add variety, but this study suggests that vendor variety alone did not remove the shared preference map in this bounded task.

So a stronger workflow keeps a trace of the candidate set and the final selections. Then a reviewer can inspect which relevant sources were repeatedly skipped, where perspectives thinned out, and whether a subject-matter expert sees a missing branch of the literature. This is an operating check, not a proven cure for selection bias.

Watch next

But the paper is a v1 preprint, and the result comes from one 120-paper technical corpus under a hard citation budget. The benchmark measures choices among papers already shown to a model. It does not measure retrieval, publication decisions, or what authors do with a generated review. The code and data were linked by the authors but were not independently run here.

Replication across fields and retrieval systems would show whether the pattern travels. Until then, a clean bibliography should be treated as the beginning of the coverage check.

Primary source: Alemohammad et al., arXiv:2608.19230v1