GlossMath

Public scorecard

How often the link is right

GlossMath claims one thing: that the symbol under your cursor is linked to the place this paper defined it. This page is where we keep that claim honest. Every number below is either measured or marked n/a. A metric we have not measured is never shown as a zero.

Last evaluated 2026-08-11 10:31:39 UTC

Papers evaluated

11

checked link by link

Links checked

424

individual symbol→definition links

Links correct

95.8% 406/424

the link points at the real definition

Abstention rate

77.7% 3104/3997

3104/3997 across the papers listed

Flag rate

0.0% 0/893

reader reports, over the links we showed

What we count as correct

A link is correct when the place it jumps to is where this paper actually introduces that symbol: the sentence that names it, the equation that sets it, or the line that gives its type. A link that lands on a later re-use of the symbol counts as wrong, even though the symbol is there. A link that lands in the right paragraph but on the wrong symbol counts as wrong.

Abstention is a feature, and we would rather it went up than down

When we cannot find where a symbol was defined, we list it under “not linked” with the reason and we do not link it. That is the abstention rate above. We could push it to zero tomorrow by guessing, and the correctness number would fall. The number we optimise is correctness among the links we do show.

What we link, and what we hold back

The model grades every definition it finds. It calls a passage high when the paper states the meaning outright at that spot, and medium when the meaning is clear from context but never actually said. We link only the first kind. 422 located passages across the papers below are held back for being the second kind.

That is a coverage decision and it costs real links. We made it because “medium” is defined, in the same prompt, as a meaning the paper never states outright — and this product’s entire claim is that a link points at where the paper defines the symbol. A link we would have to argue for is not one of those.

What a person said

Not run yet. The numbers above come from an independent model, which is a real check with a real blind spot: it reads the same passage our own pass read and answers the same question about it. Neither has tried to read the paper. A 400-question human evaluation is built and waiting on a labeller; nothing here will claim a human number until it has one.

Per paper

Paper Symbols Linked Checked Correct Flags
Attention Is All You Need arXiv:1706.03762 54 18 18 100.0% 18/18 0
Denoising Diffusion Probabilistic Models arXiv:2006.11239 56 10 10 100.0% 10/10 0
Auto-Encoding Variational Bayes arXiv:1312.6114 90 33 34 97.1% 33/34 0
Neural Ordinary Differential Equations arXiv:1806.07366 69 10 10 100.0% 10/10 0
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding arXiv:1810.04805 25 7 7 100.0% 7/7 0
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale arXiv:2010.11929 58 13 n/a n/a 0
Deep Residual Learning for Image Recognition arXiv:1512.03385 9 4 n/a n/a 0
Generative Adversarial Networks arXiv:1406.2661 59 6 n/a n/a 0
Language Models are Few-Shot Learners arXiv:2005.14165 99 8 n/a n/a 0
Distilling the Knowledge in a Neural Network arXiv:1503.02531 39 15 n/a n/a 0
Neural Machine Translation by Jointly Learning to Align and Translate arXiv:1409.0473 103 54 n/a n/a 0
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift arXiv:1502.03167 65 16 n/a n/a 0
The entropy formula for the Ricci flow and its geometric applications arXiv:math/0211159 210 48 n/a n/a 0
Calculation of prompt diphoton production cross sections at Tevatron and LHC energies arXiv:0704.0001 282 58 n/a n/a 0
Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning arXiv:1811.12808 106 42 n/a n/a 0
Learning Transferable Visual Models From Natural Language Supervision arXiv:2103.00020 31 2 n/a n/a 0
Layer Normalization arXiv:1607.06450 95 27 n/a n/a 0
Efficient Estimation of Word Representations in Vector Space arXiv:1301.3781 33 8 n/a n/a 0
Deep contextualized word representations arXiv:1802.05365 44 16 n/a n/a 0
Densely Connected Convolutional Networks arXiv:1608.06993 16 9 n/a n/a 0
Playing Atari with Deep Reinforcement Learning arXiv:1312.5602 40 11 n/a n/a 0
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models arXiv:2201.11903 16 1 n/a n/a 0
Neural Discrete Representation Learning arXiv:1711.00937 33 11 n/a n/a 0
An Algorithm for Optimal Partitioning of Data on an Interval arXiv:math/0309285 32 10 n/a n/a 0
Universal Model Routing for Efficient LLM Inference arXiv:2502.08773 190 66 n/a n/a 0
U-Net: Convolutional Networks for Biomedical Image Segmentation arXiv:1505.04597 22 6 n/a n/a 0
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks arXiv:1908.10084 19 11 n/a n/a 0
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks arXiv:1703.03400 36 14 n/a n/a 0
Differential Transformer arXiv:2410.05258 74 22 n/a n/a 0
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits arXiv:2402.17764 13 0 n/a n/a 0
Improving neural networks by preventing co-adaptation of feature detectors arXiv:1207.0580 33 4 n/a n/a 0
Very Deep Convolutional Networks for Large-Scale Image Recognition arXiv:1409.1556 7 4 n/a n/a 0
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks arXiv:1511.06434 3 1 n/a n/a 0
A sparsity result for the Dynamical Mordell-Lang Conjecture in positive characteristic arXiv:2012.13711 73 22 22 100.0% 22/22 0
Projected Stochastic Gradient Langevin Algorithms for Constrained Sampling and Non-Convex Learning arXiv:2012.12137 198 43 49 87.8% 43/49 0
Effective approximation of heat flow evolution of the Riemann ξ function, and a new upper bound for the de Bruijn-Newman constant arXiv:1904.12438 285 47 49 95.9% 47/49 0
Supermembranes and domain walls in 𝒩=1, D=4 SYM arXiv:1905.02743 269 47 48 97.9% 47/48 0
Dirichlet forms and polymer models based on stable processes arXiv:1905.00181 295 90 91 98.9% 90/91 0
Localizing virtual cycles for Donaldson-Thomas invariants of Calabi-Yau 4-folds arXiv:2012.13167 402 79 86 91.9% 79/86 0
Extremal eigenvalues of critical Erdős-Rényi graphs arXiv:1905.03243 414 n/a n/a n/a 0
Total across 40 papers 3997 893 424 95.8% 406/424 0

“Symbols” is every symbol we found. “Linked” is the subset we were confident enough to link. “Checked” is how many of those links an independent model from a different family has judged, span by span, without being shown our gloss or our confidence. It also still counts every link that check rejected and we then took off the page, so cleaning the showcase cannot raise this number.

What we do not know yet

Found a link that is wrong? Every symbol card has a quiet wrong? button. Those reports come straight here as the flag rate.