Public scorecard
How often the link is right
GlossMath claims one thing: that the symbol under your cursor is linked to the place this paper defined it. This page is where we keep that claim honest. Every number below is either measured or marked n/a. A metric we have not measured is never shown as a zero.
Last evaluated 2026-08-11 10:31:39 UTC
Papers evaluated
11
checked link by link
Links checked
424
individual symbol→definition links
Links correct
95.8% 406/424
the link points at the real definition
Abstention rate
77.7% 3104/3997
3104/3997 across the papers listed
Flag rate
0.0% 0/893
reader reports, over the links we showed
What we count as correct
A link is correct when the place it jumps to is where this paper actually introduces that symbol: the sentence that names it, the equation that sets it, or the line that gives its type. A link that lands on a later re-use of the symbol counts as wrong, even though the symbol is there. A link that lands in the right paragraph but on the wrong symbol counts as wrong.
Abstention is a feature, and we would rather it went up than down
When we cannot find where a symbol was defined, we list it under “not linked” with the reason and we do not link it. That is the abstention rate above. We could push it to zero tomorrow by guessing, and the correctness number would fall. The number we optimise is correctness among the links we do show.
What we link, and what we hold back
The model grades every definition it finds. It calls a passage high when the paper states the meaning outright at that spot, and medium when the meaning is clear from context but never actually said. We link only the first kind. 422 located passages across the papers below are held back for being the second kind.
That is a coverage decision and it costs real links. We made it because “medium” is defined, in the same prompt, as a meaning the paper never states outright — and this product’s entire claim is that a link points at where the paper defines the symbol. A link we would have to argue for is not one of those.
What a person said
Not run yet. The numbers above come from an independent model, which is a real check with a real blind spot: it reads the same passage our own pass read and answers the same question about it. Neither has tried to read the paper. A 400-question human evaluation is built and waiting on a labeller; nothing here will claim a human number until it has one.
Per paper
“Symbols” is every symbol we found. “Linked” is the subset we were confident enough to link. “Checked” is how many of those links an independent model from a different family has judged, span by span, without being shown our gloss or our confidence. It also still counts every link that check rejected and we then took off the page, so cleaning the showcase cannot raise this number.
What we do not know yet
- Whether accuracy holds outside the papers listed here. Everything measured so far is machine learning and adjacent maths, mostly recent arXiv preprints, and notation habits vary a lot by field.
- How we do on papers with no LaTeX source, where we fall back to a rougher conversion and the symbol boundaries are less reliable.
- Whether a definition we picked is the one a reader wanted. A symbol can be introduced loosely in the intro and pinned down properly three sections later. We prefer the pinned-down one, and we have not measured how often readers agree.
Found a link that is wrong? Every symbol card has a quiet wrong? button. Those reports come straight here as the flag rate.