Each evaluation runs three times and reports the median. The published score is the exact median run, not an average of reworked numbers.
The prompt is hash-locked, so a change to the prompt is a versioned change rather than a silent one. Long documents are chunked on a fixed threshold, and data-heavy passages such as numeric tables are excluded from language scoring.
Truverai scores are reproducible within a published lock: engine v4.7, prompt hash 1a6c25c7e25ddbf8, model google/gemini-3-flash-preview at temperature 0, and a median of three runs. That is reproducibility within a fixed configuration, not a claim that a different language model would return the same number.
This page needs JavaScript for the full interactive view. The summary above is the same information in plain HTML.