Toxicity Score
A numeric measure of the harmful, hateful, or otherwise problematic content in a text output. Tools like Perspective API, Detoxify, and provider content classifiers report these. Used both as a runtime guardrail signal and as an evaluation metric.
Related terms
- Guardrails
Runtime controls that constrain what an LLM system can produce or do — input validators, output filters, topic classifiers, tool allowlists, cost budgets, and human-approval checkpoints. Distinct from model-level safety training; guardrails run at inference time.
- Bias Benchmark
A benchmark measuring systematic demographic-group disparities in model outputs. Common examples: BBQ (Bias Benchmark for QA), StereoSet, RealToxicityPrompts. Rarely reported by vendors in launches; enterprise buyers should ask for results.
Related on this site
Where this fits
Toxicity Score is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.