Glossary › Evaluation

BLEU Score

N-gram precision metric for machine translation: measures overlap between generated and reference text. Range 0-1. Limitations: doesn't capture meaning, sensitive to tokenization. ROUGE is recall-focused variant.

More in Evaluation

LLM-as-Judge · Golden Set · Regression Testing · A/B Testing

Learn this by building

DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary