Glossary › Evaluation

LLM-as-Judge

Using a strong LLM (e.g. GPT-4) to evaluate outputs of another model. Cheaper than human eval, scalable. Risks: bias toward own style, position bias, verbosity bias. Mitigate with rubrics and multiple judges.

Where this is taught

More in Evaluation

Golden Set · Regression Testing · A/B Testing · BLEU Score

Learn this by building

DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary