Use a second model call as the judge when the correct answer is open-ended, writing a rubric the grader applies consistently, and knowing where LLM-as-judge is unreliable.
I
IBM TechnologyVideo
6 minutes
Direct Assessment and Pairwise Comparison in LLM-as-a-Judge Evaluation