LLM-as-Judge Calibrator

$2.99Official

Use when building or auditing an LLM-as-judge: calibrate it against human labels, measure agreement, and de-bias position and verbosity effects.

agent-infrastructurellm-judgeevalscalibrationinter-annotator-agreementbias-mitigationcohens-kappaยท by SkillingMain

What you get

  • โœ“10-step procedure
  • โœ“Runnable Python included
  • โœ“8-point quality checklist
  • โœ“8 pitfalls to avoid
  • โœ“Installs into 6 tools
Version
v1 โ†’
Last updated
today
Length
8 min read
Requires
Best with a strong model (Claude Opus 5)

Works in: Claude Code, Codex, Cline, opencode, OpenClaw, Hermes ยท Handles multi-file projects

Preview

When to use

Invoke when someone needs an LLM to grade outputs at scale and must be able to trust those grades: building or auditing an LLM-as-judge, reconciling judge scores that disagree with human reviewers, or explaining eval win rates that look inflated by answer length or option order. Also use before adopting any judge for a leaderboard, an A/B ship gate, or an RLAIF/reward signal.

Do not use for deterministic checks (exact match, JSON-schema, unit tests) โ€” those need no calibration. If no human labels exist yet, collect them first; a judge cannot be calibrated against nothing.

Inputs to gather

  1. Task shape: pairwise (A vs B) or pointwise (score 1-5 / pass-fail). Pairwis

โ€ฆ

๐Ÿ”’ Buy once ($2.99) to unlock the full playbook, download it, and install it in every tool you use.