Skip to content

Add "Your LLM-as-Judge Is Lying to You" to §8 (LLM-as-judge & verifiers)#46

Open
loopandretry wants to merge 1 commit into
benchflow-ai:mainfrom
loopandretry:add-loop-retry-target5
Open

Add "Your LLM-as-Judge Is Lying to You" to §8 (LLM-as-judge & verifiers)#46
loopandretry wants to merge 1 commit into
benchflow-ai:mainfrom
loopandretry:add-loop-retry-target5

Conversation

@loopandretry

Copy link
Copy Markdown

Adds one practitioner essay to §8 · LLM-as-judge & verifiers.

Link: https://loopandretry.github.io/posts/llm-as-judge-is-lying-to-you/

Why it clears the bar (show your work, not a generic 'you need evals' take): the post is a worked mechanism piece. It shows why raw judge-vs-human agreement is misleading — an 84% agreement number collapses to Cohen's kappa 0.50 (only moderate) once you subtract chance agreement, and an always-'pass' judge that reads nothing scores kappa 0.0 — with the arithmetic and code shown inline. It then covers hardening that actually moves kappa: prefer pairwise over absolute scores, and require the judge to quote its evidence before scoring. Same practitioner lane as the existing Eugene Yan / Hamel Husain / Han-Chung Lee entries in this section.

Disclosure: I write this blog; submitting because it's genuinely on-topic for §8, not a marketing page. Link is live (200); single focused diff.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant