Failure Modes Specific to LLM-as-Judge Systems
How judge models systematically distort evaluations through structural flaws.
Marcus Ołtarzewski
Staff Writer & Data Correspondent
Marcus trained as a computational linguist and spent years annotating and auditing large evaluation datasets for academic consortia before moving into technical writing in 2013. He brings a skeptic's eye to inter-rater agreement, calibration drift, and the reproducibility of judge verdicts across model versions.
1 story
How judge models systematically distort evaluations through structural flaws.