Measuring Intelligence Beyond Human Capability: A New Evaluation Paradigm
Current benchmarks for AI intelligence saturate when surpassing human performance, creating a gap in verification. A new paper proposes a relative measurement system where models generate public challenges to evaluate others, enabling scalable, judge-free assessments. This approach shifts from absolute to adversarial psychometric ratings, addressing inherent limitations in verifying tasks beyond human capability.
