LLMScholarBench Tests 22 Models and Finds an Accuracy-Representation Trade-Off
The benchmark treats AI-generated expert lists as recommender systems to audit, showing that factual correctness and balanced representation cannot yet be improved together reliably.
3 min