Christopher Manning
@chrmanning
The researchers in university groups are unequalled at the task of contributing to “a much broader and more ingenious stable of evaluations.” University research emphasizes novel discovery and rigorous and skeptical verification. Students who are relatively new to a field but studying it intensively have again and again been the people who come up with novel ideas that break outside received good practice, but which turn out to be good ideas that drive progress. University groups would do the best job of finding issues with model alignment and other socially-negative model behaviors. A successful example of university evaluation is in medical research: pharmaceutical companies often turn to university researchers to validate drugs, for their independence, expertise, and because of the life-or-death importance of getting it right.
One objection that could be made is that universities would be good for this role for a few years, while it is being explored, but not as a permanent assignment. This might actually be right. But I would argue that the next few years are going to be crucially important, and so it would be worth making use of their advantages in this important initial period. On the other hand, there are cases where universities have done such a job for a very long time. For instance, you might not be aware that the University of Nebraska–Lincoln has been testing the performance of tractors for over 100 years (by state government mandate).
And then, within universities, @stanfordnlp would be the right choice. In general, it is the leading group for NLP and LLM research. Particularly in this context, Stanford NLP has done much of the prominent university work on evaluating AI systems and the alignment of AI systems, examining jailbreaking systems, and in studying the positive and negative effects of human beings interacting with AI systems. Stanford NLP has also been prominent in the study of model interpretability, emphasizing causal modeling. And, finally, the proposal for third-party evaluators lies within the thesis of the institutional view of benchmarking presented in our recent @PNASNews article. (References below.)