Publications

For more information about my research, see the Cognition, Language, Interaction, & Computation (CLIC) lab website, or my Google Scholar profile.

* denotes equal contribution.

Journal Articles

Jones, C. R. & Bergen, B. K. (2026). Large language models pass a standard three-party Turing test. Proceedings of the National Academy of Sciences, 123(21), e2524472123. [paper]

Jones, C. R. & Bergen, B. K. (2026). Lies, damned lies, and language statistics: A comprehensive review of risks from manipulation, persuasion, and deception with large language models. Artificial Intelligence Review, 59(4), 116. [paper]

Moore, J., Overmark, R., Cooper, N., Cibralic, B., Haber, N., & Jones, C. R. (2026). Large language models persuade without planning theory of mind. Cognitive Science (accepted). [preprint]

Jones, C. R., Bergen, B. K., & Trott, S. (2024). Do multimodal large language models and humans ground language similarly? Computational Linguistics. [paper]

Jones, C. R., Trott, S., & Bergen, B. K. (2024). Comparing humans and large language models on an experimental protocol inventory for theory of mind evaluation (EPITOME). Transactions of the Association for Computational Linguistics. [paper]

Jones, C. R. & Bergen, B. K. (2024). Does word knowledge account for the effect of world knowledge on pronoun interpretation? Language and Cognition. [paper]

Trott, S.*, Jones, C. R.*, Michaelov, J. A., Chang, T. A., & Bergen, B. K. (2023). Do large language models know what humans know? Cognitive Science, 47(7). [paper]

Conference Proceedings

Jones, C. R., Lombardi, A., Mahowald, K., & Bergen, B. K. (2026). LLMs and people both learn to form conventions — just not with each other. Proceedings of the Annual Meeting of the Cognitive Science Society, 48. [paper]

Trott, S., Taylor, S., Jones, C. R., Michaelov, J. A., & Rivière, P. D. (2026). Language statistics and false belief reasoning: Evidence from 41 open-weight LMs. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 39816–39833. [paper]

Moore, J., Cooper, N., Overmark, R., Cibralic, B., Haber, N., & Jones, C. R. (2025). Do large language models have a planning theory of mind? Evidence from MindGames: A multi-step persuasion task. Conference on Language Modeling (COLM). [paper]

Jones, C. R., Rathi, I., Taylor, S., & Bergen, B. K. (2025). People cannot distinguish GPT-4 from a human in a Turing test. ACM Conference on Fairness, Accountability, and Transparency (FAccT). [paper]

Pi, Z., Vadaparty, A., Bergen, B. K., & Jones, C. R. (2025). Dissecting the Ullman variations with a SCALPEL: Why do LLMs fail at trivial alterations to the false belief task? Proceedings of the Annual Meeting of the Cognitive Science Society, 47. [paper]

Rathi, I., Bergen, B. K., & Jones, C. R. (2025). Judging the judges: Displacing and inverting the Turing test to investigate the interrogator. Proceedings of the Annual Meeting of the Cognitive Science Society, 47. [paper]

Rivière, P. D., Parkinson-Coombs, O., Jones, C. R., & Trott, S. (2025). Does language stabilize quantity representations in vision transformers? Proceedings of the Annual Meeting of the Cognitive Science Society, 47. [paper]

Rathi, I., Taylor, S., Bergen, B. K., & Jones, C. R. (2025). GPT-4 is judged more human than humans in displaced and inverted Turing tests. Workshop on Detecting AI-Generated Content, COLING. Best Paper Award. [paper]

Jones, C. R. & Bergen, B. K. (2024). Does GPT-4 pass the Turing test? Proceedings of NAACL. [paper]

Jones, C. R. & Trott, S. (2024). Multimodal language models show evidence of embodied simulation. LREC-COLING. [paper]

Jones, C. R., Chang, T. A., Coulson, S., Michaelov, J. A., Trott, S., & Bergen, B. K. (2022). Distributional semantics still can't account for affordances. Proceedings of the Annual Meeting of the Cognitive Science Society, 44. [paper]

Jones, C. R. & Bergen, B. K. (2021). The role of physical inference in pronoun resolution. Proceedings of the Annual Meeting of the Cognitive Science Society, 43. [paper]

Binder, F. J.*, Jones, C. R.*, Kaufman, R. A., Lin, N. T., Poole, C. R., & Vul, E. (2021). Cognitive cost and information gain trade-off in a large-scale number guessing game. Proceedings of the Annual Meeting of the Cognitive Science Society, 43. [paper]

Jones, C. & Kirby, S. (2018). The effect of biasing information on a transmission chain of short texts. 2nd Conference of the Cultural Evolution Society, Tempe, AZ.

Reports

Bengio, Y., Clare, S., Prunkl, C., … Jones, C., et al. (2026). International AI Safety Report 2026. [report]

Bengio, Y., Clare, S., Prunkl, C., … Jones, C., et al. (2025). International AI Safety Report 2025: First Key Update — Capabilities and Risk Implications. [report]

Under Review & Preprints

Schoenegger, P., Salvi, F., Liu, J., Nan, X., Debnath, R., Fasolo, B., Jones, C. R., … & Karger, E. When large language models are more persuasive than incentivized humans, and why. [preprint]

Rivière, P. D., Jones, C. R., & Trott, S. Developmental trajectories of situation modeling and mentalizing in transformer language models. [preprint]

Hager, W., Rathi, I., Hasan, M., & Jones, C. R. Inverse Turing Bench: Evaluating language models as judges of human vs. AI dialogue. [preprint]

Zeng, P., Paige, A. J., Li, W., Brennan, S. E., Rambow, O., & Jones, C. R. Implicit vs. explicit prompting strategies for LVLMs in referential communication. [preprint]

Michaelov, J. A., Arnett, C., Chang, T. A., Rivière, P. D., Taylor, S. M., Jones, C. R., … & Altman, M. How open must language models be to enable reliable scientific inference? [preprint]

Yang, M., Casper, S., Stray, J., Li, J., Jones, C. R., Gausen, A., … & Pelrine, K. AI epistemic risks: Emerging mechanisms and evidence. [preprint]

Schoenegger, P., Jones, C. R., Tetlock, P. E., & Mellers, B. Prompt engineering large language models' forecasting capabilities. [preprint]