Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics
by
Chung, Daniel J H
, Li, Tianyi
, Tadepalli, Sai Chaitanya
, Rudolph, Maja
, Kvasiuk, Yurii
, Gao, Zhiqi
, Münchmeyer, Moritz
, Sala, Frederic
in
Benchmarks
/ Datasets
/ Energy theory
/ Failure modes
/ large language models
/ Physics
/ reasoning
/ Theoretical physics
2025
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics
by
Chung, Daniel J H
, Li, Tianyi
, Tadepalli, Sai Chaitanya
, Rudolph, Maja
, Kvasiuk, Yurii
, Gao, Zhiqi
, Münchmeyer, Moritz
, Sala, Frederic
in
Benchmarks
/ Datasets
/ Energy theory
/ Failure modes
/ large language models
/ Physics
/ reasoning
/ Theoretical physics
2025
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics
by
Chung, Daniel J H
, Li, Tianyi
, Tadepalli, Sai Chaitanya
, Rudolph, Maja
, Kvasiuk, Yurii
, Gao, Zhiqi
, Münchmeyer, Moritz
, Sala, Frederic
in
Benchmarks
/ Datasets
/ Energy theory
/ Failure modes
/ large language models
/ Physics
/ reasoning
/ Theoretical physics
2025
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics
Journal Article
Theoretical physics benchmark (TPBench)—a dataset and study of AI reasoning capabilities in theoretical physics
2025
Request Book From Autostore
and Choose the Collection Method
Overview
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics (TP), focusing on high-energy theory and cosmology. The first iteration of our benchmark consists of 57 problems of varying difficulty, from undergraduate to research level. These problems are novel in the sense that they do not come from public problem collections. We evaluate our data set on various open and closed language models, including o3-mini, o1, DeepSeek-R1, GPT-4o and versions of Llama and Qwen. While we find impressive progress in model performance with the most recent models, our research-level difficulty problems are mostly unsolved. We address challenges of auto-verifiability and grading, and discuss common failure modes. While currently state-of-the art models are still of limited use for researchers, our results show that AI assisted TP research may become possible in the near future. We discuss the main obstacles towards this goal and possible strategies to overcome them. The public problems and solutions, results for various models, and updates to the data set and score distribution, are available on the website of the dataset tpbench.org .
Publisher
IOP Publishing
Subject
MBRLCatalogueRelatedBooks
Related Items
Related Items
We currently cannot retrieve any items related to this title. Kindly check back at a later time.
This website uses cookies to ensure you get the best experience on our website.