You need to agree to share your contact information to access this dataset

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this dataset content.

InferNet: Data Card for Econometric AI Agent Testset

InferNet is a project aimed at evaluating and building up the AI capability for social science research related to empirical studies. We collect the world’s largest dataset on published empirical tasks in 36 top journals mainly across 2022-2025, over economics, finance, accounting, management, operations research and information system. We have achieved over 2644 samples of triples of empirical task description, data published along with the article, and replication code packages provided by authors. With this data, we are able to establish a benchmark for evaluating LLM and AI Agent’s ability for executing empirical studies in large scale, and provide a good dataset for training AI’s capability for further development of its empirical skills.

We build benchmark for such evaluation of different LLMs, and the lead table will be available publicly for any AI models or agents that are build towards solving social science empirical problems. Based on The University of Hong Kong, Nanyang Technological University, and Shenzhen Loop Area Institute, we would like to launch a global InferNet challenge for competing the lead board of InferNet. This effort will greatly enhance the AI development for social science study, and build up talents in the interdisciplinary field of social science and AI.

Selected_1000 Dataset

Selected_1000 is the 1,000-task benchmark sample. Each task follows a standard format of information includeing the target variable, explanatory variable, control variables, data source, econometric method, task requirements, expected answer, with other metadata including journal, article and task location.

Some statistics information are shown in the following. The task-type distribution below groups detailed method labels into five basic categories: Ordinary Least Squares(OLS), Differences-in-Differences(DID), Instrument Variable(IV), Regression Discontinuity Designs(RDD) and others.

Selected_1000 task type distribution

The tasks are drawn from 10 top-tier business and economics journals including Mangement Science, Econometrica, American Economic Review, etc..

Selected_1000 journal distribution

The field distribution classifies each unique paper into one of five fields based on both journal and article content(for reference only).

Selected_1000 field distribution

MetricsAI Sample Test

This sample test dataset pertains to the study titled Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks. Our goal is to establish a standardized benchmark to evaluate the ability of artificial intelligence models and(or) agents in executing econometric tasks. We invite contributions to expand this econometric task repository and to assess the performance of various large language models (LLMs) and AI agents. If you use our dataset, please kindly cite the original source.

Citation Information

@misc{chen2026aimastereconometricsevidence,
      title={Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks}, 
      author={Qiang Chen and Tianyang Han and Jin Li and Ye Luo and Zigan Wang and Yuxiao Wu and Xiaowei Zhang and Tuo Zhou},
      year={2026},
      eprint={2506.00856},
      archivePrefix={arXiv},
      primaryClass={econ.EM},
      url={https://arxiv.org/abs/2506.00856}, 
}
Downloads last month
65

Paper for troyhan/Infernet