1. Ifargan and colleagues developed a platform that guides LLMs to annotate research data and produce comprehensive research papers.
2. The platform demonstrated reliability for simpler tasks, such as creating research papers. Human intervention remained critical as data complexity increased.
Evidence Rating Level: 3 (Average)
Study Rundown: Conducting research and compiling results is challenging and labor-intensive. LLMs have recently shown they can design and run experiments, but their ability to produce accurate findings while maintaining scientific transparency remains uncertain. Ifargan and colleagues developed an automation platform that systemically guides LLMs to complete research manuscripts using annotated data. The platform formulates hypotheses, develops and executes analysis code, interprets findings, and writes papers, either autonomously or with human feedback. The platform’s performance across five case studies involving different fields was evaluated. The primary outcomes included the correctness of analyses and interpretations, reproduction of published findings, and traceability of reported results. The study found that 8 of the 10 manuscripts generated autonomously from two relatively simple datasets contained no major errors. However, the platform’s performance deteriorated with more complex datasets or broader research objectives, and targeted human feedback was required to correct these errors. This study demonstrated that artificial intelligence (AI) could produce traceable research manuscripts from less complex data, but oversight for more complex analyses is still needed.
Click here to read the study in NEJM AI
Relevant Reading: The Future of Research in an Artificial Intelligence-Driven World
In-Depth [case series]: This study evaluated a 17-step workflow using ChatGPT-family models. Two modes were examined in the study: an Autopilot mode that proceeded without human feedback and a Copilot mode that incorporated reviewer comments. The platform’s performance across the following fields was evaluated: health indicators, social networks, infections, neonatal treatment policy, and pediatric intubation. During initial testing, five manuscripts were generated from a health-indicator dataset, and another five were generated from a congressional social-network dataset. Researchers manually reviewed analysis code and manuscript text for major errors, imperfections, and correct interpretations. The primary outcomes included the correctness of analyses and interpretations, reproduction of published findings, and traceability of reported results. The platform correctly reproduced their analysis outputs in tables, with 8 of the 10 manuscripts containing no major errors. Four autonomous analyses using a more complex longitudinal infection dataset led to major data-handling errors that required human feedback. In ten neonatal treatment-policy analyses, all were reproduced correctly, but two papers contained interpretation errors. For analyses involving pediatric intubations, the broad research task produced a 90% error rate. This study was limited by the small number of selected case studies and dependence on expert feedback. Overall, this study demonstrated that AI has the potential to accelerate biomedical research, but human oversight remains crucial to ensure accuracy.
Image: PD
©2026 2 Minute Medicine, Inc. All rights reserved. No works may be reproduced without expressed written consent from 2 Minute Medicine, Inc. Inquire about licensing here. No article should be construed as medical advice and is not intended as such by the authors or by 2 Minute Medicine, Inc.




