With the help of AI to output scientific research ideas, who gets the good ideas?
“In the past, when I was recruiting research assistants, I would pay more attention to programming ability and execution efficiency. Now AI has significantly lowered the threshold for these abilities.” said Zhu Chen, a professor at the School of Economics and Management at China Agricultural University. Around the Spring Festival this year, she built her own scientific research agent system, breaking down the process of empirical economic research into seven links and handing them over to different AI agents for execution.
The changes are significant. Zhu Chen found that the agent can automatically generate candidate research questions based on a local database. In the 14 runs she recorded, the system generated a total of 79 candidate questions, 87% of which met the three basic conditions of the existence of variables, matching of research design and data structure, and feasible measurement methods. “Formally speaking, it’s passed the test,” she commented.
But how far is “passing” from “excellent”?
A researcher who has long been paying attention to China’s macroeconomic policies and micro-foundations has browsed almost all of the more than 200 AI-generated papers published by the Yanagizawa-Drott team at the University of Zurich. His evaluation was quite restrained: “Only a very small number of selected topics are worthy of further advancement, and most of them are just entry-level graduate students. If they were sent to me, I would reject them directly.”
The problem is depth of analysis. Taking the difference-in-difference method commonly used in economics as an example, the researcher pointed out that a mature empirical paper must not only complete baseline regression, but also conduct parallel trend testing, sensitivity analysis and heterogeneous treatment effect analysis. However, in the current public AI generation papers, these key steps “are often missing, or appear only in a formal way.”
This points to a deeper question: When the topic itself can be generated by an algorithm, where does the originality of scientific research come from?
Xu Xun, chief researcher of BGI Group, divides AI capabilities in the life sciences into five levels. He judged that AI is currently at the L3 stage – “It can complete discoveries, but it cannot independently ask original questions, and asking questions is still an irreplaceable core value of human beings.” In his view, AI can do combinations and reasoning based on existing knowledge, but real scientific questions, those that require questioning beyond the known framework and pointing to the unknown, must still be completed by humans.
Luan Hao, a professor at the School of Cyberspace Security of Xi’an Jiaotong University, believes that the large model is essentially a knowledge searcher. What is different from traditional keyword search is that it can understand the meaning of the search term and then perform precise semantic retrieval in the existing knowledge base. This means that what AI can do is “accurately and comprehensively search existing human knowledge bases”, but “how to obtain new knowledge that humans have not mastered, large models cannot do it for us.” In his view, sometimes the ideas given by AI seem to be very new, “but they are not actually created. Others have provided relevant ideas and they have been incorporated into its own knowledge base.”
The practical experience of Ma Dongxin, associate professor of the Department of Chemistry at Tsinghua University, confirms this point. She calls AI “a good friend who has a lot of interdisciplinary knowledge, is willing to listen, loves to share, and is always available”, but insists that topic selection cannot be left entirely to AI. “AI will give a series of references and provide analysis based on citations. However, limited to the current technical level, fallacies will inevitably occur. You must trace back to the original documents and think independently and critically.” Ma Dongxin said.
She specifically mentioned that discussing scientific issues with AI is valuable in itself – when AI “doesn’t understand the topic”, it just prompts her to refine the problem more clearly. “This thinking process itself can help me clarify my ideas.”
When conducting experimental analysis with the help of AI, who is responsible for errors?
The idea comes to fruition and enters experimentation and analysis. This is the most “hard-core” link in scientific research, and it is also the link where AI is most involved and the risk is most hidden.
Li Linjing, a researcher at the Institute of Automation of the Chinese Academy of Sciences, observed that scientific research agents have been able to complete “a closed-loop autonomous scientific research process of literature reading, experimental planning, parameter iteration, and result analysis.” At the Peking University Stem Cell Research Center, an intelligent platform covering an area of 6 square meters can independently complete more than 120 steps from fingertip blood collection to stem cell cultivation, and the experimental error has been reduced from a maximum of 10% to less than 1%.
The efficiency is exciting, but the pitfalls are equally real.
DeepMind, a subsidiary of Google, used AI to predict 2.2 million new crystal structures, of which 380,000 were judged to be “stable.” However, less than 0.2% have been experimentally verified. Li Guojie, an academician of the Chinese Academy of Engineering, summarized this phenomenon as: AI has generated a large number of “seemingly reasonable” hypotheses, most of which cannot be verified, resulting in scientific research resources being wasted on low-priority options.
The four words “seemingly reasonable” accurately hit the anxiety of many front-line researchers.
Zhang Lianchong, executive deputy director of the National Earth Observation Science Data Center of the Institute of Aerospace Information Innovation of the Chinese Academy of Sciences, made a direct observation: “The general AI system lacks the ability to actively refuse to answer – even when faced with questions beyond its knowledge range, it will still try its best to generate a seemingly reasonable answer, but the accuracy of this answer is not guaranteed.”
Luan Hao explained the root cause of this “serious nonsense” from a technical perspective: when a problem exceeds the knowledge base of the large model, or the large model has a bias in its understanding of the problem, it will “force search” and give irrelevant or even wrong answers.
Wu Qunfang, an associate researcher at Nanjing University of Aeronautics and Astronautics, has repeatedly encountered similar situations in circuit design. For example, he said that the circuit scheme generated by AI “can indeed work in this way based on its basic principles” from a framework perspective, but when it comes to specific engineering indicators such as efficiency, ripple, and dynamic response, “it cannot take everything into account.” When students use the filter circuit or signal conditioning circuit provided by AI to simulate, they often find that the actual results are far from the paper plan. “It’s OK to give an idea, but in the end the actual things used still need to change a lot, and they have to simulate, optimize, and adjust themselves.”
Zhu Chen’s response was to force a recurrence while retaining people’s right to judge at key points. In her system, after each study is completed, a complete R code will be generated, and the researcher can rerun the analysis process based on the original data to confirm that the regression results are consistent with the report. “Reproduction of this step is an indispensable link in the entire process, and it is also the key to ensuring the reliability of the research and avoiding AI illusions.”
Zhang Lianchong and his team have a similar solution – exploring a solution path of “credible corpus + data governance + retrieval enhancement generation”. They have built a trusted corpus that has been strictly reviewed in the background, and all answers from intelligent customer service are limited to this controllable data source.
This model of “human-machine collaboration but with the person ultimately in charge” is becoming the operating method of more and more laboratories. But the problem is: when the output of AI is a “black box”, researchers can neither fully understand the internal logic of the algorithm, but are also required to take full responsibility for the results. Is the division of this responsibility boundary reasonable?
“At present, the actual deployment of large models in enterprises is slow, and this is the core obstacle.” In Luan Hao’s view, the accuracy of answers given by large models cannot be accurately quantified in theory, which is the fundamental bottleneck for its large-scale implementation in the industry. The same goes for scientific research scenarios and industrial control: AI can assist, but the decision-making power must remain in human hands.
Gao Zefeng, an associate researcher at the School of Physics at Renmin University of China, drew a more specific boundary. He encourages students to use AI to assist in writing code, but when it comes to result analysis, students are required to complete it independently. “You must be responsible for all the data involved in your academic achievements. You must not outsource it and act as a ‘hands-off shopkeeper’ yourself.” Gao Zefeng said.
When writing a paper with the help of AI, where should the boundaries be drawn?
The experiment is over and it enters the writing stage. This is the most common and controversial link where AI is involved. A doctoral student in chemistry at a university admitted to reporters that the frequency of AI used by the research team is “almost 100%”, from topic selection, first draft to polishing.
The high-frequency use of “100%” also amplifies two problems: the blurring of ghostwriting boundaries and the proliferation of false references.
“Students sent in English papers with standard sentence grammar and idiomatic wording, but they were not suitable for academic expression.” Ma Dongxin encountered a rather subtle situation. “I asked AI to translate back to Chinese and found that the translation was smooth. I called the student to ask, and it turned out that the Chinese was written first and then translated into English.”
In addition to criticizing and educating, Ma Dongxin is also thinking about where the boundaries are: “AI can assist in building a logical framework, sorting out paper ideas, modifying grammar, and polishing language, but it can never be a ghostwriter.” Gao Zefeng further requires students to complete the paper writing stage independently and cannot use AI tools for assistance.
As a reviewer for academic journals, Wu Qunfang confirmed this from the perspective of a reviewer: “The core of review is to look at innovation – how much is it improved compared to existing technology? How is it improved? What is the technical solution? What is the final verification? These are precisely the areas where AI cannot replace professional judgment.”
Luan Hao used “the finishing touch” to summarize the irreplaceability of people. “A large model may be able to help you complete 80% or even 90% of the work, but the finishing touch is often not possible.” He used cooking as an analogy: the dishes made by AI are like the big pot of rice in the canteen, and are completed based on experience and mechanized processes. In Luan Hao’s view, the most critical stroke in scientific research comes precisely from this “finishing touch” that only people can accomplish.
The costs of relying on AI don’t stop there.
Yang Pingjian, a researcher at the Chinese Academy of Environmental Sciences, discovered that some students no longer read English literature seriously after getting AI, “This is very harmful to students’ scientific research abilities.” He now requires students to read the literature themselves, and after reading it can use AI to translate and summarize it.
“I read at least five papers a week related to my academic direction to understand the most cutting-edge indicators. The field of power electronics is very technical, and I know whether the efficiency can be 99% or 99.1%.” In Wu Qunfang’s view, this continuous sensitivity to the core indicators of the field constitutes the baseline for judging the credibility of AI output.
False citations are another, more subtle trap.
Yang Pingjian gave an example: AI output “China’s carbon market is about to launch” during a content review, and members of the research team found no problem. “China’s carbon market has already been launched in 2021. The database version retrieved by AI is relatively old and does not include the latest data.” In his view, if researchers lack professional sensitivity to common sense and cutting-edge knowledge in their own fields, errors in AI output will quietly enter the official text.
Not long ago, the academic preprint platform arXiv released a systematic review of 2.5 million papers and 111 million references. The results showed that in 2025 alone, there were nearly 150,000 false references fabricated by AI in the four major platforms arXiv, bioRxiv, SSRN and PubMed Central.
Experts analyze that the root cause lies in the way generative AI operates – it tends to generate content that “looks like literature” rather than retrieving real existing data.
“Many databases in professional fields are not open to general AI. DeepSeek or Doubao cannot find the latest cutting-edge literature. Moreover, some of the references it generates actually do not exist at all.” Yang Pingjian couldn’t help but sigh, “What’s the point of this kind of literature review written using AI?”
Wu Qunfang used her own field as an example to support Yang Pingjian’s point of view: “The power electronics discipline is highly dependent on professional databases such as IEEE and IET. However, these databases are ‘authorized and require copyright’, and many contents are not connected to general large models. Good things are out of reach of AI, so the generated content is also available to the public, and there is no real innovation at all.”
This also suggests that the key to solving AI’s false output may not be the model itself, but the controllability of the data source – if researchers can “feed” AI with reviewed professional literature and own data, rather than letting it arbitrarily grab it from the open network, the risk of false citations will be greatly reduced.