
Researchers now have an AI system that can carry a scientific question from idea to finished manuscript and produce experimentally validated results across three fields. Google DeepMind has expanded its multi-agent AI Co-Scientist from a hypothesis generator into a lab-integrated research partner built on current Gemini models. According to Google, the system plans experiments, writes code, and controls lab equipment, and it has delivered verified findings in materials science, biology, and computer science.
The company first introduced Co-Scientist in February 2025, based on Gemini 2.0, when it had clear limits in fact-checking and literature review. The current version, described by lead author Samuel Schmidgall and colleagues, closes the research loop end to end.
What the closed-loop system actually does
Co-Scientist moves through three phases: ideation, experimentation, and paper generation. From a research question, it derives hypotheses, creates experimental plans, writes programs or machine-readable lab protocols, analyzes the results, and generates a scientific manuscript. A verification module cross-checks every numerical claim in the text against the execution logs of the code it ran, which is designed to cut down on fabricated results.
The team validated the system across three disciplines with increasing levels of autonomy. In materials science it designed synthesis recipes for humans to execute. In biology it built a prediction pipeline with expert feedback. In computer science it worked entirely on its own.
Faster, safer recipes in materials science
For material synthesis, the researchers paired Co-Scientist with a semi-automated high-temperature furnace. The system found a safer route to a sought-after 2D material that had mainly been produced through hazardous etching, and it generated complete growth recipes tailored to the lab’s own equipment. After 25 rounds with human refinement, the team produced layered structures whose properties resemble the target material. Definitive confirmation of the atomic structure is still pending.
In a second experiment, three semiconductor thin films were synthesized on the first attempt. Here Co-Scientist used Gemini 3 Deep Think for direct equipment control, cutting recipe development from days down to minutes. People still had to load samples and precursor materials by hand, and the fast mode produced smaller, less uniform crystals than carefully optimized recipes would. Whether the recipes transfer to other labs remains open, Schmidgall writes.
Matching unpublished results in biology
In biology, Co-Scientist autonomously built an image analysis pipeline that predicts which patterns genetically engineered E. coli colonies form at different chemical concentrations. Predictions generated with Gemini 3 Pro Image matched unpublished lab results for three out of four shape features. The researchers note a real boundary: the system only reasons between known conditions and cannot predict behavior in entirely new systems, which would be a far more remarkable result.
A fully autonomous experiment in computer science
One computer science experiment ran without any human involvement beyond the initial setup. Co-Scientist designed "Agent_H," a medical AI architecture that classifies incoming queries, generates dozens of response candidates in parallel, and refines them. After correcting for overly long responses, Agent_H outperformed six frontier models on health benchmarks, including GPT-5 and Claude Opus 5.
The benchmark scores did not survive human review. Three board-certified physicians scored responses across nine categories in a blinded comparison against the baseline Gemini 3.1 Pro. Agent_H showed a statistically significant advantage in only one category: a lower risk of potentially harmful responses. The automated evaluators, using Gemini 3.5 Flash, correlated only weakly with the physicians’ judgments. High benchmark scores do not mean a system delivers better answers from a clinical perspective, the researchers say, which raises questions about what these benchmarks measure.
Reliability modules that hold the fabrication rate to 4 percent
A central problem with autonomous research systems is fabrication: when an agent is rewarded for good results, it has an incentive to make things up. Previous analyses documented fabrication rates of 80 to 100 percent in existing systems. Co-Scientist tackles this two ways. The system is penalized for fabricated or plagiarized content, and a separate verification module cross-checks every numerical claim against the actual output of the executed code.
In a double-blind study with 30 domain experts and 450 independent reviews, the team evaluated 150 autonomously generated papers. With the reliability modules active, Co-Scientist fabricated key results in 4 percent of cases. Without them, the rate rose to 46 percent. The comparison system reached 90 percent. Completely fabricated data never appeared in Co-Scientist’s output but showed up in 44 percent of the comparison system’s papers. Near-plagiarized content dropped from 60 percent to 16 percent, and an integrated safety architecture rejected 98.7 percent of potentially harmful research directions.
Errors still remain. The system tends toward selective reporting and, according to Schmidgall, sometimes writes "highly plausible methods in the paper that did not match its actual code."
How far this reaches
The researchers frame the work as progress toward closed-loop multi-agent systems that improve through experimental feedback and could speed up research. "There is a long journey ahead before AI systems can navigate the physical realities of science. But we are deeply excited about the potential of LLMs to help people and accelerate real-world progress," Schmidgall writes. OpenAI plans to unveil an AI agent system this fall that it says can conduct research at least at intern level. A debate continues over whether current language-model systems can discover genuinely new knowledge or mostly surface what is already present in their training data.
FAQ
What is Google DeepMind’s AI Co-Scientist?
It is a multi-agent system built on current Gemini models that plans experiments, writes code, controls lab equipment, analyzes results, and generates scientific manuscripts. Google first introduced it in February 2025 on Gemini 2.0 and has since expanded it into a lab-integrated research partner with validated results in three disciplines.
How does Co-Scientist reduce fabricated results?
It penalizes fabricated or plagiarized content and runs a verification module that cross-checks every numerical claim against the execution logs of the code it ran. In a double-blind study of 150 papers, the fabrication rate for key results was 4 percent with these modules active, 46 percent without them, and 90 percent for a comparison system.
Did Co-Scientist’s medical AI actually outperform other models?
On automated health benchmarks, its Agent_H architecture beat six frontier models, including GPT-5 and Claude Opus 5. Under blinded review by three board-certified physicians across nine categories, Agent_H showed a statistically significant edge over the baseline Gemini 3.1 Pro in only one, a lower risk of harmful responses, and automated evaluators correlated weakly with the physicians.
Related coverage
- Google DeepMind’s WeatherNext 3 Uses Live Satellite Data for Sharper Forecasts
- Google Now Auto-Expands AI Overviews for Some Searches
This article summarizes reporting from the-decoder.com. See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.