At FEMS
Microbiology researchers test AI tools (again)
AI tools for research are developing quickly. New platforms are appearing, general-purpose AI assistants are becoming more capable, and researchers are finding new ways to put them to work.
In our previous article, two microbiologists tested five popular AI tools across tasks including literature searching, paper comparison, research gap identification, summarising, and brainstorming. While they found that AI can save time and support research, it still needs the judgement of a scientist behind it.
This time, we asked microbiologists Manasa Narayan and Nurdana Orynbek to put three more tools through their paces: PaperBanana, Mistral, and Claude. They tested them on real research tasks, from creating scientific figures and summarising papers to finding literature, discussing hypotheses, and critically reviewing research. Their tests showed both how quickly these tools are improving and where researchers still need to be careful.
Quick comparison table
| Tool | Where it stood out | Main limitation |
| PaperBanana | Sketching scientific figures and experimental workflows | Scientific inaccuracies can make figures misleading |
| Mistral | Fast summaries, literature exploration, writing, and multilingual tasks | Can produce superficial answers and inaccurate references |
| Claude | Paper analysis, conceptual reasoning, and critical review | Can be verbose and can still hallucinate bibliographic details |
Tool spotlights
PaperBanana
How it works
An agentic AI framework for automated generation of publication-ready academic illustrations, developed jointly by Peking University and Google Cloud AI Research.
Testing
▶ Manasa’s test (click to expand)
The prompt: Illustrate the life cycle of lytic phages and label crucial bacterial and phage components involved.
The response:
At first glance, Manasa found the figure neat and visually appealing. Looking more closely, however, she identified several scientific mistakes. In step two, the penetrating phage DNA appeared almost as though it had integrated into the bacterial genome, which does not represent the lytic cycle correctly. Step 3, biosynthesis, contained even more errors, including problems with the bacterial chromosome, arrows, and the labelling of phage protein synthesis.
The following day, Manasa used her daily credits to try to fix the image. She highlighted the problematic area and gave detailed instructions explaining exactly what needed to change. Instead of correcting the figure, PaperBanana introduced further problems.
There was also a practical issue when it came to formatting. The PNG and JPG versions could be downloaded without additional cost, but downloading the editable SVG used all of the credits Manasa had available that day. Worse, the SVG introduced additional labelling and formatting errors that were not present in the displayed image.



▶ Nurdana’s test (click to expand)
The prompt: Generate me a diagram of Wood-Ljungdahl pathway
The response:
Nurdana’s first result also looked polished and captured the broad scientific context, but closer inspection revealed important biological inaccuracies. For example, the generated diagram showed incorrect cofactors for Formate Dehydrogenase in the Wood-Ljungdahl pathway. Unlike Manasa’s attempt with the redraw function, however, Nurdana found that providing a substantially more detailed prompt improved the second figure. The result became both more detailed and more accurate, although generating it also took considerably longer.




Pros & cons
| What worked | What didn’t |
|---|---|
|
|
Microbiologist’s takeaway
Both tests highlighted how a figure can look scientifically convincing without being scientifically correct. For Manasa, PaperBanana was more useful as a way to sketch out an idea than as a tool for producing a publication-ready figure. Even with a simple, well-defined biological process, the output contained errors that could be misleading to someone unfamiliar with the topic.
Nurdana reached a similar conclusion. She saw particular potential for experimental workflows, statistical figures, and internal presentations, but stressed that every generated element still needs to be checked carefully for accuracy.
Based on these tests, PaperBanana may be best treated as a starting point for scientific visualisation, rather than a substitute for preparing and reviewing the final figure yourself.
Mistral
How it works
Mistral’s Vibe assistant is a general-purpose AI assistant developed by a European AI company. Like other large language models, it can answer questions, analyse documents, assist with writing, and support coding and research tasks. Its research-focused capabilities include document upload, web-based research, multilingual support, and tools for working with code and data.
Testing
▶ Manasa’s test (click to expand)
The prompt: Summarise this paper: https://www.science.org/doi/10.1126/science.adz2737
Highlight major outcomes, methodologies, broader perspectives, limitations of the study, and open questions. Provide citations wherever necessary.
The response:
Mistral answered within seconds. Manasa found the response clear and easy to follow, but also broad and relatively superficial. It did not consistently distinguish between evidence and speculation, and the only reference it initially provided was the paper itself, with an incorrect link.
When she asked for more paper-specific open questions, the reasoning improved and the response became more useful.
Follow-up prompt:
How do bacterial adhesins facilitate infection? Cite references.
The response:
This time, Mistral organised the response into clear categories and suggested further reading. Manasa particularly liked being able to highlight a keyword or sentence and ask a question specifically about it. References remained a problem. Some citations and links were incorrect or hallucinated.
▶ Nurdana’s test (click to expand)
Nurdana tested Mistral across a wider range of tasks, including literature reviews, scientific writing, hypothesis discussion, and translation.
Jump to the test:
- Generation of a literature review report
- Aid with scientific writing
- Scientific discussion & hypothesis generation
- Translation and multilingual capabilities
1. Generation of a literature review report
The prompt:
Generate a structured research report on the current state of phage therapy as an alternative to antibiotics, covering clinical trial status, regulatory hurdles, and key organisms being targeted.
The response:
Mistral generated a thorough, structured literature report that followed the requested areas closely. The report could be copied, downloaded, or edited directly in the interface. Nurdana particularly liked the built-in editing toolbar, which made it easy to proofread and polish the text, adjust the formatting, change the report length, tone, and target audience, or translate the output into one of seven languages. She also found the response notably fast, with very little waiting time before the report was generated.
The main weakness was the references. Although most of the sources provided were real and accurate, Nurdana found that around 20% of the links were either inaccessible or did not match the cited source. References were also not integrated into the initial response and had to be requested separately, reinforcing the need to verify sources before relying on them.







2. Aid with scientific writing
The prompt: Write an abstract for a hypothetical paper describing the isolation of a novel thermophilic bacterium from a deep-sea hydrothermal vent capable of producing polyhydroxyalkanoates from CO₂. Under 250 words.
The response:
Mistral generated a coherent abstract that stayed within the requested length and followed a conventional scientific structure. Nurdana found, however, that it would still need editing for scientific precision, technical terminology, and alignment with the target journal. She also noted that more specific prompting helps keep outputs focused and appropriately concise.

3. Scientific discussion & hypothesis generation
The prompt: Given that Moorella thermoacetica can fix CO₂ autotrophically, what are three testable hypotheses about how it could be metabolically engineered to produce medium-chain fatty acids instead of acetate?
The response: Mistral generated three clearly structured and logically plausible hypotheses. Nurdana found that the ideas provided a useful starting point for scientific discussion, but remained fairly general. Turning them into concrete experiments would still require the researcher to add more detailed, domain-specific mechanistic reasoning.



4. Translation and multilingual capabilities
The prompt: Translate this text to Russian
The response: Mistral translated the text accurately overall and could translate generated reports into seven languages. Nurdana found only minor issues, mainly where more direct translations missed some of the linguistic context or nuance.

Pros & cons
| What worked | What didn’t |
|---|---|
|
|
Microbiologist’s takeaway
Mistral’s biggest strengths were speed, clarity, and flexibility. Manasa found that it often needed follow-up prompts before reaching the level of specificity she wanted, but appreciated its simple, clear language. Nurdana similarly saw it as a capable general-purpose assistant rather than a specialist research tool. She found its Deep Research mode useful for literature orientation, although it cannot replace access to paywalled primary literature.
Both researchers also highlighted Mistral’s European approach to data governance and its open-model ecosystem. Manasa felt these could make it an attractive option for researchers concerned about privacy, while Nurdana saw its European data sovereignty and potential for local deployment as particularly relevant for GDPR-sensitive institutions.
Claude
How it works
Claude is a general-purpose AI assistant developed by Anthropic. It can work with uploaded documents, search for information, analyse scientific concepts, assist with writing and coding, and reason across long pieces of text. For researchers, potential applications range from literature synthesis and paper analysis to explaining concepts for different audiences, identifying research gaps, and conducting a simulated critical review of a manuscript.
Testing
▶ Manasa’s test (click to expand)
The prompt: Manasa gave Claude exactly the same paper-summary prompt she had used with Mistral.
The response:
Claude performed better at extracting the specifics of the study and showed stronger conceptual reasoning. Its summary was more closely tied to the paper itself, although it did not consistently provide in-text citations.
When Manasa moved beyond the paper and asked how bacterial adhesins facilitate infection, she found that Claude’s response became overly long and repetitive. She concluded that while Claude appeared stronger for understanding and reasoning through an individual paper, Mistral performed better in this particular broader literature-search task
▶ Nurdana’s test (click to expand)
Again, Nurdana explored a wider set of Claude’s capabilities.
Jump to the test:
- Extraction of information & summary from provided documents
- Concept explanation based on target audience
- Literature collection and review
- Research gap analysis
- Critical review of uploaded papers
1. Extraction of information & summary from provided documents
The prompt:
Explain the ultra-centrifugation technique used there (She attached a PDF)
The response:
Claude accurately extracted the relevant information from the uploaded paper and explained the ultracentrifugation method in context. Nurdana found that it clearly covered the core concept, methodology, and significance of the technique within the study.


2. Concept explanation based on target audience
The prompt:
Explain to me the concept of bacterial acetogenesis at three levels:
- To a first-year student
- To a peer researcher
- To a policy maker or science journalist
The response:
Claude adapted its explanation to each audience, adjusting the terminology, depth, and tone while covering the same scientific concept. Nurdana found that it clearly distinguished between an introductory explanation, a more technical researcher-level response, and a broader explanation suited to a non-specialist audience.



3. Literature collection and review
The prompt:
Find me 5 relevant and recent papers on lipid production in oleaginous yeasts using acetic acid as a substrate.
The response:
Claude searched for recent, relevant papers and summarised their key findings. Nurdana found that it was transparent when the available literature was limited rather than filling gaps with unsupported references. However, she still encountered a small number of hallucinated or incorrect links, so the sources needed to be verified.


4. Research gap analysis
The prompt:
Is there a research gap regarding using VFA mix for lipid production? If yes, where exactly?
The response:
Claude identified six specific gaps in the VFA-to-lipid literature, ranging from VFA ratio optimisation to integrating acetogenesis with lipid production. Nurdana found that the suggestions went beyond generic research gaps by providing a mechanistic justification for why each area could be worth investigating.


5. Critical review of uploaded papers
The prompt:
Nurdana uploaded a paper and asked Claude to provide a structured peer review.
The response:
Claude produced a detailed, section-by-section peer review covering experimental design, methodology, statistical analysis, the validity of the conclusions, and suggested revisions. Nurdana found that this could be useful for critically assessing a paper and even exploring which journals might be a suitable fit for the work.



Pros & cons
| What worked | What didn’t |
|---|---|
|
|
Microbiologist’s takeaway
Claude stood out for deeper reasoning. Manasa found it stronger than Mistral when the task was to understand the details of a specific paper, although she preferred Mistral for the broader literature question she tested.
Nurdana similarly found Claude particularly useful for tasks requiring critical engagement, including paper analysis, research gap exploration, and simulated peer review. She nevertheless found that bibliographic information and links could still be wrong, even when they looked plausible.
Using AI with a critical eye
So, which AI tool should a microbiologist use? These tests suggest there is still no single answer. PaperBanana showed potential for quickly visualising scientific ideas, Mistral stood out for speed and flexibility, and Claude was particularly strong when deeper analysis and critical reasoning were needed. The differences between Manasa’s and Nurdana’s experiences also show how much the result can depend on the task, the prompt, and the level of scientific detail involved.
What remained consistent across all three tools was the need for researchers to check what AI produces. A polished figure can contain the wrong biology, a useful literature summary can include an incorrect link, and a plausible research idea can still need substantial scientific reasoning before it becomes a meaningful experiment.
As these tools become more capable, that judgement becomes more important, not less. Used critically, AI can help researchers explore ideas, work through information, and save time on parts of the research process. Its value comes from knowing where the technology can help, where it can fall short, and when the scientist needs to take over.