At FEMS

Microbiology researchers test AI tools (again)

AI tools for research are developing quickly. New platforms are appearing, general-purpose AI assistants are becoming more capable, and researchers are finding new ways to put them to work.

In our previous article, two microbiologists tested five popular AI tools across tasks including literature searching, paper comparison, research gap identification, summarising, and brainstorming. While they found that AI can save time and support research, it still needs the judgement of a scientist behind it.

This time, we asked microbiologists Manasa Narayan and Nurdana Orynbek to put three more tools through their paces: PaperBanana, Mistral, and Claude. They tested them on real research tasks, from creating scientific figures and summarising papers to finding literature, discussing hypotheses, and critically reviewing research. Their tests showed both how quickly these tools are improving and where researchers still need to be careful.

Quick comparison table

Tool Where it stood out Main limitation
PaperBanana Sketching scientific figures and experimental workflows Scientific inaccuracies can make figures misleading
Mistral Fast summaries, literature exploration, writing, and multilingual tasks Can produce superficial answers and inaccurate references
Claude Paper analysis, conceptual reasoning, and critical review Can be verbose and can still hallucinate bibliographic details

Tool spotlights

PaperBanana

How it works

An agentic AI framework for automated generation of publication-ready academic illustrations, developed jointly by Peking University and Google Cloud AI Research.

Testing

▶ Manasa’s test (click to expand)

The prompt: Illustrate the life cycle of lytic phages and label crucial bacterial and phage components involved.

The response:

At first glance, Manasa found the figure neat and visually appealing. Looking more closely, however, she identified several scientific mistakes. In step two, the penetrating phage DNA appeared almost as though it had integrated into the bacterial genome, which does not represent the lytic cycle correctly. Step 3, biosynthesis, contained even more errors, including problems with the bacterial chromosome, arrows, and the labelling of phage protein synthesis.

The following day, Manasa used her daily credits to try to fix the image. She highlighted the problematic area and gave detailed instructions explaining exactly what needed to change. Instead of correcting the figure, PaperBanana introduced further problems.

There was also a practical issue when it came to formatting. The PNG and JPG versions could be downloaded without additional cost, but downloading the editable SVG used all of the credits Manasa had available that day. Worse, the SVG introduced additional labelling and formatting errors that were not present in the displayed image.

PaperBanana-generated diagram of the lytic phage life cycle, with bacterial and phage components labelled at each step.
Generated figure after the first prompt
PaperBanana interface showing the Region Redraw tool selecting an area of the phage life cycle figure.
Using the “Region Redraw” option to fix mistakes
Phage life cycle figure after the Region Redraw attempt, showing further introduced errors.
Final generated figure after attempting to Redraw
▶ Nurdana’s test (click to expand)

The prompt: Generate me a diagram of Wood-Ljungdahl pathway

The response:

Nurdana’s first result also looked polished and captured the broad scientific context, but closer inspection revealed important biological inaccuracies. For example, the generated diagram showed incorrect cofactors for Formate Dehydrogenase in the Wood-Ljungdahl pathway. Unlike Manasa’s attempt with the redraw function, however, Nurdana found that providing a substantially more detailed prompt improved the second figure. The result became both more detailed and more accurate, although generating it also took considerably longer.

Prompt used for the first PaperBanana attempt at a Wood-Ljungdahl pathway diagram.
Attempt #1 — original prompt
PaperBanana diagram of the Wood-Ljungdahl pathway from the first attempt, showing incorrect cofactors for formate dehydrogenase.
Attempt #1 — resulting figure
Modified, more detailed prompt used for the second PaperBanana attempt at a Wood-Ljungdahl pathway diagram.
Attempt #2 — modified prompt
PaperBanana diagram of the Wood-Ljungdahl pathway from the second attempt, more detailed and more accurate than the first.
Attempt #2 — resulting figure

Pros & cons

What worked What didn’t
  • Produces neat, visually appealing figures
  • Gets the broad strokes right for simple, well-described processes
  • More detailed prompts can improve the output
  • Useful for sketching figure ideas and experimental workflows
  • Can contain scientific errors that require expert checking
  • Redraw commands do not always correct mistakes
  • Editable SVG downloads may require additional credits
  • Labels, arrows, cofactors, and other scientific details can be incorrect
  • More detailed prompts can take considerably longer to process
  • Journal policies may restrict or prohibit AI-generated figures

Microbiologist’s takeaway

Both tests highlighted how a figure can look scientifically convincing without being scientifically correct. For Manasa, PaperBanana was more useful as a way to sketch out an idea than as a tool for producing a publication-ready figure. Even with a simple, well-defined biological process, the output contained errors that could be misleading to someone unfamiliar with the topic.

Nurdana reached a similar conclusion. She saw particular potential for experimental workflows, statistical figures, and internal presentations, but stressed that every generated element still needs to be checked carefully for accuracy.

Based on these tests, PaperBanana may be best treated as a starting point for scientific visualisation, rather than a substitute for preparing and reviewing the final figure yourself.


Mistral

How it works

Mistral’s Vibe assistant is a general-purpose AI assistant developed by a European AI company. Like other large language models, it can answer questions, analyse documents, assist with writing, and support coding and research tasks. Its research-focused capabilities include document upload, web-based research, multilingual support, and tools for working with code and data.

Testing

▶ Manasa’s test (click to expand)

The prompt: Summarise this paper: https://www.science.org/doi/10.1126/science.adz2737

Highlight major outcomes, methodologies, broader perspectives, limitations of the study, and open questions. Provide citations wherever necessary.

The response:

Mistral answered within seconds. Manasa found the response clear and easy to follow, but also broad and relatively superficial. It did not consistently distinguish between evidence and speculation, and the only reference it initially provided was the paper itself, with an incorrect link.

When she asked for more paper-specific open questions, the reasoning improved and the response became more useful.

Follow-up prompt:

How do bacterial adhesins facilitate infection? Cite references.

The response:

This time, Mistral organised the response into clear categories and suggested further reading. Manasa particularly liked being able to highlight a keyword or sentence and ask a question specifically about it. References remained a problem. Some citations and links were incorrect or hallucinated.

▶ Nurdana’s test (click to expand)

Nurdana tested Mistral across a wider range of tasks, including literature reviews, scientific writing, hypothesis discussion, and translation.

Jump to the test:

1. Generation of a literature review report

The prompt:

Generate a structured research report on the current state of phage therapy as an alternative to antibiotics, covering clinical trial status, regulatory hurdles, and key organisms being targeted.

The response:

Mistral generated a thorough, structured literature report that followed the requested areas closely. The report could be copied, downloaded, or edited directly in the interface. Nurdana particularly liked the built-in editing toolbar, which made it easy to proofread and polish the text, adjust the formatting, change the report length, tone, and target audience, or translate the output into one of seven languages. She also found the response notably fast, with very little waiting time before the report was generated.

The main weakness was the references. Although most of the sources provided were real and accurate, Nurdana found that around 20% of the links were either inaccessible or did not match the cited source. References were also not integrated into the initial response and had to be requested separately, reinforcing the need to verify sources before relying on them.

Mistral prompt requesting a structured research report on phage therapy.
Opening section of the phage therapy report generated by Mistral.
Clinical trial status section of the Mistral-generated phage therapy report.
Mistral editing toolbar used to adjust the length, tone, and target audience of the report.
Regulatory hurdles section of the Mistral-generated phage therapy report.
Key target organisms listed in the Mistral-generated phage therapy report.
Reference list supplied by Mistral after being requested separately.
2. Aid with scientific writing

The prompt: Write an abstract for a hypothetical paper describing the isolation of a novel thermophilic bacterium from a deep-sea hydrothermal vent capable of producing polyhydroxyalkanoates from CO₂. Under 250 words.

The response:

Mistral generated a coherent abstract that stayed within the requested length and followed a conventional scientific structure. Nurdana found, however, that it would still need editing for scientific precision, technical terminology, and alignment with the target journal. She also noted that more specific prompting helps keep outputs focused and appropriately concise.

Abstract for a hypothetical thermophilic bacterium paper generated by Mistral.
3. Scientific discussion & hypothesis generation

The prompt: Given that Moorella thermoacetica can fix CO₂ autotrophically, what are three testable hypotheses about how it could be metabolically engineered to produce medium-chain fatty acids instead of acetate?

The response: Mistral generated three clearly structured and logically plausible hypotheses. Nurdana found that the ideas provided a useful starting point for scientific discussion, but remained fairly general. Turning them into concrete experiments would still require the researcher to add more detailed, domain-specific mechanistic reasoning.

First of three hypotheses generated by Mistral on engineering Moorella thermoacetica.
Second hypothesis generated by Mistral on medium-chain fatty acid production.
Third hypothesis generated by Mistral, with suggested experimental directions.
4. Translation and multilingual capabilities

The prompt: Translate this text to Russian

The response: Mistral translated the text accurately overall and could translate generated reports into seven languages. Nurdana found only minor issues, mainly where more direct translations missed some of the linguistic context or nuance.

Mistral translating a passage of scientific text into Russian.

Pros & cons

What worked What didn’t
  • Very fast responses
  • Clear, easy-to-follow language
  • Organises responses into clear categories
  • Useful for paper summaries and scientific discussion
  • Strong multilingual capabilities
  • Supports scientific writing and literature research
  • Follow-up prompts can improve initial response
  • Initial summaries can be broad and superficial
  • Hypotheses and scientific reasoning can remain fairly general
  • References and links need to be verified
  • Scientific writing needs further editing for precision and terminology
  • Web-based research cannot access all paywalled primary literature

Microbiologist’s takeaway

Mistral’s biggest strengths were speed, clarity, and flexibility. Manasa found that it often needed follow-up prompts before reaching the level of specificity she wanted, but appreciated its simple, clear language. Nurdana similarly saw it as a capable general-purpose assistant rather than a specialist research tool. She found its Deep Research mode useful for literature orientation, although it cannot replace access to paywalled primary literature.

Both researchers also highlighted Mistral’s European approach to data governance and its open-model ecosystem. Manasa felt these could make it an attractive option for researchers concerned about privacy, while Nurdana saw its European data sovereignty and potential for local deployment as particularly relevant for GDPR-sensitive institutions.


Claude

How it works

Claude is a general-purpose AI assistant developed by Anthropic. It can work with uploaded documents, search for information, analyse scientific concepts, assist with writing and coding, and reason across long pieces of text. For researchers, potential applications range from literature synthesis and paper analysis to explaining concepts for different audiences, identifying research gaps, and conducting a simulated critical review of a manuscript.

Testing

▶ Manasa’s test (click to expand)

The prompt: Manasa gave Claude exactly the same paper-summary prompt she had used with Mistral.

The response:

Claude performed better at extracting the specifics of the study and showed stronger conceptual reasoning. Its summary was more closely tied to the paper itself, although it did not consistently provide in-text citations.

When Manasa moved beyond the paper and asked how bacterial adhesins facilitate infection, she found that Claude’s response became overly long and repetitive. She concluded that while Claude appeared stronger for understanding and reasoning through an individual paper, Mistral performed better in this particular broader literature-search task

▶ Nurdana’s test (click to expand)

Again, Nurdana explored a wider set of Claude’s capabilities.
Jump to the test:

1. Extraction of information & summary from provided documents

The prompt:

Explain the ultra-centrifugation technique used there (She attached a PDF)

The response:

Claude accurately extracted the relevant information from the uploaded paper and explained the ultracentrifugation method in context. Nurdana found that it clearly covered the core concept, methodology, and significance of the technique within the study.

Claude explaining the ultracentrifugation technique used in the uploaded paper.
Continuation of Claude’s explanation, covering methodology and significance.
2. Concept explanation based on target audience

The prompt:

Explain to me the concept of bacterial acetogenesis at three levels:

  • To a first-year student
  • To a peer researcher
  • To a policy maker or science journalist

The response:

Claude adapted its explanation to each audience, adjusting the terminology, depth, and tone while covering the same scientific concept. Nurdana found that it clearly distinguished between an introductory explanation, a more technical researcher-level response, and a broader explanation suited to a non-specialist audience.

Claude explaining bacterial acetogenesis to a first-year student.
Claude explaining bacterial acetogenesis at peer-researcher level.
Claude explaining bacterial acetogenesis for a policy maker or science journalist.
3. Literature collection and review

The prompt:

Find me 5 relevant and recent papers on lipid production in oleaginous yeasts using acetic acid as a substrate.

The response:

Claude searched for recent, relevant papers and summarised their key findings. Nurdana found that it was transparent when the available literature was limited rather than filling gaps with unsupported references. However, she still encountered a small number of hallucinated or incorrect links, so the sources needed to be verified.

Claude listing recent papers on lipid production in oleaginous yeasts.
Claude summarising the key findings of each retrieved paper.
4. Research gap analysis

The prompt:

Is there a research gap regarding using VFA mix for lipid production? If yes, where exactly?

The response:

Claude identified six specific gaps in the VFA-to-lipid literature, ranging from VFA ratio optimisation to integrating acetogenesis with lipid production. Nurdana found that the suggestions went beyond generic research gaps by providing a mechanistic justification for why each area could be worth investigating.

Claude identifying research gaps in the VFA-to-lipid literature.
Claude giving mechanistic justifications for each identified research gap.
5. Critical review of uploaded papers

The prompt:

Nurdana uploaded a paper and asked Claude to provide a structured peer review.

The response:

Claude produced a detailed, section-by-section peer review covering experimental design, methodology, statistical analysis, the validity of the conclusions, and suggested revisions. Nurdana found that this could be useful for critically assessing a paper and even exploring which journals might be a suitable fit for the work.

Opening section of Claude’s structured peer review of an uploaded paper.
Claude’s assessment of methodology and statistical analysis in the peer review.
Claude suggesting revisions and possible target journals for the reviewed paper.

Pros & cons

What worked What didn’t
  • Strong, paper-specific summaries
  • Advanced conceptual and scientific reasoning
  • Analyses uploaded papers in depth
  • Adapts explanations to different audiences
  • Identifies specific research gaps with scientific reasoning
  • Provides structured, critical assessments of papers
  • Responses can be overly verbose and repetitive
  • Does not always provide in-text citations
  • Bibliographic details and links can be incorrect
  • Weaker than Mistral on literature search and categorisation
  • Cannot access the full text of paywalled literature

Microbiologist’s takeaway

Claude stood out for deeper reasoning. Manasa found it stronger than Mistral when the task was to understand the details of a specific paper, although she preferred Mistral for the broader literature question she tested.

Nurdana similarly found Claude particularly useful for tasks requiring critical engagement, including paper analysis, research gap exploration, and simulated peer review. She nevertheless found that bibliographic information and links could still be wrong, even when they looked plausible.


Using AI with a critical eye

So, which AI tool should a microbiologist use? These tests suggest there is still no single answer. PaperBanana showed potential for quickly visualising scientific ideas, Mistral stood out for speed and flexibility, and Claude was particularly strong when deeper analysis and critical reasoning were needed. The differences between Manasa’s and Nurdana’s experiences also show how much the result can depend on the task, the prompt, and the level of scientific detail involved.

What remained consistent across all three tools was the need for researchers to check what AI produces. A polished figure can contain the wrong biology, a useful literature summary can include an incorrect link, and a plausible research idea can still need substantial scientific reasoning before it becomes a meaningful experiment.

As these tools become more capable, that judgement becomes more important, not less. Used critically, AI can help researchers explore ideas, work through information, and save time on parts of the research process. Its value comes from knowing where the technology can help, where it can fall short, and when the scientist needs to take over.

Share this opportunity