How AI could help scientists find viruses to fight drug-resistant infections

Источник: TECHCOMMUNITY.MICROSOFT.COM

How AI could help scientists find viruses to fight drug-resistant infections

Source: TECHCOMMUNITY.MICROSOFT.COM

The post How AI could help scientists find viruses to fight drug-resistant infections appeared first on Source .

•Updated: September 30, 2026

Drug-resistant infections are a major threat to global health, driving an urgent search for treatments when antibiotics no longer work. Bacteriophages, or phages, are viruses that infect bacteria and can be repurposed against drug-resistant microbes. Although phages are highly abundant throughout the environment, identifying the right phages for treating a specific bacterial infection is difficult. Effectiveness depends on phage genetics, bacterial antiphage defenses, and infection dynamics.

We used the Microsoft Discovery app to assist scientists in developing a principled bioinformatics workflow to search for useful phages against bacterial pathogens in large environmental datasets. The pipeline needed to provide explanatory statistics using accepted methodologies so scientists could make informed decisions about which phages to test in a laboratory against a specific infection.

As a data scientist who has partnered with researchers to develop bioinformatics workflows, I have seen how much work is required to turn promising scientific questions into rigorous computational processes. The challenge is not simply identifying the right analyses, but requires connecting data, tools, validation criteria, and scientific judgment in a way that scientists can inspect, trust, and reuse.

In this exercise, our multidisciplinary team explored how a human-agent team could partner to develop a sophisticated and cloud-deployable bioinformatics pipeline for phage therapy. Drawing on expertise in genomics, bioinformatics, and cloud architectures, we worked collaboratively with agents in the Microsoft Discovery app to design the workflow, assess its recommendations, and validate its outputs at each stage. Our goal was not to claim a treatment result, but to develop a reproducible computational approach that scientists could use to guide their work.

Using the Microsoft Discovery app, we turned the pipeline development process into a research plan that scientists could inspect and refine. The app helped organize scientific literature, multi-omics data sources, analytical tools, methodological recommendations, and recommended test-cases and produced validation steps on public biological data. This process was iterative and guided by expert review.

We asked the app to help design a computational pipeline with the following requirements:

  • accept the genome of a bacterial pathogen as input
  • verify antibiotic resistance markers and produce a resistance profile
  • estimate the types of antiphage defenses that may be present
  • search environmental data for novel phages that may infect the bacteria
  • propose a mixture of one or more therapeutic phages
  • explain the specific computational steps and evidence leading to the conclusion
  • and recommend follow-on laboratory experiments

The Discovery app translated these requirements into a set of research tasks, produced inspectable methods at each research stage, and linked each step’s research outputs to the next step’s research inputs in a traceable chain. Humans and agents collaboratively reviewed phage therapy research and bioinformatics methodologies before shifting to complex code writing.

We described the scientific goal in standard scientific language, and the Discovery app converted that description into a research assistant with a specific persona (called a “purpose”), expected technical outcomes, and an inspectable “task” tree. The task tree captured the sequence of research tasks for the human-agent team, and scientists had fine-grained control over these tasks and the overall research approach.

Figure 1: The Microsoft Discovery App generates a reviewable workflow hierarchy. This task tree structure enables scientists to review and refine each step of the proposed workflow.

The task-tree view made our collaboration transparent by showing how the original problem was decomposed and which evidence and tools supported each step. We then could redirect the workflow when judgment called for a different approach, allowing us to iteratively guide the investigation rather than acting as a reviewer of its final conclusions.

Although agents can broadly search the internet for information, we focused the research process on a curated collection of articles selected by scientists. The Microsoft Discovery Bookshelf feature provides a specific place for storing, indexing, and searching trusted background material separately from the broader internet.

As the agent worked through its tasks collaboratively, we incrementally expanded its access to GitHub repositories for other tools and data sources that we wanted to incorporate. The agent developed and integrated MCP-based adapters into the workflow, expanding access to resources such as PubMed, Europe PMC, NCBI BLAST, Semantic Scholar, ClinicalTrials.gov, AlphaFold, and biomarker databases available for the investigation. This significantly simplified custom tool integration: the process took minutes and avoided what would otherwise have required substantial engineering effort.

Figure 2: Six local Model Context Protocol (MCP) servers give agents access to genomic records, scientific literature, predicted protein structures, clinical trial records, and antibiotic resistance data.

The Discovery app connected these external sources with the curated Bookshelf and task tree, enabling the workflow and agents to draw on additional knowledge and data as needed. As our team interacted with the system, Discovery’s agents refined the research approach in response to new evidence and expert feedback. For example, it proposed a revised plan combining literature search, genome analysis, and protein structure prediction, which the team reviewed and approved before the affected research tasks proceeded.

Figure 3: The Microsoft Discovery App organizes the project into tasks with defined outcomes. The first task reviews the initial curated scientific bookshelf containing 170 articles from the scientists’ collection before expanding to other knowledge and data sources.

In less than four hours, the project progressed from task-tree generation to a draft report. A significant portion of this time included sourcing genomic references and running non-trivial bioinformatics workflows created via the human-agent collaboration.

The app proposed a six-module computational architecture with defined inputs, outputs, quality and safety checks, alternative approaches, datasets for positive and negative controls of the bioinformatics workflow, a laboratory validation plan, and supporting scientific references. These artifacts provided a concrete basis for review and recommended where experimental validation would be required.

The proposed pipeline follows a simple idea: begin with a specific drug-resistant bacterium, search environmental data for phages that may attack it, narrow the list using multiple sources of scientific evidence, and send candidates prioritized for further evaluation forward for laboratory testing.

We used the API prototype the app created to validate that the proposed computational pipeline could operate as intended. Working closely with bioinformatics experts on our team, we reviewed each stage of the workflow, examined its inputs and outputs, and confirmed that information passed coherently between steps.

This expert review was essential for distinguishing scientifically grounded results from plausible-sounding but unsupported AI output and for identifying areas that required refinement before the pipeline could be trusted and reused.

Figure 4: Left: Schematic of phage therapy computational pipeline (left); Right: List of inspectable artifacts generated by pipeline at each step.

Understand the target bacterium

The workflow first reads the bacterium’s genetic blueprint. It looks for antibiotic-resistance markers, possible entry points that a phage could use, and defenses that might prevent a phage from working. This creates a profile of what a successful phage would realistically need to overcome.

Search wastewater and other environmental data

Next, the pipeline searches metagenomic data for phage DNA by querying NCBI and Europe PMC for existing phage genome records and phage-isolation and wastewater-related literature and selected metadata. The pipeline reconstructs possible phage genomes, removes incomplete or potentially unsafe options, and gathers clues about which bacteria each phage may infect.

Predict which phages are the best match

No single predictor is sufficient to show that a phage will work. The pipeline therefore combines DNA similarity, known phage–bacterium relationships, predicted fit with the bacterial surface, and the bacterium’s internal defenses to rank candidates for further review. These combined signals can support a more informed shortlist, but they do not establish effectiveness without expert assessment and laboratory testing.

Rank candidates for effectiveness and safety

Each candidate is ranked by its predicted ability to infect and kill the target, its safety, and its practicality for production. The ranking also considers whether phage escape through changes to bacterial receptors could restore antibiotic sensitivity. Temperate or lysogenic candidates are excluded because integration can cause lysogenic conversion and horizontal gene transfer.

Consider engineering only when needed

If no natural phage meets the requirements, the workflow can propose carefully limited changes to help a well-understood phage recognize the target bacterium. Any engineered option must pass strict checks for safety, stability, and unintended effects before it moves forward.

Confirm the results in the laboratory

Computational predictions are only the beginning. A scientists would confirm via wet lab experiments that the phage can infect and kill the target bacterium, measure how quickly resistance develops, test whether it works well with antibiotics, and evaluate safety, stability, and production. Those results would then be used as feedback to improve future predictions.

Together, these stages concentrate detailed analysis and laboratory resources on the strongest candidates while maintaining safety and expert review throughout.

By organizing evidence, tools, and computational steps into a reproducible and reviewable workflow, the Microsoft Discovery app helped our team move from a broad scientific question to an implementation-ready research pipeline.

The outcome was not a validated therapeutic result. It was a structured plan, a working cloud-based implementation, and a clearer understanding of the evidence gaps that still require expert review and laboratory validation.

My prior work with bioinformatics researchers has shown me that translating scientific questions into trustworthy computational workflows requires both domain expertise and significant engineering effort. In this project, the Discovery app amplified that expertise rather than replace it, helping our team to build a deterministic, reusable pipeline while retaining human judgment at every consequential decision point.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.