Poster ID
P-18
Poster Title
Conversational Data Discovery: Developing AI-assisted workflows for responsible and trusted science
Authors
Mitchell Shiell¹, Jill Mesirov²,*, Helga Thorvaldsdottir³, Michael Reich², Jon Eubank¹, Justin Richardsson¹, Rakesh Mistry¹, Patrick Dos Santos¹, Robin Haw¹, Melanie Courtot¹,4,5, MSigDB Team² and the Overture Team¹
¹Ontario Institute for Cancer Research, Toronto, Canada. ²Sanford Burnham Prebys Medical Discovery Institute. ³Broad Institute, Cambridge, MA. 4University of Toronto Department of Medical Biophysics, Toronto, Canada. 5University of Toronto Department of Computer Science, Toronto, Canada.
¹Ontario Institute for Cancer Research, Toronto, Canada. ²Sanford Burnham Prebys Medical Discovery Institute. ³Broad Institute, Cambridge, MA. 4University of Toronto Department of Medical Biophysics, Toronto, Canada. 5University of Toronto Department of Computer Science, Toronto, Canada.
Abstract
Overture-powered research portals deliver a data exploration page with a faceted search interface that lets researchers connect and collaborate over their data. However, over time we have identified two core issues. For researchers: they cannot query across distributed deployments through the same interface, and in-portal analysis is non-existent, forcing manual export to separate tools. For developers: extending each portal with new tools, visualizations, and data sources creates an "integration trap," where every new connection requires costly bespoke development, yields low reuse, and adds maintenance effort that scales with each addition.
Three AI infrastructure advances offer a solution: (1) self-hosted LLMs keep sensitive data on-premises; (2) the Model Context Protocol (MCP) lets each resource speak one standard protocol; and (3) conversational hosts give researchers a flexible canvas. Unlike the fixed views of a purpose-built UI, conversational interfaces paired with LLMs can render whatever a query calls for, including tables, plots, summaries, or generated code. For development, reuse is high and integration effort stays flat as MCP servers are added.
We are building an Overture MCP server that exposes our data for conversational discovery. Targeting modestly sized open-weight models (~4–70B parameters, Q4-quantized), the server uses introspection endpoints that make the API self-describing, plus a compact tool set spanning the full query lifecycle: the model identifies the relevant dataset, loads its field definitions, classifies the request (answerable, ambiguous, unanswerable, or improper), then confirms a plain-language query before execution.
We will demonstrate this with OICR's Drug Discovery Portal (405M records, 20K genes, 32 cancer types across mutation, expression, correlation, and protein datasets), collaborating with the Mesirov Lab to run our MCP server alongside their MSigDB MCP server, which brings in 50,000+ curated human and mouse gene sets. Together, these servers let a researcher ask "which genes are frequently mutated in colorectal cancer?", get live results from Overture's search API, then refine with "what pathways are these genes involved in?", pulling results from MSigDB, all without switching tools or losing context. The payoff is simple but powerful: once data is in the conversation, it becomes a substrate for diverse downstream uses, including analysis, visualization, annotation, and reproducible code artifacts.
Three AI infrastructure advances offer a solution: (1) self-hosted LLMs keep sensitive data on-premises; (2) the Model Context Protocol (MCP) lets each resource speak one standard protocol; and (3) conversational hosts give researchers a flexible canvas. Unlike the fixed views of a purpose-built UI, conversational interfaces paired with LLMs can render whatever a query calls for, including tables, plots, summaries, or generated code. For development, reuse is high and integration effort stays flat as MCP servers are added.
We are building an Overture MCP server that exposes our data for conversational discovery. Targeting modestly sized open-weight models (~4–70B parameters, Q4-quantized), the server uses introspection endpoints that make the API self-describing, plus a compact tool set spanning the full query lifecycle: the model identifies the relevant dataset, loads its field definitions, classifies the request (answerable, ambiguous, unanswerable, or improper), then confirms a plain-language query before execution.
We will demonstrate this with OICR's Drug Discovery Portal (405M records, 20K genes, 32 cancer types across mutation, expression, correlation, and protein datasets), collaborating with the Mesirov Lab to run our MCP server alongside their MSigDB MCP server, which brings in 50,000+ curated human and mouse gene sets. Together, these servers let a researcher ask "which genes are frequently mutated in colorectal cancer?", get live results from Overture's search API, then refine with "what pathways are these genes involved in?", pulling results from MSigDB, all without switching tools or losing context. The payoff is simple but powerful: once data is in the conversation, it becomes a substrate for diverse downstream uses, including analysis, visualization, annotation, and reproducible code artifacts.
Digital Poster