The application behind a live market opportunity assessment engine

From Market Signal to Portfolio Decision

In this article, we map out the system behind a live opportunity assessment engine, and how it learns which signals matter to a decision and which are just noise.

Read more
filters
Market & Competitive Intelligence
The Context
The hard part is no longer getting the data

An early-pipeline opportunity assessment used to occupy several analysts for weeks. As our companion piece describes, agentic AI can now do most of that gathering in minutes, reading across trial registries, regulatory records, the genetic databases, and the published literature, and keeping the picture current as the market moves. The retrieval problem is, for the most part, behind us.

What comes next is harder. Once a system can pull almost any signal on demand, the work becomes sorting what matters from everything else. That is an easy thing to underestimate, and it is where most of the difficulty now lives.

The Context
The hard part is no longer getting the data

An early-pipeline opportunity assessment used to occupy several analysts for weeks. As our companion piece describes, agentic AI can now do most of that gathering in minutes, reading across trial registries, regulatory records, the genetic databases, and the published literature, and keeping the picture current as the market moves. The retrieval problem is, for the most part, behind us.

What comes next is harder. Once a system can pull almost any signal on demand, the work becomes sorting what matters from everything else. That is an easy thing to underestimate, and it is where most of the difficulty now lives.

The Context
The hard part is no longer getting the data

An early-pipeline opportunity assessment used to occupy several analysts for weeks. As our companion piece describes, agentic AI can now do most of that gathering in minutes, reading across trial registries, regulatory records, the genetic databases, and the published literature, and keeping the picture current as the market moves. The retrieval problem is, for the most part, behind us.

What comes next is harder. Once a system can pull almost any signal on demand, the work becomes sorting what matters from everything else. That is an easy thing to underestimate, and it is where most of the difficulty now lives.

The Context

The hard part is no longer getting the data

An early-pipeline opportunity assessment used to occupy several analysts for weeks. As our companion piece describes, agentic AI can now do most of that gathering in minutes, reading across trial registries, regulatory records, the genetic databases, and the published literature, and keeping the picture current as the market moves. The retrieval problem is, for the most part, behind us.

What comes next is harder. Once a system can pull almost any signal on demand, the work becomes sorting what matters from everything else. That is an easy thing to underestimate, and it is where most of the difficulty now lives.

The Context

The hard part is no longer getting the data

An early-pipeline opportunity assessment used to occupy several analysts for weeks. As our companion piece describes, agentic AI can now do most of that gathering in minutes, reading across trial registries, regulatory records, the genetic databases, and the published literature, and keeping the picture current as the market moves. The retrieval problem is, for the most part, behind us.

What comes next is harder. Once a system can pull almost any signal on demand, the work becomes sorting what matters from everything else. That is an easy thing to underestimate, and it is where most of the difficulty now lives.

The hard part is no longer getting the data

The Context
The Example
A running example: early Alzheimer's

A single example will run through this article. A neurology team has narrowed twenty candidate indications to three, and early Alzheimer's is one of them. The team is now focusing on monitoring competitive intensity. One morning, a competitor posts a Phase 3 readout for an anti-amyloid asset in early Alzheimer's. The same morning brings a Phase 1 study opening in an indication the team set aside weeks ago, and a handful of unrelated competitor headlines. Only one of those signals should reach the people making the decision, and knowing which one is the whole problem.

The Example
A running example: early Alzheimer's

A single example will run through this article. A neurology team has narrowed twenty candidate indications to three, and early Alzheimer's is one of them. The team is now focusing on monitoring competitive intensity. One morning, a competitor posts a Phase 3 readout for an anti-amyloid asset in early Alzheimer's. The same morning brings a Phase 1 study opening in an indication the team set aside weeks ago, and a handful of unrelated competitor headlines. Only one of those signals should reach the people making the decision, and knowing which one is the whole problem.

The Example
A running example: early Alzheimer's

A single example will run through this article. A neurology team has narrowed twenty candidate indications to three, and early Alzheimer's is one of them. The team is now focusing on monitoring competitive intensity. One morning, a competitor posts a Phase 3 readout for an anti-amyloid asset in early Alzheimer's. The same morning brings a Phase 1 study opening in an indication the team set aside weeks ago, and a handful of unrelated competitor headlines. Only one of those signals should reach the people making the decision, and knowing which one is the whole problem.

The Example

A running example: early Alzheimer's

A single example will run through this article. A neurology team has narrowed twenty candidate indications to three, and early Alzheimer's is one of them. The team is now focusing on monitoring competitive intensity. One morning, a competitor posts a Phase 3 readout for an anti-amyloid asset in early Alzheimer's. The same morning brings a Phase 1 study opening in an indication the team set aside weeks ago, and a handful of unrelated competitor headlines. Only one of those signals should reach the people making the decision, and knowing which one is the whole problem.

The Example

A running example: early Alzheimer's

A single example will run through this article. A neurology team has narrowed twenty candidate indications to three, and early Alzheimer's is one of them. The team is now focusing on monitoring competitive intensity. One morning, a competitor posts a Phase 3 readout for an anti-amyloid asset in early Alzheimer's. The same morning brings a Phase 1 study opening in an indication the team set aside weeks ago, and a handful of unrelated competitor headlines. Only one of those signals should reach the people making the decision, and knowing which one is the whole problem.

A running example: early Alzheimer's

The Example
Why Now
The stack is conventional; trust and calibration are the work

An organization can have all the signals in the world and be no better off for it, unless the system knows which signals could change a decision. If a competitor is one of the large players, a system that surfaces ten items a day about them is easy to build and of very little use. It becomes a second inbox, and people soon learn to stop reading it.

The reason is that a signal has no value on its own. Its value depends on the decision in front of you. The Phase 3 readout matters because early Alzheimer's is a live bet, while the Phase 1 study in a discarded space does not. Other fields know this pattern well — security teams call it alert fatigue, clinical teams alarm fatigue — and the lesson is always the same: a system that surfaces everything teaches people to ignore it.

This is also a good moment to build it, and not because the idea itself is new. Through 2026, major data platforms, such as Snowflake and Databricks, rebuilt their messaging around exactly this problem, treating agentic AI as a question of context and governance rather than raw model capability. The more practical point is that the components the build depends on have only become production-grade in the space of roughly a year.

A year ago, a team would have had to hand-build almost all of it: the governed metric definitions, the connectors into each source, the guardrails, and the harness to check the system was behaving. Today, the semantic layer is configurable within major platform technologies, and its definitions are becoming portable. The integration activities of connecting agents to registries and internal systems have largely standardized around common protocols, runtime governance and evaluation are increasingly built in, and the cost of running the models continuously has fallen sharply. The models have also only recently become reliable enough at reading and synthesizing unstructured sources, such as the literature and HTA records, to be trusted with the part of the work that was hardest to automate. The plumbing that once consumed a build now comes largely off the shelf, which leaves the calibration and judgment described below as the place where the real effort goes.

Why Now
The stack is conventional; trust and calibration are the work

An organization can have all the signals in the world and be no better off for it, unless the system knows which signals could change a decision. If a competitor is one of the large players, a system that surfaces ten items a day about them is easy to build and of very little use. It becomes a second inbox, and people soon learn to stop reading it.

The reason is that a signal has no value on its own. Its value depends on the decision in front of you. The Phase 3 readout matters because early Alzheimer's is a live bet, while the Phase 1 study in a discarded space does not. Other fields know this pattern well — security teams call it alert fatigue, clinical teams alarm fatigue — and the lesson is always the same: a system that surfaces everything teaches people to ignore it.

This is also a good moment to build it, and not because the idea itself is new. Through 2026, major data platforms, such as Snowflake and Databricks, rebuilt their messaging around exactly this problem, treating agentic AI as a question of context and governance rather than raw model capability. The more practical point is that the components the build depends on have only become production-grade in the space of roughly a year.

A year ago, a team would have had to hand-build almost all of it: the governed metric definitions, the connectors into each source, the guardrails, and the harness to check the system was behaving. Today, the semantic layer is configurable within major platform technologies, and its definitions are becoming portable. The integration activities of connecting agents to registries and internal systems have largely standardized around common protocols, runtime governance and evaluation are increasingly built in, and the cost of running the models continuously has fallen sharply. The models have also only recently become reliable enough at reading and synthesizing unstructured sources, such as the literature and HTA records, to be trusted with the part of the work that was hardest to automate. The plumbing that once consumed a build now comes largely off the shelf, which leaves the calibration and judgment described below as the place where the real effort goes.

Why Now
The stack is conventional; trust and calibration are the work

An organization can have all the signals in the world and be no better off for it, unless the system knows which signals could change a decision. If a competitor is one of the large players, a system that surfaces ten items a day about them is easy to build and of very little use. It becomes a second inbox, and people soon learn to stop reading it.

The reason is that a signal has no value on its own. Its value depends on the decision in front of you. The Phase 3 readout matters because early Alzheimer's is a live bet, while the Phase 1 study in a discarded space does not. Other fields know this pattern well — security teams call it alert fatigue, clinical teams alarm fatigue — and the lesson is always the same: a system that surfaces everything teaches people to ignore it.

This is also a good moment to build it, and not because the idea itself is new. Through 2026, major data platforms, such as Snowflake and Databricks, rebuilt their messaging around exactly this problem, treating agentic AI as a question of context and governance rather than raw model capability. The more practical point is that the components the build depends on have only become production-grade in the space of roughly a year.

A year ago, a team would have had to hand-build almost all of it: the governed metric definitions, the connectors into each source, the guardrails, and the harness to check the system was behaving. Today, the semantic layer is configurable within major platform technologies, and its definitions are becoming portable. The integration activities of connecting agents to registries and internal systems have largely standardized around common protocols, runtime governance and evaluation are increasingly built in, and the cost of running the models continuously has fallen sharply. The models have also only recently become reliable enough at reading and synthesizing unstructured sources, such as the literature and HTA records, to be trusted with the part of the work that was hardest to automate. The plumbing that once consumed a build now comes largely off the shelf, which leaves the calibration and judgment described below as the place where the real effort goes.

Why Now

The stack is conventional; trust and calibration are the work

An organization can have all the signals in the world and be no better off for it, unless the system knows which signals could change a decision. If a competitor is one of the large players, a system that surfaces ten items a day about them is easy to build and of very little use. It becomes a second inbox, and people soon learn to stop reading it.

The reason is that a signal has no value on its own. Its value depends on the decision in front of you. The Phase 3 readout matters because early Alzheimer's is a live bet, while the Phase 1 study in a discarded space does not. Other fields know this pattern well — security teams call it alert fatigue, clinical teams alarm fatigue — and the lesson is always the same: a system that surfaces everything teaches people to ignore it.

This is also a good moment to build it, and not because the idea itself is new. Through 2026, major data platforms, such as Snowflake and Databricks, rebuilt their messaging around exactly this problem, treating agentic AI as a question of context and governance rather than raw model capability. The more practical point is that the components the build depends on have only become production-grade in the space of roughly a year.

A year ago, a team would have had to hand-build almost all of it: the governed metric definitions, the connectors into each source, the guardrails, and the harness to check the system was behaving. Today, the semantic layer is configurable within major platform technologies, and its definitions are becoming portable. The integration activities of connecting agents to registries and internal systems have largely standardized around common protocols, runtime governance and evaluation are increasingly built in, and the cost of running the models continuously has fallen sharply. The models have also only recently become reliable enough at reading and synthesizing unstructured sources, such as the literature and HTA records, to be trusted with the part of the work that was hardest to automate. The plumbing that once consumed a build now comes largely off the shelf, which leaves the calibration and judgment described below as the place where the real effort goes.

Why Now

The stack is conventional; trust and calibration are the work

An organization can have all the signals in the world and be no better off for it, unless the system knows which signals could change a decision. If a competitor is one of the large players, a system that surfaces ten items a day about them is easy to build and of very little use. It becomes a second inbox, and people soon learn to stop reading it.

The reason is that a signal has no value on its own. Its value depends on the decision in front of you. The Phase 3 readout matters because early Alzheimer's is a live bet, while the Phase 1 study in a discarded space does not. Other fields know this pattern well — security teams call it alert fatigue, clinical teams alarm fatigue — and the lesson is always the same: a system that surfaces everything teaches people to ignore it.

This is also a good moment to build it, and not because the idea itself is new. Through 2026, major data platforms, such as Snowflake and Databricks, rebuilt their messaging around exactly this problem, treating agentic AI as a question of context and governance rather than raw model capability. The more practical point is that the components the build depends on have only become production-grade in the space of roughly a year.

A year ago, a team would have had to hand-build almost all of it: the governed metric definitions, the connectors into each source, the guardrails, and the harness to check the system was behaving. Today, the semantic layer is configurable within major platform technologies, and its definitions are becoming portable. The integration activities of connecting agents to registries and internal systems have largely standardized around common protocols, runtime governance and evaluation are increasingly built in, and the cost of running the models continuously has fallen sharply. The models have also only recently become reliable enough at reading and synthesizing unstructured sources, such as the literature and HTA records, to be trusted with the part of the work that was hardest to automate. The plumbing that once consumed a build now comes largely off the shelf, which leaves the calibration and judgment described below as the place where the real effort goes.

Why Now

The stack is conventional; trust and calibration are the work

Why Now
The Reframe
Tuning the system to the decision

The aim is to tune the system to the decision at hand. Instead of reacting to everything it could pick up, it should react to what would move a metric across a threshold on a decision that is currently open, and stay quiet on the rest. A broad indication scan and a late-stage business case have very different tolerances for noise, so how sensitive the system is has to change as a decision matures.

The Reframe
Tuning the system to the decision

The aim is to tune the system to the decision at hand. Instead of reacting to everything it could pick up, it should react to what would move a metric across a threshold on a decision that is currently open, and stay quiet on the rest. A broad indication scan and a late-stage business case have very different tolerances for noise, so how sensitive the system is has to change as a decision matures.

The Reframe
Tuning the system to the decision

The aim is to tune the system to the decision at hand. Instead of reacting to everything it could pick up, it should react to what would move a metric across a threshold on a decision that is currently open, and stay quiet on the rest. A broad indication scan and a late-stage business case have very different tolerances for noise, so how sensitive the system is has to change as a decision matures.

The Reframe

Tuning the system to the decision

The aim is to tune the system to the decision at hand. Instead of reacting to everything it could pick up, it should react to what would move a metric across a threshold on a decision that is currently open, and stay quiet on the rest. A broad indication scan and a late-stage business case have very different tolerances for noise, so how sensitive the system is has to change as a decision matures.

The Reframe

Tuning the system to the decision

The aim is to tune the system to the decision at hand. Instead of reacting to everything it could pick up, it should react to what would move a metric across a threshold on a decision that is currently open, and stay quiet on the rest. A broad indication scan and a late-stage business case have very different tolerances for noise, so how sensitive the system is has to change as a decision matures.

Tuning the system to the decision

The Reframe
The Trust Question
A conventional stack, with one demanding question

The underlying system is a conventional data stack: sources feed ingestion agents, signals land in a governed store, a transformation layer scores them, and a consumption layer puts them in front of a person. The architecture is not what makes this hard. The real question is why anyone should trust a number that came out of a probabilistic model.

The first concern is usually reproducibility: if a model can answer the same question two ways, how can a portfolio decision rest on it? This is what the semantic layer is for. It holds governed definitions of each metric – competitive intensity, say, defined as approved assets plus assets in Phase 3 from named sources — so that however a question is phrased, it resolves to the same result. This is no longer bespoke work; it's now included within common data platforms, from Snowflake's Cortex Analyst to Databricks' Genie Ontology to dbt's Semantic Layer. The definitions are even becoming portable between tools through Apache Ossie (formerly known as the Open Semantic Interchange), a cross-vendor standard whose first specification appeared in early 2026. A financial-services working group has already formed under it, which suggests a life-sciences equivalent is not far off.

The less structured work of synthesizing a mechanism's evidence base from the literature, or reading why a body such as NICE rejected a comparator, is where models genuinely extend what a team can do under time pressure. Reproducibility there comes from asking the model to return a fixed structure rather than free text, and from grounding every claim in a citation. The reasoning stays probabilistic, but the output can be checked. That kind of reproducibility is ultimately what allows a team to believe that a change from one month to the next reflects the market rather than the model.

Once the numbers can be trusted, the next question is which of them are worth anyone's attention. That is what calibration handles, and in our experience it is the part teams most often struggle with. Calibration is the logic that decides which signals to look for and which to put in front of someone. The temptation is to write that logic into the agents' prompts, but prompts drift over time, and control logic buried inside them brings back the unpredictability the semantic layer was meant to remove.

We prefer to treat calibration as policy: a single versioned configuration, closer to settings than to code, that the system reads both at ingestion, to decide what to pick up, and at consumption, to decide what is worth surfacing. Because once policy governs both, a single change carries through consistently. Set the threshold for competitive events in early Alzheimer's to Phase 3 and above, and the system adjusts what it detects and what it shows together. The platforms are converging on the same idea; Databricks' Unity AI Gateway now enforces this kind of policy at the moment an agent acts, rather than only at design time.

The people using the system never see the configuration. They work with a settings view: which indications are live, which competitors to watch, what counts as material, and whether a crossed threshold should be logged, quietly re-scored, or raised as an alert. Each choice writes a line of a versioned, attributed policy, so every change to sensitivity carries an owner and a date. That is also what keeps the analysis auditable.

The Trust Question
A conventional stack, with one demanding question

The underlying system is a conventional data stack: sources feed ingestion agents, signals land in a governed store, a transformation layer scores them, and a consumption layer puts them in front of a person. The architecture is not what makes this hard. The real question is why anyone should trust a number that came out of a probabilistic model.

The first concern is usually reproducibility: if a model can answer the same question two ways, how can a portfolio decision rest on it? This is what the semantic layer is for. It holds governed definitions of each metric – competitive intensity, say, defined as approved assets plus assets in Phase 3 from named sources — so that however a question is phrased, it resolves to the same result. This is no longer bespoke work; it's now included within common data platforms, from Snowflake's Cortex Analyst to Databricks' Genie Ontology to dbt's Semantic Layer. The definitions are even becoming portable between tools through Apache Ossie (formerly known as the Open Semantic Interchange), a cross-vendor standard whose first specification appeared in early 2026. A financial-services working group has already formed under it, which suggests a life-sciences equivalent is not far off.

The less structured work of synthesizing a mechanism's evidence base from the literature, or reading why a body such as NICE rejected a comparator, is where models genuinely extend what a team can do under time pressure. Reproducibility there comes from asking the model to return a fixed structure rather than free text, and from grounding every claim in a citation. The reasoning stays probabilistic, but the output can be checked. That kind of reproducibility is ultimately what allows a team to believe that a change from one month to the next reflects the market rather than the model.

Once the numbers can be trusted, the next question is which of them are worth anyone's attention. That is what calibration handles, and in our experience it is the part teams most often struggle with. Calibration is the logic that decides which signals to look for and which to put in front of someone. The temptation is to write that logic into the agents' prompts, but prompts drift over time, and control logic buried inside them brings back the unpredictability the semantic layer was meant to remove.

We prefer to treat calibration as policy: a single versioned configuration, closer to settings than to code, that the system reads both at ingestion, to decide what to pick up, and at consumption, to decide what is worth surfacing. Because once policy governs both, a single change carries through consistently. Set the threshold for competitive events in early Alzheimer's to Phase 3 and above, and the system adjusts what it detects and what it shows together. The platforms are converging on the same idea; Databricks' Unity AI Gateway now enforces this kind of policy at the moment an agent acts, rather than only at design time.

The people using the system never see the configuration. They work with a settings view: which indications are live, which competitors to watch, what counts as material, and whether a crossed threshold should be logged, quietly re-scored, or raised as an alert. Each choice writes a line of a versioned, attributed policy, so every change to sensitivity carries an owner and a date. That is also what keeps the analysis auditable.

The Trust Question
A conventional stack, with one demanding question

The underlying system is a conventional data stack: sources feed ingestion agents, signals land in a governed store, a transformation layer scores them, and a consumption layer puts them in front of a person. The architecture is not what makes this hard. The real question is why anyone should trust a number that came out of a probabilistic model.

The first concern is usually reproducibility: if a model can answer the same question two ways, how can a portfolio decision rest on it? This is what the semantic layer is for. It holds governed definitions of each metric – competitive intensity, say, defined as approved assets plus assets in Phase 3 from named sources — so that however a question is phrased, it resolves to the same result. This is no longer bespoke work; it's now included within common data platforms, from Snowflake's Cortex Analyst to Databricks' Genie Ontology to dbt's Semantic Layer. The definitions are even becoming portable between tools through Apache Ossie (formerly known as the Open Semantic Interchange), a cross-vendor standard whose first specification appeared in early 2026. A financial-services working group has already formed under it, which suggests a life-sciences equivalent is not far off.

The less structured work of synthesizing a mechanism's evidence base from the literature, or reading why a body such as NICE rejected a comparator, is where models genuinely extend what a team can do under time pressure. Reproducibility there comes from asking the model to return a fixed structure rather than free text, and from grounding every claim in a citation. The reasoning stays probabilistic, but the output can be checked. That kind of reproducibility is ultimately what allows a team to believe that a change from one month to the next reflects the market rather than the model.

Once the numbers can be trusted, the next question is which of them are worth anyone's attention. That is what calibration handles, and in our experience it is the part teams most often struggle with. Calibration is the logic that decides which signals to look for and which to put in front of someone. The temptation is to write that logic into the agents' prompts, but prompts drift over time, and control logic buried inside them brings back the unpredictability the semantic layer was meant to remove.

We prefer to treat calibration as policy: a single versioned configuration, closer to settings than to code, that the system reads both at ingestion, to decide what to pick up, and at consumption, to decide what is worth surfacing. Because once policy governs both, a single change carries through consistently. Set the threshold for competitive events in early Alzheimer's to Phase 3 and above, and the system adjusts what it detects and what it shows together. The platforms are converging on the same idea; Databricks' Unity AI Gateway now enforces this kind of policy at the moment an agent acts, rather than only at design time.

The people using the system never see the configuration. They work with a settings view: which indications are live, which competitors to watch, what counts as material, and whether a crossed threshold should be logged, quietly re-scored, or raised as an alert. Each choice writes a line of a versioned, attributed policy, so every change to sensitivity carries an owner and a date. That is also what keeps the analysis auditable.

The Trust Question

A conventional stack, with one demanding question

The underlying system is a conventional data stack: sources feed ingestion agents, signals land in a governed store, a transformation layer scores them, and a consumption layer puts them in front of a person. The architecture is not what makes this hard. The real question is why anyone should trust a number that came out of a probabilistic model.

The first concern is usually reproducibility: if a model can answer the same question two ways, how can a portfolio decision rest on it? This is what the semantic layer is for. It holds governed definitions of each metric – competitive intensity, say, defined as approved assets plus assets in Phase 3 from named sources — so that however a question is phrased, it resolves to the same result. This is no longer bespoke work; it's now included within common data platforms, from Snowflake's Cortex Analyst to Databricks' Genie Ontology to dbt's Semantic Layer. The definitions are even becoming portable between tools through Apache Ossie (formerly known as the Open Semantic Interchange), a cross-vendor standard whose first specification appeared in early 2026. A financial-services working group has already formed under it, which suggests a life-sciences equivalent is not far off.

The less structured work of synthesizing a mechanism's evidence base from the literature, or reading why a body such as NICE rejected a comparator, is where models genuinely extend what a team can do under time pressure. Reproducibility there comes from asking the model to return a fixed structure rather than free text, and from grounding every claim in a citation. The reasoning stays probabilistic, but the output can be checked. That kind of reproducibility is ultimately what allows a team to believe that a change from one month to the next reflects the market rather than the model.

Once the numbers can be trusted, the next question is which of them are worth anyone's attention. That is what calibration handles, and in our experience it is the part teams most often struggle with. Calibration is the logic that decides which signals to look for and which to put in front of someone. The temptation is to write that logic into the agents' prompts, but prompts drift over time, and control logic buried inside them brings back the unpredictability the semantic layer was meant to remove.

We prefer to treat calibration as policy: a single versioned configuration, closer to settings than to code, that the system reads both at ingestion, to decide what to pick up, and at consumption, to decide what is worth surfacing. Because once policy governs both, a single change carries through consistently. Set the threshold for competitive events in early Alzheimer's to Phase 3 and above, and the system adjusts what it detects and what it shows together. The platforms are converging on the same idea; Databricks' Unity AI Gateway now enforces this kind of policy at the moment an agent acts, rather than only at design time.

The people using the system never see the configuration. They work with a settings view: which indications are live, which competitors to watch, what counts as material, and whether a crossed threshold should be logged, quietly re-scored, or raised as an alert. Each choice writes a line of a versioned, attributed policy, so every change to sensitivity carries an owner and a date. That is also what keeps the analysis auditable.

The Trust Question

A conventional stack, with one demanding question

The underlying system is a conventional data stack: sources feed ingestion agents, signals land in a governed store, a transformation layer scores them, and a consumption layer puts them in front of a person. The architecture is not what makes this hard. The real question is why anyone should trust a number that came out of a probabilistic model.

The first concern is usually reproducibility: if a model can answer the same question two ways, how can a portfolio decision rest on it? This is what the semantic layer is for. It holds governed definitions of each metric – competitive intensity, say, defined as approved assets plus assets in Phase 3 from named sources — so that however a question is phrased, it resolves to the same result. This is no longer bespoke work; it's now included within common data platforms, from Snowflake's Cortex Analyst to Databricks' Genie Ontology to dbt's Semantic Layer. The definitions are even becoming portable between tools through Apache Ossie (formerly known as the Open Semantic Interchange), a cross-vendor standard whose first specification appeared in early 2026. A financial-services working group has already formed under it, which suggests a life-sciences equivalent is not far off.

The less structured work of synthesizing a mechanism's evidence base from the literature, or reading why a body such as NICE rejected a comparator, is where models genuinely extend what a team can do under time pressure. Reproducibility there comes from asking the model to return a fixed structure rather than free text, and from grounding every claim in a citation. The reasoning stays probabilistic, but the output can be checked. That kind of reproducibility is ultimately what allows a team to believe that a change from one month to the next reflects the market rather than the model.

Once the numbers can be trusted, the next question is which of them are worth anyone's attention. That is what calibration handles, and in our experience it is the part teams most often struggle with. Calibration is the logic that decides which signals to look for and which to put in front of someone. The temptation is to write that logic into the agents' prompts, but prompts drift over time, and control logic buried inside them brings back the unpredictability the semantic layer was meant to remove.

We prefer to treat calibration as policy: a single versioned configuration, closer to settings than to code, that the system reads both at ingestion, to decide what to pick up, and at consumption, to decide what is worth surfacing. Because once policy governs both, a single change carries through consistently. Set the threshold for competitive events in early Alzheimer's to Phase 3 and above, and the system adjusts what it detects and what it shows together. The platforms are converging on the same idea; Databricks' Unity AI Gateway now enforces this kind of policy at the moment an agent acts, rather than only at design time.

The people using the system never see the configuration. They work with a settings view: which indications are live, which competitors to watch, what counts as material, and whether a crossed threshold should be logged, quietly re-scored, or raised as an alert. Each choice writes a line of a versioned, attributed policy, so every change to sensitivity carries an owner and a date. That is also what keeps the analysis auditable.

A conventional stack, with one demanding question

The Trust Question
Back to the Example
A Tuesday in neurology

With those pieces in place, the morning resolves cleanly. At ingestion, the policy already knows early Alzheimer's is live and watched at Phase 3 and above, so the anti-amyloid readout is picked up while the Phase 1 study is not. The semantic layer resolves the readout to the governed Phase 3 count, the model confirms whether the endpoint was met, and a citation is attached. Because the competitor's asset is meaningfully differentiated, the new signal has changed a critical score in the competitive intensity dimension, going from medium to high. This fundamentally changes the asset's evaluation, and the team needs to react. Because of the change in score, the asset moves within the framework and the portfolio manager receives an alert.

What happens next is deliberately not automatic. The manager reviews the change, sees the evidence behind it, and decides whether to accept it and publish the updated assessment or to hold it. Scores do not rewrite themselves. Every update to the living view passes through a person. This matters as much for trust as for control, since an assessment that re-scored itself silently overnight would be neither stable enough to plan against nor easy to audit afterwards. The competitor headlines from the same morning, by contrast, never reach the manager: none touches a live decision or clears a threshold, so they are logged and left there.

Back to the Example
A Tuesday in neurology

With those pieces in place, the morning resolves cleanly. At ingestion, the policy already knows early Alzheimer's is live and watched at Phase 3 and above, so the anti-amyloid readout is picked up while the Phase 1 study is not. The semantic layer resolves the readout to the governed Phase 3 count, the model confirms whether the endpoint was met, and a citation is attached. Because the competitor's asset is meaningfully differentiated, the new signal has changed a critical score in the competitive intensity dimension, going from medium to high. This fundamentally changes the asset's evaluation, and the team needs to react. Because of the change in score, the asset moves within the framework and the portfolio manager receives an alert.

What happens next is deliberately not automatic. The manager reviews the change, sees the evidence behind it, and decides whether to accept it and publish the updated assessment or to hold it. Scores do not rewrite themselves. Every update to the living view passes through a person. This matters as much for trust as for control, since an assessment that re-scored itself silently overnight would be neither stable enough to plan against nor easy to audit afterwards. The competitor headlines from the same morning, by contrast, never reach the manager: none touches a live decision or clears a threshold, so they are logged and left there.

Back to the Example
A Tuesday in neurology

With those pieces in place, the morning resolves cleanly. At ingestion, the policy already knows early Alzheimer's is live and watched at Phase 3 and above, so the anti-amyloid readout is picked up while the Phase 1 study is not. The semantic layer resolves the readout to the governed Phase 3 count, the model confirms whether the endpoint was met, and a citation is attached. Because the competitor's asset is meaningfully differentiated, the new signal has changed a critical score in the competitive intensity dimension, going from medium to high. This fundamentally changes the asset's evaluation, and the team needs to react. Because of the change in score, the asset moves within the framework and the portfolio manager receives an alert.

What happens next is deliberately not automatic. The manager reviews the change, sees the evidence behind it, and decides whether to accept it and publish the updated assessment or to hold it. Scores do not rewrite themselves. Every update to the living view passes through a person. This matters as much for trust as for control, since an assessment that re-scored itself silently overnight would be neither stable enough to plan against nor easy to audit afterwards. The competitor headlines from the same morning, by contrast, never reach the manager: none touches a live decision or clears a threshold, so they are logged and left there.

Back to the Example

A Tuesday in neurology

With those pieces in place, the morning resolves cleanly. At ingestion, the policy already knows early Alzheimer's is live and watched at Phase 3 and above, so the anti-amyloid readout is picked up while the Phase 1 study is not. The semantic layer resolves the readout to the governed Phase 3 count, the model confirms whether the endpoint was met, and a citation is attached. Because the competitor's asset is meaningfully differentiated, the new signal has changed a critical score in the competitive intensity dimension, going from medium to high. This fundamentally changes the asset's evaluation, and the team needs to react. Because of the change in score, the asset moves within the framework and the portfolio manager receives an alert.

What happens next is deliberately not automatic. The manager reviews the change, sees the evidence behind it, and decides whether to accept it and publish the updated assessment or to hold it. Scores do not rewrite themselves. Every update to the living view passes through a person. This matters as much for trust as for control, since an assessment that re-scored itself silently overnight would be neither stable enough to plan against nor easy to audit afterwards. The competitor headlines from the same morning, by contrast, never reach the manager: none touches a live decision or clears a threshold, so they are logged and left there.

Back to the Example

A Tuesday in neurology

With those pieces in place, the morning resolves cleanly. At ingestion, the policy already knows early Alzheimer's is live and watched at Phase 3 and above, so the anti-amyloid readout is picked up while the Phase 1 study is not. The semantic layer resolves the readout to the governed Phase 3 count, the model confirms whether the endpoint was met, and a citation is attached. Because the competitor's asset is meaningfully differentiated, the new signal has changed a critical score in the competitive intensity dimension, going from medium to high. This fundamentally changes the asset's evaluation, and the team needs to react. Because of the change in score, the asset moves within the framework and the portfolio manager receives an alert.

What happens next is deliberately not automatic. The manager reviews the change, sees the evidence behind it, and decides whether to accept it and publish the updated assessment or to hold it. Scores do not rewrite themselves. Every update to the living view passes through a person. This matters as much for trust as for control, since an assessment that re-scored itself silently overnight would be neither stable enough to plan against nor easy to audit afterwards. The competitor headlines from the same morning, by contrast, never reach the manager: none touches a live decision or clears a threshold, so they are logged and left there.

A Tuesday in neurology

Back to the Example
The Human Checkpoint
The loop narrows attention and concentrates human judgment

A decision, once made, should change what the system pays attention to. Beforehand, the team watches every criterion because it does not yet know which will prove decisive; afterwards, attention should narrow onto the risk the team has chosen to carry. Take a decision to pursue an asset in the obesity space, where pricing and access were the hard part and the crowded field was accepted going in. Recording that decision raises sensitivity to price and payer signals and lowers it on competitive noise, which is why another GLP-1 entrant is close to irrelevant here: the crowding was already priced in. A simple detector would keep reporting every new competitor because it can see them, whereas a calibrated one stays quiet on the settled question and attentive to the open one. Lowering sensitivity is not the same as switching it off. A genuine shift, such as a rival winning a differentiated label, would still clear the higher bar and resurface.

None of these removes the need for judgment; it concentrates it. Trust rests on grounding each synthesized claim in a citation, keeping a human eye on the inputs that are only partly automatable, and being honest about where the data comes from. Sources fall into three tiers, and each is a different kind of problem:

1. Public sources such as ClinicalTrials.gov, PubMed and the HTA sites, which are essentially API calls.

2. Licensed subscriptions such as EvaluatePharma, Citeline and IQVIA, which are external but gated by contract.

3. Internal data such as CRM relationships, safety data and HCP overlap, which is an integration problem and holds the knowledge no crawl will reach.

It is also worth being able to show that the calibration works rather than assert it, by backtesting against past decisions: would the system have surfaced what mattered and stayed quiet on what did not? This mirrors a wider lesson from building agents, where the reasoning loop turns out to be a small part of the work, and most of the effort goes into evaluation, monitoring, and governance.

The Human Checkpoint
The loop narrows attention and concentrates human judgment

A decision, once made, should change what the system pays attention to. Beforehand, the team watches every criterion because it does not yet know which will prove decisive; afterwards, attention should narrow onto the risk the team has chosen to carry. Take a decision to pursue an asset in the obesity space, where pricing and access were the hard part and the crowded field was accepted going in. Recording that decision raises sensitivity to price and payer signals and lowers it on competitive noise, which is why another GLP-1 entrant is close to irrelevant here: the crowding was already priced in. A simple detector would keep reporting every new competitor because it can see them, whereas a calibrated one stays quiet on the settled question and attentive to the open one. Lowering sensitivity is not the same as switching it off. A genuine shift, such as a rival winning a differentiated label, would still clear the higher bar and resurface.

None of these removes the need for judgment; it concentrates it. Trust rests on grounding each synthesized claim in a citation, keeping a human eye on the inputs that are only partly automatable, and being honest about where the data comes from. Sources fall into three tiers, and each is a different kind of problem:

1. Public sources such as ClinicalTrials.gov, PubMed and the HTA sites, which are essentially API calls.

2. Licensed subscriptions such as EvaluatePharma, Citeline and IQVIA, which are external but gated by contract.

3. Internal data such as CRM relationships, safety data and HCP overlap, which is an integration problem and holds the knowledge no crawl will reach.

It is also worth being able to show that the calibration works rather than assert it, by backtesting against past decisions: would the system have surfaced what mattered and stayed quiet on what did not? This mirrors a wider lesson from building agents, where the reasoning loop turns out to be a small part of the work, and most of the effort goes into evaluation, monitoring, and governance.

The Human Checkpoint
The loop narrows attention and concentrates human judgment

A decision, once made, should change what the system pays attention to. Beforehand, the team watches every criterion because it does not yet know which will prove decisive; afterwards, attention should narrow onto the risk the team has chosen to carry. Take a decision to pursue an asset in the obesity space, where pricing and access were the hard part and the crowded field was accepted going in. Recording that decision raises sensitivity to price and payer signals and lowers it on competitive noise, which is why another GLP-1 entrant is close to irrelevant here: the crowding was already priced in. A simple detector would keep reporting every new competitor because it can see them, whereas a calibrated one stays quiet on the settled question and attentive to the open one. Lowering sensitivity is not the same as switching it off. A genuine shift, such as a rival winning a differentiated label, would still clear the higher bar and resurface.

None of these removes the need for judgment; it concentrates it. Trust rests on grounding each synthesized claim in a citation, keeping a human eye on the inputs that are only partly automatable, and being honest about where the data comes from. Sources fall into three tiers, and each is a different kind of problem:

1. Public sources such as ClinicalTrials.gov, PubMed and the HTA sites, which are essentially API calls.

2. Licensed subscriptions such as EvaluatePharma, Citeline and IQVIA, which are external but gated by contract.

3. Internal data such as CRM relationships, safety data and HCP overlap, which is an integration problem and holds the knowledge no crawl will reach.

It is also worth being able to show that the calibration works rather than assert it, by backtesting against past decisions: would the system have surfaced what mattered and stayed quiet on what did not? This mirrors a wider lesson from building agents, where the reasoning loop turns out to be a small part of the work, and most of the effort goes into evaluation, monitoring, and governance.

The Human Checkpoint

The loop narrows attention and concentrates human judgment

A decision, once made, should change what the system pays attention to. Beforehand, the team watches every criterion because it does not yet know which will prove decisive; afterwards, attention should narrow onto the risk the team has chosen to carry. Take a decision to pursue an asset in the obesity space, where pricing and access were the hard part and the crowded field was accepted going in. Recording that decision raises sensitivity to price and payer signals and lowers it on competitive noise, which is why another GLP-1 entrant is close to irrelevant here: the crowding was already priced in. A simple detector would keep reporting every new competitor because it can see them, whereas a calibrated one stays quiet on the settled question and attentive to the open one. Lowering sensitivity is not the same as switching it off. A genuine shift, such as a rival winning a differentiated label, would still clear the higher bar and resurface.

None of these removes the need for judgment; it concentrates it. Trust rests on grounding each synthesized claim in a citation, keeping a human eye on the inputs that are only partly automatable, and being honest about where the data comes from. Sources fall into three tiers, and each is a different kind of problem:

1. Public sources such as ClinicalTrials.gov, PubMed and the HTA sites, which are essentially API calls.

2. Licensed subscriptions such as EvaluatePharma, Citeline and IQVIA, which are external but gated by contract.

3. Internal data such as CRM relationships, safety data and HCP overlap, which is an integration problem and holds the knowledge no crawl will reach.

It is also worth being able to show that the calibration works rather than assert it, by backtesting against past decisions: would the system have surfaced what mattered and stayed quiet on what did not? This mirrors a wider lesson from building agents, where the reasoning loop turns out to be a small part of the work, and most of the effort goes into evaluation, monitoring, and governance.

The Human Checkpoint

The loop narrows attention and concentrates human judgment

A decision, once made, should change what the system pays attention to. Beforehand, the team watches every criterion because it does not yet know which will prove decisive; afterwards, attention should narrow onto the risk the team has chosen to carry. Take a decision to pursue an asset in the obesity space, where pricing and access were the hard part and the crowded field was accepted going in. Recording that decision raises sensitivity to price and payer signals and lowers it on competitive noise, which is why another GLP-1 entrant is close to irrelevant here: the crowding was already priced in. A simple detector would keep reporting every new competitor because it can see them, whereas a calibrated one stays quiet on the settled question and attentive to the open one. Lowering sensitivity is not the same as switching it off. A genuine shift, such as a rival winning a differentiated label, would still clear the higher bar and resurface.

None of these removes the need for judgment; it concentrates it. Trust rests on grounding each synthesized claim in a citation, keeping a human eye on the inputs that are only partly automatable, and being honest about where the data comes from. Sources fall into three tiers, and each is a different kind of problem:

1. Public sources such as ClinicalTrials.gov, PubMed and the HTA sites, which are essentially API calls.

2. Licensed subscriptions such as EvaluatePharma, Citeline and IQVIA, which are external but gated by contract.

3. Internal data such as CRM relationships, safety data and HCP overlap, which is an integration problem and holds the knowledge no crawl will reach.

It is also worth being able to show that the calibration works rather than assert it, by backtesting against past decisions: would the system have surfaced what mattered and stayed quiet on what did not? This mirrors a wider lesson from building agents, where the reasoning loop turns out to be a small part of the work, and most of the effort goes into evaluation, monitoring, and governance.

The loop narrows attention and concentrates human judgment

The Human Checkpoint
The Bottom Line
Where the real advantage lies

The genuinely hard parts of this build are not the crawling, the storage, or the dashboards, which are largely solved and steadily getting cheaper. The difficulty lies in the translations: turning a strategic decision into a measurable signal, and a signal into a calibrated threshold. Deciding that regulatory risk is fairly captured by the count of accelerated approvals, that a Phase 3 event matters here while a Phase 1 event does not, or that the crowding in obesity was already accounted for, are judgments a data engineer cannot make alone and a model cannot be trusted to make unsupervised.

They call for therapy-area understanding and technical fluency in the same room. As the tooling matures, that combination is what will separate the companies building a real advantage from those simply automating faster.

The Bottom Line
Where the real advantage lies

The genuinely hard parts of this build are not the crawling, the storage, or the dashboards, which are largely solved and steadily getting cheaper. The difficulty lies in the translations: turning a strategic decision into a measurable signal, and a signal into a calibrated threshold. Deciding that regulatory risk is fairly captured by the count of accelerated approvals, that a Phase 3 event matters here while a Phase 1 event does not, or that the crowding in obesity was already accounted for, are judgments a data engineer cannot make alone and a model cannot be trusted to make unsupervised.

They call for therapy-area understanding and technical fluency in the same room. As the tooling matures, that combination is what will separate the companies building a real advantage from those simply automating faster.

The Bottom Line
Where the real advantage lies

The genuinely hard parts of this build are not the crawling, the storage, or the dashboards, which are largely solved and steadily getting cheaper. The difficulty lies in the translations: turning a strategic decision into a measurable signal, and a signal into a calibrated threshold. Deciding that regulatory risk is fairly captured by the count of accelerated approvals, that a Phase 3 event matters here while a Phase 1 event does not, or that the crowding in obesity was already accounted for, are judgments a data engineer cannot make alone and a model cannot be trusted to make unsupervised.

They call for therapy-area understanding and technical fluency in the same room. As the tooling matures, that combination is what will separate the companies building a real advantage from those simply automating faster.

The Bottom Line

Where the real advantage lies

The genuinely hard parts of this build are not the crawling, the storage, or the dashboards, which are largely solved and steadily getting cheaper. The difficulty lies in the translations: turning a strategic decision into a measurable signal, and a signal into a calibrated threshold. Deciding that regulatory risk is fairly captured by the count of accelerated approvals, that a Phase 3 event matters here while a Phase 1 event does not, or that the crowding in obesity was already accounted for, are judgments a data engineer cannot make alone and a model cannot be trusted to make unsupervised.

They call for therapy-area understanding and technical fluency in the same room. As the tooling matures, that combination is what will separate the companies building a real advantage from those simply automating faster.

The Bottom Line

Where the real advantage lies

The genuinely hard parts of this build are not the crawling, the storage, or the dashboards, which are largely solved and steadily getting cheaper. The difficulty lies in the translations: turning a strategic decision into a measurable signal, and a signal into a calibrated threshold. Deciding that regulatory risk is fairly captured by the count of accelerated approvals, that a Phase 3 event matters here while a Phase 1 event does not, or that the crowding in obesity was already accounted for, are judgments a data engineer cannot make alone and a model cannot be trusted to make unsupervised.

They call for therapy-area understanding and technical fluency in the same room. As the tooling matures, that combination is what will separate the companies building a real advantage from those simply automating faster.

Where the real advantage lies

The Bottom Line
Whitepaper

The whitepaper breaks down the full governance model: intake and prioritization criteria, decision rights, funding models, a value framework covering financial, capacity, and risk outcomes, and the kill triggers that let leaders stop work without stigma.

Curious how to calibrate agentic AI to your own portfolio decisions?

Get in touch to talk through what that looks like for your setup – whether that's scoping a semantic layer for your existing data platform, working through calibration logic for a specific portfolio call, or just trading notes on where others have started.

Managing Director, Intellishore CH
Mikkel Møller Andersen

Managing Director, Intellishore CH
Mikkel Møller Andersen
Mikkel Møller Andersen
This is the default text value
Senior Manager - Intellishore CH
Rebecca Bub

Senior Manager - Intellishore CH
Rebecca Bub
Rebecca Bub
This is the default text value
Consultant, Intellishore CH
Thibaud Mottier

Consultant, Intellishore CH
Thibaud Mottier
Thibaud Mottier
This is the default text value
filters
All
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Mastering CX in Pharma: Why Data is the Backbone for Success

In this four-part series, “Mastering CX in Pharma”, we aim to share our perspective and approaches on how to succeed in designing a winning customer experience. We will deep dive into engagement model design, data as a key enabler, how to design an operating model that supports your strategic ambitions, and the role of change management.

Click to read more
Customer Experience in Pharma
Mastering CX in Pharma (Part 2)
Mastering CX in Pharma (Part 2)
Customer Engagement & Experience
AI Transformation
Personalization at Scale
Mastering CX in Pharma: Don’t Forget About the Operating Model

In this four-part series, “Mastering CX in Pharma”, we aim to share our perspective and approaches on how to succeed in designing a winning customer experience. We will deep dive into engagement model design, data as a key enabler, how to design an operating model that supports your strategic ambitions, and the role of change management.

Click to read more
Customer Experience in Pharma
Mastering CX in Pharma (Part 3)
Mastering CX in Pharma (Part 3)
Customer Engagement & Experience
Personalization at Scale
AI Transformation
Mastering CX in Pharma: Supporting Culture Shift with Strong Change Management

In this four-part series, “Mastering CX in Pharma”, we aim to share our perspective and approaches on how to succeed in designing a winning customer experience. We will deep dive into engagement model design, data as a key enabler, how to design an operating model that supports your strategic ambitions, and the role of change management.

Click to read more
Customer Experience in Pharma
Mastering CX in Pharma (Part 4)
Mastering CX in Pharma (Part 4)
Customer Engagement & Experience
AI Transformation
Personalization at Scale
The Next Frontier: Transforming Commercial Pharma with Agentic AI

Sign up for the whitepaper and learn how Agentic AI is set to disrupt pharma’s commercial engagement - and why acting now is key to staying ahead. Explore how a strategic approach to AI adoption can help pharma scale impact, optimize operations, and unlock lasting value.

Click to read more
Agentic AI
Transforming Commercial Pharma with Agentic AI
Transforming Commercial Pharma with Agentic AI
AI Transformation
Customer Engagement & Experience
Personalization at Scale