Claims in this piece are observable and methodological — no client or benchmark data used.
A procurement lead at a mid-sized pharmaceutical company has a problem that has quietly changed shape. They need a manufacturing partner for a sterile injectable — EU GMP, a defensible inspection history, technology-transfer capability — and they have perhaps a week before the first internal review. Ten years ago the shortlist came from directories, colleagues, and the exhibitor list at CPhI. Today, an evaluator may also open ChatGPT, Gemini, Perplexity, Claude, or Copilot and ask a question about potential suppliers.
That instinct is reasonable. It’s also where good procurement people get into trouble — not because AI is useless for supplier research, but because it is useful in a narrow, specific way that is easy to overestimate. The teams that get value from AI in supplier evaluation are the ones who are precise about what it can and cannot tell them. This article is about drawing that line clearly.
What AI is actually good for at the research stage
Used well, an AI assistant is a fast way to do three things at the very front of supplier discovery: widen the field, surface language, and expose gaps.
Widening the field means generating a longer list of candidate manufacturers than your existing network would produce — including organisations you hadn’t considered, in geographies you don’t usually source from. Surfacing language means learning how a capability is described across the industry, so your own requirement is phrased in terms that match how manufacturers present themselves. Exposing gaps means noticing where an assistant can say a great deal about one company and almost nothing about another — which tells you something about how clearly each has made its capabilities public, though not, on its own, about the quality of the underlying operation.
None of these is a verdict. Each is a starting input to a process that still runs on technical evaluation, quality assessment, regulatory review, and — in almost every real case — an audit. That distinction is the whole discipline.
Why the confident answer is the dangerous one
Here is the trap. Ask an AI assistant to “recommend the best CDMOs for sterile injectables in regulated markets,” and it will give you a clean, confident, well-written answer with named companies and plausible reasons. The fluency is the risk. That answer is assembled from whatever public information the model can access and interpret — which means it systematically favours manufacturers who have made themselves legible, not necessarily those who are best.
A superbly capable manufacturer with a thin or poorly-structured public presence can be entirely absent from that confident list. A more ordinary one with a clear, specific, consistent public footprint can feature prominently. The AI is not necessarily wrong; it is working from what it can find and interpret. But a procurement team that reads the output as a ranking of quality rather than a map of public legibility will draw precisely the wrong conclusion — and may quietly exclude a strong partner who simply communicates poorly.
This is the single most important thing to understand about using AI in supplier evaluation, and it is worth stating plainly.
The Emerivo evaluation principle: treat AI output as a research lead, never a reference check
For procurement, the useful mental model is a division your team already understands from other contexts: the difference between a research lead and a reference check.
A research lead points you somewhere worth looking. It is allowed to be incomplete, uneven, even occasionally wrong, because its only job is to generate candidates and questions for you to pursue. A reference check is different: it is evidence you rely on to make a decision, and it must be verifiable, attributable, and current.
The moment a team treats an assistant’s confident recommendation as evidence of capability — rather than as a prompt to go and verify capability through proper channels — it has crossed a line that matters, and it has outsourced judgement to a system that cannot be held accountable for it.
Everything practical follows from holding that line.
How to use AI in supplier evaluation, in practice
For a procurement team that wants the benefit without the trap, a disciplined sequence looks like this:
- Use AI to widen, then verify to narrow. Let assistants generate a broad candidate list; do the narrowing yourself, against your real requirements, through direct evidence.
- Run the same question across several assistants. Ask ChatGPT, Gemini, Perplexity, Claude, and Copilot the same query and compare. Divergence between them is a signal that the public information is thin or inconsistent — useful to know, and a reason to verify harder, not less.
- Ask for reasons, not just names. A recommendation with a stated rationale (“suggested because they publicly describe lyophilised fill-finish for regulated markets”) is more useful than a bare list, because you can check the reason.
- Treat absence as a question, not a verdict. If a manufacturer you know to be strong doesn’t appear, that tells you about their public legibility — worth noting — not about their capability.
- Verify everything that will inform a decision through regulatory databases, direct capability discussions, quality documentation, and audits. The AI stage ends where the qualification stage begins.
That sequence keeps AI in the role it’s good at — fast, broad, front-of-funnel research — and out of the role it’s dangerous in: standing in for evidence.
If our procurement team quietly dropped a genuinely strong manufacturer from a shortlist because an AI assistant couldn’t describe them well, would we ever know we’d done it — and what would that mistake have cost us?
A demonstration you can run before you rely on it
Before trusting AI output in any live evaluation, it’s worth seeing its variability for yourself. Take a real requirement — say, “Recommend contract manufacturers experienced in sterile injectables for EU and US regulated markets” — and run it, unchanged, across ChatGPT, Gemini, Perplexity, Claude, and Copilot.
Then compare. Notice how much the named companies differ between platforms. Notice how the stated reasons vary. Notice how confident each answer sounds regardless of how much it actually seems to know. That exercise is the fastest way to internalise why AI output is a research lead and not a reference check: the same question can produce overlapping-but-different answers across the five assistants — none of which is a ranking of quality, all of which are maps of what each system could find and interpret.
This is a demonstration, not a benchmark. Results differ by platform and change as models evolve. The point is not any specific answer; it’s the pattern of variability itself.
What this means for the other side of the table
There is a mirror to all of this that procurement teams should understand, because it affects the quality of their own research. If AI output favours manufacturers who are publicly legible, then the manufacturers missing from your AI-assisted research are not necessarily weaker — they may simply be worse at communicating publicly. The strongest partner for your specific programme might be the one an assistant can barely describe.
That’s why the disciplined team treats AI as the beginning of discovery, not the end. It widens the aperture, then does the real work of evaluation the way it has always been done — with evidence, scrutiny, and audits. Understanding why some capable manufacturers are hard for AI to surface is itself part of being a sophisticated buyer. (It’s also, from the manufacturer’s side, exactly the problem of being shortlisted before the first sales call — the same phenomenon seen from the other chair.)
Where Emerivo fits
A clarification worth making plainly: Emerivo is a specialist advisory for pharmaceutical manufacturing and research organisations — TPMs, CMOs, CDMOs, and CROs. We do not serve procurement teams as buyers. This article explains the buyer’s side precisely because it now shapes what happens to you, the manufacturer, before you know an evaluation is underway.
Our focus is the other side of the equation: how AI assistants find, interpret, and represent pharmaceutical organisations when buyers ask questions about them. The AI Discovery Audit™ tests that visibility against realistic buyer prompts and identifies where the public record is clear, incomplete, inconsistent, or difficult for AI to interpret — to identify opportunities for improvement, not to promise rankings or recommendations.
For CDMOs, the practical lesson is simple: if procurement is using AI to decide who deserves a closer look, your first task is not to persuade the buyer. It is to make sure the right facts about your organisation are available to be understood.
For the manufacturer’s side of this dynamic, see Why Some Pharmaceutical Manufacturers Are Shortlisted Before the First Sales Call and Seven Trust Signals Every Pharmaceutical Website Should Demonstrate.
Frequently asked questions
No. AI output reflects what is publicly discoverable and interpretable, which favours legible manufacturers over necessarily better ones. Use it to generate candidates; decide on quality through your own technical, regulatory, and audit process.
No — used as a research lead, it’s genuinely useful for widening the field and surfacing industry language quickly. The mistake is treating its output as evidence rather than as a prompt to verify.
Because they access and interpret public information differently, and models change over time. That variability is exactly why no single answer should be treated as authoritative — and why running the same query across several is more informative than trusting one.
It can speed up the earliest research stage. It does not speed up — or replace — qualification itself, which still requires technical evaluation, quality assessment, regulatory review, and audits.