Miner extracting evidence from layers of podcast transcripts, timestamps and research themes.

When a Podcast Archive Becomes a Research System

For years, I treated each podcast interview as a discrete piece of content.

Research the guest. Record the conversation. Edit it. Publish it. Promote it. Then move on to the next one.

The interviews remained available, of course, but the unit of value was always the episode.

That changed when five researchers: Stefan Seuring, Jannik Neuberger, Sharfah Ahmad Qazi, Lara Schilling and Andrea S. Patrucco — analysed 52 episodes of my Sustainable Supply Chain podcast for a peer-reviewed paper, Bridging Digital and Sustainable Supply Chains: Insights From a Podcast Analysis.

Their work examined how technologies including artificial intelligence, IoT, blockchain and cloud services were connected with sustainable supply-chain practices and outcomes.

I have written separately about what they found.

What interested me just as much was what their research implied.

They had treated podcast interviews not as media artefacts, but as qualitative research material.

That raised a larger question.

If meaningful insights could emerge from analysing 52 interviews together, what might be visible across the hundreds of other conversations I had recorded, but impossible to see when each episode was considered in isolation?

That question has changed how I think about the entire podcast archive.

From individual interviews to accumulated evidence

An individual interview can be valuable.

A supply-chain executive may explain why an AI deployment struggled. A technology provider may describe how a particular problem was solved. An energy specialist may outline how regulation is affecting investment. A sustainability leader may describe the operational consequences of new reporting requirements.

But viewed independently, each conversation remains one perspective at one point in time.

The analytical possibilities change when hundreds of dated interviews can be examined together.

Patterns can be compared across industries.

Claims made by technology providers can be tested against the experiences of operators.

Recurring implementation problems can be traced across otherwise unrelated conversations.

Ideas that were discussed enthusiastically several years ago can be compared with what people say about them after real-world deployment.

Apparent consensus can be tested by deliberately looking for dissent.

And because every conversation took place at a particular moment, changes in language, priorities and confidence can be examined over time.

That is when an interview collection begins to become something more than a back catalogue.

It becomes a potential source of longitudinal qualitative evidence.

The challenge is not generating answers

The obvious way to analyse hundreds of transcripts today would be to give them to an AI system and ask for the major trends.

That can produce an impressive answer.

The difficulty is establishing whether the answer deserves to be believed.

A plausible synthesis can combine genuine patterns, isolated anecdotes and the model’s own assumptions into something that sounds much more certain than the underlying evidence warrants.

So the problem I have been trying to solve is not simply how to extract answers from the transcripts.

It is how to retain enough provenance that those answers can be challenged.

If an analysis suggests that poor data foundations repeatedly undermine supply-chain AI initiatives, I need to know where that conclusion came from.

Who described the problem?

What role were they in?

Which organisation were they speaking for?

When did the interview take place?

What was the surrounding context?

And where, exactly, in the conversation was the relevant passage?

Without that chain back to the original source, the system may be useful for brainstorming.

It is much less useful for research.

First, I needed the complete record

There was a practical problem before any of this analysis could work properly.

I did not have transcripts for every episode.

For more recent interviews, that was not an issue. As transcription technology improved, transcription became part of my normal editing workflow, so I already had text versions of a large proportion of the collection.

The earlier years were different. When many of those episodes were produced, accurate automated transcription was either unavailable to me, expensive, or not practical enough to make it part of the process.

That left a significant hole in the historical record.

The solution came, somewhat unexpectedly, from Buzzsprout, the podcast hosting platform I use.

On a recent episode of their Buzzcast podcast, the team mentioned that creators could download transcripts of their back catalogue. I contacted their support team because many of my older episodes predated the point at which I had started generating transcripts myself.

They were able to run a process across the back catalogue so that those earlier episodes were transcribed as well.

That mattered far more than simply making old episodes searchable.

It meant I could finally bring the full collection of more than 800 interviews across both of my podcasts, into the same research environment.

Without that completeness, longitudinal analysis would always carry an awkward qualification: entire periods of the podcasts would effectively be invisible to the system.

With the historical transcripts available, the archive could be treated as a much more coherent record.

Then the harder analytical work could begin.

Making hundreds of conversations searchable

The first practical step is to divide long transcripts into smaller, usable passages.

An hour-long interview is too broad a unit for precise analysis. At the same time, fragments that are too small lose the context needed to understand what the speaker meant.

The goal is therefore to create passages that are specific enough to retrieve but substantial enough to preserve the argument.

Each passage remains connected to its source information: the episode, guest, date, speaker and, where available, timestamp.

That source connection is crucial.

It means analysis can move from hundreds of interviews to a specific piece of evidence and then back to the original conversation.

The next challenge is language.

Experts often describe the same underlying problem in completely different ways.

One person may talk about fragmented master data. Another may describe systems that cannot communicate. A third may explain that planners spend hours reconciling spreadsheets before they trust the numbers.

A conventional keyword search may treat those as unrelated statements.

Conceptually, they may be describing closely connected problems.

To address that, each passage can be represented mathematically according to its meaning. These representations, known as embeddings, make it possible to search for conceptually similar passages even when the vocabulary differs.

That makes a much larger proportion of the interview collection discoverable.

But it introduces another risk.

Retrieval is not evidence

Finding relevant passages is not the same as establishing a conclusion.

This distinction is fundamental.

Suppose a search retrieves 20 passages supporting one explanation and three that contradict it.

It would be tempting to interpret that as strong evidence that the first view dominates.

It is not.

The number of passages returned can be affected by the wording of the search, the structure of the transcripts, the retrieval method, repetition within individual interviews and the composition of the guest base.

Search frequency is not prevalence.

Ranking is not representativeness.

Model confidence is not evidence strength.

The retrieval system therefore has a narrower role: it helps locate material that deserves examination.

The analytical work begins after that.

Trying to disprove the finding

One of the most important parts of the process is deliberately searching for evidence that challenges an emerging conclusion.

Suppose several interviews suggest that supply-chain AI initiatives struggle to progress beyond pilots because the underlying data environment is fragmented.

A weak research process would continue searching for more examples of poor data.

A stronger one asks what would make that interpretation wrong.

Are there successful deployments that operated despite weak data?

Do practitioners identify a different constraint?

Do technology providers and operators describe the same failure differently?

Does the pattern persist across industries?

Has the explanation changed over time?

Is the apparent finding being driven disproportionately by a small number of guests or episodes?

The purpose is not to manufacture balance where none exists.

It is to understand how much weight a conclusion can legitimately carry.

Some patterns may be strongly supported across unrelated conversations.

Others may be suggestive but incomplete.

Some may turn out to depend on a particular sector, technology or type of contributor.

And sometimes the appropriate conclusion is simply that the available interviews do not provide enough evidence.

That is preferable to producing a confident answer that the material cannot support.

The importance of returning to the source

AI is extremely useful in this process because it can help interrogate a collection far larger than any individual could reliably hold in memory.

But I do not want the AI system itself to become the authority.

The authority remains the underlying evidence.

If a finding matters, it should be possible to identify the people whose accounts support it, the interviews in which they made those observations, the circumstances they were discussing and the passages on which the interpretation rests.

Contradictory evidence should be equally visible.

That is why timestamps and source metadata matter.

The goal is not to arrive at statements such as “the AI found that…”

It is to be able to say:

This pattern appears across these interviews.

These contributors describe the mechanism in these terms.

These other interviews challenge or qualify it.

And these are the limits of what the evidence allows us to conclude.

That is the methodological principle behind the Executive Briefings I have begun publishing from the interview collection.

A research base that keeps changing

There is another characteristic that makes this collection unusual.

It never really closes.

I continue to publish new interviews every week across my Resilient Supply Chain and Climate Confident podcasts.

Each new conversation becomes another dated contribution to the evidence base.

That creates the possibility of revisiting earlier conclusions rather than treating them as permanent.

A technology discussed largely in terms of promise several years ago may later be discussed in terms of deployment problems, economics, organisational resistance or measurable operating performance.

Themes can emerge, strengthen, weaken or change character.

The evolution of the supply-chain podcast itself reflects some of those shifts. It began as Digital Supply Chain, became Sustainable Supply Chain as decarbonisation and responsible sourcing moved closer to the centre of business strategy, and later became Resilient Supply Chain as disruption, geopolitical risk and operational continuity became harder to treat as separate concerns.

Those changes reflect both shifts in the business environment and my own editorial focus, which is important context when interpreting trends across the archive.

Differences between rhetoric and experience can become visible.

Predictions can eventually be compared with outcomes.

Time becomes part of the analysis.

And the relationship can operate in both directions.

Research into earlier interviews can expose contradictions, gaps and unresolved questions. Those gaps can then influence what I ask future guests.

The interviews produce research.

The research identifies missing evidence.

That missing evidence produces better questions.

And the resulting interviews strengthen the research base.

That feedback loop may ultimately prove more important than the underlying technology.

What this research cannot tell us

There is an important boundary around all of this.

The podcast collection is not a representative sample of the supply-chain, climate or energy industries.

Guests are selected rather than randomly sampled.

Editorial priorities influence which topics receive attention.

Some sectors, technologies and types of organisation are better represented than others.

Technology vendors feature prominently.

Public interviews may also be more likely to contain successful implementations than failed ones.

The academic researchers who analysed the 52 Sustainable Supply Chain episodes encountered a similar limitation. They noted that the material contained primarily positive accounts of digital technology use, making it difficult to assess unintended consequences and what they described as the darker side of digitalisation.

Those limitations do not invalidate the material.

They simply define what can legitimately be inferred from it.

The collection should not be used to claim that “most companies” believe something simply because a theme appears frequently.

It is much better suited to identifying recurring mechanisms, disagreements, implementation experiences, emerging weak signals, changing narratives and questions that warrant further investigation.

It is qualitative evidence, not an industry survey.

That distinction matters.

From content archive to research infrastructure

This work is what led me to establish the Research section of this site and begin producing Executive Briefings based on structured analysis across multiple interviews.

The objective is not to summarise old podcast episodes more efficiently.

It is to ask questions that individual interviews cannot answer.

Why do apparently successful AI pilots fail to become durable operating capabilities?

Where do technology providers and practitioners describe the same problem differently?

Which operational obstacles recur despite successive generations of technology?

Which widely accepted assumptions have surprisingly little evidence behind them?

How do priorities change as technologies move from experimentation towards deployment?

And what do the interviews tell us now that they could not have told us when they were originally published?

There are now more than 800 episodes across the wider podcast collection, with new interviews being added every week.

For years, I thought I was building a back catalogue.

The academic paper forced me to recognise something else.

I had also been accumulating a dated record of how hundreds of senior business leaders working inside major technological, operational and environmental transitions understood those changes while they were happening.

Making that record usable required more than AI. It required complete source material, reliable provenance and the ability to return every important finding to the people and conversations behind it.

The opportunity now is to stop treating that record simply as published content.

And start asking what it can teach us.


Discover more from Tom Raftery.com

Subscribe to get the latest posts sent to your email.