John Snow Labs | How NLP Will Change Clinical Trial Recruitment for the Better
Healthcare Tech Outlook
Home >> >> Vendor Viewpoint >>

This article is part of HealthCare Tech Outlook Innovation Insights series featuring expert contributions nominated by our subscribers and reviewed by our editorial team.

How NLP Will Change Clinical Trial Recruitment for the Better

David Talby, CTO , John Snow Labs

The COVID-19 vaccine development is one of the fastest in history, and will be available to parts of the population before the holidays. Past presidents and President Elect Joe Bidden have promised to take the vaccine televised to ease constituents’ concerns about rushing the process, but still, many are skeptical about the efficacy and safety of the treatment. This shouldn’t come as any surprise—after all, the average time for development is around 15 years, and unlike coronavirus, the severity of the disease is irrelevant. Whether it’s asthma or cancer, launching new drugs takes a long time.

While adverse reactions, immediate side effects, and potential long-term effects may come to mind when assessing the challenges with drug development, there’s one surprising culprit that significantly hinders the process: recruiting patients for clinical trials. This may come as a shock, but With too few participants, 86 percent of clinical trials are delayed and half of all trial sites fail. For some drugs, this equates to hundreds of lives and billions of dollars lost. Too boot, when the treatment finally does hit the market, prices skyrocket.

The availability and willingness of participants isn’t the problem—recruitments must meet very specific requirements for each trial, and narrowing down eligibility from hundreds of thousands of patients is no easy task. In order to find participants, pharmaceutical companies and the clinicians they partner with must get patient data, generate search criteria, find a match, and get patients on board. Unfortunately, this is not as simple as typing in a search term and receiving accurate matches at the click of a button.

Just consider the nuances of medical jargon: often, different terms are used by different doctors in different practices. This gets even more complex when you consider the source of data, which can come from a disorganized collection of reports and images spread across disparate systems. Increasingly complex enrollment criteria is just the cherry on top of a host of other challenges faced when recruiting for a clinical trial. Fortunately, while medical practitioners and researchers are hard at work bringing new treatments to light, data scientists are making strides in technology that will help ease the logistical hurdles.

Natural Language Processing (NLP) enables data scientists to disambiguate clinical and human language, making it easier to find viable clinical trial participants in a faster, more efficient manner. It does this by detecting and refining relevant clinical facts in a sea of highly specific and contextual clinical language.

However, this can become tricky when you account for human error, such as misspellings, incorrect punctuation, and acronyms that mean different things across different healthcare environments. To address this, healthcare-specific NLP models were created to go beyond generic biomedical NLP to extract specific entities and facts required by each medical specialty.

You can see how complicated this can get through a real-world example of recruiting patients with triple-negative breast cancer (TNBC) for a clinical trial. The acronym ‘TNCB’ doesn’t seem like it would be too hard to identify, but there are many other ways to denote this in medical records. Er-/pr-/h2-, (er pr her2) negative, tested negative for the following: er, pr, h2, and triple negative neoplasm of the upper left breast, are just a few of the ways this diagnosis appears in records. This can result in thousands of search results for even a small subset of patients.

Thankfully, healthcare-specific NLP models that are trained on breast cancer data will not only be able to identify variations of these terms, but also when they are in context. As we’ve seen, the same terms can have different meanings in a different specialty, while acronyms (like TNBC)or spelling errors can also muddy the waters. NLP technology is already grappling with these concepts, and is only getting more accurate with time and practice. Just look at clinical pipelines vs. traditional text mining to see how far we’ve already come.

A clinical NLP pipeline can go beyond recognizing terms to also identify what is being said about them. Since most breast cancer patients will have the terms mentioned in their records, it’s important to tell where an assertion made is positive, negative, inconclusive, or about someone else relevant to the patient’s medical history. Traditional text mining algorithms could not even differentiate between these common variations at reasonable accuracy until very recently. Accuracy improvements made just over the past two years have saved a significant amount of time finding the right candidates for clinical trials.

Expecting doctors and other medical professionals to read through trial descriptions and make patient matches manually, is no longer a reasonable or time-effective method. There is simply too much data, too much room for error, and too many other responsibilities to make this worthwhile. It’s vital for healthcare organizations to start leveraging the power of NLP to aid in clinical trial recruitment. With thousands of cutting-edge treatments for serious illnesses behind the gates of clinical trials, NLP is one of the fastest, easiest ways to get these lifesaving treatments to market.

MORE FROM Innovation Insights


The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.