Machine learning to optimise a new vaginal inflammation test
Machine learning is already having a transformative effect on disease diagnosis, from heart disease to Alzheimer’s. Now, the researchers behind the Genital Inflammation Test (GIFT) are using machine learning to optimise this medical device’s ability to detect levels of inflammation that point to bacterial vaginosis (BV) or sexually transmitted infections (STIs). In low- and middle-income settings, where most asymptomatic STI and BV cases currently go untreated, this could have significant impacts on women’s sexual and reproductive health and agency.
The bacteria responsible for some STIs and BV are often silent stowaways, yet their health impacts can be devastating. These frequently asymptomatic conditions cause vaginal inflammation, increasing the risk of HIV infections, fertility issues, and premature birth. The GIFT device can help healthcare workers in under-resourced settings detect asymptomatic STIs and BV by detecting genital inflammation.
GIFT is a lateral flow device, that measures three immune proteins, called cytokines, that act as markers of inflammation. Cytokines transmit instructions from the immune system to regulate inflammation. While many cytokines are involved in the vagina’s inflammatory response, the GIFT researchers found that the presence of just three specific cytokines can indicate the presence of STIs or BV, even in the absence of clear symptoms. To researchers, these cytokines are called interleukin (IL)-1α, IL-1β, and Interferon Gamma-Induced Protein 10 (IP-10). Dr Musalula Sinkala, a lecturer at the University of Cape Town’s Computational Biology Unit, has partnered with the GIFT team to lead the data analysis of the device’s performance in the field, using machine learning. One of the goals is to establish the cytokine cut-off values for the GIFT device, which will be critical to interpret the results. This is more complex than it might seem.
GIFT as a screening tool
Selecting cytokine biomarker cut-off values that balance sensitivity of the GIFT device (the number of STI/BV cases correctly classified) and specificity (number of cases incorrectly classified) is one of the most important aims of GIFT’s machine-learning component.
If we think of GIFT as a ‘magnifying glass’ for vaginal inflammation, then high sensitivity would be equivalent to a high magnification, where all inflammation is clearly visible. In this case, the test would likely ‘magnify’ most asymptomatic cases, along with some false positives. Conversely, a magnifying glass with lower magnification (a highly specific test where only the most obvious inflammation is visible) would avoid false positives but also miss many women with infections or BV.
The team are positioning the GIFT device as a screening test. This means that women receiving a positive GIFT result for vaginal inflammation would then require a follow-on diagnostic test to determine the underlying infection, if any, before receiving treatment. GIFT would be part of a two-step testing algorithm. For this reason, the focus of Sinkala’s analysis will be to increase the magnification (or optimise it). “Missing a positive case is worse than having to confirm whether the infection is actually there,” he says.
Accelerating problem solving
Machine learning uses algorithms to trawl large datasets for patterns between input and output data. The GIFT device relies on the relationship between the three cytokines to distinguish between women with and without vaginal inflammation. To determine exactly what this relationship looks like, Sinkala and his students will feed supervised machine learning models with historical data (collected from the three GIFT clinical study sites in South Africa, Zimbabwe, and Madagascar) so it can learn to connect cytokine levels with inflammation. By comparing its inflammation predictions with the actual STI and BV rates measured at the clinical study sites, it can iteratively improve its accuracy.
The process is known as ‘target profiling’, where the target is the variable (inflammation associated with BV or STIs) that the algorithm seeks to predict. Models such as decision trees can effectively analyse the cytokines’ combined prediction for inflammation, considering interactions between them that simpler statistical models might miss. The researchers will use optimisation algorithms to refine the results and determine cut-off points that balance sensitivity, specificity and accuracy.
Since the GIFT device analyses the combination of three cytokines—IL-1α, IL-1β, and IP-10—modelling all three cut-off points (and the corresponding values for sensitivity and specificity) becomes complex. “That’s where other statistical methods like machine learning are extremely powerful, because now we can model a larger space,” Sinkala says. He likens the problem to finding the highest points of Cape Town. Rather than trying to plot the topography on a 2D graph, you would model it in the three dimensions. Then, machine learning would allow you to explore all possible peaks (combinations of cytokine cut-off points) simultaneously.
“The algorithms will take 1000 orders of magnitude less time than you would, if you were to try searching for the highest bit yourself.”
Homegrown healthcare innovations
Besides speed and efficiency, machine learning has the advantage of continuing to learn as more data is added, ensuring that predictions remain accurate for different populations. This is particularly important in the GIFT context, as the cytokine profiles of women with and without BV or STIs may change depending on age, HIV status, geography or other factors. Machine learning models sometimes perform inconsistently across ethnic groups due to biases arising from the data collection, the data it has been trained on, or algorithm design. Because the input variables are limited to cytokine levels in the GIFT study, the possibility for bias in this instance is limited. However, to ensure the models are robust and unbiased, the team will test them on datasets from different African country settings.
“Our job is to understand which model performs the best on a particular data set, and then we can tweak or change parameters in that module to keep it performing consistently across different datasets, and to make sure it does not deteriorate in its performance,” says Sinkala.
The GIFT team is excited about its new partnership with Sinkala and his UCT team. Beyond contributing to the robustness of the GIFT device, Sinkala’s artificial intelligence expertise will help expand local capacity in machine learning expertise, accelerating African healthcare research and innovation.
