New Delhi: AI is making its way into pathology labs, where software can help detect cancer, analyse tissue samples and classify blood cells. But doctors deciding whether to use these tools may have little public evidence to tell them how well they actually work.
A new study of 77 AI and machine-learning-based diagnostic software devices found that performance evidence was publicly available for only 30 products, or 39 per cent of them. For the other 47, researchers could not find publicly accessible, device-specific evidence showing how accurately they performed.
Published in npj Digital Medicine, the study examined software authorised for digital pathology and haematology morphology in the European Union and the US. The products were designed for tasks including tissue assessment, cancer detection, biomarker scoring and blood-cell analysis. Researchers searched regulatory databases, scientific literature, conference proceedings and manufacturer websites.
Of the 77 devices, 68 were designed for digital pathology and nine for analysing blood or bone-marrow cell morphology. All had entered the European market through either the current In Vitro Diagnostic Regulation (IVDR) or the older In Vitro Diagnostic Directive (IVDD). Only eight were authorised in both the EU and US.
The evidence gap was particularly large among products sold only in Europe. All eight devices authorised in both jurisdictions had publicly available performance evidence, compared with just 22 of the 69 EU-only devices, about 32 per cent.
That does not necessarily mean the remaining devices have never been tested. Manufacturers may hold validation data in technical files or submit it to notified bodies without making it public. But researchers said this lack of transparency makes it harder for clinicians, laboratory directors and procurement teams to independently assess whether a product is appropriate for their patients and workflows.
The evidence that was available was also difficult to compare. The 30 products with public data reported 15 different types of performance measures, including sensitivity, specificity, agreement, correlation and prognostic outcomes. Studies differed in the clinical tasks they examined, the samples used, reference standards and comparison groups.
Most of the studies were retrospective: 25 of the 30. For 28 devices where researchers could extract sample sizes, the median study size was 211 patients, slides or specimens. More than half involved fewer than 300 units.
Also read: Mumbai came to pull apart an 85,000-brick Lego wall. Then a govt visit cut the event short
Needs more testing
There was also limited evidence on whether AI actually improves doctors’ performance. Only five devices had studies directly comparing pathologists working with and without AI assistance. All reported some improvement with AI, although the size and statistical significance of the gains varied. In one study, the time taken to read prostate biopsy cases fell from 55.7 seconds to 36.8 seconds with AI assistance.
The researchers called for more consistent reporting of AI performance, including complementary measures such as sensitivity and specificity. They also recommended more prospective, multicentre studies that test these tools in real clinical workflows.
For pathology labs, then, regulatory approval may be only the starting point. It can establish that a product meets market-access requirements, but the study suggests it does not necessarily give doctors enough publicly available information to decide whether an AI tool is reliable, useful or worth adopting.
(Edited by Ratan Priya)
