Global Tech News Technology. People. A more open tomorrow.
Science

Explainable AI Raises Doctors' Lung-Cancer Forecast Sensitivity from 0.72 to 0.87

A radiologist reviewing computed tomography scans on two computer monitors

An international research team has found that adding an explainable artificial-intelligence forecast to a case review helped doctors identify more patients whose advanced non-small-cell lung cancer was controlled by immunotherapy. In a Nature Medicine paper published on September 13, physicians' sensitivity on this task rose from 0.72 to 0.87, while accuracy rose from 0.57 to 0.65.

Sensitivity is the share of patients with actual disease control whom the doctors identified correctly. Disease control means that the cancer shrank or remained stable after treatment. The AI tool was designed to support that judgement and to estimate survival, rather than choose a treatment on its own.

The I3LUNG study assembled retrospective records from 2,396 patients treated at six centres in six countries between 2012 and 2023. The wider project tested models built from routine clinical and blood data, computed-tomography scans, digital pathology and genomic information. The physician experiment used the model based on routine clinical and blood data; it did not use the multimodal models.

Twenty physicians, split evenly between lung-cancer specialists and other oncologists or resident doctors, reviewed 100 patient cases. Each doctor assessed ten cases first with the available clinical data and images. The doctor then saw the model's forecast and a SHAP explanation, which showed how individual patient factors raised or lowered that forecast, before making a second assessment.

With AI support, sensitivity increased from 0.68 to 0.83 for specialists and from 0.77 to 0.90 for non-specialists. Agreement between the two groups rose from a kappa score of 0.11 to 0.48. The doctors adopted 74.5% of correct AI suggestions, but they also followed incorrect suggestions: specialists did so in 72.2% of those cases and non-specialists in 63.6%. The gain in sensitivity also came with a small reduction in specificity, the ability to identify patients whose disease was not controlled.

The routine-data models reached an area under the curve of up to 0.77 in a held-out test cohort. This measure is 0.5 for chance-level ranking and 1.0 for perfect ranking. Adding scans, pathology and genomic features produced a value as high as 0.88 during cross-validation, but that improvement did not consistently carry into the test and external-validation cohorts.

The study used retrospective records, and only 339 patients had complete data across all modalities. The external validation also came from one geographically distinct cohort. The team is now conducting a silent prospective validation in 2,000 patients, followed by a planned randomized trial. Those tests will determine whether the tool improves real treatment decisions without causing doctors to rely on a confident but incorrect forecast.

Sources