Skip to content
Connect Career

TEXT2WIN

LLM-Assisted Digital Twins for Predicting Therapy Response in Triple-Negative Breast Cancer

Triple-negative breast cancer (TNBC) is one of the most aggressive forms of breast cancer. Although chemotherapy can be highly effective for some patients, only a minority achieve a complete response to treatment. Currently, clinicians cannot reliably predict which patients will benefit most from therapy before treatment begins. As a result, many patients are exposed to toxic treatments without knowing whether they will be effective.

The Text2Win project explores a new artificial intelligence (AI)-based approach to improve personalized treatment decisions for TNBC patients. Hospitals generate large amounts of valuable information in pathology reports, but much of this information exists as unstructured text and cannot easily be used by conventional predictive models. We aim to unlock this hidden knowledge using large language models (LLMs) and integrate it into digital twin–like models that can help predict therapy response before treatment starts.

To achieve this goal, we will analyze publicly available TNBC datasets that contain pathology reports and clinical outcome information. Advanced LLMs will be used to automatically extract clinically relevant features from pathology text, including tumor characteristics, immune-related information, and other predictive biomarkers. These extracted features will then be combined with clinical data to create patient-specific digital representations that can be used to estimate the likelihood of treatment success. A key component of the project is the systematic evaluation of whether information derived from pathology text can improve prediction performance compared with models that rely solely on structured clinical variables. We will also investigate how interpretable these AI-generated predictions are, helping clinicians understand which factors contribute most strongly to treatment response. By focusing on transparency and validation, the project aims to generate clinically meaningful insights rather than producing “black box” predictions.

In collaboration with experts in biomedical data science, oncology, and pathology at CBmed, the project will provide important proof of concept for combining large language models with digital twin methodologies in cancer research. If successful, this approach could improve prediction of pre-treatment therapy response, reduce unnecessary toxic treatments, and support more individualized treatment strategies for patients with TNBC. Furthermore, the methods developed within Text2Win could be extended to other cancer types and contribute to the broader adoption of AI-driven precision medicine.

Together, these activities will help determine whether clinically relevant knowledge hidden in pathology reports can be transformed into actionable predictive models, bringing us one step closer to truly personalized cancer treatment and the future of precision medicine.

This project is funded by the Land Steiermark UFO Program