Narrative Medicine complements structured clinical information by providing access to patients’ and healthcare professionals’ lived experiences. Sentiment analysis is a promising tool, yet its application is hindered by textual noise and by the limited availability of robust Italian-language resources.We study Large Language Models as controlled preprocessing modules for Italian narratives, using constrained correction and Italian-to-English translation designed to preserve meaning. Viewing narratives as an unstructured textual modality, we investigate how sentiment polarity and confidence estimates depend on preprocessing choices and generation-parameter calibration.We introduce a three-stage paired pipeline that computes sentiment on variants derived from the same input: preprocessed Italian text, LLM-corrected Italian text, and its English translation, enabling comparison between Italian-based and English-translation-based sentiment classification. The study is conducted on two labeled Italian benchmarks and an unlabeled narrative medicine corpus, contributing: (i) a controlled pipeline for LLM-assisted correction and translation in Italian sentiment analysis, (ii) a comparison between Italian-specific and translation-based multilingual inference, and (iii) an analysis of generation-temperature effects and prediction stability for decision-support use.
Controlled LLM Correction and Multilingual Inference for Italian Sentiment Analysis in Narrative Medicine
Franzoni, Valentina
Supervision
;Polticchia, MattiaMembro del Collaboration Group
;Saetta, DanielaMembro del Collaboration Group
;Florindi, EmanueleMembro del Collaboration Group
2026
Abstract
Narrative Medicine complements structured clinical information by providing access to patients’ and healthcare professionals’ lived experiences. Sentiment analysis is a promising tool, yet its application is hindered by textual noise and by the limited availability of robust Italian-language resources.We study Large Language Models as controlled preprocessing modules for Italian narratives, using constrained correction and Italian-to-English translation designed to preserve meaning. Viewing narratives as an unstructured textual modality, we investigate how sentiment polarity and confidence estimates depend on preprocessing choices and generation-parameter calibration.We introduce a three-stage paired pipeline that computes sentiment on variants derived from the same input: preprocessed Italian text, LLM-corrected Italian text, and its English translation, enabling comparison between Italian-based and English-translation-based sentiment classification. The study is conducted on two labeled Italian benchmarks and an unlabeled narrative medicine corpus, contributing: (i) a controlled pipeline for LLM-assisted correction and translation in Italian sentiment analysis, (ii) a comparison between Italian-specific and translation-based multilingual inference, and (iii) an analysis of generation-temperature effects and prediction stability for decision-support use.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


