Influence of Functional Magnetic Resonance Imaging Data Preprocessing Pipelines on the Accuracy of Schizophrenia Classification Using Machine Learning Methods.
The aim of the study was to analyze the influence of various pipelines for preprocessing raw functional magnetic resonance imaging (fMRI) data on the accuracy of classification of subjects into schizophrenia patients and healthy controls using machine learning methods, and to give recommendations for optimizing the data preprocessing pipeline for this task.
The study used fMRI data from 72 subjects acquired on a Siemens Magnetom Verio 3T MRI scanner (Siemens Healthineers, Germany) at the National Research Centre "Kurchatov Institute". Seven different preprocessing pipelines were used on each dataset. Feature vectors for classification were constructed from each preprocessed dataset. We applied the following three algorithms for feature vector construction: ReHo, FCM, and FHR. Classification was performed using 15 machine learning methods from the scikit-learn package for each feature set and each preprocessing pipeline. The final accuracy was determined as the maximum accuracy among all machine learning methods.
No single optimal preprocessing pipeline for all feature vector construction algorithms was found. Smoothing increased accuracy for the voxel metrics ReHo and FHR, but not for the regional metric FCM. Spatial smoothing was the only preprocessing step that affected classification accuracy for the FHR method. The FHR feature set without smoothing demonstrated accuracy close to chance level. The steps of frequency filtering and median normalization substantially increased the accuracy for both the ReHo and FCM methods, both with and without smoothing. The steps of inhomogeneity correction and slice timing correction did not increase accuracy for the FCM feature set but did improve it by several percent for the ReHo feature set. Application of ICA filtering changed accuracy within a range of 5%, in either a positive or negative direction.
Based on the obtained results, the following recommendations can be given for fMRI data preprocessing pipelines for binary classification of subjects into schizophrenia patients and healthy controls using machine learning methods and taking into account the feature vector construction algorithm. The most suitable for the ReHo metric pipeline includes motion correction, normalization, inhomogeneity correction, slice timing correction, frequency filtering, median normalization, and spatial smoothing. The most suitable for the FCM metric pipeline includes motion correction, normalization, inhomogeneity correction, slice timing correction, frequency filtering, and median normalization. Spatial smoothing is the most important factor for the FHR metric. The ICA filtering step does not provide a clear benefit and should therefore be applied with caution.
The study used fMRI data from 72 subjects acquired on a Siemens Magnetom Verio 3T MRI scanner (Siemens Healthineers, Germany) at the National Research Centre "Kurchatov Institute". Seven different preprocessing pipelines were used on each dataset. Feature vectors for classification were constructed from each preprocessed dataset. We applied the following three algorithms for feature vector construction: ReHo, FCM, and FHR. Classification was performed using 15 machine learning methods from the scikit-learn package for each feature set and each preprocessing pipeline. The final accuracy was determined as the maximum accuracy among all machine learning methods.
No single optimal preprocessing pipeline for all feature vector construction algorithms was found. Smoothing increased accuracy for the voxel metrics ReHo and FHR, but not for the regional metric FCM. Spatial smoothing was the only preprocessing step that affected classification accuracy for the FHR method. The FHR feature set without smoothing demonstrated accuracy close to chance level. The steps of frequency filtering and median normalization substantially increased the accuracy for both the ReHo and FCM methods, both with and without smoothing. The steps of inhomogeneity correction and slice timing correction did not increase accuracy for the FCM feature set but did improve it by several percent for the ReHo feature set. Application of ICA filtering changed accuracy within a range of 5%, in either a positive or negative direction.
Based on the obtained results, the following recommendations can be given for fMRI data preprocessing pipelines for binary classification of subjects into schizophrenia patients and healthy controls using machine learning methods and taking into account the feature vector construction algorithm. The most suitable for the ReHo metric pipeline includes motion correction, normalization, inhomogeneity correction, slice timing correction, frequency filtering, median normalization, and spatial smoothing. The most suitable for the FCM metric pipeline includes motion correction, normalization, inhomogeneity correction, slice timing correction, frequency filtering, and median normalization. Spatial smoothing is the most important factor for the FHR metric. The ICA filtering step does not provide a clear benefit and should therefore be applied with caution.
Authors
Poyda Poyda, Orlov Orlov, Zhemchuzhnikov Zhemchuzhnikov, Kozlov Kozlov, Kartashov Kartashov, Bravve Bravve, Kaydan Kaydan, Kostyuk Kostyuk
View on Pubmed