{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro","metadata":{}},{"cell_type":"markdown","source":"We are delighted to introduce the newest project taken up by the students of the educational platform: [Practicum](https://practicum.com/).\n\nAdvancing to our third project, the journey so far has been exhilarating. Our student base, which has gained substantial experience from the past projects, is now ready to confront more intricate challenges, and with impressive results at that. Additionally, we've welcomed some new members, who are eager to kickstart their exploration into the world of data science. To explore more about our past projects, please, check my notebooks here: \n- [NFL Big Data Bowl by Practicum](https://www.kaggle.com/code/ernestglukhov/nfl-big-data-bowl-by-practicum)\n- [IceCube by Practicum](https://www.kaggle.com/code/ernestglukhov/icecube-by-practicum-all-notebooks)\n\nIn this report, we'll provide insights into the methodology, roles assumed by the students, and the expected outcomes from this project.\n\n## Project Insight:\n\nThe focal point of our current endeavor is to develop a machine learning model to detect Freezing of Gait (FOG), a severe symptom common among Parkinson's disease patients. This model will be trained on data sourced from wearable 3D sensors placed on the lower back. This significant undertaking promises to enhance our understanding of FOG episodes, thereby enabling medical professionals to evaluate, monitor, and ideally, prevent such distressing events.\n\n## Team Structure and Strategy:\n\nOur team is an amalgamation of seasoned students from previous projects and enthusiastic newcomers all set to apply their nascent data science skills. Each participant is given the freedom to navigate their project within the boundaries of the Kaggle competition, thereby encouraging a variety of approaches and fostering a rich, innovative environment.\n\nOur weekly meetings are platforms where students showcase their progress, exchange ideas, and troubleshoot any issues that might have cropped up. Such interactions encourage learning from shared experiences and promote collaboration. As the guiding force, my role as an instructor is to steer the team in the right direction, offer necessary guidance, and provide constructive feedback.\n\n## Role of Students:\n\nIn sync with our previous projects, students are given the responsibility to craft a project within the competition's ambit. They have to conceptualize, code their ideas, and critically assess the results. Each student is accountable for the accuracy and reliability of their models, and are always motivated to seek advice and feedback from their peers and me.\n\nIn conclusion, we anticipate that this project will not only be a platform for the students to apply their data science skills in a meaningful context, but also contribute significantly to a crucial area in medical research. We eagerly await to share the fruits of our labor and remain grateful for the relentless support from the Practicum community.","metadata":{}},{"cell_type":"markdown","source":"# Literature review\n\n#### [Paper Overview: ML Methods for Parkinson’s Disease](https://www.kaggle.com/code/lyalindmitriy/paper-overview-ml-methods-for-parkinson-s-disease)\nPaper Overview: Internet of Things Technologies and Machine Learning Methods for Parkinson’s Disease Diagnosis, Monitoring and Management: A Systematic Review\" summarizes 112 studies conducted in the last decade. These studies explore the use of machine learning and IoT technologies for addressing Parkinson's Disease-related problems. The review concludes that ML models and IoT technologies have the potential to revolutionize PD diagnosis and treatment, supporting clinicians in decision-making processes and potentially reducing healthcare costs.\n\n\n# EDA\n\n#### [Parkinson's FoG: basic EDA](https://www.kaggle.com/code/averkovanika/parkinson-s-fog-basic-eda)\nThis notebook aims to perform some basic EDA to get insights about datasets and uncover patterns for event detection. So, we are using a combination of statistical measures and visualization techniques to gain a deeper understanding of the data and potentially uncover relationships and trends that may be useful in developing a machine learning model for event detection.\n\n\n#### [Mutual Information Feature Selection](https://www.kaggle.com/code/averkovanika/mutual-information-feature-selection) in TOP-10 most voted notebooks in this competition!\nIn this notebook we explore the application of Mutual Information as a feature selection method for our model development. Our goal is to leverage Mutual Information to identify the most informative predictors and eliminate redundant ones in our dataset. Throughout this notebook, we demonstrate how Mutual Information can serve as a powerful metric for feature selection, thereby reducing overfitting and enhancing the generalization performance of our model.\n\n\n#### [Parkinson's: Tasks analysis w/ feature extraction](https://www.kaggle.com/code/averkovanika/parkinson-s-tasks-analysis-w-feature-extraction)\nThe goal of this notebook is to explore the temporal patterns and relationships between preceding tasks and specific types of FoG episodes. To achieve this goal, we investigate the influence of different Task values on the occurrence and type of FoG episodes, as well as uncover other temporal patterns that can aid in predicting FoG episodes and their types using task information. By analyzing these patterns and identifying consistent task patterns preceding FoG episodes, we extract meaningful features that enhance our machine learning model.\n\n\n#### [To Concatenate or not to Concatenate?](https://www.kaggle.com/code/konstantinsamolinov/to-concatenate-or-not-to-concatenate)\nInitially, the feasibility of combining the provided data into a cohesive dataset is investigated. This requires careful analysis of the data's structure and characteristics across various sources. Merging the data could uncover new patterns and relationships, but it's vital to ensure consistency in terms of metrics, units, and scales, while also checking for potential redundancies or discrepancies. This process lays a strong foundation for subsequent analysis stages.\n\n\n#### [EDA Parkinson](https://www.kaggle.com/code/donottalk/eda-parkinson)\nDuring the Exploratory Data Analysis (EDA), key features such as Id, Subject, Visit, Medication, Time, Init, Completion, AccV, AccML, AccAP, StartHesitation, Turn, and Walking were identified and their interrelations were examined. The objective was to detect freezing of gait episodes during walking within the tdcsfog and defog datasets. To achieve this, data concerning temporal steps, accelerations across three axes, and event types (recorded in the Type column) were utilized. It's worth mentioning that only event annotations marked as true were used in the defog dataset. The tdcsfog_metadata, defog_metadata, and events metadata provided details about laboratory visits, tests conducted, and medications administered, which were instrumental in analyzing the results.\n\n\n#### [Locating steps. Wavelets -> step rate feature](https://www.kaggle.com/code/vrbaryshev/locating-steps-wavelets-step-rate-feature) in TOP-10 most voted notebooks in this competition!\nSteps are the core part of any walking. I believe and that step rate itself should be an important feature for identifying FoG and other features calculated over individual steps can be important too. However, correctly splitting the process of walking into individual steps solely with acceleration data might be a complicated task. In this article I've tried to create, thoroughly explain and discuss a reliable solution for that task.\n\n#### [Daily unsupervised learning](https://www.kaggle.com/code/dmiitroamelin/daily-unsupervised-learning)\nThis notebook contains discription of several time series analysis methods. Here we explore unsupervised tools - segmentation of series, patterns, chains and anomaly searching in daily dataset.\n\n\n#### [PD FOG look at AccV, AccML and AccP](https://www.kaggle.com/code/mikhailseregin2309/pd-fog-look-at-accv-accml-and-accp)\nIn this notebook we test hypotises about statistical significance between tDCS FOG (tdcsfog) dataset and  DeFOG (defog) dataset: we combine all data in two major datasets and look at distribution function.\n\n\n# Infrastructure\n#### [CNN skeleton](https://www.kaggle.com/code/ernestglukhov/practicum-cnn-skeleton)\nIn this notebook, a fundamental Convolutional Neural Network (CNN) model is developed as a baseline for the ongoing competition. The primary objective is to detect Freezing of Gait (FOG) episodes in Parkinson's disease patients, utilizing data from wearable 3D lower back sensors. The CNN model, renowned for its prowess in pattern detection within spatially distributed data, serves as an excellent starting point. This notebook covers the initial stages of designing, training, and testing the baseline model, setting the groundwork for further enhancements and optimizations in subsequent stages.\n\n\n# Solutions\n#### [ML to detect FOG episodes & type](https://www.kaggle.com/code/donottalk/ml-to-detect-fog-episodes-type)\nThe model leverages patient information, medication data, FOG measurements, and acceleration data to achieve this goal. We split our data into features and the target variable, followed by building a pipeline for each target variable column using an XGBClassifier with specific hyperparameters to account for distinct characteristics of each mobility impairment episode. This diverse feature set and personalized training for each target variable class improve the prediction quality, leading to personalized approaches to Parkinson's disease treatment and monitoring.\n\n\n#### [Support Vector Machine for FoG prediction](https://www.kaggle.com/code/konstantinsamolinov/support-vector-machine-for-fog-prediction)\nThis notebook show the way to use Support Vector Machine for classification. \nHere i describe basic principles of SVM, parameters of ML method and principal formulas for kernel while finding \"best hypeprlane\" of target data.\nShould be careful to use SVM, because it's really long time to calculate.\n\n#### [FoG: F for Forest](https://www.kaggle.com/code/alexanderskachkov/fog-f-for-forest)\nIn this notebook, the design of basic classification models of various types of the FoG based on a random forest model trained on various samples is carried out. A study was performed on the correlation of main features of patients on the appearance of the FoG . The Train attribute was selected as the attribute that determines the division into training samples. A function has also been prepared that excludes the double manifestation of an event at one time. \n\n#### [CatBoost vs. LightGBM vs. XGBoost](https://www.kaggle.com/code/evgeniidvornikov/catboost-vs-lightgbm-vs-xgboost)\n\nIn this notebook gradient boosting models XGBoost, Light GBM, Cat Boots  are exploring. The result of this research is to find the most perspective gradient boosting models for FoG prediction.\n\n\n#### [Final 1D-CNN solution](https://www.kaggle.com/code/ernestglukhov/practicum-final-1d-cnn-solution)\nIn this Kaggle notebook, we present a comprehensive approach aimed at predicting Freezing of Gait (FoG) using an ensemble of 1D-CNN models and a sophisticated segmentation process. The segmentation process helps capture distinct patterns in the acceleration data, while the ensemble strategy involves training multiple 1D-CNN models using K-fold cross-validation. Our model architecture consists of multiple Convolutional Blocks, each containing convolutional layers with different kernel sizes and dilations. The models in the ensemble are trained separately on the \"tdcsfog\" and \"defog\" datasets. By treating the tdcsfog and defog datasets separately, we aimed to optimize the models' performance for each specific type of data.","metadata":{}}]}