{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"d4aba448-69fd-4d47-ba29-7e3eed678ab9","_cell_guid":"d631d75e-5818-41d6-9dc0-831c8639ffdf","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-03-29T08:18:48.659776Z","iopub.execute_input":"2023-03-29T08:18:48.660139Z","iopub.status.idle":"2023-03-29T08:18:48.894436Z","shell.execute_reply.started":"2023-03-29T08:18:48.660101Z","shell.execute_reply":"2023-03-29T08:18:48.893240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install sklearn","metadata":{"execution":{"iopub.status.busy":"2023-03-29T08:18:48.896183Z","iopub.execute_input":"2023-03-29T08:18:48.896571Z","iopub.status.idle":"2023-03-29T08:19:01.231052Z","shell.execute_reply.started":"2023-03-29T08:18:48.896541Z","shell.execute_reply":"2023-03-29T08:19:01.229510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![DdataSet](https://file.notion.so/f/s/c7130637-1f40-451c-862d-5c61f2b9fb2d/Dataset.png?id=932dacb4-e2f1-405e-94c5-3d1060ca822a&table=block&spaceId=f3942938-2f00-4a70-a117-b2579a423ac8&expirationTimestamp=1680150872126&signature=l9JEC0XSjt0VKCrtCuwprPeclxtUwszWq_SqYHTZdEo&downloadName=Dataset.png)","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\nimport pandas as pd\n\n# Load the data\n# df = pd.read_csv('003f117e14.csv')\ntrain_tdcsfog_example_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog/003f117e14.csv\")\ndf = train_tdcsfog_example_df\n\n# Plot the accelerometerS dataS\nfig = px.line(df, x='Time', y=['AccV', 'AccML', 'AccAP'],\n              title='Accelerometer Data')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:07:31.879959Z","iopub.execute_input":"2023-03-29T11:07:31.880315Z","iopub.status.idle":"2023-03-29T11:07:31.957288Z","shell.execute_reply.started":"2023-03-29T11:07:31.880283Z","shell.execute_reply":"2023-03-29T11:07:31.956360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\nscaler = StandardScaler()\ntrain_tdcsfog_example_df[['AccV', 'AccML', 'AccAP']] = scaler.fit_transform(train_tdcsfog_example_df[['AccV', 'AccML', 'AccAP']])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:08:25.575530Z","iopub.execute_input":"2023-03-29T11:08:25.575917Z","iopub.status.idle":"2023-03-29T11:08:25.587838Z","shell.execute_reply.started":"2023-03-29T11:08:25.575881Z","shell.execute_reply":"2023-03-29T11:08:25.586725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = train_tdcsfog_example_df\n\n# Plot the accelerometer data\nfig = px.line(df, x='Time', y=['AccV', 'AccML', 'AccAP'],\n              title='Accelerometer Data')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:08:27.256046Z","iopub.execute_input":"2023-03-29T11:08:27.256414Z","iopub.status.idle":"2023-03-29T11:08:27.332607Z","shell.execute_reply.started":"2023-03-29T11:08:27.256382Z","shell.execute_reply":"2023-03-29T11:08:27.331516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape[0] # tDCS data size == len(df)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T10:04:11.353534Z","iopub.execute_input":"2023-03-29T10:04:11.353896Z","iopub.status.idle":"2023-03-29T10:04:11.361388Z","shell.execute_reply.started":"2023-03-29T10:04:11.353868Z","shell.execute_reply":"2023-03-29T10:04:11.359948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":" df.shape[1] # tDCS data features","metadata":{"execution":{"iopub.status.busy":"2023-03-29T08:20:03.641241Z","iopub.execute_input":"2023-03-29T08:20:03.641634Z","iopub.status.idle":"2023-03-29T08:20:03.649020Z","shell.execute_reply.started":"2023-03-29T08:20:03.641600Z","shell.execute_reply":"2023-03-29T08:20:03.647586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.describe # 0인 값들만 보인다. 다른 값이 있는 부분들만 보고 싶음","metadata":{"execution":{"iopub.status.busy":"2023-03-29T08:21:08.571360Z","iopub.execute_input":"2023-03-29T08:21:08.571768Z","iopub.status.idle":"2023-03-29T08:21:08.583346Z","shell.execute_reply.started":"2023-03-29T08:21:08.571731Z","shell.execute_reply":"2023-03-29T08:21:08.582270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.where(df[\"AccV\"] != 0) # 0은 그냥 놔두면 될 거 같긴한데 ","metadata":{"execution":{"iopub.status.busy":"2023-03-29T08:32:51.381301Z","iopub.execute_input":"2023-03-29T08:32:51.381677Z","iopub.status.idle":"2023-03-29T08:32:51.402167Z","shell.execute_reply.started":"2023-03-29T08:32:51.381648Z","shell.execute_reply":"2023-03-29T08:32:51.400844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filter data\n# df = df[(df['AccV'] > -2) & (df['AccV'] < 2) & (df['AccML'] > -2) & (df['AccML'] < 2) & (df['AccAP'] > 8)]\n\n# Segment data\nsegments = []\nfor i in range(0, len(df) - 100, 50):\n    x = df['AccV'].values[i:i+100]\n    y = df['AccML'].values[i:i+100]\n    z = df['AccAP'].values[i:i+100]\n    segments.append([x, y, z])\n\n# Extract features\nfeatures = []\nfor segment in segments:\n    x = segment[0]\n    y = segment[1]\n    z = segment[2]\n    features.append([np.mean(x), np.mean(y), np.mean(z), np.std(x), np.std(y), np.std(z), np.max(x), np.max(y), np.max(z), np.min(x), np.min(y), np.min(z)])\n\n# Convert features to dataframe\nfeatures_df = pd.DataFrame(features, columns=['mean_x', 'mean_y', 'mean_z', 'std_x', 'std_y', 'std_z', 'max_x', 'max_y', 'max_z', 'min_x', 'min_y', 'min_z'])","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:23:30.748197Z","iopub.execute_input":"2023-03-29T11:23:30.748692Z","iopub.status.idle":"2023-03-29T11:23:30.771794Z","shell.execute_reply.started":"2023-03-29T11:23:30.748650Z","shell.execute_reply":"2023-03-29T11:23:30.770589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier\n\n# Load preprocessed accelerometer data\n# features_df = pd.read_csv('features.csv')\ntrain_tdcsfog_example_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog/003f117e14.csv\")\n\ntest_tdcsfog_example_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/tdcsfog/003f117e14.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:25:55.280740Z","iopub.execute_input":"2023-03-29T11:25:55.281149Z","iopub.status.idle":"2023-03-29T11:25:55.304372Z","shell.execute_reply.started":"2023-03-29T11:25:55.281113Z","shell.execute_reply":"2023-03-29T11:25:55.302643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"train과 test가 이미 구분되어져 있고 안의 내용물이 다른데 어떻게 하지 train에 있는 다른 내용물들을 다 버리면 되는걸까\n아니 그냥 test가 의미없는 거 아님??","metadata":{}},{"cell_type":"code","source":"# Load preprocessed accelerometer data\ntrain_tdcsfog_example_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/tdcsfog/003f117e14.csv\")\ntest_tdcsfog_example_df = pd.read_csv(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/test/tdcsfog/003f117e14.csv\")\n\n# Split data into training and testing sets\nX_train, X_test, y_train, y_test = train_test_split(train_tdcsfog_example_df.iloc[:, :-1], train_tdcsfog_example_df.iloc[:, -1], test_size=0.2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train a random forest classifier\nclf = RandomForestClassifier(n_estimators=100)\nclf.fit(X_train, y_train)\n\n# Evaluate the classifier on testing set\nscore = clf.score(X_test, y_test)\nprint('Accuracy:', score)","metadata":{"execution":{"iopub.status.busy":"2023-03-29T11:31:14.664368Z","iopub.execute_input":"2023-03-29T11:31:14.664764Z","iopub.status.idle":"2023-03-29T11:31:14.816633Z","shell.execute_reply.started":"2023-03-29T11:31:14.664729Z","shell.execute_reply":"2023-03-29T11:31:14.815455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Collect accelerometer data from the subjects. You can use a smartphone or wearable device to collect the data.\n\nPreprocess the data by filtering, segmenting, and extracting features. You can use Python libraries such as Pandas, Numpy, and Scipy to preprocess the data.\n\nTrain a machine learning model to predict freezing of gait using the preprocessed data. You can use Python libraries such as Scikit-learn, Keras, and Tensorflow to train the model.\n\nHere is an example code for preprocessing accelerometer data using Python Pandas library:","metadata":{}},{"cell_type":"markdown","source":"In this context, deFOG data refers to data recordings from subjects who wore a 3D accelerometer on their lower back during motor assessments at their home environment. The protocol included two visits at Off and On medication states where participants were evaluated. During these visits, participants performed various tasks such as walking tests and turning tasks. The acceleration units are provided in [g].\n\n","metadata":{}},{"cell_type":"markdown","source":"Observe your dataset at a high level. Start by determining the size of your dataset, including how many rows and columns it has. This can help you predict any future issues you might have with your data.\n\nCheck for missing values in your dataset. Missing values can cause issues with your analysis, so it’s important to identify them early on.\n\nCheck for outliers in your dataset. Outliers can skew your analysis and make it difficult to draw meaningful conclusions.\n\nVisualize your data using histograms, scatter plots, and other graphs. This can help you identify patterns and relationships in your data.\n\nCalculate summary statistics such as mean, median, mode, standard deviation, and variance for each variable in your dataset.\n\nUse correlation analysis to identify relationships between variables in your dataset","metadata":{}}]}