{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**※このノートは主に日本語で作成しています**","metadata":{}},{"cell_type":"markdown","source":"# 概要\n\n**Parkinson's FOG Prediction**\n\n- パーキンソン病の症状の一つであるFOG（すくみ足）を予測する事が目的\n  - [パーキンソン病の歩行障害を定量的に評価する方法](https://www.jstage.jst.go.jp/article/rigaku/advpub/0/advpub_11047/_article/-char/ja/)\n- 患者の腰部に装着した加速度センサー情報を元に推定を行う\n- 教師データは録画映像から専門家がAnnotationを行い作成した。\n- データの収集はLabとHomeで実施し、それぞれデータの形式が異なる","metadata":{}},{"cell_type":"code","source":"import os\nimport random\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\ncmap = plt.get_cmap(\"tab10\")\n\ninput_dir = Path(\"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction\")\n\ndaily_metadata = pd.read_csv(input_dir / 'daily_metadata.csv')\ndefog_metadata = pd.read_csv(input_dir / 'defog_metadata.csv')\ntdcsfog_metadata = pd.read_csv(input_dir / 'tdcsfog_metadata.csv')\nevents = pd.read_csv(input_dir / 'events.csv')\ntasks = pd.read_csv(input_dir / 'tasks.csv')\nsubjects = pd.read_csv(input_dir / 'subjects.csv')\nsample_submission = pd.read_csv(input_dir / 'sample_submission.csv')\n\n# 患者のID\nsubject_ids = subjects.Subject.unique()\n\nmetadatas = {\n    'daily_metadata' : daily_metadata, \n    'defog_metadata' : defog_metadata, \n    'tdcsfog_metadata' : tdcsfog_metadata,\n}\nindexes = {\n    'events' : events, \n    'tasks' : tasks, \n    'sample_submission' : sample_submission, \n}","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:22:45.457899Z","iopub.execute_input":"2023-05-07T13:22:45.458239Z","iopub.status.idle":"2023-05-07T13:22:45.982724Z","shell.execute_reply.started":"2023-05-07T13:22:45.458212Z","shell.execute_reply":"2023-05-07T13:22:45.981597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# どのようなデータがあるか\n\n### 重要な識別\n\n- `Id` : 各記録に紐づくID\n- `Subject` : 各患者に紐づくID\n\n### 配布された時系列データ\n\n- 各時系列データは\"Id\"毎に`train/{datatype}/{Id}.csv`として保存されている\n  - **datatype**は下記のように分類される\n  - tdcsfog : Labデータ\n  - defog: Homeデータ\n  - notype: Annotationしていないデータ","metadata":{}},{"cell_type":"code","source":"!tree -L 2 /kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-05-07T13:22:45.984435Z","iopub.execute_input":"2023-05-07T13:22:45.984740Z","iopub.status.idle":"2023-05-07T13:22:47.113885Z","shell.execute_reply.started":"2023-05-07T13:22:45.984714Z","shell.execute_reply":"2023-05-07T13:22:47.112513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"csv_files = {}\nfor key1 in ['train', 'test']:\n    for key2 in ['defog', 'tdcsfog', 'notype']:\n        files = [f for f in (input_dir / key1 / key2).glob('*.csv')]\n        if files:\n            print(f'{key1}/{key2}: {len(files)} files')\n            csv_files[(key1, key2)] = files\n            \nparquet_files = [f for f in (input_dir / 'unlabeled').glob('*.parquet')]\nprint(f'unlabeled: {len(parquet_files)} files')","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:22:47.118746Z","iopub.execute_input":"2023-05-07T13:22:47.119117Z","iopub.status.idle":"2023-05-07T13:22:47.233801Z","shell.execute_reply.started":"2023-05-07T13:22:47.119086Z","shell.execute_reply":"2023-05-07T13:22:47.232686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![image.png](https://inb195734681.s3.ap-northeast-1.amazonaws.com/nvkoslhoub5291.png)","metadata":{}},{"cell_type":"markdown","source":"## 各CSVの内容\n\n- 時系列データ(`train/*/*.csv`)\n  - AccV,AccML,AccAP: **説明変数**, 各時刻に置けるセンサーの加速度\n  - StartHesitation, Turn, Walking: **目的変数**, Annotationされた行動\n  - Event, Valid, Task: Annotation前の情報\n- 補足情報\n  - \\*_metadata.csv: 各データの予備情報（投薬の有無など）\n  - subjects.csv: 患者情報\n  - events.csv: ビデオ判定の結果\n  - tasks.csv: 患者に与えられた指示","metadata":{}},{"cell_type":"markdown","source":"# 何を推定して提出するか\n\n- {Id}_{Time}の各Annotationを推定する\n- Notebookを提出し、test/*/*.csvに格納されたデータを用いて推定を実施する","metadata":{}},{"cell_type":"code","source":"sample_submission","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:22:47.237122Z","iopub.execute_input":"2023-05-07T13:22:47.237573Z","iopub.status.idle":"2023-05-07T13:22:47.269174Z","shell.execute_reply.started":"2023-05-07T13:22:47.237537Z","shell.execute_reply":"2023-05-07T13:22:47.268270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n\n# 各データの可視化\n\n各データを下記のように確認することができる","metadata":{}},{"cell_type":"markdown","source":"![image.png](https://inb195734681.s3.ap-northeast-1.amazonaws.com/hgoiurehoi1841590.png)","metadata":{}},{"cell_type":"code","source":"# csvファイルの一覧を作成\ndatas_ = {'train':{}, 'test':{}}\nfor sample_id in subject_ids:\n    for key1, metadata in metadatas.items():\n        df = metadata[metadata.Subject==sample_id]\n        for ind in df.index:\n            id_ = df.Id[ind]\n            for train_test in ['train', 'test']:\n                for key2 in ['defog', 'tdcsfog', 'notype']:\n                    f = input_dir / train_test / key2 / f'{id_}.csv'\n                    if f.is_file():\n                        if not id_ in datas_[train_test].keys():\n                            datas_[train_test][id_] = f\n                        else:\n                            print('Duplicated')\n                            print(datas_[train_test][id_], 'vs.', f)\n                            \ndef make_fig(id_, test=False):\n    \"\"\"各測定IDに関する情報を描画\"\"\"\n    file_name = datas_['train'][id_] if not test else datas_['test'][id_]\n    #print(file_name)\n    \n    df_pulse = pd.read_csv(file_name, index_col=0)\n    # tdcsfogのみ128Hzのデータであることに注意\n    metadata_key = file_name.as_posix().split('/')[-2]\n    Hz = 128 if metadata_key=='tdcsfog' else 100\n    df_pulse.index /= Hz\n    \n    fig, axs = plt.subplots(4, 1, figsize=(10,8), sharex=True)\n\n    for i, key in enumerate(['AccV', 'AccML', 'AccAP']):\n        axs[i].plot(df_pulse.index, df_pulse[key], label=key, lw=.5, c=cmap(i))\n        axs[i].legend(loc='upper left', fontsize=8)\n\n    # ４つ目の表ではBoolを描画\n    ax=axs[3]\n    lines = ['StartHesitation', 'Turn', 'Walking', 'Event', 'Valid', 'Task']\n    for n, name in enumerate(lines):\n        y = 6-n\n        ax.axhline(y, c=cmap(n), label=name, lw=.5)\n        if name in df_pulse.columns:\n            df_tmp = df_pulse[df_pulse[name].astype(int)>0]\n            ax.scatter(df_tmp.index, [y]*len(df_tmp), c='k', s=3, marker='|')\n    ax.axhline(0, c='k', label='Task_metadata', lw=.5)\n    \n    usr_info = None\n    if metadata_key != 'notype':\n        df_metadata = metadatas[f'{metadata_key}_metadata']\n        ser_metadata = df_metadata[df_metadata.Id==id_].iloc[0]\n\n        Subject = ser_metadata.Subject\n        Medication_ = 'Med:On' if ser_metadata.Medication else 'Med:Off'\n        Visit = ser_metadata.Visit\n        Test_ = '?'\n        if 'Test' in ser_metadata.index:\n            Test_ = ser_metadata.Test\n\n        df_events = events[events.Id==id_]\n        df_tasks = tasks[tasks.Id==id_]\n\n        ser_subject = subjects[(subjects.Subject==Subject)]\n        if len(ser_subject) == 1:\n            ser_subject = ser_subject.iloc[0]\n        else:\n            ser_subject = ser_subject[ser_subject.Visit==Visit].iloc[0]\n\n        def get_int(s):\n            try:\n                return str(int(s))\n            except:\n                return '?'\n            \n        # ユーザー情報を取得\n        try:\n            Subject_ = ser_subject.Subject\n            Visit_ = get_int(ser_subject.Visit)\n            Age_ = get_int(ser_subject.Age)\n            Sex_ = ser_subject.Sex\n            YearsSinceDx_ = get_int(ser_subject.YearsSinceDx)\n            UPDRSIII_On_ = get_int(ser_subject.UPDRSIII_On)\n            UPDRSIII_Off_ = get_int(ser_subject.UPDRSIII_Off)\n            NFOGQ_ = get_int(ser_subject.NFOGQ)\n            usr_info = f'{Subject_}({Age_}yrs/Sex:{Sex_}/Visit:{Visit_}/Test:{Test_}/{Medication_})/snc:{YearsSinceDx_}/On:{UPDRSIII_On_}/Off:{UPDRSIII_Off_}/NFOGQ:{NFOGQ_}'\n        except:\n            pass\n        \n        # events.csvからKinetic/Akineticを取得描画\n        for ind in df_events.index:\n            ser = df_events.loc[ind]\n            start = ser.Init\n            end = ser.Completion\n            try:\n                y = 6-['StartHesitation', 'Turn', 'Walking'].index(ser.Type)\n                ec='r' if ser.Kinetic else 'k'\n                ax.plot((start,end), (y-.2,y-.2), c=ec, lw=1)\n                ax.plot((start,end), (y+.3,y+.3), c=ec, lw=1)\n            except:\n                pass\n\n        # tasks.csvから患者指示行動を取得描画\n        for ind in df_tasks.index:\n            ser = df_tasks.loc[ind]\n            start = ser.Begin\n            end = ser.End\n            y=0\n            ax.plot([start,end], [y,y], c='r', lw=5)\n            ax.annotate(ser.Task, [start,y-.5], rotation=45, fontsize=6)\n\n    ax.legend(loc='upper left', fontsize=5)\n    ax.set_ylim(-1, 7)\n    ax.set_xlabel('Time [s]')\n    ax.set_yticks([])\n\n    title = '/'.join(file_name.as_posix().split('/')[-3:])\n    if usr_info:\n        title = f'{title} | {usr_info}'\n    fig.suptitle(title, fontsize=8)\n    fig.tight_layout(h_pad=0.95)\n\n    return fig, ax, file_name, df_pulse","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-07T13:22:47.270888Z","iopub.execute_input":"2023-05-07T13:22:47.271213Z","iopub.status.idle":"2023-05-07T13:22:50.691352Z","shell.execute_reply.started":"2023-05-07T13:22:47.271185Z","shell.execute_reply":"2023-05-07T13:22:50.690211Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 全てのCSVに対して下記描画処理を実行","metadata":{}},{"cell_type":"code","source":"# !mkdir /kaggle/working/pics","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:22:50.695178Z","iopub.execute_input":"2023-05-07T13:22:50.696462Z","iopub.status.idle":"2023-05-07T13:22:51.794195Z","shell.execute_reply.started":"2023-05-07T13:22:50.696419Z","shell.execute_reply":"2023-05-07T13:22:51.792680Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# event_ids_train = list(datas_['train'].keys())\n# \n# for sample_id in tqdm(event_ids_train):\n#     fig, ax, file_name, df_pulse = make_fig(sample_id)\n#     key = file_name.as_posix().split('/')[-2]\n#     fig.savefig(f'/kaggle/working/pics/{key}_{id_}.png')\n#     plt.clf()\n#     plt.close()","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-05-07T13:22:51.796755Z","iopub.execute_input":"2023-05-07T13:22:51.797223Z","iopub.status.idle":"2023-05-07T13:36:58.504276Z","shell.execute_reply.started":"2023-05-07T13:22:51.797177Z","shell.execute_reply":"2023-05-07T13:36:58.503250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# !cd /kaggle/working/ && tar -czf output_images.tgz pics/","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2023-05-07T13:36:58.506103Z","iopub.execute_input":"2023-05-07T13:36:58.506705Z","iopub.status.idle":"2023-05-07T13:36:59.777093Z","shell.execute_reply.started":"2023-05-07T13:36:58.506671Z","shell.execute_reply":"2023-05-07T13:36:59.775573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n---\n\n# 各分類ごとに確認\n\n---\n\n## データ例①: train/tdcsfog\n\n- 833 csv files\n- Labにおける測定データ","metadata":{}},{"cell_type":"code","source":"sample_id = csv_files[('train', 'tdcsfog')][0].name.split('.')[0]\nfig, ax, file_name, df_pulse = make_fig(sample_id)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:36:59.779721Z","iopub.execute_input":"2023-05-07T13:36:59.780210Z","iopub.status.idle":"2023-05-07T13:37:01.305469Z","shell.execute_reply.started":"2023-05-07T13:36:59.780167Z","shell.execute_reply":"2023-05-07T13:37:01.304366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pulse","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:01.309516Z","iopub.execute_input":"2023-05-07T13:37:01.310271Z","iopub.status.idle":"2023-05-07T13:37:01.327352Z","shell.execute_reply.started":"2023-05-07T13:37:01.310229Z","shell.execute_reply":"2023-05-07T13:37:01.326170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## データ例②: train/defog\n\n- 91 csv files\n- 患者の自宅における観測データ","metadata":{}},{"cell_type":"code","source":"sample_id = csv_files[('train', 'defog')][0].name.split('.')[0]\nfig, ax, file_name, df_pulse = make_fig(sample_id)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:01.329054Z","iopub.execute_input":"2023-05-07T13:37:01.329543Z","iopub.status.idle":"2023-05-07T13:37:03.319318Z","shell.execute_reply.started":"2023-05-07T13:37:01.329507Z","shell.execute_reply":"2023-05-07T13:37:03.318258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pulse","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:03.320927Z","iopub.execute_input":"2023-05-07T13:37:03.321310Z","iopub.status.idle":"2023-05-07T13:37:03.338509Z","shell.execute_reply.started":"2023-05-07T13:37:03.321260Z","shell.execute_reply":"2023-05-07T13:37:03.337791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## データ例③: train/notype\n\n- 46 csv files\n- 未分類データ","metadata":{}},{"cell_type":"code","source":"sample_id = csv_files[('train', 'notype')][0].name.split('.')[0]\nfig, ax, file_name, df_pulse = make_fig(sample_id)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:03.339802Z","iopub.execute_input":"2023-05-07T13:37:03.340237Z","iopub.status.idle":"2023-05-07T13:37:05.662133Z","shell.execute_reply.started":"2023-05-07T13:37:03.340212Z","shell.execute_reply":"2023-05-07T13:37:05.661431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pulse","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:05.663117Z","iopub.execute_input":"2023-05-07T13:37:05.663879Z","iopub.status.idle":"2023-05-07T13:37:05.679678Z","shell.execute_reply.started":"2023-05-07T13:37:05.663850Z","shell.execute_reply":"2023-05-07T13:37:05.678648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"daily_metadata[daily_metadata.Id==sample_id]","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:05.680910Z","iopub.execute_input":"2023-05-07T13:37:05.681197Z","iopub.status.idle":"2023-05-07T13:37:05.692461Z","shell.execute_reply.started":"2023-05-07T13:37:05.681172Z","shell.execute_reply":"2023-05-07T13:37:05.691774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## データ例④: test/tdcsfog\n\n- 1 csv files\n- Labデータ","metadata":{}},{"cell_type":"code","source":"sample_id = csv_files[('test', 'tdcsfog')][0].name.split('.')[0]\nfig, ax, file_name, df_pulse = make_fig(sample_id, test=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:05.693637Z","iopub.execute_input":"2023-05-07T13:37:05.694379Z","iopub.status.idle":"2023-05-07T13:37:06.613564Z","shell.execute_reply.started":"2023-05-07T13:37:05.694351Z","shell.execute_reply":"2023-05-07T13:37:06.612453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## データ例⑤: test/defog\n\n- 1 csv files\n- 自宅データ","metadata":{}},{"cell_type":"code","source":"sample_id = csv_files[('test', 'defog')][0].name.split('.')[0]\nfig, ax, file_name, df_pulse = make_fig(sample_id, test=True)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:06.615176Z","iopub.execute_input":"2023-05-07T13:37:06.615610Z","iopub.status.idle":"2023-05-07T13:37:08.512275Z","shell.execute_reply.started":"2023-05-07T13:37:06.615577Z","shell.execute_reply":"2023-05-07T13:37:08.511271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_pulse","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:08.513722Z","iopub.execute_input":"2023-05-07T13:37:08.514199Z","iopub.status.idle":"2023-05-07T13:37:08.528568Z","shell.execute_reply.started":"2023-05-07T13:37:08.514168Z","shell.execute_reply":"2023-05-07T13:37:08.527650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n\n# 追加データ parquet ファイル\n\n- 65 files\n- このうち45人の患者はtrain/defog/にデータを持ち、残りはtestに割り振られている","metadata":{}},{"cell_type":"code","source":"ids_in_csv     = [f.name.split('.')[0] for f in csv_files[('test', 'defog')]]\nids_in_parquet = [f.name.split('.')[0] for f in parquet_files]\n\nf_parquet = parquet_files[0]\ndf_parquet = pd.read_parquet(f_parquet)\nprint(f_parquet)\nprint(df_parquet.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:08.529967Z","iopub.execute_input":"2023-05-07T13:37:08.530339Z","iopub.status.idle":"2023-05-07T13:37:29.536699Z","shell.execute_reply.started":"2023-05-07T13:37:08.530313Z","shell.execute_reply":"2023-05-07T13:37:29.535636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# for f_parquet in parquet_files:\n#     df_parquet = pd.read_parquet(f_parquet)\n#     sample_id = f_parquet.name.split('.')[0]\n#     ser = daily_metadata[daily_metadata.Id==sample_id].iloc[0]\n#     ser_subject = subjects[subjects.Subject==ser.Subject].iloc[0]\n#     print(f_parquet.name, ser_subject)\n#     break","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:29.538404Z","iopub.execute_input":"2023-05-07T13:37:29.539184Z","iopub.status.idle":"2023-05-07T13:37:34.018918Z","shell.execute_reply.started":"2023-05-07T13:37:29.539144Z","shell.execute_reply":"2023-05-07T13:37:34.018152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Kaggle Notebookで描画するには大きい\n# df_parquet.set_index('Time').plot()","metadata":{"execution":{"iopub.status.busy":"2023-05-07T13:37:34.020597Z","iopub.execute_input":"2023-05-07T13:37:34.020972Z","iopub.status.idle":"2023-05-07T13:37:34.025120Z","shell.execute_reply.started":"2023-05-07T13:37:34.020934Z","shell.execute_reply":"2023-05-07T13:37:34.024315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baselines (Editing)\n\n- [Simple LGBM multi-class classification baseline (by DATAMANYO)](https://www.kaggle.com/code/kimtaehun/simple-lgbm-multi-class-classification-baseline)\n  - 前処理(defog)\n    - (Task==1)&(Valid==1)のみを使用\n    - metadataをcombine\n    - columns-> Id,Subject,Visit,Medication,Time,AccV,AccML,AccAP,StartHesitation,Turn,Walking\n  - 前処理(tdcsfog)\n    - 全部combine\n  - FeatureEngineering\n    - reading","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}