{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"比赛结束之后的复盘，只是自己看看，把用于理解的注释分享出来（注意只有中文注释😅）\n\nAfter the competition, the post-mortem is just for self-reference. I will share the comments in Chinese that are used for understanding.","metadata":{}},{"cell_type":"markdown","source":"## infoes\nFeature Column TimeSeries grouping adapted from https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\n\nSubject-wise GroupKFold splitting adapted from https://www.kaggle.com/code/xzj19013742/groupkfold-cross-validation-tsflex\n\n\n[原始notebook链接](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model)\n\n### In this Notebook\n- Tensorflow Model with Conv1D blocks \n- Models trained separately for defog and tdcsfog data\n- Event Stratified Subject Grouped KFold splitting\n\n可以了解到的：\n1. test文件夹中也给出了整个csv文件。说明这个预测不是实时的。不需要考虑“因果”这块的问题。\n2. test数据集中，tdcsfog和defog文件也是分开的。这个方案也是分开建模分开预测。不需要对数据集进行合并\n","metadata":{}},{"cell_type":"markdown","source":"### 自己的记录\n\nkaggle_PkFoG3 这块的查看。对应方案:\n1. [52nd Place Solution + Code](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416797)\n2. [57th place solution: Conv1d model with aux head](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416100)\n\n这个baseline的处理：\n1. 没有使用其余任何额外的字段等\n2. 只使用defog数据集的 `Valid` 与 `Task` 标记为1的数据集，这点应该是遵循[evaluation](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation)中的说明。\n3. 这里将一个类型的所有数据均拼接成一个\n    1. 随机筛选_idx 作为batch数据中筛选的起点，\n    2. 随后进一步基于每个idx点 past_future 跳着选时间点作为样本时间点，每条数据的总时间长度是 windows_size\n4. 每个类型基于subject字段作分组标准，分10折训练模型，后续预测结果做均值输出。\n5. 训练过程中\"手动实现早停\"，每折fold训练全部的X个epoch，但会选择其中得分最高的那个epoch对应模型保存下来\n\n关于baseline中的模型：\n1. 这个模型本质上就是by时间帧分别预测的一个分类模型。从整体流程角度，其实就是和lgbm是一样的。\n    1. 注意到在`FOGSequence` 类的 `__getitem__` 中，若 `self.split` 设定为 'test'，那么给到模型预测的数据是顺序给到的。\n2. 线上模型训练时间: defog done in 2.78 hrs; tdcsfog done in 4.53 hrs\n\n其他可以了解的：\n1. 作者指出 `AveragePrecision` 和 `calculate_precision` 这两个函数计算结果并不一致，前者用于给到TensorFlow的计算，后者用于选择最佳模型的依据（数值不一样，但趋势一样的）\n    1. **按照官方的说法，他们应该以`calculate_precision`为准**，https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation\n2. 不同的kfold之间，val_avg_precision之间的差异非常大，这可能是由于样本标签不平衡导致（defog这类数据的 只有turn标签会多一些）\n3. 这个baseline，使用sub作为skf的切分标准，从线下线上数据分布一致性的角度，是否是合适的？这个讨论中也是使用subject作为切分，https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395391\n\n\n","metadata":{}},{"cell_type":"markdown","source":"### 其他可以探索的\n\n1. 更多的模型：https://keras.io/examples/timeseries/timeseries_transformer_classification/\n2. 不平衡问题的讨论：https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/406139","metadata":{}},{"cell_type":"markdown","source":"## Imports and Config","metadata":{}},{"cell_type":"code","source":"IF_KAGGLE = True","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.125595Z","iopub.execute_input":"2023-10-12T00:37:36.126299Z","iopub.status.idle":"2023-10-12T00:37:36.131005Z","shell.execute_reply.started":"2023-10-12T00:37:36.126265Z","shell.execute_reply":"2023-10-12T00:37:36.129686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Constants\nif IF_KAGGLE:\n    BASE_DIR = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction\"\nelse:\n    BASE_DIR = \"\"","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.136263Z","iopub.execute_input":"2023-10-12T00:37:36.136927Z","iopub.status.idle":"2023-10-12T00:37:36.142502Z","shell.execute_reply.started":"2023-10-12T00:37:36.136890Z","shell.execute_reply":"2023-10-12T00:37:36.141185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport numpy as np\nfrom numpy.random import default_rng\nimport pandas as pd\nfrom tqdm.auto import tqdm\nfrom glob import glob\nfrom os.path import basename, dirname, join, exists\nfrom time import perf_counter\nfrom collections import defaultdict as dd\nfrom functools import partial\n\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, StratifiedGroupKFold\nfrom sklearn.metrics import average_precision_score\nfrom sklearn.preprocessing import StandardScaler as Scaler\nfrom scipy.special import expit\n\nimport tensorflow as tf\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-12T00:37:36.144443Z","iopub.execute_input":"2023-10-12T00:37:36.145436Z","iopub.status.idle":"2023-10-12T00:37:36.155919Z","shell.execute_reply.started":"2023-10-12T00:37:36.145398Z","shell.execute_reply":"2023-10-12T00:37:36.154920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"TF version: {tf.__version__}\")\nAUTO = tf.data.experimental.AUTOTUNE","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.157254Z","iopub.execute_input":"2023-10-12T00:37:36.158304Z","iopub.status.idle":"2023-10-12T00:37:36.170526Z","shell.execute_reply.started":"2023-10-12T00:37:36.158269Z","shell.execute_reply":"2023-10-12T00:37:36.169398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TRAIN_DIR = join(BASE_DIR, \"train\")\nTEST_DIR = join(BASE_DIR, \"test\")\n\nIS_PUBLIC = len(glob(join(TEST_DIR, \"*/*.csv\")))==2","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.173321Z","iopub.execute_input":"2023-10-12T00:37:36.174212Z","iopub.status.idle":"2023-10-12T00:37:36.182684Z","shell.execute_reply.started":"2023-10-12T00:37:36.174176Z","shell.execute_reply":"2023-10-12T00:37:36.181601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Config:\n    train_sub_dirs = [\n        join(TRAIN_DIR, \"defog\"),\n        join(TRAIN_DIR, \"tdcsfog\")]\n    \n    metadata_paths = [\n        join(BASE_DIR, \"defog_metadata.csv\"),\n        join(BASE_DIR, \"tdcsfog_metadata.csv\")]\n    \n    splits = 10\n\n    batch_size = 1024\n    window_size = 64\n    window_future = 16\n    window_past = window_size - window_future # Includes current value\n    \n    wx = 8\n    \n    model_dropout = 0.2\n    model_hidden = 128\n    model_nblocks = 3\n    \n    lr = 0.00015\n    num_epochs = 5\n    \n    feature_list = ['AccV', 'AccML', 'AccAP']  # 特征feas\n    label_list = ['StartHesitation', 'Turn', 'Walking']  # y值\n    \n    n_features = len(feature_list)\n    n_labels = len(label_list)    \n    ","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.184089Z","iopub.execute_input":"2023-10-12T00:37:36.185011Z","iopub.status.idle":"2023-10-12T00:37:36.192978Z","shell.execute_reply.started":"2023-10-12T00:37:36.184977Z","shell.execute_reply":"2023-10-12T00:37:36.192103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cfg = Config()","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.195877Z","iopub.execute_input":"2023-10-12T00:37:36.196810Z","iopub.status.idle":"2023-10-12T00:37:36.203247Z","shell.execute_reply.started":"2023-10-12T00:37:36.196775Z","shell.execute_reply":"2023-10-12T00:37:36.202245Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Stratified Group K Fold\n\n包含tdcsfog与defog两部分数据，应该是基于Sub或者Mod来做Kfold。深度学习中一般K设定的比较大？\n\n[Do we have metadata in the test set?](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/396959)","metadata":{}},{"cell_type":"code","source":"# Create Mapping between Id and Subject\nid2sub_df = pd.concat([\n    pd.read_csv(f, usecols=['Id', 'Subject']).\n    assign(Module=basename(f).split('_')[0]) for f in cfg.metadata_paths]).astype(\"category\").set_index(\"Id\")\n\nprint(f\"id2sub_df length: {len(id2sub_df)}, unique Ids: {id2sub_df.index.nunique()}, unique Subjects: {id2sub_df.Subject.nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.204628Z","iopub.execute_input":"2023-10-12T00:37:36.205752Z","iopub.status.idle":"2023-10-12T00:37:36.225584Z","shell.execute_reply.started":"2023-10-12T00:37:36.205717Z","shell.execute_reply":"2023-10-12T00:37:36.224572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"id2sub_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.226956Z","iopub.execute_input":"2023-10-12T00:37:36.227634Z","iopub.status.idle":"2023-10-12T00:37:36.238283Z","shell.execute_reply.started":"2023-10-12T00:37:36.227599Z","shell.execute_reply":"2023-10-12T00:37:36.237137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## data proc\n\n生成`train_df`对象，这个对象只有三个y值，聚合的y值event，subject字段","metadata":{}},{"cell_type":"code","source":"# # np.select(condlist, choicelist, default=0)\n\n# # 这里只说明condlist与choicelist维度不相同时候的情况\n# # choicelist是一维的，那么它的长度必须与condlist中的每个条件数组的维度一致\n\n# import numpy as np\n\n# # 创建条件数组\n# condlist = [np.array([True, True, False, False]), \n#             np.array([False, True, False, True])]\n# #                      1       1      0      2\n\n# # 创建选择数组\n# choicelist = np.array([1, 2])\n\n# # 使用np.select进行选择\n# result = np.select(condlist, choicelist, default=0)\n\n# print(result)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T01:40:45.007490Z","iopub.execute_input":"2023-10-12T01:40:45.008105Z","iopub.status.idle":"2023-10-12T01:40:45.013867Z","shell.execute_reply.started":"2023-10-12T01:40:45.008050Z","shell.execute_reply":"2023-10-12T01:40:45.012198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read csv files and add metadata (Id, Subject, Event)\ndef reader(filepath, \n           usecols, \n           getid=False, \n           getsub=False, \n           getevent=False, \n           dtype=None, \n           exclude=['notype']):\n    \n    fog_type = basename(dirname(filepath))\n    \n    if fog_type in exclude:\n        return None\n    \n    df = pd.read_csv(filepath, index_col=\"Time\", usecols=usecols, dtype=dtype)\n    \n    if getid:\n        df['Id'] = basename(filepath).split('.')[0] + '_' + df.index.astype(str)\n        \n    if getsub:\n        df['Subject'] = id2sub_df.loc[basename(filepath).split('.')[0], 'Subject']\n        \n    if getevent:\n        df['Event'] = np.select(\n            [df[col].astype(bool) for col in cfg.label_list], \n            np.arange(1,cfg.n_labels+1), default=0\n        ).astype('int8')\n        \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.252443Z","iopub.execute_input":"2023-10-12T00:37:36.252964Z","iopub.status.idle":"2023-10-12T00:37:36.261380Z","shell.execute_reply.started":"2023-10-12T00:37:36.252939Z","shell.execute_reply":"2023-10-12T00:37:36.260608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create common train Dataframe\ntrain_paths = glob(join(TRAIN_DIR, '*/*.csv'))\ndtype = {col:'int8' for col in cfg.label_list}\ndtype['Time'] = 'int32'","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.262892Z","iopub.execute_input":"2023-10-12T00:37:36.263626Z","iopub.status.idle":"2023-10-12T00:37:36.277144Z","shell.execute_reply.started":"2023-10-12T00:37:36.263592Z","shell.execute_reply":"2023-10-12T00:37:36.276040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"usecols = ['Time', *cfg.label_list]\n\nprint('model_used cols:', usecols)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.278789Z","iopub.execute_input":"2023-10-12T00:37:36.279505Z","iopub.status.idle":"2023-10-12T00:37:36.285404Z","shell.execute_reply.started":"2023-10-12T00:37:36.279465Z","shell.execute_reply":"2023-10-12T00:37:36.284276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_reader = partial(reader, usecols=usecols, dtype=dtype, \n                       getsub=True, getevent=True)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.287010Z","iopub.execute_input":"2023-10-12T00:37:36.288026Z","iopub.status.idle":"2023-10-12T00:37:36.294657Z","shell.execute_reply.started":"2023-10-12T00:37:36.287968Z","shell.execute_reply":"2023-10-12T00:37:36.293985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.concat([train_reader(f) for f in tqdm(train_paths)]).reset_index(drop=True)\ntrain_df.Subject = train_df.Subject.astype('category')\n","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:36.295873Z","iopub.execute_input":"2023-10-12T00:37:36.296913Z","iopub.status.idle":"2023-10-12T00:37:54.012094Z","shell.execute_reply.started":"2023-10-12T00:37:36.296879Z","shell.execute_reply":"2023-10-12T00:37:54.011036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:54.013432Z","iopub.execute_input":"2023-10-12T00:37:54.014078Z","iopub.status.idle":"2023-10-12T00:37:54.029288Z","shell.execute_reply.started":"2023-10-12T00:37:54.014039Z","shell.execute_reply":"2023-10-12T00:37:54.028177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(train_df.Event.value_counts().to_frame().style.background_gradient())","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:54.030576Z","iopub.execute_input":"2023-10-12T00:37:54.031472Z","iopub.status.idle":"2023-10-12T00:37:54.122320Z","shell.execute_reply.started":"2023-10-12T00:37:54.031436Z","shell.execute_reply":"2023-10-12T00:37:54.121129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### save info by kfold\n\n这里生成的应该还是配置对象","metadata":{}},{"cell_type":"code","source":"# Save paths for each Stratified Group Fold for defog and tdcsfog separately\nsgkf = StratifiedGroupKFold(n_splits=cfg.splits, random_state=42, shuffle=True)\nprint('splits=', cfg.splits)\n\n# 基于kfold进行划分的train与valid\nfold_train_fpaths = {'defog': [], 'tdcsfog':[]}\nfold_valid_fpaths = {'defog': [], 'tdcsfog':[]}\n\ndf_paths = {'defog':glob(join(cfg.train_sub_dirs[0],'*.csv')), \n            'tdcsfog': glob(join(cfg.train_sub_dirs[1],'*.csv'))}\nprint(cfg.train_sub_dirs)","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:54.123944Z","iopub.execute_input":"2023-10-12T00:37:54.124684Z","iopub.status.idle":"2023-10-12T00:37:54.137434Z","shell.execute_reply.started":"2023-10-12T00:37:54.124643Z","shell.execute_reply":"2023-10-12T00:37:54.135938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# path = \"/path/to/file.txt\"\n# basename_ = os.path.basename(path)\n# print(basename_)  # 输出: file.txt","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:54.139170Z","iopub.execute_input":"2023-10-12T00:37:54.139543Z","iopub.status.idle":"2023-10-12T00:37:54.145755Z","shell.execute_reply.started":"2023-10-12T00:37:54.139505Z","shell.execute_reply":"2023-10-12T00:37:54.144423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 外层的两个循环：defog, tdcsfog\nfor module, paths in df_paths.items():\n    print(f\"{module}:\")\n    sub_train_df = train_df[train_df.Subject.isin(id2sub_df.loc[id2sub_df.Module==module, 'Subject'])].reset_index(drop=True)\n    \n    for i, (train_index, test_index) in enumerate(sgkf.split(sub_train_df.index, sub_train_df.Event, \n                                                             groups=sub_train_df.Subject)):\n        print(f\"\\tFold {i}:\", end=\" \")\n        train_subs = sub_train_df.loc[train_index, 'Subject'].unique()\n        test_subs = sub_train_df.loc[test_index, 'Subject'].unique()\n        \n        print(f\"Subjects->train:{len(train_subs)}|test:{len(test_subs)}\")\n        train_ids = set(id2sub_df[id2sub_df.Subject.isin(train_subs)].index)\n        test_ids = set(id2sub_df[id2sub_df.Subject.isin(test_subs)].index)\n        fold_train_fpaths[module].append([f for f in paths if basename(f).split('.')[0] in train_ids])\n        fold_valid_fpaths[module].append([f for f in paths if basename(f).split('.')[0] in test_ids])\n        del train_subs, test_subs, train_ids, test_ids\n        gc.collect()\n        \n    del sub_train_df\n    \ndel train_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:37:54.147408Z","iopub.execute_input":"2023-10-12T00:37:54.148452Z","iopub.status.idle":"2023-10-12T00:39:08.540798Z","shell.execute_reply.started":"2023-10-12T00:37:54.148382Z","shell.execute_reply":"2023-10-12T00:39:08.539725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fold_train_fpaths  # {'defog': [[fold1:path, fold2:path, .....]], 'tdcsfog': [[fold1:path, fold2:path, .....]]}","metadata":{"execution":{"iopub.status.busy":"2023-10-12T00:36:21.012419Z","iopub.execute_input":"2023-10-12T00:36:21.013141Z","iopub.status.idle":"2023-10-12T00:36:21.032142Z","shell.execute_reply.started":"2023-10-12T00:36:21.013106Z","shell.execute_reply":"2023-10-12T00:36:21.031067Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset","metadata":{}},{"cell_type":"code","source":"# # 创建一个 2x3 的数组\n# values = np.array([[1, 2, 3], [4, 5, 6]])\n\n# # 在行维度上分别在前面添加 1 行，后面添加 2 行，在列维度上不进行填充\n# padded_values = np.pad(values, ((1, 2), (0, 0)), mode='edge')\n\n# print(padded_values)\n","metadata":{"execution":{"iopub.status.busy":"2023-10-12T01:40:22.225790Z","iopub.execute_input":"2023-10-12T01:40:22.226167Z","iopub.status.idle":"2023-10-12T01:40:22.231739Z","shell.execute_reply.started":"2023-10-12T01:40:22.226136Z","shell.execute_reply":"2023-10-12T01:40:22.230411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"class FOGSequence(tf.keras.utils.Sequence) 应该就是和pytorch中的DataSet类是类似的\n\n这里最主要需要了解的是`self.mapping` 对象\n\n这里给到模型的数据对象:\n1. (indices_batch, time_windows_length, feas)\n\n","metadata":{}},{"cell_type":"markdown","source":"需要注意，按照ChatGPT的说法，在TensorFlow中, `__getitem__`  函数通常情况下**应该返回一个批次的数据，而不是单个数据**。但在pytorch中，`__getitem__`函数通常**应该返回单个样本的数据，而不是带有批次维度的数据**。PyTorch内部会负责将单个样本数据组织成批次数据并传递给模型。\n\n在TensorFlow中，`__getitem__`函数的**idx参数通常用于指定要获取的数据批次的索引**。这个索引通常是一个整数，表示数据集中的第几个批次。你可以根据需要自定义 `__getitem__` 函数，根据idx参数获取对应的数据批次。例如，可以使用idx参数来确定从数据集中的哪个位置开始获取数据，然后返回一个批次的特征和标签。\n\n```python\n\ndef __getitem__(self, idx):\n    # 根据idx参数确定要获取的数据批次的起始位置\n    start_idx = idx * self.batch_size\n    end_idx = start_idx + self.batch_size\n\n    # 获取对应的数据批次\n    batch_X = self.features[start_idx:end_idx]\n    batch_y = self.labels[start_idx:end_idx]\n\n    return batch_X, batch_y\n```\n\n\n在PyTorch中，`__getitem__`函数的idx参数通常用于指定要获取的数据集中的单个样本的索引。这个索引通常是一个整数，表示数据集中的第几个样本。你可以根据需要自定义`__getitem__`函数，根据idx参数获取对应的单个样本。例如，可以使用idx参数来确定从数据集中的哪个位置获取样本，然后返回该样本的特征和标签。\n\n```python\ndef __getitem__(self, idx):\n    # 根据idx参数确定要获取的样本的位置\n    sample_X = self.features[idx]\n    sample_y = self.labels[idx]\n\n    return sample_X, sample_y\n\n```\n\n\n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"TensorFlow的`__getitem__`函数中的idx所代表的的含义是批次的索引，那么对应的`__len__`应该返回的是一个epoch的批次总数","metadata":{}},{"cell_type":"code","source":"# Adapted from FOGDataset of https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\nclass FOGSequence(tf.keras.utils.Sequence):\n\n    def __init__(self, df_paths, cfg=cfg, split=\"train\"):\n        _time = perf_counter()\n        \n        self.rng = default_rng(42)  # from numpy.random import default_rng\n        self.cfg = cfg\n        self.split = split\n        \n        self.past_pad = self.cfg.wx*(self.cfg.window_past-1)\n        self.future_pad = self.cfg.wx*self.cfg.window_future\n        \n        if self.split == \"test\":\n            self.Ids = []\n        _values = [self._read(f) for f in df_paths]  \n        \n        # by Id.csv 统计，剔除past_pad和future_pad的部分\n        self.mapping = []\n        _length = 0\n        for _value in _values:\n            _shape = _value.shape[0]\n            self.mapping.extend(range(_length+self.past_pad, _length+_shape-self.future_pad))\n            _length += _shape\n        \n        self.values = np.concatenate(_values, axis=0)  # 全量数据集\n        self.mapping = np.array(self.mapping)\n        \n        if self.split != \"test\":\n            # Keep only vaild and task rows\n            _valid_pos = self.values[self.mapping, self.valid_position] > 0\n            _task_pos = self.values[self.mapping, self.task_position] > 0\n            self.mapping = self.mapping[_valid_pos&_task_pos]\n        self.length = self.mapping.shape[0]\n        \n        print(f\"Valid Dataset of size {self.length:,} initialized in {perf_counter() - _time:.3f} secs!\")\n        gc.collect()\n    \n    def _read(self, path):\n        \n        # tdcsfog 与 defog均使用该函数读取，前者则补充缺失的字段Valid与Task\n        # 这里return的 df应该是两维的: (time, feas)\n        \n        _is_tdcs = basename(dirname(path)).startswith('tdcs')\n        df = pd.read_csv(path)\n        \n        if self.split == \"test\":\n            _ids = basename(path).split('.')[0] + '_' + df.Time.astype(str)\n            self.Ids.extend(_ids.tolist())\n            return self._df_to_array(df, self.cfg.feature_list)\n        \n        _cols = [*self.cfg.feature_list, *self.cfg.label_list, 'Valid', 'Task']\n        \n        self.valid_position = self.cfg.n_features + self.cfg.n_labels  # \"Valid\" 字段所处的数据集列位置\n        self.task_position = self.valid_position + 1  # \"task\" 字段所处数据集列位置\n        \n        if _is_tdcs:\n            # Fill Valid and Task columns for tdcsfog\n            df['Valid'] = 1\n            df['Task'] = 1\n            \n        return self._df_to_array(df, _cols)\n    \n    def _df_to_array(self, df, cols):\n        # Pads past and future rows to dataframe values for indexing \n        _values = df[cols].values.astype(np.float16)\n        \n        # 这里是在第0维度的前后分别填充，第1维度不填充\n        return np.pad(_values, ((self.past_pad, self.future_pad),(0,0)), 'edge')\n    \n    def __len__(self):\n        return int(np.ceil(self.length / self.cfg.batch_size))\n    \n    def __getitem__(self, idx):\n        \"\"\"\n        \n        :param idx: \n        :return: \n        \"\"\"\n        \n        # 基于mapping列表，筛选batch数据的起点索引\n        if self.split == \"train\":\n            # Onlt train set has randomly selected batches\n            # _idxs 是一维的 [batch_size]\n            # 这样写是可以理解的，在抽取每个批次样本时保持随机性，这点和pytorch中不一样。后者每个epoch所有的样本都会被确定抽到一次，但深度学习中这点没有必要保证。\n            _idxs = self.rng.choice(self.mapping, size=self.cfg.batch_size, replace=False)\n        else:\n            _idxs = self._get_indices(idx)\n            \n        # For test return only features\n        if self.split == \"test\":\n            return self._get_X(_idxs)\n        # For train and val splits return y also\n        return self._get_X_y(_idxs)\n    \n    def _get_indices(self, idx):\n        _low = idx * self.cfg.batch_size\n        # Cap high at self.length so overflow does not occur\n        _high = min(_low + self.cfg.batch_size, self.length)\n        return self.mapping[_low:_high]\n    \n    def _get_X(self, indices):\n        \n        # _X: (indices_batch, time_windows_length, feas)\n        \n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            # 这里的数据还是“跳着”取的，注意语法\n            _X[i] = self.values[idx-self.past_pad:idx+self.future_pad+1:self.cfg.wx, :self.cfg.n_features]\n            \n        return _X\n    \n    def _get_X_y(self, indices):\n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            _X[i] = self.values[idx-self.past_pad: idx+self.future_pad+1:self.cfg.wx, :self.cfg.n_features]\n        return _X, self.values[indices, self.cfg.n_features:self.cfg.n_features+self.cfg.n_labels]","metadata":{"execution":{"iopub.status.busy":"2023-10-12T01:42:41.189578Z","iopub.execute_input":"2023-10-12T01:42:41.190046Z","iopub.status.idle":"2023-10-12T01:42:41.225009Z","shell.execute_reply.started":"2023-10-12T01:42:41.190004Z","shell.execute_reply":"2023-10-12T01:42:41.223859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train_ds = FOGSequence(fold_train_fpaths['defog'][0])\n\n# 基于mapping列表，筛选batch数据的起点索引\n# train_ds.rng.choice(train_ds.mapping, size=train_ds.cfg.batch_size, replace=False)\n\n# by idx 筛选batch数据\n# idx = 417965\n# train_ds.values[idx-train_ds.past_pad:idx+train_ds.future_pad+1:train_ds.cfg.wx, :train_ds.cfg.n_features].shape\n\n# train_ds.mapping[1200:1250]\n\n# train_ds.values.shape","metadata":{"execution":{"iopub.status.busy":"2023-10-12T02:47:55.591493Z","iopub.execute_input":"2023-10-12T02:47:55.592111Z","iopub.status.idle":"2023-10-12T02:47:55.600710Z","shell.execute_reply.started":"2023-10-12T02:47:55.592074Z","shell.execute_reply":"2023-10-12T02:47:55.599615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model","metadata":{}},{"cell_type":"code","source":"\ndef calculate_precision(y_true, y_pred):\n    \n    # average_precision_score with positive sample added if no true positive cases are present\n    # y_true 的维度 (batch, 3_y_vals)  \n    # 至少有一条数据有为1的label，那么就无需填充，否则就做一次填充，确保 average_precision_score 计算不报错\n    pad_width = ((0,0),(0,0)) if y_true.any(axis=0).all() else ((1,0),(0,0))\n    \n    y_true = np.pad(y_true, pad_width, constant_values=1)\n    y_pred = np.pad(y_pred, pad_width, constant_values=1)\n    \n    return average_precision_score(y_true, y_pred)  # from sklearn.metrics import average_precision_score","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:25:31.403074Z","iopub.execute_input":"2023-04-04T09:25:31.40347Z","iopub.status.idle":"2023-04-04T09:25:31.419007Z","shell.execute_reply.started":"2023-04-04T09:25:31.403431Z","shell.execute_reply":"2023-04-04T09:25:31.417772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"用于计算平均精确度（Average Precision）的自定义指标类:\n1. 它继承自tf.keras.metrics.Metric类，可以通过调用update_state方法来更新指标的状态，通过调用result方法来获取最终的指标结果，以及通过调用reset_state方法来重置指标的状态。\n2. 在初始化函数 \\_\\_init\\_\\_ 中，需要传入num_classes参数表示类别的数量，以及可选的thresholds参数表示阈值。\n    1. 在初始化函数中，首先调用父类的初始化方法来设置指标的名称和其他参数。\n    2. 然后，创建了一个长度为num_classes的列表self.class_precision，其中每个元素都是一个tf.keras.metrics.Precision对象，用于计算每个类别的精确度。\n3. 在update_state方法中，通过遍历self.class_precision列表，逐个调用其update_state方法来更新每个类别的精确度。该方法接受三个参数，y_true表示真实标签，y_pred表示预测标签，sample_weight表示样本权重。在这里，通过使用enumerate函数来同时获取类别的索引i和对应的精确度对象precision，然后将对应的真实标签和预测标签传递给precision的update_state方法进行更新。\n4. 在result方法中，通过列表推导式和tf.math.reduce_mean函数来计算所有类别精确度的平均值。\n    1. 首先，使用列表推导式遍历self.class_precision列表，获取每个类别的精确度结果，\n    2. 然后使用tf.math.reduce_mean函数来计算这些结果的平均值。\n5. 在reset_state方法中，通过遍历self.class_precision列表，逐个调用其reset_state方法来重置每个类别的精确度对象的状态。该方法没有参数，调用reset_state方法会将精确度对象的统计指标重置为初始状态，以便在下一次计算中重新开始统计。","metadata":{}},{"cell_type":"code","source":"# Note: Not same result as average_precision_score\nclass AveragePrecision(tf.keras.metrics.Metric):\n\n    def __init__(self, num_classes, thresholds=None, name='avg_precision', **kwargs):\n        super(AveragePrecision, self).__init__(name=name, **kwargs)\n        self.class_precision = [tf.keras.metrics.Precision(thresholds) for _ in range(num_classes)]\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        for i, precision in enumerate(self.class_precision):\n            precision.update_state(y_true[...,i], y_pred[..., i])\n\n    def result(self):\n        # 这里计算 mAP\n        return tf.math.reduce_mean([precision.result() for precision in self.class_precision])\n    \n    def reset_state(self):\n        for precision in self.class_precision:\n            precision.reset_state()","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:25:31.422959Z","iopub.execute_input":"2023-04-04T09:25:31.423419Z","iopub.status.idle":"2023-04-04T09:25:31.434144Z","shell.execute_reply.started":"2023-04-04T09:25:31.423376Z","shell.execute_reply":"2023-04-04T09:25:31.432935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model adapted from https://keras.io/examples/timeseries/timeseries_classification_from_scratch/\ndef get_model(checkpoint_path = None):\n    model = tf.keras.models.Sequential()\n    \n    # 定义输入维度\n    model.add(tf.keras.Input(shape=(cfg.window_size, cfg.n_features), dtype='float16'))\n    \n    for _ in range(cfg.model_nblocks):\n        model.add(tf.keras.layers.Conv1D(filters=cfg.model_hidden, kernel_size=15, padding=\"same\"))\n        model.add(tf.keras.layers.BatchNormalization())\n        model.add(tf.keras.layers.ReLU())\n        model.add(tf.keras.layers.Dropout(cfg.model_dropout))\n        \n    model.add(tf.keras.layers.GlobalAveragePooling1D())\n    model.add(tf.keras.layers.Dense(cfg.n_labels, activation=None))\n\n    if checkpoint_path is not None:\n        model.load_weights(checkpoint_path)\n        \n    model.compile(\n        tf.keras.optimizers.Adam(learning_rate=cfg.lr), \n        loss = tf.keras.losses.BinaryCrossentropy(from_logits=True),\n        metrics=[AveragePrecision(cfg.n_labels, thresholds=0.0)]\n    )\n    \n    return model\n","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:25:31.435863Z","iopub.execute_input":"2023-04-04T09:25:31.436273Z","iopub.status.idle":"2023-04-04T09:25:33.903424Z","shell.execute_reply.started":"2023-04-04T09:25:31.436235Z","shell.execute_reply":"2023-04-04T09:25:33.902554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_model().summary()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train","metadata":{}},{"cell_type":"code","source":"def train_loop(train_paths, valid_paths, fold, model_save_dir=''):\n    \n    \n    gc.collect()\n    \n    # train_paths, test_paths = train_test_split(train_paths, test_size=0.4)\n    train_ds = FOGSequence(train_paths)\n    val_ds = FOGSequence(valid_paths, split=\"val\")\n    # test_ds = FOGSequence(test_paths, split=\"val\")\n    \n    model = get_model()\n    \n    ckpt = tf.keras.callbacks.ModelCheckpoint(join(model_save_dir, f\"fold{fold}_model_\"+\"{epoch:02d}.h5\"), \n                                              save_weights_only=True)\n    \n    history = model.fit(train_ds, \n                        epochs=cfg.num_epochs, \n                        verbose=2, \n                        workers=5, \n                        validation_data=val_ds, \n                        use_multiprocessing=True, \n                        callbacks=[ckpt])\n    # 这里应该是所有的epoch都存了一份结果，因此后续在 predict_select_model 函数需要做遍历计算\n    \n    best_model_path = predict_select_model(fold, val_ds, model_save_dir)\n    # score = calculate_precision(test_ds.values[test_ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], expit(get_model(best_model_path).predict(test_ds, verbose=0)))\n    \n    del train_ds, val_ds, model, ckpt, history\n    gc.collect()\n    return best_model_path","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:25:33.91929Z","iopub.execute_input":"2023-04-04T09:25:33.919841Z","iopub.status.idle":"2023-04-04T09:25:33.933125Z","shell.execute_reply.started":"2023-04-04T09:25:33.919799Z","shell.execute_reply":"2023-04-04T09:25:33.932263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def predict_select_model(fold, ds, model_save_dir=''):\n    best_path = None\n    best_score = -1\n    \n    print(f\"Validation for fold{fold}:\")\n    for model_path in sorted(glob(join(model_save_dir, f\"fold{fold}_*.h5\"))):\n        pred_time = perf_counter()\n        gc.collect()\n        \n        score = calculate_precision(\n            ds.values[ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], \n            expit(get_model(model_path).predict(ds, verbose=0)) # expit converts to sigmoid output\n        )\n        \n        if best_score < score:\n            best_score = score\n            best_path = model_path\n            \n        gc.collect()\n        print(\"\\t\", basename(model_path), f\": score-{score:.4f} in {(perf_counter()-pred_time)/60:.2f} mins\")\n        \n    print(basename(best_path), \"selected with score\", best_score)\n    return best_path","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:25:33.904664Z","iopub.execute_input":"2023-04-04T09:25:33.905087Z","iopub.status.idle":"2023-04-04T09:25:33.917924Z","shell.execute_reply.started":"2023-04-04T09:25:33.905047Z","shell.execute_reply":"2023-04-04T09:25:33.917057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"by_module by_kfold 训练。默认设定kfold=10，训练产出20个模型","metadata":{}},{"cell_type":"code","source":"# Main training loop\nmodel_paths = {'defog': [], 'tdcsfog':[]}\n\nfor module in model_paths:\n    \n    module_start = perf_counter()\n    print(f\"***Training {module}{'*'*75}\")\n    if not exists(module): \n        os.mkdir(module)\n        \n    for fold, (train_fpaths, valid_fpaths) in enumerate(zip(fold_train_fpaths[module], fold_valid_fpaths[module])):\n        fold_start = perf_counter()\n        print(f\"Fold {fold}{'-'*25}\")\n        model_paths[module].append(train_loop(train_fpaths, valid_fpaths, fold, model_save_dir=module))\n        print(f\"Fold {fold} done in {(perf_counter()-fold_start)/60:.2f} min\")\n        \n    print(f\"***{module} done in {(perf_counter()-module_start)/3600:.2f} hrs{'*'*50}\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-04-04T08:47:09.041219Z","iopub.execute_input":"2023-04-04T08:47:09.042066Z","iopub.status.idle":"2023-04-04T09:06:01.256473Z","shell.execute_reply.started":"2023-04-04T08:47:09.04203Z","shell.execute_reply":"2023-04-04T09:06:01.255137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"markdown","source":"test数据集: Only the Time, AccV, AccML, and AccAP fields are provided for the test series. 在 [evaluation](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/overview/evaluation) 界面告知，赛题方在test数据集中，对defog的数据部分，只会使用`Valid`和`task`均标为1的数据。","metadata":{}},{"cell_type":"code","source":"test_defog_paths = glob(join(TEST_DIR, \"defog/*.csv\"))\ntest_tdcsfog_paths = glob(join(TEST_DIR, \"tdcsfog/*.csv\"))\n\ntest_ds_dict = {\n    'defog':FOGSequence(test_defog_paths, split=\"test\"), \n    'tdcsfog':FOGSequence(test_tdcsfog_paths, split=\"test\")\n}\n\n# Get test predictions\ndf_list = []\nfor module, test_ds in test_ds_dict.items():\n    y_pred_list = []\n    for model_path in model_paths[module]:\n        model = get_model(model_path)\n        y_pred_list.append(expit(model.predict(test_ds, verbose=0)))  # expit converts to sigmoid output\n    y_pred = np.mean(y_pred_list, axis=0)\n    df_list.append(pd.DataFrame(\n        {'Id': test_ds.Ids, 'StartHesitation': y_pred[:,0], 'Turn': y_pred[:,1], 'Walking': y_pred[:,2]}))","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:06:01.270977Z","iopub.execute_input":"2023-04-04T09:06:01.271624Z","iopub.status.idle":"2023-04-04T09:06:11.801757Z","shell.execute_reply.started":"2023-04-04T09:06:01.271586Z","shell.execute_reply":"2023-04-04T09:06:11.800667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate Prediction to DataFrames\nsubmission = pd.concat(df_list)\n\n# Only keep Ids in sample_submission\nsample_submission = pd.read_csv(join(BASE_DIR, \"sample_submission.csv\"))\nsubmission = pd.merge(sample_submission[['Id']], submission, how='left', on='Id').fillna(0.0)\nsubmission.to_csv(\"submission.csv\", index=False, float_format='%.5f') # round to 5 decimal places while keeping point notation","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:06:11.80431Z","iopub.execute_input":"2023-04-04T09:06:11.805087Z","iopub.status.idle":"2023-04-04T09:06:13.9527Z","shell.execute_reply.started":"2023-04-04T09:06:11.805043Z","shell.execute_reply":"2023-04-04T09:06:13.951598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!head -5 submission.csv","metadata":{"execution":{"iopub.status.busy":"2023-04-04T09:06:13.954138Z","iopub.execute_input":"2023-04-04T09:06:13.955169Z","iopub.status.idle":"2023-04-04T09:06:15.016791Z","shell.execute_reply.started":"2023-04-04T09:06:13.955128Z","shell.execute_reply":"2023-04-04T09:06:15.015494Z"},"trusted":true},"execution_count":null,"outputs":[]}]}