{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<br>\n\n**References**:\n* `coderRKJ`'s notebook: <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" style=\"text-decoration:none\">Parkinson FoG Pred Conv1D Separate TF Model</a>\n\n<br>\n<br>\n\n**Main ideas of the referenced notebook**：\n* Using 3 input features: `AccV, AccML, AccAP` with simple np.pad().\n* Training defog and tdcsfog datasets with Conv1D NN model seperately. (2 Conv1D NN model with same structure)\n* StratifiedGroupKFold (10 folds, groups=Subject) and each fold training with 5 epochs.\n* Each fold saves a best model with weights-only.\n* Final prediction derived from average results of 10 best models corresponding to 10 folds.\n\n<br>\n\n**Possible Improvements**:\n* Data preprocessing like signal denosing, standardizing, etc.\n* Feature engineering to create more new features like foot-steps.\n* Different Conv1D NN models for defog and tdcsfog, like different input shape.\n* Try 5 folds but with more epochs.\n* others like this <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416106#2294598\" style=\"text-decoration:none\">20th place solution</a>\n\n\n\n<br>\n<br>\n\nThe above content is written by Ping. The below is basically the referenced notebook content, except few code adjustments and some explanatory comments.\n\nThanks coderRKJ for sharing this wonderful notebook.\n\n\n---------------------","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"**Conv1D TF Model with Separate Pipelines for defog and tdcsfog data**\n\nFeature Column TimeSeries grouping adapted from: <a href=\"https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\" style=\"text-decoration:none\">PyTorch FOG End-to-End Baseline [LB 0.254]</a>\n\nSubject-wise GroupKFold splitting adapted from: <a href=\"https://www.kaggle.com/code/xzj19013742/groupkfold-cross-validation-tsflex\" style=\"text-decoration:none\">GroupKfold Cross-Validation tsflex 🚀</a>\n\n<br>\n\nIn this Notebook:\n* Tensorflow Model with Conv1D blocks\n* Models trained separately for defog and tdcsfog data\n* Event Stratified Subject Grouped KFold splitting","metadata":{}},{"cell_type":"markdown","source":"<br>\n\n# Imports and Config","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\n\nimport os\nimport gc\nimport glob\nfrom tqdm.auto import tqdm\nfrom time import perf_counter\nfrom collections import defaultdict\n\nimport numpy as np\nimport pandas as pd\nfrom numpy.random import default_rng\n\nfrom functools import partial\nfrom scipy.special import expit\n\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, StratifiedGroupKFold\nfrom sklearn.metrics import average_precision_score\nfrom sklearn.preprocessing import StandardScaler\n\n\nimport tensorflow as tf\nprint(f\"TF version: {tf.__version__}\")\nAUTO = tf.data.experimental.AUTOTUNE","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:22.188433Z","iopub.execute_input":"2023-06-09T23:51:22.189246Z","iopub.status.idle":"2023-06-09T23:51:25.917436Z","shell.execute_reply.started":"2023-06-09T23:51:22.189210Z","shell.execute_reply":"2023-06-09T23:51:25.916242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Constants\n\nBASE_DIR = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction\"\nTRAIN_DIR = os.path.join(BASE_DIR, \"train\")\nTEST_DIR = os.path.join(BASE_DIR, \"test\")\n\nIS_PUBLIC = len(glob.glob(os.path.join(TEST_DIR, \"*/*.csv\")))==2","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:25.923250Z","iopub.execute_input":"2023-06-09T23:51:25.923989Z","iopub.status.idle":"2023-06-09T23:51:25.932132Z","shell.execute_reply.started":"2023-06-09T23:51:25.923954Z","shell.execute_reply":"2023-06-09T23:51:25.931233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Config:\n    train_sub_dirs = [os.path.join(TRAIN_DIR, \"defog\"), os.path.join(TRAIN_DIR, \"tdcsfog\")]\n    \n    metadata_paths = [os.path.join(BASE_DIR, \"defog_metadata.csv\"), os.path.join(BASE_DIR, \"tdcsfog_metadata.csv\")]\n    \n    splits = 10\n\n    batch_size = 1024\n    window_size = 64\n    window_future = 16\n    window_past = window_size - window_future     # Includes current value\n    \n    wx = 8    # 【Ping】 window_size multiples\n    \n    model_dropout = 0.2\n    model_hidden = 128\n    model_nblocks = 3\n    \n    lr = 0.00015\n    num_epochs = 5\n    \n    feature_list = ['AccV', 'AccML', 'AccAP']\n    label_list = ['StartHesitation', 'Turn', 'Walking']\n    \n    n_features = len(feature_list)\n    n_labels = len(label_list)    \n    \n\n\ncfg = Config()","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:25.934338Z","iopub.execute_input":"2023-06-09T23:51:25.935324Z","iopub.status.idle":"2023-06-09T23:51:25.943675Z","shell.execute_reply.started":"2023-06-09T23:51:25.935291Z","shell.execute_reply":"2023-06-09T23:51:25.942590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br>\n\n# Stratified Group K Fold","metadata":{}},{"cell_type":"code","source":"# Create Mapping between Id and Subject\n#【Ping】 from `defog_metadata.csv` and `tdcsfog_metadata.csv`\nid2sub_df = pd.concat([pd.read_csv(f, usecols=['Id', 'Subject'])\\\n                         .assign(Module=os.path.basename(f).split('_')[0]) for f in cfg.metadata_paths]) \\\n              .astype(\"category\") \\\n              .set_index(\"Id\")\nprint(f\"id2sub_df length: {len(id2sub_df)}, unique Ids: {id2sub_df.index.nunique()}, unique Subjects: {id2sub_df.Subject.nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:25.947150Z","iopub.execute_input":"2023-06-09T23:51:25.947724Z","iopub.status.idle":"2023-06-09T23:51:25.970283Z","shell.execute_reply.started":"2023-06-09T23:51:25.947692Z","shell.execute_reply":"2023-06-09T23:51:25.969409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(id2sub_df.dtypes)\n\nid2sub_df","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:25.971344Z","iopub.execute_input":"2023-06-09T23:51:25.971604Z","iopub.status.idle":"2023-06-09T23:51:25.990094Z","shell.execute_reply.started":"2023-06-09T23:51:25.971582Z","shell.execute_reply":"2023-06-09T23:51:25.988619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"```python\nfp = '/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog/02ea782681.csv'\nprint(os.path.dirname(fp))                     # /kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/train/defog\nprint(os.path.basename(os.path.dirname(fp)))   # defog\n\nprint()\nprint(os.path.basename(fp))                    # 02ea782681.csv\n```","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:25.991007Z","iopub.execute_input":"2023-06-09T23:51:25.991248Z","iopub.status.idle":"2023-06-09T23:51:25.999458Z","shell.execute_reply.started":"2023-06-09T23:51:25.991227Z","shell.execute_reply":"2023-06-09T23:51:25.998351Z"}}},{"cell_type":"code","source":"# Read csv files and add metadata (Id, Subject, Event)\ndef reader(filepath, usecols, getid=False, getsub=False, getevent=False, dtype=None, exclude=['notype']):\n    fog_type = os.path.basename(os.path.dirname(filepath))\n    if fog_type in exclude:\n        return None\n    df = pd.read_csv(filepath, index_col=\"Time\", usecols=usecols, dtype=dtype)\n    if getid:\n        df['Id'] = os.path.basename(filepath).split('.')[0] + '_' + df.index.astype(str)\n    if getsub:\n        df['Subject'] = id2sub_df.loc[os.path.basename(filepath).split('.')[0], 'Subject']\n    if getevent:\n        df['Event'] = np.select([df[col].astype(bool) for col in cfg.label_list], \n                                np.arange(1, cfg.n_labels+1), \n                                default=0).astype('int8')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:26.001272Z","iopub.execute_input":"2023-06-09T23:51:26.002005Z","iopub.status.idle":"2023-06-09T23:51:26.010559Z","shell.execute_reply.started":"2023-06-09T23:51:26.001973Z","shell.execute_reply":"2023-06-09T23:51:26.009690Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create common train Dataframe\ntrain_paths = glob.glob(os.path.join(TRAIN_DIR, '*/*.csv'))\ndtype = {col:'int8' for col in cfg.label_list}\ndtype['Time'] = 'int32'\nusecols = ['Time', *cfg.label_list]\n\ntrain_reader = partial(reader, usecols=usecols, dtype=dtype, getsub=True, getevent=True)\ntrain_df = pd.concat([train_reader(f) for f in tqdm(train_paths)]).reset_index(drop=True)\ntrain_df.Subject = train_df.Subject.astype('category')\ndisplay(train_df.Event.value_counts().to_frame().style.background_gradient())\n\n\n#【Ping】\n# This code cell reads all `.csv` filse in the `TRAIN_DIR` except those in the `train/notye/` subdir.\n# And concatenate them into one pandas DataFrame named `train_df` with added cols: Id, Subject, Event.","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:26.012100Z","iopub.execute_input":"2023-06-09T23:51:26.012812Z","iopub.status.idle":"2023-06-09T23:51:48.147341Z","shell.execute_reply.started":"2023-06-09T23:51:26.012754Z","shell.execute_reply":"2023-06-09T23:51:48.146345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:48.148908Z","iopub.execute_input":"2023-06-09T23:51:48.149277Z","iopub.status.idle":"2023-06-09T23:51:48.155449Z","shell.execute_reply.started":"2023-06-09T23:51:48.149244Z","shell.execute_reply":"2023-06-09T23:51:48.154557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:48.156746Z","iopub.execute_input":"2023-06-09T23:51:48.157489Z","iopub.status.idle":"2023-06-09T23:51:48.173547Z","shell.execute_reply.started":"2023-06-09T23:51:48.157454Z","shell.execute_reply":"2023-06-09T23:51:48.172620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save paths for each Stratified Group Fold for defog and tdcsfog separately\nsgkf = StratifiedGroupKFold(n_splits=cfg.splits, random_state=42, shuffle=True)\n\nfold_train_fpaths = {'defog': [], 'tdcsfog':[]}\nfold_valid_fpaths = {'defog': [], 'tdcsfog':[]}\ndf_paths = {'defog'  : glob.glob(os.path.join(cfg.train_sub_dirs[0],'*.csv')), \n            'tdcsfog': glob.glob(os.path.join(cfg.train_sub_dirs[1],'*.csv'))}\n\n\nfor module, paths in df_paths.items():\n    print(f\"{module}:\")\n    \n    #【Ping】Get Subjects of the corresponding module (defog, tdcsfog) in train_df,\n    #       and collect their related info as sub_train_df.\n    sub_train_df = train_df[train_df.Subject.isin(id2sub_df.loc[id2sub_df.Module==module, 'Subject'])].reset_index(drop=True)\n    \n    for i, (train_index, test_index) in enumerate(sgkf.split(sub_train_df.index, sub_train_df.Event, \n                                                             groups=sub_train_df.Subject)):\n        print(f\"\\tFold {i}:\", end=\" \")\n        train_subs = sub_train_df.loc[train_index, 'Subject'].unique()\n        test_subs = sub_train_df.loc[test_index, 'Subject'].unique()\n        print(f\"Subjects->train:{len(train_subs)}|test:{len(test_subs)}\")\n        train_ids = set(id2sub_df[id2sub_df.Subject.isin(train_subs)].index)\n        test_ids = set(id2sub_df[id2sub_df.Subject.isin(test_subs)].index)\n        fold_train_fpaths[module].append([f for f in paths if os.path.basename(f).split('.')[0] in train_ids])\n        fold_valid_fpaths[module].append([f for f in paths if os.path.basename(f).split('.')[0] in test_ids])\n        \n        del train_subs, test_subs, train_ids, test_ids\n        gc.collect()\n    del sub_train_df\n    \ndel train_df\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:51:48.175069Z","iopub.execute_input":"2023-06-09T23:51:48.175381Z","iopub.status.idle":"2023-06-09T23:53:55.219367Z","shell.execute_reply.started":"2023-06-09T23:51:48.175352Z","shell.execute_reply":"2023-06-09T23:53:55.218262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br>\n\n## Dataset","metadata":{}},{"cell_type":"code","source":"# Adapted from FOGDataset of https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\n\nclass FOGSequence(tf.keras.utils.Sequence):\n    \"\"\"【Ping】 Read datas from `.csv` in `train/` dir, only including those in ·train/defog/· and `trian/tdcsfog/`.  \"\"\"\n\n    def __init__(self, df_paths, cfg=cfg, split=\"train\"):\n        _time = perf_counter()\n        \n        self.rng = default_rng(42)\n        self.cfg = cfg\n        self.split = split\n        \n        self.past_pad = self.cfg.wx*(self.cfg.window_past-1)\n        self.future_pad = self.cfg.wx*self.cfg.window_future\n        \n        if self.split == \"test\":\n            self.Ids = []\n        \n        #【Ping】Read the training `.csv`, transforming into 2D array with np.pad() and collect them into a list\n        _values = [self._read(f) for f in df_paths]\n        \n        self.mapping = []\n        _length = 0\n        for _value in _values:          #【Ping】 _value is a padded array transformed from a .csv of training folder\n            _shape = _value.shape[0]    #【Ping】 _shape is the length of this padded array, also is the length of time points\n            \n            #【Ping】 self.mapping stores the indices of the original .csv dataframe in its corresponding np.pad() array, \n            #        and these indices are orderd like all of these padded arrays are concatenated into one array.\n            self.mapping.extend(range(_length+self.past_pad, _length+_shape-self.future_pad))\n            _length += _shape\n        \n        self.values = np.concatenate(_values, axis=0)\n        self.mapping = np.array(self.mapping)\n        if self.split != \"test\":\n            # Keep only vaild and task rows\n            _valid_pos = self.values[self.mapping, self.valid_position] > 0\n            _task_pos  = self.values[self.mapping, self.task_position] > 0\n            self.mapping = self.mapping[_valid_pos&_task_pos]    #【Ping】filtering self.mapping\n        self.length = self.mapping.shape[0]\n        \n        print(f\"Valid Dataset of size {self.length:,} initialized in {perf_counter() - _time:.3f} secs!\")\n        gc.collect()\n        \n        #【Ping】 Till now, we have \n        #        `self.values` which is a padded numpy array concateanted from all training .csv.\n        #        `self.mapping` which stores row indices of the original .csv dataframe in the cacatenated padded array\n        #                       and the rows restricted to having Valid & Task > 0.\n        #        `self.length` is the length of the final `self.mapping`.\n    \n    \n    \n    def _read(self, path):\n        _is_tdcs = os.path.basename(os.path.dirname(path)).startswith('tdcs')\n        df = pd.read_csv(path)\n        \n        if self.split == \"test\":\n            _ids = os.path.basename(path).split('.')[0] + '_' + df.Time.astype(str)\n            self.Ids.extend(_ids.tolist())\n            return self._df_to_array(df, self.cfg.feature_list)\n        \n        _cols = [*self.cfg.feature_list, *self.cfg.label_list, 'Valid', 'Task']\n        self.valid_position = self.cfg.n_features + self.cfg.n_labels\n        self.task_position = self.valid_position + 1\n        \n        if _is_tdcs:\n            # Fill Valid and Task columns for tdcsfog\n            df['Valid'] = 1\n            df['Task'] = 1\n            \n        return self._df_to_array(df, _cols)\n    \n    \n    \n    def _df_to_array(self, df, cols):\n        # Pads past and future rows to dataframe values for indexing \n        _values = df[cols].values.astype(np.float16)\n        return np.pad(_values, ((self.past_pad, self.future_pad), (0,0)), 'edge')\n    \n    \n    \n    def __len__(self):\n        #【Ping】number of batch_size in the actual dataset used for training. \n        # Its length is `self.mapping.shape[0]`, namely `self.length`\n        return int(np.ceil(self.length / self.cfg.batch_size))\n    \n    \n    \n    def __getitem__(self, idx):\n        \n        if self.split == \"train\":\n            # Only train set has randomly selected batches\n            _idxs = self.rng.choice(self.mapping, size=self.cfg.batch_size, replace=False)\n        else:\n            _idxs = self._get_indices(idx)\n            \n        # For test returns only features\n        if self.split == \"test\":\n            return self._get_X(_idxs)\n        # For train and val splits return y also\n        return self._get_X_y(_idxs)\n    \n    \n    \n    def _get_indices(self, idx):\n        _low = idx * self.cfg.batch_size\n        # Cap high at self.length so overflow does not occur\n        _high = min(_low + self.cfg.batch_size, self.length)\n        return self.mapping[_low:_high]\n    \n    \n    \n    def _get_X(self, indices):\n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            _X[i] = self.values[(idx-self.past_pad):(idx+self.future_pad+1):(self.cfg.wx), :self.cfg.n_features]\n        return _X\n    \n    \n    \n    def _get_X_y(self, indices):\n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            _X[i] = self.values[(idx-self.past_pad):(idx+self.future_pad+1):(self.cfg.wx), :self.cfg.n_features]\n        return _X, self.values[indices, self.cfg.n_features:self.cfg.n_features+self.cfg.n_labels]","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:55.221194Z","iopub.execute_input":"2023-06-09T23:53:55.221606Z","iopub.status.idle":"2023-06-09T23:53:55.244935Z","shell.execute_reply.started":"2023-06-09T23:53:55.221561Z","shell.execute_reply":"2023-06-09T23:53:55.243908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br>\n\n## Model","metadata":{}},{"cell_type":"code","source":"# average_precision_score with positive sample added if no true positive cases are present\ndef calculate_precision(y_true, y_pred):\n    pad_width = ((0,0),(0,0)) if y_true.any(axis=0).all() else ((1,0),(0,0))\n    y_true = np.pad(y_true, pad_width, constant_values=1)\n    y_pred = np.pad(y_pred, pad_width, constant_values=1)\n    return average_precision_score(y_true, y_pred)","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:55.249423Z","iopub.execute_input":"2023-06-09T23:53:55.249945Z","iopub.status.idle":"2023-06-09T23:53:55.261382Z","shell.execute_reply.started":"2023-06-09T23:53:55.249918Z","shell.execute_reply":"2023-06-09T23:53:55.260326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Note: Not same result as average_precision_score\nclass AveragePrecision(tf.keras.metrics.Metric):\n\n    def __init__(self, num_classes, thresholds=None, name='avg_precision', **kwargs):\n        super(AveragePrecision, self).__init__(name=name, **kwargs)\n        self.class_precision = [tf.keras.metrics.Precision(thresholds) for _ in range(num_classes)]\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        for i, precision in enumerate(self.class_precision):\n            precision.update_state(y_true[...,i], y_pred[..., i])\n\n    def result(self):\n        return tf.math.reduce_mean([precision.result() for precision in self.class_precision])\n    \n    def reset_state(self):\n        for precision in self.class_precision:\n            precision.reset_state()","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:55.263117Z","iopub.execute_input":"2023-06-09T23:53:55.263538Z","iopub.status.idle":"2023-06-09T23:53:55.276786Z","shell.execute_reply.started":"2023-06-09T23:53:55.263506Z","shell.execute_reply":"2023-06-09T23:53:55.275771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model adapted from https://keras.io/examples/timeseries/timeseries_classification_from_scratch/\ndef get_model(checkpoint_path = None):\n    model = tf.keras.models.Sequential()\n    model.add(tf.keras.Input(shape=(cfg.window_size, cfg.n_features), dtype='float16'))\n    for _ in range(cfg.model_nblocks):\n        model.add(tf.keras.layers.Conv1D(filters=cfg.model_hidden, kernel_size=15, padding=\"same\"))\n        model.add(tf.keras.layers.BatchNormalization())\n        model.add(tf.keras.layers.ReLU())\n        model.add(tf.keras.layers.Dropout(cfg.model_dropout))\n    model.add(tf.keras.layers.GlobalAveragePooling1D())\n    model.add(tf.keras.layers.Dense(cfg.n_labels, activation=None))\n    \n    if checkpoint_path is not None:\n        model.load_weights(checkpoint_path)\n    model.compile(tf.keras.optimizers.Adam(learning_rate=cfg.lr), \n                  loss = tf.keras.losses.BinaryCrossentropy(from_logits=True),\n                  metrics=[AveragePrecision(cfg.n_labels, thresholds=0.0)]\n                 )\n    \n    return model\n\n\nget_model().summary()","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:55.278490Z","iopub.execute_input":"2023-06-09T23:53:55.278851Z","iopub.status.idle":"2023-06-09T23:53:56.845739Z","shell.execute_reply.started":"2023-06-09T23:53:55.278820Z","shell.execute_reply":"2023-06-09T23:53:56.844997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br>\n\n## Train","metadata":{}},{"cell_type":"code","source":"def predict_select_model(fold, ds, model_save_dir=''):\n    best_path, best_score = None, -1\n    print(f\"Validation for fold{fold}:\")\n    for model_path in sorted(glob.glob(os.path.join(model_save_dir, f\"fold{fold}_*.h5\"))):\n        pred_time = perf_counter()\n        gc.collect()\n        score = calculate_precision(ds.values[ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], \n                                    expit(get_model(model_path).predict(ds, verbose=0)) # expit converts to sigmoid output\n                                   )\n        if best_score < score:\n            best_score = score\n            best_path = model_path\n        \n        gc.collect()\n        \n        print(\"\\t\", os.path.basename(model_path), f\": score-{score:.4f} in {(perf_counter()-pred_time)/60:.2f} mins\")\n        \n    print(os.path.basename(best_path), \"selected with score\", best_score)\n    \n    return best_path","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:56.846887Z","iopub.execute_input":"2023-06-09T23:53:56.847221Z","iopub.status.idle":"2023-06-09T23:53:56.860536Z","shell.execute_reply.started":"2023-06-09T23:53:56.847190Z","shell.execute_reply":"2023-06-09T23:53:56.859787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_loop(train_paths, valid_paths, fold, model_save_dir=''):\n    gc.collect()\n    \n    # train_paths, test_paths = train_test_split(train_paths, test_size=0.4)\n    train_ds = FOGSequence(train_paths)\n    val_ds = FOGSequence(valid_paths, split=\"val\")\n    # test_ds = FOGSequence(test_paths, split=\"val\")\n    \n    model = get_model()\n    ckpt = tf.keras.callbacks.ModelCheckpoint(os.path.join(model_save_dir, f\"fold{fold}_model_\"+\"{epoch:02d}.h5\"), \n                                              save_weights_only=True)\n    history = model.fit(train_ds, \n                        epochs=cfg.num_epochs, \n                        verbose=2, \n                        workers=5, \n                        validation_data=val_ds, \n                        use_multiprocessing=True, \n                        callbacks=[ckpt])\n    \n    best_model_path = predict_select_model(fold, val_ds, model_save_dir)\n    # score = calculate_precision(test_ds.values[test_ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], expit(get_model(best_model_path).predict(test_ds, verbose=0)))\n    \n    del train_ds, val_ds, model, ckpt, history\n    gc.collect()\n    \n    return best_model_path","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:56.861859Z","iopub.execute_input":"2023-06-09T23:53:56.862181Z","iopub.status.idle":"2023-06-09T23:53:56.874941Z","shell.execute_reply.started":"2023-06-09T23:53:56.862148Z","shell.execute_reply":"2023-06-09T23:53:56.873712Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Main training loop\n#【Ping】10-folds training for each module, each fold trains with 5 epochs \n#       and each epoch save model with weights-only. model_paths only stores best_model_path for each fold each module.\n\n#【Ping】Comment out the code below after training is done and defore submitting.\n#       and manually define `model_paths` like the code cell below.\n\nmodel_paths = {'defog': [], 'tdcsfog':[]}\nfor module in model_paths:\n    module_start = perf_counter()\n    print(f\"*** Training {module} {'*'*75}\")\n    if not os.path.exists(module): \n        os.mkdir(module)\n    for fold, (train_fpaths, valid_fpaths) in enumerate(zip(fold_train_fpaths[module], fold_valid_fpaths[module])):\n        fold_start = perf_counter()\n        print(f\"Fold {fold}{'-'*25}\")\n        model_paths[module].append(train_loop(train_fpaths, valid_fpaths, fold, model_save_dir=module))\n        print(f\"Fold {fold} done in {(perf_counter()-fold_start)/60:.2f} min\")\n        print()\n    print(f\"***{module} done in {(perf_counter()-module_start)/3600:.2f} hrs{'*'*50}\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-06-09T23:53:56.876621Z","iopub.execute_input":"2023-06-09T23:53:56.877033Z","iopub.status.idle":"2023-06-10T06:20:05.309794Z","shell.execute_reply.started":"2023-06-09T23:53:56.877001Z","shell.execute_reply":"2023-06-10T06:20:05.308633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_paths","metadata":{"execution":{"iopub.status.busy":"2023-06-10T06:25:24.837385Z","iopub.execute_input":"2023-06-10T06:25:24.837976Z","iopub.status.idle":"2023-06-10T06:25:24.846756Z","shell.execute_reply.started":"2023-06-10T06:25:24.837936Z","shell.execute_reply":"2023-06-10T06:25:24.845730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#【Ping】 When comment out the `Main training loop`, add this code cell before submitting.\n\n# model_paths = {'defog': ['defog/fold0_model_05.h5',\n#                          'defog/fold1_model_03.h5',\n#                          'defog/fold2_model_01.h5',\n#                          'defog/fold3_model_02.h5',\n#                          'defog/fold4_model_05.h5',\n#                          'defog/fold5_model_02.h5',\n#                          'defog/fold6_model_02.h5',\n#                          'defog/fold7_model_03.h5',\n#                          'defog/fold8_model_04.h5',\n#                          'defog/fold9_model_01.h5'],\n               \n#                'tdcsfog': ['tdcsfog/fold0_model_01.h5',\n#                            'tdcsfog/fold1_model_05.h5',\n#                            'tdcsfog/fold2_model_01.h5',\n#                            'tdcsfog/fold3_model_04.h5',\n#                            'tdcsfog/fold4_model_05.h5',\n#                            'tdcsfog/fold5_model_05.h5',\n#                            'tdcsfog/fold6_model_01.h5',\n#                            'tdcsfog/fold7_model_01.h5',\n#                            'tdcsfog/fold8_model_03.h5',\n#                            'tdcsfog/fold9_model_05.h5']}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<br>\n\n# Submission","metadata":{}},{"cell_type":"code","source":"test_defog_paths = glob.glob(os.path.join(TEST_DIR, \"defog/*.csv\"))\ntest_tdcsfog_paths = glob.glob(os.path.join(TEST_DIR, \"tdcsfog/*.csv\"))\n\ntest_ds_dict = {'defog':FOGSequence(test_defog_paths, split=\"test\"), \n                'tdcsfog':FOGSequence(test_tdcsfog_paths, split=\"test\")}\n\n# Get test predictions\ndf_list = []\nfor module, test_ds in test_ds_dict.items():\n    y_pred_list = []\n    for model_path in model_paths[module]:\n        model = get_model(model_path)\n        y_pred_list.append(expit(model.predict(test_ds, verbose=0)))  # expit converts to sigmoid output\n    y_pred = np.mean(y_pred_list, axis=0)\n    df_list.append(pd.DataFrame({'Id': test_ds.Ids, \n                                 'StartHesitation': y_pred[:,0], \n                                 'Turn': y_pred[:,1], \n                                 'Walking': y_pred[:,2]}))","metadata":{"execution":{"iopub.status.busy":"2023-06-10T06:26:12.020692Z","iopub.execute_input":"2023-06-10T06:26:12.021712Z","iopub.status.idle":"2023-06-10T06:26:57.209763Z","shell.execute_reply.started":"2023-06-10T06:26:12.021663Z","shell.execute_reply":"2023-06-10T06:26:57.208721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate Prediction to DataFrames\nsubmission = pd.concat(df_list)\n\n# Only keep Ids in sample_submission\nsample_submission = pd.read_csv(os.path.join(BASE_DIR, \"sample_submission.csv\"))\nsubmission = pd.merge(sample_submission[['Id']], submission, how='left', on='Id').fillna(0.0)\n\n# round to 5 decimal places while keeping point notation\nsubmission.to_csv(\"submission.csv\", index=False, float_format='%.5f') ","metadata":{"execution":{"iopub.status.busy":"2023-06-10T06:28:21.742768Z","iopub.execute_input":"2023-06-10T06:28:21.743976Z","iopub.status.idle":"2023-06-10T06:28:25.267970Z","shell.execute_reply.started":"2023-06-10T06:28:21.743930Z","shell.execute_reply":"2023-06-10T06:28:25.266970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!head -5 submission.csv","metadata":{"execution":{"iopub.status.busy":"2023-06-10T06:28:30.504545Z","iopub.execute_input":"2023-06-10T06:28:30.505827Z","iopub.status.idle":"2023-06-10T06:28:31.751486Z","shell.execute_reply.started":"2023-06-10T06:28:30.505784Z","shell.execute_reply":"2023-06-10T06:28:31.750250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}