{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Parkinson's Freezing of Gait Prediction\nAn estimated 7 to 10 million people around the world have Parkinson’s disease, many of whom suffer from freezing of gait (FOG). During a FOG episode, a patient's feet are “glued” to the ground, preventing them from moving forward despite their attempts. FOG has a profound negative impact on health-related quality of life—people who suffer from FOG are often depressed, have an increased risk of falling, are likelier to be confined to wheelchair use, and have restricted independence.\n\nWhile researchers have multiple theories to explain when, why, and in whom FOG occurs, there is still no clear understanding of its causes. The ability to objectively and accurately quantify FOG is one of the keys to advancing its understanding and treatment. Collection and analysis of FOG events, such as with your data science skills, could lead to potential treatments.\n\nThere are many methods of evaluating FOG, though most involve FOG-provoking protocols. People with FOG are filmed while performing certain tasks that are likely to increase its occurrence. Experts then review the video to score each frame, indicating when FOG occurred. While scoring in this manner is relatively reliable and sensitive, it is extremely time-consuming and requires specific expertise. Another method involves augmenting FOG-provoking testing with wearable devices. With more sensors, the detection of FOG becomes easier, however, compliance and usability may be reduced. Therefore, a combination of these two methods may be the best approach. When combined with machine learning methods, the accuracy of detecting FOG from a lower back accelerometer is relatively high. However, the datasets used to train and test these algorithms have been relatively small and generalizability is limited to date. Furthermore, the emphasis has been on achieving high levels of accuracy, while precision, for example, has largely been ignored.","metadata":{}},{"cell_type":"code","source":"# Basic Imports\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n# Basic OS Imports\nimport os\nfrom os.path import basename, dirname, join, exists\n\n# Extra Basic Imports\nimport gc\nfrom glob import glob\n# Default Random Number Generator\nfrom numpy.random import default_rng\n# TQDM Imports\nfrom tqdm.auto import tqdm\n\n# Time libraries\nfrom time import perf_counter\n\n# Function & Collection Libraries\nfrom functools import partial\nfrom collections import defaultdict as dd\n\n# Algorithm Imports\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, StratifiedGroupKFold\nfrom sklearn.metrics import average_precision_score\nfrom sklearn.preprocessing import StandardScaler as Scaler\n# SciPy Imports \nfrom scipy.special import expit\n# Tensor Flow Imports\nimport tensorflow as tf\nprint(f\"TF version: {tf.__version__}\")\nAUTO = tf.data.experimental.AUTOTUNE","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:29.621336Z","iopub.execute_input":"2023-04-15T10:02:29.621614Z","iopub.status.idle":"2023-04-15T10:02:40.198486Z","shell.execute_reply.started":"2023-04-15T10:02:29.621581Z","shell.execute_reply":"2023-04-15T10:02:40.196815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration\nIn this configuration we define and get the training & test dataset details from input data. We try to fetch it via figuring out the common details among the input data.","metadata":{}},{"cell_type":"code","source":"# Directory defines \n_base_dir = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction\"\n# Train directory deifne\n_train_dir = join(_base_dir, \"train\")\n# Test Directory define\n_test_dir = join(_base_dir, \"test\")\n\n# Config Class \nclass Config:\n    train_dirs = [\n        join(_train_dir, \"defog\"),\n        join(_train_dir, \"tdcsfog\")\n    ]\n    \n    metadata_dirs = [\n        join(_base_dir, \"defog_metadata.csv\"),\n        join(_base_dir, \"tdcsfog_metadata.csv\")\n    ]\n    # Split for Kfold\n    splits = 10\n    # Hyper parameter declarations\n    batch_size = 1024\n    window_size = 64\n    window_future = 16\n    window_past = window_size - window_future\n    \n    wx = 8\n    # Drop out & Block variable\n    model_dropout = 0.2\n    model_hidden = 128\n    model_nblocks = 3\n    # Learning Rate & epochs\n    lr = 0.00015\n    num_epochs = 5\n    # Lists\n    feature_list = ['AccV', 'AccML', 'AccAP']\n    label_list = ['StartHesitation', 'Turn', 'Walking']\n    # Numbers of lists\n    n_features = len(feature_list)\n    n_labels = len(label_list)","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.203322Z","iopub.execute_input":"2023-04-15T10:02:40.204453Z","iopub.status.idle":"2023-04-15T10:02:40.21793Z","shell.execute_reply.started":"2023-04-15T10:02:40.204412Z","shell.execute_reply":"2023-04-15T10:02:40.216747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calling the Config Class\ncfg = Config()","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.21988Z","iopub.execute_input":"2023-04-15T10:02:40.220679Z","iopub.status.idle":"2023-04-15T10:02:40.228013Z","shell.execute_reply.started":"2023-04-15T10:02:40.220635Z","shell.execute_reply":"2023-04-15T10:02:40.226576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Engineering\nAs from the data field it is mentioned that we have to split the data. This test set contains about 250 data series. The series from the `tdcsfog` and `defog` are in a proportion similar to that of the training set. These series have subjects that are entirely distinct from those in the training set. This also holds true for the public / private test split. To split this we use **StratifiedKFold Algorithm**.\n","metadata":{}},{"cell_type":"code","source":"# Map Id and Subject fields\ndf_id_subject = pd.concat([pd.read_csv(f, usecols=['Id', 'Subject']).assign(Module=basename(f).split('_')[0]) for f in cfg.metadata_dirs\n]).astype(\"category\").set_index(\"Id\")\nprint(f\" df_id_subject length: {len(df_id_subject)}\\n unique Ids: {df_id_subject.index.nunique()}\\n unique Subjects: {df_id_subject.Subject.nunique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.232843Z","iopub.execute_input":"2023-04-15T10:02:40.23349Z","iopub.status.idle":"2023-04-15T10:02:40.283014Z","shell.execute_reply.started":"2023-04-15T10:02:40.233445Z","shell.execute_reply":"2023-04-15T10:02:40.281774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read csv files and add metadata\ndef read_add_metadata(filepath, usecols, getid=False, getsub=False, getevent=False, dtype=None, exclude=['notype']):\n    fog_type = basename(dirname(filepath))\n    if fog_type in exclude:\n        return None\n    df = pd.read_csv(filepath, index_col=\"Time\", usecols=usecols, dtype=dtype)\n    if getid:\n        df['Id'] = basename(filepath).split('.')[0] + '_' + df.index.astype(str)\n    if getsub:\n        df['Subject'] = df_id_subject.loc[basename(filepath).split('.')[0], 'Subject']\n    if getevent:\n        df['Event'] = np.select(\n            [df[col].astype(bool) for col in cfg.label_list], \n            np.arange(1,cfg.n_labels+1), default=0\n        ).astype('int8')\n    return df","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.284505Z","iopub.execute_input":"2023-04-15T10:02:40.284875Z","iopub.status.idle":"2023-04-15T10:02:40.297919Z","shell.execute_reply.started":"2023-04-15T10:02:40.284845Z","shell.execute_reply":"2023-04-15T10:02:40.296665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Train dataframe details\ntrain_dirs = glob(join(_train_dir, '*/*.csv'))\ndtype = {col:'int8' for col in cfg.label_list}\ndtype['Time'] = 'int32'\nusecols = ['Time', *cfg.label_list]","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.29976Z","iopub.execute_input":"2023-04-15T10:02:40.300054Z","iopub.status.idle":"2023-04-15T10:02:40.440562Z","shell.execute_reply.started":"2023-04-15T10:02:40.300026Z","shell.execute_reply":"2023-04-15T10:02:40.43957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Call reader & add metadata function\nvar_train_reader = partial(read_add_metadata, usecols=usecols, dtype=dtype, getsub=True, getevent=True)\ntrain_df = pd.concat([var_train_reader(f) for f in tqdm(train_dirs)]).reset_index(drop=True)\ntrain_df.Subject = train_df.Subject.astype('category')\n\n# Display the train dataframe\ndisplay(train_df.Event.value_counts().to_frame().style.background_gradient())","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:02:40.443679Z","iopub.execute_input":"2023-04-15T10:02:40.443969Z","iopub.status.idle":"2023-04-15T10:03:19.78078Z","shell.execute_reply.started":"2023-04-15T10:02:40.443943Z","shell.execute_reply":"2023-04-15T10:03:19.77967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save paths for each Stratified Group Fold for defog and tdcsfog separately\n# Call StratifiedGroupKFold algorithm functions\nstr_gkf = StratifiedGroupKFold(n_splits=cfg.splits, random_state=42, shuffle=True)\n\n# Different train/validation dataframe paths\nfold_train_fpaths, fold_valid_fpaths = {'defog': [], 'tdcsfog':[]}, {'defog': [], 'tdcsfog':[]}\n\n# Data frame paths\ndf_paths = {'defog':glob(join(cfg.train_dirs[0],'*.csv')), 'tdcsfog': glob(join(cfg.train_dirs[1],'*.csv'))}\n\nfor module, paths in df_paths.items():\n    print(f\"{module}:\")\n    sub_train_df = train_df[train_df.Subject.isin(df_id_subject.loc[df_id_subject.Module==module, 'Subject'])].reset_index(drop=True)\n    \n    for i, (train_index, test_index) in enumerate(str_gkf.split(sub_train_df.index, sub_train_df.Event, groups=sub_train_df.Subject)):\n        print(f\"\\tFold {i}:\", end=\" \")\n        train_subs = sub_train_df.loc[train_index, 'Subject'].unique()\n        test_subs = sub_train_df.loc[test_index, 'Subject'].unique()\n        print(f\"Subjects->train:{len(train_subs)}|test:{len(test_subs)}\")\n        train_ids = set(df_id_subject[df_id_subject.Subject.isin(train_subs)].index)\n        test_ids = set(df_id_subject[df_id_subject.Subject.isin(test_subs)].index)\n        fold_train_fpaths[module].append([f for f in paths if basename(f).split('.')[0] in train_ids])\n        fold_valid_fpaths[module].append([f for f in paths if basename(f).split('.')[0] in test_ids])\n        del train_subs, test_subs, train_ids, test_ids\n        \n        # Garbage Collector\n        gc.collect()\n    del sub_train_df\ndel train_df\n\n# Garbage Collector\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:03:19.782219Z","iopub.execute_input":"2023-04-15T10:03:19.78297Z","iopub.status.idle":"2023-04-15T10:04:35.888282Z","shell.execute_reply.started":"2023-04-15T10:03:19.78293Z","shell.execute_reply":"2023-04-15T10:04:35.887186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Adapted from FOGDataset of \n# https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\nclass FOGSequence(tf.keras.utils.Sequence):\n\n    def __init__(self, df_paths, cfg=cfg, split=\"train\"):\n        _time = perf_counter()\n        \n        self.rng = default_rng(42)\n        self.cfg = cfg\n        self.split = split\n        \n        self.past_pad = self.cfg.wx*(self.cfg.window_past-1)\n        self.future_pad = self.cfg.wx*self.cfg.window_future\n        \n        if self.split == \"test\":\n            self.Ids = []\n        _values = [self._read(f) for f in df_paths]\n        \n        self.mapping = []\n        _length = 0\n        for _value in _values:\n            _shape = _value.shape[0]\n            self.mapping.extend(range(_length+self.past_pad, _length+_shape-self.future_pad))\n            _length += _shape\n            \n        self.values = np.concatenate(_values, axis=0)\n        self.mapping = np.array(self.mapping)\n        if self.split != \"test\":\n            # Keep only vaild and task rows\n            _valid_pos = self.values[self.mapping,self.valid_position] > 0\n            _task_pos = self.values[self.mapping,self.task_position] > 0\n            self.mapping = self.mapping[_valid_pos&_task_pos]\n        self.length = self.mapping.shape[0]\n        \n        print(f\"Valid Dataset of size {self.length:,} initialized in {perf_counter() - _time:.3f} secs!\")\n        gc.collect()\n    \n    def _read(self, path):\n        _is_tdcs = basename(dirname(path)).startswith('tdcs')\n        df = pd.read_csv(path)\n        \n        if self.split == \"test\":\n            _ids = basename(path).split('.')[0] + '_' + df.Time.astype(str)\n            self.Ids.extend(_ids.tolist())\n            return self._df_to_array(df, self.cfg.feature_list)\n        \n        _cols = [*self.cfg.feature_list, *self.cfg.label_list, 'Valid', 'Task']\n        self.valid_position = self.cfg.n_features + self.cfg.n_labels\n        self.task_position = self.valid_position + 1\n        \n        if _is_tdcs:\n            # Fill Valid and Task columns for tdcsfog\n            df['Valid'] = 1\n            df['Task'] = 1\n            \n        return self._df_to_array(df, _cols)\n    \n    def _df_to_array(self, df, cols):\n        # Pads past and future rows to dataframe values for indexing \n        _values = df[cols].values.astype(np.float16)\n        return np.pad(_values, ((self.past_pad, self.future_pad),(0,0)), 'edge')\n    \n    def __len__(self):\n        return int(np.ceil(self.length / self.cfg.batch_size))\n    \n    def __getitem__(self, idx):\n        \n        if self.split == \"train\":\n            # Onlt train set has randomly selected batches\n            _idxs = self.rng.choice(self.mapping, size=self.cfg.batch_size, replace=False)\n        else:\n            _idxs = self._get_indices(idx)\n            \n        # For test return only features\n        if self.split == \"test\":\n            return self._get_X(_idxs)\n        # For train and val splits return y also\n        return self._get_X_y(_idxs)\n    \n    def _get_indices(self, idx):\n        _low = idx * self.cfg.batch_size\n        # Cap high at self.length so overflow does not occur\n        _high = min(_low + self.cfg.batch_size, self.length)\n        return self.mapping[_low:_high]\n    \n    def _get_X(self, indices):\n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            _X[i] = self.values[idx-self.past_pad:idx+self.future_pad+1:self.cfg.wx, :self.cfg.n_features]\n        return _X\n    \n    def _get_X_y(self, indices):\n        _X = np.empty((len(indices), self.cfg.window_size, self.cfg.n_features), dtype=np.float16)\n        for i, idx in enumerate(indices):\n            _X[i] = self.values[idx-self.past_pad: idx+self.future_pad+1:self.cfg.wx, :self.cfg.n_features]\n        return _X, self.values[indices, self.cfg.n_features:self.cfg.n_features+self.cfg.n_labels]","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:35.890142Z","iopub.execute_input":"2023-04-15T10:04:35.890545Z","iopub.status.idle":"2023-04-15T10:04:35.912066Z","shell.execute_reply.started":"2023-04-15T10:04:35.890506Z","shell.execute_reply":"2023-04-15T10:04:35.911005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Model","metadata":{}},{"cell_type":"code","source":"# average_precision_score with positive sample added if no true positive cases are present\ndef calculate_precision(y_true, y_pred):\n    pad_width = ((0,0),(0,0)) if y_true.any(axis=0).all() else ((1,0),(0,0))\n    y_true, y_pred = np.pad(y_true, pad_width, constant_values=1), np.pad(y_pred, pad_width, constant_values=1)\n    return average_precision_score(y_true, y_pred)","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:35.916478Z","iopub.execute_input":"2023-04-15T10:04:35.916826Z","iopub.status.idle":"2023-04-15T10:04:35.927599Z","shell.execute_reply.started":"2023-04-15T10:04:35.916786Z","shell.execute_reply":"2023-04-15T10:04:35.926504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Note: Not same result as average_precision_score\nclass AveragePrecision(tf.keras.metrics.Metric):\n\n    def __init__(self, num_classes, thresholds=None, name='avg_precision', **kwargs):\n        super(AveragePrecision, self).__init__(name=name, **kwargs)\n        self.class_precision = [tf.keras.metrics.Precision(thresholds) for _ in range(num_classes)]\n\n    def update_state(self, y_true, y_pred, sample_weight=None):\n        for i, precision in enumerate(self.class_precision):\n            precision.update_state(y_true[...,i], y_pred[..., i])\n\n    def result(self):\n        return tf.math.reduce_mean([precision.result() for precision in self.class_precision])\n    \n    def reset_state(self):\n        for precision in self.class_precision:\n            precision.reset_state()","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:35.9296Z","iopub.execute_input":"2023-04-15T10:04:35.930206Z","iopub.status.idle":"2023-04-15T10:04:35.942368Z","shell.execute_reply.started":"2023-04-15T10:04:35.930058Z","shell.execute_reply":"2023-04-15T10:04:35.941091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model adapted from \n# https://keras.io/examples/timeseries/timeseries_classification_from_scratch/\ndef get_model(checkpoint_path = None):\n    model = tf.keras.models.Sequential()\n    model.add(tf.keras.Input(shape=(cfg.window_size, cfg.n_features), dtype='float16'))\n    for _ in range(cfg.model_nblocks):\n        model.add(tf.keras.layers.Conv1D(filters=cfg.model_hidden, kernel_size=15, padding=\"same\"))\n        model.add(tf.keras.layers.BatchNormalization())\n        model.add(tf.keras.layers.ReLU())\n        model.add(tf.keras.layers.Dropout(cfg.model_dropout))\n    model.add(tf.keras.layers.GlobalAveragePooling1D())\n    model.add(tf.keras.layers.Dense(cfg.n_labels, activation=None))\n\n    if checkpoint_path is not None:\n        model.load_weights(checkpoint_path)\n    model.compile(\n        tf.keras.optimizers.Adam(learning_rate=cfg.lr), \n        loss = tf.keras.losses.BinaryCrossentropy(from_logits=True),\n        metrics=[AveragePrecision(cfg.n_labels, thresholds=0.0)]\n    )\n    return model\n\nget_model().summary()","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:35.944196Z","iopub.execute_input":"2023-04-15T10:04:35.944632Z","iopub.status.idle":"2023-04-15T10:04:39.882963Z","shell.execute_reply.started":"2023-04-15T10:04:35.944556Z","shell.execute_reply":"2023-04-15T10:04:39.882115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Model","metadata":{}},{"cell_type":"code","source":"def predict_select_model(fold, ds, model_save_dir=''):\n    best_path, best_score = None, -1\n    print(f\"Validation for fold{fold}:\")\n    for model_path in sorted(glob(join(model_save_dir, f\"fold{fold}_*.h5\"))):\n        pred_time = perf_counter()\n        gc.collect()\n        score = calculate_precision(\n            ds.values[ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], \n            expit(get_model(model_path).predict(ds, verbose=0)) # expit converts to sigmoid output\n        )\n        if best_score < score:\n            best_score = score\n            best_path = model_path\n        gc.collect()\n        print(\"\\t\", basename(model_path), f\": score-{score:.4f} in {(perf_counter()-pred_time)/60:.2f} mins\")\n    print(basename(best_path), \"selected with score\", best_score)\n    return best_path","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:39.884055Z","iopub.execute_input":"2023-04-15T10:04:39.88458Z","iopub.status.idle":"2023-04-15T10:04:39.901886Z","shell.execute_reply.started":"2023-04-15T10:04:39.884549Z","shell.execute_reply":"2023-04-15T10:04:39.90112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_loop(train_paths, valid_paths, fold, model_save_dir=''):\n    gc.collect()\n    \n    # train_paths, test_paths = train_test_split(train_paths, test_size=0.4)\n    train_ds = FOGSequence(train_paths)\n    val_ds = FOGSequence(valid_paths, split=\"val\")\n    # test_ds = FOGSequence(test_paths, split=\"val\")\n    \n    model = get_model()\n    ckpt = tf.keras.callbacks.ModelCheckpoint(join(model_save_dir, f\"fold{fold}_model_\"+\"{epoch:02d}.h5\"), save_weights_only=True)\n    history = model.fit(train_ds, epochs=cfg.num_epochs, verbose=2, workers=5, validation_data=val_ds, use_multiprocessing=True, callbacks=[ckpt])\n    \n    best_model_path = predict_select_model(fold, val_ds, model_save_dir)\n    # score = calculate_precision(test_ds.values[test_ds.mapping, cfg.n_features:cfg.n_features + cfg.n_labels], expit(get_model(best_model_path).predict(test_ds, verbose=0)))\n    \n    del train_ds, val_ds, model, ckpt, history\n    gc.collect()\n    return best_model_path","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:39.902845Z","iopub.execute_input":"2023-04-15T10:04:39.903212Z","iopub.status.idle":"2023-04-15T10:04:39.915002Z","shell.execute_reply.started":"2023-04-15T10:04:39.903176Z","shell.execute_reply":"2023-04-15T10:04:39.913779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Main training loop\nmodel_paths = {'defog': [], 'tdcsfog':[]}\nfor module in model_paths:\n    module_start = perf_counter()\n    print(f\"***Training {module}{'*'*75}\")\n    if not exists(module): \n        os.mkdir(module)\n    for fold, (train_fpaths, valid_fpaths) in enumerate(zip(fold_train_fpaths[module], fold_valid_fpaths[module])):\n        fold_start = perf_counter()\n        print(f\"Fold {fold}{'-'*25}\")\n        model_paths[module].append(train_loop(train_fpaths, valid_fpaths, fold, model_save_dir=module))\n        print(f\"Fold {fold} done in {(perf_counter()-fold_start)/60:.2f} min\")\n    print(f\"***{module} done in {(perf_counter()-module_start)/3600:.2f} hrs{'*'*50}\\n\")","metadata":{"execution":{"iopub.status.busy":"2023-04-15T10:04:39.916395Z","iopub.execute_input":"2023-04-15T10:04:39.916703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction & Submission","metadata":{}},{"cell_type":"code","source":"test_defog_paths = glob(join(_test_dir, \"defog/*.csv\"))\ntest_tdcsfog_paths = glob(join(_test_dir, \"tdcsfog/*.csv\"))\n\ntest_ds_dict = {\n    'defog':FOGSequence(test_defog_paths, split=\"test\"), \n    'tdcsfog':FOGSequence(test_tdcsfog_paths, split=\"test\")\n}\n\n# Get test predictions\ndf_list = []\nfor module, test_ds in test_ds_dict.items():\n    y_pred_list = []\n    for model_path in model_paths[module]:\n        model = get_model(model_path)\n        y_pred_list.append(expit(model.predict(test_ds, verbose=0)))  # expit converts to sigmoid output\n    y_pred = np.mean(y_pred_list, axis=0)\n    df_list.append(pd.DataFrame(\n        {'Id': test_ds.Ids, 'StartHesitation': y_pred[:,0], 'Turn': y_pred[:,1], 'Walking': y_pred[:,2]}))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Concatenate Prediction to DataFrames\nsubmission = pd.concat(df_list)\n\n# Only keep Ids in sample_submission\nsample_submission = pd.read_csv(join(_base_dir, \"sample_submission.csv\"))\nsubmission = pd.merge(sample_submission[['Id']], submission, how='left', on='Id').fillna(0.0)\nsubmission.to_csv(\"submission.csv\", index=False, float_format='%.5f')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}