{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Keras Template\nI want to share this template for using keras models to possibly help and get helped on good and bad usage of the keras library in time-series context.\n\n*If you think something is incorrect or could be improved, please comment*, in particular regarding:\n* model *architecture* and usage\n* approach to time-series *subsequencing*\n* approach to *out-of-core* training\n* why the model is not learining and approach to extreme *unbalancement*","metadata":{"_uuid":"7c741a89-73d3-4244-bbee-6a699989139f","_cell_guid":"338b00fd-8109-44e0-990b-3f0a8a6f20ec","trusted":true}},{"cell_type":"markdown","source":"# Notebook Phylosophy\n\n* **keras** based: for now using an LSTM but easily convertible to other architectures\n* **multioutput** model: 4 target columns (no event included for softmax convenience)\n* **out-of-core learning**: load one time-series (csv file) at a time, extract features, partial fit; this is to avoid memory errors. The current dataset is borderline and may not need this approach but I want to use it anyway for personal learning\n* **subsequences approach**: models need fixed-length input sequences, thus we need to split each time-series (csv file) in subsequences (windows) of equal length, being careful to not generate subsequences that span on two different time-series. Is also important to assign the correct target to each specific subsequence: in this case I wanted to use as features N previous measures and as target the one of the last measure of the sequence.\n\n    For this purpose I used the function `numpy.lib.stride_tricks.sliding_window_view`, alternatively `tf.keras.preprocessing.sequence.TimeseriesGenerator` may be used, but is [deprecated](https://www.tensorflow.org/api_docs/python/tf/keras/preprocessing/sequence/TimeseriesGenerator) so I decided to avoid it .\n* **separate models** for de and tdcs dataset: this might be changed but need an upsampling of the de, or a downsampling of the tdcs datasets\n\n* no usage of features based on **Time** progression: I'm not here to win and I think its usage wuold be harmful for the competition host and the beautiful aim of this challenge.\n\n* **metadata** not included, my first try lead to a decay in submission score","metadata":{"_uuid":"1dd82129-e665-47cc-ad02-e587542b6248","_cell_guid":"6c501188-942c-49f8-84a5-4be8ad6411ef","trusted":true}},{"cell_type":"markdown","source":"# Problems so far\n* model is **not learning**, outputting basically just 0s; possible solutions: \n    * better dataset balancement (tried with bad results)\n    * weight samples with rare events (tried with bad results)\n* using GPU cause an **out-of-memory** error on the hidden test set, even though I'm using out-of-core feature extraction and prediction and the `model.predict` function has a fixed batch size: probably because when enabling GPU, RAM is reduced and test time-series are very long\n* training and prediction **speed** may benefit from decorating some function with `tf.function` (see [this](https://www.tensorflow.org/guide/intro_to_graphs))","metadata":{"_uuid":"fb12c999-468f-43cd-8d65-c1f66782f816","_cell_guid":"8f140a61-d28b-4e96-b363-4456eab01e2c","trusted":true}},{"cell_type":"code","source":"from typing import Iterable, List, Tuple\nfrom joblib import Parallel, delayed, parallel_backend\nfrom sklearn.preprocessing import StandardScaler\nfrom tqdm import tqdm\nimport numpy as np, pandas as pd\nfrom numpy.lib.stride_tricks import sliding_window_view\nimport os\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nimport math\n\nimport keras.models","metadata":{"_uuid":"454aacd4-89ea-4286-a60f-522f0df462d3","_cell_guid":"36e49b72-92a3-468c-8261-49f8be320477","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:53:58.599627Z","iopub.execute_input":"2023-05-12T11:53:58.601003Z","iopub.status.idle":"2023-05-12T11:54:09.995596Z","shell.execute_reply.started":"2023-05-12T11:53:58.600940Z","shell.execute_reply":"2023-05-12T11:54:09.994104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Helper functions","metadata":{"_uuid":"8765d337-4782-4484-9627-b7de41d41a98","_cell_guid":"2e84aea6-800d-45e9-b1cd-da797054d357","trusted":true}},{"cell_type":"markdown","source":"## Data Loading\n* **de dataset**: only **correctly labeled** (Task and Valid are True) rows are loaded\n* **tdcs dataset**: units converted from **m/s2 -> g** (also a downsampling of this dataset or an upsampling of the de dataset should be performed for training the two datasets on a single model)","metadata":{"_uuid":"ce8c67e9-5094-4df4-b7d2-b75204c0fabc","_cell_guid":"4f8e4f7b-4c12-4dc7-af37-d9607836ef6f","trusted":true}},{"cell_type":"code","source":"def load_single_timeseries(ts_id:str) -> pd.DataFrame:\n    assert '.csv' not in ts_id\n    train_or_test, de_or_tdcs = find(ts_id)\n    path_to_timeseries = f\"{path_to_data}/{train_or_test}/{de_or_tdcs}fog/{ts_id}.csv\"\n    df = pd.read_csv(path_to_timeseries)\n    df['Id'] = ts_id\n    df = df.set_index('Id')\n    df = df.set_index('Time', append=True)\n    if train_or_test == 'train' and de_or_tdcs == 'de':\n        df = df[(df['Valid']) & df['Task']]\n        df = df.drop(columns=['Valid', 'Task'])\n    if de_or_tdcs == 'tdcs':\n        df.loc[:, ['AccV', 'AccML', 'AccAP']] = df[['AccV', 'AccML', 'AccAP']] / G_MS2\n    return df\n\n\ndef load_timeseries(ts_ids:Iterable[str], n_jobs:int=5) -> pd.DataFrame:\n    with parallel_backend('loky'):\n        ts = pd.concat(\n            Parallel(n_jobs=n_jobs)(delayed(load_single_timeseries)(ts_id) for ts_id in ts_ids)\n        )\n    return ts\n\n\ndef find(ts_id: str) -> (str, str):\n    if ts_id in get_ids(train_or_test='train', de_or_tdcs='de'):\n        train_or_test, de_or_tdcs = 'train', 'de'\n    elif ts_id in get_ids(train_or_test='train', de_or_tdcs='tdcs'):\n        train_or_test, de_or_tdcs = 'train', 'tdcs'\n    elif ts_id in get_ids(train_or_test='test', de_or_tdcs='de'):\n        train_or_test, de_or_tdcs = 'test', 'de'\n    elif ts_id in get_ids(train_or_test='test', de_or_tdcs='tdcs'):\n        train_or_test, de_or_tdcs = 'test', 'tdcs'\n    else:\n        raise Exception(f'{ts_id} not found!')\n    return train_or_test, de_or_tdcs\n\n\ndef get_ids(train_or_test:str, de_or_tdcs:str) -> List[str]:\n    this_folder = f\"{path_to_data}/{train_or_test}/{de_or_tdcs}fog\"\n    files = get_files_in_folder(this_folder)\n    ts_ids = [f.replace('.csv', '') for f in files]\n    return ts_ids\n\n\ndef get_files_in_folder(folder:str) -> List[str]:\n    return list(os.walk(folder))[0][2]","metadata":{"_uuid":"c2786ce0-9481-4741-9699-f3b09032b37c","_cell_guid":"7eb089cd-5520-412e-b712-75b19f5628e3","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:09.998277Z","iopub.execute_input":"2023-05-12T11:54:09.999435Z","iopub.status.idle":"2023-05-12T11:54:10.018343Z","shell.execute_reply.started":"2023-05-12T11:54:09.999379Z","shell.execute_reply":"2023-05-12T11:54:10.017409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset Balancement\n\n* just dropping time series with no events helped the score\n* balancing by randomly sampling rows of the dominant classes and repeating rows of rarer classes didn't help the score","metadata":{}},{"cell_type":"code","source":"def get_dataset_stats(de_or_tdcs:str) -> pd.DataFrame:\n    ts_ids = get_ids(train_or_test='train', de_or_tdcs=de_or_tdcs)\n    dataset = load_timeseries(ts_ids, n_jobs=-1)\n    out = dataset.groupby('Id')[['StartHesitation', 'Turn', 'Walking']].sum()\n    out = out.join(dataset.groupby('Id')['StartHesitation'].agg(count='count'))\n    out = out.assign(no_events=out['count'] - out['StartHesitation'] - out['Turn'] - out['Walking'])\n    out = out[['no_events'] + out.columns.drop('no_events').to_list()]\n    return out\n\n\ndef get_ids_with_events(de_or_tdcs:str, max_noevents_fraction:float=1) -> np.ndarray:  # this is the only balancing strategy that improved my score\n    \"\"\"\n    this is to get ids of time series that have at most the provided fraction of rows without positive events\n    (basically to drop ts without positive events)\n    \"\"\"\n    assert 0 < max_noevents_fraction <= 1  # 1 equals all ids\n    stats = get_dataset_stats(de_or_tdcs)\n    stats['no_events_fraction'] = stats['no_events'] / stats['count']\n    with_events = stats[stats['no_events_fraction'] <= max_noevents_fraction]\n    return with_events.index.values\n\n\ndef balance_classes(X:np.ndarray, y:pd.DataFrame, target_global_fractions:dict):  # this balancing strategy unfortunately worsen my score\n    n_targets = len(y.columns)\n    y_in = y.reset_index().set_index(y.index.names, append=True)  # keep track of progressive index to correctly select X\n    y_out = pd.DataFrame()\n    for target in y_in.columns:\n        correction = 1/(n_targets * target_global_fractions[target])\n        if target_global_fractions[target] >= 1 / n_targets:\n            balanced_class = y_in[y_in[target].eq(1)].sample(frac=correction)\n        else:\n            n_complete_repeats = math.floor(correction)\n            reminder = correction - n_complete_repeats\n            positive_events = y_in[y_in[target].eq(1)]\n            balanced_class = pd.concat([positive_events] * n_complete_repeats + [positive_events.sample(frac=reminder)])\n        y_out = pd.concat([y_out, balanced_class])\n    y_out = y_out.sort_index()\n    X_out = X[y_out.index.get_level_values(0).to_list(), :, :]\n    y_out = y_out.droplevel(0)\n    return X_out, y_out","metadata":{"execution":{"iopub.status.busy":"2023-05-12T11:54:10.031350Z","iopub.execute_input":"2023-05-12T11:54:10.031691Z","iopub.status.idle":"2023-05-12T11:54:10.051957Z","shell.execute_reply.started":"2023-05-12T11:54:10.031661Z","shell.execute_reply":"2023-05-12T11:54:10.050481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Subsequences generation","metadata":{"_uuid":"d66070a4-243d-474c-83da-7e2c7c7123c2","_cell_guid":"5e85d060-3958-4ade-8078-5cdcd7050e1c","trusted":true}},{"cell_type":"code","source":"def split_ts_in_subsequences(\n        ts:pd.DataFrame, stride:int, windows_length_samples:int\n) -> (np.ndarray, pd.DataFrame):\n    X = ts.reset_index()  # to keep track of the index\n    X = X.drop(columns=target_colnames, errors='ignore')\n    X = sliding_window_view(X, window_shape=windows_length_samples, axis=0)[:-1][::stride]\n    X = np.swapaxes(X, 1, 2)  # axes must be (sample, time, feature), numpy functions outputs (sample, feature, time)\n    last_index_of_sequence = pd.DataFrame(data=X[:, -1, :2], columns=ts.index.names)\n    for c in last_index_of_sequence:\n        last_index_of_sequence[c] = last_index_of_sequence[c].astype(ts.index.get_level_values(c).dtype)\n    last_index_of_sequence = pd.MultiIndex.from_frame(last_index_of_sequence)  # to correctly assign target\n    X = X[:, :, 2:].astype(float)  # remove index\n    if all(col in ts.columns for col in target_colnames):\n        y = ts.loc[last_index_of_sequence][target_colnames]\n    else:  # if unlabeled data returns an empty dataframe to keep track of the correct indices of the sequences\n        y = pd.DataFrame(index=last_index_of_sequence, columns=target_colnames)\n    return X, y","metadata":{"_uuid":"f7d2b88b-a13d-4818-b51d-8ffdc71f0d2d","_cell_guid":"16914b94-beae-4a6b-83af-a4d2fe0a9ae0","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.054015Z","iopub.execute_input":"2023-05-12T11:54:10.054421Z","iopub.status.idle":"2023-05-12T11:54:10.068113Z","shell.execute_reply.started":"2023-05-12T11:54:10.054388Z","shell.execute_reply":"2023-05-12T11:54:10.066800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature extraction\nI use as features a sequence of N previous measures and as target the label of the last measure","metadata":{"_uuid":"3d5358db-d326-4ef0-b26a-8796ae1184f2","_cell_guid":"6d99f213-8e27-4b6c-9106-3a4a477d2d6b","trusted":true}},{"cell_type":"code","source":"def scale(ts:pd.DataFrame, scaler_fitted) -> pd.DataFrame:\n    ts_out = ts.copy()\n    features_column_names = ts_out.columns.drop(target_colnames, errors='ignore').to_list()\n    X = ts_out[features_column_names]\n    ts_out.loc[:, features_column_names] = scaler_fitted.transform(X)\n    return ts_out","metadata":{"execution":{"iopub.status.busy":"2023-05-12T11:54:10.069966Z","iopub.execute_input":"2023-05-12T11:54:10.070362Z","iopub.status.idle":"2023-05-12T11:54:10.081643Z","shell.execute_reply.started":"2023-05-12T11:54:10.070332Z","shell.execute_reply":"2023-05-12T11:54:10.080399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_X_y(ts:pd.DataFrame, scaler_fitted, windows_length_samples:int, stride:int) -> (np.ndarray, pd.DataFrame):\n    ts = scale(ts, scaler_fitted)\n    X, y = split_ts_in_subsequences(\n        ts=ts, windows_length_samples=windows_length_samples, stride=stride\n    )\n    y['no_events'] = (y['StartHesitation'].eq(0) & y['Turn'].eq(0) & y['Walking'].eq(0)).astype(int)\n    return X, y","metadata":{"_uuid":"8e78f393-fd23-4c6f-887c-b2dc7f28c298","_cell_guid":"de91b2e2-5f9b-4190-91b7-49726266840a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.083300Z","iopub.execute_input":"2023-05-12T11:54:10.083796Z","iopub.status.idle":"2023-05-12T11:54:10.093340Z","shell.execute_reply.started":"2023-05-12T11:54:10.083765Z","shell.execute_reply":"2023-05-12T11:54:10.092096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training\n* **architecture**: \"standard\" LSTM, just as a baseline, then any keras architecture may be used\n* **regression or classification**: choosing the loss function MSE vs BinaryCrossentropy\n* **output probabilities**: adding the target column \"no event\" and using *softmax*\n* **out-of-core learning**: using the `model.train_on_batch` method for training separately on each time-series (csv file)","metadata":{"_uuid":"79ffbfc5-3dc7-4f53-8d86-f6879c091112","_cell_guid":"d576c380-ff5e-4e16-9796-21396776c0a9","trusted":true}},{"cell_type":"code","source":"def make_lstm(\n        input_shape:tuple, output_shape:int, lstm_n_units:int=64, learning_rate:float=1e-3, dropout_rate:float=.2\n) -> keras.Model:\n    import keras.layers, keras.activations, keras.optimizers, keras.losses, keras.metrics\n    input_layer = keras.layers.Input(input_shape)\n    lstm = keras.layers.Bidirectional(\n        keras.layers.LSTM(units=lstm_n_units)\n    )(input_layer)\n    dropout = keras.layers.Dropout(rate=dropout_rate)(lstm)\n    dense = keras.layers.Dense(units=output_shape)(dropout)\n    output_layer = keras.activations.softmax(dense)\n    model = keras.models.Model(inputs=input_layer, outputs=output_layer)\n    print(model.summary())\n    model.compile(\n        optimizer=keras.optimizers.Adam(learning_rate=learning_rate),\n        loss=keras.losses.BinaryCrossentropy(),\n        metrics=[keras.metrics.MAE, keras.metrics.AUC(curve='PR', multi_label=True, num_labels=4)]\n    )\n    return model","metadata":{"_uuid":"cb0624b4-32fb-4137-8e8f-a40820515201","_cell_guid":"08097682-f3b5-43be-8cbd-0827d558d526","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.099938Z","iopub.execute_input":"2023-05-12T11:54:10.100489Z","iopub.status.idle":"2023-05-12T11:54:10.112515Z","shell.execute_reply.started":"2023-05-12T11:54:10.100446Z","shell.execute_reply":"2023-05-12T11:54:10.111190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_on_one_ts(model, ts_id:str, scaler_fitted, windows_length_samples:int, stride:int) -> dict:\n    ts = load_single_timeseries(ts_id)\n    X, y = get_X_y(ts=ts, scaler_fitted=scaler_fitted, windows_length_samples=windows_length_samples, stride=stride)\n    if do_balance:\n        X, y = balance_classes(X=X, y=y, target_global_fractions=target_global_fractions[de_or_tdcs])  # this worsen my score\n    sample_weight = np.where(\n        y['StartHesitation'].eq(1), .9,\n        np.where(\n            y['Walking'].eq(1), .9,\n            np.where(\n                y['Turn'].eq(1), .6,\n                .1\n            )\n        )\n    )\n    metrics_dict = model.train_on_batch(X, y.values, return_dict=True)  # using sample_weights worsen the score\n    for target in y.columns:\n        metrics_dict[f\"n_{target}\"] = len(y[y[target].eq(1)])\n    metrics_dict['n_samples'] = len(y)\n    return metrics_dict","metadata":{"_uuid":"cb0624b4-32fb-4137-8e8f-a40820515201","_cell_guid":"08097682-f3b5-43be-8cbd-0827d558d526","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.114369Z","iopub.execute_input":"2023-05-12T11:54:10.114813Z","iopub.status.idle":"2023-05-12T11:54:10.126426Z","shell.execute_reply.started":"2023-05-12T11:54:10.114768Z","shell.execute_reply":"2023-05-12T11:54:10.125176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I wanted to parallelize for loops in those functions but I think is not possible to update models weights in parallel","metadata":{"_uuid":"01cdd953-8170-47d6-8809-446b79d4ec0d","_cell_guid":"5b9c89a2-ba87-42c8-b7e6-64745fbbe3e1","jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-04-23T13:12:30.801704Z","iopub.execute_input":"2023-04-23T13:12:30.802041Z","iopub.status.idle":"2023-04-23T13:12:30.816981Z","shell.execute_reply.started":"2023-04-23T13:12:30.802010Z","shell.execute_reply":"2023-04-23T13:12:30.814300Z"}}},{"cell_type":"code","source":"def train_scaler(ts_ids:Iterable[str]):\n    scaler = StandardScaler()\n    for ts_id in tqdm(ts_ids, desc='fit scaler'):\n        ts = load_single_timeseries(ts_id)\n        scaler.partial_fit(ts[acc_colnames])\n    return scaler","metadata":{"_uuid":"f8d2888f-483c-4aef-a45c-5410d1354aa1","_cell_guid":"2632589b-ad5c-45b2-84ee-a6f325f78fb0","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.128030Z","iopub.execute_input":"2023-05-12T11:54:10.128419Z","iopub.status.idle":"2023-05-12T11:54:10.144641Z","shell.execute_reply.started":"2023-05-12T11:54:10.128386Z","shell.execute_reply":"2023-05-12T11:54:10.142901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def weighted_average(df:pd.DataFrame, values_column_name:str, weights_column_name:str) -> pd.DataFrame:\n    return sum(df[weights_column_name] * df[values_column_name]) / df[weights_column_name].sum()","metadata":{"_uuid":"0c046fbb-5a77-4fa7-9ac8-ccc9bec5e119","_cell_guid":"771023e9-20f4-4e02-8ade-d8a13e324736","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.146978Z","iopub.execute_input":"2023-05-12T11:54:10.147404Z","iopub.status.idle":"2023-05-12T11:54:10.164233Z","shell.execute_reply.started":"2023-05-12T11:54:10.147373Z","shell.execute_reply":"2023-05-12T11:54:10.162903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train(ts_ids:Iterable[str], windows_length_samples:int, stride:int):\n    from keras import backend as K\n    scaler = train_scaler(ts_ids)\n    model = make_lstm(\n        input_shape=(windows_length_samples, len(acc_colnames)), output_shape=4,\n        lstm_n_units = lstm_configs['n_units'], learning_rate=lstm_configs['learning_rate'],\n        dropout_rate=lstm_configs['dropout_rate']\n    )\n    progress = tqdm(range(n_epochs_max),  desc='epoch')\n    for n_epoch in progress:\n        for n_ts, ts_id in enumerate(ts_ids):  # you can't parallelize partial fits\n            metrics_this_ts = train_on_one_ts(\n                model=model, ts_id=ts_id, scaler_fitted=scaler,\n                windows_length_samples=windows_length_samples, stride=stride,\n            )\n            metrics_this_ts = pd.DataFrame(\n                index=pd.MultiIndex.from_arrays([[ts_id], [n_epoch]], names=['Id', 'n_epoch']),\n                data=[metrics_this_ts]\n            )\n            metrics = metrics_this_ts.copy() if ts_id == ts_ids[0] and n_epoch == 0 else pd.concat([metrics, metrics_this_ts])\n        auc_colname = metrics.columns[metrics.columns.str.contains('auc')][0]  # unfortunately it varies with the fold\n        this_epoch_loss = weighted_average(metrics.xs(n_epoch, level='n_epoch'), 'loss', 'n_samples')\n        this_epoch_auc = weighted_average(metrics.xs(n_epoch, level='n_epoch'), auc_colname, 'n_samples')\n        if n_epoch == 0:\n            print(\"n samples in each class:\")\n            print({target: metrics[f\"n_{target}\"].sum() for target in ['no_events'] + target_colnames})\n        progress.set_postfix_str(f\"loss={this_epoch_loss:.3f}, auc={this_epoch_auc:.3f}\")\n        if n_epoch != 0:\n            metrics_sofar = metrics.drop(n_epoch, level='n_epoch')\n            best_loss_sofar = metrics_sofar.groupby('n_epoch').apply(weighted_average, 'loss', 'n_samples').min()\n            if best_loss_sofar - this_epoch_loss < min_loss_improv:  # TODO: train stop on mean_average_precision?\n                break\n            # K.set_value(model.optimizer.learning_rate, ?)  # TODO: reduce learning rate on condition?\n    return scaler, model, metrics","metadata":{"_uuid":"0c046fbb-5a77-4fa7-9ac8-ccc9bec5e119","_cell_guid":"771023e9-20f4-4e02-8ade-d8a13e324736","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.146978Z","iopub.execute_input":"2023-05-12T11:54:10.147404Z","iopub.status.idle":"2023-05-12T11:54:10.164233Z","shell.execute_reply.started":"2023-05-12T11:54:10.147373Z","shell.execute_reply":"2023-05-12T11:54:10.162903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualization","metadata":{}},{"cell_type":"code","source":"def plot_history(history:pd.DataFrame, title:str) -> go.Figure:\n    metrics_to_plot = [c for c in history.columns if \"n_\" not in c]\n    fig = make_subplots(rows=len(metrics_to_plot), cols=1, shared_xaxes='all', vertical_spacing=1e-5)\n    for i_col, col in enumerate(metrics_to_plot):\n        x_label, y_label =  'ts_id', col\n        multiindex = history[col].index.to_frame()\n        multiindex['cat'] = multiindex['Id'] + '__' + multiindex['n_epoch'].astype(str)\n        fig.add_trace(\n            go.Scatter(\n                x=multiindex['cat'].values, y=history[col].values,\n                name=y_label, mode='markers',\n                hovertemplate=x_label + ': %{x} <br>' + y_label + ': %{y}<extra></extra>',\n            ), row=i_col+1, col=1\n        )\n        last_ts_in_epoch = multiindex.groupby(level='n_epoch')['cat'].tail(1)\n        last_ts_in_epoch = pd.concat([pd.Series(history.index[0][0] + f\"__{history.index[0][1]}\"), last_ts_in_epoch])\n        loss_by_epoch = history[col].groupby('n_epoch').mean()\n        loss_by_epoch = pd.concat([history[col].head(1), loss_by_epoch])\n        fig.add_trace(\n            go.Scatter(\n                x=last_ts_in_epoch, y=loss_by_epoch.values,\n                name=y_label, mode='lines',\n                hovertemplate=x_label + ': %{x} <br>' + y_label + ': %{y}<extra></extra>',\n            ), row=i_col+1, col=1\n        )\n        fig.update_layout(\n            template='plotly_white', xaxis_title=x_label, title=title\n        )\n    return fig","metadata":{"_uuid":"c4357697-0e0a-4142-9698-964c97a51787","_cell_guid":"9f870526-07f2-49cf-a21b-14f0837ff1ce","collapsed":false,"_kg_hide-input":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.165899Z","iopub.execute_input":"2023-05-12T11:54:10.166306Z","iopub.status.idle":"2023-05-12T11:54:10.182714Z","shell.execute_reply.started":"2023-05-12T11:54:10.166273Z","shell.execute_reply":"2023-05-12T11:54:10.181065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prediction\nPredictions are **slower than training** (note though that during prediction you need stride = 1)","metadata":{"_uuid":"185a10de-27fb-4ea5-8499-4d3a6afcec8c","_cell_guid":"d32ed634-fcd7-4b9f-8742-526096b6e1b7","trusted":true}},{"cell_type":"code","source":"def predict(scaler_fitted, model_fitted, ts_ids:Iterable[str], windows_length_samples:int, stride:int=1):  # stride must be 1 to predict all rows in test dataset\n    y = pd.DataFrame()\n    for ts_id in tqdm(ts_ids, desc='predict'):\n        ts = load_single_timeseries(ts_id)\n        X, _ = get_X_y(ts=ts, scaler_fitted=scaler_fitted, windows_length_samples=windows_length_samples, stride=stride)\n        y_tf = model_fitted.predict(X, batch_size=5000, verbose=0)\n        y_this_ts = pd.DataFrame(\n            index=_.index, columns=target_colnames + ['no_events'], data=y_tf\n        )\n        y_this_ts = y_this_ts.drop(columns='no_events')\n        y_this_ts = y_this_ts.reindex(ts.index)  # this is to maintain all the original samples (even those dropped by windows generation) and their order\n        y = pd.concat([y, y_this_ts])\n    return y","metadata":{"_uuid":"2f4afb0d-0124-4eb0-8853-feeca6a7b150","_cell_guid":"2a33e7eb-987f-43db-a41e-b00cb80eaf99","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:58:15.399811Z","iopub.execute_input":"2023-05-12T12:58:15.400218Z","iopub.status.idle":"2023-05-12T12:58:15.409705Z","shell.execute_reply.started":"2023-05-12T12:58:15.400187Z","shell.execute_reply":"2023-05-12T12:58:15.408127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def format_output_for_competition(y:pd.DataFrame) -> pd.DataFrame:\n    y_out = y.reset_index()\n    y_out['Id'] = y_out['Id'] + '_' + y_out['Time'].astype(str)\n    y_out = y_out.drop(columns='Time')\n    return y_out","metadata":{"execution":{"iopub.status.busy":"2023-05-12T11:54:10.202415Z","iopub.execute_input":"2023-05-12T11:54:10.203148Z","iopub.status.idle":"2023-05-12T11:54:10.216713Z","shell.execute_reply.started":"2023-05-12T11:54:10.203104Z","shell.execute_reply":"2023-05-12T11:54:10.215588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configs","metadata":{"_uuid":"06f3193f-5c66-4639-a7ea-ba63f151815e","_cell_guid":"9b2338c1-81b3-422f-a275-1fe4504fdd3a","trusted":true}},{"cell_type":"code","source":"path_to_data = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction\"\nacc_colnames = ['AccV', 'AccML', 'AccAP']\ntarget_colnames = ['StartHesitation', 'Turn', 'Walking']\ntdcs_sampling = 128  # Hz\nde_sampling = 100  # Hz\nG_MS2 = 9.8066\ntarget_global_fractions = {  # this is true only when series with no positive events have been already dropped (max_noevents_fraction = .99)\n    'tdcs': {'no_events': .526 , 'StartHesitation': .066, 'Turn': .363, 'Walking': .045},\n    'de': {'no_events': .8093 , 'StartHesitation': .0001, 'Turn': .1632, 'Walking': .0274},\n}","metadata":{"_uuid":"758036bc-238e-4681-b3f3-25afd77333a2","_cell_guid":"9d2f44f1-3bc3-445d-9df5-e7b600b71eca","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.218382Z","iopub.execute_input":"2023-05-12T11:54:10.219042Z","iopub.status.idle":"2023-05-12T11:54:10.228521Z","shell.execute_reply.started":"2023-05-12T11:54:10.218999Z","shell.execute_reply":"2023-05-12T11:54:10.227334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" # Choices","metadata":{"_uuid":"e0172624-a6b4-4504-ad3a-52d9421605bf","_cell_guid":"a6830722-d934-47bc-b257-c5848bb65bf2","trusted":true}},{"cell_type":"code","source":"n_epochs_max = 100\nmin_loss_improv = .0005\nwindow_size_seconds = .5\noverlapping_fraction_during_training = .75\nlstm_configs = dict(\n    learning_rate = 1e-2,\n    dropout_rate = 0, \n    n_units = 32\n)\nmax_noevents_fraction = .99  # == .99 otherwise global fractions are incorrect\ndo_balance = False","metadata":{"_uuid":"85115143-cdfd-4fb6-8721-0eec23d17a89","_cell_guid":"1cb65b85-8502-4fee-a34b-c1919510b9ed","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.229955Z","iopub.execute_input":"2023-05-12T11:54:10.230334Z","iopub.status.idle":"2023-05-12T11:54:10.242071Z","shell.execute_reply.started":"2023-05-12T11:54:10.230301Z","shell.execute_reply":"2023-05-12T11:54:10.240995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training","metadata":{"_uuid":"a9e42e1d-7ad7-4a7b-9c5a-98a19dd502df","_cell_guid":"aaf0e2b1-7b38-46c0-b864-ee7dcaea859a","trusted":true}},{"cell_type":"code","source":"train_or_test = 'train'","metadata":{"_uuid":"413b9fd9-fc04-470f-b82c-ede02dbd3d80","_cell_guid":"b485d2c0-ce08-4b69-a8a5-83be68015406","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T11:54:10.243567Z","iopub.execute_input":"2023-05-12T11:54:10.244470Z","iopub.status.idle":"2023-05-12T11:54:10.260141Z","shell.execute_reply.started":"2023-05-12T11:54:10.244432Z","shell.execute_reply":"2023-05-12T11:54:10.258484Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler, model, metrics = {}, {}, {}\nfor de_or_tdcs in ['de', 'tdcs']:\n    sampling = tdcs_sampling if de_or_tdcs == 'tdcs' else de_sampling\n    windows_length_samples = max(1, round(window_size_seconds * sampling))\n    stride = max(1, int(windows_length_samples * (1-overlapping_fraction_during_training)))\n    ts_ids = get_ids_with_events(de_or_tdcs, max_noevents_fraction)\n    scaler[de_or_tdcs], model[de_or_tdcs], metrics[de_or_tdcs] = train(\n        ts_ids, windows_length_samples=windows_length_samples, stride=stride\n    )\n    plot_history(metrics[de_or_tdcs], de_or_tdcs).show()\n    #model.save(f'lstm_{de_or_tdcs}_model')","metadata":{"_uuid":"6297f405-0d7a-4b15-8cd3-6d56e43ca5d1","_cell_guid":"aa063bf2-99be-4852-ac90-989188e2743b","jupyter":{"outputs_hidden":false},"collapsed":false,"execution":{"iopub.status.busy":"2023-05-12T11:54:10.262754Z","iopub.execute_input":"2023-05-12T11:54:10.263304Z","iopub.status.idle":"2023-05-12T12:55:45.062388Z","shell.execute_reply.started":"2023-05-12T11:54:10.263256Z","shell.execute_reply":"2023-05-12T12:55:45.061010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Time-series with high losses are the ones with many events.\nstrategies to avoid this:\n* selecting training time-series based on the number of events\n* weighting samples with an event","metadata":{"_uuid":"d913db5e-2a6d-4e21-9d62-3fc4c91b8559","_cell_guid":"901d64ba-4e33-4c2d-8f11-0f8bf179ad9f","trusted":true}},{"cell_type":"markdown","source":"# Prediction","metadata":{"_uuid":"67fe6ede-0909-4a15-a66d-a9cf101c17f9","_cell_guid":"7f4324f1-eb27-401f-91e0-4ffad2319ced","trusted":true}},{"cell_type":"code","source":"train_or_test = 'test'","metadata":{"_uuid":"cfcc46a8-d9d3-4a2f-b980-66bbfc81333e","_cell_guid":"fe1fb984-48c2-415f-a8f2-96626384e5b0","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:55:45.064329Z","iopub.execute_input":"2023-05-12T12:55:45.064966Z","iopub.status.idle":"2023-05-12T12:55:45.071243Z","shell.execute_reply.started":"2023-05-12T12:55:45.064927Z","shell.execute_reply":"2023-05-12T12:55:45.069883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_pred = {}\nfor de_or_tdcs in ['de', 'tdcs']:\n    sampling = tdcs_sampling if de_or_tdcs == 'tdcs' else de_sampling\n    windows_length_samples = max(1, round(window_size_seconds * sampling))\n    ts_ids = get_ids(train_or_test, de_or_tdcs)\n    y_test_pred[de_or_tdcs] = predict(\n        scaler_fitted=scaler[de_or_tdcs], model_fitted=model[de_or_tdcs], ts_ids=ts_ids, windows_length_samples=windows_length_samples\n    )\n    \ny_test_pred = pd.concat([y_test_pred['de'], y_test_pred['tdcs']])\ny_test_pred","metadata":{"_uuid":"426488df-2f38-4213-b07e-f7715e674fcb","_cell_guid":"d7538fb6-8c56-4170-a0e6-972c3208da3c","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:58:25.343999Z","iopub.execute_input":"2023-05-12T12:58:25.344385Z","iopub.status.idle":"2023-05-12T12:58:45.119334Z","shell.execute_reply.started":"2023-05-12T12:58:25.344356Z","shell.execute_reply":"2023-05-12T12:58:45.118171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_pred = y_test_pred.fillna(0.)  # nans values are from measures where no complete sequence is possible","metadata":{"execution":{"iopub.status.busy":"2023-05-12T12:58:48.835530Z","iopub.execute_input":"2023-05-12T12:58:48.835961Z","iopub.status.idle":"2023-05-12T12:58:48.844717Z","shell.execute_reply.started":"2023-05-12T12:58:48.835927Z","shell.execute_reply":"2023-05-12T12:58:48.843607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = format_output_for_competition(y_test_pred)\nsubmission","metadata":{"_uuid":"996a45b2-09f0-4085-a603-7ba2bc5791e4","_cell_guid":"85d7667b-67b8-4fc1-a053-69301fe6e33b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:58:51.217799Z","iopub.execute_input":"2023-05-12T12:58:51.218216Z","iopub.status.idle":"2023-05-12T12:58:51.556255Z","shell.execute_reply.started":"2023-05-12T12:58:51.218183Z","shell.execute_reply":"2023-05-12T12:58:51.555061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.max()","metadata":{"_uuid":"d4913c39-3a94-420e-aaea-a47d2d0ab71d","_cell_guid":"a61ecddc-8e46-4b39-8191-38b806fa894a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:59:31.412548Z","iopub.execute_input":"2023-05-12T12:59:31.412980Z","iopub.status.idle":"2023-05-12T12:59:31.529211Z","shell.execute_reply.started":"2023-05-12T12:59:31.412938Z","shell.execute_reply":"2023-05-12T12:59:31.528158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the very small max values in predictions suggest the model is not learning and just throwing 0s","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"_uuid":"7edeb634-bfd5-4da2-a6b7-7fcd5dac4d1c","_cell_guid":"48e353b3-af7b-4e0b-96a2-f92b491b756b","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2023-05-12T12:59:36.570976Z","iopub.execute_input":"2023-05-12T12:59:36.571717Z","iopub.status.idle":"2023-05-12T12:59:38.574787Z","shell.execute_reply.started":"2023-05-12T12:59:36.571682Z","shell.execute_reply":"2023-05-12T12:59:38.573935Z"},"trusted":true},"execution_count":null,"outputs":[]}]}