{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://i.imgur.com/8CvOOGQ.png\">\n\n<center><h1> Problem Understanding + </h1></center>\n<center><h1> XGBoost is all you need (: </h1></center>\n\n> 📌 **Competition Scope**: detect freezing of gait (FOG), a debilitating symptom that afflicts many people with Parkinson’s disease, by developing a machine learning model trained on data collected from a wearable 3D lower back sensor.\n\n🧠**Parkinson's Disease**: is a *brain disorder* that causes *unintended or uncontrollable movements* (shaking, stiffness, and difficulty with balance and coordination).\n\n🧠**Freezing Of Gate**: is defined as a *brief, episodic absence* or marked reduction of *forward progression of the feet* despite the intention to walk. It is one of the most debilitating motor symptoms in patients with Parkinson's disease as it may lead to falls and a loss of independence.\n\n🧠**Detecting FOG**: A way to detect FOG is by using *sensors* while doing FOG-provoking testing. Another is by documenting the moment FOG appared within a patient by watching a *video*, frame by frame. While the ML trained on this data has good accuracy, the *datasets are farily small* and the focus has been on recall, while *precission has been ignored*.\n\n### ○ Libraries","metadata":{}},{"cell_type":"code","source":"# General Libraries\nimport os\nimport re\nimport gc\nimport wandb\nimport random\nimport math\nfrom glob import glob\nfrom PIL import Image\nfrom tqdm import tqdm\nfrom pprint import pprint\nfrom time import time\nfrom datetime import datetime\nimport itertools\nimport warnings\nimport pandas as pd\nimport numpy as np\n\n# For the Visuals\nimport seaborn as sns\nimport matplotlib as mpl\nfrom matplotlib import cm\nimport matplotlib.patches as patches\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.ticker import MaxNLocator\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\nfrom matplotlib.patches import Rectangle\nfrom IPython.display import display_html\nplt.rcParams.update({'font.size': 16})\nimport plotly.graph_objects as go\n\nimport cudf\nimport cupy\n\n# Environment check\nwarnings.filterwarnings(\"ignore\", category=DeprecationWarning) \nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'fog_2023', '_wandb_kernel': 'aot'}\n\n# Custom colors\nclass clr:\n    S = '\\033[1m' + '\\033[95m'\n    E = '\\033[0m'\n    \nmy_colors = [\"#761D80\", \"#9926A6\", \"#9C69C9\",\n             \"#6C91BF\", \"#58BCC6\", \"#4AD1B2\",\n             \"#4BF1B2\"]\n\nCMAP1 = ListedColormap(my_colors)\n\nprint(clr.S+\"Notebook Color Schemes:\"+clr.E)\nsns.palplot(sns.color_palette(my_colors))\nplt.show()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-25T17:56:01.023986Z","iopub.execute_input":"2023-04-25T17:56:01.024413Z","iopub.status.idle":"2023-04-25T17:56:06.726708Z","shell.execute_reply.started":"2023-04-25T17:56:01.024356Z","shell.execute_reply":"2023-04-25T17:56:06.724957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# I am doing this at the beginning of every notebook\n# so when I move to my local workstation it is easy to change the paths to the data\n# without iterating through each cell of code\n# cuz it's annoying!\n# ::smiley face::\n# sanyam you there?\nBASE_PATH = \"/kaggle/input/tlvmc-parkinsons-freezing-gait-prediction/\"\nPREP_BASE_PATH = \"/kaggle/input/fog-2023-prepped-competition-data/\"","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:56:06.734533Z","iopub.execute_input":"2023-04-25T17:56:06.738031Z","iopub.status.idle":"2023-04-25T17:56:06.747681Z","shell.execute_reply.started":"2023-04-25T17:56:06.737968Z","shell.execute_reply":"2023-04-25T17:56:06.746388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🐝 W&B Fork & Run\n\nIn order to run this notebook you will need to input your own **secret API key** within the `! wandb login $secret_value_0` line. \n\n🐝**How do you get your own API key?**\n\nSuper simple! Go to **https://wandb.ai/site** -> Login -> Click on your profile in the top right corner -> Settings -> Scroll down to API keys -> copy your very own key (for more info check [this amazing notebook for ML Experiment Tracking on Kaggle](https://www.kaggle.com/ayuraj/experiment-tracking-with-weights-and-biases)).\n\n<center><img src=\"https://i.imgur.com/fFccmoS.png\" width=500></center>","metadata":{}},{"cell_type":"code","source":"# 🐝 Secrets\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"wandb\")\n\n! wandb login $secret_value_0","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:56:06.752734Z","iopub.execute_input":"2023-04-25T17:56:06.754675Z","iopub.status.idle":"2023-04-25T17:56:10.749954Z","shell.execute_reply.started":"2023-04-25T17:56:06.754618Z","shell.execute_reply":"2023-04-25T17:56:10.748617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ○ Helpers","metadata":{}},{"cell_type":"code","source":"# === General Functions ===\ndef jitter(values,j):\n    return values + np.random.normal(j,0.1,values.shape)\n\n\ndef show_values_on_bars(axs, h_v=\"v\", space=0.4):\n    '''Plots the value at the end of the a seaborn barplot.\n    axs: the ax of the plot\n    h_v: weather or not the barplot is vertical/ horizontal'''\n    \n    def _show_on_single_plot(ax):\n        if h_v == \"v\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() / 2\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_height())\n                ax.text(_x, _y, format(value, ','), ha=\"center\") \n        elif h_v == \"h\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() + float(space)\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_width())\n                ax.text(_x, _y, format(value, ','), ha=\"left\")\n\n    if isinstance(axs, np.ndarray):\n        for idx, ax in np.ndenumerate(axs):\n            _show_on_single_plot(ax)\n    else:\n        _show_on_single_plot(axs)\n        \n        \n# === 🐝 W&B ===\ndef save_dataset_artifact(run_name, artifact_name, path, data_type=\"dataset\"):\n    '''Saves dataset to W&B Artifactory.\n    run_name: name of the experiment\n    artifact_name: under what name should the dataset be stored\n    path: path to the dataset'''\n    \n    run = wandb.init(project='fog_2023', \n                     name=run_name, \n                     config=CONFIG)\n    artifact = wandb.Artifact(name=artifact_name, \n                              type=data_type)\n    artifact.add_file(path)\n\n    wandb.log_artifact(artifact)\n    wandb.finish()\n    print(f\"🐝Artifact {artifact_name} has been saved successfully.\")\n    \n    \ndef create_wandb_plot(x_data=None, y_data=None, x_name=None, y_name=None, title=None, log=None, plot=\"line\"):\n    '''Create and save lineplot/barplot in W&B Environment.\n    x_data & y_data: Pandas Series containing x & y data\n    x_name & y_name: strings containing axis names\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[label, val] for (label, val) in zip(x_data, y_data)]\n    table = wandb.Table(data=data, columns = [x_name, y_name])\n    \n    if plot == \"line\":\n        wandb.log({log : wandb.plot.line(table, x_name, y_name, title=title)})\n    elif plot == \"bar\":\n        wandb.log({log : wandb.plot.bar(table, x_name, y_name, title=title)})\n    elif plot == \"scatter\":\n        wandb.log({log : wandb.plot.scatter(table, x_name, y_name, title=title)})\n        \n        \ndef create_wandb_hist(x_data=None, x_name=None, title=None, log=None):\n    '''Create and save histogram in W&B Environment.\n    x_data: Pandas Series containing x values\n    x_name: strings containing axis name\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[x] for x in x_data]\n    table = wandb.Table(data=data, columns=[x_name])\n    wandb.log({log : wandb.plot.histogram(table, x_name, title=title)})","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T17:56:10.755121Z","iopub.execute_input":"2023-04-25T17:56:10.755453Z","iopub.status.idle":"2023-04-25T17:56:10.773776Z","shell.execute_reply.started":"2023-04-25T17:56:10.755419Z","shell.execute_reply":"2023-04-25T17:56:10.772676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Bonus: Cover Photo\nrun = wandb.init(project='fog_2023', name='CoverPhoto', config=CONFIG)\ncover = plt.imread(PREP_BASE_PATH + \"8CvOOGQ.png\")\nwandb.log({\"cover\": wandb.Image(cover)})\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:56:10.775201Z","iopub.execute_input":"2023-04-25T17:56:10.775600Z","iopub.status.idle":"2023-04-25T17:57:00.975687Z","shell.execute_reply.started":"2023-04-25T17:56:10.775547Z","shell.execute_reply":"2023-04-25T17:57:00.974419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Complete Data Understanding\n\n* **The Data Contains**: lower-back 3D accelerometer data\n* **Objective** or **target variable**: detect the *start* and *stop* of each freezing episode and the *occurrence* in these series of three types of FOG events: \n    1. Start Hesitation\n    2. Turn\n    3. Walking\n    \n## Data Description\n\nThere are **3 types of data** (collected in 3 distinct circumstances):\n* `tdcsfog` -> data collected in lab while subjects completed FOG provoking protocol.\n* `defog` -> data collected in subject's home while completing FOG provoking protocol\n* `daily` -> data collected in continuous 24/7 recordings from 65 subjects\n    * 45/65 subjects exhibit FOG symptoms (also shown in `defog` data)\n    * 20/65 subjects didn't exhibit FOG symptoms","metadata":{}},{"cell_type":"code","source":"def read_data(meta_path, data_type=\"tdcsfog\", tests=True, verbose=True):\n    \n    # 🐝 New Experiment\n    run = wandb.init(project='fog_2023', name=f'{data_type}_basics', config=CONFIG)\n    \n    # Read in data\n    df_meta = pd.read_csv(meta_path)\n\n    if verbose:\n        print(clr.S+\"Data Shape:\"+clr.E, len(df_meta))\n        print(clr.S+\"Data Cols:\"+clr.E, df_meta.columns.tolist())\n        print(clr.S+\"No. Missing Values:\"+clr.E, df_meta.isna().sum().sum(), \"\\n\")\n        print(clr.S+\"No. of unique subjects:\"+clr.E, df_meta[\"Subject\"].nunique(), \"\\n\")\n\n        print(clr.S+\"--- Visits ---\"+clr.E)\n        print(clr.S+\"Max no. of visits:\"+clr.E, df_meta[\"Visit\"].max())\n        print(clr.S+\"Avg no. of visits:\"+clr.E, round(df_meta[\"Visit\"].mean()))\n        print(clr.S+\"Min no. of visits:\"+clr.E, df_meta[\"Visit\"].min(), \"\\n\")\n\n        if tests:\n            print(clr.S+\"--- Tests ---\"+clr.E)\n            print(clr.S+\"Tests frequencies:\"+clr.E)\n            print(df_meta[\"Test\"].value_counts(), \"\\n\")\n\n        print(clr.S+\"--- Medication Flag ---\"+clr.E)\n        print(clr.S+\"Flag frequencies:\"+clr.E)\n        print(df_meta[\"Medication\"].value_counts(), \"\\n\")\n    \n    wandb.log(\n        {\"data_shape\": len(df_meta),\n         \"max_visits\": df_meta[\"Visit\"].max(),\n         \"avg_visits\": round(df_meta[\"Visit\"].mean()),\n         \"min_visits\": df_meta[\"Visit\"].min(),\n         \"missing_values\": df_meta.isna().sum().sum()\n        }\n    )\n\n    # Get csv paths\n    base_path = BASE_PATH + \"train\"\n    df_meta[\"path\"] = df_meta[\"Id\"].apply(lambda x: f\"{base_path}/{data_type}/{x}.csv\")\n    print(clr.S+\"--- Created path to .csv info ---\"+clr.E)\n    \n    # 🐝 Experiment end\n    wandb.finish()\n    print(\"🐝 Info saved to dashboard.\")\n\n    return df_meta","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-04-25T17:57:00.978829Z","iopub.execute_input":"2023-04-25T17:57:00.979532Z","iopub.status.idle":"2023-04-25T17:57:00.992805Z","shell.execute_reply.started":"2023-04-25T17:57:00.979486Z","shell.execute_reply":"2023-04-25T17:57:00.991758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.1 TDCS FOG (lab data)\n\n* `tdcsfog_metadata.csv`\n    * `Id` -> unique ID of the session?\n    * `Subject` -> Unique ID of the patient\n    * `Visit` -> Number of the visit in the lab (1 baseline, 2 post treatment, 1 follow up)\n    * `Test` -> which out of 3 tests was performed (*3 was the most challenging*)\n    * `Medication` -> flag to indicate if the patient was *on* or *off* anti-parkinson medication","metadata":{}},{"cell_type":"code","source":"path = BASE_PATH + \"tdcsfog_metadata.csv\"\ntdcsfog_meta = read_data(meta_path=path,\n                         data_type=\"tdcsfog\",\n                         tests=True)\n\ntdcsfog_meta.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:57:00.995594Z","iopub.execute_input":"2023-04-25T17:57:00.996068Z","iopub.status.idle":"2023-04-25T17:57:43.109430Z","shell.execute_reply.started":"2023-04-25T17:57:00.996027Z","shell.execute_reply":"2023-04-25T17:57:43.108243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the labels for the nodes and links\nlabels = ['OFF', 'ON', 'Test 1', 'Test 2', 'Test 3']\n\n# Define source, target\nsource = [0, 0, 0, 1, 1, 1]\ntarget = [2, 4, 3, 3, 2, 4]\n# Define values for each link\nvalue = tdcsfog_meta.groupby(\"Medication\")[\"Test\"].value_counts().values.tolist()\n\n# Node colors\nnode_colors = [my_colors[0], my_colors[2], \n               my_colors[4], my_colors[3], my_colors[5]]\n\n# Link colors\nlink_colors = ['rgba(118, 29, 128, 0.4)', 'rgba(118, 29, 128, 0.4)', 'rgba(118, 29, 128, 0.4)', \n               'rgba(156, 105, 201, 0.4)', 'rgba(156, 105, 201, 0.4)', 'rgba(156, 105, 201, 0.4)']\n\n\n# Create the Sankey diagram figure\nfig = go.Figure(data=[go.Sankey(\n    node=dict(\n        pad=15,\n        thickness=20,\n        line=dict(color=\"white\", width=0.5),\n        label=labels,\n        color=node_colors\n    ),\n    link=dict(\n        source=source,\n        target=target,\n        value=value,\n        color = link_colors\n    ))])\n\n# Create annotation\nannotations = [\n    go.Annotation(\n        x=-0.1,\n        y=1.12,\n        showarrow=False,\n        text='Whether the patient is<br> on anti-parkinson medication',\n        font=dict(size=16, color='black', family='Arial, sans-serif'),\n),\n    go.Annotation(\n        x=1.05,\n        y=1.12,\n        showarrow=False,\n        text='Type of test<br> (3 is the hardest)',\n        font=dict(size=16, color='black', family='Arial, sans-serif'),\n),\n    go.Annotation(\n        x=-0.115,\n        y=0.53,\n        showarrow=False,\n        text='the division<br> between<br> medication<br> and tests<br> looks balanced',\n        font=dict(size=13, color='#545454', family='Arial, sans-serif'),\n),\n    go.Annotation(\n            x=1.11,\n            y=0.53,\n            showarrow=False,\n            text='the test count<br> is distributed<br> equally',\n            font=dict(size=13, color='#545454', family='Arial, sans-serif'),\n)]\n\n# Set the title and layout of the figure\nfig.update_layout(\n    title_text=\"[tdcsfog] Medication | Test Percentages\",\n    font=dict(size=16, color='black', family='Arial, sans-serif'),\n    annotations=annotations\n)\n\n# Display the figure\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T17:57:43.111942Z","iopub.execute_input":"2023-04-25T17:57:43.112336Z","iopub.status.idle":"2023-04-25T17:57:43.585679Z","shell.execute_reply.started":"2023-04-25T17:57:43.112298Z","shell.execute_reply":"2023-04-25T17:57:43.584434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**📌 Note**:\n* The tests are equally distributed between the visits made by the patient\n* There are more tests made in the first sessions than in the latter sessions","metadata":{}},{"cell_type":"code","source":"dt = tdcsfog_meta.groupby([\"Test\", \"Visit\"])[\"Medication\"].count().reset_index()\n\nnp.random.seed(41)\n\nplt.figure(figsize=(24, 13))\nax = sns.scatterplot(x=jitter(tdcsfog_meta[\"Test\"], 0.05), y=jitter(tdcsfog_meta[\"Visit\"], 0.05), \n                size=tdcsfog_meta[\"Visit\"], alpha=0.7, sizes=(500, 500),\n                hue=tdcsfog_meta[\"Visit\"],\n                palette = my_colors,\n                edgecolor=\"black\", linewidth=1)\n\nplt.title(\"Test Type x Visit Number: Correlation\", weight=\"bold\", size=25)\nplt.ylabel(\"Visit\", size = 20, weight=\"bold\")\nplt.xlabel(\"Test\", size = 20, weight=\"bold\")\n\nax.yaxis.set_major_locator(MaxNLocator(integer=True))\nax.xaxis.set_major_locator(MaxNLocator(integer=True))\n\ntext_cols = my_colors*3\nfor k in range(len(dt)):\n    row = dt.iloc[k, :]\n    ax.text(x=row[\"Test\"]+0.36, y=row[\"Visit\"], \n            s=f\"Visit {row['Visit']}: {row['Medication']} tests\", \n            color=text_cols[k], weight=\"bold\",\n            size=18)\n\n# --- FIN ---\nsns.despine(right=True, top=True, left=False)\nplt.legend([]);","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T17:57:43.587176Z","iopub.execute_input":"2023-04-25T17:57:43.587651Z","iopub.status.idle":"2023-04-25T17:57:44.362875Z","shell.execute_reply.started":"2023-04-25T17:57:43.587611Z","shell.execute_reply":"2023-04-25T17:57:44.361703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### TDCS FOG (lab data) - metadata for each session\n\n**Each session has a train metadata stored within a `.csv` file containing:**\n* `Time` -> recorded at 128Hz (128 timestamps per second)\n* `Acc` columns -> acceleration from a lower back senzor in units of m/s^2\n    * `AccV` -> vertical ax\n    * `AccML` -> mediolateral ax\n    * `AccAP` -> anteroposterior ax \n* *[target variables]* `StartHesitation`, `Turn`, `Walking` -> flags the occurrence of each of the event type\n    \n> 📌 **Note**: the vertical axis runs from head to toe, the mediolateral axis runs from side to side, and the anteroposterior axis runs from front to back.","metadata":{}},{"cell_type":"code","source":"pd.read_csv(tdcsfog_meta[\"path\"][0]).head()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:57:44.366900Z","iopub.execute_input":"2023-04-25T17:57:44.367852Z","iopub.status.idle":"2023-04-25T17:57:44.396743Z","shell.execute_reply.started":"2023-04-25T17:57:44.367814Z","shell.execute_reply":"2023-04-25T17:57:44.395707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_acc_lineplot(path, verbose=True):\n    ex = pd.read_csv(path)\n    ex = ex.melt(id_vars=[\"Time\", \"StartHesitation\", \"Turn\", \"Walking\"],\n                 value_vars=[\"AccV\", \"AccML\", \"AccAP\"],\n                 var_name=\"Acc\")\n    experiment_name = path.split(\"/\")[-1].split(\".\")[0]\n\n    # Plot\n    plt.figure(figsize=(24, 12))\n    plt.title(f\"[{experiment_name}] Acceleration senzors data\", weight=\"bold\", size=25)\n    sns.lineplot(data=ex, x=\"Time\", y=\"value\", hue=\"Acc\", \n                 palette=[my_colors[1], my_colors[3], my_colors[5]],\n                 lw=2)\n    plt.xlabel(\"Time\", size = 18, weight=\"bold\")\n    plt.ylabel(\"Acc Values\", size = 18, weight=\"bold\")\n    sns.despine(right=True, top=True, left=True)\n    \n    # Add turn, hesitation, walking\n    targets = [\"StartHesitation\", \"Turn\", \"Walking\"]\n    target_cols = [\"#9C829F\", \"#8797AC\", \"#7DA49A\"]\n\n    for k, target in enumerate(targets):\n\n        arr = ex.loc[ex[target]==1][\"Time\"]\n\n        if arr.sum() > 0:\n            # Find the indices where the difference between adjacent elements is not 1\n            idx = np.where(np.diff(arr) != 1)[0] + 1\n\n            # Split the array into sub-arrays based on the indices\n            sub_arrays = np.split(arr, idx)\n\n            # Select the min and max values for each sub-array\n            min_max_vals = list(set([(sub_arr.min(), sub_arr.max()) for sub_arr in sub_arrays]))\n\n            for start, end in min_max_vals:\n                # Add vertical lines to mark the window\n                plt.axvspan(start, end, facecolor=target_cols[k], alpha=0.4)\n                if verbose:\n                    # Write text\n                    q1 = ex.loc[ex[\"Acc\"]==\"AccV\"][\"value\"].quantile(0.001)\n                    plt.text((start + end)/2, q1, target, ha='center', va='center', \n                             fontsize=18, weight=\"bold\")\n\n    plt.tight_layout()\n    plt.show();","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T17:57:44.398333Z","iopub.execute_input":"2023-04-25T17:57:44.398764Z","iopub.status.idle":"2023-04-25T17:57:44.413586Z","shell.execute_reply.started":"2023-04-25T17:57:44.398726Z","shell.execute_reply":"2023-04-25T17:57:44.410425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for p in tdcsfog_meta[\"path\"].sample(5, random_state=25).values:\n    get_acc_lineplot(path=p)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:57:44.415090Z","iopub.execute_input":"2023-04-25T17:57:44.415495Z","iopub.status.idle":"2023-04-25T17:57:48.230428Z","shell.execute_reply.started":"2023-04-25T17:57:44.415459Z","shell.execute_reply":"2023-04-25T17:57:48.229284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.2 DE FOG (home data)\n\n* `defog_metadata.csv`\n    * `Id` -> unique ID of the test?\n    * `Subject` -> Unique ID of the patient\n    * `Visit` -> Number of the visit in the lab (1 baseline, 2 post treatment, 1 follow up)\n    * `Medication` -> flag to indicate if the patient was *on* or *off* anti-parkinson medication","metadata":{}},{"cell_type":"code","source":"path = BASE_PATH + \"defog_metadata.csv\"\ndefog_meta = read_data(meta_path=path,\n                       data_type=\"defog\",\n                       tests=False)\n\ndefog_meta.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:57:48.232002Z","iopub.execute_input":"2023-04-25T17:57:48.232703Z","iopub.status.idle":"2023-04-25T17:58:30.216405Z","shell.execute_reply.started":"2023-04-25T17:57:48.232663Z","shell.execute_reply":"2023-04-25T17:58:30.215324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define the labels for the nodes and links\nlabels = ['OFF', 'ON', 'Visit 1', 'Visit 2']\n\n# Define source, target\nsource = [0, 0, 1, 1]\ntarget = [2, 3, 2, 3]\n# Define values for each link\nvalue = defog_meta.groupby(\"Medication\")[\"Visit\"].value_counts().values.tolist()\n\n# Node colors\nnode_colors = [my_colors[0], my_colors[2], \n               my_colors[4], my_colors[3]]\n\n# Link colors\nlink_colors = ['rgba(118, 29, 128, 0.4)', 'rgba(118, 29, 128, 0.4)', \n               'rgba(156, 105, 201, 0.4)', 'rgba(156, 105, 201, 0.4)']\n\n\n# Create the Sankey diagram figure\nfig = go.Figure(data=[go.Sankey(\n    node=dict(\n        pad=15,\n        thickness=20,\n        line=dict(color=\"white\", width=0.5),\n        label=labels,\n        color=node_colors\n    ),\n    link=dict(\n        source=source,\n        target=target,\n        value=value,\n        color = link_colors\n    ))])\n\n# Create annotation\nannotations = [\n    go.Annotation(\n        x=-0.1,\n        y=1.12,\n        showarrow=False,\n        text='Whether the patient is<br> on anti-parkinson medication',\n        font=dict(size=16, color='black', family='Arial, sans-serif'),\n),\n    go.Annotation(\n        x=1.05,\n        y=1.12,\n        showarrow=False,\n        text='Visit Number<br> (only 2 visits for this dataset)',\n        font=dict(size=16, color='black', family='Arial, sans-serif'),\n),\n    go.Annotation(\n        x=-0.115,\n        y=0.53,\n        showarrow=False,\n        text='the division<br> between<br> medication<br> and visits<br> looks balanced',\n        font=dict(size=13, color='#545454', family='Arial, sans-serif'),\n)]\n\n# Set the title and layout of the figure\nfig.update_layout(\n    title_text=\"[defog] Medication | Visit Percentages\",\n    font=dict(size=16, color='black', family='Arial, sans-serif'),\n    annotations=annotations\n)\n\n# Display the figure\nfig.show()","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T17:58:30.218116Z","iopub.execute_input":"2023-04-25T17:58:30.218513Z","iopub.status.idle":"2023-04-25T17:58:30.243360Z","shell.execute_reply.started":"2023-04-25T17:58:30.218471Z","shell.execute_reply":"2023-04-25T17:58:30.242310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### DE FOG (home data) - metadata for each session\n\n> 📌 **Note**: Out of 137 rows, *only 91* observations also have metadata available.","metadata":{}},{"cell_type":"code","source":"defog_paths = glob(BASE_PATH + \"train/defog/*\")\nprint(clr.S+\"Number of available metadata files:\"+clr.E, len(defog_paths), \"out of 137\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:30.244910Z","iopub.execute_input":"2023-04-25T17:58:30.245723Z","iopub.status.idle":"2023-04-25T17:58:30.271771Z","shell.execute_reply.started":"2023-04-25T17:58:30.245686Z","shell.execute_reply":"2023-04-25T17:58:30.270363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Replace all paths where there is no metadata with NaN\nmetadata_ids = [x.split(\"/\")[-1].split(\".\")[0] for x in defog_paths]\n\nfor current_id in defog_meta[\"Id\"].unique():\n    # Create mask and filter only cases where ID is in path\n    if current_id not in metadata_ids:\n        defog_meta.loc[defog_meta[\"Id\"]==current_id, \"path\"] = np.nan","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:30.273323Z","iopub.execute_input":"2023-04-25T17:58:30.273774Z","iopub.status.idle":"2023-04-25T17:58:30.305379Z","shell.execute_reply.started":"2023-04-25T17:58:30.273732Z","shell.execute_reply":"2023-04-25T17:58:30.304465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Each session has a train metadata stored within a `.csv` file containing:**\n* `Time` -> recorded at 100Hz (100 timestamps per second)\n* `Acc` columns -> acceleration from a lower back senzor in units of m/s^2\n    * `AccV` -> vertical ax\n    * `AccML` -> mediolateral ax\n    * `AccAP` -> anteroposterior ax \n* *[target variables]* `StartHesitation`, `Turn`, `Walking` -> flags the occurrence of each of the event type\n* `Valid` -> due to a possible human error (some videos were hard to annotate, as the annotators coudlnt not tell exactly if the subject stopped voluntarily or not) this column was introduced. Only where the `flag=True` we know for sure that the labeling is accurate.\n* `Task` -> `True` is the series were annotated - `False` for portions that were not annotated.","metadata":{}},{"cell_type":"code","source":"pd.read_csv(defog_meta[\"path\"][1]).head()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:30.306941Z","iopub.execute_input":"2023-04-25T17:58:30.307297Z","iopub.status.idle":"2023-04-25T17:58:30.540341Z","shell.execute_reply.started":"2023-04-25T17:58:30.307262Z","shell.execute_reply":"2023-04-25T17:58:30.539212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**📌 Note**:\n* compared to the lab data, the *home data* is much more segmented between the 3 targets (walking, turn, start hesitation) and the episodes are **much shorter**.\n* this might be because in the lab the tests were controlled vs. at home activity.","metadata":{}},{"cell_type":"code","source":"for p in defog_meta.dropna()[\"path\"].sample(5, random_state=25).values:\n    print()\n    get_acc_lineplot(path=p, verbose=False)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:30.541822Z","iopub.execute_input":"2023-04-25T17:58:30.542836Z","iopub.status.idle":"2023-04-25T17:58:42.233909Z","shell.execute_reply.started":"2023-04-25T17:58:30.542793Z","shell.execute_reply":"2023-04-25T17:58:42.232398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Create the full metadata dataset\n\n🧠 **Steps**:\n1. Created a small function that ingests all `.csv` files within our 3 folders (\"defog\", \"notype\", \"tdcsfog\") and concatenates them\n2. Saved the full metadata in `.parquet` format\n3. Created a [dataset with the metadata files here](https://www.kaggle.com/datasets/andradaolteanu/fog-2023-prepped-competition-data)\n4. Saved to 🐝W&B as artifacts for versioning purposes","metadata":{}},{"cell_type":"code","source":"def easy_read(path):\n    '''\n    Read the data and append the Id.\n    '''\n    df = pd.read_csv(path)\n    df[\"Id\"] = path.split(\"/\")[-1].split(\".\")[0]\n    \n    return df\n\ndef create_full_metadata(folder_name):\n    '''\n    Read in all metadata within a folder and saves it.\n    '''\n    paths = glob(BASE_PATH + f\"train/{folder_name}/*\")\n    final_df = pd.concat([easy_read(p) for p in paths],\n                         axis=0)\n    final_df[\"data_type\"] = folder_name\n    final_df.to_parquet(f\"{folder_name}_meta.parquet\", index=False)\n    \n    print(clr.S+f\"{folder_name} metadata saved successfully.\"+clr.E)\n    gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:42.235764Z","iopub.execute_input":"2023-04-25T17:58:42.236444Z","iopub.status.idle":"2023-04-25T17:58:42.245111Z","shell.execute_reply.started":"2023-04-25T17:58:42.236407Z","shell.execute_reply":"2023-04-25T17:58:42.243775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Takes a few seconds to run\nfolder_names = [\"defog\", \"notype\", \"tdcsfog\"]\n\nfor name in folder_names:\n    create_full_metadata(folder_name=name)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:58:42.246590Z","iopub.execute_input":"2023-04-25T17:58:42.247007Z","iopub.status.idle":"2023-04-25T17:59:49.828933Z","shell.execute_reply.started":"2023-04-25T17:58:42.246971Z","shell.execute_reply":"2023-04-25T17:59:49.827787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Save artifacts\nfor name in [\"defog\", \"notype\", \"tdcsfog\"]:\n    save_dataset_artifact(run_name=f\"save_{name}\",\n                          artifact_name=name, \n                          path=PREP_BASE_PATH + f\"{name}_meta.parquet\",\n                          data_type=\"dataset\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T17:59:49.830495Z","iopub.execute_input":"2023-04-25T17:59:49.831206Z","iopub.status.idle":"2023-04-25T18:01:58.656158Z","shell.execute_reply.started":"2023-04-25T17:59:49.831160Z","shell.execute_reply":"2023-04-25T18:01:58.655043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Concatenate & create new features\n\n## 3.1 Creating our base training data\n\n🧠 **The functions contain the following steps:**\n1. for each type of data we read in the concatenated .csv information + the metadata within the dataframes\n2. we then concatenate these 2 types of data into 1 to obtain our `full_train` base dataframe\n3. then create the `target` variable:\n    * for each event type (start hesitation, turn and walking) we use the flag to mark if the event occured\n    * if no event was present we mark it with \"Normal\" label","metadata":{}},{"cell_type":"code","source":"# 🐝 New experiment\nrun = wandb.init(project='fog_2023', name='train_preprocessing', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:01:58.657977Z","iopub.execute_input":"2023-04-25T18:01:58.658399Z","iopub.status.idle":"2023-04-25T18:02:32.586554Z","shell.execute_reply.started":"2023-04-25T18:01:58.658338Z","shell.execute_reply":"2023-04-25T18:02:32.585398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_and_concat_data(data_type):\n    '''\n    Reads folder and df metadata and concatenates into 1 dataframe.\n    '''\n    \n    # Read in .csv folder metadata\n    path = PREP_BASE_PATH + f\"{data_type}_meta.parquet\"\n    meta = cudf.read_parquet(path)\n\n    if data_type == \"defog\":\n        meta = meta[meta['Valid']==1].reset_index(drop=True)\n\n    # Read in df_metadata\n    df_meta = cudf.read_csv(BASE_PATH + f\"{data_type}_metadata.csv\")\n\n    # Concatenate\n    df_meta = df_meta.merge(meta, how='inner', on=\"Id\")\n    if data_type == \"defog\":\n        df_meta.drop(columns=[\"Valid\", \"Task\"], axis=1, inplace=True)\n\n    print(clr.S+f\"{data_type} Shape:\"+clr.E, df_meta.shape)\n    gc.collect()\n    \n    return df_meta\n\n\ndef get_train_data():\n    '''\n    Reads in defog and tdcsfog data, merges them and creates the final target column.\n    '''\n    \n    tdcsfog_train = read_and_concat_data(data_type=\"tdcsfog\")\n\n    defog_train = read_and_concat_data(data_type=\"defog\")\n    # Create Test data & reorder the same as for tdcsfog\n    defog_train[\"Test\"] = 0\n    defog_train = defog_train.reindex(columns=tdcsfog_train.columns)\n\n    # Create full training data\n    train_full = cudf.concat([tdcsfog_train, defog_train])\n    print(clr.S+f\"--- final Shape ---\"+clr.E, train_full.shape)\n    \n    # Create multiclass event type column\n    train_full[\"target\"] = \"Normal\"\n    targets = [\"StartHesitation\", \"Turn\", \"Walking\"]\n    for target in targets:\n        train_full.loc[train_full[target] == 1, \"target\"] = target\n        \n    # Keep only columns that we need\n    train_full = train_full[[\"Time\", \"AccV\", \"AccML\", \"AccAP\", \"data_type\", \"target\"]]\n    \n    gc.collect()\n    return train_full","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:32.591826Z","iopub.execute_input":"2023-04-25T18:02:32.594257Z","iopub.status.idle":"2023-04-25T18:02:34.279029Z","shell.execute_reply.started":"2023-04-25T18:02:32.594216Z","shell.execute_reply":"2023-04-25T18:02:34.278021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get train data\ntrain_full = get_train_data()\ntrain_full.sample(5, random_state=25)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:34.284185Z","iopub.execute_input":"2023-04-25T18:02:34.286827Z","iopub.status.idle":"2023-04-25T18:02:42.322331Z","shell.execute_reply.started":"2023-04-25T18:02:34.286779Z","shell.execute_reply":"2023-04-25T18:02:42.321097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> **📌 Note:** I wanted to see (now that we have all data together) what is the frequency for our target label.\n* `turn` is the most frequent\n* `start hesitation` and `walking` have much lower frequency","metadata":{}},{"cell_type":"code","source":"# Get the data\ndt = train_full[\"target\"].value_counts().reset_index().to_pandas()\ndt.columns = [\"target\", \"frequency\"]\n\n# Plot\nplt.figure(figsize=(20, 10))\n\nfigure = sns.barplot(data=dt,\n                     x=\"target\", y=\"frequency\", palette=my_colors[1:])\nshow_values_on_bars(figure, h_v=\"v\", space=0.4)\nplt.title('[train] Target Variable distribution', weight=\"bold\", size=20)\n\nplt.xlabel(\"Target Value\", size = 18, weight=\"bold\")\nplt.ylabel(\"Count\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2023-04-25T18:02:42.324429Z","iopub.execute_input":"2023-04-25T18:02:42.325144Z","iopub.status.idle":"2023-04-25T18:02:44.426027Z","shell.execute_reply.started":"2023-04-25T18:02:42.325104Z","shell.execute_reply":"2023-04-25T18:02:44.424821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Log in plot\ncreate_wandb_plot(x_data=dt[\"target\"],\n                  y_data=dt[\"frequency\"],\n                  x_name=\"Target Name\", y_name=\"Count\",\n                  title=\"[train] Target Variable distribution\",\n                  log=\"bar_target\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:44.431835Z","iopub.execute_input":"2023-04-25T18:02:44.434673Z","iopub.status.idle":"2023-04-25T18:02:46.044802Z","shell.execute_reply.started":"2023-04-25T18:02:44.434630Z","shell.execute_reply":"2023-04-25T18:02:46.043606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.2 More Features & Time variable\n\n*Work in Progress*","metadata":{}},{"cell_type":"code","source":"# ...","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:46.050609Z","iopub.execute_input":"2023-04-25T18:02:46.050890Z","iopub.status.idle":"2023-04-25T18:02:47.786316Z","shell.execute_reply.started":"2023-04-25T18:02:46.050862Z","shell.execute_reply":"2023-04-25T18:02:47.785273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Experiment End\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:47.795698Z","iopub.execute_input":"2023-04-25T18:02:47.796057Z","iopub.status.idle":"2023-04-25T18:02:56.624020Z","shell.execute_reply.started":"2023-04-25T18:02:47.796019Z","shell.execute_reply":"2023-04-25T18:02:56.622800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. XGBoost\n\n> 🧠 **RAPIDS**: We'll use the `cuml` library from `rapids` because it's faster and enables GPU usage (:\n\n### ○ Libraries","metadata":{}},{"cell_type":"code","source":"# Cuml libraries\nfrom cuml.preprocessing import LabelEncoder\nfrom cuml.model_selection import train_test_split\nimport xgboost as xgb","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:56.625582Z","iopub.execute_input":"2023-04-25T18:02:56.625962Z","iopub.status.idle":"2023-04-25T18:02:57.824471Z","shell.execute_reply.started":"2023-04-25T18:02:56.625921Z","shell.execute_reply":"2023-04-25T18:02:57.823300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.1 Data Validation\n\n🧠 **Steps**:\n1. Encode the `target` variable\n2. Split the data (for now with the usual `train_test_split()` function)","metadata":{}},{"cell_type":"code","source":"def encode_target(df):\n    '''\n    Encode the target variable.\n    return:: encoded target, encoder\n    '''\n    # Create target encoding\n    target_encoder = LabelEncoder()\n    df[\"target\"] = target_encoder.fit_transform(df[\"target\"])\n    \n    return df[\"target\"], target_encoder\n\n\ndef get_simple_data_validation(train_df):\n    '''\n    Encode & split the data in a simple manner.\n    return:: training, validation and encoder objects\n    '''\n    # Encode the target\n    train_df[\"target\"], target_encoder = encode_target(df=train_df)\n    \n    # Split the data\n    X_train, X_valid, y_train, y_valid = train_test_split(train_df, \"target\",\n                                                          train_size=0.8, \n                                                          random_state=42)\n    \n    # Convert\n    dtrain = cudf.DataFrame(X_train)\n    dtrain['target'] = y_train\n    dvalid = cudf.DataFrame(X_valid)\n    dvalid['target'] = y_valid\n    \n    return dtrain, dvalid, target_encoder","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:57.825759Z","iopub.execute_input":"2023-04-25T18:02:57.826735Z","iopub.status.idle":"2023-04-25T18:02:57.833990Z","shell.execute_reply.started":"2023-04-25T18:02:57.826703Z","shell.execute_reply":"2023-04-25T18:02:57.832794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.2 Training Function\n\n🧠 **Steps**:\n* Create the training and validation data matrixes (this is the input XGBoost expects)\n* Train using the desired hyperparameters\n* Evaluate the model & compute the accuracy","metadata":{}},{"cell_type":"code","source":"def train_xgboost(dtrain, dvalid, config):\n    '''\n    Train the XGBoost model.\n    '''    \n    feature_cols = [\"AccV\", \"AccML\", \"AccAP\"]\n    params = {\n        'objective': config.objective,\n        'eval_metric': config.eval_metric,\n        'num_class': config.num_class,\n        'tree_method': config.tree_method,\n        \"random_state\": config.random_state,\n        \"learning_rate\": config.learning_rate,\n        \"max_depth\": config.max_depth,\n        \"min_child_weight\": config.min_child_weight,\n    }\n\n    # Matrix\n    dtrain_matrix = xgb.DMatrix(dtrain[feature_cols], label=dtrain[\"target\"])\n    dvalid_matrix = xgb.DMatrix(dvalid[feature_cols], label=dvalid[\"target\"])\n    \n    # Training ...\n    model = xgb.train(params, dtrain_matrix, \n                      evals=[(dvalid_matrix, 'test')], \n                      num_boost_round=100,\n                      verbose_eval=False)\n    \n    # Evaluate ...\n    y_pred = model.predict(dvalid_matrix)\n    y_pred = cupy.asarray(y_pred)\n\n    # Compute accuracy\n    y_true = cupy.asarray(dvalid[\"target\"])\n    accuracy = cupy.sum(y_pred == y_true) / len(y_pred)\n    wandb.log({\"accuracy\": np.float(accuracy)})\n    \n    print(clr.S+f\"Accuracy: {accuracy:.4f}\"+clr.E)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:57.835535Z","iopub.execute_input":"2023-04-25T18:02:57.836165Z","iopub.status.idle":"2023-04-25T18:02:57.848527Z","shell.execute_reply.started":"2023-04-25T18:02:57.836127Z","shell.execute_reply":"2023-04-25T18:02:57.847344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.3 Training pipeline ...","metadata":{}},{"cell_type":"code","source":"def train_pipeline():\n    \n    dtrain, dvalid, _ = get_simple_data_validation(train_df=train_full.sample(100000,\n                                                                              random_state=55))\n    \n    # XGBoost hyperparameters\n    config_defaults = {\n        'objective': 'multi:softmax',\n        'eval_metric': 'mlogloss',\n        'num_class': 4,\n        'tree_method': 'gpu_hist',\n        \"random_state\": 24,\n        \"learning_rate\": 0.1,\n        \"max_depth\": 1,\n        \"min_child_weight\": 1,\n    }\n    \n    # 🐝 W&B Experiment\n    config_defaults.update(CONFIG)\n    run = wandb.init(project='fog_2023', config=config_defaults)\n    config = wandb.config\n    \n    train_xgboost(dtrain, dvalid, config)\n    \n    # 🐝\n    wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:57.850298Z","iopub.execute_input":"2023-04-25T18:02:57.850746Z","iopub.status.idle":"2023-04-25T18:02:57.862828Z","shell.execute_reply.started":"2023-04-25T18:02:57.850705Z","shell.execute_reply":"2023-04-25T18:02:57.861847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ⦿ Baseline\n\nFirst model (with random hyperparameters):","metadata":{}},{"cell_type":"code","source":"train_pipeline()","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:02:57.864341Z","iopub.execute_input":"2023-04-25T18:02:57.864776Z","iopub.status.idle":"2023-04-25T18:03:43.660160Z","shell.execute_reply.started":"2023-04-25T18:02:57.864741Z","shell.execute_reply":"2023-04-25T18:03:43.659023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ⦿ Sweeps\n\n> 📌 **Note**: For the hyperparameter tuning part I will be using the [W&B integrated Sweeps for XGBoost](https://docs.wandb.ai/guides/integrations/xgboost) method to log all my experiments. \n\n*This is in my opinion one of their best features - I always used to \"lose\" my experiments or would not be able to reproduce them - now I have all the comparisons neatly & automatically put in my dashboard. :)*\n\n> 🧹 **Example Sweep [vfqfnnoh](https://wandb.ai/andrada/fog_2023/sweeps/vfqfnnoh?workspace=user-andrada)**:\n<center><img src=\"https://i.imgur.com/oSqobVs.png\"></center>","metadata":{}},{"cell_type":"code","source":"# Sweep Config\nsweep_config = {\n    \"method\": \"random\", # grid for all\n    \"metric\": {\n      \"name\": \"accuracy\",\n      \"goal\": \"maximize\"   \n    },\n    \"parameters\": {\n        \"max_depth\": {\n            \"values\": [1, 4, 6, 10, 15, 20]\n        },\n        \"min_child_weight\": {\n            \"values\": [1, 2, 3, 4, 5, 8, 10]\n        },\n        \"learning_rate\": {\n            \"values\": [0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.5, 0.7]\n        },\n        \"random_state\": {\n            \"values\": [10, 24, 30, 45, 50, 75, 80, 100]\n        }\n    }\n}\n\n# Sweep ID\nsweep_id = wandb.sweep(sweep_config, project=\"fog_2023\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:03:43.662083Z","iopub.execute_input":"2023-04-25T18:03:43.662975Z","iopub.status.idle":"2023-04-25T18:03:45.043436Z","shell.execute_reply.started":"2023-04-25T18:03:43.662912Z","shell.execute_reply":"2023-04-25T18:03:45.042134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 RUN SWEEPS\nstart = time()\n\n# count = the number of trials/experiments to run\nwandb.agent(sweep_id, train_pipeline, count=15)\nprint(\"Sweeping took:\", round((time()-start)/60, 1), \"mins\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:03:45.045420Z","iopub.execute_input":"2023-04-25T18:03:45.046170Z","iopub.status.idle":"2023-04-25T18:19:09.156272Z","shell.execute_reply.started":"2023-04-25T18:03:45.046128Z","shell.execute_reply":"2023-04-25T18:19:09.155069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ⦿ Further local training\n\n> 📌 **Note**: the current pipeline trains with only a sample of the data. I am doing the full training locally (ZBook Studio Mobile) as it happens faster and I can also do more sweeps. I am posting the results below and the best model [can be found here](https://www.kaggle.com/datasets/andradaolteanu/fog-2023-prepped-competition-data).\n\n**Best Sweep: [v4epryf2](https://wandb.ai/andrada/fog_2023/sweeps/v4epryf2?workspace=user-andrada)**\n\n<center><img src=\"https://i.imgur.com/TU66EeQ.png\"></center>","metadata":{}},{"cell_type":"markdown","source":"# 5. Final Model\n\nTraining the final model with the entire data using the most optimum hyperparameters:","metadata":{}},{"cell_type":"code","source":"def train_pipeline_final(config):\n    '''\n    Function to train and save the final model.\n    To be used after hyperparameter tuning.\n    '''\n\n    start = time()\n\n    dtrain, dvalid, _ = get_simple_data_validation(train_df=train_full)\n    total_train = cudf.concat([dtrain, dvalid], axis=0).reset_index(drop=True)\n\n    feature_cols = [\"AccV\", \"AccML\", \"AccAP\"]\n    # Matrix\n    dtrain_matrix = xgb.DMatrix(total_train[feature_cols],\n                                label=total_train[\"target\"])\n\n    # Training ...\n    model = xgb.train(config, dtrain_matrix, \n                      num_boost_round=100,\n                      verbose_eval=False)\n    print(\"The model has finished training.\")\n\n    model.save_model(\"baseline.model\")\n    print(\"The model has been saved successfully.\")\n\n    print(\"-----------------------------------\")\n    print(\"Total time:\", round((time()-start)/60, 1), \"mins.\")","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:19:09.158298Z","iopub.execute_input":"2023-04-25T18:19:09.159779Z","iopub.status.idle":"2023-04-25T18:19:09.168395Z","shell.execute_reply.started":"2023-04-25T18:19:09.159733Z","shell.execute_reply":"2023-04-25T18:19:09.166913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Based on the sweeps the best combo of hyperparameters is:\nfinal_configs = {\n        'objective': 'multi:softmax',\n        'eval_metric': 'mlogloss',\n        'num_class': 4,\n        'tree_method': 'gpu_hist',\n        \"random_state\": 10,\n        \"learning_rate\": 0.2,\n        \"max_depth\": 15,\n        \"min_child_weight\": 5,\n}\n\n# Training the final model on full data\ntrain_pipeline_final(final_configs)","metadata":{"execution":{"iopub.status.busy":"2023-04-25T18:19:09.170682Z","iopub.execute_input":"2023-04-25T18:19:09.171474Z","iopub.status.idle":"2023-04-25T18:21:07.381395Z","shell.execute_reply.started":"2023-04-25T18:19:09.171434Z","shell.execute_reply":"2023-04-25T18:21:07.380230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 🧠 Submission\n\n*Work in progress: this will be done in another notebook, as this one is a code competition and Internet cannot be enabled to make a submission. I will post a link to it when I have a final baseline model.*","metadata":{}},{"cell_type":"markdown","source":"### 🐝 [W&B Dashboard - find all the logs and sweeps here :)](https://wandb.ai/andrada/fog_2023)\n    \n<center><img src=\"https://i.imgur.com/9okWacK.png\"></center>\n\n------\n\n<center><img src=\"https://i.imgur.com/knxTRkO.png\"></center>\n\n### My Specs\n\n* 🖥 Z8 G4 Workstation\n* 💾 2 CPUs & 96GB Memory\n* 🎮 2x NVIDIA A6000\n* 💻 Zbook Studio G7 on the go","metadata":{}}]}