{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":46105,"databundleVersionId":5087314,"sourceType":"competition"},{"sourceId":5082760,"sourceType":"datasetVersion","datasetId":2950885}],"dockerImageVersionId":30407,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nThis notebook is created for [Kaggle's Sign Language Classifier Competition](https://www.kaggle.com/competitions/asl-signs/code)\n\nIn this notebook, I used notebooks from various sources and modify them to get this notebook. Please give upvotes to them if you also find it useful:\n\n- I used some of the data visualization of the landmark from [Sign Language EDA & Visualization](https://www.kaggle.com/code/mayukh18/sign-language-eda-visualization).\n\n- I used the preprocessed tensorflow Dataset from [tfdataset-of-google-isl-recognition-data](https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data).\n\n- I train my model following the notebook. [Submission for variable length time-series model](https://www.kaggle.com/code/aapokossi/submission-for-variable-length-time-series-model). I then tweak the layers and the epoch to increase the accuracy.\n\nI also included comments and links to help study the model further.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"contents\"></a>\n# Contents\n1. [Import Libraries and Set File Directories](#section-one)\n2. [Visualize data](#section-two)\n3. [Load Data](#section-three)\n4. [Train Model](#section-four)\n5. [Submit Model](#section-five)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"section-one\"></a>\n# Import Libraries and Set File Directories","metadata":{}},{"cell_type":"code","source":"# import libraries\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nfrom tensorflow.keras import layers, optimizers","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-03-15T15:43:10.125587Z","iopub.execute_input":"2023-03-15T15:43:10.125901Z","iopub.status.idle":"2023-03-15T15:43:18.926505Z","shell.execute_reply.started":"2023-03-15T15:43:10.125873Z","shell.execute_reply":"2023-03-15T15:43:18.925293Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# set files directories\nLANDMARK_FILES_DIR = \"/kaggle/input/asl-signs/train_landmark_files\"\nTRAIN_FILE = \"/kaggle/input/asl-signs/train.csv\"","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:18.929437Z","iopub.execute_input":"2023-03-15T15:43:18.930584Z","iopub.status.idle":"2023-03-15T15:43:18.936067Z","shell.execute_reply.started":"2023-03-15T15:43:18.930543Z","shell.execute_reply":"2023-03-15T15:43:18.934806Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"section-two\"></a>\n# Visualize data","metadata":{}},{"cell_type":"markdown","source":"Slightly different from a dataframe, in this competition we will be using a [parquet](https://towardsdatascience.com/demystifying-the-parquet-file-format-13adb0206705). We will read it using the command `read_parquet`.","metadata":{}},{"cell_type":"code","source":"# read the data, the type of the data\nsample = pd.read_parquet(\"/kaggle/input/asl-signs/train_landmark_files/16069/100015657.parquet\")\nsample.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:18.937533Z","iopub.execute_input":"2023-03-15T15:43:18.937956Z","iopub.status.idle":"2023-03-15T15:43:19.129122Z","shell.execute_reply.started":"2023-03-15T15:43:18.937917Z","shell.execute_reply":"2023-03-15T15:43:19.127943Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Everyone likes to visualize things right? Since the data is about sequences of hand movement, we might not be able to see it from tabular data only. Let us instead plot the movement into matplotlib.","metadata":{}},{"cell_type":"code","source":"# pick the left hand and right hand points\nsample_left_hand = sample[sample.type == \"left_hand\"]\nsample_right_hand = sample[sample.type == \"right_hand\"]\n\n# edges that represents the hand edges\nedges = [(0,1),(1,2),(2,3),(3,4),(0,5),(0,17),(5,6),(6,7),(7,8),(5,9),(9,10),(10,11),(11,12),\n         (9,13),(13,14),(14,15),(15,16),(13,17),(17,18),(18,19),(19,20)]\n\n# plotting a single frame into matplotlib\ndef plot_frame(df, frame_id, ax):\n    df = df[df.frame == frame_id].sort_values(['landmark_index'])\n    x = list(df.x)\n    y = list(df.y)\n    \n    # plotting the points\n    ax.scatter(df.x, df.y, color='dodgerblue')\n    for i in range(len(x)):\n        ax.text(x[i], y[i], str(i))\n    \n    # plotting the edges that represents the hand\n    for edge in edges:\n        ax.plot([x[edge[0]], x[edge[1]]], [y[edge[0]], y[edge[1]]], color='salmon')\n        ax.set_xlabel(f\"Frame no. {frame_id}\")\n        ax.set_xticks([])\n        ax.set_yticks([])\n        ax.set_xticklabels([])\n        ax.set_yticklabels([])\n\n# plotting the multiple frames\ndef plot_frame_seq(df, frame_range, n_frames):\n    frames = np.linspace(frame_range[0],frame_range[1],n_frames, dtype = int, endpoint=True)\n    fig, ax = plt.subplots(n_frames, 1, figsize=(5,25))\n    for i in range(n_frames):\n        plot_frame(df, frames[i], ax[i])\n        \n    plt.show()\n\nplot_frame_seq(sample_left_hand, (178,186), 10)","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:19.132091Z","iopub.execute_input":"2023-03-15T15:43:19.132905Z","iopub.status.idle":"2023-03-15T15:43:20.595271Z","shell.execute_reply.started":"2023-03-15T15:43:19.132866Z","shell.execute_reply":"2023-03-15T15:43:20.594277Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"section-three\"></a>\n# Load Data","metadata":{}},{"cell_type":"code","source":"# Set constants and pick important landmarks\nLANDMARK_IDX = [0,9,11,13,14,17,117,118,119,199,346,347,348] + list(range(468,543))\nDATA_PATH = \"/kaggle/input/saved-tfdataset-of-google-isl-recognition-data/GoogleISLDatasetBatched\"\nDS_CARDINALITY = 185\nVAL_SIZE  = 18\nN_SIGNS = 250\nROWS_PER_FRAME = 543","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:20.596411Z","iopub.execute_input":"2023-03-15T15:43:20.596738Z","iopub.status.idle":"2023-03-15T15:43:20.603814Z","shell.execute_reply.started":"2023-03-15T15:43:20.596707Z","shell.execute_reply":"2023-03-15T15:43:20.602744Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"To keep it simple, we will use the preprocessed [tf.Dataset](https://www.tensorflow.org/api_docs/python/tf/data/Dataset) from [tfdataset-of-google-isl-recognition-data](https://www.kaggle.com/datasets/aapokossi/saved-tfdataset-of-google-isl-recognition-data).","metadata":{}},{"cell_type":"code","source":"def preprocess(ragged_batch, labels):\n    ragged_batch = tf.gather(ragged_batch, LANDMARK_IDX, axis=2)\n    ragged_batch = tf.where(tf.math.is_nan(ragged_batch), tf.zeros_like(ragged_batch), ragged_batch)\n    return tf.concat([ragged_batch[...,i] for i in range(3)],-1), labels\n\ndataset = tf.data.Dataset.load(DATA_PATH)\ndataset = dataset.map(preprocess)\nval_ds = dataset.take(VAL_SIZE).cache().prefetch(tf.data.AUTOTUNE)\ntrain_ds = dataset.skip(VAL_SIZE).cache().shuffle(20).prefetch(tf.data.AUTOTUNE)","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:20.605775Z","iopub.execute_input":"2023-03-15T15:43:20.606154Z","iopub.status.idle":"2023-03-15T15:43:24.216343Z","shell.execute_reply.started":"2023-03-15T15:43:20.606124Z","shell.execute_reply":"2023-03-15T15:43:24.215100Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"section-four\"></a>\n# Train Model","metadata":{}},{"cell_type":"code","source":"# include early stopping and reducelr\ndef get_callbacks():\n    return [\n            tf.keras.callbacks.EarlyStopping(\n            monitor=\"val_accuracy\",\n            patience = 10,\n            restore_best_weights=True\n        ),\n        tf.keras.callbacks.ReduceLROnPlateau(\n            monitor = \"val_accuracy\",\n            factor = 0.5,\n            patience = 3\n        ),\n    ]\n\n# a single dense block followed by a normalization block and relu activation\ndef dense_block(units, name):\n    fc = layers.Dense(units)\n    norm = layers.LayerNormalization()\n    act = layers.Activation(\"relu\")\n    return lambda x: act(norm(fc(x)))\n\n# the lstm block with the final dense block for the classification\ndef classifier(lstm_units):\n    lstm = layers.LSTM(lstm_units)\n    out = layers.Dense(N_SIGNS, activation=\"softmax\")\n    drop = layers.Dropout(0.2)\n    return lambda x: drop(out(lstm(x)))","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:24.217648Z","iopub.execute_input":"2023-03-15T15:43:24.217995Z","iopub.status.idle":"2023-03-15T15:43:24.225729Z","shell.execute_reply.started":"2023-03-15T15:43:24.217956Z","shell.execute_reply":"2023-03-15T15:43:24.224596Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now let us get to the fun part, training the model!","metadata":{}},{"cell_type":"code","source":"# choose the number of nodes per layer\nencoder_units = [512, 256] # tune this\nlstm_units = 250 # tune this\n\n#define the inputs (ragged batches of time series of landmark coordinates)\ninputs = tf.keras.Input(shape=(None,3*len(LANDMARK_IDX)), ragged=True)\n\n# dense encoder model\nx = inputs\nfor i, n in enumerate(encoder_units):\n    x = dense_block(n, f\"encoder_{i}\")(x)\n\n# classifier model\nout = classifier(lstm_units)(x)\n\nmodel = tf.keras.Model(inputs=inputs, outputs=out)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:24.227353Z","iopub.execute_input":"2023-03-15T15:43:24.228292Z","iopub.status.idle":"2023-03-15T15:43:24.745258Z","shell.execute_reply.started":"2023-03-15T15:43:24.228257Z","shell.execute_reply":"2023-03-15T15:43:24.744428Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# add a decreasing learning rate scheduler to help convergence\nsteps_per_epoch = DS_CARDINALITY - VAL_SIZE\nboundaries = [steps_per_epoch * n for n in [20,40,60]]\nvalues = [1e-3,1e-4,1e-5,1e-6]\nlr_sched = optimizers.schedules.PiecewiseConstantDecay(boundaries, values)\noptimizer = optimizers.Adam(lr_sched)\n\nmodel.compile(optimizer=optimizer,\n              loss=\"sparse_categorical_crossentropy\",\n              metrics=[\"accuracy\",\"sparse_top_k_categorical_accuracy\"])","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:24.746518Z","iopub.execute_input":"2023-03-15T15:43:24.746939Z","iopub.status.idle":"2023-03-15T15:43:24.807006Z","shell.execute_reply.started":"2023-03-15T15:43:24.746911Z","shell.execute_reply":"2023-03-15T15:43:24.806075Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# fit the model with 100 epochs iteration\nmodel.fit(train_ds,\n          validation_data = val_ds,\n          callbacks = get_callbacks(),\n          epochs = 80)","metadata":{"execution":{"iopub.status.busy":"2023-03-15T15:43:24.810536Z","iopub.execute_input":"2023-03-15T15:43:24.810808Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<a id=\"section-five\"></a>\n# Submit Model","metadata":{}},{"cell_type":"markdown","source":"Now it is time to submit. In this competition, we should submit the model itself.","metadata":{}},{"cell_type":"code","source":"model.summary(expand_nested=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def get_inference_model(model):\n    inputs = tf.keras.Input(shape=(ROWS_PER_FRAME,3), name=\"inputs\")\n    \n    # drop most of the face mesh\n    x = tf.gather(inputs, LANDMARK_IDX, axis=1)\n\n    # fill nan\n    x = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n\n    # flatten landmark xyz coordinates ()\n    x = tf.concat([x[...,i] for i in range(3)], -1)\n\n    x = tf.expand_dims(x,0)\n    \n    # call trained model\n    out = model(x)\n    \n    # explicitly name the final (identity) layer for the submission format\n    out = layers.Activation(\"linear\", name=\"outputs\")(out)\n    \n    inference_model = tf.keras.Model(inputs=inputs, outputs=out)\n    inference_model.compile(loss=\"sparse_categorical_crossentropy\",\n                            metrics=\"accuracy\")\n    return inference_model","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"inference_model = get_inference_model(model)\ninference_model.summary(expand_nested=True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# save the model\nconverter = tf.lite.TFLiteConverter.from_keras_model(inference_model)\ntflite_model = converter.convert()\nmodel_path = \"model.tflite\"\n\n# submit the model\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)\n!zip submission.zip $model_path","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{},"outputs":[],"execution_count":null}]}