{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<a id=\"top\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:center;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Isolated Sign Language Recognition - Evidential ِDeep Learning</b></div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"text-align:center;\">\n    <img src=\"https://pbs.twimg.com/media/Fpq2dGiXoAUAmFw?format=png&name=4096x4096\" alt=\"isolated sign language recognition\" style=\"width:60%;\">\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Isolated Sign Language Recognition with Evidential Deep Learning\n\n\n\nIn this notebook I have used [Isolated Sign Language Recognition with DNN](https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn) by [Lonnie](https://www.kaggle.com/lonnieqin) as a base and tried to demonstrate the use of Evidential Deep Learning for Google Isolated Sign Language Recognition Competition.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"top\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Table of content</b></div>\n\n<div style=\"background-color:aliceblue; padding:30px; font-size:15px;color:#034914\">\n    \n<a id=\"TOC\"></a>\n## Table of Content\n* [Initial Configuration](#conf)\n* [Importing Required Libraries](#lib)\n* [Utilities](#utils)\n* [Loading Data](#data)\n* [Modelling](#model)\n    * [Defining Loss Functions for Evidential Deep Learning](#edl)\n* [Creating Model for inference](#minf)\n* [Creating Submission File](#sub)\n* [Making Predictions](#pred)\n* [Reference](#ref)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"conf\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Initial Configuration</b></div> ","metadata":{}},{"cell_type":"code","source":"class CFG:\n    data_path = \"../input/asl-signs/\"\n    quick_experiment = False\n    is_training = True\n    use_aggregation_dataset = True\n    num_classes = 250\n    rows_per_frame = 543 ","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:12.421159Z","iopub.execute_input":"2023-03-02T15:04:12.421432Z","iopub.status.idle":"2023-03-02T15:04:12.446628Z","shell.execute_reply.started":"2023-03-02T15:04:12.421405Z","shell.execute_reply":"2023-03-02T15:04:12.445707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"lib\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Importing Required Libraries</b></div>","metadata":{}},{"cell_type":"code","source":"!pip install pySankey -q","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:12.448463Z","iopub.execute_input":"2023-03-02T15:04:12.449129Z","iopub.status.idle":"2023-03-02T15:04:25.640251Z","shell.execute_reply.started":"2023-03-02T15:04:12.449090Z","shell.execute_reply":"2023-03-02T15:04:25.639038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tqdm import tqdm\nimport json\nimport os\nimport gc\nfrom sklearn.model_selection import train_test_split\n\n# Rquired Libraries \nimport matplotlib.pyplot as plt\nimport cv2\nfrom PIL import Image\n\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom keras.models import Sequential\nfrom keras.layers import Conv2D, MaxPool2D, Dense, Flatten, Dropout\nfrom scipy import ndimage\nfrom tensorflow.keras import backend as K\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:25.643776Z","iopub.execute_input":"2023-03-02T15:04:25.644610Z","iopub.status.idle":"2023-03-02T15:04:33.096792Z","shell.execute_reply.started":"2023-03-02T15:04:25.644574Z","shell.execute_reply":"2023-03-02T15:04:33.095718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"utils\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Utilities</b></div> ","metadata":{}},{"cell_type":"code","source":"def load_relevant_data_subset_with_imputation(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    data.replace(np.nan, 0, inplace=True)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef read_dict(file_path):\n    path = os.path.expanduser(file_path)\n    with open(path, \"r\") as f:\n        dic = json.load(f)\n    return dic","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.098385Z","iopub.execute_input":"2023-03-02T15:04:33.099391Z","iopub.status.idle":"2023-03-02T15:04:33.109691Z","shell.execute_reply.started":"2023-03-02T15:04:33.099350Z","shell.execute_reply":"2023-03-02T15:04:33.108566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"data\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Loading Data</b></div>","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(f\"{CFG.data_path}train.csv\")\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.112805Z","iopub.execute_input":"2023-03-02T15:04:33.113186Z","iopub.status.idle":"2023-03-02T15:04:33.311227Z","shell.execute_reply.started":"2023-03-02T15:04:33.113150Z","shell.execute_reply":"2023-03-02T15:04:33.310248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 21 participants. Each of them create about 3000 to 5000 training records.","metadata":{}},{"cell_type":"code","source":"train.participant_id.value_counts().plot(kind=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.312872Z","iopub.execute_input":"2023-03-02T15:04:33.313233Z","iopub.status.idle":"2023-03-02T15:04:33.635755Z","shell.execute_reply.started":"2023-03-02T15:04:33.313198Z","shell.execute_reply":"2023-03-02T15:04:33.634695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 94477 training samples in total.","metadata":{}},{"cell_type":"markdown","source":"There are 250 kinds of sign languages that we need to make prediction on. Each kind of sign languages contains about 300 to 400 samples.","metadata":{}},{"cell_type":"code","source":"label_index = read_dict(f\"{CFG.data_path}sign_to_prediction_index_map.json\")\nindex_label = dict([(label_index[key], key) for key in label_index])\nprint(label_index)\ntrain[\"label\"] = train[\"sign\"].map(lambda sign: label_index[sign])\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.637047Z","iopub.execute_input":"2023-03-02T15:04:33.637709Z","iopub.status.idle":"2023-03-02T15:04:33.686695Z","shell.execute_reply.started":"2023-03-02T15:04:33.637671Z","shell.execute_reply":"2023-03-02T15:04:33.685675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"model\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Modelling</b></div>","metadata":{}},{"cell_type":"markdown","source":"<a id = \"edl\"></a>\n## Defining Loss Functions for Evidential Deep Learning\n\nThere are three different loss functions defined in the paper:\n\n#### 1) Integrating out the class probabilities from posterior of Dirichlet prior & Multinomial likelihood - will be mentioned as *Eqn. 3* (as in the paper)\n\n$$\n\\mathcal{L}_i(\\Theta) =\n- log ( \\int \\prod_{j=1}^K p_{ij}^{y_{ij}} \\frac{1}{B(\\alpha_i)} \\prod_{j=1}^K p_{ij}^{\\alpha_{ij} -1 } d\\boldsymbol{p}_i )\n= \\sum_{j=1}^K y_{ij} (log(S_i) - log(\\alpha_{ij}))\n$$\n\n#### 2) Using cross-entropy loss - will be mentioned as *Eqn. 4* (as in the paper)\n\n$$\n\\mathcal{L}_i(\\Theta) =\n\\int [\\sum_{j=1}^K -y_{ij} log(p_{ij})] \\frac{1}{B(\\alpha_i)} \\prod_{j=1}^K p_{ij}^{\\alpha_{ij} -1 } d\\boldsymbol{p}_i \n= \\sum_{j=1}^K y_{ij} (\\psi(S_i) - \\psi(\\alpha_{ij}))\n$$\n\n#### 3) Using sum of squares loss - will be mentioned as *Eqn. 5* (as in the paper)\n\n$$\n\\mathcal{L}_i(\\Theta) =\n\\int ||\\boldsymbol{y}_i - \\boldsymbol{p}_i||_2^2 \\frac{1}{B(\\alpha_i)} \\prod_{j=1}^K p_{ij}^{\\alpha_{ij} -1 } d\\boldsymbol{p}_i \n= \\sum_{j=1}^K \\mathbb{E}[(y_{ij} - p_{ij})^2]\n$$\n\n$$\n= \\sum_{j=1}^K \\mathbb{E}[y_{ij}^2 - 2 y_{ij}p_{ij} + p_{ij}^2] \n= \\sum_{j=1}^K (y_{ij}^2 - 2 y_{ij}\\mathbb{E}[p_{ij}] + \\mathbb{E}[p_{ij}^2])\n$$\n\n$$\n= \\sum_{j=1}^K (y_{ij}^2 - 2 y_{ij}\\mathbb{E}[p_{ij}] + \\mathbb{E}[p_{ij}]^2 + \\text{Var}(p_{ij}))\n= \\sum_{j=1}^K (y_{ij} - \\mathbb{E}[p_{ij}])^2 + \\text{Var}(p_{ij})\n$$\n\n$$\n= \\sum_{j=1}^K (y_{ij}^2 - 2 y_{ij}\\mathbb{E}[p_{ij}] + \\mathbb{E}[p_{ij}]^2 + \\text{Var}(p_{ij}))\n= \\sum_{j=1}^K (y_{ij} - \\mathbb{E}[p_{ij}])^2 + \\text{Var}(p_{ij})\n$$\n\n$$\n= \\sum_{j=1}^K (y_{ij} - \\frac{\\alpha_{ij}}{S_i})^2 + \\frac{\\alpha_{ij}(S_i - \\alpha_{ij})}{S_i^2(S_i + 1)}\n$$\n\n$$\n= \\sum_{j=1}^K (y_{ij} - \\hat{p}_{ij})^2 + \\frac{\\hat{p}_{ij}(1 - \\hat{p}_{ij})}{(S_i + 1)}\n$$\n\n#### Source: https://github.com/atilberk/evidential-deep-learning-to-quantify-classification-uncertainty\n#### Paper: https://arxiv.org/pdf/1806.01768.pdf","metadata":{}},{"cell_type":"code","source":"lgamma = tf.math.lgamma\ndigamma = tf.math.digamma\n\nepochs = [1]\n\ndef KL(alpha, num_classes=250):\n    one = K.constant(np.ones((1,num_classes)),dtype=tf.float32)\n    S = K.sum(alpha,axis=1,keepdims=True)  \n\n    kl = lgamma(S) - K.sum(lgamma(alpha),axis=1,keepdims=True) +\\\n    K.sum(lgamma(one),axis=1,keepdims=True) - lgamma(K.sum(one,axis=1,keepdims=True)) +\\\n    K.sum((alpha - one)*(digamma(alpha)-digamma(S)),axis=1,keepdims=True)\n          \n    return kl\n\n\ndef loss_func(y_true, output):\n    y_evidence = K.relu(output)\n    alpha = y_evidence+1\n    S = K.sum(alpha,axis=1,keepdims=True)\n    p = alpha / S  \n\n    y_true = tf.one_hot(tf.cast(y_true, tf.int32), depth=250)\n\n    err = K.sum(K.pow((y_true-p),2),axis=1,keepdims=True)\n    var = K.sum(alpha*(S-alpha)/(S*S*(S+1)),axis=1,keepdims=True)\n    \n    l =  K.sum(err + var,axis=1,keepdims=True)\n    l = K.sum(l)\n    \n    \n    kl =  K.minimum(1.0, epochs[0]/50) * K.sum(KL((1-y_true)*(alpha)+y_true))\n    return l + kl","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.687875Z","iopub.execute_input":"2023-03-02T15:04:33.688167Z","iopub.status.idle":"2023-03-02T15:04:33.701249Z","shell.execute_reply.started":"2023-03-02T15:04:33.688139Z","shell.execute_reply":"2023-03-02T15:04:33.700193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if CFG.is_training:\n    if CFG.use_aggregation_dataset == False:\n        xs = []\n        ys = []\n        num_frames = np.zeros(len(train))\n        for i in tqdm(range(len(train))):\n            path = f\"{CFG.data_path}{train.iloc[i].path}\"\n            data = load_relevant_data_subset_with_imputation(path)\n            ## Mean Aggregation\n            xs.append(np.mean(data, axis=0))\n            ys.append(train.iloc[i].label)\n            num_frames[i] = data.shape[0]\n            if CFG.quick_experiment and i == 4999:\n                break\n        ## Save number of frames of each training sample for data analysis\n        train[\"num_frames\"] = num_frames\n        X = np.array(xs)\n        y = np.array(ys)\n        print(train[\"num_frames\"].describe())\n        train.to_csv(\"train.csv\", index=False)\n    else:\n        X = np.load(\"/kaggle/input/isolated-sign-language-aggregation-dataset/X.npy\")\n        y = np.load(\"/kaggle/input/isolated-sign-language-aggregation-dataset/y.npy\")\n    print(X.shape, y.shape)","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:04:33.702785Z","iopub.execute_input":"2023-03-02T15:04:33.703694Z","iopub.status.idle":"2023-03-02T15:04:36.613174Z","shell.execute_reply.started":"2023-03-02T15:04:33.703623Z","shell.execute_reply":"2023-03-02T15:04:36.612035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_model():\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32)\n    vector = tf.keras.layers.Dense(512, activation=\"relu\")(inputs)\n    vector = tf.keras.layers.Dense(128, activation=\"relu\")(vector)\n    vector = tf.keras.layers.Dense(64, activation=\"relu\")(vector)\n    vector = tf.keras.layers.Dense(32, activation=\"relu\")(vector)\n    vector = tf.keras.layers.Dense(16, activation=\"relu\")(vector)\n    vector = tf.keras.layers.Flatten()(vector)\n    output = tf.keras.layers.Dense(250, activation=\"softmax\")(vector)\n    model = tf.keras.Model(inputs=inputs, outputs=output)\n    \n    model.compile(loss=[loss_func, tf.keras.losses.SparseCategoricalCrossentropy()], \n                  metrics=[\"accuracy\", \n                           tf.keras.metrics.SparseTopKCategoricalAccuracy(k=5, name=\"top-5-accuracy\"),\n                           tf.keras.metrics.SparseTopKCategoricalAccuracy(k=10, name=\"top-10-accuracy\")\n                          ])\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:32:01.335783Z","iopub.execute_input":"2023-03-02T15:32:01.336299Z","iopub.status.idle":"2023-03-02T15:32:01.347234Z","shell.execute_reply.started":"2023-03-02T15:32:01.336257Z","shell.execute_reply":"2023-03-02T15:32:01.345943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if CFG.is_training:\n    X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)\n    print(X_train.shape, y_train.shape, X_val.shape, y_val.shape)\n    del X, y\n    gc.collect()\n    model = get_model()\n    callbacks = [tf.keras.callbacks.ModelCheckpoint(\"model.h5\")]\n    model.fit(X_train, y_train, epochs=30, validation_data=(X_val, y_val), batch_size=128, callbacks=callbacks)\nelse:\n    model = tf.keras.models.load_model(\"/kaggle/input/sign-language-prediction-model/model.h5\")","metadata":{"execution":{"iopub.status.busy":"2023-03-02T15:32:37.198623Z","iopub.execute_input":"2023-03-02T15:32:37.198986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"minf\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Creating Model for inference</b></div>","metadata":{}},{"cell_type":"code","source":"def get_inference_model(model):\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32, name=\"inputs\")\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    x = tf.reduce_mean(x, axis=0, keepdims=True)\n    for i in range(1, len(model.layers)):\n        x = model.layers[i](x)\n    output = tf.keras.layers.Activation(activation=\"linear\", name=\"outputs\")(x)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=output) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inference_model = get_inference_model(model)\ninference_model.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"sub\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Creating Submission File</b></div>","metadata":{}},{"cell_type":"code","source":"converter = tf.lite.TFLiteConverter.from_keras_model(inference_model)\ntflite_model = converter.convert()\nmodel_path = \"model.tflite\"\n# Save the model.\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!zip submission.zip $model_path","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"pred\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>Making Predictions</b></div>","metadata":{}},{"cell_type":"code","source":"!pip install tflite-runtime","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The performance is not optimal so far. However it can make correct prediction sometimes.","metadata":{}},{"cell_type":"code","source":"import tflite_runtime.interpreter as tflite\ninterpreter = tflite.Interpreter(model_path)\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\nfor i in range(100):\n    frames = load_relevant_data_subset(f'/kaggle/input/asl-signs/{train.iloc[i].path}')\n    output = prediction_fn(inputs=frames)\n    sign = np.argmax(output[\"outputs\"])\n    print(f\"Predicted label: {index_label[sign]}, Actual Label: {train.iloc[i].sign}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"ref\"></a>\n# <div style=\"padding:20px;color:white;margin:0;font-size:35px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#003d99;overflow:hidden\"><b>References</b></div>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:aliceblue; padding:30px; font-size:15px;color:#034914\">\n\n* In this notebook I have used [Isolated Sign Language Recognition with DNN](https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn) by [Lonnie](https://www.kaggle.com/lonnieqin) as a base and tried to demonstrate the use of Evidential Deep Learning for Google Isolated Sign Language Recognition Competition.","metadata":{}},{"cell_type":"markdown","source":"<center> <a href=\"#TOC\" role=\"button\" aria-pressed=\"true\" >⬆️Back to Table of Contents ⬆️</a>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"border-radius:10px;border:#034914 solid;padding: 15px;background-color:aliceblue;font-size:90%;text-align:left\">\n\n<h4><b>Authors :</b> Koorosh Aslansefat </h4>  \n    \n<center> <strong> If you liked this Notebook, please do upvote. </strong>\n    \n<center> <strong> If you have any questions, feel free to put a comment! </strong>","metadata":{}},{"cell_type":"markdown","source":"<center> <img src=\"https://gregcfuzion.files.wordpress.com/2022/01/kind-regards-2.png\" style='width: 600px; height: 300px;'>","metadata":{}}]}