{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Isolated Sign Language Recognition with DNN\n\nIn this notebook, I will create Isolated Sign Language Recognition model for [Google - Isolated Sign Language Recognition Competition](https://www.kaggle.com/competitions/asl-signs) using DNN.  This dataset's training records have different number of frames, in order to make it easy to start with and train faster, I will calcuate the mean frame for each training data. This model will have input shape (n, 253, 3) and output shape (n, 250). During inference, it's very tricky. I will create an inference model to do data preprocessing like calculate mean frame of test file with input shape (None, 253, 3) and imputation, finally it will have output shape (250) to match the requirement of this competiton.\n\nThis model can get about 0.35 CV and 0.33 LB. In the beggining, the LB was 0. I guessed that was because of missing value. I added following code in inference model to replace missing value with 0, luckily it  matched my hypothesis and works.\n\n```python\nx = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n```","metadata":{}},{"cell_type":"markdown","source":"## Configuration","metadata":{}},{"cell_type":"code","source":"class CFG:\n    aggregation_data_path = \"../input/isolated-sign-language-aggregation-dataset/\"\n    data_path = \"../input/asl-signs/\"\n    quick_experiment = False\n    is_training = True\n    use_aggregation_dataset = True\n    num_classes = 250\n    rows_per_frame = 543 ","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:49.491165Z","iopub.execute_input":"2023-05-12T14:11:49.491605Z","iopub.status.idle":"2023-05-12T14:11:49.523117Z","shell.execute_reply.started":"2023-05-12T14:11:49.491570Z","shell.execute_reply":"2023-05-12T14:11:49.521477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Import Packages","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tqdm.notebook import tqdm\nimport json\nimport os\nimport gc\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:49.526572Z","iopub.execute_input":"2023-05-12T14:11:49.527453Z","iopub.status.idle":"2023-05-12T14:11:59.302382Z","shell.execute_reply.started":"2023-05-12T14:11:49.527413Z","shell.execute_reply":"2023-05-12T14:11:59.301228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Utilities","metadata":{}},{"cell_type":"code","source":"def load_relevant_data_subset_with_imputation(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    data.replace(np.nan, 0, inplace=True)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef read_dict(file_path):\n    path = os.path.expanduser(file_path)\n    with open(path, \"r\") as f:\n        dic = json.load(f)\n    return dic","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.304119Z","iopub.execute_input":"2023-05-12T14:11:59.305000Z","iopub.status.idle":"2023-05-12T14:11:59.315525Z","shell.execute_reply.started":"2023-05-12T14:11:59.304954Z","shell.execute_reply":"2023-05-12T14:11:59.314375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(f\"{CFG.aggregation_data_path}train.csv\")\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.318914Z","iopub.execute_input":"2023-05-12T14:11:59.319394Z","iopub.status.idle":"2023-05-12T14:11:59.584963Z","shell.execute_reply.started":"2023-05-12T14:11:59.319354Z","shell.execute_reply":"2023-05-12T14:11:59.583954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 21 participants. Each of them create about 3000 to 5000 training records.","metadata":{}},{"cell_type":"code","source":"train.participant_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.586375Z","iopub.execute_input":"2023-05-12T14:11:59.587263Z","iopub.status.idle":"2023-05-12T14:11:59.599648Z","shell.execute_reply.started":"2023-05-12T14:11:59.587220Z","shell.execute_reply":"2023-05-12T14:11:59.598466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.participant_id.value_counts().plot(kind=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.601640Z","iopub.execute_input":"2023-05-12T14:11:59.602561Z","iopub.status.idle":"2023-05-12T14:11:59.945363Z","shell.execute_reply.started":"2023-05-12T14:11:59.602520Z","shell.execute_reply":"2023-05-12T14:11:59.944284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 94477 training samples in total.","metadata":{}},{"cell_type":"code","source":"len(train)","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.947149Z","iopub.execute_input":"2023-05-12T14:11:59.947579Z","iopub.status.idle":"2023-05-12T14:11:59.956172Z","shell.execute_reply.started":"2023-05-12T14:11:59.947530Z","shell.execute_reply":"2023-05-12T14:11:59.954957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 250 kinds of sign languages. Each kind of sign languages contains about 300 to 400 samples.","metadata":{}},{"cell_type":"code","source":"label_index = read_dict(f\"{CFG.data_path}sign_to_prediction_index_map.json\")\nindex_label = dict([(label_index[key], key) for key in label_index])\nprint(label_index)\ntrain[\"label\"] = train[\"sign\"].map(lambda sign: label_index[sign])\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:11:59.957418Z","iopub.execute_input":"2023-05-12T14:11:59.958134Z","iopub.status.idle":"2023-05-12T14:12:00.034969Z","shell.execute_reply.started":"2023-05-12T14:11:59.958094Z","shell.execute_reply":"2023-05-12T14:12:00.033600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"sign\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:12:00.036427Z","iopub.execute_input":"2023-05-12T14:12:00.037576Z","iopub.status.idle":"2023-05-12T14:12:00.246631Z","shell.execute_reply.started":"2023-05-12T14:12:00.037531Z","shell.execute_reply":"2023-05-12T14:12:00.245441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here are descriptive statistics for number of frames.","metadata":{}},{"cell_type":"code","source":"train.num_frames.describe()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:12:00.251939Z","iopub.execute_input":"2023-05-12T14:12:00.252301Z","iopub.status.idle":"2023-05-12T14:12:00.271045Z","shell.execute_reply.started":"2023-05-12T14:12:00.252265Z","shell.execute_reply":"2023-05-12T14:12:00.269733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Modeling\nI am still exploring how to handle this dataset. In order to make it easy to start with and train faster, I use mean frame as training input data. Loading data takes 30 minutes. To speedup training further, I will load mean frame data from another dataset which use the same preprocessing method.","metadata":{}},{"cell_type":"code","source":"if CFG.use_aggregation_dataset == False:\n    xs = []\n    ys = []\n    num_frames = np.zeros(len(train))\n    for i in tqdm(range(len(train))):\n        path = f\"{CFG.data_path}{train.iloc[i].path}\"\n        data = load_relevant_data_subset_with_imputation(path)\n        ## Mean Aggregation\n        xs.append(np.mean(data, axis=0))\n        ys.append(train.iloc[i].label)\n        num_frames[i] = data.shape[0]\n        if CFG.quick_experiment and i == 4999:\n            break\n    ## Save number of frames of each training sample for data analysis\n    train[\"num_frames\"] = num_frames\n    X = np.array(xs)\n    y = np.array(ys)\n    print(train[\"num_frames\"].describe())\n    train.to_csv(\"train.csv\", index=False)\nelse:\n    X = np.load(f\"{CFG.aggregation_data_path}X.npy\")\n    y = np.load(f\"{CFG.aggregation_data_path}y.npy\")\nprint(X.shape, y.shape)","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:12:00.272952Z","iopub.execute_input":"2023-05-12T14:12:00.273362Z","iopub.status.idle":"2023-05-12T14:12:06.339349Z","shell.execute_reply.started":"2023-05-12T14:12:00.273308Z","shell.execute_reply":"2023-05-12T14:12:06.337144Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_model():\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32)\n    vector = tf.keras.layers.Dense(128, activation=\"swish\")(inputs)\n    vector = tf.keras.layers.Dense(64, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dense(32, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dense(16, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dropout(0.1)(vector)\n    vector = tf.keras.layers.Flatten()(vector)\n    output = tf.keras.layers.Dense(250, activation=\"softmax\")(vector)\n    model = tf.keras.Model(inputs=inputs, outputs=output)\n    model.compile(\n        loss=tf.keras.losses.SparseCategoricalCrossentropy(), \n        metrics=[\n            \"accuracy\", \n            tf.keras.metrics.SparseTopKCategoricalAccuracy(k=5, name=\"top-5-accuracy\"),\n            tf.keras.metrics.SparseTopKCategoricalAccuracy(k=10, name=\"top-10-accuracy\")\n        ]\n    )\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:12:06.341159Z","iopub.execute_input":"2023-05-12T14:12:06.341563Z","iopub.status.idle":"2023-05-12T14:12:06.353679Z","shell.execute_reply.started":"2023-05-12T14:12:06.341522Z","shell.execute_reply":"2023-05-12T14:12:06.352395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_name = \"model.h5\"\nif CFG.is_training:\n    X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)\n    print(X_train.shape, y_train.shape, X_val.shape, y_val.shape)\n    del X, y\n    gc.collect()\n    model = get_model()\n    callbacks = [\n        tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, restore_best_weights=True, monitor=\"val_accuracy\", mode=\"max\")\n    ]\n    model.fit(X_train, y_train, epochs=50, validation_data=(X_val, y_val), batch_size=128, callbacks=callbacks)\n    model.load_weights(model_name)\nelse:\n    model = tf.keras.models.load_model(f\"../input/sign-language-prediction-model/{model_name}\")\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:12:06.355977Z","iopub.execute_input":"2023-05-12T14:12:06.356855Z","iopub.status.idle":"2023-05-12T14:19:33.906594Z","shell.execute_reply.started":"2023-05-12T14:12:06.356812Z","shell.execute_reply":"2023-05-12T14:19:33.905744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create an inference model\nThis inference model wraps the previous trained DNN model and do following preprocessing:\n* Replace nan value with 0\n* calcuate mean frame of the input tensor with (None, 543, 3) shape and convert to (1, 543, 3) shape.","metadata":{}},{"cell_type":"code","source":"def get_inference_model(model):\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32, name=\"inputs\")\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    x = tf.reduce_mean(x, axis=0, keepdims=True)\n    x = model(x)\n    output = tf.keras.layers.Activation(activation=\"linear\", name=\"outputs\")(x)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=output) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:33.907815Z","iopub.execute_input":"2023-05-12T14:19:33.908184Z","iopub.status.idle":"2023-05-12T14:19:33.926737Z","shell.execute_reply.started":"2023-05-12T14:19:33.908143Z","shell.execute_reply":"2023-05-12T14:19:33.925775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inference_model = get_inference_model(model)\ninference_model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:33.927834Z","iopub.execute_input":"2023-05-12T14:19:33.928222Z","iopub.status.idle":"2023-05-12T14:19:34.033910Z","shell.execute_reply.started":"2023-05-12T14:19:33.928182Z","shell.execute_reply":"2023-05-12T14:19:34.033052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Create submission file\nUnlike a csv file in our previous competitions, the submission file has to be a compressed tflite file named submissiom.zip.","metadata":{}},{"cell_type":"code","source":"converter = tf.lite.TFLiteConverter.from_keras_model(inference_model)\ntflite_model = converter.convert()\nmodel_path = \"model.tflite\"\n# Save the model.\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:34.035087Z","iopub.execute_input":"2023-05-12T14:19:34.035482Z","iopub.status.idle":"2023-05-12T14:19:38.268172Z","shell.execute_reply.started":"2023-05-12T14:19:34.035440Z","shell.execute_reply":"2023-05-12T14:19:38.266994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!zip submission.zip $model_path","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:38.269652Z","iopub.execute_input":"2023-05-12T14:19:38.270036Z","iopub.status.idle":"2023-05-12T14:19:39.719637Z","shell.execute_reply.started":"2023-05-12T14:19:38.269996Z","shell.execute_reply":"2023-05-12T14:19:39.718329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Making inferences\nLet's make inferences using TFLite interpreter, our submission file will be used to make inferences like following ways. To test inference speed, I will make inferences with 10000 samples. The converted TFLite model should be able to make at least 10 inferences per second.","metadata":{}},{"cell_type":"code","source":"!pip install tflite-runtime","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:39.722909Z","iopub.execute_input":"2023-05-12T14:19:39.723350Z","iopub.status.idle":"2023-05-12T14:19:52.936557Z","shell.execute_reply.started":"2023-05-12T14:19:39.723286Z","shell.execute_reply":"2023-05-12T14:19:52.935214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tflite_runtime.interpreter as tflite\ninterpreter = tflite.Interpreter(model_path)\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\nfor i in tqdm(range(10000)):\n    frames = load_relevant_data_subset(f'/kaggle/input/asl-signs/{train.iloc[i].path}')\n    output = prediction_fn(inputs=frames)\n    sign = np.argmax(output[\"outputs\"])\n    if i % 100 == 0:\n        print(f\"Predicted Label: {index_label[sign]}, Actual Label: {train.iloc[i].sign}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-12T14:19:52.939936Z","iopub.execute_input":"2023-05-12T14:19:52.940397Z","iopub.status.idle":"2023-05-12T14:23:21.211598Z","shell.execute_reply.started":"2023-05-12T14:19:52.940348Z","shell.execute_reply":"2023-05-12T14:23:21.210390Z"},"trusted":true},"execution_count":null,"outputs":[]}]}