{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"I will use DNN in this notebook to develop an isolated sign language recognition model for Google's competition in isolated sign language recognition. Because the training records for this dataset contain a variety of frame counts, I will calculate the mean frame for each set of training records to make it simpler to get started and train more quickly. The input shape for this model will be (n, 253, 3) and the output shape (n, 250). It's very difficult to do during inference. In order to meet the requirements of this competition, I will develop an inference model to perform data preprocessing such as calculate the mean frame of the test file with input shape (None, 253, 3) and imputation. Ultimately, it will have output shape (250).","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class CFG:\n    aggregation_data_path = \"../input/isolated-sign-language-aggregation-dataset/\"\n    data_path = \"../input/asl-signs/\"\n    quick_experiment = False\n    is_training = True\n    use_aggregation_dataset = True\n    num_classes = 250\n    rows_per_frame = 543 ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tqdm.notebook import tqdm\nimport json\nimport os\nimport gc\nfrom sklearn.model_selection import train_test_split","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_relevant_data_subset_with_imputation(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    data.replace(np.nan, 0, inplace=True)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / CFG.rows_per_frame)\n    data = data.values.reshape(n_frames, CFG.rows_per_frame, len(data_columns))\n    return data.astype(np.float32)\n\ndef read_dict(file_path):\n    path = os.path.expanduser(file_path)\n    with open(path, \"r\") as f:\n        dic = json.load(f)\n    return dic","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(f\"{CFG.aggregation_data_path}train.csv\")\ntrain.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 21 people taking part. They each produce between 3,000 and 5,000 training records.","metadata":{}},{"cell_type":"code","source":"train.participant_id.nunique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.participant_id.value_counts().plot(kind=\"bar\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 250 distinct sign languages for which we must make projections. Each type of sign language has between 300 and 400 examples.","metadata":{}},{"cell_type":"code","source":"label_index = read_dict(f\"{CFG.data_path}sign_to_prediction_index_map.json\")\nindex_label = dict([(label_index[key], key) for key in label_index])\nprint(label_index)\ntrain[\"label\"] = train[\"sign\"].map(lambda sign: label_index[sign])\ntrain.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"sign\"].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These statistics are descriptive of the number of frames.","metadata":{}},{"cell_type":"code","source":"train.num_frames.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"> Modeling\n","metadata":{}},{"cell_type":"markdown","source":"I continue to investigate how to manage this dataset. I use mean frame as training input data to make it simple to start with and train more quickly. It takes 30 minutes to load the data. I'll load mean frame data from another dataset that uses the same preprocessing technique to speed up training even more.","metadata":{}},{"cell_type":"code","source":"if CFG.use_aggregation_dataset == False:\n    xs = []\n    ys = []\n    num_frames = np.zeros(len(train))\n    for i in tqdm(range(len(train))):\n        path = f\"{CFG.data_path}{train.iloc[i].path}\"\n        data = load_relevant_data_subset_with_imputation(path)\n        ## Mean Aggregation\n        xs.append(np.mean(data, axis=0))\n        ys.append(train.iloc[i].label)\n        num_frames[i] = data.shape[0]\n        if CFG.quick_experiment and i == 4999:\n            break\n    ## Save number of frames of each training sample for data analysis\n    train[\"num_frames\"] = num_frames\n    X = np.array(xs)\n    y = np.array(ys)\n    print(train[\"num_frames\"].describe())\n    train.to_csv(\"train.csv\", index=False)\nelse:\n    X = np.load(f\"{CFG.aggregation_data_path}X.npy\")\n    y = np.load(f\"{CFG.aggregation_data_path}y.npy\")\nprint(X.shape, y.shape)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_model():\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32)\n    vector = tf.keras.layers.Dense(128, activation=\"swish\")(inputs)\n    vector = tf.keras.layers.Dense(64, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dense(32, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dense(16, activation=\"swish\")(vector)\n    vector = tf.keras.layers.Dropout(0.1)(vector)\n    vector = tf.keras.layers.Flatten()(vector)\n    output = tf.keras.layers.Dense(250, activation=\"softmax\")(vector)\n    model = tf.keras.Model(inputs=inputs, outputs=output)\n    model.compile(\n        loss=tf.keras.losses.SparseCategoricalCrossentropy(), \n        metrics=[\n            \"accuracy\", \n            tf.keras.metrics.SparseTopKCategoricalAccuracy(k=5, name=\"top-5-accuracy\"),\n            tf.keras.metrics.SparseTopKCategoricalAccuracy(k=10, name=\"top-10-accuracy\")\n        ]\n    )\n    return model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_name = \"model.h5\"\nif CFG.is_training:\n    X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)\n    print(X_train.shape, y_train.shape, X_val.shape, y_val.shape)\n    del X, y\n    gc.collect()\n    model = get_model()\n    callbacks = [\n        tf.keras.callbacks.ModelCheckpoint(model_name, save_best_only=True, restore_best_weights=True, monitor=\"val_accuracy\", mode=\"max\")\n    ]\n    model.fit(X_train, y_train, epochs=50, validation_data=(X_val, y_val), batch_size=128, callbacks=callbacks)\n    model.load_weights(model_name)\nelse:\n    model = tf.keras.models.load_model(f\"../input/sign-language-prediction-model/{model_name}\")\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> CREATE AN INFERENCE MODEL","metadata":{}},{"cell_type":"markdown","source":"This conclusion The model encloses the previously trained DNN model and performs the preprocessing listed below:\n\nIn order to calculate the mean frame of the input tensor with (None, 543, 3) shape and convert it to (1, 543, 3) shape, replace the nan value with 0.","metadata":{}},{"cell_type":"code","source":"def get_inference_model(model):\n    inputs = tf.keras.Input((543, 3), dtype=tf.float32, name=\"inputs\")\n    x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\n    x = tf.reduce_mean(x, axis=0, keepdims=True)\n    x = model(x)\n    output = tf.keras.layers.Activation(activation=\"linear\", name=\"outputs\")(x)\n    inference_model = tf.keras.Model(inputs=inputs, outputs=output) \n    inference_model.compile(loss=tf.keras.losses.SparseCategoricalCrossentropy(), metrics=[\"accuracy\"])\n    return inference_model","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inference_model = get_inference_model(model)\ninference_model.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> CREATE A SUBMISSION FILE","metadata":{}},{"cell_type":"markdown","source":"The submission file must be a compressed Tflite file with the name submission.zip, as opposed to a csv file in our previous competitions.","metadata":{}},{"cell_type":"code","source":"converter = tf.lite.TFLiteConverter.from_keras_model(inference_model)\ntflite_model = converter.convert()\nmodel_path = \"model.tflite\"\n# Save the model.\nwith open(model_path, 'wb') as f:\n    f.write(tflite_model)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!zip submission.zip $model_path","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">MAKING INFERENCES ****","metadata":{}},{"cell_type":"markdown","source":"Let's use the TFLite interpreter to draw conclusions about our submission file in the ways described below. I'll make inferences using 10,000 samples to test the speed of inference. At least ten inferences per second should be achievable with the converted TFLite model.","metadata":{}},{"cell_type":"code","source":"!pip install tflite-runtime","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tflite_runtime.interpreter as tflite\ninterpreter = tflite.Interpreter(model_path)\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\nfor i in tqdm(range(10000)):\n    frames = load_relevant_data_subset(f'/kaggle/input/asl-signs/{train.iloc[i].path}')\n    output = prediction_fn(inputs=frames)\n    sign = np.argmax(output[\"outputs\"])\n    if i % 100 == 0:\n        print(f\"Predicted Label: {index_label[sign]}, Actual Label: {train.iloc[i].sign}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}