{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Context**\n\n**Every day, 33 babies are born with permanent hearing loss in the U.S.**\nWithout sign language, deaf babies are at risk of Language Deprivation Syndrome. This syndrome is characterized by a lack of access to naturally occurring language acquisition during their critical language-learning years. It can cause serious impacts on different aspects of their lives, such as relationships, education, and employment.\n\n**Learning sign language is challenging.**\n\nLearning American Sign Language is as difficult for English speakers as learning Japanese. (jstor.org) It takes time and resources, which many parents don't have. PopSign is a smartphone game app that makes learning American Sign Language fun, interactive, and accessible. Players match videos of ASL signs with bubbles containing written English words to pop them.\n\n**Data Processing**\n\nOnly lips, hands and arm pose coordinates are used.\n\nA custom Tensorflow layer handles the data processing. In short, it filters all frames without coordinates for the hands and downsamples the input to 32 frames if it is too long.\n\n**Model**\n\nA transformer based model is used. The embedding layer makes an ambedding per landmark(lips/left hand/right hand/arm pose) and merges these embedding with fully connected layers. The transformer consists of just 2 blocks with a simple mean pooling and fully connected layers for classification.\n\nThis notebook is been taken reference from [@MARK WIJKHUIZEN](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training).","metadata":{}},{"cell_type":"code","source":"# Default imports\nimport numpy as np\nimport pandas as pd\n\n#Basic Imports\nimport os\nimport gc\nimport sys\nimport glob\nimport math\n\n# Progress Imports\nfrom tqdm.notebook import tqdm\n\n# Plotting Imports\nimport matplotlib.pyplot as plt\nimport matplotlib as mpl\nimport seaborn as sn\n\n# Algorithm Imports\nimport scipy\nimport sklearn\nfrom sklearn.model_selection import train_test_split, GroupShuffleSplit\n\n# Tensor Flow imports\nimport tensorflow as tf\nimport tensorflow_addons as tfa\n\nprint(f'Tensorflow V{tf.__version__}')\nprint(f'Keras V{tf.keras.__version__}')\nprint(f'Python V{sys.version}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:14.147159Z","iopub.execute_input":"2023-04-18T16:43:14.147703Z","iopub.status.idle":"2023-04-18T16:43:24.017749Z","shell.execute_reply.started":"2023-04-18T16:43:14.147662Z","shell.execute_reply":"2023-04-18T16:43:24.015465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configuration\nDetailed configuration for the model and valid hyper parameters defined also here.","metadata":{}},{"cell_type":"code","source":"# True: processing data from scratch, False: loads preprocessed data\n_preprocess_data = False\n_train_model = True\n# Session is interactive\n_is_interactive = os.environ['KAGGLE_KERNEL_RUN_TYPE'] == 'Interactive'\n# Session is verbose\n_verbose = 1 if _is_interactive else 2\n\n# Variable to define validation/training data set.\n# True: use 10% of participants as validation set\n# False: use all data for training\n_use_val = False\n\n# Number of rows\n_n_rows = 543\n# Numbers of dimenions\n_n_dims = 3\n# Dimensions\n_dim_names = ['x', 'y', 'z']\n# Number of seeds\n_seed = 42\n# Number of classes \n_num_classes = 250\n\n# Other Hyper parameters\n# Number of inputs\n_input_size = 64\n# Number of all batch signs\n_batch_all_signs = 4\n# Batch Size\n_batch_size = 256\n# Number of Epochs\n_n_epochs = 100\n# Learning Rate\n_lr_max = 1e-3\n# Number of warm-up epochs\n_n_warmup_epochs = 0\n# Relative Weight Decay\n_wd_ratio = 0.05\n# Mask Value\n_mask_val = 4237","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:31.698202Z","iopub.execute_input":"2023-04-18T16:43:31.698914Z","iopub.status.idle":"2023-04-18T16:43:31.705511Z","shell.execute_reply.started":"2023-04-18T16:43:31.698877Z","shell.execute_reply":"2023-04-18T16:43:31.704472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Engineering\n\nHere we define functions or functionalities that define or modify data for training the model.","metadata":{}},{"cell_type":"code","source":"# Read Training Data\nif _is_interactive or not _preprocess_data:\n    train_df = pd.read_csv('/kaggle/input/asl-signs/train.csv').sample(int(5e3), random_state=_seed)\nelse:\n    train_df = pd.read_csv('/kaggle/input/asl-signs/train.csv')\n\n_n_samples = len(train_df)\nprint(f'_n_samples: {_n_samples}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:35.522788Z","iopub.execute_input":"2023-04-18T16:43:35.523403Z","iopub.status.idle":"2023-04-18T16:43:35.756888Z","shell.execute_reply.started":"2023-04-18T16:43:35.523364Z","shell.execute_reply":"2023-04-18T16:43:35.755701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get complete file path to file\ndef get_file_path(path):\n    return f'/kaggle/input/asl-signs/{path}'\n\ntrain_df['file_path'] = train_df['path'].apply(get_file_path)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:46.965992Z","iopub.execute_input":"2023-04-18T16:43:46.966693Z","iopub.status.idle":"2023-04-18T16:43:46.980557Z","shell.execute_reply.started":"2023-04-18T16:43:46.966655Z","shell.execute_reply":"2023-04-18T16:43:46.979385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ordinally Encode Sign\nAdd the ordinally encode to training data frames.","metadata":{}},{"cell_type":"code","source":"# Add ordinally Encoded Sign\ntrain_df['sign_ord'] = train_df['sign'].astype('category').cat.codes\n\n# Dictionaries to translate sign <-> ordinal encoded sign\nsign_2_ord_df = train_df[['sign', 'sign_ord']].set_index('sign').squeeze().to_dict()\nord_2_sign_df = train_df[['sign_ord', 'sign']].set_index('sign_ord').squeeze().to_dict()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:48.464679Z","iopub.execute_input":"2023-04-18T16:43:48.465742Z","iopub.status.idle":"2023-04-18T16:43:48.487177Z","shell.execute_reply.started":"2023-04-18T16:43:48.465697Z","shell.execute_reply":"2023-04-18T16:43:48.486059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Details of training data frame\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:49.031506Z","iopub.execute_input":"2023-04-18T16:43:49.032063Z","iopub.status.idle":"2023-04-18T16:43:49.049884Z","shell.execute_reply.started":"2023-04-18T16:43:49.032028Z","shell.execute_reply":"2023-04-18T16:43:49.048982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Video Statistics -> Video data details\n_n = int(1e3) if (_is_interactive or not _preprocess_data) else int(10e3)\n_n_unique_frames = np.zeros(_n, dtype=np.uint16)\n_n_missing_frames = np.zeros(_n, dtype=np.uint16)\n_max_frame = np.zeros(_n, dtype=np.uint16)\n\n_percentiles = [0.01, 0.05, 0.25, 0.50, 0.75, 0.95, 0.99, 0.999]\n\nfor idx, file_path in enumerate(tqdm(train_df['file_path'].sample(_n, random_state=_seed))):\n    df = pd.read_parquet(file_path)\n    _n_unique_frames[idx] = df['frame'].nunique()\n    _n_missing_frames[idx] = (df['frame'].max() - df['frame'].min()) - df['frame'].nunique() + 1\n    _max_frame[idx] = df['frame'].max()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:43:49.713838Z","iopub.execute_input":"2023-04-18T16:43:49.714753Z","iopub.status.idle":"2023-04-18T16:44:13.403544Z","shell.execute_reply.started":"2023-04-18T16:43:49.714713Z","shell.execute_reply":"2023-04-18T16:44:13.402428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of unique frames in each video\nprint(\"======> Unique frames in each video\\n\")\ndisplay(pd.Series(_n_unique_frames).describe(percentiles=_percentiles).to_frame('_n_unique_frames'))\n\n# Number of missing frames, consecutive frames with missing intermediate frame, i.e. 1,2,4,5 -> 3 is missing\nprint(\"\\n======> Missing frames in each video\\n\")\ndisplay(pd.Series(_n_missing_frames).describe(percentiles=_percentiles).to_frame('_n_missing_frames'))\n\n# Maximum frame number\nprint(\"\\n======> Maximum frames in each video\\n\")\ndisplay(pd.Series(_max_frame).describe(percentiles=_percentiles).to_frame('_max_frame'))","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:13.407160Z","iopub.execute_input":"2023-04-18T16:44:13.407661Z","iopub.status.idle":"2023-04-18T16:44:13.442513Z","shell.execute_reply.started":"2023-04-18T16:44:13.407629Z","shell.execute_reply":"2023-04-18T16:44:13.441479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Landmark Indices\n\nLandmark indices are the indices of hand, lips, and face etc. which helps us to know the sign according to ASL. Here we are having the indices which can be pssible on standard body.","metadata":{}},{"cell_type":"code","source":"_use_types = ['left_hand', 'pose', 'right_hand']\n_start_idxs = 468\n_lips_idxs0 = np.array([\n        61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n        291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n        78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n        95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    ])\n# Landmark indices in original data\n_left_hand_idxs0 = np.arange(468,489)\n_right_hand_idxs0 = np.arange(522,543)\n_left_pose_idxs0 = np.array([502, 504, 506, 508, 510])\n_right_pose_idxs0 = np.array([503, 505, 507, 509, 511])\n_landmark_left_idxs_dominant0 = np.concatenate((_lips_idxs0, _left_hand_idxs0, _left_pose_idxs0))\n_landmark_right_idxs_dominant0 = np.concatenate((_lips_idxs0, _right_hand_idxs0, _right_pose_idxs0))\n_hands_idxs0 = np.concatenate((_left_hand_idxs0, _right_hand_idxs0), axis=0)\n_n_cols = _landmark_left_idxs_dominant0.size\n# Landmark indices in processed data\n_lips_idxs = np.argwhere(np.isin(_landmark_left_idxs_dominant0, _lips_idxs0)).squeeze()\n_left_hand_idxs = np.argwhere(np.isin(_landmark_left_idxs_dominant0, _left_hand_idxs0)).squeeze()\n_right_hand_idxs = np.argwhere(np.isin(_landmark_right_idxs_dominant0, _right_hand_idxs0)).squeeze() # if value come less change it\n_hand_idxs = np.argwhere(np.isin(_landmark_left_idxs_dominant0, _hands_idxs0)).squeeze()\n_pose_idxs = np.argwhere(np.isin(_landmark_left_idxs_dominant0, _left_pose_idxs0)).squeeze()\n\nprint(f'# HAND INDICES: {len(_hand_idxs)}, Number_of_columns: {_n_cols}')\n\n# Final Landmark data\n_lips_start = 0\n_left_hand_start = _lips_idxs.size\n_right_hand_start = _left_hand_start + _left_hand_idxs.size\n_pose_start = _right_hand_start + _right_hand_idxs.size\n\nprint(f'LIPS_START: {_lips_start}, LEFT_HAND_START: {_left_hand_start}, RIGHT_HAND_START: {_right_hand_start}, POSE_START: {_pose_start}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:15.858727Z","iopub.execute_input":"2023-04-18T16:44:15.859437Z","iopub.status.idle":"2023-04-18T16:44:15.876551Z","shell.execute_reply.started":"2023-04-18T16:44:15.859397Z","shell.execute_reply":"2023-04-18T16:44:15.875404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Engineering\n\nProcess data to be useful for Tensor Flow framework and models.","metadata":{}},{"cell_type":"code","source":"# Source: https://www.kaggle.com/competitions/asl-signs/overview/evaluation\n# Number of landmark per frame -> Number of rows -> 543(In ASL signs case)\n_rows_per_frame = _n_rows\n\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / _rows_per_frame)\n    data = data.values.reshape(n_frames, _rows_per_frame, len(data_columns))\n    return data.astype(np.float32)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:17.909083Z","iopub.execute_input":"2023-04-18T16:44:17.909743Z","iopub.status.idle":"2023-04-18T16:44:17.915855Z","shell.execute_reply.started":"2023-04-18T16:44:17.909705Z","shell.execute_reply":"2023-04-18T16:44:17.914585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n    Tensorflow layer to process data in TFLite\n    Data needs to be processed in the model itself, so we can not use Python\n\"\"\" \nclass PreprocessLayer(tf.keras.layers.Layer):\n    def __init__(self):\n        super(PreprocessLayer, self).__init__()\n        normalisation_correction = tf.constant([\n                    # Add 0.50 to left hand (original right hand) and substract 0.50 of right hand (original left hand)\n                    [0] * len(_lips_idxs) + [0.50] * len(_left_hand_idxs) + [0.50] * len(_pose_idxs),\n                    # Y coordinates stay intact\n                    [0] * len(_landmark_left_idxs_dominant0),\n                    # Z coordinates stay intact\n                    [0] * len(_landmark_left_idxs_dominant0),\n                ],\n                dtype=tf.float32,\n            )\n        self.normalisation_correction = tf.transpose(normalisation_correction, [1,0])\n        \n    def pad_edge(self, t, repeats, side):\n        if side == 'LEFT':\n            return tf.concat((tf.repeat(t[:1], repeats=repeats, axis=0), t), axis=0)\n        elif side == 'RIGHT':\n            return tf.concat((t, tf.repeat(t[-1:], repeats=repeats, axis=0)), axis=0)\n    \n    @tf.function(\n        input_signature=(tf.TensorSpec(shape=[None,_n_rows,_n_dims], dtype=tf.float32),),\n    )\n    def call(self, data0):\n        # Number of Frames in Video\n        _n_frames0 = tf.shape(data0)[0]\n        \n        # Find dominant hand by comparing summed absolute coordinates\n        left_hand_sum = tf.math.reduce_sum(tf.where(tf.math.is_nan(tf.gather(data0, _left_hand_idxs0, axis=1)), 0, 1))\n        right_hand_sum = tf.math.reduce_sum(tf.where(tf.math.is_nan(tf.gather(data0, _right_hand_idxs0, axis=1)), 0, 1))\n        left_dominant = left_hand_sum >= right_hand_sum\n        \n        # Count non NaN Hand values in each frame for the dominant hand\n        if left_dominant:\n            frames_hands_non_nan_sum = tf.math.reduce_sum(\n                    tf.where(tf.math.is_nan(tf.gather(data0, _left_hand_idxs0, axis=1)), 0, 1),\n                    axis=[1, 2],\n                )\n        else:\n            frames_hands_non_nan_sum = tf.math.reduce_sum(\n                    tf.where(tf.math.is_nan(tf.gather(data0, _right_hand_idxs0, axis=1)), 0, 1),\n                    axis=[1, 2],\n                )\n        \n        # Find frames indices with coordinates of dominant hand\n        non_empty_frames_idxs = tf.where(frames_hands_non_nan_sum > 0)\n        non_empty_frames_idxs = tf.squeeze(non_empty_frames_idxs, axis=1)\n        # Filter frames\n        data = tf.gather(data0, non_empty_frames_idxs, axis=0)\n        \n        # Cast Indices in float32 to be compatible with Tensorflow Lite\n        non_empty_frames_idxs = tf.cast(non_empty_frames_idxs, tf.float32)\n        # Normalize to start with 0\n        non_empty_frames_idxs -= tf.reduce_min(non_empty_frames_idxs)\n        \n        # Number of Frames in Filtered Video\n        _n_frames = tf.shape(data)[0]\n        \n        # Gather Relevant Landmark Columns\n        if left_dominant:\n            data = tf.gather(data, _landmark_left_idxs_dominant0, axis=1)\n        else:\n            data = tf.gather(data, _landmark_right_idxs_dominant0, axis=1)\n            data = (\n                    self.normalisation_correction + (\n                        (data - self.normalisation_correction) * tf.where(self.normalisation_correction != 0, -1.0, 1.0))\n                )\n        \n        # Video fits in INPUT_SIZE\n        if _n_frames < _input_size:\n            # Pad With -1 to indicate padding\n            non_empty_frames_idxs = tf.pad(non_empty_frames_idxs, [[0, _input_size - _n_frames]], constant_values=-1)\n            # Pad Data With Zeros\n            data = tf.pad(data, [[0, _input_size - _n_frames], [0,0], [0,0]], constant_values=0)\n            # Fill NaN Values With 0\n            data = tf.where(tf.math.is_nan(data), 0.0, data)\n            return data, non_empty_frames_idxs\n        \n        # Video needs to be downsampled to INPUT_SIZE\n        else:\n            # Repeat\n            if _n_frames < _input_size**2:\n                repeats = tf.math.floordiv(_input_size * _input_size, _n_frames0)\n                data = tf.repeat(data, repeats=repeats, axis=0)\n                non_empty_frames_idxs = tf.repeat(non_empty_frames_idxs, repeats=repeats, axis=0)\n\n            # Pad To Multiple Of Input Size\n            pool_size = tf.math.floordiv(len(data), _input_size)\n            if tf.math.mod(len(data), _input_size) > 0:\n                pool_size += 1\n\n            if pool_size == 1:\n                pad_size = (pool_size * _input_size) - len(data)\n            else:\n                pad_size = (pool_size * _input_size) % len(data)\n\n            # Pad Start/End with Start/End value\n            pad_left = tf.math.floordiv(pad_size, 2) + tf.math.floordiv(_input_size, 2)\n            pad_right = tf.math.floordiv(pad_size, 2) + tf.math.floordiv(_input_size, 2)\n            if tf.math.mod(pad_size, 2) > 0:\n                pad_right += 1\n\n            # Pad By Concatenating Left/Right Edge Values\n            data = self.pad_edge(data, pad_left, 'LEFT')\n            data = self.pad_edge(data, pad_right, 'RIGHT')\n\n            # Pad Non Empty Frame Indices\n            non_empty_frames_idxs = self.pad_edge(non_empty_frames_idxs, pad_left, 'LEFT')\n            non_empty_frames_idxs = self.pad_edge(non_empty_frames_idxs, pad_right, 'RIGHT')\n\n            # Reshape to Mean Pool\n            data = tf.reshape(data, [_input_size, -1, _n_cols, _n_dims])\n            non_empty_frames_idxs = tf.reshape(non_empty_frames_idxs, [_input_size, -1])\n\n            # Mean Pool\n            data = tf.experimental.numpy.nanmean(data, axis=1)\n            non_empty_frames_idxs = tf.experimental.numpy.nanmean(non_empty_frames_idxs, axis=1)\n\n            # Fill NaN Values With 0\n            data = tf.where(tf.math.is_nan(data), 0.0, data)\n            \n            return data, non_empty_frames_idxs\n    \npreprocess_layer = PreprocessLayer()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:19.069660Z","iopub.execute_input":"2023-04-18T16:44:19.071580Z","iopub.status.idle":"2023-04-18T16:44:21.903101Z","shell.execute_reply.started":"2023-04-18T16:44:19.071531Z","shell.execute_reply":"2023-04-18T16:44:21.901587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get file path function\n\"\"\"\n    face: 0:468\n    left_hand: 468:489\n    pose: 489:522\n    right_hand: 522:544\n        \n\"\"\"\ndef get_data(file_path):\n    # Load Raw Data\n    data = load_relevant_data_subset(file_path)\n    # Process Data Using Tensorflow\n    data = preprocess_layer(data)\n    \n    return data","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:21.905419Z","iopub.execute_input":"2023-04-18T16:44:21.907532Z","iopub.status.idle":"2023-04-18T16:44:21.913119Z","shell.execute_reply.started":"2023-04-18T16:44:21.907490Z","shell.execute_reply":"2023-04-18T16:44:21.912241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Dataset\nCreate dataset for training and validation process","metadata":{}},{"cell_type":"code","source":"# Get the full dataset\ndef preprocess_data():\n    # Create arrays to save data\n    X = np.zeros([_n_samples, _input_size, _n_cols, _n_dims], dtype=np.float32)\n    y = np.zeros([_n_samples], dtype=np.int32)\n    non_empty_frame_idxs = np.full([_n_samples, _input_size], -1, dtype=np.float32)\n\n    # Fill X/y\n    for row_idx, (file_path, sign_ord) in enumerate(tqdm(train_df[['file_path', 'sign_ord']].values)):\n        # Log message every 5000 samples\n        if row_idx % 5000 == 0:\n            print(f'Generated {row_idx}/{_n_samples}')\n\n        data, non_empty_frame_idxs = get_data(file_path)\n        X[row_idx] = data\n        y[row_idx] = sign_ord\n        non_empty_frame_idxs[row_idx] = non_empty_frame_idxs\n        # Sanity check, data should not contain NaN values\n        if np.isnan(data).sum() > 0:\n            print(row_idx)\n            return data\n\n    # Save X/y\n    np.save('X.npy', X)\n    np.save('y.npy', y)\n    np.save('non_empty_frame_idxs.npy', non_empty_frame_idxs)\n    \n    # Save Validation\n    splitter = GroupShuffleSplit(test_size=0.10, n_splits=2, random_state=_seed)\n    _participant_ids = train_df['participant_id'].values\n    train_idxs, val_idxs = next(splitter.split(X, y, groups=_participant_ids))\n\n    # Save Train\n    X_train = X[train_idxs]\n    non_empty_frame_idxs_train = non_empty_frame_idxs[train_idxs]\n    y_train = y[train_idxs]\n    np.save('X_train.npy', X_train)\n    np.save('y_train.npy', y_train)\n    np.save('non_empty_frame_idxs_train.npy', non_empty_frame_idxs_train)\n    \n    # Save Validation\n    X_val = X[val_idxs]\n    non_empty_frame_idxs_val = non_empty_frame_idxs[val_idxs]\n    y_val = y[val_idxs]\n    np.save('X_val.npy', X_val)\n    np.save('y_val.npy', y_val)\n    np.save('non_empty_frame_idxs_val.npy', non_empty_frame_idxs_val)\n    \n    # Split Statistics\n    print(f'Patient ID Intersection Train/Val: {set(_participant_ids[train_idxs]).intersection(_participant_ids[val_idxs])}')\n    print(f'X_train shape: {X_train.shape}, X_val shape: {X_val.shape}')\n    print(f'y_train shape: {y_train.shape}, y_val shape: {y_val.shape}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:23.405161Z","iopub.execute_input":"2023-04-18T16:44:23.405868Z","iopub.status.idle":"2023-04-18T16:44:23.417156Z","shell.execute_reply.started":"2023-04-18T16:44:23.405829Z","shell.execute_reply":"2023-04-18T16:44:23.415863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Preprocess All Data From Scratch\nif _preprocess_data:\n    preprocess_data()\n    _root_dir = '.'\nelse:\n    _root_dir = '/kaggle/input/gislr-dataset-public'\n    \n# Load Data\nif _use_val:\n    # Load Train\n    X_train = np.load(f'{_root_dir}/X_train.npy')\n    y_train = np.load(f'{_root_dir}/y_train.npy')\n    non_empty_frame_idxs_train = np.load(f'{_root_dir}/NON_EMPTY_FRAME_IDXS_TRAIN.npy')\n    # Load Val\n    X_val = np.load(f'{_root_dir}/X_val.npy')\n    y_val = np.load(f'{_root_dir}/y_val.npy')\n    non_empty_frame_idxs_val = np.load(f'{_root_dir}/NON_EMPTY_FRAME_IDXS_VAL.npy')\n    # Define validation Data\n    validation_data = ({ 'frames': X_val, 'non_empty_frame_idxs': non_empty_frame_idxs_val }, y_val)\nelse:\n    X_train = np.load(f'{_root_dir}/X.npy')\n    y_train = np.load(f'{_root_dir}/y.npy')\n    non_empty_frame_idxs_train = np.load(f'{_root_dir}/NON_EMPTY_FRAME_IDXS.npy')\n    validation_data = None\n\n# Train \n#print_shape_dtype([X_train, y_train, NON_EMPTY_FRAME_IDXS_TRAIN], ['X_train', 'y_train', 'NON_EMPTY_FRAME_IDXS_TRAIN'])\n# Val\n#if USE_VAL:\n    #print_shape_dtype([X_val, y_val, NON_EMPTY_FRAME_IDXS_VAL], ['X_val', 'y_val', 'NON_EMPTY_FRAME_IDXS_VAL'])\n# Sanity Check\nprint(f'# NaN Values X_train: {np.isnan(X_train).sum()}')\n\n# Class Count\ndisplay(pd.Series(y_train).value_counts().to_frame('Class Count').iloc[[0,1,2,3,4, -5,-4,-3,-2,-1]])","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:44:24.650430Z","iopub.execute_input":"2023-04-18T16:44:24.651357Z","iopub.status.idle":"2023-04-18T16:45:03.349613Z","shell.execute_reply.started":"2023-04-18T16:44:24.651307Z","shell.execute_reply":"2023-04-18T16:45:03.348524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Number Of Frames","metadata":{}},{"cell_type":"code","source":"# Vast majority of samples fits has less than 32 non empty frames\nn_empty_frames = (non_empty_frame_idxs_train != -1).sum(axis=1) \nn_empty_frames_waterfall = []\nfor n in tqdm(range(1,_input_size + 1)):\n    n_empty_frames_waterfall.append(sum(n_empty_frames >= n) / len(non_empty_frame_idxs_train) * 100)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:03.351777Z","iopub.execute_input":"2023-04-18T16:45:03.352162Z","iopub.status.idle":"2023-04-18T16:45:14.961348Z","shell.execute_reply.started":"2023-04-18T16:45:03.352125Z","shell.execute_reply":"2023-04-18T16:45:14.960285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Sampling\n\nHere we data sampling function to get the batch containing data of similar signs.","metadata":{}},{"cell_type":"code","source":"# Custom sampler to get a batch containing N times all signs\ndef get_train_batch_all_signs(X, y, non_empty_frame_idxs, n=_batch_all_signs):\n    # Arrays to store batch in\n    X_batch = np.zeros([_num_classes*n, _input_size, _n_cols, _n_dims], dtype=np.float32)\n    y_batch = np.arange(0, _num_classes, step=1/n, dtype=np.float32).astype(np.int64)\n    non_empty_frame_idxs_batch = np.zeros([_num_classes*n, _input_size], dtype=np.float32)\n    \n    # Dictionary mapping ordinally encoded sign to corresponding sample indices\n    clss2idxs = {}\n    for i in range(_num_classes):\n        clss2idxs[i] = np.argwhere(y == i).squeeze().astype(np.int32)\n            \n    while True:\n        # Fill batch arrays\n        for i in range(_num_classes):\n            idxs = np.random.choice(clss2idxs[i], n)\n            X_batch[i*n:(i+1)*n] = X[idxs]\n            non_empty_frame_idxs_batch[i*n:(i+1)*n] = non_empty_frame_idxs[idxs]\n        \n        yield { 'frames': X_batch, 'non_empty_frame_idxs': non_empty_frame_idxs_batch }, y_batch\n        \n# Calling batch function\ndummy_dataset = get_train_batch_all_signs(X_train, y_train, non_empty_frame_idxs_train)\nX_batch, y_batch = next(dummy_dataset)\n\nfor k, v in X_batch.items():\n    print(f'{k} shape: {v.shape}, dtype: {v.dtype}')\n\n# Batch shape/dtype\nprint(f'y_batch shape: {y_batch.shape}, dtype: {y_batch.dtype}')\n# Verify each batch contains each sign exactly N times\ndisplay(pd.Series(y_batch).value_counts().to_frame('Counts'))","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:14.962903Z","iopub.execute_input":"2023-04-18T16:45:14.963974Z","iopub.status.idle":"2023-04-18T16:45:15.043801Z","shell.execute_reply.started":"2023-04-18T16:45:14.963935Z","shell.execute_reply":"2023-04-18T16:45:15.042739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mean Values\nFunction to get the mean vlues for lips, hands and pose.","metadata":{}},{"cell_type":"code","source":"# Function for LIPS mean\ndef get_lips_mean_std():\n    # LIPS\n    _lips_mean_X = np.zeros([_lips_idxs.size], dtype=np.float32)\n    _lips_mean_Y = np.zeros([_lips_idxs.size], dtype=np.float32)\n    _lips_std_X = np.zeros([_lips_idxs.size], dtype=np.float32)\n    _lips_std_Y = np.zeros([_lips_idxs.size], dtype=np.float32)\n\n    for col, ll in enumerate(tqdm( np.transpose(X_train[:,:,_lips_idxs], [2,3,0,1]).reshape([_lips_idxs.size, _n_dims, -1]) )):\n        for dim, l in enumerate(ll):\n            v = l[np.nonzero(l)]\n            if dim == 0: # X\n                _lips_mean_X[col] = v.mean()\n                _lips_std_X[col] = v.std()\n            if dim == 1: # Y\n                _lips_mean_Y[col] = v.mean()\n                _lips_std_Y[col] = v.std()\n\n    lips_mean = np.array([_lips_mean_X, _lips_mean_Y]).T\n    lips_std = np.array([_lips_std_X, _lips_std_X]).T\n    \n    return lips_mean, lips_std\n\nlips_mean, lips_std = get_lips_mean_std()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:15.046418Z","iopub.execute_input":"2023-04-18T16:45:15.046795Z","iopub.status.idle":"2023-04-18T16:45:30.024278Z","shell.execute_reply.started":"2023-04-18T16:45:15.046758Z","shell.execute_reply":"2023-04-18T16:45:30.023183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function for Hands mean values\ndef get_left_right_hand_mean_std():\n    # LEFT HAND\n    _left_hand_mean_X = np.zeros([_left_hand_idxs.size], dtype=np.float32)\n    _left_hand_mean_Y = np.zeros([_left_hand_idxs.size], dtype=np.float32)\n    _left_hand_std_X = np.zeros([_left_hand_idxs.size], dtype=np.float32)\n    _left_hand_std_Y = np.zeros([_left_hand_idxs.size], dtype=np.float32)\n\n    for col, ll in enumerate(tqdm( np.transpose(X_train[:,:,_left_hand_idxs], [2,3,0,1]).reshape([_left_hand_idxs.size, _n_dims, -1]) )):\n        for dim, l in enumerate(ll):\n            v = l[np.nonzero(l)]\n            if dim == 0: # X\n                _left_hand_mean_X[col] = v.mean()\n                _left_hand_std_X[col] = v.std()\n            if dim == 1: # Y\n                _left_hand_mean_Y[col] = v.mean()\n                _left_hand_std_Y[col] = v.std()\n\n    left_hand_mean = np.array([_left_hand_mean_X, _left_hand_mean_Y]).T\n    left_hand_std = np.array([_left_hand_std_X, _left_hand_std_Y]).T\n    \n    return left_hand_mean, left_hand_std\n\nleft_hand_mean, left_hand_std = get_left_right_hand_mean_std()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:30.025933Z","iopub.execute_input":"2023-04-18T16:45:30.026599Z","iopub.status.idle":"2023-04-18T16:45:37.773591Z","shell.execute_reply.started":"2023-04-18T16:45:30.026558Z","shell.execute_reply":"2023-04-18T16:45:37.772356Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function to get Pose mean values\ndef get_pose_mean_std():\n    # POSE\n    _pose_mean_X = np.zeros([_pose_idxs.size], dtype=np.float32)\n    _pose_mean_Y = np.zeros([_pose_idxs.size], dtype=np.float32)\n    _pose_std_X = np.zeros([_pose_idxs.size], dtype=np.float32)\n    _pose_std_Y = np.zeros([_pose_idxs.size], dtype=np.float32)\n\n    for col, ll in enumerate(tqdm( np.transpose(X_train[:,:,_pose_idxs], [2,3,0,1]).reshape([_pose_idxs.size, _n_dims, -1]) )):\n        for dim, l in enumerate(ll):\n            v = l[np.nonzero(l)]\n            if dim == 0: # X\n                _pose_mean_X[col] = v.mean()\n                _pose_std_X[col] = v.std()\n            if dim == 1: # Y\n                _pose_mean_Y[col] = v.mean()\n                _pose_std_Y[col] = v.std()\n\n    pose_mean = np.array([_pose_mean_X, _pose_mean_Y]).T\n    pose_std = np.array([_pose_std_X, _pose_std_Y]).T\n    \n    return pose_mean, pose_std\n\npose_mean, pose_std = get_pose_mean_std()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:37.775252Z","iopub.execute_input":"2023-04-18T16:45:37.775910Z","iopub.status.idle":"2023-04-18T16:45:39.418482Z","shell.execute_reply.started":"2023-04-18T16:45:37.775870Z","shell.execute_reply":"2023-04-18T16:45:39.417357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model Configuration\n\nDefine extra model configuartions and hyperparameters need for modeling.","metadata":{}},{"cell_type":"code","source":"# Epsilon value for layer normalisation\n_LAYER_NORM_EPS = 1e-6\n\n# Dense layer units for landmarks\n_LIPS_UNITS = 384\n_HANDS_UNITS = 384\n_POSE_UNITS = 384\n\n# final embedding and transformer embedding size\n_UNITS = 512\n\n# Transformer\n_NUM_BLOCKS = 2\n_MLP_RATIO = 2\n\n# Dropout\n_EMBEDDING_DROPOUT = 0.00\n_MLP_DROPOUT_RATIO = 0.30\n_CLASSIFIER_DROPOUT_RATIO = 0.10\n\n# Initiailizers\n_INIT_HE_UNIFORM = tf.keras.initializers.he_uniform\n_INIT_GLOROT_UNIFORM = tf.keras.initializers.glorot_uniform\n_INIT_ZEROS = tf.keras.initializers.constant(0.0)\n\n# Activations\n_GELU = tf.keras.activations.gelu\n\nprint(f'UNITS: {_UNITS}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:43.654713Z","iopub.execute_input":"2023-04-18T16:45:43.655422Z","iopub.status.idle":"2023-04-18T16:45:43.663869Z","shell.execute_reply.started":"2023-04-18T16:45:43.655385Z","shell.execute_reply":"2023-04-18T16:45:43.662611Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Transformer\n\nNeed to implement transformer from scratch as TFLite does not support the native TF implementation of MultiHeadAttention.","metadata":{}},{"cell_type":"code","source":"# based on: https://stackoverflow.com/questions/67342988/verifying-the-implementation-of-multihead-attention-in-transformer\n# replaced softmax with softmax layer to support masked softmax\ndef scaled_dot_product(q,k,v, softmax, attention_mask):\n    #calculates Q . K(transpose)\n    qkt = tf.matmul(q,k,transpose_b=True)\n    #caculates scaling factor\n    dk = tf.math.sqrt(tf.cast(q.shape[-1],dtype=tf.float32))\n    scaled_qkt = qkt/dk\n    softmax = softmax(scaled_qkt, mask=attention_mask)\n    \n    z = tf.matmul(softmax,v)\n    #shape: (m,Tx,depth), same shape as q,k,v\n    return z\n\nclass MultiHeadAttention(tf.keras.layers.Layer):\n    def __init__(self,d_model,num_of_heads):\n        super(MultiHeadAttention,self).__init__()\n        self.d_model = d_model\n        self.num_of_heads = num_of_heads\n        self.depth = d_model//num_of_heads\n        self.wq = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wk = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wv = [tf.keras.layers.Dense(self.depth) for i in range(num_of_heads)]\n        self.wo = tf.keras.layers.Dense(d_model)\n        self.softmax = tf.keras.layers.Softmax()\n        \n    def call(self,x, attention_mask):\n        \n        multi_attn = []\n        for i in range(self.num_of_heads):\n            Q = self.wq[i](x)\n            K = self.wk[i](x)\n            V = self.wv[i](x)\n            multi_attn.append(scaled_dot_product(Q,K,V, self.softmax, attention_mask))\n            \n        multi_head = tf.concat(multi_attn,axis=-1)\n        multi_head_attention = self.wo(multi_head)\n        return multi_head_attention","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:45.187058Z","iopub.execute_input":"2023-04-18T16:45:45.187770Z","iopub.status.idle":"2023-04-18T16:45:45.200092Z","shell.execute_reply.started":"2023-04-18T16:45:45.187726Z","shell.execute_reply":"2023-04-18T16:45:45.198853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Full Transformer\nclass Transformer(tf.keras.Model):\n    def __init__(self, num_blocks):\n        super(Transformer, self).__init__(name='transformer')\n        self.num_blocks = num_blocks\n    \n    def build(self, input_shape):\n        self.ln_1s = []\n        self.mhas = []\n        self.ln_2s = []\n        self.mlps = []\n        # Make Transformer Blocks\n        for i in range(self.num_blocks):\n            # Multi Head Attention\n            self.mhas.append(MultiHeadAttention(_UNITS, 8))\n            # Multi Layer Perception\n            self.mlps.append(tf.keras.Sequential([\n                tf.keras.layers.Dense(_UNITS * _MLP_RATIO, activation=_GELU, kernel_initializer=_INIT_GLOROT_UNIFORM),\n                tf.keras.layers.Dropout(_MLP_DROPOUT_RATIO),\n                tf.keras.layers.Dense(_UNITS, kernel_initializer=_INIT_HE_UNIFORM),\n            ]))\n        \n    def call(self, x, attention_mask):\n        # Iterate input over transformer blocks\n        for mha, mlp in zip(self.mhas, self.mlps):\n            x = x + mha(x, attention_mask)\n            x = x + mlp(x)\n    \n        return x","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:45.833255Z","iopub.execute_input":"2023-04-18T16:45:45.834400Z","iopub.status.idle":"2023-04-18T16:45:45.844461Z","shell.execute_reply.started":"2023-04-18T16:45:45.834350Z","shell.execute_reply":"2023-04-18T16:45:45.843397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Landmark Embedding","metadata":{}},{"cell_type":"code","source":"class LandmarkEmbedding(tf.keras.Model):\n    def __init__(self, units, name):\n        super(LandmarkEmbedding, self).__init__(name=f'{name}_embedding')\n        self.units = units\n        \n    def build(self, input_shape):\n        # Embedding for missing landmark in frame, initizlied with zeros\n        self.empty_embedding = self.add_weight(\n            name=f'{self.name}_empty_embedding',\n            shape=[self.units],\n            initializer=_INIT_ZEROS,\n        )\n        # Embedding\n        self.dense = tf.keras.Sequential([\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_1', use_bias=False, kernel_initializer=_INIT_GLOROT_UNIFORM),\n            tf.keras.layers.Activation(_GELU),\n            tf.keras.layers.Dense(self.units, name=f'{self.name}_dense_2', use_bias=False, kernel_initializer=_INIT_HE_UNIFORM),\n        ], name=f'{self.name}_dense')\n\n    def call(self, x):\n        return tf.where(\n                # Checks whether landmark is missing in frame\n                tf.reduce_sum(x, axis=2, keepdims=True) == 0,\n                # If so, the empty embedding is used\n                self.empty_embedding,\n                # Otherwise the landmark data is embedded\n                self.dense(x),\n            )","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:46.941979Z","iopub.execute_input":"2023-04-18T16:45:46.942659Z","iopub.status.idle":"2023-04-18T16:45:46.950845Z","shell.execute_reply.started":"2023-04-18T16:45:46.942621Z","shell.execute_reply":"2023-04-18T16:45:46.949714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Embedding(tf.keras.Model):\n    def __init__(self):\n        super(Embedding, self).__init__()\n        \n    def get_diffs(self, l):\n        S = l.shape[2]\n        other = tf.expand_dims(l, 3)\n        other = tf.repeat(other, S, axis=3)\n        other = tf.transpose(other, [0,1,3,2])\n        diffs = tf.expand_dims(l, 3) - other\n        diffs = tf.reshape(diffs, [-1, _input_size, S*S])\n        return diffs\n\n    def build(self, input_shape):\n        # Positional Embedding, initialized with zeros\n        self.positional_embedding = tf.keras.layers.Embedding(_input_size+1, _UNITS, embeddings_initializer=_INIT_ZEROS)\n        # Embedding layer for Landmarks\n        self.lips_embedding = LandmarkEmbedding(_LIPS_UNITS, 'lips')\n        self.left_hand_embedding = LandmarkEmbedding(_HANDS_UNITS, 'left_hand')\n        self.pose_embedding = LandmarkEmbedding(_POSE_UNITS, 'pose')\n        # Landmark Weights\n        self.landmark_weights = tf.Variable(tf.zeros([3], dtype=tf.float32), name='landmark_weights')\n        # Fully Connected Layers for combined landmarks\n        self.fc = tf.keras.Sequential([\n            tf.keras.layers.Dense(_UNITS, name='fully_connected_1', use_bias=False, kernel_initializer=_INIT_GLOROT_UNIFORM),\n            tf.keras.layers.Activation(_GELU),\n            tf.keras.layers.Dense(_UNITS, name='fully_connected_2', use_bias=False, kernel_initializer=_INIT_HE_UNIFORM),\n        ], name='fc')\n\n\n    def call(self, lips0, left_hand0, pose0, non_empty_frame_idxs, training=False):\n        # Lips\n        lips_embedding = self.lips_embedding(lips0)\n        # Left Hand\n        left_hand_embedding = self.left_hand_embedding(left_hand0)\n        # Pose\n        pose_embedding = self.pose_embedding(pose0)\n        # Merge Embeddings of all landmarks with mean pooling\n        x = tf.stack((\n            lips_embedding, left_hand_embedding, pose_embedding,\n        ), axis=3)\n        x = x * tf.nn.softmax(self.landmark_weights)\n        x = tf.reduce_sum(x, axis=3)\n        # Fully Connected Layers\n        x = self.fc(x)\n        # Add Positional Embedding\n        max_frame_idxs = tf.clip_by_value(\n                tf.reduce_max(non_empty_frame_idxs, axis=1, keepdims=True),\n                1,\n                np.PINF,\n            )\n        normalised_non_empty_frame_idxs = tf.where(\n            tf.math.equal(non_empty_frame_idxs, -1.0),\n            _input_size,\n            tf.cast(\n                non_empty_frame_idxs / max_frame_idxs * _input_size,\n                tf.int32,\n            ),\n        )\n        x = x + self.positional_embedding(normalised_non_empty_frame_idxs)\n        \n        return x","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:47.626324Z","iopub.execute_input":"2023-04-18T16:45:47.626685Z","iopub.status.idle":"2023-04-18T16:45:47.639697Z","shell.execute_reply.started":"2023-04-18T16:45:47.626652Z","shell.execute_reply":"2023-04-18T16:45:47.638430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Augmentation","metadata":{}},{"cell_type":"code","source":"# Not used, adds random X/y translation to input on samples level\nclass Augmentation(tf.keras.layers.Layer):\n    def __init__(self, noise_std):\n        super(Augmentation, self).__init__()\n        self.noise_std = noise_std\n    \n    def add_noise(self, t):\n        B = tf.shape(t)[0]\n        return tf.where(\n            t == 0.0,\n            0.0,\n            t + tf.random.normal([B,1,1,tf.shape(t)[3]], 0, self.noise_std),\n        )\n    \n    def call(self, lips0, left_hand0, pose0, training=False):\n        if training:\n            # Lips\n            lips0 = self.add_noise(lips0)\n            # Left Hand\n            left_hand0 = self.add_noise(left_hand0)\n            # Pose\n            pose0 = self.add_noise(pose0)\n        \n        return lips0, left_hand0, pose0","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:48.484035Z","iopub.execute_input":"2023-04-18T16:45:48.484815Z","iopub.status.idle":"2023-04-18T16:45:48.492934Z","shell.execute_reply.started":"2023-04-18T16:45:48.484765Z","shell.execute_reply":"2023-04-18T16:45:48.491788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Sparse Categorical Crossentropy With Label Smoothing","metadata":{}},{"cell_type":"code","source":"# source:: https://stackoverflow.com/questions/60689185/label-smoothing-for-sparse-categorical-crossentropy\ndef scce_with_ls(y_true, y_pred):\n    # One Hot Encode Sparsely Encoded Target Sign\n    y_true = tf.cast(y_true, tf.int32)\n    y_true = tf.one_hot(y_true, _num_classes, axis=1)\n    y_true = tf.squeeze(y_true, axis=2)\n    # Categorical Crossentropy with native label smoothing support\n    return tf.keras.losses.categorical_crossentropy(y_true, y_pred, label_smoothing=0.25)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:49.964698Z","iopub.execute_input":"2023-04-18T16:45:49.965500Z","iopub.status.idle":"2023-04-18T16:45:49.974045Z","shell.execute_reply.started":"2023-04-18T16:45:49.965463Z","shell.execute_reply":"2023-04-18T16:45:49.972980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model\n\nGet the actual model to train on data.","metadata":{}},{"cell_type":"code","source":"def get_model():\n    # Inputs\n    frames = tf.keras.layers.Input([_input_size, _n_cols, _n_dims], dtype=tf.float32, name='frames')\n    non_empty_frame_idxs = tf.keras.layers.Input([_input_size], dtype=tf.float32, name='non_empty_frame_idxs')\n    # Padding Mask\n    mask0 = tf.cast(tf.math.not_equal(non_empty_frame_idxs, -1), tf.float32)\n    mask0 = tf.expand_dims(mask0, axis=2)\n    # Random Frame Masking\n    mask = tf.where(\n        (tf.random.uniform(tf.shape(mask0)) > 0.25) & tf.math.not_equal(mask0, 0.0),\n        1.0,\n        0.0,\n    )\n    # Correct Samples Which are all masked now...\n    mask = tf.where(\n        tf.math.equal(tf.reduce_sum(mask, axis=[1,2], keepdims=True), 0.0),\n        mask0,\n        mask,\n    )\n    \n    \n    \"\"\"\n        left_hand: 468:489\n        pose: 489:522\n        right_hand: 522:543\n    \"\"\"\n    x = frames\n    x = tf.slice(x, [0,0,0,0], [-1,_input_size, _n_cols, 2])\n    # LIPS\n    lips = tf.slice(x, [0,0,_lips_start,0], [-1,_input_size, 40, 2])\n    lips = tf.where(\n            tf.math.equal(lips, 0.0),\n            0.0,\n            (lips - lips_mean) / lips_std,\n        )\n    # LEFT HAND\n    left_hand = tf.slice(x, [0,0,40,0], [-1,_input_size, 21, 2])\n    left_hand = tf.where(\n            tf.math.equal(left_hand, 0.0),\n            0.0,\n            (left_hand - left_hand_mean) / left_hand_std,\n        )\n    # POSE\n    pose = tf.slice(x, [0,0,61,0], [-1,_input_size, 5, 2])\n    pose = tf.where(\n            tf.math.equal(pose, 0.0),\n            0.0,\n            (pose - pose_mean) / pose_std,\n        )\n    \n    # Flatten\n    lips = tf.reshape(lips, [-1, _input_size, 40*2])\n    left_hand = tf.reshape(left_hand, [-1, _input_size, 21*2])\n    pose = tf.reshape(pose, [-1, _input_size, 5*2])\n        \n    # Embedding\n    x = Embedding()(lips, left_hand, pose, non_empty_frame_idxs)\n    \n    # Encoder Transformer Blocks\n    x = Transformer(_NUM_BLOCKS)(x, mask)\n    \n    # Pooling\n    x = tf.reduce_sum(x * mask, axis=1) / tf.reduce_sum(mask, axis=1)\n    # Classifier Dropout\n    x = tf.keras.layers.Dropout(_CLASSIFIER_DROPOUT_RATIO)(x)\n    # Classification Layer\n    x = tf.keras.layers.Dense(_num_classes, activation=tf.keras.activations.softmax, kernel_initializer=_INIT_GLOROT_UNIFORM)(x)\n    \n    outputs = x\n    \n    # Create Tensorflow Model\n    model = tf.keras.models.Model(inputs=[frames, non_empty_frame_idxs], outputs=outputs)\n    \n    # Sparse Categorical Cross Entropy With Label Smoothing\n    loss = scce_with_ls\n    \n    # Adam Optimizer with weight decay\n    optimizer = tfa.optimizers.AdamW(learning_rate=1e-3, weight_decay=1e-5, clipnorm=1.0)\n    \n    # TopK Metrics\n    metrics = [\n        tf.keras.metrics.SparseCategoricalAccuracy(name='acc'),\n        tf.keras.metrics.SparseTopKCategoricalAccuracy(k=5, name='top_5_acc'),\n        tf.keras.metrics.SparseTopKCategoricalAccuracy(k=10, name='top_10_acc'),\n    ]\n    \n    model.compile(loss=loss, optimizer=optimizer, metrics=metrics)\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:52.010393Z","iopub.execute_input":"2023-04-18T16:45:52.010935Z","iopub.status.idle":"2023-04-18T16:45:52.056092Z","shell.execute_reply.started":"2023-04-18T16:45:52.010896Z","shell.execute_reply":"2023-04-18T16:45:52.054771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()\n\nmodel = get_model()","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:53.044444Z","iopub.execute_input":"2023-04-18T16:45:53.044803Z","iopub.status.idle":"2023-04-18T16:45:55.250535Z","shell.execute_reply.started":"2023-04-18T16:45:53.044763Z","shell.execute_reply":"2023-04-18T16:45:55.249471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plot model summary\nmodel.summary(expand_nested=True)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:45:56.726347Z","iopub.execute_input":"2023-04-18T16:45:56.728921Z","iopub.status.idle":"2023-04-18T16:45:56.873167Z","shell.execute_reply.started":"2023-04-18T16:45:56.728882Z","shell.execute_reply":"2023-04-18T16:45:56.872366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Learning Rate Scheduler","metadata":{}},{"cell_type":"code","source":"def lrfn(current_step, num_warmup_steps, lr_max, num_cycles=0.50, num_training_steps=_n_epochs):\n    \n    if current_step < num_warmup_steps:\n        if WARMUP_METHOD == 'log':\n            return lr_max * 0.10 ** (num_warmup_steps - current_step)\n        else:\n            return lr_max * 2 ** -(num_warmup_steps - current_step)\n    else:\n        progress = float(current_step - num_warmup_steps) / float(max(1, num_training_steps - num_warmup_steps))\n\n        return max(0.0, 0.5 * (1.0 + math.cos(math.pi * float(num_cycles) * 2.0 * progress))) * lr_max","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:01.653401Z","iopub.execute_input":"2023-04-18T16:46:01.654424Z","iopub.status.idle":"2023-04-18T16:46:01.661248Z","shell.execute_reply.started":"2023-04-18T16:46:01.654384Z","shell.execute_reply":"2023-04-18T16:46:01.659857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Learning rate for encoder\nlr_schedule = [lrfn(step, num_warmup_steps=_n_warmup_epochs, lr_max=_lr_max, num_cycles=0.50) for step in range(_n_epochs)]\n\n# Learning Rate Callback\nlr_callback = tf.keras.callbacks.LearningRateScheduler(lambda step: lr_schedule[step], verbose=1)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:04.405383Z","iopub.execute_input":"2023-04-18T16:46:04.405746Z","iopub.status.idle":"2023-04-18T16:46:04.413986Z","shell.execute_reply.started":"2023-04-18T16:46:04.405713Z","shell.execute_reply":"2023-04-18T16:46:04.412802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Weight Decay Callback","metadata":{}},{"cell_type":"code","source":"# Custom callback to update weight decay with learning rate\nclass WeightDecayCallback(tf.keras.callbacks.Callback):\n    def __init__(self, wd_ratio=_wd_ratio):\n        self.step_counter = 0\n        self.wd_ratio = wd_ratio\n    \n    def on_epoch_begin(self, epoch, logs=None):\n        model.optimizer.weight_decay = model.optimizer.learning_rate * self.wd_ratio\n        print(f'learning rate: {model.optimizer.learning_rate.numpy():.2e}, weight decay: {model.optimizer.weight_decay.numpy():.2e}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:07.030787Z","iopub.execute_input":"2023-04-18T16:46:07.031408Z","iopub.status.idle":"2023-04-18T16:46:07.037479Z","shell.execute_reply.started":"2023-04-18T16:46:07.031367Z","shell.execute_reply":"2023-04-18T16:46:07.036366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Performance Benchmark","metadata":{}},{"cell_type":"code","source":"%%timeit -n 100\nif _train_model:\n    # Verify model prediction is <<<100ms\n    model.predict_on_batch({ 'frames': X_train[:1], 'non_empty_frame_idxs': non_empty_frame_idxs_train[:1] })\n    pass","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:08.729926Z","iopub.execute_input":"2023-04-18T16:46:08.730948Z","iopub.status.idle":"2023-04-18T16:46:18.392030Z","shell.execute_reply.started":"2023-04-18T16:46:08.730896Z","shell.execute_reply":"2023-04-18T16:46:18.390974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate Initialzied Model","metadata":{}},{"cell_type":"code","source":"# Sanity Check\nif _train_model and _use_val:\n    _ = model.evaluate(*validation_data, verbose=2)","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:32.111156Z","iopub.execute_input":"2023-04-18T16:46:32.111854Z","iopub.status.idle":"2023-04-18T16:46:32.116910Z","shell.execute_reply.started":"2023-04-18T16:46:32.111815Z","shell.execute_reply":"2023-04-18T16:46:32.115563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train Model\n\nStart training the model","metadata":{}},{"cell_type":"code","source":"if _train_model:\n    # Clear all models in GPU\n    tf.keras.backend.clear_session()\n\n    # Get new fresh model\n    model = get_model()\n    \n    # Sanity Check\n    model.summary()\n\n    # Actual Training\n    history = model.fit(\n            x=get_train_batch_all_signs(X_train, y_train, non_empty_frame_idxs_train),\n            steps_per_epoch=len(X_train) // (_num_classes * _batch_all_signs),\n            epochs=_n_epochs,\n            # Only used for validation data since training data is a generator\n            batch_size=_batch_size,\n            validation_data=validation_data,\n            callbacks=[\n                lr_callback,\n                WeightDecayCallback(),\n            ],\n            verbose = _verbose,\n        )","metadata":{"execution":{"iopub.status.busy":"2023-04-18T16:46:40.855821Z","iopub.execute_input":"2023-04-18T16:46:40.856436Z","iopub.status.idle":"2023-04-18T18:12:16.261781Z","shell.execute_reply.started":"2023-04-18T16:46:40.856398Z","shell.execute_reply":"2023-04-18T18:12:16.260597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Save Model Weights\nmodel.save_weights('model.h5')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:12:16.264579Z","iopub.execute_input":"2023-04-18T18:12:16.264998Z","iopub.status.idle":"2023-04-18T18:12:16.430653Z","shell.execute_reply.started":"2023-04-18T18:12:16.264956Z","shell.execute_reply":"2023-04-18T18:12:16.429590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if _use_val:\n    # Validation Predictions\n    y_val_pred = model.predict({ 'frames': X_val, 'non_empty_frame_idxs': non_empty_frame_idx_val }, verbose=2).argmax(axis=1)\n    # Label\n    labels = [ord_2_sign_df.get(i).replace(' ', '_') for i in range(_num_classes)]","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:12:16.432093Z","iopub.execute_input":"2023-04-18T18:12:16.432498Z","iopub.status.idle":"2023-04-18T18:12:16.441214Z","shell.execute_reply.started":"2023-04-18T18:12:16.432458Z","shell.execute_reply":"2023-04-18T18:12:16.440156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Landmark Attention Weights","metadata":{}},{"cell_type":"code","source":"# Landmark Weights\nfor w in model.get_layer('embedding').weights:\n    if 'landmark_weights' in w.name:\n        weights = scipy.special.softmax(w)\n\nlandmarks = ['lips_embedding', 'left_hand_embedding', 'pose_embedding']\n\nfor w, lm in zip(weights, landmarks):\n    print(f'{lm} weight: {(w*100):.1f}%')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:12:16.444187Z","iopub.execute_input":"2023-04-18T18:12:16.444823Z","iopub.status.idle":"2023-04-18T18:12:16.456139Z","shell.execute_reply.started":"2023-04-18T18:12:16.444785Z","shell.execute_reply":"2023-04-18T18:12:16.454960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission\n\nSubmission code loosley based on [this notebook](https://www.kaggle.com/code/dschettler8845/gislr-learn-eda-baseline#baseline) by [Darien Schettler\n](https://www.kaggle.com/dschettler8845)","metadata":{}},{"cell_type":"code","source":"# TFLite model for submission\nclass TFLiteModel(tf.Module):\n    def __init__(self, model):\n        super(TFLiteModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.model = model\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, _n_rows, _n_dims], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Preprocess Data\n        x, non_empty_frame_idxs = self.preprocess_layer(inputs)\n        # Add Batch Dimension\n        x = tf.expand_dims(x, axis=0)\n        non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=0)\n        # Make Prediction\n        outputs = self.model({ 'frames': x, 'non_empty_frame_idxs': non_empty_frame_idxs })\n        # Squeeze Output 1x250 -> 250\n        outputs = tf.squeeze(outputs, axis=0)\n\n        # Return a dictionary with the output tensor\n        return {'outputs': outputs}\n\n# Define TF Lite Model\ntflite_keras_model = TFLiteModel(model)\n\n# Sanity Check\ndemo_raw_data = load_relevant_data_subset(train_df['file_path'].values[5])\nprint(f'demo_raw_data shape: {demo_raw_data.shape}, dtype: {demo_raw_data.dtype}')\ndemo_output = tflite_keras_model(demo_raw_data)[\"outputs\"]\nprint(f'demo_output shape: {demo_output.shape}, dtype: {demo_output.dtype}')\ndemo_prediction = demo_output.numpy().argmax()\nprint(f'demo_prediction: {demo_prediction}, correct: {train_df.iloc[0][\"sign_ord\"]}')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:36:59.605097Z","iopub.execute_input":"2023-04-18T18:36:59.605695Z","iopub.status.idle":"2023-04-18T18:37:01.961485Z","shell.execute_reply.started":"2023-04-18T18:36:59.605658Z","shell.execute_reply":"2023-04-18T18:37:01.960274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create Model Converter\nkeras_model_converter = tf.lite.TFLiteConverter.from_keras_model(tflite_keras_model)\n# Convert Model\ntflite_model = keras_model_converter.convert()\n# Write Model\nwith open('/kaggle/working/model.tflite', 'wb') as f:\n    f.write(tflite_model)\n    \n# Zip Model\n!zip submission.zip /kaggle/working/model.tflite","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:37:08.655028Z","iopub.execute_input":"2023-04-18T18:37:08.655528Z","iopub.status.idle":"2023-04-18T18:37:48.848644Z","shell.execute_reply.started":"2023-04-18T18:37:08.655490Z","shell.execute_reply":"2023-04-18T18:37:48.847212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Verify TFLite model can be loaded and used for prediction\n!pip install tflite-runtime\nimport tflite_runtime.interpreter as tflite\n\ninterpreter = tflite.Interpreter(\"/kaggle/working/model.tflite\")\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\n\noutput = prediction_fn(inputs=demo_raw_data)\nsign = output['outputs'].argmax()\n\nprint(\"PRED : \", ord_2_sign_df.get(sign), f'[{sign}]')\nprint(\"TRUE : \", train_df.sign.values[0], f'[{train_df.sign_ord.values[0]}]')","metadata":{"execution":{"iopub.status.busy":"2023-04-18T18:38:30.260661Z","iopub.execute_input":"2023-04-18T18:38:30.261045Z","iopub.status.idle":"2023-04-18T18:38:40.294576Z","shell.execute_reply.started":"2023-04-18T18:38:30.261011Z","shell.execute_reply":"2023-04-18T18:38:40.293079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}