{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nIsolated sign language recognition is an important application of computer vision and machine learning. It allows individuals with hearing impairments to communicate more easily and effectively. PyTorch is a popular deep learning framework that can be used for developing sign language recognition models. Recently, there has been growing interest in deploying machine learning models on mobile and embedded devices. To achieve this, the models need to be converted to a format that can be used on these devices. TensorFlow Lite (TFLite) is a framework that enables deployment of machine learning models on mobile and embedded devices. In this project, we will develop an isolated sign language recognition model using PyTorch and then convert it to TFLite format for deployment on mobile and embedded devices.","metadata":{}},{"cell_type":"markdown","source":"# Load Data\n\n**train_landmark_files/[participant_id]/[sequence_id]**.parquet The landmark data. The landmarks were extracted from raw videos with the MediaPipe holistic model. Not all of the frames necessarily had visible hands or hands that could be detected by the model.\n\nLandmark data should not be used to identify or re-identify an individual. Landmark data is not intended to enable any form of identity recognition or store any unique biometric identification\n* frame - The frame number in the raw video.\n* row_id - A unique identifier for the row.\n* type - The type of landmark. One of ['face', 'left_hand', 'pose', 'right_hand'].\n* landmark_index - The landmark index number. Details of the hand landmark locations can be found here.\n* [x/y/z] - The normalized spatial coordinates of the landmark. These are the only columns that will be provided to your submitted model for inference. The MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values.\n","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nimport tqdm\nimport random\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nLANDMARK_FILES_DIR = \"/kaggle/input/asl-signs/train_landmark_files\"\nTRAIN_FILE = \"/kaggle/input/asl-signs/train.csv\"\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-20T17:32:14.612828Z","iopub.execute_input":"2023-04-20T17:32:14.613456Z","iopub.status.idle":"2023-04-20T17:32:14.685583Z","shell.execute_reply.started":"2023-04-20T17:32:14.613392Z","shell.execute_reply":"2023-04-20T17:32:14.683599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"participants = os.listdir(LANDMARK_FILES_DIR)\nprint(f\"Total number of participants = {len(participants)}\")\nprint(f\"Average number of sequences per participant = {len(glob.glob(LANDMARK_FILES_DIR + '/*/*.parquet'))/len(participants)}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:14.690012Z","iopub.execute_input":"2023-04-20T17:32:14.693578Z","iopub.status.idle":"2023-04-20T17:32:19.795733Z","shell.execute_reply.started":"2023-04-20T17:32:14.693537Z","shell.execute_reply":"2023-04-20T17:32:19.794440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore Dataset\n\ntrain.csv\n\n* path - The path to the landmark file.\n* participant_id - A unique identifier for the data contributor.\n* sequence_id - A unique identifier for the landmark sequence.\n* sign - The label for the landmark sequence.\n","metadata":{}},{"cell_type":"code","source":"data_dir = '/kaggle/input/asl-signs/'\ntrain_df = pd.read_csv(TRAIN_FILE)\n\ntrain_df[\"path\"] = data_dir + \"/\" + train_df[\"path\"]\ndisplay(train_df.head(2)), len(train_df)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:19.800440Z","iopub.execute_input":"2023-04-20T17:32:19.801253Z","iopub.status.idle":"2023-04-20T17:32:20.155515Z","shell.execute_reply.started":"2023-04-20T17:32:19.801212Z","shell.execute_reply":"2023-04-20T17:32:20.134728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sign to prediction index","metadata":{}},{"cell_type":"code","source":"import json\ndef read_json(path):\n    with open(path, \"r\") as file:\n        json_data = json.load(file)\n    return json_data","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:20.157482Z","iopub.execute_input":"2023-04-20T17:32:20.159962Z","iopub.status.idle":"2023-04-20T17:32:20.179625Z","shell.execute_reply.started":"2023-04-20T17:32:20.159873Z","shell.execute_reply":"2023-04-20T17:32:20.171678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s2p_map = read_json(os.path.join(data_dir, \"sign_to_prediction_index_map.json\"))\np2s_map = {v: k for k, v in s2p_map.items()}\n\nencoder = lambda x: s2p_map.get(x)\ndecoder = lambda x: p2s_map.get(x)\n\ntrain_df[\"label\"] = train_df[\"sign\"].map(encoder)\nprint(f\"shape = {train_df.shape}\")\n\ntrain_df.head(2)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:20.194122Z","iopub.execute_input":"2023-04-20T17:32:20.198488Z","iopub.status.idle":"2023-04-20T17:32:20.333890Z","shell.execute_reply.started":"2023-04-20T17:32:20.198151Z","shell.execute_reply":"2023-04-20T17:32:20.332774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualize Top and Button count sign","metadata":{}},{"cell_type":"code","source":"train_df['sign'].value_counts().head(50).sort_values(ascending=True).plot(\n    kind='barh', figsize=(10, 8), title='Top 50 Sign'\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:20.336317Z","iopub.execute_input":"2023-04-20T17:32:20.336722Z","iopub.status.idle":"2023-04-20T17:32:22.118549Z","shell.execute_reply.started":"2023-04-20T17:32:20.336683Z","shell.execute_reply":"2023-04-20T17:32:22.116940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['sign'].value_counts().tail(50).sort_values(ascending=False).plot(\n    kind='barh', figsize=(10, 8), title='Buttom 50 Sign', color='salmon'\n)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:22.122287Z","iopub.execute_input":"2023-04-20T17:32:22.122616Z","iopub.status.idle":"2023-04-20T17:32:23.210913Z","shell.execute_reply.started":"2023-04-20T17:32:22.122584Z","shell.execute_reply":"2023-04-20T17:32:23.209250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sequence Landmarks Data¶\nLets have a look at the dataframe of one sample sequence!\n\n","metadata":{}},{"cell_type":"code","source":"example_fn = train_df.query('sign == \"look\"')[\"path\"].values[0]\n\nexample_df = pd.read_parquet(example_fn)\nprint(f\"Sample shape = {example_df.shape}\")\nexample_df.sample(10)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:23.212123Z","iopub.execute_input":"2023-04-20T17:32:23.212469Z","iopub.status.idle":"2023-04-20T17:32:23.419769Z","shell.execute_reply.started":"2023-04-20T17:32:23.212437Z","shell.execute_reply":"2023-04-20T17:32:23.418269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"All different types of landmark = {example_df.type.unique()}\")","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:23.422088Z","iopub.execute_input":"2023-04-20T17:32:23.422711Z","iopub.status.idle":"2023-04-20T17:32:23.432660Z","shell.execute_reply.started":"2023-04-20T17:32:23.422664Z","shell.execute_reply":"2023-04-20T17:32:23.430885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_left_hand = example_df[example_df.type == \"left_hand\"]\nsample_right_hand = example_df[example_df.type == \"right_hand\"]\n\nprint(f\"Percentage of nulls in Left Hand data = {100*np.mean(sample_left_hand['x'].isnull()):.02f} %\")\nprint(f\"Percentage of nulls in Right Hand data = {100*np.mean(sample_right_hand['x'].isnull()):.02f} %\")","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:23.436646Z","iopub.execute_input":"2023-04-20T17:32:23.437598Z","iopub.status.idle":"2023-04-20T17:32:23.452668Z","shell.execute_reply.started":"2023-04-20T17:32:23.437554Z","shell.execute_reply":"2023-04-20T17:32:23.449910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_right_hand.describe()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:23.456780Z","iopub.execute_input":"2023-04-20T17:32:23.457713Z","iopub.status.idle":"2023-04-20T17:32:23.494780Z","shell.execute_reply.started":"2023-04-20T17:32:23.457668Z","shell.execute_reply":"2023-04-20T17:32:23.493447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualizing\nVisualizing the hand sequences will enable us to understand the tabular data in a better way. You can see there are lots of rotations and movements going on in these frames. With the coordinates all jumbled up, depth (z axis) surely comes into play although organizers provided some caution regarding it.","metadata":{}},{"cell_type":"code","source":"edges = [(0,1),(1,2),(2,3),(3,4),(0,5),(0,17),(5,6),(6,7),(7,8),(5,9),(9,10),(10,11),(11,12),\n         (9,13),(13,14),(14,15),(15,16),(13,17),(17,18),(18,19),(19,20)]\n\ndef plot_frame(df, frame_id, ax):\n    df = df[df.frame == frame_id].sort_values(['landmark_index'])\n    x = list(df.x)\n    y = list(df.y)\n    \n    ax.scatter(df.x, df.y, color='dodgerblue')\n    for i in range(len(x)):\n        ax.text(x[i], y[i], str(i))\n        \n    for edge in edges:\n        ax.plot([x[edge[0]], x[edge[1]]], [y[edge[0]], y[edge[1]]], color='salmon')\n        ax.set_xlabel(f\"Frame no. {frame_id}\")\n        ax.set_xticks([])\n        ax.set_yticks([])\n        ax.set_xticklabels([])\n        ax.set_yticklabels([])\n\n    \ndef plot_frame_seq(df, frame_range, n_frames):\n    frames = np.linspace(frame_range[0],frame_range[1],n_frames, dtype = int, endpoint=True)\n    fig, ax = plt.subplots(n_frames, 1, figsize=(5,25))\n    for i in range(n_frames):\n        plot_frame(df, frames[i], ax[i])\n        \n    plt.show()\n\n    \nplot_frame_seq(sample_right_hand, (28,41), 5)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:23.499031Z","iopub.execute_input":"2023-04-20T17:32:23.499407Z","iopub.status.idle":"2023-04-20T17:32:24.555327Z","shell.execute_reply.started":"2023-04-20T17:32:23.499365Z","shell.execute_reply":"2023-04-20T17:32:24.554214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Dataset","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport numpy as np\n\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:24.557032Z","iopub.execute_input":"2023-04-20T17:32:24.557703Z","iopub.status.idle":"2023-04-20T17:32:29.148021Z","shell.execute_reply.started":"2023-04-20T17:32:24.557665Z","shell.execute_reply":"2023-04-20T17:32:29.146830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport numpy as np\n\nclass FeatureGen(nn.Module):\n    def __init__(self):\n        super(FeatureGen, self).__init__()\n\n    def forward(self, x):\n        x = torch.from_numpy(x)\n        x = torch.where(torch.isnan(x), torch.zeros_like(x), x)\n        x = torch.mean(x, axis=0)\n        return x\n    \nfeature_converter = FeatureGen()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.155783Z","iopub.execute_input":"2023-04-20T17:32:29.156881Z","iopub.status.idle":"2023-04-20T17:32:29.164437Z","shell.execute_reply.started":"2023-04-20T17:32:29.156835Z","shell.execute_reply":"2023-04-20T17:32:29.163342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_length = (len(train_df))\ndata_lenght_experiment = int(len(train_df)/5)\n\nprint(\"Lenght of data for modeling :\", data_length)\nprint(f\"Percentage of total data {data_lenght_experiment/data_length*100:.1f}%\")\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.166462Z","iopub.execute_input":"2023-04-20T17:32:29.167250Z","iopub.status.idle":"2023-04-20T17:32:29.180222Z","shell.execute_reply.started":"2023-04-20T17:32:29.167208Z","shell.execute_reply":"2023-04-20T17:32:29.178792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_relevant_data_subset(pq_path):\n    data_columns = [\"x\", \"y\", \"z\"]\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, len(data_columns))\n    return data.astype(np.float32)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.181791Z","iopub.execute_input":"2023-04-20T17:32:29.182282Z","iopub.status.idle":"2023-04-20T17:32:29.191600Z","shell.execute_reply.started":"2023-04-20T17:32:29.182244Z","shell.execute_reply":"2023-04-20T17:32:29.190631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def convert_row(row):\n    x = load_relevant_data_subset(os.path.join(\"/kaggle/input/asl-signs\", row.path))\n    x = feature_converter(x)\n    return x, row.label","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.193598Z","iopub.execute_input":"2023-04-20T17:32:29.194018Z","iopub.status.idle":"2023-04-20T17:32:29.209611Z","shell.execute_reply.started":"2023-04-20T17:32:29.193981Z","shell.execute_reply":"2023-04-20T17:32:29.208516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\n\nROWS_PER_FRAME = 543\n\ndef convert_and_save_data():\n    np_features = np.zeros((data_lenght_experiment, ROWS_PER_FRAME, 3))\n    np_labels = np.zeros(data_lenght_experiment)\n\n    print(f\"Total data to processe : {data_lenght_experiment}\")\n    for index, row in tqdm(train_df.iterrows()):\n        if index > data_lenght_experiment - 1:\n            break\n\n        data = load_relevant_data_subset(row.path)\n        feature, label = convert_row(row)\n        np_features[index, :, :] = feature.to(torch.double)\n        np_labels[index] = label\n\n    np.save(\"features.npy\", np_features)\n    np.save(\"labels.npy\", np_labels)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.211111Z","iopub.execute_input":"2023-04-20T17:32:29.211974Z","iopub.status.idle":"2023-04-20T17:32:29.222431Z","shell.execute_reply.started":"2023-04-20T17:32:29.211930Z","shell.execute_reply":"2023-04-20T17:32:29.221284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!rm -rf /kaggle/working/*.npy","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:29.224112Z","iopub.execute_input":"2023-04-20T17:32:29.224719Z","iopub.status.idle":"2023-04-20T17:32:30.275281Z","shell.execute_reply.started":"2023-04-20T17:32:29.224679Z","shell.execute_reply":"2023-04-20T17:32:30.273706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"try:\n    features = torch.from_numpy(np.load(\"/kaggle/working/features.npy\")).float()\n    labels = torch.from_numpy(np.load(\"/kaggle/working/labels.npy\")).long()\nexcept:\n    convert_and_save_data()\nfinally:\n    features = torch.from_numpy(np.load(\"/kaggle/working/features.npy\")).float()\n    labels = torch.from_numpy(np.load(\"/kaggle/working/labels.npy\")).long()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:32:30.277360Z","iopub.execute_input":"2023-04-20T17:32:30.277804Z","iopub.status.idle":"2023-04-20T17:41:00.038117Z","shell.execute_reply.started":"2023-04-20T17:32:30.277759Z","shell.execute_reply":"2023-04-20T17:41:00.036947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Loader","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(\n    features, labels, test_size=0.2, stratify=labels, random_state=42\n)\n\nprint(\"Training Shape\", X_train.shape, y_train.shape)\nprint(\"Testing Shape\", X_val.shape, y_val.shape) \n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.039962Z","iopub.execute_input":"2023-04-20T17:41:00.040369Z","iopub.status.idle":"2023-04-20T17:41:00.800307Z","shell.execute_reply.started":"2023-04-20T17:41:00.040325Z","shell.execute_reply":"2023-04-20T17:41:00.799032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Dataset(torch.utils.data.Dataset):\n    def __init__(self, X, y):\n        self.X = X\n        self.y = y\n\n    def __len__(self):\n        return len(self.y)\n\n    def __getitem__(self, i):\n        return self.X[i], self.y[i]","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.802351Z","iopub.execute_input":"2023-04-20T17:41:00.803177Z","iopub.status.idle":"2023-04-20T17:41:00.810482Z","shell.execute_reply.started":"2023-04-20T17:41:00.803132Z","shell.execute_reply":"2023-04-20T17:41:00.809177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_dataset = Dataset(X_train, y_train)\nval_dataset = Dataset(X_val, y_val)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.812358Z","iopub.execute_input":"2023-04-20T17:41:00.812995Z","iopub.status.idle":"2023-04-20T17:41:00.821223Z","shell.execute_reply.started":"2023-04-20T17:41:00.812954Z","shell.execute_reply":"2023-04-20T17:41:00.820143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_workers = 1\nbatch_size = 128\n\ntrain_dataloader = torch.utils.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=True,num_workers=num_workers, pin_memory=True, drop_last=True)\nval_dataloader = torch.utils.data.DataLoader(train_dataset, batch_size=batch_size, shuffle=False,num_workers=num_workers, pin_memory=True, drop_last=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.822701Z","iopub.execute_input":"2023-04-20T17:41:00.824041Z","iopub.status.idle":"2023-04-20T17:41:00.838904Z","shell.execute_reply.started":"2023-04-20T17:41:00.824011Z","shell.execute_reply":"2023-04-20T17:41:00.837523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nfrom torch.autograd import Variable ","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.840632Z","iopub.execute_input":"2023-04-20T17:41:00.841350Z","iopub.status.idle":"2023-04-20T17:41:00.850277Z","shell.execute_reply.started":"2023-04-20T17:41:00.841165Z","shell.execute_reply":"2023-04-20T17:41:00.848955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        self.layer1 = nn.Linear(3, 128)\n        self.layer2 = nn.Linear(128, 64)\n        self.layer3 = nn.Linear(64, 32)\n        self.layer4 = nn.Linear(32, 16)\n        self.layer5 = nn.Linear(16 * 543, 250)\n        self.relu = nn.ReLU()\n        self.flatten = nn.Flatten()\n\n    def forward(self, x):\n        x = self.relu(self.layer1(x))\n        x = self.relu(self.layer2(x))\n        x = self.relu(self.layer3(x))\n        x = self.relu(self.layer4(x))\n        x = self.flatten(x)\n        x = self.layer5(x)\n        return x\n\nmodel = Model()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.852125Z","iopub.execute_input":"2023-04-20T17:41:00.852655Z","iopub.status.idle":"2023-04-20T17:41:00.907381Z","shell.execute_reply.started":"2023-04-20T17:41:00.852616Z","shell.execute_reply":"2023-04-20T17:41:00.906378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.to(device)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:00.908722Z","iopub.execute_input":"2023-04-20T17:41:00.909076Z","iopub.status.idle":"2023-04-20T17:41:05.342428Z","shell.execute_reply.started":"2023-04-20T17:41:00.909039Z","shell.execute_reply":"2023-04-20T17:41:05.341208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learning_rate = 0.0001 #0.001 lr\n\ncriterion = torch.nn.CrossEntropyLoss()    # mean-squared error for regression\noptimizer = torch.optim.Adam(model.parameters(), lr=learning_rate) ","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:05.343834Z","iopub.execute_input":"2023-04-20T17:41:05.344692Z","iopub.status.idle":"2023-04-20T17:41:05.351756Z","shell.execute_reply.started":"2023-04-20T17:41:05.344648Z","shell.execute_reply":"2023-04-20T17:41:05.350384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Setup Logger","metadata":{}},{"cell_type":"code","source":"from torch.utils.tensorboard import SummaryWriter\n\n\nPATH_LOG = '/kaggle/working/asl_sign'\nwriter = SummaryWriter(PATH_LOG)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:05.353396Z","iopub.execute_input":"2023-04-20T17:41:05.354095Z","iopub.status.idle":"2023-04-20T17:41:22.138079Z","shell.execute_reply.started":"2023-04-20T17:41:05.353951Z","shell.execute_reply":"2023-04-20T17:41:22.136808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train","metadata":{}},{"cell_type":"code","source":"num_epochs = 500\nfor epoch in range(num_epochs):\n    train_loss, train_correct, train_n, val_loss, val_correct, val_n = 0,0,0,0,0,0\n    model.train()\n    \n    for ibatch, (X, y) in enumerate(train_dataloader):\n        X, y = X.to(device), y.to(device)\n        optimizer.zero_grad()\n        y_pred = model(X)\n        loss = criterion(y_pred, y)\n        \n        train_n += y.size(0)\n        train_loss += loss.item()\n        train_correct += (y_pred.argmax(1) == y).type(torch.float).sum().item()\n\n        loss.backward()\n        optimizer.step()\n        \n    train_loss /= ibatch   \n    train_correct /= train_n\n    \n    model.eval()\n    \n    for ibatch, (X, y) in enumerate(val_dataloader):\n        X, y = X.to(device), y.to(device)\n        with torch.no_grad():\n            y_pred = model(X)\n            loss = criterion(y_pred, y)\n            val_n += y.size(0)\n            val_loss += loss.item()\n            val_correct += (y_pred.argmax(1) == y).type(torch.float).sum().item()\n\n    val_loss /= ibatch   \n    val_correct /= val_n\n    \n    # write log\n    writer.add_scalar(\n        'loss',\n        train_loss,\n        epoch + 1\n    )\n    \n    writer.add_scalar(\n        'accuracy',\n        train_correct,\n        epoch + 1\n    )\n    \n    writer.add_scalar(\n        'val_loss',\n        val_loss,\n        epoch + 1\n    )\n    writer.add_scalar(\n        'val_accuracy',\n        val_correct,\n        epoch + 1\n    )\n        \n    print('Epoch %d/%d loss:%.4f accuracy:%.4f val_loss:%.4f val_accuracy:%.4f' %(epoch + 1, num_epochs, train_loss, train_correct, val_loss, val_correct))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:41:22.140018Z","iopub.execute_input":"2023-04-20T17:41:22.141385Z","iopub.status.idle":"2023-04-20T17:56:12.187533Z","shell.execute_reply.started":"2023-04-20T17:41:22.141339Z","shell.execute_reply":"2023-04-20T17:56:12.185914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Save Inference Model","metadata":{}},{"cell_type":"code","source":"torch.save(model.state_dict(), 'model.pth')","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:12.189678Z","iopub.execute_input":"2023-04-20T17:56:12.190012Z","iopub.status.idle":"2023-04-20T17:56:12.237765Z","shell.execute_reply.started":"2023-04-20T17:56:12.189980Z","shell.execute_reply":"2023-04-20T17:56:12.236472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from torchinfo import summary\n\nsummary(model=model, input_size=(128, 543, 3))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:12.239725Z","iopub.execute_input":"2023-04-20T17:56:12.240174Z","iopub.status.idle":"2023-04-20T17:56:12.295550Z","shell.execute_reply.started":"2023-04-20T17:56:12.240131Z","shell.execute_reply":"2023-04-20T17:56:12.294217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!ls /kaggle/working/asl_sign","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:12.297147Z","iopub.execute_input":"2023-04-20T17:56:12.297538Z","iopub.status.idle":"2023-04-20T17:56:13.443766Z","shell.execute_reply.started":"2023-04-20T17:56:12.297507Z","shell.execute_reply":"2023-04-20T17:56:13.442309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualize result training","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom matplotlib import pyplot as plt\nimport tensorflow as tf\nimport struct\nimport glob\nfrom tensorflow.python.summary.summary_iterator import summary_iterator\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:13.446142Z","iopub.execute_input":"2023-04-20T17:56:13.447252Z","iopub.status.idle":"2023-04-20T17:56:13.454061Z","shell.execute_reply.started":"2023-04-20T17:56:13.447201Z","shell.execute_reply":"2023-04-20T17:56:13.452700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #Load losses\n\n# log_event = glob.glob(f'{PATH_LOG}/*')[0]\n\n# train_loss, train_correct, train_n, val_loss, val_correct, val_n = [],[],[],[],[],[]\n# steps=[]\n# for e in summary_iterator(log_event):\n#     for v in e.summary.value:\n#         if v.tag == 'accuracy':       \n#             train_correct.append(v.simple_value)\n#         if v.tag == 'val_accuracy':       \n#             val_correct.append(v.simple_value)\n#         if v.tag == 'loss':\n#             train_loss.append(v.simple_value)\n#         if v.tag == 'val_loss':\n#             val_loss.append(v.simple_value)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:13.455908Z","iopub.execute_input":"2023-04-20T17:56:13.456535Z","iopub.status.idle":"2023-04-20T17:56:13.545335Z","shell.execute_reply.started":"2023-04-20T17:56:13.456496Z","shell.execute_reply":"2023-04-20T17:56:13.544160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# plt.rcParams[\"figure.figsize\"] = [7.50, 3.50]\n# plt.rcParams[\"figure.autolayout\"] = True\n\n# ax1 = plt.subplot()\n# l1, = ax1.plot(val_correct, color='red')\n# ax2 = ax1.twinx()\n# l2, = ax2.plot(train_correct, color='blue')\n\n# plt.legend([l1, l2], [\"Accuracy\", \"Val Accuracy\"])\n\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:13.546891Z","iopub.execute_input":"2023-04-20T17:56:13.547202Z","iopub.status.idle":"2023-04-20T17:56:13.920395Z","shell.execute_reply.started":"2023-04-20T17:56:13.547156Z","shell.execute_reply":"2023-04-20T17:56:13.919201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# plt.rcParams[\"figure.figsize\"] = [7.50, 3.50]\n# plt.rcParams[\"figure.autolayout\"] = True\n\n# ax1 = plt.subplot()\n# l1, = ax1.plot(val_loss, color='red')\n# ax2 = ax1.twinx()\n# l2, = ax2.plot(train_loss, color='blue')\n\n# plt.legend([l1, l2], [\"Val Loss\", \"Train Loss\"])\n\n# plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:13.922055Z","iopub.execute_input":"2023-04-20T17:56:13.922776Z","iopub.status.idle":"2023-04-20T17:56:14.302643Z","shell.execute_reply.started":"2023-04-20T17:56:13.922734Z","shell.execute_reply":"2023-04-20T17:56:14.301243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Export Model TfLite","metadata":{}},{"cell_type":"markdown","source":"## Inference Model","metadata":{}},{"cell_type":"code","source":"class Model_infe(Model):\n    \n    def __init__(self):\n        super().__init__()\n        self.softmax = nn.Softmax()\n        \n    def forward(self, x):\n        x = torch.where(torch.isnan(x), torch.tensor(0.0, dtype=torch.float32).to(device), x)\n        x = torch.mean(x, dim=0, keepdim=False)\n        x = self.relu(self.layer1(x))\n\n        x = self.relu(self.layer2(x))\n        x = self.relu(self.layer3(x))\n        x = self.relu(self.layer4(x))\n        x = self.flatten(x)\n        x = self.layer5(x)\n        return self.softmax(x)\n    \nmodel_infe = Model_infe()","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.305471Z","iopub.execute_input":"2023-04-20T17:56:14.306349Z","iopub.status.idle":"2023-04-20T17:56:14.338588Z","shell.execute_reply.started":"2023-04-20T17:56:14.306302Z","shell.execute_reply":"2023-04-20T17:56:14.337546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_infe.load_state_dict(torch.load('/kaggle/working/model.pth'), strict=False)\nmodel_infe = model_infe.to(device)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.340258Z","iopub.execute_input":"2023-04-20T17:56:14.340654Z","iopub.status.idle":"2023-04-20T17:56:14.369460Z","shell.execute_reply.started":"2023-04-20T17:56:14.340615Z","shell.execute_reply":"2023-04-20T17:56:14.368412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Export ONNX","metadata":{}},{"cell_type":"code","source":"input_size = (1, 543, 3)\nbatch_size = 1\nsaved_onnx = 'model.onnx'\n\ndummy_input = torch.rand((batch_size, *input_size)).to(device)\n\nmodel_infe.eval()\npreds = model_infe(dummy_input)\npreds","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.371089Z","iopub.execute_input":"2023-04-20T17:56:14.371703Z","iopub.status.idle":"2023-04-20T17:56:14.465452Z","shell.execute_reply.started":"2023-04-20T17:56:14.371663Z","shell.execute_reply":"2023-04-20T17:56:14.464258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.onnx.export(\n    model_infe,\n    dummy_input, \n    saved_onnx,\n    verbose=False,\n    input_names=['inputs'],\n    output_names=['outputs'],\n#     export_params=True,\n    opset_version=11\n)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.473875Z","iopub.execute_input":"2023-04-20T17:56:14.474965Z","iopub.status.idle":"2023-04-20T17:56:14.672745Z","shell.execute_reply.started":"2023-04-20T17:56:14.474924Z","shell.execute_reply":"2023-04-20T17:56:14.671573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# verify onnx\nimport onnx\n\n# Load the ONNX model\nmodel_onnx = onnx.load(saved_onnx)\n\n# Check that the model is well formed\nonnx.checker.check_model(model_onnx)\n\n# Print a human readable representation of the graph\nprint(onnx.helper.printable_graph(model_onnx.graph))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.674413Z","iopub.execute_input":"2023-04-20T17:56:14.674811Z","iopub.status.idle":"2023-04-20T17:56:14.816599Z","shell.execute_reply.started":"2023-04-20T17:56:14.674777Z","shell.execute_reply":"2023-04-20T17:56:14.815229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Export TfLite","metadata":{"execution":{"iopub.status.busy":"2023-04-03T17:01:20.891820Z","iopub.execute_input":"2023-04-03T17:01:20.892316Z","iopub.status.idle":"2023-04-03T17:01:20.901324Z","shell.execute_reply.started":"2023-04-03T17:01:20.892244Z","shell.execute_reply":"2023-04-03T17:01:20.899691Z"}}},{"cell_type":"code","source":"# Install library\n!pip install onnx-tf","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:14.817973Z","iopub.execute_input":"2023-04-20T17:56:14.818374Z","iopub.status.idle":"2023-04-20T17:56:27.830127Z","shell.execute_reply.started":"2023-04-20T17:56:14.818330Z","shell.execute_reply":"2023-04-20T17:56:27.828773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from onnx_tf.backend import prepare\n\ntf_rep = prepare(model_onnx)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:27.832784Z","iopub.execute_input":"2023-04-20T17:56:27.833266Z","iopub.status.idle":"2023-04-20T17:56:33.198086Z","shell.execute_reply.started":"2023-04-20T17:56:27.833217Z","shell.execute_reply":"2023-04-20T17:56:33.196838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### onnx to Tensorflow FrozenGraph(.pb)\n\nNow that a tf_rep variable has been created, the converted model can be exported to a .pb file and stored within this notebook.","metadata":{}},{"cell_type":"code","source":"pb_path = \"model.pb\"\ntf_rep.export_graph(pb_path)\n\nassert os.path.exists(pb_path)\nprint(\".pb model converted successfully.\")","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:33.199707Z","iopub.execute_input":"2023-04-20T17:56:33.200097Z","iopub.status.idle":"2023-04-20T17:56:36.935126Z","shell.execute_reply.started":"2023-04-20T17:56:33.200055Z","shell.execute_reply":"2023-04-20T17:56:36.933258Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### TensorFlow FrozenGraph (.pb) -> TensorFlow Lite (.tflite)\n\nNow a .pb model has been stored within the notebook, it can be prepared for Android/iOS deployment using the .tflite model format.","metadata":{}},{"cell_type":"markdown","source":"To use the TFLite converter to convert a FrozenGraph (.pb) file, the input and output nodes of the graph must be explicitly specified. The names of these nodes can be accessed easily using the existing tf_rep object created in **Section 2**.","metadata":{}},{"cell_type":"code","source":"input_nodes = tf_rep.inputs\noutput_nodes = tf_rep.outputs\nprint(\"The names of the input nodes are: {}\".format(input_nodes))\nprint(\"The names of the output nodes are: {}\".format(output_nodes))","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:36.937399Z","iopub.execute_input":"2023-04-20T17:56:36.938338Z","iopub.status.idle":"2023-04-20T17:56:36.946544Z","shell.execute_reply.started":"2023-04-20T17:56:36.938292Z","shell.execute_reply":"2023-04-20T17:56:36.945236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"converter = tf.lite.TFLiteConverter.from_saved_model(pb_path)\ntflite_rep = converter.convert()\n\ntflite_model_path = 'model.tflite'\nwith open(tflite_model_path, 'wb') as f:\n    f.write(tflite_rep)","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:36.948466Z","iopub.execute_input":"2023-04-20T17:56:36.948916Z","iopub.status.idle":"2023-04-20T17:56:37.836021Z","shell.execute_reply.started":"2023-04-20T17:56:36.948829Z","shell.execute_reply":"2023-04-20T17:56:37.834454Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction","metadata":{}},{"cell_type":"code","source":"!zip submission.zip $tflite_model_path","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:37.837569Z","iopub.execute_input":"2023-04-20T17:56:37.838000Z","iopub.status.idle":"2023-04-20T17:56:39.542402Z","shell.execute_reply.started":"2023-04-20T17:56:37.837953Z","shell.execute_reply":"2023-04-20T17:56:39.540471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Install library\n!pip3 install tflite_runtime","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:39.548436Z","iopub.execute_input":"2023-04-20T17:56:39.553496Z","iopub.status.idle":"2023-04-20T17:56:51.219742Z","shell.execute_reply.started":"2023-04-20T17:56:39.553449Z","shell.execute_reply":"2023-04-20T17:56:51.218351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import tflite_runtime.interpreter as tflite","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:51.222921Z","iopub.execute_input":"2023-04-20T17:56:51.223430Z","iopub.status.idle":"2023-04-20T17:56:51.241200Z","shell.execute_reply.started":"2023-04-20T17:56:51.223382Z","shell.execute_reply":"2023-04-20T17:56:51.240012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interpreter = tflite.Interpreter(tflite_model_path)\nfound_signatures = list(interpreter.get_signature_list().keys())\nprediction_fn = interpreter.get_signature_runner(\"serving_default\")\n\nlist_label = train_df['sign'].unique()\n\nfor i in range(100):\n    frames = load_relevant_data_subset(f'{train_df.iloc[i].path}')\n    output = prediction_fn(inputs=frames)\n    sign = np.argmax(output[\"outputs\"])\n\n    print(f\"Predicted label: {p2s_map[sign]}, Actual Label: {train_df.iloc[i].sign}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-04-20T17:56:51.243526Z","iopub.execute_input":"2023-04-20T17:56:51.244009Z","iopub.status.idle":"2023-04-20T17:56:53.192285Z","shell.execute_reply.started":"2023-04-20T17:56:51.243936Z","shell.execute_reply":"2023-04-20T17:56:53.190638Z"},"trusted":true},"execution_count":null,"outputs":[]}]}