{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<figure>\n  <img src=\"https://drive.google.com/uc?id=1Ugxmu8nlTVR3Egt7fjvjeRwEVqxrqxLy\" alt=\"Trulli\" style=\"width=100%\">\n</figure>\n","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">Introduction</div>\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Objectives</b>: In creating this notebook, my objectives are:(在創建這個筆記本時，我的目標是：</span>\n    \n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li>To learn about the data by exploration and visualization(通過探索和可視化了解數據</li>\n\n<li>To perform some processing techniques for further development(執行一些處理技術以進行進一步開發</li>\n\n</span></ul>\n\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Isolated Sign Language (ISL):</b> The signs in the dataset represent 250 of the first concepts taught to infants in any language. The goal is to create an isolated sign recognizer to incorporate into educational games for helping hearing parents of Deaf children learn American Sign Language (ASL) (Isolated Sign Language (ISL)：數據集中的符號代表以任何語言向嬰兒教授的第一個概念中的 250 個。 目標是創建一個獨立的手語識別器，以融入教育遊戲，幫助聾啞兒童的健聽父母學習美國手語 (ASL) [G1]<a href=\"https://www.kaggle.com/competitions/asl-signs/overview/data-card\">[G1]</a>\n<br> The 5 parameters of ASL are (ASL的5個參數是[G2]：<a href=\"https://www.mtsac.edu/llc/passportrewards/languagepartners/5ParametersofASL.pdf\">[G2]</a>:</span>\n\n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li>Handshapes(手型</li>\n\n<li>Palm Orientations(手掌方向</li>\n    \n<li>Locations(地點</li>\n    \n<li>Movements(動作</li>\n    \n<li>Non-Manual Signals (NMS)(非人工信號 (NMS)</li>\n\n</span></ul>\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Landmarks Files:</b> The landmarks were extracted from raw videos with the MediaPipe holistic model(地標文件：地標是使用 MediaPipe 整體模型 [G3] 從原始視頻中提取的。 <a href=\"https://google.github.io/mediapipe/solutions/holistic.html\">[G3]</a>. Not all of the frames necessarily had visible hands or hands that could be detected by the model(並非所有的框架都必須有可見的手或模型 [G4] 可以檢測到的手。 <a href=\"https://www.kaggle.com/competitions/asl-signs/data\">[G4]</a>. The spatial coordinates of the landmark are normalized to 0 and 1. Any points that are outside of [0, 1] are Mediapipe artifacts (界標的空間坐標被歸一化為 0 和 1。[0, 1] 之外的任何點都是 Mediapipe 工件 [G5]。<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/392286\">[G5]</a>.</span>\n    ","metadata":{}},{"cell_type":"code","source":"# Install\n!pip install -q itables 2> /dev/null\n!pip install -q flatbuffers 2> /dev/null\n!pip install -q mediapipe 2> /dev/null","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:01:19.406399Z","iopub.execute_input":"2023-05-02T18:01:19.406858Z","iopub.status.idle":"2023-05-02T18:01:57.500367Z","shell.execute_reply.started":"2023-05-02T18:01:19.406816Z","shell.execute_reply":"2023-05-02T18:01:57.498672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* itables：Jupyter notebooks 中的互動式表格庫\n* Flatbuffers：用於高效序列化和數據通信的庫\n* Mediapipe：用於構建多模式（例如，視頻、音訊等）應用的機器學習管道的框架。","metadata":{}},{"cell_type":"code","source":"# Imports\n\n# activate interactive mode of pd.dataframe\nimport pandas as pd\nfrom itables import init_notebook_mode\ninit_notebook_mode(all_interactive=True, connected=True)\n\nimport os\n\nimport json\nfrom tqdm import tqdm\nimport numpy as np\nimport itertools\n\nimport tensorflow as tf\n\n#pytorch model\nimport torch\nimport torch.nn.functional as F\nimport torch.nn as nn\n\nimport seaborn as sns\nimport mediapipe as mp\nimport matplotlib.pyplot as plt\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\n\nfrom matplotlib import animation\nfrom pathlib import Path\nimport IPython\nfrom IPython import display\nfrom IPython.core.display import display, HTML, Javascript\nfrom IPython.display import Markdown as md\n\nimport mediapipe as mp\nfrom mediapipe.framework.formats import landmark_pb2\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:01:57.503035Z","iopub.execute_input":"2023-05-02T18:01:57.503408Z","iopub.status.idle":"2023-05-02T18:02:09.957313Z","shell.execute_reply.started":"2023-05-02T18:01:57.503360Z","shell.execute_reply":"2023-05-02T18:02:09.956241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* pandas（PD）：數據處理庫\n* itables：Jupyter notebooks 中的互動式表格庫\n* OS：提供使用作業系統相關功能的可移植方式的模組\n* json：一個處理JSON（JavaScript物件表示法）數據的模組\n* TQDM：允許將進度條添加到迴圈和可反覆運算物件的模組\n* NumPy （NP）：一個數值計算庫\n* Itertools：一個模組，提供各種函數，這些函數在反覆運算器上工作以生成複雜的反覆運算器\n* TensorFlow（TF）：一個開源的機器學習框架\n* Torch：PyTorch 的機器學習框架\n* Seaborn：基於Matplotlib的數據可視化庫\n* MediaPipe （MP）：用於構建多模式（例如，視頻、音訊等）的框架 應用機器學習管道\n* matplotlib.pyplot（PLT）：一個提供類似MATLAB繪圖框架的模組\n* plotly.express（px）：一個用於在Python中創建互動式可視化的高級介面\n* plotly.graph_objects：用於在 Plotly 中創建複雜可視化的模組\n* make_subplots：用於在 Plotly 中創建子圖的功能\n* animation(動畫：用於在 Matplotlib 中創建動畫的模組\n* pathlib：用於處理文件系統路徑的模組\n* IPython：Python 的互動式計算環境\n* IPython.display：一個模組，提供用於在IPython中顯示各種類型的物件的類\n* HTML：用於創建 HTML 文件的類\n* Javascript：一個用於創建JavaScript代碼片段的類\n* landmark_pb2：包含地標 protobuf 定義的模組。\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# Config\n\nclass Cfg:\n    INPUT_ROOT = Path('/kaggle/input/asl-signs/')\n    OUTPUT_ROOT = Path('kaggle/working')\n    INDEX_MAP_FILE = INPUT_ROOT / 'sign_to_prediction_index_map.json'\n    TRAN_FILE = INPUT_ROOT / 'train.csv'\n    INDEX = 'sequence_id'\n    ROW_ID = 'row_id'\n    \nLANDMARK_FILES_DIR = \"/kaggle/input/asl-signs/train_landmark_files\"\nlabel_map = json.load(open(\"/kaggle/input/asl-signs/sign_to_prediction_index_map.json\", \"r\"))\n\ntrain_df = pd.read_csv(\"/kaggle/input/gislr-extended-train-dataframe/extended_train.csv\")\ntrain_df['label'] = train_df['sign'].map(label_map)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:09.958934Z","iopub.execute_input":"2023-05-02T18:02:09.959691Z","iopub.status.idle":"2023-05-02T18:02:10.833490Z","shell.execute_reply.started":"2023-05-02T18:02:09.959645Z","shell.execute_reply":"2023-05-02T18:02:10.832422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* INPUT_ROOT：指向 ASL 標誌數據所在的輸入目錄的路徑\n* OUTPUT_ROOT：代碼將在其中寫入資料的輸出目錄的路徑\n* INDEX_MAP_FILE：將 ASL 標誌映射到預測索引的 JSON 檔案的路徑\n* TRAN_FILE：包含訓練數據的 CSV 檔案的路徑\n* INDEX：訓練數據中的一列，用作每個序列的唯一標識符\n* ROW_ID：訓練數據中的一列，用作每行的唯一標識符\n* \n* LANDMARK_FILES_DIR：包含訓練數據中每個序列的地標檔的目錄的路徑\n* label_map：將每個 ASL 符號映射到預測索引的字典\n* train_df：包含訓練數據的 Pandas 數據幀，其中包含一個名為“label”的新列，用於將每個符號映射到其相應的預測索引。","metadata":{}},{"cell_type":"code","source":"# Helpers\n\nROWS_PER_FRAME = 543\ndef load_relevant_data_subset(pq_path):\n    data_columns = ['x', 'y', 'z']\n    data = pd.read_parquet(pq_path, columns=data_columns)\n    n_frames = int(len(data) / ROWS_PER_FRAME)\n    data = data.values.reshape(n_frames, ROWS_PER_FRAME, len(data_columns))\n    return data.astype(np.float32)\n\n# https://www.kaggle.com/code/ted0071/gislr-visualization\ndef read_index_map(file_path=Cfg.INDEX_MAP_FILE):\n    \"\"\"Reads the sign to predict as json file.\"\"\"\n    with open(file_path, \"r\") as f:\n        result = json.load(f)\n    return result    \n\ndef read_train(file_path=Cfg.TRAN_FILE):\n    \"\"\"Reads the train csv as pandas data frame.\"\"\"\n    return pd.read_csv(file_path).set_index(Cfg.INDEX)\n\ndef read_landmark_data_by_path(file_path, input_root=Cfg.INPUT_ROOT):\n    \"\"\"Reads landmak data by the given file path.\"\"\"\n    data = pd.read_parquet(input_root / file_path)\n    return data.set_index(Cfg.ROW_ID)\n\ndef read_landmark_data_by_id(sequence_id, train_data):\n    \"\"\"Reads the landmark data by the given sequence id.\"\"\"\n    file_path = train_data.loc[sequence_id]['path']\n    return read_landmark_data_by_path(file_path)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:10.836935Z","iopub.execute_input":"2023-05-02T18:02:10.837714Z","iopub.status.idle":"2023-05-02T18:02:10.849102Z","shell.execute_reply.started":"2023-05-02T18:02:10.837659Z","shell.execute_reply":"2023-05-02T18:02:10.847745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* load_relevant_data_subset：此函數載入特定序列的相關數據子集，給定包含所有數據的 Parquet 檔。它返回一個 NumPy 陣列，其中包含序列中每個幀的相關數據。\n* read_index_map：此函數讀取將每個 ASL 符號映射到預測索引的 JSON 檔。\n* read_train：此函數讀取包含訓練數據的 CSV 檔案，並將索引設置為 。Cfg.INDEX\n* read_landmark_data_by_path：此函數在給定檔路徑的情況下，從 Parquet 檔中讀取特定序列的地標數據。\n* read_landmark_data_by_id：此函數在給定序列ID和訓練資料的情況下，從Parquet檔中讀取特定序列的地標數據。它用於從檔中讀取數據。read_landmark_data_by_path","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">Metadata Analysis(元數據分析)</div>\n\n<span style=\"font-size:18px; font-family:Georgia;\">In this section, I want to explore the number of sequences, frames, and average frames per sign, and participant. Thus, I split the train data into 2 levels for further analysis:(在本節中，我想探索每個標誌和參與者的序列數、幀數和平均幀數。 因此，我將訓練數據分為 2 個級別進行進一步分析： </span>\n\n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li><b>ISL Level:</b> I group the train dataframe by <b>sign</b>. Then, I aggregrate the number of sequence ids and total number of frames. This level is to help me explore the number of sequences, frames, and average frames per sign.(我按符號對訓練數據幀進行分組。 然後，我匯總了序列 ID 的數量和幀總數。 這個級別是為了幫助我探索每個標誌的序列數、幀數和平均幀數。</li>\n\n<li><b>Participant Level:</b> I group the train dataframe by <b>participant_id</b> and <b>sign</b>. Then, I aggregrate the number of sequence ids and total number of frames. This level is to help me explore the number of sequences, frames, and average frames per sign per participant.(我按 participant_id 和符號對訓練數據框進行分組。 然後，我匯總了序列 ID 的數量和幀總數。 這個級別是為了幫助我探索每個參與者的每個符號的序列數、幀數和平均幀數。</li>\n\n</span></ul>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>ISL Level</b></span>\n","metadata":{}},{"cell_type":"code","source":"meta_data_df = train_df.groupby('sign').agg({'sequence_id': 'count',\n                                             'total_frames': 'sum'})\nmeta_data_df.columns = ['num_seq','num_frames']\nmeta_data_df['avg_frames'] = np.round(meta_data_df['num_frames']/meta_data_df['num_seq'])\n\nmeta_data_df","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:10.850651Z","iopub.execute_input":"2023-05-02T18:02:10.851034Z","iopub.status.idle":"2023-05-02T18:02:10.922930Z","shell.execute_reply.started":"2023-05-02T18:02:10.850997Z","shell.execute_reply":"2023-05-02T18:02:10.921590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"此代碼按「sign」列對訓練數據進行分組，並聚合「sequence_id」和「total_frames」列。生成的數據幀包含每個符號的序列數和幀總數。它還計算每個符號的每個序列的平均幀數，並將其作為新列添加到數據幀中。","metadata":{}},{"cell_type":"code","source":"int(meta_data_df.num_seq.sum()/250), int(meta_data_df.avg_frames.sum()/250)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:10.925075Z","iopub.execute_input":"2023-05-02T18:02:10.926048Z","iopub.status.idle":"2023-05-02T18:02:10.935868Z","shell.execute_reply.started":"2023-05-02T18:02:10.925995Z","shell.execute_reply":"2023-05-02T18:02:10.934772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. 此代碼計算訓練期間將處理的數據批次數。它將訓練數據中的序列總數除以 250，這是代碼中使用的批大小。\n1. 同樣，它計算訓練數據中所有符號的每個序列的平均幀數，並將其除以 250 以獲得每批的平均幀數。","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n\ncols_name = meta_data_df.columns.values.tolist()\ncolors = [\"#0F9D58\",\"#4285F4\",\"#F4B400\"]\n\nfor color,col in zip(colors,cols_name):\n    # store tmp df\n    tmp = meta_data_df.sort_values(col)\n    fig.add_trace(go.Bar(x=tmp.index, \n                         y=tmp[col],\n                         width=0.5,name=col,marker_color=color))\n\nfig.update_layout(\n    title={\n            'text': \"ISL Distribution: All Traces\",\n            'font': dict(size=20,family=\"Georgia\",color=colors[1]),\n            'y':0.87,\n            'x':0.035,\n            'xanchor': 'left',\n            'yanchor': 'top'},\n    template=\"plotly_white\",\n    xaxis_tickangle=-45,\n    width= 4000,\n    xaxis=dict(title='Sign', fixedrange=True),\n    yaxis=dict(title='Count',fixedrange=True),\n    showlegend=True,\n  \n    updatemenus=[\n        dict(\n            # customize dropdown\n            active=0,\n            direction=\"down\",\n            pad={\"r\": 50, \"t\": 25},\n            showactive=True,\n            x=0.005,\n            xanchor=\"right\",\n            y=1.2,\n            yanchor=\"top\",\n            \n            # customize button\n            buttons=list([\n                dict(label=\"All\",\n                     method=\"update\",\n                     args=[{\"visible\": [True, True,True]},\n                           {\"title\": \"ISL Distribution: All Traces\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Sequences\",\n                     method=\"update\",\n                     args=[{\"visible\": [True, False,False]},\n                           {\"title\": \"Distribution of Number of Sequences per Sign\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Frames\",\n                     method=\"update\",\n                     args=[{\"visible\": [False,True, False]},\n                           {\"title\": \"Distribution of Number of Frames per Sign\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Avg frames\",\n                     method=\"update\",\n                     args=[{\"visible\": [False, False,True]},\n                           {\"title\": \"Distribution of Average Frames per Sequence per Sign\",\n                            \"legend\":True,\n                            }]),\n            ]),\n        ),\n    ])\n\nfig.show(config= dict(displayModeBar = False))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:10.938101Z","iopub.execute_input":"2023-05-02T18:02:10.939019Z","iopub.status.idle":"2023-05-02T18:02:12.572542Z","shell.execute_reply.started":"2023-05-02T18:02:10.938967Z","shell.execute_reply":"2023-05-02T18:02:12.571361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"此代碼使用 Plotly 庫創建條形圖，以顯示 ISL（印度手語）數據的分佈。它使用來自PD數據幀的數據meta_data_df\n\n該圖表有一個下拉功能表，允許使用者在數據的不同檢視之間切換：\n\n* 全部：顯示每個符號的所有三個指標（序列數、幀數和每個序列的平均幀數）的分佈。\n* 序列：顯示每個符號序列數的分佈。\n* 幀數：顯示每個標誌的幀總數的分佈。\n* 平均幀數：顯示每個符號每個序列的平均幀數分佈。\n圖表中的每個條形代表一個不同的符號，並根據所顯示的指標進行著色。x 軸顯示符號名稱，y 軸顯示所選指標的計數。圖表標題為“ISL 分佈：所有跡線”，圖例顯示在圖表頂部。","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:white;font-size:22px;font-family:Georgia;border-style: solid;border-color: #F4B400;border-width:5px;padding:20px;margin: 0px;color:black;overflow:hidden\">\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Observations:</b> </span>\n    \n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li>Number of sequences are evenly distributed for each sign(每個符號的序列數均勻分佈</li>\n\n<li>Sign \"mitten\" has the most number of frames, and average frames per sequence (符號“連指手套”的幀數最多，每個序列的平均幀數最多</li>\n    \n<li>On average, there are <b>377</b> sequences per sign, and <b>37</b> frames per sequence(平均每個符號有 377 個序列，每個序列有 37 個幀</li>\n\n</span></ul>\n\n</div>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>Participant Level</b></span>","metadata":{}},{"cell_type":"code","source":"participant_level_df = train_df.groupby(['participant_id','sign']).agg({'sequence_id': 'count',\n                                                                        'total_frames': 'sum'})\nparticipant_level_df['avg_frames'] = np.round(participant_level_df.total_frames/participant_level_df.sequence_id)\nparticipant_level_df.columns = ['num_seq','num_frames','avg_frames']\nparticipant_level_df = participant_level_df.reset_index()\nparticipant_level_df","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:12.573605Z","iopub.execute_input":"2023-05-02T18:02:12.573952Z","iopub.status.idle":"2023-05-02T18:02:12.681573Z","shell.execute_reply.started":"2023-05-02T18:02:12.573891Z","shell.execute_reply":"2023-05-02T18:02:12.680554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"此代碼創建一個新的 DataFrame，通過按 和 對 DataFrame 進行分組來調用。 生成的 DataFrame 具有以下列：participant_level_dftrain_dfparticipant_idsign\n\n* participant_id：參與者的ID。\n* sign：序列對應的符號。\n* num_seq：符號和參與者的序列數。\n* num_frames：標誌和參與者的總幀數。\n* avg_frames：標誌和參與者每個序列的平均幀數。\n\n該方法用於將多個聚合函數應用於分組數據。 該函數應用於列以計算每個組的序列數。 該函數應用於列以計算每個組的總幀數。\n最後，應用該函數用於重命名結果 DataFrame.agg()count()sequence_idsum()total_framesnp.round() 的列","metadata":{}},{"cell_type":"code","source":"int(participant_level_df.num_seq.sum()/21), int(participant_level_df.num_frames.sum()/21)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:12.682799Z","iopub.execute_input":"2023-05-02T18:02:12.684118Z","iopub.status.idle":"2023-05-02T18:02:12.691884Z","shell.execute_reply.started":"2023-05-02T18:02:12.684079Z","shell.execute_reply":"2023-05-02T18:02:12.690845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* 每個參與者的平均序列數約為54，每個參與者的平均幀數約為11765。","metadata":{}},{"cell_type":"code","source":"fig = go.Figure()\n\ncols_name = participant_level_df.columns.values.tolist()[2:]\ncolors = [\"#0F9D58\",\"#4285F4\",\"#F4B400\"]\nmethods = ['sum','sum','sum']\n\nfor color,col,method in zip(colors,cols_name,methods):\n    # store tmp series object\n    tmp = participant_level_df.groupby(['participant_id'])\\\n                              .agg({col: method})[col]\\\n                              .sort_values(ascending=True)\n    fig.add_trace(go.Bar(x=tmp.index.astype('str'), \n                         y=tmp.values,\n                         width=0.8,name=col,marker_color=color))\nfig.update_layout(\n    title={\n            'text': \"Participant Distribution: All Traces\",\n            'font': dict(size=20,family=\"Georgia\",color=colors[1]),\n            'y':0.87,\n            'x':0.18,\n            'xanchor': 'left',\n            'yanchor': 'top'},\n    template=\"plotly_white\",\n    xaxis_tickangle=-45,\n    width= 800,\n    height=500,\n    xaxis=dict(title='Participant_ID', fixedrange=True),\n    yaxis=dict(title='Count',fixedrange=True),\n    showlegend=True,\n    \n    updatemenus=[\n        dict(\n            # customize dropdown\n            active=0,\n            direction=\"down\",\n            pad={\"r\": 50, \"t\": 25},\n            showactive=True,\n            x=0.01,\n            xanchor=\"right\",\n            y=1.20,\n            yanchor=\"top\",\n            \n            # customize button\n            buttons=list([\n                dict(label=\"All\",\n                     method=\"update\",\n                     args=[{\"visible\": [True, True,True]},\n                           {\"title\": \"Participant Distribution: All Traces\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Sequences\",\n                     method=\"update\",\n                     args=[{\"visible\": [True, False,False]},\n                           {\"title\": \"Distribution of Number of Sequences per Participant\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Frames\",\n                     method=\"update\",\n                     args=[{\"visible\": [False,True, False]},\n                           {\"title\": \"Distribution of Number of Frames per Participant\",\n                            \"legend\":True,\n                            }]),\n                dict(label=\"Avg frames\",\n                     method=\"update\",\n                     args=[{\"visible\": [False, False,True]},\n                           {\"title\": \"Distribution of Average Frames per Participant\",\n                            \"legend\":True,\n                            }]),\n            ]),\n        ),\n    ])\n\nfig.show(config= dict(displayModeBar = False))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:12.695762Z","iopub.execute_input":"2023-05-02T18:02:12.696199Z","iopub.status.idle":"2023-05-02T18:02:12.760444Z","shell.execute_reply.started":"2023-05-02T18:02:12.696166Z","shell.execute_reply":"2023-05-02T18:02:12.759288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"此代碼塊創建一個條形圖，顯示每個參與者的序列數、幀數和每個序列的平均幀數的分佈。 數據是從 participant_level_df DataFrame 獲得的，該數據幀之前是通過按參與者 ID 和標誌對 train_df DataFrame 進行分組並聚合每個標誌的序列數和幀數而創建的。 該代碼使用 Plotly 創建一個交互式條形圖，其中包含用於在不同分佈之間切換的下拉菜單選項。\n\n該代碼首先定義了 fig 對象並設置了佈局和模板。 cols_name 列表存儲不同分佈的列名（即 num_seq、num_frames 和 avg_frames）。 顏色列表存儲條形的顏色代碼，方法列表存儲按參與者 ID 分組時使用的聚合方法。\n\n然後代碼使用 for 循環為每個分佈創建條形圖。 在循環中，tmp 變量存儲一個臨時系列對象，該對像是通過按參與者 ID 對 participant_level_df DataFrame 進行分組並使用指定方法聚合相應列而創建的。 然後調用 go.Bar() 函數為給定分佈創建條形圖，使用參與者 ID 作為 x 軸，使用聚合值作為 y 軸。 使用 add_trace() 方法將跟踪添加到 fig 對象。\n\n然後代碼更新佈局以包含下拉菜單選項。 updatemenus 列表包含一個定義下拉菜單的字典。 每個菜單選項都由按鈕列表中的字典對象定義。 label 鍵定義要在下拉菜單選項上顯示的文本，而 method 鍵定義選擇該選項時要執行的操作。 方法字典中的 args 鍵定義了傳遞給 update_layout() 方法的參數，該方法根據所選選項更新圖表標題和圖例。\n\n最後，調用 fig.show() 方法來顯示圖表，配置參數設置為 dict(displayModeBar = False) 以隱藏模式欄。","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:white;font-size:22px;font-family:Georgia;border-style: solid;border-color: #F4B400;border-width:5px;padding:20px;margin: 0px;color:black;overflow:hidden\">\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Observations:</b> </span>\n    \n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li>3 participants have the least number of sequences(3 個參與者的序列數最少</li>\n\n<li>Participant ID-49445 has the most number of frames, whereas participant ID-37779 has the least number of frames(參與者 ID-49445 的幀數最多，而參與者 ID-37779 的幀數最少 </li>\n    \n<li>On average, there are <b>4498</b> sequences per participant, and <b>170666</b> frames per participant.(平均而言，每個參與者有 4498 個序列，每個參與者有 170666 個幀。</li>\n\n</span></ul>\n\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">Landmarks Analysis</div>\n\n\n<span style=\"font-size:18px; font-family:Georgia;\">In this section, I want to visualize and gain more insights from individual frames for each sequence. Thus, I adjust some helper functions from Roland Abel <a href=\"https://www.kaggle.com/code/ted0071/gislr-visualization\">[C1]</a> and create another helper function for interactive landmarks visualization.<br>Below are some references for hand, full body, and face landmarks: (在本節中，我想可視化每個序列的各個幀並從中獲得更多見解。 因此，我調整了 Roland Abel [C1] 的一些輔助函數，並為交互式地標可視化創建了另一個輔助函數。\n以下是手部、全身和面部標誌的一些參考：</span>\n\n<span style=\"font-size:25px; font-family:Georgia;\"><b>Hand Landmarks</b></span>\n\n![Hand Landmarks](https://developers.google.com/static/mediapipe/images/solutions/hand-landmarks.png)","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>Full Body Landmarks</b></span>\n![Full Body Landmarks](https://mediapipe.dev/images/mobile/pose_tracking_full_body_landmarks.png)\n\n\n<span style=\"font-size:18px; font-family:Georgia;\">For ASL, the upper body landmarks are more important than the lower body landmarks.(對於 ASL，上身標誌比下身標誌更重要\n</span>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>Face Landmarks (Contours)</b></span>\n\n\n<span style=\"font-size:18px; font-family:Georgia;\">For ASL, the non-manual signals such as head nod, eyebrows, nose, eyes, and lips are used to convey additional meaning with a sign. Thus, instead of visualizing the facemesh tesselation, I decide to only use facemesh contours.(對於美國手語，點頭、眉毛、鼻子、眼睛和嘴唇等非手動信號用於通過符號傳達額外的含義。 因此，我決定只使用面部網格輪廓，而不是可視化面部網格曲面細分。<br> Face Landmarks to use:(要使用的面部地標： </span>\n\n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li><b>Tensors to face landmarks(面對地標的張量</b> <a href=\"https://github.com/google/mediapipe/blob/master/mediapipe/modules/face_landmark/tensors_to_face_landmarks_with_attention.pbtxt\">[G6]</a></li>\n\n<li><b>Facemesh connections(Facemesh 連接</b> <a href=\"https://github.com/google/mediapipe/blob/master/mediapipe/python/solutions/face_mesh_connections.py\">[G7]</a></li>\n\n</span></ul>\n","metadata":{}},{"cell_type":"code","source":"# Helper Functions\n# [C1] adjusted from Roland Abel: https://www.kaggle.com/code/ted0071/gislr-visualization\ntrain_data = read_train()\n\nmp_drawing = mp.solutions.drawing_utils\nmp_hands = mp.solutions.hands\nmp_face_mesh = mp.solutions.face_mesh\nmp_pose = mp.solutions.pose\n\n# contour connections\nCONTOURS = list(itertools.chain(*mp_face_mesh.FACEMESH_CONTOURS))\n\ndef create_blank_image(height, width):\n    return np.zeros((height, width, 3), np.uint8)\n\ndef draw_landmarks(data, image, frame_id, \n                   landmark_type, connection_type, \n                   landmark_color=(255, 0, 0), connection_color=(0, 20, 255), \n                   thickness=2, circle_radius=1):\n    \"\"\"Draws landmarks\"\"\"\n    df = data.groupby(['frame', 'type']).get_group((frame_id, landmark_type)).copy()\n    if landmark_type == 'face':\n        df.loc[~df['landmark_index'].isin(CONTOURS),'x'] = float('NaN') #-1*df[~df['landmark_index'].isin(CONTOURS)]['x'].values\n\n        \n    landmarks = [landmark_pb2.NormalizedLandmark(x=lm.x, y=lm.y, z=lm.z) for idx, lm in df.iterrows()]\n    landmark_list = landmark_pb2.NormalizedLandmarkList(landmark = landmarks)\n    #print(len(landmark_list.landmark))\n    mp_drawing.draw_landmarks(\n        image=image,\n        landmark_list=landmark_list, \n        connections=connection_type,\n        landmark_drawing_spec=mp_drawing.DrawingSpec(\n            color=landmark_color, \n            thickness=thickness, \n            circle_radius=circle_radius),\n        connection_drawing_spec=mp_drawing.DrawingSpec(\n            color=connection_color, \n            thickness=thickness, \n            circle_radius=circle_radius))\n    return image","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:12.761997Z","iopub.execute_input":"2023-05-02T18:02:12.762432Z","iopub.status.idle":"2023-05-02T18:02:12.967706Z","shell.execute_reply.started":"2023-05-02T18:02:12.762388Z","shell.execute_reply":"2023-05-02T18:02:12.966741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"看起來這段代碼定義了一些説明程式函數，用於使用 MediaPipe 在圖像上繪製地標。具體來說，它定義了以下函數：\n\n1. create_blank_image(height, width)：創建具有給定尺寸的空白（黑色）圖像。\n1. draw_landmarks(data, image, frame_id, landmark_type, connection_type, landmark_color=(255, 0, 0), connection_color=(0, 20, 255), thickness=2, circle_radius=1)：使用給定數據幀中的地標數據在給定圖像上繪製地標。地標是為給定的 和繪製的，地標之間的連接是使用給定的 繪製的。地標和連接的顏色、粗細和圓半徑可以選擇使用、、 和 參數進行自定義。","metadata":{}},{"cell_type":"code","source":"def get_ids(df, row):\n    participant_id = df.participant_id.values[row]\n    sequence_id = df.sequence_id.values[row]\n    \n    return participant_id, sequence_id\n\ndef draw_data(participant_id, sequence_id, train_data):\n    height = 700\n    width = 500\n\n    # Read and get frames\n    data = read_landmark_data_by_id(sequence_id, train_data)\n    frame_ids = data.frame.unique().tolist()\n    buttons_ids = []\n    buttons_seq_ids = []\n    buttons=[]\n\n    fig = make_subplots(rows=2, cols=3,\n                    specs=[[{}, {},{\"rowspan\": 2}],\n                           [{}, {},None]],\n                    vertical_spacing=0.1,\n                    subplot_titles=('Face',  'Pose',\n                                    'All',  'Left Hand',\n                                    'Right Hand'),\n                    print_grid=False)\n\n    buttons_seq_ids.append(dict(label=f\"{sequence_id}\",\n                                method=\"restyle\",\n                                args=[{\"visible\": None}]\n                                ))\n    buttons_ids.append(dict(label=f\"{participant_id}\",\n                                method=\"restyle\",\n                                args=[{\"visible\": None}]\n                                ))\n\n    for i,frame_id in enumerate(frame_ids): \n        r_hand = draw_landmarks(data, image=create_blank_image(height, width ), \n                              frame_id=frame_id,\n                              landmark_type = 'right_hand', \n                              connection_type = mp_hands.HAND_CONNECTIONS,\n                              landmark_color=(255, 0, 0),\n                              connection_color=(0, 20, 255), \n                              thickness=3, \n                              circle_radius=3)\n\n\n        l_hand = draw_landmarks(data, image=create_blank_image(height, width), \n                              frame_id=frame_id,\n                              landmark_type = 'left_hand', \n                              connection_type = mp_hands.HAND_CONNECTIONS,\n                              landmark_color=(255, 0, 0),\n                              connection_color=(0, 20, 255), \n                              thickness=3, \n                              circle_radius=3)\n\n\n\n        face = draw_landmarks(data, image=create_blank_image(height, width), \n                              frame_id=frame_id,\n                              landmark_type='face', \n                              connection_type= mp_face_mesh.FACEMESH_CONTOURS,\n                              landmark_color=(255, 255, 255),\n                              connection_color=(0, 255, 0),\n                              thickness=1, \n                              circle_radius=1)\n\n        pose = draw_landmarks(data, image=create_blank_image(height, width), \n                               frame_id=frame_id,\n                               landmark_type='pose', \n                               connection_type= mp_pose.POSE_CONNECTIONS,\n                               landmark_color=(255, 255, 255),\n                               connection_color=(255, 0, 0),\n                               thickness=2, \n                               circle_radius=2)\n\n        fig.add_trace(px.imshow(face).data[0], row=1, col=1)\n        fig.add_trace(px.imshow(pose).data[0], row=1, col=2)\n        fig.add_trace(px.imshow(l_hand).data[0], row=2, col=1)\n        fig.add_trace(px.imshow(r_hand).data[0], row=2, col=2)\n        fig.add_trace(px.imshow(face+pose+l_hand+r_hand, aspect='auto').data[0], row=1, col=3)\n\n        visible=[False,False,False,False,False]*len(frame_ids)\n        visible[i*5:i*5+5]=[True]*5\n        buttons.append(dict(label=f\"{frame_id}\",\n                            method=\"update\",\n                            args=[{\"visible\": visible}]))  \n\n    sign = train_df.query('sequence_id == @sequence_id')['sign'].values[0]\n\n    fig.update_layout(\n        title={\n            'text': f'<b>Sign: {sign}',\n            'font': dict(size=20,family=\"Georgia\",color=colors[1]),\n            'y':0.98,\n            'x':0.5,\n            'xanchor': 'center',\n            'yanchor': 'top'},\n\n\n        template=\"plotly_white\",\n        width= 800,\n        height=600,\n        showlegend=True,\n\n\n        updatemenus=[\n            # Participant_ID\n            dict(\n                # customize dropdown\n                active=0,\n                direction=\"down\",\n                pad={\"r\": 50, \"t\": 25},\n                showactive=True,\n                x=0.1,\n                xanchor=\"left\",\n                y=1.2,\n                yanchor=\"top\",\n\n                # customize button      \n                buttons=buttons_ids),\n\n            # Sequence_ID\n            dict(\n                # customize dropdown\n                active=0,\n                direction=\"down\",\n                pad={\"r\": 50, \"t\": 25},\n                showactive=True,\n                x=0.43,\n                xanchor=\"left\",\n                y=1.2,\n                yanchor=\"top\",\n\n                # customize button      \n                buttons=buttons_seq_ids),\n\n            # Frames_ID\n            dict(\n                # customize dropdown\n                active=0,\n                direction=\"down\",\n                pad={\"r\": 50, \"t\": 25},\n                showactive=True,\n                x=0.8,\n                xanchor=\"left\",\n                y=1.2,\n                yanchor=\"top\",\n\n                # customize button      \n                buttons=buttons),\n\n        ])\n\n    fig.update_xaxes(showticklabels=False,fixedrange=True)\n    fig.update_yaxes(showticklabels=False,fixedrange=True)\n\n    fig.add_annotation(text=\"Participant_ID\", x=-0.05, xref=\"paper\", y=1.12, yref=\"paper\",\n                       align=\"left\", showarrow=False)\n    fig.add_annotation(text=\"Sequence_ID\", x=0.35, xref=\"paper\", y=1.125, yref=\"paper\",\n                       align=\"left\", showarrow=False)\n    fig.add_annotation(text=\"Frame_ID\", x=0.78, xref=\"paper\", y=1.13, yref=\"paper\",\n                       align=\"left\", showarrow=False)\n    \n    return fig","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-05-02T18:02:12.969558Z","iopub.execute_input":"2023-05-02T18:02:12.970161Z","iopub.status.idle":"2023-05-02T18:02:13.001399Z","shell.execute_reply.started":"2023-05-02T18:02:12.970111Z","shell.execute_reply":"2023-05-02T18:02:13.000186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"這是一個 Python 代碼，它使用 Plotly 生成一個圖來可視化手勢數據。 該圖使用地標檢測模型中的 x、y 坐標數據顯示五個子圖，包括面部、姿勢、左手、右手和所有手勢。\n\n這個 get_ids 函數接受一個數據框 df 和一個行索引 row，並返回相應行的參與者 ID 和序列 ID。\n\n這個 draw_data 函數獲取參與者 ID、序列 ID 和訓練數據，並返回一個 Plotly 圖形對象。 該函數使用 read_landmark_data_by_id 函數從訓練數據中讀取特定序列的地標數據，然後使用 draw_landmarks 函數為面部、姿勢、左手、右手和所有手勢創建五個子圖。 它還添加了三個下拉菜單，分別用於選擇參與者 ID、序列 ID 和幀 ID。 然後返回圖形對象。\n\n總的來說，這段代碼允許我們可視化手勢序列的地標數據，並選擇特定的幀、參與者和序列進行觀察。","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>Landmarks Visualization(地標可視化)</b></span>","metadata":{}},{"cell_type":"code","source":"sign_table = {}\nfor i in range(25):\n    sign_table[f'{i}'] = train_df.sign.unique()[i*10:i*10+10].tolist()\nsign_table = pd.DataFrame(sign_table)\nsign_table","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:13.003008Z","iopub.execute_input":"2023-05-02T18:02:13.003408Z","iopub.status.idle":"2023-05-02T18:02:13.236594Z","shell.execute_reply.started":"2023-05-02T18:02:13.003375Z","shell.execute_reply":"2023-05-02T18:02:13.235133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"上面的代碼塊創建了一個字典，其中鍵是從 0 到 24 的數位字串，值是使用該方法提取的數據幀中的 10 個唯一符號名稱的清單。sign_tabletrain_dfunique()\n\n然後，它從字典創建一個 Pandas 資料幀，其中每列是 10 個符號名稱的清單。生成的 DataFrame 將有 25 列，字典中的每個鍵對應一列，每列將包含 10 個唯一符號名稱的清單。sign_tablesign_table\n\n請注意，生成的數據幀可能在某些單元格中具有值，因為並非所有符號類別在數據幀中可能具有相同數量的唯一值。NaNtrain_df","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:20px; font-family:Georgia;\"><b>Example 1: Wait</b></span>","metadata":{}},{"cell_type":"markdown","source":"> <span style=\"font-size:18px; font-family:Georgia;\"><b>Disclaimer:</b> I want to produce an interactive visualization based on Participant_ID, Sequence_ID, and Frame_ID. However, I haven't figured out the way to do the chained dropdown callback without using Dash Plottly yet. Thus, only the <b>Frame_ID</b> dropdown is <b>active</b>, the <b>Participant_ID</b> and <b>Sequence_ID</b> dropdowns are currently <b>inactive</b> ^^'.(免責聲明：我想生成基於 Participant_ID、Sequence_ID 和 Frame_ID 的交互式可視化。 但是，我還沒有想出在不使用 Dash Plottly 的情況下進行鍊式下拉回調的方法。 因此，只有 Frame_ID 下拉列表處於活動狀態，Participant_ID 和 Sequence_ID 下拉列表當前處於非活動狀態 ^^'。 </span>","metadata":{}},{"cell_type":"code","source":"participant_id, sequence_id = get_ids(train_df[train_df['sign'] == 'wait'], 10)\nfig = draw_data(participant_id,sequence_id,train_data)\nfig.show(config= dict(displayModeBar = False))","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:13.237895Z","iopub.execute_input":"2023-05-02T18:02:13.239051Z","iopub.status.idle":"2023-05-02T18:02:48.527546Z","shell.execute_reply.started":"2023-05-02T18:02:13.239008Z","shell.execute_reply":"2023-05-02T18:02:48.525799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"font-size:20px; font-family:Georgia;\"><b>Example 2: Cloud</b></span>","metadata":{}},{"cell_type":"code","source":"participant_id, sequence_id = get_ids(train_df[train_df['sign'] == 'cloud'], 10 )\nfig = draw_data(participant_id,sequence_id,train_data)\nfig.show(config= dict(displayModeBar = False))","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:48.528804Z","iopub.execute_input":"2023-05-02T18:02:48.529180Z","iopub.status.idle":"2023-05-02T18:02:54.734160Z","shell.execute_reply.started":"2023-05-02T18:02:48.529145Z","shell.execute_reply":"2023-05-02T18:02:54.732989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\"font-size:20px; font-family:Georgia;\"><b>Example 3: Flower</b></span>","metadata":{}},{"cell_type":"code","source":"participant_id, sequence_id = get_ids(train_df[train_df['sign'] == 'flower'], 10 )\nfig = draw_data(participant_id,sequence_id,train_data)\nfig.show(config= dict(displayModeBar = False))","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2023-05-02T18:02:54.735387Z","iopub.execute_input":"2023-05-02T18:02:54.735703Z","iopub.status.idle":"2023-05-02T18:03:00.333659Z","shell.execute_reply.started":"2023-05-02T18:02:54.735671Z","shell.execute_reply":"2023-05-02T18:03:00.332557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background-color:white;font-size:22px;font-family:Georgia;border-style: solid;border-color: #F4B400;border-width:5px;padding:20px;margin: 0px;color:black;overflow:hidden\">\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Observations:</b> </span>\n\n<ul style=“list-style-type:circle;”><span style='font-size:18px; font-family:Georgia;'>\n\n<li>Some face frames are NaN, only pose landmarks are always available(有些人臉框是 NaN，只有姿勢界標始是可用的</li>\n\n<li>Either left hand or right hand landmarks are available(左手或右手地標可用 </li>\n    \n<li>Key frames aren't always from first and last frames(關鍵幀並不總是來自第一幀和最後一幀</li>\n\n<li>Lots of empty hand frames in a sequence(一個序列中有很多空手框</li>\n\n</span></ul>\n\n<span style=\"font-size:18px; font-family:Georgia;\">=> Besides Mediapipe artifacts (normalized coordinates out of range [0,1]), we need to deal with NaN coordinates as well!(除了 Mediapipe 工件（標準化坐標超出範圍 [0,1]），我們還需要處理 NaN 坐標！</span>\n</div>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">Feature Processing(特徵處理)</div>\n\n<span style=\"font-size:25px; font-family:Georgia;\"><b>Processing</b>: </span>\n\n<span style=\"font-size:18px; font-family:Georgia;\">In this section, I want to visualize and implement a feature preprocessing, inspired by Robert Hatch <a href=\"https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders/notebook\">[C2]</a>, Darien Schettler <a href=\"https://www.kaggle.com/code/dschettler8845/gislr-learn-eda-baseline\">[C3]</a> , and Heng CK <a href=\"https://www.kaggle.com/code/hengck23/lb-0-62-pytorch-transformer-solution/notebook\">[C4]</a> public notebooks. (在本節中，我想可視化並實施特徵預處理，靈感來自 Robert Hatch [C2]、Darien Schettler [C3] 和 Heng CK [C4] 公共筆記本。<br>There are 3 main steps in my approach (Fig 1): (我的方法有 3 個主要步驟（圖 1）： </span>\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Step 1</b>: Data Normalization(第 1 步：數據規範化</span>\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Step 2</b>: Landmarks Reduction(第 2 步：減少地標</span>\n\n<span style=\"font-size:18px; font-family:Georgia;\"><b>Step 3</b>: Frames Interpolation(第 3 步：幀插值</span>\n\n![](https://drive.google.com/uc?id=1UxXL7aTg_YpSo8-nPV019hk6rLmXZRiz)\n\n","metadata":{}},{"cell_type":"code","source":"# Clean representation for frame idx map inspired by Darien Schettler [C3]\nIDX_MAP = {\"contours\"       : list(set(CONTOURS)),\n           \"left_hand\"      : np.arange(468, 489).tolist(),\n           \"upper_body\"     : np.arange(489, 511).tolist(),\n           \"right_hand\"     : np.arange(522, 543).tolist()}\n\nFIXED_FRAMES = 37 # based on the above observations(# 基於上述觀察結果)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:03:00.335186Z","iopub.execute_input":"2023-05-02T18:03:00.336029Z","iopub.status.idle":"2023-05-02T18:03:00.342592Z","shell.execute_reply.started":"2023-05-02T18:03:00.335988Z","shell.execute_reply":"2023-05-02T18:03:00.341372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* 該字典將主體的不同部分映射到陣列中的相應索引範圍。","metadata":{}},{"cell_type":"code","source":"class FeaturePreprocess(nn.Module):\n    def __init__(self):\n        super().__init__()\n        \n    def forward(self, x_in):\n        n_frames = x_in.shape[0]\n\n        # Normalization to a common mean by Heng CK [C4]\n        x_in = x_in - x_in[~torch.isnan(x_in)].mean(0,keepdim=True) \n        x_in = x_in / x_in[~torch.isnan(x_in)].std(0, keepdim=True)\n\n        # Landmarks reduction\n        contours = x_in[:, IDX_MAP['contours']]\n        lhand    = x_in[:, IDX_MAP['left_hand']]\n        pose     = x_in[:, IDX_MAP['upper_body']]\n        rhand    = x_in[:, IDX_MAP['right_hand']]\n       \n        x_in = torch.cat([contours,\n                          lhand,\n                          pose,\n                          rhand], 1) # (n_frames, 192, 3)\n        \n        # Replace nan with 0 before Interpolation\n        x_in[torch.isnan(x_in)] = 0\n        \n        # Frames interpolation inspired by Robert Hatch [C2]\n        # If n_frames < k, use linear interpolation,\n        # else, use nearest neighbor interpolation\n        x_in = x_in.permute(2,1,0) #(3, 192, n_frames)\n        if n_frames < FIXED_FRAMES:\n            x_in = F.interpolate(x_in, size=(FIXED_FRAMES), mode= 'linear')\n        else:\n            x_in = F.interpolate(x_in, size=(FIXED_FRAMES), mode= 'nearest-exact')\n        \n        return x_in.permute(2,1,0)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:03:00.344003Z","iopub.execute_input":"2023-05-02T18:03:00.344378Z","iopub.status.idle":"2023-05-02T18:03:00.360574Z","shell.execute_reply.started":"2023-05-02T18:03:00.344341Z","shell.execute_reply":"2023-05-02T18:03:00.359478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"這是一個 PyTorch 模組，用於對一批輸入序列執行預處理。它將一批表示手語手勢的手部和身體地標的 3D 點序列作為輸入，並執行以下操作：\n\n1. 歸一化：減去非 NaN 元素的平均值並除以輸入張量的非 NaN 元素的標準差。\n1. 地標約簡：選擇代表身體不同部分（輪廓、左手、上半身、右手）的輸入張量的子集，並將它們連接成單個張量。\n1. NaN 處理：將輸入張量中的所有 NaN 值替換為 0。\n1. 幀插值：將序列重新採樣為固定數量的幀 （37），如果序列短於 37 幀，則通過線性插值，如果序列長於 37 幀，則通過最近鄰插值。\n\n輸出是形狀為 （batch_size， 37， 192， 3） 的張量，其中最後一個維度表示每個地標的 x、y、z 座標。","metadata":{}},{"cell_type":"code","source":"x_in = torch.tensor(load_relevant_data_subset(train_df.path[0]))\nfeature_preprocess = FeaturePreprocess()\nfeature_preprocess(x_in).shape, x_in[0]","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:03:00.361901Z","iopub.execute_input":"2023-05-02T18:03:00.362287Z","iopub.status.idle":"2023-05-02T18:03:00.518721Z","shell.execute_reply.started":"2023-05-02T18:03:00.362251Z","shell.execute_reply":"2023-05-02T18:03:00.517366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* x_in 的輸出形狀將是(37, 192, 3)","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:25px; font-family:Georgia;\"><b>Save Preprocessed Features</b></span>","metadata":{}},{"cell_type":"code","source":"# adapted and adjusted from Robert Hatch [C2]\nright_handed_signer = [26734, 28656, 25571, 62590, 29302, \n                       49445, 53618, 18796,  4718,  2044, \n                       37779, 30680]\nleft_handed_signer  = [16069, 32319, 36257, 22343, 27610, \n                       61333, 34503, 55372, ]\nboth_hands_signer   = [37055 ]\nmessy = [29302 ]\n\ndef convert_row(row, right_handed=True):\n    x = torch.tensor(load_relevant_data_subset(row[1].path))\n    x = feature_preprocess(x).cpu().numpy()\n    return x, row[1].label\n\ndef convert_and_save_data(df):\n    total = df.shape[0]\n    npdata = np.zeros((total, 37, 192 ,3))\n    nplabels = np.zeros(total)\n    for i, row in tqdm(enumerate(df.iterrows()), total=total):\n        (x,y) = convert_row(row)\n        npdata[i,:,:,:] = x\n        nplabels[i] = y\n    \n    np.save(\"feature_data.npy\", npdata)\n    np.save(\"feature_labels.npy\", nplabels)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:03:00.522170Z","iopub.execute_input":"2023-05-02T18:03:00.522546Z","iopub.status.idle":"2023-05-02T18:03:00.531999Z","shell.execute_reply.started":"2023-05-02T18:03:00.522512Z","shell.execute_reply":"2023-05-02T18:03:00.530845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"該函數接收包含訓練數據的數據幀，並將數據轉換為適合用於訓練機器學習模型的 numpy 數位格式。對於輸入數據幀中的每一行，該函數調用從視頻檔載入相關數據子集的函數，並對數據應用預處理步驟，例如規範化和插值。然後將生成的預處理數據與其相應的標籤一起以 numpy 陣組格式保存。該函數不返回任何內容，並將預處理的數據和標籤分別作為“feature_data.npy”和“feature_labels.npy”保存到磁碟。\n\n列表分別指定哪些簽名者是右手、左手、雙手或雜亂簽名。這些清單用作函數的輸入，以在輸入資料框中創建列。該函數獲取輸入數據幀的一行，並返回預處理的數據及其相應的標籤。","metadata":{}},{"cell_type":"code","source":"convert_and_save_data(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:03:00.533417Z","iopub.execute_input":"2023-05-02T18:03:00.533754Z","iopub.status.idle":"2023-05-02T18:46:26.429278Z","shell.execute_reply.started":"2023-05-02T18:03:00.533720Z","shell.execute_reply":"2023-05-02T18:46:26.426781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X = np.load(\"feature_data.npy\")\ny = np.load(\"feature_labels.npy\")\nprint(X.shape, y.shape)\nprint(X[0].shape,y[0])","metadata":{"execution":{"iopub.status.busy":"2023-05-02T18:46:26.438765Z","iopub.execute_input":"2023-05-02T18:46:26.439338Z","iopub.status.idle":"2023-05-02T18:47:52.941911Z","shell.execute_reply.started":"2023-05-02T18:46:26.439276Z","shell.execute_reply":"2023-05-02T18:47:52.940314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"輸出顯示載入的特徵資料和標籤陣列的形狀，以及特徵數據的第一個實例及其相應標籤的形狀：\n\n(880, 37, 192, 3) (880,)\n(37, 192, 3) 2.0\n\n這表示要素數據中有 880 個實例，每個實例有 37 幀 192 個 3D 地標點。\n這第二個 print 語句顯示特徵數據的第一個實例的形狀，即 37 幀 192 個 3D 地標點。此實例的相應標籤為 2.0，表示具有左撇子支配地位的簽名者。","metadata":{}},{"cell_type":"markdown","source":"<div style=\"background-color:white;font-size:22px;font-family:Georgia;border-style: solid;border-color: #0F9D58;border-width:5px;padding:20px;margin: 0px;color:black;overflow:hidden\">\n<h4><center><br>If you find this notebook useful, do give me an upvote, it motivates me a lot.<br>This notebook is still a work in progress. Keep checking for further developments!😊<br> Thank you!</center></h4>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">Acknowledgement</div>\n\n<span style=\"font-size:18px; font-family:Georgia;\">I want to thank <a href=\"https://www.kaggle.com/dschettler8845\">Darien Schettler</a>, <a href=\"https://www.kaggle.com/roberthatch\">Robert Hatch</a>, <a href=\"https://www.kaggle.com/ted0071\">Roland Abel</a>, <a href=\"https://www.kaggle.com/hengck23\">Heng CK</a>, and many other Kagglers for their contribution to this competition. Their work has helped me better understand the data and how to take it to the next step.</span>","metadata":{}},{"cell_type":"markdown","source":"# <div style=\"padding:20px;color:white;margin:0;font-size:30px;font-family:Georgia;text-align:left;display:fill;border-radius:5px;background-color:#4285F4;overflow:hidden\">References</div>\n\n**General References (G):**\n\n[G1] [GISLR: Data Card](https://www.kaggle.com/competitions/asl-signs/overview/data-card)\n\n[G2] [The 5 Parameters of ASL](https://www.mtsac.edu/llc/passportrewards/languagepartners/5ParametersofASL.pdf)\n\n[G3] [MediaPipe holistic model](https://google.github.io/mediapipe/solutions/holistic.html)\n\n[G4] [GISLR: Dataset Description](https://www.kaggle.com/competitions/asl-signs/data)\n\n[G5] [Discussion: X-Y coordinates aren't normalized to 0-1](https://www.kaggle.com/competitions/asl-signs/discussion/392286)\n\n[G6] [Mediapipe: Tensors to Face Landmarks](https://github.com/google/mediapipe/blob/master/mediapipe/modules/face_landmark/tensors_to_face_landmarks_with_attention.pbtxt)\n\n[G7] [Mediapipe: Facemesh Connections](https://github.com/google/mediapipe/blob/master/mediapipe/python/solutions/face_mesh_connections.py)\n\n**Code References (C):**\n\n[C1] [Roland Abel: GISLR | Visualization 🤟](https://www.kaggle.com/code/ted0071/gislr-visualization)\n\n[C2] [Robert Hatch: GISLR Feature Data: On the Shoulders](https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders/notebook)\n\n[C3] [Darien Schettler: 🤟 GISLR 🤟 - 📚Learn – 🔭EDA – 🤖Baseline](https://www.kaggle.com/code/dschettler8845/gislr-learn-eda-baseline)\n\n[C4] [Heng CK: [LB 0.62] pytorch transformer solution](https://www.kaggle.com/code/hengck23/lb-0-62-pytorch-transformer-solution/notebook)\n","metadata":{}}]}