{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Train csv EDA\nIn this notebook, we'll perform a very quick exploration of the train csv file as it already contains a lot of useful information.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\nfrom matplotlib import pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-03-04T10:37:07.395653Z","iopub.execute_input":"2023-03-04T10:37:07.396086Z","iopub.status.idle":"2023-03-04T10:37:08.355232Z","shell.execute_reply.started":"2023-03-04T10:37:07.395998Z","shell.execute_reply":"2023-03-04T10:37:08.354043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_csv_path = \"/kaggle/input/asl-signs/train.csv\"\n\ntrain_df = pd.read_csv(train_csv_path)\ntrain_df.columns","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:49:53.525553Z","iopub.execute_input":"2023-03-02T19:49:53.526741Z","iopub.status.idle":"2023-03-02T19:49:53.709021Z","shell.execute_reply.started":"2023-03-02T19:49:53.526690Z","shell.execute_reply":"2023-03-02T19:49:53.707013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok so now we've loaded the data into a dataframe. Let's check what we can find in it.\n\n* path: First of all, in the path column, you see the path to the parquet file corresponding to a certain clip. The files contain the x,y and z coordinates of certain body landmarks from the face, body and hands. The path is a combination of the train_landmark_files folder and the participant_id and sequence_id.\n* participand_id: This is the unique id of the person that was recorded. We can already see here that there are multiple ids. We'll check later on how many there are.\n* sequence_id: A unique identifier for each sequence or clip of landmark movements.\n* sign: The label for the landmark sequence. This is the target in this challenge.\n\nWe can see we have 94477 rows so there should be 94477 unique sequences.","metadata":{}},{"cell_type":"code","source":"train_df","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:49:56.198042Z","iopub.execute_input":"2023-03-02T19:49:56.198563Z","iopub.status.idle":"2023-03-02T19:49:56.229753Z","shell.execute_reply.started":"2023-03-02T19:49:56.198525Z","shell.execute_reply":"2023-03-02T19:49:56.228341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It's usually a good idea to check for null values or nan's. In this case, we can see that there are none so the data seems complete.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:49:58.128657Z","iopub.execute_input":"2023-03-02T19:49:58.129273Z","iopub.status.idle":"2023-03-02T19:49:58.152595Z","shell.execute_reply.started":"2023-03-02T19:49:58.129226Z","shell.execute_reply":"2023-03-02T19:49:58.151204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now check all the columns in more detail.\n\n## path\nThere isn't that much to be said about this column. As said before, it contains the path to the parquet files with all the sequence data. The path structure is: train_landmark_files/[participant_id]/[sequence_id].parquet. We would assume that all the paths are unique but we check it to be sure and see that this is the case.","metadata":{}},{"cell_type":"code","source":"len(train_df[\"path\"])-len(train_df[\"path\"].drop_duplicates())","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:50:16.146673Z","iopub.execute_input":"2023-03-02T19:50:16.147225Z","iopub.status.idle":"2023-03-02T19:50:16.193981Z","shell.execute_reply.started":"2023-03-02T19:50:16.147184Z","shell.execute_reply":"2023-03-02T19:50:16.192500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## participant_id\nThis column shows the unique id of the contributor who's performance can be seen in the landmark sequence. We see that there are only 21 unique participant id's.","metadata":{}},{"cell_type":"code","source":"train_df[\"participant_id\"].nunique()","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:50:18.265429Z","iopub.execute_input":"2023-03-02T19:50:18.265930Z","iopub.status.idle":"2023-03-02T19:50:18.278694Z","shell.execute_reply.started":"2023-03-02T19:50:18.265891Z","shell.execute_reply":"2023-03-02T19:50:18.277131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's have a closer look and see how much each participant contributed. \nWe see that all contributors have well over 3000 sequences with a maximum of 4968 sequences for the participant with id 49445. This is still fairly balanced. We should check later on to see if certain participants contribute more to certain labels.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8,6))\ntrain_df[\"participant_id\"].value_counts(ascending=True).plot.barh(ax=ax)\nax.bar_label(ax.containers[0], label_type='edge')\nax.set_title(\"Number of participant_id occurrences\", fontsize=16, fontweight=\"bold\", pad=20)\nax.spines[['right', 'top']].set_visible(False)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-02T19:50:21.873922Z","iopub.execute_input":"2023-03-02T19:50:21.874833Z","iopub.status.idle":"2023-03-02T19:50:22.398916Z","shell.execute_reply.started":"2023-03-02T19:50:21.874791Z","shell.execute_reply":"2023-03-02T19:50:22.397501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## sequence_id\nThis column holds the unique identifiers for the landmark sequence. As the description suggests, we shouldn't see any duplicate values here which is the case.","metadata":{}},{"cell_type":"code","source":"len(train_df[\"path\"])-len(train_df[\"path\"].drop_duplicates())","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:50:53.198762Z","iopub.execute_input":"2023-03-02T19:50:53.199289Z","iopub.status.idle":"2023-03-02T19:50:53.229175Z","shell.execute_reply.started":"2023-03-02T19:50:53.199248Z","shell.execute_reply":"2023-03-02T19:50:53.227625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There also doesn't seem to be much logic in the ids.","metadata":{}},{"cell_type":"code","source":"print(train_df[\"sequence_id\"].min())\nprint(train_df[\"sequence_id\"].max())","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:50:54.845816Z","iopub.execute_input":"2023-03-02T19:50:54.846946Z","iopub.status.idle":"2023-03-02T19:50:54.855500Z","shell.execute_reply.started":"2023-03-02T19:50:54.846892Z","shell.execute_reply":"2023-03-02T19:50:54.854096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## sign\nThis column contains the labels. This is the target value that needs to be predicted. Here we have 250 unique values and we'll print them all so we can examine them.","metadata":{}},{"cell_type":"code","source":"nunique = train_df[\"sign\"].nunique()\nprint(f\"We have {nunique} unique labels in our dataset\")\nprint(train_df[\"sign\"].unique())","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:51:00.867426Z","iopub.execute_input":"2023-03-02T19:51:00.868492Z","iopub.status.idle":"2023-03-02T19:51:00.892689Z","shell.execute_reply.started":"2023-03-02T19:51:00.868441Z","shell.execute_reply":"2023-03-02T19:51:00.891599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see several labels that might be related to eachother like a plane, a helicopter, a bird, a duck, an owl,... or several types of persons,... It might be a good idea to check videos of all these signs online if we can find them.\n\nWe have a lot more labels here so let's quickly explore the distribution of the labels before plotting it in a bar plot.","metadata":{}},{"cell_type":"code","source":"min = train_df[\"sign\"].value_counts().min()\nmax = train_df[\"sign\"].value_counts().max()\n\nprint(f\"Listen occurs the most often in the dataset with {max} times.\")\nprint(f\"Zipper occurs the least amount of times in the dataset with only {min} times.\\n\\n\")\n\nprint(train_df[\"sign\"].value_counts())","metadata":{"execution":{"iopub.status.busy":"2023-03-02T19:51:04.359499Z","iopub.execute_input":"2023-03-02T19:51:04.360467Z","iopub.status.idle":"2023-03-02T19:51:04.386647Z","shell.execute_reply.started":"2023-03-02T19:51:04.360419Z","shell.execute_reply":"2023-03-02T19:51:04.384907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we plot the distribution of the number of occurrences, we see that it still seems fairly balanced. There are some differences but probably not something to worry about at this stage.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(10,2))\nsns.boxplot(x=train_df[\"sign\"].value_counts(), ax=ax)\nplt.xlim(0,450)\nax.set_title(\"Sign occurrence distribution\", fontsize=16, fontweight=\"bold\", pad=20)\nax.spines[[\"left\", \"right\", \"top\"]].set_visible(False)\nax.set(xlabel=None)\nplt.tick_params(left = False)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-02T19:51:05.487213Z","iopub.execute_input":"2023-03-02T19:51:05.488505Z","iopub.status.idle":"2023-03-02T19:51:05.713649Z","shell.execute_reply.started":"2023-03-02T19:51:05.488444Z","shell.execute_reply":"2023-03-02T19:51:05.712373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It won't be the prettiest graph but let's just plot a bar chart with the number of occurrences per sign for completeness.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(14,30))\ntrain_df[\"sign\"].value_counts(ascending=True).plot.barh(ax=ax)\nax.bar_label(ax.containers[0], label_type='edge', fontsize=8)\nax.set_title(\"Number of sign occurrences\", fontsize=16, fontweight=\"bold\", pad=20)\nax.spines[['right', 'top']].set_visible(False)\nax.tick_params(axis='y', which='major', labelsize=8)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-02T19:51:07.899428Z","iopub.execute_input":"2023-03-02T19:51:07.899966Z","iopub.status.idle":"2023-03-02T19:51:12.799911Z","shell.execute_reply.started":"2023-03-02T19:51:07.899920Z","shell.execute_reply":"2023-03-02T19:51:12.798435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Did all participants perform all signs?\nWe know we have only 21 participants in the training set while we have more than 299 examples of each sign. Let's check if each participant performed each sign or if there was a concentration of certain signs by certain participants.\n\nWe can see in the table bellow that most participants performed all of the 250 different signs. Only participants 25571 and 30680 did only 242 and 238 signs respectively.","metadata":{}},{"cell_type":"code","source":"train_df.groupby(\"participant_id\")[[\"sign\"]].nunique().sort_values(by=\"sign\", ascending=False)","metadata":{"execution":{"iopub.status.busy":"2023-03-02T20:02:05.299862Z","iopub.execute_input":"2023-03-02T20:02:05.300438Z","iopub.status.idle":"2023-03-02T20:02:05.345273Z","shell.execute_reply.started":"2023-03-02T20:02:05.300388Z","shell.execute_reply":"2023-03-02T20:02:05.344194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok so we know that all participants perform (almost) all signs. This can be useful for the future as we might want to cross validation during training and use each participant in turns as the validation set to make sure our model generalises to unseen participants as we don't know what we have in the test data on the leaderboard.","metadata":{}},{"cell_type":"code","source":"heatmap_df = pd.pivot_table(train_df, index=\"sign\", columns=\"participant_id\", aggfunc=\"count\", values=\"path\")\nheatmap_df","metadata":{"execution":{"iopub.status.busy":"2023-03-02T20:48:59.525416Z","iopub.execute_input":"2023-03-02T20:48:59.525906Z","iopub.status.idle":"2023-03-02T20:48:59.609868Z","shell.execute_reply.started":"2023-03-02T20:48:59.525862Z","shell.execute_reply":"2023-03-02T20:48:59.608358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So in the heatmap, it looks like taking out a single participant as a cross validation set might be a good idea as the signs would be well represented amongst the training set then.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(18,80))\nheatmap = sns.heatmap(heatmap_df,\n            cmap='coolwarm',\n            annot=True,\n            cbar=False)\nax.set_title(\"Heatmap showing number of times a sign was performed by each participant\", fontsize=20, pad=20);","metadata":{"execution":{"iopub.status.busy":"2023-03-02T20:52:17.414014Z","iopub.execute_input":"2023-03-02T20:52:17.414589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}