{
  "id": 406673,
  "title": "4th place solution",
  "url": "/competitions/asl-signs/writeups/ohkawa3-4th-place-solution",
  "author_name": "",
  "post_date": "2023-05-03T11:55:16.353Z",
  "votes": 32,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I would like to thank the organizers and all the competitors.   <br>\nThe overall picture of my solution is shown in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1630583%2F2d2c527274d7918f8c3a447a7ffe1863%2Foverview.png?generation=1683110270790469&amp;alt=media\" alt=\"\"></p>\n<h1>Preprocessing</h1>\n<ul>\n<li>Use XY coordinates</li>\n<li>Normalize the coordinates between the eyebrows to (0,0).</li>\n<li>Compare the number of frames detected for the right and left hands, and flip only when the number of frames detected for the right hand is less than that of the left hand (x=-x).</li>\n<li>Use XY coordinates of 21 feature points of the right hand (flip left hand) and 40 feature points of the lips, for a total of 122 dimensions.</li>\n<li>Delete frames in which the feature points of the hand have not been detected. Therefore, the number of frames for the hand and the number of frames for the lips will be different.\"</li>\n</ul>\n<h1>Modeling</h1>\n<ul>\n<li>I have constructed two types of 1D CNNs. </li>\n<li>The first is a model that classifies fixed-length sequences (1DCNN-FixLen). </li>\n<li>The second is a model that classifies variable-length sequences (1DCNN-VariableLen). </li>\n<li>Both models have a hand backbone, a lips backbone, and a head for classification. </li>\n<li>The hand and lips backbones are similar, except that the convolution dimension differs: 128 for the hand backbone and 64 for the lips backbone.</li>\n</ul>\n<h2>1DCNN-FixLen</h2>\n<ul>\n<li>Interpolate the sequence to 96. This process is done for both hand and lip feature points.</li>\n</ul>\n<h3>Backbone</h3>\n<ul>\n<li>Apply conv 11 times and max_pooling 3 times to reduce sequence length to 12 (=96/(2**3)).</li>\n<li>Apply conv to increase the dimension to 512, then apply global_max_pooling.<ul>\n<li>Changing global_max_pooling to global_avg_pooling resulted in a significant performance drop. This is probably because important features in the sequence are only included in a few frames. We believe that with global_avg_pooling, these features are averaged out and disappear.</li>\n<li>I tried other gating mechanisms such as Gated Linear Unit, but global_max_pooling was the best.</li></ul></li>\n</ul>\n<h3>Head</h3>\n<ul>\n<li>Sum the hand features (512 dim) and lips features (512 dim).</li>\n<li>Apply 6 fully connected layers (shallow version, stochastic_depth=0.1) or 18 fully connected layers (deep version, stochastic_depth=0.5).<ul>\n<li>The deep version ties the parameters of the fully connected layers to prevent model bloat.</li></ul></li>\n<li>Classify into 250 classes.</li>\n</ul>\n<h2>1DCNN-VariableLen</h2>\n<h3>Backbone</h3>\n<ul>\n<li>Apply conv (kernel_size=3) 5 times in total and conv (kernel_size=1) 6 times in total. Unlike above, pooling is not applied.<ul>\n<li>Therefore, the size of the receptive field is 11 frames. Since reducing the number of frames further worsened the situation, and increasing the number of frames worsened the situation, it seems that the meaningful sequence was about 11 frames, even if the sequence was long.</li></ul></li>\n<li>Apply global_max_pooling after increasing the dimension to 512 with conv.<ul>\n<li>During training, the output is masked by sequence length and global_max_pooling is applied. This is because 0-padding is applied to normalize by the maximum length in the batch.</li></ul></li>\n</ul>\n<h3>Head</h3>\n<ul>\n<li>Same as FixLen with 1DCNN.</li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>3-fold participant CV.</li>\n<li>300 epochs, apply SWA after 15 epochs.</li>\n<li>AdamW optimizer.</li>\n<li>ArcMarginProduct is used, but only norm normalization of features and weight is applied since margin m=0.</li>\n<li>Stochastic Depth, 0.1 or 0.5.</li>\n<li>Label Smoothing Loss, epsilon=0.5.</li>\n</ul>\n<h1>Data Augmentation</h1>\n<ul>\n<li>Randomly drop frames (p=0.3).</li>\n<li>Augment hand position, size, and angle.</li>\n<li>Sequence length input to 1DCNN-FixLen is between 64 and 128.</li>\n</ul>\n<h1>CleanLab</h1>\n<ul>\n<li>CleanLab was used to remove approximately 5,000 scenes that were considered noise.<ul>\n<li>Calculated posterior probabilities with 21-participant-fold and used filter_by=\"both\".</li></ul></li>\n<li>The model trained on the cleaned dataset increased LB in the stand-alone model, but not much when ensembling.<ul>\n<li>Is it because the data was too clean and the diversity of the models was reduced? I honestly don't know.</li></ul></li>\n<li>Also, the gap between CV and LB appeared, so I tried not to be too overconfident.<ul>\n<li>In the final submission, CleanLab was applied to only 2 of the 6 ensembles.</li></ul></li>\n</ul>\n<h1>Extra Experiments</h1>\n<p>I did some experiments, including some that weren't included in the final submission. The following two points can be made from this evaluation:</p>\n<ul>\n<li>CleanLab is effective (+0.003) (compare A and B, D and E).</li>\n<li>Large effect of ensembling FixedLen and VariableLen (+0.01) (A0+D).</li>\n<li>I couldn't get the prize just by ensembling fixedlen.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>Length</th>\n<th>head size</th>\n<th>cleanlab</th>\n<th>seed</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A0</td>\n<td>fixed</td>\n<td>deep</td>\n<td>no</td>\n<td>0</td>\n<td>0.8661</td>\n<td>0.7843</td>\n</tr>\n<tr>\n<td>A1</td>\n<td>fixed</td>\n<td>deep</td>\n<td>no</td>\n<td>1</td>\n<td>0.8682</td>\n<td>0.7842</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>fixed</td>\n<td>deep</td>\n<td>yes</td>\n<td>5</td>\n<td>0.8721</td>\n<td>0.7848</td>\n</tr>\n<tr>\n<td>B1</td>\n<td>fixed</td>\n<td>deep</td>\n<td>yes</td>\n<td>6</td>\n<td>0.8702</td>\n<td>0.7862</td>\n</tr>\n<tr>\n<td>C0</td>\n<td>fixed</td>\n<td>shallow</td>\n<td>no</td>\n<td>5</td>\n<td>0.8651</td>\n<td>0.7794</td>\n</tr>\n<tr>\n<td>C1</td>\n<td>fixed</td>\n<td>shallow</td>\n<td>no</td>\n<td>6</td>\n<td>0.8665</td>\n<td>0.7825</td>\n</tr>\n<tr>\n<td>D</td>\n<td>variable</td>\n<td>deep</td>\n<td>no</td>\n<td>25</td>\n<td>0.8653</td>\n<td>0.7812</td>\n</tr>\n<tr>\n<td>E</td>\n<td>variable</td>\n<td>deep</td>\n<td>yes</td>\n<td>35</td>\n<td>0.8688</td>\n<td>0.7840</td>\n</tr>\n<tr>\n<td>F</td>\n<td>variable</td>\n<td>shallow</td>\n<td>no</td>\n<td>150</td>\n<td>0.8647</td>\n<td>0.7794</td>\n</tr>\n<tr>\n<td>A0+A1</td>\n<td>fixed</td>\n<td>-</td>\n<td>no</td>\n<td>-</td>\n<td>0.8722</td>\n<td>0.7905</td>\n</tr>\n<tr>\n<td>B0+B1</td>\n<td>fixed</td>\n<td>-</td>\n<td>yes</td>\n<td>-</td>\n<td>0.8761</td>\n<td>0.7935</td>\n</tr>\n<tr>\n<td>A0+D</td>\n<td>-</td>\n<td>-</td>\n<td>no</td>\n<td>-</td>\n<td>0.8766</td>\n<td>0.7945</td>\n</tr>\n<tr>\n<td>fixedlen only (A0+A1+B0+B1+C0+C1)</td>\n<td>fixed</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.8774</td>\n<td>0.7962</td>\n</tr>\n<tr>\n<td>best sub(A0+B0+B1+C0+E+F)</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.8824</td>\n<td>0.7999</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2243970",
      "postDate": "05/03/2023 10:41:50",
      "content": "<p>I would like to thank the organizers and all the competitors.   <br>\nThe overall picture of my solution is shown in the figure below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1630583%2F2d2c527274d7918f8c3a447a7ffe1863%2Foverview.png?generation=1683110270790469&amp;alt=media\" alt=\"\"></p>\n<h1>Preprocessing</h1>\n<ul>\n<li>Use XY coordinates</li>\n<li>Normalize the coordinates between the eyebrows to (0,0).</li>\n<li>Compare the number of frames detected for the right and left hands, and flip only when the number of frames detected for the right hand is less than that of the left hand (x=-x).</li>\n<li>Use XY coordinates of 21 feature points of the right hand (flip left hand) and 40 feature points of the lips, for a total of 122 dimensions.</li>\n<li>Delete frames in which the feature points of the hand have not been detected. Therefore, the number of frames for the hand and the number of frames for the lips will be different.\"</li>\n</ul>\n<h1>Modeling</h1>\n<ul>\n<li>I have constructed two types of 1D CNNs. </li>\n<li>The first is a model that classifies fixed-length sequences (1DCNN-FixLen). </li>\n<li>The second is a model that classifies variable-length sequences (1DCNN-VariableLen). </li>\n<li>Both models have a hand backbone, a lips backbone, and a head for classification. </li>\n<li>The hand and lips backbones are similar, except that the convolution dimension differs: 128 for the hand backbone and 64 for the lips backbone.</li>\n</ul>\n<h2>1DCNN-FixLen</h2>\n<ul>\n<li>Interpolate the sequence to 96. This process is done for both hand and lip feature points.</li>\n</ul>\n<h3>Backbone</h3>\n<ul>\n<li>Apply conv 11 times and max_pooling 3 times to reduce sequence length to 12 (=96/(2**3)).</li>\n<li>Apply conv to increase the dimension to 512, then apply global_max_pooling.<ul>\n<li>Changing global_max_pooling to global_avg_pooling resulted in a significant performance drop. This is probably because important features in the sequence are only included in a few frames. We believe that with global_avg_pooling, these features are averaged out and disappear.</li>\n<li>I tried other gating mechanisms such as Gated Linear Unit, but global_max_pooling was the best.</li></ul></li>\n</ul>\n<h3>Head</h3>\n<ul>\n<li>Sum the hand features (512 dim) and lips features (512 dim).</li>\n<li>Apply 6 fully connected layers (shallow version, stochastic_depth=0.1) or 18 fully connected layers (deep version, stochastic_depth=0.5).<ul>\n<li>The deep version ties the parameters of the fully connected layers to prevent model bloat.</li></ul></li>\n<li>Classify into 250 classes.</li>\n</ul>\n<h2>1DCNN-VariableLen</h2>\n<h3>Backbone</h3>\n<ul>\n<li>Apply conv (kernel_size=3) 5 times in total and conv (kernel_size=1) 6 times in total. Unlike above, pooling is not applied.<ul>\n<li>Therefore, the size of the receptive field is 11 frames. Since reducing the number of frames further worsened the situation, and increasing the number of frames worsened the situation, it seems that the meaningful sequence was about 11 frames, even if the sequence was long.</li></ul></li>\n<li>Apply global_max_pooling after increasing the dimension to 512 with conv.<ul>\n<li>During training, the output is masked by sequence length and global_max_pooling is applied. This is because 0-padding is applied to normalize by the maximum length in the batch.</li></ul></li>\n</ul>\n<h3>Head</h3>\n<ul>\n<li>Same as FixLen with 1DCNN.</li>\n</ul>\n<h1>Training</h1>\n<ul>\n<li>3-fold participant CV.</li>\n<li>300 epochs, apply SWA after 15 epochs.</li>\n<li>AdamW optimizer.</li>\n<li>ArcMarginProduct is used, but only norm normalization of features and weight is applied since margin m=0.</li>\n<li>Stochastic Depth, 0.1 or 0.5.</li>\n<li>Label Smoothing Loss, epsilon=0.5.</li>\n</ul>\n<h1>Data Augmentation</h1>\n<ul>\n<li>Randomly drop frames (p=0.3).</li>\n<li>Augment hand position, size, and angle.</li>\n<li>Sequence length input to 1DCNN-FixLen is between 64 and 128.</li>\n</ul>\n<h1>CleanLab</h1>\n<ul>\n<li>CleanLab was used to remove approximately 5,000 scenes that were considered noise.<ul>\n<li>Calculated posterior probabilities with 21-participant-fold and used filter_by=\"both\".</li></ul></li>\n<li>The model trained on the cleaned dataset increased LB in the stand-alone model, but not much when ensembling.<ul>\n<li>Is it because the data was too clean and the diversity of the models was reduced? I honestly don't know.</li></ul></li>\n<li>Also, the gap between CV and LB appeared, so I tried not to be too overconfident.<ul>\n<li>In the final submission, CleanLab was applied to only 2 of the 6 ensembles.</li></ul></li>\n</ul>\n<h1>Extra Experiments</h1>\n<p>I did some experiments, including some that weren't included in the final submission. The following two points can be made from this evaluation:</p>\n<ul>\n<li>CleanLab is effective (+0.003) (compare A and B, D and E).</li>\n<li>Large effect of ensembling FixedLen and VariableLen (+0.01) (A0+D).</li>\n<li>I couldn't get the prize just by ensembling fixedlen.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>Length</th>\n<th>head size</th>\n<th>cleanlab</th>\n<th>seed</th>\n<th>Private</th>\n<th>Public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A0</td>\n<td>fixed</td>\n<td>deep</td>\n<td>no</td>\n<td>0</td>\n<td>0.8661</td>\n<td>0.7843</td>\n</tr>\n<tr>\n<td>A1</td>\n<td>fixed</td>\n<td>deep</td>\n<td>no</td>\n<td>1</td>\n<td>0.8682</td>\n<td>0.7842</td>\n</tr>\n<tr>\n<td>B0</td>\n<td>fixed</td>\n<td>deep</td>\n<td>yes</td>\n<td>5</td>\n<td>0.8721</td>\n<td>0.7848</td>\n</tr>\n<tr>\n<td>B1</td>\n<td>fixed</td>\n<td>deep</td>\n<td>yes</td>\n<td>6</td>\n<td>0.8702</td>\n<td>0.7862</td>\n</tr>\n<tr>\n<td>C0</td>\n<td>fixed</td>\n<td>shallow</td>\n<td>no</td>\n<td>5</td>\n<td>0.8651</td>\n<td>0.7794</td>\n</tr>\n<tr>\n<td>C1</td>\n<td>fixed</td>\n<td>shallow</td>\n<td>no</td>\n<td>6</td>\n<td>0.8665</td>\n<td>0.7825</td>\n</tr>\n<tr>\n<td>D</td>\n<td>variable</td>\n<td>deep</td>\n<td>no</td>\n<td>25</td>\n<td>0.8653</td>\n<td>0.7812</td>\n</tr>\n<tr>\n<td>E</td>\n<td>variable</td>\n<td>deep</td>\n<td>yes</td>\n<td>35</td>\n<td>0.8688</td>\n<td>0.7840</td>\n</tr>\n<tr>\n<td>F</td>\n<td>variable</td>\n<td>shallow</td>\n<td>no</td>\n<td>150</td>\n<td>0.8647</td>\n<td>0.7794</td>\n</tr>\n<tr>\n<td>A0+A1</td>\n<td>fixed</td>\n<td>-</td>\n<td>no</td>\n<td>-</td>\n<td>0.8722</td>\n<td>0.7905</td>\n</tr>\n<tr>\n<td>B0+B1</td>\n<td>fixed</td>\n<td>-</td>\n<td>yes</td>\n<td>-</td>\n<td>0.8761</td>\n<td>0.7935</td>\n</tr>\n<tr>\n<td>A0+D</td>\n<td>-</td>\n<td>-</td>\n<td>no</td>\n<td>-</td>\n<td>0.8766</td>\n<td>0.7945</td>\n</tr>\n<tr>\n<td>fixedlen only (A0+A1+B0+B1+C0+C1)</td>\n<td>fixed</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.8774</td>\n<td>0.7962</td>\n</tr>\n<tr>\n<td>best sub(A0+B0+B1+C0+E+F)</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.8824</td>\n<td>0.7999</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I would like to thank the organizers and all the competitors.   \nThe overall picture of my solution is shown in the figure below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1630583%2F2d2c527274d7918f8c3a447a7ffe1863%2Foverview.png?generation=1683110270790469&alt=media)\n\n# Preprocessing\n+ Use XY coordinates\n+ Normalize the coordinates between the eyebrows to (0,0).\n+ Compare the number of frames detected for the right and left hands, and flip only when the number of frames detected for the right hand is less than that of the left hand (x=-x).\n+ Use XY coordinates of 21 feature points of the right hand (flip left hand) and 40 feature points of the lips, for a total of 122 dimensions.\n+ Delete frames in which the feature points of the hand have not been detected. Therefore, the number of frames for the hand and the number of frames for the lips will be different.\"\n\n# Modeling\n+ I have constructed two types of 1D CNNs. \n+ The first is a model that classifies fixed-length sequences (1DCNN-FixLen). \n+ The second is a model that classifies variable-length sequences (1DCNN-VariableLen). \n+ Both models have a hand backbone, a lips backbone, and a head for classification. \n+ The hand and lips backbones are similar, except that the convolution dimension differs: 128 for the hand backbone and 64 for the lips backbone.\n## 1DCNN-FixLen\n+ Interpolate the sequence to 96. This process is done for both hand and lip feature points.\n### Backbone\n+ Apply conv 11 times and max_pooling 3 times to reduce sequence length to 12 (=96/(2**3)).\n+ Apply conv to increase the dimension to 512, then apply global_max_pooling.\n  + Changing global_max_pooling to global_avg_pooling resulted in a significant performance drop. This is probably because important features in the sequence are only included in a few frames. We believe that with global_avg_pooling, these features are averaged out and disappear.\n  + I tried other gating mechanisms such as Gated Linear Unit, but global_max_pooling was the best.\n### Head\n+ Sum the hand features (512 dim) and lips features (512 dim).\n+ Apply 6 fully connected layers (shallow version, stochastic_depth=0.1) or 18 fully connected layers (deep version, stochastic_depth=0.5).\n  + The deep version ties the parameters of the fully connected layers to prevent model bloat.\n+ Classify into 250 classes.\n## 1DCNN-VariableLen\n### Backbone\n+ Apply conv (kernel_size=3) 5 times in total and conv (kernel_size=1) 6 times in total. Unlike above, pooling is not applied.\n    + Therefore, the size of the receptive field is 11 frames. Since reducing the number of frames further worsened the situation, and increasing the number of frames worsened the situation, it seems that the meaningful sequence was about 11 frames, even if the sequence was long.\n+ Apply global_max_pooling after increasing the dimension to 512 with conv.\n    + During training, the output is masked by sequence length and global_max_pooling is applied. This is because 0-padding is applied to normalize by the maximum length in the batch.\n### Head\n+ Same as FixLen with 1DCNN.\n\n# Training\n+ 3-fold participant CV.\n+ 300 epochs, apply SWA after 15 epochs.\n+ AdamW optimizer.\n+ ArcMarginProduct is used, but only norm normalization of features and weight is applied since margin m=0.\n+ Stochastic Depth, 0.1 or 0.5.\n+ Label Smoothing Loss, epsilon=0.5.\n\n# Data Augmentation\n+ Randomly drop frames (p=0.3).\n+ Augment hand position, size, and angle.\n+ Sequence length input to 1DCNN-FixLen is between 64 and 128.\n\n# CleanLab\n+ CleanLab was used to remove approximately 5,000 scenes that were considered noise.\n  + Calculated posterior probabilities with 21-participant-fold and used filter_by=\"both\".\n+ The model trained on the cleaned dataset increased LB in the stand-alone model, but not much when ensembling.\n  + Is it because the data was too clean and the diversity of the models was reduced? I honestly don't know.\n+ Also, the gap between CV and LB appeared, so I tried not to be too overconfident.\n  + In the final submission, CleanLab was applied to only 2 of the 6 ensembles.\n\n# Extra Experiments\nI did some experiments, including some that weren't included in the final submission. The following two points can be made from this evaluation:\n+ CleanLab is effective (+0.003) (compare A and B, D and E).\n+ Large effect of ensembling FixedLen and VariableLen (+0.01) (A0+D).\n+ I couldn't get the prize just by ensembling fixedlen.\n\n|model|Length|head size|cleanlab|seed|Private|Public|\n|:----|:----|:----|:----|:----|:----|:----|\n|A0|fixed|deep|no|0|0.8661|0.7843|\n|A1|fixed|deep|no|1|0.8682|0.7842|\n|B0|fixed|deep|yes|5|0.8721|0.7848|\n|B1|fixed|deep|yes|6|0.8702|0.7862|\n|C0|fixed|shallow|no|5|0.8651|0.7794|\n|C1|fixed|shallow|no|6|0.8665|0.7825|\n|D|variable|deep|no|25|0.8653|0.7812|\n|E|variable|deep|yes|35|0.8688|0.7840|\n|F|variable|shallow|no|150|0.8647|0.7794|\n|A0+A1|fixed|-|no|-|0.8722|0.7905|\n|B0+B1|fixed|-|yes|-|0.8761|0.7935|\n|A0+D|-|-|no|-|0.8766|0.7945|\n|fixedlen only (A0+A1+B0+B1+C0+C1)|fixed|-|-|-|0.8774|0.7962|\n|best sub(A0+B0+B1+C0+E+F)|-|-|-|-|0.8824|0.7999|",
      "votes": null
    },
    {
      "id": "2246975",
      "postDate": "05/05/2023 15:27:39",
      "content": "<p>Hello! This is an awesome solution!<br>\nCan you tell me why the number of frames for the right hand should be less than the number of frames for the left hand?</p>",
      "rawMarkdown": "Hello! This is an awesome solution!\nCan you tell me why the number of frames for the right hand should be less than the number of frames for the left hand?",
      "votes": null
    },
    {
      "id": "2247764",
      "postDate": "05/06/2023 09:09:24",
      "content": "<p><a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a> </p>\n<p>The sign language for this competition was only one-handed.<br>\nFor this reason, I thought it would be better to focus only on the hands where sign language is being used.<br>\nWe found that in most cases, the non-signing hand detected fewer frames than the signed hand.</p>\n<p>Therefore, if there are more right hand frames than left hand frames, the right hand is used for recognition,<br>\nConversely, if there are more left-hand frames than right-hand frames, flip the X coordinate before recognition.<br>\nBy doing this, we ensure that all sequences are signed with the right hand.</p>",
      "rawMarkdown": "ivanisaev \n\nThe sign language for this competition was only one-handed.\nFor this reason, I thought it would be better to focus only on the hands where sign language is being used.\nWe found that in most cases, the non-signing hand detected fewer frames than the signed hand.\n\nTherefore, if there are more right hand frames than left hand frames, the right hand is used for recognition,\nConversely, if there are more left-hand frames than right-hand frames, flip the X coordinate before recognition.\nBy doing this, we ensure that all sequences are signed with the right hand.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2246975,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/05/2023 15:27:39",
      "content": "<p>Hello! This is an awesome solution!<br>\nCan you tell me why the number of frames for the right hand should be less than the number of frames for the left hand?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2247764,
          "author_name": "chack3",
          "author_url": "",
          "post_date": "05/06/2023 09:09:24",
          "content": "<p><a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a> </p>\n<p>The sign language for this competition was only one-handed.<br>\nFor this reason, I thought it would be better to focus only on the hands where sign language is being used.<br>\nWe found that in most cases, the non-signing hand detected fewer frames than the signed hand.</p>\n<p>Therefore, if there are more right hand frames than left hand frames, the right hand is used for recognition,<br>\nConversely, if there are more left-hand frames than right-hand frames, flip the X coordinate before recognition.<br>\nBy doing this, we ensure that all sequences are signed with the right hand.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2243970": "I would like to thank the organizers and all the competitors.   \nThe overall picture of my solution is shown in the figure below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1630583%2F2d2c527274d7918f8c3a447a7ffe1863%2Foverview.png?generation=1683110270790469&alt=media)\n\n# Preprocessing\n+ Use XY coordinates\n+ Normalize the coordinates between the eyebrows to (0,0).\n+ Compare the number of frames detected for the right and left hands, and flip only when the number of frames detected for the right hand is less than that of the left hand (x=-x).\n+ Use XY coordinates of 21 feature points of the right hand (flip left hand) and 40 feature points of the lips, for a total of 122 dimensions.\n+ Delete frames in which the feature points of the hand have not been detected. Therefore, the number of frames for the hand and the number of frames for the lips will be different.\"\n\n# Modeling\n+ I have constructed two types of 1D CNNs. \n+ The first is a model that classifies fixed-length sequences (1DCNN-FixLen). \n+ The second is a model that classifies variable-length sequences (1DCNN-VariableLen). \n+ Both models have a hand backbone, a lips backbone, and a head for classification. \n+ The hand and lips backbones are similar, except that the convolution dimension differs: 128 for the hand backbone and 64 for the lips backbone.\n## 1DCNN-FixLen\n+ Interpolate the sequence to 96. This process is done for both hand and lip feature points.\n### Backbone\n+ Apply conv 11 times and max_pooling 3 times to reduce sequence length to 12 (=96/(2**3)).\n+ Apply conv to increase the dimension to 512, then apply global_max_pooling.\n  + Changing global_max_pooling to global_avg_pooling resulted in a significant performance drop. This is probably because important features in the sequence are only included in a few frames. We believe that with global_avg_pooling, these features are averaged out and disappear.\n  + I tried other gating mechanisms such as Gated Linear Unit, but global_max_pooling was the best.\n### Head\n+ Sum the hand features (512 dim) and lips features (512 dim).\n+ Apply 6 fully connected layers (shallow version, stochastic_depth=0.1) or 18 fully connected layers (deep version, stochastic_depth=0.5).\n  + The deep version ties the parameters of the fully connected layers to prevent model bloat.\n+ Classify into 250 classes.\n## 1DCNN-VariableLen\n### Backbone\n+ Apply conv (kernel_size=3) 5 times in total and conv (kernel_size=1) 6 times in total. Unlike above, pooling is not applied.\n    + Therefore, the size of the receptive field is 11 frames. Since reducing the number of frames further worsened the situation, and increasing the number of frames worsened the situation, it seems that the meaningful sequence was about 11 frames, even if the sequence was long.\n+ Apply global_max_pooling after increasing the dimension to 512 with conv.\n    + During training, the output is masked by sequence length and global_max_pooling is applied. This is because 0-padding is applied to normalize by the maximum length in the batch.\n### Head\n+ Same as FixLen with 1DCNN.\n\n# Training\n+ 3-fold participant CV.\n+ 300 epochs, apply SWA after 15 epochs.\n+ AdamW optimizer.\n+ ArcMarginProduct is used, but only norm normalization of features and weight is applied since margin m=0.\n+ Stochastic Depth, 0.1 or 0.5.\n+ Label Smoothing Loss, epsilon=0.5.\n\n# Data Augmentation\n+ Randomly drop frames (p=0.3).\n+ Augment hand position, size, and angle.\n+ Sequence length input to 1DCNN-FixLen is between 64 and 128.\n\n# CleanLab\n+ CleanLab was used to remove approximately 5,000 scenes that were considered noise.\n  + Calculated posterior probabilities with 21-participant-fold and used filter_by=\"both\".\n+ The model trained on the cleaned dataset increased LB in the stand-alone model, but not much when ensembling.\n  + Is it because the data was too clean and the diversity of the models was reduced? I honestly don't know.\n+ Also, the gap between CV and LB appeared, so I tried not to be too overconfident.\n  + In the final submission, CleanLab was applied to only 2 of the 6 ensembles.\n\n# Extra Experiments\nI did some experiments, including some that weren't included in the final submission. The following two points can be made from this evaluation:\n+ CleanLab is effective (+0.003) (compare A and B, D and E).\n+ Large effect of ensembling FixedLen and VariableLen (+0.01) (A0+D).\n+ I couldn't get the prize just by ensembling fixedlen.\n\n|model|Length|head size|cleanlab|seed|Private|Public|\n|:----|:----|:----|:----|:----|:----|:----|\n|A0|fixed|deep|no|0|0.8661|0.7843|\n|A1|fixed|deep|no|1|0.8682|0.7842|\n|B0|fixed|deep|yes|5|0.8721|0.7848|\n|B1|fixed|deep|yes|6|0.8702|0.7862|\n|C0|fixed|shallow|no|5|0.8651|0.7794|\n|C1|fixed|shallow|no|6|0.8665|0.7825|\n|D|variable|deep|no|25|0.8653|0.7812|\n|E|variable|deep|yes|35|0.8688|0.7840|\n|F|variable|shallow|no|150|0.8647|0.7794|\n|A0+A1|fixed|-|no|-|0.8722|0.7905|\n|B0+B1|fixed|-|yes|-|0.8761|0.7935|\n|A0+D|-|-|no|-|0.8766|0.7945|\n|fixedlen only (A0+A1+B0+B1+C0+C1)|fixed|-|-|-|0.8774|0.7962|\n|best sub(A0+B0+B1+C0+E+F)|-|-|-|-|0.8824|0.7999|",
    "2246975": "Hello! This is an awesome solution!\nCan you tell me why the number of frames for the right hand should be less than the number of frames for the left hand?",
    "2247764": "ivanisaev \n\nThe sign language for this competition was only one-handed.\nFor this reason, I thought it would be better to focus only on the hands where sign language is being used.\nWe found that in most cases, the non-signing hand detected fewer frames than the signed hand.\n\nTherefore, if there are more right hand frames than left hand frames, the right hand is used for recognition,\nConversely, if there are more left-hand frames than right-hand frames, flip the X coordinate before recognition.\nBy doing this, we ensure that all sequences are signed with the right hand."
  },
  "source": "meta"
}