{
  "id": 406411,
  "title": "8th place solution ... close but no cigar",
  "url": "/competitions/asl-signs/writeups/darragh-8th-place-solution-close-but-no-cigar",
  "author_name": "",
  "post_date": "2023-05-02T22:11:42.467Z",
  "votes": 29,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Here is a quick overview of the 8th place solution. </p>\n<p>3 transformers models, 2 layers each (384 hidden, 512 hidden ffn), with an ffn encoder (512-&gt;384), trained from scratch. LR 8e-4 with cosine schedule trained for ~300 epochs, dropout 0.1, batch size 1024, label smoothing 0.1. Using hands, lips and pose (above waist only). On one transformer all pose and a subset of lips were used for diversity.  </p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>most important was sequence cutout. On each sample, and each body part (left hand, right hand, lips, pose) with a 0.4 proba convert to nan 5 random slices of 0.15 x SequenceLength. It was hard to overfit with this in.</li>\n<li>mirror left</li>\n<li>random rotate. </li>\n</ul>\n<p><strong>Preprocessing</strong></p>\n<ul>\n<li>Linear interpolation of longer sequences to max length of 96.  </li>\n<li>Normalise each body part, using min max - I found this better than mean/std. In one model I used mean/std for diversity. </li>\n<li>Create time shift delta features on a subset of points, using time shifts of <code>[1, 2, 3, 4, 6, 8, 12, 16]</code>. It was important to forward fill NANs as opposed to 0-fill. </li>\n<li>Angle features between xyz position, and xy positions, on a few different joints - corner of mouth, hand, arms. </li>\n<li>Some point to point distances, interestingly this did not help much. </li>\n</ul>\n<p><strong>Tflite</strong><br>\nTraining was all in pytorch. Converting to tflite, the feature preprocessing was rewritten in tensorflow, and the base models were converted via onnx to tf (this turned out to be a mistake from looking at #2 solution, I should have rewritten the transformer encoder). It was a great opportunity to learn tensorflow. I like it. <br>\nA funny thing… doing the following on the pytorch model <code>model = model.half().float()</code> before converting to onnx, gave a good speed up in the final tflite inference. I tried quantizing the pytorch model and it did not further increase the speed. </p>\n<p>In the last week I found in preprocessing, normalising points across the whole sequence channel wise gave a good lift. Particularly what I tried to do was for each body part get the channel wise range (max-min), then get the average range over all frames, and use this to normalise the time deltas and the raw coordinates - as opposed to normalising each frame independently. Every time I submitted this it gave a very low score after ~25 mins. I rewrote it in different ways a number of times and it worked with tfilte in kaggle kernels but not when submitted. It was very frustrating at the end. However, I do not think the boost would have brought me into the prize range which was the goal 🤑</p>\n<p><strong>Failed attempts</strong></p>\n<ul>\n<li>Mixup worked ( worked for #9 team), I tried it on the embeddings and it worked ok, . See comments below on mixup.</li>\n<li>CNNs with mixup - tried efficientnet and edgenext. I tried on raw points and it did not work, looked like I had an incorrect normalisation from #2 solution. It worked well with an FFN encoder, but was too slow in inference. Looks like I missed to rewrite it into tf, or train in tf.  </li>\n</ul>",
  "messages": [
    {
      "id": "2242371",
      "postDate": "05/02/2023 08:08:30",
      "content": "<p>Here is a quick overview of the 8th place solution. </p>\n<p>3 transformers models, 2 layers each (384 hidden, 512 hidden ffn), with an ffn encoder (512-&gt;384), trained from scratch. LR 8e-4 with cosine schedule trained for ~300 epochs, dropout 0.1, batch size 1024, label smoothing 0.1. Using hands, lips and pose (above waist only). On one transformer all pose and a subset of lips were used for diversity.  </p>\n<p><strong>Augmentations</strong></p>\n<ul>\n<li>most important was sequence cutout. On each sample, and each body part (left hand, right hand, lips, pose) with a 0.4 proba convert to nan 5 random slices of 0.15 x SequenceLength. It was hard to overfit with this in.</li>\n<li>mirror left</li>\n<li>random rotate. </li>\n</ul>\n<p><strong>Preprocessing</strong></p>\n<ul>\n<li>Linear interpolation of longer sequences to max length of 96.  </li>\n<li>Normalise each body part, using min max - I found this better than mean/std. In one model I used mean/std for diversity. </li>\n<li>Create time shift delta features on a subset of points, using time shifts of <code>[1, 2, 3, 4, 6, 8, 12, 16]</code>. It was important to forward fill NANs as opposed to 0-fill. </li>\n<li>Angle features between xyz position, and xy positions, on a few different joints - corner of mouth, hand, arms. </li>\n<li>Some point to point distances, interestingly this did not help much. </li>\n</ul>\n<p><strong>Tflite</strong><br>\nTraining was all in pytorch. Converting to tflite, the feature preprocessing was rewritten in tensorflow, and the base models were converted via onnx to tf (this turned out to be a mistake from looking at #2 solution, I should have rewritten the transformer encoder). It was a great opportunity to learn tensorflow. I like it. <br>\nA funny thing… doing the following on the pytorch model <code>model = model.half().float()</code> before converting to onnx, gave a good speed up in the final tflite inference. I tried quantizing the pytorch model and it did not further increase the speed. </p>\n<p>In the last week I found in preprocessing, normalising points across the whole sequence channel wise gave a good lift. Particularly what I tried to do was for each body part get the channel wise range (max-min), then get the average range over all frames, and use this to normalise the time deltas and the raw coordinates - as opposed to normalising each frame independently. Every time I submitted this it gave a very low score after ~25 mins. I rewrote it in different ways a number of times and it worked with tfilte in kaggle kernels but not when submitted. It was very frustrating at the end. However, I do not think the boost would have brought me into the prize range which was the goal 🤑</p>\n<p><strong>Failed attempts</strong></p>\n<ul>\n<li>Mixup worked ( worked for #9 team), I tried it on the embeddings and it worked ok, . See comments below on mixup.</li>\n<li>CNNs with mixup - tried efficientnet and edgenext. I tried on raw points and it did not work, looked like I had an incorrect normalisation from #2 solution. It worked well with an FFN encoder, but was too slow in inference. Looks like I missed to rewrite it into tf, or train in tf.  </li>\n</ul>",
      "rawMarkdown": "Here is a quick overview of the 8th place solution. \n\n3 transformers models, 2 layers each (384 hidden, 512 hidden ffn), with an ffn encoder (512->384), trained from scratch. LR 8e-4 with cosine schedule trained for ~300 epochs, dropout 0.1, batch size 1024, label smoothing 0.1. Using hands, lips and pose (above waist only). On one transformer all pose and a subset of lips were used for diversity.  \n\n**Augmentations**\n- most important was sequence cutout. On each sample, and each body part (left hand, right hand, lips, pose) with a 0.4 proba convert to nan 5 random slices of 0.15 x SequenceLength. It was hard to overfit with this in.\n- mirror left\n- random rotate. \n\n**Preprocessing**\n- Linear interpolation of longer sequences to max length of 96.  \n- Normalise each body part, using min max - I found this better than mean/std. In one model I used mean/std for diversity. \n- Create time shift delta features on a subset of points, using time shifts of `[1, 2, 3, 4, 6, 8, 12, 16]`. It was important to forward fill NANs as opposed to 0-fill. \n- Angle features between xyz position, and xy positions, on a few different joints - corner of mouth, hand, arms. \n- Some point to point distances, interestingly this did not help much. \n\n**Tflite**\nTraining was all in pytorch. Converting to tflite, the feature preprocessing was rewritten in tensorflow, and the base models were converted via onnx to tf (this turned out to be a mistake from looking at #2 solution, I should have rewritten the transformer encoder). It was a great opportunity to learn tensorflow. I like it. \nA funny thing... doing the following on the pytorch model `model = model.half().float()` before converting to onnx, gave a good speed up in the final tflite inference. I tried quantizing the pytorch model and it did not further increase the speed. \n\nIn the last week I found in preprocessing, normalising points across the whole sequence channel wise gave a good lift. Particularly what I tried to do was for each body part get the channel wise range (max-min), then get the average range over all frames, and use this to normalise the time deltas and the raw coordinates - as opposed to normalising each frame independently. Every time I submitted this it gave a very low score after ~25 mins. I rewrote it in different ways a number of times and it worked with tfilte in kaggle kernels but not when submitted. It was very frustrating at the end. However, I do not think the boost would have brought me into the prize range which was the goal 🤑\n\n**Failed attempts**\n- Mixup worked (~~as mentioned by #9 team~~ worked for #9 team), I tried it on the embeddings and it worked ok, ~~but I did not finetune the result on the data without mixup, which probably would have helped~~. See comments below on mixup.\n- CNNs with mixup - tried efficientnet and edgenext. I tried on raw points and it did not work, looked like I had an incorrect normalisation from #2 solution. It worked well with an FFN encoder, but was too slow in inference. Looks like I missed to rewrite it into tf, or train in tf.",
      "votes": null
    },
    {
      "id": "2242415",
      "postDate": "05/02/2023 08:55:46",
      "content": "<p>Congratulations! Actually it is very impressive that you could find out what caused a worse result in such a short time period!</p>",
      "rawMarkdown": "Congratulations! Actually it is very impressive that you could find out what caused a worse result in such a short time period!",
      "votes": null
    },
    {
      "id": "2242424",
      "postDate": "05/02/2023 09:12:18",
      "content": "<p>We didn't fine-tune without mixup, beta was chosen to be default=1.0 (thus uniform lambda factor sampling) and it worked well, - gray curve with mixup, green curve without it (loss is BCE w/ label-smoothing 0.8)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F49828424985feacbedc991a356c4f9d1%2Fmixup_example.jpg?generation=1683018709427551&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "We didn't fine-tune without mixup, beta was chosen to be default=1.0 (thus uniform lambda factor sampling) and it worked well, - gray curve with mixup, green curve without it (loss is BCE w/ label-smoothing 0.8)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F49828424985feacbedc991a356c4f9d1%2Fmixup_example.jpg?generation=1683018709427551&alt=media)",
      "votes": null
    },
    {
      "id": "2242431",
      "postDate": "05/02/2023 09:18:23",
      "content": "<p>Label smoothing 0.8 ! 😎  … I did not try that. <br>\nUnderstood on the mixup - I tried mixing up the FFN output, I think it helped my CV but hurt public so I gave up. Your chart looks compelling though. Thanks for sharing.  </p>",
      "rawMarkdown": "Label smoothing 0.8 ! 😎  ... I did not try that. \nUnderstood on the mixup - I tried mixing up the FFN output, I think it helped my CV but hurt public so I gave up. Your chart looks compelling though. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2242460",
      "postDate": "05/02/2023 09:50:19",
      "content": "<p>Congrats with a gold medal!</p>\n<blockquote>\n  <p>looked like I had an incorrect normalisation from #2 solution</p>\n</blockquote>\n<p>Normalization for us improves with only ~0.005 LB. Maybe you weren't lucky enough with starting hyper parameters 😔</p>\n<blockquote>\n  <p>Looks like I missed to rewrite it into tf, or train in tf</p>\n</blockquote>\n<p>Also efficientnet with basic torch-ONNX-TFLite converters has 60-70ms inference. But you need to use some specific converters like <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">this</a> or <a href=\"https://github.com/MPolaris/onnx2tflite\" target=\"_blank\">this</a> because they dealing with HWC tensor. While the ordinal pipeline use transposes before each Conv layer which make it veeery slow. (You can check the graph on netron)</p>",
      "rawMarkdown": "Congrats with a gold medal!\n\n> looked like I had an incorrect normalisation from #2 solution\n\nNormalization for us improves with only ~0.005 LB. Maybe you weren't lucky enough with starting hyper parameters 😔\n\n>  Looks like I missed to rewrite it into tf, or train in tf\n\nAlso efficientnet with basic torch-ONNX-TFLite converters has 60-70ms inference. But you need to use some specific converters like [this](https://github.com/AlexanderLutsenko/nobuco) or [this](https://github.com/MPolaris/onnx2tflite) because they dealing with HWC tensor. While the ordinal pipeline use transposes before each Conv layer which make it veeery slow. (You can check the graph on netron)",
      "votes": null
    },
    {
      "id": "2242504",
      "postDate": "05/02/2023 10:24:48",
      "content": "<p>Looking at the logs for efficientnet-b1 it did as well as the transformers using 1e-3 lr with mixup; but this was passing points thru an ffn first which has extra overhead. If I remember correctly when I submitted it was getting 3 or 4 samples per second which was way too slow. Will know next time if I am doing this work with onnx to check netron and the links you mentioned. I gave up too early here. </p>",
      "rawMarkdown": "Looking at the logs for efficientnet-b1 it did as well as the transformers using 1e-3 lr with mixup; but this was passing points thru an ffn first which has extra overhead. If I remember correctly when I submitted it was getting 3 or 4 samples per second which was way too slow. Will know next time if I am doing this work with onnx to check netron and the links you mentioned. I gave up too early here.",
      "votes": null
    },
    {
      "id": "2242821",
      "postDate": "05/02/2023 14:29:24",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> congratulations with gold medal!🎉 Thank you for sharing explanation of your solution - in particular it very helpful for someone like me to improve skills!👍</p>",
      "rawMarkdown": "darraghdog congratulations with gold medal!🎉 Thank you for sharing explanation of your solution - in particular it very helpful for someone like me to improve skills!👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2242415,
      "author_name": "skybookreader",
      "author_url": "",
      "post_date": "05/02/2023 08:55:46",
      "content": "<p>Congratulations! Actually it is very impressive that you could find out what caused a worse result in such a short time period!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2242424,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "05/02/2023 09:12:18",
      "content": "<p>We didn't fine-tune without mixup, beta was chosen to be default=1.0 (thus uniform lambda factor sampling) and it worked well, - gray curve with mixup, green curve without it (loss is BCE w/ label-smoothing 0.8)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F49828424985feacbedc991a356c4f9d1%2Fmixup_example.jpg?generation=1683018709427551&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2242431,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/02/2023 09:18:23",
          "content": "<p>Label smoothing 0.8 ! 😎  … I did not try that. <br>\nUnderstood on the mixup - I tried mixing up the FFN output, I think it helped my CV but hurt public so I gave up. Your chart looks compelling though. Thanks for sharing.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2242460,
      "author_name": "kolyaforrat",
      "author_url": "",
      "post_date": "05/02/2023 09:50:19",
      "content": "<p>Congrats with a gold medal!</p>\n<blockquote>\n  <p>looked like I had an incorrect normalisation from #2 solution</p>\n</blockquote>\n<p>Normalization for us improves with only ~0.005 LB. Maybe you weren't lucky enough with starting hyper parameters 😔</p>\n<blockquote>\n  <p>Looks like I missed to rewrite it into tf, or train in tf</p>\n</blockquote>\n<p>Also efficientnet with basic torch-ONNX-TFLite converters has 60-70ms inference. But you need to use some specific converters like <a href=\"https://github.com/AlexanderLutsenko/nobuco\" target=\"_blank\">this</a> or <a href=\"https://github.com/MPolaris/onnx2tflite\" target=\"_blank\">this</a> because they dealing with HWC tensor. While the ordinal pipeline use transposes before each Conv layer which make it veeery slow. (You can check the graph on netron)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2242504,
          "author_name": "darraghdog",
          "author_url": "",
          "post_date": "05/02/2023 10:24:48",
          "content": "<p>Looking at the logs for efficientnet-b1 it did as well as the transformers using 1e-3 lr with mixup; but this was passing points thru an ffn first which has extra overhead. If I remember correctly when I submitted it was getting 3 or 4 samples per second which was way too slow. Will know next time if I am doing this work with onnx to check netron and the links you mentioned. I gave up too early here. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2242821,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/02/2023 14:29:24",
      "content": "<p><a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> congratulations with gold medal!🎉 Thank you for sharing explanation of your solution - in particular it very helpful for someone like me to improve skills!👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242371": "Here is a quick overview of the 8th place solution. \n\n3 transformers models, 2 layers each (384 hidden, 512 hidden ffn), with an ffn encoder (512->384), trained from scratch. LR 8e-4 with cosine schedule trained for ~300 epochs, dropout 0.1, batch size 1024, label smoothing 0.1. Using hands, lips and pose (above waist only). On one transformer all pose and a subset of lips were used for diversity.  \n\n**Augmentations**\n- most important was sequence cutout. On each sample, and each body part (left hand, right hand, lips, pose) with a 0.4 proba convert to nan 5 random slices of 0.15 x SequenceLength. It was hard to overfit with this in.\n- mirror left\n- random rotate. \n\n**Preprocessing**\n- Linear interpolation of longer sequences to max length of 96.  \n- Normalise each body part, using min max - I found this better than mean/std. In one model I used mean/std for diversity. \n- Create time shift delta features on a subset of points, using time shifts of `[1, 2, 3, 4, 6, 8, 12, 16]`. It was important to forward fill NANs as opposed to 0-fill. \n- Angle features between xyz position, and xy positions, on a few different joints - corner of mouth, hand, arms. \n- Some point to point distances, interestingly this did not help much. \n\n**Tflite**\nTraining was all in pytorch. Converting to tflite, the feature preprocessing was rewritten in tensorflow, and the base models were converted via onnx to tf (this turned out to be a mistake from looking at #2 solution, I should have rewritten the transformer encoder). It was a great opportunity to learn tensorflow. I like it. \nA funny thing... doing the following on the pytorch model `model = model.half().float()` before converting to onnx, gave a good speed up in the final tflite inference. I tried quantizing the pytorch model and it did not further increase the speed. \n\nIn the last week I found in preprocessing, normalising points across the whole sequence channel wise gave a good lift. Particularly what I tried to do was for each body part get the channel wise range (max-min), then get the average range over all frames, and use this to normalise the time deltas and the raw coordinates - as opposed to normalising each frame independently. Every time I submitted this it gave a very low score after ~25 mins. I rewrote it in different ways a number of times and it worked with tfilte in kaggle kernels but not when submitted. It was very frustrating at the end. However, I do not think the boost would have brought me into the prize range which was the goal 🤑\n\n**Failed attempts**\n- Mixup worked (~~as mentioned by #9 team~~ worked for #9 team), I tried it on the embeddings and it worked ok, ~~but I did not finetune the result on the data without mixup, which probably would have helped~~. See comments below on mixup.\n- CNNs with mixup - tried efficientnet and edgenext. I tried on raw points and it did not work, looked like I had an incorrect normalisation from #2 solution. It worked well with an FFN encoder, but was too slow in inference. Looks like I missed to rewrite it into tf, or train in tf.",
    "2242415": "Congratulations! Actually it is very impressive that you could find out what caused a worse result in such a short time period!",
    "2242424": "We didn't fine-tune without mixup, beta was chosen to be default=1.0 (thus uniform lambda factor sampling) and it worked well, - gray curve with mixup, green curve without it (loss is BCE w/ label-smoothing 0.8)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2F49828424985feacbedc991a356c4f9d1%2Fmixup_example.jpg?generation=1683018709427551&alt=media)",
    "2242431": "Label smoothing 0.8 ! 😎  ... I did not try that. \nUnderstood on the mixup - I tried mixing up the FFN output, I think it helped my CV but hurt public so I gave up. Your chart looks compelling though. Thanks for sharing.",
    "2242460": "Congrats with a gold medal!\n\n> looked like I had an incorrect normalisation from #2 solution\n\nNormalization for us improves with only ~0.005 LB. Maybe you weren't lucky enough with starting hyper parameters 😔\n\n>  Looks like I missed to rewrite it into tf, or train in tf\n\nAlso efficientnet with basic torch-ONNX-TFLite converters has 60-70ms inference. But you need to use some specific converters like [this](https://github.com/AlexanderLutsenko/nobuco) or [this](https://github.com/MPolaris/onnx2tflite) because they dealing with HWC tensor. While the ordinal pipeline use transposes before each Conv layer which make it veeery slow. (You can check the graph on netron)",
    "2242504": "Looking at the logs for efficientnet-b1 it did as well as the transformers using 1e-3 lr with mixup; but this was passing points thru an ffn first which has extra overhead. If I remember correctly when I submitted it was getting 3 or 4 samples per second which was way too slow. Will know next time if I am doing this work with onnx to check netron and the links you mentioned. I gave up too early here.",
    "2242821": "darraghdog congratulations with gold medal!🎉 Thank you for sharing explanation of your solution - in particular it very helpful for someone like me to improve skills!👍"
  },
  "source": "meta"
}