{
  "id": 406733,
  "title": "30th Place Solution - ChatGPT can be helpful for quick idea realization",
  "url": "/competitions/asl-signs/discussion/406733",
  "author_name": "Rob Mulla",
  "post_date": "2023-05-03T15:23:01.677000",
  "votes": 22,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Big thanks to Kaggle and the organizers for hosting such an engaging and challenging competition. Congratulations to all the top finishers!</p>\n<p>Throughout the competition, I experimented with various setups, conducted hundreds of experiments and still had numerous other ideas that I wanted to try. My best solution ended up being inspired by this <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">fantastic public notebook</a> similar to others.</p>\n<p>To summarize, my solution involved:</p>\n<ul>\n<li>Ensemble of two models with distinct feature sets and landmarks and parameters. Diversity in the ensemble helped a lot.</li>\n<li>Removed extremely noisy examples based on OOF scores. I attempted various thresholding methods, but it was challenging to determine if it was helpful since it naturally increased local CV (and LB) but could have resulted in overfitting.</li>\n<li>Lips, pose, and eye features, palm-normalized finger angle features, distance features (hands only), and motion features.</li>\n<li>Two sets of feature normalization: at the landmark level and globally.</li>\n<li>For hands I found including the Z-dimension were useful, epsically for distance features. All other landmarks used only x,y.</li>\n<li>Language-aware label smoothing for similar words, in addition to heavy label smoothing. I determined these by finding common false positives that appeared to be similar and then watching youtube videos of the signs.</li>\n</ul>\n<pre><code>    [98, 23, 67], # Goose, Bird, Duck\n    [233, 11], # Wake, awake\n    [167, 168], # Pen, Pencil\n    [111, 135], # Hear, Listen\n    [193, 136], # See, Look\n    [37, 249], # Carrot, Snack\n    [252,217], # Think / Hair\n    [6, 7], # Another, Any\n    [154, 137], # Noisy, Loud\n    [150, 199, 200], # Nap, Sleep, Sleepy\n    [179, 58], # Puppy, Dog\n    [95, 92], # Give, Gift\n    [149, 134], # Mouth, Lips\n    [191, 43], # Say, Chin\n    [192, 53], # Scissors, cut\n</code></pre>\n<ul>\n<li>Three-stage scheduled late dropout.</li>\n<li>Parallel transformers trained on the same embeddings.</li>\n<li>Tuned all parameters- biggest thing was finding a good balance between model size and performance, as larger models always performed better but would quickly reach the submission time limit.</li>\n</ul>\n<p>My best single-fold submission achieved a 0.76 public LB score, but my best ensemble couldn’t break 0.78 (It was very close at 0.7791).</p>\n<p>One of the challenges I faced was trying to fit a third model into my ensemble and still have it finish within the 1-hour time limit, as the runtime seemed to vary randomly. Committing the same ensemble in a notebook (tested on 4000 training examples) would finish in 3 minutes or 5 minutes depending on the run. Most of my final submissions ended up timing out and I never got the boost of the 3rd model. I became very familiar with using the <a href=\"https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\" target=\"_blank\">tflite benchmark tool</a> to find where the bottlenecks were. I should have planned better and started testing my final ensemble earlier.</p>\n<p>Things that didn’t work:</p>\n<ul>\n<li>Splitting very long examples into sub-examples based on large gaps in hand landmarks. I had a feeling these were actually multiple signs where the signer accidentally recorded in a single video.</li>\n<li>Focusing on the 6 frame examples (~12% of the training dataset contained exactly 6 frames). I'm guessing this is when the signer tapped the record button instead of holding it down.</li>\n<li>Merging similar targets during training and then applying a random prediction between the two at inference time.</li>\n<li>Model pruning.</li>\n<li>GCN models.</li>\n<li>Many other things.</li>\n</ul>\n<p>Something new I tried for the first time in this competition was consulting ChatGPT to help implement ideas. It was horrible with suggesting novel ideas, but it did provide decent starter code for some of my concepts and was able to modify existing code. For example:</p>\n<ul>\n<li>Creating a custom loss function for similar words by using a penalty matrix.</li>\n<li>Modifying my processing layers to allow for a batch.</li>\n<li>Suggesting changes to my custom preprocessing layers to not use use non-custom ops.</li>\n<li>Explaining the differences between specific pytorch and tensorflow/keras syntax.</li>\n</ul>\n<p>Again, a big congratulations to the winners and thank everyone who shared their great ideas, particularly <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a></p>",
  "messages": [
    {
      "id": 2244350,
      "postDate": "2023-05-03T15:23:01.677Z",
      "content": "<p>Big thanks to Kaggle and the organizers for hosting such an engaging and challenging competition. Congratulations to all the top finishers!</p>\n<p>Throughout the competition, I experimented with various setups, conducted hundreds of experiments and still had numerous other ideas that I wanted to try. My best solution ended up being inspired by this <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">fantastic public notebook</a> similar to others.</p>\n<p>To summarize, my solution involved:</p>\n<ul>\n<li>Ensemble of two models with distinct feature sets and landmarks and parameters. Diversity in the ensemble helped a lot.</li>\n<li>Removed extremely noisy examples based on OOF scores. I attempted various thresholding methods, but it was challenging to determine if it was helpful since it naturally increased local CV (and LB) but could have resulted in overfitting.</li>\n<li>Lips, pose, and eye features, palm-normalized finger angle features, distance features (hands only), and motion features.</li>\n<li>Two sets of feature normalization: at the landmark level and globally.</li>\n<li>For hands I found including the Z-dimension were useful, epsically for distance features. All other landmarks used only x,y.</li>\n<li>Language-aware label smoothing for similar words, in addition to heavy label smoothing. I determined these by finding common false positives that appeared to be similar and then watching youtube videos of the signs.</li>\n</ul>\n<pre><code>    [98, 23, 67], # Goose, Bird, Duck\n    [233, 11], # Wake, awake\n    [167, 168], # Pen, Pencil\n    [111, 135], # Hear, Listen\n    [193, 136], # See, Look\n    [37, 249], # Carrot, Snack\n    [252,217], # Think / Hair\n    [6, 7], # Another, Any\n    [154, 137], # Noisy, Loud\n    [150, 199, 200], # Nap, Sleep, Sleepy\n    [179, 58], # Puppy, Dog\n    [95, 92], # Give, Gift\n    [149, 134], # Mouth, Lips\n    [191, 43], # Say, Chin\n    [192, 53], # Scissors, cut\n</code></pre>\n<ul>\n<li>Three-stage scheduled late dropout.</li>\n<li>Parallel transformers trained on the same embeddings.</li>\n<li>Tuned all parameters- biggest thing was finding a good balance between model size and performance, as larger models always performed better but would quickly reach the submission time limit.</li>\n</ul>\n<p>My best single-fold submission achieved a 0.76 public LB score, but my best ensemble couldn’t break 0.78 (It was very close at 0.7791).</p>\n<p>One of the challenges I faced was trying to fit a third model into my ensemble and still have it finish within the 1-hour time limit, as the runtime seemed to vary randomly. Committing the same ensemble in a notebook (tested on 4000 training examples) would finish in 3 minutes or 5 minutes depending on the run. Most of my final submissions ended up timing out and I never got the boost of the 3rd model. I became very familiar with using the <a href=\"https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark\" target=\"_blank\">tflite benchmark tool</a> to find where the bottlenecks were. I should have planned better and started testing my final ensemble earlier.</p>\n<p>Things that didn’t work:</p>\n<ul>\n<li>Splitting very long examples into sub-examples based on large gaps in hand landmarks. I had a feeling these were actually multiple signs where the signer accidentally recorded in a single video.</li>\n<li>Focusing on the 6 frame examples (~12% of the training dataset contained exactly 6 frames). I'm guessing this is when the signer tapped the record button instead of holding it down.</li>\n<li>Merging similar targets during training and then applying a random prediction between the two at inference time.</li>\n<li>Model pruning.</li>\n<li>GCN models.</li>\n<li>Many other things.</li>\n</ul>\n<p>Something new I tried for the first time in this competition was consulting ChatGPT to help implement ideas. It was horrible with suggesting novel ideas, but it did provide decent starter code for some of my concepts and was able to modify existing code. For example:</p>\n<ul>\n<li>Creating a custom loss function for similar words by using a penalty matrix.</li>\n<li>Modifying my processing layers to allow for a batch.</li>\n<li>Suggesting changes to my custom preprocessing layers to not use use non-custom ops.</li>\n<li>Explaining the differences between specific pytorch and tensorflow/keras syntax.</li>\n</ul>\n<p>Again, a big congratulations to the winners and thank everyone who shared their great ideas, particularly <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a></p>",
      "rawMarkdown": "Big thanks to Kaggle and the organizers for hosting such an engaging and challenging competition. Congratulations to all the top finishers!\n\nThroughout the competition, I experimented with various setups, conducted hundreds of experiments and still had numerous other ideas that I wanted to try. My best solution ended up being inspired by this [fantastic public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training) similar to others.\n\nTo summarize, my solution involved:\n- Ensemble of two models with distinct feature sets and landmarks and parameters. Diversity in the ensemble helped a lot.\n- Removed extremely noisy examples based on OOF scores. I attempted various thresholding methods, but it was challenging to determine if it was helpful since it naturally increased local CV (and LB) but could have resulted in overfitting.\n- Lips, pose, and eye features, palm-normalized finger angle features, distance features (hands only), and motion features.\n- Two sets of feature normalization: at the landmark level and globally.\n- For hands I found including the Z-dimension were useful, epsically for distance features. All other landmarks used only x,y.\n- Language-aware label smoothing for similar words, in addition to heavy label smoothing. I determined these by finding common false positives that appeared to be similar and then watching youtube videos of the signs.\n```\n    [98, 23, 67], # Goose, Bird, Duck\n    [233, 11], # Wake, awake\n    [167, 168], # Pen, Pencil\n    [111, 135], # Hear, Listen\n    [193, 136], # See, Look\n    [37, 249], # Carrot, Snack\n    [252,217], # Think / Hair\n    [6, 7], # Another, Any\n    [154, 137], # Noisy, Loud\n    [150, 199, 200], # Nap, Sleep, Sleepy\n    [179, 58], # Puppy, Dog\n    [95, 92], # Give, Gift\n    [149, 134], # Mouth, Lips\n    [191, 43], # Say, Chin\n    [192, 53], # Scissors, cut\n```\n- Three-stage scheduled late dropout.\n- Parallel transformers trained on the same embeddings.\n- Tuned all parameters- biggest thing was finding a good balance between model size and performance, as larger models always performed better but would quickly reach the submission time limit.\n\nMy best single-fold submission achieved a 0.76 public LB score, but my best ensemble couldn’t break 0.78 (It was very close at 0.7791).\n\nOne of the challenges I faced was trying to fit a third model into my ensemble and still have it finish within the 1-hour time limit, as the runtime seemed to vary randomly. Committing the same ensemble in a notebook (tested on 4000 training examples) would finish in 3 minutes or 5 minutes depending on the run. Most of my final submissions ended up timing out and I never got the boost of the 3rd model. I became very familiar with using the [tflite benchmark tool](https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark) to find where the bottlenecks were. I should have planned better and started testing my final ensemble earlier.\n\nThings that didn’t work:\n- Splitting very long examples into sub-examples based on large gaps in hand landmarks. I had a feeling these were actually multiple signs where the signer accidentally recorded in a single video.\n- Focusing on the 6 frame examples (~12% of the training dataset contained exactly 6 frames). I'm guessing this is when the signer tapped the record button instead of holding it down.\n- Merging similar targets during training and then applying a random prediction between the two at inference time.\n- Model pruning.\n- GCN models.\n- Many other things.\n\nSomething new I tried for the first time in this competition was consulting ChatGPT to help implement ideas. It was horrible with suggesting novel ideas, but it did provide decent starter code for some of my concepts and was able to modify existing code. For example:\n\n- Creating a custom loss function for similar words by using a penalty matrix.\n- Modifying my processing layers to allow for a batch.\n- Suggesting changes to my custom preprocessing layers to not use use non-custom ops.\n- Explaining the differences between specific pytorch and tensorflow/keras syntax.\n\nAgain, a big congratulations to the winners and thank everyone who shared their great ideas, particularly @hengck23 and @markwijkhuizen\n",
      "votes": 22
    },
    {
      "id": 2246921,
      "postDate": "2023-05-05T14:45:00.863Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2259900,
          "postDate": "2023-05-15T10:03:41.820Z",
          "content": "<p>The submissions were required to be quantized to a tflite model.</p>",
          "rawMarkdown": "The submissions were required to be quantized to a tflite model."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2246921,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-05T14:45:00.863000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2259900,
          "author_name": "Swapnil Tripathi",
          "author_url": "",
          "post_date": "2023-05-15T10:03:41.820000",
          "content": "<p>The submissions were required to be quantized to a tflite model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2244350": "Big thanks to Kaggle and the organizers for hosting such an engaging and challenging competition. Congratulations to all the top finishers!\n\nThroughout the competition, I experimented with various setups, conducted hundreds of experiments and still had numerous other ideas that I wanted to try. My best solution ended up being inspired by this [fantastic public notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training) similar to others.\n\nTo summarize, my solution involved:\n- Ensemble of two models with distinct feature sets and landmarks and parameters. Diversity in the ensemble helped a lot.\n- Removed extremely noisy examples based on OOF scores. I attempted various thresholding methods, but it was challenging to determine if it was helpful since it naturally increased local CV (and LB) but could have resulted in overfitting.\n- Lips, pose, and eye features, palm-normalized finger angle features, distance features (hands only), and motion features.\n- Two sets of feature normalization: at the landmark level and globally.\n- For hands I found including the Z-dimension were useful, epsically for distance features. All other landmarks used only x,y.\n- Language-aware label smoothing for similar words, in addition to heavy label smoothing. I determined these by finding common false positives that appeared to be similar and then watching youtube videos of the signs.\n```\n    [98, 23, 67], # Goose, Bird, Duck\n    [233, 11], # Wake, awake\n    [167, 168], # Pen, Pencil\n    [111, 135], # Hear, Listen\n    [193, 136], # See, Look\n    [37, 249], # Carrot, Snack\n    [252,217], # Think / Hair\n    [6, 7], # Another, Any\n    [154, 137], # Noisy, Loud\n    [150, 199, 200], # Nap, Sleep, Sleepy\n    [179, 58], # Puppy, Dog\n    [95, 92], # Give, Gift\n    [149, 134], # Mouth, Lips\n    [191, 43], # Say, Chin\n    [192, 53], # Scissors, cut\n```\n- Three-stage scheduled late dropout.\n- Parallel transformers trained on the same embeddings.\n- Tuned all parameters- biggest thing was finding a good balance between model size and performance, as larger models always performed better but would quickly reach the submission time limit.\n\nMy best single-fold submission achieved a 0.76 public LB score, but my best ensemble couldn’t break 0.78 (It was very close at 0.7791).\n\nOne of the challenges I faced was trying to fit a third model into my ensemble and still have it finish within the 1-hour time limit, as the runtime seemed to vary randomly. Committing the same ensemble in a notebook (tested on 4000 training examples) would finish in 3 minutes or 5 minutes depending on the run. Most of my final submissions ended up timing out and I never got the boost of the 3rd model. I became very familiar with using the [tflite benchmark tool](https://github.com/tensorflow/tensorflow/tree/master/tensorflow/lite/tools/benchmark) to find where the bottlenecks were. I should have planned better and started testing my final ensemble earlier.\n\nThings that didn’t work:\n- Splitting very long examples into sub-examples based on large gaps in hand landmarks. I had a feeling these were actually multiple signs where the signer accidentally recorded in a single video.\n- Focusing on the 6 frame examples (~12% of the training dataset contained exactly 6 frames). I'm guessing this is when the signer tapped the record button instead of holding it down.\n- Merging similar targets during training and then applying a random prediction between the two at inference time.\n- Model pruning.\n- GCN models.\n- Many other things.\n\nSomething new I tried for the first time in this competition was consulting ChatGPT to help implement ideas. It was horrible with suggesting novel ideas, but it did provide decent starter code for some of my concepts and was able to modify existing code. For example:\n\n- Creating a custom loss function for similar words by using a penalty matrix.\n- Modifying my processing layers to allow for a batch.\n- Suggesting changes to my custom preprocessing layers to not use use non-custom ops.\n- Explaining the differences between specific pytorch and tensorflow/keras syntax.\n\nAgain, a big congratulations to the winners and thank everyone who shared their great ideas, particularly @hengck23 and @markwijkhuizen\n",
    "2246921": ""
  }
}