{
  "id": 234680,
  "title": "Sharing some results",
  "url": "/competitions/bms-molecular-translation/discussion/234680",
  "author_name": "",
  "post_date": "2021-04-25T15:22:08.068310500Z",
  "votes": 43,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I wanted to share some of my experiment results. Training time is very high for this competition and it is hard to run many tests, so I figured I would share what I have tried so far. I hope this is useful and can help guide your experiments. I will update this post as I run more tests. (Currently trying pseudo-labeling and EffNetV2 as encoder).</p>\n<p>The below results are without data augmentation. I am using square images without any sort of cropping / padding. All models were trained with cosine learning rate for 15 epochs. <code>Time</code> column is training and validation time for a single epoch. All models use a single GRU layer.</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Decoder</th>\n<th>Image size</th>\n<th>Batch size</th>\n<th>CV (no beam)</th>\n<th>CV (bs=5)</th>\n<th>LB (bs=5)</th>\n<th>Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet34</td>\n<td>GRU</td>\n<td>224</td>\n<td>256</td>\n<td>6.2229</td>\n<td>-</td>\n<td>-</td>\n<td>151 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>LSTM</td>\n<td>224</td>\n<td>256</td>\n<td>2.8435</td>\n<td>-</td>\n<td>-</td>\n<td>230 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>224</td>\n<td>256</td>\n<td>2.7913</td>\n<td>2.6509</td>\n<td>3.22</td>\n<td>228 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>256</td>\n<td>256</td>\n<td>2.3048</td>\n<td>2.2346</td>\n<td>3.19</td>\n<td>280 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>288</td>\n<td>128</td>\n<td>2.0617</td>\n<td>-</td>\n<td>-</td>\n<td>345 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>320</td>\n<td>128</td>\n<td>1.9793</td>\n<td>1.9403</td>\n<td>2.76</td>\n<td>413 mins</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, all models were trained on a Quadro RTX 8000.</p>",
  "messages": [
    {
      "id": "1284126",
      "postDate": "04/25/2021 15:22:08",
      "content": "<p>I wanted to share some of my experiment results. Training time is very high for this competition and it is hard to run many tests, so I figured I would share what I have tried so far. I hope this is useful and can help guide your experiments. I will update this post as I run more tests. (Currently trying pseudo-labeling and EffNetV2 as encoder).</p>\n<p>The below results are without data augmentation. I am using square images without any sort of cropping / padding. All models were trained with cosine learning rate for 15 epochs. <code>Time</code> column is training and validation time for a single epoch. All models use a single GRU layer.</p>\n<table>\n<thead>\n<tr>\n<th>Encoder</th>\n<th>Decoder</th>\n<th>Image size</th>\n<th>Batch size</th>\n<th>CV (no beam)</th>\n<th>CV (bs=5)</th>\n<th>LB (bs=5)</th>\n<th>Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet34</td>\n<td>GRU</td>\n<td>224</td>\n<td>256</td>\n<td>6.2229</td>\n<td>-</td>\n<td>-</td>\n<td>151 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>LSTM</td>\n<td>224</td>\n<td>256</td>\n<td>2.8435</td>\n<td>-</td>\n<td>-</td>\n<td>230 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>224</td>\n<td>256</td>\n<td>2.7913</td>\n<td>2.6509</td>\n<td>3.22</td>\n<td>228 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>256</td>\n<td>256</td>\n<td>2.3048</td>\n<td>2.2346</td>\n<td>3.19</td>\n<td>280 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>288</td>\n<td>128</td>\n<td>2.0617</td>\n<td>-</td>\n<td>-</td>\n<td>345 mins</td>\n</tr>\n<tr>\n<td>Pruned EffNetB3</td>\n<td>GRU</td>\n<td>320</td>\n<td>128</td>\n<td>1.9793</td>\n<td>1.9403</td>\n<td>2.76</td>\n<td>413 mins</td>\n</tr>\n</tbody>\n</table>\n<p>For reference, all models were trained on a Quadro RTX 8000.</p>",
      "rawMarkdown": "I wanted to share some of my experiment results. Training time is very high for this competition and it is hard to run many tests, so I figured I would share what I have tried so far. I hope this is useful and can help guide your experiments. I will update this post as I run more tests. (Currently trying pseudo-labeling and EffNetV2 as encoder).\n\nThe below results are without data augmentation. I am using square images without any sort of cropping / padding. All models were trained with cosine learning rate for 15 epochs. `Time` column is training and validation time for a single epoch. All models use a single GRU layer.\n\n| Encoder | Decoder | Image size | Batch size | CV (no beam) | CV (bs=5) | LB (bs=5) | Time |\n| -------- | --------- | ----------- | ----------- |  --- | ------------ | ------- | ------ |\n| ResNet34 | GRU | 224 | 256 | 6.2229 | - | - |151 mins|\n| Pruned EffNetB3 | LSTM| 224 | 256 | 2.8435 | - | - | 230 mins|\n| Pruned EffNetB3 | GRU| 224 | 256 | 2.7913 | 2.6509 | 3.22 | 228 mins|\n| Pruned EffNetB3 | GRU| 256| 256 | 2.3048| 2.2346| 3.19| 280 mins|\n| Pruned EffNetB3 | GRU| 288| 128 | 2.0617 | - | - | 345 mins|\n| Pruned EffNetB3 | GRU| 320| 128 | 1.9793| 1.9403| 2.76 | 413 mins| \n\nFor reference, all models were trained on a Quadro RTX 8000.",
      "votes": null
    },
    {
      "id": "1284369",
      "postDate": "04/25/2021 20:33:04",
      "content": "<p>Thanks for sharing. I guess it is better to compare the experiments based on total training time instead of the number of epochs. Therefore, it is very nice to see the epoch times in the table.</p>",
      "rawMarkdown": "Thanks for sharing. I guess it is better to compare the experiments based on total training time instead of the number of epochs. Therefore, it is very nice to see the epoch times in the table.",
      "votes": null
    },
    {
      "id": "1285102",
      "postDate": "04/26/2021 15:20:04",
      "content": "<p>Thanks for the great share. Is that single fold or full data training? </p>",
      "rawMarkdown": "Thanks for the great share. Is that single fold or full data training?",
      "votes": null
    },
    {
      "id": "1285106",
      "postDate": "04/26/2021 15:21:30",
      "content": "<p>Single fold (5 fold split stratified by sequence length). </p>",
      "rawMarkdown": "Single fold (5 fold split stratified by sequence length).",
      "votes": null
    },
    {
      "id": "1285651",
      "postDate": "04/27/2021 06:45:54",
      "content": "<p>What were your learning rates for these attempt for the encoder and decoder? I'm observing that EffNetB1 is stabilizing at around 7.6, without further decrease.</p>",
      "rawMarkdown": "What were your learning rates for these attempt for the encoder and decoder? I'm observing that EffNetB1 is stabilizing at around 7.6, without further decrease.",
      "votes": null
    },
    {
      "id": "1285937",
      "postDate": "04/27/2021 11:58:50",
      "content": "<p>I use the same initial learning rate for both encoder and decoder: <code>5e-4</code>. If I end up training with smaller batches, I will change this to <code>1e-4</code>.</p>",
      "rawMarkdown": "I use the same initial learning rate for both encoder and decoder: `5e-4`. If I end up training with smaller batches, I will change this to `1e-4`.",
      "votes": null
    },
    {
      "id": "1286545",
      "postDate": "04/28/2021 05:41:52",
      "content": "<p>my experiment: effb0 + 1 LSTM decoder layer 10 epoch 224x224 no beam search cv 6.4 LB 6.8. </p>",
      "rawMarkdown": "my experiment: effb0 + 1 LSTM decoder layer 10 epoch 224x224 no beam search cv 6.4 LB 6.8.",
      "votes": null
    },
    {
      "id": "1286583",
      "postDate": "04/28/2021 07:08:08",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for sharing.<br>\nI have some doubts----</p>\n<ul>\n<li>Which pre-trained is better (IMAGENET or noisy-student)</li>\n<li>Pruned model means compressed model without the weights which activates less?<br>\n~thanks</li>\n</ul>",
      "rawMarkdown": "Hey @tuckerarrants thanks for sharing.\nI have some doubts----\n- Which pre-trained is better (IMAGENET or noisy-student)\n- Pruned model means compressed model without the weights which activates less?\n~thanks",
      "votes": null
    },
    {
      "id": "1286707",
      "postDate": "04/28/2021 10:34:31",
      "content": "<p>I have not tried noisy student weights - I do not think there would be much difference between the two. </p>\n<p>Yes, pruned models are more computationally efficient (fewer parameters). They remove unimportant filters and their feature maps.</p>",
      "rawMarkdown": "I have not tried noisy student weights - I do not think there would be much difference between the two. \n\nYes, pruned models are more computationally efficient (fewer parameters). They remove unimportant filters and their feature maps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1284369,
      "author_name": "sorkun",
      "author_url": "",
      "post_date": "04/25/2021 20:33:04",
      "content": "<p>Thanks for sharing. I guess it is better to compare the experiments based on total training time instead of the number of epochs. Therefore, it is very nice to see the epoch times in the table.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1285102,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "04/26/2021 15:20:04",
      "content": "<p>Thanks for the great share. Is that single fold or full data training? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1285106,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "04/26/2021 15:21:30",
          "content": "<p>Single fold (5 fold split stratified by sequence length). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1285651,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "04/27/2021 06:45:54",
      "content": "<p>What were your learning rates for these attempt for the encoder and decoder? I'm observing that EffNetB1 is stabilizing at around 7.6, without further decrease.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1285937,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "04/27/2021 11:58:50",
          "content": "<p>I use the same initial learning rate for both encoder and decoder: <code>5e-4</code>. If I end up training with smaller batches, I will change this to <code>1e-4</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1286545,
      "author_name": "wantsu",
      "author_url": "",
      "post_date": "04/28/2021 05:41:52",
      "content": "<p>my experiment: effb0 + 1 LSTM decoder layer 10 epoch 224x224 no beam search cv 6.4 LB 6.8. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1286583,
      "author_name": "akhileshdkapse",
      "author_url": "",
      "post_date": "04/28/2021 07:08:08",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> thanks for sharing.<br>\nI have some doubts----</p>\n<ul>\n<li>Which pre-trained is better (IMAGENET or noisy-student)</li>\n<li>Pruned model means compressed model without the weights which activates less?<br>\n~thanks</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1286707,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "04/28/2021 10:34:31",
          "content": "<p>I have not tried noisy student weights - I do not think there would be much difference between the two. </p>\n<p>Yes, pruned models are more computationally efficient (fewer parameters). They remove unimportant filters and their feature maps.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1284126": "I wanted to share some of my experiment results. Training time is very high for this competition and it is hard to run many tests, so I figured I would share what I have tried so far. I hope this is useful and can help guide your experiments. I will update this post as I run more tests. (Currently trying pseudo-labeling and EffNetV2 as encoder).\n\nThe below results are without data augmentation. I am using square images without any sort of cropping / padding. All models were trained with cosine learning rate for 15 epochs. `Time` column is training and validation time for a single epoch. All models use a single GRU layer.\n\n| Encoder | Decoder | Image size | Batch size | CV (no beam) | CV (bs=5) | LB (bs=5) | Time |\n| -------- | --------- | ----------- | ----------- |  --- | ------------ | ------- | ------ |\n| ResNet34 | GRU | 224 | 256 | 6.2229 | - | - |151 mins|\n| Pruned EffNetB3 | LSTM| 224 | 256 | 2.8435 | - | - | 230 mins|\n| Pruned EffNetB3 | GRU| 224 | 256 | 2.7913 | 2.6509 | 3.22 | 228 mins|\n| Pruned EffNetB3 | GRU| 256| 256 | 2.3048| 2.2346| 3.19| 280 mins|\n| Pruned EffNetB3 | GRU| 288| 128 | 2.0617 | - | - | 345 mins|\n| Pruned EffNetB3 | GRU| 320| 128 | 1.9793| 1.9403| 2.76 | 413 mins| \n\nFor reference, all models were trained on a Quadro RTX 8000.",
    "1284369": "Thanks for sharing. I guess it is better to compare the experiments based on total training time instead of the number of epochs. Therefore, it is very nice to see the epoch times in the table.",
    "1285102": "Thanks for the great share. Is that single fold or full data training?",
    "1285106": "Single fold (5 fold split stratified by sequence length).",
    "1285651": "What were your learning rates for these attempt for the encoder and decoder? I'm observing that EffNetB1 is stabilizing at around 7.6, without further decrease.",
    "1285937": "I use the same initial learning rate for both encoder and decoder: `5e-4`. If I end up training with smaller batches, I will change this to `1e-4`.",
    "1286545": "my experiment: effb0 + 1 LSTM decoder layer 10 epoch 224x224 no beam search cv 6.4 LB 6.8.",
    "1286583": "Hey @tuckerarrants thanks for sharing.\nI have some doubts----\n- Which pre-trained is better (IMAGENET or noisy-student)\n- Pruned model means compressed model without the weights which activates less?\n~thanks",
    "1286707": "I have not tried noisy student weights - I do not think there would be much difference between the two. \n\nYes, pruned models are more computationally efficient (fewer parameters). They remove unimportant filters and their feature maps."
  },
  "source": "meta"
}