{
  "id": 231172,
  "title": "Effect of image sizes",
  "url": "/competitions/bms-molecular-translation/discussion/231172",
  "author_name": "",
  "post_date": "2021-04-07T09:47:58.227500800Z",
  "votes": 41,
  "comment_count": 33,
  "views": 0,
  "content": "<p>I have trained now on multiple image sizes, the results here:</p>\n<pre><code>encoder: resnet101d\ndecoder: lstm with attention\ntrain_samples: full\nepoch: 15\naugmentation: rotate90\nsize: 224x224 CV: 3.5 LB: 4.5\nsize: 288x288 CV: 3.0 LB: 3.5\nsize: 320x320 CV: 2.9 LB: 2.94 (with beam search k=4)\n</code></pre>\n<p>So training on bigger images is better as expected, the downside is slower :)</p>",
  "messages": [
    {
      "id": "1265893",
      "postDate": "04/07/2021 09:47:58",
      "content": "<p>I have trained now on multiple image sizes, the results here:</p>\n<pre><code>encoder: resnet101d\ndecoder: lstm with attention\ntrain_samples: full\nepoch: 15\naugmentation: rotate90\nsize: 224x224 CV: 3.5 LB: 4.5\nsize: 288x288 CV: 3.0 LB: 3.5\nsize: 320x320 CV: 2.9 LB: 2.94 (with beam search k=4)\n</code></pre>\n<p>So training on bigger images is better as expected, the downside is slower :)</p>",
      "rawMarkdown": "I have trained now on multiple image sizes, the results here:\n```\nencoder: resnet101d\ndecoder: lstm with attention\ntrain_samples: full\nepoch: 15\naugmentation: rotate90\nsize: 224x224 CV: 3.5 LB: 4.5\nsize: 288x288 CV: 3.0 LB: 3.5\nsize: 320x320 CV: 2.9 LB: 2.94 (with beam search k=4)\n```\n\nSo training on bigger images is better as expected, the downside is slower :)",
      "votes": null
    },
    {
      "id": "1265913",
      "postDate": "04/07/2021 10:06:01",
      "content": "<p>Thanks for sharing. Looks like it comes to computing power at the end :)</p>",
      "rawMarkdown": "Thanks for sharing. Looks like it comes to computing power at the end :)",
      "votes": null
    },
    {
      "id": "1265947",
      "postDate": "04/07/2021 11:00:40",
      "content": "<p>Are you simply resizing images or you keep the original ratio? Also how do you deal with really large images, some samples have 2048 width?</p>",
      "rawMarkdown": "Are you simply resizing images or you keep the original ratio? Also how do you deal with really large images, some samples have 2048 width?",
      "votes": null
    },
    {
      "id": "1265966",
      "postDate": "04/07/2021 11:28:43",
      "content": "<p>Currently, I am ignoring the aspect ratio :)</p>",
      "rawMarkdown": "Currently, I am ignoring the aspect ratio :)",
      "votes": null
    },
    {
      "id": "1265972",
      "postDate": "04/07/2021 11:36:14",
      "content": "<p>thanks!</p>\n<p>apart from image size, i am also interested in the number of layers in LSTM. So far in my experiments, it seems to have little effect.</p>\n<p>but i do confirm larger network is better, e.g. resnet200d is better than 101d …. but takes more time</p>\n<p>training is one issue …. inference is another. but there are good inference engine for image encoder for standard network. i think we may need torch jit from timm models</p>",
      "rawMarkdown": "thanks!\n\napart from image size, i am also interested in the number of layers in LSTM. So far in my experiments, it seems to have little effect.\n\nbut i do confirm larger network is better, e.g. resnet200d is better than 101d .... but takes more time\n\ntraining is one issue .... inference is another. but there are good inference engine for image encoder for standard network. i think we may need torch jit from timm models",
      "votes": null
    },
    {
      "id": "1265984",
      "postDate": "04/07/2021 11:44:09",
      "content": "<p>Thanks for sharing!!<br>\n<strong>What scheduler do you use?</strong><br>\nIn my settings, I'm using cosine but it takes a lot of <strong><em><em></em></em></strong><strong></strong>epochs to converge.</p>\n<p>encoder : b0<br>\ndecoder : lstm<br>\nepoch : 15<br>\nscheduler : cosine annealing<br>\nimage size : 256<br>\ncv : 4.XX</p>",
      "rawMarkdown": "Thanks for sharing!!\n**What scheduler do you use?**\nIn my settings, I'm using cosine but it takes a lot of ************epochs to converge.\n\nencoder : b0\ndecoder : lstm\nepoch : 15\nscheduler : cosine annealing\nimage size : 256\ncv : 4.XX",
      "votes": null
    },
    {
      "id": "1265988",
      "postDate": "04/07/2021 11:50:11",
      "content": "<p>The step scheduler seems to be much safer than cosine.</p>",
      "rawMarkdown": "The step scheduler seems to be much safer than cosine.",
      "votes": null
    },
    {
      "id": "1266006",
      "postDate": "04/07/2021 12:05:26",
      "content": "<p>i used lb scheduling</p>\n<ul>\n<li>start with lr 0.001. when there is no improvement make lb submission</li>\n<li>take a nap and train for a while, make lb submission again. if there is no improvement, decrease rate by 0.1. else take a nap and let it train for a while again.</li>\n</ul>",
      "rawMarkdown": "i used lb scheduling\n- start with lr 0.001. when there is no improvement make lb submission\n- take a nap and train for a while, make lb submission again. if there is no improvement, decrease rate by 0.1. else take a nap and let it train for a while again.",
      "votes": null
    },
    {
      "id": "1266012",
      "postDate": "04/07/2021 12:10:46",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> so manual step scheduler? :D</p>",
      "rawMarkdown": "hengck23 so manual step scheduler? :D",
      "votes": null
    },
    {
      "id": "1266029",
      "postDate": "04/07/2021 12:29:07",
      "content": "<p>Bigger images, models give me a headache lol</p>",
      "rawMarkdown": "Bigger images, models give me a headache lol",
      "votes": null
    },
    {
      "id": "1266058",
      "postDate": "04/07/2021 12:57:23",
      "content": "<p>I use <code>one_cycle</code>, (warm up to maximum learning rate  and then cosine decay) … Maybe I should try to use combined step/nap approach =)</p>",
      "rawMarkdown": "I use `one_cycle`, (warm up to maximum learning rate  and then cosine decay) ... Maybe I should try to use combined step/nap approach =)",
      "votes": null
    },
    {
      "id": "1266062",
      "postDate": "04/07/2021 12:58:47",
      "content": "<p>Wow! Are you using pretrained backbone or simply random initialization? So far for it seems that random works better.</p>",
      "rawMarkdown": "Wow! Are you using pretrained backbone or simply random initialization? So far for it seems that random works better.",
      "votes": null
    },
    {
      "id": "1266085",
      "postDate": "04/07/2021 13:21:20",
      "content": "<p>I do not have GPU how can I train this like this 😔</p>",
      "rawMarkdown": "I do not have GPU how can I train this like this 😔",
      "votes": null
    },
    {
      "id": "1266091",
      "postDate": "04/07/2021 13:27:57",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=0e0z28wAWfg\" target=\"_blank\">https://www.youtube.com/watch?v=0e0z28wAWfg</a> .. good old technique with pen and paper =)</p>",
      "rawMarkdown": "https://www.youtube.com/watch?v=0e0z28wAWfg .. good old technique with pen and paper =)",
      "votes": null
    },
    {
      "id": "1266111",
      "postDate": "04/07/2021 13:45:09",
      "content": "<p>High level reply 😛</p>",
      "rawMarkdown": "High level reply 😛",
      "votes": null
    },
    {
      "id": "1266118",
      "postDate": "04/07/2021 13:52:36",
      "content": "<p>Well if you get tired or bored you can earn some money on side by doing this easy steps: <a href=\"https://www.youtube.com/watch?v=y3dqhixzGVo\" target=\"_blank\">https://www.youtube.com/watch?v=y3dqhixzGVo</a> =)</p>",
      "rawMarkdown": "Well if you get tired or bored you can earn some money on side by doing this easy steps: https://www.youtube.com/watch?v=y3dqhixzGVo =)",
      "votes": null
    },
    {
      "id": "1266120",
      "postDate": "04/07/2021 13:55:09",
      "content": "<p>i train my deep network to do backprop and bitcoin mining by letting it watch youtube video</p>",
      "rawMarkdown": "i train my deep network to do backprop and bitcoin mining by letting it watch youtube video",
      "votes": null
    },
    {
      "id": "1266198",
      "postDate": "04/07/2021 14:53:07",
      "content": "<p>\"Thanks for sharing. Looks like it comes to computing power at the end :)\"</p>\n<p>I think this competition is about BOTH hardware and software stack.<br>\nyou need a good SW inference engine as well</p>",
      "rawMarkdown": "\"Thanks for sharing. Looks like it comes to computing power at the end :)\"\n\nI think this competition is about BOTH hardware and software stack.\nyou need a good SW inference engine as well",
      "votes": null
    },
    {
      "id": "1266227",
      "postDate": "04/07/2021 15:18:26",
      "content": "<p>I agree on that, this competition is good for learning some of these things from that point of view…</p>",
      "rawMarkdown": "I agree on that, this competition is good for learning some of these things from that point of view...",
      "votes": null
    },
    {
      "id": "1266603",
      "postDate": "04/07/2021 22:59:34",
      "content": "<p>fyi, transformer resnet101d 224x224 has LB 3.93/CV 3.15<br>\nno augmentation</p>",
      "rawMarkdown": "fyi, transformer resnet101d 224x224 has LB 3.93/CV 3.15\nno augmentation",
      "votes": null
    },
    {
      "id": "1266608",
      "postDate": "04/07/2021 23:12:41",
      "content": "<p>Great !<br>\nDid you manage to speed up your transformer inference ?</p>\n<p>I would likely get &lt; 4.xx on LB with my current Transformer Checkpoint. But slow ineference keeps me from submitting ^^</p>",
      "rawMarkdown": "Great !\nDid you manage to speed up your transformer inference ?\n\nI would likely get < 4.xx on LB with my current Transformer Checkpoint. But slow ineference keeps me from submitting ^^",
      "votes": null
    },
    {
      "id": "1266661",
      "postDate": "04/08/2021 01:11:50",
      "content": "<p>i am using this code: <br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>\"[2021-apr-06a]<br>\ndirty code for torch.jit.script model<br>\nif i sort the validation samples by length, i can complete 40_000 test samples in 18 min\"</p>\n<p>i sort the test samples by length predicted by the previous submission and use 4 GPUs Ti1080. <br>\nIt takes 3.2hr to run through 1.6 million test samples for me</p>",
      "rawMarkdown": "i am using this code: \nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n\n\"[2021-apr-06a]\ndirty code for torch.jit.script model\nif i sort the validation samples by length, i can complete 40_000 test samples in 18 min\"\n\ni sort the test samples by length predicted by the previous submission and use 4 GPUs Ti1080. \nIt takes 3.2hr to run through 1.6 million test samples for me",
      "votes": null
    },
    {
      "id": "1266667",
      "postDate": "04/08/2021 01:26:11",
      "content": "<p>What is the SW inference engine?</p>",
      "rawMarkdown": "What is the SW inference engine?",
      "votes": null
    },
    {
      "id": "1266686",
      "postDate": "04/08/2021 01:59:04",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , what's your batch_size of training pipeline?</p>",
      "rawMarkdown": "tugstugi , what's your batch_size of training pipeline?",
      "votes": null
    },
    {
      "id": "1267336",
      "postDate": "04/08/2021 13:27:18",
      "content": "<p>8x batch_size=64</p>",
      "rawMarkdown": "8x batch_size=64",
      "votes": null
    },
    {
      "id": "1267430",
      "postDate": "04/08/2021 14:06:18",
      "content": "<p>Thank you for sharing ! Larger image is better, it means we need more GPUs.</p>",
      "rawMarkdown": "Thank you for sharing ! Larger image is better, it means we need more GPUs.",
      "votes": null
    },
    {
      "id": "1267632",
      "postDate": "04/08/2021 16:57:48",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , <br>\nif you dont mind can you please tell what you mean by 8x? <br>\nI am assuming it is gradient accumulation steps. Am I right??</p>",
      "rawMarkdown": "tugstugi , \nif you dont mind can you please tell what you mean by 8x? \nI am assuming it is gradient accumulation steps. Am I right??",
      "votes": null
    },
    {
      "id": "1268595",
      "postDate": "04/09/2021 15:10:24",
      "content": "<p>FYI, 224 is enough to reach LB: 2.xx :)</p>",
      "rawMarkdown": "FYI, 224 is enough to reach LB: 2.xx :)",
      "votes": null
    },
    {
      "id": "1268627",
      "postDate": "04/09/2021 15:40:55",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Thanks ! Your information gives me a motivation for one more effort.</p>",
      "rawMarkdown": "yasufuminakama Thanks ! Your information gives me a motivation for one more effort.",
      "votes": null
    },
    {
      "id": "1268634",
      "postDate": "04/09/2021 15:45:55",
      "content": "<p>It's also means if you want to get 1.xx, you should use lager size. :)</p>",
      "rawMarkdown": "It's also means if you want to get 1.xx, you should use lager size. :)",
      "votes": null
    },
    {
      "id": "1268893",
      "postDate": "04/09/2021 23:07:38",
      "content": "<p>😂😂😂😂😂😂😂😂😂😂😂😂</p>",
      "rawMarkdown": "😂😂😂😂😂😂😂😂😂😂😂😂",
      "votes": null
    },
    {
      "id": "1269567",
      "postDate": "04/10/2021 16:51:32",
      "content": "<p>Probably 8 GPU cards</p>",
      "rawMarkdown": "Probably 8 GPU cards",
      "votes": null
    },
    {
      "id": "1269821",
      "postDate": "04/11/2021 00:15:56",
      "content": "<p>we definitely have to go to large images<br>\nthe question is \"do we go to large images now or after we finish experiments on small images first?\"</p>\n<p>this is a resource poblem</p>",
      "rawMarkdown": "we definitely have to go to large images\nthe question is \"do we go to large images now or after we finish experiments on small images first?\"\n\nthis is a resource poblem",
      "votes": null
    },
    {
      "id": "1269957",
      "postDate": "04/11/2021 05:52:38",
      "content": "<p>no augmentation</p>\n<p>transformer in transformer:    <br>\n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)<br>\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)</p>",
      "rawMarkdown": "no augmentation\n\ntransformer in transformer:    \n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1265913,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "04/07/2021 10:06:01",
      "content": "<p>Thanks for sharing. Looks like it comes to computing power at the end :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1266198,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 14:53:07",
          "content": "<p>\"Thanks for sharing. Looks like it comes to computing power at the end :)\"</p>\n<p>I think this competition is about BOTH hardware and software stack.<br>\nyou need a good SW inference engine as well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266227,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "04/07/2021 15:18:26",
          "content": "<p>I agree on that, this competition is good for learning some of these things from that point of view…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266667,
          "author_name": "shainirock",
          "author_url": "",
          "post_date": "04/08/2021 01:26:11",
          "content": "<p>What is the SW inference engine?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265947,
      "author_name": "skull8888888",
      "author_url": "",
      "post_date": "04/07/2021 11:00:40",
      "content": "<p>Are you simply resizing images or you keep the original ratio? Also how do you deal with really large images, some samples have 2048 width?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1265966,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 11:28:43",
          "content": "<p>Currently, I am ignoring the aspect ratio :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266062,
          "author_name": "skull8888888",
          "author_url": "",
          "post_date": "04/07/2021 12:58:47",
          "content": "<p>Wow! Are you using pretrained backbone or simply random initialization? So far for it seems that random works better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265972,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/07/2021 11:36:14",
      "content": "<p>thanks!</p>\n<p>apart from image size, i am also interested in the number of layers in LSTM. So far in my experiments, it seems to have little effect.</p>\n<p>but i do confirm larger network is better, e.g. resnet200d is better than 101d …. but takes more time</p>\n<p>training is one issue …. inference is another. but there are good inference engine for image encoder for standard network. i think we may need torch jit from timm models</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1265984,
      "author_name": "gwanghan",
      "author_url": "",
      "post_date": "04/07/2021 11:44:09",
      "content": "<p>Thanks for sharing!!<br>\n<strong>What scheduler do you use?</strong><br>\nIn my settings, I'm using cosine but it takes a lot of <strong><em><em></em></em></strong><strong></strong>epochs to converge.</p>\n<p>encoder : b0<br>\ndecoder : lstm<br>\nepoch : 15<br>\nscheduler : cosine annealing<br>\nimage size : 256<br>\ncv : 4.XX</p>",
      "votes": null,
      "replies": [
        {
          "id": 1265988,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 11:50:11",
          "content": "<p>The step scheduler seems to be much safer than cosine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266006,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 12:05:26",
          "content": "<p>i used lb scheduling</p>\n<ul>\n<li>start with lr 0.001. when there is no improvement make lb submission</li>\n<li>take a nap and train for a while, make lb submission again. if there is no improvement, decrease rate by 0.1. else take a nap and let it train for a while again.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266012,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 12:10:46",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> so manual step scheduler? :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266058,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/07/2021 12:57:23",
          "content": "<p>I use <code>one_cycle</code>, (warm up to maximum learning rate  and then cosine decay) … Maybe I should try to use combined step/nap approach =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266029,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "04/07/2021 12:29:07",
      "content": "<p>Bigger images, models give me a headache lol</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1266085,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "04/07/2021 13:21:20",
      "content": "<p>I do not have GPU how can I train this like this 😔</p>",
      "votes": null,
      "replies": [
        {
          "id": 1266091,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/07/2021 13:27:57",
          "content": "<p><a href=\"https://www.youtube.com/watch?v=0e0z28wAWfg\" target=\"_blank\">https://www.youtube.com/watch?v=0e0z28wAWfg</a> .. good old technique with pen and paper =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266111,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "04/07/2021 13:45:09",
          "content": "<p>High level reply 😛</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266118,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/07/2021 13:52:36",
          "content": "<p>Well if you get tired or bored you can earn some money on side by doing this easy steps: <a href=\"https://www.youtube.com/watch?v=y3dqhixzGVo\" target=\"_blank\">https://www.youtube.com/watch?v=y3dqhixzGVo</a> =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266120,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 13:55:09",
          "content": "<p>i train my deep network to do backprop and bitcoin mining by letting it watch youtube video</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1268893,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "04/09/2021 23:07:38",
          "content": "<p>😂😂😂😂😂😂😂😂😂😂😂😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266603,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/07/2021 22:59:34",
      "content": "<p>fyi, transformer resnet101d 224x224 has LB 3.93/CV 3.15<br>\nno augmentation</p>",
      "votes": null,
      "replies": [
        {
          "id": 1266608,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "04/07/2021 23:12:41",
          "content": "<p>Great !<br>\nDid you manage to speed up your transformer inference ?</p>\n<p>I would likely get &lt; 4.xx on LB with my current Transformer Checkpoint. But slow ineference keeps me from submitting ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266661,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/08/2021 01:11:50",
          "content": "<p>i am using this code: <br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>\n<p>\"[2021-apr-06a]<br>\ndirty code for torch.jit.script model<br>\nif i sort the validation samples by length, i can complete 40_000 test samples in 18 min\"</p>\n<p>i sort the test samples by length predicted by the previous submission and use 4 GPUs Ti1080. <br>\nIt takes 3.2hr to run through 1.6 million test samples for me</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266686,
      "author_name": "shainirock",
      "author_url": "",
      "post_date": "04/08/2021 01:59:04",
      "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , what's your batch_size of training pipeline?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1267336,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/08/2021 13:27:18",
          "content": "<p>8x batch_size=64</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1267632,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "04/08/2021 16:57:48",
          "content": "<p><a href=\"https://www.kaggle.com/tugstugi\" target=\"_blank\">@tugstugi</a> , <br>\nif you dont mind can you please tell what you mean by 8x? <br>\nI am assuming it is gradient accumulation steps. Am I right??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1269567,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "04/10/2021 16:51:32",
          "content": "<p>Probably 8 GPU cards</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1267430,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "04/08/2021 14:06:18",
      "content": "<p>Thank you for sharing ! Larger image is better, it means we need more GPUs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1268595,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "04/09/2021 15:10:24",
          "content": "<p>FYI, 224 is enough to reach LB: 2.xx :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1268627,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "04/09/2021 15:40:55",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Thanks ! Your information gives me a motivation for one more effort.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1268634,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "04/09/2021 15:45:55",
          "content": "<p>It's also means if you want to get 1.xx, you should use lager size. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1269821,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/11/2021 00:15:56",
          "content": "<p>we definitely have to go to large images<br>\nthe question is \"do we go to large images now or after we finish experiments on small images first?\"</p>\n<p>this is a resource poblem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1269957,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/11/2021 05:52:38",
          "content": "<p>no augmentation</p>\n<p>transformer in transformer:    <br>\n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)<br>\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1265893": "I have trained now on multiple image sizes, the results here:\n```\nencoder: resnet101d\ndecoder: lstm with attention\ntrain_samples: full\nepoch: 15\naugmentation: rotate90\nsize: 224x224 CV: 3.5 LB: 4.5\nsize: 288x288 CV: 3.0 LB: 3.5\nsize: 320x320 CV: 2.9 LB: 2.94 (with beam search k=4)\n```\n\nSo training on bigger images is better as expected, the downside is slower :)",
    "1265913": "Thanks for sharing. Looks like it comes to computing power at the end :)",
    "1265947": "Are you simply resizing images or you keep the original ratio? Also how do you deal with really large images, some samples have 2048 width?",
    "1265966": "Currently, I am ignoring the aspect ratio :)",
    "1265972": "thanks!\n\napart from image size, i am also interested in the number of layers in LSTM. So far in my experiments, it seems to have little effect.\n\nbut i do confirm larger network is better, e.g. resnet200d is better than 101d .... but takes more time\n\ntraining is one issue .... inference is another. but there are good inference engine for image encoder for standard network. i think we may need torch jit from timm models",
    "1265984": "Thanks for sharing!!\n**What scheduler do you use?**\nIn my settings, I'm using cosine but it takes a lot of ************epochs to converge.\n\nencoder : b0\ndecoder : lstm\nepoch : 15\nscheduler : cosine annealing\nimage size : 256\ncv : 4.XX",
    "1265988": "The step scheduler seems to be much safer than cosine.",
    "1266006": "i used lb scheduling\n- start with lr 0.001. when there is no improvement make lb submission\n- take a nap and train for a while, make lb submission again. if there is no improvement, decrease rate by 0.1. else take a nap and let it train for a while again.",
    "1266012": "hengck23 so manual step scheduler? :D",
    "1266029": "Bigger images, models give me a headache lol",
    "1266058": "I use `one_cycle`, (warm up to maximum learning rate  and then cosine decay) ... Maybe I should try to use combined step/nap approach =)",
    "1266062": "Wow! Are you using pretrained backbone or simply random initialization? So far for it seems that random works better.",
    "1266085": "I do not have GPU how can I train this like this 😔",
    "1266091": "https://www.youtube.com/watch?v=0e0z28wAWfg .. good old technique with pen and paper =)",
    "1266111": "High level reply 😛",
    "1266118": "Well if you get tired or bored you can earn some money on side by doing this easy steps: https://www.youtube.com/watch?v=y3dqhixzGVo =)",
    "1266120": "i train my deep network to do backprop and bitcoin mining by letting it watch youtube video",
    "1266198": "\"Thanks for sharing. Looks like it comes to computing power at the end :)\"\n\nI think this competition is about BOTH hardware and software stack.\nyou need a good SW inference engine as well",
    "1266227": "I agree on that, this competition is good for learning some of these things from that point of view...",
    "1266603": "fyi, transformer resnet101d 224x224 has LB 3.93/CV 3.15\nno augmentation",
    "1266608": "Great !\nDid you manage to speed up your transformer inference ?\n\nI would likely get < 4.xx on LB with my current Transformer Checkpoint. But slow ineference keeps me from submitting ^^",
    "1266661": "i am using this code: \nhttps://www.kaggle.com/c/bms-molecular-translation/discussion/231190\n\n\"[2021-apr-06a]\ndirty code for torch.jit.script model\nif i sort the validation samples by length, i can complete 40_000 test samples in 18 min\"\n\ni sort the test samples by length predicted by the previous submission and use 4 GPUs Ti1080. \nIt takes 3.2hr to run through 1.6 million test samples for me",
    "1266667": "What is the SW inference engine?",
    "1266686": "tugstugi , what's your batch_size of training pipeline?",
    "1267336": "8x batch_size=64",
    "1267430": "Thank you for sharing ! Larger image is better, it means we need more GPUs.",
    "1267632": "tugstugi , \nif you dont mind can you please tell what you mean by 8x? \nI am assuming it is gradient accumulation steps. Am I right??",
    "1268595": "FYI, 224 is enough to reach LB: 2.xx :)",
    "1268627": "yasufuminakama Thanks ! Your information gives me a motivation for one more effort.",
    "1268634": "It's also means if you want to get 1.xx, you should use lager size. :)",
    "1268893": "😂😂😂😂😂😂😂😂😂😂😂😂",
    "1269567": "Probably 8 GPU cards",
    "1269821": "we definitely have to go to large images\nthe question is \"do we go to large images now or after we finish experiments on small images first?\"\n\nthis is a resource poblem",
    "1269957": "no augmentation\n\ntransformer in transformer:    \n     - TNT-S-224-16: LB 3.11/CV 1.96 (inference = 80 min on 4xti1080)\n     - TNT-S-320-16: LB 2.27/CV 1.55 (inference = 140 min on 4xti1080)"
  },
  "source": "meta"
}