{
  "id": 102796,
  "title": "Unable to Improve with Old Data",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102796",
  "author_name": "",
  "post_date": "2019-08-05T05:24:22.901613800Z",
  "votes": 9,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I find this very confusing. My current highest LB score is trained on only the new competition data, with B4 and TTA. When I pretrain on old data, my LB score decreases. Looking through discussions, I notice that the agreed-upon best strategy is pretraining on old data, then finetuning on new. However, no matter what combination I've tried, I am never able to beat my new-data-only score of .792. I do slightly worse (~.77-.78).</p>\n\n<p>I use Ben's preprocessing, img size 256 for both new and old images. The old images I use are the pre-cropped and resized ones. I've even tried the suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. </p>\n\n<p>I don't believe that using exclusively new data is the best way to go. There must be something I'm doing very wrong. Any help would be appreciated.</p>",
  "messages": [
    {
      "id": "592270",
      "postDate": "08/05/2019 05:24:22",
      "content": "<p>I find this very confusing. My current highest LB score is trained on only the new competition data, with B4 and TTA. When I pretrain on old data, my LB score decreases. Looking through discussions, I notice that the agreed-upon best strategy is pretraining on old data, then finetuning on new. However, no matter what combination I've tried, I am never able to beat my new-data-only score of .792. I do slightly worse (~.77-.78).</p>\n\n<p>I use Ben's preprocessing, img size 256 for both new and old images. The old images I use are the pre-cropped and resized ones. I've even tried the suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. </p>\n\n<p>I don't believe that using exclusively new data is the best way to go. There must be something I'm doing very wrong. Any help would be appreciated.</p>",
      "rawMarkdown": "I find this very confusing. My current highest LB score is trained on only the new competition data, with B4 and TTA. When I pretrain on old data, my LB score decreases. Looking through discussions, I notice that the agreed-upon best strategy is pretraining on old data, then finetuning on new. However, no matter what combination I've tried, I am never able to beat my new-data-only score of .792. I do slightly worse (~.77-.78).\n\nI use Ben's preprocessing, img size 256 for both new and old images. The old images I use are the pre-cropped and resized ones. I've even tried the suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. \n\nI don't believe that using exclusively new data is the best way to go. There must be something I'm doing very wrong. Any help would be appreciated.",
      "votes": null
    },
    {
      "id": "592327",
      "postDate": "08/05/2019 06:31:53",
      "content": "<p>I  think    I  have  found  the  same thing with  you </p>",
      "rawMarkdown": "I  think    I  have  found  the  same thing with  you",
      "votes": null
    },
    {
      "id": "592342",
      "postDate": "08/05/2019 07:14:11",
      "content": "<p>The same with you, follow the idea of two stage training my LB become worse</p>",
      "rawMarkdown": "The same with you, follow the idea of two stage training my LB become worse",
      "votes": null
    },
    {
      "id": "592343",
      "postDate": "08/05/2019 07:15:19",
      "content": "<p>but you got 55th dude</p>",
      "rawMarkdown": "but you got 55th dude",
      "votes": null
    },
    {
      "id": "592371",
      "postDate": "08/05/2019 07:54:41",
      "content": "<p>Did you try combining both old and current data and train them in one go?</p>",
      "rawMarkdown": "Did you try combining both old and current data and train them in one go?",
      "votes": null
    },
    {
      "id": "592375",
      "postDate": "08/05/2019 08:02:16",
      "content": "<p>Ive tried, no improvement how about you</p>",
      "rawMarkdown": "Ive tried, no improvement how about you",
      "votes": null
    },
    {
      "id": "592380",
      "postDate": "08/05/2019 08:04:24",
      "content": "<p>I actually have, and that does verrryyyy slightly worse.</p>",
      "rawMarkdown": "I actually have, and that does verrryyyy slightly worse.",
      "votes": null
    },
    {
      "id": "592387",
      "postDate": "08/05/2019 08:09:11",
      "content": "<p>Did you do some type of data augmentation? Because even with 37k of images model can still overfit easily.</p>",
      "rawMarkdown": "Did you do some type of data augmentation? Because even with 37k of images model can still overfit easily.",
      "votes": null
    },
    {
      "id": "592388",
      "postDate": "08/05/2019 08:11:11",
      "content": "<p>I  think  you  can  get to this ,  when  you get  55th  , you  will  find   that  how  easy  it is ,  just  keep  studying , xiongdi</p>",
      "rawMarkdown": "I  think  you  can  get to this ,  when  you get  55th  , you  will  find   that  how  easy  it is ,  just  keep  studying , xiongdi",
      "votes": null
    },
    {
      "id": "592390",
      "postDate": "08/05/2019 08:13:19",
      "content": "<p>My current setup uses Ben's to preprocess, then random rotations and flips.</p>",
      "rawMarkdown": "My current setup uses Ben's to preprocess, then random rotations and flips.",
      "votes": null
    },
    {
      "id": "592391",
      "postDate": "08/05/2019 08:14:34",
      "content": "<p>You can try something in this article. something works well.\n<a href=\"https://arxiv.org/pdf/1812.01187.pdf\">https://arxiv.org/pdf/1812.01187.pdf</a></p>",
      "rawMarkdown": "You can try something in this article. something works well.\nhttps://arxiv.org/pdf/1812.01187.pdf",
      "votes": null
    },
    {
      "id": "592393",
      "postDate": "08/05/2019 08:18:31",
      "content": "<p>I  have tried  mix up  and  cosine_lr   in  regression   but  works  worse  ,  do  you  get  a better   result? </p>",
      "rawMarkdown": "I  have tried  mix up  and  cosine_lr   in  regression   but  works  worse  ,  do  you  get  a better   result?",
      "votes": null
    },
    {
      "id": "592398",
      "postDate": "08/05/2019 08:26:01",
      "content": "<p>If you do preprocessing (in this case Ben's preprocessing), generally you would use a smaller model than you normally would without preprocessing. I think B4 might be overkill (or worse) :D</p>",
      "rawMarkdown": "If you do preprocessing (in this case Ben's preprocessing), generally you would use a smaller model than you normally would without preprocessing. I think B4 might be overkill (or worse) :D",
      "votes": null
    },
    {
      "id": "592401",
      "postDate": "08/05/2019 08:31:14",
      "content": "<p>I find that B4 outperforms B0 - but also B5. Maybe that indicates I should try something in between...</p>",
      "rawMarkdown": "I find that B4 outperforms B0 - but also B5. Maybe that indicates I should try something in between...",
      "votes": null
    },
    {
      "id": "592404",
      "postDate": "08/05/2019 08:34:29",
      "content": "<p>how about warm up lr,  or no bias decay? </p>",
      "rawMarkdown": "how about warm up lr,  or no bias decay?",
      "votes": null
    },
    {
      "id": "592405",
      "postDate": "08/05/2019 08:37:40",
      "content": "<p>agree with you ! how about b0-b3 with bens preprocess?</p>",
      "rawMarkdown": "agree with you ! how about b0-b3 with bens preprocess?",
      "votes": null
    },
    {
      "id": "592406",
      "postDate": "08/05/2019 08:38:51",
      "content": "<p>Agreed :D. Because doing heavy preprocessing enables easier learning</p>",
      "rawMarkdown": "Agreed :D. Because doing heavy preprocessing enables easier learning",
      "votes": null
    },
    {
      "id": "592408",
      "postDate": "08/05/2019 08:40:47",
      "content": "<p>is those training tricks change the LB score a lot when using the same model?</p>",
      "rawMarkdown": "is those training tricks change the LB score a lot when using the same model?",
      "votes": null
    },
    {
      "id": "592414",
      "postDate": "08/05/2019 08:46:31",
      "content": "<p>thanks daxiongdi, i will keep on going</p>",
      "rawMarkdown": "thanks daxiongdi, i will keep on going",
      "votes": null
    },
    {
      "id": "592439",
      "postDate": "08/05/2019 09:44:22",
      "content": "<p>I  used    CosineAnnealingWarmRestarts       ,  it  not as  good  as    Adam.  They  are  experiment  in  the  kernel [LB 0.777]   which  you  published   .   I   don't  know  exactly  what  you're  talking  about <code>bias  decay</code> ?         Is  it <code>weight  decay</code>?</p>",
      "rawMarkdown": "I  used    CosineAnnealingWarmRestarts       ,  it  not as  good  as    Adam.  They  are  experiment  in  the  kernel [LB 0.777]   which  you  published   .   I   don't  know  exactly  what  you're  talking  about ` bias  decay` ?         Is  it `  weight  decay`?",
      "votes": null
    },
    {
      "id": "592496",
      "postDate": "08/05/2019 11:47:19",
      "content": "<p>Warm up learning rate in this article should be like this. warm restarts may cause overfitting.\n<a href=\"https://github.com/ildoonet/pytorch-gradual-warmup-lr\">https://github.com/ildoonet/pytorch-gradual-warmup-lr</a></p>\n\n<p>Just sharing same tips does not work well right now.\n- regression \n1. Only CosineAnnealinglr\n2. Mix-up (it is unstable)\n3. suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. </p>\n\n<ul>\n<li>classification(label smoothing is useful: 0.774-&gt;0.798 TTA=10)\n<ol><li>kfolds (I don't know why)</li>\n<li>using deeper model</li></ol></li>\n</ul>",
      "rawMarkdown": "Warm up learning rate in this article should be like this. warm restarts may cause overfitting.\nhttps://github.com/ildoonet/pytorch-gradual-warmup-lr\n\nJust sharing same tips does not work well right now.\n- regression \n1. Only CosineAnnealinglr\n2. Mix-up (it is unstable)\n3. suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. \n\n- classification(label smoothing is useful: 0.774-&gt;0.798 TTA=10)\n1. kfolds (I don't know why)\n2. using deeper model",
      "votes": null
    },
    {
      "id": "592515",
      "postDate": "08/05/2019 12:08:51",
      "content": "<p>Thanks ,mate ,   if  I have  any  improvement  ,  I will inform  you </p>",
      "rawMarkdown": "Thanks ,mate ,   if  I have  any  improvement  ,  I will inform  you",
      "votes": null
    },
    {
      "id": "592532",
      "postDate": "08/05/2019 12:43:30",
      "content": "<p>Are you using old data?</p>",
      "rawMarkdown": "Are you using old data?",
      "votes": null
    },
    {
      "id": "592546",
      "postDate": "08/05/2019 13:03:26",
      "content": "<p>Go deeper will not work if current model overfits. It just cause more trouble and waste of time.</p>",
      "rawMarkdown": "Go deeper will not work if current model overfits. It just cause more trouble and waste of time.",
      "votes": null
    },
    {
      "id": "592579",
      "postDate": "08/05/2019 13:57:33",
      "content": "<p>Yes....</p>",
      "rawMarkdown": "Yes....",
      "votes": null
    },
    {
      "id": "593010",
      "postDate": "08/06/2019 05:23:23",
      "content": "<p>My score isn't the best but rn its B3 trained on new and old data, no tricks: no custom layers, no tta, very few augmentations, no finetuning on new dataset, no optimizing rounder coeficcients, etc.</p>",
      "rawMarkdown": "My score isn't the best but rn its B3 trained on new and old data, no tricks: no custom layers, no tta, very few augmentations, no finetuning on new dataset, no optimizing rounder coeficcients, etc.",
      "votes": null
    },
    {
      "id": "593125",
      "postDate": "08/06/2019 07:51:28",
      "content": "<p>Do you start with imagenet weights?</p>",
      "rawMarkdown": "Do you start with imagenet weights?",
      "votes": null
    },
    {
      "id": "594036",
      "postDate": "08/07/2019 13:32:21",
      "content": "<p>I found that even slight variations of the learning process can cause significant changes in the LB score (+/- 1%)\nIt seems true, that the B5 model is not necessarily better than B3. I got my highest score with B3.</p>",
      "rawMarkdown": "I found that even slight variations of the learning process can cause significant changes in the LB score (+/- 1%)\nIt seems true, that the B5 model is not necessarily better than B3. I got my highest score with B3.",
      "votes": null
    },
    {
      "id": "594051",
      "postDate": "08/07/2019 13:48:38",
      "content": "<p>Do  you  use  B3  with  image size:256   or  bigger ?</p>",
      "rawMarkdown": "Do  you  use  B3  with  image size:256   or  bigger ?",
      "votes": null
    },
    {
      "id": "594069",
      "postDate": "08/07/2019 14:12:01",
      "content": "<p>hi! what is the strategy of freezing all layers ? only study on top classifier layer?</p>",
      "rawMarkdown": "hi! what is the strategy of freezing all layers ? only study on top classifier layer?",
      "votes": null
    },
    {
      "id": "594219",
      "postDate": "08/07/2019 18:07:11",
      "content": "<p>yup imagenet weights</p>",
      "rawMarkdown": "yup imagenet weights",
      "votes": null
    },
    {
      "id": "594465",
      "postDate": "08/08/2019 04:26:43",
      "content": "<p>I think that the training set should do some enlargement and reduction.</p>",
      "rawMarkdown": "I think that the training set should do some enlargement and reduction.",
      "votes": null
    },
    {
      "id": "594556",
      "postDate": "08/08/2019 06:51:06",
      "content": "<p>Just curious, how do u got such high score with 2019 data only?</p>",
      "rawMarkdown": "Just curious, how do u got such high score with 2019 data only?",
      "votes": null
    },
    {
      "id": "594567",
      "postDate": "08/08/2019 07:03:44",
      "content": "<p>Just curious,  why  can't  ?</p>",
      "rawMarkdown": "Just curious,  why  can't  ?",
      "votes": null
    },
    {
      "id": "595404",
      "postDate": "08/09/2019 07:32:29",
      "content": "<p>I used CosineAnnealingWarmRestarts after I finished training with Adam.\nIt did add some improvement. But it's not a magic tool...</p>",
      "rawMarkdown": "I used CosineAnnealingWarmRestarts after I finished training with Adam.\nIt did add some improvement. But it's not a magic tool...",
      "votes": null
    },
    {
      "id": "595405",
      "postDate": "08/09/2019 07:33:15",
      "content": "<p>Can you please also some some tips, which actually works too? ;-)</p>",
      "rawMarkdown": "Can you please also some some tips, which actually works too? ;-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 592327,
      "author_name": "xujingzhao",
      "author_url": "",
      "post_date": "08/05/2019 06:31:53",
      "content": "<p>I  think    I  have  found  the  same thing with  you </p>",
      "votes": null,
      "replies": [
        {
          "id": 592343,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/05/2019 07:15:19",
          "content": "<p>but you got 55th dude</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592388,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/05/2019 08:11:11",
          "content": "<p>I  think  you  can  get to this ,  when  you get  55th  , you  will  find   that  how  easy  it is ,  just  keep  studying , xiongdi</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592414,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/05/2019 08:46:31",
          "content": "<p>thanks daxiongdi, i will keep on going</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 592342,
      "author_name": "leixiang",
      "author_url": "",
      "post_date": "08/05/2019 07:14:11",
      "content": "<p>The same with you, follow the idea of two stage training my LB become worse</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 592371,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/05/2019 07:54:41",
      "content": "<p>Did you try combining both old and current data and train them in one go?</p>",
      "votes": null,
      "replies": [
        {
          "id": 592375,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/05/2019 08:02:16",
          "content": "<p>Ive tried, no improvement how about you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592380,
          "author_name": "dreimd",
          "author_url": "",
          "post_date": "08/05/2019 08:04:24",
          "content": "<p>I actually have, and that does verrryyyy slightly worse.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592387,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/05/2019 08:09:11",
          "content": "<p>Did you do some type of data augmentation? Because even with 37k of images model can still overfit easily.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592390,
          "author_name": "dreimd",
          "author_url": "",
          "post_date": "08/05/2019 08:13:19",
          "content": "<p>My current setup uses Ben's to preprocess, then random rotations and flips.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592398,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/05/2019 08:26:01",
          "content": "<p>If you do preprocessing (in this case Ben's preprocessing), generally you would use a smaller model than you normally would without preprocessing. I think B4 might be overkill (or worse) :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592401,
          "author_name": "dreimd",
          "author_url": "",
          "post_date": "08/05/2019 08:31:14",
          "content": "<p>I find that B4 outperforms B0 - but also B5. Maybe that indicates I should try something in between...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592405,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/05/2019 08:37:40",
          "content": "<p>agree with you ! how about b0-b3 with bens preprocess?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592406,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/05/2019 08:38:51",
          "content": "<p>Agreed :D. Because doing heavy preprocessing enables easier learning</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 592391,
      "author_name": "chanhu",
      "author_url": "",
      "post_date": "08/05/2019 08:14:34",
      "content": "<p>You can try something in this article. something works well.\n<a href=\"https://arxiv.org/pdf/1812.01187.pdf\">https://arxiv.org/pdf/1812.01187.pdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 592393,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/05/2019 08:18:31",
          "content": "<p>I  have tried  mix up  and  cosine_lr   in  regression   but  works  worse  ,  do  you  get  a better   result? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592404,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/05/2019 08:34:29",
          "content": "<p>how about warm up lr,  or no bias decay? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592408,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/05/2019 08:40:47",
          "content": "<p>is those training tricks change the LB score a lot when using the same model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592439,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/05/2019 09:44:22",
          "content": "<p>I  used    CosineAnnealingWarmRestarts       ,  it  not as  good  as    Adam.  They  are  experiment  in  the  kernel [LB 0.777]   which  you  published   .   I   don't  know  exactly  what  you're  talking  about <code>bias  decay</code> ?         Is  it <code>weight  decay</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592496,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/05/2019 11:47:19",
          "content": "<p>Warm up learning rate in this article should be like this. warm restarts may cause overfitting.\n<a href=\"https://github.com/ildoonet/pytorch-gradual-warmup-lr\">https://github.com/ildoonet/pytorch-gradual-warmup-lr</a></p>\n\n<p>Just sharing same tips does not work well right now.\n- regression \n1. Only CosineAnnealinglr\n2. Mix-up (it is unstable)\n3. suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. </p>\n\n<ul>\n<li>classification(label smoothing is useful: 0.774-&gt;0.798 TTA=10)\n<ol><li>kfolds (I don't know why)</li>\n<li>using deeper model</li></ol></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592515,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/05/2019 12:08:51",
          "content": "<p>Thanks ,mate ,   if  I have  any  improvement  ,  I will inform  you </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592532,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/05/2019 12:43:30",
          "content": "<p>Are you using old data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592546,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/05/2019 13:03:26",
          "content": "<p>Go deeper will not work if current model overfits. It just cause more trouble and waste of time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592579,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/05/2019 13:57:33",
          "content": "<p>Yes....</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 594069,
          "author_name": "ggbrother",
          "author_url": "",
          "post_date": "08/07/2019 14:12:01",
          "content": "<p>hi! what is the strategy of freezing all layers ? only study on top classifier layer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 595404,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/09/2019 07:32:29",
          "content": "<p>I used CosineAnnealingWarmRestarts after I finished training with Adam.\nIt did add some improvement. But it's not a magic tool...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 595405,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/09/2019 07:33:15",
          "content": "<p>Can you please also some some tips, which actually works too? ;-)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 593010,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "08/06/2019 05:23:23",
      "content": "<p>My score isn't the best but rn its B3 trained on new and old data, no tricks: no custom layers, no tta, very few augmentations, no finetuning on new dataset, no optimizing rounder coeficcients, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 593125,
          "author_name": "dreimd",
          "author_url": "",
          "post_date": "08/06/2019 07:51:28",
          "content": "<p>Do you start with imagenet weights?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 594219,
          "author_name": "sidhanthholalkere",
          "author_url": "",
          "post_date": "08/07/2019 18:07:11",
          "content": "<p>yup imagenet weights</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 594036,
      "author_name": "nemethpeti",
      "author_url": "",
      "post_date": "08/07/2019 13:32:21",
      "content": "<p>I found that even slight variations of the learning process can cause significant changes in the LB score (+/- 1%)\nIt seems true, that the B5 model is not necessarily better than B3. I got my highest score with B3.</p>",
      "votes": null,
      "replies": [
        {
          "id": 594051,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/07/2019 13:48:38",
          "content": "<p>Do  you  use  B3  with  image size:256   or  bigger ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 594465,
      "author_name": "a18974761777",
      "author_url": "",
      "post_date": "08/08/2019 04:26:43",
      "content": "<p>I think that the training set should do some enlargement and reduction.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 594556,
      "author_name": "mikelkl",
      "author_url": "",
      "post_date": "08/08/2019 06:51:06",
      "content": "<p>Just curious, how do u got such high score with 2019 data only?</p>",
      "votes": null,
      "replies": [
        {
          "id": 594567,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "08/08/2019 07:03:44",
          "content": "<p>Just curious,  why  can't  ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "592270": "I find this very confusing. My current highest LB score is trained on only the new competition data, with B4 and TTA. When I pretrain on old data, my LB score decreases. Looking through discussions, I notice that the agreed-upon best strategy is pretraining on old data, then finetuning on new. However, no matter what combination I've tried, I am never able to beat my new-data-only score of .792. I do slightly worse (~.77-.78).\n\nI use Ben's preprocessing, img size 256 for both new and old images. The old images I use are the pre-cropped and resized ones. I've even tried the suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. \n\nI don't believe that using exclusively new data is the best way to go. There must be something I'm doing very wrong. Any help would be appreciated.",
    "592327": "I  think    I  have  found  the  same thing with  you",
    "592342": "The same with you, follow the idea of two stage training my LB become worse",
    "592343": "but you got 55th dude",
    "592371": "Did you try combining both old and current data and train them in one go?",
    "592375": "Ive tried, no improvement how about you",
    "592380": "I actually have, and that does verrryyyy slightly worse.",
    "592387": "Did you do some type of data augmentation? Because even with 37k of images model can still overfit easily.",
    "592388": "I  think  you  can  get to this ,  when  you get  55th  , you  will  find   that  how  easy  it is ,  just  keep  studying , xiongdi",
    "592390": "My current setup uses Ben's to preprocess, then random rotations and flips.",
    "592391": "You can try something in this article. something works well.\nhttps://arxiv.org/pdf/1812.01187.pdf",
    "592393": "I  have tried  mix up  and  cosine_lr   in  regression   but  works  worse  ,  do  you  get  a better   result?",
    "592398": "If you do preprocessing (in this case Ben's preprocessing), generally you would use a smaller model than you normally would without preprocessing. I think B4 might be overkill (or worse) :D",
    "592401": "I find that B4 outperforms B0 - but also B5. Maybe that indicates I should try something in between...",
    "592404": "how about warm up lr,  or no bias decay?",
    "592405": "agree with you ! how about b0-b3 with bens preprocess?",
    "592406": "Agreed :D. Because doing heavy preprocessing enables easier learning",
    "592408": "is those training tricks change the LB score a lot when using the same model?",
    "592414": "thanks daxiongdi, i will keep on going",
    "592439": "I  used    CosineAnnealingWarmRestarts       ,  it  not as  good  as    Adam.  They  are  experiment  in  the  kernel [LB 0.777]   which  you  published   .   I   don't  know  exactly  what  you're  talking  about ` bias  decay` ?         Is  it `  weight  decay`?",
    "592496": "Warm up learning rate in this article should be like this. warm restarts may cause overfitting.\nhttps://github.com/ildoonet/pytorch-gradual-warmup-lr\n\nJust sharing same tips does not work well right now.\n- regression \n1. Only CosineAnnealinglr\n2. Mix-up (it is unstable)\n3. suggested strategy of freezing all layers for 5 epochs, training for another 15 on old data, then training for 5 more on new. \n\n- classification(label smoothing is useful: 0.774-&gt;0.798 TTA=10)\n1. kfolds (I don't know why)\n2. using deeper model",
    "592515": "Thanks ,mate ,   if  I have  any  improvement  ,  I will inform  you",
    "592532": "Are you using old data?",
    "592546": "Go deeper will not work if current model overfits. It just cause more trouble and waste of time.",
    "592579": "Yes....",
    "593010": "My score isn't the best but rn its B3 trained on new and old data, no tricks: no custom layers, no tta, very few augmentations, no finetuning on new dataset, no optimizing rounder coeficcients, etc.",
    "593125": "Do you start with imagenet weights?",
    "594036": "I found that even slight variations of the learning process can cause significant changes in the LB score (+/- 1%)\nIt seems true, that the B5 model is not necessarily better than B3. I got my highest score with B3.",
    "594051": "Do  you  use  B3  with  image size:256   or  bigger ?",
    "594069": "hi! what is the strategy of freezing all layers ? only study on top classifier layer?",
    "594219": "yup imagenet weights",
    "594465": "I think that the training set should do some enlargement and reduction.",
    "594556": "Just curious, how do u got such high score with 2019 data only?",
    "594567": "Just curious,  why  can't  ?",
    "595404": "I used CosineAnnealingWarmRestarts after I finished training with Adam.\nIt did add some improvement. But it's not a magic tool...",
    "595405": "Can you please also some some tips, which actually works too? ;-)"
  },
  "source": "meta"
}