{
  "id": 105100,
  "title": "how to use 2015 data to improve accuracy?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/105100",
  "author_name": "",
  "post_date": "2019-08-21T07:51:28.018081100Z",
  "votes": 14,
  "comment_count": 48,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "604270",
      "postDate": "08/21/2019 07:51:28",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "604272",
      "postDate": "08/21/2019 07:52:22",
      "content": "<p>Hi All,</p>\n\n<p>I just combined 2015 and 2019 data and tried to build a model but accuracy drops as compared to only 2019 data. Any tips would be welcome. Thanks !</p>",
      "rawMarkdown": "Hi All,\n\nI just combined 2015 and 2019 data and tried to build a model but accuracy drops as compared to only 2019 data. Any tips would be welcome. Thanks !",
      "votes": null
    },
    {
      "id": "604298",
      "postDate": "08/21/2019 08:40:45",
      "content": "<p>I suggest building validation set from 2019 data only. Since 2019 training data might have closer distribution to 2019 test data.</p>",
      "rawMarkdown": "I suggest building validation set from 2019 data only. Since 2019 training data might have closer distribution to 2019 test data.",
      "votes": null
    },
    {
      "id": "604325",
      "postDate": "08/21/2019 09:16:15",
      "content": "<p>Maybe a careful study in the 2015 data, trying to find a same perfil data sounds good, after that merging it in to a big dataset for better training purposes sounds something interesting to try.</p>",
      "rawMarkdown": "Maybe a careful study in the 2015 data, trying to find a same perfil data sounds good, after that merging it in to a big dataset for better training purposes sounds something interesting to try.",
      "votes": null
    },
    {
      "id": "604401",
      "postDate": "08/21/2019 10:50:10",
      "content": "<p>great work</p>",
      "rawMarkdown": "great work",
      "votes": null
    },
    {
      "id": "604560",
      "postDate": "08/21/2019 14:14:37",
      "content": "<p><a href=\"/anuragtr\">@anuragtr</a> are you using 2015 test data? My LB lowered drastically after adding 2015 test data. I got better result using 2015 training data with current competition data and i think <a href=\"/quandapro\">@quandapro</a> suggested a good strategy for building validation data</p>",
      "rawMarkdown": "anuragtr are you using 2015 test data? My LB lowered drastically after adding 2015 test data. I got better result using 2015 training data with current competition data and i think @quandapro suggested a good strategy for building validation data",
      "votes": null
    },
    {
      "id": "604656",
      "postDate": "08/21/2019 16:09:48",
      "content": "<p>thanks <a href=\"/brunhs\">@brunhs</a> will try</p>",
      "rawMarkdown": "thanks @brunhs will try",
      "votes": null
    },
    {
      "id": "604657",
      "postDate": "08/21/2019 16:10:29",
      "content": "<p>thanks <a href=\"/quandapro\">@quandapro</a> that's a good suggestion</p>",
      "rawMarkdown": "thanks @quandapro that's a good suggestion",
      "votes": null
    },
    {
      "id": "604658",
      "postDate": "08/21/2019 16:11:13",
      "content": "<p>thanks <a href=\"/sabbiracoustic1006\">@sabbiracoustic1006</a> for the experience you had</p>",
      "rawMarkdown": "thanks @sabbiracoustic1006 for the experience you had",
      "votes": null
    },
    {
      "id": "604678",
      "postDate": "08/21/2019 16:35:40",
      "content": "<p>You can pretrain on the 2015 data then finetune on  2019 data, or you can  merging it in to a big dataset(use part of 2015 data). Two ways work for me, but the first is better. </p>",
      "rawMarkdown": "You can pretrain on the 2015 data then finetune on  2019 data, or you can  merging it in to a big dataset(use part of 2015 data). Two ways work for me, but the first is better.",
      "votes": null
    },
    {
      "id": "605109",
      "postDate": "08/22/2019 04:38:19",
      "content": "<p>thanks for the good advice <a href=\"/buaazijian\">@buaazijian</a> </p>",
      "rawMarkdown": "thanks for the good advice @buaazijian",
      "votes": null
    },
    {
      "id": "605150",
      "postDate": "08/22/2019 05:29:35",
      "content": "<p>hi, JIANJIAN. do you use both train and test of 2015 data for pretrain or only train data? pretrain on all 2015 data ( total about 88000 images) seems doesn't work for me</p>",
      "rawMarkdown": "hi, JIANJIAN. do you use both train and test of 2015 data for pretrain or only train data? pretrain on all 2015 data ( total about 88000 images) seems doesn't work for me",
      "votes": null
    },
    {
      "id": "605156",
      "postDate": "08/22/2019 05:35:29",
      "content": "<p>yes, I use both train and test and it works for me. </p>",
      "rawMarkdown": "yes, I use both train and test and it works for me.",
      "votes": null
    },
    {
      "id": "605162",
      "postDate": "08/22/2019 05:50:36",
      "content": "<p>thank you for your info</p>",
      "rawMarkdown": "thank you for your info",
      "votes": null
    },
    {
      "id": "605339",
      "postDate": "08/22/2019 09:13:58",
      "content": "<p>where is the 2015 data</p>",
      "rawMarkdown": "where is the 2015 data",
      "votes": null
    },
    {
      "id": "605360",
      "postDate": "08/22/2019 09:43:32",
      "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a></p>",
      "rawMarkdown": "https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized",
      "votes": null
    },
    {
      "id": "605431",
      "postDate": "08/22/2019 11:22:07",
      "content": "<p>I find it!, Thank you, and how many epoch to train when use previous data! sincerely appreciate！</p>",
      "rawMarkdown": "I find it!, Thank you, and how many epoch to train when use previous data! sincerely appreciate！",
      "votes": null
    },
    {
      "id": "605440",
      "postDate": "08/22/2019 11:30:48",
      "content": "<p>monitor your validation loss until it converges and stop the training. Generally it should converge within 10 epochs</p>",
      "rawMarkdown": "monitor your validation loss until it converges and stop the training. Generally it should converge within 10 epochs",
      "votes": null
    },
    {
      "id": "605467",
      "postDate": "08/22/2019 12:00:57",
      "content": "<p>the validation data also split from the previous data?</p>",
      "rawMarkdown": "the validation data also split from the previous data?",
      "votes": null
    },
    {
      "id": "605486",
      "postDate": "08/22/2019 12:16:59",
      "content": "<p>Currently i am using validation data from both this competition and previous competitions data but i am planning to use the strategy suggested by <a href=\"/quandapro\">@quandapro</a> to use only this competitions data in the validation set.</p>",
      "rawMarkdown": "Currently i am using validation data from both this competition and previous competitions data but i am planning to use the strategy suggested by @quandapro to use only this competitions data in the validation set.",
      "votes": null
    },
    {
      "id": "605492",
      "postDate": "08/22/2019 12:23:38",
      "content": "<p>thank you for you advance, I think that use this competitions data fine tune model may be better</p>",
      "rawMarkdown": "thank you for you advance, I think that use this competitions data fine tune model may be better",
      "votes": null
    },
    {
      "id": "605596",
      "postDate": "08/22/2019 14:25:12",
      "content": "<p>when I try use the previous data in kernel, and batchsize=1, I get a memory error<code>CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.03 GiB already allocated; 1.88 MiB free; 190.01 MiB cached)</code></p>",
      "rawMarkdown": "when I try use the previous data in kernel, and batchsize=1, I get a memory error`CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.03 GiB already allocated; 1.88 MiB free; 190.01 MiB cached)`",
      "votes": null
    },
    {
      "id": "605896",
      "postDate": "08/22/2019 23:21:22",
      "content": "<p>Hi what's the diference between \"pretrain\" and \"finetune\"?  Does your first approach means training on 2015 data then stop, throw the optimizer weights away and just keep model weights, and further training on 19 data? I try to train all the 15 and 19 data together but seems not work well. </p>",
      "rawMarkdown": "Hi what's the diference between \"pretrain\" and \"finetune\"?  Does your first approach means training on 2015 data then stop, throw the optimizer weights away and just keep model weights, and further training on 19 data? I try to train all the 15 and 19 data together but seems not work well.",
      "votes": null
    },
    {
      "id": "605917",
      "postDate": "08/23/2019 00:52:08",
      "content": "<p>It depends on how you do it. There are many ways to do pretrain and finetune.</p>",
      "rawMarkdown": "It depends on how you do it. There are many ways to do pretrain and finetune.",
      "votes": null
    },
    {
      "id": "605989",
      "postDate": "08/23/2019 03:51:05",
      "content": "<p>it depends on your model and input size and has nothing to do with using old data. If the code worked before, it should be fine now as well!!</p>",
      "rawMarkdown": "it depends on your model and input size and has nothing to do with using old data. If the code worked before, it should be fine now as well!!",
      "votes": null
    },
    {
      "id": "605991",
      "postDate": "08/23/2019 03:56:09",
      "content": "<p>the previous have black picture, when I process the picture ,run error, for example,1986_left.jpeg, I will find all bad picture</p>",
      "rawMarkdown": "the previous have black picture, when I process the picture ,run error, for example,1986_left.jpeg, I will find all bad picture",
      "votes": null
    },
    {
      "id": "606063",
      "postDate": "08/23/2019 06:28:32",
      "content": "<p>4 or 5 images might be corrupted (imread does not load the array) and should not be included in the training or validation set.</p>",
      "rawMarkdown": "4 or 5 images might be corrupted (imread does not load the array) and should not be included in the training or validation set.",
      "votes": null
    },
    {
      "id": "606093",
      "postDate": "08/23/2019 07:18:23",
      "content": "<p>I have counted 14 damaged pictures in train_resize data</p>",
      "rawMarkdown": "I have counted 14 damaged pictures in train_resize data",
      "votes": null
    },
    {
      "id": "606115",
      "postDate": "08/23/2019 08:00:35",
      "content": "<p>use 2015 data as pretrained model and you can easily get LB beyond 0.8</p>",
      "rawMarkdown": "use 2015 data as pretrained model and you can easily get LB beyond 0.8",
      "votes": null
    },
    {
      "id": "606153",
      "postDate": "08/23/2019 09:12:35",
      "content": "<p><a href=\"/jiangkun2\">@jiangkun2</a>  seems you have only used current year data and are under top 100 rank, would you like to share few tips to improve accuracy :) </p>",
      "rawMarkdown": "jiangkun2  seems you have only used current year data and are under top 100 rank, would you like to share few tips to improve accuracy :)",
      "votes": null
    },
    {
      "id": "606169",
      "postDate": "08/23/2019 09:35:54",
      "content": "<p>hi <a href=\"/buaazijian\">@buaazijian</a>  thanks for letting us know,did you use kaggle kernels to train your finalmodel or your own pc?</p>",
      "rawMarkdown": "hi @buaazijian  thanks for letting us know,did you use kaggle kernels to train your finalmodel or your own pc?",
      "votes": null
    },
    {
      "id": "606191",
      "postDate": "08/23/2019 09:59:06",
      "content": "<p>own pc</p>",
      "rawMarkdown": "own pc",
      "votes": null
    },
    {
      "id": "606460",
      "postDate": "08/23/2019 16:00:22",
      "content": "<p><a href=\"/leixiang\">@leixiang</a> Did you train on 2015 data as regression or classification? If classification, then multiclass classification or multilabel classification ? Did you use any preprocessing like cropping or ben's processing ?  And what augmentations did you use ?   Sorry for so many questions but it will be really helpful if you answer them . Thanks </p>",
      "rawMarkdown": "leixiang Did you train on 2015 data as regression or classification? If classification, then multiclass classification or multilabel classification ? Did you use any preprocessing like cropping or ben's processing ?  And what augmentations did you use ?   Sorry for so many questions but it will be really helpful if you answer them . Thanks",
      "votes": null
    },
    {
      "id": "606709",
      "postDate": "08/23/2019 23:58:04",
      "content": "<p>just concat the datasets and using 2019 as val. ensembling some models can easily get 0.815</p>",
      "rawMarkdown": "just concat the datasets and using 2019 as val. ensembling some models can easily get 0.815",
      "votes": null
    },
    {
      "id": "608365",
      "postDate": "08/26/2019 17:04:18",
      "content": "<ol>\n<li>2015  pretrained and finetune on 2019 treated as regression problem</li>\n<li>cycle crop and no bens</li>\n<li>i use really heavy augments like flip max zoom，random contrast/crop ，hsv shift etc.</li>\n</ol>",
      "rawMarkdown": "1. 2015  pretrained and finetune on 2019 treated as regression problem\n2. cycle crop and no bens\n3. i use really heavy augments like flip max zoom，random contrast/crop ，hsv shift etc.",
      "votes": null
    },
    {
      "id": "608735",
      "postDate": "08/27/2019 05:50:01",
      "content": "<p>When pretrain, is it necessary to do augments?</p>",
      "rawMarkdown": "When pretrain, is it necessary to do augments?",
      "votes": null
    },
    {
      "id": "608737",
      "postDate": "08/27/2019 05:57:34",
      "content": "<p>And also to cycle crop for 2019 test data?</p>",
      "rawMarkdown": "And also to cycle crop for 2019 test data?",
      "votes": null
    },
    {
      "id": "608901",
      "postDate": "08/27/2019 08:59:23",
      "content": "<p>Need to freeze and unfreeze for pretrain? Or just train whole layer with a consistent lr?</p>",
      "rawMarkdown": "Need to freeze and unfreeze for pretrain? Or just train whole layer with a consistent lr?",
      "votes": null
    },
    {
      "id": "608991",
      "postDate": "08/27/2019 10:17:58",
      "content": "<p>yep the same crop both on 2015 and 2019</p>",
      "rawMarkdown": "yep the same crop both on 2015 and 2019",
      "votes": null
    },
    {
      "id": "609078",
      "postDate": "08/27/2019 11:55:07",
      "content": "<p>One question: transform first or resize first? Because it seems that transformed pictures are not reasonable to save as a checkpoint for later training process since there are some randomness within the transform process... If resize first then transform, will this behavior different to transform then resize?</p>",
      "rawMarkdown": "One question: transform first or resize first? Because it seems that transformed pictures are not reasonable to save as a checkpoint for later training process since there are some randomness within the transform process... If resize first then transform, will this behavior different to transform then resize?",
      "votes": null
    },
    {
      "id": "609627",
      "postDate": "08/27/2019 23:56:31",
      "content": "<p><a href=\"/leixiang\">@leixiang</a> what is cycle crop?</p>",
      "rawMarkdown": "leixiang what is cycle crop?",
      "votes": null
    },
    {
      "id": "609678",
      "postDate": "08/28/2019 01:58:29",
      "content": "<p>I think he meant circle crop</p>",
      "rawMarkdown": "I think he meant circle crop",
      "votes": null
    },
    {
      "id": "609685",
      "postDate": "08/28/2019 02:16:55",
      "content": "<p>in my operator , I unfreeze the pretrain and use new data to retain from scratch, and use 0.5*lr every 5 steps</p>",
      "rawMarkdown": "in my operator , I unfreeze the pretrain and use new data to retain from scratch, and use 0.5*lr every 5 steps",
      "votes": null
    },
    {
      "id": "614776",
      "postDate": "09/01/2019 03:44:59",
      "content": "<p>thanks <a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> for the good advice</p>",
      "rawMarkdown": "thanks @sidhanthholalkere for the good advice",
      "votes": null
    },
    {
      "id": "614854",
      "postDate": "09/01/2019 06:47:47",
      "content": "<p>Do you resize the image or apply some kind of cropping.</p>",
      "rawMarkdown": "Do you resize the image or apply some kind of cropping.",
      "votes": null
    },
    {
      "id": "615209",
      "postDate": "09/01/2019 15:57:14",
      "content": "<p>i think everyone is cropping it. I do as it gave me significant boost</p>",
      "rawMarkdown": "i think everyone is cropping it. I do as it gave me significant boost",
      "votes": null
    },
    {
      "id": "615647",
      "postDate": "09/02/2019 07:52:07",
      "content": "<p><a href=\"/buaazijian\">@buaazijian</a>  may need your help. My kernel run around 8 hours. To speed up, I upload the pretrained model(train on 2015 and finetune on 2019 data) as Dataset, the wired thing is that the prediction result differs from my fully run model(fit_generator and predict). Any clue of this? Thanks a lot.</p>",
      "rawMarkdown": "buaazijian  may need your help. My kernel run around 8 hours. To speed up, I upload the pretrained model(train on 2015 and finetune on 2019 data) as Dataset, the wired thing is that the prediction result differs from my fully run model(fit_generator and predict). Any clue of this? Thanks a lot.",
      "votes": null
    },
    {
      "id": "615855",
      "postDate": "09/02/2019 12:41:09",
      "content": "<p>I  remove the black edges, but  other people's public kernel does't works for me.</p>",
      "rawMarkdown": "I  remove the black edges, but  other people's public kernel does't works for me.",
      "votes": null
    },
    {
      "id": "615864",
      "postDate": "09/02/2019 12:52:18",
      "content": "<p>Are you sure the parameters are exactly the same? You must miss something, like random seed or optimizer parameters? </p>",
      "rawMarkdown": "Are you sure the parameters are exactly the same? You must miss something, like random seed or optimizer parameters?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 604272,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "08/21/2019 07:52:22",
      "content": "<p>Hi All,</p>\n\n<p>I just combined 2015 and 2019 data and tried to build a model but accuracy drops as compared to only 2019 data. Any tips would be welcome. Thanks !</p>",
      "votes": null,
      "replies": [
        {
          "id": 604298,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/21/2019 08:40:45",
          "content": "<p>I suggest building validation set from 2019 data only. Since 2019 training data might have closer distribution to 2019 test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 604560,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/21/2019 14:14:37",
          "content": "<p><a href=\"/anuragtr\">@anuragtr</a> are you using 2015 test data? My LB lowered drastically after adding 2015 test data. I got better result using 2015 training data with current competition data and i think <a href=\"/quandapro\">@quandapro</a> suggested a good strategy for building validation data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 604657,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "08/21/2019 16:10:29",
          "content": "<p>thanks <a href=\"/quandapro\">@quandapro</a> that's a good suggestion</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 604658,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "08/21/2019 16:11:13",
          "content": "<p>thanks <a href=\"/sabbiracoustic1006\">@sabbiracoustic1006</a> for the experience you had</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 604325,
      "author_name": "brunhs",
      "author_url": "",
      "post_date": "08/21/2019 09:16:15",
      "content": "<p>Maybe a careful study in the 2015 data, trying to find a same perfil data sounds good, after that merging it in to a big dataset for better training purposes sounds something interesting to try.</p>",
      "votes": null,
      "replies": [
        {
          "id": 604656,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "08/21/2019 16:09:48",
          "content": "<p>thanks <a href=\"/brunhs\">@brunhs</a> will try</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 604401,
      "author_name": "eshaan1",
      "author_url": "",
      "post_date": "08/21/2019 10:50:10",
      "content": "<p>great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 604678,
      "author_name": "buaazijian",
      "author_url": "",
      "post_date": "08/21/2019 16:35:40",
      "content": "<p>You can pretrain on the 2015 data then finetune on  2019 data, or you can  merging it in to a big dataset(use part of 2015 data). Two ways work for me, but the first is better. </p>",
      "votes": null,
      "replies": [
        {
          "id": 605109,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "08/22/2019 04:38:19",
          "content": "<p>thanks for the good advice <a href=\"/buaazijian\">@buaazijian</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605150,
          "author_name": "littlebiglittle",
          "author_url": "",
          "post_date": "08/22/2019 05:29:35",
          "content": "<p>hi, JIANJIAN. do you use both train and test of 2015 data for pretrain or only train data? pretrain on all 2015 data ( total about 88000 images) seems doesn't work for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605156,
          "author_name": "buaazijian",
          "author_url": "",
          "post_date": "08/22/2019 05:35:29",
          "content": "<p>yes, I use both train and test and it works for me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605162,
          "author_name": "littlebiglittle",
          "author_url": "",
          "post_date": "08/22/2019 05:50:36",
          "content": "<p>thank you for your info</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605896,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "08/22/2019 23:21:22",
          "content": "<p>Hi what's the diference between \"pretrain\" and \"finetune\"?  Does your first approach means training on 2015 data then stop, throw the optimizer weights away and just keep model weights, and further training on 19 data? I try to train all the 15 and 19 data together but seems not work well. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605917,
          "author_name": "buaazijian",
          "author_url": "",
          "post_date": "08/23/2019 00:52:08",
          "content": "<p>It depends on how you do it. There are many ways to do pretrain and finetune.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 606169,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "08/23/2019 09:35:54",
          "content": "<p>hi <a href=\"/buaazijian\">@buaazijian</a>  thanks for letting us know,did you use kaggle kernels to train your finalmodel or your own pc?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 606191,
          "author_name": "buaazijian",
          "author_url": "",
          "post_date": "08/23/2019 09:59:06",
          "content": "<p>own pc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 614854,
          "author_name": "vishnus",
          "author_url": "",
          "post_date": "09/01/2019 06:47:47",
          "content": "<p>Do you resize the image or apply some kind of cropping.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615209,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/01/2019 15:57:14",
          "content": "<p>i think everyone is cropping it. I do as it gave me significant boost</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615647,
          "author_name": "tonyzl",
          "author_url": "",
          "post_date": "09/02/2019 07:52:07",
          "content": "<p><a href=\"/buaazijian\">@buaazijian</a>  may need your help. My kernel run around 8 hours. To speed up, I upload the pretrained model(train on 2015 and finetune on 2019 data) as Dataset, the wired thing is that the prediction result differs from my fully run model(fit_generator and predict). Any clue of this? Thanks a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615855,
          "author_name": "buaazijian",
          "author_url": "",
          "post_date": "09/02/2019 12:41:09",
          "content": "<p>I  remove the black edges, but  other people's public kernel does't works for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 615864,
          "author_name": "buaazijian",
          "author_url": "",
          "post_date": "09/02/2019 12:52:18",
          "content": "<p>Are you sure the parameters are exactly the same? You must miss something, like random seed or optimizer parameters? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 605339,
      "author_name": "jiangkun2",
      "author_url": "",
      "post_date": "08/22/2019 09:13:58",
      "content": "<p>where is the 2015 data</p>",
      "votes": null,
      "replies": [
        {
          "id": 605360,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/22/2019 09:43:32",
          "content": "<p><a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605431,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/22/2019 11:22:07",
          "content": "<p>I find it!, Thank you, and how many epoch to train when use previous data! sincerely appreciate！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605440,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/22/2019 11:30:48",
          "content": "<p>monitor your validation loss until it converges and stop the training. Generally it should converge within 10 epochs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605467,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/22/2019 12:00:57",
          "content": "<p>the validation data also split from the previous data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605486,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/22/2019 12:16:59",
          "content": "<p>Currently i am using validation data from both this competition and previous competitions data but i am planning to use the strategy suggested by <a href=\"/quandapro\">@quandapro</a> to use only this competitions data in the validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605492,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/22/2019 12:23:38",
          "content": "<p>thank you for you advance, I think that use this competitions data fine tune model may be better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605596,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/22/2019 14:25:12",
          "content": "<p>when I try use the previous data in kernel, and batchsize=1, I get a memory error<code>CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.03 GiB already allocated; 1.88 MiB free; 190.01 MiB cached)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605989,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/23/2019 03:51:05",
          "content": "<p>it depends on your model and input size and has nothing to do with using old data. If the code worked before, it should be fine now as well!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605991,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/23/2019 03:56:09",
          "content": "<p>the previous have black picture, when I process the picture ,run error, for example,1986_left.jpeg, I will find all bad picture</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 606063,
          "author_name": "sabbiracoustic1006",
          "author_url": "",
          "post_date": "08/23/2019 06:28:32",
          "content": "<p>4 or 5 images might be corrupted (imread does not load the array) and should not be included in the training or validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 606093,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/23/2019 07:18:23",
          "content": "<p>I have counted 14 damaged pictures in train_resize data</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608901,
          "author_name": "zhan2019",
          "author_url": "",
          "post_date": "08/27/2019 08:59:23",
          "content": "<p>Need to freeze and unfreeze for pretrain? Or just train whole layer with a consistent lr?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609685,
          "author_name": "jiangkun2",
          "author_url": "",
          "post_date": "08/28/2019 02:16:55",
          "content": "<p>in my operator , I unfreeze the pretrain and use new data to retain from scratch, and use 0.5*lr every 5 steps</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 606115,
      "author_name": "leixiang",
      "author_url": "",
      "post_date": "08/23/2019 08:00:35",
      "content": "<p>use 2015 data as pretrained model and you can easily get LB beyond 0.8</p>",
      "votes": null,
      "replies": [
        {
          "id": 606460,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/23/2019 16:00:22",
          "content": "<p><a href=\"/leixiang\">@leixiang</a> Did you train on 2015 data as regression or classification? If classification, then multiclass classification or multilabel classification ? Did you use any preprocessing like cropping or ben's processing ?  And what augmentations did you use ?   Sorry for so many questions but it will be really helpful if you answer them . Thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608365,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/26/2019 17:04:18",
          "content": "<ol>\n<li>2015  pretrained and finetune on 2019 treated as regression problem</li>\n<li>cycle crop and no bens</li>\n<li>i use really heavy augments like flip max zoom，random contrast/crop ，hsv shift etc.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608735,
          "author_name": "zhan2019",
          "author_url": "",
          "post_date": "08/27/2019 05:50:01",
          "content": "<p>When pretrain, is it necessary to do augments?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608737,
          "author_name": "zhan2019",
          "author_url": "",
          "post_date": "08/27/2019 05:57:34",
          "content": "<p>And also to cycle crop for 2019 test data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608991,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/27/2019 10:17:58",
          "content": "<p>yep the same crop both on 2015 and 2019</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609078,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "08/27/2019 11:55:07",
          "content": "<p>One question: transform first or resize first? Because it seems that transformed pictures are not reasonable to save as a checkpoint for later training process since there are some randomness within the transform process... If resize first then transform, will this behavior different to transform then resize?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609627,
          "author_name": "yousof9",
          "author_url": "",
          "post_date": "08/27/2019 23:56:31",
          "content": "<p><a href=\"/leixiang\">@leixiang</a> what is cycle crop?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609678,
          "author_name": "quandapro",
          "author_url": "",
          "post_date": "08/28/2019 01:58:29",
          "content": "<p>I think he meant circle crop</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 606153,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "08/23/2019 09:12:35",
      "content": "<p><a href=\"/jiangkun2\">@jiangkun2</a>  seems you have only used current year data and are under top 100 rank, would you like to share few tips to improve accuracy :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 606709,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "08/23/2019 23:58:04",
      "content": "<p>just concat the datasets and using 2019 as val. ensembling some models can easily get 0.815</p>",
      "votes": null,
      "replies": [
        {
          "id": 614776,
          "author_name": "anuragtr",
          "author_url": "",
          "post_date": "09/01/2019 03:44:59",
          "content": "<p>thanks <a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> for the good advice</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "604270": "",
    "604272": "Hi All,\n\nI just combined 2015 and 2019 data and tried to build a model but accuracy drops as compared to only 2019 data. Any tips would be welcome. Thanks !",
    "604298": "I suggest building validation set from 2019 data only. Since 2019 training data might have closer distribution to 2019 test data.",
    "604325": "Maybe a careful study in the 2015 data, trying to find a same perfil data sounds good, after that merging it in to a big dataset for better training purposes sounds something interesting to try.",
    "604401": "great work",
    "604560": "anuragtr are you using 2015 test data? My LB lowered drastically after adding 2015 test data. I got better result using 2015 training data with current competition data and i think @quandapro suggested a good strategy for building validation data",
    "604656": "thanks @brunhs will try",
    "604657": "thanks @quandapro that's a good suggestion",
    "604658": "thanks @sabbiracoustic1006 for the experience you had",
    "604678": "You can pretrain on the 2015 data then finetune on  2019 data, or you can  merging it in to a big dataset(use part of 2015 data). Two ways work for me, but the first is better.",
    "605109": "thanks for the good advice @buaazijian",
    "605150": "hi, JIANJIAN. do you use both train and test of 2015 data for pretrain or only train data? pretrain on all 2015 data ( total about 88000 images) seems doesn't work for me",
    "605156": "yes, I use both train and test and it works for me.",
    "605162": "thank you for your info",
    "605339": "where is the 2015 data",
    "605360": "https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized",
    "605431": "I find it!, Thank you, and how many epoch to train when use previous data! sincerely appreciate！",
    "605440": "monitor your validation loss until it converges and stop the training. Generally it should converge within 10 epochs",
    "605467": "the validation data also split from the previous data?",
    "605486": "Currently i am using validation data from both this competition and previous competitions data but i am planning to use the strategy suggested by @quandapro to use only this competitions data in the validation set.",
    "605492": "thank you for you advance, I think that use this competitions data fine tune model may be better",
    "605596": "when I try use the previous data in kernel, and batchsize=1, I get a memory error`CUDA out of memory. Tried to allocate 2.00 MiB (GPU 0; 15.90 GiB total capacity; 15.03 GiB already allocated; 1.88 MiB free; 190.01 MiB cached)`",
    "605896": "Hi what's the diference between \"pretrain\" and \"finetune\"?  Does your first approach means training on 2015 data then stop, throw the optimizer weights away and just keep model weights, and further training on 19 data? I try to train all the 15 and 19 data together but seems not work well.",
    "605917": "It depends on how you do it. There are many ways to do pretrain and finetune.",
    "605989": "it depends on your model and input size and has nothing to do with using old data. If the code worked before, it should be fine now as well!!",
    "605991": "the previous have black picture, when I process the picture ,run error, for example,1986_left.jpeg, I will find all bad picture",
    "606063": "4 or 5 images might be corrupted (imread does not load the array) and should not be included in the training or validation set.",
    "606093": "I have counted 14 damaged pictures in train_resize data",
    "606115": "use 2015 data as pretrained model and you can easily get LB beyond 0.8",
    "606153": "jiangkun2  seems you have only used current year data and are under top 100 rank, would you like to share few tips to improve accuracy :)",
    "606169": "hi @buaazijian  thanks for letting us know,did you use kaggle kernels to train your finalmodel or your own pc?",
    "606191": "own pc",
    "606460": "leixiang Did you train on 2015 data as regression or classification? If classification, then multiclass classification or multilabel classification ? Did you use any preprocessing like cropping or ben's processing ?  And what augmentations did you use ?   Sorry for so many questions but it will be really helpful if you answer them . Thanks",
    "606709": "just concat the datasets and using 2019 as val. ensembling some models can easily get 0.815",
    "608365": "1. 2015  pretrained and finetune on 2019 treated as regression problem\n2. cycle crop and no bens\n3. i use really heavy augments like flip max zoom，random contrast/crop ，hsv shift etc.",
    "608735": "When pretrain, is it necessary to do augments?",
    "608737": "And also to cycle crop for 2019 test data?",
    "608901": "Need to freeze and unfreeze for pretrain? Or just train whole layer with a consistent lr?",
    "608991": "yep the same crop both on 2015 and 2019",
    "609078": "One question: transform first or resize first? Because it seems that transformed pictures are not reasonable to save as a checkpoint for later training process since there are some randomness within the transform process... If resize first then transform, will this behavior different to transform then resize?",
    "609627": "leixiang what is cycle crop?",
    "609678": "I think he meant circle crop",
    "609685": "in my operator , I unfreeze the pretrain and use new data to retain from scratch, and use 0.5*lr every 5 steps",
    "614776": "thanks @sidhanthholalkere for the good advice",
    "614854": "Do you resize the image or apply some kind of cropping.",
    "615209": "i think everyone is cropping it. I do as it gave me significant boost",
    "615647": "buaazijian  may need your help. My kernel run around 8 hours. To speed up, I upload the pretrained model(train on 2015 and finetune on 2019 data) as Dataset, the wired thing is that the prediction result differs from my fully run model(fit_generator and predict). Any clue of this? Thanks a lot.",
    "615855": "I  remove the black edges, but  other people's public kernel does't works for me.",
    "615864": "Are you sure the parameters are exactly the same? You must miss something, like random seed or optimizer parameters?"
  },
  "source": "meta"
}