{
  "id": 173724,
  "title": "Which external dataset to use as pytorch user ? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173724",
  "author_name": "",
  "post_date": "2020-08-10T12:12:45.757909500Z",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Which external dataset to use as Pytorch user ? I have been usings Romans External dataset till now (<a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">https://www.kaggle.com/nroman/melanoma-external-malignant-256</a>) with EfNet b4 and got .92 CV and .91 LB .\nI have seen people recommending Chris Deotte's Dataset , But its in tfrecord format , How to use it in Pytorch ? \nAlso i have been stuck on .91 lb for long time , Is there any public Kernel of pytorch that has higher LB that that ?</p>",
  "messages": [
    {
      "id": "965145",
      "postDate": "08/10/2020 12:12:45",
      "content": "<p>Which external dataset to use as Pytorch user ? I have been usings Romans External dataset till now (<a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">https://www.kaggle.com/nroman/melanoma-external-malignant-256</a>) with EfNet b4 and got .92 CV and .91 LB .\nI have seen people recommending Chris Deotte's Dataset , But its in tfrecord format , How to use it in Pytorch ? \nAlso i have been stuck on .91 lb for long time , Is there any public Kernel of pytorch that has higher LB that that ?</p>",
      "rawMarkdown": "Which external dataset to use as Pytorch user ? I have been usings Romans External dataset till now (https://www.kaggle.com/nroman/melanoma-external-malignant-256) with EfNet b4 and got .92 CV and .91 LB .\nI have seen people recommending Chris Deotte's Dataset , But its in tfrecord format , How to use it in Pytorch ? \nAlso i have been stuck on .91 lb for long time , Is there any public Kernel of pytorch that has higher LB that that ?",
      "votes": null
    },
    {
      "id": "965175",
      "postDate": "08/10/2020 12:38:03",
      "content": "<p>Hi! Have a look at <a href=\"/cdeotte\">@cdeotte</a> datasets again. He also provides the same datasets as jpeg. If you want the same stratification you have to use the .csv file he provides to see that the data comes into correct strata.</p>",
      "rawMarkdown": "Hi! Have a look at @cdeotte datasets again. He also provides the same datasets as jpeg. If you want the same stratification you have to use the .csv file he provides to see that the data comes into correct strata.",
      "votes": null
    },
    {
      "id": "965181",
      "postDate": "08/10/2020 12:46:02",
      "content": "<p>Yep, as <a href=\"https://www.kaggle.com/andersericssongnosco\" target=\"_blank\">@andersericssongnosco</a> said, Chris has provided the data in JPEG, 256 is <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\" target=\"_blank\">here</a>. :)</p>",
      "rawMarkdown": "Yep, as @andersericssongnosco said, Chris has provided the data in JPEG, 256 is [here](https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256). :)",
      "votes": null
    },
    {
      "id": "965187",
      "postDate": "08/10/2020 12:51:53",
      "content": "<p>I typically try to get all \"unique/non duplicate\" images in 1 single train folder and then let the data-loader do all the data generation, folds etc.</p>",
      "rawMarkdown": "I typically try to get all \"unique/non duplicate\" images in 1 single train folder and then let the data-loader do all the data generation, folds etc.",
      "votes": null
    },
    {
      "id": "965190",
      "postDate": "08/10/2020 12:58:02",
      "content": "<p>Thanks for your reply.\nI have recheked .There is a jpeg folder containing 580 this years Melannoma Images but external images are in tfrecord format .\n<a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">https://www.kaggle.com/cdeotte/melanoma-256x256</a></p>",
      "rawMarkdown": "Thanks for your reply.\nI have recheked .There is a jpeg folder containing 580 this years Melannoma Images but external images are in tfrecord format .\nhttps://www.kaggle.com/cdeotte/melanoma-256x256",
      "votes": null
    },
    {
      "id": "965191",
      "postDate": "08/10/2020 13:01:21",
      "content": "<p>Yeah i have been using this one only <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\">https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256</a> .\nBut this doesnt have any external data \n<a href=\"/sarques\">@sarques</a> </p>",
      "rawMarkdown": "Yeah i have been using this one only https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256 .\nBut this doesnt have any external data \n@sarques",
      "votes": null
    },
    {
      "id": "965193",
      "postDate": "08/10/2020 13:02:55",
      "content": "<p>Which external dataset you use ? </p>",
      "rawMarkdown": "Which external dataset you use ?",
      "votes": null
    },
    {
      "id": "965197",
      "postDate": "08/10/2020 13:05:46",
      "content": "<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#958676\" target=\"_blank\">Here you go!</a> :)</p>",
      "rawMarkdown": "[Here you go!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#958676) :)",
      "votes": null
    },
    {
      "id": "965229",
      "postDate": "08/10/2020 13:36:08",
      "content": "<p>Thanks , Didn't know 2019 data's were provided as JPEG's too .\nHowever one question ,  2020 data has 15 tfrecords , 2019 has 30 tfrecords . \nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?\n<a href=\"/sarques\">@sarques</a> </p>",
      "rawMarkdown": "Thanks , Didn't know 2019 data's were provided as JPEG's too .\nHowever one question ,  2020 data has 15 tfrecords , 2019 has 30 tfrecords . \nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?\n@sarques",
      "votes": null
    },
    {
      "id": "965242",
      "postDate": "08/10/2020 13:45:43",
      "content": "<p>Although this idea seems good to me, I don't think I am qualified enough to comment on data leakage, what I can do is to give you another idea if you want only the malignant images, it should not concern data leakage I guess! :)</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">Here</a>.</p>",
      "rawMarkdown": "Although this idea seems good to me, I don't think I am qualified enough to comment on data leakage, what I can do is to give you another idea if you want only the malignant images, it should not concern data leakage I guess! :)\n\n[Here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130).",
      "votes": null
    },
    {
      "id": "965286",
      "postDate": "08/10/2020 14:22:56",
      "content": "<p>I am using ISIC 2019 (inclusive of 2018, 2017) , and the other external dataset of 580 odd positive melanoma images. </p>",
      "rawMarkdown": "I am using ISIC 2019 (inclusive of 2018, 2017) , and the other external dataset of 580 odd positive melanoma images.",
      "votes": null
    },
    {
      "id": "965288",
      "postDate": "08/10/2020 14:25:13",
      "content": "<p>Okay , Let me try :D \nThanks for such informative and awesome Reply's . </p>",
      "rawMarkdown": "Okay , Let me try :D \nThanks for such informative and awesome Reply's .",
      "votes": null
    },
    {
      "id": "965294",
      "postDate": "08/10/2020 14:26:34",
      "content": "<p>How do you create your validation dataset ? Validation data should have only images of 2020 data right ?\n2020 data has 15 tfrecords , 2019 has 30 tfrecords .\nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?</p>",
      "rawMarkdown": "How do you create your validation dataset ? Validation data should have only images of 2020 data right ?\n2020 data has 15 tfrecords , 2019 has 30 tfrecords .\nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?",
      "votes": null
    },
    {
      "id": "965447",
      "postDate": "08/10/2020 16:37:01",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a></p>",
      "rawMarkdown": "https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg",
      "votes": null
    },
    {
      "id": "966105",
      "postDate": "08/11/2020 06:21:05",
      "content": "<p>Thanks for Reply , will have a look :D </p>",
      "rawMarkdown": "Thanks for Reply , will have a look :D",
      "votes": null
    },
    {
      "id": "966108",
      "postDate": "08/11/2020 06:28:37",
      "content": "<p>All my TFRecords datasets have a corresponding dataset with just JPEGs for PyTorch users. There is list of sizes for 2020 comp data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\" target=\"_blank\">here</a>. And list of sizes for 2019 2018 2017 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. Use the datasets that say \"JPEGs\". (They are the exact same jpegs that are inside my TFRecords).</p>",
      "rawMarkdown": "All my TFRecords datasets have a corresponding dataset with just JPEGs for PyTorch users. There is list of sizes for 2020 comp data [here][1]. And list of sizes for 2019 2018 2017 [here][2]. Use the datasets that say \"JPEGs\". (They are the exact same jpegs that are inside my TFRecords).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
      "votes": null
    },
    {
      "id": "966127",
      "postDate": "08/11/2020 06:48:52",
      "content": "<blockquote>\n  <p>Thanks , Didn't know 2019 data's were provided as JPEG's too .<br>\n  However one question , 2020 data has 15 tfrecords , 2019 has 30 tfrecords .<br>\n  How to build a train and validation fold ?<br>\n  I am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .<br>\n  However i am concerned about data leakage. Any suggestion ?<br>\n  <a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> </p>\n</blockquote>\n<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Any insights on this? :)</p>",
      "rawMarkdown": "&gt; Thanks , Didn't know 2019 data's were provided as JPEG's too .\n&gt; However one question , 2020 data has 15 tfrecords , 2019 has 30 tfrecords .\n&gt; How to build a train and validation fold ?\n&gt; I am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\n&gt; However i am concerned about data leakage. Any suggestion ?\n&gt; @sarques \n\nHey @cdeotte Any insights on this? :)",
      "votes": null
    },
    {
      "id": "966436",
      "postDate": "08/11/2020 12:19:57",
      "content": "<p>I am following the simple approach. In my val loader, which loads image IDs (per fold) from a DataFrame. I put a condition saying \"Is 2020\" == 1.  Note that \"Is 2020\" is a column I manually created in the loader DF which is 1 for all 2020 images. So now what happens is during validation only 2020 images are used, while in training all images are used. </p>",
      "rawMarkdown": "I am following the simple approach. In my val loader, which loads image IDs (per fold) from a DataFrame. I put a condition saying \"Is 2020\" == 1.  Note that \"Is 2020\" is a column I manually created in the loader DF which is 1 for all 2020 images. So now what happens is during validation only 2020 images are used, while in training all images are used.",
      "votes": null
    },
    {
      "id": "967668",
      "postDate": "08/12/2020 12:13:09",
      "content": "<p>Cool , That seems a nice idea</p>",
      "rawMarkdown": "Cool , That seems a nice idea",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 965187,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "08/10/2020 12:51:53",
      "content": "<p>I typically try to get all \"unique/non duplicate\" images in 1 single train folder and then let the data-loader do all the data generation, folds etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 965193,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 13:02:55",
          "content": "<p>Which external dataset you use ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965286,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/10/2020 14:22:56",
          "content": "<p>I am using ISIC 2019 (inclusive of 2018, 2017) , and the other external dataset of 580 odd positive melanoma images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965294,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 14:26:34",
          "content": "<p>How do you create your validation dataset ? Validation data should have only images of 2020 data right ?\n2020 data has 15 tfrecords , 2019 has 30 tfrecords .\nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 966436,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/11/2020 12:19:57",
          "content": "<p>I am following the simple approach. In my val loader, which loads image IDs (per fold) from a DataFrame. I put a condition saying \"Is 2020\" == 1.  Note that \"Is 2020\" is a column I manually created in the loader DF which is 1 for all 2020 images. So now what happens is during validation only 2020 images are used, while in training all images are used. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 967668,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/12/2020 12:13:09",
          "content": "<p>Cool , That seems a nice idea</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 966108,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/11/2020 06:28:37",
      "content": "<p>All my TFRecords datasets have a corresponding dataset with just JPEGs for PyTorch users. There is list of sizes for 2020 comp data <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\" target=\"_blank\">here</a>. And list of sizes for 2019 2018 2017 <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">here</a>. Use the datasets that say \"JPEGs\". (They are the exact same jpegs that are inside my TFRecords).</p>",
      "votes": null,
      "replies": [
        {
          "id": 966127,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/11/2020 06:48:52",
          "content": "<blockquote>\n  <p>Thanks , Didn't know 2019 data's were provided as JPEG's too .<br>\n  However one question , 2020 data has 15 tfrecords , 2019 has 30 tfrecords .<br>\n  How to build a train and validation fold ?<br>\n  I am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .<br>\n  However i am concerned about data leakage. Any suggestion ?<br>\n  <a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> </p>\n</blockquote>\n<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Any insights on this? :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 965175,
      "author_name": "andersericssongnosco",
      "author_url": "",
      "post_date": "08/10/2020 12:38:03",
      "content": "<p>Hi! Have a look at <a href=\"/cdeotte\">@cdeotte</a> datasets again. He also provides the same datasets as jpeg. If you want the same stratification you have to use the .csv file he provides to see that the data comes into correct strata.</p>",
      "votes": null,
      "replies": [
        {
          "id": 965181,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/10/2020 12:46:02",
          "content": "<p>Yep, as <a href=\"https://www.kaggle.com/andersericssongnosco\" target=\"_blank\">@andersericssongnosco</a> said, Chris has provided the data in JPEG, 256 is <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\" target=\"_blank\">here</a>. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965190,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 12:58:02",
          "content": "<p>Thanks for your reply.\nI have recheked .There is a jpeg folder containing 580 this years Melannoma Images but external images are in tfrecord format .\n<a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">https://www.kaggle.com/cdeotte/melanoma-256x256</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965191,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 13:01:21",
          "content": "<p>Yeah i have been using this one only <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256\">https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256</a> .\nBut this doesnt have any external data \n<a href=\"/sarques\">@sarques</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965197,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/10/2020 13:05:46",
          "content": "<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#958676\" target=\"_blank\">Here you go!</a> :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965229,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 13:36:08",
          "content": "<p>Thanks , Didn't know 2019 data's were provided as JPEG's too .\nHowever one question ,  2020 data has 15 tfrecords , 2019 has 30 tfrecords . \nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?\n<a href=\"/sarques\">@sarques</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965242,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/10/2020 13:45:43",
          "content": "<p>Although this idea seems good to me, I don't think I am qualified enough to comment on data leakage, what I can do is to give you another idea if you want only the malignant images, it should not concern data leakage I guess! :)</p>\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130\" target=\"_blank\">Here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965288,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/10/2020 14:25:13",
          "content": "<p>Okay , Let me try :D \nThanks for such informative and awesome Reply's . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 965447,
      "author_name": "hamonk",
      "author_url": "",
      "post_date": "08/10/2020 16:37:01",
      "content": "<p><a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 966105,
          "author_name": "haideralishuvo",
          "author_url": "",
          "post_date": "08/11/2020 06:21:05",
          "content": "<p>Thanks for Reply , will have a look :D </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "965145": "Which external dataset to use as Pytorch user ? I have been usings Romans External dataset till now (https://www.kaggle.com/nroman/melanoma-external-malignant-256) with EfNet b4 and got .92 CV and .91 LB .\nI have seen people recommending Chris Deotte's Dataset , But its in tfrecord format , How to use it in Pytorch ? \nAlso i have been stuck on .91 lb for long time , Is there any public Kernel of pytorch that has higher LB that that ?",
    "965175": "Hi! Have a look at @cdeotte datasets again. He also provides the same datasets as jpeg. If you want the same stratification you have to use the .csv file he provides to see that the data comes into correct strata.",
    "965181": "Yep, as @andersericssongnosco said, Chris has provided the data in JPEG, 256 is [here](https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256). :)",
    "965187": "I typically try to get all \"unique/non duplicate\" images in 1 single train folder and then let the data-loader do all the data generation, folds etc.",
    "965190": "Thanks for your reply.\nI have recheked .There is a jpeg folder containing 580 this years Melannoma Images but external images are in tfrecord format .\nhttps://www.kaggle.com/cdeotte/melanoma-256x256",
    "965191": "Yeah i have been using this one only https://www.kaggle.com/cdeotte/jpeg-melanoma-256x256 .\nBut this doesnt have any external data \n@sarques",
    "965193": "Which external dataset you use ?",
    "965197": "[Here you go!](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154296#958676) :)",
    "965229": "Thanks , Didn't know 2019 data's were provided as JPEG's too .\nHowever one question ,  2020 data has 15 tfrecords , 2019 has 30 tfrecords . \nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?\n@sarques",
    "965242": "Although this idea seems good to me, I don't think I am qualified enough to comment on data leakage, what I can do is to give you another idea if you want only the malignant images, it should not concern data leakage I guess! :)\n\n[Here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173130).",
    "965286": "I am using ISIC 2019 (inclusive of 2018, 2017) , and the other external dataset of 580 odd positive melanoma images.",
    "965288": "Okay , Let me try :D \nThanks for such informative and awesome Reply's .",
    "965294": "How do you create your validation dataset ? Validation data should have only images of 2020 data right ?\n2020 data has 15 tfrecords , 2019 has 30 tfrecords .\nHow to build a train and validation fold ?\nI am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\nHowever i am concerned about data leakage . Any suggestion ?",
    "965447": "https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg",
    "966105": "Thanks for Reply , will have a look :D",
    "966108": "All my TFRecords datasets have a corresponding dataset with just JPEGs for PyTorch users. There is list of sizes for 2020 comp data [here][1]. And list of sizes for 2019 2018 2017 [here][2]. Use the datasets that say \"JPEGs\". (They are the exact same jpegs that are inside my TFRecords).\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910",
    "966127": "&gt; Thanks , Didn't know 2019 data's were provided as JPEG's too .\n&gt; However one question , 2020 data has 15 tfrecords , 2019 has 30 tfrecords .\n&gt; How to build a train and validation fold ?\n&gt; I am thinking of taking 12 tfrecords from 2020 data in trainfold , 3 tfrecord from 2020 data in validfold and in trainfold add all malignants of previous years .\n&gt; However i am concerned about data leakage. Any suggestion ?\n&gt; @sarques \n\nHey @cdeotte Any insights on this? :)",
    "966436": "I am following the simple approach. In my val loader, which loads image IDs (per fold) from a DataFrame. I put a condition saying \"Is 2020\" == 1.  Note that \"Is 2020\" is a column I manually created in the loader DF which is 1 for all 2020 images. So now what happens is during validation only 2020 images are used, while in training all images are used.",
    "967668": "Cool , That seems a nice idea"
  },
  "source": "meta"
}