{
  "id": 170179,
  "title": "external data and overfitting",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170179",
  "author_name": "",
  "post_date": "2020-07-26T17:40:55.931788200Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>What are your experiences with external data - ISIC 2019?\nI read that people see higher scores with it, I agree.\nHowever.\nIn deep learning the way to DECREASE overfitting is to add more data.\nBy using ISIC 2019 we are adding more data.\nBut the result is INCREASED overfitting, not decreased.\nTo clarify - I see better score in train, worse in valid.\nHow is that possible?</p>",
  "messages": [
    {
      "id": "946644",
      "postDate": "07/26/2020 17:40:55",
      "content": "<p>What are your experiences with external data - ISIC 2019?\nI read that people see higher scores with it, I agree.\nHowever.\nIn deep learning the way to DECREASE overfitting is to add more data.\nBy using ISIC 2019 we are adding more data.\nBut the result is INCREASED overfitting, not decreased.\nTo clarify - I see better score in train, worse in valid.\nHow is that possible?</p>",
      "rawMarkdown": "What are your experiences with external data - ISIC 2019?\nI read that people see higher scores with it, I agree.\nHowever.\nIn deep learning the way to DECREASE overfitting is to add more data.\nBy using ISIC 2019 we are adding more data.\nBut the result is INCREASED overfitting, not decreased.\nTo clarify - I see better score in train, worse in valid.\nHow is that possible?",
      "votes": null
    },
    {
      "id": "946751",
      "postDate": "07/26/2020 19:37:58",
      "content": "<p>Maybe You have in your validation dataset only samples from ISIC2020, so you train model on external and internal, but measure quality only on internal dataset.\nAlso I heard, that ISIC2019 very vary from ISIC2020 and ISIC2018, but ISIC2018 and ISIC2020 look similar. </p>",
      "rawMarkdown": "Maybe You have in your validation dataset only samples from ISIC2020, so you train model on external and internal, but measure quality only on internal dataset.\nAlso I heard, that ISIC2019 very vary from ISIC2020 and ISIC2018, but ISIC2018 and ISIC2020 look similar.",
      "votes": null
    },
    {
      "id": "946774",
      "postDate": "07/26/2020 20:14:05",
      "content": "<p>That is correct, I validate only on 2020.\nWhat's the point to validate on 2019 if test data is from 2020?</p>\n\n<p>About 2018:\n2019 contains 2018.</p>",
      "rawMarkdown": "That is correct, I validate only on 2020.\nWhat's the point to validate on 2019 if test data is from 2020?\n\nAbout 2018:\n2019 contains 2018.",
      "votes": null
    },
    {
      "id": "946777",
      "postDate": "07/26/2020 20:20:08",
      "content": "<p>Then your validation dataset is not the same as trainnig dataset and it is not fair comparison. \nMaybe your overfitting is not so big, you should compare model only on data from ISIC2020 or add some External samples to your validation.\nYes ISIC2019 contains all previous ISICs, I mean ISIC2019 is new samples for previous competitions</p>",
      "rawMarkdown": "Then your validation dataset is not the same as trainnig dataset and it is not fair comparison. \nMaybe your overfitting is not so big, you should compare model only on data from ISIC2020 or add some External samples to your validation.\nYes ISIC2019 contains all previous ISICs, I mean ISIC2019 is new samples for previous competitions",
      "votes": null
    },
    {
      "id": "946803",
      "postDate": "07/26/2020 21:11:06",
      "content": "<p>The ISIC 2019/2018/2017 data is different enough from the 2020 data that your model is preferentially training the larger external dataset. So it does better. When you check your model with your Validation data (only 2020) the model is overfitted to external data.</p>\n\n<p>People have reported mixed results with the external data.</p>",
      "rawMarkdown": "The ISIC 2019/2018/2017 data is different enough from the 2020 data that your model is preferentially training the larger external dataset. So it does better. When you check your model with your Validation data (only 2020) the model is overfitted to external data.\n\nPeople have reported mixed results with the external data.",
      "votes": null
    },
    {
      "id": "946805",
      "postDate": "07/26/2020 21:16:41",
      "content": "<p>if the data in 2019 is different then what's the point to use it in the competition?</p>",
      "rawMarkdown": "if the data in 2019 is different then what's the point to use it in the competition?",
      "votes": null
    },
    {
      "id": "946827",
      "postDate": "07/26/2020 21:46:13",
      "content": "<p>I think many competitors thinks it helps. I am using external data, but not simply adding it all as more data. I think you might need to be more strategic in how you use it. I haven't proven it is helping me, but I think it is.</p>",
      "rawMarkdown": "I think many competitors thinks it helps. I am using external data, but not simply adding it all as more data. I think you might need to be more strategic in how you use it. I haven't proven it is helping me, but I think it is.",
      "votes": null
    },
    {
      "id": "946847",
      "postDate": "07/26/2020 22:29:49",
      "content": "<p>Read carefully what <a href=\"/aybatov\">@aybatov</a> writes. Your model it's more robust with more data, that's the point to use 2019 and 2018 data. In resume you should use data from 2018 and 2019 to validate as well.</p>",
      "rawMarkdown": "Read carefully what @aybatov writes. Your model it's more robust with more data, that's the point to use 2019 and 2018 data. In resume you should use data from 2018 and 2019 to validate as well.",
      "votes": null
    },
    {
      "id": "946868",
      "postDate": "07/26/2020 22:58:53",
      "content": "<p>I just ran adverserial validation on external data vs train as well as external data vs test (Image Data). For 2019, AUC is even greater than 0.99!! So, I feel it's risky to include 2019 data. </p>",
      "rawMarkdown": "I just ran adverserial validation on external data vs train as well as external data vs test (Image Data). For 2019, AUC is even greater than 0.99!! So, I feel it's risky to include 2019 data.",
      "votes": null
    },
    {
      "id": "946913",
      "postDate": "07/26/2020 23:58:24",
      "content": "<p>Omg, I heard but not tested for myself, 0.99 AUC is really scared</p>",
      "rawMarkdown": "Omg, I heard but not tested for myself, 0.99 AUC is really scared",
      "votes": null
    },
    {
      "id": "946917",
      "postDate": "07/27/2020 00:14:08",
      "content": "<p>thanks for info</p>",
      "rawMarkdown": "thanks for info",
      "votes": null
    },
    {
      "id": "946921",
      "postDate": "07/27/2020 00:17:18",
      "content": "<p><a href=\"/hiramcho\">@hiramcho</a> no I still don't understand\nthis is competition for 2020 data, score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation? how it helps in anything?</p>",
      "rawMarkdown": "hiramcho no I still don't understand\nthis is competition for 2020 data, score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation? how it helps in anything?",
      "votes": null
    },
    {
      "id": "946963",
      "postDate": "07/27/2020 01:39:33",
      "content": "<p>&gt;  \"this is competition for 2020 data\" </p>\n\n<p>Answer: Yes but, there are a big class imbalance due the fact that ther're a way more images without <strong>melanoma</strong> so using more data, that includes images with melanoma from another datasets can be helpful.</p>\n\n<p>&gt; \"score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation\"</p>\n\n<p>Answer: Yes, but you use 2020 data AND 2018-2019 to do the validation,  the point of use 2019 data to do validation it's to solve the problem that you write in this post. <a href=\"https://www.kaggle.com/ks2019/adverserial-validation-on-image-data\">This</a> can be helpful. BTW thanks <a href=\"/ks2019\">@ks2019</a> for share.</p>",
      "rawMarkdown": "&gt;  \"this is competition for 2020 data\" \n\nAnswer: Yes but, there are a big class imbalance due the fact that ther're a way more images without **melanoma** so using more data, that includes images with melanoma from another datasets can be helpful.\n\n&gt; \"score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation\"\n\nAnswer: Yes, but you use 2020 data AND 2018-2019 to do the validation,  the point of use 2019 data to do validation it's to solve the problem that you write in this post. [This](https://www.kaggle.com/ks2019/adverserial-validation-on-image-data) can be helpful. BTW thanks @ks2019 for share.",
      "votes": null
    },
    {
      "id": "947184",
      "postDate": "07/27/2020 05:52:59",
      "content": "<p>256x256 eff-b1 with external 2017+2018+2019: CV 0.901 LB 0.9374\n256x256 eff-b1 with external 2017+2018: CV 0.895 LB 0.9395</p>\n\n<p>To me adding ISIC 2019 increases CV but decreases LB.</p>",
      "rawMarkdown": "256x256 eff-b1 with external 2017+2018+2019: CV 0.901 LB 0.9374\n256x256 eff-b1 with external 2017+2018: CV 0.895 LB 0.9395\n\nTo me adding ISIC 2019 increases CV but decreases LB.",
      "votes": null
    },
    {
      "id": "947420",
      "postDate": "07/27/2020 08:59:05",
      "content": "<p>Maybe I will explain:\n- assumption that adding 2019 is helpful is a theory\n- in practice by validating on 2020 we see it's not</p>\n\n<p>Of course it needs more research and experiments, but my initial thought is that pracice is more important than theory</p>",
      "rawMarkdown": "Maybe I will explain:\n- assumption that adding 2019 is helpful is a theory\n- in practice by validating on 2020 we see it's not\n\nOf course it needs more research and experiments, but my initial thought is that pracice is more important than theory",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 946751,
      "author_name": "aybatov",
      "author_url": "",
      "post_date": "07/26/2020 19:37:58",
      "content": "<p>Maybe You have in your validation dataset only samples from ISIC2020, so you train model on external and internal, but measure quality only on internal dataset.\nAlso I heard, that ISIC2019 very vary from ISIC2020 and ISIC2018, but ISIC2018 and ISIC2020 look similar. </p>",
      "votes": null,
      "replies": [
        {
          "id": 946774,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/26/2020 20:14:05",
          "content": "<p>That is correct, I validate only on 2020.\nWhat's the point to validate on 2019 if test data is from 2020?</p>\n\n<p>About 2018:\n2019 contains 2018.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946777,
          "author_name": "aybatov",
          "author_url": "",
          "post_date": "07/26/2020 20:20:08",
          "content": "<p>Then your validation dataset is not the same as trainnig dataset and it is not fair comparison. \nMaybe your overfitting is not so big, you should compare model only on data from ISIC2020 or add some External samples to your validation.\nYes ISIC2019 contains all previous ISICs, I mean ISIC2019 is new samples for previous competitions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946805,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/26/2020 21:16:41",
          "content": "<p>if the data in 2019 is different then what's the point to use it in the competition?</p>",
          "votes": null,
          "replies": [
            {
              "id": 946827,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "07/26/2020 21:46:13",
              "content": "<p>I think many competitors thinks it helps. I am using external data, but not simply adding it all as more data. I think you might need to be more strategic in how you use it. I haven't proven it is helping me, but I think it is.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 946847,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "07/26/2020 22:29:49",
          "content": "<p>Read carefully what <a href=\"/aybatov\">@aybatov</a> writes. Your model it's more robust with more data, that's the point to use 2019 and 2018 data. In resume you should use data from 2018 and 2019 to validate as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946921,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/27/2020 00:17:18",
          "content": "<p><a href=\"/hiramcho\">@hiramcho</a> no I still don't understand\nthis is competition for 2020 data, score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation? how it helps in anything?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946963,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "07/27/2020 01:39:33",
          "content": "<p>&gt;  \"this is competition for 2020 data\" </p>\n\n<p>Answer: Yes but, there are a big class imbalance due the fact that ther're a way more images without <strong>melanoma</strong> so using more data, that includes images with melanoma from another datasets can be helpful.</p>\n\n<p>&gt; \"score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation\"</p>\n\n<p>Answer: Yes, but you use 2020 data AND 2018-2019 to do the validation,  the point of use 2019 data to do validation it's to solve the problem that you write in this post. <a href=\"https://www.kaggle.com/ks2019/adverserial-validation-on-image-data\">This</a> can be helpful. BTW thanks <a href=\"/ks2019\">@ks2019</a> for share.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947420,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/27/2020 08:59:05",
          "content": "<p>Maybe I will explain:\n- assumption that adding 2019 is helpful is a theory\n- in practice by validating on 2020 we see it's not</p>\n\n<p>Of course it needs more research and experiments, but my initial thought is that pracice is more important than theory</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 946803,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "07/26/2020 21:11:06",
      "content": "<p>The ISIC 2019/2018/2017 data is different enough from the 2020 data that your model is preferentially training the larger external dataset. So it does better. When you check your model with your Validation data (only 2020) the model is overfitted to external data.</p>\n\n<p>People have reported mixed results with the external data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 946868,
      "author_name": "ks2019",
      "author_url": "",
      "post_date": "07/26/2020 22:58:53",
      "content": "<p>I just ran adverserial validation on external data vs train as well as external data vs test (Image Data). For 2019, AUC is even greater than 0.99!! So, I feel it's risky to include 2019 data. </p>",
      "votes": null,
      "replies": [
        {
          "id": 946913,
          "author_name": "aybatov",
          "author_url": "",
          "post_date": "07/26/2020 23:58:24",
          "content": "<p>Omg, I heard but not tested for myself, 0.99 AUC is really scared</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946917,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/27/2020 00:14:08",
          "content": "<p>thanks for info</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 947184,
      "author_name": "waylongo",
      "author_url": "",
      "post_date": "07/27/2020 05:52:59",
      "content": "<p>256x256 eff-b1 with external 2017+2018+2019: CV 0.901 LB 0.9374\n256x256 eff-b1 with external 2017+2018: CV 0.895 LB 0.9395</p>\n\n<p>To me adding ISIC 2019 increases CV but decreases LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "946644": "What are your experiences with external data - ISIC 2019?\nI read that people see higher scores with it, I agree.\nHowever.\nIn deep learning the way to DECREASE overfitting is to add more data.\nBy using ISIC 2019 we are adding more data.\nBut the result is INCREASED overfitting, not decreased.\nTo clarify - I see better score in train, worse in valid.\nHow is that possible?",
    "946751": "Maybe You have in your validation dataset only samples from ISIC2020, so you train model on external and internal, but measure quality only on internal dataset.\nAlso I heard, that ISIC2019 very vary from ISIC2020 and ISIC2018, but ISIC2018 and ISIC2020 look similar.",
    "946774": "That is correct, I validate only on 2020.\nWhat's the point to validate on 2019 if test data is from 2020?\n\nAbout 2018:\n2019 contains 2018.",
    "946777": "Then your validation dataset is not the same as trainnig dataset and it is not fair comparison. \nMaybe your overfitting is not so big, you should compare model only on data from ISIC2020 or add some External samples to your validation.\nYes ISIC2019 contains all previous ISICs, I mean ISIC2019 is new samples for previous competitions",
    "946803": "The ISIC 2019/2018/2017 data is different enough from the 2020 data that your model is preferentially training the larger external dataset. So it does better. When you check your model with your Validation data (only 2020) the model is overfitted to external data.\n\nPeople have reported mixed results with the external data.",
    "946805": "if the data in 2019 is different then what's the point to use it in the competition?",
    "946827": "I think many competitors thinks it helps. I am using external data, but not simply adding it all as more data. I think you might need to be more strategic in how you use it. I haven't proven it is helping me, but I think it is.",
    "946847": "Read carefully what @aybatov writes. Your model it's more robust with more data, that's the point to use 2019 and 2018 data. In resume you should use data from 2018 and 2019 to validate as well.",
    "946868": "I just ran adverserial validation on external data vs train as well as external data vs test (Image Data). For 2019, AUC is even greater than 0.99!! So, I feel it's risky to include 2019 data.",
    "946913": "Omg, I heard but not tested for myself, 0.99 AUC is really scared",
    "946917": "thanks for info",
    "946921": "hiramcho no I still don't understand\nthis is competition for 2020 data, score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation? how it helps in anything?",
    "946963": "&gt;  \"this is competition for 2020 data\" \n\nAnswer: Yes but, there are a big class imbalance due the fact that ther're a way more images without **melanoma** so using more data, that includes images with melanoma from another datasets can be helpful.\n\n&gt; \"score will depend how the model perform on 2020 data\nwhat's the point to use 2019 data for validation\"\n\nAnswer: Yes, but you use 2020 data AND 2018-2019 to do the validation,  the point of use 2019 data to do validation it's to solve the problem that you write in this post. [This](https://www.kaggle.com/ks2019/adverserial-validation-on-image-data) can be helpful. BTW thanks @ks2019 for share.",
    "947184": "256x256 eff-b1 with external 2017+2018+2019: CV 0.901 LB 0.9374\n256x256 eff-b1 with external 2017+2018: CV 0.895 LB 0.9395\n\nTo me adding ISIC 2019 increases CV but decreases LB.",
    "947420": "Maybe I will explain:\n- assumption that adding 2019 is helpful is a theory\n- in practice by validating on 2020 we see it's not\n\nOf course it needs more research and experiments, but my initial thought is that pracice is more important than theory"
  },
  "source": "meta"
}