{
  "id": 162010,
  "title": "model overffitting because of the extremely in-banlanced training data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/162010",
  "author_name": "Matrix",
  "post_date": "2020-06-27T03:13:58.960000",
  "votes": 0,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I found the malicious skin data only hold  2% in training datasets, causing my model overfitting and produce zero for all test data, though I take the CV during training procedure. \nGuys, please help me! thanks</p>",
  "messages": [
    {
      "id": 909391,
      "postDate": "2020-06-30T15:05:17.730Z",
      "content": "<p>You should find ways to create \"synthetic\" data to upsample the imbalanced classed. \nFor example, simply creating flipped versions of the images will give you double of the number of images </p>\n\n<p>Happy Kaggling :) </p>",
      "rawMarkdown": "You should find ways to create \"synthetic\" data to upsample the imbalanced classed. \nFor example, simply creating flipped versions of the images will give you double of the number of images \n\nHappy Kaggling :) ",
      "replies": [
        {
          "id": 910135,
          "postDate": "2020-07-01T02:27:28.387Z",
          "content": "<p>hi，thanks for your help. It seems that  to get more positive samples is the best way for me to handle this problem. I’ve try more sophisticated CV strategy，classweight，and many other tips，totally failed！very depressing</p>",
          "rawMarkdown": "hi，thanks for your help. It seems that  to get more positive samples is the best way for me to handle this problem. I’ve try more sophisticated CV strategy，classweight，and many other tips，totally failed！very depressing"
        },
        {
          "id": 910262,
          "postDate": "2020-07-01T04:36:17.720Z",
          "content": "<p>Oh, misleading here.  Classweight is helpful, but not enough. </p>",
          "rawMarkdown": "Oh, misleading here.  Classweight is helpful, but not enough. "
        }
      ]
    },
    {
      "id": 903859,
      "postDate": "2020-06-27T06:57:09.307Z",
      "content": "<p>Yes it's a common thing in competitions. The most basic way to deal with it is you can start with Kfold cv validation. Anyway look at the works by others, many have contributed great notebooks about preprocessing the data and how to get good cv score.</p>",
      "rawMarkdown": "Yes it's a common thing in competitions. The most basic way to deal with it is you can start with Kfold cv validation. Anyway look at the works by others, many have contributed great notebooks about preprocessing the data and how to get good cv score.",
      "replies": [
        {
          "id": 903985,
          "postDate": "2020-06-27T09:01:42.300Z",
          "content": "<p>Thanks for your suggestion! That's really helpful</p>",
          "rawMarkdown": "Thanks for your suggestion! That's really helpful"
        }
      ]
    },
    {
      "id": 903732,
      "postDate": "2020-06-27T04:24:44.430Z",
      "content": "<p>Class inbalance is very typical issue in Kaggle competitions, if you are just starting I highly recommend to read previous deep learning competitions results to learn.</p>",
      "rawMarkdown": "Class inbalance is very typical issue in Kaggle competitions, if you are just starting I highly recommend to read previous deep learning competitions results to learn.",
      "replies": [
        {
          "id": 903994,
          "postDate": "2020-06-27T09:08:54.680Z",
          "content": "<p>👍 Thanks!</p>",
          "rawMarkdown": "👍 Thanks!"
        }
      ]
    },
    {
      "id": 903663,
      "postDate": "2020-06-27T03:19:45.527Z",
      "content": "<p>You may want to take a look at some public notebooks illustrating how to set up a reliable cross-validation on this data set. Here is, for example, my notebook: <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512?scriptVersionId=37515894\">EfficientNet BN+Tabular Features TF CV5 512x512</a>. You can use it as a starting point. And here is the discussion topic with more details about my approach: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395\">TF/ TPU: from .tfrec to cross-validation. Step-by-step</a>. </p>",
      "rawMarkdown": "You may want to take a look at some public notebooks illustrating how to set up a reliable cross-validation on this data set. Here is, for example, my notebook: [EfficientNet BN+Tabular Features TF CV5 512x512](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512?scriptVersionId=37515894). You can use it as a starting point. And here is the discussion topic with more details about my approach: [TF/ TPU: from .tfrec to cross-validation. Step-by-step](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395). ",
      "replies": [
        {
          "id": 903674,
          "postDate": "2020-06-27T03:30:10.480Z",
          "content": "<p>Thanks for your help. I'll try</p>",
          "rawMarkdown": "Thanks for your help. I'll try"
        }
      ]
    },
    {
      "id": 903661,
      "postDate": "2020-06-27T03:13:58.960Z",
      "content": "<p>I found the malicious skin data only hold  2% in training datasets, causing my model overfitting and produce zero for all test data, though I take the CV during training procedure. \nGuys, please help me! thanks</p>",
      "rawMarkdown": "I found the malicious skin data only hold  2% in training datasets, causing my model overfitting and produce zero for all test data, though I take the CV during training procedure. \nGuys, please help me! thanks"
    },
    {
      "id": 909621,
      "postDate": "2020-06-30T17:40:13.970Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 910125,
          "postDate": "2020-07-01T02:21:11.560Z",
          "content": "<p>very appreciate your sharing，friend！</p>",
          "rawMarkdown": "very appreciate your sharing，friend！"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 909391,
      "author_name": "AaronWard",
      "author_url": "",
      "post_date": "2020-06-30T15:05:17.730000",
      "content": "<p>You should find ways to create \"synthetic\" data to upsample the imbalanced classed. \nFor example, simply creating flipped versions of the images will give you double of the number of images </p>\n\n<p>Happy Kaggling :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 910135,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-07-01T02:27:28.387000",
          "content": "<p>hi，thanks for your help. It seems that  to get more positive samples is the best way for me to handle this problem. I’ve try more sophisticated CV strategy，classweight，and many other tips，totally failed！very depressing</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 910262,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-07-01T04:36:17.720000",
          "content": "<p>Oh, misleading here.  Classweight is helpful, but not enough. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903859,
      "author_name": "sayak",
      "author_url": "",
      "post_date": "2020-06-27T06:57:09.307000",
      "content": "<p>Yes it's a common thing in competitions. The most basic way to deal with it is you can start with Kfold cv validation. Anyway look at the works by others, many have contributed great notebooks about preprocessing the data and how to get good cv score.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 903985,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-06-27T09:01:42.300000",
          "content": "<p>Thanks for your suggestion! That's really helpful</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903732,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-06-27T04:24:44.430000",
      "content": "<p>Class inbalance is very typical issue in Kaggle competitions, if you are just starting I highly recommend to read previous deep learning competitions results to learn.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 903994,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-06-27T09:08:54.680000",
          "content": "<p>👍 Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 903663,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-06-27T03:19:45.527000",
      "content": "<p>You may want to take a look at some public notebooks illustrating how to set up a reliable cross-validation on this data set. Here is, for example, my notebook: <a href=\"https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512?scriptVersionId=37515894\">EfficientNet BN+Tabular Features TF CV5 512x512</a>. You can use it as a starting point. And here is the discussion topic with more details about my approach: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395\">TF/ TPU: from .tfrec to cross-validation. Step-by-step</a>. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 903674,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-06-27T03:30:10.480000",
          "content": "<p>Thanks for your help. I'll try</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 909621,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-30T17:40:13.970000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 910125,
          "author_name": "Matrix",
          "author_url": "",
          "post_date": "2020-07-01T02:21:11.560000",
          "content": "<p>very appreciate your sharing，friend！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "909391": "You should find ways to create \"synthetic\" data to upsample the imbalanced classed. \nFor example, simply creating flipped versions of the images will give you double of the number of images \n\nHappy Kaggling :) ",
    "903859": "Yes it's a common thing in competitions. The most basic way to deal with it is you can start with Kfold cv validation. Anyway look at the works by others, many have contributed great notebooks about preprocessing the data and how to get good cv score.",
    "903732": "Class inbalance is very typical issue in Kaggle competitions, if you are just starting I highly recommend to read previous deep learning competitions results to learn.",
    "903663": "You may want to take a look at some public notebooks illustrating how to set up a reliable cross-validation on this data set. Here is, for example, my notebook: [EfficientNet BN+Tabular Features TF CV5 512x512](https://www.kaggle.com/graf10a/efficientnet-bn-tabular-features-tf-cv5-512x512?scriptVersionId=37515894). You can use it as a starting point. And here is the discussion topic with more details about my approach: [TF/ TPU: from .tfrec to cross-validation. Step-by-step](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395). ",
    "903661": "I found the malicious skin data only hold  2% in training datasets, causing my model overfitting and produce zero for all test data, though I take the CV during training procedure. \nGuys, please help me! thanks",
    "909621": ""
  }
}