{
  "id": 164407,
  "title": "Looking for the optimal cross validation",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164407",
  "author_name": "cayala",
  "post_date": "2020-07-06T06:54:11.344000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I'm focused on making a realiable cross-validation (CV). I have experimented with 3, 4 and 5 folds by using a Stratified Gruped K-Fold strategy. Despite I'm getting a similar performance in train and valdations sets, the results obtained for test set are far from leaderboard. What I'm suppose to do? repeat the CV until getting the expected results?</p>\n\n<p>Here are some of my CV performance locally and their corresponding public ladeboard score:</p>\n\n<p>| local | public LB |\n|-------|-----------|\n| 0.629 | 0.686     |\n| 0.774 | 0.670     |\n| 0.800 | 0.652     |\n| 0.615 | 0.639     |\n| 0.740 | 0.664     |\n| 0.779 | 0.686     |\n| 0.794 | 0.638     |\n| 0.815 | 0.649     |\n| 0.835 | 0.633     |</p>\n\n<p>Thank you in advance.</p>",
  "messages": [
    {
      "id": 918803,
      "postDate": "2020-07-07T13:59:50.730Z",
      "content": "<p>The exact same thing is happening to us! We don't know what the problem is, there's apparently no leakage between our training and validation sets (we coded a loop that went though all the arrays on both sets and not two equal images were found). Let me know if you solve the problem, we'll do the same :)</p>",
      "rawMarkdown": "The exact same thing is happening to us! We don't know what the problem is, there's apparently no leakage between our training and validation sets (we coded a loop that went though all the arrays on both sets and not two equal images were found). Let me know if you solve the problem, we'll do the same :)",
      "votes": 2,
      "replies": [
        {
          "id": 920526,
          "postDate": "2020-07-08T16:43:49.333Z",
          "content": "<p>Until the next week I won't be able to try some ideas I have. Probably averagin multiple epochs would produce more estable results. However this approach would affect negatively to the score. </p>",
          "rawMarkdown": "Until the next week I won't be able to try some ideas I have. Probably averagin multiple epochs would produce more estable results. However this approach would affect negatively to the score. "
        },
        {
          "id": 920662,
          "postDate": "2020-07-08T18:33:06.713Z",
          "content": "<p>the instability is inherent to the evaluation metric and class imbalance with very few positive examples, a single false negative can have a big effect on the score.</p>",
          "rawMarkdown": "the instability is inherent to the evaluation metric and class imbalance with very few positive examples, a single false negative can have a big effect on the score."
        },
        {
          "id": 921836,
          "postDate": "2020-07-09T15:41:49.633Z",
          "content": "<p>i usually use mean of 5 fold. Not just predict of 1 fold. In my opinion, it's better to know about data, not overfit</p>",
          "rawMarkdown": "i usually use mean of 5 fold. Not just predict of 1 fold. In my opinion, it's better to know about data, not overfit"
        }
      ]
    },
    {
      "id": 917043,
      "postDate": "2020-07-06T07:43:55.097Z",
      "content": "<p>You could have a leak between training and validation and are overfitting the validation set as a result of it.</p>\n\n<p>I've used standard KFold CV and the results seem pretty stable (from last 5 single model submissions):</p>\n\n<p>| CV | public LB |\n| --- | --- |\n|  0.928 | 0.924 |\n| 0.915 | 0.915 |\n|0.922 | 0.922 |\n| 0.901 | 0.910 |\n|0.905 | 0.900 |</p>\n\n<p>One would expect patient-wise to be more robust.</p>\n\n<p>Edit: \nIn fact, the model is fairly low performing so maybe it is something else. Could you give a description of the model and if it's a deep learning model could you post a training curve?</p>",
      "rawMarkdown": "You could have a leak between training and validation and are overfitting the validation set as a result of it.\n\nI've used standard KFold CV and the results seem pretty stable (from last 5 single model submissions):\n\n| CV | public LB |\n| --- | --- |\n|  0.928 | 0.924 |\n| 0.915 | 0.915 |\n|0.922 | 0.922 |\n| 0.901 | 0.910 |\n|0.905 | 0.900 |\n\nOne would expect patient-wise to be more robust.\n\nEdit: \nIn fact, the model is fairly low performing so maybe it is something else. Could you give a description of the model and if it's a deep learning model could you post a training curve?",
      "votes": 2,
      "replies": [
        {
          "id": 917059,
          "postDate": "2020-07-06T07:55:56.470Z",
          "content": "<p>I'm training a ResNet-34 on 128x128 images during 3 epochs. I'm also using Adam(1e-3) as optimizer and FocalLoss(2.0, 0.25) as loss function. Moreover, I'm not applying any regularization technique.</p>\n\n<p>The CV has been built using <em>sex</em> and *anatom_site_general_challenge* as feature columns as well as \"target\" and \"patient_id\" (following <a href=\"/meemr5\">@meemr5</a> <a href=\"https://www.kaggle.com/meemr5/melanoma-classification-eda-sgkfold-logisticr/notebook#5.-Validation-Strategy---Stratified-Group-K-Fold\">notebook</a>)</p>",
          "rawMarkdown": "I'm training a ResNet-34 on 128x128 images during 3 epochs. I'm also using Adam(1e-3) as optimizer and FocalLoss(2.0, 0.25) as loss function. Moreover, I'm not applying any regularization technique.\n\nThe CV has been built using *sex* and *anatom_site_general_challenge* as feature columns as well as \"target\" and \"patient_id\" (following @meemr5 [notebook]( https://www.kaggle.com/meemr5/melanoma-classification-eda-sgkfold-logisticr/notebook#5.-Validation-Strategy---Stratified-Group-K-Fold))"
        },
        {
          "id": 917087,
          "postDate": "2020-07-06T08:25:57.253Z",
          "content": "<p>Thanks, that's useful. Look at your training/validation loss curve  and try training for 15 epochs and see if that improves things; I'm guessing the number of epochs is too small.</p>",
          "rawMarkdown": "Thanks, that's useful. Look at your training/validation loss curve  and try training for 15 epochs and see if that improves things; I'm guessing the number of epochs is too small."
        },
        {
          "id": 917154,
          "postDate": "2020-07-06T09:21:31.413Z",
          "content": "<p>Hey, <a href=\"/cayala\">@cayala</a>!\nI'm using SGKFOLD as you mentioned. </p>\n\n<hr>\n\n<p>Here are my CV and LB Score:\n| CV - SGKFOLD |LB |\n| --- | --- |\n|0.8980|0.859|\n|  0.912220621897585|  0.884|\n|0.9165890794989842|0.903|\n|0.9059320092201233|0.910|\n|0.8970596194267273|0.894|</p>\n\n<p>I'm using TTA-10 for test-set predictions. Can't say where you are making a mistake.</p>",
          "rawMarkdown": "Hey, @cayala!\nI'm using SGKFOLD as you mentioned. \n\n---\n\nHere are my CV and LB Score:\n| CV - SGKFOLD |LB |\n| --- | --- |\n|0.8980|0.859|\n|  0.912220621897585|  0.884|\n|0.9165890794989842|0.903|\n|0.9059320092201233|0.910|\n|0.8970596194267273|0.894|\n\nI'm using TTA-10 for test-set predictions. Can't say where you are making a mistake."
        },
        {
          "id": 917162,
          "postDate": "2020-07-06T09:26:29.453Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 917187,
          "postDate": "2020-07-06T09:49:19.107Z",
          "content": "<p>I'm using standard Kfold at the TFRecord level using the records <a href=\"/cdeotte\">@cdeotte</a> made (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>). </p>\n\n<p>There are some results I don't understand (e.g., CV : 0.901 and LB 0.910).</p>\n\n<p>Chris may have grouped patients into the same TFRecord file, in that case I would have unintentionally been doing patient-wise cross validation, I will check this today.</p>\n\n<p>EDIT: </p>\n\n<p>Just checked, the TFRecords are not separated by patient_id so I am just using standard KFold. Definitely not recommended to use, I see no reason not to use patient-wise CV.</p>",
          "rawMarkdown": "I'm using standard Kfold at the TFRecord level using the records @cdeotte made ([here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579)). \n\nThere are some results I don't understand (e.g., CV : 0.901 and LB 0.910).\n\nChris may have grouped patients into the same TFRecord file, in that case I would have unintentionally been doing patient-wise cross validation, I will check this today.\n\nEDIT: \n\nJust checked, the TFRecords are not separated by patient_id so I am just using standard KFold. Definitely not recommended to use, I see no reason not to use patient-wise CV.",
          "votes": 1
        }
      ]
    },
    {
      "id": 916984,
      "postDate": "2020-07-06T06:54:11.343Z",
      "content": "<p>I'm focused on making a realiable cross-validation (CV). I have experimented with 3, 4 and 5 folds by using a Stratified Gruped K-Fold strategy. Despite I'm getting a similar performance in train and valdations sets, the results obtained for test set are far from leaderboard. What I'm suppose to do? repeat the CV until getting the expected results?</p>\n\n<p>Here are some of my CV performance locally and their corresponding public ladeboard score:</p>\n\n<p>| local | public LB |\n|-------|-----------|\n| 0.629 | 0.686     |\n| 0.774 | 0.670     |\n| 0.800 | 0.652     |\n| 0.615 | 0.639     |\n| 0.740 | 0.664     |\n| 0.779 | 0.686     |\n| 0.794 | 0.638     |\n| 0.815 | 0.649     |\n| 0.835 | 0.633     |</p>\n\n<p>Thank you in advance.</p>",
      "rawMarkdown": "I'm focused on making a realiable cross-validation (CV). I have experimented with 3, 4 and 5 folds by using a Stratified Gruped K-Fold strategy. Despite I'm getting a similar performance in train and valdations sets, the results obtained for test set are far from leaderboard. What I'm suppose to do? repeat the CV until getting the expected results?\n\nHere are some of my CV performance locally and their corresponding public ladeboard score:\n\n| local | public LB |\n|-------|-----------|\n| 0.629 | 0.686     |\n| 0.774 | 0.670     |\n| 0.800 | 0.652     |\n| 0.615 | 0.639     |\n| 0.740 | 0.664     |\n| 0.779 | 0.686     |\n| 0.794 | 0.638     |\n| 0.815 | 0.649     |\n| 0.835 | 0.633     |\n\nThank you in advance.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 918803,
      "author_name": "Fernando Gaston Codony",
      "author_url": "",
      "post_date": "2020-07-07T13:59:50.730000",
      "content": "<p>The exact same thing is happening to us! We don't know what the problem is, there's apparently no leakage between our training and validation sets (we coded a loop that went though all the arrays on both sets and not two equal images were found). Let me know if you solve the problem, we'll do the same :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 920526,
          "author_name": "cayala",
          "author_url": "",
          "post_date": "2020-07-08T16:43:49.333000",
          "content": "<p>Until the next week I won't be able to try some ideas I have. Probably averagin multiple epochs would produce more estable results. However this approach would affect negatively to the score. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 920662,
          "author_name": "alaa",
          "author_url": "",
          "post_date": "2020-07-08T18:33:06.713000",
          "content": "<p>the instability is inherent to the evaluation metric and class imbalance with very few positive examples, a single false negative can have a big effect on the score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 921836,
          "author_name": "Manh Lab",
          "author_url": "",
          "post_date": "2020-07-09T15:41:49.633000",
          "content": "<p>i usually use mean of 5 fold. Not just predict of 1 fold. In my opinion, it's better to know about data, not overfit</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 917043,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2020-07-06T07:43:55.097000",
      "content": "<p>You could have a leak between training and validation and are overfitting the validation set as a result of it.</p>\n\n<p>I've used standard KFold CV and the results seem pretty stable (from last 5 single model submissions):</p>\n\n<p>| CV | public LB |\n| --- | --- |\n|  0.928 | 0.924 |\n| 0.915 | 0.915 |\n|0.922 | 0.922 |\n| 0.901 | 0.910 |\n|0.905 | 0.900 |</p>\n\n<p>One would expect patient-wise to be more robust.</p>\n\n<p>Edit: \nIn fact, the model is fairly low performing so maybe it is something else. Could you give a description of the model and if it's a deep learning model could you post a training curve?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 917059,
          "author_name": "cayala",
          "author_url": "",
          "post_date": "2020-07-06T07:55:56.470000",
          "content": "<p>I'm training a ResNet-34 on 128x128 images during 3 epochs. I'm also using Adam(1e-3) as optimizer and FocalLoss(2.0, 0.25) as loss function. Moreover, I'm not applying any regularization technique.</p>\n\n<p>The CV has been built using <em>sex</em> and *anatom_site_general_challenge* as feature columns as well as \"target\" and \"patient_id\" (following <a href=\"/meemr5\">@meemr5</a> <a href=\"https://www.kaggle.com/meemr5/melanoma-classification-eda-sgkfold-logisticr/notebook#5.-Validation-Strategy---Stratified-Group-K-Fold\">notebook</a>)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917087,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-07-06T08:25:57.253000",
          "content": "<p>Thanks, that's useful. Look at your training/validation loss curve  and try training for 15 epochs and see if that improves things; I'm guessing the number of epochs is too small.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917154,
          "author_name": "Meet Ranoliya",
          "author_url": "",
          "post_date": "2020-07-06T09:21:31.413000",
          "content": "<p>Hey, <a href=\"/cayala\">@cayala</a>!\nI'm using SGKFOLD as you mentioned. </p>\n\n<hr>\n\n<p>Here are my CV and LB Score:\n| CV - SGKFOLD |LB |\n| --- | --- |\n|0.8980|0.859|\n|  0.912220621897585|  0.884|\n|0.9165890794989842|0.903|\n|0.9059320092201233|0.910|\n|0.8970596194267273|0.894|</p>\n\n<p>I'm using TTA-10 for test-set predictions. Can't say where you are making a mistake.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917162,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-06T09:26:29.453000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 917187,
          "author_name": "FChmiel",
          "author_url": "",
          "post_date": "2020-07-06T09:49:19.107000",
          "content": "<p>I'm using standard Kfold at the TFRecord level using the records <a href=\"/cdeotte\">@cdeotte</a> made (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>). </p>\n\n<p>There are some results I don't understand (e.g., CV : 0.901 and LB 0.910).</p>\n\n<p>Chris may have grouped patients into the same TFRecord file, in that case I would have unintentionally been doing patient-wise cross validation, I will check this today.</p>\n\n<p>EDIT: </p>\n\n<p>Just checked, the TFRecords are not separated by patient_id so I am just using standard KFold. Definitely not recommended to use, I see no reason not to use patient-wise CV.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "918803": "The exact same thing is happening to us! We don't know what the problem is, there's apparently no leakage between our training and validation sets (we coded a loop that went though all the arrays on both sets and not two equal images were found). Let me know if you solve the problem, we'll do the same :)",
    "917043": "You could have a leak between training and validation and are overfitting the validation set as a result of it.\n\nI've used standard KFold CV and the results seem pretty stable (from last 5 single model submissions):\n\n| CV | public LB |\n| --- | --- |\n|  0.928 | 0.924 |\n| 0.915 | 0.915 |\n|0.922 | 0.922 |\n| 0.901 | 0.910 |\n|0.905 | 0.900 |\n\nOne would expect patient-wise to be more robust.\n\nEdit: \nIn fact, the model is fairly low performing so maybe it is something else. Could you give a description of the model and if it's a deep learning model could you post a training curve?",
    "916984": "I'm focused on making a realiable cross-validation (CV). I have experimented with 3, 4 and 5 folds by using a Stratified Gruped K-Fold strategy. Despite I'm getting a similar performance in train and valdations sets, the results obtained for test set are far from leaderboard. What I'm suppose to do? repeat the CV until getting the expected results?\n\nHere are some of my CV performance locally and their corresponding public ladeboard score:\n\n| local | public LB |\n|-------|-----------|\n| 0.629 | 0.686     |\n| 0.774 | 0.670     |\n| 0.800 | 0.652     |\n| 0.615 | 0.639     |\n| 0.740 | 0.664     |\n| 0.779 | 0.686     |\n| 0.794 | 0.638     |\n| 0.815 | 0.649     |\n| 0.835 | 0.633     |\n\nThank you in advance."
  }
}