{
  "id": 79620,
  "title": "Validation AUC score?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/79620",
  "author_name": "David J. Slate",
  "post_date": "2019-02-06T01:55:34.770000",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Has anyone calculated AUC scores between raw numeric predictions and labels for validation data?  I'm getting what look like ridiculously high AUC scores, like 0.9849, together with MCC scores like 0.712 that are much higher than my LB scores.  Sounds like this should be due to classic overfitting of validation data and/or some leak from training to validation, but so far I haven't found the cause.</p>",
  "messages": [
    {
      "id": 467093,
      "postDate": "2019-02-06T12:48:48.553Z",
      "content": "<p>Do you use early stop of some kind? I discovered that It's pretty easy to overfit on validation when you have only 8K samples, thousands noisy features and very sensitive and unstable metric. But maybe it's just a lack of my skills.</p>",
      "rawMarkdown": "Do you use early stop of some kind? I discovered that It's pretty easy to overfit on validation when you have only 8K samples, thousands noisy features and very sensitive and unstable metric. But maybe it's just a lack of my skills.",
      "votes": 5,
      "replies": [
        {
          "id": 467196,
          "postDate": "2019-02-06T16:07:56.280Z",
          "content": "<p>Same thing ! I tried using acc loss instead of MMC to detect the early stopping, but it gives me similar results : it's very unstable.</p>",
          "rawMarkdown": "Same thing ! I tried using acc loss instead of MMC to detect the early stopping, but it gives me similar results : it's very unstable.",
          "votes": 1
        }
      ]
    },
    {
      "id": 466823,
      "postDate": "2019-02-06T01:55:34.770Z",
      "content": "<p>Has anyone calculated AUC scores between raw numeric predictions and labels for validation data?  I'm getting what look like ridiculously high AUC scores, like 0.9849, together with MCC scores like 0.712 that are much higher than my LB scores.  Sounds like this should be due to classic overfitting of validation data and/or some leak from training to validation, but so far I haven't found the cause.</p>",
      "rawMarkdown": "Has anyone calculated AUC scores between raw numeric predictions and labels for validation data?  I'm getting what look like ridiculously high AUC scores, like 0.9849, together with MCC scores like 0.712 that are much higher than my LB scores.  Sounds like this should be due to classic overfitting of validation data and/or some leak from training to validation, but so far I haven't found the cause.",
      "votes": 2
    },
    {
      "id": 466930,
      "postDate": "2019-02-06T08:18:31.973Z",
      "content": "<p>maybe your validation is  \"in sample\".  I mean measurement_id is spitted between training set and validation set. </p>",
      "rawMarkdown": "maybe your validation is  \"in sample\".  I mean measurement_id is spitted between training set and validation set. ",
      "replies": [
        {
          "id": 466996,
          "postDate": "2019-02-06T10:00:07.567Z",
          "content": "<p>Thanks Jonas, but I split the data so that the signals from each measurement are kept together in either the training or validation set.  I was hoping the problem was that simple, but I did some printouts and it looks ok.</p>",
          "rawMarkdown": "Thanks Jonas, but I split the data so that the signals from each measurement are kept together in either the training or validation set.  I was hoping the problem was that simple, but I did some printouts and it looks ok."
        },
        {
          "id": 467046,
          "postDate": "2019-02-06T11:33:07.763Z",
          "content": "<p>This competition is really mysterious :) And it seems for all, many already commented the same problem... I can tell one more my secret - the best LB score I get after 1 epoch . </p>",
          "rawMarkdown": "This competition is really mysterious :) And it seems for all, many already commented the same problem... I can tell one more my secret - the best LB score I get after 1 epoch . "
        }
      ]
    },
    {
      "id": 466923,
      "postDate": "2019-02-06T08:02:05.897Z",
      "content": "<p>I am also very uncertain about my validation score vs lb score, I just had 2 models, both trained on slightly different features, both had ~.72 MCC validation score. One scored 0.64 on lb, other 0.42. </p>",
      "rawMarkdown": "I am also very uncertain about my validation score vs lb score, I just had 2 models, both trained on slightly different features, both had ~.72 MCC validation score. One scored 0.64 on lb, other 0.42. "
    },
    {
      "id": 468869,
      "postDate": "2019-02-09T22:33:34.527Z",
      "content": "<p>Same for me, with local CV I get AUC &gt; 0.97 and MCC &gt; 0.68 while on LB MCC = 0.52 or less.</p>\n\n<p>Currently I am working with simple FE and XGB but is does not explain the discrepancy.\nStill working on it.\nMaybe I made a bad rule for selection of the threshold to binarize predictions?\nMaybe proportion of PD signals is much lower in test set ?</p>",
      "rawMarkdown": "Same for me, with local CV I get AUC &gt; 0.97 and MCC &gt; 0.68 while on LB MCC = 0.52 or less.\n\nCurrently I am working with simple FE and XGB but is does not explain the discrepancy.\nStill working on it.\nMaybe I made a bad rule for selection of the threshold to binarize predictions?\nMaybe proportion of PD signals is much lower in test set ?",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 467093,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2019-02-06T12:48:48.553000",
      "content": "<p>Do you use early stop of some kind? I discovered that It's pretty easy to overfit on validation when you have only 8K samples, thousands noisy features and very sensitive and unstable metric. But maybe it's just a lack of my skills.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 467196,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2019-02-06T16:07:56.280000",
          "content": "<p>Same thing ! I tried using acc loss instead of MMC to detect the early stopping, but it gives me similar results : it's very unstable.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 466930,
      "author_name": "Jonas Matuzas",
      "author_url": "",
      "post_date": "2019-02-06T08:18:31.973000",
      "content": "<p>maybe your validation is  \"in sample\".  I mean measurement_id is spitted between training set and validation set. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 466996,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2019-02-06T10:00:07.567000",
          "content": "<p>Thanks Jonas, but I split the data so that the signals from each measurement are kept together in either the training or validation set.  I was hoping the problem was that simple, but I did some printouts and it looks ok.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467046,
          "author_name": "Jonas Matuzas",
          "author_url": "",
          "post_date": "2019-02-06T11:33:07.763000",
          "content": "<p>This competition is really mysterious :) And it seems for all, many already commented the same problem... I can tell one more my secret - the best LB score I get after 1 epoch . </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 466923,
      "author_name": "MartijnNanne",
      "author_url": "",
      "post_date": "2019-02-06T08:02:05.897000",
      "content": "<p>I am also very uncertain about my validation score vs lb score, I just had 2 models, both trained on slightly different features, both had ~.72 MCC validation score. One scored 0.64 on lb, other 0.42. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 468869,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-09T22:33:34.527000",
      "content": "<p>Same for me, with local CV I get AUC &gt; 0.97 and MCC &gt; 0.68 while on LB MCC = 0.52 or less.</p>\n\n<p>Currently I am working with simple FE and XGB but is does not explain the discrepancy.\nStill working on it.\nMaybe I made a bad rule for selection of the threshold to binarize predictions?\nMaybe proportion of PD signals is much lower in test set ?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "467093": "Do you use early stop of some kind? I discovered that It's pretty easy to overfit on validation when you have only 8K samples, thousands noisy features and very sensitive and unstable metric. But maybe it's just a lack of my skills.",
    "466823": "Has anyone calculated AUC scores between raw numeric predictions and labels for validation data?  I'm getting what look like ridiculously high AUC scores, like 0.9849, together with MCC scores like 0.712 that are much higher than my LB scores.  Sounds like this should be due to classic overfitting of validation data and/or some leak from training to validation, but so far I haven't found the cause.",
    "466930": "maybe your validation is  \"in sample\".  I mean measurement_id is spitted between training set and validation set. ",
    "466923": "I am also very uncertain about my validation score vs lb score, I just had 2 models, both trained on slightly different features, both had ~.72 MCC validation score. One scored 0.64 on lb, other 0.42. ",
    "468869": "Same for me, with local CV I get AUC &gt; 0.97 and MCC &gt; 0.68 while on LB MCC = 0.52 or less.\n\nCurrently I am working with simple FE and XGB but is does not explain the discrepancy.\nStill working on it.\nMaybe I made a bad rule for selection of the threshold to binarize predictions?\nMaybe proportion of PD signals is much lower in test set ?"
  }
}