{
  "id": 16669,
  "title": "Anyone else having trouble with cross-validation?",
  "url": "/competitions/dato-native/discussion/16669",
  "author_name": "",
  "post_date": "2015-09-25T19:56:16.413Z",
  "votes": 2,
  "comment_count": 9,
  "views": 1619,
  "content": "<p>My internal CV is giving me ~.94, but I keep getting around .68 on the leaderboard.</p>\n\n<p>I'm starting to thing I screwed something up when partitioning the files....</p>",
  "messages": [
    {
      "id": "93438",
      "postDate": "09/25/2015 19:56:16",
      "content": "<p>My internal CV is giving me ~.94, but I keep getting around .68 on the leaderboard.</p>\n\n<p>I'm starting to thing I screwed something up when partitioning the files....</p>",
      "rawMarkdown": "My internal CV is giving me ~.94, but I keep getting around .68 on the leaderboard.\r\n\r\nI'm starting to thing I screwed something up when partitioning the files....",
      "votes": null
    },
    {
      "id": "93439",
      "postDate": "09/25/2015 20:07:45",
      "content": "<p>I use stratified 10-fold cross-validation, and my LB score is slightly higher than the average of those folds, and has been for every submission before and after the competition reset.</p>\n\n<p>The difference between the highest and lowest scoring folds is less than 0.005.</p>",
      "rawMarkdown": "I use stratified 10-fold cross-validation, and my LB score is slightly higher than the average of those folds, and has been for every submission before and after the competition reset.\r\n\r\nThe difference between the highest and lowest scoring folds is less than 0.005.",
      "votes": null
    },
    {
      "id": "93441",
      "postDate": "09/25/2015 20:17:22",
      "content": "<p>Ok, I have to check my pre-processing code.  I must have something wrong.</p>",
      "rawMarkdown": "Ok, I have to check my pre-processing code.  I must have something wrong.",
      "votes": null
    },
    {
      "id": "93474",
      "postDate": "09/26/2015 12:05:44",
      "content": "<p>Man, this is frustrating.  I re-built my whole dataset from scratch, CV-d a a simple model, and got an AUC of .89 on my CV and on the leaderboard.</p>\n\n<p>Same dataset, more complex model, .94 CV and .89 leaderboard.  I'm not even using any crazy features: it's basically the set <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">from this post</a>, plus some SVD features made from a sparse bag-of-words matrix on all the HTML files.</p>",
      "rawMarkdown": "Man, this is frustrating.  I re-built my whole dataset from scratch, CV-d a a simple model, and got an AUC of .89 on my CV and on the leaderboard.\r\n\r\nSame dataset, more complex model, .94 CV and .89 leaderboard.  I'm not even using any crazy features: it's basically the set [from this post][1], plus some SVD features made from a sparse bag-of-words matrix on all the HTML files.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
      "votes": null
    },
    {
      "id": "93514",
      "postDate": "09/26/2015 23:25:30",
      "content": "<p>I'm getting similar scores on both so far.. difference floats around 0.005, same as mortehu's</p>",
      "rawMarkdown": "I'm getting similar scores on both so far.. difference floats around 0.005, same as mortehu's",
      "votes": null
    },
    {
      "id": "93668",
      "postDate": "09/29/2015 22:54:47",
      "content": "<p>I took a new approach, and now my CV scores are very close to my actual leaderboard scores.</p>\n\n<p>I suspect there may have been a bug in my old code, or perhaps I had a very interesting way of over-fitting.</p>",
      "rawMarkdown": "I took a new approach, and now my CV scores are very close to my actual leaderboard scores.\r\n\r\nI suspect there may have been a bug in my old code, or perhaps I had a very interesting way of over-fitting.",
      "votes": null
    },
    {
      "id": "95834",
      "postDate": "10/12/2015 15:35:48",
      "content": "<p>I'm intrigued, would you mind sharing what you've found as interesting way of over-fitting?</p>\n\n<p>In exchange, here's my usual suspects:</p>\n\n<p>I often encountered that roc_auc is somewhat too easy to get high (pun?), especially for certain models that can usually reach 1.0 by its nature, like xgboost, libffm, CRFs, etc. So firstly I would switch to other metrics like <strong>log loss</strong>.</p>\n\n<p>Secondly when evaluate CV result there are some <strong>assumptions might be problematic</strong>, e.g. <a href=\"https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables\">https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables</a> or (im-)balanced classes.</p>\n\n<p>Finally, as for my &quot;favorite&quot; one, especially in NLP, something is just <strong>unseen</strong>, or at least not well recognized and eventually became friendly noise (I'm sure there's no such term but hopefully you will get it).</p>\n\n<p>In the end all the above is just the same thing to me, but I am always anxious about that I could be too naive or even totally wrong.</p>",
      "rawMarkdown": "I'm intrigued, would you mind sharing what you've found as interesting way of over-fitting?\r\n\r\nIn exchange, here's my usual suspects:\r\n\r\nI often encountered that roc_auc is somewhat too easy to get high (pun?), especially for certain models that can usually reach 1.0 by its nature, like xgboost, libffm, CRFs, etc. So firstly I would switch to other metrics like **log loss**.\r\n\r\nSecondly when evaluate CV result there are some **assumptions might be problematic**, e.g. https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables or (im-)balanced classes.\r\n\r\nFinally, as for my \"favorite\" one, especially in NLP, something is just **unseen**, or at least not well recognized and eventually became friendly noise (I'm sure there's no such term but hopefully you will get it).\r\n\r\nIn the end all the above is just the same thing to me, but I am always anxious about that I could be too naive or even totally wrong.",
      "votes": null
    },
    {
      "id": "95839",
      "postDate": "10/12/2015 15:52:24",
      "content": "<p>So I found that bag-of-words + SVD + a linear model resulted in overfitting, but bag-of-words + a linear model did not.  In the former case I got an AUC of .95 on cross-validation and .68 on the leaderboard, in the latter case I got .95 on cross-validation and .94 on the leaderboard.</p>\n\n<p>I can't figure out why this would be the case, and it seems counter-intuitive to me.  The only possibility I can think of at the moment is I screwed something up with the ordering of my dataset during the SVD model, but I don't have time to go back and confirm.</p>",
      "rawMarkdown": "So I found that bag-of-words + SVD + a linear model resulted in overfitting, but bag-of-words + a linear model did not.  In the former case I got an AUC of .95 on cross-validation and .68 on the leaderboard, in the latter case I got .95 on cross-validation and .94 on the leaderboard.\r\n\r\nI can't figure out why this would be the case, and it seems counter-intuitive to me.  The only possibility I can think of at the moment is I screwed something up with the ordering of my dataset during the SVD model, but I don't have time to go back and confirm.",
      "votes": null
    },
    {
      "id": "95846",
      "postDate": "10/12/2015 16:30:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zachmayer\">@Zach</a>:</p>\n\n<p>I see, thank you for your kind explanation. Speaking of SVD I have an obviously unchecked theory. It reminds me that when word2vec/GloVe got hot, people wondered how come such approach so similar to SVD could result in better performance for certain applications. Probably a half year later a brilliant researcher figured out that word2vec/GloVe just got a (lucky) bias which fit the task. On the other hand, the famous libffm is also a kind of matrix decomposition, and it overfits a lot. So, my wild guess is that SVD gave you a bad bias.</p>",
      "rawMarkdown": "Hi [@Zach][1]:\r\n\r\nI see, thank you for your kind explanation. Speaking of SVD I have an obviously unchecked theory. It reminds me that when word2vec/GloVe got hot, people wondered how come such approach so similar to SVD could result in better performance for certain applications. Probably a half year later a brilliant researcher figured out that word2vec/GloVe just got a (lucky) bias which fit the task. On the other hand, the famous libffm is also a kind of matrix decomposition, and it overfits a lot. So, my wild guess is that SVD gave you a bad bias.\r\n\r\n  [1]: https://www.kaggle.com/zachmayer",
      "votes": null
    },
    {
      "id": "95904",
      "postDate": "10/12/2015 22:20:03",
      "content": "<p>Hi Zach,</p>\n\n<p>When you calculated your SVD, did you do it on a combined table of train + test data, or on just the train data? I think the right way to do SVD is to calculate on the train data, and then rotate the test data. If you run SVD on the combined data table, then you are finding features that are in both data-sets, but you only have valid classification signal for the train data. It makes sense to me that you would over-fit to the train data in a situation like that.</p>",
      "rawMarkdown": "Hi Zach,\r\n\r\nWhen you calculated your SVD, did you do it on a combined table of train + test data, or on just the train data? I think the right way to do SVD is to calculate on the train data, and then rotate the test data. If you run SVD on the combined data table, then you are finding features that are in both data-sets, but you only have valid classification signal for the train data. It makes sense to me that you would over-fit to the train data in a situation like that.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 93439,
      "author_name": "mortehu",
      "author_url": "",
      "post_date": "09/25/2015 20:07:45",
      "content": "<p>I use stratified 10-fold cross-validation, and my LB score is slightly higher than the average of those folds, and has been for every submission before and after the competition reset.</p>\n\n<p>The difference between the highest and lowest scoring folds is less than 0.005.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93441,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "09/25/2015 20:17:22",
      "content": "<p>Ok, I have to check my pre-processing code.  I must have something wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93474,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "09/26/2015 12:05:44",
      "content": "<p>Man, this is frustrating.  I re-built my whole dataset from scratch, CV-d a a simple model, and got an AUC of .89 on my CV and on the leaderboard.</p>\n\n<p>Same dataset, more complex model, .94 CV and .89 leaderboard.  I'm not even using any crazy features: it's basically the set <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">from this post</a>, plus some SVD features made from a sparse bag-of-words matrix on all the HTML files.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93514,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "09/26/2015 23:25:30",
      "content": "<p>I'm getting similar scores on both so far.. difference floats around 0.005, same as mortehu's</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93668,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "09/29/2015 22:54:47",
      "content": "<p>I took a new approach, and now my CV scores are very close to my actual leaderboard scores.</p>\n\n<p>I suspect there may have been a bug in my old code, or perhaps I had a very interesting way of over-fitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95834,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/12/2015 15:35:48",
      "content": "<p>I'm intrigued, would you mind sharing what you've found as interesting way of over-fitting?</p>\n\n<p>In exchange, here's my usual suspects:</p>\n\n<p>I often encountered that roc_auc is somewhat too easy to get high (pun?), especially for certain models that can usually reach 1.0 by its nature, like xgboost, libffm, CRFs, etc. So firstly I would switch to other metrics like <strong>log loss</strong>.</p>\n\n<p>Secondly when evaluate CV result there are some <strong>assumptions might be problematic</strong>, e.g. <a href=\"https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables\">https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables</a> or (im-)balanced classes.</p>\n\n<p>Finally, as for my &quot;favorite&quot; one, especially in NLP, something is just <strong>unseen</strong>, or at least not well recognized and eventually became friendly noise (I'm sure there's no such term but hopefully you will get it).</p>\n\n<p>In the end all the above is just the same thing to me, but I am always anxious about that I could be too naive or even totally wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95839,
      "author_name": "zachmayer",
      "author_url": "",
      "post_date": "10/12/2015 15:52:24",
      "content": "<p>So I found that bag-of-words + SVD + a linear model resulted in overfitting, but bag-of-words + a linear model did not.  In the former case I got an AUC of .95 on cross-validation and .68 on the leaderboard, in the latter case I got .95 on cross-validation and .94 on the leaderboard.</p>\n\n<p>I can't figure out why this would be the case, and it seems counter-intuitive to me.  The only possibility I can think of at the moment is I screwed something up with the ordering of my dataset during the SVD model, but I don't have time to go back and confirm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95846,
      "author_name": "tmjiang",
      "author_url": "",
      "post_date": "10/12/2015 16:30:17",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zachmayer\">@Zach</a>:</p>\n\n<p>I see, thank you for your kind explanation. Speaking of SVD I have an obviously unchecked theory. It reminds me that when word2vec/GloVe got hot, people wondered how come such approach so similar to SVD could result in better performance for certain applications. Probably a half year later a brilliant researcher figured out that word2vec/GloVe just got a (lucky) bias which fit the task. On the other hand, the famous libffm is also a kind of matrix decomposition, and it overfits a lot. So, my wild guess is that SVD gave you a bad bias.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95904,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "10/12/2015 22:20:03",
      "content": "<p>Hi Zach,</p>\n\n<p>When you calculated your SVD, did you do it on a combined table of train + test data, or on just the train data? I think the right way to do SVD is to calculate on the train data, and then rotate the test data. If you run SVD on the combined data table, then you are finding features that are in both data-sets, but you only have valid classification signal for the train data. It makes sense to me that you would over-fit to the train data in a situation like that.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "93438": "My internal CV is giving me ~.94, but I keep getting around .68 on the leaderboard.\r\n\r\nI'm starting to thing I screwed something up when partitioning the files....",
    "93439": "I use stratified 10-fold cross-validation, and my LB score is slightly higher than the average of those folds, and has been for every submission before and after the competition reset.\r\n\r\nThe difference between the highest and lowest scoring folds is less than 0.005.",
    "93441": "Ok, I have to check my pre-processing code.  I must have something wrong.",
    "93474": "Man, this is frustrating.  I re-built my whole dataset from scratch, CV-d a a simple model, and got an AUC of .89 on my CV and on the leaderboard.\r\n\r\nSame dataset, more complex model, .94 CV and .89 leaderboard.  I'm not even using any crazy features: it's basically the set [from this post][1], plus some SVD features made from a sparse bag-of-words matrix on all the HTML files.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
    "93514": "I'm getting similar scores on both so far.. difference floats around 0.005, same as mortehu's",
    "93668": "I took a new approach, and now my CV scores are very close to my actual leaderboard scores.\r\n\r\nI suspect there may have been a bug in my old code, or perhaps I had a very interesting way of over-fitting.",
    "95834": "I'm intrigued, would you mind sharing what you've found as interesting way of over-fitting?\r\n\r\nIn exchange, here's my usual suspects:\r\n\r\nI often encountered that roc_auc is somewhat too easy to get high (pun?), especially for certain models that can usually reach 1.0 by its nature, like xgboost, libffm, CRFs, etc. So firstly I would switch to other metrics like **log loss**.\r\n\r\nSecondly when evaluate CV result there are some **assumptions might be problematic**, e.g. https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables or (im-)balanced classes.\r\n\r\nFinally, as for my \"favorite\" one, especially in NLP, something is just **unseen**, or at least not well recognized and eventually became friendly noise (I'm sure there's no such term but hopefully you will get it).\r\n\r\nIn the end all the above is just the same thing to me, but I am always anxious about that I could be too naive or even totally wrong.",
    "95839": "So I found that bag-of-words + SVD + a linear model resulted in overfitting, but bag-of-words + a linear model did not.  In the former case I got an AUC of .95 on cross-validation and .68 on the leaderboard, in the latter case I got .95 on cross-validation and .94 on the leaderboard.\r\n\r\nI can't figure out why this would be the case, and it seems counter-intuitive to me.  The only possibility I can think of at the moment is I screwed something up with the ordering of my dataset during the SVD model, but I don't have time to go back and confirm.",
    "95846": "Hi [@Zach][1]:\r\n\r\nI see, thank you for your kind explanation. Speaking of SVD I have an obviously unchecked theory. It reminds me that when word2vec/GloVe got hot, people wondered how come such approach so similar to SVD could result in better performance for certain applications. Probably a half year later a brilliant researcher figured out that word2vec/GloVe just got a (lucky) bias which fit the task. On the other hand, the famous libffm is also a kind of matrix decomposition, and it overfits a lot. So, my wild guess is that SVD gave you a bad bias.\r\n\r\n  [1]: https://www.kaggle.com/zachmayer",
    "95904": "Hi Zach,\r\n\r\nWhen you calculated your SVD, did you do it on a combined table of train + test data, or on just the train data? I think the right way to do SVD is to calculate on the train data, and then rotate the test data. If you run SVD on the combined data table, then you are finding features that are in both data-sets, but you only have valid classification signal for the train data. It makes sense to me that you would over-fit to the train data in a situation like that."
  },
  "source": "meta"
}