{
  "id": 193348,
  "title": "Evaluating model?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/193348",
  "author_name": "",
  "post_date": "2020-10-26T16:34:50.873627400Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>So, my model that's clearly overfit has a better public score than my meticulously not overfitted model. The not-overfitted model had less binary log-loss and better AUC and performed worse. So how can I actually tell if my model is getting better? </p>\n<p>Btw, my CV set is the last records of each student</p>",
  "messages": [
    {
      "id": "1060953",
      "postDate": "10/26/2020 16:34:50",
      "content": "<p>So, my model that's clearly overfit has a better public score than my meticulously not overfitted model. The not-overfitted model had less binary log-loss and better AUC and performed worse. So how can I actually tell if my model is getting better? </p>\n<p>Btw, my CV set is the last records of each student</p>",
      "rawMarkdown": "So, my model that's clearly overfit has a better public score than my meticulously not overfitted model. The not-overfitted model had less binary log-loss and better AUC and performed worse. So how can I actually tell if my model is getting better? \n\nBtw, my CV set is the last records of each student",
      "votes": null
    },
    {
      "id": "1060966",
      "postDate": "10/26/2020 16:41:48",
      "content": "<p>If the difference is small then it is just due to negligible randomness. Many competitors have shared <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">here</a> that a significant boost in CV translates to a boost on LB.</p>",
      "rawMarkdown": "If the difference is small then it is just due to negligible randomness. Many competitors have shared [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919) that a significant boost in CV translates to a boost on LB.",
      "votes": null
    },
    {
      "id": "1061013",
      "postDate": "10/26/2020 17:34:41",
      "content": "<p>It's big enough to jump some positions in the leaderboard (+0.02 auc_roc local). And I did use a fix random_state everywhere. </p>",
      "rawMarkdown": "It's big enough to jump some positions in the leaderboard (+0.02 auc_roc local). And I did use a fix random_state everywhere.",
      "votes": null
    },
    {
      "id": "1061015",
      "postDate": "10/26/2020 17:36:37",
      "content": "<p>I'm thinking it could be outliers or something. I am training on the entire dataset (except the CV set).</p>",
      "rawMarkdown": "I'm thinking it could be outliers or something. I am training on the entire dataset (except the CV set).",
      "votes": null
    },
    {
      "id": "1061094",
      "postDate": "10/26/2020 18:28:18",
      "content": "<p>It's hard to evaluate this without more information on what types of features you're using, but I have a few thoughts. First, why do you say that one model is clearly overfit vs. the other? If anything, it sounds like your model that appears to be performing better locally may be overfit to your local validation. Another distinct possibility is that your model that looks better locally should be performing better, but there is an issue with your data preparation. For example, maybe you impute nulls differently on test than on train or your features are otherwise misaligned. .02 AUC seems like a pretty large jump to me, which could point in the data preparation issue direction (I and other competitors are finding local vs. LB performance to be pretty stable within a few 3rd decimal places).</p>\n<p>It can be challenging to identify one problem vs. another, but my main suggestion is to try to make your local validation setup more reliable. The single last record is probably not enough (especially since some students may only have 1 record), and the very end of a user's history may have other issues as well that could introduce bias into your model. Once you get your validation right, you can be a lot more comfortable with drawing clear conclusions from your local work that will make it easier to identify submission problems that could be completely separate from the model quality itself. </p>",
      "rawMarkdown": "It's hard to evaluate this without more information on what types of features you're using, but I have a few thoughts. First, why do you say that one model is clearly overfit vs. the other? If anything, it sounds like your model that appears to be performing better locally may be overfit to your local validation. Another distinct possibility is that your model that looks better locally should be performing better, but there is an issue with your data preparation. For example, maybe you impute nulls differently on test than on train or your features are otherwise misaligned. .02 AUC seems like a pretty large jump to me, which could point in the data preparation issue direction (I and other competitors are finding local vs. LB performance to be pretty stable within a few 3rd decimal places).\n\nIt can be challenging to identify one problem vs. another, but my main suggestion is to try to make your local validation setup more reliable. The single last record is probably not enough (especially since some students may only have 1 record), and the very end of a user's history may have other issues as well that could introduce bias into your model. Once you get your validation right, you can be a lot more comfortable with drawing clear conclusions from your local work that will make it easier to identify submission problems that could be completely separate from the model quality itself.",
      "votes": null
    },
    {
      "id": "1061133",
      "postDate": "10/26/2020 19:07:31",
      "content": "<p>In one model, train error  was decreasing, while CV error started increasing after a while. This one performed better than the model that I used early stopping.</p>\n<p>I changed some parameters and it looks like it worked, both dont overfit now. I'll try to get a better validation set anyway. Thanks!</p>\n<p>\" and the very end of a user's history may have other issues as well that could introduce bias into your model.\"</p>\n<p>What issues for example?</p>",
      "rawMarkdown": "In one model, train error  was decreasing, while CV error started increasing after a while. This one performed better than the model that I used early stopping.\n\nI changed some parameters and it looks like it worked, both dont overfit now. I'll try to get a better validation set anyway. Thanks!\n\n\" and the very end of a user's history may have other issues as well that could introduce bias into your model.\"\n\nWhat issues for example?",
      "votes": null
    },
    {
      "id": "1061146",
      "postDate": "10/26/2020 19:14:55",
      "content": "<p>Depends on how the data was prepared which we don't know, unfortunately, but one thing that comes to mind is if there are users that are only in train / last train record is the end of their entire history. If they stopped using the app after that, maybe they had a bad experience or faced a particularly tough question, etc. Or maybe the app sent them an easier question to try to keep them engaged. </p>\n<p>All just speculation, but even the risk of that sort of effect would deter me from using too specific of a record in the histories for validation.</p>",
      "rawMarkdown": "Depends on how the data was prepared which we don't know, unfortunately, but one thing that comes to mind is if there are users that are only in train / last train record is the end of their entire history. If they stopped using the app after that, maybe they had a bad experience or faced a particularly tough question, etc. Or maybe the app sent them an easier question to try to keep them engaged. \n\nAll just speculation, but even the risk of that sort of effect would deter me from using too specific of a record in the histories for validation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1060966,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/26/2020 16:41:48",
      "content": "<p>If the difference is small then it is just due to negligible randomness. Many competitors have shared <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">here</a> that a significant boost in CV translates to a boost on LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1061013,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "10/26/2020 17:34:41",
          "content": "<p>It's big enough to jump some positions in the leaderboard (+0.02 auc_roc local). And I did use a fix random_state everywhere. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061015,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "10/26/2020 17:36:37",
          "content": "<p>I'm thinking it could be outliers or something. I am training on the entire dataset (except the CV set).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1061094,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/26/2020 18:28:18",
      "content": "<p>It's hard to evaluate this without more information on what types of features you're using, but I have a few thoughts. First, why do you say that one model is clearly overfit vs. the other? If anything, it sounds like your model that appears to be performing better locally may be overfit to your local validation. Another distinct possibility is that your model that looks better locally should be performing better, but there is an issue with your data preparation. For example, maybe you impute nulls differently on test than on train or your features are otherwise misaligned. .02 AUC seems like a pretty large jump to me, which could point in the data preparation issue direction (I and other competitors are finding local vs. LB performance to be pretty stable within a few 3rd decimal places).</p>\n<p>It can be challenging to identify one problem vs. another, but my main suggestion is to try to make your local validation setup more reliable. The single last record is probably not enough (especially since some students may only have 1 record), and the very end of a user's history may have other issues as well that could introduce bias into your model. Once you get your validation right, you can be a lot more comfortable with drawing clear conclusions from your local work that will make it easier to identify submission problems that could be completely separate from the model quality itself. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1061133,
          "author_name": "iuryck",
          "author_url": "",
          "post_date": "10/26/2020 19:07:31",
          "content": "<p>In one model, train error  was decreasing, while CV error started increasing after a while. This one performed better than the model that I used early stopping.</p>\n<p>I changed some parameters and it looks like it worked, both dont overfit now. I'll try to get a better validation set anyway. Thanks!</p>\n<p>\" and the very end of a user's history may have other issues as well that could introduce bias into your model.\"</p>\n<p>What issues for example?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061146,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/26/2020 19:14:55",
          "content": "<p>Depends on how the data was prepared which we don't know, unfortunately, but one thing that comes to mind is if there are users that are only in train / last train record is the end of their entire history. If they stopped using the app after that, maybe they had a bad experience or faced a particularly tough question, etc. Or maybe the app sent them an easier question to try to keep them engaged. </p>\n<p>All just speculation, but even the risk of that sort of effect would deter me from using too specific of a record in the histories for validation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1060953": "So, my model that's clearly overfit has a better public score than my meticulously not overfitted model. The not-overfitted model had less binary log-loss and better AUC and performed worse. So how can I actually tell if my model is getting better? \n\nBtw, my CV set is the last records of each student",
    "1060966": "If the difference is small then it is just due to negligible randomness. Many competitors have shared [here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919) that a significant boost in CV translates to a boost on LB.",
    "1061013": "It's big enough to jump some positions in the leaderboard (+0.02 auc_roc local). And I did use a fix random_state everywhere.",
    "1061015": "I'm thinking it could be outliers or something. I am training on the entire dataset (except the CV set).",
    "1061094": "It's hard to evaluate this without more information on what types of features you're using, but I have a few thoughts. First, why do you say that one model is clearly overfit vs. the other? If anything, it sounds like your model that appears to be performing better locally may be overfit to your local validation. Another distinct possibility is that your model that looks better locally should be performing better, but there is an issue with your data preparation. For example, maybe you impute nulls differently on test than on train or your features are otherwise misaligned. .02 AUC seems like a pretty large jump to me, which could point in the data preparation issue direction (I and other competitors are finding local vs. LB performance to be pretty stable within a few 3rd decimal places).\n\nIt can be challenging to identify one problem vs. another, but my main suggestion is to try to make your local validation setup more reliable. The single last record is probably not enough (especially since some students may only have 1 record), and the very end of a user's history may have other issues as well that could introduce bias into your model. Once you get your validation right, you can be a lot more comfortable with drawing clear conclusions from your local work that will make it easier to identify submission problems that could be completely separate from the model quality itself.",
    "1061133": "In one model, train error  was decreasing, while CV error started increasing after a while. This one performed better than the model that I used early stopping.\n\nI changed some parameters and it looks like it worked, both dont overfit now. I'll try to get a better validation set anyway. Thanks!\n\n\" and the very end of a user's history may have other issues as well that could introduce bias into your model.\"\n\nWhat issues for example?",
    "1061146": "Depends on how the data was prepared which we don't know, unfortunately, but one thing that comes to mind is if there are users that are only in train / last train record is the end of their entire history. If they stopped using the app after that, maybe they had a bad experience or faced a particularly tough question, etc. Or maybe the app sent them an easier question to try to keep them engaged. \n\nAll just speculation, but even the risk of that sort of effect would deter me from using too specific of a record in the histories for validation."
  },
  "source": "meta"
}