{
  "id": 52354,
  "title": "Why different AUC local train/test sets and leaderboard set",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52354",
  "author_name": "",
  "post_date": "2018-03-19T13:25:20.003437500Z",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I keep getting much lower results on leaderboard than on local train/test sets. \nI get the local train/test sets from the original train set provided and the sklearn train_test_split function.</p>\n\n<p>Here is how a typical training goes :  </p>\n\n<pre><code>[0] train-auc:0.930613  valid-auc:0.927525 \n\n[1] train-auc:0.930659  valid-auc:0.927598  \n\n[2] train-auc:0.930659  valid-auc:0.927598  \n\n[3] train-auc:0.930659  valid-auc:0.927598\n\n[4] train-auc:0.931786  valid-auc:0.929019\n\n[5] train-auc:0.93162   valid-auc:0.928857  \n\n[6] train-auc:0.933939  valid-auc:0.931206\n\n[7] train-auc:0.934027  valid-auc:0.931289\n\n[8] train-auc:0.941249  valid-auc:0.938918\n\n[9] train-auc:0.941249  valid-auc:0.938918\n</code></pre>\n\n<p>Score on leaderboard with this model : 0.87</p>\n\n<p>I get similar values for local training and test sets , which indicates there is no overfitting occurring, but the leaderboard score is completely different, which I do not get. </p>\n\n<p>It is like the distribution of (X,Y) in the leaderboard set is different from the one is the train set.  Has this happened to you and do you know what could be the reason of it ?</p>",
  "messages": [
    {
      "id": "298376",
      "postDate": "03/19/2018 13:25:20",
      "content": "<p>I keep getting much lower results on leaderboard than on local train/test sets. \nI get the local train/test sets from the original train set provided and the sklearn train_test_split function.</p>\n\n<p>Here is how a typical training goes :  </p>\n\n<pre><code>[0] train-auc:0.930613  valid-auc:0.927525 \n\n[1] train-auc:0.930659  valid-auc:0.927598  \n\n[2] train-auc:0.930659  valid-auc:0.927598  \n\n[3] train-auc:0.930659  valid-auc:0.927598\n\n[4] train-auc:0.931786  valid-auc:0.929019\n\n[5] train-auc:0.93162   valid-auc:0.928857  \n\n[6] train-auc:0.933939  valid-auc:0.931206\n\n[7] train-auc:0.934027  valid-auc:0.931289\n\n[8] train-auc:0.941249  valid-auc:0.938918\n\n[9] train-auc:0.941249  valid-auc:0.938918\n</code></pre>\n\n<p>Score on leaderboard with this model : 0.87</p>\n\n<p>I get similar values for local training and test sets , which indicates there is no overfitting occurring, but the leaderboard score is completely different, which I do not get. </p>\n\n<p>It is like the distribution of (X,Y) in the leaderboard set is different from the one is the train set.  Has this happened to you and do you know what could be the reason of it ?</p>",
      "rawMarkdown": "I keep getting much lower results on leaderboard than on local train/test sets. \nI get the local train/test sets from the original train set provided and the sklearn train_test_split function.\n\nHere is how a typical training goes :  \n\n    [0]\ttrain-auc:0.930613\tvalid-auc:0.927525 \n\n    [1]\ttrain-auc:0.930659\tvalid-auc:0.927598  \n\n    [2]\ttrain-auc:0.930659\tvalid-auc:0.927598  \n\n    [3]\ttrain-auc:0.930659\tvalid-auc:0.927598\n\n    [4]\ttrain-auc:0.931786\tvalid-auc:0.929019\n\n    [5]\ttrain-auc:0.93162\tvalid-auc:0.928857  \n\n    [6]\ttrain-auc:0.933939\tvalid-auc:0.931206\n\n    [7]\ttrain-auc:0.934027\tvalid-auc:0.931289\n\n    [8]\ttrain-auc:0.941249\tvalid-auc:0.938918\n\n    [9]\ttrain-auc:0.941249\tvalid-auc:0.938918\n\nScore on leaderboard with this model : 0.87\n\nI get similar values for local training and test sets , which indicates there is no overfitting occurring, but the leaderboard score is completely different, which I do not get. \n\nIt is like the distribution of (X,Y) in the leaderboard set is different from the one is the train set.  Has this happened to you and do you know what could be the reason of it ?",
      "votes": null
    },
    {
      "id": "298381",
      "postDate": "03/19/2018 13:42:19",
      "content": "<p>I think it's related to the size of your test set size, if you increase the size it will be closer to the leaderboard score.</p>",
      "rawMarkdown": "I think it's related to the size of your test set size, if you increase the size it will be closer to the leaderboard score.",
      "votes": null
    },
    {
      "id": "298385",
      "postDate": "03/19/2018 13:45:43",
      "content": "<p>Thank you for your quick reply, I will try to increase the current size of the test set (10% of 20 million rows) and let you know if it makes a difference</p>",
      "rawMarkdown": "Thank you for your quick reply, I will try to increase the current size of the test set (10% of 20 million rows) and let you know if it makes a difference",
      "votes": null
    },
    {
      "id": "298386",
      "postDate": "03/19/2018 13:48:01",
      "content": "<p>I'd also look at this other <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492\">validation post</a>, which has a discussion of the time series nature of the data. Basically, since this competition has time series data, you should consider holding out the last X% of the data instead of doing random sampling, since random sampling will leak information from the future and give overly optimistic estimates of performance.</p>",
      "rawMarkdown": "I'd also look at this other [validation post](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492), which has a discussion of the time series nature of the data. Basically, since this competition has time series data, you should consider holding out the last X% of the data instead of doing random sampling, since random sampling will leak information from the future and give overly optimistic estimates of performance.",
      "votes": null
    },
    {
      "id": "298390",
      "postDate": "03/19/2018 13:59:12",
      "content": "<p>unfortunately rising the % of rows in test set up to 50%  did not make any change</p>",
      "rawMarkdown": "unfortunately rising the % of rows in test set up to 50%  did not make any change",
      "votes": null
    },
    {
      "id": "298393",
      "postDate": "03/19/2018 14:00:37",
      "content": "<p>thank you for those insights. I'll have a deeper look at those posts</p>",
      "rawMarkdown": "thank you for those insights. I'll have a deeper look at those posts",
      "votes": null
    },
    {
      "id": "298395",
      "postDate": "03/19/2018 14:03:26",
      "content": "<p>have u checked that test ids are not exactly ascending ? they are: 0     1     2     3     4     5     6     7     9     8    10    11    12    13    14    15    16    20    19    17    18    21    22    23    24\nthis is why my first sub scored 0.8958. after fixing that it moved to 0.9619.</p>",
      "rawMarkdown": "have u checked that test ids are not exactly ascending ? they are: 0     1     2     3     4     5     6     7     9     8    10    11    12    13    14    15    16    20    19    17    18    21    22    23    24\nthis is why my first sub scored 0.8958. after fixing that it moved to 0.9619.",
      "votes": null
    },
    {
      "id": "298415",
      "postDate": "03/19/2018 14:45:41",
      "content": "<p>That is the answer thank you, I might have kept going like this for a while without your reply.</p>",
      "rawMarkdown": "That is the answer thank you, I might have kept going like this for a while without your reply.",
      "votes": null
    },
    {
      "id": "298429",
      "postDate": "03/19/2018 15:00:58",
      "content": "<p>Sorry ! It's strange that it worked for me ... \nI don't have the answer if I find out I let you know !</p>",
      "rawMarkdown": "Sorry ! It's strange that it worked for me ... \nI don't have the answer if I find out I let you know !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 298381,
      "author_name": "nathanlauga",
      "author_url": "",
      "post_date": "03/19/2018 13:42:19",
      "content": "<p>I think it's related to the size of your test set size, if you increase the size it will be closer to the leaderboard score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 298385,
          "author_name": "guillaumeligner",
          "author_url": "",
          "post_date": "03/19/2018 13:45:43",
          "content": "<p>Thank you for your quick reply, I will try to increase the current size of the test set (10% of 20 million rows) and let you know if it makes a difference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298390,
          "author_name": "guillaumeligner",
          "author_url": "",
          "post_date": "03/19/2018 13:59:12",
          "content": "<p>unfortunately rising the % of rows in test set up to 50%  did not make any change</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 298429,
          "author_name": "nathanlauga",
          "author_url": "",
          "post_date": "03/19/2018 15:00:58",
          "content": "<p>Sorry ! It's strange that it worked for me ... \nI don't have the answer if I find out I let you know !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 298386,
      "author_name": "chotch",
      "author_url": "",
      "post_date": "03/19/2018 13:48:01",
      "content": "<p>I'd also look at this other <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492\">validation post</a>, which has a discussion of the time series nature of the data. Basically, since this competition has time series data, you should consider holding out the last X% of the data instead of doing random sampling, since random sampling will leak information from the future and give overly optimistic estimates of performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 298393,
          "author_name": "guillaumeligner",
          "author_url": "",
          "post_date": "03/19/2018 14:00:37",
          "content": "<p>thank you for those insights. I'll have a deeper look at those posts</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 298395,
      "author_name": "mjahrer",
      "author_url": "",
      "post_date": "03/19/2018 14:03:26",
      "content": "<p>have u checked that test ids are not exactly ascending ? they are: 0     1     2     3     4     5     6     7     9     8    10    11    12    13    14    15    16    20    19    17    18    21    22    23    24\nthis is why my first sub scored 0.8958. after fixing that it moved to 0.9619.</p>",
      "votes": null,
      "replies": [
        {
          "id": 298415,
          "author_name": "guillaumeligner",
          "author_url": "",
          "post_date": "03/19/2018 14:45:41",
          "content": "<p>That is the answer thank you, I might have kept going like this for a while without your reply.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "298376": "I keep getting much lower results on leaderboard than on local train/test sets. \nI get the local train/test sets from the original train set provided and the sklearn train_test_split function.\n\nHere is how a typical training goes :  \n\n    [0]\ttrain-auc:0.930613\tvalid-auc:0.927525 \n\n    [1]\ttrain-auc:0.930659\tvalid-auc:0.927598  \n\n    [2]\ttrain-auc:0.930659\tvalid-auc:0.927598  \n\n    [3]\ttrain-auc:0.930659\tvalid-auc:0.927598\n\n    [4]\ttrain-auc:0.931786\tvalid-auc:0.929019\n\n    [5]\ttrain-auc:0.93162\tvalid-auc:0.928857  \n\n    [6]\ttrain-auc:0.933939\tvalid-auc:0.931206\n\n    [7]\ttrain-auc:0.934027\tvalid-auc:0.931289\n\n    [8]\ttrain-auc:0.941249\tvalid-auc:0.938918\n\n    [9]\ttrain-auc:0.941249\tvalid-auc:0.938918\n\nScore on leaderboard with this model : 0.87\n\nI get similar values for local training and test sets , which indicates there is no overfitting occurring, but the leaderboard score is completely different, which I do not get. \n\nIt is like the distribution of (X,Y) in the leaderboard set is different from the one is the train set.  Has this happened to you and do you know what could be the reason of it ?",
    "298381": "I think it's related to the size of your test set size, if you increase the size it will be closer to the leaderboard score.",
    "298385": "Thank you for your quick reply, I will try to increase the current size of the test set (10% of 20 million rows) and let you know if it makes a difference",
    "298386": "I'd also look at this other [validation post](https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492), which has a discussion of the time series nature of the data. Basically, since this competition has time series data, you should consider holding out the last X% of the data instead of doing random sampling, since random sampling will leak information from the future and give overly optimistic estimates of performance.",
    "298390": "unfortunately rising the % of rows in test set up to 50%  did not make any change",
    "298393": "thank you for those insights. I'll have a deeper look at those posts",
    "298395": "have u checked that test ids are not exactly ascending ? they are: 0     1     2     3     4     5     6     7     9     8    10    11    12    13    14    15    16    20    19    17    18    21    22    23    24\nthis is why my first sub scored 0.8958. after fixing that it moved to 0.9619.",
    "298415": "That is the answer thank you, I might have kept going like this for a while without your reply.",
    "298429": "Sorry ! It's strange that it worked for me ... \nI don't have the answer if I find out I let you know !"
  },
  "source": "meta"
}