{
  "id": 54723,
  "title": "Accidental interesting finding about validation",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54723",
  "author_name": "",
  "post_date": "2018-04-17T05:47:15.127486500Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Because of lack of consideration, I run a training with the attribution rate about 0.25% in train set (similar to the rate of the universe) and a validation set of attribution rate about 0.12%.</p>\n\n<p>In the whole train time, firstly quite odd, the auc for validation set is higher than that for training set. Of course, the auc on LB is much less.</p>\n\n<p>My ballpark theory is because of less positives in validation set, the result for true positive looks higher than it should be and on the other hand more negatives in validation set artificially makes false positive lower. The overall effect is pulling the roc curve to the top left  corner, therefore resulting higher auc.</p>\n\n<p>The take-away can be always create validation set with stratifying the classes.</p>",
  "messages": [
    {
      "id": "315441",
      "postDate": "04/17/2018 05:47:15",
      "content": "<p>Because of lack of consideration, I run a training with the attribution rate about 0.25% in train set (similar to the rate of the universe) and a validation set of attribution rate about 0.12%.</p>\n\n<p>In the whole train time, firstly quite odd, the auc for validation set is higher than that for training set. Of course, the auc on LB is much less.</p>\n\n<p>My ballpark theory is because of less positives in validation set, the result for true positive looks higher than it should be and on the other hand more negatives in validation set artificially makes false positive lower. The overall effect is pulling the roc curve to the top left  corner, therefore resulting higher auc.</p>\n\n<p>The take-away can be always create validation set with stratifying the classes.</p>",
      "rawMarkdown": "Because of lack of consideration, I run a training with the attribution rate about 0.25% in train set (similar to the rate of the universe) and a validation set of attribution rate about 0.12%.\n\nIn the whole train time, firstly quite odd, the auc for validation set is higher than that for training set. Of course, the auc on LB is much less.\n\nMy ballpark theory is because of less positives in validation set, the result for true positive looks higher than it should be and on the other hand more negatives in validation set artificially makes false positive lower. The overall effect is pulling the roc curve to the top left  corner, therefore resulting higher auc.\n\nThe take-away can be always create validation set with stratifying the classes.",
      "votes": null
    },
    {
      "id": "315454",
      "postDate": "04/17/2018 06:17:08",
      "content": "<p>I had the same sometimes with a validation set with 10m rows from the training set for training and the next 2m for validation. I did not check the attribution rate though.</p>",
      "rawMarkdown": "I had the same sometimes with a validation set with 10m rows from the training set for training and the next 2m for validation. I did not check the attribution rate though.",
      "votes": null
    },
    {
      "id": "315457",
      "postDate": "04/17/2018 06:21:20",
      "content": "<p>Don't have time to verify what I said. But it's highly likely true. In the next training, I stratified the classes so the validation set has same attribution rate as the training set. Validation auc then comes under training auc during training.</p>",
      "rawMarkdown": "Don't have time to verify what I said. But it's highly likely true. In the next training, I stratified the classes so the validation set has same attribution rate as the training set. Validation auc then comes under training auc during training.",
      "votes": null
    },
    {
      "id": "315461",
      "postDate": "04/17/2018 06:27:46",
      "content": "<p>How did you stratify ?\nWhenever I get more free tm I'll tryto check more thoroughly your theory ;)</p>",
      "rawMarkdown": "How did you stratify ?\nWhenever I get more free tm I'll tryto check more thoroughly your theory ;)",
      "votes": null
    },
    {
      "id": "315472",
      "postDate": "04/17/2018 06:38:54",
      "content": "<p>I set the stratify parameter in sklearn train_test_split:</p>\n\n<p><code>\ntrain, valid = train_test_split(dat, test_size=0.2, stratify=dat.is_attributed)\n</code></p>",
      "rawMarkdown": "I set the stratify parameter in sklearn train_test_split:\n\n```\ntrain, valid = train_test_split(dat, test_size=0.2, stratify=dat.is_attributed)\n```",
      "votes": null
    },
    {
      "id": "316331",
      "postDate": "04/18/2018 18:39:36",
      "content": "<p>As @Fei specified you need to stratify the test/validation set.</p>\n\n<p>Alternatively, if you only use pandas, you can shuffle the data before picking a test/validation set - it would give a similar percentage of is_attributed as of the entire set (0.247% of data has i_attributed = 1)\nI'm doing something like this</p>\n\n<p>test_size = len(train_df) // 5    #(I have this fixed to 2000000)</p>\n\n<p>train_df = train_df.sample(frac=1.0)</p>\n\n<p>train_data = train_df[:-test_size]</p>\n\n<p>test_data = train_df[len(train_df)-test_size:]</p>\n\n<p>Obviously if you do not fully retrain the model you'll have to persist the generated train/test_data and reload if needed.</p>",
      "rawMarkdown": "As @Fei specified you need to stratify the test/validation set.\n\nAlternatively, if you only use pandas, you can shuffle the data before picking a test/validation set - it would give a similar percentage of is_attributed as of the entire set (0.247% of data has i_attributed = 1)\nI'm doing something like this\n\ntest_size = len(train_df) // 5    #(I have this fixed to 2000000)\n\ntrain_df = train_df.sample(frac=1.0)\n\ntrain_data = train_df[:-test_size]\n\ntest_data = train_df[len(train_df)-test_size:]\n\nObviously if you do not fully retrain the model you'll have to persist the generated train/test_data and reload if needed.",
      "votes": null
    },
    {
      "id": "316919",
      "postDate": "04/20/2018 07:29:00",
      "content": "<p>A double check with value_counts would be handy:</p>\n\n<p><code>\nvalid.is_attributed.value_counts(normalize=True)\n</code></p>\n\n<p>Which gives the percentage of each category based on counts.</p>",
      "rawMarkdown": "A double check with value_counts would be handy:\n\n\n```\nvalid.is_attributed.value_counts(normalize=True)\n```\n\nWhich gives the percentage of each category based on counts.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 315454,
      "author_name": "thepathofd",
      "author_url": "",
      "post_date": "04/17/2018 06:17:08",
      "content": "<p>I had the same sometimes with a validation set with 10m rows from the training set for training and the next 2m for validation. I did not check the attribution rate though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 315457,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/17/2018 06:21:20",
          "content": "<p>Don't have time to verify what I said. But it's highly likely true. In the next training, I stratified the classes so the validation set has same attribution rate as the training set. Validation auc then comes under training auc during training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 315461,
          "author_name": "propanon",
          "author_url": "",
          "post_date": "04/17/2018 06:27:46",
          "content": "<p>How did you stratify ?\nWhenever I get more free tm I'll tryto check more thoroughly your theory ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 315472,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/17/2018 06:38:54",
          "content": "<p>I set the stratify parameter in sklearn train_test_split:</p>\n\n<p><code>\ntrain, valid = train_test_split(dat, test_size=0.2, stratify=dat.is_attributed)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316331,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "04/18/2018 18:39:36",
          "content": "<p>As @Fei specified you need to stratify the test/validation set.</p>\n\n<p>Alternatively, if you only use pandas, you can shuffle the data before picking a test/validation set - it would give a similar percentage of is_attributed as of the entire set (0.247% of data has i_attributed = 1)\nI'm doing something like this</p>\n\n<p>test_size = len(train_df) // 5    #(I have this fixed to 2000000)</p>\n\n<p>train_df = train_df.sample(frac=1.0)</p>\n\n<p>train_data = train_df[:-test_size]</p>\n\n<p>test_data = train_df[len(train_df)-test_size:]</p>\n\n<p>Obviously if you do not fully retrain the model you'll have to persist the generated train/test_data and reload if needed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316919,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/20/2018 07:29:00",
          "content": "<p>A double check with value_counts would be handy:</p>\n\n<p><code>\nvalid.is_attributed.value_counts(normalize=True)\n</code></p>\n\n<p>Which gives the percentage of each category based on counts.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "315441": "Because of lack of consideration, I run a training with the attribution rate about 0.25% in train set (similar to the rate of the universe) and a validation set of attribution rate about 0.12%.\n\nIn the whole train time, firstly quite odd, the auc for validation set is higher than that for training set. Of course, the auc on LB is much less.\n\nMy ballpark theory is because of less positives in validation set, the result for true positive looks higher than it should be and on the other hand more negatives in validation set artificially makes false positive lower. The overall effect is pulling the roc curve to the top left  corner, therefore resulting higher auc.\n\nThe take-away can be always create validation set with stratifying the classes.",
    "315454": "I had the same sometimes with a validation set with 10m rows from the training set for training and the next 2m for validation. I did not check the attribution rate though.",
    "315457": "Don't have time to verify what I said. But it's highly likely true. In the next training, I stratified the classes so the validation set has same attribution rate as the training set. Validation auc then comes under training auc during training.",
    "315461": "How did you stratify ?\nWhenever I get more free tm I'll tryto check more thoroughly your theory ;)",
    "315472": "I set the stratify parameter in sklearn train_test_split:\n\n```\ntrain, valid = train_test_split(dat, test_size=0.2, stratify=dat.is_attributed)\n```",
    "316331": "As @Fei specified you need to stratify the test/validation set.\n\nAlternatively, if you only use pandas, you can shuffle the data before picking a test/validation set - it would give a similar percentage of is_attributed as of the entire set (0.247% of data has i_attributed = 1)\nI'm doing something like this\n\ntest_size = len(train_df) // 5    #(I have this fixed to 2000000)\n\ntrain_df = train_df.sample(frac=1.0)\n\ntrain_data = train_df[:-test_size]\n\ntest_data = train_df[len(train_df)-test_size:]\n\nObviously if you do not fully retrain the model you'll have to persist the generated train/test_data and reload if needed.",
    "316919": "A double check with value_counts would be handy:\n\n\n```\nvalid.is_attributed.value_counts(normalize=True)\n```\n\nWhich gives the percentage of each category based on counts."
  },
  "source": "meta"
}