{
  "id": 334113,
  "title": "The label distribution of train set will strongly impact model metric",
  "url": "/competitions/amex-default-prediction/discussion/334113",
  "author_name": "",
  "post_date": "2022-06-29T20:12:25.296012Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I notice if there is a perfect model that makes correct predictions, it still may get a low metric score due to train set label distribution</p>\n<p>Here is a quick example:</p>\n<pre><code># 01 assuming a train set consists of 90% of '1' and 10% of '0'\n\nx = [1 if i &lt; 9000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.5644444444444444\n\n# 02 assuming a train set consists of 10% of '1' and 90% of '0'\nx = [1 if i &lt; 1000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n\n\n# 03 assuming a train set consists of 50% of '1' and 50% of '0'\nx = [1 if i &lt; 5000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.9199999999999999\n\n\n# 04 assuming a train set consists of 25% of '1' and 75% of '0' (the original train set distribution)\nx = [1 if i &lt; 2500 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n</code></pre>\n<p>So, I suggest that if anyone generates datasets with random seeds but without a fixed label distribution, it is better to recheck your results. Model results based on different datasets may mislead you to make wrong decisions. For example, a cause of the difference between your CV result and LB score is that your train set includes too many '0' labels</p>",
  "messages": [
    {
      "id": "1837706",
      "postDate": "06/29/2022 20:12:25",
      "content": "<p>I notice if there is a perfect model that makes correct predictions, it still may get a low metric score due to train set label distribution</p>\n<p>Here is a quick example:</p>\n<pre><code># 01 assuming a train set consists of 90% of '1' and 10% of '0'\n\nx = [1 if i &lt; 9000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.5644444444444444\n\n# 02 assuming a train set consists of 10% of '1' and 90% of '0'\nx = [1 if i &lt; 1000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n\n\n# 03 assuming a train set consists of 50% of '1' and 50% of '0'\nx = [1 if i &lt; 5000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.9199999999999999\n\n\n# 04 assuming a train set consists of 25% of '1' and 75% of '0' (the original train set distribution)\nx = [1 if i &lt; 2500 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n</code></pre>\n<p>So, I suggest that if anyone generates datasets with random seeds but without a fixed label distribution, it is better to recheck your results. Model results based on different datasets may mislead you to make wrong decisions. For example, a cause of the difference between your CV result and LB score is that your train set includes too many '0' labels</p>",
      "rawMarkdown": "I notice if there is a perfect model that makes correct predictions, it still may get a low metric score due to train set label distribution\n\nHere is a quick example:\n\n```\n# 01 assuming a train set consists of 90% of '1' and 10% of '0'\n\nx = [1 if i < 9000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.5644444444444444\n\n# 02 assuming a train set consists of 10% of '1' and 90% of '0'\nx = [1 if i < 1000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n\n\n# 03 assuming a train set consists of 50% of '1' and 50% of '0'\nx = [1 if i < 5000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.9199999999999999\n\n\n# 04 assuming a train set consists of 25% of '1' and 75% of '0' (the original train set distribution)\nx = [1 if i < 2500 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n```\nSo, I suggest that if anyone generates datasets with random seeds but without a fixed label distribution, it is better to recheck your results. Model results based on different datasets may mislead you to make wrong decisions. For example, a cause of the difference between your CV result and LB score is that your train set includes too many '0' labels",
      "votes": null
    },
    {
      "id": "1837774",
      "postDate": "06/29/2022 22:14:16",
      "content": "<p>I think the idea is the <code>1</code>s are very rare and can be captured fully in the top 4% of the highest scoring population if the model is perfect. </p>",
      "rawMarkdown": "I think the idea is the `1`s are very rare and can be captured fully in the top 4% of the highest scoring population if the model is perfect.",
      "votes": null
    },
    {
      "id": "1838037",
      "postDate": "06/30/2022 06:43:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carloszonetgmailcom\" target=\"_blank\">@carloszonetgmailcom</a> That's exactly the purpose of <code>StratifiedKFold</code>: It ensures that all your splits have the same ratio of '0' and '1'.</p>",
      "rawMarkdown": "Hi @carloszonetgmailcom That's exactly the purpose of `StratifiedKFold`: It ensures that all your splits have the same ratio of '0' and '1'.",
      "votes": null
    },
    {
      "id": "1838703",
      "postDate": "06/30/2022 18:32:28",
      "content": "<p>The distribution of labels in the train set can definitely impact the model metric.<br>\nTo combat this, use <code>StratifiedKFold</code>. it make sure the distribution of the labels is perserved through the different folds.</p>",
      "rawMarkdown": "The distribution of labels in the train set can definitely impact the model metric.\nTo combat this, use `StratifiedKFold`. it make sure the distribution of the labels is perserved through the different folds.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1837774,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "06/29/2022 22:14:16",
      "content": "<p>I think the idea is the <code>1</code>s are very rare and can be captured fully in the top 4% of the highest scoring population if the model is perfect. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838037,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "06/30/2022 06:43:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carloszonetgmailcom\" target=\"_blank\">@carloszonetgmailcom</a> That's exactly the purpose of <code>StratifiedKFold</code>: It ensures that all your splits have the same ratio of '0' and '1'.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838703,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "06/30/2022 18:32:28",
      "content": "<p>The distribution of labels in the train set can definitely impact the model metric.<br>\nTo combat this, use <code>StratifiedKFold</code>. it make sure the distribution of the labels is perserved through the different folds.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1837706": "I notice if there is a perfect model that makes correct predictions, it still may get a low metric score due to train set label distribution\n\nHere is a quick example:\n\n```\n# 01 assuming a train set consists of 90% of '1' and 10% of '0'\n\nx = [1 if i < 9000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.5644444444444444\n\n# 02 assuming a train set consists of 10% of '1' and 90% of '0'\nx = [1 if i < 1000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n\n\n# 03 assuming a train set consists of 50% of '1' and 50% of '0'\nx = [1 if i < 5000 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n0.9199999999999999\n\n\n# 04 assuming a train set consists of 25% of '1' and 75% of '0' (the original train set distribution)\nx = [1 if i < 2500 else 0 for i in range(10000)] \n\ntrue_values = pd.DataFrame(x, columns = ['target'])\npre_values = pd.DataFrame(x, columns = ['prediction'])\namex_metric(true_values, pre_values)\n\n1.0\n```\nSo, I suggest that if anyone generates datasets with random seeds but without a fixed label distribution, it is better to recheck your results. Model results based on different datasets may mislead you to make wrong decisions. For example, a cause of the difference between your CV result and LB score is that your train set includes too many '0' labels",
    "1837774": "I think the idea is the `1`s are very rare and can be captured fully in the top 4% of the highest scoring population if the model is perfect.",
    "1838037": "Hi @carloszonetgmailcom That's exactly the purpose of `StratifiedKFold`: It ensures that all your splits have the same ratio of '0' and '1'.",
    "1838703": "The distribution of labels in the train set can definitely impact the model metric.\nTo combat this, use `StratifiedKFold`. it make sure the distribution of the labels is perserved through the different folds."
  },
  "source": "meta"
}