{
  "id": 341886,
  "title": "Length of training labels vs length of training data",
  "url": "/competitions/amex-default-prediction/discussion/341886",
  "author_name": "",
  "post_date": "2022-08-04T15:44:19.841294500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm relatively new to the big data sphere but the length of the training labels is 458913 and the length of the training dataset is 5531451. I am trying to fit my Catboost model to the entire training dataset if possible. Is there something I am completely missing? Is it that multiple customer IDs can exist in the training data set but in the labeled dataset there exists only 1 target label?</p>\n<p>Thank you for any helpful insights,<br>\nKarl</p>",
  "messages": [
    {
      "id": "1884728",
      "postDate": "08/04/2022 15:44:19",
      "content": "<p>Hi everyone,</p>\n<p>I'm relatively new to the big data sphere but the length of the training labels is 458913 and the length of the training dataset is 5531451. I am trying to fit my Catboost model to the entire training dataset if possible. Is there something I am completely missing? Is it that multiple customer IDs can exist in the training data set but in the labeled dataset there exists only 1 target label?</p>\n<p>Thank you for any helpful insights,<br>\nKarl</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm relatively new to the big data sphere but the length of the training labels is 458913 and the length of the training dataset is 5531451. I am trying to fit my Catboost model to the entire training dataset if possible. Is there something I am completely missing? Is it that multiple customer IDs can exist in the training data set but in the labeled dataset there exists only 1 target label?\n\nThank you for any helpful insights,\nKarl",
      "votes": null
    },
    {
      "id": "1884862",
      "postDate": "08/04/2022 17:45:21",
      "content": "<p>There are 458913 customers for which we predict labels. Customers have up to 13 credit card statements each, which are in the train file. The choice is to do one of these: 1) group the data from the start by <code>customer_ID</code> and aggregate their values across all statements, which will give 458913 samples; 2) train and predict for 5+ million original data points, group by <code>customer_ID</code> and average their predictions, which will also give you 458913 in the end.</p>\n<p>Most people are doing option #1 because it is faster during the training stage and is less likely to overfit. I suggest you look through the posted kernels in the code section and it should be pretty clear what pre-processing needs to be done to go with option #1.</p>",
      "rawMarkdown": "There are 458913 customers for which we predict labels. Customers have up to 13 credit card statements each, which are in the train file. The choice is to do one of these: 1) group the data from the start by `customer_ID` and aggregate their values across all statements, which will give 458913 samples; 2) train and predict for 5+ million original data points, group by `customer_ID` and average their predictions, which will also give you 458913 in the end.\n\nMost people are doing option #1 because it is faster during the training stage and is less likely to overfit. I suggest you look through the posted kernels in the code section and it should be pretty clear what pre-processing needs to be done to go with option #1.",
      "votes": null
    },
    {
      "id": "1885003",
      "postDate": "08/04/2022 19:35:15",
      "content": "<p>Another thing if doing option #2, your CV needs to split customers completely. Basic KFold on the 5.5 million rows will essentially leak information into the validation set, it won't be independent. </p>\n<p>If you do your KFold split on the 458k target data, it will avoid this issue. Then you just need to merge both target and fold number into the 5.5 million rows, merging on customer ID. </p>",
      "rawMarkdown": "Another thing if doing option #2, your CV needs to split customers completely. Basic KFold on the 5.5 million rows will essentially leak information into the validation set, it won't be independent. \n\nIf you do your KFold split on the 458k target data, it will avoid this issue. Then you just need to merge both target and fold number into the 5.5 million rows, merging on customer ID.",
      "votes": null
    },
    {
      "id": "1890502",
      "postDate": "08/08/2022 19:27:25",
      "content": "<p>Thank you for the Insight!!</p>",
      "rawMarkdown": "Thank you for the Insight!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1884862,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/04/2022 17:45:21",
      "content": "<p>There are 458913 customers for which we predict labels. Customers have up to 13 credit card statements each, which are in the train file. The choice is to do one of these: 1) group the data from the start by <code>customer_ID</code> and aggregate their values across all statements, which will give 458913 samples; 2) train and predict for 5+ million original data points, group by <code>customer_ID</code> and average their predictions, which will also give you 458913 in the end.</p>\n<p>Most people are doing option #1 because it is faster during the training stage and is less likely to overfit. I suggest you look through the posted kernels in the code section and it should be pretty clear what pre-processing needs to be done to go with option #1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1885003,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/04/2022 19:35:15",
          "content": "<p>Another thing if doing option #2, your CV needs to split customers completely. Basic KFold on the 5.5 million rows will essentially leak information into the validation set, it won't be independent. </p>\n<p>If you do your KFold split on the 458k target data, it will avoid this issue. Then you just need to merge both target and fold number into the 5.5 million rows, merging on customer ID. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1890502,
          "author_name": "karlpetz",
          "author_url": "",
          "post_date": "08/08/2022 19:27:25",
          "content": "<p>Thank you for the Insight!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1884728": "Hi everyone,\n\nI'm relatively new to the big data sphere but the length of the training labels is 458913 and the length of the training dataset is 5531451. I am trying to fit my Catboost model to the entire training dataset if possible. Is there something I am completely missing? Is it that multiple customer IDs can exist in the training data set but in the labeled dataset there exists only 1 target label?\n\nThank you for any helpful insights,\nKarl",
    "1884862": "There are 458913 customers for which we predict labels. Customers have up to 13 credit card statements each, which are in the train file. The choice is to do one of these: 1) group the data from the start by `customer_ID` and aggregate their values across all statements, which will give 458913 samples; 2) train and predict for 5+ million original data points, group by `customer_ID` and average their predictions, which will also give you 458913 in the end.\n\nMost people are doing option #1 because it is faster during the training stage and is less likely to overfit. I suggest you look through the posted kernels in the code section and it should be pretty clear what pre-processing needs to be done to go with option #1.",
    "1885003": "Another thing if doing option #2, your CV needs to split customers completely. Basic KFold on the 5.5 million rows will essentially leak information into the validation set, it won't be independent. \n\nIf you do your KFold split on the 458k target data, it will avoid this issue. Then you just need to merge both target and fold number into the 5.5 million rows, merging on customer ID.",
    "1890502": "Thank you for the Insight!!"
  },
  "source": "meta"
}