{
  "id": 336133,
  "title": "Customers with 13 rows?",
  "url": "/competitions/amex-default-prediction/discussion/336133",
  "author_name": "",
  "post_date": "2022-07-09T13:56:35.621334200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, <br>\nI want to know if someone would be willing to help me, with a few questions?</p>\n<ol>\n<li>Does each customer have exactly 13 rows in the train and test data?</li>\n</ol>\n<p>I preprocessed the train data and found customers with less than 13 rows (12 to 1 row per customer).<br>\nHere is a summary of what I found.<br>\nCUSTOMER ROWS 1 DEFAULT 0.34 %<br>\nCUSTOMER ROWS 2 DEFAULT 0.32 %<br>\nCUSTOMER ROWS 3 DEFAULT 0.36 %<br>\nCUSTOMER ROWS 4 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 5 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 6 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 7 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 8 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 9 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 10 DEFAULT 0.46 %<br>\nCUSTOMER ROWS 11 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 12 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 13 DEFAULT 0.23 %</p>\n<ol>\n<li>How would you account for this in feature engineering, or would you need to account for this?</li>\n</ol>\n<p>For example the XGB Starter Notebook used features like max, min, and std, columns with 1 row<br>\nMax == Min and STD == 0, wouldn’t algorithms like XGBoost be confused?</p>\n<p>Note: I don't know if I messed up with preprocessing</p>",
  "messages": [
    {
      "id": "1849424",
      "postDate": "07/09/2022 13:56:35",
      "content": "<p>Hi, <br>\nI want to know if someone would be willing to help me, with a few questions?</p>\n<ol>\n<li>Does each customer have exactly 13 rows in the train and test data?</li>\n</ol>\n<p>I preprocessed the train data and found customers with less than 13 rows (12 to 1 row per customer).<br>\nHere is a summary of what I found.<br>\nCUSTOMER ROWS 1 DEFAULT 0.34 %<br>\nCUSTOMER ROWS 2 DEFAULT 0.32 %<br>\nCUSTOMER ROWS 3 DEFAULT 0.36 %<br>\nCUSTOMER ROWS 4 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 5 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 6 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 7 DEFAULT 0.42 %<br>\nCUSTOMER ROWS 8 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 9 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 10 DEFAULT 0.46 %<br>\nCUSTOMER ROWS 11 DEFAULT 0.45 %<br>\nCUSTOMER ROWS 12 DEFAULT 0.39 %<br>\nCUSTOMER ROWS 13 DEFAULT 0.23 %</p>\n<ol>\n<li>How would you account for this in feature engineering, or would you need to account for this?</li>\n</ol>\n<p>For example the XGB Starter Notebook used features like max, min, and std, columns with 1 row<br>\nMax == Min and STD == 0, wouldn’t algorithms like XGBoost be confused?</p>\n<p>Note: I don't know if I messed up with preprocessing</p>",
      "rawMarkdown": "Hi, \nI want to know if someone would be willing to help me, with a few questions?\n\n1. Does each customer have exactly 13 rows in the train and test data?\n\nI preprocessed the train data and found customers with less than 13 rows (12 to 1 row per customer).\nHere is a summary of what I found.\nCUSTOMER ROWS 1 DEFAULT 0.34 %\nCUSTOMER ROWS 2 DEFAULT 0.32 %\nCUSTOMER ROWS 3 DEFAULT 0.36 %\nCUSTOMER ROWS 4 DEFAULT 0.42 %\nCUSTOMER ROWS 5 DEFAULT 0.39 %\nCUSTOMER ROWS 6 DEFAULT 0.39 %\nCUSTOMER ROWS 7 DEFAULT 0.42 %\nCUSTOMER ROWS 8 DEFAULT 0.45 %\nCUSTOMER ROWS 9 DEFAULT 0.45 %\nCUSTOMER ROWS 10 DEFAULT 0.46 %\nCUSTOMER ROWS 11 DEFAULT 0.45 %\nCUSTOMER ROWS 12 DEFAULT 0.39 %\nCUSTOMER ROWS 13 DEFAULT 0.23 %\n\n2.\tHow would you account for this in feature engineering, or would you need to account for this?\n\nFor example the XGB Starter Notebook used features like max, min, and std, columns with 1 row\nMax == Min and STD == 0, wouldn’t algorithms like XGBoost be confused?\n\nNote: I don't know if I messed up with preprocessing",
      "votes": null
    },
    {
      "id": "1849689",
      "postDate": "07/09/2022 17:57:29",
      "content": "<p>There's a number of discussions you can find on this topic.  Yes - there are customers with less than 13.  Some are missing the early statements, some are missing the ending statements and some missing in the middle.  </p>",
      "rawMarkdown": "There's a number of discussions you can find on this topic.  Yes - there are customers with less than 13.  Some are missing the early statements, some are missing the ending statements and some missing in the middle.",
      "votes": null
    },
    {
      "id": "1849958",
      "postDate": "07/10/2022 00:47:27",
      "content": "<p>There are customers with less than 13, you are correct that it is not ideal calculating some of the features on such a small set of samples (such as std). <br>\nThere are some ways to improve this approach.</p>",
      "rawMarkdown": "There are customers with less than 13, you are correct that it is not ideal calculating some of the features on such a small set of samples (such as std). \nThere are some ways to improve this approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1849689,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/09/2022 17:57:29",
      "content": "<p>There's a number of discussions you can find on this topic.  Yes - there are customers with less than 13.  Some are missing the early statements, some are missing the ending statements and some missing in the middle.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1849958,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/10/2022 00:47:27",
      "content": "<p>There are customers with less than 13, you are correct that it is not ideal calculating some of the features on such a small set of samples (such as std). <br>\nThere are some ways to improve this approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1849424": "Hi, \nI want to know if someone would be willing to help me, with a few questions?\n\n1. Does each customer have exactly 13 rows in the train and test data?\n\nI preprocessed the train data and found customers with less than 13 rows (12 to 1 row per customer).\nHere is a summary of what I found.\nCUSTOMER ROWS 1 DEFAULT 0.34 %\nCUSTOMER ROWS 2 DEFAULT 0.32 %\nCUSTOMER ROWS 3 DEFAULT 0.36 %\nCUSTOMER ROWS 4 DEFAULT 0.42 %\nCUSTOMER ROWS 5 DEFAULT 0.39 %\nCUSTOMER ROWS 6 DEFAULT 0.39 %\nCUSTOMER ROWS 7 DEFAULT 0.42 %\nCUSTOMER ROWS 8 DEFAULT 0.45 %\nCUSTOMER ROWS 9 DEFAULT 0.45 %\nCUSTOMER ROWS 10 DEFAULT 0.46 %\nCUSTOMER ROWS 11 DEFAULT 0.45 %\nCUSTOMER ROWS 12 DEFAULT 0.39 %\nCUSTOMER ROWS 13 DEFAULT 0.23 %\n\n2.\tHow would you account for this in feature engineering, or would you need to account for this?\n\nFor example the XGB Starter Notebook used features like max, min, and std, columns with 1 row\nMax == Min and STD == 0, wouldn’t algorithms like XGBoost be confused?\n\nNote: I don't know if I messed up with preprocessing",
    "1849689": "There's a number of discussions you can find on this topic.  Yes - there are customers with less than 13.  Some are missing the early statements, some are missing the ending statements and some missing in the middle.",
    "1849958": "There are customers with less than 13, you are correct that it is not ideal calculating some of the features on such a small set of samples (such as std). \nThere are some ways to improve this approach."
  },
  "source": "meta"
}