{
  "id": 327629,
  "title": "Cross-validation",
  "url": "/competitions/amex-default-prediction/discussion/327629",
  "author_name": "Gunes Evitan",
  "post_date": "2022-05-28T08:53:50.884000",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I saw some of the notebooks are using stratifiedkfold but I would prefer to use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html\" target=\"_blank\">StratifiedGroupKFold</a> here since test set has unseen customer IDs.</p>\n<pre><code>2022-05-28 11:41:07 INFO validation - create_folds: Dataset split into 5 folds (Seed: 42)\n2022-05-28 11:41:08 INFO validation - create_folds: Fold 1 (1106140, 196) - Target Mean: 0.2485 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:09 INFO validation - create_folds: Fold 2 (1106518, 196) - Target Mean: 0.2487 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:10 INFO validation - create_folds: Fold 3 (1105984, 196) - Target Mean: 0.2483 Std: 0.4320 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 4 (1106699, 196) - Target Mean: 0.2495 Std: 0.4328 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 5 (1106110, 196) - Target Mean: 0.2504 Std: 0.4333 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: 0 overlapping customer_IDs\n</code></pre>\n<p>Those are my folds after splitting the dataset with stratifiedgroupkfold.</p>",
  "messages": [
    {
      "id": 1803858,
      "postDate": "2022-05-28T08:53:50.883Z",
      "content": "<p>I saw some of the notebooks are using stratifiedkfold but I would prefer to use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html\" target=\"_blank\">StratifiedGroupKFold</a> here since test set has unseen customer IDs.</p>\n<pre><code>2022-05-28 11:41:07 INFO validation - create_folds: Dataset split into 5 folds (Seed: 42)\n2022-05-28 11:41:08 INFO validation - create_folds: Fold 1 (1106140, 196) - Target Mean: 0.2485 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:09 INFO validation - create_folds: Fold 2 (1106518, 196) - Target Mean: 0.2487 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:10 INFO validation - create_folds: Fold 3 (1105984, 196) - Target Mean: 0.2483 Std: 0.4320 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 4 (1106699, 196) - Target Mean: 0.2495 Std: 0.4328 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 5 (1106110, 196) - Target Mean: 0.2504 Std: 0.4333 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: 0 overlapping customer_IDs\n</code></pre>\n<p>Those are my folds after splitting the dataset with stratifiedgroupkfold.</p>",
      "rawMarkdown": "I saw some of the notebooks are using stratifiedkfold but I would prefer to use [StratifiedGroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html) here since test set has unseen customer IDs.\n\n```\n2022-05-28 11:41:07 INFO validation - create_folds: Dataset split into 5 folds (Seed: 42)\n2022-05-28 11:41:08 INFO validation - create_folds: Fold 1 (1106140, 196) - Target Mean: 0.2485 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:09 INFO validation - create_folds: Fold 2 (1106518, 196) - Target Mean: 0.2487 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:10 INFO validation - create_folds: Fold 3 (1105984, 196) - Target Mean: 0.2483 Std: 0.4320 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 4 (1106699, 196) - Target Mean: 0.2495 Std: 0.4328 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 5 (1106110, 196) - Target Mean: 0.2504 Std: 0.4333 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: 0 overlapping customer_IDs\n```\n\nThose are my folds after splitting the dataset with stratifiedgroupkfold.",
      "votes": 10
    },
    {
      "id": 1803905,
      "postDate": "2022-05-28T09:38:29.930Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Most of these notebooks compute one prediction per customer rather than one prediction per statement. They group the data by customer_ID before starting cross-validation. If our training data consists of one row per customer_ID, the StratifiedKFold suffices.</p>",
      "rawMarkdown": "Hi @gunesevitan Most of these notebooks compute one prediction per customer rather than one prediction per statement. They group the data by customer_ID before starting cross-validation. If our training data consists of one row per customer_ID, the StratifiedKFold suffices.",
      "votes": 6,
      "replies": [
        {
          "id": 1803914,
          "postDate": "2022-05-28T09:54:53.460Z",
          "content": "<p>You are right. I didn't notice that part. Do you think training customers as single data points makes more sense rather than aggregating raw data point predictions?</p>",
          "rawMarkdown": "You are right. I didn't notice that part. Do you think training customers as single data points makes more sense rather than aggregating raw data point predictions?",
          "votes": 2
        },
        {
          "id": 1803929,
          "postDate": "2022-05-28T10:29:51.033Z",
          "content": "<p>Treating customers as single data points is the only way to go: Let's assume we have only one feature \\(x_{i, j}\\) where \\(i\\) is the number of the customer and \\(j \\in 1,…,13\\) is the number of the statement. If the label of customer \\(i\\) depends on \\(\\sum_{j=1,12}{x_{i,j}} - x_{i,13}\\) (that is, the effect of the most recent statement is inverted), a model can learn this dependency only if it sees the customer as a whole.</p>",
          "rawMarkdown": "Treating customers as single data points is the only way to go: Let's assume we have only one feature \\\\(x_{i, j}\\\\) where \\\\(i\\\\) is the number of the customer and \\\\(j \\in 1,...,13\\\\) is the number of the statement. If the label of customer \\\\(i\\\\) depends on \\\\(\\sum_{j=1,12}{x_{i,j}} - x_{i,13}\\\\) (that is, the effect of the most recent statement is inverted), a model can learn this dependency only if it sees the customer as a whole.",
          "votes": 4
        },
        {
          "id": 1803938,
          "postDate": "2022-05-28T10:57:20.053Z",
          "content": "<p>You are absolutely right. Thanks for explaining the statement concept. I have some questions. </p>\n<p>Is it confirmed that the last statement is correctly ordered for every customer? What is the temporal relationship between those statements per customer? Are they sorted by time? </p>",
          "rawMarkdown": "You are absolutely right. Thanks for explaining the statement concept. I have some questions. \n\nIs it confirmed that the last statement is correctly ordered for every customer? What is the temporal relationship between those statements per customer? Are they sorted by time? ",
          "votes": 1
        },
        {
          "id": 1804001,
          "postDate": "2022-05-28T12:21:57.393Z",
          "content": "<p>Statements can be sorted by the statement date(S_2), wherein April 2018 is the statement on which Default is observed, for details check <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327595\" target=\"_blank\">this</a></p>\n<p>For validation, Ideal case would have been to have multiple statement time periods in the train set similar to test set, Check <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327602\" target=\"_blank\">this</a></p>",
          "rawMarkdown": "Statements can be sorted by the statement date(S_2), wherein April 2018 is the statement on which Default is observed, for details check [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327595)\n\n\nFor validation, Ideal case would have been to have multiple statement time periods in the train set similar to test set, Check [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327602)\n",
          "votes": 2
        },
        {
          "id": 1804002,
          "postDate": "2022-05-28T12:22:20.250Z",
          "content": "<p>Yes, a customer's statements are sorted by time. A customer gets one statement per month, and for every customer we see between 1 and 13 months.</p>",
          "rawMarkdown": "Yes, a customer's statements are sorted by time. A customer gets one statement per month, and for every customer we see between 1 and 13 months.",
          "votes": 1
        },
        {
          "id": 1804009,
          "postDate": "2022-05-28T12:35:52.540Z",
          "content": "<p>Yes, 13 is the max number of months for someone to go 120+ delinquent in 18th month</p>",
          "rawMarkdown": "Yes, 13 is the max number of months for someone to go 120+ delinquent in 18th month"
        },
        {
          "id": 1808518,
          "postDate": "2022-06-01T23:26:05.213Z",
          "content": "<p><a href=\"https://www.kaggle.com/taran8727\" target=\"_blank\">@taran8727</a> hmm since the train set provided is 13 months so either way that's the max number of months we can have, but theoretically should it be 14 the max number of months to go 120 days delinquent in 18 months window? since 14 + 4 (120 days) = 18 instead of 13 being the max?</p>",
          "rawMarkdown": "@taran8727 hmm since the train set provided is 13 months so either way that's the max number of months we can have, but theoretically should it be 14 the max number of months to go 120 days delinquent in 18 months window? since 14 + 4 (120 days) = 18 instead of 13 being the max?"
        },
        {
          "id": 1814514,
          "postDate": "2022-06-08T02:05:58.043Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1803905,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2022-05-28T09:38:29.930000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> Most of these notebooks compute one prediction per customer rather than one prediction per statement. They group the data by customer_ID before starting cross-validation. If our training data consists of one row per customer_ID, the StratifiedKFold suffices.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1803914,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-05-28T09:54:53.460000",
          "content": "<p>You are right. I didn't notice that part. Do you think training customers as single data points makes more sense rather than aggregating raw data point predictions?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1803929,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-05-28T10:29:51.033000",
          "content": "<p>Treating customers as single data points is the only way to go: Let's assume we have only one feature \\(x_{i, j}\\) where \\(i\\) is the number of the customer and \\(j \\in 1,…,13\\) is the number of the statement. If the label of customer \\(i\\) depends on \\(\\sum_{j=1,12}{x_{i,j}} - x_{i,13}\\) (that is, the effect of the most recent statement is inverted), a model can learn this dependency only if it sees the customer as a whole.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1803938,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-05-28T10:57:20.053000",
          "content": "<p>You are absolutely right. Thanks for explaining the statement concept. I have some questions. </p>\n<p>Is it confirmed that the last statement is correctly ordered for every customer? What is the temporal relationship between those statements per customer? Are they sorted by time? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1804001,
          "author_name": "Seeker",
          "author_url": "",
          "post_date": "2022-05-28T12:21:57.393000",
          "content": "<p>Statements can be sorted by the statement date(S_2), wherein April 2018 is the statement on which Default is observed, for details check <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327595\" target=\"_blank\">this</a></p>\n<p>For validation, Ideal case would have been to have multiple statement time periods in the train set similar to test set, Check <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327602\" target=\"_blank\">this</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1804002,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-05-28T12:22:20.250000",
          "content": "<p>Yes, a customer's statements are sorted by time. A customer gets one statement per month, and for every customer we see between 1 and 13 months.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1804009,
          "author_name": "Seeker",
          "author_url": "",
          "post_date": "2022-05-28T12:35:52.540000",
          "content": "<p>Yes, 13 is the max number of months for someone to go 120+ delinquent in 18th month</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1808518,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "2022-06-01T23:26:05.213000",
          "content": "<p><a href=\"https://www.kaggle.com/taran8727\" target=\"_blank\">@taran8727</a> hmm since the train set provided is 13 months so either way that's the max number of months we can have, but theoretically should it be 14 the max number of months to go 120 days delinquent in 18 months window? since 14 + 4 (120 days) = 18 instead of 13 being the max?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1814514,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-08T02:05:58.043000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1803858": "I saw some of the notebooks are using stratifiedkfold but I would prefer to use [StratifiedGroupKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedGroupKFold.html) here since test set has unseen customer IDs.\n\n```\n2022-05-28 11:41:07 INFO validation - create_folds: Dataset split into 5 folds (Seed: 42)\n2022-05-28 11:41:08 INFO validation - create_folds: Fold 1 (1106140, 196) - Target Mean: 0.2485 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:09 INFO validation - create_folds: Fold 2 (1106518, 196) - Target Mean: 0.2487 Std: 0.4322 (91782 Unique Customer IDs)\n2022-05-28 11:41:10 INFO validation - create_folds: Fold 3 (1105984, 196) - Target Mean: 0.2483 Std: 0.4320 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 4 (1106699, 196) - Target Mean: 0.2495 Std: 0.4328 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: Fold 5 (1106110, 196) - Target Mean: 0.2504 Std: 0.4333 (91783 Unique Customer IDs)\n2022-05-28 11:41:11 INFO validation - create_folds: 0 overlapping customer_IDs\n```\n\nThose are my folds after splitting the dataset with stratifiedgroupkfold.",
    "1803905": "Hi @gunesevitan Most of these notebooks compute one prediction per customer rather than one prediction per statement. They group the data by customer_ID before starting cross-validation. If our training data consists of one row per customer_ID, the StratifiedKFold suffices."
  }
}