{
  "id": 339426,
  "title": "Stratification by P_2_LST can significantly reduce your CV variance",
  "url": "/competitions/amex-default-prediction/discussion/339426",
  "author_name": "Vladislav Bogorod",
  "post_date": "2022-07-24T18:37:54.803000",
  "votes": 43,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Recently I've shared that noticed high variance of CV|LB scores <br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957</a><br>\nThe general thought that were discussed that we have a problem with data and label representation.</p>\n<p>In my model the product feature from P_2_LST has <strong>more than x10 higher</strong> gain score than others.<br>\nIt seems that P_2_LST and its products can make score unstable. <br>\nGenerally I think that we have undersampling data based on this feature and so it can be so sensitive.</p>\n<p>Also I change metric of optimisation from max(auc) to max(auc-stdv) to make score more conservative.<br>\nAnd finally I used several random seed to find best number of iterations not to be overfitted to a single seed that used for optimization.</p>\n<p><strong>Finally I reduce auc-stdv from average 0.00048 to 0.00022.</strong></p>\n<p>The code of optimization process here:<br>\n<a href=\"https://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification\" target=\"_blank\">https://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification</a></p>",
  "messages": [
    {
      "id": 1869428,
      "postDate": "2022-07-24T18:37:54.803Z",
      "content": "<p>Recently I've shared that noticed high variance of CV|LB scores <br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/336957</a><br>\nThe general thought that were discussed that we have a problem with data and label representation.</p>\n<p>In my model the product feature from P_2_LST has <strong>more than x10 higher</strong> gain score than others.<br>\nIt seems that P_2_LST and its products can make score unstable. <br>\nGenerally I think that we have undersampling data based on this feature and so it can be so sensitive.</p>\n<p>Also I change metric of optimisation from max(auc) to max(auc-stdv) to make score more conservative.<br>\nAnd finally I used several random seed to find best number of iterations not to be overfitted to a single seed that used for optimization.</p>\n<p><strong>Finally I reduce auc-stdv from average 0.00048 to 0.00022.</strong></p>\n<p>The code of optimization process here:<br>\n<a href=\"https://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification\" target=\"_blank\">https://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification</a></p>",
      "rawMarkdown": "Recently I've shared that noticed high variance of CV|LB scores \nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\nThe general thought that were discussed that we have a problem with data and label representation.\n\nIn my model the product feature from P_2_LST has **more than x10 higher** gain score than others.\nIt seems that P_2_LST and its products can make score unstable. \nGenerally I think that we have undersampling data based on this feature and so it can be so sensitive.\n\nAlso I change metric of optimisation from max(auc) to max(auc-stdv) to make score more conservative.\nAnd finally I used several random seed to find best number of iterations not to be overfitted to a single seed that used for optimization.\n\n**Finally I reduce auc-stdv from average 0.00048 to 0.00022.**\n\nThe code of optimization process here:\nhttps://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification",
      "votes": 41
    },
    {
      "id": 1869497,
      "postDate": "2022-07-24T19:45:09.630Z",
      "content": "<p>I would say 'CV bias\" is not correct term - should be CV variance. Does this help to reduce <code>D</code> variance in the metric?</p>",
      "rawMarkdown": "I would say 'CV bias\" is not correct term - should be CV variance. Does this help to reduce `D` variance in the metric?\n",
      "votes": 4,
      "replies": [
        {
          "id": 1869509,
          "postDate": "2022-07-24T19:51:32.083Z",
          "content": "<p>Yes, it does. You are right. Thank you for correcting. <br>\nI have almost the same average AUC but variance reduced on hold-out CV (for every tested seed).</p>",
          "rawMarkdown": "Yes, it does. You are right. Thank you for correcting. \nI have almost the same average AUC but variance reduced on hold-out CV (for every tested seed).",
          "votes": 4
        },
        {
          "id": 1869518,
          "postDate": "2022-07-24T20:01:00.047Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1871501,
      "postDate": "2022-07-26T09:57:29.213Z",
      "content": "<p>Thank you for sharing this! I was also having issues with high variance of the CV. I am going to try this for sure.</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Thank you for sharing this! I was also having issues with high variance of the CV. I am going to try this for sure.\n\nThe Devastator.\n",
      "votes": -2
    },
    {
      "id": 1884836,
      "postDate": "2022-08-04T17:19:34.373Z",
      "content": "<p>Does \"P_2_lst\" refer to \"the value of P_2 on the last record?\"</p>",
      "rawMarkdown": "Does \"P_2_lst\" refer to \"the value of P_2 on the last record?\"",
      "replies": [
        {
          "id": 1885385,
          "postDate": "2022-08-05T06:37:18.547Z",
          "content": "<p>If we have a row of data [JAN, FEF, …, DEC] and the last available month is DEC and REPORT MONTH is DEC then P_2_LST refer to DEC. <br>\nBut in my case of feature processing it looks like <br>\n<strong>values[~np.isnan(values)][-1] if len(~np.isnan(values)) else np.nan</strong> # your may be different<br>\n, so IF DEC IS NULL THEN NOV if exist, eg. the last available</p>",
          "rawMarkdown": "If we have a row of data [JAN, FEF, ..., DEC] and the last available month is DEC and REPORT MONTH is DEC then P_2_LST refer to DEC. \nBut in my case of feature processing it looks like \n**values[~np.isnan(values)][-1] if len(~np.isnan(values)) else np.nan** # your may be different\n, so IF DEC IS NULL THEN NOV if exist, eg. the last available"
        }
      ]
    },
    {
      "id": 1876173,
      "postDate": "2022-07-29T16:09:45.023Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1877636,
          "postDate": "2022-07-30T21:22:32.810Z",
          "content": "<p>Took me awhile to parse it. I'm not expert but 99% sure that bucket_id is equivalent to labels for stratification by StratifiedKFold. </p>\n<p>So if 5 folds, then exactly five rows will get bucket_id 1, 5 more get bucket_id 2, and so on. He also separates by target 0 vs 1, so another 5 will also get bucket_id 1, but later he adds 1,000,000 to all bucket IDs for target=1. </p>\n<p>So separate buckets for each. </p>",
          "rawMarkdown": "Took me awhile to parse it. I'm not expert but 99% sure that bucket_id is equivalent to labels for stratification by StratifiedKFold. \n\nSo if 5 folds, then exactly five rows will get bucket_id 1, 5 more get bucket_id 2, and so on. He also separates by target 0 vs 1, so another 5 will also get bucket_id 1, but later he adds 1,000,000 to all bucket IDs for target=1. \n\nSo separate buckets for each. "
        },
        {
          "id": 1892943,
          "postDate": "2022-08-10T12:49:44.047Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1898155,
      "postDate": "2022-08-14T11:02:56.237Z",
      "content": "<p>Thanks for sharing </p>",
      "rawMarkdown": "Thanks for sharing "
    },
    {
      "id": 1897689,
      "postDate": "2022-08-14T02:22:47.130Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing"
    },
    {
      "id": 1892708,
      "postDate": "2022-08-10T09:16:10.597Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing"
    },
    {
      "id": 1887969,
      "postDate": "2022-08-07T08:26:49.027Z",
      "content": "<p>Thank you for sharing this!</p>",
      "rawMarkdown": "Thank you for sharing this!"
    },
    {
      "id": 1882112,
      "postDate": "2022-08-03T04:50:13.740Z",
      "content": "<p>Thank you for sharing this! </p>",
      "rawMarkdown": "Thank you for sharing this! "
    }
  ],
  "comments": [
    {
      "id": 1869497,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-07-24T19:45:09.630000",
      "content": "<p>I would say 'CV bias\" is not correct term - should be CV variance. Does this help to reduce <code>D</code> variance in the metric?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1869509,
          "author_name": "Vladislav Bogorod",
          "author_url": "",
          "post_date": "2022-07-24T19:51:32.083000",
          "content": "<p>Yes, it does. You are right. Thank you for correcting. <br>\nI have almost the same average AUC but variance reduced on hold-out CV (for every tested seed).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1869518,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-24T20:01:00.047000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1871501,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-26T09:57:29.213000",
      "content": "<p>Thank you for sharing this! I was also having issues with high variance of the CV. I am going to try this for sure.</p>\n<p>The Devastator.</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 1884836,
      "author_name": "DavidIRudel",
      "author_url": "",
      "post_date": "2022-08-04T17:19:34.373000",
      "content": "<p>Does \"P_2_lst\" refer to \"the value of P_2 on the last record?\"</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1885385,
          "author_name": "Vladislav Bogorod",
          "author_url": "",
          "post_date": "2022-08-05T06:37:18.547000",
          "content": "<p>If we have a row of data [JAN, FEF, …, DEC] and the last available month is DEC and REPORT MONTH is DEC then P_2_LST refer to DEC. <br>\nBut in my case of feature processing it looks like <br>\n<strong>values[~np.isnan(values)][-1] if len(~np.isnan(values)) else np.nan</strong> # your may be different<br>\n, so IF DEC IS NULL THEN NOV if exist, eg. the last available</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1876173,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-29T16:09:45.023000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1877636,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-30T21:22:32.810000",
          "content": "<p>Took me awhile to parse it. I'm not expert but 99% sure that bucket_id is equivalent to labels for stratification by StratifiedKFold. </p>\n<p>So if 5 folds, then exactly five rows will get bucket_id 1, 5 more get bucket_id 2, and so on. He also separates by target 0 vs 1, so another 5 will also get bucket_id 1, but later he adds 1,000,000 to all bucket IDs for target=1. </p>\n<p>So separate buckets for each. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1892943,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-10T12:49:44.047000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1898155,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T11:02:56.237000",
      "content": "<p>Thanks for sharing </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897689,
      "author_name": "ruir",
      "author_url": "",
      "post_date": "2022-08-14T02:22:47.130000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1892708,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-10T09:16:10.597000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1887969,
      "author_name": "coco",
      "author_url": "",
      "post_date": "2022-08-07T08:26:49.027000",
      "content": "<p>Thank you for sharing this!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1882112,
      "author_name": "Wynmolly",
      "author_url": "",
      "post_date": "2022-08-03T04:50:13.740000",
      "content": "<p>Thank you for sharing this! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1869428": "Recently I've shared that noticed high variance of CV|LB scores \nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/336957\nThe general thought that were discussed that we have a problem with data and label representation.\n\nIn my model the product feature from P_2_LST has **more than x10 higher** gain score than others.\nIt seems that P_2_LST and its products can make score unstable. \nGenerally I think that we have undersampling data based on this feature and so it can be so sensitive.\n\nAlso I change metric of optimisation from max(auc) to max(auc-stdv) to make score more conservative.\nAnd finally I used several random seed to find best number of iterations not to be overfitted to a single seed that used for optimization.\n\n**Finally I reduce auc-stdv from average 0.00048 to 0.00022.**\n\nThe code of optimization process here:\nhttps://www.kaggle.com/code/bogorodvo/bayesianoptimization-p-2-lst-stratification",
    "1869497": "I would say 'CV bias\" is not correct term - should be CV variance. Does this help to reduce `D` variance in the metric?\n",
    "1871501": "Thank you for sharing this! I was also having issues with high variance of the CV. I am going to try this for sure.\n\nThe Devastator.\n",
    "1884836": "Does \"P_2_lst\" refer to \"the value of P_2 on the last record?\"",
    "1876173": "",
    "1898155": "Thanks for sharing ",
    "1897689": "Thank you for sharing",
    "1892708": "Thank you for sharing",
    "1887969": "Thank you for sharing this!",
    "1882112": "Thank you for sharing this! "
  }
}