{
  "id": 343004,
  "title": "## Using Encoding Target method",
  "url": "/competitions/amex-default-prediction/discussion/343004",
  "author_name": "",
  "post_date": "2022-08-09T15:12:01.893415400Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everybody, I have some questions about using the Encoding Target method with some way of regularization encoding Target, After using those ways, I found my CV go up above 0.95 (Amex score), but my score in public LeaderBoard got down less than 0.6. I know the Encoding Target method is 100% overfitting or leakage method, but the huge variance between my CV and LB is not the only cause of using Encoding Target, maybe there is some reason for this result, such as test data and train data are completely different. If anyone has any advice about this problem, please show us your reason for this problem. and thanks.</p>",
  "messages": [
    {
      "id": "1891669",
      "postDate": "08/09/2022 15:12:01",
      "content": "<p>Hello everybody, I have some questions about using the Encoding Target method with some way of regularization encoding Target, After using those ways, I found my CV go up above 0.95 (Amex score), but my score in public LeaderBoard got down less than 0.6. I know the Encoding Target method is 100% overfitting or leakage method, but the huge variance between my CV and LB is not the only cause of using Encoding Target, maybe there is some reason for this result, such as test data and train data are completely different. If anyone has any advice about this problem, please show us your reason for this problem. and thanks.</p>",
      "rawMarkdown": "Hello everybody, I have some questions about using the Encoding Target method with some way of regularization encoding Target, After using those ways, I found my CV go up above 0.95 (Amex score), but my score in public LeaderBoard got down less than 0.6. I know the Encoding Target method is 100% overfitting or leakage method, but the huge variance between my CV and LB is not the only cause of using Encoding Target, maybe there is some reason for this result, such as test data and train data are completely different. If anyone has any advice about this problem, please show us your reason for this problem. and thanks.",
      "votes": null
    },
    {
      "id": "1891826",
      "postDate": "08/09/2022 17:12:26",
      "content": "<p>It is absolutely possible that the difference you see is only because of overfitting by target encoding. Test and train data are not completely different, or else we would not be getting good correlation between CV and LB scores.</p>",
      "rawMarkdown": "It is absolutely possible that the difference you see is only because of overfitting by target encoding. Test and train data are not completely different, or else we would not be getting good correlation between CV and LB scores.",
      "votes": null
    },
    {
      "id": "1891833",
      "postDate": "08/09/2022 17:18:37",
      "content": "<p>Yeah, maybe you're right, but I added 90% of the noise to my encoding target, and the problem still exists.</p>",
      "rawMarkdown": "Yeah, maybe you're right, but I added 90% of the noise to my encoding target, and the problem still exists.",
      "votes": null
    },
    {
      "id": "1891848",
      "postDate": "08/09/2022 17:29:04",
      "content": "<p>For this competition (and this metric) it may not be possible at all to do it without overfitting, but this could be worth a try:</p>\n<p><a href=\"https://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features</a></p>",
      "rawMarkdown": "For this competition (and this metric) it may not be possible at all to do it without overfitting, but this could be worth a try:\n\nhttps://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features",
      "votes": null
    },
    {
      "id": "1892067",
      "postDate": "08/09/2022 20:15:15",
      "content": "<p>Target encoding must be performed under k-fold to prevent overfitting, complicated model like LGBM or NNs are smart enough to get the information even you have already added laplace smoothing or any regularization/noise.</p>",
      "rawMarkdown": "Target encoding must be performed under k-fold to prevent overfitting, complicated model like LGBM or NNs are smart enough to get the information even you have already added laplace smoothing or any regularization/noise.",
      "votes": null
    },
    {
      "id": "1892079",
      "postDate": "08/09/2022 20:29:42",
      "content": "<p>I did that before, but I think I need to try something else for stopping the leakage of data when I use Target encoding.</p>",
      "rawMarkdown": "I did that before, but I think I need to try something else for stopping the leakage of data when I use Target encoding.",
      "votes": null
    },
    {
      "id": "1892126",
      "postDate": "08/09/2022 21:31:09",
      "content": "<p>That sounds like a more fundamental problem on test inference. </p>\n<p>Either some huge unrelated bug that oddly got introduced with your latest changes (scrambled customer IDs might still manage 0.6). Or a mistake in the methodology of target encoding of the test data (or train data) like just not encoding the test data at all, or an issue in the implementation. </p>",
      "rawMarkdown": "That sounds like a more fundamental problem on test inference. \n\nEither some huge unrelated bug that oddly got introduced with your latest changes (scrambled customer IDs might still manage 0.6). Or a mistake in the methodology of target encoding of the test data (or train data) like just not encoding the test data at all, or an issue in the implementation.",
      "votes": null
    },
    {
      "id": "1892152",
      "postDate": "08/09/2022 22:10:43",
      "content": "<p>Target encoding for me is like reverse engineering in a simple way, So normally we discover the association or correlation or causality from the features to the Target.  but with Encoding target, we start from Target to the features, <br>\nSo, the first error or mistake is by giving the signs to the features which never affect the Target,  <br>\nthe second error is when we have different distribution between train set and test set, the error is by leakage of the test set from train set, So we need to do some regularization such as CV methods, <br>\nSo, we need to add some methods to reduce the second problem, <br>\nFor me, I did that, So, I need to test my hypothesis about the test set and train set if they have a different distribution.</p>",
      "rawMarkdown": "Target encoding for me is like reverse engineering in a simple way, So normally we discover the association or correlation or causality from the features to the Target.  but with Encoding target, we start from Target to the features, \nSo, the first error or mistake is by giving the signs to the features which never affect the Target,  \nthe second error is when we have different distribution between train set and test set, the error is by leakage of the test set from train set, So we need to do some regularization such as CV methods, \nSo, we need to add some methods to reduce the second problem, \nFor me, I did that, So, I need to test my hypothesis about the test set and train set if they have a different distribution.",
      "votes": null
    },
    {
      "id": "1892177",
      "postDate": "08/09/2022 22:58:18",
      "content": "<p>Let me put it this way: when I make a mistake in my code, I'd feel pretty lucky if I get a score of 0.6. It's much harder to know the root issue if I get a .791 or even a .780. But a 0.6 leaves no realistic possibility that target encoding was implemented correctly but doesn't work for this dataset. It's narrowed down to a bug in the code. </p>\n<p>On a prior snapshot without target encoding, try taking a few entire columns, on the test data only, and scramble those columns. I bet you still score over 0.75. With entire bogus columns. Unless you scramble P_2 or something extra critical. Point is, the issue isn't train vs test issue, I think. </p>",
      "rawMarkdown": "Let me put it this way: when I make a mistake in my code, I'd feel pretty lucky if I get a score of 0.6. It's much harder to know the root issue if I get a .791 or even a .780. But a 0.6 leaves no realistic possibility that target encoding was implemented correctly but doesn't work for this dataset. It's narrowed down to a bug in the code. \n\nOn a prior snapshot without target encoding, try taking a few entire columns, on the test data only, and scramble those columns. I bet you still score over 0.75. With entire bogus columns. Unless you scramble P_2 or something extra critical. Point is, the issue isn't train vs test issue, I think.",
      "votes": null
    },
    {
      "id": "1895968",
      "postDate": "08/12/2022 13:21:33",
      "content": "<p>What was output? Was it a class '0' and '1'  or probablity?</p>",
      "rawMarkdown": "What was output? Was it a class '0' and '1'  or probablity?",
      "votes": null
    },
    {
      "id": "1896506",
      "postDate": "08/12/2022 21:18:57",
      "content": "<p>Mean of Target equal to 1, or frequency of Target equal to 1</p>",
      "rawMarkdown": "Mean of Target equal to 1, or frequency of Target equal to 1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1891826,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "08/09/2022 17:12:26",
      "content": "<p>It is absolutely possible that the difference you see is only because of overfitting by target encoding. Test and train data are not completely different, or else we would not be getting good correlation between CV and LB scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1891833,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/09/2022 17:18:37",
          "content": "<p>Yeah, maybe you're right, but I added 90% of the noise to my encoding target, and the problem still exists.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1891848,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "08/09/2022 17:29:04",
          "content": "<p>For this competition (and this metric) it may not be possible at all to do it without overfitting, but this could be worth a try:</p>\n<p><a href=\"https://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features\" target=\"_blank\">https://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1892067,
      "author_name": "andrew60909",
      "author_url": "",
      "post_date": "08/09/2022 20:15:15",
      "content": "<p>Target encoding must be performed under k-fold to prevent overfitting, complicated model like LGBM or NNs are smart enough to get the information even you have already added laplace smoothing or any regularization/noise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1892079,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/09/2022 20:29:42",
          "content": "<p>I did that before, but I think I need to try something else for stopping the leakage of data when I use Target encoding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1892126,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "08/09/2022 21:31:09",
      "content": "<p>That sounds like a more fundamental problem on test inference. </p>\n<p>Either some huge unrelated bug that oddly got introduced with your latest changes (scrambled customer IDs might still manage 0.6). Or a mistake in the methodology of target encoding of the test data (or train data) like just not encoding the test data at all, or an issue in the implementation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1892152,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/09/2022 22:10:43",
          "content": "<p>Target encoding for me is like reverse engineering in a simple way, So normally we discover the association or correlation or causality from the features to the Target.  but with Encoding target, we start from Target to the features, <br>\nSo, the first error or mistake is by giving the signs to the features which never affect the Target,  <br>\nthe second error is when we have different distribution between train set and test set, the error is by leakage of the test set from train set, So we need to do some regularization such as CV methods, <br>\nSo, we need to add some methods to reduce the second problem, <br>\nFor me, I did that, So, I need to test my hypothesis about the test set and train set if they have a different distribution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1892177,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/09/2022 22:58:18",
          "content": "<p>Let me put it this way: when I make a mistake in my code, I'd feel pretty lucky if I get a score of 0.6. It's much harder to know the root issue if I get a .791 or even a .780. But a 0.6 leaves no realistic possibility that target encoding was implemented correctly but doesn't work for this dataset. It's narrowed down to a bug in the code. </p>\n<p>On a prior snapshot without target encoding, try taking a few entire columns, on the test data only, and scramble those columns. I bet you still score over 0.75. With entire bogus columns. Unless you scramble P_2 or something extra critical. Point is, the issue isn't train vs test issue, I think. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1895968,
      "author_name": "ravisinghiitbhu",
      "author_url": "",
      "post_date": "08/12/2022 13:21:33",
      "content": "<p>What was output? Was it a class '0' and '1'  or probablity?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1896506,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/12/2022 21:18:57",
          "content": "<p>Mean of Target equal to 1, or frequency of Target equal to 1</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1891669": "Hello everybody, I have some questions about using the Encoding Target method with some way of regularization encoding Target, After using those ways, I found my CV go up above 0.95 (Amex score), but my score in public LeaderBoard got down less than 0.6. I know the Encoding Target method is 100% overfitting or leakage method, but the huge variance between my CV and LB is not the only cause of using Encoding Target, maybe there is some reason for this result, such as test data and train data are completely different. If anyone has any advice about this problem, please show us your reason for this problem. and thanks.",
    "1891826": "It is absolutely possible that the difference you see is only because of overfitting by target encoding. Test and train data are not completely different, or else we would not be getting good correlation between CV and LB scores.",
    "1891833": "Yeah, maybe you're right, but I added 90% of the noise to my encoding target, and the problem still exists.",
    "1891848": "For this competition (and this metric) it may not be possible at all to do it without overfitting, but this could be worth a try:\n\nhttps://www.kaggle.com/code/ogrellier/python-target-encoding-for-categorical-features",
    "1892067": "Target encoding must be performed under k-fold to prevent overfitting, complicated model like LGBM or NNs are smart enough to get the information even you have already added laplace smoothing or any regularization/noise.",
    "1892079": "I did that before, but I think I need to try something else for stopping the leakage of data when I use Target encoding.",
    "1892126": "That sounds like a more fundamental problem on test inference. \n\nEither some huge unrelated bug that oddly got introduced with your latest changes (scrambled customer IDs might still manage 0.6). Or a mistake in the methodology of target encoding of the test data (or train data) like just not encoding the test data at all, or an issue in the implementation.",
    "1892152": "Target encoding for me is like reverse engineering in a simple way, So normally we discover the association or correlation or causality from the features to the Target.  but with Encoding target, we start from Target to the features, \nSo, the first error or mistake is by giving the signs to the features which never affect the Target,  \nthe second error is when we have different distribution between train set and test set, the error is by leakage of the test set from train set, So we need to do some regularization such as CV methods, \nSo, we need to add some methods to reduce the second problem, \nFor me, I did that, So, I need to test my hypothesis about the test set and train set if they have a different distribution.",
    "1892177": "Let me put it this way: when I make a mistake in my code, I'd feel pretty lucky if I get a score of 0.6. It's much harder to know the root issue if I get a .791 or even a .780. But a 0.6 leaves no realistic possibility that target encoding was implemented correctly but doesn't work for this dataset. It's narrowed down to a bug in the code. \n\nOn a prior snapshot without target encoding, try taking a few entire columns, on the test data only, and scramble those columns. I bet you still score over 0.75. With entire bogus columns. Unless you scramble P_2 or something extra critical. Point is, the issue isn't train vs test issue, I think.",
    "1895968": "What was output? Was it a class '0' and '1'  or probablity?",
    "1896506": "Mean of Target equal to 1, or frequency of Target equal to 1"
  },
  "source": "meta"
}