{
  "id": 346617,
  "title": "Amex Metric Score - 0.99145420 with test data but public leaderboard score - 0.57",
  "url": "/competitions/amex-default-prediction/discussion/346617",
  "author_name": "Neha",
  "post_date": "2022-08-20T14:58:18.180000",
  "votes": -5,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Amex Metric Score is coming to be 0.99145420 (test data) and 0.999958889 (training data) but the competition public leaderboard score shows 0.57.</p>\n<p>Can someone help to understand the possible reason behind this?</p>",
  "messages": [
    {
      "id": 1907886,
      "postDate": "2022-08-21T07:04:36.097Z",
      "content": "<p>Just look into your model's feature importance, you can use Shapley summary (beeswarm) plot for example <a href=\"https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html\" target=\"_blank\">https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html</a> to identify most important features, my guess is the most important one, or first top N, contain the leak.</p>\n<p>Other option is to try to build single feature models and see if any model is much better than it should be.</p>\n<p>Also when <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> suggests to have a look at feature names:</p>\n<pre><code>print([col for col in X_train.columns if 'target' in col])\n</code></pre>\n<p>it prints every column that has the substring \"target\" as part of it, for example \"encoded_target\" would match, while your code</p>\n<pre><code>'target' in S.columns\n</code></pre>\n<p>would not work, as it only prints the column name if it is the exact text \"target\".</p>\n<p>I believe you either leak the target in the features, or if you train an ensemble, you might be training it on train + test and then evaluating on test.</p>\n<p>Good luck with finding the issue.</p>",
      "rawMarkdown": "Just look into your model's feature importance, you can use Shapley summary (beeswarm) plot for example https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html to identify most important features, my guess is the most important one, or first top N, contain the leak.\n\nOther option is to try to build single feature models and see if any model is much better than it should be.\n\nAlso when @imeintanis suggests to have a look at feature names:\n```\nprint([col for col in X_train.columns if 'target' in col])\n```\nit prints every column that has the substring \"target\" as part of it, for example \"encoded_target\" would match, while your code\n```\n'target' in S.columns\n```\nwould not work, as it only prints the column name if it is the exact text \"target\".\n\nI believe you either leak the target in the features, or if you train an ensemble, you might be training it on train + test and then evaluating on test.\n\nGood luck with finding the issue.",
      "votes": 5,
      "replies": [
        {
          "id": 1908042,
          "postDate": "2022-08-21T10:00:41.717Z",
          "content": "<p>Exactly!! Thank you for making it clear</p>\n<p>I bet it contains somethink like \"target_diff\", \"target_mean\", etc or \"target_x\" after multiple merges</p>",
          "rawMarkdown": "Exactly!! Thank you for making it clear\n\nI bet it contains somethink like \"target_diff\", \"target_mean\", etc or \"target_x\" after multiple merges",
          "votes": 1
        },
        {
          "id": 1908084,
          "postDate": "2022-08-21T10:56:45.957Z",
          "content": "<p>Thank you for your responses. <br>\nTarget has not been chased at all to contain any column like these.<br>\nTrain and test data are also separate. </p>\n<p>Maybe the issue is multicollinear columns.</p>",
          "rawMarkdown": "Thank you for your responses. \nTarget has not been chased at all to contain any column like these.\nTrain and test data are also separate. \n\nMaybe the issue is multicollinear columns."
        },
        {
          "id": 1908086,
          "postDate": "2022-08-21T10:59:22.967Z",
          "content": "<p>print([col for col in X_train.columns if 'target' in col])<br>\nOutput for this is an empty list.</p>",
          "rawMarkdown": "print([col for col in X_train.columns if 'target' in col])\nOutput for this is an empty list.\n"
        }
      ]
    },
    {
      "id": 1907537,
      "postDate": "2022-08-20T21:36:45.243Z",
      "content": "<p>Do you have target-encoding features in your training data?</p>",
      "rawMarkdown": "Do you have target-encoding features in your training data?",
      "votes": 1,
      "replies": [
        {
          "id": 1907594,
          "postDate": "2022-08-20T23:39:33.850Z",
          "content": "<p>Hey,<br>\nno, haven't tried target-encoding.<br>\nBut have done one hot encoding.</p>",
          "rawMarkdown": "Hey,\nno, haven't tried target-encoding.\nBut have done one hot encoding."
        }
      ]
    },
    {
      "id": 1907324,
      "postDate": "2022-08-20T17:07:50.320Z",
      "content": "<p>U might be having target info in your training data. </p>",
      "rawMarkdown": "U might be having target info in your training data. ",
      "votes": 1,
      "replies": [
        {
          "id": 1907397,
          "postDate": "2022-08-20T18:22:48.970Z",
          "content": "<p>No, target data has been removed before splitting the dataset.</p>",
          "rawMarkdown": "No, target data has been removed before splitting the dataset."
        },
        {
          "id": 1907558,
          "postDate": "2022-08-20T22:28:26.490Z",
          "content": "<p>Certainly you have target col (or derivative) in your features, search better.. I'd suggest to print out and check your features again</p>",
          "rawMarkdown": "Certainly you have target col (or derivative) in your features, search better.. I'd suggest to print out and check your features again",
          "votes": 1
        },
        {
          "id": 1907595,
          "postDate": "2022-08-20T23:40:51.737Z",
          "content": "<p>target col has been removed from X_train and X_test. </p>",
          "rawMarkdown": "target col has been removed from X_train and X_test. "
        },
        {
          "id": 1907661,
          "postDate": "2022-08-21T00:46:04.887Z",
          "content": "<p>I bet some target info is there! check each feature you have created one by one! <br>\nor use this print statement just before you pass it to fit<br>\n<code>print([col for col in X_train.columns if 'target' in col])</code></p>\n<p>PS: there is no \"real\" model that can score 99% :)  </p>",
          "rawMarkdown": "I bet some target info is there! check each feature you have created one by one! \nor use this print statement just before you pass it to fit\n`print([col for col in X_train.columns if 'target' in col])`\n\nPS: there is no \"real\" model that can score 99% :)  ",
          "votes": 1
        },
        {
          "id": 1907669,
          "postDate": "2022-08-21T00:53:29.473Z",
          "content": "<p><strong>I ran this statement.</strong><br>\n'target' in S.columns</p>\n<p>Output: <strong>False</strong><br>\nS is my training data.</p>",
          "rawMarkdown": "**I ran this statement.**\n'target' in S.columns\n\nOutput: **False**\nS is my training data.\n"
        },
        {
          "id": 1908141,
          "postDate": "2022-08-21T11:50:55.043Z",
          "content": "<p>Maybe you have changed the name or something else? Search for the values in your dataframe. Maybe you have used the target column and aggregated it within a different feature etc. etc.</p>\n<p>There is no way to score this high and that low on public lb</p>\n<p>Good luck!</p>",
          "rawMarkdown": "Maybe you have changed the name or something else? Search for the values in your dataframe. Maybe you have used the target column and aggregated it within a different feature etc. etc.\n\nThere is no way to score this high and that low on public lb\n\nGood luck!"
        }
      ]
    },
    {
      "id": 1907189,
      "postDate": "2022-08-20T15:08:11.087Z",
      "content": "<p>Looks like you have a leak in your training. Is the target in your training data?</p>",
      "rawMarkdown": "Looks like you have a leak in your training. Is the target in your training data?",
      "votes": 1,
      "replies": [
        {
          "id": 1907398,
          "postDate": "2022-08-20T18:23:34.570Z",
          "content": "<p>No, target is not part of X_train and X_test</p>",
          "rawMarkdown": "No, target is not part of X_train and X_test"
        },
        {
          "id": 1907633,
          "postDate": "2022-08-21T00:13:44.700Z",
          "content": "<p>You definitively have a leak somewhere in your pipeline. 0.99 is impossible to reach on the test set.</p>",
          "rawMarkdown": "You definitively have a leak somewhere in your pipeline. 0.99 is impossible to reach on the test set.",
          "votes": 2
        },
        {
          "id": 1907639,
          "postDate": "2022-08-21T00:24:06.733Z",
          "content": "<p>Trying to figure it out.</p>",
          "rawMarkdown": "Trying to figure it out."
        }
      ]
    },
    {
      "id": 1907671,
      "postDate": "2022-08-21T00:54:48.403Z",
      "content": "<p>I ran this statement.<br>\n'target' in S.columns</p>\n<p>Output: False<br>\nS is my training data.</p>",
      "rawMarkdown": "I ran this statement.\n'target' in S.columns\n\nOutput: False\nS is my training data.",
      "votes": -1,
      "replies": [
        {
          "id": 1908137,
          "postDate": "2022-08-21T11:48:07.510Z",
          "content": "<p>Dump out the column names in your dataset to a CSV dataset and go through them one by one.</p>",
          "rawMarkdown": "Dump out the column names in your dataset to a CSV dataset and go through them one by one."
        }
      ]
    },
    {
      "id": 1907177,
      "postDate": "2022-08-20T14:58:18.180Z",
      "content": "<p>Amex Metric Score is coming to be 0.99145420 (test data) and 0.999958889 (training data) but the competition public leaderboard score shows 0.57.</p>\n<p>Can someone help to understand the possible reason behind this?</p>",
      "rawMarkdown": "Amex Metric Score is coming to be 0.99145420 (test data) and 0.999958889 (training data) but the competition public leaderboard score shows 0.57.\n\nCan someone help to understand the possible reason behind this?",
      "votes": -5
    },
    {
      "id": 1907906,
      "postDate": "2022-08-21T07:16:42.650Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1907886,
      "author_name": "MartinBarus",
      "author_url": "",
      "post_date": "2022-08-21T07:04:36.097000",
      "content": "<p>Just look into your model's feature importance, you can use Shapley summary (beeswarm) plot for example <a href=\"https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html\" target=\"_blank\">https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html</a> to identify most important features, my guess is the most important one, or first top N, contain the leak.</p>\n<p>Other option is to try to build single feature models and see if any model is much better than it should be.</p>\n<p>Also when <a href=\"https://www.kaggle.com/imeintanis\" target=\"_blank\">@imeintanis</a> suggests to have a look at feature names:</p>\n<pre><code>print([col for col in X_train.columns if 'target' in col])\n</code></pre>\n<p>it prints every column that has the substring \"target\" as part of it, for example \"encoded_target\" would match, while your code</p>\n<pre><code>'target' in S.columns\n</code></pre>\n<p>would not work, as it only prints the column name if it is the exact text \"target\".</p>\n<p>I believe you either leak the target in the features, or if you train an ensemble, you might be training it on train + test and then evaluating on test.</p>\n<p>Good luck with finding the issue.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1908042,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-08-21T10:00:41.717000",
          "content": "<p>Exactly!! Thank you for making it clear</p>\n<p>I bet it contains somethink like \"target_diff\", \"target_mean\", etc or \"target_x\" after multiple merges</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1908084,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-21T10:56:45.957000",
          "content": "<p>Thank you for your responses. <br>\nTarget has not been chased at all to contain any column like these.<br>\nTrain and test data are also separate. </p>\n<p>Maybe the issue is multicollinear columns.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1908086,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-21T10:59:22.967000",
          "content": "<p>print([col for col in X_train.columns if 'target' in col])<br>\nOutput for this is an empty list.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1907537,
      "author_name": "Rasoul Mojtahedzadeh",
      "author_url": "",
      "post_date": "2022-08-20T21:36:45.243000",
      "content": "<p>Do you have target-encoding features in your training data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1907594,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-20T23:39:33.850000",
          "content": "<p>Hey,<br>\nno, haven't tried target-encoding.<br>\nBut have done one hot encoding.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1907324,
      "author_name": "kushal Agrawal",
      "author_url": "",
      "post_date": "2022-08-20T17:07:50.320000",
      "content": "<p>U might be having target info in your training data. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1907397,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-20T18:22:48.970000",
          "content": "<p>No, target data has been removed before splitting the dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1907558,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-08-20T22:28:26.490000",
          "content": "<p>Certainly you have target col (or derivative) in your features, search better.. I'd suggest to print out and check your features again</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1907595,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-20T23:40:51.737000",
          "content": "<p>target col has been removed from X_train and X_test. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1907661,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-08-21T00:46:04.887000",
          "content": "<p>I bet some target info is there! check each feature you have created one by one! <br>\nor use this print statement just before you pass it to fit<br>\n<code>print([col for col in X_train.columns if 'target' in col])</code></p>\n<p>PS: there is no \"real\" model that can score 99% :)  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1907669,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-21T00:53:29.473000",
          "content": "<p><strong>I ran this statement.</strong><br>\n'target' in S.columns</p>\n<p>Output: <strong>False</strong><br>\nS is my training data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1908141,
          "author_name": "ToriGlori",
          "author_url": "",
          "post_date": "2022-08-21T11:50:55.043000",
          "content": "<p>Maybe you have changed the name or something else? Search for the values in your dataframe. Maybe you have used the target column and aggregated it within a different feature etc. etc.</p>\n<p>There is no way to score this high and that low on public lb</p>\n<p>Good luck!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1907189,
      "author_name": "Jake",
      "author_url": "",
      "post_date": "2022-08-20T15:08:11.087000",
      "content": "<p>Looks like you have a leak in your training. Is the target in your training data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1907398,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-20T18:23:34.570000",
          "content": "<p>No, target is not part of X_train and X_test</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1907633,
          "author_name": "Fritz Cremer",
          "author_url": "",
          "post_date": "2022-08-21T00:13:44.700000",
          "content": "<p>You definitively have a leak somewhere in your pipeline. 0.99 is impossible to reach on the test set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1907639,
          "author_name": "Neha",
          "author_url": "",
          "post_date": "2022-08-21T00:24:06.733000",
          "content": "<p>Trying to figure it out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1907671,
      "author_name": "Neha",
      "author_url": "",
      "post_date": "2022-08-21T00:54:48.403000",
      "content": "<p>I ran this statement.<br>\n'target' in S.columns</p>\n<p>Output: False<br>\nS is my training data.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1908137,
          "author_name": "LoneMiner",
          "author_url": "",
          "post_date": "2022-08-21T11:48:07.510000",
          "content": "<p>Dump out the column names in your dataset to a CSV dataset and go through them one by one.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1907906,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-21T07:16:42.650000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1907886": "Just look into your model's feature importance, you can use Shapley summary (beeswarm) plot for example https://shap.readthedocs.io/en/latest/example_notebooks/api_examples/plots/beeswarm.html to identify most important features, my guess is the most important one, or first top N, contain the leak.\n\nOther option is to try to build single feature models and see if any model is much better than it should be.\n\nAlso when @imeintanis suggests to have a look at feature names:\n```\nprint([col for col in X_train.columns if 'target' in col])\n```\nit prints every column that has the substring \"target\" as part of it, for example \"encoded_target\" would match, while your code\n```\n'target' in S.columns\n```\nwould not work, as it only prints the column name if it is the exact text \"target\".\n\nI believe you either leak the target in the features, or if you train an ensemble, you might be training it on train + test and then evaluating on test.\n\nGood luck with finding the issue.",
    "1907537": "Do you have target-encoding features in your training data?",
    "1907324": "U might be having target info in your training data. ",
    "1907189": "Looks like you have a leak in your training. Is the target in your training data?",
    "1907671": "I ran this statement.\n'target' in S.columns\n\nOutput: False\nS is my training data.",
    "1907177": "Amex Metric Score is coming to be 0.99145420 (test data) and 0.999958889 (training data) but the competition public leaderboard score shows 0.57.\n\nCan someone help to understand the possible reason behind this?",
    "1907906": ""
  }
}