{
  "id": 536496,
  "title": "Data leakage when handling PCIAT_Total",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536496",
  "author_name": "",
  "post_date": "2024-09-27T19:37:20.145737Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Say I train a 1st model using train data <strong>with</strong> PCIAT_Total to fill up <strong>blank</strong> PCIAT_Total in train data, LGBM.<br>\nAnd then I use all this data to train a 2nd LGBM model to predict PCIAT_Total in test data, and convert them into sii 0-3 and submit.<br>\nIs it a valid method or there is data leakage? </p>\n<p>Mean Train Cohen's kappa: 1.0000<br>\nMean Validation Cohen's kappa: 0.3040<br>\nPublic Score: 0.153 😂</p>",
  "messages": [
    {
      "id": "3000591",
      "postDate": "09/27/2024 19:37:20",
      "content": "<p>Say I train a 1st model using train data <strong>with</strong> PCIAT_Total to fill up <strong>blank</strong> PCIAT_Total in train data, LGBM.<br>\nAnd then I use all this data to train a 2nd LGBM model to predict PCIAT_Total in test data, and convert them into sii 0-3 and submit.<br>\nIs it a valid method or there is data leakage? </p>\n<p>Mean Train Cohen's kappa: 1.0000<br>\nMean Validation Cohen's kappa: 0.3040<br>\nPublic Score: 0.153 😂</p>",
      "rawMarkdown": "Say I train a 1st model using train data **with** PCIAT_Total to fill up **blank** PCIAT_Total in train data, LGBM.\nAnd then I use all this data to train a 2nd LGBM model to predict PCIAT_Total in test data, and convert them into sii 0-3 and submit.\nIs it a valid method or there is data leakage? \n\nMean Train Cohen's kappa: 1.0000\nMean Validation Cohen's kappa: 0.3040\nPublic Score: 0.153 😂",
      "votes": null
    },
    {
      "id": "3000595",
      "postDate": "09/27/2024 19:44:17",
      "content": "<p><code>Mean Train Cohen's kappa: 1.0000</code> - this is the main issue <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> <br>\nI am sure your target column is present in the train data (features)</p>",
      "rawMarkdown": "`Mean Train Cohen's kappa: 1.0000` - this is the main issue @tomyuen \nI am sure your target column is present in the train data (features)",
      "votes": null
    },
    {
      "id": "3000730",
      "postDate": "09/28/2024 02:20:16",
      "content": "<p>The target column ('sii') is derived from the total PCIAT questionnaire scores, which is causing the leakage. </p>",
      "rawMarkdown": "The target column ('sii') is derived from the total PCIAT questionnaire scores, which is causing the leakage.",
      "votes": null
    },
    {
      "id": "3000867",
      "postDate": "09/28/2024 07:05:35",
      "content": "<p>Agreed, just ensure none of the <code>PCIAT and sii columns</code> are present in the train dataset while training models <a href=\"https://www.kaggle.com/ethanp345\" target=\"_blank\">@ethanp345</a> <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> </p>",
      "rawMarkdown": "Agreed, just ensure none of the `PCIAT and sii columns` are present in the train dataset while training models @ethanp345 @tomyuen",
      "votes": null
    },
    {
      "id": "3001469",
      "postDate": "09/28/2024 21:09:29",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <a href=\"https://www.kaggle.com/ethanp345\" target=\"_blank\">@ethanp345</a> <br>\nThanks both! I've tried it another day but still in vain. <br>\nI dropped sii and all PCIAT column (except Total) once I load the train_df and I've used the standard X Y split<br>\n<code>X = train.drop(['PCIAT-PCIAT_Total'], axis=1)</code><br>\n<code>y = train['PCIAT-PCIAT_Total']</code></p>\n<p>I guess it's the 1st step, when I trained a LGBM to fill in missed PCIAT_Total in train_df, that there're too many missed value but too few valid sample, that the 1st model memorize the pattern and filled in some toxic PCIAT_Total prediction.<br>\nSo when I trained the 2nd LGBM, the train score is so high, but of course it doesn't generalize to unseen data.</p>\n<p>Most likely I'll drop this idea of 'filling blank PCIAT_Total in train_df' at the moment 😅, and reference how others are doing.<br>\nThanks again.</p>",
      "rawMarkdown": "ravi20076 @ethanp345 \nThanks both! I've tried it another day but still in vain. \nI dropped sii and all PCIAT column (except Total) once I load the train_df and I've used the standard X Y split\n`    X = train.drop(['PCIAT-PCIAT_Total'], axis=1)`\n`    y = train['PCIAT-PCIAT_Total']`\n\nI guess it's the 1st step, when I trained a LGBM to fill in missed PCIAT_Total in train_df, that there're too many missed value but too few valid sample, that the 1st model memorize the pattern and filled in some toxic PCIAT_Total prediction.\nSo when I trained the 2nd LGBM, the train score is so high, but of course it doesn't generalize to unseen data.\n\nMost likely I'll drop this idea of 'filling blank PCIAT_Total in train_df' at the moment 😅, and reference how others are doing.\nThanks again.",
      "votes": null
    },
    {
      "id": "3001485",
      "postDate": "09/28/2024 21:39:08",
      "content": "<p>You need to drop all PCIAT columns (inclusive of the total column), since that column specifically determines our target ('sii'). </p>\n<pre><code>target = \n\ntrain_cols = (train.)\ntest_cols = (test.)\n\n\ncols_to_remove = train_cols - test_cols \n\n\ncols_to_remove(target)\n\n\ntrain = train( = cols_to_remove)\n</code></pre>\n<p>Now feel free to split to X,y and begin modeling. </p>\n<p>Hope this helps! </p>",
      "rawMarkdown": "You need to drop all PCIAT columns (inclusive of the total column), since that column specifically determines our target ('sii'). \n\n```\ntarget = 'sii'\n\ntrain_cols = set(train.columns)\ntest_cols = set(test.columns)\n\n'Finding columns that need to be removed'\ncols_to_remove = train_cols - test_cols \n\n'Removing our target variable from this list'\ncols_to_remove.remove(target)\n\n'Removing columns'\ntrain = train.drop(columns = cols_to_remove)\n\n```\n\nNow feel free to split to X,y and begin modeling. \n\nHope this helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3000595,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "09/27/2024 19:44:17",
      "content": "<p><code>Mean Train Cohen's kappa: 1.0000</code> - this is the main issue <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> <br>\nI am sure your target column is present in the train data (features)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3000730,
      "author_name": "ethanp345",
      "author_url": "",
      "post_date": "09/28/2024 02:20:16",
      "content": "<p>The target column ('sii') is derived from the total PCIAT questionnaire scores, which is causing the leakage. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3000867,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "09/28/2024 07:05:35",
          "content": "<p>Agreed, just ensure none of the <code>PCIAT and sii columns</code> are present in the train dataset while training models <a href=\"https://www.kaggle.com/ethanp345\" target=\"_blank\">@ethanp345</a> <a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 3001469,
              "author_name": "tomyuen",
              "author_url": "",
              "post_date": "09/28/2024 21:09:29",
              "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <a href=\"https://www.kaggle.com/ethanp345\" target=\"_blank\">@ethanp345</a> <br>\nThanks both! I've tried it another day but still in vain. <br>\nI dropped sii and all PCIAT column (except Total) once I load the train_df and I've used the standard X Y split<br>\n<code>X = train.drop(['PCIAT-PCIAT_Total'], axis=1)</code><br>\n<code>y = train['PCIAT-PCIAT_Total']</code></p>\n<p>I guess it's the 1st step, when I trained a LGBM to fill in missed PCIAT_Total in train_df, that there're too many missed value but too few valid sample, that the 1st model memorize the pattern and filled in some toxic PCIAT_Total prediction.<br>\nSo when I trained the 2nd LGBM, the train score is so high, but of course it doesn't generalize to unseen data.</p>\n<p>Most likely I'll drop this idea of 'filling blank PCIAT_Total in train_df' at the moment 😅, and reference how others are doing.<br>\nThanks again.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3001485,
                  "author_name": "ethanp345",
                  "author_url": "",
                  "post_date": "09/28/2024 21:39:08",
                  "content": "<p>You need to drop all PCIAT columns (inclusive of the total column), since that column specifically determines our target ('sii'). </p>\n<pre><code>target = \n\ntrain_cols = (train.)\ntest_cols = (test.)\n\n\ncols_to_remove = train_cols - test_cols \n\n\ncols_to_remove(target)\n\n\ntrain = train( = cols_to_remove)\n</code></pre>\n<p>Now feel free to split to X,y and begin modeling. </p>\n<p>Hope this helps! </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3000591": "Say I train a 1st model using train data **with** PCIAT_Total to fill up **blank** PCIAT_Total in train data, LGBM.\nAnd then I use all this data to train a 2nd LGBM model to predict PCIAT_Total in test data, and convert them into sii 0-3 and submit.\nIs it a valid method or there is data leakage? \n\nMean Train Cohen's kappa: 1.0000\nMean Validation Cohen's kappa: 0.3040\nPublic Score: 0.153 😂",
    "3000595": "`Mean Train Cohen's kappa: 1.0000` - this is the main issue @tomyuen \nI am sure your target column is present in the train data (features)",
    "3000730": "The target column ('sii') is derived from the total PCIAT questionnaire scores, which is causing the leakage.",
    "3000867": "Agreed, just ensure none of the `PCIAT and sii columns` are present in the train dataset while training models @ethanp345 @tomyuen",
    "3001469": "ravi20076 @ethanp345 \nThanks both! I've tried it another day but still in vain. \nI dropped sii and all PCIAT column (except Total) once I load the train_df and I've used the standard X Y split\n`    X = train.drop(['PCIAT-PCIAT_Total'], axis=1)`\n`    y = train['PCIAT-PCIAT_Total']`\n\nI guess it's the 1st step, when I trained a LGBM to fill in missed PCIAT_Total in train_df, that there're too many missed value but too few valid sample, that the 1st model memorize the pattern and filled in some toxic PCIAT_Total prediction.\nSo when I trained the 2nd LGBM, the train score is so high, but of course it doesn't generalize to unseen data.\n\nMost likely I'll drop this idea of 'filling blank PCIAT_Total in train_df' at the moment 😅, and reference how others are doing.\nThanks again.",
    "3001485": "You need to drop all PCIAT columns (inclusive of the total column), since that column specifically determines our target ('sii'). \n\n```\ntarget = 'sii'\n\ntrain_cols = set(train.columns)\ntest_cols = set(test.columns)\n\n'Finding columns that need to be removed'\ncols_to_remove = train_cols - test_cols \n\n'Removing our target variable from this list'\ncols_to_remove.remove(target)\n\n'Removing columns'\ntrain = train.drop(columns = cols_to_remove)\n\n```\n\nNow feel free to split to X,y and begin modeling. \n\nHope this helps!"
  },
  "source": "meta"
}