{
  "id": 349408,
  "title": "Does PCA with all training data cause leaks?",
  "url": "/competitions/open-problems-multimodal/discussion/349408",
  "author_name": "",
  "post_date": "2022-09-01T09:15:33.323614100Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In creating a model, I need to create my model based on appropriate validation data to avoid overfitting to LB.<br>\nIn the cite seq data,<br>\ntrain data contains donor{13176, 31800, 32606} and<br>\ntest data contains donor{13176, <strong>27678</strong>, 31800, 32606}.<br>\nJudging from this, I think it is better to do group kfold based on donor.<br>\nOne thing I am wondering is that if we do PCA based on all train data, it may cause leaks.</p>\n<p>What kind of validation data should be created?<br>\nShould I do a separate PCA for each fold?</p>",
  "messages": [
    {
      "id": "1922105",
      "postDate": "09/01/2022 09:15:33",
      "content": "<p>In creating a model, I need to create my model based on appropriate validation data to avoid overfitting to LB.<br>\nIn the cite seq data,<br>\ntrain data contains donor{13176, 31800, 32606} and<br>\ntest data contains donor{13176, <strong>27678</strong>, 31800, 32606}.<br>\nJudging from this, I think it is better to do group kfold based on donor.<br>\nOne thing I am wondering is that if we do PCA based on all train data, it may cause leaks.</p>\n<p>What kind of validation data should be created?<br>\nShould I do a separate PCA for each fold?</p>",
      "rawMarkdown": "In creating a model, I need to create my model based on appropriate validation data to avoid overfitting to LB.\nIn the cite seq data,\ntrain data contains donor{13176, 31800, 32606} and\ntest data contains donor{13176, **27678**, 31800, 32606}.\nJudging from this, I think it is better to do group kfold based on donor.\nOne thing I am wondering is that if we do PCA based on all train data, it may cause leaks.\n\nWhat kind of validation data should be created?\nShould I do a separate PCA for each fold?",
      "votes": null
    },
    {
      "id": "1922990",
      "postDate": "09/01/2022 22:27:53",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a> , </p>\n<p>Firstly, let's call X is your whole training dataset and X_k is data for each fold in K-fold. </p>\n<ol>\n<li>Why do we apply PCA:</li>\n</ol>\n<ul>\n<li>Reduce dimensionality of dataset.</li>\n<li>It is reasonable that  when applying PCA to your data, you can lose some information of your data. However, the lost information is may not necessary or noise some time. There are a variety of applications using PCA as a method for abnormal detection where we consider abnormal data as outliers and noise. <br>\n-&gt; From points above, using PCA may \"leak\"  some information but it can help traning model faster and make model more \"robust\" to noise.<br>\nI have done a notebook using PCA for abnormal detection, you can check it out for its performance. It is still better than some supervised existing learning approaches. <br>\n<a href=\"https://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection\" target=\"_blank\">https://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection</a> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174022%2Fa934d0d00280049237bd83c02df5c42c%2FCapture.PNG?generation=1662071110034461&amp;alt=media\" alt=\"\"></li>\n</ul>\n<ol>\n<li>Should we use in X or X-k?<br>\nBecause PCA try to learn the principal components from a dataset, then data will be projected on to main components with smallest variance. <br>\n-&gt; We should use PCA for X because X have more data. The more data, the more accurate PCA is. </li>\n</ol>",
      "rawMarkdown": "Hello @hiroakifukuse , \n\nFirstly, let's call X is your whole training dataset and X_k is data for each fold in K-fold. \n\n1. Why do we apply PCA:\n- Reduce dimensionality of dataset.\n- It is reasonable that  when applying PCA to your data, you can lose some information of your data. However, the lost information is may not necessary or noise some time. There are a variety of applications using PCA as a method for abnormal detection where we consider abnormal data as outliers and noise. \n-> From points above, using PCA may \"leak\"  some information but it can help traning model faster and make model more \"robust\" to noise.\nI have done a notebook using PCA for abnormal detection, you can check it out for its performance. It is still better than some supervised existing learning approaches. \nhttps://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174022%2Fa934d0d00280049237bd83c02df5c42c%2FCapture.PNG?generation=1662071110034461&alt=media)\n\n2. Should we use in X or X-k?\nBecause PCA try to learn the principal components from a dataset, then data will be projected on to main components with smallest variance. \n-> We should use PCA for X because X have more data. The more data, the more accurate PCA is.",
      "votes": null
    },
    {
      "id": "1923003",
      "postDate": "09/01/2022 22:51:22",
      "content": "<p>It doesn't cause target leak. Since PCA is an unsupervised task, using the full train/test set to do PCA doesn't encode any additional feature information into your training set.</p>",
      "rawMarkdown": "It doesn't cause target leak. Since PCA is an unsupervised task, using the full train/test set to do PCA doesn't encode any additional feature information into your training set.",
      "votes": null
    },
    {
      "id": "1923117",
      "postDate": "09/02/2022 02:36:30",
      "content": "<p>Thanks for the comment.<br>\nI agree that target leak do not happen.<br>\nUsually when we normalize data, we learn how to normalize on train data and then normalize the test data in the same way as the train data.<br>\n・example</p>\n<pre><code>X_train_minmax = min_max_scaler.fit_transform(X_train)\nX_test_minmax = min_max_scaler.transform(X_test)\n</code></pre>\n<p>cited:  <a href=\"https://scikit-learn.org/stable/modules/preprocessing.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/preprocessing.html</a></p>\n<p>I think this is so that the distribution of the test data does not affect the train data.<br>\nI was wondering if the same thing applies to PCA.</p>",
      "rawMarkdown": "Thanks for the comment.\nI agree that target leak do not happen.\nUsually when we normalize data, we learn how to normalize on train data and then normalize the test data in the same way as the train data.\n・example\n```\nX_train_minmax = min_max_scaler.fit_transform(X_train)\nX_test_minmax = min_max_scaler.transform(X_test)\n```\ncited:  https://scikit-learn.org/stable/modules/preprocessing.html\n\nI think this is so that the distribution of the test data does not affect the train data.\nI was wondering if the same thing applies to PCA.",
      "votes": null
    },
    {
      "id": "1923126",
      "postDate": "09/02/2022 02:40:07",
      "content": "<p>Thanks for the comment.<br>\nIt helped me understand.</p>",
      "rawMarkdown": "Thanks for the comment.\nIt helped me understand.",
      "votes": null
    },
    {
      "id": "1923967",
      "postDate": "09/02/2022 16:31:59",
      "content": "<p>Do you mean to make PCA for train+test togather? <br>\nIf so - it seems to me that question is discussed many times,<br>\nit is not exactly dateleak, but it is Kaggle trick which might improve score, but fail on \"production\" or \"kernel-only\" competitions.</p>",
      "rawMarkdown": "Do you mean to make PCA for train+test togather? \nIf so - it seems to me that question is discussed many times,\nit is not exactly dateleak, but it is Kaggle trick which might improve score, but fail on \"production\" or \"kernel-only\" competitions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1922990,
      "author_name": "leokaka",
      "author_url": "",
      "post_date": "09/01/2022 22:27:53",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/hiroakifukuse\" target=\"_blank\">@hiroakifukuse</a> , </p>\n<p>Firstly, let's call X is your whole training dataset and X_k is data for each fold in K-fold. </p>\n<ol>\n<li>Why do we apply PCA:</li>\n</ol>\n<ul>\n<li>Reduce dimensionality of dataset.</li>\n<li>It is reasonable that  when applying PCA to your data, you can lose some information of your data. However, the lost information is may not necessary or noise some time. There are a variety of applications using PCA as a method for abnormal detection where we consider abnormal data as outliers and noise. <br>\n-&gt; From points above, using PCA may \"leak\"  some information but it can help traning model faster and make model more \"robust\" to noise.<br>\nI have done a notebook using PCA for abnormal detection, you can check it out for its performance. It is still better than some supervised existing learning approaches. <br>\n<a href=\"https://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection\" target=\"_blank\">https://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection</a> <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174022%2Fa934d0d00280049237bd83c02df5c42c%2FCapture.PNG?generation=1662071110034461&amp;alt=media\" alt=\"\"></li>\n</ul>\n<ol>\n<li>Should we use in X or X-k?<br>\nBecause PCA try to learn the principal components from a dataset, then data will be projected on to main components with smallest variance. <br>\n-&gt; We should use PCA for X because X have more data. The more data, the more accurate PCA is. </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1923126,
          "author_name": "hiroakifukuse",
          "author_url": "",
          "post_date": "09/02/2022 02:40:07",
          "content": "<p>Thanks for the comment.<br>\nIt helped me understand.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1923003,
      "author_name": "sebastianvangerwen",
      "author_url": "",
      "post_date": "09/01/2022 22:51:22",
      "content": "<p>It doesn't cause target leak. Since PCA is an unsupervised task, using the full train/test set to do PCA doesn't encode any additional feature information into your training set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1923117,
          "author_name": "hiroakifukuse",
          "author_url": "",
          "post_date": "09/02/2022 02:36:30",
          "content": "<p>Thanks for the comment.<br>\nI agree that target leak do not happen.<br>\nUsually when we normalize data, we learn how to normalize on train data and then normalize the test data in the same way as the train data.<br>\n・example</p>\n<pre><code>X_train_minmax = min_max_scaler.fit_transform(X_train)\nX_test_minmax = min_max_scaler.transform(X_test)\n</code></pre>\n<p>cited:  <a href=\"https://scikit-learn.org/stable/modules/preprocessing.html\" target=\"_blank\">https://scikit-learn.org/stable/modules/preprocessing.html</a></p>\n<p>I think this is so that the distribution of the test data does not affect the train data.<br>\nI was wondering if the same thing applies to PCA.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1923967,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/02/2022 16:31:59",
      "content": "<p>Do you mean to make PCA for train+test togather? <br>\nIf so - it seems to me that question is discussed many times,<br>\nit is not exactly dateleak, but it is Kaggle trick which might improve score, but fail on \"production\" or \"kernel-only\" competitions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1922105": "In creating a model, I need to create my model based on appropriate validation data to avoid overfitting to LB.\nIn the cite seq data,\ntrain data contains donor{13176, 31800, 32606} and\ntest data contains donor{13176, **27678**, 31800, 32606}.\nJudging from this, I think it is better to do group kfold based on donor.\nOne thing I am wondering is that if we do PCA based on all train data, it may cause leaks.\n\nWhat kind of validation data should be created?\nShould I do a separate PCA for each fold?",
    "1922990": "Hello @hiroakifukuse , \n\nFirstly, let's call X is your whole training dataset and X_k is data for each fold in K-fold. \n\n1. Why do we apply PCA:\n- Reduce dimensionality of dataset.\n- It is reasonable that  when applying PCA to your data, you can lose some information of your data. However, the lost information is may not necessary or noise some time. There are a variety of applications using PCA as a method for abnormal detection where we consider abnormal data as outliers and noise. \n-> From points above, using PCA may \"leak\"  some information but it can help traning model faster and make model more \"robust\" to noise.\nI have done a notebook using PCA for abnormal detection, you can check it out for its performance. It is still better than some supervised existing learning approaches. \nhttps://www.kaggle.com/code/leokaka/pca-90-accuracy-for-attack-detection \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8174022%2Fa934d0d00280049237bd83c02df5c42c%2FCapture.PNG?generation=1662071110034461&alt=media)\n\n2. Should we use in X or X-k?\nBecause PCA try to learn the principal components from a dataset, then data will be projected on to main components with smallest variance. \n-> We should use PCA for X because X have more data. The more data, the more accurate PCA is.",
    "1923003": "It doesn't cause target leak. Since PCA is an unsupervised task, using the full train/test set to do PCA doesn't encode any additional feature information into your training set.",
    "1923117": "Thanks for the comment.\nI agree that target leak do not happen.\nUsually when we normalize data, we learn how to normalize on train data and then normalize the test data in the same way as the train data.\n・example\n```\nX_train_minmax = min_max_scaler.fit_transform(X_train)\nX_test_minmax = min_max_scaler.transform(X_test)\n```\ncited:  https://scikit-learn.org/stable/modules/preprocessing.html\n\nI think this is so that the distribution of the test data does not affect the train data.\nI was wondering if the same thing applies to PCA.",
    "1923126": "Thanks for the comment.\nIt helped me understand.",
    "1923967": "Do you mean to make PCA for train+test togather? \nIf so - it seems to me that question is discussed many times,\nit is not exactly dateleak, but it is Kaggle trick which might improve score, but fail on \"production\" or \"kernel-only\" competitions."
  },
  "source": "meta"
}