{
  "id": 359865,
  "title": "Joining train & test inputs data before applying PCA?",
  "url": "/competitions/open-problems-multimodal/discussion/359865",
  "author_name": "",
  "post_date": "2022-10-14T02:16:48.009070800Z",
  "votes": 10,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey I have a question. Would  joining both train &amp; test  inputs sets and then applying PCA lead to overfitting? Even though we're not giving our models any information about the test targets, we're still giving them information about the test inputs by using them to apply PCA. </p>",
  "messages": [
    {
      "id": "1986300",
      "postDate": "10/14/2022 02:16:48",
      "content": "<p>Hey I have a question. Would  joining both train &amp; test  inputs sets and then applying PCA lead to overfitting? Even though we're not giving our models any information about the test targets, we're still giving them information about the test inputs by using them to apply PCA. </p>",
      "rawMarkdown": "Hey I have a question. Would  joining both train & test  inputs sets and then applying PCA lead to overfitting? Even though we're not giving our models any information about the test targets, we're still giving them information about the test inputs by using them to apply PCA.",
      "votes": null
    },
    {
      "id": "1986714",
      "postDate": "10/14/2022 08:24:40",
      "content": "<p>That my friend would be Data Leakage, it is important to test a predictor on data held-out from training, preprocessing (such as standardization, feature selection, etc.) and similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction.</p>",
      "rawMarkdown": "That my friend would be Data Leakage, it is important to test a predictor on data held-out from training, preprocessing (such as standardization, feature selection, etc.) and similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction.",
      "votes": null
    },
    {
      "id": "1986741",
      "postDate": "10/14/2022 08:43:58",
      "content": "<p>It is data leakage which you SHOULD DO on Kaggle to improve score, when competition is NOT \"kernel only\".<br>\nPS<br>\nThat question has been discussed many times.</p>",
      "rawMarkdown": "It is data leakage which you SHOULD DO on Kaggle to improve score, when competition is NOT \"kernel only\".\nPS\nThat question has been discussed many times.",
      "votes": null
    },
    {
      "id": "1987074",
      "postDate": "10/14/2022 13:42:33",
      "content": "<p>Excellent.Thank you for your analysis and time</p>",
      "rawMarkdown": "Excellent.Thank you for your analysis and time",
      "votes": null
    },
    {
      "id": "1988066",
      "postDate": "10/15/2022 04:09:25",
      "content": "<p>If you have the test data available when you train a model, there is no reason not to use that information when training your models. Even when a competition is \"kernel only\", if you can train your models within the time limit, it is wise to use the test data if you can. </p>\n<p>If you want to be extreme and not use the test data to train models, then you should also not use the validation data when doing pca either. </p>",
      "rawMarkdown": "If you have the test data available when you train a model, there is no reason not to use that information when training your models. Even when a competition is \"kernel only\", if you can train your models within the time limit, it is wise to use the test data if you can. \n\nIf you want to be extreme and not use the test data to train models, then you should also not use the validation data when doing pca either.",
      "votes": null
    },
    {
      "id": "1988097",
      "postDate": "10/15/2022 04:44:08",
      "content": "<p>Nice ! <br>\nBut there is some detail about \"kernel only\" - the test set would be different - thus using oringal test set for PCA or like that - during your crossvalidation and choice of the best params for the model is not a good idea, which might lead to shake-up.<br>\nJust because - private  test would be different, but model was fitted for the other test.<br>\nMore safe is not to use it, thus gaining more generalazibility for the model and less rely on the public test which will disappear.</p>\n<p>The same is when one thinks about data science in real products - one should not use test for anything otherwise great chance for overfit. It is data leakage, which might be explored on Kaggle, but do not do such things in production.</p>\n<p>PS </p>\n<p>Again that question can be asked for any competition, not only that in particular,<br>\n and has been discussed many times. </p>",
      "rawMarkdown": "Nice ! \nBut there is some detail about \"kernel only\" - the test set would be different - thus using oringal test set for PCA or like that - during your crossvalidation and choice of the best params for the model is not a good idea, which might lead to shake-up.\nJust because - private  test would be different, but model was fitted for the other test.\nMore safe is not to use it, thus gaining more generalazibility for the model and less rely on the public test which will disappear.\n\nThe same is when one thinks about data science in real products - one should not use test for anything otherwise great chance for overfit. It is data leakage, which might be explored on Kaggle, but do not do such things in production.\n\nPS \n\nAgain that question can be asked for any competition, not only that in particular,\n and has been discussed many times.",
      "votes": null
    },
    {
      "id": "1988106",
      "postDate": "10/15/2022 04:57:41",
      "content": "<p>I agree with what you are saying if I meant the public test data, but when I say \"kernel only\", I am assuming you have access to all the private test inputs in the submission notebook. You won't be able to \"see it\" in real time or react to it, but you could use it for feature reduction. I believe this was true in the MOA competition and others. In that case, you could run your entire training pipeline from scratch in your submission notebook, using test set inputs when doing PCA.</p>\n<p>Even in production, if you could see the test inputs and have time to train models, wouldn't you want to use this information if you could? Otherwise, I agree it is overfitting. A proper validation strategy not include using the validation inputs in a way that you can't use the test data (production data).</p>",
      "rawMarkdown": "I agree with what you are saying if I meant the public test data, but when I say \"kernel only\", I am assuming you have access to all the private test inputs in the submission notebook. You won't be able to \"see it\" in real time or react to it, but you could use it for feature reduction. I believe this was true in the MOA competition and others. In that case, you could run your entire training pipeline from scratch in your submission notebook, using test set inputs when doing PCA.\n\nEven in production, if you could see the test inputs and have time to train models, wouldn't you want to use this information if you could? Otherwise, I agree it is overfitting. A proper validation strategy not include using the validation inputs in a way that you can't use the test data (production data).",
      "votes": null
    },
    {
      "id": "1988112",
      "postDate": "10/15/2022 05:04:04",
      "content": "<p>You are right ! <br>\nSorry for misunderstanding.<br>\nI meant , typically you do not have enough time to react on the new test data.<br>\nIf you have - then it is Okay !<br>\nThe model training is typically long and inference should be fast.  <br>\nThus time consuming  selection of the parameters for the model via cross-validation , typically cannot be done when new test data appear, thus we need more rely on generalizability of the model and thus train it for that, thus not including test to PCA.</p>",
      "rawMarkdown": "You are right ! \nSorry for misunderstanding.\nI meant , typically you do not have enough time to react on the new test data.\nIf you have - then it is Okay !\nThe model training is typically long and inference should be fast.  \nThus time consuming  selection of the parameters for the model via cross-validation , typically cannot be done when new test data appear, thus we need more rely on generalizability of the model and thus train it for that, thus not including test to PCA.",
      "votes": null
    },
    {
      "id": "1988116",
      "postDate": "10/15/2022 05:08:37",
      "content": "<p>Absolutely, I agree 100%</p>",
      "rawMarkdown": "Absolutely, I agree 100%",
      "votes": null
    },
    {
      "id": "2029093",
      "postDate": "11/14/2022 12:41:19",
      "content": "<p>THANKS GREAT WORK</p>",
      "rawMarkdown": "THANKS GREAT WORK",
      "votes": null
    },
    {
      "id": "2029792",
      "postDate": "11/14/2022 23:27:21",
      "content": "<p>This is great! Thanks :)</p>",
      "rawMarkdown": "This is great! Thanks :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1986714,
      "author_name": "robertobonilla",
      "author_url": "",
      "post_date": "10/14/2022 08:24:40",
      "content": "<p>That my friend would be Data Leakage, it is important to test a predictor on data held-out from training, preprocessing (such as standardization, feature selection, etc.) and similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1986741,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "10/14/2022 08:43:58",
          "content": "<p>It is data leakage which you SHOULD DO on Kaggle to improve score, when competition is NOT \"kernel only\".<br>\nPS<br>\nThat question has been discussed many times.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1987074,
      "author_name": "farimahvatankhah",
      "author_url": "",
      "post_date": "10/14/2022 13:42:33",
      "content": "<p>Excellent.Thank you for your analysis and time</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1988066,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "10/15/2022 04:09:25",
      "content": "<p>If you have the test data available when you train a model, there is no reason not to use that information when training your models. Even when a competition is \"kernel only\", if you can train your models within the time limit, it is wise to use the test data if you can. </p>\n<p>If you want to be extreme and not use the test data to train models, then you should also not use the validation data when doing pca either. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1988097,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "10/15/2022 04:44:08",
          "content": "<p>Nice ! <br>\nBut there is some detail about \"kernel only\" - the test set would be different - thus using oringal test set for PCA or like that - during your crossvalidation and choice of the best params for the model is not a good idea, which might lead to shake-up.<br>\nJust because - private  test would be different, but model was fitted for the other test.<br>\nMore safe is not to use it, thus gaining more generalazibility for the model and less rely on the public test which will disappear.</p>\n<p>The same is when one thinks about data science in real products - one should not use test for anything otherwise great chance for overfit. It is data leakage, which might be explored on Kaggle, but do not do such things in production.</p>\n<p>PS </p>\n<p>Again that question can be asked for any competition, not only that in particular,<br>\n and has been discussed many times. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1988106,
          "author_name": "chrisrichardmiles",
          "author_url": "",
          "post_date": "10/15/2022 04:57:41",
          "content": "<p>I agree with what you are saying if I meant the public test data, but when I say \"kernel only\", I am assuming you have access to all the private test inputs in the submission notebook. You won't be able to \"see it\" in real time or react to it, but you could use it for feature reduction. I believe this was true in the MOA competition and others. In that case, you could run your entire training pipeline from scratch in your submission notebook, using test set inputs when doing PCA.</p>\n<p>Even in production, if you could see the test inputs and have time to train models, wouldn't you want to use this information if you could? Otherwise, I agree it is overfitting. A proper validation strategy not include using the validation inputs in a way that you can't use the test data (production data).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1988112,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "10/15/2022 05:04:04",
          "content": "<p>You are right ! <br>\nSorry for misunderstanding.<br>\nI meant , typically you do not have enough time to react on the new test data.<br>\nIf you have - then it is Okay !<br>\nThe model training is typically long and inference should be fast.  <br>\nThus time consuming  selection of the parameters for the model via cross-validation , typically cannot be done when new test data appear, thus we need more rely on generalizability of the model and thus train it for that, thus not including test to PCA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1988116,
          "author_name": "chrisrichardmiles",
          "author_url": "",
          "post_date": "10/15/2022 05:08:37",
          "content": "<p>Absolutely, I agree 100%</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2029093,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2022 12:41:19",
      "content": "<p>THANKS GREAT WORK</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2029792,
      "author_name": "",
      "author_url": "",
      "post_date": "11/14/2022 23:27:21",
      "content": "<p>This is great! Thanks :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1986300": "Hey I have a question. Would  joining both train & test  inputs sets and then applying PCA lead to overfitting? Even though we're not giving our models any information about the test targets, we're still giving them information about the test inputs by using them to apply PCA.",
    "1986714": "That my friend would be Data Leakage, it is important to test a predictor on data held-out from training, preprocessing (such as standardization, feature selection, etc.) and similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction.",
    "1986741": "It is data leakage which you SHOULD DO on Kaggle to improve score, when competition is NOT \"kernel only\".\nPS\nThat question has been discussed many times.",
    "1987074": "Excellent.Thank you for your analysis and time",
    "1988066": "If you have the test data available when you train a model, there is no reason not to use that information when training your models. Even when a competition is \"kernel only\", if you can train your models within the time limit, it is wise to use the test data if you can. \n\nIf you want to be extreme and not use the test data to train models, then you should also not use the validation data when doing pca either.",
    "1988097": "Nice ! \nBut there is some detail about \"kernel only\" - the test set would be different - thus using oringal test set for PCA or like that - during your crossvalidation and choice of the best params for the model is not a good idea, which might lead to shake-up.\nJust because - private  test would be different, but model was fitted for the other test.\nMore safe is not to use it, thus gaining more generalazibility for the model and less rely on the public test which will disappear.\n\nThe same is when one thinks about data science in real products - one should not use test for anything otherwise great chance for overfit. It is data leakage, which might be explored on Kaggle, but do not do such things in production.\n\nPS \n\nAgain that question can be asked for any competition, not only that in particular,\n and has been discussed many times.",
    "1988106": "I agree with what you are saying if I meant the public test data, but when I say \"kernel only\", I am assuming you have access to all the private test inputs in the submission notebook. You won't be able to \"see it\" in real time or react to it, but you could use it for feature reduction. I believe this was true in the MOA competition and others. In that case, you could run your entire training pipeline from scratch in your submission notebook, using test set inputs when doing PCA.\n\nEven in production, if you could see the test inputs and have time to train models, wouldn't you want to use this information if you could? Otherwise, I agree it is overfitting. A proper validation strategy not include using the validation inputs in a way that you can't use the test data (production data).",
    "1988112": "You are right ! \nSorry for misunderstanding.\nI meant , typically you do not have enough time to react on the new test data.\nIf you have - then it is Okay !\nThe model training is typically long and inference should be fast.  \nThus time consuming  selection of the parameters for the model via cross-validation , typically cannot be done when new test data appear, thus we need more rely on generalizability of the model and thus train it for that, thus not including test to PCA.",
    "1988116": "Absolutely, I agree 100%",
    "2029093": "THANKS GREAT WORK",
    "2029792": "This is great! Thanks :)"
  },
  "source": "meta"
}