{
  "id": 56550,
  "title": "How to handle FE with a train/val/test split?",
  "url": "/competitions/avito-demand-prediction/discussion/56550",
  "author_name": "",
  "post_date": "2018-05-11T02:14:44.942508300Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Sorry for the n00b question, I'm new to data science and kaggle :O</p>\n\n<p>So my approach for this competition has been to engineer features by setting up functions and then running those functions on train/val/test datasets separately. Now this strategy makes sense to me with features that are row independent (like day of the week, # of letters in the title) but I'm getting confused for what the established best practice is for handling features that are fit to an entire dataset. For instance if I want to do label encoding on city, region, and param1-3 like I saw in some of the kernels, from what I can think of, I have 3 options using scikitlearn's preprocessing.LabelEncoder() </p>\n\n<p><strong>1. Fit-transform on each of train/val/test separately:</strong> This works for a competition since the test set is finite, but I don't think this is what practicing data scientists do since you're trying to generalize a model on the train set, right? </p>\n\n<p><strong>2. Fit on train, then transform on train, on val, and on test:</strong> So this is the strategy that I feel like a practicing data scientist would use as you're fitting the feature engineering on train data, then applying that fit by transforming validation and test data; in other words you're generalizing the model on train data. The issue I run into with this is what if val or train have a category that isn't represented in train? Then do I just fill that in with an nan during feature generation?  </p>\n\n<p><strong>3. Fit on train+val+test, then transform on train, on val, and on test:</strong> This is the strategy I saw in SRK and sban's kernels use. This doesn't make sense to me though because you're fitting your model on val and test data in addition to train. But doesn't that leak val and test information into your train setup and cause the model to overfit? </p>\n\n<p>Would appreciate any suggestions or wisdom on this. I'm also trying to implement features using tf-idf with feature reduction, and I'm running into the same question. How should I be fitting the tfidf vectorizer and then transforming my datasets? I imagine it would follow the same workflow as with label encoding. </p>",
  "messages": [
    {
      "id": "327198",
      "postDate": "05/11/2018 02:14:44",
      "content": "<p>Sorry for the n00b question, I'm new to data science and kaggle :O</p>\n\n<p>So my approach for this competition has been to engineer features by setting up functions and then running those functions on train/val/test datasets separately. Now this strategy makes sense to me with features that are row independent (like day of the week, # of letters in the title) but I'm getting confused for what the established best practice is for handling features that are fit to an entire dataset. For instance if I want to do label encoding on city, region, and param1-3 like I saw in some of the kernels, from what I can think of, I have 3 options using scikitlearn's preprocessing.LabelEncoder() </p>\n\n<p><strong>1. Fit-transform on each of train/val/test separately:</strong> This works for a competition since the test set is finite, but I don't think this is what practicing data scientists do since you're trying to generalize a model on the train set, right? </p>\n\n<p><strong>2. Fit on train, then transform on train, on val, and on test:</strong> So this is the strategy that I feel like a practicing data scientist would use as you're fitting the feature engineering on train data, then applying that fit by transforming validation and test data; in other words you're generalizing the model on train data. The issue I run into with this is what if val or train have a category that isn't represented in train? Then do I just fill that in with an nan during feature generation?  </p>\n\n<p><strong>3. Fit on train+val+test, then transform on train, on val, and on test:</strong> This is the strategy I saw in SRK and sban's kernels use. This doesn't make sense to me though because you're fitting your model on val and test data in addition to train. But doesn't that leak val and test information into your train setup and cause the model to overfit? </p>\n\n<p>Would appreciate any suggestions or wisdom on this. I'm also trying to implement features using tf-idf with feature reduction, and I'm running into the same question. How should I be fitting the tfidf vectorizer and then transforming my datasets? I imagine it would follow the same workflow as with label encoding. </p>",
      "rawMarkdown": "Sorry for the n00b question, I'm new to data science and kaggle :O\n\nSo my approach for this competition has been to engineer features by setting up functions and then running those functions on train/val/test datasets separately. Now this strategy makes sense to me with features that are row independent (like day of the week, # of letters in the title) but I'm getting confused for what the established best practice is for handling features that are fit to an entire dataset. For instance if I want to do label encoding on city, region, and param1-3 like I saw in some of the kernels, from what I can think of, I have 3 options using scikitlearn's preprocessing.LabelEncoder() \n\n**1. Fit-transform on each of train/val/test separately:** This works for a competition since the test set is finite, but I don't think this is what practicing data scientists do since you're trying to generalize a model on the train set, right? \n\n**2. Fit on train, then transform on train, on val, and on test:** So this is the strategy that I feel like a practicing data scientist would use as you're fitting the feature engineering on train data, then applying that fit by transforming validation and test data; in other words you're generalizing the model on train data. The issue I run into with this is what if val or train have a category that isn't represented in train? Then do I just fill that in with an nan during feature generation?  \n\n**3. Fit on train+val+test, then transform on train, on val, and on test:** This is the strategy I saw in SRK and sban's kernels use. This doesn't make sense to me though because you're fitting your model on val and test data in addition to train. But doesn't that leak val and test information into your train setup and cause the model to overfit? \n\nWould appreciate any suggestions or wisdom on this. I'm also trying to implement features using tf-idf with feature reduction, and I'm running into the same question. How should I be fitting the tfidf vectorizer and then transforming my datasets? I imagine it would follow the same workflow as with label encoding.",
      "votes": null
    },
    {
      "id": "327211",
      "postDate": "05/11/2018 03:10:43",
      "content": "<p>I have not looked closely at this competition so cannot say anything specific about it. I can say in general, how you treat train - valid - test set varies according to your dataset, your problem and the features you want to generate. \nApproach 1 is actually recommended by many people, as it ensure no future information leaked to training set. \nApproach 2 should be used in your case (maybe?) About your problem with new category, you should see if the category information is really that necessary. If not, drop it entirely. If it's important, you should try to group them into bigger groups, like group 1, group 2, group 3, others... according to some criteria that you defines (can be according to the relation between the categorical variable and some other numerical variables)\nApproach 3 is still ok if your data points are not related, or your features are not \"aggregated features\" like sum, mean, median etc. However if your data points are related, for example time-series data, then you should be careful.</p>",
      "rawMarkdown": "I have not looked closely at this competition so cannot say anything specific about it. I can say in general, how you treat train - valid - test set varies according to your dataset, your problem and the features you want to generate. \nApproach 1 is actually recommended by many people, as it ensure no future information leaked to training set. \nApproach 2 should be used in your case (maybe?) About your problem with new category, you should see if the category information is really that necessary. If not, drop it entirely. If it's important, you should try to group them into bigger groups, like group 1, group 2, group 3, others... according to some criteria that you defines (can be according to the relation between the categorical variable and some other numerical variables)\nApproach 3 is still ok if your data points are not related, or your features are not \"aggregated features\" like sum, mean, median etc. However if your data points are related, for example time-series data, then you should be careful.",
      "votes": null
    },
    {
      "id": "327212",
      "postDate": "05/11/2018 03:11:59",
      "content": "<p>Deciding between 2 and 3 is a matter of debate. I don't know the right answer. I lean toward 2, but I've seen people do 3 and be fine. ...But definitely don't do 1.</p>",
      "rawMarkdown": "Deciding between 2 and 3 is a matter of debate. I don't know the right answer. I lean toward 2, but I've seen people do 3 and be fine. ...But definitely don't do 1.",
      "votes": null
    },
    {
      "id": "327687",
      "postDate": "05/12/2018 07:16:07",
      "content": "<p>Many things/tricks work in Kaggle only, try that in real life and you should probably get fired :D <br>\nDo whatever you want, as long as you a) or b): <br>\na) Work as if the test set never existed <br>\nb) You will make sure that real life inputs will be enforced to a given standard.</p>",
      "rawMarkdown": "Many things/tricks work in Kaggle only, try that in real life and you should probably get fired :D  \nDo whatever you want, as long as you a) or b):  \na) Work as if the test set never existed  \nb) You will make sure that real life inputs will be enforced to a given standard.",
      "votes": null
    },
    {
      "id": "327760",
      "postDate": "05/12/2018 11:57:51",
      "content": "<p>I don't advise you to use label encorder for your categorical variables, ask about target encoding (there are lots of topics on the forum).\nYour score on the leaderboard will be better.</p>\n\n<p>To make label encoder I advise you the third option</p>",
      "rawMarkdown": "I don't advise you to use label encorder for your categorical variables, ask about target encoding (there are lots of topics on the forum).\nYour score on the leaderboard will be better.\n\nTo make label encoder I advise you the third option",
      "votes": null
    },
    {
      "id": "327812",
      "postDate": "05/12/2018 14:47:08",
      "content": "<p>Hi Ed! Nice to see you here :)</p>\n\n<p>You'll probably want to treat different types of transformations slightly differently, but the overarching principle should be that the <strong>format of the information must be the same at train time and prediction time</strong>. This means that the features should be the same, the category labels should be the same, the scale of numerical features should be the same, etc. Otherwise,  your model will either just fail to run or will misinterpret the data it sees at prediction time.</p>\n\n<p>So for example, if you fit one LabelEncoder on train data and another on validation data, it's very unlikely that the labels will be consistent between the two. At prediction time, your model will see a category and interpret it as a different category from what it actually is, bad news. So that's why it's convenient and safe to fit the LabelEncoder on all of the data you'll be using - the other option would be to fit one label encoder on train, and use that same encoder to transform on validation/test (handling categories missing in train is a separate issue from this). One thing that hasn't been mentioned yet is that LabelEncoder isn't actually engineering anything, it's just choosing arbitrary category labels so there's no need to worry about any leakage.</p>\n\n<p>What about for tf-idf? This is trickier. So first, if you fit tf-idf separately on train / validation / test, you'll just get features that don't align. Say you get 1000 word columns in train, you might get 600 in validation and then your model will literally just not work because it won't accept the feature format. Here again you can fit tf-idf on train and then use the same tf-idf model to transform validation/test. If you do that, you'll end up just ignoring words that occur in val/test but not train, which is reasonable since it's harder to get useful information about those words. But a workaround could be to actually fit tf-idf on all of the data then run PCA/SVD (get an LSA model) to link together words that occur only in train and val/test. For example, let's say puppy and dog co-occur in train, dog and cat co-occur in test. When you run SVD, you can extract similarity between puppy and dog, dog and cat and then end up linking together puppy and cat by projecting them onto similar dimensions.</p>\n\n<p>A final example would be numeric feature scaling - usual best practice here is to fit on train, use same fit to transform on validation/test. That way your model interprets everything relative to the same scale (again can think of this as \"format\"). </p>\n\n<p>Re: the broader question about whether any of this is realistic. Well, it depends. There's nothing intrinsically invalid with using some information from the data you want to predict on for model preprocessing, but the problem is that it then requires you to fit a new model after the processing is done. If you want a maintainable model that doesn't need to be constantly updated and re-validated, your preprocessing needs to be agnostic to the exact information in your prediction data. For example, if you had a real-world model based on tf-idf features and needed to predict unseen data on a daily basis, it would be pretty cumbersome to incorporate new words that occur in that unseen data and recreate that model every time you want to predict. But here on kaggle, you have a fixed dataset that you need to predict on and might easily benefit from explicitly incorporating prediction-time information. Say e.g. if there are categories that occur in train but not in the test set, it could be very useful to exclude them from train as well.   </p>",
      "rawMarkdown": "Hi Ed! Nice to see you here :)\n\nYou'll probably want to treat different types of transformations slightly differently, but the overarching principle should be that the **format of the information must be the same at train time and prediction time**. This means that the features should be the same, the category labels should be the same, the scale of numerical features should be the same, etc. Otherwise,  your model will either just fail to run or will misinterpret the data it sees at prediction time.\n\nSo for example, if you fit one LabelEncoder on train data and another on validation data, it's very unlikely that the labels will be consistent between the two. At prediction time, your model will see a category and interpret it as a different category from what it actually is, bad news. So that's why it's convenient and safe to fit the LabelEncoder on all of the data you'll be using - the other option would be to fit one label encoder on train, and use that same encoder to transform on validation/test (handling categories missing in train is a separate issue from this). One thing that hasn't been mentioned yet is that LabelEncoder isn't actually engineering anything, it's just choosing arbitrary category labels so there's no need to worry about any leakage.\n\nWhat about for tf-idf? This is trickier. So first, if you fit tf-idf separately on train / validation / test, you'll just get features that don't align. Say you get 1000 word columns in train, you might get 600 in validation and then your model will literally just not work because it won't accept the feature format. Here again you can fit tf-idf on train and then use the same tf-idf model to transform validation/test. If you do that, you'll end up just ignoring words that occur in val/test but not train, which is reasonable since it's harder to get useful information about those words. But a workaround could be to actually fit tf-idf on all of the data then run PCA/SVD (get an LSA model) to link together words that occur only in train and val/test. For example, let's say puppy and dog co-occur in train, dog and cat co-occur in test. When you run SVD, you can extract similarity between puppy and dog, dog and cat and then end up linking together puppy and cat by projecting them onto similar dimensions.\n\nA final example would be numeric feature scaling - usual best practice here is to fit on train, use same fit to transform on validation/test. That way your model interprets everything relative to the same scale (again can think of this as \"format\"). \n\nRe: the broader question about whether any of this is realistic. Well, it depends. There's nothing intrinsically invalid with using some information from the data you want to predict on for model preprocessing, but the problem is that it then requires you to fit a new model after the processing is done. If you want a maintainable model that doesn't need to be constantly updated and re-validated, your preprocessing needs to be agnostic to the exact information in your prediction data. For example, if you had a real-world model based on tf-idf features and needed to predict unseen data on a daily basis, it would be pretty cumbersome to incorporate new words that occur in that unseen data and recreate that model every time you want to predict. But here on kaggle, you have a fixed dataset that you need to predict on and might easily benefit from explicitly incorporating prediction-time information. Say e.g. if there are categories that occur in train but not in the test set, it could be very useful to exclude them from train as well.",
      "votes": null
    },
    {
      "id": "328246",
      "postDate": "05/13/2018 19:13:44",
      "content": "<p>Haha, thanks so much Joe! When I was typing up the question I was thinking to myself what a shame it is to no longer have Joe around and now here you are 🦇🦇🦇 I like you’re idea behind the SVD workaround with tf-idf. I will definitely mess around with that, and thanks for clearing up the rest!</p>",
      "rawMarkdown": "Haha, thanks so much Joe! When I was typing up the question I was thinking to myself what a shame it is to no longer have Joe around and now here you are 🦇🦇🦇 I like you’re idea behind the SVD workaround with tf-idf. I will definitely mess around with that, and thanks for clearing up the rest!",
      "votes": null
    },
    {
      "id": "329006",
      "postDate": "05/15/2018 14:14:44",
      "content": "<p>Can I get examples of things that will get me fired? For a friend of course.</p>",
      "rawMarkdown": "Can I get examples of things that will get me fired? For a friend of course.",
      "votes": null
    },
    {
      "id": "329277",
      "postDate": "05/16/2018 06:19:50",
      "content": "<p>@ edlardi \nWelcome to Kaggle. I have been using this platform for the last 5 months for 3 competitions and am a relative rookie myself. My answer is based on my experience in these competitions.</p>\n\n<p>Approach 2</p>\n\n<ul>\n<li><p>I would say Approach 2 would simulate what Data Scientists do in real world problems.  No test data is used \nfor building the final model. </p></li>\n<li><p>Be it word embeddings or numerical count calculations, fitting on training data and using those vectors to transform \nnew incoming test data would be a clean approach. </p></li>\n<li><p>As an intuition, this makes perfect sense. A machine learning model is built on known data without any influence from \nother data points. </p>\n\n<p>For ex:  Let's say, a word embedding model is built on a corpus of 1Mil item descriptions. A new set of test data points \nmay have  a completely different word distribution because of the nature of the product ( Say a specific 'NICHE' item \ncategory ) and when the original training data based vectors are used to predict the output for this test set, the results \nmay be significantly off the mark. Simply because, if a majority of the words in the test set were not available in the \noriginal training set,  the TFIDF and Count vector values become zero for the test record - thereby <strong><em>making it sparse</em></strong> \nand thereby not generalizing to the model that was built. </p>\n\n<p>This actually would confirm the intuition used to build the model which is that - We did not generalize the model for \nthis new category and it actually makes sense that the predictions are off the mark. </p></li>\n</ul>\n\n<p>Approach 3</p>\n\n<p>We have the luxury of using Approach 3 in Data Science competitions but may not be best practice in real world problems. </p>\n\n<p>For ex: In the Mercari Price Challenge - Combining train and test to generate TFIDF vectors as opposed to fitting on train and transforming on test and validation, improved the score in the Public leader board. Most of the top performers in Mercari competition advised against combining train and test to generate word vectors - from a best practice perspective and also that it doesn't give you too much of a boost .</p>\n\n<p>Interesting side story - I crashed out of Mercari since I vectorized on train and test combined and Stage II of that competition increased the test set size. My memory overloaded and I ended up not getting a Private Leaderboard score. \nSad but that will not be the case in Avito since this is not a kernel's only competition ( Thankfully ) :)</p>\n\n<p>Approach 1 </p>\n\n<p>One may not have to use this I think. Again as an intuition we use the training set to vectorize test. So independently creating test vectors would mean that we are ignoring the training set altogether. </p>\n\n<p>Hope this helps. Feel free to discuss more - any questions or any erroneous assumptions I may have made. </p>",
      "rawMarkdown": "edlardi \nWelcome to Kaggle. I have been using this platform for the last 5 months for 3 competitions and am a relative rookie myself. My answer is based on my experience in these competitions.\n\nApproach 2\n\n- I would say Approach 2 would simulate what Data Scientists do in real world problems.  No test data is used \n   for building the final model. \n\n-  Be it word embeddings or numerical count calculations, fitting on training data and using those vectors to transform \n   new incoming test data would be a clean approach. \n\n-  As an intuition, this makes perfect sense. A machine learning model is built on known data without any influence from \n   other data points. \n \n   For ex:  Let's say, a word embedding model is built on a corpus of 1Mil item descriptions. A new set of test data points \n  may have  a completely different word distribution because of the nature of the product ( Say a specific 'NICHE' item \n  category ) and when the original training data based vectors are used to predict the output for this test set, the results \n  may be significantly off the mark. Simply because, if a majority of the words in the test set were not available in the \n  original training set,  the TFIDF and Count vector values become zero for the test record - thereby ***making it sparse*** \n  and thereby not generalizing to the model that was built. \n\n  This actually would confirm the intuition used to build the model which is that - We did not generalize the model for \n   this new category and it actually makes sense that the predictions are off the mark. \n\nApproach 3\n\nWe have the luxury of using Approach 3 in Data Science competitions but may not be best practice in real world problems. \n\nFor ex: In the Mercari Price Challenge - Combining train and test to generate TFIDF vectors as opposed to fitting on train and transforming on test and validation, improved the score in the Public leader board. Most of the top performers in Mercari competition advised against combining train and test to generate word vectors - from a best practice perspective and also that it doesn't give you too much of a boost .\n\nInteresting side story - I crashed out of Mercari since I vectorized on train and test combined and Stage II of that competition increased the test set size. My memory overloaded and I ended up not getting a Private Leaderboard score. \nSad but that will not be the case in Avito since this is not a kernel's only competition ( Thankfully ) :)\n\nApproach 1 \n\nOne may not have to use this I think. Again as an intuition we use the training set to vectorize test. So independently creating test vectors would mean that we are ignoring the training set altogether. \n\nHope this helps. Feel free to discuss more - any questions or any erroneous assumptions I may have made.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 327211,
      "author_name": "haphamhoang",
      "author_url": "",
      "post_date": "05/11/2018 03:10:43",
      "content": "<p>I have not looked closely at this competition so cannot say anything specific about it. I can say in general, how you treat train - valid - test set varies according to your dataset, your problem and the features you want to generate. \nApproach 1 is actually recommended by many people, as it ensure no future information leaked to training set. \nApproach 2 should be used in your case (maybe?) About your problem with new category, you should see if the category information is really that necessary. If not, drop it entirely. If it's important, you should try to group them into bigger groups, like group 1, group 2, group 3, others... according to some criteria that you defines (can be according to the relation between the categorical variable and some other numerical variables)\nApproach 3 is still ok if your data points are not related, or your features are not \"aggregated features\" like sum, mean, median etc. However if your data points are related, for example time-series data, then you should be careful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327212,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "05/11/2018 03:11:59",
      "content": "<p>Deciding between 2 and 3 is a matter of debate. I don't know the right answer. I lean toward 2, but I've seen people do 3 and be fine. ...But definitely don't do 1.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327687,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "05/12/2018 07:16:07",
      "content": "<p>Many things/tricks work in Kaggle only, try that in real life and you should probably get fired :D <br>\nDo whatever you want, as long as you a) or b): <br>\na) Work as if the test set never existed <br>\nb) You will make sure that real life inputs will be enforced to a given standard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 329006,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "05/15/2018 14:14:44",
          "content": "<p>Can I get examples of things that will get me fired? For a friend of course.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327760,
      "author_name": "adilztn",
      "author_url": "",
      "post_date": "05/12/2018 11:57:51",
      "content": "<p>I don't advise you to use label encorder for your categorical variables, ask about target encoding (there are lots of topics on the forum).\nYour score on the leaderboard will be better.</p>\n\n<p>To make label encoder I advise you the third option</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327812,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/12/2018 14:47:08",
      "content": "<p>Hi Ed! Nice to see you here :)</p>\n\n<p>You'll probably want to treat different types of transformations slightly differently, but the overarching principle should be that the <strong>format of the information must be the same at train time and prediction time</strong>. This means that the features should be the same, the category labels should be the same, the scale of numerical features should be the same, etc. Otherwise,  your model will either just fail to run or will misinterpret the data it sees at prediction time.</p>\n\n<p>So for example, if you fit one LabelEncoder on train data and another on validation data, it's very unlikely that the labels will be consistent between the two. At prediction time, your model will see a category and interpret it as a different category from what it actually is, bad news. So that's why it's convenient and safe to fit the LabelEncoder on all of the data you'll be using - the other option would be to fit one label encoder on train, and use that same encoder to transform on validation/test (handling categories missing in train is a separate issue from this). One thing that hasn't been mentioned yet is that LabelEncoder isn't actually engineering anything, it's just choosing arbitrary category labels so there's no need to worry about any leakage.</p>\n\n<p>What about for tf-idf? This is trickier. So first, if you fit tf-idf separately on train / validation / test, you'll just get features that don't align. Say you get 1000 word columns in train, you might get 600 in validation and then your model will literally just not work because it won't accept the feature format. Here again you can fit tf-idf on train and then use the same tf-idf model to transform validation/test. If you do that, you'll end up just ignoring words that occur in val/test but not train, which is reasonable since it's harder to get useful information about those words. But a workaround could be to actually fit tf-idf on all of the data then run PCA/SVD (get an LSA model) to link together words that occur only in train and val/test. For example, let's say puppy and dog co-occur in train, dog and cat co-occur in test. When you run SVD, you can extract similarity between puppy and dog, dog and cat and then end up linking together puppy and cat by projecting them onto similar dimensions.</p>\n\n<p>A final example would be numeric feature scaling - usual best practice here is to fit on train, use same fit to transform on validation/test. That way your model interprets everything relative to the same scale (again can think of this as \"format\"). </p>\n\n<p>Re: the broader question about whether any of this is realistic. Well, it depends. There's nothing intrinsically invalid with using some information from the data you want to predict on for model preprocessing, but the problem is that it then requires you to fit a new model after the processing is done. If you want a maintainable model that doesn't need to be constantly updated and re-validated, your preprocessing needs to be agnostic to the exact information in your prediction data. For example, if you had a real-world model based on tf-idf features and needed to predict unseen data on a daily basis, it would be pretty cumbersome to incorporate new words that occur in that unseen data and recreate that model every time you want to predict. But here on kaggle, you have a fixed dataset that you need to predict on and might easily benefit from explicitly incorporating prediction-time information. Say e.g. if there are categories that occur in train but not in the test set, it could be very useful to exclude them from train as well.   </p>",
      "votes": null,
      "replies": [
        {
          "id": 328246,
          "author_name": "edleardi",
          "author_url": "",
          "post_date": "05/13/2018 19:13:44",
          "content": "<p>Haha, thanks so much Joe! When I was typing up the question I was thinking to myself what a shame it is to no longer have Joe around and now here you are 🦇🦇🦇 I like you’re idea behind the SVD workaround with tf-idf. I will definitely mess around with that, and thanks for clearing up the rest!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329277,
      "author_name": "shanth84",
      "author_url": "",
      "post_date": "05/16/2018 06:19:50",
      "content": "<p>@ edlardi \nWelcome to Kaggle. I have been using this platform for the last 5 months for 3 competitions and am a relative rookie myself. My answer is based on my experience in these competitions.</p>\n\n<p>Approach 2</p>\n\n<ul>\n<li><p>I would say Approach 2 would simulate what Data Scientists do in real world problems.  No test data is used \nfor building the final model. </p></li>\n<li><p>Be it word embeddings or numerical count calculations, fitting on training data and using those vectors to transform \nnew incoming test data would be a clean approach. </p></li>\n<li><p>As an intuition, this makes perfect sense. A machine learning model is built on known data without any influence from \nother data points. </p>\n\n<p>For ex:  Let's say, a word embedding model is built on a corpus of 1Mil item descriptions. A new set of test data points \nmay have  a completely different word distribution because of the nature of the product ( Say a specific 'NICHE' item \ncategory ) and when the original training data based vectors are used to predict the output for this test set, the results \nmay be significantly off the mark. Simply because, if a majority of the words in the test set were not available in the \noriginal training set,  the TFIDF and Count vector values become zero for the test record - thereby <strong><em>making it sparse</em></strong> \nand thereby not generalizing to the model that was built. </p>\n\n<p>This actually would confirm the intuition used to build the model which is that - We did not generalize the model for \nthis new category and it actually makes sense that the predictions are off the mark. </p></li>\n</ul>\n\n<p>Approach 3</p>\n\n<p>We have the luxury of using Approach 3 in Data Science competitions but may not be best practice in real world problems. </p>\n\n<p>For ex: In the Mercari Price Challenge - Combining train and test to generate TFIDF vectors as opposed to fitting on train and transforming on test and validation, improved the score in the Public leader board. Most of the top performers in Mercari competition advised against combining train and test to generate word vectors - from a best practice perspective and also that it doesn't give you too much of a boost .</p>\n\n<p>Interesting side story - I crashed out of Mercari since I vectorized on train and test combined and Stage II of that competition increased the test set size. My memory overloaded and I ended up not getting a Private Leaderboard score. \nSad but that will not be the case in Avito since this is not a kernel's only competition ( Thankfully ) :)</p>\n\n<p>Approach 1 </p>\n\n<p>One may not have to use this I think. Again as an intuition we use the training set to vectorize test. So independently creating test vectors would mean that we are ignoring the training set altogether. </p>\n\n<p>Hope this helps. Feel free to discuss more - any questions or any erroneous assumptions I may have made. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "327198": "Sorry for the n00b question, I'm new to data science and kaggle :O\n\nSo my approach for this competition has been to engineer features by setting up functions and then running those functions on train/val/test datasets separately. Now this strategy makes sense to me with features that are row independent (like day of the week, # of letters in the title) but I'm getting confused for what the established best practice is for handling features that are fit to an entire dataset. For instance if I want to do label encoding on city, region, and param1-3 like I saw in some of the kernels, from what I can think of, I have 3 options using scikitlearn's preprocessing.LabelEncoder() \n\n**1. Fit-transform on each of train/val/test separately:** This works for a competition since the test set is finite, but I don't think this is what practicing data scientists do since you're trying to generalize a model on the train set, right? \n\n**2. Fit on train, then transform on train, on val, and on test:** So this is the strategy that I feel like a practicing data scientist would use as you're fitting the feature engineering on train data, then applying that fit by transforming validation and test data; in other words you're generalizing the model on train data. The issue I run into with this is what if val or train have a category that isn't represented in train? Then do I just fill that in with an nan during feature generation?  \n\n**3. Fit on train+val+test, then transform on train, on val, and on test:** This is the strategy I saw in SRK and sban's kernels use. This doesn't make sense to me though because you're fitting your model on val and test data in addition to train. But doesn't that leak val and test information into your train setup and cause the model to overfit? \n\nWould appreciate any suggestions or wisdom on this. I'm also trying to implement features using tf-idf with feature reduction, and I'm running into the same question. How should I be fitting the tfidf vectorizer and then transforming my datasets? I imagine it would follow the same workflow as with label encoding.",
    "327211": "I have not looked closely at this competition so cannot say anything specific about it. I can say in general, how you treat train - valid - test set varies according to your dataset, your problem and the features you want to generate. \nApproach 1 is actually recommended by many people, as it ensure no future information leaked to training set. \nApproach 2 should be used in your case (maybe?) About your problem with new category, you should see if the category information is really that necessary. If not, drop it entirely. If it's important, you should try to group them into bigger groups, like group 1, group 2, group 3, others... according to some criteria that you defines (can be according to the relation between the categorical variable and some other numerical variables)\nApproach 3 is still ok if your data points are not related, or your features are not \"aggregated features\" like sum, mean, median etc. However if your data points are related, for example time-series data, then you should be careful.",
    "327212": "Deciding between 2 and 3 is a matter of debate. I don't know the right answer. I lean toward 2, but I've seen people do 3 and be fine. ...But definitely don't do 1.",
    "327687": "Many things/tricks work in Kaggle only, try that in real life and you should probably get fired :D  \nDo whatever you want, as long as you a) or b):  \na) Work as if the test set never existed  \nb) You will make sure that real life inputs will be enforced to a given standard.",
    "327760": "I don't advise you to use label encorder for your categorical variables, ask about target encoding (there are lots of topics on the forum).\nYour score on the leaderboard will be better.\n\nTo make label encoder I advise you the third option",
    "327812": "Hi Ed! Nice to see you here :)\n\nYou'll probably want to treat different types of transformations slightly differently, but the overarching principle should be that the **format of the information must be the same at train time and prediction time**. This means that the features should be the same, the category labels should be the same, the scale of numerical features should be the same, etc. Otherwise,  your model will either just fail to run or will misinterpret the data it sees at prediction time.\n\nSo for example, if you fit one LabelEncoder on train data and another on validation data, it's very unlikely that the labels will be consistent between the two. At prediction time, your model will see a category and interpret it as a different category from what it actually is, bad news. So that's why it's convenient and safe to fit the LabelEncoder on all of the data you'll be using - the other option would be to fit one label encoder on train, and use that same encoder to transform on validation/test (handling categories missing in train is a separate issue from this). One thing that hasn't been mentioned yet is that LabelEncoder isn't actually engineering anything, it's just choosing arbitrary category labels so there's no need to worry about any leakage.\n\nWhat about for tf-idf? This is trickier. So first, if you fit tf-idf separately on train / validation / test, you'll just get features that don't align. Say you get 1000 word columns in train, you might get 600 in validation and then your model will literally just not work because it won't accept the feature format. Here again you can fit tf-idf on train and then use the same tf-idf model to transform validation/test. If you do that, you'll end up just ignoring words that occur in val/test but not train, which is reasonable since it's harder to get useful information about those words. But a workaround could be to actually fit tf-idf on all of the data then run PCA/SVD (get an LSA model) to link together words that occur only in train and val/test. For example, let's say puppy and dog co-occur in train, dog and cat co-occur in test. When you run SVD, you can extract similarity between puppy and dog, dog and cat and then end up linking together puppy and cat by projecting them onto similar dimensions.\n\nA final example would be numeric feature scaling - usual best practice here is to fit on train, use same fit to transform on validation/test. That way your model interprets everything relative to the same scale (again can think of this as \"format\"). \n\nRe: the broader question about whether any of this is realistic. Well, it depends. There's nothing intrinsically invalid with using some information from the data you want to predict on for model preprocessing, but the problem is that it then requires you to fit a new model after the processing is done. If you want a maintainable model that doesn't need to be constantly updated and re-validated, your preprocessing needs to be agnostic to the exact information in your prediction data. For example, if you had a real-world model based on tf-idf features and needed to predict unseen data on a daily basis, it would be pretty cumbersome to incorporate new words that occur in that unseen data and recreate that model every time you want to predict. But here on kaggle, you have a fixed dataset that you need to predict on and might easily benefit from explicitly incorporating prediction-time information. Say e.g. if there are categories that occur in train but not in the test set, it could be very useful to exclude them from train as well.",
    "328246": "Haha, thanks so much Joe! When I was typing up the question I was thinking to myself what a shame it is to no longer have Joe around and now here you are 🦇🦇🦇 I like you’re idea behind the SVD workaround with tf-idf. I will definitely mess around with that, and thanks for clearing up the rest!",
    "329006": "Can I get examples of things that will get me fired? For a friend of course.",
    "329277": "edlardi \nWelcome to Kaggle. I have been using this platform for the last 5 months for 3 competitions and am a relative rookie myself. My answer is based on my experience in these competitions.\n\nApproach 2\n\n- I would say Approach 2 would simulate what Data Scientists do in real world problems.  No test data is used \n   for building the final model. \n\n-  Be it word embeddings or numerical count calculations, fitting on training data and using those vectors to transform \n   new incoming test data would be a clean approach. \n\n-  As an intuition, this makes perfect sense. A machine learning model is built on known data without any influence from \n   other data points. \n \n   For ex:  Let's say, a word embedding model is built on a corpus of 1Mil item descriptions. A new set of test data points \n  may have  a completely different word distribution because of the nature of the product ( Say a specific 'NICHE' item \n  category ) and when the original training data based vectors are used to predict the output for this test set, the results \n  may be significantly off the mark. Simply because, if a majority of the words in the test set were not available in the \n  original training set,  the TFIDF and Count vector values become zero for the test record - thereby ***making it sparse*** \n  and thereby not generalizing to the model that was built. \n\n  This actually would confirm the intuition used to build the model which is that - We did not generalize the model for \n   this new category and it actually makes sense that the predictions are off the mark. \n\nApproach 3\n\nWe have the luxury of using Approach 3 in Data Science competitions but may not be best practice in real world problems. \n\nFor ex: In the Mercari Price Challenge - Combining train and test to generate TFIDF vectors as opposed to fitting on train and transforming on test and validation, improved the score in the Public leader board. Most of the top performers in Mercari competition advised against combining train and test to generate word vectors - from a best practice perspective and also that it doesn't give you too much of a boost .\n\nInteresting side story - I crashed out of Mercari since I vectorized on train and test combined and Stage II of that competition increased the test set size. My memory overloaded and I ended up not getting a Private Leaderboard score. \nSad but that will not be the case in Avito since this is not a kernel's only competition ( Thankfully ) :)\n\nApproach 1 \n\nOne may not have to use this I think. Again as an intuition we use the training set to vectorize test. So independently creating test vectors would mean that we are ignoring the training set altogether. \n\nHope this helps. Feel free to discuss more - any questions or any erroneous assumptions I may have made."
  },
  "source": "meta"
}