{
  "id": 53595,
  "title": "Model Leakage",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53595",
  "author_name": "",
  "post_date": "2018-04-02T07:00:25.587896100Z",
  "votes": 12,
  "comment_count": 20,
  "views": 0,
  "content": "<p>At the beginning I will say that this competition is my first on Kaggle. Perhaps this is a normal practice here, but browsing the public kernels I noticed almost every one of them is leaking. People deliberately introduce information during training that undermines the whole reason of creating machine learning models. </p>\n\n<ul>\n<li>Validation for selected hours that will be included in the test set</li>\n<li>Creating classification variables such as belongs to the most common app or device from a test set</li>\n</ul>\n\n<p>These are just a few examples. Perhaps this approach allows you to deliberately fit better into the test data set, but it certainly does not create a better model that you can use in practice. I know that you will probably say I needlessly get angry, but it seems to me a little pointless.</p>",
  "messages": [
    {
      "id": "307645",
      "postDate": "04/02/2018 07:00:25",
      "content": "<p>At the beginning I will say that this competition is my first on Kaggle. Perhaps this is a normal practice here, but browsing the public kernels I noticed almost every one of them is leaking. People deliberately introduce information during training that undermines the whole reason of creating machine learning models. </p>\n\n<ul>\n<li>Validation for selected hours that will be included in the test set</li>\n<li>Creating classification variables such as belongs to the most common app or device from a test set</li>\n</ul>\n\n<p>These are just a few examples. Perhaps this approach allows you to deliberately fit better into the test data set, but it certainly does not create a better model that you can use in practice. I know that you will probably say I needlessly get angry, but it seems to me a little pointless.</p>",
      "rawMarkdown": "At the beginning I will say that this competition is my first on Kaggle. Perhaps this is a normal practice here, but browsing the public kernels I noticed almost every one of them is leaking. People deliberately introduce information during training that undermines the whole reason of creating machine learning models. \n\n - Validation for selected hours that will be included in the test set\n - Creating classification variables such as belongs to the most common app or device from a test set\n\nThese are just a few examples. Perhaps this approach allows you to deliberately fit better into the test data set, but it certainly does not create a better model that you can use in practice. I know that you will probably say I needlessly get angry, but it seems to me a little pointless.",
      "votes": null
    },
    {
      "id": "307657",
      "postDate": "04/02/2018 07:27:51",
      "content": "<p>Generally, you are right, but it's a part of the any competition.</p>",
      "rawMarkdown": "Generally, you are right, but it's a part of the any competition.",
      "votes": null
    },
    {
      "id": "307670",
      "postDate": "04/02/2018 07:52:04",
      "content": "<p>I thought so. Apparently I will have to get used to the fact that everything I learned on this subject on Kaggle is not valid. :)</p>\n\n<p>Is this not against any rules? The company sponsors the competition with the hope that it will receive a model in return that will help them solve a problem. In fact, such a model with leaks will be much weaker, or it will have to be modified due to problems associated with it.</p>",
      "rawMarkdown": "I thought so. Apparently I will have to get used to the fact that everything I learned on this subject on Kaggle is not valid. :)\n\nIs this not against any rules? The company sponsors the competition with the hope that it will receive a model in return that will help them solve a problem. In fact, such a model with leaks will be much weaker, or it will have to be modified due to problems associated with it.",
      "votes": null
    },
    {
      "id": "307681",
      "postDate": "04/02/2018 08:03:12",
      "content": "<p>Yeah, but the best models will be very complex and company cannot use it in production. But they can find some interesting ideas or interesting people. So, have fun at this competition and learn something new. </p>\n\n<p>Good luck :)</p>",
      "rawMarkdown": "Yeah, but the best models will be very complex and company cannot use it in production. But they can find some interesting ideas or interesting people. So, have fun at this competition and learn something new. \n\nGood luck :)",
      "votes": null
    },
    {
      "id": "307689",
      "postDate": "04/02/2018 08:43:02",
      "content": "<p>You assume that they want to detect the downloads immediately in real life. Maybe for them, it is enough to detect them a day after, since they are actually looking for frauds, not actual download probabilities. So in this case, we are even handicapped compared to the real life scenario where they also have the labels for all other rows for that day.</p>",
      "rawMarkdown": "You assume that they want to detect the downloads immediately in real life. Maybe for them, it is enough to detect them a day after, since they are actually looking for frauds, not actual download probabilities. So in this case, we are even handicapped compared to the real life scenario where they also have the labels for all other rows for that day.",
      "votes": null
    },
    {
      "id": "307701",
      "postDate": "04/02/2018 08:59:50",
      "content": "<p>I agree with you.  Good news maybe is that good ML practice is generally good on Kaggle competitions. </p>\n\n<p>For instance, using test data distribution as a feature is indeed weird from a ML perspective.  It could be valid if test data distribution is similar to train data distribution.  But if this is the case, then we could use train data distribution instead of test data.</p>\n\n<p>Also, don't worry too much about public kernels.  They often (not always) overfit to the public test data because LB score is used to tune the model in them.  Given authors seldom describe how they tune their model there can be doubt about their value.  kernels with validation included are much more valuable.</p>\n\n<p>This said we know the test data, and if using that information can  lead to better prediction on test data then it's part of the game.  This is where Kaggle departs from real world.  In some Kaggle competitions, the private test data is only disclosed close to the competition end, and you must finalize your model before that.  This is like in real world.  The current Data Science Bowl competition is one of these.</p>",
      "rawMarkdown": "I agree with you.  Good news maybe is that good ML practice is generally good on Kaggle competitions. \n\nFor instance, using test data distribution as a feature is indeed weird from a ML perspective.  It could be valid if test data distribution is similar to train data distribution.  But if this is the case, then we could use train data distribution instead of test data.\n\nAlso, don't worry too much about public kernels.  They often (not always) overfit to the public test data because LB score is used to tune the model in them.  Given authors seldom describe how they tune their model there can be doubt about their value.  kernels with validation included are much more valuable.\n\nThis said we know the test data, and if using that information can  lead to better prediction on test data then it's part of the game.  This is where Kaggle departs from real world.  In some Kaggle competitions, the private test data is only disclosed close to the competition end, and you must finalize your model before that.  This is like in real world.  The current Data Science Bowl competition is one of these.",
      "votes": null
    },
    {
      "id": "307703",
      "postDate": "04/02/2018 09:04:57",
      "content": "<p>It really was my assumption. I admit it. But if we use features such as those I mentioned, the model will not require re-training every time we want to predict something? </p>\n\n<p>belongs to the most common app or device from a test set -&gt; we train the model for the selected training set</p>",
      "rawMarkdown": "It really was my assumption. I admit it. But if we use features such as those I mentioned, the model will not require re-training every time we want to predict something? \n\nbelongs to the most common app or device from a test set -&gt; we train the model for the selected training set",
      "votes": null
    },
    {
      "id": "307708",
      "postDate": "04/02/2018 09:12:56",
      "content": "<p>As I mentioned, this is my first Kaggle competition. So far, my activity on this platform has focused on the creation of small kernels. I started because I wanted to develop and learn something new. And of course I learn a lot here. But ... There are also traps. It is worth paying attention to these aspects of the competitions. That they are different from the real problem. Because for someone who learns, just like me, certain things are not so obvious. :) Nevertheless, Kaggle is great.</p>",
      "rawMarkdown": "As I mentioned, this is my first Kaggle competition. So far, my activity on this platform has focused on the creation of small kernels. I started because I wanted to develop and learn something new. And of course I learn a lot here. But ... There are also traps. It is worth paying attention to these aspects of the competitions. That they are different from the real problem. Because for someone who learns, just like me, certain things are not so obvious. :) Nevertheless, Kaggle is great.",
      "votes": null
    },
    {
      "id": "307745",
      "postDate": "04/02/2018 11:12:37",
      "content": "<p>For me Kaggle is a place to practice, learn and experiment. It is not a one to one correspondance with ML practices in real-world business settings. As far as you understand this, I don't see a problem with spending some time on Kaggle. I hope that also companies that organize competitions also understand that it wont be reasonable to implement the winning solution straight into production (even if possible). Again the idea is different - practice, explore, learn and experiment</p>",
      "rawMarkdown": "For me Kaggle is a place to practice, learn and experiment. It is not a one to one correspondance with ML practices in real-world business settings. As far as you understand this, I don't see a problem with spending some time on Kaggle. I hope that also companies that organize competitions also understand that it wont be reasonable to implement the winning solution straight into production (even if possible). Again the idea is different - practice, explore, learn and experiment",
      "votes": null
    },
    {
      "id": "307761",
      "postDate": "04/02/2018 12:06:17",
      "content": "<blockquote>\n  <p>belongs to the most common app or device from a test set -&gt; we train the model for the selected training set</p>\n</blockquote>\n\n<p>I am not very familiar with app stores and such, since I do not have a smartphone, but I doubt these two variables will change much. The most common apps and devices is not something that will change overnight, but with time. I assume a lot of time, so much that re-training would have been necessary either way.</p>\n\n<p>I don't think this is an issue, although I agree that validating for the hours that will be in the test set is not the best of ML practices.</p>",
      "rawMarkdown": "&gt; belongs to the most common app or device from a test set -&gt; we train the model for the selected training set\n\nI am not very familiar with app stores and such, since I do not have a smartphone, but I doubt these two variables will change much. The most common apps and devices is not something that will change overnight, but with time. I assume a lot of time, so much that re-training would have been necessary either way.\n\nI don't think this is an issue, although I agree that validating for the hours that will be in the test set is not the best of ML practices.",
      "votes": null
    },
    {
      "id": "307766",
      "postDate": "04/02/2018 12:10:24",
      "content": "<p>Good observation. And it's great that there is something to comment on those basic b*tch public kernels in terms of best practices in machine learning compared to lucky number kernels. I just hope that you at least appreciate their time and effort to build such basic b**ch models.  <strong>Please accept the fact that not everyone on Kaggle is professional data scientist.</strong>  </p>\n\n<p>Although I have became shameless to learn and  gladly accept constructive criticism but I think some of your wordings are bit harsh and feel that many kernel authors may get demotivated who are spending their time on competition to learn something from fellow Kagglers and data scientists with industry knowledge. </p>",
      "rawMarkdown": "Good observation. And it's great that there is something to comment on those basic b*tch public kernels in terms of best practices in machine learning compared to lucky number kernels. I just hope that you at least appreciate their time and effort to build such basic b**ch models.  **Please accept the fact that not everyone on Kaggle is professional data scientist.**  \n\nAlthough I have became shameless to learn and  gladly accept constructive criticism but I think some of your wordings are bit harsh and feel that many kernel authors may get demotivated who are spending their time on competition to learn something from fellow Kagglers and data scientists with industry knowledge.",
      "votes": null
    },
    {
      "id": "307817",
      "postDate": "04/02/2018 14:30:05",
      "content": "<p>For me, it was not about being disrespectful to anyone. I understand that they put their time, as we all do. And this competition, which is my first, has taught me so far so far.</p>\n\n<blockquote>\n  <p>Please accept the fact that not everyone on Kaggle is professional data scientist.</p>\n</blockquote>\n\n<p>I am also not a data scientist, but I would like to be one, one day.</p>",
      "rawMarkdown": "For me, it was not about being disrespectful to anyone. I understand that they put their time, as we all do. And this competition, which is my first, has taught me so far so far.\n\n&gt; Please accept the fact that not everyone on Kaggle is professional data scientist.\n\nI am also not a data scientist, but I would like to be one, one day.",
      "votes": null
    },
    {
      "id": "307832",
      "postDate": "04/02/2018 15:16:31",
      "content": "<p>I agree with a lot of your sentiment, but I'd encourage you to reframe it a bit. The problem framing isn't that realistic (look-back models, only modeling for certain times of day), but once you accept that framing I believe that proper ML practices still apply for the sake of doing well in this competition.</p>\n\n<p><strong>Validation for select hours that will be included in the test set</strong>: once you accept that the task you're being asked to do here is to predict on these hours, of course it is proper machine learning to select a validation set that is most representative of the task. This is just a good idea given a weirdly constrained problem, not kaggle magic. As a quick terminology point, people typically wouldn't call this \"leakage\" but instead a form of overfitting in the broader context of a model that would work for the whole day.</p>\n\n<p><strong>Creating classification variables such as belongs to the most common app or device from a test set</strong>: When I saw these in the kernels it also bothered me a lot. Things like raw counts across all 4 days did not seem like proper feature engineering. So I've made a very careful point of <em>not</em> using features like that and only using information that's available at prediction time. So far I think it is working better for me than using features that \"cheat\". </p>\n\n<p>In short, if you think people are violating best ML practices in the kernels/discussions, there's a very good chance that course-correcting will improve your results, not make them worse. </p>",
      "rawMarkdown": "I agree with a lot of your sentiment, but I'd encourage you to reframe it a bit. The problem framing isn't that realistic (look-back models, only modeling for certain times of day), but once you accept that framing I believe that proper ML practices still apply for the sake of doing well in this competition.\n\n**Validation for select hours that will be included in the test set**: once you accept that the task you're being asked to do here is to predict on these hours, of course it is proper machine learning to select a validation set that is most representative of the task. This is just a good idea given a weirdly constrained problem, not kaggle magic. As a quick terminology point, people typically wouldn't call this \"leakage\" but instead a form of overfitting in the broader context of a model that would work for the whole day.\n\n**Creating classification variables such as belongs to the most common app or device from a test set**: When I saw these in the kernels it also bothered me a lot. Things like raw counts across all 4 days did not seem like proper feature engineering. So I've made a very careful point of *not* using features like that and only using information that's available at prediction time. So far I think it is working better for me than using features that \"cheat\". \n\nIn short, if you think people are violating best ML practices in the kernels/discussions, there's a very good chance that course-correcting will improve your results, not make them worse.",
      "votes": null
    },
    {
      "id": "307869",
      "postDate": "04/02/2018 16:23:01",
      "content": "<p>Regarding the first point it seems to me quite a reasonable assumption that we will make predictions the next day. And we will have the entire set of data from yesterday. As for the second point you can not hide that being creative is paying off for you - 25th palce. Out of curiosity, the result 0.9734 is single model or stacking?</p>",
      "rawMarkdown": "Regarding the first point it seems to me quite a reasonable assumption that we will make predictions the next day. And we will have the entire set of data from yesterday. As for the second point you can not hide that being creative is paying off for you - 25th palce. Out of curiosity, the result 0.9734 is single model or stacking?",
      "votes": null
    },
    {
      "id": "307887",
      "postDate": "04/02/2018 16:48:12",
      "content": "<p>Cool. No worries! </p>",
      "rawMarkdown": "Cool. No worries!",
      "votes": null
    },
    {
      "id": "307903",
      "postDate": "04/02/2018 17:11:10",
      "content": "<p>Yeah, here the unrealistic restriction to certain hours for testing arises from kaggle's platform constraints - if you read other threads, apparently the original (full day) test data was too large for their current platform to accommodate. I think everyone can agree that's unfortunate, but we should still make the correct modeling choices to adapt to the situation (i.e. use corresponding hours for validation), as we would in real life ML if we were faced with a similar constraint for some reason.</p>\n\n<p>My 0.9734 score is a single model.</p>",
      "rawMarkdown": "Yeah, here the unrealistic restriction to certain hours for testing arises from kaggle's platform constraints - if you read other threads, apparently the original (full day) test data was too large for their current platform to accommodate. I think everyone can agree that's unfortunate, but we should still make the correct modeling choices to adapt to the situation (i.e. use corresponding hours for validation), as we would in real life ML if we were faced with a similar constraint for some reason.\n\nMy 0.9734 score is a single model.",
      "votes": null
    },
    {
      "id": "307914",
      "postDate": "04/02/2018 17:24:45",
      "content": "<p>In this situation, I have to make some corrections in my model. That is, set validation for selected hours from the last day, and not for the whole last day. What more, the conversation with you convinced me to continue looking for new features. Because there are many things that can be found. Still.</p>",
      "rawMarkdown": "In this situation, I have to make some corrections in my model. That is, set validation for selected hours from the last day, and not for the whole last day. What more, the conversation with you convinced me to continue looking for new features. Because there are many things that can be found. Still.",
      "votes": null
    },
    {
      "id": "307915",
      "postDate": "04/02/2018 17:26:34",
      "content": "<p>Oh yes I think there's lots still to be found. I feel that I'm only scratching the surface.</p>",
      "rawMarkdown": "Oh yes I think there's lots still to be found. I feel that I'm only scratching the surface.",
      "votes": null
    },
    {
      "id": "307916",
      "postDate": "04/02/2018 17:30:48",
      "content": "<p>I simply lack experience. And I do not even know what should come to my mind and what could work. I am waiting a bit for the end of the competition, I want to see what was possible and what I did not get into.</p>",
      "rawMarkdown": "I simply lack experience. And I do not even know what should come to my mind and what could work. I am waiting a bit for the end of the competition, I want to see what was possible and what I did not get into.",
      "votes": null
    },
    {
      "id": "308171",
      "postDate": "04/03/2018 04:18:37",
      "content": "<p>Your opinion might be true but we just simply lack the information of how the company are going to use the model. Moreover, no feature restriction so far.</p>\n\n<p>IF, talking data wants to use the model in the end of the day (23:59) to evaluate all clicks in that day, then all the public kernel did a good job so far.</p>\n\n<p>IF, talking data wants to use the model in real-time then most of public kernel is unacceptable.</p>",
      "rawMarkdown": "Your opinion might be true but we just simply lack the information of how the company are going to use the model. Moreover, no feature restriction so far.\n\nIF, talking data wants to use the model in the end of the day (23:59) to evaluate all clicks in that day, then all the public kernel did a good job so far.\n\nIF, talking data wants to use the model in real-time then most of public kernel is unacceptable.",
      "votes": null
    },
    {
      "id": "311740",
      "postDate": "04/10/2018 15:30:17",
      "content": "<p>Interesting observation. I've worried about the same. How are others sub-setting data to create features? For example, I've seen kernels that use all days available as Joe Eddy mentioned. Are others just deriving features on the day of or perhaps even just the day prior, etc?</p>",
      "rawMarkdown": "Interesting observation. I've worried about the same. How are others sub-setting data to create features? For example, I've seen kernels that use all days available as Joe Eddy mentioned. Are others just deriving features on the day of or perhaps even just the day prior, etc?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307657,
      "author_name": "nvarganov",
      "author_url": "",
      "post_date": "04/02/2018 07:27:51",
      "content": "<p>Generally, you are right, but it's a part of the any competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 307670,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 07:52:04",
          "content": "<p>I thought so. Apparently I will have to get used to the fact that everything I learned on this subject on Kaggle is not valid. :)</p>\n\n<p>Is this not against any rules? The company sponsors the competition with the hope that it will receive a model in return that will help them solve a problem. In fact, such a model with leaks will be much weaker, or it will have to be modified due to problems associated with it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307681,
          "author_name": "nvarganov",
          "author_url": "",
          "post_date": "04/02/2018 08:03:12",
          "content": "<p>Yeah, but the best models will be very complex and company cannot use it in production. But they can find some interesting ideas or interesting people. So, have fun at this competition and learn something new. </p>\n\n<p>Good luck :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307689,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "04/02/2018 08:43:02",
      "content": "<p>You assume that they want to detect the downloads immediately in real life. Maybe for them, it is enough to detect them a day after, since they are actually looking for frauds, not actual download probabilities. So in this case, we are even handicapped compared to the real life scenario where they also have the labels for all other rows for that day.</p>",
      "votes": null,
      "replies": [
        {
          "id": 307703,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 09:04:57",
          "content": "<p>It really was my assumption. I admit it. But if we use features such as those I mentioned, the model will not require re-training every time we want to predict something? </p>\n\n<p>belongs to the most common app or device from a test set -&gt; we train the model for the selected training set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307761,
          "author_name": "antmarakis",
          "author_url": "",
          "post_date": "04/02/2018 12:06:17",
          "content": "<blockquote>\n  <p>belongs to the most common app or device from a test set -&gt; we train the model for the selected training set</p>\n</blockquote>\n\n<p>I am not very familiar with app stores and such, since I do not have a smartphone, but I doubt these two variables will change much. The most common apps and devices is not something that will change overnight, but with time. I assume a lot of time, so much that re-training would have been necessary either way.</p>\n\n<p>I don't think this is an issue, although I agree that validating for the hours that will be in the test set is not the best of ML practices.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307701,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/02/2018 08:59:50",
      "content": "<p>I agree with you.  Good news maybe is that good ML practice is generally good on Kaggle competitions. </p>\n\n<p>For instance, using test data distribution as a feature is indeed weird from a ML perspective.  It could be valid if test data distribution is similar to train data distribution.  But if this is the case, then we could use train data distribution instead of test data.</p>\n\n<p>Also, don't worry too much about public kernels.  They often (not always) overfit to the public test data because LB score is used to tune the model in them.  Given authors seldom describe how they tune their model there can be doubt about their value.  kernels with validation included are much more valuable.</p>\n\n<p>This said we know the test data, and if using that information can  lead to better prediction on test data then it's part of the game.  This is where Kaggle departs from real world.  In some Kaggle competitions, the private test data is only disclosed close to the competition end, and you must finalize your model before that.  This is like in real world.  The current Data Science Bowl competition is one of these.</p>",
      "votes": null,
      "replies": [
        {
          "id": 307708,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 09:12:56",
          "content": "<p>As I mentioned, this is my first Kaggle competition. So far, my activity on this platform has focused on the creation of small kernels. I started because I wanted to develop and learn something new. And of course I learn a lot here. But ... There are also traps. It is worth paying attention to these aspects of the competitions. That they are different from the real problem. Because for someone who learns, just like me, certain things are not so obvious. :) Nevertheless, Kaggle is great.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307745,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/02/2018 11:12:37",
      "content": "<p>For me Kaggle is a place to practice, learn and experiment. It is not a one to one correspondance with ML practices in real-world business settings. As far as you understand this, I don't see a problem with spending some time on Kaggle. I hope that also companies that organize competitions also understand that it wont be reasonable to implement the winning solution straight into production (even if possible). Again the idea is different - practice, explore, learn and experiment</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307766,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/02/2018 12:10:24",
      "content": "<p>Good observation. And it's great that there is something to comment on those basic b*tch public kernels in terms of best practices in machine learning compared to lucky number kernels. I just hope that you at least appreciate their time and effort to build such basic b**ch models.  <strong>Please accept the fact that not everyone on Kaggle is professional data scientist.</strong>  </p>\n\n<p>Although I have became shameless to learn and  gladly accept constructive criticism but I think some of your wordings are bit harsh and feel that many kernel authors may get demotivated who are spending their time on competition to learn something from fellow Kagglers and data scientists with industry knowledge. </p>",
      "votes": null,
      "replies": [
        {
          "id": 307817,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 14:30:05",
          "content": "<p>For me, it was not about being disrespectful to anyone. I understand that they put their time, as we all do. And this competition, which is my first, has taught me so far so far.</p>\n\n<blockquote>\n  <p>Please accept the fact that not everyone on Kaggle is professional data scientist.</p>\n</blockquote>\n\n<p>I am also not a data scientist, but I would like to be one, one day.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307887,
          "author_name": "pranav84",
          "author_url": "",
          "post_date": "04/02/2018 16:48:12",
          "content": "<p>Cool. No worries! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307832,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/02/2018 15:16:31",
      "content": "<p>I agree with a lot of your sentiment, but I'd encourage you to reframe it a bit. The problem framing isn't that realistic (look-back models, only modeling for certain times of day), but once you accept that framing I believe that proper ML practices still apply for the sake of doing well in this competition.</p>\n\n<p><strong>Validation for select hours that will be included in the test set</strong>: once you accept that the task you're being asked to do here is to predict on these hours, of course it is proper machine learning to select a validation set that is most representative of the task. This is just a good idea given a weirdly constrained problem, not kaggle magic. As a quick terminology point, people typically wouldn't call this \"leakage\" but instead a form of overfitting in the broader context of a model that would work for the whole day.</p>\n\n<p><strong>Creating classification variables such as belongs to the most common app or device from a test set</strong>: When I saw these in the kernels it also bothered me a lot. Things like raw counts across all 4 days did not seem like proper feature engineering. So I've made a very careful point of <em>not</em> using features like that and only using information that's available at prediction time. So far I think it is working better for me than using features that \"cheat\". </p>\n\n<p>In short, if you think people are violating best ML practices in the kernels/discussions, there's a very good chance that course-correcting will improve your results, not make them worse. </p>",
      "votes": null,
      "replies": [
        {
          "id": 307869,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 16:23:01",
          "content": "<p>Regarding the first point it seems to me quite a reasonable assumption that we will make predictions the next day. And we will have the entire set of data from yesterday. As for the second point you can not hide that being creative is paying off for you - 25th palce. Out of curiosity, the result 0.9734 is single model or stacking?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307903,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/02/2018 17:11:10",
          "content": "<p>Yeah, here the unrealistic restriction to certain hours for testing arises from kaggle's platform constraints - if you read other threads, apparently the original (full day) test data was too large for their current platform to accommodate. I think everyone can agree that's unfortunate, but we should still make the correct modeling choices to adapt to the situation (i.e. use corresponding hours for validation), as we would in real life ML if we were faced with a similar constraint for some reason.</p>\n\n<p>My 0.9734 score is a single model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307914,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 17:24:45",
          "content": "<p>In this situation, I have to make some corrections in my model. That is, set validation for selected hours from the last day, and not for the whole last day. What more, the conversation with you convinced me to continue looking for new features. Because there are many things that can be found. Still.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307915,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "04/02/2018 17:26:34",
          "content": "<p>Oh yes I think there's lots still to be found. I feel that I'm only scratching the surface.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307916,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/02/2018 17:30:48",
          "content": "<p>I simply lack experience. And I do not even know what should come to my mind and what could work. I am waiting a bit for the end of the competition, I want to see what was possible and what I did not get into.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 308171,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "04/03/2018 04:18:37",
      "content": "<p>Your opinion might be true but we just simply lack the information of how the company are going to use the model. Moreover, no feature restriction so far.</p>\n\n<p>IF, talking data wants to use the model in the end of the day (23:59) to evaluate all clicks in that day, then all the public kernel did a good job so far.</p>\n\n<p>IF, talking data wants to use the model in real-time then most of public kernel is unacceptable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311740,
      "author_name": "learnmower",
      "author_url": "",
      "post_date": "04/10/2018 15:30:17",
      "content": "<p>Interesting observation. I've worried about the same. How are others sub-setting data to create features? For example, I've seen kernels that use all days available as Joe Eddy mentioned. Are others just deriving features on the day of or perhaps even just the day prior, etc?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "307645": "At the beginning I will say that this competition is my first on Kaggle. Perhaps this is a normal practice here, but browsing the public kernels I noticed almost every one of them is leaking. People deliberately introduce information during training that undermines the whole reason of creating machine learning models. \n\n - Validation for selected hours that will be included in the test set\n - Creating classification variables such as belongs to the most common app or device from a test set\n\nThese are just a few examples. Perhaps this approach allows you to deliberately fit better into the test data set, but it certainly does not create a better model that you can use in practice. I know that you will probably say I needlessly get angry, but it seems to me a little pointless.",
    "307657": "Generally, you are right, but it's a part of the any competition.",
    "307670": "I thought so. Apparently I will have to get used to the fact that everything I learned on this subject on Kaggle is not valid. :)\n\nIs this not against any rules? The company sponsors the competition with the hope that it will receive a model in return that will help them solve a problem. In fact, such a model with leaks will be much weaker, or it will have to be modified due to problems associated with it.",
    "307681": "Yeah, but the best models will be very complex and company cannot use it in production. But they can find some interesting ideas or interesting people. So, have fun at this competition and learn something new. \n\nGood luck :)",
    "307689": "You assume that they want to detect the downloads immediately in real life. Maybe for them, it is enough to detect them a day after, since they are actually looking for frauds, not actual download probabilities. So in this case, we are even handicapped compared to the real life scenario where they also have the labels for all other rows for that day.",
    "307701": "I agree with you.  Good news maybe is that good ML practice is generally good on Kaggle competitions. \n\nFor instance, using test data distribution as a feature is indeed weird from a ML perspective.  It could be valid if test data distribution is similar to train data distribution.  But if this is the case, then we could use train data distribution instead of test data.\n\nAlso, don't worry too much about public kernels.  They often (not always) overfit to the public test data because LB score is used to tune the model in them.  Given authors seldom describe how they tune their model there can be doubt about their value.  kernels with validation included are much more valuable.\n\nThis said we know the test data, and if using that information can  lead to better prediction on test data then it's part of the game.  This is where Kaggle departs from real world.  In some Kaggle competitions, the private test data is only disclosed close to the competition end, and you must finalize your model before that.  This is like in real world.  The current Data Science Bowl competition is one of these.",
    "307703": "It really was my assumption. I admit it. But if we use features such as those I mentioned, the model will not require re-training every time we want to predict something? \n\nbelongs to the most common app or device from a test set -&gt; we train the model for the selected training set",
    "307708": "As I mentioned, this is my first Kaggle competition. So far, my activity on this platform has focused on the creation of small kernels. I started because I wanted to develop and learn something new. And of course I learn a lot here. But ... There are also traps. It is worth paying attention to these aspects of the competitions. That they are different from the real problem. Because for someone who learns, just like me, certain things are not so obvious. :) Nevertheless, Kaggle is great.",
    "307745": "For me Kaggle is a place to practice, learn and experiment. It is not a one to one correspondance with ML practices in real-world business settings. As far as you understand this, I don't see a problem with spending some time on Kaggle. I hope that also companies that organize competitions also understand that it wont be reasonable to implement the winning solution straight into production (even if possible). Again the idea is different - practice, explore, learn and experiment",
    "307761": "&gt; belongs to the most common app or device from a test set -&gt; we train the model for the selected training set\n\nI am not very familiar with app stores and such, since I do not have a smartphone, but I doubt these two variables will change much. The most common apps and devices is not something that will change overnight, but with time. I assume a lot of time, so much that re-training would have been necessary either way.\n\nI don't think this is an issue, although I agree that validating for the hours that will be in the test set is not the best of ML practices.",
    "307766": "Good observation. And it's great that there is something to comment on those basic b*tch public kernels in terms of best practices in machine learning compared to lucky number kernels. I just hope that you at least appreciate their time and effort to build such basic b**ch models.  **Please accept the fact that not everyone on Kaggle is professional data scientist.**  \n\nAlthough I have became shameless to learn and  gladly accept constructive criticism but I think some of your wordings are bit harsh and feel that many kernel authors may get demotivated who are spending their time on competition to learn something from fellow Kagglers and data scientists with industry knowledge.",
    "307817": "For me, it was not about being disrespectful to anyone. I understand that they put their time, as we all do. And this competition, which is my first, has taught me so far so far.\n\n&gt; Please accept the fact that not everyone on Kaggle is professional data scientist.\n\nI am also not a data scientist, but I would like to be one, one day.",
    "307832": "I agree with a lot of your sentiment, but I'd encourage you to reframe it a bit. The problem framing isn't that realistic (look-back models, only modeling for certain times of day), but once you accept that framing I believe that proper ML practices still apply for the sake of doing well in this competition.\n\n**Validation for select hours that will be included in the test set**: once you accept that the task you're being asked to do here is to predict on these hours, of course it is proper machine learning to select a validation set that is most representative of the task. This is just a good idea given a weirdly constrained problem, not kaggle magic. As a quick terminology point, people typically wouldn't call this \"leakage\" but instead a form of overfitting in the broader context of a model that would work for the whole day.\n\n**Creating classification variables such as belongs to the most common app or device from a test set**: When I saw these in the kernels it also bothered me a lot. Things like raw counts across all 4 days did not seem like proper feature engineering. So I've made a very careful point of *not* using features like that and only using information that's available at prediction time. So far I think it is working better for me than using features that \"cheat\". \n\nIn short, if you think people are violating best ML practices in the kernels/discussions, there's a very good chance that course-correcting will improve your results, not make them worse.",
    "307869": "Regarding the first point it seems to me quite a reasonable assumption that we will make predictions the next day. And we will have the entire set of data from yesterday. As for the second point you can not hide that being creative is paying off for you - 25th palce. Out of curiosity, the result 0.9734 is single model or stacking?",
    "307887": "Cool. No worries!",
    "307903": "Yeah, here the unrealistic restriction to certain hours for testing arises from kaggle's platform constraints - if you read other threads, apparently the original (full day) test data was too large for their current platform to accommodate. I think everyone can agree that's unfortunate, but we should still make the correct modeling choices to adapt to the situation (i.e. use corresponding hours for validation), as we would in real life ML if we were faced with a similar constraint for some reason.\n\nMy 0.9734 score is a single model.",
    "307914": "In this situation, I have to make some corrections in my model. That is, set validation for selected hours from the last day, and not for the whole last day. What more, the conversation with you convinced me to continue looking for new features. Because there are many things that can be found. Still.",
    "307915": "Oh yes I think there's lots still to be found. I feel that I'm only scratching the surface.",
    "307916": "I simply lack experience. And I do not even know what should come to my mind and what could work. I am waiting a bit for the end of the competition, I want to see what was possible and what I did not get into.",
    "308171": "Your opinion might be true but we just simply lack the information of how the company are going to use the model. Moreover, no feature restriction so far.\n\nIF, talking data wants to use the model in the end of the day (23:59) to evaluate all clicks in that day, then all the public kernel did a good job so far.\n\nIF, talking data wants to use the model in real-time then most of public kernel is unacceptable.",
    "311740": "Interesting observation. I've worried about the same. How are others sub-setting data to create features? For example, I've seen kernels that use all days available as Joe Eddy mentioned. Are others just deriving features on the day of or perhaps even just the day prior, etc?"
  },
  "source": "meta"
}