{
  "id": 58415,
  "title": "100% Correlation",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/58415",
  "author_name": "",
  "post_date": "2018-06-07T16:03:48.312042600Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>If we do feature engineering, and turn the attributed_time variable to binary i.e. 0 if not downloaded and 1 if downloaded then it becomes the same as the dependent variable. Am I making some mistake over here?</p>",
  "messages": [
    {
      "id": "339776",
      "postDate": "06/07/2018 16:03:48",
      "content": "<p>If we do feature engineering, and turn the attributed_time variable to binary i.e. 0 if not downloaded and 1 if downloaded then it becomes the same as the dependent variable. Am I making some mistake over here?</p>",
      "rawMarkdown": "If we do feature engineering, and turn the attributed_time variable to binary i.e. 0 if not downloaded and 1 if downloaded then it becomes the same as the dependent variable. Am I making some mistake over here?",
      "votes": null
    },
    {
      "id": "340680",
      "postDate": "06/10/2018 02:14:01",
      "content": "<p>You should not need to consider the variable 'attributed_time' for following reasons:\n 1. This column is not present in test data. How would you generate this feature in test data set?\n 2. What if this information was given in the test dataset? Though it is not possible but somehow we got this information. This would be a classic case of data leakage. Any variable included in the training set and generated after the or with the event of interest causes data leakage. For example, if our variable of interest is \"whether the customer will make a purchase or not\" and we are given with the price paid for the purchase. On this information, we may divide the given data into train, CV, and test set and check for our accuracy. Strong predictor would be \"price of purchase\". We will get the perfect model. However, after deployment, the model will fail because we will not have this strong predictor variable. </p>\n\n<p>Also, in most of the competitions, you can get to know about such variables by looking at the train and test datasets. Some competitions have failed because of this data leakage. Most of the times, the winning solutions have data leakage.</p>\n\n<p>You can search on the internet to read more about it. </p>",
      "rawMarkdown": "You should not need to consider the variable 'attributed_time' for following reasons:\n 1. This column is not present in test data. How would you generate this feature in test data set?\n 2. What if this information was given in the test dataset? Though it is not possible but somehow we got this information. This would be a classic case of data leakage. Any variable included in the training set and generated after the or with the event of interest causes data leakage. For example, if our variable of interest is \"whether the customer will make a purchase or not\" and we are given with the price paid for the purchase. On this information, we may divide the given data into train, CV, and test set and check for our accuracy. Strong predictor would be \"price of purchase\". We will get the perfect model. However, after deployment, the model will fail because we will not have this strong predictor variable. \n\nAlso, in most of the competitions, you can get to know about such variables by looking at the train and test datasets. Some competitions have failed because of this data leakage. Most of the times, the winning solutions have data leakage.\n\nYou can search on the internet to read more about it.",
      "votes": null
    },
    {
      "id": "347916",
      "postDate": "06/25/2018 17:12:45",
      "content": "<p>Thanks Gaurav! Did not know test data did not have the 'attributed_time' feature. Also, I have read about data leakage but for the first time I came across one in a dataset. Hence did not realize that this can be a case of data leakage. Thanks for clearing it up.</p>",
      "rawMarkdown": "Thanks Gaurav! Did not know test data did not have the 'attributed_time' feature. Also, I have read about data leakage but for the first time I came across one in a dataset. Hence did not realize that this can be a case of data leakage. Thanks for clearing it up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 340680,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "06/10/2018 02:14:01",
      "content": "<p>You should not need to consider the variable 'attributed_time' for following reasons:\n 1. This column is not present in test data. How would you generate this feature in test data set?\n 2. What if this information was given in the test dataset? Though it is not possible but somehow we got this information. This would be a classic case of data leakage. Any variable included in the training set and generated after the or with the event of interest causes data leakage. For example, if our variable of interest is \"whether the customer will make a purchase or not\" and we are given with the price paid for the purchase. On this information, we may divide the given data into train, CV, and test set and check for our accuracy. Strong predictor would be \"price of purchase\". We will get the perfect model. However, after deployment, the model will fail because we will not have this strong predictor variable. </p>\n\n<p>Also, in most of the competitions, you can get to know about such variables by looking at the train and test datasets. Some competitions have failed because of this data leakage. Most of the times, the winning solutions have data leakage.</p>\n\n<p>You can search on the internet to read more about it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 347916,
          "author_name": "aayushd94",
          "author_url": "",
          "post_date": "06/25/2018 17:12:45",
          "content": "<p>Thanks Gaurav! Did not know test data did not have the 'attributed_time' feature. Also, I have read about data leakage but for the first time I came across one in a dataset. Hence did not realize that this can be a case of data leakage. Thanks for clearing it up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "339776": "If we do feature engineering, and turn the attributed_time variable to binary i.e. 0 if not downloaded and 1 if downloaded then it becomes the same as the dependent variable. Am I making some mistake over here?",
    "340680": "You should not need to consider the variable 'attributed_time' for following reasons:\n 1. This column is not present in test data. How would you generate this feature in test data set?\n 2. What if this information was given in the test dataset? Though it is not possible but somehow we got this information. This would be a classic case of data leakage. Any variable included in the training set and generated after the or with the event of interest causes data leakage. For example, if our variable of interest is \"whether the customer will make a purchase or not\" and we are given with the price paid for the purchase. On this information, we may divide the given data into train, CV, and test set and check for our accuracy. Strong predictor would be \"price of purchase\". We will get the perfect model. However, after deployment, the model will fail because we will not have this strong predictor variable. \n\nAlso, in most of the competitions, you can get to know about such variables by looking at the train and test datasets. Some competitions have failed because of this data leakage. Most of the times, the winning solutions have data leakage.\n\nYou can search on the internet to read more about it.",
    "347916": "Thanks Gaurav! Did not know test data did not have the 'attributed_time' feature. Also, I have read about data leakage but for the first time I came across one in a dataset. Hence did not realize that this can be a case of data leakage. Thanks for clearing it up."
  },
  "source": "meta"
}