{
  "id": 54968,
  "title": "May i do feature engineering on train+test dataset same time ?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54968",
  "author_name": "",
  "post_date": "2018-04-20T08:56:34.107987900Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "316956",
      "postDate": "04/20/2018 08:56:34",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "316962",
      "postDate": "04/20/2018 09:14:31",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508</a></p>",
      "rawMarkdown": "See https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508",
      "votes": null
    },
    {
      "id": "316963",
      "postDate": "04/20/2018 09:16:20",
      "content": "<p>Don't think there is any rule against that.</p>",
      "rawMarkdown": "Don't think there is any rule against that.",
      "votes": null
    },
    {
      "id": "316966",
      "postDate": "04/20/2018 09:26:45",
      "content": "<p>Of course I like your suggestion. Good to follows.</p>",
      "rawMarkdown": "Of course I like your suggestion. Good to follows.",
      "votes": null
    },
    {
      "id": "316969",
      "postDate": "04/20/2018 09:33:22",
      "content": "<p>Of course we can but it won't valid. Ok tell me how can we predict future click? how can you make features(example : ip_day_hour_count, ip_app_count) which we did for training and testing(submission)  for future click?  </p>",
      "rawMarkdown": "Of course we can but it won't valid. Ok tell me how can we predict future click? how can you make features(example : ip_day_hour_count, ip_app_count) which we did for training and testing(submission)  for future click?",
      "votes": null
    },
    {
      "id": "317061",
      "postDate": "04/20/2018 15:56:20",
      "content": "<p>That assumes your counting all the items in the set ... But you can always count all the items up to that item ;-) I wouldn't call it a \"count\" I would call it an index ... but it can be viewed as a count. This way you can look for correlations. It also shows some interesting facts about the data.</p>",
      "rawMarkdown": "That assumes your counting all the items in the set ... But you can always count all the items up to that item ;-) I wouldn't call it a \"count\" I would call it an index ... but it can be viewed as a count. This way you can look for correlations. It also shows some interesting facts about the data.",
      "votes": null
    },
    {
      "id": "317737",
      "postDate": "04/22/2018 12:24:19",
      "content": "<p>As i wrote before:</p>\n\n<blockquote>\n  <p>As Andy once said in the comment of a kernel(Sorry, i can't remember which kernel since i have try more than 20 until now): There is nothing to stop us from doing that. I think that makes sense because when we generate some new features using frequency,mean,var,count,unique etc. , we want to get the accurate numerical number as much as we can which is close to the hidden population distribution statistically in this 4 days , test set is also a sample of this population distribution in this 4 days, So we NEED to use train/test set together to generate the statistical features which can help us get close to the hidden population distribution. And as you can imagine, If we can use the entire train/test set to compute the features, that certainly would be more accurate to train our model.</p>\n</blockquote>\n\n<p>If you are new to kaggle, i strongly recommand you to see these kernals below, that would be the fastest way to get started:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\">TalkingData EDA plus time patterns</a></li>\n<li><a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">How to Work with BIG Datasets on 16G RAM (+Dask)</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">LightGBM (Fixing unbalanced data)</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">TalkingData: EDA to Model Evaluation | LB: 0.9683</a></li>\n</ul>\n\n<p>After read, try, and comprehend these kernels, you will be able to invent your own stuff and rolling in the deep with this competition!</p>",
      "rawMarkdown": "As i wrote before:\n\n&gt; As Andy once said in the comment of a kernel(Sorry, i can't remember which kernel since i have try more than 20 until now): There is nothing to stop us from doing that. I think that makes sense because when we generate some new features using frequency,mean,var,count,unique etc. , we want to get the accurate numerical number as much as we can which is close to the hidden population distribution statistically in this 4 days , test set is also a sample of this population distribution in this 4 days, So we NEED to use train/test set together to generate the statistical features which can help us get close to the hidden population distribution. And as you can imagine, If we can use the entire train/test set to compute the features, that certainly would be more accurate to train our model.\n\nIf you are new to kaggle, i strongly recommand you to see these kernals below, that would be the fastest way to get started:\n\n- [TalkingData EDA plus time patterns][1]\n- [How to Work with BIG Datasets on 16G RAM (+Dask)][2]\n- [LightGBM (Fixing unbalanced data)][3]\n- [TalkingData: EDA to Model Evaluation | LB: 0.9683][4]\n\nAfter read, try, and comprehend these kernels, you will be able to invent your own stuff and rolling in the deep with this competition!\n  [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\n  [2]: https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\n  [3]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\n  [4]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683",
      "votes": null
    },
    {
      "id": "317740",
      "postDate": "04/22/2018 12:41:13",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 316962,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/20/2018 09:14:31",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 316966,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "04/20/2018 09:26:45",
          "content": "<p>Of course I like your suggestion. Good to follows.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316963,
      "author_name": "enfeizhan",
      "author_url": "",
      "post_date": "04/20/2018 09:16:20",
      "content": "<p>Don't think there is any rule against that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 316969,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "04/20/2018 09:33:22",
          "content": "<p>Of course we can but it won't valid. Ok tell me how can we predict future click? how can you make features(example : ip_day_hour_count, ip_app_count) which we did for training and testing(submission)  for future click?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 317061,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "04/20/2018 15:56:20",
          "content": "<p>That assumes your counting all the items in the set ... But you can always count all the items up to that item ;-) I wouldn't call it a \"count\" I would call it an index ... but it can be viewed as a count. This way you can look for correlations. It also shows some interesting facts about the data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 317737,
      "author_name": "wenjiebai",
      "author_url": "",
      "post_date": "04/22/2018 12:24:19",
      "content": "<p>As i wrote before:</p>\n\n<blockquote>\n  <p>As Andy once said in the comment of a kernel(Sorry, i can't remember which kernel since i have try more than 20 until now): There is nothing to stop us from doing that. I think that makes sense because when we generate some new features using frequency,mean,var,count,unique etc. , we want to get the accurate numerical number as much as we can which is close to the hidden population distribution statistically in this 4 days , test set is also a sample of this population distribution in this 4 days, So we NEED to use train/test set together to generate the statistical features which can help us get close to the hidden population distribution. And as you can imagine, If we can use the entire train/test set to compute the features, that certainly would be more accurate to train our model.</p>\n</blockquote>\n\n<p>If you are new to kaggle, i strongly recommand you to see these kernals below, that would be the fastest way to get started:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\">TalkingData EDA plus time patterns</a></li>\n<li><a href=\"https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\">How to Work with BIG Datasets on 16G RAM (+Dask)</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">LightGBM (Fixing unbalanced data)</a></li>\n<li><a href=\"https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\">TalkingData: EDA to Model Evaluation | LB: 0.9683</a></li>\n</ul>\n\n<p>After read, try, and comprehend these kernels, you will be able to invent your own stuff and rolling in the deep with this competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 317740,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "04/22/2018 12:41:13",
          "content": "<p>Thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "316956": "",
    "316962": "See https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251#latest-316508",
    "316963": "Don't think there is any rule against that.",
    "316966": "Of course I like your suggestion. Good to follows.",
    "316969": "Of course we can but it won't valid. Ok tell me how can we predict future click? how can you make features(example : ip_day_hour_count, ip_app_count) which we did for training and testing(submission)  for future click?",
    "317061": "That assumes your counting all the items in the set ... But you can always count all the items up to that item ;-) I wouldn't call it a \"count\" I would call it an index ... but it can be viewed as a count. This way you can look for correlations. It also shows some interesting facts about the data.",
    "317737": "As i wrote before:\n\n&gt; As Andy once said in the comment of a kernel(Sorry, i can't remember which kernel since i have try more than 20 until now): There is nothing to stop us from doing that. I think that makes sense because when we generate some new features using frequency,mean,var,count,unique etc. , we want to get the accurate numerical number as much as we can which is close to the hidden population distribution statistically in this 4 days , test set is also a sample of this population distribution in this 4 days, So we NEED to use train/test set together to generate the statistical features which can help us get close to the hidden population distribution. And as you can imagine, If we can use the entire train/test set to compute the features, that certainly would be more accurate to train our model.\n\nIf you are new to kaggle, i strongly recommand you to see these kernals below, that would be the fastest way to get started:\n\n- [TalkingData EDA plus time patterns][1]\n- [How to Work with BIG Datasets on 16G RAM (+Dask)][2]\n- [LightGBM (Fixing unbalanced data)][3]\n- [TalkingData: EDA to Model Evaluation | LB: 0.9683][4]\n\nAfter read, try, and comprehend these kernels, you will be able to invent your own stuff and rolling in the deep with this competition!\n  [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\n  [2]: https://www.kaggle.com/yuliagm/how-to-work-with-big-datasets-on-16g-ram-dask\n  [3]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683\n  [4]: https://www.kaggle.com/pranav84/talkingdata-eda-to-model-evaluation-lb-0-9683",
    "317740": "Thank you."
  },
  "source": "meta"
}