{
  "id": 53246,
  "title": "features to use",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53246",
  "author_name": "",
  "post_date": "2018-03-28T14:27:39.564413800Z",
  "votes": 1,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Can someone give me few ideas of how to make intelligent features?</p>",
  "messages": [
    {
      "id": "305147",
      "postDate": "03/28/2018 14:27:39",
      "content": "<p>Can someone give me few ideas of how to make intelligent features?</p>",
      "rawMarkdown": "Can someone give me few ideas of how to make intelligent features?",
      "votes": null
    },
    {
      "id": "305151",
      "postDate": "03/28/2018 14:39:12",
      "content": "<p>What about: thinking?</p>\n\n<p>Sorry if this sounds rude, but that's really the best answer I can think of.</p>",
      "rawMarkdown": "What about: thinking?\n\nSorry if this sounds rude, but that's really the best answer I can think of.",
      "votes": null
    },
    {
      "id": "305156",
      "postDate": "03/28/2018 14:45:58",
      "content": "<p>I have seen kernels with people creating different features using combination of feature but i am not able to have intution of why they are doing it?</p>",
      "rawMarkdown": "I have seen kernels with people creating different features using combination of feature but i am not able to have intution of why they are doing it?",
      "votes": null
    },
    {
      "id": "305167",
      "postDate": "03/28/2018 14:54:18",
      "content": "<p>Also i was thinking to make an ip count feature but am not able to differentiate whether should i consider the whole dataset and make the count feature or should i first split it ito train and validation and make count feature based on train set?</p>",
      "rawMarkdown": "Also i was thinking to make an ip count feature but am not able to differentiate whether should i consider the whole dataset and make the count feature or should i first split it ito train and validation and make count feature based on train set?",
      "votes": null
    },
    {
      "id": "305177",
      "postDate": "03/28/2018 15:05:12",
      "content": "<p>Feature engineering needs a lot experience and trial/error. <br>What I do as a beginner is: <br>1. Try to make possible combination of variables like frequency counts, interactions i.e count of Var 1 by Var 2 and try them one by one for training. <br> 2. If you are tasked to identify two classes with certain set of features, which combinations can help you better classify the class? Think like this and try to make some combinations  based on your intuition for ex. I can say that if any Ad  is clicked by same IP million or more times in an hour, this may help me classify the click as fraud. Think outside of the box and go with your instinct.<br> 3. Read top solutions of past similar kaggle competitions for ex. <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction\">Porto</a>. This practice might not help you find golden features for this competition but it will surely help you develop a mindset. <br> Feature Engineering is an Art, which we can only master with practice and experience.     </p>",
      "rawMarkdown": "Feature engineering needs a lot experience and trial/error. <br>What I do as a beginner is: <br>1. Try to make possible combination of variables like frequency counts, interactions i.e count of Var 1 by Var 2 and try them one by one for training. <br> 2. If you are tasked to identify two classes with certain set of features, which combinations can help you better classify the class? Think like this and try to make some combinations  based on your intuition for ex. I can say that if any Ad  is clicked by same IP million or more times in an hour, this may help me classify the click as fraud. Think outside of the box and go with your instinct.<br> 3. Read top solutions of past similar kaggle competitions for ex. [Porto][1]. This practice might not help you find golden features for this competition but it will surely help you develop a mindset. <br> Feature Engineering is an Art, which we can only master with practice and experience.     \n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction",
      "votes": null
    },
    {
      "id": "305180",
      "postDate": "03/28/2018 15:07:39",
      "content": "<p>magic feature for sale!! $1 for lb score 0.0001 improvement, $100 for lb score 0.01 improvement!! Spots are limited!! Please call 123456789 for more details!!</p>",
      "rawMarkdown": "magic feature for sale!! $1 for lb score 0.0001 improvement, $100 for lb score 0.01 improvement!! Spots are limited!! Please call 123456789 for more details!!",
      "votes": null
    },
    {
      "id": "305347",
      "postDate": "03/28/2018 18:43:26",
      "content": "<p>My advice was a bit harsh, and certainly not helpful.   Read what others do, here, and also in other competitions.  Google 'feature engineering'.  Take online course or university degrees.  There are plenty of ways to learn about this topic.</p>",
      "rawMarkdown": "My advice was a bit harsh, and certainly not helpful.   Read what others do, here, and also in other competitions.  Google 'feature engineering'.  Take online course or university degrees.  There are plenty of ways to learn about this topic.",
      "votes": null
    },
    {
      "id": "305359",
      "postDate": "03/28/2018 19:05:31",
      "content": "<p>Grab small subset of data and test out your theories :)</p>",
      "rawMarkdown": "Grab small subset of data and test out your theories :)",
      "votes": null
    },
    {
      "id": "305578",
      "postDate": "03/29/2018 05:07:38",
      "content": "<p>I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?</p>",
      "rawMarkdown": "I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?",
      "votes": null
    },
    {
      "id": "307249",
      "postDate": "04/01/2018 07:20:36",
      "content": "<p>Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...</p>\n\n<p>I will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.</p>",
      "rawMarkdown": "Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...\n\nI will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.",
      "votes": null
    },
    {
      "id": "311372",
      "postDate": "04/09/2018 23:46:51",
      "content": "<p>are we suppose to use click_id?  is it important?</p>",
      "rawMarkdown": "are we suppose to use click_id?  is it important?",
      "votes": null
    },
    {
      "id": "311438",
      "postDate": "04/10/2018 04:14:53",
      "content": "<p>One approach is to create features that make intuitive sense. For example, if I suspect that there may be a relation between box office sales and the release data of a film and I have the theater release date in my data, I may consider creating a feature that corresponds to a summer release date. </p>\n\n<p>For this particular data, what would be some of the most intuitive features to engineer? Maybe combinations of the operating system and device type? How about a feature for the time of day the click occurred?</p>",
      "rawMarkdown": "One approach is to create features that make intuitive sense. For example, if I suspect that there may be a relation between box office sales and the release data of a film and I have the theater release date in my data, I may consider creating a feature that corresponds to a summer release date. \n\nFor this particular data, what would be some of the most intuitive features to engineer? Maybe combinations of the operating system and device type? How about a feature for the time of day the click occurred?",
      "votes": null
    },
    {
      "id": "317330",
      "postDate": "04/21/2018 08:12:41",
      "content": "<p>Keep torturing the data it will confess. Ask different questions from data. For example, ask data that what was the trend during the later quarter of the day. \nThere can be many ways in which you can think, first write down all ideas then apply the filter that whether the features will be useful or not then weigh all those features importance via applying the ML algorithms and noting the differences in results with there inclusion/exclusion. \nBe creative, you can read discussions and solutions of previous competitions and can come up with good ideas that can be the blend of many points. \nThanks</p>\n\n<p><img src=\"https://i.stack.imgur.com/arpEK.jpg\" alt=\"Keep Torturing the Data\"></p>",
      "rawMarkdown": "Keep torturing the data it will confess. Ask different questions from data. For example, ask data that what was the trend during the later quarter of the day. \nThere can be many ways in which you can think, first write down all ideas then apply the filter that whether the features will be useful or not then weigh all those features importance via applying the ML algorithms and noting the differences in results with there inclusion/exclusion. \nBe creative, you can read discussions and solutions of previous competitions and can come up with good ideas that can be the blend of many points. \nThanks\n\n![Keep Torturing the Data][1]\n  [1]: https://i.stack.imgur.com/arpEK.jpg",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 305151,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/28/2018 14:39:12",
      "content": "<p>What about: thinking?</p>\n\n<p>Sorry if this sounds rude, but that's really the best answer I can think of.</p>",
      "votes": null,
      "replies": [
        {
          "id": 305156,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/28/2018 14:45:58",
          "content": "<p>I have seen kernels with people creating different features using combination of feature but i am not able to have intution of why they are doing it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 305167,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/28/2018 14:54:18",
          "content": "<p>Also i was thinking to make an ip count feature but am not able to differentiate whether should i consider the whole dataset and make the count feature or should i first split it ito train and validation and make count feature based on train set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 305347,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "03/28/2018 18:43:26",
          "content": "<p>My advice was a bit harsh, and certainly not helpful.   Read what others do, here, and also in other competitions.  Google 'feature engineering'.  Take online course or university degrees.  There are plenty of ways to learn about this topic.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 305177,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/28/2018 15:05:12",
      "content": "<p>Feature engineering needs a lot experience and trial/error. <br>What I do as a beginner is: <br>1. Try to make possible combination of variables like frequency counts, interactions i.e count of Var 1 by Var 2 and try them one by one for training. <br> 2. If you are tasked to identify two classes with certain set of features, which combinations can help you better classify the class? Think like this and try to make some combinations  based on your intuition for ex. I can say that if any Ad  is clicked by same IP million or more times in an hour, this may help me classify the click as fraud. Think outside of the box and go with your instinct.<br> 3. Read top solutions of past similar kaggle competitions for ex. <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction\">Porto</a>. This practice might not help you find golden features for this competition but it will surely help you develop a mindset. <br> Feature Engineering is an Art, which we can only master with practice and experience.     </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 305180,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "03/28/2018 15:07:39",
      "content": "<p>magic feature for sale!! $1 for lb score 0.0001 improvement, $100 for lb score 0.01 improvement!! Spots are limited!! Please call 123456789 for more details!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 305359,
      "author_name": "michaelsnell",
      "author_url": "",
      "post_date": "03/28/2018 19:05:31",
      "content": "<p>Grab small subset of data and test out your theories :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 305578,
          "author_name": "nitishsingh41",
          "author_url": "",
          "post_date": "03/29/2018 05:07:38",
          "content": "<p>I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307249,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "04/01/2018 07:20:36",
      "content": "<p>Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...</p>\n\n<p>I will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311372,
      "author_name": "karankatiyar",
      "author_url": "",
      "post_date": "04/09/2018 23:46:51",
      "content": "<p>are we suppose to use click_id?  is it important?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 311438,
      "author_name": "attackgnome",
      "author_url": "",
      "post_date": "04/10/2018 04:14:53",
      "content": "<p>One approach is to create features that make intuitive sense. For example, if I suspect that there may be a relation between box office sales and the release data of a film and I have the theater release date in my data, I may consider creating a feature that corresponds to a summer release date. </p>\n\n<p>For this particular data, what would be some of the most intuitive features to engineer? Maybe combinations of the operating system and device type? How about a feature for the time of day the click occurred?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 317330,
      "author_name": "muhammadzaman",
      "author_url": "",
      "post_date": "04/21/2018 08:12:41",
      "content": "<p>Keep torturing the data it will confess. Ask different questions from data. For example, ask data that what was the trend during the later quarter of the day. \nThere can be many ways in which you can think, first write down all ideas then apply the filter that whether the features will be useful or not then weigh all those features importance via applying the ML algorithms and noting the differences in results with there inclusion/exclusion. \nBe creative, you can read discussions and solutions of previous competitions and can come up with good ideas that can be the blend of many points. \nThanks</p>\n\n<p><img src=\"https://i.stack.imgur.com/arpEK.jpg\" alt=\"Keep Torturing the Data\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "305147": "Can someone give me few ideas of how to make intelligent features?",
    "305151": "What about: thinking?\n\nSorry if this sounds rude, but that's really the best answer I can think of.",
    "305156": "I have seen kernels with people creating different features using combination of feature but i am not able to have intution of why they are doing it?",
    "305167": "Also i was thinking to make an ip count feature but am not able to differentiate whether should i consider the whole dataset and make the count feature or should i first split it ito train and validation and make count feature based on train set?",
    "305177": "Feature engineering needs a lot experience and trial/error. <br>What I do as a beginner is: <br>1. Try to make possible combination of variables like frequency counts, interactions i.e count of Var 1 by Var 2 and try them one by one for training. <br> 2. If you are tasked to identify two classes with certain set of features, which combinations can help you better classify the class? Think like this and try to make some combinations  based on your intuition for ex. I can say that if any Ad  is clicked by same IP million or more times in an hour, this may help me classify the click as fraud. Think outside of the box and go with your instinct.<br> 3. Read top solutions of past similar kaggle competitions for ex. [Porto][1]. This practice might not help you find golden features for this competition but it will surely help you develop a mindset. <br> Feature Engineering is an Art, which we can only master with practice and experience.     \n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction",
    "305180": "magic feature for sale!! $1 for lb score 0.0001 improvement, $100 for lb score 0.01 improvement!! Spots are limited!! Please call 123456789 for more details!!",
    "305347": "My advice was a bit harsh, and certainly not helpful.   Read what others do, here, and also in other competitions.  Google 'feature engineering'.  Take online course or university degrees.  There are plenty of ways to learn about this topic.",
    "305359": "Grab small subset of data and test out your theories :)",
    "305578": "I took ip and ip count feature and it gave better prediction than the model with just ip count feature. I was wondering as ip and ip count feature are dependent on each other then shouldn't it decrease my model accuracy?",
    "307249": "Surely, just experimenting will result in finding some good features and you will practice coding. Just trial and error. Try this combination, see the result, try that check if the LB score increases. Var on ip+os+attribute_time per hour? Day of the week as a feature where there are only 4 days in the whole data? Sure, just throw everything...\n\nI will go against all those things. Again - you will become better coder, but I dont think you will become better and feature engineering. My advice is - try to think what makes sense in real world ! What are the existing features, what is the problem, what do you want to predict? How would you do it if you had 1 shot? First try those things, and then of course a bit of experimenatition wouldn't hurt, but it should be supplementary and not instead of logical thinking and understanding the real business problem.",
    "311372": "are we suppose to use click_id?  is it important?",
    "311438": "One approach is to create features that make intuitive sense. For example, if I suspect that there may be a relation between box office sales and the release data of a film and I have the theater release date in my data, I may consider creating a feature that corresponds to a summer release date. \n\nFor this particular data, what would be some of the most intuitive features to engineer? Maybe combinations of the operating system and device type? How about a feature for the time of day the click occurred?",
    "317330": "Keep torturing the data it will confess. Ask different questions from data. For example, ask data that what was the trend during the later quarter of the day. \nThere can be many ways in which you can think, first write down all ideas then apply the filter that whether the features will be useful or not then weigh all those features importance via applying the ML algorithms and noting the differences in results with there inclusion/exclusion. \nBe creative, you can read discussions and solutions of previous competitions and can come up with good ideas that can be the blend of many points. \nThanks\n\n![Keep Torturing the Data][1]\n  [1]: https://i.stack.imgur.com/arpEK.jpg"
  },
  "source": "meta"
}