{
  "id": 55346,
  "title": "In real scenario, these features wouldn't be used",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55346",
  "author_name": "",
  "post_date": "2018-04-25T13:35:39.879888600Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi all</p>\n\n<p>I saw most people use lots of frequency features(including me) in their models. However, in real scenario, these features wouldn't be used, at least, not in this way. In online advertisement scenario, anit-spam/anti-fraud detection is usually done in real-time mode, which means you can't get those frequency features, especially for those TEST instances. So, the LB score is not very practical for real use.</p>\n\n<p>UPDATE: What I say here depend on company's strategy and business scenario. From my experiment , the real-time advertisement bidding, DSP or ad exchange need predict if a click even an impression is fraud which would affect the AD effect and cost. For long term analysis or post-validation, it's ok to use all data. Just like the AD conversion, the click and conversion do not happen at the same time, it may happens hours or days after,  you need join this data in a rolling window, such as 3 or 5 days. But, before the real conversion happen, you need do some prediction for ranking or recommendation, how can you use such information?  I am not saying you can't use those features, just saying it depends on the scenario, specially some which require online/real-time prediction, offline analysis is not in this range. By the way,  I just say that most of kernel's count features are not useful in real life, because you have all test set.</p>\n\n<p>Best Regards,</p>",
  "messages": [
    {
      "id": "319193",
      "postDate": "04/25/2018 13:35:39",
      "content": "<p>Hi all</p>\n\n<p>I saw most people use lots of frequency features(including me) in their models. However, in real scenario, these features wouldn't be used, at least, not in this way. In online advertisement scenario, anit-spam/anti-fraud detection is usually done in real-time mode, which means you can't get those frequency features, especially for those TEST instances. So, the LB score is not very practical for real use.</p>\n\n<p>UPDATE: What I say here depend on company's strategy and business scenario. From my experiment , the real-time advertisement bidding, DSP or ad exchange need predict if a click even an impression is fraud which would affect the AD effect and cost. For long term analysis or post-validation, it's ok to use all data. Just like the AD conversion, the click and conversion do not happen at the same time, it may happens hours or days after,  you need join this data in a rolling window, such as 3 or 5 days. But, before the real conversion happen, you need do some prediction for ranking or recommendation, how can you use such information?  I am not saying you can't use those features, just saying it depends on the scenario, specially some which require online/real-time prediction, offline analysis is not in this range. By the way,  I just say that most of kernel's count features are not useful in real life, because you have all test set.</p>\n\n<p>Best Regards,</p>",
      "rawMarkdown": "Hi all\n\nI saw most people use lots of frequency features(including me) in their models. However, in real scenario, these features wouldn't be used, at least, not in this way. In online advertisement scenario, anit-spam/anti-fraud detection is usually done in real-time mode, which means you can't get those frequency features, especially for those TEST instances. So, the LB score is not very practical for real use.\n\nUPDATE: What I say here depend on company's strategy and business scenario. From my experiment , the real-time advertisement bidding, DSP or ad exchange need predict if a click even an impression is fraud which would affect the AD effect and cost. For long term analysis or post-validation, it's ok to use all data. Just like the AD conversion, the click and conversion do not happen at the same time, it may happens hours or days after,  you need join this data in a rolling window, such as 3 or 5 days. But, before the real conversion happen, you need do some prediction for ranking or recommendation, how can you use such information?  I am not saying you can't use those features, just saying it depends on the scenario, specially some which require online/real-time prediction, offline analysis is not in this range. By the way,  I just say that most of kernel's count features are not useful in real life, because you have all test set.\n \nBest Regards,",
      "votes": null
    },
    {
      "id": "319196",
      "postDate": "04/25/2018 13:37:33",
      "content": "<p>Welcome to kaggle! </p>",
      "rawMarkdown": "Welcome to kaggle!",
      "votes": null
    },
    {
      "id": "319227",
      "postDate": "04/25/2018 14:52:28",
      "content": "<p>Yes you can use most off them.\nYou can recreate features daily (rolling windows of the last 3 days) and update your model to predict the current day. You have the clicks for the last 3 days, so you can learn over time without deteriorated performance.\nVoila.</p>",
      "rawMarkdown": "Yes you can use most off them.\nYou can recreate features daily (rolling windows of the last 3 days) and update your model to predict the current day. You have the clicks for the last 3 days, so you can learn over time without deteriorated performance.\nVoila.",
      "votes": null
    },
    {
      "id": "319287",
      "postDate": "04/25/2018 17:41:28",
      "content": "<p>I have seen posts like these pop up before, but I don't really understand why they pop up. Unless someone is an insider to the company, how can they know for sure that certain features can't be used?</p>\n\n<p>Even if someone is an expert in a field, commenting on another company's workings yields educated guesses at best.</p>",
      "rawMarkdown": "I have seen posts like these pop up before, but I don't really understand why they pop up. Unless someone is an insider to the company, how can they know for sure that certain features can't be used?\n\nEven if someone is an expert in a field, commenting on another company's workings yields educated guesses at best.",
      "votes": null
    },
    {
      "id": "319288",
      "postDate": "04/25/2018 17:44:52",
      "content": "<p>So the attributed time is exactly the same as click_time?  You are just plain wrong - they must be using stats post click_time</p>",
      "rawMarkdown": "So the attributed time is exactly the same as click_time?  You are just plain wrong - they must be using stats post click_time",
      "votes": null
    },
    {
      "id": "319305",
      "postDate": "04/25/2018 18:30:20",
      "content": "<p>There is causality. In real time you can't use features from the future. For example using data from the end of test to predict the start of test. For example count frequency of ip using all of the supplement test.</p>",
      "rawMarkdown": "There is causality. In real time you can't use features from the future. For example using data from the end of test to predict the start of test. For example count frequency of ip using all of the supplement test.",
      "votes": null
    },
    {
      "id": "319311",
      "postDate": "04/25/2018 18:48:06",
      "content": "<p>If you were interested, I also discussed this topic a few weeks ago. And I admit that at the beginning I also had a very similar opinion to yours. But the truth is that companies can implement such a model, for example, the next day. And then they have very similar data like us. Look here: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645</a></p>",
      "rawMarkdown": "If you were interested, I also discussed this topic a few weeks ago. And I admit that at the beginning I also had a very similar opinion to yours. But the truth is that companies can implement such a model, for example, the next day. And then they have very similar data like us. Look here: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645",
      "votes": null
    },
    {
      "id": "319355",
      "postDate": "04/25/2018 21:13:05",
      "content": "<p>I believe that fitting a model for frequency and updating it from time to time would be the way these features would appear in industry.</p>\n\n<p>Anyway, it's an interesting discussion. Of course that for the possible leakage part of it you would never be able to use, but that's not particular to the kind of feature engineering here...</p>",
      "rawMarkdown": "I believe that fitting a model for frequency and updating it from time to time would be the way these features would appear in industry.\n\nAnyway, it's an interesting discussion. Of course that for the possible leakage part of it you would never be able to use, but that's not particular to the kind of feature engineering here...",
      "votes": null
    },
    {
      "id": "319414",
      "postDate": "04/26/2018 02:17:07",
      "content": "<p>yes. what I mean is that such as count feature is not practical in sponsor advertisement or display advertisement field. In financial  field, such as credit card theft detection, it's ok to use those count feature, because it's offline detection.  </p>\n\n<p>Actually, in industry, we use some count features also, but not in this way.</p>",
      "rawMarkdown": "yes. what I mean is that such as count feature is not practical in sponsor advertisement or display advertisement field. In financial  field, such as credit card theft detection, it's ok to use those count feature, because it's offline detection.  \n\nActually, in industry, we use some count features also, but not in this way.",
      "votes": null
    },
    {
      "id": "319415",
      "postDate": "04/26/2018 02:19:56",
      "content": "<p>Thank you for the reply.  You are right, it depends on what strategy the company choose. If company don't need a real-time detection, it's ok. However, in real online advertisement company ,such as Google, or other DSP company, it's always real-time, which means they don't have such TEST information like us.</p>",
      "rawMarkdown": "Thank you for the reply.  You are right, it depends on what strategy the company choose. If company don't need a real-time detection, it's ok. However, in real online advertisement company ,such as Google, or other DSP company, it's always real-time, which means they don't have such TEST information like us.",
      "votes": null
    },
    {
      "id": "319417",
      "postDate": "04/26/2018 02:20:57",
      "content": "<p>Yes, you are right, it's what  I mean. I just say that most of kernel's count features are not useful in real life.</p>",
      "rawMarkdown": "Yes, you are right, it's what  I mean. I just say that most of kernel's count features are not useful in real life.",
      "votes": null
    },
    {
      "id": "319419",
      "postDate": "04/26/2018 02:26:07",
      "content": "<p>It's a good question, just like the conversion rate prediction in online advertisement. The click and conversion does not happen at the same time. What my point is how to design those count feature. It's ok for this competition, because you have all data, both of train and test. In real life, how can we get those TEST information? Or, it's not a real-time prediction, then it's ok.</p>",
      "rawMarkdown": "It's a good question, just like the conversion rate prediction in online advertisement. The click and conversion does not happen at the same time. What my point is how to design those count feature. It's ok for this competition, because you have all data, both of train and test. In real life, how can we get those TEST information? Or, it's not a real-time prediction, then it's ok.",
      "votes": null
    },
    {
      "id": "319482",
      "postDate": "04/26/2018 06:48:54",
      "content": "<p>I know for sure that Google adsense publishers have report with invalid clicks, which is calculated month after actual ads clicks, just before payments. There is a mechanism to analyze clicks history and this is far from realtime.</p>",
      "rawMarkdown": "I know for sure that Google adsense publishers have report with invalid clicks, which is calculated month after actual ads clicks, just before payments. There is a mechanism to analyze clicks history and this is far from realtime.",
      "votes": null
    },
    {
      "id": "319513",
      "postDate": "04/26/2018 07:55:38",
      "content": "<p>I thought so. In this situation, the used features seem to be quite valid.</p>",
      "rawMarkdown": "I thought so. In this situation, the used features seem to be quite valid.",
      "votes": null
    },
    {
      "id": "319552",
      "postDate": "04/26/2018 09:31:47",
      "content": "<p>interesting knowledge. Anyway, it depends on company strategy. For real-time bidding, DSP or Ad exchange need to know if a click or an impression is fraud, which would affect the advertisement bidding. </p>",
      "rawMarkdown": "interesting knowledge. Anyway, it depends on company strategy. For real-time bidding, DSP or Ad exchange need to know if a click or an impression is fraud, which would affect the advertisement bidding.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319196,
      "author_name": "hireme",
      "author_url": "",
      "post_date": "04/25/2018 13:37:33",
      "content": "<p>Welcome to kaggle! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319227,
      "author_name": "mrbeer",
      "author_url": "",
      "post_date": "04/25/2018 14:52:28",
      "content": "<p>Yes you can use most off them.\nYou can recreate features daily (rolling windows of the last 3 days) and update your model to predict the current day. You have the clicks for the last 3 days, so you can learn over time without deteriorated performance.\nVoila.</p>",
      "votes": null,
      "replies": [
        {
          "id": 319417,
          "author_name": "jdxyw2004",
          "author_url": "",
          "post_date": "04/26/2018 02:20:57",
          "content": "<p>Yes, you are right, it's what  I mean. I just say that most of kernel's count features are not useful in real life.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319287,
      "author_name": "antmarakis",
      "author_url": "",
      "post_date": "04/25/2018 17:41:28",
      "content": "<p>I have seen posts like these pop up before, but I don't really understand why they pop up. Unless someone is an insider to the company, how can they know for sure that certain features can't be used?</p>\n\n<p>Even if someone is an expert in a field, commenting on another company's workings yields educated guesses at best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 319305,
          "author_name": "mrbeer",
          "author_url": "",
          "post_date": "04/25/2018 18:30:20",
          "content": "<p>There is causality. In real time you can't use features from the future. For example using data from the end of test to predict the start of test. For example count frequency of ip using all of the supplement test.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319288,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "04/25/2018 17:44:52",
      "content": "<p>So the attributed time is exactly the same as click_time?  You are just plain wrong - they must be using stats post click_time</p>",
      "votes": null,
      "replies": [
        {
          "id": 319419,
          "author_name": "jdxyw2004",
          "author_url": "",
          "post_date": "04/26/2018 02:26:07",
          "content": "<p>It's a good question, just like the conversion rate prediction in online advertisement. The click and conversion does not happen at the same time. What my point is how to design those count feature. It's ok for this competition, because you have all data, both of train and test. In real life, how can we get those TEST information? Or, it's not a real-time prediction, then it's ok.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319311,
      "author_name": "skalskip",
      "author_url": "",
      "post_date": "04/25/2018 18:48:06",
      "content": "<p>If you were interested, I also discussed this topic a few weeks ago. And I admit that at the beginning I also had a very similar opinion to yours. But the truth is that companies can implement such a model, for example, the next day. And then they have very similar data like us. Look here: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 319415,
          "author_name": "jdxyw2004",
          "author_url": "",
          "post_date": "04/26/2018 02:19:56",
          "content": "<p>Thank you for the reply.  You are right, it depends on what strategy the company choose. If company don't need a real-time detection, it's ok. However, in real online advertisement company ,such as Google, or other DSP company, it's always real-time, which means they don't have such TEST information like us.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319482,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/26/2018 06:48:54",
          "content": "<p>I know for sure that Google adsense publishers have report with invalid clicks, which is calculated month after actual ads clicks, just before payments. There is a mechanism to analyze clicks history and this is far from realtime.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319513,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "04/26/2018 07:55:38",
          "content": "<p>I thought so. In this situation, the used features seem to be quite valid.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319552,
          "author_name": "jdxyw2004",
          "author_url": "",
          "post_date": "04/26/2018 09:31:47",
          "content": "<p>interesting knowledge. Anyway, it depends on company strategy. For real-time bidding, DSP or Ad exchange need to know if a click or an impression is fraud, which would affect the advertisement bidding. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 319355,
      "author_name": "lgmoneda",
      "author_url": "",
      "post_date": "04/25/2018 21:13:05",
      "content": "<p>I believe that fitting a model for frequency and updating it from time to time would be the way these features would appear in industry.</p>\n\n<p>Anyway, it's an interesting discussion. Of course that for the possible leakage part of it you would never be able to use, but that's not particular to the kind of feature engineering here...</p>",
      "votes": null,
      "replies": [
        {
          "id": 319414,
          "author_name": "jdxyw2004",
          "author_url": "",
          "post_date": "04/26/2018 02:17:07",
          "content": "<p>yes. what I mean is that such as count feature is not practical in sponsor advertisement or display advertisement field. In financial  field, such as credit card theft detection, it's ok to use those count feature, because it's offline detection.  </p>\n\n<p>Actually, in industry, we use some count features also, but not in this way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "319193": "Hi all\n\nI saw most people use lots of frequency features(including me) in their models. However, in real scenario, these features wouldn't be used, at least, not in this way. In online advertisement scenario, anit-spam/anti-fraud detection is usually done in real-time mode, which means you can't get those frequency features, especially for those TEST instances. So, the LB score is not very practical for real use.\n\nUPDATE: What I say here depend on company's strategy and business scenario. From my experiment , the real-time advertisement bidding, DSP or ad exchange need predict if a click even an impression is fraud which would affect the AD effect and cost. For long term analysis or post-validation, it's ok to use all data. Just like the AD conversion, the click and conversion do not happen at the same time, it may happens hours or days after,  you need join this data in a rolling window, such as 3 or 5 days. But, before the real conversion happen, you need do some prediction for ranking or recommendation, how can you use such information?  I am not saying you can't use those features, just saying it depends on the scenario, specially some which require online/real-time prediction, offline analysis is not in this range. By the way,  I just say that most of kernel's count features are not useful in real life, because you have all test set.\n \nBest Regards,",
    "319196": "Welcome to kaggle!",
    "319227": "Yes you can use most off them.\nYou can recreate features daily (rolling windows of the last 3 days) and update your model to predict the current day. You have the clicks for the last 3 days, so you can learn over time without deteriorated performance.\nVoila.",
    "319287": "I have seen posts like these pop up before, but I don't really understand why they pop up. Unless someone is an insider to the company, how can they know for sure that certain features can't be used?\n\nEven if someone is an expert in a field, commenting on another company's workings yields educated guesses at best.",
    "319288": "So the attributed time is exactly the same as click_time?  You are just plain wrong - they must be using stats post click_time",
    "319305": "There is causality. In real time you can't use features from the future. For example using data from the end of test to predict the start of test. For example count frequency of ip using all of the supplement test.",
    "319311": "If you were interested, I also discussed this topic a few weeks ago. And I admit that at the beginning I also had a very similar opinion to yours. But the truth is that companies can implement such a model, for example, the next day. And then they have very similar data like us. Look here: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53595#307645",
    "319355": "I believe that fitting a model for frequency and updating it from time to time would be the way these features would appear in industry.\n\nAnyway, it's an interesting discussion. Of course that for the possible leakage part of it you would never be able to use, but that's not particular to the kind of feature engineering here...",
    "319414": "yes. what I mean is that such as count feature is not practical in sponsor advertisement or display advertisement field. In financial  field, such as credit card theft detection, it's ok to use those count feature, because it's offline detection.  \n\nActually, in industry, we use some count features also, but not in this way.",
    "319415": "Thank you for the reply.  You are right, it depends on what strategy the company choose. If company don't need a real-time detection, it's ok. However, in real online advertisement company ,such as Google, or other DSP company, it's always real-time, which means they don't have such TEST information like us.",
    "319417": "Yes, you are right, it's what  I mean. I just say that most of kernel's count features are not useful in real life.",
    "319419": "It's a good question, just like the conversion rate prediction in online advertisement. The click and conversion does not happen at the same time. What my point is how to design those count feature. It's ok for this competition, because you have all data, both of train and test. In real life, how can we get those TEST information? Or, it's not a real-time prediction, then it's ok.",
    "319482": "I know for sure that Google adsense publishers have report with invalid clicks, which is calculated month after actual ads clicks, just before payments. There is a mechanism to analyze clicks history and this is far from realtime.",
    "319513": "I thought so. In this situation, the used features seem to be quite valid.",
    "319552": "interesting knowledge. Anyway, it depends on company strategy. For real-time bidding, DSP or Ad exchange need to know if a click or an impression is fraud, which would affect the advertisement bidding."
  },
  "source": "meta"
}