{
  "id": 53485,
  "title": "How data was prepared for this competition?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53485",
  "author_name": "",
  "post_date": "2018-03-31T09:54:49.720928Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Description of the competition says that TalkingData processes 3 billion clicks per day and we have about 200 millions for 4 days. </p>\n\n<p>So it would be good to know how the data was prepared considering 60-times data minimization.\nIf the procedure is competely random, then we can miss a lot of information, it is like seeing only one frame per second for a video. My guess is that data somehow filtered by locations or ips or devices or in some other way.</p>\n\n<p>I came to this question after checking that about 94% clicks are from the same device type, which is unreasonably high. But also 67% downloads happened from the same device type - also unrealistic value. \nHence I may conclude that either competition data was filtered somehow or goal of this competition is not really good as app downloads cannot be good indication of fraud/real user.</p>",
  "messages": [
    {
      "id": "306915",
      "postDate": "03/31/2018 09:54:49",
      "content": "<p>Description of the competition says that TalkingData processes 3 billion clicks per day and we have about 200 millions for 4 days. </p>\n\n<p>So it would be good to know how the data was prepared considering 60-times data minimization.\nIf the procedure is competely random, then we can miss a lot of information, it is like seeing only one frame per second for a video. My guess is that data somehow filtered by locations or ips or devices or in some other way.</p>\n\n<p>I came to this question after checking that about 94% clicks are from the same device type, which is unreasonably high. But also 67% downloads happened from the same device type - also unrealistic value. \nHence I may conclude that either competition data was filtered somehow or goal of this competition is not really good as app downloads cannot be good indication of fraud/real user.</p>",
      "rawMarkdown": "Description of the competition says that TalkingData processes 3 billion clicks per day and we have about 200 millions for 4 days. \n\nSo it would be good to know how the data was prepared considering 60-times data minimization.\nIf the procedure is competely random, then we can miss a lot of information, it is like seeing only one frame per second for a video. My guess is that data somehow filtered by locations or ips or devices or in some other way.\n\nI came to this question after checking that about 94% clicks are from the same device type, which is unreasonably high. But also 67% downloads happened from the same device type - also unrealistic value. \nHence I may conclude that either competition data was filtered somehow or goal of this competition is not really good as app downloads cannot be good indication of fraud/real user.",
      "votes": null
    },
    {
      "id": "306949",
      "postDate": "03/31/2018 11:57:32",
      "content": "<p>One of the data scientist from TalkingData and who usually <a href=\"https://www.kaggle.com/aaronyin/discussion?sortBy=latestPost&amp;group=commentsAndTopics&amp;page=1&amp;pageSize=20\">answers the questions</a>   posted on discussion forum, once <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/51142/welcome?forumMessageId=292492#post292492\">mentioned</a>  that:</p>\n\n<pre><code>The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days.\n</code></pre>\n\n<p>Glad that this question is asked. It would be great if they can provide <strong>more details</strong>  on data preparation part. </p>",
      "rawMarkdown": "One of the data scientist from TalkingData and who usually [answers the questions][1]   posted on discussion forum, once [mentioned][2]  that:\n\n\n    The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days.\n\nGlad that this question is asked. It would be great if they can provide **more details**  on data preparation part. \n\n  [1]: https://www.kaggle.com/aaronyin/discussion?sortBy=latestPost&amp;group=commentsAndTopics&amp;page=1&amp;pageSize=20\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/51142/welcome?forumMessageId=292492#post292492",
      "votes": null
    },
    {
      "id": "307025",
      "postDate": "03/31/2018 16:43:44",
      "content": "<p>Too much and then chopped down ;)</p>",
      "rawMarkdown": "Too much and then chopped down ;)",
      "votes": null
    },
    {
      "id": "307271",
      "postDate": "04/01/2018 08:47:47",
      "content": "<p>Pranav, thank you for citation. \nAgree with you that any additional details about data preparation would be helpful.</p>\n\n<p>My current conclusion is that TalkingData is interested in some subset of clicks, and it looks like this subset is related to single device type (i.e. we have very norrow view to real picture).</p>",
      "rawMarkdown": "Pranav, thank you for citation. \nAgree with you that any additional details about data preparation would be helpful.\n\nMy current conclusion is that TalkingData is interested in some subset of clicks, and it looks like this subset is related to single device type (i.e. we have very norrow view to real picture).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306949,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "03/31/2018 11:57:32",
      "content": "<p>One of the data scientist from TalkingData and who usually <a href=\"https://www.kaggle.com/aaronyin/discussion?sortBy=latestPost&amp;group=commentsAndTopics&amp;page=1&amp;pageSize=20\">answers the questions</a>   posted on discussion forum, once <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/51142/welcome?forumMessageId=292492#post292492\">mentioned</a>  that:</p>\n\n<pre><code>The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days.\n</code></pre>\n\n<p>Glad that this question is asked. It would be great if they can provide <strong>more details</strong>  on data preparation part. </p>",
      "votes": null,
      "replies": [
        {
          "id": 307271,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/01/2018 08:47:47",
          "content": "<p>Pranav, thank you for citation. \nAgree with you that any additional details about data preparation would be helpful.</p>\n\n<p>My current conclusion is that TalkingData is interested in some subset of clicks, and it looks like this subset is related to single device type (i.e. we have very norrow view to real picture).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307025,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "03/31/2018 16:43:44",
      "content": "<p>Too much and then chopped down ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "306915": "Description of the competition says that TalkingData processes 3 billion clicks per day and we have about 200 millions for 4 days. \n\nSo it would be good to know how the data was prepared considering 60-times data minimization.\nIf the procedure is competely random, then we can miss a lot of information, it is like seeing only one frame per second for a video. My guess is that data somehow filtered by locations or ips or devices or in some other way.\n\nI came to this question after checking that about 94% clicks are from the same device type, which is unreasonably high. But also 67% downloads happened from the same device type - also unrealistic value. \nHence I may conclude that either competition data was filtered somehow or goal of this competition is not really good as app downloads cannot be good indication of fraud/real user.",
    "306949": "One of the data scientist from TalkingData and who usually [answers the questions][1]   posted on discussion forum, once [mentioned][2]  that:\n\n\n    The training data is built by extract all clicks, from a random selection across IP address of interest, for a period of time over a few days.\n\nGlad that this question is asked. It would be great if they can provide **more details**  on data preparation part. \n\n  [1]: https://www.kaggle.com/aaronyin/discussion?sortBy=latestPost&amp;group=commentsAndTopics&amp;page=1&amp;pageSize=20\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/forums/t/51142/welcome?forumMessageId=292492#post292492",
    "307025": "Too much and then chopped down ;)",
    "307271": "Pranav, thank you for citation. \nAgree with you that any additional details about data preparation would be helpful.\n\nMy current conclusion is that TalkingData is interested in some subset of clicks, and it looks like this subset is related to single device type (i.e. we have very norrow view to real picture)."
  },
  "source": "meta"
}