{
  "id": 54605,
  "title": "The importance of test_sup",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54605",
  "author_name": "",
  "post_date": "2018-04-15T18:27:21.470582500Z",
  "votes": -5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I didn't realize till now that the test_sup is really important. I was using some methods that didn't work because of the test dataset released by TalkingData. It is clear that there are some huge gap between hours:</p>\n\n<pre><code>dbadmin =&gt; select hour(click_time), count(*) from AdFraudDetection_f_test group by 1 order by 1;\n hour |  count  \n------+---------\n    4 | 3344125\n    5 | 2858427\n    6 |     381\n    9 | 2984808\n   10 | 3127993\n   11 |     413\n   13 | 3212566\n   14 | 3261257\n   15 |     499\n(9 rows)\n</code></pre>\n\n<p>How can you release something like this? It can destroy some implementation methods. We are loosing too much information. If they want a code that they can generalize for Fraud Detection, they must be more serious... Do they want a random algorithms or something that you can use in the future?</p>",
  "messages": [
    {
      "id": "314521",
      "postDate": "04/15/2018 18:27:21",
      "content": "<p>I didn't realize till now that the test_sup is really important. I was using some methods that didn't work because of the test dataset released by TalkingData. It is clear that there are some huge gap between hours:</p>\n\n<pre><code>dbadmin =&gt; select hour(click_time), count(*) from AdFraudDetection_f_test group by 1 order by 1;\n hour |  count  \n------+---------\n    4 | 3344125\n    5 | 2858427\n    6 |     381\n    9 | 2984808\n   10 | 3127993\n   11 |     413\n   13 | 3212566\n   14 | 3261257\n   15 |     499\n(9 rows)\n</code></pre>\n\n<p>How can you release something like this? It can destroy some implementation methods. We are loosing too much information. If they want a code that they can generalize for Fraud Detection, they must be more serious... Do they want a random algorithms or something that you can use in the future?</p>",
      "rawMarkdown": "I didn't realize till now that the test_sup is really important. I was using some methods that didn't work because of the test dataset released by TalkingData. It is clear that there are some huge gap between hours:\n\n    dbadmin =&gt; select hour(click_time), count(*) from AdFraudDetection_f_test group by 1 order by 1;\n     hour |  count  \n    ------+---------\n        4 | 3344125\n        5 | 2858427\n        6 |     381\n        9 | 2984808\n       10 | 3127993\n       11 |     413\n       13 | 3212566\n       14 | 3261257\n       15 |     499\n    (9 rows)\n\nHow can you release something like this? It can destroy some implementation methods. We are loosing too much information. If they want a code that they can generalize for Fraud Detection, they must be more serious... Do they want a random algorithms or something that you can use in the future?",
      "votes": null
    },
    {
      "id": "314530",
      "postDate": "04/15/2018 18:48:21",
      "content": "<p>They discussed in a separate thread, but tl;dr basically it was too much data for the kernels and scorer to handle. They hope in the future as kaggle moves onto it's new platform to be able to support this sort of thing. Hence test_sup released as a Do Your Own Analysis dset.</p>\n\n<p>It was mentioned in the discussions though.</p>",
      "rawMarkdown": "They discussed in a separate thread, but tl;dr basically it was too much data for the kernels and scorer to handle. They hope in the future as kaggle moves onto it's new platform to be able to support this sort of thing. Hence test_sup released as a Do Your Own Analysis dset.\n\nIt was mentioned in the discussions though.",
      "votes": null
    },
    {
      "id": "314611",
      "postDate": "04/15/2018 23:53:54",
      "content": "<p>In that case,\nYou can increase the train set and use some successive hours for the test. In this kind of dataset ALL the clicks are important, they should consider the prediction in like 5 succ hours, it would be better.</p>",
      "rawMarkdown": "In that case,\nYou can increase the train set and use some successive hours for the test. In this kind of dataset ALL the clicks are important, they should consider the prediction in like 5 succ hours, it would be better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 314530,
      "author_name": "authman",
      "author_url": "",
      "post_date": "04/15/2018 18:48:21",
      "content": "<p>They discussed in a separate thread, but tl;dr basically it was too much data for the kernels and scorer to handle. They hope in the future as kaggle moves onto it's new platform to be able to support this sort of thing. Hence test_sup released as a Do Your Own Analysis dset.</p>\n\n<p>It was mentioned in the discussions though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 314611,
          "author_name": "oualibadr",
          "author_url": "",
          "post_date": "04/15/2018 23:53:54",
          "content": "<p>In that case,\nYou can increase the train set and use some successive hours for the test. In this kind of dataset ALL the clicks are important, they should consider the prediction in like 5 succ hours, it would be better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "314521": "I didn't realize till now that the test_sup is really important. I was using some methods that didn't work because of the test dataset released by TalkingData. It is clear that there are some huge gap between hours:\n\n    dbadmin =&gt; select hour(click_time), count(*) from AdFraudDetection_f_test group by 1 order by 1;\n     hour |  count  \n    ------+---------\n        4 | 3344125\n        5 | 2858427\n        6 |     381\n        9 | 2984808\n       10 | 3127993\n       11 |     413\n       13 | 3212566\n       14 | 3261257\n       15 |     499\n    (9 rows)\n\nHow can you release something like this? It can destroy some implementation methods. We are loosing too much information. If they want a code that they can generalize for Fraud Detection, they must be more serious... Do they want a random algorithms or something that you can use in the future?",
    "314530": "They discussed in a separate thread, but tl;dr basically it was too much data for the kernels and scorer to handle. They hope in the future as kaggle moves onto it's new platform to be able to support this sort of thing. Hence test_sup released as a Do Your Own Analysis dset.\n\nIt was mentioned in the discussions though.",
    "314611": "In that case,\nYou can increase the train set and use some successive hours for the test. In this kind of dataset ALL the clicks are important, they should consider the prediction in like 5 succ hours, it would be better."
  },
  "source": "meta"
}