{
  "id": 53067,
  "title": "What will the private 82% test set look like?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53067",
  "author_name": "",
  "post_date": "2018-03-26T18:14:38.210009100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Attached is a screenshot of the following 3 visuals. </p>\n\n<ol>\n<li>hourly distribution of <strong>click_time</strong> on 10,000,000 randomly sampled rows from the <strong>training</strong> set.</li>\n<li>hourly distribution of <strong>click_time</strong> on 10,000,000 randomly sampled rows from <strong>test_supplement</strong>.</li>\n<li>hourly distribution of <strong>click_time</strong> on the full <strong>public test</strong> set. </li>\n</ol>\n\n<p>I am not sure if this question was asked or there is an answer to it that I didn't find, but would appreciate any thoughts on this.</p>\n\n<p>Will the private 82% test set follow the behavior of the <strong>test_supplement</strong> or <strong>test</strong>?</p>",
  "messages": [
    {
      "id": "303853",
      "postDate": "03/26/2018 18:14:38",
      "content": "<p>Attached is a screenshot of the following 3 visuals. </p>\n\n<ol>\n<li>hourly distribution of <strong>click_time</strong> on 10,000,000 randomly sampled rows from the <strong>training</strong> set.</li>\n<li>hourly distribution of <strong>click_time</strong> on 10,000,000 randomly sampled rows from <strong>test_supplement</strong>.</li>\n<li>hourly distribution of <strong>click_time</strong> on the full <strong>public test</strong> set. </li>\n</ol>\n\n<p>I am not sure if this question was asked or there is an answer to it that I didn't find, but would appreciate any thoughts on this.</p>\n\n<p>Will the private 82% test set follow the behavior of the <strong>test_supplement</strong> or <strong>test</strong>?</p>",
      "rawMarkdown": "Attached is a screenshot of the following 3 visuals. \n\n1. hourly distribution of **click_time** on 10,000,000 randomly sampled rows from the **training** set.\n2. hourly distribution of **click_time** on 10,000,000 randomly sampled rows from **test_supplement**.\n3. hourly distribution of **click_time** on the full **public test** set. \n\nI am not sure if this question was asked or there is an answer to it that I didn't find, but would appreciate any thoughts on this.\n\nWill the private 82% test set follow the behavior of the **test_supplement** or **test**?",
      "votes": null
    },
    {
      "id": "303896",
      "postDate": "03/26/2018 18:43:37",
      "content": "<p>The 82% is a subset of the <strong>test</strong> set (i.e. not from the <strong>test_supplement</strong>), so it will have the distribution you presented in your image (#3) minus the 18% of records used for <strong>public</strong> scoring. I haven't seen it mentioned anywhere whether the split is random, time-based, or based on some other feature, however, which might cause <strong>public</strong> and <strong>private</strong> to have different distributions.</p>",
      "rawMarkdown": "The 82% is a subset of the **test** set (i.e. not from the **test_supplement**), so it will have the distribution you presented in your image (#3) minus the 18% of records used for **public** scoring. I haven't seen it mentioned anywhere whether the split is random, time-based, or based on some other feature, however, which might cause **public** and **private** to have different distributions.",
      "votes": null
    },
    {
      "id": "303905",
      "postDate": "03/26/2018 18:54:58",
      "content": "<p>Thanks a lot for the clarification. I somehow assumed that the test set available to us is 18% of the original test and the remaining 82% is not available to us. </p>\n\n<p>Now this discussion makes sense: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833</a></p>\n\n<p>Due to it the evaluation on the public leaderboard is based on the first hour (i.e. the hour 4) of the test set. Still, have to confirm this myself though.</p>",
      "rawMarkdown": "Thanks a lot for the clarification. I somehow assumed that the test set available to us is 18% of the original test and the remaining 82% is not available to us. \n\nNow this discussion makes sense: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833\n\nDue to it the evaluation on the public leaderboard is based on the first hour (i.e. the hour 4) of the test set. Still, have to confirm this myself though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 303896,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "03/26/2018 18:43:37",
      "content": "<p>The 82% is a subset of the <strong>test</strong> set (i.e. not from the <strong>test_supplement</strong>), so it will have the distribution you presented in your image (#3) minus the 18% of records used for <strong>public</strong> scoring. I haven't seen it mentioned anywhere whether the split is random, time-based, or based on some other feature, however, which might cause <strong>public</strong> and <strong>private</strong> to have different distributions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 303905,
          "author_name": "araksstepanyan",
          "author_url": "",
          "post_date": "03/26/2018 18:54:58",
          "content": "<p>Thanks a lot for the clarification. I somehow assumed that the test set available to us is 18% of the original test and the remaining 82% is not available to us. </p>\n\n<p>Now this discussion makes sense: <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833</a></p>\n\n<p>Due to it the evaluation on the public leaderboard is based on the first hour (i.e. the hour 4) of the test set. Still, have to confirm this myself though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "303853": "Attached is a screenshot of the following 3 visuals. \n\n1. hourly distribution of **click_time** on 10,000,000 randomly sampled rows from the **training** set.\n2. hourly distribution of **click_time** on 10,000,000 randomly sampled rows from **test_supplement**.\n3. hourly distribution of **click_time** on the full **public test** set. \n\nI am not sure if this question was asked or there is an answer to it that I didn't find, but would appreciate any thoughts on this.\n\nWill the private 82% test set follow the behavior of the **test_supplement** or **test**?",
    "303896": "The 82% is a subset of the **test** set (i.e. not from the **test_supplement**), so it will have the distribution you presented in your image (#3) minus the 18% of records used for **public** scoring. I haven't seen it mentioned anywhere whether the split is random, time-based, or based on some other feature, however, which might cause **public** and **private** to have different distributions.",
    "303905": "Thanks a lot for the clarification. I somehow assumed that the test set available to us is 18% of the original test and the remaining 82% is not available to us. \n\nNow this discussion makes sense: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52833\n\nDue to it the evaluation on the public leaderboard is based on the first hour (i.e. the hour 4) of the test set. Still, have to confirm this myself though."
  },
  "source": "meta"
}