{
  "id": 24107,
  "title": "Welcome!",
  "url": "/competitions/outbrain-click-prediction/discussion/24107",
  "author_name": "",
  "post_date": "2016-10-05T17:34:28.723Z",
  "votes": 17,
  "comment_count": 9,
  "views": 3803,
  "content": "<p>Welcome to Outbrain's Click Prediction competition! A few words to kick off:</p>\n\n<ul>\n<li>This is a very large dataset (100GB+), but you can get started on the\nproblem without the full set. Most of the size is contained in the\nuser-level page_views table. It's still possible to make predictions without using this table if you don't have the computing horsepower.</li>\n<li>Because of the data/computing limits on Kernels, the Kernels container only has the page views sample at the moment. </li>\n<li>Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our <a href=\"https://www.kaggle.com/wiki/ANoteOnTorrents\">note on torrents</a> before you suggest using torrents.</li>\n<li>You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the &quot;time machine rule&quot;, it was an intentional tradeoff. You may use all the available data to make predictions.</li>\n</ul>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": "137861",
      "postDate": "10/05/2016 17:34:28",
      "content": "<p>Welcome to Outbrain's Click Prediction competition! A few words to kick off:</p>\n\n<ul>\n<li>This is a very large dataset (100GB+), but you can get started on the\nproblem without the full set. Most of the size is contained in the\nuser-level page_views table. It's still possible to make predictions without using this table if you don't have the computing horsepower.</li>\n<li>Because of the data/computing limits on Kernels, the Kernels container only has the page views sample at the moment. </li>\n<li>Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our <a href=\"https://www.kaggle.com/wiki/ANoteOnTorrents\">note on torrents</a> before you suggest using torrents.</li>\n<li>You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the &quot;time machine rule&quot;, it was an intentional tradeoff. You may use all the available data to make predictions.</li>\n</ul>\n\n<p>Good luck!</p>",
      "rawMarkdown": "Welcome to Outbrain's Click Prediction competition! A few words to kick off:\r\n\r\n - This is a very large dataset (100GB+), but you can get started on the\r\n   problem without the full set. Most of the size is contained in the\r\n   user-level page_views table. It's still possible to make predictions without using this table if you don't have the computing horsepower.\r\n - Because of the data/computing limits on Kernels, the Kernels container only has the page views sample at the moment. \r\n - Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our [note on torrents][1] before you suggest using torrents.\r\n - You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the \"time machine rule\", it was an intentional tradeoff. You may use all the available data to make predictions.\r\n\r\nGood luck!\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/ANoteOnTorrents",
      "votes": null
    },
    {
      "id": "137865",
      "postDate": "10/05/2016 18:15:26",
      "content": "<p>[quote=William Cukierski;137861]</p>\n\n<ul>\n<li>You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the &quot;time machine rule&quot;, it was an intentional tradeoff. You may use all the available data to make predictions.</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>I just love this part! Please make it standard for future competitions :D</p>",
      "rawMarkdown": "[quote=William Cukierski;137861]\r\n\r\n - You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the \"time machine rule\", it was an intentional tradeoff. You may use all the available data to make predictions.\r\n\r\n[/quote]\r\n\r\nI just love this part! Please make it standard for future competitions :D",
      "votes": null
    },
    {
      "id": "137878",
      "postDate": "10/05/2016 20:38:01",
      "content": "<p>Wow, that's a lot of data. Does Kaggle have a secret conspiracy with RAM manufacturers? :P</p>",
      "rawMarkdown": "Wow, that's a lot of data. Does Kaggle have a secret conspiracy with RAM manufacturers? :P",
      "votes": null
    },
    {
      "id": "137889",
      "postDate": "10/05/2016 21:14:24",
      "content": "<p>You could split the big file in parts (say, 10). This would make it much easier to download it.</p>",
      "rawMarkdown": "You could split the big file in parts (say, 10). This would make it much easier to download it.",
      "votes": null
    },
    {
      "id": "137894",
      "postDate": "10/05/2016 21:24:36",
      "content": "<p>@anokas - and maybe intel ;)  Skylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32.  It's not like they've added enough CPU power to justify the upgrade...</p>\n\n<p>Also the road to hell, er, leaks, is paved by random non-time splits...</p>",
      "rawMarkdown": "anokas - and maybe intel ;)  Skylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32.  It's not like they've added enough CPU power to justify the upgrade...\r\n\r\nAlso the road to hell, er, leaks, is paved by random non-time splits...",
      "votes": null
    },
    {
      "id": "137897",
      "postDate": "10/05/2016 21:40:33",
      "content": "<p>[quote=happycube;137894]\nSkylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32. \n[/quote]</p>\n\n<p>The <a href=\"http://ark.intel.com/products/63698\">i7-3820</a> from +4 years ago seems to disagree with that :P</p>",
      "rawMarkdown": "[quote=happycube;137894]\r\nSkylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32. \r\n[/quote]\r\n\r\nThe [i7-3820][1] from +4 years ago seems to disagree with that :P\r\n\r\n\r\n  [1]: http://ark.intel.com/products/63698",
      "votes": null
    },
    {
      "id": "137903",
      "postDate": "10/05/2016 22:14:49",
      "content": "<p>I meant regular desktop series, not 4-memory-channel Xeons in sheep's clothing ;)</p>\n\n<p>A $300 HEDT Skylake-E would be interesting if intel bothers to make one, it might have AVX512...</p>",
      "rawMarkdown": "I meant regular desktop series, not 4-memory-channel Xeons in sheep's clothing ;)\r\n\r\nA $300 HEDT Skylake-E would be interesting if intel bothers to make one, it might have AVX512...",
      "votes": null
    },
    {
      "id": "137906",
      "postDate": "10/05/2016 22:57:40",
      "content": "<p>[quote=NxGTR;137897]</p>\n\n<p>The <a href=\"http://ark.intel.com/products/63698\">i7-3820</a> from +4 years ago seems to disagree with that :P</p>\n\n<p>[/quote]</p>\n\n<p>In the details, it's a Sandy Bridge-E :p</p>",
      "rawMarkdown": "[quote=NxGTR;137897]\r\n\r\nThe [i7-3820][1] from +4 years ago seems to disagree with that :P\r\n\r\n  [1]: http://ark.intel.com/products/63698\r\n\r\n[/quote]\r\n\r\nIn the details, it's a Sandy Bridge-E :p",
      "votes": null
    },
    {
      "id": "137939",
      "postDate": "10/06/2016 05:32:25",
      "content": "<p>I have a Xeon &quot;laptop&quot; with 64GB RAM. Ironically, ever since I bought it, I haven't taken part in any Kaggle competitions.</p>",
      "rawMarkdown": "I have a Xeon \"laptop\" with 64GB RAM. Ironically, ever since I bought it, I haven't taken part in any Kaggle competitions.",
      "votes": null
    },
    {
      "id": "138610",
      "postDate": "10/09/2016 19:12:18",
      "content": "<p>Is the clicks_test file already in order of decreasing ad click probability or not ?</p>\n\n<p>Also, how do we compute the Mean Average Precision @12  score with respect to a test set ? That is, would we simply hold out a portion of the clicks_train and use that for testing ?</p>",
      "rawMarkdown": "Is the clicks_test file already in order of decreasing ad click probability or not ?\r\n\r\nAlso, how do we compute the Mean Average Precision @12  score with respect to a test set ? That is, would we simply hold out a portion of the clicks_train and use that for testing ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 137865,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "10/05/2016 18:15:26",
      "content": "<p>[quote=William Cukierski;137861]</p>\n\n<ul>\n<li>You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the &quot;time machine rule&quot;, it was an intentional tradeoff. You may use all the available data to make predictions.</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>I just love this part! Please make it standard for future competitions :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137878,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "10/05/2016 20:38:01",
      "content": "<p>Wow, that's a lot of data. Does Kaggle have a secret conspiracy with RAM manufacturers? :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137889,
      "author_name": "alantus",
      "author_url": "",
      "post_date": "10/05/2016 21:14:24",
      "content": "<p>You could split the big file in parts (say, 10). This would make it much easier to download it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137894,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "10/05/2016 21:24:36",
      "content": "<p>@anokas - and maybe intel ;)  Skylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32.  It's not like they've added enough CPU power to justify the upgrade...</p>\n\n<p>Also the road to hell, er, leaks, is paved by random non-time splits...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137897,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "10/05/2016 21:40:33",
      "content": "<p>[quote=happycube;137894]\nSkylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32. \n[/quote]</p>\n\n<p>The <a href=\"http://ark.intel.com/products/63698\">i7-3820</a> from +4 years ago seems to disagree with that :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137903,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "10/05/2016 22:14:49",
      "content": "<p>I meant regular desktop series, not 4-memory-channel Xeons in sheep's clothing ;)</p>\n\n<p>A $300 HEDT Skylake-E would be interesting if intel bothers to make one, it might have AVX512...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137906,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "10/05/2016 22:57:40",
      "content": "<p>[quote=NxGTR;137897]</p>\n\n<p>The <a href=\"http://ark.intel.com/products/63698\">i7-3820</a> from +4 years ago seems to disagree with that :P</p>\n\n<p>[/quote]</p>\n\n<p>In the details, it's a Sandy Bridge-E :p</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137939,
      "author_name": "innerproduct",
      "author_url": "",
      "post_date": "10/06/2016 05:32:25",
      "content": "<p>I have a Xeon &quot;laptop&quot; with 64GB RAM. Ironically, ever since I bought it, I haven't taken part in any Kaggle competitions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 138610,
      "author_name": "samwhitehill",
      "author_url": "",
      "post_date": "10/09/2016 19:12:18",
      "content": "<p>Is the clicks_test file already in order of decreasing ad click probability or not ?</p>\n\n<p>Also, how do we compute the Mean Average Precision @12  score with respect to a test set ? That is, would we simply hold out a portion of the clicks_train and use that for testing ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137861": "Welcome to Outbrain's Click Prediction competition! A few words to kick off:\r\n\r\n - This is a very large dataset (100GB+), but you can get started on the\r\n   problem without the full set. Most of the size is contained in the\r\n   user-level page_views table. It's still possible to make predictions without using this table if you don't have the computing horsepower.\r\n - Because of the data/computing limits on Kernels, the Kernels container only has the page views sample at the moment. \r\n - Expect a traffic jam to download the 30 GB file during the launch gold rush (sorry!). Use a download manager with resuming capabilities and we promise you'll eventually get the file. There is also a small sample available. Please read our [note on torrents][1] before you suggest using torrents.\r\n - You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the \"time machine rule\", it was an intentional tradeoff. You may use all the available data to make predictions.\r\n\r\nGood luck!\r\n\r\n\r\n  [1]: https://www.kaggle.com/wiki/ANoteOnTorrents",
    "137865": "[quote=William Cukierski;137861]\r\n\r\n - You'll note that this problem has a time component, but the split is not purely time based. Outbrain has intentionally split the train/test to cover the time period in their desired manner. The public/private split is uniformly random within the test set. While this technically violates the \"time machine rule\", it was an intentional tradeoff. You may use all the available data to make predictions.\r\n\r\n[/quote]\r\n\r\nI just love this part! Please make it standard for future competitions :D",
    "137878": "Wow, that's a lot of data. Does Kaggle have a secret conspiracy with RAM manufacturers? :P",
    "137889": "You could split the big file in parts (say, 10). This would make it much easier to download it.",
    "137894": "anokas - and maybe intel ;)  Skylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32.  It's not like they've added enough CPU power to justify the upgrade...\r\n\r\nAlso the road to hell, er, leaks, is paved by random non-time splits...",
    "137897": "[quote=happycube;137894]\r\nSkylake (core 6xxx) is the first regular desktop series that can handle 64GB, whereas Sandy Bridge to Broadwell can do 32. \r\n[/quote]\r\n\r\nThe [i7-3820][1] from +4 years ago seems to disagree with that :P\r\n\r\n\r\n  [1]: http://ark.intel.com/products/63698",
    "137903": "I meant regular desktop series, not 4-memory-channel Xeons in sheep's clothing ;)\r\n\r\nA $300 HEDT Skylake-E would be interesting if intel bothers to make one, it might have AVX512...",
    "137906": "[quote=NxGTR;137897]\r\n\r\nThe [i7-3820][1] from +4 years ago seems to disagree with that :P\r\n\r\n  [1]: http://ark.intel.com/products/63698\r\n\r\n[/quote]\r\n\r\nIn the details, it's a Sandy Bridge-E :p",
    "137939": "I have a Xeon \"laptop\" with 64GB RAM. Ironically, ever since I bought it, I haven't taken part in any Kaggle competitions.",
    "138610": "Is the clicks_test file already in order of decreasing ad click probability or not ?\r\n\r\nAlso, how do we compute the Mean Average Precision @12  score with respect to a test set ? That is, would we simply hold out a portion of the clicks_train and use that for testing ?"
  },
  "source": "meta"
}