{
  "id": 56271,
  "title": "Let us talk about data size",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56271",
  "author_name": "Scirpus",
  "post_date": "2018-05-08T06:59:19.209000",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>With the controversy with Dirk releasing his late kernel that annoyed so many people - why did it have such an impact compared to other late submissions?</p>\n\n<p>I believe strongly the impact was correlated with data size.</p>\n\n<p>2 years ago I was competitive with 6gb of RAM (Bronze)\n1 year ago I was competitive with 24Gb of RAM (Silver)\nNow I use 128 Gb of RAM (Silver)</p>\n\n<p>It was very easy to do well in this competition with a huge amount of memory with very little effort just brute force on all the data (minus day 6 which was dodgy).  Dirk ran his on a large memory machine and as a result got a great score.</p>\n\n<p>Kaggle need to start thinking about how much data they want us to work on - obviously most people have laptops with 16gb RAM which would be a major hindrance in doing well in this competition.</p>",
  "messages": [
    {
      "id": 325141,
      "postDate": "2018-05-08T06:59:19.210Z",
      "content": "<p>With the controversy with Dirk releasing his late kernel that annoyed so many people - why did it have such an impact compared to other late submissions?</p>\n\n<p>I believe strongly the impact was correlated with data size.</p>\n\n<p>2 years ago I was competitive with 6gb of RAM (Bronze)\n1 year ago I was competitive with 24Gb of RAM (Silver)\nNow I use 128 Gb of RAM (Silver)</p>\n\n<p>It was very easy to do well in this competition with a huge amount of memory with very little effort just brute force on all the data (minus day 6 which was dodgy).  Dirk ran his on a large memory machine and as a result got a great score.</p>\n\n<p>Kaggle need to start thinking about how much data they want us to work on - obviously most people have laptops with 16gb RAM which would be a major hindrance in doing well in this competition.</p>",
      "rawMarkdown": "With the controversy with Dirk releasing his late kernel that annoyed so many people - why did it have such an impact compared to other late submissions?\n\nI believe strongly the impact was correlated with data size.\n\n2 years ago I was competitive with 6gb of RAM (Bronze)\n1 year ago I was competitive with 24Gb of RAM (Silver)\nNow I use 128 Gb of RAM (Silver)\n\nIt was very easy to do well in this competition with a huge amount of memory with very little effort just brute force on all the data (minus day 6 which was dodgy).  Dirk ran his on a large memory machine and as a result got a great score.\n\nKaggle need to start thinking about how much data they want us to work on - obviously most people have laptops with 16gb RAM which would be a major hindrance in doing well in this competition.\n",
      "votes": 15
    },
    {
      "id": 325160,
      "postDate": "2018-05-08T07:20:21.477Z",
      "content": "<p>I find this kind of topics about data size and memory / resources to all fall under the category ... \"Have you tried figuring out a way of lazy loading the data to limit the memory usage?\" I agree, having all the data in memory has a speed-up in training ... But, we have a saying in Romania that translates into ... \"Cut your coat according to your cloth\" </p>\n\n<p>I understand that everybody wants to use out-of-the-box solutions to win competitions ... I honestly think that figuring out solutions for this kind of problems (talking about resource usage) is more of a value to participants than winning by using a set of libraries and costly resources. </p>\n\n<p>My personal opinion is there are some interests in having \"resource starving\" out-of-the-box solution as it would sell \"cloud computing\". This doesn't mean there are no alternative solutions ... </p>\n\n<p>The more labeled data you have, the better your model can get so I would expect the future to present even bigger sets of data for competitions. Exploration of alternative ways of loading the data and altering the algorithms to be memory efficient will provide an edge for getting ahead in this kind of competitions. (or having a access to resources :-))</p>\n\n<p>If we refer to the popular \"LightGBM\" ... I see it has \"Optimization in Parallel Learning\" (<a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel\">https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel</a>) that supports splitting data in chunks ... obviously for parallel processing ... But, it means one could create a variant that instead of doing \"parallel\" it emulates batching ... because the same logic can apply. </p>\n\n<p>This:</p>\n\n<ol>\n<li>Partition data horizontally</li>\n<li>Workers use local data to construct local histograms</li>\n<li>Merge global histograms from all local histograms</li>\n<li>Find best split from merged global histograms, then perform splits</li>\n</ol>\n\n<p>could be translated into</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Iterate through <em>partition data</em> to construct <em>partition</em> histograms</p></li>\n<li><p>Merge global histograms from all <em>partition</em> histograms</p></li>\n<li><p>Find best split from merged global histograms, then perform splits</p></li>\n</ol>",
      "rawMarkdown": "I find this kind of topics about data size and memory / resources to all fall under the category ... \"Have you tried figuring out a way of lazy loading the data to limit the memory usage?\" I agree, having all the data in memory has a speed-up in training ... But, we have a saying in Romania that translates into ... \"Cut your coat according to your cloth\" \n\nI understand that everybody wants to use out-of-the-box solutions to win competitions ... I honestly think that figuring out solutions for this kind of problems (talking about resource usage) is more of a value to participants than winning by using a set of libraries and costly resources. \n\nMy personal opinion is there are some interests in having \"resource starving\" out-of-the-box solution as it would sell \"cloud computing\". This doesn't mean there are no alternative solutions ... \n\nThe more labeled data you have, the better your model can get so I would expect the future to present even bigger sets of data for competitions. Exploration of alternative ways of loading the data and altering the algorithms to be memory efficient will provide an edge for getting ahead in this kind of competitions. (or having a access to resources :-))\n\nIf we refer to the popular \"LightGBM\" ... I see it has \"Optimization in Parallel Learning\" (https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel) that supports splitting data in chunks ... obviously for parallel processing ... But, it means one could create a variant that instead of doing \"parallel\" it emulates batching ... because the same logic can apply. \n\nThis:\n\n1. Partition data horizontally\n2. Workers use local data to construct local histograms\n3. Merge global histograms from all local histograms\n4. Find best split from merged global histograms, then perform splits\n\ncould be translated into\n\n1. Partition data horizontally\n\n2. Iterate through _partition data_ to construct _partition_ histograms\n\n3. Merge global histograms from all _partition_ histograms\n\n4. Find best split from merged global histograms, then perform splits\n\n",
      "votes": 4,
      "replies": [
        {
          "id": 325173,
          "postDate": "2018-05-08T07:31:01.083Z",
          "content": "<p>My thoughts were aimed at why this kernel had such an impact - I am not saying you cannot do well by being clever just that huge datasets give people with huge memory an enormous advantage.  If someone releases the submission file after using the entire data especially on the last day this will make people with 16gb relatively uncompetitive. </p>",
          "rawMarkdown": "My thoughts were aimed at why this kernel had such an impact - I am not saying you cannot do well by being clever just that huge datasets give people with huge memory an enormous advantage.  If someone releases the submission file after using the entire data especially on the last day this will make people with 16gb relatively uncompetitive. ",
          "votes": 1
        },
        {
          "id": 325179,
          "postDate": "2018-05-08T07:41:00.993Z",
          "content": "<p>You state that the impact is because of data size ... I'm pointing out that there are ways around data size ... Not out of the box ways ... but ways ... </p>\n\n<p>The impact of the .9811 kernel to \"data size\"? .... 0 ... it had the submission file associated ... it didn't require to run to see the results ... </p>",
          "rawMarkdown": "You state that the impact is because of data size ... I'm pointing out that there are ways around data size ... Not out of the box ways ... but ways ... \n\nThe impact of the .9811 kernel to \"data size\"? .... 0 ... it had the submission file associated ... it didn't require to run to see the results ... "
        },
        {
          "id": 325184,
          "postDate": "2018-05-08T07:45:38.933Z",
          "content": "<p>Precisely that is why I stated \"If someone releases the submission file\" ;)</p>",
          "rawMarkdown": "Precisely that is why I stated \"If someone releases the submission file\" ;)"
        }
      ]
    },
    {
      "id": 325178,
      "postDate": "2018-05-08T07:35:53.993Z",
      "content": "<p>Aye. Besides, I think those who have big RAMs could do more trials than those who don't have. Sometimes you have to wait for a result to consider what to do next, which means you might only have 2 or 3 chance per day. So if you work with 8G or 16G RAM with such big dataset, it will definitely be painful. We need impressive competition, but we also need fairness in some way.</p>",
      "rawMarkdown": "Aye. Besides, I think those who have big RAMs could do more trials than those who don't have. Sometimes you have to wait for a result to consider what to do next, which means you might only have 2 or 3 chance per day. So if you work with 8G or 16G RAM with such big dataset, it will definitely be painful. We need impressive competition, but we also need fairness in some way.",
      "votes": 1
    },
    {
      "id": 338495,
      "postDate": "2018-06-05T07:14:25.623Z",
      "content": "<p>Maybe set another kernel leaderboard would solve this problem.</p>",
      "rawMarkdown": "Maybe set another kernel leaderboard would solve this problem."
    },
    {
      "id": 326484,
      "postDate": "2018-05-09T19:58:10.520Z",
      "content": "<p>Very good point! I also think a smaller dataset should make this condition better.</p>",
      "rawMarkdown": "Very good point! I also think a smaller dataset should make this condition better."
    },
    {
      "id": 326012,
      "postDate": "2018-05-09T06:53:14.547Z",
      "content": "<p>One way to solve the problem and put all people into same conditions is to force to use only kaggle kernels</p>",
      "rawMarkdown": "One way to solve the problem and put all people into same conditions is to force to use only kaggle kernels"
    },
    {
      "id": 325149,
      "postDate": "2018-05-08T07:09:31.750Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 325160,
      "author_name": "Mihai Cvasnievschi",
      "author_url": "",
      "post_date": "2018-05-08T07:20:21.477000",
      "content": "<p>I find this kind of topics about data size and memory / resources to all fall under the category ... \"Have you tried figuring out a way of lazy loading the data to limit the memory usage?\" I agree, having all the data in memory has a speed-up in training ... But, we have a saying in Romania that translates into ... \"Cut your coat according to your cloth\" </p>\n\n<p>I understand that everybody wants to use out-of-the-box solutions to win competitions ... I honestly think that figuring out solutions for this kind of problems (talking about resource usage) is more of a value to participants than winning by using a set of libraries and costly resources. </p>\n\n<p>My personal opinion is there are some interests in having \"resource starving\" out-of-the-box solution as it would sell \"cloud computing\". This doesn't mean there are no alternative solutions ... </p>\n\n<p>The more labeled data you have, the better your model can get so I would expect the future to present even bigger sets of data for competitions. Exploration of alternative ways of loading the data and altering the algorithms to be memory efficient will provide an edge for getting ahead in this kind of competitions. (or having a access to resources :-))</p>\n\n<p>If we refer to the popular \"LightGBM\" ... I see it has \"Optimization in Parallel Learning\" (<a href=\"https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel\">https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel</a>) that supports splitting data in chunks ... obviously for parallel processing ... But, it means one could create a variant that instead of doing \"parallel\" it emulates batching ... because the same logic can apply. </p>\n\n<p>This:</p>\n\n<ol>\n<li>Partition data horizontally</li>\n<li>Workers use local data to construct local histograms</li>\n<li>Merge global histograms from all local histograms</li>\n<li>Find best split from merged global histograms, then perform splits</li>\n</ol>\n\n<p>could be translated into</p>\n\n<ol>\n<li><p>Partition data horizontally</p></li>\n<li><p>Iterate through <em>partition data</em> to construct <em>partition</em> histograms</p></li>\n<li><p>Merge global histograms from all <em>partition</em> histograms</p></li>\n<li><p>Find best split from merged global histograms, then perform splits</p></li>\n</ol>",
      "votes": 4,
      "replies": [
        {
          "id": 325173,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-05-08T07:31:01.083000",
          "content": "<p>My thoughts were aimed at why this kernel had such an impact - I am not saying you cannot do well by being clever just that huge datasets give people with huge memory an enormous advantage.  If someone releases the submission file after using the entire data especially on the last day this will make people with 16gb relatively uncompetitive. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 325179,
          "author_name": "Mihai Cvasnievschi",
          "author_url": "",
          "post_date": "2018-05-08T07:41:00.993000",
          "content": "<p>You state that the impact is because of data size ... I'm pointing out that there are ways around data size ... Not out of the box ways ... but ways ... </p>\n\n<p>The impact of the .9811 kernel to \"data size\"? .... 0 ... it had the submission file associated ... it didn't require to run to see the results ... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325184,
          "author_name": "Scirpus",
          "author_url": "",
          "post_date": "2018-05-08T07:45:38.933000",
          "content": "<p>Precisely that is why I stated \"If someone releases the submission file\" ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325178,
      "author_name": "Laevatein",
      "author_url": "",
      "post_date": "2018-05-08T07:35:53.993000",
      "content": "<p>Aye. Besides, I think those who have big RAMs could do more trials than those who don't have. Sometimes you have to wait for a result to consider what to do next, which means you might only have 2 or 3 chance per day. So if you work with 8G or 16G RAM with such big dataset, it will definitely be painful. We need impressive competition, but we also need fairness in some way.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 338495,
      "author_name": "interstellar",
      "author_url": "",
      "post_date": "2018-06-05T07:14:25.623000",
      "content": "<p>Maybe set another kernel leaderboard would solve this problem.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326484,
      "author_name": "Pengyue Wang",
      "author_url": "",
      "post_date": "2018-05-09T19:58:10.520000",
      "content": "<p>Very good point! I also think a smaller dataset should make this condition better.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 326012,
      "author_name": "Insaf Ashrapov",
      "author_url": "",
      "post_date": "2018-05-09T06:53:14.547000",
      "content": "<p>One way to solve the problem and put all people into same conditions is to force to use only kaggle kernels</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325149,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T07:09:31.750000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325141": "With the controversy with Dirk releasing his late kernel that annoyed so many people - why did it have such an impact compared to other late submissions?\n\nI believe strongly the impact was correlated with data size.\n\n2 years ago I was competitive with 6gb of RAM (Bronze)\n1 year ago I was competitive with 24Gb of RAM (Silver)\nNow I use 128 Gb of RAM (Silver)\n\nIt was very easy to do well in this competition with a huge amount of memory with very little effort just brute force on all the data (minus day 6 which was dodgy).  Dirk ran his on a large memory machine and as a result got a great score.\n\nKaggle need to start thinking about how much data they want us to work on - obviously most people have laptops with 16gb RAM which would be a major hindrance in doing well in this competition.\n",
    "325160": "I find this kind of topics about data size and memory / resources to all fall under the category ... \"Have you tried figuring out a way of lazy loading the data to limit the memory usage?\" I agree, having all the data in memory has a speed-up in training ... But, we have a saying in Romania that translates into ... \"Cut your coat according to your cloth\" \n\nI understand that everybody wants to use out-of-the-box solutions to win competitions ... I honestly think that figuring out solutions for this kind of problems (talking about resource usage) is more of a value to participants than winning by using a set of libraries and costly resources. \n\nMy personal opinion is there are some interests in having \"resource starving\" out-of-the-box solution as it would sell \"cloud computing\". This doesn't mean there are no alternative solutions ... \n\nThe more labeled data you have, the better your model can get so I would expect the future to present even bigger sets of data for competitions. Exploration of alternative ways of loading the data and altering the algorithms to be memory efficient will provide an edge for getting ahead in this kind of competitions. (or having a access to resources :-))\n\nIf we refer to the popular \"LightGBM\" ... I see it has \"Optimization in Parallel Learning\" (https://github.com/Microsoft/LightGBM/blob/master/docs/Features.rst#data-parallel) that supports splitting data in chunks ... obviously for parallel processing ... But, it means one could create a variant that instead of doing \"parallel\" it emulates batching ... because the same logic can apply. \n\nThis:\n\n1. Partition data horizontally\n2. Workers use local data to construct local histograms\n3. Merge global histograms from all local histograms\n4. Find best split from merged global histograms, then perform splits\n\ncould be translated into\n\n1. Partition data horizontally\n\n2. Iterate through _partition data_ to construct _partition_ histograms\n\n3. Merge global histograms from all _partition_ histograms\n\n4. Find best split from merged global histograms, then perform splits\n\n",
    "325178": "Aye. Besides, I think those who have big RAMs could do more trials than those who don't have. Sometimes you have to wait for a result to consider what to do next, which means you might only have 2 or 3 chance per day. So if you work with 8G or 16G RAM with such big dataset, it will definitely be painful. We need impressive competition, but we also need fairness in some way.",
    "338495": "Maybe set another kernel leaderboard would solve this problem.",
    "326484": "Very good point! I also think a smaller dataset should make this condition better.",
    "326012": "One way to solve the problem and put all people into same conditions is to force to use only kaggle kernels",
    "325149": ""
  }
}