{
  "id": 20878,
  "title": "What is the best machine configuration for competitions such as this one (with huge data sizes)  to be able to participate?",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20878",
  "author_name": "",
  "post_date": "2016-05-12T03:20:59.723Z",
  "votes": null,
  "comment_count": 7,
  "views": 1130,
  "content": "<p>My machine is a 8GB, 2.4Ghz, 64bit, windows machine. Not sure if it is apt for such a competition.\nOr are there alternatives, such as some techniques or online platforms that help me analyze data so huge.</p>",
  "messages": [
    {
      "id": "119650",
      "postDate": "05/12/2016 03:20:59",
      "content": "<p>My machine is a 8GB, 2.4Ghz, 64bit, windows machine. Not sure if it is apt for such a competition.\nOr are there alternatives, such as some techniques or online platforms that help me analyze data so huge.</p>",
      "rawMarkdown": "My machine is a 8GB, 2.4Ghz, 64bit, windows machine. Not sure if it is apt for such a competition.\r\nOr are there alternatives, such as some techniques or online platforms that help me analyze data so huge.",
      "votes": null
    },
    {
      "id": "119661",
      "postDate": "05/12/2016 05:02:02",
      "content": "<p>You can compete with your current machine as long as everything fits on your hard drive.</p>",
      "rawMarkdown": "You can compete with your current machine as long as everything fits on your hard drive.",
      "votes": null
    },
    {
      "id": "119665",
      "postDate": "05/12/2016 05:30:40",
      "content": "<p>The data can be handled in your current config.\nThe images just need space. You don't have to load all the images in memory anyway.</p>",
      "rawMarkdown": "The data can be handled in your current config.\r\nThe images just need space. You don't have to load all the images in memory anyway.",
      "votes": null
    },
    {
      "id": "119839",
      "postDate": "05/13/2016 02:13:43",
      "content": "<p>so, that means trusting the validation from a subset of data?</p>",
      "rawMarkdown": "so, that means trusting the validation from a subset of data?",
      "votes": null
    },
    {
      "id": "119906",
      "postDate": "05/13/2016 17:10:32",
      "content": "<p>@iLL-Logistic on this dataset a stratified validation set (i recommend 20%) should perform just as well as cross-validation, as the data is rather uniform.</p>\n\n<p>However, expect your submission results to be lower as the sets are sampled at different points in time.</p>",
      "rawMarkdown": "iLL-Logistic on this dataset a stratified validation set (i recommend 20%) should perform just as well as cross-validation, as the data is rather uniform.\r\n\r\nHowever, expect your submission results to be lower as the sets are sampled at different points in time.",
      "votes": null
    },
    {
      "id": "119910",
      "postDate": "05/13/2016 17:26:30",
      "content": "<p>[quote=iLL-Logistic;119839]</p>\n\n<p>so, that means trusting the validation from a subset of data?</p>\n\n<p>[/quote]\nThat means you don't need to load all the data into memory. You process the images in batches (that fit in your memory) and extract features from them. The resulting dataset is much smaller in size and should fit in 8GB RAM just fine. I suppose you can compete without even using the images, just working with the NLP features. You wouldn't get a competitive result in the end, though.</p>",
      "rawMarkdown": "[quote=iLL-Logistic;119839]\r\n\r\nso, that means trusting the validation from a subset of data?\r\n\r\n[/quote]\r\nThat means you don't need to load all the data into memory. You process the images in batches (that fit in your memory) and extract features from them. The resulting dataset is much smaller in size and should fit in 8GB RAM just fine. I suppose you can compete without even using the images, just working with the NLP features. You wouldn't get a competitive result in the end, though.",
      "votes": null
    },
    {
      "id": "120248",
      "postDate": "05/16/2016 18:18:27",
      "content": "<p>Let me add some information for starters, new to Python data handling</p>\n\n<p>1) You can deal with the data on your desktop</p>\n\n<p>2) Use pandas merge. E.g Say when handling the images - think how you can create information and then use merge to get the information (like hash and diff) into the same dataframe. Using proper merging (of pair of image ids and hashes and differences between hashes) will help you cut down days of processing to few minutes !.</p>\n\n<p>3) When using apply function and lambda (e.g df.apply(lambda x: f(x), axis=1)) , make sure f(x) is not accessing other large dataframes. This makes things incredibly slow.  To avoid it - use 2)</p>\n\n<p>4) If you have multiple cores, you can use multiprocessing like below    </p>\n\n<pre><code>  import multiprocessing as mp\n\n# pool = mp.Pool(initializer=init_worker, initargs=(somearg)) OR \n  pool = mp.Pool(8)\n  values = [an array of values which you want to split across the processes]\n  results = pool.map(f,values)\n</code></pre>\n\n<p>5) If you cannot load something using pd.read_csv directly - use iterator=True, error_bad_lines=False, and chunksize=1000 (e.g). </p>\n\n<pre><code> df = pd.read_csv('bigfile.csv', iterator=True,error_bad_lines=False, chunksize=1000)\n df = pd.concat(list(df), ignore_index=True)\n</code></pre>\n\n<p>Hope that helps someone</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Let me add some information for starters, new to Python data handling\r\n\r\n1) You can deal with the data on your desktop\r\n\r\n2) Use pandas merge. E.g Say when handling the images - think how you can create information and then use merge to get the information (like hash and diff) into the same dataframe. Using proper merging (of pair of image ids and hashes and differences between hashes) will help you cut down days of processing to few minutes !.\r\n\r\n3) When using apply function and lambda (e.g df.apply(lambda x: f(x), axis=1)) , make sure f(x) is not accessing other large dataframes. This makes things incredibly slow.  To avoid it - use 2)\r\n\r\n4) If you have multiple cores, you can use multiprocessing like below    \r\n\r\n      import multiprocessing as mp\r\n     \r\n    # pool = mp.Pool(initializer=init_worker, initargs=(somearg)) OR \r\n      pool = mp.Pool(8)\r\n      values = [an array of values which you want to split across the processes]\r\n      results = pool.map(f,values)\r\n\r\n5) If you cannot load something using pd.read_csv directly - use iterator=True, error_bad_lines=False, and chunksize=1000 (e.g). \r\n\r\n     df = pd.read_csv('bigfile.csv', iterator=True,error_bad_lines=False, chunksize=1000)\r\n     df = pd.concat(list(df), ignore_index=True)\r\n\r\nHope that helps someone\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "120308",
      "postDate": "05/17/2016 09:54:33",
      "content": "<p>I am using R for preprocessing the data and Python for the learning algorithms. My configuration is the same as yours (with 16GB of RAM though, 8 cores).</p>\n\n<p>R reached 12GB of memory usage for the most memory consuming parts. I am not even using the data.table package, so I think this can be reduced. Besides, preprocessing can be done on chunks of the data easily.</p>\n\n<p>On the other hand, the learning part in Python does not go beyond 4GB of memory (12 features, on the whole dataset) and XGBoost takes around 10 minutes to run (500 trees).</p>\n\n<p>So your computer (and a little bit of efforts ;) ) is more than enough for this competition !</p>",
      "rawMarkdown": "I am using R for preprocessing the data and Python for the learning algorithms. My configuration is the same as yours (with 16GB of RAM though, 8 cores).\r\n\r\nR reached 12GB of memory usage for the most memory consuming parts. I am not even using the data.table package, so I think this can be reduced. Besides, preprocessing can be done on chunks of the data easily.\r\n\r\nOn the other hand, the learning part in Python does not go beyond 4GB of memory (12 features, on the whole dataset) and XGBoost takes around 10 minutes to run (500 trees).\r\n\r\nSo your computer (and a little bit of efforts ;) ) is more than enough for this competition !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119661,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "05/12/2016 05:02:02",
      "content": "<p>You can compete with your current machine as long as everything fits on your hard drive.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119665,
      "author_name": "sonnylaskar",
      "author_url": "",
      "post_date": "05/12/2016 05:30:40",
      "content": "<p>The data can be handled in your current config.\nThe images just need space. You don't have to load all the images in memory anyway.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119839,
      "author_name": "sircausticmail",
      "author_url": "",
      "post_date": "05/13/2016 02:13:43",
      "content": "<p>so, that means trusting the validation from a subset of data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119906,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "05/13/2016 17:10:32",
      "content": "<p>@iLL-Logistic on this dataset a stratified validation set (i recommend 20%) should perform just as well as cross-validation, as the data is rather uniform.</p>\n\n<p>However, expect your submission results to be lower as the sets are sampled at different points in time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119910,
      "author_name": "fernandoprocy",
      "author_url": "",
      "post_date": "05/13/2016 17:26:30",
      "content": "<p>[quote=iLL-Logistic;119839]</p>\n\n<p>so, that means trusting the validation from a subset of data?</p>\n\n<p>[/quote]\nThat means you don't need to load all the data into memory. You process the images in batches (that fit in your memory) and extract features from them. The resulting dataset is much smaller in size and should fit in 8GB RAM just fine. I suppose you can compete without even using the images, just working with the NLP features. You wouldn't get a competitive result in the end, though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120248,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "05/16/2016 18:18:27",
      "content": "<p>Let me add some information for starters, new to Python data handling</p>\n\n<p>1) You can deal with the data on your desktop</p>\n\n<p>2) Use pandas merge. E.g Say when handling the images - think how you can create information and then use merge to get the information (like hash and diff) into the same dataframe. Using proper merging (of pair of image ids and hashes and differences between hashes) will help you cut down days of processing to few minutes !.</p>\n\n<p>3) When using apply function and lambda (e.g df.apply(lambda x: f(x), axis=1)) , make sure f(x) is not accessing other large dataframes. This makes things incredibly slow.  To avoid it - use 2)</p>\n\n<p>4) If you have multiple cores, you can use multiprocessing like below    </p>\n\n<pre><code>  import multiprocessing as mp\n\n# pool = mp.Pool(initializer=init_worker, initargs=(somearg)) OR \n  pool = mp.Pool(8)\n  values = [an array of values which you want to split across the processes]\n  results = pool.map(f,values)\n</code></pre>\n\n<p>5) If you cannot load something using pd.read_csv directly - use iterator=True, error_bad_lines=False, and chunksize=1000 (e.g). </p>\n\n<pre><code> df = pd.read_csv('bigfile.csv', iterator=True,error_bad_lines=False, chunksize=1000)\n df = pd.concat(list(df), ignore_index=True)\n</code></pre>\n\n<p>Hope that helps someone</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120308,
      "author_name": "rejulien",
      "author_url": "",
      "post_date": "05/17/2016 09:54:33",
      "content": "<p>I am using R for preprocessing the data and Python for the learning algorithms. My configuration is the same as yours (with 16GB of RAM though, 8 cores).</p>\n\n<p>R reached 12GB of memory usage for the most memory consuming parts. I am not even using the data.table package, so I think this can be reduced. Besides, preprocessing can be done on chunks of the data easily.</p>\n\n<p>On the other hand, the learning part in Python does not go beyond 4GB of memory (12 features, on the whole dataset) and XGBoost takes around 10 minutes to run (500 trees).</p>\n\n<p>So your computer (and a little bit of efforts ;) ) is more than enough for this competition !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119650": "My machine is a 8GB, 2.4Ghz, 64bit, windows machine. Not sure if it is apt for such a competition.\r\nOr are there alternatives, such as some techniques or online platforms that help me analyze data so huge.",
    "119661": "You can compete with your current machine as long as everything fits on your hard drive.",
    "119665": "The data can be handled in your current config.\r\nThe images just need space. You don't have to load all the images in memory anyway.",
    "119839": "so, that means trusting the validation from a subset of data?",
    "119906": "iLL-Logistic on this dataset a stratified validation set (i recommend 20%) should perform just as well as cross-validation, as the data is rather uniform.\r\n\r\nHowever, expect your submission results to be lower as the sets are sampled at different points in time.",
    "119910": "[quote=iLL-Logistic;119839]\r\n\r\nso, that means trusting the validation from a subset of data?\r\n\r\n[/quote]\r\nThat means you don't need to load all the data into memory. You process the images in batches (that fit in your memory) and extract features from them. The resulting dataset is much smaller in size and should fit in 8GB RAM just fine. I suppose you can compete without even using the images, just working with the NLP features. You wouldn't get a competitive result in the end, though.",
    "120248": "Let me add some information for starters, new to Python data handling\r\n\r\n1) You can deal with the data on your desktop\r\n\r\n2) Use pandas merge. E.g Say when handling the images - think how you can create information and then use merge to get the information (like hash and diff) into the same dataframe. Using proper merging (of pair of image ids and hashes and differences between hashes) will help you cut down days of processing to few minutes !.\r\n\r\n3) When using apply function and lambda (e.g df.apply(lambda x: f(x), axis=1)) , make sure f(x) is not accessing other large dataframes. This makes things incredibly slow.  To avoid it - use 2)\r\n\r\n4) If you have multiple cores, you can use multiprocessing like below    \r\n\r\n      import multiprocessing as mp\r\n     \r\n    # pool = mp.Pool(initializer=init_worker, initargs=(somearg)) OR \r\n      pool = mp.Pool(8)\r\n      values = [an array of values which you want to split across the processes]\r\n      results = pool.map(f,values)\r\n\r\n5) If you cannot load something using pd.read_csv directly - use iterator=True, error_bad_lines=False, and chunksize=1000 (e.g). \r\n\r\n     df = pd.read_csv('bigfile.csv', iterator=True,error_bad_lines=False, chunksize=1000)\r\n     df = pd.concat(list(df), ignore_index=True)\r\n\r\nHope that helps someone\r\n\r\nThanks",
    "120308": "I am using R for preprocessing the data and Python for the learning algorithms. My configuration is the same as yours (with 16GB of RAM though, 8 cores).\r\n\r\nR reached 12GB of memory usage for the most memory consuming parts. I am not even using the data.table package, so I think this can be reduced. Besides, preprocessing can be done on chunks of the data easily.\r\n\r\nOn the other hand, the learning part in Python does not go beyond 4GB of memory (12 features, on the whole dataset) and XGBoost takes around 10 minutes to run (500 trees).\r\n\r\nSo your computer (and a little bit of efforts ;) ) is more than enough for this competition !"
  },
  "source": "meta"
}