{
  "id": 41193,
  "title": "Loading Data blazing fast with MongoDB (<20ms~)",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/41193",
  "author_name": "",
  "post_date": "2017-10-14T04:40:38.584386500Z",
  "votes": 23,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Loading data was always a problem for my workstation, HDD + 16Gb of Ram was usually not enough for big databases, specially this one.</p>\n\n<p>Leveraging my code with MongoDB, I was able to get loading speeds from 15ms~ for the first hit to 1.7ms~ on the second hit while using less then 8Gb of ram. This was enough to keep my GTX 1070 feed to 100%~.</p>\n\n<p>So let's get started with mongoDB!\nFirst you will need to load the *.bson files into MongoDB, assuming that you already installed Mongo.\nRun this comands on your console:</p>\n\n<pre><code>mongorestore -d Cdiscount -c test test.bson\nmongorestore -d Cdiscount -c train train.bson\n</code></pre>\n\n<p>You can check if the data was loaded correctly with:</p>\n\n<pre><code>mongo\nshow dbs\n</code></pre>\n\n<p>Should have some like this:\n<img src=\"https://imgur.com/ltESIzF.png\" alt=\"\" title=\"\"></p>\n\n<p>Great! Now we have to access that with python and Pymongo!</p>\n\n<pre><code>import io\nimport pymongo\nfrom pymongo import MongoClient\n\nclient = MongoClient(connect=False) #Makes it \"good enough\" for our multi-threaded use case. \ntrain = client.Cdiscount['train']\ntest = client.Cdiscount['test']\n\nindex = 0 #Witch idx Do you want to load?\ndata = train.find_one({'_id': index})\nimg = imread(io.BytesIO(data['imgs'][0]['picture'])) Get the first picture from data\nlabel = data['category_id']\n</code></pre>\n\n<p>Pytorch Data-loader, first hit:\n<img src=\"https://imgur.com/9I3jH20.png\" alt=\"\" title=\"\"></p>\n\n<p>Second hit (on adjacent idx):\n<img src=\"https://imgur.com/f1WQIll.png\" alt=\"\" title=\"\"></p>\n\n<p>Pytorch-MongoDB Starter kit coming soon =D</p>",
  "messages": [
    {
      "id": "231247",
      "postDate": "10/14/2017 04:40:38",
      "content": "<p>Loading data was always a problem for my workstation, HDD + 16Gb of Ram was usually not enough for big databases, specially this one.</p>\n\n<p>Leveraging my code with MongoDB, I was able to get loading speeds from 15ms~ for the first hit to 1.7ms~ on the second hit while using less then 8Gb of ram. This was enough to keep my GTX 1070 feed to 100%~.</p>\n\n<p>So let's get started with mongoDB!\nFirst you will need to load the *.bson files into MongoDB, assuming that you already installed Mongo.\nRun this comands on your console:</p>\n\n<pre><code>mongorestore -d Cdiscount -c test test.bson\nmongorestore -d Cdiscount -c train train.bson\n</code></pre>\n\n<p>You can check if the data was loaded correctly with:</p>\n\n<pre><code>mongo\nshow dbs\n</code></pre>\n\n<p>Should have some like this:\n<img src=\"https://imgur.com/ltESIzF.png\" alt=\"\" title=\"\"></p>\n\n<p>Great! Now we have to access that with python and Pymongo!</p>\n\n<pre><code>import io\nimport pymongo\nfrom pymongo import MongoClient\n\nclient = MongoClient(connect=False) #Makes it \"good enough\" for our multi-threaded use case. \ntrain = client.Cdiscount['train']\ntest = client.Cdiscount['test']\n\nindex = 0 #Witch idx Do you want to load?\ndata = train.find_one({'_id': index})\nimg = imread(io.BytesIO(data['imgs'][0]['picture'])) Get the first picture from data\nlabel = data['category_id']\n</code></pre>\n\n<p>Pytorch Data-loader, first hit:\n<img src=\"https://imgur.com/9I3jH20.png\" alt=\"\" title=\"\"></p>\n\n<p>Second hit (on adjacent idx):\n<img src=\"https://imgur.com/f1WQIll.png\" alt=\"\" title=\"\"></p>\n\n<p>Pytorch-MongoDB Starter kit coming soon =D</p>",
      "rawMarkdown": "Loading data was always a problem for my workstation, HDD + 16Gb of Ram was usually not enough for big databases, specially this one.\n\nLeveraging my code with MongoDB, I was able to get loading speeds from 15ms~ for the first hit to 1.7ms~ on the second hit while using less then 8Gb of ram. This was enough to keep my GTX 1070 feed to 100%~.\n\nSo let's get started with mongoDB!\nFirst you will need to load the *.bson files into MongoDB, assuming that you already installed Mongo.\nRun this comands on your console:\n\n    mongorestore -d Cdiscount -c test test.bson\n    mongorestore -d Cdiscount -c train train.bson\n\nYou can check if the data was loaded correctly with:\n\n    mongo\n    show dbs\n\nShould have some like this:\n![]\n(https://imgur.com/ltESIzF.png)\n\n\nGreat! Now we have to access that with python and Pymongo!\n\n    import io\n    import pymongo\n    from pymongo import MongoClient\n    \n    client = MongoClient(connect=False) #Makes it \"good enough\" for our multi-threaded use case. \n    train = client.Cdiscount['train']\n    test = client.Cdiscount['test']\n    \n    index = 0 #Witch idx Do you want to load?\n    data = train.find_one({'_id': index})\n    img = imread(io.BytesIO(data['imgs'][0]['picture'])) Get the first picture from data\n    label = data['category_id']\nPytorch Data-loader, first hit:\n![]\n(https://imgur.com/9I3jH20.png)\n\nSecond hit (on adjacent idx):\n![]\n(https://imgur.com/f1WQIll.png)\n\nPytorch-MongoDB Starter kit coming soon =D",
      "votes": null
    },
    {
      "id": "231290",
      "postDate": "10/14/2017 09:38:09",
      "content": "<p>Hello Felipe,</p>\n\n<p>This is interesting, I also tried mongodb without success. Did you test randomly access ids instead of consecutive ones?</p>",
      "rawMarkdown": "Hello Felipe,\n\nThis is interesting, I also tried mongodb without success. Did you test randomly access ids instead of consecutive ones?",
      "votes": null
    },
    {
      "id": "231293",
      "postDate": "10/14/2017 09:45:24",
      "content": "<p>I am using this approach and for me random access is really no problem. As far as I know this is even one of mongodb's strengths!</p>",
      "rawMarkdown": "I am using this approach and for me random access is really no problem. As far as I know this is even one of mongodb's strengths!",
      "votes": null
    },
    {
      "id": "231296",
      "postDate": "10/14/2017 10:10:13",
      "content": "<p>Ok maybe I was not setting it up correctly, but I don't quite understand how it can do better since random read is mainly limited by HDD I/O performance?</p>",
      "rawMarkdown": "Ok maybe I was not setting it up correctly, but I don't quite understand how it can do better since random read is mainly limited by HDD I/O performance?",
      "votes": null
    },
    {
      "id": "231300",
      "postDate": "10/14/2017 10:26:48",
      "content": "<p>Ah sure you are right. Obviously without SSD you may still be limited. </p>",
      "rawMarkdown": "Ah sure you are right. Obviously without SSD you may still be limited.",
      "votes": null
    },
    {
      "id": "231322",
      "postDate": "10/14/2017 12:38:10",
      "content": "<p>I,ve see similar ideas in <a href=\"https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py\">https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py</a></p>",
      "rawMarkdown": "I,ve see similar ideas in https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py",
      "votes": null
    },
    {
      "id": "231367",
      "postDate": "10/14/2017 15:35:49",
      "content": "<p>Yes! Actually I'm even more impressed with this, not sure how MongoDB is able to do it. To be fair, I'm running my tests on the test dataset that is relatively small compered to the training one. Here are the results for 10k random reads! Just under 7ms (on HDD)!!!</p>\n\n<p>(Not sure if I din't messed up with timing, please correct me if I did something wrong)</p>\n\n<p>timing*\n<img src=\"https://imgur.com/m0CKixv.png\" alt=\"\" title=\"\"></p>",
      "rawMarkdown": "Yes! Actually I'm even more impressed with this, not sure how MongoDB is able to do it. To be fair, I'm running my tests on the test dataset that is relatively small compered to the training one. Here are the results for 10k random reads! Just under 7ms (on HDD)!!!\n\n(Not sure if I din't messed up with timing, please correct me if I did something wrong)\n\ntiming*\n![]\n(https://imgur.com/m0CKixv.png)",
      "votes": null
    },
    {
      "id": "231372",
      "postDate": "10/14/2017 15:50:11",
      "content": "<p>Sequencial reads are much faster, reading 100k sequentially yields this results:</p>\n\n<p><img src=\"https://imgur.com/wLBTLtY.png\" alt=\"\" title=\"\"></p>",
      "rawMarkdown": "Sequencial reads are much faster, reading 100k sequentially yields this results:\n\n![]\n(https://imgur.com/wLBTLtY.png)",
      "votes": null
    },
    {
      "id": "232208",
      "postDate": "10/17/2017 04:52:56",
      "content": "<p>Hi Felipe, I tried your implementation into my current generator function. However, I seem to notice certain index doesn't return any values ( which returns None type ).  Do you encounter this issues in your side as well? </p>",
      "rawMarkdown": "Hi Felipe, I tried your implementation into my current generator function. However, I seem to notice certain index doesn't return any values ( which returns None type ).  Do you encounter this issues in your side as well?",
      "votes": null
    },
    {
      "id": "232210",
      "postDate": "10/17/2017 04:55:14",
      "content": "<p>My current workaround is simply ignoring those None results. I hope to know if anyone has any experience working with mongodb can help explain whether this was normal or not?</p>",
      "rawMarkdown": "My current workaround is simply ignoring those None results. I hope to know if anyone has any experience working with mongodb can help explain whether this was normal or not?",
      "votes": null
    },
    {
      "id": "232282",
      "postDate": "10/17/2017 09:25:51",
      "content": "<p>I do not use this implementation, but I never had any None results in mine.</p>",
      "rawMarkdown": "I do not use this implementation, but I never had any None results in mine.",
      "votes": null
    },
    {
      "id": "232287",
      "postDate": "10/17/2017 09:31:47",
      "content": "<p>I think there are two possible causes.</p>\n\n<ol>\n<li><p>not converting index to int.</p></li>\n<li><p>not using product_id for index.</p></li>\n</ol>",
      "rawMarkdown": "I think there are two possible causes.\n\n1. not converting index to int.\n\n1. not using product_id for index.",
      "votes": null
    },
    {
      "id": "232346",
      "postDate": "10/17/2017 11:33:35",
      "content": "<p>I faced the same problem, I pre processed the data creating a list with every _id. Then I used this list for accessing the database. I will continue this post after my mid term exams that are happening this week =P</p>",
      "rawMarkdown": "I faced the same problem, I pre processed the data creating a list with every _id. Then I used this list for accessing the database. I will continue this post after my mid term exams that are happening this week =P",
      "votes": null
    },
    {
      "id": "239666",
      "postDate": "11/04/2017 04:05:48",
      "content": "<p>The None result is because we are assuming the id for each product is assigned in order. In fact, it turns out that the _id are randomly keeping number (maybe it gose to the test.bson). I show some id over here:</p>\n\n<ul>\n<li>index          _id</li>\n<li>7069891  23620452</li>\n<li>7069892  23620453</li>\n<li>7069895  23620455</li>\n<li>7069894  23620457</li>\n<li>7069893  23620460</li>\n</ul>\n\n<p>where 23620452 &gt;&gt; 7069891</p>",
      "rawMarkdown": "The None result is because we are assuming the id for each product is assigned in order. In fact, it turns out that the _id are randomly keeping number (maybe it gose to the test.bson). I show some id over here:\n\n - index          _id\n - 7069891  23620452\n - 7069892  23620453\n - 7069895  23620455\n - 7069894  23620457\n - 7069893  23620460\n\nwhere 23620452 &gt;&gt; 7069891",
      "votes": null
    },
    {
      "id": "242692",
      "postDate": "11/12/2017 11:55:26",
      "content": "<p>How did you manage to put the whole train+test data in 24GB?? Is this real life?</p>",
      "rawMarkdown": "How did you manage to put the whole train+test data in 24GB?? Is this real life?",
      "votes": null
    },
    {
      "id": "1081168",
      "postDate": "11/16/2020 21:54:26",
      "content": "<p>Hi i have a json document and i have opened it in kaggle in phyton and i want to use mongodb codes but It doesn’t identify the codes « db » what can i do?? I don’t have mongodb and phyton on my pc and also my pc doesn’t support SSh.</p>",
      "rawMarkdown": "Hi i have a json document and i have opened it in kaggle in phyton and i want to use mongodb codes but It doesn’t identify the codes « db » what can i do?? I don’t have mongodb and phyton on my pc and also my pc doesn’t support SSh.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1081168,
      "author_name": "negarta",
      "author_url": "",
      "post_date": "11/16/2020 21:54:26",
      "content": "<p>Hi i have a json document and i have opened it in kaggle in phyton and i want to use mongodb codes but It doesn’t identify the codes « db » what can i do?? I don’t have mongodb and phyton on my pc and also my pc doesn’t support SSh.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 231290,
      "author_name": "lamdang",
      "author_url": "",
      "post_date": "10/14/2017 09:38:09",
      "content": "<p>Hello Felipe,</p>\n\n<p>This is interesting, I also tried mongodb without success. Did you test randomly access ids instead of consecutive ones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 231293,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 09:45:24",
          "content": "<p>I am using this approach and for me random access is really no problem. As far as I know this is even one of mongodb's strengths!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231296,
          "author_name": "lamdang",
          "author_url": "",
          "post_date": "10/14/2017 10:10:13",
          "content": "<p>Ok maybe I was not setting it up correctly, but I don't quite understand how it can do better since random read is mainly limited by HDD I/O performance?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231300,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/14/2017 10:26:48",
          "content": "<p>Ah sure you are right. Obviously without SSD you may still be limited. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231322,
          "author_name": "jpizarrom",
          "author_url": "",
          "post_date": "10/14/2017 12:38:10",
          "content": "<p>I,ve see similar ideas in <a href=\"https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py\">https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231367,
          "author_name": "felipesens",
          "author_url": "",
          "post_date": "10/14/2017 15:35:49",
          "content": "<p>Yes! Actually I'm even more impressed with this, not sure how MongoDB is able to do it. To be fair, I'm running my tests on the test dataset that is relatively small compered to the training one. Here are the results for 10k random reads! Just under 7ms (on HDD)!!!</p>\n\n<p>(Not sure if I din't messed up with timing, please correct me if I did something wrong)</p>\n\n<p>timing*\n<img src=\"https://imgur.com/m0CKixv.png\" alt=\"\" title=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 231372,
          "author_name": "felipesens",
          "author_url": "",
          "post_date": "10/14/2017 15:50:11",
          "content": "<p>Sequencial reads are much faster, reading 100k sequentially yields this results:</p>\n\n<p><img src=\"https://imgur.com/wLBTLtY.png\" alt=\"\" title=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 232208,
      "author_name": "theblackcat",
      "author_url": "",
      "post_date": "10/17/2017 04:52:56",
      "content": "<p>Hi Felipe, I tried your implementation into my current generator function. However, I seem to notice certain index doesn't return any values ( which returns None type ).  Do you encounter this issues in your side as well? </p>",
      "votes": null,
      "replies": [
        {
          "id": 232210,
          "author_name": "theblackcat",
          "author_url": "",
          "post_date": "10/17/2017 04:55:14",
          "content": "<p>My current workaround is simply ignoring those None results. I hope to know if anyone has any experience working with mongodb can help explain whether this was normal or not?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 232282,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/17/2017 09:25:51",
          "content": "<p>I do not use this implementation, but I never had any None results in mine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 232287,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "10/17/2017 09:31:47",
          "content": "<p>I think there are two possible causes.</p>\n\n<ol>\n<li><p>not converting index to int.</p></li>\n<li><p>not using product_id for index.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 232346,
          "author_name": "felipesens",
          "author_url": "",
          "post_date": "10/17/2017 11:33:35",
          "content": "<p>I faced the same problem, I pre processed the data creating a list with every _id. Then I used this list for accessing the database. I will continue this post after my mid term exams that are happening this week =P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 239666,
          "author_name": "auroralht",
          "author_url": "",
          "post_date": "11/04/2017 04:05:48",
          "content": "<p>The None result is because we are assuming the id for each product is assigned in order. In fact, it turns out that the _id are randomly keeping number (maybe it gose to the test.bson). I show some id over here:</p>\n\n<ul>\n<li>index          _id</li>\n<li>7069891  23620452</li>\n<li>7069892  23620453</li>\n<li>7069895  23620455</li>\n<li>7069894  23620457</li>\n<li>7069893  23620460</li>\n</ul>\n\n<p>where 23620452 &gt;&gt; 7069891</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 242692,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "11/12/2017 11:55:26",
      "content": "<p>How did you manage to put the whole train+test data in 24GB?? Is this real life?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "231247": "Loading data was always a problem for my workstation, HDD + 16Gb of Ram was usually not enough for big databases, specially this one.\n\nLeveraging my code with MongoDB, I was able to get loading speeds from 15ms~ for the first hit to 1.7ms~ on the second hit while using less then 8Gb of ram. This was enough to keep my GTX 1070 feed to 100%~.\n\nSo let's get started with mongoDB!\nFirst you will need to load the *.bson files into MongoDB, assuming that you already installed Mongo.\nRun this comands on your console:\n\n    mongorestore -d Cdiscount -c test test.bson\n    mongorestore -d Cdiscount -c train train.bson\n\nYou can check if the data was loaded correctly with:\n\n    mongo\n    show dbs\n\nShould have some like this:\n![]\n(https://imgur.com/ltESIzF.png)\n\n\nGreat! Now we have to access that with python and Pymongo!\n\n    import io\n    import pymongo\n    from pymongo import MongoClient\n    \n    client = MongoClient(connect=False) #Makes it \"good enough\" for our multi-threaded use case. \n    train = client.Cdiscount['train']\n    test = client.Cdiscount['test']\n    \n    index = 0 #Witch idx Do you want to load?\n    data = train.find_one({'_id': index})\n    img = imread(io.BytesIO(data['imgs'][0]['picture'])) Get the first picture from data\n    label = data['category_id']\nPytorch Data-loader, first hit:\n![]\n(https://imgur.com/9I3jH20.png)\n\nSecond hit (on adjacent idx):\n![]\n(https://imgur.com/f1WQIll.png)\n\nPytorch-MongoDB Starter kit coming soon =D",
    "231290": "Hello Felipe,\n\nThis is interesting, I also tried mongodb without success. Did you test randomly access ids instead of consecutive ones?",
    "231293": "I am using this approach and for me random access is really no problem. As far as I know this is even one of mongodb's strengths!",
    "231296": "Ok maybe I was not setting it up correctly, but I don't quite understand how it can do better since random read is mainly limited by HDD I/O performance?",
    "231300": "Ah sure you are right. Obviously without SSD you may still be limited.",
    "231322": "I,ve see similar ideas in https://github.com/xkumiyu/cdiscount-kernel/blob/master/dataset.py",
    "231367": "Yes! Actually I'm even more impressed with this, not sure how MongoDB is able to do it. To be fair, I'm running my tests on the test dataset that is relatively small compered to the training one. Here are the results for 10k random reads! Just under 7ms (on HDD)!!!\n\n(Not sure if I din't messed up with timing, please correct me if I did something wrong)\n\ntiming*\n![]\n(https://imgur.com/m0CKixv.png)",
    "231372": "Sequencial reads are much faster, reading 100k sequentially yields this results:\n\n![]\n(https://imgur.com/wLBTLtY.png)",
    "232208": "Hi Felipe, I tried your implementation into my current generator function. However, I seem to notice certain index doesn't return any values ( which returns None type ).  Do you encounter this issues in your side as well?",
    "232210": "My current workaround is simply ignoring those None results. I hope to know if anyone has any experience working with mongodb can help explain whether this was normal or not?",
    "232282": "I do not use this implementation, but I never had any None results in mine.",
    "232287": "I think there are two possible causes.\n\n1. not converting index to int.\n\n1. not using product_id for index.",
    "232346": "I faced the same problem, I pre processed the data creating a list with every _id. Then I used this list for accessing the database. I will continue this post after my mid term exams that are happening this week =P",
    "239666": "The None result is because we are assuming the id for each product is assigned in order. In fact, it turns out that the _id are randomly keeping number (maybe it gose to the test.bson). I show some id over here:\n\n - index          _id\n - 7069891  23620452\n - 7069892  23620453\n - 7069895  23620455\n - 7069894  23620457\n - 7069893  23620460\n\nwhere 23620452 &gt;&gt; 7069891",
    "242692": "How did you manage to put the whole train+test data in 24GB?? Is this real life?",
    "1081168": "Hi i have a json document and i have opened it in kaggle in phyton and i want to use mongodb codes but It doesn’t identify the codes « db » what can i do?? I don’t have mongodb and phyton on my pc and also my pc doesn’t support SSh."
  },
  "source": "meta"
}