{
  "id": 272204,
  "title": "Handling a huge size of image data.  Are you ready Newbies?  ",
  "url": "/competitions/wikipedia-image-caption/discussion/272204",
  "author_name": "",
  "post_date": "2021-09-14T14:47:16.664270300Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What I learned from my first Kaggle Notebook in this competition:<br>\nI was just following my guts since I don't have any background. </p>\n<p>First issue, I would have to deal with URL. There are plenty of work with images (jpg) in Kaggle, not with URL.  After reading Radmir's (it was the only Notebook, I highly recommend it) to have an insight about to perform with URLs.</p>\n<p>How to open an image from the URL in PIL?  After searching in Google:<br>\n\"Write Url with file name in URLLIB. request. urlretrieve() method.\"  Then Bla, Bla, Bla, Blaba. </p>\n<p>Simple? It seemed to be, though after several attempts and answers in StackOverflow  I've only got: \"display: unable to open X server\" </p>\n<p>If you want to open a file with many rows apply: nrows=10000 or 1000 or whatever. Thank you Udbhav.</p>\n<p>Then I found: \"Accessing Lots of Images in Python\" by Rebecca Stone: (2019 no date)<br>\n<a href=\"https://realpython.com/storing-images-in-python/\" target=\"_blank\">https://realpython.com/storing-images-in-python/</a></p>\n<p>\"The number of images required for a given task is getting larger and larger. Algorithms like convolutional neural networks,  CNNs, can handle enormous datasets of images and even learn from them.\" </p>\n<p>You can use Pillow for the image manipulation. I tried but I was not well succeeded.</p>\n<p>\"Use LMDB which is a B+ tree, which basically means that it is a tree-like graph structure stored in memory where each key-value element is a node, and nodes can have many children. Nodes on the same level are linked to one another for fast traversal.\"  I didn't applied that.</p>\n<p>What about HDF5?  NASA’s blurb on HDF5.   If it's good for Nasa, probably it won't help a Newbie.</p>\n<p>!pip install h5py<br>\nimport h5py</p>\n<p>You can store that huge data in a disk, in LMDB and of course In GOOGLE Cloud, or other cloud shh! </p>\n<p>How long did all of that storing take?  I didn't even think about that.</p>\n<p>You'll have to adjust the code for many images, then experiment for reading them with timeit:</p>\n<p>from timeit import timeit</p>\n<p>\"The write time is often less critical than the read time.\"  And, plot all the read and write timings. </p>\n<p>Parallel Access</p>\n<p>\"With such large datasets, you may want to speed up your operation through parallelization.\"  </p>\n<p>In fact, I just intended to plot some images, without coping others code, not to go so far. As a beginner we can't absorb or even understand that level of information. </p>\n<p><a href=\"https://realpython.com/storing-images-in-python/\" target=\"_blank\">https://realpython.com/storing-images-in-python/</a></p>\n<p>Concurrency</p>\n<p>\"In the majority of cases, you won’t be interested in reading parts of the same image at the same time, but you will want to read multiple images at once. With this definition of concurrency, storing to disk as .png files actually allows for complete concurrency.\" (Rebecca Stone) </p>\n<p>You can learn more about how Convnets (CNNs) can be used for ranking selfies. </p>\n<p>Better choice: simply MAKE LESS SELFIES since not everybody wants to see your face over and over again . The Concurrence is too high.</p>\n<p>Don't forget to read all the Codes. Another great opportunity to learn how to deal with URL and this large amount of images.  Have fun in this Playground!   </p>",
  "messages": [
    {
      "id": "1512752",
      "postDate": "09/14/2021 14:47:16",
      "content": "<p>What I learned from my first Kaggle Notebook in this competition:<br>\nI was just following my guts since I don't have any background. </p>\n<p>First issue, I would have to deal with URL. There are plenty of work with images (jpg) in Kaggle, not with URL.  After reading Radmir's (it was the only Notebook, I highly recommend it) to have an insight about to perform with URLs.</p>\n<p>How to open an image from the URL in PIL?  After searching in Google:<br>\n\"Write Url with file name in URLLIB. request. urlretrieve() method.\"  Then Bla, Bla, Bla, Blaba. </p>\n<p>Simple? It seemed to be, though after several attempts and answers in StackOverflow  I've only got: \"display: unable to open X server\" </p>\n<p>If you want to open a file with many rows apply: nrows=10000 or 1000 or whatever. Thank you Udbhav.</p>\n<p>Then I found: \"Accessing Lots of Images in Python\" by Rebecca Stone: (2019 no date)<br>\n<a href=\"https://realpython.com/storing-images-in-python/\" target=\"_blank\">https://realpython.com/storing-images-in-python/</a></p>\n<p>\"The number of images required for a given task is getting larger and larger. Algorithms like convolutional neural networks,  CNNs, can handle enormous datasets of images and even learn from them.\" </p>\n<p>You can use Pillow for the image manipulation. I tried but I was not well succeeded.</p>\n<p>\"Use LMDB which is a B+ tree, which basically means that it is a tree-like graph structure stored in memory where each key-value element is a node, and nodes can have many children. Nodes on the same level are linked to one another for fast traversal.\"  I didn't applied that.</p>\n<p>What about HDF5?  NASA’s blurb on HDF5.   If it's good for Nasa, probably it won't help a Newbie.</p>\n<p>!pip install h5py<br>\nimport h5py</p>\n<p>You can store that huge data in a disk, in LMDB and of course In GOOGLE Cloud, or other cloud shh! </p>\n<p>How long did all of that storing take?  I didn't even think about that.</p>\n<p>You'll have to adjust the code for many images, then experiment for reading them with timeit:</p>\n<p>from timeit import timeit</p>\n<p>\"The write time is often less critical than the read time.\"  And, plot all the read and write timings. </p>\n<p>Parallel Access</p>\n<p>\"With such large datasets, you may want to speed up your operation through parallelization.\"  </p>\n<p>In fact, I just intended to plot some images, without coping others code, not to go so far. As a beginner we can't absorb or even understand that level of information. </p>\n<p><a href=\"https://realpython.com/storing-images-in-python/\" target=\"_blank\">https://realpython.com/storing-images-in-python/</a></p>\n<p>Concurrency</p>\n<p>\"In the majority of cases, you won’t be interested in reading parts of the same image at the same time, but you will want to read multiple images at once. With this definition of concurrency, storing to disk as .png files actually allows for complete concurrency.\" (Rebecca Stone) </p>\n<p>You can learn more about how Convnets (CNNs) can be used for ranking selfies. </p>\n<p>Better choice: simply MAKE LESS SELFIES since not everybody wants to see your face over and over again . The Concurrence is too high.</p>\n<p>Don't forget to read all the Codes. Another great opportunity to learn how to deal with URL and this large amount of images.  Have fun in this Playground!   </p>",
      "rawMarkdown": "What I learned from my first Kaggle Notebook in this competition:\nI was just following my guts since I don't have any background. \n\nFirst issue, I would have to deal with URL. There are plenty of work with images (jpg) in Kaggle, not with URL.  After reading Radmir's (it was the only Notebook, I highly recommend it) to have an insight about to perform with URLs.\n\nHow to open an image from the URL in PIL?  After searching in Google:\n\"Write Url with file name in URLLIB. request. urlretrieve() method.\"  Then Bla, Bla, Bla, Blaba. \n\nSimple? It seemed to be, though after several attempts and answers in StackOverflow  I've only got: \"display: unable to open X server\" \n\nIf you want to open a file with many rows apply: nrows=10000 or 1000 or whatever. Thank you Udbhav.\n\nThen I found: \"Accessing Lots of Images in Python\" by Rebecca Stone: (2019 no date)\nhttps://realpython.com/storing-images-in-python/\n\n \"The number of images required for a given task is getting larger and larger. Algorithms like convolutional neural networks,  CNNs, can handle enormous datasets of images and even learn from them.\" \n\nYou can use Pillow for the image manipulation. I tried but I was not well succeeded.\n\n\"Use LMDB which is a B+ tree, which basically means that it is a tree-like graph structure stored in memory where each key-value element is a node, and nodes can have many children. Nodes on the same level are linked to one another for fast traversal.\"  I didn't applied that.\n\nWhat about HDF5?  NASA’s blurb on HDF5.   If it's good for Nasa, probably it won't help a Newbie.\n\n!pip install h5py\nimport h5py\n\nYou can store that huge data in a disk, in LMDB and of course In GOOGLE Cloud, or other cloud shh! \n\nHow long did all of that storing take?  I didn't even think about that.\n\nYou'll have to adjust the code for many images, then experiment for reading them with timeit:\n\nfrom timeit import timeit\n\n\"The write time is often less critical than the read time.\"  And, plot all the read and write timings. \n\nParallel Access\n\n\"With such large datasets, you may want to speed up your operation through parallelization.\"  \n\n In fact, I just intended to plot some images, without coping others code, not to go so far. As a beginner we can't absorb or even understand that level of information. \n\nhttps://realpython.com/storing-images-in-python/\n\nConcurrency\n\n\"In the majority of cases, you won’t be interested in reading parts of the same image at the same time, but you will want to read multiple images at once. With this definition of concurrency, storing to disk as .png files actually allows for complete concurrency.\" (Rebecca Stone) \n\nYou can learn more about how Convnets (CNNs) can be used for ranking selfies. \n\nBetter choice: simply MAKE LESS SELFIES since not everybody wants to see your face over and over again . The Concurrence is too high.\n\nDon't forget to read all the Codes. Another great opportunity to learn how to deal with URL and this large amount of images.  Have fun in this Playground!",
      "votes": null
    },
    {
      "id": "1512895",
      "postDate": "09/14/2021 17:10:26",
      "content": "<p><a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>  well believe it or not it is easy to work with HDF5 format.</p>",
      "rawMarkdown": "mpwolke  well believe it or not it is easy to work with HDF5 format.",
      "votes": null
    },
    {
      "id": "1513190",
      "postDate": "09/14/2021 23:59:39",
      "content": "<p>For me it was not easy to work with any of them. Besides, I wasn't able to plot a single URL image, which it's a little bit frustating. </p>\n<p>Thank you again Muhammad. </p>",
      "rawMarkdown": "For me it was not easy to work with any of them. Besides, I wasn't able to plot a single URL image, which it's a little bit frustating. \n\nThank you again Muhammad.",
      "votes": null
    },
    {
      "id": "1532536",
      "postDate": "10/03/2021 05:22:33",
      "content": "<p>Looking forward to this</p>",
      "rawMarkdown": "Looking forward to this",
      "votes": null
    },
    {
      "id": "1532817",
      "postDate": "10/03/2021 12:17:00",
      "content": "<p>Thank you for commenting CR C1 10P.</p>",
      "rawMarkdown": "Thank you for commenting CR C1 10P.",
      "votes": null
    },
    {
      "id": "1534059",
      "postDate": "10/04/2021 15:09:09",
      "content": "<p>Any idea how to open the bytes64 encoded images? I have tried a few methods but I am not able to open them.<br>\nIf there is a script to do it, this would be better than downloading each image from the urls.</p>",
      "rawMarkdown": "Any idea how to open the bytes64 encoded images? I have tried a few methods but I am not able to open them.\nIf there is a script to do it, this would be better than downloading each image from the urls.",
      "votes": null
    },
    {
      "id": "1534348",
      "postDate": "10/04/2021 20:00:30",
      "content": "<p>Hi Psyopus,</p>\n<p>I found that code below in StackOverflow. Though if it's not what your question, you can open a Discussion topic and maybe some skilled Kaggler could answer your doubt.</p>\n<p>import base64<br>\nimgdata = base64.b64decode(imgstring)<br>\nfilename = 'some_image.jpg'  # I assume you have a way of picking unique filenames<br>\nwith open(filename, 'wb') as f:<br>\n    f.write(imgdata)</p>\n<p>f gets closed when you exit the with statement</p>\n<p>Now save the value of filename to your database</p>\n<p><a href=\"https://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280\" target=\"_blank\">https://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280</a></p>",
      "rawMarkdown": "Hi Psyopus,\n\nI found that code below in StackOverflow. Though if it's not what your question, you can open a Discussion topic and maybe some skilled Kaggler could answer your doubt.\n\nimport base64\nimgdata = base64.b64decode(imgstring)\nfilename = 'some_image.jpg'  # I assume you have a way of picking unique filenames\nwith open(filename, 'wb') as f:\n    f.write(imgdata)\n\n f gets closed when you exit the with statement\n\n Now save the value of filename to your database\n\nhttps://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1512895,
      "author_name": "saadsikander",
      "author_url": "",
      "post_date": "09/14/2021 17:10:26",
      "content": "<p><a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>  well believe it or not it is easy to work with HDF5 format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1513190,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "09/14/2021 23:59:39",
          "content": "<p>For me it was not easy to work with any of them. Besides, I wasn't able to plot a single URL image, which it's a little bit frustating. </p>\n<p>Thank you again Muhammad. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532536,
      "author_name": "chiragtagadiya",
      "author_url": "",
      "post_date": "10/03/2021 05:22:33",
      "content": "<p>Looking forward to this</p>",
      "votes": null,
      "replies": [
        {
          "id": 1532817,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "10/03/2021 12:17:00",
          "content": "<p>Thank you for commenting CR C1 10P.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1534059,
      "author_name": "psyopus777",
      "author_url": "",
      "post_date": "10/04/2021 15:09:09",
      "content": "<p>Any idea how to open the bytes64 encoded images? I have tried a few methods but I am not able to open them.<br>\nIf there is a script to do it, this would be better than downloading each image from the urls.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1534348,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "10/04/2021 20:00:30",
          "content": "<p>Hi Psyopus,</p>\n<p>I found that code below in StackOverflow. Though if it's not what your question, you can open a Discussion topic and maybe some skilled Kaggler could answer your doubt.</p>\n<p>import base64<br>\nimgdata = base64.b64decode(imgstring)<br>\nfilename = 'some_image.jpg'  # I assume you have a way of picking unique filenames<br>\nwith open(filename, 'wb') as f:<br>\n    f.write(imgdata)</p>\n<p>f gets closed when you exit the with statement</p>\n<p>Now save the value of filename to your database</p>\n<p><a href=\"https://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280\" target=\"_blank\">https://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1512752": "What I learned from my first Kaggle Notebook in this competition:\nI was just following my guts since I don't have any background. \n\nFirst issue, I would have to deal with URL. There are plenty of work with images (jpg) in Kaggle, not with URL.  After reading Radmir's (it was the only Notebook, I highly recommend it) to have an insight about to perform with URLs.\n\nHow to open an image from the URL in PIL?  After searching in Google:\n\"Write Url with file name in URLLIB. request. urlretrieve() method.\"  Then Bla, Bla, Bla, Blaba. \n\nSimple? It seemed to be, though after several attempts and answers in StackOverflow  I've only got: \"display: unable to open X server\" \n\nIf you want to open a file with many rows apply: nrows=10000 or 1000 or whatever. Thank you Udbhav.\n\nThen I found: \"Accessing Lots of Images in Python\" by Rebecca Stone: (2019 no date)\nhttps://realpython.com/storing-images-in-python/\n\n \"The number of images required for a given task is getting larger and larger. Algorithms like convolutional neural networks,  CNNs, can handle enormous datasets of images and even learn from them.\" \n\nYou can use Pillow for the image manipulation. I tried but I was not well succeeded.\n\n\"Use LMDB which is a B+ tree, which basically means that it is a tree-like graph structure stored in memory where each key-value element is a node, and nodes can have many children. Nodes on the same level are linked to one another for fast traversal.\"  I didn't applied that.\n\nWhat about HDF5?  NASA’s blurb on HDF5.   If it's good for Nasa, probably it won't help a Newbie.\n\n!pip install h5py\nimport h5py\n\nYou can store that huge data in a disk, in LMDB and of course In GOOGLE Cloud, or other cloud shh! \n\nHow long did all of that storing take?  I didn't even think about that.\n\nYou'll have to adjust the code for many images, then experiment for reading them with timeit:\n\nfrom timeit import timeit\n\n\"The write time is often less critical than the read time.\"  And, plot all the read and write timings. \n\nParallel Access\n\n\"With such large datasets, you may want to speed up your operation through parallelization.\"  \n\n In fact, I just intended to plot some images, without coping others code, not to go so far. As a beginner we can't absorb or even understand that level of information. \n\nhttps://realpython.com/storing-images-in-python/\n\nConcurrency\n\n\"In the majority of cases, you won’t be interested in reading parts of the same image at the same time, but you will want to read multiple images at once. With this definition of concurrency, storing to disk as .png files actually allows for complete concurrency.\" (Rebecca Stone) \n\nYou can learn more about how Convnets (CNNs) can be used for ranking selfies. \n\nBetter choice: simply MAKE LESS SELFIES since not everybody wants to see your face over and over again . The Concurrence is too high.\n\nDon't forget to read all the Codes. Another great opportunity to learn how to deal with URL and this large amount of images.  Have fun in this Playground!",
    "1512895": "mpwolke  well believe it or not it is easy to work with HDF5 format.",
    "1513190": "For me it was not easy to work with any of them. Besides, I wasn't able to plot a single URL image, which it's a little bit frustating. \n\nThank you again Muhammad.",
    "1532536": "Looking forward to this",
    "1532817": "Thank you for commenting CR C1 10P.",
    "1534059": "Any idea how to open the bytes64 encoded images? I have tried a few methods but I am not able to open them.\nIf there is a script to do it, this would be better than downloading each image from the urls.",
    "1534348": "Hi Psyopus,\n\nI found that code below in StackOverflow. Though if it's not what your question, you can open a Discussion topic and maybe some skilled Kaggler could answer your doubt.\n\nimport base64\nimgdata = base64.b64decode(imgstring)\nfilename = 'some_image.jpg'  # I assume you have a way of picking unique filenames\nwith open(filename, 'wb') as f:\n    f.write(imgdata)\n\n f gets closed when you exit the with statement\n\n Now save the value of filename to your database\n\nhttps://stackoverflow.com/questions/16214190/how-to-convert-base64-string-to-image/16214280"
  },
  "source": "meta"
}