{
  "id": 105600,
  "title": "A script to join x6 channel images into x1 RGB image (multicore)",
  "url": "/competitions/recursion-cellular-image-classification/discussion/105600",
  "author_name": "",
  "post_date": "2019-08-24T15:28:03.155037500Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Created <a href=\"https://www.kaggle.com/purplejester/joining-channels-into-rgb-images-parallel\">a simple script</a> that traverses the directory with the data and converts 6 standalone images into a single RGB image. (Using all available cores of your machine). Probably could be helpful for someone to start with the competition. Not sure how efficient it is to join 6 channels of data into 3 RGB channels, though. But a good start point, I believe.</p>",
  "messages": [
    {
      "id": "607080",
      "postDate": "08/24/2019 15:28:03",
      "content": "<p>Created <a href=\"https://www.kaggle.com/purplejester/joining-channels-into-rgb-images-parallel\">a simple script</a> that traverses the directory with the data and converts 6 standalone images into a single RGB image. (Using all available cores of your machine). Probably could be helpful for someone to start with the competition. Not sure how efficient it is to join 6 channels of data into 3 RGB channels, though. But a good start point, I believe.</p>",
      "rawMarkdown": "Created [a simple script](https://www.kaggle.com/purplejester/joining-channels-into-rgb-images-parallel) that traverses the directory with the data and converts 6 standalone images into a single RGB image. (Using all available cores of your machine). Probably could be helpful for someone to start with the competition. Not sure how efficient it is to join 6 channels of data into 3 RGB channels, though. But a good start point, I believe.",
      "votes": null
    },
    {
      "id": "607447",
      "postDate": "08/25/2019 09:11:02",
      "content": "<p>What is the advantage over using the script <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py\">provided by the host</a>?</p>",
      "rawMarkdown": "What is the advantage over using the script [provided by the host](https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py)?",
      "votes": null
    },
    {
      "id": "607677",
      "postDate": "08/25/2019 17:48:26",
      "content": "<p>There is also a very good notebook by <a href=\"/xhlulu\">@xhlulu</a> using the script provided by the host + additional code to perform resizing and saving images to an easily retrievable new path <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images</a> .</p>",
      "rawMarkdown": "There is also a very good notebook by @xhlulu using the script provided by the host + additional code to perform resizing and saving images to an easily retrievable new path https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images .",
      "votes": null
    },
    {
      "id": "608363",
      "postDate": "08/26/2019 17:03:23",
      "content": "<p>Probably I've misunderstood, but the host script seems to pull the data from a remote bucket, not from the local files. And <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">this kernel</a> doesn't use multi-processing, as I can see. On my machine, the parallel script gets approximately x12 speed-up. So probably someone could adapt this code for their need. Anyway, decided to share just in case :)</p>",
      "rawMarkdown": "Probably I've misunderstood, but the host script seems to pull the data from a remote bucket, not from the local files. And [this kernel](https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images) doesn't use multi-processing, as I can see. On my machine, the parallel script gets approximately x12 speed-up. So probably someone could adapt this code for their need. Anyway, decided to share just in case :)",
      "votes": null
    },
    {
      "id": "609006",
      "postDate": "08/27/2019 10:26:55",
      "content": "<p>Updated the script to use only local data without querying remote buckets provided by the host. Before it used <code>rio.combine_metadata</code>. Could be helpful if you run this script on your local machine, it should be capable to finish the processing within 1-2 hours, depending on the number of cores you have. (In my case, it was around ~30 minutes to merge all samples from the train subset).</p>\n\n<p>Probably it is not too useful for interactive kernels, but could be valuable for dataset preparation.</p>",
      "rawMarkdown": "Updated the script to use only local data without querying remote buckets provided by the host. Before it used `rio.combine_metadata`. Could be helpful if you run this script on your local machine, it should be capable to finish the processing within 1-2 hours, depending on the number of cores you have. (In my case, it was around ~30 minutes to merge all samples from the train subset).\n\nProbably it is not too useful for interactive kernels, but could be valuable for dataset preparation.",
      "votes": null
    },
    {
      "id": "609522",
      "postDate": "08/27/2019 19:56:15",
      "content": "<p>If you check <code>rio.combine_metadata</code> in the github repo, you'll find that you can change the base_path for it to be a local one.</p>\n\n<p><code>\ndef combine_metadata(base_path=DEFAULT_METADATA_BASE_PATH,\n                     include_controls=True):\n</code>\n<code>DEFAULT_METADATA_BASE_PATH</code> is <code>gs://rxrx1-us-central1/metadata</code> indeed, but you can change it.</p>\n\n<p>source: <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233</a></p>",
      "rawMarkdown": "If you check `rio.combine_metadata` in the github repo, you'll find that you can change the base_path for it to be a local one.\n\n```\ndef combine_metadata(base_path=DEFAULT_METADATA_BASE_PATH,\n                     include_controls=True):\n```\n`DEFAULT_METADATA_BASE_PATH` is `gs://rxrx1-us-central1/metadata` indeed, but you can change it.\n\nsource: https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233",
      "votes": null
    },
    {
      "id": "609794",
      "postDate": "08/28/2019 05:35:07",
      "content": "<p>Ah makes sense! Didn't check it. </p>",
      "rawMarkdown": "Ah makes sense! Didn't check it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 607447,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/25/2019 09:11:02",
      "content": "<p>What is the advantage over using the script <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py\">provided by the host</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 607677,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/25/2019 17:48:26",
          "content": "<p>There is also a very good notebook by <a href=\"/xhlulu\">@xhlulu</a> using the script provided by the host + additional code to perform resizing and saving images to an easily retrievable new path <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images</a> .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 608363,
      "author_name": "purplejester",
      "author_url": "",
      "post_date": "08/26/2019 17:03:23",
      "content": "<p>Probably I've misunderstood, but the host script seems to pull the data from a remote bucket, not from the local files. And <a href=\"https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images\">this kernel</a> doesn't use multi-processing, as I can see. On my machine, the parallel script gets approximately x12 speed-up. So probably someone could adapt this code for their need. Anyway, decided to share just in case :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 609006,
      "author_name": "purplejester",
      "author_url": "",
      "post_date": "08/27/2019 10:26:55",
      "content": "<p>Updated the script to use only local data without querying remote buckets provided by the host. Before it used <code>rio.combine_metadata</code>. Could be helpful if you run this script on your local machine, it should be capable to finish the processing within 1-2 hours, depending on the number of cores you have. (In my case, it was around ~30 minutes to merge all samples from the train subset).</p>\n\n<p>Probably it is not too useful for interactive kernels, but could be valuable for dataset preparation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 609522,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/27/2019 19:56:15",
          "content": "<p>If you check <code>rio.combine_metadata</code> in the github repo, you'll find that you can change the base_path for it to be a local one.</p>\n\n<p><code>\ndef combine_metadata(base_path=DEFAULT_METADATA_BASE_PATH,\n                     include_controls=True):\n</code>\n<code>DEFAULT_METADATA_BASE_PATH</code> is <code>gs://rxrx1-us-central1/metadata</code> indeed, but you can change it.</p>\n\n<p>source: <a href=\"https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233\">https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609794,
          "author_name": "purplejester",
          "author_url": "",
          "post_date": "08/28/2019 05:35:07",
          "content": "<p>Ah makes sense! Didn't check it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "607080": "Created [a simple script](https://www.kaggle.com/purplejester/joining-channels-into-rgb-images-parallel) that traverses the directory with the data and converts 6 standalone images into a single RGB image. (Using all available cores of your machine). Probably could be helpful for someone to start with the competition. Not sure how efficient it is to join 6 channels of data into 3 RGB channels, though. But a good start point, I believe.",
    "607447": "What is the advantage over using the script [provided by the host](https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py)?",
    "607677": "There is also a very good notebook by @xhlulu using the script provided by the host + additional code to perform resizing and saving images to an easily retrievable new path https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images .",
    "608363": "Probably I've misunderstood, but the host script seems to pull the data from a remote bucket, not from the local files. And [this kernel](https://www.kaggle.com/xhlulu/recursion-2019-load-resize-and-save-images) doesn't use multi-processing, as I can see. On my machine, the parallel script gets approximately x12 speed-up. So probably someone could adapt this code for their need. Anyway, decided to share just in case :)",
    "609006": "Updated the script to use only local data without querying remote buckets provided by the host. Before it used `rio.combine_metadata`. Could be helpful if you run this script on your local machine, it should be capable to finish the processing within 1-2 hours, depending on the number of cores you have. (In my case, it was around ~30 minutes to merge all samples from the train subset).\n\nProbably it is not too useful for interactive kernels, but could be valuable for dataset preparation.",
    "609522": "If you check `rio.combine_metadata` in the github repo, you'll find that you can change the base_path for it to be a local one.\n\n```\ndef combine_metadata(base_path=DEFAULT_METADATA_BASE_PATH,\n                     include_controls=True):\n```\n`DEFAULT_METADATA_BASE_PATH` is `gs://rxrx1-us-central1/metadata` indeed, but you can change it.\n\nsource: https://github.com/recursionpharma/rxrx1-utils/blob/master/rxrx/io.py#L233",
    "609794": "Ah makes sense! Didn't check it."
  },
  "source": "meta"
}