{
  "id": 19077,
  "title": "Image processing times with skimage",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19077",
  "author_name": "",
  "post_date": "2016-02-19T15:35:25.727Z",
  "votes": null,
  "comment_count": 4,
  "views": 867,
  "content": "<p>Hello everyone. </p>\n\n<p>Are people are using skimage for their data augmentation? If so, what are some reasonable processing times. My augmentation takes about 230 seconds per iteration which is really killing me! </p>\n\n<p>I'm using the transform._warps_cy._warp_fast() function which I know was faster about a year ago but I was wondering if that changed or if there is a faster way now?</p>\n\n<p>Thanks to all and good luck in the final stretch of the competition!</p>",
  "messages": [
    {
      "id": "108744",
      "postDate": "02/19/2016 15:35:25",
      "content": "<p>Hello everyone. </p>\n\n<p>Are people are using skimage for their data augmentation? If so, what are some reasonable processing times. My augmentation takes about 230 seconds per iteration which is really killing me! </p>\n\n<p>I'm using the transform._warps_cy._warp_fast() function which I know was faster about a year ago but I was wondering if that changed or if there is a faster way now?</p>\n\n<p>Thanks to all and good luck in the final stretch of the competition!</p>",
      "rawMarkdown": "Hello everyone. \r\n\r\nAre people are using skimage for their data augmentation? If so, what are some reasonable processing times. My augmentation takes about 230 seconds per iteration which is really killing me! \r\n\r\nI'm using the transform._warps_cy._warp_fast() function which I know was faster about a year ago but I was wondering if that changed or if there is a faster way now?\r\n\r\nThanks to all and good luck in the final stretch of the competition!",
      "votes": null
    },
    {
      "id": "108752",
      "postDate": "02/19/2016 16:09:39",
      "content": "<p>i am using skimage but as pre-processing step, and don't have the additional timelag that you refer...</p>",
      "rawMarkdown": "i am using skimage but as pre-processing step, and don't have the additional timelag that you refer...",
      "votes": null
    },
    {
      "id": "108759",
      "postDate": "02/19/2016 16:56:57",
      "content": "<p>I'm using other framework. </p>\n\n<p>But, are you using a GPU to process the data with the last drivers installed ? <br>\nAnd, if you are using a GPU, how are you transfering the data to the GPU ?</p>",
      "rawMarkdown": "I'm using other framework. \r\n\r\nBut, are you using a GPU to process the data with the last drivers installed ?  \r\nAnd, if you are using a GPU, how are you transfering the data to the GPU ?",
      "votes": null
    },
    {
      "id": "108760",
      "postDate": "02/19/2016 17:05:28",
      "content": "<p>I'm using a GTX 980. I have the data in RAM -&gt; random augmentations before each batch -&gt; augmented images to GPU. When I've done this on RGB images it only adds about 20 seconds to each full iteration. But, I guess since I'm treating these as 30 channel images instead of 3 it'll be 10x the processing time. </p>",
      "rawMarkdown": "I'm using a GTX 980. I have the data in RAM -> random augmentations before each batch -> augmented images to GPU. When I've done this on RGB images it only adds about 20 seconds to each full iteration. But, I guess since I'm treating these as 30 channel images instead of 3 it'll be 10x the processing time.",
      "votes": null
    },
    {
      "id": "108767",
      "postDate": "02/19/2016 18:10:56",
      "content": "<p>Are you sure that the <strong>all the data you need</strong> are in the GPU memory ? </p>\n\n<p>It's not the same to put 128 MB into the GPU than 1 GB.\nRemember, if the GPU need to go into the RAM to take the date it will take time.</p>\n\n<p>And if you have put some variable in the RAM that should be in the GPU it should be a problem.</p>\n\n<p>Another thing, i don't know what calculus are you doing with the data, but remember to split the calculus to be done into the CPU and another to be done with the GPU. And parallelism need to be done with some wisdom, because if you split so much the calculus you will make loose time. It's called Amdahl's law.\nMay be you have some problem  there.</p>\n\n<p>And about the another variables ?</p>\n\n<p>How much time take to load and save the data from the disk  ?</p>\n\n<p>All your framework are running in many platform (like Python for the main code and frameworks in C) ?\nCommunicate into different platforms take time too.</p>\n\n<p>Are you sure that you framework have enabled the GPU support and it's suporting all the resources you need ?</p>\n\n<p>I don't  know if it's helping.</p>",
      "rawMarkdown": "Are you sure that the **all the data you need** are in the GPU memory ? \r\n\r\nIt's not the same to put 128 MB into the GPU than 1 GB.\r\nRemember, if the GPU need to go into the RAM to take the date it will take time.\r\n\r\nAnd if you have put some variable in the RAM that should be in the GPU it should be a problem.\r\n\r\nAnother thing, i don't know what calculus are you doing with the data, but remember to split the calculus to be done into the CPU and another to be done with the GPU. And parallelism need to be done with some wisdom, because if you split so much the calculus you will make loose time. It's called Amdahl's law.\r\nMay be you have some problem  there.\r\n\r\nAnd about the another variables ?\r\n\r\nHow much time take to load and save the data from the disk  ?\r\n\r\nAll your framework are running in many platform (like Python for the main code and frameworks in C) ?\r\nCommunicate into different platforms take time too.\r\n\r\nAre you sure that you framework have enabled the GPU support and it's suporting all the resources you need ?\r\n\r\nI don't  know if it's helping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 108752,
      "author_name": "wouterd1",
      "author_url": "",
      "post_date": "02/19/2016 16:09:39",
      "content": "<p>i am using skimage but as pre-processing step, and don't have the additional timelag that you refer...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108759,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/19/2016 16:56:57",
      "content": "<p>I'm using other framework. </p>\n\n<p>But, are you using a GPU to process the data with the last drivers installed ? <br>\nAnd, if you are using a GPU, how are you transfering the data to the GPU ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108760,
      "author_name": "florianm",
      "author_url": "",
      "post_date": "02/19/2016 17:05:28",
      "content": "<p>I'm using a GTX 980. I have the data in RAM -&gt; random augmentations before each batch -&gt; augmented images to GPU. When I've done this on RGB images it only adds about 20 seconds to each full iteration. But, I guess since I'm treating these as 30 channel images instead of 3 it'll be 10x the processing time. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108767,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/19/2016 18:10:56",
      "content": "<p>Are you sure that the <strong>all the data you need</strong> are in the GPU memory ? </p>\n\n<p>It's not the same to put 128 MB into the GPU than 1 GB.\nRemember, if the GPU need to go into the RAM to take the date it will take time.</p>\n\n<p>And if you have put some variable in the RAM that should be in the GPU it should be a problem.</p>\n\n<p>Another thing, i don't know what calculus are you doing with the data, but remember to split the calculus to be done into the CPU and another to be done with the GPU. And parallelism need to be done with some wisdom, because if you split so much the calculus you will make loose time. It's called Amdahl's law.\nMay be you have some problem  there.</p>\n\n<p>And about the another variables ?</p>\n\n<p>How much time take to load and save the data from the disk  ?</p>\n\n<p>All your framework are running in many platform (like Python for the main code and frameworks in C) ?\nCommunicate into different platforms take time too.</p>\n\n<p>Are you sure that you framework have enabled the GPU support and it's suporting all the resources you need ?</p>\n\n<p>I don't  know if it's helping.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "108744": "Hello everyone. \r\n\r\nAre people are using skimage for their data augmentation? If so, what are some reasonable processing times. My augmentation takes about 230 seconds per iteration which is really killing me! \r\n\r\nI'm using the transform._warps_cy._warp_fast() function which I know was faster about a year ago but I was wondering if that changed or if there is a faster way now?\r\n\r\nThanks to all and good luck in the final stretch of the competition!",
    "108752": "i am using skimage but as pre-processing step, and don't have the additional timelag that you refer...",
    "108759": "I'm using other framework. \r\n\r\nBut, are you using a GPU to process the data with the last drivers installed ?  \r\nAnd, if you are using a GPU, how are you transfering the data to the GPU ?",
    "108760": "I'm using a GTX 980. I have the data in RAM -> random augmentations before each batch -> augmented images to GPU. When I've done this on RGB images it only adds about 20 seconds to each full iteration. But, I guess since I'm treating these as 30 channel images instead of 3 it'll be 10x the processing time.",
    "108767": "Are you sure that the **all the data you need** are in the GPU memory ? \r\n\r\nIt's not the same to put 128 MB into the GPU than 1 GB.\r\nRemember, if the GPU need to go into the RAM to take the date it will take time.\r\n\r\nAnd if you have put some variable in the RAM that should be in the GPU it should be a problem.\r\n\r\nAnother thing, i don't know what calculus are you doing with the data, but remember to split the calculus to be done into the CPU and another to be done with the GPU. And parallelism need to be done with some wisdom, because if you split so much the calculus you will make loose time. It's called Amdahl's law.\r\nMay be you have some problem  there.\r\n\r\nAnd about the another variables ?\r\n\r\nHow much time take to load and save the data from the disk  ?\r\n\r\nAll your framework are running in many platform (like Python for the main code and frameworks in C) ?\r\nCommunicate into different platforms take time too.\r\n\r\nAre you sure that you framework have enabled the GPU support and it's suporting all the resources you need ?\r\n\r\nI don't  know if it's helping."
  },
  "source": "meta"
}