{
  "id": 15025,
  "title": "How would you batch process images without waiting for hours and hours?",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/15025",
  "author_name": "",
  "post_date": "2015-07-03T06:22:10.170Z",
  "votes": null,
  "comment_count": 3,
  "views": 1393,
  "content": "<p>Hi all,</p>\n<p>How&nbsp;do you&nbsp;batch process the raw images (resize, white balance, etc...)? There are 80k images.</p>\n<p>I use imagemagick convert to&nbsp;do that but it takes many hours to complete even one run. Suppose I need to do another experimentation (tuned some parameter), it is a loooooong wait.</p>\n<p>How does everyone do this? I'd really love to learn if there's better ways.</p>\n\n<p>Current way of doing it:</p>\n<p>-&nbsp;Use the largest instance on EC2 to parallel process the images (python joblib). This is the&nbsp;technique I use and this takes a lot of time, which I am looking for improvements.</p>\n<p>- Graphlab to parallelize on multiple nodes (?). The starter code seems to use a&nbsp;GraphLab to do so. I haven't tried it, does this make it a lot faster?</p>",
  "messages": [
    {
      "id": "83308",
      "postDate": "07/03/2015 06:22:10",
      "content": "<p>Hi all,</p>\n<p>How&nbsp;do you&nbsp;batch process the raw images (resize, white balance, etc...)? There are 80k images.</p>\n<p>I use imagemagick convert to&nbsp;do that but it takes many hours to complete even one run. Suppose I need to do another experimentation (tuned some parameter), it is a loooooong wait.</p>\n<p>How does everyone do this? I'd really love to learn if there's better ways.</p>\n\n<p>Current way of doing it:</p>\n<p>-&nbsp;Use the largest instance on EC2 to parallel process the images (python joblib). This is the&nbsp;technique I use and this takes a lot of time, which I am looking for improvements.</p>\n<p>- Graphlab to parallelize on multiple nodes (?). The starter code seems to use a&nbsp;GraphLab to do so. I haven't tried it, does this make it a lot faster?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "83315",
      "postDate": "07/03/2015 08:10:53",
      "content": "<p>One thing that can help is to&nbsp;already downscale the images before and keep using those as input. If there are any other &quot;static&quot; transformations, you can also do those.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "83662",
      "postDate": "07/07/2015 13:55:51",
      "content": "<p>It is an &quot;embarrassingly parallel&quot; thing, so you don't even need any helper libraries.\nJust write your python to take in value ranges from the command line, and then write a bash script with something like:</p>\n\n<pre><code>#!/bin/bash\npython bananas.py 1..1000 &amp;\npython bananas.py 1001..2000 &amp;\n...etc...\nwait\necho &quot;OMG IT IS DONE&quot;\n</code></pre>\n\n<p>The wait means the bash script will wait until the children in there are done, and then whatever you want to echo can help you remind you were you are at. \nRun that in one terminal window, and in another you can use something like htop and see how many cores it is using and other resource.</p>\n\n<p>I am doing something similar in Perl with a 32-core EC2 instance - the slow part for me are the other parts that use the changes I am doing - but this part has been easy enough, and you don't have to worry about parallel complications with shared resources.\n(unless your N scripts are writing to the same file - then you will need to sort out locking/unlocking before/after writes, and reads if order matters for what you are doing)</p>",
      "rawMarkdown": "It is an \"embarrassingly parallel\" thing, so you don't even need any helper libraries.\r\nJust write your python to take in value ranges from the command line, and then write a bash script with something like:\r\n\r\n    #!/bin/bash\r\n    python bananas.py 1..1000 &\r\n    python bananas.py 1001..2000 &\r\n    ...etc...\r\n    wait\r\n    echo \"OMG IT IS DONE\"\r\n\r\nThe wait means the bash script will wait until the children in there are done, and then whatever you want to echo can help you remind you were you are at. \r\nRun that in one terminal window, and in another you can use something like htop and see how many cores it is using and other resource.\r\n\r\nI am doing something similar in Perl with a 32-core EC2 instance - the slow part for me are the other parts that use the changes I am doing - but this part has been easy enough, and you don't have to worry about parallel complications with shared resources.\r\n(unless your N scripts are writing to the same file - then you will need to sort out locking/unlocking before/after writes, and reads if order matters for what you are doing)",
      "votes": null
    },
    {
      "id": "83809",
      "postDate": "07/08/2015 16:33:09",
      "content": "<p>I know it is embarrassing parallel. I am already scaling it to multiple threads. It is just that it is still very slow. I am now resizing it before doing - it's fine.</p>",
      "rawMarkdown": "I know it is embarrassing parallel. I am already scaling it to multiple threads. It is just that it is still very slow. I am now resizing it before doing - it's fine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 83315,
      "author_name": "jeffreydf",
      "author_url": "",
      "post_date": "07/03/2015 08:10:53",
      "content": "<p>One thing that can help is to&nbsp;already downscale the images before and keep using those as input. If there are any other &quot;static&quot; transformations, you can also do those.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83662,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "07/07/2015 13:55:51",
      "content": "<p>It is an &quot;embarrassingly parallel&quot; thing, so you don't even need any helper libraries.\nJust write your python to take in value ranges from the command line, and then write a bash script with something like:</p>\n\n<pre><code>#!/bin/bash\npython bananas.py 1..1000 &amp;\npython bananas.py 1001..2000 &amp;\n...etc...\nwait\necho &quot;OMG IT IS DONE&quot;\n</code></pre>\n\n<p>The wait means the bash script will wait until the children in there are done, and then whatever you want to echo can help you remind you were you are at. \nRun that in one terminal window, and in another you can use something like htop and see how many cores it is using and other resource.</p>\n\n<p>I am doing something similar in Perl with a 32-core EC2 instance - the slow part for me are the other parts that use the changes I am doing - but this part has been easy enough, and you don't have to worry about parallel complications with shared resources.\n(unless your N scripts are writing to the same file - then you will need to sort out locking/unlocking before/after writes, and reads if order matters for what you are doing)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83809,
      "author_name": "ericchio",
      "author_url": "",
      "post_date": "07/08/2015 16:33:09",
      "content": "<p>I know it is embarrassing parallel. I am already scaling it to multiple threads. It is just that it is still very slow. I am now resizing it before doing - it's fine.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "83308": "",
    "83315": "",
    "83662": "It is an \"embarrassingly parallel\" thing, so you don't even need any helper libraries.\r\nJust write your python to take in value ranges from the command line, and then write a bash script with something like:\r\n\r\n    #!/bin/bash\r\n    python bananas.py 1..1000 &\r\n    python bananas.py 1001..2000 &\r\n    ...etc...\r\n    wait\r\n    echo \"OMG IT IS DONE\"\r\n\r\nThe wait means the bash script will wait until the children in there are done, and then whatever you want to echo can help you remind you were you are at. \r\nRun that in one terminal window, and in another you can use something like htop and see how many cores it is using and other resource.\r\n\r\nI am doing something similar in Perl with a 32-core EC2 instance - the slow part for me are the other parts that use the changes I am doing - but this part has been easy enough, and you don't have to worry about parallel complications with shared resources.\r\n(unless your N scripts are writing to the same file - then you will need to sort out locking/unlocking before/after writes, and reads if order matters for what you are doing)",
    "83809": "I know it is embarrassing parallel. I am already scaling it to multiple threads. It is just that it is still very slow. I am now resizing it before doing - it's fine."
  },
  "source": "meta"
}