{
  "id": 241447,
  "title": "Speed Up saving/resizing dicoms with joblib 🔥",
  "url": "/competitions/siim-covid19-detection/discussion/241447",
  "author_name": "Moein",
  "post_date": "2021-05-24T15:10:20.066000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>As most of us are using the nice code snippet from <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> to resize and save images (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/239918\" target=\"_blank\">here </a>is the related discussion post), with the following few lines of code you can speed it up at least 2 times faster. This will help if your inference notebook first resizes and saves all test images and then runs the model on .png or .jpg images.</p>\n<p><code>from joblib import Parallel, delayed</code></p>\n<pre><code>split = 'test'\nsave_dir = f'/kaggle/tmp/{split}/'\n\nos.makedirs(save_dir, exist_ok=True)\n\nsave_dir = f'/kaggle/tmp/{split}/study/'\nos.makedirs(save_dir, exist_ok=True)\n\n\ndef save_png(dirname, file):\n    xray = read_xray(os.path.join(dirname, file))\n    im = resize(xray, size=512) \n    study = dirname.split('/')[-2] + '_study.png'\n    im.save(os.path.join(save_dir, study))\n\npath = f'../input/siim-covid19-detection/{split}'\n\noutputs = Parallel(n_jobs=2)(delayed(save_png)(d, f) \n          for d, _, fnames  in tqdm(os.walk(path)) for f in fnames )\n</code></pre>",
  "messages": [
    {
      "id": 1321248,
      "postDate": "2021-05-24T15:10:20.067Z",
      "content": "<p>As most of us are using the nice code snippet from <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> to resize and save images (<a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/239918\" target=\"_blank\">here </a>is the related discussion post), with the following few lines of code you can speed it up at least 2 times faster. This will help if your inference notebook first resizes and saves all test images and then runs the model on .png or .jpg images.</p>\n<p><code>from joblib import Parallel, delayed</code></p>\n<pre><code>split = 'test'\nsave_dir = f'/kaggle/tmp/{split}/'\n\nos.makedirs(save_dir, exist_ok=True)\n\nsave_dir = f'/kaggle/tmp/{split}/study/'\nos.makedirs(save_dir, exist_ok=True)\n\n\ndef save_png(dirname, file):\n    xray = read_xray(os.path.join(dirname, file))\n    im = resize(xray, size=512) \n    study = dirname.split('/')[-2] + '_study.png'\n    im.save(os.path.join(save_dir, study))\n\npath = f'../input/siim-covid19-detection/{split}'\n\noutputs = Parallel(n_jobs=2)(delayed(save_png)(d, f) \n          for d, _, fnames  in tqdm(os.walk(path)) for f in fnames )\n</code></pre>",
      "rawMarkdown": "As most of us are using the nice code snippet from @xhlulu to resize and save images ([here ](https://www.kaggle.com/c/siim-covid19-detection/discussion/239918)is the related discussion post), with the following few lines of code you can speed it up at least 2 times faster. This will help if your inference notebook first resizes and saves all test images and then runs the model on .png or .jpg images.\n\n\n```from joblib import Parallel, delayed```\n\n```\nsplit = 'test'\nsave_dir = f'/kaggle/tmp/{split}/'\n\nos.makedirs(save_dir, exist_ok=True)\n\nsave_dir = f'/kaggle/tmp/{split}/study/'\nos.makedirs(save_dir, exist_ok=True)\n\n\ndef save_png(dirname, file):\n    xray = read_xray(os.path.join(dirname, file))\n    im = resize(xray, size=512) \n    study = dirname.split('/')[-2] + '_study.png'\n    im.save(os.path.join(save_dir, study))\n\npath = f'../input/siim-covid19-detection/{split}'\n\noutputs = Parallel(n_jobs=2)(delayed(save_png)(d, f) \n          for d, _, fnames  in tqdm(os.walk(path)) for f in fnames )\n\n```\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1321248": "As most of us are using the nice code snippet from @xhlulu to resize and save images ([here ](https://www.kaggle.com/c/siim-covid19-detection/discussion/239918)is the related discussion post), with the following few lines of code you can speed it up at least 2 times faster. This will help if your inference notebook first resizes and saves all test images and then runs the model on .png or .jpg images.\n\n\n```from joblib import Parallel, delayed```\n\n```\nsplit = 'test'\nsave_dir = f'/kaggle/tmp/{split}/'\n\nos.makedirs(save_dir, exist_ok=True)\n\nsave_dir = f'/kaggle/tmp/{split}/study/'\nos.makedirs(save_dir, exist_ok=True)\n\n\ndef save_png(dirname, file):\n    xray = read_xray(os.path.join(dirname, file))\n    im = resize(xray, size=512) \n    study = dirname.split('/')[-2] + '_study.png'\n    im.save(os.path.join(save_dir, study))\n\npath = f'../input/siim-covid19-detection/{split}'\n\noutputs = Parallel(n_jobs=2)(delayed(save_png)(d, f) \n          for d, _, fnames  in tqdm(os.walk(path)) for f in fnames )\n\n```\n"
  }
}