{
  "id": 103778,
  "title": "Fastest way to crop all chars and save them to a new file",
  "url": "/competitions/kuzushiji-recognition/discussion/103778",
  "author_name": "",
  "post_date": "2019-08-11T21:48:50.911926800Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey friends,\nI would like to crop all images and save them to a new file each for an EDA-Kernel</p>\n\n<p>The first version of the code was really slow and the estimated execution time was about 8 hours. I worked on my code and it now executes in less than 15 minutes. I'm happy with the improvement but I would like to get the code even faster!</p>\n\n<p>Do you see anything I could improve?</p>\n\n<p>Thanks, Chris</p>\n\n<p>```\nimport os\nimport shutil\nfrom pathlib import Path</p>\n\n<p>import pandas as pd\nimport numpy as np</p>\n\n<p>from PIL import Image\nimport cv2</p>\n\n<p>from tqdm import tnrange, tqdm_notebook</p>\n\n<h1>all important folders</h1>\n\n<p>input_dir = Path(\"../input\")\ntest_dir = input_dir/'test_images'\ntrain_dir = input_dir/'train_images'</p>\n\n<h1>create folder for all chacters</h1>\n\n<p>char_dir = Path(\"../test_chars\")\nos.makedirs(char_dir)</p>\n\n<p>unicode_map = {codepoint: char for codepoint, char in pd.read_csv(input_dir/'unicode_translation.csv').values}\nunicode_list = list(unicode_map)</p>\n\n<h1>unicode to int conversion</h1>\n\n<h1>unique identifier for every unicode character</h1>\n\n<p>def unicodeToInt(unicode):\n    return unicode_list.index(unicode)</p>\n\n<p>train_df = pd.read_csv(\"../input/train.csv\")</p>\n\n<p>pbar = tqdm_notebook(total=len(os.listdir(train_dir)))\nfor loop1_index, (list_index, row) in enumerate(train_df.iterrows()):\n    labels = row.labels</p>\n\n<pre><code>if isinstance(labels, float):\n    pbar.update(1)\n    continue\n\nimg = cv2.imread(str(f\"{str(train_dir)}/{row.image_id}.jpg\"))\nimg = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\nlabels = labels.split(\" \")\n</code></pre>\n\n<h1>change unicode char to integer representing the class</h1>\n\n<pre><code>labels[::5] = map(unicodeToInt, labels[::5])\n\nlabels = np.array(labels, dtype=np.int16)\n\nlabels = labels.reshape(-1, 5)\n\nlabels[:, 3] = np.sum(a=labels[:,[1,3]], axis=1)\nlabels[:, 4] = np.sum(a=labels[:,[2,4]], axis=1)\n\n[Image.fromarray(img[label[2]:label[4], label[1]:label[3]]).save(f\"{char_dir}/{label[0]}_{loop1_index}-{loop2_index}.jpg\") for loop2_index, label in enumerate(labels)]\n\npbar.update(1)\n</code></pre>\n\n<p>pbar.close()\n```</p>",
  "messages": [
    {
      "id": "597124",
      "postDate": "08/11/2019 21:48:50",
      "content": "<p>Hey friends,\nI would like to crop all images and save them to a new file each for an EDA-Kernel</p>\n\n<p>The first version of the code was really slow and the estimated execution time was about 8 hours. I worked on my code and it now executes in less than 15 minutes. I'm happy with the improvement but I would like to get the code even faster!</p>\n\n<p>Do you see anything I could improve?</p>\n\n<p>Thanks, Chris</p>\n\n<p>```\nimport os\nimport shutil\nfrom pathlib import Path</p>\n\n<p>import pandas as pd\nimport numpy as np</p>\n\n<p>from PIL import Image\nimport cv2</p>\n\n<p>from tqdm import tnrange, tqdm_notebook</p>\n\n<h1>all important folders</h1>\n\n<p>input_dir = Path(\"../input\")\ntest_dir = input_dir/'test_images'\ntrain_dir = input_dir/'train_images'</p>\n\n<h1>create folder for all chacters</h1>\n\n<p>char_dir = Path(\"../test_chars\")\nos.makedirs(char_dir)</p>\n\n<p>unicode_map = {codepoint: char for codepoint, char in pd.read_csv(input_dir/'unicode_translation.csv').values}\nunicode_list = list(unicode_map)</p>\n\n<h1>unicode to int conversion</h1>\n\n<h1>unique identifier for every unicode character</h1>\n\n<p>def unicodeToInt(unicode):\n    return unicode_list.index(unicode)</p>\n\n<p>train_df = pd.read_csv(\"../input/train.csv\")</p>\n\n<p>pbar = tqdm_notebook(total=len(os.listdir(train_dir)))\nfor loop1_index, (list_index, row) in enumerate(train_df.iterrows()):\n    labels = row.labels</p>\n\n<pre><code>if isinstance(labels, float):\n    pbar.update(1)\n    continue\n\nimg = cv2.imread(str(f\"{str(train_dir)}/{row.image_id}.jpg\"))\nimg = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n\nlabels = labels.split(\" \")\n</code></pre>\n\n<h1>change unicode char to integer representing the class</h1>\n\n<pre><code>labels[::5] = map(unicodeToInt, labels[::5])\n\nlabels = np.array(labels, dtype=np.int16)\n\nlabels = labels.reshape(-1, 5)\n\nlabels[:, 3] = np.sum(a=labels[:,[1,3]], axis=1)\nlabels[:, 4] = np.sum(a=labels[:,[2,4]], axis=1)\n\n[Image.fromarray(img[label[2]:label[4], label[1]:label[3]]).save(f\"{char_dir}/{label[0]}_{loop1_index}-{loop2_index}.jpg\") for loop2_index, label in enumerate(labels)]\n\npbar.update(1)\n</code></pre>\n\n<p>pbar.close()\n```</p>",
      "rawMarkdown": "Hey friends,\nI would like to crop all images and save them to a new file each for an EDA-Kernel\n\nThe first version of the code was really slow and the estimated execution time was about 8 hours. I worked on my code and it now executes in less than 15 minutes. I'm happy with the improvement but I would like to get the code even faster!\n\nDo you see anything I could improve?\n\nThanks, Chris\n\n```\nimport os\nimport shutil\nfrom pathlib import Path\n\nimport pandas as pd\nimport numpy as np\n\nfrom PIL import Image\nimport cv2\n\nfrom tqdm import tnrange, tqdm_notebook\n\n\n# all important folders\ninput_dir = Path(\"../input\")\ntest_dir = input_dir/'test_images'\ntrain_dir = input_dir/'train_images'\n\n# create folder for all chacters\nchar_dir = Path(\"../test_chars\")\nos.makedirs(char_dir)\n\nunicode_map = {codepoint: char for codepoint, char in pd.read_csv(input_dir/'unicode_translation.csv').values}\nunicode_list = list(unicode_map)\n\n# unicode to int conversion\n# unique identifier for every unicode character\ndef unicodeToInt(unicode):\n    return unicode_list.index(unicode)\n\ntrain_df = pd.read_csv(\"../input/train.csv\")\n\npbar = tqdm_notebook(total=len(os.listdir(train_dir)))\nfor loop1_index, (list_index, row) in enumerate(train_df.iterrows()):\n    labels = row.labels\n    \n    if isinstance(labels, float):\n        pbar.update(1)\n        continue\n    \n    img = cv2.imread(str(f\"{str(train_dir)}/{row.image_id}.jpg\"))\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    \n    labels = labels.split(\" \")\n    \n#     change unicode char to integer representing the class\n    labels[::5] = map(unicodeToInt, labels[::5])\n    \n    labels = np.array(labels, dtype=np.int16)\n\n    labels = labels.reshape(-1, 5)\n    \n    labels[:, 3] = np.sum(a=labels[:,[1,3]], axis=1)\n    labels[:, 4] = np.sum(a=labels[:,[2,4]], axis=1)\n    \n    [Image.fromarray(img[label[2]:label[4], label[1]:label[3]]).save(f\"{char_dir}/{label[0]}_{loop1_index}-{loop2_index}.jpg\") for loop2_index, label in enumerate(labels)]\n    \n    pbar.update(1)\npbar.close()\n```",
      "votes": null
    },
    {
      "id": "597744",
      "postDate": "08/12/2019 18:03:18",
      "content": "<p>If you refactor the code inside <code>enumerate(train_df.iterrows())</code> to be a function, you can use <a href=\"https://docs.python.org/3/library/multiprocessing.html\">https://docs.python.org/3/library/multiprocessing.html</a> to write all the images in parallel. Theoretical speedup should be the number of threads on your CPU. I use this pattern regularly when doing data transformations on a lot of files. You would have to the train dataframe into a list.</p>",
      "rawMarkdown": "If you refactor the code inside `enumerate(train_df.iterrows())` to be a function, you can use https://docs.python.org/3/library/multiprocessing.html to write all the images in parallel. Theoretical speedup should be the number of threads on your CPU. I use this pattern regularly when doing data transformations on a lot of files. You would have to the train dataframe into a list.",
      "votes": null
    },
    {
      "id": "598226",
      "postDate": "08/13/2019 09:51:46",
      "content": "<p>thanks a lot for the tip</p>",
      "rawMarkdown": "thanks a lot for the tip",
      "votes": null
    },
    {
      "id": "600941",
      "postDate": "08/16/2019 19:27:41",
      "content": "<p>Hey friends, to speed up the cropping, I tried the Python multiprocessing module as suggested by \nJeffrey Scholz in the comment below. But it seems like there is only a small improvement of around 2 secs in execution time per iteration. Did I do anything wrong, or is this usual and expected?</p>\n\n<p>Here is my kernel:\n<a href=\"https://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images\">https://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images</a></p>",
      "rawMarkdown": "Hey friends, to speed up the cropping, I tried the Python multiprocessing module as suggested by \nJeffrey Scholz in the comment below. But it seems like there is only a small improvement of around 2 secs in execution time per iteration. Did I do anything wrong, or is this usual and expected?\n\nHere is my kernel:\nhttps://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images",
      "votes": null
    },
    {
      "id": "616728",
      "postDate": "09/03/2019 11:58:39",
      "content": "<p>You should probably profile the (single process) code first to see what exactly is causing the slowdown. One thing I would suggest is only spawning as many processes as your system can handle (at or around the number of cores) as context-switching can slow things down.  From what I see, there shouldn't be any major issues with data passing between the parent and child processes since you are using manually initialized the processes instead of using a process pool. Long story short, you should always start by profiling. </p>",
      "rawMarkdown": "You should probably profile the (single process) code first to see what exactly is causing the slowdown. One thing I would suggest is only spawning as many processes as your system can handle (at or around the number of cores) as context-switching can slow things down.  From what I see, there shouldn't be any major issues with data passing between the parent and child processes since you are using manually initialized the processes instead of using a process pool. Long story short, you should always start by profiling.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 597744,
      "author_name": "jeffreyscholz",
      "author_url": "",
      "post_date": "08/12/2019 18:03:18",
      "content": "<p>If you refactor the code inside <code>enumerate(train_df.iterrows())</code> to be a function, you can use <a href=\"https://docs.python.org/3/library/multiprocessing.html\">https://docs.python.org/3/library/multiprocessing.html</a> to write all the images in parallel. Theoretical speedup should be the number of threads on your CPU. I use this pattern regularly when doing data transformations on a lot of files. You would have to the train dataframe into a list.</p>",
      "votes": null,
      "replies": [
        {
          "id": 598226,
          "author_name": "christianwallenwein",
          "author_url": "",
          "post_date": "08/13/2019 09:51:46",
          "content": "<p>thanks a lot for the tip</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 600941,
      "author_name": "christianwallenwein",
      "author_url": "",
      "post_date": "08/16/2019 19:27:41",
      "content": "<p>Hey friends, to speed up the cropping, I tried the Python multiprocessing module as suggested by \nJeffrey Scholz in the comment below. But it seems like there is only a small improvement of around 2 secs in execution time per iteration. Did I do anything wrong, or is this usual and expected?</p>\n\n<p>Here is my kernel:\n<a href=\"https://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images\">https://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 616728,
          "author_name": "deasmhumnha",
          "author_url": "",
          "post_date": "09/03/2019 11:58:39",
          "content": "<p>You should probably profile the (single process) code first to see what exactly is causing the slowdown. One thing I would suggest is only spawning as many processes as your system can handle (at or around the number of cores) as context-switching can slow things down.  From what I see, there shouldn't be any major issues with data passing between the parent and child processes since you are using manually initialized the processes instead of using a process pool. Long story short, you should always start by profiling. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "597124": "Hey friends,\nI would like to crop all images and save them to a new file each for an EDA-Kernel\n\nThe first version of the code was really slow and the estimated execution time was about 8 hours. I worked on my code and it now executes in less than 15 minutes. I'm happy with the improvement but I would like to get the code even faster!\n\nDo you see anything I could improve?\n\nThanks, Chris\n\n```\nimport os\nimport shutil\nfrom pathlib import Path\n\nimport pandas as pd\nimport numpy as np\n\nfrom PIL import Image\nimport cv2\n\nfrom tqdm import tnrange, tqdm_notebook\n\n\n# all important folders\ninput_dir = Path(\"../input\")\ntest_dir = input_dir/'test_images'\ntrain_dir = input_dir/'train_images'\n\n# create folder for all chacters\nchar_dir = Path(\"../test_chars\")\nos.makedirs(char_dir)\n\nunicode_map = {codepoint: char for codepoint, char in pd.read_csv(input_dir/'unicode_translation.csv').values}\nunicode_list = list(unicode_map)\n\n# unicode to int conversion\n# unique identifier for every unicode character\ndef unicodeToInt(unicode):\n    return unicode_list.index(unicode)\n\ntrain_df = pd.read_csv(\"../input/train.csv\")\n\npbar = tqdm_notebook(total=len(os.listdir(train_dir)))\nfor loop1_index, (list_index, row) in enumerate(train_df.iterrows()):\n    labels = row.labels\n    \n    if isinstance(labels, float):\n        pbar.update(1)\n        continue\n    \n    img = cv2.imread(str(f\"{str(train_dir)}/{row.image_id}.jpg\"))\n    img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)\n    \n    labels = labels.split(\" \")\n    \n#     change unicode char to integer representing the class\n    labels[::5] = map(unicodeToInt, labels[::5])\n    \n    labels = np.array(labels, dtype=np.int16)\n\n    labels = labels.reshape(-1, 5)\n    \n    labels[:, 3] = np.sum(a=labels[:,[1,3]], axis=1)\n    labels[:, 4] = np.sum(a=labels[:,[2,4]], axis=1)\n    \n    [Image.fromarray(img[label[2]:label[4], label[1]:label[3]]).save(f\"{char_dir}/{label[0]}_{loop1_index}-{loop2_index}.jpg\") for loop2_index, label in enumerate(labels)]\n    \n    pbar.update(1)\npbar.close()\n```",
    "597744": "If you refactor the code inside `enumerate(train_df.iterrows())` to be a function, you can use https://docs.python.org/3/library/multiprocessing.html to write all the images in parallel. Theoretical speedup should be the number of threads on your CPU. I use this pattern regularly when doing data transformations on a lot of files. You would have to the train dataframe into a list.",
    "598226": "thanks a lot for the tip",
    "600941": "Hey friends, to speed up the cropping, I tried the Python multiprocessing module as suggested by \nJeffrey Scholz in the comment below. But it seems like there is only a small improvement of around 2 secs in execution time per iteration. Did I do anything wrong, or is this usual and expected?\n\nHere is my kernel:\nhttps://www.kaggle.com/christianwallenwein/fastest-way-to-crop-all-images",
    "616728": "You should probably profile the (single process) code first to see what exactly is causing the slowdown. One thing I would suggest is only spawning as many processes as your system can handle (at or around the number of cores) as context-switching can slow things down.  From what I see, there shouldn't be any major issues with data passing between the parent and child processes since you are using manually initialized the processes instead of using a process pool. Long story short, you should always start by profiling."
  },
  "source": "meta"
}