{
  "id": 223423,
  "title": "Here's why ihelon's image_id_2_path function is genius (and might save you 20 min)",
  "url": "/competitions/bms-molecular-translation/discussion/223423",
  "author_name": "",
  "post_date": "2021-03-03T18:34:50.899741400Z",
  "votes": 60,
  "comment_count": 5,
  "views": 0,
  "content": "<p>You've probably read <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s great <a href=\"https://www.kaggle.com/ihelon/molecular-translation-exploratory-data-analysis\" target=\"_blank\">EDA notebook</a> by now! However, I wanted to highlight one incredible part of his notebook: his <code>convert_image_id_2_path</code> function.</p>\n<h1>The <code>os.walk</code> approach</h1>\n<p>Before we think about pandas, let's consider <code>os.walk</code>. After all, it's a built-in function and it's shown in every new notebook, so it's a good choice, right? </p>\n<p>That's what i initially thought, and here's what I wrote to compute all the paths:</p>\n<pre><code>import os\nimport tqdm\n\npaths = {\"train\": [], \"test\": []}\n\nfor split in ['train', 'test']:\n    load_dir = f'../input/bms-molecular-translation/{split}/'\n    for dirname, _, filenames in tqdm(os.walk(load_dir)):\n        for filename in filenames:\n            paths[split].append(os.path.join(dirname, filename))\n</code></pre>\n<p>But not so fast! <strong>Actually, the loop above is terrible and will take you 20 minutes to run</strong> .The reason for that is that <code>os.walk</code> will compute the directory tree by listing the content of a directory with <code>os.scandir</code> and recursively traverse each sub-directory (well, that's roughly how it works, you can check the <a href=\"https://docs.python.org/3/library/os.html#os.walk\" target=\"_blank\">references</a> for the exact details). </p>\n<p>Notice how the directories have a fixed structure: the first level directory has the name of the first character of your image, the second level directory has the second character, and so forth. For example, if your image ID is <code>abcd123</code> then the path should be <code>train/a/b/c/abcd123.png</code>. Since you already know ahead of time what the directory structure is, you don't need <code>os.walk</code> anymore! This brings us to the simple string formatting used by <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>.</p>\n<h1><a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s super fast conversion function</h1>\n<p>The function below (taken from his notebook) converts a image ID to the path of the image:</p>\n<pre><code>def convert_image_id_2_path(image_id: str) -&gt; str:\n    return \"../input/bms-molecular-translation/train/{}/{}/{}/{}.png\".format(\n        image_id[0], image_id[1], image_id[2], image_id \n    )\n</code></pre>\n<p>You can apply it to your dataframe to get all the paths for either the train or test dataframe:</p>\n<pre><code>import pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntrain_paths = train.image_id.apply(convert_image_id_2_path)\n</code></pre>\n<p><strong>Notice that it takes around 2 seconds per dataframe</strong>, which is a tiny fraction of the original 20 minutes it took with <code>os.walk</code>! The reason why it's so fast is because pandas can efficiently apply the function to each element in <code>df.image_id</code>, which should be (in theory) faster than for loops.</p>\n<h1>Bonus</h1>\n<p>I'm updating <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s function to also work with the test set:</p>\n<pre><code>def convert_to_path(split: str):\n    # https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\n    def aux(image_id: str) -&gt; str:\n        return \"../input/bms-molecular-translation/{}/{}/{}/{}/{}.png\".format(\n            split, image_id[0], image_id[1], image_id[2], image_id \n        )\n\n    return aux\n</code></pre>\n<p>Which you can use like this:</p>\n<pre><code>import pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntest = pd.read_csv('../input/bms-molecular-translation/sample_submission.csv')\n\ntrain_paths = train.image_id.apply(convert_to_path('train'))\ntest_paths = test.image_id.apply(convert_to_path('test'))\n</code></pre>\n<p>Hope this helps!</p>",
  "messages": [
    {
      "id": "1225607",
      "postDate": "03/03/2021 18:34:50",
      "content": "<p>You've probably read <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s great <a href=\"https://www.kaggle.com/ihelon/molecular-translation-exploratory-data-analysis\" target=\"_blank\">EDA notebook</a> by now! However, I wanted to highlight one incredible part of his notebook: his <code>convert_image_id_2_path</code> function.</p>\n<h1>The <code>os.walk</code> approach</h1>\n<p>Before we think about pandas, let's consider <code>os.walk</code>. After all, it's a built-in function and it's shown in every new notebook, so it's a good choice, right? </p>\n<p>That's what i initially thought, and here's what I wrote to compute all the paths:</p>\n<pre><code>import os\nimport tqdm\n\npaths = {\"train\": [], \"test\": []}\n\nfor split in ['train', 'test']:\n    load_dir = f'../input/bms-molecular-translation/{split}/'\n    for dirname, _, filenames in tqdm(os.walk(load_dir)):\n        for filename in filenames:\n            paths[split].append(os.path.join(dirname, filename))\n</code></pre>\n<p>But not so fast! <strong>Actually, the loop above is terrible and will take you 20 minutes to run</strong> .The reason for that is that <code>os.walk</code> will compute the directory tree by listing the content of a directory with <code>os.scandir</code> and recursively traverse each sub-directory (well, that's roughly how it works, you can check the <a href=\"https://docs.python.org/3/library/os.html#os.walk\" target=\"_blank\">references</a> for the exact details). </p>\n<p>Notice how the directories have a fixed structure: the first level directory has the name of the first character of your image, the second level directory has the second character, and so forth. For example, if your image ID is <code>abcd123</code> then the path should be <code>train/a/b/c/abcd123.png</code>. Since you already know ahead of time what the directory structure is, you don't need <code>os.walk</code> anymore! This brings us to the simple string formatting used by <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>.</p>\n<h1><a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s super fast conversion function</h1>\n<p>The function below (taken from his notebook) converts a image ID to the path of the image:</p>\n<pre><code>def convert_image_id_2_path(image_id: str) -&gt; str:\n    return \"../input/bms-molecular-translation/train/{}/{}/{}/{}.png\".format(\n        image_id[0], image_id[1], image_id[2], image_id \n    )\n</code></pre>\n<p>You can apply it to your dataframe to get all the paths for either the train or test dataframe:</p>\n<pre><code>import pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntrain_paths = train.image_id.apply(convert_image_id_2_path)\n</code></pre>\n<p><strong>Notice that it takes around 2 seconds per dataframe</strong>, which is a tiny fraction of the original 20 minutes it took with <code>os.walk</code>! The reason why it's so fast is because pandas can efficiently apply the function to each element in <code>df.image_id</code>, which should be (in theory) faster than for loops.</p>\n<h1>Bonus</h1>\n<p>I'm updating <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a>'s function to also work with the test set:</p>\n<pre><code>def convert_to_path(split: str):\n    # https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\n    def aux(image_id: str) -&gt; str:\n        return \"../input/bms-molecular-translation/{}/{}/{}/{}/{}.png\".format(\n            split, image_id[0], image_id[1], image_id[2], image_id \n        )\n\n    return aux\n</code></pre>\n<p>Which you can use like this:</p>\n<pre><code>import pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntest = pd.read_csv('../input/bms-molecular-translation/sample_submission.csv')\n\ntrain_paths = train.image_id.apply(convert_to_path('train'))\ntest_paths = test.image_id.apply(convert_to_path('test'))\n</code></pre>\n<p>Hope this helps!</p>",
      "rawMarkdown": "You've probably read @ihelon's great [EDA notebook](https://www.kaggle.com/ihelon/molecular-translation-exploratory-data-analysis) by now! However, I wanted to highlight one incredible part of his notebook: his `convert_image_id_2_path` function.\n\n# The `os.walk` approach\n\nBefore we think about pandas, let's consider `os.walk`. After all, it's a built-in function and it's shown in every new notebook, so it's a good choice, right? \n\nThat's what i initially thought, and here's what I wrote to compute all the paths:\n```\nimport os\nimport tqdm\n\npaths = {\"train\": [], \"test\": []}\n\nfor split in ['train', 'test']:\n    load_dir = f'../input/bms-molecular-translation/{split}/'\n    for dirname, _, filenames in tqdm(os.walk(load_dir)):\n        for filename in filenames:\n            paths[split].append(os.path.join(dirname, filename))\n```\n\nBut not so fast! **Actually, the loop above is terrible and will take you 20 minutes to run** .The reason for that is that `os.walk` will compute the directory tree by listing the content of a directory with `os.scandir` and recursively traverse each sub-directory (well, that's roughly how it works, you can check the [references](https://docs.python.org/3/library/os.html#os.walk) for the exact details). \n\nNotice how the directories have a fixed structure: the first level directory has the name of the first character of your image, the second level directory has the second character, and so forth. For example, if your image ID is `abcd123` then the path should be `train/a/b/c/abcd123.png`. Since you already know ahead of time what the directory structure is, you don't need `os.walk` anymore! This brings us to the simple string formatting used by @ihelon.\n\n# @ihelon's super fast conversion function\n\nThe function below (taken from his notebook) converts a image ID to the path of the image:\n```\ndef convert_image_id_2_path(image_id: str) -> str:\n    return \"../input/bms-molecular-translation/train/{}/{}/{}/{}.png\".format(\n        image_id[0], image_id[1], image_id[2], image_id \n    )\n```\n\nYou can apply it to your dataframe to get all the paths for either the train or test dataframe:\n```\nimport pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntrain_paths = train.image_id.apply(convert_image_id_2_path)\n```\n\n**Notice that it takes around 2 seconds per dataframe**, which is a tiny fraction of the original 20 minutes it took with `os.walk`! The reason why it's so fast is because pandas can efficiently apply the function to each element in `df.image_id`, which should be (in theory) faster than for loops.\n\n# Bonus\n\nI'm updating @ihelon's function to also work with the test set:\n```\ndef convert_to_path(split: str):\n    # https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\n    def aux(image_id: str) -> str:\n        return \"../input/bms-molecular-translation/{}/{}/{}/{}/{}.png\".format(\n            split, image_id[0], image_id[1], image_id[2], image_id \n        )\n    \n    return aux\n```\n\nWhich you can use like this:\n```\nimport pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntest = pd.read_csv('../input/bms-molecular-translation/sample_submission.csv')\n\ntrain_paths = train.image_id.apply(convert_to_path('train'))\ntest_paths = test.image_id.apply(convert_to_path('test'))\n```\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "1226086",
      "postDate": "03/04/2021 08:11:59",
      "content": "<p>It's because of these bits and tricks why we beginners should join and follow the competitions even though we may not know that much and it serves as a good starting point!!<br>\nGreat find out btw</p>",
      "rawMarkdown": "It's because of these bits and tricks why we beginners should join and follow the competitions even though we may not know that much and it serves as a good starting point!!\nGreat find out btw",
      "votes": null
    },
    {
      "id": "1226625",
      "postDate": "03/04/2021 17:43:39",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1228699",
      "postDate": "03/06/2021 16:50:38",
      "content": "<p>Thanks! Was pleasantly surprised <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> was able to figure that out so quickly… Guess that's what it means to work smarter not harder!</p>",
      "rawMarkdown": "Thanks! Was pleasantly surprised @ihelon was able to figure that out so quickly... Guess that's what it means to work smarter not harder!",
      "votes": null
    },
    {
      "id": "1258173",
      "postDate": "03/31/2021 12:05:51",
      "content": "<p>Genius thoughts!</p>",
      "rawMarkdown": "Genius thoughts!",
      "votes": null
    },
    {
      "id": "1287626",
      "postDate": "04/29/2021 08:37:35",
      "content": "<p>Now, just use pandarallel's parallel_apply and it will be much faster than this.</p>",
      "rawMarkdown": "Now, just use pandarallel's parallel_apply and it will be much faster than this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1226086,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/04/2021 08:11:59",
      "content": "<p>It's because of these bits and tricks why we beginners should join and follow the competitions even though we may not know that much and it serves as a good starting point!!<br>\nGreat find out btw</p>",
      "votes": null,
      "replies": [
        {
          "id": 1228699,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "03/06/2021 16:50:38",
          "content": "<p>Thanks! Was pleasantly surprised <a href=\"https://www.kaggle.com/ihelon\" target=\"_blank\">@ihelon</a> was able to figure that out so quickly… Guess that's what it means to work smarter not harder!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1226625,
      "author_name": "rajkumarsinghakonwar",
      "author_url": "",
      "post_date": "03/04/2021 17:43:39",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1258173,
      "author_name": "appleinca",
      "author_url": "",
      "post_date": "03/31/2021 12:05:51",
      "content": "<p>Genius thoughts!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1287626,
      "author_name": "abhishekvermasg1",
      "author_url": "",
      "post_date": "04/29/2021 08:37:35",
      "content": "<p>Now, just use pandarallel's parallel_apply and it will be much faster than this.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1225607": "You've probably read @ihelon's great [EDA notebook](https://www.kaggle.com/ihelon/molecular-translation-exploratory-data-analysis) by now! However, I wanted to highlight one incredible part of his notebook: his `convert_image_id_2_path` function.\n\n# The `os.walk` approach\n\nBefore we think about pandas, let's consider `os.walk`. After all, it's a built-in function and it's shown in every new notebook, so it's a good choice, right? \n\nThat's what i initially thought, and here's what I wrote to compute all the paths:\n```\nimport os\nimport tqdm\n\npaths = {\"train\": [], \"test\": []}\n\nfor split in ['train', 'test']:\n    load_dir = f'../input/bms-molecular-translation/{split}/'\n    for dirname, _, filenames in tqdm(os.walk(load_dir)):\n        for filename in filenames:\n            paths[split].append(os.path.join(dirname, filename))\n```\n\nBut not so fast! **Actually, the loop above is terrible and will take you 20 minutes to run** .The reason for that is that `os.walk` will compute the directory tree by listing the content of a directory with `os.scandir` and recursively traverse each sub-directory (well, that's roughly how it works, you can check the [references](https://docs.python.org/3/library/os.html#os.walk) for the exact details). \n\nNotice how the directories have a fixed structure: the first level directory has the name of the first character of your image, the second level directory has the second character, and so forth. For example, if your image ID is `abcd123` then the path should be `train/a/b/c/abcd123.png`. Since you already know ahead of time what the directory structure is, you don't need `os.walk` anymore! This brings us to the simple string formatting used by @ihelon.\n\n# @ihelon's super fast conversion function\n\nThe function below (taken from his notebook) converts a image ID to the path of the image:\n```\ndef convert_image_id_2_path(image_id: str) -> str:\n    return \"../input/bms-molecular-translation/train/{}/{}/{}/{}.png\".format(\n        image_id[0], image_id[1], image_id[2], image_id \n    )\n```\n\nYou can apply it to your dataframe to get all the paths for either the train or test dataframe:\n```\nimport pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntrain_paths = train.image_id.apply(convert_image_id_2_path)\n```\n\n**Notice that it takes around 2 seconds per dataframe**, which is a tiny fraction of the original 20 minutes it took with `os.walk`! The reason why it's so fast is because pandas can efficiently apply the function to each element in `df.image_id`, which should be (in theory) faster than for loops.\n\n# Bonus\n\nI'm updating @ihelon's function to also work with the test set:\n```\ndef convert_to_path(split: str):\n    # https://www.kaggle.com/c/bms-molecular-translation/discussion/223423\n    def aux(image_id: str) -> str:\n        return \"../input/bms-molecular-translation/{}/{}/{}/{}/{}.png\".format(\n            split, image_id[0], image_id[1], image_id[2], image_id \n        )\n    \n    return aux\n```\n\nWhich you can use like this:\n```\nimport pandas as pd\n\ntrain = pd.read_csv('../input/bms-molecular-translation/train_labels.csv')\ntest = pd.read_csv('../input/bms-molecular-translation/sample_submission.csv')\n\ntrain_paths = train.image_id.apply(convert_to_path('train'))\ntest_paths = test.image_id.apply(convert_to_path('test'))\n```\n\nHope this helps!",
    "1226086": "It's because of these bits and tricks why we beginners should join and follow the competitions even though we may not know that much and it serves as a good starting point!!\nGreat find out btw",
    "1226625": "Thank you for sharing!",
    "1228699": "Thanks! Was pleasantly surprised @ihelon was able to figure that out so quickly... Guess that's what it means to work smarter not harder!",
    "1258173": "Genius thoughts!",
    "1287626": "Now, just use pandarallel's parallel_apply and it will be much faster than this."
  },
  "source": "meta"
}