{
  "id": 227492,
  "title": "🔥🔥🔥 Useful utility: normalize metadata JSONs  to extract content as a dataframe",
  "url": "/competitions/iwildcam2021-fgvc8/discussion/227492",
  "author_name": "",
  "post_date": "2021-03-20T19:06:02.390285Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<h3>💥Introduction</h3>\n<p>The annotations for train data, including the bounding boxes for detected animals, the metadata for train and test are provided in this competition in JSON format.</p>\n<p>I provide here a fast and simple method to bulk process these JSON files, detect content to extract and dynamically create dataframe (with the relevant name) to contain the extracted content from these JSON files.</p>\n<p>The processing has just few steps:</p>\n<ul>\n<li>Browse the json files;</li>\n<li>For each json file, load (using json.load)<ul>\n<li>Browse the items in the current json;</li>\n<li>For each item:<ul>\n<li>Generate a dataframe using pandas.json_normalize</li>\n<li>Create a dynamical variable with a name composed from json name and the item name, to store the currently extracted content</li>\n<li>Assign the dataframe to the dynamical variable</li>\n<li>Save (to the output) the dataframe</li></ul></li></ul></li>\n</ul>\n<h3>💥Process the data</h3>\n<pre><code>json_folder_path = \"/kaggle/input/iwildcam2021-fgvc8/metadata\"\nlist_of_files = list(os.listdir(json_folder_path))\n\nfor file_name in list_of_files:\n    json_path = os.path.join(json_folder_path, file_name)\n    print(f\"Current json processed: {file_name}\")\n    with open(json_path) as json_file:\n        # read each json\n        json_data = json.load(json_file)\n        # for each item in the json\n        for item in json_data.items():\n            # prepare the dataframe name\n            file_name_split = file_name.split(\".\")[0]\n            file_name_split = file_name_split.split(\"_\")\n            file_name_str = file_name_split[1] + \"_\" + file_name_split[2]\n            print(f\"\\tjson item: {item[0]} length: {len(item[1])}\")\n            df_name = f\"{file_name_str}_{item[0]}_df\"\n            print(f\"\\tDynamic dataframe created: {data_frame_name}\")\n            # dynamic creation of a dataframe, using vars()[df_name]\n            vars()[df_name] = pd.json_normalize(json_data.get(item[0]))\n            # output the dataframe\n            vars()[df_name].to_csv(f\"{df_name}\", index=False)\n</code></pre>\n<h3>💥Resulted data</h3>\n<p>As a result of running this script, several dataframes are generated, containing the normalized information from the JSONs.</p>\n<pre><code>megadetector_results_info_df\nmegadetector_results_images_df\nmegadetector_results_detection_categories_df\ntest_information_images_df\ntrain_annotations_images_df\ntrain_annotations_annotations_df\ntrain_annotations_categories_df\n</code></pre>\n<p>Also, these dataframes are saved.</p>\n<h3>💥Notebook</h3>\n<p>You can find this code in the notebook: <a href=\"https://www.kaggle.com/gpreda/iwildcam2021-normalize-metadata-jsons\" target=\"_blank\">iWildCam2021 Normalize Metadata JSONs</a></p>",
  "messages": [
    {
      "id": "1246471",
      "postDate": "03/20/2021 19:06:02",
      "content": "<h3>💥Introduction</h3>\n<p>The annotations for train data, including the bounding boxes for detected animals, the metadata for train and test are provided in this competition in JSON format.</p>\n<p>I provide here a fast and simple method to bulk process these JSON files, detect content to extract and dynamically create dataframe (with the relevant name) to contain the extracted content from these JSON files.</p>\n<p>The processing has just few steps:</p>\n<ul>\n<li>Browse the json files;</li>\n<li>For each json file, load (using json.load)<ul>\n<li>Browse the items in the current json;</li>\n<li>For each item:<ul>\n<li>Generate a dataframe using pandas.json_normalize</li>\n<li>Create a dynamical variable with a name composed from json name and the item name, to store the currently extracted content</li>\n<li>Assign the dataframe to the dynamical variable</li>\n<li>Save (to the output) the dataframe</li></ul></li></ul></li>\n</ul>\n<h3>💥Process the data</h3>\n<pre><code>json_folder_path = \"/kaggle/input/iwildcam2021-fgvc8/metadata\"\nlist_of_files = list(os.listdir(json_folder_path))\n\nfor file_name in list_of_files:\n    json_path = os.path.join(json_folder_path, file_name)\n    print(f\"Current json processed: {file_name}\")\n    with open(json_path) as json_file:\n        # read each json\n        json_data = json.load(json_file)\n        # for each item in the json\n        for item in json_data.items():\n            # prepare the dataframe name\n            file_name_split = file_name.split(\".\")[0]\n            file_name_split = file_name_split.split(\"_\")\n            file_name_str = file_name_split[1] + \"_\" + file_name_split[2]\n            print(f\"\\tjson item: {item[0]} length: {len(item[1])}\")\n            df_name = f\"{file_name_str}_{item[0]}_df\"\n            print(f\"\\tDynamic dataframe created: {data_frame_name}\")\n            # dynamic creation of a dataframe, using vars()[df_name]\n            vars()[df_name] = pd.json_normalize(json_data.get(item[0]))\n            # output the dataframe\n            vars()[df_name].to_csv(f\"{df_name}\", index=False)\n</code></pre>\n<h3>💥Resulted data</h3>\n<p>As a result of running this script, several dataframes are generated, containing the normalized information from the JSONs.</p>\n<pre><code>megadetector_results_info_df\nmegadetector_results_images_df\nmegadetector_results_detection_categories_df\ntest_information_images_df\ntrain_annotations_images_df\ntrain_annotations_annotations_df\ntrain_annotations_categories_df\n</code></pre>\n<p>Also, these dataframes are saved.</p>\n<h3>💥Notebook</h3>\n<p>You can find this code in the notebook: <a href=\"https://www.kaggle.com/gpreda/iwildcam2021-normalize-metadata-jsons\" target=\"_blank\">iWildCam2021 Normalize Metadata JSONs</a></p>",
      "rawMarkdown": "### 💥Introduction\n\nThe annotations for train data, including the bounding boxes for detected animals, the metadata for train and test are provided in this competition in JSON format.\n\nI provide here a fast and simple method to bulk process these JSON files, detect content to extract and dynamically create dataframe (with the relevant name) to contain the extracted content from these JSON files.\n\nThe processing has just few steps:\n* Browse the json files;\n* For each json file, load (using json.load)\n   * Browse the items in the current json;\n    * For each item:\n       * Generate a dataframe using pandas.json_normalize\n       * Create a dynamical variable with a name composed from json name and the item name, to store the currently extracted content\n       * Assign the dataframe to the dynamical variable\n       * Save (to the output) the dataframe\n\n### 💥Process the data\n\n```\njson_folder_path = \"/kaggle/input/iwildcam2021-fgvc8/metadata\"\nlist_of_files = list(os.listdir(json_folder_path))\n\nfor file_name in list_of_files:\n    json_path = os.path.join(json_folder_path, file_name)\n    print(f\"Current json processed: {file_name}\")\n    with open(json_path) as json_file:\n        # read each json\n        json_data = json.load(json_file)\n        # for each item in the json\n        for item in json_data.items():\n            # prepare the dataframe name\n            file_name_split = file_name.split(\".\")[0]\n            file_name_split = file_name_split.split(\"_\")\n            file_name_str = file_name_split[1] + \"_\" + file_name_split[2]\n            print(f\"\\tjson item: {item[0]} length: {len(item[1])}\")\n            df_name = f\"{file_name_str}_{item[0]}_df\"\n            print(f\"\\tDynamic dataframe created: {data_frame_name}\")\n            # dynamic creation of a dataframe, using vars()[df_name]\n            vars()[df_name] = pd.json_normalize(json_data.get(item[0]))\n            # output the dataframe\n            vars()[df_name].to_csv(f\"{df_name}\", index=False)\n```\n\n### 💥Resulted data\n\nAs a result of running this script, several dataframes are generated, containing the normalized information from the JSONs.\n\n```\nmegadetector_results_info_df\nmegadetector_results_images_df\nmegadetector_results_detection_categories_df\ntest_information_images_df\ntrain_annotations_images_df\ntrain_annotations_annotations_df\ntrain_annotations_categories_df\n```\n\nAlso, these dataframes are saved.\n\n### 💥Notebook\n\nYou can find this code in the notebook: [iWildCam2021 Normalize Metadata JSONs](https://www.kaggle.com/gpreda/iwildcam2021-normalize-metadata-jsons)",
      "votes": null
    },
    {
      "id": "1250958",
      "postDate": "03/24/2021 11:45:34",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a> for sharing. This is an important toolset for me.</p>",
      "rawMarkdown": "Thank you @gpreda for sharing. This is an important toolset for me.",
      "votes": null
    },
    {
      "id": "1250969",
      "postDate": "03/24/2021 11:54:40",
      "content": "<p>This is an important toolset for me . Thank you <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>  for sharing</p>",
      "rawMarkdown": "This is an important toolset for me . Thank you @gpreda  for sharing",
      "votes": null
    },
    {
      "id": "1785622",
      "postDate": "05/12/2022 09:18:57",
      "content": "<p>Hello. I need json files for Farm Animals. My final project is Farm Animal detector with YoloV5 </p>",
      "rawMarkdown": "Hello. I need json files for Farm Animals. My final project is Farm Animal detector with YoloV5",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1250958,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/24/2021 11:45:34",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a> for sharing. This is an important toolset for me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1250969,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/24/2021 11:54:40",
      "content": "<p>This is an important toolset for me . Thank you <a href=\"https://www.kaggle.com/gpreda\" target=\"_blank\">@gpreda</a>  for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1785622,
      "author_name": "beatricegeaninarus",
      "author_url": "",
      "post_date": "05/12/2022 09:18:57",
      "content": "<p>Hello. I need json files for Farm Animals. My final project is Farm Animal detector with YoloV5 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1246471": "### 💥Introduction\n\nThe annotations for train data, including the bounding boxes for detected animals, the metadata for train and test are provided in this competition in JSON format.\n\nI provide here a fast and simple method to bulk process these JSON files, detect content to extract and dynamically create dataframe (with the relevant name) to contain the extracted content from these JSON files.\n\nThe processing has just few steps:\n* Browse the json files;\n* For each json file, load (using json.load)\n   * Browse the items in the current json;\n    * For each item:\n       * Generate a dataframe using pandas.json_normalize\n       * Create a dynamical variable with a name composed from json name and the item name, to store the currently extracted content\n       * Assign the dataframe to the dynamical variable\n       * Save (to the output) the dataframe\n\n### 💥Process the data\n\n```\njson_folder_path = \"/kaggle/input/iwildcam2021-fgvc8/metadata\"\nlist_of_files = list(os.listdir(json_folder_path))\n\nfor file_name in list_of_files:\n    json_path = os.path.join(json_folder_path, file_name)\n    print(f\"Current json processed: {file_name}\")\n    with open(json_path) as json_file:\n        # read each json\n        json_data = json.load(json_file)\n        # for each item in the json\n        for item in json_data.items():\n            # prepare the dataframe name\n            file_name_split = file_name.split(\".\")[0]\n            file_name_split = file_name_split.split(\"_\")\n            file_name_str = file_name_split[1] + \"_\" + file_name_split[2]\n            print(f\"\\tjson item: {item[0]} length: {len(item[1])}\")\n            df_name = f\"{file_name_str}_{item[0]}_df\"\n            print(f\"\\tDynamic dataframe created: {data_frame_name}\")\n            # dynamic creation of a dataframe, using vars()[df_name]\n            vars()[df_name] = pd.json_normalize(json_data.get(item[0]))\n            # output the dataframe\n            vars()[df_name].to_csv(f\"{df_name}\", index=False)\n```\n\n### 💥Resulted data\n\nAs a result of running this script, several dataframes are generated, containing the normalized information from the JSONs.\n\n```\nmegadetector_results_info_df\nmegadetector_results_images_df\nmegadetector_results_detection_categories_df\ntest_information_images_df\ntrain_annotations_images_df\ntrain_annotations_annotations_df\ntrain_annotations_categories_df\n```\n\nAlso, these dataframes are saved.\n\n### 💥Notebook\n\nYou can find this code in the notebook: [iWildCam2021 Normalize Metadata JSONs](https://www.kaggle.com/gpreda/iwildcam2021-normalize-metadata-jsons)",
    "1250958": "Thank you @gpreda for sharing. This is an important toolset for me.",
    "1250969": "This is an important toolset for me . Thank you @gpreda  for sharing",
    "1785622": "Hello. I need json files for Farm Animals. My final project is Farm Animal detector with YoloV5"
  },
  "source": "meta"
}