{
  "id": 225237,
  "title": "Quick Way to Convert JSON to CSV + CSV Dataset Link",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/225237",
  "author_name": "Tom M",
  "post_date": "2021-03-11T11:34:05.201000",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>This is a quick way to convert the important information from the JSON file to a dataframe.</p>\n<pre><code>import pandas as pd\nimport json\n\nwith open(\"../input/herbarium-2021-fgvc8/train/metadata.json\") as json_meta:\n    data = json.load(json_meta)\n\nannotations_df = pd.json_normalize(data, record_path = ['annotations'])\nimages_df = pd.json_normalize(data, record_path = ['images'])\n\ntrain_df = annotations_df.merge(images_df, on = \"id\")\n\ncategories_df = pd.json_normalize(data, record_path = [\"categories\"])\ncategories_df['category_id'] = categories_df['id']\ncategories_df = categories_df.drop(['id'], axis = 1)\n\nall_df = train_df.merge(categories_df, on = \"category_id\")\n</code></pre>\n<p><strong>I have made this output available in kaggle datasets for ease:</strong> Simply add it to your kernel in the add data tab on the top right.  Here's the link.  <a href=\"https://www.kaggle.com/tpmeli/herbarium-traincsv\" target=\"_blank\">https://www.kaggle.com/tpmeli/herbarium-traincsv</a></p>\n<p><strong>Demonstration Starter:</strong> <a href=\"https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu\" target=\"_blank\">https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu</a></p>",
  "messages": [
    {
      "id": 1234587,
      "postDate": "2021-03-11T11:34:05.203Z",
      "content": "<p>This is a quick way to convert the important information from the JSON file to a dataframe.</p>\n<pre><code>import pandas as pd\nimport json\n\nwith open(\"../input/herbarium-2021-fgvc8/train/metadata.json\") as json_meta:\n    data = json.load(json_meta)\n\nannotations_df = pd.json_normalize(data, record_path = ['annotations'])\nimages_df = pd.json_normalize(data, record_path = ['images'])\n\ntrain_df = annotations_df.merge(images_df, on = \"id\")\n\ncategories_df = pd.json_normalize(data, record_path = [\"categories\"])\ncategories_df['category_id'] = categories_df['id']\ncategories_df = categories_df.drop(['id'], axis = 1)\n\nall_df = train_df.merge(categories_df, on = \"category_id\")\n</code></pre>\n<p><strong>I have made this output available in kaggle datasets for ease:</strong> Simply add it to your kernel in the add data tab on the top right.  Here's the link.  <a href=\"https://www.kaggle.com/tpmeli/herbarium-traincsv\" target=\"_blank\">https://www.kaggle.com/tpmeli/herbarium-traincsv</a></p>\n<p><strong>Demonstration Starter:</strong> <a href=\"https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu\" target=\"_blank\">https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu</a></p>",
      "rawMarkdown": "This is a quick way to convert the important information from the JSON file to a dataframe.\n\n```\nimport pandas as pd\nimport json\n\nwith open(\"../input/herbarium-2021-fgvc8/train/metadata.json\") as json_meta:\n    data = json.load(json_meta)\n\nannotations_df = pd.json_normalize(data, record_path = ['annotations'])\nimages_df = pd.json_normalize(data, record_path = ['images'])\n\ntrain_df = annotations_df.merge(images_df, on = \"id\")\n\ncategories_df = pd.json_normalize(data, record_path = [\"categories\"])\ncategories_df['category_id'] = categories_df['id']\ncategories_df = categories_df.drop(['id'], axis = 1)\n\nall_df = train_df.merge(categories_df, on = \"category_id\")\n\n```\n\n**I have made this output available in kaggle datasets for ease:** Simply add it to your kernel in the add data tab on the top right.  Here's the link.  [https://www.kaggle.com/tpmeli/herbarium-traincsv](https://www.kaggle.com/tpmeli/herbarium-traincsv)\n\n**Demonstration Starter:** [https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu](https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu)",
      "votes": 14
    },
    {
      "id": 1235386,
      "postDate": "2021-03-12T05:25:09.640Z",
      "content": "<p><a href=\"https://www.kaggle.com/tpmeli\" target=\"_blank\">@tpmeli</a> Thanks so much for the utility function to convert json to csv file . Very helpful . Looking forward to your completed notebook</p>",
      "rawMarkdown": "@tpmeli Thanks so much for the utility function to convert json to csv file . Very helpful . Looking forward to your completed notebook",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1235386,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-03-12T05:25:09.640000",
      "content": "<p><a href=\"https://www.kaggle.com/tpmeli\" target=\"_blank\">@tpmeli</a> Thanks so much for the utility function to convert json to csv file . Very helpful . Looking forward to your completed notebook</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1234587": "This is a quick way to convert the important information from the JSON file to a dataframe.\n\n```\nimport pandas as pd\nimport json\n\nwith open(\"../input/herbarium-2021-fgvc8/train/metadata.json\") as json_meta:\n    data = json.load(json_meta)\n\nannotations_df = pd.json_normalize(data, record_path = ['annotations'])\nimages_df = pd.json_normalize(data, record_path = ['images'])\n\ntrain_df = annotations_df.merge(images_df, on = \"id\")\n\ncategories_df = pd.json_normalize(data, record_path = [\"categories\"])\ncategories_df['category_id'] = categories_df['id']\ncategories_df = categories_df.drop(['id'], axis = 1)\n\nall_df = train_df.merge(categories_df, on = \"category_id\")\n\n```\n\n**I have made this output available in kaggle datasets for ease:** Simply add it to your kernel in the add data tab on the top right.  Here's the link.  [https://www.kaggle.com/tpmeli/herbarium-traincsv](https://www.kaggle.com/tpmeli/herbarium-traincsv)\n\n**Demonstration Starter:** [https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu](https://www.kaggle.com/tpmeli/herbarium-starter-efficientnet-tf-keras-gpu)",
    "1235386": "@tpmeli Thanks so much for the utility function to convert json to csv file . Very helpful . Looking forward to your completed notebook"
  }
}