{
  "id": 252594,
  "title": "Parsing json data for beginners.",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/252594",
  "author_name": "",
  "post_date": "2021-07-13T04:15:19.535302300Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm a beginner in data science and python. I found it very difficult for me to understand and deal with json format data.</p>\n<p>This post is for beginners who have the same problem. </p>\n<p>We first load the data</p>\n<pre><code>train = pd.read_csv('data/train.csv')\n</code></pre>\n<p>Then start parsing:</p>\n<pre><code># You can put any column you want to parse here:\ncolumn_to_parse=['games']\n\n# Parsing start:\nfor col in column_to_parse:\n    print(f\"Parse columns {col}\")\n\n    # Start a dataframe to capture the parsed data.\n    target_df = pd.DataFrame()\n    target_df['date'] = train['date']\n\n    for i in range(train.shape[0]):\n        # choose each cell in a certain column.\n        cell = train.iloc[i][col]\n\n        if cell is not None:\n            # the try/except will help you skip the NaN values.\n            try:\n                value = json.loads(cell)\n            except TypeError:\n                continue\n\n            for k, v in enumerate(value):\n                prefix_col = f'{col}_{k}'\n\n                for key, values in v.items():\n                    new_col = f'{prefix_col}_{key}'\n\n                    # this trt/except will help you deal with different data types.\n                    try:\n                        target_df.at[i, new_col] = values\n                    except ValueError:\n                        target_df[new_col] = target_df[new_col].astype(object)\n                        target_df.at[i, new_col] = values\n\n    target_df.to_csv('data/train_vs/'+ col + '.csv')\n</code></pre>",
  "messages": [
    {
      "id": "1385812",
      "postDate": "07/13/2021 04:15:19",
      "content": "<p>I'm a beginner in data science and python. I found it very difficult for me to understand and deal with json format data.</p>\n<p>This post is for beginners who have the same problem. </p>\n<p>We first load the data</p>\n<pre><code>train = pd.read_csv('data/train.csv')\n</code></pre>\n<p>Then start parsing:</p>\n<pre><code># You can put any column you want to parse here:\ncolumn_to_parse=['games']\n\n# Parsing start:\nfor col in column_to_parse:\n    print(f\"Parse columns {col}\")\n\n    # Start a dataframe to capture the parsed data.\n    target_df = pd.DataFrame()\n    target_df['date'] = train['date']\n\n    for i in range(train.shape[0]):\n        # choose each cell in a certain column.\n        cell = train.iloc[i][col]\n\n        if cell is not None:\n            # the try/except will help you skip the NaN values.\n            try:\n                value = json.loads(cell)\n            except TypeError:\n                continue\n\n            for k, v in enumerate(value):\n                prefix_col = f'{col}_{k}'\n\n                for key, values in v.items():\n                    new_col = f'{prefix_col}_{key}'\n\n                    # this trt/except will help you deal with different data types.\n                    try:\n                        target_df.at[i, new_col] = values\n                    except ValueError:\n                        target_df[new_col] = target_df[new_col].astype(object)\n                        target_df.at[i, new_col] = values\n\n    target_df.to_csv('data/train_vs/'+ col + '.csv')\n</code></pre>",
      "rawMarkdown": "I'm a beginner in data science and python. I found it very difficult for me to understand and deal with json format data.\n\nThis post is for beginners who have the same problem. \n\nWe first load the data\n```\ntrain = pd.read_csv('data/train.csv')\n```\n\nThen start parsing:\n```\n# You can put any column you want to parse here:\ncolumn_to_parse=['games']\n\n# Parsing start:\nfor col in column_to_parse:\n    print(f\"Parse columns {col}\")\n\n    # Start a dataframe to capture the parsed data.\n    target_df = pd.DataFrame()\n    target_df['date'] = train['date']\n\n    for i in range(train.shape[0]):\n        # choose each cell in a certain column.\n        cell = train.iloc[i][col]\n\n        if cell is not None:\n            # the try/except will help you skip the NaN values.\n            try:\n                value = json.loads(cell)\n            except TypeError:\n                continue\n\n            for k, v in enumerate(value):\n                prefix_col = f'{col}_{k}'\n\n                for key, values in v.items():\n                    new_col = f'{prefix_col}_{key}'\n\n                    # this trt/except will help you deal with different data types.\n                    try:\n                        target_df.at[i, new_col] = values\n                    except ValueError:\n                        target_df[new_col] = target_df[new_col].astype(object)\n                        target_df.at[i, new_col] = values\n\n    target_df.to_csv('data/train_vs/'+ col + '.csv')\n```",
      "votes": null
    },
    {
      "id": "1386711",
      "postDate": "07/13/2021 16:32:52",
      "content": "<p><a href=\"https://www.kaggle.com/yingzhou0510\" target=\"_blank\">@yingzhou0510</a> Pandas has a read_json function that you can use to read in the json formatted data, see <a href=\"https://pandas.pydata.org/pandas-docs/version/1.1.3/reference/api/pandas.read_json.html\" target=\"_blank\">here</a> for the documentation. An example of using the function with the MLB data can be found in the <br>\n<a href=\"https://www.kaggle.com/alokpattani/mlb-player-digital-engagement-data-exploration/comments#Read-in-Kaggle-Data-Files\" target=\"_blank\">MLB Player Digital Engagement Data Exploration notebook</a>.</p>",
      "rawMarkdown": "yingzhou0510 Pandas has a read_json function that you can use to read in the json formatted data, see [here](https://pandas.pydata.org/pandas-docs/version/1.1.3/reference/api/pandas.read_json.html) for the documentation. An example of using the function with the MLB data can be found in the \n[MLB Player Digital Engagement Data Exploration notebook](https://www.kaggle.com/alokpattani/mlb-player-digital-engagement-data-exploration/comments#Read-in-Kaggle-Data-Files).",
      "votes": null
    },
    {
      "id": "1387084",
      "postDate": "07/14/2021 00:06:29",
      "content": "<p>Thanks for doing this.  I wish someone could help me on the R side with this like you have helped on the python side.  I'm able to make the parsing work in Kaggle notebooks but not on my own RStudio setup :-(</p>",
      "rawMarkdown": "Thanks for doing this.  I wish someone could help me on the R side with this like you have helped on the python side.  I'm able to make the parsing work in Kaggle notebooks but not on my own RStudio setup :-(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1386711,
      "author_name": "bryanmariscal",
      "author_url": "",
      "post_date": "07/13/2021 16:32:52",
      "content": "<p><a href=\"https://www.kaggle.com/yingzhou0510\" target=\"_blank\">@yingzhou0510</a> Pandas has a read_json function that you can use to read in the json formatted data, see <a href=\"https://pandas.pydata.org/pandas-docs/version/1.1.3/reference/api/pandas.read_json.html\" target=\"_blank\">here</a> for the documentation. An example of using the function with the MLB data can be found in the <br>\n<a href=\"https://www.kaggle.com/alokpattani/mlb-player-digital-engagement-data-exploration/comments#Read-in-Kaggle-Data-Files\" target=\"_blank\">MLB Player Digital Engagement Data Exploration notebook</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1387084,
      "author_name": "sbushman",
      "author_url": "",
      "post_date": "07/14/2021 00:06:29",
      "content": "<p>Thanks for doing this.  I wish someone could help me on the R side with this like you have helped on the python side.  I'm able to make the parsing work in Kaggle notebooks but not on my own RStudio setup :-(</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1385812": "I'm a beginner in data science and python. I found it very difficult for me to understand and deal with json format data.\n\nThis post is for beginners who have the same problem. \n\nWe first load the data\n```\ntrain = pd.read_csv('data/train.csv')\n```\n\nThen start parsing:\n```\n# You can put any column you want to parse here:\ncolumn_to_parse=['games']\n\n# Parsing start:\nfor col in column_to_parse:\n    print(f\"Parse columns {col}\")\n\n    # Start a dataframe to capture the parsed data.\n    target_df = pd.DataFrame()\n    target_df['date'] = train['date']\n\n    for i in range(train.shape[0]):\n        # choose each cell in a certain column.\n        cell = train.iloc[i][col]\n\n        if cell is not None:\n            # the try/except will help you skip the NaN values.\n            try:\n                value = json.loads(cell)\n            except TypeError:\n                continue\n\n            for k, v in enumerate(value):\n                prefix_col = f'{col}_{k}'\n\n                for key, values in v.items():\n                    new_col = f'{prefix_col}_{key}'\n\n                    # this trt/except will help you deal with different data types.\n                    try:\n                        target_df.at[i, new_col] = values\n                    except ValueError:\n                        target_df[new_col] = target_df[new_col].astype(object)\n                        target_df.at[i, new_col] = values\n\n    target_df.to_csv('data/train_vs/'+ col + '.csv')\n```",
    "1386711": "yingzhou0510 Pandas has a read_json function that you can use to read in the json formatted data, see [here](https://pandas.pydata.org/pandas-docs/version/1.1.3/reference/api/pandas.read_json.html) for the documentation. An example of using the function with the MLB data can be found in the \n[MLB Player Digital Engagement Data Exploration notebook](https://www.kaggle.com/alokpattani/mlb-player-digital-engagement-data-exploration/comments#Read-in-Kaggle-Data-Files).",
    "1387084": "Thanks for doing this.  I wish someone could help me on the R side with this like you have helped on the python side.  I'm able to make the parsing work in Kaggle notebooks but not on my own RStudio setup :-("
  },
  "source": "meta"
}