{
  "id": 252982,
  "title": "Issue with \"events\" table extracted from the train dataset",
  "url": "/competitions/mlb-player-digital-engagement-forecasting/discussion/252982",
  "author_name": "Serge Bushman",
  "post_date": "2021-07-14T14:17:34.803000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm a noob, and want to describe my extraction process from JSON.  I am working in R.</p>\n<p>I've borrowed heavily from Long Je at <a href=\"https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r\" target=\"_blank\">https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r</a> in the unpacking process.  I'm using Kaggle Notebook for the first time, using the R Kernel.  It has taken a little while to get accustomed to, as my normal IDE is RStudio Pro.  </p>\n<p>Yesterday I observed that Kaggle Notebook took a long time to process the MLB data, and that maybe Kaggle Notebook should be used primarily for presenting and sharing results rather than the nitty gritty work.  I therefore have started working in RStudio with the assumption that I would subsequently port my work to Kaggle Notebook for publication.  </p>\n<p>I've since discovered that my RStudio setup isn't much faster processing the data, though it is a more familiar setup.  I also noticed that simply copying and pasting kaggle Notebook doesn't make 1:1 sense for either an RScript or an Rmd, so I'm not sure how useful this will be going forward.</p>\n<p>But on to the topic of this post:  While working in RStudio I loaded the competition data and started manipulating it.  I borrowed Long Je's function for generating dataframes from the JSON. (<a href=\"https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r\" target=\"_blank\">https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r</a>)</p>\n<p>It didn't work.  Why not?  I kept getting an error that looked like this:  </p>\n<p>Error: parse error: premature EOF<br>\n                                                 [{\"gamePk\":634327,\"gameDate\":\"2<br>\n                     (right here) ------^</p>\n<p>There is 'some' documentation on this, but I found it confusing and instead relied on some recent experiences trying to publish text in a shiny app.  I learned that in order to make it work, I simply had to add a 'return' at the end of the file, thus creating an empty row at the end.</p>\n<p>I needed to remember that JSON files are essentially long text files.  I extracted the first cell using:</p>\n<p>events1 &lt; - fromJSON(events[1]) </p>\n<p>and learned that 'gamePK' is the name of the first column and 'gameDate' is the name of the second column.  the parse error appeared to point at the curly bracket near the beginning.  </p>\n<p>Looking closely at the data in 'gamePk' I figured out that it represents all the games played on a particular day.  I also figured out that 634327 could be found in the final row of the 'event' table.  This got me thinking I should look at the very end of the final entry to see if the final entry had been cropped in some way.</p>\n<p>Since there are 536 rows in the events database, I compared rows 535 and 536 to see if anything might be cut off.  I isolated these with the following:</p>\n<p>my535 &lt; - events[535]<br>\nmy536 &lt; - events[536]</p>\n<p>I wasn't sure how to look at these entries in RStudio, so I saved them as .txt files and downloaded the two documents.  Then I looked at them in my local notepad.</p>\n<p>Scrolling to the bottom I discovered that in fact, the 535 document concluded with \"}], and the 536 document had not.  In notepad, I added \"}]\" to 536, and saved it locally.  Then I uploaded it to RStudio Server Pro and loaded it into my environment.</p>\n<p>Then I extracted a dataframe from the new-formed JSON.</p>\n<p>my536 &lt; - fromJSON(my536)</p>\n<p>It worked!</p>",
  "messages": [
    {
      "id": 1387925,
      "postDate": "2021-07-14T14:17:34.803Z",
      "content": "<p>I'm a noob, and want to describe my extraction process from JSON.  I am working in R.</p>\n<p>I've borrowed heavily from Long Je at <a href=\"https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r\" target=\"_blank\">https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r</a> in the unpacking process.  I'm using Kaggle Notebook for the first time, using the R Kernel.  It has taken a little while to get accustomed to, as my normal IDE is RStudio Pro.  </p>\n<p>Yesterday I observed that Kaggle Notebook took a long time to process the MLB data, and that maybe Kaggle Notebook should be used primarily for presenting and sharing results rather than the nitty gritty work.  I therefore have started working in RStudio with the assumption that I would subsequently port my work to Kaggle Notebook for publication.  </p>\n<p>I've since discovered that my RStudio setup isn't much faster processing the data, though it is a more familiar setup.  I also noticed that simply copying and pasting kaggle Notebook doesn't make 1:1 sense for either an RScript or an Rmd, so I'm not sure how useful this will be going forward.</p>\n<p>But on to the topic of this post:  While working in RStudio I loaded the competition data and started manipulating it.  I borrowed Long Je's function for generating dataframes from the JSON. (<a href=\"https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r\" target=\"_blank\">https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r</a>)</p>\n<p>It didn't work.  Why not?  I kept getting an error that looked like this:  </p>\n<p>Error: parse error: premature EOF<br>\n                                                 [{\"gamePk\":634327,\"gameDate\":\"2<br>\n                     (right here) ------^</p>\n<p>There is 'some' documentation on this, but I found it confusing and instead relied on some recent experiences trying to publish text in a shiny app.  I learned that in order to make it work, I simply had to add a 'return' at the end of the file, thus creating an empty row at the end.</p>\n<p>I needed to remember that JSON files are essentially long text files.  I extracted the first cell using:</p>\n<p>events1 &lt; - fromJSON(events[1]) </p>\n<p>and learned that 'gamePK' is the name of the first column and 'gameDate' is the name of the second column.  the parse error appeared to point at the curly bracket near the beginning.  </p>\n<p>Looking closely at the data in 'gamePk' I figured out that it represents all the games played on a particular day.  I also figured out that 634327 could be found in the final row of the 'event' table.  This got me thinking I should look at the very end of the final entry to see if the final entry had been cropped in some way.</p>\n<p>Since there are 536 rows in the events database, I compared rows 535 and 536 to see if anything might be cut off.  I isolated these with the following:</p>\n<p>my535 &lt; - events[535]<br>\nmy536 &lt; - events[536]</p>\n<p>I wasn't sure how to look at these entries in RStudio, so I saved them as .txt files and downloaded the two documents.  Then I looked at them in my local notepad.</p>\n<p>Scrolling to the bottom I discovered that in fact, the 535 document concluded with \"}], and the 536 document had not.  In notepad, I added \"}]\" to 536, and saved it locally.  Then I uploaded it to RStudio Server Pro and loaded it into my environment.</p>\n<p>Then I extracted a dataframe from the new-formed JSON.</p>\n<p>my536 &lt; - fromJSON(my536)</p>\n<p>It worked!</p>",
      "rawMarkdown": "I'm a noob, and want to describe my extraction process from JSON.  I am working in R.\n\nI've borrowed heavily from Long Je at https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r in the unpacking process.  I'm using Kaggle Notebook for the first time, using the R Kernel.  It has taken a little while to get accustomed to, as my normal IDE is RStudio Pro.  \n\nYesterday I observed that Kaggle Notebook took a long time to process the MLB data, and that maybe Kaggle Notebook should be used primarily for presenting and sharing results rather than the nitty gritty work.  I therefore have started working in RStudio with the assumption that I would subsequently port my work to Kaggle Notebook for publication.  \n\nI've since discovered that my RStudio setup isn't much faster processing the data, though it is a more familiar setup.  I also noticed that simply copying and pasting kaggle Notebook doesn't make 1:1 sense for either an RScript or an Rmd, so I'm not sure how useful this will be going forward.\n\nBut on to the topic of this post:  While working in RStudio I loaded the competition data and started manipulating it.  I borrowed Long Je's function for generating dataframes from the JSON. (https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r)\n\nIt didn't work.  Why not?  I kept getting an error that looked like this:  \n\nError: parse error: premature EOF\n                                                 [{\"gamePk\":634327,\"gameDate\":\"2\n                     (right here) ------^\n\nThere is 'some' documentation on this, but I found it confusing and instead relied on some recent experiences trying to publish text in a shiny app.  I learned that in order to make it work, I simply had to add a 'return' at the end of the file, thus creating an empty row at the end.\n\nI needed to remember that JSON files are essentially long text files.  I extracted the first cell using:\n\nevents1 < - fromJSON(events[1]) \n\nand learned that 'gamePK' is the name of the first column and 'gameDate' is the name of the second column.  the parse error appeared to point at the curly bracket near the beginning.  \n\nLooking closely at the data in 'gamePk' I figured out that it represents all the games played on a particular day.  I also figured out that 634327 could be found in the final row of the 'event' table.  This got me thinking I should look at the very end of the final entry to see if the final entry had been cropped in some way.\n\nSince there are 536 rows in the events database, I compared rows 535 and 536 to see if anything might be cut off.  I isolated these with the following:\n\nmy535 < - events[535]\nmy536 < - events[536]\n\nI wasn't sure how to look at these entries in RStudio, so I saved them as .txt files and downloaded the two documents.  Then I looked at them in my local notepad.\n\nScrolling to the bottom I discovered that in fact, the 535 document concluded with \"}], and the 536 document had not.  In notepad, I added \"}]\" to 536, and saved it locally.  Then I uploaded it to RStudio Server Pro and loaded it into my environment.\n\nThen I extracted a dataframe from the new-formed JSON.\n\nmy536 < - fromJSON(my536)\n\nIt worked!\n\n\n  \n\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1387925": "I'm a noob, and want to describe my extraction process from JSON.  I am working in R.\n\nI've borrowed heavily from Long Je at https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r in the unpacking process.  I'm using Kaggle Notebook for the first time, using the R Kernel.  It has taken a little while to get accustomed to, as my normal IDE is RStudio Pro.  \n\nYesterday I observed that Kaggle Notebook took a long time to process the MLB data, and that maybe Kaggle Notebook should be used primarily for presenting and sharing results rather than the nitty gritty work.  I therefore have started working in RStudio with the assumption that I would subsequently port my work to Kaggle Notebook for publication.  \n\nI've since discovered that my RStudio setup isn't much faster processing the data, though it is a more familiar setup.  I also noticed that simply copying and pasting kaggle Notebook doesn't make 1:1 sense for either an RScript or an Rmd, so I'm not sure how useful this will be going forward.\n\nBut on to the topic of this post:  While working in RStudio I loaded the competition data and started manipulating it.  I borrowed Long Je's function for generating dataframes from the JSON. (https://www.kaggle.com/dhlongle/mlb-digital-engagement-unpack-train-csv-with-r)\n\nIt didn't work.  Why not?  I kept getting an error that looked like this:  \n\nError: parse error: premature EOF\n                                                 [{\"gamePk\":634327,\"gameDate\":\"2\n                     (right here) ------^\n\nThere is 'some' documentation on this, but I found it confusing and instead relied on some recent experiences trying to publish text in a shiny app.  I learned that in order to make it work, I simply had to add a 'return' at the end of the file, thus creating an empty row at the end.\n\nI needed to remember that JSON files are essentially long text files.  I extracted the first cell using:\n\nevents1 < - fromJSON(events[1]) \n\nand learned that 'gamePK' is the name of the first column and 'gameDate' is the name of the second column.  the parse error appeared to point at the curly bracket near the beginning.  \n\nLooking closely at the data in 'gamePk' I figured out that it represents all the games played on a particular day.  I also figured out that 634327 could be found in the final row of the 'event' table.  This got me thinking I should look at the very end of the final entry to see if the final entry had been cropped in some way.\n\nSince there are 536 rows in the events database, I compared rows 535 and 536 to see if anything might be cut off.  I isolated these with the following:\n\nmy535 < - events[535]\nmy536 < - events[536]\n\nI wasn't sure how to look at these entries in RStudio, so I saved them as .txt files and downloaded the two documents.  Then I looked at them in my local notepad.\n\nScrolling to the bottom I discovered that in fact, the 535 document concluded with \"}], and the 536 document had not.  In notepad, I added \"}]\" to 536, and saved it locally.  Then I uploaded it to RStudio Server Pro and loaded it into my environment.\n\nThen I extracted a dataframe from the new-formed JSON.\n\nmy536 < - fromJSON(my536)\n\nIt worked!\n\n\n  \n\n\n"
  }
}