{
  "id": 138111,
  "title": "About the stage 2 data",
  "url": "/competitions/march-madness-analytics-2020/discussion/138111",
  "author_name": "",
  "post_date": "2020-03-23T20:22:58.775887300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone, here's an update regarding the stage 2 data.  I am making this same post for all three contests. </p>\n\n<p>I have brought the season 2020 data up to the present, with its ragged termination around day 128-129, and you will see new Stage 2 zip files on the Data page.  Thiese data updates are reflected in the various files such as RegularSeasonCompactResults, RegularSeasonDetailedResults, and GameCities, along with several others.  In addition, I decided to rebuild the entire history of play-by-play, including for seasons prior to 2020.  The play-by-play now includes a file for the 2020 season, up through its termination at day 128/129.</p>\n\n<p>Therefore there is a new men’s data files zip (MDataFiles_Stage2.zip), a new men’s play-by-play zip (MPlayByPlay_Stage2.zip), and the same two files on the women’s side (WDataFiles_Stage2.zip and WPlayByPlay_Stage2.zip).  The file names indicate that they are stage 2 data files.  So each prediction contest has two new files available on its Data page, and the analytics contest (which covers both men’s and women’s) has four new files available on its Data page.</p>\n\n<p>I decided to refresh the entire play-by-play data (including previous years) because I know there were some bug fixes upstream of my work, that would hopefully result in the running score (as the game proceeded) being populated rather than always zero, as well as a significant problem that was fixed in the elapsed seconds of the women’s data.  My understanding is that these fixes were implemented after I pulled the play-by-play data for stage 1, and so they should be fixed in this latest stage 2 release, now that I've pulled the data again.  The play-by-play files from previous years seem a bit larger now, so I expect there are a few more games covered than before - unless it’s just that there are nonzero running scores in there rather than zeroes, taking up a few extra bytes per row.</p>\n\n<p>The Massey Ordinals (men’s only) are only populated up through Day 128 - usually there is a Day 133 but the season ended early.  I think there may be one more round of ordinals posted on the site but I am not sure how complete that data is.</p>\n\n<p>Everything should be up to date now, although obviously there are no new NCAA Tourney seeds or results or anything like that.  We don’t even get to know which regions were W, X, Y, or Z!  Anyway, these updates should provide useful extra information for anyone still working on the analytics contest, and or people wrapping up their work on the predictions contest.  Please post any issues you find in these data releases, to this thread for your contest, and hopefully I can address them.</p>",
  "messages": [
    {
      "id": "783956",
      "postDate": "03/23/2020 20:22:58",
      "content": "<p>Hi everyone, here's an update regarding the stage 2 data.  I am making this same post for all three contests. </p>\n\n<p>I have brought the season 2020 data up to the present, with its ragged termination around day 128-129, and you will see new Stage 2 zip files on the Data page.  Thiese data updates are reflected in the various files such as RegularSeasonCompactResults, RegularSeasonDetailedResults, and GameCities, along with several others.  In addition, I decided to rebuild the entire history of play-by-play, including for seasons prior to 2020.  The play-by-play now includes a file for the 2020 season, up through its termination at day 128/129.</p>\n\n<p>Therefore there is a new men’s data files zip (MDataFiles_Stage2.zip), a new men’s play-by-play zip (MPlayByPlay_Stage2.zip), and the same two files on the women’s side (WDataFiles_Stage2.zip and WPlayByPlay_Stage2.zip).  The file names indicate that they are stage 2 data files.  So each prediction contest has two new files available on its Data page, and the analytics contest (which covers both men’s and women’s) has four new files available on its Data page.</p>\n\n<p>I decided to refresh the entire play-by-play data (including previous years) because I know there were some bug fixes upstream of my work, that would hopefully result in the running score (as the game proceeded) being populated rather than always zero, as well as a significant problem that was fixed in the elapsed seconds of the women’s data.  My understanding is that these fixes were implemented after I pulled the play-by-play data for stage 1, and so they should be fixed in this latest stage 2 release, now that I've pulled the data again.  The play-by-play files from previous years seem a bit larger now, so I expect there are a few more games covered than before - unless it’s just that there are nonzero running scores in there rather than zeroes, taking up a few extra bytes per row.</p>\n\n<p>The Massey Ordinals (men’s only) are only populated up through Day 128 - usually there is a Day 133 but the season ended early.  I think there may be one more round of ordinals posted on the site but I am not sure how complete that data is.</p>\n\n<p>Everything should be up to date now, although obviously there are no new NCAA Tourney seeds or results or anything like that.  We don’t even get to know which regions were W, X, Y, or Z!  Anyway, these updates should provide useful extra information for anyone still working on the analytics contest, and or people wrapping up their work on the predictions contest.  Please post any issues you find in these data releases, to this thread for your contest, and hopefully I can address them.</p>",
      "rawMarkdown": "Hi everyone, here's an update regarding the stage 2 data.  I am making this same post for all three contests. \n\nI have brought the season 2020 data up to the present, with its ragged termination around day 128-129, and you will see new Stage 2 zip files on the Data page.  Thiese data updates are reflected in the various files such as RegularSeasonCompactResults, RegularSeasonDetailedResults, and GameCities, along with several others.  In addition, I decided to rebuild the entire history of play-by-play, including for seasons prior to 2020.  The play-by-play now includes a file for the 2020 season, up through its termination at day 128/129.\n\nTherefore there is a new men’s data files zip (MDataFiles_Stage2.zip), a new men’s play-by-play zip (MPlayByPlay_Stage2.zip), and the same two files on the women’s side (WDataFiles_Stage2.zip and WPlayByPlay_Stage2.zip).  The file names indicate that they are stage 2 data files.  So each prediction contest has two new files available on its Data page, and the analytics contest (which covers both men’s and women’s) has four new files available on its Data page.\n\nI decided to refresh the entire play-by-play data (including previous years) because I know there were some bug fixes upstream of my work, that would hopefully result in the running score (as the game proceeded) being populated rather than always zero, as well as a significant problem that was fixed in the elapsed seconds of the women’s data.  My understanding is that these fixes were implemented after I pulled the play-by-play data for stage 1, and so they should be fixed in this latest stage 2 release, now that I've pulled the data again.  The play-by-play files from previous years seem a bit larger now, so I expect there are a few more games covered than before - unless it’s just that there are nonzero running scores in there rather than zeroes, taking up a few extra bytes per row.\n\nThe Massey Ordinals (men’s only) are only populated up through Day 128 - usually there is a Day 133 but the season ended early.  I think there may be one more round of ordinals posted on the site but I am not sure how complete that data is.\n\nEverything should be up to date now, although obviously there are no new NCAA Tourney seeds or results or anything like that.  We don’t even get to know which regions were W, X, Y, or Z!  Anyway, these updates should provide useful extra information for anyone still working on the analytics contest, and or people wrapping up their work on the predictions contest.  Please post any issues you find in these data releases, to this thread for your contest, and hopefully I can address them.",
      "votes": null
    },
    {
      "id": "792787",
      "postDate": "03/31/2020 14:57:25",
      "content": "<p><a href=\"/jeffsonas\">@jeffsonas</a> Thanks for update. </p>",
      "rawMarkdown": "jeffsonas Thanks for update.",
      "votes": null
    },
    {
      "id": "795766",
      "postDate": "04/03/2020 02:58:44",
      "content": "<p>Very obscure issue with the data, but the LCurrentScore and WCurrentScore values in the play by play data are swapped for a game between New Mexico St and Colorado St that happened early this season on the men's side. Strangely, its not for the entire game, just part of overtime. The final few play by play entries have Colorado St winning 78 - 70 but it should be reverse. I'm sure other mistakes like this could be found by just searching for games where LCurrentScore is higher than WCurrentScore for the final entry.</p>\n\n<p>Another such example is Penn St vs Syracuse this season. WCurrentScore and LCurrentScore are swapped for the final 2 minutes of the game.</p>",
      "rawMarkdown": "Very obscure issue with the data, but the LCurrentScore and WCurrentScore values in the play by play data are swapped for a game between New Mexico St and Colorado St that happened early this season on the men's side. Strangely, its not for the entire game, just part of overtime. The final few play by play entries have Colorado St winning 78 - 70 but it should be reverse. I'm sure other mistakes like this could be found by just searching for games where LCurrentScore is higher than WCurrentScore for the final entry.\n\nAnother such example is Penn St vs Syracuse this season. WCurrentScore and LCurrentScore are swapped for the final 2 minutes of the game.",
      "votes": null
    },
    {
      "id": "796909",
      "postDate": "04/04/2020 04:06:07",
      "content": "<p><a href=\"/hmtessier\">@hmtessier</a>: many thanks for taking the time to investigate this issue. <a href=\"/jeffsonas\">@jeffsonas</a>: kindly investigate this one. Play-By-Play data is highly critical in this analytics competitions.</p>",
      "rawMarkdown": "hmtessier: many thanks for taking the time to investigate this issue. @jeffsonas: kindly investigate this one. Play-By-Play data is highly critical in this analytics competitions.",
      "votes": null
    },
    {
      "id": "796983",
      "postDate": "04/04/2020 06:45:20",
      "content": "<p>This is not data that I made any changes to during the import process - it is straight from the source.  It is possible to calculate WCurrentScore and LCurrentScore yourself and so perhaps people may want to do that themselves, to ensure higher quality data if this does seem extensive.  But there isn't anything I can easily do to resolve it, other than calculating it myself from the game events, and that has its own problems.  I don't think there is anything we should do here for this.</p>",
      "rawMarkdown": "This is not data that I made any changes to during the import process - it is straight from the source.  It is possible to calculate WCurrentScore and LCurrentScore yourself and so perhaps people may want to do that themselves, to ensure higher quality data if this does seem extensive.  But there isn't anything I can easily do to resolve it, other than calculating it myself from the game events, and that has its own problems.  I don't think there is anything we should do here for this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 792787,
      "author_name": "anshumoudgil",
      "author_url": "",
      "post_date": "03/31/2020 14:57:25",
      "content": "<p><a href=\"/jeffsonas\">@jeffsonas</a> Thanks for update. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 795766,
      "author_name": "hmtessier",
      "author_url": "",
      "post_date": "04/03/2020 02:58:44",
      "content": "<p>Very obscure issue with the data, but the LCurrentScore and WCurrentScore values in the play by play data are swapped for a game between New Mexico St and Colorado St that happened early this season on the men's side. Strangely, its not for the entire game, just part of overtime. The final few play by play entries have Colorado St winning 78 - 70 but it should be reverse. I'm sure other mistakes like this could be found by just searching for games where LCurrentScore is higher than WCurrentScore for the final entry.</p>\n\n<p>Another such example is Penn St vs Syracuse this season. WCurrentScore and LCurrentScore are swapped for the final 2 minutes of the game.</p>",
      "votes": null,
      "replies": [
        {
          "id": 796909,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "04/04/2020 04:06:07",
          "content": "<p><a href=\"/hmtessier\">@hmtessier</a>: many thanks for taking the time to investigate this issue. <a href=\"/jeffsonas\">@jeffsonas</a>: kindly investigate this one. Play-By-Play data is highly critical in this analytics competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 796983,
          "author_name": "jeffsonas",
          "author_url": "",
          "post_date": "04/04/2020 06:45:20",
          "content": "<p>This is not data that I made any changes to during the import process - it is straight from the source.  It is possible to calculate WCurrentScore and LCurrentScore yourself and so perhaps people may want to do that themselves, to ensure higher quality data if this does seem extensive.  But there isn't anything I can easily do to resolve it, other than calculating it myself from the game events, and that has its own problems.  I don't think there is anything we should do here for this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "783956": "Hi everyone, here's an update regarding the stage 2 data.  I am making this same post for all three contests. \n\nI have brought the season 2020 data up to the present, with its ragged termination around day 128-129, and you will see new Stage 2 zip files on the Data page.  Thiese data updates are reflected in the various files such as RegularSeasonCompactResults, RegularSeasonDetailedResults, and GameCities, along with several others.  In addition, I decided to rebuild the entire history of play-by-play, including for seasons prior to 2020.  The play-by-play now includes a file for the 2020 season, up through its termination at day 128/129.\n\nTherefore there is a new men’s data files zip (MDataFiles_Stage2.zip), a new men’s play-by-play zip (MPlayByPlay_Stage2.zip), and the same two files on the women’s side (WDataFiles_Stage2.zip and WPlayByPlay_Stage2.zip).  The file names indicate that they are stage 2 data files.  So each prediction contest has two new files available on its Data page, and the analytics contest (which covers both men’s and women’s) has four new files available on its Data page.\n\nI decided to refresh the entire play-by-play data (including previous years) because I know there were some bug fixes upstream of my work, that would hopefully result in the running score (as the game proceeded) being populated rather than always zero, as well as a significant problem that was fixed in the elapsed seconds of the women’s data.  My understanding is that these fixes were implemented after I pulled the play-by-play data for stage 1, and so they should be fixed in this latest stage 2 release, now that I've pulled the data again.  The play-by-play files from previous years seem a bit larger now, so I expect there are a few more games covered than before - unless it’s just that there are nonzero running scores in there rather than zeroes, taking up a few extra bytes per row.\n\nThe Massey Ordinals (men’s only) are only populated up through Day 128 - usually there is a Day 133 but the season ended early.  I think there may be one more round of ordinals posted on the site but I am not sure how complete that data is.\n\nEverything should be up to date now, although obviously there are no new NCAA Tourney seeds or results or anything like that.  We don’t even get to know which regions were W, X, Y, or Z!  Anyway, these updates should provide useful extra information for anyone still working on the analytics contest, and or people wrapping up their work on the predictions contest.  Please post any issues you find in these data releases, to this thread for your contest, and hopefully I can address them.",
    "792787": "jeffsonas Thanks for update.",
    "795766": "Very obscure issue with the data, but the LCurrentScore and WCurrentScore values in the play by play data are swapped for a game between New Mexico St and Colorado St that happened early this season on the men's side. Strangely, its not for the entire game, just part of overtime. The final few play by play entries have Colorado St winning 78 - 70 but it should be reverse. I'm sure other mistakes like this could be found by just searching for games where LCurrentScore is higher than WCurrentScore for the final entry.\n\nAnother such example is Penn St vs Syracuse this season. WCurrentScore and LCurrentScore are swapped for the final 2 minutes of the game.",
    "796909": "hmtessier: many thanks for taking the time to investigate this issue. @jeffsonas: kindly investigate this one. Play-By-Play data is highly critical in this analytics competitions.",
    "796983": "This is not data that I made any changes to during the import process - it is straight from the source.  It is possible to calculate WCurrentScore and LCurrentScore yourself and so perhaps people may want to do that themselves, to ensure higher quality data if this does seem extensive.  But there isn't anything I can easily do to resolve it, other than calculating it myself from the game events, and that has its own problems.  I don't think there is anything we should do here for this."
  },
  "source": "meta"
}