{
  "id": 121395,
  "title": "When the PlayerTrackData is grouped by PlayKey and sliced by max time to find the end of the play, doing a value_counts of the PlayKeys shows that there are duplicates?",
  "url": "/competitions/nfl-playing-surface-analytics/discussion/121395",
  "author_name": "",
  "post_date": "2019-12-13T01:14:19.428010200Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.</p>",
  "messages": [
    {
      "id": "693944",
      "postDate": "12/13/2019 01:14:19",
      "content": "<p>Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.</p>",
      "rawMarkdown": "Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.",
      "votes": null
    },
    {
      "id": "694775",
      "postDate": "12/14/2019 04:00:26",
      "content": "<p>For each play, player variables are measured 10 times per second for the duration of the play.</p>",
      "rawMarkdown": "For each play, player variables are measured 10 times per second for the duration of the play.",
      "votes": null
    },
    {
      "id": "694781",
      "postDate": "12/14/2019 04:11:53",
      "content": "<p>Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.</p>",
      "rawMarkdown": "Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.",
      "votes": null
    },
    {
      "id": "695232",
      "postDate": "12/14/2019 20:08:16",
      "content": "<p>Would you like to show the code you use? If you group by PlayKey there shouldn't be duplicate PlayKeys. When I run the following code I get a return value of 0.</p>\n\n<pre><code>import pandas as pd\npd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv', \n    usecols=['PlayKey', 'time']).duplicated().sum()\n</code></pre>",
      "rawMarkdown": "Would you like to show the code you use? If you group by PlayKey there shouldn't be duplicate PlayKeys. When I run the following code I get a return value of 0.\n\n    import pandas as pd\n    pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv', \n        usecols=['PlayKey', 'time']).duplicated().sum()",
      "votes": null
    },
    {
      "id": "695246",
      "postDate": "12/14/2019 20:44:15",
      "content": "<p>players=pd.read_pickle('/content/drive/My Drive/Colab Notebooks/NFLInjuries/PlayerTrackData.pkl')</p>\n\n<p>idx=players.groupby(['PlayKey'])['time'].transform(max) == players['time'] </p>\n\n<p>players[idx]['PlayKey'].value_counts()</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F246949869deb86d6de04fa4ee28080c1%2FScreen%20Shot%202019-12-14%20at%2012.42.58%20PM.png?generation=1576356237393509&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "players=pd.read_pickle('/content/drive/My Drive/Colab Notebooks/NFLInjuries/PlayerTrackData.pkl')\n\nidx=players.groupby(['PlayKey'])['time'].transform(max) == players['time'] \n\nplayers[idx]['PlayKey'].value_counts()\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F246949869deb86d6de04fa4ee28080c1%2FScreen%20Shot%202019-12-14%20at%2012.42.58%20PM.png?generation=1576356237393509&amp;alt=media)",
      "votes": null
    },
    {
      "id": "695301",
      "postDate": "12/14/2019 23:08:50",
      "content": "<p>I suspect data in the pkl file has changed somehow. There are 266,960 unique PlayKeys in the dataset.</p>",
      "rawMarkdown": "I suspect data in the pkl file has changed somehow. There are 266,960 unique PlayKeys in the dataset.",
      "votes": null
    },
    {
      "id": "695322",
      "postDate": "12/15/2019 00:23:41",
      "content": "<p>Proven, thanks for the tip.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F7b0cf237abcefe5d11dccc854aa8c9a3%2FScreen%20Shot%202019-12-14%20at%204.22.43%20PM.png?generation=1576369402931270&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Proven, thanks for the tip.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F7b0cf237abcefe5d11dccc854aa8c9a3%2FScreen%20Shot%202019-12-14%20at%204.22.43%20PM.png?generation=1576369402931270&amp;alt=media)",
      "votes": null
    },
    {
      "id": "695347",
      "postDate": "12/15/2019 02:08:49",
      "content": "<p>Been there - it's the kind of thing that can drive you crazy for days!</p>",
      "rawMarkdown": "Been there - it's the kind of thing that can drive you crazy for days!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694775,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "12/14/2019 04:00:26",
      "content": "<p>For each play, player variables are measured 10 times per second for the duration of the play.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 694781,
      "author_name": "erictwabe",
      "author_url": "",
      "post_date": "12/14/2019 04:11:53",
      "content": "<p>Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.</p>",
      "votes": null,
      "replies": [
        {
          "id": 695232,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "12/14/2019 20:08:16",
          "content": "<p>Would you like to show the code you use? If you group by PlayKey there shouldn't be duplicate PlayKeys. When I run the following code I get a return value of 0.</p>\n\n<pre><code>import pandas as pd\npd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv', \n    usecols=['PlayKey', 'time']).duplicated().sum()\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 695246,
      "author_name": "erictwabe",
      "author_url": "",
      "post_date": "12/14/2019 20:44:15",
      "content": "<p>players=pd.read_pickle('/content/drive/My Drive/Colab Notebooks/NFLInjuries/PlayerTrackData.pkl')</p>\n\n<p>idx=players.groupby(['PlayKey'])['time'].transform(max) == players['time'] </p>\n\n<p>players[idx]['PlayKey'].value_counts()</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F246949869deb86d6de04fa4ee28080c1%2FScreen%20Shot%202019-12-14%20at%2012.42.58%20PM.png?generation=1576356237393509&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 695301,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "12/14/2019 23:08:50",
          "content": "<p>I suspect data in the pkl file has changed somehow. There are 266,960 unique PlayKeys in the dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695322,
          "author_name": "erictwabe",
          "author_url": "",
          "post_date": "12/15/2019 00:23:41",
          "content": "<p>Proven, thanks for the tip.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F7b0cf237abcefe5d11dccc854aa8c9a3%2FScreen%20Shot%202019-12-14%20at%204.22.43%20PM.png?generation=1576369402931270&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695347,
          "author_name": "jpmiller",
          "author_url": "",
          "post_date": "12/15/2019 02:08:49",
          "content": "<p>Been there - it's the kind of thing that can drive you crazy for days!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "693944": "Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.",
    "694775": "For each play, player variables are measured 10 times per second for the duration of the play.",
    "694781": "Let me rephrase. If placed in a dataframe, and grouping by playkeys and slicing by time transform for max(we are looking for the last row with max time when the play terminates), and then we find the value_counts() for the playkeys, there are duplicate playkeys. Across 7 million rows, there are playkeys that are repeated twice meaning there are plays that occur twice.",
    "695232": "Would you like to show the code you use? If you group by PlayKey there shouldn't be duplicate PlayKeys. When I run the following code I get a return value of 0.\n\n    import pandas as pd\n    pd.read_csv('../input/nfl-playing-surface-analytics/PlayerTrackData.csv', \n        usecols=['PlayKey', 'time']).duplicated().sum()",
    "695246": "players=pd.read_pickle('/content/drive/My Drive/Colab Notebooks/NFLInjuries/PlayerTrackData.pkl')\n\nidx=players.groupby(['PlayKey'])['time'].transform(max) == players['time'] \n\nplayers[idx]['PlayKey'].value_counts()\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F246949869deb86d6de04fa4ee28080c1%2FScreen%20Shot%202019-12-14%20at%2012.42.58%20PM.png?generation=1576356237393509&amp;alt=media)",
    "695301": "I suspect data in the pkl file has changed somehow. There are 266,960 unique PlayKeys in the dataset.",
    "695322": "Proven, thanks for the tip.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2573206%2F7b0cf237abcefe5d11dccc854aa8c9a3%2FScreen%20Shot%202019-12-14%20at%204.22.43%20PM.png?generation=1576369402931270&amp;alt=media)",
    "695347": "Been there - it's the kind of thing that can drive you crazy for days!"
  },
  "source": "meta"
}