{
  "id": 279278,
  "title": "Processing the Tracking data",
  "url": "/competitions/nfl-big-data-bowl-2022/discussion/279278",
  "author_name": "",
  "post_date": "2021-10-17T13:47:53.902854500Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I'm having some trouble processing the tracking data. I'd like to bring it all into a DataFrame but only look at one specific type of play, like punts or kickoffs, and keep running into memory usage issues.</p>\n<p>How I'd think of handling it would be to bring in the plays data and create a primary key by combining the playId and gameId columns (since the playId isn't unique across games - perhaps I could do this with a multi-index?), and then filtering out all of the non-punts (or whichever).</p>\n<p>Then bringing in all of the tracking data and making the same key, inner-joining on the two, so that I have a combined dataset with plays and tracking data for only punts. However this kills the memory.</p>\n<p>I'm looking for a way to bring in the tracking data but only for those rows of data that are the specific play type automatically. For example, as pd.read_csv is reading the data, immediately dropping those rows that don't pertain to punts by creating and checking the unique play key (or perhaps by chunking). Would that work and how would I implement this? </p>\n<p>Is there a more elegant way to do this? I took a look at the <a href=\"https://www.kaggle.com/dhritiyandapally/analyzing-place-kicker-offset-from-center-python\" target=\"_blank\">demo</a> but the filtering of the tracking data there doesn't seem as complex as filtering for a playKey.</p>\n<p>Decently experienced with python but my formal computer science knowledge is definitely lacking. I'd appreciate kicks in the right direction. Thanks!</p>",
  "messages": [
    {
      "id": "1547692",
      "postDate": "10/17/2021 13:47:53",
      "content": "<p>Hello,</p>\n<p>I'm having some trouble processing the tracking data. I'd like to bring it all into a DataFrame but only look at one specific type of play, like punts or kickoffs, and keep running into memory usage issues.</p>\n<p>How I'd think of handling it would be to bring in the plays data and create a primary key by combining the playId and gameId columns (since the playId isn't unique across games - perhaps I could do this with a multi-index?), and then filtering out all of the non-punts (or whichever).</p>\n<p>Then bringing in all of the tracking data and making the same key, inner-joining on the two, so that I have a combined dataset with plays and tracking data for only punts. However this kills the memory.</p>\n<p>I'm looking for a way to bring in the tracking data but only for those rows of data that are the specific play type automatically. For example, as pd.read_csv is reading the data, immediately dropping those rows that don't pertain to punts by creating and checking the unique play key (or perhaps by chunking). Would that work and how would I implement this? </p>\n<p>Is there a more elegant way to do this? I took a look at the <a href=\"https://www.kaggle.com/dhritiyandapally/analyzing-place-kicker-offset-from-center-python\" target=\"_blank\">demo</a> but the filtering of the tracking data there doesn't seem as complex as filtering for a playKey.</p>\n<p>Decently experienced with python but my formal computer science knowledge is definitely lacking. I'd appreciate kicks in the right direction. Thanks!</p>",
      "rawMarkdown": "Hello,\n\nI'm having some trouble processing the tracking data. I'd like to bring it all into a DataFrame but only look at one specific type of play, like punts or kickoffs, and keep running into memory usage issues.\n\nHow I'd think of handling it would be to bring in the plays data and create a primary key by combining the playId and gameId columns (since the playId isn't unique across games - perhaps I could do this with a multi-index?), and then filtering out all of the non-punts (or whichever).\n\nThen bringing in all of the tracking data and making the same key, inner-joining on the two, so that I have a combined dataset with plays and tracking data for only punts. However this kills the memory.\n\nI'm looking for a way to bring in the tracking data but only for those rows of data that are the specific play type automatically. For example, as pd.read_csv is reading the data, immediately dropping those rows that don't pertain to punts by creating and checking the unique play key (or perhaps by chunking). Would that work and how would I implement this? \n\nIs there a more elegant way to do this? I took a look at the [demo](https://www.kaggle.com/dhritiyandapally/analyzing-place-kicker-offset-from-center-python) but the filtering of the tracking data there doesn't seem as complex as filtering for a playKey.\n\nDecently experienced with python but my formal computer science knowledge is definitely lacking. I'd appreciate kicks in the right direction. Thanks!",
      "votes": null
    },
    {
      "id": "1548021",
      "postDate": "10/17/2021 20:44:40",
      "content": "<p>Perhaps this stackoverflow could help:</p>\n<p><a href=\"https://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function\" target=\"_blank\">https://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function</a></p>",
      "rawMarkdown": "Perhaps this stackoverflow could help:\n\nhttps://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function",
      "votes": null
    },
    {
      "id": "1553781",
      "postDate": "10/22/2021 14:11:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danmel\" target=\"_blank\">@danmel</a> - for the tracking data, if you're just wanting the rows where a special teams play occurred, you could write a function that loads the tracking data as a dataframe, reduces/filters it to only those rows where an event occurred (i.e. tracking_df.loc[~tracking_df['event'].isnull()], and have the function return this reduced dataframe</p>",
      "rawMarkdown": "Hi @danmel - for the tracking data, if you're just wanting the rows where a special teams play occurred, you could write a function that loads the tracking data as a dataframe, reduces/filters it to only those rows where an event occurred (i.e. tracking_df.loc[~tracking_df['event'].isnull()], and have the function return this reduced dataframe",
      "votes": null
    },
    {
      "id": "1563846",
      "postDate": "10/28/2021 16:24:56",
      "content": "<p>Super cool example here for chunking with filter</p>\n<p>import pandas as pd<br>\niter_csv = pd.read_csv('file.csv', iterator=True, chunksize=1000)<br>\ndf = pd.concat([chunk[chunk['field'] &gt; constant] for chunk in iter_csv])</p>",
      "rawMarkdown": "Super cool example here for chunking with filter\n\nimport pandas as pd\niter_csv = pd.read_csv('file.csv', iterator=True, chunksize=1000)\ndf = pd.concat([chunk[chunk['field'] > constant] for chunk in iter_csv])",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1548021,
      "author_name": "tombliss",
      "author_url": "",
      "post_date": "10/17/2021 20:44:40",
      "content": "<p>Perhaps this stackoverflow could help:</p>\n<p><a href=\"https://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function\" target=\"_blank\">https://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1563846,
          "author_name": "sfurlong",
          "author_url": "",
          "post_date": "10/28/2021 16:24:56",
          "content": "<p>Super cool example here for chunking with filter</p>\n<p>import pandas as pd<br>\niter_csv = pd.read_csv('file.csv', iterator=True, chunksize=1000)<br>\ndf = pd.concat([chunk[chunk['field'] &gt; constant] for chunk in iter_csv])</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1553781,
      "author_name": "valbauman",
      "author_url": "",
      "post_date": "10/22/2021 14:11:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/danmel\" target=\"_blank\">@danmel</a> - for the tracking data, if you're just wanting the rows where a special teams play occurred, you could write a function that loads the tracking data as a dataframe, reduces/filters it to only those rows where an event occurred (i.e. tracking_df.loc[~tracking_df['event'].isnull()], and have the function return this reduced dataframe</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1547692": "Hello,\n\nI'm having some trouble processing the tracking data. I'd like to bring it all into a DataFrame but only look at one specific type of play, like punts or kickoffs, and keep running into memory usage issues.\n\nHow I'd think of handling it would be to bring in the plays data and create a primary key by combining the playId and gameId columns (since the playId isn't unique across games - perhaps I could do this with a multi-index?), and then filtering out all of the non-punts (or whichever).\n\nThen bringing in all of the tracking data and making the same key, inner-joining on the two, so that I have a combined dataset with plays and tracking data for only punts. However this kills the memory.\n\nI'm looking for a way to bring in the tracking data but only for those rows of data that are the specific play type automatically. For example, as pd.read_csv is reading the data, immediately dropping those rows that don't pertain to punts by creating and checking the unique play key (or perhaps by chunking). Would that work and how would I implement this? \n\nIs there a more elegant way to do this? I took a look at the [demo](https://www.kaggle.com/dhritiyandapally/analyzing-place-kicker-offset-from-center-python) but the filtering of the tracking data there doesn't seem as complex as filtering for a playKey.\n\nDecently experienced with python but my formal computer science knowledge is definitely lacking. I'd appreciate kicks in the right direction. Thanks!",
    "1548021": "Perhaps this stackoverflow could help:\n\nhttps://stackoverflow.com/questions/13651117/how-can-i-filter-lines-on-load-in-pandas-read-csv-function",
    "1553781": "Hi @danmel - for the tracking data, if you're just wanting the rows where a special teams play occurred, you could write a function that loads the tracking data as a dataframe, reduces/filters it to only those rows where an event occurred (i.e. tracking_df.loc[~tracking_df['event'].isnull()], and have the function return this reduced dataframe",
    "1563846": "Super cool example here for chunking with filter\n\nimport pandas as pd\niter_csv = pd.read_csv('file.csv', iterator=True, chunksize=1000)\ndf = pd.concat([chunk[chunk['field'] > constant] for chunk in iter_csv])"
  },
  "source": "meta"
}