{
  "id": 373971,
  "title": "running time",
  "url": "/competitions/nfl-player-contact-detection/discussion/373971",
  "author_name": "",
  "post_date": "2022-12-24T14:24:05.104722400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have a question, </p>\n<p>I'm trying to merge train_player_tracking.csv file with train_labels.csv so I have the position of both players in one row in the new dataframe along with other variable such as speed, direction etc and the target variable (contact) all in one dataframe to train in this data and preprocess it. However, it takes me around 50 hours to create this dataframe because it's a huge data. I'm working on this with my own gpu, I'm wondering if submit later, is this going to give me run time error? </p>",
  "messages": [
    {
      "id": "2074712",
      "postDate": "12/24/2022 14:24:05",
      "content": "<p>I have a question, </p>\n<p>I'm trying to merge train_player_tracking.csv file with train_labels.csv so I have the position of both players in one row in the new dataframe along with other variable such as speed, direction etc and the target variable (contact) all in one dataframe to train in this data and preprocess it. However, it takes me around 50 hours to create this dataframe because it's a huge data. I'm working on this with my own gpu, I'm wondering if submit later, is this going to give me run time error? </p>",
      "rawMarkdown": "I have a question, \n\nI'm trying to merge train_player_tracking.csv file with train_labels.csv so I have the position of both players in one row in the new dataframe along with other variable such as speed, direction etc and the target variable (contact) all in one dataframe to train in this data and preprocess it. However, it takes me around 50 hours to create this dataframe because it's a huge data. I'm working on this with my own gpu, I'm wondering if submit later, is this going to give me run time error?",
      "votes": null
    },
    {
      "id": "2074719",
      "postDate": "12/24/2022 14:39:26",
      "content": "<p>There is no way, something might be wrong on your code. We would need to see the code, but if I had to guess, you are probably merging on the wrong keys and this generates a massive amount of end rows.</p>",
      "rawMarkdown": "There is no way, something might be wrong on your code. We would need to see the code, but if I had to guess, you are probably merging on the wrong keys and this generates a massive amount of end rows.",
      "votes": null
    },
    {
      "id": "2074749",
      "postDate": "12/24/2022 14:59:56",
      "content": "<p>my code basically loop through train_labels.csv row by row, look in the train_player_tracking.csv based on the step and game_play and nfl_player_id1 which returns to me the info about that player and then I do the same thing for nfl_player_id2. Then I merge these two into one row and populate in a dataframe that I created. However, the problem is I need to loop through around 4 million rows to do this and create my new dataframe. </p>",
      "rawMarkdown": "my code basically loop through train_labels.csv row by row, look in the train_player_tracking.csv based on the step and game_play and nfl_player_id1 which returns to me the info about that player and then I do the same thing for nfl_player_id2. Then I merge these two into one row and populate in a dataframe that I created. However, the problem is I need to loop through around 4 million rows to do this and create my new dataframe.",
      "votes": null
    },
    {
      "id": "2075046",
      "postDate": "12/25/2022 01:12:07",
      "content": "<p>You should use joins instead of looping through each row</p>",
      "rawMarkdown": "You should use joins instead of looping through each row",
      "votes": null
    },
    {
      "id": "2080988",
      "postDate": "12/30/2022 17:37:50",
      "content": "<p>Here is a snippet that does what you need using Pandas joins. I suggest getting more familiar with it using this <a href=\"https://www.kaggle.com/learn/pandas\" target=\"_blank\">Kaggle course</a>, as loops aren't effective at doing most operations on dataframes.</p>\n<pre><code>import pandas as pd\n\nfrom pathlib import Path\n\n\n## Source: https://www.kaggle.com/code/robikscube/nfl-player-contact-detection-getting-started\ndef add_contact_id(df):\n\n    # Create contact ids\n    df[\"contact_id\"] = (\n        df[\"game_play\"]\n        + \"_\"\n        + df[\"step\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_1\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_2\"].astype(\"str\")\n    )\n    return df\n\nlabels = pd.read_csv(BASE_DIR/\"train_labels.csv\", parse_dates=[\"datetime\"])\n\ntr_tracking = pd.read_csv(\n    BASE_DIR/\"train_player_tracking.csv\", parse_dates=[\"datetime\"]\n)\n\ndf_combo = (\n        labels.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", 'datetime', \"nfl_player_id\", \"x_position\", \"y_position\"]\n            ],\n            left_on=[\"game_play\", 'datetime', \"nfl_player_id_1\"],\n            right_on=[\"game_play\", 'datetime', \"nfl_player_id\"],\n            how=\"left\",\n        )\n)\n</code></pre>",
      "rawMarkdown": "Here is a snippet that does what you need using Pandas joins. I suggest getting more familiar with it using this [Kaggle course](https://www.kaggle.com/learn/pandas), as loops aren't effective at doing most operations on dataframes.\n\n```\nimport pandas as pd\n\nfrom pathlib import Path\n\n\n## Source: https://www.kaggle.com/code/robikscube/nfl-player-contact-detection-getting-started\ndef add_contact_id(df):\n\n    # Create contact ids\n    df[\"contact_id\"] = (\n        df[\"game_play\"]\n        + \"_\"\n        + df[\"step\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_1\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_2\"].astype(\"str\")\n    )\n    return df\n\nlabels = pd.read_csv(BASE_DIR/\"train_labels.csv\", parse_dates=[\"datetime\"])\n\ntr_tracking = pd.read_csv(\n    BASE_DIR/\"train_player_tracking.csv\", parse_dates=[\"datetime\"]\n)\n\ndf_combo = (\n        labels.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", 'datetime', \"nfl_player_id\", \"x_position\", \"y_position\"]\n            ],\n            left_on=[\"game_play\", 'datetime', \"nfl_player_id_1\"],\n            right_on=[\"game_play\", 'datetime', \"nfl_player_id\"],\n            how=\"left\",\n        )\n)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2074719,
      "author_name": "carloshuertas",
      "author_url": "",
      "post_date": "12/24/2022 14:39:26",
      "content": "<p>There is no way, something might be wrong on your code. We would need to see the code, but if I had to guess, you are probably merging on the wrong keys and this generates a massive amount of end rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2074749,
          "author_name": "naifalkhunaizi3",
          "author_url": "",
          "post_date": "12/24/2022 14:59:56",
          "content": "<p>my code basically loop through train_labels.csv row by row, look in the train_player_tracking.csv based on the step and game_play and nfl_player_id1 which returns to me the info about that player and then I do the same thing for nfl_player_id2. Then I merge these two into one row and populate in a dataframe that I created. However, the problem is I need to loop through around 4 million rows to do this and create my new dataframe. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2075046,
              "author_name": "ryancaldwell",
              "author_url": "",
              "post_date": "12/25/2022 01:12:07",
              "content": "<p>You should use joins instead of looping through each row</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2080988,
              "author_name": "samir95",
              "author_url": "",
              "post_date": "12/30/2022 17:37:50",
              "content": "<p>Here is a snippet that does what you need using Pandas joins. I suggest getting more familiar with it using this <a href=\"https://www.kaggle.com/learn/pandas\" target=\"_blank\">Kaggle course</a>, as loops aren't effective at doing most operations on dataframes.</p>\n<pre><code>import pandas as pd\n\nfrom pathlib import Path\n\n\n## Source: https://www.kaggle.com/code/robikscube/nfl-player-contact-detection-getting-started\ndef add_contact_id(df):\n\n    # Create contact ids\n    df[\"contact_id\"] = (\n        df[\"game_play\"]\n        + \"_\"\n        + df[\"step\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_1\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_2\"].astype(\"str\")\n    )\n    return df\n\nlabels = pd.read_csv(BASE_DIR/\"train_labels.csv\", parse_dates=[\"datetime\"])\n\ntr_tracking = pd.read_csv(\n    BASE_DIR/\"train_player_tracking.csv\", parse_dates=[\"datetime\"]\n)\n\ndf_combo = (\n        labels.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", 'datetime', \"nfl_player_id\", \"x_position\", \"y_position\"]\n            ],\n            left_on=[\"game_play\", 'datetime', \"nfl_player_id_1\"],\n            right_on=[\"game_play\", 'datetime', \"nfl_player_id\"],\n            how=\"left\",\n        )\n)\n</code></pre>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2074712": "I have a question, \n\nI'm trying to merge train_player_tracking.csv file with train_labels.csv so I have the position of both players in one row in the new dataframe along with other variable such as speed, direction etc and the target variable (contact) all in one dataframe to train in this data and preprocess it. However, it takes me around 50 hours to create this dataframe because it's a huge data. I'm working on this with my own gpu, I'm wondering if submit later, is this going to give me run time error?",
    "2074719": "There is no way, something might be wrong on your code. We would need to see the code, but if I had to guess, you are probably merging on the wrong keys and this generates a massive amount of end rows.",
    "2074749": "my code basically loop through train_labels.csv row by row, look in the train_player_tracking.csv based on the step and game_play and nfl_player_id1 which returns to me the info about that player and then I do the same thing for nfl_player_id2. Then I merge these two into one row and populate in a dataframe that I created. However, the problem is I need to loop through around 4 million rows to do this and create my new dataframe.",
    "2075046": "You should use joins instead of looping through each row",
    "2080988": "Here is a snippet that does what you need using Pandas joins. I suggest getting more familiar with it using this [Kaggle course](https://www.kaggle.com/learn/pandas), as loops aren't effective at doing most operations on dataframes.\n\n```\nimport pandas as pd\n\nfrom pathlib import Path\n\n\n## Source: https://www.kaggle.com/code/robikscube/nfl-player-contact-detection-getting-started\ndef add_contact_id(df):\n\n    # Create contact ids\n    df[\"contact_id\"] = (\n        df[\"game_play\"]\n        + \"_\"\n        + df[\"step\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_1\"].astype(\"str\")\n        + \"_\"\n        + df[\"nfl_player_id_2\"].astype(\"str\")\n    )\n    return df\n\nlabels = pd.read_csv(BASE_DIR/\"train_labels.csv\", parse_dates=[\"datetime\"])\n\ntr_tracking = pd.read_csv(\n    BASE_DIR/\"train_player_tracking.csv\", parse_dates=[\"datetime\"]\n)\n\ndf_combo = (\n        labels.astype({\"nfl_player_id_1\": \"str\"})\n        .merge(\n            tr_tracking.astype({\"nfl_player_id\": \"str\"})[\n                [\"game_play\", 'datetime', \"nfl_player_id\", \"x_position\", \"y_position\"]\n            ],\n            left_on=[\"game_play\", 'datetime', \"nfl_player_id_1\"],\n            right_on=[\"game_play\", 'datetime', \"nfl_player_id\"],\n            how=\"left\",\n        )\n)\n```"
  },
  "source": "meta"
}