{
  "id": 382248,
  "title": "Question about reconstructing the data files.",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/382248",
  "author_name": "",
  "post_date": "2023-01-30T09:00:26.508805900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all! I'm interested in this competition but new (very new) to data science and AI. </p>\n<p>I have an idea that doesn't know whether is suitable or not. I want to ignore the event_id and regard one whole batch as a \"time series\" of data (since there is indeed \"time\" in the datasets), and use LSTM to recognize the azimuth and zenith angle in each period of time. In other words, I would like to design a network for real-time recognization of neutrino events rather than fitting data statically.</p>\n<p>But I met up with a problem. I want to reconstruct the data for LSTM, which should include:<br>\n<code>time | x | y | z | charge</code></p>\n<p>and the output should have the structure(also predict whether the datapoint is auxiliary or not):<br>\n<code>auxiliary | azimuth | zenith</code></p>\n<p>So I need to put time, charge, x, y, z, auxiliary, azimuth, and zenith inside one dataset according to sensor_id and event_id and then split it into X and y. I used <code>groupby</code> and <code>apply</code> like this:</p>\n<pre><code>_df = train_batch.groupby('event_id').apply(lambda x : pd.merge(x, train_meta, on=\"event_id\"))\n_df.reset_index(drop=True, inplace=True)\ndf = _df.groupby('sensor_id').apply(lambda x : pd.merge(x, sensor_geometry, on=\"sensor_id\"))\ndf.reset_index(drop=True, inplace=True)\n</code></pre>\n<p>But this is incredibly slow in the <code>_df</code> part. I was wondering whether there is any simple and fast way to reconstruct the files.</p>\n<p>Thank you very much!</p>",
  "messages": [
    {
      "id": "2121370",
      "postDate": "01/30/2023 09:00:26",
      "content": "<p>Hi all! I'm interested in this competition but new (very new) to data science and AI. </p>\n<p>I have an idea that doesn't know whether is suitable or not. I want to ignore the event_id and regard one whole batch as a \"time series\" of data (since there is indeed \"time\" in the datasets), and use LSTM to recognize the azimuth and zenith angle in each period of time. In other words, I would like to design a network for real-time recognization of neutrino events rather than fitting data statically.</p>\n<p>But I met up with a problem. I want to reconstruct the data for LSTM, which should include:<br>\n<code>time | x | y | z | charge</code></p>\n<p>and the output should have the structure(also predict whether the datapoint is auxiliary or not):<br>\n<code>auxiliary | azimuth | zenith</code></p>\n<p>So I need to put time, charge, x, y, z, auxiliary, azimuth, and zenith inside one dataset according to sensor_id and event_id and then split it into X and y. I used <code>groupby</code> and <code>apply</code> like this:</p>\n<pre><code>_df = train_batch.groupby('event_id').apply(lambda x : pd.merge(x, train_meta, on=\"event_id\"))\n_df.reset_index(drop=True, inplace=True)\ndf = _df.groupby('sensor_id').apply(lambda x : pd.merge(x, sensor_geometry, on=\"sensor_id\"))\ndf.reset_index(drop=True, inplace=True)\n</code></pre>\n<p>But this is incredibly slow in the <code>_df</code> part. I was wondering whether there is any simple and fast way to reconstruct the files.</p>\n<p>Thank you very much!</p>",
      "rawMarkdown": "Hi all! I'm interested in this competition but new (very new) to data science and AI. \n\nI have an idea that doesn't know whether is suitable or not. I want to ignore the event_id and regard one whole batch as a \"time series\" of data (since there is indeed \"time\" in the datasets), and use LSTM to recognize the azimuth and zenith angle in each period of time. In other words, I would like to design a network for real-time recognization of neutrino events rather than fitting data statically.\n\nBut I met up with a problem. I want to reconstruct the data for LSTM, which should include:\n`time | x | y | z | charge`\n\nand the output should have the structure(also predict whether the datapoint is auxiliary or not):\n`auxiliary | azimuth | zenith`\n\nSo I need to put time, charge, x, y, z, auxiliary, azimuth, and zenith inside one dataset according to sensor_id and event_id and then split it into X and y. I used `groupby` and `apply` like this:\n``` \n_df = train_batch.groupby('event_id').apply(lambda x : pd.merge(x, train_meta, on=\"event_id\"))\n_df.reset_index(drop=True, inplace=True)\ndf = _df.groupby('sensor_id').apply(lambda x : pd.merge(x, sensor_geometry, on=\"sensor_id\"))\ndf.reset_index(drop=True, inplace=True)\n```\nBut this is incredibly slow in the `_df` part. I was wondering whether there is any simple and fast way to reconstruct the files.\n\nThank you very much!",
      "votes": null
    },
    {
      "id": "2121787",
      "postDate": "01/30/2023 14:50:05",
      "content": "<p>You need to 'left' merge the train labels and sensor data without calling groupby. </p>\n<p>This is how I do it in <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">my notebook</a>:</p>\n<pre><code>    \n    batch = batch.reset_index().merge(sensor, how=, on=, left_index=).set_index()\n</code></pre>",
      "rawMarkdown": "You need to 'left' merge the train labels and sensor data without calling groupby. \n\nThis is how I do it in [my notebook](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars):\n```python\n    ## Merge in sensor x,y,z data\n    batch = batch.reset_index().merge(sensor, how='left', on='sensor_id', left_index=False).set_index('event_id')\n```",
      "votes": null
    },
    {
      "id": "2122413",
      "postDate": "01/30/2023 23:14:45",
      "content": "<p>It works! Only took about minutes to finish the merge. Thank you very much!</p>",
      "rawMarkdown": "It works! Only took about minutes to finish the merge. Thank you very much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2121787,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "01/30/2023 14:50:05",
      "content": "<p>You need to 'left' merge the train labels and sensor data without calling groupby. </p>\n<p>This is how I do it in <a href=\"https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars\" target=\"_blank\">my notebook</a>:</p>\n<pre><code>    \n    batch = batch.reset_index().merge(sensor, how=, on=, left_index=).set_index()\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2122413,
          "author_name": "thenotfish",
          "author_url": "",
          "post_date": "01/30/2023 23:14:45",
          "content": "<p>It works! Only took about minutes to finish the merge. Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2121370": "Hi all! I'm interested in this competition but new (very new) to data science and AI. \n\nI have an idea that doesn't know whether is suitable or not. I want to ignore the event_id and regard one whole batch as a \"time series\" of data (since there is indeed \"time\" in the datasets), and use LSTM to recognize the azimuth and zenith angle in each period of time. In other words, I would like to design a network for real-time recognization of neutrino events rather than fitting data statically.\n\nBut I met up with a problem. I want to reconstruct the data for LSTM, which should include:\n`time | x | y | z | charge`\n\nand the output should have the structure(also predict whether the datapoint is auxiliary or not):\n`auxiliary | azimuth | zenith`\n\nSo I need to put time, charge, x, y, z, auxiliary, azimuth, and zenith inside one dataset according to sensor_id and event_id and then split it into X and y. I used `groupby` and `apply` like this:\n``` \n_df = train_batch.groupby('event_id').apply(lambda x : pd.merge(x, train_meta, on=\"event_id\"))\n_df.reset_index(drop=True, inplace=True)\ndf = _df.groupby('sensor_id').apply(lambda x : pd.merge(x, sensor_geometry, on=\"sensor_id\"))\ndf.reset_index(drop=True, inplace=True)\n```\nBut this is incredibly slow in the `_df` part. I was wondering whether there is any simple and fast way to reconstruct the files.\n\nThank you very much!",
    "2121787": "You need to 'left' merge the train labels and sensor data without calling groupby. \n\nThis is how I do it in [my notebook](https://www.kaggle.com/code/roberthatch/lb-1-183-lightning-fast-baseline-with-polars):\n```python\n    ## Merge in sensor x,y,z data\n    batch = batch.reset_index().merge(sensor, how='left', on='sensor_id', left_index=False).set_index('event_id')\n```",
    "2122413": "It works! Only took about minutes to finish the merge. Thank you very much!"
  },
  "source": "meta"
}