{
  "id": 402402,
  "title": "Preprocessed data",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402402",
  "author_name": "Viktor Cikojevic",
  "post_date": "2023-04-18T08:27:03.946000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all, even though it's only few days until the competition ends, I share with you the preprocessed dataset, which might be useful for fast event loading: <a href=\"https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023\" target=\"_blank\">https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023</a></p>\n<p>Each row contains features for a given <code>event_id</code>, together with <code>azimuth</code> and <code>zenith</code>. I take first 96 signals, as inspired from this kernel: <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu</a>.<br>\nI've taken random 30% of the total train dataset, ie 198 batches out of 660.</p>\n<p>Coordinates are rescaled as follows </p>\n<pre><code>x = x / 500\ny = y / 500\nz = z / 500\n</code></pre>\n<p>Time is rescaled as <code>t = t / 16000</code>, charge is rescaled as <code>charge = charge / 2</code>, and it is clipped to values between 0 and 2.5. I've added <code>charge_clipped</code> field that denotes whether the charge was clipped. </p>\n<p>I've done the preprocessing in BigQuery, but I also have the PyTorch code to preprocess the data is here: <a href=\"https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode\" target=\"_blank\">https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode</a> (kernel will be available after the competition ends).</p>",
  "messages": [
    {
      "id": 2225535,
      "postDate": "2023-04-18T08:27:03.947Z",
      "content": "<p>Hi all, even though it's only few days until the competition ends, I share with you the preprocessed dataset, which might be useful for fast event loading: <a href=\"https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023\" target=\"_blank\">https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023</a></p>\n<p>Each row contains features for a given <code>event_id</code>, together with <code>azimuth</code> and <code>zenith</code>. I take first 96 signals, as inspired from this kernel: <a href=\"https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu\" target=\"_blank\">https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu</a>.<br>\nI've taken random 30% of the total train dataset, ie 198 batches out of 660.</p>\n<p>Coordinates are rescaled as follows </p>\n<pre><code>x = x / 500\ny = y / 500\nz = z / 500\n</code></pre>\n<p>Time is rescaled as <code>t = t / 16000</code>, charge is rescaled as <code>charge = charge / 2</code>, and it is clipped to values between 0 and 2.5. I've added <code>charge_clipped</code> field that denotes whether the charge was clipped. </p>\n<p>I've done the preprocessing in BigQuery, but I also have the PyTorch code to preprocess the data is here: <a href=\"https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode\" target=\"_blank\">https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode</a> (kernel will be available after the competition ends).</p>",
      "rawMarkdown": "Hi all, even though it's only few days until the competition ends, I share with you the preprocessed dataset, which might be useful for fast event loading: https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023\n\nEach row contains features for a given `event_id`, together with `azimuth` and `zenith`. I take first 96 signals, as inspired from this kernel: https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu.\nI've taken random 30% of the total train dataset, ie 198 batches out of 660.\n\nCoordinates are rescaled as follows \n\n```\nx = x / 500\ny = y / 500\nz = z / 500\n```\n\nTime is rescaled as `t = t / 16000`, charge is rescaled as `charge = charge / 2`, and it is clipped to values between 0 and 2.5. I've added `charge_clipped` field that denotes whether the charge was clipped. \n\nI've done the preprocessing in BigQuery, but I also have the PyTorch code to preprocess the data is here: https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode (kernel will be available after the competition ends).\n\n\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2225535": "Hi all, even though it's only few days until the competition ends, I share with you the preprocessed dataset, which might be useful for fast event loading: https://www.kaggle.com/datasets/viktorcikojevic/preprocessed-icecube-2023\n\nEach row contains features for a given `event_id`, together with `azimuth` and `zenith`. I take first 96 signals, as inspired from this kernel: https://www.kaggle.com/code/rsmits/tensorflow-lstm-model-training-tpu.\nI've taken random 30% of the total train dataset, ie 198 batches out of 660.\n\nCoordinates are rescaled as follows \n\n```\nx = x / 500\ny = y / 500\nz = z / 500\n```\n\nTime is rescaled as `t = t / 16000`, charge is rescaled as `charge = charge / 2`, and it is clipped to values between 0 and 2.5. I've added `charge_clipped` field that denotes whether the charge was clipped. \n\nI've done the preprocessing in BigQuery, but I also have the PyTorch code to preprocess the data is here: https://www.kaggle.com/code/viktorcikojevic/lstm-cnn-transformer-mode (kernel will be available after the competition ends).\n\n\n"
  }
}