{
  "id": 381126,
  "title": "scoring with hidden test set. Should I use event_id in sample_submission.parquet or test_meta.parquet",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/381126",
  "author_name": "",
  "post_date": "2023-01-25T11:33:35.960661Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Well, I know in this competition, when I submit notebook, sample test set(test_meta.parquet) and will be replaced by actual test set. But reading this parquet file may be very big, so I think reading event_id from sample_submission.parquet will be better (it has only 3 columns). </p>\n<p>But sample_submission.parquet will be replace with actual <code>submission.parquet</code> file with all event_id in actual test set ? <br>\nIf <code>sample_submission.parquet</code> doesn't get replace, so if I use this file as base event_ids, my final result will miss all event_ids in actual test set. </p>\n<p>I doubt it, so if anyone know about this mechanics, please share with me</p>",
  "messages": [
    {
      "id": "2114974",
      "postDate": "01/25/2023 11:33:35",
      "content": "<p>Well, I know in this competition, when I submit notebook, sample test set(test_meta.parquet) and will be replaced by actual test set. But reading this parquet file may be very big, so I think reading event_id from sample_submission.parquet will be better (it has only 3 columns). </p>\n<p>But sample_submission.parquet will be replace with actual <code>submission.parquet</code> file with all event_id in actual test set ? <br>\nIf <code>sample_submission.parquet</code> doesn't get replace, so if I use this file as base event_ids, my final result will miss all event_ids in actual test set. </p>\n<p>I doubt it, so if anyone know about this mechanics, please share with me</p>",
      "rawMarkdown": "Well, I know in this competition, when I submit notebook, sample test set(test_meta.parquet) and will be replaced by actual test set. But reading this parquet file may be very big, so I think reading event_id from sample_submission.parquet will be better (it has only 3 columns). \n\nBut sample_submission.parquet will be replace with actual `submission.parquet` file with all event_id in actual test set ? \nIf `sample_submission.parquet` doesn't get replace, so if I use this file as base event_ids, my final result will miss all event_ids in actual test set. \n\nI doubt it, so if anyone know about this mechanics, please share with me",
      "votes": null
    },
    {
      "id": "2116336",
      "postDate": "01/26/2023 12:30:53",
      "content": "<p>Quoting from competition data page <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data</a></p>\n<blockquote>\n  <p>When your submitted notebook is scored the actual test data (including a full length <strong>sample submission</strong>) will be made available to your notebook.</p>\n</blockquote>\n<p>I haven't tried this myself though.</p>\n<p>Note that <code>pandas.read_parquet()</code> method is accepting <code>columns</code> parameter, which can be used to define which columns (names) that should be read, other columns won't be read, hence it can be used to save some memory.</p>",
      "rawMarkdown": "Quoting from competition data page https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\n\n>When your submitted notebook is scored the actual test data (including a full length **sample submission**) will be made available to your notebook.\n\nI haven't tried this myself though.\n\nNote that `pandas.read_parquet()` method is accepting `columns` parameter, which can be used to define which columns (names) that should be read, other columns won't be read, hence it can be used to save some memory.",
      "votes": null
    },
    {
      "id": "2119464",
      "postDate": "01/28/2023 19:27:37",
      "content": "<p>You can use event_id from any source out of sample_submission.parquet, test_meta.parquet, or even the option to get event_ids directly from all the batch_N.parquet files in test/<em>.</em></p>\n<p>In all three cases, the hidden files used for scoring will have all the full hidden test data, all the event_ids.</p>\n<p>The key thing is to make sure you end up writing a submission.csv that has the three columns event_id,azimuth,zenith for every event_id in the test data.</p>",
      "rawMarkdown": "You can use event_id from any source out of sample_submission.parquet, test_meta.parquet, or even the option to get event_ids directly from all the batch_N.parquet files in test/*.*\n\nIn all three cases, the hidden files used for scoring will have all the full hidden test data, all the event_ids.\n\nThe key thing is to make sure you end up writing a submission.csv that has the three columns event_id,azimuth,zenith for every event_id in the test data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2116336,
      "author_name": "thariqnugrohotomo",
      "author_url": "",
      "post_date": "01/26/2023 12:30:53",
      "content": "<p>Quoting from competition data page <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\" target=\"_blank\">https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data</a></p>\n<blockquote>\n  <p>When your submitted notebook is scored the actual test data (including a full length <strong>sample submission</strong>) will be made available to your notebook.</p>\n</blockquote>\n<p>I haven't tried this myself though.</p>\n<p>Note that <code>pandas.read_parquet()</code> method is accepting <code>columns</code> parameter, which can be used to define which columns (names) that should be read, other columns won't be read, hence it can be used to save some memory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2119464,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "01/28/2023 19:27:37",
      "content": "<p>You can use event_id from any source out of sample_submission.parquet, test_meta.parquet, or even the option to get event_ids directly from all the batch_N.parquet files in test/<em>.</em></p>\n<p>In all three cases, the hidden files used for scoring will have all the full hidden test data, all the event_ids.</p>\n<p>The key thing is to make sure you end up writing a submission.csv that has the three columns event_id,azimuth,zenith for every event_id in the test data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2114974": "Well, I know in this competition, when I submit notebook, sample test set(test_meta.parquet) and will be replaced by actual test set. But reading this parquet file may be very big, so I think reading event_id from sample_submission.parquet will be better (it has only 3 columns). \n\nBut sample_submission.parquet will be replace with actual `submission.parquet` file with all event_id in actual test set ? \nIf `sample_submission.parquet` doesn't get replace, so if I use this file as base event_ids, my final result will miss all event_ids in actual test set. \n\nI doubt it, so if anyone know about this mechanics, please share with me",
    "2116336": "Quoting from competition data page https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/data\n\n>When your submitted notebook is scored the actual test data (including a full length **sample submission**) will be made available to your notebook.\n\nI haven't tried this myself though.\n\nNote that `pandas.read_parquet()` method is accepting `columns` parameter, which can be used to define which columns (names) that should be read, other columns won't be read, hence it can be used to save some memory.",
    "2119464": "You can use event_id from any source out of sample_submission.parquet, test_meta.parquet, or even the option to get event_ids directly from all the batch_N.parquet files in test/*.*\n\nIn all three cases, the hidden files used for scoring will have all the full hidden test data, all the event_ids.\n\nThe key thing is to make sure you end up writing a submission.csv that has the three columns event_id,azimuth,zenith for every event_id in the test data."
  },
  "source": "meta"
}