{
  "id": 253079,
  "title": "Competition Relaunch - New Data!",
  "url": "/competitions/seti-breakthrough-listen/discussion/253079",
  "author_name": "inversion",
  "post_date": "2021-07-14T22:40:50.040000",
  "votes": 91,
  "comment_count": 38,
  "views": 0,
  "content": "<p>First of all, thanks for everyone's patience, and many thanks to those who highlighted the issues with the previous dataset.</p>\n<p>Immediately after posting this, I'll be swapping in the new dataset to the Data Page. The updated dataset has a folder <code>old_leaky_data</code> which contains the previous competition data as well as the corresponding test labels. You should not assume the old data is comparable to the new dataset. (It might help, it might not.)</p>\n<p>While we've taken efforts to remove the leakage artifacts that were present in the previous dataset, it's difficult to know whether or not the new dataset has some undetected surprises. Remember, we don't actually know what a \"true\" ET signal is supposed to look like. This competition is about building anomaly detection methods that the Breakthrough Listen team and others can use to efficiently and effectively process the vast amounts of data they receive every year. This competition is a way for Kagglers to help contribute to that cause.</p>\n<p>With the new dataset, the competition leaderboard will be reset. This typically takes a few minutes, and you might see some strange LB behavior during the process. </p>\n<p>Finally, the competition deadline has been extended to Aug 18.</p>\n<p>Again, apologies for the inconveniences!</p>",
  "messages": [
    {
      "id": 1388398,
      "postDate": "2021-07-14T22:40:50.040Z",
      "content": "<p>First of all, thanks for everyone's patience, and many thanks to those who highlighted the issues with the previous dataset.</p>\n<p>Immediately after posting this, I'll be swapping in the new dataset to the Data Page. The updated dataset has a folder <code>old_leaky_data</code> which contains the previous competition data as well as the corresponding test labels. You should not assume the old data is comparable to the new dataset. (It might help, it might not.)</p>\n<p>While we've taken efforts to remove the leakage artifacts that were present in the previous dataset, it's difficult to know whether or not the new dataset has some undetected surprises. Remember, we don't actually know what a \"true\" ET signal is supposed to look like. This competition is about building anomaly detection methods that the Breakthrough Listen team and others can use to efficiently and effectively process the vast amounts of data they receive every year. This competition is a way for Kagglers to help contribute to that cause.</p>\n<p>With the new dataset, the competition leaderboard will be reset. This typically takes a few minutes, and you might see some strange LB behavior during the process. </p>\n<p>Finally, the competition deadline has been extended to Aug 18.</p>\n<p>Again, apologies for the inconveniences!</p>",
      "rawMarkdown": "First of all, thanks for everyone's patience, and many thanks to those who highlighted the issues with the previous dataset.\n\nImmediately after posting this, I'll be swapping in the new dataset to the Data Page. The updated dataset has a folder `old_leaky_data` which contains the previous competition data as well as the corresponding test labels. You should not assume the old data is comparable to the new dataset. (It might help, it might not.)\n\nWhile we've taken efforts to remove the leakage artifacts that were present in the previous dataset, it's difficult to know whether or not the new dataset has some undetected surprises. Remember, we don't actually know what a \"true\" ET signal is supposed to look like. This competition is about building anomaly detection methods that the Breakthrough Listen team and others can use to efficiently and effectively process the vast amounts of data they receive every year. This competition is a way for Kagglers to help contribute to that cause.\n\nWith the new dataset, the competition leaderboard will be reset. This typically takes a few minutes, and you might see some strange LB behavior during the process. \n\nFinally, the competition deadline has been extended to Aug 18.\n\nAgain, apologies for the inconveniences!",
      "votes": 91
    },
    {
      "id": 1389223,
      "postDate": "2021-07-15T14:30:25.157Z",
      "content": "<p>Funny side effect: finding best public score is a challenge now with all notebooks with score 1.0 from before the relaunch.  LOL.</p>",
      "rawMarkdown": "Funny side effect: finding best public score is a challenge now with all notebooks with score 1.0 from before the relaunch.  LOL.",
      "votes": 9
    },
    {
      "id": 1388724,
      "postDate": "2021-07-15T07:06:36.723Z",
      "content": "<p>why  deadline has been extended to Aug 18 only? i think when this competition was launched first it had deadline of around 60+ days, Aug 18 is very aggressive deadline,,with new data we expected 60+ days deadline like it was before with old data,,,,considering other good active kaggle competitions like covid19,commonlit this new deadline is aggressive  :(</p>",
      "rawMarkdown": "why  deadline has been extended to Aug 18 only? i think when this competition was launched first it had deadline of around 60+ days, Aug 18 is very aggressive deadline,,with new data we expected 60+ days deadline like it was before with old data,,,,considering other good active kaggle competitions like covid19,commonlit this new deadline is aggressive  :(",
      "votes": 9,
      "replies": [
        {
          "id": 1388829,
          "postDate": "2021-07-15T08:50:58.650Z",
          "content": "<p>Why would you extend the competition to the same full time span? That does not make any sense. The leak was found ~4 weeks ago, and now it was extended ~3 weeks. I think this is more than enough, specifically given that you could continue working on the problem the whole time.</p>",
          "rawMarkdown": "Why would you extend the competition to the same full time span? That does not make any sense. The leak was found ~4 weeks ago, and now it was extended ~3 weeks. I think this is more than enough, specifically given that you could continue working on the problem the whole time.",
          "votes": 11
        },
        {
          "id": 1388841,
          "postDate": "2021-07-15T08:59:07.503Z",
          "content": "<p>Looks fair to me as well. The new data is also sufficiently similar to the old data, so there was a possibility to continue experimenting even after the leak was found.</p>",
          "rawMarkdown": "Looks fair to me as well. The new data is also sufficiently similar to the old data, so there was a possibility to continue experimenting even after the leak was found.",
          "votes": 7
        },
        {
          "id": 1389234,
          "postDate": "2021-07-15T14:38:16.580Z",
          "content": "<p>We have 5 weeks left, which looks like what was left when the leaks were disclosed. Not sure what the issue is.</p>",
          "rawMarkdown": "We have 5 weeks left, which looks like what was left when the leaks were disclosed. Not sure what the issue is.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1388407,
      "postDate": "2021-07-14T23:05:35.323Z",
      "content": "<p>good job! I will continue this competition</p>",
      "rawMarkdown": "good job! I will continue this competition",
      "votes": 5
    },
    {
      "id": 1388499,
      "postDate": "2021-07-15T02:40:26.873Z",
      "content": "<p>Good to know this, let's find 👽 AGAIN!</p>",
      "rawMarkdown": "Good to know this, let's find 👽 AGAIN!",
      "votes": 6,
      "replies": [
        {
          "id": 1480475,
          "postDate": "2021-08-19T03:27:30.517Z",
          "content": "<p>Let's find!</p>",
          "rawMarkdown": "Let's find!"
        }
      ]
    },
    {
      "id": 1388788,
      "postDate": "2021-07-15T08:21:33.290Z",
      "content": "<p>Can you separate the download links of the new and old data? It is just unreal to have 145 GB data downloaded and then extracted on one machine. The data set is just TOO LARGE!</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Can you separate the download links of the new and old data? It is just unreal to have 145 GB data downloaded and then extracted on one machine. The data set is just TOO LARGE!\n\nThank you!",
      "votes": 3,
      "replies": [
        {
          "id": 1388791,
          "postDate": "2021-07-15T08:25:09.503Z",
          "content": "<p>You can simply click the folders/files that you want to download on the data page without the need to download everything at once.</p>",
          "rawMarkdown": "You can simply click the folders/files that you want to download on the data page without the need to download everything at once."
        },
        {
          "id": 1388795,
          "postDate": "2021-07-15T08:29:35.337Z",
          "content": "<p>I am using cloud GPU servers and using the Kaggle API to download the dataset…can we choose which folders/files to download using command line?</p>",
          "rawMarkdown": "I am using cloud GPU servers and using the Kaggle API to download the dataset...can we choose which folders/files to download using command line?"
        },
        {
          "id": 1388806,
          "postDate": "2021-07-15T08:38:24.987Z",
          "content": "<p>It seems that Kaggle API allows to download files separately, have a look on it's GitHub page <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">https://github.com/Kaggle/kaggle-api</a>. Just list the files available in the competition with <code>kaggle competitions files [-h] [-v] [-q] [competition]</code> and download needed files with <code>kaggle competitions download [-h] [-f FILE_NAME] [-p PATH] [-w] [-o] [-q] [competition]</code>.</p>",
          "rawMarkdown": "It seems that Kaggle API allows to download files separately, have a look on it's GitHub page https://github.com/Kaggle/kaggle-api. Just list the files available in the competition with `kaggle competitions files [-h] [-v] [-q] [competition]` and download needed files with `kaggle competitions download [-h] [-f FILE_NAME] [-p PATH] [-w] [-o] [-q] [competition]`."
        },
        {
          "id": 1388809,
          "postDate": "2021-07-15T08:40:10.793Z",
          "content": "<p>I think you can download selected files using:</p>\n<pre><code>kaggle competitions download [-f FILE_NAME] \n</code></pre>\n<p>Check out Kaggle API documentation <a href=\"https://github.com/Kaggle/kaggle-api#competitions\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "I think you can download selected files using:\n\n```\nkaggle competitions download [-f FILE_NAME] \n```\n\nCheck out Kaggle API documentation [here](https://github.com/Kaggle/kaggle-api#competitions).",
          "votes": 2
        },
        {
          "id": 1388811,
          "postDate": "2021-07-15T08:41:21.707Z",
          "content": "<p>I'm not sure that downloading files individually is possible since there's a per day limit on the number of calls to the Kaggle API. I tried <a href=\"https://www.kaggle.com/product-feedback/246258\" target=\"_blank\">that once</a> but ended with a 429 - too many requests.</p>",
          "rawMarkdown": "I'm not sure that downloading files individually is possible since there's a per day limit on the number of calls to the Kaggle API. I tried [that once](https://www.kaggle.com/product-feedback/246258) but ended with a 429 - too many requests."
        },
        {
          "id": 1388828,
          "postDate": "2021-07-15T08:50:22.030Z",
          "content": "<p>Thank you all!<br>\nI will check out the documentation carefully~</p>",
          "rawMarkdown": "Thank you all!\nI will check out the documentation carefully~"
        },
        {
          "id": 1392560,
          "postDate": "2021-07-18T20:06:40.080Z",
          "content": "<p>Option 1:</p>\n<p>Make a new private dataset on kaggle with only new data. This is easy. In kaggle kernel, import competition data &amp; copy all the data to /tmp/. Then use kaggle api to create private dataset. Then on cloud, use kaggle api to download private dataset. </p>\n<p>Option 2:</p>\n<p>Every Kaggle Dataset has some gcs path associated with it. You can get the GCS path in kaggle kernel using following command:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-covid-tfr-512')\n</code></pre>\n<p>Use the GCS path to download specific folder in the competition dataset<br>\n<code>gsutil -m cp -r GCS_PATH/folder_you_want_to_download destination</code></p>",
          "rawMarkdown": "Option 1:\n\nMake a new private dataset on kaggle with only new data. This is easy. In kaggle kernel, import competition data & copy all the data to /tmp/. Then use kaggle api to create private dataset. Then on cloud, use kaggle api to download private dataset. \n\nOption 2:\n\nEvery Kaggle Dataset has some gcs path associated with it. You can get the GCS path in kaggle kernel using following command:\n```\n\nfrom kaggle_datasets import KaggleDatasets\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-covid-tfr-512')\n```\n\nUse the GCS path to download specific folder in the competition dataset\n`gsutil -m cp -r GCS_PATH/folder_you_want_to_download destination`\n",
          "votes": 5
        },
        {
          "id": 1393150,
          "postDate": "2021-07-19T12:24:56.927Z",
          "content": "<p>Thanks for these tips can be used for the new notebooks,</p>",
          "rawMarkdown": "Thanks for these tips can be used for the new notebooks,"
        }
      ]
    },
    {
      "id": 1389178,
      "postDate": "2021-07-15T13:56:32.843Z",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>: Thanks for your efforts!</p>",
      "rawMarkdown": "@inversion: Thanks for your efforts!",
      "votes": 4,
      "replies": [
        {
          "id": 1389219,
          "postDate": "2021-07-15T14:29:10.157Z",
          "content": "<p>I concur, lots of work done in a short time frame.  Host can be thanked too!</p>",
          "rawMarkdown": "I concur, lots of work done in a short time frame.  Host can be thanked too!",
          "votes": 5
        },
        {
          "id": 1389566,
          "postDate": "2021-07-15T20:30:34.950Z",
          "content": "<p>I concur as well!! Thanks, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>!</p>",
          "rawMarkdown": "I concur as well!! Thanks, @inversion!"
        },
        {
          "id": 1393148,
          "postDate": "2021-07-19T12:23:54.850Z",
          "content": "<p>I concur too, thanks <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> </p>",
          "rawMarkdown": "I concur too, thanks @felipekitamura ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1480395,
      "postDate": "2021-08-19T02:25:54.393Z",
      "content": "<p>Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:</p>\n<p>One account per participant<br>\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.<br>\nNo private sharing outside teams<br>\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.<br>\nbut our results were cancelled. Is there a staff member who can explain?</p>",
      "rawMarkdown": "Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:\n\nOne account per participant\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.\nNo private sharing outside teams\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.\nbut our results were cancelled. Is there a staff member who can explain?",
      "votes": 1
    },
    {
      "id": 1389772,
      "postDate": "2021-07-16T04:17:19.797Z",
      "content": "<p>I try it!<br>\nThe is my first try for kaggle competition!<br>\nPlease cheer up and upvote me plz!</p>",
      "rawMarkdown": "I try it!\nThe is my first try for kaggle competition!\nPlease cheer up and upvote me plz!"
    },
    {
      "id": 1388572,
      "postDate": "2021-07-15T04:45:38.910Z",
      "content": "<p>Awesome, Martians we are coming for you 👽 👌👀😄😍😎🤠</p>",
      "rawMarkdown": "Awesome, Martians we are coming for you 👽 👌👀😄😍😎🤠",
      "votes": -1
    },
    {
      "id": 1465620,
      "postDate": "2021-08-11T05:47:26.107Z",
      "content": "<p>Thanks for this! This is my first competition in kaggle and I'm enjoying it!</p>",
      "rawMarkdown": "Thanks for this! This is my first competition in kaggle and I'm enjoying it!"
    },
    {
      "id": 1393147,
      "postDate": "2021-07-19T12:23:24.757Z",
      "content": "<p>Thanks for letting us know and also making the changes.</p>\n<p>I still see notebooks with scores close to 100%, is it possible not all notebooks have been reset and still needs some more time before we can see the new changes take effect?</p>",
      "rawMarkdown": "Thanks for letting us know and also making the changes.\n\nI still see notebooks with scores close to 100%, is it possible not all notebooks have been reset and still needs some more time before we can see the new changes take effect?",
      "replies": [
        {
          "id": 1393474,
          "postDate": "2021-07-19T17:17:07.757Z",
          "content": "<p>Unfortunately, it's a known bug, and we don't have timing for when it will be resolved.</p>",
          "rawMarkdown": "Unfortunately, it's a known bug, and we don't have timing for when it will be resolved.",
          "votes": 2
        },
        {
          "id": 1394041,
          "postDate": "2021-07-20T06:02:59.580Z",
          "content": "<p>okay no worries, thanks for letting us know</p>",
          "rawMarkdown": "okay no worries, thanks for letting us know"
        }
      ]
    },
    {
      "id": 1392163,
      "postDate": "2021-07-18T12:01:03.227Z",
      "content": "<p>Great job !!</p>",
      "rawMarkdown": "Great job !!"
    },
    {
      "id": 1390188,
      "postDate": "2021-07-16T12:39:59.893Z",
      "content": "<p>comment 101</p>",
      "rawMarkdown": "comment 101\n"
    },
    {
      "id": 1389574,
      "postDate": "2021-07-15T20:42:17.030Z",
      "content": "<p>Thanks for the this action and your effort.</p>",
      "rawMarkdown": "Thanks for the this action and your effort."
    },
    {
      "id": 1406552,
      "postDate": "2021-07-31T23:30:20.817Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 1401118,
      "postDate": "2021-07-27T02:46:24.843Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 1391581,
      "postDate": "2021-07-17T18:20:41.293Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1390209,
      "postDate": "2021-07-16T13:00:38.443Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1479537,
      "postDate": "2021-08-18T14:11:36.753Z",
      "content": "<p>Thanks for updating</p>",
      "rawMarkdown": "Thanks for updating"
    },
    {
      "id": 1408564,
      "postDate": "2021-08-02T14:00:22.130Z",
      "content": "<p>Thanks, will be part of it</p>",
      "rawMarkdown": "Thanks, will be part of it"
    },
    {
      "id": 1393400,
      "postDate": "2021-07-19T15:53:40.893Z",
      "content": "<p>Thank you for updating</p>",
      "rawMarkdown": "Thank you for updating"
    }
  ],
  "comments": [
    {
      "id": 1389223,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-07-15T14:30:25.157000",
      "content": "<p>Funny side effect: finding best public score is a challenge now with all notebooks with score 1.0 from before the relaunch.  LOL.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1388724,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2021-07-15T07:06:36.723000",
      "content": "<p>why  deadline has been extended to Aug 18 only? i think when this competition was launched first it had deadline of around 60+ days, Aug 18 is very aggressive deadline,,with new data we expected 60+ days deadline like it was before with old data,,,,considering other good active kaggle competitions like covid19,commonlit this new deadline is aggressive  :(</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1388829,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-07-15T08:50:58.650000",
          "content": "<p>Why would you extend the competition to the same full time span? That does not make any sense. The leak was found ~4 weeks ago, and now it was extended ~3 weeks. I think this is more than enough, specifically given that you could continue working on the problem the whole time.</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1388841,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-07-15T08:59:07.503000",
          "content": "<p>Looks fair to me as well. The new data is also sufficiently similar to the old data, so there was a possibility to continue experimenting even after the leak was found.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1389234,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-15T14:38:16.580000",
          "content": "<p>We have 5 weeks left, which looks like what was left when the leaks were disclosed. Not sure what the issue is.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1388407,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-07-14T23:05:35.323000",
      "content": "<p>good job! I will continue this competition</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1388499,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2021-07-15T02:40:26.873000",
      "content": "<p>Good to know this, let's find 👽 AGAIN!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1480475,
          "author_name": "Ying Peng",
          "author_url": "",
          "post_date": "2021-08-19T03:27:30.517000",
          "content": "<p>Let's find!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1388788,
      "author_name": "HugoHuya",
      "author_url": "",
      "post_date": "2021-07-15T08:21:33.290000",
      "content": "<p>Can you separate the download links of the new and old data? It is just unreal to have 145 GB data downloaded and then extracted on one machine. The data set is just TOO LARGE!</p>\n<p>Thank you!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1388791,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-07-15T08:25:09.503000",
          "content": "<p>You can simply click the folders/files that you want to download on the data page without the need to download everything at once.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1388795,
          "author_name": "HugoHuya",
          "author_url": "",
          "post_date": "2021-07-15T08:29:35.337000",
          "content": "<p>I am using cloud GPU servers and using the Kaggle API to download the dataset…can we choose which folders/files to download using command line?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1388806,
          "author_name": "Oleg Panichev",
          "author_url": "",
          "post_date": "2021-07-15T08:38:24.987000",
          "content": "<p>It seems that Kaggle API allows to download files separately, have a look on it's GitHub page <a href=\"https://github.com/Kaggle/kaggle-api\" target=\"_blank\">https://github.com/Kaggle/kaggle-api</a>. Just list the files available in the competition with <code>kaggle competitions files [-h] [-v] [-q] [competition]</code> and download needed files with <code>kaggle competitions download [-h] [-f FILE_NAME] [-p PATH] [-w] [-o] [-q] [competition]</code>.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1388809,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-07-15T08:40:10.793000",
          "content": "<p>I think you can download selected files using:</p>\n<pre><code>kaggle competitions download [-f FILE_NAME] \n</code></pre>\n<p>Check out Kaggle API documentation <a href=\"https://github.com/Kaggle/kaggle-api#competitions\" target=\"_blank\">here</a>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1388811,
          "author_name": "FabienDaniel",
          "author_url": "",
          "post_date": "2021-07-15T08:41:21.707000",
          "content": "<p>I'm not sure that downloading files individually is possible since there's a per day limit on the number of calls to the Kaggle API. I tried <a href=\"https://www.kaggle.com/product-feedback/246258\" target=\"_blank\">that once</a> but ended with a 429 - too many requests.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1388828,
          "author_name": "HugoHuya",
          "author_url": "",
          "post_date": "2021-07-15T08:50:22.030000",
          "content": "<p>Thank you all!<br>\nI will check out the documentation carefully~</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1392560,
          "author_name": "Kumar Shubham",
          "author_url": "",
          "post_date": "2021-07-18T20:06:40.080000",
          "content": "<p>Option 1:</p>\n<p>Make a new private dataset on kaggle with only new data. This is easy. In kaggle kernel, import competition data &amp; copy all the data to /tmp/. Then use kaggle api to create private dataset. Then on cloud, use kaggle api to download private dataset. </p>\n<p>Option 2:</p>\n<p>Every Kaggle Dataset has some gcs path associated with it. You can get the GCS path in kaggle kernel using following command:</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nGCS_PATH = KaggleDatasets().get_gcs_path('siim-covid-tfr-512')\n</code></pre>\n<p>Use the GCS path to download specific folder in the competition dataset<br>\n<code>gsutil -m cp -r GCS_PATH/folder_you_want_to_download destination</code></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1393150,
          "author_name": "Mani Sarkar",
          "author_url": "",
          "post_date": "2021-07-19T12:24:56.927000",
          "content": "<p>Thanks for these tips can be used for the new notebooks,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1389178,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2021-07-15T13:56:32.843000",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>: Thanks for your efforts!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1389219,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-15T14:29:10.157000",
          "content": "<p>I concur, lots of work done in a short time frame.  Host can be thanked too!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1389566,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-07-15T20:30:34.950000",
          "content": "<p>I concur as well!! Thanks, <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1393148,
          "author_name": "Mani Sarkar",
          "author_url": "",
          "post_date": "2021-07-19T12:23:54.850000",
          "content": "<p>I concur too, thanks <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1480395,
      "author_name": "Ctrl_CV",
      "author_url": "",
      "post_date": "2021-08-19T02:25:54.393000",
      "content": "<p>Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:</p>\n<p>One account per participant<br>\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.<br>\nNo private sharing outside teams<br>\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.<br>\nbut our results were cancelled. Is there a staff member who can explain?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1389772,
      "author_name": "DaehyeonKyeong",
      "author_url": "",
      "post_date": "2021-07-16T04:17:19.797000",
      "content": "<p>I try it!<br>\nThe is my first try for kaggle competition!<br>\nPlease cheer up and upvote me plz!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1388572,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-07-15T04:45:38.910000",
      "content": "<p>Awesome, Martians we are coming for you 👽 👌👀😄😍😎🤠</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1465620,
      "author_name": "SafeAndSound",
      "author_url": "",
      "post_date": "2021-08-11T05:47:26.107000",
      "content": "<p>Thanks for this! This is my first competition in kaggle and I'm enjoying it!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1393147,
      "author_name": "Mani Sarkar",
      "author_url": "",
      "post_date": "2021-07-19T12:23:24.757000",
      "content": "<p>Thanks for letting us know and also making the changes.</p>\n<p>I still see notebooks with scores close to 100%, is it possible not all notebooks have been reset and still needs some more time before we can see the new changes take effect?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1393474,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2021-07-19T17:17:07.757000",
          "content": "<p>Unfortunately, it's a known bug, and we don't have timing for when it will be resolved.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1394041,
          "author_name": "Mani Sarkar",
          "author_url": "",
          "post_date": "2021-07-20T06:02:59.580000",
          "content": "<p>okay no worries, thanks for letting us know</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1392163,
      "author_name": "Mohamed Hany",
      "author_url": "",
      "post_date": "2021-07-18T12:01:03.227000",
      "content": "<p>Great job !!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1390188,
      "author_name": "Zack Hingley",
      "author_url": "",
      "post_date": "2021-07-16T12:39:59.893000",
      "content": "<p>comment 101</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1389574,
      "author_name": "Minati Ghosh",
      "author_url": "",
      "post_date": "2021-07-15T20:42:17.030000",
      "content": "<p>Thanks for the this action and your effort.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1406552,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-31T23:30:20.817000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1401118,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-27T02:46:24.843000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 1391581,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-17T18:20:41.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1390209,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-16T13:00:38.443000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1479537,
      "author_name": "Ying Peng",
      "author_url": "",
      "post_date": "2021-08-18T14:11:36.753000",
      "content": "<p>Thanks for updating</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1408564,
      "author_name": "Sham",
      "author_url": "",
      "post_date": "2021-08-02T14:00:22.130000",
      "content": "<p>Thanks, will be part of it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1393400,
      "author_name": "Harsh Bansal",
      "author_url": "",
      "post_date": "2021-07-19T15:53:40.893000",
      "content": "<p>Thank you for updating</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1388398": "First of all, thanks for everyone's patience, and many thanks to those who highlighted the issues with the previous dataset.\n\nImmediately after posting this, I'll be swapping in the new dataset to the Data Page. The updated dataset has a folder `old_leaky_data` which contains the previous competition data as well as the corresponding test labels. You should not assume the old data is comparable to the new dataset. (It might help, it might not.)\n\nWhile we've taken efforts to remove the leakage artifacts that were present in the previous dataset, it's difficult to know whether or not the new dataset has some undetected surprises. Remember, we don't actually know what a \"true\" ET signal is supposed to look like. This competition is about building anomaly detection methods that the Breakthrough Listen team and others can use to efficiently and effectively process the vast amounts of data they receive every year. This competition is a way for Kagglers to help contribute to that cause.\n\nWith the new dataset, the competition leaderboard will be reset. This typically takes a few minutes, and you might see some strange LB behavior during the process. \n\nFinally, the competition deadline has been extended to Aug 18.\n\nAgain, apologies for the inconveniences!",
    "1389223": "Funny side effect: finding best public score is a challenge now with all notebooks with score 1.0 from before the relaunch.  LOL.",
    "1388724": "why  deadline has been extended to Aug 18 only? i think when this competition was launched first it had deadline of around 60+ days, Aug 18 is very aggressive deadline,,with new data we expected 60+ days deadline like it was before with old data,,,,considering other good active kaggle competitions like covid19,commonlit this new deadline is aggressive  :(",
    "1388407": "good job! I will continue this competition",
    "1388499": "Good to know this, let's find 👽 AGAIN!",
    "1388788": "Can you separate the download links of the new and old data? It is just unreal to have 145 GB data downloaded and then extracted on one machine. The data set is just TOO LARGE!\n\nThank you!",
    "1389178": "@inversion: Thanks for your efforts!",
    "1480395": "Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:\n\nOne account per participant\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.\nNo private sharing outside teams\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.\nbut our results were cancelled. Is there a staff member who can explain?",
    "1389772": "I try it!\nThe is my first try for kaggle competition!\nPlease cheer up and upvote me plz!",
    "1388572": "Awesome, Martians we are coming for you 👽 👌👀😄😍😎🤠",
    "1465620": "Thanks for this! This is my first competition in kaggle and I'm enjoying it!",
    "1393147": "Thanks for letting us know and also making the changes.\n\nI still see notebooks with scores close to 100%, is it possible not all notebooks have been reset and still needs some more time before we can see the new changes take effect?",
    "1392163": "Great job !!",
    "1390188": "comment 101\n",
    "1389574": "Thanks for the this action and your effort.",
    "1406552": "",
    "1401118": "",
    "1391581": "",
    "1390209": "",
    "1479537": "Thanks for updating",
    "1408564": "Thanks, will be part of it",
    "1393400": "Thank you for updating"
  }
}