{
  "id": 248755,
  "title": "(Partial) Competition Update",
  "url": "/competitions/seti-breakthrough-listen/discussion/248755",
  "author_name": "inversion",
  "post_date": "2021-06-24T19:38:18.762000",
  "votes": 68,
  "comment_count": 45,
  "views": 0,
  "content": "<p>Hi everyone -</p>\n<p>Here's an update of where things stand on the competition reset:</p>\n<ul>\n<li>A new test set has been generated, taking into account the fantastic feedback we've received from the community. (Thank you for all of the insights!)</li>\n<li>We're in the process of scrutinizing the data, and are targeting a relaunch mid-next week (but may need more time if we find any issues along the way).</li>\n<li>The new training set will include the current training <em>and</em> test set (with the test set labels)</li>\n<li>We'll adjust the competition timeline at that time to reflect the time lost with the reset</li>\n</ul>\n<p>An additional bit of detail on the time stamp leakage failure point (which I did a post-mortem for):</p>\n<p>Kaggle actually has a triple-redundant system for time stamp leakage, where all points failed: I <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782\" target=\"_blank\">previously mentioned</a> our best practices of random file saves and secondary checks, but our system <em>should</em> have been tolerant to any human failure points, since our backend actually re-creates time stamps for our competition files on the Data page after upload. In this instance, because our zipping process was erroring out, the \"Download All\" zip file was manually put into position, rather than being automatically created from the uploaded competition files. (Which, by the way, is the reason the timestamp leakage <em>didn't</em> show up when the data in the Notebooks environment, because it was processed normally, and only when using the downloaded zip file from the competition page, which wasn't.) </p>\n<p>It's a good lesson on the importances of sticking to best practices. :-)</p>",
  "messages": [
    {
      "id": 1364280,
      "postDate": "2021-06-24T19:38:18.763Z",
      "content": "<p>Hi everyone -</p>\n<p>Here's an update of where things stand on the competition reset:</p>\n<ul>\n<li>A new test set has been generated, taking into account the fantastic feedback we've received from the community. (Thank you for all of the insights!)</li>\n<li>We're in the process of scrutinizing the data, and are targeting a relaunch mid-next week (but may need more time if we find any issues along the way).</li>\n<li>The new training set will include the current training <em>and</em> test set (with the test set labels)</li>\n<li>We'll adjust the competition timeline at that time to reflect the time lost with the reset</li>\n</ul>\n<p>An additional bit of detail on the time stamp leakage failure point (which I did a post-mortem for):</p>\n<p>Kaggle actually has a triple-redundant system for time stamp leakage, where all points failed: I <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782\" target=\"_blank\">previously mentioned</a> our best practices of random file saves and secondary checks, but our system <em>should</em> have been tolerant to any human failure points, since our backend actually re-creates time stamps for our competition files on the Data page after upload. In this instance, because our zipping process was erroring out, the \"Download All\" zip file was manually put into position, rather than being automatically created from the uploaded competition files. (Which, by the way, is the reason the timestamp leakage <em>didn't</em> show up when the data in the Notebooks environment, because it was processed normally, and only when using the downloaded zip file from the competition page, which wasn't.) </p>\n<p>It's a good lesson on the importances of sticking to best practices. :-)</p>",
      "rawMarkdown": "Hi everyone -\n\nHere's an update of where things stand on the competition reset:\n- A new test set has been generated, taking into account the fantastic feedback we've received from the community. (Thank you for all of the insights!)\n- We're in the process of scrutinizing the data, and are targeting a relaunch mid-next week (but may need more time if we find any issues along the way).\n- The new training set will include the current training _and_ test set (with the test set labels)\n- We'll adjust the competition timeline at that time to reflect the time lost with the reset\n\nAn additional bit of detail on the time stamp leakage failure point (which I did a post-mortem for):\n\nKaggle actually has a triple-redundant system for time stamp leakage, where all points failed: I [previously mentioned](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782) our best practices of random file saves and secondary checks, but our system *should* have been tolerant to any human failure points, since our backend actually re-creates time stamps for our competition files on the Data page after upload. In this instance, because our zipping process was erroring out, the \"Download All\" zip file was manually put into position, rather than being automatically created from the uploaded competition files. (Which, by the way, is the reason the timestamp leakage *didn't* show up when the data in the Notebooks environment, because it was processed normally, and only when using the downloaded zip file from the competition page, which wasn't.) \n\nIt's a good lesson on the importances of sticking to best practices. :-)",
      "votes": 67
    },
    {
      "id": 1373457,
      "postDate": "2021-07-02T13:57:15.993Z",
      "content": "<p>Were we currently stand: The data generation has taken a bit longer after it was determined the training set would also need to be re-generated, but it is nearly complete. We're going into a US long weekend, and I don't expect significant work over the next 4 days, but I'll pick things back up next Wednesday. With processing and checking, I'd imagine things will be ready for relaunch early the week of July 12.</p>",
      "rawMarkdown": "Were we currently stand: The data generation has taken a bit longer after it was determined the training set would also need to be re-generated, but it is nearly complete. We're going into a US long weekend, and I don't expect significant work over the next 4 days, but I'll pick things back up next Wednesday. With processing and checking, I'd imagine things will be ready for relaunch early the week of July 12.",
      "votes": 35,
      "replies": [
        {
          "id": 1386038,
          "postDate": "2021-07-13T07:45:46.483Z",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Any updates ?</p>",
          "rawMarkdown": "@inversion Any updates ?",
          "votes": 1
        },
        {
          "id": 1386062,
          "postDate": "2021-07-13T08:13:07.793Z",
          "content": "<p><strong>early the week of July 12</strong> , which means July 13-July 15 is possible</p>",
          "rawMarkdown": "**early the week of July 12** , which means July 13-July 15 is possible",
          "votes": 3
        }
      ]
    },
    {
      "id": 1371593,
      "postDate": "2021-07-01T05:50:18.417Z",
      "content": "<p>Posting as a top-level comment for visibility:</p>\n<p>Just investigated whether renormalizing the current training set is able to remove the statistics leak, and unfortunately the truncation from casting into FP16 caused minute differences in the mean and standard deviations between positive and negative samples that we weren't able to remove.</p>\n<p>I will be discussing with <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> to decide on the best way to mitigate this. We will most likely be generating a new training set without leaks.</p>\n<p>Average of Means from On and Off (blue is needles, orange is new haystack)</p>\n<blockquote><a href=\"//imgur.com/i1qe3pE\">Average Means from On and Off</a></blockquote>\n\n<p>Average of STDs from On and Off (blue is needles, orange is new haystack)</p>\n<blockquote><a href=\"//imgur.com/a/ZVKIYAG\">Average of STDs from On and Off</a></blockquote>\n",
      "rawMarkdown": "Posting as a top-level comment for visibility:\n\nJust investigated whether renormalizing the current training set is able to remove the statistics leak, and unfortunately the truncation from casting into FP16 caused minute differences in the mean and standard deviations between positive and negative samples that we weren't able to remove.\n\nI will be discussing with @inversion to decide on the best way to mitigate this. We will most likely be generating a new training set without leaks.\n\nAverage of Means from On and Off (blue is needles, orange is new haystack)\n<blockquote class=\"imgur-embed-pub\" lang=\"en\" data-id=\"i1qe3pE\"  ><a href=\"//imgur.com/i1qe3pE\">Average Means from On and Off</a></blockquote><script async src=\"//s.imgur.com/min/embed.js\" charset=\"utf-8\"></script>\n\nAverage of STDs from On and Off (blue is needles, orange is new haystack)\n<blockquote class=\"imgur-embed-pub\" lang=\"en\" data-id=\"a/ZVKIYAG\"  ><a href=\"//imgur.com/a/ZVKIYAG\">Average of STDs from On and Off</a></blockquote><script async src=\"//s.imgur.com/min/embed.js\" charset=\"utf-8\"></script>",
      "votes": 18,
      "replies": [
        {
          "id": 1371944,
          "postDate": "2021-07-01T10:16:54.833Z",
          "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> - if you would create a new training set, would you create the exact same current train / test data? Otherwise I think people will still try to re-use current data to make the data larger, and then I am wondering whether it could also be an option for people to try to remove the inherent bias themselves. Of course as long as the test data has no bias.<br>\nCurious what others think.</p>",
          "rawMarkdown": "Thanks for the update @yuhongc - if you would create a new training set, would you create the exact same current train / test data? Otherwise I think people will still try to re-use current data to make the data larger, and then I am wondering whether it could also be an option for people to try to remove the inherent bias themselves. Of course as long as the test data has no bias.\nCurious what others think.",
          "votes": 7
        },
        {
          "id": 1372004,
          "postDate": "2021-07-01T10:55:43.763Z",
          "content": "<p>Does this mean that it will take more time for the dataset to be updated ? If so, Do you have an estimate on when it will go live.</p>",
          "rawMarkdown": "Does this mean that it will take more time for the dataset to be updated ? If so, Do you have an estimate on when it will go live."
        },
        {
          "id": 1372041,
          "postDate": "2021-07-01T11:26:58.937Z",
          "content": "<p>Thanks for sharing the info with us!</p>\n<p>A small statistical leak in the training data is not as critical as in the test data.<br>\nOf course, it needs to be verified, but if it is difficult to prepare new training data, it may be enough to re-normalize the current training data even if there are still small leaks.<br>\nAs <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> pointed out, even if you prepare completely new training data, many participants will try to use the old training data along with it, and in that case, participants will eventually have to deal with this small leak.</p>",
          "rawMarkdown": "Thanks for sharing the info with us!\n\nA small statistical leak in the training data is not as critical as in the test data.\nOf course, it needs to be verified, but if it is difficult to prepare new training data, it may be enough to re-normalize the current training data even if there are still small leaks.\nAs @philippsinger pointed out, even if you prepare completely new training data, many participants will try to use the old training data along with it, and in that case, participants will eventually have to deal with this small leak.",
          "votes": 1
        },
        {
          "id": 1372060,
          "postDate": "2021-07-01T11:51:39.547Z",
          "content": "<p>Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.</p>",
          "rawMarkdown": "Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.",
          "votes": 7
        },
        {
          "id": 1373353,
          "postDate": "2021-07-02T12:32:51.197Z",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a>! Do you think generating the entire data (train and test) as uint16 could help?</p>",
          "rawMarkdown": "Thank you, @yuhongc! Do you think generating the entire data (train and test) as uint16 could help?",
          "votes": 1
        },
        {
          "id": 1373401,
          "postDate": "2021-07-02T13:17:29.127Z",
          "content": "<blockquote>\n  <p>Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.</p>\n</blockquote>\n<p>I agree.</p>\n<p>However, this assumes that any randomness used for generating signals can be reproduced.  I hope all random seeds were saved…</p>",
          "rawMarkdown": "> Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.\n\nI agree.\n\nHowever, this assumes that any randomness used for generating signals can be reproduced.  I hope all random seeds were saved...",
          "votes": 3
        },
        {
          "id": 1374969,
          "postDate": "2021-07-03T18:35:46.200Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1365400,
      "postDate": "2021-06-25T17:17:51.893Z",
      "content": "<p>Can it be a code competition with the new data?</p>",
      "rawMarkdown": "Can it be a code competition with the new data?",
      "votes": 11,
      "replies": [
        {
          "id": 1370229,
          "postDate": "2021-06-30T03:05:15.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Please Take this into consideration </p>",
          "rawMarkdown": "@inversion Please Take this into consideration ",
          "votes": 1
        },
        {
          "id": 1370483,
          "postDate": "2021-06-30T07:35:39.150Z",
          "content": "<p>I am really glad that more and more people see the advantages of code competitions :) I always try to argue for it.</p>",
          "rawMarkdown": "I am really glad that more and more people see the advantages of code competitions :) I always try to argue for it.",
          "votes": 8
        },
        {
          "id": 1370836,
          "postDate": "2021-06-30T12:48:07.573Z",
          "content": "<p>I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.</p>",
          "rawMarkdown": "I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.",
          "votes": 1
        },
        {
          "id": 1370838,
          "postDate": "2021-06-30T12:49:12.073Z",
          "content": "<p>Main pro for code competition is to make it harder to use test data distribution to overfit models to test data.</p>",
          "rawMarkdown": "Main pro for code competition is to make it harder to use test data distribution to overfit models to test data."
        },
        {
          "id": 1370850,
          "postDate": "2021-06-30T12:57:19.890Z",
          "content": "<blockquote>\n  <p>I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.</p>\n</blockquote>\n<p>It will be difficult for them to afford anything of a substantial upgrade to the P100. Maybe use a T4 cause it`s better in Mixed_float16 ?</p>",
          "rawMarkdown": "> I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.\n\n\n  It will be difficult for them to afford anything of a substantial upgrade to the P100. Maybe use a T4 cause it`s better in Mixed_float16 ?"
        },
        {
          "id": 1370911,
          "postDate": "2021-06-30T13:37:45.320Z",
          "content": "<p>For inference P100 is more than enough, not sure why you would need something better.</p>",
          "rawMarkdown": "For inference P100 is more than enough, not sure why you would need something better.",
          "votes": 3
        },
        {
          "id": 1371050,
          "postDate": "2021-06-30T15:53:05.347Z",
          "content": "<p>Small GPU memory can be an issue in some cases.  </p>",
          "rawMarkdown": "Small GPU memory can be an issue in some cases.  "
        },
        {
          "id": 1371054,
          "postDate": "2021-06-30T15:54:33.797Z",
          "content": "<p>I dont know of any model not fitting into P100 memory, apart from some ultra heavy transformer nlp models that are basically not used on kaggle anyways. </p>",
          "rawMarkdown": "I dont know of any model not fitting into P100 memory, apart from some ultra heavy transformer nlp models that are basically not used on kaggle anyways. ",
          "votes": 3
        },
        {
          "id": 1371104,
          "postDate": "2021-06-30T16:52:16.440Z",
          "content": "<p>Not all models are NNs.  I got stuck when using RAPIDS once.</p>",
          "rawMarkdown": "Not all models are NNs.  I got stuck when using RAPIDS once."
        },
        {
          "id": 1371884,
          "postDate": "2021-07-01T09:39:22.257Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  no offense but kaggle code competition do replicate real life condition where you do not always  have access to the Best hardware for deployment.</p>",
          "rawMarkdown": "@cpmpml  no offense but kaggle code competition do replicate real life condition where you do not always  have access to the Best hardware for deployment.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1364351,
      "postDate": "2021-06-24T22:03:12.797Z",
      "content": "<p>Good news, thanks.  Will you also provide instructions on how to remove the distribution leak from existing training an d test data? Indeed, if we train on data with the leak and predict on new test without leak then results may be irrelevant.</p>\n<p>I suggested <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194\" target=\"_blank\">some way</a>, but there maybe other, better ways.</p>",
      "rawMarkdown": "Good news, thanks.  Will you also provide instructions on how to remove the distribution leak from existing training an d test data? Indeed, if we train on data with the leak and predict on new test without leak then results may be irrelevant.\n\nI suggested [some way](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194), but there maybe other, better ways.",
      "votes": 4,
      "replies": [
        {
          "id": 1364358,
          "postDate": "2021-06-24T22:09:50.863Z",
          "content": "<p>If possible, it'd be great if instead of just the original (zip) data, when the new training set is provided (current train + current test), they are re-processed according to best practices. That way individual contestants don't need to do any further processing.</p>",
          "rawMarkdown": "If possible, it'd be great if instead of just the original (zip) data, when the new training set is provided (current train + current test), they are re-processed according to best practices. That way individual contestants don't need to do any further processing."
        },
        {
          "id": 1364374,
          "postDate": "2021-06-24T22:33:58.717Z",
          "content": "<p>Fully fixing the distribution leak in the current dataset may be difficult because the data was already in float16 pre-injection. We really didn't see the truncation issue coming 😔.</p>\n<p>We may consider making a new training set with the truncation issue fixed, but some internal discussion is needed before I can say for sure.</p>",
          "rawMarkdown": "Fully fixing the distribution leak in the current dataset may be difficult because the data was already in float16 pre-injection. We really didn't see the truncation issue coming 😔.\n\nWe may consider making a new training set with the truncation issue fixed, but some internal discussion is needed before I can say for sure.",
          "votes": 10
        },
        {
          "id": 1364709,
          "postDate": "2021-06-25T06:49:16.003Z",
          "content": "<p>But you will fix it in new test set, right <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> ?</p>",
          "rawMarkdown": "But you will fix it in new test set, right @yuhongc ?",
          "votes": 5
        },
        {
          "id": 1364965,
          "postDate": "2021-06-25T10:55:53.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> Fixing only in new test isn't enough as it introduces a distribution shift between new train and new test.</p>\n<p>I get that creating a new train set is a lot of work, which is why I am asking if the prepossessing I proposed was enough to remove this distribution shift.</p>\n<p>And thanks for confirming my guess that your performed standardization in FP16.</p>",
          "rawMarkdown": "@yuhongc Fixing only in new test isn't enough as it introduces a distribution shift between new train and new test.\n\nI get that creating a new train set is a lot of work, which is why I am asking if the prepossessing I proposed was enough to remove this distribution shift.\n\nAnd thanks for confirming my guess that your performed standardization in FP16.",
          "votes": 2
        },
        {
          "id": 1364971,
          "postDate": "2021-06-25T11:03:31.737Z",
          "content": "<p>Fixing only new test should be enough, if you standardize properly yourself in training / validation.</p>",
          "rawMarkdown": "Fixing only new test should be enough, if you standardize properly yourself in training / validation.",
          "votes": 2
        },
        {
          "id": 1364993,
          "postDate": "2021-06-25T11:26:07.033Z",
          "content": "<blockquote>\n  <p>if you standardize properly yourself</p>\n</blockquote>\n<p>This is what I proposed, and this is what I am asking confirmation for.</p>\n<p>Given some files were standardized in FP32 and others (snippets with message) in FP16 in current training set, there maybe distribution differences that cannot be removed.</p>",
          "rawMarkdown": ">  if you standardize properly yourself\n\nThis is what I proposed, and this is what I am asking confirmation for.\n\nGiven some files were standardized in FP32 and others (snippets with message) in FP16 in current training set, there maybe distribution differences that cannot be removed.",
          "votes": 5
        },
        {
          "id": 1365465,
          "postDate": "2021-06-25T18:55:32.440Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Really appreciate your comments and support. I will investigate whether the standardization you did or some sort of other preprocessing is enough to fix the distribution difference from quantization.</p>\n<p>If it's indeed irreversible, I think we'll release a replacement training set, (while keeping the old data available out of fairness), but again can't say for sure without internal discussion.</p>",
          "rawMarkdown": "@cpmpml Really appreciate your comments and support. I will investigate whether the standardization you did or some sort of other preprocessing is enough to fix the distribution difference from quantization.\n\nIf it's indeed irreversible, I think we'll release a replacement training set, (while keeping the old data available out of fairness), but again can't say for sure without internal discussion.",
          "votes": 4
        },
        {
          "id": 1365473,
          "postDate": "2021-06-25T19:03:09.737Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> We believe it's fixed in the new test. We'll be verifying this over the next few days. The worry with the current train+test right now is that casting back and forth between FP16 and FP32 (mea culpa) may have introduced tiny irreversible differences in the statistics because of the difference in accuracy between the two data formats.</p>",
          "rawMarkdown": "@philippsinger We believe it's fixed in the new test. We'll be verifying this over the next few days. The worry with the current train+test right now is that casting back and forth between FP16 and FP32 (mea culpa) may have introduced tiny irreversible differences in the statistics because of the difference in accuracy between the two data formats.",
          "votes": 6
        },
        {
          "id": 1365856,
          "postDate": "2021-06-26T07:41:56.873Z",
          "content": "<p><a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> thanks you.  Having a responsive host means a lot for competitors.</p>",
          "rawMarkdown": "@yuhongc thanks you.  Having a responsive host means a lot for competitors.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1381061,
      "postDate": "2021-07-08T15:30:59.720Z",
      "content": "<p>hello <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, <br>\ncould you please provide an update on the datasets? also will the timeline for submission be extened? </p>",
      "rawMarkdown": "hello @inversion, \ncould you please provide an update on the datasets? also will the timeline for submission be extened? ",
      "votes": 1
    },
    {
      "id": 1364437,
      "postDate": "2021-06-25T02:05:44.763Z",
      "content": "<p>Thank you for updating information. I really appreciate your hard work.</p>\n<p>I've explored several methods while the competition was suspended.<br>\nI'm looking forward to trying them out on new datasets😃</p>",
      "rawMarkdown": "Thank you for updating information. I really appreciate your hard work.\n\nI've explored several methods while the competition was suspended.\nI'm looking forward to trying them out on new datasets😃",
      "votes": 2
    },
    {
      "id": 1370452,
      "postDate": "2021-06-30T07:16:51.023Z",
      "content": "<p>Can it be a code competition with the new data?</p>",
      "rawMarkdown": "Can it be a code competition with the new data?",
      "votes": -1
    },
    {
      "id": 1371177,
      "postDate": "2021-06-30T18:19:51.503Z",
      "content": "<p>Hi, can it be a code competition with a new dataset</p>",
      "rawMarkdown": "Hi, can it be a code competition with a new dataset\n",
      "votes": -1
    },
    {
      "id": 1480398,
      "postDate": "2021-08-19T02:26:35.560Z",
      "content": "<p>Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:</p>\n<p>One account per participant<br>\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.<br>\nNo private sharing outside teams<br>\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.<br>\nbut our results were cancelled. Is there a staff member who can explain?</p>",
      "rawMarkdown": "Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:\n\nOne account per participant\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.\nNo private sharing outside teams\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.\nbut our results were cancelled. Is there a staff member who can explain?\n"
    },
    {
      "id": 1371216,
      "postDate": "2021-06-30T18:58:45.170Z",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , One doubt since the competition is going to restart does it means that the public Leaderboard as well all submissions made till now by all participants for this Competition will reset ?</p>",
      "rawMarkdown": "@inversion , One doubt since the competition is going to restart does it means that the public Leaderboard as well all submissions made till now by all participants for this Competition will reset ?"
    },
    {
      "id": 1364559,
      "postDate": "2021-06-25T04:51:20.143Z",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Could you also please check the How do you zip the files. As your adding More files to the train data it will become impossible to use this dataset with Colab Pro. So a well zipped dataset will surely help</p>",
      "rawMarkdown": "@inversion Could you also please check the How do you zip the files. As your adding More files to the train data it will become impossible to use this dataset with Colab Pro. So a well zipped dataset will surely help",
      "replies": [
        {
          "id": 1365015,
          "postDate": "2021-06-25T11:46:31.497Z",
          "content": "<p>use GCS directly (if you use tensorflow)<br>\nyou can access data even if you use colab w/o pro</p>\n<p>but one thing clear is training time will be increase almost 1.7</p>",
          "rawMarkdown": "use GCS directly (if you use tensorflow)\nyou can access data even if you use colab w/o pro\n\nbut one thing clear is training time will be increase almost 1.7",
          "votes": 4
        },
        {
          "id": 1365366,
          "postDate": "2021-06-25T16:43:44.987Z",
          "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> first of all you have pay a lot for GCS. And my training times are like already one hour and increase by 1.7 will mean i have a nearly 2 hour training Time which honestly a lot </p>",
          "rawMarkdown": "@assign first of all you have pay a lot for GCS. And my training times are like already one hour and increase by 1.7 will mean i have a nearly 2 hour training Time which honestly a lot "
        },
        {
          "id": 1365631,
          "postDate": "2021-06-26T00:59:29.067Z",
          "content": "<p>I means</p>\n<ol>\n<li>make it as kaggle dataset </li>\n<li>get gcs bucket id</li>\n<li>read by tensorflow record format</li>\n</ol>\n<p>1 preprocess dataset as tfrec format</p>\n<p>2 Kaggle kernel after add the dataset</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path(&lt;target dataset name&gt;)# without ../input/\n</code></pre>\n<p><a href=\"https://www.kaggle.com/assign/get-data-info\" target=\"_blank\">example of it</a></p>\n<p>example of output:<br>\n'gs://kds-0752b66b41214b9aeeb3741bbd6b1dcf6c443007693b57cc9f4e2ea5'</p>\n<p>3 Colab</p>\n<pre><code>GCS_PATH = &lt;output of upper code&gt;\nfiles = list(np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))) # depend on your file name\nds = tf.data.TFRecordDataset(files, num_parallel_reads = len(files))\nds = ds.map(&lt;your parsing function&gt;)\nds = ds.batch(512)\nfor I in ds: ~~~\n</code></pre>\n<p>you need to preprocess dataset as serialized tensor, <br>\nand design suitable parsing function.<br>\nI recommend to read tf.io documents</p>\n<p>I don't know whether this is google's design for using gcs freely. <br>\nBut, we can use gcs without cost</p>\n<p>well, with this approach, using TPU much easier and we can expect speed up(1 epoch: 17min -&gt; 1min)</p>\n<p>I'll delete it if it violate kaggle community guidelines</p>",
          "rawMarkdown": "I means\n1. make it as kaggle dataset \n2. get gcs bucket id\n3. read by tensorflow record format\n\n\n\n1 preprocess dataset as tfrec format\n\n2 Kaggle kernel after add the dataset\n```\nfrom kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path(<target dataset name>)# without ../input/\n```\n[example of it](https://www.kaggle.com/assign/get-data-info)\n\nexample of output:\n'gs://kds-0752b66b41214b9aeeb3741bbd6b1dcf6c443007693b57cc9f4e2ea5'\n\n\n3 Colab\n```\nGCS_PATH = <output of upper code>\nfiles = list(np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))) # depend on your file name\nds = tf.data.TFRecordDataset(files, num_parallel_reads = len(files))\nds = ds.map(<your parsing function>)\nds = ds.batch(512)\nfor I in ds: ~~~\n```\nyou need to preprocess dataset as serialized tensor, \nand design suitable parsing function.\nI recommend to read tf.io documents\n\nI don't know whether this is google's design for using gcs freely. \nBut, we can use gcs without cost\n\nwell, with this approach, using TPU much easier and we can expect speed up(1 epoch: 17min -> 1min)\n\n\nI'll delete it if it violate kaggle community guidelines",
          "votes": 2
        }
      ]
    },
    {
      "id": 1370957,
      "postDate": "2021-06-30T14:12:53.690Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1370699,
      "postDate": "2021-06-30T10:44:01.613Z",
      "content": "<p>Thank you for the update 😊</p>",
      "rawMarkdown": "Thank you for the update 😊"
    },
    {
      "id": 1364380,
      "postDate": "2021-06-24T22:52:59.537Z",
      "content": "<p>Thank you for the update 😊</p>",
      "rawMarkdown": "Thank you for the update 😊",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1373457,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2021-07-02T13:57:15.993000",
      "content": "<p>Were we currently stand: The data generation has taken a bit longer after it was determined the training set would also need to be re-generated, but it is nearly complete. We're going into a US long weekend, and I don't expect significant work over the next 4 days, but I'll pick things back up next Wednesday. With processing and checking, I'd imagine things will be ready for relaunch early the week of July 12.</p>",
      "votes": 35,
      "replies": [
        {
          "id": 1386038,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-13T07:45:46.483000",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Any updates ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1386062,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-07-13T08:13:07.793000",
          "content": "<p><strong>early the week of July 12</strong> , which means July 13-July 15 is possible</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1371593,
      "author_name": "Yuhong Chen",
      "author_url": "",
      "post_date": "2021-07-01T05:50:18.417000",
      "content": "<p>Posting as a top-level comment for visibility:</p>\n<p>Just investigated whether renormalizing the current training set is able to remove the statistics leak, and unfortunately the truncation from casting into FP16 caused minute differences in the mean and standard deviations between positive and negative samples that we weren't able to remove.</p>\n<p>I will be discussing with <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> to decide on the best way to mitigate this. We will most likely be generating a new training set without leaks.</p>\n<p>Average of Means from On and Off (blue is needles, orange is new haystack)</p>\n<blockquote><a href=\"//imgur.com/i1qe3pE\">Average Means from On and Off</a></blockquote>\n\n<p>Average of STDs from On and Off (blue is needles, orange is new haystack)</p>\n<blockquote><a href=\"//imgur.com/a/ZVKIYAG\">Average of STDs from On and Off</a></blockquote>\n",
      "votes": 18,
      "replies": [
        {
          "id": 1371944,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-07-01T10:16:54.833000",
          "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> - if you would create a new training set, would you create the exact same current train / test data? Otherwise I think people will still try to re-use current data to make the data larger, and then I am wondering whether it could also be an option for people to try to remove the inherent bias themselves. Of course as long as the test data has no bias.<br>\nCurious what others think.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1372004,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-01T10:55:43.763000",
          "content": "<p>Does this mean that it will take more time for the dataset to be updated ? If so, Do you have an estimate on when it will go live.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1372041,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-07-01T11:26:58.937000",
          "content": "<p>Thanks for sharing the info with us!</p>\n<p>A small statistical leak in the training data is not as critical as in the test data.<br>\nOf course, it needs to be verified, but if it is difficult to prepare new training data, it may be enough to re-normalize the current training data even if there are still small leaks.<br>\nAs <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> pointed out, even if you prepare completely new training data, many participants will try to use the old training data along with it, and in that case, participants will eventually have to deal with this small leak.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1372060,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-07-01T11:51:39.547000",
          "content": "<p>Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1373353,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-07-02T12:32:51.197000",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a>! Do you think generating the entire data (train and test) as uint16 could help?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1373401,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-02T13:17:29.127000",
          "content": "<blockquote>\n  <p>Indeed, so I think the issue could only be solved if you re-create the exact same current train/test data without leak.</p>\n</blockquote>\n<p>I agree.</p>\n<p>However, this assumes that any randomness used for generating signals can be reproduced.  I hope all random seeds were saved…</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1374969,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-07-03T18:35:46.200000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1365400,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-06-25T17:17:51.893000",
      "content": "<p>Can it be a code competition with the new data?</p>",
      "votes": 11,
      "replies": [
        {
          "id": 1370229,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-30T03:05:15.877000",
          "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Please Take this into consideration </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1370483,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-30T07:35:39.150000",
          "content": "<p>I am really glad that more and more people see the advantages of code competitions :) I always try to argue for it.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1370836,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-30T12:48:07.573000",
          "content": "<p>I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1370838,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-30T12:49:12.073000",
          "content": "<p>Main pro for code competition is to make it harder to use test data distribution to overfit models to test data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1370850,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-30T12:57:19.890000",
          "content": "<blockquote>\n  <p>I'll be a full supporter of code competitions when we have GPU more recent than 5 years old P100.</p>\n</blockquote>\n<p>It will be difficult for them to afford anything of a substantial upgrade to the P100. Maybe use a T4 cause it`s better in Mixed_float16 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1370911,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-30T13:37:45.320000",
          "content": "<p>For inference P100 is more than enough, not sure why you would need something better.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1371050,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-30T15:53:05.347000",
          "content": "<p>Small GPU memory can be an issue in some cases.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1371054,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-30T15:54:33.797000",
          "content": "<p>I dont know of any model not fitting into P100 memory, apart from some ultra heavy transformer nlp models that are basically not used on kaggle anyways. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1371104,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-30T16:52:16.440000",
          "content": "<p>Not all models are NNs.  I got stuck when using RAPIDS once.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1371884,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-07-01T09:39:22.257000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  no offense but kaggle code competition do replicate real life condition where you do not always  have access to the Best hardware for deployment.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1364351,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-24T22:03:12.797000",
      "content": "<p>Good news, thanks.  Will you also provide instructions on how to remove the distribution leak from existing training an d test data? Indeed, if we train on data with the leak and predict on new test without leak then results may be irrelevant.</p>\n<p>I suggested <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194\" target=\"_blank\">some way</a>, but there maybe other, better ways.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1364358,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-06-24T22:09:50.863000",
          "content": "<p>If possible, it'd be great if instead of just the original (zip) data, when the new training set is provided (current train + current test), they are re-processed according to best practices. That way individual contestants don't need to do any further processing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1364374,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-06-24T22:33:58.717000",
          "content": "<p>Fully fixing the distribution leak in the current dataset may be difficult because the data was already in float16 pre-injection. We really didn't see the truncation issue coming 😔.</p>\n<p>We may consider making a new training set with the truncation issue fixed, but some internal discussion is needed before I can say for sure.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1364709,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-25T06:49:16.003000",
          "content": "<p>But you will fix it in new test set, right <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> ?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1364965,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-25T10:55:53.927000",
          "content": "<p><a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> Fixing only in new test isn't enough as it introduces a distribution shift between new train and new test.</p>\n<p>I get that creating a new train set is a lot of work, which is why I am asking if the prepossessing I proposed was enough to remove this distribution shift.</p>\n<p>And thanks for confirming my guess that your performed standardization in FP16.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1364971,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-25T11:03:31.737000",
          "content": "<p>Fixing only new test should be enough, if you standardize properly yourself in training / validation.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1364993,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-25T11:26:07.033000",
          "content": "<blockquote>\n  <p>if you standardize properly yourself</p>\n</blockquote>\n<p>This is what I proposed, and this is what I am asking confirmation for.</p>\n<p>Given some files were standardized in FP32 and others (snippets with message) in FP16 in current training set, there maybe distribution differences that cannot be removed.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1365465,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-06-25T18:55:32.440000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Really appreciate your comments and support. I will investigate whether the standardization you did or some sort of other preprocessing is enough to fix the distribution difference from quantization.</p>\n<p>If it's indeed irreversible, I think we'll release a replacement training set, (while keeping the old data available out of fairness), but again can't say for sure without internal discussion.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1365473,
          "author_name": "Yuhong Chen",
          "author_url": "",
          "post_date": "2021-06-25T19:03:09.737000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> We believe it's fixed in the new test. We'll be verifying this over the next few days. The worry with the current train+test right now is that casting back and forth between FP16 and FP32 (mea culpa) may have introduced tiny irreversible differences in the statistics because of the difference in accuracy between the two data formats.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1365856,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-26T07:41:56.873000",
          "content": "<p><a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> thanks you.  Having a responsive host means a lot for competitors.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1381061,
      "author_name": "VikramanK",
      "author_url": "",
      "post_date": "2021-07-08T15:30:59.720000",
      "content": "<p>hello <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a>, <br>\ncould you please provide an update on the datasets? also will the timeline for submission be extened? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1364437,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-06-25T02:05:44.763000",
      "content": "<p>Thank you for updating information. I really appreciate your hard work.</p>\n<p>I've explored several methods while the competition was suspended.<br>\nI'm looking forward to trying them out on new datasets😃</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1370452,
      "author_name": "Gia Huy Nguyễn",
      "author_url": "",
      "post_date": "2021-06-30T07:16:51.023000",
      "content": "<p>Can it be a code competition with the new data?</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1371177,
      "author_name": "shiv",
      "author_url": "",
      "post_date": "2021-06-30T18:19:51.503000",
      "content": "<p>Hi, can it be a code competition with a new dataset</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1480398,
      "author_name": "Ctrl_CV",
      "author_url": "",
      "post_date": "2021-08-19T02:26:35.560000",
      "content": "<p>Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:</p>\n<p>One account per participant<br>\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.<br>\nNo private sharing outside teams<br>\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.<br>\nbut our results were cancelled. Is there a staff member who can explain?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1371216,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2021-06-30T18:58:45.170000",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> , One doubt since the competition is going to restart does it means that the public Leaderboard as well all submissions made till now by all participants for this Competition will reset ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1364559,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-25T04:51:20.143000",
      "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> Could you also please check the How do you zip the files. As your adding More files to the train data it will become impossible to use this dataset with Colab Pro. So a well zipped dataset will surely help</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1365015,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-25T11:46:31.497000",
          "content": "<p>use GCS directly (if you use tensorflow)<br>\nyou can access data even if you use colab w/o pro</p>\n<p>but one thing clear is training time will be increase almost 1.7</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1365366,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-25T16:43:44.987000",
          "content": "<p><a href=\"https://www.kaggle.com/assign\" target=\"_blank\">@assign</a> first of all you have pay a lot for GCS. And my training times are like already one hour and increase by 1.7 will mean i have a nearly 2 hour training Time which honestly a lot </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1365631,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-26T00:59:29.067000",
          "content": "<p>I means</p>\n<ol>\n<li>make it as kaggle dataset </li>\n<li>get gcs bucket id</li>\n<li>read by tensorflow record format</li>\n</ol>\n<p>1 preprocess dataset as tfrec format</p>\n<p>2 Kaggle kernel after add the dataset</p>\n<pre><code>from kaggle_datasets import KaggleDatasets\nKaggleDatasets().get_gcs_path(&lt;target dataset name&gt;)# without ../input/\n</code></pre>\n<p><a href=\"https://www.kaggle.com/assign/get-data-info\" target=\"_blank\">example of it</a></p>\n<p>example of output:<br>\n'gs://kds-0752b66b41214b9aeeb3741bbd6b1dcf6c443007693b57cc9f4e2ea5'</p>\n<p>3 Colab</p>\n<pre><code>GCS_PATH = &lt;output of upper code&gt;\nfiles = list(np.sort(np.array(tf.io.gfile.glob(GCS_PATH + '/train*.tfrec')))) # depend on your file name\nds = tf.data.TFRecordDataset(files, num_parallel_reads = len(files))\nds = ds.map(&lt;your parsing function&gt;)\nds = ds.batch(512)\nfor I in ds: ~~~\n</code></pre>\n<p>you need to preprocess dataset as serialized tensor, <br>\nand design suitable parsing function.<br>\nI recommend to read tf.io documents</p>\n<p>I don't know whether this is google's design for using gcs freely. <br>\nBut, we can use gcs without cost</p>\n<p>well, with this approach, using TPU much easier and we can expect speed up(1 epoch: 17min -&gt; 1min)</p>\n<p>I'll delete it if it violate kaggle community guidelines</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1370957,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-30T14:12:53.690000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1370699,
      "author_name": "Mateusz Ostrowski",
      "author_url": "",
      "post_date": "2021-06-30T10:44:01.613000",
      "content": "<p>Thank you for the update 😊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1364380,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-24T22:52:59.537000",
      "content": "<p>Thank you for the update 😊</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1364280": "Hi everyone -\n\nHere's an update of where things stand on the competition reset:\n- A new test set has been generated, taking into account the fantastic feedback we've received from the community. (Thank you for all of the insights!)\n- We're in the process of scrutinizing the data, and are targeting a relaunch mid-next week (but may need more time if we find any issues along the way).\n- The new training set will include the current training _and_ test set (with the test set labels)\n- We'll adjust the competition timeline at that time to reflect the time lost with the reset\n\nAn additional bit of detail on the time stamp leakage failure point (which I did a post-mortem for):\n\nKaggle actually has a triple-redundant system for time stamp leakage, where all points failed: I [previously mentioned](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782) our best practices of random file saves and secondary checks, but our system *should* have been tolerant to any human failure points, since our backend actually re-creates time stamps for our competition files on the Data page after upload. In this instance, because our zipping process was erroring out, the \"Download All\" zip file was manually put into position, rather than being automatically created from the uploaded competition files. (Which, by the way, is the reason the timestamp leakage *didn't* show up when the data in the Notebooks environment, because it was processed normally, and only when using the downloaded zip file from the competition page, which wasn't.) \n\nIt's a good lesson on the importances of sticking to best practices. :-)",
    "1373457": "Were we currently stand: The data generation has taken a bit longer after it was determined the training set would also need to be re-generated, but it is nearly complete. We're going into a US long weekend, and I don't expect significant work over the next 4 days, but I'll pick things back up next Wednesday. With processing and checking, I'd imagine things will be ready for relaunch early the week of July 12.",
    "1371593": "Posting as a top-level comment for visibility:\n\nJust investigated whether renormalizing the current training set is able to remove the statistics leak, and unfortunately the truncation from casting into FP16 caused minute differences in the mean and standard deviations between positive and negative samples that we weren't able to remove.\n\nI will be discussing with @inversion to decide on the best way to mitigate this. We will most likely be generating a new training set without leaks.\n\nAverage of Means from On and Off (blue is needles, orange is new haystack)\n<blockquote class=\"imgur-embed-pub\" lang=\"en\" data-id=\"i1qe3pE\"  ><a href=\"//imgur.com/i1qe3pE\">Average Means from On and Off</a></blockquote><script async src=\"//s.imgur.com/min/embed.js\" charset=\"utf-8\"></script>\n\nAverage of STDs from On and Off (blue is needles, orange is new haystack)\n<blockquote class=\"imgur-embed-pub\" lang=\"en\" data-id=\"a/ZVKIYAG\"  ><a href=\"//imgur.com/a/ZVKIYAG\">Average of STDs from On and Off</a></blockquote><script async src=\"//s.imgur.com/min/embed.js\" charset=\"utf-8\"></script>",
    "1365400": "Can it be a code competition with the new data?",
    "1364351": "Good news, thanks.  Will you also provide instructions on how to remove the distribution leak from existing training an d test data? Indeed, if we train on data with the leak and predict on new test without leak then results may be irrelevant.\n\nI suggested [some way](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/248194), but there maybe other, better ways.",
    "1381061": "hello @inversion, \ncould you please provide an update on the datasets? also will the timeline for submission be extened? ",
    "1364437": "Thank you for updating information. I really appreciate your hard work.\n\nI've explored several methods while the competition was suspended.\nI'm looking forward to trying them out on new datasets😃",
    "1370452": "Can it be a code competition with the new data?",
    "1371177": "Hi, can it be a code competition with a new dataset\n",
    "1480398": "Why was our team cancelled? After communicating with my teammates, I found that we did not violate two principles:\n\nOne account per participant\nYou cannot sign up to Kaggle from multiple accounts and therefore you cannot submit from multiple accounts.\nNo private sharing outside teams\nPrivately sharing code or data outside of teams is not permitted. It's okay to share code if made available to all participants on the forums.\nbut our results were cancelled. Is there a staff member who can explain?\n",
    "1371216": "@inversion , One doubt since the competition is going to restart does it means that the public Leaderboard as well all submissions made till now by all participants for this Competition will reset ?",
    "1364559": "@inversion Could you also please check the How do you zip the files. As your adding More files to the train data it will become impossible to use this dataset with Colab Pro. So a well zipped dataset will surely help",
    "1370957": "",
    "1370699": "Thank you for the update 😊",
    "1364380": "Thank you for the update 😊"
  }
}