{
  "id": 266312,
  "title": "Top1 LB:0.965? Magic score!",
  "url": "/competitions/seti-breakthrough-listen/discussion/266312",
  "author_name": "Ctrl_CV",
  "post_date": "2021-08-18T17:09:38.272000",
  "votes": 31,
  "comment_count": 31,
  "views": 0,
  "content": "<p>I was very surprised by such scores. It seemed that the first place almost completely eliminated the distribution gap between training and testing. I was curious about which magic methods they used？</p>",
  "messages": [
    {
      "id": 1479852,
      "postDate": "2021-08-18T17:09:38.273Z",
      "content": "<p>I was very surprised by such scores. It seemed that the first place almost completely eliminated the distribution gap between training and testing. I was curious about which magic methods they used？</p>",
      "rawMarkdown": "I was very surprised by such scores. It seemed that the first place almost completely eliminated the distribution gap between training and testing. I was curious about which magic methods they used？",
      "votes": 31
    },
    {
      "id": 1480292,
      "postDate": "2021-08-18T23:58:30.807Z",
      "content": "<p>My bet is they did not use an image classification approach.  Host has published a method where they predict new few time frames and messages are surprises in this prediction.  Maybe they did something like that: reconstruct images and messages are what is not well reconstructed.  I wanted to try and did not have time or forget, hard to tell now.</p>\n<p>I really hope it is not a leak or reverse engineering of the data generator.</p>\n<p>Whichever it is, they did an awesome job.</p>",
      "rawMarkdown": "My bet is they did not use an image classification approach.  Host has published a method where they predict new few time frames and messages are surprises in this prediction.  Maybe they did something like that: reconstruct images and messages are what is not well reconstructed.  I wanted to try and did not have time or forget, hard to tell now.\n\nI really hope it is not a leak or reverse engineering of the data generator.\n\nWhichever it is, they did an awesome job.",
      "votes": 7
    },
    {
      "id": 1480036,
      "postDate": "2021-08-18T18:55:07.770Z",
      "content": "<p>Looks like real black magic.</p>",
      "rawMarkdown": "Looks like real black magic.",
      "votes": 8
    },
    {
      "id": 1480120,
      "postDate": "2021-08-18T20:11:12.563Z",
      "content": "<p>It seems like the leaderboard wasn't updated for them after the leakage fix. 😒😒</p>",
      "rawMarkdown": "It seems like the leaderboard wasn't updated for them after the leakage fix. 😒😒",
      "votes": 5,
      "replies": [
        {
          "id": 1480138,
          "postDate": "2021-08-18T20:20:26.997Z",
          "content": "<p>That's one way to explain it. 😁 </p>",
          "rawMarkdown": "That's one way to explain it. 😁 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1479884,
      "postDate": "2021-08-18T17:23:55.157Z",
      "content": "<p>Feels like external data (or leakage). Will find out soon!</p>",
      "rawMarkdown": "Feels like external data (or leakage). Will find out soon!",
      "votes": 5,
      "replies": [
        {
          "id": 1479912,
          "postDate": "2021-08-18T17:37:36.240Z",
          "content": "<p>Wait and see！！</p>",
          "rawMarkdown": "Wait and see！！",
          "votes": -1
        },
        {
          "id": 1480115,
          "postDate": "2021-08-18T20:07:45.763Z",
          "content": "<blockquote>\n  <p>Feels like external data (or leakage)</p>\n</blockquote>\n<p>I've got the same feeling, either that or they find a way to generate synth data.. <br>\nWhat I'd prefer though is to see a nice, well-tuned semi-supervised approach  </p>\n<p>ps: I haven't returned to the comp after reset, but now I can't wait to see the PVT LB and read that \"alien\" post :)</p>",
          "rawMarkdown": "> Feels like external data (or leakage)\n\nI've got the same feeling, either that or they find a way to generate synth data.. \nWhat I'd prefer though is to see a nice, well-tuned semi-supervised approach  \n\nps: I haven't returned to the comp after reset, but now I can't wait to see the PVT LB and read that \"alien\" post :)"
        },
        {
          "id": 1480629,
          "postDate": "2021-08-19T05:29:23.260Z",
          "content": "<p>Having read the first place write-up, while interesting in some respects, it does show why code competitions have more value, are more fair as private test is unseen.  And that having access to the full test set private and public opens up the probability for solutions that can exploit the test data, but would they be able to be useful in real world applications? Not sure.<br>\n e.g.<br>\n\"We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape signal -- by using a randomized signal generator. \"<br>\n\"To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from <a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a>) which adds those signals to the training set.” <br>\n\"One obvious aspect in this competition is that the background between train and test data is very different.\"</p>\n<p>Anyways this was more like a Kafka-esq competition and perfect for added lockdown eternity torture. </p>",
          "rawMarkdown": "Having read the first place write-up, while interesting in some respects, it does show why code competitions have more value, are more fair as private test is unseen.  And that having access to the full test set private and public opens up the probability for solutions that can exploit the test data, but would they be able to be useful in real world applications? Not sure.\n e.g.\n\"We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape signal -- by using a randomized signal generator. \"\n\"To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from https://github.com/bbrzycki/setigen) which adds those signals to the training set.” \n\"One obvious aspect in this competition is that the background between train and test data is very different.\"\n\nAnyways this was more like a Kafka-esq competition and perfect for added lockdown eternity torture. "
        }
      ]
    },
    {
      "id": 1480302,
      "postDate": "2021-08-19T00:13:53.150Z",
      "content": "<p>I'm absolutely amazed at the difference between 1st and 2nd place, I've never seen anything like this</p>",
      "rawMarkdown": "I'm absolutely amazed at the difference between 1st and 2nd place, I've never seen anything like this",
      "votes": 1
    },
    {
      "id": 1480251,
      "postDate": "2021-08-18T22:31:22.470Z",
      "content": "<p>I try training on the testset, and It was work.<br>\nIf we have large batchsize, we can training with uncertain labels.<br>\nHowever, there are so many noise in channel 2,4,5 and 1,3,5. <br>\nso we can't directly use single model on test set.<br>\nWe need many models ensemble. on testing dataset.</p>",
      "rawMarkdown": "I try training on the testset, and It was work.\nIf we have large batchsize, we can training with uncertain labels.\nHowever, there are so many noise in channel 2,4,5 and 1,3,5. \nso we can't directly use single model on test set.\nWe need many models ensemble. on testing dataset.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 1480276,
          "postDate": "2021-08-18T23:23:36.377Z",
          "content": "<p>thanks for sharing!!<br>\nIs that mean large batch size makes pseudo labeling stable generally?</p>",
          "rawMarkdown": "thanks for sharing!!\nIs that mean large batch size makes pseudo labeling stable generally?"
        }
      ]
    },
    {
      "id": 1479960,
      "postDate": "2021-08-18T18:03:37.020Z",
      "content": "<p>Crazy!         </p>",
      "rawMarkdown": "Crazy!         ",
      "votes": 1,
      "replies": [
        {
          "id": 1479962,
          "postDate": "2021-08-18T18:04:23Z",
          "content": "<p>Looking forward to seeing psi solutions！</p>",
          "rawMarkdown": "Looking forward to seeing psi solutions！",
          "votes": 1
        },
        {
          "id": 1480297,
          "postDate": "2021-08-19T00:08:52.650Z",
          "content": "<p>GM now, bro. </p>",
          "rawMarkdown": "GM now, bro. ",
          "votes": 1
        },
        {
          "id": 1480299,
          "postDate": "2021-08-19T00:13:15.730Z",
          "content": "<p>Congratulations! our new GM!</p>",
          "rawMarkdown": "Congratulations! our new GM!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1479934,
      "postDate": "2021-08-18T17:48:48.373Z",
      "content": "<p>If these three guys in the competition all the others are fighting for 2nd place without options :-)</p>",
      "rawMarkdown": "If these three guys in the competition all the others are fighting for 2nd place without options :-)",
      "votes": 1
    },
    {
      "id": 1480265,
      "postDate": "2021-08-18T23:01:56.733Z",
      "content": "<p>I tried pseudo label on test set and it works for me.<br>\nThe tricky thing is, it doesn't always work. I trained plenty of models, dropping some data I thought it was bad, only got two satisfied single fold(0.776 and 0.781) with somewhat luck.<br>\nThe leakage of the whole dataset is obvious.</p>",
      "rawMarkdown": "I tried pseudo label on test set and it works for me.\nThe tricky thing is, it doesn't always work. I trained plenty of models, dropping some data I thought it was bad, only got two satisfied single fold(0.776 and 0.781) with somewhat luck.\nThe leakage of the whole dataset is obvious.",
      "votes": 2
    },
    {
      "id": 1479973,
      "postDate": "2021-08-18T18:11:47.150Z",
      "content": "<p>I pray that this is not another leakage…</p>",
      "rawMarkdown": "I pray that this is not another leakage...",
      "votes": 2
    },
    {
      "id": 1479896,
      "postDate": "2021-08-18T17:32:38.050Z",
      "content": "<p>It's not funny especially if it's a leakage or reverse engineering of data creation…</p>",
      "rawMarkdown": "It's not funny especially if it's a leakage or reverse engineering of data creation...",
      "votes": 2,
      "replies": [
        {
          "id": 1479910,
          "postDate": "2021-08-18T17:37:25.503Z",
          "content": "<p>Wait and see！</p>",
          "rawMarkdown": "Wait and see！",
          "votes": -1
        },
        {
          "id": 1479947,
          "postDate": "2021-08-18T17:57:29.153Z",
          "content": "<p>Before the LB reset, I attempted to reverse engineering the data creation process based on the expected <a href=\"https://www.kaggle.com/tentotheminus9/seti-simulated-signals-data-exploration\" target=\"_blank\">six most popular</a> categories. This dataset and kernel were both shared pretty early in the comp. I had trained a model using only negative images with my numba augmentation plotter randomly creating 10% alien samples and was able to get 0.983416 CV—I don't recall LB since the submissions page just states ERROR now for all those subs. After the LB reset, no matter how I tinkered with the parameters of the generating algorithm, I was never able to get it to perform better than using the actual train images. In fact, even just adding alien technosignatures to the existing train data hurt performance. I was sure it'd be a magic bullet too.</p>\n<p>Anyhow, I've already steeled myself and expecting to get crushed in private lb.</p>",
          "rawMarkdown": "Before the LB reset, I attempted to reverse engineering the data creation process based on the expected [six most popular](https://www.kaggle.com/tentotheminus9/seti-simulated-signals-data-exploration) categories. This dataset and kernel were both shared pretty early in the comp. I had trained a model using only negative images with my numba augmentation plotter randomly creating 10% alien samples and was able to get 0.983416 CV—I don't recall LB since the submissions page just states ERROR now for all those subs. After the LB reset, no matter how I tinkered with the parameters of the generating algorithm, I was never able to get it to perform better than using the actual train images. In fact, even just adding alien technosignatures to the existing train data hurt performance. I was sure it'd be a magic bullet too.\n\nAnyhow, I've already steeled myself and expecting to get crushed in private lb.",
          "votes": 2
        },
        {
          "id": 1479955,
          "postDate": "2021-08-18T18:01:06.647Z",
          "content": "<p>Hope you get the result you want</p>",
          "rawMarkdown": "Hope you get the result you want"
        }
      ]
    },
    {
      "id": 1480216,
      "postDate": "2021-08-18T21:31:44.760Z",
      "content": "<p>Since the CV/LB split is most certainly due to more weak signals in the (public) test data, I would bet that this is either a leak or a way to reverse engineer the data creation process. No way that this is a legitimate solution that would work on real world (ok, newly generated) data.</p>",
      "rawMarkdown": "Since the CV/LB split is most certainly due to more weak signals in the (public) test data, I would bet that this is either a leak or a way to reverse engineer the data creation process. No way that this is a legitimate solution that would work on real world (ok, newly generated) data."
    },
    {
      "id": 1480201,
      "postDate": "2021-08-18T21:11:07.880Z",
      "content": "<p>they got hint from ALIEN to build new model.  I am telling you.</p>",
      "rawMarkdown": "they got hint from ALIEN to build new model.  I am telling you.",
      "votes": 1
    },
    {
      "id": 1480301,
      "postDate": "2021-08-19T00:13:40.577Z",
      "content": "<p>I am curious as well, anyways congrats to them!</p>",
      "rawMarkdown": " I am curious as well, anyways congrats to them!"
    },
    {
      "id": 1480261,
      "postDate": "2021-08-18T22:44:40.453Z",
      "content": "<p>Even though I have not participated in the competition, I am very curious to learn as well 👀</p>",
      "rawMarkdown": "Even though I have not participated in the competition, I am very curious to learn as well 👀"
    },
    {
      "id": 1480151,
      "postDate": "2021-08-18T20:27:15.443Z",
      "content": "<p>Can't even get 0.96 on CV ….<br>\nHope this is a result of good modeling and not another leakage</p>",
      "rawMarkdown": "Can't even get 0.96 on CV ....\nHope this is a result of good modeling and not another leakage"
    },
    {
      "id": 1480105,
      "postDate": "2021-08-18T19:59:00.850Z",
      "content": "<p>Waiting for the trick to be revealed soon. Hopefully not another leak. 👌 </p>",
      "rawMarkdown": "Waiting for the trick to be revealed soon. Hopefully not another leak. 👌 "
    },
    {
      "id": 1480083,
      "postDate": "2021-08-18T19:31:24.910Z",
      "content": "<p>They are aliens for sure 😅</p>",
      "rawMarkdown": "They are aliens for sure 😅"
    },
    {
      "id": 1479880,
      "postDate": "2021-08-18T17:21:31.043Z",
      "content": "<p>give me chakra magic</p>",
      "rawMarkdown": "give me chakra magic",
      "replies": [
        {
          "id": 1479913,
          "postDate": "2021-08-18T17:37:47.240Z",
          "content": "<p>Wait and see！！！</p>",
          "rawMarkdown": "Wait and see！！！"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1480292,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-18T23:58:30.807000",
      "content": "<p>My bet is they did not use an image classification approach.  Host has published a method where they predict new few time frames and messages are surprises in this prediction.  Maybe they did something like that: reconstruct images and messages are what is not well reconstructed.  I wanted to try and did not have time or forget, hard to tell now.</p>\n<p>I really hope it is not a leak or reverse engineering of the data generator.</p>\n<p>Whichever it is, they did an awesome job.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1480036,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2021-08-18T18:55:07.770000",
      "content": "<p>Looks like real black magic.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1480120,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2021-08-18T20:11:12.563000",
      "content": "<p>It seems like the leaderboard wasn't updated for them after the leakage fix. 😒😒</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1480138,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2021-08-18T20:20:26.997000",
          "content": "<p>That's one way to explain it. 😁 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1479884,
      "author_name": "Μαριος Μιχαηλιδης KazAnova",
      "author_url": "",
      "post_date": "2021-08-18T17:23:55.157000",
      "content": "<p>Feels like external data (or leakage). Will find out soon!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1479912,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-08-18T17:37:36.240000",
          "content": "<p>Wait and see！！</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1480115,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2021-08-18T20:07:45.763000",
          "content": "<blockquote>\n  <p>Feels like external data (or leakage)</p>\n</blockquote>\n<p>I've got the same feeling, either that or they find a way to generate synth data.. <br>\nWhat I'd prefer though is to see a nice, well-tuned semi-supervised approach  </p>\n<p>ps: I haven't returned to the comp after reset, but now I can't wait to see the PVT LB and read that \"alien\" post :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1480629,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2021-08-19T05:29:23.260000",
          "content": "<p>Having read the first place write-up, while interesting in some respects, it does show why code competitions have more value, are more fair as private test is unseen.  And that having access to the full test set private and public opens up the probability for solutions that can exploit the test data, but would they be able to be useful in real world applications? Not sure.<br>\n e.g.<br>\n\"We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape signal -- by using a randomized signal generator. \"<br>\n\"To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from <a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a>) which adds those signals to the training set.” <br>\n\"One obvious aspect in this competition is that the background between train and test data is very different.\"</p>\n<p>Anyways this was more like a Kafka-esq competition and perfect for added lockdown eternity torture. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480302,
      "author_name": "HinePo",
      "author_url": "",
      "post_date": "2021-08-19T00:13:53.150000",
      "content": "<p>I'm absolutely amazed at the difference between 1st and 2nd place, I've never seen anything like this</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480251,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-08-18T22:31:22.470000",
      "content": "<p>I try training on the testset, and It was work.<br>\nIf we have large batchsize, we can training with uncertain labels.<br>\nHowever, there are so many noise in channel 2,4,5 and 1,3,5. <br>\nso we can't directly use single model on test set.<br>\nWe need many models ensemble. on testing dataset.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1480276,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2021-08-18T23:23:36.377000",
          "content": "<p>thanks for sharing!!<br>\nIs that mean large batch size makes pseudo labeling stable generally?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1479960,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-08-18T18:03:37.020000",
      "content": "<p>Crazy!         </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1479962,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-08-18T18:04:23",
          "content": "<p>Looking forward to seeing psi solutions！</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1480297,
          "author_name": "Tian",
          "author_url": "",
          "post_date": "2021-08-19T00:08:52.650000",
          "content": "<p>GM now, bro. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1480299,
          "author_name": "Matthew Wu",
          "author_url": "",
          "post_date": "2021-08-19T00:13:15.730000",
          "content": "<p>Congratulations! our new GM!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1479934,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2021-08-18T17:48:48.373000",
      "content": "<p>If these three guys in the competition all the others are fighting for 2nd place without options :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480265,
      "author_name": "Last Scene",
      "author_url": "",
      "post_date": "2021-08-18T23:01:56.733000",
      "content": "<p>I tried pseudo label on test set and it works for me.<br>\nThe tricky thing is, it doesn't always work. I trained plenty of models, dropping some data I thought it was bad, only got two satisfied single fold(0.776 and 0.781) with somewhat luck.<br>\nThe leakage of the whole dataset is obvious.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1479973,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2021-08-18T18:11:47.150000",
      "content": "<p>I pray that this is not another leakage…</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1479896,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-08-18T17:32:38.050000",
      "content": "<p>It's not funny especially if it's a leakage or reverse engineering of data creation…</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1479910,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-08-18T17:37:25.503000",
          "content": "<p>Wait and see！</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1479947,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-08-18T17:57:29.153000",
          "content": "<p>Before the LB reset, I attempted to reverse engineering the data creation process based on the expected <a href=\"https://www.kaggle.com/tentotheminus9/seti-simulated-signals-data-exploration\" target=\"_blank\">six most popular</a> categories. This dataset and kernel were both shared pretty early in the comp. I had trained a model using only negative images with my numba augmentation plotter randomly creating 10% alien samples and was able to get 0.983416 CV—I don't recall LB since the submissions page just states ERROR now for all those subs. After the LB reset, no matter how I tinkered with the parameters of the generating algorithm, I was never able to get it to perform better than using the actual train images. In fact, even just adding alien technosignatures to the existing train data hurt performance. I was sure it'd be a magic bullet too.</p>\n<p>Anyhow, I've already steeled myself and expecting to get crushed in private lb.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1479955,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-08-18T18:01:06.647000",
          "content": "<p>Hope you get the result you want</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480216,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-18T21:31:44.760000",
      "content": "<p>Since the CV/LB split is most certainly due to more weak signals in the (public) test data, I would bet that this is either a leak or a way to reverse engineer the data creation process. No way that this is a legitimate solution that would work on real world (ok, newly generated) data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480201,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2021-08-18T21:11:07.880000",
      "content": "<p>they got hint from ALIEN to build new model.  I am telling you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480301,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2021-08-19T00:13:40.577000",
      "content": "<p>I am curious as well, anyways congrats to them!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480261,
      "author_name": "Sinan Calisir",
      "author_url": "",
      "post_date": "2021-08-18T22:44:40.453000",
      "content": "<p>Even though I have not participated in the competition, I am very curious to learn as well 👀</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480151,
      "author_name": "Amed",
      "author_url": "",
      "post_date": "2021-08-18T20:27:15.443000",
      "content": "<p>Can't even get 0.96 on CV ….<br>\nHope this is a result of good modeling and not another leakage</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480105,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-08-18T19:59:00.850000",
      "content": "<p>Waiting for the trick to be revealed soon. Hopefully not another leak. 👌 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480083,
      "author_name": "Parth Dhameliya",
      "author_url": "",
      "post_date": "2021-08-18T19:31:24.910000",
      "content": "<p>They are aliens for sure 😅</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1479880,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2021-08-18T17:21:31.043000",
      "content": "<p>give me chakra magic</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1479913,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-08-18T17:37:47.240000",
          "content": "<p>Wait and see！！！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1479852": "I was very surprised by such scores. It seemed that the first place almost completely eliminated the distribution gap between training and testing. I was curious about which magic methods they used？",
    "1480292": "My bet is they did not use an image classification approach.  Host has published a method where they predict new few time frames and messages are surprises in this prediction.  Maybe they did something like that: reconstruct images and messages are what is not well reconstructed.  I wanted to try and did not have time or forget, hard to tell now.\n\nI really hope it is not a leak or reverse engineering of the data generator.\n\nWhichever it is, they did an awesome job.",
    "1480036": "Looks like real black magic.",
    "1480120": "It seems like the leaderboard wasn't updated for them after the leakage fix. 😒😒",
    "1479884": "Feels like external data (or leakage). Will find out soon!",
    "1480302": "I'm absolutely amazed at the difference between 1st and 2nd place, I've never seen anything like this",
    "1480251": "I try training on the testset, and It was work.\nIf we have large batchsize, we can training with uncertain labels.\nHowever, there are so many noise in channel 2,4,5 and 1,3,5. \nso we can't directly use single model on test set.\nWe need many models ensemble. on testing dataset.\n\n",
    "1479960": "Crazy!         ",
    "1479934": "If these three guys in the competition all the others are fighting for 2nd place without options :-)",
    "1480265": "I tried pseudo label on test set and it works for me.\nThe tricky thing is, it doesn't always work. I trained plenty of models, dropping some data I thought it was bad, only got two satisfied single fold(0.776 and 0.781) with somewhat luck.\nThe leakage of the whole dataset is obvious.",
    "1479973": "I pray that this is not another leakage...",
    "1479896": "It's not funny especially if it's a leakage or reverse engineering of data creation...",
    "1480216": "Since the CV/LB split is most certainly due to more weak signals in the (public) test data, I would bet that this is either a leak or a way to reverse engineer the data creation process. No way that this is a legitimate solution that would work on real world (ok, newly generated) data.",
    "1480201": "they got hint from ALIEN to build new model.  I am telling you.",
    "1480301": " I am curious as well, anyways congrats to them!",
    "1480261": "Even though I have not participated in the competition, I am very curious to learn as well 👀",
    "1480151": "Can't even get 0.96 on CV ....\nHope this is a result of good modeling and not another leakage",
    "1480105": "Waiting for the trick to be revealed soon. Hopefully not another leak. 👌 ",
    "1480083": "They are aliens for sure 😅",
    "1479880": "give me chakra magic"
  }
}