{
  "id": 266385,
  "title": "1st Place Solution",
  "url": "/competitions/seti-breakthrough-listen/discussion/266385",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2021-08-19T00:20:44.062000",
  "votes": 277,
  "comment_count": 113,
  "views": 0,
  "content": "<p>Thanks to Kaggle and Berkeley SETI Research Center for this interesting competition. In the following, we want to give a summary of the winning solution of Team Watercooled. As always, thanks to all team members contributing equally to the solution.<br>\n<a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a></p>\n<h1>Summary</h1>\n<p>Our solution is based on large state-of-the-art classification models that were fine-tuned for this specific task. We pre-processed images by cleaning the backgrounds to boost signal to noise ratio. During training, we employed heavy augmentation in the form of Mixup. For faster training iterations, we only used the ON-channels. We only rely on provided competition data and do not utilize any external data. We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape” signal -- by using a randomized signal generator. </p>\n<h1>Cross-Validation and Preprocessing</h1>\n<p>The CV setup for this competition was quite straightforward, using a 5 fold split on the new training data. We also found that the old leaky data was still giving us a significant boost if used for training, so we added the full old train and old test data to each training fold. To reduce the impact of leakage in the training as much as possible, we cast to float32 and subsequently applied a channel normalization after loading an image. </p>\n<h1>CV/LB gap</h1>\n<p>As most participants, we experienced major differences between CV and LB scores early on in the competition. While this is in theory not necessarily always an issue, LB might just be harder (and it certainly is), we did not observe clear correlation between higher CV and higher LB. This led us to explore different ways of understanding these differences and trying to account for them. Our big gap to second place on LB appears to mostly be based on these insights and we elaborate them next.</p>\n<h1>The magic #1</h1>\n<p>As mentioned, we observed that certain models exhibit significant different LB scores even though their CV scores were similar. That led us to start investigating single predictions, where different models with different LB scores disagree. After comparing several models, we observed that better LB models had much higher probabilities for images with an “s-shape” signal as shown below. </p>\n<p><img src=\"https://i.imgur.com/9mE6ptH.png\" alt=\"s-shape in test\"></p>\n<p>That type of signal only appears in the test set and some models were able to identify the signal as atypical, while others were not. It was later also confirmed by hosts that test data contains additional type(s) of signal. To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from <a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a>) which adds those signals to the training set. We tuned the probability (p=0.01) and randomized the shape and signal to noise ratio for maximum LB score while making sure that CV score stays constant. An example of such an injected signal is shown in the plot below. </p>\n<p><img src=\"https://i.imgur.com/SCsiLeo.png\" alt=\"s-shape injected\"></p>\n<h1>The magic #2</h1>\n<p>Many noticed a large jump in our public leaderboard score a few weeks before the competition ended. We were already at a decently high score close to 0.800, which was 2nd place on LB at that time, and we were quite confident to further improve our solution as we were only subbing very simple single-fold blends and haven’t even turned our attention to further advancement like pseudo tagging which appeared to hold much promise in this competition. To further improve our solution, we continued to investigate the large CV/LB gap that was still present after injecting the s-shaped signal. </p>\n<p>One obvious aspect in this competition is that the background between train and test data is very different. This is not only imminent from visual inspection or simple binary classifiers that can easily distinguish between train and test, but also from early pseudo tagging experiments we conducted. Here, we only added very certain target=1 samples from test to train, and while CV looked good, our models suddenly only predicted target=1 for the whole test, meaning that the models perfectly learned that target=1 is always from test. </p>\n<p>So we started to speculate that the models sometimes focus too much on the background rather than the signal. In the plot below, we exemplary show the histogram of our output logits grouped by train similarity. Here, train similarity is just based on image data (mean, std, unique counts, min, max), but it was clear that the models are more certain if they “know the background”. Thus, we explored methods to reduce the impact of the background noise or reduce overfitting on the background noise.</p>\n<p><img src=\"https://i.imgur.com/V9BxbeT.png\" alt=\"train_like vs not_train_like\"></p>\n<p>Further investigations revealed that there even was significant overlap from one image to another. This means that the model could also potentially overfit to single backgrounds if it only sees target=0 or target=1 for that background. As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential. </p>\n<p>However, It is far from trivial to utilize this property due to image-wise channel normalization. One cannot just compare raw values from one image to another. To allow for comparison between images, we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns. Due to rounding errors and possible signals in the columns, we received column distances greater than 0. We utilized the overlaps to increase signal to noise ratio, by replacing the original data with the difference of the normed regions between multiple matched samples. This induces shadowing (difference is much smaller than 0) from signals that are present in the matched samples, which our models were able to distinguish from actual signals. In the plot below, the difference of the normalized overlapping region between two samples is shown. The yellow needle (values &gt; 0) originates from sample 1, and the blue needle (values &lt; 0) originates from the matched sample. </p>\n<p><img src=\"https://i.imgur.com/XxC4lDI.png\" alt=\"match shadow\"></p>\n<p>Based on this process, we attempted to clean all images where possible; see the following examples:</p>\n<p>Original sample with signal:<br>\n<img src=\"https://i.imgur.com/1y9kpMs.png\" alt=\"Original sample with signal\"><br>\nCleaned sample with signal:<br>\n<img src=\"https://i.imgur.com/CemLefI.png\" alt=\"Cleaned sample with signal\"></p>\n<p>Original sample without signal:<br>\n<img src=\"https://i.imgur.com/VSE7nxC.png\" alt=\"Original sample without signal\"><br>\nCleaned sample without signal:<br>\n<img src=\"https://i.imgur.com/STPDRyi.png\" alt=\"Cleaned sample without signal\"></p>\n<p>Originally, when we started cleaning the data, we just quickly attempted to use existing models trained on the uncleaned data and augment inference with the cleaned data. Immediately, we saw big boosts on LB giving us the 0.853 score. As this was already a significant difference to the next highest score on LB, we contacted Kaggle to elaborate on our approach and to confirm that we can continue using it. After that, we started to directly utilize the cleaned data in our modeling approach, to further boost our performance significantly.</p>\n<h1>Models</h1>\n<p>As we were very confident about our solution, we decided against submitting huge blends or usage of pseudo labels, thus rendering our solution much better suited for actual use in the field. Our best final submission is thus based on a single fit on the full data.</p>\n<p>Our model architecture and training rechime works both for uncleaned and cleaned data (i.e. nearly duplicating training data). Consequently, we decided to train the data on all data available:</p>\n<ul>\n<li>Old train (uncleaned &amp; cleaned)</li>\n<li>Old test (uncleaned &amp; cleaned)</li>\n<li>New train (uncleaned &amp; cleaned)</li>\n</ul>\n<p>For inference, we always pick the cleaned image, if available, and otherwise the uncleaned image.</p>\n<p>Early, we discovered that norm-free models exhibit slightly better CV/LB correlations, probably also due to the lack of batch norm that can be influenced by data distributional differences. Our final best model used an eca_nfnet_l2 backbone and trainable GeM pooling. </p>\n<p>Model and resolution wise we saw on CV that the bigger the better. Hence, as mentioned earlier, we feed the full image with only concatenated ON channels in and do not do any resizing. <br>\nTo further “enlarge” the images we change the model stride of the first conv layer to (1,2), which is a common trick used in computer vision competitions to basically further enlarge image resolution. That allows the model to train on a higher resolution through the backbone. Christof and Philipp used the same trick in ALASKA2 Steganalysis competition, motivated by the very similar data setup (find hidden data in an image). Training time of this “high resolution” model is about 15h on 8xV100 using DDP.</p>\n<p><img src=\"https://i.imgur.com/WdbhFTB.png\" alt=\"model\"></p>\n<p>We train the model for 20 epochs, and apart from random vertical flip, we also employ mixup. We randomly mix two images with drawing from the beta distribution with alpha=beta=5, so mostly equal blend. We always take the maximum of both labels as the final label should always be target=1 as soon as there is some form of a signal in the data. We only mix between uncleaned and cleaned images respectively. What we also found helpful, is to re-normalize the images after mixing, to keep the nature of the data intact for inference. We also employ 4xTTA (regular, hflip, vflip, hflip+vflip).</p>\n<p>Our final single fold validation score is 0.976 with a public LB score of 0.965 and private LB score of 0.964 (nearly closing the cv/lb gap). The fullfit (trained on all data with the same parameters) scores 0.967 public LB and 0.967 private LB.</p>\n<h1>Outlook</h1>\n<p>We already had a very good solution with a public LB score close to 0.800 before starting to additionally clean the backgrounds in images. We quickly realized that this has a major impact on the solution and focussed our efforts on utilizing the information to its fullest extent. Thus, bound by long training times, we did not explore every approach: In the early stage of the competition, we also trained models with a kind of self attention based on the OFF-channel activations. While not yielding better scores on its own, this proved to be valuable when used in blends as it seems to add diversity. One other thing we noticed is that pseudo labeling has a positive impact as it may help with the train/test domain shift. We also explored additional techniques and research to tackle domain shift.</p>",
  "messages": [
    {
      "id": 1480304,
      "postDate": "2021-08-19T00:20:44.063Z",
      "content": "<p>Thanks to Kaggle and Berkeley SETI Research Center for this interesting competition. In the following, we want to give a summary of the winning solution of Team Watercooled. As always, thanks to all team members contributing equally to the solution.<br>\n<a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a></p>\n<h1>Summary</h1>\n<p>Our solution is based on large state-of-the-art classification models that were fine-tuned for this specific task. We pre-processed images by cleaning the backgrounds to boost signal to noise ratio. During training, we employed heavy augmentation in the form of Mixup. For faster training iterations, we only used the ON-channels. We only rely on provided competition data and do not utilize any external data. We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape” signal -- by using a randomized signal generator. </p>\n<h1>Cross-Validation and Preprocessing</h1>\n<p>The CV setup for this competition was quite straightforward, using a 5 fold split on the new training data. We also found that the old leaky data was still giving us a significant boost if used for training, so we added the full old train and old test data to each training fold. To reduce the impact of leakage in the training as much as possible, we cast to float32 and subsequently applied a channel normalization after loading an image. </p>\n<h1>CV/LB gap</h1>\n<p>As most participants, we experienced major differences between CV and LB scores early on in the competition. While this is in theory not necessarily always an issue, LB might just be harder (and it certainly is), we did not observe clear correlation between higher CV and higher LB. This led us to explore different ways of understanding these differences and trying to account for them. Our big gap to second place on LB appears to mostly be based on these insights and we elaborate them next.</p>\n<h1>The magic #1</h1>\n<p>As mentioned, we observed that certain models exhibit significant different LB scores even though their CV scores were similar. That led us to start investigating single predictions, where different models with different LB scores disagree. After comparing several models, we observed that better LB models had much higher probabilities for images with an “s-shape” signal as shown below. </p>\n<p><img src=\"https://i.imgur.com/9mE6ptH.png\" alt=\"s-shape in test\"></p>\n<p>That type of signal only appears in the test set and some models were able to identify the signal as atypical, while others were not. It was later also confirmed by hosts that test data contains additional type(s) of signal. To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from <a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a>) which adds those signals to the training set. We tuned the probability (p=0.01) and randomized the shape and signal to noise ratio for maximum LB score while making sure that CV score stays constant. An example of such an injected signal is shown in the plot below. </p>\n<p><img src=\"https://i.imgur.com/SCsiLeo.png\" alt=\"s-shape injected\"></p>\n<h1>The magic #2</h1>\n<p>Many noticed a large jump in our public leaderboard score a few weeks before the competition ended. We were already at a decently high score close to 0.800, which was 2nd place on LB at that time, and we were quite confident to further improve our solution as we were only subbing very simple single-fold blends and haven’t even turned our attention to further advancement like pseudo tagging which appeared to hold much promise in this competition. To further improve our solution, we continued to investigate the large CV/LB gap that was still present after injecting the s-shaped signal. </p>\n<p>One obvious aspect in this competition is that the background between train and test data is very different. This is not only imminent from visual inspection or simple binary classifiers that can easily distinguish between train and test, but also from early pseudo tagging experiments we conducted. Here, we only added very certain target=1 samples from test to train, and while CV looked good, our models suddenly only predicted target=1 for the whole test, meaning that the models perfectly learned that target=1 is always from test. </p>\n<p>So we started to speculate that the models sometimes focus too much on the background rather than the signal. In the plot below, we exemplary show the histogram of our output logits grouped by train similarity. Here, train similarity is just based on image data (mean, std, unique counts, min, max), but it was clear that the models are more certain if they “know the background”. Thus, we explored methods to reduce the impact of the background noise or reduce overfitting on the background noise.</p>\n<p><img src=\"https://i.imgur.com/V9BxbeT.png\" alt=\"train_like vs not_train_like\"></p>\n<p>Further investigations revealed that there even was significant overlap from one image to another. This means that the model could also potentially overfit to single backgrounds if it only sees target=0 or target=1 for that background. As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential. </p>\n<p>However, It is far from trivial to utilize this property due to image-wise channel normalization. One cannot just compare raw values from one image to another. To allow for comparison between images, we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns. Due to rounding errors and possible signals in the columns, we received column distances greater than 0. We utilized the overlaps to increase signal to noise ratio, by replacing the original data with the difference of the normed regions between multiple matched samples. This induces shadowing (difference is much smaller than 0) from signals that are present in the matched samples, which our models were able to distinguish from actual signals. In the plot below, the difference of the normalized overlapping region between two samples is shown. The yellow needle (values &gt; 0) originates from sample 1, and the blue needle (values &lt; 0) originates from the matched sample. </p>\n<p><img src=\"https://i.imgur.com/XxC4lDI.png\" alt=\"match shadow\"></p>\n<p>Based on this process, we attempted to clean all images where possible; see the following examples:</p>\n<p>Original sample with signal:<br>\n<img src=\"https://i.imgur.com/1y9kpMs.png\" alt=\"Original sample with signal\"><br>\nCleaned sample with signal:<br>\n<img src=\"https://i.imgur.com/CemLefI.png\" alt=\"Cleaned sample with signal\"></p>\n<p>Original sample without signal:<br>\n<img src=\"https://i.imgur.com/VSE7nxC.png\" alt=\"Original sample without signal\"><br>\nCleaned sample without signal:<br>\n<img src=\"https://i.imgur.com/STPDRyi.png\" alt=\"Cleaned sample without signal\"></p>\n<p>Originally, when we started cleaning the data, we just quickly attempted to use existing models trained on the uncleaned data and augment inference with the cleaned data. Immediately, we saw big boosts on LB giving us the 0.853 score. As this was already a significant difference to the next highest score on LB, we contacted Kaggle to elaborate on our approach and to confirm that we can continue using it. After that, we started to directly utilize the cleaned data in our modeling approach, to further boost our performance significantly.</p>\n<h1>Models</h1>\n<p>As we were very confident about our solution, we decided against submitting huge blends or usage of pseudo labels, thus rendering our solution much better suited for actual use in the field. Our best final submission is thus based on a single fit on the full data.</p>\n<p>Our model architecture and training rechime works both for uncleaned and cleaned data (i.e. nearly duplicating training data). Consequently, we decided to train the data on all data available:</p>\n<ul>\n<li>Old train (uncleaned &amp; cleaned)</li>\n<li>Old test (uncleaned &amp; cleaned)</li>\n<li>New train (uncleaned &amp; cleaned)</li>\n</ul>\n<p>For inference, we always pick the cleaned image, if available, and otherwise the uncleaned image.</p>\n<p>Early, we discovered that norm-free models exhibit slightly better CV/LB correlations, probably also due to the lack of batch norm that can be influenced by data distributional differences. Our final best model used an eca_nfnet_l2 backbone and trainable GeM pooling. </p>\n<p>Model and resolution wise we saw on CV that the bigger the better. Hence, as mentioned earlier, we feed the full image with only concatenated ON channels in and do not do any resizing. <br>\nTo further “enlarge” the images we change the model stride of the first conv layer to (1,2), which is a common trick used in computer vision competitions to basically further enlarge image resolution. That allows the model to train on a higher resolution through the backbone. Christof and Philipp used the same trick in ALASKA2 Steganalysis competition, motivated by the very similar data setup (find hidden data in an image). Training time of this “high resolution” model is about 15h on 8xV100 using DDP.</p>\n<p><img src=\"https://i.imgur.com/WdbhFTB.png\" alt=\"model\"></p>\n<p>We train the model for 20 epochs, and apart from random vertical flip, we also employ mixup. We randomly mix two images with drawing from the beta distribution with alpha=beta=5, so mostly equal blend. We always take the maximum of both labels as the final label should always be target=1 as soon as there is some form of a signal in the data. We only mix between uncleaned and cleaned images respectively. What we also found helpful, is to re-normalize the images after mixing, to keep the nature of the data intact for inference. We also employ 4xTTA (regular, hflip, vflip, hflip+vflip).</p>\n<p>Our final single fold validation score is 0.976 with a public LB score of 0.965 and private LB score of 0.964 (nearly closing the cv/lb gap). The fullfit (trained on all data with the same parameters) scores 0.967 public LB and 0.967 private LB.</p>\n<h1>Outlook</h1>\n<p>We already had a very good solution with a public LB score close to 0.800 before starting to additionally clean the backgrounds in images. We quickly realized that this has a major impact on the solution and focussed our efforts on utilizing the information to its fullest extent. Thus, bound by long training times, we did not explore every approach: In the early stage of the competition, we also trained models with a kind of self attention based on the OFF-channel activations. While not yielding better scores on its own, this proved to be valuable when used in blends as it seems to add diversity. One other thing we noticed is that pseudo labeling has a positive impact as it may help with the train/test domain shift. We also explored additional techniques and research to tackle domain shift.</p>",
      "rawMarkdown": "Thanks to Kaggle and Berkeley SETI Research Center for this interesting competition. In the following, we want to give a summary of the winning solution of Team Watercooled. As always, thanks to all team members contributing equally to the solution.\n@philippsinger @christofhenkel @ilu000\n\n# Summary\nOur solution is based on large state-of-the-art classification models that were fine-tuned for this specific task. We pre-processed images by cleaning the backgrounds to boost signal to noise ratio. During training, we employed heavy augmentation in the form of Mixup. For faster training iterations, we only used the ON-channels. We only rely on provided competition data and do not utilize any external data. We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape” signal -- by using a randomized signal generator. \n\n# Cross-Validation and Preprocessing\nThe CV setup for this competition was quite straightforward, using a 5 fold split on the new training data. We also found that the old leaky data was still giving us a significant boost if used for training, so we added the full old train and old test data to each training fold. To reduce the impact of leakage in the training as much as possible, we cast to float32 and subsequently applied a channel normalization after loading an image. \n\n# CV/LB gap\nAs most participants, we experienced major differences between CV and LB scores early on in the competition. While this is in theory not necessarily always an issue, LB might just be harder (and it certainly is), we did not observe clear correlation between higher CV and higher LB. This led us to explore different ways of understanding these differences and trying to account for them. Our big gap to second place on LB appears to mostly be based on these insights and we elaborate them next.\n\n# The magic #1\nAs mentioned, we observed that certain models exhibit significant different LB scores even though their CV scores were similar. That led us to start investigating single predictions, where different models with different LB scores disagree. After comparing several models, we observed that better LB models had much higher probabilities for images with an “s-shape” signal as shown below. \n\n![s-shape in test](https://i.imgur.com/9mE6ptH.png)\n\nThat type of signal only appears in the test set and some models were able to identify the signal as atypical, while others were not. It was later also confirmed by hosts that test data contains additional type(s) of signal. To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from https://github.com/bbrzycki/setigen) which adds those signals to the training set. We tuned the probability (p=0.01) and randomized the shape and signal to noise ratio for maximum LB score while making sure that CV score stays constant. An example of such an injected signal is shown in the plot below. \n\n![s-shape injected](https://i.imgur.com/SCsiLeo.png)\n\n# The magic #2\nMany noticed a large jump in our public leaderboard score a few weeks before the competition ended. We were already at a decently high score close to 0.800, which was 2nd place on LB at that time, and we were quite confident to further improve our solution as we were only subbing very simple single-fold blends and haven’t even turned our attention to further advancement like pseudo tagging which appeared to hold much promise in this competition. To further improve our solution, we continued to investigate the large CV/LB gap that was still present after injecting the s-shaped signal. \n\nOne obvious aspect in this competition is that the background between train and test data is very different. This is not only imminent from visual inspection or simple binary classifiers that can easily distinguish between train and test, but also from early pseudo tagging experiments we conducted. Here, we only added very certain target=1 samples from test to train, and while CV looked good, our models suddenly only predicted target=1 for the whole test, meaning that the models perfectly learned that target=1 is always from test. \n\nSo we started to speculate that the models sometimes focus too much on the background rather than the signal. In the plot below, we exemplary show the histogram of our output logits grouped by train similarity. Here, train similarity is just based on image data (mean, std, unique counts, min, max), but it was clear that the models are more certain if they “know the background”. Thus, we explored methods to reduce the impact of the background noise or reduce overfitting on the background noise.\n\n![train_like vs not_train_like](https://i.imgur.com/V9BxbeT.png)\n\nFurther investigations revealed that there even was significant overlap from one image to another. This means that the model could also potentially overfit to single backgrounds if it only sees target=0 or target=1 for that background. As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential. \n\nHowever, It is far from trivial to utilize this property due to image-wise channel normalization. One cannot just compare raw values from one image to another. To allow for comparison between images, we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns. Due to rounding errors and possible signals in the columns, we received column distances greater than 0. We utilized the overlaps to increase signal to noise ratio, by replacing the original data with the difference of the normed regions between multiple matched samples. This induces shadowing (difference is much smaller than 0) from signals that are present in the matched samples, which our models were able to distinguish from actual signals. In the plot below, the difference of the normalized overlapping region between two samples is shown. The yellow needle (values > 0) originates from sample 1, and the blue needle (values < 0) originates from the matched sample. \n\n![match shadow](https://i.imgur.com/XxC4lDI.png)\n\nBased on this process, we attempted to clean all images where possible; see the following examples:\n\nOriginal sample with signal:\n![Original sample with signal](https://i.imgur.com/1y9kpMs.png)\nCleaned sample with signal:\n![Cleaned sample with signal](https://i.imgur.com/CemLefI.png)\n\nOriginal sample without signal:\n![Original sample without signal](https://i.imgur.com/VSE7nxC.png)\nCleaned sample without signal:\n![Cleaned sample without signal](https://i.imgur.com/STPDRyi.png)\n\nOriginally, when we started cleaning the data, we just quickly attempted to use existing models trained on the uncleaned data and augment inference with the cleaned data. Immediately, we saw big boosts on LB giving us the 0.853 score. As this was already a significant difference to the next highest score on LB, we contacted Kaggle to elaborate on our approach and to confirm that we can continue using it. After that, we started to directly utilize the cleaned data in our modeling approach, to further boost our performance significantly.\n\n# Models\nAs we were very confident about our solution, we decided against submitting huge blends or usage of pseudo labels, thus rendering our solution much better suited for actual use in the field. Our best final submission is thus based on a single fit on the full data.\n\nOur model architecture and training rechime works both for uncleaned and cleaned data (i.e. nearly duplicating training data). Consequently, we decided to train the data on all data available:\n- Old train (uncleaned & cleaned)\n- Old test (uncleaned & cleaned)\n- New train (uncleaned & cleaned)\n\nFor inference, we always pick the cleaned image, if available, and otherwise the uncleaned image.\n\nEarly, we discovered that norm-free models exhibit slightly better CV/LB correlations, probably also due to the lack of batch norm that can be influenced by data distributional differences. Our final best model used an eca_nfnet_l2 backbone and trainable GeM pooling. \n\nModel and resolution wise we saw on CV that the bigger the better. Hence, as mentioned earlier, we feed the full image with only concatenated ON channels in and do not do any resizing. \nTo further “enlarge” the images we change the model stride of the first conv layer to (1,2), which is a common trick used in computer vision competitions to basically further enlarge image resolution. That allows the model to train on a higher resolution through the backbone. Christof and Philipp used the same trick in ALASKA2 Steganalysis competition, motivated by the very similar data setup (find hidden data in an image). Training time of this “high resolution” model is about 15h on 8xV100 using DDP.\n\n![model](https://i.imgur.com/WdbhFTB.png)\n\nWe train the model for 20 epochs, and apart from random vertical flip, we also employ mixup. We randomly mix two images with drawing from the beta distribution with alpha=beta=5, so mostly equal blend. We always take the maximum of both labels as the final label should always be target=1 as soon as there is some form of a signal in the data. We only mix between uncleaned and cleaned images respectively. What we also found helpful, is to re-normalize the images after mixing, to keep the nature of the data intact for inference. We also employ 4xTTA (regular, hflip, vflip, hflip+vflip).\n\nOur final single fold validation score is 0.976 with a public LB score of 0.965 and private LB score of 0.964 (nearly closing the cv/lb gap). The fullfit (trained on all data with the same parameters) scores 0.967 public LB and 0.967 private LB.\n\n# Outlook\nWe already had a very good solution with a public LB score close to 0.800 before starting to additionally clean the backgrounds in images. We quickly realized that this has a major impact on the solution and focussed our efforts on utilizing the information to its fullest extent. Thus, bound by long training times, we did not explore every approach: In the early stage of the competition, we also trained models with a kind of self attention based on the OFF-channel activations. While not yielding better scores on its own, this proved to be valuable when used in blends as it seems to add diversity. One other thing we noticed is that pseudo labeling has a positive impact as it may help with the train/test domain shift. We also explored additional techniques and research to tackle domain shift.\n",
      "votes": 277
    },
    {
      "id": 1481215,
      "postDate": "2021-08-19T11:24:40.047Z",
      "content": "<p>Just one more note on magic #2 for further elaboration. </p>\n<p>Imagine the following simplistic use case:<br>\nYou have a set of images with certain backgrounds (white, green, striped, constant, etc.) and then you have black dots on those images and you want your model to identify if an image contains a black spot (anomaly). Unfortunately, you have some backgrounds that always exhibit a black spot, or maybe never exhibit one. If you are unlucky, your model might decide to overfit on the actual background, instead of trying to identify the raw signal of the black spot. And here in this competition we observed quite obvious overfit on the background, and a challenge was to steer the models to learn the actual patterns. </p>\n<p>One thing most people did was to do mixup, so mixing different backgrounds together and having the signals in different intensity. One other thing was to utilize pseudo tagging, that then also better learns to adjust to the new background (distribution) of the test data. But the best solution is to process the images in a way that they increase the signal to noise ratio, which we did by cleaning common backgrounds among images. In the example above, it would be obviously better to completely remove all background, and just keep the actual signal.</p>\n<p>I do personally not see how this is necessarily only limited to this data and problem. Of course, the data is synthetic here, so the effects are emphasized, but I have observed similar effects before, in other problems, competitions, and projects, but have never so thoroughly thought about it like here. There are many use cases where this is an imminent issue, such as fault prediction in machinery, anomaly prediction, steganalysis, or even chest x-ray prediction where often models prefer to overfit on certain device characteristics.</p>\n<p>Will the same technique applied here work for all problems? Probably not, but the thought process can be the same. So I would encourage you to see it as a learning, at least that's what I am doing, and try to approach these issues in the future in different projects. I can imagine, that clever modeling approaches can also attempt to deal with these issues directly, for example via attention like 3rd place used. There is also plenty of research in the area of domain adaption, where I have seen little being applied on Kaggle, and I have only started to look into it:<br>\n<a href=\"https://paperswithcode.com/task/domain-adaptation\" target=\"_blank\">https://paperswithcode.com/task/domain-adaptation</a></p>",
      "rawMarkdown": "Just one more note on magic #2 for further elaboration. \n\nImagine the following simplistic use case:\nYou have a set of images with certain backgrounds (white, green, striped, constant, etc.) and then you have black dots on those images and you want your model to identify if an image contains a black spot (anomaly). Unfortunately, you have some backgrounds that always exhibit a black spot, or maybe never exhibit one. If you are unlucky, your model might decide to overfit on the actual background, instead of trying to identify the raw signal of the black spot. And here in this competition we observed quite obvious overfit on the background, and a challenge was to steer the models to learn the actual patterns. \n\nOne thing most people did was to do mixup, so mixing different backgrounds together and having the signals in different intensity. One other thing was to utilize pseudo tagging, that then also better learns to adjust to the new background (distribution) of the test data. But the best solution is to process the images in a way that they increase the signal to noise ratio, which we did by cleaning common backgrounds among images. In the example above, it would be obviously better to completely remove all background, and just keep the actual signal.\n\nI do personally not see how this is necessarily only limited to this data and problem. Of course, the data is synthetic here, so the effects are emphasized, but I have observed similar effects before, in other problems, competitions, and projects, but have never so thoroughly thought about it like here. There are many use cases where this is an imminent issue, such as fault prediction in machinery, anomaly prediction, steganalysis, or even chest x-ray prediction where often models prefer to overfit on certain device characteristics.\n\nWill the same technique applied here work for all problems? Probably not, but the thought process can be the same. So I would encourage you to see it as a learning, at least that's what I am doing, and try to approach these issues in the future in different projects. I can imagine, that clever modeling approaches can also attempt to deal with these issues directly, for example via attention like 3rd place used. There is also plenty of research in the area of domain adaption, where I have seen little being applied on Kaggle, and I have only started to look into it:\nhttps://paperswithcode.com/task/domain-adaptation\n",
      "votes": 24,
      "replies": [
        {
          "id": 1481268,
          "postDate": "2021-08-19T11:54:58.773Z",
          "content": "<p>I agree with most of this except on one thing.  You say data is synthetic.  The messages are synthetic for sure, but the background does not seem to be synthetic according to the data description page:</p>\n<blockquote>\n  <p>Breakthrough Listen generates similar spectrograms to the one shown above, but typically spanning several GHz of the radio spectrum (rather than the approx. 2 MHz shown above). The data are stored either as filterbank format or HDF5 format files, but essentially are arrays of intensity as a function of frequency and time, accompanied by headers containing metadata such as the direction the telescope was pointed in, the frequency scale, and so on. We generate over 1 PB of spectrograms per year; individual filterbank files can be tens of GB in size. For the purposes of the Kaggle challenge, we have discarded the majority of the metadata and are simply presenting numpy arrays consisting of small regions of the spectrograms that we refer to as “snippets”.</p>\n</blockquote>",
          "rawMarkdown": "I agree with most of this except on one thing.  You say data is synthetic.  The messages are synthetic for sure, but the background does not seem to be synthetic according to the data description page:\n\n> Breakthrough Listen generates similar spectrograms to the one shown above, but typically spanning several GHz of the radio spectrum (rather than the approx. 2 MHz shown above). The data are stored either as filterbank format or HDF5 format files, but essentially are arrays of intensity as a function of frequency and time, accompanied by headers containing metadata such as the direction the telescope was pointed in, the frequency scale, and so on. We generate over 1 PB of spectrograms per year; individual filterbank files can be tens of GB in size. For the purposes of the Kaggle challenge, we have discarded the majority of the metadata and are simply presenting numpy arrays consisting of small regions of the spectrograms that we refer to as “snippets”.",
          "votes": 3
        },
        {
          "id": 1481346,
          "postDate": "2021-08-19T12:58:09.653Z",
          "content": "<p>First - congratulations on your incredible result! Since data preparation/cleaning is also my main interest, I would really like to understand your magic #2 in detail. Based on the infomation in this thread, I haven't yet managed to. Will you add some explanations and/or code later on? Thanks!</p>",
          "rawMarkdown": "First - congratulations on your incredible result! Since data preparation/cleaning is also my main interest, I would really like to understand your magic #2 in detail. Based on the infomation in this thread, I haven't yet managed to. Will you add some explanations and/or code later on? Thanks!",
          "votes": 1
        },
        {
          "id": 1481504,
          "postDate": "2021-08-19T14:49:39.303Z",
          "content": "<p>I am so grateful that you figured this out and shared it. <br>\nImagine a possible and boring parallel universe where no one in the competition found these magics!</p>",
          "rawMarkdown": "I am so grateful that you figured this out and shared it. \nImagine a possible and boring parallel universe where no one in the competition found these magics!",
          "votes": 4
        },
        {
          "id": 1486177,
          "postDate": "2021-08-22T17:59:06.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> could u through some light on cleaning methods u used to improve SNR ..<br>\n<code>\" As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential.\"</code><br>\nthanks in advance..</p>",
          "rawMarkdown": "@philippsinger  @christofhenkel could u through some light on cleaning methods u used to improve SNR ..\n`\" As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential.\"`\nthanks in advance.."
        }
      ]
    },
    {
      "id": 1480415,
      "postDate": "2021-08-19T02:45:02.700Z",
      "content": "<p>Huge congratulations! I've never seen a win with such a big lead! </p>\n<p>It never occurred to me that the same background image, by itself and the image into which it was injected with a signal, would be in the training set at the same time.</p>\n<p>Such a dataset does not occur in industry or real world, so this is completely beyond my imagination.</p>\n<p>I feel shame on sharing our solution this time so I'll just skip it. 😅</p>",
      "rawMarkdown": "Huge congratulations! I've never seen a win with such a big lead! \n\nIt never occurred to me that the same background image, by itself and the image into which it was injected with a signal, would be in the training set at the same time.\n\nSuch a dataset does not occur in industry or real world, so this is completely beyond my imagination.\n\nI feel shame on sharing our solution this time so I'll just skip it. 😅",
      "votes": 17,
      "replies": [
        {
          "id": 1481082,
          "postDate": "2021-08-19T09:51:52.270Z",
          "content": "<blockquote>\n  <p>Such a dataset does not occur in industry or real world, so this is completely beyond my imagination.</p>\n</blockquote>\n<p>The organizer like it though - so you may have to rethink about this!</p>\n<p>It seems like they wanted you to try and solve that little puzzle in there</p>",
          "rawMarkdown": "> \nSuch a dataset does not occur in industry or real world, so this is completely beyond my imagination.\n\n\n\n\n\n\nThe organizer like it though - so you may have to rethink about this!\n\nIt seems like they wanted you to try and solve that little puzzle in there\n\n",
          "votes": 6
        },
        {
          "id": 1481099,
          "postDate": "2021-08-19T10:08:14.467Z",
          "content": "<p>E.g . The organizer encouraged you to find the leaks they injected in the data - This has gone downhill from the time that we were wondering whether solutions with leaks should be viable to encouraging people using them .</p>\n<p><strong>Jut to be clear - the winning team ,did the right thing raising that and reporting it</strong></p>",
          "rawMarkdown": "E.g . The organizer encouraged you to find the leaks they injected in the data - This has gone downhill from the time that we were wondering whether solutions with leaks should be viable to encouraging people using them .\n\n**Jut to be clear - the winning team ,did the right thing raising that and reporting it**",
          "votes": 4
        },
        {
          "id": 1481101,
          "postDate": "2021-08-19T10:10:30.257Z",
          "content": "<p>Poor us, we were trying to build something that could work in the real world (at least to some extend)</p>",
          "rawMarkdown": "Poor us, we were trying to build something that could work in the real world (at least to some extend)",
          "votes": 8
        },
        {
          "id": 1481748,
          "postDate": "2021-08-19T16:46:38.260Z",
          "content": "<blockquote>\n  <p>The organizer like it though - so you may have to rethink about this!<br>\n  It seems like they wanted you to try and solve that little puzzle in there</p>\n</blockquote>\n<p>I'm a MLE. Personally, I came here mainly to analyse other datasets (kaggling and reading others post) to find something that works for the datasets I work on in my company.</p>\n<p>Thanks for that so far I have found many methods or tricks to improve the performance of my company's models ;)</p>\n<blockquote>\n  <p>Jut to be clear - the winning team ,did the right thing raising that and reporting it</p>\n</blockquote>\n<p>I agree. So I don't actually feel unhappy about the competition, whether the puzzle was deliberately left in place by the organisers or was an oversight.</p>\n<p>Even if I knew about the puzzle beforehand, it shouldn't have affected my main objective - to try out various methods and tricks in an effort to make it work for both the competition and my own work.</p>\n<p>This is my way of kaggling, but I'm not against other ways of kaggling, as long as they don't break the rules.</p>",
          "rawMarkdown": "> The organizer like it though - so you may have to rethink about this!\n> It seems like they wanted you to try and solve that little puzzle in there\n\nI'm a MLE. Personally, I came here mainly to analyse other datasets (kaggling and reading others post) to find something that works for the datasets I work on in my company.\n\nThanks for that so far I have found many methods or tricks to improve the performance of my company's models ;)\n\n> Jut to be clear - the winning team ,did the right thing raising that and reporting it\n\nI agree. So I don't actually feel unhappy about the competition, whether the puzzle was deliberately left in place by the organisers or was an oversight.\n\nEven if I knew about the puzzle beforehand, it shouldn't have affected my main objective - to try out various methods and tricks in an effort to make it work for both the competition and my own work.\n\nThis is my way of kaggling, but I'm not against other ways of kaggling, as long as they don't break the rules.\n\n",
          "votes": 3
        },
        {
          "id": 1481943,
          "postDate": "2021-08-19T18:42:18.343Z",
          "content": "<p>Fair enough. I guess I was hoping to have a leak-free competition  the second time after what happened in the first, but that was wishful thinking for my part. </p>",
          "rawMarkdown": "Fair enough. I guess I was hoping to have a leak-free competition  the second time after what happened in the first, but that was wishful thinking for my part. ",
          "votes": 2
        },
        {
          "id": 1482381,
          "postDate": "2021-08-20T03:55:24.430Z",
          "content": "<p>Haha, I still believed it was leek-free when the LB was at 0.85, but eventually found out that was wrong.</p>\n<p>But the point I have to worry about is that while this is an obvious leek in the MLE's view, it's not easy to convince people who don't know much about machine learning ;)</p>\n<p>This is not meant to be negative, but the fact is that there are a number of competition organisers whose directors don't know much about machine learning.</p>",
          "rawMarkdown": "Haha, I still believed it was leek-free when the LB was at 0.85, but eventually found out that was wrong.\n\nBut the point I have to worry about is that while this is an obvious leek in the MLE's view, it's not easy to convince people who don't know much about machine learning ;)\n\nThis is not meant to be negative, but the fact is that there are a number of competition organisers whose directors don't know much about machine learning.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1480321,
      "postDate": "2021-08-19T00:48:46.370Z",
      "content": "<p>Congratulations, we had found the magic#1 but not the magic#2 even if we've spent some time on pre-processing/cleaning.</p>",
      "rawMarkdown": "Congratulations, we had found the magic#1 but not the magic#2 even if we've spent some time on pre-processing/cleaning.",
      "votes": 14,
      "replies": [
        {
          "id": 1480376,
          "postDate": "2021-08-19T02:06:06.373Z",
          "content": "<p>Thank you. Congratulation to becoming a grandmaster <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> . Huge achievement.</p>",
          "rawMarkdown": "Thank you. Congratulation to becoming a grandmaster @mpware . Huge achievement.",
          "votes": 5
        },
        {
          "id": 1480652,
          "postDate": "2021-08-19T05:53:48.697Z",
          "content": "<p>Huge Congratulations <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> on becoming grandmaster and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> &amp; team on the win , we also found magic1 but it did not elate our score by a large margin mostly because we were not inducing signal properly in train like <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> does , we were not looking into anything else in last few days as we were totally biased that magic 1 was reason for the huge boost as said by the hosts , but this solution is absolutely mind blowing </p>",
          "rawMarkdown": "Huge Congratulations @mpware on becoming grandmaster and @ilu000 & team on the win , we also found magic1 but it did not elate our score by a large margin mostly because we were not inducing signal properly in train like @ilu000 does , we were not looking into anything else in last few days as we were totally biased that magic 1 was reason for the huge boost as said by the hosts , but this solution is absolutely mind blowing ",
          "votes": 2
        },
        {
          "id": 1480805,
          "postDate": "2021-08-19T07:22:15.330Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> ! We were also exited when we discovered magic#1 because we expected it would close the gap but after a few submissions we realized that it will not and something more was needed. Writeup of our solution/findings is coming …</p>",
          "rawMarkdown": "Thanks @christofhenkel, @tanulsingh077 ! We were also exited when we discovered magic#1 because we expected it would close the gap but after a few submissions we realized that it will not and something more was needed. Writeup of our solution/findings is coming ...",
          "votes": 2
        }
      ]
    },
    {
      "id": 1480422,
      "postDate": "2021-08-19T02:52:15.873Z",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <br>\nCongratulations &amp; thanks for sharing the solution.</p>\n<p>I'm not sure of the process of magic #2.</p>\n<blockquote>\n  <p>we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns.</p>\n</blockquote>\n<p>Does it mean the process below?</p>\n<ol>\n<li>normalize column vector for each images</li>\n<li>calculate pixel difference from the first column per each image</li>\n<li>take the mean of 2 (column-wise), which give us the one feature vector</li>\n<li>using feature vector given on 3, match the identical background image in the all image space (using kNN etc.)</li>\n<li>take pixel difference from the matched image, and clear the background</li>\n</ol>\n<p>Example: process for 3x3, 1-channel image</p>\n<pre><code># original image\nimg1 == [\n  [1, 4, 3],\n  [2, 3, 3],\n  [3, 2, 3],\n]\n\n# after column-wise normalization\nimg1 == [\n  [-1, 1, 0],\n  [0, 0, 0],\n  [1, -1, 0],\n]\n\n# after taking column-wise difference from the first column\nimg1 == [\n  [0, 2, 1],\n  [0, 0, 0],\n  [0, -2, -1],\n]\n\n# after calculating mean pixel difference\nimg1 == [\n  [1],\n  [0],\n  [-1],\n]\n</code></pre>",
      "rawMarkdown": "@ilu000 \nCongratulations & thanks for sharing the solution.\n\nI'm not sure of the process of magic #2.\n\n> we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns.\n\nDoes it mean the process below?\n1. normalize column vector for each images\n2. calculate pixel difference from the first column per each image\n3. take the mean of 2 (column-wise), which give us the one feature vector\n4. using feature vector given on 3, match the identical background image in the all image space (using kNN etc.)\n5. take pixel difference from the matched image, and clear the background\n\nExample: process for 3x3, 1-channel image\n```\n# original image\nimg1 == [\n  [1, 4, 3],\n  [2, 3, 3],\n  [3, 2, 3],\n]\n\n# after column-wise normalization\nimg1 == [\n  [-1, 1, 0],\n  [0, 0, 0],\n  [1, -1, 0],\n]\n\n# after taking column-wise difference from the first column\nimg1 == [\n  [0, 2, 1],\n  [0, 0, 0],\n  [0, -2, -1],\n]\n\n# after calculating mean pixel difference\nimg1 == [\n  [1],\n  [0],\n  [-1],\n]\n```",
      "votes": 7,
      "replies": [
        {
          "id": 1480937,
          "postDate": "2021-08-19T08:32:08.990Z",
          "content": "<p>I've read the explanation of magic #2 by the winning team multiple times and I still don't get it at all.. I understand your process on a technical level - but how does this help to remove the background? That's what I don't get.</p>\n<p>Is it basically a leak because images appear multiple times in the train set or what? I'm hoping for some code later on.</p>",
          "rawMarkdown": "I've read the explanation of magic #2 by the winning team multiple times and I still don't get it at all.. I understand your process on a technical level - but how does this help to remove the background? That's what I don't get.\n\nIs it basically a leak because images appear multiple times in the train set or what? I'm hoping for some code later on.",
          "votes": 1
        },
        {
          "id": 1481462,
          "postDate": "2021-08-19T14:15:03.770Z",
          "content": "<p>I believe it works as follows. Let's consider only column 1. For every image (actually i mean image channel 276x256), you normalize column 1. Then for each image individually, the sum of column 1 equals 0. Next you compute the average of all column 1s. (Each column 1 is 276x1, so the average of all columns 1's is 276x1).  Let's call it <code>global column 1 average</code>. This is a vector of length 276, i.e. height of image). Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the <code>global column 1 average</code>.</p>\n<p>The reason this works is as follows. Imagine that we have 10 images with the same background and no signals yet. Now imagine that we inject strong signal into column number 89 of images 3,5,7. At this point all images' column 1 are still the same.</p>\n<p>Now the competition host normalizes each image individually. At this point column 1 of images 3,5,7 are different from the others. Next we normalize every image column 1 individually. Now all column 1s are the same again.</p>\n<p>Finally we subtract each image's normalized column 1 from <code>global column 1 average</code>. Now all column 1s are zero.</p>",
          "rawMarkdown": "I believe it works as follows. Let's consider only column 1. For every image (actually i mean image channel 276x256), you normalize column 1. Then for each image individually, the sum of column 1 equals 0. Next you compute the average of all column 1s. (Each column 1 is 276x1, so the average of all columns 1's is 276x1).  Let's call it `global column 1 average`. This is a vector of length 276, i.e. height of image). Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the `global column 1 average`.\n\nThe reason this works is as follows. Imagine that we have 10 images with the same background and no signals yet. Now imagine that we inject strong signal into column number 89 of images 3,5,7. At this point all images' column 1 are still the same.\n\nNow the competition host normalizes each image individually. At this point column 1 of images 3,5,7 are different from the others. Next we normalize every image column 1 individually. Now all column 1s are the same again.\n\nFinally we subtract each image's normalized column 1 from `global column 1 average`. Now all column 1s are zero.",
          "votes": 7
        },
        {
          "id": 1481501,
          "postDate": "2021-08-19T14:45:54.040Z",
          "content": "<p>Thanks, Chris! It's starting to get clearer, but I'm still confused. Let me just try to explain where I get lost:</p>\n<blockquote>\n  <p>For every image, you normalize column 1. Then for each image individually, the sum of column 1 equals 0.</p>\n</blockquote>\n<p>Got it. One question on terminology: \"image\" means one 273x256 block? So each sample contains 6 images?</p>\n<blockquote>\n  <p>Next you compute the average of all column 1s. Let's call it global column 1 average. This is a vector of length height of image</p>\n</blockquote>\n<p><br>\nEdit: understood it now. You average pixels across images, therefore you get a vector of lenght image height.</p>\n<blockquote>\n  <p>Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the global column 1 average.</p>\n</blockquote>\n<p>Here I am fully lost. They should all be 0? And what do I do with columns 2-256?</p>\n<p>Thanks a lot for your help!</p>",
          "rawMarkdown": "Thanks, Chris! It's starting to get clearer, but I'm still confused. Let me just try to explain where I get lost:\n\n> For every image, you normalize column 1. Then for each image individually, the sum of column 1 equals 0.\n\nGot it. One question on terminology: \"image\" means one 273x256 block? So each sample contains 6 images?\n\n> Next you compute the average of all column 1s. Let's call it global column 1 average. This is a vector of length height of image\n\n~~This average should be 0. I just normalized them? And if I do this for *all* column 1s, shouldn't it be a vector of length number of samples * 6?~~\nEdit: understood it now. You average pixels across images, therefore you get a vector of lenght image height.\n\n > Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the global column 1 average.\n\nHere I am fully lost. They should all be 0? And what do I do with columns 2-256?\n\nThanks a lot for your help!\n",
          "votes": 1
        },
        {
          "id": 1481511,
          "postDate": "2021-08-19T14:52:46.673Z",
          "content": "<p>Update. I believe they group images into similar clusters. And they subtract normalized column 1 with a \"cluster's global column 1 average\" computed on only the similar images. So <code>global column 1 average</code> is actually a <code>cluster global column 1 average</code> of all images that are similar to the current image.</p>\n<p>If images are collected at similar points in time then their columns are similar. Because 1 column represents the energy at one fixed frequency. (But they are not similar until applying column normalization because the host applied image channel normalization).</p>",
          "rawMarkdown": "Update. I believe they group images into similar clusters. And they subtract normalized column 1 with a \"cluster's global column 1 average\" computed on only the similar images. So `global column 1 average` is actually a `cluster global column 1 average` of all images that are similar to the current image.\n\nIf images are collected at similar points in time then their columns are similar. Because 1 column represents the energy at one fixed frequency. (But they are not similar until applying column normalization because the host applied image channel normalization).",
          "votes": 1
        },
        {
          "id": 1481524,
          "postDate": "2021-08-19T14:59:27.417Z",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> </p>\n<blockquote>\n  <p>\"image\" means one 273x256 block</p>\n</blockquote>\n<p>yes</p>\n<blockquote>\n  <p>This average should be 0. I just normalized them? And if I do this for all column 1s, shouldn't it be a vector of length number of samples * 6?</p>\n</blockquote>\n<p>Each \"image\" is 273 rows by 256 columns. Therefore specifically column 1 is <code>273x1</code>. When i take the average of all column 1s (note these are only column 1s. Not column 2 nor column 3 etc), the result is also <code>273x1</code>.</p>\n<p>Here is a toy example. Here are 2 columns (of length 4). After normalizing, here they are <code>[1, -1, 1, -1]</code> and <code>[2, 2, -2, -2]</code>. Each column has sum 0 by itself because it is normalized. The average (i.e. <code>global column 1 average</code>) of these two columns is <code>[1.5, 0.5, -0.5, -1.5]</code>. </p>\n<p>We then replace the first vector with <code>[1-1.5, -1-0.5, 1-(-0.5), -1-(-1.5)] = [-0.5, 1.5, 1.5, 0.5]</code>.</p>",
          "rawMarkdown": "@friedchips \n\n>\"image\" means one 273x256 block\n\nyes\n\n>This average should be 0. I just normalized them? And if I do this for all column 1s, shouldn't it be a vector of length number of samples * 6?\n\nEach \"image\" is 273 rows by 256 columns. Therefore specifically column 1 is `273x1`. When i take the average of all column 1s (note these are only column 1s. Not column 2 nor column 3 etc), the result is also `273x1`.\n\nHere is a toy example. Here are 2 columns (of length 4). After normalizing, here they are `[1, -1, 1, -1]` and `[2, 2, -2, -2]`. Each column has sum 0 by itself because it is normalized. The average (i.e. `global column 1 average`) of these two columns is `[1.5, 0.5, -0.5, -1.5]`. \n\nWe then replace the first vector with `[1-1.5, -1-0.5, 1-(-0.5), -1-(-1.5)] = [-0.5, 1.5, 1.5, 0.5]`.",
          "votes": 1
        },
        {
          "id": 1481597,
          "postDate": "2021-08-19T15:39:35.267Z",
          "content": "<p>Got it. Thanks again! I think I finally understand magic #2. Even the original explanation by the winning team is beginning to make sense now.</p>\n<p>But: The way I understand it now is that they first normalized every column of every image separately. Then, they took the first (normalized) column of the image they wanted to clean and searched for a match for this column among <em>all</em> other normalized columns in <em>all</em> the data. Why? Because they suspected that the background would be duplicated somewhere. If they found a match, they knew that the columns \"to the right\" of the match would <em>also</em> match the columns to the right of column 1 in the image to be cleaned. Then they could simply subtract column-wise and finally re-normalize the whole image.</p>\n<p>If this is correct, and I strongly suspect so and will check myself, then this only works because the background images repeat (several times?) in the dataset.</p>",
          "rawMarkdown": "Got it. Thanks again! I think I finally understand magic #2. Even the original explanation by the winning team is beginning to make sense now.\n\nBut: The way I understand it now is that they first normalized every column of every image separately. Then, they took the first (normalized) column of the image they wanted to clean and searched for a match for this column among *all* other normalized columns in *all* the data. Why? Because they suspected that the background would be duplicated somewhere. If they found a match, they knew that the columns \"to the right\" of the match would *also* match the columns to the right of column 1 in the image to be cleaned. Then they could simply subtract column-wise and finally re-normalize the whole image.\n\nIf this is correct, and I strongly suspect so and will check myself, then this only works because the background images repeat (several times?) in the dataset.",
          "votes": 2,
          "replies": [
            {
              "id": 2176304,
              "postDate": "2023-03-10T14:40:38.883Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2176306,
              "postDate": "2023-03-10T14:41:45.487Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 1480479,
      "postDate": "2021-08-19T03:30:50.607Z",
      "content": "<p>Congratulations on such huge lead win <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, your 2 magics are so brilliant, I thought you must implement magic #1 to find extra class/pattern that only exist in test data, but your magic #2 is far beyond my imagination, I guess it`s difficult for me to replicate it even after read your solution 😂. <br>\nBTW, \"15h on 8xV100 using DDP\", what does DDP mean here, something like AMP(Auto Mixed Precision)? </p>",
      "rawMarkdown": "Congratulations on such huge lead win @ilu000, your 2 magics are so brilliant, I thought you must implement magic #1 to find extra class/pattern that only exist in test data, but your magic #2 is far beyond my imagination, I guess it`s difficult for me to replicate it even after read your solution 😂. \nBTW, \"15h on 8xV100 using DDP\", what does DDP mean here, something like AMP(Auto Mixed Precision)? ",
      "votes": 5,
      "replies": [
        {
          "id": 1480626,
          "postDate": "2021-08-19T05:27:11.630Z",
          "content": "<p>DDP stands for distributed data parallel which is used to train efficiently on a multi-GPU regime. </p>\n<blockquote>\n  <p>Distributed Data-Parallel Training (DDP) is a widely adopted single-program multiple-data training paradigm. With DDP, the model is replicated on every process, and every model replica will be fed with a different set of input data samples. DDP takes care of gradient communications to keep model replicas synchronized and overlaps it with the gradient computations to speed up training.</p>\n</blockquote>\n<p>(from <a href=\"https://pytorch.org/tutorials/beginner/dist_overview.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/dist_overview.html</a>)</p>",
          "rawMarkdown": "DDP stands for distributed data parallel which is used to train efficiently on a multi-GPU regime. \n\n> Distributed Data-Parallel Training (DDP) is a widely adopted single-program multiple-data training paradigm. With DDP, the model is replicated on every process, and every model replica will be fed with a different set of input data samples. DDP takes care of gradient communications to keep model replicas synchronized and overlaps it with the gradient computations to speed up training.\n\n(from https://pytorch.org/tutorials/beginner/dist_overview.html)",
          "votes": 6
        },
        {
          "id": 1480868,
          "postDate": "2021-08-19T07:55:37.503Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>  for the detailed explaination, looks more computation resouces are more and more important for Kaggle competition.</p>",
          "rawMarkdown": "Thanks @christofhenkel  for the detailed explaination, looks more computation resouces are more and more important for Kaggle competition."
        }
      ]
    },
    {
      "id": 1481571,
      "postDate": "2021-08-19T15:27:03.313Z",
      "content": "<p>Brilliant solution team! Congratulations on 1st place and your amazing lead over 2nd.</p>\n<p>Applying column normalization is brilliant. I knew that the host divided each image channel by different number and subtracted from each image channel a different number but i couldn't think how to correct this imbalanced normalization. Fantastic work with column normalization.</p>\n<p>To detect anomalies we need control images. It was brilliant to use one \"on\" cadence image as control for another \"on\" cadence image because both \"on\" cadence images point to the same point in space. Great idea!</p>",
      "rawMarkdown": "Brilliant solution team! Congratulations on 1st place and your amazing lead over 2nd.\n\nApplying column normalization is brilliant. I knew that the host divided each image channel by different number and subtracted from each image channel a different number but i couldn't think how to correct this imbalanced normalization. Fantastic work with column normalization.\n\nTo detect anomalies we need control images. It was brilliant to use one \"on\" cadence image as control for another \"on\" cadence image because both \"on\" cadence images point to the same point in space. Great idea!",
      "votes": 3
    },
    {
      "id": 1480982,
      "postDate": "2021-08-19T08:58:03.750Z",
      "content": "<p>Would you mind disclosing the Id of your example \"Original sample with signal\" for \"Magic #2\" ?</p>",
      "rawMarkdown": "Would you mind disclosing the Id of your example \"Original sample with signal\" for \"Magic #2\" ?",
      "votes": 3,
      "replies": [
        {
          "id": 1481841,
          "postDate": "2021-08-19T17:37:27.590Z",
          "content": "<p>sure, it's \"0000799a2b2c42d\"</p>",
          "rawMarkdown": "sure, it's \"0000799a2b2c42d\"",
          "votes": 1
        }
      ]
    },
    {
      "id": 1480334,
      "postDate": "2021-08-19T01:11:41.413Z",
      "content": "<p>Congrats, this is an amazing solution. I was following from afar and it reminded me of the <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018\" target=\"_blank\">PLAsTiCC competition</a> (coincidentally, also an astronomy setting) - there, the biggest challenge was also CV/LB gap and the test set had an extra class not found in the training set. The winner in that competition was an astronomer with limited ML experience who used a clever idea to augment the training data to match the test data and won solo with a single LGB against teams using much more sophisticated modelling techniques. I wasn't too familiar with the ins and outs of this competition but after noticing similarities to PLAsTiCC I had a feeling a data approach would make the difference here as well. </p>\n<p>That being said, this solution is even more impressive IMO. It's clear from the write up just how much thought was put into each component. Winning with ~0.97 AUC when second place is ~0.81 is unprecedented as far as I'm aware.</p>",
      "rawMarkdown": "Congrats, this is an amazing solution. I was following from afar and it reminded me of the [PLAsTiCC competition](https://www.kaggle.com/c/PLAsTiCC-2018) (coincidentally, also an astronomy setting) - there, the biggest challenge was also CV/LB gap and the test set had an extra class not found in the training set. The winner in that competition was an astronomer with limited ML experience who used a clever idea to augment the training data to match the test data and won solo with a single LGB against teams using much more sophisticated modelling techniques. I wasn't too familiar with the ins and outs of this competition but after noticing similarities to PLAsTiCC I had a feeling a data approach would make the difference here as well. \n\nThat being said, this solution is even more impressive IMO. It's clear from the write up just how much thought was put into each component. Winning with ~0.97 AUC when second place is ~0.81 is unprecedented as far as I'm aware.",
      "votes": 3,
      "replies": [
        {
          "id": 1481002,
          "postDate": "2021-08-19T09:07:27.940Z",
          "content": "<p>Is .97 too realistic AUC?</p>",
          "rawMarkdown": "Is .97 too realistic AUC?"
        }
      ]
    },
    {
      "id": 1480311,
      "postDate": "2021-08-19T00:34:45.047Z",
      "content": "<p>Thank you for the detailed solution! This is an absolutely beautiful data-centric approach to this competition. Looking forward to learning a lot more from the Watercooled team in the future!</p>",
      "rawMarkdown": "Thank you for the detailed solution! This is an absolutely beautiful data-centric approach to this competition. Looking forward to learning a lot more from the Watercooled team in the future!",
      "votes": 3
    },
    {
      "id": 1481286,
      "postDate": "2021-08-19T12:13:51.617Z",
      "content": "<p>Congrats on the clever use of all data available.  I would love to see more details on how you trained eca-nfnet-l2.</p>\n<p>Let me play  the devil advocate for now.</p>\n<p>Would you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?</p>\n<p>Do you have an idea of where you would end without magic 2?  I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.</p>",
      "rawMarkdown": "Congrats on the clever use of all data available.  I would love to see more details on how you trained eca-nfnet-l2.\n\nLet me play  the devil advocate for now.\n\nWould you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?\n\nDo you have an idea of where you would end without magic 2?  I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.",
      "votes": 4,
      "replies": [
        {
          "id": 1481297,
          "postDate": "2021-08-19T12:23:11.273Z",
          "content": "<blockquote>\n  <p>I would love to see more details on how you trained eca-nfnet-l2.</p>\n</blockquote>\n<p>What would you like to know exactly?</p>\n<blockquote>\n  <p>Would you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?</p>\n</blockquote>\n<p>It is really hard to say honestly. Maybe yes, maybe no. Magic 2 can be totally found only looking at train, or even old train, and is not specific to test data. But often times these findings also involve a bit of luck, so I do not know. That said, I think you know my opinion regarding csv vs. code competitions :)</p>\n<blockquote>\n  <p>Do you have an idea of where you would end without magic 2? I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.</p>\n</blockquote>\n<p>Again, I cant make a clear statement here. But we were second place, basically tied with first before finding it, and we were really confident to further improve that solution, as we still were working with single folds, simpler models, and havent looked into pseudo tagging for example yet where we expected significant gains (as also imminent from other top solutions). What I can say though is that this, at that point of time, second place solution, would still be in gold area now. Seeing the movement of our direct competitors at that point, I would assume we would have ended somewhere in top 3 area, but who knows.</p>",
          "rawMarkdown": "> I would love to see more details on how you trained eca-nfnet-l2.\n\nWhat would you like to know exactly?\n\n> Would you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?\n\nIt is really hard to say honestly. Maybe yes, maybe no. Magic 2 can be totally found only looking at train, or even old train, and is not specific to test data. But often times these findings also involve a bit of luck, so I do not know. That said, I think you know my opinion regarding csv vs. code competitions :)\n\n> Do you have an idea of where you would end without magic 2? I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.\n\nAgain, I cant make a clear statement here. But we were second place, basically tied with first before finding it, and we were really confident to further improve that solution, as we still were working with single folds, simpler models, and havent looked into pseudo tagging for example yet where we expected significant gains (as also imminent from other top solutions). What I can say though is that this, at that point of time, second place solution, would still be in gold area now. Seeing the movement of our direct competitors at that point, I would assume we would have ended somewhere in top 3 area, but who knows.",
          "votes": 5
        },
        {
          "id": 1481417,
          "postDate": "2021-08-19T13:42:54.123Z",
          "content": "<blockquote>\n  <p>What would you like to know exactly?</p>\n</blockquote>\n<p>training hyperparameters like learning rate, optimizer,  scheduler, drop rate, drop path rate, use of amp, etc.</p>",
          "rawMarkdown": "> What would you like to know exactly?\n\ntraining hyperparameters like learning rate, optimizer,  scheduler, drop rate, drop path rate, use of amp, etc.",
          "votes": 1
        },
        {
          "id": 1481621,
          "postDate": "2021-08-19T15:50:01.180Z",
          "content": "<p>Sure:</p>\n<p>backbone = \"eca_nfnet_l2\"</p>\n<p>epochs = 20<br>\nlr = 0.000025<br>\noptimizer = \"AdamW\"<br>\nweight_decay = 1e-4<br>\nbatch_size = 8</p>\n<p>drop_rate = 0.1<br>\ndrop_path_rate = 0.1</p>\n<p>stride = (1,2)</p>\n<p>aug = \"vflip\" + mixup (p=1.0, beta=5, max_target)</p>\n<p>tta = [[\"hflip\"], [\"vflip\"], [\"hflip\", \"vflip\"]]</p>\n<p>cuda.amp with DDP</p>",
          "rawMarkdown": "Sure:\n\nbackbone = \"eca_nfnet_l2\"\n\nepochs = 20\nlr = 0.000025\noptimizer = \"AdamW\"\nweight_decay = 1e-4\nbatch_size = 8\n\ndrop_rate = 0.1\ndrop_path_rate = 0.1\n\nstride = (1,2)\n\naug = \"vflip\" + mixup (p=1.0, beta=5, max_target)\n\ntta = [[\"hflip\"], [\"vflip\"], [\"hflip\", \"vflip\"]]\n\ncuda.amp with DDP",
          "votes": 5
        },
        {
          "id": 1481641,
          "postDate": "2021-08-19T16:01:56.820Z",
          "content": "<p>Psi, if I may ask and since you changed the first conv stride: pretrained or from scratch ? </p>",
          "rawMarkdown": "Psi, if I may ask and since you changed the first conv stride: pretrained or from scratch ? "
        },
        {
          "id": 1481649,
          "postDate": "2021-08-19T16:07:37.483Z",
          "content": "<p>pretrained, I honestly havent seen a case where pretrained is not helpful, at least in terms of training time needed</p>",
          "rawMarkdown": "pretrained, I honestly havent seen a case where pretrained is not helpful, at least in terms of training time needed",
          "votes": 4
        },
        {
          "id": 1481681,
          "postDate": "2021-08-19T16:16:42.910Z",
          "content": "<p>But the fisrt Conv layer is not part of backbone, right? How did you get pretrained weights for this layer? In my unstanding it`s like to use Conv layer to downsample the input image, instead of normal resize which would lose some information.</p>",
          "rawMarkdown": "But the fisrt Conv layer is not part of backbone, right? How did you get pretrained weights for this layer? In my unstanding it`s like to use Conv layer to downsample the input image, instead of normal resize which would lose some information."
        },
        {
          "id": 1481694,
          "postDate": "2021-08-19T16:23:26.063Z",
          "content": "<p>With the first conv layer, we are referring to the first conv layer of the backbone. We modified the stride of that one. </p>\n<p>I am sorry about the image in the above post, which is indeed a bit misleading in this case.</p>",
          "rawMarkdown": "With the first conv layer, we are referring to the first conv layer of the backbone. We modified the stride of that one. \n\nI am sorry about the image in the above post, which is indeed a bit misleading in this case.",
          "votes": 1
        },
        {
          "id": 1481715,
          "postDate": "2021-08-19T16:33:48.840Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> ,  lesson learned. Then I'm wondering if you modified the first Conv layer's stride number of backbone, are the pretained weights still match, wouldn't prompt any error/warning when training?</p>",
          "rawMarkdown": "Thanks @ilu000 ,  lesson learned. Then I'm wondering if you modified the first Conv layer's stride number of backbone, are the pretained weights still match, wouldn't prompt any error/warning when training?"
        },
        {
          "id": 1481718,
          "postDate": "2021-08-19T16:35:55.917Z",
          "content": "<p>Yeah, changing the stride doesnt change anything, kernel size etc stays intact.</p>",
          "rawMarkdown": "Yeah, changing the stride doesnt change anything, kernel size etc stays intact.",
          "votes": 1
        },
        {
          "id": 1481734,
          "postDate": "2021-08-19T16:40:51.390Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> , I just realized change stride didn`t change paramters number, very good trick to replace resize and does not bring any extra computation cost.</p>",
          "rawMarkdown": "Thanks @philippsinger , I just realized change stride didn`t change paramters number, very good trick to replace resize and does not bring any extra computation cost.",
          "votes": -1
        },
        {
          "id": 1481773,
          "postDate": "2021-08-19T16:59:00.167Z",
          "content": "<p>Oh it does bring huge extra computational cost. Similar to increasing image size two times, three times, etc.</p>",
          "rawMarkdown": "Oh it does bring huge extra computational cost. Similar to increasing image size two times, three times, etc.",
          "votes": 1
        },
        {
          "id": 1482062,
          "postDate": "2021-08-19T21:04:47.537Z",
          "content": "<p>When you do mixup with batch size 8. Do you have two dataloaders each supplying 8 images and you mixup those two batches to make one batch? Or do you have one dataloader and mixup the images within the one batch?</p>",
          "rawMarkdown": "When you do mixup with batch size 8. Do you have two dataloaders each supplying 8 images and you mixup those two batches to make one batch? Or do you have one dataloader and mixup the images within the one batch?"
        },
        {
          "id": 1482231,
          "postDate": "2021-08-20T00:55:49.703Z",
          "content": "<p>One dataloader and mix within batch</p>",
          "rawMarkdown": "One dataloader and mix within batch",
          "votes": 2
        },
        {
          "id": 1482237,
          "postDate": "2021-08-20T01:06:30.553Z",
          "content": "<p>Thanks for clarification. I notice this is what everyone at Kaggle does. However the original research paper <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">here</a> (on page 3) uses 2 dataloaders. I suspect for most cases it doesn't matter, but as batch size decreases and alpha increases, it may make a difference.</p>\n<pre><code># y1, y2 should be one-hot vectors\nfor (x1, y1), (x2, y2) in zip(loader1, loader2):\n    lam = numpy.random.beta(alpha, alpha)\n    x = Variable(lam * x1 + (1. - lam) * x2)\n    y = Variable(lam * y1 + (1. - lam) * y2)\n    optimizer.zero_grad()\n    loss(net(x), y).backward()\n    optimizer.step()\n</code></pre>",
          "rawMarkdown": "Thanks for clarification. I notice this is what everyone at Kaggle does. However the original research paper [here][1] (on page 3) uses 2 dataloaders. I suspect for most cases it doesn't matter, but as batch size decreases and alpha increases, it may make a difference.\n\n    # y1, y2 should be one-hot vectors\n    for (x1, y1), (x2, y2) in zip(loader1, loader2):\n        lam = numpy.random.beta(alpha, alpha)\n        x = Variable(lam * x1 + (1. - lam) * x2)\n        y = Variable(lam * y1 + (1. - lam) * y2)\n        optimizer.zero_grad()\n        loss(net(x), y).backward()\n        optimizer.step()\n\n[1]: https://arxiv.org/abs/1710.09412"
        },
        {
          "id": 1482363,
          "postDate": "2021-08-20T03:36:36.290Z",
          "content": "<p>2 dataloaders may negatively impact training time.  I guess none of us saw an upside compared to in batch mixup.</p>",
          "rawMarkdown": "2 dataloaders may negatively impact training time.  I guess none of us saw an upside compared to in batch mixup."
        },
        {
          "id": 1482548,
          "postDate": "2021-08-20T06:27:01.833Z",
          "content": "<p>Actually in this competition we mix across whole data, not within batch. Usually it doesnt make much difference. So we do it directly in the data loader, just mixing a random image from the full data in.</p>\n<p>I personally prefer this version, also due to reasons pointed out by Chris. But if data loading is a bottleneck (e.g., large images), I also switch to batch mixing.</p>\n<p>One thing we do here though, is that we re-normalize after mixing, so that mixed images exhibit same data characteristics as images that you actually score on.</p>",
          "rawMarkdown": "Actually in this competition we mix across whole data, not within batch. Usually it doesnt make much difference. So we do it directly in the data loader, just mixing a random image from the full data in.\n\nI personally prefer this version, also due to reasons pointed out by Chris. But if data loading is a bottleneck (e.g., large images), I also switch to batch mixing.\n\nOne thing we do here though, is that we re-normalize after mixing, so that mixed images exhibit same data characteristics as images that you actually score on.",
          "votes": 1
        },
        {
          "id": 1482804,
          "postDate": "2021-08-20T09:40:29.037Z",
          "content": "<blockquote>\n  <p>So we do it directly in the data loader, just mixing a random image from the full data in.</p>\n</blockquote>\n<p>We tried that too, didn't see much a difference.  But we did not have your cleaned data.</p>",
          "rawMarkdown": ">  So we do it directly in the data loader, just mixing a random image from the full data in.\n\nWe tried that too, didn't see much a difference.  But we did not have your cleaned data.\n\n"
        },
        {
          "id": 1482812,
          "postDate": "2021-08-20T09:45:29.317Z",
          "content": "<p>As said, it usually makes little to no difference. We already used it before cleaned data. I think one minor key is to re-normalize afterwards. And I also believe that within batch mixing can actually be sometimes better with batchnorm models, but I have not investigated this hypothesis yet.</p>",
          "rawMarkdown": "As said, it usually makes little to no difference. We already used it before cleaned data. I think one minor key is to re-normalize afterwards. And I also believe that within batch mixing can actually be sometimes better with batchnorm models, but I have not investigated this hypothesis yet."
        },
        {
          "id": 1483916,
          "postDate": "2021-08-21T01:09:38.637Z",
          "content": "<p>I understand changing stride (2, 2) -&gt; (1, 2) is equal to enlarging input size, but why only magnified hight? I think magnifying both height and width is also possible: (2, 2) -&gt; (1, 1).</p>\n<p>Are there any reason that height(time) dimension is more important than width(frequency) dimension?</p>",
          "rawMarkdown": "I understand changing stride (2, 2) -> (1, 2) is equal to enlarging input size, but why only magnified hight? I think magnifying both height and width is also possible: (2, 2) -> (1, 1).\n\nAre there any reason that height(time) dimension is more important than width(frequency) dimension?",
          "votes": 1
        },
        {
          "id": 1484348,
          "postDate": "2021-08-21T08:30:53.883Z",
          "content": "<p>In this competition it was more important, we found (1,3) to be already really good, and (1,2) to be slightly better. Then it is also a matter of runtime, (1,1) is already really crazy for those larger models. You also saw most other competitors enlarging height, which has similar effects (but should be a bit worse as you lose information).</p>",
          "rawMarkdown": "In this competition it was more important, we found (1,3) to be already really good, and (1,2) to be slightly better. Then it is also a matter of runtime, (1,1) is already really crazy for those larger models. You also saw most other competitors enlarging height, which has similar effects (but should be a bit worse as you lose information)."
        },
        {
          "id": 1484375,
          "postDate": "2021-08-21T08:58:52.047Z",
          "content": "<p>I see. It’ was computational cost trade-off.</p>\n<p>But I still wonder why height is more important than the width. In this competition, signals tend to be longer in height, and tilted signal is positive whereas vertical straight line is negative. On this condition, I think it’s natural that width information is more important to distinguish signal: low resolution in width make it hard to distinguish near-vertical line signal from vertical line noise. Did you run through the condition of (3, 1) or (2, 1), and still (1, 2) is the best?</p>",
          "rawMarkdown": "I see. It’ was computational cost trade-off.\n\nBut I still wonder why height is more important than the width. In this competition, signals tend to be longer in height, and tilted signal is positive whereas vertical straight line is negative. On this condition, I think it’s natural that width information is more important to distinguish signal: low resolution in width make it hard to distinguish near-vertical line signal from vertical line noise. Did you run through the condition of (3, 1) or (2, 1), and still (1, 2) is the best?"
        },
        {
          "id": 1484386,
          "postDate": "2021-08-21T09:19:17.207Z",
          "content": "<p>Yes, changing the stride was one of the first things we tried (even before the reset) and we tested several combinations and for long time we sticked to (1,3) and only in the end reduced that to (1,2).</p>",
          "rawMarkdown": "Yes, changing the stride was one of the first things we tried (even before the reset) and we tested several combinations and for long time we sticked to (1,3) and only in the end reduced that to (1,2).",
          "votes": 1
        },
        {
          "id": 1485616,
          "postDate": "2021-08-22T09:19:11.247Z",
          "content": "<p>I see. Thanks.</p>",
          "rawMarkdown": "I see. Thanks."
        }
      ]
    },
    {
      "id": 1480388,
      "postDate": "2021-08-19T02:19:17.690Z",
      "content": "<p>This read makes me feel like having accomplished nothing haha</p>",
      "rawMarkdown": "This read makes me feel like having accomplished nothing haha",
      "votes": 4
    },
    {
      "id": 1484663,
      "postDate": "2021-08-21T13:39:16.717Z",
      "content": "<p>Nice job. Congratulation</p>",
      "rawMarkdown": "Nice job. Congratulation",
      "votes": 1
    },
    {
      "id": 1481756,
      "postDate": "2021-08-19T16:51:11.893Z",
      "content": "<p>Congratulations Team Watercooled on your outstanding work. Well deserved win!</p>",
      "rawMarkdown": "Congratulations Team Watercooled on your outstanding work. Well deserved win!",
      "votes": 1
    },
    {
      "id": 1481302,
      "postDate": "2021-08-19T12:25:10.587Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>  outstanding delivery and success. Well deserved!</p>",
      "rawMarkdown": "Congratulations @philippsinger @christofhenkel and @ilu000  outstanding delivery and success. Well deserved!",
      "votes": 1
    },
    {
      "id": 1480693,
      "postDate": "2021-08-19T06:24:28.357Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for the big win.</p>",
      "rawMarkdown": "Congratulations @philippsinger @christofhenkel @ilu000 for the big win.",
      "votes": 1
    },
    {
      "id": 1480319,
      "postDate": "2021-08-19T00:45:40.903Z",
      "content": "<p>Awesome job, congrats!</p>",
      "rawMarkdown": "Awesome job, congrats!",
      "votes": 1,
      "replies": [
        {
          "id": 1480369,
          "postDate": "2021-08-19T01:59:06.323Z",
          "content": "<p>Them, yes, however I have to stress that people should  NOT mislead on the forums and even worse if it comes from the <strong>organizer</strong>. I am referring to this post:</p>\n<p><a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405282\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405282</a></p>\n<p>and this post:</p>\n<p><a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405732\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405732</a></p>\n<p><strong>where you hinted that the high score was due to (probably) the extra class</strong> in the test and \"you are not worried\"(when clearly the <code>0.853</code> score was after the \"data cleaning\" method) .</p>\n<p>And you also became conclusive about it via saying:</p>\n<blockquote>\n  <p>Sorry I spoiled it for you, didn't want people worrying about another reset .</p>\n</blockquote>\n<p><strong>Even worse that they contacted you about this</strong>. And even if your post came before them contacting you, still it does not matter because the misleading post is there - you should have made amends. </p>\n<p>I understand that you did not want another reset, but it would have been better if you hadn't said anything. </p>\n<p>Kaggle will need to look into this in the future : <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>",
          "rawMarkdown": "Them, yes, however I have to stress that people should  NOT mislead on the forums and even worse if it comes from the **organizer**. I am referring to this post:\n\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405282\n\nand this post:\n\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405732\n\n **where you hinted that the high score was due to (probably) the extra class** in the test and \"you are not worried\"(when clearly the `0.853` score was after the \"data cleaning\" method) .\n\nAnd you also became conclusive about it via saying:\n\n> Sorry I spoiled it for you, didn't want people worrying about another reset .\n\n  **Even worse that they contacted you about this**. And even if your post came before them contacting you, still it does not matter because the misleading post is there - you should have made amends. \n\nI understand that you did not want another reset, but it would have been better if you hadn't said anything. \n\nKaggle will need to look into this in the future : @inversion \n\n",
          "votes": 19
        },
        {
          "id": 1480406,
          "postDate": "2021-08-19T02:31:38.457Z",
          "content": "<p>We tried Arcface and many unsupervised methods to discover the \"unseen class\" and we couldn't find it. We also found that many of the samples in the training data are simply not learnable by a CNN or discoverable by eyes. The data cleaning is indeed brilliant and impressive.</p>",
          "rawMarkdown": "We tried Arcface and many unsupervised methods to discover the \"unseen class\" and we couldn't find it. We also found that many of the samples in the training data are simply not learnable by a CNN or discoverable by eyes. The data cleaning is indeed brilliant and impressive.",
          "votes": 6
        },
        {
          "id": 1480446,
          "postDate": "2021-08-19T03:09:31.600Z",
          "content": "<p>I'm pretty sure it was not extra class, but 'cleaning', and i agree with <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "rawMarkdown": "I'm pretty sure it was not extra class, but 'cleaning', and i agree with @kazanova ",
          "votes": 2
        },
        {
          "id": 1480882,
          "postDate": "2021-08-19T08:02:32.777Z",
          "content": "<blockquote>\n  <p>That led us to start investigating single predictions, where different models with different LB scores disagree</p>\n</blockquote>\n<p>So top1 decided it is s-shape new class and set it to label one for all examples using human eye. Well, is your pipeline automatic for any other new class with other tricky signal and do not require human eye intervention? In other words can model identify other new pattern without performance degradation? If no, i think it is similar to handy labeled test data and rule violation? <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> what do you think guys?  </p>",
          "rawMarkdown": "> That led us to start investigating single predictions, where different models with different LB scores disagree\n\nSo top1 decided it is s-shape new class and set it to label one for all examples using human eye. Well, is your pipeline automatic for any other new class with other tricky signal and do not require human eye intervention? In other words can model identify other new pattern without performance degradation? If no, i think it is similar to handy labeled test data and rule violation? @inversion @yuhongc @kazanova what do you think guys?  ",
          "votes": -5
        },
        {
          "id": 1480977,
          "postDate": "2021-08-19T08:56:03.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> I would appreciate you not making unfounded accusations and if you would read our solution post you would obviously see that we definitely did not hand label any test data.</p>",
          "rawMarkdown": "@sggpls I would appreciate you not making unfounded accusations and if you would read our solution post you would obviously see that we definitely did not hand label any test data.",
          "votes": 5
        },
        {
          "id": 1480995,
          "postDate": "2021-08-19T09:04:58.573Z",
          "content": "<blockquote>\n  <p>So top1 decided it is s-shape new class and set it to label one for all examples using human eye</p>\n</blockquote>\n<p>To clarify on this, of course we did NOT hand label any test data. We used a signal generator and randomly applied that (together with target=1) on the train set. </p>\n<p>Regarding your point on generalization to new signals/classes, this is obviously a huge challenge in all computer vision tasks and not really solved, yet. Supervised ML can only reliably learn what is already in the training set. </p>",
          "rawMarkdown": ">So top1 decided it is s-shape new class and set it to label one for all examples using human eye\n\nTo clarify on this, of course we did NOT hand label any test data. We used a signal generator and randomly applied that (together with target=1) on the train set. \n\nRegarding your point on generalization to new signals/classes, this is obviously a huge challenge in all computer vision tasks and not really solved, yet. Supervised ML can only reliably learn what is already in the training set. "
        },
        {
          "id": 1481019,
          "postDate": "2021-08-19T09:17:15.633Z",
          "content": "<p>I said <code>decided</code> that s-shape is one-target by human. Because not all your models said that it is one-target. So you decided it is one-target for all samples not your model decide. Am I right?   </p>",
          "rawMarkdown": "I said `decided` that s-shape is one-target by human. Because not all your models said that it is one-target. So you decided it is one-target for all samples not your model decide. Am I right?   ",
          "votes": 3
        },
        {
          "id": 1481026,
          "postDate": "2021-08-19T09:21:11.827Z",
          "content": "<p>Anyway I would like to hear the official answer from the admins for the future. I would also like to understand if the generated data is external data or not.</p>",
          "rawMarkdown": "Anyway I would like to hear the official answer from the admins for the future. I would also like to understand if the generated data is external data or not.",
          "votes": 2
        },
        {
          "id": 1481027,
          "postDate": "2021-08-19T09:21:20.333Z",
          "content": "<p>Anyway, you guys did a great job, congratulations, I take off my hat to you!</p>",
          "rawMarkdown": "Anyway, you guys did a great job, congratulations, I take off my hat to you!",
          "votes": 1
        },
        {
          "id": 1481036,
          "postDate": "2021-08-19T09:24:14.017Z",
          "content": "<p><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> think s-shape is not target. It is not only in 0,2,4, but also in 1,3,5. I found it too, but I didn't know how to use it.</p>\n<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a></p>",
          "rawMarkdown": "@sggpls think s-shape is not target. It is not only in 0,2,4, but also in 1,3,5. I found it too, but I didn't know how to use it.\n\nCongratulations @philippsinger @christofhenkel @ilu000",
          "votes": 1
        }
      ]
    },
    {
      "id": 1497724,
      "postDate": "2021-08-31T11:45:04.533Z",
      "content": "<p>Congrats =))</p>",
      "rawMarkdown": "Congrats =))"
    },
    {
      "id": 1480442,
      "postDate": "2021-08-19T03:06:06.713Z",
      "content": "<p>but why do stealth submitting?</p>",
      "rawMarkdown": "but why do stealth submitting?",
      "votes": 1
    },
    {
      "id": 1481058,
      "postDate": "2021-08-19T09:39:58.473Z",
      "content": "<p>Great Work!. So the keypoint is using signal generator.<br>\n<a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a><br>\nIt makes model fitting on the signal than background.<br>\nThank you for sharing. I will try to review and apply this generator.</p>",
      "rawMarkdown": "Great Work!. So the keypoint is using signal generator.\nhttps://github.com/bbrzycki/setigen\nIt makes model fitting on the signal than background.\nThank you for sharing. I will try to review and apply this generator.",
      "votes": -1,
      "replies": [
        {
          "id": 1481217,
          "postDate": "2021-08-19T11:24:55.930Z",
          "content": "<p>By reading the solution I wouldn't say that's the only \"keypoint\", there are many key points to take into consideration imo…</p>",
          "rawMarkdown": "By reading the solution I wouldn't say that's the only \"keypoint\", there are many key points to take into consideration imo...",
          "votes": 3
        },
        {
          "id": 1481441,
          "postDate": "2021-08-19T13:57:30.490Z",
          "content": "<p>My understanding is that magic #1 (signal generator) gets them to LB 812 first place. And magic #2 gets them to LB 968 first place. I think magic #2 was the powerhouse! </p>",
          "rawMarkdown": "My understanding is that magic #1 (signal generator) gets them to LB 812 first place. And magic #2 gets them to LB 968 first place. I think magic #2 was the powerhouse! ",
          "votes": 2
        },
        {
          "id": 1481457,
          "postDate": "2021-08-19T14:09:02.010Z",
          "content": "<p>they got to 0.800 with magic 1, then found magic 2.  We (Giba) found  magic 1 too, as several other top teams.  It proves that the key was to find magic 2.</p>",
          "rawMarkdown": "they got to 0.800 with magic 1, then found magic 2.  We (Giba) found  magic 1 too, as several other top teams.  It proves that the key was to find magic 2.",
          "votes": 6
        }
      ]
    },
    {
      "id": 2183881,
      "postDate": "2023-03-16T02:11:44.583Z",
      "content": "<p>This is an interesting solution. </p>",
      "rawMarkdown": "This is an interesting solution. "
    },
    {
      "id": 1499027,
      "postDate": "2021-09-01T11:53:06.077Z",
      "content": "<p>congratulations , good job</p>",
      "rawMarkdown": "congratulations , good job"
    },
    {
      "id": 1498422,
      "postDate": "2021-09-01T01:25:56.003Z",
      "content": "<p>Great Work</p>",
      "rawMarkdown": "Great Work"
    },
    {
      "id": 1496159,
      "postDate": "2021-08-30T06:39:40.117Z",
      "content": "<p>Thank you for sharing this awesome solution!</p>",
      "rawMarkdown": "Thank you for sharing this awesome solution!"
    },
    {
      "id": 1488170,
      "postDate": "2021-08-24T06:21:28.197Z",
      "content": "<p>Congrat guys! Visualisation parts is nice<br>\nThanks for sharing 👍💥</p>",
      "rawMarkdown": "Congrat guys! Visualisation parts is nice\nThanks for sharing 👍💥"
    },
    {
      "id": 1487920,
      "postDate": "2021-08-24T00:45:07.580Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 1487831,
      "postDate": "2021-08-23T21:42:50.047Z",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, how computationally intense is magic #2? Iterating each column seems to take forever even using rapids.ai. Any trick to run this analysis faster?</p>",
      "rawMarkdown": "@ilu000, how computationally intense is magic #2? Iterating each column seems to take forever even using rapids.ai. Any trick to run this analysis faster?",
      "replies": [
        {
          "id": 1488111,
          "postDate": "2021-08-24T05:42:29.033Z",
          "content": "<ol>\n<li>compare only part of the columns, e.g. first 100 rows.</li>\n<li>load all images into RAM (but only use first 100 rows for each image and only OFF channels)</li>\n<li>use array operations on your single 60,000 x 100 x … array</li>\n</ol>",
          "rawMarkdown": "1. compare only part of the columns, e.g. first 100 rows.\n2. load all images into RAM (but only use first 100 rows for each image and only OFF channels)\n3. use array operations on your single 60,000 x 100 x ... array",
          "votes": 3
        },
        {
          "id": 1496435,
          "postDate": "2021-08-30T11:30:32.197Z",
          "content": "<p>Thanks! Did you share the code for magic #2? Would you mind sharing it?</p>",
          "rawMarkdown": "Thanks! Did you share the code for magic #2? Would you mind sharing it?"
        }
      ]
    },
    {
      "id": 1486856,
      "postDate": "2021-08-23T08:57:22.930Z",
      "content": "<p>awesome, thx for sharing. </p>",
      "rawMarkdown": "awesome, thx for sharing. "
    },
    {
      "id": 1486631,
      "postDate": "2021-08-23T05:41:56.417Z",
      "content": "<p>Congratulations, guys! This is one of the most wonderful works on signal processing I have ever read. Clean and effective. Good luck with your future works!</p>",
      "rawMarkdown": "Congratulations, guys! This is one of the most wonderful works on signal processing I have ever read. Clean and effective. Good luck with your future works!"
    },
    {
      "id": 1486080,
      "postDate": "2021-08-22T16:38:07.927Z",
      "content": "<p>congratulations! huge achievement.</p>",
      "rawMarkdown": "congratulations! huge achievement."
    },
    {
      "id": 1485394,
      "postDate": "2021-08-22T04:19:47.327Z",
      "content": "<p>Amazing Data Analysis you did .</p>",
      "rawMarkdown": "Amazing Data Analysis you did ."
    },
    {
      "id": 1484382,
      "postDate": "2021-08-21T09:14:13.007Z",
      "content": "<p>Nice work. And very interesting to read. Thank you from Austria</p>",
      "rawMarkdown": "Nice work. And very interesting to read. Thank you from Austria"
    },
    {
      "id": 1483308,
      "postDate": "2021-08-20T15:13:20.323Z",
      "content": "<p>Amazing solution Watercooled team, huge congratulations!! Could you share how you measure signal to noise ratio in the images?</p>",
      "rawMarkdown": "Amazing solution Watercooled team, huge congratulations!! Could you share how you measure signal to noise ratio in the images?"
    },
    {
      "id": 1483138,
      "postDate": "2021-08-20T13:16:09.453Z",
      "content": "<p><a href=\"https://www.kaggle.com/llu\" target=\"_blank\">@llu</a> There is so much to grasp and understand . Clealy, you worked hard. Grasping one by one </p>",
      "rawMarkdown": "@llu There is so much to grasp and understand . Clealy, you worked hard. Grasping one by one "
    },
    {
      "id": 1482172,
      "postDate": "2021-08-19T23:06:59.933Z",
      "content": "<p>Congratulations and thanks for sharing! Very interesting!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! Very interesting!"
    },
    {
      "id": 1482064,
      "postDate": "2021-08-19T21:05:24.773Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! "
    },
    {
      "id": 1481796,
      "postDate": "2021-08-19T17:06:31.913Z",
      "content": "<p>Congratulations, and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations, and thanks for sharing."
    },
    {
      "id": 1481786,
      "postDate": "2021-08-19T17:01:17.237Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 1481189,
      "postDate": "2021-08-19T11:13:02.313Z",
      "content": "<p>thx for sharing!</p>",
      "rawMarkdown": "thx for sharing!\n\n"
    },
    {
      "id": 1480461,
      "postDate": "2021-08-19T03:20:46.043Z",
      "content": "<p>Marvelous! <br>\nthank you for the detailed explanations.</p>",
      "rawMarkdown": "Marvelous! \nthank you for the detailed explanations."
    },
    {
      "id": 1480426,
      "postDate": "2021-08-19T02:54:44.193Z",
      "content": "<p>Congratulations! Your work is so brilliant! </p>",
      "rawMarkdown": "Congratulations! Your work is so brilliant! "
    },
    {
      "id": 1480377,
      "postDate": "2021-08-19T02:06:58.483Z",
      "content": "<p>brilliant!!!</p>",
      "rawMarkdown": "brilliant!!!"
    },
    {
      "id": 1480339,
      "postDate": "2021-08-19T01:25:26.947Z",
      "content": "<p>What a solution!<br>\n天秀！</p>",
      "rawMarkdown": "What a solution!\n天秀！"
    },
    {
      "id": 1480314,
      "postDate": "2021-08-19T00:37:24.670Z",
      "content": "<p>Incredibly outstanding!👋👋👋Respect!</p>",
      "rawMarkdown": "Incredibly outstanding!👋👋👋Respect!"
    },
    {
      "id": 1480309,
      "postDate": "2021-08-19T00:34:29.290Z",
      "content": "<p>thx for sharing!</p>",
      "rawMarkdown": "thx for sharing!"
    },
    {
      "id": 1495845,
      "postDate": "2021-08-29T20:40:25.617Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1495842,
      "postDate": "2021-08-29T20:39:47.277Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1495835,
      "postDate": "2021-08-29T20:34:54.460Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1480503,
      "postDate": "2021-08-19T03:53:56.547Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1498221,
      "postDate": "2021-08-31T19:03:58.847Z",
      "content": "<p>Thanks for the expalnation</p>",
      "rawMarkdown": "Thanks for the expalnation\n\n"
    },
    {
      "id": 1480382,
      "postDate": "2021-08-19T02:13:47.777Z",
      "content": "<p>Thank you for the detailed solution!  </p>",
      "rawMarkdown": "Thank you for the detailed solution!  "
    }
  ],
  "comments": [
    {
      "id": 1481215,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-08-19T11:24:40.047000",
      "content": "<p>Just one more note on magic #2 for further elaboration. </p>\n<p>Imagine the following simplistic use case:<br>\nYou have a set of images with certain backgrounds (white, green, striped, constant, etc.) and then you have black dots on those images and you want your model to identify if an image contains a black spot (anomaly). Unfortunately, you have some backgrounds that always exhibit a black spot, or maybe never exhibit one. If you are unlucky, your model might decide to overfit on the actual background, instead of trying to identify the raw signal of the black spot. And here in this competition we observed quite obvious overfit on the background, and a challenge was to steer the models to learn the actual patterns. </p>\n<p>One thing most people did was to do mixup, so mixing different backgrounds together and having the signals in different intensity. One other thing was to utilize pseudo tagging, that then also better learns to adjust to the new background (distribution) of the test data. But the best solution is to process the images in a way that they increase the signal to noise ratio, which we did by cleaning common backgrounds among images. In the example above, it would be obviously better to completely remove all background, and just keep the actual signal.</p>\n<p>I do personally not see how this is necessarily only limited to this data and problem. Of course, the data is synthetic here, so the effects are emphasized, but I have observed similar effects before, in other problems, competitions, and projects, but have never so thoroughly thought about it like here. There are many use cases where this is an imminent issue, such as fault prediction in machinery, anomaly prediction, steganalysis, or even chest x-ray prediction where often models prefer to overfit on certain device characteristics.</p>\n<p>Will the same technique applied here work for all problems? Probably not, but the thought process can be the same. So I would encourage you to see it as a learning, at least that's what I am doing, and try to approach these issues in the future in different projects. I can imagine, that clever modeling approaches can also attempt to deal with these issues directly, for example via attention like 3rd place used. There is also plenty of research in the area of domain adaption, where I have seen little being applied on Kaggle, and I have only started to look into it:<br>\n<a href=\"https://paperswithcode.com/task/domain-adaptation\" target=\"_blank\">https://paperswithcode.com/task/domain-adaptation</a></p>",
      "votes": 24,
      "replies": [
        {
          "id": 1481268,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T11:54:58.773000",
          "content": "<p>I agree with most of this except on one thing.  You say data is synthetic.  The messages are synthetic for sure, but the background does not seem to be synthetic according to the data description page:</p>\n<blockquote>\n  <p>Breakthrough Listen generates similar spectrograms to the one shown above, but typically spanning several GHz of the radio spectrum (rather than the approx. 2 MHz shown above). The data are stored either as filterbank format or HDF5 format files, but essentially are arrays of intensity as a function of frequency and time, accompanied by headers containing metadata such as the direction the telescope was pointed in, the frequency scale, and so on. We generate over 1 PB of spectrograms per year; individual filterbank files can be tens of GB in size. For the purposes of the Kaggle challenge, we have discarded the majority of the metadata and are simply presenting numpy arrays consisting of small regions of the spectrograms that we refer to as “snippets”.</p>\n</blockquote>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1481346,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T12:58:09.653000",
          "content": "<p>First - congratulations on your incredible result! Since data preparation/cleaning is also my main interest, I would really like to understand your magic #2 in detail. Based on the infomation in this thread, I haven't yet managed to. Will you add some explanations and/or code later on? Thanks!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481504,
          "author_name": "dan",
          "author_url": "",
          "post_date": "2021-08-19T14:49:39.303000",
          "content": "<p>I am so grateful that you figured this out and shared it. <br>\nImagine a possible and boring parallel universe where no one in the competition found these magics!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1486177,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2021-08-22T17:59:06.793000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> could u through some light on cleaning methods u used to improve SNR ..<br>\n<code>\" As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential.\"</code><br>\nthanks in advance..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480415,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2021-08-19T02:45:02.700000",
      "content": "<p>Huge congratulations! I've never seen a win with such a big lead! </p>\n<p>It never occurred to me that the same background image, by itself and the image into which it was injected with a signal, would be in the training set at the same time.</p>\n<p>Such a dataset does not occur in industry or real world, so this is completely beyond my imagination.</p>\n<p>I feel shame on sharing our solution this time so I'll just skip it. 😅</p>",
      "votes": 17,
      "replies": [
        {
          "id": 1481082,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2021-08-19T09:51:52.270000",
          "content": "<blockquote>\n  <p>Such a dataset does not occur in industry or real world, so this is completely beyond my imagination.</p>\n</blockquote>\n<p>The organizer like it though - so you may have to rethink about this!</p>\n<p>It seems like they wanted you to try and solve that little puzzle in there</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1481099,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2021-08-19T10:08:14.467000",
          "content": "<p>E.g . The organizer encouraged you to find the leaks they injected in the data - This has gone downhill from the time that we were wondering whether solutions with leaks should be viable to encouraging people using them .</p>\n<p><strong>Jut to be clear - the winning team ,did the right thing raising that and reporting it</strong></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1481101,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2021-08-19T10:10:30.257000",
          "content": "<p>Poor us, we were trying to build something that could work in the real world (at least to some extend)</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1481748,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2021-08-19T16:46:38.260000",
          "content": "<blockquote>\n  <p>The organizer like it though - so you may have to rethink about this!<br>\n  It seems like they wanted you to try and solve that little puzzle in there</p>\n</blockquote>\n<p>I'm a MLE. Personally, I came here mainly to analyse other datasets (kaggling and reading others post) to find something that works for the datasets I work on in my company.</p>\n<p>Thanks for that so far I have found many methods or tricks to improve the performance of my company's models ;)</p>\n<blockquote>\n  <p>Jut to be clear - the winning team ,did the right thing raising that and reporting it</p>\n</blockquote>\n<p>I agree. So I don't actually feel unhappy about the competition, whether the puzzle was deliberately left in place by the organisers or was an oversight.</p>\n<p>Even if I knew about the puzzle beforehand, it shouldn't have affected my main objective - to try out various methods and tricks in an effort to make it work for both the competition and my own work.</p>\n<p>This is my way of kaggling, but I'm not against other ways of kaggling, as long as they don't break the rules.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1481943,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2021-08-19T18:42:18.343000",
          "content": "<p>Fair enough. I guess I was hoping to have a leak-free competition  the second time after what happened in the first, but that was wishful thinking for my part. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1482381,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2021-08-20T03:55:24.430000",
          "content": "<p>Haha, I still believed it was leek-free when the LB was at 0.85, but eventually found out that was wrong.</p>\n<p>But the point I have to worry about is that while this is an obvious leek in the MLE's view, it's not easy to convince people who don't know much about machine learning ;)</p>\n<p>This is not meant to be negative, but the fact is that there are a number of competition organisers whose directors don't know much about machine learning.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1480321,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-08-19T00:48:46.370000",
      "content": "<p>Congratulations, we had found the magic#1 but not the magic#2 even if we've spent some time on pre-processing/cleaning.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 1480376,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-08-19T02:06:06.373000",
          "content": "<p>Thank you. Congratulation to becoming a grandmaster <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> . Huge achievement.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1480652,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-08-19T05:53:48.697000",
          "content": "<p>Huge Congratulations <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> on becoming grandmaster and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> &amp; team on the win , we also found magic1 but it did not elate our score by a large margin mostly because we were not inducing signal properly in train like <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> does , we were not looking into anything else in last few days as we were totally biased that magic 1 was reason for the huge boost as said by the hosts , but this solution is absolutely mind blowing </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1480805,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-08-19T07:22:15.330000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> ! We were also exited when we discovered magic#1 because we expected it would close the gap but after a few submissions we realized that it will not and something more was needed. Writeup of our solution/findings is coming …</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1480422,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2021-08-19T02:52:15.873000",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> <br>\nCongratulations &amp; thanks for sharing the solution.</p>\n<p>I'm not sure of the process of magic #2.</p>\n<blockquote>\n  <p>we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns.</p>\n</blockquote>\n<p>Does it mean the process below?</p>\n<ol>\n<li>normalize column vector for each images</li>\n<li>calculate pixel difference from the first column per each image</li>\n<li>take the mean of 2 (column-wise), which give us the one feature vector</li>\n<li>using feature vector given on 3, match the identical background image in the all image space (using kNN etc.)</li>\n<li>take pixel difference from the matched image, and clear the background</li>\n</ol>\n<p>Example: process for 3x3, 1-channel image</p>\n<pre><code># original image\nimg1 == [\n  [1, 4, 3],\n  [2, 3, 3],\n  [3, 2, 3],\n]\n\n# after column-wise normalization\nimg1 == [\n  [-1, 1, 0],\n  [0, 0, 0],\n  [1, -1, 0],\n]\n\n# after taking column-wise difference from the first column\nimg1 == [\n  [0, 2, 1],\n  [0, 0, 0],\n  [0, -2, -1],\n]\n\n# after calculating mean pixel difference\nimg1 == [\n  [1],\n  [0],\n  [-1],\n]\n</code></pre>",
      "votes": 7,
      "replies": [
        {
          "id": 1480937,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T08:32:08.990000",
          "content": "<p>I've read the explanation of magic #2 by the winning team multiple times and I still don't get it at all.. I understand your process on a technical level - but how does this help to remove the background? That's what I don't get.</p>\n<p>Is it basically a leak because images appear multiple times in the train set or what? I'm hoping for some code later on.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481462,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T14:15:03.770000",
          "content": "<p>I believe it works as follows. Let's consider only column 1. For every image (actually i mean image channel 276x256), you normalize column 1. Then for each image individually, the sum of column 1 equals 0. Next you compute the average of all column 1s. (Each column 1 is 276x1, so the average of all columns 1's is 276x1).  Let's call it <code>global column 1 average</code>. This is a vector of length 276, i.e. height of image). Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the <code>global column 1 average</code>.</p>\n<p>The reason this works is as follows. Imagine that we have 10 images with the same background and no signals yet. Now imagine that we inject strong signal into column number 89 of images 3,5,7. At this point all images' column 1 are still the same.</p>\n<p>Now the competition host normalizes each image individually. At this point column 1 of images 3,5,7 are different from the others. Next we normalize every image column 1 individually. Now all column 1s are the same again.</p>\n<p>Finally we subtract each image's normalized column 1 from <code>global column 1 average</code>. Now all column 1s are zero.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1481501,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T14:45:54.040000",
          "content": "<p>Thanks, Chris! It's starting to get clearer, but I'm still confused. Let me just try to explain where I get lost:</p>\n<blockquote>\n  <p>For every image, you normalize column 1. Then for each image individually, the sum of column 1 equals 0.</p>\n</blockquote>\n<p>Got it. One question on terminology: \"image\" means one 273x256 block? So each sample contains 6 images?</p>\n<blockquote>\n  <p>Next you compute the average of all column 1s. Let's call it global column 1 average. This is a vector of length height of image</p>\n</blockquote>\n<p><br>\nEdit: understood it now. You average pixels across images, therefore you get a vector of lenght image height.</p>\n<blockquote>\n  <p>Finally you replace every image's column 1 with the difference between their individual normalized column 1 and the the global column 1 average.</p>\n</blockquote>\n<p>Here I am fully lost. They should all be 0? And what do I do with columns 2-256?</p>\n<p>Thanks a lot for your help!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481511,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T14:52:46.673000",
          "content": "<p>Update. I believe they group images into similar clusters. And they subtract normalized column 1 with a \"cluster's global column 1 average\" computed on only the similar images. So <code>global column 1 average</code> is actually a <code>cluster global column 1 average</code> of all images that are similar to the current image.</p>\n<p>If images are collected at similar points in time then their columns are similar. Because 1 column represents the energy at one fixed frequency. (But they are not similar until applying column normalization because the host applied image channel normalization).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481524,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T14:59:27.417000",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> </p>\n<blockquote>\n  <p>\"image\" means one 273x256 block</p>\n</blockquote>\n<p>yes</p>\n<blockquote>\n  <p>This average should be 0. I just normalized them? And if I do this for all column 1s, shouldn't it be a vector of length number of samples * 6?</p>\n</blockquote>\n<p>Each \"image\" is 273 rows by 256 columns. Therefore specifically column 1 is <code>273x1</code>. When i take the average of all column 1s (note these are only column 1s. Not column 2 nor column 3 etc), the result is also <code>273x1</code>.</p>\n<p>Here is a toy example. Here are 2 columns (of length 4). After normalizing, here they are <code>[1, -1, 1, -1]</code> and <code>[2, 2, -2, -2]</code>. Each column has sum 0 by itself because it is normalized. The average (i.e. <code>global column 1 average</code>) of these two columns is <code>[1.5, 0.5, -0.5, -1.5]</code>. </p>\n<p>We then replace the first vector with <code>[1-1.5, -1-0.5, 1-(-0.5), -1-(-1.5)] = [-0.5, 1.5, 1.5, 0.5]</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481597,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T15:39:35.267000",
          "content": "<p>Got it. Thanks again! I think I finally understand magic #2. Even the original explanation by the winning team is beginning to make sense now.</p>\n<p>But: The way I understand it now is that they first normalized every column of every image separately. Then, they took the first (normalized) column of the image they wanted to clean and searched for a match for this column among <em>all</em> other normalized columns in <em>all</em> the data. Why? Because they suspected that the background would be duplicated somewhere. If they found a match, they knew that the columns \"to the right\" of the match would <em>also</em> match the columns to the right of column 1 in the image to be cleaned. Then they could simply subtract column-wise and finally re-normalize the whole image.</p>\n<p>If this is correct, and I strongly suspect so and will check myself, then this only works because the background images repeat (several times?) in the dataset.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2176304,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-10T14:40:38.883000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2176306,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-10T14:41:45.487000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1480479,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-08-19T03:30:50.607000",
      "content": "<p>Congratulations on such huge lead win <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, your 2 magics are so brilliant, I thought you must implement magic #1 to find extra class/pattern that only exist in test data, but your magic #2 is far beyond my imagination, I guess it`s difficult for me to replicate it even after read your solution 😂. <br>\nBTW, \"15h on 8xV100 using DDP\", what does DDP mean here, something like AMP(Auto Mixed Precision)? </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1480626,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-08-19T05:27:11.630000",
          "content": "<p>DDP stands for distributed data parallel which is used to train efficiently on a multi-GPU regime. </p>\n<blockquote>\n  <p>Distributed Data-Parallel Training (DDP) is a widely adopted single-program multiple-data training paradigm. With DDP, the model is replicated on every process, and every model replica will be fed with a different set of input data samples. DDP takes care of gradient communications to keep model replicas synchronized and overlaps it with the gradient computations to speed up training.</p>\n</blockquote>\n<p>(from <a href=\"https://pytorch.org/tutorials/beginner/dist_overview.html\" target=\"_blank\">https://pytorch.org/tutorials/beginner/dist_overview.html</a>)</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1480868,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-19T07:55:37.503000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>  for the detailed explaination, looks more computation resouces are more and more important for Kaggle competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1481571,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-19T15:27:03.313000",
      "content": "<p>Brilliant solution team! Congratulations on 1st place and your amazing lead over 2nd.</p>\n<p>Applying column normalization is brilliant. I knew that the host divided each image channel by different number and subtracted from each image channel a different number but i couldn't think how to correct this imbalanced normalization. Fantastic work with column normalization.</p>\n<p>To detect anomalies we need control images. It was brilliant to use one \"on\" cadence image as control for another \"on\" cadence image because both \"on\" cadence images point to the same point in space. Great idea!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1480982,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2021-08-19T08:58:03.750000",
      "content": "<p>Would you mind disclosing the Id of your example \"Original sample with signal\" for \"Magic #2\" ?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1481841,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-19T17:37:27.590000",
          "content": "<p>sure, it's \"0000799a2b2c42d\"</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1480334,
      "author_name": "Ming Pan",
      "author_url": "",
      "post_date": "2021-08-19T01:11:41.413000",
      "content": "<p>Congrats, this is an amazing solution. I was following from afar and it reminded me of the <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018\" target=\"_blank\">PLAsTiCC competition</a> (coincidentally, also an astronomy setting) - there, the biggest challenge was also CV/LB gap and the test set had an extra class not found in the training set. The winner in that competition was an astronomer with limited ML experience who used a clever idea to augment the training data to match the test data and won solo with a single LGB against teams using much more sophisticated modelling techniques. I wasn't too familiar with the ins and outs of this competition but after noticing similarities to PLAsTiCC I had a feeling a data approach would make the difference here as well. </p>\n<p>That being said, this solution is even more impressive IMO. It's clear from the write up just how much thought was put into each component. Winning with ~0.97 AUC when second place is ~0.81 is unprecedented as far as I'm aware.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1481002,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2021-08-19T09:07:27.940000",
          "content": "<p>Is .97 too realistic AUC?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480311,
      "author_name": "hide on bread",
      "author_url": "",
      "post_date": "2021-08-19T00:34:45.047000",
      "content": "<p>Thank you for the detailed solution! This is an absolutely beautiful data-centric approach to this competition. Looking forward to learning a lot more from the Watercooled team in the future!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1481286,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-19T12:13:51.617000",
      "content": "<p>Congrats on the clever use of all data available.  I would love to see more details on how you trained eca-nfnet-l2.</p>\n<p>Let me play  the devil advocate for now.</p>\n<p>Would you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?</p>\n<p>Do you have an idea of where you would end without magic 2?  I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1481297,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T12:23:11.273000",
          "content": "<blockquote>\n  <p>I would love to see more details on how you trained eca-nfnet-l2.</p>\n</blockquote>\n<p>What would you like to know exactly?</p>\n<blockquote>\n  <p>Would you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?</p>\n</blockquote>\n<p>It is really hard to say honestly. Maybe yes, maybe no. Magic 2 can be totally found only looking at train, or even old train, and is not specific to test data. But often times these findings also involve a bit of luck, so I do not know. That said, I think you know my opinion regarding csv vs. code competitions :)</p>\n<blockquote>\n  <p>Do you have an idea of where you would end without magic 2? I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.</p>\n</blockquote>\n<p>Again, I cant make a clear statement here. But we were second place, basically tied with first before finding it, and we were really confident to further improve that solution, as we still were working with single folds, simpler models, and havent looked into pseudo tagging for example yet where we expected significant gains (as also imminent from other top solutions). What I can say though is that this, at that point of time, second place solution, would still be in gold area now. Seeing the movement of our direct competitors at that point, I would assume we would have ended somewhere in top 3 area, but who knows.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1481417,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T13:42:54.123000",
          "content": "<blockquote>\n  <p>What would you like to know exactly?</p>\n</blockquote>\n<p>training hyperparameters like learning rate, optimizer,  scheduler, drop rate, drop path rate, use of amp, etc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481621,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T15:50:01.180000",
          "content": "<p>Sure:</p>\n<p>backbone = \"eca_nfnet_l2\"</p>\n<p>epochs = 20<br>\nlr = 0.000025<br>\noptimizer = \"AdamW\"<br>\nweight_decay = 1e-4<br>\nbatch_size = 8</p>\n<p>drop_rate = 0.1<br>\ndrop_path_rate = 0.1</p>\n<p>stride = (1,2)</p>\n<p>aug = \"vflip\" + mixup (p=1.0, beta=5, max_target)</p>\n<p>tta = [[\"hflip\"], [\"vflip\"], [\"hflip\", \"vflip\"]]</p>\n<p>cuda.amp with DDP</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1481641,
          "author_name": "eagle4",
          "author_url": "",
          "post_date": "2021-08-19T16:01:56.820000",
          "content": "<p>Psi, if I may ask and since you changed the first conv stride: pretrained or from scratch ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481649,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T16:07:37.483000",
          "content": "<p>pretrained, I honestly havent seen a case where pretrained is not helpful, at least in terms of training time needed</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1481681,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-19T16:16:42.910000",
          "content": "<p>But the fisrt Conv layer is not part of backbone, right? How did you get pretrained weights for this layer? In my unstanding it`s like to use Conv layer to downsample the input image, instead of normal resize which would lose some information.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481694,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-19T16:23:26.063000",
          "content": "<p>With the first conv layer, we are referring to the first conv layer of the backbone. We modified the stride of that one. </p>\n<p>I am sorry about the image in the above post, which is indeed a bit misleading in this case.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481715,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-19T16:33:48.840000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> ,  lesson learned. Then I'm wondering if you modified the first Conv layer's stride number of backbone, are the pretained weights still match, wouldn't prompt any error/warning when training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481718,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T16:35:55.917000",
          "content": "<p>Yeah, changing the stride doesnt change anything, kernel size etc stays intact.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481734,
          "author_name": "Hao",
          "author_url": "",
          "post_date": "2021-08-19T16:40:51.390000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> , I just realized change stride didn`t change paramters number, very good trick to replace resize and does not bring any extra computation cost.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1481773,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T16:59:00.167000",
          "content": "<p>Oh it does bring huge extra computational cost. Similar to increasing image size two times, three times, etc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482062,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T21:04:47.537000",
          "content": "<p>When you do mixup with batch size 8. Do you have two dataloaders each supplying 8 images and you mixup those two batches to make one batch? Or do you have one dataloader and mixup the images within the one batch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482231,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-08-20T00:55:49.703000",
          "content": "<p>One dataloader and mix within batch</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1482237,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T01:06:30.553000",
          "content": "<p>Thanks for clarification. I notice this is what everyone at Kaggle does. However the original research paper <a href=\"https://arxiv.org/abs/1710.09412\" target=\"_blank\">here</a> (on page 3) uses 2 dataloaders. I suspect for most cases it doesn't matter, but as batch size decreases and alpha increases, it may make a difference.</p>\n<pre><code># y1, y2 should be one-hot vectors\nfor (x1, y1), (x2, y2) in zip(loader1, loader2):\n    lam = numpy.random.beta(alpha, alpha)\n    x = Variable(lam * x1 + (1. - lam) * x2)\n    y = Variable(lam * y1 + (1. - lam) * y2)\n    optimizer.zero_grad()\n    loss(net(x), y).backward()\n    optimizer.step()\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482363,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T03:36:36.290000",
          "content": "<p>2 dataloaders may negatively impact training time.  I guess none of us saw an upside compared to in batch mixup.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482548,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T06:27:01.833000",
          "content": "<p>Actually in this competition we mix across whole data, not within batch. Usually it doesnt make much difference. So we do it directly in the data loader, just mixing a random image from the full data in.</p>\n<p>I personally prefer this version, also due to reasons pointed out by Chris. But if data loading is a bottleneck (e.g., large images), I also switch to batch mixing.</p>\n<p>One thing we do here though, is that we re-normalize after mixing, so that mixed images exhibit same data characteristics as images that you actually score on.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482804,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T09:40:29.037000",
          "content": "<blockquote>\n  <p>So we do it directly in the data loader, just mixing a random image from the full data in.</p>\n</blockquote>\n<p>We tried that too, didn't see much a difference.  But we did not have your cleaned data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482812,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T09:45:29.317000",
          "content": "<p>As said, it usually makes little to no difference. We already used it before cleaned data. I think one minor key is to re-normalize afterwards. And I also believe that within batch mixing can actually be sometimes better with batchnorm models, but I have not investigated this hypothesis yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483916,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-21T01:09:38.637000",
          "content": "<p>I understand changing stride (2, 2) -&gt; (1, 2) is equal to enlarging input size, but why only magnified hight? I think magnifying both height and width is also possible: (2, 2) -&gt; (1, 1).</p>\n<p>Are there any reason that height(time) dimension is more important than width(frequency) dimension?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1484348,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-21T08:30:53.883000",
          "content": "<p>In this competition it was more important, we found (1,3) to be already really good, and (1,2) to be slightly better. Then it is also a matter of runtime, (1,1) is already really crazy for those larger models. You also saw most other competitors enlarging height, which has similar effects (but should be a bit worse as you lose information).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1484375,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-21T08:58:52.047000",
          "content": "<p>I see. It’ was computational cost trade-off.</p>\n<p>But I still wonder why height is more important than the width. In this competition, signals tend to be longer in height, and tilted signal is positive whereas vertical straight line is negative. On this condition, I think it’s natural that width information is more important to distinguish signal: low resolution in width make it hard to distinguish near-vertical line signal from vertical line noise. Did you run through the condition of (3, 1) or (2, 1), and still (1, 2) is the best?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1484386,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-21T09:19:17.207000",
          "content": "<p>Yes, changing the stride was one of the first things we tried (even before the reset) and we tested several combinations and for long time we sticked to (1,3) and only in the end reduced that to (1,2).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1485616,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2021-08-22T09:19:11.247000",
          "content": "<p>I see. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1480388,
      "author_name": "dan",
      "author_url": "",
      "post_date": "2021-08-19T02:19:17.690000",
      "content": "<p>This read makes me feel like having accomplished nothing haha</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1484663,
      "author_name": "Tien Hoang",
      "author_url": "",
      "post_date": "2021-08-21T13:39:16.717000",
      "content": "<p>Nice job. Congratulation</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1481756,
      "author_name": "nalewkoz",
      "author_url": "",
      "post_date": "2021-08-19T16:51:11.893000",
      "content": "<p>Congratulations Team Watercooled on your outstanding work. Well deserved win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1481302,
      "author_name": "Kalilur Rahman",
      "author_url": "",
      "post_date": "2021-08-19T12:25:10.587000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>  outstanding delivery and success. Well deserved!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480693,
      "author_name": "Solomon Kimunyu",
      "author_url": "",
      "post_date": "2021-08-19T06:24:28.357000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> for the big win.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480319,
      "author_name": "Yuhong Chen",
      "author_url": "",
      "post_date": "2021-08-19T00:45:40.903000",
      "content": "<p>Awesome job, congrats!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1480369,
          "author_name": "Μαριος Μιχαηλιδης KazAnova",
          "author_url": "",
          "post_date": "2021-08-19T01:59:06.323000",
          "content": "<p>Them, yes, however I have to stress that people should  NOT mislead on the forums and even worse if it comes from the <strong>organizer</strong>. I am referring to this post:</p>\n<p><a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405282\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405282</a></p>\n<p>and this post:</p>\n<p><a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405732\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/256043#1405732</a></p>\n<p><strong>where you hinted that the high score was due to (probably) the extra class</strong> in the test and \"you are not worried\"(when clearly the <code>0.853</code> score was after the \"data cleaning\" method) .</p>\n<p>And you also became conclusive about it via saying:</p>\n<blockquote>\n  <p>Sorry I spoiled it for you, didn't want people worrying about another reset .</p>\n</blockquote>\n<p><strong>Even worse that they contacted you about this</strong>. And even if your post came before them contacting you, still it does not matter because the misleading post is there - you should have made amends. </p>\n<p>I understand that you did not want another reset, but it would have been better if you hadn't said anything. </p>\n<p>Kaggle will need to look into this in the future : <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> </p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1480406,
          "author_name": "sin",
          "author_url": "",
          "post_date": "2021-08-19T02:31:38.457000",
          "content": "<p>We tried Arcface and many unsupervised methods to discover the \"unseen class\" and we couldn't find it. We also found that many of the samples in the training data are simply not learnable by a CNN or discoverable by eyes. The data cleaning is indeed brilliant and impressive.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1480446,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T03:09:31.600000",
          "content": "<p>I'm pretty sure it was not extra class, but 'cleaning', and i agree with <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1480882,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-08-19T08:02:32.777000",
          "content": "<blockquote>\n  <p>That led us to start investigating single predictions, where different models with different LB scores disagree</p>\n</blockquote>\n<p>So top1 decided it is s-shape new class and set it to label one for all examples using human eye. Well, is your pipeline automatic for any other new class with other tricky signal and do not require human eye intervention? In other words can model identify other new pattern without performance degradation? If no, i think it is similar to handy labeled test data and rule violation? <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> <a href=\"https://www.kaggle.com/yuhongc\" target=\"_blank\">@yuhongc</a> <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> what do you think guys?  </p>",
          "votes": -5,
          "replies": []
        },
        {
          "id": 1480977,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-19T08:56:03.793000",
          "content": "<p><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> I would appreciate you not making unfounded accusations and if you would read our solution post you would obviously see that we definitely did not hand label any test data.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1480995,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-19T09:04:58.573000",
          "content": "<blockquote>\n  <p>So top1 decided it is s-shape new class and set it to label one for all examples using human eye</p>\n</blockquote>\n<p>To clarify on this, of course we did NOT hand label any test data. We used a signal generator and randomly applied that (together with target=1) on the train set. </p>\n<p>Regarding your point on generalization to new signals/classes, this is obviously a huge challenge in all computer vision tasks and not really solved, yet. Supervised ML can only reliably learn what is already in the training set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481019,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-08-19T09:17:15.633000",
          "content": "<p>I said <code>decided</code> that s-shape is one-target by human. Because not all your models said that it is one-target. So you decided it is one-target for all samples not your model decide. Am I right?   </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1481026,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-08-19T09:21:11.827000",
          "content": "<p>Anyway I would like to hear the official answer from the admins for the future. I would also like to understand if the generated data is external data or not.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481027,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-08-19T09:21:20.333000",
          "content": "<p>Anyway, you guys did a great job, congratulations, I take off my hat to you!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481036,
          "author_name": "Johnny Lee",
          "author_url": "",
          "post_date": "2021-08-19T09:24:14.017000",
          "content": "<p><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> think s-shape is not target. It is not only in 0,2,4, but also in 1,3,5. I found it too, but I didn't know how to use it.</p>\n<p>Congratulations <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1497724,
      "author_name": "John D",
      "author_url": "",
      "post_date": "2021-08-31T11:45:04.533000",
      "content": "<p>Congrats =))</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480442,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-19T03:06:06.713000",
      "content": "<p>but why do stealth submitting?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1481058,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-08-19T09:39:58.473000",
      "content": "<p>Great Work!. So the keypoint is using signal generator.<br>\n<a href=\"https://github.com/bbrzycki/setigen\" target=\"_blank\">https://github.com/bbrzycki/setigen</a><br>\nIt makes model fitting on the signal than background.<br>\nThank you for sharing. I will try to review and apply this generator.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1481217,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-08-19T11:24:55.930000",
          "content": "<p>By reading the solution I wouldn't say that's the only \"keypoint\", there are many key points to take into consideration imo…</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1481441,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T13:57:30.490000",
          "content": "<p>My understanding is that magic #1 (signal generator) gets them to LB 812 first place. And magic #2 gets them to LB 968 first place. I think magic #2 was the powerhouse! </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481457,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T14:09:02.010000",
          "content": "<p>they got to 0.800 with magic 1, then found magic 2.  We (Giba) found  magic 1 too, as several other top teams.  It proves that the key was to find magic 2.</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 2183881,
      "author_name": "Dac-Thanh Van",
      "author_url": "",
      "post_date": "2023-03-16T02:11:44.583000",
      "content": "<p>This is an interesting solution. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1499027,
      "author_name": "DhirajBajracharya",
      "author_url": "",
      "post_date": "2021-09-01T11:53:06.077000",
      "content": "<p>congratulations , good job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1498422,
      "author_name": "Roma Hambar",
      "author_url": "",
      "post_date": "2021-09-01T01:25:56.003000",
      "content": "<p>Great Work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1496159,
      "author_name": "Adrian",
      "author_url": "",
      "post_date": "2021-08-30T06:39:40.117000",
      "content": "<p>Thank you for sharing this awesome solution!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1488170,
      "author_name": "VK",
      "author_url": "",
      "post_date": "2021-08-24T06:21:28.197000",
      "content": "<p>Congrat guys! Visualisation parts is nice<br>\nThanks for sharing 👍💥</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1487920,
      "author_name": "Brian_0207",
      "author_url": "",
      "post_date": "2021-08-24T00:45:07.580000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1487831,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2021-08-23T21:42:50.047000",
      "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a>, how computationally intense is magic #2? Iterating each column seems to take forever even using rapids.ai. Any trick to run this analysis faster?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1488111,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2021-08-24T05:42:29.033000",
          "content": "<ol>\n<li>compare only part of the columns, e.g. first 100 rows.</li>\n<li>load all images into RAM (but only use first 100 rows for each image and only OFF channels)</li>\n<li>use array operations on your single 60,000 x 100 x … array</li>\n</ol>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1496435,
          "author_name": "FelipeKitamura, MD, PhD",
          "author_url": "",
          "post_date": "2021-08-30T11:30:32.197000",
          "content": "<p>Thanks! Did you share the code for magic #2? Would you mind sharing it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1486856,
      "author_name": "Jackie Lee",
      "author_url": "",
      "post_date": "2021-08-23T08:57:22.930000",
      "content": "<p>awesome, thx for sharing. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1486631,
      "author_name": "Dac-Thanh Van",
      "author_url": "",
      "post_date": "2021-08-23T05:41:56.417000",
      "content": "<p>Congratulations, guys! This is one of the most wonderful works on signal processing I have ever read. Clean and effective. Good luck with your future works!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1486080,
      "author_name": "bahareh keshavarz",
      "author_url": "",
      "post_date": "2021-08-22T16:38:07.927000",
      "content": "<p>congratulations! huge achievement.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1485394,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-22T04:19:47.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1484382,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-21T09:14:13.007000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1483308,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-20T15:13:20.323000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1483138,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-20T13:16:09.453000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1482172,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T23:06:59.933000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1482064,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T21:05:24.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1481796,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T17:06:31.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1481786,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T17:01:17.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1481189,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T11:13:02.313000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480461,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T03:20:46.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480426,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T02:54:44.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480377,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T02:06:58.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480339,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T01:25:26.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480314,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T00:37:24.670000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480309,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T00:34:29.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1495845,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-29T20:40:25.617000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1495842,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-29T20:39:47.277000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1495835,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-29T20:34:54.460000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1480503,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T03:53:56.547000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1498221,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-31T19:03:58.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1480382,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-19T02:13:47.777000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1480304": "Thanks to Kaggle and Berkeley SETI Research Center for this interesting competition. In the following, we want to give a summary of the winning solution of Team Watercooled. As always, thanks to all team members contributing equally to the solution.\n@philippsinger @christofhenkel @ilu000\n\n# Summary\nOur solution is based on large state-of-the-art classification models that were fine-tuned for this specific task. We pre-processed images by cleaning the backgrounds to boost signal to noise ratio. During training, we employed heavy augmentation in the form of Mixup. For faster training iterations, we only used the ON-channels. We only rely on provided competition data and do not utilize any external data. We additionally augment the training data with an extra signal that only appears in test files -- an “s-shape” signal -- by using a randomized signal generator. \n\n# Cross-Validation and Preprocessing\nThe CV setup for this competition was quite straightforward, using a 5 fold split on the new training data. We also found that the old leaky data was still giving us a significant boost if used for training, so we added the full old train and old test data to each training fold. To reduce the impact of leakage in the training as much as possible, we cast to float32 and subsequently applied a channel normalization after loading an image. \n\n# CV/LB gap\nAs most participants, we experienced major differences between CV and LB scores early on in the competition. While this is in theory not necessarily always an issue, LB might just be harder (and it certainly is), we did not observe clear correlation between higher CV and higher LB. This led us to explore different ways of understanding these differences and trying to account for them. Our big gap to second place on LB appears to mostly be based on these insights and we elaborate them next.\n\n# The magic #1\nAs mentioned, we observed that certain models exhibit significant different LB scores even though their CV scores were similar. That led us to start investigating single predictions, where different models with different LB scores disagree. After comparing several models, we observed that better LB models had much higher probabilities for images with an “s-shape” signal as shown below. \n\n![s-shape in test](https://i.imgur.com/9mE6ptH.png)\n\nThat type of signal only appears in the test set and some models were able to identify the signal as atypical, while others were not. It was later also confirmed by hosts that test data contains additional type(s) of signal. To better predict these signals and better generalize to the test data, we added a signal generator (code base adjusted from https://github.com/bbrzycki/setigen) which adds those signals to the training set. We tuned the probability (p=0.01) and randomized the shape and signal to noise ratio for maximum LB score while making sure that CV score stays constant. An example of such an injected signal is shown in the plot below. \n\n![s-shape injected](https://i.imgur.com/SCsiLeo.png)\n\n# The magic #2\nMany noticed a large jump in our public leaderboard score a few weeks before the competition ended. We were already at a decently high score close to 0.800, which was 2nd place on LB at that time, and we were quite confident to further improve our solution as we were only subbing very simple single-fold blends and haven’t even turned our attention to further advancement like pseudo tagging which appeared to hold much promise in this competition. To further improve our solution, we continued to investigate the large CV/LB gap that was still present after injecting the s-shaped signal. \n\nOne obvious aspect in this competition is that the background between train and test data is very different. This is not only imminent from visual inspection or simple binary classifiers that can easily distinguish between train and test, but also from early pseudo tagging experiments we conducted. Here, we only added very certain target=1 samples from test to train, and while CV looked good, our models suddenly only predicted target=1 for the whole test, meaning that the models perfectly learned that target=1 is always from test. \n\nSo we started to speculate that the models sometimes focus too much on the background rather than the signal. In the plot below, we exemplary show the histogram of our output logits grouped by train similarity. Here, train similarity is just based on image data (mean, std, unique counts, min, max), but it was clear that the models are more certain if they “know the background”. Thus, we explored methods to reduce the impact of the background noise or reduce overfitting on the background noise.\n\n![train_like vs not_train_like](https://i.imgur.com/V9BxbeT.png)\n\nFurther investigations revealed that there even was significant overlap from one image to another. This means that the model could also potentially overfit to single backgrounds if it only sees target=0 or target=1 for that background. As a result of this investigation, we decided to attempt to remove some common background artifacts from the data (i.e. clean the data), to increase the signal and reduce the noise and overfit potential. \n\nHowever, It is far from trivial to utilize this property due to image-wise channel normalization. One cannot just compare raw values from one image to another. To allow for comparison between images, we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns. Due to rounding errors and possible signals in the columns, we received column distances greater than 0. We utilized the overlaps to increase signal to noise ratio, by replacing the original data with the difference of the normed regions between multiple matched samples. This induces shadowing (difference is much smaller than 0) from signals that are present in the matched samples, which our models were able to distinguish from actual signals. In the plot below, the difference of the normalized overlapping region between two samples is shown. The yellow needle (values > 0) originates from sample 1, and the blue needle (values < 0) originates from the matched sample. \n\n![match shadow](https://i.imgur.com/XxC4lDI.png)\n\nBased on this process, we attempted to clean all images where possible; see the following examples:\n\nOriginal sample with signal:\n![Original sample with signal](https://i.imgur.com/1y9kpMs.png)\nCleaned sample with signal:\n![Cleaned sample with signal](https://i.imgur.com/CemLefI.png)\n\nOriginal sample without signal:\n![Original sample without signal](https://i.imgur.com/VSE7nxC.png)\nCleaned sample without signal:\n![Cleaned sample without signal](https://i.imgur.com/STPDRyi.png)\n\nOriginally, when we started cleaning the data, we just quickly attempted to use existing models trained on the uncleaned data and augment inference with the cleaned data. Immediately, we saw big boosts on LB giving us the 0.853 score. As this was already a significant difference to the next highest score on LB, we contacted Kaggle to elaborate on our approach and to confirm that we can continue using it. After that, we started to directly utilize the cleaned data in our modeling approach, to further boost our performance significantly.\n\n# Models\nAs we were very confident about our solution, we decided against submitting huge blends or usage of pseudo labels, thus rendering our solution much better suited for actual use in the field. Our best final submission is thus based on a single fit on the full data.\n\nOur model architecture and training rechime works both for uncleaned and cleaned data (i.e. nearly duplicating training data). Consequently, we decided to train the data on all data available:\n- Old train (uncleaned & cleaned)\n- Old test (uncleaned & cleaned)\n- New train (uncleaned & cleaned)\n\nFor inference, we always pick the cleaned image, if available, and otherwise the uncleaned image.\n\nEarly, we discovered that norm-free models exhibit slightly better CV/LB correlations, probably also due to the lack of batch norm that can be influenced by data distributional differences. Our final best model used an eca_nfnet_l2 backbone and trainable GeM pooling. \n\nModel and resolution wise we saw on CV that the bigger the better. Hence, as mentioned earlier, we feed the full image with only concatenated ON channels in and do not do any resizing. \nTo further “enlarge” the images we change the model stride of the first conv layer to (1,2), which is a common trick used in computer vision competitions to basically further enlarge image resolution. That allows the model to train on a higher resolution through the backbone. Christof and Philipp used the same trick in ALASKA2 Steganalysis competition, motivated by the very similar data setup (find hidden data in an image). Training time of this “high resolution” model is about 15h on 8xV100 using DDP.\n\n![model](https://i.imgur.com/WdbhFTB.png)\n\nWe train the model for 20 epochs, and apart from random vertical flip, we also employ mixup. We randomly mix two images with drawing from the beta distribution with alpha=beta=5, so mostly equal blend. We always take the maximum of both labels as the final label should always be target=1 as soon as there is some form of a signal in the data. We only mix between uncleaned and cleaned images respectively. What we also found helpful, is to re-normalize the images after mixing, to keep the nature of the data intact for inference. We also employ 4xTTA (regular, hflip, vflip, hflip+vflip).\n\nOur final single fold validation score is 0.976 with a public LB score of 0.965 and private LB score of 0.964 (nearly closing the cv/lb gap). The fullfit (trained on all data with the same parameters) scores 0.967 public LB and 0.967 private LB.\n\n# Outlook\nWe already had a very good solution with a public LB score close to 0.800 before starting to additionally clean the backgrounds in images. We quickly realized that this has a major impact on the solution and focussed our efforts on utilizing the information to its fullest extent. Thus, bound by long training times, we did not explore every approach: In the early stage of the competition, we also trained models with a kind of self attention based on the OFF-channel activations. While not yielding better scores on its own, this proved to be valuable when used in blends as it seems to add diversity. One other thing we noticed is that pseudo labeling has a positive impact as it may help with the train/test domain shift. We also explored additional techniques and research to tackle domain shift.\n",
    "1481215": "Just one more note on magic #2 for further elaboration. \n\nImagine the following simplistic use case:\nYou have a set of images with certain backgrounds (white, green, striped, constant, etc.) and then you have black dots on those images and you want your model to identify if an image contains a black spot (anomaly). Unfortunately, you have some backgrounds that always exhibit a black spot, or maybe never exhibit one. If you are unlucky, your model might decide to overfit on the actual background, instead of trying to identify the raw signal of the black spot. And here in this competition we observed quite obvious overfit on the background, and a challenge was to steer the models to learn the actual patterns. \n\nOne thing most people did was to do mixup, so mixing different backgrounds together and having the signals in different intensity. One other thing was to utilize pseudo tagging, that then also better learns to adjust to the new background (distribution) of the test data. But the best solution is to process the images in a way that they increase the signal to noise ratio, which we did by cleaning common backgrounds among images. In the example above, it would be obviously better to completely remove all background, and just keep the actual signal.\n\nI do personally not see how this is necessarily only limited to this data and problem. Of course, the data is synthetic here, so the effects are emphasized, but I have observed similar effects before, in other problems, competitions, and projects, but have never so thoroughly thought about it like here. There are many use cases where this is an imminent issue, such as fault prediction in machinery, anomaly prediction, steganalysis, or even chest x-ray prediction where often models prefer to overfit on certain device characteristics.\n\nWill the same technique applied here work for all problems? Probably not, but the thought process can be the same. So I would encourage you to see it as a learning, at least that's what I am doing, and try to approach these issues in the future in different projects. I can imagine, that clever modeling approaches can also attempt to deal with these issues directly, for example via attention like 3rd place used. There is also plenty of research in the area of domain adaption, where I have seen little being applied on Kaggle, and I have only started to look into it:\nhttps://paperswithcode.com/task/domain-adaptation\n",
    "1480415": "Huge congratulations! I've never seen a win with such a big lead! \n\nIt never occurred to me that the same background image, by itself and the image into which it was injected with a signal, would be in the training set at the same time.\n\nSuch a dataset does not occur in industry or real world, so this is completely beyond my imagination.\n\nI feel shame on sharing our solution this time so I'll just skip it. 😅",
    "1480321": "Congratulations, we had found the magic#1 but not the magic#2 even if we've spent some time on pre-processing/cleaning.",
    "1480422": "@ilu000 \nCongratulations & thanks for sharing the solution.\n\nI'm not sure of the process of magic #2.\n\n> we normalized per column and calculated the mean pixel difference of the first column of each image to all other images and columns.\n\nDoes it mean the process below?\n1. normalize column vector for each images\n2. calculate pixel difference from the first column per each image\n3. take the mean of 2 (column-wise), which give us the one feature vector\n4. using feature vector given on 3, match the identical background image in the all image space (using kNN etc.)\n5. take pixel difference from the matched image, and clear the background\n\nExample: process for 3x3, 1-channel image\n```\n# original image\nimg1 == [\n  [1, 4, 3],\n  [2, 3, 3],\n  [3, 2, 3],\n]\n\n# after column-wise normalization\nimg1 == [\n  [-1, 1, 0],\n  [0, 0, 0],\n  [1, -1, 0],\n]\n\n# after taking column-wise difference from the first column\nimg1 == [\n  [0, 2, 1],\n  [0, 0, 0],\n  [0, -2, -1],\n]\n\n# after calculating mean pixel difference\nimg1 == [\n  [1],\n  [0],\n  [-1],\n]\n```",
    "1480479": "Congratulations on such huge lead win @ilu000, your 2 magics are so brilliant, I thought you must implement magic #1 to find extra class/pattern that only exist in test data, but your magic #2 is far beyond my imagination, I guess it`s difficult for me to replicate it even after read your solution 😂. \nBTW, \"15h on 8xV100 using DDP\", what does DDP mean here, something like AMP(Auto Mixed Precision)? ",
    "1481571": "Brilliant solution team! Congratulations on 1st place and your amazing lead over 2nd.\n\nApplying column normalization is brilliant. I knew that the host divided each image channel by different number and subtracted from each image channel a different number but i couldn't think how to correct this imbalanced normalization. Fantastic work with column normalization.\n\nTo detect anomalies we need control images. It was brilliant to use one \"on\" cadence image as control for another \"on\" cadence image because both \"on\" cadence images point to the same point in space. Great idea!",
    "1480982": "Would you mind disclosing the Id of your example \"Original sample with signal\" for \"Magic #2\" ?",
    "1480334": "Congrats, this is an amazing solution. I was following from afar and it reminded me of the [PLAsTiCC competition](https://www.kaggle.com/c/PLAsTiCC-2018) (coincidentally, also an astronomy setting) - there, the biggest challenge was also CV/LB gap and the test set had an extra class not found in the training set. The winner in that competition was an astronomer with limited ML experience who used a clever idea to augment the training data to match the test data and won solo with a single LGB against teams using much more sophisticated modelling techniques. I wasn't too familiar with the ins and outs of this competition but after noticing similarities to PLAsTiCC I had a feeling a data approach would make the difference here as well. \n\nThat being said, this solution is even more impressive IMO. It's clear from the write up just how much thought was put into each component. Winning with ~0.97 AUC when second place is ~0.81 is unprecedented as far as I'm aware.",
    "1480311": "Thank you for the detailed solution! This is an absolutely beautiful data-centric approach to this competition. Looking forward to learning a lot more from the Watercooled team in the future!",
    "1481286": "Congrats on the clever use of all data available.  I would love to see more details on how you trained eca-nfnet-l2.\n\nLet me play  the devil advocate for now.\n\nWould you have found any magic if test set was hidden as in a code competition? magic 1 looks impossible to find without looking at test images, but what about magic 2?\n\nDo you have an idea of where you would end without magic 2?  I know it is hard to answer because if you had not found it you probably would have spent time on other ways to clean data.",
    "1480388": "This read makes me feel like having accomplished nothing haha",
    "1484663": "Nice job. Congratulation",
    "1481756": "Congratulations Team Watercooled on your outstanding work. Well deserved win!",
    "1481302": "Congratulations @philippsinger @christofhenkel and @ilu000  outstanding delivery and success. Well deserved!",
    "1480693": "Congratulations @philippsinger @christofhenkel @ilu000 for the big win.",
    "1480319": "Awesome job, congrats!",
    "1497724": "Congrats =))",
    "1480442": "but why do stealth submitting?",
    "1481058": "Great Work!. So the keypoint is using signal generator.\nhttps://github.com/bbrzycki/setigen\nIt makes model fitting on the signal than background.\nThank you for sharing. I will try to review and apply this generator.",
    "2183881": "This is an interesting solution. ",
    "1499027": "congratulations , good job",
    "1498422": "Great Work",
    "1496159": "Thank you for sharing this awesome solution!",
    "1488170": "Congrat guys! Visualisation parts is nice\nThanks for sharing 👍💥",
    "1487920": "Congratulations!",
    "1487831": "@ilu000, how computationally intense is magic #2? Iterating each column seems to take forever even using rapids.ai. Any trick to run this analysis faster?",
    "1486856": "awesome, thx for sharing. ",
    "1486631": "Congratulations, guys! This is one of the most wonderful works on signal processing I have ever read. Clean and effective. Good luck with your future works!",
    "1486080": "congratulations! huge achievement.",
    "1485394": "Amazing Data Analysis you did .",
    "1484382": "Nice work. And very interesting to read. Thank you from Austria",
    "1483308": "Amazing solution Watercooled team, huge congratulations!! Could you share how you measure signal to noise ratio in the images?",
    "1483138": "@llu There is so much to grasp and understand . Clealy, you worked hard. Grasping one by one ",
    "1482172": "Congratulations and thanks for sharing! Very interesting!",
    "1482064": "Congratulations! ",
    "1481796": "Congratulations, and thanks for sharing.",
    "1481786": "Congratulations!",
    "1481189": "thx for sharing!\n\n",
    "1480461": "Marvelous! \nthank you for the detailed explanations.",
    "1480426": "Congratulations! Your work is so brilliant! ",
    "1480377": "brilliant!!!",
    "1480339": "What a solution!\n天秀！",
    "1480314": "Incredibly outstanding!👋👋👋Respect!",
    "1480309": "thx for sharing!",
    "1495845": "",
    "1495842": "",
    "1495835": "",
    "1480503": "",
    "1498221": "Thanks for the expalnation\n\n",
    "1480382": "Thank you for the detailed solution!  "
  }
}