{
  "id": 266534,
  "title": "86th - Bronze Medal - with Silver Gold Magic Explained",
  "url": "/competitions/seti-breakthrough-listen/discussion/266534",
  "author_name": "Chris Deotte",
  "post_date": "2021-08-19T12:58:51.798000",
  "votes": 54,
  "comment_count": 19,
  "views": 0,
  "content": "<h1>Exciting Competition!</h1>\n<p>What an interesting competition and difficult puzzle! Thank you Kaggle and SETI for hosting. I started this competition last Wednesday and worked hard every day. In this time, I closed CV LB gap from 0.125 to 0.104. I would have liked to have more time and continue the puzzle of closing the CV LB gap. Below is my journey over the past 7 days.</p>\n<h1>My Model - CV 0.890 - LB 0.765 - GAP = 0.125 (Days 1-3)</h1>\n<p>Last Wednesday, Thursday, and Friday, I built dozens of my own models without reading public discussions nor notebooks. I was excited when my best model achieved CV 0.890 but surprised and disappointed when I submitted and received LB 0.765 🙁This was a 0.125 gap between CV and LB. Something is going on!</p>\n<p>My best model used only \"on\" cadence in a spatial stack resized to 768x768 and fed into a EfficientNetB4 backbone. The only augmentation was horizontal flip and vertical flip. It used cosine learning schedule and 15 epochs.</p>\n<h1>Alternative Ideas (Days 3-4)</h1>\n<p>At this point, I thought the gap was caused by new pattens in test data. So, on Friday and Saturday, i starting building more creative models that used more than just classify train patterns.</p>\n<ul>\n<li>Use \"off\" cadence train images as more <code>target=0</code>. Increased CV 0.003 but decreased LB 0.010</li>\n<li>Use \"off\" cadence test images as more <code>target=0</code>. Did not affect CV but decreased LB like 0.050!</li>\n<li>Use test pseudo labels. Decrease LB like 0.030!</li>\n<li>Train model to classify \"train\" image versus \"test\" image. So discard target column and use <code>test image as target=1</code> and <code>train image as target=0</code>. The hope was that this model would find any new patterns in test data because that would be the difference between train and test. To my surprise the model achieved CV 0.99 but LB was terrible 🙁</li>\n<li>Train model to classify \"on\" cadence <code>img[::2]</code> versus \"off\" cadence <code>img[1::2]</code>. So discard target column and use <code>cadence on as target=1</code> and <code>cadence off as target=0</code>. The advantage here is that we can train directly on test data without train data!</li>\n<li>Sliding window over \"on\" cadence and compute cosine similarity with sliding window over \"off\" cadence. Take minimum cosine similarity value over all crops of \"off\" cadence. Then take maximum (of these minimums) over all crops of \"on\" cadence. One advantage here is that we can apply this to test data without using train data!</li>\n</ul>\n<p>My favorite model was the last which achieved LB 0.550 and did not use any train data. It simply compared crops of test image \"on\" cadence with crops of test image \"off\" cadence and searched for dissimilarity.</p>\n<h1>Public Notebook - CV 884 - LB 0.779 - GAP = 0.105 (Days 5-6)</h1>\n<p>By Saturday, i was frustrated that I couldn't reach Bronze zone. How was everyone doing it?  </p>\n<p>I began reading discussions and notebooks. I found <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> amazing \"SETI - Learned Image Resizing\" notebook <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing?scriptVersionId=68644223\" target=\"_blank\">here</a>. What impressed me most was that his <strong>GAP was only 0.102</strong>! His CV was 0.846 and his LB was 0.744. This was a smaller gap than my model. </p>\n<p>Without understanding why, I quickly ran his notebook locally with larger models and larger image sizes to secure bronze position. The following table are results for fold 0 only. </p>\n<p>Note that Tucker's notebook has a resize module before the EfficientNet, so we put <code>image_in</code> into the resize module and <code>image_out</code> comes out. We then feed <code>image_out</code> into the EfficientNet backbone:</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Image In</th>\n<th>Image Out</th>\n<th>CV - fold 0</th>\n<th>LB - fold 0</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EB3</td>\n<td>512</td>\n<td>256</td>\n<td>0.859</td>\n<td>0.764</td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>640</td>\n<td>320</td>\n<td>0.870</td>\n<td>0.768</td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>768</td>\n<td>384</td>\n<td>0.870</td>\n<td></td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>orig</td>\n<td>448</td>\n<td>0.872</td>\n<td></td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>640</td>\n<td>320</td>\n<td>0.873</td>\n<td>0.770</td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>768</td>\n<td>384</td>\n<td>0.872</td>\n<td></td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>orig</td>\n<td>448</td>\n<td>0.869</td>\n<td></td>\n</tr>\n<tr>\n<td>EB6</td>\n<td>640</td>\n<td>320</td>\n<td>0.872</td>\n<td>0.770</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td>0.883</td>\n<td>0.777</td>\n</tr>\n</tbody>\n</table>\n<p>Then 5-Fold was LB 0.778 and post process got LB 0.779</p>\n<h1>Nvidia 4xV100 GPUs</h1>\n<p>For the past few days, Nvidia V100 GPUs were running constantly to train the models in the table above. I am learning PyTorch and discovered a great trick <a href=\"https://pytorch.org/tutorials/beginner/blitz/data_parallel_tutorial.html#create-model-and-dataparallel\" target=\"_blank\">here</a>. Using Tucker's notebook, you can add 1 line of code to train PyTorch with multiple GPUs (for large backbones and batch sizes). Just replace the one line <code>model.to(device)</code> with the three lines:</p>\n<pre><code>import torch.nn as nn\nmodel = nn.DataParallel(model)\nmodel.to(device)\n</code></pre>\n<p>That's it! With 1 line of code, we can train larger models and larger batch sizes by using multiple Nvidia GPU. </p>\n<h1>The Secret of Timm Backbones</h1>\n<p>Many people (including <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>) were confused why they could not train efficientnet B4 with Tucker's notebook. That is because timm does not contain an imagenet pretrained efficientnet B4. (So if you use <code>model_name='efficientnet_b4'</code> you train with a not-pretrained effnet) </p>\n<p>Timm only contains pretrains of EB4, 5, 6, 7 of the TF converted weights, <code>tf_efficientnet_b4</code>, <code>tf_efficientnet_b4_ns</code>, <code>tf_efficientnet_b4_ap</code>. The first is original efficient net, the next two are pretrained with noisy student and adv prop.  To view the full list of available timm pretrained backbones, type</p>\n<pre><code>import timm\nfrom pprint import pprint\nmodel_names = timm.list_models(pretrained=True)\npprint(model_names)\n</code></pre>\n<h1>Model Difference - Mysterious \"S\" Shape! (Day 7)</h1>\n<p>After securing bronze medal position, I began to investigate the CV LB gap again. Four hours before competition deadline, I plotted the difference between my model and public notebook model. I sorted by test images where the difference between my model prediction and public model prediction were greatest. To my surprise, I found a mysterious \"S\" shaped test pattern!</p>\n<p>The image on the left is original comp test data \"on\" cadence vertical stack. With time on x axis and frequency on y axis. The second image is pixel histogram. The third image is \"on\" cadence with emboss, blur, and adaptive histogram equalization. (Filter code shown below) The fourth image (right image) is the \"off\" cadence.</p>\n<pre><code>import cv2\nclahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\n# FILTERS FOR BETTER IMAGE DISPLAY\nimg = np.vstack(cadence[::2,])\nimg = img[1:,1:] - img[:-1,:-1] #emboss\nimg -= np.min(img)\nimg /= np.max(img)\nimg = (img*255).astype('uint8')\nimg = cv2.GaussianBlur(img,(5,5),0)\nimg = clahe.apply(img)\n</code></pre>\n<p>Above the left image we see the public notebook prediction value. And above the third image we see my model's prediction value.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex1b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex2b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex3b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex4b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex5b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex6b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex7b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex8b.png\" alt=\"\"></p>\n<h1>Post Process - LB +0.001 - GAP = 0.104 (Last 4 hours!)</h1>\n<p>With only 4 hours remaining, I trained a model on the images that my model and public notebook model disagreed most on. I then inferred the entire test data. For all images with <code>p&gt;0.6</code>, I increased their submission prediction. This post processed closed the CV LB gap by <code>0.001</code> and increased LB by <code>0.001</code>. Thus there is something special about these \"S\" patterns!</p>\n<h1>UPDATE - Mixup is the Silver Medal Magic!</h1>\n<p>After the comp ended, i have been continuing with experiments to determine why my original model's CV LB is large. It appears that mixup is very important to teach your model to generalize to test data which is different than the train data. It both closes the CV LB gap and boosts CV LB. </p>\n<p>Using 1 fold of EB4 and image size 768x768 with only Hflip, Vflip, and Mixup (alpha=3, max target). One can achieve Silver medal with public LB 0.786 and private LB 0.781. Notebook <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a>. (Also important is large image size and large backbone).<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"></p>\n<h1>UPDATE - \"S-shape\" and \"Remove Noise\" is Gold Medal Magic!</h1>\n<p>To climb from Silver medal to Gold medal, we need to detect \"s-shapes\" and remove noise. The first place winning solution <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> explains how. Additionally, repeatedly using test pseudo labels also helps climb to top of Gold Medals! (and using old data before reset can help too).</p>\n<h1>UPDATE - Grad Cam</h1>\n<p>I posted a discussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/268314\" target=\"_blank\">here</a> about grad cam and a notebook <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a> using grad cam. Using grad cam can help us understand what makes a particular image a <code>target=1</code>. For example, our model can make a test prediction and then circle what it thinks is causing a <code>target=1</code> prediction. The image below is 100% generated by code and no human labeling!<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1481348,
      "postDate": "2021-08-19T12:58:51.797Z",
      "content": "<h1>Exciting Competition!</h1>\n<p>What an interesting competition and difficult puzzle! Thank you Kaggle and SETI for hosting. I started this competition last Wednesday and worked hard every day. In this time, I closed CV LB gap from 0.125 to 0.104. I would have liked to have more time and continue the puzzle of closing the CV LB gap. Below is my journey over the past 7 days.</p>\n<h1>My Model - CV 0.890 - LB 0.765 - GAP = 0.125 (Days 1-3)</h1>\n<p>Last Wednesday, Thursday, and Friday, I built dozens of my own models without reading public discussions nor notebooks. I was excited when my best model achieved CV 0.890 but surprised and disappointed when I submitted and received LB 0.765 🙁This was a 0.125 gap between CV and LB. Something is going on!</p>\n<p>My best model used only \"on\" cadence in a spatial stack resized to 768x768 and fed into a EfficientNetB4 backbone. The only augmentation was horizontal flip and vertical flip. It used cosine learning schedule and 15 epochs.</p>\n<h1>Alternative Ideas (Days 3-4)</h1>\n<p>At this point, I thought the gap was caused by new pattens in test data. So, on Friday and Saturday, i starting building more creative models that used more than just classify train patterns.</p>\n<ul>\n<li>Use \"off\" cadence train images as more <code>target=0</code>. Increased CV 0.003 but decreased LB 0.010</li>\n<li>Use \"off\" cadence test images as more <code>target=0</code>. Did not affect CV but decreased LB like 0.050!</li>\n<li>Use test pseudo labels. Decrease LB like 0.030!</li>\n<li>Train model to classify \"train\" image versus \"test\" image. So discard target column and use <code>test image as target=1</code> and <code>train image as target=0</code>. The hope was that this model would find any new patterns in test data because that would be the difference between train and test. To my surprise the model achieved CV 0.99 but LB was terrible 🙁</li>\n<li>Train model to classify \"on\" cadence <code>img[::2]</code> versus \"off\" cadence <code>img[1::2]</code>. So discard target column and use <code>cadence on as target=1</code> and <code>cadence off as target=0</code>. The advantage here is that we can train directly on test data without train data!</li>\n<li>Sliding window over \"on\" cadence and compute cosine similarity with sliding window over \"off\" cadence. Take minimum cosine similarity value over all crops of \"off\" cadence. Then take maximum (of these minimums) over all crops of \"on\" cadence. One advantage here is that we can apply this to test data without using train data!</li>\n</ul>\n<p>My favorite model was the last which achieved LB 0.550 and did not use any train data. It simply compared crops of test image \"on\" cadence with crops of test image \"off\" cadence and searched for dissimilarity.</p>\n<h1>Public Notebook - CV 884 - LB 0.779 - GAP = 0.105 (Days 5-6)</h1>\n<p>By Saturday, i was frustrated that I couldn't reach Bronze zone. How was everyone doing it?  </p>\n<p>I began reading discussions and notebooks. I found <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> amazing \"SETI - Learned Image Resizing\" notebook <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing?scriptVersionId=68644223\" target=\"_blank\">here</a>. What impressed me most was that his <strong>GAP was only 0.102</strong>! His CV was 0.846 and his LB was 0.744. This was a smaller gap than my model. </p>\n<p>Without understanding why, I quickly ran his notebook locally with larger models and larger image sizes to secure bronze position. The following table are results for fold 0 only. </p>\n<p>Note that Tucker's notebook has a resize module before the EfficientNet, so we put <code>image_in</code> into the resize module and <code>image_out</code> comes out. We then feed <code>image_out</code> into the EfficientNet backbone:</p>\n<table>\n<thead>\n<tr>\n<th>Backbone</th>\n<th>Image In</th>\n<th>Image Out</th>\n<th>CV - fold 0</th>\n<th>LB - fold 0</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>EB3</td>\n<td>512</td>\n<td>256</td>\n<td>0.859</td>\n<td>0.764</td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>640</td>\n<td>320</td>\n<td>0.870</td>\n<td>0.768</td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>768</td>\n<td>384</td>\n<td>0.870</td>\n<td></td>\n</tr>\n<tr>\n<td>EB3</td>\n<td>orig</td>\n<td>448</td>\n<td>0.872</td>\n<td></td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>640</td>\n<td>320</td>\n<td>0.873</td>\n<td>0.770</td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>768</td>\n<td>384</td>\n<td>0.872</td>\n<td></td>\n</tr>\n<tr>\n<td>EB4</td>\n<td>orig</td>\n<td>448</td>\n<td>0.869</td>\n<td></td>\n</tr>\n<tr>\n<td>EB6</td>\n<td>640</td>\n<td>320</td>\n<td>0.872</td>\n<td>0.770</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td></td>\n<td></td>\n<td>0.883</td>\n<td>0.777</td>\n</tr>\n</tbody>\n</table>\n<p>Then 5-Fold was LB 0.778 and post process got LB 0.779</p>\n<h1>Nvidia 4xV100 GPUs</h1>\n<p>For the past few days, Nvidia V100 GPUs were running constantly to train the models in the table above. I am learning PyTorch and discovered a great trick <a href=\"https://pytorch.org/tutorials/beginner/blitz/data_parallel_tutorial.html#create-model-and-dataparallel\" target=\"_blank\">here</a>. Using Tucker's notebook, you can add 1 line of code to train PyTorch with multiple GPUs (for large backbones and batch sizes). Just replace the one line <code>model.to(device)</code> with the three lines:</p>\n<pre><code>import torch.nn as nn\nmodel = nn.DataParallel(model)\nmodel.to(device)\n</code></pre>\n<p>That's it! With 1 line of code, we can train larger models and larger batch sizes by using multiple Nvidia GPU. </p>\n<h1>The Secret of Timm Backbones</h1>\n<p>Many people (including <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a>) were confused why they could not train efficientnet B4 with Tucker's notebook. That is because timm does not contain an imagenet pretrained efficientnet B4. (So if you use <code>model_name='efficientnet_b4'</code> you train with a not-pretrained effnet) </p>\n<p>Timm only contains pretrains of EB4, 5, 6, 7 of the TF converted weights, <code>tf_efficientnet_b4</code>, <code>tf_efficientnet_b4_ns</code>, <code>tf_efficientnet_b4_ap</code>. The first is original efficient net, the next two are pretrained with noisy student and adv prop.  To view the full list of available timm pretrained backbones, type</p>\n<pre><code>import timm\nfrom pprint import pprint\nmodel_names = timm.list_models(pretrained=True)\npprint(model_names)\n</code></pre>\n<h1>Model Difference - Mysterious \"S\" Shape! (Day 7)</h1>\n<p>After securing bronze medal position, I began to investigate the CV LB gap again. Four hours before competition deadline, I plotted the difference between my model and public notebook model. I sorted by test images where the difference between my model prediction and public model prediction were greatest. To my surprise, I found a mysterious \"S\" shaped test pattern!</p>\n<p>The image on the left is original comp test data \"on\" cadence vertical stack. With time on x axis and frequency on y axis. The second image is pixel histogram. The third image is \"on\" cadence with emboss, blur, and adaptive histogram equalization. (Filter code shown below) The fourth image (right image) is the \"off\" cadence.</p>\n<pre><code>import cv2\nclahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\n# FILTERS FOR BETTER IMAGE DISPLAY\nimg = np.vstack(cadence[::2,])\nimg = img[1:,1:] - img[:-1,:-1] #emboss\nimg -= np.min(img)\nimg /= np.max(img)\nimg = (img*255).astype('uint8')\nimg = cv2.GaussianBlur(img,(5,5),0)\nimg = clahe.apply(img)\n</code></pre>\n<p>Above the left image we see the public notebook prediction value. And above the third image we see my model's prediction value.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex1b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex2b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex3b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex4b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex5b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex6b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex7b.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex8b.png\" alt=\"\"></p>\n<h1>Post Process - LB +0.001 - GAP = 0.104 (Last 4 hours!)</h1>\n<p>With only 4 hours remaining, I trained a model on the images that my model and public notebook model disagreed most on. I then inferred the entire test data. For all images with <code>p&gt;0.6</code>, I increased their submission prediction. This post processed closed the CV LB gap by <code>0.001</code> and increased LB by <code>0.001</code>. Thus there is something special about these \"S\" patterns!</p>\n<h1>UPDATE - Mixup is the Silver Medal Magic!</h1>\n<p>After the comp ended, i have been continuing with experiments to determine why my original model's CV LB is large. It appears that mixup is very important to teach your model to generalize to test data which is different than the train data. It both closes the CV LB gap and boosts CV LB. </p>\n<p>Using 1 fold of EB4 and image size 768x768 with only Hflip, Vflip, and Mixup (alpha=3, max target). One can achieve Silver medal with public LB 0.786 and private LB 0.781. Notebook <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a>. (Also important is large image size and large backbone).<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"></p>\n<h1>UPDATE - \"S-shape\" and \"Remove Noise\" is Gold Medal Magic!</h1>\n<p>To climb from Silver medal to Gold medal, we need to detect \"s-shapes\" and remove noise. The first place winning solution <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> explains how. Additionally, repeatedly using test pseudo labels also helps climb to top of Gold Medals! (and using old data before reset can help too).</p>\n<h1>UPDATE - Grad Cam</h1>\n<p>I posted a discussion <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/268314\" target=\"_blank\">here</a> about grad cam and a notebook <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a> using grad cam. Using grad cam can help us understand what makes a particular image a <code>target=1</code>. For example, our model can make a test prediction and then circle what it thinks is causing a <code>target=1</code> prediction. The image below is 100% generated by code and no human labeling!<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png\" alt=\"\"></p>",
      "rawMarkdown": "# Exciting Competition!\nWhat an interesting competition and difficult puzzle! Thank you Kaggle and SETI for hosting. I started this competition last Wednesday and worked hard every day. In this time, I closed CV LB gap from 0.125 to 0.104. I would have liked to have more time and continue the puzzle of closing the CV LB gap. Below is my journey over the past 7 days.\n\n# My Model - CV 0.890 - LB 0.765 - GAP = 0.125 (Days 1-3)\nLast Wednesday, Thursday, and Friday, I built dozens of my own models without reading public discussions nor notebooks. I was excited when my best model achieved CV 0.890 but surprised and disappointed when I submitted and received LB 0.765 🙁This was a 0.125 gap between CV and LB. Something is going on!\n\nMy best model used only \"on\" cadence in a spatial stack resized to 768x768 and fed into a EfficientNetB4 backbone. The only augmentation was horizontal flip and vertical flip. It used cosine learning schedule and 15 epochs.\n\n# Alternative Ideas (Days 3-4)\nAt this point, I thought the gap was caused by new pattens in test data. So, on Friday and Saturday, i starting building more creative models that used more than just classify train patterns.\n* Use \"off\" cadence train images as more `target=0`. Increased CV 0.003 but decreased LB 0.010\n* Use \"off\" cadence test images as more `target=0`. Did not affect CV but decreased LB like 0.050!\n* Use test pseudo labels. Decrease LB like 0.030!\n* Train model to classify \"train\" image versus \"test\" image. So discard target column and use `test image as target=1` and `train image as target=0`. The hope was that this model would find any new patterns in test data because that would be the difference between train and test. To my surprise the model achieved CV 0.99 but LB was terrible 🙁\n* Train model to classify \"on\" cadence `img[::2]` versus \"off\" cadence `img[1::2]`. So discard target column and use `cadence on as target=1` and `cadence off as target=0`. The advantage here is that we can train directly on test data without train data!\n* Sliding window over \"on\" cadence and compute cosine similarity with sliding window over \"off\" cadence. Take minimum cosine similarity value over all crops of \"off\" cadence. Then take maximum (of these minimums) over all crops of \"on\" cadence. One advantage here is that we can apply this to test data without using train data!\n\nMy favorite model was the last which achieved LB 0.550 and did not use any train data. It simply compared crops of test image \"on\" cadence with crops of test image \"off\" cadence and searched for dissimilarity.\n\n# Public Notebook - CV 884 - LB 0.779 - GAP = 0.105 (Days 5-6)\nBy Saturday, i was frustrated that I couldn't reach Bronze zone. How was everyone doing it?  \n\nI began reading discussions and notebooks. I found @tuckerarrants amazing \"SETI - Learned Image Resizing\" notebook [here][1]. What impressed me most was that his **GAP was only 0.102**! His CV was 0.846 and his LB was 0.744. This was a smaller gap than my model. \n\nWithout understanding why, I quickly ran his notebook locally with larger models and larger image sizes to secure bronze position. The following table are results for fold 0 only. \n\nNote that Tucker's notebook has a resize module before the EfficientNet, so we put `image_in` into the resize module and `image_out` comes out. We then feed `image_out` into the EfficientNet backbone:\n| Backbone | Image In | Image Out |CV - fold 0 | LB - fold 0 |\n| --- | --- | --- | --- | --- |\n| EB3 | 512 | 256 | 0.859 | 0.764 |\n| EB3 | 640 | 320 | 0.870 | 0.768 |\n| EB3 | 768 | 384 | 0.870 | |\n| EB3 | orig | 448 | 0.872 | |\n| EB4 | 640 | 320 | 0.873 | 0.770 |\n| EB4 | 768 | 384 | 0.872 | |\n| EB4 | orig | 448 | 0.869 | |\n| EB6 | 640 | 320 | 0.872 | 0.770 |\n| Ensemble| | | 0.883 | 0.777 |\n\nThen 5-Fold was LB 0.778 and post process got LB 0.779\n\n# Nvidia 4xV100 GPUs\nFor the past few days, Nvidia V100 GPUs were running constantly to train the models in the table above. I am learning PyTorch and discovered a great trick [here][2]. Using Tucker's notebook, you can add 1 line of code to train PyTorch with multiple GPUs (for large backbones and batch sizes). Just replace the one line `model.to(device)` with the three lines:\n\n    import torch.nn as nn\n    model = nn.DataParallel(model)\n    model.to(device)\n\nThat's it! With 1 line of code, we can train larger models and larger batch sizes by using multiple Nvidia GPU. \n\n# The Secret of Timm Backbones\nMany people (including @dragonzhang) were confused why they could not train efficientnet B4 with Tucker's notebook. That is because timm does not contain an imagenet pretrained efficientnet B4. (So if you use `model_name='efficientnet_b4'` you train with a not-pretrained effnet) \n\nTimm only contains pretrains of EB4, 5, 6, 7 of the TF converted weights, `tf_efficientnet_b4`, `tf_efficientnet_b4_ns`, `tf_efficientnet_b4_ap`. The first is original efficient net, the next two are pretrained with noisy student and adv prop.  To view the full list of available timm pretrained backbones, type\n\n    import timm\n    from pprint import pprint\n    model_names = timm.list_models(pretrained=True)\n    pprint(model_names)\n\n# Model Difference - Mysterious \"S\" Shape! (Day 7)\nAfter securing bronze medal position, I began to investigate the CV LB gap again. Four hours before competition deadline, I plotted the difference between my model and public notebook model. I sorted by test images where the difference between my model prediction and public model prediction were greatest. To my surprise, I found a mysterious \"S\" shaped test pattern!\n\nThe image on the left is original comp test data \"on\" cadence vertical stack. With time on x axis and frequency on y axis. The second image is pixel histogram. The third image is \"on\" cadence with emboss, blur, and adaptive histogram equalization. (Filter code shown below) The fourth image (right image) is the \"off\" cadence.\n \n    import cv2\n    clahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\n    # FILTERS FOR BETTER IMAGE DISPLAY\n    img = np.vstack(cadence[::2,])\n    img = img[1:,1:] - img[:-1,:-1] #emboss\n    img -= np.min(img)\n    img /= np.max(img)\n    img = (img*255).astype('uint8')\n    img = cv2.GaussianBlur(img,(5,5),0)\n    img = clahe.apply(img)\n\nAbove the left image we see the public notebook prediction value. And above the third image we see my model's prediction value.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex1b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex2b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex3b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex4b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex5b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex6b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex7b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex8b.png)\n\n# Post Process - LB +0.001 - GAP = 0.104 (Last 4 hours!)\nWith only 4 hours remaining, I trained a model on the images that my model and public notebook model disagreed most on. I then inferred the entire test data. For all images with `p>0.6`, I increased their submission prediction. This post processed closed the CV LB gap by `0.001` and increased LB by `0.001`. Thus there is something special about these \"S\" patterns!\n\n# UPDATE - Mixup is the Silver Medal Magic!\nAfter the comp ended, i have been continuing with experiments to determine why my original model's CV LB is large. It appears that mixup is very important to teach your model to generalize to test data which is different than the train data. It both closes the CV LB gap and boosts CV LB. \n\nUsing 1 fold of EB4 and image size 768x768 with only Hflip, Vflip, and Mixup (alpha=3, max target). One can achieve Silver medal with public LB 0.786 and private LB 0.781. Notebook [here][5]. (Also important is large image size and large backbone).\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\n\n# UPDATE - \"S-shape\" and \"Remove Noise\" is Gold Medal Magic!\nTo climb from Silver medal to Gold medal, we need to detect \"s-shapes\" and remove noise. The first place winning solution [here][3] explains how. Additionally, repeatedly using test pseudo labels also helps climb to top of Gold Medals! (and using old data before reset can help too).\n\n# UPDATE - Grad Cam\nI posted a discussion [here][6] about grad cam and a notebook [here][5] using grad cam. Using grad cam can help us understand what makes a particular image a `target=1`. For example, our model can make a test prediction and then circle what it thinks is causing a `target=1` prediction. The image below is 100% generated by code and no human labeling!\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png)\n\n[1]: https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing?scriptVersionId=68644223\n[2]: https://pytorch.org/tutorials/beginner/blitz/data_parallel_tutorial.html#create-model-and-dataparallel\n[3]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\n[4]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266513#1483010\n[5]: https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\n[6]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/268314",
      "votes": 54
    },
    {
      "id": 1481535,
      "postDate": "2021-08-19T15:07:49.567Z",
      "content": "<p>Great write up Chris and congratulations on getting bronze. I'm glad my notebook was useful during your 7 day journey!</p>",
      "rawMarkdown": "Great write up Chris and congratulations on getting bronze. I'm glad my notebook was useful during your 7 day journey!",
      "votes": 3
    },
    {
      "id": 1482912,
      "postDate": "2021-08-20T10:50:09.637Z",
      "content": "<p>Looks like we lost another one to PyTorch. 😄<br>\nThanks for sharing Chris. A pleasure as always.</p>",
      "rawMarkdown": "Looks like we lost another one to PyTorch. 😄\nThanks for sharing Chris. A pleasure as always.",
      "votes": 4
    },
    {
      "id": 1493974,
      "postDate": "2021-08-28T10:03:04.370Z",
      "content": "<p>Thank you for your detailed and clear explanation of your own thinking process, thank you for your open source code, and all kinds of thoughtful notes in the article. These are very friendly to novices like me, and I can always learn a lot from you.😄</p>",
      "rawMarkdown": "Thank you for your detailed and clear explanation of your own thinking process, thank you for your open source code, and all kinds of thoughtful notes in the article. These are very friendly to novices like me, and I can always learn a lot from you.😄",
      "votes": 1
    },
    {
      "id": 1488383,
      "postDate": "2021-08-24T08:51:33.210Z",
      "content": "<p>Wow! It's a grand post. Brilliantly done and congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Wow! It's a grand post. Brilliantly done and congratulations @cdeotte ",
      "votes": 1
    },
    {
      "id": 1484548,
      "postDate": "2021-08-21T12:07:13.687Z",
      "content": "<p>Hi Chris, congratulations and thanks for this writeup, this really helps me to learn. Just two quick questions:</p>\n<ul>\n<li>What exactly do the \"Image in\" and \"Image out\" resolutions mean?</li>\n<li>I used Keras EfficientNets which expect inputs in the 0-255 range. Do you also prepare your data to be in that range or is there a way to feed EffNets with ~mean 0, std 1 data?<br>\nThanks!</li>\n</ul>",
      "rawMarkdown": "Hi Chris, congratulations and thanks for this writeup, this really helps me to learn. Just two quick questions:\n- What exactly do the \"Image in\" and \"Image out\" resolutions mean?\n- I used Keras EfficientNets which expect inputs in the 0-255 range. Do you also prepare your data to be in that range or is there a way to feed EffNets with ~mean 0, std 1 data?\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1484566,
          "postDate": "2021-08-21T12:35:32.410Z",
          "content": "<p>Hi Markus. My model is more than just an EfficientNet. It has a special module beforehand which learns to resize the image by itself. Here's a diagram of just the module. (The efficientnet is not in the picture):<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/resize.png\" alt=\"\"><br>\nSo first i load the Numpy array and resize it to <code>image_in</code> with <code>img = np.vstack( img[::2] )</code> and <code>img = cv2.resize(img,(image_in,image_in))</code>. </p>\n<p>Then it gets fed into the module. Then the module learns to resize it to <code>image_out</code>. Finally that result is fed into the EfficientNet. I used Tucker's public notebook <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">here</a></p>\n<p>Note that i used this public notebook because i had a few days left and could not figure out how to close the CV LB gap with my own model. I am now investigating why that notebook has a smaller CV LB gap than my model. I have discovered that this special module is not needed. I believe the secret of that notebook is that it uses mixup augmentation whereas my model did not.</p>",
          "rawMarkdown": "Hi Markus. My model is more than just an EfficientNet. It has a special module beforehand which learns to resize the image by itself. Here's a diagram of just the module. (The efficientnet is not in the picture):\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/resize.png)\nSo first i load the Numpy array and resize it to `image_in` with `img = np.vstack( img[::2] )` and `img = cv2.resize(img,(image_in,image_in))`. \n\nThen it gets fed into the module. Then the module learns to resize it to `image_out`. Finally that result is fed into the EfficientNet. I used Tucker's public notebook [here][1]\n\nNote that i used this public notebook because i had a few days left and could not figure out how to close the CV LB gap with my own model. I am now investigating why that notebook has a smaller CV LB gap than my model. I have discovered that this special module is not needed. I believe the secret of that notebook is that it uses mixup augmentation whereas my model did not.\n\n[1]: https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing",
          "votes": 2
        },
        {
          "id": 1484572,
          "postDate": "2021-08-21T12:42:40.013Z",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> There are two efficientnet's you can use for TensorFlow. You can use the one in <code>https://keras.io/api/applications/</code> which was pretrained with image pixels of 0 thru 255. Or you can use <code>pip install efficientnet</code> <a href=\"https://pypi.org/project/efficientnet/\" target=\"_blank\">here</a> which was pretrained with image pixels normalized with mean 0 and std 1. </p>\n<p>If you use the PyTorch efficientnet's. They are pretrained with pixels mean 0 and std 1.</p>\n<p>It will help your efficientnet converge faster to use the same pixel range that the pretraining used. However if it you get it wrong, the model should still converge and learn to adjust to what you are feeding it.</p>",
          "rawMarkdown": "@friedchips There are two efficientnet's you can use for TensorFlow. You can use the one in `https://keras.io/api/applications/` which was pretrained with image pixels of 0 thru 255. Or you can use `pip install efficientnet` [here][1] which was pretrained with image pixels normalized with mean 0 and std 1. \n\nIf you use the PyTorch efficientnet's. They are pretrained with pixels mean 0 and std 1.\n\nIt will help your efficientnet converge faster to use the same pixel range that the pretraining used. However if it you get it wrong, the model should still converge and learn to adjust to what you are feeding it.\n\n[1]: https://pypi.org/project/efficientnet/",
          "votes": 2
        },
        {
          "id": 1484743,
          "postDate": "2021-08-21T14:48:05.130Z",
          "content": "<p>Thanks! I had overlooked your image resizing stage. I just rescaled the original input images with a Recaling layer before feeding them into my EffNetB2 and also saw that the bigger the better, finally going with 768x768 as input to the EffNet.</p>\n<p>Yes, Mixup was extremely helpful, especially with alpha very large (5.0). It just never occurred to me to not mix the labels…</p>\n<p>Thanks for the other efficientnet implementation. My Keras efficientnet converged much worse (CV 0.82 instead of 0.87) when I fed it with 0-1 data before realizing that it expects 0-255.</p>",
          "rawMarkdown": "Thanks! I had overlooked your image resizing stage. I just rescaled the original input images with a Recaling layer before feeding them into my EffNetB2 and also saw that the bigger the better, finally going with 768x768 as input to the EffNet.\n\nYes, Mixup was extremely helpful, especially with alpha very large (5.0). It just never occurred to me to not mix the labels...\n\nThanks for the other efficientnet implementation. My Keras efficientnet converged much worse (CV 0.82 instead of 0.87) when I fed it with 0-1 data before realizing that it expects 0-255.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1487265,
      "postDate": "2021-08-23T14:06:19.443Z",
      "content": "<p>Update. I have been conducting more experiments. I have updated my post to identify the Silver Medal Magic. And the Gold Medal Magic!</p>",
      "rawMarkdown": "Update. I have been conducting more experiments. I have updated my post to identify the Silver Medal Magic. And the Gold Medal Magic!",
      "votes": 2,
      "replies": [
        {
          "id": 1491516,
          "postDate": "2021-08-26T12:50:15.690Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great insights as usual. </p>\n<p>I realise many contestants plot pixel differences like you did in image comps. May I know if you can share the code you used to plot the above diagrams so I can compare if I’m doing it right  </p>",
          "rawMarkdown": "Thanks @cdeotte great insights as usual. \n\nI realise many contestants plot pixel differences like you did in image comps. May I know if you can share the code you used to plot the above diagrams so I can compare if I’m doing it right  ",
          "votes": 2
        },
        {
          "id": 1491537,
          "postDate": "2021-08-26T12:59:34.140Z",
          "content": "<p>The code is posted above. For \"on\" cadence, i would read NumPy like </p>\n<pre><code>file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\nON_CADENCE = file[::2,]\n</code></pre>\n<p>And for the same \"off\" cadence, use the same <code>name</code> and load</p>\n<pre><code>file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\nOFF_CADENCE = file[1::2,]\n</code></pre>\n<p>Next, i apply the following filter to make the spectrograms easier to see with human eye</p>\n<pre><code>import cv2\nclahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\nimg = np.vstack(CADENCE)\nimg = img[1:,1:] - img[:-1,:-1] #emboss\nimg -= np.min(img)\nimg /= np.max(img)\nimg = (img*255).astype('uint8')\nimg = cv2.GaussianBlur(img,(5,5),0)\nimg = clahe.apply(img)\n</code></pre>\n<p>To plot 4 tiles in a row, you use <code>subplot</code> like</p>\n<pre><code>plt.figure(figsize=(20,5)\nplt.subplot(1,4,1)\nplt.imshow(img1)\nplt.subplot(1,4,2)\nplt.imshow(img2)\nplt.subplot(1,4,3)\nplt.imshow(img3)\nplt.subplot(1,4,4)\nplt.imshow(img4)\nplt.show()\n</code></pre>",
          "rawMarkdown": "The code is posted above. For \"on\" cadence, i would read NumPy like \n\n    file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\n    ON_CADENCE = file[::2,]\n\nAnd for the same \"off\" cadence, use the same `name` and load\n\n    file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\n    OFF_CADENCE = file[1::2,]\n\nNext, i apply the following filter to make the spectrograms easier to see with human eye\n\n    import cv2\n    clahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\n    img = np.vstack(CADENCE)\n    img = img[1:,1:] - img[:-1,:-1] #emboss\n    img -= np.min(img)\n    img /= np.max(img)\n    img = (img*255).astype('uint8')\n    img = cv2.GaussianBlur(img,(5,5),0)\n    img = clahe.apply(img)\n\nTo plot 4 tiles in a row, you use `subplot` like\n\n    plt.figure(figsize=(20,5)\n    plt.subplot(1,4,1)\n    plt.imshow(img1)\n    plt.subplot(1,4,2)\n    plt.imshow(img2)\n    plt.subplot(1,4,3)\n    plt.imshow(img3)\n    plt.subplot(1,4,4)\n    plt.imshow(img4)\n    plt.show()\n",
          "votes": 1
        },
        {
          "id": 1492932,
          "postDate": "2021-08-27T14:33:13.683Z",
          "content": "<p>Thanks! Got it. </p>",
          "rawMarkdown": "Thanks! Got it. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1483185,
      "postDate": "2021-08-20T13:47:22.503Z",
      "content": "<p>Nice write up as always Chris, thanks for sharing your findings. I'm excited to see your future pytorch experiments too, you did pretty good with pytorch written scripts in our last competition, I'm sure you'll get great at that in no time too!</p>",
      "rawMarkdown": "Nice write up as always Chris, thanks for sharing your findings. I'm excited to see your future pytorch experiments too, you did pretty good with pytorch written scripts in our last competition, I'm sure you'll get great at that in no time too!",
      "votes": 2,
      "replies": [
        {
          "id": 1483218,
          "postDate": "2021-08-20T14:07:47.837Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Congrats earning Silver medal. I was watching your team. At one point, i saw you jump from LB 775 to LB 783. What discovery did you make at that point?</p>\n<p>I'd love to hear about your solution. What was your strongest single model? And what preprocessing and/or post processing techniques did you use?</p>",
          "rawMarkdown": "Thanks @datafan07 Congrats earning Silver medal. I was watching your team. At one point, i saw you jump from LB 775 to LB 783. What discovery did you make at that point?\n\nI'd love to hear about your solution. What was your strongest single model? And what preprocessing and/or post processing techniques did you use?",
          "votes": 1
        },
        {
          "id": 1483268,
          "postDate": "2021-08-20T14:40:06.850Z",
          "content": "<p>Ah I remember that jump, funny enough it was from single fold model predictions. We were building different models, some team members building tf tpu models and I was working with pytorch and that single model comes from pytorch trial. I remember adding heavy augmentations with mixup + cutmix and changing backbone to  'EfficientNetV2-M' helped me a lot. It was our best single model with almost 0.90cv.</p>\n<p>After that I didn't have enough resources and time to make trials on rest of the folds so we went for ensembling for last days. I tried test same setup on TPU's with effnetv2 etc. to get faster results but didn't get same score as pytorch one. cv was decent but lb failed with that one…</p>",
          "rawMarkdown": "Ah I remember that jump, funny enough it was from single fold model predictions. We were building different models, some team members building tf tpu models and I was working with pytorch and that single model comes from pytorch trial. I remember adding heavy augmentations with mixup + cutmix and changing backbone to  'EfficientNetV2-M' helped me a lot. It was our best single model with almost 0.90cv.\n\nAfter that I didn't have enough resources and time to make trials on rest of the folds so we went for ensembling for last days. I tried test same setup on TPU's with effnetv2 etc. to get faster results but didn't get same score as pytorch one. cv was decent but lb failed with that one...",
          "votes": 2
        },
        {
          "id": 1483304,
          "postDate": "2021-08-20T15:07:46.733Z",
          "content": "<p>Wow, CV 0.90, that's great. I would to see code that scores over LB 0.780. To me, achieving over 0.780 is still magic. No team posted code that achieves over LB 0.780 yet. (Second place posted code but it throws an error if you run it).</p>",
          "rawMarkdown": "Wow, CV 0.90, that's great. I would to see code that scores over LB 0.780. To me, achieving over 0.780 is still magic. No team posted code that achieves over LB 0.780 yet. (Second place posted code but it throws an error if you run it).",
          "votes": 3
        },
        {
          "id": 1483350,
          "postDate": "2021-08-20T15:35:37.697Z",
          "content": "<p>Like I did in commonlit I actually shared baseline version of my training <a href=\"https://www.kaggle.com/datafan07/pytorch-lightning-single-fold-training-lb-0-97\" target=\"_blank\">long time ago here</a>. I just added some more updates based on that pipeline, like cutmix, different backbones etc. and more epochs. I didn't update the public code after leak reset but main part still stands same…(There might be some errors in this one too since some packages got important updates, like checkpoint code)</p>",
          "rawMarkdown": "Like I did in commonlit I actually shared baseline version of my training [long time ago here](https://www.kaggle.com/datafan07/pytorch-lightning-single-fold-training-lb-0-97). I just added some more updates based on that pipeline, like cutmix, different backbones etc. and more epochs. I didn't update the public code after leak reset but main part still stands same...(There might be some errors in this one too since some packages got important updates, like checkpoint code)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1482418,
      "postDate": "2021-08-20T04:47:34.590Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> … the evolution of thought and how you proceed with different alternatives is helpful on how to think through the steps and try out multiple approaches! </p>",
      "rawMarkdown": "Thanks for sharing @cdeotte ... the evolution of thought and how you proceed with different alternatives is helpful on how to think through the steps and try out multiple approaches! ",
      "votes": 2
    },
    {
      "id": 1481643,
      "postDate": "2021-08-19T16:03:21.090Z",
      "content": "<p>Thanks as always for your kind explanation. :) <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks as always for your kind explanation. :) @cdeotte ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1481535,
      "author_name": "Tucker Arrants",
      "author_url": "",
      "post_date": "2021-08-19T15:07:49.567000",
      "content": "<p>Great write up Chris and congratulations on getting bronze. I'm glad my notebook was useful during your 7 day journey!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1482912,
      "author_name": "Jared Savage",
      "author_url": "",
      "post_date": "2021-08-20T10:50:09.637000",
      "content": "<p>Looks like we lost another one to PyTorch. 😄<br>\nThanks for sharing Chris. A pleasure as always.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1493974,
      "author_name": "Gainover",
      "author_url": "",
      "post_date": "2021-08-28T10:03:04.370000",
      "content": "<p>Thank you for your detailed and clear explanation of your own thinking process, thank you for your open source code, and all kinds of thoughtful notes in the article. These are very friendly to novices like me, and I can always learn a lot from you.😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1488383,
      "author_name": "Kalilur Rahman",
      "author_url": "",
      "post_date": "2021-08-24T08:51:33.210000",
      "content": "<p>Wow! It's a grand post. Brilliantly done and congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1484548,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-21T12:07:13.687000",
      "content": "<p>Hi Chris, congratulations and thanks for this writeup, this really helps me to learn. Just two quick questions:</p>\n<ul>\n<li>What exactly do the \"Image in\" and \"Image out\" resolutions mean?</li>\n<li>I used Keras EfficientNets which expect inputs in the 0-255 range. Do you also prepare your data to be in that range or is there a way to feed EffNets with ~mean 0, std 1 data?<br>\nThanks!</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1484566,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T12:35:32.410000",
          "content": "<p>Hi Markus. My model is more than just an EfficientNet. It has a special module beforehand which learns to resize the image by itself. Here's a diagram of just the module. (The efficientnet is not in the picture):<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/resize.png\" alt=\"\"><br>\nSo first i load the Numpy array and resize it to <code>image_in</code> with <code>img = np.vstack( img[::2] )</code> and <code>img = cv2.resize(img,(image_in,image_in))</code>. </p>\n<p>Then it gets fed into the module. Then the module learns to resize it to <code>image_out</code>. Finally that result is fed into the EfficientNet. I used Tucker's public notebook <a href=\"https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing\" target=\"_blank\">here</a></p>\n<p>Note that i used this public notebook because i had a few days left and could not figure out how to close the CV LB gap with my own model. I am now investigating why that notebook has a smaller CV LB gap than my model. I have discovered that this special module is not needed. I believe the secret of that notebook is that it uses mixup augmentation whereas my model did not.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1484572,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T12:42:40.013000",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> There are two efficientnet's you can use for TensorFlow. You can use the one in <code>https://keras.io/api/applications/</code> which was pretrained with image pixels of 0 thru 255. Or you can use <code>pip install efficientnet</code> <a href=\"https://pypi.org/project/efficientnet/\" target=\"_blank\">here</a> which was pretrained with image pixels normalized with mean 0 and std 1. </p>\n<p>If you use the PyTorch efficientnet's. They are pretrained with pixels mean 0 and std 1.</p>\n<p>It will help your efficientnet converge faster to use the same pixel range that the pretraining used. However if it you get it wrong, the model should still converge and learn to adjust to what you are feeding it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1484743,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-21T14:48:05.130000",
          "content": "<p>Thanks! I had overlooked your image resizing stage. I just rescaled the original input images with a Recaling layer before feeding them into my EffNetB2 and also saw that the bigger the better, finally going with 768x768 as input to the EffNet.</p>\n<p>Yes, Mixup was extremely helpful, especially with alpha very large (5.0). It just never occurred to me to not mix the labels…</p>\n<p>Thanks for the other efficientnet implementation. My Keras efficientnet converged much worse (CV 0.82 instead of 0.87) when I fed it with 0-1 data before realizing that it expects 0-255.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1487265,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-23T14:06:19.443000",
      "content": "<p>Update. I have been conducting more experiments. I have updated my post to identify the Silver Medal Magic. And the Gold Medal Magic!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1491516,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-08-26T12:50:15.690000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> great insights as usual. </p>\n<p>I realise many contestants plot pixel differences like you did in image comps. May I know if you can share the code you used to plot the above diagrams so I can compare if I’m doing it right  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1491537,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-26T12:59:34.140000",
          "content": "<p>The code is posted above. For \"on\" cadence, i would read NumPy like </p>\n<pre><code>file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\nON_CADENCE = file[::2,]\n</code></pre>\n<p>And for the same \"off\" cadence, use the same <code>name</code> and load</p>\n<pre><code>file = numpy .load(path+name[0]+'/'+name+'.npy').astype('float32')\nOFF_CADENCE = file[1::2,]\n</code></pre>\n<p>Next, i apply the following filter to make the spectrograms easier to see with human eye</p>\n<pre><code>import cv2\nclahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\nimg = np.vstack(CADENCE)\nimg = img[1:,1:] - img[:-1,:-1] #emboss\nimg -= np.min(img)\nimg /= np.max(img)\nimg = (img*255).astype('uint8')\nimg = cv2.GaussianBlur(img,(5,5),0)\nimg = clahe.apply(img)\n</code></pre>\n<p>To plot 4 tiles in a row, you use <code>subplot</code> like</p>\n<pre><code>plt.figure(figsize=(20,5)\nplt.subplot(1,4,1)\nplt.imshow(img1)\nplt.subplot(1,4,2)\nplt.imshow(img2)\nplt.subplot(1,4,3)\nplt.imshow(img3)\nplt.subplot(1,4,4)\nplt.imshow(img4)\nplt.show()\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1492932,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-08-27T14:33:13.683000",
          "content": "<p>Thanks! Got it. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1483185,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2021-08-20T13:47:22.503000",
      "content": "<p>Nice write up as always Chris, thanks for sharing your findings. I'm excited to see your future pytorch experiments too, you did pretty good with pytorch written scripts in our last competition, I'm sure you'll get great at that in no time too!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1483218,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T14:07:47.837000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> Congrats earning Silver medal. I was watching your team. At one point, i saw you jump from LB 775 to LB 783. What discovery did you make at that point?</p>\n<p>I'd love to hear about your solution. What was your strongest single model? And what preprocessing and/or post processing techniques did you use?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483268,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-08-20T14:40:06.850000",
          "content": "<p>Ah I remember that jump, funny enough it was from single fold model predictions. We were building different models, some team members building tf tpu models and I was working with pytorch and that single model comes from pytorch trial. I remember adding heavy augmentations with mixup + cutmix and changing backbone to  'EfficientNetV2-M' helped me a lot. It was our best single model with almost 0.90cv.</p>\n<p>After that I didn't have enough resources and time to make trials on rest of the folds so we went for ensembling for last days. I tried test same setup on TPU's with effnetv2 etc. to get faster results but didn't get same score as pytorch one. cv was decent but lb failed with that one…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483304,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T15:07:46.733000",
          "content": "<p>Wow, CV 0.90, that's great. I would to see code that scores over LB 0.780. To me, achieving over 0.780 is still magic. No team posted code that achieves over LB 0.780 yet. (Second place posted code but it throws an error if you run it).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1483350,
          "author_name": "Ertuğrul Demir",
          "author_url": "",
          "post_date": "2021-08-20T15:35:37.697000",
          "content": "<p>Like I did in commonlit I actually shared baseline version of my training <a href=\"https://www.kaggle.com/datafan07/pytorch-lightning-single-fold-training-lb-0-97\" target=\"_blank\">long time ago here</a>. I just added some more updates based on that pipeline, like cutmix, different backbones etc. and more epochs. I didn't update the public code after leak reset but main part still stands same…(There might be some errors in this one too since some packages got important updates, like checkpoint code)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1482418,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2021-08-20T04:47:34.590000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> … the evolution of thought and how you proceed with different alternatives is helpful on how to think through the steps and try out multiple approaches! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1481643,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-08-19T16:03:21.090000",
      "content": "<p>Thanks as always for your kind explanation. :) <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1481348": "# Exciting Competition!\nWhat an interesting competition and difficult puzzle! Thank you Kaggle and SETI for hosting. I started this competition last Wednesday and worked hard every day. In this time, I closed CV LB gap from 0.125 to 0.104. I would have liked to have more time and continue the puzzle of closing the CV LB gap. Below is my journey over the past 7 days.\n\n# My Model - CV 0.890 - LB 0.765 - GAP = 0.125 (Days 1-3)\nLast Wednesday, Thursday, and Friday, I built dozens of my own models without reading public discussions nor notebooks. I was excited when my best model achieved CV 0.890 but surprised and disappointed when I submitted and received LB 0.765 🙁This was a 0.125 gap between CV and LB. Something is going on!\n\nMy best model used only \"on\" cadence in a spatial stack resized to 768x768 and fed into a EfficientNetB4 backbone. The only augmentation was horizontal flip and vertical flip. It used cosine learning schedule and 15 epochs.\n\n# Alternative Ideas (Days 3-4)\nAt this point, I thought the gap was caused by new pattens in test data. So, on Friday and Saturday, i starting building more creative models that used more than just classify train patterns.\n* Use \"off\" cadence train images as more `target=0`. Increased CV 0.003 but decreased LB 0.010\n* Use \"off\" cadence test images as more `target=0`. Did not affect CV but decreased LB like 0.050!\n* Use test pseudo labels. Decrease LB like 0.030!\n* Train model to classify \"train\" image versus \"test\" image. So discard target column and use `test image as target=1` and `train image as target=0`. The hope was that this model would find any new patterns in test data because that would be the difference between train and test. To my surprise the model achieved CV 0.99 but LB was terrible 🙁\n* Train model to classify \"on\" cadence `img[::2]` versus \"off\" cadence `img[1::2]`. So discard target column and use `cadence on as target=1` and `cadence off as target=0`. The advantage here is that we can train directly on test data without train data!\n* Sliding window over \"on\" cadence and compute cosine similarity with sliding window over \"off\" cadence. Take minimum cosine similarity value over all crops of \"off\" cadence. Then take maximum (of these minimums) over all crops of \"on\" cadence. One advantage here is that we can apply this to test data without using train data!\n\nMy favorite model was the last which achieved LB 0.550 and did not use any train data. It simply compared crops of test image \"on\" cadence with crops of test image \"off\" cadence and searched for dissimilarity.\n\n# Public Notebook - CV 884 - LB 0.779 - GAP = 0.105 (Days 5-6)\nBy Saturday, i was frustrated that I couldn't reach Bronze zone. How was everyone doing it?  \n\nI began reading discussions and notebooks. I found @tuckerarrants amazing \"SETI - Learned Image Resizing\" notebook [here][1]. What impressed me most was that his **GAP was only 0.102**! His CV was 0.846 and his LB was 0.744. This was a smaller gap than my model. \n\nWithout understanding why, I quickly ran his notebook locally with larger models and larger image sizes to secure bronze position. The following table are results for fold 0 only. \n\nNote that Tucker's notebook has a resize module before the EfficientNet, so we put `image_in` into the resize module and `image_out` comes out. We then feed `image_out` into the EfficientNet backbone:\n| Backbone | Image In | Image Out |CV - fold 0 | LB - fold 0 |\n| --- | --- | --- | --- | --- |\n| EB3 | 512 | 256 | 0.859 | 0.764 |\n| EB3 | 640 | 320 | 0.870 | 0.768 |\n| EB3 | 768 | 384 | 0.870 | |\n| EB3 | orig | 448 | 0.872 | |\n| EB4 | 640 | 320 | 0.873 | 0.770 |\n| EB4 | 768 | 384 | 0.872 | |\n| EB4 | orig | 448 | 0.869 | |\n| EB6 | 640 | 320 | 0.872 | 0.770 |\n| Ensemble| | | 0.883 | 0.777 |\n\nThen 5-Fold was LB 0.778 and post process got LB 0.779\n\n# Nvidia 4xV100 GPUs\nFor the past few days, Nvidia V100 GPUs were running constantly to train the models in the table above. I am learning PyTorch and discovered a great trick [here][2]. Using Tucker's notebook, you can add 1 line of code to train PyTorch with multiple GPUs (for large backbones and batch sizes). Just replace the one line `model.to(device)` with the three lines:\n\n    import torch.nn as nn\n    model = nn.DataParallel(model)\n    model.to(device)\n\nThat's it! With 1 line of code, we can train larger models and larger batch sizes by using multiple Nvidia GPU. \n\n# The Secret of Timm Backbones\nMany people (including @dragonzhang) were confused why they could not train efficientnet B4 with Tucker's notebook. That is because timm does not contain an imagenet pretrained efficientnet B4. (So if you use `model_name='efficientnet_b4'` you train with a not-pretrained effnet) \n\nTimm only contains pretrains of EB4, 5, 6, 7 of the TF converted weights, `tf_efficientnet_b4`, `tf_efficientnet_b4_ns`, `tf_efficientnet_b4_ap`. The first is original efficient net, the next two are pretrained with noisy student and adv prop.  To view the full list of available timm pretrained backbones, type\n\n    import timm\n    from pprint import pprint\n    model_names = timm.list_models(pretrained=True)\n    pprint(model_names)\n\n# Model Difference - Mysterious \"S\" Shape! (Day 7)\nAfter securing bronze medal position, I began to investigate the CV LB gap again. Four hours before competition deadline, I plotted the difference between my model and public notebook model. I sorted by test images where the difference between my model prediction and public model prediction were greatest. To my surprise, I found a mysterious \"S\" shaped test pattern!\n\nThe image on the left is original comp test data \"on\" cadence vertical stack. With time on x axis and frequency on y axis. The second image is pixel histogram. The third image is \"on\" cadence with emboss, blur, and adaptive histogram equalization. (Filter code shown below) The fourth image (right image) is the \"off\" cadence.\n \n    import cv2\n    clahe = cv2.createCLAHE(clipLimit=16.0, tileGridSize=(8,8))\n    # FILTERS FOR BETTER IMAGE DISPLAY\n    img = np.vstack(cadence[::2,])\n    img = img[1:,1:] - img[:-1,:-1] #emboss\n    img -= np.min(img)\n    img /= np.max(img)\n    img = (img*255).astype('uint8')\n    img = cv2.GaussianBlur(img,(5,5),0)\n    img = clahe.apply(img)\n\nAbove the left image we see the public notebook prediction value. And above the third image we see my model's prediction value.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex1b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex2b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex3b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex4b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex5b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex6b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex7b.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/ex8b.png)\n\n# Post Process - LB +0.001 - GAP = 0.104 (Last 4 hours!)\nWith only 4 hours remaining, I trained a model on the images that my model and public notebook model disagreed most on. I then inferred the entire test data. For all images with `p>0.6`, I increased their submission prediction. This post processed closed the CV LB gap by `0.001` and increased LB by `0.001`. Thus there is something special about these \"S\" patterns!\n\n# UPDATE - Mixup is the Silver Medal Magic!\nAfter the comp ended, i have been continuing with experiments to determine why my original model's CV LB is large. It appears that mixup is very important to teach your model to generalize to test data which is different than the train data. It both closes the CV LB gap and boosts CV LB. \n\nUsing 1 fold of EB4 and image size 768x768 with only Hflip, Vflip, and Mixup (alpha=3, max target). One can achieve Silver medal with public LB 0.786 and private LB 0.781. Notebook [here][5]. (Also important is large image size and large backbone).\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\n\n# UPDATE - \"S-shape\" and \"Remove Noise\" is Gold Medal Magic!\nTo climb from Silver medal to Gold medal, we need to detect \"s-shapes\" and remove noise. The first place winning solution [here][3] explains how. Additionally, repeatedly using test pseudo labels also helps climb to top of Gold Medals! (and using old data before reset can help too).\n\n# UPDATE - Grad Cam\nI posted a discussion [here][6] about grad cam and a notebook [here][5] using grad cam. Using grad cam can help us understand what makes a particular image a `target=1`. For example, our model can make a test prediction and then circle what it thinks is causing a `target=1` prediction. The image below is 100% generated by code and no human labeling!\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png)\n\n[1]: https://www.kaggle.com/tuckerarrants/seti-learned-image-resizing?scriptVersionId=68644223\n[2]: https://pytorch.org/tutorials/beginner/blitz/data_parallel_tutorial.html#create-model-and-dataparallel\n[3]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\n[4]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266513#1483010\n[5]: https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\n[6]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/268314",
    "1481535": "Great write up Chris and congratulations on getting bronze. I'm glad my notebook was useful during your 7 day journey!",
    "1482912": "Looks like we lost another one to PyTorch. 😄\nThanks for sharing Chris. A pleasure as always.",
    "1493974": "Thank you for your detailed and clear explanation of your own thinking process, thank you for your open source code, and all kinds of thoughtful notes in the article. These are very friendly to novices like me, and I can always learn a lot from you.😄",
    "1488383": "Wow! It's a grand post. Brilliantly done and congratulations @cdeotte ",
    "1484548": "Hi Chris, congratulations and thanks for this writeup, this really helps me to learn. Just two quick questions:\n- What exactly do the \"Image in\" and \"Image out\" resolutions mean?\n- I used Keras EfficientNets which expect inputs in the 0-255 range. Do you also prepare your data to be in that range or is there a way to feed EffNets with ~mean 0, std 1 data?\nThanks!",
    "1487265": "Update. I have been conducting more experiments. I have updated my post to identify the Silver Medal Magic. And the Gold Medal Magic!",
    "1483185": "Nice write up as always Chris, thanks for sharing your findings. I'm excited to see your future pytorch experiments too, you did pretty good with pytorch written scripts in our last competition, I'm sure you'll get great at that in no time too!",
    "1482418": "Thanks for sharing @cdeotte ... the evolution of thought and how you proceed with different alternatives is helpful on how to think through the steps and try out multiple approaches! ",
    "1481643": "Thanks as always for your kind explanation. :) @cdeotte "
  }
}