{
  "id": 266619,
  "title": "8th place solution",
  "url": "/competitions/seti-breakthrough-listen/discussion/266619",
  "author_name": "Ilya Makarov",
  "post_date": "2021-08-19T18:26:05.016000",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>TL;DR</h1>\n<p>My best scoring prediction is an average of eca_nfnet_l1 (384x384 and 420x420), EfficientNet-B2 (384x384), -B6 (512x512 and 640x640) and -B7 (640x640) predictions with 2x tta (original + HFlip). The models were trained with FastAI locally on my 3070 and later 3080Ti (eca_nfnet_l1 and B2) and with TF on Kaggle TPU (B6 and B7) with LAMB as optimizer, SGD with restarts as lr schedule and applying slight weighing in cross-entropy loss function (2). New train data, old train and test data normalized per tile as well as soft labeled new test data were used for training with MixUp and HFlip as augmentations. All 6 tiles of the cadence were stacked and resized in two steps: first, upsized to 1638x1638 with INTER_NEAREST interpolation and then downsized with INTER_AREA.</p>\n<h1>In detail</h1>\n<h3>Selecting models</h3>\n<p>As was noted by other teams too, not all models with good local CV had correspondingly good public LB. The two families of models that turned out to perform well were eca_nfnet_l's and EfficientNets while resnets (regular and d ones), ecaresnets, resnests and some other had a lower LB than I would anticipate from their CV scores. Unfortunately, I did not manage to extract much useful insights from this discrepancy between the LB and CV… Also, unsurprisingly larger models performed slightly better than the smaller ones, at the same time, excluding the smaller models from the final ensemble did worsen the score a bit. The best single model was a 4-fold EfficientNet-B7 (640x640) with a public LB of 0.80092 and private LB 0.79648.</p>\n<h3>Image size and resizing</h3>\n<p>For the training I used all the 6 tiles of cadence because before the restart of the competition such an approach performed best. While after the restart I kept using the same approach as I saw no gain to a slight negative decline in the score when using just the 3 positive tiles. Also, very early in the competition I discovered that the score was pretty sensitive to the resizing method. When downsizing the images to 256x256, INTER_AREA interpolation was the method of choice as any other interpolation available in OpenCV was consistently giving a 0.02-0.03 lower AUC for any model tested. In addition, changing two dimensions at the same time (like 1638x256 -&gt; 320x320) I could never obtain a better score than with 256x256 inputs regardless of the interpolation. The only way for me to get a better score for the images with a frequency dimension greater than 256 px was to first scale the original inputs to 1638x1638 with INTER_NEAREST interpolation and then resize these large images to any smaller size. When upscaling the inputs in this two-step way, the larger input sizes always resulted in larger CV and LB scores. At the end, the input sizes I used in my training was a compromise between the available resources and the score.</p>\n<h3>Optimizers and lr schedules</h3>\n<p>LAMB optimizer turned out to give consistently better score (both with TF and FastAI) in comparison with any other optimizer I tried even though LAMB was the slowest of all them. However, the consistent gain in score was worth the 10-20% slowdown per epoch. Regarding the lr schedules, SGD with warm restarts and training for 7 epochs (1 + 2 + 4) gave the best results in turns of training time and the CV score when compared to the 1-cycle policy or cosine decay with warm-up as well as training for more epochs. Finally, using the cross-entropy with a light weighing of positive examples outperformed simple cross-entropy and the focal loss.</p>\n<h3>Train data</h3>\n<p>In my experiments adding the normalized old train and test data to the new train data resulted in a consistent increase by 0.02-0.03 in the local CV. As it was obvious that the new test data differed a lot from the train sets, using the soft labeled test data helped further improve the public LB score (at the same time it slightly decreased the local CV but at this point I trusted the LB score more that my CV).</p>\n<h3>Tools</h3>\n<p>I did almost all my experiments in FastAI on my local computer with one RTX 3070. It was only in the last couple of weeks of the competition when I found a good way to upscale the inputs and that's when I ran out of my local resources (even though I upgraded my GPU to 3080Ti) and switched to TensorFlow and TPUs on Kaggle. In general, I find the amount of the train data quite suitable for fast experimentations even on a such relatively small machine. Also it was interesting to see that FastAI turned out to be faster than Lightning in several tests (at least in my hands :D), so I decided not to search for any alternatives and kept using FastAI almost exclusively. Finally, to follow all the experiments I used wandb and I should say it helped me a lot not to lose track of all the hundreds of experiments I ran.</p>",
  "messages": [
    {
      "id": 1481926,
      "postDate": "2021-08-19T18:26:05.017Z",
      "content": "<h1>TL;DR</h1>\n<p>My best scoring prediction is an average of eca_nfnet_l1 (384x384 and 420x420), EfficientNet-B2 (384x384), -B6 (512x512 and 640x640) and -B7 (640x640) predictions with 2x tta (original + HFlip). The models were trained with FastAI locally on my 3070 and later 3080Ti (eca_nfnet_l1 and B2) and with TF on Kaggle TPU (B6 and B7) with LAMB as optimizer, SGD with restarts as lr schedule and applying slight weighing in cross-entropy loss function (2). New train data, old train and test data normalized per tile as well as soft labeled new test data were used for training with MixUp and HFlip as augmentations. All 6 tiles of the cadence were stacked and resized in two steps: first, upsized to 1638x1638 with INTER_NEAREST interpolation and then downsized with INTER_AREA.</p>\n<h1>In detail</h1>\n<h3>Selecting models</h3>\n<p>As was noted by other teams too, not all models with good local CV had correspondingly good public LB. The two families of models that turned out to perform well were eca_nfnet_l's and EfficientNets while resnets (regular and d ones), ecaresnets, resnests and some other had a lower LB than I would anticipate from their CV scores. Unfortunately, I did not manage to extract much useful insights from this discrepancy between the LB and CV… Also, unsurprisingly larger models performed slightly better than the smaller ones, at the same time, excluding the smaller models from the final ensemble did worsen the score a bit. The best single model was a 4-fold EfficientNet-B7 (640x640) with a public LB of 0.80092 and private LB 0.79648.</p>\n<h3>Image size and resizing</h3>\n<p>For the training I used all the 6 tiles of cadence because before the restart of the competition such an approach performed best. While after the restart I kept using the same approach as I saw no gain to a slight negative decline in the score when using just the 3 positive tiles. Also, very early in the competition I discovered that the score was pretty sensitive to the resizing method. When downsizing the images to 256x256, INTER_AREA interpolation was the method of choice as any other interpolation available in OpenCV was consistently giving a 0.02-0.03 lower AUC for any model tested. In addition, changing two dimensions at the same time (like 1638x256 -&gt; 320x320) I could never obtain a better score than with 256x256 inputs regardless of the interpolation. The only way for me to get a better score for the images with a frequency dimension greater than 256 px was to first scale the original inputs to 1638x1638 with INTER_NEAREST interpolation and then resize these large images to any smaller size. When upscaling the inputs in this two-step way, the larger input sizes always resulted in larger CV and LB scores. At the end, the input sizes I used in my training was a compromise between the available resources and the score.</p>\n<h3>Optimizers and lr schedules</h3>\n<p>LAMB optimizer turned out to give consistently better score (both with TF and FastAI) in comparison with any other optimizer I tried even though LAMB was the slowest of all them. However, the consistent gain in score was worth the 10-20% slowdown per epoch. Regarding the lr schedules, SGD with warm restarts and training for 7 epochs (1 + 2 + 4) gave the best results in turns of training time and the CV score when compared to the 1-cycle policy or cosine decay with warm-up as well as training for more epochs. Finally, using the cross-entropy with a light weighing of positive examples outperformed simple cross-entropy and the focal loss.</p>\n<h3>Train data</h3>\n<p>In my experiments adding the normalized old train and test data to the new train data resulted in a consistent increase by 0.02-0.03 in the local CV. As it was obvious that the new test data differed a lot from the train sets, using the soft labeled test data helped further improve the public LB score (at the same time it slightly decreased the local CV but at this point I trusted the LB score more that my CV).</p>\n<h3>Tools</h3>\n<p>I did almost all my experiments in FastAI on my local computer with one RTX 3070. It was only in the last couple of weeks of the competition when I found a good way to upscale the inputs and that's when I ran out of my local resources (even though I upgraded my GPU to 3080Ti) and switched to TensorFlow and TPUs on Kaggle. In general, I find the amount of the train data quite suitable for fast experimentations even on a such relatively small machine. Also it was interesting to see that FastAI turned out to be faster than Lightning in several tests (at least in my hands :D), so I decided not to search for any alternatives and kept using FastAI almost exclusively. Finally, to follow all the experiments I used wandb and I should say it helped me a lot not to lose track of all the hundreds of experiments I ran.</p>",
      "rawMarkdown": "# TL;DR\nMy best scoring prediction is an average of eca_nfnet_l1 (384x384 and 420x420), EfficientNet-B2 (384x384), -B6 (512x512 and 640x640) and -B7 (640x640) predictions with 2x tta (original + HFlip). The models were trained with FastAI locally on my 3070 and later 3080Ti (eca_nfnet_l1 and B2) and with TF on Kaggle TPU (B6 and B7) with LAMB as optimizer, SGD with restarts as lr schedule and applying slight weighing in cross-entropy loss function (2). New train data, old train and test data normalized per tile as well as soft labeled new test data were used for training with MixUp and HFlip as augmentations. All 6 tiles of the cadence were stacked and resized in two steps: first, upsized to 1638x1638 with INTER_NEAREST interpolation and then downsized with INTER_AREA.\n\n# In detail\n### Selecting models\nAs was noted by other teams too, not all models with good local CV had correspondingly good public LB. The two families of models that turned out to perform well were eca_nfnet_l's and EfficientNets while resnets (regular and d ones), ecaresnets, resnests and some other had a lower LB than I would anticipate from their CV scores. Unfortunately, I did not manage to extract much useful insights from this discrepancy between the LB and CV... Also, unsurprisingly larger models performed slightly better than the smaller ones, at the same time, excluding the smaller models from the final ensemble did worsen the score a bit. The best single model was a 4-fold EfficientNet-B7 (640x640) with a public LB of 0.80092 and private LB 0.79648.\n\n### Image size and resizing\nFor the training I used all the 6 tiles of cadence because before the restart of the competition such an approach performed best. While after the restart I kept using the same approach as I saw no gain to a slight negative decline in the score when using just the 3 positive tiles. Also, very early in the competition I discovered that the score was pretty sensitive to the resizing method. When downsizing the images to 256x256, INTER_AREA interpolation was the method of choice as any other interpolation available in OpenCV was consistently giving a 0.02-0.03 lower AUC for any model tested. In addition, changing two dimensions at the same time (like 1638x256 -> 320x320) I could never obtain a better score than with 256x256 inputs regardless of the interpolation. The only way for me to get a better score for the images with a frequency dimension greater than 256 px was to first scale the original inputs to 1638x1638 with INTER_NEAREST interpolation and then resize these large images to any smaller size. When upscaling the inputs in this two-step way, the larger input sizes always resulted in larger CV and LB scores. At the end, the input sizes I used in my training was a compromise between the available resources and the score.\n\n### Optimizers and lr schedules\nLAMB optimizer turned out to give consistently better score (both with TF and FastAI) in comparison with any other optimizer I tried even though LAMB was the slowest of all them. However, the consistent gain in score was worth the 10-20% slowdown per epoch. Regarding the lr schedules, SGD with warm restarts and training for 7 epochs (1 + 2 + 4) gave the best results in turns of training time and the CV score when compared to the 1-cycle policy or cosine decay with warm-up as well as training for more epochs. Finally, using the cross-entropy with a light weighing of positive examples outperformed simple cross-entropy and the focal loss.\n\n### Train data\nIn my experiments adding the normalized old train and test data to the new train data resulted in a consistent increase by 0.02-0.03 in the local CV. As it was obvious that the new test data differed a lot from the train sets, using the soft labeled test data helped further improve the public LB score (at the same time it slightly decreased the local CV but at this point I trusted the LB score more that my CV).\n\n### Tools\nI did almost all my experiments in FastAI on my local computer with one RTX 3070. It was only in the last couple of weeks of the competition when I found a good way to upscale the inputs and that's when I ran out of my local resources (even though I upgraded my GPU to 3080Ti) and switched to TensorFlow and TPUs on Kaggle. In general, I find the amount of the train data quite suitable for fast experimentations even on a such relatively small machine. Also it was interesting to see that FastAI turned out to be faster than Lightning in several tests (at least in my hands :D), so I decided not to search for any alternatives and kept using FastAI almost exclusively. Finally, to follow all the experiments I used wandb and I should say it helped me a lot not to lose track of all the hundreds of experiments I ran.",
      "votes": 18
    },
    {
      "id": 1482023,
      "postDate": "2021-08-19T19:53:45.783Z",
      "content": "<p>FastAI needs to sponsor you as you won 2 gold medals with FastAI 😄</p>",
      "rawMarkdown": "FastAI needs to sponsor you as you won 2 gold medals with FastAI 😄",
      "votes": 4,
      "replies": [
        {
          "id": 1482638,
          "postDate": "2021-08-20T07:13:57.240Z",
          "content": "<p>Just winning gold medals without any sponsorship should also be fine 😆 </p>",
          "rawMarkdown": "Just winning gold medals without any sponsorship should also be fine 😆 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1483361,
      "postDate": "2021-08-20T15:40:14.953Z",
      "content": "<p>Well done! Nice demonstration of showing solo gold winning with such limited computing resource. </p>",
      "rawMarkdown": "Well done! Nice demonstration of showing solo gold winning with such limited computing resource. ",
      "votes": 1,
      "replies": [
        {
          "id": 1483564,
          "postDate": "2021-08-20T17:50:08.643Z",
          "content": "<p>I guess that 3070/80 is a threshold.   </p>",
          "rawMarkdown": "I guess that 3070/80 is a threshold.   "
        }
      ]
    },
    {
      "id": 1482209,
      "postDate": "2021-08-20T00:23:34.963Z",
      "content": "<p>Congrats on the solo gold!</p>",
      "rawMarkdown": "Congrats on the solo gold!",
      "votes": 2,
      "replies": [
        {
          "id": 1482640,
          "postDate": "2021-08-20T07:15:18.060Z",
          "content": "<p>Thanks! Hope more medals to come 😃</p>",
          "rawMarkdown": "Thanks! Hope more medals to come 😃",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1482023,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-08-19T19:53:45.783000",
      "content": "<p>FastAI needs to sponsor you as you won 2 gold medals with FastAI 😄</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1482638,
          "author_name": "Ilya Makarov",
          "author_url": "",
          "post_date": "2021-08-20T07:13:57.240000",
          "content": "<p>Just winning gold medals without any sponsorship should also be fine 😆 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1483361,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2021-08-20T15:40:14.953000",
      "content": "<p>Well done! Nice demonstration of showing solo gold winning with such limited computing resource. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1483564,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-08-20T17:50:08.643000",
          "content": "<p>I guess that 3070/80 is a threshold.   </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1482209,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-20T00:23:34.963000",
      "content": "<p>Congrats on the solo gold!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1482640,
          "author_name": "Ilya Makarov",
          "author_url": "",
          "post_date": "2021-08-20T07:15:18.060000",
          "content": "<p>Thanks! Hope more medals to come 😃</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1481926": "# TL;DR\nMy best scoring prediction is an average of eca_nfnet_l1 (384x384 and 420x420), EfficientNet-B2 (384x384), -B6 (512x512 and 640x640) and -B7 (640x640) predictions with 2x tta (original + HFlip). The models were trained with FastAI locally on my 3070 and later 3080Ti (eca_nfnet_l1 and B2) and with TF on Kaggle TPU (B6 and B7) with LAMB as optimizer, SGD with restarts as lr schedule and applying slight weighing in cross-entropy loss function (2). New train data, old train and test data normalized per tile as well as soft labeled new test data were used for training with MixUp and HFlip as augmentations. All 6 tiles of the cadence were stacked and resized in two steps: first, upsized to 1638x1638 with INTER_NEAREST interpolation and then downsized with INTER_AREA.\n\n# In detail\n### Selecting models\nAs was noted by other teams too, not all models with good local CV had correspondingly good public LB. The two families of models that turned out to perform well were eca_nfnet_l's and EfficientNets while resnets (regular and d ones), ecaresnets, resnests and some other had a lower LB than I would anticipate from their CV scores. Unfortunately, I did not manage to extract much useful insights from this discrepancy between the LB and CV... Also, unsurprisingly larger models performed slightly better than the smaller ones, at the same time, excluding the smaller models from the final ensemble did worsen the score a bit. The best single model was a 4-fold EfficientNet-B7 (640x640) with a public LB of 0.80092 and private LB 0.79648.\n\n### Image size and resizing\nFor the training I used all the 6 tiles of cadence because before the restart of the competition such an approach performed best. While after the restart I kept using the same approach as I saw no gain to a slight negative decline in the score when using just the 3 positive tiles. Also, very early in the competition I discovered that the score was pretty sensitive to the resizing method. When downsizing the images to 256x256, INTER_AREA interpolation was the method of choice as any other interpolation available in OpenCV was consistently giving a 0.02-0.03 lower AUC for any model tested. In addition, changing two dimensions at the same time (like 1638x256 -> 320x320) I could never obtain a better score than with 256x256 inputs regardless of the interpolation. The only way for me to get a better score for the images with a frequency dimension greater than 256 px was to first scale the original inputs to 1638x1638 with INTER_NEAREST interpolation and then resize these large images to any smaller size. When upscaling the inputs in this two-step way, the larger input sizes always resulted in larger CV and LB scores. At the end, the input sizes I used in my training was a compromise between the available resources and the score.\n\n### Optimizers and lr schedules\nLAMB optimizer turned out to give consistently better score (both with TF and FastAI) in comparison with any other optimizer I tried even though LAMB was the slowest of all them. However, the consistent gain in score was worth the 10-20% slowdown per epoch. Regarding the lr schedules, SGD with warm restarts and training for 7 epochs (1 + 2 + 4) gave the best results in turns of training time and the CV score when compared to the 1-cycle policy or cosine decay with warm-up as well as training for more epochs. Finally, using the cross-entropy with a light weighing of positive examples outperformed simple cross-entropy and the focal loss.\n\n### Train data\nIn my experiments adding the normalized old train and test data to the new train data resulted in a consistent increase by 0.02-0.03 in the local CV. As it was obvious that the new test data differed a lot from the train sets, using the soft labeled test data helped further improve the public LB score (at the same time it slightly decreased the local CV but at this point I trusted the LB score more that my CV).\n\n### Tools\nI did almost all my experiments in FastAI on my local computer with one RTX 3070. It was only in the last couple of weeks of the competition when I found a good way to upscale the inputs and that's when I ran out of my local resources (even though I upgraded my GPU to 3080Ti) and switched to TensorFlow and TPUs on Kaggle. In general, I find the amount of the train data quite suitable for fast experimentations even on a such relatively small machine. Also it was interesting to see that FastAI turned out to be faster than Lightning in several tests (at least in my hands :D), so I decided not to search for any alternatives and kept using FastAI almost exclusively. Finally, to follow all the experiments I used wandb and I should say it helped me a lot not to lose track of all the hundreds of experiments I ran.",
    "1482023": "FastAI needs to sponsor you as you won 2 gold medals with FastAI 😄",
    "1483361": "Well done! Nice demonstration of showing solo gold winning with such limited computing resource. ",
    "1482209": "Congrats on the solo gold!"
  }
}