{
  "id": 266884,
  "title": "Simple solution for solo silver (and one old GTX1080Ti)",
  "url": "/competitions/seti-breakthrough-listen/discussion/266884",
  "author_name": "Allie K.",
  "post_date": "2021-08-20T19:14:58.179000",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks to everybody who presented their great solutions to enable to others to learn from them!<br>\nAs I didn't notice anybody with similar simple tricks which I used, I decided to share my first humble overview.</p>\n<p>I didn't use ON channels only, but I cropped images with \"collars\" of OFF channels' parts. The first symmetric version:</p>\n<pre><code>img = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:] ))\nimg = img.transpose(1, 0) \n</code></pre>\n<p>and the second non-symmetric version (then couldn't use HFlip):</p>\n<pre><code>img = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:], img[5][:65][:] ))\nimg = img.transpose(1, 0) \n</code></pre>\n<p>It helps the network to exclude potential false positives - detecting a signal which continues to or from OFF channels.<br>\nAs I was struggling with limited image size (max. 512x512 for EfficientNet-B4) for my GTX, I was trying to add as few more pixels as possible. </p>\n<p>In the first part of the competition my models where suffering from high number of false negatives (partly due to low resolution) and I was considering generating some additional positive samples. After the competition reset there was hardly any (computation) time for doing this, so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.</p>\n<p>I used pytorch, StratifiedKfold(5) and after several short experiments I trained networks with Resnet34d and EfficientNetB4 backbones, pretrained weights, with mixup (alpha=0.19 or 0.42), HFlip, VFlip, ShiftScaleRotate, MotionBlur, RBContrast.<br>\nI tried some post-train ensembles but in the end my best model was single EfficientNet-B4 with non-symetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).</p>",
  "messages": [
    {
      "id": 1483682,
      "postDate": "2021-08-20T19:14:58.180Z",
      "content": "<p>Thanks to everybody who presented their great solutions to enable to others to learn from them!<br>\nAs I didn't notice anybody with similar simple tricks which I used, I decided to share my first humble overview.</p>\n<p>I didn't use ON channels only, but I cropped images with \"collars\" of OFF channels' parts. The first symmetric version:</p>\n<pre><code>img = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:] ))\nimg = img.transpose(1, 0) \n</code></pre>\n<p>and the second non-symmetric version (then couldn't use HFlip):</p>\n<pre><code>img = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:], img[5][:65][:] ))\nimg = img.transpose(1, 0) \n</code></pre>\n<p>It helps the network to exclude potential false positives - detecting a signal which continues to or from OFF channels.<br>\nAs I was struggling with limited image size (max. 512x512 for EfficientNet-B4) for my GTX, I was trying to add as few more pixels as possible. </p>\n<p>In the first part of the competition my models where suffering from high number of false negatives (partly due to low resolution) and I was considering generating some additional positive samples. After the competition reset there was hardly any (computation) time for doing this, so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.</p>\n<p>I used pytorch, StratifiedKfold(5) and after several short experiments I trained networks with Resnet34d and EfficientNetB4 backbones, pretrained weights, with mixup (alpha=0.19 or 0.42), HFlip, VFlip, ShiftScaleRotate, MotionBlur, RBContrast.<br>\nI tried some post-train ensembles but in the end my best model was single EfficientNet-B4 with non-symetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).</p>",
      "rawMarkdown": "Thanks to everybody who presented their great solutions to enable to others to learn from them!\nAs I didn't notice anybody with similar simple tricks which I used, I decided to share my first humble overview.\n\nI didn't use ON channels only, but I cropped images with \"collars\" of OFF channels' parts. The first symmetric version:\n```\nimg = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:] ))\nimg = img.transpose(1, 0) \n```\nand the second non-symmetric version (then couldn't use HFlip):\n```\nimg = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:], img[5][:65][:] ))\nimg = img.transpose(1, 0) \n```\nIt helps the network to exclude potential false positives - detecting a signal which continues to or from OFF channels.\nAs I was struggling with limited image size (max. 512x512 for EfficientNet-B4) for my GTX, I was trying to add as few more pixels as possible. \n\nIn the first part of the competition my models where suffering from high number of false negatives (partly due to low resolution) and I was considering generating some additional positive samples. After the competition reset there was hardly any (computation) time for doing this, so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.\n \nI used pytorch, StratifiedKfold(5) and after several short experiments I trained networks with Resnet34d and EfficientNetB4 backbones, pretrained weights, with mixup (alpha=0.19 or 0.42), HFlip, VFlip, ShiftScaleRotate, MotionBlur, RBContrast.\nI tried some post-train ensembles but in the end my best model was single EfficientNet-B4 with non-symetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\n",
      "votes": 10
    },
    {
      "id": 1485742,
      "postDate": "2021-08-22T11:38:08.280Z",
      "content": "<p>Congratulations. Great solo finish, well done!</p>\n<p>When you say </p>\n<blockquote>\n  <p>so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.</p>\n</blockquote>\n<p>Does this mean that you used more <code>target=1</code> from the data before competition reset? That's a smart idea.</p>",
      "rawMarkdown": "Congratulations. Great solo finish, well done!\n\nWhen you say \n>so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.\n\nDoes this mean that you used more `target=1` from the data before competition reset? That's a smart idea.",
      "votes": 2,
      "replies": [
        {
          "id": 1485784,
          "postDate": "2021-08-22T12:27:33.067Z",
          "content": "<p>Thank you for your interest.<br>\nThis exactly means that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.<br>\nI think that together with \"in batch mixup\" with batch size=8 it gave higher variety of how \"needles\" can look to the network. My results started getting significantly better when I introduced this new train set.</p>",
          "rawMarkdown": "Thank you for your interest.\nThis exactly means that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.\nI think that together with \"in batch mixup\" with batch size=8 it gave higher variety of how \"needles\" can look to the network. My results started getting significantly better when I introduced this new train set.",
          "votes": 3
        },
        {
          "id": 1485806,
          "postDate": "2021-08-22T12:42:51.313Z",
          "content": "<p>Yes that is a great idea. I was planning on exploring the original data but I ran out of time. Thanks for sharing.</p>",
          "rawMarkdown": "Yes that is a great idea. I was planning on exploring the original data but I ran out of time. Thanks for sharing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1485346,
      "postDate": "2021-08-22T03:02:51.643Z",
      "content": "<p><a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a> - would you be willing to share more information with Chris Deotte <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266813#1484947\" target=\"_blank\">here</a> ?  He was looking for models that scored over 0.78 and had just read your post so put a link to here as this seemed a novel idea.</p>\n<p>Mentioned your idea of non symmetric cropping and private 0.78191 result for efn b4 and he had some questions. Like public score for it. And scores for models selected.  Thanks!    </p>",
      "rawMarkdown": "@blankaf - would you be willing to share more information with Chris Deotte [here](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266813#1484947) ?  He was looking for models that scored over 0.78 and had just read your post so put a link to here as this seemed a novel idea.\n\nMentioned your idea of non symmetric cropping and private 0.78191 result for efn b4 and he had some questions. Like public score for it. And scores for models selected.  Thanks!    ",
      "votes": 2
    },
    {
      "id": 1483769,
      "postDate": "2021-08-20T20:55:23.150Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1485742,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-22T11:38:08.280000",
      "content": "<p>Congratulations. Great solo finish, well done!</p>\n<p>When you say </p>\n<blockquote>\n  <p>so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.</p>\n</blockquote>\n<p>Does this mean that you used more <code>target=1</code> from the data before competition reset? That's a smart idea.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1485784,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2021-08-22T12:27:33.067000",
          "content": "<p>Thank you for your interest.<br>\nThis exactly means that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.<br>\nI think that together with \"in batch mixup\" with batch size=8 it gave higher variety of how \"needles\" can look to the network. My results started getting significantly better when I introduced this new train set.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1485806,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-22T12:42:51.313000",
          "content": "<p>Yes that is a great idea. I was planning on exploring the original data but I ran out of time. Thanks for sharing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1485346,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-08-22T03:02:51.643000",
      "content": "<p><a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a> - would you be willing to share more information with Chris Deotte <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266813#1484947\" target=\"_blank\">here</a> ?  He was looking for models that scored over 0.78 and had just read your post so put a link to here as this seemed a novel idea.</p>\n<p>Mentioned your idea of non symmetric cropping and private 0.78191 result for efn b4 and he had some questions. Like public score for it. And scores for models selected.  Thanks!    </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1483769,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-20T20:55:23.150000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1483682": "Thanks to everybody who presented their great solutions to enable to others to learn from them!\nAs I didn't notice anybody with similar simple tricks which I used, I decided to share my first humble overview.\n\nI didn't use ON channels only, but I cropped images with \"collars\" of OFF channels' parts. The first symmetric version:\n```\nimg = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:] ))\nimg = img.transpose(1, 0) \n```\nand the second non-symmetric version (then couldn't use HFlip):\n```\nimg = np.vstack((img[0][:][:], img[1][:35][:], img[1][-35:][:], img[2][:][:],                             img[3][:35][:], img[3][-35:][:], img[4][:][:], img[5][:65][:] ))\nimg = img.transpose(1, 0) \n```\nIt helps the network to exclude potential false positives - detecting a signal which continues to or from OFF channels.\nAs I was struggling with limited image size (max. 512x512 for EfficientNet-B4) for my GTX, I was trying to add as few more pixels as possible. \n\nIn the first part of the competition my models where suffering from high number of false negatives (partly due to low resolution) and I was considering generating some additional positive samples. After the competition reset there was hardly any (computation) time for doing this, so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.\n \nI used pytorch, StratifiedKfold(5) and after several short experiments I trained networks with Resnet34d and EfficientNetB4 backbones, pretrained weights, with mixup (alpha=0.19 or 0.42), HFlip, VFlip, ShiftScaleRotate, MotionBlur, RBContrast.\nI tried some post-train ensembles but in the end my best model was single EfficientNet-B4 with non-symetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\n",
    "1485742": "Congratulations. Great solo finish, well done!\n\nWhen you say \n>so I decided to simply add (cast and normalized) positive only original train and test samples. I got an even bigger CV-LB gap but it helped climbing the leaderboard.\n\nDoes this mean that you used more `target=1` from the data before competition reset? That's a smart idea.",
    "1485346": "@blankaf - would you be willing to share more information with Chris Deotte [here](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266813#1484947) ?  He was looking for models that scored over 0.78 and had just read your post so put a link to here as this seemed a novel idea.\n\nMentioned your idea of non symmetric cropping and private 0.78191 result for efn b4 and he had some questions. Like public score for it. And scores for models selected.  Thanks!    ",
    "1483769": ""
  }
}