{
  "id": 244461,
  "title": " Seti-Data dataset to augment positive examples",
  "url": "/competitions/seti-breakthrough-listen/discussion/244461",
  "author_name": "AleNic",
  "post_date": "2021-06-06T21:56:23.819000",
  "votes": 38,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I found this interesting dataset:</p>\n<p><a href=\"https://www.kaggle.com/tentotheminus9/seti-data\" target=\"_blank\">https://www.kaggle.com/tentotheminus9/seti-data</a></p>\n<p>I used the dataset images (but noise class) together with other negative slices to augment the number of positive examples, and seems that it works well!<br>\nHope it can help! :)</p>",
  "messages": [
    {
      "id": 1338944,
      "postDate": "2021-06-06T21:56:23.820Z",
      "content": "<p>I found this interesting dataset:</p>\n<p><a href=\"https://www.kaggle.com/tentotheminus9/seti-data\" target=\"_blank\">https://www.kaggle.com/tentotheminus9/seti-data</a></p>\n<p>I used the dataset images (but noise class) together with other negative slices to augment the number of positive examples, and seems that it works well!<br>\nHope it can help! :)</p>",
      "rawMarkdown": "I found this interesting dataset:\n\nhttps://www.kaggle.com/tentotheminus9/seti-data\n\nI used the dataset images (but noise class) together with other negative slices to augment the number of positive examples, and seems that it works well!\nHope it can help! :)",
      "votes": 38
    },
    {
      "id": 1342188,
      "postDate": "2021-06-09T08:47:36.483Z",
      "content": "<p>I've just tried it and got similar results for a single model:</p>\n<ul>\n<li>Without external, OOF = 0.9902</li>\n<li>With external, OOF = 0.9899</li>\n</ul>\n<p>Brightpixel removed too. Interesting anyway.</p>",
      "rawMarkdown": "I've just tried it and got similar results for a single model:\n- Without external, OOF = 0.9902\n- With external, OOF = 0.9899\n\nBrightpixel removed too. Interesting anyway.",
      "votes": 3,
      "replies": [
        {
          "id": 1342204,
          "postDate": "2021-06-09T09:07:40.123Z",
          "content": "<p>in my case, oof score decreased, 0.993 -&gt; .983<br>\nand with small lr ( 0.17 -&gt; 0.02) .983 -&gt; .995 (but LB .96)</p>",
          "rawMarkdown": "in my case, oof score decreased, 0.993 -> .983\nand with small lr ( 0.17 -> 0.02) .983 -> .995 (but LB .96)",
          "votes": 2
        },
        {
          "id": 1342302,
          "postDate": "2021-06-09T10:36:20.507Z",
          "content": "<p>After getting 0.995 oof, you must have been disappointed when you saw 0.96 LB :D<br>\nDo you use 0.17 learning rate? Isn't that a little huge?</p>",
          "rawMarkdown": "After getting 0.995 oof, you must have been disappointed when you saw 0.96 LB :D\nDo you use 0.17 learning rate? Isn't that a little huge?",
          "votes": 1
        },
        {
          "id": 1342572,
          "postDate": "2021-06-09T14:53:58.233Z",
          "content": "<p>I guess 0.995 is CV (taking into account the external data in validation score). What about if you remove external data when you compute OOF (after everything is trained)? So you can really compare with your previous models. It should drop from what you're reporting on LB.</p>",
          "rawMarkdown": "I guess 0.995 is CV (taking into account the external data in validation score). What about if you remove external data when you compute OOF (after everything is trained)? So you can really compare with your previous models. It should drop from what you're reporting on LB.",
          "votes": 3
        },
        {
          "id": 1343249,
          "postDate": "2021-06-10T06:00:09.633Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> in my case, I use SGD with batch size 2^10, just 46 steps per epochs<br>\ndue to large batchsize, I use large lr.</p>\n<p>in addition, without lr scheduling but only increase momentum .9 -&gt; .995</p>",
          "rawMarkdown": "@nofreewill in my case, I use SGD with batch size 2^10, just 46 steps per epochs\ndue to large batchsize, I use large lr.\n\nin addition, without lr scheduling but only increase momentum .9 -> .995",
          "votes": 1
        }
      ]
    },
    {
      "id": 1341475,
      "postDate": "2021-06-08T16:56:21.167Z",
      "content": "<p>Thx for sharing!</p>",
      "rawMarkdown": "Thx for sharing!",
      "votes": 1
    },
    {
      "id": 1343151,
      "postDate": "2021-06-10T04:38:11.467Z",
      "content": "<p>Thanks for your sharing! I would like to ask whether the images in all folders（narrowband、narrowbanddrd、squarepulsednarrowband、squiggle and squigglesquarepulsednarrowband ) should be input into our datasets as positive samples？</p>",
      "rawMarkdown": "Thanks for your sharing! I would like to ask whether the images in all folders（narrowband、narrowbanddrd、squarepulsednarrowband、squiggle and squigglesquarepulsednarrowband ) should be input into our datasets as positive samples？"
    },
    {
      "id": 1340446,
      "postDate": "2021-06-08T00:28:09.527Z",
      "content": "<p>Thank you for sharing valuable information. It must be useful.<br>\nAre there png files only? If so, how do you read them as 1 channel numpy data?<br>\nIn addition, do you use brightpixel class? I am not sure they are in the competition dataset.</p>",
      "rawMarkdown": "Thank you for sharing valuable information. It must be useful.\nAre there png files only? If so, how do you read them as 1 channel numpy data?\nIn addition, do you use brightpixel class? I am not sure they are in the competition dataset.",
      "replies": [
        {
          "id": 1340613,
          "postDate": "2021-06-08T06:20:25.303Z",
          "content": "<p>There are only png images, and you can use it to combine them with the other (transformed from npy to images) of the original dataset to generate other positives. I don't use brightpixel and noise classes.</p>",
          "rawMarkdown": "There are only png images, and you can use it to combine them with the other (transformed from npy to images) of the original dataset to generate other positives. I don't use brightpixel and noise classes.",
          "votes": 2
        },
        {
          "id": 1341009,
          "postDate": "2021-06-08T11:11:54.130Z",
          "content": "<p>Thanks. I will try it.</p>",
          "rawMarkdown": "Thanks. I will try it."
        }
      ]
    },
    {
      "id": 1338949,
      "postDate": "2021-06-06T22:00:35.213Z",
      "content": "<p>This looks helpful, thank you!</p>",
      "rawMarkdown": "This looks helpful, thank you!",
      "votes": 1
    },
    {
      "id": 1344956,
      "postDate": "2021-06-11T08:17:35.357Z",
      "content": "<p>This looks helpful, thank you👍</p>",
      "rawMarkdown": "This looks helpful, thank you👍"
    }
  ],
  "comments": [
    {
      "id": 1342188,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-06-09T08:47:36.483000",
      "content": "<p>I've just tried it and got similar results for a single model:</p>\n<ul>\n<li>Without external, OOF = 0.9902</li>\n<li>With external, OOF = 0.9899</li>\n</ul>\n<p>Brightpixel removed too. Interesting anyway.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1342204,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-09T09:07:40.123000",
          "content": "<p>in my case, oof score decreased, 0.993 -&gt; .983<br>\nand with small lr ( 0.17 -&gt; 0.02) .983 -&gt; .995 (but LB .96)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1342302,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-09T10:36:20.507000",
          "content": "<p>After getting 0.995 oof, you must have been disappointed when you saw 0.96 LB :D<br>\nDo you use 0.17 learning rate? Isn't that a little huge?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1342572,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2021-06-09T14:53:58.233000",
          "content": "<p>I guess 0.995 is CV (taking into account the external data in validation score). What about if you remove external data when you compute OOF (after everything is trained)? So you can really compare with your previous models. It should drop from what you're reporting on LB.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1343249,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-10T06:00:09.633000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> in my case, I use SGD with batch size 2^10, just 46 steps per epochs<br>\ndue to large batchsize, I use large lr.</p>\n<p>in addition, without lr scheduling but only increase momentum .9 -&gt; .995</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1341475,
      "author_name": "cadog",
      "author_url": "",
      "post_date": "2021-06-08T16:56:21.167000",
      "content": "<p>Thx for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1343151,
      "author_name": "maotree",
      "author_url": "",
      "post_date": "2021-06-10T04:38:11.467000",
      "content": "<p>Thanks for your sharing! I would like to ask whether the images in all folders（narrowband、narrowbanddrd、squarepulsednarrowband、squiggle and squigglesquarepulsednarrowband ) should be input into our datasets as positive samples？</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1340446,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-06-08T00:28:09.527000",
      "content": "<p>Thank you for sharing valuable information. It must be useful.<br>\nAre there png files only? If so, how do you read them as 1 channel numpy data?<br>\nIn addition, do you use brightpixel class? I am not sure they are in the competition dataset.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1340613,
          "author_name": "AleNic",
          "author_url": "",
          "post_date": "2021-06-08T06:20:25.303000",
          "content": "<p>There are only png images, and you can use it to combine them with the other (transformed from npy to images) of the original dataset to generate other positives. I don't use brightpixel and noise classes.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1341009,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-08T11:11:54.130000",
          "content": "<p>Thanks. I will try it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1338949,
      "author_name": "nofreewill42",
      "author_url": "",
      "post_date": "2021-06-06T22:00:35.213000",
      "content": "<p>This looks helpful, thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1344956,
      "author_name": "Mohamed Bakrey Mahmoud",
      "author_url": "",
      "post_date": "2021-06-11T08:17:35.357000",
      "content": "<p>This looks helpful, thank you👍</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1338944": "I found this interesting dataset:\n\nhttps://www.kaggle.com/tentotheminus9/seti-data\n\nI used the dataset images (but noise class) together with other negative slices to augment the number of positive examples, and seems that it works well!\nHope it can help! :)",
    "1342188": "I've just tried it and got similar results for a single model:\n- Without external, OOF = 0.9902\n- With external, OOF = 0.9899\n\nBrightpixel removed too. Interesting anyway.",
    "1341475": "Thx for sharing!",
    "1343151": "Thanks for your sharing! I would like to ask whether the images in all folders（narrowband、narrowbanddrd、squarepulsednarrowband、squiggle and squigglesquarepulsednarrowband ) should be input into our datasets as positive samples？",
    "1340446": "Thank you for sharing valuable information. It must be useful.\nAre there png files only? If so, how do you read them as 1 channel numpy data?\nIn addition, do you use brightpixel class? I am not sure they are in the competition dataset.",
    "1338949": "This looks helpful, thank you!",
    "1344956": "This looks helpful, thank you👍"
  }
}