{
  "id": 265251,
  "title": "what's the key to this competition",
  "url": "/competitions/seti-breakthrough-listen/discussion/265251",
  "author_name": "Yi Wu",
  "post_date": "2021-08-15T07:35:54.222000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am very frustrated to figure out the key to this competition, and I think ensemble cannot take us to bronze area. according previous discussion, I think open set recognition should be extremely important.</p>",
  "messages": [
    {
      "id": 1473281,
      "postDate": "2021-08-15T13:37:48.360Z",
      "content": "<p>Agree that open set regcognition is very important, without it I believe it's hard to reach gold area, as it's told by orgnizer there is a pattern of target only exist in test data. And IMHO another <br>\n key is that the data is super noisy, especially in new data many target is not visible by human eye. I guess the winning solution should be using composite model like this <a href=\"https://arxiv.org/pdf/1901.04636v1.pdf\" target=\"_blank\">paper</a>, pure CNN model has limitaion for this competition. Let`s wait and see after competition.</p>",
      "rawMarkdown": "Agree that open set regcognition is very important, without it I believe it's hard to reach gold area, as it's told by orgnizer there is a pattern of target only exist in test data. And IMHO another \n key is that the data is super noisy, especially in new data many target is not visible by human eye. I guess the winning solution should be using composite model like this [paper](https://arxiv.org/pdf/1901.04636v1.pdf), pure CNN model has limitaion for this competition. Let`s wait and see after competition.",
      "votes": 5
    },
    {
      "id": 1473544,
      "postDate": "2021-08-15T15:45:17.440Z",
      "content": "<p>Basically, the train dataset and test dataset are consist of<br>\nneedle image + noise. <br>\nThe pattern of the noise is different between train and test.</p>\n<p>The effect of the noise is reduced by convolution layer.<br>\nThis is why large model well work then small model.</p>\n<p>train_old, test_old, train_new actually same pattern.<br>\ntrain_new have additional needle pattern than these two old dataset. <br>\n(like bright pixel pattern)</p>\n<p>Test dataset have different noise pattern and brightness than others.<br>\ni.e. test dataset have more needle images that in decision boundary.<br>\n(low standard deviation than trainning dataset before standardization)</p>\n<p>Ensemble make decision boundary much better.<br>\nWe can use 90 rotation image for another training set.</p>\n<p>Also, I think large batch size is more stable. <br>\nI believe that the different of performance come from batchsize.<br>\nTransfer learning on image size is well work. <br>\nSo, I think extend the image size and keep large batch size is the key point.<br>\nWe also can try SGD with gradient cliping than AdamW.</p>\n<p>Finally, the thickness of the needle image is too small,<br>\nSo we can make them bigger by taking low resolution with max pooling or<br>\nsomething denoising filter.</p>\n<p>I also don't know the key point.<br>\nI am trying to modify squeeze layer in efficientnet <br>\nbecause It take mean.. But It's too difficult to me.</p>",
      "rawMarkdown": "Basically, the train dataset and test dataset are consist of\nneedle image + noise. \nThe pattern of the noise is different between train and test.\n\nThe effect of the noise is reduced by convolution layer.\nThis is why large model well work then small model.\n\ntrain_old, test_old, train_new actually same pattern.\ntrain_new have additional needle pattern than these two old dataset. \n(like bright pixel pattern)\n\nTest dataset have different noise pattern and brightness than others.\ni.e. test dataset have more needle images that in decision boundary.\n(low standard deviation than trainning dataset before standardization)\n\nEnsemble make decision boundary much better.\nWe can use 90 rotation image for another training set.\n\nAlso, I think large batch size is more stable. \nI believe that the different of performance come from batchsize.\nTransfer learning on image size is well work. \nSo, I think extend the image size and keep large batch size is the key point.\nWe also can try SGD with gradient cliping than AdamW.\n\nFinally, the thickness of the needle image is too small,\nSo we can make them bigger by taking low resolution with max pooling or\nsomething denoising filter.\n\nI also don't know the key point.\nI am trying to modify squeeze layer in efficientnet \nbecause It take mean.. But It's too difficult to me.",
      "votes": 3,
      "replies": [
        {
          "id": 1475073,
          "postDate": "2021-08-16T12:31:46.933Z",
          "content": "<blockquote>\n  <p>Test dataset have different noise pattern and brightness than others.<br>\n  … (low standard deviation than trainning dataset before standardization)</p>\n</blockquote>\n<p>std is the same in new test and new train, see this:  <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253613\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253613</a></p>\n<p>Can yo be more specific about the difference you see?</p>",
          "rawMarkdown": "> Test dataset have different noise pattern and brightness than others.\n> ... (low standard deviation than trainning dataset before standardization)\n\nstd is the same in new test and new train, see this:  https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253613\n\nCan yo be more specific about the difference you see?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1472881,
      "postDate": "2021-08-15T07:35:54.223Z",
      "content": "<p>I am very frustrated to figure out the key to this competition, and I think ensemble cannot take us to bronze area. according previous discussion, I think open set recognition should be extremely important.</p>",
      "rawMarkdown": "I am very frustrated to figure out the key to this competition, and I think ensemble cannot take us to bronze area. according previous discussion, I think open set recognition should be extremely important.",
      "votes": 3
    },
    {
      "id": 1475738,
      "postDate": "2021-08-16T20:13:50.027Z",
      "content": "<p>I wish I knew.</p>",
      "rawMarkdown": "I wish I knew.",
      "votes": 1
    },
    {
      "id": 1473026,
      "postDate": "2021-08-15T09:48:47.350Z",
      "content": "<p>The ensemble can give you a bronze zone at this stage. Moreover, an estimate of 0.771 can be obtained even for one model, in which there is nothing special. The key importance in this competition is primarily validation and augmentation.</p>",
      "rawMarkdown": "The ensemble can give you a bronze zone at this stage. Moreover, an estimate of 0.771 can be obtained even for one model, in which there is nothing special. The key importance in this competition is primarily validation and augmentation.",
      "votes": 2,
      "replies": [
        {
          "id": 1473280,
          "postDate": "2021-08-15T13:37:30.097Z",
          "content": "<p>Climbing the LB is very easy.</p>\n<p>The problem is, everyone else keeps climbing too… :) Why can't they stay still!</p>",
          "rawMarkdown": "Climbing the LB is very easy.\n\nThe problem is, everyone else keeps climbing too... :) Why can't they stay still!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1473449,
      "postDate": "2021-08-15T15:15:12.733Z",
      "content": "<p>I don't understand the key to this competition, too……</p>\n<p>However, single 5fold model using only new train dataset lead to 0.772 in LB and ensemble didn't improve score largely.<br>\nI believe that other model capacity and large shake will not happen.</p>",
      "rawMarkdown": "I don't understand the key to this competition, too......\n\nHowever, single 5fold model using only new train dataset lead to 0.772 in LB and ensemble didn't improve score largely.\nI believe that other model capacity and large shake will not happen."
    },
    {
      "id": 1473451,
      "postDate": "2021-08-15T15:16:55.623Z",
      "content": "<p>I think the key is combined CNN+LSTM model, like one described here <a href=\"https://www.nature.com/articles/s41598-019-45748-1\" target=\"_blank\">https://www.nature.com/articles/s41598-019-45748-1</a>.</p>",
      "rawMarkdown": "I think the key is combined CNN+LSTM model, like one described here https://www.nature.com/articles/s41598-019-45748-1."
    }
  ],
  "comments": [
    {
      "id": 1473281,
      "author_name": "Hao",
      "author_url": "",
      "post_date": "2021-08-15T13:37:48.360000",
      "content": "<p>Agree that open set regcognition is very important, without it I believe it's hard to reach gold area, as it's told by orgnizer there is a pattern of target only exist in test data. And IMHO another <br>\n key is that the data is super noisy, especially in new data many target is not visible by human eye. I guess the winning solution should be using composite model like this <a href=\"https://arxiv.org/pdf/1901.04636v1.pdf\" target=\"_blank\">paper</a>, pure CNN model has limitaion for this competition. Let`s wait and see after competition.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1473544,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-08-15T15:45:17.440000",
      "content": "<p>Basically, the train dataset and test dataset are consist of<br>\nneedle image + noise. <br>\nThe pattern of the noise is different between train and test.</p>\n<p>The effect of the noise is reduced by convolution layer.<br>\nThis is why large model well work then small model.</p>\n<p>train_old, test_old, train_new actually same pattern.<br>\ntrain_new have additional needle pattern than these two old dataset. <br>\n(like bright pixel pattern)</p>\n<p>Test dataset have different noise pattern and brightness than others.<br>\ni.e. test dataset have more needle images that in decision boundary.<br>\n(low standard deviation than trainning dataset before standardization)</p>\n<p>Ensemble make decision boundary much better.<br>\nWe can use 90 rotation image for another training set.</p>\n<p>Also, I think large batch size is more stable. <br>\nI believe that the different of performance come from batchsize.<br>\nTransfer learning on image size is well work. <br>\nSo, I think extend the image size and keep large batch size is the key point.<br>\nWe also can try SGD with gradient cliping than AdamW.</p>\n<p>Finally, the thickness of the needle image is too small,<br>\nSo we can make them bigger by taking low resolution with max pooling or<br>\nsomething denoising filter.</p>\n<p>I also don't know the key point.<br>\nI am trying to modify squeeze layer in efficientnet <br>\nbecause It take mean.. But It's too difficult to me.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1475073,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-16T12:31:46.933000",
          "content": "<blockquote>\n  <p>Test dataset have different noise pattern and brightness than others.<br>\n  … (low standard deviation than trainning dataset before standardization)</p>\n</blockquote>\n<p>std is the same in new test and new train, see this:  <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253613\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253613</a></p>\n<p>Can yo be more specific about the difference you see?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1475738,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-08-16T20:13:50.027000",
      "content": "<p>I wish I knew.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1473026,
      "author_name": "Aristarkh_BFG",
      "author_url": "",
      "post_date": "2021-08-15T09:48:47.350000",
      "content": "<p>The ensemble can give you a bronze zone at this stage. Moreover, an estimate of 0.771 can be obtained even for one model, in which there is nothing special. The key importance in this competition is primarily validation and augmentation.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1473280,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-08-15T13:37:30.097000",
          "content": "<p>Climbing the LB is very easy.</p>\n<p>The problem is, everyone else keeps climbing too… :) Why can't they stay still!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1473449,
      "author_name": "imori",
      "author_url": "",
      "post_date": "2021-08-15T15:15:12.733000",
      "content": "<p>I don't understand the key to this competition, too……</p>\n<p>However, single 5fold model using only new train dataset lead to 0.772 in LB and ensemble didn't improve score largely.<br>\nI believe that other model capacity and large shake will not happen.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1473451,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2021-08-15T15:16:55.623000",
      "content": "<p>I think the key is combined CNN+LSTM model, like one described here <a href=\"https://www.nature.com/articles/s41598-019-45748-1\" target=\"_blank\">https://www.nature.com/articles/s41598-019-45748-1</a>.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1473281": "Agree that open set regcognition is very important, without it I believe it's hard to reach gold area, as it's told by orgnizer there is a pattern of target only exist in test data. And IMHO another \n key is that the data is super noisy, especially in new data many target is not visible by human eye. I guess the winning solution should be using composite model like this [paper](https://arxiv.org/pdf/1901.04636v1.pdf), pure CNN model has limitaion for this competition. Let`s wait and see after competition.",
    "1473544": "Basically, the train dataset and test dataset are consist of\nneedle image + noise. \nThe pattern of the noise is different between train and test.\n\nThe effect of the noise is reduced by convolution layer.\nThis is why large model well work then small model.\n\ntrain_old, test_old, train_new actually same pattern.\ntrain_new have additional needle pattern than these two old dataset. \n(like bright pixel pattern)\n\nTest dataset have different noise pattern and brightness than others.\ni.e. test dataset have more needle images that in decision boundary.\n(low standard deviation than trainning dataset before standardization)\n\nEnsemble make decision boundary much better.\nWe can use 90 rotation image for another training set.\n\nAlso, I think large batch size is more stable. \nI believe that the different of performance come from batchsize.\nTransfer learning on image size is well work. \nSo, I think extend the image size and keep large batch size is the key point.\nWe also can try SGD with gradient cliping than AdamW.\n\nFinally, the thickness of the needle image is too small,\nSo we can make them bigger by taking low resolution with max pooling or\nsomething denoising filter.\n\nI also don't know the key point.\nI am trying to modify squeeze layer in efficientnet \nbecause It take mean.. But It's too difficult to me.",
    "1472881": "I am very frustrated to figure out the key to this competition, and I think ensemble cannot take us to bronze area. according previous discussion, I think open set recognition should be extremely important.",
    "1475738": "I wish I knew.",
    "1473026": "The ensemble can give you a bronze zone at this stage. Moreover, an estimate of 0.771 can be obtained even for one model, in which there is nothing special. The key importance in this competition is primarily validation and augmentation.",
    "1473449": "I don't understand the key to this competition, too......\n\nHowever, single 5fold model using only new train dataset lead to 0.772 in LB and ensemble didn't improve score largely.\nI believe that other model capacity and large shake will not happen.",
    "1473451": "I think the key is combined CNN+LSTM model, like one described here https://www.nature.com/articles/s41598-019-45748-1."
  }
}