{
  "id": 244379,
  "title": "Does TestSet contain \"unknown\" messages? ",
  "url": "/competitions/seti-breakthrough-listen/discussion/244379",
  "author_name": "Tawara",
  "post_date": "2021-06-06T14:04:10.611000",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>From the problem setup, we are asked to find messages from the unknown.<br>\nSo isn't it natural that \"unknown\" messages are included in the Private Test Set?</p>\n<p>The first thing I tired is to compare distributions of TrainSet and TestSet:  <br>\n<a href=\"https://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings\" target=\"_blank\">https://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings</a></p>\n<p>I extract images' embeddings by using my baseline then perform dimensionality reduction on them by TSNE.</p>\n<p>The first thing I found is that positive examples(\"needles\") and negative examples in TrainSet are almost separable as the following plots:  </p>\n<p><img src=\"https://pbs.twimg.com/media/E3NGp6WVEAEy7Xe?format=png\" alt=\"needles v.s. non-needles\"></p>\n<p>The second thing is that distribution of image embeddings for TrainSet and TestSet are almost identical as follows:  </p>\n<p><img src=\"https://pbs.twimg.com/media/E3NGundVEAIEt7W?format=jpg\" alt=\"Train v.s. Test\"><br>\n<br></p>\n<p>There seems no difference between TrainSet and TestSet, at least <strong>for the distribution of image embeddings</strong> extracted by CNN.  <br>\nI think, however, it is possible that there are unknown messages hidden in the TestSet.</p>",
  "messages": [
    {
      "id": 1338515,
      "postDate": "2021-06-06T14:04:10.610Z",
      "content": "<p>From the problem setup, we are asked to find messages from the unknown.<br>\nSo isn't it natural that \"unknown\" messages are included in the Private Test Set?</p>\n<p>The first thing I tired is to compare distributions of TrainSet and TestSet:  <br>\n<a href=\"https://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings\" target=\"_blank\">https://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings</a></p>\n<p>I extract images' embeddings by using my baseline then perform dimensionality reduction on them by TSNE.</p>\n<p>The first thing I found is that positive examples(\"needles\") and negative examples in TrainSet are almost separable as the following plots:  </p>\n<p><img src=\"https://pbs.twimg.com/media/E3NGp6WVEAEy7Xe?format=png\" alt=\"needles v.s. non-needles\"></p>\n<p>The second thing is that distribution of image embeddings for TrainSet and TestSet are almost identical as follows:  </p>\n<p><img src=\"https://pbs.twimg.com/media/E3NGundVEAIEt7W?format=jpg\" alt=\"Train v.s. Test\"><br>\n<br></p>\n<p>There seems no difference between TrainSet and TestSet, at least <strong>for the distribution of image embeddings</strong> extracted by CNN.  <br>\nI think, however, it is possible that there are unknown messages hidden in the TestSet.</p>",
      "rawMarkdown": "From the problem setup, we are asked to find messages from the unknown.\nSo isn't it natural that \"unknown\" messages are included in the Private Test Set?\n\nThe first thing I tired is to compare distributions of TrainSet and TestSet:  \nhttps://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings\n\nI extract images' embeddings by using my baseline then perform dimensionality reduction on them by TSNE.\n\nThe first thing I found is that positive examples(\"needles\") and negative examples in TrainSet are almost separable as the following plots:  \n\n![needles v.s. non-needles](https://pbs.twimg.com/media/E3NGp6WVEAEy7Xe?format=png)\n\nThe second thing is that distribution of image embeddings for TrainSet and TestSet are almost identical as follows:  \n\n![Train v.s. Test](https://pbs.twimg.com/media/E3NGundVEAIEt7W?format=jpg)\n<br>\n\nThere seems no difference between TrainSet and TestSet, at least **for the distribution of image embeddings** extracted by CNN.  \nI think, however, it is possible that there are unknown messages hidden in the TestSet.",
      "votes": 24
    },
    {
      "id": 1339990,
      "postDate": "2021-06-07T15:21:48.557Z",
      "content": "<p>I have made an adversarial test on train/test data and it seems that distributions are nearly completely the same. ROC AUC ~ 51%</p>",
      "rawMarkdown": "I have made an adversarial test on train/test data and it seems that distributions are nearly completely the same. ROC AUC ~ 51%\n",
      "votes": 3
    },
    {
      "id": 1338538,
      "postDate": "2021-06-06T14:34:35.950Z",
      "content": "<p>It feels like there's something different between train and test, to me. Because certain model types give me 0.98 every time I tried, and certain give me 0.97, with almost identical cross-validation performance.</p>",
      "rawMarkdown": "It feels like there's something different between train and test, to me. Because certain model types give me 0.98 every time I tried, and certain give me 0.97, with almost identical cross-validation performance.\n",
      "votes": 3,
      "replies": [
        {
          "id": 1338565,
          "postDate": "2021-06-06T14:48:29.190Z",
          "content": "<p>One possible reason I think is that the Public Test is biased and either one of those model types is (not) good at the \"biased\" Public Test.</p>\n<p>Did you compare the prediction from different model types?</p>",
          "rawMarkdown": "One possible reason I think is that the Public Test is biased and either one of those model types is (not) good at the \"biased\" Public Test.\n\nDid you compare the prediction from different model types?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1477409,
      "postDate": "2021-08-17T13:17:57.960Z",
      "content": "<p>I don't say a lot, but test distribution is far from train<br>\nthere maybe interesting stuff at final LB</p>",
      "rawMarkdown": "I don't say a lot, but test distribution is far from train\nthere maybe interesting stuff at final LB",
      "votes": -1,
      "replies": [
        {
          "id": 1477659,
          "postDate": "2021-08-17T15:15:50.233Z",
          "content": "<p>The host says public/private split is complete random, do you think still there will be something intersting?</p>",
          "rawMarkdown": "The host says public/private split is complete random, do you think still there will be something intersting?"
        },
        {
          "id": 1477846,
          "postDate": "2021-08-17T17:06:24.333Z",
          "content": "<p>yeah, let's see</p>",
          "rawMarkdown": "yeah, let's see",
          "votes": -1
        },
        {
          "id": 1481574,
          "postDate": "2021-08-19T15:29:14.947Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryunosukeishizaki\" target=\"_blank\">@ryunosukeishizaki</a>  After all, what was \"interesting\"? I would like you to find out why I shake down and let me know so that I can make use of it next time.</p>",
          "rawMarkdown": "@ryunosukeishizaki  After all, what was \"interesting\"? I would like you to find out why I shake down and let me know so that I can make use of it next time."
        },
        {
          "id": 1481611,
          "postDate": "2021-08-19T15:45:27.177Z",
          "content": "<p>Hi,<br>\nsince I joined this competition 5days ago, I believed reverse engineering to generate unknown message (based on hosts paper) was the way to easily get higher, private dataset was mentioned to contain more unknown signals, &amp; test distribution was usually much different from train(for weighted average histogram-based feature, train data consists of 2 types of distribution, and at test was domained by mainly 1 of distribution.<br>\nthen what I was thinking was guys who directly optimizing unknown \"needle\" information will be  still robustly at top, &amp; guys who don't mind about it will relatively shakedown,<br>\nbut it was still funny for me to see 1st place denoising solution,</p>",
          "rawMarkdown": "Hi,\nsince I joined this competition 5days ago, I believed reverse engineering to generate unknown message (based on hosts paper) was the way to easily get higher, private dataset was mentioned to contain more unknown signals, & test distribution was usually much different from train(for weighted average histogram-based feature, train data consists of 2 types of distribution, and at test was domained by mainly 1 of distribution.\nthen what I was thinking was guys who directly optimizing unknown \"needle\" information will be  still robustly at top, & guys who don't mind about it will relatively shakedown,\nbut it was still funny for me to see 1st place denoising solution,"
        }
      ]
    },
    {
      "id": 1338545,
      "postDate": "2021-06-06T14:39:05.607Z",
      "content": "<p>Isn't it because the model is well trained?<br>\nIf it shows different aspects in the input space of train and test, it can be seen that it has been trained incorrectly…<br>\nand my model make significantly different distribution (LB .94 model)</p>",
      "rawMarkdown": "Isn't it because the model is well trained?\nIf it shows different aspects in the input space of train and test, it can be seen that it has been trained incorrectly...\nand my model make significantly different distribution (LB .94 model)",
      "replies": [
        {
          "id": 1338579,
          "postDate": "2021-06-06T14:55:42.943Z",
          "content": "<p>Thanks.</p>\n<blockquote>\n  <p>my model make significantly different distribution</p>\n</blockquote>\n<p>By your model, are the distributions of Train and Test far apart?</p>\n<p>P.S.  <br>\nBecause the plots above are generated by TSNE, the shape of the distribution in 2D-space is meaningless.</p>",
          "rawMarkdown": "Thanks.\n\n> my model make significantly different distribution\n\nBy your model, are the distributions of Train and Test far apart?\n\n\nP.S.  \nBecause the plots above are generated by TSNE, the shape of the distribution in 2D-space is meaningless."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1339990,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2021-06-07T15:21:48.557000",
      "content": "<p>I have made an adversarial test on train/test data and it seems that distributions are nearly completely the same. ROC AUC ~ 51%</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1338538,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-06-06T14:34:35.950000",
      "content": "<p>It feels like there's something different between train and test, to me. Because certain model types give me 0.98 every time I tried, and certain give me 0.97, with almost identical cross-validation performance.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1338565,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-06T14:48:29.190000",
          "content": "<p>One possible reason I think is that the Public Test is biased and either one of those model types is (not) good at the \"biased\" Public Test.</p>\n<p>Did you compare the prediction from different model types?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1477409,
      "author_name": "Ryunosuke Ishizaki",
      "author_url": "",
      "post_date": "2021-08-17T13:17:57.960000",
      "content": "<p>I don't say a lot, but test distribution is far from train<br>\nthere maybe interesting stuff at final LB</p>",
      "votes": -1,
      "replies": [
        {
          "id": 1477659,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2021-08-17T15:15:50.233000",
          "content": "<p>The host says public/private split is complete random, do you think still there will be something intersting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1477846,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-08-17T17:06:24.333000",
          "content": "<p>yeah, let's see</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1481574,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2021-08-19T15:29:14.947000",
          "content": "<p><a href=\"https://www.kaggle.com/ryunosukeishizaki\" target=\"_blank\">@ryunosukeishizaki</a>  After all, what was \"interesting\"? I would like you to find out why I shake down and let me know so that I can make use of it next time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481611,
          "author_name": "Ryunosuke Ishizaki",
          "author_url": "",
          "post_date": "2021-08-19T15:45:27.177000",
          "content": "<p>Hi,<br>\nsince I joined this competition 5days ago, I believed reverse engineering to generate unknown message (based on hosts paper) was the way to easily get higher, private dataset was mentioned to contain more unknown signals, &amp; test distribution was usually much different from train(for weighted average histogram-based feature, train data consists of 2 types of distribution, and at test was domained by mainly 1 of distribution.<br>\nthen what I was thinking was guys who directly optimizing unknown \"needle\" information will be  still robustly at top, &amp; guys who don't mind about it will relatively shakedown,<br>\nbut it was still funny for me to see 1st place denoising solution,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1338545,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-06-06T14:39:05.607000",
      "content": "<p>Isn't it because the model is well trained?<br>\nIf it shows different aspects in the input space of train and test, it can be seen that it has been trained incorrectly…<br>\nand my model make significantly different distribution (LB .94 model)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1338579,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-06T14:55:42.943000",
          "content": "<p>Thanks.</p>\n<blockquote>\n  <p>my model make significantly different distribution</p>\n</blockquote>\n<p>By your model, are the distributions of Train and Test far apart?</p>\n<p>P.S.  <br>\nBecause the plots above are generated by TSNE, the shape of the distribution in 2D-space is meaningless.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1338515": "From the problem setup, we are asked to find messages from the unknown.\nSo isn't it natural that \"unknown\" messages are included in the Private Test Set?\n\nThe first thing I tired is to compare distributions of TrainSet and TestSet:  \nhttps://www.kaggle.com/ttahara/eda-seti-e-t-train-v-s-test-by-cnn-embeddings\n\nI extract images' embeddings by using my baseline then perform dimensionality reduction on them by TSNE.\n\nThe first thing I found is that positive examples(\"needles\") and negative examples in TrainSet are almost separable as the following plots:  \n\n![needles v.s. non-needles](https://pbs.twimg.com/media/E3NGp6WVEAEy7Xe?format=png)\n\nThe second thing is that distribution of image embeddings for TrainSet and TestSet are almost identical as follows:  \n\n![Train v.s. Test](https://pbs.twimg.com/media/E3NGundVEAIEt7W?format=jpg)\n<br>\n\nThere seems no difference between TrainSet and TestSet, at least **for the distribution of image embeddings** extracted by CNN.  \nI think, however, it is possible that there are unknown messages hidden in the TestSet.",
    "1339990": "I have made an adversarial test on train/test data and it seems that distributions are nearly completely the same. ROC AUC ~ 51%\n",
    "1338538": "It feels like there's something different between train and test, to me. Because certain model types give me 0.98 every time I tried, and certain give me 0.97, with almost identical cross-validation performance.\n",
    "1477409": "I don't say a lot, but test distribution is far from train\nthere maybe interesting stuff at final LB",
    "1338545": "Isn't it because the model is well trained?\nIf it shows different aspects in the input space of train and test, it can be seen that it has been trained incorrectly...\nand my model make significantly different distribution (LB .94 model)"
  }
}