{
  "id": 246566,
  "title": "Why I can get oof score 0.61 only with off-target images?",
  "url": "/competitions/seti-breakthrough-listen/discussion/246566",
  "author_name": "tomoo inubushi",
  "post_date": "2021-06-15T23:55:39.117000",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/tomooinubushi/train-model-only-with-off-target-images\" target=\"_blank\">this notebook</a>, I trained the model only with off-taget images (i.e., img[[1,3,5]]) to check these images do not have any information about needles.</p>\n<p>Interestingly, I got the <strong>oof score 0.610164</strong> and <strong>LB score 0.53</strong>. It means <strong>CV score is better than chance level, and shows some over-fitting</strong>. Do you know the reason why? My hypothesis is that our CV strategy (e.g., train-val split) is not appropriate, but I am not sure.</p>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>metric</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.602645</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.613821</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.613558</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.612594</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.612873</td>\n</tr>\n<tr>\n<td>oof</td>\n<td>0.610164</td>\n</tr>\n<tr>\n<td>LB</td>\n<td>0.53</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 1350930,
      "postDate": "2021-06-15T23:55:39.117Z",
      "content": "<p>In <a href=\"https://www.kaggle.com/tomooinubushi/train-model-only-with-off-target-images\" target=\"_blank\">this notebook</a>, I trained the model only with off-taget images (i.e., img[[1,3,5]]) to check these images do not have any information about needles.</p>\n<p>Interestingly, I got the <strong>oof score 0.610164</strong> and <strong>LB score 0.53</strong>. It means <strong>CV score is better than chance level, and shows some over-fitting</strong>. Do you know the reason why? My hypothesis is that our CV strategy (e.g., train-val split) is not appropriate, but I am not sure.</p>\n<table>\n<thead>\n<tr>\n<th>fold</th>\n<th>metric</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>0.602645</td>\n</tr>\n<tr>\n<td>1</td>\n<td>0.613821</td>\n</tr>\n<tr>\n<td>2</td>\n<td>0.613558</td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.612594</td>\n</tr>\n<tr>\n<td>4</td>\n<td>0.612873</td>\n</tr>\n<tr>\n<td>oof</td>\n<td>0.610164</td>\n</tr>\n<tr>\n<td>LB</td>\n<td>0.53</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "In [this notebook](https://www.kaggle.com/tomooinubushi/train-model-only-with-off-target-images), I trained the model only with off-taget images (i.e., img[[1,3,5]]) to check these images do not have any information about needles.\n\nInterestingly, I got the **oof score 0.610164** and **LB score 0.53**. It means **CV score is better than chance level, and shows some over-fitting**. Do you know the reason why? My hypothesis is that our CV strategy (e.g., train-val split) is not appropriate, but I am not sure.\n\n| fold | metric |\n| --- | --- |\n|  0| \t0.602645 |\n|  1 |\t0.613821 |\n|  2| \t0.613558 |\n|  3| \t0.612594 |\n|  4| \t0.612873 |\n|  oof| \t0.610164 |\n|  LB| \t0.53 |",
      "votes": 7
    },
    {
      "id": 1356969,
      "postDate": "2021-06-19T11:32:19.450Z",
      "content": "<p>because mean and variance are different for images with target =1 from images with target 0.</p>",
      "rawMarkdown": "because mean and variance are different for images with target =1 from images with target 0.",
      "votes": 1,
      "replies": [
        {
          "id": 1357040,
          "postDate": "2021-06-19T12:25:43.910Z",
          "content": "<p>please can you help clarify a bit, are you saying that mean and variance of off-target images with 0 label are different from off-target images with 1 label?</p>",
          "rawMarkdown": "please can you help clarify a bit, are you saying that mean and variance of off-target images with 0 label are different from off-target images with 1 label?",
          "votes": 1
        },
        {
          "id": 1357093,
          "postDate": "2021-06-19T13:26:05.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> have you checked <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\" target=\"_blank\">this</a>?<br>\nThe differences of mean and variance for off-target images are not clear compared with those for on-target images, but I could see they have some predictive power.</p>",
          "rawMarkdown": "@ankitsajwan have you checked [this](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901) and [this](https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995)?\nThe differences of mean and variance for off-target images are not clear compared with those for on-target images, but I could see they have some predictive power.",
          "votes": 1
        },
        {
          "id": 1357260,
          "postDate": "2021-06-19T15:11:57.383Z",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> yes, see <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901#1354331\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901#1354331</a></p>",
          "rawMarkdown": "@ankitsajwan yes, see https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901#1354331",
          "votes": 1
        }
      ]
    },
    {
      "id": 1356383,
      "postDate": "2021-06-19T01:46:10.040Z",
      "content": "<p>Deep learning models work with the principle of \"garbage in garbage out\". Your experiment is really interesting however, given the unexplainable nature of the model we cannot be certain if there's a necessary signal in the off-target spectrograms. </p>\n<p>CNNs are feature extractors and given a supervised target, it will find something in the image to map it to the target. There's a paper that actually investigates this - </p>\n<p>The authors randomly shuffled the ground truth labels for the training data, thus an image of a cat might have a label elephant. They trained the model with this mislabeled dataset and to surprise the training accuracy reached, say, 90%. If you train the same model with correct labels, the training accuracy is also close to 90%. However, as expected the model trained with bad labels generalized POORly on the test dataset.</p>\n<p>TL;DR, deep learning models can extract signals in order to map them to a target label, but it necessarily doesn't mean that we have relevant signal in the image. </p>\n<p>I might be wrong but I will conduct the same experiment on my experimental setup. :)</p>",
      "rawMarkdown": "Deep learning models work with the principle of \"garbage in garbage out\". Your experiment is really interesting however, given the unexplainable nature of the model we cannot be certain if there's a necessary signal in the off-target spectrograms. \n\nCNNs are feature extractors and given a supervised target, it will find something in the image to map it to the target. There's a paper that actually investigates this - \n\nThe authors randomly shuffled the ground truth labels for the training data, thus an image of a cat might have a label elephant. They trained the model with this mislabeled dataset and to surprise the training accuracy reached, say, 90%. If you train the same model with correct labels, the training accuracy is also close to 90%. However, as expected the model trained with bad labels generalized POORly on the test dataset.\n\nTL;DR, deep learning models can extract signals in order to map them to a target label, but it necessarily doesn't mean that we have relevant signal in the image. \n\nI might be wrong but I will conduct the same experiment on my experimental setup. :)",
      "votes": 1,
      "replies": [
        {
          "id": 1357077,
          "postDate": "2021-06-19T13:06:03.430Z",
          "content": "<p>Thank you for your comment.<br>\nYes, the input was a garbage, but the output was not just a garbage. It was an <em>interesting</em> garbage. <br>\nI could reach to the <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901\" target=\"_blank\">statistical anomaly of the dataset</a> on my way to digging into this strange result. I did not know why this statistical anomaly happened, and could not notice it was caused by a leakage. I am feeling sorry for it.</p>\n<p>One thing I am still wondering is why the LB score was not as good as 0.61, because this statistical leakage gives me higher LB score. I will find the answer when the new dataset is released.</p>",
          "rawMarkdown": "Thank you for your comment.\nYes, the input was a garbage, but the output was not just a garbage. It was an *interesting* garbage. \nI could reach to the [statistical anomaly of the dataset](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901) on my way to digging into this strange result. I did not know why this statistical anomaly happened, and could not notice it was caused by a leakage. I am feeling sorry for it.\n\nOne thing I am still wondering is why the LB score was not as good as 0.61, because this statistical leakage gives me higher LB score. I will find the answer when the new dataset is released.",
          "votes": 1
        },
        {
          "id": 1357095,
          "postDate": "2021-06-19T13:31:29.760Z",
          "content": "<p>I agree it's an interesting garbage.</p>",
          "rawMarkdown": "I agree it's an interesting garbage."
        }
      ]
    },
    {
      "id": 1351287,
      "postDate": "2021-06-16T07:29:08.130Z",
      "content": "<p>(just personal idea)</p>\n<ol>\n<li>Since the number of detected artifact's signal will be limited, information about nearby signals(=off) is also limited to similar areas.</li>\n<li>test dataset and train dataset give signal from different object &amp; each dataset contain same object / different time captures</li>\n<li>but, human launches artifacts to somewhere interest. and there are common factor (p &lt; 0.05) of RF signal from interest places (eg planet?, 90 degree from earth orbit?)</li>\n</ol>",
      "rawMarkdown": "(just personal idea)\n1. Since the number of detected artifact's signal will be limited, information about nearby signals(=off) is also limited to similar areas.\n2. test dataset and train dataset give signal from different object & each dataset contain same object / different time captures\n3. but, human launches artifacts to somewhere interest. and there are common factor (p < 0.05) of RF signal from interest places (eg planet?, 90 degree from earth orbit?)",
      "votes": 2,
      "replies": [
        {
          "id": 1351461,
          "postDate": "2021-06-16T10:32:34.453Z",
          "content": "<p>Thank you for your reply. <br>\nIt is reasonable that aliens should live in a specific area of the universe.<br>\nI believe this strange property is related to how they 'measure' the data.</p>",
          "rawMarkdown": "Thank you for your reply. \nIt is reasonable that aliens should live in a specific area of the universe.\nI believe this strange property is related to how they 'measure' the data."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1356969,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-19T11:32:19.450000",
      "content": "<p>because mean and variance are different for images with target =1 from images with target 0.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1357040,
          "author_name": "sajwankit",
          "author_url": "",
          "post_date": "2021-06-19T12:25:43.910000",
          "content": "<p>please can you help clarify a bit, are you saying that mean and variance of off-target images with 0 label are different from off-target images with 1 label?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1357093,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-19T13:26:05.403000",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> have you checked <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\" target=\"_blank\">this</a>?<br>\nThe differences of mean and variance for off-target images are not clear compared with those for on-target images, but I could see they have some predictive power.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1357260,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-19T15:11:57.383000",
          "content": "<p><a href=\"https://www.kaggle.com/ankitsajwan\" target=\"_blank\">@ankitsajwan</a> yes, see <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901#1354331\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901#1354331</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1356383,
      "author_name": "Ayush Thakur",
      "author_url": "",
      "post_date": "2021-06-19T01:46:10.040000",
      "content": "<p>Deep learning models work with the principle of \"garbage in garbage out\". Your experiment is really interesting however, given the unexplainable nature of the model we cannot be certain if there's a necessary signal in the off-target spectrograms. </p>\n<p>CNNs are feature extractors and given a supervised target, it will find something in the image to map it to the target. There's a paper that actually investigates this - </p>\n<p>The authors randomly shuffled the ground truth labels for the training data, thus an image of a cat might have a label elephant. They trained the model with this mislabeled dataset and to surprise the training accuracy reached, say, 90%. If you train the same model with correct labels, the training accuracy is also close to 90%. However, as expected the model trained with bad labels generalized POORly on the test dataset.</p>\n<p>TL;DR, deep learning models can extract signals in order to map them to a target label, but it necessarily doesn't mean that we have relevant signal in the image. </p>\n<p>I might be wrong but I will conduct the same experiment on my experimental setup. :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1357077,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-19T13:06:03.430000",
          "content": "<p>Thank you for your comment.<br>\nYes, the input was a garbage, but the output was not just a garbage. It was an <em>interesting</em> garbage. <br>\nI could reach to the <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246901\" target=\"_blank\">statistical anomaly of the dataset</a> on my way to digging into this strange result. I did not know why this statistical anomaly happened, and could not notice it was caused by a leakage. I am feeling sorry for it.</p>\n<p>One thing I am still wondering is why the LB score was not as good as 0.61, because this statistical leakage gives me higher LB score. I will find the answer when the new dataset is released.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1357095,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-19T13:31:29.760000",
          "content": "<p>I agree it's an interesting garbage.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1351287,
      "author_name": "assign",
      "author_url": "",
      "post_date": "2021-06-16T07:29:08.130000",
      "content": "<p>(just personal idea)</p>\n<ol>\n<li>Since the number of detected artifact's signal will be limited, information about nearby signals(=off) is also limited to similar areas.</li>\n<li>test dataset and train dataset give signal from different object &amp; each dataset contain same object / different time captures</li>\n<li>but, human launches artifacts to somewhere interest. and there are common factor (p &lt; 0.05) of RF signal from interest places (eg planet?, 90 degree from earth orbit?)</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 1351461,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-16T10:32:34.453000",
          "content": "<p>Thank you for your reply. <br>\nIt is reasonable that aliens should live in a specific area of the universe.<br>\nI believe this strange property is related to how they 'measure' the data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1350930": "In [this notebook](https://www.kaggle.com/tomooinubushi/train-model-only-with-off-target-images), I trained the model only with off-taget images (i.e., img[[1,3,5]]) to check these images do not have any information about needles.\n\nInterestingly, I got the **oof score 0.610164** and **LB score 0.53**. It means **CV score is better than chance level, and shows some over-fitting**. Do you know the reason why? My hypothesis is that our CV strategy (e.g., train-val split) is not appropriate, but I am not sure.\n\n| fold | metric |\n| --- | --- |\n|  0| \t0.602645 |\n|  1 |\t0.613821 |\n|  2| \t0.613558 |\n|  3| \t0.612594 |\n|  4| \t0.612873 |\n|  oof| \t0.610164 |\n|  LB| \t0.53 |",
    "1356969": "because mean and variance are different for images with target =1 from images with target 0.",
    "1356383": "Deep learning models work with the principle of \"garbage in garbage out\". Your experiment is really interesting however, given the unexplainable nature of the model we cannot be certain if there's a necessary signal in the off-target spectrograms. \n\nCNNs are feature extractors and given a supervised target, it will find something in the image to map it to the target. There's a paper that actually investigates this - \n\nThe authors randomly shuffled the ground truth labels for the training data, thus an image of a cat might have a label elephant. They trained the model with this mislabeled dataset and to surprise the training accuracy reached, say, 90%. If you train the same model with correct labels, the training accuracy is also close to 90%. However, as expected the model trained with bad labels generalized POORly on the test dataset.\n\nTL;DR, deep learning models can extract signals in order to map them to a target label, but it necessarily doesn't mean that we have relevant signal in the image. \n\nI might be wrong but I will conduct the same experiment on my experimental setup. :)",
    "1351287": "(just personal idea)\n1. Since the number of detected artifact's signal will be limited, information about nearby signals(=off) is also limited to similar areas.\n2. test dataset and train dataset give signal from different object & each dataset contain same object / different time captures\n3. but, human launches artifacts to somewhere interest. and there are common factor (p < 0.05) of RF signal from interest places (eg planet?, 90 degree from earth orbit?)"
  }
}