{
  "id": 246901,
  "title": "Yet another leakage? (LB 0.995 without timestamp)",
  "url": "/competitions/seti-breakthrough-listen/discussion/246901",
  "author_name": "nyanp",
  "post_date": "2021-06-17T09:56:26.561000",
  "votes": 63,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n<p>I found a way to get a score of 0.995 by blending a simple statistic with the existing public kernel, without using the leak that <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> found.</p>\n<p><a href=\"https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\" target=\"_blank\">https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995</a></p>\n<p>This leak probably indicates a problem with the data generation process in the simulation.</p>\n<p>Some of the models of the top participants may already implicitly take advantage of this distributional difference, but I hope that this leak will be fixed along with the timestamp leak.</p>",
  "messages": [
    {
      "id": 1353955,
      "postDate": "2021-06-17T09:56:26.560Z",
      "content": "<p>Hi Kagglers,</p>\n<p>I found a way to get a score of 0.995 by blending a simple statistic with the existing public kernel, without using the leak that <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> found.</p>\n<p><a href=\"https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\" target=\"_blank\">https://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995</a></p>\n<p>This leak probably indicates a problem with the data generation process in the simulation.</p>\n<p>Some of the models of the top participants may already implicitly take advantage of this distributional difference, but I hope that this leak will be fixed along with the timestamp leak.</p>",
      "rawMarkdown": "Hi Kagglers,\n\nI found a way to get a score of 0.995 by blending a simple statistic with the existing public kernel, without using the leak that @kazanova found.\n\nhttps://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\n\nThis leak probably indicates a problem with the data generation process in the simulation.\n\nSome of the models of the top participants may already implicitly take advantage of this distributional difference, but I hope that this leak will be fixed along with the timestamp leak.",
      "votes": 63
    },
    {
      "id": 1354157,
      "postDate": "2021-06-17T12:11:27.977Z",
      "content": "<p>I figured out - <strong>problem in downcast truncating.</strong></p>\n<p><code>image = np.load(\"../input/train/0/001c619bdf53.npy\").astype(np.float64)</code><br>\n<code>image.std(axis=(1,2))</code><br>\n<code>Out: array([1.00001716, 0.99999972, 1.00001509, 1.0000001 , 1.00000851,\n       0.99999885])</code></p>\n<p><code>image_scaled = (image.transpose(1, 2, 0) - image.mean(axis=(1, 2))) / image.std(axis=(1, 2))</code><br>\n<code>image_scaled = image_scaled.transpose(2, 0, 1)</code><br>\n<code>image_scaled.std(axis=(1, 2))</code><br>\n<code>Out: array([1., 1., 1., 1., 1., 1.])</code></p>\n<p><code>image_scaled.astype(np.float16).astype(np.float64).std(axis=(1, 2))</code><br>\n<code>Out: array([1.00001714, 0.99999972, 1.00001499, 0.9999999 , 1.00000858,\n       1.00000199])</code></p>\n<p>It is just example of one statistic. But all on-off distributions changed and helps to find targeted samples.   </p>",
      "rawMarkdown": "I figured out - **problem in downcast truncating.**\n\n```image = np.load(\"../input/train/0/001c619bdf53.npy\").astype(np.float64)```\n```image.std(axis=(1,2))```\n```Out: array([1.00001716, 0.99999972, 1.00001509, 1.0000001 , 1.00000851,\n       0.99999885])```\n\n```image_scaled = (image.transpose(1, 2, 0) - image.mean(axis=(1, 2))) / image.std(axis=(1, 2))```\n```image_scaled = image_scaled.transpose(2, 0, 1)```\n```image_scaled.std(axis=(1, 2))```\n```Out: array([1., 1., 1., 1., 1., 1.])```\n\n```image_scaled.astype(np.float16).astype(np.float64).std(axis=(1, 2))```\n```Out: array([1.00001714, 0.99999972, 1.00001499, 0.9999999 , 1.00000858,\n       1.00000199])```\n\nIt is just example of one statistic. But all on-off distributions changed and helps to find targeted samples.   ",
      "votes": 25,
      "replies": [
        {
          "id": 1354172,
          "postDate": "2021-06-17T12:21:59.710Z",
          "content": "<p>This explains why normalization was degrading a bit my CV and LB…</p>",
          "rawMarkdown": "This explains why normalization was degrading a bit my CV and LB...",
          "votes": 8
        },
        {
          "id": 1354177,
          "postDate": "2021-06-17T12:28:12.117Z",
          "content": "<p>Yep. But there is another challenge - how to prepare data leak free. Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..</p>",
          "rawMarkdown": "Yep. But there is another challenge - how to prepare data leak free. Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..",
          "votes": 3
        },
        {
          "id": 1354198,
          "postDate": "2021-06-17T12:49:39.123Z",
          "content": "<p>Ah, this is a reasonable explanation. They might have to quantize the target appropriately or add random floa64 noise to background.</p>",
          "rawMarkdown": "Ah, this is a reasonable explanation. They might have to quantize the target appropriately or add random floa64 noise to background.",
          "votes": 4
        },
        {
          "id": 1354214,
          "postDate": "2021-06-17T13:05:12.490Z",
          "content": "<p>does it means label is related to sub digits?<br>\nor can I ask more explains?</p>",
          "rawMarkdown": "does it means label is related to sub digits?\nor can I ask more explains?",
          "votes": 1
        },
        {
          "id": 1354268,
          "postDate": "2021-06-17T13:26:18.773Z",
          "content": "<p>Finally an explanation. Thanks. </p>",
          "rawMarkdown": "Finally an explanation. Thanks. "
        },
        {
          "id": 1354275,
          "postDate": "2021-06-17T13:34:17.367Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1354322,
          "postDate": "2021-06-17T14:31:41.427Z",
          "content": "<blockquote>\n  <p>Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..</p>\n</blockquote>\n<p>If all data is normalized the same way then why would this leak info?</p>",
          "rawMarkdown": "> Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..\n\nIf all data is normalized the same way then why would this leak info?",
          "votes": 1
        },
        {
          "id": 1354899,
          "postDate": "2021-06-18T01:19:34.340Z",
          "content": "<p>What does this mean? How can we pick out the target image by seeing the channel std?</p>",
          "rawMarkdown": "What does this mean? How can we pick out the target image by seeing the channel std?"
        },
        {
          "id": 1362712,
          "postDate": "2021-06-23T15:23:35.483Z",
          "content": "<p>I'll try to summarize the key points of this discussion below, please let me know if I am wrong at any place.</p>\n<ul>\n<li>Ideally all the images in this competition should have mean of around 0 and standard deviation of 1.</li>\n<li>It has been found that images with target 0 have lesser variance when compared to target 1 images.</li>\n<li>There is a positive bias in the mean of target 1 images when compared to target 0 images.</li>\n<li>The signal is added to the images(constituting target 1) artificially. So there should be good care taken  to make sure that there is no difference in the distributions images belonging to the two classes. Otherwise any simple machine learning model can be a strong classifier.</li>\n<li>However the data description of the competition said  'After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array'. So we have a contradictory scenario here.</li>\n<li><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> provided the rationale for increase in the variance, normalization of image in its float64 form is totally fine, all the 6 images have a standard deviation of 1.0 but when downcasted to float16 the standard deviation does not remain exactly 1.0 anymore, they are slightly deviated around 1.0. Hence the increase in variance.</li>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> has probably identified the fault, some of the images were normalized in their float16 form while others were normalized in float64 form and then downcasted to float16, this introduced the difference in the distributions of the two classes.</li>\n</ul>",
          "rawMarkdown": "I'll try to summarize the key points of this discussion below, please let me know if I am wrong at any place.\n\n- Ideally all the images in this competition should have mean of around 0 and standard deviation of 1.\n- It has been found that images with target 0 have lesser variance when compared to target 1 images.\n- There is a positive bias in the mean of target 1 images when compared to target 0 images.\n- The signal is added to the images(constituting target 1) artificially. So there should be good care taken  to make sure that there is no difference in the distributions images belonging to the two classes. Otherwise any simple machine learning model can be a strong classifier.\n- However the data description of the competition said  'After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array'. So we have a contradictory scenario here.\n- @sggpls provided the rationale for increase in the variance, normalization of image in its float64 form is totally fine, all the 6 images have a standard deviation of 1.0 but when downcasted to float16 the standard deviation does not remain exactly 1.0 anymore, they are slightly deviated around 1.0. Hence the increase in variance.\n- @cpmpml has probably identified the fault, some of the images were normalized in their float16 form while others were normalized in float64 form and then downcasted to float16, this introduced the difference in the distributions of the two classes.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1354024,
      "postDate": "2021-06-17T10:45:20.960Z",
      "content": "<p>I am sorry for this. This competition was my first participation in Kaggle.<br>\nOn yesterday, while exploring the training data features, I found the relationship between timestamps and targets.<br>\nI trained on training data by grouping the data by timestamp, and it was successful.</p>\n<p>I wasn't sure about the relationship until I submit the 0.995 answer.<br>\nI sincerely apologize for not revealing the timestamp effect on the target.<br>\nGreed ruined me.</p>\n<p>I apologize for not revealing the data timestamp leakage.</p>",
      "rawMarkdown": "I am sorry for this. This competition was my first participation in Kaggle.\nOn yesterday, while exploring the training data features, I found the relationship between timestamps and targets.\nI trained on training data by grouping the data by timestamp, and it was successful.\n\nI wasn't sure about the relationship until I submit the 0.995 answer.\nI sincerely apologize for not revealing the timestamp effect on the target.\nGreed ruined me.\n\nI apologize for not revealing the data timestamp leakage.",
      "votes": 23,
      "replies": [
        {
          "id": 1354037,
          "postDate": "2021-06-17T10:59:34.220Z",
          "content": "<p>Thanks for being clear.  :)<br>\nI couldn't sleep due to making idea about how your score would rise dramatically.</p>",
          "rawMarkdown": "Thanks for being clear.  :)\nI couldn't sleep due to making idea about how your score would rise dramatically.",
          "votes": 1
        },
        {
          "id": 1354041,
          "postDate": "2021-06-17T11:05:50.553Z",
          "content": "<p>I am really sorry.<br>\nI hope you get good results in this competition. </p>",
          "rawMarkdown": "I am really sorry.\nI hope you get good results in this competition. ",
          "votes": 6
        }
      ]
    },
    {
      "id": 1354211,
      "postDate": "2021-06-17T13:01:14.493Z",
      "content": "<p>If leakage happens not only in timestamp but also in the image level, then organizers should replace both train and test datasets. Replacing or fixing only testset will create a bias between the train and the test images. </p>",
      "rawMarkdown": "If leakage happens not only in timestamp but also in the image level, then organizers should replace both train and test datasets. Replacing or fixing only testset will create a bias between the train and the test images. ",
      "votes": 19,
      "replies": [
        {
          "id": 1354285,
          "postDate": "2021-06-17T13:47:34.247Z",
          "content": "<p>Think organisers did state that both are being replaced. Recall seeing a post by a staff member/organiser stating that they may be adding the current test dataset to the future train dataset.</p>\n<p>EDIT : You can find it in the comments section here - <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782</a></p>",
          "rawMarkdown": "Think organisers did state that both are being replaced. Recall seeing a post by a staff member/organiser stating that they may be adding the current test dataset to the future train dataset.\n\nEDIT : You can find it in the comments section here - https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782",
          "votes": 2
        }
      ]
    },
    {
      "id": 1354014,
      "postDate": "2021-06-17T10:34:49.877Z",
      "content": "<p>Thank you for reporting. I never expected to see GBDT in CV competition except for stacking 😅</p>\n<p>If your hypothesis is correct, train dataset also needs updating?</p>",
      "rawMarkdown": "Thank you for reporting. I never expected to see GBDT in CV competition except for stacking 😅\n\nIf your hypothesis is correct, train dataset also needs updating?",
      "votes": 9,
      "replies": [
        {
          "id": 1354029,
          "postDate": "2021-06-17T10:52:01.713Z",
          "content": "<blockquote>\n  <p>If your hypothesis is correct, train dataset also needs updating?</p>\n</blockquote>\n<p>I think so, but it may be enough to simply re-normalize the current training data.</p>",
          "rawMarkdown": "> If your hypothesis is correct, train dataset also needs updating?\n\nI think so, but it may be enough to simply re-normalize the current training data.\n",
          "votes": 9
        },
        {
          "id": 1354165,
          "postDate": "2021-06-17T12:14:45.820Z",
          "content": "<p>Indeed, if they are going to release a new test test properly normalized then better normalize train as well.</p>",
          "rawMarkdown": "Indeed, if they are going to release a new test test properly normalized then better normalize train as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1354013,
      "postDate": "2021-06-17T10:33:52.253Z",
      "content": "<p>I found the same and was about to share.  Thanks for sharing!</p>",
      "rawMarkdown": "I found the same and was about to share.  Thanks for sharing!",
      "votes": 8,
      "replies": [
        {
          "id": 1354331,
          "postDate": "2021-06-17T14:36:15.667Z",
          "content": "<p>To be clear, I had found the difference in distribution, but I didn't train any model on it.  It would have been next step if the leak hadn't been disclosed here first.</p>\n<p>Example of visualizations I did, red are target=1.</p>\n<p><img src=\"https://i.imgur.com/yxhzUm7.png\" alt=\"distribution\"></p>",
          "rawMarkdown": "To be clear, I had found the difference in distribution, but I didn't train any model on it.  It would have been next step if the leak hadn't been disclosed here first.\n\nExample of visualizations I did, red are target=1.\n\n![distribution](https://i.imgur.com/yxhzUm7.png)",
          "votes": 4
        },
        {
          "id": 1354389,
          "postDate": "2021-06-17T14:59:03.610Z",
          "content": "<p>What are the axes on the plot you showed above?</p>",
          "rawMarkdown": "What are the axes on the plot you showed above?",
          "votes": 5
        },
        {
          "id": 1354428,
          "postDate": "2021-06-17T15:36:00.563Z",
          "content": "<p>I know I should not have shared it…</p>\n<p>It is the sum of the means of on snippets, and the sum of the means of off snippets.</p>\n<p>We see that variance difference comes from the snippets containing messages.</p>\n<p>If your model captures this then no wonder it does well.</p>",
          "rawMarkdown": "I know I should not have shared it...\n\nIt is the sum of the means of on snippets, and the sum of the means of off snippets.\n\nWe see that variance difference comes from the snippets containing messages.\n\nIf your model captures this then no wonder it does well.",
          "votes": 3
        },
        {
          "id": 1354437,
          "postDate": "2021-06-17T15:42:58.630Z",
          "content": "<p>My guess is that the messages were added using FP16, then normalized using FP16 while the rest of the data had been normalized with FP32 or FP64 before downcasting.</p>",
          "rawMarkdown": "My guess is that the messages were added using FP16, then normalized using FP16 while the rest of the data had been normalized with FP32 or FP64 before downcasting.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1354121,
      "postDate": "2021-06-17T11:46:52.667Z",
      "content": "<blockquote>\n  <p>Perhaps the host added the target signal to the image after normalizing the background noise. (Of course, to prevent leakage, the image should be normalized after the target signal is added.</p>\n</blockquote>\n<p>I was theorizing the same, but in the data description they say:</p>\n<blockquote>\n  <p>After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.</p>\n</blockquote>\n<p>So I didn't even check… </p>",
      "rawMarkdown": "> Perhaps the host added the target signal to the image after normalizing the background noise. (Of course, to prevent leakage, the image should be normalized after the target signal is added.\n\nI was theorizing the same, but in the data description they say:\n\n> After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.\n\nSo I didn't even check... ",
      "votes": 7,
      "replies": [
        {
          "id": 1354187,
          "postDate": "2021-06-17T12:39:46.620Z",
          "content": "<p>Perhaps they injected a signal into the image and then normalized it with the mean and variance of the background noise 🤔</p>",
          "rawMarkdown": "Perhaps they injected a signal into the image and then normalized it with the mean and variance of the background noise 🤔",
          "votes": 1
        }
      ]
    },
    {
      "id": 1354090,
      "postDate": "2021-06-17T11:29:55.753Z",
      "content": "<p>Yep, it seems my custom nn <em>implicitly</em> used information about on-off channel distribution. Interesting competition :sarcasm: At least it cool to learn different way to find and exploit the leaks. Who knows what else the aliens have prepared for us?</p>",
      "rawMarkdown": "Yep, it seems my custom nn *implicitly* used information about on-off channel distribution. Interesting competition :sarcasm: At least it cool to learn different way to find and exploit the leaks. Who knows what else the aliens have prepared for us?\n ",
      "votes": 5
    },
    {
      "id": 1354009,
      "postDate": "2021-06-17T10:26:05.070Z",
      "content": "<p>Although it is clear this finding is real, I think the organizers wanted to make sure this did not happen, according to this statement in the Data Information tab:</p>\n<p><code>After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.</code></p>",
      "rawMarkdown": "Although it is clear this finding is real, I think the organizers wanted to make sure this did not happen, according to this statement in the Data Information tab:\n\n`After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array. `",
      "votes": 3,
      "replies": [
        {
          "id": 1354026,
          "postDate": "2021-06-17T10:47:26.293Z",
          "content": "<blockquote>\n  <p>I think the organizers wanted to make sure this did not happen</p>\n</blockquote>\n<p>Definitely. There must be something wrong with their simulation code.</p>",
          "rawMarkdown": "> I think the organizers wanted to make sure this did not happen\n\nDefinitely. There must be something wrong with their simulation code.",
          "votes": 2
        },
        {
          "id": 1354128,
          "postDate": "2021-06-17T11:51:29.237Z",
          "content": "<p>Right, I also saw this and didn't bother to check… never trust anything I guess :)</p>",
          "rawMarkdown": "Right, I also saw this and didn't bother to check... never trust anything I guess :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 1354007,
      "postDate": "2021-06-17T10:25:42.347Z",
      "content": "<p>I noticed this point recently, but I am not sure I should call it a leakage.<br>\nI think it is OK if \"real\" alien signals have such statistical properties.</p>",
      "rawMarkdown": "I noticed this point recently, but I am not sure I should call it a leakage.\nI think it is OK if \"real\" alien signals have such statistical properties.",
      "votes": 1,
      "replies": [
        {
          "id": 1354019,
          "postDate": "2021-06-17T10:40:49.540Z",
          "content": "<p>I agree that it is not clear if this is a leak or not. However, to prevent another competition reset, I thought this information should be made public.</p>",
          "rawMarkdown": "I agree that it is not clear if this is a leak or not. However, to prevent another competition reset, I thought this information should be made public.",
          "votes": 9
        },
        {
          "id": 1354035,
          "postDate": "2021-06-17T10:58:56.143Z",
          "content": "<p>I agree with you. One competition reset is enough.</p>",
          "rawMarkdown": "I agree with you. One competition reset is enough."
        }
      ]
    },
    {
      "id": 1354946,
      "postDate": "2021-06-18T03:00:35.900Z",
      "content": "<p>Very interesting to see so many old timers flocking to this competition. It definitely looks like this is the most interesting one of those currently being ran. For now though, time to check out other comps until we get an updated train/test…</p>",
      "rawMarkdown": "Very interesting to see so many old timers flocking to this competition. It definitely looks like this is the most interesting one of those currently being ran. For now though, time to check out other comps until we get an updated train/test..."
    }
  ],
  "comments": [
    {
      "id": 1354157,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-06-17T12:11:27.977000",
      "content": "<p>I figured out - <strong>problem in downcast truncating.</strong></p>\n<p><code>image = np.load(\"../input/train/0/001c619bdf53.npy\").astype(np.float64)</code><br>\n<code>image.std(axis=(1,2))</code><br>\n<code>Out: array([1.00001716, 0.99999972, 1.00001509, 1.0000001 , 1.00000851,\n       0.99999885])</code></p>\n<p><code>image_scaled = (image.transpose(1, 2, 0) - image.mean(axis=(1, 2))) / image.std(axis=(1, 2))</code><br>\n<code>image_scaled = image_scaled.transpose(2, 0, 1)</code><br>\n<code>image_scaled.std(axis=(1, 2))</code><br>\n<code>Out: array([1., 1., 1., 1., 1., 1.])</code></p>\n<p><code>image_scaled.astype(np.float16).astype(np.float64).std(axis=(1, 2))</code><br>\n<code>Out: array([1.00001714, 0.99999972, 1.00001499, 0.9999999 , 1.00000858,\n       1.00000199])</code></p>\n<p>It is just example of one statistic. But all on-off distributions changed and helps to find targeted samples.   </p>",
      "votes": 25,
      "replies": [
        {
          "id": 1354172,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T12:21:59.710000",
          "content": "<p>This explains why normalization was degrading a bit my CV and LB…</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1354177,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-06-17T12:28:12.117000",
          "content": "<p>Yep. But there is another challenge - how to prepare data leak free. Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1354198,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-06-17T12:49:39.123000",
          "content": "<p>Ah, this is a reasonable explanation. They might have to quantize the target appropriately or add random floa64 noise to background.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1354214,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-17T13:05:12.490000",
          "content": "<p>does it means label is related to sub digits?<br>\nor can I ask more explains?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1354268,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-06-17T13:26:18.773000",
          "content": "<p>Finally an explanation. Thanks. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1354275,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-06-17T13:34:17.367000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1354322,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T14:31:41.427000",
          "content": "<blockquote>\n  <p>Even if we have totally normalized data in float64 there is a door to truncate data by downcasting..</p>\n</blockquote>\n<p>If all data is normalized the same way then why would this leak info?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1354899,
          "author_name": "Coin",
          "author_url": "",
          "post_date": "2021-06-18T01:19:34.340000",
          "content": "<p>What does this mean? How can we pick out the target image by seeing the channel std?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1362712,
          "author_name": "Inumellonium",
          "author_url": "",
          "post_date": "2021-06-23T15:23:35.483000",
          "content": "<p>I'll try to summarize the key points of this discussion below, please let me know if I am wrong at any place.</p>\n<ul>\n<li>Ideally all the images in this competition should have mean of around 0 and standard deviation of 1.</li>\n<li>It has been found that images with target 0 have lesser variance when compared to target 1 images.</li>\n<li>There is a positive bias in the mean of target 1 images when compared to target 0 images.</li>\n<li>The signal is added to the images(constituting target 1) artificially. So there should be good care taken  to make sure that there is no difference in the distributions images belonging to the two classes. Otherwise any simple machine learning model can be a strong classifier.</li>\n<li>However the data description of the competition said  'After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array'. So we have a contradictory scenario here.</li>\n<li><a href=\"https://www.kaggle.com/sggpls\" target=\"_blank\">@sggpls</a> provided the rationale for increase in the variance, normalization of image in its float64 form is totally fine, all the 6 images have a standard deviation of 1.0 but when downcasted to float16 the standard deviation does not remain exactly 1.0 anymore, they are slightly deviated around 1.0. Hence the increase in variance.</li>\n<li><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> has probably identified the fault, some of the images were normalized in their float16 form while others were normalized in float64 form and then downcasted to float16, this introduced the difference in the distributions of the two classes.</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1354024,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-17T10:45:20.960000",
      "content": "<p>I am sorry for this. This competition was my first participation in Kaggle.<br>\nOn yesterday, while exploring the training data features, I found the relationship between timestamps and targets.<br>\nI trained on training data by grouping the data by timestamp, and it was successful.</p>\n<p>I wasn't sure about the relationship until I submit the 0.995 answer.<br>\nI sincerely apologize for not revealing the timestamp effect on the target.<br>\nGreed ruined me.</p>\n<p>I apologize for not revealing the data timestamp leakage.</p>",
      "votes": 23,
      "replies": [
        {
          "id": 1354037,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-17T10:59:34.220000",
          "content": "<p>Thanks for being clear.  :)<br>\nI couldn't sleep due to making idea about how your score would rise dramatically.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1354041,
          "author_name": "WOOSUNG YOON",
          "author_url": "",
          "post_date": "2021-06-17T11:05:50.553000",
          "content": "<p>I am really sorry.<br>\nI hope you get good results in this competition. </p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1354211,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2021-06-17T13:01:14.493000",
      "content": "<p>If leakage happens not only in timestamp but also in the image level, then organizers should replace both train and test datasets. Replacing or fixing only testset will create a bias between the train and the test images. </p>",
      "votes": 19,
      "replies": [
        {
          "id": 1354285,
          "author_name": "maze508",
          "author_url": "",
          "post_date": "2021-06-17T13:47:34.247000",
          "content": "<p>Think organisers did state that both are being replaced. Recall seeing a post by a staff member/organiser stating that they may be adding the current test dataset to the future train dataset.</p>\n<p>EDIT : You can find it in the comments section here - <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/246782</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1354014,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-06-17T10:34:49.877000",
      "content": "<p>Thank you for reporting. I never expected to see GBDT in CV competition except for stacking 😅</p>\n<p>If your hypothesis is correct, train dataset also needs updating?</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1354029,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-06-17T10:52:01.713000",
          "content": "<blockquote>\n  <p>If your hypothesis is correct, train dataset also needs updating?</p>\n</blockquote>\n<p>I think so, but it may be enough to simply re-normalize the current training data.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1354165,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T12:14:45.820000",
          "content": "<p>Indeed, if they are going to release a new test test properly normalized then better normalize train as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1354013,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-17T10:33:52.253000",
      "content": "<p>I found the same and was about to share.  Thanks for sharing!</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1354331,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T14:36:15.667000",
          "content": "<p>To be clear, I had found the difference in distribution, but I didn't train any model on it.  It would have been next step if the leak hadn't been disclosed here first.</p>\n<p>Example of visualizations I did, red are target=1.</p>\n<p><img src=\"https://i.imgur.com/yxhzUm7.png\" alt=\"distribution\"></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1354389,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-06-17T14:59:03.610000",
          "content": "<p>What are the axes on the plot you showed above?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1354428,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T15:36:00.563000",
          "content": "<p>I know I should not have shared it…</p>\n<p>It is the sum of the means of on snippets, and the sum of the means of off snippets.</p>\n<p>We see that variance difference comes from the snippets containing messages.</p>\n<p>If your model captures this then no wonder it does well.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1354437,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-17T15:42:58.630000",
          "content": "<p>My guess is that the messages were added using FP16, then normalized using FP16 while the rest of the data had been normalized with FP32 or FP64 before downcasting.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1354121,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-06-17T11:46:52.667000",
      "content": "<blockquote>\n  <p>Perhaps the host added the target signal to the image after normalizing the background noise. (Of course, to prevent leakage, the image should be normalized after the target signal is added.</p>\n</blockquote>\n<p>I was theorizing the same, but in the data description they say:</p>\n<blockquote>\n  <p>After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.</p>\n</blockquote>\n<p>So I didn't even check… </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1354187,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-06-17T12:39:46.620000",
          "content": "<p>Perhaps they injected a signal into the image and then normalized it with the mean and variance of the background noise 🤔</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1354090,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-06-17T11:29:55.753000",
      "content": "<p>Yep, it seems my custom nn <em>implicitly</em> used information about on-off channel distribution. Interesting competition :sarcasm: At least it cool to learn different way to find and exploit the leaks. Who knows what else the aliens have prepared for us?</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1354009,
      "author_name": "FelipeKitamura, MD, PhD",
      "author_url": "",
      "post_date": "2021-06-17T10:26:05.070000",
      "content": "<p>Although it is clear this finding is real, I think the organizers wanted to make sure this did not happen, according to this statement in the Data Information tab:</p>\n<p><code>After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.</code></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1354026,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-06-17T10:47:26.293000",
          "content": "<blockquote>\n  <p>I think the organizers wanted to make sure this did not happen</p>\n</blockquote>\n<p>Definitely. There must be something wrong with their simulation code.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1354128,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-06-17T11:51:29.237000",
          "content": "<p>Right, I also saw this and didn't bother to check… never trust anything I guess :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1354007,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-06-17T10:25:42.347000",
      "content": "<p>I noticed this point recently, but I am not sure I should call it a leakage.<br>\nI think it is OK if \"real\" alien signals have such statistical properties.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1354019,
          "author_name": "nyanp",
          "author_url": "",
          "post_date": "2021-06-17T10:40:49.540000",
          "content": "<p>I agree that it is not clear if this is a leak or not. However, to prevent another competition reset, I thought this information should be made public.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1354035,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-17T10:58:56.143000",
          "content": "<p>I agree with you. One competition reset is enough.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1354946,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-06-18T03:00:35.900000",
      "content": "<p>Very interesting to see so many old timers flocking to this competition. It definitely looks like this is the most interesting one of those currently being ran. For now though, time to check out other comps until we get an updated train/test…</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1353955": "Hi Kagglers,\n\nI found a way to get a score of 0.995 by blending a simple statistic with the existing public kernel, without using the leak that @kazanova found.\n\nhttps://www.kaggle.com/nyanpn/yet-another-leakage-lb0-995\n\nThis leak probably indicates a problem with the data generation process in the simulation.\n\nSome of the models of the top participants may already implicitly take advantage of this distributional difference, but I hope that this leak will be fixed along with the timestamp leak.",
    "1354157": "I figured out - **problem in downcast truncating.**\n\n```image = np.load(\"../input/train/0/001c619bdf53.npy\").astype(np.float64)```\n```image.std(axis=(1,2))```\n```Out: array([1.00001716, 0.99999972, 1.00001509, 1.0000001 , 1.00000851,\n       0.99999885])```\n\n```image_scaled = (image.transpose(1, 2, 0) - image.mean(axis=(1, 2))) / image.std(axis=(1, 2))```\n```image_scaled = image_scaled.transpose(2, 0, 1)```\n```image_scaled.std(axis=(1, 2))```\n```Out: array([1., 1., 1., 1., 1., 1.])```\n\n```image_scaled.astype(np.float16).astype(np.float64).std(axis=(1, 2))```\n```Out: array([1.00001714, 0.99999972, 1.00001499, 0.9999999 , 1.00000858,\n       1.00000199])```\n\nIt is just example of one statistic. But all on-off distributions changed and helps to find targeted samples.   ",
    "1354024": "I am sorry for this. This competition was my first participation in Kaggle.\nOn yesterday, while exploring the training data features, I found the relationship between timestamps and targets.\nI trained on training data by grouping the data by timestamp, and it was successful.\n\nI wasn't sure about the relationship until I submit the 0.995 answer.\nI sincerely apologize for not revealing the timestamp effect on the target.\nGreed ruined me.\n\nI apologize for not revealing the data timestamp leakage.",
    "1354211": "If leakage happens not only in timestamp but also in the image level, then organizers should replace both train and test datasets. Replacing or fixing only testset will create a bias between the train and the test images. ",
    "1354014": "Thank you for reporting. I never expected to see GBDT in CV competition except for stacking 😅\n\nIf your hypothesis is correct, train dataset also needs updating?",
    "1354013": "I found the same and was about to share.  Thanks for sharing!",
    "1354121": "> Perhaps the host added the target signal to the image after normalizing the background noise. (Of course, to prevent leakage, the image should be normalized after the target signal is added.\n\nI was theorizing the same, but in the data description they say:\n\n> After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array.\n\nSo I didn't even check... ",
    "1354090": "Yep, it seems my custom nn *implicitly* used information about on-off channel distribution. Interesting competition :sarcasm: At least it cool to learn different way to find and exploit the leaks. Who knows what else the aliens have prepared for us?\n ",
    "1354009": "Although it is clear this finding is real, I think the organizers wanted to make sure this did not happen, according to this statement in the Data Information tab:\n\n`After we perform the signal injections, we normalize each snippet, so you probably can’t identify most of the needles just by looking for excess energy in the corresponding array. `",
    "1354007": "I noticed this point recently, but I am not sure I should call it a leakage.\nI think it is OK if \"real\" alien signals have such statistical properties.",
    "1354946": "Very interesting to see so many old timers flocking to this competition. It definitely looks like this is the most interesting one of those currently being ran. For now though, time to check out other comps until we get an updated train/test..."
  }
}