{
  "id": 266513,
  "title": "Is second magic an artifact (bad) or is it a consequence of physics (good)?",
  "url": "/competitions/seti-breakthrough-listen/discussion/266513",
  "author_name": "CPMP",
  "post_date": "2021-08-19T11:30:07.710000",
  "votes": 24,
  "comment_count": 39,
  "views": 0,
  "content": "<p>I have been thinking a lot about why background could be cancelled by matching images as done by top team.  I, like most participants, assumed that background images were form real data hence had no reason to be nearly identical to others.  Actually, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> on our team spent some time with an auto encoder then clustering on latent vectors.  He did find that images could be clustered, and when looking, image background looked very similar.  This was near the end and we did not thought of canceling background by, say, removing the average of these common images from each of the images.</p>\n<p>Here is an example of group Giba found:</p>\n<p><img src=\"https://i.imgur.com/WEJI4mm.png\" alt=\"group\"></p>\n<p>This common background can easily be dismissed about something that does not occur in real world.  Furthermore, using it to boost model score would be useless in same real world.  Several made this point in various comments.</p>\n<p>However there may be physics reasons for this common background pattern. </p>\n<ol>\n<li>If you point the telescope again in the same direction then you should get a very similar background.  </li>\n<li>Related to 1, you may take several cadences before moving the radio telescope to a different region</li>\n<li>The data may span a larger frequency range that what we got, and host cropped smaller band for us.  Think of random cropping data augmentation: crops overlap.  Same here.</li>\n<li>etc.</li>\n</ol>\n<p>If these reason are valid, then the background cancellation devised by top team could be useful.  Indeed, host has meta data that record the direction the telescope is pointed to.  Then matching images taken from very similar directions could enable background cancellation.  I think it is why host encouraged top team to keep doing it when they were asked.</p>\n<p>TL;DR it maybe that second magic is a clever use of the way real data was sampled.  In that case the top solution is useful in production.</p>\n<p>I'd love to get host take on this.</p>",
  "messages": [
    {
      "id": 1481227,
      "postDate": "2021-08-19T11:30:07.710Z",
      "content": "<p>I have been thinking a lot about why background could be cancelled by matching images as done by top team.  I, like most participants, assumed that background images were form real data hence had no reason to be nearly identical to others.  Actually, <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> on our team spent some time with an auto encoder then clustering on latent vectors.  He did find that images could be clustered, and when looking, image background looked very similar.  This was near the end and we did not thought of canceling background by, say, removing the average of these common images from each of the images.</p>\n<p>Here is an example of group Giba found:</p>\n<p><img src=\"https://i.imgur.com/WEJI4mm.png\" alt=\"group\"></p>\n<p>This common background can easily be dismissed about something that does not occur in real world.  Furthermore, using it to boost model score would be useless in same real world.  Several made this point in various comments.</p>\n<p>However there may be physics reasons for this common background pattern. </p>\n<ol>\n<li>If you point the telescope again in the same direction then you should get a very similar background.  </li>\n<li>Related to 1, you may take several cadences before moving the radio telescope to a different region</li>\n<li>The data may span a larger frequency range that what we got, and host cropped smaller band for us.  Think of random cropping data augmentation: crops overlap.  Same here.</li>\n<li>etc.</li>\n</ol>\n<p>If these reason are valid, then the background cancellation devised by top team could be useful.  Indeed, host has meta data that record the direction the telescope is pointed to.  Then matching images taken from very similar directions could enable background cancellation.  I think it is why host encouraged top team to keep doing it when they were asked.</p>\n<p>TL;DR it maybe that second magic is a clever use of the way real data was sampled.  In that case the top solution is useful in production.</p>\n<p>I'd love to get host take on this.</p>",
      "rawMarkdown": "I have been thinking a lot about why background could be cancelled by matching images as done by top team.  I, like most participants, assumed that background images were form real data hence had no reason to be nearly identical to others.  Actually, @titericz on our team spent some time with an auto encoder then clustering on latent vectors.  He did find that images could be clustered, and when looking, image background looked very similar.  This was near the end and we did not thought of canceling background by, say, removing the average of these common images from each of the images.\n\nHere is an example of group Giba found:\n\n![group](https://i.imgur.com/WEJI4mm.png)\n\nThis common background can easily be dismissed about something that does not occur in real world.  Furthermore, using it to boost model score would be useless in same real world.  Several made this point in various comments.\n\nHowever there may be physics reasons for this common background pattern. \n1. If you point the telescope again in the same direction then you should get a very similar background.  \n2. Related to 1, you may take several cadences before moving the radio telescope to a different region\n3. The data may span a larger frequency range that what we got, and host cropped smaller band for us.  Think of random cropping data augmentation: crops overlap.  Same here.\n4. etc.\n\nIf these reason are valid, then the background cancellation devised by top team could be useful.  Indeed, host has meta data that record the direction the telescope is pointed to.  Then matching images taken from very similar directions could enable background cancellation.  I think it is why host encouraged top team to keep doing it when they were asked.\n\nTL;DR it maybe that second magic is a clever use of the way real data was sampled.  In that case the top solution is useful in production.\n\nI'd love to get host take on this.\n",
      "votes": 23
    },
    {
      "id": 1481256,
      "postDate": "2021-08-19T11:47:17.583Z",
      "content": "<p>It is a bad artefact - that whole competition (both pre and after reset) was full of puzzleish things that was forcing you to look at the data in ways that you wouldn't in a real-life application and trying to take advantage of how the data was generated rather than solving the problem at hand.</p>\n<p>I really hate post-discussions that try to relabel the leaks as non-leaks. In my case, looking the data in that way, would have been out of the question. </p>\n<p>In any case ,as I have stated before, kaggle encourages you to look the data in that way. The host liked it for the problem to be seen in that way, the winning team did raise it in a timely manner…so… . Anyway - I have moaned a lot about this and THIS will be my last comment. </p>",
      "rawMarkdown": "It is a bad artefact - that whole competition (both pre and after reset) was full of puzzleish things that was forcing you to look at the data in ways that you wouldn't in a real-life application and trying to take advantage of how the data was generated rather than solving the problem at hand.\n\nI really hate post-discussions that try to relabel the leaks as non-leaks. In my case, looking the data in that way, would have been out of the question. \n\nIn any case ,as I have stated before, kaggle encourages you to look the data in that way. The host liked it for the problem to be seen in that way, the winning team did raise it in a timely manner...so... . Anyway - I have moaned a lot about this and THIS will be my last comment. ",
      "votes": 10,
      "replies": [
        {
          "id": 1481264,
          "postDate": "2021-08-19T11:51:33.163Z",
          "content": "<blockquote>\n  <p>It is a bad artefact</p>\n</blockquote>\n<p>You may be true but at this point you cannot know for sure.  Only the host can tell us which is which.</p>",
          "rawMarkdown": "> It is a bad artefact\n\nYou may be true but at this point you cannot know for sure.  Only the host can tell us which is which.",
          "votes": 2
        },
        {
          "id": 1481595,
          "postDate": "2021-08-19T15:39:12.420Z",
          "content": "<p>Once again i doubled <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> . This competition should be unsupervised problem. As a competitive team we were forced to do data mining (especially after Watercooled jump) instead of solving problem. To be clear i have no problem with Watercooled abusing synthetics.</p>",
          "rawMarkdown": "Once again i doubled @kazanova . This competition should be unsupervised problem. As a competitive team we were forced to do data mining (especially after Watercooled jump) instead of solving problem. To be clear i have no problem with Watercooled abusing synthetics.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1482106,
      "postDate": "2021-08-19T21:52:52.163Z",
      "content": "<p>An update to the discussion below: I've managed to recreate a version of \"magic #2\" in <a href=\"https://www.kaggle.com/friedchips/magic-2-an-explanation\" target=\"_blank\">this notebook</a>. Cleaned data by finding a clone of a sample with needle looks like this:<br>\n<img src=\"https://www.kaggleusercontent.com/kf/72439533/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..xNe4-aAn6LI9r92BJKFUBg.vxv7YobvU-0p2e3hf1HfoF9JzcrxsQ3SPa-qcYZdTIoqrrNrdK84_7OLh7i1H8xDY_L5qcaK_IfiqHfLiA9ab6Q5VNNI7DLW_I0rCxdrb5Q6rkmz-iN-FRQchMuOUDaQKoMPJtqnNx2DsK_5fd1K_AJCEVA8vs9WurGdKYi2zs0LCYYwt8YdAua9j5B9lxlepBV_MH0yZYfqK_GY4oJfwx3tjxBgA-XN61UTYmXNd60u9jySm9hWIKojI1aBwcFljhQfL-M5yRFdAwJPPbbF5u8lgEytq3QCn9h-ldXuYeoTjfqrzqD2Cq_va-qHsm1kmBmqXolyxqrnjKqFTreIFbLbg-GvDKCPuzzKn2i3hUJMKbGQT-CLDxPDq9GMMeUhKlkEbNJnTlYyUUn4JE3Og4uZJ-es2FvUlbHJdw-l7v6uCeO0gMD-mx4Rl4H-NYinxjqLksOqPLdVAaAvzwj-0pcN5BOCAHFIyWpQ76tTYbglIcAsyz97MRRQn9iPMPXozFe8g_LD986xElWPYnz2-1QEYtiDnREzcDZlkaP7gA3eXTrkeWV1j4pOha8ZlqF0HWoaLO-lJKzdugrcgGn-zRC6Z1hT9OUXBKcJcpCofg2sp51Cxp9PhghO7DHH9WgOZZubcMYa8aDTx1aTjgn-aD2YQ9OTlNObs6NbyezDP-k.gDmnbxaHQ_tilG1GmeruVQ/__results___files/__results___16_0.png\" alt=\"\"></p>\n<p>You can even see the rectangles where the competition hosts inserted the needles.</p>",
      "rawMarkdown": "An update to the discussion below: I've managed to recreate a version of \"magic #2\" in [this notebook](https://www.kaggle.com/friedchips/magic-2-an-explanation). Cleaned data by finding a clone of a sample with needle looks like this:\n![](https://www.kaggleusercontent.com/kf/72439533/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..xNe4-aAn6LI9r92BJKFUBg.vxv7YobvU-0p2e3hf1HfoF9JzcrxsQ3SPa-qcYZdTIoqrrNrdK84_7OLh7i1H8xDY_L5qcaK_IfiqHfLiA9ab6Q5VNNI7DLW_I0rCxdrb5Q6rkmz-iN-FRQchMuOUDaQKoMPJtqnNx2DsK_5fd1K_AJCEVA8vs9WurGdKYi2zs0LCYYwt8YdAua9j5B9lxlepBV_MH0yZYfqK_GY4oJfwx3tjxBgA-XN61UTYmXNd60u9jySm9hWIKojI1aBwcFljhQfL-M5yRFdAwJPPbbF5u8lgEytq3QCn9h-ldXuYeoTjfqrzqD2Cq_va-qHsm1kmBmqXolyxqrnjKqFTreIFbLbg-GvDKCPuzzKn2i3hUJMKbGQT-CLDxPDq9GMMeUhKlkEbNJnTlYyUUn4JE3Og4uZJ-es2FvUlbHJdw-l7v6uCeO0gMD-mx4Rl4H-NYinxjqLksOqPLdVAaAvzwj-0pcN5BOCAHFIyWpQ76tTYbglIcAsyz97MRRQn9iPMPXozFe8g_LD986xElWPYnz2-1QEYtiDnREzcDZlkaP7gA3eXTrkeWV1j4pOha8ZlqF0HWoaLO-lJKzdugrcgGn-zRC6Z1hT9OUXBKcJcpCofg2sp51Cxp9PhghO7DHH9WgOZZubcMYa8aDTx1aTjgn-aD2YQ9OTlNObs6NbyezDP-k.gDmnbxaHQ_tilG1GmeruVQ/__results___files/__results___16_0.png)\n\nYou can even see the rectangles where the competition hosts inserted the needles.\n",
      "votes": 7
    },
    {
      "id": 1481782,
      "postDate": "2021-08-19T17:00:43.123Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a><br>\nIt's not physics based, it's a leak. Full samples are <em>exactly</em> repeated twice or more in the data set, simply shifted by some pixels in the frequency direction (the 256-size axis). Proof:<br>\nHere is image 0028a35de92941d, containing a needle (images are transposed for space reasons, time is horizontal):<br>\n<img src=\"https://i.imgur.com/epyD6Pn.png\" alt=\"img\"><br>\nHere is image e26ddbf40e0f281:<br>\n<img src=\"https://i.imgur.com/Rgi12YR.png\" alt=\"img\"><br>\nThey are identical where they overlap. Compare relative to the horizontal line. Just roll the 2nd image 135 pixels along the frequency axis and subtract from the 1st image:<br>\n<img src=\"https://i.imgur.com/W8aNcRe.png\" alt=\"\"><br>\nEven without any renormalization it's obvious. If you add the magic #2 renormalization, you get a perfect fit.<br>\nMost images have such a clone in the data, as I'm finding out right now…</p>",
      "rawMarkdown": "@cpmpml, @cdeotte, @kazanova\nIt's not physics based, it's a leak. Full samples are *exactly* repeated twice or more in the data set, simply shifted by some pixels in the frequency direction (the 256-size axis). Proof:\nHere is image 0028a35de92941d, containing a needle (images are transposed for space reasons, time is horizontal):\n![img](https://i.imgur.com/epyD6Pn.png)\nHere is image e26ddbf40e0f281:\n![img](https://i.imgur.com/Rgi12YR.png)\nThey are identical where they overlap. Compare relative to the horizontal line. Just roll the 2nd image 135 pixels along the frequency axis and subtract from the 1st image:\n![](https://i.imgur.com/W8aNcRe.png)\nEven without any renormalization it's obvious. If you add the magic #2 renormalization, you get a perfect fit.\nMost images have such a clone in the data, as I'm finding out right now...\n",
      "votes": 7,
      "replies": [
        {
          "id": 1481921,
          "postDate": "2021-08-19T18:19:41.933Z",
          "content": "<p>This is an example of my case 3: overlapping crop.  It would be a leak if one of them had a message in the overlapping region while the other didn't.  Is it the case here?</p>",
          "rawMarkdown": "This is an example of my case 3: overlapping crop.  It would be a leak if one of them had a message in the overlapping region while the other didn't.  Is it the case here?"
        },
        {
          "id": 1481960,
          "postDate": "2021-08-19T18:55:56.327Z",
          "content": "<p>Not in this example (it was the first I found), but in many others. One example which is a clear leak:<br>\n83b61505803557d contains a needle:<br>\n<img src=\"https://i.imgur.com/fGO8ytz.png\" alt=\"\"><br>\nwhich can be perfectly isolated by subtracting 3fe446de262d513:<br>\n<img src=\"https://i.imgur.com/3EOTMkQ.png\" alt=\"\">.</p>\n<p>I wouldn't mind if such an overlap happened occasionally, but they are everywhere. I was sure it was a leak once I saw this image from the 1st place post:<br>\n<img src=\"https://i.imgur.com/XxC4lDI.png\" alt=\"\"><br>\nThis is <em>perfect</em> noise removal, which doesn't happen in reality. The matching sample even also had a needle, which appears in negative after the subtraction.</p>\n<p>This isn't noise removal, it's about finding the sample again in the dataset, this time without a signal.</p>",
          "rawMarkdown": "Not in this example (it was the first I found), but in many others. One example which is a clear leak:\n83b61505803557d contains a needle:\n![](https://i.imgur.com/fGO8ytz.png)\nwhich can be perfectly isolated by subtracting 3fe446de262d513:\n![](https://i.imgur.com/3EOTMkQ.png).\n\nI wouldn't mind if such an overlap happened occasionally, but they are everywhere. I was sure it was a leak once I saw this image from the 1st place post:\n![](https://i.imgur.com/XxC4lDI.png)\nThis is *perfect* noise removal, which doesn't happen in reality. The matching sample even also had a needle, which appears in negative after the subtraction.\n\nThis isn't noise removal, it's about finding the sample again in the dataset, this time without a signal.",
          "votes": 10
        },
        {
          "id": 1481986,
          "postDate": "2021-08-19T19:24:54.653Z",
          "content": "<p>Great detective work <a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a>. The competition data does appear to contains leaks. </p>\n<p>However i'm curious if this technique would also work on real data as a method of comparing \"on\" cadence to another \"on\" cadence in the same direction with slightly different time. Would comparing \"on\" to \"on\" be better than comparing \"on\" to \"off\"? (Since \"off\" is pointed in another direction? and \"off\" has a different time too anyway.)</p>",
          "rawMarkdown": "Great detective work @friedchips. The competition data does appear to contains leaks. \n\nHowever i'm curious if this technique would also work on real data as a method of comparing \"on\" cadence to another \"on\" cadence in the same direction with slightly different time. Would comparing \"on\" to \"on\" be better than comparing \"on\" to \"off\"? (Since \"off\" is pointed in another direction? and \"off\" has a different time too anyway.)"
        },
        {
          "id": 1482051,
          "postDate": "2021-08-19T20:38:34.583Z",
          "content": "<p>Thanks - this was only possible thanks to your explanations.</p>\n<p>No, this technique wouldn't really work because these are spectrograms. Both the signal and the noise are time-dependent and non-periodic. Noise removal only works well when <em>both</em> are stationary or at least periodic. And what the winning team has achieved is to detect signals with incredibly low SNRs. This was only possible because they knew the noise and its time-dependence perfectly.</p>\n<p>Or, to put it in another way: if the intensity in the spectrogram at a specific frequency changes, is it due to noise or a signal? You don't know, otherwise finding SETI signals would be trivial.</p>\n<p>One remark: I do definitely not want to critizise the winning team for doing this, their work was absolutely brilliant. But (unfortunate for science) in a \"master detective\" way, not in a way that will be helpful for SETI.</p>",
          "rawMarkdown": "Thanks - this was only possible thanks to your explanations.\n\nNo, this technique wouldn't really work because these are spectrograms. Both the signal and the noise are time-dependent and non-periodic. Noise removal only works well when *both* are stationary or at least periodic. And what the winning team has achieved is to detect signals with incredibly low SNRs. This was only possible because they knew the noise and its time-dependence perfectly.\n\nOr, to put it in another way: if the intensity in the spectrogram at a specific frequency changes, is it due to noise or a signal? You don't know, otherwise finding SETI signals would be trivial.\n\nOne remark: I do definitely not want to critizise the winning team for doing this, their work was absolutely brilliant. But (unfortunate for science) in a \"master detective\" way, not in a way that will be helpful for SETI.",
          "votes": 8
        },
        {
          "id": 1482210,
          "postDate": "2021-08-20T00:26:52.640Z",
          "content": "<p>OK, you convinced me they reused same background over and over.</p>\n<p>I don't get why as they have an almost unlimited supply of data.</p>\n<p>As a result they get a top solution which is partly not useful.</p>",
          "rawMarkdown": "OK, you convinced me they reused same background over and over.\n\nI don't get why as they have an almost unlimited supply of data.\n\nAs a result they get a top solution which is partly not useful.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1481237,
      "postDate": "2021-08-19T11:34:58.347Z",
      "content": "<p>Thanks for the post <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - I definitely agree with the physics reasons and am also curious about learning more about it, we discussed this also before internally with similar hypotheses.</p>\n<blockquote>\n  <p>This common background can easily be dismissed about something that does not occur in real world.</p>\n</blockquote>\n<p>I even think that it might be useful in different problem settings, although obviously the effects might not be as significant as we have it here. See a comment I made here:<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215</a> </p>",
      "rawMarkdown": "Thanks for the post @cpmpml - I definitely agree with the physics reasons and am also curious about learning more about it, we discussed this also before internally with similar hypotheses.\n\n> This common background can easily be dismissed about something that does not occur in real world.\n\nI even think that it might be useful in different problem settings, although obviously the effects might not be as significant as we have it here. See a comment I made here:\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215 ",
      "votes": 6
    },
    {
      "id": 1481492,
      "postDate": "2021-08-19T14:42:13.987Z",
      "content": "<p>I believe this property can be physics based and not a leak. The host does not need to repeat any backgrounds. Here's how it works.</p>\n<p>We collect \"on\" cadence signals from time=0-1, 2-3, and 4-5. (And \"off\" cadence from 1-2, 3-4, 5-6). The problem with using the \"off\" cadence as control images is that they are different direction in space.</p>\n<p>Better control images would be \"on\" cadence signals from time=6-7, 8-9, 10-11. So it is best to compare one image's \"on\" cadence with another image's \"on\" cadence like first place team did.</p>\n<p>(Note that a column has fixed frequency and an interval of time. So comparing <code>column x</code> from one image \"on\" cadence to <code>column x</code> from another image \"on\" cadence is like comparing the frequency energy at the same point in space over two different times)</p>",
      "rawMarkdown": "I believe this property can be physics based and not a leak. The host does not need to repeat any backgrounds. Here's how it works.\n\nWe collect \"on\" cadence signals from time=0-1, 2-3, and 4-5. (And \"off\" cadence from 1-2, 3-4, 5-6). The problem with using the \"off\" cadence as control images is that they are different direction in space.\n\nBetter control images would be \"on\" cadence signals from time=6-7, 8-9, 10-11. So it is best to compare one image's \"on\" cadence with another image's \"on\" cadence like first place team did.\n\n(Note that a column has fixed frequency and an interval of time. So comparing `column x` from one image \"on\" cadence to `column x` from another image \"on\" cadence is like comparing the frequency energy at the same point in space over two different times)\n",
      "votes": 3,
      "replies": [
        {
          "id": 1481608,
          "postDate": "2021-08-19T15:44:34.753Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I believe ON and OFF are from the same direction, but collect in different time because eyeballing both ON and OFF images we notice same kind of patterns. And thats the idea about anomaly detection, find patterns in ON that are not present in OFF images.</p>",
          "rawMarkdown": "@cdeotte I believe ON and OFF are from the same direction, but collect in different time because eyeballing both ON and OFF images we notice same kind of patterns. And thats the idea about anomaly detection, find patterns in ON that are not present in OFF images."
        },
        {
          "id": 1481679,
          "postDate": "2021-08-19T16:13:40.403Z",
          "content": "<p><a href=\"https://www.kaggle.com/giba\" target=\"_blank\">@giba</a> ON and OFF are from slightly different directions so that background is similar but message appears only in one of them.</p>",
          "rawMarkdown": "@giba ON and OFF are from slightly different directions so that background is similar but message appears only in one of them.",
          "votes": 1
        },
        {
          "id": 1481870,
          "postDate": "2021-08-19T17:43:31.800Z",
          "content": "<p>From the data description <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/overview/data-information\" target=\"_blank\">here</a> is says</p>\n<blockquote>\n  <p>One method we use to isolate candidate technosignatures from RFI is to look for signals that appear to be coming from particular positions on the sky. Typically we do this by alternating observations of our primary target star with observations of three nearby stars: 5 minutes on star “A”, then 5 minutes on star “B”, then back to star “A” for 5 minutes, then “C”, then back to “A”, then finishing with 5 minutes on star “D”. One set of six observations (ABACAD) is referred to as a “cadence”. Since we’re just giving you a small range of frequencies for each cadence, we refer to the datasets you’ll be analyzing as “cadence snippets”.</p>\n</blockquote>\n<p>Furthermore, if you train adversarial validation comparing train \"on\" cadence to train \"off\" cadence, it achieves AUC 0.99 or something. Even using only target=0.</p>",
          "rawMarkdown": "From the data description [here][1] is says\n\n> One method we use to isolate candidate technosignatures from RFI is to look for signals that appear to be coming from particular positions on the sky. Typically we do this by alternating observations of our primary target star with observations of three nearby stars: 5 minutes on star “A”, then 5 minutes on star “B”, then back to star “A” for 5 minutes, then “C”, then back to “A”, then finishing with 5 minutes on star “D”. One set of six observations (ABACAD) is referred to as a “cadence”. Since we’re just giving you a small range of frequencies for each cadence, we refer to the datasets you’ll be analyzing as “cadence snippets”.\n\nFurthermore, if you train adversarial validation comparing train \"on\" cadence to train \"off\" cadence, it achieves AUC 0.99 or something. Even using only target=0.\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/overview/data-information"
        },
        {
          "id": 1482257,
          "postDate": "2021-08-20T01:22:40.117Z",
          "content": "<p>I read that description many times, but did you eyeballed ON and OFF images of the same id ? Patterns are very similar.</p>",
          "rawMarkdown": "I read that description many times, but did you eyeballed ON and OFF images of the same id ? Patterns are very similar.",
          "votes": 1
        },
        {
          "id": 1482276,
          "postDate": "2021-08-20T02:00:36.053Z",
          "content": "<p>Yes, they are very similar. My point is that comparing two \"on\" cadence is even <strong>more</strong> similar because this is fake data. </p>\n<p>Let me explain. Imagine that we did \"ABACAD\" and immediately after did another \"ABACAD\". Call the first sequence \"A1, B2, A3, C4, A5, D6\" and call the second sequence \"A7, B8, A9, C10, A11, D12\". The first A1 is recorded from time=0 to time=1. Then B2 is recorded from from time=1 to time=2. etc etc. And finally D12 is recorded from time=11 to time=12.</p>\n<p>In real life if there is an alien signal in A1, A3, and A5. Then it would most likely be in A7, A9, A11 which is only a few seconds later. (So in real life, we cannot compare A1 to A7 to detect alien. We must compare A1 to B1).</p>\n<p>This is not real life. This is fake data. The host puts alien in A1, A3, and A5. But the host does <strong>not</strong> put alien in A7, A9, A11. Therefore, it is best to compare A1, A3 and A5 with A7, A9, and A11. Because all of these are pointed in the same direction and occur only seconds from each other.</p>\n<p>The fake data gives us an unnatural opportunity to get unnatural comparisons which are better than anything we can do in real life. (In real life we would compare A1 to B1. Or A3 to B4 which differ in both direction and time).</p>",
          "rawMarkdown": "Yes, they are very similar. My point is that comparing two \"on\" cadence is even **more** similar because this is fake data. \n\nLet me explain. Imagine that we did \"ABACAD\" and immediately after did another \"ABACAD\". Call the first sequence \"A1, B2, A3, C4, A5, D6\" and call the second sequence \"A7, B8, A9, C10, A11, D12\". The first A1 is recorded from time=0 to time=1. Then B2 is recorded from from time=1 to time=2. etc etc. And finally D12 is recorded from time=11 to time=12.\n\nIn real life if there is an alien signal in A1, A3, and A5. Then it would most likely be in A7, A9, A11 which is only a few seconds later. (So in real life, we cannot compare A1 to A7 to detect alien. We must compare A1 to B1).\n\nThis is not real life. This is fake data. The host puts alien in A1, A3, and A5. But the host does **not** put alien in A7, A9, A11. Therefore, it is best to compare A1, A3 and A5 with A7, A9, and A11. Because all of these are pointed in the same direction and occur only seconds from each other.\n\nThe fake data gives us an unnatural opportunity to get unnatural comparisons which are better than anything we can do in real life. (In real life we would compare A1 to B1. Or A3 to B4 which differ in both direction and time).",
          "votes": 2
        },
        {
          "id": 1482560,
          "postDate": "2021-08-20T06:38:16.313Z",
          "content": "<p>In real life the second observation (A7 - D12) wouldn't even exist. Telescope time is too valuable to look at the same star twice in  a short time.<br>\nThe only reason why the adversarial validation works is because many \"cadence snippets\" come from the same observation, even overlapping massively.</p>",
          "rawMarkdown": "In real life the second observation (A7 - D12) wouldn't even exist. Telescope time is too valuable to look at the same star twice in  a short time.\nThe only reason why the adversarial validation works is because many \"cadence snippets\" come from the same observation, even overlapping massively."
        },
        {
          "id": 1482881,
          "postDate": "2021-08-20T10:33:59.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> Are there any examples in the data that overlap partially in the time dimension? Above you  have provided examples that overlap partially in the frequency dimension. </p>\n<p>For example, the data contains background time=0-6, freq=1000-2000 and the data contains time=0-6, freq 1500-2500. But does the data contain overlap like time=0-6, freq=1000-2000 and also time=3-9, freq=1000-2000? </p>\n<p>(Or does the data contain my example of \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" which is time=0-6, freq=1000-2000 and also time=6-12, freq=1000-2000)</p>\n<p>Because if a telescope doesn't point in the same direction in real life, why would there be clusters in the dataset? I'm thinking this is a second leak that is different than the first leak.</p>\n<ul>\n<li>leak 1 - overlapping backgrounds (and we use noise cancelation)</li>\n<li>leak 2 - providing \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" where first is <strong>with</strong> alien and second is <strong>without</strong> alien. (and we use cluster analysis)</li>\n</ul>",
          "rawMarkdown": "@friedchips Are there any examples in the data that overlap partially in the time dimension? Above you  have provided examples that overlap partially in the frequency dimension. \n\nFor example, the data contains background time=0-6, freq=1000-2000 and the data contains time=0-6, freq 1500-2500. But does the data contain overlap like time=0-6, freq=1000-2000 and also time=3-9, freq=1000-2000? \n\n(Or does the data contain my example of \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" which is time=0-6, freq=1000-2000 and also time=6-12, freq=1000-2000)\n\nBecause if a telescope doesn't point in the same direction in real life, why would there be clusters in the dataset? I'm thinking this is a second leak that is different than the first leak.\n\n* leak 1 - overlapping backgrounds (and we use noise cancelation)\n* leak 2 - providing \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" where first is **with** alien and second is **without** alien. (and we use cluster analysis)",
          "votes": 1
        },
        {
          "id": 1483037,
          "postDate": "2021-08-20T12:11:28.197Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> No, they don't overlap in the time direction because they physically can't. Those 273x6 time steps are all there is. One ABACAD sequence is one real measurement process in time. The 273 time steps per image correspond to ~5min in real time. Then ~2min of data is missing, because the telescope is moving to the next target. Then again 5min of consecutive measurements. and so on. One ABACAD sample is about 40+ minutes of real time.</p>\n<p>A full sequence has definitely only 273x6 time steps and -as far as I can tell- maybe around 1000 frequency bins. Out of this data they cut out the \"cadence snippets\" of size (273x6, 256). And -probably to get more samples- they used significant overlap in the frequency dimension.</p>\n<p>But all samples from one cadence are perfectly aligned in time.</p>",
          "rawMarkdown": "@cdeotte No, they don't overlap in the time direction because they physically can't. Those 273x6 time steps are all there is. One ABACAD sequence is one real measurement process in time. The 273 time steps per image correspond to ~5min in real time. Then ~2min of data is missing, because the telescope is moving to the next target. Then again 5min of consecutive measurements. and so on. One ABACAD sample is about 40+ minutes of real time.\n\nA full sequence has definitely only 273x6 time steps and -as far as I can tell- maybe around 1000 frequency bins. Out of this data they cut out the \"cadence snippets\" of size (273x6, 256). And -probably to get more samples- they used significant overlap in the frequency dimension.\n\nBut all samples from one cadence are perfectly aligned in time.\n\n\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1481345,
      "postDate": "2021-08-19T12:57:59.783Z",
      "content": "<p>That is a good question <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, but looking at the time axis in the competition images I would say NO, frequency spectrum is not periodic, the background patterns repeats but at random intensities, frequencies and time periods. So in real world I bet it couldn't be \"perfectly\" matched if you point to the same region again.</p>",
      "rawMarkdown": "That is a good question @cpmpml, but looking at the time axis in the competition images I would say NO, frequency spectrum is not periodic, the background patterns repeats but at random intensities, frequencies and time periods. So in real world I bet it couldn't be \"perfectly\" matched if you point to the same region again.",
      "votes": 3,
      "replies": [
        {
          "id": 1481424,
          "postDate": "2021-08-19T13:47:45.293Z",
          "content": "<p>Good point, time patterns are indeed supporting they reused same data for image with and without message.</p>",
          "rawMarkdown": "Good point, time patterns are indeed supporting they reused same data for image with and without message.",
          "votes": 2
        },
        {
          "id": 1481771,
          "postDate": "2021-08-19T16:58:46.997Z",
          "content": "<p>Recall \"TGS Salt Identification Challenge\"</p>",
          "rawMarkdown": "Recall \"TGS Salt Identification Challenge\"",
          "votes": 1
        },
        {
          "id": 1482211,
          "postDate": "2021-08-20T00:27:56.697Z",
          "content": "<p>I thought of TGS indeed.  And also other cases where \"reconstructing the mosaic\" was the key.</p>",
          "rawMarkdown": "I thought of TGS indeed.  And also other cases where \"reconstructing the mosaic\" was the key."
        }
      ]
    },
    {
      "id": 1482770,
      "postDate": "2021-08-20T09:07:38.967Z",
      "content": "<p>Somehow i completely forgot most important part: before restart with the same clustering we find one static RFI pattern, which was included in old_train about ~5000 times (with slightest variations). <strong>These samples contains 26% anomalies</strong>. So out of 6000 total anomalies there was 1300 on that specific sample (we didnt even check percentage in old_test, for reasons).  Probably this was just a way organizers tried to fool algorithms on training with background, not signals. After restart they probably did the same thing, but bit differently. </p>",
      "rawMarkdown": "Somehow i completely forgot most important part: before restart with the same clustering we find one static RFI pattern, which was included in old_train about ~5000 times (with slightest variations). **These samples contains 26% anomalies**. So out of 6000 total anomalies there was 1300 on that specific sample (we didnt even check percentage in old_test, for reasons).  Probably this was just a way organizers tried to fool algorithms on training with background, not signals. After restart they probably did the same thing, but bit differently. ",
      "votes": 4,
      "replies": [
        {
          "id": 1482966,
          "postDate": "2021-08-20T11:19:30.567Z",
          "content": "<p>This is interesting <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> . Maybe the host did this on purpose. Maybe certain backgrounds in train data have a high percentage of positive targets and then those same backgrounds in test have a lower percentage. This may be the cause of the CV LB gap.</p>\n<p>Maybe the host did this to find ML algorithms that can avoid learning spurious correlations. Avoiding spurious correlation and detecting true signal is a difficult and important task.</p>",
          "rawMarkdown": "This is interesting @bakeryproducts . Maybe the host did this on purpose. Maybe certain backgrounds in train data have a high percentage of positive targets and then those same backgrounds in test have a lower percentage. This may be the cause of the CV LB gap.\n\nMaybe the host did this to find ML algorithms that can avoid learning spurious correlations. Avoiding spurious correlation and detecting true signal is a difficult and important task.",
          "votes": 1
        },
        {
          "id": 1482967,
          "postDate": "2021-08-20T11:21:26.297Z",
          "content": "<p>This supports one of my experiments. During training, if you add the \"off\" cadence train images, i.e. <code>np.vstack( img[1::2])</code> with <code>target=0</code> as extra train images, it increases CV by +0.003! But it decreases LB by -0.010. I couldn't understand why. But you may have explained why.</p>",
          "rawMarkdown": "This supports one of my experiments. During training, if you add the \"off\" cadence train images, i.e. `np.vstack( img[1::2])` with `target=0` as extra train images, it increases CV by +0.003! But it decreases LB by -0.010. I couldn't understand why. But you may have explained why."
        },
        {
          "id": 1482984,
          "postDate": "2021-08-20T11:34:19.307Z",
          "content": "<p>That spurious correlation (on background) is what I wanted to emphasize in this post:<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215</a><br>\nWe also discussed this overfit on background in our solution post, and it made us think more about that.</p>",
          "rawMarkdown": "That spurious correlation (on background) is what I wanted to emphasize in this post:\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215\nWe also discussed this overfit on background in our solution post, and it made us think more about that.",
          "votes": 1
        },
        {
          "id": 1482996,
          "postDate": "2021-08-20T11:42:18.890Z",
          "content": "<p>I added off images as yous aid and it improved CV and LB for me.</p>",
          "rawMarkdown": "I added off images as yous aid and it improved CV and LB for me.",
          "votes": 2
        },
        {
          "id": 1483010,
          "postDate": "2021-08-20T11:55:48.607Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I read and thought about your post all day yesterday. IMO, it is one of the most important areas needing growth in ML and DL. It is related to the topic of cause and effect versus correlation. Humans can generalize better than computers because humans understand cause and effect.</p>\n<p>It is not a simple problem of removing backgrounds. Sometimes backgrounds <strong>cause</strong> the the foreground. For example, if a model is classifying animals, it is helpful to see water in the background (i.e. the animal is <strong>caused</strong> to be fish) versus seeing land in the background (i.e. the animal is <strong>caused</strong> to be land animal).</p>\n<p>So when does the background contain signal and when does it not? You might say let's use CV. But what if training data contains a correlation but then test data does not? How can a model ignore background in train data if the background helps increase CV score in train data like this comp?</p>",
          "rawMarkdown": "@philippsinger I read and thought about your post all day yesterday. IMO, it is one of the most important areas needing growth in ML and DL. It is related to the topic of cause and effect versus correlation. Humans can generalize better than computers because humans understand cause and effect.\n\nIt is not a simple problem of removing backgrounds. Sometimes backgrounds **cause** the the foreground. For example, if a model is classifying animals, it is helpful to see water in the background (i.e. the animal is **caused** to be fish) versus seeing land in the background (i.e. the animal is **caused** to be land animal).\n\nSo when does the background contain signal and when does it not? You might say let's use CV. But what if training data contains a correlation but then test data does not? How can a model ignore background in train data if the background helps increase CV score in train data like this comp?",
          "votes": 6
        },
        {
          "id": 1483065,
          "postDate": "2021-08-20T12:35:34.897Z",
          "content": "<p>Yes, it is a very tricky topic, and I also do not have an answer to that. I have many times observed that the models are learning certain correlation, which I as a human would see as spurious. So take chest x-rays for example, there often is a huge data bias in there, because different devices are used and they exhibit different types of x-rays, sometimes even special markers on the images. And the models often pick these effects up, because they contain information about the target (e.g., different devices for different hospitals focusing on different diseases). </p>\n<p>As a human, I would be concerned about my models learning these effects, because I might want the models to generalize to new hospitals / devices, but from a modeling perspective, and specifically on kaggle, you need to keep these biases in, because they help a ton on the metric.</p>\n<p>And then you have problems, where they actually try to force you to generalize to new distributions of data, like partly also in this competition. And then it can be also useful for the actual test metric to remove these correlation effects.</p>\n<p>In best case you might want to have models that are flexible for certain scenarios, but you potentially need to steer them into that direction.</p>\n<p>But overall there is no easy answer to it, and I totally agree with you that it is a very important area that warrants more research and thought. And at least to me that is a big takeaway from this competition.</p>",
          "rawMarkdown": "Yes, it is a very tricky topic, and I also do not have an answer to that. I have many times observed that the models are learning certain correlation, which I as a human would see as spurious. So take chest x-rays for example, there often is a huge data bias in there, because different devices are used and they exhibit different types of x-rays, sometimes even special markers on the images. And the models often pick these effects up, because they contain information about the target (e.g., different devices for different hospitals focusing on different diseases). \n\nAs a human, I would be concerned about my models learning these effects, because I might want the models to generalize to new hospitals / devices, but from a modeling perspective, and specifically on kaggle, you need to keep these biases in, because they help a ton on the metric.\n\nAnd then you have problems, where they actually try to force you to generalize to new distributions of data, like partly also in this competition. And then it can be also useful for the actual test metric to remove these correlation effects.\n\nIn best case you might want to have models that are flexible for certain scenarios, but you potentially need to steer them into that direction.\n\nBut overall there is no easy answer to it, and I totally agree with you that it is a very important area that warrants more research and thought. And at least to me that is a big takeaway from this competition.",
          "votes": 5
        },
        {
          "id": 1483175,
          "postDate": "2021-08-20T13:39:55.740Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Totally agree, I usually go with boats and water example on this topic, and have no answer as well</p>",
          "rawMarkdown": "@cdeotte Totally agree, I usually go with boats and water example on this topic, and have no answer as well",
          "votes": 1
        }
      ]
    },
    {
      "id": 1486709,
      "postDate": "2021-08-23T07:03:44.800Z",
      "content": "<p>(I don't want to criticize 1st team nor hosts. It's just my opinion.)</p>\n<p>As for usefulness for physics, I think magic #1 is also should be discussed.<br>\nI think this competition should be solved as a task of anomaly detection, not the classification.</p>\n<p>Because the alien signal is not the one we can predict, making any hypothesis by human (like S-like signal is a signal) causes overlooking possibly unknown alien signals.</p>\n<p>Many competitor (including me) solves the issue as a classification task. It helps to detect the known signal, but it is no use to detect unknown anomaly signals.</p>\n<p>I wonder what is the hosts wanted to get from the competition.</p>",
      "rawMarkdown": "(I don't want to criticize 1st team nor hosts. It's just my opinion.)\n\nAs for usefulness for physics, I think magic #1 is also should be discussed.\nI think this competition should be solved as a task of anomaly detection, not the classification.\n\nBecause the alien signal is not the one we can predict, making any hypothesis by human (like S-like signal is a signal) causes overlooking possibly unknown alien signals.\n\nMany competitor (including me) solves the issue as a classification task. It helps to detect the known signal, but it is no use to detect unknown anomaly signals.\n\nI wonder what is the hosts wanted to get from the competition.",
      "votes": 2
    },
    {
      "id": 1481576,
      "postDate": "2021-08-19T15:29:31.830Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  I did the same thing with clustering, and sample that you post is just \"tightest\" cluster distant-wise (with total N~90). Total number of clusters i get ~200-300 with N~=30 samples in each. Moreover if you do frequency norm and mean on any M subsamples (or all test images) from test you will get <strong>same</strong> static RFI pattern (also valid for train, but pattern will not be the same). So all test /or/ train signals share same basic pattern. SNR of static pattern ~.01 of image SNR. We thought that this was adversarial difference between train and test. </p>",
      "rawMarkdown": "@cpmpml  I did the same thing with clustering, and sample that you post is just \"tightest\" cluster distant-wise (with total N~90). Total number of clusters i get ~200-300 with N~=30 samples in each. Moreover if you do frequency norm and mean on any M subsamples (or all test images) from test you will get **same** static RFI pattern (also valid for train, but pattern will not be the same). So all test /or/ train signals share same basic pattern. SNR of static pattern ~.01 of image SNR. We thought that this was adversarial difference between train and test. ",
      "votes": 2,
      "replies": [
        {
          "id": 1481586,
          "postDate": "2021-08-19T15:34:23.027Z",
          "content": "<p>tx for sharing.  We found about 200 clusters indeed.</p>",
          "rawMarkdown": "tx for sharing.  We found about 200 clusters indeed.",
          "votes": 1
        },
        {
          "id": 1481616,
          "postDate": "2021-08-19T15:47:37.397Z",
          "content": "<p>And as for train clusters, anomaly percentage in them certainly was above average. I have no exact numbers, but on random ten from cluster of 30 there was 3-8 anomalies</p>",
          "rawMarkdown": "And as for train clusters, anomaly percentage in them certainly was above average. I have no exact numbers, but on random ten from cluster of 30 there was 3-8 anomalies"
        }
      ]
    },
    {
      "id": 1481323,
      "postDate": "2021-08-19T12:43:01.183Z",
      "content": "<p>For 3rd reason, I thought cropped data do not overlap or at least such data should have the same needles (if they have). <br>\nDo you really find any overlap? Are there exactly the same background noise in different data?<br>\nAs I have not fully understood the matching and cancelling methods of the 1st and your teams, correct me if I am wrong.</p>",
      "rawMarkdown": "For 3rd reason, I thought cropped data do not overlap or at least such data should have the same needles (if they have). \nDo you really find any overlap? Are there exactly the same background noise in different data?\nAs I have not fully understood the matching and cancelling methods of the 1st and your teams, correct me if I am wrong.",
      "votes": 2,
      "replies": [
        {
          "id": 1481327,
          "postDate": "2021-08-19T12:49:08.563Z",
          "content": "<p>Our team method is very different, I think <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> will explain it in our writeup.  And if we could match images we did not use the matching at all anyway!</p>\n<p>But net result is indeed that there ar every similar goups of images, for instance one found by Giba:</p>\n<p><img src=\"https://i.imgur.com/WEJI4mm.png\" alt=\"similar\"></p>",
          "rawMarkdown": "Our team method is very different, I think @titericz will explain it in our writeup.  And if we could match images we did not use the matching at all anyway!\n\nBut net result is indeed that there ar every similar goups of images, for instance one found by Giba:\n\n![similar](https://i.imgur.com/WEJI4mm.png)",
          "votes": 2
        },
        {
          "id": 1481418,
          "postDate": "2021-08-19T13:43:50.933Z",
          "content": "<p>Thank you for your comment. I checked the data you mentioned and found the differences are not zero but very small.  It could be a bad artifact and I have to agree with <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>. Anyway, noticing and utilizing this magic must be the result of their strong DS/ML skills. Congrats for 1st and your teams!</p>\n<pre><code>path0 = '../input/seti-breakthrough-listen/test/0/0b5883540795e27.npy'\npath1 = '../input/seti-breakthrough-listen/test/0/024b8089ea7a610.npy'\npath2 = '../input/seti-breakthrough-listen/test/0/08e8b07f8acf858.npy'\npath3 = '../input/seti-breakthrough-listen/test/0/0ec1162c0d34b3a.npy'\npath4 = '../input/seti-breakthrough-listen/test/1/17c269bbb395b58.npy'\npath5 = '../input/seti-breakthrough-listen/test/1/1ba3f5f8503c390.npy'\npath6 = '../input/seti-breakthrough-listen/test/2/20e0677310e055d.npy'\n</code></pre>",
          "rawMarkdown": "Thank you for your comment. I checked the data you mentioned and found the differences are not zero but very small.  It could be a bad artifact and I have to agree with @kazanova. Anyway, noticing and utilizing this magic must be the result of their strong DS/ML skills. Congrats for 1st and your teams!\n```\npath0 = '../input/seti-breakthrough-listen/test/0/0b5883540795e27.npy'\npath1 = '../input/seti-breakthrough-listen/test/0/024b8089ea7a610.npy'\npath2 = '../input/seti-breakthrough-listen/test/0/08e8b07f8acf858.npy'\npath3 = '../input/seti-breakthrough-listen/test/0/0ec1162c0d34b3a.npy'\npath4 = '../input/seti-breakthrough-listen/test/1/17c269bbb395b58.npy'\npath5 = '../input/seti-breakthrough-listen/test/1/1ba3f5f8503c390.npy'\npath6 = '../input/seti-breakthrough-listen/test/2/20e0677310e055d.npy'\n```",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1481256,
      "author_name": "Μαριος Μιχαηλιδης KazAnova",
      "author_url": "",
      "post_date": "2021-08-19T11:47:17.583000",
      "content": "<p>It is a bad artefact - that whole competition (both pre and after reset) was full of puzzleish things that was forcing you to look at the data in ways that you wouldn't in a real-life application and trying to take advantage of how the data was generated rather than solving the problem at hand.</p>\n<p>I really hate post-discussions that try to relabel the leaks as non-leaks. In my case, looking the data in that way, would have been out of the question. </p>\n<p>In any case ,as I have stated before, kaggle encourages you to look the data in that way. The host liked it for the problem to be seen in that way, the winning team did raise it in a timely manner…so… . Anyway - I have moaned a lot about this and THIS will be my last comment. </p>",
      "votes": 10,
      "replies": [
        {
          "id": 1481264,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T11:51:33.163000",
          "content": "<blockquote>\n  <p>It is a bad artefact</p>\n</blockquote>\n<p>You may be true but at this point you cannot know for sure.  Only the host can tell us which is which.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481595,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T15:39:12.420000",
          "content": "<p>Once again i doubled <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a> . This competition should be unsupervised problem. As a competitive team we were forced to do data mining (especially after Watercooled jump) instead of solving problem. To be clear i have no problem with Watercooled abusing synthetics.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1482106,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-19T21:52:52.163000",
      "content": "<p>An update to the discussion below: I've managed to recreate a version of \"magic #2\" in <a href=\"https://www.kaggle.com/friedchips/magic-2-an-explanation\" target=\"_blank\">this notebook</a>. Cleaned data by finding a clone of a sample with needle looks like this:<br>\n<img src=\"https://www.kaggleusercontent.com/kf/72439533/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..xNe4-aAn6LI9r92BJKFUBg.vxv7YobvU-0p2e3hf1HfoF9JzcrxsQ3SPa-qcYZdTIoqrrNrdK84_7OLh7i1H8xDY_L5qcaK_IfiqHfLiA9ab6Q5VNNI7DLW_I0rCxdrb5Q6rkmz-iN-FRQchMuOUDaQKoMPJtqnNx2DsK_5fd1K_AJCEVA8vs9WurGdKYi2zs0LCYYwt8YdAua9j5B9lxlepBV_MH0yZYfqK_GY4oJfwx3tjxBgA-XN61UTYmXNd60u9jySm9hWIKojI1aBwcFljhQfL-M5yRFdAwJPPbbF5u8lgEytq3QCn9h-ldXuYeoTjfqrzqD2Cq_va-qHsm1kmBmqXolyxqrnjKqFTreIFbLbg-GvDKCPuzzKn2i3hUJMKbGQT-CLDxPDq9GMMeUhKlkEbNJnTlYyUUn4JE3Og4uZJ-es2FvUlbHJdw-l7v6uCeO0gMD-mx4Rl4H-NYinxjqLksOqPLdVAaAvzwj-0pcN5BOCAHFIyWpQ76tTYbglIcAsyz97MRRQn9iPMPXozFe8g_LD986xElWPYnz2-1QEYtiDnREzcDZlkaP7gA3eXTrkeWV1j4pOha8ZlqF0HWoaLO-lJKzdugrcgGn-zRC6Z1hT9OUXBKcJcpCofg2sp51Cxp9PhghO7DHH9WgOZZubcMYa8aDTx1aTjgn-aD2YQ9OTlNObs6NbyezDP-k.gDmnbxaHQ_tilG1GmeruVQ/__results___files/__results___16_0.png\" alt=\"\"></p>\n<p>You can even see the rectangles where the competition hosts inserted the needles.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1481782,
      "author_name": "Markus Frank",
      "author_url": "",
      "post_date": "2021-08-19T17:00:43.123000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a><br>\nIt's not physics based, it's a leak. Full samples are <em>exactly</em> repeated twice or more in the data set, simply shifted by some pixels in the frequency direction (the 256-size axis). Proof:<br>\nHere is image 0028a35de92941d, containing a needle (images are transposed for space reasons, time is horizontal):<br>\n<img src=\"https://i.imgur.com/epyD6Pn.png\" alt=\"img\"><br>\nHere is image e26ddbf40e0f281:<br>\n<img src=\"https://i.imgur.com/Rgi12YR.png\" alt=\"img\"><br>\nThey are identical where they overlap. Compare relative to the horizontal line. Just roll the 2nd image 135 pixels along the frequency axis and subtract from the 1st image:<br>\n<img src=\"https://i.imgur.com/W8aNcRe.png\" alt=\"\"><br>\nEven without any renormalization it's obvious. If you add the magic #2 renormalization, you get a perfect fit.<br>\nMost images have such a clone in the data, as I'm finding out right now…</p>",
      "votes": 7,
      "replies": [
        {
          "id": 1481921,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T18:19:41.933000",
          "content": "<p>This is an example of my case 3: overlapping crop.  It would be a leak if one of them had a message in the overlapping region while the other didn't.  Is it the case here?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481960,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T18:55:56.327000",
          "content": "<p>Not in this example (it was the first I found), but in many others. One example which is a clear leak:<br>\n83b61505803557d contains a needle:<br>\n<img src=\"https://i.imgur.com/fGO8ytz.png\" alt=\"\"><br>\nwhich can be perfectly isolated by subtracting 3fe446de262d513:<br>\n<img src=\"https://i.imgur.com/3EOTMkQ.png\" alt=\"\">.</p>\n<p>I wouldn't mind if such an overlap happened occasionally, but they are everywhere. I was sure it was a leak once I saw this image from the 1st place post:<br>\n<img src=\"https://i.imgur.com/XxC4lDI.png\" alt=\"\"><br>\nThis is <em>perfect</em> noise removal, which doesn't happen in reality. The matching sample even also had a needle, which appears in negative after the subtraction.</p>\n<p>This isn't noise removal, it's about finding the sample again in the dataset, this time without a signal.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1481986,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T19:24:54.653000",
          "content": "<p>Great detective work <a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a>. The competition data does appear to contains leaks. </p>\n<p>However i'm curious if this technique would also work on real data as a method of comparing \"on\" cadence to another \"on\" cadence in the same direction with slightly different time. Would comparing \"on\" to \"on\" be better than comparing \"on\" to \"off\"? (Since \"off\" is pointed in another direction? and \"off\" has a different time too anyway.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482051,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-19T20:38:34.583000",
          "content": "<p>Thanks - this was only possible thanks to your explanations.</p>\n<p>No, this technique wouldn't really work because these are spectrograms. Both the signal and the noise are time-dependent and non-periodic. Noise removal only works well when <em>both</em> are stationary or at least periodic. And what the winning team has achieved is to detect signals with incredibly low SNRs. This was only possible because they knew the noise and its time-dependence perfectly.</p>\n<p>Or, to put it in another way: if the intensity in the spectrogram at a specific frequency changes, is it due to noise or a signal? You don't know, otherwise finding SETI signals would be trivial.</p>\n<p>One remark: I do definitely not want to critizise the winning team for doing this, their work was absolutely brilliant. But (unfortunate for science) in a \"master detective\" way, not in a way that will be helpful for SETI.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1482210,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T00:26:52.640000",
          "content": "<p>OK, you convinced me they reused same background over and over.</p>\n<p>I don't get why as they have an almost unlimited supply of data.</p>\n<p>As a result they get a top solution which is partly not useful.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1481237,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-08-19T11:34:58.347000",
      "content": "<p>Thanks for the post <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> - I definitely agree with the physics reasons and am also curious about learning more about it, we discussed this also before internally with similar hypotheses.</p>\n<blockquote>\n  <p>This common background can easily be dismissed about something that does not occur in real world.</p>\n</blockquote>\n<p>I even think that it might be useful in different problem settings, although obviously the effects might not be as significant as we have it here. See a comment I made here:<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215</a> </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1481492,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-19T14:42:13.987000",
      "content": "<p>I believe this property can be physics based and not a leak. The host does not need to repeat any backgrounds. Here's how it works.</p>\n<p>We collect \"on\" cadence signals from time=0-1, 2-3, and 4-5. (And \"off\" cadence from 1-2, 3-4, 5-6). The problem with using the \"off\" cadence as control images is that they are different direction in space.</p>\n<p>Better control images would be \"on\" cadence signals from time=6-7, 8-9, 10-11. So it is best to compare one image's \"on\" cadence with another image's \"on\" cadence like first place team did.</p>\n<p>(Note that a column has fixed frequency and an interval of time. So comparing <code>column x</code> from one image \"on\" cadence to <code>column x</code> from another image \"on\" cadence is like comparing the frequency energy at the same point in space over two different times)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1481608,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2021-08-19T15:44:34.753000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I believe ON and OFF are from the same direction, but collect in different time because eyeballing both ON and OFF images we notice same kind of patterns. And thats the idea about anomaly detection, find patterns in ON that are not present in OFF images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1481679,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T16:13:40.403000",
          "content": "<p><a href=\"https://www.kaggle.com/giba\" target=\"_blank\">@giba</a> ON and OFF are from slightly different directions so that background is similar but message appears only in one of them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481870,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T17:43:31.800000",
          "content": "<p>From the data description <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/overview/data-information\" target=\"_blank\">here</a> is says</p>\n<blockquote>\n  <p>One method we use to isolate candidate technosignatures from RFI is to look for signals that appear to be coming from particular positions on the sky. Typically we do this by alternating observations of our primary target star with observations of three nearby stars: 5 minutes on star “A”, then 5 minutes on star “B”, then back to star “A” for 5 minutes, then “C”, then back to “A”, then finishing with 5 minutes on star “D”. One set of six observations (ABACAD) is referred to as a “cadence”. Since we’re just giving you a small range of frequencies for each cadence, we refer to the datasets you’ll be analyzing as “cadence snippets”.</p>\n</blockquote>\n<p>Furthermore, if you train adversarial validation comparing train \"on\" cadence to train \"off\" cadence, it achieves AUC 0.99 or something. Even using only target=0.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482257,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2021-08-20T01:22:40.117000",
          "content": "<p>I read that description many times, but did you eyeballed ON and OFF images of the same id ? Patterns are very similar.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482276,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T02:00:36.053000",
          "content": "<p>Yes, they are very similar. My point is that comparing two \"on\" cadence is even <strong>more</strong> similar because this is fake data. </p>\n<p>Let me explain. Imagine that we did \"ABACAD\" and immediately after did another \"ABACAD\". Call the first sequence \"A1, B2, A3, C4, A5, D6\" and call the second sequence \"A7, B8, A9, C10, A11, D12\". The first A1 is recorded from time=0 to time=1. Then B2 is recorded from from time=1 to time=2. etc etc. And finally D12 is recorded from time=11 to time=12.</p>\n<p>In real life if there is an alien signal in A1, A3, and A5. Then it would most likely be in A7, A9, A11 which is only a few seconds later. (So in real life, we cannot compare A1 to A7 to detect alien. We must compare A1 to B1).</p>\n<p>This is not real life. This is fake data. The host puts alien in A1, A3, and A5. But the host does <strong>not</strong> put alien in A7, A9, A11. Therefore, it is best to compare A1, A3 and A5 with A7, A9, and A11. Because all of these are pointed in the same direction and occur only seconds from each other.</p>\n<p>The fake data gives us an unnatural opportunity to get unnatural comparisons which are better than anything we can do in real life. (In real life we would compare A1 to B1. Or A3 to B4 which differ in both direction and time).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1482560,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-20T06:38:16.313000",
          "content": "<p>In real life the second observation (A7 - D12) wouldn't even exist. Telescope time is too valuable to look at the same star twice in  a short time.<br>\nThe only reason why the adversarial validation works is because many \"cadence snippets\" come from the same observation, even overlapping massively.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482881,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T10:33:59.637000",
          "content": "<p><a href=\"https://www.kaggle.com/friedchips\" target=\"_blank\">@friedchips</a> Are there any examples in the data that overlap partially in the time dimension? Above you  have provided examples that overlap partially in the frequency dimension. </p>\n<p>For example, the data contains background time=0-6, freq=1000-2000 and the data contains time=0-6, freq 1500-2500. But does the data contain overlap like time=0-6, freq=1000-2000 and also time=3-9, freq=1000-2000? </p>\n<p>(Or does the data contain my example of \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" which is time=0-6, freq=1000-2000 and also time=6-12, freq=1000-2000)</p>\n<p>Because if a telescope doesn't point in the same direction in real life, why would there be clusters in the dataset? I'm thinking this is a second leak that is different than the first leak.</p>\n<ul>\n<li>leak 1 - overlapping backgrounds (and we use noise cancelation)</li>\n<li>leak 2 - providing \"A1, B2, A3, C4, A5, D6\" and \"A7, B8, A9, C10, A11, D12\" where first is <strong>with</strong> alien and second is <strong>without</strong> alien. (and we use cluster analysis)</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483037,
          "author_name": "Markus Frank",
          "author_url": "",
          "post_date": "2021-08-20T12:11:28.197000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> No, they don't overlap in the time direction because they physically can't. Those 273x6 time steps are all there is. One ABACAD sequence is one real measurement process in time. The 273 time steps per image correspond to ~5min in real time. Then ~2min of data is missing, because the telescope is moving to the next target. Then again 5min of consecutive measurements. and so on. One ABACAD sample is about 40+ minutes of real time.</p>\n<p>A full sequence has definitely only 273x6 time steps and -as far as I can tell- maybe around 1000 frequency bins. Out of this data they cut out the \"cadence snippets\" of size (273x6, 256). And -probably to get more samples- they used significant overlap in the frequency dimension.</p>\n<p>But all samples from one cadence are perfectly aligned in time.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1481345,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2021-08-19T12:57:59.783000",
      "content": "<p>That is a good question <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>, but looking at the time axis in the competition images I would say NO, frequency spectrum is not periodic, the background patterns repeats but at random intensities, frequencies and time periods. So in real world I bet it couldn't be \"perfectly\" matched if you point to the same region again.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1481424,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T13:47:45.293000",
          "content": "<p>Good point, time patterns are indeed supporting they reused same data for image with and without message.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481771,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-08-19T16:58:46.997000",
          "content": "<p>Recall \"TGS Salt Identification Challenge\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482211,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T00:27:56.697000",
          "content": "<p>I thought of TGS indeed.  And also other cases where \"reconstructing the mosaic\" was the key.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1482770,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-20T09:07:38.967000",
      "content": "<p>Somehow i completely forgot most important part: before restart with the same clustering we find one static RFI pattern, which was included in old_train about ~5000 times (with slightest variations). <strong>These samples contains 26% anomalies</strong>. So out of 6000 total anomalies there was 1300 on that specific sample (we didnt even check percentage in old_test, for reasons).  Probably this was just a way organizers tried to fool algorithms on training with background, not signals. After restart they probably did the same thing, but bit differently. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1482966,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T11:19:30.567000",
          "content": "<p>This is interesting <a href=\"https://www.kaggle.com/bakeryproducts\" target=\"_blank\">@bakeryproducts</a> . Maybe the host did this on purpose. Maybe certain backgrounds in train data have a high percentage of positive targets and then those same backgrounds in test have a lower percentage. This may be the cause of the CV LB gap.</p>\n<p>Maybe the host did this to find ML algorithms that can avoid learning spurious correlations. Avoiding spurious correlation and detecting true signal is a difficult and important task.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482967,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T11:21:26.297000",
          "content": "<p>This supports one of my experiments. During training, if you add the \"off\" cadence train images, i.e. <code>np.vstack( img[1::2])</code> with <code>target=0</code> as extra train images, it increases CV by +0.003! But it decreases LB by -0.010. I couldn't understand why. But you may have explained why.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482984,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T11:34:19.307000",
          "content": "<p>That spurious correlation (on background) is what I wanted to emphasize in this post:<br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215</a><br>\nWe also discussed this overfit on background in our solution post, and it made us think more about that.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482996,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-20T11:42:18.890000",
          "content": "<p>I added off images as yous aid and it improved CV and LB for me.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483010,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T11:55:48.607000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I read and thought about your post all day yesterday. IMO, it is one of the most important areas needing growth in ML and DL. It is related to the topic of cause and effect versus correlation. Humans can generalize better than computers because humans understand cause and effect.</p>\n<p>It is not a simple problem of removing backgrounds. Sometimes backgrounds <strong>cause</strong> the the foreground. For example, if a model is classifying animals, it is helpful to see water in the background (i.e. the animal is <strong>caused</strong> to be fish) versus seeing land in the background (i.e. the animal is <strong>caused</strong> to be land animal).</p>\n<p>So when does the background contain signal and when does it not? You might say let's use CV. But what if training data contains a correlation but then test data does not? How can a model ignore background in train data if the background helps increase CV score in train data like this comp?</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1483065,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T12:35:34.897000",
          "content": "<p>Yes, it is a very tricky topic, and I also do not have an answer to that. I have many times observed that the models are learning certain correlation, which I as a human would see as spurious. So take chest x-rays for example, there often is a huge data bias in there, because different devices are used and they exhibit different types of x-rays, sometimes even special markers on the images. And the models often pick these effects up, because they contain information about the target (e.g., different devices for different hospitals focusing on different diseases). </p>\n<p>As a human, I would be concerned about my models learning these effects, because I might want the models to generalize to new hospitals / devices, but from a modeling perspective, and specifically on kaggle, you need to keep these biases in, because they help a ton on the metric.</p>\n<p>And then you have problems, where they actually try to force you to generalize to new distributions of data, like partly also in this competition. And then it can be also useful for the actual test metric to remove these correlation effects.</p>\n<p>In best case you might want to have models that are flexible for certain scenarios, but you potentially need to steer them into that direction.</p>\n<p>But overall there is no easy answer to it, and I totally agree with you that it is a very important area that warrants more research and thought. And at least to me that is a big takeaway from this competition.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1483175,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-20T13:39:55.740000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Totally agree, I usually go with boats and water example on this topic, and have no answer as well</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1486709,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2021-08-23T07:03:44.800000",
      "content": "<p>(I don't want to criticize 1st team nor hosts. It's just my opinion.)</p>\n<p>As for usefulness for physics, I think magic #1 is also should be discussed.<br>\nI think this competition should be solved as a task of anomaly detection, not the classification.</p>\n<p>Because the alien signal is not the one we can predict, making any hypothesis by human (like S-like signal is a signal) causes overlooking possibly unknown alien signals.</p>\n<p>Many competitor (including me) solves the issue as a classification task. It helps to detect the known signal, but it is no use to detect unknown anomaly signals.</p>\n<p>I wonder what is the hosts wanted to get from the competition.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1481576,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-19T15:29:31.830000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  I did the same thing with clustering, and sample that you post is just \"tightest\" cluster distant-wise (with total N~90). Total number of clusters i get ~200-300 with N~=30 samples in each. Moreover if you do frequency norm and mean on any M subsamples (or all test images) from test you will get <strong>same</strong> static RFI pattern (also valid for train, but pattern will not be the same). So all test /or/ train signals share same basic pattern. SNR of static pattern ~.01 of image SNR. We thought that this was adversarial difference between train and test. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1481586,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T15:34:23.027000",
          "content": "<p>tx for sharing.  We found about 200 clusters indeed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1481616,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-19T15:47:37.397000",
          "content": "<p>And as for train clusters, anomaly percentage in them certainly was above average. I have no exact numbers, but on random ten from cluster of 30 there was 3-8 anomalies</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1481323,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-08-19T12:43:01.183000",
      "content": "<p>For 3rd reason, I thought cropped data do not overlap or at least such data should have the same needles (if they have). <br>\nDo you really find any overlap? Are there exactly the same background noise in different data?<br>\nAs I have not fully understood the matching and cancelling methods of the 1st and your teams, correct me if I am wrong.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1481327,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-19T12:49:08.563000",
          "content": "<p>Our team method is very different, I think <a href=\"https://www.kaggle.com/titericz\" target=\"_blank\">@titericz</a> will explain it in our writeup.  And if we could match images we did not use the matching at all anyway!</p>\n<p>But net result is indeed that there ar every similar goups of images, for instance one found by Giba:</p>\n<p><img src=\"https://i.imgur.com/WEJI4mm.png\" alt=\"similar\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1481418,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-08-19T13:43:50.933000",
          "content": "<p>Thank you for your comment. I checked the data you mentioned and found the differences are not zero but very small.  It could be a bad artifact and I have to agree with <a href=\"https://www.kaggle.com/kazanova\" target=\"_blank\">@kazanova</a>. Anyway, noticing and utilizing this magic must be the result of their strong DS/ML skills. Congrats for 1st and your teams!</p>\n<pre><code>path0 = '../input/seti-breakthrough-listen/test/0/0b5883540795e27.npy'\npath1 = '../input/seti-breakthrough-listen/test/0/024b8089ea7a610.npy'\npath2 = '../input/seti-breakthrough-listen/test/0/08e8b07f8acf858.npy'\npath3 = '../input/seti-breakthrough-listen/test/0/0ec1162c0d34b3a.npy'\npath4 = '../input/seti-breakthrough-listen/test/1/17c269bbb395b58.npy'\npath5 = '../input/seti-breakthrough-listen/test/1/1ba3f5f8503c390.npy'\npath6 = '../input/seti-breakthrough-listen/test/2/20e0677310e055d.npy'\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1481227": "I have been thinking a lot about why background could be cancelled by matching images as done by top team.  I, like most participants, assumed that background images were form real data hence had no reason to be nearly identical to others.  Actually, @titericz on our team spent some time with an auto encoder then clustering on latent vectors.  He did find that images could be clustered, and when looking, image background looked very similar.  This was near the end and we did not thought of canceling background by, say, removing the average of these common images from each of the images.\n\nHere is an example of group Giba found:\n\n![group](https://i.imgur.com/WEJI4mm.png)\n\nThis common background can easily be dismissed about something that does not occur in real world.  Furthermore, using it to boost model score would be useless in same real world.  Several made this point in various comments.\n\nHowever there may be physics reasons for this common background pattern. \n1. If you point the telescope again in the same direction then you should get a very similar background.  \n2. Related to 1, you may take several cadences before moving the radio telescope to a different region\n3. The data may span a larger frequency range that what we got, and host cropped smaller band for us.  Think of random cropping data augmentation: crops overlap.  Same here.\n4. etc.\n\nIf these reason are valid, then the background cancellation devised by top team could be useful.  Indeed, host has meta data that record the direction the telescope is pointed to.  Then matching images taken from very similar directions could enable background cancellation.  I think it is why host encouraged top team to keep doing it when they were asked.\n\nTL;DR it maybe that second magic is a clever use of the way real data was sampled.  In that case the top solution is useful in production.\n\nI'd love to get host take on this.\n",
    "1481256": "It is a bad artefact - that whole competition (both pre and after reset) was full of puzzleish things that was forcing you to look at the data in ways that you wouldn't in a real-life application and trying to take advantage of how the data was generated rather than solving the problem at hand.\n\nI really hate post-discussions that try to relabel the leaks as non-leaks. In my case, looking the data in that way, would have been out of the question. \n\nIn any case ,as I have stated before, kaggle encourages you to look the data in that way. The host liked it for the problem to be seen in that way, the winning team did raise it in a timely manner...so... . Anyway - I have moaned a lot about this and THIS will be my last comment. ",
    "1482106": "An update to the discussion below: I've managed to recreate a version of \"magic #2\" in [this notebook](https://www.kaggle.com/friedchips/magic-2-an-explanation). Cleaned data by finding a clone of a sample with needle looks like this:\n![](https://www.kaggleusercontent.com/kf/72439533/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..xNe4-aAn6LI9r92BJKFUBg.vxv7YobvU-0p2e3hf1HfoF9JzcrxsQ3SPa-qcYZdTIoqrrNrdK84_7OLh7i1H8xDY_L5qcaK_IfiqHfLiA9ab6Q5VNNI7DLW_I0rCxdrb5Q6rkmz-iN-FRQchMuOUDaQKoMPJtqnNx2DsK_5fd1K_AJCEVA8vs9WurGdKYi2zs0LCYYwt8YdAua9j5B9lxlepBV_MH0yZYfqK_GY4oJfwx3tjxBgA-XN61UTYmXNd60u9jySm9hWIKojI1aBwcFljhQfL-M5yRFdAwJPPbbF5u8lgEytq3QCn9h-ldXuYeoTjfqrzqD2Cq_va-qHsm1kmBmqXolyxqrnjKqFTreIFbLbg-GvDKCPuzzKn2i3hUJMKbGQT-CLDxPDq9GMMeUhKlkEbNJnTlYyUUn4JE3Og4uZJ-es2FvUlbHJdw-l7v6uCeO0gMD-mx4Rl4H-NYinxjqLksOqPLdVAaAvzwj-0pcN5BOCAHFIyWpQ76tTYbglIcAsyz97MRRQn9iPMPXozFe8g_LD986xElWPYnz2-1QEYtiDnREzcDZlkaP7gA3eXTrkeWV1j4pOha8ZlqF0HWoaLO-lJKzdugrcgGn-zRC6Z1hT9OUXBKcJcpCofg2sp51Cxp9PhghO7DHH9WgOZZubcMYa8aDTx1aTjgn-aD2YQ9OTlNObs6NbyezDP-k.gDmnbxaHQ_tilG1GmeruVQ/__results___files/__results___16_0.png)\n\nYou can even see the rectangles where the competition hosts inserted the needles.\n",
    "1481782": "@cpmpml, @cdeotte, @kazanova\nIt's not physics based, it's a leak. Full samples are *exactly* repeated twice or more in the data set, simply shifted by some pixels in the frequency direction (the 256-size axis). Proof:\nHere is image 0028a35de92941d, containing a needle (images are transposed for space reasons, time is horizontal):\n![img](https://i.imgur.com/epyD6Pn.png)\nHere is image e26ddbf40e0f281:\n![img](https://i.imgur.com/Rgi12YR.png)\nThey are identical where they overlap. Compare relative to the horizontal line. Just roll the 2nd image 135 pixels along the frequency axis and subtract from the 1st image:\n![](https://i.imgur.com/W8aNcRe.png)\nEven without any renormalization it's obvious. If you add the magic #2 renormalization, you get a perfect fit.\nMost images have such a clone in the data, as I'm finding out right now...\n",
    "1481237": "Thanks for the post @cpmpml - I definitely agree with the physics reasons and am also curious about learning more about it, we discussed this also before internally with similar hypotheses.\n\n> This common background can easily be dismissed about something that does not occur in real world.\n\nI even think that it might be useful in different problem settings, although obviously the effects might not be as significant as we have it here. See a comment I made here:\nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385#1481215 ",
    "1481492": "I believe this property can be physics based and not a leak. The host does not need to repeat any backgrounds. Here's how it works.\n\nWe collect \"on\" cadence signals from time=0-1, 2-3, and 4-5. (And \"off\" cadence from 1-2, 3-4, 5-6). The problem with using the \"off\" cadence as control images is that they are different direction in space.\n\nBetter control images would be \"on\" cadence signals from time=6-7, 8-9, 10-11. So it is best to compare one image's \"on\" cadence with another image's \"on\" cadence like first place team did.\n\n(Note that a column has fixed frequency and an interval of time. So comparing `column x` from one image \"on\" cadence to `column x` from another image \"on\" cadence is like comparing the frequency energy at the same point in space over two different times)\n",
    "1481345": "That is a good question @cpmpml, but looking at the time axis in the competition images I would say NO, frequency spectrum is not periodic, the background patterns repeats but at random intensities, frequencies and time periods. So in real world I bet it couldn't be \"perfectly\" matched if you point to the same region again.",
    "1482770": "Somehow i completely forgot most important part: before restart with the same clustering we find one static RFI pattern, which was included in old_train about ~5000 times (with slightest variations). **These samples contains 26% anomalies**. So out of 6000 total anomalies there was 1300 on that specific sample (we didnt even check percentage in old_test, for reasons).  Probably this was just a way organizers tried to fool algorithms on training with background, not signals. After restart they probably did the same thing, but bit differently. ",
    "1486709": "(I don't want to criticize 1st team nor hosts. It's just my opinion.)\n\nAs for usefulness for physics, I think magic #1 is also should be discussed.\nI think this competition should be solved as a task of anomaly detection, not the classification.\n\nBecause the alien signal is not the one we can predict, making any hypothesis by human (like S-like signal is a signal) causes overlooking possibly unknown alien signals.\n\nMany competitor (including me) solves the issue as a classification task. It helps to detect the known signal, but it is no use to detect unknown anomaly signals.\n\nI wonder what is the hosts wanted to get from the competition.",
    "1481576": "@cpmpml  I did the same thing with clustering, and sample that you post is just \"tightest\" cluster distant-wise (with total N~90). Total number of clusters i get ~200-300 with N~=30 samples in each. Moreover if you do frequency norm and mean on any M subsamples (or all test images) from test you will get **same** static RFI pattern (also valid for train, but pattern will not be the same). So all test /or/ train signals share same basic pattern. SNR of static pattern ~.01 of image SNR. We thought that this was adversarial difference between train and test. ",
    "1481323": "For 3rd reason, I thought cropped data do not overlap or at least such data should have the same needles (if they have). \nDo you really find any overlap? Are there exactly the same background noise in different data?\nAs I have not fully understood the matching and cancelling methods of the 1st and your teams, correct me if I am wrong."
  }
}