{
  "id": 256369,
  "title": "Alternative input spaces thread",
  "url": "/competitions/seti-breakthrough-listen/discussion/256369",
  "author_name": "Vlad Vaduva",
  "post_date": "2021-08-01T08:26:12.156000",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hello guys,</p>\n<p>Besides the normal 6 or 3 dimensions aligned and used as a input for the convolutional models there can be created different and more complex methodologies and architectures.<br>\nI want this to be a thread with alternative approaches for input spaces that the community tried so far. It does not matter if the concept leaded to good or bad results, <strong>the idea</strong> is the only thing that matters.<br>\nI will start sharing my current ideas.<br>\nI am in the process of testing two alternatives.<br>\nAs one of the main characteristic of a positive signal is the fact that it is present in the ON channels and absent in the OFF channels, the first alternative approach is to calculate the pixels mean of all 3 ON channels, calculate the pixels mean of all OFF channels, calculate the difference and use it as a input feature., I will use the absolute value of this difference because we are interested on the value of the difference (as the resulted pixel value will grow, it will be a strong sign that is a big difference in that region). Practically it will be one channel representative for the difference.<br>\nWe can use this new channel as the only input space or as an add on to the 3 or 6 classical channel input spaces.</p>\n<pre><code>   positive=np.mean( np.array([ image[0], image[2],image[4] ]), axis=0 )\n   negative=np.mean( np.array([ image[1], image[3],image[5] ]), axis=0 )\n   inputSpace = abs(positive - negative)\n</code></pre>\n<p>Another option will be to calculate the difference per consecutive channels. In this way we will not lose information by averaging channels (the useful signal can be in just one channel. if we average with the other 2 ON channels we can lose it's intensity)</p>\n<pre><code>        inputSpace1 = abs(image[0] -image[1])\n        inputSpace2 = abs(image[2] -image[3])\n        inputSpace3 = abs(image[4] -image[5])\n</code></pre>\n<p>For a negative sample, the input space will look like this:<br>\nOriginal:<br>\n<img src=\"https://i.postimg.cc/V6h4tThK/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"><br>\nMean of differences:<br>\n<img src=\"https://i.postimg.cc/QN4VQD08/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"><br>\nDifferences on consecutive channels:<br>\n<img src=\"https://i.postimg.cc/WpBcLf4m/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"></p>\n<p>For a positive sample, the input space will look like this:<br>\nOriginal:<br>\n<img src=\"https://i.postimg.cc/Hk4gt4dw/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"><br>\nMean of differences:<br>\n<img src=\"https://i.postimg.cc/3NvM02tD/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"><br>\nDifferences on consecutive channels:<br>\n<img src=\"https://i.postimg.cc/wBjjybLy/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"></p>\n<p>Is had to say just by looking on the images if a good model will find a better trend compared with the classical 3/6 channels but for sure is a interesting approach to test. <br>\nOther ideea is to use computer vision feature descriptors like histogram of oriented gradients and sift(Scale-invariant feature transform). With HOG we can describe an input space by the histogram of the angles values.The technique counts occurrences of gradient orientation in localized portions of an image. In computer vision it is usually used with SVM for recognizing objects by their specific angles which are defining the shape of the objects.</p>",
  "messages": [
    {
      "id": 1406828,
      "postDate": "2021-08-01T08:26:12.157Z",
      "content": "<p>Hello guys,</p>\n<p>Besides the normal 6 or 3 dimensions aligned and used as a input for the convolutional models there can be created different and more complex methodologies and architectures.<br>\nI want this to be a thread with alternative approaches for input spaces that the community tried so far. It does not matter if the concept leaded to good or bad results, <strong>the idea</strong> is the only thing that matters.<br>\nI will start sharing my current ideas.<br>\nI am in the process of testing two alternatives.<br>\nAs one of the main characteristic of a positive signal is the fact that it is present in the ON channels and absent in the OFF channels, the first alternative approach is to calculate the pixels mean of all 3 ON channels, calculate the pixels mean of all OFF channels, calculate the difference and use it as a input feature., I will use the absolute value of this difference because we are interested on the value of the difference (as the resulted pixel value will grow, it will be a strong sign that is a big difference in that region). Practically it will be one channel representative for the difference.<br>\nWe can use this new channel as the only input space or as an add on to the 3 or 6 classical channel input spaces.</p>\n<pre><code>   positive=np.mean( np.array([ image[0], image[2],image[4] ]), axis=0 )\n   negative=np.mean( np.array([ image[1], image[3],image[5] ]), axis=0 )\n   inputSpace = abs(positive - negative)\n</code></pre>\n<p>Another option will be to calculate the difference per consecutive channels. In this way we will not lose information by averaging channels (the useful signal can be in just one channel. if we average with the other 2 ON channels we can lose it's intensity)</p>\n<pre><code>        inputSpace1 = abs(image[0] -image[1])\n        inputSpace2 = abs(image[2] -image[3])\n        inputSpace3 = abs(image[4] -image[5])\n</code></pre>\n<p>For a negative sample, the input space will look like this:<br>\nOriginal:<br>\n<img src=\"https://i.postimg.cc/V6h4tThK/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"><br>\nMean of differences:<br>\n<img src=\"https://i.postimg.cc/QN4VQD08/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"><br>\nDifferences on consecutive channels:<br>\n<img src=\"https://i.postimg.cc/WpBcLf4m/00a1d9ce5045cd5-npy-0-0.jpg\" alt=\"image\"></p>\n<p>For a positive sample, the input space will look like this:<br>\nOriginal:<br>\n<img src=\"https://i.postimg.cc/Hk4gt4dw/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"><br>\nMean of differences:<br>\n<img src=\"https://i.postimg.cc/3NvM02tD/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"><br>\nDifferences on consecutive channels:<br>\n<img src=\"https://i.postimg.cc/wBjjybLy/00b2a4425260a05-npy-1-0.jpg\" alt=\"image\"></p>\n<p>Is had to say just by looking on the images if a good model will find a better trend compared with the classical 3/6 channels but for sure is a interesting approach to test. <br>\nOther ideea is to use computer vision feature descriptors like histogram of oriented gradients and sift(Scale-invariant feature transform). With HOG we can describe an input space by the histogram of the angles values.The technique counts occurrences of gradient orientation in localized portions of an image. In computer vision it is usually used with SVM for recognizing objects by their specific angles which are defining the shape of the objects.</p>",
      "rawMarkdown": "Hello guys,\n\nBesides the normal 6 or 3 dimensions aligned and used as a input for the convolutional models there can be created different and more complex methodologies and architectures.\nI want this to be a thread with alternative approaches for input spaces that the community tried so far. It does not matter if the concept leaded to good or bad results, **the idea** is the only thing that matters.\nI will start sharing my current ideas.\nI am in the process of testing two alternatives.\nAs one of the main characteristic of a positive signal is the fact that it is present in the ON channels and absent in the OFF channels, the first alternative approach is to calculate the pixels mean of all 3 ON channels, calculate the pixels mean of all OFF channels, calculate the difference and use it as a input feature., I will use the absolute value of this difference because we are interested on the value of the difference (as the resulted pixel value will grow, it will be a strong sign that is a big difference in that region). Practically it will be one channel representative for the difference.\nWe can use this new channel as the only input space or as an add on to the 3 or 6 classical channel input spaces.\n  ```\n   positive=np.mean( np.array([ image[0], image[2],image[4] ]), axis=0 )\n   negative=np.mean( np.array([ image[1], image[3],image[5] ]), axis=0 )\n   inputSpace = abs(positive - negative)\n```\nAnother option will be to calculate the difference per consecutive channels. In this way we will not lose information by averaging channels (the useful signal can be in just one channel. if we average with the other 2 ON channels we can lose it's intensity)\n```\n        inputSpace1 = abs(image[0] -image[1])\n        inputSpace2 = abs(image[2] -image[3])\n        inputSpace3 = abs(image[4] -image[5])\n```\n\n\nFor a negative sample, the input space will look like this:\nOriginal:\n![image](https://i.postimg.cc/V6h4tThK/00a1d9ce5045cd5-npy-0-0.jpg)\nMean of differences:\n![image](https://i.postimg.cc/QN4VQD08/00a1d9ce5045cd5-npy-0-0.jpg)\nDifferences on consecutive channels:\n![image](https://i.postimg.cc/WpBcLf4m/00a1d9ce5045cd5-npy-0-0.jpg)\n\nFor a positive sample, the input space will look like this:\nOriginal:\n![image](https://i.postimg.cc/Hk4gt4dw/00b2a4425260a05-npy-1-0.jpg)\nMean of differences:\n![image](https://i.postimg.cc/3NvM02tD/00b2a4425260a05-npy-1-0.jpg)\nDifferences on consecutive channels:\n![image](https://i.postimg.cc/wBjjybLy/00b2a4425260a05-npy-1-0.jpg)\n\nIs had to say just by looking on the images if a good model will find a better trend compared with the classical 3/6 channels but for sure is a interesting approach to test. \nOther ideea is to use computer vision feature descriptors like histogram of oriented gradients and sift(Scale-invariant feature transform). With HOG we can describe an input space by the histogram of the angles values.The technique counts occurrences of gradient orientation in localized portions of an image. In computer vision it is usually used with SVM for recognizing objects by their specific angles which are defining the shape of the objects.\n\n",
      "votes": 14
    },
    {
      "id": 1406864,
      "postDate": "2021-08-01T09:09:35.133Z",
      "content": "<p>I tried the difference approach without abs. <br>\nI compared the two approaches with minimal augmentations and with fixed backbone and epochs<br>\nAugmentation Used - Flip LR + Mixup<br>\nBackbone EfficientNet B0 <br>\nEpochs 10</p>\n<blockquote>\n  <p>Approach 1 (Concat 0, 2, 4) - CV 0.822 LB 0.726<br>\n  Approach 2 (Concat (1-0),(3-2),(5-4)) - CV 0.798 LB 0.702</p>\n</blockquote>",
      "rawMarkdown": "I tried the difference approach without abs. \nI compared the two approaches with minimal augmentations and with fixed backbone and epochs\nAugmentation Used - Flip LR + Mixup\nBackbone EfficientNet B0 \nEpochs 10\n> Approach 1 (Concat 0, 2, 4) - CV 0.822 LB 0.726\n> Approach 2 (Concat (1-0),(3-2),(5-4)) - CV 0.798 LB 0.702",
      "votes": 3,
      "replies": [
        {
          "id": 1406877,
          "postDate": "2021-08-01T09:19:33.237Z",
          "content": "<p>Good feedback <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nI choose abs with the purpose to make the model recognition pattern simpler, as the difference from OFF to ON will be bigger the value is bigger, no matter in what direction will be that difference.</p>",
          "rawMarkdown": "Good feedback @ks2019 \nI choose abs with the purpose to make the model recognition pattern simpler, as the difference from OFF to ON will be bigger the value is bigger, no matter in what direction will be that difference.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 1409525,
      "postDate": "2021-08-02T17:00:38.817Z",
      "content": "<p>This is an example of why I don't think the proposed strategies will work (on channels only):</p>\n<p><img src=\"https://i.imgur.com/Nxftfhm.png\" alt=\"example\"></p>\n<p>In an ideal world, when the cadence returns back to the [ON] subject of interest, the raw values recorded should look very similar. But in practice, (at least with our procedural generated dataset here) the observed amplitude of the signal can be all over the place. Subtracting a mean across on channels, or off channels, or even between neighboring channels generates a lot of artifacts in these cases.</p>\n<p>Anyhow, I tried doing this sort of thing back before the reset. To get around the above issue, I would store my regular images in one channel, then my 'processed' images in another channel (since we have 3 inputs to the pretrained model). It did not improve the model and caused it to overfit quicker.</p>",
      "rawMarkdown": "This is an example of why I don't think the proposed strategies will work (on channels only):\n\n![example](https://i.imgur.com/Nxftfhm.png)\n\nIn an ideal world, when the cadence returns back to the [ON] subject of interest, the raw values recorded should look very similar. But in practice, (at least with our procedural generated dataset here) the observed amplitude of the signal can be all over the place. Subtracting a mean across on channels, or off channels, or even between neighboring channels generates a lot of artifacts in these cases.\n\nAnyhow, I tried doing this sort of thing back before the reset. To get around the above issue, I would store my regular images in one channel, then my 'processed' images in another channel (since we have 3 inputs to the pretrained model). It did not improve the model and caused it to overfit quicker.",
      "votes": 1
    },
    {
      "id": 1407809,
      "postDate": "2021-08-02T06:34:47.587Z",
      "content": "<p>Something I looked at on the original data before the leak, was doing image align for the pairs ON/OFF( 0-1, 2-3, 4-5) using cv2.findTransformECC. So if it could find align True or False for each pair, and if found the difference in the alignment.  But it takes some time to run all these and prepare the data.  Was looking at each pair separately because the signal can be present in some or all the A (ON) channels. <br>\nWhat might be useful with this idea is if you have OOF, where it does well (or not) in predicting then you could see if image align difference predicted does help.  If that makes sense.   </p>",
      "rawMarkdown": "Something I looked at on the original data before the leak, was doing image align for the pairs ON/OFF( 0-1, 2-3, 4-5) using cv2.findTransformECC. So if it could find align True or False for each pair, and if found the difference in the alignment.  But it takes some time to run all these and prepare the data.  Was looking at each pair separately because the signal can be present in some or all the A (ON) channels. \nWhat might be useful with this idea is if you have OOF, where it does well (or not) in predicting then you could see if image align difference predicted does help.  If that makes sense.   ",
      "votes": 1
    },
    {
      "id": 1449460,
      "postDate": "2021-08-04T21:59:06.303Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1449488,
          "postDate": "2021-08-04T22:08:00.650Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1449718,
          "postDate": "2021-08-05T00:16:58Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1418595,
      "postDate": "2021-08-03T01:36:00.087Z",
      "content": "<p>Thanks for your idea.👍</p>",
      "rawMarkdown": "Thanks for your idea.👍"
    }
  ],
  "comments": [
    {
      "id": 1406864,
      "author_name": "Kumar Shubham",
      "author_url": "",
      "post_date": "2021-08-01T09:09:35.133000",
      "content": "<p>I tried the difference approach without abs. <br>\nI compared the two approaches with minimal augmentations and with fixed backbone and epochs<br>\nAugmentation Used - Flip LR + Mixup<br>\nBackbone EfficientNet B0 <br>\nEpochs 10</p>\n<blockquote>\n  <p>Approach 1 (Concat 0, 2, 4) - CV 0.822 LB 0.726<br>\n  Approach 2 (Concat (1-0),(3-2),(5-4)) - CV 0.798 LB 0.702</p>\n</blockquote>",
      "votes": 3,
      "replies": [
        {
          "id": 1406877,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2021-08-01T09:19:33.237000",
          "content": "<p>Good feedback <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> <br>\nI choose abs with the purpose to make the model recognition pattern simpler, as the difference from OFF to ON will be bigger the value is bigger, no matter in what direction will be that difference.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1409525,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-08-02T17:00:38.817000",
      "content": "<p>This is an example of why I don't think the proposed strategies will work (on channels only):</p>\n<p><img src=\"https://i.imgur.com/Nxftfhm.png\" alt=\"example\"></p>\n<p>In an ideal world, when the cadence returns back to the [ON] subject of interest, the raw values recorded should look very similar. But in practice, (at least with our procedural generated dataset here) the observed amplitude of the signal can be all over the place. Subtracting a mean across on channels, or off channels, or even between neighboring channels generates a lot of artifacts in these cases.</p>\n<p>Anyhow, I tried doing this sort of thing back before the reset. To get around the above issue, I would store my regular images in one channel, then my 'processed' images in another channel (since we have 3 inputs to the pretrained model). It did not improve the model and caused it to overfit quicker.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1407809,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-08-02T06:34:47.587000",
      "content": "<p>Something I looked at on the original data before the leak, was doing image align for the pairs ON/OFF( 0-1, 2-3, 4-5) using cv2.findTransformECC. So if it could find align True or False for each pair, and if found the difference in the alignment.  But it takes some time to run all these and prepare the data.  Was looking at each pair separately because the signal can be present in some or all the A (ON) channels. <br>\nWhat might be useful with this idea is if you have OOF, where it does well (or not) in predicting then you could see if image align difference predicted does help.  If that makes sense.   </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1449460,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-04T21:59:06.303000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1449488,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-04T22:08:00.650000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1449718,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-08-05T00:16:58",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1418595,
      "author_name": "Gainover",
      "author_url": "",
      "post_date": "2021-08-03T01:36:00.087000",
      "content": "<p>Thanks for your idea.👍</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1406828": "Hello guys,\n\nBesides the normal 6 or 3 dimensions aligned and used as a input for the convolutional models there can be created different and more complex methodologies and architectures.\nI want this to be a thread with alternative approaches for input spaces that the community tried so far. It does not matter if the concept leaded to good or bad results, **the idea** is the only thing that matters.\nI will start sharing my current ideas.\nI am in the process of testing two alternatives.\nAs one of the main characteristic of a positive signal is the fact that it is present in the ON channels and absent in the OFF channels, the first alternative approach is to calculate the pixels mean of all 3 ON channels, calculate the pixels mean of all OFF channels, calculate the difference and use it as a input feature., I will use the absolute value of this difference because we are interested on the value of the difference (as the resulted pixel value will grow, it will be a strong sign that is a big difference in that region). Practically it will be one channel representative for the difference.\nWe can use this new channel as the only input space or as an add on to the 3 or 6 classical channel input spaces.\n  ```\n   positive=np.mean( np.array([ image[0], image[2],image[4] ]), axis=0 )\n   negative=np.mean( np.array([ image[1], image[3],image[5] ]), axis=0 )\n   inputSpace = abs(positive - negative)\n```\nAnother option will be to calculate the difference per consecutive channels. In this way we will not lose information by averaging channels (the useful signal can be in just one channel. if we average with the other 2 ON channels we can lose it's intensity)\n```\n        inputSpace1 = abs(image[0] -image[1])\n        inputSpace2 = abs(image[2] -image[3])\n        inputSpace3 = abs(image[4] -image[5])\n```\n\n\nFor a negative sample, the input space will look like this:\nOriginal:\n![image](https://i.postimg.cc/V6h4tThK/00a1d9ce5045cd5-npy-0-0.jpg)\nMean of differences:\n![image](https://i.postimg.cc/QN4VQD08/00a1d9ce5045cd5-npy-0-0.jpg)\nDifferences on consecutive channels:\n![image](https://i.postimg.cc/WpBcLf4m/00a1d9ce5045cd5-npy-0-0.jpg)\n\nFor a positive sample, the input space will look like this:\nOriginal:\n![image](https://i.postimg.cc/Hk4gt4dw/00b2a4425260a05-npy-1-0.jpg)\nMean of differences:\n![image](https://i.postimg.cc/3NvM02tD/00b2a4425260a05-npy-1-0.jpg)\nDifferences on consecutive channels:\n![image](https://i.postimg.cc/wBjjybLy/00b2a4425260a05-npy-1-0.jpg)\n\nIs had to say just by looking on the images if a good model will find a better trend compared with the classical 3/6 channels but for sure is a interesting approach to test. \nOther ideea is to use computer vision feature descriptors like histogram of oriented gradients and sift(Scale-invariant feature transform). With HOG we can describe an input space by the histogram of the angles values.The technique counts occurrences of gradient orientation in localized portions of an image. In computer vision it is usually used with SVM for recognizing objects by their specific angles which are defining the shape of the objects.\n\n",
    "1406864": "I tried the difference approach without abs. \nI compared the two approaches with minimal augmentations and with fixed backbone and epochs\nAugmentation Used - Flip LR + Mixup\nBackbone EfficientNet B0 \nEpochs 10\n> Approach 1 (Concat 0, 2, 4) - CV 0.822 LB 0.726\n> Approach 2 (Concat (1-0),(3-2),(5-4)) - CV 0.798 LB 0.702",
    "1409525": "This is an example of why I don't think the proposed strategies will work (on channels only):\n\n![example](https://i.imgur.com/Nxftfhm.png)\n\nIn an ideal world, when the cadence returns back to the [ON] subject of interest, the raw values recorded should look very similar. But in practice, (at least with our procedural generated dataset here) the observed amplitude of the signal can be all over the place. Subtracting a mean across on channels, or off channels, or even between neighboring channels generates a lot of artifacts in these cases.\n\nAnyhow, I tried doing this sort of thing back before the reset. To get around the above issue, I would store my regular images in one channel, then my 'processed' images in another channel (since we have 3 inputs to the pretrained model). It did not improve the model and caused it to overfit quicker.",
    "1407809": "Something I looked at on the original data before the leak, was doing image align for the pairs ON/OFF( 0-1, 2-3, 4-5) using cv2.findTransformECC. So if it could find align True or False for each pair, and if found the difference in the alignment.  But it takes some time to run all these and prepare the data.  Was looking at each pair separately because the signal can be present in some or all the A (ON) channels. \nWhat might be useful with this idea is if you have OOF, where it does well (or not) in predicting then you could see if image align difference predicted does help.  If that makes sense.   ",
    "1449460": "",
    "1418595": "Thanks for your idea.👍"
  }
}