{
  "id": 245152,
  "title": "[Comparative Study] What's the best data representation and the effect of Mixup?",
  "url": "/competitions/seti-breakthrough-listen/discussion/245152",
  "author_name": "Ayush Thakur",
  "post_date": "2021-06-09T23:16:34.236000",
  "votes": 55,
  "comment_count": 19,
  "views": 0,
  "content": "<h1>Update</h1>\n<p>Here's the kernel to run the experiments yourself: <a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">http://wandb.me/kernel-seti-exp</a> <br>\n(If you find any errors do let me know.)</p>\n<h1>TL;DR</h1>\n<p><strong>For interacting with the results check out the W&amp;B report</strong>: <a href=\"http://wandb.me/seti-img-mixup-exp\" target=\"_blank\">http://wandb.me/seti-img-mixup-exp</a></p>\n<h1>Introduction</h1>\n<p>We are one month in the competition and there are many teams in the 0.97x-0.99x LB score range. I still wanted to run a few experiments of my own to answer two very important questions:</p>\n<ul>\n<li><p><strong>What's the best way to use the cadence snippet?</strong> Should we use it **channel-wise or spatially? **Should we use <strong>all 6 spectrograms or just the ones with aliens' signal?</strong></p></li>\n<li><p><strong>Mixup</strong> is giving a significant performance boost. But what's the <strong>gain in percentage?</strong> <strong>How much are the models trained with Mixup dependent on random initialization?</strong></p></li>\n</ul>\n<h1>Experimentation Setup</h1>\n<ul>\n<li><strong>Framework</strong>: TensorFlow</li>\n<li><strong>Data</strong>: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</li>\n<li><strong>Backbone architecture</strong>: EfficientNetB0</li>\n<li><strong>Experiment Tracking</strong>: I used Weights and Biases for tracking all the experiments. </li>\n<li><strong>Trained for</strong>: Each experiment was run 3 times to get the mean and standard deviation. </li>\n<li><strong>Regularization</strong>: Trained with early stopping with the patience of 5 epochs. </li>\n<li><strong>Other Configs</strong>: Adam optimizer was used. Refer to figure 1 for more config settings. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/xyplZrm.png\" alt=\"img\"></p>\n<h1>Best way to use the Cadence snippet?</h1>\n<ul>\n<li>Channel-wise arrangement of spectrograms means stacking them on-top-of-each other. For 6 spectrograms arranged channel-wise the resulting shape would be <code>(Height, Width, 6)</code>.</li>\n<li>Spatial arrangement of spectrograms means stacking them side-by-side. For 6 spectrograms arranged spatially, the resulting shape would be <code>(Height, Width*6, 1)</code>.</li>\n</ul>\n<h3>Arrange all 6 spectrograms <strong>channel-wise vs spatially</strong></h3>\n<p>Note: The 6 channel image was reduced to 3 channels by using a <code>Conv2D</code> layer. This was then fed to the backbone model. </p>\n<ul>\n<li>Spatial arrangement gives a gain of ~6% for the validation ROC-AUC metric.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/MH7kSK9.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h3>Arrange <strong>only target</strong> spectrograms <strong>channel-wise vs spatially</strong></h3>\n<p>Note: Target spectrograms contain alien signals. </p>\n<ul>\n<li>We can clearly see a significant improvement in the channel-wise arrangement when <strong>only target spectrograms are used</strong>. </li>\n<li>There is about ~1% improvement in the spatial arrangement when only target spectrograms are used. </li>\n<li><strong>However, note that the standard deviation is much higher with just target spectrograms arranged spatially.</strong></li>\n</ul>\n<p><img src=\"https://i.imgur.com/US8yvhd.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h3>Normalize individual spectrograms vs clip pixels and then normalize</h3>\n<p>Note: The code snippet below shows the difference between image-level normalization vs clip and then normalize</p>\n<pre><code># Normalize\ndata = ((data - np.mean(data, axis=0)) / np.std(data, axis=0))\n</code></pre>\n<pre><code># Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n</code></pre>\n<ul>\n<li>Clearly the standard deviation reduced.</li>\n<li>There is ~1% improvement in the score. </li>\n<li><strong>So far, the target-only spectrograms arranged spatially with clip and then normalize gave the best mean ROC-AUC score with the least standard deviation in the score.</strong> </li>\n</ul>\n<p><img src=\"https://i.imgur.com/4wxLybP.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h1>Mixup</h1>\n<p>Mixup augmentation is used by almost every team in this competition. And the reason will be obvious from the results of my experiments. By the way, if you want to learn more about Mixup (Cutmix, Augmix, etc) here's a <a href=\"http://wandb.me/ayut-augmentation\" target=\"_blank\">blog post that I have written</a>. </p>\n<ul>\n<li><strong>There's a huge ~7% gain in the score</strong> compared to the best score in the previous section. </li>\n<li>The standard deviation is also better compared to the previous best.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/bVBVAbl.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h1>Final Thoughts</h1>\n<ul>\n<li>Use spatial arrangement of spectrograms to get the best single-model or k-fold models.</li>\n<li>Use Mixup augmentation. </li>\n</ul>\n<p>In the next post, I will share the effect of <code>alpha</code>, which is a Mixup-based hyperparameter. I will also share the dependence of the number of trainable parameters on the ROC score.</p>",
  "messages": [
    {
      "id": 1342996,
      "postDate": "2021-06-09T23:16:34.237Z",
      "content": "<h1>Update</h1>\n<p>Here's the kernel to run the experiments yourself: <a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">http://wandb.me/kernel-seti-exp</a> <br>\n(If you find any errors do let me know.)</p>\n<h1>TL;DR</h1>\n<p><strong>For interacting with the results check out the W&amp;B report</strong>: <a href=\"http://wandb.me/seti-img-mixup-exp\" target=\"_blank\">http://wandb.me/seti-img-mixup-exp</a></p>\n<h1>Introduction</h1>\n<p>We are one month in the competition and there are many teams in the 0.97x-0.99x LB score range. I still wanted to run a few experiments of my own to answer two very important questions:</p>\n<ul>\n<li><p><strong>What's the best way to use the cadence snippet?</strong> Should we use it **channel-wise or spatially? **Should we use <strong>all 6 spectrograms or just the ones with aliens' signal?</strong></p></li>\n<li><p><strong>Mixup</strong> is giving a significant performance boost. But what's the <strong>gain in percentage?</strong> <strong>How much are the models trained with Mixup dependent on random initialization?</strong></p></li>\n</ul>\n<h1>Experimentation Setup</h1>\n<ul>\n<li><strong>Framework</strong>: TensorFlow</li>\n<li><strong>Data</strong>: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.</li>\n<li><strong>Backbone architecture</strong>: EfficientNetB0</li>\n<li><strong>Experiment Tracking</strong>: I used Weights and Biases for tracking all the experiments. </li>\n<li><strong>Trained for</strong>: Each experiment was run 3 times to get the mean and standard deviation. </li>\n<li><strong>Regularization</strong>: Trained with early stopping with the patience of 5 epochs. </li>\n<li><strong>Other Configs</strong>: Adam optimizer was used. Refer to figure 1 for more config settings. </li>\n</ul>\n<p><img src=\"https://i.imgur.com/xyplZrm.png\" alt=\"img\"></p>\n<h1>Best way to use the Cadence snippet?</h1>\n<ul>\n<li>Channel-wise arrangement of spectrograms means stacking them on-top-of-each other. For 6 spectrograms arranged channel-wise the resulting shape would be <code>(Height, Width, 6)</code>.</li>\n<li>Spatial arrangement of spectrograms means stacking them side-by-side. For 6 spectrograms arranged spatially, the resulting shape would be <code>(Height, Width*6, 1)</code>.</li>\n</ul>\n<h3>Arrange all 6 spectrograms <strong>channel-wise vs spatially</strong></h3>\n<p>Note: The 6 channel image was reduced to 3 channels by using a <code>Conv2D</code> layer. This was then fed to the backbone model. </p>\n<ul>\n<li>Spatial arrangement gives a gain of ~6% for the validation ROC-AUC metric.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/MH7kSK9.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h3>Arrange <strong>only target</strong> spectrograms <strong>channel-wise vs spatially</strong></h3>\n<p>Note: Target spectrograms contain alien signals. </p>\n<ul>\n<li>We can clearly see a significant improvement in the channel-wise arrangement when <strong>only target spectrograms are used</strong>. </li>\n<li>There is about ~1% improvement in the spatial arrangement when only target spectrograms are used. </li>\n<li><strong>However, note that the standard deviation is much higher with just target spectrograms arranged spatially.</strong></li>\n</ul>\n<p><img src=\"https://i.imgur.com/US8yvhd.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h3>Normalize individual spectrograms vs clip pixels and then normalize</h3>\n<p>Note: The code snippet below shows the difference between image-level normalization vs clip and then normalize</p>\n<pre><code># Normalize\ndata = ((data - np.mean(data, axis=0)) / np.std(data, axis=0))\n</code></pre>\n<pre><code># Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n</code></pre>\n<ul>\n<li>Clearly the standard deviation reduced.</li>\n<li>There is ~1% improvement in the score. </li>\n<li><strong>So far, the target-only spectrograms arranged spatially with clip and then normalize gave the best mean ROC-AUC score with the least standard deviation in the score.</strong> </li>\n</ul>\n<p><img src=\"https://i.imgur.com/4wxLybP.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h1>Mixup</h1>\n<p>Mixup augmentation is used by almost every team in this competition. And the reason will be obvious from the results of my experiments. By the way, if you want to learn more about Mixup (Cutmix, Augmix, etc) here's a <a href=\"http://wandb.me/ayut-augmentation\" target=\"_blank\">blog post that I have written</a>. </p>\n<ul>\n<li><strong>There's a huge ~7% gain in the score</strong> compared to the best score in the previous section. </li>\n<li>The standard deviation is also better compared to the previous best.</li>\n</ul>\n<p><img src=\"https://i.imgur.com/bVBVAbl.png\" alt=\"img\"><br>\n(<a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">Source</a>)</p>\n<h1>Final Thoughts</h1>\n<ul>\n<li>Use spatial arrangement of spectrograms to get the best single-model or k-fold models.</li>\n<li>Use Mixup augmentation. </li>\n</ul>\n<p>In the next post, I will share the effect of <code>alpha</code>, which is a Mixup-based hyperparameter. I will also share the dependence of the number of trainable parameters on the ROC score.</p>",
      "rawMarkdown": "# Update\n\nHere's the kernel to run the experiments yourself: http://wandb.me/kernel-seti-exp \n(If you find any errors do let me know.)\n\n# TL;DR\n\n**For interacting with the results check out the W&B report**: http://wandb.me/seti-img-mixup-exp\n\n# Introduction\n\nWe are one month in the competition and there are many teams in the 0.97x-0.99x LB score range. I still wanted to run a few experiments of my own to answer two very important questions:\n\n* **What's the best way to use the cadence snippet?** Should we use it **channel-wise or spatially? **Should we use **all 6 spectrograms or just the ones with aliens' signal?**\n\n* **Mixup** is giving a significant performance boost. But what's the **gain in percentage?** **How much are the models trained with Mixup dependent on random initialization?**\n\n# Experimentation Setup\n\n* **Framework**: TensorFlow\n* **Data**: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.\n* **Backbone architecture**: EfficientNetB0\n* **Experiment Tracking**: I used Weights and Biases for tracking all the experiments. \n* **Trained for**: Each experiment was run 3 times to get the mean and standard deviation. \n* **Regularization**: Trained with early stopping with the patience of 5 epochs. \n* **Other Configs**: Adam optimizer was used. Refer to figure 1 for more config settings. \n\n![img](https://i.imgur.com/xyplZrm.png)\n\n# Best way to use the Cadence snippet?\n\n* Channel-wise arrangement of spectrograms means stacking them on-top-of-each other. For 6 spectrograms arranged channel-wise the resulting shape would be `(Height, Width, 6)`.\n*  Spatial arrangement of spectrograms means stacking them side-by-side. For 6 spectrograms arranged spatially, the resulting shape would be `(Height, Width*6, 1)`.\n\n### Arrange all 6 spectrograms **channel-wise vs spatially** \n\nNote: The 6 channel image was reduced to 3 channels by using a `Conv2D` layer. This was then fed to the backbone model. \n\n* Spatial arrangement gives a gain of ~6% for the validation ROC-AUC metric.\n\n![img](https://i.imgur.com/MH7kSK9.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n### Arrange **only target** spectrograms **channel-wise vs spatially** \n\nNote: Target spectrograms contain alien signals. \n\n* We can clearly see a significant improvement in the channel-wise arrangement when **only target spectrograms are used**. \n* There is about ~1% improvement in the spatial arrangement when only target spectrograms are used. \n* **However, note that the standard deviation is much higher with just target spectrograms arranged spatially.**\n\n![img](https://i.imgur.com/US8yvhd.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n### Normalize individual spectrograms vs clip pixels and then normalize\n\nNote: The code snippet below shows the difference between image-level normalization vs clip and then normalize\n\n```\n# Normalize\ndata = ((data - np.mean(data, axis=0)) / np.std(data, axis=0))\n```\n```\n# Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n```\n\n* Clearly the standard deviation reduced.\n* There is ~1% improvement in the score. \n* **So far, the target-only spectrograms arranged spatially with clip and then normalize gave the best mean ROC-AUC score with the least standard deviation in the score.** \n\n![img](https://i.imgur.com/4wxLybP.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n# Mixup \n\nMixup augmentation is used by almost every team in this competition. And the reason will be obvious from the results of my experiments. By the way, if you want to learn more about Mixup (Cutmix, Augmix, etc) here's a [blog post that I have written](http://wandb.me/ayut-augmentation). \n\n* **There's a huge ~7% gain in the score** compared to the best score in the previous section. \n* The standard deviation is also better compared to the previous best.\n\n![img](https://i.imgur.com/bVBVAbl.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n# Final Thoughts\n\n* Use spatial arrangement of spectrograms to get the best single-model or k-fold models.\n* Use Mixup augmentation. \n\nIn the next post, I will share the effect of `alpha`, which is a Mixup-based hyperparameter. I will also share the dependence of the number of trainable parameters on the ROC score.\n",
      "votes": 55
    },
    {
      "id": 1348641,
      "postDate": "2021-06-14T06:56:32.290Z",
      "content": "<p>Great work, thanks for the insights! In your experience, how do experiment results on a subsample translate to the full dataset in computer vision? I guess early stopping is a must here. </p>",
      "rawMarkdown": "Great work, thanks for the insights! In your experience, how do experiment results on a subsample translate to the full dataset in computer vision? I guess early stopping is a must here. ",
      "votes": 1,
      "replies": [
        {
          "id": 1348947,
          "postDate": "2021-06-14T12:03:38.650Z",
          "content": "<p>That's a great question <a href=\"https://www.kaggle.com/leventelippenszky\" target=\"_blank\">@leventelippenszky</a> :) </p>\n<p>By doing experiments on a subsample (this subsample should make sense):</p>\n<ul>\n<li>We can get away with useful insights quicker than training on the entire dataset.</li>\n<li>We can save a lot of computing costs. </li>\n<li>We can iterate over ideas quickly. </li>\n</ul>\n<p>The results may not translate linearly since a lot of factors affect an experiment but in my experience, an incremental improvement usually translated well to a full dataset. But it's super important that the subsample should be carefully created.</p>\n<p>In this experiment, since we have over 50k images randomly sampling 5k images will not shift the distribution of the ground truth labels by a lot. Suppose I manually curate a subsample with easy to classify images, surely results on such a subsample should not confidently translate well to the full dataset. </p>",
          "rawMarkdown": "That's a great question @leventelippenszky :) \n\nBy doing experiments on a subsample (this subsample should make sense):\n* We can get away with useful insights quicker than training on the entire dataset.\n* We can save a lot of computing costs. \n* We can iterate over ideas quickly. \n\nThe results may not translate linearly since a lot of factors affect an experiment but in my experience, an incremental improvement usually translated well to a full dataset. But it's super important that the subsample should be carefully created.\n\nIn this experiment, since we have over 50k images randomly sampling 5k images will not shift the distribution of the ground truth labels by a lot. Suppose I manually curate a subsample with easy to classify images, surely results on such a subsample should not confidently translate well to the full dataset. \n",
          "votes": 2
        },
        {
          "id": 1349056,
          "postDate": "2021-06-14T13:43:20.733Z",
          "content": "<p>Thanks for the detailed answer:)</p>",
          "rawMarkdown": "Thanks for the detailed answer:)"
        }
      ]
    },
    {
      "id": 1343477,
      "postDate": "2021-06-10T08:22:55.563Z",
      "content": "<p>Very nice work! 👍<br>\nIf you have time to wrap up your insights as a notebook, I would be pretty interested to study that </p>",
      "rawMarkdown": "Very nice work! 👍\nIf you have time to wrap up your insights as a notebook, I would be pretty interested to study that ",
      "votes": 1,
      "replies": [
        {
          "id": 1343629,
          "postDate": "2021-06-10T10:36:55.433Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a>, thanks. I have updated the post with the link to the kernel. </p>",
          "rawMarkdown": "Hey @alexlwh, thanks. I have updated the post with the link to the kernel. ",
          "votes": 1
        },
        {
          "id": 1343819,
          "postDate": "2021-06-10T13:15:12.693Z",
          "content": "<p>awesome! 👍</p>",
          "rawMarkdown": "awesome! 👍"
        }
      ]
    },
    {
      "id": 1343015,
      "postDate": "2021-06-10T01:04:44.790Z",
      "content": "<p>Great post with extreme clarity. I do recommend you also share this in a notebook if possible! I myself am interested.</p>\n<p>In particular, the experiments you did are very insightful, for example, the code to get the standard deviation etc.</p>",
      "rawMarkdown": "Great post with extreme clarity. I do recommend you also share this in a notebook if possible! I myself am interested.\n\nIn particular, the experiments you did are very insightful, for example, the code to get the standard deviation etc.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1343628,
          "postDate": "2021-06-10T10:36:03.133Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>, I am glad you liked it. I have updated the post with the link to the <a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">kernel</a>. </p>\n<p>Also just to bring this to your notice, I didn't code to get the standard deviation or any other plot. I used Weights and Biases to track my experiments and it created the charts automatically. :)</p>",
          "rawMarkdown": "Hey @reighns, I am glad you liked it. I have updated the post with the link to the [kernel](http://wandb.me/kernel-seti-exp). \n\nAlso just to bring this to your notice, I didn't code to get the standard deviation or any other plot. I used Weights and Biases to track my experiments and it created the charts automatically. :)",
          "votes": 2
        },
        {
          "id": 1343660,
          "postDate": "2021-06-10T11:01:56.087Z",
          "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> Oh wow, I should give wb a try then, I have been using Neptune, and I don't know how to use it effectively.</p>",
          "rawMarkdown": "@ayuraj Oh wow, I should give wb a try then, I have been using Neptune, and I don't know how to use it effectively.",
          "votes": 1
        },
        {
          "id": 1343786,
          "postDate": "2021-06-10T12:44:58.917Z",
          "content": "<p>You should give W&amp;B a try. Also, feel free to reach out to me if you need any help regarding this. :)</p>",
          "rawMarkdown": "You should give W&B a try. Also, feel free to reach out to me if you need any help regarding this. :)"
        }
      ]
    },
    {
      "id": 1350935,
      "postDate": "2021-06-16T00:02:48.917Z",
      "content": "<p>thx Ayush . good work!!!<br>\nI just dont understand why the convert_image should do any normalization?</p>\n<h1>Clip</h1>\n<p>data = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)</p>\n<h1>Normalize</h1>\n<p>data = tf.image.convert_image_dtype(data, tf.float32)</p>\n<p>can you pls explain?<br>\nthx<br>\nRoman</p>",
      "rawMarkdown": "thx Ayush . good work!!!\nI just dont understand why the convert_image should do any normalization?\n\n# Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n\ncan you pls explain?\nthx\nRoman",
      "replies": [
        {
          "id": 1350940,
          "postDate": "2021-06-16T00:26:51.763Z",
          "content": "<p>From the <code>tf.image.convert_image_dtype</code> <a href=\"https://www.tensorflow.org/api_docs/python/tf/image/convert_image_dtype\" target=\"_blank\">documentation page</a>,</p>\n<blockquote>\n  <p>Convert image to dtype, scaling its values if needed.</p>\n</blockquote>\n<p>You see we tend to use the word normalization and scaling loosely in the context of the image. Basically this function scale down the pixel value from <code>[0-255]</code> to <code>[0-1]</code>. </p>\n<p>Note: At times we use the word scaling to resize the image and at times to scale the pixels. We at times use the word normalize to mean that the image pixel mean is 0 with std 1, while at times we use this term to mean that we are downscaling the pixels. </p>",
          "rawMarkdown": "From the `tf.image.convert_image_dtype` [documentation page](https://www.tensorflow.org/api_docs/python/tf/image/convert_image_dtype),\n\n> Convert image to dtype, scaling its values if needed.\n\nYou see we tend to use the word normalization and scaling loosely in the context of the image. Basically this function scale down the pixel value from `[0-255]` to `[0-1]`. \n\nNote: At times we use the word scaling to resize the image and at times to scale the pixels. We at times use the word normalize to mean that the image pixel mean is 0 with std 1, while at times we use this term to mean that we are downscaling the pixels. "
        }
      ]
    },
    {
      "id": 1349019,
      "postDate": "2021-06-14T13:05:38.567Z",
      "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> thanks very much for the well-written and informative blog post. As a non-specialist in this field, I learned plenty from it - I may even start to understand more of the jargon used on Kaggle. </p>",
      "rawMarkdown": "@ayuraj thanks very much for the well-written and informative blog post. As a non-specialist in this field, I learned plenty from it - I may even start to understand more of the jargon used on Kaggle. ",
      "replies": [
        {
          "id": 1350941,
          "postDate": "2021-06-16T00:27:05.307Z",
          "content": "<p>Glad it helped. :)</p>",
          "rawMarkdown": "Glad it helped. :)"
        }
      ]
    },
    {
      "id": 1347948,
      "postDate": "2021-06-13T15:27:17.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/ayusdas\" target=\"_blank\">@ayusdas</a> could you share how you submitted test results with this notebooks ? i am struggling with it </p>",
      "rawMarkdown": "@ayusdas could you share how you submitted test results with this notebooks ? i am struggling with it ",
      "replies": [
        {
          "id": 1348489,
          "postDate": "2021-06-14T05:13:35.350Z",
          "content": "<p>I didn't get your question <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>? What do you refer to as test results here?</p>",
          "rawMarkdown": "I didn't get your question @mithilsalunkhe? What do you refer to as test results here?"
        },
        {
          "id": 1348495,
          "postDate": "2021-06-14T05:20:43.067Z",
          "content": "<p><a href=\"https://www.kaggle.com/ayusdas\" target=\"_blank\">@ayusdas</a> I am asking if you could share your test pipeline.</p>",
          "rawMarkdown": "@ayusdas I am asking if you could share your test pipeline."
        }
      ]
    },
    {
      "id": 1343102,
      "postDate": "2021-06-10T03:16:05.723Z",
      "content": "<p>Thank you for showing interesting results.<br>\nIt is very interesting to see the results of channel-wise target is comparable to those of spatials with less std. I suspect that the conv2d layer placed before backbone is doing something bad.</p>",
      "rawMarkdown": "Thank you for showing interesting results.\nIt is very interesting to see the results of channel-wise target is comparable to those of spatials with less std. I suspect that the conv2d layer placed before backbone is doing something bad.",
      "replies": [
        {
          "id": 1343626,
          "postDate": "2021-06-10T10:33:52.297Z",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>. Maybe that conv2d layer is doing something bad but the reason I am not very confident with this assertion is that, for spatial arrangements where the shape of the input tensor is <code>(Height, Width, 1)</code>, I am still using a conv2d layer to expand the channels going into efficientnet to 3. </p>",
          "rawMarkdown": "Thanks, @tomooinubushi. Maybe that conv2d layer is doing something bad but the reason I am not very confident with this assertion is that, for spatial arrangements where the shape of the input tensor is `(Height, Width, 1)`, I am still using a conv2d layer to expand the channels going into efficientnet to 3. ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1348641,
      "author_name": "_lev_lipinski",
      "author_url": "",
      "post_date": "2021-06-14T06:56:32.290000",
      "content": "<p>Great work, thanks for the insights! In your experience, how do experiment results on a subsample translate to the full dataset in computer vision? I guess early stopping is a must here. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1348947,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-14T12:03:38.650000",
          "content": "<p>That's a great question <a href=\"https://www.kaggle.com/leventelippenszky\" target=\"_blank\">@leventelippenszky</a> :) </p>\n<p>By doing experiments on a subsample (this subsample should make sense):</p>\n<ul>\n<li>We can get away with useful insights quicker than training on the entire dataset.</li>\n<li>We can save a lot of computing costs. </li>\n<li>We can iterate over ideas quickly. </li>\n</ul>\n<p>The results may not translate linearly since a lot of factors affect an experiment but in my experience, an incremental improvement usually translated well to a full dataset. But it's super important that the subsample should be carefully created.</p>\n<p>In this experiment, since we have over 50k images randomly sampling 5k images will not shift the distribution of the ground truth labels by a lot. Suppose I manually curate a subsample with easy to classify images, surely results on such a subsample should not confidently translate well to the full dataset. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1349056,
          "author_name": "_lev_lipinski",
          "author_url": "",
          "post_date": "2021-06-14T13:43:20.733000",
          "content": "<p>Thanks for the detailed answer:)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1343477,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-06-10T08:22:55.563000",
      "content": "<p>Very nice work! 👍<br>\nIf you have time to wrap up your insights as a notebook, I would be pretty interested to study that </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1343629,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-10T10:36:55.433000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/alexlwh\" target=\"_blank\">@alexlwh</a>, thanks. I have updated the post with the link to the kernel. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1343819,
          "author_name": "Alex Lau",
          "author_url": "",
          "post_date": "2021-06-10T13:15:12.693000",
          "content": "<p>awesome! 👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1343015,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-06-10T01:04:44.790000",
      "content": "<p>Great post with extreme clarity. I do recommend you also share this in a notebook if possible! I myself am interested.</p>\n<p>In particular, the experiments you did are very insightful, for example, the code to get the standard deviation etc.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1343628,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-10T10:36:03.133000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a>, I am glad you liked it. I have updated the post with the link to the <a href=\"http://wandb.me/kernel-seti-exp\" target=\"_blank\">kernel</a>. </p>\n<p>Also just to bring this to your notice, I didn't code to get the standard deviation or any other plot. I used Weights and Biases to track my experiments and it created the charts automatically. :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1343660,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-10T11:01:56.087000",
          "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> Oh wow, I should give wb a try then, I have been using Neptune, and I don't know how to use it effectively.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1343786,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-10T12:44:58.917000",
          "content": "<p>You should give W&amp;B a try. Also, feel free to reach out to me if you need any help regarding this. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1350935,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2021-06-16T00:02:48.917000",
      "content": "<p>thx Ayush . good work!!!<br>\nI just dont understand why the convert_image should do any normalization?</p>\n<h1>Clip</h1>\n<p>data = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)</p>\n<h1>Normalize</h1>\n<p>data = tf.image.convert_image_dtype(data, tf.float32)</p>\n<p>can you pls explain?<br>\nthx<br>\nRoman</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1350940,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-16T00:26:51.763000",
          "content": "<p>From the <code>tf.image.convert_image_dtype</code> <a href=\"https://www.tensorflow.org/api_docs/python/tf/image/convert_image_dtype\" target=\"_blank\">documentation page</a>,</p>\n<blockquote>\n  <p>Convert image to dtype, scaling its values if needed.</p>\n</blockquote>\n<p>You see we tend to use the word normalization and scaling loosely in the context of the image. Basically this function scale down the pixel value from <code>[0-255]</code> to <code>[0-1]</code>. </p>\n<p>Note: At times we use the word scaling to resize the image and at times to scale the pixels. We at times use the word normalize to mean that the image pixel mean is 0 with std 1, while at times we use this term to mean that we are downscaling the pixels. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1349019,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-06-14T13:05:38.567000",
      "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> thanks very much for the well-written and informative blog post. As a non-specialist in this field, I learned plenty from it - I may even start to understand more of the jargon used on Kaggle. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1350941,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-16T00:27:05.307000",
          "content": "<p>Glad it helped. :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1347948,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-13T15:27:17.847000",
      "content": "<p><a href=\"https://www.kaggle.com/ayusdas\" target=\"_blank\">@ayusdas</a> could you share how you submitted test results with this notebooks ? i am struggling with it </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1348489,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-14T05:13:35.350000",
          "content": "<p>I didn't get your question <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a>? What do you refer to as test results here?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1348495,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-14T05:20:43.067000",
          "content": "<p><a href=\"https://www.kaggle.com/ayusdas\" target=\"_blank\">@ayusdas</a> I am asking if you could share your test pipeline.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1343102,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-06-10T03:16:05.723000",
      "content": "<p>Thank you for showing interesting results.<br>\nIt is very interesting to see the results of channel-wise target is comparable to those of spatials with less std. I suspect that the conv2d layer placed before backbone is doing something bad.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1343626,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-06-10T10:33:52.297000",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>. Maybe that conv2d layer is doing something bad but the reason I am not very confident with this assertion is that, for spatial arrangements where the shape of the input tensor is <code>(Height, Width, 1)</code>, I am still using a conv2d layer to expand the channels going into efficientnet to 3. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1342996": "# Update\n\nHere's the kernel to run the experiments yourself: http://wandb.me/kernel-seti-exp \n(If you find any errors do let me know.)\n\n# TL;DR\n\n**For interacting with the results check out the W&B report**: http://wandb.me/seti-img-mixup-exp\n\n# Introduction\n\nWe are one month in the competition and there are many teams in the 0.97x-0.99x LB score range. I still wanted to run a few experiments of my own to answer two very important questions:\n\n* **What's the best way to use the cadence snippet?** Should we use it **channel-wise or spatially? **Should we use **all 6 spectrograms or just the ones with aliens' signal?**\n\n* **Mixup** is giving a significant performance boost. But what's the **gain in percentage?** **How much are the models trained with Mixup dependent on random initialization?**\n\n# Experimentation Setup\n\n* **Framework**: TensorFlow\n* **Data**: 5000 examples randomly shuffled. The resulting data distribution is close to the full training data distribution.\n* **Backbone architecture**: EfficientNetB0\n* **Experiment Tracking**: I used Weights and Biases for tracking all the experiments. \n* **Trained for**: Each experiment was run 3 times to get the mean and standard deviation. \n* **Regularization**: Trained with early stopping with the patience of 5 epochs. \n* **Other Configs**: Adam optimizer was used. Refer to figure 1 for more config settings. \n\n![img](https://i.imgur.com/xyplZrm.png)\n\n# Best way to use the Cadence snippet?\n\n* Channel-wise arrangement of spectrograms means stacking them on-top-of-each other. For 6 spectrograms arranged channel-wise the resulting shape would be `(Height, Width, 6)`.\n*  Spatial arrangement of spectrograms means stacking them side-by-side. For 6 spectrograms arranged spatially, the resulting shape would be `(Height, Width*6, 1)`.\n\n### Arrange all 6 spectrograms **channel-wise vs spatially** \n\nNote: The 6 channel image was reduced to 3 channels by using a `Conv2D` layer. This was then fed to the backbone model. \n\n* Spatial arrangement gives a gain of ~6% for the validation ROC-AUC metric.\n\n![img](https://i.imgur.com/MH7kSK9.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n### Arrange **only target** spectrograms **channel-wise vs spatially** \n\nNote: Target spectrograms contain alien signals. \n\n* We can clearly see a significant improvement in the channel-wise arrangement when **only target spectrograms are used**. \n* There is about ~1% improvement in the spatial arrangement when only target spectrograms are used. \n* **However, note that the standard deviation is much higher with just target spectrograms arranged spatially.**\n\n![img](https://i.imgur.com/US8yvhd.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n### Normalize individual spectrograms vs clip pixels and then normalize\n\nNote: The code snippet below shows the difference between image-level normalization vs clip and then normalize\n\n```\n# Normalize\ndata = ((data - np.mean(data, axis=0)) / np.std(data, axis=0))\n```\n```\n# Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n```\n\n* Clearly the standard deviation reduced.\n* There is ~1% improvement in the score. \n* **So far, the target-only spectrograms arranged spatially with clip and then normalize gave the best mean ROC-AUC score with the least standard deviation in the score.** \n\n![img](https://i.imgur.com/4wxLybP.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n# Mixup \n\nMixup augmentation is used by almost every team in this competition. And the reason will be obvious from the results of my experiments. By the way, if you want to learn more about Mixup (Cutmix, Augmix, etc) here's a [blog post that I have written](http://wandb.me/ayut-augmentation). \n\n* **There's a huge ~7% gain in the score** compared to the best score in the previous section. \n* The standard deviation is also better compared to the previous best.\n\n![img](https://i.imgur.com/bVBVAbl.png)\n([Source](http://wandb.me/kernel-seti-exp))\n\n# Final Thoughts\n\n* Use spatial arrangement of spectrograms to get the best single-model or k-fold models.\n* Use Mixup augmentation. \n\nIn the next post, I will share the effect of `alpha`, which is a Mixup-based hyperparameter. I will also share the dependence of the number of trainable parameters on the ROC score.\n",
    "1348641": "Great work, thanks for the insights! In your experience, how do experiment results on a subsample translate to the full dataset in computer vision? I guess early stopping is a must here. ",
    "1343477": "Very nice work! 👍\nIf you have time to wrap up your insights as a notebook, I would be pretty interested to study that ",
    "1343015": "Great post with extreme clarity. I do recommend you also share this in a notebook if possible! I myself am interested.\n\nIn particular, the experiments you did are very insightful, for example, the code to get the standard deviation etc.\n\n",
    "1350935": "thx Ayush . good work!!!\nI just dont understand why the convert_image should do any normalization?\n\n# Clip \ndata = ((np.clip(data, -1, 3) + 1) / 4 * 255).astype(np.uint8)\n# Normalize\ndata = tf.image.convert_image_dtype(data, tf.float32)\n\ncan you pls explain?\nthx\nRoman",
    "1349019": "@ayuraj thanks very much for the well-written and informative blog post. As a non-specialist in this field, I learned plenty from it - I may even start to understand more of the jargon used on Kaggle. ",
    "1347948": "@ayusdas could you share how you submitted test results with this notebooks ? i am struggling with it ",
    "1343102": "Thank you for showing interesting results.\nIt is very interesting to see the results of channel-wise target is comparable to those of spatials with less std. I suspect that the conv2d layer placed before backbone is doing something bad."
  }
}