{
  "id": 241195,
  "title": "Image Size vs Score",
  "url": "/competitions/seti-breakthrough-listen/discussion/241195",
  "author_name": "Tawara",
  "post_date": "2021-05-23T14:00:35.917000",
  "votes": 75,
  "comment_count": 36,
  "views": 0,
  "content": "<p>Here I compare the models trained under the same experimental settings, except for the input image size.</p>\n<h4>settings</h4>\n<ul>\n<li>Model: resnet18d</li>\n<li>Initial Weights: ImageNet</li>\n<li>Train-Val Split: StratifiedKFold(K=5)</li>\n<li>Data Augmentation: HorizontalFlip, VerticalFlip, ShiftScaleRotate, RandomResizedCrop</li>\n<li>Loss: Binary Cross Entropy</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n</ul>\n<h4>results</h4>\n<table>\n<thead>\n<tr>\n<th>Input Image Size(Frequency x Time)</th>\n<th>OOF Loss</th>\n<th>OOF Score</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>256x256</td>\n<td>0.05147</td>\n<td>0.9818</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>384x384</td>\n<td>0.04485</td>\n<td>0.9847</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>512x512</td>\n<td>0.04180</td>\n<td>0.9868</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>640x640</td>\n<td>0.04188</td>\n<td>0.9863</td>\n<td><strong>0.97</strong></td>\n</tr>\n<tr>\n<td>768x768(new)</td>\n<td><strong>0.04038</strong></td>\n<td><strong>0.9868</strong></td>\n<td>0.97</td>\n</tr>\n</tbody>\n</table>\n<p>The Public score order of them is as follows:  <br>\n<strong>640x640</strong> &gt; 768x768(new) &gt; 512x512 &gt; 384x384 &gt; 256x256</p>\n<p>Maybe, 512x512 is good ?</p>\n<p><strong>UPDATE1</strong>: add 768x768 result to the table above.</p>",
  "messages": [
    {
      "id": 1319803,
      "postDate": "2021-05-23T14:00:35.917Z",
      "content": "<p>Here I compare the models trained under the same experimental settings, except for the input image size.</p>\n<h4>settings</h4>\n<ul>\n<li>Model: resnet18d</li>\n<li>Initial Weights: ImageNet</li>\n<li>Train-Val Split: StratifiedKFold(K=5)</li>\n<li>Data Augmentation: HorizontalFlip, VerticalFlip, ShiftScaleRotate, RandomResizedCrop</li>\n<li>Loss: Binary Cross Entropy</li>\n<li>Optimizer: AdamW</li>\n<li>Scheduler: OneCycleLR</li>\n</ul>\n<h4>results</h4>\n<table>\n<thead>\n<tr>\n<th>Input Image Size(Frequency x Time)</th>\n<th>OOF Loss</th>\n<th>OOF Score</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>256x256</td>\n<td>0.05147</td>\n<td>0.9818</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>384x384</td>\n<td>0.04485</td>\n<td>0.9847</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>512x512</td>\n<td>0.04180</td>\n<td>0.9868</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>640x640</td>\n<td>0.04188</td>\n<td>0.9863</td>\n<td><strong>0.97</strong></td>\n</tr>\n<tr>\n<td>768x768(new)</td>\n<td><strong>0.04038</strong></td>\n<td><strong>0.9868</strong></td>\n<td>0.97</td>\n</tr>\n</tbody>\n</table>\n<p>The Public score order of them is as follows:  <br>\n<strong>640x640</strong> &gt; 768x768(new) &gt; 512x512 &gt; 384x384 &gt; 256x256</p>\n<p>Maybe, 512x512 is good ?</p>\n<p><strong>UPDATE1</strong>: add 768x768 result to the table above.</p>",
      "rawMarkdown": "Here I compare the models trained under the same experimental settings, except for the input image size.\n\n#### settings\n\n* Model: resnet18d\n* Initial Weights: ImageNet\n* Train-Val Split: StratifiedKFold(K=5)\n* Data Augmentation: HorizontalFlip, VerticalFlip, ShiftScaleRotate, RandomResizedCrop\n* Loss: Binary Cross Entropy\n* Optimizer: AdamW\n* Scheduler: OneCycleLR\n\n#### results\n\n|Input Image Size(Frequency x Time) | OOF Loss | OOF Score | Public LB Score |\n|:-----------------------------:|:---:|:----:|\n| 256x256 | 0.05147 | 0.9818 | 0.97 |\n| 384x384 | 0.04485 | 0.9847 | 0.97 |\n| 512x512 | 0.04180 | 0.9868 | 0.97 |\n| 640x640 | 0.04188 | 0.9863 | **0.97** |\n| 768x768(new) | **0.04038** | **0.9868** | 0.97 |\n\nThe Public score order of them is as follows:  \n**640x640** > 768x768(new) > 512x512 > 384x384 > 256x256\n\nMaybe, 512x512 is good ?\n\n**UPDATE1**: add 768x768 result to the table above.",
      "votes": 73
    },
    {
      "id": 1326165,
      "postDate": "2021-05-28T09:15:03.153Z",
      "content": "<p>I did the same experiments for <code>resnet34d</code>.</p>\n<table>\n<thead>\n<tr>\n<th>Input Image Size(Frequency x Time)</th>\n<th>OOF Loss</th>\n<th>OOF Score</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>384x384</td>\n<td>0.04230</td>\n<td>0.9862</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>512x512</td>\n<td>0.03971</td>\n<td>0.9869</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>640x640</td>\n<td>0.04003</td>\n<td>0.9876</td>\n<td><strong>0.97</strong></td>\n</tr>\n<tr>\n<td>768x768</td>\n<td><strong>0.03860</strong></td>\n<td><strong>0.9886</strong></td>\n<td>0.97</td>\n</tr>\n</tbody>\n</table>\n<p>The Public score order of them is as follows:<br>\n<strong>640x640</strong> &gt; 768x768 &gt; 512x512 &gt; 384x384</p>",
      "rawMarkdown": "I did the same experiments for `resnet34d`.\n\n|Input Image Size(Frequency x Time) | OOF Loss | OOF Score | Public LB Score |\n|:-----------------------------:|:---:|:----:|\n| 384x384 | 0.04230 | 0.9862 | 0.97 |\n| 512x512 | 0.03971 | 0.9869 | 0.97 |\n| 640x640 | 0.04003 | 0.9876  | **0.97** |\n| 768x768 | **0.03860** |**0.9886** | 0.97 |\n\nThe Public score order of them is as follows:\n**640x640** > 768x768 > 512x512 > 384x384",
      "votes": 5
    },
    {
      "id": 1323660,
      "postDate": "2021-05-26T11:39:36.040Z",
      "content": "<p>Update: I added 768x768 result.</p>",
      "rawMarkdown": "Update: I added 768x768 result.",
      "votes": 4
    },
    {
      "id": 1323184,
      "postDate": "2021-05-26T04:24:15.240Z",
      "content": "<p>Thanks for sharing.<br>\nBased on your results, I hypothesis that upsampling freq axis more than original (=256) size helps, and it may be true (but it's rough experiments).<br>\n(1, 6 * 273, 256) -&gt; CV=0.97688<br>\n(1, 6 * 273, 256) -&gt; resize (1, 512, 512) -&gt; CV=0.98290<br>\nI will try experiments about sensibility of time (=273) axis.</p>",
      "rawMarkdown": "Thanks for sharing.\nBased on your results, I hypothesis that upsampling freq axis more than original (=256) size helps, and it may be true (but it's rough experiments).\n(1, 6 * 273, 256) -> CV=0.97688\n(1, 6 * 273, 256) -> resize (1, 512, 512) -> CV=0.98290\nI will try experiments about sensibility of time (=273) axis.",
      "votes": 4,
      "replies": [
        {
          "id": 1323215,
          "postDate": "2021-05-26T04:58:19.423Z",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>I hypothesis that upsampling freq axis more than original (=256) size helps</p>\n</blockquote>\n<p>I have  similar thoughts with you, and I tried the cases where frequency axis is bigger than time axis.</p>\n<table>\n<thead>\n<tr>\n<th>Size (Freq x Time)</th>\n<th>val score (fold 0)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512x512 (for comparison)</td>\n<td><strong>0.9860</strong></td>\n</tr>\n<tr>\n<td><strong>512x256</strong></td>\n<td>0.9769</td>\n</tr>\n<tr>\n<td>640x640 (for comparison)</td>\n<td><strong>0.9866</strong></td>\n</tr>\n<tr>\n<td><strong>640x320</strong></td>\n<td>0.9799</td>\n</tr>\n<tr>\n<td>768x768 (for comparison)</td>\n<td><strong>0.9871</strong></td>\n</tr>\n<tr>\n<td><strong>768x384</strong></td>\n<td>0.9805</td>\n</tr>\n</tbody>\n</table>\n<p>This is very rough experiment but square inputs may be better?</p>",
          "rawMarkdown": "Thanks!\n\n> I hypothesis that upsampling freq axis more than original (=256) size helps\n\nI have  similar thoughts with you, and I tried the cases where frequency axis is bigger than time axis.\n\n| Size (Freq x Time) | val score (fold 0) |\n|:-------------------:|:-----------------:|\n| 512x512 (for comparison) | **0.9860** |\n| **512x256** | 0.9769 |\n| 640x640 (for comparison) | **0.9866** |\n| **640x320** | 0.9799 |\n| 768x768 (for comparison) | **0.9871** |\n| **768x384** | 0.9805 |\n\nThis is very rough experiment but square inputs may be better?",
          "votes": 3
        },
        {
          "id": 1323302,
          "postDate": "2021-05-26T06:23:50.090Z",
          "content": "<p>Thanks for sharing your result, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>\n<p>I have the same hypothesis as <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiments, CV-score is following:<br>\n(512, 512) &gt; (6*273, 256) &gt; (224, 224)</p>",
          "rawMarkdown": "Thanks for sharing your result, @yukia18 @ttahara \n\nI have the same hypothesis as @yukia18 \nIn my experiments, CV-score is following:\n(512, 512) > (6*273, 256) > (224, 224)",
          "votes": 2
        },
        {
          "id": 1323646,
          "postDate": "2021-05-26T11:21:12.933Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> and <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiment, <br>\nFor Fold 0, I resized 512x512 and trained EfficientNet-B0 for 25 epochs and validation score was 0.972 but submission was around 0.95.<br>\nI used ScaleShiftRotate, VerticalFlip, HorizontalFlip, Mixup augmentations as well.</p>",
          "rawMarkdown": "Hi @ttahara and @yukia18 \nIn my experiment, \nFor Fold 0, I resized 512x512 and trained EfficientNet-B0 for 25 epochs and validation score was 0.972 but submission was around 0.95.\nI used ScaleShiftRotate, VerticalFlip, HorizontalFlip, Mixup augmentations as well.",
          "votes": 3
        },
        {
          "id": 1324678,
          "postDate": "2021-05-27T06:27:43.157Z",
          "content": "<p>Yes same thing is happening with me i get fold 0 score 0.9702 and fold 1 score 0.9741 and cv from these 2 folds come 0.9723 but lb is only 0.95.<br>\nI used effentB0 with input shape (3,244,244) used vstacking.</p>",
          "rawMarkdown": "Yes same thing is happening with me i get fold 0 score 0.9702 and fold 1 score 0.9741 and cv from these 2 folds come 0.9723 but lb is only 0.95.\nI used effentB0 with input shape (3,244,244) used vstacking.",
          "votes": -1
        },
        {
          "id": 1325779,
          "postDate": "2021-05-28T02:07:54.407Z",
          "content": "<p>Updated:</p>\n<table>\n<thead>\n<tr>\n<th>image size (time x freq)</th>\n<th>CV</th>\n<th>Rank on LB (all 0.97x)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512 x 512</td>\n<td>0.98735</td>\n<td>4</td>\n</tr>\n<tr>\n<td>768 x 768</td>\n<td>0.98943</td>\n<td>2</td>\n</tr>\n<tr>\n<td><strong>1024 x 512</strong></td>\n<td><strong>0.98999</strong></td>\n<td><strong>1</strong></td>\n</tr>\n<tr>\n<td>512 x 1024</td>\n<td>0.98871</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>As I expected, upsampling time axis contributes a lot. On the other hand, expanding freq axis more than 512 still helps. I know that large image size always helps in the almost competition, but I didn't know that it is applicable to upsample image more than original size…</p>",
          "rawMarkdown": "Updated:\n| image size (time x freq)  | CV | Rank on LB (all 0.97x) |\n| --- | --- | --- |\n| 512 x 512 | 0.98735 | 4 |\n| 768 x 768 | 0.98943 | 2 |\n| **1024 x 512** | **0.98999** | **1** |\n| 512 x 1024 | 0.98871 | 3 |\n\nAs I expected, upsampling time axis contributes a lot. On the other hand, expanding freq axis more than 512 still helps. I know that large image size always helps in the almost competition, but I didn't know that it is applicable to upsample image more than original size...",
          "votes": 10
        },
        {
          "id": 1325797,
          "postDate": "2021-05-28T02:38:04.943Z",
          "content": "<p>Number of Epochs?</p>",
          "rawMarkdown": "Number of Epochs?",
          "votes": -1
        },
        {
          "id": 1325806,
          "postDate": "2021-05-28T02:49:32.523Z",
          "content": "<p>10 epochs.</p>",
          "rawMarkdown": "10 epochs.",
          "votes": 1
        },
        {
          "id": 1325856,
          "postDate": "2021-05-28T04:27:16.667Z",
          "content": "<p><a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a><br>\nThank you for sharing. I've not tried yet cases where frequency axis is smaller than time axis.<br>\nNow I have started training that case. I'll report the result later.</p>\n<blockquote>\n  <p>but I didn't know that it is applicable to upsample image more than original size…</p>\n</blockquote>\n<p>IMO, expanding freq axis emphasizes E.T. Signals' frequency change over time. It may make signals easier to be find.</p>",
          "rawMarkdown": "@yukia18\nThank you for sharing. I've not tried yet cases where frequency axis is smaller than time axis.\nNow I have started training that case. I'll report the result later.\n\n> but I didn't know that it is applicable to upsample image more than original size…\n\nIMO, expanding freq axis emphasizes E.T. Signals' frequency change over time. It may make signals easier to be find.",
          "votes": 3
        },
        {
          "id": 1328193,
          "postDate": "2021-05-30T03:08:59.550Z",
          "content": "<p>I tried (Freq,Time) = (512, 1024) and (512, 768). But they got lower CV than 512x512 …🤔</p>",
          "rawMarkdown": "I tried (Freq,Time) = (512, 1024) and (512, 768). But they got lower CV than 512x512 ...🤔",
          "votes": 2
        },
        {
          "id": 1329296,
          "postDate": "2021-05-31T03:28:29.880Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yabea\" target=\"_blank\">@yabea</a> did you use data aug in training phase, other than resizing and normalization?</p>",
          "rawMarkdown": "Hi @yabea did you use data aug in training phase, other than resizing and normalization?"
        },
        {
          "id": 1340467,
          "postDate": "2021-06-08T01:52:16.583Z",
          "content": "<p>RandomAffine and Mixup. I'm sure Mixup helps both CV and LB as others mentioned.<br>\nTranspose and GaussianBlur got worse dramatically in my case.</p>",
          "rawMarkdown": "RandomAffine and Mixup. I'm sure Mixup helps both CV and LB as others mentioned.\nTranspose and GaussianBlur got worse dramatically in my case.",
          "votes": 3
        },
        {
          "id": 1340471,
          "postDate": "2021-06-08T02:05:29.430Z",
          "content": "<p>Thank you for the report, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiments, when I use mixup, only VerticalFlip and HorizontalFlip improve my score but others worsen :(</p>",
          "rawMarkdown": "Thank you for the report, @yukia18 \nIn my experiments, when I use mixup, only VerticalFlip and HorizontalFlip improve my score but others worsen :(",
          "votes": 2
        },
        {
          "id": 1340491,
          "postDate": "2021-06-08T03:09:48.003Z",
          "content": "<p><a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> How did you implement RandomAffine ? . I cant seem to find it anywhere . Edit - So i found it. For those who want to implement this just use <code>albumentations.augmentations.geometric.transforms.Affine(p=0.5)</code></p>",
          "rawMarkdown": "@yukia18 How did you implement RandomAffine ? . I cant seem to find it anywhere . Edit - So i found it. For those who want to implement this just use `albumentations.augmentations.geometric.transforms.Affine(p=0.5)`"
        },
        {
          "id": 1340555,
          "postDate": "2021-06-08T05:10:59.557Z",
          "content": "<p><a href=\"https://pytorch.org/vision/stable/transforms.html?highlight=randomaffine#torchvision.transforms.RandomAffine\" target=\"_blank\">https://pytorch.org/vision/stable/transforms.html?highlight=randomaffine#torchvision.transforms.RandomAffine</a><br>\nI use torchvision.transforms instead of albumentations. I think RandomAffine is corresponding to ShiftScaleRotate of albumentations.</p>",
          "rawMarkdown": "https://pytorch.org/vision/stable/transforms.html?highlight=randomaffine#torchvision.transforms.RandomAffine\nI use torchvision.transforms instead of albumentations. I think RandomAffine is corresponding to ShiftScaleRotate of albumentations."
        }
      ]
    },
    {
      "id": 1324440,
      "postDate": "2021-05-27T00:24:22.497Z",
      "content": "<p>(1, 6 * 273, 256) -&gt; resize (1, 512, 512) -&gt; OOF=0.982170 <br>\n5 Fold Average LB : 0.97 </p>",
      "rawMarkdown": "(1, 6 * 273, 256) -> resize (1, 512, 512) -> OOF=0.982170 \n5 Fold Average LB : 0.97 ",
      "votes": 1
    },
    {
      "id": 1322707,
      "postDate": "2021-05-25T15:58:16.200Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <br>\nimage = np.vstack(image).transpose((1, 0)) or any type of stacking will stack 6 channels in 1 channel.<br>\nSo you are saying you resized stacked channels to these shapes like 256x256, 512x512 etc?</p>",
      "rawMarkdown": "Hi @ttahara \nimage = np.vstack(image).transpose((1, 0)) or any type of stacking will stack 6 channels in 1 channel.\nSo you are saying you resized stacked channels to these shapes like 256x256, 512x512 etc?",
      "votes": 2,
      "replies": [
        {
          "id": 1322995,
          "postDate": "2021-05-25T21:19:28.107Z",
          "content": "<p>Yes, I resize cadence snippets after stacking its channels into 1 channel along time axis and transposing it.</p>",
          "rawMarkdown": "Yes, I resize cadence snippets after stacking its channels into 1 channel along time axis and transposing it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1345980,
      "postDate": "2021-06-12T02:55:49.507Z",
      "content": "<p>Thanks for your discussion，I have benefited a lot.🎉</p>",
      "rawMarkdown": "Thanks for your discussion，I have benefited a lot.🎉"
    },
    {
      "id": 1337961,
      "postDate": "2021-06-06T04:25:11.130Z",
      "content": "<p>Sorry for the rudimentary question, but what kind of code do I need to write to resize from (6, 273, 256) to those sizes?<br>\nI'm a beginner in image identification.  :(</p>",
      "rawMarkdown": "Sorry for the rudimentary question, but what kind of code do I need to write to resize from (6, 273, 256) to those sizes?\nI'm a beginner in image identification.  :(",
      "replies": [
        {
          "id": 1338081,
          "postDate": "2021-06-06T06:49:40.363Z",
          "content": "<p>Here is one example when reading a cadence snippet file:</p>\n<pre><code># # read a cadence snippet file\nimg = np.load(file_path)  # shape: (position, time, frequency)=(6, 273, 256)\n# # stack 6 positions along time-axis\nimg = np.vstack(img)      # shape: (time, frequency)=(1638, 256)\n# #  switch axis\nimg = img.transpose(1, 0) # shape: (frequency, time)=(256, 1638)\n</code></pre>\n<p>※Note: switching axis is not necessary. </p>\n<p>After that I  resize the <code>img</code> by <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> as a part of augmentation.</p>\n<p>I recommend you <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's great notebook as a starter:<br>\n<a href=\"https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training</a></p>",
          "rawMarkdown": "Here is one example when reading a cadence snippet file:\n\n```python\n# # read a cadence snippet file\nimg = np.load(file_path)  # shape: (position, time, frequency)=(6, 273, 256)\n# # stack 6 positions along time-axis\nimg = np.vstack(img)      # shape: (time, frequency)=(1638, 256)\n# #  switch axis\nimg = img.transpose(1, 0) # shape: (frequency, time)=(256, 1638)\n```\n※Note: switching axis is not necessary. \n\nAfter that I  resize the `img` by [albumentations](https://github.com/albumentations-team/albumentations) as a part of augmentation.\n\nI recommend you @yasufuminakama 's great notebook as a starter:\nhttps://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training",
          "votes": 3
        },
        {
          "id": 1338377,
          "postDate": "2021-06-06T11:49:53.993Z",
          "content": "<p>Thank you very much!!<br>\nI'll do my best :)</p>",
          "rawMarkdown": "Thank you very much!!\nI'll do my best :)"
        },
        {
          "id": 1338394,
          "postDate": "2021-06-06T12:16:26.633Z",
          "content": "<p>Good Luck 👍</p>",
          "rawMarkdown": "Good Luck 👍",
          "votes": 3
        }
      ]
    },
    {
      "id": 1321200,
      "postDate": "2021-05-24T14:55:08.787Z",
      "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> Thanks for sharing :) </p>",
      "rawMarkdown": "@ttahara Thanks for sharing :) ",
      "replies": [
        {
          "id": 1322980,
          "postDate": "2021-05-25T21:00:38.910Z",
          "content": "<p>You're welcome :)</p>",
          "rawMarkdown": "You're welcome :)"
        }
      ]
    },
    {
      "id": 1319807,
      "postDate": "2021-05-23T14:02:53.713Z",
      "content": "<p>Can you please add spatial image size in this to compare ?</p>",
      "rawMarkdown": "Can you please add spatial image size in this to compare ?",
      "replies": [
        {
          "id": 1319819,
          "postDate": "2021-05-23T14:17:27.073Z",
          "content": "<p>Sorry I'm not sure I understand what you mean, but the image sizes above are input size of CNN model.</p>",
          "rawMarkdown": "Sorry I'm not sure I understand what you mean, but the image sizes above are input size of CNN model.",
          "votes": 1
        },
        {
          "id": 1319823,
          "postDate": "2021-05-23T14:19:29.830Z",
          "content": "<p>I mean the size (256, 1638, 3) you can convert them with this code `    list_x = [id_to_path(x, self.is_train) for x in batch_ids]<br>\n        batch_x = []<br>\n        for i in list_x:<br>\n            image = np.load(i).astype(float)<br>\n            image = np.vstack(image).transpose((1, 0))<br>\n            x = np.zeros(shape=(256, 1638, 3))</p>\n<pre><code>        x[:, :, 0] = image\n        x[:, :, 1] = image\n        x[:, :, 2] = image`\n</code></pre>",
          "rawMarkdown": "I mean the size (256, 1638, 3) you can convert them with this code `\tlist_x = [id_to_path(x, self.is_train) for x in batch_ids]\n\t\tbatch_x = []\n\t\tfor i in list_x:\n\t\t\timage = np.load(i).astype(float)\n\t\t\timage = np.vstack(image).transpose((1, 0))\n\t\t\tx = np.zeros(shape=(256, 1638, 3))\n\n\t\t\tx[:, :, 0] = image\n\t\t\tx[:, :, 1] = image\n\t\t\tx[:, :, 2] = image`"
        },
        {
          "id": 1319829,
          "postDate": "2021-05-23T14:26:35.677Z",
          "content": "<p>I used models which input images have only <strong>1</strong> channel.<br>\nSo image sizes above mean (Height, Width).</p>",
          "rawMarkdown": "I used models which input images have only **1** channel.\nSo image sizes above mean (Height, Width).",
          "votes": 3
        },
        {
          "id": 1324673,
          "postDate": "2021-05-27T06:20:31.433Z",
          "content": "<p>Can you tell the model because resnet18 takes 3 channels?</p>",
          "rawMarkdown": "Can you tell the model because resnet18 takes 3 channels?"
        },
        {
          "id": 1325351,
          "postDate": "2021-05-27T17:16:38.547Z",
          "content": "<p>I use <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">Pytorch Image Models (timm)</a>.create_model api for 1 channel model like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's <a href=\"https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training\" target=\"_blank\">starter notebook</a>.<br>\nThe following code is an example. <code>in_chans</code> specifies the number of channels for model input.</p>\n<pre><code>import timm\n\nclass SingleChannelModel(nn.Module):\n\n    def __init__(self, base_name, n_targets, pretrained=False):\n        super().__init__()\n        self.backbone =timm.create_model(\n            base_name, num_classes=0, pretrained=pretrained, in_chans=1)\n        in_features = self.backbone.num_features\n\n        self.head_fc = nn.Linear(in_features, n_targets)\n\n    def forward(self, x):\n        h = self.backbone(x)\n        y = self.head_fc(h)\n        return y\n\n# # create model\nmodel = SingleChannelModel(\"resnet18d\", 1, True)\n</code></pre>\n<p><br><br>\nAccording to <a href=\"https://fastai.github.io/timmdocs/models#Case-1:-When-the-number-of-input-channels-is-1\" target=\"_blank\">docs</a>: </p>\n<blockquote>\n  <p>If the number of input channels is 1, timm simply sums the 3 channel weights into a single channel to update the shape of conv1.weight to be [64, 1, 7, 7].</p>\n</blockquote>",
          "rawMarkdown": "I use [Pytorch Image Models (timm)](https://github.com/rwightman/pytorch-image-models).create_model api for 1 channel model like @yasufuminakama 's [starter notebook](https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training).\nThe following code is an example. `in_chans` specifies the number of channels for model input.\n\n```python\n\nimport timm\n\nclass SingleChannelModel(nn.Module):\n\n    def __init__(self, base_name, n_targets, pretrained=False):\n        super().__init__()\n        self.backbone =timm.create_model(\n            base_name, num_classes=0, pretrained=pretrained, in_chans=1)\n        in_features = self.backbone.num_features\n\n        self.head_fc = nn.Linear(in_features, n_targets)\n\n    def forward(self, x):\n        h = self.backbone(x)\n        y = self.head_fc(h)\n        return y\n\n# # create model\nmodel = SingleChannelModel(\"resnet18d\", 1, True)\n```\n<br>\nAccording to [docs](https://fastai.github.io/timmdocs/models#Case-1:-When-the-number-of-input-channels-is-1): \n> If the number of input channels is 1, timm simply sums the 3 channel weights into a single channel to update the shape of conv1.weight to be [64, 1, 7, 7].\n",
          "votes": 2
        },
        {
          "id": 1325446,
          "postDate": "2021-05-27T18:48:39.937Z",
          "rawMarkdown": "",
          "votes": 4,
          "isDeleted": true
        },
        {
          "id": 1325674,
          "postDate": "2021-05-27T23:48:32.623Z",
          "content": "<p>Thanks for information.</p>\n<p>I usually create <code>head_fc</code> out of <code>backbone</code> for customizing head layers.</p>",
          "rawMarkdown": "Thanks for information.\n\nI usually create `head_fc` out of `backbone` for customizing head layers.",
          "votes": 1
        },
        {
          "id": 1328265,
          "postDate": "2021-05-30T05:29:35.353Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
          "rawMarkdown": "thanks @ttahara "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1326165,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-05-28T09:15:03.153000",
      "content": "<p>I did the same experiments for <code>resnet34d</code>.</p>\n<table>\n<thead>\n<tr>\n<th>Input Image Size(Frequency x Time)</th>\n<th>OOF Loss</th>\n<th>OOF Score</th>\n<th>Public LB Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>384x384</td>\n<td>0.04230</td>\n<td>0.9862</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>512x512</td>\n<td>0.03971</td>\n<td>0.9869</td>\n<td>0.97</td>\n</tr>\n<tr>\n<td>640x640</td>\n<td>0.04003</td>\n<td>0.9876</td>\n<td><strong>0.97</strong></td>\n</tr>\n<tr>\n<td>768x768</td>\n<td><strong>0.03860</strong></td>\n<td><strong>0.9886</strong></td>\n<td>0.97</td>\n</tr>\n</tbody>\n</table>\n<p>The Public score order of them is as follows:<br>\n<strong>640x640</strong> &gt; 768x768 &gt; 512x512 &gt; 384x384</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1323660,
      "author_name": "Tawara",
      "author_url": "",
      "post_date": "2021-05-26T11:39:36.040000",
      "content": "<p>Update: I added 768x768 result.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1323184,
      "author_name": "yabea",
      "author_url": "",
      "post_date": "2021-05-26T04:24:15.240000",
      "content": "<p>Thanks for sharing.<br>\nBased on your results, I hypothesis that upsampling freq axis more than original (=256) size helps, and it may be true (but it's rough experiments).<br>\n(1, 6 * 273, 256) -&gt; CV=0.97688<br>\n(1, 6 * 273, 256) -&gt; resize (1, 512, 512) -&gt; CV=0.98290<br>\nI will try experiments about sensibility of time (=273) axis.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1323215,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-26T04:58:19.423000",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>I hypothesis that upsampling freq axis more than original (=256) size helps</p>\n</blockquote>\n<p>I have  similar thoughts with you, and I tried the cases where frequency axis is bigger than time axis.</p>\n<table>\n<thead>\n<tr>\n<th>Size (Freq x Time)</th>\n<th>val score (fold 0)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512x512 (for comparison)</td>\n<td><strong>0.9860</strong></td>\n</tr>\n<tr>\n<td><strong>512x256</strong></td>\n<td>0.9769</td>\n</tr>\n<tr>\n<td>640x640 (for comparison)</td>\n<td><strong>0.9866</strong></td>\n</tr>\n<tr>\n<td><strong>640x320</strong></td>\n<td>0.9799</td>\n</tr>\n<tr>\n<td>768x768 (for comparison)</td>\n<td><strong>0.9871</strong></td>\n</tr>\n<tr>\n<td><strong>768x384</strong></td>\n<td>0.9805</td>\n</tr>\n</tbody>\n</table>\n<p>This is very rough experiment but square inputs may be better?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1323302,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2021-05-26T06:23:50.090000",
          "content": "<p>Thanks for sharing your result, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>\n<p>I have the same hypothesis as <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiments, CV-score is following:<br>\n(512, 512) &gt; (6*273, 256) &gt; (224, 224)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1323646,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-05-26T11:21:12.933000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> and <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiment, <br>\nFor Fold 0, I resized 512x512 and trained EfficientNet-B0 for 25 epochs and validation score was 0.972 but submission was around 0.95.<br>\nI used ScaleShiftRotate, VerticalFlip, HorizontalFlip, Mixup augmentations as well.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1324678,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-05-27T06:27:43.157000",
          "content": "<p>Yes same thing is happening with me i get fold 0 score 0.9702 and fold 1 score 0.9741 and cv from these 2 folds come 0.9723 but lb is only 0.95.<br>\nI used effentB0 with input shape (3,244,244) used vstacking.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1325779,
          "author_name": "yabea",
          "author_url": "",
          "post_date": "2021-05-28T02:07:54.407000",
          "content": "<p>Updated:</p>\n<table>\n<thead>\n<tr>\n<th>image size (time x freq)</th>\n<th>CV</th>\n<th>Rank on LB (all 0.97x)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>512 x 512</td>\n<td>0.98735</td>\n<td>4</td>\n</tr>\n<tr>\n<td>768 x 768</td>\n<td>0.98943</td>\n<td>2</td>\n</tr>\n<tr>\n<td><strong>1024 x 512</strong></td>\n<td><strong>0.98999</strong></td>\n<td><strong>1</strong></td>\n</tr>\n<tr>\n<td>512 x 1024</td>\n<td>0.98871</td>\n<td>3</td>\n</tr>\n</tbody>\n</table>\n<p>As I expected, upsampling time axis contributes a lot. On the other hand, expanding freq axis more than 512 still helps. I know that large image size always helps in the almost competition, but I didn't know that it is applicable to upsample image more than original size…</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1325797,
          "author_name": "Salman",
          "author_url": "",
          "post_date": "2021-05-28T02:38:04.943000",
          "content": "<p>Number of Epochs?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1325806,
          "author_name": "yabea",
          "author_url": "",
          "post_date": "2021-05-28T02:49:32.523000",
          "content": "<p>10 epochs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1325856,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-28T04:27:16.667000",
          "content": "<p><a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a><br>\nThank you for sharing. I've not tried yet cases where frequency axis is smaller than time axis.<br>\nNow I have started training that case. I'll report the result later.</p>\n<blockquote>\n  <p>but I didn't know that it is applicable to upsample image more than original size…</p>\n</blockquote>\n<p>IMO, expanding freq axis emphasizes E.T. Signals' frequency change over time. It may make signals easier to be find.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1328193,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-30T03:08:59.550000",
          "content": "<p>I tried (Freq,Time) = (512, 1024) and (512, 768). But they got lower CV than 512x512 …🤔</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1329296,
          "author_name": "Yi Wu",
          "author_url": "",
          "post_date": "2021-05-31T03:28:29.880000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yabea\" target=\"_blank\">@yabea</a> did you use data aug in training phase, other than resizing and normalization?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1340467,
          "author_name": "yabea",
          "author_url": "",
          "post_date": "2021-06-08T01:52:16.583000",
          "content": "<p>RandomAffine and Mixup. I'm sure Mixup helps both CV and LB as others mentioned.<br>\nTranspose and GaussianBlur got worse dramatically in my case.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1340471,
          "author_name": "Yamame🐟",
          "author_url": "",
          "post_date": "2021-06-08T02:05:29.430000",
          "content": "<p>Thank you for the report, <a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> <br>\nIn my experiments, when I use mixup, only VerticalFlip and HorizontalFlip improve my score but others worsen :(</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1340491,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-06-08T03:09:48.003000",
          "content": "<p><a href=\"https://www.kaggle.com/yukia18\" target=\"_blank\">@yukia18</a> How did you implement RandomAffine ? . I cant seem to find it anywhere . Edit - So i found it. For those who want to implement this just use <code>albumentations.augmentations.geometric.transforms.Affine(p=0.5)</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1340555,
          "author_name": "yabea",
          "author_url": "",
          "post_date": "2021-06-08T05:10:59.557000",
          "content": "<p><a href=\"https://pytorch.org/vision/stable/transforms.html?highlight=randomaffine#torchvision.transforms.RandomAffine\" target=\"_blank\">https://pytorch.org/vision/stable/transforms.html?highlight=randomaffine#torchvision.transforms.RandomAffine</a><br>\nI use torchvision.transforms instead of albumentations. I think RandomAffine is corresponding to ShiftScaleRotate of albumentations.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1324440,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2021-05-27T00:24:22.497000",
      "content": "<p>(1, 6 * 273, 256) -&gt; resize (1, 512, 512) -&gt; OOF=0.982170 <br>\n5 Fold Average LB : 0.97 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1322707,
      "author_name": "Salman",
      "author_url": "",
      "post_date": "2021-05-25T15:58:16.200000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> <br>\nimage = np.vstack(image).transpose((1, 0)) or any type of stacking will stack 6 channels in 1 channel.<br>\nSo you are saying you resized stacked channels to these shapes like 256x256, 512x512 etc?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1322995,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-25T21:19:28.107000",
          "content": "<p>Yes, I resize cadence snippets after stacking its channels into 1 channel along time axis and transposing it.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1345980,
      "author_name": "Gainover",
      "author_url": "",
      "post_date": "2021-06-12T02:55:49.507000",
      "content": "<p>Thanks for your discussion，I have benefited a lot.🎉</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1337961,
      "author_name": "regichan",
      "author_url": "",
      "post_date": "2021-06-06T04:25:11.130000",
      "content": "<p>Sorry for the rudimentary question, but what kind of code do I need to write to resize from (6, 273, 256) to those sizes?<br>\nI'm a beginner in image identification.  :(</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1338081,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-06T06:49:40.363000",
          "content": "<p>Here is one example when reading a cadence snippet file:</p>\n<pre><code># # read a cadence snippet file\nimg = np.load(file_path)  # shape: (position, time, frequency)=(6, 273, 256)\n# # stack 6 positions along time-axis\nimg = np.vstack(img)      # shape: (time, frequency)=(1638, 256)\n# #  switch axis\nimg = img.transpose(1, 0) # shape: (frequency, time)=(256, 1638)\n</code></pre>\n<p>※Note: switching axis is not necessary. </p>\n<p>After that I  resize the <code>img</code> by <a href=\"https://github.com/albumentations-team/albumentations\" target=\"_blank\">albumentations</a> as a part of augmentation.</p>\n<p>I recommend you <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's great notebook as a starter:<br>\n<a href=\"https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training\" target=\"_blank\">https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1338377,
          "author_name": "regichan",
          "author_url": "",
          "post_date": "2021-06-06T11:49:53.993000",
          "content": "<p>Thank you very much!!<br>\nI'll do my best :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1338394,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-06-06T12:16:26.633000",
          "content": "<p>Good Luck 👍</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1321200,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2021-05-24T14:55:08.787000",
      "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> Thanks for sharing :) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1322980,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-25T21:00:38.910000",
          "content": "<p>You're welcome :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1319807,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-23T14:02:53.713000",
      "content": "<p>Can you please add spatial image size in this to compare ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1319819,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-23T14:17:27.073000",
          "content": "<p>Sorry I'm not sure I understand what you mean, but the image sizes above are input size of CNN model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1319823,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-05-23T14:19:29.830000",
          "content": "<p>I mean the size (256, 1638, 3) you can convert them with this code `    list_x = [id_to_path(x, self.is_train) for x in batch_ids]<br>\n        batch_x = []<br>\n        for i in list_x:<br>\n            image = np.load(i).astype(float)<br>\n            image = np.vstack(image).transpose((1, 0))<br>\n            x = np.zeros(shape=(256, 1638, 3))</p>\n<pre><code>        x[:, :, 0] = image\n        x[:, :, 1] = image\n        x[:, :, 2] = image`\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1319829,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-23T14:26:35.677000",
          "content": "<p>I used models which input images have only <strong>1</strong> channel.<br>\nSo image sizes above mean (Height, Width).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1324673,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-05-27T06:20:31.433000",
          "content": "<p>Can you tell the model because resnet18 takes 3 channels?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325351,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-27T17:16:38.547000",
          "content": "<p>I use <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">Pytorch Image Models (timm)</a>.create_model api for 1 channel model like <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> 's <a href=\"https://www.kaggle.com/yasufuminakama/seti-nfnet-l0-starter-training\" target=\"_blank\">starter notebook</a>.<br>\nThe following code is an example. <code>in_chans</code> specifies the number of channels for model input.</p>\n<pre><code>import timm\n\nclass SingleChannelModel(nn.Module):\n\n    def __init__(self, base_name, n_targets, pretrained=False):\n        super().__init__()\n        self.backbone =timm.create_model(\n            base_name, num_classes=0, pretrained=pretrained, in_chans=1)\n        in_features = self.backbone.num_features\n\n        self.head_fc = nn.Linear(in_features, n_targets)\n\n    def forward(self, x):\n        h = self.backbone(x)\n        y = self.head_fc(h)\n        return y\n\n# # create model\nmodel = SingleChannelModel(\"resnet18d\", 1, True)\n</code></pre>\n<p><br><br>\nAccording to <a href=\"https://fastai.github.io/timmdocs/models#Case-1:-When-the-number-of-input-channels-is-1\" target=\"_blank\">docs</a>: </p>\n<blockquote>\n  <p>If the number of input channels is 1, timm simply sums the 3 channel weights into a single channel to update the shape of conv1.weight to be [64, 1, 7, 7].</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1325446,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-27T18:48:39.937000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1325674,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-05-27T23:48:32.623000",
          "content": "<p>Thanks for information.</p>\n<p>I usually create <code>head_fc</code> out of <code>backbone</code> for customizing head layers.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1328265,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-05-30T05:29:35.353000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1319803": "Here I compare the models trained under the same experimental settings, except for the input image size.\n\n#### settings\n\n* Model: resnet18d\n* Initial Weights: ImageNet\n* Train-Val Split: StratifiedKFold(K=5)\n* Data Augmentation: HorizontalFlip, VerticalFlip, ShiftScaleRotate, RandomResizedCrop\n* Loss: Binary Cross Entropy\n* Optimizer: AdamW\n* Scheduler: OneCycleLR\n\n#### results\n\n|Input Image Size(Frequency x Time) | OOF Loss | OOF Score | Public LB Score |\n|:-----------------------------:|:---:|:----:|\n| 256x256 | 0.05147 | 0.9818 | 0.97 |\n| 384x384 | 0.04485 | 0.9847 | 0.97 |\n| 512x512 | 0.04180 | 0.9868 | 0.97 |\n| 640x640 | 0.04188 | 0.9863 | **0.97** |\n| 768x768(new) | **0.04038** | **0.9868** | 0.97 |\n\nThe Public score order of them is as follows:  \n**640x640** > 768x768(new) > 512x512 > 384x384 > 256x256\n\nMaybe, 512x512 is good ?\n\n**UPDATE1**: add 768x768 result to the table above.",
    "1326165": "I did the same experiments for `resnet34d`.\n\n|Input Image Size(Frequency x Time) | OOF Loss | OOF Score | Public LB Score |\n|:-----------------------------:|:---:|:----:|\n| 384x384 | 0.04230 | 0.9862 | 0.97 |\n| 512x512 | 0.03971 | 0.9869 | 0.97 |\n| 640x640 | 0.04003 | 0.9876  | **0.97** |\n| 768x768 | **0.03860** |**0.9886** | 0.97 |\n\nThe Public score order of them is as follows:\n**640x640** > 768x768 > 512x512 > 384x384",
    "1323660": "Update: I added 768x768 result.",
    "1323184": "Thanks for sharing.\nBased on your results, I hypothesis that upsampling freq axis more than original (=256) size helps, and it may be true (but it's rough experiments).\n(1, 6 * 273, 256) -> CV=0.97688\n(1, 6 * 273, 256) -> resize (1, 512, 512) -> CV=0.98290\nI will try experiments about sensibility of time (=273) axis.",
    "1324440": "(1, 6 * 273, 256) -> resize (1, 512, 512) -> OOF=0.982170 \n5 Fold Average LB : 0.97 ",
    "1322707": "Hi @ttahara \nimage = np.vstack(image).transpose((1, 0)) or any type of stacking will stack 6 channels in 1 channel.\nSo you are saying you resized stacked channels to these shapes like 256x256, 512x512 etc?",
    "1345980": "Thanks for your discussion，I have benefited a lot.🎉",
    "1337961": "Sorry for the rudimentary question, but what kind of code do I need to write to resize from (6, 273, 256) to those sizes?\nI'm a beginner in image identification.  :(",
    "1321200": "@ttahara Thanks for sharing :) ",
    "1319807": "Can you please add spatial image size in this to compare ?"
  }
}