{
  "id": 266813,
  "title": "Can someone post code that scores Over LB 0.780?",
  "url": "/competitions/seti-breakthrough-listen/discussion/266813",
  "author_name": "Chris Deotte",
  "post_date": "2021-08-20T15:17:12.717000",
  "votes": 13,
  "comment_count": 26,
  "views": 0,
  "content": "<h3>How Does Code achieve 0.780+?</h3>\n<p>I would love to see code that scores over LB 0.780. I would like to learn what my models are missing. </p>\n<p>Does anyone have code, that trains with new data only without data preprocess. Where they use only \"on\" cadence <code>np.vstack( img[::2] )</code> and feed 1 channel images into a CNN? And then just uses augmentation and learning schedule and achieves LB 0.780+</p>\n<p>(I see that second place posted their code, but it throws an error and I don't want to bother them to debug it for me).</p>\n<h3>UPDATE: Mixup is needed to achieve 0.780+</h3>\n<p>Thanks everyone for sharing your code with me. I have learned that mixup is the secret ingredient to achieve Silver Medal (and large backbone with large image size)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! </p>\n<p>The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
  "messages": [
    {
      "id": 1483318,
      "postDate": "2021-08-20T15:17:12.717Z",
      "content": "<h3>How Does Code achieve 0.780+?</h3>\n<p>I would love to see code that scores over LB 0.780. I would like to learn what my models are missing. </p>\n<p>Does anyone have code, that trains with new data only without data preprocess. Where they use only \"on\" cadence <code>np.vstack( img[::2] )</code> and feed 1 channel images into a CNN? And then just uses augmentation and learning schedule and achieves LB 0.780+</p>\n<p>(I see that second place posted their code, but it throws an error and I don't want to bother them to debug it for me).</p>\n<h3>UPDATE: Mixup is needed to achieve 0.780+</h3>\n<p>Thanks everyone for sharing your code with me. I have learned that mixup is the secret ingredient to achieve Silver Medal (and large backbone with large image size)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! </p>\n<p>The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "rawMarkdown": "### How Does Code achieve 0.780+?\nI would love to see code that scores over LB 0.780. I would like to learn what my models are missing. \n\nDoes anyone have code, that trains with new data only without data preprocess. Where they use only \"on\" cadence `np.vstack( img[::2] )` and feed 1 channel images into a CNN? And then just uses augmentation and learning schedule and achieves LB 0.780+\n\n(I see that second place posted their code, but it throws an error and I don't want to bother them to debug it for me).\n\n### UPDATE: Mixup is needed to achieve 0.780+\nThanks everyone for sharing your code with me. I have learned that mixup is the secret ingredient to achieve Silver Medal (and large backbone with large image size)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! \n\nThe score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!",
      "votes": 13
    },
    {
      "id": 1483346,
      "postDate": "2021-08-20T15:34:14.217Z",
      "content": "<p>What you already tried Chris? </p>\n<p>Things that helps score close to 0.79:  double frequency axis size from 256-&gt;512(enlarge lines),  mixup, dropout, weightedsampling (increase to 20-30% positives during training), blend of 5 folds with TTA(hflip, vflip, hvflip, resize), train with my target generator applied randomly to control images (this double the examples to train). </p>",
      "rawMarkdown": "What you already tried Chris? \n\nThings that helps score close to 0.79:  double frequency axis size from 256->512(enlarge lines),  mixup, dropout, weightedsampling (increase to 20-30% positives during training), blend of 5 folds with TTA(hflip, vflip, hvflip, resize), train with my target generator applied randomly to control images (this double the examples to train). ",
      "votes": 5,
      "replies": [
        {
          "id": 1483542,
          "postDate": "2021-08-20T17:37:07.400Z",
          "content": "<p>I built 100+ different models in the past week! I thought I tried everything 😄</p>\n<p>I'm thinking that i implemented something wrong. I plan to revisit my mixup experiments which did not increase CV but I believe should. I will also try some of Statking's suggestions above to see if they close my CV LB gap.</p>\n<p>I like your idea of upsampling positive. I will try that thanks. I noticed that train is 10% positive and test is 20% positive, so perhaps that will help. (However with AUC, i'm not sure).</p>\n<p>Eventually, i will try your generator. I like that, but first i want to achieve LB 0.780 without external data nor preprocess.</p>",
          "rawMarkdown": "I built 100+ different models in the past week! I thought I tried everything 😄\n\nI'm thinking that i implemented something wrong. I plan to revisit my mixup experiments which did not increase CV but I believe should. I will also try some of Statking's suggestions above to see if they close my CV LB gap.\n\nI like your idea of upsampling positive. I will try that thanks. I noticed that train is 10% positive and test is 20% positive, so perhaps that will help. (However with AUC, i'm not sure).\n\nEventually, i will try your generator. I like that, but first i want to achieve LB 0.780 without external data nor preprocess."
        },
        {
          "id": 1483651,
          "postDate": "2021-08-20T18:45:45.357Z",
          "content": "<p>But why limit yourself to new data? Old data gave us very significant boosts.</p>",
          "rawMarkdown": "But why limit yourself to new data? Old data gave us very significant boosts.",
          "votes": 2
        },
        {
          "id": 1483802,
          "postDate": "2021-08-20T21:52:37Z",
          "content": "<p>At this point, i just want to know what my model is missing and why my CV LB gap is large. I think there is an important learning lesson for me here.</p>\n<p>During the competition, i was planning to try old data but there was too many other things to explore in my short time.</p>",
          "rawMarkdown": "At this point, i just want to know what my model is missing and why my CV LB gap is large. I think there is an important learning lesson for me here.\n\nDuring the competition, i was planning to try old data but there was too many other things to explore in my short time.",
          "votes": 2
        },
        {
          "id": 1484343,
          "postDate": "2021-08-21T08:28:26.560Z",
          "content": "<p>Makes sense, I am quite sure that using mixup is important.</p>",
          "rawMarkdown": "Makes sense, I am quite sure that using mixup is important.",
          "votes": 1
        },
        {
          "id": 1484611,
          "postDate": "2021-08-21T13:00:14.727Z",
          "content": "<p>I think you are right. I am sure the secret (to closing CV LB gap) is to teach the model to learn the signal and ignore the background like we discussed in the other thread. Mixup (with max target) seems important to accomplish this.</p>",
          "rawMarkdown": "I think you are right. I am sure the secret (to closing CV LB gap) is to teach the model to learn the signal and ignore the background like we discussed in the other thread. Mixup (with max target) seems important to accomplish this.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1484601,
      "postDate": "2021-08-21T12:56:55.620Z",
      "content": "<p>UPDATE: Thanks everyone for your help. I'm excited to say that my CV LB gap is closing by using the suggestions in this thread. I made 8 changes to my model and it boosted CV and boosted LB and closed the gap! I believe using 5 folds (with small backbone) or 1 fold (with large backbone) then the single model will achieve <code>Silver Medal</code>! </p>\n<p>I am now making one change at a time to see which is the most important. Here is the list of 6 changes I made</p>\n<ul>\n<li>Change <code>20 epochs cosine without warmup</code> to <code>40 epochs cosine with warmup</code>.</li>\n<li>Use  <code>x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)</code> to convert 1 channel to 3 channel instead of <code>x = tf.keras.layers.Concatenate(axis=-1)([inp,inp,inp])</code></li>\n<li>Add <code>dropout(0.15)</code> between efficientnet embeddings and final sigmoid output</li>\n<li>Add albumentations <code>coarse dropout with max_holes=32, max_height=size/10</code> </li>\n<li>Add mixup within batch with <code>alpha=3</code></li>\n<li>After <code>np.vstack()</code> apply normalize with subtract mean, divide std.</li>\n<li>Use random brightness augmentation, i.e. <code>img += np.random.uniform(-1,1)</code></li>\n<li>Apply augmentation to image separately before mixup instead of after mixup</li>\n</ul>\n<p>I will also display GradCam with my old model and new model to see what the model is learning differently.</p>",
      "rawMarkdown": "UPDATE: Thanks everyone for your help. I'm excited to say that my CV LB gap is closing by using the suggestions in this thread. I made 8 changes to my model and it boosted CV and boosted LB and closed the gap! I believe using 5 folds (with small backbone) or 1 fold (with large backbone) then the single model will achieve `Silver Medal`! \n\nI am now making one change at a time to see which is the most important. Here is the list of 6 changes I made\n* Change `20 epochs cosine without warmup` to `40 epochs cosine with warmup`.\n* Use  `x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)` to convert 1 channel to 3 channel instead of `x = tf.keras.layers.Concatenate(axis=-1)([inp,inp,inp])`\n* Add `dropout(0.15)` between efficientnet embeddings and final sigmoid output\n* Add albumentations `coarse dropout with max_holes=32, max_height=size/10` \n* Add mixup within batch with `alpha=3`\n* After `np.vstack()` apply normalize with subtract mean, divide std.\n* Use random brightness augmentation, i.e. `img += np.random.uniform(-1,1)`\n* Apply augmentation to image separately before mixup instead of after mixup\n\nI will also display GradCam with my old model and new model to see what the model is learning differently.",
      "votes": 4
    },
    {
      "id": 1487186,
      "postDate": "2021-08-23T13:18:12.970Z",
      "content": "<h3>Mystery solved! Mixup can achieve LB 0.780+!</h3>\n<p>(and large backbone with large image size)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make to beat LB 780 and get Silver medal. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "rawMarkdown": "### Mystery solved! Mixup can achieve LB 0.780+! \n(and large backbone with large image size)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make to beat LB 780 and get Silver medal. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385",
      "votes": 1
    },
    {
      "id": 1484197,
      "postDate": "2021-08-21T05:50:24.307Z",
      "content": "<p>Maybe check out this silver post by <a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a> <br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884</a></p>\n<p>in the end my best model was single EfficientNet-B4 with non-symmetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\"</p>\n<p>not sure if the code is available or could be, but perhaps the non-symmetric cropping idea? had not see that before. </p>",
      "rawMarkdown": "Maybe check out this silver post by @blankaf \nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\n\nin the end my best model was single EfficientNet-B4 with non-symmetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\"\n\nnot sure if the code is available or could be, but perhaps the non-symmetric cropping idea? had not see that before. ",
      "votes": 1,
      "replies": [
        {
          "id": 1484947,
          "postDate": "2021-08-21T17:14:57.393Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> . What was the public score for your EB4 that had private LB 0.78191? And what was the public and private scores of the final submission you did select?</p>",
          "rawMarkdown": "Thanks @something4kag . What was the public score for your EB4 that had private LB 0.78191? And what was the public and private scores of the final submission you did select?"
        },
        {
          "id": 1485657,
          "postDate": "2021-08-22T10:28:35.320Z",
          "content": "<p>Chris, <br>\nyou are actually asking about my solutions <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\" target=\"_blank\">here</a>.<br>\nAs I have written there, I trained (with the same data \"tricks\") 2 models: Resnet34d and EfficientNet-B4. For my finals I chose 2 different blends of them, the better which scored me 47th private was private LB 0.78172 and public LB 0.78618. Apparently all blends or models' (hyper)parameters combinations where not very significant for private LB results because I had 26 submissions scoring over 0.780 private. All the submissions with Resnet34d were overfitting the public LB (as you see about 0.005) but all \"pure\" EfficientNet-B4 had nearly equal private and public LB score. My mentioned best solution even had better private: 0.78191 than public: 0.78175. <br>\nSo the main trick was really in the data processing: non-symmetric cropping and training with old positive samples. And EfficientNet was better than Resnet for better generalization. I believe that with higher EfficientNet model with higher resolution the results could be even much better.  </p>",
          "rawMarkdown": "Chris, \nyou are actually asking about my solutions [here](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884).\nAs I have written there, I trained (with the same data \"tricks\") 2 models: Resnet34d and EfficientNet-B4. For my finals I chose 2 different blends of them, the better which scored me 47th private was private LB 0.78172 and public LB 0.78618. Apparently all blends or models' (hyper)parameters combinations where not very significant for private LB results because I had 26 submissions scoring over 0.780 private. All the submissions with Resnet34d were overfitting the public LB (as you see about 0.005) but all \"pure\" EfficientNet-B4 had nearly equal private and public LB score. My mentioned best solution even had better private: 0.78191 than public: 0.78175. \nSo the main trick was really in the data processing: non-symmetric cropping and training with old positive samples. And EfficientNet was better than Resnet for better generalization. I believe that with higher EfficientNet model with higher resolution the results could be even much better.  ",
          "votes": 1
        },
        {
          "id": 1485741,
          "postDate": "2021-08-22T11:35:08.593Z",
          "content": "<p>Thanks. That's weird how Resnet34d overfit more. When you say \"old positive\" samples. Do you mean the train data with <code>target=1</code> from the data <strong>before</strong> competition reset?</p>",
          "rawMarkdown": "Thanks. That's weird how Resnet34d overfit more. When you say \"old positive\" samples. Do you mean the train data with `target=1` from the data **before** competition reset?"
        },
        {
          "id": 1485787,
          "postDate": "2021-08-22T12:28:38.677Z",
          "content": "<p>I exactly mean that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.</p>",
          "rawMarkdown": "I exactly mean that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1483387,
      "postDate": "2021-08-20T15:48:39.967Z",
      "content": "<p>I could easily achieve over LB 780 with a single model ( 5 FOLD SUBMISSION BLENDING).</p>\n<p><strong>Preprocess</strong> : only new data, pick 3 channels(0,2,4) and normalize by subtracting mean and by dividing std each channel<br>\nthen the size becomes [256,819]<br>\nWhen I made tfrecords, i loaded pure numpy files , and serialized it(i didn't use cv.imencode because it removes  some signals)</p>\n<p><strong>Training</strong> : At least IMAGE_SIZE [720,720]. and efficientnet &gt;= 4 with mixup lambda = 1<br>\n                      (INPUT IMAGE SIZE [720,720,1] and by tf.keras.layers.CONV2D, size became [720,720,3], after that input goes into the efficientnet),<br>\n<strong>Optimizer</strong> : CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS</p>\n<p><strong>AUGMENTATION</strong> : randomcropresize(0.9,1.0), flip vertical, horizontal, a slight coarse dropout, brightnesschange(since input channel is 1, i just added some number from uniform distribution)</p>\n<p><strong>TTA NUMS</strong> : 11 (randomcropandresize, flip vertical, horizontal, coarsedropout, brightnesschange)</p>\n<p>BUT My best single model score was LB0.78990  and i tried a lot of things, i coudn't achieve a score of LB &gt; 0.78990. with this method. (EFFICIENTNET 6, SIZE 720x720)</p>",
      "rawMarkdown": "I could easily achieve over LB 780 with a single model ( 5 FOLD SUBMISSION BLENDING).\n\n**Preprocess** : only new data, pick 3 channels(0,2,4) and normalize by subtracting mean and by dividing std each channel\nthen the size becomes [256,819]\nWhen I made tfrecords, i loaded pure numpy files , and serialized it(i didn't use cv.imencode because it removes  some signals)\n\n**Training** : At least IMAGE_SIZE [720,720]. and efficientnet >= 4 with mixup lambda = 1\n                      (INPUT IMAGE SIZE [720,720,1] and by tf.keras.layers.CONV2D, size became [720,720,3], after that input goes into the efficientnet),\n**Optimizer** : CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS\n\n**AUGMENTATION** : randomcropresize(0.9,1.0), flip vertical, horizontal, a slight coarse dropout, brightnesschange(since input channel is 1, i just added some number from uniform distribution)\n\n**TTA NUMS** : 11 (randomcropandresize, flip vertical, horizontal, coarsedropout, brightnesschange)\n \nBUT My best single model score was LB0.78990  and i tried a lot of things, i coudn't achieve a score of LB > 0.78990. with this method. (EFFICIENTNET 6, SIZE 720x720)\n",
      "votes": 1,
      "replies": [
        {
          "id": 1483399,
          "postDate": "2021-08-20T16:00:53.410Z",
          "content": "<p>This is great thanks. This is nearly identical to my best model which gets CV 0.887. It uses 768x768 and EffNetB4. But for some reason my model gets LB 0.765. If it had the typical CV LB gap, then it would also achieve LB 0.780+. I need to debug this model of mine. Something must not be right.</p>",
          "rawMarkdown": "This is great thanks. This is nearly identical to my best model which gets CV 0.887. It uses 768x768 and EffNetB4. But for some reason my model gets LB 0.765. If it had the typical CV LB gap, then it would also achieve LB 0.780+. I need to debug this model of mine. Something must not be right.",
          "votes": 2
        },
        {
          "id": 1483419,
          "postDate": "2021-08-20T16:13:33.807Z",
          "content": "<p>This code is from your MIXUP CODE(FLOWER TPU), i changed some</p>\n<pre><code>import random\nclass CoarseDropout:\n  def __init__(self, max_holes=32, size=0.06):\n    self.size = size\n    self.max_holes = max_holes\n\n  def __call__(self, image):\n      P = random.uniform(0,1)\n      height = image.shape[0]\n      width = image.shape[1]\n      for _n in range(self.max_holes):\n          hole_height = height * self.size * P\n          hole_width = width * self.size * P\n          hole_height = int(hole_height)\n          hole_width = int(hole_width)\n          y1 = random.randint(0, int(height - hole_height))\n          x1 = random.randint(0, int(width- hole_width))\n          y2 = y1 + hole_height\n          x2 = x1 + hole_width\n\n          one = image[y1:y2,0:x1,:]\n          two = tf.zeros([y2-y1,x2-x1,1], dtype=tf.float32) \n          three = image[y1:y2,x2:width,:]\n          middle = tf.concat([one,two,three],axis=1)\n          image = tf.concat([image[0:y1,:,:][tf.newaxis,...],middle[tf.newaxis,...],image[y2:height,:,:][tf.newaxis,...]],axis=1)[0]\n\n      image = tf.cast(image, dtype=tf.float32)\n      return image\n\nclass RandomBrightness:\n    def __init__(self, magnitude=0.7):\n        self.magnitude = magnitude\n\n  def __call__(self, image):\n      P = tf.random.uniform([],-1,1,dtype=tf.float32)\n      image = tf.add(P*self.magnitude, image)\n      return image\n</code></pre>\n<pre><code>def mixup(image, label, PROBABILITY = 1.0):\n    # initial image size : [256,819]\n    AUG_BATCH = CFG.BATCH_SIZE\n    DIM1 = CFG.OBJ_HEIGHT\n    DIM2 = CFG.OBJ_WIDTH\n    CLASSES = CFG.NUMBER_OF_CLASSES\n\n    imgs = []; labs = []\n    image = tf.image.resize(image, size=(DIM1, DIM2))\n\n\n    coarse = CoarseDropout(CFG.COARSE, size=0.02)\n    brightness = RandomBrightness(magnitude=1.0)\n    randomcrop = RandomResizedCrop(CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, scale=(0.9, 1.0))\n\n    for j in range(AUG_BATCH):\n\n        k = tf.cast( tf.random.uniform([],0,AUG_BATCH),tf.int32)\n        mixup_lambda = random.uniform(0.3,1.2)\n        a = tf.cast(np.random.beta(mixup_lambda,mixup_lambda), dtype=tf.float32)\n        img1 = image[j,]\n        img2 = image[k,]\n        img1 = brightness(img1)\n        img2 = brightness(img2)\n        img1 = randomcrop(img1)\n        img2 = randomcrop(img2)\n\n        mixup_image = (1-a)*img1 + a*img2\n        p_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        p_v_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        if p_flip &gt; 0.5:\n            mixup_image = tf.image.flip_left_right(mixup_image)\n        if p_v_flip &gt; 0.5:\n            mixup_image = tf.image.flip_up_down(mixup_image)\n        mixup_image = coarse(mixup_image)\n\n        mixup_image = tf.cast(mixup_image, tf.float32)\n        imgs.append(mixup_image)\n\n        lab1 = label[j,]\n        lab2 = label[k,]\n        labs.append((1-a)*lab1 + a*lab2)\n\n    image2 = tf.reshape(tf.stack(imgs),(AUG_BATCH,CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH,CFG.CHANNELS))\n    label2 = tf.reshape(tf.stack(labs),(AUG_BATCH,CLASSES))\n    return image2,label2\n</code></pre>",
          "rawMarkdown": "This code is from your MIXUP CODE(FLOWER TPU), i changed some\n```\nimport random\nclass CoarseDropout:\n  def __init__(self, max_holes=32, size=0.06):\n    self.size = size\n    self.max_holes = max_holes\n\n  def __call__(self, image):\n      P = random.uniform(0,1)\n      height = image.shape[0]\n      width = image.shape[1]\n      for _n in range(self.max_holes):\n          hole_height = height * self.size * P\n          hole_width = width * self.size * P\n          hole_height = int(hole_height)\n          hole_width = int(hole_width)\n          y1 = random.randint(0, int(height - hole_height))\n          x1 = random.randint(0, int(width- hole_width))\n          y2 = y1 + hole_height\n          x2 = x1 + hole_width\n        \n          one = image[y1:y2,0:x1,:]\n          two = tf.zeros([y2-y1,x2-x1,1], dtype=tf.float32) \n          three = image[y1:y2,x2:width,:]\n          middle = tf.concat([one,two,three],axis=1)\n          image = tf.concat([image[0:y1,:,:][tf.newaxis,...],middle[tf.newaxis,...],image[y2:height,:,:][tf.newaxis,...]],axis=1)[0]\n          \n      image = tf.cast(image, dtype=tf.float32)\n      return image\n\nclass RandomBrightness:\n    def __init__(self, magnitude=0.7):\n        self.magnitude = magnitude\n\n  def __call__(self, image):\n      P = tf.random.uniform([],-1,1,dtype=tf.float32)\n      image = tf.add(P*self.magnitude, image)\n      return image\n```\n\n\n```\ndef mixup(image, label, PROBABILITY = 1.0):\n    # initial image size : [256,819]\n    AUG_BATCH = CFG.BATCH_SIZE\n    DIM1 = CFG.OBJ_HEIGHT\n    DIM2 = CFG.OBJ_WIDTH\n    CLASSES = CFG.NUMBER_OF_CLASSES\n    \n    imgs = []; labs = []\n    image = tf.image.resize(image, size=(DIM1, DIM2))\n\n    \n    coarse = CoarseDropout(CFG.COARSE, size=0.02)\n    brightness = RandomBrightness(magnitude=1.0)\n    randomcrop = RandomResizedCrop(CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, scale=(0.9, 1.0))\n\n    for j in range(AUG_BATCH):\n\n        k = tf.cast( tf.random.uniform([],0,AUG_BATCH),tf.int32)\n        mixup_lambda = random.uniform(0.3,1.2)\n        a = tf.cast(np.random.beta(mixup_lambda,mixup_lambda), dtype=tf.float32)\n        img1 = image[j,]\n        img2 = image[k,]\n        img1 = brightness(img1)\n        img2 = brightness(img2)\n        img1 = randomcrop(img1)\n        img2 = randomcrop(img2)\n        \n        mixup_image = (1-a)*img1 + a*img2\n        p_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        p_v_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        if p_flip > 0.5:\n            mixup_image = tf.image.flip_left_right(mixup_image)\n        if p_v_flip > 0.5:\n            mixup_image = tf.image.flip_up_down(mixup_image)\n        mixup_image = coarse(mixup_image)\n\n        mixup_image = tf.cast(mixup_image, tf.float32)\n        imgs.append(mixup_image)\n\n        lab1 = label[j,]\n        lab2 = label[k,]\n        labs.append((1-a)*lab1 + a*lab2)\n            \n    image2 = tf.reshape(tf.stack(imgs),(AUG_BATCH,CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH,CFG.CHANNELS))\n    label2 = tf.reshape(tf.stack(labs),(AUG_BATCH,CLASSES))\n    return image2,label2\n```",
          "votes": 2
        },
        {
          "id": 1483444,
          "postDate": "2021-08-20T16:36:13.227Z",
          "content": "<p>Your cv is almost the same as mine. my cv scores were cluttered at [0.886,0.895]. I also experienced similar things like your cv (my high cv 905, but LB 76x) when I added this augmentation code in  TTA. (training is okay and improves cv a lot but when I added it to TTA, it degraded the LB score. As soon as I removed this method FROM TTA(but included in Training), LB score became normal). In my case, This was the point where the difference between CV scores and LB scores occurred.</p>\n<pre><code>import random\nclass RandomRollingShiftX:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP &gt; 0.5): # LEFT SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[:,RollStart:CFG.OBJ_WIDTH,:], image[:,0:RollStart,:]], axis=1)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[:,CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:], image_copy[:,RollStart:CFG.OBJ_WIDTH,:]], axis=1)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n\n\nclass RandomRollingShiftY:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      #print(RollStart)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP &gt; 0.5): # UP SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[RollStart:CFG.OBJ_WIDTH,:,:], image[0:RollStart,:,:]], axis=0)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:,:], image_copy[RollStart:CFG.OBJ_WIDTH,:,:]], axis=0)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n</code></pre>",
          "rawMarkdown": "Your cv is almost the same as mine. my cv scores were cluttered at [0.886,0.895]. I also experienced similar things like your cv (my high cv 905, but LB 76x) when I added this augmentation code in  TTA. (training is okay and improves cv a lot but when I added it to TTA, it degraded the LB score. As soon as I removed this method FROM TTA(but included in Training), LB score became normal). In my case, This was the point where the difference between CV scores and LB scores occurred.\n\n```\nimport random\nclass RandomRollingShiftX:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP > 0.5): # LEFT SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[:,RollStart:CFG.OBJ_WIDTH,:], image[:,0:RollStart,:]], axis=1)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[:,CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:], image_copy[:,RollStart:CFG.OBJ_WIDTH,:]], axis=1)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n\n\nclass RandomRollingShiftY:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      #print(RollStart)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP > 0.5): # UP SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[RollStart:CFG.OBJ_WIDTH,:,:], image[0:RollStart,:,:]], axis=0)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:,:], image_copy[RollStart:CFG.OBJ_WIDTH,:,:]], axis=0)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n```",
          "votes": 2
        },
        {
          "id": 1483493,
          "postDate": "2021-08-20T17:11:15.600Z",
          "content": "<p>Thanks so much for this information <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> I'm going to add this stuff to my current model to see if it fixes the LB score. Can you share your code for your learning scheduler and optimizer?</p>\n<pre><code>CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS\n</code></pre>\n<p>I like how you apply augmentation to each image separately before using mixup. (I augmented after mixup). That's smart, i didn't do that. Also i think your random brightness is helpful. I didn't do that. (When looking at histogram of train versus test, their brightness differs).</p>\n<p>And you use <code>tf.keras.layers.CONV2D</code> to convert from 1 channel to 3 channels. That's good. I use <code>tf.keras.layers.Concatenate()([input, input, input])</code>. But i think your method is better.</p>",
          "rawMarkdown": "Thanks so much for this information @deepkim I'm going to add this stuff to my current model to see if it fixes the LB score. Can you share your code for your learning scheduler and optimizer?\n\n    CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS\n\nI like how you apply augmentation to each image separately before using mixup. (I augmented after mixup). That's smart, i didn't do that. Also i think your random brightness is helpful. I didn't do that. (When looking at histogram of train versus test, their brightness differs).\n\nAnd you use `tf.keras.layers.CONV2D` to convert from 1 channel to 3 channels. That's good. I use `tf.keras.layers.Concatenate()([input, input, input])`. But i think your method is better.",
          "votes": 1
        },
        {
          "id": 1483523,
          "postDate": "2021-08-20T17:27:07.203Z",
          "content": "<p>thx for sharing. I forked some notebook, similar CV/LB gap. however, I only have free kaggle/colab resource, not able to try big image size. </p>",
          "rawMarkdown": "thx for sharing. I forked some notebook, similar CV/LB gap. however, I only have free kaggle/colab resource, not able to try big image size. ",
          "votes": 1
        },
        {
          "id": 1483554,
          "postDate": "2021-08-20T17:42:18.107Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> In your notebook version 19 <a href=\"https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71898164\" target=\"_blank\">here</a>, if you use <code>tf_efficientnet_b5</code> instead of <code>efficientnet_b5</code> then it will train correctly. And in your version 12 <a href=\"https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71782408\" target=\"_blank\">here</a>, if you use <code>tf_efficientnet_b4</code> it will train correctly. There are no image pretrained backbones for <code>efficientnet &gt;=4</code> without using the prefix <code>tf</code>.</p>",
          "rawMarkdown": "Hi @dragonzhang In your notebook version 19 [here][1], if you use `tf_efficientnet_b5` instead of `efficientnet_b5` then it will train correctly. And in your version 12 [here][2], if you use `tf_efficientnet_b4` it will train correctly. There are no image pretrained backbones for `efficientnet >=4` without using the prefix `tf`.\n\n[1]: https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71898164\n[2]: https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71782408",
          "votes": 2
        },
        {
          "id": 1483648,
          "postDate": "2021-08-20T18:45:02.220Z",
          "content": "<p>this is CosineDecayRAdam (i inherited RectifiedAdam class and changed some)</p>\n<pre><code>from tensorflow_addons.utils.types import FloatTensorLike\nfrom typing import Union, Callable, Dict\nfrom typeguard import typechecked\n\n\nclass CosineDecayRAdam(tfa.optimizers.RectifiedAdam):\n    def _resource_apply_dense(self, grad, var):\n        var_dtype = var.dtype.base_dtype\n        lr_t = self._decayed_lr(var_dtype)\n        wd_t = self._decayed_wd(var_dtype)\n        m = self.get_slot(var, \"m\")\n        v = self.get_slot(var, \"v\")\n        beta_1_t = self._get_hyper(\"beta_1\", var_dtype)\n        beta_2_t = self._get_hyper(\"beta_2\", var_dtype)\n        epsilon_t = tf.convert_to_tensor(self.epsilon, var_dtype)\n        local_step = tf.cast(self.iterations + 1, var_dtype)\n        beta_1_power = tf.pow(beta_1_t, local_step)\n        beta_2_power = tf.pow(beta_2_t, local_step)\n\n        if self._initial_total_steps &gt; 0:\n            total_steps = self._get_hyper(\"total_steps\", var_dtype)\n            warmup_steps = total_steps * self._get_hyper(\"warmup_proportion\", var_dtype)\n            min_lr = self._get_hyper(\"min_lr\", var_dtype)\n            decay_steps = tf.maximum(total_steps - warmup_steps, 1)\n            decay_rate = (min_lr - lr_t) / decay_steps\n            pi = tf.constant(3.141592)\n            cos = tf.math.cos(pi * ((local_step - warmup_steps) / (total_steps - warmup_steps))) + tf.constant(1.)\n            lr_t = tf.where(\n                local_step &lt;= warmup_steps,\n                lr_t * (local_step / warmup_steps),\n                min_lr + (lr_t - min_lr) / 2. * cos\n            )\n\n        sma_inf = 2.0 / (1.0 - beta_2_t) - 1.0\n        sma_t = sma_inf - 2.0 * local_step * beta_2_power / (1.0 - beta_2_power)\n\n        m_t = m.assign(\n            beta_1_t * m + (1.0 - beta_1_t) * grad, use_locking=self._use_locking\n        )\n        m_corr_t = m_t / (1.0 - beta_1_power)\n\n        v_t = v.assign(\n            beta_2_t * v + (1.0 - beta_2_t) * tf.square(grad),\n            use_locking=self._use_locking,\n        )\n        if self.amsgrad:\n            vhat = self.get_slot(var, \"vhat\")\n            vhat_t = vhat.assign(tf.maximum(vhat, v_t), use_locking=self._use_locking)\n            v_corr_t = tf.sqrt(vhat_t / (1.0 - beta_2_power))\n        else:\n            vhat_t = None\n            v_corr_t = tf.sqrt(v_t / (1.0 - beta_2_power))\n\n        r_t = tf.sqrt(\n            (sma_t - 4.0)\n            / (sma_inf - 4.0)\n            * (sma_t - 2.0)\n            / (sma_inf - 2.0)\n            * sma_inf\n            / sma_t\n        )\n\n        sma_threshold = self._get_hyper(\"sma_threshold\", var_dtype)\n        var_t = tf.where(\n            sma_t &gt;= sma_threshold, r_t * m_corr_t / (v_corr_t + epsilon_t), m_corr_t\n        )\n\n        if self._has_weight_decay:\n            var_t += wd_t * var\n\n        var_update = var.assign_sub(lr_t * var_t, use_locking=self._use_locking)\n\n        updates = [var_update, m_t, v_t]\n        if self.amsgrad:\n            updates.append(vhat_t)\n        return tf.group(*updates)\n</code></pre>\n<p>My Model</p>\n<pre><code>def get_model2(NET): \n        inp = tf.keras.layers.Input(shape = (CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH, CFG.CHANNELS), name = 'inp1')\n        x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)\n        effnet = effnets[NET](weights = 'noisy-student', include_top = False, pooling='avg')\n\n        x0 = effnet(x)\n\n        x = tf.keras.layers.Dropout(0.15)(x0)\n        x = tf.keras.layers.Dense(CFG.NUMBER_OF_CLASSES, activation='sigmoid', dtype='float32')(x)\n\n        model = tf.keras.models.Model(inputs = inp, outputs = x)\n        opt = CosineDecayRAdam(learning_rate=CFG.LEARNING_RATE, total_steps=int(STEPS_PER_EPOCH*CFG.EPOCHS), warmup_proportion=0.1, min_lr=5e-6)        \n        opt = tfa.optimizers.Lookahead(opt)\n        model.compile(\n            optimizer = opt,\n            loss = 'binary_crossentropy',\n            metrics = [tf.keras.metrics.AUC()]\n            ) \n\n        return model\n</code></pre>",
          "rawMarkdown": "this is CosineDecayRAdam (i inherited RectifiedAdam class and changed some)\n\n```\nfrom tensorflow_addons.utils.types import FloatTensorLike\nfrom typing import Union, Callable, Dict\nfrom typeguard import typechecked\n \n \nclass CosineDecayRAdam(tfa.optimizers.RectifiedAdam):\n    def _resource_apply_dense(self, grad, var):\n        var_dtype = var.dtype.base_dtype\n        lr_t = self._decayed_lr(var_dtype)\n        wd_t = self._decayed_wd(var_dtype)\n        m = self.get_slot(var, \"m\")\n        v = self.get_slot(var, \"v\")\n        beta_1_t = self._get_hyper(\"beta_1\", var_dtype)\n        beta_2_t = self._get_hyper(\"beta_2\", var_dtype)\n        epsilon_t = tf.convert_to_tensor(self.epsilon, var_dtype)\n        local_step = tf.cast(self.iterations + 1, var_dtype)\n        beta_1_power = tf.pow(beta_1_t, local_step)\n        beta_2_power = tf.pow(beta_2_t, local_step)\n \n        if self._initial_total_steps > 0:\n            total_steps = self._get_hyper(\"total_steps\", var_dtype)\n            warmup_steps = total_steps * self._get_hyper(\"warmup_proportion\", var_dtype)\n            min_lr = self._get_hyper(\"min_lr\", var_dtype)\n            decay_steps = tf.maximum(total_steps - warmup_steps, 1)\n            decay_rate = (min_lr - lr_t) / decay_steps\n            pi = tf.constant(3.141592)\n            cos = tf.math.cos(pi * ((local_step - warmup_steps) / (total_steps - warmup_steps))) + tf.constant(1.)\n            lr_t = tf.where(\n                local_step <= warmup_steps,\n                lr_t * (local_step / warmup_steps),\n                min_lr + (lr_t - min_lr) / 2. * cos\n            )\n \n        sma_inf = 2.0 / (1.0 - beta_2_t) - 1.0\n        sma_t = sma_inf - 2.0 * local_step * beta_2_power / (1.0 - beta_2_power)\n \n        m_t = m.assign(\n            beta_1_t * m + (1.0 - beta_1_t) * grad, use_locking=self._use_locking\n        )\n        m_corr_t = m_t / (1.0 - beta_1_power)\n \n        v_t = v.assign(\n            beta_2_t * v + (1.0 - beta_2_t) * tf.square(grad),\n            use_locking=self._use_locking,\n        )\n        if self.amsgrad:\n            vhat = self.get_slot(var, \"vhat\")\n            vhat_t = vhat.assign(tf.maximum(vhat, v_t), use_locking=self._use_locking)\n            v_corr_t = tf.sqrt(vhat_t / (1.0 - beta_2_power))\n        else:\n            vhat_t = None\n            v_corr_t = tf.sqrt(v_t / (1.0 - beta_2_power))\n \n        r_t = tf.sqrt(\n            (sma_t - 4.0)\n            / (sma_inf - 4.0)\n            * (sma_t - 2.0)\n            / (sma_inf - 2.0)\n            * sma_inf\n            / sma_t\n        )\n \n        sma_threshold = self._get_hyper(\"sma_threshold\", var_dtype)\n        var_t = tf.where(\n            sma_t >= sma_threshold, r_t * m_corr_t / (v_corr_t + epsilon_t), m_corr_t\n        )\n \n        if self._has_weight_decay:\n            var_t += wd_t * var\n \n        var_update = var.assign_sub(lr_t * var_t, use_locking=self._use_locking)\n \n        updates = [var_update, m_t, v_t]\n        if self.amsgrad:\n            updates.append(vhat_t)\n        return tf.group(*updates)\n```\n\nMy Model\n\n```\ndef get_model2(NET): \n        inp = tf.keras.layers.Input(shape = (CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH, CFG.CHANNELS), name = 'inp1')\n        x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)\n        effnet = effnets[NET](weights = 'noisy-student', include_top = False, pooling='avg')\n        \n        x0 = effnet(x)\n        \n        x = tf.keras.layers.Dropout(0.15)(x0)\n        x = tf.keras.layers.Dense(CFG.NUMBER_OF_CLASSES, activation='sigmoid', dtype='float32')(x)\n \n        model = tf.keras.models.Model(inputs = inp, outputs = x)\n        opt = CosineDecayRAdam(learning_rate=CFG.LEARNING_RATE, total_steps=int(STEPS_PER_EPOCH*CFG.EPOCHS), warmup_proportion=0.1, min_lr=5e-6)        \n        opt = tfa.optimizers.Lookahead(opt)\n        model.compile(\n            optimizer = opt,\n            loss = 'binary_crossentropy',\n            metrics = [tf.keras.metrics.AUC()]\n            ) \n        \n        return model\n```",
          "votes": 4
        }
      ]
    },
    {
      "id": 1483538,
      "postDate": "2021-08-20T17:35:18.347Z",
      "content": "<p>I've inference code on Kaggle for single model reaching 0.780. I can share it but I've not the training part uploaded.</p>",
      "rawMarkdown": "I've inference code on Kaggle for single model reaching 0.780. I can share it but I've not the training part uploaded.",
      "votes": 2,
      "replies": [
        {
          "id": 1484580,
          "postDate": "2021-08-21T12:45:27.410Z",
          "content": "<p>Thanks Mpware. I won't need it. I have figured out why my model has a large CV LB gap using Statking's suggestions below.</p>",
          "rawMarkdown": "Thanks Mpware. I won't need it. I have figured out why my model has a large CV LB gap using Statking's suggestions below.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1483417,
      "postDate": "2021-08-20T16:11:23.787Z",
      "content": "<p>That code is the actual script I was using, copied and pasted into a notebook for readability. So it may not work in the notebook. (I also made some changes around omegaconf)<br>\nI also used this Docker image as the execution environment.<br>\n<a href=\"https://hub.docker.com/r/hirune924/pikachu\" target=\"_blank\">https://hub.docker.com/r/hirune924/pikachu</a></p>",
      "rawMarkdown": "That code is the actual script I was using, copied and pasted into a notebook for readability. So it may not work in the notebook. (I also made some changes around omegaconf)\nI also used this Docker image as the execution environment.\nhttps://hub.docker.com/r/hirune924/pikachu",
      "replies": [
        {
          "id": 1484581,
          "postDate": "2021-08-21T12:45:46.643Z",
          "content": "<p>Thanks         </p>",
          "rawMarkdown": "Thanks         "
        }
      ]
    },
    {
      "id": 1629906,
      "postDate": "2021-12-26T17:59:20.060Z",
      "content": "<p>It depends on how you transform data the best strategy is to use interrelated features to create new features such A*B = AB in order to achieve the highest possible accuracy and cross validation results</p>",
      "rawMarkdown": "It depends on how you transform data the best strategy is to use interrelated features to create new features such A*B = AB in order to achieve the highest possible accuracy and cross validation results"
    }
  ],
  "comments": [
    {
      "id": 1483346,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2021-08-20T15:34:14.217000",
      "content": "<p>What you already tried Chris? </p>\n<p>Things that helps score close to 0.79:  double frequency axis size from 256-&gt;512(enlarge lines),  mixup, dropout, weightedsampling (increase to 20-30% positives during training), blend of 5 folds with TTA(hflip, vflip, hvflip, resize), train with my target generator applied randomly to control images (this double the examples to train). </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1483542,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T17:37:07.400000",
          "content": "<p>I built 100+ different models in the past week! I thought I tried everything 😄</p>\n<p>I'm thinking that i implemented something wrong. I plan to revisit my mixup experiments which did not increase CV but I believe should. I will also try some of Statking's suggestions above to see if they close my CV LB gap.</p>\n<p>I like your idea of upsampling positive. I will try that thanks. I noticed that train is 10% positive and test is 20% positive, so perhaps that will help. (However with AUC, i'm not sure).</p>\n<p>Eventually, i will try your generator. I like that, but first i want to achieve LB 0.780 without external data nor preprocess.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483651,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-20T18:45:45.357000",
          "content": "<p>But why limit yourself to new data? Old data gave us very significant boosts.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483802,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T21:52:37",
          "content": "<p>At this point, i just want to know what my model is missing and why my CV LB gap is large. I think there is an important learning lesson for me here.</p>\n<p>During the competition, i was planning to try old data but there was too many other things to explore in my short time.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1484343,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-08-21T08:28:26.560000",
          "content": "<p>Makes sense, I am quite sure that using mixup is important.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1484611,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T13:00:14.727000",
          "content": "<p>I think you are right. I am sure the secret (to closing CV LB gap) is to teach the model to learn the signal and ignore the background like we discussed in the other thread. Mixup (with max target) seems important to accomplish this.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1484601,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-21T12:56:55.620000",
      "content": "<p>UPDATE: Thanks everyone for your help. I'm excited to say that my CV LB gap is closing by using the suggestions in this thread. I made 8 changes to my model and it boosted CV and boosted LB and closed the gap! I believe using 5 folds (with small backbone) or 1 fold (with large backbone) then the single model will achieve <code>Silver Medal</code>! </p>\n<p>I am now making one change at a time to see which is the most important. Here is the list of 6 changes I made</p>\n<ul>\n<li>Change <code>20 epochs cosine without warmup</code> to <code>40 epochs cosine with warmup</code>.</li>\n<li>Use  <code>x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)</code> to convert 1 channel to 3 channel instead of <code>x = tf.keras.layers.Concatenate(axis=-1)([inp,inp,inp])</code></li>\n<li>Add <code>dropout(0.15)</code> between efficientnet embeddings and final sigmoid output</li>\n<li>Add albumentations <code>coarse dropout with max_holes=32, max_height=size/10</code> </li>\n<li>Add mixup within batch with <code>alpha=3</code></li>\n<li>After <code>np.vstack()</code> apply normalize with subtract mean, divide std.</li>\n<li>Use random brightness augmentation, i.e. <code>img += np.random.uniform(-1,1)</code></li>\n<li>Apply augmentation to image separately before mixup instead of after mixup</li>\n</ul>\n<p>I will also display GradCam with my old model and new model to see what the model is learning differently.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1487186,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-23T13:18:12.970000",
      "content": "<h3>Mystery solved! Mixup can achieve LB 0.780+!</h3>\n<p>(and large backbone with large image size)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make to beat LB 780 and get Silver medal. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1484197,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-08-21T05:50:24.307000",
      "content": "<p>Maybe check out this silver post by <a href=\"https://www.kaggle.com/blankaf\" target=\"_blank\">@blankaf</a> <br>\n<a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\" target=\"_blank\">https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884</a></p>\n<p>in the end my best model was single EfficientNet-B4 with non-symmetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\"</p>\n<p>not sure if the code is available or could be, but perhaps the non-symmetric cropping idea? had not see that before. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1484947,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T17:14:57.393000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/something4kag\" target=\"_blank\">@something4kag</a> . What was the public score for your EB4 that had private LB 0.78191? And what was the public and private scores of the final submission you did select?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1485657,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2021-08-22T10:28:35.320000",
          "content": "<p>Chris, <br>\nyou are actually asking about my solutions <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\" target=\"_blank\">here</a>.<br>\nAs I have written there, I trained (with the same data \"tricks\") 2 models: Resnet34d and EfficientNet-B4. For my finals I chose 2 different blends of them, the better which scored me 47th private was private LB 0.78172 and public LB 0.78618. Apparently all blends or models' (hyper)parameters combinations where not very significant for private LB results because I had 26 submissions scoring over 0.780 private. All the submissions with Resnet34d were overfitting the public LB (as you see about 0.005) but all \"pure\" EfficientNet-B4 had nearly equal private and public LB score. My mentioned best solution even had better private: 0.78191 than public: 0.78175. <br>\nSo the main trick was really in the data processing: non-symmetric cropping and training with old positive samples. And EfficientNet was better than Resnet for better generalization. I believe that with higher EfficientNet model with higher resolution the results could be even much better.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1485741,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-22T11:35:08.593000",
          "content": "<p>Thanks. That's weird how Resnet34d overfit more. When you say \"old positive\" samples. Do you mean the train data with <code>target=1</code> from the data <strong>before</strong> competition reset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1485787,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2021-08-22T12:28:38.677000",
          "content": "<p>I exactly mean that I added cast and normalized train and test samples with target=1 from the datasets before competition reset. Finally I had 70012 samles in my new train set including 16012 samples with target=1.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1483387,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2021-08-20T15:48:39.967000",
      "content": "<p>I could easily achieve over LB 780 with a single model ( 5 FOLD SUBMISSION BLENDING).</p>\n<p><strong>Preprocess</strong> : only new data, pick 3 channels(0,2,4) and normalize by subtracting mean and by dividing std each channel<br>\nthen the size becomes [256,819]<br>\nWhen I made tfrecords, i loaded pure numpy files , and serialized it(i didn't use cv.imencode because it removes  some signals)</p>\n<p><strong>Training</strong> : At least IMAGE_SIZE [720,720]. and efficientnet &gt;= 4 with mixup lambda = 1<br>\n                      (INPUT IMAGE SIZE [720,720,1] and by tf.keras.layers.CONV2D, size became [720,720,3], after that input goes into the efficientnet),<br>\n<strong>Optimizer</strong> : CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS</p>\n<p><strong>AUGMENTATION</strong> : randomcropresize(0.9,1.0), flip vertical, horizontal, a slight coarse dropout, brightnesschange(since input channel is 1, i just added some number from uniform distribution)</p>\n<p><strong>TTA NUMS</strong> : 11 (randomcropandresize, flip vertical, horizontal, coarsedropout, brightnesschange)</p>\n<p>BUT My best single model score was LB0.78990  and i tried a lot of things, i coudn't achieve a score of LB &gt; 0.78990. with this method. (EFFICIENTNET 6, SIZE 720x720)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1483399,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T16:00:53.410000",
          "content": "<p>This is great thanks. This is nearly identical to my best model which gets CV 0.887. It uses 768x768 and EffNetB4. But for some reason my model gets LB 0.765. If it had the typical CV LB gap, then it would also achieve LB 0.780+. I need to debug this model of mine. Something must not be right.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483419,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-08-20T16:13:33.807000",
          "content": "<p>This code is from your MIXUP CODE(FLOWER TPU), i changed some</p>\n<pre><code>import random\nclass CoarseDropout:\n  def __init__(self, max_holes=32, size=0.06):\n    self.size = size\n    self.max_holes = max_holes\n\n  def __call__(self, image):\n      P = random.uniform(0,1)\n      height = image.shape[0]\n      width = image.shape[1]\n      for _n in range(self.max_holes):\n          hole_height = height * self.size * P\n          hole_width = width * self.size * P\n          hole_height = int(hole_height)\n          hole_width = int(hole_width)\n          y1 = random.randint(0, int(height - hole_height))\n          x1 = random.randint(0, int(width- hole_width))\n          y2 = y1 + hole_height\n          x2 = x1 + hole_width\n\n          one = image[y1:y2,0:x1,:]\n          two = tf.zeros([y2-y1,x2-x1,1], dtype=tf.float32) \n          three = image[y1:y2,x2:width,:]\n          middle = tf.concat([one,two,three],axis=1)\n          image = tf.concat([image[0:y1,:,:][tf.newaxis,...],middle[tf.newaxis,...],image[y2:height,:,:][tf.newaxis,...]],axis=1)[0]\n\n      image = tf.cast(image, dtype=tf.float32)\n      return image\n\nclass RandomBrightness:\n    def __init__(self, magnitude=0.7):\n        self.magnitude = magnitude\n\n  def __call__(self, image):\n      P = tf.random.uniform([],-1,1,dtype=tf.float32)\n      image = tf.add(P*self.magnitude, image)\n      return image\n</code></pre>\n<pre><code>def mixup(image, label, PROBABILITY = 1.0):\n    # initial image size : [256,819]\n    AUG_BATCH = CFG.BATCH_SIZE\n    DIM1 = CFG.OBJ_HEIGHT\n    DIM2 = CFG.OBJ_WIDTH\n    CLASSES = CFG.NUMBER_OF_CLASSES\n\n    imgs = []; labs = []\n    image = tf.image.resize(image, size=(DIM1, DIM2))\n\n\n    coarse = CoarseDropout(CFG.COARSE, size=0.02)\n    brightness = RandomBrightness(magnitude=1.0)\n    randomcrop = RandomResizedCrop(CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, scale=(0.9, 1.0))\n\n    for j in range(AUG_BATCH):\n\n        k = tf.cast( tf.random.uniform([],0,AUG_BATCH),tf.int32)\n        mixup_lambda = random.uniform(0.3,1.2)\n        a = tf.cast(np.random.beta(mixup_lambda,mixup_lambda), dtype=tf.float32)\n        img1 = image[j,]\n        img2 = image[k,]\n        img1 = brightness(img1)\n        img2 = brightness(img2)\n        img1 = randomcrop(img1)\n        img2 = randomcrop(img2)\n\n        mixup_image = (1-a)*img1 + a*img2\n        p_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        p_v_flip = tf.random.uniform([], 0, 1, dtype=tf.float32)\n        if p_flip &gt; 0.5:\n            mixup_image = tf.image.flip_left_right(mixup_image)\n        if p_v_flip &gt; 0.5:\n            mixup_image = tf.image.flip_up_down(mixup_image)\n        mixup_image = coarse(mixup_image)\n\n        mixup_image = tf.cast(mixup_image, tf.float32)\n        imgs.append(mixup_image)\n\n        lab1 = label[j,]\n        lab2 = label[k,]\n        labs.append((1-a)*lab1 + a*lab2)\n\n    image2 = tf.reshape(tf.stack(imgs),(AUG_BATCH,CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH,CFG.CHANNELS))\n    label2 = tf.reshape(tf.stack(labs),(AUG_BATCH,CLASSES))\n    return image2,label2\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483444,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-08-20T16:36:13.227000",
          "content": "<p>Your cv is almost the same as mine. my cv scores were cluttered at [0.886,0.895]. I also experienced similar things like your cv (my high cv 905, but LB 76x) when I added this augmentation code in  TTA. (training is okay and improves cv a lot but when I added it to TTA, it degraded the LB score. As soon as I removed this method FROM TTA(but included in Training), LB score became normal). In my case, This was the point where the difference between CV scores and LB scores occurred.</p>\n<pre><code>import random\nclass RandomRollingShiftX:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP &gt; 0.5): # LEFT SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[:,RollStart:CFG.OBJ_WIDTH,:], image[:,0:RollStart,:]], axis=1)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[:,CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:], image_copy[:,RollStart:CFG.OBJ_WIDTH,:]], axis=1)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n\n\nclass RandomRollingShiftY:\n\n  def __call__(self, image):\n      RollP = tf.random.uniform([],0,0.4,dtype=tf.float32)\n      RollStart = int(RollP*CFG.OBJ_WIDTH)\n      #print(RollStart)\n      Roll_XYP = tf.random.uniform([],0,1.0,dtype=tf.float32)\n      if (Roll_XYP &gt; 0.5): # UP SHIFT\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image[RollStart:CFG.OBJ_WIDTH,:,:], image[0:RollStart,:,:]], axis=0)\n      else:\n          image_copy = tf.identity(image)\n          image_copy = tf.concat([image_copy[CFG.OBJ_WIDTH-RollStart:CFG.OBJ_WIDTH,:,:], image_copy[RollStart:CFG.OBJ_WIDTH,:,:]], axis=0)\n      image_copy = tf.reshape(image_copy, [CFG.OBJ_HEIGHT, CFG.OBJ_WIDTH, CFG.CHANNELS])\n      return image_copy\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483493,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T17:11:15.600000",
          "content": "<p>Thanks so much for this information <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> I'm going to add this stuff to my current model to see if it fixes the LB score. Can you share your code for your learning scheduler and optimizer?</p>\n<pre><code>CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS\n</code></pre>\n<p>I like how you apply augmentation to each image separately before using mixup. (I augmented after mixup). That's smart, i didn't do that. Also i think your random brightness is helpful. I didn't do that. (When looking at histogram of train versus test, their brightness differs).</p>\n<p>And you use <code>tf.keras.layers.CONV2D</code> to convert from 1 channel to 3 channels. That's good. I use <code>tf.keras.layers.Concatenate()([input, input, input])</code>. But i think your method is better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483523,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2021-08-20T17:27:07.203000",
          "content": "<p>thx for sharing. I forked some notebook, similar CV/LB gap. however, I only have free kaggle/colab resource, not able to try big image size. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1483554,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T17:42:18.107000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> In your notebook version 19 <a href=\"https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71898164\" target=\"_blank\">here</a>, if you use <code>tf_efficientnet_b5</code> instead of <code>efficientnet_b5</code> then it will train correctly. And in your version 12 <a href=\"https://www.kaggle.com/dragonzhang/seti-learned-image-resizing?scriptVersionId=71782408\" target=\"_blank\">here</a>, if you use <code>tf_efficientnet_b4</code> it will train correctly. There are no image pretrained backbones for <code>efficientnet &gt;=4</code> without using the prefix <code>tf</code>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1483648,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2021-08-20T18:45:02.220000",
          "content": "<p>this is CosineDecayRAdam (i inherited RectifiedAdam class and changed some)</p>\n<pre><code>from tensorflow_addons.utils.types import FloatTensorLike\nfrom typing import Union, Callable, Dict\nfrom typeguard import typechecked\n\n\nclass CosineDecayRAdam(tfa.optimizers.RectifiedAdam):\n    def _resource_apply_dense(self, grad, var):\n        var_dtype = var.dtype.base_dtype\n        lr_t = self._decayed_lr(var_dtype)\n        wd_t = self._decayed_wd(var_dtype)\n        m = self.get_slot(var, \"m\")\n        v = self.get_slot(var, \"v\")\n        beta_1_t = self._get_hyper(\"beta_1\", var_dtype)\n        beta_2_t = self._get_hyper(\"beta_2\", var_dtype)\n        epsilon_t = tf.convert_to_tensor(self.epsilon, var_dtype)\n        local_step = tf.cast(self.iterations + 1, var_dtype)\n        beta_1_power = tf.pow(beta_1_t, local_step)\n        beta_2_power = tf.pow(beta_2_t, local_step)\n\n        if self._initial_total_steps &gt; 0:\n            total_steps = self._get_hyper(\"total_steps\", var_dtype)\n            warmup_steps = total_steps * self._get_hyper(\"warmup_proportion\", var_dtype)\n            min_lr = self._get_hyper(\"min_lr\", var_dtype)\n            decay_steps = tf.maximum(total_steps - warmup_steps, 1)\n            decay_rate = (min_lr - lr_t) / decay_steps\n            pi = tf.constant(3.141592)\n            cos = tf.math.cos(pi * ((local_step - warmup_steps) / (total_steps - warmup_steps))) + tf.constant(1.)\n            lr_t = tf.where(\n                local_step &lt;= warmup_steps,\n                lr_t * (local_step / warmup_steps),\n                min_lr + (lr_t - min_lr) / 2. * cos\n            )\n\n        sma_inf = 2.0 / (1.0 - beta_2_t) - 1.0\n        sma_t = sma_inf - 2.0 * local_step * beta_2_power / (1.0 - beta_2_power)\n\n        m_t = m.assign(\n            beta_1_t * m + (1.0 - beta_1_t) * grad, use_locking=self._use_locking\n        )\n        m_corr_t = m_t / (1.0 - beta_1_power)\n\n        v_t = v.assign(\n            beta_2_t * v + (1.0 - beta_2_t) * tf.square(grad),\n            use_locking=self._use_locking,\n        )\n        if self.amsgrad:\n            vhat = self.get_slot(var, \"vhat\")\n            vhat_t = vhat.assign(tf.maximum(vhat, v_t), use_locking=self._use_locking)\n            v_corr_t = tf.sqrt(vhat_t / (1.0 - beta_2_power))\n        else:\n            vhat_t = None\n            v_corr_t = tf.sqrt(v_t / (1.0 - beta_2_power))\n\n        r_t = tf.sqrt(\n            (sma_t - 4.0)\n            / (sma_inf - 4.0)\n            * (sma_t - 2.0)\n            / (sma_inf - 2.0)\n            * sma_inf\n            / sma_t\n        )\n\n        sma_threshold = self._get_hyper(\"sma_threshold\", var_dtype)\n        var_t = tf.where(\n            sma_t &gt;= sma_threshold, r_t * m_corr_t / (v_corr_t + epsilon_t), m_corr_t\n        )\n\n        if self._has_weight_decay:\n            var_t += wd_t * var\n\n        var_update = var.assign_sub(lr_t * var_t, use_locking=self._use_locking)\n\n        updates = [var_update, m_t, v_t]\n        if self.amsgrad:\n            updates.append(vhat_t)\n        return tf.group(*updates)\n</code></pre>\n<p>My Model</p>\n<pre><code>def get_model2(NET): \n        inp = tf.keras.layers.Input(shape = (CFG.OBJ_HEIGHT,CFG.OBJ_WIDTH, CFG.CHANNELS), name = 'inp1')\n        x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)\n        effnet = effnets[NET](weights = 'noisy-student', include_top = False, pooling='avg')\n\n        x0 = effnet(x)\n\n        x = tf.keras.layers.Dropout(0.15)(x0)\n        x = tf.keras.layers.Dense(CFG.NUMBER_OF_CLASSES, activation='sigmoid', dtype='float32')(x)\n\n        model = tf.keras.models.Model(inputs = inp, outputs = x)\n        opt = CosineDecayRAdam(learning_rate=CFG.LEARNING_RATE, total_steps=int(STEPS_PER_EPOCH*CFG.EPOCHS), warmup_proportion=0.1, min_lr=5e-6)        \n        opt = tfa.optimizers.Lookahead(opt)\n        model.compile(\n            optimizer = opt,\n            loss = 'binary_crossentropy',\n            metrics = [tf.keras.metrics.AUC()]\n            ) \n\n        return model\n</code></pre>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1483538,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-08-20T17:35:18.347000",
      "content": "<p>I've inference code on Kaggle for single model reaching 0.780. I can share it but I've not the training part uploaded.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1484580,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T12:45:27.410000",
          "content": "<p>Thanks Mpware. I won't need it. I have figured out why my model has a large CV LB gap using Statking's suggestions below.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1483417,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "2021-08-20T16:11:23.787000",
      "content": "<p>That code is the actual script I was using, copied and pasted into a notebook for readability. So it may not work in the notebook. (I also made some changes around omegaconf)<br>\nI also used this Docker image as the execution environment.<br>\n<a href=\"https://hub.docker.com/r/hirune924/pikachu\" target=\"_blank\">https://hub.docker.com/r/hirune924/pikachu</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1484581,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-21T12:45:46.643000",
          "content": "<p>Thanks         </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1629906,
      "author_name": "Muhammad Ammar Jamshed",
      "author_url": "",
      "post_date": "2021-12-26T17:59:20.060000",
      "content": "<p>It depends on how you transform data the best strategy is to use interrelated features to create new features such A*B = AB in order to achieve the highest possible accuracy and cross validation results</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1483318": "### How Does Code achieve 0.780+?\nI would love to see code that scores over LB 0.780. I would like to learn what my models are missing. \n\nDoes anyone have code, that trains with new data only without data preprocess. Where they use only \"on\" cadence `np.vstack( img[::2] )` and feed 1 channel images into a CNN? And then just uses augmentation and learning schedule and achieves LB 0.780+\n\n(I see that second place posted their code, but it throws an error and I don't want to bother them to debug it for me).\n\n### UPDATE: Mixup is needed to achieve 0.780+\nThanks everyone for sharing your code with me. I have learned that mixup is the secret ingredient to achieve Silver Medal (and large backbone with large image size)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! \n\nThe score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!",
    "1483346": "What you already tried Chris? \n\nThings that helps score close to 0.79:  double frequency axis size from 256->512(enlarge lines),  mixup, dropout, weightedsampling (increase to 20-30% positives during training), blend of 5 folds with TTA(hflip, vflip, hvflip, resize), train with my target generator applied randomly to control images (this double the examples to train). ",
    "1484601": "UPDATE: Thanks everyone for your help. I'm excited to say that my CV LB gap is closing by using the suggestions in this thread. I made 8 changes to my model and it boosted CV and boosted LB and closed the gap! I believe using 5 folds (with small backbone) or 1 fold (with large backbone) then the single model will achieve `Silver Medal`! \n\nI am now making one change at a time to see which is the most important. Here is the list of 6 changes I made\n* Change `20 epochs cosine without warmup` to `40 epochs cosine with warmup`.\n* Use  `x = tf.keras.layers.Conv2D(3,3,strides=1,padding='same')(inp)` to convert 1 channel to 3 channel instead of `x = tf.keras.layers.Concatenate(axis=-1)([inp,inp,inp])`\n* Add `dropout(0.15)` between efficientnet embeddings and final sigmoid output\n* Add albumentations `coarse dropout with max_holes=32, max_height=size/10` \n* Add mixup within batch with `alpha=3`\n* After `np.vstack()` apply normalize with subtract mean, divide std.\n* Use random brightness augmentation, i.e. `img += np.random.uniform(-1,1)`\n* Apply augmentation to image separately before mixup instead of after mixup\n\nI will also display GradCam with my old model and new model to see what the model is learning differently.",
    "1487186": "### Mystery solved! Mixup can achieve LB 0.780+! \n(and large backbone with large image size)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make to beat LB 780 and get Silver medal. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. And fold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385",
    "1484197": "Maybe check out this silver post by @blankaf \nhttps://www.kaggle.com/c/seti-breakthrough-listen/discussion/266884\n\nin the end my best model was single EfficientNet-B4 with non-symmetric version cropping, 16 epochs (1 fold training for 15 hours), private LB 0.78191 (unfortunately not chosen for final).\"\n\nnot sure if the code is available or could be, but perhaps the non-symmetric cropping idea? had not see that before. ",
    "1483387": "I could easily achieve over LB 780 with a single model ( 5 FOLD SUBMISSION BLENDING).\n\n**Preprocess** : only new data, pick 3 channels(0,2,4) and normalize by subtracting mean and by dividing std each channel\nthen the size becomes [256,819]\nWhen I made tfrecords, i loaded pure numpy files , and serialized it(i didn't use cv.imencode because it removes  some signals)\n\n**Training** : At least IMAGE_SIZE [720,720]. and efficientnet >= 4 with mixup lambda = 1\n                      (INPUT IMAGE SIZE [720,720,1] and by tf.keras.layers.CONV2D, size became [720,720,3], after that input goes into the efficientnet),\n**Optimizer** : CosineDecay RAdam(7e-4 * REPLICAS(8)), WARMUP 0.1 and decay to 1e-6, 50 EPOCHS\n\n**AUGMENTATION** : randomcropresize(0.9,1.0), flip vertical, horizontal, a slight coarse dropout, brightnesschange(since input channel is 1, i just added some number from uniform distribution)\n\n**TTA NUMS** : 11 (randomcropandresize, flip vertical, horizontal, coarsedropout, brightnesschange)\n \nBUT My best single model score was LB0.78990  and i tried a lot of things, i coudn't achieve a score of LB > 0.78990. with this method. (EFFICIENTNET 6, SIZE 720x720)\n",
    "1483538": "I've inference code on Kaggle for single model reaching 0.780. I can share it but I've not the training part uploaded.",
    "1483417": "That code is the actual script I was using, copied and pasted into a notebook for readability. So it may not work in the notebook. (I also made some changes around omegaconf)\nI also used this Docker image as the execution environment.\nhttps://hub.docker.com/r/hirune924/pikachu",
    "1629906": "It depends on how you transform data the best strategy is to use interrelated features to create new features such A*B = AB in order to achieve the highest possible accuracy and cross validation results"
  }
}