{
  "id": 265777,
  "title": "Don't follow the top-scoring kernel \"blindly\"",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/265777",
  "author_name": "Aman Arora",
  "post_date": "2021-08-16T23:17:03.833000",
  "votes": 48,
  "comment_count": 25,
  "views": 0,
  "content": "<p>So I just started playing around with this competition. And I found the top-scoring kernel <a href=\"https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type\" target=\"_blank\">Efficientnet3D with one MRI type\n</a> to be at rank-44! Amazing, right?</p>\n<p>And I was like \"WOW\"! Only to realize that the model doesn't really train at all! For example, check the training logs from the top-scoring kernel that uses <code>Conv3d</code>, the valid loss just rotates around 0.69. In fact, it goes up in the later epochs!! <strong>(So, you're actually overfitting)</strong></p>\n<p>Also, look closely and you'll see training loss doesn't really go below 0.68! So, <strong>the model isn't actually training</strong>. </p>\n<p>I suspected the model doesn't train due to zero-padding used in <code>load_dicom_images_3d</code>. The model doesn't really have a way to differentiate b/w zero-padding and the actual pixel values. (which is strange - it shouldn't be like that!)</p>\n<p>So, I thought let's fix that. One could always use \"reflective\" padding. Or ignore zero-padded channels/images when calculating loss. (Just some ideas FYI)</p>\n<p>Only to realize - that now my training loss goes down to 0.2076, but the validation loss shoots up to ~1.1.</p>\n<p>IMHO, using Conv3d - simply following the approach in the top-scoring kernel might not be the best strategy. Let's try something new perhaps?</p>\n<blockquote>\n  <p>NOTE: All loss values reported are absolute BCE loss values.</p>\n</blockquote>",
  "messages": [
    {
      "id": 1475969,
      "postDate": "2021-08-16T23:17:03.833Z",
      "content": "<p>So I just started playing around with this competition. And I found the top-scoring kernel <a href=\"https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type\" target=\"_blank\">Efficientnet3D with one MRI type\n</a> to be at rank-44! Amazing, right?</p>\n<p>And I was like \"WOW\"! Only to realize that the model doesn't really train at all! For example, check the training logs from the top-scoring kernel that uses <code>Conv3d</code>, the valid loss just rotates around 0.69. In fact, it goes up in the later epochs!! <strong>(So, you're actually overfitting)</strong></p>\n<p>Also, look closely and you'll see training loss doesn't really go below 0.68! So, <strong>the model isn't actually training</strong>. </p>\n<p>I suspected the model doesn't train due to zero-padding used in <code>load_dicom_images_3d</code>. The model doesn't really have a way to differentiate b/w zero-padding and the actual pixel values. (which is strange - it shouldn't be like that!)</p>\n<p>So, I thought let's fix that. One could always use \"reflective\" padding. Or ignore zero-padded channels/images when calculating loss. (Just some ideas FYI)</p>\n<p>Only to realize - that now my training loss goes down to 0.2076, but the validation loss shoots up to ~1.1.</p>\n<p>IMHO, using Conv3d - simply following the approach in the top-scoring kernel might not be the best strategy. Let's try something new perhaps?</p>\n<blockquote>\n  <p>NOTE: All loss values reported are absolute BCE loss values.</p>\n</blockquote>",
      "rawMarkdown": "So I just started playing around with this competition. And I found the top-scoring kernel [Efficientnet3D with one MRI type\n](https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type) to be at rank-44! Amazing, right?\n\nAnd I was like \"WOW\"! Only to realize that the model doesn't really train at all! For example, check the training logs from the top-scoring kernel that uses `Conv3d`, the valid loss just rotates around 0.69. In fact, it goes up in the later epochs!! **(So, you're actually overfitting)**\n\nAlso, look closely and you'll see training loss doesn't really go below 0.68! So, **the model isn't actually training**. \n\nI suspected the model doesn't train due to zero-padding used in `load_dicom_images_3d`. The model doesn't really have a way to differentiate b/w zero-padding and the actual pixel values. (which is strange - it shouldn't be like that!)\n\nSo, I thought let's fix that. One could always use \"reflective\" padding. Or ignore zero-padded channels/images when calculating loss. (Just some ideas FYI)\n\nOnly to realize - that now my training loss goes down to 0.2076, but the validation loss shoots up to ~1.1.\n\nIMHO, using Conv3d - simply following the approach in the top-scoring kernel might not be the best strategy. Let's try something new perhaps?\n\n> NOTE: All loss values reported are absolute BCE loss values.",
      "votes": 47
    },
    {
      "id": 1488985,
      "postDate": "2021-08-24T16:12:44.187Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a>, were you able to decrease the validation loss?<br>\nSame as you, I seem to be overfitting a lot. Validation loss does not go below 0.67 even though training loss went as low as 0.32</p>",
      "rawMarkdown": "Hey @aroraaman, were you able to decrease the validation loss?\nSame as you, I seem to be overfitting a lot. Validation loss does not go below 0.67 even though training loss went as low as 0.32",
      "votes": 5
    },
    {
      "id": 1476661,
      "postDate": "2021-08-17T07:46:00.977Z",
      "content": "<p>It's useful as a starting point. But I just assumed that this part of the code was humour on the part of kernel author, a special gift to those who just fork and submit. </p>",
      "rawMarkdown": " It's useful as a starting point. But I just assumed that this part of the code was humour on the part of kernel author, a special gift to those who just fork and submit. ",
      "votes": 6,
      "replies": [
        {
          "id": 1476838,
          "postDate": "2021-08-17T09:05:30.687Z",
          "content": "<p>Over 600 of them!</p>",
          "rawMarkdown": "Over 600 of them!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1476366,
      "postDate": "2021-08-17T05:21:15.597Z",
      "content": "<p>Hey, can you share how to do reflective padding and during calculating loss how can we ignore the zero-padded images, And your decreasing loss was only because of this or any other tricks,<br>\nThanks</p>",
      "rawMarkdown": "Hey, can you share how to do reflective padding and during calculating loss how can we ignore the zero-padded images, And your decreasing loss was only because of this or any other tricks,\nThanks",
      "votes": 3,
      "replies": [
        {
          "id": 1477389,
          "postDate": "2021-08-17T13:11:02.627Z",
          "content": "<p>Sure. So you could maybe add the image slices (in reverse to make it reflective) instead of zero padding. </p>\n<p>Basically, if the 3-d matrix has 20 images, and you’re trying to select all 3-d matrices of size 64, then perhaps copy the 20images 4 times and now select 64 out of 80. This gets you 3-d matrix with 64 images. </p>\n<p>Instead of like having 20 images and everything else zero padded. </p>",
          "rawMarkdown": "Sure. So you could maybe add the image slices (in reverse to make it reflective) instead of zero padding. \n\nBasically, if the 3-d matrix has 20 images, and you’re trying to select all 3-d matrices of size 64, then perhaps copy the 20images 4 times and now select 64 out of 80. This gets you 3-d matrix with 64 images. \n\nInstead of like having 20 images and everything else zero padded. ",
          "votes": 3
        },
        {
          "id": 1477390,
          "postDate": "2021-08-17T13:11:34.180Z",
          "content": "<p>I added LR scheduler and used a different learning rate!</p>",
          "rawMarkdown": "I added LR scheduler and used a different learning rate!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1476249,
      "postDate": "2021-08-17T03:55:31.300Z",
      "content": "<p>The top scoring public kernal is not very good. The public dataset is small hence noisy. You can expect a couple of notebooks to score perfect score in the next 2 months due to lb probing. Trust your CV, LB is meaningless</p>",
      "rawMarkdown": "The top scoring public kernal is not very good. The public dataset is small hence noisy. You can expect a couple of notebooks to score perfect score in the next 2 months due to lb probing. Trust your CV, LB is meaningless",
      "votes": 3,
      "replies": [
        {
          "id": 1476252,
          "postDate": "2021-08-17T03:58:07.110Z",
          "content": "<p>Absolutely! </p>",
          "rawMarkdown": "Absolutely! "
        }
      ]
    },
    {
      "id": 1535955,
      "postDate": "2021-10-06T11:02:49.373Z",
      "content": "<p>Yes, you are right. This problem is really harder then fit and predict famous arch from kernel. I implemented three articles and also incorporate my ideas and I cant still beat 620 AUC validation 😄</p>",
      "rawMarkdown": "Yes, you are right. This problem is really harder then fit and predict famous arch from kernel. I implemented three articles and also incorporate my ideas and I cant still beat 620 AUC validation 😄",
      "votes": 1
    },
    {
      "id": 1505661,
      "postDate": "2021-09-07T13:09:07.843Z",
      "content": "<p>instead of \"blindly\" forking…would adding some preprocessing improve the training? if yes, what preprocessing methods are recommended besides reflective padding.</p>",
      "rawMarkdown": "instead of \"blindly\" forking...would adding some preprocessing improve the training? if yes, what preprocessing methods are recommended besides reflective padding.",
      "votes": 1
    },
    {
      "id": 1503282,
      "postDate": "2021-09-05T08:17:47.523Z",
      "content": "<p>Haha, though the same when going through the public notebooks. Saw fluctuations, no training whatsoever</p>",
      "rawMarkdown": "Haha, though the same when going through the public notebooks. Saw fluctuations, no training whatsoever",
      "votes": 1
    },
    {
      "id": 1485334,
      "postDate": "2021-08-22T02:38:26.503Z",
      "content": "<p>Great findings. I tried \"reflective\" padding and LR scheduler. However, it failed, the training loss still oscillated around 0.69 which indicates no training at all. Can you share your detailed parameters if possible. Also, I wonder LB score when training loss goes down to about 0.2.</p>",
      "rawMarkdown": "Great findings. I tried \"reflective\" padding and LR scheduler. However, it failed, the training loss still oscillated around 0.69 which indicates no training at all. Can you share your detailed parameters if possible. Also, I wonder LB score when training loss goes down to about 0.2.",
      "votes": 1
    },
    {
      "id": 1478521,
      "postDate": "2021-08-18T03:27:17.807Z",
      "content": "<p>One thing I'm curious about is, how was the notebook able to reach a high score? </p>",
      "rawMarkdown": "One thing I'm curious about is, how was the notebook able to reach a high score? ",
      "votes": 1,
      "replies": [
        {
          "id": 1478532,
          "postDate": "2021-08-18T03:37:46.760Z",
          "content": "<p>as another poster pointed out, it looks like a high score but it's actually not. </p>",
          "rawMarkdown": "as another poster pointed out, it looks like a high score but it's actually not. ",
          "votes": 1
        },
        {
          "id": 1478546,
          "postDate": "2021-08-18T03:44:47.940Z",
          "content": "<p>Small public test set. </p>",
          "rawMarkdown": "Small public test set. "
        },
        {
          "id": 1480225,
          "postDate": "2021-08-18T21:40:56.897Z",
          "content": "<p>Overfit on the public dataset/lucky coincidence. </p>",
          "rawMarkdown": "Overfit on the public dataset/lucky coincidence. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1477338,
      "postDate": "2021-08-17T12:49:53.043Z",
      "content": "<p>Yes I agree, public LB scores can be very noisy. Best thing you can do to ensure position on private leaderboard is do some K-fold validation. I still think the public notebooks are useful for helping people get started so that they don't have to go off an empty notebook.</p>",
      "rawMarkdown": "Yes I agree, public LB scores can be very noisy. Best thing you can do to ensure position on private leaderboard is do some K-fold validation. I still think the public notebooks are useful for helping people get started so that they don't have to go off an empty notebook.",
      "votes": 2,
      "replies": [
        {
          "id": 1477395,
          "postDate": "2021-08-17T13:12:36.833Z",
          "content": "<p>Trust your CV. I agree :)<br>\nEspecially for this comp - public nb has 87 studies I guess?</p>",
          "rawMarkdown": "Trust your CV. I agree :)\nEspecially for this comp - public nb has 87 studies I guess?",
          "votes": 1
        },
        {
          "id": 1477546,
          "postDate": "2021-08-17T14:24:16.573Z",
          "content": "<p>What do you mean by \"studies\", do you mean forks i.e. copies of the public notebook?</p>",
          "rawMarkdown": "What do you mean by \"studies\", do you mean forks i.e. copies of the public notebook?",
          "votes": 1
        },
        {
          "id": 1477555,
          "postDate": "2021-08-17T14:28:13.227Z",
          "content": "<p>There are only 87 samples in the public test set.</p>",
          "rawMarkdown": "There are only 87 samples in the public test set.",
          "votes": 2
        },
        {
          "id": 1477582,
          "postDate": "2021-08-17T14:39:25.340Z",
          "content": "<p>Ah yes that is very small, even the training set is only 550 samples or so. </p>",
          "rawMarkdown": "Ah yes that is very small, even the training set is only 550 samples or so. \n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1476964,
      "postDate": "2021-08-17T10:07:48.367Z",
      "content": "<p>What do you mean by \"a stronger SE module\"?<br>\nEfficientNet already uses SE, do you mean improving the existing SE module by changing the reduction value?</p>",
      "rawMarkdown": "What do you mean by \"a stronger SE module\"?\nEfficientNet already uses SE, do you mean improving the existing SE module by changing the reduction value?",
      "votes": 2,
      "replies": [
        {
          "id": 1477383,
          "postDate": "2021-08-17T13:07:44.103Z",
          "content": "<p>Ah, my bad. I misunderstood the expansion ratio in SE module to be some kind of attention parameter. Thanks for pointing this out!</p>",
          "rawMarkdown": "Ah, my bad. I misunderstood the expansion ratio in SE module to be some kind of attention parameter. Thanks for pointing this out!",
          "votes": 2
        },
        {
          "id": 1477392,
          "postDate": "2021-08-17T13:11:51.320Z",
          "content": "<p>Corrected. :)</p>",
          "rawMarkdown": "Corrected. :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1488304,
      "postDate": "2021-08-24T08:06:52.353Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1488985,
      "author_name": "Pranshu15",
      "author_url": "",
      "post_date": "2021-08-24T16:12:44.187000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aroraaman\" target=\"_blank\">@aroraaman</a>, were you able to decrease the validation loss?<br>\nSame as you, I seem to be overfitting a lot. Validation loss does not go below 0.67 even though training loss went as low as 0.32</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1476661,
      "author_name": "andy jennings",
      "author_url": "",
      "post_date": "2021-08-17T07:46:00.977000",
      "content": "<p>It's useful as a starting point. But I just assumed that this part of the code was humour on the part of kernel author, a special gift to those who just fork and submit. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 1476838,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T09:05:30.687000",
          "content": "<p>Over 600 of them!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1476366,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2021-08-17T05:21:15.597000",
      "content": "<p>Hey, can you share how to do reflective padding and during calculating loss how can we ignore the zero-padded images, And your decreasing loss was only because of this or any other tricks,<br>\nThanks</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1477389,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T13:11:02.627000",
          "content": "<p>Sure. So you could maybe add the image slices (in reverse to make it reflective) instead of zero padding. </p>\n<p>Basically, if the 3-d matrix has 20 images, and you’re trying to select all 3-d matrices of size 64, then perhaps copy the 20images 4 times and now select 64 out of 80. This gets you 3-d matrix with 64 images. </p>\n<p>Instead of like having 20 images and everything else zero padded. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1477390,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T13:11:34.180000",
          "content": "<p>I added LR scheduler and used a different learning rate!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1476249,
      "author_name": "Sarthak Bhatt",
      "author_url": "",
      "post_date": "2021-08-17T03:55:31.300000",
      "content": "<p>The top scoring public kernal is not very good. The public dataset is small hence noisy. You can expect a couple of notebooks to score perfect score in the next 2 months due to lb probing. Trust your CV, LB is meaningless</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1476252,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T03:58:07.110000",
          "content": "<p>Absolutely! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1535955,
      "author_name": "Sergey Bryansky",
      "author_url": "",
      "post_date": "2021-10-06T11:02:49.373000",
      "content": "<p>Yes, you are right. This problem is really harder then fit and predict famous arch from kernel. I implemented three articles and also incorporate my ideas and I cant still beat 620 AUC validation 😄</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1505661,
      "author_name": "hagid chida",
      "author_url": "",
      "post_date": "2021-09-07T13:09:07.843000",
      "content": "<p>instead of \"blindly\" forking…would adding some preprocessing improve the training? if yes, what preprocessing methods are recommended besides reflective padding.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1503282,
      "author_name": "Md Zarif Ul Alam",
      "author_url": "",
      "post_date": "2021-09-05T08:17:47.523000",
      "content": "<p>Haha, though the same when going through the public notebooks. Saw fluctuations, no training whatsoever</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1485334,
      "author_name": "aaronjie",
      "author_url": "",
      "post_date": "2021-08-22T02:38:26.503000",
      "content": "<p>Great findings. I tried \"reflective\" padding and LR scheduler. However, it failed, the training loss still oscillated around 0.69 which indicates no training at all. Can you share your detailed parameters if possible. Also, I wonder LB score when training loss goes down to about 0.2.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1478521,
      "author_name": "Jaechan Lee",
      "author_url": "",
      "post_date": "2021-08-18T03:27:17.807000",
      "content": "<p>One thing I'm curious about is, how was the notebook able to reach a high score? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1478532,
          "author_name": "andy jennings",
          "author_url": "",
          "post_date": "2021-08-18T03:37:46.760000",
          "content": "<p>as another poster pointed out, it looks like a high score but it's actually not. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1478546,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-18T03:44:47.940000",
          "content": "<p>Small public test set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1480225,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-18T21:40:56.897000",
          "content": "<p>Overfit on the public dataset/lucky coincidence. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1477338,
      "author_name": "Daniel Chen",
      "author_url": "",
      "post_date": "2021-08-17T12:49:53.043000",
      "content": "<p>Yes I agree, public LB scores can be very noisy. Best thing you can do to ensure position on private leaderboard is do some K-fold validation. I still think the public notebooks are useful for helping people get started so that they don't have to go off an empty notebook.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1477395,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T13:12:36.833000",
          "content": "<p>Trust your CV. I agree :)<br>\nEspecially for this comp - public nb has 87 studies I guess?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1477546,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-17T14:24:16.573000",
          "content": "<p>What do you mean by \"studies\", do you mean forks i.e. copies of the public notebook?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1477555,
          "author_name": "novice03",
          "author_url": "",
          "post_date": "2021-08-17T14:28:13.227000",
          "content": "<p>There are only 87 samples in the public test set.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1477582,
          "author_name": "Daniel Chen",
          "author_url": "",
          "post_date": "2021-08-17T14:39:25.340000",
          "content": "<p>Ah yes that is very small, even the training set is only 550 samples or so. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1476964,
      "author_name": "Pranshu15",
      "author_url": "",
      "post_date": "2021-08-17T10:07:48.367000",
      "content": "<p>What do you mean by \"a stronger SE module\"?<br>\nEfficientNet already uses SE, do you mean improving the existing SE module by changing the reduction value?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1477383,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T13:07:44.103000",
          "content": "<p>Ah, my bad. I misunderstood the expansion ratio in SE module to be some kind of attention parameter. Thanks for pointing this out!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1477392,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2021-08-17T13:11:51.320000",
          "content": "<p>Corrected. :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1488304,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-24T08:06:52.353000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1475969": "So I just started playing around with this competition. And I found the top-scoring kernel [Efficientnet3D with one MRI type\n](https://www.kaggle.com/rluethy/efficientnet3d-with-one-mri-type) to be at rank-44! Amazing, right?\n\nAnd I was like \"WOW\"! Only to realize that the model doesn't really train at all! For example, check the training logs from the top-scoring kernel that uses `Conv3d`, the valid loss just rotates around 0.69. In fact, it goes up in the later epochs!! **(So, you're actually overfitting)**\n\nAlso, look closely and you'll see training loss doesn't really go below 0.68! So, **the model isn't actually training**. \n\nI suspected the model doesn't train due to zero-padding used in `load_dicom_images_3d`. The model doesn't really have a way to differentiate b/w zero-padding and the actual pixel values. (which is strange - it shouldn't be like that!)\n\nSo, I thought let's fix that. One could always use \"reflective\" padding. Or ignore zero-padded channels/images when calculating loss. (Just some ideas FYI)\n\nOnly to realize - that now my training loss goes down to 0.2076, but the validation loss shoots up to ~1.1.\n\nIMHO, using Conv3d - simply following the approach in the top-scoring kernel might not be the best strategy. Let's try something new perhaps?\n\n> NOTE: All loss values reported are absolute BCE loss values.",
    "1488985": "Hey @aroraaman, were you able to decrease the validation loss?\nSame as you, I seem to be overfitting a lot. Validation loss does not go below 0.67 even though training loss went as low as 0.32",
    "1476661": " It's useful as a starting point. But I just assumed that this part of the code was humour on the part of kernel author, a special gift to those who just fork and submit. ",
    "1476366": "Hey, can you share how to do reflective padding and during calculating loss how can we ignore the zero-padded images, And your decreasing loss was only because of this or any other tricks,\nThanks",
    "1476249": "The top scoring public kernal is not very good. The public dataset is small hence noisy. You can expect a couple of notebooks to score perfect score in the next 2 months due to lb probing. Trust your CV, LB is meaningless",
    "1535955": "Yes, you are right. This problem is really harder then fit and predict famous arch from kernel. I implemented three articles and also incorporate my ideas and I cant still beat 620 AUC validation 😄",
    "1505661": "instead of \"blindly\" forking...would adding some preprocessing improve the training? if yes, what preprocessing methods are recommended besides reflective padding.",
    "1503282": "Haha, though the same when going through the public notebooks. Saw fluctuations, no training whatsoever",
    "1485334": "Great findings. I tried \"reflective\" padding and LR scheduler. However, it failed, the training loss still oscillated around 0.69 which indicates no training at all. Can you share your detailed parameters if possible. Also, I wonder LB score when training loss goes down to about 0.2.",
    "1478521": "One thing I'm curious about is, how was the notebook able to reach a high score? ",
    "1477338": "Yes I agree, public LB scores can be very noisy. Best thing you can do to ensure position on private leaderboard is do some K-fold validation. I still think the public notebooks are useful for helping people get started so that they don't have to go off an empty notebook.",
    "1476964": "What do you mean by \"a stronger SE module\"?\nEfficientNet already uses SE, do you mean improving the existing SE module by changing the reduction value?",
    "1488304": ""
  }
}