{
  "id": 165070,
  "title": "Low Single Fold CV/LB with Keras",
  "url": "/competitions/alaska2-image-steganalysis/discussion/165070",
  "author_name": "",
  "post_date": "2020-07-08T13:04:46.767161400Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>I reproduced the highest score pytorch public efficientnet-b3 notebook on Keras. Augmentation includes the same horizontal and vertical flips. I tried label smoothing but made things worse. I also tried reduceOnPlanteau but didn't work either. Now I'm using custom stepLR and normal crossentropy. However, my single fold model score with even b5 and b6 are lower than the notebook. </p>\n\n<p>Did you guys encountered the same problem in Keras?</p>\n\n<p>Also, if you have a single fold small model score with Keras greater than the LB of current highest public notebook (0.922), please contact me to team up! I am using 5 TPUv3x8, meaning that I can do 5 experiments at the same time with giant models to get good results very very quickly. I need some good insights on hyperparameters, augmentations, loss, etc. We will make good progress together!</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "920245",
      "postDate": "07/08/2020 13:04:46",
      "content": "<p>Hi All,</p>\n\n<p>I reproduced the highest score pytorch public efficientnet-b3 notebook on Keras. Augmentation includes the same horizontal and vertical flips. I tried label smoothing but made things worse. I also tried reduceOnPlanteau but didn't work either. Now I'm using custom stepLR and normal crossentropy. However, my single fold model score with even b5 and b6 are lower than the notebook. </p>\n\n<p>Did you guys encountered the same problem in Keras?</p>\n\n<p>Also, if you have a single fold small model score with Keras greater than the LB of current highest public notebook (0.922), please contact me to team up! I am using 5 TPUv3x8, meaning that I can do 5 experiments at the same time with giant models to get good results very very quickly. I need some good insights on hyperparameters, augmentations, loss, etc. We will make good progress together!</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi All,\n\nI reproduced the highest score pytorch public efficientnet-b3 notebook on Keras. Augmentation includes the same horizontal and vertical flips. I tried label smoothing but made things worse. I also tried reduceOnPlanteau but didn't work either. Now I'm using custom stepLR and normal crossentropy. However, my single fold model score with even b5 and b6 are lower than the notebook. \n\nDid you guys encountered the same problem in Keras?\n\nAlso, if you have a single fold small model score with Keras greater than the LB of current highest public notebook (0.922), please contact me to team up! I am using 5 TPUv3x8, meaning that I can do 5 experiments at the same time with giant models to get good results very very quickly. I need some good insights on hyperparameters, augmentations, loss, etc. We will make good progress together!\n\nThanks!",
      "votes": null
    },
    {
      "id": "921121",
      "postDate": "07/09/2020 04:57:32",
      "content": "<p>Hi,\ncan you tell us how many epochs did you train the model? I am curious to know how many epochs in Keras it takes to get to the baseline. </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\ncan you tell us how many epochs did you train the model? I am curious to know how many epochs in Keras it takes to get to the baseline. \n\nThanks",
      "votes": null
    },
    {
      "id": "926387",
      "postDate": "07/12/2020 17:17:27",
      "content": "<p><strong>ALOT</strong></p>",
      "rawMarkdown": "**ALOT**",
      "votes": null
    },
    {
      "id": "926390",
      "postDate": "07/12/2020 17:19:21",
      "content": "<p>I'll be more specific. If you look at the public GPU kernel, it trained more more than 48 hours. </p>\n\n<p>So if you use Keras and TF you should expect similar results if you follow similar training procedures. </p>\n\n<p>See <a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">https://www.kaggle.com/shonenkov/train-inference-gpu-baseline</a> </p>",
      "rawMarkdown": "I'll be more specific. If you look at the public GPU kernel, it trained more more than 48 hours. \n\nSo if you use Keras and TF you should expect similar results if you follow similar training procedures. \n\nSee https://www.kaggle.com/shonenkov/train-inference-gpu-baseline",
      "votes": null
    },
    {
      "id": "926856",
      "postDate": "07/13/2020 03:44:00",
      "content": "<p>Yeah, I followed the same kernel and did training for 34 epochs. But, could get better results. I need to do something in my training procedure I think. Thanks for the reply. </p>",
      "rawMarkdown": "Yeah, I followed the same kernel and did training for 34 epochs. But, could get better results. I need to do something in my training procedure I think. Thanks for the reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 921121,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "07/09/2020 04:57:32",
      "content": "<p>Hi,\ncan you tell us how many epochs did you train the model? I am curious to know how many epochs in Keras it takes to get to the baseline. </p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 926387,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/12/2020 17:17:27",
          "content": "<p><strong>ALOT</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926390,
          "author_name": "hooong",
          "author_url": "",
          "post_date": "07/12/2020 17:19:21",
          "content": "<p>I'll be more specific. If you look at the public GPU kernel, it trained more more than 48 hours. </p>\n\n<p>So if you use Keras and TF you should expect similar results if you follow similar training procedures. </p>\n\n<p>See <a href=\"https://www.kaggle.com/shonenkov/train-inference-gpu-baseline\">https://www.kaggle.com/shonenkov/train-inference-gpu-baseline</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926856,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "07/13/2020 03:44:00",
          "content": "<p>Yeah, I followed the same kernel and did training for 34 epochs. But, could get better results. I need to do something in my training procedure I think. Thanks for the reply. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920245": "Hi All,\n\nI reproduced the highest score pytorch public efficientnet-b3 notebook on Keras. Augmentation includes the same horizontal and vertical flips. I tried label smoothing but made things worse. I also tried reduceOnPlanteau but didn't work either. Now I'm using custom stepLR and normal crossentropy. However, my single fold model score with even b5 and b6 are lower than the notebook. \n\nDid you guys encountered the same problem in Keras?\n\nAlso, if you have a single fold small model score with Keras greater than the LB of current highest public notebook (0.922), please contact me to team up! I am using 5 TPUv3x8, meaning that I can do 5 experiments at the same time with giant models to get good results very very quickly. I need some good insights on hyperparameters, augmentations, loss, etc. We will make good progress together!\n\nThanks!",
    "921121": "Hi,\ncan you tell us how many epochs did you train the model? I am curious to know how many epochs in Keras it takes to get to the baseline. \n\nThanks",
    "926387": "**ALOT**",
    "926390": "I'll be more specific. If you look at the public GPU kernel, it trained more more than 48 hours. \n\nSo if you use Keras and TF you should expect similar results if you follow similar training procedures. \n\nSee https://www.kaggle.com/shonenkov/train-inference-gpu-baseline",
    "926856": "Yeah, I followed the same kernel and did training for 34 epochs. But, could get better results. I need to do something in my training procedure I think. Thanks for the reply."
  },
  "source": "meta"
}