{
  "id": 75803,
  "title": "Is there anything wrong with using Keras for this problem?",
  "url": "/competitions/humpback-whale-identification/discussion/75803",
  "author_name": "",
  "post_date": "2018-12-26T18:54:34.511602100Z",
  "votes": 1,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I'm starting to think I'm just just opening door with a brick wall before I even enter. Yes, that's what I feel now whenever I update my code and run it and check for any result. I have implemented a lot, directly implementing what worked for others in this problem and getting nowhere. Please do note that the same thing worked wonders for someone implementing it in fastai.</p>\n\n<p>Here is a link to my unfortunate kernel -- </p>\n\n<p><a href=\"https://www.kaggle.com/whatvermawhat/resnet50-128x128\">https://www.kaggle.com/whatvermawhat/resnet50-128x128</a></p>\n\n<p>I'm a beginner, this is my first competition on kaggle and I have nowhere to ask else. In a short summary here's what I'm doing so that you don't have open and read every line -- </p>\n\n<ol>\n<li><p>Importing images, resizing them to 128x128 and making a numpy array out of it, using 0.1% for validation (removing new_whale class).</p></li>\n<li><p>Using ResNet50 with last 6 Conv layers trainable. (the first layer, is preprocessing layer so that after augmentation, they get preprocessed according to resnet50 preprocessing provided by keras)</p></li>\n<li><p>Using data augmentation with your usual shifting, rotation and horizontal flips.</p></li>\n<li><p>Training with learning rate of 1e-3, Adam optimizer, </p></li>\n<li><p>Using a threshold of 0.38 to insert new_whale class. </p></li>\n</ol>\n\n<p>With this configuration, I get a score 13.8 which is half of 27.6 which is the score you obtain if you put new_whale class as predicted class every time, so that implies that there's no improvement over a dumb model where you just predict the new_whale class every time. Just that instead of predicting as the top class, I predict it as the second top class and that's why my prediction is half of the dumb model predictions. </p>\n\n<p>So now, my one and only question is....Is there something I'm doing that you think I shouldn't be or is there something I'm not doing, and you feel that obviously I be doing?</p>",
  "messages": [
    {
      "id": "445576",
      "postDate": "12/26/2018 18:54:34",
      "content": "<p>I'm starting to think I'm just just opening door with a brick wall before I even enter. Yes, that's what I feel now whenever I update my code and run it and check for any result. I have implemented a lot, directly implementing what worked for others in this problem and getting nowhere. Please do note that the same thing worked wonders for someone implementing it in fastai.</p>\n\n<p>Here is a link to my unfortunate kernel -- </p>\n\n<p><a href=\"https://www.kaggle.com/whatvermawhat/resnet50-128x128\">https://www.kaggle.com/whatvermawhat/resnet50-128x128</a></p>\n\n<p>I'm a beginner, this is my first competition on kaggle and I have nowhere to ask else. In a short summary here's what I'm doing so that you don't have open and read every line -- </p>\n\n<ol>\n<li><p>Importing images, resizing them to 128x128 and making a numpy array out of it, using 0.1% for validation (removing new_whale class).</p></li>\n<li><p>Using ResNet50 with last 6 Conv layers trainable. (the first layer, is preprocessing layer so that after augmentation, they get preprocessed according to resnet50 preprocessing provided by keras)</p></li>\n<li><p>Using data augmentation with your usual shifting, rotation and horizontal flips.</p></li>\n<li><p>Training with learning rate of 1e-3, Adam optimizer, </p></li>\n<li><p>Using a threshold of 0.38 to insert new_whale class. </p></li>\n</ol>\n\n<p>With this configuration, I get a score 13.8 which is half of 27.6 which is the score you obtain if you put new_whale class as predicted class every time, so that implies that there's no improvement over a dumb model where you just predict the new_whale class every time. Just that instead of predicting as the top class, I predict it as the second top class and that's why my prediction is half of the dumb model predictions. </p>\n\n<p>So now, my one and only question is....Is there something I'm doing that you think I shouldn't be or is there something I'm not doing, and you feel that obviously I be doing?</p>",
      "rawMarkdown": "I'm starting to think I'm just just opening door with a brick wall before I even enter. Yes, that's what I feel now whenever I update my code and run it and check for any result. I have implemented a lot, directly implementing what worked for others in this problem and getting nowhere. Please do note that the same thing worked wonders for someone implementing it in fastai.\n\nHere is a link to my unfortunate kernel -- \n\nhttps://www.kaggle.com/whatvermawhat/resnet50-128x128\n\nI'm a beginner, this is my first competition on kaggle and I have nowhere to ask else. In a short summary here's what I'm doing so that you don't have open and read every line -- \n\n1. Importing images, resizing them to 128x128 and making a numpy array out of it, using 0.1% for validation (removing new_whale class).\n\n2. Using ResNet50 with last 6 Conv layers trainable. (the first layer, is preprocessing layer so that after augmentation, they get preprocessed according to resnet50 preprocessing provided by keras)\n\n3. Using data augmentation with your usual shifting, rotation and horizontal flips.\n\n4. Training with learning rate of 1e-3, Adam optimizer, \n\n5. Using a threshold of 0.38 to insert new_whale class. \n\nWith this configuration, I get a score 13.8 which is half of 27.6 which is the score you obtain if you put new_whale class as predicted class every time, so that implies that there's no improvement over a dumb model where you just predict the new_whale class every time. Just that instead of predicting as the top class, I predict it as the second top class and that's why my prediction is half of the dumb model predictions. \n\nSo now, my one and only question is....Is there something I'm doing that you think I shouldn't be or is there something I'm not doing, and you feel that obviously I be doing?",
      "votes": null
    },
    {
      "id": "445602",
      "postDate": "12/26/2018 19:43:22",
      "content": "<p>It's better to understand your model rather than to apply anecdotal strategies that you don't understand. I would start with a smaller model and dataset with no augmentation or regularization. Take two whales of different classes and get your smaller model to overfit on them. Then start adding more data/classes and growing your model from there. Add in augmentation and regularization when you start overfitting to much larger data.</p>",
      "rawMarkdown": "It's better to understand your model rather than to apply anecdotal strategies that you don't understand. I would start with a smaller model and dataset with no augmentation or regularization. Take two whales of different classes and get your smaller model to overfit on them. Then start adding more data/classes and growing your model from there. Add in augmentation and regularization when you start overfitting to much larger data.",
      "votes": null
    },
    {
      "id": "445655",
      "postDate": "12/26/2018 22:05:39",
      "content": "<p>I came here just to ask this very question. I'm having some troubles as well in order to get Keras to work properly building similar models to the public kernels implemented in pytorch+fastai. I've already begun studying fastai to implement my solutions using this framework. It seems from my experiences that the simplest models in fastai are already better than more elaborated ones in Keras. </p>",
      "rawMarkdown": "I came here just to ask this very question. I'm having some troubles as well in order to get Keras to work properly building similar models to the public kernels implemented in pytorch+fastai. I've already begun studying fastai to implement my solutions using this framework. It seems from my experiences that the simplest models in fastai are already better than more elaborated ones in Keras.",
      "votes": null
    },
    {
      "id": "445902",
      "postDate": "12/27/2018 07:29:26",
      "content": "<p>Hello Simeon,</p>\n\n<p>First of all, thank you for your response!\nI never said I didn't understand what others implemented, and yes, I've implemented stuff from the ground up. From 5 conv layers with some dropout, then implementing it without new_whale class, then using a pretrained model, first training only the dense classifier and then trying out fine tuning, with and without new_whale class. Then using Data Augmentation, on 100x100, 128x128, 224x224 images and I'm still yet to see any score above 30%. I'm still experiencing horrible overfitting.</p>\n\n<p>My problem was why even the simple models (Let's face it, fine tuning a pretrained model is very simple in keras) doesn't work for me while the exact same implementation worked for someone who implemented it in fastai?</p>",
      "rawMarkdown": "Hello Simeon,\n\nFirst of all, thank you for your response!\nI never said I didn't understand what others implemented, and yes, I've implemented stuff from the ground up. From 5 conv layers with some dropout, then implementing it without new_whale class, then using a pretrained model, first training only the dense classifier and then trying out fine tuning, with and without new_whale class. Then using Data Augmentation, on 100x100, 128x128, 224x224 images and I'm still yet to see any score above 30%. I'm still experiencing horrible overfitting.\n\nMy problem was why even the simple models (Let's face it, fine tuning a pretrained model is very simple in keras) doesn't work for me while the exact same implementation worked for someone who implemented it in fastai?",
      "votes": null
    },
    {
      "id": "445905",
      "postDate": "12/27/2018 07:34:20",
      "content": "<p>Exactly! and that fact alone (simple models of fastai &gt; keras) is quite frustrating because I spent a good amount of time experimenting on Keras. The only thing that's now left to experiment is using an image size greater than 224x224 and aside from that I don't think there's difference between mine and the baseline fastai resnet50 model which got 62% LB</p>",
      "rawMarkdown": "Exactly! and that fact alone (simple models of fastai &gt; keras) is quite frustrating because I spent a good amount of time experimenting on Keras. The only thing that's now left to experiment is using an image size greater than 224x224 and aside from that I don't think there's difference between mine and the baseline fastai resnet50 model which got 62% LB",
      "votes": null
    },
    {
      "id": "445906",
      "postDate": "12/27/2018 07:36:27",
      "content": "<p>Also, may I ask where are you learning fastai from? </p>",
      "rawMarkdown": "Also, may I ask where are you learning fastai from?",
      "votes": null
    },
    {
      "id": "446055",
      "postDate": "12/27/2018 12:27:07",
      "content": "<p>Actually, I'm just reading other kernels written in fast.ai, since the deep learning principles are the same and doesn't change between frameworks and we just need to learn another set of \"grammar\". Whenever I'm in doubt, I go at <a href=\"https://course.fast.ai/lessons/lesson2.html\">fastai site</a> (or just google the specific issue). </p>",
      "rawMarkdown": "Actually, I'm just reading other kernels written in fast.ai, since the deep learning principles are the same and doesn't change between frameworks and we just need to learn another set of \"grammar\". Whenever I'm in doubt, I go at [fastai site](https://course.fast.ai/lessons/lesson2.html) (or just google the specific issue).",
      "votes": null
    },
    {
      "id": "446062",
      "postDate": "12/27/2018 12:43:37",
      "content": "<p>\"while the exact same implementation worked for someone who implemented it in fastai?'\n0)Very important: \nrotation_range=60, from your kernel is huuuuge. Try 15 or 20 degrees. Or no rotation at all.\n1)Most important, fast.ai uses carefully tuned fancy one_cycle policy:\n<a href=\"https://docs.fast.ai/basic_train.html#fit_one_cycle\">https://docs.fast.ai/basic_train.html#fit_one_cycle</a>\nIt is way better, than plain Adam. </p>",
      "rawMarkdown": "\"while the exact same implementation worked for someone who implemented it in fastai?'\n0)Very important: \nrotation_range=60, from your kernel is huuuuge. Try 15 or 20 degrees. Or no rotation at all.\n1)Most important, fast.ai uses carefully tuned fancy one_cycle policy:\nhttps://docs.fast.ai/basic_train.html#fit_one_cycle\nIt is way better, than plain Adam.",
      "votes": null
    },
    {
      "id": "446066",
      "postDate": "12/27/2018 12:56:43",
      "content": "<p>Yea I think that's the way to go about. I started with some documentations but then later realised Kaggle kernels don't support fastai v1 and now started with the videos. Reading a baseline fastai model now, let's get started with fastai now haha. Thanks for the input!</p>",
      "rawMarkdown": "Yea I think that's the way to go about. I started with some documentations but then later realised Kaggle kernels don't support fastai v1 and now started with the videos. Reading a baseline fastai model now, let's get started with fastai now haha. Thanks for the input!",
      "votes": null
    },
    {
      "id": "446068",
      "postDate": "12/27/2018 13:00:12",
      "content": "<p>I see, I didn't know about the onecycle policy, seems interesting. About the rotation range, I did start from 20 and went upto 60 in an attempt to cure the overfitting. I'm now starting with fastai, while trying out some minor attempts at the keras implementation. Thank you for your input :)</p>",
      "rawMarkdown": "I see, I didn't know about the onecycle policy, seems interesting. About the rotation range, I did start from 20 and went upto 60 in an attempt to cure the overfitting. I'm now starting with fastai, while trying out some minor attempts at the keras implementation. Thank you for your input :)",
      "votes": null
    },
    {
      "id": "446688",
      "postDate": "12/28/2018 14:11:32",
      "content": "<p>I already checked 256*256*1 ,, I still not able to achieve above 0.25. I tried everything, I thought there is something wrong with batching, then I converted data into TFRecord. and tried to use Dataset API with Keras. I even written Resnet50 model in tensorflow, and tried to train this model on 2 1080Ti, took me lot of time to optimize it.  And still not able to go above  .27. I was about to give up today. Then I thought to check out in discussions. Now I starting to think to shift to fast.ai. Never tried it before, but if this is working then why not.</p>",
      "rawMarkdown": "I already checked 256*256*1 ,, I still not able to achieve above 0.25. I tried everything, I thought there is something wrong with batching, then I converted data into TFRecord. and tried to use Dataset API with Keras. I even written Resnet50 model in tensorflow, and tried to train this model on 2 1080Ti, took me lot of time to optimize it.  And still not able to go above  .27. I was about to give up today. Then I thought to check out in discussions. Now I starting to think to shift to fast.ai. Never tried it before, but if this is working then why not.",
      "votes": null
    },
    {
      "id": "446699",
      "postDate": "12/28/2018 14:30:22",
      "content": "<p>I guess the faster we shift, the better. Although fastai is not quite documented properly I'd say, but if that's working, It'll do for me.</p>",
      "rawMarkdown": "I guess the faster we shift, the better. Although fastai is not quite documented properly I'd say, but if that's working, It'll do for me.",
      "votes": null
    },
    {
      "id": "447369",
      "postDate": "12/29/2018 18:10:01",
      "content": "<p>You can program your own learning rate scheduler, such as one cycle, quite easily in Keras. I don't really see this as a compelling reason to switch to fastai, other than the fact that it holds your hand more than Keras? You lose the learning gained from understanding why the fastai training is more effective than the vanilla Keras implementation by just switching to \"what works\". That's fine if you just want to power past it, but you won't understand what was actually done behind the scenes. This is exactly what I was talking about when I said, \"It's better to understand your model rather than to apply anecdotal strategies that you don't understand.\" Best of luck.</p>",
      "rawMarkdown": "You can program your own learning rate scheduler, such as one cycle, quite easily in Keras. I don't really see this as a compelling reason to switch to fastai, other than the fact that it holds your hand more than Keras? You lose the learning gained from understanding why the fastai training is more effective than the vanilla Keras implementation by just switching to \"what works\". That's fine if you just want to power past it, but you won't understand what was actually done behind the scenes. This is exactly what I was talking about when I said, \"It's better to understand your model rather than to apply anecdotal strategies that you don't understand.\" Best of luck.",
      "votes": null
    },
    {
      "id": "447461",
      "postDate": "12/29/2018 21:33:16",
      "content": "<p>Hello Abhinav Verma,\nI think its always good to learn something new like fastai, but I wouldn't give up with your current code. I think there should be something wrong in your implementation. So did you implement a kernel in keras that is written with fastai? Thats a great think to do.\nIs the loss right? Is the network really the same? Did you give the right labels to the network? Did you shuffle the features but not the labels? Maybe you can post both codes, so we can compare.</p>\n\n<p>Maybe, you could try to train cifar10 or mnist with your script. And then you can find the solution step by step</p>",
      "rawMarkdown": "Hello Abhinav Verma,\nI think its always good to learn something new like fastai, but I wouldn't give up with your current code. I think there should be something wrong in your implementation. So did you implement a kernel in keras that is written with fastai? Thats a great think to do.\nIs the loss right? Is the network really the same? Did you give the right labels to the network? Did you shuffle the features but not the labels? Maybe you can post both codes, so we can compare.\n\nMaybe, you could try to train cifar10 or mnist with your script. And then you can find the solution step by step",
      "votes": null
    },
    {
      "id": "447876",
      "postDate": "12/30/2018 18:40:32",
      "content": "<p>There is nothing wrong with Keras.</p>\n\n<p>The \"problem\" is that fast.ai does a lot of magic under the hood, which is cool if you either are already familiar with the magic or you don't really care about understanding about what is going on. If you implement the magic in Keras (not really hard) you should be able to obtain the same results.</p>\n\n<p>My current best submission is based on an internal (from work) framework build with TensorFlow+Keras and is on pair with the highest published (as far as I know) fast.ai (1.0) solution: <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/74647\">https://www.kaggle.com/c/humpback-whale-identification/discussion/74647</a>.\nThat fast.ai solution is, in comparison, a far more complex pipeline (including an oversampled dataset which probably has a relevant positive impact in the score).</p>\n\n<p>I can not share the code or use a Kernel but I would be happy to answer questions about the Keras code needed to replicate the result.</p>\n\n<p>My settings are:</p>\n\n<ul>\n<li>Input data:</li>\n</ul>\n\n<p>The same bounding box dataset that everybody is using in the public kernels. new_whale class is ignored. </p>\n\n<ul>\n<li>Model: </li>\n</ul>\n\n<p>Plain ResNet50 with \"big head\" using weights converted from keras-applications. The big head comes from replacing the last fully connected of the original model with:</p>\n\n<p>(starting from Global Average Pooling) -&gt; BatchNorm -&gt; Dropout(0.25) -&gt; Dense (2048) -&gt; BatchNorm -&gt; Dropout(0.5) -&gt; Dense(5004)</p>\n\n<ul>\n<li><p>Batch size: 20</p></li>\n<li><p>Input size: 200x600</p></li>\n<li><p>Preprocesing:</p></li>\n</ul>\n\n<p>Rescale to range(-1, 1) -&gt; random rotation (20 degree) -&gt; random brightness (0.2 delta) -&gt; random displace (5 pixels)</p>\n\n<ul>\n<li><p>Epochs: 24</p></li>\n<li><p>Optimizer: </p></li>\n</ul>\n\n<p>A modified version of Keras Adam adding support for \"True weight decay\" and \"Learning rate multipliers\" just like fast.ai. Settings:</p>\n\n<p>Clip grad norm to 0.1</p>\n\n<p>\"True\" Weight decay: 0.0001</p>\n\n<p>Multipliers (the learning rate gets multiplied by this value for certain layers):</p>\n\n<p>0.01 for fist blocks.\n0.1 for middle blocks</p>\n\n<ul>\n<li>Scheduler:</li>\n</ul>\n\n<p>One cycle just like fast.ai (1.0) with defaults params. Note that the fast.ai one cycle has some relevant details that some public Keras implementations don't.</p>\n\n<p>Maximum learning rate of 0.001. </p>",
      "rawMarkdown": "There is nothing wrong with Keras.\n\nThe \"problem\" is that fast.ai does a lot of magic under the hood, which is cool if you either are already familiar with the magic or you don't really care about understanding about what is going on. If you implement the magic in Keras (not really hard) you should be able to obtain the same results.\n\nMy current best submission is based on an internal (from work) framework build with TensorFlow+Keras and is on pair with the highest published (as far as I know) fast.ai (1.0) solution: https://www.kaggle.com/c/humpback-whale-identification/discussion/74647.\nThat fast.ai solution is, in comparison, a far more complex pipeline (including an oversampled dataset which probably has a relevant positive impact in the score).\n\nI can not share the code or use a Kernel but I would be happy to answer questions about the Keras code needed to replicate the result.\n\nMy settings are:\n\n- Input data:\n\nThe same bounding box dataset that everybody is using in the public kernels. new_whale class is ignored. \n\n- Model: \n\nPlain ResNet50 with \"big head\" using weights converted from keras-applications. The big head comes from replacing the last fully connected of the original model with:\n\n(starting from Global Average Pooling) -&gt; BatchNorm -&gt; Dropout(0.25) -&gt; Dense (2048) -&gt; BatchNorm -&gt; Dropout(0.5) -&gt; Dense(5004)\n\n- Batch size: 20\n\n- Input size: 200x600\n\n- Preprocesing:\n\nRescale to range(-1, 1) -&gt; random rotation (20 degree) -&gt; random brightness (0.2 delta) -&gt; random displace (5 pixels)\n\n- Epochs: 24\n\n- Optimizer: \n\nA modified version of Keras Adam adding support for \"True weight decay\" and \"Learning rate multipliers\" just like fast.ai. Settings:\n\nClip grad norm to 0.1\n\n\"True\" Weight decay: 0.0001\n\nMultipliers (the learning rate gets multiplied by this value for certain layers):\n\n0.01 for fist blocks.\n0.1 for middle blocks\n\n- Scheduler:\n\nOne cycle just like fast.ai (1.0) with defaults params. Note that the fast.ai one cycle has some relevant details that some public Keras implementations don't.\n\nMaximum learning rate of 0.001.",
      "votes": null
    },
    {
      "id": "447931",
      "postDate": "12/30/2018 21:33:27",
      "content": "<p>What is the reasoning for not using bias in the Dense layer? </p>",
      "rawMarkdown": "What is the reasoning for not using bias in the Dense layer?",
      "votes": null
    },
    {
      "id": "447942",
      "postDate": "12/30/2018 21:55:57",
      "content": "<p>Hi David,</p>\n\n<p>Thank you for such an informative comment! \nAfter getting a little frustrated with Keras, I decided to try out fast.ai by trying out their MOOC. After learning about what fastai does under the hood, I'm now beginning to understand why it's performing better and how what I implemented thinking that it's as same as the fastai solution, is really not even close. Personally, I don't prefer the fast.ai approach (Not really understanding what's going on under the hood) and now trying out the different techniques that fast.ai implemented, on Keras. Let's hope for the best :) </p>",
      "rawMarkdown": "Hi David,\n\nThank you for such an informative comment! \nAfter getting a little frustrated with Keras, I decided to try out fast.ai by trying out their MOOC. After learning about what fastai does under the hood, I'm now beginning to understand why it's performing better and how what I implemented thinking that it's as same as the fastai solution, is really not even close. Personally, I don't prefer the fast.ai approach (Not really understanding what's going on under the hood) and now trying out the different techniques that fast.ai implemented, on Keras. Let's hope for the best :)",
      "votes": null
    },
    {
      "id": "447943",
      "postDate": "12/30/2018 21:57:58",
      "content": "<p>Hi corner200, </p>\n\n<p>After trying out a little of fastai, I'm back to Keras and now trying out different things (Like SGD with restarts). Let's see if things work out for the best :)</p>",
      "rawMarkdown": "Hi corner200, \n\nAfter trying out a little of fastai, I'm back to Keras and now trying out different things (Like SGD with restarts). Let's see if things work out for the best :)",
      "votes": null
    },
    {
      "id": "447953",
      "postDate": "12/30/2018 22:27:55",
      "content": "<p>since I am not active in this competition I just glanced over your code and couldn't find the mistake. but I am sure it should work with keras. If I can help, let me know. Try to start from something very simple, step by step you will find the solution!</p>",
      "rawMarkdown": "since I am not active in this competition I just glanced over your code and couldn't find the mistake. but I am sure it should work with keras. If I can help, let me know. Try to start from something very simple, step by step you will find the solution!",
      "votes": null
    },
    {
      "id": "447955",
      "postDate": "12/30/2018 22:30:55",
      "content": "<p>My mistake. use_bias was set to True which is the default</p>",
      "rawMarkdown": "My mistake. use_bias was set to True which is the default",
      "votes": null
    },
    {
      "id": "448154",
      "postDate": "12/31/2018 11:06:42",
      "content": "<p>Just changing:</p>\n\n<ul>\n<li>Batch size: 10</li>\n<li>Input size: 300x900</li>\n</ul>\n\n<p>Raises the score to 0.788</p>\n\n<p>Inserting new_whale at threshold 0.8. \nThis threshold is quite relevant and properly tuning it would boost the score. So far I didn't find a way to tune this parameter locally without using the public scores. I believe that using the leaderboard would result in overfit to the 20% of the test set.</p>",
      "rawMarkdown": "Just changing:\n\n- Batch size: 10\n- Input size: 300x900\n\nRaises the score to 0.788\n\nInserting new_whale at threshold 0.8. \nThis threshold is quite relevant and properly tuning it would boost the score. So far I didn't find a way to tune this parameter locally without using the public scores. I believe that using the leaderboard would result in overfit to the 20% of the test set.",
      "votes": null
    },
    {
      "id": "458910",
      "postDate": "01/20/2019 19:43:18",
      "content": "<p>Hey, thanks for the advice. I customised Keras Adam optimiser, added weight decays and multipliers, and achieved .66 LB with 300 300 1 and 0.4 as threshold. </p>\n\n<p>Obviously I think  with higher h and w of images and optimizing the threshold, it able to achieve much higher LB.  </p>\n\n<p>But yeah optimizing threshold by LB definitely result in overfit. </p>",
      "rawMarkdown": "Hey, thanks for the advice. I customised Keras Adam optimiser, added weight decays and multipliers, and achieved .66 LB with 300 300 1 and 0.4 as threshold. \n\nObviously I think  with higher h and w of images and optimizing the threshold, it able to achieve much higher LB.  \n\nBut yeah optimizing threshold by LB definitely result in overfit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 445602,
      "author_name": "badtyprr",
      "author_url": "",
      "post_date": "12/26/2018 19:43:22",
      "content": "<p>It's better to understand your model rather than to apply anecdotal strategies that you don't understand. I would start with a smaller model and dataset with no augmentation or regularization. Take two whales of different classes and get your smaller model to overfit on them. Then start adding more data/classes and growing your model from there. Add in augmentation and regularization when you start overfitting to much larger data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 445902,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/27/2018 07:29:26",
          "content": "<p>Hello Simeon,</p>\n\n<p>First of all, thank you for your response!\nI never said I didn't understand what others implemented, and yes, I've implemented stuff from the ground up. From 5 conv layers with some dropout, then implementing it without new_whale class, then using a pretrained model, first training only the dense classifier and then trying out fine tuning, with and without new_whale class. Then using Data Augmentation, on 100x100, 128x128, 224x224 images and I'm still yet to see any score above 30%. I'm still experiencing horrible overfitting.</p>\n\n<p>My problem was why even the simple models (Let's face it, fine tuning a pretrained model is very simple in keras) doesn't work for me while the exact same implementation worked for someone who implemented it in fastai?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446062,
          "author_name": "oldufo",
          "author_url": "",
          "post_date": "12/27/2018 12:43:37",
          "content": "<p>\"while the exact same implementation worked for someone who implemented it in fastai?'\n0)Very important: \nrotation_range=60, from your kernel is huuuuge. Try 15 or 20 degrees. Or no rotation at all.\n1)Most important, fast.ai uses carefully tuned fancy one_cycle policy:\n<a href=\"https://docs.fast.ai/basic_train.html#fit_one_cycle\">https://docs.fast.ai/basic_train.html#fit_one_cycle</a>\nIt is way better, than plain Adam. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446068,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/27/2018 13:00:12",
          "content": "<p>I see, I didn't know about the onecycle policy, seems interesting. About the rotation range, I did start from 20 and went upto 60 in an attempt to cure the overfitting. I'm now starting with fastai, while trying out some minor attempts at the keras implementation. Thank you for your input :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447369,
          "author_name": "badtyprr",
          "author_url": "",
          "post_date": "12/29/2018 18:10:01",
          "content": "<p>You can program your own learning rate scheduler, such as one cycle, quite easily in Keras. I don't really see this as a compelling reason to switch to fastai, other than the fact that it holds your hand more than Keras? You lose the learning gained from understanding why the fastai training is more effective than the vanilla Keras implementation by just switching to \"what works\". That's fine if you just want to power past it, but you won't understand what was actually done behind the scenes. This is exactly what I was talking about when I said, \"It's better to understand your model rather than to apply anecdotal strategies that you don't understand.\" Best of luck.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 445655,
      "author_name": "hrmello",
      "author_url": "",
      "post_date": "12/26/2018 22:05:39",
      "content": "<p>I came here just to ask this very question. I'm having some troubles as well in order to get Keras to work properly building similar models to the public kernels implemented in pytorch+fastai. I've already begun studying fastai to implement my solutions using this framework. It seems from my experiences that the simplest models in fastai are already better than more elaborated ones in Keras. </p>",
      "votes": null,
      "replies": [
        {
          "id": 445905,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/27/2018 07:34:20",
          "content": "<p>Exactly! and that fact alone (simple models of fastai &gt; keras) is quite frustrating because I spent a good amount of time experimenting on Keras. The only thing that's now left to experiment is using an image size greater than 224x224 and aside from that I don't think there's difference between mine and the baseline fastai resnet50 model which got 62% LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445906,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/27/2018 07:36:27",
          "content": "<p>Also, may I ask where are you learning fastai from? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446055,
          "author_name": "hrmello",
          "author_url": "",
          "post_date": "12/27/2018 12:27:07",
          "content": "<p>Actually, I'm just reading other kernels written in fast.ai, since the deep learning principles are the same and doesn't change between frameworks and we just need to learn another set of \"grammar\". Whenever I'm in doubt, I go at <a href=\"https://course.fast.ai/lessons/lesson2.html\">fastai site</a> (or just google the specific issue). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446066,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/27/2018 12:56:43",
          "content": "<p>Yea I think that's the way to go about. I started with some documentations but then later realised Kaggle kernels don't support fastai v1 and now started with the videos. Reading a baseline fastai model now, let's get started with fastai now haha. Thanks for the input!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446688,
          "author_name": "kakush30101991",
          "author_url": "",
          "post_date": "12/28/2018 14:11:32",
          "content": "<p>I already checked 256*256*1 ,, I still not able to achieve above 0.25. I tried everything, I thought there is something wrong with batching, then I converted data into TFRecord. and tried to use Dataset API with Keras. I even written Resnet50 model in tensorflow, and tried to train this model on 2 1080Ti, took me lot of time to optimize it.  And still not able to go above  .27. I was about to give up today. Then I thought to check out in discussions. Now I starting to think to shift to fast.ai. Never tried it before, but if this is working then why not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446699,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/28/2018 14:30:22",
          "content": "<p>I guess the faster we shift, the better. Although fastai is not quite documented properly I'd say, but if that's working, It'll do for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447461,
      "author_name": "corner200",
      "author_url": "",
      "post_date": "12/29/2018 21:33:16",
      "content": "<p>Hello Abhinav Verma,\nI think its always good to learn something new like fastai, but I wouldn't give up with your current code. I think there should be something wrong in your implementation. So did you implement a kernel in keras that is written with fastai? Thats a great think to do.\nIs the loss right? Is the network really the same? Did you give the right labels to the network? Did you shuffle the features but not the labels? Maybe you can post both codes, so we can compare.</p>\n\n<p>Maybe, you could try to train cifar10 or mnist with your script. And then you can find the solution step by step</p>",
      "votes": null,
      "replies": [
        {
          "id": 447943,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/30/2018 21:57:58",
          "content": "<p>Hi corner200, </p>\n\n<p>After trying out a little of fastai, I'm back to Keras and now trying out different things (Like SGD with restarts). Let's see if things work out for the best :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447953,
          "author_name": "corner200",
          "author_url": "",
          "post_date": "12/30/2018 22:27:55",
          "content": "<p>since I am not active in this competition I just glanced over your code and couldn't find the mistake. but I am sure it should work with keras. If I can help, let me know. Try to start from something very simple, step by step you will find the solution!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447876,
      "author_name": "daavoo",
      "author_url": "",
      "post_date": "12/30/2018 18:40:32",
      "content": "<p>There is nothing wrong with Keras.</p>\n\n<p>The \"problem\" is that fast.ai does a lot of magic under the hood, which is cool if you either are already familiar with the magic or you don't really care about understanding about what is going on. If you implement the magic in Keras (not really hard) you should be able to obtain the same results.</p>\n\n<p>My current best submission is based on an internal (from work) framework build with TensorFlow+Keras and is on pair with the highest published (as far as I know) fast.ai (1.0) solution: <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/74647\">https://www.kaggle.com/c/humpback-whale-identification/discussion/74647</a>.\nThat fast.ai solution is, in comparison, a far more complex pipeline (including an oversampled dataset which probably has a relevant positive impact in the score).</p>\n\n<p>I can not share the code or use a Kernel but I would be happy to answer questions about the Keras code needed to replicate the result.</p>\n\n<p>My settings are:</p>\n\n<ul>\n<li>Input data:</li>\n</ul>\n\n<p>The same bounding box dataset that everybody is using in the public kernels. new_whale class is ignored. </p>\n\n<ul>\n<li>Model: </li>\n</ul>\n\n<p>Plain ResNet50 with \"big head\" using weights converted from keras-applications. The big head comes from replacing the last fully connected of the original model with:</p>\n\n<p>(starting from Global Average Pooling) -&gt; BatchNorm -&gt; Dropout(0.25) -&gt; Dense (2048) -&gt; BatchNorm -&gt; Dropout(0.5) -&gt; Dense(5004)</p>\n\n<ul>\n<li><p>Batch size: 20</p></li>\n<li><p>Input size: 200x600</p></li>\n<li><p>Preprocesing:</p></li>\n</ul>\n\n<p>Rescale to range(-1, 1) -&gt; random rotation (20 degree) -&gt; random brightness (0.2 delta) -&gt; random displace (5 pixels)</p>\n\n<ul>\n<li><p>Epochs: 24</p></li>\n<li><p>Optimizer: </p></li>\n</ul>\n\n<p>A modified version of Keras Adam adding support for \"True weight decay\" and \"Learning rate multipliers\" just like fast.ai. Settings:</p>\n\n<p>Clip grad norm to 0.1</p>\n\n<p>\"True\" Weight decay: 0.0001</p>\n\n<p>Multipliers (the learning rate gets multiplied by this value for certain layers):</p>\n\n<p>0.01 for fist blocks.\n0.1 for middle blocks</p>\n\n<ul>\n<li>Scheduler:</li>\n</ul>\n\n<p>One cycle just like fast.ai (1.0) with defaults params. Note that the fast.ai one cycle has some relevant details that some public Keras implementations don't.</p>\n\n<p>Maximum learning rate of 0.001. </p>",
      "votes": null,
      "replies": [
        {
          "id": 447931,
          "author_name": "badtyprr",
          "author_url": "",
          "post_date": "12/30/2018 21:33:27",
          "content": "<p>What is the reasoning for not using bias in the Dense layer? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447942,
          "author_name": "whatvermawhat",
          "author_url": "",
          "post_date": "12/30/2018 21:55:57",
          "content": "<p>Hi David,</p>\n\n<p>Thank you for such an informative comment! \nAfter getting a little frustrated with Keras, I decided to try out fast.ai by trying out their MOOC. After learning about what fastai does under the hood, I'm now beginning to understand why it's performing better and how what I implemented thinking that it's as same as the fastai solution, is really not even close. Personally, I don't prefer the fast.ai approach (Not really understanding what's going on under the hood) and now trying out the different techniques that fast.ai implemented, on Keras. Let's hope for the best :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447955,
          "author_name": "daavoo",
          "author_url": "",
          "post_date": "12/30/2018 22:30:55",
          "content": "<p>My mistake. use_bias was set to True which is the default</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448154,
          "author_name": "daavoo",
          "author_url": "",
          "post_date": "12/31/2018 11:06:42",
          "content": "<p>Just changing:</p>\n\n<ul>\n<li>Batch size: 10</li>\n<li>Input size: 300x900</li>\n</ul>\n\n<p>Raises the score to 0.788</p>\n\n<p>Inserting new_whale at threshold 0.8. \nThis threshold is quite relevant and properly tuning it would boost the score. So far I didn't find a way to tune this parameter locally without using the public scores. I believe that using the leaderboard would result in overfit to the 20% of the test set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 458910,
          "author_name": "kakush30101991",
          "author_url": "",
          "post_date": "01/20/2019 19:43:18",
          "content": "<p>Hey, thanks for the advice. I customised Keras Adam optimiser, added weight decays and multipliers, and achieved .66 LB with 300 300 1 and 0.4 as threshold. </p>\n\n<p>Obviously I think  with higher h and w of images and optimizing the threshold, it able to achieve much higher LB.  </p>\n\n<p>But yeah optimizing threshold by LB definitely result in overfit. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "445576": "I'm starting to think I'm just just opening door with a brick wall before I even enter. Yes, that's what I feel now whenever I update my code and run it and check for any result. I have implemented a lot, directly implementing what worked for others in this problem and getting nowhere. Please do note that the same thing worked wonders for someone implementing it in fastai.\n\nHere is a link to my unfortunate kernel -- \n\nhttps://www.kaggle.com/whatvermawhat/resnet50-128x128\n\nI'm a beginner, this is my first competition on kaggle and I have nowhere to ask else. In a short summary here's what I'm doing so that you don't have open and read every line -- \n\n1. Importing images, resizing them to 128x128 and making a numpy array out of it, using 0.1% for validation (removing new_whale class).\n\n2. Using ResNet50 with last 6 Conv layers trainable. (the first layer, is preprocessing layer so that after augmentation, they get preprocessed according to resnet50 preprocessing provided by keras)\n\n3. Using data augmentation with your usual shifting, rotation and horizontal flips.\n\n4. Training with learning rate of 1e-3, Adam optimizer, \n\n5. Using a threshold of 0.38 to insert new_whale class. \n\nWith this configuration, I get a score 13.8 which is half of 27.6 which is the score you obtain if you put new_whale class as predicted class every time, so that implies that there's no improvement over a dumb model where you just predict the new_whale class every time. Just that instead of predicting as the top class, I predict it as the second top class and that's why my prediction is half of the dumb model predictions. \n\nSo now, my one and only question is....Is there something I'm doing that you think I shouldn't be or is there something I'm not doing, and you feel that obviously I be doing?",
    "445602": "It's better to understand your model rather than to apply anecdotal strategies that you don't understand. I would start with a smaller model and dataset with no augmentation or regularization. Take two whales of different classes and get your smaller model to overfit on them. Then start adding more data/classes and growing your model from there. Add in augmentation and regularization when you start overfitting to much larger data.",
    "445655": "I came here just to ask this very question. I'm having some troubles as well in order to get Keras to work properly building similar models to the public kernels implemented in pytorch+fastai. I've already begun studying fastai to implement my solutions using this framework. It seems from my experiences that the simplest models in fastai are already better than more elaborated ones in Keras.",
    "445902": "Hello Simeon,\n\nFirst of all, thank you for your response!\nI never said I didn't understand what others implemented, and yes, I've implemented stuff from the ground up. From 5 conv layers with some dropout, then implementing it without new_whale class, then using a pretrained model, first training only the dense classifier and then trying out fine tuning, with and without new_whale class. Then using Data Augmentation, on 100x100, 128x128, 224x224 images and I'm still yet to see any score above 30%. I'm still experiencing horrible overfitting.\n\nMy problem was why even the simple models (Let's face it, fine tuning a pretrained model is very simple in keras) doesn't work for me while the exact same implementation worked for someone who implemented it in fastai?",
    "445905": "Exactly! and that fact alone (simple models of fastai &gt; keras) is quite frustrating because I spent a good amount of time experimenting on Keras. The only thing that's now left to experiment is using an image size greater than 224x224 and aside from that I don't think there's difference between mine and the baseline fastai resnet50 model which got 62% LB",
    "445906": "Also, may I ask where are you learning fastai from?",
    "446055": "Actually, I'm just reading other kernels written in fast.ai, since the deep learning principles are the same and doesn't change between frameworks and we just need to learn another set of \"grammar\". Whenever I'm in doubt, I go at [fastai site](https://course.fast.ai/lessons/lesson2.html) (or just google the specific issue).",
    "446062": "\"while the exact same implementation worked for someone who implemented it in fastai?'\n0)Very important: \nrotation_range=60, from your kernel is huuuuge. Try 15 or 20 degrees. Or no rotation at all.\n1)Most important, fast.ai uses carefully tuned fancy one_cycle policy:\nhttps://docs.fast.ai/basic_train.html#fit_one_cycle\nIt is way better, than plain Adam.",
    "446066": "Yea I think that's the way to go about. I started with some documentations but then later realised Kaggle kernels don't support fastai v1 and now started with the videos. Reading a baseline fastai model now, let's get started with fastai now haha. Thanks for the input!",
    "446068": "I see, I didn't know about the onecycle policy, seems interesting. About the rotation range, I did start from 20 and went upto 60 in an attempt to cure the overfitting. I'm now starting with fastai, while trying out some minor attempts at the keras implementation. Thank you for your input :)",
    "446688": "I already checked 256*256*1 ,, I still not able to achieve above 0.25. I tried everything, I thought there is something wrong with batching, then I converted data into TFRecord. and tried to use Dataset API with Keras. I even written Resnet50 model in tensorflow, and tried to train this model on 2 1080Ti, took me lot of time to optimize it.  And still not able to go above  .27. I was about to give up today. Then I thought to check out in discussions. Now I starting to think to shift to fast.ai. Never tried it before, but if this is working then why not.",
    "446699": "I guess the faster we shift, the better. Although fastai is not quite documented properly I'd say, but if that's working, It'll do for me.",
    "447369": "You can program your own learning rate scheduler, such as one cycle, quite easily in Keras. I don't really see this as a compelling reason to switch to fastai, other than the fact that it holds your hand more than Keras? You lose the learning gained from understanding why the fastai training is more effective than the vanilla Keras implementation by just switching to \"what works\". That's fine if you just want to power past it, but you won't understand what was actually done behind the scenes. This is exactly what I was talking about when I said, \"It's better to understand your model rather than to apply anecdotal strategies that you don't understand.\" Best of luck.",
    "447461": "Hello Abhinav Verma,\nI think its always good to learn something new like fastai, but I wouldn't give up with your current code. I think there should be something wrong in your implementation. So did you implement a kernel in keras that is written with fastai? Thats a great think to do.\nIs the loss right? Is the network really the same? Did you give the right labels to the network? Did you shuffle the features but not the labels? Maybe you can post both codes, so we can compare.\n\nMaybe, you could try to train cifar10 or mnist with your script. And then you can find the solution step by step",
    "447876": "There is nothing wrong with Keras.\n\nThe \"problem\" is that fast.ai does a lot of magic under the hood, which is cool if you either are already familiar with the magic or you don't really care about understanding about what is going on. If you implement the magic in Keras (not really hard) you should be able to obtain the same results.\n\nMy current best submission is based on an internal (from work) framework build with TensorFlow+Keras and is on pair with the highest published (as far as I know) fast.ai (1.0) solution: https://www.kaggle.com/c/humpback-whale-identification/discussion/74647.\nThat fast.ai solution is, in comparison, a far more complex pipeline (including an oversampled dataset which probably has a relevant positive impact in the score).\n\nI can not share the code or use a Kernel but I would be happy to answer questions about the Keras code needed to replicate the result.\n\nMy settings are:\n\n- Input data:\n\nThe same bounding box dataset that everybody is using in the public kernels. new_whale class is ignored. \n\n- Model: \n\nPlain ResNet50 with \"big head\" using weights converted from keras-applications. The big head comes from replacing the last fully connected of the original model with:\n\n(starting from Global Average Pooling) -&gt; BatchNorm -&gt; Dropout(0.25) -&gt; Dense (2048) -&gt; BatchNorm -&gt; Dropout(0.5) -&gt; Dense(5004)\n\n- Batch size: 20\n\n- Input size: 200x600\n\n- Preprocesing:\n\nRescale to range(-1, 1) -&gt; random rotation (20 degree) -&gt; random brightness (0.2 delta) -&gt; random displace (5 pixels)\n\n- Epochs: 24\n\n- Optimizer: \n\nA modified version of Keras Adam adding support for \"True weight decay\" and \"Learning rate multipliers\" just like fast.ai. Settings:\n\nClip grad norm to 0.1\n\n\"True\" Weight decay: 0.0001\n\nMultipliers (the learning rate gets multiplied by this value for certain layers):\n\n0.01 for fist blocks.\n0.1 for middle blocks\n\n- Scheduler:\n\nOne cycle just like fast.ai (1.0) with defaults params. Note that the fast.ai one cycle has some relevant details that some public Keras implementations don't.\n\nMaximum learning rate of 0.001.",
    "447931": "What is the reasoning for not using bias in the Dense layer?",
    "447942": "Hi David,\n\nThank you for such an informative comment! \nAfter getting a little frustrated with Keras, I decided to try out fast.ai by trying out their MOOC. After learning about what fastai does under the hood, I'm now beginning to understand why it's performing better and how what I implemented thinking that it's as same as the fastai solution, is really not even close. Personally, I don't prefer the fast.ai approach (Not really understanding what's going on under the hood) and now trying out the different techniques that fast.ai implemented, on Keras. Let's hope for the best :)",
    "447943": "Hi corner200, \n\nAfter trying out a little of fastai, I'm back to Keras and now trying out different things (Like SGD with restarts). Let's see if things work out for the best :)",
    "447953": "since I am not active in this competition I just glanced over your code and couldn't find the mistake. but I am sure it should work with keras. If I can help, let me know. Try to start from something very simple, step by step you will find the solution!",
    "447955": "My mistake. use_bias was set to True which is the default",
    "448154": "Just changing:\n\n- Batch size: 10\n- Input size: 300x900\n\nRaises the score to 0.788\n\nInserting new_whale at threshold 0.8. \nThis threshold is quite relevant and properly tuning it would boost the score. So far I didn't find a way to tune this parameter locally without using the public scores. I believe that using the leaderboard would result in overfit to the 20% of the test set.",
    "458910": "Hey, thanks for the advice. I customised Keras Adam optimiser, added weight decays and multipliers, and achieved .66 LB with 300 300 1 and 0.4 as threshold. \n\nObviously I think  with higher h and w of images and optimizing the threshold, it able to achieve much higher LB.  \n\nBut yeah optimizing threshold by LB definitely result in overfit."
  },
  "source": "meta"
}