{
  "id": 199920,
  "title": "Misconceptions about fastai and PyTorch - Too Much Magic",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/199920",
  "author_name": "",
  "post_date": "2020-11-28T00:36:46.701192200Z",
  "votes": 19,
  "comment_count": 9,
  "views": 0,
  "content": "<p>There are many misconceptions that you must either go one way or the other when choosing to use PyTorch or fastai, and fastai has too much magic to perfectly reproduce what you've done in PyTorch inside the framework. In my kernel <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai/\" target=\"_blank\">here</a> I did my best to show an example where that <em>isn't</em> true.</p>\n<p>Some key factors people tend to glaze over (IMO):</p>\n<ul>\n<li>Make sure your augmentations are the same. Every library does things a bit differently. For example fastai does it's Hue and Saturation on RGB, while Albumentations does it on HSV. As a result the two have very different logics</li>\n<li>Make sure your schedulers are the same. In the example here torch's <code>CosineAnnealingWarmRestarts</code> is equivalent to running <code>fit_flat_cos</code> with a <code>start_pct</code> of 0 (just cosine annealing schedule)</li>\n</ul>\n<p>I'd like to use this discussion as a Q/A from Kagglers for fastai-related implementations. What are you having trouble with? What could there be more examples of? Did this example notebook help clear your thoughts on a few things, or did it just enact more confusion?</p>\n<p>Thanks :) </p>",
  "messages": [
    {
      "id": "1093718",
      "postDate": "11/28/2020 00:36:46",
      "content": "<p>There are many misconceptions that you must either go one way or the other when choosing to use PyTorch or fastai, and fastai has too much magic to perfectly reproduce what you've done in PyTorch inside the framework. In my kernel <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai/\" target=\"_blank\">here</a> I did my best to show an example where that <em>isn't</em> true.</p>\n<p>Some key factors people tend to glaze over (IMO):</p>\n<ul>\n<li>Make sure your augmentations are the same. Every library does things a bit differently. For example fastai does it's Hue and Saturation on RGB, while Albumentations does it on HSV. As a result the two have very different logics</li>\n<li>Make sure your schedulers are the same. In the example here torch's <code>CosineAnnealingWarmRestarts</code> is equivalent to running <code>fit_flat_cos</code> with a <code>start_pct</code> of 0 (just cosine annealing schedule)</li>\n</ul>\n<p>I'd like to use this discussion as a Q/A from Kagglers for fastai-related implementations. What are you having trouble with? What could there be more examples of? Did this example notebook help clear your thoughts on a few things, or did it just enact more confusion?</p>\n<p>Thanks :) </p>",
      "rawMarkdown": "There are many misconceptions that you must either go one way or the other when choosing to use PyTorch or fastai, and fastai has too much magic to perfectly reproduce what you've done in PyTorch inside the framework. In my kernel [here](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai/) I did my best to show an example where that *isn't* true.\n\nSome key factors people tend to glaze over (IMO):\n* Make sure your augmentations are the same. Every library does things a bit differently. For example fastai does it's Hue and Saturation on RGB, while Albumentations does it on HSV. As a result the two have very different logics\n* Make sure your schedulers are the same. In the example here torch's `CosineAnnealingWarmRestarts` is equivalent to running `fit_flat_cos` with a `start_pct` of 0 (just cosine annealing schedule)\n\nI'd like to use this discussion as a Q/A from Kagglers for fastai-related implementations. What are you having trouble with? What could there be more examples of? Did this example notebook help clear your thoughts on a few things, or did it just enact more confusion?\n\nThanks :)",
      "votes": null
    },
    {
      "id": "1093725",
      "postDate": "11/28/2020 00:59:13",
      "content": "<p>I think another major difference between other approaches to TTA and the fastai's version is that fastai interpolates the x number of predictions with a prediction on the center crop while other TTA approaches don't include it at all (they only the x number of predictions on random augs of data)… </p>\n<p>Setting beta=0 (I did this in my <a href=\"https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter\" target=\"_blank\">kernel</a>) will change it so be similar behavior as most kaggle kernels (but it still runs the center crop).</p>",
      "rawMarkdown": "I think another major difference between other approaches to TTA and the fastai's version is that fastai interpolates the x number of predictions with a prediction on the center crop while other TTA approaches don't include it at all (they only the x number of predictions on random augs of data)... \n\nSetting beta=0 (I did this in my [kernel](https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter)) will change it so be similar behavior as most kaggle kernels (but it still runs the center crop).",
      "votes": null
    },
    {
      "id": "1094215",
      "postDate": "11/28/2020 12:17:41",
      "content": "<p><a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> I know this is beat to death but I'd like to see the learning rate be more science than art (or magic ;)). Might be fun and useful to build this unless I'm missing it and it's been built but I haven't seen it incorporated in any notebooks. </p>\n<p>And, where is there more information on the wwf library you created? I was on your page but maybe I'm missing what it does? Is it pulling in all the functions to run timm?</p>",
      "rawMarkdown": "muellerzr I know this is beat to death but I'd like to see the learning rate be more science than art (or magic ;)). Might be fun and useful to build this unless I'm missing it and it's been built but I haven't seen it incorporated in any notebooks. \n\nAnd, where is there more information on the wwf library you created? I was on your page but maybe I'm missing what it does? Is it pulling in all the functions to run timm?",
      "votes": null
    },
    {
      "id": "1094216",
      "postDate": "11/28/2020 12:19:08",
      "content": "<p>Really good notebook btw. I was just about to build one from scratch and saw yours.. Always great to learn a new thing or two! And love that you and Zach used Timm. </p>",
      "rawMarkdown": "Really good notebook btw. I was just about to build one from scratch and saw yours.. Always great to learn a new thing or two! And love that you and Zach used Timm.",
      "votes": null
    },
    {
      "id": "1094390",
      "postDate": "11/28/2020 15:22:03",
      "content": "<p>Well, it is science! It's a published paper: <a href=\"https://arxiv.org/abs/1506.01186\" target=\"_blank\">https://arxiv.org/abs/1506.01186</a>. The science is picking the steepest slope found. We can algorithmically get this number back, which there are a number of different methods to do so. fastai has suggestions=True, and we have another example by Andrew Chang which goes into even more detail: <a href=\"https://forums.fast.ai/t/automated-learning-rate-suggester/44199\" target=\"_blank\">https://forums.fast.ai/t/automated-learning-rate-suggester/44199</a></p>",
      "rawMarkdown": "Well, it is science! It's a published paper: https://arxiv.org/abs/1506.01186. The science is picking the steepest slope found. We can algorithmically get this number back, which there are a number of different methods to do so. fastai has suggestions=True, and we have another example by Andrew Chang which goes into even more detail: https://forums.fast.ai/t/automated-learning-rate-suggester/44199",
      "votes": null
    },
    {
      "id": "1094391",
      "postDate": "11/28/2020 15:23:32",
      "content": "<p>walkwithfastai.com I wrote the timm example to show how fastai is working internally with its <code>cnn_learner</code> and adapting it for <code>timm</code>. The same code covered in the <code>timm</code> article is the ones being exported to <code>wwf.timm</code> through the power of nbdev</p>",
      "rawMarkdown": "walkwithfastai.com I wrote the timm example to show how fastai is working internally with its `cnn_learner` and adapting it for `timm`. The same code covered in the `timm` article is the ones being exported to `wwf.timm` through the power of nbdev",
      "votes": null
    },
    {
      "id": "1094460",
      "postDate": "11/28/2020 16:24:18",
      "content": "<p>Awesome. I will give that a try too.. I was adding the functions. But great that this was added. Appreciate the work. </p>\n<p>I had come across Andrew Chang's code. I'll put that into a version of my notebook :)</p>",
      "rawMarkdown": "Awesome. I will give that a try too.. I was adding the functions. But great that this was added. Appreciate the work. \n\nI had come across Andrew Chang's code. I'll put that into a version of my notebook :)",
      "votes": null
    },
    {
      "id": "1136881",
      "postDate": "01/03/2021 14:07:23",
      "content": "<p>Did anybody get TPUs to work with fastai? If so, do you upgrade to fastai2 on TPU-environment or downgrade to fastai1 on GPU-env. Or can I even use a learn.export()ed &amp; load_learner across both versions?</p>",
      "rawMarkdown": "Did anybody get TPUs to work with fastai? If so, do you upgrade to fastai2 on TPU-environment or downgrade to fastai1 on GPU-env. Or can I even use a learn.export()ed & load_learner across both versions?",
      "votes": null
    },
    {
      "id": "1137358",
      "postDate": "01/03/2021 21:35:29",
      "content": "<p>fastai doesn't currently have TPU support (but it is in somewhat active development).</p>",
      "rawMarkdown": "fastai doesn't currently have TPU support (but it is in somewhat active development).",
      "votes": null
    },
    {
      "id": "1137382",
      "postDate": "01/03/2021 22:03:49",
      "content": "<p>Just found this: <a href=\"https://www.kaggle.com/johnyquest/tpu-fastai-notebook\" target=\"_blank\">https://www.kaggle.com/johnyquest/tpu-fastai-notebook</a> <br>\nBut it didn't improve speed (maybe because BS is small in the example.). </p>",
      "rawMarkdown": "Just found this: https://www.kaggle.com/johnyquest/tpu-fastai-notebook \nBut it didn't improve speed (maybe because BS is small in the example.).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1093725,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "11/28/2020 00:59:13",
      "content": "<p>I think another major difference between other approaches to TTA and the fastai's version is that fastai interpolates the x number of predictions with a prediction on the center crop while other TTA approaches don't include it at all (they only the x number of predictions on random augs of data)… </p>\n<p>Setting beta=0 (I did this in my <a href=\"https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter\" target=\"_blank\">kernel</a>) will change it so be similar behavior as most kaggle kernels (but it still runs the center crop).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1094216,
          "author_name": "crained",
          "author_url": "",
          "post_date": "11/28/2020 12:19:08",
          "content": "<p>Really good notebook btw. I was just about to build one from scratch and saw yours.. Always great to learn a new thing or two! And love that you and Zach used Timm. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1094215,
      "author_name": "crained",
      "author_url": "",
      "post_date": "11/28/2020 12:17:41",
      "content": "<p><a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> I know this is beat to death but I'd like to see the learning rate be more science than art (or magic ;)). Might be fun and useful to build this unless I'm missing it and it's been built but I haven't seen it incorporated in any notebooks. </p>\n<p>And, where is there more information on the wwf library you created? I was on your page but maybe I'm missing what it does? Is it pulling in all the functions to run timm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1094390,
          "author_name": "muellerzr",
          "author_url": "",
          "post_date": "11/28/2020 15:22:03",
          "content": "<p>Well, it is science! It's a published paper: <a href=\"https://arxiv.org/abs/1506.01186\" target=\"_blank\">https://arxiv.org/abs/1506.01186</a>. The science is picking the steepest slope found. We can algorithmically get this number back, which there are a number of different methods to do so. fastai has suggestions=True, and we have another example by Andrew Chang which goes into even more detail: <a href=\"https://forums.fast.ai/t/automated-learning-rate-suggester/44199\" target=\"_blank\">https://forums.fast.ai/t/automated-learning-rate-suggester/44199</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094391,
          "author_name": "muellerzr",
          "author_url": "",
          "post_date": "11/28/2020 15:23:32",
          "content": "<p>walkwithfastai.com I wrote the timm example to show how fastai is working internally with its <code>cnn_learner</code> and adapting it for <code>timm</code>. The same code covered in the <code>timm</code> article is the ones being exported to <code>wwf.timm</code> through the power of nbdev</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094460,
          "author_name": "crained",
          "author_url": "",
          "post_date": "11/28/2020 16:24:18",
          "content": "<p>Awesome. I will give that a try too.. I was adding the functions. But great that this was added. Appreciate the work. </p>\n<p>I had come across Andrew Chang's code. I'll put that into a version of my notebook :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136881,
      "author_name": "joatom",
      "author_url": "",
      "post_date": "01/03/2021 14:07:23",
      "content": "<p>Did anybody get TPUs to work with fastai? If so, do you upgrade to fastai2 on TPU-environment or downgrade to fastai1 on GPU-env. Or can I even use a learn.export()ed &amp; load_learner across both versions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1137358,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "01/03/2021 21:35:29",
          "content": "<p>fastai doesn't currently have TPU support (but it is in somewhat active development).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1137382,
          "author_name": "joatom",
          "author_url": "",
          "post_date": "01/03/2021 22:03:49",
          "content": "<p>Just found this: <a href=\"https://www.kaggle.com/johnyquest/tpu-fastai-notebook\" target=\"_blank\">https://www.kaggle.com/johnyquest/tpu-fastai-notebook</a> <br>\nBut it didn't improve speed (maybe because BS is small in the example.). </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1093718": "There are many misconceptions that you must either go one way or the other when choosing to use PyTorch or fastai, and fastai has too much magic to perfectly reproduce what you've done in PyTorch inside the framework. In my kernel [here](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai/) I did my best to show an example where that *isn't* true.\n\nSome key factors people tend to glaze over (IMO):\n* Make sure your augmentations are the same. Every library does things a bit differently. For example fastai does it's Hue and Saturation on RGB, while Albumentations does it on HSV. As a result the two have very different logics\n* Make sure your schedulers are the same. In the example here torch's `CosineAnnealingWarmRestarts` is equivalent to running `fit_flat_cos` with a `start_pct` of 0 (just cosine annealing schedule)\n\nI'd like to use this discussion as a Q/A from Kagglers for fastai-related implementations. What are you having trouble with? What could there be more examples of? Did this example notebook help clear your thoughts on a few things, or did it just enact more confusion?\n\nThanks :)",
    "1093725": "I think another major difference between other approaches to TTA and the fastai's version is that fastai interpolates the x number of predictions with a prediction on the center crop while other TTA approaches don't include it at all (they only the x number of predictions on random augs of data)... \n\nSetting beta=0 (I did this in my [kernel](https://www.kaggle.com/tanlikesmath/cassava-classification-eda-fastai-starter)) will change it so be similar behavior as most kaggle kernels (but it still runs the center crop).",
    "1094215": "muellerzr I know this is beat to death but I'd like to see the learning rate be more science than art (or magic ;)). Might be fun and useful to build this unless I'm missing it and it's been built but I haven't seen it incorporated in any notebooks. \n\nAnd, where is there more information on the wwf library you created? I was on your page but maybe I'm missing what it does? Is it pulling in all the functions to run timm?",
    "1094216": "Really good notebook btw. I was just about to build one from scratch and saw yours.. Always great to learn a new thing or two! And love that you and Zach used Timm.",
    "1094390": "Well, it is science! It's a published paper: https://arxiv.org/abs/1506.01186. The science is picking the steepest slope found. We can algorithmically get this number back, which there are a number of different methods to do so. fastai has suggestions=True, and we have another example by Andrew Chang which goes into even more detail: https://forums.fast.ai/t/automated-learning-rate-suggester/44199",
    "1094391": "walkwithfastai.com I wrote the timm example to show how fastai is working internally with its `cnn_learner` and adapting it for `timm`. The same code covered in the `timm` article is the ones being exported to `wwf.timm` through the power of nbdev",
    "1094460": "Awesome. I will give that a try too.. I was adding the functions. But great that this was added. Appreciate the work. \n\nI had come across Andrew Chang's code. I'll put that into a version of my notebook :)",
    "1136881": "Did anybody get TPUs to work with fastai? If so, do you upgrade to fastai2 on TPU-environment or downgrade to fastai1 on GPU-env. Or can I even use a learn.export()ed & load_learner across both versions?",
    "1137358": "fastai doesn't currently have TPU support (but it is in somewhat active development).",
    "1137382": "Just found this: https://www.kaggle.com/johnyquest/tpu-fastai-notebook \nBut it didn't improve speed (maybe because BS is small in the example.)."
  },
  "source": "meta"
}