{
  "id": 224582,
  "title": "Any hard and fast rules on the addition of custom layers",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/224582",
  "author_name": "",
  "post_date": "2021-03-09T05:50:37.979153200Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Dear all, this is a more generic question on transfer learning, in general, if you are <strong>finetuning</strong> your model, say, an efficientnet from <code>timm</code>; do you further add custom layers on top of the last layer? That is to say, are there any anecdotally better \"layers\" than others. Like add a dense layer with swish activation?</p>",
  "messages": [
    {
      "id": "1231627",
      "postDate": "03/09/2021 05:50:37",
      "content": "<p>Dear all, this is a more generic question on transfer learning, in general, if you are <strong>finetuning</strong> your model, say, an efficientnet from <code>timm</code>; do you further add custom layers on top of the last layer? That is to say, are there any anecdotally better \"layers\" than others. Like add a dense layer with swish activation?</p>",
      "rawMarkdown": "Dear all, this is a more generic question on transfer learning, in general, if you are **finetuning** your model, say, an efficientnet from `timm`; do you further add custom layers on top of the last layer? That is to say, are there any anecdotally better \"layers\" than others. Like add a dense layer with swish activation?",
      "votes": null
    },
    {
      "id": "1231759",
      "postDate": "03/09/2021 08:18:53",
      "content": "<p>There's definitely quite a few options out there, but I've not seen a systematic study of the topic. From what they describe in their book (see below), it looks like the <code>fastai</code> team/Jeremy Howard experimented on this, but I've only found their high-level summary of their findings. Here's some of the options I'm aware of:</p>\n<ul>\n<li>It looks to me like e.g. the EfficientNet-B0 seems to just do <strong>Pooling</strong> (not clear which, max? avg? It's AdaptiveAvgPool2d in the <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py\" target=\"_blank\">EfficientNet-PyTorch repo</a>) -&gt; <strong>Flatten</strong> -&gt; <strong>Linear to number of targets</strong> (see <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">arXiv paper</a>), without even a drop-out in there (see <a href=\"https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/efficientnet_model.py\" target=\"_blank\">GitHub</a>, while in the <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py\" target=\"_blank\">EfficientNet-PyTorch repo</a> there is dropout there). Given that that paper is all about optimizing network architecture and SoTA, you might think they would not have skipped easy gains via a head, but who knows.</li>\n<li>There's also the same thing, just with AdaptiveConcatPool2d (concatenating average and max pooling - and of course the variant with just max pooling), with and without dropout that seems to get used a good bit.</li>\n<li>In the Cassava competition, I think I saw some suggestion that <strong>Pooling</strong> -&gt; <strong>Flatten</strong> -&gt; <strong>Linear(512)</strong> -&gt; <strong>BatchNorm</strong> -&gt; <strong>SiLU</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear to number of targets</strong> could be a good head.</li>\n<li>The <code>fastai</code> team seem to have looked into that topic in some depth. E.g. the book chapter 15 of their book (see the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">github for the book</a>) describes  <strong>AdaptiveConcatPool2d or AdaptiveAvgPool2d</strong> -&gt; <strong>flattening</strong> -&gt; <strong>BN</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear</strong> -&gt; <strong>ReLU</strong> -&gt; <strong>BN</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear to number of classes</strong>. They say that this two-Linear-layer type of head tends to be a bit better than simpler alternatives. The particular set and order of layers is just one of the options in their <a href=\"https://docs.fast.ai/vision.learner.html#create_head\" target=\"_blank\">create_head</a> function, which I found quite interesting to look at and I've tried a few of the options in there (e.g. a BN first, or a final BN before the output, which they suggest sometimes helps).</li>\n</ul>\n<p>I have the impression from my experience of cross-validating across two vision competitions (this one and Cassava - so take this limited experience with a grain of salt) that very often the extra layers à la fastai do help a little, but only with the right training strategy (e.g. you really need to train the head of the network without unfreezing the layers below, if you have a more complex head, while it seems less important with super simple heads), sufficient dropout for regularization and with a sufficiently large batch size (otherwise the BN layers could be problematic - e.g. in this competition I've failed to make them work on top of ResNet200D, I suspect because I had to train with a small batch size, but that's just my speculation).</p>",
      "rawMarkdown": "There's definitely quite a few options out there, but I've not seen a systematic study of the topic. From what they describe in their book (see below), it looks like the `fastai` team/Jeremy Howard experimented on this, but I've only found their high-level summary of their findings. Here's some of the options I'm aware of:\n* It looks to me like e.g. the EfficientNet-B0 seems to just do **Pooling** (not clear which, max? avg? It's AdaptiveAvgPool2d in the [EfficientNet-PyTorch repo](https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py)) -> **Flatten** -> **Linear to number of targets** (see [arXiv paper](https://arxiv.org/abs/1905.11946)), without even a drop-out in there (see [GitHub](https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/efficientnet_model.py), while in the [EfficientNet-PyTorch repo](https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py) there is dropout there). Given that that paper is all about optimizing network architecture and SoTA, you might think they would not have skipped easy gains via a head, but who knows.\n* There's also the same thing, just with AdaptiveConcatPool2d (concatenating average and max pooling - and of course the variant with just max pooling), with and without dropout that seems to get used a good bit.\n* In the Cassava competition, I think I saw some suggestion that **Pooling** -> **Flatten** -> **Linear(512)** -> **BatchNorm** -> **SiLU** -> **Dropout** -> **Linear to number of targets** could be a good head.\n* The `fastai` team seem to have looked into that topic in some depth. E.g. the book chapter 15 of their book (see the [github for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb)) describes  **AdaptiveConcatPool2d or AdaptiveAvgPool2d** -> **flattening** -> **BN** -> **Dropout** -> **Linear** -> **ReLU** -> **BN** -> **Dropout** -> **Linear to number of classes**. They say that this two-Linear-layer type of head tends to be a bit better than simpler alternatives. The particular set and order of layers is just one of the options in their [create_head](https://docs.fast.ai/vision.learner.html#create_head) function, which I found quite interesting to look at and I've tried a few of the options in there (e.g. a BN first, or a final BN before the output, which they suggest sometimes helps).\n\nI have the impression from my experience of cross-validating across two vision competitions (this one and Cassava - so take this limited experience with a grain of salt) that very often the extra layers à la fastai do help a little, but only with the right training strategy (e.g. you really need to train the head of the network without unfreezing the layers below, if you have a more complex head, while it seems less important with super simple heads), sufficient dropout for regularization and with a sufficiently large batch size (otherwise the BN layers could be problematic - e.g. in this competition I've failed to make them work on top of ResNet200D, I suspect because I had to train with a small batch size, but that's just my speculation).",
      "votes": null
    },
    {
      "id": "1232084",
      "postDate": "03/09/2021 13:29:49",
      "content": "<p>Thanks for such a detailed reply, always good to hear insights like this. I agree, lately I have been trying to follow a non-beginner course for deep learning, do you think fastai is a good one?</p>",
      "rawMarkdown": "Thanks for such a detailed reply, always good to hear insights like this. I agree, lately I have been trying to follow a non-beginner course for deep learning, do you think fastai is a good one?",
      "votes": null
    },
    {
      "id": "1232137",
      "postDate": "03/09/2021 13:58:28",
      "content": "<p>I love the course, but that's partially because I like their teaching philosophy (i.e. start with the problem/application, then increasingly dig into the details/theory - if slowly building up the theory first works better for you, then the style may not be your thing) and because it was my second learning resource after François Chollet's book 3 years ago. The course is intended for DL beginners with a coding background, but moves pretty fast and gets to reasonably advanced topics. I've heard plenty of interviews with Kagglers, where they've stated that even despite being more experienced they found the course useful and even watch each new iteration of it. Given that you've done well in several vision competitions, I have no real idea how well it will suit you. The great thing is, of course, that you can just start <a href=\"https://course.fast.ai/\" target=\"_blank\">listening to the videos</a> for free or take a peek at the <a href=\"https://github.com/fastai/fastbook/\" target=\"_blank\">book in the GitHub repo</a> to get an impression.</p>",
      "rawMarkdown": "I love the course, but that's partially because I like their teaching philosophy (i.e. start with the problem/application, then increasingly dig into the details/theory - if slowly building up the theory first works better for you, then the style may not be your thing) and because it was my second learning resource after François Chollet's book 3 years ago. The course is intended for DL beginners with a coding background, but moves pretty fast and gets to reasonably advanced topics. I've heard plenty of interviews with Kagglers, where they've stated that even despite being more experienced they found the course useful and even watch each new iteration of it. Given that you've done well in several vision competitions, I have no real idea how well it will suit you. The great thing is, of course, that you can just start [listening to the videos](https://course.fast.ai/) for free or take a peek at the [book in the GitHub repo](https://github.com/fastai/fastbook/) to get an impression.",
      "votes": null
    },
    {
      "id": "1232192",
      "postDate": "03/09/2021 14:41:46",
      "content": "<p>Thanks a bunch. I realised you have a PhD in math, while I’m just a degree holder in math. What’s the best course or book that can come hand in hand with fastai to complenent the mathematical aspects </p>",
      "rawMarkdown": "Thanks a bunch. I realised you have a PhD in math, while I’m just a degree holder in math. What’s the best course or book that can come hand in hand with fastai to complenent the mathematical aspects",
      "votes": null
    },
    {
      "id": "1232228",
      "postDate": "03/09/2021 15:19:40",
      "content": "<p>That's of course an endless debate - i.e. what extent of a maths background does one need for ML/DL etc., I'm sure it helps, but how much so? I certainly have not needed my measure theory from my undergraduate degree for a long time (and I did go into a much more applied statistics direction with the PhD I did after working in industry for some years). A key question is what's the opportunity cost, when you could strengthen your maths vs. doing something else like understanding issues around sampling/populations/causal inference vs. getting practical experience? There's some minimum level of mathematical understanding/background, where the lack of it would really get in the way all the time, but I'm not really sure whether those with a quantitative background / degree in some subject like maths/physics/computer science have a problem there. In any case, I guess it depends a lot on your ambitions (I'm sure there's theoretical proofs to be derived about optimization for neural networks that require a super-strong mathematical background, but unless that's your goal…).</p>\n<p>To complement fast.ai, there's always <a href=\"https://www.fast.ai/2017/07/17/num-lin-alg/\" target=\"_blank\">this option</a> that has a heavy ML focus.</p>",
      "rawMarkdown": "That's of course an endless debate - i.e. what extent of a maths background does one need for ML/DL etc., I'm sure it helps, but how much so? I certainly have not needed my measure theory from my undergraduate degree for a long time (and I did go into a much more applied statistics direction with the PhD I did after working in industry for some years). A key question is what's the opportunity cost, when you could strengthen your maths vs. doing something else like understanding issues around sampling/populations/causal inference vs. getting practical experience? There's some minimum level of mathematical understanding/background, where the lack of it would really get in the way all the time, but I'm not really sure whether those with a quantitative background / degree in some subject like maths/physics/computer science have a problem there. In any case, I guess it depends a lot on your ambitions (I'm sure there's theoretical proofs to be derived about optimization for neural networks that require a super-strong mathematical background, but unless that's your goal...).\n\nTo complement fast.ai, there's always [this option](https://www.fast.ai/2017/07/17/num-lin-alg/) that has a heavy ML focus.",
      "votes": null
    },
    {
      "id": "1232282",
      "postDate": "03/09/2021 16:04:02",
      "content": "<p>Its really hard to say for me, sometimes  it works and sometimes it doesnt.. One thing for sure is that the bigger the head the longer it takes to converge (in my experience). And taking longer to converge isnt a bad thing if you can get better results out of it but when comparing these two heads i often see that the bigger head requires more training to get to the same loss.</p>",
      "rawMarkdown": "Its really hard to say for me, sometimes  it works and sometimes it doesnt.. One thing for sure is that the bigger the head the longer it takes to converge (in my experience). And taking longer to converge isnt a bad thing if you can get better results out of it but when comparing these two heads i often see that the bigger head requires more training to get to the same loss.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1231759,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/09/2021 08:18:53",
      "content": "<p>There's definitely quite a few options out there, but I've not seen a systematic study of the topic. From what they describe in their book (see below), it looks like the <code>fastai</code> team/Jeremy Howard experimented on this, but I've only found their high-level summary of their findings. Here's some of the options I'm aware of:</p>\n<ul>\n<li>It looks to me like e.g. the EfficientNet-B0 seems to just do <strong>Pooling</strong> (not clear which, max? avg? It's AdaptiveAvgPool2d in the <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py\" target=\"_blank\">EfficientNet-PyTorch repo</a>) -&gt; <strong>Flatten</strong> -&gt; <strong>Linear to number of targets</strong> (see <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">arXiv paper</a>), without even a drop-out in there (see <a href=\"https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/efficientnet_model.py\" target=\"_blank\">GitHub</a>, while in the <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py\" target=\"_blank\">EfficientNet-PyTorch repo</a> there is dropout there). Given that that paper is all about optimizing network architecture and SoTA, you might think they would not have skipped easy gains via a head, but who knows.</li>\n<li>There's also the same thing, just with AdaptiveConcatPool2d (concatenating average and max pooling - and of course the variant with just max pooling), with and without dropout that seems to get used a good bit.</li>\n<li>In the Cassava competition, I think I saw some suggestion that <strong>Pooling</strong> -&gt; <strong>Flatten</strong> -&gt; <strong>Linear(512)</strong> -&gt; <strong>BatchNorm</strong> -&gt; <strong>SiLU</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear to number of targets</strong> could be a good head.</li>\n<li>The <code>fastai</code> team seem to have looked into that topic in some depth. E.g. the book chapter 15 of their book (see the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">github for the book</a>) describes  <strong>AdaptiveConcatPool2d or AdaptiveAvgPool2d</strong> -&gt; <strong>flattening</strong> -&gt; <strong>BN</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear</strong> -&gt; <strong>ReLU</strong> -&gt; <strong>BN</strong> -&gt; <strong>Dropout</strong> -&gt; <strong>Linear to number of classes</strong>. They say that this two-Linear-layer type of head tends to be a bit better than simpler alternatives. The particular set and order of layers is just one of the options in their <a href=\"https://docs.fast.ai/vision.learner.html#create_head\" target=\"_blank\">create_head</a> function, which I found quite interesting to look at and I've tried a few of the options in there (e.g. a BN first, or a final BN before the output, which they suggest sometimes helps).</li>\n</ul>\n<p>I have the impression from my experience of cross-validating across two vision competitions (this one and Cassava - so take this limited experience with a grain of salt) that very often the extra layers à la fastai do help a little, but only with the right training strategy (e.g. you really need to train the head of the network without unfreezing the layers below, if you have a more complex head, while it seems less important with super simple heads), sufficient dropout for regularization and with a sufficiently large batch size (otherwise the BN layers could be problematic - e.g. in this competition I've failed to make them work on top of ResNet200D, I suspect because I had to train with a small batch size, but that's just my speculation).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1232084,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 13:29:49",
          "content": "<p>Thanks for such a detailed reply, always good to hear insights like this. I agree, lately I have been trying to follow a non-beginner course for deep learning, do you think fastai is a good one?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1232137,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/09/2021 13:58:28",
          "content": "<p>I love the course, but that's partially because I like their teaching philosophy (i.e. start with the problem/application, then increasingly dig into the details/theory - if slowly building up the theory first works better for you, then the style may not be your thing) and because it was my second learning resource after François Chollet's book 3 years ago. The course is intended for DL beginners with a coding background, but moves pretty fast and gets to reasonably advanced topics. I've heard plenty of interviews with Kagglers, where they've stated that even despite being more experienced they found the course useful and even watch each new iteration of it. Given that you've done well in several vision competitions, I have no real idea how well it will suit you. The great thing is, of course, that you can just start <a href=\"https://course.fast.ai/\" target=\"_blank\">listening to the videos</a> for free or take a peek at the <a href=\"https://github.com/fastai/fastbook/\" target=\"_blank\">book in the GitHub repo</a> to get an impression.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1232192,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/09/2021 14:41:46",
          "content": "<p>Thanks a bunch. I realised you have a PhD in math, while I’m just a degree holder in math. What’s the best course or book that can come hand in hand with fastai to complenent the mathematical aspects </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1232228,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/09/2021 15:19:40",
          "content": "<p>That's of course an endless debate - i.e. what extent of a maths background does one need for ML/DL etc., I'm sure it helps, but how much so? I certainly have not needed my measure theory from my undergraduate degree for a long time (and I did go into a much more applied statistics direction with the PhD I did after working in industry for some years). A key question is what's the opportunity cost, when you could strengthen your maths vs. doing something else like understanding issues around sampling/populations/causal inference vs. getting practical experience? There's some minimum level of mathematical understanding/background, where the lack of it would really get in the way all the time, but I'm not really sure whether those with a quantitative background / degree in some subject like maths/physics/computer science have a problem there. In any case, I guess it depends a lot on your ambitions (I'm sure there's theoretical proofs to be derived about optimization for neural networks that require a super-strong mathematical background, but unless that's your goal…).</p>\n<p>To complement fast.ai, there's always <a href=\"https://www.fast.ai/2017/07/17/num-lin-alg/\" target=\"_blank\">this option</a> that has a heavy ML focus.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1232282,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "03/09/2021 16:04:02",
      "content": "<p>Its really hard to say for me, sometimes  it works and sometimes it doesnt.. One thing for sure is that the bigger the head the longer it takes to converge (in my experience). And taking longer to converge isnt a bad thing if you can get better results out of it but when comparing these two heads i often see that the bigger head requires more training to get to the same loss.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1231627": "Dear all, this is a more generic question on transfer learning, in general, if you are **finetuning** your model, say, an efficientnet from `timm`; do you further add custom layers on top of the last layer? That is to say, are there any anecdotally better \"layers\" than others. Like add a dense layer with swish activation?",
    "1231759": "There's definitely quite a few options out there, but I've not seen a systematic study of the topic. From what they describe in their book (see below), it looks like the `fastai` team/Jeremy Howard experimented on this, but I've only found their high-level summary of their findings. Here's some of the options I'm aware of:\n* It looks to me like e.g. the EfficientNet-B0 seems to just do **Pooling** (not clear which, max? avg? It's AdaptiveAvgPool2d in the [EfficientNet-PyTorch repo](https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py)) -> **Flatten** -> **Linear to number of targets** (see [arXiv paper](https://arxiv.org/abs/1905.11946)), without even a drop-out in there (see [GitHub](https://github.com/tensorflow/tpu/blob/master/models/official/efficientnet/efficientnet_model.py), while in the [EfficientNet-PyTorch repo](https://github.com/lukemelas/EfficientNet-PyTorch/blob/master/efficientnet_pytorch/model.py) there is dropout there). Given that that paper is all about optimizing network architecture and SoTA, you might think they would not have skipped easy gains via a head, but who knows.\n* There's also the same thing, just with AdaptiveConcatPool2d (concatenating average and max pooling - and of course the variant with just max pooling), with and without dropout that seems to get used a good bit.\n* In the Cassava competition, I think I saw some suggestion that **Pooling** -> **Flatten** -> **Linear(512)** -> **BatchNorm** -> **SiLU** -> **Dropout** -> **Linear to number of targets** could be a good head.\n* The `fastai` team seem to have looked into that topic in some depth. E.g. the book chapter 15 of their book (see the [github for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb)) describes  **AdaptiveConcatPool2d or AdaptiveAvgPool2d** -> **flattening** -> **BN** -> **Dropout** -> **Linear** -> **ReLU** -> **BN** -> **Dropout** -> **Linear to number of classes**. They say that this two-Linear-layer type of head tends to be a bit better than simpler alternatives. The particular set and order of layers is just one of the options in their [create_head](https://docs.fast.ai/vision.learner.html#create_head) function, which I found quite interesting to look at and I've tried a few of the options in there (e.g. a BN first, or a final BN before the output, which they suggest sometimes helps).\n\nI have the impression from my experience of cross-validating across two vision competitions (this one and Cassava - so take this limited experience with a grain of salt) that very often the extra layers à la fastai do help a little, but only with the right training strategy (e.g. you really need to train the head of the network without unfreezing the layers below, if you have a more complex head, while it seems less important with super simple heads), sufficient dropout for regularization and with a sufficiently large batch size (otherwise the BN layers could be problematic - e.g. in this competition I've failed to make them work on top of ResNet200D, I suspect because I had to train with a small batch size, but that's just my speculation).",
    "1232084": "Thanks for such a detailed reply, always good to hear insights like this. I agree, lately I have been trying to follow a non-beginner course for deep learning, do you think fastai is a good one?",
    "1232137": "I love the course, but that's partially because I like their teaching philosophy (i.e. start with the problem/application, then increasingly dig into the details/theory - if slowly building up the theory first works better for you, then the style may not be your thing) and because it was my second learning resource after François Chollet's book 3 years ago. The course is intended for DL beginners with a coding background, but moves pretty fast and gets to reasonably advanced topics. I've heard plenty of interviews with Kagglers, where they've stated that even despite being more experienced they found the course useful and even watch each new iteration of it. Given that you've done well in several vision competitions, I have no real idea how well it will suit you. The great thing is, of course, that you can just start [listening to the videos](https://course.fast.ai/) for free or take a peek at the [book in the GitHub repo](https://github.com/fastai/fastbook/) to get an impression.",
    "1232192": "Thanks a bunch. I realised you have a PhD in math, while I’m just a degree holder in math. What’s the best course or book that can come hand in hand with fastai to complenent the mathematical aspects",
    "1232228": "That's of course an endless debate - i.e. what extent of a maths background does one need for ML/DL etc., I'm sure it helps, but how much so? I certainly have not needed my measure theory from my undergraduate degree for a long time (and I did go into a much more applied statistics direction with the PhD I did after working in industry for some years). A key question is what's the opportunity cost, when you could strengthen your maths vs. doing something else like understanding issues around sampling/populations/causal inference vs. getting practical experience? There's some minimum level of mathematical understanding/background, where the lack of it would really get in the way all the time, but I'm not really sure whether those with a quantitative background / degree in some subject like maths/physics/computer science have a problem there. In any case, I guess it depends a lot on your ambitions (I'm sure there's theoretical proofs to be derived about optimization for neural networks that require a super-strong mathematical background, but unless that's your goal...).\n\nTo complement fast.ai, there's always [this option](https://www.fast.ai/2017/07/17/num-lin-alg/) that has a heavy ML focus.",
    "1232282": "Its really hard to say for me, sometimes  it works and sometimes it doesnt.. One thing for sure is that the bigger the head the longer it takes to converge (in my experience). And taking longer to converge isnt a bad thing if you can get better results out of it but when comparing these two heads i often see that the bigger head requires more training to get to the same loss."
  },
  "source": "meta"
}