{
  "id": 213027,
  "title": "Any lightweight solutions?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/213027",
  "author_name": "",
  "post_date": "2021-01-21T08:14:35.195872400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Hello</h2>\n<p>This is the first non-playground competition I'm taking seriously, so please don't mind if I say something stupid. </p>\n<p>As far as I'm concerned, the most utilized model family in this competition providing up to 0.89 LB scores with not so much effort are EfficientNets (and B4 model specifically).<br>\nHowever, in one of my first entries, I used a rather simple convnet (<code>SeparableConv2D</code> layers) of only 90K parameters which gave me a 0.82 LB score. On the other hand, EfficientNetB4 is a pretty complex model of 17M parameters, and it seems like a huge cost for not so much accuracy gain (with everything else left the same, jumping from 90K convnet to EfficientNetB4 scored 0.88 LB).</p>\n<p>So I wonder maybe someone else tried to use some lightweight solutions and whether they did or didn't work?</p>",
  "messages": [
    {
      "id": "1162554",
      "postDate": "01/21/2021 08:14:35",
      "content": "<h2>Hello</h2>\n<p>This is the first non-playground competition I'm taking seriously, so please don't mind if I say something stupid. </p>\n<p>As far as I'm concerned, the most utilized model family in this competition providing up to 0.89 LB scores with not so much effort are EfficientNets (and B4 model specifically).<br>\nHowever, in one of my first entries, I used a rather simple convnet (<code>SeparableConv2D</code> layers) of only 90K parameters which gave me a 0.82 LB score. On the other hand, EfficientNetB4 is a pretty complex model of 17M parameters, and it seems like a huge cost for not so much accuracy gain (with everything else left the same, jumping from 90K convnet to EfficientNetB4 scored 0.88 LB).</p>\n<p>So I wonder maybe someone else tried to use some lightweight solutions and whether they did or didn't work?</p>",
      "rawMarkdown": "## Hello\n\nThis is the first non-playground competition I'm taking seriously, so please don't mind if I say something stupid. \n\nAs far as I'm concerned, the most utilized model family in this competition providing up to 0.89 LB scores with not so much effort are EfficientNets (and B4 model specifically).\nHowever, in one of my first entries, I used a rather simple convnet (`SeparableConv2D` layers) of only 90K parameters which gave me a 0.82 LB score. On the other hand, EfficientNetB4 is a pretty complex model of 17M parameters, and it seems like a huge cost for not so much accuracy gain (with everything else left the same, jumping from 90K convnet to EfficientNetB4 scored 0.88 LB).\n\nSo I wonder maybe someone else tried to use some lightweight solutions and whether they did or didn't work?",
      "votes": null
    },
    {
      "id": "1162652",
      "postDate": "01/21/2021 08:55:16",
      "content": "<p>You really think jumping from 0.82 to 0.88 in not so much accuracy gain ? </p>\n<p>I wonder now what those guys are doing with all these Imagenet SOTA the last 8 years ^^</p>",
      "rawMarkdown": "You really think jumping from 0.82 to 0.88 in not so much accuracy gain ? \n\nI wonder now what those guys are doing with all these Imagenet SOTA the last 8 years ^^",
      "votes": null
    },
    {
      "id": "1162713",
      "postDate": "01/21/2021 09:38:11",
      "content": "<p>I see this exaggeration now make look like I am intended to underestimate the immense work done by EfficientNet and other models' creators, which is definitely not so. <br>\nBut let me rather paraphrase the question: do you from your experience consider trying out here (and possibly on other competitions) lightweight models a waste of time?</p>",
      "rawMarkdown": "I see this exaggeration now make look like I am intended to underestimate the immense work done by EfficientNet and other models' creators, which is definitely not so. \nBut let me rather paraphrase the question: do you from your experience consider trying out here (and possibly on other competitions) lightweight models a waste of time?",
      "votes": null
    },
    {
      "id": "1162734",
      "postDate": "01/21/2021 10:01:26",
      "content": "<p>There's two different perspectives on this 1) if you want to beat the leaderboard in Kaggle and 2) for practical applications.</p>\n<p>From the first perspective, the difference between 0.89 and 0.91 is enormous. Closing that gap really requires a huge amount of work and a lot of tries - potentially with new or at least non-standard approaches, one's own tweaks of existing approaches etc. (and sure, sometimes something surprisingly simple might work great). As a side effect this also creates a great automatic testing ground for ML approaches, where good methods can get popularized quickly (some great examples are e.g. xgboost, LightGBM, and using embeddings for high cardinality categorical features in tabular neural networks) and become widely used standard methods.</p>\n<p>The difference between 0.89 and 0.91 in accuracy may not matter that much from the second perspective and other considerations such as fairness, interpretability, building the solution into the practical workflow, allowing for an appeal of wrong decisions to humans etc. may matter more than a little bit of accuracy (sure, if every extra click on an ad is a few more cents you earn, small improvements matter, but many use cases are not like that..). I'd say 0.82 vs. 0.88 does start to get relatively big. So, learning the widely accepted good ideas that almost always work (e.g. transfer learning, what's currently good architectures, data augmentation etc.) is important. </p>\n<p>My minimum ambition in a competition is that I learn enough so that I can implement a solution, for which the gap to top solutions is not so meaningful in a practical sense (even if it's a huge number of ranks on the leaderboard). At least then I've learnt something useful and can deal with such a problem in practice. Of course, medals would be nice. ;)</p>",
      "rawMarkdown": "There's two different perspectives on this 1) if you want to beat the leaderboard in Kaggle and 2) for practical applications.\n\nFrom the first perspective, the difference between 0.89 and 0.91 is enormous. Closing that gap really requires a huge amount of work and a lot of tries - potentially with new or at least non-standard approaches, one's own tweaks of existing approaches etc. (and sure, sometimes something surprisingly simple might work great). As a side effect this also creates a great automatic testing ground for ML approaches, where good methods can get popularized quickly (some great examples are e.g. xgboost, LightGBM, and using embeddings for high cardinality categorical features in tabular neural networks) and become widely used standard methods.\n\nThe difference between 0.89 and 0.91 in accuracy may not matter that much from the second perspective and other considerations such as fairness, interpretability, building the solution into the practical workflow, allowing for an appeal of wrong decisions to humans etc. may matter more than a little bit of accuracy (sure, if every extra click on an ad is a few more cents you earn, small improvements matter, but many use cases are not like that..). I'd say 0.82 vs. 0.88 does start to get relatively big. So, learning the widely accepted good ideas that almost always work (e.g. transfer learning, what's currently good architectures, data augmentation etc.) is important. \n\nMy minimum ambition in a competition is that I learn enough so that I can implement a solution, for which the gap to top solutions is not so meaningful in a practical sense (even if it's a huge number of ranks on the leaderboard). At least then I've learnt something useful and can deal with such a problem in practice. Of course, medals would be nice. ;)",
      "votes": null
    },
    {
      "id": "1162776",
      "postDate": "01/21/2021 10:38:09",
      "content": "<p>I was about to reply but <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> did it much better I could do. </p>",
      "rawMarkdown": "I was about to reply but @bjoernholzhauer did it much better I could do.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1162652,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/21/2021 08:55:16",
      "content": "<p>You really think jumping from 0.82 to 0.88 in not so much accuracy gain ? </p>\n<p>I wonder now what those guys are doing with all these Imagenet SOTA the last 8 years ^^</p>",
      "votes": null,
      "replies": [
        {
          "id": 1162713,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "01/21/2021 09:38:11",
          "content": "<p>I see this exaggeration now make look like I am intended to underestimate the immense work done by EfficientNet and other models' creators, which is definitely not so. <br>\nBut let me rather paraphrase the question: do you from your experience consider trying out here (and possibly on other competitions) lightweight models a waste of time?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1162776,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/21/2021 10:38:09",
          "content": "<p>I was about to reply but <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> did it much better I could do. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1162734,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "01/21/2021 10:01:26",
      "content": "<p>There's two different perspectives on this 1) if you want to beat the leaderboard in Kaggle and 2) for practical applications.</p>\n<p>From the first perspective, the difference between 0.89 and 0.91 is enormous. Closing that gap really requires a huge amount of work and a lot of tries - potentially with new or at least non-standard approaches, one's own tweaks of existing approaches etc. (and sure, sometimes something surprisingly simple might work great). As a side effect this also creates a great automatic testing ground for ML approaches, where good methods can get popularized quickly (some great examples are e.g. xgboost, LightGBM, and using embeddings for high cardinality categorical features in tabular neural networks) and become widely used standard methods.</p>\n<p>The difference between 0.89 and 0.91 in accuracy may not matter that much from the second perspective and other considerations such as fairness, interpretability, building the solution into the practical workflow, allowing for an appeal of wrong decisions to humans etc. may matter more than a little bit of accuracy (sure, if every extra click on an ad is a few more cents you earn, small improvements matter, but many use cases are not like that..). I'd say 0.82 vs. 0.88 does start to get relatively big. So, learning the widely accepted good ideas that almost always work (e.g. transfer learning, what's currently good architectures, data augmentation etc.) is important. </p>\n<p>My minimum ambition in a competition is that I learn enough so that I can implement a solution, for which the gap to top solutions is not so meaningful in a practical sense (even if it's a huge number of ranks on the leaderboard). At least then I've learnt something useful and can deal with such a problem in practice. Of course, medals would be nice. ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1162554": "## Hello\n\nThis is the first non-playground competition I'm taking seriously, so please don't mind if I say something stupid. \n\nAs far as I'm concerned, the most utilized model family in this competition providing up to 0.89 LB scores with not so much effort are EfficientNets (and B4 model specifically).\nHowever, in one of my first entries, I used a rather simple convnet (`SeparableConv2D` layers) of only 90K parameters which gave me a 0.82 LB score. On the other hand, EfficientNetB4 is a pretty complex model of 17M parameters, and it seems like a huge cost for not so much accuracy gain (with everything else left the same, jumping from 90K convnet to EfficientNetB4 scored 0.88 LB).\n\nSo I wonder maybe someone else tried to use some lightweight solutions and whether they did or didn't work?",
    "1162652": "You really think jumping from 0.82 to 0.88 in not so much accuracy gain ? \n\nI wonder now what those guys are doing with all these Imagenet SOTA the last 8 years ^^",
    "1162713": "I see this exaggeration now make look like I am intended to underestimate the immense work done by EfficientNet and other models' creators, which is definitely not so. \nBut let me rather paraphrase the question: do you from your experience consider trying out here (and possibly on other competitions) lightweight models a waste of time?",
    "1162734": "There's two different perspectives on this 1) if you want to beat the leaderboard in Kaggle and 2) for practical applications.\n\nFrom the first perspective, the difference between 0.89 and 0.91 is enormous. Closing that gap really requires a huge amount of work and a lot of tries - potentially with new or at least non-standard approaches, one's own tweaks of existing approaches etc. (and sure, sometimes something surprisingly simple might work great). As a side effect this also creates a great automatic testing ground for ML approaches, where good methods can get popularized quickly (some great examples are e.g. xgboost, LightGBM, and using embeddings for high cardinality categorical features in tabular neural networks) and become widely used standard methods.\n\nThe difference between 0.89 and 0.91 in accuracy may not matter that much from the second perspective and other considerations such as fairness, interpretability, building the solution into the practical workflow, allowing for an appeal of wrong decisions to humans etc. may matter more than a little bit of accuracy (sure, if every extra click on an ad is a few more cents you earn, small improvements matter, but many use cases are not like that..). I'd say 0.82 vs. 0.88 does start to get relatively big. So, learning the widely accepted good ideas that almost always work (e.g. transfer learning, what's currently good architectures, data augmentation etc.) is important. \n\nMy minimum ambition in a competition is that I learn enough so that I can implement a solution, for which the gap to top solutions is not so meaningful in a practical sense (even if it's a huge number of ranks on the leaderboard). At least then I've learnt something useful and can deal with such a problem in practice. Of course, medals would be nice. ;)",
    "1162776": "I was about to reply but @bjoernholzhauer did it much better I could do."
  },
  "source": "meta"
}