{
  "id": 344577,
  "title": "Top scores on leaderboard, Stacking/Ensembling or single model?",
  "url": "/competitions/amex-default-prediction/discussion/344577",
  "author_name": "",
  "post_date": "2022-08-15T18:28:24.301613800Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm a little curious,</p>\n<p>Almost all new notebooks with high scores are stacking/ensembling techniques, but all are stuck at 0.799.<br>\nAre the highest scores on the leaderboard using these methods or single models with deeper/different feature engineering and/or different loss/models?</p>",
  "messages": [
    {
      "id": "1900143",
      "postDate": "08/15/2022 18:28:24",
      "content": "<p>I'm a little curious,</p>\n<p>Almost all new notebooks with high scores are stacking/ensembling techniques, but all are stuck at 0.799.<br>\nAre the highest scores on the leaderboard using these methods or single models with deeper/different feature engineering and/or different loss/models?</p>",
      "rawMarkdown": "I'm a little curious,\n\nAlmost all new notebooks with high scores are stacking/ensembling techniques, but all are stuck at 0.799.\nAre the highest scores on the leaderboard using these methods or single models with deeper/different feature engineering and/or different loss/models?",
      "votes": null
    },
    {
      "id": "1900295",
      "postDate": "08/15/2022 21:57:04",
      "content": "<p>I don't think it's impossible that there's a single model that scores in the top 10 on public LB (especially because of how noisy the metric is), but I would be pretty surprised. If you can build one model with better features / architectures / general design you can probably get more than one of them and fruitfully ensemble, which is what I assume everyone at the top is doing. Sometimes you'll see competitions where people report that ensembling didn't work (talkingdata from some years back comes to mind), but I think that's much more likely with large training data (yes the raw data is large, but 450k training data labels is actually pretty small). In my view, noisy metric + small data = friendly to ensembling.</p>",
      "rawMarkdown": "I don't think it's impossible that there's a single model that scores in the top 10 on public LB (especially because of how noisy the metric is), but I would be pretty surprised. If you can build one model with better features / architectures / general design you can probably get more than one of them and fruitfully ensemble, which is what I assume everyone at the top is doing. Sometimes you'll see competitions where people report that ensembling didn't work (talkingdata from some years back comes to mind), but I think that's much more likely with large training data (yes the raw data is large, but 450k training data labels is actually pretty small). In my view, noisy metric + small data = friendly to ensembling.",
      "votes": null
    },
    {
      "id": "1901083",
      "postDate": "08/16/2022 13:05:18",
      "content": "<p>Thinking about becoming the top solutions productive, ensembling is not the best way to do it, right?<br>\nFrom the Amex point of view, I don't think it's the best solution, because you have to transform the new data, for each model, and infer for each model, and then do the ensembling…</p>",
      "rawMarkdown": "Thinking about becoming the top solutions productive, ensembling is not the best way to do it, right?\nFrom the Amex point of view, I don't think it's the best solution, because you have to transform the new data, for each model, and infer for each model, and then do the ensembling...",
      "votes": null
    },
    {
      "id": "1901128",
      "postDate": "08/16/2022 13:27:26",
      "content": "<p>That is a recurring theme going back to the Netflix Prize. The maintenance complexity of large ensembles rarely justifies the marginal performance improvement they allow. Small ensembles are a different question, and much more likely to be useable - this is probably already what Amex uses internally for this problem (see this <a href=\"https://arxiv.org/pdf/2012.15330.pdf\" target=\"_blank\">paper</a>). </p>\n<p>Generally I don't think hosts would expect to use solutions as-is, but instead can use them to pick up a few useful new ideas for feature engineering, preprocessing, model design, etc. A few nice FE ideas might easily be worth well over 100k to them in savings. Also, you could argue that the actual solutions are less important to the host than the marketing and recruiting opportunity the competition offers (I'm sure this varies host to host).</p>",
      "rawMarkdown": "That is a recurring theme going back to the Netflix Prize. The maintenance complexity of large ensembles rarely justifies the marginal performance improvement they allow. Small ensembles are a different question, and much more likely to be useable - this is probably already what Amex uses internally for this problem (see this [paper](https://arxiv.org/pdf/2012.15330.pdf)). \n\nGenerally I don't think hosts would expect to use solutions as-is, but instead can use them to pick up a few useful new ideas for feature engineering, preprocessing, model design, etc. A few nice FE ideas might easily be worth well over 100k to them in savings. Also, you could argue that the actual solutions are less important to the host than the marketing and recruiting opportunity the competition offers (I'm sure this varies host to host).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1900295,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "08/15/2022 21:57:04",
      "content": "<p>I don't think it's impossible that there's a single model that scores in the top 10 on public LB (especially because of how noisy the metric is), but I would be pretty surprised. If you can build one model with better features / architectures / general design you can probably get more than one of them and fruitfully ensemble, which is what I assume everyone at the top is doing. Sometimes you'll see competitions where people report that ensembling didn't work (talkingdata from some years back comes to mind), but I think that's much more likely with large training data (yes the raw data is large, but 450k training data labels is actually pretty small). In my view, noisy metric + small data = friendly to ensembling.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1901083,
          "author_name": "nandodmelo",
          "author_url": "",
          "post_date": "08/16/2022 13:05:18",
          "content": "<p>Thinking about becoming the top solutions productive, ensembling is not the best way to do it, right?<br>\nFrom the Amex point of view, I don't think it's the best solution, because you have to transform the new data, for each model, and infer for each model, and then do the ensembling…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1901128,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "08/16/2022 13:27:26",
          "content": "<p>That is a recurring theme going back to the Netflix Prize. The maintenance complexity of large ensembles rarely justifies the marginal performance improvement they allow. Small ensembles are a different question, and much more likely to be useable - this is probably already what Amex uses internally for this problem (see this <a href=\"https://arxiv.org/pdf/2012.15330.pdf\" target=\"_blank\">paper</a>). </p>\n<p>Generally I don't think hosts would expect to use solutions as-is, but instead can use them to pick up a few useful new ideas for feature engineering, preprocessing, model design, etc. A few nice FE ideas might easily be worth well over 100k to them in savings. Also, you could argue that the actual solutions are less important to the host than the marketing and recruiting opportunity the competition offers (I'm sure this varies host to host).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1900143": "I'm a little curious,\n\nAlmost all new notebooks with high scores are stacking/ensembling techniques, but all are stuck at 0.799.\nAre the highest scores on the leaderboard using these methods or single models with deeper/different feature engineering and/or different loss/models?",
    "1900295": "I don't think it's impossible that there's a single model that scores in the top 10 on public LB (especially because of how noisy the metric is), but I would be pretty surprised. If you can build one model with better features / architectures / general design you can probably get more than one of them and fruitfully ensemble, which is what I assume everyone at the top is doing. Sometimes you'll see competitions where people report that ensembling didn't work (talkingdata from some years back comes to mind), but I think that's much more likely with large training data (yes the raw data is large, but 450k training data labels is actually pretty small). In my view, noisy metric + small data = friendly to ensembling.",
    "1901083": "Thinking about becoming the top solutions productive, ensembling is not the best way to do it, right?\nFrom the Amex point of view, I don't think it's the best solution, because you have to transform the new data, for each model, and infer for each model, and then do the ensembling...",
    "1901128": "That is a recurring theme going back to the Netflix Prize. The maintenance complexity of large ensembles rarely justifies the marginal performance improvement they allow. Small ensembles are a different question, and much more likely to be useable - this is probably already what Amex uses internally for this problem (see this [paper](https://arxiv.org/pdf/2012.15330.pdf)). \n\nGenerally I don't think hosts would expect to use solutions as-is, but instead can use them to pick up a few useful new ideas for feature engineering, preprocessing, model design, etc. A few nice FE ideas might easily be worth well over 100k to them in savings. Also, you could argue that the actual solutions are less important to the host than the marketing and recruiting opportunity the competition offers (I'm sure this varies host to host)."
  },
  "source": "meta"
}