{
  "id": 582801,
  "title": "A Better Way To Ensemble",
  "url": "/competitions/waveform-inversion/discussion/582801",
  "author_name": "",
  "post_date": "2025-06-02T21:48:06.340450600Z",
  "votes": 51,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I feel I have to give back a bit after all the goodies <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> has been generously sharing.</p>\n<p>Here is a better implementation of the ensemble model from his notebooks. Using the median is better than using the mean when the objective is MAE. </p>\n<p>This implementation does not use <code>torch.median()</code> because there is a little catch. When the number of entries is even, then there are two possible values of the median. For instance, the median of [1, 2, 3, 4, 5, 6] can be either 3 or 4. When using <code>quantile=0.5</code> the average of the two medians is taken.</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .models = nn.ModuleList(models).()\n\n     ():\n        output = []\n\n         m  .models:\n            logits = m(x)\n\n            output.append(logits)\n\n        output = torch.stack(output)\n        output = torch.quantile(output, , dim=)\n         output\n</code></pre>",
  "messages": [
    {
      "id": "3215937",
      "postDate": "06/02/2025 21:48:06",
      "content": "<p>I feel I have to give back a bit after all the goodies <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> has been generously sharing.</p>\n<p>Here is a better implementation of the ensemble model from his notebooks. Using the median is better than using the mean when the objective is MAE. </p>\n<p>This implementation does not use <code>torch.median()</code> because there is a little catch. When the number of entries is even, then there are two possible values of the median. For instance, the median of [1, 2, 3, 4, 5, 6] can be either 3 or 4. When using <code>quantile=0.5</code> the average of the two medians is taken.</p>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        .models = nn.ModuleList(models).()\n\n     ():\n        output = []\n\n         m  .models:\n            logits = m(x)\n\n            output.append(logits)\n\n        output = torch.stack(output)\n        output = torch.quantile(output, , dim=)\n         output\n</code></pre>",
      "rawMarkdown": "I feel I have to give back a bit after all the goodies @brendanartley has been generously sharing.\n\nHere is a better implementation of the ensemble model from his notebooks. Using the median is better than using the mean when the objective is MAE. \n\nThis implementation does not use `torch.median()` because there is a little catch. When the number of entries is even, then there are two possible values of the median. For instance, the median of [1, 2, 3, 4, 5, 6] can be either 3 or 4. When using `quantile=0.5` the average of the two medians is taken.\n\n```python\nclass EnsembleModel(nn.Module):\n    def __init__(self, models):\n        super().__init__()\n        self.models = nn.ModuleList(models).eval()\n\n    def forward(self, x):\n        output = []\n        \n        for m in self.models:\n            logits = m(x)\n            \n            output.append(logits)\n                \n        output = torch.stack(output)\n        output = torch.quantile(output, 0.5, dim=0)\n        return output\n```",
      "votes": null
    },
    {
      "id": "3219656",
      "postDate": "06/08/2025 06:19:15",
      "content": "<p>Thank you for your sharing! But why using the median is better than using the mean when the objective is MAE?</p>",
      "rawMarkdown": "Thank you for your sharing! But why using the median is better than using the mean when the objective is MAE?",
      "votes": null
    },
    {
      "id": "3219722",
      "postDate": "06/08/2025 08:43:26",
      "content": "<p>Thank you for insight :) Then is it better to use the mean for MSE and the median for MAE?</p>",
      "rawMarkdown": "Thank you for insight :) Then is it better to use the mean for MSE and the median for MAE?",
      "votes": null
    },
    {
      "id": "3219732",
      "postDate": "06/08/2025 08:57:53",
      "content": "<p>Good question.</p>\n<p>Let a_1, a_2, …, a_n be some floating points, and f(x) = sum_i |x - a_i|</p>\n<p>Then the median of {a_1, a_2, …, a_n } minimizes x.</p>\n<p>Proof. </p>\n<p>Let's compute the gradient of f. Let's start with one entry equal to 0.</p>\n<p>Let f(x) = |x|<br>\ndf/dx(x) = sign(x)</p>\n<p>In the general case:<br>\nLet f(x) = sum_i |x - a_i|</p>\n<p>df/dx(x) = sum_i sign(x - a_i) = |{a_i such that a_i &lt; x}| -  |{a_i such that a_i &gt; x}|</p>\n<p>df/dx(x)  is zero when the number of a_i smaller than x is equal to the number of a_i greater than x, i.e. when x is the median.</p>",
      "rawMarkdown": "Good question.\n\nLet a_1, a_2, ..., a_n be some floating points, and f(x) = sum_i |x - a_i|\n\nThen the median of {a_1, a_2, ..., a_n } minimizes x.\n\nProof. \n\nLet's compute the gradient of f. Let's start with one entry equal to 0.\n\nLet f(x) = |x|\ndf/dx(x) = sign(x)\n\nIn the general case:\nLet f(x) = sum_i |x - a_i|\n\ndf/dx(x) = sum_i sign(x - a_i) = |{a_i such that a_i < x}| -  |{a_i such that a_i > x}|\n\ndf/dx(x)  is zero when the number of a_i smaller than x is equal to the number of a_i greater than x, i.e. when x is the median.",
      "votes": null
    },
    {
      "id": "3219828",
      "postDate": "06/08/2025 11:06:53",
      "content": "<p>See <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> answer for a direct derivation. But I want to add that this related to the fact that MAE is a (strictly) proper scoring rule for the median, whereas MSE is a (strictly) proper scoring rule for the expected value.</p>",
      "rawMarkdown": "See @cpmpml answer for a direct derivation. But I want to add that this related to the fact that MAE is a (strictly) proper scoring rule for the median, whereas MSE is a (strictly) proper scoring rule for the expected value.",
      "votes": null
    },
    {
      "id": "3219995",
      "postDate": "06/08/2025 15:57:59",
      "content": "<p>Yes, it is.</p>",
      "rawMarkdown": "Yes, it is.",
      "votes": null
    },
    {
      "id": "3225765",
      "postDate": "06/16/2025 20:13:51",
      "content": "<p>i haven't tried this but a weighted ensemble version could be:</p>\n<p>median(sorted(a,a,a,b,b,b,,c,))</p>\n<p>repeats model predicted by its importance weightt</p>",
      "rawMarkdown": "i haven't tried this but a weighted ensemble version could be:\n\nmedian(sorted(a,a,a,b,b,b,,c,))\n\nrepeats model predicted by its importance weightt",
      "votes": null
    },
    {
      "id": "3226426",
      "postDate": "06/17/2025 16:22:57",
      "content": "<p>It is rather easy to check that median(sorted(a,a,a,b,b,b,,c,)) equals median([a,b,c]), see below.</p>\n<p>Median is tricky, and repeating occurrences  can be counter intuitive.</p>\n<p>In your case, a and b play the same role, we can therefore assume  a &lt; b.</p>\n<p>There are three cases if a,b, and c are pairwise different.</p>\n<p>c &lt; a &lt; b:<br>\nsorted numbers are c, a, a, a, b, b, b, and median is a</p>\n<p>a &lt; c &lt; b:<br>\nsorted numbers are a, a, a, c, b, b, b, and median is c</p>\n<p>a &lt; b &lt; c:<br>\nsorted numbers are a, a, a, b, b, b, c, and median is b</p>",
      "rawMarkdown": "It is rather easy to check that median(sorted(a,a,a,b,b,b,,c,)) equals median([a,b,c]), see below.\n\nMedian is tricky, and repeating occurrences  can be counter intuitive.\n\nIn your case, a and b play the same role, we can therefore assume  a < b.\n\nThere are three cases if a,b, and c are pairwise different.\n\nc < a < b:\nsorted numbers are c, a, a, a, b, b, b, and median is a\n\na < c < b:\nsorted numbers are a, a, a, c, b, b, b, and median is c\n\n a < b < c:\nsorted numbers are a, a, a, b, b, b, c, and median is b",
      "votes": null
    },
    {
      "id": "3227226",
      "postDate": "06/18/2025 17:05:52",
      "content": "<p>Maybe just use forward wave reconstruction error to choose among solutions or different ensembling solutions?</p>",
      "rawMarkdown": "Maybe just use forward wave reconstruction error to choose among solutions or different ensembling solutions?",
      "votes": null
    },
    {
      "id": "3227230",
      "postDate": "06/18/2025 17:08:18",
      "content": "<p>Thanks, your argument make sense. Need to think of  new ways how lb scores of each model can help in ensembling.</p>",
      "rawMarkdown": "Thanks, your argument make sense. Need to think of  new ways how lb scores of each model can help in ensembling.",
      "votes": null
    },
    {
      "id": "3238558",
      "postDate": "07/02/2025 01:14:43",
      "content": "<p>I tested it and it was a bit worse than median. Did you use it?</p>",
      "rawMarkdown": "I tested it and it was a bit worse than median. Did you use it?",
      "votes": null
    },
    {
      "id": "3238563",
      "postDate": "07/02/2025 01:35:48",
      "content": "<p>It has used in our solution</p>",
      "rawMarkdown": "It has used in our solution",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3219656,
      "author_name": "xunden",
      "author_url": "",
      "post_date": "06/08/2025 06:19:15",
      "content": "<p>Thank you for your sharing! But why using the median is better than using the mean when the objective is MAE?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219732,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/08/2025 08:57:53",
          "content": "<p>Good question.</p>\n<p>Let a_1, a_2, …, a_n be some floating points, and f(x) = sum_i |x - a_i|</p>\n<p>Then the median of {a_1, a_2, …, a_n } minimizes x.</p>\n<p>Proof. </p>\n<p>Let's compute the gradient of f. Let's start with one entry equal to 0.</p>\n<p>Let f(x) = |x|<br>\ndf/dx(x) = sign(x)</p>\n<p>In the general case:<br>\nLet f(x) = sum_i |x - a_i|</p>\n<p>df/dx(x) = sum_i sign(x - a_i) = |{a_i such that a_i &lt; x}| -  |{a_i such that a_i &gt; x}|</p>\n<p>df/dx(x)  is zero when the number of a_i smaller than x is equal to the number of a_i greater than x, i.e. when x is the median.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3219828,
          "author_name": "hoffmanns",
          "author_url": "",
          "post_date": "06/08/2025 11:06:53",
          "content": "<p>See <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> answer for a direct derivation. But I want to add that this related to the fact that MAE is a (strictly) proper scoring rule for the median, whereas MSE is a (strictly) proper scoring rule for the expected value.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3219722,
      "author_name": "changhyouns",
      "author_url": "",
      "post_date": "06/08/2025 08:43:26",
      "content": "<p>Thank you for insight :) Then is it better to use the mean for MSE and the median for MAE?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3219995,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/08/2025 15:57:59",
          "content": "<p>Yes, it is.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3225765,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/16/2025 20:13:51",
      "content": "<p>i haven't tried this but a weighted ensemble version could be:</p>\n<p>median(sorted(a,a,a,b,b,b,,c,))</p>\n<p>repeats model predicted by its importance weightt</p>",
      "votes": null,
      "replies": [
        {
          "id": 3226426,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/17/2025 16:22:57",
          "content": "<p>It is rather easy to check that median(sorted(a,a,a,b,b,b,,c,)) equals median([a,b,c]), see below.</p>\n<p>Median is tricky, and repeating occurrences  can be counter intuitive.</p>\n<p>In your case, a and b play the same role, we can therefore assume  a &lt; b.</p>\n<p>There are three cases if a,b, and c are pairwise different.</p>\n<p>c &lt; a &lt; b:<br>\nsorted numbers are c, a, a, a, b, b, b, and median is a</p>\n<p>a &lt; c &lt; b:<br>\nsorted numbers are a, a, a, c, b, b, b, and median is c</p>\n<p>a &lt; b &lt; c:<br>\nsorted numbers are a, a, a, b, b, b, c, and median is b</p>",
          "votes": null,
          "replies": [
            {
              "id": 3227230,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "06/18/2025 17:08:18",
              "content": "<p>Thanks, your argument make sense. Need to think of  new ways how lb scores of each model can help in ensembling.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3227226,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/18/2025 17:05:52",
      "content": "<p>Maybe just use forward wave reconstruction error to choose among solutions or different ensembling solutions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3238558,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/02/2025 01:14:43",
          "content": "<p>I tested it and it was a bit worse than median. Did you use it?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3238563,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "07/02/2025 01:35:48",
              "content": "<p>It has used in our solution</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3215937": "I feel I have to give back a bit after all the goodies @brendanartley has been generously sharing.\n\nHere is a better implementation of the ensemble model from his notebooks. Using the median is better than using the mean when the objective is MAE. \n\nThis implementation does not use `torch.median()` because there is a little catch. When the number of entries is even, then there are two possible values of the median. For instance, the median of [1, 2, 3, 4, 5, 6] can be either 3 or 4. When using `quantile=0.5` the average of the two medians is taken.\n\n```python\nclass EnsembleModel(nn.Module):\n    def __init__(self, models):\n        super().__init__()\n        self.models = nn.ModuleList(models).eval()\n\n    def forward(self, x):\n        output = []\n        \n        for m in self.models:\n            logits = m(x)\n            \n            output.append(logits)\n                \n        output = torch.stack(output)\n        output = torch.quantile(output, 0.5, dim=0)\n        return output\n```",
    "3219656": "Thank you for your sharing! But why using the median is better than using the mean when the objective is MAE?",
    "3219722": "Thank you for insight :) Then is it better to use the mean for MSE and the median for MAE?",
    "3219732": "Good question.\n\nLet a_1, a_2, ..., a_n be some floating points, and f(x) = sum_i |x - a_i|\n\nThen the median of {a_1, a_2, ..., a_n } minimizes x.\n\nProof. \n\nLet's compute the gradient of f. Let's start with one entry equal to 0.\n\nLet f(x) = |x|\ndf/dx(x) = sign(x)\n\nIn the general case:\nLet f(x) = sum_i |x - a_i|\n\ndf/dx(x) = sum_i sign(x - a_i) = |{a_i such that a_i < x}| -  |{a_i such that a_i > x}|\n\ndf/dx(x)  is zero when the number of a_i smaller than x is equal to the number of a_i greater than x, i.e. when x is the median.",
    "3219828": "See @cpmpml answer for a direct derivation. But I want to add that this related to the fact that MAE is a (strictly) proper scoring rule for the median, whereas MSE is a (strictly) proper scoring rule for the expected value.",
    "3219995": "Yes, it is.",
    "3225765": "i haven't tried this but a weighted ensemble version could be:\n\nmedian(sorted(a,a,a,b,b,b,,c,))\n\nrepeats model predicted by its importance weightt",
    "3226426": "It is rather easy to check that median(sorted(a,a,a,b,b,b,,c,)) equals median([a,b,c]), see below.\n\nMedian is tricky, and repeating occurrences  can be counter intuitive.\n\nIn your case, a and b play the same role, we can therefore assume  a < b.\n\nThere are three cases if a,b, and c are pairwise different.\n\nc < a < b:\nsorted numbers are c, a, a, a, b, b, b, and median is a\n\na < c < b:\nsorted numbers are a, a, a, c, b, b, b, and median is c\n\n a < b < c:\nsorted numbers are a, a, a, b, b, b, c, and median is b",
    "3227226": "Maybe just use forward wave reconstruction error to choose among solutions or different ensembling solutions?",
    "3227230": "Thanks, your argument make sense. Need to think of  new ways how lb scores of each model can help in ensembling.",
    "3238558": "I tested it and it was a bit worse than median. Did you use it?",
    "3238563": "It has used in our solution"
  },
  "source": "meta"
}