{
  "id": 245459,
  "title": "LB 0.98: Forward Ensembling Technique (Credits to Chris)",
  "url": "/competitions/seti-breakthrough-listen/discussion/245459",
  "author_name": "gao-hongnan",
  "post_date": "2021-06-11T01:47:07.283000",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/reighns/lb-0-98-cv-0-9915-forward-ensembling-technique?scriptVersionId=65595160\" target=\"_blank\">LB 0.98 CV | Forward Ensembling Technique</a></p>\n<p>I started Deep Learning journey a year back, just when Melanoma comp was rampant. I learnt a lot from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ; and I kept his Forward Ensembling tricks till now. I made his code into a <code>class</code> as I like to reuse them for many use cases. The code may not be optimized, but it is what I can write at this level (not CS major).</p>\n<p>Here is the code:</p>\n<pre><code>class ForwardEnsemble:\n    def __init__(\n        self,\n        dir: str,\n        oof: pd.DataFrame,\n        weight_interval: int,\n        patience: int,\n        min_increase: float,\n        target_column_names: List[str],\n        pred_column_names: List[str],\n    ):\n        super().__init__()\n        self.dir = dir\n        FILES = os.listdir(dir)\n        self.oof_list = np.sort([f for f in FILES if \"oof\" in f])\n        self.num_oofs = len(self.oof_list)\n\n        self.oof = oof  # the oof csv with n rows m columns where n is the number of images in the dataset, and m be the number of target columns * number of oof you have\n        self.weight_interval = weight_interval\n        self.patience = patience\n        self.min_increase = min_increase\n        self.target_column_names = (\n            target_column_names  # target_cols = oof[0].iloc[:, 1:12].columns.tolist()\n        )\n        self.pred_column_names = (\n            pred_column_names  # pred_cols = oof[0].iloc[:, 15:].columns.tolist()\n        )\n\n        self.col_len = len(target_column_names)\n\n        self.num_test_images = len(oof[0])\n\n        # get ground truth\n        self.y_true = y_true_df['target'].values\n\n        self.all_oof_preds = np.zeros(\n            (self.num_test_images, self.num_oofs * self.col_len)\n        )\n\n        # append all oof preds to all_oof_preds: for example - k=0 -&gt; all_oof_preds[:,0:11] = self.oof[0][['ETT - Abnormal OOF', etc]].values\n        for k in range(self.num_oofs):\n            self.all_oof_preds[\n                :,\n                int(k * self.col_len) : int((k + 1) * self.col_len),\n            ] = oof[k][pred_column_names].values\n\n        print(self.all_oof_preds)\n        print(self.num_oofs)\n\n        self.model_i_score, self.model_i_index, self.model_i_weight = 0, 0, 0\n\n    def __len__(self):\n        return len(\n            self.column_names\n        )  # get number of prediction columns, in multi-label, should have more than 1 column, while in binary, there is only 1\n\n    def macro_multilabel_auc(self, label, pred):\n        \"\"\" Also works for binary AUC like Melanoma\"\"\"\n        aucs = []\n        aucs.append(roc_auc_score(label, pred))\n        return np.mean(aucs)\n\n    def compute_best_oof(self):\n        _all = []\n        for k in range(self.num_oofs):\n            print(self.all_oof_preds[:, 0])\n            auc = self.macro_multilabel_auc(\n                self.y_true,\n                self.all_oof_preds[\n                    :,k\n                ],\n            )\n            _all.append(auc)\n            print(\"Model %i has OOF AUC = %.4f\" % (k, auc))\n        best_auc, best_oof_index = np.max(_all), np.argmax(_all)\n        return best_auc, best_oof_index\n\n    def forward_ensemble(self):\n        DUPLICATES = False\n        old_best_auc, best_oof_index = self.compute_best_oof()\n        chosen_model = [best_oof_index]\n        optimal_weights = []\n        for oof_index in range(self.num_oofs):\n            curr_model = self.all_oof_preds[\n                :,\n                int(best_oof_index * self.col_len) : int(\n                    (best_oof_index + 1) * self.col_len\n                ),\n            ]\n            for i, k in enumerate(chosen_model[1:]):\n                # this step is confusing because it overwrites curr_model in the previous step. basically curr_model is reset to the best oof model initially, and then loop through to get the best oof\n                curr_model = (\n                    optimal_weights[i]\n                    * self.all_oof_preds[\n                        :, int(k * self.col_len) : int((k + 1) * self.col_len)\n                    ]\n                    + (1 - optimal_weights[i]) * curr_model\n                )\n\n            print(\"Searching for best model to add\")\n\n            # try add each model\n            for i in range(self.num_oofs):\n                print(i, \", \", end=\"\")\n                if not DUPLICATES and (i in chosen_model):\n                    continue\n                best_weight_index, best_score, patience_counter = 0, 0, 0\n                for j in range(self.weight_interval):\n                    temp = (j / self.weight_interval) * self.all_oof_preds[\n                        :, int(i * self.col_len) : int((i + 1) * self.col_len)\n                    ] + (1 - j / self.weight_interval) * curr_model\n                    auc = self.macro_multilabel_auc(self.y_true, temp)\n\n                    if auc &gt; best_score:\n                        best_score = auc\n                        best_weight_index = j / self.weight_interval\n                    else:\n                        patience_counter += 1\n                    if patience_counter &gt; self.patience:\n                        break\n                    if best_score &gt; self.model_i_score:\n                        self.model_i_score = best_score\n                        self.model_i_index = i\n                        self.model_i_weights = best_weight_index\n\n            increment = self.model_i_score - old_best_auc\n            if increment &lt;= self.min_increase:\n                print(\"No more significant increase\")\n                break\n\n            old_best_auc = self.model_i_score\n            chosen_model.append(self.model_i_index)\n            optimal_weights.append(self.model_i_weights)\n        return chosen_model, optimal_weights\n</code></pre>\n<p>Here I present a small use case in how one can ensemble small models (thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>) to achieve 0.98 easily. I used 3 models, namely <code>efficientnetb0</code> and <code>resnet34d</code> and I forgot the other one to achieve this score. Each individual model stands around 0.977-0.979 ~~ in LB (my estimate)</p>\n<p>One phenomenon I find in this competition is that ensemble of different models do not lead to a boost in LB…..? Why? I don't know. There's cases whereby ensembling decreases the model by a lot. I need to figure out why.</p>\n<p>You can find <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> notebook here:<br>\n<a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private</a></p>\n<p>Edit: I will post my pipeline in time to come. Need to tidy up a bit. But if you are already eager to try, tawara pipeline is good enough. </p>",
  "messages": [
    {
      "id": 1344768,
      "postDate": "2021-06-11T06:03:19.457Z",
      "content": "<p>First of all, thank you for sharing, but I think you can share the idea of reaching 0.98 without opening up submiss.csv, because 0.98 also takes a long time to train, which is not fair for the current competition, because some people can get high marks if they see it first</p>",
      "rawMarkdown": "First of all, thank you for sharing, but I think you can share the idea of reaching 0.98 without opening up submiss.csv, because 0.98 also takes a long time to train, which is not fair for the current competition, because some people can get high marks if they see it first",
      "votes": 13,
      "replies": [
        {
          "id": 1344774,
          "postDate": "2021-06-11T06:08:57.967Z",
          "content": "<p>Thanks for highlighting this. I’m very much aware on this sensitive issue as I posted a high scoring notebook 1 month prior to competition end, and I received many backlash and negative remarks from the community. </p>\n<p>This time therefore, I published it with almost 2 months left.</p>",
          "rawMarkdown": "Thanks for highlighting this. I’m very much aware on this sensitive issue as I posted a high scoring notebook 1 month prior to competition end, and I received many backlash and negative remarks from the community. \n\nThis time therefore, I published it with almost 2 months left.",
          "votes": 1
        },
        {
          "id": 1344788,
          "postDate": "2021-06-11T06:24:35.720Z",
          "content": "<p>Just don't think it's fair to the people who spend time optimizing the training model，no offense meant，hh🙏</p>",
          "rawMarkdown": "Just don't think it's fair to the people who spend time optimizing the training model，no offense meant，hh🙏",
          "votes": 6
        },
        {
          "id": 1344794,
          "postDate": "2021-06-11T06:28:22.630Z",
          "content": "<p>Thanks. No offence taken. This isn’t the first time I received backlash on this (I adhere closely to the rules that one should not publish any high scoring notebooks 2 weeks prior - even though the banter only appears 1 week before). But still people aren’t pleased on this (not saying you but the previous competition). I however, concur that posting a meaningless high scoring notebook is frowned upon. I’m not sure if my notebook is meaningless as it does explores a technique on forward ensembling (hill climbing). (Credits to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as usual as he helped me learnt a lot) </p>",
          "rawMarkdown": "Thanks. No offence taken. This isn’t the first time I received backlash on this (I adhere closely to the rules that one should not publish any high scoring notebooks 2 weeks prior - even though the banter only appears 1 week before). But still people aren’t pleased on this (not saying you but the previous competition). I however, concur that posting a meaningless high scoring notebook is frowned upon. I’m not sure if my notebook is meaningless as it does explores a technique on forward ensembling (hill climbing). (Credits to @cdeotte as usual as he helped me learnt a lot) ",
          "votes": 8
        },
        {
          "id": 1345289,
          "postDate": "2021-06-11T13:23:52.620Z",
          "content": "<p>open high score .csv submission is big issue?<br>\nI think users will not submit other's submissions</p>",
          "rawMarkdown": "open high score .csv submission is big issue?\nI think users will not submit other's submissions",
          "votes": -4
        },
        {
          "id": 1345307,
          "postDate": "2021-06-11T13:40:34.300Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1344525,
      "postDate": "2021-06-11T01:47:07.283Z",
      "content": "<p><a href=\"https://www.kaggle.com/reighns/lb-0-98-cv-0-9915-forward-ensembling-technique?scriptVersionId=65595160\" target=\"_blank\">LB 0.98 CV | Forward Ensembling Technique</a></p>\n<p>I started Deep Learning journey a year back, just when Melanoma comp was rampant. I learnt a lot from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ; and I kept his Forward Ensembling tricks till now. I made his code into a <code>class</code> as I like to reuse them for many use cases. The code may not be optimized, but it is what I can write at this level (not CS major).</p>\n<p>Here is the code:</p>\n<pre><code>class ForwardEnsemble:\n    def __init__(\n        self,\n        dir: str,\n        oof: pd.DataFrame,\n        weight_interval: int,\n        patience: int,\n        min_increase: float,\n        target_column_names: List[str],\n        pred_column_names: List[str],\n    ):\n        super().__init__()\n        self.dir = dir\n        FILES = os.listdir(dir)\n        self.oof_list = np.sort([f for f in FILES if \"oof\" in f])\n        self.num_oofs = len(self.oof_list)\n\n        self.oof = oof  # the oof csv with n rows m columns where n is the number of images in the dataset, and m be the number of target columns * number of oof you have\n        self.weight_interval = weight_interval\n        self.patience = patience\n        self.min_increase = min_increase\n        self.target_column_names = (\n            target_column_names  # target_cols = oof[0].iloc[:, 1:12].columns.tolist()\n        )\n        self.pred_column_names = (\n            pred_column_names  # pred_cols = oof[0].iloc[:, 15:].columns.tolist()\n        )\n\n        self.col_len = len(target_column_names)\n\n        self.num_test_images = len(oof[0])\n\n        # get ground truth\n        self.y_true = y_true_df['target'].values\n\n        self.all_oof_preds = np.zeros(\n            (self.num_test_images, self.num_oofs * self.col_len)\n        )\n\n        # append all oof preds to all_oof_preds: for example - k=0 -&gt; all_oof_preds[:,0:11] = self.oof[0][['ETT - Abnormal OOF', etc]].values\n        for k in range(self.num_oofs):\n            self.all_oof_preds[\n                :,\n                int(k * self.col_len) : int((k + 1) * self.col_len),\n            ] = oof[k][pred_column_names].values\n\n        print(self.all_oof_preds)\n        print(self.num_oofs)\n\n        self.model_i_score, self.model_i_index, self.model_i_weight = 0, 0, 0\n\n    def __len__(self):\n        return len(\n            self.column_names\n        )  # get number of prediction columns, in multi-label, should have more than 1 column, while in binary, there is only 1\n\n    def macro_multilabel_auc(self, label, pred):\n        \"\"\" Also works for binary AUC like Melanoma\"\"\"\n        aucs = []\n        aucs.append(roc_auc_score(label, pred))\n        return np.mean(aucs)\n\n    def compute_best_oof(self):\n        _all = []\n        for k in range(self.num_oofs):\n            print(self.all_oof_preds[:, 0])\n            auc = self.macro_multilabel_auc(\n                self.y_true,\n                self.all_oof_preds[\n                    :,k\n                ],\n            )\n            _all.append(auc)\n            print(\"Model %i has OOF AUC = %.4f\" % (k, auc))\n        best_auc, best_oof_index = np.max(_all), np.argmax(_all)\n        return best_auc, best_oof_index\n\n    def forward_ensemble(self):\n        DUPLICATES = False\n        old_best_auc, best_oof_index = self.compute_best_oof()\n        chosen_model = [best_oof_index]\n        optimal_weights = []\n        for oof_index in range(self.num_oofs):\n            curr_model = self.all_oof_preds[\n                :,\n                int(best_oof_index * self.col_len) : int(\n                    (best_oof_index + 1) * self.col_len\n                ),\n            ]\n            for i, k in enumerate(chosen_model[1:]):\n                # this step is confusing because it overwrites curr_model in the previous step. basically curr_model is reset to the best oof model initially, and then loop through to get the best oof\n                curr_model = (\n                    optimal_weights[i]\n                    * self.all_oof_preds[\n                        :, int(k * self.col_len) : int((k + 1) * self.col_len)\n                    ]\n                    + (1 - optimal_weights[i]) * curr_model\n                )\n\n            print(\"Searching for best model to add\")\n\n            # try add each model\n            for i in range(self.num_oofs):\n                print(i, \", \", end=\"\")\n                if not DUPLICATES and (i in chosen_model):\n                    continue\n                best_weight_index, best_score, patience_counter = 0, 0, 0\n                for j in range(self.weight_interval):\n                    temp = (j / self.weight_interval) * self.all_oof_preds[\n                        :, int(i * self.col_len) : int((i + 1) * self.col_len)\n                    ] + (1 - j / self.weight_interval) * curr_model\n                    auc = self.macro_multilabel_auc(self.y_true, temp)\n\n                    if auc &gt; best_score:\n                        best_score = auc\n                        best_weight_index = j / self.weight_interval\n                    else:\n                        patience_counter += 1\n                    if patience_counter &gt; self.patience:\n                        break\n                    if best_score &gt; self.model_i_score:\n                        self.model_i_score = best_score\n                        self.model_i_index = i\n                        self.model_i_weights = best_weight_index\n\n            increment = self.model_i_score - old_best_auc\n            if increment &lt;= self.min_increase:\n                print(\"No more significant increase\")\n                break\n\n            old_best_auc = self.model_i_score\n            chosen_model.append(self.model_i_index)\n            optimal_weights.append(self.model_i_weights)\n        return chosen_model, optimal_weights\n</code></pre>\n<p>Here I present a small use case in how one can ensemble small models (thanks <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>) to achieve 0.98 easily. I used 3 models, namely <code>efficientnetb0</code> and <code>resnet34d</code> and I forgot the other one to achieve this score. Each individual model stands around 0.977-0.979 ~~ in LB (my estimate)</p>\n<p>One phenomenon I find in this competition is that ensemble of different models do not lead to a boost in LB…..? Why? I don't know. There's cases whereby ensembling decreases the model by a lot. I need to figure out why.</p>\n<p>You can find <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> notebook here:<br>\n<a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private</a></p>\n<p>Edit: I will post my pipeline in time to come. Need to tidy up a bit. But if you are already eager to try, tawara pipeline is good enough. </p>",
      "rawMarkdown": "[LB 0.98 CV | Forward Ensembling Technique](https://www.kaggle.com/reighns/lb-0-98-cv-0-9915-forward-ensembling-technique?scriptVersionId=65595160)\n\n\nI started Deep Learning journey a year back, just when Melanoma comp was rampant. I learnt a lot from @cdeotte ; and I kept his Forward Ensembling tricks till now. I made his code into a `class` as I like to reuse them for many use cases. The code may not be optimized, but it is what I can write at this level (not CS major).\n\nHere is the code:\n\n```\nclass ForwardEnsemble:\n    def __init__(\n        self,\n        dir: str,\n        oof: pd.DataFrame,\n        weight_interval: int,\n        patience: int,\n        min_increase: float,\n        target_column_names: List[str],\n        pred_column_names: List[str],\n    ):\n        super().__init__()\n        self.dir = dir\n        FILES = os.listdir(dir)\n        self.oof_list = np.sort([f for f in FILES if \"oof\" in f])\n        self.num_oofs = len(self.oof_list)\n\n        self.oof = oof  # the oof csv with n rows m columns where n is the number of images in the dataset, and m be the number of target columns * number of oof you have\n        self.weight_interval = weight_interval\n        self.patience = patience\n        self.min_increase = min_increase\n        self.target_column_names = (\n            target_column_names  # target_cols = oof[0].iloc[:, 1:12].columns.tolist()\n        )\n        self.pred_column_names = (\n            pred_column_names  # pred_cols = oof[0].iloc[:, 15:].columns.tolist()\n        )\n\n        self.col_len = len(target_column_names)\n\n        self.num_test_images = len(oof[0])\n\n        # get ground truth\n        self.y_true = y_true_df['target'].values\n\n        self.all_oof_preds = np.zeros(\n            (self.num_test_images, self.num_oofs * self.col_len)\n        )\n\n        # append all oof preds to all_oof_preds: for example - k=0 -> all_oof_preds[:,0:11] = self.oof[0][['ETT - Abnormal OOF', etc]].values\n        for k in range(self.num_oofs):\n            self.all_oof_preds[\n                :,\n                int(k * self.col_len) : int((k + 1) * self.col_len),\n            ] = oof[k][pred_column_names].values\n            \n        print(self.all_oof_preds)\n        print(self.num_oofs)\n        \n        self.model_i_score, self.model_i_index, self.model_i_weight = 0, 0, 0\n\n    def __len__(self):\n        return len(\n            self.column_names\n        )  # get number of prediction columns, in multi-label, should have more than 1 column, while in binary, there is only 1\n\n    def macro_multilabel_auc(self, label, pred):\n        \"\"\" Also works for binary AUC like Melanoma\"\"\"\n        aucs = []\n        aucs.append(roc_auc_score(label, pred))\n        return np.mean(aucs)\n\n    def compute_best_oof(self):\n        _all = []\n        for k in range(self.num_oofs):\n            print(self.all_oof_preds[:, 0])\n            auc = self.macro_multilabel_auc(\n                self.y_true,\n                self.all_oof_preds[\n                    :,k\n                ],\n            )\n            _all.append(auc)\n            print(\"Model %i has OOF AUC = %.4f\" % (k, auc))\n        best_auc, best_oof_index = np.max(_all), np.argmax(_all)\n        return best_auc, best_oof_index\n\n    def forward_ensemble(self):\n        DUPLICATES = False\n        old_best_auc, best_oof_index = self.compute_best_oof()\n        chosen_model = [best_oof_index]\n        optimal_weights = []\n        for oof_index in range(self.num_oofs):\n            curr_model = self.all_oof_preds[\n                :,\n                int(best_oof_index * self.col_len) : int(\n                    (best_oof_index + 1) * self.col_len\n                ),\n            ]\n            for i, k in enumerate(chosen_model[1:]):\n                # this step is confusing because it overwrites curr_model in the previous step. basically curr_model is reset to the best oof model initially, and then loop through to get the best oof\n                curr_model = (\n                    optimal_weights[i]\n                    * self.all_oof_preds[\n                        :, int(k * self.col_len) : int((k + 1) * self.col_len)\n                    ]\n                    + (1 - optimal_weights[i]) * curr_model\n                )\n\n            print(\"Searching for best model to add\")\n\n            # try add each model\n            for i in range(self.num_oofs):\n                print(i, \", \", end=\"\")\n                if not DUPLICATES and (i in chosen_model):\n                    continue\n                best_weight_index, best_score, patience_counter = 0, 0, 0\n                for j in range(self.weight_interval):\n                    temp = (j / self.weight_interval) * self.all_oof_preds[\n                        :, int(i * self.col_len) : int((i + 1) * self.col_len)\n                    ] + (1 - j / self.weight_interval) * curr_model\n                    auc = self.macro_multilabel_auc(self.y_true, temp)\n\n                    if auc > best_score:\n                        best_score = auc\n                        best_weight_index = j / self.weight_interval\n                    else:\n                        patience_counter += 1\n                    if patience_counter > self.patience:\n                        break\n                    if best_score > self.model_i_score:\n                        self.model_i_score = best_score\n                        self.model_i_index = i\n                        self.model_i_weights = best_weight_index\n\n            increment = self.model_i_score - old_best_auc\n            if increment <= self.min_increase:\n                print(\"No more significant increase\")\n                break\n\n            old_best_auc = self.model_i_score\n            chosen_model.append(self.model_i_index)\n            optimal_weights.append(self.model_i_weights)\n        return chosen_model, optimal_weights\n```\n\nHere I present a small use case in how one can ensemble small models (thanks @ttahara) to achieve 0.98 easily. I used 3 models, namely `efficientnetb0` and `resnet34d` and I forgot the other one to achieve this score. Each individual model stands around 0.977-0.979 ~~ in LB (my estimate)\n\nOne phenomenon I find in this competition is that ensemble of different models do not lead to a boost in LB.....? Why? I don't know. There's cases whereby ensembling decreases the model by a lot. I need to figure out why.\n\nYou can find @cdeotte notebook here:\nhttps://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\n\nEdit: I will post my pipeline in time to come. Need to tidy up a bit. But if you are already eager to try, tawara pipeline is good enough. \n\n",
      "votes": 16
    },
    {
      "id": 1345169,
      "postDate": "2021-06-11T11:17:21.263Z",
      "content": "<p>I learned a lot from this post. Thanks.<br>\nI applied the ensemble using xgboost before.<br>\nHowever, there was a problem of overfitting in the auc score.</p>\n<p>Before I read your post, I did not know the reason for dividing the data into valid, and test dataset.<br>\nIt seems that the method of learning test dataset through oof is pretty good.<br>\nI will try to use this method in final phase.</p>",
      "rawMarkdown": "I learned a lot from this post. Thanks.\nI applied the ensemble using xgboost before.\nHowever, there was a problem of overfitting in the auc score.\n\nBefore I read your post, I did not know the reason for dividing the data into valid, and test dataset.\nIt seems that the method of learning test dataset through oof is pretty good.\nI will try to use this method in final phase.",
      "votes": 1
    },
    {
      "id": 1346007,
      "postDate": "2021-06-12T03:21:57.127Z",
      "content": "<p>Nice Work. Do take a look at my notebooks and give feedback.</p>",
      "rawMarkdown": "Nice Work. Do take a look at my notebooks and give feedback.",
      "votes": -14
    },
    {
      "id": 1347431,
      "postDate": "2021-06-13T07:37:01.097Z",
      "content": "<p>Thank you for sharing excellent method! But i have a question about oof! XX numbers is what numbers? What does it mean? </p>",
      "rawMarkdown": "Thank you for sharing excellent method! But i have a question about oof! XX numbers is what numbers? What does it mean? "
    }
  ],
  "comments": [
    {
      "id": 1344768,
      "author_name": "xuxu_sky",
      "author_url": "",
      "post_date": "2021-06-11T06:03:19.457000",
      "content": "<p>First of all, thank you for sharing, but I think you can share the idea of reaching 0.98 without opening up submiss.csv, because 0.98 also takes a long time to train, which is not fair for the current competition, because some people can get high marks if they see it first</p>",
      "votes": 13,
      "replies": [
        {
          "id": 1344774,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-11T06:08:57.967000",
          "content": "<p>Thanks for highlighting this. I’m very much aware on this sensitive issue as I posted a high scoring notebook 1 month prior to competition end, and I received many backlash and negative remarks from the community. </p>\n<p>This time therefore, I published it with almost 2 months left.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1344788,
          "author_name": "xuxu_sky",
          "author_url": "",
          "post_date": "2021-06-11T06:24:35.720000",
          "content": "<p>Just don't think it's fair to the people who spend time optimizing the training model，no offense meant，hh🙏</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1344794,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-06-11T06:28:22.630000",
          "content": "<p>Thanks. No offence taken. This isn’t the first time I received backlash on this (I adhere closely to the rules that one should not publish any high scoring notebooks 2 weeks prior - even though the banter only appears 1 week before). But still people aren’t pleased on this (not saying you but the previous competition). I however, concur that posting a meaningless high scoring notebook is frowned upon. I’m not sure if my notebook is meaningless as it does explores a technique on forward ensembling (hill climbing). (Credits to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> as usual as he helped me learnt a lot) </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1345289,
          "author_name": "assign",
          "author_url": "",
          "post_date": "2021-06-11T13:23:52.620000",
          "content": "<p>open high score .csv submission is big issue?<br>\nI think users will not submit other's submissions</p>",
          "votes": -4,
          "replies": []
        },
        {
          "id": 1345307,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-06-11T13:40:34.300000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1345169,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-11T11:17:21.263000",
      "content": "<p>I learned a lot from this post. Thanks.<br>\nI applied the ensemble using xgboost before.<br>\nHowever, there was a problem of overfitting in the auc score.</p>\n<p>Before I read your post, I did not know the reason for dividing the data into valid, and test dataset.<br>\nIt seems that the method of learning test dataset through oof is pretty good.<br>\nI will try to use this method in final phase.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1346007,
      "author_name": "Arnab Dey",
      "author_url": "",
      "post_date": "2021-06-12T03:21:57.127000",
      "content": "<p>Nice Work. Do take a look at my notebooks and give feedback.</p>",
      "votes": -14,
      "replies": []
    },
    {
      "id": 1347431,
      "author_name": "Jackie Mai",
      "author_url": "",
      "post_date": "2021-06-13T07:37:01.097000",
      "content": "<p>Thank you for sharing excellent method! But i have a question about oof! XX numbers is what numbers? What does it mean? </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1344768": "First of all, thank you for sharing, but I think you can share the idea of reaching 0.98 without opening up submiss.csv, because 0.98 also takes a long time to train, which is not fair for the current competition, because some people can get high marks if they see it first",
    "1344525": "[LB 0.98 CV | Forward Ensembling Technique](https://www.kaggle.com/reighns/lb-0-98-cv-0-9915-forward-ensembling-technique?scriptVersionId=65595160)\n\n\nI started Deep Learning journey a year back, just when Melanoma comp was rampant. I learnt a lot from @cdeotte ; and I kept his Forward Ensembling tricks till now. I made his code into a `class` as I like to reuse them for many use cases. The code may not be optimized, but it is what I can write at this level (not CS major).\n\nHere is the code:\n\n```\nclass ForwardEnsemble:\n    def __init__(\n        self,\n        dir: str,\n        oof: pd.DataFrame,\n        weight_interval: int,\n        patience: int,\n        min_increase: float,\n        target_column_names: List[str],\n        pred_column_names: List[str],\n    ):\n        super().__init__()\n        self.dir = dir\n        FILES = os.listdir(dir)\n        self.oof_list = np.sort([f for f in FILES if \"oof\" in f])\n        self.num_oofs = len(self.oof_list)\n\n        self.oof = oof  # the oof csv with n rows m columns where n is the number of images in the dataset, and m be the number of target columns * number of oof you have\n        self.weight_interval = weight_interval\n        self.patience = patience\n        self.min_increase = min_increase\n        self.target_column_names = (\n            target_column_names  # target_cols = oof[0].iloc[:, 1:12].columns.tolist()\n        )\n        self.pred_column_names = (\n            pred_column_names  # pred_cols = oof[0].iloc[:, 15:].columns.tolist()\n        )\n\n        self.col_len = len(target_column_names)\n\n        self.num_test_images = len(oof[0])\n\n        # get ground truth\n        self.y_true = y_true_df['target'].values\n\n        self.all_oof_preds = np.zeros(\n            (self.num_test_images, self.num_oofs * self.col_len)\n        )\n\n        # append all oof preds to all_oof_preds: for example - k=0 -> all_oof_preds[:,0:11] = self.oof[0][['ETT - Abnormal OOF', etc]].values\n        for k in range(self.num_oofs):\n            self.all_oof_preds[\n                :,\n                int(k * self.col_len) : int((k + 1) * self.col_len),\n            ] = oof[k][pred_column_names].values\n            \n        print(self.all_oof_preds)\n        print(self.num_oofs)\n        \n        self.model_i_score, self.model_i_index, self.model_i_weight = 0, 0, 0\n\n    def __len__(self):\n        return len(\n            self.column_names\n        )  # get number of prediction columns, in multi-label, should have more than 1 column, while in binary, there is only 1\n\n    def macro_multilabel_auc(self, label, pred):\n        \"\"\" Also works for binary AUC like Melanoma\"\"\"\n        aucs = []\n        aucs.append(roc_auc_score(label, pred))\n        return np.mean(aucs)\n\n    def compute_best_oof(self):\n        _all = []\n        for k in range(self.num_oofs):\n            print(self.all_oof_preds[:, 0])\n            auc = self.macro_multilabel_auc(\n                self.y_true,\n                self.all_oof_preds[\n                    :,k\n                ],\n            )\n            _all.append(auc)\n            print(\"Model %i has OOF AUC = %.4f\" % (k, auc))\n        best_auc, best_oof_index = np.max(_all), np.argmax(_all)\n        return best_auc, best_oof_index\n\n    def forward_ensemble(self):\n        DUPLICATES = False\n        old_best_auc, best_oof_index = self.compute_best_oof()\n        chosen_model = [best_oof_index]\n        optimal_weights = []\n        for oof_index in range(self.num_oofs):\n            curr_model = self.all_oof_preds[\n                :,\n                int(best_oof_index * self.col_len) : int(\n                    (best_oof_index + 1) * self.col_len\n                ),\n            ]\n            for i, k in enumerate(chosen_model[1:]):\n                # this step is confusing because it overwrites curr_model in the previous step. basically curr_model is reset to the best oof model initially, and then loop through to get the best oof\n                curr_model = (\n                    optimal_weights[i]\n                    * self.all_oof_preds[\n                        :, int(k * self.col_len) : int((k + 1) * self.col_len)\n                    ]\n                    + (1 - optimal_weights[i]) * curr_model\n                )\n\n            print(\"Searching for best model to add\")\n\n            # try add each model\n            for i in range(self.num_oofs):\n                print(i, \", \", end=\"\")\n                if not DUPLICATES and (i in chosen_model):\n                    continue\n                best_weight_index, best_score, patience_counter = 0, 0, 0\n                for j in range(self.weight_interval):\n                    temp = (j / self.weight_interval) * self.all_oof_preds[\n                        :, int(i * self.col_len) : int((i + 1) * self.col_len)\n                    ] + (1 - j / self.weight_interval) * curr_model\n                    auc = self.macro_multilabel_auc(self.y_true, temp)\n\n                    if auc > best_score:\n                        best_score = auc\n                        best_weight_index = j / self.weight_interval\n                    else:\n                        patience_counter += 1\n                    if patience_counter > self.patience:\n                        break\n                    if best_score > self.model_i_score:\n                        self.model_i_score = best_score\n                        self.model_i_index = i\n                        self.model_i_weights = best_weight_index\n\n            increment = self.model_i_score - old_best_auc\n            if increment <= self.min_increase:\n                print(\"No more significant increase\")\n                break\n\n            old_best_auc = self.model_i_score\n            chosen_model.append(self.model_i_index)\n            optimal_weights.append(self.model_i_weights)\n        return chosen_model, optimal_weights\n```\n\nHere I present a small use case in how one can ensemble small models (thanks @ttahara) to achieve 0.98 easily. I used 3 models, namely `efficientnetb0` and `resnet34d` and I forgot the other one to achieve this score. Each individual model stands around 0.977-0.979 ~~ in LB (my estimate)\n\nOne phenomenon I find in this competition is that ensemble of different models do not lead to a boost in LB.....? Why? I don't know. There's cases whereby ensembling decreases the model by a lot. I need to figure out why.\n\nYou can find @cdeotte notebook here:\nhttps://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\n\nEdit: I will post my pipeline in time to come. Need to tidy up a bit. But if you are already eager to try, tawara pipeline is good enough. \n\n",
    "1345169": "I learned a lot from this post. Thanks.\nI applied the ensemble using xgboost before.\nHowever, there was a problem of overfitting in the auc score.\n\nBefore I read your post, I did not know the reason for dividing the data into valid, and test dataset.\nIt seems that the method of learning test dataset through oof is pretty good.\nI will try to use this method in final phase.",
    "1346007": "Nice Work. Do take a look at my notebooks and give feedback.",
    "1347431": "Thank you for sharing excellent method! But i have a question about oof! XX numbers is what numbers? What does it mean? "
  }
}