{
  "id": 394376,
  "title": "The first 0.67 lb notebook is introduced",
  "url": "/competitions/asl-signs/discussion/394376",
  "author_name": "Andrij",
  "post_date": "2023-03-13T09:55:03.470000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi!<br>\nAs I wrote earlier, it seems that instead of complicating the model by its size is ineffective. <br>\nA much more efficient ensemble of several models with different preprocessing.<br>\nIn principle, there is nothing new, but in the conditions of the limits of the model in terms of size and execution time, this approach seems to be optimal. <br>\nIf you look at my notebook, there is an ensemble of three approaches, each of which contains models for the folds part. I can't use all the models because of the runtime limit. But all the same, using one model at a time folds is more effective in terms of gain lb/performance time penalty in an ensemble of different models<br>\n  than using an ensemble of 5 or 7 folds of the same model.<br>\nHere is my best notebook: <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67</a>. </p>\n<p>As you can see, the ensemble is as simple as possible:</p>\n<pre><code>    x = self.prep_inputs(tf.cast(inputs, dtype=tf.float32))\n    outputs=[]\n    for asl_model in self.asl_models:\n        outputs.append(asl_model(x)[0, :])\n    outputs_3 = tf.keras.layers.Average()(outputs)\n\n\n    x1 = self.prep_inputs_1(tf.cast(inputs, dtype=tf.float32))\n    x2 = self.prep_inputs_2(tf.cast(inputs, dtype=tf.float32))\n\n    outputs_1 = tf.keras.layers.Average()([_model(x1) for _model in self.models_1])\n\n    outputs_2=[]\n    for gwg_model in self.models_2:\n        outputs_2.append(gwg_model(x2))\n    outputs_2 = tf.keras.layers.Average()(outputs_2)\n\n    outputs = tf.multiply(0.25, outputs_1)+tf.multiply(0.45, outputs_2)+tf.multiply(0.3, outputs_3).\n</code></pre>",
  "messages": [
    {
      "id": 2179652,
      "postDate": "2023-03-13T09:55:03.470Z",
      "content": "<p>Hi!<br>\nAs I wrote earlier, it seems that instead of complicating the model by its size is ineffective. <br>\nA much more efficient ensemble of several models with different preprocessing.<br>\nIn principle, there is nothing new, but in the conditions of the limits of the model in terms of size and execution time, this approach seems to be optimal. <br>\nIf you look at my notebook, there is an ensemble of three approaches, each of which contains models for the folds part. I can't use all the models because of the runtime limit. But all the same, using one model at a time folds is more effective in terms of gain lb/performance time penalty in an ensemble of different models<br>\n  than using an ensemble of 5 or 7 folds of the same model.<br>\nHere is my best notebook: <a href=\"https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67\" target=\"_blank\">https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67</a>. </p>\n<p>As you can see, the ensemble is as simple as possible:</p>\n<pre><code>    x = self.prep_inputs(tf.cast(inputs, dtype=tf.float32))\n    outputs=[]\n    for asl_model in self.asl_models:\n        outputs.append(asl_model(x)[0, :])\n    outputs_3 = tf.keras.layers.Average()(outputs)\n\n\n    x1 = self.prep_inputs_1(tf.cast(inputs, dtype=tf.float32))\n    x2 = self.prep_inputs_2(tf.cast(inputs, dtype=tf.float32))\n\n    outputs_1 = tf.keras.layers.Average()([_model(x1) for _model in self.models_1])\n\n    outputs_2=[]\n    for gwg_model in self.models_2:\n        outputs_2.append(gwg_model(x2))\n    outputs_2 = tf.keras.layers.Average()(outputs_2)\n\n    outputs = tf.multiply(0.25, outputs_1)+tf.multiply(0.45, outputs_2)+tf.multiply(0.3, outputs_3).\n</code></pre>",
      "rawMarkdown": "Hi!\nAs I wrote earlier, it seems that instead of complicating the model by its size is ineffective. \nA much more efficient ensemble of several models with different preprocessing.\nIn principle, there is nothing new, but in the conditions of the limits of the model in terms of size and execution time, this approach seems to be optimal. \nIf you look at my notebook, there is an ensemble of three approaches, each of which contains models for the folds part. I can't use all the models because of the runtime limit. But all the same, using one model at a time folds is more effective in terms of gain lb/performance time penalty in an ensemble of different models\n  than using an ensemble of 5 or 7 folds of the same model.\nHere is my best notebook: https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67. \n\nAs you can see, the ensemble is as simple as possible:\n\n        x = self.prep_inputs(tf.cast(inputs, dtype=tf.float32))\n        outputs=[]\n        for asl_model in self.asl_models:\n            outputs.append(asl_model(x)[0, :])\n        outputs_3 = tf.keras.layers.Average()(outputs)\n        \n        \n        x1 = self.prep_inputs_1(tf.cast(inputs, dtype=tf.float32))\n        x2 = self.prep_inputs_2(tf.cast(inputs, dtype=tf.float32))\n        \n        outputs_1 = tf.keras.layers.Average()([_model(x1) for _model in self.models_1])\n        \n        outputs_2=[]\n        for gwg_model in self.models_2:\n            outputs_2.append(gwg_model(x2))\n        outputs_2 = tf.keras.layers.Average()(outputs_2)\n        \n        outputs = tf.multiply(0.25, outputs_1)+tf.multiply(0.45, outputs_2)+tf.multiply(0.3, outputs_3).\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2179652": "Hi!\nAs I wrote earlier, it seems that instead of complicating the model by its size is ineffective. \nA much more efficient ensemble of several models with different preprocessing.\nIn principle, there is nothing new, but in the conditions of the limits of the model in terms of size and execution time, this approach seems to be optimal. \nIf you look at my notebook, there is an ensemble of three approaches, each of which contains models for the folds part. I can't use all the models because of the runtime limit. But all the same, using one model at a time folds is more effective in terms of gain lb/performance time penalty in an ensemble of different models\n  than using an ensemble of 5 or 7 folds of the same model.\nHere is my best notebook: https://www.kaggle.com/code/aikhmelnytskyy/gislr-tf-on-the-shoulders-ensamble-v2-0-67. \n\nAs you can see, the ensemble is as simple as possible:\n\n        x = self.prep_inputs(tf.cast(inputs, dtype=tf.float32))\n        outputs=[]\n        for asl_model in self.asl_models:\n            outputs.append(asl_model(x)[0, :])\n        outputs_3 = tf.keras.layers.Average()(outputs)\n        \n        \n        x1 = self.prep_inputs_1(tf.cast(inputs, dtype=tf.float32))\n        x2 = self.prep_inputs_2(tf.cast(inputs, dtype=tf.float32))\n        \n        outputs_1 = tf.keras.layers.Average()([_model(x1) for _model in self.models_1])\n        \n        outputs_2=[]\n        for gwg_model in self.models_2:\n            outputs_2.append(gwg_model(x2))\n        outputs_2 = tf.keras.layers.Average()(outputs_2)\n        \n        outputs = tf.multiply(0.25, outputs_1)+tf.multiply(0.45, outputs_2)+tf.multiply(0.3, outputs_3).\n"
  }
}