{
  "id": 275872,
  "title": "Models' economics ranking",
  "url": "/competitions/landmark-recognition-2021/discussion/275872",
  "author_name": "",
  "post_date": "2021-10-02T00:47:27.911503800Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Since there are always questions about hardware in these types of competitions, I will share my most economical model: relatively good performance at an affordable price.</p>\n<p>Probably the best candidate on my side is the B3 model. I used the instance g4dn.12xlarge, and the model trained for about 48 hours (10 epochs on 1.6m images). The instance cost is 3.9 USD / hour, which yields a total training cost of 187 USD. The private score of this model is 0.357, which would be 24th place on the LB.</p>\n<p>Instance specification and pricing:<br>\n<a href=\"https://aws.amazon.com/ec2/instance-types/g4/\" target=\"_blank\">https://aws.amazon.com/ec2/instance-types/g4/</a></p>",
  "messages": [
    {
      "id": "1531409",
      "postDate": "10/02/2021 00:47:27",
      "content": "<p>Since there are always questions about hardware in these types of competitions, I will share my most economical model: relatively good performance at an affordable price.</p>\n<p>Probably the best candidate on my side is the B3 model. I used the instance g4dn.12xlarge, and the model trained for about 48 hours (10 epochs on 1.6m images). The instance cost is 3.9 USD / hour, which yields a total training cost of 187 USD. The private score of this model is 0.357, which would be 24th place on the LB.</p>\n<p>Instance specification and pricing:<br>\n<a href=\"https://aws.amazon.com/ec2/instance-types/g4/\" target=\"_blank\">https://aws.amazon.com/ec2/instance-types/g4/</a></p>",
      "rawMarkdown": "Since there are always questions about hardware in these types of competitions, I will share my most economical model: relatively good performance at an affordable price.\n\nProbably the best candidate on my side is the B3 model. I used the instance g4dn.12xlarge, and the model trained for about 48 hours (10 epochs on 1.6m images). The instance cost is 3.9 USD / hour, which yields a total training cost of 187 USD. The private score of this model is 0.357, which would be 24th place on the LB.\n\nInstance specification and pricing:\nhttps://aws.amazon.com/ec2/instance-types/g4/",
      "votes": null
    },
    {
      "id": "1531411",
      "postDate": "10/02/2021 00:52:27",
      "content": "<p>Awesome to know <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I was just looking at that instance earlier. </p>",
      "rawMarkdown": "Awesome to know @narsil I was just looking at that instance earlier.",
      "votes": null
    },
    {
      "id": "1531616",
      "postDate": "10/02/2021 07:26:46",
      "content": "<p>To be fair, it is not so much about length of training the final model, but hardware requirements to get to that final model, i.e. how many experiments did you need, run, etc.</p>",
      "rawMarkdown": "To be fair, it is not so much about length of training the final model, but hardware requirements to get to that final model, i.e. how many experiments did you need, run, etc.",
      "votes": null
    },
    {
      "id": "1531631",
      "postDate": "10/02/2021 07:51:07",
      "content": "<p>Well, I wrote more about training costs (rather than training time) which capture the hardware requirements. And it indeed only measures the final model, not the road to achieve it. But it is still an interesting efficiency KPI to compare between models. E.g. when evaluating models for production which would have to be retrained frequently (typical working setting I operate in on daily basis, albeit not for CV models, but for more for events and time-series) this would exactly be the measure that I would monitor</p>",
      "rawMarkdown": "Well, I wrote more about training costs (rather than training time) which capture the hardware requirements. And it indeed only measures the final model, not the road to achieve it. But it is still an interesting efficiency KPI to compare between models. E.g. when evaluating models for production which would have to be retrained frequently (typical working setting I operate in on daily basis, albeit not for CV models, but for more for events and time-series) this would exactly be the measure that I would monitor",
      "votes": null
    },
    {
      "id": "1532238",
      "postDate": "10/02/2021 18:38:39",
      "content": "<p>I trained my model on kaggle tpu. So It was for free. But I just used effectnet 5. With a bigger model I am sure I would have scored higher</p>",
      "rawMarkdown": "I trained my model on kaggle tpu. So It was for free. But I just used effectnet 5. With a bigger model I am sure I would have scored higher",
      "votes": null
    },
    {
      "id": "1532439",
      "postDate": "10/03/2021 01:52:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/lucamtb\" target=\"_blank\">@lucamtb</a> , may I ask what's your effnetb5 validation acc? cause somehow I can't get it work.<br>\nbtw, thanks for your notebook, it helps me a lot.</p>",
      "rawMarkdown": "Hi @lucamtb , may I ask what's your effnetb5 validation acc? cause somehow I can't get it work.\nbtw, thanks for your notebook, it helps me a lot.",
      "votes": null
    },
    {
      "id": "1533236",
      "postDate": "10/03/2021 19:24:26",
      "content": "<p>TPU on Kaggle is free, but on standard pricing it is 8$/hour, so we could also compare the costs of models' training on TPU.</p>\n<p><a href=\"https://cloud.google.com/tpu/pricing\" target=\"_blank\">https://cloud.google.com/tpu/pricing</a></p>",
      "rawMarkdown": "TPU on Kaggle is free, but on standard pricing it is 8$/hour, so we could also compare the costs of models' training on TPU.\n\nhttps://cloud.google.com/tpu/pricing",
      "votes": null
    },
    {
      "id": "2269866",
      "postDate": "05/22/2023 18:54:11",
      "content": "<p>perfect.<br>\ndo you mean bazel commitee by b3?</p>",
      "rawMarkdown": "perfect.\ndo you mean bazel commitee by b3?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1531411,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "10/02/2021 00:52:27",
      "content": "<p>Awesome to know <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> I was just looking at that instance earlier. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1531616,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "10/02/2021 07:26:46",
      "content": "<p>To be fair, it is not so much about length of training the final model, but hardware requirements to get to that final model, i.e. how many experiments did you need, run, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1531631,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "10/02/2021 07:51:07",
          "content": "<p>Well, I wrote more about training costs (rather than training time) which capture the hardware requirements. And it indeed only measures the final model, not the road to achieve it. But it is still an interesting efficiency KPI to compare between models. E.g. when evaluating models for production which would have to be retrained frequently (typical working setting I operate in on daily basis, albeit not for CV models, but for more for events and time-series) this would exactly be the measure that I would monitor</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1532238,
      "author_name": "lucamtb",
      "author_url": "",
      "post_date": "10/02/2021 18:38:39",
      "content": "<p>I trained my model on kaggle tpu. So It was for free. But I just used effectnet 5. With a bigger model I am sure I would have scored higher</p>",
      "votes": null,
      "replies": [
        {
          "id": 1532439,
          "author_name": "gdoong",
          "author_url": "",
          "post_date": "10/03/2021 01:52:06",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/lucamtb\" target=\"_blank\">@lucamtb</a> , may I ask what's your effnetb5 validation acc? cause somehow I can't get it work.<br>\nbtw, thanks for your notebook, it helps me a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1533236,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "10/03/2021 19:24:26",
          "content": "<p>TPU on Kaggle is free, but on standard pricing it is 8$/hour, so we could also compare the costs of models' training on TPU.</p>\n<p><a href=\"https://cloud.google.com/tpu/pricing\" target=\"_blank\">https://cloud.google.com/tpu/pricing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2269866,
      "author_name": "batoolhosseini",
      "author_url": "",
      "post_date": "05/22/2023 18:54:11",
      "content": "<p>perfect.<br>\ndo you mean bazel commitee by b3?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1531409": "Since there are always questions about hardware in these types of competitions, I will share my most economical model: relatively good performance at an affordable price.\n\nProbably the best candidate on my side is the B3 model. I used the instance g4dn.12xlarge, and the model trained for about 48 hours (10 epochs on 1.6m images). The instance cost is 3.9 USD / hour, which yields a total training cost of 187 USD. The private score of this model is 0.357, which would be 24th place on the LB.\n\nInstance specification and pricing:\nhttps://aws.amazon.com/ec2/instance-types/g4/",
    "1531411": "Awesome to know @narsil I was just looking at that instance earlier.",
    "1531616": "To be fair, it is not so much about length of training the final model, but hardware requirements to get to that final model, i.e. how many experiments did you need, run, etc.",
    "1531631": "Well, I wrote more about training costs (rather than training time) which capture the hardware requirements. And it indeed only measures the final model, not the road to achieve it. But it is still an interesting efficiency KPI to compare between models. E.g. when evaluating models for production which would have to be retrained frequently (typical working setting I operate in on daily basis, albeit not for CV models, but for more for events and time-series) this would exactly be the measure that I would monitor",
    "1532238": "I trained my model on kaggle tpu. So It was for free. But I just used effectnet 5. With a bigger model I am sure I would have scored higher",
    "1532439": "Hi @lucamtb , may I ask what's your effnetb5 validation acc? cause somehow I can't get it work.\nbtw, thanks for your notebook, it helps me a lot.",
    "1533236": "TPU on Kaggle is free, but on standard pricing it is 8$/hour, so we could also compare the costs of models' training on TPU.\n\nhttps://cloud.google.com/tpu/pricing",
    "2269866": "perfect.\ndo you mean bazel commitee by b3?"
  },
  "source": "meta"
}