{
  "id": 227328,
  "title": "Long prediction time",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/227328",
  "author_name": "Luigi Saetta",
  "post_date": "2021-03-19T23:31:13.998000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>One of the complications in this competition is that the test set is really large (243020 images).<br>\nAnother source of complication is that for each image predict produces 64500 probs (one for each class).</p>\n<p>Therefore, at least in my approach, it is not possible to do the prediction in one call (I get \"socket closed\" on TPU), therefore, I have been obliged to break the prediction into batches. This makes rather complicated to do k-fold (I need to predict with 5 different model)… without breaking into batches it is easy to get out of memory.</p>\n<p>The entire prediction takes on TPU 4 hours, maybe I can find some optimizations. In the end, it is faster to train than predicts</p>\n<p>As soon I consolidate something, I'll publish the Notebooks. For now, the result in the leaderboard has been obtained using only 600K images to train…. for sure I can get better results by increasing the train set.</p>",
  "messages": [
    {
      "id": 1245797,
      "postDate": "2021-03-20T07:04:42.440Z",
      "content": "<p>My experiments were with GPU + Resnet18 (<code>FOLDS = 0</code>), and it took 6 hours to train for a single epoch for the whole training set with minimal augmentations. </p>\n<p>The inference kernel took 70 minutes though. Why did it take 4 hours on predictions though ?? 🧐🧐</p>",
      "rawMarkdown": "My experiments were with GPU + Resnet18 (`FOLDS = 0`), and it took 6 hours to train for a single epoch for the whole training set with minimal augmentations. \n\nThe inference kernel took 70 minutes though. Why did it take 4 hours on predictions though ?? 🧐🧐",
      "votes": 1,
      "replies": [
        {
          "id": 1245838,
          "postDate": "2021-03-20T08:26:30.183Z",
          "content": "<p>I need to undertsnad. Probably because my model is moch more… heavvy, I have used EfficientNet B$, second I do 5 fold CV, so I need to evaluate on 5 different models and then do average of the predictions. Average is complicated because the entire np array is huge (240Kx64500 float numbers and therefore I need to find a smart way to handle it)</p>",
          "rawMarkdown": "I need to undertsnad. Probably because my model is moch more... heavvy, I have used EfficientNet B$, second I do 5 fold CV, so I need to evaluate on 5 different models and then do average of the predictions. Average is complicated because the entire np array is huge (240Kx64500 float numbers and therefore I need to find a smart way to handle it)"
        }
      ]
    },
    {
      "id": 1245585,
      "postDate": "2021-03-19T23:31:14Z",
      "content": "<p>One of the complications in this competition is that the test set is really large (243020 images).<br>\nAnother source of complication is that for each image predict produces 64500 probs (one for each class).</p>\n<p>Therefore, at least in my approach, it is not possible to do the prediction in one call (I get \"socket closed\" on TPU), therefore, I have been obliged to break the prediction into batches. This makes rather complicated to do k-fold (I need to predict with 5 different model)… without breaking into batches it is easy to get out of memory.</p>\n<p>The entire prediction takes on TPU 4 hours, maybe I can find some optimizations. In the end, it is faster to train than predicts</p>\n<p>As soon I consolidate something, I'll publish the Notebooks. For now, the result in the leaderboard has been obtained using only 600K images to train…. for sure I can get better results by increasing the train set.</p>",
      "rawMarkdown": "One of the complications in this competition is that the test set is really large (243020 images).\nAnother source of complication is that for each image predict produces 64500 probs (one for each class).\n\nTherefore, at least in my approach, it is not possible to do the prediction in one call (I get \"socket closed\" on TPU), therefore, I have been obliged to break the prediction into batches. This makes rather complicated to do k-fold (I need to predict with 5 different model)... without breaking into batches it is easy to get out of memory.\n\nThe entire prediction takes on TPU 4 hours, maybe I can find some optimizations. In the end, it is faster to train than predicts\n\nAs soon I consolidate something, I'll publish the Notebooks. For now, the result in the leaderboard has been obtained using only 600K images to train.... for sure I can get better results by increasing the train set.\n\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1245797,
      "author_name": "Saurav Maheshkar ☕️",
      "author_url": "",
      "post_date": "2021-03-20T07:04:42.440000",
      "content": "<p>My experiments were with GPU + Resnet18 (<code>FOLDS = 0</code>), and it took 6 hours to train for a single epoch for the whole training set with minimal augmentations. </p>\n<p>The inference kernel took 70 minutes though. Why did it take 4 hours on predictions though ?? 🧐🧐</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1245838,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2021-03-20T08:26:30.183000",
          "content": "<p>I need to undertsnad. Probably because my model is moch more… heavvy, I have used EfficientNet B$, second I do 5 fold CV, so I need to evaluate on 5 different models and then do average of the predictions. Average is complicated because the entire np array is huge (240Kx64500 float numbers and therefore I need to find a smart way to handle it)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1245797": "My experiments were with GPU + Resnet18 (`FOLDS = 0`), and it took 6 hours to train for a single epoch for the whole training set with minimal augmentations. \n\nThe inference kernel took 70 minutes though. Why did it take 4 hours on predictions though ?? 🧐🧐",
    "1245585": "One of the complications in this competition is that the test set is really large (243020 images).\nAnother source of complication is that for each image predict produces 64500 probs (one for each class).\n\nTherefore, at least in my approach, it is not possible to do the prediction in one call (I get \"socket closed\" on TPU), therefore, I have been obliged to break the prediction into batches. This makes rather complicated to do k-fold (I need to predict with 5 different model)... without breaking into batches it is easy to get out of memory.\n\nThe entire prediction takes on TPU 4 hours, maybe I can find some optimizations. In the end, it is faster to train than predicts\n\nAs soon I consolidate something, I'll publish the Notebooks. For now, the result in the leaderboard has been obtained using only 600K images to train.... for sure I can get better results by increasing the train set.\n\n"
  }
}