{
  "id": 228588,
  "title": "Why difficult?",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/228588",
  "author_name": "Luigi Saetta",
  "post_date": "2021-03-25T11:59:25.514000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>It seems that there are not so many people working, at least till now, on this competition.</p>\n<ul>\n<li>It could be boring.</li>\n<li>It could be because there are other competitions similar now finishing (then we should expect more people entering in the near future)</li>\n<li>It could be because for some reason the barrier to entry is high.</li>\n</ul>\n<p>I want to share my opinion: probably I've not chosen till now a very smart approach, I'm thinking on it. But for sure one problem is that the training set and even the test set (240K + images) are huge.</p>\n<p>I'm currently using a CNN (EFNET4) and TPU, with K-fold cross-validation and the training process takes about 8 hours. Even if I have packed all images (resized to 256x256) in TFRecords files.</p>\n<p>In addition, after having trained the five model (Folds = 5), since the test set is huge, to do the predictions I was completed to cleverly (hope so) split in batches and the prediction takes a separate run, during about 1.5 hours (for now I use only 3 out of five of the trained models)</p>\n<p>To keep the story short, I have run out of TPU, and I'm waiting for the replenishment.</p>\n<p>I have discovered that I can do, in a reasonable time, predictions on GPU… but still for training I need TPU (so, waiting and thinking).</p>\n<p>Ah, consider that for now I'm using only 60% of the training images… otherwise, I would be compelled to exclude K-old.</p>\n<p>Would be interesting, even for newbies, to share what you think about these aspects of complexities and eventually share how are you managing them.</p>\n<p>Last but not least: it is nice to see that a CNN can achieve 90% accuracy classifying in 64500 classes… much more than ILSVRC.</p>",
  "messages": [
    {
      "id": 1252067,
      "postDate": "2021-03-25T11:59:25.513Z",
      "content": "<p>It seems that there are not so many people working, at least till now, on this competition.</p>\n<ul>\n<li>It could be boring.</li>\n<li>It could be because there are other competitions similar now finishing (then we should expect more people entering in the near future)</li>\n<li>It could be because for some reason the barrier to entry is high.</li>\n</ul>\n<p>I want to share my opinion: probably I've not chosen till now a very smart approach, I'm thinking on it. But for sure one problem is that the training set and even the test set (240K + images) are huge.</p>\n<p>I'm currently using a CNN (EFNET4) and TPU, with K-fold cross-validation and the training process takes about 8 hours. Even if I have packed all images (resized to 256x256) in TFRecords files.</p>\n<p>In addition, after having trained the five model (Folds = 5), since the test set is huge, to do the predictions I was completed to cleverly (hope so) split in batches and the prediction takes a separate run, during about 1.5 hours (for now I use only 3 out of five of the trained models)</p>\n<p>To keep the story short, I have run out of TPU, and I'm waiting for the replenishment.</p>\n<p>I have discovered that I can do, in a reasonable time, predictions on GPU… but still for training I need TPU (so, waiting and thinking).</p>\n<p>Ah, consider that for now I'm using only 60% of the training images… otherwise, I would be compelled to exclude K-old.</p>\n<p>Would be interesting, even for newbies, to share what you think about these aspects of complexities and eventually share how are you managing them.</p>\n<p>Last but not least: it is nice to see that a CNN can achieve 90% accuracy classifying in 64500 classes… much more than ILSVRC.</p>",
      "rawMarkdown": "It seems that there are not so many people working, at least till now, on this competition.\n\n- It could be boring.\n- It could be because there are other competitions similar now finishing (then we should expect more people entering in the near future)\n- It could be because for some reason the barrier to entry is high.\n\nI want to share my opinion: probably I've not chosen till now a very smart approach, I'm thinking on it. But for sure one problem is that the training set and even the test set (240K + images) are huge.\n\nI'm currently using a CNN (EFNET4) and TPU, with K-fold cross-validation and the training process takes about 8 hours. Even if I have packed all images (resized to 256x256) in TFRecords files.\n\nIn addition, after having trained the five model (Folds = 5), since the test set is huge, to do the predictions I was completed to cleverly (hope so) split in batches and the prediction takes a separate run, during about 1.5 hours (for now I use only 3 out of five of the trained models)\n\nTo keep the story short, I have run out of TPU, and I'm waiting for the replenishment.\n\nI have discovered that I can do, in a reasonable time, predictions on GPU... but still for training I need TPU (so, waiting and thinking).\n\nAh, consider that for now I'm using only 60% of the training images... otherwise, I would be compelled to exclude K-old.\n\nWould be interesting, even for newbies, to share what you think about these aspects of complexities and eventually share how are you managing them.\n\nLast but not least: it is nice to see that a CNN can achieve 90% accuracy classifying in 64500 classes... much more than ILSVRC.",
      "votes": 3
    },
    {
      "id": 1292073,
      "postDate": "2021-05-03T16:07:12.893Z",
      "content": "<p>It's good to read other's views on this challenge.  Although I could not disagree more with point number 1 - the topic is not boring at all!<br>\nIn my case the problem is that training on the entire training set is too much for the kernel and it dies after some time (9 hours? I'm not sure). I've just tried to load the dataset in Colab and even there there is not enough storage for the entire dataset 😓<br>\nSo what I'm doing now is training very basic models (resnet18, 34 and 50) with 1/100th of the dataset. I'm scared of trying with bigger models or even unfreezing some layers of the model apart from the FC.<br>\nI'm also planning to do some sort of \"smart\" subsampling\" so every category gets at least 5 images and share the CSV file I get.<br>\nDoes anyone know if there is such \"filtered CSV\" out there?</p>",
      "rawMarkdown": "It's good to read other's views on this challenge.  Although I could not disagree more with point number 1 - the topic is not boring at all!\nIn my case the problem is that training on the entire training set is too much for the kernel and it dies after some time (9 hours? I'm not sure). I've just tried to load the dataset in Colab and even there there is not enough storage for the entire dataset 😓\nSo what I'm doing now is training very basic models (resnet18, 34 and 50) with 1/100th of the dataset. I'm scared of trying with bigger models or even unfreezing some layers of the model apart from the FC.\nI'm also planning to do some sort of \"smart\" subsampling\" so every category gets at least 5 images and share the CSV file I get.\nDoes anyone know if there is such \"filtered CSV\" out there?",
      "votes": 1
    },
    {
      "id": 1252171,
      "postDate": "2021-03-25T13:23:57.717Z",
      "content": "<p>Honestly, I think point 3 (barrier to entry is high) is the most important, especially for people who are planning to train on computers that are available to them locally/their own computers. Maybe I'm biased about point 1 (boring) because I have an interest in both botany and AI 😅 </p>\n<p>If people wanted to use their own locally available machine and train a SotA model like EFNet at a high resolution with &gt; 16 batch size they'd have to do gradient accumulation. Then, it becomes hard to do fail-fast testing for people who do have to do gradient accumulation, since training a model (at least in my testing on last years dataset) took almost a week for 12 epochs on a single 1080Ti.</p>\n<p>I'm expecting people who have more experience to quickly overtake me in accuracy, so I'm constantly trying new things and testing new methodologies in order to see how well they work (I don't think I'm smart nor knowledgeable enough to make something cool like TResNet).</p>\n<p>And yeah, the 240K * 64500 problem is a big pain and I'm not sure how to solve it other than processing it in batches or saving the array to disk and loading it partially to memory (which is basically processing it in batches). I don't do CV in my submissions but maybe it will be key in this competition.</p>",
      "rawMarkdown": "Honestly, I think point 3 (barrier to entry is high) is the most important, especially for people who are planning to train on computers that are available to them locally/their own computers. Maybe I'm biased about point 1 (boring) because I have an interest in both botany and AI 😅 \n\nIf people wanted to use their own locally available machine and train a SotA model like EFNet at a high resolution with > 16 batch size they'd have to do gradient accumulation. Then, it becomes hard to do fail-fast testing for people who do have to do gradient accumulation, since training a model (at least in my testing on last years dataset) took almost a week for 12 epochs on a single 1080Ti.\n\nI'm expecting people who have more experience to quickly overtake me in accuracy, so I'm constantly trying new things and testing new methodologies in order to see how well they work (I don't think I'm smart nor knowledgeable enough to make something cool like TResNet).\n\nAnd yeah, the 240K * 64500 problem is a big pain and I'm not sure how to solve it other than processing it in batches or saving the array to disk and loading it partially to memory (which is basically processing it in batches). I don't do CV in my submissions but maybe it will be key in this competition.",
      "votes": 1,
      "replies": [
        {
          "id": 1252212,
          "postDate": "2021-03-25T13:58:05.273Z",
          "content": "<p>Well, regarding K-fold CV… it is difficult for the above-mentioned reasons. I have taken the approach of dividing into 5 the training set, but then actually I stop after the third fold. It will take me another two hours for 5 fold and TPU session max duration is not enough.</p>\n<p>As soon as I get more TPU time I'll try with EF2 (or lower) to see what changes (maybe less time, same results)</p>",
          "rawMarkdown": "Well, regarding K-fold CV... it is difficult for the above-mentioned reasons. I have taken the approach of dividing into 5 the training set, but then actually I stop after the third fold. It will take me another two hours for 5 fold and TPU session max duration is not enough.\n\nAs soon as I get more TPU time I'll try with EF2 (or lower) to see what changes (maybe less time, same results)\n\n"
        }
      ]
    },
    {
      "id": 1252336,
      "postDate": "2021-03-25T15:41:40.803Z",
      "content": "<p>I'm probably also biased for point 1 :)</p>\n<p>From an objective point of view, I agree that the size of the dataset is probably a high barrier to entry for many participants but I also think it is a large part of the interest. Such a large number of species is exciting and can potentially lead to new interesting and creative solutions, especially when paired with the hierarchical labels!</p>\n<p>For running the predictions on the test set, probably doing so in batches is the easiest solution. </p>\n<p>An additional tip, if you would want to experiment with a smaller training set and iterate more quickly through your ideas, you can always create your own subset of the training data and experiment on that first, before training your model on the full training set. </p>\n<p>Anyways I'm happy that you are giving it a good try! Good luck with the experimentation and very excited to see what creative solutions come out of it! </p>",
      "rawMarkdown": "I'm probably also biased for point 1 :)\n\nFrom an objective point of view, I agree that the size of the dataset is probably a high barrier to entry for many participants but I also think it is a large part of the interest. Such a large number of species is exciting and can potentially lead to new interesting and creative solutions, especially when paired with the hierarchical labels!\n\nFor running the predictions on the test set, probably doing so in batches is the easiest solution. \n\nAn additional tip, if you would want to experiment with a smaller training set and iterate more quickly through your ideas, you can always create your own subset of the training data and experiment on that first, before training your model on the full training set. \n\n\nAnyways I'm happy that you are giving it a good try! Good luck with the experimentation and very excited to see what creative solutions come out of it! \n ",
      "votes": 2,
      "replies": [
        {
          "id": 1252347,
          "postDate": "2021-03-25T15:54:07.710Z",
          "content": "<p>Hi Riccardo, thanks for your suggestions… the hierarchical labels is definitively something a should look into carefully.</p>\n<p>I do think this competition is really interesting, my post was to stimulate discussions and some more participation and ideas.</p>",
          "rawMarkdown": "Hi Riccardo, thanks for your suggestions... the hierarchical labels is definitively something a should look into carefully.\n\nI do think this competition is really interesting, my post was to stimulate discussions and some more participation and ideas.\n ",
          "votes": 1
        },
        {
          "id": 1252352,
          "postDate": "2021-03-25T15:56:45.287Z",
          "content": "<p>Of course, and thank you for animating the discussion forum!</p>",
          "rawMarkdown": "Of course, and thank you for animating the discussion forum!"
        },
        {
          "id": 1252367,
          "postDate": "2021-03-25T16:13:13.517Z",
          "content": "<p>With the hierarchical labels - I agree most definitely. It's something that my mentor and I have been thinking about for a while, just don't know how to put it into a model/models \"elegantly.\" </p>\n<p>I'm sure someone will figure it out, and if it performs well that would be super awesome to see!</p>",
          "rawMarkdown": "With the hierarchical labels - I agree most definitely. It's something that my mentor and I have been thinking about for a while, just don't know how to put it into a model/models \"elegantly.\" \n\nI'm sure someone will figure it out, and if it performs well that would be super awesome to see!"
        },
        {
          "id": 1298696,
          "postDate": "2021-05-09T05:59:58.810Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1292073,
      "author_name": "Ignacio Hernández",
      "author_url": "",
      "post_date": "2021-05-03T16:07:12.893000",
      "content": "<p>It's good to read other's views on this challenge.  Although I could not disagree more with point number 1 - the topic is not boring at all!<br>\nIn my case the problem is that training on the entire training set is too much for the kernel and it dies after some time (9 hours? I'm not sure). I've just tried to load the dataset in Colab and even there there is not enough storage for the entire dataset 😓<br>\nSo what I'm doing now is training very basic models (resnet18, 34 and 50) with 1/100th of the dataset. I'm scared of trying with bigger models or even unfreezing some layers of the model apart from the FC.<br>\nI'm also planning to do some sort of \"smart\" subsampling\" so every category gets at least 5 images and share the CSV file I get.<br>\nDoes anyone know if there is such \"filtered CSV\" out there?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1252171,
      "author_name": "Dax Ledesma",
      "author_url": "",
      "post_date": "2021-03-25T13:23:57.717000",
      "content": "<p>Honestly, I think point 3 (barrier to entry is high) is the most important, especially for people who are planning to train on computers that are available to them locally/their own computers. Maybe I'm biased about point 1 (boring) because I have an interest in both botany and AI 😅 </p>\n<p>If people wanted to use their own locally available machine and train a SotA model like EFNet at a high resolution with &gt; 16 batch size they'd have to do gradient accumulation. Then, it becomes hard to do fail-fast testing for people who do have to do gradient accumulation, since training a model (at least in my testing on last years dataset) took almost a week for 12 epochs on a single 1080Ti.</p>\n<p>I'm expecting people who have more experience to quickly overtake me in accuracy, so I'm constantly trying new things and testing new methodologies in order to see how well they work (I don't think I'm smart nor knowledgeable enough to make something cool like TResNet).</p>\n<p>And yeah, the 240K * 64500 problem is a big pain and I'm not sure how to solve it other than processing it in batches or saving the array to disk and loading it partially to memory (which is basically processing it in batches). I don't do CV in my submissions but maybe it will be key in this competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1252212,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2021-03-25T13:58:05.273000",
          "content": "<p>Well, regarding K-fold CV… it is difficult for the above-mentioned reasons. I have taken the approach of dividing into 5 the training set, but then actually I stop after the third fold. It will take me another two hours for 5 fold and TPU session max duration is not enough.</p>\n<p>As soon as I get more TPU time I'll try with EF2 (or lower) to see what changes (maybe less time, same results)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1252336,
      "author_name": "Riccardo de Lutio",
      "author_url": "",
      "post_date": "2021-03-25T15:41:40.803000",
      "content": "<p>I'm probably also biased for point 1 :)</p>\n<p>From an objective point of view, I agree that the size of the dataset is probably a high barrier to entry for many participants but I also think it is a large part of the interest. Such a large number of species is exciting and can potentially lead to new interesting and creative solutions, especially when paired with the hierarchical labels!</p>\n<p>For running the predictions on the test set, probably doing so in batches is the easiest solution. </p>\n<p>An additional tip, if you would want to experiment with a smaller training set and iterate more quickly through your ideas, you can always create your own subset of the training data and experiment on that first, before training your model on the full training set. </p>\n<p>Anyways I'm happy that you are giving it a good try! Good luck with the experimentation and very excited to see what creative solutions come out of it! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1252347,
          "author_name": "Luigi Saetta",
          "author_url": "",
          "post_date": "2021-03-25T15:54:07.710000",
          "content": "<p>Hi Riccardo, thanks for your suggestions… the hierarchical labels is definitively something a should look into carefully.</p>\n<p>I do think this competition is really interesting, my post was to stimulate discussions and some more participation and ideas.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1252352,
          "author_name": "Riccardo de Lutio",
          "author_url": "",
          "post_date": "2021-03-25T15:56:45.287000",
          "content": "<p>Of course, and thank you for animating the discussion forum!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252367,
          "author_name": "Dax Ledesma",
          "author_url": "",
          "post_date": "2021-03-25T16:13:13.517000",
          "content": "<p>With the hierarchical labels - I agree most definitely. It's something that my mentor and I have been thinking about for a while, just don't know how to put it into a model/models \"elegantly.\" </p>\n<p>I'm sure someone will figure it out, and if it performs well that would be super awesome to see!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298696,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-05-09T05:59:58.810000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1252067": "It seems that there are not so many people working, at least till now, on this competition.\n\n- It could be boring.\n- It could be because there are other competitions similar now finishing (then we should expect more people entering in the near future)\n- It could be because for some reason the barrier to entry is high.\n\nI want to share my opinion: probably I've not chosen till now a very smart approach, I'm thinking on it. But for sure one problem is that the training set and even the test set (240K + images) are huge.\n\nI'm currently using a CNN (EFNET4) and TPU, with K-fold cross-validation and the training process takes about 8 hours. Even if I have packed all images (resized to 256x256) in TFRecords files.\n\nIn addition, after having trained the five model (Folds = 5), since the test set is huge, to do the predictions I was completed to cleverly (hope so) split in batches and the prediction takes a separate run, during about 1.5 hours (for now I use only 3 out of five of the trained models)\n\nTo keep the story short, I have run out of TPU, and I'm waiting for the replenishment.\n\nI have discovered that I can do, in a reasonable time, predictions on GPU... but still for training I need TPU (so, waiting and thinking).\n\nAh, consider that for now I'm using only 60% of the training images... otherwise, I would be compelled to exclude K-old.\n\nWould be interesting, even for newbies, to share what you think about these aspects of complexities and eventually share how are you managing them.\n\nLast but not least: it is nice to see that a CNN can achieve 90% accuracy classifying in 64500 classes... much more than ILSVRC.",
    "1292073": "It's good to read other's views on this challenge.  Although I could not disagree more with point number 1 - the topic is not boring at all!\nIn my case the problem is that training on the entire training set is too much for the kernel and it dies after some time (9 hours? I'm not sure). I've just tried to load the dataset in Colab and even there there is not enough storage for the entire dataset 😓\nSo what I'm doing now is training very basic models (resnet18, 34 and 50) with 1/100th of the dataset. I'm scared of trying with bigger models or even unfreezing some layers of the model apart from the FC.\nI'm also planning to do some sort of \"smart\" subsampling\" so every category gets at least 5 images and share the CSV file I get.\nDoes anyone know if there is such \"filtered CSV\" out there?",
    "1252171": "Honestly, I think point 3 (barrier to entry is high) is the most important, especially for people who are planning to train on computers that are available to them locally/their own computers. Maybe I'm biased about point 1 (boring) because I have an interest in both botany and AI 😅 \n\nIf people wanted to use their own locally available machine and train a SotA model like EFNet at a high resolution with > 16 batch size they'd have to do gradient accumulation. Then, it becomes hard to do fail-fast testing for people who do have to do gradient accumulation, since training a model (at least in my testing on last years dataset) took almost a week for 12 epochs on a single 1080Ti.\n\nI'm expecting people who have more experience to quickly overtake me in accuracy, so I'm constantly trying new things and testing new methodologies in order to see how well they work (I don't think I'm smart nor knowledgeable enough to make something cool like TResNet).\n\nAnd yeah, the 240K * 64500 problem is a big pain and I'm not sure how to solve it other than processing it in batches or saving the array to disk and loading it partially to memory (which is basically processing it in batches). I don't do CV in my submissions but maybe it will be key in this competition.",
    "1252336": "I'm probably also biased for point 1 :)\n\nFrom an objective point of view, I agree that the size of the dataset is probably a high barrier to entry for many participants but I also think it is a large part of the interest. Such a large number of species is exciting and can potentially lead to new interesting and creative solutions, especially when paired with the hierarchical labels!\n\nFor running the predictions on the test set, probably doing so in batches is the easiest solution. \n\nAn additional tip, if you would want to experiment with a smaller training set and iterate more quickly through your ideas, you can always create your own subset of the training data and experiment on that first, before training your model on the full training set. \n\n\nAnyways I'm happy that you are giving it a good try! Good luck with the experimentation and very excited to see what creative solutions come out of it! \n "
  }
}