{
  "id": 307615,
  "title": "839772 images in the training dataset!!",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/307615",
  "author_name": "Hamza",
  "post_date": "2022-02-15T01:25:14.570000",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hello Guys,</p>\n<p>I'm surprised, as I found that we have 839 772 images in the training dataset…<br>\nI think we will struggle in this competition to train our models using Kaggle kernels!<br>\nDo Kaggle kernels allow us to train a model on 839 772 images without <b>memory error<b>?🤔!!!</b></b></p>",
  "messages": [
    {
      "id": 1690522,
      "postDate": "2022-02-15T01:25:14.570Z",
      "content": "<p>Hello Guys,</p>\n<p>I'm surprised, as I found that we have 839 772 images in the training dataset…<br>\nI think we will struggle in this competition to train our models using Kaggle kernels!<br>\nDo Kaggle kernels allow us to train a model on 839 772 images without <b>memory error<b>?🤔!!!</b></b></p>",
      "rawMarkdown": "Hello Guys,\n\nI'm surprised, as I found that we have 839 772 images in the training dataset...\nI think we will struggle in this competition to train our models using Kaggle kernels!\nDo Kaggle kernels allow us to train a model on 839 772 images without <b>memory error<b>?🤔!!!",
      "votes": 11
    },
    {
      "id": 1692712,
      "postDate": "2022-02-16T07:44:11.543Z",
      "content": "<p>we are successfully running EfficientNet D0 with batch size 128 and downscale images to input size 224px<br>\n<a href=\"https://www.kaggle.com/jirkaborovec/herbarium-eda-baseline-flash-efficientnet\" target=\"_blank\">🌿Herbarium: EDA 🔎 &amp; baseline Flash⚡EfficientNet</a></p>",
      "rawMarkdown": "we are successfully running EfficientNet D0 with batch size 128 and downscale images to input size 224px\n[🌿Herbarium: EDA 🔎 & baseline Flash⚡EfficientNet](https://www.kaggle.com/jirkaborovec/herbarium-eda-baseline-flash-efficientnet)",
      "votes": 3,
      "replies": [
        {
          "id": 1693243,
          "postDate": "2022-02-16T14:43:47.100Z",
          "content": "<p>Great, thanks for sharing</p>",
          "rawMarkdown": "Great, thanks for sharing"
        }
      ]
    },
    {
      "id": 1691146,
      "postDate": "2022-02-15T08:58:04.800Z",
      "content": "<p>Small batches, small architectures - the solution! 😄Good luck, <a href=\"https://www.kaggle.com/hamzaghanmi\" target=\"_blank\">@hamzaghanmi</a>! </p>",
      "rawMarkdown": "Small batches, small architectures - the solution! 😄Good luck, @hamzaghanmi! ",
      "votes": 3,
      "replies": [
        {
          "id": 1691166,
          "postDate": "2022-02-15T09:09:08.540Z",
          "content": "<p>…and don't forget data sampling! When sampling the data, try to include a reasonable amount of images for each class.</p>",
          "rawMarkdown": "...and don't forget data sampling! When sampling the data, try to include a reasonable amount of images for each class.",
          "votes": 2
        },
        {
          "id": 1691235,
          "postDate": "2022-02-15T10:04:46.710Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a> </p>",
          "rawMarkdown": "Thank you @vad13irt ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1691184,
      "postDate": "2022-02-15T09:23:47.130Z",
      "content": "<p>Plus no points towards tiers/rankings :(</p>",
      "rawMarkdown": "Plus no points towards tiers/rankings :(",
      "votes": 1,
      "replies": [
        {
          "id": 1691209,
          "postDate": "2022-02-15T09:48:21.973Z",
          "content": "<p>Last year comp had just 80 participants.</p>",
          "rawMarkdown": "Last year comp had just 80 participants.",
          "votes": 1
        },
        {
          "id": 1691253,
          "postDate": "2022-02-15T10:09:23.210Z",
          "content": "<p>ooooh ☹️ I didn't notice that before</p>",
          "rawMarkdown": "ooooh ☹️ I didn't notice that before"
        }
      ]
    },
    {
      "id": 1691174,
      "postDate": "2022-02-15T09:15:27.263Z",
      "content": "<p>there is also a good chance to try contrastive / zero/few shot learning since the number of classes is quite absurd </p>",
      "rawMarkdown": "there is also a good chance to try contrastive / zero/few shot learning since the number of classes is quite absurd ",
      "votes": 1,
      "replies": [
        {
          "id": 1691263,
          "postDate": "2022-02-15T10:15:40.770Z",
          "content": "<p>15,501 classes !!! 😨</p>",
          "rawMarkdown": "15,501 classes !!! 😨"
        }
      ]
    },
    {
      "id": 1691161,
      "postDate": "2022-02-15T09:06:51.990Z",
      "content": "<p>Managing memory is a key skill for data scientists it would seem! Personally I enjoy it when there are other challenges alongside just making a model, it forces me to learn new things 😃</p>",
      "rawMarkdown": "Managing memory is a key skill for data scientists it would seem! Personally I enjoy it when there are other challenges alongside just making a model, it forces me to learn new things 😃",
      "votes": 2,
      "replies": [
        {
          "id": 1691261,
          "postDate": "2022-02-15T10:14:48.333Z",
          "content": "<p>I do agree with you, learning new things is the most important part 💯</p>",
          "rawMarkdown": "I do agree with you, learning new things is the most important part 💯",
          "votes": 1
        }
      ]
    },
    {
      "id": 1690814,
      "postDate": "2022-02-15T05:38:33.130Z",
      "content": "<p>i think like the previous runs of this competition you will need to sample the data or have your own supercomputer 😂</p>",
      "rawMarkdown": "i think like the previous runs of this competition you will need to sample the data or have your own supercomputer 😂",
      "votes": 2,
      "replies": [
        {
          "id": 1691021,
          "postDate": "2022-02-15T07:32:15.977Z",
          "content": "<p>I agree - you'll have to sample data. Also I'd use tfrecords and TPU here.</p>",
          "rawMarkdown": "I agree - you'll have to sample data. Also I'd use tfrecords and TPU here.",
          "votes": 1
        },
        {
          "id": 1691124,
          "postDate": "2022-02-15T08:40:04.673Z",
          "content": "<p>Using tfrecords and TPU is a good idea but again it will be very slow with +800000 images… <br> We need a supercomputer 💯! </p>",
          "rawMarkdown": "Using tfrecords and TPU is a good idea but again it will be very slow with +800000 images... <br> We need a supercomputer 💯! "
        },
        {
          "id": 1691162,
          "postDate": "2022-02-15T09:07:05.830Z",
          "content": "<p>I believe it's possible to sample data to achieve good model accuracy. I don't think full data training is an option here. And I doubt that a supercomputer will be useful here - it's very specific equipment requiring special programming and used for specific tasks.</p>",
          "rawMarkdown": "I believe it's possible to sample data to achieve good model accuracy. I don't think full data training is an option here. And I doubt that a supercomputer will be useful here - it's very specific equipment requiring special programming and used for specific tasks.",
          "votes": 1
        },
        {
          "id": 1691232,
          "postDate": "2022-02-15T10:03:54.490Z",
          "content": "<p>I Totally agree with you <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> 💯</p>",
          "rawMarkdown": "I Totally agree with you @atamazian 💯"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1692712,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2022-02-16T07:44:11.543000",
      "content": "<p>we are successfully running EfficientNet D0 with batch size 128 and downscale images to input size 224px<br>\n<a href=\"https://www.kaggle.com/jirkaborovec/herbarium-eda-baseline-flash-efficientnet\" target=\"_blank\">🌿Herbarium: EDA 🔎 &amp; baseline Flash⚡EfficientNet</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1693243,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-16T14:43:47.100000",
          "content": "<p>Great, thanks for sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1691146,
      "author_name": "Vadim Irtlach",
      "author_url": "",
      "post_date": "2022-02-15T08:58:04.800000",
      "content": "<p>Small batches, small architectures - the solution! 😄Good luck, <a href=\"https://www.kaggle.com/hamzaghanmi\" target=\"_blank\">@hamzaghanmi</a>! </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1691166,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2022-02-15T09:09:08.540000",
          "content": "<p>…and don't forget data sampling! When sampling the data, try to include a reasonable amount of images for each class.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1691235,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T10:04:46.710000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/vad13irt\" target=\"_blank\">@vad13irt</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1691184,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2022-02-15T09:23:47.130000",
      "content": "<p>Plus no points towards tiers/rankings :(</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1691209,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2022-02-15T09:48:21.973000",
          "content": "<p>Last year comp had just 80 participants.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1691253,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T10:09:23.210000",
          "content": "<p>ooooh ☹️ I didn't notice that before</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1691174,
      "author_name": "Loy Juncheng",
      "author_url": "",
      "post_date": "2022-02-15T09:15:27.263000",
      "content": "<p>there is also a good chance to try contrastive / zero/few shot learning since the number of classes is quite absurd </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1691263,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T10:15:40.770000",
          "content": "<p>15,501 classes !!! 😨</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1691161,
      "author_name": "Jake",
      "author_url": "",
      "post_date": "2022-02-15T09:06:51.990000",
      "content": "<p>Managing memory is a key skill for data scientists it would seem! Personally I enjoy it when there are other challenges alongside just making a model, it forces me to learn new things 😃</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1691261,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T10:14:48.333000",
          "content": "<p>I do agree with you, learning new things is the most important part 💯</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1690814,
      "author_name": "Loy Juncheng",
      "author_url": "",
      "post_date": "2022-02-15T05:38:33.130000",
      "content": "<p>i think like the previous runs of this competition you will need to sample the data or have your own supercomputer 😂</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1691021,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2022-02-15T07:32:15.977000",
          "content": "<p>I agree - you'll have to sample data. Also I'd use tfrecords and TPU here.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1691124,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T08:40:04.673000",
          "content": "<p>Using tfrecords and TPU is a good idea but again it will be very slow with +800000 images… <br> We need a supercomputer 💯! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1691162,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2022-02-15T09:07:05.830000",
          "content": "<p>I believe it's possible to sample data to achieve good model accuracy. I don't think full data training is an option here. And I doubt that a supercomputer will be useful here - it's very specific equipment requiring special programming and used for specific tasks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1691232,
          "author_name": "Hamza",
          "author_url": "",
          "post_date": "2022-02-15T10:03:54.490000",
          "content": "<p>I Totally agree with you <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> 💯</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1690522": "Hello Guys,\n\nI'm surprised, as I found that we have 839 772 images in the training dataset...\nI think we will struggle in this competition to train our models using Kaggle kernels!\nDo Kaggle kernels allow us to train a model on 839 772 images without <b>memory error<b>?🤔!!!",
    "1692712": "we are successfully running EfficientNet D0 with batch size 128 and downscale images to input size 224px\n[🌿Herbarium: EDA 🔎 & baseline Flash⚡EfficientNet](https://www.kaggle.com/jirkaborovec/herbarium-eda-baseline-flash-efficientnet)",
    "1691146": "Small batches, small architectures - the solution! 😄Good luck, @hamzaghanmi! ",
    "1691184": "Plus no points towards tiers/rankings :(",
    "1691174": "there is also a good chance to try contrastive / zero/few shot learning since the number of classes is quite absurd ",
    "1691161": "Managing memory is a key skill for data scientists it would seem! Personally I enjoy it when there are other challenges alongside just making a model, it forces me to learn new things 😃",
    "1690814": "i think like the previous runs of this competition you will need to sample the data or have your own supercomputer 😂"
  }
}