{
  "id": 312696,
  "title": "Sharing baselines ?",
  "url": "/competitions/geolifeclef-2022-lifeclef-2022-fgvc9/discussion/312696",
  "author_name": "",
  "post_date": "2022-03-13T15:04:15.191481900Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>There doesn't seems to be much involvement from the community in this competition. I was wondering about two things:</p>\n<ul>\n<li>I don't see much sharing from the community (only one notebook shared in the 2021 edition). Usually kaggle is nice for sharing EDA / insightfull baselines. Is there any plan to take notebooks into account even if they don't reach good scores ? </li>\n<li>There seems to be some advanced baselines on the LB, including advanced Computer Vision pytorch stuff. Wouldn't it be better to share those ? (maybe share those a month in ?)  </li>\n</ul>",
  "messages": [
    {
      "id": "1721295",
      "postDate": "03/13/2022 15:04:15",
      "content": "<p>There doesn't seems to be much involvement from the community in this competition. I was wondering about two things:</p>\n<ul>\n<li>I don't see much sharing from the community (only one notebook shared in the 2021 edition). Usually kaggle is nice for sharing EDA / insightfull baselines. Is there any plan to take notebooks into account even if they don't reach good scores ? </li>\n<li>There seems to be some advanced baselines on the LB, including advanced Computer Vision pytorch stuff. Wouldn't it be better to share those ? (maybe share those a month in ?)  </li>\n</ul>",
      "rawMarkdown": "There doesn't seems to be much involvement from the community in this competition. I was wondering about two things:\n- I don't see much sharing from the community (only one notebook shared in the 2021 edition). Usually kaggle is nice for sharing EDA / insightfull baselines. Is there any plan to take notebooks into account even if they don't reach good scores ? \n- There seems to be some advanced baselines on the LB, including advanced Computer Vision pytorch stuff. Wouldn't it be better to share those ? (maybe share those a month in ?)",
      "votes": null
    },
    {
      "id": "1721347",
      "postDate": "03/13/2022 15:51:23",
      "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> Thanks for bringing this competition to the notice of Kagglers. Hope it sees more participation.</p>",
      "rawMarkdown": "lucasmorin Thanks for bringing this competition to the notice of Kagglers. Hope it sees more participation.",
      "votes": null
    },
    {
      "id": "1723683",
      "postDate": "03/15/2022 16:10:00",
      "content": "<p>Well the competition only started, it is not very standard data and training of convolutional neural networks can take a bit of time (around 24h for each of the ResNet-50 baselines), but I'm sure there will be more and more momentum as time goes by 😉</p>\n<p>I'm sorry but I did not get your first question: what do you mean by \"take notebooks into account\"? (Into account for what exactly?)</p>\n<p>For the second point, it's actually very basic computer vision baselines: finetuning a pre-trained neural network has been used for years now and there is nothing particular additionally done in the baselines provided (moreover, all the hyperparameters used are specified on the leaderboard).<br>\nParticipants familiar with computer vision and CNNs should be easily able to reproduce them using the Pytorch data loader we provide on the GitHub here: <a href=\"https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py\" target=\"_blank\">https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py</a></p>\n<p>I'm not able to share the code for those right now as it is not the cleanest in the world and we don't necessarily want participants to start from it and only make marginal changes: we rather have participants come up with their ideas and try new things 🙂<br>\nMaybe I'll share this code latter on in the competition if it makes sense.<br>\nOr if people really struggle to reproduce the baselines on their own, I could give a little more details if I forgot to mention something important (if it's the case, do not hesitate to ask me if you have a doubt on something and I'll gladly answer 🙂).<br>\nAs of now, one participant, <a href=\"https://www.kaggle.com/juansensio\" target=\"_blank\">@juansensio</a>, managed to reproduce the first CNN baseline, so it looks feasible.</p>",
      "rawMarkdown": "Well the competition only started, it is not very standard data and training of convolutional neural networks can take a bit of time (around 24h for each of the ResNet-50 baselines), but I'm sure there will be more and more momentum as time goes by 😉\n\nI'm sorry but I did not get your first question: what do you mean by \"take notebooks into account\"? (Into account for what exactly?)\n\nFor the second point, it's actually very basic computer vision baselines: finetuning a pre-trained neural network has been used for years now and there is nothing particular additionally done in the baselines provided (moreover, all the hyperparameters used are specified on the leaderboard).\nParticipants familiar with computer vision and CNNs should be easily able to reproduce them using the Pytorch data loader we provide on the GitHub here: https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py\n\nI'm not able to share the code for those right now as it is not the cleanest in the world and we don't necessarily want participants to start from it and only make marginal changes: we rather have participants come up with their ideas and try new things 🙂\nMaybe I'll share this code latter on in the competition if it makes sense.\nOr if people really struggle to reproduce the baselines on their own, I could give a little more details if I forgot to mention something important (if it's the case, do not hesitate to ask me if you have a doubt on something and I'll gladly answer 🙂).\nAs of now, one participant, @juansensio, managed to reproduce the first CNN baseline, so it looks feasible.",
      "votes": null
    },
    {
      "id": "1723766",
      "postDate": "03/15/2022 17:37:42",
      "content": "<p>Hmmm ok. Maybe I was misled from the 'baseline' name and their appearance on the kaggle LB. From what you say I get that the CNN baseline do not run on kaggle notebooks (at least from time constraints). Even the tabular baseline doesn't seems to immediately run on kaggle (from memory constraints). But you are right, not sharing them immediately will help us try different things. </p>\n<p>Regarding the first question: I understand that description of models are published in a paper. Is htere a need for data exploration or just models ?</p>",
      "rawMarkdown": "Hmmm ok. Maybe I was misled from the 'baseline' name and their appearance on the kaggle LB. From what you say I get that the CNN baseline do not run on kaggle notebooks (at least from time constraints). Even the tabular baseline doesn't seems to immediately run on kaggle (from memory constraints). But you are right, not sharing them immediately will help us try different things. \n\nRegarding the first question: I understand that description of models are published in a paper. Is htere a need for data exploration or just models ?",
      "votes": null
    },
    {
      "id": "1724531",
      "postDate": "03/16/2022 09:54:50",
      "content": "<p>Yes, you are right, most of the solutions (baselines included) will not work using Kaggle's notebooks due to memory constraints.<br>\nThis competition is thus harder to enter for participants who do not have access to computing power of their own, hence the smaller involvement of the community you were referring to.<br>\n(Technically, the Random Forest baseline could take less memory but current scikit-learn implementation of Classification Decision Trees uses a dense array (<code>tree_.value</code>) which happens to be huge (of size # nodes x # classes) but very sparse in our case.)</p>\n<p><strong>Participants are encouraged to describe their solutions in working notes paper, whatever their final ranking is!</strong>  🙂<br>\nAdditionally providing data exploration in that paper is also highly valuable for us!<br>\nThe only restricting is that the provided solutions must have tried something new, i.e., not strictly reproducing the provided baselines (which would be of very limited scientific interest…).<br>\nIn particular, we are interested in understanding more precisely the gap between methods using more traditional machine learning methods on environmental vectors or others (illustrated by the baseline of Random Forest) and deep learning methods using image patches (illustrated by the two CNN baselines).<br>\nThus, if, for instance, a team only beats the Random Forest baseline but using other traditional machine learning models on environmental vectors (or other), we will be very interested in knowing their solution, and highly encourage them to submit a working notes paper. 😀</p>",
      "rawMarkdown": "Yes, you are right, most of the solutions (baselines included) will not work using Kaggle's notebooks due to memory constraints.\nThis competition is thus harder to enter for participants who do not have access to computing power of their own, hence the smaller involvement of the community you were referring to.\n(Technically, the Random Forest baseline could take less memory but current scikit-learn implementation of Classification Decision Trees uses a dense array (`tree_.value`) which happens to be huge (of size # nodes x # classes) but very sparse in our case.)\n\n**Participants are encouraged to describe their solutions in working notes paper, whatever their final ranking is!**  🙂\nAdditionally providing data exploration in that paper is also highly valuable for us!\nThe only restricting is that the provided solutions must have tried something new, i.e., not strictly reproducing the provided baselines (which would be of very limited scientific interest...).\nIn particular, we are interested in understanding more precisely the gap between methods using more traditional machine learning methods on environmental vectors or others (illustrated by the baseline of Random Forest) and deep learning methods using image patches (illustrated by the two CNN baselines).\nThus, if, for instance, a team only beats the Random Forest baseline but using other traditional machine learning models on environmental vectors (or other), we will be very interested in knowing their solution, and highly encourage them to submit a working notes paper. 😀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1721347,
      "author_name": "gianetan",
      "author_url": "",
      "post_date": "03/13/2022 15:51:23",
      "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> Thanks for bringing this competition to the notice of Kagglers. Hope it sees more participation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1723683,
      "author_name": "tlorieul",
      "author_url": "",
      "post_date": "03/15/2022 16:10:00",
      "content": "<p>Well the competition only started, it is not very standard data and training of convolutional neural networks can take a bit of time (around 24h for each of the ResNet-50 baselines), but I'm sure there will be more and more momentum as time goes by 😉</p>\n<p>I'm sorry but I did not get your first question: what do you mean by \"take notebooks into account\"? (Into account for what exactly?)</p>\n<p>For the second point, it's actually very basic computer vision baselines: finetuning a pre-trained neural network has been used for years now and there is nothing particular additionally done in the baselines provided (moreover, all the hyperparameters used are specified on the leaderboard).<br>\nParticipants familiar with computer vision and CNNs should be easily able to reproduce them using the Pytorch data loader we provide on the GitHub here: <a href=\"https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py\" target=\"_blank\">https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py</a></p>\n<p>I'm not able to share the code for those right now as it is not the cleanest in the world and we don't necessarily want participants to start from it and only make marginal changes: we rather have participants come up with their ideas and try new things 🙂<br>\nMaybe I'll share this code latter on in the competition if it makes sense.<br>\nOr if people really struggle to reproduce the baselines on their own, I could give a little more details if I forgot to mention something important (if it's the case, do not hesitate to ask me if you have a doubt on something and I'll gladly answer 🙂).<br>\nAs of now, one participant, <a href=\"https://www.kaggle.com/juansensio\" target=\"_blank\">@juansensio</a>, managed to reproduce the first CNN baseline, so it looks feasible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1723766,
          "author_name": "lucasmorin",
          "author_url": "",
          "post_date": "03/15/2022 17:37:42",
          "content": "<p>Hmmm ok. Maybe I was misled from the 'baseline' name and their appearance on the kaggle LB. From what you say I get that the CNN baseline do not run on kaggle notebooks (at least from time constraints). Even the tabular baseline doesn't seems to immediately run on kaggle (from memory constraints). But you are right, not sharing them immediately will help us try different things. </p>\n<p>Regarding the first question: I understand that description of models are published in a paper. Is htere a need for data exploration or just models ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1724531,
          "author_name": "tlorieul",
          "author_url": "",
          "post_date": "03/16/2022 09:54:50",
          "content": "<p>Yes, you are right, most of the solutions (baselines included) will not work using Kaggle's notebooks due to memory constraints.<br>\nThis competition is thus harder to enter for participants who do not have access to computing power of their own, hence the smaller involvement of the community you were referring to.<br>\n(Technically, the Random Forest baseline could take less memory but current scikit-learn implementation of Classification Decision Trees uses a dense array (<code>tree_.value</code>) which happens to be huge (of size # nodes x # classes) but very sparse in our case.)</p>\n<p><strong>Participants are encouraged to describe their solutions in working notes paper, whatever their final ranking is!</strong>  🙂<br>\nAdditionally providing data exploration in that paper is also highly valuable for us!<br>\nThe only restricting is that the provided solutions must have tried something new, i.e., not strictly reproducing the provided baselines (which would be of very limited scientific interest…).<br>\nIn particular, we are interested in understanding more precisely the gap between methods using more traditional machine learning methods on environmental vectors or others (illustrated by the baseline of Random Forest) and deep learning methods using image patches (illustrated by the two CNN baselines).<br>\nThus, if, for instance, a team only beats the Random Forest baseline but using other traditional machine learning models on environmental vectors (or other), we will be very interested in knowing their solution, and highly encourage them to submit a working notes paper. 😀</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1721295": "There doesn't seems to be much involvement from the community in this competition. I was wondering about two things:\n- I don't see much sharing from the community (only one notebook shared in the 2021 edition). Usually kaggle is nice for sharing EDA / insightfull baselines. Is there any plan to take notebooks into account even if they don't reach good scores ? \n- There seems to be some advanced baselines on the LB, including advanced Computer Vision pytorch stuff. Wouldn't it be better to share those ? (maybe share those a month in ?)",
    "1721347": "lucasmorin Thanks for bringing this competition to the notice of Kagglers. Hope it sees more participation.",
    "1723683": "Well the competition only started, it is not very standard data and training of convolutional neural networks can take a bit of time (around 24h for each of the ResNet-50 baselines), but I'm sure there will be more and more momentum as time goes by 😉\n\nI'm sorry but I did not get your first question: what do you mean by \"take notebooks into account\"? (Into account for what exactly?)\n\nFor the second point, it's actually very basic computer vision baselines: finetuning a pre-trained neural network has been used for years now and there is nothing particular additionally done in the baselines provided (moreover, all the hyperparameters used are specified on the leaderboard).\nParticipants familiar with computer vision and CNNs should be easily able to reproduce them using the Pytorch data loader we provide on the GitHub here: https://github.com/maximiliense/GLC/blob/master/data_loading/pytorch_dataset.py\n\nI'm not able to share the code for those right now as it is not the cleanest in the world and we don't necessarily want participants to start from it and only make marginal changes: we rather have participants come up with their ideas and try new things 🙂\nMaybe I'll share this code latter on in the competition if it makes sense.\nOr if people really struggle to reproduce the baselines on their own, I could give a little more details if I forgot to mention something important (if it's the case, do not hesitate to ask me if you have a doubt on something and I'll gladly answer 🙂).\nAs of now, one participant, @juansensio, managed to reproduce the first CNN baseline, so it looks feasible.",
    "1723766": "Hmmm ok. Maybe I was misled from the 'baseline' name and their appearance on the kaggle LB. From what you say I get that the CNN baseline do not run on kaggle notebooks (at least from time constraints). Even the tabular baseline doesn't seems to immediately run on kaggle (from memory constraints). But you are right, not sharing them immediately will help us try different things. \n\nRegarding the first question: I understand that description of models are published in a paper. Is htere a need for data exploration or just models ?",
    "1724531": "Yes, you are right, most of the solutions (baselines included) will not work using Kaggle's notebooks due to memory constraints.\nThis competition is thus harder to enter for participants who do not have access to computing power of their own, hence the smaller involvement of the community you were referring to.\n(Technically, the Random Forest baseline could take less memory but current scikit-learn implementation of Classification Decision Trees uses a dense array (`tree_.value`) which happens to be huge (of size # nodes x # classes) but very sparse in our case.)\n\n**Participants are encouraged to describe their solutions in working notes paper, whatever their final ranking is!**  🙂\nAdditionally providing data exploration in that paper is also highly valuable for us!\nThe only restricting is that the provided solutions must have tried something new, i.e., not strictly reproducing the provided baselines (which would be of very limited scientific interest...).\nIn particular, we are interested in understanding more precisely the gap between methods using more traditional machine learning methods on environmental vectors or others (illustrated by the baseline of Random Forest) and deep learning methods using image patches (illustrated by the two CNN baselines).\nThus, if, for instance, a team only beats the Random Forest baseline but using other traditional machine learning models on environmental vectors (or other), we will be very interested in knowing their solution, and highly encourage them to submit a working notes paper. 😀"
  },
  "source": "meta"
}