{
  "id": 314831,
  "title": "Is Private Public Split randomly ?",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/314831",
  "author_name": "",
  "post_date": "2022-03-24T16:26:14.211063500Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Is Private Public Split randomly  or is the split done intetionally different ?</p>",
  "messages": [
    {
      "id": "1733822",
      "postDate": "03/24/2022 16:26:14",
      "content": "<p>Is Private Public Split randomly  or is the split done intetionally different ?</p>",
      "rawMarkdown": "Is Private Public Split randomly  or is the split done intetionally different ?",
      "votes": null
    },
    {
      "id": "1733841",
      "postDate": "03/24/2022 16:44:41",
      "content": "<p>You can read more about the dataset specifics in the paper presented about the dataset at FGVC last year -- <a href=\"https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view\" target=\"_blank\">https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view</a>. Specifically, this is the relevant bit on the dataset split:</p>\n<blockquote>\n  <p>Each cultivar was grown in two separate plots in the TERRA-REF field … to account for extremely local field or soil conditions that might impact the growth of plants in one particular plot. We leverage this natural split in the data when dividing our dataset between train and test – images for a given cultivar in the training dataset come from one plot, while the test images from that same cultivar come from the other plot. This means that a model cannot achieve high performance by memorizing features that aren’t meaningful phenotypes (e.g., by memorizing patterns observed in the dirt).</p>\n</blockquote>",
      "rawMarkdown": "You can read more about the dataset specifics in the paper presented about the dataset at FGVC last year -- https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view. Specifically, this is the relevant bit on the dataset split:\n\n> Each cultivar was grown in two separate plots in the TERRA-REF field ... to account for extremely local field or soil conditions that might impact the growth of plants in one particular plot. We leverage this natural split in the data when dividing our dataset between train and test – images for a given cultivar in the training dataset come from one plot, while the test images from that same cultivar come from the other plot. This means that a model cannot achieve high performance by memorizing features that aren’t meaningful phenotypes (e.g., by memorizing patterns observed in the dirt).",
      "votes": null
    },
    {
      "id": "1733844",
      "postDate": "03/24/2022 16:46:45",
      "content": "<p>I know the difference between the train and test split ; to be more specific i am asking about test public and test private lb split </p>",
      "rawMarkdown": "I know the difference between the train and test split ; to be more specific i am asking about test public and test private lb split",
      "votes": null
    },
    {
      "id": "1733852",
      "postDate": "03/24/2022 16:51:15",
      "content": "<p>Our standard practice is to not disclose how the public/private split was defined.</p>",
      "rawMarkdown": "Our standard practice is to not disclose how the public/private split was defined.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1733841,
      "author_name": "abbystylianou",
      "author_url": "",
      "post_date": "03/24/2022 16:44:41",
      "content": "<p>You can read more about the dataset specifics in the paper presented about the dataset at FGVC last year -- <a href=\"https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view\" target=\"_blank\">https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view</a>. Specifically, this is the relevant bit on the dataset split:</p>\n<blockquote>\n  <p>Each cultivar was grown in two separate plots in the TERRA-REF field … to account for extremely local field or soil conditions that might impact the growth of plants in one particular plot. We leverage this natural split in the data when dividing our dataset between train and test – images for a given cultivar in the training dataset come from one plot, while the test images from that same cultivar come from the other plot. This means that a model cannot achieve high performance by memorizing features that aren’t meaningful phenotypes (e.g., by memorizing patterns observed in the dirt).</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1733844,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "03/24/2022 16:46:45",
          "content": "<p>I know the difference between the train and test split ; to be more specific i am asking about test public and test private lb split </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1733852,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "03/24/2022 16:51:15",
          "content": "<p>Our standard practice is to not disclose how the public/private split was defined.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1733822": "Is Private Public Split randomly  or is the split done intetionally different ?",
    "1733841": "You can read more about the dataset specifics in the paper presented about the dataset at FGVC last year -- https://drive.google.com/file/d/18A1bibKMsuqUK7E2ZMYPZiNu3kT7JmgC/view. Specifically, this is the relevant bit on the dataset split:\n\n> Each cultivar was grown in two separate plots in the TERRA-REF field ... to account for extremely local field or soil conditions that might impact the growth of plants in one particular plot. We leverage this natural split in the data when dividing our dataset between train and test – images for a given cultivar in the training dataset come from one plot, while the test images from that same cultivar come from the other plot. This means that a model cannot achieve high performance by memorizing features that aren’t meaningful phenotypes (e.g., by memorizing patterns observed in the dirt).",
    "1733844": "I know the difference between the train and test split ; to be more specific i am asking about test public and test private lb split",
    "1733852": "Our standard practice is to not disclose how the public/private split was defined."
  },
  "source": "meta"
}