{
  "id": 67806,
  "title": "if anybody can clarify",
  "url": "/competitions/PLAsTiCC-2018/discussion/67806",
  "author_name": "",
  "post_date": "2018-10-05T16:18:02.966066Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>how come there are only 7848 rows in the train set metadata ( 1.42 million rows in training set  ). and why test set is so huge in size</p>",
  "messages": [
    {
      "id": "399337",
      "postDate": "10/05/2018 16:18:02",
      "content": "<p>how come there are only 7848 rows in the train set metadata ( 1.42 million rows in training set  ). and why test set is so huge in size</p>",
      "rawMarkdown": "how come there are only 7848 rows in the train set metadata ( 1.42 million rows in training set  ). and why test set is so huge in size",
      "votes": null
    },
    {
      "id": "399458",
      "postDate": "10/05/2018 20:28:15",
      "content": "<p>Hi Philidor,</p>\n\n<p>I would guess, that observation time is expensive. Hence, you use rather smaller but high quality data sets for calibrating your observations and then apply your learnings of the object population statistically on a larger data set.</p>\n\n<p>Astrophysics is a mainly a statistical science.</p>\n\n<p>Cheers,\nJeffrey</p>",
      "rawMarkdown": "Hi Philidor,\n\nI would guess, that observation time is expensive. Hence, you use rather smaller but high quality data sets for calibrating your observations and then apply your learnings of the object population statistically on a larger data set.\n\nAstrophysics is a mainly a statistical science.\n\nCheers,\nJeffrey",
      "votes": null
    },
    {
      "id": "399500",
      "postDate": "10/05/2018 23:07:52",
      "content": "<p>Since it's \"simulated observations\", I doubt that's the reason.  </p>",
      "rawMarkdown": "Since it's \"simulated observations\", I doubt that's the reason.",
      "votes": null
    },
    {
      "id": "399681",
      "postDate": "10/06/2018 13:24:55",
      "content": "<p>Hi Todd,</p>\n\n<p>Let me try to elaborate my answer.</p>\n\n<p>Of course, here at the competition, we deal with completely simulated data.</p>\n\n<p>For real observations, it is not feasible to study each object of your observation in detail.  Spectroscopy is a highly time consuming task, while photometric measurements can be executed within a reasonable time frame. By adopting assumptions about the objects spectrum one can approximate the redshift of the object. With this approach, one unlocks lots of data which qualifies for multiplicities of analytical purposes.</p>\n\n<p>In case, one wants to measure the spectrum of highly distant objects, it becomes even worse, because due to our Earths atmosphere, we have blind spots in the infrared regime, as parts of the incoming radiation gets absorbed. Observations with telescopes at high altitude or with orbital telescopes can help overcome this issue. The Hubble Telescope has been upgraded to its maximum extend for measurements in the infrared, the James Webb Space Telescope will take over.</p>\n\n<p>The COSMOS survey (see below for reference) is a wonderful example for the procedure of calibrating lots of data with few high quality observations. From this survey, we have learned quite a lot about the pro's and con's of this procedure. Other future surveys will adopt the same procedure, e.g., the Euclid survey for cosmological and gravitational lensing studies.</p>\n\n<p>COSMOS Photometric Redshifts with 30-bands for 2-deg2 by O. Ilbert et al. 2008, <a href=\"https://arxiv.org/pdf/0809.2101.pdf\">https://arxiv.org/pdf/0809.2101.pdf</a></p>\n\n<p>I am happy to discuss this topic further. </p>\n\n<p>Cheers, Jeffrey</p>",
      "rawMarkdown": "Hi Todd,\n\nLet me try to elaborate my answer.\n\nOf course, here at the competition, we deal with completely simulated data.\n\nFor real observations, it is not feasible to study each object of your observation in detail.  Spectroscopy is a highly time consuming task, while photometric measurements can be executed within a reasonable time frame. By adopting assumptions about the objects spectrum one can approximate the redshift of the object. With this approach, one unlocks lots of data which qualifies for multiplicities of analytical purposes.\n\nIn case, one wants to measure the spectrum of highly distant objects, it becomes even worse, because due to our Earths atmosphere, we have blind spots in the infrared regime, as parts of the incoming radiation gets absorbed. Observations with telescopes at high altitude or with orbital telescopes can help overcome this issue. The Hubble Telescope has been upgraded to its maximum extend for measurements in the infrared, the James Webb Space Telescope will take over.\n\nThe COSMOS survey (see below for reference) is a wonderful example for the procedure of calibrating lots of data with few high quality observations. From this survey, we have learned quite a lot about the pro's and con's of this procedure. Other future surveys will adopt the same procedure, e.g., the Euclid survey for cosmological and gravitational lensing studies.\n\nCOSMOS Photometric Redshifts with 30-bands for 2-deg2 by O. Ilbert et al. 2008, https://arxiv.org/pdf/0809.2101.pdf\n\nI am happy to discuss this topic further. \n\nCheers, Jeffrey",
      "votes": null
    },
    {
      "id": "409192",
      "postDate": "10/23/2018 23:48:08",
      "content": "<p>The training set meta data has a one to many relationship to the training set data. The metadata does not change for each object, while the training data has many observations over time to track the changing flux of the object. </p>\n\n<p>I’m guessing that test set is huge as this is closer to what is expected in the real world. The quantity of data being run through the model at inference time will be much larger than the quantity of data the model is trained on.</p>",
      "rawMarkdown": "The training set meta data has a one to many relationship to the training set data. The metadata does not change for each object, while the training data has many observations over time to track the changing flux of the object. \n\nI’m guessing that test set is huge as this is closer to what is expected in the real world. The quantity of data being run through the model at inference time will be much larger than the quantity of data the model is trained on.",
      "votes": null
    },
    {
      "id": "409548",
      "postDate": "10/24/2018 13:26:56",
      "content": "<p>Thanks Jack</p>",
      "rawMarkdown": "Thanks Jack",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 399458,
      "author_name": "jeffkk",
      "author_url": "",
      "post_date": "10/05/2018 20:28:15",
      "content": "<p>Hi Philidor,</p>\n\n<p>I would guess, that observation time is expensive. Hence, you use rather smaller but high quality data sets for calibrating your observations and then apply your learnings of the object population statistically on a larger data set.</p>\n\n<p>Astrophysics is a mainly a statistical science.</p>\n\n<p>Cheers,\nJeffrey</p>",
      "votes": null,
      "replies": [
        {
          "id": 399500,
          "author_name": "depmountaineer",
          "author_url": "",
          "post_date": "10/05/2018 23:07:52",
          "content": "<p>Since it's \"simulated observations\", I doubt that's the reason.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 399681,
          "author_name": "jeffkk",
          "author_url": "",
          "post_date": "10/06/2018 13:24:55",
          "content": "<p>Hi Todd,</p>\n\n<p>Let me try to elaborate my answer.</p>\n\n<p>Of course, here at the competition, we deal with completely simulated data.</p>\n\n<p>For real observations, it is not feasible to study each object of your observation in detail.  Spectroscopy is a highly time consuming task, while photometric measurements can be executed within a reasonable time frame. By adopting assumptions about the objects spectrum one can approximate the redshift of the object. With this approach, one unlocks lots of data which qualifies for multiplicities of analytical purposes.</p>\n\n<p>In case, one wants to measure the spectrum of highly distant objects, it becomes even worse, because due to our Earths atmosphere, we have blind spots in the infrared regime, as parts of the incoming radiation gets absorbed. Observations with telescopes at high altitude or with orbital telescopes can help overcome this issue. The Hubble Telescope has been upgraded to its maximum extend for measurements in the infrared, the James Webb Space Telescope will take over.</p>\n\n<p>The COSMOS survey (see below for reference) is a wonderful example for the procedure of calibrating lots of data with few high quality observations. From this survey, we have learned quite a lot about the pro's and con's of this procedure. Other future surveys will adopt the same procedure, e.g., the Euclid survey for cosmological and gravitational lensing studies.</p>\n\n<p>COSMOS Photometric Redshifts with 30-bands for 2-deg2 by O. Ilbert et al. 2008, <a href=\"https://arxiv.org/pdf/0809.2101.pdf\">https://arxiv.org/pdf/0809.2101.pdf</a></p>\n\n<p>I am happy to discuss this topic further. </p>\n\n<p>Cheers, Jeffrey</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409192,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "10/23/2018 23:48:08",
      "content": "<p>The training set meta data has a one to many relationship to the training set data. The metadata does not change for each object, while the training data has many observations over time to track the changing flux of the object. </p>\n\n<p>I’m guessing that test set is huge as this is closer to what is expected in the real world. The quantity of data being run through the model at inference time will be much larger than the quantity of data the model is trained on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 409548,
          "author_name": "mks2192",
          "author_url": "",
          "post_date": "10/24/2018 13:26:56",
          "content": "<p>Thanks Jack</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "399337": "how come there are only 7848 rows in the train set metadata ( 1.42 million rows in training set  ). and why test set is so huge in size",
    "399458": "Hi Philidor,\n\nI would guess, that observation time is expensive. Hence, you use rather smaller but high quality data sets for calibrating your observations and then apply your learnings of the object population statistically on a larger data set.\n\nAstrophysics is a mainly a statistical science.\n\nCheers,\nJeffrey",
    "399500": "Since it's \"simulated observations\", I doubt that's the reason.",
    "399681": "Hi Todd,\n\nLet me try to elaborate my answer.\n\nOf course, here at the competition, we deal with completely simulated data.\n\nFor real observations, it is not feasible to study each object of your observation in detail.  Spectroscopy is a highly time consuming task, while photometric measurements can be executed within a reasonable time frame. By adopting assumptions about the objects spectrum one can approximate the redshift of the object. With this approach, one unlocks lots of data which qualifies for multiplicities of analytical purposes.\n\nIn case, one wants to measure the spectrum of highly distant objects, it becomes even worse, because due to our Earths atmosphere, we have blind spots in the infrared regime, as parts of the incoming radiation gets absorbed. Observations with telescopes at high altitude or with orbital telescopes can help overcome this issue. The Hubble Telescope has been upgraded to its maximum extend for measurements in the infrared, the James Webb Space Telescope will take over.\n\nThe COSMOS survey (see below for reference) is a wonderful example for the procedure of calibrating lots of data with few high quality observations. From this survey, we have learned quite a lot about the pro's and con's of this procedure. Other future surveys will adopt the same procedure, e.g., the Euclid survey for cosmological and gravitational lensing studies.\n\nCOSMOS Photometric Redshifts with 30-bands for 2-deg2 by O. Ilbert et al. 2008, https://arxiv.org/pdf/0809.2101.pdf\n\nI am happy to discuss this topic further. \n\nCheers, Jeffrey",
    "409192": "The training set meta data has a one to many relationship to the training set data. The metadata does not change for each object, while the training data has many observations over time to track the changing flux of the object. \n\nI’m guessing that test set is huge as this is closer to what is expected in the real world. The quantity of data being run through the model at inference time will be much larger than the quantity of data the model is trained on.",
    "409548": "Thanks Jack"
  },
  "source": "meta"
}